The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.
All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.
So the shape now is one entry point and one screen:
Campaigns -> Campaign -> Story
State · Knowledge · Context · Save Points · Settings
Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.
Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.
Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.
The two defects worth the space:
A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.
And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.
Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.
`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.
Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.
Backend, and only what the browser could not otherwise reach:
AdventureCreate.opening a start action could only come from a Scenario, so
every campaign made in the new setup flow opened on
a blank page. Same node, same code path.
canon_rules campaign_canon has been the highest authority in a
campaign since M5, read by the prompt builder and
the state validator, and had no API at all — a
fixture had to write it with SQL.
a 401 and a 429 message the last user-facing text describing a hosted
deployment. One told the reader to check an API key
that has not existed since M2.
No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.
The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.
They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.
A verification pass over all of it then found three more, each by driving the
product rather than reading it:
Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.
The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.
And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.
Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.
`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.
Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:
The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.
Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.
The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.
The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.
Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
327 lines
20 KiB
Markdown
327 lines
20 KiB
Markdown
# Adventure Storyteller
|
||
|
||
[](LICENSE)
|
||
|
||
An interactive storytelling app that runs entirely on your own machine, with your own model.
|
||
Start a campaign from a short form and play an open-ended story where a local LLM narrates the
|
||
world, keeps track of what is true, and remembers what happened. Since M8 the browser has one
|
||
entry point — a campaign library — and one natural-language input; the scenario gallery and its
|
||
editor are gone from the interface, though the AI Dungeon-compatible scenario *format* is still
|
||
supported for import and export.
|
||
|
||
This is the **Adventure Storyteller** fork of [AI-DnD](https://github.com/parththakkar106/AI-DnD).
|
||
It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model
|
||
you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support,
|
||
the Postgres path, and the JavaScript scripting engine have all been removed rather than
|
||
disabled. What is left is a storyteller you can run offline.
|
||
|
||
> **Local-only, by design.** The app talks to one place — an Ollama-compatible endpoint on this
|
||
> machine or on a machine you control on your own network — and it refuses to be pointed at a
|
||
> public address. There is no telemetry, no account, no cloud inference, and nothing is fetched
|
||
> at runtime from the Internet.
|
||
>
|
||
> For the internals, read [`planning/TECHNICAL-DESIGN.md`](planning/TECHNICAL-DESIGN.md) and
|
||
> [`planning/CONTEXT-AND-MEMORY.md`](planning/CONTEXT-AND-MEMORY.md), which cover the context
|
||
> budgeting, the state model and the memory bank as this fork builds them.
|
||
|
||
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing
|
||
everything in one SQLite file.
|
||
|
||
On the play screen, the left rail carries live world state. The AI proposes changes each turn
|
||
and a Python engine decides what actually sticks; the chip under the narration reports what
|
||
changed. The `‹ 2/2 ›` under a turn steps between the takes it has, and writing below a take
|
||
that isn't the live one starts a new branch.
|
||
|
||
## Features
|
||
|
||
- **The full play loop, in one box.** You write what you do or say in a single
|
||
natural-language field — an action and a piece of quoted dialogue are both just what you
|
||
wrote — with **Continue** for a beat you do not act in and a **Story direction** toggle for
|
||
speaking to the narrator rather than in the story. Responses stream (SSE), and Undo, Redo,
|
||
Retry and Edit sit beside the box. Correcting narrator prose does not overwrite it: the
|
||
correction becomes a new continuation carrying the state it implies, and the original
|
||
narration keeps its own future as retained history. Reasoning models are supported: the
|
||
narrator's thinking streams into a collapsible panel with its own token budget.
|
||
- **Retained history, without a tree to manage.** Underneath, the story is a tree: any turn
|
||
can hold more than one **take**, and `‹ 2/4 ›` steps between them. Stepping is free — the
|
||
story below simply empties, and the server is told nothing. Writing below a take that is not
|
||
the live one is what starts a different continuation. Branches borrow their ancestors' turns
|
||
instead of copying them, so one costs about 100 bytes, and a 20-fork story loads within 1% of
|
||
the same story flat (`backend/app/tree.py`, `backend/app/context/lineage.py`).
|
||
|
||
**None of that vocabulary reaches the reader.** M8 removed the branch panel and the tree
|
||
overlay from the browser: what you get is Undo, Redo, Retry, takes, Save Points and Restore,
|
||
and nothing on screen says branch, fork, node or head. The mechanism is unchanged and still
|
||
fully tested — this is a decision about what you are asked to understand, not about what the
|
||
product can do.
|
||
- **Authoritative narrative state, and the application owns it.** The story tracks who exists,
|
||
where they are, what they hold, what is true, how they are tied to each other, and what is
|
||
still open — as generic entities, facts, relationships and threads, with no genre baked in.
|
||
The same schema holds a silver key in an abbey and a data crystal on an orbital station.
|
||
The AI proposes **typed events with absolute values** (`set_possession`, `add_fact`,
|
||
`set_current_location` …), and a Python validator decides what is accepted: unknown event
|
||
types are refused, references must resolve, campaign canon outranks the narration, and the
|
||
machine-readable block never reaches the reader (`backend/app/narrative/`). Every accepted
|
||
change is recorded with what it was before and which turn caused it, so the Story State panel
|
||
can show what changed and why. You can correct it by hand, and your correction outranks the
|
||
story.
|
||
- **AI Dungeon-compatible context engine.** Memory, author's note, and story cards (world
|
||
info) are triggered by keywords in recent story text, then assembled under a token budget
|
||
(`backend/app/context/builder.py`). Story cards are the inherited authored-lore primitive and
|
||
are kept; they are **not** the knowledge library below, which is a first-class subsystem with
|
||
its own classification, provenance, chunking and index.
|
||
- **Total prompt transparency.** Every turn stores the exact prompt sent to the model.
|
||
**Inspect context** on any narrator turn opens a readable account of what it was given —
|
||
what it remembered, what it read, what it believes, and what each part cost — with the
|
||
assembled prompt itself kept as an advanced section rather than opening on a wall of text.
|
||
A passage that came from an imported file links back to the file it came from.
|
||
- **Auto-summarization and Memory Bank.** The modern AI Dungeon memory system: AI-generated
|
||
memories every few actions, a running story summary, and embedding-based retrieval that
|
||
pulls old-but-relevant facts back into context, with similarity scores visible in the
|
||
context inspector
|
||
(`backend/app/memorybank.py`).
|
||
- **An imported knowledge library, classified by how much authority it has.** Import your own
|
||
local `.txt` and `.md` files — a setting bible, character notes, research, a passage whose
|
||
voice you want the prose to have — as **Canon**, **Reference** or **Inspiration**. The class
|
||
is not a label: it decides the words the passage is framed with in the prompt, the weight it
|
||
carries when passages are ranked, and which budget it competes in when the context is tight.
|
||
Canon can establish what is true; Reference informs detail without establishing anything;
|
||
Inspiration influences tone and introduces no facts at all. Retrieval is **hybrid and local**:
|
||
a SQLite FTS5 index finds the names and invented terms an embedding is worst at, local Ollama
|
||
embeddings find what you meant when your words differ from the file's, and the two are merged,
|
||
de-duplicated and reranked by relevance × class. Lexical search is a supported production
|
||
path, not a fallback — the library works with no embedding model at all. Canon you mark
|
||
**always include** is supplied on every turn whether or not the scene resembles it, and Canon
|
||
you mark **narrator only** is given to the narrator with instructions not to let the
|
||
protagonist know it. Every passage that reaches a prompt is listed in the context inspector with its file,
|
||
class, heading, passage number, scores and token cost, and that record is kept in the turn, so
|
||
deleting a source never erases the evidence of what an old turn was shown
|
||
(`backend/app/knowledge/`).
|
||
- **Imported text is data, never instruction.** Every imported passage is delimited in the
|
||
prompt as untrusted data with the authority order stated in words, so "ignore all previous
|
||
instructions" inside a file is a sentence in a file. Nothing is fetched: a URL in a source is
|
||
text, a remote Markdown image never loads, and no endpoint anywhere takes a filesystem path —
|
||
a source arrives as an upload, so there is no path for a traversal to escape from. Imported
|
||
content is displayed as inert text and never rendered as HTML.
|
||
- **Undo, Redo, and retry that roll back state and delete nothing.** Undo moves where the story
|
||
is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped
|
||
over. Both restore the world state from a per-node snapshot rather than just the text, and a
|
||
memory derived from a turn now behind the head stops being retrieved without being deleted or
|
||
re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the
|
||
displaced future stays on the line it was written for, and ordinary Redo stops offering it.
|
||
Nothing a retry replaces is discarded either — the old attempt stays as another take of that
|
||
turn, one keystroke and one click from becoming a branch of its own.
|
||
- **Save Points.** Name a moment — "Before entering the abbey" — keep playing,
|
||
restart the app, and come back to it. Restoring one moves the story back to
|
||
that moment and deletes nothing: the turns you wrote after it stay, Redo still
|
||
walks forward into them, and writing something different from the Save Point
|
||
is what starts a new line while the old one is kept. A Save Point is a name for
|
||
a position and holds no copy of the story, so restoring it is the same
|
||
movement Undo makes (`backend/app/routers/adventures/checkpoints.py`,
|
||
`backend/app/head.py`). They last until *you* delete them: deleting one deletes
|
||
no story, and deleting a branch a Save Point is kept on is refused until you
|
||
remove the Save Point yourself, so nothing takes a named moment away behind
|
||
your back.
|
||
- **Import and export.** AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree:
|
||
every branch, every take, the fork points, which branches the story has left behind, the Save
|
||
Points and the position it is being read at — all of them chosen rather than computed, which is
|
||
the rule for what a bundle carries. A campaign opens where its head says, never at a Save Point
|
||
merely because it has one. A campaign exported after two Undos imports still undone, with its
|
||
retained future intact, instead of silently reopening at its newest turn. Files that predate
|
||
the head position, and files saved in the old single-line format, still import.
|
||
- **Single user, no accounts.** There is no sign-up, no login, no session and no API key
|
||
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
|
||
design, because the only person who can reach it is the person running it. A new install
|
||
starts with a short pre-played adventure, so the first screen shows real turns and their
|
||
world-state changes rather than an empty page (`backend/app/starter.py`).
|
||
- **A refusal you can rely on.** The inference endpoint is checked against an address
|
||
allowlist when you save it and again before every request, so a public endpoint is refused
|
||
even if the setting is edited in the database directly. TLS verification is never traded
|
||
against reachability: a privately issued certificate is verified against your machine's own
|
||
trust store, and there is no bypass switch.
|
||
|
||
## Screenshots
|
||
|
||
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
|
||
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
|
||
removed rather than left standing as a picture of a product that no longer exists. The M4
|
||
closeout drove the real application in a real browser, so the screens exist and work; taking
|
||
presentable screenshots of them is a job for the UI pass in M8.
|
||
|
||
## Quick start
|
||
|
||
### Docker (any OS)
|
||
|
||
```sh
|
||
docker compose up --build
|
||
```
|
||
|
||
Open http://localhost:8000. Your data persists in a named volume across restarts.
|
||
|
||
### Windows
|
||
|
||
```powershell
|
||
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
|
||
cd frontend; npm install; cd ..
|
||
.\start.ps1
|
||
```
|
||
|
||
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
|
||
|
||
### macOS / Linux
|
||
|
||
```sh
|
||
./start.sh # creates the venv and installs dependencies on first run
|
||
```
|
||
|
||
Open http://localhost:5173.
|
||
|
||
## Connect a model
|
||
|
||
Ollama is the inference backend v1 supports. Open **Settings** in the app and point it at one:
|
||
|
||
| Where Ollama runs | Endpoint URL | Notes |
|
||
|---|---|---|
|
||
| The same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works |
|
||
| A machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | explicitly configured; see below |
|
||
|
||
Model name, generation parameters, and (optionally) summary and embedding models for the
|
||
Memory Bank are configured there too. No config files and no rebuild are needed. There is no
|
||
API key field, because there is nothing to authenticate to.
|
||
|
||
The adapter underneath speaks the OpenAI-compatible protocol, because that is what Ollama
|
||
serves. That is an implementation detail, not a promise of support for arbitrary local
|
||
servers that happen to speak the same protocol. Public and cloud inference endpoints are
|
||
prohibited outright — see `planning/DECISIONS/002-ollama-only-v1.md` and
|
||
`planning/DECISIONS/011-local-inference-endpoint-policy.md`.
|
||
|
||
### What the endpoint policy allows
|
||
|
||
The address is checked when you save it and again before every request. Only loopback and
|
||
private-network addresses are accepted; every public address is refused, by address rather than
|
||
by hostname, so a name that resolves outward is refused too. A well-known cloud inference host
|
||
is named in the error message only so the refusal says *why*.
|
||
|
||
Running the model on a second machine you control is supported and expected — that machine
|
||
does the inference while the storyteller itself stays bound to loopback on yours. If that
|
||
machine serves HTTPS with a certificate from a CA you installed, it works: certificates are
|
||
verified against your operating system's trust store as well as the bundled one. Verification
|
||
itself is never relaxed, and there is no option to turn it off.
|
||
|
||
### Playing against a local shim (development only)
|
||
|
||
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787`
|
||
backed by a command-line tool, which is useful for testing the turn engine against a stronger
|
||
model. Each request spawns one process, which suits the engine: the app assembles the whole
|
||
prompt every turn and expects a stateless endpoint.
|
||
|
||
```sh
|
||
cd backend
|
||
.venv/bin/python tools/claude_shim.py # listens on 127.0.0.1:8787
|
||
```
|
||
|
||
Set the base URL to `http://127.0.0.1:8787/v1` and pick a model the tool offers. Set the
|
||
reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field that
|
||
some models reject. Embeddings are not served — leave the embedding model blank, or point the
|
||
Memory Bank at an endpoint that serves one.
|
||
|
||
The shim has no authentication and spends whatever quota backs it, so run it on loopback and
|
||
leave it there.
|
||
|
||
## How a turn works
|
||
|
||
```
|
||
player input
|
||
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
|
||
+ [plot essentials] + [story summary] + [retrieved memories]
|
||
+ [triggered story cards] + [retrieved imported knowledge,
|
||
framed by class and bounded by its own budget]
|
||
+ [history along this branch, token-budgeted]
|
||
+ [author's note] + [player action]
|
||
→ snapshot context (Insights)
|
||
→ provider adapter → AI (streamed)
|
||
→ extract + referee the world-state delta block, strip it from the prose
|
||
→ store & render
|
||
```
|
||
|
||
## Architecture
|
||
|
||
```
|
||
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||
├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug
|
||
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource
|
||
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting)
|
||
├─ endpoints.py the inference-endpoint address policy
|
||
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
|
||
├─ tree.py forking, promotion, and where a node is placed
|
||
├─ head.py the active head: where the story is read, and what moving it costs
|
||
├─ checkpoints Save Points: durable names for positions, in routers/adventures/
|
||
├─ attempts.py the takes of one turn, grouped by parent
|
||
├─ context/ prompt assembly under a token budget + lineage/history windowing
|
||
├─ narrative/ the authoritative state: typed events, validation, snapshots
|
||
├─ worldstate/ the inherited RPG stat engine — legacy, no longer authoritative
|
||
├─ memorybank.py auto-summarization + embedding retrieval
|
||
├─ knowledge/ the imported library: import, chunk, FTS5, embed, rank, inject
|
||
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
|
||
├─ providers/ OpenAI-compatible adapter, streaming
|
||
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
|
||
```
|
||
|
||
In production the backend serves the built SPA from one port (see `Dockerfile`). In
|
||
development, Vite proxies `/api` to FastAPI.
|
||
|
||
## Tests
|
||
|
||
920 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||
the model provider mocked. They run with no route to the Internet, which is a requirement
|
||
rather than a convenience — an offline claim proved on a machine that has been online once
|
||
proves nothing. A further handful need a real local model and skip without one; they exist
|
||
because a mocked provider can leave the production wiring dead while the suite stays green,
|
||
which this project has shipped twice.
|
||
|
||
```sh
|
||
cd backend && pip install -r requirements.txt -r requirements-dev.txt
|
||
python -m pytest tests/
|
||
```
|
||
|
||
## Notes on performance
|
||
|
||
Two of the tests exist because of bugs that were measured rather than guessed at. They are the
|
||
most interesting engineering in the repo.
|
||
|
||
- **Database egress, cut about 189x.** Every adventure load pulled `Action.context_snapshot`,
|
||
the entire assembled prompt at about 74 KB per turn, to read two small fields off it. Moving
|
||
those fields into their own columns and marking the heavy ones `deferred` took one adventure
|
||
load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL, so the
|
||
old data never crosses the wire. `tests/test_egress.py` hooks into SQLAlchemy's cursor events
|
||
and fails if a bulk load ever names those columns again.
|
||
- **Turn cost, made flat.** Assembling a turn walked the whole story, so cost scaled with story
|
||
length: 839 KB of reads at turn 200. `backend/app/context/history.py` now serves tails and
|
||
slices from SQL and measures what it fetched. The same turn now costs 129 KB and stops
|
||
growing at around turn 50.
|
||
- **Branching that costs nothing to read.** A branch stores no turns. It stores where it left
|
||
its parent, and borrows everything above that. A 40-turn story forked twenty times loads in
|
||
31,652 bytes against 31,433 bytes for the same story flat: a 1.007x ratio, or about 103 bytes
|
||
per branch. Reads stay cheap because the lineage is windowed the same way the history is, so
|
||
the number of SQL clauses is bounded by the context window rather than by the number of
|
||
forks.
|
||
|
||
## Repo notes
|
||
|
||
- `planning/` is this fork's own package: the product specification, the architecture
|
||
decisions, the milestone plan, the acceptance contract, and a review report for every
|
||
milestone shipped. Start at [`planning/README.md`](planning/README.md).
|
||
- [`planning/archive/`](planning/archive/README.md) holds the Phase 0 research that chose this
|
||
base and the completed milestone reports. It is history, not instruction.
|
||
- [`DEVELOPMENT.md`](DEVELOPMENT.md) is how to set the project up, point it at a model, and run
|
||
the tests. [`PROVENANCE.md`](PROVENANCE.md) records what came from upstream and what changed.
|
||
- `backend/.env.example` lists the two environment variables the backend reads. Everything
|
||
about the model is a runtime setting on the Settings page instead.
|
||
- Upstream's own `plan/` build log and `docs/` project site were removed in the 2026-09-03
|
||
documentation pass: they described AI-DnD's hosted, scripted, multi-user product. Both are
|
||
still in Git history, and in upstream.
|
||
|
||
## License
|
||
|
||
[MIT](LICENSE)
|