Files
interactive-story/README.md
T
JesseMarkowitzandClaude Opus 5 144406cd48 M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was
whether any of the earlier evidence meant what it said. M8 measured a deployment
enforcing a 4,096-token input window while the application budgeted 16,384.
Every request returned 200. What Ollama does with the excess is drop the oldest
tokens, and the oldest tokens here are the system block — the narrator's rules
and the campaign canon. A hundred-turn certification against that server would
have looked perfect and proved nothing, which is why this milestone could not
begin with a hundred turns.

So the application asks now. Ollama's window is a property of how a model was
loaded rather than of the request — sending num_ctx is accepted, ignored, and
worse, reloads the model at the server's own default — so the only honest move
is to find out and then tell the truth about it. /api/ps reports what a resident
model is being served with, /api/show what an unloaded one will load with, both
on the same host inference already uses, through the same endpoint policy and
the same TLS trust store. A verified window is a ceiling on the budget; an
unverified one leaves the budget alone and is recorded as unverified in the
turn's own provenance, so an old turn can be asked afterwards whether it was
built against a checked window. There is no third behaviour, and in particular
no hard-coded 4,096: a number the server did not say would be right on one
machine and wrong on the next.

The proof that this is doing something is a campaign whose canon sits at the
front of the prompt, 120 turns of history, and a 4,096-token window. The canon
is still there afterwards and the oldest history is gone. The same campaign
built the old way produces a prompt more than twice the window — the defect,
reproduced, so the fix is measured against it rather than asserted.

Two defects the validation found on its own, and they are the same defect twice:
something was true and nobody was told. A manual state correction of four
changes with one bad reference applied three, returned 201, and said nothing —
while recording the refusal on the audit row nobody reads. It came to light
because the identity diagnostic's own fixture was refused that way and the whole
run proceeded on a campaign with no scene, which would have read as a model
failure. And the narration-length setting moved no number: brief, medium and
long each became one English sentence, while the numeric hint the model actually
reads was derived from the global reply cap and said the same thing for all
three. Both now say what they did.

The other two post-M8 findings are closed as well. The tab said AI D&D, which no
document had ever claimed it did not; it says Interactive Story now, with the
open campaign first, and the name is the owner's decision rather than a
find-and-replace to something narrower than the engine. After an Undo the reader
could not tell where they had landed; the control row now ends with
"Moment 11 · later story ahead", from the server's own answer, in the word the
transcript already uses, with none of head, branch or depth anywhere near it.

The identity diagnostic exists and the root cause does not. That campaign was
destroyed, so no cause can be established — what M11 owes the finding is
something that can classify the next occurrence, and a diagnostic that makes only
the judgements a program can honestly make: duplicate keys, shared names,
protagonist drift, state and context disagreeing. Whether prose misattributed a
line is left to a person reading it beside its prompt, because a regex cannot
read dialogue and one that pretended to would produce exactly the confident wrong
answer this finding is about. Its detectors are proved to fire against a planted
second Alice.

Two entities may still share a display name. That was checked first, as the
finding asked, and left permitted: a mother and a daughter, or a stranger giving
a false name, are ordinary fiction, and refusing them to guard against a model
mistake would refuse the wrong thing. What was missing was that it happened
silently. It is reported now.

Evidence, not inference: a hundred accepted turns against a real narrator with
genuine process restarts; a real browser against the built SPA; a container with
no network at all; a campaign moved into a data directory that never existed.
Each was discarded and re-run whenever the product changed under it, and the runs
that were thrown away are listed in the report with the reason, along with ten
defects in the harnesses themselves — because a harness that has only ever
agreed with itself is not evidence, and two of M8's five harness defects were
masking real ones.

No dependency was added, removed or upgraded. No acceptance test was retired,
relaxed or reclassified. M11 is implemented and verified; it is not accepted, and
there is no release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 14:01:20 -04:00

23 KiB
Raw Blame History

Adventure Storyteller

License: MIT

An interactive storytelling app that runs entirely on your own machine, with your own model. Start a campaign from a short form and play an open-ended story where a local LLM narrates the world, keeps track of what is true, and remembers what happened. Since M8 the browser has one entry point — a campaign library — and one natural-language input; the scenario gallery and its editor are gone from the interface, though the AI Dungeon-compatible scenario format is still supported for import and export.

This is the Adventure Storyteller fork of AI-DnD. It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support, the Postgres path, and the JavaScript scripting engine have all been removed rather than disabled. What is left is a storyteller you can run offline.

Local-only, by design. The app talks to one place — an Ollama-compatible endpoint on this machine or on a machine you control on your own network — and it refuses to be pointed at a public address. There is no telemetry, no account, no cloud inference, and nothing is fetched at runtime from the Internet.

For the internals, read planning/TECHNICAL-DESIGN.md and planning/CONTEXT-AND-MEMORY.md, which cover the context budgeting, the state model and the memory bank as this fork builds them.

Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing everything in one SQLite file.

On the play screen, the left rail carries live world state. The AI proposes changes each turn and a Python engine decides what actually sticks; the chip under the narration reports what changed. The ‹ 2/2 › under a turn steps between the takes it has, and writing below a take that isn't the live one starts a new branch.

Features

  • The full play loop, in one box. You write what you do or say in a single natural-language field — an action and a piece of quoted dialogue are both just what you wrote — with Continue for a beat you do not act in and a Story direction toggle for speaking to the narrator rather than in the story. Responses stream (SSE), and Undo, Redo, Retry and Edit sit beside the box. Correcting narrator prose does not overwrite it: the correction becomes a new continuation carrying the state it implies, and the original narration keeps its own future as retained history. Reasoning models are supported: the narrator's thinking streams into a collapsible panel with its own token budget.

  • Retained history, without a tree to manage. Underneath, the story is a tree: any turn can hold more than one take, and ‹ 2/4 › steps between them. Stepping is free — the story below simply empties, and the server is told nothing. Writing below a take that is not the live one is what starts a different continuation. Branches borrow their ancestors' turns instead of copying them, so one costs about 100 bytes, and a 20-fork story loads within 1% of the same story flat (backend/app/tree.py, backend/app/context/lineage.py).

    None of that vocabulary reaches the reader. M8 removed the branch panel and the tree overlay from the browser: what you get is Undo, Redo, Retry, takes, Save Points and Restore, and nothing on screen says branch, fork, node or head. The mechanism is unchanged and still fully tested — this is a decision about what you are asked to understand, not about what the product can do.

  • Authoritative narrative state, and the application owns it. The story tracks who exists, where they are, what they hold, what is true, how they are tied to each other, and what is still open — as generic entities, facts, relationships and threads, with no genre baked in. The same schema holds a silver key in an abbey and a data crystal on an orbital station. The AI proposes typed events with absolute values (set_possession, add_fact, set_current_location …), and a Python validator decides what is accepted: unknown event types are refused, references must resolve, campaign canon outranks the narration, and the machine-readable block never reaches the reader (backend/app/narrative/). Every accepted change is recorded with what it was before and which turn caused it, so the Story State panel can show what changed and why. You can correct it by hand, and your correction outranks the story.

  • A context engine you can account for. Memory, the author's note, the campaign's own rules, the authoritative state, the summary that applies here, and the retrieved imported passages are assembled under one token budget, in an order chosen so that a section which changes does not re-price the cached prefix above it (backend/app/context/builder.py).

    And the budget is the one your server will actually read. Ollama enforces a context window of its own — 4,096 by default on a machine with no VRAM — and a larger prompt is not refused, it is silently trimmed from the oldest end, which here is the narrator's rules and your campaign's canon. The application asks the server what window your model gets and caps the prompt to it, so what a small window costs is history rather than the canon at the front (backend/app/contextwindow.py). If it cannot check, it says so instead of assuming.

    Story cards — AI Dungeon's world-info primitive, inherited with the fork — are kept as legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched card used to arrive in front of it as a world fact with no class, no visibility, no source and nothing to switch it off, competing with your imported Canon for the same budget; the knowledge library below replaces it, and does all of that explicitly.

  • Total prompt transparency. Every turn stores the exact prompt sent to the model. Inspect context on any narrator turn opens a readable account of what it was given — what it remembered, what it read, what it believes, and what each part cost — with the assembled prompt itself kept as an advanced section rather than opening on a wall of text. A passage that came from an imported file links back to the file it came from.

  • Auto-summarization and Memory Bank. The modern AI Dungeon memory system: AI-generated memories every few actions, a running story summary, and embedding-based retrieval that pulls old-but-relevant facts back into context, with similarity scores visible in the context inspector (backend/app/memorybank.py).

  • An imported knowledge library, classified by how much authority it has. Import your own local .txt and .md files — a setting bible, character notes, research, a passage whose voice you want the prose to have — as Canon, Reference or Inspiration. The class is not a label: it decides the words the passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. Canon can establish what is true; Reference informs detail without establishing anything; Inspiration influences tone and introduces no facts at all. Retrieval is hybrid and local: a SQLite FTS5 index finds the names and invented terms an embedding is worst at, local Ollama embeddings find what you meant when your words differ from the file's, and the two are merged, de-duplicated and reranked by relevance × class. Lexical search is a supported production path, not a fallback — the library works with no embedding model at all. Canon you mark always include is supplied on every turn whether or not the scene resembles it, and Canon you mark narrator only is given to the narrator with instructions not to let the protagonist know it. Every passage that reaches a prompt is listed in the context inspector with its file, class, heading, passage number, scores and token cost, and that record is kept in the turn, so deleting a source never erases the evidence of what an old turn was shown (backend/app/knowledge/).

  • Imported text is data, never instruction. Every imported passage is delimited in the prompt as untrusted data with the authority order stated in words, so "ignore all previous instructions" inside a file is a sentence in a file. Nothing is fetched: a URL in a source is text, a remote Markdown image never loads, and no endpoint anywhere takes a filesystem path — a source arrives as an upload, so there is no path for a traversal to escape from. Imported content is displayed as inert text and never rendered as HTML.

  • Undo, Redo, and retry that roll back state and delete nothing. Undo moves where the story is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped over. Both restore the world state from a per-node snapshot rather than just the text, and a memory derived from a turn now behind the head stops being retrieved without being deleted or re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the displaced future stays on the line it was written for, and ordinary Redo stops offering it. Nothing a retry replaces is discarded either — the old attempt stays as another take of that turn, one keystroke and one click from becoming a branch of its own.

  • Save Points. Name a moment — "Before entering the abbey" — keep playing, restart the app, and come back to it. Restoring one moves the story back to that moment and deletes nothing: the turns you wrote after it stay, Redo still walks forward into them, and writing something different from the Save Point is what starts a new line while the old one is kept. A Save Point is a name for a position and holds no copy of the story, so restoring it is the same movement Undo makes (backend/app/routers/adventures/checkpoints.py, backend/app/head.py). They last until you delete them: deleting one deletes no story, and deleting a branch a Save Point is kept on is refused until you remove the Save Point yourself, so nothing takes a named moment away behind your back.

  • Import and export, as a recovery contract. A campaign exports as one JSON file, ai-dnd-adventure-v3, and imports into a clean install on another machine. It carries the whole tree — every branch, every take, the fork points, which branches the story has left behind, the Save Points and the position it is being read at — and, since it is meant to be recovery rather than a copy of the text, everything that explains that story: the authoritative state and the typed events behind it, the exact prompt each turn was given and the passages it was shown, the summaries with the coordinates that decide whether they still apply, and your imported files with their classifications. A restored campaign can still answer "why does the state say this?" and "what was the narrator actually told?" — after the source file has been deleted and the canon edited since.

    A campaign opens where its head says, never at a Save Point merely because it has one. Exported after two Undos, it imports still undone, with its retained future intact. Search indexes are not carried: they are rebuilt from the content, before the import returns. Nothing about your machine travels — no endpoint, no model, no path — so importing somebody's campaign never reconfigures your inference, and a campaign imports whether or not you have the model that wrote it. Older files still import: the flat single-line format, files that predate the head position, and files that predate everything above. AI Dungeon-compatible scenario format is still read and written for scenarios and story cards.

  • A verified backup of everything, taken while you play. Settings → Back up everything on this machine writes a copy of the whole database through SQLite's online backup API — not a file copy, which of a live database can read one page before a transaction and another after it and produce a file that opens and is quietly missing rows. It is checked with PRAGMA quick_check before it is kept, and an existing backup is never overwritten (backend/app/backup.py). Restoring one is a documented stop-move-start procedure in DEVELOPMENT.md, deliberately not a button: replacing the file the running application has open is how you lose both copies.

  • Single user, no accounts. There is no sign-up, no login, no session and no API key anywhere in the product. The storyteller API binds to loopback and is unauthenticated by design, because the only person who can reach it is the person running it. A new install starts with a short pre-played adventure, so the first screen shows real turns and their world-state changes rather than an empty page (backend/app/starter.py).

  • A refusal you can rely on. The inference endpoint is checked against an address allowlist when you save it and again before every request, so a public endpoint is refused even if the setting is edited in the database directly. TLS verification is never traded against reachability: a privately issued certificate is verified against your machine's own trust store, and there is no bypass switch.

Screenshots

None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up, a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were removed rather than left standing as a picture of a product that no longer exists. The M4 closeout drove the real application in a real browser, so the screens exist and work; taking presentable screenshots of them is a job for the UI pass in M8.

Quick start

Docker (any OS)

docker compose up --build

Open http://localhost:8000. Your data persists in a named volume across restarts.

Windows

cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
cd frontend; npm install; cd ..
.\start.ps1

Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).

macOS / Linux

./start.sh   # creates the venv and installs dependencies on first run

Open http://localhost:5173.

Connect a model

Ollama is the inference backend v1 supports. Open Settings in the app and point it at one:

Where Ollama runs Endpoint URL Notes
The same machine http://localhost:11434/v1 the default, and the simplest thing that works
A machine on your own network http://<host>:11434/v1 or https://<host>/v1 explicitly configured; see below

Model name, generation parameters, and (optionally) summary and embedding models for the Memory Bank are configured there too. No config files and no rebuild are needed. There is no API key field, because there is nothing to authenticate to.

The adapter underneath speaks the OpenAI-compatible protocol, because that is what Ollama serves. That is an implementation detail, not a promise of support for arbitrary local servers that happen to speak the same protocol. Public and cloud inference endpoints are prohibited outright — see planning/DECISIONS/002-ollama-only-v1.md and planning/DECISIONS/011-local-inference-endpoint-policy.md.

What the endpoint policy allows

The address is checked when you save it and again before every request. Only loopback and private-network addresses are accepted; every public address is refused, by address rather than by hostname, so a name that resolves outward is refused too. A well-known cloud inference host is named in the error message only so the refusal says why.

Running the model on a second machine you control is supported and expected — that machine does the inference while the storyteller itself stays bound to loopback on yours. If that machine serves HTTPS with a certificate from a CA you installed, it works: certificates are verified against your operating system's trust store as well as the bundled one. Verification itself is never relaxed, and there is no option to turn it off.

Playing against a local shim (development only)

backend/tools/claude_shim.py serves an OpenAI-compatible endpoint on 127.0.0.1:8787 backed by a command-line tool, which is useful for testing the turn engine against a stronger model. Each request spawns one process, which suits the engine: the app assembles the whole prompt every turn and expects a stateless endpoint.

cd backend
.venv/bin/python tools/claude_shim.py      # listens on 127.0.0.1:8787

Set the base URL to http://127.0.0.1:8787/v1 and pick a model the tool offers. Set the reasoning budget to 0 or -1: a positive budget sends a reasoning.max_tokens field that some models reject. Embeddings are not served — leave the embedding model blank, or point the Memory Bank at an endpoint that serves one.

The shim has no authentication and spends whatever quota backs it, so run it on loopback and leave it there.

How a turn works

player input
  → assemble context:  [narrator prompt] + [world state + stat guide] + [AI instructions]
                       + [plot essentials] + [story summary] + [retrieved memories]
                       + [retrieved imported knowledge, framed by class and
                         bounded by its own budget]
                       + [history along this branch, token-budgeted]
                       + [author's note] + [player action]
  → snapshot context (Insights)
  → provider adapter → AI (streamed)
  → extract + referee the world-state delta block, strip it from the prose
  → store & render

Architecture

frontend/   React + Vite SPA  ──HTTP/SSE──►  backend/  FastAPI
                                              ├─ routers/      scenarios, adventures, knowledge, story cards, chat, settings, debug
                                              ├─ models.py     SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile
                                              ├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting)
                                              ├─ endpoints.py  the inference-endpoint address policy
                                              ├─ contextwindow.py  what the server will actually accept, and the cap
                                              ├─ tlstrust.py   one TLS context: the OS trust store unioned with certifi's
                                              ├─ tree.py       forking, promotion, and where a node is placed
                                              ├─ head.py       the active head: where the story is read, and what moving it costs
                                              ├─ checkpoints   Save Points: durable names for positions, in routers/adventures/
                                              ├─ attempts.py   the takes of one turn, grouped by parent
                                              ├─ context/      prompt assembly under a token budget + lineage/history windowing
                                              ├─ narrative/    the authoritative state: typed events, validation, snapshots
                                              ├─ worldstate/   the inherited RPG stat engine — legacy, no longer authoritative
                                              ├─ memorybank.py auto-summarization + embedding retrieval
                                              ├─ knowledge/    the imported library: import, chunk, FTS5, embed, rank, inject
                                              ├─ bundle.py     the export/import formats: v3, and readers for v2 and v1
                                              ├─ media/        the future-media seam: scene packets, visual profiles, provider contracts
                                              ├─ backup.py     a verified whole-database copy, via SQLite's backup API
                                              ├─ providers/    OpenAI-compatible adapter, streaming
                                              └─ data.db       SQLite (path overridable via AIDND_DB_PATH)

In production the backend serves the built SPA from one port (see Dockerfile). In development, Vite proxies /api to FastAPI.

Tests

1,191 backend tests: unit tests plus full HTTP integration through the real turn engine, with the model provider mocked. They run with no route to the Internet, which is a requirement rather than a convenience — an offline claim proved on a machine that has been online once proves nothing. A further handful need a real local model and skip without one; they exist because a mocked provider can leave the production wiring dead while the suite stays green, which this project has shipped twice.

cd backend && pip install -r requirements.txt -r requirements-dev.txt
python -m pytest tests/

Notes on performance

Two of the tests exist because of bugs that were measured rather than guessed at. They are the most interesting engineering in the repo.

  • Database egress, cut about 189x. Every adventure load pulled Action.context_snapshot, the entire assembled prompt at about 74 KB per turn, to read two small fields off it. Moving those fields into their own columns and marking the heavy ones deferred took one adventure load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL, so the old data never crosses the wire. tests/test_egress.py hooks into SQLAlchemy's cursor events and fails if a bulk load ever names those columns again.
  • Turn cost, made flat. Assembling a turn walked the whole story, so cost scaled with story length: 839 KB of reads at turn 200. backend/app/context/history.py now serves tails and slices from SQL and measures what it fetched. The same turn now costs 129 KB and stops growing at around turn 50.
  • Branching that costs nothing to read. A branch stores no turns. It stores where it left its parent, and borrows everything above that. A 40-turn story forked twenty times loads in 31,652 bytes against 31,433 bytes for the same story flat: a 1.007x ratio, or about 103 bytes per branch. Reads stay cheap because the lineage is windowed the same way the history is, so the number of SQL clauses is bounded by the context window rather than by the number of forks.

Repo notes

  • planning/ is this fork's own package: the product specification, the architecture decisions, the milestone plan, the acceptance contract, and a review report for every milestone shipped. Start at planning/README.md.
  • planning/archive/ holds the Phase 0 research that chose this base and the completed milestone reports. It is history, not instruction.
  • DEVELOPMENT.md is how to set the project up, point it at a model, and run the tests. PROVENANCE.md records what came from upstream and what changed.
  • backend/.env.example lists the two environment variables the backend reads. Everything about the model is a runtime setting on the Settings page instead.
  • Upstream's own plan/ build log and docs/ project site were removed in the 2026-09-03 documentation pass: they described AI-DnD's hosted, scripted, multi-user product. Both are still in Git history, and in upstream.

License

MIT