Files
interactive-story/README.md
T
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00

380 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Adventure Storyteller
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
An interactive storytelling app that runs entirely on your own machine, with your own model.
Start a campaign from a short form and play an open-ended story where a local LLM narrates the
world, keeps track of what is true, and remembers what happened. Since M8 the browser has one
entry point — a campaign library — and one natural-language input; the scenario gallery and its
editor are gone from the interface, though the AI Dungeon-compatible scenario *format* is still
supported for import and export.
This is the **Adventure Storyteller** fork of [AI-DnD](https://github.com/parththakkar106/AI-DnD).
It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model
you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support,
the Postgres path, and the JavaScript scripting engine have all been removed rather than
disabled. What is left is a storyteller you can run offline.
> **Local-only, by design.** The app talks to one place — an Ollama-compatible endpoint on this
> machine or on a machine you control on your own network — and it refuses to be pointed at a
> public address. There is no telemetry, no account, no cloud inference, and nothing is fetched
> at runtime from the Internet.
>
> For the internals, read [`planning/TECHNICAL-DESIGN.md`](planning/TECHNICAL-DESIGN.md) and
> [`planning/CONTEXT-AND-MEMORY.md`](planning/CONTEXT-AND-MEMORY.md), which cover the context
> budgeting, the state model and the memory bank as this fork builds them.
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing
everything in one SQLite file.
On the play screen, the left rail carries live world state. The AI proposes changes each turn
and a Python engine decides what actually sticks; the chip under the narration reports what
changed. The `‹ 2/2 ›` under a turn steps between the takes it has, and writing below a take
that isn't the live one starts a new branch.
## Features
- **The full play loop, in one box.** You write what you do or say in a single
natural-language field — an action and a piece of quoted dialogue are both just what you
wrote — with **Continue** for a beat you do not act in and a **Story direction** toggle for
speaking to the narrator rather than in the story. Responses stream (SSE), and Undo, Redo,
Retry and Edit sit beside the box. Correcting narrator prose does not overwrite it: the
correction becomes a new continuation carrying the state it implies, and the original
narration keeps its own future as retained history. Reasoning models are supported: the
narrator's thinking streams into a collapsible panel with its own token budget.
- **Retained history, without a tree to manage.** Underneath, the story is a tree: any turn
can hold more than one **take**, and `‹ 2/4 ›` steps between them. Stepping is free — the
story below simply empties, and the server is told nothing. Writing below a take that is not
the live one is what starts a different continuation. Branches borrow their ancestors' turns
instead of copying them, so one costs about 100 bytes, and a 20-fork story loads within 1% of
the same story flat (`backend/app/tree.py`, `backend/app/context/lineage.py`).
**None of that vocabulary reaches the reader.** M8 removed the branch panel and the tree
overlay from the browser: what you get is Undo, Redo, Retry, takes, Save Points and Restore,
and nothing on screen says branch, fork, node or head. The mechanism is unchanged and still
fully tested — this is a decision about what you are asked to understand, not about what the
product can do.
- **Authoritative narrative state, and the application owns it.** The story tracks who exists,
where they are, what they hold, what is true, how they are tied to each other, and what is
still open — as generic entities, facts, relationships and threads, with no genre baked in.
The same schema holds a silver key in an abbey and a data crystal on an orbital station.
The AI proposes **typed events with absolute values** (`set_possession`, `add_fact`,
`set_current_location` …), and a Python validator decides what is accepted: unknown event
types are refused, references must resolve, campaign canon outranks the narration, and the
machine-readable block never reaches the reader (`backend/app/narrative/`). Every accepted
change is recorded with what it was before and which turn caused it, so the Story State panel
can show what changed and why. You can correct it by hand, and your correction outranks the
story.
- **A context engine you can account for.** Memory, the author's note, the campaign's own
rules, the authoritative state, the summary that applies here, and the retrieved imported
passages are assembled under one token budget, in an order chosen so that a section which
changes does not re-price the cached prefix above it (`backend/app/context/builder.py`).
**And the budget is the one your server will actually read.** Ollama enforces a context
window of its own — 4,096 by default on a machine with no VRAM — and a larger prompt is not
refused, it is silently trimmed from the *oldest* end, which here is the narrator's rules and
your campaign's canon. The application asks the server what window your model gets and caps
the prompt to it, so what a small window costs is history rather than the canon at the front
(`backend/app/contextwindow.py`). If it cannot check, it says so instead of assuming — and
on a server it cannot ask, which is any server that is not Ollama, `context_window_override`
in settings lets you state the window so the prompt is still capped. A window the server
itself reported always wins over that, and a declared one is never reported as verified.
Since v1.1 the prompt also stops short of that window on purpose. It leaves
`max(256, 5% of the window)` tokens free, because your model counts tokens differently from
the application, and the v1 evidence came within 23 tokens of the edge. After each turn,
the server's own count of what it read is compared with what was sent. A turn the server
appears to have truncated is kept, flagged and shown in the context inspector, not left to
pass silently. A model that isn't loaded yet, and so cannot report its window, is loaded
once before the turn is built, so the first turn of a session gets the real window too.
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
card used to arrive in front of it as a world fact with no class, no visibility, no source and
nothing to switch it off, competing with your imported Canon for the same budget; the knowledge
library below replaces it, and does all of that explicitly.
- **Total prompt transparency.** Every turn stores the exact prompt sent to the model.
**Inspect context** on any narrator turn opens a readable account of what it was given —
what it remembered, what it read, what it believes, and what each part cost — with the
assembled prompt itself kept as an advanced section rather than opening on a wall of text.
A passage that came from an imported file links back to the file it came from.
- **Auto-summarization and Memory Bank.** The modern AI Dungeon memory system: AI-generated
memories every few actions, a running story summary, and embedding-based retrieval that
pulls old-but-relevant facts back into context, with similarity scores visible in the
context inspector
(`backend/app/memorybank.py`).
- **An imported knowledge library, classified by how much authority it has.** Import your own
local `.txt` and `.md` files — a setting bible, character notes, research, a passage whose
voice you want the prose to have — as **Canon**, **Reference** or **Inspiration**. The class
is not a label: it decides the words the passage is framed with in the prompt, the weight it
carries when passages are ranked, and which budget it competes in when the context is tight.
Canon can establish what is true; Reference informs detail without establishing anything;
Inspiration influences tone and introduces no facts at all. Retrieval is **hybrid and local**:
a SQLite FTS5 index finds the names and invented terms an embedding is worst at, local Ollama
embeddings find what you meant when your words differ from the file's, and the two are merged,
de-duplicated and reranked by relevance × class. Lexical search is a supported production
path, not a fallback — the library works with no embedding model at all. Canon you mark
**always include** is supplied on every turn whether or not the scene resembles it, and Canon
you mark **narrator only** is given to the narrator with instructions not to let the
protagonist know it. Every passage that reaches a prompt is listed in the context inspector with its file,
class, heading, passage number, scores and token cost, and that record is kept in the turn, so
deleting a source never erases the evidence of what an old turn was shown
(`backend/app/knowledge/`).
- **Imported text is data, never instruction.** Every imported passage is delimited in the
prompt as untrusted data with the authority order stated in words, so "ignore all previous
instructions" inside a file is a sentence in a file. Nothing is fetched: a URL in a source is
text, a remote Markdown image never loads, and no endpoint anywhere takes a filesystem path —
a source arrives as an upload, so there is no path for a traversal to escape from. Imported
content is displayed as inert text and never rendered as HTML.
- **Undo, Redo, and retry that roll back state and delete nothing.** Undo moves where the story
is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped
over. Both restore the world state from a per-node snapshot rather than just the text, and a
memory derived from a turn now behind the head stops being retrieved without being deleted or
re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the
displaced future stays on the line it was written for, and ordinary Redo stops offering it.
Nothing a retry replaces is discarded either — the old attempt stays as another take of that
turn, one keystroke and one click from becoming a branch of its own.
- **Save Points.** Name a moment — "Before entering the abbey" — keep playing,
restart the app, and come back to it. Restoring one moves the story back to
that moment and deletes nothing: the turns you wrote after it stay, Redo still
walks forward into them, and writing something different from the Save Point
is what starts a new line while the old one is kept. A Save Point is a name for
a position and holds no copy of the story, so restoring it is the same
movement Undo makes (`backend/app/routers/adventures/checkpoints.py`,
`backend/app/head.py`). They last until *you* delete them: deleting one deletes
no story, and deleting a branch a Save Point is kept on is refused until you
remove the Save Point yourself, so nothing takes a named moment away behind
your back.
- **Import and export, as a recovery contract.** A campaign exports as one JSON file,
`ai-dnd-adventure-v3`, and imports into a clean install on another machine. It carries the whole
tree — every branch, every take, the fork points, which branches the story has left behind, the
Save Points and the position it is being read at — and, since it is meant to be *recovery*
rather than a copy of the text, everything that explains that story: the authoritative state
and the typed events behind it, **the exact prompt each turn was given and the passages it was
shown**, the summaries with the coordinates that decide whether they still apply, and your
imported files with their classifications. A restored campaign can still answer "why does the
state say this?" and "what was the narrator actually told?" — after the source file has been
deleted and the canon edited since.
A campaign opens where its head says, never at a Save Point merely because it has one. Exported
after two Undos, it imports still undone, with its retained future intact. Search indexes are
not carried: they are rebuilt from the content, before the import returns. Nothing about your
machine travels — no endpoint, no model, no path — so importing somebody's campaign never
reconfigures your inference, and a campaign imports whether or not you have the model that
wrote it. Older files still import: the flat single-line format, files that predate the head
position, and files that predate everything above. AI Dungeon-compatible scenario format is
still read and written for scenarios and story cards.
- **A verified backup of everything, taken while you play.** Settings → *Back up everything on
this machine* writes a copy of the whole database through SQLite's online backup API — not a
file copy, which of a live database can read one page before a transaction and another after it
and produce a file that opens and is quietly missing rows. It is checked with `PRAGMA
quick_check` before it is kept, and an existing backup is never overwritten
(`backend/app/backup.py`). Restoring one is a documented stop-move-start procedure in
`DEVELOPMENT.md`, deliberately not a button: replacing the file the running application has
open is how you lose both copies.
- **Single user, no accounts.** There is no sign-up, no login, no session and no API key
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
design, because the only person who can reach it is the person running it. A new install
starts with a short pre-played adventure, so the first screen shows real turns and their
world-state changes rather than an empty page (`backend/app/starter.py`).
- **A refusal you can rely on.** The inference endpoint is checked against an address
allowlist when you save it and again before every request, so a public endpoint is refused
even if the setting is edited in the database directly. TLS verification is never traded
against reachability: a privately issued certificate is verified against your machine's own
trust store, and there is no bypass switch.
## Screenshots
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
removed rather than left standing as a picture of a product that no longer exists. The
screens exist and are driven in a real browser by the release harness
(`backend/tools/m11_browser.py`). Presentable screenshots of them have not been taken.
## Quick start
### Docker (any OS)
```sh
docker compose up --build
```
Open http://localhost:8000. Your data persists in a named volume across restarts.
### Windows
```powershell
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
cd frontend; npm install; cd ..
.\start.ps1
```
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
### macOS / Linux
```sh
./start.sh # creates the venv and installs dependencies on first run
```
Open http://localhost:5173.
## Connect a model
Ollama is the inference backend v1 supports. Open **Settings** in the app and point it at one:
| Where Ollama runs | Endpoint URL | Notes |
|---|---|---|
| The same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works |
| A machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | explicitly configured; see below |
Model name, generation parameters, and (optionally) summary and embedding models for the
Memory Bank are configured there too. No config files and no rebuild are needed. There is no
API key field, because there is nothing to authenticate to.
The adapter underneath speaks the OpenAI-compatible protocol, because that is what Ollama
serves. That is an implementation detail, not a promise of support for arbitrary local
servers that happen to speak the same protocol. Public and cloud inference endpoints are
prohibited outright — see `planning/DECISIONS/002-ollama-only-v1.md` and
`planning/DECISIONS/011-local-inference-endpoint-policy.md`.
### What the endpoint policy allows
The address is checked when you save it and again before every request. Only loopback and
private-network addresses are accepted; every public address is refused, by address rather than
by hostname, so a name that resolves outward is refused too. A well-known cloud inference host
is named in the error message only so the refusal says *why*.
Running the model on a second machine you control is supported and expected — that machine
does the inference while the storyteller itself stays bound to loopback on yours. If that
machine serves HTTPS with a certificate from a CA you installed, it works: certificates are
verified against your operating system's trust store as well as the bundled one. Verification
itself is never relaxed, and there is no option to turn it off.
### Playing against a local shim (development only)
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787`
backed by a command-line tool, which is useful for testing the turn engine against a stronger
model. Each request spawns one process, which suits the engine: the app assembles the whole
prompt every turn and expects a stateless endpoint.
```sh
cd backend
.venv/bin/python tools/claude_shim.py # listens on 127.0.0.1:8787
```
Set the base URL to `http://127.0.0.1:8787/v1` and pick a model the tool offers. Set the
reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field that
some models reject. Embeddings are not served — leave the embedding model blank, or point the
Memory Bank at an endpoint that serves one.
The shim has no authentication and spends whatever quota backs it, so run it on loopback and
leave it there.
## How a turn works
```
player input
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
+ [plot essentials] + [story summary] + [retrieved memories]
+ [retrieved imported knowledge, framed by class and
bounded by its own budget]
+ [history along this branch, token-budgeted]
+ [author's note] + [player action]
→ snapshot context (Insights)
→ provider adapter → AI (streamed)
→ extract + referee the world-state delta block, strip it from the prose
→ store & render
```
## Architecture
```
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting)
├─ endpoints.py the inference-endpoint address policy
├─ contextwindow.py what the server will actually accept, and the cap
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
├─ tree.py forking, promotion, and where a node is placed
├─ head.py the active head: where the story is read, and what moving it costs
├─ checkpoints Save Points: durable names for positions, in routers/adventures/
├─ attempts.py the takes of one turn, grouped by parent
├─ context/ prompt assembly under a token budget + lineage/history windowing
├─ narrative/ the authoritative state: typed events, validation, snapshots
├─ worldstate/ the inherited RPG stat engine — legacy, no longer authoritative
├─ memorybank.py auto-summarization + embedding retrieval
├─ knowledge/ the imported library: import, chunk, FTS5, embed, rank, inject
├─ bundle.py the export/import formats: v3, and readers for v2 and v1
├─ media/ the future-media seam: scene packets, visual profiles, provider contracts
├─ backup.py a verified whole-database copy, via SQLite's backup API
├─ providers/ OpenAI-compatible adapter, streaming
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
```
In production the backend serves the built SPA from one port (see `Dockerfile`). In
development, Vite proxies `/api` to FastAPI.
## Tests
1,524 backend tests (1,507 run everywhere, 17 need a real local model and skip without one): unit
tests plus full HTTP integration through the real turn engine, with
the model provider mocked. They run with no route to the Internet, which is a requirement
rather than a convenience — an offline claim proved on a machine that has been online once
proves nothing. A further handful need a real local model and skip without one; they exist
because a mocked provider can leave the production wiring dead while the suite stays green,
which this project has shipped twice.
```sh
cd backend && pip install -r requirements.txt -r requirements-dev.txt
python -m pytest tests/
```
## Notes on performance
Two of the tests exist because of bugs that were measured rather than guessed at. They are the
most interesting engineering in the repo.
- **Database egress, cut about 189x.** Every adventure load pulled `Action.context_snapshot`,
the entire assembled prompt at about 74 KB per turn, to read two small fields off it. Moving
those fields into their own columns and marking the heavy ones `deferred` took one adventure
load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL, so the
old data never crosses the wire. `tests/test_egress.py` hooks into SQLAlchemy's cursor events
and fails if a bulk load ever names those columns again.
- **Turn cost, made flat.** Assembling a turn walked the whole story, so cost scaled with story
length: 839 KB of reads at turn 200. `backend/app/context/history.py` now serves tails and
slices from SQL and measures what it fetched. The same turn now costs 129 KB and stops
growing at around turn 50.
- **Branching that costs nothing to read.** A branch stores no turns. It stores where it left
its parent, and borrows everything above that. A 40-turn story forked twenty times loads in
31,652 bytes against 31,433 bytes for the same story flat: a 1.007x ratio, or about 103 bytes
per branch. Reads stay cheap because the lineage is windowed the same way the history is, so
the number of SQL clauses is bounded by the context window rather than by the number of
forks.
## Repo notes
- **Status:** **v1.0.0 was released on 2026-09-14.** Milestones M1-M11 are
complete, and the v1 release gate passed (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md),
§T). The signed tag `v1.0.0` and `main` both point at the signed release
commit `432f041`. v1.1 development has begun on the `v1.1-development` branch;
its plan is [`planning/V1.1-PLAN.md`](planning/V1.1-PLAN.md).
- `planning/` is this fork's own package: the product specification, the architecture
decisions, the milestone plan, the acceptance contract, and a review report for every
milestone shipped. Start at [`planning/README.md`](planning/README.md).
- [`planning/archive/`](planning/archive/README.md) holds the Phase 0 research that chose this
base and the completed milestone reports. It is history, not instruction.
- [`DEVELOPMENT.md`](DEVELOPMENT.md) is how to set the project up, point it at a model, and run
the tests. [`PROVENANCE.md`](PROVENANCE.md) records what came from upstream and what changed.
- `backend/.env.example` lists the two environment variables the backend reads. Everything
about the model is a runtime setting on the Settings page instead.
- Upstream's own `plan/` build log and `docs/` project site were removed in the 2026-09-03
documentation pass: they described AI-DnD's hosted, scripted, multi-user product. Both are
still in Git history, and in upstream.
## License
[MIT](LICENSE)