Files
interactive-story/README.md
T
JesseMarkowitzandClaude Opus 5 e08d49c3eb M4: add durable named Save Points
A Save Point is a name for a story position, and restoring one is head
movement. That is the whole architecture, and it is what ADR 012 and
BUILD-MILESTONES' note on M4 asked for: M3 made the head a stored
(branch, depth) and made arriving at one a row lookup plus a state restore,
so a Save Point needs no restore machinery of its own.

What the user gets:

- Name the moment they are reading, keep playing, restart the app, and come
  back to it. Restoring moves the story back and deletes nothing: the later
  turns stay, Redo still walks forward into them, and writing something
  different is what starts a new line while the old one is kept.
- Rename, delete, and a list, in a Save Points panel beside the branch panel,
  with a Save Point button next to Undo and Redo. Both confirmations say what
  is *not* destroyed, because that is the part the screen cannot show.
- Save Points survive export and import.

What was deliberately not built:

- No second restore path. `head.move_to_node` is the only new movement: its
  depth half is M3's `head.move_to` unchanged, and its branch half is the
  single assignment `switch_branch` already makes. No head field is written
  in the checkpoint router, nothing reconstructs state, nothing prunes a
  memory, nothing copies or deletes a turn, and restore never forks — the
  first write below the restored head does, through `fork_if_behind_head`.
- No automatic cleanup. A Save Point behind the head, or naming a line the
  story left, is doing its job (STORY-BRANCH-SEMANTICS §19). The one removal
  is a cascade: deleting a branch takes its Save Points, as it takes its
  memories, because the story they named went with it.
- No new ADR. ADR 012 already decides the architecture, and a table is not a
  decision.

The one call the planning package did not already make: restore moves the
branch half of the head only when the coordinate is off the path being read.
Doing it unconditionally would quietly hand back an abandoned continuation
whenever a Save Point in a shared prefix was restored; never doing it would
make a Save Point on a departed line unrestorable, which contradicts §19.
TECHNICAL-DESIGN §8.8 records it.

Schema: a `checkpoints` table holding a name, an optional note and a
(branch, depth) coordinate — no copy of any story. `create_all` builds it as
it did `memories` and `branches`; migration 80 adds the index. No backfill,
because nobody had named a position before M4.

The coordinate is deliberately not an action id: one coordinate holds every
attempt at a turn and exactly one is live, so a coordinate follows a retry
where a row id would pin a take the story no longer tells.

Tests: 680 pass (638 before). 42 new in tests/test_save_points.py covering
D11-D14, I04, L03, E-series lineage and memory isolation after restore and
divergence, the edge cases, and an M3-database migration. One pre-existing
fixture in test_tree_migration.py needed `checkpoints` added to its drop
list — SQLite refuses to drop a table another table references.

Not verified: the browser. No session has had a usable one, so the Save
Point panel's DOM behaviour is unobserved — as M3's Redo control still is.
The twenty-step sequence was driven over HTTP against a live server with a
real process restart instead, and all seventeen checks pass. M4 is
implemented, not accepted: no review has been written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-03 18:48:54 -04:00

274 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Adventure Storyteller
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
An interactive storytelling app that runs entirely on your own machine, with your own model.
Create scenarios and play open-ended adventures where a local LLM narrates the world, keeps
track of what is true, and remembers what happened.
This is the **Adventure Storyteller** fork of [AI-DnD](https://github.com/parththakkar106/AI-DnD).
It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model
you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support,
the Postgres path, and the JavaScript scripting engine have all been removed rather than
disabled. What is left is a storyteller you can run offline.
> **Local-only, by design.** The app talks to one place — an Ollama-compatible endpoint on this
> machine or on a machine you control on your own network — and it refuses to be pointed at a
> public address. There is no telemetry, no account, no cloud inference, and nothing is fetched
> at runtime from the Internet.
>
> For the internals, read [`planning/TECHNICAL-DESIGN.md`](planning/TECHNICAL-DESIGN.md) and
> [`planning/CONTEXT-AND-MEMORY.md`](planning/CONTEXT-AND-MEMORY.md), which cover the context
> budgeting, the state model and the memory bank as this fork builds them.
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing
everything in one SQLite file.
On the play screen, the left rail carries live world state. The AI proposes changes each turn
and a Python engine decides what actually sticks; the chip under the narration reports what
changed. The `‹ 2/2 ›` under a turn steps between the takes it has, and writing below a take
that isn't the live one starts a new branch.
## Features
- **The full play loop.** Do / Say / Story / Continue actions, streamed AI responses (SSE),
retry, undo, redo, and edit. Reasoning models are supported: "thinking" streams into a collapsible
💭 panel with its own token budget.
- **A branching story tree.** The story is a tree, not a list. Any turn can hold more than one
**take**, and `‹ 2/4 ›` steps between them. Stepping is free: the story below simply empties,
and the server is told nothing. Writing below a take that isn't the live one is what makes a
branch. Branches borrow their ancestors' turns instead of copying them, so a fork costs about
100 bytes, and a 20-fork story loads within 1% of the same story flat. Switching restores that
line's world state, script state, and cooldown clocks. A branch panel switches, renames, and
deletes; **⌗ See the tree** draws every line against the story's own clock
(`backend/app/tree.py`, `backend/app/context/lineage.py`).
- **An RPG world-state engine.** A scenario can declare stats, flags, milestones, and a named
cast; the adventure carries their live values. The AI proposes deltas, and a Python engine
referees them: it clamps values to range, enforces per-turn caps and cooldowns, keeps counters
monotonic and milestones sticky, then strips the machine-readable block out of the prose
(`backend/app/worldstate/`: `apply.py` clamps, `parse.py` reads the block back).
Word-labeled bands (`40–60: minor damage`) make the
model reliable at it. No dice and no scripting are required.
- **AI Dungeon-compatible context engine.** Memory, author's note, and story cards (world
info) are triggered by keywords in recent story text, then assembled under a token budget
(`backend/app/context/builder.py`).
- **Insights: total prompt transparency.** Every turn stores the exact prompt sent to the
model. Open 🔍 on any AI action to see each context component, its token cost, and why it was
included.
- **Auto-summarization and Memory Bank.** The modern AI Dungeon memory system: AI-generated
memories every few actions, a running story summary, and embedding-based retrieval that
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
(`backend/app/memorybank.py`).
- **Undo, Redo, and retry that roll back state and delete nothing.** Undo moves where the story
is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped
over. Both restore the world state from a per-node snapshot rather than just the text, and a
memory derived from a turn now behind the head stops being retrieved without being deleted or
re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the
displaced future stays on the line it was written for, and ordinary Redo stops offering it.
Nothing a retry replaces is discarded either — the old attempt stays as another take of that
turn, one keystroke and one click from becoming a branch of its own.
- **Save Points.** Name a moment — "Before entering the abbey" — keep playing,
restart the app, and come back to it. Restoring one moves the story back to
that moment and deletes nothing: the turns you wrote after it stay, Redo still
walks forward into them, and writing something different from the Save Point
is what starts a new line while the old one is kept. A Save Point is a name for
a position and holds no copy of the story, so restoring it is the same
movement Undo makes (`backend/app/routers/adventures/checkpoints.py`,
`backend/app/head.py`). They last until you delete them, and deleting one
deletes no story.
- **Import and export.** AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree:
every branch, every take, the fork points, which branches the story has left behind, the Save
Points and the position it is being read at — all of them chosen rather than computed, which is
the rule for what a bundle carries. A campaign opens where its head says, never at a Save Point
merely because it has one. A campaign exported after two Undos imports still undone, with its
retained future intact, instead of silently reopening at its newest turn. Files that predate
the head position, and files saved in the old single-line format, still import.
- **Single user, no accounts.** There is no sign-up, no login, no session and no API key
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
design, because the only person who can reach it is the person running it. A new install
starts with a short pre-played adventure, so the first screen shows real turns and their
world-state changes rather than an empty page (`backend/app/starter.py`).
- **A refusal you can rely on.** The inference endpoint is checked against an address
allowlist when you save it and again before every request, so a public endpoint is refused
even if the setting is edited in the database directly. TLS verification is never traded
against reachability: a privately issued certificate is verified against your machine's own
trust store, and there is no bypass switch.
## Screenshots
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
removed rather than left standing as a picture of a product that no longer exists. New ones
are taken when the browser smoke test M3 still owes is run.
## Quick start
### Docker (any OS)
```sh
docker compose up --build
```
Open http://localhost:8000. Your data persists in a named volume across restarts.
### Windows
```powershell
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
cd frontend; npm install; cd ..
.\start.ps1
```
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
### macOS / Linux
```sh
./start.sh # creates the venv and installs dependencies on first run
```
Open http://localhost:5173.
## Connect a model
Ollama is the inference backend v1 supports. Open **Settings** in the app and point it at one:
| Where Ollama runs | Endpoint URL | Notes |
|---|---|---|
| The same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works |
| A machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | explicitly configured; see below |
Model name, generation parameters, and (optionally) summary and embedding models for the
Memory Bank are configured there too. No config files and no rebuild are needed. There is no
API key field, because there is nothing to authenticate to.
The adapter underneath speaks the OpenAI-compatible protocol, because that is what Ollama
serves. That is an implementation detail, not a promise of support for arbitrary local
servers that happen to speak the same protocol. Public and cloud inference endpoints are
prohibited outright — see `planning/DECISIONS/002-ollama-only-v1.md` and
`planning/DECISIONS/011-local-inference-endpoint-policy.md`.
### What the endpoint policy allows
The address is checked when you save it and again before every request. Only loopback and
private-network addresses are accepted; every public address is refused, by address rather than
by hostname, so a name that resolves outward is refused too. A well-known cloud inference host
is named in the error message only so the refusal says *why*.
Running the model on a second machine you control is supported and expected — that machine
does the inference while the storyteller itself stays bound to loopback on yours. If that
machine serves HTTPS with a certificate from a CA you installed, it works: certificates are
verified against your operating system's trust store as well as the bundled one. Verification
itself is never relaxed, and there is no option to turn it off.
### Playing against a local shim (development only)
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787`
backed by a command-line tool, which is useful for testing the turn engine against a stronger
model. Each request spawns one process, which suits the engine: the app assembles the whole
prompt every turn and expects a stateless endpoint.
```sh
cd backend
.venv/bin/python tools/claude_shim.py # listens on 127.0.0.1:8787
```
Set the base URL to `http://127.0.0.1:8787/v1` and pick a model the tool offers. Set the
reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field that
some models reject. Embeddings are not served — leave the embedding model blank, or point the
Memory Bank at an endpoint that serves one.
The shim has no authentication and spends whatever quota backs it, so run it on loopback and
leave it there.
## How a turn works
```
player input
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
+ [plot essentials] + [story summary] + [retrieved memories]
+ [triggered story cards] + [history along this branch, token-budgeted]
+ [author's note] + [player action]
→ snapshot context (Insights)
→ provider adapter → AI (streamed)
→ extract + referee the world-state delta block, strip it from the prose
→ store & render
```
## Architecture
```
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ scenarios, adventures, story cards, chat, settings, debug
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (79 and counting)
├─ endpoints.py the inference-endpoint address policy
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
├─ tree.py forking, promotion, and where a node is placed
├─ head.py the active head: where the story is read, and what moving it costs
├─ checkpoints Save Points: durable names for positions, in routers/adventures/
├─ attempts.py the takes of one turn, grouped by parent
├─ context/ prompt assembly under a token budget + lineage/history windowing
├─ worldstate/ the stat engine: clamps, cooldowns, bands, milestones
├─ memorybank.py auto-summarization + embedding retrieval
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
├─ providers/ OpenAI-compatible adapter, streaming
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
```
In production the backend serves the built SPA from one port (see `Dockerfile`). In
development, Vite proxies `/api` to FastAPI.
## Tests
680 backend tests: unit tests plus full HTTP integration through the real turn engine, with
the model provider mocked. They run with no route to the Internet, which is a requirement
rather than a convenience — an offline claim proved on a machine that has been online once
proves nothing.
```sh
cd backend && pip install -r requirements.txt -r requirements-dev.txt
python -m pytest tests/
```
## Notes on performance
Two of the tests exist because of bugs that were measured rather than guessed at. They are the
most interesting engineering in the repo.
- **Database egress, cut about 189x.** Every adventure load pulled `Action.context_snapshot`,
the entire assembled prompt at about 74 KB per turn, to read two small fields off it. Moving
those fields into their own columns and marking the heavy ones `deferred` took one adventure
load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL, so the
old data never crosses the wire. `tests/test_egress.py` hooks into SQLAlchemy's cursor events
and fails if a bulk load ever names those columns again.
- **Turn cost, made flat.** Assembling a turn walked the whole story, so cost scaled with story
length: 839 KB of reads at turn 200. `backend/app/context/history.py` now serves tails and
slices from SQL and measures what it fetched. The same turn now costs 129 KB and stops
growing at around turn 50.
- **Branching that costs nothing to read.** A branch stores no turns. It stores where it left
its parent, and borrows everything above that. A 40-turn story forked twenty times loads in
31,652 bytes against 31,433 bytes for the same story flat: a 1.007x ratio, or about 103 bytes
per branch. Reads stay cheap because the lineage is windowed the same way the history is, so
the number of SQL clauses is bounded by the context window rather than by the number of
forks.
## Repo notes
- `planning/` is this fork's own package: the product specification, the architecture
decisions, the milestone plan, the acceptance contract, and a review report for every
milestone shipped. Start at [`planning/README.md`](planning/README.md).
- [`planning/archive/`](planning/archive/README.md) holds the Phase 0 research that chose this
base and the completed milestone reports. It is history, not instruction.
- [`DEVELOPMENT.md`](DEVELOPMENT.md) is how to set the project up, point it at a model, and run
the tests. [`PROVENANCE.md`](PROVENANCE.md) records what came from upstream and what changed.
- `backend/.env.example` lists the two environment variables the backend reads. Everything
about the model is a runtime setting on the Settings page instead.
- Upstream's own `plan/` build log and `docs/` project site were removed in the 2026-09-03
documentation pass: they described AI-DnD's hosted, scripted, multi-user product. Both are
still in Git history, and in upstream.
## License
[MIT](LICENSE)