A Save Point is a name for a story position, and restoring one is head movement. That is the whole architecture, and it is what ADR 012 and BUILD-MILESTONES' note on M4 asked for: M3 made the head a stored (branch, depth) and made arriving at one a row lookup plus a state restore, so a Save Point needs no restore machinery of its own. What the user gets: - Name the moment they are reading, keep playing, restart the app, and come back to it. Restoring moves the story back and deletes nothing: the later turns stay, Redo still walks forward into them, and writing something different is what starts a new line while the old one is kept. - Rename, delete, and a list, in a Save Points panel beside the branch panel, with a Save Point button next to Undo and Redo. Both confirmations say what is *not* destroyed, because that is the part the screen cannot show. - Save Points survive export and import. What was deliberately not built: - No second restore path. `head.move_to_node` is the only new movement: its depth half is M3's `head.move_to` unchanged, and its branch half is the single assignment `switch_branch` already makes. No head field is written in the checkpoint router, nothing reconstructs state, nothing prunes a memory, nothing copies or deletes a turn, and restore never forks — the first write below the restored head does, through `fork_if_behind_head`. - No automatic cleanup. A Save Point behind the head, or naming a line the story left, is doing its job (STORY-BRANCH-SEMANTICS §19). The one removal is a cascade: deleting a branch takes its Save Points, as it takes its memories, because the story they named went with it. - No new ADR. ADR 012 already decides the architecture, and a table is not a decision. The one call the planning package did not already make: restore moves the branch half of the head only when the coordinate is off the path being read. Doing it unconditionally would quietly hand back an abandoned continuation whenever a Save Point in a shared prefix was restored; never doing it would make a Save Point on a departed line unrestorable, which contradicts §19. TECHNICAL-DESIGN §8.8 records it. Schema: a `checkpoints` table holding a name, an optional note and a (branch, depth) coordinate — no copy of any story. `create_all` builds it as it did `memories` and `branches`; migration 80 adds the index. No backfill, because nobody had named a position before M4. The coordinate is deliberately not an action id: one coordinate holds every attempt at a turn and exactly one is live, so a coordinate follows a retry where a row id would pin a take the story no longer tells. Tests: 680 pass (638 before). 42 new in tests/test_save_points.py covering D11-D14, I04, L03, E-series lineage and memory isolation after restore and divergence, the edge cases, and an M3-database migration. One pre-existing fixture in test_tree_migration.py needed `checkpoints` added to its drop list — SQLite refuses to drop a table another table references. Not verified: the browser. No session has had a usable one, so the Save Point panel's DOM behaviour is unobserved — as M3's Redo control still is. The twenty-step sequence was driven over HTTP against a live server with a real process restart instead, and all seventeen checks pass. M4 is implemented, not accepted: no review has been written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
274 lines
16 KiB
Markdown
274 lines
16 KiB
Markdown
# Adventure Storyteller
|
||
|
||
[](LICENSE)
|
||
|
||
An interactive storytelling app that runs entirely on your own machine, with your own model.
|
||
Create scenarios and play open-ended adventures where a local LLM narrates the world, keeps
|
||
track of what is true, and remembers what happened.
|
||
|
||
This is the **Adventure Storyteller** fork of [AI-DnD](https://github.com/parththakkar106/AI-DnD).
|
||
It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model
|
||
you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support,
|
||
the Postgres path, and the JavaScript scripting engine have all been removed rather than
|
||
disabled. What is left is a storyteller you can run offline.
|
||
|
||
> **Local-only, by design.** The app talks to one place — an Ollama-compatible endpoint on this
|
||
> machine or on a machine you control on your own network — and it refuses to be pointed at a
|
||
> public address. There is no telemetry, no account, no cloud inference, and nothing is fetched
|
||
> at runtime from the Internet.
|
||
>
|
||
> For the internals, read [`planning/TECHNICAL-DESIGN.md`](planning/TECHNICAL-DESIGN.md) and
|
||
> [`planning/CONTEXT-AND-MEMORY.md`](planning/CONTEXT-AND-MEMORY.md), which cover the context
|
||
> budgeting, the state model and the memory bank as this fork builds them.
|
||
|
||
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing
|
||
everything in one SQLite file.
|
||
|
||
On the play screen, the left rail carries live world state. The AI proposes changes each turn
|
||
and a Python engine decides what actually sticks; the chip under the narration reports what
|
||
changed. The `‹ 2/2 ›` under a turn steps between the takes it has, and writing below a take
|
||
that isn't the live one starts a new branch.
|
||
|
||
## Features
|
||
|
||
- **The full play loop.** Do / Say / Story / Continue actions, streamed AI responses (SSE),
|
||
retry, undo, redo, and edit. Reasoning models are supported: "thinking" streams into a collapsible
|
||
💭 panel with its own token budget.
|
||
- **A branching story tree.** The story is a tree, not a list. Any turn can hold more than one
|
||
**take**, and `‹ 2/4 ›` steps between them. Stepping is free: the story below simply empties,
|
||
and the server is told nothing. Writing below a take that isn't the live one is what makes a
|
||
branch. Branches borrow their ancestors' turns instead of copying them, so a fork costs about
|
||
100 bytes, and a 20-fork story loads within 1% of the same story flat. Switching restores that
|
||
line's world state, script state, and cooldown clocks. A branch panel switches, renames, and
|
||
deletes; **⌗ See the tree** draws every line against the story's own clock
|
||
(`backend/app/tree.py`, `backend/app/context/lineage.py`).
|
||
- **An RPG world-state engine.** A scenario can declare stats, flags, milestones, and a named
|
||
cast; the adventure carries their live values. The AI proposes deltas, and a Python engine
|
||
referees them: it clamps values to range, enforces per-turn caps and cooldowns, keeps counters
|
||
monotonic and milestones sticky, then strips the machine-readable block out of the prose
|
||
(`backend/app/worldstate/`: `apply.py` clamps, `parse.py` reads the block back).
|
||
Word-labeled bands (`40–60: minor damage`) make the
|
||
model reliable at it. No dice and no scripting are required.
|
||
- **AI Dungeon-compatible context engine.** Memory, author's note, and story cards (world
|
||
info) are triggered by keywords in recent story text, then assembled under a token budget
|
||
(`backend/app/context/builder.py`).
|
||
- **Insights: total prompt transparency.** Every turn stores the exact prompt sent to the
|
||
model. Open 🔍 on any AI action to see each context component, its token cost, and why it was
|
||
included.
|
||
- **Auto-summarization and Memory Bank.** The modern AI Dungeon memory system: AI-generated
|
||
memories every few actions, a running story summary, and embedding-based retrieval that
|
||
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
|
||
(`backend/app/memorybank.py`).
|
||
- **Undo, Redo, and retry that roll back state and delete nothing.** Undo moves where the story
|
||
is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped
|
||
over. Both restore the world state from a per-node snapshot rather than just the text, and a
|
||
memory derived from a turn now behind the head stops being retrieved without being deleted or
|
||
re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the
|
||
displaced future stays on the line it was written for, and ordinary Redo stops offering it.
|
||
Nothing a retry replaces is discarded either — the old attempt stays as another take of that
|
||
turn, one keystroke and one click from becoming a branch of its own.
|
||
- **Save Points.** Name a moment — "Before entering the abbey" — keep playing,
|
||
restart the app, and come back to it. Restoring one moves the story back to
|
||
that moment and deletes nothing: the turns you wrote after it stay, Redo still
|
||
walks forward into them, and writing something different from the Save Point
|
||
is what starts a new line while the old one is kept. A Save Point is a name for
|
||
a position and holds no copy of the story, so restoring it is the same
|
||
movement Undo makes (`backend/app/routers/adventures/checkpoints.py`,
|
||
`backend/app/head.py`). They last until you delete them, and deleting one
|
||
deletes no story.
|
||
- **Import and export.** AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree:
|
||
every branch, every take, the fork points, which branches the story has left behind, the Save
|
||
Points and the position it is being read at — all of them chosen rather than computed, which is
|
||
the rule for what a bundle carries. A campaign opens where its head says, never at a Save Point
|
||
merely because it has one. A campaign exported after two Undos imports still undone, with its
|
||
retained future intact, instead of silently reopening at its newest turn. Files that predate
|
||
the head position, and files saved in the old single-line format, still import.
|
||
- **Single user, no accounts.** There is no sign-up, no login, no session and no API key
|
||
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
|
||
design, because the only person who can reach it is the person running it. A new install
|
||
starts with a short pre-played adventure, so the first screen shows real turns and their
|
||
world-state changes rather than an empty page (`backend/app/starter.py`).
|
||
- **A refusal you can rely on.** The inference endpoint is checked against an address
|
||
allowlist when you save it and again before every request, so a public endpoint is refused
|
||
even if the setting is edited in the database directly. TLS verification is never traded
|
||
against reachability: a privately issued certificate is verified against your machine's own
|
||
trust store, and there is no bypass switch.
|
||
|
||
## Screenshots
|
||
|
||
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
|
||
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
|
||
removed rather than left standing as a picture of a product that no longer exists. New ones
|
||
are taken when the browser smoke test M3 still owes is run.
|
||
|
||
## Quick start
|
||
|
||
### Docker (any OS)
|
||
|
||
```sh
|
||
docker compose up --build
|
||
```
|
||
|
||
Open http://localhost:8000. Your data persists in a named volume across restarts.
|
||
|
||
### Windows
|
||
|
||
```powershell
|
||
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
|
||
cd frontend; npm install; cd ..
|
||
.\start.ps1
|
||
```
|
||
|
||
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
|
||
|
||
### macOS / Linux
|
||
|
||
```sh
|
||
./start.sh # creates the venv and installs dependencies on first run
|
||
```
|
||
|
||
Open http://localhost:5173.
|
||
|
||
## Connect a model
|
||
|
||
Ollama is the inference backend v1 supports. Open **Settings** in the app and point it at one:
|
||
|
||
| Where Ollama runs | Endpoint URL | Notes |
|
||
|---|---|---|
|
||
| The same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works |
|
||
| A machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | explicitly configured; see below |
|
||
|
||
Model name, generation parameters, and (optionally) summary and embedding models for the
|
||
Memory Bank are configured there too. No config files and no rebuild are needed. There is no
|
||
API key field, because there is nothing to authenticate to.
|
||
|
||
The adapter underneath speaks the OpenAI-compatible protocol, because that is what Ollama
|
||
serves. That is an implementation detail, not a promise of support for arbitrary local
|
||
servers that happen to speak the same protocol. Public and cloud inference endpoints are
|
||
prohibited outright — see `planning/DECISIONS/002-ollama-only-v1.md` and
|
||
`planning/DECISIONS/011-local-inference-endpoint-policy.md`.
|
||
|
||
### What the endpoint policy allows
|
||
|
||
The address is checked when you save it and again before every request. Only loopback and
|
||
private-network addresses are accepted; every public address is refused, by address rather than
|
||
by hostname, so a name that resolves outward is refused too. A well-known cloud inference host
|
||
is named in the error message only so the refusal says *why*.
|
||
|
||
Running the model on a second machine you control is supported and expected — that machine
|
||
does the inference while the storyteller itself stays bound to loopback on yours. If that
|
||
machine serves HTTPS with a certificate from a CA you installed, it works: certificates are
|
||
verified against your operating system's trust store as well as the bundled one. Verification
|
||
itself is never relaxed, and there is no option to turn it off.
|
||
|
||
### Playing against a local shim (development only)
|
||
|
||
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787`
|
||
backed by a command-line tool, which is useful for testing the turn engine against a stronger
|
||
model. Each request spawns one process, which suits the engine: the app assembles the whole
|
||
prompt every turn and expects a stateless endpoint.
|
||
|
||
```sh
|
||
cd backend
|
||
.venv/bin/python tools/claude_shim.py # listens on 127.0.0.1:8787
|
||
```
|
||
|
||
Set the base URL to `http://127.0.0.1:8787/v1` and pick a model the tool offers. Set the
|
||
reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field that
|
||
some models reject. Embeddings are not served — leave the embedding model blank, or point the
|
||
Memory Bank at an endpoint that serves one.
|
||
|
||
The shim has no authentication and spends whatever quota backs it, so run it on loopback and
|
||
leave it there.
|
||
|
||
## How a turn works
|
||
|
||
```
|
||
player input
|
||
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
|
||
+ [plot essentials] + [story summary] + [retrieved memories]
|
||
+ [triggered story cards] + [history along this branch, token-budgeted]
|
||
+ [author's note] + [player action]
|
||
→ snapshot context (Insights)
|
||
→ provider adapter → AI (streamed)
|
||
→ extract + referee the world-state delta block, strip it from the prose
|
||
→ store & render
|
||
```
|
||
|
||
## Architecture
|
||
|
||
```
|
||
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||
├─ routers/ scenarios, adventures, story cards, chat, settings, debug
|
||
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory
|
||
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (79 and counting)
|
||
├─ endpoints.py the inference-endpoint address policy
|
||
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
|
||
├─ tree.py forking, promotion, and where a node is placed
|
||
├─ head.py the active head: where the story is read, and what moving it costs
|
||
├─ checkpoints Save Points: durable names for positions, in routers/adventures/
|
||
├─ attempts.py the takes of one turn, grouped by parent
|
||
├─ context/ prompt assembly under a token budget + lineage/history windowing
|
||
├─ worldstate/ the stat engine: clamps, cooldowns, bands, milestones
|
||
├─ memorybank.py auto-summarization + embedding retrieval
|
||
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
|
||
├─ providers/ OpenAI-compatible adapter, streaming
|
||
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
|
||
```
|
||
|
||
In production the backend serves the built SPA from one port (see `Dockerfile`). In
|
||
development, Vite proxies `/api` to FastAPI.
|
||
|
||
## Tests
|
||
|
||
680 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||
the model provider mocked. They run with no route to the Internet, which is a requirement
|
||
rather than a convenience — an offline claim proved on a machine that has been online once
|
||
proves nothing.
|
||
|
||
```sh
|
||
cd backend && pip install -r requirements.txt -r requirements-dev.txt
|
||
python -m pytest tests/
|
||
```
|
||
|
||
## Notes on performance
|
||
|
||
Two of the tests exist because of bugs that were measured rather than guessed at. They are the
|
||
most interesting engineering in the repo.
|
||
|
||
- **Database egress, cut about 189x.** Every adventure load pulled `Action.context_snapshot`,
|
||
the entire assembled prompt at about 74 KB per turn, to read two small fields off it. Moving
|
||
those fields into their own columns and marking the heavy ones `deferred` took one adventure
|
||
load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL, so the
|
||
old data never crosses the wire. `tests/test_egress.py` hooks into SQLAlchemy's cursor events
|
||
and fails if a bulk load ever names those columns again.
|
||
- **Turn cost, made flat.** Assembling a turn walked the whole story, so cost scaled with story
|
||
length: 839 KB of reads at turn 200. `backend/app/context/history.py` now serves tails and
|
||
slices from SQL and measures what it fetched. The same turn now costs 129 KB and stops
|
||
growing at around turn 50.
|
||
- **Branching that costs nothing to read.** A branch stores no turns. It stores where it left
|
||
its parent, and borrows everything above that. A 40-turn story forked twenty times loads in
|
||
31,652 bytes against 31,433 bytes for the same story flat: a 1.007x ratio, or about 103 bytes
|
||
per branch. Reads stay cheap because the lineage is windowed the same way the history is, so
|
||
the number of SQL clauses is bounded by the context window rather than by the number of
|
||
forks.
|
||
|
||
## Repo notes
|
||
|
||
- `planning/` is this fork's own package: the product specification, the architecture
|
||
decisions, the milestone plan, the acceptance contract, and a review report for every
|
||
milestone shipped. Start at [`planning/README.md`](planning/README.md).
|
||
- [`planning/archive/`](planning/archive/README.md) holds the Phase 0 research that chose this
|
||
base and the completed milestone reports. It is history, not instruction.
|
||
- [`DEVELOPMENT.md`](DEVELOPMENT.md) is how to set the project up, point it at a model, and run
|
||
the tests. [`PROVENANCE.md`](PROVENANCE.md) records what came from upstream and what changed.
|
||
- `backend/.env.example` lists the two environment variables the backend reads. Everything
|
||
about the model is a runtime setting on the Settings page instead.
|
||
- Upstream's own `plan/` build log and `docs/` project site were removed in the 2026-09-03
|
||
documentation pass: they described AI-DnD's hosted, scripted, multi-user product. Both are
|
||
still in Git history, and in upstream.
|
||
|
||
## License
|
||
|
||
[MIT](LICENSE)
|