M3's review recommended planning changes and, following the M2 pattern, reported rather than applied them. This applies them, and adds the ADR the review asked for. ADR 012 records the architecture rather than the requirement. ADR 005 already says that going backward must preserve abandoned history and that the user sees Undo/Redo/Retry rather than branch management; it names a movable active head as the direction and stops. What M3 settled is the shape: the head is stored rather than derived, every read of the story is capped at it in one place, one mechanism moves it, the state of a position comes off the node rather than from a replay, the first write below a moved-back head is the divergence, and whether Redo exists is decided by the lineage rather than by a flag that could be stale. The last of those is the property worth keeping — a flag can be wrong and make the story wrong; a lineage cannot. Two semantics are ratified in STORY-BRANCH-SEMANTICS.md, both of them reversals or narrowings that a reader would otherwise take for bugs. Undo now crosses fork points and continues to the campaign opening, because refusing at the fork was a consequence of deleting rows the parent line was also reading, and nothing is deleted any more. And the system refuses to switch which take is live while a later story is off screen, because doing it quietly would leave retained history continuing from words the story no longer says. A new §14A covers editing in place. §14-15 describe the finished behaviour — the edit becomes authoritative, the state it implies is re-evaluated, a new continuation is created, the original is retained — and that requirement is intact and explicitly not weakened here. It is also not built, because re-evaluating state from prose a user typed needs M5's extraction pass. §14A says what exists in the meantime and why refusing is the minimum that holds the invariant rather than the destination. TECHNICAL-DESIGN.md gains §8.7 and §9.1, recording the implemented model and the bundle behaviour as fact in the way §5.2 records M1 and M2. §10.4 gains a constraint that is easy to lose: the snapshot half of the hybrid state model is a requirement, not an optimization. Head movement is a row lookup plus a restore, which is why Undo, Redo and Save Point restore cost the same at any distance into a campaign; a state model recoverable only by replaying from the opening would make all three proportional to campaign length, on exactly the long campaigns this product is for. DATA-MODEL.md records the head as stored on the campaign rather than derived from its newest turn — two campaigns holding identical turns can be read at different places, and nothing about the turns can tell them apart — and the branch disposition as implemented: the depth a divergent write left the branch at, deliberately advisory, and carried through export because every row of an abandoned line is exported either way. BUILD-MILESTONES.md marks M3 complete and states the one condition still open. M4 is told a Save Point is a durable pointer and that restoring one is head movement with a bounds check, not a restore system: a second mover is the specific failure to avoid, because the two paths would silently disagree about what restore means. M5 gets three constraints — keep state efficiently recoverable, move the test instrumentation rather than the assertions when the world-state protocol goes, and finish the narrator edit §14A defers. V1-ACCEPTANCE-TESTS.md clarifies ownership without lowering a bar. D10 keeps all three pass conditions and is explicitly recorded as *not* satisfied at the end of M3; what changed is that the document now says which milestone delivers which condition. D03's result is recorded as a full pass rather than the partial the text allowed for, I07 gains the pre-M3 bundle clause, and L01 gains the note that resolves its apparent conflict with A05 — a failed turn does advance the head by one, onto the player's retained input, and that is A05 working rather than L01 failing. README.md described a different application: a hosted demo, guest accounts, cloud providers, Postgres, a Render blueprint, an analytics dashboard, a QuickJS scripting engine, and 549 tests. M2 removed all of that and the README was never updated — a gap M2's own debt table missed. It now describes what this fork is, including the endpoint policy and the TLS behaviour, and the numbers in it are the current ones. M3's report is included here as its own evidence record: no separate baseline report was produced, so it carries the raw counts and runtime observations as well as the review, and §W records this closeout. SPECIFICATION.md and SECURITY-THREAT-MODEL.md are unchanged. M3 altered no product requirement and touched no path in the threat model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u
265 lines
15 KiB
Markdown
265 lines
15 KiB
Markdown
# AI D&D
|
||
|
||
[](https://github.com/parththakkar106/AI-DnD/actions/workflows/ci.yml)
|
||
[](LICENSE)
|
||
|
||
An interactive storytelling app that runs entirely on your own machine, with your own model.
|
||
Create scenarios and play open-ended adventures where a local LLM narrates the world, keeps
|
||
track of what is true, and remembers what happened.
|
||
|
||
This is the **Adventure Storyteller** fork of [AI-DnD](https://github.com/parththakkar106/AI-DnD).
|
||
It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model
|
||
you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support,
|
||
the Postgres path, and the JavaScript scripting engine have all been removed rather than
|
||
disabled. What is left is a storyteller you can run offline.
|
||
|
||
> **Local-only, by design.** The app talks to one place — an Ollama-compatible endpoint on this
|
||
> machine or on a machine you control on your own network — and it refuses to be pointed at a
|
||
> public address. There is no telemetry, no account, no cloud inference, and nothing is fetched
|
||
> at runtime from the Internet.
|
||
>
|
||
> For the internals, read the **[design notes](docs/GUIDE.md)**. They walk through the context
|
||
> budgeting, the world-state referee, and the memory bank, and state the reasoning behind each
|
||
> one. Some sections still describe upstream subsystems this fork has removed.
|
||
|
||
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing
|
||
everything in one SQLite file.
|
||
|
||

|
||
|
||
*The play screen. The left rail shows live world state. The AI proposes changes each turn, and
|
||
a Python engine decides what actually sticks. The chip under the narration reports what
|
||
changed. The `‹ 2/2 ›` under a turn steps between the takes it has. Writing below a take that
|
||
isn't the live one starts a new branch.*
|
||
|
||
## Features
|
||
|
||
- **The full play loop.** Do / Say / Story / Continue actions, streamed AI responses (SSE),
|
||
retry, undo, redo, and edit. Reasoning models are supported: "thinking" streams into a collapsible
|
||
💭 panel with its own token budget.
|
||
- **A branching story tree.** The story is a tree, not a list. Any turn can hold more than one
|
||
**take**, and `‹ 2/4 ›` steps between them. Stepping is free: the story below simply empties,
|
||
and the server is told nothing. Writing below a take that isn't the live one is what makes a
|
||
branch. Branches borrow their ancestors' turns instead of copying them, so a fork costs about
|
||
100 bytes, and a 20-fork story loads within 1% of the same story flat. Switching restores that
|
||
line's world state, script state, and cooldown clocks. A branch panel switches, renames, and
|
||
deletes; **⌗ See the tree** draws every line against the story's own clock
|
||
(`backend/app/tree.py`, `backend/app/context/lineage.py`).
|
||
- **An RPG world-state engine.** A scenario can declare stats, flags, milestones, and a named
|
||
cast; the adventure carries their live values. The AI proposes deltas, and a Python engine
|
||
referees them: it clamps values to range, enforces per-turn caps and cooldowns, keeps counters
|
||
monotonic and milestones sticky, then strips the machine-readable block out of the prose
|
||
(`backend/app/worldstate/engine.py`). Word-labeled bands (`40–60: minor damage`) make the
|
||
model reliable at it. No dice and no scripting are required.
|
||
- **AI Dungeon-compatible context engine.** Memory, author's note, and story cards (world
|
||
info) are triggered by keywords in recent story text, then assembled under a token budget
|
||
(`backend/app/context/builder.py`).
|
||
- **Insights: total prompt transparency.** Every turn stores the exact prompt sent to the
|
||
model. Open 🔍 on any AI action to see each context component, its token cost, and why it was
|
||
included.
|
||
- **Auto-summarization and Memory Bank.** The modern AI Dungeon memory system: AI-generated
|
||
memories every few actions, a running story summary, and embedding-based retrieval that
|
||
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
|
||
(`backend/app/memorybank.py`).
|
||
- **Undo, Redo, and retry that roll back state and delete nothing.** Undo moves where the story
|
||
is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped
|
||
over. Both restore the world state from a per-node snapshot rather than just the text, and a
|
||
memory derived from a turn now behind the head stops being retrieved without being deleted or
|
||
re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the
|
||
displaced future stays on the line it was written for, and ordinary Redo stops offering it.
|
||
Nothing a retry replaces is discarded either — the old attempt stays as another take of that
|
||
turn, one keystroke and one click from becoming a branch of its own.
|
||
- **Import and export.** AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree:
|
||
every branch, every take, the fork points, which branches the story has left behind, and the
|
||
position it is being read at — all of them chosen rather than computed, which is the rule for
|
||
what a bundle carries. A campaign exported after two Undos imports still undone, with its
|
||
retained future intact, instead of silently reopening at its newest turn. Files that predate
|
||
the head position, and files saved in the old single-line format, still import.
|
||
- **Single user, no accounts.** There is no sign-up, no login, no session and no API key
|
||
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
|
||
design, because the only person who can reach it is the person running it. A new install
|
||
starts with a short pre-played adventure, so the first screen shows real turns and their
|
||
world-state changes rather than an empty page (`backend/app/starter.py`).
|
||
- **A refusal you can rely on.** The inference endpoint is checked against an address
|
||
allowlist when you save it and again before every request, so a public endpoint is refused
|
||
even if the setting is edited in the database directly. TLS verification is never traded
|
||
against reachability: a privately issued certificate is verified against your machine's own
|
||
trust store, and there is no bypass switch.
|
||
|
||
## Screenshots
|
||
|
||
| | |
|
||
|---|---|
|
||
|  |  |
|
||
| **Insights**: the exact prompt for the next turn, broken into components with token counts and the trigger word that pulled each story card in. | **Authoring**: stats with ranges, per-turn caps, cooldowns, and word-labeled bands; NPCs the AI addresses by id. |
|
||
|  |  |
|
||
| **Home**: continue a story in progress or start from a scenario. | **The tree**: one lane per line, from the moment it left its parent to the moment it ends. The horizontal axis is the story's own clock, so a short branch reads as short. |
|
||
|  | |
|
||
| **Branches**: every line the story has taken, and the three things you can do to one. A line the one you're reading was forked from can't be deleted, and says so. | |
|
||
|
||
## Quick start
|
||
|
||
### Docker (any OS)
|
||
|
||
```sh
|
||
docker compose up --build
|
||
```
|
||
|
||
Open http://localhost:8000. Your data persists in a named volume across restarts.
|
||
|
||
### Windows
|
||
|
||
```powershell
|
||
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
|
||
cd frontend; npm install; cd ..
|
||
.\start.ps1
|
||
```
|
||
|
||
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
|
||
|
||
### macOS / Linux
|
||
|
||
```sh
|
||
./start.sh # creates the venv and installs dependencies on first run
|
||
```
|
||
|
||
Open http://localhost:5173.
|
||
|
||
## Connect a model
|
||
|
||
Open **Settings** in the app and point it at a local Ollama-compatible endpoint:
|
||
|
||
| Where the model runs | Endpoint URL | Notes |
|
||
|---|---|---|
|
||
| Ollama, same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works |
|
||
| Ollama, a machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | see below |
|
||
| LM Studio, same machine | `http://localhost:1234/v1` | |
|
||
|
||
Model name, generation parameters, and (optionally) summary and embedding models for the
|
||
Memory Bank are configured there too. No config files and no rebuild are needed. There is no
|
||
API key field, because there is nothing to authenticate to.
|
||
|
||
### What the endpoint policy allows
|
||
|
||
The address is checked when you save it and again before every request. Only loopback and
|
||
private-network addresses are accepted; every public address is refused, by address rather than
|
||
by hostname, so a name that resolves outward is refused too. A well-known cloud inference host
|
||
is named in the error message only so the refusal says *why*.
|
||
|
||
Running the model on a second machine you control is supported and expected — that machine
|
||
does the inference while the storyteller itself stays bound to loopback on yours. If that
|
||
machine serves HTTPS with a certificate from a CA you installed, it works: certificates are
|
||
verified against your operating system's trust store as well as the bundled one. Verification
|
||
itself is never relaxed, and there is no option to turn it off.
|
||
|
||
### Playing against a local shim
|
||
|
||
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787`
|
||
backed by a command-line tool, which is useful for testing the turn engine against a stronger
|
||
model. Each request spawns one process, which suits the engine: the app assembles the whole
|
||
prompt every turn and expects a stateless endpoint.
|
||
|
||
```sh
|
||
cd backend
|
||
.venv/bin/python tools/claude_shim.py # listens on 127.0.0.1:8787
|
||
```
|
||
|
||
Set the base URL to `http://127.0.0.1:8787/v1` and pick a model the tool offers. Set the
|
||
reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field that
|
||
some models reject. Embeddings are not served — leave the embedding model blank, or point the
|
||
Memory Bank at an endpoint that serves one.
|
||
|
||
The shim has no authentication and spends whatever quota backs it, so run it on loopback and
|
||
leave it there.
|
||
|
||
## How a turn works
|
||
|
||
```
|
||
player input
|
||
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
|
||
+ [plot essentials] + [story summary] + [retrieved memories]
|
||
+ [triggered story cards] + [history along this branch, token-budgeted]
|
||
+ [author's note] + [player action]
|
||
→ snapshot context (Insights)
|
||
→ provider adapter → AI (streamed)
|
||
→ extract + referee the world-state delta block, strip it from the prose
|
||
→ store & render
|
||
```
|
||
|
||
## Architecture
|
||
|
||
```
|
||
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||
├─ routers/ scenarios, adventures, story cards, chat, settings, debug
|
||
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory
|
||
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (79 and counting)
|
||
├─ endpoints.py the inference-endpoint address policy
|
||
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
|
||
├─ tree.py forking, promotion, and where a node is placed
|
||
├─ head.py the active head: where the story is read, and what moving it costs
|
||
├─ attempts.py the takes of one turn, grouped by parent
|
||
├─ context/ prompt assembly under a token budget + lineage/history windowing
|
||
├─ worldstate/ the stat engine: clamps, cooldowns, bands, milestones
|
||
├─ memorybank.py auto-summarization + embedding retrieval
|
||
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
|
||
├─ providers/ OpenAI-compatible adapter, streaming
|
||
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
|
||
```
|
||
|
||
In production the backend serves the built SPA from one port (see `Dockerfile`). In
|
||
development, Vite proxies `/api` to FastAPI.
|
||
|
||
## Tests
|
||
|
||
638 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||
the model provider mocked. They run with no route to the Internet, which is a requirement
|
||
rather than a convenience — an offline claim proved on a machine that has been online once
|
||
proves nothing.
|
||
|
||
```sh
|
||
cd backend && pip install -r requirements.txt -r requirements-dev.txt
|
||
python -m pytest tests/
|
||
```
|
||
|
||
## Notes on performance
|
||
|
||
Two of the tests exist because of bugs that were measured rather than guessed at. They are the
|
||
most interesting engineering in the repo.
|
||
|
||
- **Database egress, cut about 189x.** Every adventure load pulled `Action.context_snapshot`,
|
||
the entire assembled prompt at about 74 KB per turn, to read two small fields off it. Moving
|
||
those fields into their own columns and marking the heavy ones `deferred` took one adventure
|
||
load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL, so the
|
||
old data never crosses the wire. `tests/test_egress.py` hooks into SQLAlchemy's cursor events
|
||
and fails if a bulk load ever names those columns again.
|
||
- **Turn cost, made flat.** Assembling a turn walked the whole story, so cost scaled with story
|
||
length: 839 KB of reads at turn 200. `backend/app/context/history.py` now serves tails and
|
||
slices from SQL and measures what it fetched. The same turn now costs 129 KB and stops
|
||
growing at around turn 50.
|
||
- **Branching that costs nothing to read.** A branch stores no turns. It stores where it left
|
||
its parent, and borrows everything above that. A 40-turn story forked twenty times loads in
|
||
31,652 bytes against 31,433 bytes for the same story flat: a 1.007x ratio, or about 103 bytes
|
||
per branch. Reads stay cheap because the lineage is windowed the same way the history is, so
|
||
the number of SQL clauses is bounded by the context window rather than by the number of
|
||
forks.
|
||
|
||
## Repo notes
|
||
|
||
- `planning/` is this fork's own package: the product specification, the architecture
|
||
decisions, the milestone plan, the acceptance contract, and a review report for every
|
||
milestone shipped. Start at [`planning/README.md`](planning/README.md).
|
||
- `plan/` holds the *upstream* project's phased implementation plan, kept as a build log. The
|
||
later files (11, 12, 14) still serve as design notes for the state-revert, world-state, and
|
||
story-tree work this fork inherited.
|
||
- [`docs/GUIDE.md`](docs/GUIDE.md) holds upstream's design notes: how each subsystem works and
|
||
why it was built that way, with the measurements behind the decisions. Sections covering
|
||
scripting, accounts and hosted deployment describe subsystems this fork removed.
|
||
- `backend/.env.example` lists the two environment variables the backend reads. Everything
|
||
about the model is a runtime setting on the Settings page instead.
|
||
- [`docs/self-review.md`](docs/self-review.md) records a full-codebase self-review pass and
|
||
what came out of it. All correctness findings are resolved.
|
||
|
||
## License
|
||
|
||
[MIT](LICENSE)
|