Planning: close M3 and record active-head architecture

M3's review recommended planning changes and, following the M2 pattern,
reported rather than applied them. This applies them, and adds the ADR the
review asked for.

ADR 012 records the architecture rather than the requirement. ADR 005
already says that going backward must preserve abandoned history and that
the user sees Undo/Redo/Retry rather than branch management; it names a
movable active head as the direction and stops. What M3 settled is the
shape: the head is stored rather than derived, every read of the story is
capped at it in one place, one mechanism moves it, the state of a position
comes off the node rather than from a replay, the first write below a
moved-back head is the divergence, and whether Redo exists is decided by
the lineage rather than by a flag that could be stale. The last of those
is the property worth keeping — a flag can be wrong and make the story
wrong; a lineage cannot.

Two semantics are ratified in STORY-BRANCH-SEMANTICS.md, both of them
reversals or narrowings that a reader would otherwise take for bugs. Undo
now crosses fork points and continues to the campaign opening, because
refusing at the fork was a consequence of deleting rows the parent line
was also reading, and nothing is deleted any more. And the system refuses
to switch which take is live while a later story is off screen, because
doing it quietly would leave retained history continuing from words the
story no longer says.

A new §14A covers editing in place. §14-15 describe the finished
behaviour — the edit becomes authoritative, the state it implies is
re-evaluated, a new continuation is created, the original is retained —
and that requirement is intact and explicitly not weakened here. It is
also not built, because re-evaluating state from prose a user typed needs
M5's extraction pass. §14A says what exists in the meantime and why
refusing is the minimum that holds the invariant rather than the
destination.

TECHNICAL-DESIGN.md gains §8.7 and §9.1, recording the implemented model
and the bundle behaviour as fact in the way §5.2 records M1 and M2. §10.4
gains a constraint that is easy to lose: the snapshot half of the hybrid
state model is a requirement, not an optimization. Head movement is a row
lookup plus a restore, which is why Undo, Redo and Save Point restore cost
the same at any distance into a campaign; a state model recoverable only
by replaying from the opening would make all three proportional to
campaign length, on exactly the long campaigns this product is for.

DATA-MODEL.md records the head as stored on the campaign rather than
derived from its newest turn — two campaigns holding identical turns can
be read at different places, and nothing about the turns can tell them
apart — and the branch disposition as implemented: the depth a divergent
write left the branch at, deliberately advisory, and carried through
export because every row of an abandoned line is exported either way.

BUILD-MILESTONES.md marks M3 complete and states the one condition still
open. M4 is told a Save Point is a durable pointer and that restoring one
is head movement with a bounds check, not a restore system: a second
mover is the specific failure to avoid, because the two paths would
silently disagree about what restore means. M5 gets three constraints —
keep state efficiently recoverable, move the test instrumentation rather
than the assertions when the world-state protocol goes, and finish the
narrator edit §14A defers.

V1-ACCEPTANCE-TESTS.md clarifies ownership without lowering a bar. D10
keeps all three pass conditions and is explicitly recorded as *not*
satisfied at the end of M3; what changed is that the document now says
which milestone delivers which condition. D03's result is recorded as a
full pass rather than the partial the text allowed for, I07 gains the
pre-M3 bundle clause, and L01 gains the note that resolves its apparent
conflict with A05 — a failed turn does advance the head by one, onto the
player's retained input, and that is A05 working rather than L01 failing.

README.md described a different application: a hosted demo, guest
accounts, cloud providers, Postgres, a Render blueprint, an analytics
dashboard, a QuickJS scripting engine, and 549 tests. M2 removed all of
that and the README was never updated — a gap M2's own debt table missed.
It now describes what this fork is, including the endpoint policy and the
TLS behaviour, and the numbers in it are the current ones.

M3's report is included here as its own evidence record: no separate
baseline report was produced, so it carries the raw counts and runtime
observations as well as the review, and §W records this closeout.

SPECIFICATION.md and SECURITY-THREAT-MODEL.md are unchanged. M3 altered no
product requirement and touched no path in the threat model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u
This commit is contained in:
JesseMarkowitz
2026-09-03 13:54:44 -04:00
co-authored by Claude Opus 5
parent 7f082b61d8
commit c8755c21c2
10 changed files with 2216 additions and 144 deletions
+101 -129
View File
@@ -3,25 +3,27 @@
[![CI](https://github.com/parththakkar106/AI-DnD/actions/workflows/ci.yml/badge.svg)](https://github.com/parththakkar106/AI-DnD/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
An AI Dungeon-style interactive storytelling app that runs entirely on your own machine, with
your own AI model. Create scenarios, play open-ended adventures where an LLM narrates the
world, and extend the engine with **JavaScript scripts compatible with real AI Dungeon
scripting**.
An interactive storytelling app that runs entirely on your own machine, with your own model.
Create scenarios and play open-ended adventures where a local LLM narrates the world, keeps
track of what is true, and remembers what happened.
> ### ▶️ Try it live: **[parththakkar106.github.io/AI-DnD](https://parththakkar106.github.io/AI-DnD/)**
> The project page loads instantly and launches the hosted demo in one tap. Play a scenario as
> a guest: no sign-up and no API key needed. The demo runs on a free tier that sleeps, so the
> first load after it's been idle takes about 30 to 60 seconds to wake up.
This is the **Adventure Storyteller** fork of [AI-DnD](https://github.com/parththakkar106/AI-DnD).
It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model
you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support,
the Postgres path, and the JavaScript scripting engine have all been removed rather than
disabled. What is left is a storyteller you can run offline.
> **Local-only, by design.** The app talks to one place — an Ollama-compatible endpoint on this
> machine or on a machine you control on your own network — and it refuses to be pointed at a
> public address. There is no telemetry, no account, no cloud inference, and nothing is fetched
> at runtime from the Internet.
>
> For the internals, read the **[design notes](https://parththakkar106.github.io/AI-DnD/guide.html)**.
> They walk through the context budgeting, the world-state referee, and the memory bank, and
> state the reasoning behind each one ([Markdown version](docs/GUIDE.md)).
> For the internals, read the **[design notes](docs/GUIDE.md)**. They walk through the context
> budgeting, the world-state referee, and the memory bank, and state the reasoning behind each
> one. Some sections still describe upstream subsystems this fork has removed.
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, running on
SQLite locally and Postgres in the cloud. It works with **any OpenAI-compatible endpoint**:
Ollama and LM Studio locally, or OpenRouter, OpenAI, Groq, or vLLM in the cloud. Endpoint, key,
and model are all runtime settings, and OpenRouter's free-tier models make the whole experience
cost nothing.
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing
everything in one SQLite file.
![The play screen, with the world-state rail open](docs/images/play-world-state.jpg)
@@ -33,7 +35,7 @@ isn't the live one starts a new branch.*
## Features
- **The full play loop.** Do / Say / Story / Continue actions, streamed AI responses (SSE),
retry, undo, and edit. Reasoning models are supported: "thinking" streams into a collapsible
retry, undo, redo, and edit. Reasoning models are supported: "thinking" streams into a collapsible
💭 panel with its own token budget.
- **A branching story tree.** The story is a tree, not a list. Any turn can hold more than one
**take**, and `‹ 2/4 ›` steps between them. Stepping is free: the story below simply empties,
@@ -55,30 +57,34 @@ isn't the live one starts a new branch.*
- **Insights: total prompt transparency.** Every turn stores the exact prompt sent to the
model. Open 🔍 on any AI action to see each context component, its token cost, and why it was
included.
- **JavaScript scripting, AI Dungeon-compatible.** `onInput` / `onModelContext` / `onOutput`
modifiers share `state` and a `worldEntries` API, and run in an embedded quickjs sandbox
(`backend/app/scripting/`). Real AI Dungeon scripts import and run as is. An in-app CodeMirror
editor is included.
- **Auto-summarization and Memory Bank.** The modern AI Dungeon memory system: AI-generated
memories every few actions, a running story summary, and embedding-based retrieval that
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
(`backend/app/memorybank.py`).
- **Undo and retry that actually roll back state.** Undo and retry roll back the world state
and script state to a per-node snapshot, not just the text, and prune the memories that
covered the removed turns. Nothing a retry replaces is discarded: the old attempt stays as
another take of that turn, one keystroke and one click from becoming a branch of its own.
- **Import and export.** AI Dungeon-compatible formats for scripts and scenarios; JSON for
everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree:
every branch, every take, and the fork points, since those were chosen rather than computed.
Files saved in the old single-line format still import.
- **Optional accounts for hosted deployments.** By default the app is single-user with zero
auth friction. Set `AIDND_MULTI_USER=1` and visitors play instantly as guests (signed
session cookie), can register (email and password) at any point to keep their data, and each
user gets isolated data plus their own encrypted-at-rest API key. A server-funded **shared
demo key** with a daily turn cap lets people try it without bringing a key
(`backend/app/auth.py`). Each new guest is also given a copy of a short pre-played
adventure, so the first screen shows real turns and their world-state changes without
spending a demo turn (`backend/app/starter.py`).
- **Undo, Redo, and retry that roll back state and delete nothing.** Undo moves where the story
is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped
over. Both restore the world state from a per-node snapshot rather than just the text, and a
memory derived from a turn now behind the head stops being retrieved without being deleted or
re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the
displaced future stays on the line it was written for, and ordinary Redo stops offering it.
Nothing a retry replaces is discarded either — the old attempt stays as another take of that
turn, one keystroke and one click from becoming a branch of its own.
- **Import and export.** AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree:
every branch, every take, the fork points, which branches the story has left behind, and the
position it is being read at — all of them chosen rather than computed, which is the rule for
what a bundle carries. A campaign exported after two Undos imports still undone, with its
retained future intact, instead of silently reopening at its newest turn. Files that predate
the head position, and files saved in the old single-line format, still import.
- **Single user, no accounts.** There is no sign-up, no login, no session and no API key
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
design, because the only person who can reach it is the person running it. A new install
starts with a short pre-played adventure, so the first screen shows real turns and their
world-state changes rather than an empty page (`backend/app/starter.py`).
- **A refusal you can rely on.** The inference endpoint is checked against an address
allowlist when you save it and again before every request, so a public endpoint is refused
even if the setting is edited in the database directly. TLS verification is never traded
against reachability: a privately issued certificate is verified against your machine's own
trust store, and there is no bypass switch.
## Screenshots
@@ -86,10 +92,10 @@ isn't the live one starts a new branch.*
|---|---|
| ![Insights panel](docs/images/insights.jpg) | ![Scenario editor](docs/images/scenario-editor-npcs.jpg) |
| **Insights**: the exact prompt for the next turn, broken into components with token counts and the trigger word that pulled each story card in. | **Authoring**: stats with ranges, per-turn caps, cooldowns, and word-labeled bands; NPCs the AI addresses by id. |
| ![Script editor](docs/images/script-editor.jpg) | ![Home](docs/images/home.jpg) |
| **Scripting**: the three AI Dungeon hooks with shared persistent `state`, run in a quickjs sandbox. | **Home**: continue a story in progress or start from a scenario. |
| ![The branch map](docs/images/branch-map.jpg) | ![The branches panel](docs/images/branches-panel.jpg) |
| **The tree**: one lane per line, from the moment it left its parent to the moment it ends. The horizontal axis is the story's own clock, so a short branch reads as short. | **Branches**: every line the story has taken, and the three things you can do to one. A line the one you're reading was forked from can't be deleted, and says so. |
| ![Home](docs/images/home.jpg) | ![The branch map](docs/images/branch-map.jpg) |
| **Home**: continue a story in progress or start from a scenario. | **The tree**: one lane per line, from the moment it left its parent to the moment it ends. The horizontal axis is the story's own clock, so a short branch reads as short. |
| ![The branches panel](docs/images/branches-panel.jpg) | |
| **Branches**: every line the story has taken, and the three things you can do to one. A line the one you're reading was forked from can't be deleted, and says so. | |
## Quick start
@@ -121,58 +127,62 @@ Open http://localhost:5173.
## Connect a model
Open **Settings** in the app and point it at any OpenAI-compatible endpoint:
Open **Settings** in the app and point it at a local Ollama-compatible endpoint:
| Provider | Endpoint URL | Notes |
| Where the model runs | Endpoint URL | Notes |
|---|---|---|
| Ollama (local) | `http://localhost:11434/v1` | free, private; also serves embedding models for the Memory Bank (e.g. `nomic-embed-text`) |
| LM Studio (local) | `http://localhost:1234/v1` | free, private |
| OpenRouter | `https://openrouter.ai/api/v1` | `:free` models cost nothing (no embeddings on the free tier) |
| OpenAI / Groq / vLLM / … | provider's `/v1` URL | anything speaking `/v1/chat/completions` |
| Claude Code CLI (local) | `http://127.0.0.1:8787/v1` | your Claude subscription instead of an API key; see [Playing against Claude locally](#playing-against-claude-locally) |
| Ollama, same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works |
| Ollama, a machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | see below |
| LM Studio, same machine | `http://localhost:1234/v1` | |
Model name, API key, generation parameters, and (optionally) summary and embedding models for
the Memory Bank are all configured there too. No config files and no rebuild are needed.
Model name, generation parameters, and (optionally) summary and embedding models for the
Memory Bank are configured there too. No config files and no rebuild are needed. There is no
API key field, because there is nothing to authenticate to.
### Playing against Claude locally
### What the endpoint policy allows
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by the
`claude` command line tool, so you can play the demos against a real model without an
API key. Each request spawns one `claude --print` process, which suits the turn engine:
the app assembles the whole prompt every turn and expects a stateless endpoint.
The address is checked when you save it and again before every request. Only loopback and
private-network addresses are accepted; every public address is refused, by address rather than
by hostname, so a name that resolves outward is refused too. A well-known cloud inference host
is named in the error message only so the refusal says *why*.
Running the model on a second machine you control is supported and expected — that machine
does the inference while the storyteller itself stays bound to loopback on yours. If that
machine serves HTTPS with a certificate from a CA you installed, it works: certificates are
verified against your operating system's trust store as well as the bundled one. Verification
itself is never relaxed, and there is no option to turn it off.
### Playing against a local shim
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787`
backed by a command-line tool, which is useful for testing the turn engine against a stronger
model. Each request spawns one process, which suits the engine: the app assembles the whole
prompt every turn and expects a stateless endpoint.
```sh
cd backend
.venv/Scripts/python.exe tools/claude_shim.py # listens on 127.0.0.1:8787
.venv/bin/python tools/claude_shim.py # listens on 127.0.0.1:8787
```
In Settings, choose the OpenAI-compatible provider, set the base URL to
`http://127.0.0.1:8787/v1`, put any non-empty string in the API key field, and pick
`sonnet`. The shim ignores the key and authenticates as you, through the CLI. Set the
reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field
that Claude 5 models reject.
Set the base URL to `http://127.0.0.1:8787/v1` and pick a model the tool offers. Set the
reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field that
some models reject. Embeddings are not served — leave the embedding model blank, or point the
Memory Bank at an endpoint that serves one.
Embeddings are not served. Leave the embedding model blank, or point the Memory Bank at
a real endpoint.
Run it against a local backend only. The endpoint has no authentication, and anything
reaching it spends your Claude quota. `app/netguard.py` blocks localhost endpoints when
`AIDND_MULTI_USER` is set, so a deployed instance cannot be pointed at it.
The shim has no authentication and spends whatever quota backs it, so run it on loopback and
leave it there.
## How a turn works
```
player input
→ onInput script modifier
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
+ [plot essentials] + [story summary] + [retrieved memories]
+ [triggered story cards] + [history along this branch, token-budgeted]
+ [author's note] + [player action]
→ onModelContext script modifier
→ snapshot context (Insights)
→ provider adapter → AI (streamed)
→ extract + referee the world-state delta block, strip it from the prose
→ onOutput script modifier
→ store & render
```
@@ -180,18 +190,17 @@ player input
```
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ auth, scenarios, adventures, story cards, scripts, chat, settings, analytics, debug
├─ models.py SQLAlchemy: User, Scenario, Adventure, Branch, Action, StoryCard, Script, Settings, Memory
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (64 and counting)
├─ auth.py guest/registered users, sessions, shared demo key
├─ security.py password hashing, cookie signing, API-key encryption
├─ routers/ scenarios, adventures, story cards, chat, settings, debug
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (79 and counting)
├─ endpoints.py the inference-endpoint address policy
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
├─ tree.py forking, promotion, and where a node is placed
├─ head.py the active head: where the story is read, and what moving it costs
├─ attempts.py the takes of one turn, grouped by parent
├─ context/ prompt assembly under a token budget + lineage/history windowing
├─ worldstate/ the stat engine: clamps, cooldowns, bands, milestones
├─ scripting/ quickjs sandbox + AI Dungeon API surface
├─ memorybank.py auto-summarization + embedding retrieval
├─ analytics.py buffered visit counters + the owner's dashboard query
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
├─ providers/ OpenAI-compatible adapter, streaming
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
@@ -202,9 +211,10 @@ development, Vite proxies `/api` to FastAPI.
## Tests
549 backend tests: unit tests plus full HTTP integration through the real quickjs scripting
engine, with the LLM provider mocked. CI runs them on every push, alongside the frontend
lint/build and a Docker image build.
638 backend tests: unit tests plus full HTTP integration through the real turn engine, with
the model provider mocked. They run with no route to the Internet, which is a requirement
rather than a convenience — an offline claim proved on a machine that has been online once
proves nothing.
```sh
cd backend && pip install -r requirements.txt -r requirements-dev.txt
@@ -233,57 +243,19 @@ most interesting engineering in the repo.
the number of SQL clauses is bounded by the context window rather than by the number of
forks.
## Visit analytics
The hosted demo keeps its own analytics: an owner-only dashboard at `/analytics` shows
traffic, which shared scenarios get played, turns and demo-key spend, errors, and a funnel
from *visited* to *played a turn* to *signed up*. It is visible only to the emails listed in
`AIDND_ANALYTICS_EMAILS`, and the route returns 404 for everyone else.
This is built into the app rather than added with a third-party script, for reasons specific
to this project: the CSP allows only `script-src 'self'`, ad blockers block the popular
trackers, and none of those trackers can see the measurement that matters here, a turn. Counts
are aggregated in memory and flushed as UPSERTs, so a visit is a write and never a read, and
every dashboard query is a `GROUP BY` that returns tens of rows regardless of traffic volume.
That matters: see the egress note above for what reading rows per request costs on this stack.
## Deploy (Render)
The repo ships a [`render.yaml`](render.yaml) blueprint: one Docker web service that serves
the SPA and API same-origin, backed by external [Neon](https://neon.tech) Postgres. The free
Render tier has no persistent disk, so the database lives off-box.
1. Create a **Neon** project and copy its pooled connection string.
2. In Render, choose **New → Blueprint** and point it at this repo. Render reads
`render.yaml`.
3. Fill in the secrets it prompts for (`sync: false` vars): `AIDND_DATABASE_URL` (the Neon
string); `AIDND_DEMO_API_KEY` and `AIDND_DEMO_MODELS` to offer a no-signup demo; and
`AIDND_ANALYTICS_EMAILS` (your own account's email) to see the Visitors dashboard.
`AIDND_SECRET_KEY` is generated automatically and stays stable across deploys.
4. Deploy. Pushes to `main` auto-deploy after this. The health check is `/api/health`.
On the free tier the service sleeps after about 15 minutes idle, and the first request after
that takes about 30 to 60 seconds to wake it. Point any keep-warm pinger at `/api/health`,
which deliberately doesn't touch the database: waking the database around the clock costs far
more than the cold start saves.
If you put another proxy or CDN in front of Render, set `AIDND_TRUSTED_PROXY_HOPS` to the
number of proxies in the chain. It defaults to 1. The rate limiter reads the client IP that
many entries from the right of `X-Forwarded-For`, because the trusted edge appends the real
one last. Leave it at 1 behind two proxies and the limiter reads an entry the caller
supplied, so anyone can rotate the header for a fresh rate-limit bucket per request.
## Repo notes
- `plan/` holds the phased implementation plan this project was built from, kept as a build
log. All fourteen phases are complete. The later files (11, 12, 14) also serve as design
notes for the state-revert, world-state, and story-tree work.
[`plan/STATUS.md`](plan/STATUS.md) is the running thread: what shipped, what was measured,
and what is owed next.
- [`docs/GUIDE.md`](docs/GUIDE.md) holds design notes: how each subsystem works and why it was
built that way, with the measurements behind the decisions. It is also rendered as a
[reading page](https://parththakkar106.github.io/AI-DnD/guide.html).
- `backend/.env.example` lists the few environment variables the backend reads.
- `planning/` is this fork's own package: the product specification, the architecture
decisions, the milestone plan, the acceptance contract, and a review report for every
milestone shipped. Start at [`planning/README.md`](planning/README.md).
- `plan/` holds the *upstream* project's phased implementation plan, kept as a build log. The
later files (11, 12, 14) still serve as design notes for the state-revert, world-state, and
story-tree work this fork inherited.
- [`docs/GUIDE.md`](docs/GUIDE.md) holds upstream's design notes: how each subsystem works and
why it was built that way, with the measurements behind the decisions. Sections covering
scripting, accounts and hosted deployment describe subsystems this fork removed.
- `backend/.env.example` lists the two environment variables the backend reads. Everything
about the model is a runtime setting on the Settings page instead.
- [`docs/self-review.md`](docs/self-review.md) records a full-codebase self-review pass and
what came out of it. All correctness findings are resolved.