A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
20 KiB
Adventure Storyteller
An interactive storytelling app that runs entirely on your own machine, with your own model. Create scenarios and play open-ended adventures where a local LLM narrates the world, keeps track of what is true, and remembers what happened.
This is the Adventure Storyteller fork of AI-DnD. It is deliberately narrower than its upstream: single-user, local-only, and pointed at a model you run yourself. The hosted deployment, the accounts and sessions, the cloud provider support, the Postgres path, and the JavaScript scripting engine have all been removed rather than disabled. What is left is a storyteller you can run offline.
Local-only, by design. The app talks to one place — an Ollama-compatible endpoint on this machine or on a machine you control on your own network — and it refuses to be pointed at a public address. There is no telemetry, no account, no cloud inference, and nothing is fetched at runtime from the Internet.
For the internals, read
planning/TECHNICAL-DESIGN.mdandplanning/CONTEXT-AND-MEMORY.md, which cover the context budgeting, the state model and the memory bank as this fork builds them.
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing everything in one SQLite file.
On the play screen, the left rail carries live world state. The AI proposes changes each turn
and a Python engine decides what actually sticks; the chip under the narration reports what
changed. The ‹ 2/2 › under a turn steps between the takes it has, and writing below a take
that isn't the live one starts a new branch.
Features
- The full play loop. Do / Say / Story / Continue actions, streamed AI responses (SSE), retry, undo, redo, and edit. Correcting narrator prose does not overwrite it: the correction becomes a new continuation carrying the state it implies, and the original narration keeps its own future as retained history. Reasoning models are supported: "thinking" streams into a collapsible 💭 panel with its own token budget.
- A branching story tree. The story is a tree, not a list. Any turn can hold more than one
take, and
‹ 2/4 ›steps between them. Stepping is free: the story below simply empties, and the server is told nothing. Writing below a take that isn't the live one is what makes a branch. Branches borrow their ancestors' turns instead of copying them, so a fork costs about 100 bytes, and a 20-fork story loads within 1% of the same story flat. Switching restores that line's world state, script state, and cooldown clocks. A branch panel switches, renames, and deletes; ⌗ See the tree draws every line against the story's own clock (backend/app/tree.py,backend/app/context/lineage.py). - Authoritative narrative state, and the application owns it. The story tracks who exists,
where they are, what they hold, what is true, how they are tied to each other, and what is
still open — as generic entities, facts, relationships and threads, with no genre baked in.
The same schema holds a silver key in an abbey and a data crystal on an orbital station.
The AI proposes typed events with absolute values (
set_possession,add_fact,set_current_location…), and a Python validator decides what is accepted: unknown event types are refused, references must resolve, campaign canon outranks the narration, and the machine-readable block never reaches the reader (backend/app/narrative/). Every accepted change is recorded with what it was before and which turn caused it, so the Story State panel can show what changed and why. You can correct it by hand, and your correction outranks the story. - AI Dungeon-compatible context engine. Memory, author's note, and story cards (world
info) are triggered by keywords in recent story text, then assembled under a token budget
(
backend/app/context/builder.py). Story cards are the inherited authored-lore primitive and are kept; they are not the knowledge library below, which is a first-class subsystem with its own classification, provenance, chunking and index. - Insights: total prompt transparency. Every turn stores the exact prompt sent to the model. Open 🔍 on any AI action to see each context component, its token cost, and why it was included.
- Auto-summarization and Memory Bank. The modern AI Dungeon memory system: AI-generated
memories every few actions, a running story summary, and embedding-based retrieval that
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
(
backend/app/memorybank.py). - An imported knowledge library, classified by how much authority it has. Import your own
local
.txtand.mdfiles — a setting bible, character notes, research, a passage whose voice you want the prose to have — as Canon, Reference or Inspiration. The class is not a label: it decides the words the passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. Canon can establish what is true; Reference informs detail without establishing anything; Inspiration influences tone and introduces no facts at all. Retrieval is hybrid and local: a SQLite FTS5 index finds the names and invented terms an embedding is worst at, local Ollama embeddings find what you meant when your words differ from the file's, and the two are merged, de-duplicated and reranked by relevance × class. Lexical search is a supported production path, not a fallback — the library works with no embedding model at all. Canon you mark always include is supplied on every turn whether or not the scene resembles it, and Canon you mark narrator only is given to the narrator with instructions not to let the protagonist know it. Every passage that reaches a prompt is listed in Insights with its file, class, heading, passage number, scores and token cost, and that record is kept in the turn, so deleting a source never erases the evidence of what an old turn was shown (backend/app/knowledge/). - Imported text is data, never instruction. Every imported passage is delimited in the prompt as untrusted data with the authority order stated in words, so "ignore all previous instructions" inside a file is a sentence in a file. Nothing is fetched: a URL in a source is text, a remote Markdown image never loads, and no endpoint anywhere takes a filesystem path — a source arrives as an upload, so there is no path for a traversal to escape from. Imported content is displayed as inert text and never rendered as HTML.
- Undo, Redo, and retry that roll back state and delete nothing. Undo moves where the story is being read; it removes no accepted turn, so Redo can walk forward into the turns it stepped over. Both restore the world state from a per-node snapshot rather than just the text, and a memory derived from a turn now behind the head stops being retrieved without being deleted or re-embedded. Writing a new turn below a moved-back head is the moment the story forks: the displaced future stays on the line it was written for, and ordinary Redo stops offering it. Nothing a retry replaces is discarded either — the old attempt stays as another take of that turn, one keystroke and one click from becoming a branch of its own.
- Save Points. Name a moment — "Before entering the abbey" — keep playing,
restart the app, and come back to it. Restoring one moves the story back to
that moment and deletes nothing: the turns you wrote after it stay, Redo still
walks forward into them, and writing something different from the Save Point
is what starts a new line while the old one is kept. A Save Point is a name for
a position and holds no copy of the story, so restoring it is the same
movement Undo makes (
backend/app/routers/adventures/checkpoints.py,backend/app/head.py). They last until you delete them: deleting one deletes no story, and deleting a branch a Save Point is kept on is refused until you remove the Save Point yourself, so nothing takes a named moment away behind your back. - Import and export. AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as
ai-dnd-adventure-v2, which carries the whole tree: every branch, every take, the fork points, which branches the story has left behind, the Save Points and the position it is being read at — all of them chosen rather than computed, which is the rule for what a bundle carries. A campaign opens where its head says, never at a Save Point merely because it has one. A campaign exported after two Undos imports still undone, with its retained future intact, instead of silently reopening at its newest turn. Files that predate the head position, and files saved in the old single-line format, still import. - Single user, no accounts. There is no sign-up, no login, no session and no API key
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
design, because the only person who can reach it is the person running it. A new install
starts with a short pre-played adventure, so the first screen shows real turns and their
world-state changes rather than an empty page (
backend/app/starter.py). - A refusal you can rely on. The inference endpoint is checked against an address allowlist when you save it and again before every request, so a public endpoint is refused even if the setting is edited in the database directly. TLS verification is never traded against reachability: a privately issued certificate is verified against your machine's own trust store, and there is no bypass switch.
Screenshots
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up, a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were removed rather than left standing as a picture of a product that no longer exists. The M4 closeout drove the real application in a real browser, so the screens exist and work; taking presentable screenshots of them is a job for the UI pass in M8.
Quick start
Docker (any OS)
docker compose up --build
Open http://localhost:8000. Your data persists in a named volume across restarts.
Windows
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
cd frontend; npm install; cd ..
.\start.ps1
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
macOS / Linux
./start.sh # creates the venv and installs dependencies on first run
Open http://localhost:5173.
Connect a model
Ollama is the inference backend v1 supports. Open Settings in the app and point it at one:
| Where Ollama runs | Endpoint URL | Notes |
|---|---|---|
| The same machine | http://localhost:11434/v1 |
the default, and the simplest thing that works |
| A machine on your own network | http://<host>:11434/v1 or https://<host>/v1 |
explicitly configured; see below |
Model name, generation parameters, and (optionally) summary and embedding models for the Memory Bank are configured there too. No config files and no rebuild are needed. There is no API key field, because there is nothing to authenticate to.
The adapter underneath speaks the OpenAI-compatible protocol, because that is what Ollama
serves. That is an implementation detail, not a promise of support for arbitrary local
servers that happen to speak the same protocol. Public and cloud inference endpoints are
prohibited outright — see planning/DECISIONS/002-ollama-only-v1.md and
planning/DECISIONS/011-local-inference-endpoint-policy.md.
What the endpoint policy allows
The address is checked when you save it and again before every request. Only loopback and private-network addresses are accepted; every public address is refused, by address rather than by hostname, so a name that resolves outward is refused too. A well-known cloud inference host is named in the error message only so the refusal says why.
Running the model on a second machine you control is supported and expected — that machine does the inference while the storyteller itself stays bound to loopback on yours. If that machine serves HTTPS with a certificate from a CA you installed, it works: certificates are verified against your operating system's trust store as well as the bundled one. Verification itself is never relaxed, and there is no option to turn it off.
Playing against a local shim (development only)
backend/tools/claude_shim.py serves an OpenAI-compatible endpoint on 127.0.0.1:8787
backed by a command-line tool, which is useful for testing the turn engine against a stronger
model. Each request spawns one process, which suits the engine: the app assembles the whole
prompt every turn and expects a stateless endpoint.
cd backend
.venv/bin/python tools/claude_shim.py # listens on 127.0.0.1:8787
Set the base URL to http://127.0.0.1:8787/v1 and pick a model the tool offers. Set the
reasoning budget to 0 or -1: a positive budget sends a reasoning.max_tokens field that
some models reject. Embeddings are not served — leave the embedding model blank, or point the
Memory Bank at an endpoint that serves one.
The shim has no authentication and spends whatever quota backs it, so run it on loopback and leave it there.
How a turn works
player input
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
+ [plot essentials] + [story summary] + [retrieved memories]
+ [triggered story cards] + [retrieved imported knowledge,
framed by class and bounded by its own budget]
+ [history along this branch, token-budgeted]
+ [author's note] + [player action]
→ snapshot context (Insights)
→ provider adapter → AI (streamed)
→ extract + referee the world-state delta block, strip it from the prose
→ store & render
Architecture
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting)
├─ endpoints.py the inference-endpoint address policy
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
├─ tree.py forking, promotion, and where a node is placed
├─ head.py the active head: where the story is read, and what moving it costs
├─ checkpoints Save Points: durable names for positions, in routers/adventures/
├─ attempts.py the takes of one turn, grouped by parent
├─ context/ prompt assembly under a token budget + lineage/history windowing
├─ narrative/ the authoritative state: typed events, validation, snapshots
├─ worldstate/ the inherited RPG stat engine — legacy, no longer authoritative
├─ memorybank.py auto-summarization + embedding retrieval
├─ knowledge/ the imported library: import, chunk, FTS5, embed, rank, inject
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
├─ providers/ OpenAI-compatible adapter, streaming
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
In production the backend serves the built SPA from one port (see Dockerfile). In
development, Vite proxies /api to FastAPI.
Tests
920 backend tests: unit tests plus full HTTP integration through the real turn engine, with the model provider mocked. They run with no route to the Internet, which is a requirement rather than a convenience — an offline claim proved on a machine that has been online once proves nothing. A further handful need a real local model and skip without one; they exist because a mocked provider can leave the production wiring dead while the suite stays green, which this project has shipped twice.
cd backend && pip install -r requirements.txt -r requirements-dev.txt
python -m pytest tests/
Notes on performance
Two of the tests exist because of bugs that were measured rather than guessed at. They are the most interesting engineering in the repo.
- Database egress, cut about 189x. Every adventure load pulled
Action.context_snapshot, the entire assembled prompt at about 74 KB per turn, to read two small fields off it. Moving those fields into their own columns and marking the heavy onesdeferredtook one adventure load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL, so the old data never crosses the wire.tests/test_egress.pyhooks into SQLAlchemy's cursor events and fails if a bulk load ever names those columns again. - Turn cost, made flat. Assembling a turn walked the whole story, so cost scaled with story
length: 839 KB of reads at turn 200.
backend/app/context/history.pynow serves tails and slices from SQL and measures what it fetched. The same turn now costs 129 KB and stops growing at around turn 50. - Branching that costs nothing to read. A branch stores no turns. It stores where it left its parent, and borrows everything above that. A 40-turn story forked twenty times loads in 31,652 bytes against 31,433 bytes for the same story flat: a 1.007x ratio, or about 103 bytes per branch. Reads stay cheap because the lineage is windowed the same way the history is, so the number of SQL clauses is bounded by the context window rather than by the number of forks.
Repo notes
planning/is this fork's own package: the product specification, the architecture decisions, the milestone plan, the acceptance contract, and a review report for every milestone shipped. Start atplanning/README.md.planning/archive/holds the Phase 0 research that chose this base and the completed milestone reports. It is history, not instruction.DEVELOPMENT.mdis how to set the project up, point it at a model, and run the tests.PROVENANCE.mdrecords what came from upstream and what changed.backend/.env.examplelists the two environment variables the backend reads. Everything about the model is a runtime setting on the Settings page instead.- Upstream's own
plan/build log anddocs/project site were removed in the 2026-09-03 documentation pass: they described AI-DnD's hosted, scripted, multi-user product. Both are still in Git history, and in upstream.