A hand-written memory used to carry a NULL depth, described in the model as "belongs to the adventure rather than to a path". That sounds harmless and is not: a NULL is a coordinate no fork can cap, so a note typed on one line followed the reader onto branches whose story it never described. It takes the head now — the story you were reading when you wrote it — and obeys exactly the rule a summarised memory obeys. The unanchored escape clause in lineage.Path.clause existed for that single case and is deleted rather than left unused. Its docstring argued that a capped depth would drop a typed memory the moment its branch stopped being the newest entry; anchoring answers the same worry better, because the memory is not exempt from the path, it is on one. The drawer now shows the path being read and nothing else, filtered by the clause retrieval itself uses, so the bank you can see is the bank the model can see. Nothing is stranded: a memory lives on a branch, switching to that branch shows it, and deleting the branch deletes it. Pinning decides order, the path decides existence. Migration 62 lands existing NULL-depth memories at depth 0 of their branch rather than at the tip. 0 is at or before every fork point, so every memory stays visible from exactly the paths it is visible from today — nobody's bank loses a row on deploy. The tip is the tidier-sounding choice and would have emptied them out of every branch forked earlier than they were typed. This supersedes the on_path flag and the "another branch" badge from earlier today; anchoring makes them redundant, and they are removed. Four tests changed because they asserted the old contract, not because they broke. The one worth reading is the pair replacing test_a_hand_written_memory_is_not_lost_at_the_first_fork: typed on shared trunk it still survives a fork, and typed on ground the fork never travelled it no longer follows you. 402 tests. Verified on tools/branch_fixture.py: each branch's drawer holds its own memory and not the other's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
AI D&D
An AI Dungeon-style interactive storytelling app you can run entirely on your own machine — with your own AI model. Create scenarios, play open-ended adventures where an LLM narrates the world, and extend the engine with JavaScript scripts compatible with real AI Dungeon scripting.
▶️ Try it live: parththakkar106.github.io/AI-DnD
The project page loads instantly and launches the hosted demo in one tap — play a scenario as a guest, no sign-up and no API key needed. (The demo runs on a free tier that sleeps, so the first load after it's been idle takes ~30–60s to wake up.)
Want the internals? The design notes walk through the context budgeting, the world-state referee and the memory bank, and state the reasoning behind each one (Markdown version).
Built with FastAPI + SQLAlchemy on the backend and React (Vite) on the frontend, running on SQLite locally and Postgres in the cloud. Works with any OpenAI-compatible endpoint: Ollama and LM Studio locally, or OpenRouter / OpenAI / Groq / vLLM in the cloud — endpoint, key, and model are all runtime settings, and OpenRouter's free-tier models make the whole experience $0.
The play screen. The left rail is live world state — the AI proposes changes each turn and a Python engine decides what actually sticks. The chip under the narration reports what changed.
Features
- The full play loop — Do / Say / Story / Continue actions, streamed AI responses (SSE), retry, undo, and edit. Reasoning models supported: "thinking" streams into a collapsible 💭 panel with its own token budget.
- An RPG world-state engine — a scenario can declare stats, flags, milestones and a named
cast; the adventure carries their live values. The design is the AI proposes deltas and a
Python engine referees them: it clamps to range, enforces per-turn caps and cooldowns, keeps
counters monotonic and milestones sticky, then strips the machine-readable block out of the
prose (
backend/app/worldstate/engine.py). Word-labelled bands (40–60: minor damage) are what make the model reliable at it. No dice, no scripting required. - AI Dungeon-compatible context engine — memory, author's note, and story cards (world
info) triggered by keywords in recent story text, assembled under a token budget
(
backend/app/context/builder.py). - Insights: total prompt transparency — every turn stores the exact prompt sent to the model; open 🔍 on any AI action to see each context component, its token cost, and why it was included.
- JavaScript scripting, AI Dungeon-compatible —
onInput/onModelContext/onOutputmodifiers with sharedstateand aworldEntriesAPI, executed in an embedded quickjs sandbox (backend/app/scripting/). Real AI Dungeon scripts import and run. In-app CodeMirror editor included. - Auto-summarization + Memory Bank — the modern AI Dungeon memory system: AI-generated
memories every few actions, a running story summary, and embedding-based retrieval that
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
(
backend/app/memorybank.py). - Undo and retry that actually rewind — undo and retry roll back the world state and script
state to a per-action snapshot, not just the text, and prune the memories that covered the
removed turns. Retries are kept as browsable variants (
‹ 2/3 ›) rather than thrown away. - Import/export — AI Dungeon-compatible formats for scripts and scenarios; JSON for everything.
- Optional accounts for hosted deployments — by default the app is single-user with zero
auth friction; set
AIDND_MULTI_USER=1and visitors play instantly as guests (signed session cookie), can register (email + password) at any point to keep their data, and each user gets isolated data plus their own encrypted-at-rest API key. A server-funded shared demo key with a daily turn cap lets people try it without bringing a key (backend/app/auth.py).
Screenshots
Quick start
Docker (any OS)
docker compose up --build
Open http://localhost:8000. Your data persists in a named volume across restarts.
Windows
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
cd frontend; npm install; cd ..
.\start.ps1
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
macOS / Linux
./start.sh # creates the venv and installs dependencies on first run
Open http://localhost:5173.
Connect a model
Open Settings in the app and point it at any OpenAI-compatible endpoint:
| Provider | Endpoint URL | Notes |
|---|---|---|
| Ollama (local) | http://localhost:11434/v1 |
free, private; also serves embedding models for the Memory Bank (e.g. nomic-embed-text) |
| LM Studio (local) | http://localhost:1234/v1 |
free, private |
| OpenRouter | https://openrouter.ai/api/v1 |
:free models cost nothing (no embeddings on the free tier) |
| OpenAI / Groq / vLLM / … | provider's /v1 URL |
anything speaking /v1/chat/completions |
Model name, API key, generation parameters, and (optionally) summary/embedding models for the Memory Bank are all configured there too — no config files, no rebuild.
How a turn works
player input
→ onInput script modifier
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
+ [plot essentials] + [story summary] + [retrieved memories]
+ [triggered story cards] + [story history, token-budgeted]
+ [author's note] + [player action]
→ onModelContext script modifier
→ snapshot context (Insights)
→ provider adapter → AI (streamed)
→ extract + referee the world-state delta block, strip it from the prose
→ onOutput script modifier
→ store & render
Architecture
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ auth, scenarios, adventures, story cards, scripts, chat, settings, debug
├─ models.py SQLAlchemy: User, Scenario, Adventure, Action, StoryCard, Script, Settings, Memory
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (37 and counting)
├─ auth.py guest/registered users, sessions, shared demo key
├─ security.py password hashing, cookie signing, API-key encryption
├─ context/ prompt assembly under a token budget + windowed history queries
├─ worldstate/ the stat engine: clamps, cooldowns, bands, milestones
├─ scripting/ quickjs sandbox + AI Dungeon API surface
├─ memorybank.py auto-summarization + embedding retrieval
├─ providers/ OpenAI-compatible adapter, streaming
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
In production the backend serves the built SPA from one port (see Dockerfile); in
development Vite proxies /api to FastAPI.
Tests
151 backend tests — unit plus full HTTP integration through the real quickjs scripting engine, with the LLM provider mocked. CI runs them on every push, alongside the frontend lint/build and a Docker image build.
cd backend && pip install -r requirements.txt -r requirements-dev.txt
python -m pytest tests/
Notes on performance
Two of the tests exist because of bugs that were measured rather than guessed at, and they're the most interesting engineering in the repo:
- Database egress, cut ~189x. Every adventure load was pulling
Action.context_snapshot— the entire assembled prompt, ~74 KB per turn — to read two small fields off it. Moving those fields into their own columns and marking the heavy onesdeferredtook one adventure load from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL so the old data never crosses the wire.tests/test_egress.pyhooks into SQLAlchemy's cursor events and fails if a bulk load ever names those columns again. - Turn cost, made flat. Assembling a turn walked the whole story, so it was O(story length)
— 839 KB of reads at turn 200.
backend/app/context/history.pynow serves tails and slices from SQL and measures what it fetched; the same turn costs 129 KB and stops growing at around turn 50.
Deploy (Render)
The repo ships a render.yaml blueprint: one Docker web service that
serves the SPA and API same-origin, backed by external Neon Postgres
(the free tier has no persistent disk, so the database lives off-box).
- Create a Neon project and copy its pooled connection string.
- In Render: New → Blueprint, point it at this repo. Render reads
render.yaml. - Fill the secrets it prompts for (
sync: falsevars):AIDND_DATABASE_URL(the Neon string) and, to offer a no-signup demo,AIDND_DEMO_API_KEY/AIDND_DEMO_MODELS.AIDND_SECRET_KEYis generated automatically and kept stable across deploys. - Deploy. Pushes to
mainauto-deploy thereafter. Health check:/api/health.
On the free tier the service sleeps after ~15 min idle; the first request then takes
~30–60s to wake. Point any keep-warm pinger at /api/health, which deliberately doesn't
touch the database — waking the database around the clock costs far more than the cold start
is worth.
Repo notes
plan/— the phased implementation plan this was built from, kept as a build log. All twelve phases are complete; the later files (11, 12) double as design notes for the state-revert and world-state work.docs/GUIDE.md— design notes: how each subsystem works and why it was built that way, with the measurements behind the decisions. Also rendered as a reading page.backend/.env.example— the few environment variables the backend reads.docs/self-review.md— a full-codebase self-review pass and what came out of it. All correctness findings are resolved.




