Three fallout bugs from keeping the retried action row alive (906ba42),
plus two long-standing cursor bugs the same investigation turned up.
Retry context leak: the row being regenerated is still attached to the
adventure, so it was replayed as established story and the model wrote a
continuation of the attempt it was meant to replace — the story visibly
blended both takes. It leaked into four places, not one: history replay,
story-card trigger matching, in-scene NPC detection, and the memory-bank
similarity query. Adds a shared context.story_actions(exclude_action_id),
threaded through build_context and retrieve_memories.
Memory holdback: a memory could summarize the just-generated turn; retry
rewrites Action.text but memory_cursor has already advanced, so the memory
was never regenerated and went on describing narration no longer in the
story. settled_story_actions() holds the newest action back one turn —
only the last action is retryable, so that makes it unreachable. The
settled list is always a prefix, so cursors stay valid and nothing is
skipped. The run_post_turn clamp deliberately still uses the full count:
clamping to settled rewinds legacy adventures a step and double-covers an
action.
Cursor bookkeeping: memory_cursor is a position into story_actions() while
Memory.source_* are Action.index values, and the two diverge as soon as
anything is deleted. Deleting a middle action slid a never-summarized
action into the covered range, skipping it forever; and pruning a memory
left the actions it covered stranded behind the cursor. Adds
note_action_removed() (called before the delete in delete_action and
undo_turn) and a rewind in prune_dangling_memories. delete_action also
now prunes at all, which it never did.
Not addressed: editing an already-summarized action still leaves its
memory stale, and the cumulative story summary can't have one fact
un-mixed from it.
117 backend tests pass, including new test_memory_settling.py (12) and
two retry-context regression tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
AI D&D
An AI Dungeon-style interactive storytelling app you can run entirely on your own machine — with your own AI model. Create scenarios, play open-ended adventures where an LLM narrates the world, and extend the engine with JavaScript scripts compatible with real AI Dungeon scripting.
▶️ Try it live: ai-dnd-1gmp.onrender.com
Play a demo scenario as a guest — no sign-up, no API key needed. (Hosted on Render's free tier, so the first load after it's been idle takes ~30–60s to wake up.)
Built with FastAPI + SQLite on the backend and React (Vite) on the frontend. Works with any OpenAI-compatible endpoint: Ollama and LM Studio locally, or OpenRouter / OpenAI / Groq / vLLM in the cloud — endpoint, key, and model are all runtime settings, and OpenRouter's free-tier models make the whole experience $0.
📸 Screenshots and a demo GIF are coming; for now the fastest tour is running it — one command with Docker.
Features
- The full play loop — Do / Say / Story / Continue actions, streamed AI responses (SSE), retry, undo, and edit. Reasoning models supported: "thinking" streams into a collapsible 💭 panel with its own token budget.
- AI Dungeon-compatible context engine — memory, author's note, and story cards (world
info) triggered by keywords in recent story text, assembled under a token budget
(
backend/app/context/builder.py). - Insights: total prompt transparency — every turn stores the exact prompt sent to the model; open 🔍 on any AI action to see each context component and why it was included.
- JavaScript scripting, AI Dungeon-compatible —
onInput/onModelContext/onOutputmodifiers with sharedstateand aworldEntriesAPI, executed in an embedded quickjs sandbox (backend/app/scripting/). Real AI Dungeon scripts import and run. In-app CodeMirror editor included. - Auto-summarization + Memory Bank — the modern AI Dungeon memory system: AI-generated
memories every few actions, a running story summary, and embedding-based retrieval that
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
(
backend/app/memorybank.py). - Import/export — AI Dungeon-compatible formats for scripts and scenarios; JSON for everything.
- Optional accounts for hosted deployments — by default the app is single-user with zero
auth friction; set
AIDND_MULTI_USER=1and visitors play instantly as guests (signed session cookie), can register (email + password) at any point to keep their data, and each user gets isolated data plus their own encrypted-at-rest API key. A server-funded shared demo key with a daily turn cap lets people try it without bringing a key (backend/app/auth.py).
Quick start
Docker (any OS)
docker compose up --build
Open http://localhost:8000. Your data persists in a named volume across restarts.
Windows
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
cd frontend; npm install; cd ..
.\start.ps1
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
macOS / Linux
./start.sh # creates the venv and installs dependencies on first run
Open http://localhost:5173.
Connect a model
Open Settings in the app and point it at any OpenAI-compatible endpoint:
| Provider | Endpoint URL | Notes |
|---|---|---|
| Ollama (local) | http://localhost:11434/v1 |
free, private; also serves embedding models for the Memory Bank (e.g. nomic-embed-text) |
| LM Studio (local) | http://localhost:1234/v1 |
free, private |
| OpenRouter | https://openrouter.ai/api/v1 |
:free models cost nothing (no embeddings on the free tier) |
| OpenAI / Groq / vLLM / … | provider's /v1 URL |
anything speaking /v1/chat/completions |
Model name, API key, generation parameters, and (optionally) summary/embedding models for the Memory Bank are all configured there too — no config files, no rebuild.
How a turn works
player input
→ onInput script modifier
→ assemble context: [AI instructions] + [plot essentials] + [story summary]
+ [retrieved memories] + [triggered story cards]
+ [story history, token-budgeted] + [author's note] + [player action]
→ onModelContext script modifier
→ snapshot context (Insights)
→ provider adapter → AI (streamed)
→ onOutput script modifier
→ store & render
Architecture
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ auth, scenarios, adventures, story cards, scripts, settings, debug
├─ models.py SQLAlchemy: User, Scenario, Adventure, Action, StoryCard, Script, Settings, Memory
├─ auth.py guest/registered users, sessions, shared demo key
├─ security.py password hashing, cookie signing, API-key encryption
├─ context/ prompt assembly under a token budget
├─ scripting/ quickjs sandbox + AI Dungeon API surface
├─ memorybank.py auto-summarization + embedding retrieval
├─ providers/ OpenAI-compatible adapter, streaming
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
In production the backend serves the built SPA from one port (see Dockerfile); in
development Vite proxies /api to FastAPI.
Deploy (Render)
The repo ships a render.yaml blueprint: one Docker web service that
serves the SPA and API same-origin, backed by external Neon Postgres
(the free tier has no persistent disk, so the database lives off-box).
- Create a Neon project and copy its pooled connection string.
- In Render: New → Blueprint, point it at this repo. Render reads
render.yaml. - Fill the secrets it prompts for (
sync: falsevars):AIDND_DATABASE_URL(the Neon string) and, to offer a no-signup demo,AIDND_DEMO_API_KEY/AIDND_DEMO_MODELS.AIDND_SECRET_KEYis generated automatically and kept stable across deploys. - Deploy. Pushes to
mainauto-deploy thereafter. Health check:/api/health.
On the free tier the service sleeps after ~15 min idle; the first request then takes ~30–60s to wake.
Repo notes
plan/— the phased implementation plan this was built from, kept as a build log (phases 1–6 complete; 7–10 cover the public release).backend/.env.example— the few environment variables the backend reads.CODE_REVIEW_FINDINGS.md— notes from a self-review pass.