A hosted demo raises a question a local app never does: is anyone using it, and do they reach the part that matters? `/analytics` answers it — visitors, pages, referrers, countries, devices, which shared scenarios get played, turns and demo-key spend, API and turn errors, and a funnel from visited to played a turn to signed up. Not a third-party script, for reasons specific to this one. The CSP allows `script-src 'self'`, so a tracker means loosening it; adblockers eat the popular ones, which silently biases exactly the technical audience this project gets shown to; and none of them can see the measurement that actually matters here, which is a turn, not a pageview. **A visit is a write and never a read.** After the 189x egress fix it would be perverse to add a feature that reads rows per request, so counts accumulate in a process-local dict and flush every 60s as UPSERTs. Storage is a generic `(day, metric, label) -> hits` counter, so measuring something new later costs a constant rather than a migration, plus one row per visitor per day for the funnel flags. Every dashboard query is a GROUP BY returning tens of rows however much traffic sits behind it; a month reads back in a few kilobytes. The buffer's cost is that a hard restart can lose up to a minute — the flusher also runs on shutdown, and a tier that sleeps when idle sleeps on an empty buffer anyway. **The counters are anonymous; the access log beside them is not, on purpose.** A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars — one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined back to `users`, and keyed, so no client can compute one. Story content never reaches that module, and the only content it ever names is a seeded public scenario's title; a player's own titles are theirs. `accesslog.py` is the identifying half and is a separate module writing a separate table so that separation is a property of the code rather than a convention: `access_events` records sessions, sign-ins, registrations and failed attempts with address, email and device, read on a second tab of the same page behind the same gate. Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An unmetered tester is not automatically someone who should see the traffic. The route 404s and the nav link is absent for everyone else, the same treatment AI Chat gets; unset in a hosted deploy means nobody sees it, including me. Three things came out of building it that a test would not have suggested. **A failed turn is an HTTP 200 with a bad ending.** The status-code middleware cannot see one, so a demo whose model had started refusing every request would look perfectly healthy from outside. All five SSE error paths in `_generate_turn` now go through a `turn_error()` helper that counts on the way out. Error buckets elsewhere are labelled by the matched route template rather than the requested path — one bucket per endpoint instead of one per adventure id, and, the reason it isn't merely tidier, an unmatched path is entirely attacker-chosen, so labelling by it would let anyone mint rows. **The funnel counts people, not clicks.** A player who starts six adventures is one person who started an adventure. That is the whole reason the per-visitor-day table exists; its flags only ever turn on, and `is_new` is settled by the first write of a visitor's first day. **The tests run on SQLite and production is Neon.** A flush that raises is caught and logged, so a dialect mistake in the UPSERTs would have stayed invisible until the dashboard quietly never filled. `test_the_upserts_compile_for_postgres` compiles both statements against the Postgres dialect without connecting to one. Two things this leans on elsewhere. `limits._client_ip` is now public `client_ip`: the access log needs the same answer, and two functions both deciding which hop is the caller's is how one of them ends up trusting a header it shouldn't. And the cleanup sweeper now starts if *either* job has work — a deployment can keep every guest forever and still want its visitor-day rows aged out. No migration. Both tables are new and `bootstrap()` calls `create_all` on existing databases too, the route `branches` took in Phase 14, so `LATEST_VERSION` is still 64. 497 tests green, frontend lint and build clean, driven by hand against a synthetic 90-day fixture at 1568px. The narrow-screen layout follows the existing 720px block but is unverified: `resize_window` is ignored on a maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run through the test names themselves, leaving `testclient_ip_*` — still collected by pytest, which is why it passed unnoticed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
256 lines
15 KiB
Markdown
256 lines
15 KiB
Markdown
# AI D&D
|
||
|
||
[](https://github.com/parththakkar106/AI-DnD/actions/workflows/ci.yml)
|
||
[](LICENSE)
|
||
|
||
An AI Dungeon-style interactive storytelling app you can run entirely on your own machine —
|
||
with your own AI model. Create scenarios, play open-ended adventures where an LLM narrates the
|
||
world, and extend the engine with **JavaScript scripts compatible with real AI Dungeon
|
||
scripting**.
|
||
|
||
> ### ▶️ Try it live: **[parththakkar106.github.io/AI-DnD](https://parththakkar106.github.io/AI-DnD/)**
|
||
> The project page loads instantly and launches the hosted demo in one tap — play a scenario as
|
||
> a guest, no sign-up and no API key needed. (The demo runs on a free tier that sleeps, so the
|
||
> first load after it's been idle takes ~30–60s to wake up.)
|
||
>
|
||
> Want the internals? The **[design notes](https://parththakkar106.github.io/AI-DnD/guide.html)**
|
||
> walk through the context budgeting, the world-state referee and the memory bank, and state the
|
||
> reasoning behind each one ([Markdown version](docs/GUIDE.md)).
|
||
|
||
Built with FastAPI + SQLAlchemy on the backend and React (Vite) on the frontend, running on
|
||
SQLite locally and Postgres in the cloud. Works with **any OpenAI-compatible endpoint**: Ollama
|
||
and LM Studio locally, or OpenRouter / OpenAI / Groq / vLLM in the cloud — endpoint, key, and
|
||
model are all runtime settings, and OpenRouter's free-tier models make the whole experience $0.
|
||
|
||

|
||
|
||
*The play screen. The left rail is live world state — the AI proposes changes each turn and a
|
||
Python engine decides what actually sticks. The chip under the narration reports what changed.
|
||
The `‹ 2/2 ›` under a turn steps between the takes it has; writing below one that isn't the
|
||
live one is what starts a new branch.*
|
||
|
||
## Features
|
||
|
||
- **The full play loop** — Do / Say / Story / Continue actions, streamed AI responses (SSE),
|
||
retry, undo, and edit. Reasoning models supported: "thinking" streams into a collapsible 💭
|
||
panel with its own token budget.
|
||
- **A branching story tree** — the story is a tree, not a list. Any turn can hold more than one
|
||
**take**; `‹ 2/4 ›` steps between them, and stepping is free — the story below simply empties,
|
||
and the server is told nothing. **Writing below a take that isn't the live one is what makes a
|
||
branch.** Branches borrow their ancestors' turns instead of copying them, so a fork costs about
|
||
100 bytes and a 20-fork story loads within 1% of the same story flat; switching restores that
|
||
line's world state, script state and cooldown clocks. A branch panel switches, renames and
|
||
deletes; **⌗ See the tree** draws every line against the story's own clock
|
||
(`backend/app/tree.py`, `backend/app/context/lineage.py`).
|
||
- **An RPG world-state engine** — a scenario can declare stats, flags, milestones and a named
|
||
cast; the adventure carries their live values. The design is **the AI proposes deltas and a
|
||
Python engine referees them**: it clamps to range, enforces per-turn caps and cooldowns, keeps
|
||
counters monotonic and milestones sticky, then strips the machine-readable block out of the
|
||
prose (`backend/app/worldstate/engine.py`). Word-labelled bands (`40–60: minor damage`) are
|
||
what make the model reliable at it. No dice, no scripting required.
|
||
- **AI Dungeon-compatible context engine** — memory, author's note, and story cards (world
|
||
info) triggered by keywords in recent story text, assembled under a token budget
|
||
(`backend/app/context/builder.py`).
|
||
- **Insights: total prompt transparency** — every turn stores the exact prompt sent to the
|
||
model; open 🔍 on any AI action to see each context component, its token cost, and why it was
|
||
included.
|
||
- **JavaScript scripting, AI Dungeon-compatible** — `onInput` / `onModelContext` / `onOutput`
|
||
modifiers with shared `state` and a `worldEntries` API, executed in an embedded quickjs
|
||
sandbox (`backend/app/scripting/`). Real AI Dungeon scripts import and run. In-app
|
||
CodeMirror editor included.
|
||
- **Auto-summarization + Memory Bank** — the modern AI Dungeon memory system: AI-generated
|
||
memories every few actions, a running story summary, and embedding-based retrieval that
|
||
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
|
||
(`backend/app/memorybank.py`).
|
||
- **Undo and retry that actually rewind** — undo and retry roll back the world state and script
|
||
state to a per-node snapshot, not just the text, and prune the memories that covered the
|
||
removed turns. Nothing a retry replaces is thrown away: the old attempt stays as another take
|
||
of that turn, and is one keystroke and one click from being a branch of its own.
|
||
- **Import/export** — AI Dungeon-compatible formats for scripts and scenarios; JSON for
|
||
everything. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree —
|
||
every branch, every take, and the fork points, because those were chosen rather than computed.
|
||
Files saved in the old single-line format still import.
|
||
- **Optional accounts for hosted deployments** — by default the app is single-user with zero
|
||
auth friction; set `AIDND_MULTI_USER=1` and visitors play instantly as guests (signed
|
||
session cookie), can register (email + password) at any point to keep their data, and each
|
||
user gets isolated data plus their own encrypted-at-rest API key. A server-funded **shared
|
||
demo key** with a daily turn cap lets people try it without bringing a key
|
||
(`backend/app/auth.py`).
|
||
|
||
## Screenshots
|
||
|
||
| | |
|
||
|---|---|
|
||
|  |  |
|
||
| **Insights** — the exact prompt for the next turn, broken into components with token counts and the trigger word that pulled each story card in. | **Authoring** — stats with ranges, per-turn caps, cooldowns and word-labelled bands; NPCs the AI addresses by id. |
|
||
|  |  |
|
||
| **Scripting** — the three AI Dungeon hooks with shared persistent `state`, run in a quickjs sandbox. | **Home** — continue a story in progress or start from a scenario. |
|
||
|  |  |
|
||
| **The tree** — one lane per line, from the moment it left its parent to the moment it ends. The horizontal axis is the story's own clock, so a short branch reads as short. | **Branches** — every line the story has taken, and the three things you can do to one. A line the one you're reading was forked from can't be deleted, and says so. |
|
||
|
||
## Quick start
|
||
|
||
### Docker (any OS)
|
||
|
||
```sh
|
||
docker compose up --build
|
||
```
|
||
|
||
Open http://localhost:8000. Your data persists in a named volume across restarts.
|
||
|
||
### Windows
|
||
|
||
```powershell
|
||
cd backend; python -m venv .venv; .\.venv\Scripts\pip.exe install -r requirements.txt; cd ..
|
||
cd frontend; npm install; cd ..
|
||
.\start.ps1
|
||
```
|
||
|
||
Open http://localhost:5173 (dev servers; API docs at http://localhost:8000/docs).
|
||
|
||
### macOS / Linux
|
||
|
||
```sh
|
||
./start.sh # creates the venv and installs dependencies on first run
|
||
```
|
||
|
||
Open http://localhost:5173.
|
||
|
||
## Connect a model
|
||
|
||
Open **Settings** in the app and point it at any OpenAI-compatible endpoint:
|
||
|
||
| Provider | Endpoint URL | Notes |
|
||
|---|---|---|
|
||
| Ollama (local) | `http://localhost:11434/v1` | free, private; also serves embedding models for the Memory Bank (e.g. `nomic-embed-text`) |
|
||
| LM Studio (local) | `http://localhost:1234/v1` | free, private |
|
||
| OpenRouter | `https://openrouter.ai/api/v1` | `:free` models cost nothing (no embeddings on the free tier) |
|
||
| OpenAI / Groq / vLLM / … | provider's `/v1` URL | anything speaking `/v1/chat/completions` |
|
||
|
||
Model name, API key, generation parameters, and (optionally) summary/embedding models for the
|
||
Memory Bank are all configured there too — no config files, no rebuild.
|
||
|
||
## How a turn works
|
||
|
||
```
|
||
player input
|
||
→ onInput script modifier
|
||
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
|
||
+ [plot essentials] + [story summary] + [retrieved memories]
|
||
+ [triggered story cards] + [history along this branch, token-budgeted]
|
||
+ [author's note] + [player action]
|
||
→ onModelContext script modifier
|
||
→ snapshot context (Insights)
|
||
→ provider adapter → AI (streamed)
|
||
→ extract + referee the world-state delta block, strip it from the prose
|
||
→ onOutput script modifier
|
||
→ store & render
|
||
```
|
||
|
||
## Architecture
|
||
|
||
```
|
||
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||
├─ routers/ auth, scenarios, adventures, story cards, scripts, chat, settings, analytics, debug
|
||
├─ models.py SQLAlchemy: User, Scenario, Adventure, Branch, Action, StoryCard, Script, Settings, Memory
|
||
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (64 and counting)
|
||
├─ auth.py guest/registered users, sessions, shared demo key
|
||
├─ security.py password hashing, cookie signing, API-key encryption
|
||
├─ tree.py forking, promotion, and where a node is placed
|
||
├─ attempts.py the takes of one turn, grouped by parent
|
||
├─ context/ prompt assembly under a token budget + lineage/history windowing
|
||
├─ worldstate/ the stat engine: clamps, cooldowns, bands, milestones
|
||
├─ scripting/ quickjs sandbox + AI Dungeon API surface
|
||
├─ memorybank.py auto-summarization + embedding retrieval
|
||
├─ analytics.py buffered visit counters + the owner's dashboard query
|
||
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
|
||
├─ providers/ OpenAI-compatible adapter, streaming
|
||
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
|
||
```
|
||
|
||
In production the backend serves the built SPA from one port (see `Dockerfile`); in
|
||
development Vite proxies `/api` to FastAPI.
|
||
|
||
## Tests
|
||
|
||
497 backend tests — unit plus full HTTP integration through the real quickjs scripting engine,
|
||
with the LLM provider mocked. CI runs them on every push, alongside the frontend lint/build and
|
||
a Docker image build.
|
||
|
||
```sh
|
||
cd backend && pip install -r requirements.txt -r requirements-dev.txt
|
||
python -m pytest tests/
|
||
```
|
||
|
||
## Notes on performance
|
||
|
||
Two of the tests exist because of bugs that were measured rather than guessed at, and they're
|
||
the most interesting engineering in the repo:
|
||
|
||
- **Database egress, cut ~189x.** Every adventure load was pulling `Action.context_snapshot` —
|
||
the entire assembled prompt, ~74 KB per turn — to read two small fields off it. Moving those
|
||
fields into their own columns and marking the heavy ones `deferred` took one adventure load
|
||
from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL so the old
|
||
data never crosses the wire. `tests/test_egress.py` hooks into SQLAlchemy's cursor events and
|
||
fails if a bulk load ever names those columns again.
|
||
- **Turn cost, made flat.** Assembling a turn walked the whole story, so it was O(story length)
|
||
— 839 KB of reads at turn 200. `backend/app/context/history.py` now serves tails and slices
|
||
from SQL and measures what it fetched; the same turn costs 129 KB and stops growing at around
|
||
turn 50.
|
||
- **Branching that doesn't cost anything to read.** A branch stores no turns — it stores where it
|
||
left its parent, and borrows everything above that. A 40-turn story forked twenty times loads in
|
||
31,652 B against 31,433 B for the same story flat: **1.007×**, or about 103 B per branch. Reads
|
||
stay cheap because the lineage is windowed like the history is, so the number of SQL clauses is
|
||
bounded by the context window rather than by the number of forks.
|
||
|
||
## Visit analytics
|
||
|
||
The hosted demo keeps its own analytics: an owner-only dashboard at `/analytics` showing
|
||
traffic, which shared scenarios get played, turns and demo-key spend, errors, and a funnel
|
||
from *visited* to *played a turn* to *signed up*. It is visible only to the emails listed in
|
||
`AIDND_ANALYTICS_EMAILS`, and the route 404s for everyone else.
|
||
|
||
Built into the app rather than bolted on with a third-party script, for reasons specific to
|
||
this one: the CSP allows `script-src 'self'`, adblockers eat the popular trackers, and none
|
||
of them can see the measurement that actually matters here — a turn. Counts are aggregated
|
||
in memory and flushed as UPSERTs, so a visit is a write and never a read, and every dashboard
|
||
query is a `GROUP BY` returning tens of rows however much traffic sits behind it. That last
|
||
part is not incidental; see the egress note above for what reading rows per request costs on
|
||
this stack.
|
||
|
||
## Deploy (Render)
|
||
|
||
The repo ships a [`render.yaml`](render.yaml) blueprint: one Docker web service that
|
||
serves the SPA and API same-origin, backed by external [Neon](https://neon.tech) Postgres
|
||
(the free tier has no persistent disk, so the database lives off-box).
|
||
|
||
1. Create a **Neon** project and copy its pooled connection string.
|
||
2. In Render: **New → Blueprint**, point it at this repo. Render reads `render.yaml`.
|
||
3. Fill the secrets it prompts for (`sync: false` vars): `AIDND_DATABASE_URL` (the Neon
|
||
string); to offer a no-signup demo, `AIDND_DEMO_API_KEY` / `AIDND_DEMO_MODELS`; and to
|
||
see the Visitors dashboard, `AIDND_ANALYTICS_EMAILS` (your own account's email).
|
||
`AIDND_SECRET_KEY` is generated automatically and kept stable across deploys.
|
||
4. Deploy. Pushes to `main` auto-deploy thereafter. Health check: `/api/health`.
|
||
|
||
On the free tier the service sleeps after ~15 min idle; the first request then takes
|
||
~30–60s to wake. Point any keep-warm pinger at `/api/health`, which deliberately doesn't
|
||
touch the database — waking the database around the clock costs far more than the cold start
|
||
is worth.
|
||
|
||
## Repo notes
|
||
|
||
- `plan/` — the phased implementation plan this was built from, kept as a build log. All fourteen
|
||
phases are complete; the later files (11, 12, 14) double as design notes for the state-revert,
|
||
world-state and story-tree work. [`plan/STATUS.md`](plan/STATUS.md) is the running thread —
|
||
what shipped, what was measured, and what is owed next.
|
||
- [`docs/GUIDE.md`](docs/GUIDE.md) — design notes: how each subsystem works and why it was built
|
||
that way, with the measurements behind the decisions. Also rendered as a
|
||
[reading page](https://parththakkar106.github.io/AI-DnD/guide.html).
|
||
- `backend/.env.example` — the few environment variables the backend reads.
|
||
- [`docs/self-review.md`](docs/self-review.md) — a full-codebase self-review pass and what came
|
||
out of it. All correctness findings are resolved.
|
||
|
||
## License
|
||
|
||
[MIT](LICENSE)
|