Make the repo legible to someone seeing it for the first time

The README claimed screenshots were "coming" and stopped describing the
project around phase 10, so the world-state engine, the retry/undo state
revert and the egress work — the most interesting parts — were invisible.

- Screenshots of the play screen, Insights, the schema editor and the
  script editor, captured from the running app.
- Document the world-state engine, undo/retry rewind, and the two measured
  performance fixes; correct the context assembly order to match
  context/builder.py; drop the stale "phases 1-6 complete" note.
- Add CI (backend pytest, frontend lint + build, Docker image build) and
  its badge. Nothing ran the 151 tests but me.
- Move CODE_REVIEW_FINDINGS.md to docs/self-review.md and say up front
  that every correctness finding is resolved.
- Drop a dead import that the newly-wired lint flagged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015sygnUWH8WnaoKPDp7hXM1
This commit is contained in:
parththakkar106
2026-08-10 18:52:02 +05:30
co-authored by Claude Opus 5
parent f295893204
commit ebb81b8f9a
9 changed files with 159 additions and 21 deletions
+74
View File
@@ -0,0 +1,74 @@
name: CI
on:
push:
branches: [main]
pull_request:
workflow_dispatch:
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
backend:
name: Backend tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12" # matches the Dockerfile
cache: pip
cache-dependency-path: |
backend/requirements.txt
backend/requirements-dev.txt
- name: Install dependencies
working-directory: backend
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt -r requirements-dev.txt
- name: Run tests
working-directory: backend
run: python -m pytest tests/ -q
frontend:
name: Frontend lint + build
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: "24" # matches the Dockerfile
cache: npm
cache-dependency-path: frontend/package-lock.json
- name: Install dependencies
working-directory: frontend
run: npm ci
- name: Lint
working-directory: frontend
run: npm run lint
- name: Build
working-directory: frontend
run: npm run build
docker:
name: Docker image builds
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: docker/setup-buildx-action@v3
- name: Build image
uses: docker/build-push-action@v6
with:
context: .
push: false
cache-from: type=gha
cache-to: type=gha,mode=max
+75 -16
View File
@@ -1,5 +1,8 @@
# AI D&D
[![CI](https://github.com/parththakkar106/AI-DnD/actions/workflows/ci.yml/badge.svg)](https://github.com/parththakkar106/AI-DnD/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
An AI Dungeon-style interactive storytelling app you can run entirely on your own machine —
with your own AI model. Create scenarios, play open-ended adventures where an LLM narrates the
world, and extend the engine with **JavaScript scripts compatible with real AI Dungeon
@@ -9,24 +12,33 @@ scripting**.
> Play a demo scenario as a guest — no sign-up, no API key needed. (Hosted on Render's free
> tier, so the first load after it's been idle takes ~30–60s to wake up.)
Built with FastAPI + SQLite on the backend and React (Vite) on the frontend. Works with **any
OpenAI-compatible endpoint**: Ollama and LM Studio locally, or OpenRouter / OpenAI / Groq / vLLM
in the cloud — endpoint, key, and model are all runtime settings, and OpenRouter's free-tier
models make the whole experience $0.
Built with FastAPI + SQLAlchemy on the backend and React (Vite) on the frontend, running on
SQLite locally and Postgres in the cloud. Works with **any OpenAI-compatible endpoint**: Ollama
and LM Studio locally, or OpenRouter / OpenAI / Groq / vLLM in the cloud — endpoint, key, and
model are all runtime settings, and OpenRouter's free-tier models make the whole experience $0.
> 📸 *Screenshots and a demo GIF are coming; for now the fastest tour is running it — one
> command with Docker.*
![The play screen, with the world-state rail open](docs/images/play-world-state.jpg)
*The play screen. The left rail is live world state — the AI proposes changes each turn and a
Python engine decides what actually sticks. The chip under the narration reports what changed.*
## Features
- **The full play loop** — Do / Say / Story / Continue actions, streamed AI responses (SSE),
retry, undo, and edit. Reasoning models supported: "thinking" streams into a collapsible 💭
panel with its own token budget.
- **An RPG world-state engine** — a scenario can declare stats, flags, milestones and a named
cast; the adventure carries their live values. The design is **the AI proposes deltas and a
Python engine referees them**: it clamps to range, enforces per-turn caps and cooldowns, keeps
counters monotonic and milestones sticky, then strips the machine-readable block out of the
prose (`backend/app/worldstate/engine.py`). Word-labelled bands (`40–60: minor damage`) are
what make the model reliable at it. No dice, no scripting required.
- **AI Dungeon-compatible context engine** — memory, author's note, and story cards (world
info) triggered by keywords in recent story text, assembled under a token budget
(`backend/app/context/builder.py`).
- **Insights: total prompt transparency** — every turn stores the exact prompt sent to the
model; open 🔍 on any AI action to see each context component and why it was included.
model; open 🔍 on any AI action to see each context component, its token cost, and why it was
included.
- **JavaScript scripting, AI Dungeon-compatible** — `onInput` / `onModelContext` / `onOutput`
modifiers with shared `state` and a `worldEntries` API, executed in an embedded quickjs
sandbox (`backend/app/scripting/`). Real AI Dungeon scripts import and run. In-app
@@ -35,6 +47,9 @@ models make the whole experience $0.
memories every few actions, a running story summary, and embedding-based retrieval that
pulls old-but-relevant facts back into context, with similarity scores visible in Insights
(`backend/app/memorybank.py`).
- **Undo and retry that actually rewind** — undo and retry roll back the world state and script
state to a per-action snapshot, not just the text, and prune the memories that covered the
removed turns. Retries are kept as browsable variants (`‹ 2/3 ›`) rather than thrown away.
- **Import/export** — AI Dungeon-compatible formats for scripts and scenarios; JSON for
everything.
- **Optional accounts for hosted deployments** — by default the app is single-user with zero
@@ -44,6 +59,15 @@ models make the whole experience $0.
demo key** with a daily turn cap lets people try it without bringing a key
(`backend/app/auth.py`).
## Screenshots
| | |
|---|---|
| ![Insights panel](docs/images/insights.jpg) | ![Scenario editor](docs/images/scenario-editor-npcs.jpg) |
| **Insights** — the exact prompt for the next turn, broken into components with token counts and the trigger word that pulled each story card in. | **Authoring** — stats with ranges, per-turn caps, cooldowns and word-labelled bands; NPCs the AI addresses by id. |
| ![Script editor](docs/images/script-editor.jpg) | ![Home](docs/images/home.jpg) |
| **Scripting** — the three AI Dungeon hooks with shared persistent `state`, run in a quickjs sandbox. | **Home** — continue a story in progress or start from a scenario. |
## Quick start
### Docker (any OS)
@@ -91,12 +115,14 @@ Memory Bank are all configured there too — no config files, no rebuild.
```
player input
→ onInput script modifier
→ assemble context: [AI instructions] + [plot essentials] + [story summary]
+ [retrieved memories] + [triggered story cards]
+ [story history, token-budgeted] + [author's note] + [player action]
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
+ [plot essentials] + [story summary] + [retrieved memories]
+ [triggered story cards] + [story history, token-budgeted]
+ [author's note] + [player action]
→ onModelContext script modifier
→ snapshot context (Insights)
→ provider adapter → AI (streamed)
→ extract + referee the world-state delta block, strip it from the prose
→ onOutput script modifier
→ store & render
```
@@ -105,11 +131,13 @@ player input
```
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ auth, scenarios, adventures, story cards, scripts, settings, debug
├─ routers/ auth, scenarios, adventures, story cards, scripts, chat, settings, debug
├─ models.py SQLAlchemy: User, Scenario, Adventure, Action, StoryCard, Script, Settings, Memory
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (37 and counting)
├─ auth.py guest/registered users, sessions, shared demo key
├─ security.py password hashing, cookie signing, API-key encryption
├─ context/ prompt assembly under a token budget
├─ context/ prompt assembly under a token budget + windowed history queries
├─ worldstate/ the stat engine: clamps, cooldowns, bands, milestones
├─ scripting/ quickjs sandbox + AI Dungeon API surface
├─ memorybank.py auto-summarization + embedding retrieval
├─ providers/ OpenAI-compatible adapter, streaming
@@ -119,6 +147,33 @@ frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
In production the backend serves the built SPA from one port (see `Dockerfile`); in
development Vite proxies `/api` to FastAPI.
## Tests
151 backend tests — unit plus full HTTP integration through the real quickjs scripting engine,
with the LLM provider mocked. CI runs them on every push, alongside the frontend lint/build and
a Docker image build.
```sh
cd backend && pip install -r requirements.txt -r requirements-dev.txt
python -m pytest tests/
```
## Notes on performance
Two of the tests exist because of bugs that were measured rather than guessed at, and they're
the most interesting engineering in the repo:
- **Database egress, cut ~189x.** Every adventure load was pulling `Action.context_snapshot` —
the entire assembled prompt, ~74 KB per turn — to read two small fields off it. Moving those
fields into their own columns and marking the heavy ones `deferred` took one adventure load
from 38.5 MB to 0.20 MB. The backfill runs server-side with dialect-specific SQL so the old
data never crosses the wire. `tests/test_egress.py` hooks into SQLAlchemy's cursor events and
fails if a bulk load ever names those columns again.
- **Turn cost, made flat.** Assembling a turn walked the whole story, so it was O(story length)
— 839 KB of reads at turn 200. `backend/app/context/history.py` now serves tails and slices
from SQL and measures what it fetched; the same turn costs 129 KB and stops growing at around
turn 50.
## Deploy (Render)
The repo ships a [`render.yaml`](render.yaml) blueprint: one Docker web service that
@@ -133,14 +188,18 @@ serves the SPA and API same-origin, backed by external [Neon](https://neon.tech)
4. Deploy. Pushes to `main` auto-deploy thereafter. Health check: `/api/health`.
On the free tier the service sleeps after ~15 min idle; the first request then takes
~30–60s to wake.
~30–60s to wake. Point any keep-warm pinger at `/api/health`, which deliberately doesn't
touch the database — waking the database around the clock costs far more than the cold start
is worth.
## Repo notes
- `plan/` — the phased implementation plan this was built from, kept as a build log
(phases 1–6 complete; 7–10 cover the public release).
- `plan/` — the phased implementation plan this was built from, kept as a build log. All twelve
phases are complete; the later files (11, 12) double as design notes for the state-revert and
world-state work.
- `backend/.env.example` — the few environment variables the backend reads.
- `CODE_REVIEW_FINDINGS.md` — notes from a self-review pass.
- [`docs/self-review.md`](docs/self-review.md) — a full-codebase self-review pass and what came
out of it. All correctness findings are resolved.
## License
Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 157 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 110 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 59 KiB

@@ -1,8 +1,13 @@
# Code Review Findings — 2026-07-05
# Self-review log
Full-codebase review (no git history, so whole project was the scope).
Status: `pending` = not yet fixed, `fixed` = applied, `verify-failed` = finding was wrong on closer look, `skipped` = intentionally not fixed.
Resume point: fix `pending` items top-to-bottom (they are ranked by severity).
**Every correctness bug on this page is resolved** — 20 fixed, 1 intentionally skipped with the
reasoning recorded below. The record is kept because the reasoning outlives the verdicts;
several of these are traps worth remembering. The *cleanup backlog* at the bottom is a
deliberately open list of non-bugs (reuse, simplification, efficiency), not outstanding defects.
Original review: 2026-07-05, whole project in scope (no git history at the time).
Status key: `fixed` = applied, `skipped` = intentionally not fixed, `pending` = outstanding
(none remain).
**2026-07-06 update (branch `bugfix-code-review`):** every finding re-verified against
current code. #1/#2/#3/#4/#6/#12 had already been fixed in earlier sessions (statuses were
+1 -1
View File
@@ -1,7 +1,7 @@
import { useEffect, useState } from 'react'
import { useNavigate } from 'react-router-dom'
import { api } from '../api'
import { downloadJSON, pickJSONFile, useToast } from '../components'
import { pickJSONFile, useToast } from '../components'
export default function Scripts() {
const [scripts, setScripts] = useState(null)