Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was authoritative. Phase 0 execution prompts sat beside the specification; four completed milestone reports sat beside the current one; and upstream AI-DnD's own `plan/` build log and `docs/` project site still described a hosted, scripted, multi-user product with accounts — every screenshot in it showed a Scripts tab and a Sign up button, none of which has existed since M2. `planning/archive/` now holds the history and says so in its own README: `phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied. `planning/reports/` holds only the current milestone's report, because that is the one M4 planning has to read; it moves to the archive when M4's replaces it. Deleted rather than archived: the Phase 0B execution prompts and the handoff/status/summary documents, the Phase 0A discovery and triage reports, upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's template boilerplate. All of it is in Git history, and the two recommendation reports carry every conclusion the deleted research reached. Archived documents are kept verbatim. Paths written inside them point at where those files were when the document was written, which is the point: an evidence record that has been quietly edited is no longer evidence. Active documentation is corrected where it pointed at the removed trees or described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not touch" list had gone stale at M2 and claimed QuickJS scripting was still tested; its test count was 604 against an actual 638. `README.md` loses the upstream CI badge, which reported upstream's pipeline rather than this fork's, and a reference to `backend/app/worldstate/engine.py`, a file that does not exist. `planning/README.md` is rewritten as the documentation index. New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the manifest of what belongs in the ChatGPT project's Sources. Source comments referring to the deleted trees are reworded; no behaviour changes. 638 backend tests pass, frontend lints and builds, and a reference scan over all 48 tracked Markdown files reports no unresolved path in active documentation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
@@ -208,7 +208,7 @@ visible from within.
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
cd backend && .venv/bin/python -m pytest tests/ -q # 604 tests
|
||||
cd backend && .venv/bin/python -m pytest tests/ -q # 638 tests
|
||||
cd frontend && npm run lint && npm run build
|
||||
```
|
||||
|
||||
@@ -266,21 +266,24 @@ docker exec app python -c "import socket; socket.create_connection(('1.1.1.1',44
|
||||
# -> OSError: Network is unreachable, and story turns still work
|
||||
```
|
||||
|
||||
`planning/reports/M1-BASELINE-REPORT.md` records the run this procedure is
|
||||
`planning/archive/milestone-reports/M1-BASELINE-REPORT.md` records the run this procedure is
|
||||
taken from, including the packet captures.
|
||||
|
||||
## Things inherited from upstream that M1 deliberately did not touch
|
||||
## Things still inherited from upstream
|
||||
|
||||
These are M2's scope (`planning/BUILD-MILESTONES.md`), listed here so nobody
|
||||
reports them as new:
|
||||
M2 removed the hosted, cloud, account, analytics, Postgres/Render and QuickJS
|
||||
scripting surfaces outright — `PROVENANCE.md` lists exactly what went. What is
|
||||
left of upstream that a newcomer might report as a defect:
|
||||
|
||||
- hosted/multi-user/account/demo-key code, analytics tables, Postgres and
|
||||
Render deployment paths, and the OpenRouter default endpoint constant all
|
||||
still exist in the tree. None of them is reachable from a default local run,
|
||||
and none requires a cloud service.
|
||||
- `docs/*.html` is upstream's GitHub Pages project site and still links Google
|
||||
Fonts. It is not served by the application and is not part of any build.
|
||||
- `.github/workflows/ci.yml` is upstream's GitHub Actions pipeline. This
|
||||
- **Inert legacy tables and columns.** Five tables and four columns M2 emptied
|
||||
of meaning are still in the schema, unmapped, so an M1-era campaign database
|
||||
opens unchanged. Nothing reads or writes them. A cleanup migration waits for
|
||||
the schema to settle after M5 (`planning/BUILD-MILESTONES.md`).
|
||||
- **Dual-dialect migration code.** `backend/app/migrations.py` still carries
|
||||
SQLite/Postgres branches from upstream, although Postgres support itself is
|
||||
gone and SQLite is the only store. Same cleanup, same milestone.
|
||||
- **`.github/workflows/ci.yml`** is upstream's GitHub Actions pipeline. This
|
||||
repository lives on a self-hosted Gitea; the workflow is kept for provenance
|
||||
and is not what runs the tests here.
|
||||
- QuickJS campaign scripting is still present and still tested.
|
||||
- **No frontend tests.** `npm run lint && npm run build` is the whole frontend
|
||||
check. A test runner is M8's job.
|
||||
|
||||
@@ -26,9 +26,11 @@ import is a merge of the pinned commit with `--allow-unrelated-histories`, so:
|
||||
|
||||
- `git log d72f7c1bda0f34fccd84afb7a25c34eb01c901de` shows the real upstream
|
||||
history, not a squashed snapshot;
|
||||
- upstream paths are unchanged (`backend/`, `frontend/`, `docs/`, …), so a
|
||||
later upstream commit can still be fetched and cherry-picked against
|
||||
matching files;
|
||||
- upstream code paths are unchanged (`backend/`, `frontend/`, …), so a later
|
||||
upstream commit can still be fetched and cherry-picked against matching
|
||||
files. Upstream's own documentation trees, `plan/` and `docs/`, were removed
|
||||
on 2026-09-03: they described the hosted, scripted, multi-user product this
|
||||
fork is not. They remain in this repository's history and in upstream;
|
||||
- the planning package that predates the fork keeps its own history on the
|
||||
other parent of the merge.
|
||||
|
||||
@@ -104,7 +106,7 @@ unchanged. They are not product functionality and nothing reads or writes them.
|
||||
## What this fork changed in Milestone M1
|
||||
|
||||
Nothing was removed from upstream. The changes are the offline/locality
|
||||
hardening M1 called for; see `planning/reports/M1-BASELINE-REPORT.md` for the
|
||||
hardening M1 called for; see `planning/archive/milestone-reports/M1-BASELINE-REPORT.md` for the
|
||||
evidence.
|
||||
|
||||
- `backend/app/context/encoding.py` (new) and `backend/app/context/builder.py` —
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
# AI D&D
|
||||
# Adventure Storyteller
|
||||
|
||||
[](https://github.com/parththakkar106/AI-DnD/actions/workflows/ci.yml)
|
||||
[](LICENSE)
|
||||
|
||||
An interactive storytelling app that runs entirely on your own machine, with your own model.
|
||||
@@ -18,19 +17,17 @@ disabled. What is left is a storyteller you can run offline.
|
||||
> public address. There is no telemetry, no account, no cloud inference, and nothing is fetched
|
||||
> at runtime from the Internet.
|
||||
>
|
||||
> For the internals, read the **[design notes](docs/GUIDE.md)**. They walk through the context
|
||||
> budgeting, the world-state referee, and the memory bank, and state the reasoning behind each
|
||||
> one. Some sections still describe upstream subsystems this fork has removed.
|
||||
> For the internals, read [`planning/TECHNICAL-DESIGN.md`](planning/TECHNICAL-DESIGN.md) and
|
||||
> [`planning/CONTEXT-AND-MEMORY.md`](planning/CONTEXT-AND-MEMORY.md), which cover the context
|
||||
> budgeting, the state model and the memory bank as this fork builds them.
|
||||
|
||||
Built with FastAPI and SQLAlchemy on the backend and React (Vite) on the frontend, storing
|
||||
everything in one SQLite file.
|
||||
|
||||

|
||||
|
||||
*The play screen. The left rail shows live world state. The AI proposes changes each turn, and
|
||||
a Python engine decides what actually sticks. The chip under the narration reports what
|
||||
changed. The `‹ 2/2 ›` under a turn steps between the takes it has. Writing below a take that
|
||||
isn't the live one starts a new branch.*
|
||||
On the play screen, the left rail carries live world state. The AI proposes changes each turn
|
||||
and a Python engine decides what actually sticks; the chip under the narration reports what
|
||||
changed. The `‹ 2/2 ›` under a turn steps between the takes it has, and writing below a take
|
||||
that isn't the live one starts a new branch.
|
||||
|
||||
## Features
|
||||
|
||||
@@ -49,7 +46,8 @@ isn't the live one starts a new branch.*
|
||||
cast; the adventure carries their live values. The AI proposes deltas, and a Python engine
|
||||
referees them: it clamps values to range, enforces per-turn caps and cooldowns, keeps counters
|
||||
monotonic and milestones sticky, then strips the machine-readable block out of the prose
|
||||
(`backend/app/worldstate/engine.py`). Word-labeled bands (`40–60: minor damage`) make the
|
||||
(`backend/app/worldstate/`: `apply.py` clamps, `parse.py` reads the block back).
|
||||
Word-labeled bands (`40–60: minor damage`) make the
|
||||
model reliable at it. No dice and no scripting are required.
|
||||
- **AI Dungeon-compatible context engine.** Memory, author's note, and story cards (world
|
||||
info) are triggered by keywords in recent story text, then assembled under a token budget
|
||||
@@ -88,14 +86,10 @@ isn't the live one starts a new branch.*
|
||||
|
||||
## Screenshots
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
|  |  |
|
||||
| **Insights**: the exact prompt for the next turn, broken into components with token counts and the trigger word that pulled each story card in. | **Authoring**: stats with ranges, per-turn caps, cooldowns, and word-labeled bands; NPCs the AI addresses by id. |
|
||||
|  |  |
|
||||
| **Home**: continue a story in progress or start from a scenario. | **The tree**: one lane per line, from the moment it left its parent to the moment it ends. The horizontal axis is the story's own clock, so a short branch reads as short. |
|
||||
|  | |
|
||||
| **Branches**: every line the story has taken, and the three things you can do to one. A line the one you're reading was forked from can't be deleted, and says so. | |
|
||||
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
|
||||
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
|
||||
removed rather than left standing as a picture of a product that no longer exists. New ones
|
||||
are taken when the browser smoke test M3 still owes is run.
|
||||
|
||||
## Quick start
|
||||
|
||||
@@ -248,16 +242,15 @@ most interesting engineering in the repo.
|
||||
- `planning/` is this fork's own package: the product specification, the architecture
|
||||
decisions, the milestone plan, the acceptance contract, and a review report for every
|
||||
milestone shipped. Start at [`planning/README.md`](planning/README.md).
|
||||
- `plan/` holds the *upstream* project's phased implementation plan, kept as a build log. The
|
||||
later files (11, 12, 14) still serve as design notes for the state-revert, world-state, and
|
||||
story-tree work this fork inherited.
|
||||
- [`docs/GUIDE.md`](docs/GUIDE.md) holds upstream's design notes: how each subsystem works and
|
||||
why it was built that way, with the measurements behind the decisions. Sections covering
|
||||
scripting, accounts and hosted deployment describe subsystems this fork removed.
|
||||
- [`planning/archive/`](planning/archive/README.md) holds the Phase 0 research that chose this
|
||||
base and the completed milestone reports. It is history, not instruction.
|
||||
- [`DEVELOPMENT.md`](DEVELOPMENT.md) is how to set the project up, point it at a model, and run
|
||||
the tests. [`PROVENANCE.md`](PROVENANCE.md) records what came from upstream and what changed.
|
||||
- `backend/.env.example` lists the two environment variables the backend reads. Everything
|
||||
about the model is a runtime setting on the Settings page instead.
|
||||
- [`docs/self-review.md`](docs/self-review.md) records a full-codebase self-review pass and
|
||||
what came out of it. All correctness findings are resolved.
|
||||
- Upstream's own `plan/` build log and `docs/` project site were removed in the 2026-09-03
|
||||
documentation pass: they described AI-DnD's hosted, scripted, multi-user product. Both are
|
||||
still in Git history, and in upstream.
|
||||
|
||||
## License
|
||||
|
||||
|
||||
@@ -301,7 +301,8 @@ MIGRATIONS: list[tuple[int, str | dict[str, str]]] = [
|
||||
(63, "ALTER TABLE actions ADD COLUMN parent_id INTEGER REFERENCES actions(id) ON DELETE SET NULL"),
|
||||
(64, "CREATE INDEX IF NOT EXISTS ix_actions_parent ON actions (parent_id)"),
|
||||
# Phase 17: `Settings.stream` was dead state. Nothing ever read it, and every
|
||||
# turn streams. This is item S1 in `docs/self-review.md`. The table holds one
|
||||
# turn streams; upstream's self-review log flagged it as dead state. The
|
||||
# table holds one
|
||||
# row per user, so the rewrite is small and needs no VACUUM FULL.
|
||||
(65, "ALTER TABLE settings DROP COLUMN stream"),
|
||||
# Phase 17, SP8: drop the eight columns the story tree replaced. Each one was
|
||||
|
||||
@@ -4,7 +4,7 @@ Most of this file used to be about the shared demo key: an access gate on a
|
||||
"power user" email allowlist, and a pinning rule that stopped a public visitor
|
||||
reaching paid models on a server-funded key. M2 removed the hosted deployment
|
||||
those defended, so the rules they tested no longer exist to be tested. See
|
||||
`planning/reports/M2-*` for the accounting.
|
||||
`planning/archive/milestone-reports/M2-*` for the accounting.
|
||||
|
||||
What remains is what the page still does: stream a reply from the configured
|
||||
model, honour a system prompt and a per-request model override, and refuse a
|
||||
|
||||
@@ -1,5 +1,4 @@
|
||||
"""Tests for undo and retry rolling back the shared `world_state`
|
||||
(plan/11-state-revert-and-retry-fix.md).
|
||||
"""Tests for undo and retry rolling back the shared `world_state`.
|
||||
|
||||
The state being rolled back was the scripting engine's `script_state` until
|
||||
M2 removed campaign scripting. The machinery under test — `attempts.restore_state`,
|
||||
|
||||
@@ -15,7 +15,7 @@ This file must pass unmodified through SP1 (schema), SP2 (branch clause),
|
||||
and SP3 (memories on nodes). If a change here looks necessary in one of
|
||||
those subphases, the change is wrong, not the test. SP4 is the first
|
||||
subphase allowed to move it, and only for the variant-count semantics
|
||||
called out in plan/14.
|
||||
the story-tree migration changed.
|
||||
"""
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
|
||||
@@ -31,8 +31,7 @@ from app.main import app
|
||||
# The three tables as they stood at schema 45, frozen. This is a snapshot
|
||||
# of a past schema. It must not be updated to track `models.py`, because
|
||||
# the whole point is that it lacks what SP1 adds. This DDL uses SQLite
|
||||
# syntax only. The migration's Postgres half is exercised against a real
|
||||
# server at deploy time (see plan/14).
|
||||
# syntax only.
|
||||
PRE_TREE_DDL = (
|
||||
"""
|
||||
CREATE TABLE adventures (
|
||||
|
||||
@@ -4,10 +4,9 @@ Reviewing a prompt tells you what it asks for. It does not tell you what a model
|
||||
does with it. This runs the real pipeline twice over the *same* blocks of the
|
||||
*same* story, changing only the prompt, and prints the memories side by side.
|
||||
|
||||
What it showed the first time it was run, and why `MEMORY_MAX_WORDS` exists, is
|
||||
written up in `plan/18-persona-and-memory-quality.md`. Two consecutive memories
|
||||
from one story came back in two different persons, and the same model wrote 34
|
||||
words for one block and 105 for the next.
|
||||
Why `MEMORY_MAX_WORDS` exists is what it showed the first time it was run: two
|
||||
consecutive memories from one story came back in two different persons, and the
|
||||
same model wrote 34 words for one block and 105 for the next.
|
||||
|
||||
Three things make it a fair test rather than a demonstration:
|
||||
|
||||
|
||||
|
Before Width: | Height: | Size: 36 KiB |
|
Before Width: | Height: | Size: 103 KiB |
|
Before Width: | Height: | Size: 66 KiB |
|
Before Width: | Height: | Size: 157 KiB |
|
Before Width: | Height: | Size: 94 KiB |
|
Before Width: | Height: | Size: 84 KiB |
|
Before Width: | Height: | Size: 59 KiB |
@@ -1,267 +0,0 @@
|
||||
<!doctype html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||
<title>AI D&D — an AI Dungeon-style storytelling engine</title>
|
||||
<meta name="description" content="An AI Dungeon-style interactive storytelling app. FastAPI + React, any OpenAI-compatible model, an RPG world-state engine the AI proposes and Python referees, and a quickjs sandbox that runs real AI Dungeon scripts.">
|
||||
<meta property="og:title" content="AI D&D — an AI Dungeon-style storytelling engine">
|
||||
<meta property="og:description" content="Play open-ended adventures narrated by an LLM, with a world-state engine that keeps the numbers honest.">
|
||||
<meta property="og:image" content="https://parththakkar106.github.io/AI-DnD/images/play-world-state.jpg">
|
||||
<meta property="og:type" content="website">
|
||||
<meta name="twitter:card" content="summary_large_image">
|
||||
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'%3E%3Ctext y='.9em' font-size='90'%3E%E2%9A%94%3C/text%3E%3C/svg%3E">
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
|
||||
<link href="https://fonts.googleapis.com/css2?family=Cinzel:wght@500;700&family=Inter:wght@400;500;600&display=swap" rel="stylesheet">
|
||||
<style>
|
||||
:root {
|
||||
--bg: #0a0a0f;
|
||||
--panel: #131320;
|
||||
--border: #2b2b3d;
|
||||
--border-bright: #3d3d55;
|
||||
--text: #e2ddd0;
|
||||
--dim: #918c7d;
|
||||
--accent: #d4a94e;
|
||||
--accent-bright: #e8c476;
|
||||
--accent-dim: #96773a;
|
||||
--display: 'Cinzel', Georgia, serif;
|
||||
--ui: 'Inter', 'Segoe UI', system-ui, sans-serif;
|
||||
color-scheme: dark;
|
||||
}
|
||||
* { box-sizing: border-box; }
|
||||
body {
|
||||
margin: 0;
|
||||
background:
|
||||
radial-gradient(1200px 700px at 15% -10%, rgba(212,169,78,.06), transparent 60%),
|
||||
radial-gradient(1000px 600px at 90% 110%, rgba(88,76,140,.08), transparent 55%),
|
||||
var(--bg);
|
||||
background-attachment: fixed;
|
||||
color: var(--text);
|
||||
font-family: var(--ui);
|
||||
line-height: 1.65;
|
||||
-webkit-font-smoothing: antialiased;
|
||||
}
|
||||
.wrap { max-width: 1080px; margin: 0 auto; padding: 0 24px; }
|
||||
a { color: var(--accent-bright); }
|
||||
|
||||
header { padding: 72px 0 40px; text-align: center; }
|
||||
.mark { font-family: var(--display); font-size: 14px; letter-spacing: .28em; color: var(--accent); text-transform: uppercase; }
|
||||
h1 {
|
||||
font-family: var(--display); font-weight: 700;
|
||||
font-size: clamp(2.4rem, 6vw, 4rem); margin: .2em 0 .1em; letter-spacing: .02em;
|
||||
background: linear-gradient(180deg, var(--accent-bright), var(--accent));
|
||||
-webkit-background-clip: text; background-clip: text; color: transparent;
|
||||
}
|
||||
.tagline { font-size: clamp(1.05rem, 2.2vw, 1.3rem); color: var(--text); max-width: 46ch; margin: .6em auto 0; }
|
||||
.sub { color: var(--dim); max-width: 60ch; margin: 1em auto 0; font-size: .97rem; }
|
||||
|
||||
.cta { display: flex; gap: 14px; justify-content: center; flex-wrap: wrap; margin: 32px 0 10px; }
|
||||
.btn {
|
||||
display: inline-block; padding: 13px 26px; border-radius: 8px; text-decoration: none;
|
||||
font-weight: 600; font-size: 1rem; border: 1px solid var(--border-bright); transition: .18s;
|
||||
}
|
||||
.btn-primary { background: linear-gradient(180deg, var(--accent-bright), var(--accent)); color: #17130a; border-color: var(--accent); }
|
||||
.btn-primary:hover { filter: brightness(1.08); transform: translateY(-1px); }
|
||||
.btn-ghost { background: var(--panel); color: var(--text); }
|
||||
.btn-ghost:hover { border-color: var(--accent); color: var(--accent-bright); }
|
||||
.wake { color: var(--dim); font-size: .85rem; text-align: center; margin-top: 4px; }
|
||||
|
||||
figure { margin: 0; }
|
||||
figure img {
|
||||
width: 100%; height: auto; display: block; border-radius: 10px;
|
||||
border: 1px solid var(--border); box-shadow: 0 24px 60px rgba(0,0,0,.55);
|
||||
}
|
||||
figcaption { color: var(--dim); font-size: .88rem; margin-top: 12px; }
|
||||
.hero-shot { margin: 44px 0 8px; }
|
||||
|
||||
section { padding: 56px 0; border-top: 1px solid var(--border); margin-top: 56px; }
|
||||
h2 {
|
||||
font-family: var(--display); font-size: 1.6rem; font-weight: 700;
|
||||
color: var(--accent); margin: 0 0 8px; letter-spacing: .02em;
|
||||
}
|
||||
.lede { color: var(--dim); margin: 0 0 32px; max-width: 68ch; }
|
||||
|
||||
.grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(min(100%, 420px), 1fr)); gap: 32px; }
|
||||
.card { background: var(--panel); border: 1px solid var(--border); border-radius: 12px; padding: 22px; }
|
||||
.card h3 { font-family: var(--display); font-size: 1.12rem; margin: 0 0 8px; color: var(--text); }
|
||||
.card p { margin: 0 0 16px; color: var(--dim); font-size: .95rem; }
|
||||
.card p:last-child { margin-bottom: 0; }
|
||||
.card img { border-radius: 8px; border: 1px solid var(--border); width: 100%; height: auto; display: block; }
|
||||
|
||||
pre {
|
||||
background: var(--panel); border: 1px solid var(--border); border-radius: 10px;
|
||||
padding: 20px; overflow-x: auto; font-size: .86rem; line-height: 1.7; color: var(--text);
|
||||
}
|
||||
code { font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace; }
|
||||
|
||||
.stats { display: grid; grid-template-columns: repeat(auto-fit, minmax(180px, 1fr)); gap: 20px; margin-top: 8px; }
|
||||
.stat { background: var(--panel); border: 1px solid var(--border); border-radius: 12px; padding: 20px; }
|
||||
.stat .n { font-family: var(--display); font-size: 1.9rem; color: var(--accent-bright); line-height: 1.1; }
|
||||
.stat .l { color: var(--dim); font-size: .88rem; margin-top: 6px; }
|
||||
|
||||
ul.notes { padding-left: 0; list-style: none; margin: 0; }
|
||||
ul.notes li { border-left: 2px solid var(--accent-dim, #96773a); padding: 2px 0 2px 18px; margin-bottom: 22px; color: var(--dim); }
|
||||
ul.notes strong { color: var(--text); }
|
||||
|
||||
.stack { display: flex; flex-wrap: wrap; gap: 8px; margin-top: 20px; }
|
||||
.chip { background: var(--panel); border: 1px solid var(--border); border-radius: 999px; padding: 5px 14px; font-size: .85rem; color: var(--dim); }
|
||||
|
||||
footer { border-top: 1px solid var(--border); margin-top: 56px; padding: 36px 0 64px; color: var(--dim); font-size: .9rem; text-align: center; }
|
||||
|
||||
@media (max-width: 720px) {
|
||||
header { padding: 48px 0 24px; }
|
||||
section { padding: 40px 0; margin-top: 40px; }
|
||||
.btn { width: 100%; }
|
||||
}
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
|
||||
<div class="wrap">
|
||||
|
||||
<header>
|
||||
<div class="mark">⚔ Interactive fiction, refereed</div>
|
||||
<h1>AI D&D</h1>
|
||||
<p class="tagline">Open-ended adventures narrated by an LLM — with an engine that keeps the numbers honest.</p>
|
||||
<p class="sub">Create a world, play it in second person, and let the model improvise the story while a Python
|
||||
referee enforces what's actually true: hit points, an ally's trust, a raised alarm, a quest milestone.
|
||||
Bring your own model, or play the demo with none.</p>
|
||||
|
||||
<div class="cta">
|
||||
<a class="btn btn-primary" href="https://ai-dnd-1gmp.onrender.com">Launch the live demo →</a>
|
||||
<a class="btn btn-ghost" href="guide.html">Read the design notes</a>
|
||||
<a class="btn btn-ghost" href="architecture.html">HLD & LLD walkthrough</a>
|
||||
<a class="btn btn-ghost" href="https://github.com/parththakkar106/AI-DnD">View the source</a>
|
||||
</div>
|
||||
<p class="wake">No sign-up, no API key. Hosted on a free tier that sleeps — the first load takes ~30–60s to wake.</p>
|
||||
|
||||
<figure class="hero-shot">
|
||||
<img src="images/play-world-state.jpg" alt="The play screen with the world-state rail open, showing HP, mana, an NPC's trust and a raised alarm flag">
|
||||
<figcaption>The left rail is live world state. The model proposes what changed this turn; the engine decides
|
||||
what sticks, and the chip under the narration reports the result. The <code>‹ 2/2 ›</code>
|
||||
under a turn steps between the takes it has.</figcaption>
|
||||
</figure>
|
||||
</header>
|
||||
|
||||
<section>
|
||||
<h2>What makes it more than a chat wrapper</h2>
|
||||
<p class="lede">Four things a plain "talk to a model" app doesn't do.</p>
|
||||
|
||||
<div class="grid">
|
||||
<div class="card">
|
||||
<h3>The story is a tree</h3>
|
||||
<p>Any turn can hold more than one <em>take</em>. Stepping between them is free — the story below simply
|
||||
empties, and the server is told nothing. Writing below a take that isn't the live one is what makes a
|
||||
branch, and a branch stores no turns of its own: it records where it left its parent and borrows
|
||||
everything above that. Twenty forks cost 1.007× the page load of the same story flat. Switch lines and
|
||||
the world state, the script scoreboard and the cooldown clocks all come back to what that line left.</p>
|
||||
<img src="images/branch-map.jpg" alt="The branch map: one horizontal lane per line of the story, each joined to its parent by an elbow at the moment it forked">
|
||||
</div>
|
||||
|
||||
<div class="card">
|
||||
<h3>The AI proposes, Python referees</h3>
|
||||
<p>A scenario declares stats, flags, milestones and a named cast. Each turn the model appends the changes
|
||||
it thinks happened — and the engine clamps them to range, enforces per-turn caps and cooldowns, keeps
|
||||
counters monotonic and milestones sticky, then strips the machine-readable block out of the prose.
|
||||
Word-labelled bands (<code>40–60: minor damage</code>) are what make the model reliable at it.
|
||||
No dice, no scripting required.</p>
|
||||
<img src="images/scenario-editor-npcs.jpg" alt="The scenario editor showing NPC stats with ranges, per-turn caps, cooldowns and labelled bands">
|
||||
</div>
|
||||
|
||||
<div class="card">
|
||||
<h3>You can see the entire prompt</h3>
|
||||
<p>Every turn stores exactly what was sent to the model. Open Insights on any action to see each context
|
||||
component, what it cost in tokens, and why it was there — including which trigger word pulled in each
|
||||
story card and the similarity score behind each retrieved memory.</p>
|
||||
<img src="images/insights.jpg" alt="The Insights panel showing the assembled prompt broken into components with token counts">
|
||||
</div>
|
||||
|
||||
<div class="card">
|
||||
<h3>Real AI Dungeon scripts run</h3>
|
||||
<p>The three familiar hooks — <code>onInput</code>, <code>onModelContext</code>, <code>onOutput</code> —
|
||||
with shared persistent <code>state</code> and a <code>worldEntries</code> API, executed in an embedded
|
||||
quickjs sandbox. Scripts written for AI Dungeon import and work, and there's a CodeMirror editor in the app.</p>
|
||||
<img src="images/script-editor.jpg" alt="The in-app script editor showing an input hook written in JavaScript">
|
||||
</div>
|
||||
|
||||
<div class="card">
|
||||
<h3>Memory that survives a long story</h3>
|
||||
<p>The modern AI Dungeon memory system: AI-generated memories every few actions, a running story summary,
|
||||
and embedding-based retrieval that pulls an old-but-relevant fact back into context when it matters.
|
||||
Undo and retry roll the world state back to a per-action snapshot rather than only rewriting the text.
|
||||
Every story stays where you left it, and the home page opens on its most recent line.</p>
|
||||
<img src="images/home.jpg" alt="The home page, showing stories in progress alongside scenarios to start from">
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>How a turn works</h2>
|
||||
<p class="lede">Player input goes through the script pipeline, into a token-budgeted context, out to whichever
|
||||
model you configured, and back through the referee.</p>
|
||||
<pre><code>player input
|
||||
→ onInput script modifier
|
||||
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
|
||||
+ [plot essentials] + [story summary] + [retrieved memories]
|
||||
+ [triggered story cards] + [history along this branch, token-budgeted]
|
||||
+ [author's note] + [player action]
|
||||
→ onModelContext script modifier
|
||||
→ snapshot context (Insights)
|
||||
→ provider adapter → AI (streamed)
|
||||
→ extract + referee the world-state delta block, strip it from the prose
|
||||
→ onOutput script modifier
|
||||
→ store & render</code></pre>
|
||||
|
||||
<div class="stack">
|
||||
<span class="chip">FastAPI</span>
|
||||
<span class="chip">SQLAlchemy</span>
|
||||
<span class="chip">React + Vite</span>
|
||||
<span class="chip">Postgres / SQLite</span>
|
||||
<span class="chip">quickjs sandbox</span>
|
||||
<span class="chip">Server-sent events</span>
|
||||
<span class="chip">Docker</span>
|
||||
<span class="chip">Any OpenAI-compatible endpoint</span>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>Engineering notes</h2>
|
||||
<p class="lede">The parts that were measured rather than guessed at.</p>
|
||||
|
||||
<div class="stats">
|
||||
<div class="stat"><div class="n">189×</div><div class="l">less database egress per adventure load</div></div>
|
||||
<div class="stat"><div class="n">440</div><div class="l">backend tests, run by CI on every push</div></div>
|
||||
<div class="stat"><div class="n">64</div><div class="l">schema migrations, applied in order on boot</div></div>
|
||||
<div class="stat"><div class="n">$0</div><div class="l">to run it locally against Ollama</div></div>
|
||||
</div>
|
||||
|
||||
<ul class="notes" style="margin-top:34px">
|
||||
<li><strong>Database egress, cut ~189×.</strong> Every adventure load was pulling the entire assembled
|
||||
prompt — about 74 KB per turn — just to read two small fields off it. Moving those into their own columns
|
||||
and deferring the heavy ones took one load from 38.5 MB to 0.20 MB. A test hooks into SQLAlchemy's cursor
|
||||
events and fails if a bulk load ever names those columns again.</li>
|
||||
<li><strong>Turn cost, made flat.</strong> Assembling a turn walked the whole story, so it grew with story
|
||||
length — 839 KB of reads by turn 200. History is now served as tails and slices from SQL: the same turn
|
||||
costs 129 KB and stops growing at around turn 50.</li>
|
||||
<li><strong>Branching that costs 103 bytes.</strong> A branch stores where it left its parent and borrows
|
||||
every turn above that, so nothing is copied on a fork. A 40-turn story forked twenty times loads in
|
||||
31,652 B against 31,433 B for the same story flat — 1.007×. Reads stay cheap because the ancestry is
|
||||
windowed the way the history is: the number of SQL clauses is bounded by the context window, not by how
|
||||
many times the story has forked.</li>
|
||||
<li><strong>A shared demo key that can't be drained.</strong> The hosted demo funds a model for visitors, so
|
||||
model selection is pinned server-side with a structural backstop that raises if any code path tries to
|
||||
resolve a model outside the allowed set — plus a daily per-visitor turn cap.</li>
|
||||
</ul>
|
||||
</section>
|
||||
|
||||
<footer>
|
||||
<p>Built by <a href="https://github.com/parththakkar106">Parth Thakkar</a> ·
|
||||
<a href="https://github.com/parththakkar106/AI-DnD">Source on GitHub</a> ·
|
||||
<a href="https://github.com/parththakkar106/AI-DnD/blob/main/LICENSE">MIT</a></p>
|
||||
<p>Run it yourself with one command: <code>docker compose up --build</code></p>
|
||||
</footer>
|
||||
|
||||
</div>
|
||||
</body>
|
||||
</html>
|
||||
@@ -1,156 +0,0 @@
|
||||
# Self-review log
|
||||
|
||||
**Every correctness bug on this page is resolved.** 20 are fixed, and 1 is intentionally
|
||||
skipped, with the reasoning recorded below. The record is kept because the reasoning outlives
|
||||
the verdicts, and several of these are traps worth remembering. The *cleanup backlog* at the
|
||||
bottom is a deliberately open list of non-bugs (reuse, simplification, efficiency), not
|
||||
outstanding defects.
|
||||
|
||||
Original review: 2026-07-05, whole project in scope (no git history at the time).
|
||||
Status key: `fixed` means applied, `skipped` means intentionally not fixed, and `pending` means
|
||||
outstanding (none remain).
|
||||
|
||||
**2026-07-06 update (branch `bugfix-code-review`):** every finding was re-verified against
|
||||
current code. #1, #2, #3, #4, #6, and #12 had already been fixed in earlier sessions, so their
|
||||
statuses were stale. #5, #7, #8, #9, #10, #13, #14 (backend) and #15 through #21 (frontend)
|
||||
were fixed in this pass. #11 is skipped: real AI Dungeon's `addStoryCard` also returns the new
|
||||
card's index (0-falsy included) per the official scripting guidebook, so changing it would
|
||||
break compatibility. This is documented in `engine.py`'s prelude instead. The cleanup backlog
|
||||
(R/S/E/A items) below remains open by design.
|
||||
|
||||
## Correctness bugs
|
||||
|
||||
### 1. [fixed] seed_demo.py doesn't stamp schema version → server crashes on next start
|
||||
- `backend/seed_demo.py:10`
|
||||
- A fresh DB created via `Base.metadata.create_all` leaves `PRAGMA user_version` at 0. The next server start sees the tables exist and replays every ALTER TABLE migration, causing a `duplicate column name` crash.
|
||||
- Fix: stamp user_version to latest after create_all (reuse migrations.bootstrap logic).
|
||||
|
||||
### 2. [fixed] Turn-lock race: two simultaneous turns can run on the same adventure
|
||||
- `backend/app/routers/adventures.py:302`
|
||||
- `ensure_not_generating()` runs in the route handler, but `_active_turns.add()` only happens when the StreamingResponse generator is first iterated. Double-clicking Continue sends both requests past the 409 check, producing duplicate Action.index rows and interleaved generations.
|
||||
- Fix: atomically test-and-set the lock in the request phase, release in the stream's `finally`.
|
||||
|
||||
### 3. [fixed] Migration 10 renumbers indexes with a correlated subquery on the table being updated
|
||||
- `backend/app/migrations.py:34`
|
||||
- SQLite may evaluate the subquery against partially updated rows, so duplicate indexes can survive the "repair".
|
||||
- Fix: compute new indexes in Python (SELECT ordered, then UPDATE per row).
|
||||
|
||||
### 4. [fixed] SQLite foreign keys never enabled → CASCADE/SET NULL clauses are dead
|
||||
- `backend/app/database.py:8`
|
||||
- Deleting a Script leaves orphaned `scenario_scripts` rows. SQLite rowid reuse can then attach a future script to an old scenario.
|
||||
- Fix: `PRAGMA foreign_keys=ON` via engine connect event.
|
||||
|
||||
### 5. [fixed] Provider generate() silently yields nothing for non-SSE 200 responses
|
||||
- `backend/app/providers/openai_compatible.py:84`
|
||||
- A server that ignores `stream=true` and returns plain JSON produces no `data:` lines, so the result is an empty AI action with no error.
|
||||
- Fix: buffer the non-SSE body and fall back to parsing it as a single JSON completion.
|
||||
|
||||
### 6. [fixed] Fire-and-forget asyncio task can be GC'd mid-run and wedge the memory bank
|
||||
- `backend/app/memorybank.py:146`
|
||||
- The result of `asyncio.create_task` is not referenced, so the task can vanish silently and leave an adventure ID stuck in `_running`.
|
||||
- Fix: keep strong refs in a set, discard in done-callback.
|
||||
|
||||
### 7. [fixed] Memory cursors are list positions but Memory.source_start/end are Action.index values
|
||||
- `backend/app/memorybank.py:213`
|
||||
- After any action deletion, indexes keep gaps and positions shift, so summarization skips or duplicates blocks and `_update_story_summary` folds the wrong memories.
|
||||
- Fix: use one space consistently. Track cursors by Action.index (position-independent), or renumber on delete.
|
||||
|
||||
### 8. [fixed] Pinned memories don't count toward memory_top_k cap
|
||||
- `backend/app/memorybank.py:120`
|
||||
- 6 pinned plus top_k=5 injects 11 memories, blowing the token budget.
|
||||
- Fix: fill with unpinned only up to `top_k - len(pinned)` (min 0).
|
||||
|
||||
### 9. [fixed] Embedding-model change → cosine() zips different-dimension vectors silently
|
||||
- `backend/app/memorybank.py:69`
|
||||
- Old 768-dimension embeddings scored against a new 1536-dimension query produce garbage similarity with no error, and are never re-embedded.
|
||||
- Fix: return 0.0 on length mismatch (and ideally clear stale embeddings so _embed_pending redoes them).
|
||||
|
||||
### 10. [fixed] MAX_STORY_CARDS cap is a no-op for cards created in one hook
|
||||
- `backend/app/scripting/pipeline.py:59`
|
||||
- `len(existing) + len(seen_ids) < MAX...` never counts newly added cards, since seen_ids is a subset of existing. A script can insert an unbounded number of cards in one turn.
|
||||
- Fix: count inserts made during the loop.
|
||||
|
||||
### 11. [skipped] addStoryCard returns 0 (falsy) for the first card, indistinguishable from `false` rejection
|
||||
- `backend/app/scripting/engine.py:39`
|
||||
- `if (!addStoryCard(...))` misfires when the card list was empty.
|
||||
- Fix: return `storyCards.length` (1-based, always truthy) or `true`; document.
|
||||
|
||||
### 12. [fixed] scenario_id=0 truthiness bug in create_story_card
|
||||
- `backend/app/routers/story_cards.py:28`
|
||||
- `scenario_id or ...` picks the wrong owner when id is 0. Use `is not None`.
|
||||
|
||||
### 13. [fixed] test_connection 500s on non-dict JSON from /models
|
||||
- `backend/app/routers/settings.py:54`
|
||||
- Only ValueError is caught. `data.get`/`m.get` on non-dict input raises AttributeError, producing a 500 instead of `{ok:false}`.
|
||||
- Fix: catch (ValueError, AttributeError, TypeError) or validate shapes.
|
||||
|
||||
### 14. [fixed] AI Dungeon exports with `worldInformation` key lose all story cards silently
|
||||
- `backend/app/routers/scenarios.py:133`
|
||||
- Import reads only `storyCards`/`worldInfo`. `worldInformation` is in _IGNORED_KEYS, so it is dropped without being reported.
|
||||
- Fix: accept `worldInformation` as a card source too.
|
||||
|
||||
### 15. [fixed] Shared debounce timer loses edits (Play PlotPanel)
|
||||
- `frontend/src/pages/Play.jsx:28`
|
||||
- One `saveTimer` is shared by all plot fields and story-card saves. Editing a second field within 600ms cancels the first pending PATCH, causing silent data loss.
|
||||
- Fix: per-key timers (e.g. a Map keyed by field/card id).
|
||||
|
||||
### 16. [fixed] Same shared-debounce data loss in ScenarioEditor
|
||||
- `frontend/src/pages/ScenarioEditor.jsx:22`
|
||||
- Same fix as #15.
|
||||
|
||||
### 17. [fixed] Continue button silently discards typed input text
|
||||
- `frontend/src/pages/Play.jsx:469`
|
||||
- Clicking Continue with text in the box sends type 'continue' (the backend ignores the text) and clears the input.
|
||||
- Fix: don't clear input on continue (or treat non-empty input as a normal send).
|
||||
|
||||
### 18. [fixed] retry() optimistically deletes last AI action with no rollback on failure
|
||||
- `frontend/src/pages/Play.jsx:475`
|
||||
- A failed retry (409 or network error) leaves the UI missing an action that still exists server-side.
|
||||
- Fix: restore the removed action in the catch path (or only remove on first stream event).
|
||||
|
||||
### 19. [fixed] Settings test()/save() have no error handling → stuck on "Testing…"
|
||||
- `frontend/src/pages/Settings.jsx:66`
|
||||
- A rejection leaves `{pending:true}` forever and produces an unhandled rejection.
|
||||
- Fix: try/catch → setTestResult({ok:false, error:msg}).
|
||||
|
||||
### 20. [fixed] InsightsPanel race: slow earlier request overwrites newer report
|
||||
- `frontend/src/pages/Play.jsx:302`
|
||||
- There is no staleness guard, so a slow getAdventureContext call can overwrite a newer action snapshot.
|
||||
- Fix: track a request id or cancelled flag in the effect.
|
||||
|
||||
### 21. [fixed] extractPlaceholders ignores ${...} in story-card trigger keys
|
||||
- `frontend/src/pages/Scenarios.jsx:50`
|
||||
- The backend fills placeholders in card.keys, but the modal never prompts for those names, so a literal `${hero}` key never matches.
|
||||
- Fix: also scan card.keys when collecting placeholder names.
|
||||
|
||||
## Cleanup backlog (reuse / simplification / efficiency / altitude — not bugs, apply later)
|
||||
|
||||
- **R1** ~~three copies of the debounced-autosave handler.~~
|
||||
**Half applied in phase 17 (2026-08).** All three use
|
||||
`frontend/src/hooks/useDebouncedSave.js`. A shared StoryCardList component is
|
||||
still open.
|
||||
- **R2** `backend/seed_demo.py:228`: re-implements create_adventure. Call the router logic instead.
|
||||
- **R3** `backend/app/providers/openai_compatible.py:122`: complete() duplicates _request()'s body building. Add a `stream` param to _request().
|
||||
- **R4** ~~six copies of child-resource get, owner-check, and 404.~~
|
||||
**Applied in phase 17 (2026-08).** All 32 handlers take the `current_adventure`
|
||||
dependency from `routers/adventures/deps.py`.
|
||||
- **R5** `frontend/src/api.js:26`: streamSSE duplicates request()'s error extraction. Extract `throwIfNotOk(resp)`.
|
||||
- **S1** ~~`Settings.stream` is dead state (never read).~~
|
||||
**Applied in phase 17 (2026-08).** The column, both schema fields, and migration 65
|
||||
drop it. `R4` went with it: the ownership check is the `current_adventure` dependency.
|
||||
- **S2** `frontend/src/pages/Play.jsx:6`: MODES and PLAYER_TYPES are identical constants; lastIsAi/canUndo are computed twice.
|
||||
- **E1** `backend/app/context/builder.py:119`: joins and tokenizes the entire adventure history every turn for the trigger window. Walk reversed(actions) until budget instead.
|
||||
- **E2** `backend/app/scripting/pipeline.py:77`: rebuilds full history dicts, JSON, and a blocking commit per script per hook. Build once per hook, slice to HISTORY_WINDOW first, and commit once.
|
||||
- **E3** `frontend/src/pages/Play.jsx:561`: every SSE chunk re-renders all action rows. Isolate streaming text in a child component or React.memo rows.
|
||||
- **E4** `backend/app/context/builder.py:48`: Section.tokens is uncached, so the whole context gets tokenized 2 to 3 times per turn. Cache counts and sum sections.
|
||||
- **E5** `frontend/src/pages/Play.jsx:490`: the keydown effect has no dependency array, so the listener is re-registered every render.
|
||||
- **E6** `backend/app/memorybank.py:182`: catch-up summarization awaits blocks sequentially. Gather independent blocks instead.
|
||||
- **A1** ~~no UniqueConstraint('adventure_id','index'); index allocation is ad-hoc per writer.~~
|
||||
**Overtaken by phase 14 (2026-08).** `index` is a legacy column that nothing reads: ordering
|
||||
is `(branch_id, depth)` now, allocated in one place (`tree.place_action`). The column is kept
|
||||
unread for one release and then dropped, so a constraint on it would apply to a column that
|
||||
no longer does anything.
|
||||
- **A2** `backend/app/providers/openai_compatible.py:45`: CHAT_CONTINUE_HINT is appended below the budgeting layer. Assemble prompts in the context builder instead.
|
||||
- **A3** `backend/app/routers/adventures.py:390`: import endpoints hand-coerce raw dicts. Use a Pydantic bundle schema.
|
||||
- **A4** `backend/app/routers/adventures.py:207`: onModelContext flattens (system, story) and ships everything as user content if modified. Pass structure through the hook instead.
|
||||
- **A5** `frontend/src/pages/Home.jsx:87`: the client appends 'Z' to naive datetimes. Emit ISO-8601 with an offset from the API instead.
|
||||
@@ -1,16 +0,0 @@
|
||||
# React + Vite
|
||||
|
||||
This template provides a minimal setup to get React working in Vite with HMR and some Oxlint rules.
|
||||
|
||||
Currently, two official plugins are available:
|
||||
|
||||
- [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react) uses [Oxc](https://oxc.rs)
|
||||
- [@vitejs/plugin-react-swc](https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react-swc) uses [SWC](https://swc.rs/)
|
||||
|
||||
## React Compiler
|
||||
|
||||
The React Compiler is not enabled on this template because of its impact on dev & build performances. To add it, see [this documentation](https://react.dev/learn/react-compiler/installation).
|
||||
|
||||
## Expanding the Oxlint configuration
|
||||
|
||||
If you are developing a production application, use TypeScript with type-aware lint rules enabled. See the [TS template](https://github.com/vitejs/vite/tree/main/packages/create-vite/template-react-ts) for how to integrate TypeScript and Oxlint's TypeScript-related rules in your project.
|
||||
@@ -5,8 +5,8 @@
|
||||
* imports below.
|
||||
*
|
||||
* The page component still owns all of the session state. Lifting it into a
|
||||
* `usePlaySession` hook waits for Stage 5 of `plan/17-refactor.md`, which adds
|
||||
* a frontend test runner. Moving eighteen `useState` calls and seven
|
||||
* `usePlaySession` hook waits for the frontend test runner M8 adds
|
||||
* (`planning/BUILD-MILESTONES.md`). Moving eighteen `useState` calls and seven
|
||||
* `useEffect` calls is a rewrite rather than a move, and nothing would catch a
|
||||
* mistake in it today.
|
||||
*/
|
||||
|
||||
@@ -1,122 +0,0 @@
|
||||
# AI D&D — Local AI Dungeon Clone: Plan Overview
|
||||
|
||||
> **[STATUS.md](STATUS.md) — where things stand and what to pick up next.** Read that
|
||||
> first; this file is the shape of the project, not its current state.
|
||||
|
||||
A locally hosted web app replicating AI Dungeon's core experience: scenarios, adventures,
|
||||
AI-driven storytelling, AI Dungeon-style memory/context management, JavaScript scripting
|
||||
(compatible with real AI Dungeon scripts), and full transparency into what is sent to the AI.
|
||||
|
||||
## Confirmed decisions
|
||||
|
||||
| Area | Decision |
|
||||
|---|---|
|
||||
| Backend | Python — FastAPI + SQLAlchemy + SQLite (single-user, local) |
|
||||
| Frontend | React SPA (Vite), dark AI Dungeon-like theme |
|
||||
| AI provider | Provider-agnostic adapter layer; first adapter: **OpenAI-compatible** (`/v1/chat/completions`) — covers Ollama, LM Studio, OpenAI, OpenRouter, vLLM, Groq. Endpoint URL, API key, model name all configurable at runtime. |
|
||||
| Scripting | **JavaScript, AI Dungeon-compatible** (`onInput` / `onModelContext` / `onOutput` modifiers, shared `state`, `worldEntries` API) via an embedded JS engine (quickjs / py-mini-racer). Real AI Dungeon scripts should import and run. |
|
||||
| Import/export | AI Dungeon-compatible formats for scripts and scenarios; JSON export/import for everything. |
|
||||
|
||||
## Architecture at a glance
|
||||
|
||||
```
|
||||
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||||
├─ routers/ (scenarios, adventures, actions, scripts, settings, insights)
|
||||
├─ models/ (SQLAlchemy: Scenario, Adventure, Action, StoryCard, Script, Settings)
|
||||
├─ context/ (prompt assembly: memory, author's note, world info, history budget)
|
||||
├─ scripting/ (JS sandbox, AI Dungeon API surface, per-adventure state)
|
||||
├─ providers/ (base adapter + openai_compatible.py; streaming)
|
||||
└─ data.db (SQLite)
|
||||
```
|
||||
|
||||
## Core domain model
|
||||
|
||||
- **Scenario** — template: title, description, opening prompt (with `${placeholders}`), memory,
|
||||
author's note, story cards (world info), attached scripts, tags.
|
||||
- **Adventure** — a playthrough created from a scenario (or blank). Owns its own copy of memory,
|
||||
author's note, story cards, script state, and the action list.
|
||||
- **Action** — one entry in the story: type (`do` / `say` / `story` / `continue` / AI output),
|
||||
text, timestamp, plus the **context snapshot** (exact prompt sent to the AI) for Insights.
|
||||
- **Story Card / World Info** — keys (comma-separated keywords), entry text, optional type/notes.
|
||||
Injected into context only when a key matches recent story text.
|
||||
- **Script** — JS source per hook (input / context / output modifier), attachable to scenarios;
|
||||
copied into adventures with persistent `state`.
|
||||
|
||||
## The turn pipeline (heart of the app)
|
||||
|
||||
```
|
||||
player input
|
||||
→ onInput script modifier
|
||||
→ store player action
|
||||
→ assemble context: [AI instructions] + [plot essentials] + [story summary]
|
||||
+ [triggered story cards ("World Lore:")] + [story history, token-budgeted]
|
||||
+ [author's note inserted N lines from the end] + [player action]
|
||||
→ onModelContext script modifier
|
||||
→ snapshot context (Insights)
|
||||
→ provider adapter → AI (streamed)
|
||||
→ onOutput script modifier
|
||||
→ store AI action → render
|
||||
```
|
||||
|
||||
## Phases
|
||||
|
||||
1. **[Phase 1 — Foundation](01-phase-foundation.md)**: repo scaffold, FastAPI + SQLite models,
|
||||
React shell, scenario/adventure CRUD, settings (endpoint config).
|
||||
2. **[Phase 2 — Play loop + AI](02-phase-play-loop.md)**: provider adapter with streaming,
|
||||
Do/Say/Story/Continue, Retry/Undo/Edit, the adventure play screen.
|
||||
3. **[Phase 3 — Context engine + Insights](03-phase-context-insights.md)**: memory, author's note,
|
||||
story cards with keyword triggering, token budgeting, per-turn prompt snapshots + Insights UI.
|
||||
4. **[Phase 4 — Scripting](04-phase-scripting.md)**: embedded JS sandbox, AI Dungeon scripting API,
|
||||
script editor, script + scenario import/export (AI Dungeon-compatible).
|
||||
5. **[Phase 5 — Polish](05-phase-polish.md)**: AI Dungeon-like theming pass, placeholders on
|
||||
scenario start, adventure export/import, quality-of-life and hardening.
|
||||
6. **[Phase 6 — Auto Summarization + Memory Bank](06-phase-memory-bank.md)** *(optional)*:
|
||||
modern AI Dungeon memory system — AI-generated memories every 6 actions, Story Summary every
|
||||
15, embedding-based retrieval of relevant memories into context.
|
||||
|
||||
Each phase ends with the app runnable and testable end-to-end.
|
||||
|
||||
## Public release (phases 7–10)
|
||||
|
||||
Goal: public GitHub repo + hosted multi-user deployment, linkable from resume/website.
|
||||
|
||||
| Area | Decision |
|
||||
|---|---|
|
||||
| License | MIT |
|
||||
| Auth | Email + password, **optional** — guest sessions play instantly, register to keep data |
|
||||
| Signup | Open (rate-limited) |
|
||||
| LLM keys | BYOK per user + shared server-funded demo key (capped, details TBD in Phase 8) |
|
||||
| Database (hosted) | **TBD — ask at start of Phase 9** (SQLite-on-disk vs Postgres) |
|
||||
| Hosting | Render (tier TBD in Phase 10) |
|
||||
| Domain | Platform URL (custom domain later, optional) |
|
||||
| README media | Deferred — text-only first, screenshots/GIF in a later pass |
|
||||
|
||||
Open questions are recorded at the top of each phase file under "Ask before implementing".
|
||||
|
||||
7. **[Phase 7 — Public repo & portability](07-phase-public-repo.md)**: MIT license, portfolio
|
||||
README, Dockerfile + compose, cross-platform run instructions, publish to GitHub.
|
||||
8. **[Phase 8 — Optional accounts & multi-user](08-phase-accounts.md)** *(the big one)*:
|
||||
guest-first sessions, optional email+password upgrade, per-user data scoping across all
|
||||
routers/tables, per-user encrypted BYOK settings, shared demo key with caps.
|
||||
9. **[Phase 9 — Production hardening](09-phase-hardening.md)**: env-var config, quickjs
|
||||
time/memory limits, rate limiting, size/row caps, locked-down debug surface, production
|
||||
serving, database decision.
|
||||
10. **[Phase 10 — Deploy & publish](10-phase-deploy.md)**: Render blueprint + deploy, seeded
|
||||
demo scenarios, live smoke test, resume/website links and blurb.
|
||||
|
||||
## Post-launch
|
||||
|
||||
- **[State revert + retry fix](11-state-revert-and-retry-fix.md)**: undo/retry roll the shared
|
||||
`script_state` back; per-action `state_before` snapshots; undo concurrency lock.
|
||||
- **[Phase 12 — RPG world state](12-phase-rpg-world-state.md)**: structured world/player/NPC stats
|
||||
+ milestones per scenario (`stat_schema`); the AI proposes deltas, a Python engine clamps them
|
||||
(min/max, per-turn cap, cooldown, sticky milestones); band descriptions keep the model honest;
|
||||
World State drawer + Insights delta report; reuses the Phase 11 undo/retry snapshot.
|
||||
- **[Memory-bank embedding cost](13-memory-embedding-cost.md)** *(round three of the egress
|
||||
work)*: a turn fetched the whole memory bank's vectors to pick five — 96% of everything it
|
||||
read. Packed float32 + an in-process cache + SQL-side filtering; a played turn went 6.4 MB
|
||||
to 123 kB. Includes the byte-meter harness (`backend/tools/`). One step left, see STATUS.
|
||||
- **[Phase 14 — Story tree](14-phase-story-tree.md)** *(designed, not started)*: the linear
|
||||
action list becomes a branching tree, so a retry is a sibling rather than a rewrite and both
|
||||
paths survive. The real argument is bug elimination — seven bug classes trace to "the story
|
||||
is a mutable list" and all disappear when nothing is rewritten in place.
|
||||
@@ -1,38 +0,0 @@
|
||||
# Phase 1 — Foundation
|
||||
|
||||
**Goal:** runnable skeleton — backend serving a database-backed API, frontend shell with
|
||||
navigation, scenario & adventure CRUD, and a settings page for the AI endpoint. No AI calls yet.
|
||||
|
||||
## Backend
|
||||
|
||||
- [x] Project scaffold: `backend/` with FastAPI app, uvicorn entrypoint, `requirements.txt`
|
||||
(fastapi, uvicorn, sqlalchemy, pydantic, httpx, tiktoken; quickjs deferred to phase 4).
|
||||
- [x] SQLite via SQLAlchemy; auto-create `data.db` on first run.
|
||||
- [x] Models:
|
||||
- `Scenario`: id, title, description, prompt, memory, authors_note, tags, created/updated.
|
||||
- `StoryCard`: id, owner (scenario or adventure), keys, entry, type, title, description.
|
||||
- `Adventure`: id, scenario_id (nullable), title, memory, authors_note, script_state (JSON),
|
||||
created/updated.
|
||||
- `Action`: id, adventure_id, index, type (`do|say|story|continue|ai|start`), text,
|
||||
context_snapshot (JSON, nullable), created.
|
||||
- `Script`: id, name, description, input_js, context_js, output_js, library_js.
|
||||
- `Settings`: single row — endpoint_url, api_key, model, temperature, max_output_tokens,
|
||||
context_token_budget, api_mode (`chat|completion`).
|
||||
- [x] Routers: CRUD for scenarios, adventures (+ create-from-scenario copying memory/AN/cards),
|
||||
story cards, settings. Consistent JSON errors.
|
||||
- [x] CORS for the Vite dev server; production mode serves built frontend as static files.
|
||||
|
||||
## Frontend
|
||||
|
||||
- [x] Vite + React scaffold in `frontend/`; router with pages: Home (adventure list),
|
||||
Scenarios (list + editor), Play (placeholder), Settings.
|
||||
- [x] Dark base theme (AI Dungeon-like: near-black background, serif story font, gold/teal accent).
|
||||
- [x] Scenario editor: title, description, prompt, memory, author's note, story card list editor.
|
||||
- [x] Settings page: endpoint URL, API key, model name, sampling params; "Test connection" button
|
||||
(backend proxies a trivial request — wired for real in Phase 2, stub now).
|
||||
- [x] "New adventure" flow: pick scenario (or blank) → creates adventure → navigates to Play page.
|
||||
|
||||
## Exit criteria
|
||||
|
||||
Run `uvicorn` + `npm run dev`, create/edit/delete scenarios with story cards, start an adventure
|
||||
from one, see it listed on Home, and save endpoint settings — all persisted across restarts.
|
||||
@@ -1,45 +0,0 @@
|
||||
# Phase 2 — Play loop + AI provider
|
||||
|
||||
**Goal:** the core game is playable. Player acts, AI continues the story, streamed live.
|
||||
|
||||
## Provider adapter layer
|
||||
|
||||
- [x] `providers/base.py`: abstract `Provider` — `generate(prompt_parts, params) -> async stream of text`.
|
||||
Takes an assembled context object (system text + story text), so providers decide how to
|
||||
map it to their wire format.
|
||||
- [x] `providers/openai_compatible.py`:
|
||||
- Chat mode: system message carries instructions/memory; story history flows in as
|
||||
user/assistant text continuation framing suited to a chat endpoint. A configurable
|
||||
"narrator" system prompt frames the AI as a second-person storyteller continuing the text.
|
||||
- Streaming via SSE from the endpoint, re-streamed to the browser.
|
||||
- Optional raw completion mode (`/v1/completions`) for pure-continuation models.
|
||||
- [x] Errors surfaced cleanly (bad key, connection refused, model not found) with retry affordance.
|
||||
- [x] "Test connection" on Settings now real.
|
||||
|
||||
## Turn engine (`POST /adventures/{id}/actions`)
|
||||
|
||||
- [x] Input formatting per AI Dungeon conventions:
|
||||
- **Do** → `> You <text>` (normalized to second person, stripped punctuation as needed)
|
||||
- **Say** → `> You say "<text>"`
|
||||
- **Story** → raw text appended
|
||||
- **Continue** → no player text; AI just continues
|
||||
- [x] Simple context for this phase: opening prompt + full history, truncated from the top to the
|
||||
token budget (tiktoken count). Real context engine lands in Phase 3.
|
||||
- [x] Response streamed to the client via SSE; final text stored as an `ai` action.
|
||||
- [x] **Retry**: delete last AI action, regenerate with same input.
|
||||
- [x] **Undo/Erase**: delete last action pair (player + AI) or single action.
|
||||
- [x] **Edit**: PATCH any action's text in place.
|
||||
|
||||
## Play UI
|
||||
|
||||
- [x] Story view: continuous prose (not chat bubbles), player actions styled distinctly
|
||||
(`>` prefix, accent color), auto-scroll, streaming text renders token-by-token.
|
||||
- [x] Input bar with mode selector (Do / Say / Story) + Continue button; Enter to send.
|
||||
- [x] Per-turn controls: Retry, Undo, Edit (inline contenteditable or textarea swap).
|
||||
- [x] Loading/streaming state, error toast with retry.
|
||||
|
||||
## Exit criteria
|
||||
|
||||
Point Settings at any OpenAI-compatible endpoint (e.g. Ollama or LM Studio locally), start an
|
||||
adventure, and play a multi-turn story with all four input modes plus retry/undo/edit, with
|
||||
streaming output.
|
||||
@@ -1,62 +0,0 @@
|
||||
# Phase 3 — Context engine + Insights
|
||||
|
||||
**Goal:** AI Dungeon-grade context management, and full visibility into every prompt.
|
||||
|
||||
## Context assembly (`context/builder.py`)
|
||||
|
||||
Assembles the prompt each turn from AI Dungeon's **plot components**
|
||||
(per help.aidungeon.com/faq/the-memory-system):
|
||||
|
||||
```
|
||||
[AI Instructions] ↠behavioral guidance for the model (always included)
|
||||
[Plot Essentials] ↠key facts for constant recall — the classic "Memory" (always)
|
||||
[Story Summary] ↠running summary slot; manual in this phase, auto in Phase 6
|
||||
[Triggered Story Cards] ↠"World Lore: <entry>" for each triggered card (conditional)
|
||||
[Story history] ↠as many recent actions as fit the token budget
|
||||
[Author's Note] ↠injected N lines (default 3) before the end of history
|
||||
[Latest player action] ↠(+ script frontMemory right after it, Phase 4)
|
||||
```
|
||||
|
||||
- [x] **AI Instructions / Plot Essentials / Story Summary / Author's Note**: adventure-level
|
||||
free-text fields, always included, editable mid-adventure from the side panel.
|
||||
- [x] **Story cards** — five fields per official docs: **Type** (organizational, not sent to AI),
|
||||
**Name** (not sent to AI), **Entry** (sent when triggered), **Triggers**, **Notes** (not sent).
|
||||
- Triggers: comma-separated words/phrases; **case-insensitive but space-sensitive**;
|
||||
**partial-word matching** (`boat` triggers on `boats`); matched against both player input
|
||||
and AI output in the recent-story window.
|
||||
- Not instant: a card triggered mid-response only enters context on the *next* turn; once
|
||||
triggered, stays active while the triggering text remains in the context window.
|
||||
- Triggered entries injected once each, prefixed `World Lore:`; story cards are the **first
|
||||
component dropped** when context is full.
|
||||
- Editable per-adventure (copied from scenario at creation). Soft-cap sanity limit (AI Dungeon
|
||||
allows 5,000/adventure).
|
||||
- [x] **Author's Note**: inserted near the end (strongest steering position), formatted
|
||||
`[Author's note: <text>]`.
|
||||
- [x] **Token budgeting** with tiktoken: always-included components (AI Instructions, Plot
|
||||
Essentials, Story Summary, Author's Note) reserved first; story cards get a capped share
|
||||
and are dropped first when over budget; remainder goes to story history (newest first).
|
||||
Budget = `context_token_budget` setting.
|
||||
- [x] Slots for script-provided memory overrides (Phase 4): `state.memory.context` (prepended),
|
||||
`state.memory.authorsNote` (replaces/augments author's note), `state.memory.frontMemory`
|
||||
(inserted immediately after the latest player action).
|
||||
- [x] Builder returns a structured `ContextReport`: ordered sections, each with source label,
|
||||
text, token count; plus totals and a list of triggered cards (and which keyword fired).
|
||||
|
||||
## Insights
|
||||
|
||||
- [x] Every AI turn stores its `ContextReport` on the `Action` row (`context_snapshot`).
|
||||
- [x] `GET /adventures/{id}/actions/{id}/context` returns it.
|
||||
- [x] **Insights panel** in Play UI (drawer/tab):
|
||||
- Exact final prompt text as sent, sectioned and color-coded (memory / world info / history /
|
||||
author's note / input), with per-section token counts and total vs budget.
|
||||
- Which story cards triggered and on which keyword; which history got cut off.
|
||||
- Viewable for the *upcoming* turn (dry-run endpoint: "what would be sent now") and for any
|
||||
past AI action.
|
||||
- [x] Adventure side panel: edit memory, author's note, story cards mid-game (AI Dungeon's
|
||||
right-hand panel equivalent).
|
||||
|
||||
## Exit criteria
|
||||
|
||||
Create a card with key `dragon`; mention a dragon in play and see the card enter the context in
|
||||
the Insights panel (and influence the AI); verify memory and author's note appear in the snapshot
|
||||
in the right positions; long adventures visibly trim oldest history within budget.
|
||||
@@ -1,84 +0,0 @@
|
||||
# Phase 4 — Scripting (AI Dungeon-compatible JavaScript)
|
||||
|
||||
**Goal:** real AI Dungeon scripts import and run: the three modifier hooks, persistent `state`,
|
||||
and the scripting API surface.
|
||||
|
||||
## JS runtime
|
||||
|
||||
- [x] Embed a JS engine in Python: **quickjs** (preferred; check Windows wheel availability at
|
||||
implementation time; fallback: py-mini-racer, or Node subprocess as last resort).
|
||||
- [x] Sandbox matching AI Dungeon's documented limits: each hook runs **isolated**, **16 MB
|
||||
memory cap**, **2-second timeout**; no filesystem/network/process access; script errors
|
||||
captured and surfaced in the UI, never crash a turn.
|
||||
|
||||
## AI Dungeon scripting model (compatibility target)
|
||||
|
||||
*(Per official docs: help.aidungeon.com/faq/how-do-i-write-scripts-and-use-scripting)*
|
||||
|
||||
Three lifecycle hooks — `onInput`, `onModelContext`, `onOutput`. Each script defines a modifier
|
||||
and **must call it as its last line**:
|
||||
|
||||
```javascript
|
||||
const modifier = (text) => {
|
||||
// script logic
|
||||
return { text, stop }
|
||||
}
|
||||
modifier(text)
|
||||
```
|
||||
|
||||
- **onInput** — modifies player input before context construction.
|
||||
- **onModelContext** — modifies the assembled text sent to the model.
|
||||
- **onOutput** — modifies the model output before it is shown/stored.
|
||||
- **Shared Library** — code prepended to all three slots.
|
||||
|
||||
Return contract:
|
||||
- [x] `{ text, stop }`; `stop: true` from onInput prevents the AI call.
|
||||
- [x] Empty-string `text` from onInput/onOutput → user-facing error (replicate this behavior).
|
||||
|
||||
Globals provided (exact names from docs):
|
||||
|
||||
- [x] `text` — hook input (player input / context / AI response respectively).
|
||||
- [x] `state` — persisted per adventure across turns (`Adventure.script_state`); includes
|
||||
`state.memory`, `state.message` (shown as a UI notice), `state.placeholders`.
|
||||
- [x] `state.memory` slots: `context` (prepended to context), `authorsNote` (near end, before
|
||||
latest response), `frontMemory` (inserted right after the player's input).
|
||||
- [x] `history` — array of recent actions: `{ text, rawText, type }`.
|
||||
- [x] `storyCards` — array of `{ id, keys, entry, type }`, backed by the adventure's story cards.
|
||||
- [x] Story card functions: `addStoryCard(keys, entry, type)` → index (or `false` on duplicate),
|
||||
`updateStoryCard(index, keys, entry, type)` and `removeStoryCard(index)` → throw if absent.
|
||||
- [x] Legacy aliases for older scripts: `worldInfo` / `worldEntries`, `addWorldEntry`,
|
||||
`updateWorldEntry`, `removeWorldEntry` mapped onto the storyCards implementation.
|
||||
- [x] `info` — `{ actionCount, characterNames, memoryLength, maxChars }`.
|
||||
- [x] `log(message)` / `console.log` — captured per turn, shown in a script log panel.
|
||||
|
||||
## Pipeline integration
|
||||
|
||||
```
|
||||
player input → INPUT modifier → format & store
|
||||
context build → CONTEXT modifier → (snapshot includes pre- and post-script versions in Insights)
|
||||
AI response → OUTPUT modifier → store & render
|
||||
```
|
||||
|
||||
- [x] Insights (Phase 3) extended: show context before vs after the context modifier (diff view),
|
||||
and script log output per turn.
|
||||
|
||||
## Script management UI
|
||||
|
||||
- [x] Scripts page: create/edit scripts with a code editor (CodeMirror), one tab per slot
|
||||
(Library / Input / Context / Output), description field.
|
||||
- [x] Attach scripts to scenarios; adventures inherit at creation. Enable/disable per adventure.
|
||||
- [x] Test-run a script against sample text without an AI call.
|
||||
|
||||
## Import / Export
|
||||
|
||||
- [x] **Scripts**: export/import as JSON bundle `{ name, library, input, context, output }` and
|
||||
as raw `.js` files per slot (matching how AI Dungeon scripts circulate — paste or file).
|
||||
- [x] **Scenarios**: export/import JSON including prompt, memory, author's note, story cards,
|
||||
and attached scripts. Accept AI Dungeon scenario export JSON where format is known;
|
||||
map fields best-effort and report anything unmapped.
|
||||
|
||||
## Exit criteria
|
||||
|
||||
Paste a real AI Dungeon script (e.g. a simple input modifier + state counter + world entry
|
||||
manipulation) and it runs unmodified across turns; state persists; export a scenario with scripts,
|
||||
re-import it into a fresh database, and play it.
|
||||
@@ -1,38 +0,0 @@
|
||||
# Phase 5 — Polish & quality of life
|
||||
|
||||
**Goal:** the app feels like AI Dungeon — cohesive dark UI, smooth flows, safe data handling.
|
||||
|
||||
## UI/UX pass
|
||||
|
||||
- [x] Theming: refined dark palette, serif story typography, subtle textures/gradients à la
|
||||
AI Dungeon; consistent buttons, panels, modals; responsive layout.
|
||||
- [x] Home: adventure cards with scenario name, last-played time, action count; search/filter;
|
||||
scenario gallery with tags.
|
||||
- [x] Play screen: collapsible right side panel (Memory / Cards / Scripts / Insights tabs),
|
||||
keyboard shortcuts (Enter send, Ctrl+Z undo, Ctrl+R retry), smooth streaming autoscroll
|
||||
that pauses when the user scrolls up.
|
||||
- [x] Scenario **placeholders**: `${Character name}` style variables in prompt/memory prompt the
|
||||
player for values when starting an adventure (AI Dungeon behavior).
|
||||
|
||||
## Data & robustness
|
||||
|
||||
- [x] Adventure export/import (full JSON: actions, memory, cards, script state) — backup/share.
|
||||
- [x] Delete confirmations (trash/soft-delete skipped — plain confirm dialogs).
|
||||
- [x] SQLite migrations story (versioned schema bootstrap via PRAGMA user_version —
|
||||
`backend/app/migrations.py`).
|
||||
- [x] Request logging + a debug page tailing recent provider requests/responses (bodies redacted
|
||||
of API key) — "Recent AI requests" on the Settings page.
|
||||
- [x] Graceful handling: provider timeout/cancel (stop generation button), concurrent turn lock
|
||||
per adventure.
|
||||
|
||||
## Nice-to-haves (only if time/interest)
|
||||
|
||||
- [ ] Multiple provider profiles with quick switching (e.g. local Ollama vs OpenRouter).
|
||||
- [ ] Per-scenario generation params overriding global settings.
|
||||
- [ ] Retry with "give me something different" (temperature bump / anti-repeat nudge).
|
||||
- [ ] Basic light theme toggle.
|
||||
|
||||
## Exit criteria
|
||||
|
||||
A friend could sit down at `localhost`, start a scenario with placeholders, play comfortably,
|
||||
peek at Insights, and you can back up / restore everything via export files.
|
||||
@@ -1,47 +0,0 @@
|
||||
# Phase 6 — Auto Summarization + Memory Bank (optional)
|
||||
|
||||
**Goal:** replicate modern AI Dungeon's Memory System (per
|
||||
help.aidungeon.com/faq/the-memory-system): AI-generated memories, a running Story Summary, and
|
||||
embedding-based retrieval. This phase makes extra AI calls (summarization + embeddings), so it is
|
||||
opt-in per adventure and gated on the endpoint supporting it.
|
||||
|
||||
## Auto Summarization
|
||||
|
||||
- [x] **Memories**: every 6 actions (starting at action 12), summarize that block of
|
||||
player actions + AI responses into a short "memory" via a background AI call
|
||||
(same provider, cheap/configurable model override).
|
||||
- [x] **Story Summary**: every 15 actions, update the Story Summary plot component — a running
|
||||
overview of the plot — folding in recent memories; compress it when it grows too long.
|
||||
- [x] Story Summary stays **manually editable**; user edits inform future updates
|
||||
(they are the base text for the next summarization pass) but are never overwritten silently.
|
||||
- [x] Summarization failures are non-fatal: log, retry next interval.
|
||||
(Failed calls appear on the debug page; cursors only advance on success, so the
|
||||
next turn retries. Implementation: `backend/app/memorybank.py`.)
|
||||
|
||||
## Memory Bank
|
||||
|
||||
- [x] Store each memory with an **embedding vector** (OpenAI-compatible `/v1/embeddings`;
|
||||
embedding model configurable in Settings; feature disabled if unavailable).
|
||||
- [x] Each turn, embed the recent story text and rank memories by cosine similarity;
|
||||
inject the top-K "Used Memories" into context as their own component
|
||||
(between Story Summary and story cards in the layout). Pinned memories are always
|
||||
included; top-K is a setting (default 5).
|
||||
- [x] Configurable bank capacity (AI Dungeon tiers: 25–400; ours: a setting, default 200);
|
||||
when full, evict least-recently-used/least-retrieved memories ("Forgotten Memories").
|
||||
- [x] SQLite storage for vectors (JSON blob + in-memory cosine ranking — pure Python,
|
||||
no numpy needed at this scale).
|
||||
|
||||
## UI
|
||||
|
||||
- [x] Memory Bank panel: list memories (used / idle / forgotten), edit or delete, pin favorites,
|
||||
see which memories were retrieved for a given turn (🔍 on an AI action → Insights snapshot).
|
||||
- [x] Insights integration: retrieved memories shown as a context section with similarity scores.
|
||||
- [x] Adventure settings: toggle auto-summarization / memory bank (per adventure, in the Memory
|
||||
panel); summary + embedding models are chosen globally in Settings (deliberate
|
||||
simplification — one endpoint config for the whole app).
|
||||
|
||||
## Exit criteria
|
||||
|
||||
Play a 40+ action adventure: memories appear every 6 actions, the Story Summary updates every 15,
|
||||
an early-game fact that scrolled out of the raw history gets retrieved via the Memory Bank when it
|
||||
becomes relevant again, and Insights shows exactly which memories were injected and why.
|
||||
@@ -1,62 +0,0 @@
|
||||
# Phase 7 — Public repo & portability
|
||||
|
||||
**Goal:** make the repo public-worthy and runnable by anyone on any OS, so the GitHub link is
|
||||
immediately usable on a resume — before any hosted-deployment work.
|
||||
|
||||
## Decisions (confirmed)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| License | **MIT** |
|
||||
| README media (screenshots/GIF) | **Skip for now** — text-only README; visuals in a later pass |
|
||||
|
||||
**Repo name (decided): `AI-DnD`.** Ask before publishing: whether the existing commit
|
||||
history/messages are fine to publish as-is.
|
||||
|
||||
## Repo hygiene
|
||||
|
||||
- [x] Add `LICENSE` (MIT, current year, Parth Thakkar).
|
||||
- [x] Verify no secrets or user data are tracked (`openrouter_key.env`, `data.db` — already
|
||||
gitignored and never committed; re-verify before push).
|
||||
- [x] Add `backend/.env.example` documenting every env var the app reads (grows in Phase 9).
|
||||
(Currently just `AIDND_DB_PATH`, added to `database.py` for Docker/hosted volumes.)
|
||||
- [x] Decide what to do with `CODE_REVIEW_FINDINGS.md` and `plan/` — keep (shows process, good
|
||||
for a portfolio) — just give them a one-line mention in the README.
|
||||
|
||||
## README rewrite (portfolio-grade, text-only)
|
||||
|
||||
- [x] Pitch paragraph: what it is, what makes it interesting (AI Dungeon-compatible scripting,
|
||||
memory bank with embeddings, full prompt transparency/Insights, provider-agnostic).
|
||||
- [x] Feature list with pointers into the code (scripting engine, context builder, memory bank).
|
||||
- [x] Architecture diagram (reuse/refresh the one in `plan/00-OVERVIEW.md`).
|
||||
- [x] Setup instructions for **Windows (start.ps1), macOS/Linux (manual), and Docker**.
|
||||
- [x] "Bring your own model" section: Ollama / LM Studio / OpenRouter free models — emphasize it
|
||||
runs fully free.
|
||||
- [x] Placeholder section for screenshots/GIF (added in a later pass).
|
||||
|
||||
## Docker (one-command run for non-Windows users)
|
||||
|
||||
- [x] `Dockerfile`: multi-stage — build frontend (`npm run build`), then Python image serving
|
||||
FastAPI with the built SPA mounted (SPA fallback already exists in `app/main.py`).
|
||||
(3 stages: node build → pip wheel build with gcc for quickjs → slim runtime.)
|
||||
- [x] `docker-compose.yml`: single service, named volume mounted at `/data`
|
||||
(`AIDND_DB_PATH=/data/data.db`), port mapping.
|
||||
- [x] `start.sh` for macOS/Linux dev parity with `start.ps1` (optional, nice-to-have).
|
||||
- [x] Test: `docker compose up` from a clean clone → app works at `http://localhost:8000`.
|
||||
Done 2026-07-06 in WSL2 Ubuntu (Docker Engine installed there for this): health + SPA +
|
||||
deep links 200, scenario created via API survives a full `compose down`/`up` (named
|
||||
volume). Two real bugs found and fixed: package-lock.json was missing top-level
|
||||
`@emnapi/*` entries (regenerated with npm 11.18) and the build image needed node 24 to
|
||||
match the npm-11 lockfile. Note: from Windows, reach a WSL-hosted container via the WSL
|
||||
IP (`hostname -I`) — localhost forwarding didn't apply; irrelevant on normal hosts.
|
||||
|
||||
## Publish
|
||||
|
||||
- [ ] Create public GitHub repo (`gh repo create`), push `main`.
|
||||
- [ ] Add repo description, topics (`ai-dungeon`, `fastapi`, `react`, `llm`, `interactive-fiction`).
|
||||
- [ ] Confirm the GitHub rendering of README looks right.
|
||||
|
||||
## Exit criteria
|
||||
|
||||
A stranger on macOS with Docker installed can clone the repo, run one command, open the app,
|
||||
paste an OpenRouter free-tier key, and play an adventure — without asking you anything.
|
||||
@@ -1,81 +0,0 @@
|
||||
# Phase 8 — Optional accounts & multi-user ✅ (implemented 2026-07-06, branch `phase-8-accounts`)
|
||||
|
||||
**Goal:** turn the single-user app into a multi-user one where **accounts are optional**:
|
||||
a visitor can start playing instantly as a guest, and can register (email + password) at any
|
||||
point to keep their adventures across devices/browsers. This is the largest phase — it touches
|
||||
every router and most tables.
|
||||
|
||||
## Decisions (confirmed)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Auth method | **Email + password, optional** — guest sessions work without an account |
|
||||
| Signup policy | **Open signup** (rate-limited) |
|
||||
| LLM API keys | **BYOK + shared demo key** — users can paste their own key; users without one get limited turns on a server-funded key |
|
||||
| Demo key funding | **OpenRouter free models** (owner's key, `:free` whitelist; default `google/gemma-4-26b-a4b-it:free`) |
|
||||
| Demo turn cap | **20 successful turns/user/day** (failed provider calls don't count) |
|
||||
| Guest data retention | **Never delete** for v1 (no cleanup job; revisit if the DB grows) |
|
||||
| Password reset | **Skipped for v1** (no email provider; forgotten password = lost account) |
|
||||
| Login with an active guest session | Guest is **abandoned**, not merged (its data stays under the guest user) |
|
||||
| Memory bank on demo key | **Disabled** (no background AI calls on the server-funded key; visible note in the Memory panel/Insights) |
|
||||
|
||||
## Data model
|
||||
|
||||
- [x] `User` table: id, email (nullable — null means guest), password_hash (nullable),
|
||||
created_at, last_seen_at, is_guest flag, demo_turns_used + demo_turns_date.
|
||||
- [x] `user_id` FK on `Adventure`, `Scenario`, `Script`, `Settings`. Story cards/actions/
|
||||
memories inherit scope via their parent (ownership checks resolve the parent).
|
||||
- [x] Settings **per-user** (row per user_id, unique index). API key **encrypted at rest**
|
||||
(Fernet; key derived from `AIDND_SECRET_KEY` or auto-generated `secret.key` next to the
|
||||
DB). Key is write-only through the API (`has_api_key` instead of echoing it).
|
||||
- [x] Migrations 13–23: create "local user" id=1, assign all existing rows to it, unique
|
||||
index on settings.user_id; plus a Python bootstrap step that encrypts any plaintext
|
||||
api_key (`enc:` prefix marks encrypted values).
|
||||
- [x] Demo/starter scenarios: `user_id NULL` + `is_public` — everyone sees them read-only;
|
||||
`seed_demo.py` seeds the Sunken Crypt scenario as public (its scripts are unowned and
|
||||
ship with it; the sample adventure belongs to the local user).
|
||||
|
||||
## Auth & sessions
|
||||
|
||||
- [x] Guest flow: `GET /api/auth/me` with no/invalid cookie → creates guest User + signed
|
||||
long-lived httpOnly cookie (HMAC, `security.py`). Other endpoints 401 without a session;
|
||||
the frontend re-establishes via /me and retries once. No signup wall anywhere.
|
||||
- [x] Register upgrades the guest **in place** (same user_id — data kept). scrypt password
|
||||
hashing (stdlib, no extra dep).
|
||||
- [x] Login switches the session cookie to the account (guest abandoned). Logout clears it.
|
||||
- [x] Every router handler resolves `current_user`; every query filtered by user_id
|
||||
(scenarios/adventures/scripts/story-cards/settings; debug log is local-mode only since
|
||||
it's a global buffer).
|
||||
- [x] Rate limit on register/login: 10 attempts / 5 min per IP (in-memory).
|
||||
- [x] Local/self-hosted mode stays frictionless: auto-created local user, no login UI unless
|
||||
`AIDND_MULTI_USER=1`. Local installs and docker compose behave exactly as before.
|
||||
|
||||
## Shared demo key (BYOK fallback)
|
||||
|
||||
- [x] Env vars: `AIDND_DEMO_API_KEY`, `AIDND_DEMO_ENDPOINT_URL` (default OpenRouter),
|
||||
`AIDND_DEMO_MODELS` (comma whitelist), `AIDND_DEMO_TURNS_PER_DAY` (default 20).
|
||||
Demo only activates in multi-user mode.
|
||||
- [x] No API key configured → demo endpoint/key/whitelisted model; per-user per-day counter;
|
||||
429 with a friendly "add your own key in Settings" message when capped (checked before
|
||||
the turn starts so no orphaned player action).
|
||||
- [x] Memory bank + auto-summarization disabled on demo turns (decided: disable, not count).
|
||||
|
||||
## Frontend
|
||||
|
||||
- [x] Auth UI: Sign up / Log in modal (register default, toggle to login), "Playing as guest —
|
||||
sign up to keep your adventures" nudge in the header, account email + logout when
|
||||
registered. All hidden in local mode (`multi_user:false` from /me).
|
||||
- [x] `api.js`: 401 → GET /auth/me (new guest session) → retry once, for both JSON and SSE.
|
||||
- [x] Settings: demo banner ("Using the shared demo key — N of M free turns left today"),
|
||||
write-only API key field with Remove button, debug log hidden in multi-user mode.
|
||||
- [x] Public scenarios: "demo ✦" badge in the list; read-only editor (fieldset-disabled) with
|
||||
an explainer banner; Play/Export still available.
|
||||
|
||||
## Exit criteria — verified 2026-07-06
|
||||
|
||||
Two sessions (curl cookie jars + Chrome UI): each guest gets an isolated world; register
|
||||
mid-session keeps all data (same user id); logging in from the second session shows the same
|
||||
account data; duplicate email → 409; wrong password → 401; rate limiter kicks in. Demo cap
|
||||
returns 429 at 0 turns left. Migration tested on a copy of the real data.db (rows adopted by
|
||||
local user, api_key Fernet-encrypted and decrypts back to the original). Live OpenRouter turn
|
||||
through the encrypted-key path works in local mode. `vite build` + oxlint clean.
|
||||
@@ -1,93 +0,0 @@
|
||||
# Phase 9 — Production hardening
|
||||
|
||||
**Goal:** make the app safe and stable to expose to strangers on the internet: config via
|
||||
environment, resource limits on everything user-controlled, and a single-service production
|
||||
build.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Database | **Decide at start of this phase.** SQLite on a persistent disk (zero code change, but Render disks require the ~$7/mo starter tier) vs Postgres (free/cheap managed options, better resume talking point, needs SQLAlchemy URL + migration tweaks). Revisit with current Render pricing. |
|
||||
|
||||
**Ask before implementing:** the database choice above, and target monthly budget (drives
|
||||
Render tier: free tier sleeps after idle + has no persistent disk).
|
||||
|
||||
## Configuration
|
||||
|
||||
- [ ] All config via env vars with sane local defaults: `DATABASE_URL`, `SECRET_KEY`
|
||||
(sessions + API-key encryption), `MULTI_USER`, `CORS_ORIGINS`, demo-key vars (Phase 8),
|
||||
port/host. Document each in `backend/.env.example`.
|
||||
- [ ] Fail fast on missing `SECRET_KEY` when `MULTI_USER=true`.
|
||||
|
||||
## Abuse & resource limits
|
||||
|
||||
- [ ] **quickjs limits**: per-execution time limit and memory limit on the scripting engine
|
||||
(`scripting/engine.py`) — user-submitted JS must not be able to hang or OOM the server.
|
||||
- [ ] Rate limiting on expensive endpoints (turn generation, script run, auth) — per-user and
|
||||
per-IP (e.g. `slowapi`).
|
||||
- [ ] Request size limits (script source length, memory/story-card text lengths, action text).
|
||||
- [ ] Cap per-user row counts (adventures, scenarios, scripts, story cards) with friendly errors.
|
||||
- [ ] Audit debug router (`routers/debug.py`) and `/docs`: admin-only or disabled when
|
||||
`MULTI_USER=true` — debug log may contain other users' prompts.
|
||||
|
||||
## Production serving
|
||||
|
||||
- [ ] Single service: FastAPI serves the built SPA (fallback already exists) — verify the Docker
|
||||
image from Phase 7 is production-ready (no `--reload`, multiple workers or async-safe
|
||||
single worker; check SQLite + multiple workers interaction before choosing).
|
||||
- [ ] CORS locked to the deployed origin (moot if same-origin single service — verify).
|
||||
- [ ] Security headers middleware; cookies `Secure` + `SameSite`.
|
||||
- [ ] Streaming (SSE) works behind Render's proxy — verify no buffering issues.
|
||||
- [ ] Structured logging; scrub API keys from all logs and the debug page.
|
||||
|
||||
## Database (after decision)
|
||||
|
||||
- [ ] If Postgres: swap `DATABASE_URL`, verify JSON-blob columns (embeddings) and
|
||||
`migrations.py` work; test full play loop.
|
||||
- [ ] If SQLite-on-disk: confirm WAL mode + single-worker (or serialized writes) is acceptable.
|
||||
- [ ] Backup story: platform DB backups (Postgres) or a scheduled dump of the disk (SQLite).
|
||||
|
||||
## Exit criteria
|
||||
|
||||
Running the production Docker image locally with `MULTI_USER=true`: a hostile user cannot hang
|
||||
the server with a `while(true)` script, cannot see another user's data or the debug log, gets
|
||||
rate-limited instead of burning the demo key, and the app streams turns normally the whole time.
|
||||
|
||||
### Verified 2026-07-07 (uvicorn, `MULTI_USER=1`, fresh SQLite DB, curl)
|
||||
|
||||
- **Fail-fast secret:** `import app.main` with `MULTI_USER=1` and no `AIDND_SECRET_KEY` raises
|
||||
the RuntimeError as designed (won't boot).
|
||||
- **`while(true)` script:** `POST /api/scripts/{id}/test` on an `input_js` infinite loop returns
|
||||
`InternalError: interrupted` (engine time limit) — server stays responsive afterward.
|
||||
- **Cross-user isolation:** guest B sees `[]` for scripts, gets 404 on guest A's script id;
|
||||
guest A keeps its own row. No leakage.
|
||||
- **Debug log:** `GET /api/debug/requests` → 403 in multi-user mode.
|
||||
- **Rate limiting:** 12 rapid `POST /api/auth/register` → 429 after the 10th (auth scope, 10/300s).
|
||||
- **Body size:** 3 MB body to `POST /api/scenarios` → 413 (limit 2 MB) via BodySizeLimitMiddleware.
|
||||
- **Security headers:** CSP, `x-frame-options: DENY`, `x-content-type-options: nosniff`,
|
||||
`referrer-policy: same-origin` on every response — including the SSE stream.
|
||||
- **Docs disabled:** Swagger UI and OpenAPI schema not served (`/docs`, `/openapi.json` fall
|
||||
through to the SPA `index.html`; no `swagger-ui`, no API schema exposed).
|
||||
- **SSE streaming:** `POST /api/adventures/{id}/actions` streams `text/event-stream` with
|
||||
`x-accel-buffering: no`, chunked, incremental events — the pure-ASGI middlewares don't buffer.
|
||||
(No LLM key configured here, so it streams the "No model configured" error event; a *live*
|
||||
provider turn through this path was verified end-to-end in Phase 8.)
|
||||
|
||||
All Phase 9 exit criteria met.
|
||||
|
||||
### Postgres path verified 2026-07-07 (real Neon, PostgreSQL 18.4)
|
||||
|
||||
Pointed the backend at a live Neon `DATABASE_URL` (pooled, `sslmode=require&channel_binding=require`):
|
||||
|
||||
- **Driver/URL:** `postgres://`/`postgresql://` normalized to `postgresql+psycopg://`; raw
|
||||
psycopg3 connect succeeds.
|
||||
- **Bootstrap:** fresh DB → `create_all` builds all 11 tables and stamps `schema_version = 23`
|
||||
(= `LATEST_VERSION`); the SQLite-only migrations 2–23 are correctly skipped on a fresh DB.
|
||||
- **Embedding column:** `memories.embedding` is Postgres `json`; a float list round-trips
|
||||
intact (`[0.1, 0.2, 0.3, -0.4]` back as a `list`).
|
||||
- **ORM CRUD:** user → scenario → adventure → memory create/read/delete with FK cascades works.
|
||||
- **HTTP round-trip (multi-user, TestClient):** guest bootstrap → register → me → create
|
||||
scenario → list → settings → delete, all 200/201/204 against Neon.
|
||||
|
||||
Remaining Phase 10 step: the Docker production image specifically (build + run the container).
|
||||
@@ -1,56 +0,0 @@
|
||||
# Phase 10 — Deploy & publish
|
||||
|
||||
**Goal:** the app live on Render at a public URL, linked from resume/website alongside the
|
||||
GitHub repo.
|
||||
|
||||
## Decisions (confirmed)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Platform | **Render** |
|
||||
| Domain | **Platform URL is fine** (e.g. `ai-dnd.onrender.com`); custom domain can be added later anytime |
|
||||
|
||||
**Ask before implementing:** Render tier (free-with-sleep vs ~$7/mo always-on — depends on the
|
||||
Phase 9 database decision), and the exact service name (it becomes the public URL).
|
||||
|
||||
### Decisions (confirmed 2026-07-07)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Database | **Neon Postgres** (external managed; free tier has no persistent disk). Path fully verified against real Neon — see `plan/09-phase-hardening.md`. |
|
||||
| Tier | **Free** (sleeps after ~15 min idle; ~30–60s first-wake). |
|
||||
| Service name | `ai-dnd` (→ `ai-dnd.onrender.com`, adjustable in dashboard). |
|
||||
| Region | `virginia` (us-east, matches Neon us-east-1). |
|
||||
| Local Docker preflight | Skipped by choice — Render builds the same Dockerfile in the cloud; its build logs are the image test. |
|
||||
|
||||
## Deploy
|
||||
|
||||
- [x] `render.yaml` blueprint: web service from the Dockerfile, env vars (SECRET_KEY generated,
|
||||
demo-key vars, `MULTI_USER=true`), `/api/health` health check, external Neon Postgres.
|
||||
README gained a "Deploy (Render)" section.
|
||||
- [ ] Set up the Render service, connect the GitHub repo, auto-deploy on push to `main`.
|
||||
- [ ] Seed production with 2–3 good demo scenarios (public/starter scenarios from Phase 8) so
|
||||
first-time visitors have something great to click immediately.
|
||||
- [ ] Smoke test the live URL: guest play on demo key, register, BYOK flow, scripting, memory
|
||||
bank, Insights — from a device/network that isn't yours.
|
||||
- [ ] Free-tier note: if on free tier, first request after idle takes ~30–60s to wake — add a
|
||||
friendly loading state or accept it (revisit tier if it feels bad).
|
||||
|
||||
## Post-launch guardrails
|
||||
|
||||
- [ ] Watch demo-key spend/usage for the first days (OpenRouter dashboard); confirm caps hold.
|
||||
- [ ] Set up uptime monitoring (free: UptimeRobot or similar) — optional.
|
||||
- [ ] Error visibility: Render logs are enough for v1; note how to pull them.
|
||||
|
||||
## Resume / website
|
||||
|
||||
- [ ] README: add the live-demo link + "Try it" section at the top.
|
||||
- [ ] 2–3 sentence project blurb for resume/website (stack, the interesting hard parts:
|
||||
AI Dungeon-compatible JS scripting sandbox, embedding-based memory bank, prompt
|
||||
transparency, guest-first optional auth).
|
||||
- [ ] Later pass (deferred from Phase 7): screenshots/demo GIF for README and website card.
|
||||
|
||||
## Exit criteria
|
||||
|
||||
A recruiter clicks one link on your resume, lands on the live app, plays three turns of a demo
|
||||
scenario as a guest without configuring anything, and can find the GitHub repo from the page.
|
||||
@@ -1,81 +0,0 @@
|
||||
# Plan: undo/retry state revert (+ concurrency lock)
|
||||
|
||||
Fixes three linked issues around `script_state` (the shared per-adventure
|
||||
"scoreboard" scripts write to) and undo/retry.
|
||||
|
||||
## Background
|
||||
|
||||
- `script_state` is one shared dict on `Adventure` (`models.py:90`), mutated in
|
||||
exactly ONE place: `pipeline.py:89` (`self.adventure.script_state = state`).
|
||||
- Today, undo (`adventures.py:445`) and retry (`:422`) delete *actions* but never
|
||||
touch `script_state`, so state never rolls back.
|
||||
|
||||
## Issue 1 — retry double-applies state (pre-existing bug)
|
||||
|
||||
Retry deletes the last AI action and regenerates. The output hook already mutated
|
||||
`script_state` on the first attempt; regenerating runs it again, stacking the change
|
||||
(e.g. "add 10 gold" → 20 gold after one retry). Same root cause as undo not
|
||||
reverting.
|
||||
|
||||
## Issue 4 — undo has no concurrency guard
|
||||
|
||||
Turns take `acquire_turn_lock` (`:189`); undo does not, so undo can race a turn
|
||||
that is still streaming.
|
||||
|
||||
## Issue 2 — Memory Bank leftovers (smaller than expected)
|
||||
|
||||
`run_post_turn` already clamps `memory_cursor`/`summary_cursor` down to the current
|
||||
action count (`memorybank.py:178-182`), so there is NO cursor stall. The only
|
||||
remainder: a `Memory` created from a turn that was later undone stays behind, its
|
||||
`source_start/source_end` now pointing past the end of the story.
|
||||
|
||||
---
|
||||
|
||||
## The fix
|
||||
|
||||
### 1. Snapshot state per turn (Issue 1 + enables undo revert)
|
||||
|
||||
- Add column `state_before: JSON nullable` to `Action` (`models.py`).
|
||||
- Migration: append `(25, "ALTER TABLE actions ADD COLUMN state_before JSON")`
|
||||
to `migrations.py`. `JSON` is valid on both SQLite and Postgres.
|
||||
- In `run_player_turn` / `generate_turn`, when the FIRST action of a turn is
|
||||
created, stash `copy.deepcopy(adventure.script_state)` onto it — captured
|
||||
*before* any hook runs. (Player action for do/say/story; the AI action for a
|
||||
bare `continue`.)
|
||||
- Fix retry directly: before regenerating, restore
|
||||
`adventure.script_state` from the deleted AI action's `state_before` so the
|
||||
output hook starts from the pre-turn scoreboard instead of the mutated one.
|
||||
|
||||
### 2. Revert on undo (Issue depends on #1)
|
||||
|
||||
- In `undo_turn`, after deleting the popped actions, set
|
||||
`adventure.script_state = <deleted player/first action>.state_before` (fall back
|
||||
to `{}` if null, i.e. pre-migration turns), then commit.
|
||||
- Only wire this into undo + retry — NOT the arbitrary
|
||||
`delete_action` endpoint (`:830`); mid-history state revert is undefined.
|
||||
|
||||
### 3. Lock undo (Issue 4)
|
||||
|
||||
- Wrap `undo_turn` body in `acquire_turn_lock(adventure_id)` /
|
||||
`_active_turns.discard(...)` in a `finally`. It's synchronous (not SSE), so no
|
||||
`with_turn_lock` wrapper needed — just acquire and discard.
|
||||
|
||||
### 4. Clean up dangling memories (Issue 2, optional)
|
||||
|
||||
- In `undo_turn` after deleting actions, delete any `Memory` whose `source_start`
|
||||
is >= the new story-action count (i.e. summarized a turn that no longer exists).
|
||||
Cursors already self-heal, so this is polish, not correctness.
|
||||
|
||||
## Known limits (document, don't fix)
|
||||
|
||||
- Story cards a script created (`_apply_cards`, `pipeline.py:87`) are NOT reverted —
|
||||
only text + script_state roll back.
|
||||
- Pre-migration turns have `state_before = NULL` → undo falls back to `{}`.
|
||||
- Demo turn cap is not refunded on undo (intentional).
|
||||
|
||||
## Test checklist
|
||||
|
||||
- Script that increments a counter: play → undo → counter back to prior value.
|
||||
- Same script: play → retry → counter changes once, not twice.
|
||||
- Undo during an active stream returns 409, doesn't corrupt state.
|
||||
- Undo a turn old enough to have been summarized: no orphaned memory left.
|
||||
@@ -1,269 +0,0 @@
|
||||
# Phase 12 — RPG world state (AI-authored, engine-clamped)
|
||||
|
||||
**Goal:** each turn carries a structured **world state** — world stats (e.g. `day`),
|
||||
player stats (`hp`, `mana`), per-NPC stats (`health`, `trust`, …) for the NPCs currently
|
||||
in scene, and **milestones** (story-progress flags / quest objectives). After the player
|
||||
acts, the AI reads the current state, narrates, and **proposes** state changes; the
|
||||
engine **validates and clamps** them against a schema before storing. Stat meanings are
|
||||
described in words (bands) so the model reasons semantically, not arithmetically.
|
||||
|
||||
This is deliberately the "AI owns mechanics, engine enforces limits" design — NOT a
|
||||
deterministic dice engine. The AI proposes; Python is the referee.
|
||||
|
||||
## Design decisions (settled)
|
||||
|
||||
- **AI proposes, engine clamps.** The model never owns the numbers directly. It emits
|
||||
a delta; the engine applies min/max, per-turn caps, and cooldowns server-side. The
|
||||
AI cannot be trusted to obey its own frequency rules — the engine must.
|
||||
- **Band descriptions are the reliability trick.** Stats carry word ranges
|
||||
(`0–20: very weak`, `20–40: hurt`, …). The model reads "he's badly hurt" and adjusts
|
||||
down, instead of doing math it's bad at.
|
||||
- **Dedicated NPCs, each with its own stats.** NPCs are defined in a `npcs` section of
|
||||
the schema, keyed by a stable id (`gwen`). Each has a `name`, `desc`, trigger `keys`,
|
||||
and its **own** `stats` block (a dragon can have `ferocity`, a merchant `prices`) — no
|
||||
forced shared template. On adventure creation each NPC auto-creates a story card (from
|
||||
its name/keys/desc) so lore injection + in-scene detection keep working. Live values
|
||||
live in `world_state["npc"]` keyed by the NPC id; only NPCs **in scene this turn** get
|
||||
their stats injected into context. The AI addresses them as `npc.<id>.<stat>`.
|
||||
- **Schema on the scenario, live values on the adventure.** The scenario is the
|
||||
template (what stats exist, their bands + rules); the adventure holds current values.
|
||||
- **Milestones are sticky story flags.** Predefined objectives the AI marks reached via
|
||||
the same delta channel. Once reached they stay reached (revert only through the undo
|
||||
snapshot). Injected as "Goals" (pending) so the AI drives toward them and "Achieved"
|
||||
so it doesn't re-do them. Emergent/AI-invented milestones are out of scope for v1.
|
||||
- **Flags are two-way booleans.** Separate from milestones: named on/off world/character
|
||||
state (`has_key`, `disguised`, `alarm_raised`) the AI can flip either direction via
|
||||
`"flags.<name>": true|false`. Have an `initial` value and a `desc`; no clamp/cooldown.
|
||||
- **Stat guide (descriptions + band ranges).** A fixed, per-scenario legend injected each
|
||||
turn, describing each stat's `desc` and its full band ladder (`0–20 very weak, …`) —
|
||||
handled independently, so a stat may have a description, a range, both, or neither. This
|
||||
is separate from the live values block (which still shows only the *current* band label),
|
||||
giving the model the whole scale to reason across without bloating the per-turn line.
|
||||
- **One-call turn.** The AI narrates AND appends a fenced state-delta block; the engine
|
||||
parses it and strips it from the visible text. No second LLM call — matters on the
|
||||
rate-limited free-tier demo (20 req/min). Parser is forgiving of messy JSON from
|
||||
weaker free models.
|
||||
- **Memory bank / multi-memory is untouched.** (Confirmed.)
|
||||
- **Separate from `script_state`.** World state gets its own column so it never collides
|
||||
with the scripting scoreboard, and reuses the same undo/retry snapshot pattern
|
||||
(`Action.state_before`, see `plan/11`).
|
||||
|
||||
---
|
||||
|
||||
## Data model
|
||||
|
||||
### Schema definition — `Scenario.stat_schema` (JSON, nullable)
|
||||
|
||||
Migration `(26, "ALTER TABLE scenarios ADD COLUMN stat_schema JSON")`.
|
||||
|
||||
```jsonc
|
||||
{
|
||||
"world": {
|
||||
"day": { "type": "counter", "min": 1, "initial": 1, "desc": "In-game day",
|
||||
"max_delta_per_turn": 1, "cooldown": 0 }
|
||||
},
|
||||
"player": {
|
||||
"hp": { "min": 0, "max": 100, "initial": 100, "max_delta_per_turn": 30,
|
||||
"cooldown": 0, "bands": [[0,20,"very weak"],[20,40,"hurt"],
|
||||
[40,60,"minor damage"],[60,90,"healthy"],[90,100,"full health"]] },
|
||||
"mana": { "min": 0, "max": 50, "initial": 20, "max_delta_per_turn": 15 }
|
||||
},
|
||||
"npcs": { // each NPC has its OWN stats
|
||||
"gwen": {
|
||||
"name": "Gwen", "keys": "Gwen, ranger, her",
|
||||
"desc": "A loyal ranger and the player's ally.",
|
||||
"stats": {
|
||||
"health": { "min": 0, "max": 100, "initial": 100, "bands": [...] },
|
||||
"trust": { "min": -100, "max": 100, "initial": 20, "max_delta_per_turn": 20,
|
||||
"bands": [[-100,-30,"hostile"],[-30,30,"wary"],[30,100,"ally"]] }
|
||||
}
|
||||
}
|
||||
},
|
||||
"flags": { // two-way on/off booleans
|
||||
"has_key": { "desc": "Player holds the dungeon key", "initial": false },
|
||||
"alarm_raised": { "desc": "The enemy is alerted", "initial": false }
|
||||
},
|
||||
"milestones": { // sticky story-progress flags
|
||||
"rescue_gwen": { "desc": "Rescue Gwen from the bandits" },
|
||||
"reach_capital": { "desc": "Arrive at the capital city" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Per-stat rule fields (all optional, engine enforces):
|
||||
- `min` / `max` — hard clamp.
|
||||
- `initial` — value when first instantiated.
|
||||
- `max_delta_per_turn` — largest absolute change allowed in one turn (extra is clamped).
|
||||
- `cooldown` — minimum player actions between changes to this stat (0 = every turn).
|
||||
- `bands` — `[lo, hi, label]` triples, used only to describe the value to the model.
|
||||
- `type` — `"counter"` (monotonic, e.g. day) vs default numeric; counters reject
|
||||
negative deltas.
|
||||
|
||||
Milestones carry only `desc` (the objective text). They are boolean and sticky — the
|
||||
engine accepts a delta of `true` only, records the action index reached, and ignores
|
||||
attempts to re-set or un-set (undo is the only way back).
|
||||
|
||||
Each NPC in `npcs` carries `name`, `desc`, trigger `keys`, and its own `stats` block
|
||||
(same per-stat fields as above). All defined NPCs are instantiated up front; a story card
|
||||
is auto-created per NPC on adventure creation (skipped if a card with that name already
|
||||
exists) so descriptions inject as lore and in-scene detection works.
|
||||
|
||||
### Live values — `Adventure.world_state` (JSON, default `{}`)
|
||||
|
||||
Migration `(27, "ALTER TABLE adventures ADD COLUMN world_state JSON")`.
|
||||
|
||||
```jsonc
|
||||
{
|
||||
"world": { "day": 3 },
|
||||
"player": { "hp": 55, "mana": 10 },
|
||||
"npc": { "gwen": { "health": 80, "trust": 20 } }, // keyed by NPC id
|
||||
"milestones": { "rescue_gwen": { "reached": true, "at": 7 } },
|
||||
"_meta": { "last_changed": { "player.hp": 7, "npc.gwen.trust": 6 } } // action index
|
||||
}
|
||||
```
|
||||
|
||||
`_meta.last_changed` backs the `cooldown` rule. NPC blocks are instantiated up front from
|
||||
each NPC's own `stats`.
|
||||
|
||||
### Undo/retry snapshot — `Action.world_state_before` (JSON, nullable)
|
||||
|
||||
Migration `(28, "ALTER TABLE actions ADD COLUMN world_state_before JSON")`.
|
||||
Snapshotted and reverted exactly like `state_before` (Phase 11) — same call sites.
|
||||
|
||||
---
|
||||
|
||||
## Turn flow
|
||||
|
||||
```
|
||||
player input → INPUT hook (existing)
|
||||
context build → inject [World State] section (current values + band scale + emit-rule)
|
||||
AI response → narration + trailing ```state { ...delta... } ``` block
|
||||
→ parse delta → validate/clamp against schema → apply → strip block
|
||||
→ snapshot world_state onto the action (undo)
|
||||
```
|
||||
|
||||
### 1. Context injection (`context/builder.py`)
|
||||
|
||||
New always-on section `world_state`, placed with the system sections. Keep it **terse**
|
||||
(competes with story history under the default context budget, raised to 16384 in Phase 12):
|
||||
|
||||
```
|
||||
World state — day 3.
|
||||
You: HP 55/100 (minor damage), Mana 10/50.
|
||||
Gwen: health 80 (healthy), trust 20 (neutral).
|
||||
Goals: Arrive at the capital city.
|
||||
Achieved: Rescued Gwen from the bandits.
|
||||
```
|
||||
|
||||
- Only inject NPC lines for cards **triggered this turn** — reuse `triggered` /
|
||||
`card_records` already computed in `build_context` (`builder.py:120`). No extra work.
|
||||
- Milestones: list unreached ones under `Goals` and reached ones under `Achieved`
|
||||
(omit either line when empty). These are cheap and always included.
|
||||
- Append the value's band label in parentheses so the model reads meaning, not just a
|
||||
number.
|
||||
- Append a compact **emit rule** (once, in the narrator/system text). Instruct the model
|
||||
explicitly to:
|
||||
- end its reply with a fenced `state` block **only when something actually changed**;
|
||||
- **omit the block entirely** when nothing changed this turn (no empty `{}`);
|
||||
- include **only the stats that changed** as deltas — never restate unchanged stats,
|
||||
never send full/absolute values, e.g. `{"player.hp": -15, "npc.gwen.trust": +5}`.
|
||||
This keeps the emitted block tiny (saves output tokens on the free tier) and means the
|
||||
engine's clamp/cooldown logic only ever sees real changes.
|
||||
|
||||
### 2. Delta parse + validate (new `worldstate/` module)
|
||||
|
||||
New module `backend/app/worldstate/engine.py` (mirrors `scripting/` layout):
|
||||
|
||||
- `extract_delta(text) -> (clean_text, delta_dict)` — pull the trailing ```` ```state ````
|
||||
block, tolerate missing/extra fences, trailing commas, `+N` numbers; return `{}` on
|
||||
parse failure (never break the turn — same philosophy as a broken script).
|
||||
- `apply_delta(adventure, delta, action_index) -> report` — for each `path: change`:
|
||||
1. resolve `path` (`player.hp`, `world.day`, `npc.<npcId>.trust`,
|
||||
`milestones.<id>`) against the schema; unknown paths ignored (logged).
|
||||
2. **milestone path** → accept only `true`, set `{reached: true, at: action_index}`,
|
||||
ignore if already reached; skip the numeric steps below.
|
||||
3. lazily instantiate NPC stat block from template if missing.
|
||||
4. reject if `cooldown` not elapsed (`action_index - _meta.last_changed[path] < cooldown`).
|
||||
5. clamp change to `max_delta_per_turn`; counters reject negative.
|
||||
6. apply, then clamp result to `[min, max]`.
|
||||
7. record `_meta.last_changed[path] = action_index`.
|
||||
- Returns a report (applied / clamped / rejected) for the Insights panel, like the
|
||||
script report.
|
||||
|
||||
### 3. Wire into `generate_turn` (`routers/adventures.py:249`)
|
||||
|
||||
- Snapshot: `world_state_before = snapshot_world_state(adventure)` alongside the existing
|
||||
`state_before` (`:271`).
|
||||
- After the `output` hook and empty-text check (`:340`): `clean, delta =
|
||||
extract_delta(text)`, `report = apply_delta(...)`, store `text = clean`, stash the
|
||||
report into `snapshot["world_state"]`.
|
||||
- Persist `world_state_before` onto the AI `Action` (`:341` block) and commit
|
||||
`adventure.world_state`.
|
||||
- Do this only when `scenario.stat_schema` is non-empty — zero overhead for plain
|
||||
narrative adventures.
|
||||
|
||||
### 4. Undo / retry (`routers/adventures.py`)
|
||||
|
||||
Reuse the Phase 11 wiring verbatim, in parallel:
|
||||
- retry (`:449`): also restore `adventure.world_state` from the deleted AI action's
|
||||
`world_state_before`.
|
||||
- undo (`:489`): also restore from the first removed action's `world_state_before`
|
||||
(fall back to `{}`).
|
||||
|
||||
---
|
||||
|
||||
## UI
|
||||
|
||||
- **World State panel** (play view): render current `world_state` as a readable sheet —
|
||||
world / player / per-NPC, with band label and a bar for `min..max` stats, plus a
|
||||
**milestones checklist** (reached vs pending). Reuse the collapsible-tree styling from
|
||||
the existing Play State drawer.
|
||||
- **Insights**: show the parsed delta + apply/clamp/reject report per turn (next to the
|
||||
script report already there).
|
||||
- **Scenario editor**: a `stat_schema` editor. **v1 is a raw JSON editor** (CodeMirror,
|
||||
reusing the script-slot editor setup) — settled, no form builder for now. A form-based
|
||||
stat builder is a possible later nicety.
|
||||
- Graceful when `stat_schema` is empty: panel + editor hidden, app behaves exactly as
|
||||
today.
|
||||
|
||||
---
|
||||
|
||||
## Seed / demo
|
||||
|
||||
- New seeded demo scenario `seed_data/*.json` with a small `stat_schema` (player hp/mana,
|
||||
`day`, one or two NPC story cards with health/trust) so the feature is visible on the
|
||||
live demo without the player configuring anything. Keep it `:free`-model friendly.
|
||||
- Extend `seed.py` idempotently (matches existing seeder contract).
|
||||
|
||||
---
|
||||
|
||||
## Exit criteria
|
||||
|
||||
Play the demo RPG scenario: the World State panel shows hp/mana/day, an NPC's trust, and
|
||||
a milestones checklist; taking a fight action drops hp and the narration matches the new
|
||||
band; a friendly action raises an NPC's trust; completing an objective marks its
|
||||
milestone reached (and it stays reached); a change larger than `max_delta_per_turn` is
|
||||
clamped; undo rolls every stat and milestone back to the prior turn; a plain (no-schema)
|
||||
adventure is completely unaffected.
|
||||
|
||||
## Known limits (document, don't fix in v1)
|
||||
|
||||
- The AI can still narrate against the numbers occasionally; injected state + firm emit
|
||||
rule reduces but won't eliminate it. No post-narration consistency check in v1.
|
||||
- No dice / skill checks / combat resolution — this phase is stat tracking only. A
|
||||
deterministic resolver is a possible Phase 13.
|
||||
- NPC stats are keyed by the schema NPC id; editing a scenario's `npcs` between play
|
||||
sessions can orphan a live `npc` entry (harmless, ignored on read).
|
||||
- Cooldown/`max_delta_per_turn` are per-turn heuristics, not a full rules engine.
|
||||
|
||||
## Test checklist
|
||||
|
||||
- Schema with `max_delta_per_turn: 30`: a delta of `-50` applies as `-30`, clamps at `min`.
|
||||
- `cooldown: 2` on a stat: two consecutive changes → second is rejected until 2 actions pass.
|
||||
- Counter (`day`): a negative delta is rejected; `+1` advances.
|
||||
- Milestone: `true` marks it reached with `at`; a second set is a no-op; `false` ignored.
|
||||
- Each NPC instantiates its own stats at `initial`; `npc.<id>.<stat>` resolves per-NPC.
|
||||
- Malformed / missing `state` block → turn still completes, delta `{}`, no crash.
|
||||
- A turn where nothing changes emits no `state` block (and an empty `{}` is a no-op).
|
||||
- Undo after a stat change restores the prior value; retry doesn't double-apply.
|
||||
- Empty `stat_schema`: no World State section injected, no `world_state` writes.
|
||||
@@ -1,317 +0,0 @@
|
||||
# 13 — Memory-bank embedding cost (round three of the egress work)
|
||||
|
||||
**Goal:** stop every turn fetching the entire memory bank's embeddings. Measured at
|
||||
**3,024 KB per turn** on adventure 25 against **129 KB** for everything else a turn
|
||||
reads — the memory bank is ~96% of a turn's database traffic, and it is fetched fresh
|
||||
every single turn to pick `memory_top_k = 5` memories.
|
||||
|
||||
Found 2026-08-16 while designing the story tree (see `14-phase-story-tree.md`), because
|
||||
memory retrieval is the one read a tree **cannot** window — it is long-range recall by
|
||||
design, so it always spans the full path. That makes this the cost floor of a turn under
|
||||
the tree, which is why it lands first.
|
||||
|
||||
## The measurement (production, Neon SQL editor)
|
||||
|
||||
```sql
|
||||
SELECT count(*), avg(json_array_length(embedding))::int,
|
||||
avg(length(embedding::text))::int,
|
||||
pg_size_pretty(sum(length(embedding::text))::bigint)
|
||||
FROM memories WHERE embedding IS NOT NULL;
|
||||
-- 134 memories | 1536 dims | 30,971 bytes each | 4,053 kB total
|
||||
```
|
||||
|
||||
| adventure | active | fetched | per turn |
|
||||
|---|---|---|---|
|
||||
| 25 | 100 | 100 | **3,024 kB** |
|
||||
| 21 | 18 | 18 | 545 kB |
|
||||
| 12 | 12 | 12 | 363 kB |
|
||||
| 20 | 4 | 4 | 121 kB |
|
||||
|
||||
Break-even against the rest of a turn is **4.2 memories**, i.e. about action 25. Every
|
||||
adventure past that is dominated by this.
|
||||
|
||||
`active == fetched` everywhere: nothing has been evicted yet, so the Python-side
|
||||
`forgotten` filter currently costs nothing. It becomes a real leak the moment eviction
|
||||
starts.
|
||||
|
||||
## Why round two missed it
|
||||
|
||||
`retrieve_memories` needs `settings.embedding_model`, and embedding providers are
|
||||
**BYOK-only by construction** — they never touch the demo key. So the public demo never
|
||||
embeds anything, and the round-two stress harness (which had no embedding model
|
||||
configured) measured the turn loop with its heaviest read switched off. The 23.0 MB
|
||||
figure for a 200-action playthrough is the memory-bank-**off** number; with it on,
|
||||
adventure 25 is closer to 300 MB.
|
||||
|
||||
**Rule going forward: any egress measurement must run with an embedding model set.**
|
||||
|
||||
## Root cause
|
||||
|
||||
`memorybank.py:208`
|
||||
|
||||
```python
|
||||
candidates = [m for m in adventure.memories if not m.forgotten and m.embedding]
|
||||
```
|
||||
|
||||
Walks the relationship, so every memory row for the adventure loads with its embedding.
|
||||
`embedding` is `Mapped[list]` on a `JSON` column — 1536 floats serialised as text is
|
||||
~31 KB. Cosine ranking happens in Python (a deliberate choice, documented at
|
||||
`models.py:144`), so all of it must cross the wire. The comment sized it by **count**
|
||||
("fine at a few hundred") rather than by **bytes**.
|
||||
|
||||
Not the same bug as migration 36/37 — the column is not a repeating group and there is
|
||||
no denormalisation to fix. It is a *format* problem plus a *fetch-frequency* problem.
|
||||
|
||||
## Decisions (settled 2026-08-16)
|
||||
|
||||
- **Packed float32, keep 1536 dimensions.** 31 KB → 6 KB, a straight 5x, with **zero
|
||||
retrieval-quality risk**. Explicitly rejected dropping to 512/768 dims: the size fix
|
||||
plus the cache makes the extra 3x unnecessary, and it would have meant re-embedding.
|
||||
- **No re-embedding.** Dimensions are unchanged, so the migration is a pure format
|
||||
conversion of the 134 existing rows — read the JSON, write packed bytes, no API calls.
|
||||
One-time 4 MB read.
|
||||
- **In-process cache alongside the size fix**, not sequenced after it. Turns for one
|
||||
adventure arrive back-to-back, so a dict keyed by adventure id takes steady-state cost
|
||||
to ~0. 100 vectors as float32 is 600 KB of RAM — negligible.
|
||||
- **`memory_bank_capacity` 200 → ~80.** Taken on *quality* grounds as much as cost:
|
||||
ranking 200 memories to pick 5 dilutes retrieval. Note this will start evicting on
|
||||
adventure 25 immediately (it sits at 100).
|
||||
- **Not pgvector.** It is the structural answer and would keep vectors in the database
|
||||
entirely, but it breaks SQLite dev parity — which the codebase protects deliberately
|
||||
(`context/history.py:42`, the `replace()`/`trim()` dialect dance). Revisit only if the
|
||||
bank grows past what Python cosine can handle.
|
||||
|
||||
## Work, in order
|
||||
|
||||
1. **Rebuild the byte-meter harness — in the repo this time.** `backend/tools/dbmeter.py`
|
||||
+ `stress_session.py`. The originals lived outside the repo and are gone. **Default it
|
||||
to running with an embedding model configured**, since that omission is precisely what
|
||||
hid this finding. Do this first so every item below is measured, not assumed.
|
||||
**Done 2026-08-16** — see the baseline below.
|
||||
2. **Migration 38 — `memories.embedding_blob` (`LargeBinary`).** **Done.** Backfill in Python
|
||||
(`struct.pack(f"<{n}f", *vec)`); the conversion cannot be expressed in portable SQL, so
|
||||
unlike migration 36/37 this one does pay a one-time 4 MB read. Drop the old JSON column
|
||||
in a follow-up migration once verified, not in the same one.
|
||||
3. **Read path.** **Done.** `retrieve_memories` queries `memories` directly with
|
||||
`forgotten = false AND embedding_blob IS NOT NULL` in **SQL**, not Python. Unpack with
|
||||
`struct`/`numpy`. Same for `_evict_over_capacity` and `_embed_pending`, which walk the
|
||||
same relationship for a count and for the unembedded rows (see the baseline above).
|
||||
4. **Vector cache.** **Done.** Keyed by `adventure_id`, bounded to 8 adventures. It
|
||||
turned out to need no invalidation callbacks at all — see below.
|
||||
5. **Capacity default 200 → 80.** **Done** (migration 41, only rows still on the old
|
||||
default). Eviction checked at scale first: trimming a 100-memory bank to 80 costs
|
||||
0.8 kB and eight statements, and reads no vectors at all.
|
||||
6. **Infinite scroll upward in `Play.jsx`** for the remaining 423 KB page load of a
|
||||
finished adventure — the last open item from round two. Load the newest turns, fetch
|
||||
older ones as the reader scrolls up.
|
||||
|
||||
## Baseline from the harness (2026-08-16)
|
||||
|
||||
`python -m tools.stress_session`, 200 actions × 4 KB, 100 memories × 1536 dims:
|
||||
|
||||
| shape | fetched | memories' share |
|
||||
|---|---|---|
|
||||
| `GET /adventures` (index) | 0.7 kB | — |
|
||||
| `GET /adventures/{id}` (page load) | 426.7 kB | — |
|
||||
| `POST /adventures/{id}/actions` (one turn) | **3,258.7 kB** | 96% |
|
||||
| `GET /adventures/{id}/context` (Insights) | 3,223.7 kB | 97% |
|
||||
| `run_post_turn` (background) | **3,139.1 kB** | 100% |
|
||||
|
||||
It reproduces both production figures independently — 426.7 kB against the measured
|
||||
423 KB page load, 3,258.7 kB against 3,024 + 129 kB for a turn, and 31.4 KB per
|
||||
embedding against 31.0 KB. `--no-embeddings` reports 112.0 kB for the same turn, so
|
||||
the round-two blind spot is now a **29x** gap anyone can see in one flag.
|
||||
|
||||
**Two findings the SQL measurement could not have shown**, both the same root cause
|
||||
in a different caller:
|
||||
|
||||
- **`run_post_turn` fetches the whole bank again**, every turn. `_evict_over_capacity`
|
||||
walks `adventure.memories` to count the active ones, and `_embed_pending` walks it to
|
||||
find the unembedded ones. So a played turn actually costs ~6.4 MB, not 3.2 — the
|
||||
original estimate was half the real number.
|
||||
- **Insights pays it a third time**, on a page the player can open repeatedly without
|
||||
spending a turn.
|
||||
|
||||
So step 3 below is not just `retrieve_memories`: **every walk of
|
||||
`adventure.memories` has to go**. `_evict_over_capacity` wants a count and an ordering,
|
||||
`_embed_pending` wants rows where `embedding IS NULL` — neither needs a single vector,
|
||||
and both are pure SQL.
|
||||
|
||||
## After (2026-08-16, same harness, same fixture)
|
||||
|
||||
| shape | before | after (cold) | after (warm) |
|
||||
|---|---|---|---|
|
||||
| `POST .../actions` (one turn) | 3,258.7 kB | 723.4 kB | **122.3 kB** |
|
||||
| `run_post_turn` | 3,139.1 kB | 0.7 kB | 0.7 kB |
|
||||
| `GET .../context` (Insights) | 3,223.7 kB | 117.9 kB | 117.9 kB |
|
||||
| `GET .../memories` (drawer) | ~3,138 kB | 23.6 kB | 23.6 kB |
|
||||
|
||||
A played turn is turn + post_turn: **6,398 kB → 123 kB** once warm, a 52x cut. The
|
||||
targets above were ~700 kB cold and ~130 kB warm, so both were met.
|
||||
|
||||
Steps 2–5 landed together, because they are one deployable unit: the columns are no
|
||||
use unless something reads them, and deferring them breaks the old readers. Two
|
||||
additions the plan did not anticipate:
|
||||
|
||||
- **`memories.embedded`, a boolean beside the blob** (migration 39/40). Once the
|
||||
vectors are deferred, every "does this have an embedding?" check becomes a lazy
|
||||
load — an N+1 of 6 KB reads down the Memories drawer. Same shape as
|
||||
`actions.variant_count` beside `actions.variants`, and the same reason.
|
||||
- **The cache needs no invalidation callbacks.** Vectors only ever change through
|
||||
`set_vector`, which drops the one entry; everything that *removes* a memory from
|
||||
play leaves the catalogue query, and entries missing from the catalogue are dropped
|
||||
on the next read. So eviction, deletion and pruning need no hooks and cannot be
|
||||
forgotten. Vectors are held as `array("f")` — 6 KB each, matching the column;
|
||||
a list of Python floats would have been eight times the plan's RAM estimate.
|
||||
|
||||
**Step 6 landed on 2026-08-17**, along with everything else this plan left open — see
|
||||
"Closing the plan" at the end. The page load is a window of 60 actions now: 62.6 kB on
|
||||
a 600-action fixture, down from 606.0 kB, and no longer a function of the story's
|
||||
length.
|
||||
|
||||
## Guardrails to add with this work
|
||||
|
||||
- **Query-count / byte assertions per endpoint**, extending the `test_egress.py` idea:
|
||||
assert an endpoint issues at most N queries and fetches under X KB against
|
||||
production-sized fixtures. This class of bug is invisible at ten rows.
|
||||
- **Explicit column projections on read paths.** List endpoints name the fields they
|
||||
need rather than loading whole entities, so the next heavy column is opt-**in**. This
|
||||
is the structural version of what `deferred=True` does by hand.
|
||||
- **Row-width review rule.** Any new large column justifies itself or goes out-of-line.
|
||||
`actions` now carries five JSON columns.
|
||||
|
||||
Deliberately **not** taken: moving `context_snapshot` out of the database. It costs
|
||||
nothing on reads now that it is deferred, and storage is ~$0.02/mo. Revisit only if
|
||||
backups or storage start to hurt.
|
||||
|
||||
> **Revisit it.** That call weighed egress and got egress right, but it never weighed
|
||||
> the free tier's *storage* ceiling — see "Storage, which this plan did not cost" below.
|
||||
|
||||
## Verification
|
||||
|
||||
- Harness: `python -m tools.stress_session`, memory bank **on**, before and after,
|
||||
against the baseline table above. Target for the turn shape is 3,258 kB → ~700 kB
|
||||
cold, ~130 kB warm (the cache leaves only what a turn reads besides the bank).
|
||||
`run_post_turn` should fall to roughly nothing: neither of its two walks needs a
|
||||
vector at all.
|
||||
- Re-run the round-two shapes with an embedding model configured, so the 200-action
|
||||
playthrough number is finally honest.
|
||||
- The existing `test_egress.py` guard must still pass — nothing here should touch the
|
||||
deferred action columns.
|
||||
|
||||
## Verified on production, 2026-08-17
|
||||
|
||||
Two things were still taken on trust when this shipped: every measurement had run on
|
||||
SQLite, and every number came from a synthetic fixture. Both are now checked.
|
||||
|
||||
### The migration landed on real Postgres
|
||||
|
||||
`schema_version` reads **41**, matching the repo's `LATEST_VERSION`. The live schema has
|
||||
`embedding_blob bytea` and `embedded boolean`, so the `{dialect: sql}` map in migration
|
||||
38 spells BYTEA correctly against a real server — the one thing tests could not prove,
|
||||
since `test_migration_38_is_spelled_for_both_dialects` only inspects the SQL string.
|
||||
The backfill is complete: 134 memories, `embedded = 134`, `embedding_blob = 134`, no
|
||||
stragglers and no rows skipped as malformed.
|
||||
|
||||
### The 5x is real, on real vectors
|
||||
|
||||
| | bytes | per memory |
|
||||
|---|---|---|
|
||||
| `embedding` (JSON) | 4,150,121 | 30,971 |
|
||||
| `embedding_blob` (float32) | 823,296 | 6,144 |
|
||||
|
||||
**5.04x**, against the plan's predicted ~31 KB → 6,144 B. The largest real bank is 100
|
||||
memories = 614,400 B of vectors, so the old code fetched **~3.10 MB per retrieval** on
|
||||
that adventure — which is where the 3,153 kB measured on production came from. That
|
||||
figure is now fully accounted for.
|
||||
|
||||
### SQLite and Postgres agree
|
||||
|
||||
`tools.stress_session` gained an `AIDND_STRESS_DATABASE_URL` escape hatch and was run
|
||||
against a throwaway Neon database at the default fixture (200 actions, 100 memories):
|
||||
|
||||
| shape | SQLite | Postgres |
|
||||
|---|---|---|
|
||||
| index | 4.1 kB | 4.1 kB |
|
||||
| page load | 426.7 kB | 425.0 kB |
|
||||
| one turn, cold | 723.4 kB | 722.3 kB |
|
||||
| one turn, warm | 122.3 kB | **121.1 kB** |
|
||||
| Insights | 117.9 kB | 116.7 kB |
|
||||
| Memories drawer | 23.7 kB | 21.7 kB |
|
||||
| `run_post_turn` | 0.7 kB | 0.6 kB |
|
||||
|
||||
Within 0.5% everywhere. The dialect caveat in the harness docstring is real but small:
|
||||
what dominates is which columns get asked for, and the ORM decides that identically.
|
||||
The warm turn spends **1.7 kB on `memories`, 1% of the read** — the cache behaves on
|
||||
psycopg exactly as it does on SQLite.
|
||||
|
||||
### The page load is worse than modelled, for a different reason
|
||||
|
||||
The synthetic fixture is **~2x heavier per action than production**: 994 B/action real
|
||||
against ~2,133 B/action synthetic, so a real 200-action adventure is ~194 kB, not 427.
|
||||
But the largest real adventure is **607 actions**, not 200, and costs **589.5 kB** in
|
||||
one response. Step 6 is more urgent than this plan assumed, and for the opposite
|
||||
reason to the one modelled — stories get *longer* than the fixture, not heavier.
|
||||
|
||||
Worth fixing the fixture's narration size when step 6 lands, so the harness stops
|
||||
flattering the per-action figure while understating the length.
|
||||
|
||||
### Storage, which this plan did not cost
|
||||
|
||||
`context_snapshot` is **150.8 MB of uncompressed JSON across 944 actions** — ~163 kB a
|
||||
row on average, and ~232 kB a row in the largest adventure, against the ~74 KB/row the
|
||||
comment in `models.py` claims. TOAST compresses it to ~89 MB on disk, but
|
||||
`octet_length` is what would cross the wire, because Postgres decompresses before
|
||||
sending. Deferral is the only thing standing between a bulk read and a 137 MB query.
|
||||
|
||||
The database is **99.6 MB total**, of which `actions` is **88.9 MB**. Neon's free tier
|
||||
is 512 MB. At ~94 kB of disk per action that ceiling arrives at roughly **5,400
|
||||
actions**, and 944 are already stored. So the "~$0.02/mo, leave it in the database"
|
||||
call above is wrong for the tier this actually runs on — not because reads cost
|
||||
anything, but because the free tier meters *storage*, and that is the constraint with
|
||||
a cliff. Dropping the dead `memories.embedding` column reclaims 4.05 MB (4%), which
|
||||
helps and does not solve it.
|
||||
|
||||
None of the numbers above required reading a single row of anyone's content: counts,
|
||||
`octet_length` sums and catalog sizes only.
|
||||
|
||||
## Closing the plan, 2026-08-17
|
||||
|
||||
Everything above landed the same day the verification did.
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| page load, 600 actions | 606.0 kB | **62.6 kB**, and flat in story length |
|
||||
| adventures index, 6 adventures | 469.7 kB | **0.3 kB** |
|
||||
| `context_snapshot` stored | ~89 MB | ~43 MB (after a VACUUM) |
|
||||
| a played turn | 6.4 MB (2026-08-16) | 123 kB |
|
||||
|
||||
**Step 6, the window.** `GET /adventures/{id}` returns the newest `ACTION_PAGE`
|
||||
actions and the story's length; `GET /{id}/actions?before_id=` walks back from there.
|
||||
Anchored on an action rather than an offset — the offset version breaks precisely when
|
||||
a turn lands mid-scroll, handing the reader one action twice and hiding another — and
|
||||
that choice is also what makes it survive the story tree, since it compares indices to
|
||||
order a branch rather than treating them as positions.
|
||||
|
||||
**Byte assertions.** `tests/test_egress.py` now carries per-action budgets as well as
|
||||
column guards, plus one test whose only job is to fail if the fixture ever gets too
|
||||
small for the budgets to catch anything.
|
||||
|
||||
**Column projections.** `ACTION_LIST_COLUMNS` and `MEMORY_LIST_COLUMNS` name what a
|
||||
list response renders, and the adventures index selects four columns instead of the
|
||||
entity. That one was not just future-proofing: an Adventure carries seven text and JSON
|
||||
columns the index never shows.
|
||||
|
||||
**The JSON vector column is gone**, and dropping it exposed a live bug — changing your
|
||||
embedding model had stopped re-embedding the bank the day migration 38 shipped. See
|
||||
`tests/test_embedding_model_switch.py`.
|
||||
|
||||
**Storage.** `context_snapshot` is zlib-compressed through a TypeDecorator
|
||||
(`app/compression.py`, migrations 43–45), so the call sites never learned about it.
|
||||
3.5x on real Postgres. **This does not shrink anything until `VACUUM FULL actions`
|
||||
runs** — Postgres marks dropped columns rather than reclaiming them, and the backfill
|
||||
leaves a dead tuple per row.
|
||||
|
||||
Not done: nothing exercises the scroll behaviour in a browser. The frontend has no test
|
||||
runner, and prepend-and-restore-scroll is the part most likely to feel wrong even when
|
||||
it is correct.
|
||||
@@ -1,914 +0,0 @@
|
||||
# Phase 14 — Story tree (branching adventures)
|
||||
|
||||
**Goal:** replace the linear action list with a **tree**. A retry becomes a sibling
|
||||
rather than a rewrite; continuing from one makes it a branch. Players can go back to any
|
||||
turn, take a different path, and keep both — switching between them freely.
|
||||
|
||||
Depends on `13-memory-embedding-cost.md` shipping first: memory retrieval is the one read
|
||||
a tree cannot window, so it is the cost floor of a turn under this design. Fix the floor
|
||||
before building on it.
|
||||
|
||||
## Why (beyond the feature)
|
||||
|
||||
Seven bug classes stop existing, and every one traces to the same root — **the story is a
|
||||
mutable list**:
|
||||
|
||||
| Bug | Why it goes away |
|
||||
|---|---|
|
||||
| Deleting a middle action skips a later one forever (`note_action_removed`) | Cursors become node ids, not positions in a shifting list |
|
||||
| `prune_dangling_memories` orphaning actions behind the cursor | Nothing is ever removed |
|
||||
| Editing a summarised action leaves its memory stale forever — **currently unfixed** | Editing makes a new node; the old memory stays correct for the old path |
|
||||
| The one-turn memory holdback (`settled_story_actions`) | Nothing is mutated, so nothing goes stale — the concept is unnecessary |
|
||||
| The legacy-cursor no-rewind trap found while fixing that | Same |
|
||||
| "Anything reading `adventure.actions` during generation must exclude the retried action" — leaked into 4 call sites | The replaced turn is not on your path; it cannot leak |
|
||||
| `variant_count` / mirrored `text` drifting from `variants` | The denormalisation disappears entirely |
|
||||
|
||||
It also resolves the 1NF violation: `variants` as a JSON repeating group becomes rows.
|
||||
|
||||
**Scored, 2026-08-18, with SP1–SP5 shipped.** Six of the seven are gone as described.
|
||||
The seventh — the one-turn holdback — is gone too, but the reasoning above was wrong
|
||||
about *why* it could go: siblings share a coordinate, so replacing what a turn says still
|
||||
invalidates the memory covering it. What made it deletable is that the repair already
|
||||
existed for undo and delete; see the trap note below. Two rows want a footnote:
|
||||
|
||||
- **Editing a summarised action** is still unfixed. The table says editing makes a new
|
||||
node; nothing in SP1–SP5 makes it do that, and no subphase is scheduled to. `PATCH
|
||||
/actions/{id}` still writes over the text in place, and the memory covering it goes
|
||||
stale exactly as before. What *has* changed is that the machinery to fix it now exists
|
||||
— an edit could write a sibling and switch to it, which is a retry the player typed —
|
||||
so it is a small change whenever it is wanted.
|
||||
- **The 1NF violation** is resolved. Nothing writes a `variants` array any more, in the
|
||||
database or out of it — SP6 replaced the bundle that was its last producer, and the
|
||||
array survives only in the v1 *reader*, which exists so files already saved still
|
||||
import.
|
||||
|
||||
## Design decisions (settled 2026-08-16)
|
||||
|
||||
- **Full branching, with UI.** Not the "tree schema, no branch picker" middle option —
|
||||
branching ships as a feature people use.
|
||||
- **`branch_id` + `depth` on every node, not parent pointers alone.** Parent pointers
|
||||
alone mean walking N links to read a story, which throws away the round-two windowing
|
||||
work. `depth` replaces `index` as the ordering key.
|
||||
- **A `branches` table with `parent_branch_id` and `fork_depth`.** The fork point is
|
||||
**stored at fork time, never inferred**. Nothing is copied on a fork — a branch borrows
|
||||
its ancestors' turns.
|
||||
- **Lineage cached on the branch row.** `lineage = [(D,∞), (C,45), (B,30), (A,10)]`,
|
||||
computed once at fork (parent's lineage + one entry). Reads never walk to reconstruct
|
||||
it. Each ancestor is capped at the `fork_depth` of the branch beneath it.
|
||||
- **Read the lineage lazily, windowed.** Query the newest few lineage entries, measure,
|
||||
fetch more only if the context budget is not covered — the exact shape of
|
||||
`history.window_covering()`. **Clause count is bounded by the context window, not by
|
||||
fork count**, so a 200-fork story reads as cheaply as a 2-fork one.
|
||||
- **Promote to a branch only on continue.** Attempts at the tip stay as sibling leaves;
|
||||
one becomes a branch the moment a turn is played past it. Keeps the lineage chain to
|
||||
"divergences I built a story on", not "every retry ever" — the difference between a
|
||||
handful of entries and fifty.
|
||||
- **Memories and the summary attach to the node that produced them**, found by walking
|
||||
up. Shared ancestors are shared automatically, so a fork costs nothing and **nothing
|
||||
needs recreating**. A memory covering depths 37–42 hangs off that branch's node 42 and
|
||||
is invisible to any path not through it. Generalise the rule: *anything derived attaches
|
||||
to the node that produced it* — the summary included, so stop storing it per turn.
|
||||
- **Full lineage for memories, windowed lineage for the story.** Memory retrieval is
|
||||
long-range recall and cannot be windowed, but memories are sparse (~1 per 6 actions), so
|
||||
a long OR-clause returning ~33 small rows is fine. Two queries, one lineage.
|
||||
- **Never auto-prune.** Nothing is deleted without an explicit user action. **This makes
|
||||
a branch-management UI a hard dependency, not a nice-to-have** — storage grows without
|
||||
limit otherwise.
|
||||
- **Story cards stay adventure-wide.** A card invented on branch B shows on branch A.
|
||||
Already true for undo today (script card mutations are not reverted), so this is a
|
||||
documented limit, not a regression. Explicitly rejected event-sourcing card changes onto
|
||||
nodes.
|
||||
- **State carries over almost free.** `state_before` / `world_state_before` already
|
||||
snapshot the script scoreboard and RPG world state per action, and `apply_variant()`
|
||||
already restores them on a switch — a branch switch is the same move. Wrinkle: those are
|
||||
*before* pictures; a node wants the *after*. And they are NULL on pre-column rows.
|
||||
|
||||
## Schema sketch
|
||||
|
||||
```
|
||||
branches(id, adventure_id, parent_branch_id, fork_depth, lineage JSON, created_at)
|
||||
actions(id, adventure_id, branch_id, depth, type, text, reasoning,
|
||||
world_delta, state_after, world_state_after, context_snapshot, created_at)
|
||||
memories(..., branch_id, depth) -- attached to the node that produced it
|
||||
adventures(..., head_branch_id, head_depth)
|
||||
```
|
||||
|
||||
Reading branch C, tip at depth 7, lineage `[(C,7), (B,5), (A,3)]`:
|
||||
|
||||
```sql
|
||||
SELECT * FROM actions
|
||||
WHERE (branch_id='C' AND depth <= 7)
|
||||
OR (branch_id='B' AND depth <= 5)
|
||||
OR (branch_id='A' AND depth <= 3)
|
||||
ORDER BY depth DESC LIMIT 32
|
||||
```
|
||||
|
||||
→ `A0 A1 A2 A3 B4 B5 C6 C7`. Depth is a position along *a* path, not a global turn
|
||||
number — `A4` and `B4` are alternatives, not duplicates.
|
||||
|
||||
## Work
|
||||
|
||||
1. `branches` table, `branch_id`/`depth` on `actions`, `head_*` on `adventures`.
|
||||
2. Lineage computation + a **single module that owns the branch clause** — same role
|
||||
`context/history.py` plays today. Every query must go through it; one forgotten clause
|
||||
shows the wrong story, quietly.
|
||||
3. `history.py` rewritten against lineage windowing. `window_covering` keeps its shape.
|
||||
4. Memories/summary attached to nodes; delete the cursor-position machinery
|
||||
(`position_of_index`, `settled_*`, `note_action_removed`, `_rewind_cursors_to_index`).
|
||||
5. Sibling storage for un-promoted tip attempts + the promotion step.
|
||||
6. Node state moves from *before* to *after* snapshots.
|
||||
7. Migration: every existing adventure becomes branch A; `index` → `depth`; `variants`
|
||||
entries become sibling nodes; `variant_index` becomes the head pointer.
|
||||
8. Frontend: branch picker replacing `VariantPager`, plus branch management (rename,
|
||||
delete, switch) — required, given no auto-pruning.
|
||||
|
||||
## Open
|
||||
|
||||
Nothing. The last item — **`retry_of.index` reuse**, which stopped the world-state
|
||||
cooldown clock advancing on a re-run of the same turn — was closed in SP4 by reusing the
|
||||
retried node's *depth*, and pinned by a test in SP5. Everything else that stood open here
|
||||
was decided on 2026-08-17; see the next section.
|
||||
|
||||
---
|
||||
|
||||
# Implementation plan (decided 2026-08-17)
|
||||
|
||||
Four decisions, taken before any code was written:
|
||||
|
||||
| Question | Decision |
|
||||
|---|---|
|
||||
| Phased or single migration? | **Structural-first, no runtime flag.** SP1–SP4 migrate to the tree *representation* with behaviour identical to today. Branching features layer on after. |
|
||||
| How destructive is the migration? | **Keep the legacy columns for one release.** `index`, `variants`, `variant_index`, `variant_count` stay, unread, until the tree is proven live (SP8 drops them). |
|
||||
| Branch-management UI scope | **Full tree visualisation** — a spatial view of forks, not just a picker. |
|
||||
| Export bundle | **`ai-dnd-adventure-v2`**, with a v1 legacy reader so existing bundles keep importing. |
|
||||
|
||||
**Why no feature flag.** A linear story *is* a tree with one branch, so the intermediate
|
||||
states are not half-migrated — they are the same product with a superset schema
|
||||
underneath. That makes "current adventures are unaffected" a literal, testable pass
|
||||
condition for every subphase up to SP4, which a flag would have replaced with two live
|
||||
code paths through the context builder, the memory bank, undo and retry at once.
|
||||
|
||||
## The regression contract
|
||||
|
||||
`tests/test_story_tree_baseline.py` (built in SP0) drives the whole product over HTTP —
|
||||
create, play, retry, switch, undo, page with `before_id`, memories, summary, export — and
|
||||
asserts only on API responses, never internals.
|
||||
|
||||
**It must pass unmodified through SP1, SP2 and SP3.** That is the contract those
|
||||
subphases are verified against. SP4 is the first subphase permitted to change it, and
|
||||
even there only where the change is deliberate and named below.
|
||||
|
||||
## Traps found while reading (these are the ones that bite quietly)
|
||||
|
||||
- **`history._from_memory()` slices `adventure.actions`.** The "never load twice"
|
||||
shortcut returns whatever is already in the session — which under a tree is *every
|
||||
branch's* actions, not the path. It would silently assemble context from siblings.
|
||||
This is the single highest-risk line in the change; SP2 must make the shortcut
|
||||
branch-aware or delete it.
|
||||
- **`scripting/pipeline.py::_history()` and `_info()` read `adventure.actions` directly**,
|
||||
and hand it to user scripts as the documented history API. Same trap, user-visible.
|
||||
- **`Adventure.actions` is `order_by="Action.index"`.** Ordering it by `depth` is not
|
||||
enough — the collection is still every branch. Anything iterating it needs the path.
|
||||
- **`limits.check_row_cap("actions")` counts every action in the adventure.** With no
|
||||
auto-pruning, a branched adventure hits `MAX_ACTIONS_PER_ADVENTURE` while its *story*
|
||||
is far shorter. The cap has to count the tree but be explained as the tree, or move.
|
||||
- **The holdback cannot die in SP3.** `settled_story_actions` exists because retry
|
||||
mutates a row in place. Retry stops mutating in SP4, so the holdback is only safe to
|
||||
delete there — deleting it in SP3 reopens the exact bug it was written for. It survived
|
||||
SP3 as `memorybank.settled_after`, which is the `- 1` in "how much story is past the
|
||||
mark"; that subtraction is the whole of it.
|
||||
|
||||
*Closed in SP4, but not for the stated reason.* Siblings share a coordinate and the
|
||||
mark names the coordinate, so replacing what a turn says still invalidates the memory
|
||||
covering it — retry mutating a row was never the whole of the problem. What let the
|
||||
holdback go is that the right repair (`forget_node` plus a rewind) already existed for
|
||||
undo and delete, and retry and a sibling switch now make it too. **If a mark still
|
||||
needs correcting when the story changes, correct it; do not decline to make the mark.**
|
||||
|
||||
## Subphases
|
||||
|
||||
Each ships independently, on its own branch, green before the next starts.
|
||||
|
||||
### SP0 — Baseline and regression net *(no product change)*
|
||||
|
||||
- `tests/test_story_tree_baseline.py` — the contract above. 24 tests covering opening a
|
||||
windowed adventure, the turn engine (do/say/story/continue), script effects, retry and
|
||||
variant switching, undo, paging up to the start, edit/delete, memories, export/import
|
||||
round-trip, and world state.
|
||||
|
||||
**Done, 2026-08-17.** 259 existing tests green (17.8s), then **283 green** with the
|
||||
baseline added, against unmodified `main`. That is the number every subphase below is
|
||||
measured against.
|
||||
|
||||
A pre-migration (**schema 45**) fixture is needed too, built by a script rather than
|
||||
committed as a binary — but its only consumer is the SP1 migration test, so it lands
|
||||
there.
|
||||
|
||||
#### `--rich`: a correctness fixture beside the scale one
|
||||
|
||||
The measuring fixture is sized from production and leaves every column it does not weigh
|
||||
at its default. Checked against a freshly built one, that is exactly the set of columns a
|
||||
tree has to migrate: `state_before`/`world_state_before` NULL on all 600 rows, no
|
||||
scenario and so no RPG layer, no adventure scripts, both cursors 0, and 100 retry
|
||||
histories whose two attempts carry **byte-identical text with `variant_index` always 0** —
|
||||
so "which attempt is live?", the one question SP4's migration answers, had no observable
|
||||
answer.
|
||||
|
||||
`tools.stress_session --rich` fills in precisely those and nothing else:
|
||||
|
||||
```
|
||||
cd backend
|
||||
.venv/Scripts/python.exe -m tools.stress_session --rich --actions 30 --memories 12
|
||||
```
|
||||
|
||||
An RPG scenario built from the real seed schema with a world state played forward (hp
|
||||
100 → 71, flags flipping partway); per-action `state_before`/`world_state_before`
|
||||
snapshots that are monotonic, so a bad rollback reads as a wrong number rather than as
|
||||
nothing; a gold script on the adventure; story cards; non-zero memory and summary
|
||||
cursors with a real `story_summary`; pinned and forgotten memories; retry attempts with
|
||||
distinct texts, counts of 2 *and* 3, and a live attempt that is **often not the last one
|
||||
written**; and a second adventure, so "does this leak across adventures?" is answerable —
|
||||
a branch clause that forgot its adventure would still look correct on a database holding
|
||||
exactly one.
|
||||
|
||||
It is a correctness fixture, so prefer it small: what it is for is variety per row, not
|
||||
rows. **Its byte figures are not comparable to a plain run** and it does not replace the
|
||||
scale fixture — the plain one still holds the egress ceilings, and was re-measured
|
||||
unchanged (1.8 kB, actions 1.7 kB) after `--rich` was added.
|
||||
|
||||
The invariant `text == variants[variant_index]["text"]` holds on every retried row, and
|
||||
is asserted when the fixture is built. SP4's migration reads exactly that to decide which
|
||||
sibling becomes the head.
|
||||
|
||||
### SP1 — Schema and migration *(no behaviour change)*
|
||||
|
||||
| File | Change |
|
||||
|---|---|
|
||||
| `app/models.py` | New `Branch` model. `Action.branch_id`, `Action.depth`. `Adventure.head_branch_id`, `head_depth`. `Memory.branch_id`, `Memory.depth`. Legacy columns untouched. |
|
||||
| `app/migrations.py` | Migrations 46+: create `branches`; add columns; index `(branch_id, depth)`; backfill. Dialect map wherever BLOB/BYTEA-style spellings diverge. |
|
||||
|
||||
Backfill: one branch per adventure, `lineage = [(A, ∞)]`; `actions.branch_id = A`,
|
||||
`depth = index`; `adventures.head_*` from `max(index)`; memories take branch A and a
|
||||
depth derived from `source_end`.
|
||||
|
||||
**Verify:** baseline test unmodified. New `tests/test_tree_migration.py` — every action
|
||||
carries a branch and `depth == old index`, no row lost, head pointers correct, memories
|
||||
mapped. Bootstrap run twice is a no-op. `tests/test_egress.py` ceilings unmoved (two
|
||||
integers a row). **Post-deploy `VACUUM FULL actions;` is mandatory** — this rewrites
|
||||
every row, which is the 144 MB lesson at the top of `STATUS.md`.
|
||||
|
||||
**Done, 2026-08-17** (branch `sp1-tree-schema`). **297 tests green**, the 283 from SP0
|
||||
plus 14 in `test_tree_migration.py`. The baseline contract passes **unmodified**, which
|
||||
was the pass condition. Five things the plan did not anticipate, all of them found by
|
||||
building it:
|
||||
|
||||
- **The writes could not wait for SP2.** The file table above lists only `models.py` and
|
||||
`migrations.py`, but a migration never visits a row written *after* it runs — so
|
||||
shipping the columns without a writer would leave every turn played between the two
|
||||
deploys with no branch, and from SP2 on a row with no branch is a row no read can see.
|
||||
`app/tree.py` is that writer: `root_branch` / `head_branch` (get-or-create),
|
||||
`place_action`, `place_memory`, `refresh_head`. One module for the same reason SP2 gets
|
||||
one — a node written without a branch fails by *disappearing*, not by raising. Wired
|
||||
into create/turn/import/undo/delete/memory, plus `seed_demo.py` and
|
||||
`tools.stress_session` (a fixture built by `create_all` is stamped LATEST, so no
|
||||
migration ever runs against it).
|
||||
- **`adventures.head_branch_id` cannot be a foreign key.** `branches.adventure_id`
|
||||
already points the other way, and the pair is then a cycle `create_all` refuses to
|
||||
order; the fix for that is `use_alter`, which SQLite has no ALTER for. It is a plain
|
||||
integer, documented as a cache, and `head_branch` recovers onto the root if it ever
|
||||
names a branch that is gone.
|
||||
- **`lineage` is NOT NULL, because `branches` comes from `create_all`.** The backfill
|
||||
cannot insert a row and fill the lineage afterwards via a NULL marker, so it inserts
|
||||
`'[]'` and guards step two on `json_array_length(lineage) = 0` — not `= '[]'`, because
|
||||
Postgres `json` has no equality operator.
|
||||
- **SQLite cannot drop a column a foreign key names.** So a current-schema database
|
||||
cannot be rewound past `branch_id` at all, which broke the two existing tests that
|
||||
simulate an old database by rewinding only the *stamp*. Fixed properly:
|
||||
`migrations._column_already_there` makes every `ADD COLUMN` idempotent (the
|
||||
`IF NOT EXISTS` the module docstring asks for and SQLite has no syntax for), and
|
||||
`tests/schema_rewind.py` holds the inverse of the migrations that *can* be undone.
|
||||
The SP1 fixture therefore builds a **genuine schema 45** by dropping the three tables
|
||||
and recreating them from frozen pre-tree DDL, so the real ALTERs run — including the
|
||||
one that adds a foreign key.
|
||||
- **Deleting a branch takes its nodes with it**, and deleting an adventure takes its
|
||||
branch — both verified, both at the database level via `ON DELETE CASCADE` on the two
|
||||
`branch_id` columns. SP7's delete-a-branch needs no code of its own for the nodes.
|
||||
|
||||
Measured: `branches` costs **0.1 kB of a 733.5 kB turn (0 %)**, and the page load
|
||||
(62.6 kB) and index (1.8 kB) shapes are byte-identical to the figures in `STATUS.md`.
|
||||
One cost that is *not* free: the new table and its index add ~47 ms to every
|
||||
`create_all`/`drop_all` cycle on SQLite, and the suite does one per test — 20 s → 38 s.
|
||||
Test-only (DDL fsync), so no model change; if it ever matters, the fix is the test
|
||||
harness, not the schema.
|
||||
|
||||
### SP2 — The branch clause *(reads move to lineage; still one branch)*
|
||||
|
||||
One module owns the clause; every action read goes through it. A forgotten clause shows
|
||||
the wrong story, quietly, which is why it is one module and not a convention.
|
||||
|
||||
| File | Change |
|
||||
|---|---|
|
||||
| `app/context/lineage.py` *(new)* | Lineage computation and the branch clause. The only place that knows how a path is selected. |
|
||||
| `app/context/history.py` | `_filters` takes the clause; order by `depth`; `window_covering` keeps its shape. **Fix `_from_memory`.** |
|
||||
| `app/routers/adventures.py` | `action_window` anchors on the anchor's `depth`; `last_action`, `next_index`→`next_depth`, `_latest_narration`. |
|
||||
| `app/scripting/pipeline.py` | `_history()`/`_info()` read the path, not the collection. |
|
||||
|
||||
**Verify:** baseline test unmodified. `test_history_window.py`, `test_action_paging.py`
|
||||
green. New test builds a **two-branch fixture directly in the DB** and asserts the design
|
||||
doc's own example reads back as `A0 A1 A2 A3 B4 B5 C6 C7`, and that a sibling's nodes are
|
||||
invisible. Egress: a 20-fork fixture costs within a small factor of a 1-fork one —
|
||||
clause count is bounded by the context window, not by fork count.
|
||||
|
||||
**Done, 2026-08-17** (branch `sp2-branch-clause`). **317 tests green**, the 297 from SP1
|
||||
plus 20 in `test_branch_clause.py`. The baseline contract passes **unmodified**, which
|
||||
was the pass condition. Six things worth not rediscovering:
|
||||
|
||||
- **The contract forced the write side, not the read side.** The SP0 baseline writes its
|
||||
actions straight to the database and must pass unmodified — so "every writer calls
|
||||
`place_action`" could not be the invariant, because the baseline is a writer and does
|
||||
not. Neither did any of the eleven other test fixtures. The alternative was a read
|
||||
tolerant of a NULL branch, which is the quiet-wrong-story failure this subphase exists
|
||||
to make impossible. So the session enforces it instead: `tree.place_new_nodes` runs
|
||||
from `Session.before_flush` and places anything unplaced, registered in `models.py` so
|
||||
that importing the models arms it. The call sites keep their explicit calls — a node
|
||||
placed at the call site is placed *before* the code around it reads the row back.
|
||||
- **A branch could no longer be created with a flush.** `root_branch` did
|
||||
`db.add(); db.flush()` to get the id its lineage names, and a nested flush inside
|
||||
`before_flush` raises. It inserts through Core and reads the row back — same
|
||||
transaction, three statements, once per adventure ever.
|
||||
- **The identity map holds weak references, and it cost 25 % of the suite.** Resolving
|
||||
the head branch per node re-read the `branches` row for every node in the flush, because
|
||||
nothing held a strong reference between two calls: **201 SELECTs to write 200 actions**,
|
||||
and 36 s → 45 s on the same 297 tests measured back to back. The head is now resolved
|
||||
once per adventure per flush — **2 SELECTs**, and the same 297 tests then time within
|
||||
noise of SP1 (44.2 s each; this machine's load drifts by ~20 % between runs, so trust
|
||||
the statement count, not the stopwatch). Pinned by a test that counts reads of
|
||||
`branches`, because the stopwatch is all the symptom there ever was.
|
||||
- **The index screen is the one read scoped by head branch rather than by lineage.**
|
||||
`_latest_narration` picks one row per adventure for a hundred adventures at once, and a
|
||||
lineage clause each would put hundreds of OR-terms on that query. The two answers differ
|
||||
only for a branch with no nodes of its own, which cannot exist — a branch is created by
|
||||
playing a turn onto it.
|
||||
- **Two reads are deliberately left un-pathed**, both documented where they live.
|
||||
`max_action_index` allocates the legacy `index`, which must stay adventure-wide or two
|
||||
branches issue the same number; and export is a flat v1 bundle whose reader has no idea
|
||||
branches exist, which is why SP6 replaces the format rather than widening the query.
|
||||
A third is a known divergence, not a decision: the index screen's `action_count` counts
|
||||
the tree, and will overstate a branched story until SP5.
|
||||
- **`Adventure.actions` was left ordered by `index` on purpose.** Ordering the collection
|
||||
by depth would not make it a story — it is every branch's actions, and a path is a
|
||||
selection out of it. What the relationship is still for is ownership and the
|
||||
delete-orphan cascade.
|
||||
|
||||
Measured: a story forked **20 times reads its newest 32-action window for 3,178 B against
|
||||
the 2,961 B an unforked story of the same length costs (1.07×)**, naming one branch of its
|
||||
22 lineage entries. The pre-tree 600-action `--keep` fixture was migrated and then driven
|
||||
over HTTP end to end: index **1,840 B**, page load **64,149 B** — the same shapes as
|
||||
before the phase — and scrolling to the start took 9 pages and saw all 600 actions exactly
|
||||
once.
|
||||
|
||||
### SP3 — Memories and summary attach to nodes
|
||||
|
||||
Cursors stop being positions in a shifting list. `memory_cursor`/`summary_cursor` become
|
||||
node-anchored `(branch_id, depth)`.
|
||||
|
||||
Deleted: `position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
|
||||
`prune_dangling_memories`, and their call sites in `undo_turn` and `delete_action`.
|
||||
**Kept until SP4:** `settled_story_actions` and the holdback (see traps).
|
||||
|
||||
**Verify:** baseline test unmodified. `test_memory_settling.py` and
|
||||
`test_memory_retrieval.py` updated, plus a new branch-isolation test — a memory created
|
||||
on branch B is invisible from branch A, and shared ancestors are visible from both.
|
||||
Memory retrieval reads the *full* lineage (it cannot be windowed) but stays sparse:
|
||||
assert the byte cost on a deep fork.
|
||||
|
||||
**Done, 2026-08-18** (branch `sp3-node-cursors`). **330 tests green**, the 318 the branch
|
||||
started from plus 12 — 11 in `test_branch_clause`'s new sibling `test_memory_nodes.py`
|
||||
and one on migration 56. The baseline contract passes **unmodified**, which was the pass
|
||||
condition. `app/context/cursors.py` is the new module; migrations 53–56 add
|
||||
`memory_cursor_branch_id/_depth` and `summary_cursor_branch_id/_depth` and translate the
|
||||
old counts into them.
|
||||
|
||||
Seven things worth not rediscovering:
|
||||
|
||||
- **An anchor is a coordinate, not a pointer, and that is what deleted the machinery.**
|
||||
Half of this subphase was expected to be rewriting the cursor bookkeeping in depth
|
||||
terms. None of it needed rewriting: `count_after(41)` is well defined with node 41
|
||||
deleted, and deleting node 12 does not change what "past node 41" means. So
|
||||
`note_action_removed`, `_rewind_cursors_to_index`, `position_of_index` and the
|
||||
post-turn clamp did not become depth-shaped versions of themselves — they became
|
||||
nothing. **If a mark still needs correcting when the story changes, it is still a
|
||||
position.**
|
||||
- **The clamp had its own trap and it also goes.** `run_post_turn` clamped both cursors
|
||||
to the story length every pass, deliberately against the *full* count, because
|
||||
clamping to the settled count rewound a caught-up adventure a step and re-covered an
|
||||
action. An anchor past the tip is not a broken value: `settled_after` reports nothing
|
||||
to do, and the story growing back past it resumes exactly where it left off.
|
||||
- **`prune_dangling_memories` became a lookup, and got stricter by accident.** A memory
|
||||
hangs off the node its block ends on, so "what did this node produce?" is
|
||||
`(branch_id, depth)` — `memorybank.forget_node`. The scan it replaces could only ever
|
||||
notice damage *after* the fact (a covered range past `max(index)`), and could not
|
||||
notice at all when the node was deleted from the middle of a story that still had
|
||||
later actions. Withdrawing the memory is half the job: the ground it covered is still
|
||||
behind the mark, so the mark goes back to `source_start - 1` — a depth, whether or not
|
||||
a row still sits there.
|
||||
- **A memory with no node had to be spelled out in the clause.** A hand-written memory
|
||||
summarises nothing, so it carries a branch and a NULL depth. Every ancestor entry in a
|
||||
lineage clause is capped `depth <= fork`, and NULL fails that — so a typed memory would
|
||||
have become invisible at the first fork after it was written, with nothing to see but a
|
||||
prompt that stopped mentioning it. `Path.clause(unanchored=True)` is that case, and
|
||||
actions never pass it: an action with no depth is a pre-tree row no read should see.
|
||||
- **Retrieval reads the whole lineage, and it is free.** Measured on two stories of 84
|
||||
actions and 14 memories each, one flat and one forked twenty times: **1,807 B against
|
||||
1,823 B**. The clause carries 22 branch terms instead of one, and the clause is not what
|
||||
crosses the wire. The egress shapes are otherwise byte-identical to SP1's — index
|
||||
1.8 kB, page load 62.7 kB, turn 733.8 kB.
|
||||
- **Three reads stay adventure-wide, deliberately.** Embedding and eviction are facts
|
||||
about the row and about the bank, not about the path — skipping a sibling's memories
|
||||
would only mean embedding them at the moment somebody switched to them, and evicting the
|
||||
memories of a story nobody is reading is the right thing to evict first. The Memories
|
||||
drawer is management rather than retrieval, and hiding a branch's memories there would
|
||||
make them unfindable in a phase whose rule is that nothing is removed automatically.
|
||||
- **The v1 bundle still speaks positions, in exactly two places.** Export counts the
|
||||
anchor back into a position; import translates the other way, but only after the
|
||||
actions exist, because that is the one moment the two coordinate systems can be lined
|
||||
up. `cursors.position_of` and `cursors.anchor_at_position` are the whole of what still
|
||||
knows about positions, and SP6's v2 format retires them.
|
||||
|
||||
Migrations 53–56 rewrite `adventures`, not `actions` — a few hundred rows against a few
|
||||
hundred thousand — so **this deploy needs no `VACUUM FULL` of its own**. The one SP1 owes
|
||||
is still owed.
|
||||
|
||||
### SP4 — Variants become sibling nodes
|
||||
|
||||
Retry stops mutating a row. It writes a sibling leaf at the same depth.
|
||||
|
||||
Deleted: `set_variants`, `variant_of`, `apply_variant`, `VARIANT_SNAPSHOT_KEYS`, and the
|
||||
holdback. Node state moves from *before* to *after* snapshots (`state_after`,
|
||||
`world_state_after`) — a sibling needs its own outcome, so this belongs here rather than
|
||||
in its own subphase. Pre-column rows are NULL and must stay tolerated.
|
||||
|
||||
A migration converts existing `variants` JSON into sibling rows — reading the legacy
|
||||
column that decision 2 kept, which is the whole reason it was kept.
|
||||
|
||||
`/variants` and `/variant` keep their URLs and response shapes here, re-implemented over
|
||||
sibling rows, so the frontend keeps working until SP7 replaces it.
|
||||
|
||||
**Verify:** the first subphase allowed to move the baseline test, and only for
|
||||
`variant_count`/`variant_index` semantics — the retry *outcomes* must not move.
|
||||
`test_retry_variants.py` rewritten against siblings but asserting the same observable
|
||||
results, including that the gold script is not double-applied. `test_state_revert.py`
|
||||
against after-snapshots. Migration test: attempt count preserved, the active attempt
|
||||
becomes the head.
|
||||
|
||||
**Done, 2026-08-18** (branch `sp4-sibling-nodes`). **347 tests green**, the 330 SP3
|
||||
finished with plus 13 in the new `test_attempt_siblings.py`, 8 in `test_tree_migration
|
||||
.py`, and three holdback tests deleted. `app/attempts.py` is the new module; migrations
|
||||
57–60 add `live`, `state_after`, `world_state_after`, derive the after-snapshots, and
|
||||
split every `variants` list into rows.
|
||||
|
||||
**The baseline test did not have to move, and neither did `test_retry_variants.py`.**
|
||||
Both pass unmodified. SP4 was *permitted* to change the variant-count semantics and it
|
||||
turned out nothing observable needed changing — which is the strongest form the pass
|
||||
condition could have taken, and worth knowing before SP6 asks for the same licence.
|
||||
|
||||
Seven things worth not rediscovering:
|
||||
|
||||
- **A coordinate needs a `live` flag, and the branch clause is where it belongs.**
|
||||
Siblings share `(branch_id, depth)`, so the lineage clause alone returns all of them
|
||||
and the story tells itself twice. `Path.clause` adds `Action.live` for actions (and
|
||||
only for actions — memories have no siblings to lose to), which means no read of the
|
||||
story had to learn that retries exist. `app/attempts.py` is the only code that looks
|
||||
past it.
|
||||
- **The prompt has to move with the flag or retry becomes a storage multiplier.** A
|
||||
`context_snapshot` is ~163 kB of prompt that every attempt at a turn shares, and the
|
||||
old JSON list existed precisely to store it once. Giving each sibling row a copy would
|
||||
have undone that. So the invariant is: the assembled prompt lives on the attempt the
|
||||
story tells, and a superseded one keeps only its own slices (`ATTEMPT_KEYS` — the
|
||||
world-state delta, the script report, the raw reply). Measured on the pre-tree
|
||||
600-action fixture, migrated: **700 rows for the same 600-turn story, and the prompt
|
||||
archive byte-identical at 0.50 MB.**
|
||||
- **The holdback was not quite unnecessary — it was the wrong repair.** The plan said
|
||||
retry stops mutating rows so nothing goes stale. Not so: siblings share a coordinate,
|
||||
and the mark and the memory both name the *coordinate*, so replacing what a turn says
|
||||
still invalidates them. What made the holdback deletable is that the right repair
|
||||
already existed — `forget_node` plus a rewind, which undo and delete have called since
|
||||
SP3. Retry and a sibling switch now call it too, and the holdback (`settled_count`,
|
||||
`settled_after`, `settled_story_actions`, `newest_settled`) is gone. **A memory can
|
||||
now cover the newest action**, which it never could before.
|
||||
- **`state_after` needed the same flush guard `place_action` has.** The migration
|
||||
derives every existing row's outcome from the next row's `state_before`, and the tip's
|
||||
from the adventure's live state — but a row written by a fixture or a script after
|
||||
that has nobody to derive from, and the failure mode is undo silently leaving the
|
||||
scoreboard where it was. `tree.stamp_outcome` runs from `before_flush` beside
|
||||
`place_new_nodes`. It writes the truth as of the flush: a writer that changes no state
|
||||
between two nodes leaves the same state behind both.
|
||||
- **Switching attempts hands the caller a different row id**, because that is what an
|
||||
attempt being a node *means*. The endpoints are addressed by any attempt at the turn
|
||||
rather than by the live one, so a client holding a stale id still asks about the right
|
||||
turn — but `Play.jsx` matched the reply against `updated.id` and had to be changed to
|
||||
match against the action it asked about. One line, and SP7 removes the pager anyway.
|
||||
- **Deleting a turn deletes its attempts.** A discarded attempt is only reachable
|
||||
*through* its coordinate, so leaving it behind would leave a row nothing can name and
|
||||
no read can see. `delete_turn` is that rule in one place, called by undo and by
|
||||
delete-an-action, and it works whichever attempt's id the caller happens to hold.
|
||||
- **Siblings share the legacy `index`.** They are takes on one turn, so two rows now
|
||||
carry one index — which `max_action_index` (a maximum, not a count) survives, and
|
||||
which is what lets the v1 export fold a group back into a `variants` array. Export is
|
||||
now the only producer of that shape anywhere; nothing in the database holds one.
|
||||
|
||||
Measured: index **1.8 kB**, page load **62.7 kB**, one turn **734.8 kB** — the first two
|
||||
byte-identical to SP1's and SP3's, the turn up 1.0 kB (0.14 %) for the `live` column
|
||||
across a 346-row read. Migrations 57–59 each rewrite every row of `actions` and 60
|
||||
inserts one per discarded attempt, so **this deploy owes a `VACUUM FULL actions;`** —
|
||||
the SP1 one is still owed too, and one vacuum after this deploy settles both.
|
||||
|
||||
### SP5 — Fork on continue
|
||||
|
||||
Playing past a non-head sibling promotes it: a `branches` row with `parent_branch_id`,
|
||||
`fork_depth`, and a `lineage` computed once from the parent's. Nothing is copied.
|
||||
`adventures.head_*` moves. **The cooldown clock must not advance on a re-run of the same
|
||||
turn** — the one item still open above.
|
||||
|
||||
**Verify:** a fork creates exactly one branch row and copies no actions; lineage is
|
||||
correct and capped at each `fork_depth`; both branches read independently; switching
|
||||
restores the right script/world state; the cooldown test from `test_worldstate.py`
|
||||
still holds across a retry.
|
||||
|
||||
**Done, 2026-08-18** (branch `sp5-fork-on-continue`). **365 tests green**, the 347 from
|
||||
SP4 plus 18 in `test_branch_forking.py`. `tree.fork` is the whole of it; three endpoints
|
||||
(`GET /branches`, `POST /branches/{id}/switch`, `POST /actions/{id}/fork`) are what SP7's
|
||||
tree view will be drawn on.
|
||||
|
||||
**Where the promotion happens, and why not where the plan said.** The plan put it on the
|
||||
*next turn*: attempts stay leaves, and one becomes a branch when a turn is played past
|
||||
it. Promoting the winner then means moving a row off the branch the reader is standing
|
||||
on and leaving that branch to pick a new node for the depth — it disturbs a story nobody
|
||||
asked to change. The same divergence, seen from the other side, promotes the attempt you
|
||||
are *leaving for*: `POST /actions/{id}/fork` gives the discarded attempt a branch and
|
||||
moves the head to it, and the line it leaves is untouched. Branch count is identical
|
||||
either way — one per divergence somebody actually built on — and one of the two never
|
||||
rewrites a story in place. Without it, the losing attempts would also be unreachable
|
||||
forever, since only the tip can be switched: fork-on-continue alone is a one-way door.
|
||||
|
||||
Six things worth not rediscovering:
|
||||
|
||||
- **A fork must not move the derived work, and it was about to.** The first cut moved the
|
||||
memories at the forked coordinate onto the new branch and re-anchored the cursors that
|
||||
named it. Both are wrong, and for the same reason: a memory describes whichever attempt
|
||||
was *live* at that coordinate, which is the one staying on the parent. The right answer
|
||||
needs no code at all — the lineage caps the parent at `fork_depth`, one depth short of
|
||||
it, so the memory is simply out of range from the fork, invisible to both the retrieval
|
||||
clause and `Path.depth_on`. The block is summarized again, from the text this branch
|
||||
actually tells.
|
||||
- **Depth had to stop following `index`.** They agreed until now. `index` is
|
||||
adventure-wide (the v1 bundle is keyed on it) so on a story forked at depth 6 after
|
||||
twenty turns it hands the next node depth 21 and leaves a fourteen-deep hole in the
|
||||
middle of a path — which every windowing estimate then has to work around.
|
||||
`next_depth` is `head_depth + 1`; `place_action` still derives depth from index for
|
||||
fixtures and imports, which is what that default is for.
|
||||
- **Undo has to stop at the fork.** It reads the newest two nodes on the path and deletes
|
||||
the turn they make up — and on a fresh branch the second of them is borrowed from the
|
||||
parent, whose story also contains it. The guard is on the row's own `branch_id`, not on
|
||||
the fork depth, because that is the fact that decides it.
|
||||
- **The cooldown clock came out right for free.** It lives in `_meta.last_changed` inside
|
||||
the world state, and the world state is restored from the tip's `world_state_after` on
|
||||
every switch — so each branch carries its own clock without anything knowing there is
|
||||
one. The carried-over open item (a retry must not advance it) is SP4's reused depth,
|
||||
and both are pinned by tests.
|
||||
- **The session does not autoflush**, and `fork` read the sibling group after moving the
|
||||
node out of it — so the move had not been written and the node was renumbered straight
|
||||
back into the group it had just left. Read the group first. (`autoflush=False` is
|
||||
deliberate, in `database.py`; anything in this phase that mutates then queries the same
|
||||
rows has to order itself by hand.)
|
||||
- **`/fork` has to be idempotent before it is anything else**, because a fork leaves the
|
||||
promoted attempt alone on its branch: a repeated call — a double click, a retried
|
||||
request — would otherwise be told the turn it just forked has nothing to fork to. The
|
||||
"already the story" case is answered before the shape of the turn is looked at.
|
||||
|
||||
Measured, on a 40-turn story forked **twenty** times against the same story flat:
|
||||
21 branches, 140 rows, an 80-action story, and a page load of **31,652 B against
|
||||
31,433 B (1.007×)**. A branch costs **103 B** of id, parent, fork depth and cached
|
||||
ancestry. No migration, no vacuum.
|
||||
|
||||
**One gap, deliberately left for SP6.** A forked adventure has no honest `v1` export —
|
||||
the format has one story and there are two — so export emits every branch's turns
|
||||
interleaved by index, which reads as a mangled story rather than as lost data. SP6's v2
|
||||
bundle fixes it, and SP7 is where a player first gets any way to fork at all, so the
|
||||
order those two ship in is the order that matters.
|
||||
|
||||
**Also known, and not fixed here:** the two cursors are one pair on the adventure, so
|
||||
switching branches makes the mark on the branch being left unreadable from the new one
|
||||
(`Path.depth_on` answers "nothing covered", which is the safe direction — redo the work,
|
||||
never skip it). Switching back and forth therefore re-summarizes. Per-branch cursors are
|
||||
the fix if it ever matters; it costs AI calls, not correctness.
|
||||
|
||||
### SP6 — Export/import v2
|
||||
|
||||
`ai-dnd-adventure-v2` carries branches and nodes; import accepts v1 and v2, mapping a v1
|
||||
bundle's linear actions plus `variants` onto one branch with siblings.
|
||||
`limits.check_bundle_lists` learns about branches.
|
||||
|
||||
**Verify:** v2 round-trip of a branched adventure is lossless; a v1 bundle still imports;
|
||||
a bundle claiming more branches than rows is rejected rather than half-applied.
|
||||
|
||||
**Done, 2026-08-18** (branch `sp6-bundle-v2`). **381 tests green**, the 365 SP5 finished
|
||||
with plus 16 in the new `test_bundle_v2.py`. `app/bundle.py` owns both formats; the two
|
||||
endpoints in `routers/adventures.py` are a delegation and the shared plumbing, and the
|
||||
`variants` array now exists nowhere but the v1 *reader*.
|
||||
|
||||
**The rule the module is built on: a bundle carries what was *chosen*, never what is
|
||||
*derived*.** The head branch, the fork points, the live flags and the anchors are
|
||||
decisions somebody made, and they are in the file. `lineage`, the head *depth*, `index`
|
||||
and the variant ordinals are computed from those and are rebuilt on the way in. That is
|
||||
not tidiness — a bundle is a text file anybody can edit, and every derived field shipped
|
||||
beside its source is a chance for the file to disagree with itself in a way no read would
|
||||
report. It is also the answer to "is the round trip lossless?": everything omitted is
|
||||
reconstructed, and the tests assert the reconstruction rather than the bytes.
|
||||
|
||||
Six things worth not rediscovering:
|
||||
|
||||
- **`depth` cannot be the legacy `index`, and this is where that stops being academic.**
|
||||
They agreed until SP5, and a bundle is the first writer that has to fill `index` for a
|
||||
*forked* story — where two branches both have a node at depth 4. `index`'s one
|
||||
remaining job is handing the next row a number nothing else holds, which is a fact
|
||||
about the adventure rather than about a path, so the import allocates one per turn in
|
||||
bundle order: siblings share it, the way SP4 leaves them, and no two coordinates do.
|
||||
- **Validation happens before the adventure row exists.** Everything a hand-edited file
|
||||
can get wrong about the shape of a tree — a node naming a branch that is not listed, a
|
||||
fork with no depth, a branch forking from one listed after it — is a 400 raised by
|
||||
`plan`, which touches no session. The alternative is an adventure holding a story with
|
||||
a hole in it, and this whole phase exists because a story with a hole in it fails by
|
||||
going quiet.
|
||||
- **A branch may only fork from one listed before it.** That is how the export writes
|
||||
them, and requiring it buys acyclicity for the price of a comparison — a cycle in the
|
||||
parent chain would be an import that never returns rather than one that fails.
|
||||
- **`{}` and absent are different snapshots.** An empty `state_after` means "this node
|
||||
left an empty scoreboard behind"; a missing one means "nobody knows, leave the live
|
||||
state alone" (`attempts.restore_state`). Trimming empty dicts on the way out would have
|
||||
saved eighteen bytes a row and turned an undo that clears a score into one that leaves
|
||||
it standing. Only `worldDelta`, which is display, is dropped when empty.
|
||||
- **A file is allowed to be wrong about which attempt is live, and the import corrects it
|
||||
rather than refusing.** "Exactly one sibling in a group is live" is an invariant of the
|
||||
*database*, not of the format; a coordinate with none is a turn no read can see, so the
|
||||
first attempt is made live. That is a different class from a missing branch, which is
|
||||
structure, and is refused.
|
||||
- **`limits.MAX_BRANCHES_PER_ADVENTURE` (1000) has no live counterpart.** Forking is a
|
||||
POST that adds one row and has no cap of its own, so this is the one bundle cap that
|
||||
does not mirror something creation enforces. Worth closing if branch management ever
|
||||
makes forking cheap to repeat.
|
||||
|
||||
**The verify line above was slightly wrong, and the code does the honest version.** "More
|
||||
branches than rows" fails on an adventure with no actions, which legitimately has one
|
||||
branch and no rows. It became two rules instead: a cap on the branch list, and *every
|
||||
branch a node names must exist*.
|
||||
|
||||
Measured with `tools/measure_bundle.py` on the 600-action `--rich` fixture — 600 turns,
|
||||
750 nodes, because 150 of them were retried:
|
||||
|
||||
| | bytes | vs v1 |
|
||||
|---|---|---|
|
||||
| v1 shape | 587,475 | — |
|
||||
| **v2** | **911,229** | **1.551×** |
|
||||
| v2 without the outcomes | 544,318 | 0.927× |
|
||||
|
||||
**The tree is free; the outcomes are what cost.** Coordinates *save* 57.5 B a node
|
||||
against v1's turn-and-variants shape, and the entire 1.55× is `state_after` /
|
||||
`world_state_after` at 489 B a node — which are there because a bundle without them
|
||||
imports a tree nobody can switch inside. Twenty forks add 660 B to the same file, **33 B
|
||||
a branch**, so the format is as indifferent to fork count as the reads are. The longest
|
||||
adventure production holds exports at 4.3 % of `MAX_IMPORT_BODY_BYTES`.
|
||||
|
||||
**Two lines of the SP0 baseline changed, not the one it predicted.** The note in
|
||||
`test_story_tree_baseline.py` allowed for the `format` assertion; the `variants` array in
|
||||
`test_export_keeps_retry_attempts` is the same fact from the other side — a bundle with
|
||||
coordinates has no use for a repeating group. Everything else in that file still passes
|
||||
unmodified.
|
||||
|
||||
**No migration, no vacuum.** SP6 adds no column and rewrites no row.
|
||||
|
||||
**And one trap, paid for once.** `tools/measure_bundle.py` imported `app` before
|
||||
`tools.stress_session`, which is what points `AIDND_DB_PATH` at a throwaway file —
|
||||
`app.database` reads it at module scope. It fails by *working*: the first run seeded a
|
||||
synthetic user and adventure into the local `backend/data.db` and printed perfectly good
|
||||
numbers, and only the second run tripped over the unique email. Anything importing that
|
||||
harness must import it first, and the file now says so where the imports are.
|
||||
|
||||
### SP7 — Frontend: the tree becomes reachable
|
||||
|
||||
`VariantPager` is removed. A branch view replaces it, plus switch, rename and
|
||||
delete-with-confirm. `api.js` gains the branch endpoints.
|
||||
|
||||
**Verify:** this is where the standing open gap gets closed — **drive the 600-action
|
||||
`--keep` fixture in a browser by hand**, on the same scroll path that has never been
|
||||
driven and already hid one bug. A vitest + jsdom harness covers the prepend arithmetic;
|
||||
jsdom has no layout, so scroll position still needs eyes.
|
||||
|
||||
**Done, 2026-08-18** (branch `sp7-tree-ui`). **396 tests green**, the 381 SP6 finished
|
||||
with plus 15 in the new `test_branch_management.py`. Driven by hand against the `--keep`
|
||||
fixture in Chrome, which is how the one bug below was found.
|
||||
|
||||
**Shape chosen: a branch rail, not a spatial node map.** Three mockups were built and
|
||||
compared before any of it was written, and the deciding argument was not aesthetic. A
|
||||
node map is a second windowing problem — the fixture this subphase must be verified
|
||||
against is 600 actions, which is 600 nodes — so building one would have spent SP7 on the
|
||||
thing that delays the verification SP7 exists to do. The rail ships now; the map is a
|
||||
later feature and costs nothing extra to add, because both draw from the same
|
||||
`GET /branches`. The panel sits beside Plot/Memory/Scripts/Insights, which is this app's
|
||||
existing idiom for a right-hand rail rather than a new one.
|
||||
|
||||
**SP7 was not a frontend-only subphase, and the spec above did not say so.** Of the three
|
||||
operations it names, SP5 had built exactly one. `switch` existed; `rename` had no column
|
||||
and no route, `delete` had no route at all. So it opens with migration 61
|
||||
(`branches.name`), a `PATCH` and a `DELETE` — worth remembering for any future subphase
|
||||
whose one-line spec says "plus the UI for X".
|
||||
|
||||
Five things worth not rediscovering:
|
||||
|
||||
- **A name is stored; a label is derived.** `branches.name` is NULL until somebody
|
||||
chooses one, and the client draws an unnamed branch from its fork depth
|
||||
(`Fork at moment 547`). A generated "branch 4" in the column would be a lie the moment
|
||||
branch 3 is deleted and the ordinals shift under it; a fork depth is a coordinate, and
|
||||
nothing can shift it. The v2 bundle carries the name for exactly the reason SP6 gives
|
||||
for carrying the fork points — it is a decision, not something computed from one.
|
||||
- **Refusing to delete the head is only half of it.** The other half is refusing any
|
||||
branch the head was *forked from*: `parent_branch_id` cascades, so deleting an ancestor
|
||||
takes the head with it and leaves `head_branch_id` pointing at a row that is gone. One
|
||||
membership test against the head's own `lineage` covers both, because a lineage already
|
||||
names itself and every branch it borrows from.
|
||||
- **A deleted branch's cursor has to be cleared, and the reason is SQLite.** On Postgres
|
||||
a stale branch id simply never resolves. SQLite hands the freed id to the next fork, at
|
||||
which point the anchor resolves onto a branch it has never seen and reports a stretch
|
||||
of story as already summarized — losing it from the memories for good. Same class as
|
||||
the width-mismatch `cosine` returning 0.0: it reports nothing.
|
||||
- **`VariantOut` had to grow an `id`.** A fork is addressed by the node being taken. The
|
||||
group renumbers whenever an attempt is added, so an ordinal held across that points at
|
||||
a different take — the same reason SP4's note called the pager's index match "one line,
|
||||
and SP7 removes the pager anyway".
|
||||
- **Four panels reloaded on the wrong thing, and only a browser could say so.**
|
||||
`actions.length` was the refresh key for the Branches panel, the Status drawer
|
||||
(script state), Insights and the Memory Bank. **A branch switch does not change the
|
||||
length of the story** — it changes which story it is. So the Branches panel drew a
|
||||
one-branch tree while the reader was already on a second, Insights showed the prompt
|
||||
for the path just left, and the scoreboard kept the other line's numbers. The server
|
||||
was correct throughout, and nothing in 396 tests could see any of it. All four now key
|
||||
on `${actions.length}:${stateKey}`, and `stateKey` is bumped by `adoptWindow` — the two
|
||||
operations that move the head — plus branch deletion, which removes that branch's
|
||||
memories without a turn being played.
|
||||
|
||||
`tools/branch_fixture.py` exists because of this: it builds two branches of **equal
|
||||
path length**, which is the case `actions.length` cannot distinguish at all. The
|
||||
`--keep` fixture could not have found it, and neither could a fixture whose branches
|
||||
happened to differ in length.
|
||||
|
||||
**What a switch puts back was checked end to end, not just server-side.** SP5 already
|
||||
proved `restore_state` in `test_switching_restores_the_script_and_world_state`; what had
|
||||
never been looked at is whether the *screen* re-reads it. On `tools/branch_fixture.py`,
|
||||
switching between the two branches moves the World State drawer from **hp 60 to hp 95**
|
||||
live, redraws the bar, swaps the story to the other take (`hp -5`, not `hp -40`) and
|
||||
repoints Insights at the other path — "History: 5 of 5 actions", carrying the scratch and
|
||||
not the beating.
|
||||
|
||||
**Every memory is now a memory of a story, not of an adventure — migration 62.** A
|
||||
hand-written memory used to carry a NULL depth, described as "belongs to the adventure
|
||||
rather than to a path". That reads as harmless and is not: a NULL is a coordinate no fork
|
||||
can cap, so a note typed on one line followed the reader onto branches whose events it
|
||||
never described. `tree.place_memory` anchors it at the head instead — *the story you were
|
||||
reading when you wrote it* — and it then obeys exactly the rule a summarised memory obeys.
|
||||
|
||||
The whole `unanchored` escape clause in `lineage.Path.clause` existed for that one case
|
||||
and is **deleted**, not merely unused. Its docstring argued a capped `depth <= fork` would
|
||||
drop a typed memory "the moment its branch stopped being the newest entry"; anchoring is
|
||||
the better answer to the same worry, because the memory is not exempt from the path, it is
|
||||
*on* one.
|
||||
|
||||
**The drawer shows the path being read, and nothing else.** Same clause as retrieval, so
|
||||
the bank you can see is the bank the model can see — one question, one answer. The earlier
|
||||
attempt at this shipped an adventure-wide list with an `on_path` flag and an *another
|
||||
branch* badge; anchoring makes that redundant, and the field, the badge and its CSS are
|
||||
gone. Nothing is stranded by hiding: a memory lives on a branch, switching to that branch
|
||||
shows it, and deleting the branch deletes it
|
||||
(`test_deleting_a_branch_deletes_the_memories_written_on_it`).
|
||||
|
||||
Pinning is unchanged and still path-scoped: it decides *order*, the path decides
|
||||
*existence*. A pinned memory on a branch you are not reading is not sent, because the path
|
||||
clause runs before pinning is considered.
|
||||
|
||||
**Migration 62 anchors existing NULL-depth memories at depth 0 of their branch**, not at
|
||||
the tip. 0 is at or before every fork point, so every memory stays visible from exactly
|
||||
the paths it is visible from today — the anchor takes nothing out of anybody's bank on
|
||||
deploy. Anchoring at the tip would have emptied them out of every branch forked earlier
|
||||
than they were typed, on a database with real users on it.
|
||||
|
||||
**The scroll path was driven, and it holds.** Three prepends on the 602-action fixture,
|
||||
60 actions and ~16,200 px each. The same DOM node stayed at viewport top 792 → 787 — a
|
||||
**5 px drift** across the prepend — and the view stayed 48,174 px from the bottom, so PR
|
||||
#2's throw-to-the-end does not reproduce. Console clean. One note for anyone measuring it
|
||||
again: the fixture's prose repeats, so an anchor found by matching *text* lands on an
|
||||
older copy of the same sentence and reads as a huge jump. Hold the DOM node.
|
||||
|
||||
**Still open, deliberately:** no vitest + jsdom harness. The verify line offers
|
||||
hand-driving *or* the harness and this took the first. The harness remains the thing that
|
||||
would catch a prepend regression without a person in the loop, and jsdom's lack of layout
|
||||
means it would not have settled the 5 px question either way.
|
||||
|
||||
**One user-visible change nobody asked for, recorded here because a script
|
||||
author would otherwise find it by being surprised.** Moving `_history` onto
|
||||
`context_history.story_actions` fixed the branch bug it was there to fix, and
|
||||
carried a second change with it: `story_actions` drops blank-text rows, which
|
||||
`adventure.actions` did not. So a user script's `history` array and
|
||||
`info.actionCount` both got shorter for any adventure holding one. It is the
|
||||
right shape — a textless row is bookkeeping with no AI Dungeon counterpart, and
|
||||
the prompt never included one — but a script that fires "every N actions" now
|
||||
fires on different turns. Not compatible with both readings; this is the one
|
||||
chosen.
|
||||
|
||||
### SP8 — Drop the legacy columns
|
||||
|
||||
Only once the tree is proven live. Migration drops `index`, `variants`, `variant_index`,
|
||||
`variant_count`, followed by `VACUUM FULL actions;`. **Also `adventures.memory_cursor`
|
||||
and `summary_cursor`** — unread since SP3, kept only so a rolled-back build resumes from
|
||||
a real number. They are on `adventures`, so dropping them costs no vacuum.
|
||||
|
||||
**And `actions.state_before` / `world_state_before`**, unwritten and unread since SP4 for
|
||||
the same reason: a rolled-back build still finds a real snapshot on every row it wrote
|
||||
itself. They are deferred JSON on `actions`, so they go in the same rewrite as `index`
|
||||
and cost nothing extra.
|
||||
|
||||
`variant_count` and `variant_index` are the two to check before dropping: SP4 left them
|
||||
as a maintained cache of the sibling group's shape, because the pager reads both for
|
||||
every row of a page and must not pay a query per turn to get them. They are dead only
|
||||
once SP7's tree view has replaced the pager.
|
||||
|
||||
**Verify:** full suite; egress ceilings; a measured before/after size, aggregates only.
|
||||
|
||||
### SP9 — Takes, not chips: one pager, and a fork that makes a new take
|
||||
|
||||
**Why this exists.** SP7 was driven by hand and the tree was unusable. Three things were
|
||||
wrong, and only the third is a bug:
|
||||
|
||||
* **A chip did two different things.** At the tip it switched; above the tip it only
|
||||
*previewed*, and taking it needed a second button. One control, two meanings, and the
|
||||
meaning depended on where the player was standing.
|
||||
* **Nothing could fork but an AI turn that already had a second take.**
|
||||
`POST /actions/{id}/fork` answers 400 when the turn has one take, so a player's own
|
||||
message had no way to become anything else and "branch from here" did not exist.
|
||||
* **Retry only worked on the newest action.** Mid-story there was no retry at all.
|
||||
|
||||
**The model, in the player's words.** Every action can gain another *take*. On an AI node
|
||||
that means regenerate; on the player's own it means type something else. Stepping between
|
||||
takes is free — `3/3` to `1/3` is navigation, and the story below simply empties, because
|
||||
that take has no children yet. **A branch is created when you write below a take that is
|
||||
not the live one**, never before.
|
||||
|
||||
That collapses SP5's fork and SP4's retry into one operation and deletes the
|
||||
tip-versus-past distinction from the UI entirely. It survives only in the implementation,
|
||||
where it decides whether a write needs a branch at all.
|
||||
|
||||
**Takes are grouped by parent, not by coordinate.** This is the load-bearing change.
|
||||
`attempts.group()` filters `branch_id == … AND depth == …`, and the player's own example
|
||||
breaks it:
|
||||
|
||||
```
|
||||
B ── C C1 C2 <- three takes, one parent (B)
|
||||
│ └── D1' D2' <- two takes, parent C2
|
||||
└── D1 D2 D3 <- three takes, parent C1
|
||||
```
|
||||
|
||||
Standing on the C2 path at that depth must read `2/2`, not `5`. Coordinate grouping gets
|
||||
that right by accident — writing under a non-live take forks, so the two sets land on
|
||||
different branches. It gets `C` wrong: once C is forked onto its own branch it is alone
|
||||
at its coordinate and reads `1/1`, losing C1 and C2 from a pager that must still say
|
||||
`1/3`.
|
||||
|
||||
**Decision: add `actions.parent_id`.** The alternative — making a branch's fork point a
|
||||
*node* instead of a depth, so a promoted take never moves — was rejected. `lineage` exists
|
||||
precisely so a read is an OR-clause per branch rather than a walk up parent pointers, and
|
||||
re-pointing the fork at a node changes path resolution itself, which drags in `cursors`,
|
||||
memory depths and both bundle formats. `parent_id` is read only to group a turn's takes:
|
||||
one indexed lookup, never a walk, and nothing about how a path resolves changes. Per SP2's
|
||||
sizing note an integer beside `depth` is cheap, and the backfill rewrites a heap measured
|
||||
at 1.7 MB on 2026-08-18.
|
||||
|
||||
**`variant_count` / `variant_index` are reprieved, not revived.** SP8 was going to drop
|
||||
both; the pager needs the group's shape again. It needs it *per parent*, which is not what
|
||||
either column caches, so SP8 still drops them and SP9 computes the shape from `parent_id`.
|
||||
|
||||
**Verify:** the SP0 baseline passes unmodified — a linear story has one take per parent,
|
||||
and none of this is reachable without a second one. Plus: `2/2` under one take while its
|
||||
sibling holds `3/3`; a pager that still reads `1/3` after its take has been forked onto its
|
||||
own branch; a write below a non-live take forks exactly once; a fork on a player's own
|
||||
message generates a reply.
|
||||
|
||||
**Owed:** a migration adding one column, and a `VACUUM FULL actions;` after it.
|
||||
|
||||
## Standing constraints
|
||||
|
||||
- **No production data is read at any point in this phase.** Migrations, the e2e
|
||||
baseline and every egress measurement run on local SQLite and the synthetic `--keep`
|
||||
fixture. If a real Postgres is ever needed for a write path, it is a throwaway
|
||||
database whose name contains `stress`/`scratch`, and it is asked for first.
|
||||
- **Any migration that rewrites `actions` is followed by `VACUUM FULL actions;`** on the
|
||||
direct endpoint, not `-pooler`. SP1, SP4 and SP8 each rewrite every row. **As of
|
||||
2026-08-18 two are owed** (SP1's and SP4's) and neither has been deployed; one vacuum
|
||||
after the SP4 deploy settles both.
|
||||
@@ -1,117 +0,0 @@
|
||||
# Pokémon League Championship demo: playtest handover
|
||||
|
||||
Read this before you resume work on the demo scenario. It covers the current
|
||||
state of `05-league-championship.json`, what a live 10-turn playtest on
|
||||
production confirmed, and two bugs the playtest found.
|
||||
|
||||
**Last updated: 2026-08-28.**
|
||||
|
||||
> **Superseded in part on 2026-08-28.** Both bugs below were investigated against the
|
||||
> real data and **both root causes named here are wrong**. See
|
||||
> `plan/16-world-state-refusals.md` for what was actually happening and what was
|
||||
> changed. The playtest record and the "Confirmed working" section still stand.
|
||||
|
||||
---
|
||||
|
||||
## Where things stand
|
||||
|
||||
The scenario is live at `https://ai-dnd-1gmp.onrender.com` as
|
||||
**"[Demo] League Championship: Round One"**, seeded from
|
||||
`backend/app/seed_data/05-league-championship.json`. It replaced an earlier,
|
||||
weaker draft titled "Road to the Champion" — that old scenario and its stale
|
||||
test adventure (adventure id 42) are still in the production database. Delete
|
||||
them by hand from `/scenarios` and `/adventures` when convenient; a delete
|
||||
click froze the browser tab during this session behind what looked like a
|
||||
native confirm dialog, so budget time for that if you try again.
|
||||
|
||||
The schema nests all five of the player's Pokémon and Milo's Pokémon under
|
||||
`npcs`, not flattened into `player.<name>_hp` fields. Each npc entry carries
|
||||
its own `stats` map (`hp`, `status`, and for Milo, `active_pokemon`,
|
||||
`active_hp`, `active_status`, `pokemon_left`). `player.active_pokemon` is a
|
||||
text stat that names whichever of the player's Pokémon is currently out. This
|
||||
design survived a real playtest: see "Confirmed working" below.
|
||||
|
||||
Settings on the test account now point at the user's own OpenRouter key,
|
||||
endpoint `https://openrouter.ai/api/v1`, model `deepseek/deepseek-v4-flash-0731`,
|
||||
reasoning budget `-1`. The shared demo key's model
|
||||
(`google/gemma-4-26b-a4b-it:free`) was hitting persistent 429s from OpenRouter
|
||||
capacity, not from the app's own rate limiter — switch back to it only after
|
||||
confirming that model isn't still rate-limited.
|
||||
|
||||
## Confirmed working: a live 10-turn playtest
|
||||
|
||||
Adventure id 43, played turn by turn against DeepSeek V4 Flash on production.
|
||||
Milo's Graveler and Onix both fainted; Kabutops came out third. Across ten
|
||||
turns:
|
||||
|
||||
- HP tracked correctly on both sides, including sandstorm chip damage each
|
||||
turn once `sandstorm_active` flipped on.
|
||||
- `status` stayed correctly independent per Pokémon (all `none` throughout
|
||||
this run — no status move was tried).
|
||||
- `player.active_pokemon` and `npc.milo.active_pokemon` both switched
|
||||
correctly as Pokémon were sent out or fainted (Pidgeotto → Wartortle →
|
||||
Machoke; Graveler → Onix → Kabutops).
|
||||
- `player.potions` decremented correctly on use (3 → 2) and the healed
|
||||
Pokémon's HP rose by the expected amount.
|
||||
- The World State sidebar reflected every one of these changes live, without
|
||||
a manual refresh, after the model's reply finished streaming.
|
||||
|
||||
This confirms the nested-`npcs` redesign from earlier in the session was the
|
||||
right fix for "why did you flatten then" — parallel entities with independent
|
||||
stats work as npcs, not as flattened player fields.
|
||||
|
||||
## Two bugs the playtest found
|
||||
|
||||
**1. The model sometimes skips the trailing `state` block entirely.**
|
||||
|
||||
> **Unconfirmed.** Truncation at `max_output_tokens` removes the block too, and it
|
||||
> was never ruled out here. See `plan/16`.
|
||||
On the very first turn of this run, DeepSeek V4 Flash narrated a Wartortle HP
|
||||
drop but never appended the ` ```state ` block the engine parses. The engine
|
||||
correctly left the state untouched — this is model non-compliance, not an
|
||||
engine bug — but the drop was silent: no error, no visible sign in the UI
|
||||
beyond "the numbers didn't move." A `Retry` on that same turn produced the
|
||||
delta correctly. Confirmed by reading the raw action row over
|
||||
`/api/adventures/{id}/actions` — `hasState` came back `false` on the first
|
||||
attempt and `true` on the retry. If this happens often during your own play,
|
||||
it is worth an authors'-note reminder or a stronger trailing-instruction
|
||||
nudge in `engine.py`'s prompt scaffolding, not a schema change.
|
||||
|
||||
**2. The model reliably forgets `npc.milo.pokemon_left` and milestones on a
|
||||
faint, despite an explicit instruction to update both.**
|
||||
|
||||
> **Wrong on both halves.** The model emitted `pokemon_left` at *both* faints; the
|
||||
> engine clamped it to nothing and reported it as applied. And it never emitted a
|
||||
> milestone because the milestone ids were absent from the prompt entirely. Neither
|
||||
> was an attention problem. See `plan/16`. `ai_instructions` in
|
||||
the scenario file already says: *"decrement `npc.milo.pokemon_left` when one
|
||||
of his faints"* and *"Mark milestones as they happen."* Across two separate
|
||||
faints in this session (Graveler, then Onix), the model correctly reset
|
||||
`active_pokemon`/`active_hp`/`active_status` for the incoming Pokémon every
|
||||
time, but never once touched `pokemon_left` (stuck at `3/3` through both
|
||||
faints) and never checked off "Knock out Milo's lead Graveler," even though
|
||||
that milestone was unambiguously satisfied. This looks like the faint
|
||||
instruction is buried inside a longer bulleted list the model is only
|
||||
partially attending to. Worth trying: pull the faint-handling instructions
|
||||
into their own short paragraph, or add a stat-guide line for `pokemon_left`
|
||||
and the milestones that makes them as visually prominent as `hp`/`status`.
|
||||
|
||||
**Separately, not necessarily a bug:** `world.turn` (a `counter` stat defined
|
||||
in the schema) stayed at `0` for all ten turns. `ai_instructions` never tells
|
||||
the model to increment it — the instructions cover HP, status, potions,
|
||||
active_pokemon, and pokemon_left, but not turn. If you want the counter to
|
||||
mean something, add an explicit line telling the model to bump
|
||||
`world.turn` by 1 every reply.
|
||||
|
||||
## Suggested next steps
|
||||
|
||||
1. Decide whether to patch `ai_instructions` for the two gaps above, then
|
||||
redeploy and play a few more turns to confirm faints correctly decrement
|
||||
`pokemon_left` and flip milestones.
|
||||
2. Clean up the stale "Road to the Champion" scenario and adventure 42.
|
||||
3. Try a status-condition move (Ivysaur's Poison Powder or similar) — this
|
||||
playtest never exercised the `status` stat changing away from `none`, so
|
||||
it is unverified in practice even though the schema supports it.
|
||||
4. If DeepSeek keeps skipping state blocks more than rarely, consider the
|
||||
trailing-reminder wording in `engine.py` (`build_state_reminder` or
|
||||
equivalent) rather than switching models — the schema itself is sound.
|
||||
@@ -1,343 +0,0 @@
|
||||
# World-state refusals: what the engine throws away, and who gets told
|
||||
|
||||
Read this before you play the Pokémon demo again. It records why two bugs in
|
||||
`plan/15-pokemon-demo-handover.md` were diagnosed wrongly, what the engine was
|
||||
actually doing, and what changed. Everything here is merged and green. It **has** now been driven in a browser;
|
||||
read "Driven in a browser" at the end first, because three of the five changes
|
||||
did not work and one bug explains all three.
|
||||
|
||||
**Last updated: 2026-08-28.**
|
||||
|
||||
---
|
||||
|
||||
## The one sentence version
|
||||
|
||||
`apply_delta` records three outcomes for every change the model sends —
|
||||
`applied`, `clamped`, `rejected` — and everything downstream read only
|
||||
`applied`. A refused change therefore reached the player as an ordinary chip and
|
||||
reached the model, on the next turn, as a change that had succeeded.
|
||||
|
||||
## What was actually wrong
|
||||
|
||||
`plan/15` recorded two bugs and named a cause for each. Both causes were wrong,
|
||||
and the investigation is worth keeping because the same reasoning trap is easy
|
||||
to repeat: **the visible evidence was "the number did not move", and the natural
|
||||
reading of that is that the model never tried.**
|
||||
|
||||
### `pokemon_left` was emitted every time
|
||||
|
||||
Adventure 43's action rows carry `milo pokemon_left` in `world_changes` at both
|
||||
faints, each with `"delta": 0, "value": 3`. The model saw the faint and wrote
|
||||
the path. It was not forgetting anything.
|
||||
|
||||
`old == new == 3` is reachable only from a **positive** value. `pokemon_left`
|
||||
was `min 0, max 3, initial 3, max_delta_per_turn 1`, so `+2` capped to `+1`,
|
||||
reached 4, and clamped back to the ceiling of 3. Net zero.
|
||||
|
||||
So the model sent the remaining count as an absolute — "two left" — instead of
|
||||
a delta of `-1`. Milo's four stats alternate between the two conventions:
|
||||
|
||||
| stat | convention |
|
||||
|---|---|
|
||||
| `active_pokemon` | text, absolute |
|
||||
| `active_hp` | number, delta |
|
||||
| `active_status` | text, absolute |
|
||||
| `pokemon_left` | number, delta |
|
||||
|
||||
HP survives because damage is naturally phrased as a change. A count is
|
||||
naturally phrased as a state, so it got text semantics.
|
||||
|
||||
**The general rule this produces:** a numeric stat whose `initial` equals the
|
||||
boundary it moves away from turns every wrong-signed change into a silent
|
||||
no-op. Every `hp` in the scenario has that shape (`initial == max`). It has
|
||||
never fired only because damage is phrased as a decrease by luck of language.
|
||||
|
||||
### Milestones were never emitted at all
|
||||
|
||||
Zero milestone changes across nine AI turns. `EMIT_RULE` asks for
|
||||
`"milestones.<id>": true` and `apply_delta` matches `<id>` against the schema
|
||||
key, but `render_state_section` printed only the description, and
|
||||
`render_reference` skipped the milestones section entirely
|
||||
(`STAT_SECTIONS = ("world", "player")`). The string `graveler_defeated` was
|
||||
nowhere in the prompt.
|
||||
|
||||
The same playtest is its own control: `sandstorm_active` is a flag, flags
|
||||
*are* printed by name, and it worked.
|
||||
|
||||
### The replay was teaching the model to repeat itself
|
||||
|
||||
Found while deciding whether to feed refusals forward, and the most damaging of
|
||||
the three. `_history_text` re-attached each past turn's state block from
|
||||
`world_delta["delta"]` — **what the model sent**, not what was applied. So the
|
||||
turn after the faint contained the model's own block claiming
|
||||
`"npc.milo.pokemon_left": 2`, directly above a live values line reading
|
||||
`pokemon_left 3/3`, with nothing to say which was true.
|
||||
|
||||
That is a per-turn lesson that sending `2` is correct. The identical mistake at
|
||||
the second faint is what that lesson predicts.
|
||||
|
||||
## What changed
|
||||
|
||||
Five changes, on `fix-silent-clamps-and-milestone-ids`.
|
||||
|
||||
1. **`Action.world_changes` reports refusals** (`models.py`). Reads `clamped`
|
||||
and `rejected` beside `applied`. Accepted stats carry `clamped`; refusals
|
||||
become `kind: "rejected"` entries. The `fix` key is present only when the
|
||||
engine wrote one — it is empty for every accepted change, and this property
|
||||
runs for every action of every list response.
|
||||
2. **The UI distinguishes three outcomes** (`Play.jsx`, `index.css`). Clamped to
|
||||
a standstill reads `no change — at its limit` on a dashed chip; a partial
|
||||
clamp is marked `(limited)`; a rejection carries its reason. Dashed and
|
||||
dimmed rather than red: the rules refusing a change is them working.
|
||||
3. **Milestones are named to the model** (`engine.py`). The goals line is now
|
||||
`Goals (mark with milestones.<id>): graveler_defeated — Knock out Milo's lead
|
||||
Graveler; …`, the same treatment NPCs get with `(npc.milo)`.
|
||||
4. **Refusals carry a generated correction** (`engine.py`). Each rejection, and
|
||||
each clamp that moved nothing, builds a `fix` string from the stat definition
|
||||
at the point of refusal, so it quotes the real limits and lists the real
|
||||
names. `render_refusals()` renders them into the prompt directly above
|
||||
`EMIT_REMINDER`, for the previous AI turn only.
|
||||
5. **History replays what was accepted** (`builder.py`). `applied_delta()`
|
||||
rebuilds the block from `report["applied"]`, dropping any numeric entry where
|
||||
`new == old` so a change that moved nothing cannot be copied as a zero.
|
||||
|
||||
Plus the demo scenario: `pokemon_left` became `pokemon_fainted`
|
||||
(`type: counter`, `initial: 0`), which puts a wrong sign on the counter rule
|
||||
where it is rejected out loud instead of absorbed. The faint instruction moved
|
||||
into its own paragraph, and a `world.turn` line was added — it sat at 0 for the
|
||||
whole playtest because nothing ever told the model to move it.
|
||||
|
||||
### The design call worth not re-litigating
|
||||
|
||||
**A clamp that reduced a change but still moved the value says nothing.** Only
|
||||
total losses are reported. If you tell a model its 80 damage became 30, it can
|
||||
treat the shortfall as a debt and send the remaining 50 next turn — which is the
|
||||
swing `max_delta_per_turn` exists to prevent. A rejection has no partial credit
|
||||
to chase. `test_a_clamp_that_still_moved_the_value_says_nothing` pins this.
|
||||
|
||||
## How to test it
|
||||
|
||||
532 backend tests pass and the frontend builds. **The UI work has no automated
|
||||
cover** — this project has no frontend test runner, which is the standing
|
||||
reason its UI bugs are found by hand.
|
||||
|
||||
Run the backend from `backend/` with
|
||||
`.venv/Scripts/python.exe -m pytest tests/`. The new file is
|
||||
`tests/test_change_visibility.py` (23 tests). Each of the three mechanisms fails
|
||||
its own test when disabled; that was checked by sabotage, not assumed.
|
||||
|
||||
To drive it, re-seed the scenario and play the demo:
|
||||
|
||||
1. **A refusal chip.** Open the World State drawer, use its ✎ edit mode to put
|
||||
Milo's `active_hp` at full, then play a turn where he takes no damage but the
|
||||
model tries to heal him. Easier and more reliable: send a deliberately wrong
|
||||
block by editing an AI turn. What you are looking for is a dashed chip
|
||||
reading `no change — at its limit`, not a `+0`.
|
||||
2. **The milestone.** Knock out Graveler. `graveler_defeated` should tick, and
|
||||
a `✓ graveler defeated` chip should appear. This is the single clearest
|
||||
pass/fail in the whole change — it never once happened before.
|
||||
3. **The faint counter.** At the same faint, `pokemon_fainted` should go 0 → 1.
|
||||
If the model sends an absolute again, it is now refused rather than absorbed,
|
||||
and the refusal note should appear in the *next* turn's prompt. Read it under
|
||||
Insights → the turn's context snapshot, section `world_state_refusals`.
|
||||
4. **The replay.** In the same snapshot, check the replayed history: a past
|
||||
turn's `state` block should carry only the changes that were accepted.
|
||||
5. **`world.turn`** should now advance by 1 per reply.
|
||||
|
||||
Check the narrow layout too. The chips grew longer text, and `.chg` is inside
|
||||
the story column with nothing to scroll sideways — `overflow-wrap: anywhere` is
|
||||
doing the work, and it was not re-checked at 390 px. Chrome clamps its minimum
|
||||
window width to ~500 px, so relaunch with `--window-size=` rather than trying to
|
||||
resize a maximized window.
|
||||
|
||||
## Still open
|
||||
|
||||
- **The missing `state` block from `plan/15` bug 1 is unexplained.** Truncation
|
||||
at `max_output_tokens` removes the block, which `LENGTH_HEADROOM` exists to
|
||||
prevent, and it was never ruled out. The distinguishing evidence is whether
|
||||
the narration ends mid-sentence with `finish_reason: length`.
|
||||
- **Every `hp` stat still has the `initial == max` shape.** Now visible when it
|
||||
bites, rather than silent, but not designed out.
|
||||
- ~~**The stale "Road to the Champion" scenario and adventure 42** are still on
|
||||
production. Deleting them is hand-work and was deliberately not automated.~~
|
||||
Automated on 2026-08-28: `seed.py` now deletes seeded scenarios no seed file
|
||||
claims. Adventure 42 survives with a NULL `scenario_id`, and losing the cover
|
||||
art is the whole cost.
|
||||
- **The Bandit Camp demo (`04-rpg-world-state.json`) was not checked** for the
|
||||
same milestone problem. Its milestones were equally unnamed to the model
|
||||
before this change, so it is worth asking whether one has ever fired there.
|
||||
|
||||
---
|
||||
|
||||
## Driven in a browser, 2026-08-28
|
||||
|
||||
Adventure 45, four turns, on production against the demo model. Two of the five
|
||||
changes worked. Three did not reach anyone, for one reason.
|
||||
|
||||
### The bug: the stored column dropped two of the three lists
|
||||
|
||||
`world_delta_of` in `routers/adventures.py` wrote `delta` and `applied` only.
|
||||
Every consumer that distinguishes outcomes reads the other two:
|
||||
|
||||
- `Action.world_changes` builds `clamped_paths` from `world_delta["clamped"]`,
|
||||
so every chip carried `clamped: false`. `Play.jsx`'s `blocked = c.clamped &&
|
||||
d === 0` could never be true, and `(limited)` could never render.
|
||||
- With no `rejected` list, a `kind: "rejected"` chip was unreachable.
|
||||
- `worldstate.refusals` reads both lists, so `render_refusals` always returned
|
||||
an empty string and no correction ever reached the next prompt.
|
||||
|
||||
The `fix` string survived only because the engine stores it inside the `applied`
|
||||
entry. Turn 3 is the record: the snapshot's `report.clamped` held
|
||||
`npc.ivysaur.hp` and `npc.milo.active_hp` with correct `fix` text, both chips
|
||||
came back `clamped: false`, turn 4's prompt contained no correction, and the
|
||||
model repeated the same mistake.
|
||||
|
||||
`world_delta_of` now carries all three lists.
|
||||
|
||||
**Why the 21 tests passed.** The `action()` helper in
|
||||
`test_change_visibility.py` built the column by hand as `{"delta": delta,
|
||||
**report}`, with every list present. The write path was never exercised. The
|
||||
helper now fills the column through `world_delta_of`. Removing the two lines
|
||||
again fails 9 of the 23 tests; that was checked, not assumed.
|
||||
|
||||
### The two that were not code faults
|
||||
|
||||
`graveler_defeated` never fired and `world.turn` never moved. Both instructions
|
||||
are present and correct in the assembled prompt: the goals line reads `Goals
|
||||
(mark with milestones.<id>): graveler_defeated — …`, and the scenario says to
|
||||
add 1 to `world.turn` every reply. The demo model ignores both. Change 3
|
||||
landed; the model is the limit.
|
||||
|
||||
### The absolutes were coming from the scenario's own wording
|
||||
|
||||
The model sent a total, not a change, for almost every number:
|
||||
`npc.ivysaur.hp: 96`, `npc.milo.active_hp: 65` then `88`, `player.potions: 2`.
|
||||
Because every `hp` has `initial == max`, each one clamped back to where it
|
||||
started. `potions` went **up** when the player spent one.
|
||||
|
||||
`EMIT_RULE` does say "CHANGES ONLY, as deltas (not new totals)", and
|
||||
`EMIT_REMINDER` repeats "deltas only". The scenario contradicted both at closer
|
||||
range. `milo.active_hp.desc` said "**Reset this to** the newcomer's full HP",
|
||||
which is an instruction to send an absolute, and that desc is injected every
|
||||
turn. The bullets said "**Drop** the HP … and **raise** it", naming a direction
|
||||
but never a sign. The only line that said "(not a delta)" was
|
||||
`player.active_pokemon`, so naming the exception made the rule look optional.
|
||||
|
||||
The one stat whose desc used delta wording, `pokemon_fainted` ("Add 1 each
|
||||
time"), is the one that worked.
|
||||
|
||||
Fixed in `05-league-championship.json`: every `hp` desc and the potion desc now
|
||||
state the sign, the bullets do too, and the lead-in says plainly that every
|
||||
number is a change with a worked example. `milo.active_hp.max_delta_per_turn`
|
||||
went 65 → 98, because a switch legitimately moves that stat a full bar and the
|
||||
old cap made the reset unreachable in one turn.
|
||||
|
||||
**Not fixed:** the `initial == max` shape itself. It is now loud rather than
|
||||
silent, and the wording removes the usual cause, but the shape is still there.
|
||||
|
||||
---
|
||||
|
||||
## Driven against Claude, 2026-08-28
|
||||
|
||||
The four turns above ran on the free demo model, which ignored two instructions
|
||||
that were present and correct in the prompt. To separate model behavior from
|
||||
code, the same demo was played again through `backend/tools/claude_shim.py`, an
|
||||
OpenAI-compatible endpoint backed by the local `claude` command line tool. The
|
||||
model was `sonnet`. See the README for how to run it.
|
||||
|
||||
Every one of the five changes worked:
|
||||
|
||||
| Turn | Chips |
|
||||
|---|---|
|
||||
| 1 | `turn +1`, `active pokemon → Wartortle`, `wartortle hp -8`, `milo active hp -46`, dashed `type advantage used refused — no such flag` |
|
||||
| 2 | `turn +1`, `milo active hp -52`, `milo pokemon fainted +1`, `milo active pokemon → Onix`, `✓ graveler defeated` |
|
||||
| 3 | `turn +1 (limited)`, dashed `milo active hp no change — at its limit` |
|
||||
| 4 | `turn +1`, `milo active hp +32`, `✓ type advantage used` |
|
||||
|
||||
`graveler_defeated` fired for the first time in five playtests, `world.turn`
|
||||
moved every turn, `pokemon_fainted` went 0 to 1 at the faint, and every HP number
|
||||
was a signed change rather than a total. The two failures left open above were
|
||||
the demo model, not the instructions.
|
||||
|
||||
**The refusal loop is verified end to end.** Turn 3 clamped to nothing. Turn 4's
|
||||
assembled prompt, read from `GET /api/adventures/1/actions/9/context`, carried
|
||||
the correction verbatim:
|
||||
|
||||
```
|
||||
[Part of your last state block was not applied. Correct it in this turn's block:
|
||||
- `npc.milo.active_hp` did not move. It is already at its minimum of 0 (it runs
|
||||
from 0 to 98; it moves at most 98 per turn).]
|
||||
```
|
||||
|
||||
The model's next delta was `{"world.turn": 1, "npc.milo.active_hp": 32,
|
||||
"milestones.type_advantage_used": true}`. That is the loop the whole change
|
||||
exists for, and it had never been observed running.
|
||||
|
||||
### The one bug that is still a schema fault
|
||||
|
||||
At the faint the model set `npc.milo.active_pokemon` to `Onix` and sent no
|
||||
positive `active_hp` change, so Onix arrived on the field at 0 of 98 and every
|
||||
later hit was refused. Sonnet followed the rest of the scenario closely, so this
|
||||
is the schema rather than the model: `active_hp` is one stat with one `max` of
|
||||
98, shared by Graveler at 98, Onix at 90, and Kabutops at 88. A switch has to
|
||||
raise it, and nothing in the schema can raise it on the referee's side.
|
||||
|
||||
The fix is to give Milo's three Pokemon their own NPC entries, which also removes
|
||||
the `initial == max` shape from this demo. Not done: the session's remaining time
|
||||
went to the guest starter instead.
|
||||
|
||||
### Cost
|
||||
|
||||
About $0.04 a turn, billed against the Claude subscription rather than a card.
|
||||
Roughly 13k of each request's 20k prompt tokens is the command line tool's own
|
||||
overhead; the app's prompt is about 7k.
|
||||
|
||||
---
|
||||
|
||||
## What shipped alongside it, 2026-08-28
|
||||
|
||||
**A permanent local test rig.** `backend/tools/claude_shim.py` is now in the
|
||||
repo. It finds the `claude` binary through `AIDND_CLAUDE_BIN`, then `PATH`, then
|
||||
the per-user install, and takes `--port` and `--claude`.
|
||||
|
||||
**The demo says Pokemon in its title, and has a Pokeball for cover art.**
|
||||
`05-league-championship.json` is now `[Demo] Pokemon League Championship: Round
|
||||
One`. Renaming a seed exposed a fault: `seed.py` matches a seed to its row by
|
||||
title, so a rename inserted a second public scenario and stranded the first,
|
||||
which is the same failure this file records for "Road to the Champion". Seeds
|
||||
now carry `previous_titles`, and `_find_renamed` lands the rename on the
|
||||
existing row. The scenario kept its id, and no orphan appeared.
|
||||
|
||||
The cover art is a 3.2 kB PNG data URI. An SVG one was tried first and stored as
|
||||
an empty `image_url`: `app/images.py` accepts raster formats only, on purpose,
|
||||
because SVG can carry script and these bytes are served from the app's own
|
||||
origin. `backend/tools/make_pokeball.py` draws the ball with `zlib` alone, so
|
||||
regenerating it needs no image library.
|
||||
|
||||
**Seeded scenarios no seed file claims are deleted.** `previous_titles` stops a
|
||||
rename stranding a row, but the rows already stranded still needed deleting by
|
||||
hand on every deployment. `_sweep_unclaimed` removes them on the next boot. Only
|
||||
a NULL owner with `is_public` is reachable, so nothing a player made can be
|
||||
touched, and `adventures.scenario_id` is `ON DELETE SET NULL`, so an adventure
|
||||
started from a deleted demo keeps its story and loses only the artwork. The
|
||||
sweep is skipped when a seed file fails to parse, and an empty seed directory
|
||||
never reaches it: neither reads as an instruction to delete live content.
|
||||
|
||||
**Every new guest gets the played adventure.** `app/starter.py` copies a shipped
|
||||
export bundle into each new guest account, from the guest mint in
|
||||
`routers/auth.py`. The bundle is this session's adventure trimmed to its first
|
||||
two exchanges, which ends on the knockout and shows an applied change, a refused
|
||||
one, and a milestone. It stops before the Onix bug above, and the state it
|
||||
leaves has Onix at its own 90 HP so a guest can play on from it.
|
||||
|
||||
An adventure has no cover art of its own and inherits its scenario's, and a
|
||||
bundle carries no scenario id, because an id is local to one database. The
|
||||
starter file names its source under `scenarioTitle` instead, and
|
||||
`starter._link_scenario` looks it up, so the copy shows the Pokeball rather than
|
||||
a monogram tile.
|
||||
|
||||
The row building that `POST /adventures/import` did inline moved into
|
||||
`bundle.materialize`, which both callers now use. The limit and rate checks
|
||||
stayed in the endpoint: the starter writes a file the server ships, so it has no
|
||||
untrusted list to cap. `tests/test_starter_adventure.py` covers the file, the
|
||||
copy, the chips, the playable end state, and the two failure paths.
|
||||
@@ -1,582 +0,0 @@
|
||||
# Phase 17: refactor for readability
|
||||
|
||||
The code works and it is tested. It is also hard to read, because four files hold
|
||||
most of it, one schema migration is half finished, and the published guide is a
|
||||
hand-maintained copy of another file. This phase fixes those three things without
|
||||
changing what the app does.
|
||||
|
||||
## How to use this file
|
||||
|
||||
This file is both the plan and the running status. Update the progress table at
|
||||
the end of every stage. Read the table first when you pick the work back up.
|
||||
|
||||
Each stage lands as its own commit on `refactor-17`. No stage starts until the
|
||||
stage before it is green.
|
||||
|
||||
## Progress
|
||||
|
||||
| Stage | Work | Status | Landed |
|
||||
|---|---|---|---|
|
||||
| 0 | Hygiene: worktrees, branches, undocumented settings | done, except the branch deletion | 2026-08-29 |
|
||||
| 1 | Split the four largest files | done: one test setup, and all four files split | 2026-08-29 |
|
||||
| 2 | Remove duplication | done: all seven items | 2026-08-29 |
|
||||
| 3 | SP8: drop the legacy columns | code done, all eight columns gone; the deploy and `VACUUM FULL` remain | 2026-08-29 |
|
||||
| 4 | Documentation | not started | |
|
||||
| 5 | A frontend test runner | not started | |
|
||||
|
||||
## Baseline
|
||||
|
||||
Record these numbers before you start. They are how you tell a refactor from a
|
||||
rewrite.
|
||||
|
||||
- 549 backend tests pass in 103 seconds.
|
||||
- `backend/app` holds 13,486 lines across 45 files.
|
||||
- `backend/tests` holds 12,221 lines across 38 files.
|
||||
- `frontend/src` holds 9,171 lines across 21 files.
|
||||
- The app exposes 70 endpoints and applies 64 migrations on boot.
|
||||
|
||||
Run the suite with `cd backend && .venv/Scripts/python.exe -m pytest tests/ -q`.
|
||||
|
||||
## What the review found
|
||||
|
||||
### The code is clean at the statement level
|
||||
|
||||
A scan for unreferenced functions across `backend/app` returned only FastAPI route
|
||||
handlers, which the decorator references rather than the name. A scan of
|
||||
`index.css` found 5 unused class names out of 373, and every one of the 5 is either
|
||||
CodeMirror's or built from a template string. There is no dead code to delete and
|
||||
no `TODO` or `FIXME` anywhere in the tree.
|
||||
|
||||
Read that as a constraint. The gains in this phase come from moving code, not from
|
||||
finding rot.
|
||||
|
||||
### Four files hold most of the complexity
|
||||
|
||||
| File | Lines | What it holds |
|
||||
|---|---|---|
|
||||
| `backend/app/routers/adventures.py` | 2353 | Nine concerns: the turn engine, branches, takes, import and export, adventure scripts, refresh from scenario, insights, memory CRUD, and action CRUD |
|
||||
| `frontend/src/pages/Play.jsx` | 2280 | 27 components, plus a 722-line `Play()` holding 18 `useState` calls and 7 `useEffect` calls |
|
||||
| `frontend/src/index.css` | 2865 | One stylesheet. The `max-width: 720px` block at line 2655 overrides rules written 2000 lines above it |
|
||||
| `backend/app/worldstate/engine.py` | 918 | Four jobs: parse a delta, apply a delta, render the context sections, and instantiate a schema |
|
||||
|
||||
### The schema is half migrated
|
||||
|
||||
`plan/14-phase-story-tree.md` defines SP8, which drops the columns the tree
|
||||
replaced. SP8 was gated on the tree running in production. It now does. Eight
|
||||
columns are still written on every turn and read by nothing:
|
||||
|
||||
- `actions.index`, `variants`, `variant_index`, `variant_count`
|
||||
- `actions.state_before`, `world_state_before`
|
||||
- `adventures.memory_cursor`, `summary_cursor`
|
||||
|
||||
Until they go, `models.py` carries two vocabularies for one idea, and every reader
|
||||
has to be told which one is live.
|
||||
|
||||
### The published guide is a hand-written copy
|
||||
|
||||
`docs/guide.html` is 84 KB of hand-written HTML covering the same material as the
|
||||
63 KB `docs/GUIDE.md`. Nothing generates one from the other. They have already
|
||||
drifted: the HTML last changed on 2026-08-20 and the Markdown on 2026-08-28, and
|
||||
section 3.6, "Counting visits", exists only in the Markdown. GitHub Pages publishes
|
||||
the HTML, so readers get the stale copy. `docs/architecture.html` is 87 KB of
|
||||
hand-written HTML with no Markdown source at all.
|
||||
|
||||
### Duplication the tests already protect
|
||||
|
||||
- `worldstate.apply_delta` and `apply_override` each re-implement the same path
|
||||
routing for `flags.`, `milestones.`, and `npc.<id>.`. That is six parallel
|
||||
branches at `engine.py:573`, `587`, `617`, `675`, `697`, and `740`.
|
||||
- 35 of the 38 test files repeat the same 8-line temporary-database prologue. There
|
||||
is no `conftest.py`.
|
||||
- `ScriptedProvider` or `FakeProvider` is defined 12 times, once per test file that
|
||||
needs a fake model.
|
||||
- The debounced autosave helper is copied into `Play.jsx:204`,
|
||||
`ScenarioEditor.jsx:23`, and `ScriptEditor.jsx:22`.
|
||||
- `get_adventure_or_404` is called by hand in about 20 handlers. It is a function,
|
||||
not a dependency, so every handler repeats the call.
|
||||
- `routers/chat.py:22` imports `SSE_HEADERS` and `sse` from `routers/adventures.py`.
|
||||
One router reaches into another for shared plumbing.
|
||||
|
||||
### One setting is documented nowhere
|
||||
|
||||
`AIDND_TRUSTED_PROXY_HOPS` is read at `limits.py:55`. It does not appear in
|
||||
`backend/.env.example`, `README.md`, or `render.yaml`. It sets how many proxy hops
|
||||
`client_ip()` trusts, which is what stops a rotated `X-Forwarded-For` header from
|
||||
buying a fresh rate-limit bucket. If a deployment adds a proxy hop and nobody sets
|
||||
this variable, the bypass comes back silently.
|
||||
|
||||
The unprefixed `DATABASE_URL` that `database.py:24` also accepts turned out to be
|
||||
documented already, inside the `AIDND_DATABASE_URL` entry in
|
||||
`backend/.env.example`. No change needed there.
|
||||
|
||||
### `Settings.stream` is dead
|
||||
|
||||
`models.py:622` defines the column, and `schemas.py:452` and `499` expose it.
|
||||
Nothing reads it in the backend or the frontend. This is item S1 in
|
||||
`docs/self-review.md`.
|
||||
|
||||
### Local clutter
|
||||
|
||||
Everything here is gitignored, so it costs repository weight nothing. It costs disk
|
||||
and attention.
|
||||
|
||||
- Two abandoned worktrees, each a full copy of the tree:
|
||||
`.claude/worktrees/simplify-comments` and `.claude/worktrees/sp7-tree-ui`. Both
|
||||
branches landed on `main` as squashes.
|
||||
- 21 local branches. 13 read as unmerged to `git branch --no-merged`, and
|
||||
`plan/STATUS.md` already records that they landed as squashes.
|
||||
- Three stale SQLite files in `backend/`: `data.backup-2026-07-06.db`,
|
||||
`data.backup-pre-phase8.db`, and `scroll_fixture.db`, 2.9 MB together. The live
|
||||
local database, `backend/data.db`, is not one of them and is not touched.
|
||||
|
||||
## Stage 0: hygiene
|
||||
|
||||
No source changes. Do this first, so later stages run against a quiet tree.
|
||||
|
||||
1. Remove both worktrees with `git worktree remove`, then delete the branches that
|
||||
landed as squashes. Verify each one first: a branch is safe to delete when
|
||||
`git log main..<branch>` names only commits whose message already appears in
|
||||
`git log main`.
|
||||
2. Delete the three stale `.db` files. Leave `backend/data.db` alone.
|
||||
3. Add `AIDND_TRUSTED_PROXY_HOPS` to `backend/.env.example`, with the reason it
|
||||
exists. Add it to the deploy section of `README.md` too, because that is where
|
||||
somebody putting the app behind another proxy looks.
|
||||
|
||||
**Check:** the suite still passes, and `git status` is clean.
|
||||
|
||||
### What Stage 0 actually did, 2026-08-29
|
||||
|
||||
Both worktrees are gone, which freed about 104 MB. Removing `sp7-tree-ui` needed one
|
||||
extra step: a Vite dev server had been running out of that worktree since
|
||||
2026-08-18, holding `frontend/.vite` open and owning port 5173. It was serving a tree
|
||||
54 commits behind `main`. Stopping it freed the directory. Restart the real one with
|
||||
`start.ps1`.
|
||||
|
||||
The three stale `.db` files are deleted. `backend/data.db` is untouched.
|
||||
|
||||
`AIDND_TRUSTED_PROXY_HOPS` is now in `backend/.env.example` and in the README deploy
|
||||
section.
|
||||
|
||||
**Still owed.** The 19 squash-landed branches are still there. `git branch -D` is
|
||||
blocked by the permission classifier, which is a reasonable guard on a destructive
|
||||
command. All 19 are verified safe by the rule above. Run this to clear them:
|
||||
|
||||
```
|
||||
git branch -D bugfix-code-review docs-story-tree fix-prepend-autoscroll \
|
||||
fix-silent-clamps-and-milestone-ids fix-store-refusals-and-delta-wording \
|
||||
handover-2026-08-17 measure-post-vacuum-sizes phase-14-story-tree \
|
||||
phase-7-public-repo phase-8-accounts sp1-tree-schema sp10-memory-bank-eviction \
|
||||
sp2-branch-clause sp3-node-cursors sp4-sibling-nodes sp5-fork-on-continue \
|
||||
sp6-bundle-v2 sp7-tree-ui sp7b-take-pager worktree-keep-fixture-and-status \
|
||||
worktree-simplify-comments
|
||||
```
|
||||
|
||||
`worktree-keep-fixture-and-status` is the one that needed checking by hand. Its
|
||||
commit subject appears nowhere in `main`, but the `--keep` flag it adds is on `main`
|
||||
at `tools/stress_session.py:764`, along with both gotchas the message describes, at
|
||||
`:686` and `:692`. Only the message differs.
|
||||
|
||||
The remote branches are left alone. Deleting those is a separate decision.
|
||||
|
||||
## Stage 1: split the four largest files
|
||||
|
||||
Every change in this stage moves code. None of it changes behavior. The 549 tests
|
||||
are the check, and they must pass without being edited, except where this section
|
||||
says otherwise.
|
||||
|
||||
### `routers/adventures.py` becomes a package
|
||||
|
||||
Split it into `backend/app/routers/adventures/`:
|
||||
|
||||
| Module | Holds | Source lines |
|
||||
|---|---|---|
|
||||
| `__init__.py` | The `APIRouter`, the shared dependencies, and re-exports | |
|
||||
| `paging.py` | `ACTION_LIST_COLUMNS`, `ACTION_PAGE`, `action_window`, `annotate_takes`, `current_window` | 27-146, 1203-1261 |
|
||||
| `crud.py` | List, create, get, patch, delete, plus the script-state and world-state readers | 225-580 |
|
||||
| `turns.py` | The turn engine: `generate_turn`, `_generate_turn`, `run_player_turn`, the turn lock, and the SSE helpers | 581-1004 |
|
||||
| `takes.py` | Retry, variants, takes, forking, and undo | 1005-1192, 1503-1778 |
|
||||
| `branches.py` | The branch endpoints | 1193-1502 |
|
||||
| `bundle_io.py` | Export and import | 1779-1837 |
|
||||
| `scripts.py` | The per-adventure script endpoints | 1838-1941 |
|
||||
| `refresh.py` | Refresh from scenario | 1942-2152 |
|
||||
| `insights.py` | The context dry run and the per-action context | 2153-2189 |
|
||||
| `memories.py` | Memory bank CRUD | 2190-2279 |
|
||||
|
||||
Target no file above 450 lines.
|
||||
|
||||
**The tests couple to the module, so read this before you start.** Twelve test
|
||||
files call `monkeypatch.setattr(adventures, "OpenAICompatibleProvider", ...)`.
|
||||
`monkeypatch` replaces a name in the module where the calling code looks it up, so
|
||||
re-exporting from `__init__.py` does not keep those patches working. Once
|
||||
`_generate_turn` lives in `turns.py`, the target becomes
|
||||
`adventures.turns.OpenAICompatibleProvider`.
|
||||
|
||||
Do Stage 2's `conftest.py` work first if you want that to be a one-line change
|
||||
instead of twelve. Otherwise retarget all twelve here. The other patched names are
|
||||
`adventures.limits`, `adventures.check_demo_cap`, `adventures.generate_turn`, and
|
||||
`adventures._active_turns`, each used once.
|
||||
|
||||
Keep these importable from the package root, because tests import them by name:
|
||||
`ACTION_PAGE`, `SNIPPET_MAX`, `_snippet`, `world_delta_of`, `acquire_turn_lock`,
|
||||
`retry_action`, `undo_turn`, and `_active_turns`.
|
||||
|
||||
`_active_turns` is module-level mutable state guarded by a lock. It must live in
|
||||
exactly one module, `turns.py`, and every other module must import the module and
|
||||
reach through it. If two modules import the set by value, the lock guards two
|
||||
different sets and the turn lock stops working.
|
||||
|
||||
### What the router split actually did, 2026-08-29
|
||||
|
||||
`backend/app/routers/adventures.py` is now a package of 14 modules. The largest
|
||||
is `turns.py` at 443 lines. Four modules exist that the table above does not
|
||||
list, because the plan's eleven still mixed unrelated work:
|
||||
|
||||
| Extra module | Why it exists |
|
||||
|---|---|
|
||||
| `deps.py` | Holds the `APIRouter` and the ownership check. It imports nothing else in the package, so every endpoint module can import the router without importing its siblings. |
|
||||
| `scenario_text.py` | Copying a scenario's text and cards has two callers, `crud.create_adventure` and `refresh`. Leaving it in either one made the other import an endpoint module. |
|
||||
| `nodes.py` | Story-tree navigation that four modules use: `last_action`, `next_index`, `next_depth`, `stand_on`, `db_tip`, `delete_turn`. |
|
||||
| `actions.py` | The three action endpoints. They page and delete rather than play a turn, so they do not belong in `crud.py`. |
|
||||
|
||||
Two decisions differ from the plan above.
|
||||
|
||||
**The package root does not re-export `acquire_turn_lock` or `_active_turns`.**
|
||||
The plan said to keep them importable, but that makes a broken patch look like a
|
||||
working one. Rebinding `adventures.generate_turn` changes the alias and leaves
|
||||
every caller reading the original, and the test still passes. Leaving those names
|
||||
off the package root raises `AttributeError` instead. Eighteen test call sites and
|
||||
four in `backend/tools/` now say `adventures.turns.<name>`. The package root still
|
||||
re-exports the pure helpers, so `adventures.ACTION_PAGE`, `adventures.undo_turn`,
|
||||
and `chat.py`'s `from .adventures import SSE_HEADERS, sse` are unchanged.
|
||||
|
||||
**`world_delta_of` lives in `turns.py`.** It reads as a world-state helper and sat
|
||||
beside the world-state endpoints, but `_generate_turn` is its only caller.
|
||||
|
||||
`_active_turns` behaved as the plan warned. Every module reaches it as
|
||||
`turns._active_turns`, and a check confirms the four modules see one set object
|
||||
and one lock.
|
||||
|
||||
One test coupled to the router for an unrelated module: it called
|
||||
`adventures.worldstate.instantiate`. It imports `app.worldstate` directly now.
|
||||
|
||||
The split moved text rather than retyping it. An AST comparison against the
|
||||
pre-split file confirms all 86 definitions are identical, once the `turns.`
|
||||
prefix is normalized away. The 549 tests pass, and the OpenAPI schema still lists
|
||||
the same 35 operations.
|
||||
|
||||
### `worldstate/engine.py` becomes a package
|
||||
|
||||
Split `backend/app/worldstate/engine.py` into four modules under
|
||||
`backend/app/worldstate/`:
|
||||
|
||||
- `schema.py`: `has_schema`, `instantiate`, `reconcile`, `band_label`, `npc_name`,
|
||||
`npc_triggers`, `_initials`.
|
||||
- `parse.py`: `extract_delta`, `_tolerant_load`, `render_delta_block`,
|
||||
`applied_delta`, `refusals`, `render_refusals`.
|
||||
- `apply.py`: `apply_delta`, `apply_override`, and the shared path resolver Stage 2
|
||||
introduces.
|
||||
- `render.py`: `render_state_section`, `render_reference`, `_stat_line`,
|
||||
`_describe_stat`.
|
||||
|
||||
`worldstate/__init__.py` re-exports every public name it exports today, so no call
|
||||
site changes. The engine is imported as a module, not monkeypatched, so this split
|
||||
carries none of the router's test coupling.
|
||||
|
||||
### `pages/Play.jsx` becomes a directory
|
||||
|
||||
Split it into `frontend/src/pages/Play/`:
|
||||
|
||||
- `index.jsx`: the page component.
|
||||
- `usePlaySession.js`: the 18 `useState` calls and 7 `useEffect` calls that drive
|
||||
one adventure, behind one hook.
|
||||
- `panels/`: `PlotPanel`, `MemoryPanel`, `ScriptsPanel`, `InsightsPanel`,
|
||||
`BranchPanel`.
|
||||
- `drawers/`: `StatusDrawer`, `WorldStateDrawer`, and the `StatRow`, `StatGroup`,
|
||||
`StateTree`, and `StateValue` parts they use.
|
||||
- `reports/`: `StateChangeChips`, `WorldStateReport`, `ScriptReport`,
|
||||
`CacheReport`, `TokenBreakdown`.
|
||||
- `TakePager.jsx`, `RefreshModal.jsx`.
|
||||
|
||||
There is no frontend test runner until Stage 5, so this split is verified by
|
||||
`npm run lint`, `npm run build`, and by driving the Play screen in a browser. Drive
|
||||
it. `plan/STATUS.md` records that the last two Play bugs were both found by hand and
|
||||
were unreachable from any test that existed.
|
||||
|
||||
### `index.css` becomes a directory
|
||||
|
||||
Split it into `frontend/src/styles/`, imported in order from `index.css`:
|
||||
`tokens.css`, `base.css`, `nav.css`, `cards.css`, `play.css`, `drawers.css`,
|
||||
`schema-editor.css`, `insights.css`, `modals.css`, and `theme.css`.
|
||||
|
||||
Move each `@media` block next to the rules it overrides, rather than leaving one
|
||||
`max-width: 720px` block at the end. Keep the source order identical when you move
|
||||
rules, because CSS resolves ties by order and the file relies on that in at least
|
||||
two known places, recorded in `plan/STATUS.md` and in the comments at
|
||||
`index.css:1011` and `1082`.
|
||||
|
||||
**Check:** 549 tests pass, `npm run lint` and `npm run build` are clean, and the
|
||||
Play screen works in a browser at desktop and at 500 px wide.
|
||||
|
||||
### What Stage 1 actually did to the frontend, 2026-08-29
|
||||
|
||||
`Play.jsx` is now `frontend/src/pages/Play/`, twelve files. The largest is
|
||||
`index.jsx` at 781 lines. `index.css` is now an `@import` list over
|
||||
`frontend/src/styles/`, eighteen files.
|
||||
|
||||
Two decisions differ from the plan above.
|
||||
|
||||
**`usePlaySession.js` does not exist yet.** The page component still owns all of
|
||||
the session state. Moving eighteen `useState` calls and seven `useEffect` calls
|
||||
is a rewrite, not a move, and no frontend test would catch a mistake in it
|
||||
today. It waits for Stage 5.
|
||||
|
||||
**The `@media` blocks stayed where they were.** `responsive.css` still holds one
|
||||
`max-width: 720px` block, at the end of the import order. Moving a `@media` block
|
||||
next to the rules it overrides moves it earlier in the cascade, which changes
|
||||
which of two equal-specificity rules wins. Nothing in the test suite would catch
|
||||
that. Do this after Stage 5.
|
||||
|
||||
Each split is verified by a different proof, because neither one has a test:
|
||||
|
||||
- CSS: the parts rebuild `index.css` byte for byte, and the built bundle is
|
||||
identical before and after at 56686 bytes.
|
||||
- JSX: every non-blank line of the original appears exactly once, in order, across
|
||||
the twelve files. A name-resolution check confirms every identifier each file
|
||||
references is defined or imported there, with no unused imports.
|
||||
|
||||
The line split stranded a comment at six of the boundaries. A leading comment
|
||||
sits above the section it describes, so each boundary cut one loose and left it
|
||||
at the end of the file before it. All six moved to the section they describe.
|
||||
|
||||
`npm run lint` and `npm run build` are clean, and 549 tests pass. Driving the
|
||||
Play screen covered the story view, all five panels, both drawers including the
|
||||
world-state edit form, the branch map, the refresh dialog, and the take pager,
|
||||
which stepped onto a take that lives on another branch and switched to it. The
|
||||
console reported no errors. The extension cannot resize the render viewport and
|
||||
the app sends `X-Frame-Options: DENY`, so the narrow-width check ran by setting
|
||||
the `max-width` media queries to `all` in the live stylesheet. All 71 narrow
|
||||
rules found their elements: the nav collapses to one button, the panel tabs move
|
||||
onto the title row, a panel fills the screen, the composer stacks, and both
|
||||
drawers become edge tabs.
|
||||
|
||||
## Stage 2: remove duplication
|
||||
|
||||
Each item here is small and is covered by tests that already exist.
|
||||
|
||||
1. **One path resolver in `worldstate`.** `apply_delta` and `apply_override` route
|
||||
`flags.<name>`, `milestones.<id>`, `npc.<id>.<stat>`, `world.<stat>`, and
|
||||
`player.<stat>` with parallel code. Extract a resolver that returns the
|
||||
container, the stat definition, and the kind, then let the two functions differ
|
||||
only in the write rule. Their rules genuinely differ, so do not merge the
|
||||
functions themselves. `apply_override` ignores `cooldown`,
|
||||
`max_delta_per_turn`, and the rule that a counter cannot decrease, and it lets a
|
||||
milestone be un-set. `tests/test_worldstate.py` covers both.
|
||||
2. **Add `backend/tests/conftest.py`.** Move the temporary-database prologue there,
|
||||
so the other 35 files drop 8 lines each. The prologue has to run before
|
||||
`from app.main import app`, because `main.py` calls `bootstrap(engine)` at import
|
||||
time. A `conftest.py` runs before any test module, which satisfies that.
|
||||
3. **One fake provider.** Move `ScriptedProvider` and `FakeProvider` to `conftest.py`
|
||||
as one fixture, replacing 12 copies.
|
||||
4. **Move the SSE helpers out of the router.** Put `sse`, `SSE_HEADERS`, and
|
||||
`turn_error` in `backend/app/sse.py`, so `chat.py` stops importing from
|
||||
`adventures`.
|
||||
5. **Make `get_adventure_or_404` a dependency.** About 20 handlers repeat the call.
|
||||
A `Depends` removes the line from each one and puts the ownership check in the
|
||||
signature, where a reader looking for it expects it.
|
||||
6. **One `useDebouncedSave` hook** in `frontend/src/hooks/`, replacing the three
|
||||
copies.
|
||||
7. **Delete `Settings.stream`.** Remove the column, the two schema fields, and add
|
||||
the migration that drops it. This is item S1 in `docs/self-review.md`. Mark it
|
||||
applied there.
|
||||
|
||||
**Check:** 549 tests pass. Test count may drop if consolidating fixtures removes a
|
||||
duplicate case. If it does, say which case and why in the commit message.
|
||||
|
||||
### What Stage 2 actually did, 2026-08-29
|
||||
|
||||
All seven items landed. Items 2 and 3 went in early, with the test setup in
|
||||
`32cd7c1`, because the conftest had to exist before the router split could move
|
||||
any test. Items 1, 4, and 5 are `2c57b1c`. Items 6 and 7 are `e0bf2b6`.
|
||||
|
||||
**Item 1, the path resolver.** `_resolve` returns a `_Target` naming the kind,
|
||||
the section, the key, the stat definition, and the character, or a rejection.
|
||||
The two callers now differ only in the write rule, which is what the item asked
|
||||
for. Building the container is a method on `_Target` rather than part of
|
||||
resolving, because a rejected path must not leave an empty section behind.
|
||||
|
||||
A refactor here is hard to check by reading, so it was checked by running. A
|
||||
differential harness fed 3960 generated payloads through the old and the new
|
||||
implementation and compared the state and the report. Ignoring `fix`, there are
|
||||
zero differences. 674 rejections gained a `fix` string and none lost one, all of
|
||||
them in `apply_override`, which had been the terser of the two. That is an
|
||||
improvement rather than a regression: `apply_delta` already worded those
|
||||
strings, and the world-state editor renders them.
|
||||
|
||||
**Item 5, the ownership dependency.** All 32 handlers converted, not the 20 the
|
||||
review estimated. Two proofs, because tests alone would not catch a change in
|
||||
the order of the checks or in the public HTTP surface:
|
||||
|
||||
- An AST pass confirmed the `get_adventure_or_404` call was the first statement
|
||||
in every one of the 32 handlers. If it were not, hoisting it into a dependency
|
||||
would move work that used to run after something else.
|
||||
- The generated OpenAPI document was diffed against one built from a `git clone`
|
||||
at `HEAD`. The only difference is that `rename_branch` now lists `branch_id`
|
||||
before `adventure_id`, which is parameter order in the spec and not a route
|
||||
change.
|
||||
|
||||
Seven tests in `test_state_revert.py` needed updating. They call handlers as
|
||||
plain functions rather than over HTTP, so they have to pass `adventure=` now.
|
||||
That file is the only one that does this.
|
||||
|
||||
**Item 7 needed a migration guard.** Migration 65 drops `settings.stream`, and
|
||||
it is the first migration that drops a column. `_column_already_there` already
|
||||
existed for the `ADD COLUMN` case. Dropping needs the mirror, `_column_already_gone`,
|
||||
because `create_all` builds the current schema, which is already missing the
|
||||
column, and fixtures like `test_tree_migration.pre_tree` stamp an old version
|
||||
against a database built that way and replay. Verified by running migration 65
|
||||
twice, once against a database that still had the column and once against one
|
||||
that did not.
|
||||
|
||||
**One duplicate was left alone.** `FakeProvider` in `test_chat.py` is not a copy
|
||||
of `ScriptedProvider`. It records the key, model, and endpoint it was
|
||||
constructed with, which is how the chat tests assert on what would have gone
|
||||
over the wire. `fakes.ScriptedProvider` streams replies and records prompts.
|
||||
Merging them would give one class two unrelated jobs.
|
||||
|
||||
## Stage 3: SP8, drop the legacy columns
|
||||
|
||||
This is the only stage that touches the production database. Follow
|
||||
`plan/14-phase-story-tree.md`, which specifies it.
|
||||
|
||||
1. Confirm nothing reads the eight columns. `ACTION_LIST_COLUMNS` in the adventures
|
||||
package lists `index`, `variant_count`, and `variant_index` today, so that tuple
|
||||
changes here. `bundle.py:37` and `context/history.py:424` both describe `index`
|
||||
as unread; verify that rather than trusting the comment.
|
||||
2. Add the migration that drops them. Remove the columns from `models.py`, and
|
||||
remove the fields from `schemas.py` and from `ActionOut`.
|
||||
3. Deploy, then run `VACUUM FULL actions;` on the direct Neon endpoint, not the
|
||||
`-pooler` one. Dropping a column rewrites toasted values, and nothing reclaims
|
||||
that space on its own. `plan/STATUS.md` records what happened the last time
|
||||
nobody ran it: the database reached 144.2 MB against a 512 MB tier.
|
||||
4. `VACUUM FULL` takes an `ACCESS EXCLUSIVE` lock, so the app blocks on `actions`
|
||||
for the duration. It took 5.5 seconds at 144 MB.
|
||||
|
||||
**Check:** 549 tests pass, `/api/health` answers `{"ok":true}` after the deploy, and
|
||||
one existing adventure opens, takes a turn, retries it, and pages back through the
|
||||
transcript.
|
||||
|
||||
### What Stage 3 actually did, 2026-08-29
|
||||
|
||||
Steps 1 and 2 landed. Step 3, the deploy and the `VACUUM FULL`, is still open,
|
||||
because it runs against production and belongs with the release, not with the
|
||||
branch.
|
||||
|
||||
**Step 1, the read audit.** Nothing outside the migrations reads the eight
|
||||
columns. `TakePager.jsx` reads `action.take_count` and `action.take_index` only,
|
||||
which is what makes dropping the three payload fields safe. The comments in
|
||||
`bundle.py` and `context/history.py` were accurate.
|
||||
|
||||
**Step 2, the drop.** Migrations 66 to 73 drop one column each. `index` is a
|
||||
keyword in SQLite, so migration 71 quotes it. `models.py`, `schemas.py`, and
|
||||
`ACTION_LIST_COLUMNS` lost the same eight, `Adventure.actions` now orders by
|
||||
`id`, and `attempts.renumber`, `context/history.max_action_index`, and
|
||||
`nodes.next_index` are deleted.
|
||||
|
||||
Two problems came out of the migration passes rather than the DDL:
|
||||
|
||||
- `_split_variants_into_siblings` wrote through `Base.metadata.tables["actions"]`,
|
||||
the live ORM table, so it stopped compiling the moment migration 66 removed
|
||||
five of its columns. It now writes through `_ACTIONS_AT_60`, a frozen `Table`
|
||||
carrying its own `MetaData`. That declaration is a snapshot of a past schema
|
||||
and must not be updated to track `models.py`.
|
||||
- Five passes read columns that migrations 66 to 73 drop. A `create_all`
|
||||
database replays every migration against the current schema, so each pass now
|
||||
calls `_has_columns` and returns early when the columns are absent. This is
|
||||
the rule `_column_already_there` applies to DDL, applied to the passes.
|
||||
|
||||
`bootstrap` gained a `through` argument. A migration test that asserts on
|
||||
something a later migration removes stops at the version it is about, rather
|
||||
than reading a schema several versions newer.
|
||||
|
||||
**One pre-existing bug found and not fixed.** `bundle._write_nodes` never sets
|
||||
`Action.parent_id`, and `paging.annotate_takes` groups on `parent_id`, so an
|
||||
imported adventure's take pager reads 1 of 1. The frontend has read only
|
||||
`take_count` since SP9, so this predates Stage 3 and is not a regression from
|
||||
it. `test_retry_variants.py` documents it.
|
||||
|
||||
**Check:** 555 backend tests pass, up from 549. `test_tree_migration.py` gained
|
||||
eight parametrized cases asserting each column is gone after a real schema-45
|
||||
database migrates all the way. A `create_all` database would pass those without
|
||||
running the migration, which is why the fixture is a frozen pre-tree one.
|
||||
|
||||
## Stage 4: documentation
|
||||
|
||||
1. **Generate `docs/guide.html` from `docs/GUIDE.md`.** Write a small build script
|
||||
that renders the Markdown into the existing hand-written HTML shell, keeping the
|
||||
current styles and metadata. After that, one edit updates both. Note in
|
||||
`README.md` that the HTML is generated and that you edit the Markdown.
|
||||
2. **Decide what `docs/architecture.html` is.** It has no Markdown source. Either
|
||||
give it one and generate it the same way, or state at the top of the file that
|
||||
it is hand-written, so the next person does not look for a source that does not
|
||||
exist.
|
||||
3. **Split `plan/STATUS.md`.** It is 63 KB, and most of it is dated session logs.
|
||||
Keep "Pick up here", "Things worth remembering", and "Running things" in
|
||||
`STATUS.md`. Move the dated entries to `plan/history/`, newest file first.
|
||||
4. **Add `docs/DEVELOPING.md`.** Record how to run the app, how to run the tests,
|
||||
where each subsystem lives, and the invariants a newcomer breaks first:
|
||||
- Anything that reads `adventure.actions` during generation takes
|
||||
`exclude_action_id`. It leaks into four places, not one.
|
||||
- `memory_cursor` and `summary_cursor` are positions into `story_actions()`, and
|
||||
`Memory.source_start` and `source_end` are `Action.index` values. The two
|
||||
spaces diverge as soon as anything is deleted.
|
||||
- `/auth/me` must not raise. It is the SPA's bootstrap call, so anything it
|
||||
touches that can raise takes the whole frontend down.
|
||||
- The global field rule in the stylesheet keys off `input[type=...]`, so a bare
|
||||
`<input>` and any new input type get browser default styling until you add
|
||||
them.
|
||||
- A modal opened from `.side-panel` needs `createPortal`, because the panel's
|
||||
filling transform animation makes it the containing block for
|
||||
`position: fixed`.
|
||||
|
||||
**Check:** the generated HTML matches the Markdown section for section, including
|
||||
section 3.6. Every link in `README.md` and `docs/index.html` resolves.
|
||||
|
||||
## Stage 5: a frontend test runner
|
||||
|
||||
`plan/STATUS.md` names the missing runner as the reason this project keeps finding
|
||||
UI bugs by hand. Two shipped bugs were unreachable from any test that existed.
|
||||
|
||||
1. Add Vitest, React Testing Library, and jsdom. Add a `test` script to
|
||||
`frontend/package.json`.
|
||||
2. Add the step to the existing `frontend` job in `.github/workflows/ci.yml`,
|
||||
between lint and build.
|
||||
3. Write the first tests against the two bug classes that already recurred:
|
||||
- The take pager renders when a retry's reply arrives over SSE. The stream builds
|
||||
its own `ActionOut`, so it was the one payload that never carried the
|
||||
annotation.
|
||||
- Retaking a player turn does not produce `> You > You`. The editor is seeded
|
||||
with stored text that is already formatted.
|
||||
4. Add a test for `usePlaySession` from Stage 1, since that hook now holds the state
|
||||
the page used to hold inline.
|
||||
|
||||
**Check:** `npm test` passes locally and in CI.
|
||||
|
||||
## Risks
|
||||
|
||||
| Risk | Where | How you catch it |
|
||||
|---|---|---|
|
||||
| A monkeypatch silently stops patching, and a test passes while calling a real provider | Stage 1, `routers/adventures` split | After retargeting, break the fake on purpose and confirm the tests that use it fail |
|
||||
| `_active_turns` ends up imported by value into two modules, so the turn lock guards two sets | Stage 1 | `tests/test_state_revert.py` reaches for `adventures._active_turns`. Keep that import path working, and confirm a second concurrent turn still returns 409 |
|
||||
| A CSS rule changes meaning because the split reorders it | Stage 1, stylesheet split | Compare the built CSS before and after. Order within each section must not change |
|
||||
| The column drop rewrites the table and nobody reclaims the space | Stage 3 | Run `VACUUM FULL actions;` on the direct endpoint, and read sizes from `sum(octet_length(col))` rather than `n_live_tup` |
|
||||
| Splitting `Play.jsx` breaks something no test covers | Stage 1 | Drive the Play screen by hand at both widths. Stage 5 exists to shrink this risk for next time |
|
||||
|
||||
## What this phase does not do
|
||||
|
||||
- It does not replace the hand-rolled migration runner with Alembic. 64 migrations
|
||||
run correctly, and `docs/GUIDE.md` section 2.6 already explains the choice.
|
||||
- It does not move `bootstrap(engine)` out of import time in `main.py`. The current
|
||||
behavior means a failed migration means no service, which is deliberate. Stage 2's
|
||||
`conftest.py` removes the boilerplate that import-time bootstrapping forces on
|
||||
tests, which is the part that actually hurts.
|
||||
- It does not change any API shape, any database content, or any prompt.
|
||||
@@ -1,145 +0,0 @@
|
||||
# Appendix: the memory-prompt A/B, run 2 — a fresh story
|
||||
|
||||
An independent replication of `plan/18-appendix-memory-ab-run.md`, produced by
|
||||
`backend/tools/memory_ab.py` through `tools/claude_shim.py` after
|
||||
`MEMORY_MAX_WORDS` was added. The story is **newly generated**, not the one run 1
|
||||
used, so this tests the prompts against different prose rather than re-scoring
|
||||
the same text.
|
||||
|
||||
## What replicated, and what did not
|
||||
|
||||
| | run 1 before | run 2 before | run 1 after | run 2 after |
|
||||
|---|---|---|---|---|
|
||||
| memory 1 | 34 w, "you" | 89 w, "the player" | 72 w, named | 54 w, named |
|
||||
| memory 2 | 105 w, "the player" | 107 w, "the player" | 68 w, named | 55 w, named |
|
||||
|
||||
**Naming replicated cleanly.** Four control memories across two runs, and not one
|
||||
of them names the protagonist. Four treatment memories, and all four do. That is
|
||||
the change working, and it is the finding to rely on.
|
||||
|
||||
**Length replicated, and the ceiling holds.** The control ran 34, 89, 105, 107
|
||||
words — a four-fold spread with no stated budget. The treatment after
|
||||
`MEMORY_MAX_WORDS` ran 54 and 55.
|
||||
|
||||
**The person-drift did not recur, and the claim about it should be read
|
||||
narrowly.** Run 1 produced two control memories in two different persons — "You
|
||||
crept low" and "The player asked Gwen" — which is exactly the reported
|
||||
complaint. Run 2's controls were both "the player", consistently. So drifting
|
||||
between *second* and *third* person is a real thing a model does, observed once,
|
||||
not something it does every time. What is consistent across both runs is that
|
||||
the control never reaches for the character's name, because it has never been
|
||||
told one.
|
||||
|
||||
---
|
||||
|
||||
*Everything below is the raw run, unedited.*
|
||||
|
||||
One story, generated through `build_context` against `sonnet` at `http://127.0.0.1:8787/v1`. Both prompts then summarize the same blocks, so the prompt is the only variable. The control is `MEMORY_SYSTEM_PROMPT` as of `9cdcb55`.
|
||||
|
||||
| memory | arm | words | framing |
|
||||
|---|---|---|---|
|
||||
| 1 | before | 89 | "the player" |
|
||||
| 1 | after | 54 | named |
|
||||
| 2 | before | 107 | "the player" |
|
||||
| 2 | after | 55 | named |
|
||||
|
||||
## Memory 1
|
||||
|
||||
**Before:** The player and Gwen scouted a bandit camp at dawn, agreeing on a quiet approach with Gwen circling the right flank (targeting the spear-carrying guard) while the player advanced from the left; the strongbox sat hidden under an oiled tarp in the camp's center. The player searched an empty bedroll undetected, taking a belt knife, a handful of copper and silver coins, and a charcoal-marked parchment map of the camp, then began moving low toward the strongbox while the guard remained oblivious and Gwen stayed ready with her bow.
|
||||
|
||||
**After:** Kaelen and Gwen infiltrated the bandit camp at dawn, splitting flanks with Gwen covering right from the tree line, bow ready, while Kaelen approached left. Kaelen looted a bedroll near an overturned cart, taking a belt knife, coins, and a charcoal map of the camp, then began moving low toward the tarp-covered strongbox undetected.
|
||||
|
||||
## Memory 2
|
||||
|
||||
**Before:** The player crept through the bandit camp toward a tarp-covered iron strongbox, reflexively reached to pull Gwen down beside him despite her being forty feet away at the tree line, and the resulting half-second lapse nearly caused a second bandit (a wood-gatherer) to spot him at the woodpile; Gwen almost loosed an arrow but held on his signal, and the bandit left unaware. The player then waited for the spearman to turn toward the fire and moved undetected to crouch beside the strongbox, where Gwen signaled approval and pointed to the simple hasp-and-padlock lock before he asked her whether she still trusted his read on the situation.
|
||||
|
||||
**After:** Kaelen infiltrated the bandit camp toward the tarp-covered strongbox while Gwen covered him from the tree line; he instinctively reached to pull her down beside him but she was forty feet away, nearly exposing him to an unnoticed second bandit gathering wood, which frayed Gwen's trust before he safely reached the strongbox's simple hasp-and-padlock lock.
|
||||
|
||||
## The story both arms summarized
|
||||
|
||||
**start:** Dawn breaks grey over the treeline as you and Gwen crouch at the edge of the bandit camp. Smoke curls from a dying fire; three bedrolls lie empty. Somewhere ahead, the stolen caravan strongbox waits.
|
||||
|
||||
Gwen checks her bowstring and looks to you. "Quiet, or loud?"
|
||||
|
||||
**do:** go quiet and signal Gwen to circle around the right flank
|
||||
|
||||
**ai:** You lean close, voice barely a breath. "Quiet. Take the right flank, stay in the tree line — I'll come at them from the left when you're set. Don't loose until I move first, unless someone spots you."
|
||||
|
||||
Gwen's mouth quirks, half approval, half habit. "Try not to trip over anything this time." She melts sideways into the brush without a sound, her passage barely stirring a fern, and angles wide around the camp's edge, bow already loose in her hand.
|
||||
|
||||
You give her a slow count — enough time to get into position — watching the camp in the meantime. A lone bandit paces near the fire, spear resting on his shoulder, yawning like the night watch had been long and dull. Beyond him, tents sag in a loose ring, and past those, half-hidden under an oiled tarp, a squat iron-banded shape sits guarded by nothing more than habit and complacency. The strongbox.
|
||||
|
||||
You catch a flicker of movement at the tree line opposite — Gwen, settling into a crouch behind a fallen log, bow half-drawn, patient as stone. She lifts two fingers, then points toward the spear-carrier: hers if it goes loud.
|
||||
|
||||
You ease forward through the underbrush, testing each footfall before committing your weight, keeping the dying fire's smoke between you and the pacing guard's sightline. Your pulse is steady, senses sharpened — the cool mana still humming faintly under your skin from old habit, ready if you need it. No alarm yet. No eyes turned your way. The camp sleeps in its false security, and for now, the quiet is yours to keep or break.
|
||||
|
||||
You reach the shadow of an overturned cart at the camp's rim, close enough now to hear the guard mutter something to himself. Gwen is in position, arrow nocked, waiting on your signal.
|
||||
|
||||
**do:** search the nearest bedroll for anything useful
|
||||
|
||||
**ai:** You glance toward the nearest bedroll — one of the three empty ones, rumpled blankets still holding the shape of whoever slept there, maybe ten feet from the cart's shadow. The guard's back is turned, spear-shoulder rotating slowly as he paces his short loop. You judge the distance, the timing, and decide it's worth the risk.
|
||||
|
||||
You slip low across the open ground, boots finding bare dirt instead of dry leaves, and drop to a crouch beside the bedroll. Your fingers move quick and quiet through the folded blanket, a stiff leather pack shoved half-underneath. Coins clink faintly — you cup your palm over the pouch before it can rattle again — and your hand closes on a stubby belt knife, a handful of copper and silver, and a folded scrap of parchment that might be a rough map of the camp's layout, marks scratched in charcoal.
|
||||
|
||||
The guard's pacing falters. He half-turns, spear dipping, some animal instinct prickling at the edge of his awareness. You freeze mid-motion, breath held, willing yourself into the cart's shadow. Across the clearing, you catch the barest shift from Gwen — her draw hand tightening, ready to loose if this goes wrong.
|
||||
|
||||
The moment stretches. Then the guard grunts, spits into the dying fire, and resumes his slow circuit, oblivious. You let your breath go slow and silent, tucking the coin pouch and parchment into your jerkin, the knife into your belt beside your own.
|
||||
|
||||
It's not much, but it's something — coin for later, and maybe a map worth reading twice before you commit to a route through this camp. You ease back from the bedroll, retreating to the cart's cover, eyes flicking to Gwen. She gives the faintest nod: *still clear, still yours to call.*
|
||||
|
||||
The guard yawns again, spear drooping, utterly unaware how close he came to dying loud. The strongbox waits under its tarp, patient as ever. The quiet is still yours — for now.
|
||||
|
||||
**do:** move toward the strongbox, keeping low
|
||||
|
||||
**ai:** You keep to the cart's shadow a moment longer, judging the guard's rhythm — three steps out, a pause, three steps back, spear dipping low on the turn. When his back rotates fully away, you move.
|
||||
|
||||
Low and fast, you cross the gap between the cart and the nearest tent, one hand brushing the canvas to steady yourself rather than grip it — no rustle, no tug on the guy-lines. You thread between two more tents, close enough now to smell tallow and old woodsmoke, and drop into a crouch behind a stack of crated supplies not ten feet from the tarp-covered strongbox.
|
||||
|
||||
From here you can see it properly: iron bands dark with rust, a heavy lock through a hasp, the oiled tarp thrown over it more to keep off dew than to hide it. Careless. Whoever's running this camp trusts their numbers more than their diligence.
|
||||
|
||||
You risk a glance back toward the tree line. Gwen hasn't moved, but her head has tracked your progress the whole way, bow still half-drawn, patient. She catches your eye and gives the barest tilt of her chin toward the strongbox — *go on, I've got you* — before her gaze slides back to the pacing guard, who is still oblivious, still yawning his way through a boring watch.
|
||||
|
||||
Closer now, you can make out voices from one of the tents — low, unhurried, two men talking about nothing in particular. Not alarmed. Not aware. The camp's quiet holds, fragile as it is, and the strongbox sits within reach, close enough that you could touch the tarp if you stretched out a hand.
|
||||
|
||||
Your fingers itch to check the lock, to see if it needs force or finesse — but that's a decision for the next breath, not this one. For now you crouch in the crate-shadow, heartbeat steady, senses stretched wide across the camp: the pacing guard, the murmuring tent, Gwen's stillness at your back. Nothing has broken yet. The strongbox is yours to reach for.
|
||||
|
||||
**do:** grab her wrist and pull her down behind the woodpile
|
||||
|
||||
**ai:** Your hand moves on instinct before thought catches up — reaching sideways for Gwen's wrist to haul her down out of sight — but your fingers close on nothing but cool air. She isn't beside you. She's forty feet off across the clearing, still nested behind her log at the tree line, exactly where you left her.
|
||||
|
||||
The mistake costs you half a second, and half a second is enough. Boots crunch on the far side of the crates — close, closer than the pacing guard should be. A second bandit, one you hadn't accounted for, rounds the supply stack with a coil of rope over one shoulder, heading for the woodpile stacked just past your cover.
|
||||
|
||||
You drop flat behind the crates instead, breath locked in your chest, willing your travel-stained leathers to melt into the shadow. The bandit hums tunelessly, kneels at the woodpile — near enough you could count the notches in his boot heel — and starts loading split logs into his arms, utterly unaware he's sharing six feet of dirt with an exiled ranger and a very recent, very foolish reflex.
|
||||
|
||||
Across the camp, Gwen has gone rigid, bow lifting a fraction — she saw the second man appear, saw you freeze, and for one taut moment she's clearly bracing to put an arrow in his throat if he so much as glances your way. You catch her eye through the crate-gap and give the smallest shake of your head. *Not yet. Hold.*
|
||||
|
||||
The wood-gatherer straightens, arms full, and ambles off toward the fire without ever knowing how close he came. Your heart hammers against your ribs. Gwen eases her draw, but her jaw stays tight, and even at this distance you can feel the flicker of irritation under her patience — not at the bandit, at you, for a half-second of carelessness that nearly bought you both a fight you hadn't chosen.
|
||||
|
||||
The strongbox still waits behind its tarp, closer now than ever. But the quiet feels thinner than it did a moment ago.
|
||||
|
||||
**do:** wait for the bandit to turn, then move
|
||||
|
||||
**ai:** You settle back against the crates, still as the wood itself, and let your breathing fall into the same slow rhythm as the camp around you. The wood-gatherer's footsteps fade toward the fire pit, and a moment later you hear the clatter of logs dropped beside the flames, a muttered word to the pacing spearman, then silence again but for the crackle of embers catching new fuel.
|
||||
|
||||
You count another ten heartbeats before you risk a look. The spearman has stopped to warm his hands, back fully to you now, attention on the fire instead of his rounds. No one else stirs between you and the tarp.
|
||||
|
||||
This is the gap.
|
||||
|
||||
You rise from the crate-shadow in one smooth motion, low and quick, closing the last stretch of open ground between the crates and the strongbox in a handful of silent strides. The tarp is close enough to touch now — you drop to a crouch beside it, back pressed to the iron-banded corner, and finally let yourself breathe.
|
||||
|
||||
Across the clearing, Gwen's stance has eased a fraction. She tracked every step of that crossing, and though her jaw is still set from the wood-gatherer's near miss, there's the faintest approval in how she holds herself now — the tension of someone watching a mistake corrected rather than repeated. She lifts two fingers again, then points toward the lock hasp: *your move.*
|
||||
|
||||
The lock itself is heavy, old iron, a simple hasp-and-padlock rather than anything clever — the kind of thing built to deter opportunists, not a trained hand. You could pick it quiet, or put a boot to it loud. The voices still murmur from the nearby tent, unconcerned. The spearman still warms his hands at the fire. The quiet holds, thinner than before, but yours.
|
||||
|
||||
**do:** ask her whether she still trusts my read on this
|
||||
|
||||
**ai:** You crouch there a moment longer, one hand still braced on the strongbox's cold iron corner, and risk a glance across the clearing. Gwen's eyes are already on you — they usually are — and you mouth more than whisper the question, tilting your head just enough to ask it: *still with me on this?*
|
||||
|
||||
For a beat she doesn't answer, just watches the spearman warming his hands, the murmuring tent, the whole fragile balance of the camp's false calm. Then she gives a small, economical shrug, the kind that means *mostly*, and taps two fingers against her bow stave — a gesture you've learned means *your instincts, not your luck.* It isn't a full yes. The wood-gatherer's near miss still sits behind her eyes, and you can read it plainly enough: she'll follow your read, but she's watching it closer than she was an hour ago.
|
||||
|
||||
Her hand shifts, index finger drawing a small circle in the air, then pointing at the lock — *pick it, don't break it. Loud is still your last resort, not your first.* It's about as much reassurance as you're going to get from a ranger who's had one scare too many this morning, and honestly, it's fair. You nod once, slow, so she can see you take it.
|
||||
|
||||
The spearman shifts his weight at the fire but doesn't turn. The tent's murmur continues, unbroken. Whatever trust frayed at the woodpile hasn't snapped — it's just drawn tighter, watchful, waiting to see what you do with the lock in front of you. The strongbox sits solid and patient at your back, and beyond the thin canvas walls, the camp sleeps on in its dangerous, borrowed quiet.
|
||||
|
||||
You turn back to the hasp, letting your fingers find the mechanism, senses still split between the metal under your hands and the fire-lit shape of the spearman thirty feet off. The moment is yours to use well — or not.
|
||||
|
||||
@@ -1,134 +0,0 @@
|
||||
# Appendix: the memory-prompt A/B, run 2026-08-31
|
||||
|
||||
The evidence behind `plan/18-persona-and-memory-quality.md`. Reproduce it with:
|
||||
|
||||
python tools/claude_shim.py --port 8787 &
|
||||
python tools/memory_ab.py --out ab.md
|
||||
|
||||
One story, generated a turn at a time through the app's own `build_context`.
|
||||
Both prompts then summarize the **same** blocks, so the story is held constant
|
||||
and the prompt is the only variable. Every call is a separate request, so
|
||||
neither arm sees the other's output, and the model is never told what is being
|
||||
measured. The control is `MEMORY_SYSTEM_PROMPT` as of commit `9cdcb55`, read
|
||||
from git rather than pasted, so it cannot drift from what shipped.
|
||||
|
||||
The model was a Claude model, reached through `tools/claude_shim.py`. See
|
||||
"What this does not show" in plan/18 before generalising from it.
|
||||
|
||||
## What it showed
|
||||
|
||||
| memory | arm | words | how it names the protagonist |
|
||||
|---|---|---|---|
|
||||
| 1 | before | 34 | second person — "**You** crept low…" |
|
||||
| 1 | after | 72 | "**Kaelen** and Gwen crouched…" |
|
||||
| 2 | before | 105 | "**The player** asked Gwen to…" |
|
||||
| 2 | after | 68 | "**Kaelen** and Gwen infiltrated…" |
|
||||
|
||||
Two things came out of this, and only one of them was the thing being tested.
|
||||
|
||||
**The reported fault reproduced.** Two consecutive memories, from one story,
|
||||
written minutes apart, in two different persons. That is the complaint, observed
|
||||
rather than argued from the prompt text. Both after-memories name Kaelen.
|
||||
|
||||
**A fault nobody had noticed.** "1-2 plain sentences" is not a length: 34 words
|
||||
for one block, 105 for the next. `memory_top_k` injects five memories every
|
||||
turn, so the bank's running cost was set by a number never stated.
|
||||
`MEMORY_MAX_WORDS = 50` now states it. Re-run over this same story, the two
|
||||
blocks came back at **32 and 58 words**, still named, still third person, still
|
||||
carrying the camp map, the strongbox behind the second tent, and the strap
|
||||
frayed near through:
|
||||
|
||||
> Kaelen crept low toward the strongbox, now known from the map to be behind
|
||||
> the second tent, with the dozing watchman still unaware and Gwen holding her
|
||||
> covering position on the flank.
|
||||
|
||||
**A likely accuracy win, at n=1.** The player's move was `grab her wrist and
|
||||
pull her down behind the woodpile`. The before-memory recorded "The player asked
|
||||
Gwen to grab her wrist and pull her down", inverting who acted. This is the
|
||||
failure the brief predicts — with no cast there is nothing to say whose wrist
|
||||
"her wrist" is — but one sample is an observation, not a result.
|
||||
|
||||
---
|
||||
|
||||
*Everything below is the raw run, unedited.*
|
||||
|
||||
## Memory 1
|
||||
|
||||
**Before:** You crept low through the mist toward the second tent, following the map's mark toward the strongbox while Gwen held her flanking position, the sleeping watchman still dozing on his spear near the lean-to.
|
||||
|
||||
**After:** Kaelen and Gwen crouched at dawn outside the bandit camp; Kaelen signaled Gwen to circle the right flank while a drowsy watchman kept post near the lean-to, and Gwen took position unseen near the strongbox tent. Kaelen searched an abandoned bedroll, finding a whetstone, a few copper coins, dried meat, and a charcoal-scrawled camp map revealing the stolen strongbox was hidden behind the second tent, then began moving low toward that location.
|
||||
|
||||
## Memory 2
|
||||
|
||||
**Before:** The player asked Gwen to grab her wrist and pull her down behind the woodpile just before a second bandit—half-dressed, carrying a piss-bucket—stumbled out, relieved himself, and returned to bed in the first tent; Gwen lingered a beat before freeing her wrist and murmured "Good eyes." After waiting five heartbeats to confirm the bandit was asleep, the player and Gwen returned to the strongbox tent, where they found the old iron caravan lock too sturdy to pick and instead began testing the straps, discovering one frayed near through at the fold; the player then asked Gwen whether she still trusted their read on the situation.
|
||||
|
||||
**After:** Kaelen and Gwen infiltrated the bandit camp to reclaim a stolen strongbox, sheltering behind a woodpile while a second bandit relieved himself nearby before returning to sleep; afterward Kaelen returned to the strongbox tent, found its lock too sturdy to force by hand but noticed one strap frayed near through, and paused there with Gwen covering him to ask whether she still trusted his read on the situation.
|
||||
|
||||
## Summary
|
||||
|
||||
**Before:** Under cover of mist, the player and Gwen scouted a bandit camp, aiming to reach a strongbox marked on their map while a watchman dozed near a lean-to. Approaching the second tent, the player pulled Gwen down behind a woodpile just in time to avoid a half-dressed bandit who stepped out to relieve himself before returning to sleep in the first tent; Gwen lingered a moment before pulling free, murmuring "Good eyes." After waiting to confirm the camp was settled, the two returned to the strongbox tent, finding its old iron caravan lock too sturdy to pick. They shifted to testing the straps securing it instead, discovering one frayed nearly through at the fold—a possible way in. The player then turned to Gwen to ask whether she still trusted their read on the situation, leaving the moment of decision unresolved.
|
||||
|
||||
**After:** At dawn, Kaelen and Gwen scouted a bandit camp holding a stolen strongbox. Kaelen sent Gwen to circle the right flank near the strongbox tent while a drowsy watchman kept post, then searched an abandoned bedroll himself, finding a whetstone, a few copper coins, dried meat, and a charcoal map marking the strongbox's location behind the second tent. Moving low toward it, the pair sheltered behind a woodpile while a second bandit briefly woke and relieved himself before returning to sleep. Reaching the strongbox tent, Kaelen found its lock too sturdy to force by hand, but noticed one securing strap frayed nearly through. With Gwen covering him, he paused to ask whether she still trusted his read on the situation before attempting to exploit the weak strap.
|
||||
|
||||
## The story that was summarized
|
||||
|
||||
**start:** Dawn breaks grey over the treeline as you and Gwen crouch at the edge of the bandit camp. Smoke curls from a dying fire; three bedrolls lie empty. Somewhere ahead, the stolen caravan strongbox waits.
|
||||
|
||||
Gwen checks her bowstring and looks to you. "Quiet, or loud?"
|
||||
|
||||
**do:** go quiet and signal Gwen to circle around the right flank
|
||||
|
||||
**ai:** You melt back from the ridge line without a sound, boots finding moss instead of twigs, and catch Gwen's eye. Two fingers point right, then curl in a half-circle — *swing wide, come at the flank*. She reads it instantly, dips her chin, and peels off through the bracken, bow already half-raised, moving low and quiet along the tree line until the mist swallows her.
|
||||
|
||||
You hold position at the camp's edge, breath slow, watching the dying fire spit its last embers. A bandit shifts on watch near the lean-to, spear butt planted in the dirt, more bored than alert. No one else stirs. Somewhere off to the right you catch the faintest whisper of disturbed leaves — Gwen, further along than you'd expect, patient as ever. She's given herself a clean angle on the strongbox tent from the flank, out of the watchman's sightline entirely.
|
||||
|
||||
The plan is good. Quiet holds. All you need now is the first move — and the choice of who falls first, and how.
|
||||
|
||||
The watchman yawns, oblivious, close enough that a single silent step would put you in reach of his throat before he ever draws breath to shout.
|
||||
|
||||
**do:** search the nearest bedroll for anything useful
|
||||
|
||||
**ai:** You crouch by the nearest bedroll — a heap of stitched furs still holding the sleeper's warmth, though its owner is long since up and yawning by the fire. You keep one eye on the watchman as your fingers work through the folds. A whetstone, a handful of copper coins that clink softly against your palm before you still them with a fist, a scrap of oilcloth wrapped around dried meat gone hard at the edges. Nothing worth the risk so far.
|
||||
|
||||
Then your fingers catch on something stiffer, tucked beneath the bedroll's edge — a folded scrap of parchment, half-torn, charcoal-scrawled. A rough map of the camp, crude but legible: the lean-to marked with an X, a second mark deeper in camp labeled in shorthand you recognize as bandit cant for *strongbox*. It's tucked behind the second tent, not the first — further from the watchman than you'd guessed, closer to where Gwen is now waiting in the mist.
|
||||
|
||||
You fold it away inside your jerkin, pulse ticking a little faster. Small find, real value — it just turned guesswork into a plan. The watchman still hasn't stirred, spear planted, chin dropping toward his chest. Gwen's silhouette is a smudge of shadow at the flank, patient, waiting on your signal.
|
||||
|
||||
|
||||
|
||||
The camp holds still around you — watchman drowsing, fire dying, Gwen coiled and ready on the flank. The map sits warm against your ribs, the strongbox's true location now known. Your move.
|
||||
|
||||
**do:** move toward the strongbox, keeping low
|
||||
|
||||
**ai:** You drop low, weight rolling from heel to ball of each foot, threading between the empty bedroll and a stack of split firewood. The watchman's chin has sunk near to his chest, spear listing in his slack grip. You give him a wide berth anyway — no need to gamble breath on a man already halfway to sleep.
|
||||
|
||||
The second tent is smaller than the first, its flap laced shut with a crude leather cord. Just as the parchment promised, its bulk sits wrong for a tent meant to sleep anyone — square-cornered, iron-banded, too heavy for the ground it rests on. You crouch beside it, easing the lacing open one loop at a time, and the flap parts on a dull gleam of banded oak and iron: the strongbox, lock intact, straps still buckled from the caravan raid.
|
||||
|
||||
A shift of cloth behind you — not the watchman, too far behind you for that. You go still, hand dropping to your knife hilt, eyes cutting sideways through the tent's shadow toward the sound. Nothing resolves yet, just a settling, maybe wind through canvas, maybe not. Across the clearing, Gwen has gone motionless too, bow half-drawn, watching the same dark gap between tents that's caught your attention.
|
||||
|
||||
The strongbox is here, within reach, real weight under your palm instead of a rumor on a scrap of parchment. But something in that gap between the tents hasn't declared itself yet, and the watchman is still only a few careless steps away from waking to find you crouched over his camp's whole reason for being.
|
||||
|
||||
**do:** grab her wrist and pull her down behind the woodpile
|
||||
|
||||
**ai:** Your hand shoots out and closes around Gwen's wrist before she's fully turned toward the gap between the tents. She goes rigid for half a heartbeat — a killer's instinct, ready to break the grip — then reads your face and lets herself be pulled. You drag her down behind the woodpile together, shoulders pressed to split logs, breath shallow.
|
||||
|
||||
The shape that had been shifting in the shadow resolves itself: a second bandit, roused early, stumbling out from behind the tents with his breeches half-laced and a piss-bucket in hand, utterly uninterested in anything but the treeline. He doesn't look toward the strongbox tent. Doesn't look toward the woodpile. He shuffles past, empties the bucket against a stump, and turns back the way he came, already yawning.
|
||||
|
||||
Gwen's breath eases out slow beside you. She doesn't pull her wrist away immediately — a beat longer than she needs to — before easing it free herself. "Good eyes," she breathes, barely sound at all. Her gaze flicks to your hand, then back to your face, something unreadable and warm passing behind it. Watching your back instead of charging in — that's the kind of thing she remembers.
|
||||
|
||||
Through the gap in the logs you track the second bandit's path: he ducks into the first tent, not the watch line, and doesn't reemerge. Three accounted for now — the drowsing watchman, this one bedding back down, and whoever else the camp is hiding. The strongbox tent is still open behind you, lock waiting, straps waiting. Gwen settles back into her crouch, bow across her knees, eyes on you for the next signal.
|
||||
|
||||
**do:** wait for the bandit to turn, then move
|
||||
|
||||
**ai:** You go still as stone, eyes fixed on the tent flap the second bandit vanished into. Seconds crawl. Then the canvas stirs one last time — a grunt, a shifting weight settling onto a bedroll, the particular silence of a man who's decided the day can wait a while longer. You count five heartbeats to be sure, watching for any second shadow, any sign he's only feigning sleep. Nothing. Even the watchman by the fire hasn't stirred, spear still planted, chin still low.
|
||||
|
||||
You ease up from the woodpile, moving on the balls of your feet, and slip back to the strongbox tent. Gwen catches the motion and rises with you, silent as a drawn breath, bow angled to cover the gap between tents while you work. The lock is old iron, caravan-make, not bandit work — it'll need proper tools or proper force, not fingers.
|
||||
|
||||
You crouch and test the straps first, feeling for a weak buckle, a give point. One strap is frayed near through at the fold. With steady pressure it might part without the sharp crack of metal — quieter than forcing the lock, if slower. Gwen shifts her weight beside you, patient, trusting you to read the moment right, her eyes flicking once toward the dozing watchman and back.
|
||||
|
||||
The camp holds its breath around you both. The strongbox waits under your hands, and the next move is yours to make.
|
||||
|
||||
**do:** ask her whether she still trusts my read on this
|
||||
|
||||
**ai:**
|
||||
|
||||
@@ -1,655 +0,0 @@
|
||||
# The protagonist has no name, and the summarizer is told nothing
|
||||
|
||||
Two changes, in order. Phase 1 gives the adventure a persona. Phase 2 uses it,
|
||||
along with the cast, to fix the memories. Phase 1 is worth shipping on its own;
|
||||
Phase 2 depends on it and is much smaller once it lands.
|
||||
|
||||
**Both phases are built and green (631 backend tests). Phase 1 was driven in a
|
||||
browser (21/21 checks). Phase 2 was run end to end against a real model, as a
|
||||
controlled A/B on one story — see "Run with a real model". A bank written under
|
||||
the old prompt can be rewritten in place — see the last section.**
|
||||
|
||||
**Last updated: 2026-08-31.**
|
||||
|
||||
---
|
||||
|
||||
## The one sentence version
|
||||
|
||||
Memories come back inconsistent and vague because the summarizer is handed six
|
||||
actions of second-person prose and nothing else — no protagonist, no cast, no
|
||||
setting, no instruction about what person to write in — so it cannot say who
|
||||
"you" is or who "she" is, and neither can any memory it writes.
|
||||
|
||||
## The evidence
|
||||
|
||||
This is the entire prompt that writes a memory (`memorybank.py:391`):
|
||||
|
||||
```
|
||||
system: You compress interactive-fiction story excerpts into memories. Respond
|
||||
with 1-2 plain sentences in past tense stating the concrete facts and
|
||||
events (names, places, items, promises, injuries). No preamble, no
|
||||
commentary.
|
||||
|
||||
user: Story excerpt:
|
||||
|
||||
<6 actions, joined by blank lines, truncated to 2000 tokens>
|
||||
|
||||
Memory:
|
||||
```
|
||||
|
||||
Nothing else is passed. Not `adventure.memory` (the plot essentials), not the
|
||||
NPC names and descriptions the scenario already defines, not the world state,
|
||||
and not the protagonist — because until Phase 1 there is no protagonist to pass.
|
||||
|
||||
Two failures follow from that, and they are the two complaints:
|
||||
|
||||
**Inconsistent framing.** The prompt never says what person to write in. The
|
||||
model picks one per call. Across a single bank you get "You entered the crypt",
|
||||
"The player entered the crypt", and "He entered the crypt" describing the same
|
||||
kind of event.
|
||||
|
||||
**No idea who anyone is.** Given `You push the door open. She grabs your arm.`
|
||||
the only honest memory is *"You entered a room and she stopped you."* Retrieved
|
||||
forty turns later into a scene with three women in it, that memory is worse than
|
||||
nothing.
|
||||
|
||||
The summary inherits both problems, because `_update_story_summary` builds from
|
||||
the memory bullets.
|
||||
|
||||
---
|
||||
|
||||
# Phase 1 — the persona
|
||||
|
||||
## What already exists, and what does not
|
||||
|
||||
`stat_schema` already gives the player a stat block, and it is already
|
||||
namespaced beside the NPCs:
|
||||
|
||||
```
|
||||
world.day
|
||||
player.hp <- the player's stats, already sectioned
|
||||
npc.gwen.trust
|
||||
flags.alarm_raised
|
||||
milestones.escaped
|
||||
```
|
||||
|
||||
NPCs carry `name`, `keys`, and `desc` (`schema.py`, `render_reference`). The
|
||||
`player` section carries none of those. That asymmetry is the whole gap: the
|
||||
player has stats but no identity, so the block renders as `You: hp 100/100` and
|
||||
reads to the model as a floating global rather than a character.
|
||||
|
||||
## The paths do not change
|
||||
|
||||
`player.hp` stays `player.hp`. It does not become `kaelen.hp` or `user.hp`.
|
||||
|
||||
The reason is `_history_text()` in `context/builder.py`. It replays every past
|
||||
turn's stored delta back into the prompt, and those stored blobs contain literal
|
||||
strings like `{"player.hp": -15}`. A path that carries the persona's name breaks
|
||||
the moment a player renames their character: every replayed block in history
|
||||
then names a path the schema no longer defines, `applied_delta` renders it
|
||||
anyway, and the model copies the broken form. A rename would also require
|
||||
migrating every stored `world_state`, every seed scenario, and the `EMIT_RULE`
|
||||
example text, for no functional gain.
|
||||
|
||||
What changes is the **label**, using the trick NPCs already use — print the
|
||||
display name and the path together:
|
||||
|
||||
```
|
||||
before: You: hp 100/100, mana 30/50.
|
||||
after: Kaelen (player): hp 100/100, mana 30/50.
|
||||
```
|
||||
|
||||
The name is visible to the reader, the path stays copyable by the model.
|
||||
|
||||
## Where the persona lives
|
||||
|
||||
Three plain columns on `adventures`, alongside `memory` and `authors_note`:
|
||||
|
||||
```python
|
||||
persona_name: Mapped[str] = mapped_column(String(80), default="")
|
||||
persona_pronouns: Mapped[str] = mapped_column(String(40), default="")
|
||||
persona_desc: Mapped[str] = mapped_column(Text, default="")
|
||||
```
|
||||
|
||||
**Not in `stat_schema`.** Two reasons. It has to work for an adventure with no
|
||||
RPG layer at all, which is most of them, and `_initials()` in
|
||||
`worldstate/schema.py` treats every dict inside a section as a stat definition —
|
||||
a `persona` key dropped into `stat_schema.player` would be instantiated as a
|
||||
stat, rendered as a stat line in the guide, and handed an `initial` value.
|
||||
|
||||
**Adventure-level, not scenario-level.** Two people playing the same scenario
|
||||
are different characters. A scenario steers the protagonist through its plot
|
||||
essentials, which it already can. Keeping personas off `scenarios` also avoids
|
||||
having to decide what "Update from scenario" does to a persona the player has
|
||||
edited: the answer is nothing, because the scenario never had one.
|
||||
|
||||
**Not per-branch.** Branches are alternative futures within one adventure; the
|
||||
protagonist is the same person down all of them.
|
||||
|
||||
Empty `persona_name` means the feature is off and behavior is exactly what it is
|
||||
today. That is the entire backward-compatibility story — no backfill.
|
||||
|
||||
## Pronouns get their own field
|
||||
|
||||
One short string: `he/him`, `she/her`, `they/them`. It exists because Phase 2
|
||||
will tell the summarizer to write in third person. Without a stated pronoun the
|
||||
model infers one from the name, and once it infers wrong that error is baked
|
||||
into every memory it writes from then on and into the summary built from them.
|
||||
A field is cheaper than a wrong guess repeated forever.
|
||||
|
||||
Blank is a valid value. When it is blank nothing is rendered, and the Phase 2
|
||||
prompt tells the summarizer to use the name rather than a pronoun.
|
||||
|
||||
## Where it enters the prompt
|
||||
|
||||
A new static section in `build_context`, between `ai_instructions` and
|
||||
`plot_essentials`:
|
||||
|
||||
```
|
||||
Player character:
|
||||
You are Kaelen (he/him). A half-elf ranger, exiled from the northern holds.
|
||||
```
|
||||
|
||||
**It goes in the system block on purpose.** The description is user-only and
|
||||
never changes during a turn, so it sits inside the cached prefix and costs
|
||||
nothing after the first turn. This is why "only the user can change it" is not
|
||||
just a product choice — it is what keeps the section free. Anything the AI could
|
||||
rewrite would have to move below the history with the other live sections, and
|
||||
would re-price the prompt on every change.
|
||||
|
||||
It is emitted whether or not the adventure has an RPG layer. It is not gated
|
||||
behind `has_ws`.
|
||||
|
||||
## Where it enters the world state
|
||||
|
||||
`worldstate/render.py`, two edits:
|
||||
|
||||
- `render_state_section` takes the display name and prints
|
||||
`Kaelen (player): …` in place of the hardcoded `You: …`. Falls back to `You:`
|
||||
when no persona is set.
|
||||
- `render_reference` gains one line tying the name to the path, mirroring the
|
||||
NPC header it already writes:
|
||||
`- Protagonist Kaelen — stats are addressed as player.<stat>.`
|
||||
Without this the model reads "Kaelen" in the prose and invents `kaelen.hp`.
|
||||
|
||||
The persona description is **not** repeated in the stat guide. It already has
|
||||
its own section, and the guide is about paths.
|
||||
|
||||
## The AI cannot edit it
|
||||
|
||||
Nothing to build. `_resolve()` in `worldstate/apply.py` rejects any path it does
|
||||
not recognize, so a delta of `{"persona.name": "Bob"}` already falls through to
|
||||
"`persona.name` is not a tracked value", and that refusal already reaches the
|
||||
model through `render_refusals`. Worth one test to hold the behavior in place.
|
||||
|
||||
## Setting it, and changing it
|
||||
|
||||
**At the start.** `PlaceholderModal` (`components.jsx`) already exists and
|
||||
already opens before an adventure begins — but only when the scenario contains
|
||||
`${...}` tokens. It becomes a "Begin adventure" modal that always opens, with
|
||||
the three persona fields at the top and any placeholder fields below.
|
||||
|
||||
**Placeholders stay independent of the persona.** A scenario using `${Name}`
|
||||
will ask for a name twice. That is accepted: no scenario in the repo uses
|
||||
placeholders at all (grepped across `seed_data/` and `starter_data/` — zero
|
||||
occurrences), so the collision is hypothetical. If it ever shows up in practice,
|
||||
pre-filling the `Name` field from the persona is a few lines in the modal with
|
||||
no backend change.
|
||||
|
||||
**Later.** `AdventureUpdate` gains the three fields, and `PlotPanel` gets a
|
||||
Protagonist block at the top. The debounced-save wiring is already there.
|
||||
`WorldStateDrawer` shows the name as the character-sheet heading but does not
|
||||
edit it — one edit surface, not two.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Change |
|
||||
|---|---|
|
||||
| `models.py` | three columns on `Adventure` |
|
||||
| `migrations.py` | 74, 75, 76 |
|
||||
| `schemas.py` | `AdventureCreate`, `AdventureUpdate`, `AdventureOut` |
|
||||
| `context/builder.py` | `persona` section after `plot_essentials` |
|
||||
| `worldstate/render.py` | state-section label, reference line |
|
||||
| `worldstate/__init__.py` | export whatever `render.py` newly exposes |
|
||||
| `routers/adventures/crud.py` | store persona at creation |
|
||||
| `bundle.py` | export key + import read |
|
||||
| `frontend/components.jsx` | modal gains persona fields |
|
||||
| `frontend/pages/Scenarios.jsx`, `Home.jsx` | always open the modal |
|
||||
| `frontend/pages/Play/panels/PlotPanel.jsx` | Protagonist block |
|
||||
| `frontend/pages/Play/drawers/WorldStateDrawer.jsx` | heading |
|
||||
| `frontend/api.js` | pass the fields through |
|
||||
|
||||
## Tests
|
||||
|
||||
| File | What |
|
||||
|---|---|
|
||||
| new `test_persona.py` | create with persona; update; renders in the system block; works with no `stat_schema`; empty name is a no-op; AI delta at `persona.*` refused |
|
||||
| `test_worldstate.py` | state line reads `Kaelen (player):`; `player.hp` still resolves |
|
||||
| `test_prompt_caching.py` | persona is in the static prefix, above history |
|
||||
| `test_bundle_v2.py` | round-trips; a bundle without the key still imports |
|
||||
|
||||
## Bundle format
|
||||
|
||||
No `FORMAT` bump. The export gains a `persona` key and the import reads it with
|
||||
`.get()`, so a v2 bundle written before this change imports with an empty
|
||||
persona — which is the same as not having one.
|
||||
|
||||
---
|
||||
|
||||
# Phase 2 — what the summarizer is told
|
||||
|
||||
Depends on Phase 1 only for the protagonist's name. Everything else it needs is
|
||||
already in the database and simply never passed.
|
||||
|
||||
## The cast brief
|
||||
|
||||
Build one short block and prepend it to both the memory prompt and the summary
|
||||
prompt:
|
||||
|
||||
```
|
||||
Cast:
|
||||
- Kaelen (he/him) — the protagonist. A half-elf ranger, exiled from the
|
||||
northern holds...
|
||||
- Gwen — a loyal ranger and Kaelen's ally. Quick with a bow, dry-humoured.
|
||||
|
||||
Setting:
|
||||
<adventure.memory, the plot essentials>
|
||||
```
|
||||
|
||||
Sources, in order:
|
||||
|
||||
1. **Protagonist** — the Phase 1 persona. Falls back to "the player" when unset,
|
||||
so the change still helps adventures with no persona.
|
||||
2. **NPCs** — `stat_schema.npcs`: `name` + `desc`. Already there, never used
|
||||
outside the turn prompt.
|
||||
3. **Story cards** — for adventures with no RPG layer, run the existing
|
||||
`_match_cards()` over the block being summarized and take the matched cards'
|
||||
names and entries. This reuses the trigger code and yields exactly the
|
||||
entities that appear in *that block*, not the whole world.
|
||||
|
||||
## Two rules for the brief
|
||||
|
||||
**Fixed descriptions only. No live stats.** It is tempting to include
|
||||
`Gwen: trust 40 (wary)`. Do not. It makes every memory prompt different, which
|
||||
loses prompt caching, and worse, it makes the same event summarized at two
|
||||
different times come out framed differently — which is the problem being fixed.
|
||||
|
||||
**The summary gets it for free.** `_update_story_summary` builds from the memory
|
||||
bullets, so better memories produce a better summary with no further change.
|
||||
Pass the brief there too, because that function falls back to raw story text
|
||||
when memory-writing has fallen behind.
|
||||
|
||||
## The prompt changes
|
||||
|
||||
`MEMORY_SYSTEM_PROMPT` gains an explicit framing rule:
|
||||
|
||||
- write in third person, never "you"
|
||||
- refer to the protagonist by name
|
||||
- name characters rather than using bare pronouns
|
||||
|
||||
Expected effect:
|
||||
|
||||
```
|
||||
before: You entered a room and she stopped you.
|
||||
after: Kaelen bribed Gwen with fifty silver to hold the north door while
|
||||
he went down alone.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Phase 1 as built
|
||||
|
||||
Everything above describes what shipped. Three notes on where it differs from
|
||||
the sketch, and what has not been verified.
|
||||
|
||||
## Decisions taken during the build
|
||||
|
||||
**The section reads as one paragraph.** `render_persona` joins the sentence and
|
||||
the description with a space rather than `SEPARATOR`. A blank line inside a
|
||||
two-sentence section about one character reads as two unrelated notes.
|
||||
|
||||
**`Field` gained a `maxLength` prop.** The persona name is `VARCHAR(80)` and the
|
||||
schema rejects 81 characters with a 422. Without the attribute the player only
|
||||
learns that from a toast after typing, so the cap is now enforced in the input
|
||||
as well. Every other `Field` is unaffected — the prop is optional.
|
||||
|
||||
**The blank-adventure button opens the modal too.** It used to create the
|
||||
adventure immediately. It now collects a persona first, and passes
|
||||
`title: 'Blank Adventure'` explicitly, because there is no scenario to take a
|
||||
title from.
|
||||
|
||||
**`test_prompt_caching.py`'s fixture now sets a persona**, so every test in the
|
||||
file that guards the static block runs with one present.
|
||||
|
||||
## Verified
|
||||
|
||||
- Migration 73 → 76 on a real pre-existing SQLite database: columns added, the
|
||||
existing row preserved, persona empty, `render_persona` returns `""`.
|
||||
- 593 backend tests pass, including 26 new ones in `test_persona.py` and 3 in
|
||||
`test_bundle_v2.py`.
|
||||
- `npm run build` is clean.
|
||||
|
||||
## Driven in a browser
|
||||
|
||||
Chromium via Playwright, against a fresh database with the demo scenarios
|
||||
seeded. No API key needed: Insights assembles the prompt without calling a
|
||||
model, so every check below runs on the real assembled context rather than on
|
||||
a unit-test stub. 21/21 checks passed.
|
||||
|
||||
What was confirmed on screen and in the live `/context` payload:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| The modal opens for a scenario with **no** `${...}` placeholders | it is now the only way to name a character, so it can no longer be conditional |
|
||||
| The section order is real | `narrator, world_state_guide, world_state_rule, ai_instructions, **persona**, plot_essentials, …` |
|
||||
| The persona is in the system half | `You are Kaelen (he/him). A half-elf ranger…` present in `prompt.system`, absent from `prompt.story` |
|
||||
| The live stat line carries the name and the path | `Kaelen (player): hp 100/100 (full health), mana 30/50 (brimming)…` |
|
||||
| The stat guide ties the name to the path | `Protagonist Kaelen … player.<stat>` |
|
||||
| The drawer heading follows the name | rail read `KAELEN`, not `You` |
|
||||
| A rename propagates | renamed to Aria in the Plot panel → drawer read `ARIA` after a reload, and the prompt read `You are Aria (he/him).` / `Aria (player):` |
|
||||
| Clearing every field restores the old behavior exactly | no `persona` section, stat line back to `You:`, no `Protagonist` line in the guide |
|
||||
| The blank-adventure button collects a persona and keeps its title | `title='Blank Adventure' persona='Wren'` |
|
||||
| **A persona works with no RPG layer at all** | a blank adventure's sections were `['narrator', 'persona', 'length_hint']` — the case this feature was added for |
|
||||
|
||||
Two request failures in the run were the sandbox rather than the app: Google
|
||||
Fonts is blocked by the egress policy, and the analytics beacon is aborted when
|
||||
the page unloads. Neither appears with normal network access.
|
||||
|
||||
The driver script is not in the repo. There is no frontend test runner yet
|
||||
(plan/17 stage 5), and one Playwright script is not the place to start one.
|
||||
|
||||
## A note for whoever runs the tests
|
||||
|
||||
`tiktoken` downloads `cl100k_base` from `openaipublic.blob.core.windows.net` on
|
||||
first use, and 182 tests fail with a proxy error where that host is blocked.
|
||||
The encoding is reconstructable offline from the npm package `js-tiktoken`,
|
||||
whose `dist/ranks/cl100k_base.cjs` holds the same ranks in a compressed form —
|
||||
decode it, write `<base64> <rank>` per line sorted by rank, and the result
|
||||
matches the SHA-256 that `tiktoken_ext/openai_public.py` hardcodes, so the
|
||||
reconstruction is verifiable rather than trusted. Drop it at
|
||||
`$TIKTOKEN_CACHE_DIR/<sha1 of the URL>`.
|
||||
|
||||
---
|
||||
|
||||
# Phase 2 as built
|
||||
|
||||
## The cast comes from the story cards, not from `stat_schema`
|
||||
|
||||
This is the one thing the sketch above got wrong, and it made the change much
|
||||
smaller. `scenario_text.scenario_card_specs` already turns **every schema NPC
|
||||
into a story card** on the adventure at creation, deduplicated against the
|
||||
hand-written cards by name. So the cards are a single unified cast source that
|
||||
covers schema NPCs, an author's own cards, and an adventure with no RPG layer,
|
||||
through one path instead of three. `memorybank` never reads `stat_schema`.
|
||||
|
||||
## Keyword matching alone was not enough
|
||||
|
||||
The sketch said to run `_match_cards` over the block. Built that way first, and
|
||||
it failed the exact case the change exists for.
|
||||
|
||||
The block `You push the door open. She grabs your arm.` matches **no** card
|
||||
keyword, so the brief listed the protagonist and nobody else — leaving the
|
||||
summarizer guessing at precisely the moment it was handed a brief to stop
|
||||
guessing. Seed 04 gives Gwen the trigger keys `"Gwen, ranger, her"`, and even
|
||||
that does not save it: the text says "she", not "her".
|
||||
|
||||
So the roster is **matched cards first, then topped up with the other
|
||||
`character` cards** to `MAX_CAST_MEMBERS`. Places and items are not topped up —
|
||||
an unmentioned tavern is not who "she" was — but a place that *is* mentioned
|
||||
still matches normally.
|
||||
|
||||
That asymmetry with the turn prompt is deliberate. Including an untriggered card
|
||||
as lore would be wrong: it is not relevant to the next sentence. Including an
|
||||
untriggered character in a roster is right: the question the roster answers is
|
||||
"who could these pronouns be", not "what is on stage".
|
||||
|
||||
## `_match_cards` became `match_cards`
|
||||
|
||||
Two callers now run the same rule, so it is public and exported from
|
||||
`app.context`. One rule, one implementation.
|
||||
|
||||
## What actually gets sent
|
||||
|
||||
Verified against the seeded Bandit Camp scenario, with the real
|
||||
`_create_due_memories` path and a stub provider:
|
||||
|
||||
```
|
||||
Cast:
|
||||
- Kaelen (he/him) — the protagonist. A half-elf ranger, exiled from the
|
||||
northern holds for a killing he still won't explain.
|
||||
- Bandit Camp — A rough camp of bandits in a forest clearing, holding a
|
||||
stolen caravan strongbox.
|
||||
- Gwen — A loyal ranger and the player's ally. Quick with a bow,
|
||||
dry-humoured, fiercely protective. …
|
||||
- Bandit Leader — The scarred leader of the bandit camp, guarding the
|
||||
strongbox. …
|
||||
|
||||
Setting:
|
||||
The player and Gwen, a loyal ranger ally, are raiding a bandit camp to
|
||||
recover a stolen strongbox. …
|
||||
|
||||
Story excerpt:
|
||||
|
||||
… You are six paces from the strongbox when she hisses a warning. …
|
||||
|
||||
Memory:
|
||||
```
|
||||
|
||||
That last line is the case in miniature: "she" is now resolvable.
|
||||
|
||||
**With no persona set**, the roster still lists the NPCs and the setting, and
|
||||
the system prompt tells the model to call the protagonist "the player". Phase 2
|
||||
therefore improves adventures that never set a persona at all.
|
||||
|
||||
**With no persona, no cards and no plot essentials**, the user message is byte
|
||||
for byte what it was before this change — `Story excerpt:` first, no stray blank
|
||||
lines. A test holds that.
|
||||
|
||||
## Decided, from the sketch's open questions
|
||||
|
||||
**Existing memories: left alone at first, and now rewritable on demand.** The
|
||||
original answer was to let eviction age them out at `memory_bank_capacity` (80),
|
||||
on the grounds that re-summarizing would duplicate whatever was still in the
|
||||
bank, because nothing deletes the old rows. That reasoning was wrong about the
|
||||
only option: a memory can be rewritten in place rather than written again.
|
||||
See "Rewriting a bank written under the old prompt", at the end of this file.
|
||||
|
||||
**Retrieval framing mismatch: watched, not fixed.** `retrieve_memories` still
|
||||
embeds the last 4 actions raw, in second person, while new memories are third
|
||||
person and named. Embeddings handle paraphrase well, so this is speculative.
|
||||
If retrieval quality visibly dips, prepending the same brief to the query text
|
||||
is the first thing to try.
|
||||
|
||||
## Run with a real model, and what it changed
|
||||
|
||||
Run end to end against a real model. One story generated through the app's own
|
||||
`build_context` a turn at a time, then **both** memory prompts run over the
|
||||
**same** blocks, so the story is held constant and the prompt is the only
|
||||
variable. Neither arm sees the other's output, and the model is never told what
|
||||
is being measured. The control is `MEMORY_SYSTEM_PROMPT` as of commit `9cdcb55`,
|
||||
read out of git rather than pasted, so it cannot drift from what shipped.
|
||||
|
||||
The harness is `backend/tools/memory_ab.py`, and it goes through
|
||||
`OpenAICompatibleProvider` rather than calling a model directly, so the run
|
||||
exercises the provider, the streaming path and `complete()`. Pointed at
|
||||
`tools/claude_shim.py` it spends a Claude subscription instead of API credit:
|
||||
|
||||
python tools/claude_shim.py --port 8787 &
|
||||
python tools/memory_ab.py --out ab.md
|
||||
|
||||
**The full transcript, with every memory, both summaries and the story they were
|
||||
written from, is in `plan/18-appendix-memory-ab-run.md`.**
|
||||
|
||||
### The reported fault reproduced, and the fix held
|
||||
|
||||
| | words | framing |
|
||||
|---|---|---|
|
||||
| memory 1, before | 34 | second person — "**You** crept low through the mist…" |
|
||||
| memory 2, before | 105 | third person — "**The player** asked Gwen to…" |
|
||||
| memory 1, after | 72 | "**Kaelen** and Gwen crouched at dawn…" |
|
||||
| memory 2, after | 68 | "**Kaelen** and Gwen infiltrated the bandit camp…" |
|
||||
|
||||
Two consecutive memories in one bank, written from the same story minutes apart,
|
||||
in two different persons. That is the complaint, reproduced under controlled
|
||||
conditions rather than argued from the prompt text.
|
||||
|
||||
**A second run, on a freshly generated story, qualifies that.** Its two control
|
||||
memories were both "the player", consistently — the second-to-third person drift
|
||||
did not recur. So the drift is a real thing a model does, seen once, not
|
||||
something it does every time. Read the person-drift row as one observation.
|
||||
|
||||
What holds across both runs is the thing the change is actually for: **four
|
||||
control memories, and not one names the protagonist. Four treatment memories,
|
||||
and all four do.** The control has never been told a name. See
|
||||
`plan/18-appendix-memory-ab-run-2.md`.
|
||||
|
||||
The control also got a fact wrong that the treatment did not. The player's move
|
||||
was `grab her wrist and pull her down behind the woodpile`; the before-memory
|
||||
recorded "The player asked Gwen to grab her wrist and pull her down behind the
|
||||
woodpile", inverting who acted. One sample, so this is an observation rather
|
||||
than a claim — but it is the failure mode the brief predicts, since without a
|
||||
cast there is no way to tell whose wrist "her wrist" is.
|
||||
|
||||
### It also found a real problem, which is now fixed
|
||||
|
||||
"1-2 plain sentences" is not a length. The same model wrote 34 words for one
|
||||
block and 105 for the next. A 105-word memory is a paragraph, and
|
||||
`memory_top_k` injects five of them every turn, so the bank's running cost is
|
||||
set by a number nobody had ever stated.
|
||||
|
||||
`MEMORY_MAX_WORDS = 50` now states it, and the prompt says which details to keep
|
||||
when trimming: the ones a later scene could turn on. Re-run over the identical
|
||||
story, the same two blocks came back at **32 and 58 words**, still naming
|
||||
Kaelen, still third person, and still carrying every load-bearing fact — the
|
||||
camp map, the strongbox behind the second tent, the strap frayed near through.
|
||||
The second run, on different prose, came back at **54 and 55** against controls
|
||||
of 89 and 107. Overshooting 50 slightly is expected: models exceed word budgets,
|
||||
which is why `builder.length_hint` carries a `LENGTH_BUFFER` for the same reason.
|
||||
|
||||
Across both runs the control ran 34, 89, 105 and 107 words — a four-fold spread
|
||||
with no budget stated anywhere. That is the number this found.
|
||||
|
||||
The variance is the real gain. Before, across both runs: 34 to 107. After: 32
|
||||
to 58.
|
||||
|
||||
### What this does not show
|
||||
|
||||
The model behind the run is a Claude model. The app talks to an
|
||||
OpenAI-compatible endpoint, and the parser in `worldstate/parse.py` exists
|
||||
because weaker free models emit trailing commas and leading `+`. So this shows
|
||||
the prompt is followable and that the brief supplies the missing information.
|
||||
It does not show that a weaker production model complies as well. An explicit
|
||||
framing rule is usually worth *more* on a weaker model, but that is an
|
||||
expectation, not a measurement.
|
||||
|
||||
Two side observations from the same run, both pre-existing behavior working
|
||||
correctly: the model sent `{"milestones.strongbox_found": false}`, `apply_delta`
|
||||
refused it ("a milestone is sticky, so only true is accepted"), the refusal
|
||||
reached the model through `render_refusals`, and the next turn sent `true`. The
|
||||
Phase 1 persona also held across all six generated turns.
|
||||
|
||||
---
|
||||
|
||||
# Rewriting a bank written under the old prompt
|
||||
|
||||
An adventure played before this change keeps a bank of unnamed, second-person
|
||||
memories, and those are exactly the rows `memory_top_k` injects into every turn
|
||||
from now on. Ageing them out only works for an adventure that is still being
|
||||
played, and only after another 80 memories have been written. So there is a
|
||||
backfill: `backend/tools/rewrite_memories.py`.
|
||||
|
||||
```
|
||||
cd backend
|
||||
python -m tools.rewrite_memories # what would change
|
||||
python -m tools.rewrite_memories --write --limit 3 # try three of them
|
||||
python -m tools.rewrite_memories --write --embed # the whole backfill
|
||||
```
|
||||
|
||||
Without `--write` it makes no model calls and spends nothing. It reads whichever
|
||||
database the app reads — `AIDND_DB_PATH`, or `DATABASE_URL` on a hosted deploy.
|
||||
|
||||
**In place, not delete-and-regenerate.** The obvious alternative is to drop the
|
||||
bank and rewind `cursors.MEMORY`, letting the post-turn pass write it again.
|
||||
That loses everything the row carries besides its text: whether it is pinned,
|
||||
how often it has been retrieved, and the node it hangs off, which is what makes
|
||||
a fork inherit the right memories and no others. It would also trickle the bank
|
||||
back at `MAX_MEMORIES_PER_RUN` per turn, so an adventure nobody is playing would
|
||||
never recover. Rewriting `text` keeps the row and costs one call per memory.
|
||||
|
||||
**One prompt assembly, not two.** `memorybank.summarize_block` is now the single
|
||||
place a memory prompt is built, and both the post-turn pass and the tool call
|
||||
it. A backfill that assembled its own prompt would be writing memories with a
|
||||
prompt that never shipped, and nothing would report the difference.
|
||||
|
||||
**Reading the block back is the one genuinely new part.** A memory records
|
||||
`source_start`, `source_end` and `branch_id`, and `memorybank.source_block`
|
||||
turns those back into actions. The subtlety is the branch: the read has to use
|
||||
the lineage of *the branch the memory was written on*, not the branch the
|
||||
adventure is playing now. After a fork, the same depths hold different actions
|
||||
on each side, so a read through the adventure's current path would summarize the
|
||||
wrong story and say nothing about it. `test_memory_rewrite.py` builds that fork
|
||||
and holds the rule.
|
||||
|
||||
**What it will not touch:**
|
||||
|
||||
- A memory with no source range — hand-written, or migrated by 62 from before
|
||||
memories had coordinates. There is no block to rewrite it from, and the player
|
||||
may have typed it.
|
||||
- A memory whose actions have since been deleted.
|
||||
- An adventure whose owner has no API key in Settings, because summarization
|
||||
spends the user's own key by construction and never the shared demo key.
|
||||
`--api-key`, `--model` and `--endpoint` override that — the last of these
|
||||
points a run at `tools/claude_shim.py`, so a backfill can spend a Claude
|
||||
subscription instead of API credit.
|
||||
|
||||
**The vector is cleared for every memory it rewrites**, because the stored one
|
||||
describes wording that no longer exists. That takes the memory out of the ranked
|
||||
bank until something embeds the new text: the app's own post-turn pass does it
|
||||
`MAX_EMBED_BATCH` at a time, or `--embed` does it in the run. Re-embedding always
|
||||
uses the owner's own embedding model and endpoint, never `--endpoint`, because a
|
||||
vector is only meaningful against the vectors it is ranked beside.
|
||||
|
||||
**Stop the app before running with `--embed`.** A running process caches vectors
|
||||
by memory id and expects to be the only writer (`memorybank._vector_cache`), so
|
||||
a vector written from outside it can sit behind a stale cached copy until it
|
||||
restarts. Clearing alone is safe at any time: an unembedded memory leaves the
|
||||
catalogue, which is what the cache invalidates on.
|
||||
|
||||
## Running it against the hosted deploy
|
||||
|
||||
Two things make production different from a local database, and both are easy
|
||||
to get wrong quietly.
|
||||
|
||||
**It holds other people's stories, and each adventure is summarized with its
|
||||
owner's key.** An unfiltered `--write` would spend other people's money on
|
||||
memories they did not ask to have rewritten. `--email` restricts a run to named
|
||||
accounts and `--adventure` to single adventures; the dry run costs nothing and
|
||||
prints the owner of each. Guests have no email and are reachable only by id,
|
||||
which is the right amount of friction for rewriting a stranger's bank.
|
||||
|
||||
**The stored API keys are encrypted with `AIDND_SECRET_KEY`.** Render generates
|
||||
that value and holds it for the web service, so a run from a checkout has to
|
||||
carry the same one. With a different secret, `decrypt_secret` returns "" rather
|
||||
than failing, and every adventure is reported as having no key — a run that
|
||||
looks like it worked and did nothing.
|
||||
|
||||
```
|
||||
AIDND_DATABASE_URL=<the Neon URL from the Render dashboard> \
|
||||
AIDND_SECRET_KEY=<the value the web service has> \
|
||||
python -m tools.rewrite_memories --email you@example.com
|
||||
```
|
||||
|
||||
Run it from a checkout rather than from a shell on Render. The image copies
|
||||
`backend/app` alone, so `tools/` is not on the box, and the free plan has no
|
||||
shell anyway. The database is the same one either way.
|
||||
|
||||
Two smaller notes for that environment. The Neon URL to use is the direct
|
||||
endpoint, not `-pooler`, for the same reason the sizing queries in STATUS use
|
||||
it. And an adventure owned by a visitor playing on the shared demo key is
|
||||
skipped, because summarization has never spent that key.
|
||||
|
||||
**The story summary is not rewritten.** It is one text per adventure rather than
|
||||
a bank, and `_update_story_summary` hands the model the whole of it and asks for
|
||||
an updated version under the new framing rule, so the next scheduled update
|
||||
should carry it over to third person by itself. If it does not, that is a
|
||||
separate and much smaller fix than this one.
|
||||
@@ -97,8 +97,8 @@ A clean production build can start, open the browser UI, generate and persist st
|
||||
|
||||
## Status: COMPLETE
|
||||
|
||||
Accepted 2026-09-02. Evidence: `planning/reports/M1-BASELINE-REPORT.md` (run
|
||||
logs and packet captures) and `planning/reports/M1-IMPLEMENTATION-REPORT.md`
|
||||
Accepted 2026-09-02. Evidence: `planning/archive/milestone-reports/M1-BASELINE-REPORT.md` (run
|
||||
logs and packet captures) and `planning/archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md`
|
||||
(review report). A01-A06, H01-H03 and H11 all pass on runtime evidence; 648
|
||||
backend tests pass, including with no route to the Internet.
|
||||
|
||||
@@ -168,8 +168,8 @@ The codebase has a narrow single-user/local-only surface and the inherited story
|
||||
|
||||
## Status: COMPLETE
|
||||
|
||||
Accepted 2026-09-02. Evidence: `planning/reports/M2-BASELINE-REPORT.md`
|
||||
(measurements) and `planning/reports/M2-IMPLEMENTATION-REPORT.md` (review);
|
||||
Accepted 2026-09-02. Evidence: `planning/archive/milestone-reports/M2-BASELINE-REPORT.md`
|
||||
(measurements) and `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` (review);
|
||||
verdict *accept with non-blocking debt, proceed to M3*. Implementation is
|
||||
commit `8c65ae9`, and the three defects the review found are commit `8652fe7`
|
||||
— see the closeout note appended to both reports.
|
||||
@@ -193,9 +193,10 @@ commit `8c65ae9`, and the three defects the review found are commit `8652fe7`
|
||||
**Debt carried forward, none of it blocking M3:** `Settings.model` still
|
||||
defaults to `""` with nothing prompting for it (M8); inert legacy tables and
|
||||
columns await a cleanup migration once the schema settles, after M3/M5; there
|
||||
are still no frontend tests (M8); `docs/*.html`, upstream's project site and not
|
||||
served by the app, still links Google Fonts. Full table in the implementation
|
||||
report §P.
|
||||
are still no frontend tests (M8). `docs/*.html`, upstream's project site, was
|
||||
listed here as still linking Google Fonts; the whole inherited `docs/` tree was
|
||||
deleted in the 2026-09-03 documentation pass, which closes that item. Full table
|
||||
in the implementation report §P.
|
||||
|
||||
---
|
||||
|
||||
@@ -434,7 +435,7 @@ undo semantics and are expected to be rewritten by M3; that is separate from
|
||||
this instrumentation point, and they should likewise be rewritten rather than
|
||||
dropped.
|
||||
|
||||
Evidence: `planning/reports/M2-IMPLEMENTATION-REPORT.md` §K.2, §Q.
|
||||
Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §K.2, §Q.
|
||||
|
||||
## Note from M3 — three constraints this milestone must satisfy
|
||||
|
||||
@@ -529,7 +530,7 @@ M6 therefore additionally requires:
|
||||
degraded the storyteller quietly and left the transcript correct, which is the
|
||||
right failure direction — but it must also be a *visible* one.
|
||||
|
||||
Evidence: `planning/reports/M2-IMPLEMENTATION-REPORT.md` §A.1, §9.1.
|
||||
Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §A.1, §9.1.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,14 +0,0 @@
|
||||
# Codex Handoff Note
|
||||
|
||||
**Status:** Phase 0B handoff complete; historical only.
|
||||
|
||||
The earlier Phase 0B execution prompts are retained for audit/history:
|
||||
|
||||
- `PHASE-0B-CODEX-BRIEF.md`
|
||||
- `PHASE-0B-CODEX-HANDOFF.md`
|
||||
|
||||
Do **not** execute either as the next production task.
|
||||
|
||||
Phase 0B has been completed and the resulting recommendation reviewed. AI-DnD is now the selected production base, and the planning package has been revised accordingly.
|
||||
|
||||
The next production Codex prompt has **not** been prepared. It should be created only after the current planning-package revision is reviewed and approved. When authorized, the first implementation prompt should be derived from Production Milestone M1 in `BUILD-MILESTONES.md`, not from the Phase 0B validation briefs.
|
||||
@@ -57,4 +57,4 @@ Therefore:
|
||||
`http://…:11434`.
|
||||
|
||||
M1 implemented this (`backend/app/tlstrust.py`); see
|
||||
`planning/reports/M1-IMPLEMENTATION-REPORT.md` §G.
|
||||
`planning/archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md` §G.
|
||||
|
||||
@@ -115,4 +115,4 @@ itself beyond loopback, which remains out of scope for v1.
|
||||
- `SECURITY-THREAT-MODEL.md` §10A, §71A item 5, §77
|
||||
- `TECHNICAL-DESIGN.md` §5.1 item 4, §5.2
|
||||
- ADR 002 (Ollama-only, and the TLS consequence), ADR 004 (local-only production)
|
||||
- `planning/reports/M2-IMPLEMENTATION-REPORT.md` §F, §K.1
|
||||
- `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §F, §K.1
|
||||
|
||||
@@ -1,272 +0,0 @@
|
||||
# Phase 0B — Codex Initial Validation Brief
|
||||
|
||||
**Status:** COMPLETE / HISTORICAL — do not execute as a production prompt
|
||||
**Purpose:** Run a focused first round of local validation on the three finalist repositories and return a recommendation based on what was actually learned.
|
||||
|
||||
## 1. Goal
|
||||
|
||||
We are **not** asking you to build the production application yet.
|
||||
|
||||
The goal of this round is to answer one question:
|
||||
|
||||
> Which existing project is the best starting point for the local interactive-story application, and what important technical facts did we learn that should affect the next design step?
|
||||
|
||||
The three finalists are:
|
||||
|
||||
1. AI-DnD
|
||||
https://github.com/parththakkar106/AI-DnD
|
||||
|
||||
2. Open Dungeon
|
||||
https://github.com/newideas99/open-dungeon
|
||||
|
||||
3. ai-adventure
|
||||
https://github.com/CaoRuiming/ai-adventure
|
||||
|
||||
## 2. Important Product Requirements
|
||||
|
||||
Use these as the main evaluation criteria.
|
||||
|
||||
The eventual application should be:
|
||||
|
||||
- browser-first,
|
||||
- local-only in v1,
|
||||
- based on local Ollama inference,
|
||||
- single-user,
|
||||
- genre-agnostic,
|
||||
- persistent across restarts,
|
||||
- able to retain authoritative story state separately from model prose,
|
||||
- able to Undo/Redo/Retry safely,
|
||||
- able to preserve named save points/checkpoints,
|
||||
- able to retain abandoned history without immediately deleting it,
|
||||
- able to prevent abandoned-history facts/memories from leaking into the active story,
|
||||
- able to support long-term context/memory,
|
||||
- able to import local knowledge,
|
||||
- able to inspect what context was sent to the model,
|
||||
- architecturally compatible with future local image/video/TTS/STT support.
|
||||
|
||||
Do not try to implement all of these now.
|
||||
|
||||
This round is about determining which candidate already gives us the strongest foundation.
|
||||
|
||||
## 3. Read Only What You Need
|
||||
|
||||
Start with:
|
||||
|
||||
1. `README.md`
|
||||
2. `SPECIFICATION.md`
|
||||
3. `reports/PRELIMINARY-RECOMMENDATION.md`
|
||||
4. `reports/REUSE-MATRIX.md`
|
||||
|
||||
Then use these only when relevant to a specific experiment:
|
||||
|
||||
- `STORY-BRANCH-SEMANTICS.md`
|
||||
- `CONTEXT-AND-MEMORY.md`
|
||||
- `SECURITY-THREAT-MODEL.md`
|
||||
- `TEST-CAMPAIGN-FIXTURE.md`
|
||||
|
||||
Do **not** read every planning document up front unless needed.
|
||||
|
||||
## 4. Baseline Work for Each Candidate
|
||||
|
||||
For each repository:
|
||||
|
||||
1. Clone it cleanly.
|
||||
2. Record the exact commit SHA.
|
||||
3. Follow the documented install instructions.
|
||||
4. Run the existing tests.
|
||||
5. Build/start the application.
|
||||
6. Confirm the basic local story flow.
|
||||
7. Record:
|
||||
- test results,
|
||||
- storage/database technology,
|
||||
- model/provider assumptions,
|
||||
- local ports,
|
||||
- major runtime failures,
|
||||
- obvious cloud/hosted dependencies.
|
||||
|
||||
Do not spend excessive time fixing unrelated upstream problems.
|
||||
|
||||
If a project does not run cleanly, document why and continue.
|
||||
|
||||
## 5. Focused Experiment A — AI-DnD
|
||||
|
||||
We want to know whether AI-DnD can realistically serve as the production base.
|
||||
|
||||
Test:
|
||||
|
||||
- Can it run with Ollama locally?
|
||||
- Can its story-tree/history system support simple user-facing Undo/Redo/Retry behavior?
|
||||
- Does state rollback work independently of heavy RPG/stat mechanics?
|
||||
- Can RPG-specific state be left empty/minimal without breaking the useful history/state architecture?
|
||||
- Can QuickJS/scripting and hosted/cloud-oriented features be disabled without breaking local story operation?
|
||||
- Can local memory/embedding behavior work without cloud services?
|
||||
- Do memories/state respect the active history path?
|
||||
- Is its context/Insights system useful for showing what was sent to the model?
|
||||
|
||||
Use a small disposable experiment if necessary.
|
||||
|
||||
Do **not** start stripping the whole application down.
|
||||
|
||||
## 6. Focused Experiment B — Open Dungeon
|
||||
|
||||
We want to know how expensive it would be to fix its history model.
|
||||
|
||||
Test:
|
||||
|
||||
- Confirm how Retry/Edit/Erase affect stored history.
|
||||
- Identify whether old future turns are deleted.
|
||||
- Trace which parts of the application depend on that linear/destructive behavior.
|
||||
- Estimate how invasive it would be to change to:
|
||||
- parent-linked turns,
|
||||
- active head,
|
||||
- retained abandoned history,
|
||||
- named checkpoints,
|
||||
- lineage-safe summaries/state.
|
||||
|
||||
Do not implement the complete branch system.
|
||||
|
||||
Also record useful existing pieces:
|
||||
- browser UX,
|
||||
- local image generation,
|
||||
- character visual continuity,
|
||||
- any scene/media architecture worth reusing.
|
||||
|
||||
## 7. Focused Experiment C — ai-adventure
|
||||
|
||||
We want to know whether its strong state/privacy architecture can realistically become a browser-based Ollama application.
|
||||
|
||||
Test:
|
||||
|
||||
- Run the existing tests.
|
||||
- Confirm Undo/branch/checkpoint/replay behavior.
|
||||
- Identify the provider abstraction.
|
||||
- Prove one local Ollama-backed story turn using the smallest practical adapter.
|
||||
- Determine how tightly the core application logic is coupled to the CLI.
|
||||
- Assess whether the core could sit behind a browser/API layer without moving authoritative state logic.
|
||||
- Review its local lore/FTS approach for possible reuse.
|
||||
|
||||
Do not build a browser frontend.
|
||||
|
||||
## 8. Offline / Privacy Check
|
||||
|
||||
For each candidate, once dependencies/models are installed:
|
||||
|
||||
- run it with outbound Internet unavailable or blocked where practical,
|
||||
- exercise basic story generation,
|
||||
- note any unexpected network attempts.
|
||||
|
||||
We do not need a full penetration test in this round.
|
||||
|
||||
We do need to know:
|
||||
|
||||
- whether local story use truly works offline,
|
||||
- whether cloud services are required,
|
||||
- whether analytics/telemetry/remote assets are present,
|
||||
- how difficult those paths would be to remove.
|
||||
|
||||
## 9. Use the Standard Fixture Selectively
|
||||
|
||||
Use `TEST-CAMPAIGN-FIXTURE.md` where it helps answer continuity questions.
|
||||
|
||||
You do not need to execute the entire fixture against every candidate.
|
||||
|
||||
The most important checks are:
|
||||
|
||||
- possession/state consistency,
|
||||
- restore/undo behavior,
|
||||
- abandoned-path isolation,
|
||||
- whether an old discarded fact can leak into current memory/context.
|
||||
|
||||
## 10. What Not to Do
|
||||
|
||||
Do not:
|
||||
|
||||
- build the production fork,
|
||||
- merge repositories,
|
||||
- redesign the full UI,
|
||||
- implement full RAG,
|
||||
- implement complete branching in Open Dungeon,
|
||||
- remove all RPG code from AI-DnD,
|
||||
- build a browser frontend for ai-adventure,
|
||||
- add image/video/TTS/STT features,
|
||||
- write the final production milestone plan.
|
||||
|
||||
Small disposable code changes are allowed only when needed to answer the evaluation questions.
|
||||
|
||||
## 11. Final Deliverable
|
||||
|
||||
The main output from this round should be a single recommendation document:
|
||||
|
||||
```text
|
||||
PHASE-0B-RECOMMENDATION.md
|
||||
```
|
||||
|
||||
It should summarize what was learned, not just list test logs.
|
||||
|
||||
Include:
|
||||
|
||||
### A. Executive Recommendation
|
||||
|
||||
- Which repository should be the production base?
|
||||
- Confidence level: high / medium / low.
|
||||
- Did the initial AI-DnD recommendation hold up?
|
||||
|
||||
### B. What We Learned About Each Candidate
|
||||
|
||||
For each:
|
||||
- what worked,
|
||||
- what failed,
|
||||
- strongest reusable pieces,
|
||||
- major architectural problems,
|
||||
- likely amount/type of adaptation required.
|
||||
|
||||
### C. Key Technical Findings
|
||||
|
||||
Especially:
|
||||
- history/undo model,
|
||||
- state rollback,
|
||||
- memory isolation,
|
||||
- local Ollama support,
|
||||
- offline/privacy behavior,
|
||||
- browser suitability,
|
||||
- imported-knowledge potential,
|
||||
- future media extension potential.
|
||||
|
||||
### D. Important Surprises
|
||||
|
||||
Anything that contradicts the current planning assumptions.
|
||||
|
||||
### E. Recommendation for Next Step
|
||||
|
||||
Do **not** perform the next step.
|
||||
|
||||
Instead recommend what should happen next, such as:
|
||||
- fork AI-DnD and begin a controlled strip-down,
|
||||
- perform one additional experiment first,
|
||||
- reconsider Open Dungeon,
|
||||
- use ai-adventure as the base instead,
|
||||
- revise one of the product assumptions.
|
||||
|
||||
### F. Open Questions
|
||||
|
||||
List anything that could not be resolved in this round.
|
||||
|
||||
## 12. Supporting Evidence
|
||||
|
||||
You may also create concise supporting notes/logs for:
|
||||
|
||||
- baseline results,
|
||||
- AI-DnD experiment,
|
||||
- Open Dungeon history analysis,
|
||||
- ai-adventure Ollama adapter,
|
||||
- offline/network observations.
|
||||
|
||||
Keep them concise.
|
||||
|
||||
The recommendation document is the primary deliverable.
|
||||
|
||||
## 13. Stop Condition
|
||||
|
||||
When `PHASE-0B-RECOMMENDATION.md` is complete, stop.
|
||||
|
||||
We will take the findings back into the design discussion, re-examine the assumptions, and decide the next step before any production implementation begins.
|
||||
@@ -1,277 +0,0 @@
|
||||
# Phase 0B — Codex Local Validation Handoff
|
||||
|
||||
**Status:** COMPLETE / HISTORICAL — do not execute as a production prompt
|
||||
**Purpose:** Validate the Phase 0A recommendation using local builds, tests, offline runtime observation, and tightly scoped experiments.
|
||||
**Stop rule:** Do not begin production implementation.
|
||||
|
||||
## 1. Read Before Starting
|
||||
|
||||
Read the package in the order listed in `README.md`.
|
||||
|
||||
At minimum, before modifying any finalist, read:
|
||||
|
||||
1. `SPECIFICATION.md`
|
||||
2. `DATA-MODEL.md`
|
||||
3. `STORY-BRANCH-SEMANTICS.md`
|
||||
4. `CONTEXT-AND-MEMORY.md`
|
||||
5. `IMPORTED-KNOWLEDGE-DESIGN.md`
|
||||
6. `SECURITY-THREAT-MODEL.md`
|
||||
7. `MEDIA-EXTENSION-CONTRACT.md`
|
||||
8. `BROWSER-UX-SPEC.md`
|
||||
9. `TEST-CAMPAIGN-FIXTURE.md`
|
||||
10. `V1-ACCEPTANCE-TESTS.md`
|
||||
11. `reports/PRELIMINARY-RECOMMENDATION.md`
|
||||
12. `reports/REUSE-MATRIX.md`
|
||||
|
||||
Treat the detailed behavioral documents and acceptance tests as the target behavior. Treat `TECHNICAL-DESIGN.md` as provisional.
|
||||
|
||||
## 2. Finalists to Clone
|
||||
|
||||
Clone only these three primary finalists for Phase 0B:
|
||||
|
||||
1. https://github.com/parththakkar106/AI-DnD
|
||||
2. https://github.com/newideas99/open-dungeon
|
||||
3. https://github.com/CaoRuiming/ai-adventure
|
||||
|
||||
At clone time record exact commit SHA, branch/tag, date, license, dependency lockfiles, required runtimes, and documented local model/provider assumptions.
|
||||
|
||||
Keep each upstream clone clean. Use separate experiment branches/worktrees for disposable changes. Do not merge candidate repositories together.
|
||||
|
||||
## 3. Result Codes
|
||||
|
||||
Use consistently:
|
||||
|
||||
```text
|
||||
PASS
|
||||
PARTIAL
|
||||
FAIL
|
||||
NOT IMPLEMENTED
|
||||
NOT APPLICABLE
|
||||
```
|
||||
|
||||
Do not convert an untested requirement into a PASS.
|
||||
|
||||
## 4. Validation V0 — Environment and Baseline
|
||||
|
||||
For all three:
|
||||
|
||||
- install from documented instructions,
|
||||
- run existing test suite,
|
||||
- run build/lint/typecheck where applicable,
|
||||
- record failures,
|
||||
- record actual current test count,
|
||||
- record local data paths,
|
||||
- record listening ports,
|
||||
- record child processes/services,
|
||||
- record model/provider configuration,
|
||||
- record database/storage technology.
|
||||
|
||||
Deliver one baseline report per project. Do not rely on README claims for test counts or feature behavior.
|
||||
|
||||
## 5. Validation V1 — Offline and Network Behavior
|
||||
|
||||
After dependencies and local models are already installed, block outbound Internet and exercise launch, story creation, 5+ turns, restart/resume, summaries, memory/embeddings if present, Retry, Undo/rewind, checkpoint/branch features if present, `.txt`/`.md` import if present, and Open Dungeon local image generation if configured.
|
||||
|
||||
Capture open sockets, DNS attempts, HTTP(S)/WebSocket destinations, and which feature caused each request.
|
||||
|
||||
Use `SECURITY-THREAT-MODEL.md` and acceptance groups A, G, and H.
|
||||
|
||||
Pass condition for target v1 operation:
|
||||
|
||||
> Story content, imported knowledge, prompt/context data, and media prompts do not leave loopback or explicitly approved local endpoints.
|
||||
|
||||
## 6. Standard Fixture Use
|
||||
|
||||
Use `TEST-CAMPAIGN-FIXTURE.md` as the standard narrative test bed.
|
||||
|
||||
Where a finalist cannot represent the fixture directly, map it as closely as possible, document the mismatch, and do not silently change expected truth/state to suit the candidate.
|
||||
|
||||
Important checks:
|
||||
|
||||
- Mara's knowledge boundaries,
|
||||
- Silver Key ownership,
|
||||
- resurrection canon,
|
||||
- Canon vs Reference vs Inspiration authority,
|
||||
- Path A secret followed by restore/divergence into Path B,
|
||||
- abandoned-history memory isolation,
|
||||
- long-term memory plant,
|
||||
- checkpoint persistence,
|
||||
- science-fiction variant.
|
||||
|
||||
## 7. Acceptance-Test Mapping
|
||||
|
||||
Use `V1-ACCEPTANCE-TESTS.md` as the common comparison contract. Produce a gap matrix rather than forcing each candidate to fully pass v1.
|
||||
|
||||
Use these interpretations:
|
||||
|
||||
```text
|
||||
Already passes
|
||||
Passes with configuration
|
||||
Small adaptation
|
||||
Foundational redesign
|
||||
Not present
|
||||
```
|
||||
|
||||
Prioritize high-risk groups:
|
||||
|
||||
- A01-A05 — local operation/persistence
|
||||
- C01-C05 — canon/state
|
||||
- D01-D14 — Undo/Redo/Retry/checkpoints
|
||||
- E01-E04 — lineage safety
|
||||
- F01-F08 — memory/context
|
||||
- G01-G10 — imported knowledge where supported
|
||||
- H01-H10 — security/privacy
|
||||
- J01-J03 — genre independence
|
||||
- K01-K04 — media readiness
|
||||
- L01-L04 — data integrity
|
||||
|
||||
Long-run M01-M04 need not be fully executed against every candidate if disproportionate; identify production risk and existing test coverage instead.
|
||||
|
||||
## 8. Experiment V2 — AI-DnD Strip-Down Feasibility
|
||||
|
||||
Do not redesign the application.
|
||||
|
||||
Answer:
|
||||
|
||||
1. Can a scenario run with RPG stats absent, empty, or minimal?
|
||||
2. Do branch/retry/undo/tree tests operate independently of RPG mechanics?
|
||||
3. Can user-facing tree complexity be hidden behind `STORY-BRANCH-SEMANTICS.md`?
|
||||
4. Disable QuickJS scripting. What breaks?
|
||||
5. Disable/remove hosted multi-user/auth/demo/analytics paths. What breaks locally?
|
||||
6. Configure only local Ollama generation.
|
||||
7. Configure only local embeddings, preferably Ollama/local.
|
||||
8. Verify branch switching restores correct generic state.
|
||||
9. Verify memory retrieval respects active lineage.
|
||||
10. Test whether abandoned Path A facts leak into Path B.
|
||||
11. Map Story Cards/world-info to Canon / Reference / Inspiration.
|
||||
12. Determine whether prompt/context snapshots satisfy Context Inspector requirements.
|
||||
13. Determine whether visual/scene snapshot data can be added without RPG coupling.
|
||||
14. Inventory code coupled to RPG worldstate, scripting, hosted auth, analytics, remote providers, and AI Dungeon compatibility.
|
||||
|
||||
Estimate invasiveness by affected files/modules, not hours. Do not merge the experiment.
|
||||
|
||||
## 9. Experiment V3 — Open Dungeon Branch Retrofit Impact
|
||||
|
||||
Do not implement full branching.
|
||||
|
||||
Trace message CRUD, Retry, Erase, Edit, Continue, summary generation, state/character persistence, image association, and visual continuity. Confirm destructive-tail assumptions.
|
||||
|
||||
Design a minimal hypothetical persistence change supporting:
|
||||
|
||||
```text
|
||||
turn/node ID
|
||||
parent turn ID
|
||||
active head
|
||||
alternate narrator takes
|
||||
retained disposable history
|
||||
checkpoint pointer
|
||||
lineage-safe summaries/memories
|
||||
```
|
||||
|
||||
Use the standard fixture to reason through Undo/restore, Path A -> Path B divergence, stale summary/state risks, and image attachment after divergence.
|
||||
|
||||
Also inspect local image/provider patterns for reuse. Measure invasiveness; do not build the branch system.
|
||||
|
||||
## 10. Experiment V4 — ai-adventure Ollama / Service Boundary
|
||||
|
||||
1. Run existing tests unchanged.
|
||||
2. Identify provider interface.
|
||||
3. Prove one Ollama-backed turn using the smallest disposable adapter possible.
|
||||
4. Identify modules that know about the CLI.
|
||||
5. Determine whether the app/state layer can be wrapped by a browser/API service without moving authoritative logic.
|
||||
6. Verify undo, branch, checkpoint, restore, replay.
|
||||
7. Evaluate event/commit discipline for reuse.
|
||||
8. Evaluate its FTS/lore system against `IMPORTED-KNOWLEDGE-DESIGN.md`.
|
||||
9. Determine difficulty of adding semantic local retrieval while preserving lexical retrieval.
|
||||
10. Check privacy boundary with the Ollama adapter.
|
||||
|
||||
Do not build a browser UI.
|
||||
|
||||
## 11. Validation V5 — Test Quality
|
||||
|
||||
For each finalist report actual test count, categories, branch/rollback coverage, state reconstruction, migrations, summary/memory coverage, provider mocks, offline/network tests, browser tests, security tests, flaky/failing tests, and tests requiring Internet.
|
||||
|
||||
Highlight which high-risk acceptance requirements already have regression coverage.
|
||||
|
||||
## 12. Validation V6 — Imported Knowledge Gap Analysis
|
||||
|
||||
Against `IMPORTED-KNOWLEDGE-DESIGN.md`, report local `.txt`/`.md` ingestion, classifications, campaign isolation, provenance, lexical/semantic search, embedding provider, enable/disable, deletion, export/import, hidden canon, prompt-injection framing, and remote URL/image behavior.
|
||||
|
||||
Do not implement a full new RAG subsystem during Phase 0B.
|
||||
|
||||
## 13. Validation V7 — Browser UX Gap Analysis
|
||||
|
||||
Against `BROWSER-UX-SPEC.md`, report story reading/input quality, streaming, Undo/Redo/Retry UI, alternate-take selection, edit behavior, Save Points, state inspection, knowledge management, prompt/context inspection, local-model status, and advanced complexity exposed to the user.
|
||||
|
||||
Explicitly identify AI-DnD components worth retaining and Open Dungeon components worth borrowing/reimplementing. Do not redesign the frontend.
|
||||
|
||||
## 14. Validation V8 — Future Media and Speech Readiness
|
||||
|
||||
Against `MEDIA-EXTENSION-CONTRACT.md`, determine whether the architecture can support future local image generation, video generation, audio/ambience, text-to-speech, and speech-to-text.
|
||||
|
||||
For STT, verify the architecture can support:
|
||||
|
||||
```text
|
||||
local microphone/audio
|
||||
->
|
||||
local STT provider
|
||||
->
|
||||
editable draft text
|
||||
->
|
||||
normal user submission
|
||||
```
|
||||
|
||||
STT output must not bypass the normal story commit path.
|
||||
|
||||
Do not implement STT/TTS/video during Phase 0B. Open Dungeon local image behavior may be exercised because it already exists.
|
||||
|
||||
## 15. Final Acceptance Gap Matrix
|
||||
|
||||
Produce a matrix organized by acceptance-test group covering A, C, D, E, F, G, H, I, J, K, L, plus UX fit. Include production impact for each gap.
|
||||
|
||||
## 16. Final Decision Matrix
|
||||
|
||||
Return:
|
||||
|
||||
| Question | AI-DnD | Open Dungeon | ai-adventure |
|
||||
|---|---|---|---|
|
||||
| Baseline builds | | | |
|
||||
| Existing tests pass | | | |
|
||||
| Runs offline after setup | | | |
|
||||
| Ollama works | | | |
|
||||
| History semantics fit | | | |
|
||||
| State authority fits | | | |
|
||||
| Memory/lineage fits | | | |
|
||||
| Imported knowledge fit | | | |
|
||||
| Prompt inspection fit | | | |
|
||||
| Security/local-only hardening | | | |
|
||||
| Unwanted-code removal scope | | | |
|
||||
| Browser UX fit | | | |
|
||||
| Media extension fit | | | |
|
||||
| Future TTS/STT fit | | | |
|
||||
| Major blockers | | | |
|
||||
|
||||
## 17. Recommendation Report
|
||||
|
||||
The final recommendation should answer:
|
||||
|
||||
1. Which single repository should be the production base?
|
||||
2. Why?
|
||||
3. What are the top architectural risks?
|
||||
4. What must be removed?
|
||||
5. What must be generalized?
|
||||
6. Which concepts/components should be reimplemented from other candidates?
|
||||
7. Does Phase 0B change the preliminary AI-DnD recommendation?
|
||||
8. Which open questions remain before `TECHNICAL-DESIGN.md` v1.0?
|
||||
9. Are any v1 requirements likely to need reconsideration because of real technical constraints?
|
||||
10. Is unlimited Undo straightforward? If not, what practical limit exists and why?
|
||||
|
||||
Use evidence, not repository popularity or feature count.
|
||||
|
||||
## 18. Stop Condition
|
||||
|
||||
Stop after baseline reports, offline/network evidence, three scoped experiments, test-quality report, acceptance-gap matrix, decision matrix, and final recommendation.
|
||||
|
||||
Do not start the production fork conversion, implement the complete branch system, build the final browser UI, implement full RAG, add video/TTS/STT, rewrite the production technical design, or write production milestones.
|
||||
|
||||
Return reports and experiment diffs/results for review. The fork/architecture decision will be made from those results.
|
||||
@@ -1,49 +0,0 @@
|
||||
# Planning Update Summary — Post Phase 0B Review
|
||||
|
||||
**Date:** 2026-09-01
|
||||
**Purpose:** Review aid. This file summarizes planning changes made after accepting the ten architecture decisions from the Phase 0B review.
|
||||
|
||||
## New Decisions Recorded
|
||||
|
||||
1. AI-DnD is the production base at pinned Phase 0B commit `d72f7c1b...`.
|
||||
2. AI-DnD's browser/service/story-tree/memory/context foundation is retained as the ownership center.
|
||||
3. The non-destructive head-cursor Undo/Redo spike is promoted into the production design, but the disposable spike is not treated as merge-ready production code.
|
||||
4. Narrative state will use explicit typed events/absolute assignments inspired by ai-adventure, not AI-DnD's relative-delta protocol.
|
||||
5. ai-adventure is an implementation reference, not the product specification.
|
||||
6. Open Dungeon is a UX/media reference only.
|
||||
7. Abandoned history is retained and marked disposable; no automatic cleanup is required in v1.
|
||||
8. Export/import must preserve active head position as part of the history work.
|
||||
9. Imported knowledge will be a separate first-class subsystem rather than an extension of Story Cards.
|
||||
10. Future image/video/audio/TTS/STT extension contracts remain, but no media generation is required for v1.
|
||||
11. Local-only inference may span user-controlled machines: same-host Ollama is the default, but an explicitly configured trusted-LAN Ollama host is supported in v1 without exposing the storyteller UI/API to the LAN.
|
||||
|
||||
## Documents Materially Revised
|
||||
|
||||
- `README.md`
|
||||
- `SPECIFICATION.md`
|
||||
- `TECHNICAL-DESIGN.md`
|
||||
- `DATA-MODEL.md`
|
||||
- `STORY-BRANCH-SEMANTICS.md`
|
||||
- `CONTEXT-AND-MEMORY.md`
|
||||
- `IMPORTED-KNOWLEDGE-DESIGN.md`
|
||||
- `SECURITY-THREAT-MODEL.md`
|
||||
- `V1-ACCEPTANCE-TESTS.md`
|
||||
- `BUILD-MILESTONES.md`
|
||||
- `RESEARCH-PLAN.md`
|
||||
- `CODEX-HANDOFF-NOTE.md`
|
||||
- `004-local-only-production.md`
|
||||
- `005-branch-preserving-history.md`
|
||||
- `008-phase0-before-build-plan.md`
|
||||
|
||||
## New ADRs
|
||||
|
||||
- `009-ai-dnd-production-base.md`
|
||||
- `010-explicit-typed-narrative-state-events.md`
|
||||
|
||||
## Historical Evidence Intentionally Preserved
|
||||
|
||||
The Phase 0A candidate reports and the coding agent's `PHASE-0B-RECOMMENDATION.md` remain research evidence. They may contain assumptions that are superseded by the current planning documents. They are not silently rewritten to make the historical analysis appear as if it had reached later conclusions originally.
|
||||
|
||||
## Coding Prompt Status
|
||||
|
||||
No next Codex implementation prompt is included in this revision.
|
||||
@@ -0,0 +1,121 @@
|
||||
# ChatGPT Project Sources — what to upload
|
||||
|
||||
This file exists for one purpose: to tell the repository owner which files
|
||||
belong in the **InteractiveStory** ChatGPT project's Sources. It is not a
|
||||
documentation index — `planning/README.md` is that.
|
||||
|
||||
`planning/project-sources.txt` is the same list as bare paths, one per line, for
|
||||
gathering the files.
|
||||
|
||||
The rule behind the list: **upload what is authoritative now.** A project source
|
||||
is treated as current fact by whatever reads it, so a superseded document
|
||||
uploaded alongside a current one does not add context — it manufactures a
|
||||
conflict.
|
||||
|
||||
Refresh the Sources whenever the planning package is revised (`VERSION.md`
|
||||
records each revision) and whenever a milestone completes.
|
||||
|
||||
## REQUIRED PROJECT SOURCES
|
||||
|
||||
Twenty-eight files.
|
||||
|
||||
**Repository orientation**
|
||||
|
||||
```text
|
||||
README.md
|
||||
DEVELOPMENT.md
|
||||
PROVENANCE.md
|
||||
```
|
||||
|
||||
**The planning index and its version**
|
||||
|
||||
```text
|
||||
planning/README.md
|
||||
planning/VERSION.md
|
||||
```
|
||||
|
||||
**Product requirements and design**
|
||||
|
||||
```text
|
||||
planning/SPECIFICATION.md
|
||||
planning/TECHNICAL-DESIGN.md
|
||||
planning/DATA-MODEL.md
|
||||
planning/STORY-BRANCH-SEMANTICS.md
|
||||
planning/CONTEXT-AND-MEMORY.md
|
||||
planning/IMPORTED-KNOWLEDGE-DESIGN.md
|
||||
planning/SECURITY-THREAT-MODEL.md
|
||||
planning/MEDIA-EXTENSION-CONTRACT.md
|
||||
planning/BROWSER-UX-SPEC.md
|
||||
```
|
||||
|
||||
**Implementation plan and acceptance contract**
|
||||
|
||||
```text
|
||||
planning/BUILD-MILESTONES.md
|
||||
planning/V1-ACCEPTANCE-TESTS.md
|
||||
planning/TEST-CAMPAIGN-FIXTURE.md
|
||||
```
|
||||
|
||||
**The active ADRs** — all eleven. They are short, and each one closes a question
|
||||
that will otherwise be reopened.
|
||||
|
||||
```text
|
||||
planning/DECISIONS/001-browser-first.md
|
||||
planning/DECISIONS/002-ollama-only-v1.md
|
||||
planning/DECISIONS/003-authoritative-local-state.md
|
||||
planning/DECISIONS/004-local-only-production.md
|
||||
planning/DECISIONS/005-branch-preserving-history.md
|
||||
planning/DECISIONS/006-genre-agnostic-core.md
|
||||
planning/DECISIONS/007-future-media-extension.md
|
||||
planning/DECISIONS/009-ai-dnd-production-base.md
|
||||
planning/DECISIONS/010-explicit-typed-narrative-state-events.md
|
||||
planning/DECISIONS/011-local-inference-endpoint-policy.md
|
||||
planning/DECISIONS/012-active-head-non-destructive-history.md
|
||||
```
|
||||
|
||||
ADR **008** is deliberately absent: it was the process gate requiring Phase 0 to
|
||||
close before production work began, Phase 0 closed on 2026-09-01, and it is now
|
||||
in `planning/archive/decisions/`.
|
||||
|
||||
## CURRENT-MILESTONE SOURCE
|
||||
|
||||
One file, and it changes as development progresses:
|
||||
|
||||
```text
|
||||
planning/reports/M3-IMPLEMENTATION-REPORT.md
|
||||
```
|
||||
|
||||
M3 is the most recently completed milestone, and M4 is the next to be briefed.
|
||||
This report is M3's review *and* its primary evidence record — no separate M3
|
||||
baseline report was produced — so it is the only place some of what M3 left
|
||||
behind is written down, including the browser smoke test M3 still owes.
|
||||
|
||||
**Replace it, do not accumulate.** When M4's report lands, remove this one from
|
||||
the project Sources and upload M4's instead. The repository does the same thing:
|
||||
`planning/reports/` holds the current milestone's report and
|
||||
`planning/archive/milestone-reports/` holds the rest.
|
||||
|
||||
## DO NOT UPLOAD AS PROJECT SOURCES
|
||||
|
||||
Not because these are worthless — because a source is read as current fact, and
|
||||
these are not current.
|
||||
|
||||
- **Everything under `planning/archive/`.** The Phase 0 research reports, the
|
||||
Phase 0B recommendation and spikes, the completed M1 and M2 milestone reports,
|
||||
and archived ADR 008. Every conclusion they reached that still matters has
|
||||
already been applied to the active documents; what is left is superseded
|
||||
reasoning that will contradict the current package if uploaded beside it.
|
||||
- **Older milestone reports**, once their successor exists. Uploading M1, M2 and
|
||||
M3 together produces three descriptions of the same subsystem at three
|
||||
different stages.
|
||||
- **Source code.** `backend/`, `frontend/`, tests, migrations. The planning
|
||||
package describes the system; the code is read in the repository, where it can
|
||||
be searched and run.
|
||||
- **Inherited upstream documentation.** Upstream AI-DnD's `plan/` build log and
|
||||
`docs/` guides, generated HTML and screenshots were deleted from this
|
||||
repository on 2026-09-03 for exactly this reason: they describe a hosted,
|
||||
scripted, multi-user product this fork removed. Do not re-upload them from
|
||||
upstream or from Git history.
|
||||
- **Generated HTML, images and screenshots.** They add tokens, not facts.
|
||||
- **`repo-inventory.txt` or any other local scratch output.** Not tracked, not
|
||||
authoritative, stale the moment it is written.
|
||||
@@ -1,21 +1,151 @@
|
||||
# Adventure Storyteller Planning Package
|
||||
# Adventure Storyteller — Planning Package
|
||||
|
||||
**Status:** Phase 0 complete; architecture selected; **Milestones M1, M2 and M3 implemented and accepted (M3: 2026-09-03)**.
|
||||
**Production coding:** Underway, milestone by milestone. M1, M2 and M3 are done; M4 is the next milestone to brief.
|
||||
**This file is the index. Start here.**
|
||||
|
||||
This package contains the current product requirements, final Phase 0 architecture decisions, detailed subsystem designs, acceptance tests, research evidence, and the production milestone plan for the local-only interactive-story project.
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
milestones **M1, M2 and M3 implemented and accepted** (M3: 2026-09-03).
|
||||
**Next:** **M4 — named Save Points.** Its brief has not been written yet, and
|
||||
writing it is the current action.
|
||||
|
||||
## Current Decision
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
and why.
|
||||
|
||||
Phase 0A static research and Phase 0B local validation are complete.
|
||||
## Where documentation lives
|
||||
|
||||
The production starting point is:
|
||||
```text
|
||||
planning/
|
||||
active specifications and implementation planning — authoritative
|
||||
|
||||
planning/DECISIONS/
|
||||
active architectural decisions (ADRs) — authoritative
|
||||
|
||||
planning/reports/
|
||||
the current milestone's implementation report — evidence, not instruction
|
||||
|
||||
planning/archive/
|
||||
historical evidence and completed planning material;
|
||||
not authoritative for current implementation
|
||||
```
|
||||
|
||||
You do not need to open `planning/archive/` to do ordinary milestone work. Go
|
||||
there only when an active document sends you for a specific piece of historical
|
||||
evidence. If an archived document contradicts an active one, the active one is
|
||||
right.
|
||||
|
||||
Outside `planning/`: `README.md` describes the application, `DEVELOPMENT.md` is
|
||||
setup and local operation, and `PROVENANCE.md` records what came from upstream
|
||||
AI-DnD and what each milestone changed.
|
||||
|
||||
## Which document controls
|
||||
|
||||
When two documents appear to conflict, the one higher in this list wins:
|
||||
|
||||
1. `SPECIFICATION.md` — product requirements and required behavior.
|
||||
2. The detailed behavior/design documents:
|
||||
`STORY-BRANCH-SEMANTICS.md`, `CONTEXT-AND-MEMORY.md`,
|
||||
`IMPORTED-KNOWLEDGE-DESIGN.md`, `SECURITY-THREAT-MODEL.md`,
|
||||
`MEDIA-EXTENSION-CONTRACT.md`, `BROWSER-UX-SPEC.md`, `DATA-MODEL.md`.
|
||||
3. `V1-ACCEPTANCE-TESTS.md` — the observable pass/fail contract.
|
||||
4. `TECHNICAL-DESIGN.md` — the selected implementation architecture.
|
||||
5. The ADRs in `DECISIONS/`.
|
||||
6. `BUILD-MILESTONES.md` — implementation sequence. It sequences work; it does
|
||||
not override product behavior.
|
||||
7. The current milestone report in `reports/`, then `planning/archive/` —
|
||||
evidence and historical findings, authoritative over nothing.
|
||||
|
||||
Two standing qualifications:
|
||||
|
||||
- Candidate repositories and research reports are **not** specifications. In
|
||||
particular, ai-adventure is an implementation reference for selected patterns;
|
||||
its behavior does not override this package.
|
||||
- Where an ADR records the architecture chosen to implement a requirement stated
|
||||
elsewhere, both stand: ADR 005 states the history requirement and ADR 012
|
||||
states the architecture that implements it.
|
||||
|
||||
## The active documents
|
||||
|
||||
| Document | What it is for |
|
||||
| --- | --- |
|
||||
| `SPECIFICATION.md` | What the product must do. The top of the authority order. |
|
||||
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M3 built, recorded as fact. |
|
||||
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the export shape. |
|
||||
| `STORY-BRANCH-SEMANTICS.md` | Undo/Redo/Retry/branch/take behavior, including the M3 ratifications. |
|
||||
| `CONTEXT-AND-MEMORY.md` | Prompt assembly, summarization, branch-safe memory. |
|
||||
| `IMPORTED-KNOWLEDGE-DESIGN.md` | Canon / Reference / Inspiration knowledge as a first-class subsystem. |
|
||||
| `SECURITY-THREAT-MODEL.md` | The trust boundary, and the inference endpoint policy as implemented. |
|
||||
| `MEDIA-EXTENSION-CONTRACT.md` | The contract future image/video/audio/TTS/STT work must fit. |
|
||||
| `BROWSER-UX-SPEC.md` | The browser surface, and what is deliberately not in it. |
|
||||
| `BUILD-MILESTONES.md` | M1-M11, what each delivers, what is done, and the notes each milestone leaves its successors. |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | The pass/fail contract v1 is measured against. |
|
||||
| `TEST-CAMPAIGN-FIXTURE.md` | The standard campaign the acceptance tests are run on. |
|
||||
| `VERSION.md` | Package revision history: what each closeout changed. |
|
||||
| `PROJECT-SOURCES.md` | Which files to upload as ChatGPT Project Sources. |
|
||||
|
||||
## Reading order for a new coding agent
|
||||
|
||||
1. This file.
|
||||
2. `SPECIFICATION.md`
|
||||
3. `TECHNICAL-DESIGN.md`
|
||||
4. `BUILD-MILESTONES.md` — find the milestone you are being asked to do.
|
||||
5. `STORY-BRANCH-SEMANTICS.md`
|
||||
6. `DATA-MODEL.md`
|
||||
7. `CONTEXT-AND-MEMORY.md`
|
||||
8. `IMPORTED-KNOWLEDGE-DESIGN.md`
|
||||
9. `SECURITY-THREAT-MODEL.md`
|
||||
10. `BROWSER-UX-SPEC.md`
|
||||
11. `V1-ACCEPTANCE-TESTS.md`
|
||||
12. `DECISIONS/` — all of them; they are short.
|
||||
13. `reports/M3-IMPLEMENTATION-REPORT.md`, for what the last milestone actually
|
||||
left behind. Nothing in `planning/archive/` unless sent there.
|
||||
|
||||
## Architectural decisions
|
||||
|
||||
Active ADRs, all of which still constrain current or future work:
|
||||
|
||||
| ADR | Decision |
|
||||
| --- | --- |
|
||||
| `001-browser-first.md` | The UI is a browser application. |
|
||||
| `002-ollama-only-v1.md` | Ollama-compatible local inference only; trusted-LAN HTTPS with a private CA is verified, never bypassed. |
|
||||
| `003-authoritative-local-state.md` | The application owns authoritative state; the model does not. |
|
||||
| `004-local-only-production.md` | Local-only runtime; no cloud, telemetry or runtime remote assets. |
|
||||
| `005-branch-preserving-history.md` | History is preserved, not overwritten. **The requirement.** |
|
||||
| `006-genre-agnostic-core.md` | The core state model is genre-agnostic. |
|
||||
| `007-future-media-extension.md` | Media generation stays optional and decoupled. |
|
||||
| `009-ai-dnd-production-base.md` | AI-DnD at `d72f7c1b…` is the production base. |
|
||||
| `010-explicit-typed-narrative-state-events.md` | Explicit typed events / absolute assignments, not relative deltas. |
|
||||
| `011-local-inference-endpoint-policy.md` | Address allowlist, deny by default, checked twice, TLS mandatory. |
|
||||
| `012-active-head-non-destructive-history.md` | The stored active head. **The architecture implementing ADR 005.** |
|
||||
|
||||
`008-phase0-before-build-plan.md` was a process gate — do not begin production
|
||||
work until Phase 0 closes — which Phase 0 satisfied on 2026-09-01. It is in
|
||||
`archive/decisions/` and constrains nothing now. ADR numbering continues from
|
||||
012; 008 is not reused.
|
||||
|
||||
## Milestone reports
|
||||
|
||||
`reports/` holds the report for the milestone most recently completed, because
|
||||
that is the one the next milestone's planning has to consult:
|
||||
|
||||
- `reports/M3-IMPLEMENTATION-REPORT.md` — M3's review **and** its primary
|
||||
evidence record; no separate M3 baseline report was produced. M4 needs its
|
||||
§W (closeout) and §M/§W.4 (the open browser smoke test).
|
||||
|
||||
Completed earlier milestones are in `archive/milestone-reports/`. When M4's
|
||||
report lands, M3's moves there too: a milestone report is useful during the
|
||||
immediate next milestone and historical afterwards.
|
||||
|
||||
## The decision this package rests on
|
||||
|
||||
Phase 0A static research and Phase 0B local validation closed on 2026-09-01 with
|
||||
one decision:
|
||||
|
||||
> **Fork AI-DnD at upstream commit `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`.**
|
||||
|
||||
The selection is based on measured Phase 0B behavior, not feature count. AI-DnD already contains the highest-value structural machinery: browser UI, FastAPI service boundary, SQLite persistence, parent-linked story history, alternate takes, branch-aware state snapshots, local Ollama operation, branch-scoped memory, prompt/context inspection, streaming, export/import, and a substantial automated test suite.
|
||||
|
||||
The selected composition of ideas is:
|
||||
It was chosen on measured behavior, not feature count: AI-DnD already had the
|
||||
browser UI, the FastAPI service boundary, SQLite persistence, parent-linked
|
||||
story history, alternate takes, branch-aware state snapshots, local Ollama
|
||||
operation, branch-scoped memory, prompt inspection, streaming, export/import and
|
||||
a substantial test suite. The composition of ideas selected around it:
|
||||
|
||||
```text
|
||||
AI-DnD production base
|
||||
@@ -26,45 +156,20 @@ AI-DnD production base
|
||||
+ Gamentic provider-neutral media concepts
|
||||
```
|
||||
|
||||
This is **not** a repository merge. AI-DnD is the ownership center. Other projects are implementation references only unless a later milestone explicitly reimplements a compatible idea.
|
||||
This is **not** a repository merge. AI-DnD is the ownership center; the others
|
||||
are implementation references only.
|
||||
|
||||
## Phase 0B Findings That Changed the Plan
|
||||
Phase 0B also corrected several Phase 0A assumptions, and those corrections are
|
||||
now built into the active documents rather than needing to be read from the
|
||||
research: shipped Undo was destructive with no Redo (ADR 012 replaced it);
|
||||
relative-delta world state can be semantically wrong while syntactically valid
|
||||
(ADR 010); `tiktoken` and Google Fonts broke offline operation (fixed in M1);
|
||||
export had to carry the head position (M3); Story Cards were not a sufficient
|
||||
imported-knowledge store (`IMPORTED-KNOWLEDGE-DESIGN.md`).
|
||||
|
||||
Phase 0B confirmed the fork choice while correcting several Phase 0A assumptions:
|
||||
|
||||
- AI-DnD's shipped Undo was destructive and had no Redo.
|
||||
- A disposable spike proved non-destructive head-cursor Undo/Redo in three backend files while preserving branch-scoped memory isolation.
|
||||
- A new continuation written after Undo can fork from the moved-back head while retaining the abandoned future.
|
||||
- AI-DnD's current relative-delta world-state protocol can produce semantically wrong state under realistic context even when the proposal is syntactically valid.
|
||||
- Production narrative state will therefore use explicit typed events/absolute assignments inspired by ai-adventure rather than AI-DnD's relative-delta protocol.
|
||||
- AI-DnD requires offline hardening: `tiktoken` attempts a first-use CDN fetch and the SPA requests Google Fonts at runtime.
|
||||
- AI-DnD's export format must preserve the active head position; otherwise export/import can silently redo an undone story.
|
||||
- AI-DnD Story Cards are not a sufficient imported-knowledge store because they are not designed for the required classification, provenance, chunking, and lineage semantics.
|
||||
- Open Dungeon remains useful for UX/media ideas but is no longer a serious production-fork candidate.
|
||||
- ai-adventure is not the production base but is the strongest implementation reference for authoritative typed state events, head movement, checkpoints, replay, and narrow local-only behavior.
|
||||
|
||||
See `PHASE-0B-RECOMMENDATION.md` for the coding agent's evidence. That report is retained as research evidence; the planning documents in this package record the decisions made after reviewing it.
|
||||
|
||||
## Document Authority
|
||||
|
||||
Use the documents in this order when requirements appear to conflict:
|
||||
|
||||
1. `SPECIFICATION.md` — product requirements and required behavior.
|
||||
2. Detailed behavior/design documents:
|
||||
- `STORY-BRANCH-SEMANTICS.md`
|
||||
- `CONTEXT-AND-MEMORY.md`
|
||||
- `IMPORTED-KNOWLEDGE-DESIGN.md`
|
||||
- `SECURITY-THREAT-MODEL.md`
|
||||
- `MEDIA-EXTENSION-CONTRACT.md`
|
||||
- `BROWSER-UX-SPEC.md`
|
||||
- `DATA-MODEL.md`
|
||||
3. `V1-ACCEPTANCE-TESTS.md` — observable pass/fail contract.
|
||||
4. `TECHNICAL-DESIGN.md` — selected implementation architecture.
|
||||
5. Foundational ADRs (`001-...md` through the current ADR set).
|
||||
6. `BUILD-MILESTONES.md` — implementation sequence; it does not override product behavior.
|
||||
7. Phase 0 research reports — evidence and historical findings.
|
||||
|
||||
Candidate repositories and research reports are **not** specifications. In particular, ai-adventure is an implementation reference for selected patterns; its behavior does not override this package.
|
||||
The evidence is in `archive/phase0/`, and `archive/phase0/PHASE-0B-RECOMMENDATION.md`
|
||||
is the strongest single document there. It is research evidence, not a
|
||||
specification, and it was not rewritten to match later conclusions.
|
||||
|
||||
## Foundational Decisions
|
||||
|
||||
@@ -95,29 +200,10 @@ The following are settled for v1:
|
||||
- LAN inference is distinct from LAN exposure of the storyteller UI/API; the latter is not required for v1,
|
||||
- future local image/video/audio/TTS/STT support remains optional and decoupled from the story engine.
|
||||
|
||||
## Phase 0 Status
|
||||
|
||||
### Complete
|
||||
## Still open, and deliberately so
|
||||
|
||||
- candidate discovery and triage,
|
||||
- static architecture/privacy/licensing review,
|
||||
- local clone/build/test validation,
|
||||
- real Ollama testing,
|
||||
- offline/network observation,
|
||||
- AI-DnD strip-down/entanglement checks,
|
||||
- Open Dungeon history-retrofit analysis,
|
||||
- ai-adventure Ollama/service-boundary checks,
|
||||
- AI-DnD non-destructive Undo/Redo spike,
|
||||
- referee/state-protocol follow-up,
|
||||
- export/import head-position follow-up,
|
||||
- Story Card lineage review,
|
||||
- Postgres removability review,
|
||||
- production fork decision,
|
||||
- production architecture decision.
|
||||
|
||||
### Deferred to implementation/release validation
|
||||
|
||||
These do not block the architecture decision:
|
||||
These do not block the architecture and are not v1 requirements:
|
||||
|
||||
- comparative recommendation of narrator/state models for real users,
|
||||
- multi-hour/100-turn long-run behavior,
|
||||
@@ -125,78 +211,56 @@ These do not block the architecture decision:
|
||||
- actual future image/video/TTS/STT provider integration,
|
||||
- abandoned-history cleanup UI/policy (not required in v1).
|
||||
|
||||
## Recommended Reading Order for the Next Implementation Stage
|
||||
|
||||
Do not convert this into a coding prompt until the package review is approved.
|
||||
|
||||
When implementation planning resumes, read:
|
||||
|
||||
1. `SPECIFICATION.md`
|
||||
2. `TECHNICAL-DESIGN.md`
|
||||
3. `BUILD-MILESTONES.md`
|
||||
4. `STORY-BRANCH-SEMANTICS.md`
|
||||
5. `DATA-MODEL.md`
|
||||
6. `CONTEXT-AND-MEMORY.md`
|
||||
7. `IMPORTED-KNOWLEDGE-DESIGN.md`
|
||||
8. `SECURITY-THREAT-MODEL.md`
|
||||
9. `BROWSER-UX-SPEC.md`
|
||||
10. `V1-ACCEPTANCE-TESTS.md`
|
||||
11. ADRs, especially the production-base and narrative-state-event decisions
|
||||
12. Phase 0B reports only as supporting evidence
|
||||
|
||||
## Workflow From Here
|
||||
## Where the work stands
|
||||
|
||||
```text
|
||||
Phase 0 research and spikes COMPLETE
|
||||
Architecture/fork decision COMPLETE
|
||||
Planning package approved COMPLETE
|
||||
|
|
||||
v
|
||||
Architecture/fork decision COMPLETE
|
||||
Milestone M1 COMPLETE (2026-09-02)
|
||||
fork + offline baseline archive/milestone-reports/M1-*.md
|
||||
|
|
||||
v
|
||||
Planning package revision COMPLETE
|
||||
|
|
||||
v
|
||||
Approve planning package COMPLETE
|
||||
|
|
||||
v
|
||||
Milestone M1 COMPLETE (2026-09-02)
|
||||
fork + offline baseline see planning/reports/M1-*.md
|
||||
|
|
||||
v
|
||||
Milestone M2 COMPLETE (2026-09-02)
|
||||
local-only surface + endpoint see planning/reports/M2-*.md
|
||||
Milestone M2 COMPLETE (2026-09-02)
|
||||
local-only surface + endpoint archive/milestone-reports/M2-*.md
|
||||
policy
|
||||
|
|
||||
v
|
||||
Milestone M3 COMPLETE (2026-09-03)
|
||||
non-destructive undo/redo, see planning/reports/M3-*.md and ADR 012
|
||||
active-head export one open condition: the browser smoke test
|
||||
Milestone M3 COMPLETE (2026-09-03)
|
||||
non-destructive undo/redo, reports/M3-IMPLEMENTATION-REPORT.md
|
||||
active-head export and ADR 012
|
||||
|
|
||||
v
|
||||
Milestone M4 NEXT — brief not yet prepared
|
||||
Milestone M4 NEXT — brief not yet prepared
|
||||
named Save Points
|
||||
|
|
||||
v
|
||||
Implement and review milestone-by-milestone
|
||||
M5-M11, one at a time see BUILD-MILESTONES.md
|
||||
```
|
||||
|
||||
## Stop Rule
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
|
||||
M1, M2 and M3 are complete and accepted; the evidence is in `reports/M1-*.md`,
|
||||
`reports/M2-*.md` and `reports/M3-IMPLEMENTATION-REPORT.md` — the last of which
|
||||
is M3's primary evidence record as well as its review, since no separate M3
|
||||
baseline report was produced. **No M4 brief has been prepared.** The current
|
||||
action is to write one, informed by the post-M3 corrections below, by the note
|
||||
`BUILD-MILESTONES.md` now attaches to M4, and by **ADR 012**, which records the
|
||||
head-movement mechanism M4 must reuse rather than reimplement.
|
||||
**No M4 brief has been prepared.** Writing one is the current action, informed
|
||||
by the post-M3 corrections below, by the note `BUILD-MILESTONES.md` attaches to
|
||||
M4, and by **ADR 012**, which records the head-movement mechanism M4 must reuse
|
||||
rather than reimplement.
|
||||
|
||||
One M3 condition remains open and is not a blocker for M4: the required
|
||||
**browser smoke test has not been performed**, because no session in which M3
|
||||
was implemented or reviewed had a browser available. See
|
||||
One M3 condition remains open and does not block M4: the required **browser
|
||||
smoke test has not been performed**, because no session in which M3 was
|
||||
implemented or reviewed had a browser available. See
|
||||
`reports/M3-IMPLEMENTATION-REPORT.md` §M and §W.4.
|
||||
|
||||
## What each milestone closeout corrected
|
||||
|
||||
These tables are the audit trail: what implementation evidence forced back into
|
||||
the planning package, milestone by milestone. `VERSION.md` narrates the same
|
||||
changes.
|
||||
|
||||
### Post-M3 corrections applied (2026-09-03)
|
||||
|
||||
M3's review recommended planning changes and, following the M2 pattern, reported
|
||||
|
||||
@@ -184,7 +184,7 @@ controls on a network they trust — and why the endpoint is always explicitly
|
||||
configured, never discovered.
|
||||
|
||||
See ADR 002 (*Transport for a Trusted-LAN Endpoint*) and, for the demonstrated
|
||||
deployment, `planning/reports/M1-IMPLEMENTATION-REPORT.md` §F.
|
||||
deployment, `planning/archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md` §F.
|
||||
|
||||
Allowed future local paths may also include explicitly configured local media services.
|
||||
|
||||
|
||||
@@ -1,8 +1,44 @@
|
||||
# Planning Package Version
|
||||
|
||||
**Package:** Adventure Storyteller Planning Package v2.3
|
||||
**Revision date:** 2026-09-03
|
||||
**Status:** Phase 0 complete; architecture selected; **Milestones M1, M2 and M3 implemented and accepted**; M4 is next to brief.
|
||||
- **Package:** Adventure Storyteller Planning Package v2.4
|
||||
- **Revision date:** 2026-09-03
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1, M2 and M3 implemented and accepted**; M4 is next to brief.
|
||||
|
||||
## v2.4 — Documentation Consolidation (2026-09-03)
|
||||
|
||||
No product requirement, architecture decision or milestone status changed in
|
||||
this revision. It reorganises the documentation so that a new coding agent can
|
||||
tell authoritative material from evidence at a glance.
|
||||
|
||||
- `planning/archive/` is created and is **non-authoritative by declaration**
|
||||
(`archive/README.md`). It holds `phase0/` — the research that chose AI-DnD —
|
||||
`milestone-reports/` — the completed M1 and M2 reports — and `decisions/`,
|
||||
which now holds ADR **008**, the Phase-0-before-build process gate that Phase 0
|
||||
satisfied. ADR numbering continues from 012; 008 is not reused.
|
||||
- `planning/reports/` now holds **only the current milestone's report**,
|
||||
`M3-IMPLEMENTATION-REPORT.md`, because M4 planning has to consult it. It moves
|
||||
to the archive when M4's report replaces it.
|
||||
- The Phase 0B execution prompts and handoff/status/summary documents
|
||||
(`CODEX-HANDOFF-NOTE.md`, `PHASE-0B-CODEX-BRIEF.md`,
|
||||
`PHASE-0B-CODEX-HANDOFF.md`, `PLANNING-UPDATE-SUMMARY.md`) and the Phase 0A
|
||||
discovery and triage reports were **deleted**: intermediate working documents
|
||||
whose conclusions all reached the two recommendation reports, and which remain
|
||||
in Git history.
|
||||
- Upstream AI-DnD's inherited `plan/` build log and `docs/` project site
|
||||
(guides, generated HTML, screenshots) were **deleted**. They documented a
|
||||
hosted, scripted, multi-user product with accounts — every screenshot showed a
|
||||
Scripts tab and a Sign up button — which M2 removed. Both trees remain in Git
|
||||
history and in upstream.
|
||||
- `README.md`, `DEVELOPMENT.md` and `PROVENANCE.md` are corrected where they
|
||||
pointed at the removed trees or described removed capability as present.
|
||||
`DEVELOPMENT.md`'s "things M1 did not touch" section had gone stale at M2 and
|
||||
now says what is actually still inherited.
|
||||
- `planning/README.md` is rewritten as **the documentation index**: current
|
||||
milestone, the three-way active/ADR/archive split, the authority order,
|
||||
reading order, where reports live, and what comes next. The milestone
|
||||
correction tables are preserved unchanged.
|
||||
- New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
|
||||
manifest of files to upload as ChatGPT Project Sources.
|
||||
|
||||
## v2.3 — Post-M3 Closeout (2026-09-03)
|
||||
|
||||
@@ -51,7 +87,7 @@ M2 removed the hosted, cloud, account and scripting surface and added the
|
||||
inference endpoint policy. Its review recommended six planning changes and
|
||||
reported rather than applied them; all six are applied in this revision, listed
|
||||
in `README.md` § *Post-M2 corrections applied*, with the evidence in
|
||||
`reports/M2-BASELINE-REPORT.md` and `reports/M2-IMPLEMENTATION-REPORT.md`.
|
||||
`archive/milestone-reports/M2-BASELINE-REPORT.md` and `archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md`.
|
||||
|
||||
In summary:
|
||||
|
||||
@@ -84,7 +120,7 @@ the architecture selected in v2 was reversed.
|
||||
M1 implementation evidence contradicted or under-specified parts of v2. The
|
||||
corrections are recorded in the documents themselves and listed in
|
||||
`README.md` § *Post-M1 corrections applied*; the evidence behind them is in
|
||||
`reports/M1-BASELINE-REPORT.md` and `reports/M1-IMPLEMENTATION-REPORT.md`.
|
||||
`archive/milestone-reports/M1-BASELINE-REPORT.md` and `archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md`.
|
||||
|
||||
In summary:
|
||||
|
||||
|
||||
@@ -0,0 +1,67 @@
|
||||
# Archive — historical material, not authoritative
|
||||
|
||||
Everything under `planning/archive/` is **evidence and history**. None of it
|
||||
governs current implementation. If an archived document and an active planning
|
||||
document disagree, the active document is right and the archived one records
|
||||
what was believed or measured at the time.
|
||||
|
||||
Do not consult this directory during ordinary milestone work unless an active
|
||||
document sends you here for a specific piece of historical evidence.
|
||||
|
||||
## What is here
|
||||
|
||||
### `phase0/` — why AI-DnD was selected
|
||||
|
||||
Phase 0A static research and Phase 0B local validation, closed 2026-09-01.
|
||||
|
||||
| File | What it is |
|
||||
| --- | --- |
|
||||
| `RESEARCH-PLAN.md` | The question Phase 0 existed to answer, and its final dispositions. |
|
||||
| `PRELIMINARY-RECOMMENDATION.md` | The Phase 0A conclusion, from static review alone. |
|
||||
| `REUSE-MATRIX.md` | What each candidate offered against the required subsystems. |
|
||||
| `AI-DND-ANALYSIS.md` | The selected base, analysed before the fork. |
|
||||
| `AI-ADVENTURE-ANALYSIS.md` | The rejected finalist still cited as the implementation reference for typed state events. |
|
||||
| `OPEN-DUNGEON-ANALYSIS.md` | The rejected finalist still cited as the UX/future-media reference. |
|
||||
| `PHASE-0B-RECOMMENDATION.md` | **The strongest single document.** The measured Phase 0B findings and the fork decision. |
|
||||
| `PHASE-0B-BASELINE.md` | Clean clones, builds and test runs for the three finalists. |
|
||||
| `PHASE-0B-AI-DND-EXPERIMENT.md` | AI-DnD driven against a real local Ollama. |
|
||||
| `PHASE-0B-AI-ADVENTURE-OLLAMA.md` | ai-adventure's Ollama and service-boundary behaviour. |
|
||||
| `PHASE-0B-OPEN-DUNGEON-HISTORY.md` | Whether Open Dungeon could be retrofitted with history. |
|
||||
| `PHASE-0B-OFFLINE-NETWORK.md` | What each candidate reached for with no route to the Internet. |
|
||||
| `PHASE-0B-UNDO-SPIKE.md` | The disposable spike that proved non-destructive undo/redo, and became ADR 012's architecture. |
|
||||
| `PHASE-0B-FOLLOWUP-CHECKS.md` | The world-state protocol, export/head, story-card lineage and Postgres checks. |
|
||||
|
||||
Phase 0A discovery and triage material (the candidate inventory, source index,
|
||||
reference-project list, licensing and static-privacy reviews, the Phase 0A
|
||||
status page) and the Phase 0B execution prompts were deleted in the 2026-09-03
|
||||
documentation cleanup. They are intermediate working documents whose
|
||||
conclusions all reached the two recommendation reports above, and they remain in
|
||||
Git history.
|
||||
|
||||
### `milestone-reports/` — completed milestone evidence
|
||||
|
||||
`M1-BASELINE-REPORT.md`, `M1-IMPLEMENTATION-REPORT.md`,
|
||||
`M2-BASELINE-REPORT.md`, `M2-IMPLEMENTATION-REPORT.md`.
|
||||
|
||||
Every architectural conclusion these reports reached has already been applied to
|
||||
the active planning documents and the ADRs — see `planning/VERSION.md`, which
|
||||
lists the corrections each milestone produced. The reports are kept for their
|
||||
measurements and their reasoning, not as instructions.
|
||||
|
||||
The **current** milestone's report stays in `planning/reports/` while it is
|
||||
still useful for reviewing the next milestone, and moves here when it is not.
|
||||
|
||||
### `decisions/` — superseded or completed ADRs
|
||||
|
||||
`008-phase0-before-build-plan.md` — a process gate ("do not start production
|
||||
work before Phase 0 closes") that Phase 0 satisfied on 2026-09-01. It
|
||||
constrains nothing now. ADR numbering continues from 012 in
|
||||
`planning/DECISIONS/`; 008 is not reused.
|
||||
|
||||
## A note on paths inside these files
|
||||
|
||||
Archived documents are kept **verbatim**. File paths written inside them refer
|
||||
to where those files lived when the document was written — before this archive
|
||||
existed, and in some cases before files were deleted. That is deliberate: an
|
||||
evidence record that has been quietly edited is no longer evidence. Resolve any
|
||||
such path against Git history, not against the current tree.
|
||||
@@ -0,0 +1,29 @@
|
||||
README.md
|
||||
DEVELOPMENT.md
|
||||
PROVENANCE.md
|
||||
planning/README.md
|
||||
planning/VERSION.md
|
||||
planning/SPECIFICATION.md
|
||||
planning/TECHNICAL-DESIGN.md
|
||||
planning/DATA-MODEL.md
|
||||
planning/STORY-BRANCH-SEMANTICS.md
|
||||
planning/CONTEXT-AND-MEMORY.md
|
||||
planning/IMPORTED-KNOWLEDGE-DESIGN.md
|
||||
planning/SECURITY-THREAT-MODEL.md
|
||||
planning/MEDIA-EXTENSION-CONTRACT.md
|
||||
planning/BROWSER-UX-SPEC.md
|
||||
planning/BUILD-MILESTONES.md
|
||||
planning/V1-ACCEPTANCE-TESTS.md
|
||||
planning/TEST-CAMPAIGN-FIXTURE.md
|
||||
planning/DECISIONS/001-browser-first.md
|
||||
planning/DECISIONS/002-ollama-only-v1.md
|
||||
planning/DECISIONS/003-authoritative-local-state.md
|
||||
planning/DECISIONS/004-local-only-production.md
|
||||
planning/DECISIONS/005-branch-preserving-history.md
|
||||
planning/DECISIONS/006-genre-agnostic-core.md
|
||||
planning/DECISIONS/007-future-media-extension.md
|
||||
planning/DECISIONS/009-ai-dnd-production-base.md
|
||||
planning/DECISIONS/010-explicit-typed-narrative-state-events.md
|
||||
planning/DECISIONS/011-local-inference-endpoint-policy.md
|
||||
planning/DECISIONS/012-active-head-non-destructive-history.md
|
||||
planning/reports/M3-IMPLEMENTATION-REPORT.md
|
||||
@@ -1,46 +0,0 @@
|
||||
# aiMultiFool — Static Architecture Analysis
|
||||
|
||||
**Repository:** https://github.com/omgboohoo/aimultifool
|
||||
**Date reviewed:** 2026-09-01
|
||||
**Disposition:** Reference only.
|
||||
|
||||
## Useful ideas
|
||||
|
||||
aiMultiFool is a local roleplay/chat sandbox with:
|
||||
- Ollama support,
|
||||
- local inference paths,
|
||||
- Vector Chat / semantic-memory concepts,
|
||||
- save/load,
|
||||
- rewind/regenerate,
|
||||
- context inspection,
|
||||
- optional encrypted local data.
|
||||
|
||||
Those are useful implementation references for local memory tooling and diagnostics.
|
||||
|
||||
## Why it is not a fork finalist
|
||||
|
||||
### Product mismatch
|
||||
The interface is terminal/Textual-oriented and character-chat/roleplay focused rather than a browser-first persistent fiction editor.
|
||||
|
||||
### History/context mismatch
|
||||
Its documented smart-pruning strategy removes older middle messages from active chat state as context pressure grows. That is a reasonable chat optimization but not the target architecture. The target must preserve an immutable authoritative transcript and prune only the prompt representation.
|
||||
|
||||
### State model
|
||||
The project does not provide the same authoritative event/state/branch model found in AI-DnD or ai-adventure.
|
||||
|
||||
### License
|
||||
The repository is GPL-3.0. Directly copying substantial GPL code into an MIT/Apache-derived application would change licensing obligations. Unless the final project intentionally adopts GPL-compatible distribution terms, use this project for concepts rather than source copying.
|
||||
|
||||
## Recommended reuse
|
||||
|
||||
Study:
|
||||
- local embedding workflow,
|
||||
- vector inspection/debugging,
|
||||
- encrypted local payload design,
|
||||
- user-facing memory controls.
|
||||
|
||||
Do not make it a Phase 0B build finalist.
|
||||
|
||||
## Source
|
||||
|
||||
- Repository: https://github.com/omgboohoo/aimultifool
|
||||
@@ -1,68 +0,0 @@
|
||||
# Candidate Inventory and Triage
|
||||
|
||||
**Historical status:** Phase 0A triage. The final production disposition was selected after Phase 0B; see ADR 009.
|
||||
|
||||
**Phase:** 0A — Static research
|
||||
**Date:** 2026-09-01
|
||||
|
||||
## Executive result
|
||||
|
||||
Three projects should advance to local validation:
|
||||
|
||||
1. **AI-DnD** — strongest implementation of the hardest required backend capabilities.
|
||||
2. **Open Dungeon** — strongest direct product/UX fit and strongest near-term media path.
|
||||
3. **CaoRuiming/ai-adventure (Local Adventure Engine)** — strongest authoritative-state, replay, checkpoint, and privacy architecture.
|
||||
|
||||
Everything else should remain available as a design/source reference but should not consume local build-validation effort unless one of the three finalists fails.
|
||||
|
||||
## Triage table
|
||||
|
||||
| Project | Browser-first | Ollama | Durable branch/rollback | Long memory | Local knowledge | Future media | Static disposition |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| AI-DnD | Yes | Yes | **Strong** | **Strong** | Story cards + memory | Not core | **Finalist #1** |
|
||||
| Open Dungeon | **Yes** | **Yes** | Weak / destructive linear tail today | Summary-based | Limited | **Strong; local image generation already present** | **Finalist #2** |
|
||||
| ai-adventure | No; CLI | LM Studio today | **Strong** | Summary + lore FTS | **Strong deterministic local lore** | No | **Finalist #3** |
|
||||
| Chronicler | Yes | Yes | Not the focus | **Excellent memory model** | Memory-centric | No | Reference |
|
||||
| Gamentic | **Yes** | llama.cpp/OpenAI-compatible | Game-state oriented | Strong | World bible | **Excellent image/voice provider design** | Reference |
|
||||
| Interactive Fiction Framework | **Yes** | **Yes** | Not established as required story-tree model | Canon/scene/character memory | Story Bible | Not core | Reference |
|
||||
| Sonder Engine | **Yes** | **Yes** | Persistent variants/checkpoints, but much more agentic | **Very sophisticated** | Character-scoped retrieval | Not primary | Reference |
|
||||
| Corvus Story Core | **Yes** | OpenAI-compatible | No equivalent branch tree established | Summaries + state | World/state | ComfyUI + TTS | Reference |
|
||||
| aiMultiFool | No; terminal | **Yes** | Rewind, not target architecture | Vector chat | RAG-oriented | No | Reference only |
|
||||
| SillyTavern | **Yes** | Local backends | Chat-oriented | Extensions/lorebooks | **Excellent lorebook UX** | Broad extensions | Reference only |
|
||||
| RisuAI | **Yes** | Local/remote ecosystem | Chat-oriented | Hypa/SupaMemory | Lorebooks | Broad media | Reference only |
|
||||
| KoboldAI | **Yes** | Local ecosystem | Traditional save/load | Memory/World Info | World Info | Limited | Reference only |
|
||||
|
||||
## Why the shortlist is only three
|
||||
|
||||
### AI-DnD advances because
|
||||
It already implements the expensive correctness work: a parent/lineage story tree, alternate takes, non-destructive retry, state snapshots and rollback, branch-aware context, summaries, embedding retrieval, story cards, exact prompt inspection, export/import of the complete tree, and a substantial automated test suite.
|
||||
|
||||
### Open Dungeon advances because
|
||||
It is almost exactly the desired product shape: browser-first, simple interactive fiction, Ollama, local SQLite, streaming narration, visual character continuity, and local image-generation hooks. Its key weakness is architectural rather than cosmetic: its current message schema is linear and its retry/erase operation deletes the selected message and the rest of the tail.
|
||||
|
||||
### ai-adventure advances because
|
||||
Its core philosophy most closely matches the required trust model. SQLite and typed events are authoritative; the model proposes changes; validation occurs before atomic commit; undo/checkpoint/restore/branch work by replaying parent-linked history; lore is local; and the privacy documentation explicitly minimizes network and executable-extension surfaces.
|
||||
|
||||
## Projects eliminated from fork contention
|
||||
|
||||
### Chronicler
|
||||
Excellent source for memory semantics, but the application is centered on long-running roleplay and YantrikDB cognitive memory rather than the simpler interactive-story product. Its memory-tier design should be borrowed conceptually.
|
||||
|
||||
### Gamentic
|
||||
Technically impressive and very useful for future media design, but it is intentionally a multi-agent RPG with image and voice infrastructure, tuned around a heavier local stack. Forking it would mean removing more game/agent behavior than necessary.
|
||||
|
||||
### Interactive Fiction Framework
|
||||
Its Story Bible, validation, and application-owned-state design are highly relevant. However, it is oriented toward contributor-authored, schema-driven stories and planner-approved choices rather than the unrestricted natural-language story continuation and branch history required here.
|
||||
|
||||
### Sonder Engine
|
||||
Strong engineering, but its core differentiator is separate fictional minds with strict perception/knowledge boundaries and a multi-stage agent pipeline. That is substantially more complexity than v1 requires.
|
||||
|
||||
### Corvus Story Core
|
||||
Useful image/TTS and state-extraction reference, but its persistence is JSON/JSONL-oriented and the static review did not establish the required non-destructive branch/checkpoint model.
|
||||
|
||||
### aiMultiFool
|
||||
Useful local vector-memory ideas, but it is a terminal character-roleplay application, its context-pruning approach is not the desired immutable-history architecture, and GPL-3.0 complicates direct code reuse into a permissively licensed fork.
|
||||
|
||||
## Sources
|
||||
|
||||
See `SOURCE-INDEX.md` for repository/source links.
|
||||
@@ -1,75 +0,0 @@
|
||||
# Licensing and Reuse Review
|
||||
|
||||
**Date:** 2026-09-01
|
||||
**Nature:** Engineering planning summary, not legal advice.
|
||||
|
||||
## Production Selection Update
|
||||
|
||||
Phase 0B selected **AI-DnD (MIT)** as the production fork/base. The production strategy therefore remains on a permissive-license path.
|
||||
|
||||
The project may reimplement compatible ideas from other projects, but direct source copying must still be reviewed file-by-file and retain applicable notices. GPL/AGPL projects remain concept/reference sources unless a later explicit licensing decision changes that policy.
|
||||
|
||||
Maintain third-party notices from the first production milestone.
|
||||
|
||||
## Permissive finalists
|
||||
|
||||
### AI-DnD
|
||||
- License: MIT
|
||||
- Direct modification/forking is generally compatible with a permissive local application, subject to preserving required notices.
|
||||
|
||||
### Open Dungeon
|
||||
- License: MIT
|
||||
- Same practical advantage for direct reuse.
|
||||
|
||||
### ai-adventure
|
||||
- License: Apache-2.0
|
||||
- Permissive, but Apache notice/license obligations must be preserved.
|
||||
|
||||
These three can plausibly participate in a permissively licensed implementation strategy, subject to checking individual vendored/third-party files.
|
||||
|
||||
## Permissive reference projects
|
||||
|
||||
Static repository licensing indicates:
|
||||
- Chronicler: MIT
|
||||
- Interactive Fiction Framework: MIT
|
||||
- Gamentic: MIT
|
||||
- Sonder Engine: MIT
|
||||
- Corvus Story Core: MIT
|
||||
|
||||
If source is copied, retain the applicable notices and verify whether particular directories/files carry separate licenses.
|
||||
|
||||
## Copyleft references
|
||||
|
||||
### aiMultiFool
|
||||
- GPL-3.0
|
||||
- Treat as a concept/reference source unless the final project intentionally accepts GPL obligations.
|
||||
|
||||
### LettuceAI
|
||||
- AGPL-3.0
|
||||
- Reference only for this project unless there is a deliberate licensing decision.
|
||||
|
||||
Mature roleplay ecosystems such as SillyTavern/RisuAI/KoboldAI should have their exact current license verified before any code copying. No direct reuse is currently recommended.
|
||||
|
||||
## Media dependencies
|
||||
|
||||
Important distinction:
|
||||
- application code license,
|
||||
- media runtime license,
|
||||
- model-weight license
|
||||
are separate.
|
||||
|
||||
For example Gamentic documents:
|
||||
- its own code under MIT,
|
||||
- ComfyUI runtime under GPL-3.0,
|
||||
- model weights under their own terms.
|
||||
|
||||
Using a separately running local service through an API is architecturally different from copying its code into the storyteller, but distribution/bundling choices should be reviewed before release.
|
||||
|
||||
## Recommendation
|
||||
|
||||
Keep the production application's own code on a permissive-license path if possible:
|
||||
- primary fork from MIT or Apache-2.0,
|
||||
- copy code only from compatible permissive sources,
|
||||
- treat GPL/AGPL projects as design references unless a conscious license change is made,
|
||||
- keep optional media providers as external adapters/services where practical,
|
||||
- maintain a third-party notices file from the first production milestone.
|
||||
@@ -1,41 +0,0 @@
|
||||
# Phase 0A Status
|
||||
|
||||
**Historical status:** Phase 0A complete. Phase 0B is also complete; see README, ADR 009, ADR 010, and `TECHNICAL-DESIGN.md` for current architecture.
|
||||
|
||||
**Completed:** 2026-09-01
|
||||
|
||||
## Completed statically
|
||||
|
||||
- candidate discovery and triage,
|
||||
- deep source/document architecture review of the three finalists,
|
||||
- static privacy/network-surface review,
|
||||
- preliminary licensing/reuse review,
|
||||
- subsystem reuse matrix,
|
||||
- preliminary fork recommendation,
|
||||
- narrowed Codex validation plan.
|
||||
|
||||
## Preliminary decision
|
||||
|
||||
Validate **AI-DnD first as the production fork candidate**.
|
||||
|
||||
Keep:
|
||||
- **Open Dungeon** as the fallback fork and primary UI/media reference.
|
||||
- **ai-adventure** as the state/replay/privacy architecture reference and third validation candidate.
|
||||
|
||||
## Still requires local/Codex work
|
||||
|
||||
- pin exact SHAs,
|
||||
- clone/install/build,
|
||||
- run actual tests,
|
||||
- verify Ollama against the user's machine,
|
||||
- runtime network capture,
|
||||
- offline operation,
|
||||
- AI-DnD strip-down experiment,
|
||||
- Open Dungeon branch-retrofit impact experiment,
|
||||
- ai-adventure Ollama/service-boundary experiment.
|
||||
|
||||
See `PHASE-0B-CODEX-HANDOFF.md`.
|
||||
|
||||
## Phase gate
|
||||
|
||||
Do not finalize `TECHNICAL-DESIGN.md` v1.0 or production `BUILD-MILESTONES.md` until Phase 0B results are reviewed.
|
||||
@@ -1,122 +0,0 @@
|
||||
# Static Privacy and Network Review
|
||||
|
||||
**Date:** 2026-09-01
|
||||
**Scope:** Source/config/documentation review only. Runtime capture is still required in Phase 0B.
|
||||
|
||||
## Target rule
|
||||
|
||||
The final v1 should be able to operate with Internet access physically blocked, with ordinary story data traveling only:
|
||||
|
||||
```text
|
||||
Browser -> local application -> local Ollama
|
||||
```
|
||||
|
||||
Future media should similarly use explicitly configured local providers.
|
||||
|
||||
## AI-DnD
|
||||
|
||||
### Static positives
|
||||
- documented local single-user mode,
|
||||
- local SQLite,
|
||||
- local Ollama support,
|
||||
- no auth required in local mode,
|
||||
- hosted analytics are first-party application functionality rather than a required third-party browser tracker.
|
||||
|
||||
### Unwanted surfaces to remove
|
||||
- OpenRouter/OpenAI/Groq/vLLM provider support,
|
||||
- hosted account/guest flows,
|
||||
- demo API keys,
|
||||
- Render deployment,
|
||||
- Neon/Postgres cloud deployment path,
|
||||
- visit analytics,
|
||||
- QuickJS user scripting,
|
||||
- Claude CLI shim if not wanted,
|
||||
- any hosted-mode rate-limit/account code that adds no local value.
|
||||
|
||||
### Risk
|
||||
The cloud/hosted code is explicit and documented, which is good, but Phase 0B must prove it can be removed cleanly.
|
||||
|
||||
## Open Dungeon
|
||||
|
||||
### Static positives
|
||||
- Ollama loopback default,
|
||||
- local SQLite,
|
||||
- local image backend,
|
||||
- no telemetry requirement apparent in inspected package/config.
|
||||
|
||||
### Unwanted or optional surfaces
|
||||
- OpenRouter configuration,
|
||||
- arbitrary remote OpenAI-compatible endpoint support,
|
||||
- Tailscale/LAN exposure options,
|
||||
- any runtime remote assets,
|
||||
- any model/image automatic download behavior after setup.
|
||||
|
||||
### Risk
|
||||
The app is smaller, so hardening may be easier, but no runtime capture has been performed.
|
||||
|
||||
## ai-adventure
|
||||
|
||||
### Static positives
|
||||
This project most closely matches the target from the outset:
|
||||
- no telemetry,
|
||||
- no cloud account,
|
||||
- no MCP,
|
||||
- no executable plugins,
|
||||
- no shell tools,
|
||||
- loopback model endpoint default,
|
||||
- non-loopback warning,
|
||||
- imported content treated as bounded data,
|
||||
- path traversal/symlink defenses documented.
|
||||
|
||||
### Unwanted surface
|
||||
- configurable non-loopback model endpoint should be prohibited or strongly gated in the target v1.
|
||||
- LM Studio provider should be replaced/extended with Ollama.
|
||||
|
||||
## Reference projects
|
||||
|
||||
### Gamentic
|
||||
Local defaults are strong, but the project intentionally supports cloud text/image/audio dialects as alternatives. A target fork would need those disabled. Its Docker/media stack also has setup-time model acquisition concerns separate from story-time privacy.
|
||||
|
||||
### Chronicler
|
||||
Supports local Ollama but also broader providers and a separate local YantrikDB/MCP memory service. More moving parts than needed.
|
||||
|
||||
### Sonder / Corvus
|
||||
Both support local backends but also remote provider configurations; Sonder additionally has extension/optional external-service surfaces.
|
||||
|
||||
### aiMultiFool
|
||||
Primarily local, but direct code reuse is constrained by GPL considerations and it is not a fork finalist.
|
||||
|
||||
## Required Phase 0B runtime tests
|
||||
|
||||
For each finalist:
|
||||
|
||||
1. block outbound Internet access,
|
||||
2. start the app,
|
||||
3. create/load a story,
|
||||
4. generate multiple turns,
|
||||
5. trigger summarization/memory,
|
||||
6. trigger embeddings where applicable,
|
||||
7. save/restore/branch,
|
||||
8. for Open Dungeon, generate a local image,
|
||||
9. capture socket/DNS/HTTP activity,
|
||||
10. fail the test if story content leaves loopback or explicitly approved LAN endpoints.
|
||||
|
||||
Record:
|
||||
- process,
|
||||
- destination IP/hostname,
|
||||
- port,
|
||||
- trigger,
|
||||
- payload classification,
|
||||
- whether required or optional.
|
||||
|
||||
## Recommended production hardening
|
||||
|
||||
- bind app and Ollama to loopback by default,
|
||||
- allowlist provider URLs rather than accept arbitrary URLs,
|
||||
- no API-key UI in v1,
|
||||
- no remote URL ingestion,
|
||||
- no executable campaign scripts,
|
||||
- no third-party analytics,
|
||||
- bundle frontend assets locally,
|
||||
- content-security policy that rejects remote scripts/styles/images by default,
|
||||
- CI test or integration harness that runs with outbound networking disabled.
|
||||
@@ -1,111 +0,0 @@
|
||||
# Reference Project Findings
|
||||
|
||||
**Date:** 2026-09-01
|
||||
|
||||
These projects are not recommended as primary forks after static review, but each contributes a useful architectural pattern.
|
||||
|
||||
## Chronicler
|
||||
|
||||
Repository: https://github.com/yantrikos/chronicler
|
||||
|
||||
### Borrow
|
||||
Its memory model distinguishes different trust levels rather than treating all remembered text equally.
|
||||
|
||||
Useful conceptual tiers:
|
||||
- durable canon,
|
||||
- scene/recent memory,
|
||||
- heuristic/inferred memory.
|
||||
|
||||
Its anti-confabulation approach is especially relevant: retrieved hints should not automatically become established historical fact.
|
||||
|
||||
### Do not necessarily adopt
|
||||
The full YantrikDB/MCP cognitive-memory stack is heavier than v1 needs. Start with a simpler local store and preserve the trust-tier semantics.
|
||||
|
||||
## Interactive Fiction Framework
|
||||
|
||||
Repository: https://github.com/georgebutler/interactive-fiction-framework
|
||||
|
||||
### Borrow
|
||||
- Story Bible as highest-authority narrative context,
|
||||
- application owns durable state,
|
||||
- model enriches prose rather than overriding state,
|
||||
- structured output validation,
|
||||
- deterministic fallback,
|
||||
- separation of director/planner/memory/validator.
|
||||
|
||||
### Why not fork
|
||||
It is designed around contributor-authored story bundles and planner-approved choices, whereas the target is more freeform collaborative fiction with branch-preserving history.
|
||||
|
||||
## Gamentic
|
||||
|
||||
Repository: https://github.com/hec-ovi/gamentic
|
||||
|
||||
### Borrow
|
||||
This is the strongest reference found for future multimodal architecture.
|
||||
|
||||
It separates each modality behind a provider layer:
|
||||
|
||||
```text
|
||||
engine
|
||||
-> text provider
|
||||
-> image provider
|
||||
-> audio provider
|
||||
```
|
||||
|
||||
The game can continue text-first while images render asynchronously. Character image/voice identity lives in game state rather than in provider-specific code.
|
||||
|
||||
It also demonstrates an unusually strong local-project test strategy with over a thousand automated tests documented across backend/frontend/services.
|
||||
|
||||
### Why not fork
|
||||
The core product is a multi-agent RPG with significant game mechanics and a heavy local image/voice stack. That is broader than the desired v1 storyteller.
|
||||
|
||||
## Sonder Engine
|
||||
|
||||
Repository: https://github.com/N0819/Sonder_Engine
|
||||
|
||||
### Borrow later
|
||||
- one persistent commit boundary,
|
||||
- objective state distinct from character perception/belief/memory,
|
||||
- retrieval scoped by what a character may legitimately know,
|
||||
- model stages with different contexts.
|
||||
|
||||
### Why not fork
|
||||
Its defining feature is separate character minds and a multi-stage agent pipeline. That is valuable for a future sophisticated simulation but unnecessary complexity for v1.
|
||||
|
||||
## Corvus Story Core
|
||||
|
||||
Repository: https://github.com/JustLateNightAI/Corvus-Story-Core
|
||||
|
||||
### Borrow
|
||||
- hidden GM/state extraction pass,
|
||||
- scene/NPC visual descriptions,
|
||||
- ComfyUI scene art,
|
||||
- optional TTS,
|
||||
- local-first media integration.
|
||||
|
||||
### Why not fork
|
||||
Static review did not show the same robust branch/checkpoint/replay model; persistence is oriented around local JSON/JSONL rather than the desired transactional story graph.
|
||||
|
||||
## SillyTavern / RisuAI / KoboldAI
|
||||
|
||||
### Borrow
|
||||
- lorebook/world-info UX,
|
||||
- author's-note concepts,
|
||||
- context placement and triggering,
|
||||
- character/world metadata workflows.
|
||||
|
||||
### Why not fork
|
||||
They are mature but broad roleplay/chat ecosystems. Adapting them would mean carrying a large amount of unrelated general-purpose functionality.
|
||||
|
||||
## Design consequence
|
||||
|
||||
The production fork should not try to merge these projects.
|
||||
|
||||
Use a primary codebase, then deliberately implement selected patterns:
|
||||
|
||||
- AI-DnD: story tree, rollback, memory, prompt inspection.
|
||||
- ai-adventure: authoritative event/replay/privacy discipline.
|
||||
- Open Dungeon: story-focused UX and visual continuity.
|
||||
- Chronicler: memory trust tiers.
|
||||
- Gamentic: provider-neutral/asynchronous media.
|
||||
- IFF: Story Bible authority and validation.
|
||||
@@ -1,83 +0,0 @@
|
||||
# Phase 0 Source Index
|
||||
|
||||
**Status:** Phase 0 source inventory; static links plus Phase 0B pinned finalist commits.
|
||||
**Scope:** Repository/source references used during Phase 0. Runtime findings are recorded in `PHASE-0B-RECOMMENDATION.md` and supporting Phase 0B evidence.
|
||||
|
||||
## Phase 0B Production Selection
|
||||
|
||||
Selected production base:
|
||||
|
||||
- AI-DnD
|
||||
- pinned Phase 0B commit: `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
||||
- license: MIT
|
||||
|
||||
Phase 0B also evaluated:
|
||||
|
||||
- Open Dungeon `b0a79f96bf852be7b4e53908dff6a7f7c179da23`
|
||||
- ai-adventure `873ea9180d5b611576cddb155921fc16a17ae88b`
|
||||
|
||||
The source list below records the Phase 0 research set.
|
||||
|
||||
## Finalists
|
||||
|
||||
### AI-DnD
|
||||
- Repository: https://github.com/parththakkar106/AI-DnD
|
||||
- README / architecture summary: https://github.com/parththakkar106/AI-DnD/blob/main/README.md
|
||||
- Design guide: https://github.com/parththakkar106/AI-DnD/blob/main/docs/GUIDE.md
|
||||
- License: MIT
|
||||
|
||||
### Open Dungeon
|
||||
- Repository: https://github.com/newideas99/open-dungeon
|
||||
- Database layer: https://github.com/newideas99/open-dungeon/blob/main/src/lib/db.ts
|
||||
- Prompt/context layer: https://github.com/newideas99/open-dungeon/blob/main/src/lib/story-prompt.ts
|
||||
- Environment configuration: https://github.com/newideas99/open-dungeon/blob/main/.env.example
|
||||
- Package manifest: https://github.com/newideas99/open-dungeon/blob/main/package.json
|
||||
- License: MIT
|
||||
|
||||
### Local Adventure Engine / ai-adventure
|
||||
- Repository: https://github.com/CaoRuiming/ai-adventure
|
||||
- Architecture: https://github.com/CaoRuiming/ai-adventure/blob/main/docs/architecture.md
|
||||
- Privacy/security: https://github.com/CaoRuiming/ai-adventure/blob/main/docs/privacy-and-security.md
|
||||
- License: Apache-2.0
|
||||
|
||||
## High-value reference projects
|
||||
|
||||
### aiMultiFool
|
||||
- Repository: https://github.com/omgboohoo/aimultifool
|
||||
- Role: local roleplay/RAG/encryption ideas
|
||||
- License: GPL-3.0
|
||||
|
||||
### Chronicler
|
||||
- Repository: https://github.com/yantrikos/chronicler
|
||||
- Role: memory tiers, canon/heuristic/reflex separation, anti-confabulation patterns
|
||||
- License: MIT (application); YantrikDB is separately Apache-2.0
|
||||
|
||||
### Interactive Fiction Framework
|
||||
- Repository: https://github.com/georgebutler/interactive-fiction-framework
|
||||
- Role: Story Bible, model-as-prose-writer/application-as-state-owner, validation/fallback patterns
|
||||
- License: MIT
|
||||
|
||||
### Gamentic
|
||||
- Repository: https://github.com/hec-ovi/gamentic
|
||||
- Role: local text/image/voice provider abstraction, asynchronous media generation, large automated test suite
|
||||
- License: MIT
|
||||
|
||||
### Sonder Engine
|
||||
- Repository: https://github.com/N0819/Sonder_Engine
|
||||
- Role: objective truth vs perception/memory/belief, commit boundary, sophisticated character knowledge
|
||||
- License: MIT
|
||||
|
||||
### Corvus Story Core
|
||||
- Repository: https://github.com/JustLateNightAI/Corvus-Story-Core
|
||||
- Role: structured state + local ComfyUI/TTS integration
|
||||
- License: MIT
|
||||
|
||||
## Mature ecosystem references
|
||||
|
||||
- SillyTavern: https://github.com/SillyTavern/SillyTavern
|
||||
- RisuAI: https://github.com/kwaroran/RisuAI
|
||||
- KoboldAI Client: https://github.com/KoboldAI/KoboldAI-Client
|
||||
|
||||
## Important research caveat
|
||||
|
||||
Repository documentation can be stale relative to current source. Phase 0B should pin exact commit SHAs at clone time, run the projects, run their tests, and verify all network behavior locally.
|
||||