The README, the project page and the engineering guide all describe a linear story. The tree shipped two days ago. Every published surface is a phase behind, and the guide is not merely behind — it is wrong in a way that costs a reader time. Its 2.2 was "Two coordinate systems, and the bug class they create", and it explained the codebase through position_of_index, note_action_removed and settled_story_actions. All three were deleted in SP3. 2.3 explained retry through Action.variants and state_before. Somebody reading either would go looking for machinery that is not there, which is worse than a gap. So 2.2 is now "The story is a tree", written at the depth 1.2 and 1.3 are written at: the seven bugs that turned out to be one bug, the lineage clause and the two properties that make fork count free, why takes group by parent_id rather than by coordinate, cursors becoming anchors, and a closing list of what the design is honest about. 2.3 is rewritten around state_after and takes, and 1.1 and 1.5 follow, because the pipeline no longer snapshots before the call and the memory bank no longer holds an action back. The numbers were simply old: 151 tests where there are 440, 37 migrations where there are 64, twelve phases where there are fourteen. They appear in four places across the README, the project page's stat tiles and the guide's results table. The measured branch cost — 103 B, and 1.007x the page load of the same story flat — is added beside the egress and turn-cost figures it belongs with, since it is the number that answers "what does branching cost me". Three screenshots, on a new tools/shots_fixture.py: the Bandit Camp demo driven through eight written turns with written deltas, three discarded takes forked onto branches of their own, one off a branch so the map has to nest. Same reason tree_fixture.py is committed — the shots have to be reproducible and the frontend still has no test runner. play-world-state.jpg is reshot because it predates the entire tree UI; the map and the branches panel are new. Note for next time: docs/guide.html is hand-written, not generated from the Markdown, so every guide edit is two edits in two vocabularies. Both files were checked for tag balance and both pages rendered locally before this landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
1568 lines
81 KiB
HTML
1568 lines
81 KiB
HTML
<!doctype html>
|
||
<html lang="en">
|
||
<head>
|
||
<meta charset="utf-8">
|
||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||
<title>AI D&D — the engineering guide</title>
|
||
<meta name="description" content="How the AI D&D storytelling engine works and why it was built this way: context assembly under a token budget, an AI-proposes/Python-referees world-state engine, a branching story tree, an embedding memory bank, and the production concerns around a server-funded demo key.">
|
||
<meta property="og:title" content="AI D&D — the engineering guide">
|
||
<meta property="og:description" content="Design decisions, measured results, and spoken answers for every part of the engine.">
|
||
<meta property="og:type" content="article">
|
||
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'%3E%3Ctext y='.9em' font-size='90'%3E%E2%9A%94%3C/text%3E%3C/svg%3E">
|
||
<style>
|
||
/* ------------------------------------------------------------------ tokens
|
||
Single-theme by intent: this page is part of the AI D&D project site, which
|
||
is a committed dark identity. Every colour is painted explicitly so the page
|
||
holds on any host ground. */
|
||
:root {
|
||
--ink: #0b0b12;
|
||
--ink-raised: #12121b;
|
||
--panel: #16161f;
|
||
--rule: #282838;
|
||
--rule-soft: #1e1e2a;
|
||
--text: #e5e0d3;
|
||
--dim: #948d7e;
|
||
--dimmer: #6b6559;
|
||
--gold: #d4a94e;
|
||
--gold-lit: #ecc978;
|
||
--gold-deep: #8e7233;
|
||
--say: #79a892;
|
||
--say-deep: #2c4740;
|
||
--trap: #c07f66;
|
||
--trap-deep: #4a2e25;
|
||
|
||
--display: 'Iowan Old Style', 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
|
||
--body: 'Iowan Old Style', Charter, Georgia, 'Times New Roman', serif;
|
||
--ui: ui-sans-serif, system-ui, -apple-system, 'Segoe UI', sans-serif;
|
||
--mono: ui-monospace, SFMono-Regular, 'SF Mono', Menlo, Consolas, monospace;
|
||
|
||
--measure: 34rem;
|
||
color-scheme: dark;
|
||
}
|
||
|
||
*, *::before, *::after { box-sizing: border-box; }
|
||
|
||
html { -webkit-text-size-adjust: 100%; scroll-behavior: smooth; }
|
||
@media (prefers-reduced-motion: reduce) {
|
||
html { scroll-behavior: auto; }
|
||
* { animation-duration: .001ms !important; transition-duration: .001ms !important; }
|
||
}
|
||
|
||
body {
|
||
margin: 0;
|
||
background:
|
||
radial-gradient(900px 520px at 12% -6%, rgba(212,169,78,.055), transparent 62%),
|
||
radial-gradient(760px 460px at 92% 104%, rgba(96,84,150,.07), transparent 58%),
|
||
var(--ink);
|
||
background-attachment: fixed;
|
||
color: var(--text);
|
||
font-family: var(--body);
|
||
font-size: 17px;
|
||
line-height: 1.68;
|
||
-webkit-font-smoothing: antialiased;
|
||
text-rendering: optimizeLegibility;
|
||
}
|
||
|
||
/* ------------------------------------------------------------- reading rail */
|
||
#progress {
|
||
position: fixed; inset: 0 auto auto 0; height: 2px; width: 0;
|
||
background: linear-gradient(90deg, var(--gold-deep), var(--gold-lit));
|
||
z-index: 50;
|
||
}
|
||
|
||
.bar {
|
||
position: sticky; top: 0; z-index: 40;
|
||
display: flex; align-items: center; justify-content: space-between; gap: 1rem;
|
||
padding: .6rem 1.15rem;
|
||
background: rgba(11,11,18,.9);
|
||
backdrop-filter: blur(10px);
|
||
border-bottom: 1px solid var(--rule-soft);
|
||
font-family: var(--ui);
|
||
font-size: .74rem;
|
||
letter-spacing: .11em;
|
||
text-transform: uppercase;
|
||
}
|
||
.bar b { color: var(--gold); font-weight: 600; letter-spacing: .16em; }
|
||
.bar a { color: var(--dim); text-decoration: none; border-bottom: 1px solid transparent; }
|
||
.bar a:hover, .bar a:focus-visible { color: var(--gold-lit); border-bottom-color: var(--gold-deep); }
|
||
|
||
/* ------------------------------------------------------------------ layout */
|
||
.wrap { max-width: var(--measure); margin: 0 auto; padding: 0 1.15rem; }
|
||
|
||
main { padding-bottom: 5rem; }
|
||
|
||
/* --------------------------------------------------------------- masthead */
|
||
/* Only the block axis — .mast is also a .wrap, whose inline padding must survive. */
|
||
.mast { padding-top: 3.4rem; padding-bottom: 2.2rem; }
|
||
.eyebrow {
|
||
font-family: var(--ui); font-size: .68rem; letter-spacing: .3em;
|
||
text-transform: uppercase; color: var(--gold); margin: 0 0 1.1rem;
|
||
}
|
||
.mast h1 {
|
||
font-family: var(--display); font-weight: 600;
|
||
font-size: clamp(2rem, 8vw, 2.9rem); line-height: 1.12;
|
||
margin: 0 0 .5rem; letter-spacing: -.005em; text-wrap: balance;
|
||
color: var(--text);
|
||
}
|
||
.mast .lede { color: var(--dim); font-size: 1.02rem; margin: 0 0 1.4rem; text-wrap: pretty; }
|
||
.mast .meta {
|
||
font-family: var(--ui); font-size: .78rem; color: var(--dimmer);
|
||
border-top: 1px solid var(--rule-soft); padding-top: .9rem;
|
||
}
|
||
.mast .meta a { color: var(--gold); text-decoration: none; border-bottom: 1px solid var(--gold-deep); }
|
||
|
||
/* -------------------------------------------------------------------- TOC */
|
||
.toc {
|
||
border: 1px solid var(--rule); border-radius: 3px;
|
||
background: var(--ink-raised); margin: 0 0 3rem;
|
||
}
|
||
.toc > summary {
|
||
cursor: pointer; list-style: none; padding: .85rem 1.1rem;
|
||
font-family: var(--ui); font-size: .72rem; letter-spacing: .18em;
|
||
text-transform: uppercase; color: var(--gold);
|
||
display: flex; align-items: center; justify-content: space-between;
|
||
}
|
||
.toc > summary::-webkit-details-marker { display: none; }
|
||
.toc > summary::after { content: '+'; color: var(--dimmer); font-size: 1rem; }
|
||
.toc[open] > summary::after { content: '−'; }
|
||
.toc ol { list-style: none; margin: 0; padding: 0 1.1rem 1rem; font-family: var(--ui); font-size: .87rem; }
|
||
.toc li { padding: .28rem 0; border-top: 1px solid var(--rule-soft); }
|
||
.toc li:first-child { border-top: 0; }
|
||
.toc .sub { padding-left: 1.1rem; color: var(--dim); font-size: .83rem; }
|
||
.toc a { color: var(--text); text-decoration: none; }
|
||
.toc .sub a { color: var(--dim); }
|
||
.toc a:hover, .toc a:focus-visible { color: var(--gold-lit); }
|
||
|
||
/* ------------------------------------------------------------ part divider */
|
||
.part { margin: 4.5rem 0 2.2rem; scroll-margin-top: 3.5rem; }
|
||
.part .num {
|
||
font-family: var(--ui); font-size: .68rem; letter-spacing: .3em;
|
||
text-transform: uppercase; color: var(--gold-deep);
|
||
display: flex; align-items: center; gap: .8rem;
|
||
}
|
||
.part .num::after { content: ''; flex: 1; height: 1px; background: var(--rule); }
|
||
.part h2 {
|
||
font-family: var(--display); font-weight: 600;
|
||
font-size: clamp(1.55rem, 5.5vw, 2rem); line-height: 1.2;
|
||
margin: .55rem 0 0; color: var(--gold-lit); text-wrap: balance;
|
||
}
|
||
.part p { color: var(--dim); margin: .6rem 0 0; }
|
||
|
||
/* ------------------------------------------------------------------ prose */
|
||
h3 {
|
||
font-family: var(--display); font-weight: 600;
|
||
font-size: 1.32rem; line-height: 1.25; margin: 3rem 0 .2rem;
|
||
color: var(--text); text-wrap: balance; scroll-margin-top: 4rem;
|
||
}
|
||
h3 .h-num {
|
||
display: block; font-family: var(--ui); font-size: .66rem; letter-spacing: .24em;
|
||
text-transform: uppercase; color: var(--gold-deep); margin-bottom: .4rem;
|
||
}
|
||
h4 {
|
||
font-family: var(--ui); font-weight: 600; font-size: .78rem;
|
||
letter-spacing: .14em; text-transform: uppercase; color: var(--gold);
|
||
margin: 2rem 0 .5rem;
|
||
}
|
||
p { margin: 0 0 1.05rem; text-wrap: pretty; }
|
||
a { color: var(--gold-lit); text-decoration-color: var(--gold-deep); text-underline-offset: .18em; }
|
||
strong { color: #f3efe4; font-weight: 600; }
|
||
em { color: var(--text); }
|
||
hr { border: 0; border-top: 1px solid var(--rule-soft); margin: 2.6rem 0; }
|
||
|
||
ul, ol { margin: 0 0 1.05rem; padding-left: 1.25rem; }
|
||
li { margin: 0 0 .45rem; }
|
||
li::marker { color: var(--gold-deep); }
|
||
|
||
code {
|
||
font-family: var(--mono); font-size: .84em;
|
||
background: #1b1b26; border: 1px solid var(--rule-soft);
|
||
border-radius: 2px; padding: .1em .32em; color: #dfd6bd;
|
||
}
|
||
pre {
|
||
font-family: var(--mono); font-size: .78rem; line-height: 1.6;
|
||
background: var(--panel); border: 1px solid var(--rule);
|
||
border-left: 2px solid var(--gold-deep);
|
||
border-radius: 2px; padding: .95rem 1rem; margin: 0 0 1.3rem;
|
||
overflow-x: auto; color: #ccc4ae;
|
||
}
|
||
pre code { background: 0; border: 0; padding: 0; font-size: inherit; color: inherit; }
|
||
|
||
blockquote {
|
||
margin: 0 0 1.3rem; padding: 0 0 0 1rem;
|
||
border-left: 2px solid var(--rule); color: var(--dim); font-style: italic;
|
||
}
|
||
|
||
/* ------------------------------------------------------------------ tables */
|
||
.scroll { overflow-x: auto; margin: 0 0 1.4rem; border: 1px solid var(--rule); border-radius: 2px; }
|
||
table { border-collapse: collapse; width: 100%; font-family: var(--ui); font-size: .82rem; }
|
||
th, td { text-align: left; padding: .58rem .8rem; border-bottom: 1px solid var(--rule-soft); vertical-align: top; }
|
||
th {
|
||
background: var(--ink-raised); color: var(--gold); font-weight: 600;
|
||
font-size: .68rem; letter-spacing: .12em; text-transform: uppercase; white-space: nowrap;
|
||
}
|
||
tr:last-child td { border-bottom: 0; }
|
||
td code { font-size: .8em; }
|
||
.nums td { font-variant-numeric: tabular-nums; }
|
||
/* Keep the label column from collapsing to one word per line on a phone; the
|
||
figures only get nowrap once there is room for it. */
|
||
.nums td:first-child { min-width: 10rem; }
|
||
@media (min-width: 34rem) { .nums td:last-child { white-space: nowrap; } }
|
||
|
||
/* ---------------------------------------------------------------- callouts */
|
||
.trap {
|
||
margin: 1.6rem 0 1.5rem; padding: 1rem 1.1rem;
|
||
border: 1px solid var(--rule); border-radius: 2px;
|
||
background: var(--ink-raised); font-size: .95rem;
|
||
}
|
||
.trap { border-left: 2px solid var(--trap); }
|
||
.trap .tag {
|
||
display: block; font-family: var(--ui); font-size: .64rem;
|
||
letter-spacing: .2em; text-transform: uppercase; margin-bottom: .55rem;
|
||
}
|
||
.trap .tag { color: var(--trap); }
|
||
.trap p { margin: 0 0 .7rem; }
|
||
.trap p:last-child { margin: 0; }
|
||
|
||
/* -------------------------------------------------------------- decision */
|
||
.decision {
|
||
margin: 1.5rem 0; border: 1px solid var(--rule); border-radius: 2px;
|
||
background: var(--panel); overflow: hidden;
|
||
}
|
||
.decision > div { padding: .7rem 1rem; border-bottom: 1px solid var(--rule-soft); }
|
||
.decision > div:last-child { border-bottom: 0; }
|
||
.decision dt {
|
||
font-family: var(--ui); font-size: .63rem; letter-spacing: .2em;
|
||
text-transform: uppercase; color: var(--gold-deep); margin-bottom: .25rem;
|
||
}
|
||
.decision dd { margin: 0; font-size: .93rem; }
|
||
|
||
/* -------------------------------------------------------------- pipeline */
|
||
.pipe { margin: 1.5rem 0; font-family: var(--ui); font-size: .8rem; }
|
||
.pipe .step {
|
||
position: relative; padding: .55rem .8rem .55rem 1.5rem;
|
||
border: 1px solid var(--rule); border-radius: 2px;
|
||
background: var(--panel); margin-bottom: .38rem;
|
||
}
|
||
.pipe .step::before {
|
||
content: ''; position: absolute; left: .62rem; top: 50%;
|
||
width: 5px; height: 5px; margin-top: -2.5px; border-radius: 50%;
|
||
background: var(--rule); border: 1px solid var(--dimmer);
|
||
}
|
||
.pipe .step.ai::before { background: var(--gold); border-color: var(--gold-lit); }
|
||
.pipe .step.hook::before { background: var(--say-deep); border-color: var(--say); }
|
||
.pipe .step b { color: var(--text); font-weight: 600; }
|
||
.pipe .step span { color: var(--dim); }
|
||
.pipe .legend {
|
||
display: flex; flex-wrap: wrap; gap: .9rem; margin-top: .7rem;
|
||
font-size: .7rem; color: var(--dimmer); letter-spacing: .06em;
|
||
}
|
||
.pipe .legend i { display: inline-block; width: 6px; height: 6px; border-radius: 50%; margin-right: .35rem; }
|
||
|
||
/* -------------------------------------------------------------- svg figure */
|
||
figure { margin: 1.8rem 0; }
|
||
figure svg { display: block; width: 100%; height: auto; }
|
||
figcaption {
|
||
font-family: var(--ui); font-size: .74rem; color: var(--dimmer);
|
||
margin-top: .6rem; line-height: 1.5;
|
||
}
|
||
|
||
/* ------------------------------------------------------------------ stats */
|
||
.stats { display: grid; grid-template-columns: 1fr; gap: .5rem; margin: 1.6rem 0; }
|
||
@media (min-width: 34rem) { .stats { grid-template-columns: 1fr 1fr; } }
|
||
.stat {
|
||
border: 1px solid var(--rule); border-radius: 2px;
|
||
background: var(--panel); padding: .85rem 1rem;
|
||
}
|
||
.stat .v {
|
||
font-family: var(--display); font-size: 1.45rem; color: var(--gold-lit);
|
||
font-variant-numeric: tabular-nums; line-height: 1.1;
|
||
}
|
||
.stat .k {
|
||
font-family: var(--ui); font-size: .7rem; color: var(--dim);
|
||
letter-spacing: .06em; margin-top: .3rem; line-height: 1.4;
|
||
}
|
||
|
||
|
||
/* ---------------------------------------------------------------- footer */
|
||
footer {
|
||
border-top: 1px solid var(--rule); margin-top: 4rem; padding: 2rem 0 3rem;
|
||
font-family: var(--ui); font-size: .8rem; color: var(--dimmer);
|
||
}
|
||
footer a { color: var(--gold); }
|
||
|
||
:focus-visible { outline: 2px solid var(--gold); outline-offset: 2px; border-radius: 2px; }
|
||
</style>
|
||
</head>
|
||
<body>
|
||
|
||
<div id="progress"></div>
|
||
|
||
<div class="bar">
|
||
<b>⚔ AI D&D</b>
|
||
<a href="https://github.com/parththakkar106/AI-DnD">Repo →</a>
|
||
</div>
|
||
|
||
<main>
|
||
|
||
<header class="mast wrap">
|
||
<p class="eyebrow">Engineering guide</p>
|
||
<h1>How this thing works, and why it works that way</h1>
|
||
<p class="lede">An AI Dungeon-style storytelling engine. The chat loop is the boring part —
|
||
the interesting parts are the token-budget allocator, the world-state referee, the story tree
|
||
that lets a turn have more than one answer, and the memory system that decides what the model is
|
||
allowed to remember.</p>
|
||
<p class="meta">
|
||
Written to be read end to end. Every section states the decision, the reasoning behind
|
||
it, and what it cost. ·
|
||
<a href="https://parththakkar106.github.io/AI-DnD/">Project page</a> ·
|
||
<a href="https://github.com/parththakkar106/AI-DnD">Source</a>
|
||
</p>
|
||
</header>
|
||
|
||
<div class="wrap">
|
||
<details class="toc" open>
|
||
<summary>Contents</summary>
|
||
<ol>
|
||
<li><a href="#p0">Part 0 — Orientation</a></li>
|
||
<li><a href="#p1">Part 1 — The AI layer</a></li>
|
||
<li class="sub"><a href="#s11">1.1 The turn pipeline</a></li>
|
||
<li class="sub"><a href="#s12">1.2 Context assembly is a budget problem</a></li>
|
||
<li class="sub"><a href="#s13">1.3 World state: the AI proposes, Python referees</a></li>
|
||
<li class="sub"><a href="#s14">1.4 Output length, by measurement</a></li>
|
||
<li class="sub"><a href="#s15">1.5 The memory bank</a></li>
|
||
<li class="sub"><a href="#s16">1.6 Streaming</a></li>
|
||
<li class="sub"><a href="#s17">1.7 The scripting sandbox</a></li>
|
||
<li class="sub"><a href="#s18">1.8 Why there is no agent framework</a></li>
|
||
<li><a href="#p2">Part 2 — Data and correctness</a></li>
|
||
<li class="sub"><a href="#s22">2.2 The story is a tree</a></li>
|
||
<li class="sub"><a href="#s23">2.3 Undo and retry that rewind</a></li>
|
||
<li class="sub"><a href="#s25">2.5 The 189× egress fix</a></li>
|
||
<li><a href="#p3">Part 3 — Production concerns</a></li>
|
||
<li><a href="#p4">Part 4 — The web plumbing</a></li>
|
||
<li><a href="#p5">Part 5 — Results and limitations</a></li>
|
||
</ol>
|
||
</details>
|
||
</div>
|
||
|
||
<!-- ===================================================== PART 0 ========== -->
|
||
<section class="wrap part" id="p0">
|
||
<p class="num">Part 0</p>
|
||
<h2>Orientation</h2>
|
||
<p>What the thing is, in the fewest words that are still true.</p>
|
||
</section>
|
||
|
||
<div class="wrap">
|
||
|
||
<h3><span class="h-num">0.1</span>What it is</h3>
|
||
|
||
<p>An AI Dungeon clone. You write a scenario, then play an open-ended text adventure where a
|
||
language model narrates the world. You type “I open the door”, the model writes what happens
|
||
next, and it remembers what came before.</p>
|
||
|
||
<p>Four things make it more than a chat wrapper:</p>
|
||
|
||
<ol>
|
||
<li><strong>A context engine.</strong> The model has a limited input window. The app decides,
|
||
every single turn, which pieces of the story get to be in the prompt and which get dropped.</li>
|
||
<li><strong>A world-state engine.</strong> The scenario declares stats — <code>hp</code>,
|
||
<code>trust</code>, <code>day</code>. The model proposes changes each turn; a Python engine
|
||
decides what actually sticks.</li>
|
||
<li><strong>A story tree.</strong> The story is not a list. Any turn can hold more than one
|
||
take, and writing below one that isn’t the live one starts a branch — which borrows every turn
|
||
above the fork instead of copying it.</li>
|
||
<li><strong>A scripting sandbox.</strong> Real AI Dungeon JavaScript scripts import and run,
|
||
inside an embedded QuickJS interpreter.</li>
|
||
</ol>
|
||
|
||
<p>Runs locally against Ollama for free, or hosted against any OpenAI-compatible endpoint.</p>
|
||
|
||
<h3><span class="h-num">0.2</span>The stack, and what each part is doing</h3>
|
||
|
||
<div class="scroll">
|
||
<table>
|
||
<thead><tr><th>Piece</th><th>What it actually does here</th></tr></thead>
|
||
<tbody>
|
||
<tr><td><strong>FastAPI</strong></td><td>The HTTP server. Every URL like <code>/api/adventures/3/actions</code> maps to a Python function. Also does the SSE streaming.</td></tr>
|
||
<tr><td><strong>SQLAlchemy</strong></td><td>Lets you write Python classes instead of SQL. <code>Adventure</code>, <code>Action</code>, <code>Memory</code> are Python classes; SQLAlchemy turns them into tables and turns attribute access into <code>SELECT</code>s.</td></tr>
|
||
<tr><td><strong>SQLite / Postgres</strong></td><td>The database. SQLite is one file on disk (local). Postgres is a server (hosted, on Neon). Same code talks to both.</td></tr>
|
||
<tr><td><strong>React</strong></td><td>The UI. Describes what the screen should look like for a given state; when the state changes it re-renders.</td></tr>
|
||
<tr><td><strong>Vite</strong></td><td>The frontend build tool and dev server. Bundles React into plain JS the browser can load.</td></tr>
|
||
<tr><td><strong>httpx</strong></td><td>The Python HTTP client used to call the model endpoint.</td></tr>
|
||
<tr><td><strong>tiktoken</strong></td><td>Counts tokens, so the budgeting is real arithmetic and not a guess.</td></tr>
|
||
<tr><td><strong>QuickJS</strong></td><td>A small embeddable JavaScript engine, used as a sandbox for user scripts.</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>The whole thing is one process in production: FastAPI serves the API <em>and</em> the built
|
||
React files from the same port.</p>
|
||
|
||
<h4>The shape of one request</h4>
|
||
|
||
<pre><code>you tap "Do"
|
||
→ POST /api/adventures/3/actions
|
||
{type: "do", text: "open the door"}
|
||
→ check ownership, rate limit, turn lock
|
||
→ assemble the prompt ← the interesting part
|
||
→ POST to the model endpoint, stream=true
|
||
→ tokens come back one at a time
|
||
→ each is forwarded on as a Server-Sent Event
|
||
→ React appends it to the screen as it arrives
|
||
→ stream ends: parse the state block, referee
|
||
it, save the action</code></pre>
|
||
|
||
</div>
|
||
|
||
<!-- ===================================================== PART 1 ========== -->
|
||
<section class="wrap part" id="p1">
|
||
<p class="num">Part 1</p>
|
||
<h2>The AI layer</h2>
|
||
<p>Where most of the design effort went. Everything here is a decision someone could
|
||
reasonably disagree with.</p>
|
||
</section>
|
||
|
||
<div class="wrap">
|
||
|
||
<h3 id="s11"><span class="h-num">1.1</span>The turn pipeline</h3>
|
||
|
||
<p>Everything that happens between “player pressed a button” and “text is on screen”.
|
||
Source: <code>backend/app/routers/adventures.py</code>.</p>
|
||
|
||
<div class="pipe">
|
||
<div class="step hook"><b>onInput</b> <span>— user JS may rewrite or block the input</span></div>
|
||
<div class="step"><b>store the player action</b></div>
|
||
<div class="step ai"><b>retrieve memories</b> <span>— embed recent story, cosine-rank the bank</span></div>
|
||
<div class="step ai"><b>build_context()</b> <span>— the budget allocator</span></div>
|
||
<div class="step hook"><b>onModelContext</b> <span>— user JS may rewrite the whole prompt</span></div>
|
||
<div class="step"><b>snapshot the exact prompt</b> <span>— for the Insights panel</span></div>
|
||
<div class="step ai"><b>provider.generate()</b> <span>— streamed, token by token</span></div>
|
||
<div class="step hook"><b>onOutput</b></div>
|
||
<div class="step ai"><b>extract + referee the state block</b> <span>— then strip it from the prose</span></div>
|
||
<div class="step"><b>save the action</b> <span>— stamped with the state it leaves behind</span></div>
|
||
<div class="step ai"><b>background: summarize + embed</b> <span>— fire-and-forget</span></div>
|
||
<div class="legend">
|
||
<span><i style="background:var(--gold)"></i>model or prompt work</span>
|
||
<span><i style="background:var(--say)"></i>user script hook</span>
|
||
<span><i style="background:var(--dimmer)"></i>persistence</span>
|
||
</div>
|
||
</div>
|
||
|
||
<p>Two design choices are visible in that list before any of the details.</p>
|
||
|
||
<p><strong>The prompt is snapshotted, not reconstructed.</strong> Every AI action stores the
|
||
exact text that was sent to the model. That’s what powers the Insights panel — open any turn and
|
||
see each context component, its token cost, and why it was included. It’s also what makes prompt
|
||
bugs findable. The cost is storage, about 74 KB per turn, which turns into a real performance
|
||
problem later (see <a href="#s25">2.5</a>).</p>
|
||
|
||
<p><strong>Every node records the state it leaves behind.</strong> <code>state_after</code> and
|
||
<code>world_state_after</code> are stapled onto the action once its hooks and its delta have run,
|
||
so a node carries the scoreboard and the RPG stats as they stood when that turn finished. Rewinding
|
||
to <em>before</em> a turn is then a read of the node in front of it — the same move as switching to
|
||
another branch. One mechanism, and it is why undo, retry and a branch switch all put the numbers
|
||
back rather than only rewriting text.</p>
|
||
|
||
<h3 id="s12"><span class="h-num">1.2</span>Context assembly is a budget problem</h3>
|
||
|
||
<p>Source: <code>backend/app/context/builder.py</code>.</p>
|
||
|
||
<h4>The problem</h4>
|
||
|
||
<p>The model can only read so much. Say the budget is 8,000 tokens. A 200-turn adventure has far
|
||
more story than that. Something has to be dropped, and <em>what</em> gets dropped decides whether
|
||
the story stays coherent.</p>
|
||
|
||
<h4>The naive version, and why it breaks</h4>
|
||
|
||
<p>Send the last N turns. That fails in two directions: N turns of short exchanges wastes the
|
||
window, and N turns of long ones overflows it. Worse, “the last N turns” throws away the things
|
||
that matter most — the premise, the character sheet, the fact that you promised the innkeeper
|
||
you’d return.</p>
|
||
|
||
<h4>What this app does</h4>
|
||
|
||
<p>Split the prompt into <strong>fixed</strong> sections and <strong>elastic</strong> ones.</p>
|
||
|
||
<figure>
|
||
<svg viewBox="-4 -4 648 176" role="img" aria-label="A stacked bar showing the context token budget: fixed sections are reserved first, then story cards take up to 40 percent of what remains, and story history fills the rest, newest first.">
|
||
<defs>
|
||
<linearGradient id="gFixed" x1="0" y1="0" x2="1" y2="0">
|
||
<stop offset="0" stop-color="#8e7233"/><stop offset="1" stop-color="#d4a94e"/>
|
||
</linearGradient>
|
||
</defs>
|
||
<text x="0" y="14" fill="#948d7e" font-family="ui-sans-serif, system-ui" font-size="11" letter-spacing="1.4">CONTEXT TOKEN BUDGET</text>
|
||
|
||
<!-- full bar outline -->
|
||
<rect x="0" y="26" width="640" height="44" fill="#16161f" stroke="#282838"/>
|
||
|
||
<!-- fixed -->
|
||
<rect x="0" y="26" width="228" height="44" fill="url(#gFixed)" opacity="0.85"/>
|
||
<!-- cards -->
|
||
<rect x="228" y="26" width="165" height="44" fill="#2c4740" stroke="#79a892"/>
|
||
<!-- history -->
|
||
<rect x="393" y="26" width="247" height="44" fill="#1e1e2a" stroke="#3a3a50"/>
|
||
|
||
<text x="10" y="53" fill="#0b0b12" font-family="ui-sans-serif, system-ui" font-size="11" font-weight="600">RESERVED</text>
|
||
<text x="238" y="53" fill="#79a892" font-family="ui-sans-serif, system-ui" font-size="11" font-weight="600">CARDS ≤ 40%</text>
|
||
<text x="403" y="53" fill="#948d7e" font-family="ui-sans-serif, system-ui" font-size="11" font-weight="600">HISTORY, NEWEST FIRST</text>
|
||
|
||
<!-- brace for available -->
|
||
<path d="M228 78 L228 86 L640 86 L640 78" fill="none" stroke="#3a3a50"/>
|
||
<text x="434" y="102" fill="#948d7e" font-family="ui-sans-serif, system-ui" font-size="10.5" text-anchor="middle">available = budget − reserved</text>
|
||
|
||
<text x="0" y="128" fill="#6b6559" font-family="ui-monospace, monospace" font-size="10.5">narrator · stat guide · world state · emit rule · ai instructions</text>
|
||
<text x="0" y="144" fill="#6b6559" font-family="ui-monospace, monospace" font-size="10.5">plot essentials · story summary · retrieved memories</text>
|
||
<text x="0" y="160" fill="#6b6559" font-family="ui-monospace, monospace" font-size="10.5">↑ these are always included, whatever they cost</text>
|
||
</svg>
|
||
<figcaption>Fixed sections are reserved first and never dropped. What’s left is the elastic
|
||
budget: triggered story cards may take up to 40% of it, and story history spends the remainder
|
||
filling backwards from the newest turn.</figcaption>
|
||
</figure>
|
||
|
||
<p>The algorithm is three lines of arithmetic:</p>
|
||
|
||
<pre><code>reserved = every fixed section + note + hint + reminder
|
||
available = max(256, token_budget - reserved)
|
||
|
||
cards ≤ available * 0.4
|
||
history = available - cards_used, newest first</code></pre>
|
||
|
||
<h4>The details that are actually decisions</h4>
|
||
|
||
<p><strong>Cards are capped at 40% of the elastic budget.</strong> Story cards are triggered by
|
||
keyword match, so a scene mentioning six named things could pull in six lore entries and leave no
|
||
room for the story itself. The cap makes the failure mode “some lore is missing” instead of “the
|
||
model has no idea what just happened”. Cards that don’t fit are still <em>reported</em> to
|
||
Insights with <code>included: false</code>, so the UI can show the lore that got squeezed out.</p>
|
||
|
||
<p><strong>History fills newest-first and stops.</strong> Oldest turns fall out. That’s the right
|
||
direction because the old material isn’t actually lost — it’s been summarized into memories and
|
||
the running summary, which live in the fixed section.</p>
|
||
|
||
<p><strong>If even the single newest turn is over budget, it gets hard-truncated</strong> rather
|
||
than dropped. A prompt with no story at all produces nonsense; a prompt with the tail end of the
|
||
last turn produces something.</p>
|
||
|
||
<p><strong>The author’s note is injected three actions from the end</strong>, not at the top.
|
||
Instructions placed near the end of a prompt have more influence on what comes next than
|
||
instructions at the top — recency. The author’s note is a steering control (“keep it tense”), so
|
||
it goes where steering works.</p>
|
||
|
||
<p><strong>The world-state reminder goes dead last.</strong> The full emit rule lives up in the
|
||
system block, hundreds of tokens away from where the model starts writing. A one-line reminder
|
||
occupies the final slot. Same recency logic, applied to the thing most likely to be forgotten.</p>
|
||
|
||
<div class="trap">
|
||
<span class="tag">The subtle one</span>
|
||
<p><strong>Past AI turns get their state block re-attached.</strong> The block is stripped from
|
||
the text before storage, so a replayed history would show the model twenty of its own past
|
||
turns that contain <em>no</em> state block — teaching it, by imitation, to stop emitting one.
|
||
So the history builder reconstructs the block from the stored delta and re-appends it. The
|
||
model sees its own pattern and keeps following it.</p>
|
||
</div>
|
||
|
||
<h4>The performance trap hiding in this</h4>
|
||
|
||
<p>Building the context needs the newest ~6,000 tokens of story. The obvious implementation reads
|
||
<code>adventure.actions</code> — which loads every row of the adventure — then throws 90% of it
|
||
away. At turn 200 that was 839 KB of database reads to use maybe 70 KB, growing every turn.</p>
|
||
|
||
<p><code>context/history.py</code> fixes it by serving three shapes directly from SQL: a tail, a
|
||
slice, and a count. <code>window_covering()</code> fetches the newest 32 actions, measures their
|
||
real token count, and if that’s short of the budget it <em>projects</em> how many more it needs
|
||
from the average length just measured, rather than blindly doubling:</p>
|
||
|
||
<pre><code>average = tokens / len(actions)
|
||
projected = int(budget / average * 1.15) + 8</code></pre>
|
||
|
||
<p>Each round fetches only what it doesn’t already hold, so no row is read twice. The same turn
|
||
costs 129 KB instead of 839 KB, and stops growing at around turn 50 — the cost is bounded by the
|
||
context budget instead of by the length of the story.</p>
|
||
|
||
<p>There’s a second rule in that module worth naming: <strong>if the actions are already loaded
|
||
in memory, slice them instead of querying.</strong> The scripting pipeline hands the whole history
|
||
to user scripts, because AI Dungeon’s API requires it, so on a scripted adventure the rows are
|
||
already there — issuing a query beside them would mean paying twice.</p>
|
||
|
||
<h3 id="s13"><span class="h-num">1.3</span>World state: the AI proposes, Python referees</h3>
|
||
|
||
<p>Source: <code>backend/app/worldstate/engine.py</code>.</p>
|
||
|
||
<h4>The question</h4>
|
||
|
||
<p>You want an RPG layer — hit points, trust, quest progress. Who owns the numbers?</p>
|
||
|
||
<div class="decision">
|
||
<div>
|
||
<dt>Option A — a deterministic dice engine</dt>
|
||
<dd>The engine rolls and applies damage, the model narrates the result. This is what a real
|
||
RPG does. It loses here because the action space is unbounded: the player can type anything,
|
||
and mapping arbitrary natural language onto a fixed rules system is a harder problem than the
|
||
one being solved.</dd>
|
||
</div>
|
||
<div>
|
||
<dt>Option B — the model owns the numbers</dt>
|
||
<dd>Track hp in the prose and trust it. Fails immediately. Models are bad at arithmetic, worse
|
||
at holding a number across twenty turns, and completely unable to obey their own frequency
|
||
rules — tell one “only change this every 5 turns” and it changes it every turn.</dd>
|
||
</div>
|
||
<div>
|
||
<dt>Option C — chosen: propose and dispose</dt>
|
||
<dd>The model narrates and appends a JSON delta of what changed. Python validates and clamps
|
||
it before anything is stored. The model owns intent; the engine owns arithmetic.</dd>
|
||
</div>
|
||
</div>
|
||
|
||
<pre><code>narration: "The blade catches your shoulder. Gwen shouts and drags you back."
|
||
|
||
```state
|
||
{"player.hp": -15, "npc.gwen.trust": 5, "milestones.escaped": true}
|
||
```</code></pre>
|
||
|
||
<p>The engine then applies, in order:</p>
|
||
|
||
<div class="scroll">
|
||
<table>
|
||
<thead><tr><th>Rule</th><th>What it stops</th></tr></thead>
|
||
<tbody>
|
||
<tr><td>Path must exist in the schema</td><td>Hallucinated stats</td></tr>
|
||
<tr><td>Value must be the right type</td><td><code>"a lot"</code> instead of <code>-15</code></td></tr>
|
||
<tr><td>Cooldown</td><td>Changing a stat more often than the scenario allows</td></tr>
|
||
<tr><td>Counters can’t decrease</td><td>The in-game day going backwards</td></tr>
|
||
<tr><td><code>max_delta_per_turn</code></td><td>Losing 90 hp to a stubbed toe</td></tr>
|
||
<tr><td>Clamp to <code>min</code>/<code>max</code></td><td>Negative hp, trust above 100</td></tr>
|
||
<tr><td>Milestones sticky, <code>true</code> only</td><td>Un-completing a quest</td></tr>
|
||
<tr><td>Flags are two-way booleans</td><td>Deliberately unrestricted — that’s what flags are for</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>Everything rejected is <em>reported</em>, not silently swallowed. The Insights panel shows
|
||
applied, clamped and rejected paths per turn, and the chip under each narration shows what
|
||
actually changed.</p>
|
||
|
||
<h4>The reliability mechanism: word bands</h4>
|
||
|
||
<p>A stat can carry <strong>bands</strong>:</p>
|
||
|
||
<pre><code>"hp": { "min": 0, "max": 100, "initial": 100,
|
||
"bands": [[0,20,"very weak"],[20,40,"hurt"],[40,60,"minor damage"],
|
||
[60,90,"healthy"],[90,100,"full health"]] }</code></pre>
|
||
|
||
<p>Two things use them. The live state block shows the current band label —
|
||
<code>hp 55/100 (minor damage)</code> — so the model reads a <em>word</em>, not just a number. And
|
||
the stat guide shows the whole ladder once per turn, so the model can see the full scale it’s
|
||
reasoning across.</p>
|
||
|
||
<p>The point: models reason well over semantics and badly over arithmetic. “He’s badly hurt, so a
|
||
solid hit should take him to very weak” is a judgement a model can make. “55 minus 22 is 33” is
|
||
one it will get wrong often enough to matter.</p>
|
||
|
||
<h4>The failure philosophy</h4>
|
||
|
||
<p>Nothing in the world-state engine raises. A malformed delta returns <code>{}</code> and the
|
||
turn continues. The parser is deliberately tolerant — it strips trailing commas and leading
|
||
<code>+</code> signs on numbers, both of which weaker free models emit and strict JSON rejects. It
|
||
accepts a <code>state</code>, <code>json</code> or unlabelled fence, and falls back to a bare JSON
|
||
object at the end of the text, but only if it parses into something that looks like a delta, so
|
||
prose ending in <code>}</code> is never eaten.</p>
|
||
|
||
<p>This matters because the public demo runs on free-tier models. A stricter parser would mean a
|
||
good model works and a free one doesn’t.</p>
|
||
|
||
<h4>One call, not two</h4>
|
||
|
||
<p>The model narrates <em>and</em> emits the delta in a single request. The alternative — narrate,
|
||
then a second call to extract structured state — is more reliable per call and costs twice the
|
||
latency and twice the rate-limit budget. On the free tier (20 requests/minute) that would halve
|
||
the playable turn rate. The tolerant parser plus the terminal reminder was the cheaper way to buy
|
||
the same reliability.</p>
|
||
|
||
<h3 id="s14"><span class="h-num">1.4</span>Output length, by measurement</h3>
|
||
|
||
<h4>The problem</h4>
|
||
|
||
<p><code>max_output_tokens</code> is a hard wall the endpoint enforces mid-sentence. Hit it and
|
||
whatever is being written gets cut off. Since the state block is emitted <em>last</em>, the state
|
||
block is what gets lost. The turn narrates fine and silently records nothing.</p>
|
||
|
||
<h4>First attempt, and the measurement</h4>
|
||
|
||
<p>Tell the model its budget: <em>“keep this turn under about N words”</em>.</p>
|
||
|
||
<div class="stats">
|
||
<div class="stat"><div class="v">174 → 246</div><div class="k">average words per turn once the “budget” hint was added — every run longer than every unhinted run (n=5)</div></div>
|
||
<div class="stat"><div class="v">170</div><div class="k">average after rephrasing the same number as a hard ceiling</div></div>
|
||
</div>
|
||
|
||
<p>Phrased as a budget, the number reads as a <em>target to fill</em>. The hint pushed turns
|
||
toward the very wall it existed to protect.</p>
|
||
|
||
<h4>The fix</h4>
|
||
|
||
<pre><code>[Hard limit: this turn must not exceed 412 words. Write only as much as the
|
||
moment needs — a typical turn is much shorter. Finish the narration and append
|
||
the state block well inside the limit.]</code></pre>
|
||
|
||
<h4>And the arithmetic around it</h4>
|
||
|
||
<pre><code>words = int((max_output_tokens - 50) * 0.75 * 0.90)</code></pre>
|
||
|
||
<ul>
|
||
<li><code>- 50</code> — tokens held back for the state block itself.</li>
|
||
<li><code>* 0.75</code> — models can’t count their own tokens, but they do follow a word budget.
|
||
English prose is roughly 0.75 words per token.</li>
|
||
<li><code>* 0.90</code> — a word budget is a suggestion the model overshoots; the cap it protects
|
||
is a hard wall. Aim 10% short so the overshoot lands in slack.</li>
|
||
<li>Below 40 words the hint is dropped entirely — it stops earning its tokens.</li>
|
||
</ul>
|
||
|
||
<h3 id="s15"><span class="h-num">1.5</span>The memory bank</h3>
|
||
|
||
<p>Source: <code>backend/app/memorybank.py</code>.</p>
|
||
|
||
<h4>The problem</h4>
|
||
|
||
<p>Story history falls out of the context window as the adventure grows. Turn 4 said you promised
|
||
the innkeeper you’d return. At turn 90 that’s long gone from the prompt — but if you walk back
|
||
into the inn, it should come back.</p>
|
||
|
||
<h4>Three layers</h4>
|
||
|
||
<div class="scroll">
|
||
<table>
|
||
<thead><tr><th>Layer</th><th>Cadence</th><th>Purpose</th></tr></thead>
|
||
<tbody>
|
||
<tr><td><strong>Memory</strong></td><td>every 6 actions, from 12</td><td>One or two past-tense sentences of concrete fact.</td></tr>
|
||
<tr><td><strong>Story summary</strong></td><td>every 15 actions</td><td>A single ≤250-word overview, rewritten by folding in the new memories.</td></tr>
|
||
<tr><td><strong>Retrieval</strong></td><td>every turn</td><td>Embed the last 4 actions (≤600 tokens), cosine-rank the bank, inject the top 5.</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>Retrieval is what answers the innkeeper problem: the promise is a memory, the memory has a
|
||
vector, walking into the inn produces a query vector near it, and it comes back into the prompt.</p>
|
||
|
||
<h4>The decisions inside it</h4>
|
||
|
||
<p><strong>A memory hangs off the node whose block it ends on.</strong> Not off the adventure, and
|
||
not off a <em>position</em> in a list of actions — off a <code>(branch_id, depth)</code> coordinate.
|
||
That makes “which memories described this turn?” an indexed lookup rather than a scan for rows whose
|
||
covered range has fallen off the end of the story, and it is what makes memories inherit correctly
|
||
across a fork: the ones above the fork point already sit on ancestors both lines read.</p>
|
||
|
||
<p>It is also the repair. When a turn’s text is replaced or removed — a retry, an undo, a deleted
|
||
action — <code>forget_node</code> withdraws the memory hanging off that coordinate <em>and</em>
|
||
rewinds both marks to just before the stretch it covered, so that ground is summarized again from
|
||
what the story now says. An earlier version instead held the newest action back a turn so it could
|
||
never be summarized before it stopped being retryable; that is no longer needed, because the repair
|
||
exists whether or not the invalidation happens at the tip.</p>
|
||
|
||
<p><strong>Cursors only advance on success.</strong> Every AI call here is best-effort. If
|
||
summarization fails, the function returns and the cursor is unchanged, so the same block is
|
||
retried on a later turn. There’s no retry loop, no dead-letter queue, no backoff — the cadence
|
||
<em>is</em> the retry mechanism.</p>
|
||
|
||
<p><strong>Summarization is fire-and-forget, in a background task with its own DB session.</strong>
|
||
The player’s turn is already on screen; making them wait would add a second or two of latency
|
||
every sixth turn for no visible benefit. The task holds a strong reference to itself — the event
|
||
loop only keeps weak ones, so a fire-and-forget task can otherwise be garbage-collected mid-run —
|
||
and a per-adventure guard stops two from overlapping.</p>
|
||
|
||
<p><strong>Pinned memories count toward <code>top_k</code>.</strong> Pinned ones are always
|
||
injected; unpinned fill up to <code>top_k − len(pinned)</code>. Without that, 6 pinned memories
|
||
plus <code>top_k=5</code> injects 11 and blows the budget the whole context engine exists to
|
||
respect.</p>
|
||
|
||
<p><strong>A dimension mismatch scores 0.0, it doesn’t crash.</strong> If the user changes their
|
||
embedding model, old 768-dim vectors get compared against a new 1536-dim query.
|
||
<code>zip()</code> would happily truncate and score garbage, silently. An explicit length check
|
||
returns 0.0 instead.</p>
|
||
|
||
<p><strong>Eviction is LRU-ish, and evicted memories are kept.</strong> Over capacity (default
|
||
200), the least-used unpinned memories are marked <code>forgotten</code> rather than deleted — so
|
||
the UI can still show them and you can un-forget one.</p>
|
||
|
||
<div class="trap">
|
||
<span class="tag">Money trap</span>
|
||
<p><strong>Background calls never spend the shared demo key.</strong> The summarization and
|
||
embedding providers are built directly from the user’s own settings, never from the demo config,
|
||
and their call sites are skipped when the turn is running on the demo key. Summarization is
|
||
unmetered background spend; the demo key is server-funded. Both facts together would be a bill.</p>
|
||
</div>
|
||
|
||
<h3 id="s16"><span class="h-num">1.6</span>Streaming</h3>
|
||
|
||
<p>The model produces tokens one at a time. Waiting for the whole reply before showing anything
|
||
makes a 20-second generation feel broken.</p>
|
||
|
||
<p><strong>Server-Sent Events</strong> is the mechanism: an HTTP response that stays open and
|
||
pushes <code>data: {...}</code> lines as they become available. It’s one-directional
|
||
(server → browser), which is exactly the shape of this problem — WebSockets would be a
|
||
bidirectional connection for a unidirectional need.</p>
|
||
|
||
<pre><code>model endpoint --SSE--> FastAPI --SSE--> browser --> React state --> screen</code></pre>
|
||
|
||
<p>FastAPI reads the provider’s stream and for each chunk yields
|
||
<code>data: {"type":"chunk","text":"…"}</code>. The frontend reads the response body with a
|
||
<code>ReadableStream</code> reader, buffers on <code>\n\n</code> boundaries, and dispatches each
|
||
parsed event. Event types: <code>player</code>, <code>reasoning</code> (thinking-model traces,
|
||
which stream into a separate collapsible panel with their own token budget), <code>chunk</code>,
|
||
<code>stopped</code>, <code>error</code>, <code>done</code>.</p>
|
||
|
||
<p>Two production details that only show up when hosted:</p>
|
||
|
||
<ul>
|
||
<li><code>X-Accel-Buffering: no</code> — nginx-style reverse proxies buffer responses by default,
|
||
which turns a stream into one big delivery at the end. This header tells them to flush each event.</li>
|
||
<li>The security-headers and body-size middlewares are written as <strong>pure ASGI</strong>
|
||
rather than Starlette’s <code>BaseHTTPMiddleware</code>, because the latter buffers the response
|
||
body and would break streaming.</li>
|
||
</ul>
|
||
|
||
<p><strong>The empty-reply case is diagnosed, not reported as “empty”.</strong> If a reasoning
|
||
model streams thinking but no story text, it spent its whole budget thinking — the error says so
|
||
and names the three settings that fix it.</p>
|
||
|
||
<h3 id="s17"><span class="h-num">1.7</span>The scripting sandbox</h3>
|
||
|
||
<p>Real AI Dungeon scripts are JavaScript files defining <code>modifier(text)</code> and calling it
|
||
as the last line, with globals like <code>state</code>, <code>history</code>,
|
||
<code>storyCards</code>. To be compatible, this app runs the same contract in an embedded
|
||
<strong>QuickJS</strong> interpreter.</p>
|
||
|
||
<p>The safety properties are mostly structural:</p>
|
||
|
||
<div class="scroll">
|
||
<table>
|
||
<thead><tr><th>Property</th><th>How</th></tr></thead>
|
||
<tbody>
|
||
<tr><td>No filesystem, network or process access</td><td>QuickJS has none by default — nothing was removed, nothing was added</td></tr>
|
||
<tr><td>Memory cap</td><td>16 MB per run</td></tr>
|
||
<tr><td>CPU cap</td><td>2 seconds per run</td></tr>
|
||
<tr><td>No shared state between runs</td><td>A fresh context per hook execution</td></tr>
|
||
<tr><td>A broken script can’t break a turn</td><td>Every failure returns as <code>.error</code> with text, state and cards unchanged; the pipeline logs it and continues</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>Data crosses the boundary as JSON — Python serializes <code>{state, text, history, storyCards,
|
||
info}</code> in and the results out. There is no object bridge to exploit.</p>
|
||
|
||
<p>One deliberate bug-compatibility: <code>addStoryCard</code> returns the new card’s
|
||
<em>index</em>, so the first card returns <code>0</code>, which is falsy, so
|
||
<code>if (!addStoryCard(…))</code> misfires. That’s upstream AI Dungeon’s behaviour. It’s
|
||
documented in the code and left alone, because matching real scripts is the entire point of the
|
||
feature.</p>
|
||
|
||
<h3 id="s18"><span class="h-num">1.8</span>Why there is no agent framework</h3>
|
||
|
||
<p>Graph-based agent frameworks (LangGraph and similar) earn their complexity with
|
||
<strong>branching, cyclic, multi-step control flow</strong>
|
||
— a graph of nodes where the path depends on what the model decides, with loops, retries, tool
|
||
calls, and persisted state between steps.</p>
|
||
|
||
<p>This turn pipeline is a <strong>fixed linear sequence with exactly one model call</strong>.
|
||
There is no routing decision, no tool selection, no loop. Every turn takes the same path. Adding a
|
||
graph framework would mean carrying its state abstraction, its serialization model and its
|
||
debugging surface to express a straight line.</p>
|
||
|
||
<p>There’s also a specific reason a framework’s context handling wouldn’t fit here:
|
||
<strong>the budgeting logic is the product.</strong> Buffer-window and summary-memory abstractions
|
||
are opinionated about how to fit history into a window. This app shows the user every context
|
||
component, its token cost, and the trigger word that pulled it in — which means the assembly has
|
||
to be explicit and inspectable.</p>
|
||
|
||
<p><strong>When it would be the right call:</strong> if the design went toward the
|
||
two-call version — narrate, then a separate structured-extraction step, with a retry branch when
|
||
extraction fails and a tool-calling path for dice — that is a graph, and hand-rolling it would get
|
||
ugly fast.</p>
|
||
|
||
</div>
|
||
|
||
<!-- ===================================================== PART 2 ========== -->
|
||
<section class="wrap part" id="p2">
|
||
<p class="num">Part 2</p>
|
||
<h2>Data and correctness</h2>
|
||
<p>The bugs in this section are the kind that don’t crash. They just quietly produce the wrong
|
||
answer, which is why each one has a test.</p>
|
||
</section>
|
||
|
||
<div class="wrap">
|
||
|
||
<h3><span class="h-num">2.1</span>The domain model</h3>
|
||
|
||
<pre><code>User
|
||
├─ Scenario (the template) ── stat_schema, prompt, memory, author's note
|
||
│ └─ StoryCard, Script
|
||
└─ Adventure (the playthrough) ── world_state, script_state, head_branch_id/head_depth
|
||
├─ Branch (one line of it) ── parent_branch_id, fork_depth, lineage, name
|
||
├─ Action (one node) ── branch_id, depth, parent_id, live,
|
||
│ text, context_snapshot, state_after
|
||
├─ StoryCard (its own copy)
|
||
├─ Memory (text, embedding, branch_id, depth, use_count)
|
||
└─ AdventureScript</code></pre>
|
||
|
||
<p><strong>The one decision that shapes everything: template vs instance.</strong> A scenario
|
||
declares what stats <em>exist</em>; an adventure holds what they <em>are</em> right now. Creating
|
||
an adventure copies the scenario’s story cards, scripts and plot fields into it, so editing a
|
||
scenario later never mutates a game in progress. There’s an explicit opt-in “Update from scenario”
|
||
flow for when you <em>do</em> want that, which diffs the two and shows what would change.</p>
|
||
|
||
<p>Same reasoning as instantiating a class: shared definition, independent state.</p>
|
||
|
||
<h3 id="s22"><span class="h-num">2.2</span>The story is a tree</h3>
|
||
|
||
<p>The largest structural change the project has had, and the one with the most reasoning behind
|
||
it.</p>
|
||
|
||
<h4>The problem</h4>
|
||
|
||
<p>The story used to be a list, and a mutable one. Retry rewrote the last entry in place; undo and
|
||
delete removed entries from the middle. Everything derived from the story — the memories, the
|
||
running summary, the two marks saying how far each had got — was indexed by <em>position in that
|
||
list</em>, and a position means something different after anything in front of it is deleted.</p>
|
||
|
||
<p>That single fact produced a family of bugs that all looked different:</p>
|
||
|
||
<ul>
|
||
<li>Deleting a middle action slid a never-summarized action down into the “already covered”
|
||
range, so a <em>recent</em> action silently never became a memory.</li>
|
||
<li>Discarding a memory left its actions behind the mark, describing nothing.</li>
|
||
<li>Retry rewrote an action’s text after the mark had passed it, so its memory described
|
||
narration that was no longer in the story.</li>
|
||
<li>The retried row was still attached to the adventure while its replacement was being written,
|
||
so the model was shown the attempt it was meant to replace and wrote a <em>continuation</em> of
|
||
it. That exclusion had to be threaded through four separate readers.</li>
|
||
<li>Attempts lived in a JSON array on the row with a mirrored copy of the live one in the
|
||
ordinary columns — a repeating group and a denormalisation in one.</li>
|
||
</ul>
|
||
|
||
<div class="trap">
|
||
<span class="tag">The pattern</span>
|
||
<p>Each was fixed where it was found. The shape only becomes visible when you line them up:
|
||
they are all the same bug, and it is that <strong>the story is a list nobody may reorder</strong>.</p>
|
||
</div>
|
||
|
||
<h4>The shape</h4>
|
||
|
||
<p>Make the story a tree and none of them are reachable. Every action is a <strong>node</strong>
|
||
with a <code>branch_id</code> and a <code>depth</code>. A <strong>branch</strong> is one line
|
||
through the tree: it holds the nodes played on it and <em>borrows</em> everything before its fork
|
||
point from its ancestors. Nothing is ever copied, and — apart from an explicit delete — nothing is
|
||
ever removed.</p>
|
||
|
||
<pre><code>branches(id, adventure_id, parent_branch_id, fork_depth, lineage, name)
|
||
actions(id, adventure_id, branch_id, depth, parent_id, live, text, …, state_after)
|
||
memories(…, branch_id, depth)
|
||
adventures(…, head_branch_id, head_depth)</code></pre>
|
||
|
||
<p><code>depth</code> is a position along <em>a</em> path, not a global turn number:
|
||
<code>A4</code> and <code>B4</code> are two alternatives, not two turns. Reading branch C, whose
|
||
tip is at depth 7 and which left B at 5, which left A at 3:</p>
|
||
|
||
<pre><code>SELECT * FROM actions
|
||
WHERE (branch_id = 'C')
|
||
OR (branch_id = 'B' AND depth <= 5)
|
||
OR (branch_id = 'A' AND depth <= 3)
|
||
ORDER BY depth DESC LIMIT 32</code></pre>
|
||
|
||
<p>→ <code>A0 A1 A2 A3 B4 B5 C6 C7</code>.</p>
|
||
|
||
<p><strong>Why <code>branch_id</code> + <code>depth</code> rather than parent pointers alone.</strong>
|
||
Parent pointers are the obvious way to store a tree and the wrong way to read one: reading a story
|
||
would be N round trips up a chain, which throws away the windowed history work (<a href="#s12">1.2</a>)
|
||
that made a turn’s read cost flat. Depth replaces the old <code>index</code> as the ordering key, so
|
||
the reads keep the shape they already had.</p>
|
||
|
||
<h4>The lineage, and why fork count costs nothing</h4>
|
||
|
||
<p>The OR-clause above is not reconstructed per read. It is stored on the branch row as
|
||
<code>lineage</code> — <code>[(C, ∞), (B, 5), (A, 3)]</code> — computed once when the fork happens,
|
||
from the parent’s lineage plus one entry. One module knows how to turn it into a query, which is
|
||
deliberate: one forgotten clause shows the wrong story and reports nothing.</p>
|
||
|
||
<p>Two properties of the shape do the real work:</p>
|
||
|
||
<ul>
|
||
<li><strong>The ranges are disjoint and descending.</strong> A branch’s own nodes always sit
|
||
deeper than its fork point, and each ancestor is capped at the fork depth of the branch beneath
|
||
it. So ordering the whole clause by <code>depth DESC</code> reads entry 0’s nodes, then entry 1’s,
|
||
then entry 2’s — which lets a tail read use the newest few entries and stop.</li>
|
||
<li><strong>Clause count is bounded by the context window, not by fork count.</strong> A 200-fork
|
||
story whose newest branch is 40 turns long reads with <em>one</em> clause, because the window is
|
||
covered before the second entry is reached.</li>
|
||
</ul>
|
||
|
||
<div class="stats">
|
||
<div class="stat"><div class="v">1.007×</div><div class="k">page load of a 40-turn story forked twenty times, against the same story flat — 31,652 B vs 31,433 B</div></div>
|
||
<div class="stat"><div class="v">~103 B</div><div class="k">what one branch costs: an id, a parent, a fork depth and a cached ancestry. No copy, no migration, no vacuum.</div></div>
|
||
</div>
|
||
|
||
<h4>What a player actually does</h4>
|
||
|
||
<p>None of the above is what the screen shows. In the player’s words:</p>
|
||
|
||
<blockquote>
|
||
<p>Any turn can gain another <strong>take</strong>. On an AI turn that means regenerate; on your
|
||
own message it means type something else. Stepping between takes with <code>‹ 2/4 ›</code> is free
|
||
— the story below simply empties, because that take has no children yet. <strong>A branch is
|
||
created when you write below a take that is not the live one</strong>, never before.</p>
|
||
</blockquote>
|
||
|
||
<p>That rule collapses two operations into one and deletes a distinction from the UI. The first
|
||
version of this screen had a chip that <em>switched</em> at the tip and only <em>previewed</em>
|
||
above it, with a second button to take that line — one control whose meaning depended on where the
|
||
reader was standing. The rule above replaced it with a pager that only ever steps, a fork button on
|
||
every turn, and no tip-versus-past distinction at all. The distinction survives in the
|
||
implementation, where it decides whether a write needs a branch: at the tip the attempts are still
|
||
leaves nobody has built on, so taking one is a switch and no branch is created.</p>
|
||
|
||
<h4>Takes are grouped by parent, not by coordinate</h4>
|
||
|
||
<p>The load-bearing detail, and the one that isn’t obvious. The natural way to find “the other
|
||
takes of this turn” is by coordinate — same branch, same depth. It is wrong in both directions:</p>
|
||
|
||
<pre><code>B ── C C1 C2 <- three takes, one parent (B)
|
||
│ └── D1' D2' <- two takes, parent C2
|
||
└── D1 D2 D3 <- three takes, parent C1</code></pre>
|
||
|
||
<p>Standing on the C2 path at that depth must read <code>2/2</code>, not <code>5</code>. Coordinate
|
||
grouping gets that one right by accident, because writing under a non-live take forks and the two
|
||
sets land on different branches. It gets <code>C</code> wrong: once C has been forked onto a branch
|
||
of its own it is alone at its coordinate and reads <code>1/1</code>, having lost C1 and C2 from a
|
||
pager that must still say <code>1/3</code>.</p>
|
||
|
||
<p>So a node carries <code>parent_id</code>, read for nothing but this. The alternative — making a
|
||
branch’s fork point a <em>node</em> rather than a depth, so a promoted take never moves — was
|
||
rejected: the whole point of <code>lineage</code> is that a read is an OR-clause per branch instead
|
||
of a walk up parent pointers, and re-pointing the fork at a node changes path resolution itself,
|
||
dragging in the cursors, memory depths and both bundle formats. <code>parent_id</code> is one
|
||
indexed lookup, never a walk, and nothing about how a path resolves changes.</p>
|
||
|
||
<h4>Cursors become anchors</h4>
|
||
|
||
<p>The two marks — how far the memory bank has got, how far the summary has got — used to be
|
||
counts. A count is a position in a list, and every rule about sliding them, rewinding them and
|
||
translating between positions and <code>Action.index</code> existed to patch up the fact that the
|
||
list moves.</p>
|
||
|
||
<p>A cursor is now an <strong>anchor</strong>: <code>(branch_id, depth)</code>, the node up to and
|
||
including which the work is done. Deleting an action doesn’t move it, because a depth is a
|
||
coordinate along a path rather than a slot in a list. “What is not covered yet” becomes a question
|
||
about the story instead of about a list index, and it answers correctly whatever has been deleted
|
||
in front of it. The branch half is what makes it survive forking: a depth alone is ambiguous once
|
||
two branches both have a node 41.</p>
|
||
|
||
<p><code>position_of_index</code>, <code>note_action_removed</code>,
|
||
<code>settled_story_actions</code> and the cursor-rewind machinery were <strong>deleted</strong>,
|
||
not left unused.</p>
|
||
|
||
<h4>Derived work attaches to the node that produced it</h4>
|
||
|
||
<p>Generalise the rule and a lot falls out: <em>anything derived hangs off the node that produced
|
||
it</em>. A memory covering depths 37–42 hangs off that branch’s node 42 and is invisible to any path
|
||
that doesn’t run through it. Shared ancestors are therefore shared automatically, so <strong>a fork
|
||
needs nothing recreated</strong> — the memories above the fork point are already on the ancestors
|
||
both lines read.</p>
|
||
|
||
<div class="trap">
|
||
<span class="tag">The subtle case</span>
|
||
<p>The memory sitting <em>at</em> the forked coordinate. The first cut moved it onto the new
|
||
branch and re-anchored the marks naming it. Both are wrong for the same reason: that memory
|
||
describes whichever attempt was live at that coordinate, which is the one staying on the parent.</p>
|
||
<p>The right answer needs no code. The lineage caps the parent one depth short of the fork, so
|
||
the memory is simply out of range from the new branch — invisible to both the retrieval clause and
|
||
the anchor read. The new line summarizes that ground again, from the text it actually tells.</p>
|
||
</div>
|
||
|
||
<p>Hand-written memories obey the same rule. One used to carry a NULL depth, described as “belongs
|
||
to the adventure rather than to a path” — which sounds harmless and is not: a NULL is a coordinate
|
||
no fork can cap, so a note typed on one line followed the reader onto branches whose events it never
|
||
described. They are anchored at the head instead: <em>the story you were reading when you wrote
|
||
it</em>.</p>
|
||
|
||
<h4>Deleting, and why the branch UI was a hard dependency</h4>
|
||
|
||
<p>Nothing is ever auto-pruned. That is the guarantee the whole design rests on, and it is also why
|
||
branch management couldn’t be a nice-to-have: without a way to delete a line, storage grows without
|
||
limit.</p>
|
||
|
||
<p>The delete rule has two halves and the second is easy to miss. Refusing to delete the line being
|
||
read is obvious. The other half is refusing any line it was <strong>forked from</strong> —
|
||
<code>parent_branch_id</code> cascades, so deleting an ancestor takes the head with it and leaves
|
||
<code>head_branch_id</code> pointing at a row that is gone. One membership test against the head’s
|
||
own lineage covers both, because a lineage already names itself and every branch it borrows from.
|
||
The server is the authority; the client computes the same set only so a button can say so before it
|
||
is pressed.</p>
|
||
|
||
<h4>The migration, and what it deliberately did not do</h4>
|
||
|
||
<p>There is no feature flag. <strong>A linear story is a tree with one branch</strong>, so the
|
||
intermediate states weren’t half-migrated — they were the same product with a superset schema
|
||
underneath, which made “existing adventures are unaffected” a literal, testable pass condition at
|
||
every step. A flag would have bought two live code paths through the context builder, the memory
|
||
bank, undo and retry at once.</p>
|
||
|
||
<p>The legacy columns (<code>index</code>, <code>variants</code>, <code>variant_index</code>, the
|
||
two <code>*_before</code> snapshots) were kept unread for a release rather than dropped with the
|
||
migration that stopped using them, so that redeploying the previous build is still a way out.
|
||
Dropping columns is the one step that isn’t.</p>
|
||
|
||
<div class="trap">
|
||
<span class="tag">Operational, and it generalises</span>
|
||
<p>On Postgres a migration that rewrites every row of <code>actions</code> roughly doubles the
|
||
table, and only <code>VACUUM FULL</code> gives it back — 79 MB reclaimed in 5.5 s on one occasion.
|
||
But bloat scales with the <strong>heap</strong>, and <code>context_snapshot</code> is 94% of this
|
||
table and lives out of line, so a migration touching only small columns reuses the existing TOAST
|
||
pointer and costs a tenth of that.</p>
|
||
<p>Read the sizes from <code>sum(octet_length(col))</code> per column, not from
|
||
<code>n_live_tup</code> — that one is a stale estimate in exactly the direction that makes bloat
|
||
look smaller.</p>
|
||
</div>
|
||
|
||
<h4>What this is honest about</h4>
|
||
|
||
<ul>
|
||
<li><strong>The two marks are one pair on the adventure</strong>, not one per branch. Switching
|
||
lines makes the mark on the line being left unreadable from the new one, and that ground is
|
||
summarized again. It answers “nothing covered”, which is the safe direction — redo the work, never
|
||
skip it — but switching back and forth costs AI calls.</li>
|
||
<li><strong>Story cards stay adventure-wide.</strong> A card invented on branch B shows on branch
|
||
A. Event-sourcing card changes onto nodes was considered and rejected.</li>
|
||
<li><strong>Editing an already-summarized action still leaves its memory stale.</strong> The
|
||
machinery to fix it now exists — an edit could write a sibling take and switch to it, which is a
|
||
retry the player typed — but it doesn’t do that yet.</li>
|
||
</ul>
|
||
|
||
<h3 id="s23"><span class="h-num">2.3</span>Undo and retry that actually rewind</h3>
|
||
|
||
<p>Most implementations of undo delete the last message. That’s wrong here, because a turn mutates
|
||
three things: the text, the scripting scoreboard, and the RPG stats.</p>
|
||
|
||
<p><strong>The mechanism:</strong> every node carries <code>state_after</code> and
|
||
<code>world_state_after</code> — deep copies of what the adventure looked like once that turn had
|
||
played. Rewinding to before a turn is a read of the node in front of it, so undo, retry and a branch
|
||
switch are the same restore. The cooldown clock comes along for free: it lives inside the world
|
||
state, so each line of the story carries its own without anything having to know there is one.</p>
|
||
|
||
<p><strong>Nothing a retry replaces is thrown away.</strong> The old attempt stays as another
|
||
<em>take</em> of that turn — a sibling node at the same coordinate, <code>live</code> false — and the
|
||
pager steps between them. Which is to say retry isn’t a special case: it is the tree, with the branch
|
||
not yet created (<a href="#s22">2.2</a>).</p>
|
||
|
||
<p>Four details that are easy to get wrong:</p>
|
||
|
||
<ul>
|
||
<li><strong>The turn being retried is excluded from its own context.</strong> Its takes are still
|
||
attached to the adventure, so without an explicit exclusion the model would be shown the attempt
|
||
it’s replacing as established story — and would write a continuation of it. The exclusion had
|
||
leaked into four readers, not one: history replay, story-card trigger matching, in-scene NPC
|
||
detection, and the memory-bank similarity query. <em>Anything reading the story during generation
|
||
takes the exclusion.</em></li>
|
||
<li><strong>A retry reuses the turn’s depth</strong>, not the next one. Cooldowns are measured
|
||
along the path, so allocating a new depth would advance the clock the cooldown rules run on and a
|
||
retry would quietly unlock stats that should still be waiting.</li>
|
||
<li><strong><code>delete_turn</code> used to mean “every take at this coordinate”.</strong> Once a
|
||
take can be forked onto a branch of its own the group spans branches, and undo reached across and
|
||
deleted a take belonging to a line nobody asked about. Anything that reads a take group and then
|
||
<em>writes</em> has to say whether it means the turn or the coordinate.</li>
|
||
<li><strong>If the regeneration fails, the rollback is reversed.</strong> The generator is wrapped
|
||
in a <code>try/finally</code>: if it ends without saving — provider error, empty reply, a script
|
||
<code>stop</code>, or the browser hanging up — the previous take is put back in charge. Otherwise
|
||
the state on the server drifts from the text still on the user’s screen.</li>
|
||
</ul>
|
||
|
||
<h3><span class="h-num">2.4</span>The turn lock</h3>
|
||
|
||
<p>One turn at a time per adventure. Double-clicking “Continue” must not run two generations.</p>
|
||
|
||
<p>The subtlety: the check has to happen in the <strong>request phase</strong>, not when the SSE
|
||
generator first runs. A streaming response doesn’t start iterating its generator until the response
|
||
begins, so a check inside the generator lets two rapid requests both pass before either claims the
|
||
slot. And because sync FastAPI endpoints run in a threadpool, the test-and-set needs a real lock.</p>
|
||
|
||
<pre><code>def acquire_turn_lock(adventure_id): # in the request handler
|
||
with _active_turns_guard:
|
||
if adventure_id in _active_turns:
|
||
raise HTTPException(409, "A turn is already generating…")
|
||
_active_turns.add(adventure_id)
|
||
|
||
async def with_turn_lock(adventure_id, gen): # wraps the SSE generator
|
||
try:
|
||
async for event in gen: yield event
|
||
finally:
|
||
_active_turns.discard(adventure_id)</code></pre>
|
||
|
||
<p>In-memory, so it’s a single-process guarantee. That’s honest for the deployment this targets —
|
||
one Render web service. Two processes would need the lock in the database.</p>
|
||
|
||
<h3 id="s25"><span class="h-num">2.5</span>The 189× egress fix</h3>
|
||
|
||
<div class="stats">
|
||
<div class="stat"><div class="v">38.5 MB → 0.20 MB</div><div class="k">database egress for one adventure load</div></div>
|
||
<div class="stat"><div class="v">~74 KB</div><div class="k">per-turn prompt snapshot — 94% of the database</div></div>
|
||
</div>
|
||
|
||
<p><strong>The bug:</strong> <code>Action.context_snapshot</code> holds the entire assembled prompt
|
||
for a turn. Every adventure load pulled that column for every action, to read two small fields out
|
||
of it — the world-state delta for the “what changed” chip, and the applied report. SQLAlchemy loads
|
||
all columns by default.</p>
|
||
|
||
<p><strong>The fix, in three parts:</strong></p>
|
||
|
||
<ol>
|
||
<li>Move the two small things that <em>are</em> needed for every action into their own column.</li>
|
||
<li>Mark the heavy columns <code>deferred</code> — snapshot, variants, reasoning — so they’re only
|
||
fetched when explicitly asked for.</li>
|
||
<li>Backfill the new column with dialect-specific server-side SQL, so the old data is extracted
|
||
inside the database and never crosses the wire.</li>
|
||
</ol>
|
||
|
||
<p><strong>The part that makes it stick:</strong> <code>tests/test_egress.py</code> hooks into
|
||
SQLAlchemy’s <code>before_cursor_execute</code> event, captures every statement the ORM sends, and
|
||
fails if a bulk load ever names those columns again. The regression is caught by asserting on the
|
||
<em>SQL</em>, not on a timing.</p>
|
||
|
||
<p>One more detail from that test’s design: the count query is written as a real
|
||
<code>SELECT count(…)</code> rather than <code>query.count()</code>, because SQLAlchemy’s
|
||
<code>.count()</code> wraps the entity select in a subquery whose SQL names every column —
|
||
including the deferred ones. No bytes come back either way, but the database still reads them, and
|
||
a guard that greps SQL can’t tell the two apart.</p>
|
||
|
||
<p>There’s a companion denormalization for the same reason: the pager has to know how many takes a
|
||
turn has without fetching any of them, so <code>variant_index</code> and <code>variant_count</code>
|
||
are cached on the row and refreshed by exactly one function, precisely so they can’t drift and the
|
||
pager can’t lie. <code>variant_count</code> is 0 rather than 1 for a turn nobody retried, because
|
||
the question it answers is “is there anything to page through?”</p>
|
||
|
||
<h3><span class="h-num">2.6</span>Migrations, hand-rolled</h3>
|
||
|
||
<p>No Alembic. An append-only list of <code>(version, SQL)</code> pairs, with the current version
|
||
stored in SQLite’s <code>PRAGMA user_version</code> or a one-row table on Postgres. 64 versions so
|
||
far.</p>
|
||
|
||
<ul>
|
||
<li>A <strong>fresh</strong> database is created by <code>create_all()</code> — always current —
|
||
and stamped at the latest version. It never replays history.</li>
|
||
<li>An <strong>existing</strong> database runs every migration above its stored version, in order.</li>
|
||
</ul>
|
||
|
||
<p>Why this and not Alembic: for a single-file SQLite app someone may have been running for months,
|
||
the entire requirement is “add a column, don’t lose their data”. Alembic’s autogenerate, branching
|
||
and down-migrations are machinery for a team with a staging environment. This is 250 lines and you
|
||
can read all of it.</p>
|
||
|
||
<p>The constraint it creates is written at the top of the file: change <code>models.py</code> so
|
||
fresh databases are current, <em>and</em> append a pair here so existing ones upgrade. Migrations
|
||
2–23 predate Postgres support and use SQLite-only syntax — harmless, because every Postgres
|
||
database starts fresh and never replays them, but anything added since must run on both dialects.</p>
|
||
|
||
<p>One migration worth reading (repairing duplicate action indexes) uses <code>UPDATE … FROM</code>
|
||
with a window function rather than a correlated subquery, because SQLite may evaluate a correlated
|
||
subquery against partially-updated rows and produce duplicates again while “repairing” them.</p>
|
||
|
||
</div>
|
||
|
||
<!-- ===================================================== PART 3 ========== -->
|
||
<section class="wrap part" id="p3">
|
||
<p class="num">Part 3</p>
|
||
<h2>Production concerns</h2>
|
||
<p>What changes when the app stops being yours and starts being a URL strangers can open.</p>
|
||
</section>
|
||
|
||
<div class="wrap">
|
||
|
||
<h3><span class="h-num">3.1</span>Two modes, one codebase</h3>
|
||
|
||
<p><code>AIDND_MULTI_USER</code> switches the whole app between two personalities:</p>
|
||
|
||
<div class="scroll">
|
||
<table>
|
||
<thead><tr><th></th><th>Local (default)</th><th>Hosted</th></tr></thead>
|
||
<tbody>
|
||
<tr><td>Users</td><td>One auto-created local user</td><td>Guest on first visit, optional account</td></tr>
|
||
<tr><td>Auth</td><td>None — no cookies, no login UI</td><td>Signed session cookie</td></tr>
|
||
<tr><td>Rate limits</td><td>Off</td><td>On</td></tr>
|
||
<tr><td>Row caps</td><td>Off</td><td>On</td></tr>
|
||
<tr><td>API docs</td><td>On</td><td>Off</td></tr>
|
||
<tr><td>Provider</td><td>Whatever Settings points at</td><td>User’s key, or the shared demo key</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>The reasoning: someone running this on their own laptop should never be throttled by their own
|
||
app, never see a login screen, and should get the interactive API docs. A hosted deployment needs
|
||
all four to be the opposite. Rather than two builds, the differences are gated at each site.</p>
|
||
|
||
<p><strong>Guests upgrade in place.</strong> A visitor gets a guest <code>User</code> row on first
|
||
load. Registering sets <code>email</code> and <code>password_hash</code> on that <em>same row</em>
|
||
— so every adventure they played as a guest survives with no re-parenting and no migration step.
|
||
Three kinds of row share the users table: local, guest, and registered.</p>
|
||
|
||
<p><strong>Guests expire; accounts don't.</strong> One row per curious visitor adds up, so
|
||
<code>cleanup.py</code> deletes guests idle for <code>AIDND_GUEST_RETENTION_DAYS</code>
|
||
(default 5) — measured as <code>COALESCE(last_seen_at, created_at)</code>, since
|
||
<code>_touch</code> only writes <code>last_seen_at</code> hourly and a freshly minted guest
|
||
has NULL until its second request. The filter requires both <code>is_guest</code>
|
||
<em>and</em> <code>email IS NULL</code>, so upgrading in place is also how you opt out of
|
||
expiry. It sweeps once at startup — the reliable trigger on a host that sleeps — and then
|
||
every few hours.</p>
|
||
|
||
<p>It's a single Core <code>DELETE</code>, not <code>db.delete(user)</code>: the ORM path
|
||
would SELECT every adventure, action and memory into Python purely to delete them, and the
|
||
foreign keys are <code>ON DELETE CASCADE</code> from <code>users</code> all the way down, so
|
||
the database does the whole graph in one statement. Nothing a guest owns is visible to
|
||
anyone else either — <code>is_public</code> is output-only, so shared content is exactly the
|
||
seeded scenarios, which have <code>user_id NULL</code> and never match the filter.</p>
|
||
|
||
<h3><span class="h-num">3.2</span>The shared demo key</h3>
|
||
|
||
<p>The demo lets people play with no signup and no API key, on a key the server pays for. That is a
|
||
spending surface, so it’s the most defended code in the project.</p>
|
||
|
||
<p>One function makes the BYOK-vs-demo decision, and on the demo branch it pins <strong>two</strong>
|
||
things:</p>
|
||
|
||
<ul>
|
||
<li><strong>The model</strong> — to a whitelist. A caller-supplied override or a hand-edited
|
||
settings row can’t aim a server-funded key at an expensive model. Anything unrecognised falls
|
||
back to the first whitelisted model.</li>
|
||
<li><strong>The endpoint</strong> — to the configured demo URL. Otherwise the key could be
|
||
redirected to a URL the user controls and harvested.</li>
|
||
</ul>
|
||
|
||
<p>Plus a daily per-user turn cap (default 20), checked <em>before</em> the player’s input is stored
|
||
so a capped player doesn’t get their message saved with no reply, and counted only after a
|
||
successful turn.</p>
|
||
|
||
<div class="trap">
|
||
<span class="tag">A real bug, recorded in a comment</span>
|
||
<p>There’s a defensive check that raises if a demo config somehow carries a non-whitelisted
|
||
model. It tests <code>using_demo</code>, <strong>not</strong>
|
||
<code>api_key == DEMO_API_KEY</code>. Keying on the key value looks stricter but is wrong — the
|
||
demo key is an ordinary OpenRouter key, so a user can legitimately paste that same key into their
|
||
own settings as BYOK, and then every resolution raised, 500ing even <code>GET /auth/me</code> and
|
||
taking the whole SPA down. <code>using_demo</code> is what actually means “the server is paying”.</p>
|
||
</div>
|
||
|
||
<h3><span class="h-num">3.3</span>Secrets</h3>
|
||
|
||
<p>Everything derives from one server-side secret.</p>
|
||
|
||
<div class="scroll">
|
||
<table>
|
||
<thead><tr><th>Thing</th><th>Mechanism</th></tr></thead>
|
||
<tbody>
|
||
<tr><td>Passwords</td><td><code>hashlib.scrypt</code>, N=2¹⁴, r=8, p=1, per-password salt, constant-time compare. Stdlib, so no extra dependency.</td></tr>
|
||
<tr><td>Sessions</td><td><code>v1.<user_id>.<HMAC-SHA256></code>, no expiry — long-lived guest sessions are the point. A cookie can outlive a swept guest row; that resolves to a 401, which the frontend already turns into a fresh session.</td></tr>
|
||
<tr><td>Stored LLM API keys</td><td>Fernet encryption at rest, key derived from the secret, <code>enc:</code> prefix so legacy plaintext rows are recognisable and migratable.</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>The secret auto-generates into a file next to the database for local installs (zero config), but
|
||
<strong>multi-user mode refuses to start without the env var</strong> — with an error that explains
|
||
why and gives the command to generate one. Hosted filesystems are ephemeral; a regenerated secret
|
||
on every deploy would silently log out every user and orphan their stored API keys.</p>
|
||
|
||
<p>A rotated secret makes stored keys undecryptable. Decryption treats that as “unset” rather than
|
||
raising, so the user just re-enters their key instead of hitting a 500.</p>
|
||
|
||
<h3><span class="h-num">3.4</span>Abuse guards</h3>
|
||
|
||
<div class="scroll">
|
||
<table class="nums">
|
||
<thead><tr><th>Guard</th><th>Limit</th></tr></thead>
|
||
<tbody>
|
||
<tr><td>Turn generation</td><td>10 / min</td></tr>
|
||
<tr><td>Auth attempts (per IP)</td><td>10 / 5 min</td></tr>
|
||
<tr><td>Guest creation (per IP)</td><td>30 / 5 min</td></tr>
|
||
<tr><td>Script test runs</td><td>30 / min</td></tr>
|
||
<tr><td>Connection test</td><td>10 / min</td></tr>
|
||
<tr><td>Adventures / scenarios / scripts per user</td><td>100 / 200 / 200</td></tr>
|
||
<tr><td>Actions per adventure</td><td>5,000</td></tr>
|
||
<tr><td>Request body</td><td>2 MB (20 MB on import)</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>Rate limits are keyed per user when one is known — accounts survive IP changes — and per IP
|
||
otherwise, in fixed windows held in memory, with a pruning pass so the per-IP dict can’t grow
|
||
without bound. Import endpoints check bundle list lengths against the same caps live creation
|
||
enforces, otherwise the cap is trivially bypassed by uploading a file.</p>
|
||
|
||
<p>Security headers on every response: <code>nosniff</code>, <code>X-Frame-Options: DENY</code>,
|
||
<code>Referrer-Policy: same-origin</code>, and a CSP allowing exactly what the SPA uses.</p>
|
||
|
||
<h3><span class="h-num">3.5</span>Deployment</h3>
|
||
|
||
<p>One Docker web service on Render, serving the SPA and the API same-origin, with Postgres on Neon.</p>
|
||
|
||
<p>The Postgres decision was forced: Render’s free tier has no persistent disk, so a SQLite file
|
||
wouldn’t survive a deploy. The database lives off-box.</p>
|
||
|
||
<p>Two things worth knowing about the free tier:</p>
|
||
|
||
<ul>
|
||
<li>The service <strong>sleeps after ~15 minutes idle</strong>, and the first request then takes
|
||
30–60 seconds.</li>
|
||
<li><code>/api/health</code> deliberately <strong>doesn’t touch the database</strong>, so a
|
||
keep-warm pinger wakes the web service without waking the database. Waking a database around the
|
||
clock costs far more than the cold start is worth.</li>
|
||
</ul>
|
||
|
||
<p>CI runs the backend tests, the frontend lint and build, and a Docker image build on every push.</p>
|
||
|
||
</div>
|
||
|
||
<!-- ===================================================== PART 4 ========== -->
|
||
<section class="wrap part" id="p4">
|
||
<p class="num">Part 4</p>
|
||
<h2>The web plumbing, briefly</h2>
|
||
<p>The parts that are just how the web works, not decisions.</p>
|
||
</section>
|
||
|
||
<div class="wrap">
|
||
|
||
<p><strong>Frontend and backend are two programs.</strong> In development they’re two servers —
|
||
Vite on 5173 serving React, FastAPI on 8000 serving the API — and Vite proxies <code>/api</code> to
|
||
FastAPI so the browser thinks it’s all one origin, which avoids CORS entirely. In production
|
||
there’s one server: FastAPI serves the built React files as static assets from the same port.</p>
|
||
|
||
<p><strong>SPA routing.</strong> React Router handles URLs like <code>/play/3</code> in the browser
|
||
without a round trip. But if you <em>reload</em> that URL, the browser asks the server for
|
||
<code>/play/3</code>, which isn’t a file. So the static-file handler catches the 404 and returns
|
||
<code>index.html</code>, letting React take over and read the URL itself. API routes are matched
|
||
before the static mount, so they’re unaffected.</p>
|
||
|
||
<p><strong>Sessions.</strong> A cookie is a small value the browser stores and automatically
|
||
attaches to every request to that site. Here it holds <code>v1.<user_id>.<signature></code>.
|
||
The server doesn’t store sessions anywhere — it re-verifies the signature on each request, which is
|
||
why there’s no session table.</p>
|
||
|
||
<p><strong>The 401 retry.</strong> If the cookie is missing or stale, any API call returns 401. The
|
||
frontend catches that once, calls <code>/api/auth/me</code> — which mints a fresh guest session —
|
||
and retries the original request. So a returning visitor with an expired cookie never sees an
|
||
error.</p>
|
||
|
||
<p><strong>React, in one paragraph.</strong> A component is a function that returns a description
|
||
of some UI. <code>useState</code> holds a value; changing it re-renders the component. The
|
||
streaming turn is the clearest example: each SSE chunk appends to a state string, React re-renders,
|
||
and the text appears to type itself.</p>
|
||
|
||
</div>
|
||
|
||
<!-- ===================================================== PART 5 ========== -->
|
||
<section class="wrap part" id="p5">
|
||
<p class="num">Part 5</p>
|
||
<h2>Results and limitations</h2>
|
||
<p>What was measured, and what this design knowingly does not do.</p>
|
||
</section>
|
||
|
||
<div class="wrap">
|
||
|
||
<h3><span class="h-num">5.1</span>Measured results</h3>
|
||
|
||
<div class="scroll">
|
||
<table class="nums">
|
||
<tbody>
|
||
<tr><td>Database egress per adventure load</td><td>38.5 MB → 0.20 MB (~189×)</td></tr>
|
||
<tr><td>Prompt snapshot size</td><td>~74 KB/turn, 94% of the DB</td></tr>
|
||
<tr><td>Turn read cost at turn 200</td><td>839 KB → 129 KB, flat after ~turn 50</td></tr>
|
||
<tr><td>Cost of a branch</td><td>~103 B; 20 forks load at 1.007× the same story flat</td></tr>
|
||
<tr><td>Length-hint phrasing</td><td>174 → 246 words as a budget; 170 as a ceiling (n=5)</td></tr>
|
||
<tr><td>Backend tests</td><td>440, LLM mocked, real QuickJS engine</td></tr>
|
||
<tr><td>Schema versions</td><td>64</td></tr>
|
||
<tr><td>Sandbox limits</td><td>16 MB, 2 s CPU, fresh context per run</td></tr>
|
||
<tr><td>Context defaults</td><td>author’s note at depth 3; cards ≤ 40% of elastic budget</td></tr>
|
||
<tr><td>Memory cadence</td><td>memory / 6 turns, summary / 15 turns, top-5 retrieval</td></tr>
|
||
</tbody>
|
||
</table>
|
||
</div>
|
||
|
||
<p>Two of the tests encode a performance property rather than a behaviour:
|
||
<code>test_egress.py</code> asserts on the SQL the ORM emits, and
|
||
<code>test_history_window.py</code> asserts that the read cost stops growing with story
|
||
length.</p>
|
||
|
||
<h3><span class="h-num">5.2</span>Known limitations</h3>
|
||
|
||
<p>Deliberate trades for a single-user-first app that also happens to be hosted, listed so
|
||
nobody has to discover them the hard way.</p>
|
||
|
||
<ul>
|
||
<li><strong>Single process.</strong> The turn lock, the rate limiter and the summarization task
|
||
all assume one worker. A second worker would need the lock in the database — a row-level
|
||
advisory lock — and the rate limiter in Redis.</li>
|
||
<li><strong>No vector index.</strong> Retrieval does cosine similarity in Python over the whole
|
||
bank. Fine at the 200-memory cap; at 10,000 it would want pgvector.</li>
|
||
<li><strong>Prompt snapshots are heavy</strong> even after the egress fix — they’re deferred, not
|
||
smaller. Compressing them or expiring old ones is the real fix.</li>
|
||
<li><strong>In-memory rate-limit windows reset on restart</strong>, so a restart grants a brief
|
||
extra allowance.</li>
|
||
<li><strong>Background summarization is a fire-and-forget asyncio task</strong>, so it does not
|
||
survive a restart. At real load it belongs in a queue.</li>
|
||
<li><strong>The demo key depends on a free-tier provider’s daily cap</strong>, which the app can
|
||
only detect after the fact by string-matching the 429 body.</li>
|
||
<li><strong>The two memory marks are one pair on the adventure</strong>, not one per branch.
|
||
Switching lines makes the mark on the line being left unreadable from the new one, so that ground
|
||
is summarized again. It fails in the safe direction — redo, never skip — but switching back and
|
||
forth costs AI calls.</li>
|
||
<li><strong>Story cards are adventure-wide</strong>, so a card invented on one branch shows on all
|
||
of them.</li>
|
||
<li><strong>Editing an already-summarized turn leaves its memory stale.</strong> Replacing a turn
|
||
withdraws what was derived from it; editing one in place does not.</li>
|
||
</ul>
|
||
|
||
<h3><span class="h-num">5.3</span>Cleanup backlog</h3>
|
||
|
||
<p><code>docs/self-review.md</code> carries an open list of non-bugs — reuse, simplification and
|
||
efficiency items — kept deliberately separate from the correctness list, which is empty. The
|
||
largest ones:</p>
|
||
|
||
<ul>
|
||
<li><code>Section.tokens</code> is uncached, so the context gets tokenized two or three times a
|
||
turn.</li>
|
||
<li><code>onModelContext</code> flattens system and story into one string before handing it to
|
||
user scripts; if a script modifies it, the structure is gone and everything ships as user
|
||
content. Passing structure through the hook would be better but would break AI Dungeon
|
||
compatibility, which is the point of the feature.</li>
|
||
<li>The import endpoints hand-coerce raw dicts instead of using Pydantic bundle schemas.</li>
|
||
<li>The legacy pre-tree columns (<code>index</code>, <code>variants</code>,
|
||
<code>variant_index</code>, and the two <code>*_before</code> snapshots) are still on
|
||
<code>actions</code>, unread, kept for one release so redeploying the previous build remains a way
|
||
out. Dropping them is a migration that rewrites every row, so it owes a
|
||
<code>VACUUM FULL actions;</code> after it.</li>
|
||
</ul>
|
||
|
||
</div>
|
||
|
||
<footer>
|
||
<div class="wrap">
|
||
<p><strong>AI D&D</strong> — design notes ·
|
||
<a href="https://github.com/parththakkar106/AI-DnD">Source on GitHub</a> ·
|
||
<a href="https://parththakkar106.github.io/AI-DnD/">Project page</a></p>
|
||
<p>Full source for every claim here is in the repo; the file paths are named inline.</p>
|
||
</div>
|
||
</footer>
|
||
|
||
</main>
|
||
|
||
<script>
|
||
(function () {
|
||
var bar = document.getElementById('progress');
|
||
var tick = function () {
|
||
var h = document.documentElement;
|
||
var max = h.scrollHeight - h.clientHeight;
|
||
bar.style.width = (max > 0 ? (h.scrollTop / max) * 100 : 0) + '%';
|
||
};
|
||
addEventListener('scroll', tick, { passive: true });
|
||
addEventListener('resize', tick);
|
||
tick();
|
||
})();
|
||
</script>
|
||
|
||
</body>
|
||
</html>
|