The pager could only step between attempts, and stepping has nothing to say
about the thing the tree exists for: taking a path the story moved past and
keeping both. Every attempt has been its own node since SP4, so a chip is a
node now, and "take this path" forks — or simply switches, when the turn is
still the tip and its attempts are leaves nobody has built on.
Beside it, a Branches panel: every line the story has taken, with where each
left its parent and where it ends, and switch, rename and delete-with-confirm.
It sits with Plot/Memory/Scripts/Insights rather than inventing a new place to
put a rail. An unnamed branch is drawn from its fork depth, never from its
position in the list — a position shifts the moment a branch above it goes.
A spatial per-node map was considered and deliberately not built. At the size
this has to be verified against it is a second windowing problem, and it can be
added later without a new endpoint, since the rail and a map read the same
GET /branches. VariantOut grows an id because a fork is addressed by the node
being taken, not by an ordinal in a group that renumbers.
Driven by hand against the 602-action fixture, which found one bug that no test
could: the panel refreshed on actions.length, and a fork swaps a 60-action
window for another 60-action window, so it went on drawing a one-branch tree
while the story was already on the second. It keys off the counter adoptWindow
bumps now.
The scroll path was driven at the same time — three prepends of ~16,200 px, the
same node holding viewport top 792 to 787, never thrown to the end. That closes
the standing gap in this project. Console clean.
396 tests, build clean, no new lint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
SP7 needs three branch operations and SP5 built one. Switching exists;
naming and deleting had no column and no route between them.
A name is stored because a player chose it. An unnamed branch keeps NULL
rather than a generated "branch 4" — a generated label is derived, and it
would go stale the moment a branch before it is deleted and the ordinals
shift underneath. The client draws those from the fork depth, which
nothing can shift. The v2 bundle carries the name for the same reason it
carries the fork points and leaves `lineage` out: it is a decision, not
something computed from one.
Delete is what stands between a tree and unbounded growth, since nothing
prunes one on its own. It refuses two branches: the root, which holds the
turns every other branch borrows, and the one being read — including any
branch the head was forked from, which is the same mistake in disguise and
the one that would cascade the head away and leave head_branch_id pointing
at nothing. The nodes, memories and descendants go through the foreign
keys that already cascade.
A cursor standing on a deleted branch is cleared. On Postgres a stale
branch id would simply never resolve; SQLite hands the freed id to the
next fork, and then the anchor resolves onto a branch it has never seen
and calls a stretch of story already summarized.
396 tests, 15 new.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
A bundle had one list and a forked adventure has two stories, so export was
emitting every branch's turns interleaved by index — a mangled story rather
than lost data, and unreachable only because forking has no UI yet.
`ai-dnd-adventure-v2` carries the branches, the depth each one left its
parent at, which attempt at every turn is the story, and what each node
left behind. That last one is not decoration: the after-snapshots are what
a branch switch puts back, and a bundle without them imports a tree nobody
can switch inside.
`app/bundle.py` owns both formats and nothing else knows either. The v1
reader stays — those files are already on people's disks — and it is now
the only place a `variants` array exists anywhere.
The rule the module is built on is that a bundle carries what was chosen
and never what is derived. The head branch, the fork points, the live flags
and the anchors are decisions somebody made. The lineage, the head depth,
the legacy `index` and the variant ordinals are computed from those and are
rebuilt on the way in, because a bundle is a text file anybody can edit and
a derived field shipped beside its source is a chance for the file to
disagree with itself where no read would report it.
`index` is the one that stops being academic here. It agreed with `depth`
until SP5, and this is the first writer that has to fill it for a forked
story, where two branches both hold a node at depth 4. It is allocated one
per turn instead: siblings share it, no two coordinates do.
Everything a hand-edited file can get wrong about the shape of a tree is a
400 raised before the adventure row exists, because a half-applied import
is exactly the failure this phase exists to end — a story that goes quiet.
A file wrong about which attempt is live is corrected rather than refused;
that is an invariant of the database, not of the format.
Measured on the 600-action fixture: 587 kB to 911 kB, and all of the
increase is the outcomes at 489 B a node — the coordinates themselves save
57.5 B a node against the old turn-and-variants shape. Twenty forks add
660 B. 4.3% of the import body cap.
381 tests green, 16 of them new in test_bundle_v2.py. No migration, no
vacuum owed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
Attempts pile up at the tip as leaves and cost nothing. The moment the
player takes the story down one the line has already moved past, the two
futures have to coexist — so `tree.fork` gives that attempt a branch of
its own, forked at the depth just before it, and the line it leaves
keeps every turn it has.
One row is inserted and one row is moved. Nothing is copied: everything
before the fork is borrowed through the lineage cached on the branch
row. Measured on a 40-turn story forked twenty times — 21 branches, 140
rows, an 80-action story — the page load costs 31,652 B against the
31,433 B the same story flat costs, and a branch is 103 B of ancestry.
Nothing derived moves either, and that is the part worth keeping: a
memory hangs off the coordinate the parent's attempt still occupies, and
the lineage caps the parent one depth short of it. The fork simply
cannot see it, so it resummarizes that ground from the text it actually
tells, without a line of bookkeeping.
Two things had to change underneath. A new node's depth now comes from
the tip of its branch rather than from the adventure-wide `index`, which
would have left a hole in a fork's path the width of the other branch.
And undo stops at the fork — the turns before it belong to the branch
this one grew out of.
`GET /branches`, `POST /branches/{id}/switch` and
`POST /actions/{id}/fork` are the endpoints SP7's tree view is drawn on.
365 tests green, 18 of them new in test_branch_forking.py.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Every attempt at a turn is now its own row at the same (branch, depth),
with `live` naming the one the story tells. The JSON repeating group on
`actions.variants` is read one last time, by a migration that writes it
out as the sibling rows it always described, and then goes unread.
The snapshots turn around with it: an action carries the state it left
behind rather than the state it started from, because attempts at one
turn share a starting position and differ exactly in their outcome.
Rolling back is "what the node in front left behind", one lookup on the
path, and it is what undo and retry now both read.
And the memory holdback goes. It existed because retry rewrote a row
under a mark that had already moved past it; a retry writes a sibling
now, and replacing what a coordinate says withdraws what was derived
from it — the same repair undo and delete already made.
The assembled prompt is still stored once per turn: it moves with the
live flag, so a superseded attempt keeps only the few hundred bytes that
were its own. Measured on the 600-action fixture: 700 rows for the same
600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB
and page load 62.7 kB unmoved.
347 tests green. `tests/test_story_tree_baseline.py` and
`tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move
the baseline for the variant-count semantics and did not need to.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The memory bank and the story summary each kept a cursor: how many story
actions they had already covered. A count is a position in a list, and this
list moves — delete an action in front of the mark and every later one slides
down a slot, so the mark now covers one it has never read. All the cursor
bookkeeping existed to patch that up.
Both marks are now (branch_id, depth): the node up to and including which the
work is done. A depth is a coordinate along a path, not an offset into a list,
so nothing in front of it can move it. That deletes rather than rewrites
`position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
`prune_dangling_memories` and the every-pass clamp in `run_post_turn`.
A memory hangs off the node its block ends on, so a fork inherits its
ancestors' memories without copying any, and retrieval selects through the
branch clause over the *whole* lineage — recall is long-range by definition and
cannot be windowed. Measured: 1,807 B on a story forked twenty times against
1,823 B on a flat one of the same length.
Migrations 53-56 translate the old counts into nodes. They rewrite `adventures`
and not `actions`, so this one needs no VACUUM FULL.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The identity map holds weak references. A branch row nobody keeps a strong
reference to is collected between two nodes, so resolving the head inside the
placement loop read it back from the database for every node in the flush: 201
SELECTs on `branches` to write 200 actions, and 36 s -> 45 s on the same 297
tests. Nothing about any result changed, which is why only a stopwatch found
it, and why there is now a test counting the reads.
Also: the two reads left un-pathed on purpose say so where they live —
`max_action_index` allocates the legacy `index` and must stay adventure-wide
or two branches issue the same number, and export is a flat v1 bundle whose
reader has no idea branches exist. And `Adventure.actions` keeps its `index`
ordering, because ordering the collection by depth would not make it a story:
it is every branch's actions, and a path is a selection out of it.
318 tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Every read of an action now goes through a single module. `context/lineage.py`
turns a branch's stored lineage into the OR-of-ranges that is "this story", and
history, paging, the newest-action lookups, the index screen and the scripting
history API all select through it. A forgotten clause does not raise — it
quietly assembles a page, or a prompt, out of two different stories — so the
clause lives in one place rather than in a convention.
The read that mattered most was the shortcut: `_from_memory` sliced
`adventure.actions`, which is every branch's actions, not the path. It now cuts
the loaded collection down with the same predicate the SQL uses. Same trap one
layer up, and user-visible: `pipeline._history()` hands user scripts the story,
and was handing them the collection.
Tail reads window the lineage as well as the rows: the newest few entries cover
the context budget, so a story forked twenty times reads its tail with one
clause and costs 1.07x what an unforked story of the same length costs. The
estimate is depth arithmetic, and where a deleted action leaves a gap the read
notices it came up short and widens to the whole ancestry.
Ordering moves from `index` to `depth`, with `id` breaking ties. The two hold
the same numbers until retry stops mutating rows in SP4, but only one of them
is a position along a path.
One thing SP1 did not anticipate: wiring the writers was not enough. From here
a row without a branch is a row no read can see, and "every writer remembers"
has to hold for every fixture, script and test ever written — including the
SP0 baseline, which writes its actions straight to the database and must pass
unmodified. So the session enforces it: `tree.place_new_nodes` runs from
before_flush and places anything unplaced. The call sites keep their explicit
calls, because a node placed at the call site is placed before the code around
it reads it back.
316 tests green: the 297 from SP1, plus 19 in test_branch_clause.py.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a
`branches` table, `branch_id`/`depth` on actions and memories, a head pointer
on adventures, migrations 46-52, and a server-side backfill that re-reads every
existing adventure as a tree with one branch. `depth` holds the number `index`
already held, gaps included, so no story changes — a linear story *is* a tree
with one branch, which is what makes the SP0 baseline passing unmodified the
pass condition rather than a hope.
The writer had to come with it. No migration will ever visit a row written
after it ran, so columns backfilled today and populated next subphase would
leave a hole exactly the width of one deploy, and from SP2 on a row without a
branch is a row no read can see. `app/tree.py` owns that: one module, because a
node written without a branch fails by disappearing rather than by raising.
Three things the schema itself insisted on:
- `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing
both ways makes the two tables a cycle create_all cannot order, and its
escape hatch needs an ALTER SQLite does not have. It is a cache, and a head
naming a branch that is gone recovers onto the root.
- `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a
second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`,
because Postgres `json` has no equality operator.
- SQLite will not drop a column a foreign key names, which is how two existing
tests broke: they simulated an old database by rewinding the stamp while
leaving the new columns in place. Every ADD COLUMN migration is now
idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by
rebuilding three tables from frozen DDL so the real ALTERs run.
297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn;
page load and index are byte-identical to the recorded figures.
The deploy that ships this needs one `VACUUM FULL actions;` on the direct
endpoint afterwards — it rewrites every row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The stress fixture is sized from production and exists to weigh bytes, so it
leaves every column it does not weigh at its default. Checked against a freshly
built one, that is exactly the set phase 14 has to migrate: state_before and
world_state_before NULL on all 600 rows, no scenario and so no RPG layer, no
adventure scripts, both cursors 0, and a hundred retry histories whose two
attempts carry byte-identical text with variant_index pinned at 0.
That last one matters most. "Which attempt is live?" is the question SP4's
migration answers when it decides which sibling node becomes the head of a
turn, and against that fixture the question had no observable answer -- a
migration that picked wrong would have looked correct.
--rich fills in those columns and no others. A real RPG scenario read from the
seed data rather than invented, with a world state played forward so hp
declines and flags flip; monotonic per-action state snapshots, so a rollback
that does not happen reads as a wrong number instead of as nothing; a gold
script, story cards, non-zero cursors, pinned and forgotten memories; retry
attempts with distinct texts, counts of two and three, and a live attempt that
is often not the last written. Plus a second adventure, because a branch clause
that forgot its adventure still looks right on a database holding one.
The invariant SP4 reads -- text mirrors variants[variant_index] -- is asserted
at build time rather than assumed.
The plain fixture is untouched and re-measured unchanged at 1.8 kB, so the
egress ceilings stay comparable. All six shapes run clean on --rich, including
the turn path with the RPG layer and script pipeline now live. 283 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Phase 14 replaces the action list with a tree, which touches the memory bank,
the context builder, undo, retry and the UI at once. Before any of that moves,
write down what "unaffected" means and prove it holds today.
tests/test_story_tree_baseline.py drives the product over HTTP the way a player
does and asserts only on API responses, never on how anything is stored: the
turn engine, retry and variant switching, undo, paging all the way back to the
start without gaps or repeats, edit/delete, memories, export/import round-trip,
world state. 283 green with it added, from 259.
It must pass unmodified through the schema change, the branch clause and the
move of memories onto nodes. A linear story is a tree with one branch, so those
subphases have no licence to change behaviour, and this is what says so.
The plan itself records four decisions that were open: structural-first with no
runtime flag, legacy columns kept one release, full tree visualisation, and a
v2 bundle with a v1 reader. Plus four traps found by reading rather than
running -- history._from_memory and the scripting pipeline both slice
adventure.actions, which under a tree is every branch rather than the path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.
--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:
- create_all() builds the current schema but leaves the version stamp at its
default, and bootstrap() reads a populated-but-unstamped database as
ancient — it replays every migration against a schema that already has the
columns, and fails on the first.
- the fixture's user is a registered one, but local mode looks for the row
with email IS NULL and is_guest false, so without clearing the email the
app opens on an empty library.
--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.
Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The backfill issued an UPDATE per action. It runs at container start, before
uvicorn binds the port, against a database at the other end of a network --
944 rows on production, so a thousand sequential round trips standing between
the deploy and its first health check.
One executemany per batch instead. Same rows, same verification, same
transaction; nineteen round trips rather than nine hundred.
Re-verified on Postgres, since executemany binds bytea through a different
psycopg path: 720,864 B of JSON to 204,293 B, every row equal.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Opening a finished adventure fetched every action in one response: 589.5 kB on
production's longest, and nothing about that curve bends on its own, because a
story only ever gets longer. The page load now brings the newest 60 actions
and the reader pages up from there. On the harness's 600-action fixture that
is 606.0 kB down to 62.6 kB, and -- the part that matters -- it no longer
depends on how long the story is.
Paged by anchor, not by offset. `before_id` is the oldest action the caller
holds; the server returns what precedes it. An offset counted back from the
newest would shift every older position the moment a turn lands, which is
exactly when someone is likely to be scrolling, and the reader would get one
action twice and never see another. It also keeps working when the story stops
being a flat list: comparing indices to order a branch survives the story tree,
treating them as positions does not.
`has_more` comes from fetching one row past the window rather than from
counting. A deleted anchor -- undo, mid-scroll -- reports the end rather than
guessing and serving a page the reader already has.
GET /{id}/actions and POST /{id}/undo now return {actions, total, has_more}
instead of a bare list. Undo is the action most likely to be repeated several
times running, so having it re-fetch the whole story would have undone the
paging on the worst case.
The adventure payload gets its window through set_committed_value rather than
by assignment: the actions relationship cascades delete-orphan, so assigning a
60-item list to it would delete everything outside the window on the next
flush.
In Play.jsx the prepend is followed by a useLayoutEffect that restores the
scroll position, before paint, so the story does not jump. Loading starts 400px
from the top rather than at it, guarded by a ref because scroll fires far
faster than React re-renders. There is a button as well as the scroll trigger:
on a short viewport the transcript may not be tall enough to scroll at all, and
a reader who cannot scroll must still be able to reach the beginning.
Verified against a running backend and a 220-action adventure: the page load
returns 60 of 220 ending on the newest, and walking back from an anchor returns
exactly the actions before it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
One column is 89% of the database and the free tier allows 512 MB. Reads were
already solved -- the column is deferred, so a page load never touches it and
one screen fetches one row at a time -- but nothing had costed storage, and
storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per
action, so the ceiling arrives around 5,400 actions and 944 are stored.
Postgres already compresses it and only gets 1.7x. pglz is tuned for fast
decompression of data a query might filter on, and nothing has ever filtered
on an assembled prompt -- it is written once and read whole, rarely, by the
Insights viewer. zlib gets 3.5x on the same text for a decompress on a request
that already made an LLM call.
Done as a TypeDecorator rather than a second column, so every call site still
writes a dict and reads a dict back, and deferred/undefer/load_only keep
naming the same attribute. Only the storage format moves.
Migrations 43-45: add the bytea, convert into it, drop the original, rename.
The backfill is the one destructive step in the file -- 44 removes the only
other copy -- so it decompresses every row and compares it against what went
in, and a row that fails aborts the run. The whole loop is one transaction, so
an abort rolls the DROP back and the prompts are still there.
Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway
Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column
came out named context_snapshot, every snapshot compared equal and the one
NULL stayed NULL.
Postgres does not return the disk by itself: DROP COLUMN only marks the column
gone and the backfill leaves a dead tuple per row, so the table peaks near
twice its size before settling. The deploy needs one VACUUM FULL to collect
it; the migration comment says so.
The egress fixture's snapshots are prose now rather than "x" * 20_000, and the
prose generator moved to tools/fakeprose.py so the harness and the tests share
one definition. A repeated character compresses a thousandfold: against the
old fixture a compressed column looked free and the byte ceilings would have
been guarding nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
deferred=True keeps the four heavy Action columns out of a bulk read, but it
makes narrowness the thing a future column has to remember to ask for -- and
both egress blowouts this project has had were a column nobody remembered.
Listing what each list response renders inverts the default: a new column
costs nothing on these paths until someone adds it to the tuple.
The adventures index was not merely a future risk. It loaded whole Adventure
entities to render a title, a stamp and a snippet, and an Adventure carries
script_state, world_state, placeholders, story_summary, memory, authors_note
and ai_instructions -- ~15 kB a row in production, none of it on that screen,
all of it fetched once per adventure on every index load. Measured on six
adventures with 78 kB of body each: 469.7 kB entity-loaded against 318 B
projected.
The memories drawer stops walking adventure.memories. The walk is what
retrieval used to do and the reason a turn cost megabytes; a relationship load
takes whole entities, so it picks up whatever the model happens to grow.
Nothing changes today -- embedding_blob is already deferred -- which is the
point.
world_delta stays on the action list because ActionOut.world_changes is
computed from it. Leaving it off would not save the bytes, it would spend them
one row at a time as a lazy load.
Two tests cover the index: one asserts the listing query names none of the
body columns, one puts a byte ceiling on six adventures carrying 80 kB apiece.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
test_egress.py asserted which columns a statement names, which is the shape
both of this project's egress blowouts took. It would all still pass if a
response grew tenfold within the columns it is allowed to read -- and a story
that keeps getting longer does exactly that. Production's longest adventure is
607 actions where the plan assumed 200.
So dbmeter, which was built to be importable from tests and was not yet used
by any, now backs four byte ceilings: the page load, the action list, and one
action's snapshot fetched on demand. Budgets are per action rather than
absolute, so they mean the same thing whatever size the fixture is set to, and
generous -- 3 kB against a real 994 B. They are there to catch an order of
magnitude, not to freeze a byte count.
The fourth test is the one that keeps the other three honest. A ceiling proves
nothing unless the thing it excludes would breach it, so it undefers the
snapshot on purpose and asserts the same twelve rows cost more than ten times
the budget. If the fixture ever shrinks below the point where that holds, that
test fails rather than the ceilings quietly passing on nothing.
Meter grows detach() and a context manager. A script exits and takes the
wrapping with it; a test does not, and one test leaving the shared engine
metered would charge bytes to a scope nobody opened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The old fixture was wrong in both directions at once and happened to land near
the right total. Actions were modelled at ~2.1 KB against a real 886 B of
text, and adventures at 200 actions against a real 607. Width flattered,
length did not, and length is what a page load pays for.
Re-sized from the 2026-08-17 measurements: 600 actions, 1700 B of narration
alternating with a one-line player input, 232 KB of context_snapshot a row.
The page-load shape now reports 606.0 kB against the 589.5 kB measured on
production's longest adventure -- 2.8% out, where the old defaults were 28%
out on a story a third of the length.
Filler text is now generated word by word instead of one sentence repeated.
That matters for what comes next: the repeated string compresses 313x and the
generated prose 3.7x, so any compression ratio measured against the old
fixture would have been fiction, and shrinking context_snapshot is the open
question it exists to answer.
context_snapshot also gains a flag of its own rather than being hardcoded, and
the 74 KB figure in the comment -- inherited from models.py -- is corrected:
the real column averages 163 KB a row across the table.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Migration 38 left memories.embedding in place so a rollback could still find
the vectors. Production has since been verified reading from embedding_blob,
so migration 42 drops it: 4 MB of a 99.6 MB database holding nothing anyone
reads.
Removing it surfaced a live bug. Changing your embedding model is supposed to
throw the bank's vectors away and let the post-turn pass rebuild them, because
two models' vectors are not comparable. The settings route did that by nulling
memories.embedding -- correct until 38 moved the vectors, after which it
cleared the dead column and left the blob intact with `embedded` still true.
_embed_pending filters on `embedded IS FALSE`, so it never saw those rows and
the bank went on ranking against the old model's vectors permanently.
Nothing would have reported it. cosine returns 0.0 on a width mismatch, so a
different-width model scores every memory zero and retrieval returns whichever
rows happen to sort first; a same-width model scores plausible garbage.
The bulk clear now sets both columns. It stays a bulk UPDATE rather than going
through set_vector -- loading the rows is the cost that whole path exists to
avoid -- so set_vector's docstring now names it as the one caller that
legitimately writes those columns by hand. No cache invalidation is added:
clearing `embedded` drops the rows out of the catalogue query, and set_vector
evicts each entry as the re-embed puts it back.
test_embedding_blob.py now rebuilds the pre-38 schema by hand where it tests
the backfill, since create_all no longer produces the column it converts from,
and asserts 42 removes it at the end of a full bootstrap -- 38 reads that
column and 42 drops it, so an upgrade that reordered them would arrive with an
empty bank.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Everything measured so far ran on SQLite against a synthetic fixture, so two
claims were still on trust: that migration 38 spells BYTEA correctly for a real
server, and that the byte figures survive psycopg's encodings.
Both hold. Production reads schema_version 41 with embedding_blob bytea and
embedded boolean present, the backfill is complete at 134/134, and the packed
vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a
memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and
against a throwaway Neon database every shape lands within 0.5% of the SQLite
run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of
it.
The harness writes, so it refuses any target whose name does not say stress or
scratch -- pointed at the production database it stops rather than seeding it
with a fake user and 200 fake turns. It also empties a Postgres target before
building, which a fresh SQLite temp file never needed.
Two corrections fall out, both recorded in plan/13. The page-load model has the
wrong shape: real actions are half the fixture's weight but real stories run to
607 actions, not 200, so the worst real page load is 589.5 kB. And the decision
to leave context_snapshot in the database costed egress but never storage --
it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the
ceiling this deploy will hit first.
Measured with counts and octet_length sums only. No user content was read.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.
Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.
one turn 3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
run_post_turn 3,139.1 kB -> 0.7 kB
Insights 3,223.7 kB -> 117.9 kB
Memories drawer ~3.1 MB -> 23.6 kB
A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.
The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.
memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.
Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.
It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.
Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.
The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.
Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
Both egress blowouts this project has had were one query fetching a column
nobody read, and a statement count would have shown nothing wrong in either.
So the meter counts bytes, at the DBAPI cursor -- everything that crosses
that line crossed the wire.
tools/stress_session.py drives a production-shaped adventure through the real
routes with only the network faked. It reproduces both figures measured
directly on production: 426.7 kB for a 200-action page load against 423 KB,
and 3,258.7 kB for one turn against 3,153 kB.
The memory bank is on by default, which is the whole point -- the previous
harness ran without an embedding model, so retrieval returned early and the
heaviest read in a turn never happened. --no-embeddings reproduces that
deliberately, and the gap is 29x.
It also turned up two callers the production SQL could not see: run_post_turn
walks the whole bank again every turn, and Insights pays for it a third time.
A played turn costs ~6.4 MB, not 3.2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
The app had no cleanup of any kind: in multi-user mode every first visit
mints a users row, so the demo has been accumulating one permanent account
per visitor along with everything they generated.
cleanup.py sweeps guests idle for AIDND_GUEST_RETENTION_DAYS (default 5),
once at startup and then every few hours. Startup is the load-bearing
trigger — the free tier sleeps after ~15 minutes, so a long timer rarely
gets to fire.
Idle is COALESCE(last_seen_at, created_at), not last_seen_at: _touch only
writes that column hourly, and a guest minted by /auth/me has it NULL until
its second request, so the simpler query would have deleted brand-new
visitors mid-session.
It's one Core DELETE rather than db.delete(user), which would SELECT every
adventure, action and memory into Python purely to delete them — the same
egress pattern as the 189x fix. Every FK from users down is ON DELETE
CASCADE, so the database does the whole graph and returns a count.
The filter requires is_guest AND email IS NULL, so registered users (who
upgrade in place) and local mode's implicit user are both out of reach, and
is_public is output-only so a guest can never own content another user can
see. Session cookies have no expiry and can outlive a swept row; that path
401s and the frontend's existing retry re-mints a session.
Guests are told: /auth/me serves guest_retention_days and the signup modal
states the window, sourced from the server so it can't drift from what is
enforced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
The per-IP rate limits could be bypassed entirely: uvicorn ran with
--forwarded-allow-ips "*", which trusts the leftmost X-Forwarded-For value
(client-controlled), and Render forwards the inbound header rather than
stripping it. Rotating the header handed out a fresh rate-limit bucket per
request, so the login/register limit (10/5min) and guest-minting limit
(30/5min) were no throttle at all — unbounded password guessing and guest-row
creation. Confirmed live: fixed IP -> 429 after 10; rotating spoofed header ->
no 429 across 14 attempts.
Two-layer fix:
- limits._client_ip now derives the client IP from the hop the trusted edge
appends (rightmost of X-Forwarded-For), which a client can't spoof past;
tunable via AIDND_TRUSTED_PROXY_HOPS. Dropped --forwarded-allow-ips "*".
- New per-account login throttle (email-keyed, 8 fails / 15 min, cleared on
success): stops distributed guessing against one account that a per-IP limit
can't, since it can't be diluted across many source addresses.
Also close an SSRF on the BYOK endpoint_url (hosted mode only): the connection
test and turn/chat streams now refuse a URL that resolves to a non-public
address (private/loopback/link-local metadata/reserved), checked at request
time so it resists a DNS record flipping to a private IP. No-op locally, where
reaching localhost Ollama is intended.
Tests: test_ratelimit_hardening.py (8), test_netguard.py (13). 172 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
max_output_tokens is a hard wall the endpoint enforces mid-sentence. The
```state block is emitted after the narration, so a long turn hits the wall
partway through the block and the deltas are lost — silently, since nothing
reads finish_reason.
builder.length_hint() derives a word limit from the cap ((cap - 50 headroom)
* 0.75 words/token * 0.90 buffer) and injects it just above EMIT_REMINDER,
which keeps the last slot it needs. Reserved in build_context like the
reminder is.
Phrased as a ceiling, not a budget. Measured against gemma-4-26b at cap 800,
n=5 per arm: no hint 174 words, "keep this turn under about N words" 246,
"hard limit ... a typical turn is much shorter" 170. A budget reads as a
target to fill — every budget run was longer than every unhinted one, pushing
turns toward the wall the hint exists to avoid. Ceiling phrasing still works
at tight caps: at 250, unhinted hit finish_reason=length 2/6, hinted 0/6.
tests/test_length_hint.py, 11 tests; each mechanism verified by sabotage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Two reads still grew without bound after the snapshot fix.
`Action.variants` holds every discarded retry attempt, but a list response
only needs how many there are — so each retry permanently added ~5 KB to
every later load of that adventure. Defer the column and keep the count
beside it (migration 37, backfilled server-side), with set_variants() as the
one write path that keeps the two in step.
`story_actions()` walked adventure.actions, then every caller threw almost
all of it away: the builder concatenates the story and immediately cuts it
back to the token budget, the NPC check looks at the last 6, retrieval at the
last 4, the cursor clamp only wants a count. A turn on a 200-action adventure
read 839 KB to use ~70 KB, and grew with every turn played. app/context/
history.py serves those shapes from SQL; window_covering() measures the
actions it fetched and projects how many more it needs, fetching only the
part it does not already hold. Memorybank cursors move to position_of_index()
and settled_count()/settled_slice() — same arithmetic, no full list.
The scripting pipeline still receives the whole history per AI Dungeon's API,
and every helper reuses adventure.actions when it is already loaded, so a
scripted adventure pays what it always did and never twice.
Measured at production shape: retry tax 5.1 KB -> 0; turn 200 839 KB -> 129 KB
and flat from ~turn 50; a 200-turn playthrough 84.5 MB -> 23.0 MB; a delete
115 KB -> 5 KB.
Verified the window builds a byte-identical prompt to the full story across
budgets from 1K to 100K tokens, with and without the retry exclusion - this
is a cost change and nothing else. Cursor helpers checked against the old list
arithmetic, including after deleting a middle action. Counts are real
SELECT count(...): Query.count() wraps the entity select in a subquery, so the
SQL named every deferred column and the egress guard could not tell it apart
from a bulk fetch. 139 tests pass; the four new guards verified by sabotage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
The free-tier 5 GB/month network transfer allowance ran out, which blocks
connections outright. The database is only ~55 MB, so 5 GB meant the whole
thing was being pulled roughly 90 times over.
Cause: actions is 39 MB of that 55 MB -- 541 rows at ~74 KB each, almost
entirely context_snapshot, which stores the whole assembled prompt for a
turn. Every adventure load and every turn fetched all of it in order to
read two small things out of it: the world-change chips under an AI
message (Action.world_changes) and the emit block re-attached when
replaying history to the model (_history_text). The Insights viewer is
the only consumer that wants the whole snapshot, and it asks for one
action at a time.
Lifts that slice into its own small actions.world_delta column
(migration 36) and marks context_snapshot, state_before and
world_state_before deferred, so they load only when something touches
the attribute -- Insights, undo and retry, all single-action paths.
The backfill runs server-side, dialect-specific (json_extract on SQLite,
#> on Postgres), because pulling 39 MB of snapshots into Python to
rewrite a slice of each would defeat the purpose.
Measured at production shape (541 actions, 72 KB snapshots), one
adventure load goes from 38.46 MB to 0.20 MB. The traffic that consumed
5 GB would now be about 27 MB.
Deliberately not included: limiting the history query to recent actions,
and removing the redundant db.refresh(adventure) calls. Both were sized
against the old numbers; against a 0.20 MB load they would take ~27 MB a
month down to ~10 MB, which is not worth the complexity.
tests/test_egress.py hooks before_cursor_execute and asserts the emitted
SQL never names the deferred columns during a bulk load, so this cannot
regress silently. 123 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Three fallout bugs from keeping the retried action row alive (906ba42),
plus two long-standing cursor bugs the same investigation turned up.
Retry context leak: the row being regenerated is still attached to the
adventure, so it was replayed as established story and the model wrote a
continuation of the attempt it was meant to replace — the story visibly
blended both takes. It leaked into four places, not one: history replay,
story-card trigger matching, in-scene NPC detection, and the memory-bank
similarity query. Adds a shared context.story_actions(exclude_action_id),
threaded through build_context and retrieve_memories.
Memory holdback: a memory could summarize the just-generated turn; retry
rewrites Action.text but memory_cursor has already advanced, so the memory
was never regenerated and went on describing narration no longer in the
story. settled_story_actions() holds the newest action back one turn —
only the last action is retryable, so that makes it unreachable. The
settled list is always a prefix, so cursors stay valid and nothing is
skipped. The run_post_turn clamp deliberately still uses the full count:
clamping to settled rewinds legacy adventures a step and double-covers an
action.
Cursor bookkeeping: memory_cursor is a position into story_actions() while
Memory.source_* are Action.index values, and the two diverge as soon as
anything is deleted. Deleting a middle action slid a never-summarized
action into the covered range, skipping it forever; and pruning a memory
left the actions it covered stranded behind the cursor. Adds
note_action_removed() (called before the delete in delete_action and
undo_turn) and a rewind in prune_dangling_memories. delete_action also
now prunes at all, which it never did.
Not addressed: editing an already-summarized action still leaves its
memory stale, and the cumulative story summary can't have one fact
un-mixed from it.
117 backend tests pass, including new test_memory_settling.py (12) and
two retry-context regression tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Retry used to delete the last AI action and generate a replacement, so the
discarded narration was simply gone. The row survives now: each attempt is
appended to actions.variants with variant_index naming the live one, and a
ChatGPT-style pager under the message browses them.
Action.text still mirrors the active variant, so the context builder, memory
bank, summarizer and export needed no changes. A variant carries only what
differs between attempts -- the text, the reasoning, and the state it
produced -- never the assembled prompt, which is identical across attempts of
one turn and is the bulk of context_snapshot.
Only the last message can be switched, restoring the script/world state that
attempt produced; earlier turns were written as a continuation of whatever is
active there, so theirs are read-only previews.
Three things that would otherwise bite:
- generate_turn now wraps _generate_turn and watches for a save sentinel. If
the generator ends without it (provider error, empty reply, script stop,
client hangup) it re-applies the previous variant -- otherwise a failed
retry leaves rolled-back stats under un-rolled-back text.
- Retry reuses the turn's own index rather than next_index, or the clock the
world-state cooldowns run on advances on a re-run of the same turn.
- Editing a message rewrites the active variant too, or paging away and back
silently reverts the edit.
Migrations 34/35 verified as an upgrade against a populated database, not
just a fresh schema. test_state_revert's retry test asserted the old delete
behaviour and was rewritten.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.
Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.
Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
An adventure copies its scenario's plot text and story cards at creation so
later authoring never disturbs a story in progress. This is the explicit
opt-out, alongside the existing per-script "Sync from library".
GET /adventures/{id}/refresh returns a plan (per-field old/new diff, card
add/update/remove, world-state added/removed paths, and any ${...} answers
still needed); POST applies it under the turn lock so it can't race a
generating turn. The Plot panel shows the plan in a confirm modal first.
Overwrites the plot fields and scenario-derived cards. Deliberately left
alone: the opening `start` action (the story is built on it, and it is baked
into memories and the summary), the adventure's own title and summary,
player-authored story cards, and the live value of every stat the schema
still defines.
Two enablers were needed:
- adventures.placeholders (migration 32). ${...} answers were used once at
creation and discarded, so re-copying scenario text would have re-injected
a literal ${Hero}. Adventures predating the column re-prompt once via the
existing modal, then the answers are saved.
- story_cards.source_ref (migration 33), "card:<id>" / "npc:<key>", NULL for
player-authored. Adventure cards had no link back to their source, so a
rename read as delete-plus-add and player cards would have been clobbered.
Legacy cards name-match once, then adopt the ref.
World state goes through a new worldstate.reconcile(): keep values the schema
still defines, add missing ones at their initial, drop removed ones and their
cooldown bookkeeping. Not instantiate(), which would heal the player to full
and wipe their milestones.
The confirm modal is portalled to <body>: .side-panel's panel-in animation
has fill mode `both`, which makes it the containing block for position:fixed
descendants, so an overlay rendered in place was trapped in the 420px panel
and clipped by its overflow.
14 new tests in backend/tests/test_scenario_refresh.py; 85 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XzBCyXH4hBEVHqEaertcq4
The demo-key backstop added in 1dd3108 keyed its check on
`api_key == DEMO_API_KEY` rather than on `using_demo`. That looked stricter but
was wrong: the demo key is an ordinary OpenRouter key, so a user can
legitimately paste that same value into their own Settings as BYOK. The guard
then raised on every resolve_provider_config() call for that account.
Because me_payload() resolves a provider config, this 500'd GET /api/auth/me —
the SPA's bootstrap call — so the frontend's `me` never resolved and the nav
(including the AI Chat link) never rendered, on top of chat itself failing.
`using_demo` is the flag that actually means "the server is paying", and only
resolve_provider_config's demo branch sets it, so the pinning guarantee is
unchanged: server-funded turns still can't reach an off-whitelist model.
Adds a regression test for a BYOK user whose key equals the demo key value, and
corrects the test that had asserted the buggy behaviour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
AI Chat is a plain scratchpad for talking to a model directly — no story
context, scripts or world state — for poking at models, prompts and endpoints
without starting an adventure. Power users only: the router 404s (rather than
403s) for everyone else and the nav link is hidden. The conversation lives in
localStorage, so there's no new table or migration.
is_power_user() now also returns True in local mode: it's the operator's own
machine and their own key, the same reasoning that makes the provider debug log
local-only.
Alongside that, the rule keeping the shared demo key off paid models now lives
in exactly one place. It had been duplicated into the chat router, which is how
one copy eventually drifts:
- resolve_provider_config() takes an optional model_override and is the only
place the whitelist is applied, so turns, AI Chat and the connection test all
inherit it. An override is a per-request preference, never a grant.
- ProviderConfig.__post_init__ refuses to exist when api_key is the demo key
and the model isn't whitelisted. It keys on the key itself rather than the
using_demo flag, so a mislabelled config can't slip past, and it raises so a
future path that bypasses the resolver fails loudly instead of billing.
- The demo branch still pins endpoint_url too — a user-controlled endpoint
would leak the key itself, which is worse than spending it.
Provider gained chat(messages, ...) beside generate(), both delegating to a
shared _stream(url, body); completion-mode endpoints get the messages flattened
into a labelled transcript. Settings' /models fetch moved to
list_endpoint_models() and is shared with /api/chat/config.
Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub
resolve_provider_config — the point is to exercise the real BYOK-vs-demo
decision and assert on what the provider actually received: off-whitelist
override pinned, off-whitelist Settings.model pinned, redirected endpoint
pinned, BYOK passed through untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
Rework the home page into a single landing surface and give the app a
consistent visual language, chosen as "illuminated tome" over two other
pitched directions because it builds on the existing Cinzel + gold identity
instead of replacing it.
Home is now Continue (up to 4 in-progress stories, each showing where you
left off) over a scenario shelf, each section with a "See all" link. The
full adventure list moves to /adventures.
Scenario cover art has three tiers, in precedence order: an uploaded picture
(downscaled client-side to 400px WebP before storing), an emoji, or gradient
art generated from a hash of the title so no card is ever an empty box.
Adventures inherit their scenario's art. Images live in the row rather than
on disk because Render's free tier has no persistent volume, and it keeps
export bundles self-contained; list responses carry a cacheable
/api/scenarios/{id}/image URL rather than the base64.
Also: ambient drifting motes behind the app, loading skeletons, staggered
card entrance, ornamental scene breaks and a drop cap in the story, a
"Weaving" thinking indicator, and a toast system replacing every alert().
Two fixes found along the way:
- Importing a scenario bundle with no "tags" key returned a 500. Column
defaults are not applied until flush, so the attribute was still None
when the width clamp sliced it.
- Anything meaning "the story's latest narration" was missing action type
"start", which is the only text a freshly created adventure has, so new
adventures looked empty. Collected as NARRATION_TYPES.
Migrations 30 and 31 add scenarios.image and scenarios.icon; both are
additive with a '' default and were verified against a database stamped at
29. vite.config.js now reads AIDND_API_PORT so the recurring port-8000
clash with another local app needs no file edit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
The AI would sometimes stop emitting the `state` delta block once it missed
a turn. Two compounding causes: the emit rule sat only in the system block
(far from where the model generates), and the block was stripped before
storage — so every replayed history turn looked blockless, biasing the model
by imitation to stop emitting too.
- EMIT_REMINDER: a one-line reminder appended last in the prompt (strongest
recency slot), gated on has_ws and counted against the token budget.
- render_delta_block + _history_text: re-attach each past AI turn's own delta
block in replayed history (reconstructed from the stored snapshot delta), so
the model always sees its emit format. Action.text stays clean, so UI,
embeddings, and card/NPC trigger-matching are unaffected. History budgeting
counts the augmented text so it can't overflow.
Undo/retry untouched (read the separate world_state_before column). 46 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UWVyFKvqJGjfbXdibLgkMe
Players/authors can now directly correct the live values of stats the
schema already defines (health, trust, flags, milestones, etc.) without
waiting for the AI to emit a delta. New apply_override() sets values
absolutely rather than adding deltas, and — unlike the AI-facing
apply_delta() — bypasses cooldown/max_delta_per_turn and lets
milestones be un-set, since this is a deliberate correction rather
than a turn to police. Exposed via PUT /adventures/{id}/world-state
and an edit toggle in the World State drawer. Also fixes the
schema-editor stat-kind dropdown and free-text initial-value input
to size consistently with the numeric fields next to them.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Numeric stats couldn't represent things like worn armor or held items.
A stat can now be marked type: "text" — the AI reports the new full
value instead of a delta, with no clamping/bands (cooldown still
applies). Updated the scenario editor's stat kind selector, the
World State drawer's stat display, and the emit-rule prompt; added
tests and an outfit stat to the RPG demo scenario.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Tell the model to read each stat's range and band labels and keep changes
proportionate — minor events nudge a value, while large jumps or hitting a
min/max are reserved for pivotal moments. Examples are relationship/story
based (not combat) so it fits any scenario.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reword the state-emit rule to be scenario-agnostic: record any tracked
change — health/resources, time, relationships/mood, status, progress,
items/info — not just fight outcomes. Soften the demo's AI instructions to
cover relationships and objectives alongside combat.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rewrite the global emit rule to be directive: treat narration as
authoritative and record numeric changes (damage, healing, time, emotion
shifts, fights intensifying), not just easy on/off flags — updating every
stat the scene affected. Also firm up the demo scenario's AI instructions
to reflect combat outcomes in the numbers each turn. Helps weaker free
models emit numeric deltas; the emit rule applies live so all RPG
adventures benefit immediately.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Each AI action now carries a compact world_changes summary (derived from its
stored snapshot), rendered as small chips beneath the message: numeric stats
show a signed delta (green up / red down), flags show on/off, milestones show
a check. Gives an at-a-glance "what changed" without opening Insights.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Store the model's literal reply (before the world-state block is stripped)
in each turn's snapshot, and render it in the Insights panel when inspecting
a past AI message (the 🔍 button). Makes it possible to see exactly what the
model emitted — including whether it sent a state delta and with what paths —
which is the fastest way to debug under-reported or misrouted stat changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replace the single shared `npc` stat template (+ npc_card_types) with an
`npcs` section: each NPC keyed by a stable id, carrying its own name,
description, trigger keys, and its OWN stats block. The AI addresses NPCs
as npc.<id>.<stat> (id shown in context), which also fixes the old
card-id-guessing problem. On adventure creation each NPC auto-creates a
story card (name/keys/desc) for lore + in-scene detection, unless a
same-name card already exists. All NPCs instantiate up front.
- engine: npcs instantiate/apply/render/reference, npc_name/npc_triggers
- builder: _visible_npcs matches each NPC's own keys
- create_adventure: auto-create story cards from npcs
- WorldStateDrawer: render defined NPCs with their own stats + desc tooltip
- demo seed: Gwen (health/trust) + Bandit Leader (health/aggression)
- tests updated (34 pass); no new migration (npcs lives in stat_schema JSON)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Structured world/player/NPC stats, two-way flags, and sticky milestones
per scenario (stat_schema). The AI proposes a per-turn delta; a Python
engine referees it (clamp to min/max, per-turn cap, cooldown, counters).
Band word-labels plus a fixed stat guide (descriptions + full ranges)
keep the model grounded. World State drawer + Insights delta report;
undo/retry roll it back via the Phase 11 snapshot pattern.
- migrations 26-28 (scenarios.stat_schema, adventures.world_state,
actions.world_state_before); all nullable, additive, safe on existing rows
- migration 29 raises the default context budget 4096 -> 16384
(custom values preserved)
- seeded demo scenario 04-rpg-world-state.json (Bandit Camp)
- 19 new tests (33 total pass)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The shared per-adventure script_state ("scoreboard" scripts write to) was
never reverted by undo, and retry re-ran the output hook on top of the already-
mutated state, double-applying its changes (e.g. "+10 gold" became +20).
Each action now snapshots script_state as it was immediately before its own
hooks ran (new Action.state_before column, migration 25):
- undo restores the turn's first-action snapshot, prunes memories that
summarized the removed actions, and takes the turn lock against races.
- retry restores the AI action's snapshot before regenerating.
Story-card mutations are not reverted (documented limit). Adds the project's
first test suite: unit + full HTTP integration through the real scripting
engine (14 tests). See plan/11-state-revert-and-retry-fix.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Map HTTP 429 in _friendly_http_error: OpenRouter's free-models-per-day cap
gets a "daily limit, resets 00:00 UTC" message; other rate limits get a
generic "too many requests, try again" note. Avoids leaking the raw error
JSON (incl. user_id) to players.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Adventure scripts are frozen snapshots taken at creation. Add an opt-in
way to pull the latest library code into a live adventure:
- AdventureScript.source_script_id links a copy to its library Script
(migration 24; set at adventure creation)
- list endpoint flags out_of_date by diffing copy vs library
- POST /adventures/{id}/scripts/{sid}/sync overwrites the copy's code,
preserving enabled/position/script_state
- legacy copies with no link fall back to a name match, then adopt the link
- Play page shows a Sync button only when a copy differs from its library
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Trusted testers listed in AIDND_POWER_USERS get unmetered turns on the
shared demo key: demo_turns_left reports the full cap and count_demo_turn
skips them. Registered accounts only, matched case-insensitively.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5