e692780f08da1eb433431a0e2a076d9b6debb1f9
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
811368d048 |
Refresh what a switch changes, not what a turn changes
A branch switch does not change the length of the story. It changes which story it is. Four panels keyed on actions.length and so could not tell the difference: the Branches panel drew one branch while the reader was already on a second, Insights showed the prompt built for the path just left, the script-state drawer kept the other line's numbers, and the Memory Bank did not notice a deleted branch taking its memories with it. Only the world-state drawer was right, and only because it happened to carry stateKey already. They all key on the pair now, and deleting a branch bumps it too — that is the one operation that changes what is stored without a turn being played and without the story on the current path moving by a single action. tools/branch_fixture.py is the thing that could show it. The stress fixture's world state is empty, so it cannot answer whether a switch puts the scoreboard back, and its story is one branch. This builds a small bootable adventure with a stat schema, a gold script, two takes on one turn that differ by 35 hit points, a fork, and a memory on each side — with both branches the same length on purpose, because equal length is precisely the case a length-based key cannot see. Verified in a browser with the drawer open: hp 60 to 95 and back, the bar redrawn, the story swapped to the other take, Insights carrying the scratch and not the beating. The Memory Bank deliberately does not change on a switch: the drawer is adventure-wide so a memory is always findable to delete, and retrieval is the path-scoped half. 396 tests, build clean, no new lint. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g |
||
|
|
a7bf47a35e |
Let a backup carry a story that went two ways
A bundle had one list and a forked adventure has two stories, so export was emitting every branch's turns interleaved by index — a mangled story rather than lost data, and unreachable only because forking has no UI yet. `ai-dnd-adventure-v2` carries the branches, the depth each one left its parent at, which attempt at every turn is the story, and what each node left behind. That last one is not decoration: the after-snapshots are what a branch switch puts back, and a bundle without them imports a tree nobody can switch inside. `app/bundle.py` owns both formats and nothing else knows either. The v1 reader stays — those files are already on people's disks — and it is now the only place a `variants` array exists anywhere. The rule the module is built on is that a bundle carries what was chosen and never what is derived. The head branch, the fork points, the live flags and the anchors are decisions somebody made. The lineage, the head depth, the legacy `index` and the variant ordinals are computed from those and are rebuilt on the way in, because a bundle is a text file anybody can edit and a derived field shipped beside its source is a chance for the file to disagree with itself where no read would report it. `index` is the one that stops being academic here. It agreed with `depth` until SP5, and this is the first writer that has to fill it for a forked story, where two branches both hold a node at depth 4. It is allocated one per turn instead: siblings share it, no two coordinates do. Everything a hand-edited file can get wrong about the shape of a tree is a 400 raised before the adventure row exists, because a half-applied import is exactly the failure this phase exists to end — a story that goes quiet. A file wrong about which attempt is live is corrected rather than refused; that is an invariant of the database, not of the format. Measured on the 600-action fixture: 587 kB to 911 kB, and all of the increase is the outcomes at 489 B a node — the coordinates themselves save 57.5 B a node against the old turn-and-variants shape. Twenty forks add 660 B. 4.3% of the import body cap. 381 tests green, 16 of them new in test_bundle_v2.py. No migration, no vacuum owed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g |
||
|
|
0a12d9cd47 |
Make a retry a node, not a rewrite
Every attempt at a turn is now its own row at the same (branch, depth), with `live` naming the one the story tells. The JSON repeating group on `actions.variants` is read one last time, by a migration that writes it out as the sibling rows it always described, and then goes unread. The snapshots turn around with it: an action carries the state it left behind rather than the state it started from, because attempts at one turn share a starting position and differ exactly in their outcome. Rolling back is "what the node in front left behind", one lookup on the path, and it is what undo and retry now both read. And the memory holdback goes. It existed because retry rewrote a row under a mark that had already moved past it; a retry writes a sibling now, and replacing what a coordinate says withdraws what was derived from it — the same repair undo and delete already made. The assembled prompt is still stored once per turn: it moves with the live flag, so a superseded attempt keeps only the few hundred bytes that were its own. Measured on the 600-action fixture: 700 rows for the same 600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB and page load 62.7 kB unmoved. 347 tests green. `tests/test_story_tree_baseline.py` and `tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move the baseline for the variant-count semantics and did not need to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
c51531709d |
Mark the story with a node, not with a count
The memory bank and the story summary each kept a cursor: how many story actions they had already covered. A count is a position in a list, and this list moves — delete an action in front of the mark and every later one slides down a slot, so the mark now covers one it has never read. All the cursor bookkeeping existed to patch that up. Both marks are now (branch_id, depth): the node up to and including which the work is done. A depth is a coordinate along a path, not an offset into a list, so nothing in front of it can move it. That deletes rather than rewrites `position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`, `prune_dangling_memories` and the every-pass clamp in `run_post_turn`. A memory hangs off the node its block ends on, so a fork inherits its ancestors' memories without copying any, and retrieval selects through the branch clause over the *whole* lineage — recall is long-range by definition and cannot be windowed. Measured: 1,807 B on a story forked twenty times against 1,823 B on a flat one of the same length. Migrations 53-56 translate the old counts into nodes. They rewrite `adventures` and not `actions`, so this one needs no VACUUM FULL. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
d3756abdaa |
Give every action a branch and a depth
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a `branches` table, `branch_id`/`depth` on actions and memories, a head pointer on adventures, migrations 46-52, and a server-side backfill that re-reads every existing adventure as a tree with one branch. `depth` holds the number `index` already held, gaps included, so no story changes — a linear story *is* a tree with one branch, which is what makes the SP0 baseline passing unmodified the pass condition rather than a hope. The writer had to come with it. No migration will ever visit a row written after it ran, so columns backfilled today and populated next subphase would leave a hole exactly the width of one deploy, and from SP2 on a row without a branch is a row no read can see. `app/tree.py` owns that: one module, because a node written without a branch fails by disappearing rather than by raising. Three things the schema itself insisted on: - `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing both ways makes the two tables a cycle create_all cannot order, and its escape hatch needs an ALTER SQLite does not have. It is a cache, and a head naming a branch that is gone recovers onto the root. - `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`, because Postgres `json` has no equality operator. - SQLite will not drop a column a foreign key names, which is how two existing tests broke: they simulated an old database by rewinding the stamp while leaving the new columns in place. Every ADD COLUMN migration is now idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by rebuilding three tables from frozen DDL so the real ALTERs run. 297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn; page load and index are byte-identical to the recorded figures. The deploy that ships this needs one `VACUUM FULL actions;` on the direct endpoint afterwards — it rewrites every row. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
5c1bcf7305 |
Build a fixture that can witness what a tree changes
The stress fixture is sized from production and exists to weigh bytes, so it leaves every column it does not weigh at its default. Checked against a freshly built one, that is exactly the set phase 14 has to migrate: state_before and world_state_before NULL on all 600 rows, no scenario and so no RPG layer, no adventure scripts, both cursors 0, and a hundred retry histories whose two attempts carry byte-identical text with variant_index pinned at 0. That last one matters most. "Which attempt is live?" is the question SP4's migration answers when it decides which sibling node becomes the head of a turn, and against that fixture the question had no observable answer -- a migration that picked wrong would have looked correct. --rich fills in those columns and no others. A real RPG scenario read from the seed data rather than invented, with a world state played forward so hp declines and flags flip; monotonic per-action state snapshots, so a rollback that does not happen reads as a wrong number instead of as nothing; a gold script, story cards, non-zero cursors, pinned and forgotten memories; retry attempts with distinct texts, counts of two and three, and a live attempt that is often not the last written. Plus a second adventure, because a branch clause that forgot its adventure still looks right on a database holding one. The invariant SP4 reads -- text mirrors variants[variant_index] -- is asserted at build time rather than assumed. The plain fixture is untouched and re-measured unchanged at 1.8 kB, so the egress ceilings stay comparable. All six shapes run clean on --rich, including the turn path with the RPG layer and script pipeline now live. 283 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
0f1e05c808 |
Keep the fixture, so there is something to scroll (#5)
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.
--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:
- create_all() builds the current schema but leaves the version stamp at its
default, and bootstrap() reads a populated-but-unstamped database as
ancient — it replays every migration against a schema that already has the
columns, and fails on the first.
- the fixture's user is a registered one, but local mode looks for the row
with email IS NULL and is_guest false, so without clearing the email the
app opens on an empty library.
--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.
Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ae6e5af6c7 |
Store context_snapshot compressed
One column is 89% of the database and the free tier allows 512 MB. Reads were already solved -- the column is deferred, so a page load never touches it and one screen fetches one row at a time -- but nothing had costed storage, and storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per action, so the ceiling arrives around 5,400 actions and 944 are stored. Postgres already compresses it and only gets 1.7x. pglz is tuned for fast decompression of data a query might filter on, and nothing has ever filtered on an assembled prompt -- it is written once and read whole, rarely, by the Insights viewer. zlib gets 3.5x on the same text for a decompress on a request that already made an LLM call. Done as a TypeDecorator rather than a second column, so every call site still writes a dict and reads a dict back, and deferred/undefer/load_only keep naming the same attribute. Only the storage format moves. Migrations 43-45: add the bytea, convert into it, drop the original, rename. The backfill is the one destructive step in the file -- 44 removes the only other copy -- so it decompresses every row and compares it against what went in, and a row that fails aborts the run. The whole loop is one transaction, so an abort rolls the DROP back and the prompts are still there. Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column came out named context_snapshot, every snapshot compared equal and the one NULL stayed NULL. Postgres does not return the disk by itself: DROP COLUMN only marks the column gone and the backfill leaves a dead tuple per row, so the table peaks near twice its size before settling. The deploy needs one VACUUM FULL to collect it; the migration comment says so. The egress fixture's snapshots are prose now rather than "x" * 20_000, and the prose generator moved to tools/fakeprose.py so the harness and the tests share one definition. A repeated character compresses a thousandfold: against the old fixture a compressed column looked free and the byte ceilings would have been guarding nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
12d57afdac |
Put a number on what an endpoint may fetch, not just a column list
test_egress.py asserted which columns a statement names, which is the shape both of this project's egress blowouts took. It would all still pass if a response grew tenfold within the columns it is allowed to read -- and a story that keeps getting longer does exactly that. Production's longest adventure is 607 actions where the plan assumed 200. So dbmeter, which was built to be importable from tests and was not yet used by any, now backs four byte ceilings: the page load, the action list, and one action's snapshot fetched on demand. Budgets are per action rather than absolute, so they mean the same thing whatever size the fixture is set to, and generous -- 3 kB against a real 994 B. They are there to catch an order of magnitude, not to freeze a byte count. The fourth test is the one that keeps the other three honest. A ceiling proves nothing unless the thing it excludes would breach it, so it undefers the snapshot on purpose and asserts the same twelve rows cost more than ten times the budget. If the fixture ever shrinks below the point where that holds, that test fails rather than the ceilings quietly passing on nothing. Meter grows detach() and a context manager. A script exits and takes the wrapping with it; a test does not, and one test leaving the shared engine metered would charge bytes to a scope nobody opened. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
be66780a26 |
Size the stress fixture from what production actually holds
The old fixture was wrong in both directions at once and happened to land near the right total. Actions were modelled at ~2.1 KB against a real 886 B of text, and adventures at 200 actions against a real 607. Width flattered, length did not, and length is what a page load pays for. Re-sized from the 2026-08-17 measurements: 600 actions, 1700 B of narration alternating with a one-line player input, 232 KB of context_snapshot a row. The page-load shape now reports 606.0 kB against the 589.5 kB measured on production's longest adventure -- 2.8% out, where the old defaults were 28% out on a story a third of the length. Filler text is now generated word by word instead of one sentence repeated. That matters for what comes next: the repeated string compresses 313x and the generated prose 3.7x, so any compression ratio measured against the old fixture would have been fiction, and shrinking context_snapshot is the open question it exists to answer. context_snapshot also gains a flag of its own rather than being hardcoded, and the 74 KB figure in the comment -- inherited from models.py -- is corrected: the real column averages 163 KB a row across the table. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
70d024da62 |
Check the egress work against the database it actually runs on
Everything measured so far ran on SQLite against a synthetic fixture, so two claims were still on trust: that migration 38 spells BYTEA correctly for a real server, and that the byte figures survive psycopg's encodings. Both hold. Production reads schema_version 41 with embedding_blob bytea and embedded boolean present, the backfill is complete at 134/134, and the packed vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and against a throwaway Neon database every shape lands within 0.5% of the SQLite run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of it. The harness writes, so it refuses any target whose name does not say stress or scratch -- pointed at the production database it stops rather than seeding it with a fake user and 200 fake turns. It also empties a Postgres target before building, which a fresh SQLite temp file never needed. Two corrections fall out, both recorded in plan/13. The page-load model has the wrong shape: real actions are half the fixture's weight but real stories run to 607 actions, not 200, so the worst real page load is 589.5 kB. And the decision to leave context_snapshot in the database costed egress but never storage -- it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the ceiling this deploy will hit first. Measured with counts and octet_length sums only. No user content was read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
b7e53ae581 |
Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.
Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.
one turn 3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
run_post_turn 3,139.1 kB -> 0.7 kB
Insights 3,223.7 kB -> 117.9 kB
Memories drawer ~3.1 MB -> 23.6 kB
A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.
The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.
memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.
Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
|
||
|
|
c56864877a |
Store embeddings as packed float32 instead of a JSON list
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.
It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.
Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.
The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.
Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
|
||
|
|
7ee5ceea6c |
Measure database egress in bytes, from inside the repo
Both egress blowouts this project has had were one query fetching a column nobody read, and a statement count would have shown nothing wrong in either. So the meter counts bytes, at the DBAPI cursor -- everything that crosses that line crossed the wire. tools/stress_session.py drives a production-shaped adventure through the real routes with only the network faked. It reproduces both figures measured directly on production: 426.7 kB for a 200-action page load against 423 KB, and 3,258.7 kB for one turn against 3,153 kB. The memory bank is on by default, which is the whole point -- the previous harness ran without an embedding model, so retrieval returned early and the heaviest read in a turn never happened. --no-embeddings reproduces that deliberately, and the gap is 29x. It also turned up two callers the production SQL could not see: run_post_turn walks the whole bank again every turn, and Insights pays for it a third time. A played turn costs ~6.4 MB, not 3.2. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7 |