The handover said SP9 still needed a person clicking it. It has had one, and
the two bugs that came out are the point: both were in the gap between the
suite and the screen, and the frontend still has no test runner to close it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The 11 MB the vacuum gave back against a mid-hundreds guess, and the reason:
the heap is 1.7 MB and the rest is TOAST, so a migration touching small
columns bloats the heap and reuses the toast pointers. The rule keeps its
cost estimate rather than its size, and SP8 is the other shape -- it drops
toasted columns, and DROP COLUMN frees nothing until a VACUUM FULL.
And the handover moves to SP9: what a hand-driven SP7 turned out to be wrong
about, the three findings from fixing it, and the fact that the fix itself has
not been driven by hand either.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
STATUS still described SP7 as the newest thing and counted 396 tests. The
stack is open as PR #6 and the review pass is answered, so the pick-up
paragraph now says both, and part five records what the review found.
The part worth writing down is the finding that was turned down rather than
the eight that were fixed: a memory anchored to a node is meant to go when the
node goes, the root is the one exception, and the reason the exception is
narrow — protecting depth 0 outright would rebuild the dangling rows
forget_node exists to prevent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
A hand-written memory used to carry a NULL depth, described in the model as
"belongs to the adventure rather than to a path". That sounds harmless and
is not: a NULL is a coordinate no fork can cap, so a note typed on one line
followed the reader onto branches whose story it never described. It takes
the head now — the story you were reading when you wrote it — and obeys
exactly the rule a summarised memory obeys.
The unanchored escape clause in lineage.Path.clause existed for that single
case and is deleted rather than left unused. Its docstring argued that a
capped depth would drop a typed memory the moment its branch stopped being
the newest entry; anchoring answers the same worry better, because the
memory is not exempt from the path, it is on one.
The drawer now shows the path being read and nothing else, filtered by the
clause retrieval itself uses, so the bank you can see is the bank the model
can see. Nothing is stranded: a memory lives on a branch, switching to that
branch shows it, and deleting the branch deletes it. Pinning decides order,
the path decides existence.
Migration 62 lands existing NULL-depth memories at depth 0 of their branch
rather than at the tip. 0 is at or before every fork point, so every memory
stays visible from exactly the paths it is visible from today — nobody's
bank loses a row on deploy. The tip is the tidier-sounding choice and would
have emptied them out of every branch forked earlier than they were typed.
This supersedes the on_path flag and the "another branch" badge from
earlier today; anchoring makes them redundant, and they are removed.
Four tests changed because they asserted the old contract, not because
they broke. The one worth reading is the pair replacing
test_a_hand_written_memory_is_not_lost_at_the_first_fork: typed on shared
trunk it still survives a fork, and typed on ground the fork never
travelled it no longer follows you.
402 tests. Verified on tools/branch_fixture.py: each branch's drawer holds
its own memory and not the other's.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
Retrieval has been path-scoped since SP3: a memory on a branch this story
never travelled is never sent to the AI. The drawer listed every memory
alike, so on a fork you read "Fell down the cellar stairs; badly hurt" and
reasonably concluded the model knew it. It does not. That is worse than
either hiding the row or retrieving it — it is the screen claiming
something the engine contradicts.
Hiding them is not the answer either; that was the original comment's
point, and it stands. A memory nobody can list is a memory nobody can
delete, in a phase whose rule is that nothing is removed automatically.
So the whole bank still lists, and the rows off the current path are set
back, dashed, and labelled "another branch". MemoryOut.on_path carries it,
computed from the predicate retrieval itself uses rather than a second
spelling of the same idea — two spellings drift, and the failure mode here
is a badge that says the opposite of what the model gets. Pinning does not
override it: the path clause runs before pinning is considered.
The relationship is asymmetric and there is now a test that says so. A fork
borrows its ancestors, so a memory written on the parent is on the fork's
path too; the reverse never is. Worth pinning before somebody makes it
symmetric on the grounds that it looks wrong.
399 tests, three new, egress ceilings intact — the flag costs one id-only
query. Verified in a browser on tools/branch_fixture.py.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
A branch switch does not change the length of the story. It changes which
story it is. Four panels keyed on actions.length and so could not tell the
difference: the Branches panel drew one branch while the reader was already
on a second, Insights showed the prompt built for the path just left, the
script-state drawer kept the other line's numbers, and the Memory Bank did
not notice a deleted branch taking its memories with it.
Only the world-state drawer was right, and only because it happened to
carry stateKey already. They all key on the pair now, and deleting a branch
bumps it too — that is the one operation that changes what is stored
without a turn being played and without the story on the current path
moving by a single action.
tools/branch_fixture.py is the thing that could show it. The stress
fixture's world state is empty, so it cannot answer whether a switch puts
the scoreboard back, and its story is one branch. This builds a small
bootable adventure with a stat schema, a gold script, two takes on one turn
that differ by 35 hit points, a fork, and a memory on each side — with both
branches the same length on purpose, because equal length is precisely the
case a length-based key cannot see.
Verified in a browser with the drawer open: hp 60 to 95 and back, the bar
redrawn, the story swapped to the other take, Insights carrying the scratch
and not the beating. The Memory Bank deliberately does not change on a
switch: the drawer is adventure-wide so a memory is always findable to
delete, and retrieval is the path-scoped half.
396 tests, build clean, no new lint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
The pager could only step between attempts, and stepping has nothing to say
about the thing the tree exists for: taking a path the story moved past and
keeping both. Every attempt has been its own node since SP4, so a chip is a
node now, and "take this path" forks — or simply switches, when the turn is
still the tip and its attempts are leaves nobody has built on.
Beside it, a Branches panel: every line the story has taken, with where each
left its parent and where it ends, and switch, rename and delete-with-confirm.
It sits with Plot/Memory/Scripts/Insights rather than inventing a new place to
put a rail. An unnamed branch is drawn from its fork depth, never from its
position in the list — a position shifts the moment a branch above it goes.
A spatial per-node map was considered and deliberately not built. At the size
this has to be verified against it is a second windowing problem, and it can be
added later without a new endpoint, since the rail and a map read the same
GET /branches. VariantOut grows an id because a fork is addressed by the node
being taken, not by an ordinal in a group that renumbers.
Driven by hand against the 602-action fixture, which found one bug that no test
could: the panel refreshed on actions.length, and a fork swaps a 60-action
window for another 60-action window, so it went on drawing a one-branch tree
while the story was already on the second. It keys off the counter adoptWindow
bumps now.
The scroll path was driven at the same time — three prepends of ~16,200 px, the
same node holding viewport top 792 to 787, never thrown to the end. That closes
the standing gap in this project. Console clean.
396 tests, build clean, no new lint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
A bundle had one list and a forked adventure has two stories, so export was
emitting every branch's turns interleaved by index — a mangled story rather
than lost data, and unreachable only because forking has no UI yet.
`ai-dnd-adventure-v2` carries the branches, the depth each one left its
parent at, which attempt at every turn is the story, and what each node
left behind. That last one is not decoration: the after-snapshots are what
a branch switch puts back, and a bundle without them imports a tree nobody
can switch inside.
`app/bundle.py` owns both formats and nothing else knows either. The v1
reader stays — those files are already on people's disks — and it is now
the only place a `variants` array exists anywhere.
The rule the module is built on is that a bundle carries what was chosen
and never what is derived. The head branch, the fork points, the live flags
and the anchors are decisions somebody made. The lineage, the head depth,
the legacy `index` and the variant ordinals are computed from those and are
rebuilt on the way in, because a bundle is a text file anybody can edit and
a derived field shipped beside its source is a chance for the file to
disagree with itself where no read would report it.
`index` is the one that stops being academic here. It agreed with `depth`
until SP5, and this is the first writer that has to fill it for a forked
story, where two branches both hold a node at depth 4. It is allocated one
per turn instead: siblings share it, no two coordinates do.
Everything a hand-edited file can get wrong about the shape of a tree is a
400 raised before the adventure row exists, because a half-applied import
is exactly the failure this phase exists to end — a story that goes quiet.
A file wrong about which attempt is live is corrected rather than refused;
that is an invariant of the database, not of the format.
Measured on the 600-action fixture: 587 kB to 911 kB, and all of the
increase is the outcomes at 489 B a node — the coordinates themselves save
57.5 B a node against the old turn-and-variants shape. Twenty forks add
660 B. 4.3% of the import body cap.
381 tests green, 16 of them new in test_bundle_v2.py. No migration, no
vacuum owed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
SP4 and SP5 are written down: what shipped, what it measured, and the
handful of things worth not rediscovering. The bug table at the top is
scored now that six of its seven rows are gone — and the seventh, the
one-turn holdback, is gone for a different reason than the one written
there, which is the correction that matters most. Editing a summarised
action is still unfixed and now says so; no subphase is scheduled for it.
The Open section closes. `retry_of.index` was the last item and SP4
answered it by reusing the retried node's depth.
Two vacuums are owed, SP1's and SP4's, and neither has deployed. One run
after the SP4 deploy settles both. SP8 grows two more columns to drop and
a caveat about which ones are not dead yet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
The memory bank and the story summary each kept a cursor: how many story
actions they had already covered. A count is a position in a list, and this
list moves — delete an action in front of the mark and every later one slides
down a slot, so the mark now covers one it has never read. All the cursor
bookkeeping existed to patch that up.
Both marks are now (branch_id, depth): the node up to and including which the
work is done. A depth is a coordinate along a path, not an offset into a list,
so nothing in front of it can move it. That deletes rather than rewrites
`position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
`prune_dangling_memories` and the every-pass clamp in `run_post_turn`.
A memory hangs off the node its block ends on, so a fork inherits its
ancestors' memories without copying any, and retrieval selects through the
branch clause over the *whole* lineage — recall is long-range by definition and
cannot be windowed. Measured: 1,807 B on a story forked twenty times against
1,823 B on a flat one of the same length.
Migrations 53-56 translate the old counts into nodes. They rewrite `adventures`
and not `actions`, so this one needs no VACUUM FULL.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Trust the statement count, not the stopwatch: this machine's suite timings
drift about 20% between runs, so the branch-read regression is recorded as
201 SELECTs -> 2 rather than as seconds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The identity map holds weak references. A branch row nobody keeps a strong
reference to is collected between two nodes, so resolving the head inside the
placement loop read it back from the database for every node in the flush: 201
SELECTs on `branches` to write 200 actions, and 36 s -> 45 s on the same 297
tests. Nothing about any result changed, which is why only a stopwatch found
it, and why there is now a test counting the reads.
Also: the two reads left un-pathed on purpose say so where they live —
`max_action_index` allocates the legacy `index` and must stay adventure-wide
or two branches issue the same number, and export is a flat v1 bundle whose
reader has no idea branches exist. And `Adventure.actions` keeps its `index`
ordering, because ordering the collection by depth would not make it a story:
it is every branch's actions, and a path is a selection out of it.
318 tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a
`branches` table, `branch_id`/`depth` on actions and memories, a head pointer
on adventures, migrations 46-52, and a server-side backfill that re-reads every
existing adventure as a tree with one branch. `depth` holds the number `index`
already held, gaps included, so no story changes — a linear story *is* a tree
with one branch, which is what makes the SP0 baseline passing unmodified the
pass condition rather than a hope.
The writer had to come with it. No migration will ever visit a row written
after it ran, so columns backfilled today and populated next subphase would
leave a hole exactly the width of one deploy, and from SP2 on a row without a
branch is a row no read can see. `app/tree.py` owns that: one module, because a
node written without a branch fails by disappearing rather than by raising.
Three things the schema itself insisted on:
- `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing
both ways makes the two tables a cycle create_all cannot order, and its
escape hatch needs an ALTER SQLite does not have. It is a cache, and a head
naming a branch that is gone recovers onto the root.
- `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a
second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`,
because Postgres `json` has no equality operator.
- SQLite will not drop a column a foreign key names, which is how two existing
tests broke: they simulated an old database by rewinding the stamp while
leaving the new columns in place. Every ADD COLUMN migration is now
idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by
rebuilding three tables from frozen DDL so the real ALTERs run.
297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn;
page load and index are byte-identical to the recorded figures.
The deploy that ships this needs one `VACUUM FULL actions;` on the direct
endpoint afterwards — it rewrites every row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The stress fixture is sized from production and exists to weigh bytes, so it
leaves every column it does not weigh at its default. Checked against a freshly
built one, that is exactly the set phase 14 has to migrate: state_before and
world_state_before NULL on all 600 rows, no scenario and so no RPG layer, no
adventure scripts, both cursors 0, and a hundred retry histories whose two
attempts carry byte-identical text with variant_index pinned at 0.
That last one matters most. "Which attempt is live?" is the question SP4's
migration answers when it decides which sibling node becomes the head of a
turn, and against that fixture the question had no observable answer -- a
migration that picked wrong would have looked correct.
--rich fills in those columns and no others. A real RPG scenario read from the
seed data rather than invented, with a world state played forward so hp
declines and flags flip; monotonic per-action state snapshots, so a rollback
that does not happen reads as a wrong number instead of as nothing; a gold
script, story cards, non-zero cursors, pinned and forgotten memories; retry
attempts with distinct texts, counts of two and three, and a live attempt that
is often not the last written. Plus a second adventure, because a branch clause
that forgot its adventure still looks right on a database holding one.
The invariant SP4 reads -- text mirrors variants[variant_index] -- is asserted
at build time rather than assumed.
The plain fixture is untouched and re-measured unchanged at 1.8 kB, so the
egress ceilings stay comparable. All six shapes run clean on --rich, including
the turn path with the RPG layer and script pipeline now live. 283 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.
--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:
- create_all() builds the current schema but leaves the version stamp at its
default, and bootstrap() reads a populated-but-unstamped database as
ancient — it replays every migration against a schema that already has the
columns, and fails on the first.
- the fixture's user is a registered one, but local mode looks for the row
with email IS NULL and is_guest false, so without clearing the email the
app opens on an empty library.
--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.
Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The post-vacuum sizes had never been read. They were 144.2 MB - above the
99.6 MB the compression work started from, not the ~53 MB it projected. The
earlier VACUUM FULL ran before migrations 42-45, which then rewrote every row
and refilled the table with dead space that only another VACUUM FULL returns.
Ran it: 144.2 -> 65.0 MB, 79.2 MB reclaimed in 5.5s, health green after.
actions is now 52.1 MB holding 50.4 MB of live bytes. The projection was right;
nothing had reclaimed what the compression freed.
Also records where those bytes are. context_snapshot is 94% of the table at
104 kB a row, and the per-action state columns a tree would sit beside are
rounding errors - which is the number phase 14 wants before it adds more.
And a trap: n_live_tup on the TOAST relation implied half the real bloat. It
is a leftover ANALYZE estimate, stale in the flattering direction right after
a migration. Size from sum(octet_length(col)) instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
STATUS was written before the deploy and still asked for a VACUUM that has
since been run. Replaces that with what was actually verified against the
running service, and with the two things a next session would otherwise have
to rediscover.
The post-vacuum sizes were never measured, so the projection stays a
projection; the query to settle it is in the file. And the scroll gap is
promoted from a footnote to the one real open item, because re-reading that
path after shipping found a bug in it -- which is the argument that reading it
again is not the fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Records what the six changes measured, and replaces the pick-up item -- which
was step 6 -- with plan/14, since there is nothing left in 13.
Two things a future session needs and cannot infer from the code. The VACUUM:
migration 43 compresses context_snapshot but Postgres does not return the disk
by itself, so until `VACUUM FULL actions` runs the storage win exists only on
paper and the table is temporarily larger, not smaller. And the gap: nothing
exercises the scroll behaviour in a browser, because the frontend has no test
runner, and prepend-and-restore-scroll is the part most likely to feel wrong
even when it is correct.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
The previous commit recorded the service as suspended. It is not. Render
appended a suffix to the `ai-dnd` service name in render.yaml, and plain
ai-dnd.onrender.com is a different, suspended service whose 503 page reads
"suspended by its owner" -- indistinguishable from this deploy being down
unless you notice the host is wrong.
The real host answers {"ok":true} on /api/health in 0.6s, mints a guest,
returns an empty adventure list for that guest and serves the starter
scenarios. The authoritative link is the one docs/index.html points at, not
the service name in the blueprint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Everything measured so far ran on SQLite against a synthetic fixture, so two
claims were still on trust: that migration 38 spells BYTEA correctly for a real
server, and that the byte figures survive psycopg's encodings.
Both hold. Production reads schema_version 41 with embedding_blob bytea and
embedded boolean present, the backfill is complete at 134/134, and the packed
vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a
memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and
against a throwaway Neon database every shape lands within 0.5% of the SQLite
run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of
it.
The harness writes, so it refuses any target whose name does not say stress or
scratch -- pointed at the production database it stops rather than seeding it
with a fake user and 200 fake turns. It also empties a Postgres target before
building, which a fresh SQLite temp file never needed.
Two corrections fall out, both recorded in plan/13. The page-load model has the
wrong shape: real actions are half the fixture's weight but real stories run to
607 actions, not 200, so the worst real page load is 589.5 kB. And the decision
to leave context_snapshot in the database costed egress but never storage --
it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the
ceiling this deploy will hit first.
Measured with counts and octet_length sums only. No user content was read.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
plan/STATUS.md: what the last session changed, what to pick up next, and the
handful of things that were learned the hard way and would otherwise have to
be rediscovered. Linked from the overview, which now also lists plans 13 and
14 alongside the earlier phases.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7