Commit Graph
17 Commits
Author SHA1 Message Date
Parth e7d75c3b05 Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style

Rewrite the comments and docstrings across the backend core modules so they
read plainly. The previous prose was accurate but dense and figurative, which
made it slow to skim.

Applies the Google developer documentation style guide: short sentences, active
voice, present tense, American spelling, and no metaphors, idioms, or
rhetorical asides. Replaces em-dash chains with separate sentences.
2026-08-26 15:37:25 +05:30
parththakkar106andClaude Opus 5 40d2555f84 Say what the app is now, everywhere it is published
The README, the project page and the engineering guide all describe a linear
story. The tree shipped two days ago. Every published surface is a phase
behind, and the guide is not merely behind — it is wrong in a way that costs a
reader time.

Its 2.2 was "Two coordinate systems, and the bug class they create", and it
explained the codebase through position_of_index, note_action_removed and
settled_story_actions. All three were deleted in SP3. 2.3 explained retry
through Action.variants and state_before. Somebody reading either would go
looking for machinery that is not there, which is worse than a gap.

So 2.2 is now "The story is a tree", written at the depth 1.2 and 1.3 are
written at: the seven bugs that turned out to be one bug, the lineage clause
and the two properties that make fork count free, why takes group by parent_id
rather than by coordinate, cursors becoming anchors, and a closing list of what
the design is honest about. 2.3 is rewritten around state_after and takes, and
1.1 and 1.5 follow, because the pipeline no longer snapshots before the call
and the memory bank no longer holds an action back.

The numbers were simply old: 151 tests where there are 440, 37 migrations where
there are 64, twelve phases where there are fourteen. They appear in four
places across the README, the project page's stat tiles and the guide's results
table. The measured branch cost — 103 B, and 1.007x the page load of the same
story flat — is added beside the egress and turn-cost figures it belongs with,
since it is the number that answers "what does branching cost me".

Three screenshots, on a new tools/shots_fixture.py: the Bandit Camp demo driven
through eight written turns with written deltas, three discarded takes forked
onto branches of their own, one off a branch so the map has to nest. Same
reason tree_fixture.py is committed — the shots have to be reproducible and the
frontend still has no test runner. play-world-state.jpg is reshot because it
predates the entire tree UI; the map and the branches panel are new.

Note for next time: docs/guide.html is hand-written, not generated from the
Markdown, so every guide edit is two edits in two vocabularies. Both files were
checked for tag balance and both pages rendered locally before this landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-20 04:19:53 +05:30
parththakkar106andClaude Opus 5 c8e081e9d6 Draw the tree the branch list can only spell out
The Branches panel says which lines exist. It cannot say where they parted
or how much story each one is, because those are the two numbers `GET
/branches` already answers and a list has nowhere to put them. A ⌗ See the
tree button opens a map: one horizontal lane per branch, running from the
moment it left its parent to the moment it ends, joined to the parent by an
elbow at the fork. The horizontal axis is the story's own clock, so two
lanes at the same x are at the same moment and a short branch reads as
short.

This is a branch map, not the per-node map SP7 refused, and that is the
whole reason it was cheap. Lanes are bounded by branch count, not node
count, so the 600-node windowing problem never arrives. It reads the single
request the rail already made and nothing else.

`branches.js` holds the tree maths and `BranchMap.jsx` the drawing. Play.jsx
gives up its private copies of branchLabel and orderBranches: the panel and
the map now label and order a branch through the same functions, so a branch
cannot be called two things by the two views. The three operations stay in
BranchPanel and are passed down, and `run` answers whether it worked so
neither view clears a half-typed name on a refusal.

Three things came out of driving it, none of which a test could have seen.
`clientWidth` counts the canvas padding the ResizeObserver leaves out, so
the first paint drew an svg 24px wider than its box — and the observer's
initial observation never arrived here, so dropping the seed left the map
never drawing at all. Both are needed and the seed subtracts the padding.
Only the name was being clipped, not the meta line under it, so a late fork
ran its text off the right edge; both are clipped now, and a lane starting
in the right third hangs its labels back over the fork, where its own band
guarantees nothing to collide with. And the delete rule lived in the server
and in the map but not in the list, which offered Delete on a branch the
head was forked from and answered with a toast from the server's 400.
`headLineage` is the client's copy of that rule and both views use it; the
server stays the authority.

`tools/tree_fixture.py` is the counterpart to `tools/branch_fixture.py` —
four branches at three fork depths, one forked off a fork. With no frontend
test runner it is the whole of the map's coverage, and it exists to be
looked at. Nothing in the backend changed; the 440 tests pass unmoved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-20 03:38:17 +05:30
parththakkar106andClaude Opus 5 811368d048 Refresh what a switch changes, not what a turn changes
A branch switch does not change the length of the story. It changes which
story it is. Four panels keyed on actions.length and so could not tell the
difference: the Branches panel drew one branch while the reader was already
on a second, Insights showed the prompt built for the path just left, the
script-state drawer kept the other line's numbers, and the Memory Bank did
not notice a deleted branch taking its memories with it.

Only the world-state drawer was right, and only because it happened to
carry stateKey already. They all key on the pair now, and deleting a branch
bumps it too — that is the one operation that changes what is stored
without a turn being played and without the story on the current path
moving by a single action.

tools/branch_fixture.py is the thing that could show it. The stress
fixture's world state is empty, so it cannot answer whether a switch puts
the scoreboard back, and its story is one branch. This builds a small
bootable adventure with a stat schema, a gold script, two takes on one turn
that differ by 35 hit points, a fork, and a memory on each side — with both
branches the same length on purpose, because equal length is precisely the
case a length-based key cannot see.

Verified in a browser with the drawer open: hp 60 to 95 and back, the bar
redrawn, the story swapped to the other take, Insights carrying the scratch
and not the beating. The Memory Bank deliberately does not change on a
switch: the drawer is adventure-wide so a memory is always findable to
delete, and retrieval is the path-scoped half.

396 tests, build clean, no new lint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 a7bf47a35e Let a backup carry a story that went two ways
A bundle had one list and a forked adventure has two stories, so export was
emitting every branch's turns interleaved by index — a mangled story rather
than lost data, and unreachable only because forking has no UI yet.
`ai-dnd-adventure-v2` carries the branches, the depth each one left its
parent at, which attempt at every turn is the story, and what each node
left behind. That last one is not decoration: the after-snapshots are what
a branch switch puts back, and a bundle without them imports a tree nobody
can switch inside.

`app/bundle.py` owns both formats and nothing else knows either. The v1
reader stays — those files are already on people's disks — and it is now
the only place a `variants` array exists anywhere.

The rule the module is built on is that a bundle carries what was chosen
and never what is derived. The head branch, the fork points, the live flags
and the anchors are decisions somebody made. The lineage, the head depth,
the legacy `index` and the variant ordinals are computed from those and are
rebuilt on the way in, because a bundle is a text file anybody can edit and
a derived field shipped beside its source is a chance for the file to
disagree with itself where no read would report it.

`index` is the one that stops being academic here. It agreed with `depth`
until SP5, and this is the first writer that has to fill it for a forked
story, where two branches both hold a node at depth 4. It is allocated one
per turn instead: siblings share it, no two coordinates do.

Everything a hand-edited file can get wrong about the shape of a tree is a
400 raised before the adventure row exists, because a half-applied import
is exactly the failure this phase exists to end — a story that goes quiet.
A file wrong about which attempt is live is corrected rather than refused;
that is an invariant of the database, not of the format.

Measured on the 600-action fixture: 587 kB to 911 kB, and all of the
increase is the outcomes at 489 B a node — the coordinates themselves save
57.5 B a node against the old turn-and-variants shape. Twenty forks add
660 B. 4.3% of the import body cap.

381 tests green, 16 of them new in test_bundle_v2.py. No migration, no
vacuum owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 0a12d9cd47 Make a retry a node, not a rewrite
Every attempt at a turn is now its own row at the same (branch, depth),
with `live` naming the one the story tells. The JSON repeating group on
`actions.variants` is read one last time, by a migration that writes it
out as the sibling rows it always described, and then goes unread.

The snapshots turn around with it: an action carries the state it left
behind rather than the state it started from, because attempts at one
turn share a starting position and differ exactly in their outcome.
Rolling back is "what the node in front left behind", one lookup on the
path, and it is what undo and retry now both read.

And the memory holdback goes. It existed because retry rewrote a row
under a mark that had already moved past it; a retry writes a sibling
now, and replacing what a coordinate says withdraws what was derived
from it — the same repair undo and delete already made.

The assembled prompt is still stored once per turn: it moves with the
live flag, so a superseded attempt keeps only the few hundred bytes that
were its own. Measured on the 600-action fixture: 700 rows for the same
600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB
and page load 62.7 kB unmoved.

347 tests green. `tests/test_story_tree_baseline.py` and
`tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move
the baseline for the variant-count semantics and did not need to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 c51531709d Mark the story with a node, not with a count
The memory bank and the story summary each kept a cursor: how many story
actions they had already covered. A count is a position in a list, and this
list moves — delete an action in front of the mark and every later one slides
down a slot, so the mark now covers one it has never read. All the cursor
bookkeeping existed to patch that up.

Both marks are now (branch_id, depth): the node up to and including which the
work is done. A depth is a coordinate along a path, not an offset into a list,
so nothing in front of it can move it. That deletes rather than rewrites
`position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
`prune_dangling_memories` and the every-pass clamp in `run_post_turn`.

A memory hangs off the node its block ends on, so a fork inherits its
ancestors' memories without copying any, and retrieval selects through the
branch clause over the *whole* lineage — recall is long-range by definition and
cannot be windowed. Measured: 1,807 B on a story forked twenty times against
1,823 B on a flat one of the same length.

Migrations 53-56 translate the old counts into nodes. They rewrite `adventures`
and not `actions`, so this one needs no VACUUM FULL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 d3756abdaa Give every action a branch and a depth
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a
`branches` table, `branch_id`/`depth` on actions and memories, a head pointer
on adventures, migrations 46-52, and a server-side backfill that re-reads every
existing adventure as a tree with one branch. `depth` holds the number `index`
already held, gaps included, so no story changes — a linear story *is* a tree
with one branch, which is what makes the SP0 baseline passing unmodified the
pass condition rather than a hope.

The writer had to come with it. No migration will ever visit a row written
after it ran, so columns backfilled today and populated next subphase would
leave a hole exactly the width of one deploy, and from SP2 on a row without a
branch is a row no read can see. `app/tree.py` owns that: one module, because a
node written without a branch fails by disappearing rather than by raising.

Three things the schema itself insisted on:

- `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing
  both ways makes the two tables a cycle create_all cannot order, and its
  escape hatch needs an ALTER SQLite does not have. It is a cache, and a head
  naming a branch that is gone recovers onto the root.
- `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a
  second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`,
  because Postgres `json` has no equality operator.
- SQLite will not drop a column a foreign key names, which is how two existing
  tests broke: they simulated an old database by rewinding the stamp while
  leaving the new columns in place. Every ADD COLUMN migration is now
  idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by
  rebuilding three tables from frozen DDL so the real ALTERs run.

297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn;
page load and index are byte-identical to the recorded figures.

The deploy that ships this needs one `VACUUM FULL actions;` on the direct
endpoint afterwards — it rewrites every row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 5c1bcf7305 Build a fixture that can witness what a tree changes
The stress fixture is sized from production and exists to weigh bytes, so it
leaves every column it does not weigh at its default. Checked against a freshly
built one, that is exactly the set phase 14 has to migrate: state_before and
world_state_before NULL on all 600 rows, no scenario and so no RPG layer, no
adventure scripts, both cursors 0, and a hundred retry histories whose two
attempts carry byte-identical text with variant_index pinned at 0.

That last one matters most. "Which attempt is live?" is the question SP4's
migration answers when it decides which sibling node becomes the head of a
turn, and against that fixture the question had no observable answer -- a
migration that picked wrong would have looked correct.

--rich fills in those columns and no others. A real RPG scenario read from the
seed data rather than invented, with a world state played forward so hp
declines and flags flip; monotonic per-action state snapshots, so a rollback
that does not happen reads as a wrong number instead of as nothing; a gold
script, story cards, non-zero cursors, pinned and forgotten memories; retry
attempts with distinct texts, counts of two and three, and a live attempt that
is often not the last written. Plus a second adventure, because a branch clause
that forgot its adventure still looks right on a database holding one.

The invariant SP4 reads -- text mirrors variants[variant_index] -- is asserted
at build time rather than assumed.

The plain fixture is untouched and re-measured unchanged at 1.8 kB, so the
egress ceilings stay comparable. All six shapes run clean on --rich, including
the turn path with the RPG layer and script pipeline now live. 283 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
ParthandClaude Opus 5 0f1e05c808 Keep the fixture, so there is something to scroll (#5)
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.

--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:

  - create_all() builds the current schema but leaves the version stamp at its
    default, and bootstrap() reads a populated-but-unstamped database as
    ancient — it replays every migration against a schema that already has the
    columns, and fails on the first.
  - the fixture's user is a registered one, but local mode looks for the row
    with email IS NULL and is_guest false, so without clearing the email the
    app opens on an empty library.

--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.

Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.


Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 16:50:41 +05:30
parththakkar106andClaude Opus 5 ae6e5af6c7 Store context_snapshot compressed
One column is 89% of the database and the free tier allows 512 MB. Reads were
already solved -- the column is deferred, so a page load never touches it and
one screen fetches one row at a time -- but nothing had costed storage, and
storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per
action, so the ceiling arrives around 5,400 actions and 944 are stored.

Postgres already compresses it and only gets 1.7x. pglz is tuned for fast
decompression of data a query might filter on, and nothing has ever filtered
on an assembled prompt -- it is written once and read whole, rarely, by the
Insights viewer. zlib gets 3.5x on the same text for a decompress on a request
that already made an LLM call.

Done as a TypeDecorator rather than a second column, so every call site still
writes a dict and reads a dict back, and deferred/undefer/load_only keep
naming the same attribute. Only the storage format moves.

Migrations 43-45: add the bytea, convert into it, drop the original, rename.
The backfill is the one destructive step in the file -- 44 removes the only
other copy -- so it decompresses every row and compares it against what went
in, and a row that fails aborts the run. The whole loop is one transaction, so
an abort rolls the DROP back and the prompts are still there.

Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway
Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column
came out named context_snapshot, every snapshot compared equal and the one
NULL stayed NULL.

Postgres does not return the disk by itself: DROP COLUMN only marks the column
gone and the backfill leaves a dead tuple per row, so the table peaks near
twice its size before settling. The deploy needs one VACUUM FULL to collect
it; the migration comment says so.

The egress fixture's snapshots are prose now rather than "x" * 20_000, and the
prose generator moved to tools/fakeprose.py so the harness and the tests share
one definition. A repeated character compresses a thousandfold: against the
old fixture a compressed column looked free and the byte ceilings would have
been guarding nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 14:08:23 +05:30
parththakkar106andClaude Opus 5 12d57afdac Put a number on what an endpoint may fetch, not just a column list
test_egress.py asserted which columns a statement names, which is the shape
both of this project's egress blowouts took. It would all still pass if a
response grew tenfold within the columns it is allowed to read -- and a story
that keeps getting longer does exactly that. Production's longest adventure is
607 actions where the plan assumed 200.

So dbmeter, which was built to be importable from tests and was not yet used
by any, now backs four byte ceilings: the page load, the action list, and one
action's snapshot fetched on demand. Budgets are per action rather than
absolute, so they mean the same thing whatever size the fixture is set to, and
generous -- 3 kB against a real 994 B. They are there to catch an order of
magnitude, not to freeze a byte count.

The fourth test is the one that keeps the other three honest. A ceiling proves
nothing unless the thing it excludes would breach it, so it undefers the
snapshot on purpose and asserts the same twelve rows cost more than ten times
the budget. If the fixture ever shrinks below the point where that holds, that
test fails rather than the ceilings quietly passing on nothing.

Meter grows detach() and a context manager. A script exits and takes the
wrapping with it; a test does not, and one test leaving the shared engine
metered would charge bytes to a scope nobody opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 13:51:10 +05:30
parththakkar106andClaude Opus 5 be66780a26 Size the stress fixture from what production actually holds
The old fixture was wrong in both directions at once and happened to land near
the right total. Actions were modelled at ~2.1 KB against a real 886 B of
text, and adventures at 200 actions against a real 607. Width flattered,
length did not, and length is what a page load pays for.

Re-sized from the 2026-08-17 measurements: 600 actions, 1700 B of narration
alternating with a one-line player input, 232 KB of context_snapshot a row.
The page-load shape now reports 606.0 kB against the 589.5 kB measured on
production's longest adventure -- 2.8% out, where the old defaults were 28%
out on a story a third of the length.

Filler text is now generated word by word instead of one sentence repeated.
That matters for what comes next: the repeated string compresses 313x and the
generated prose 3.7x, so any compression ratio measured against the old
fixture would have been fiction, and shrinking context_snapshot is the open
question it exists to answer.

context_snapshot also gains a flag of its own rather than being hardcoded, and
the 74 KB figure in the comment -- inherited from models.py -- is corrected:
the real column averages 163 KB a row across the table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 13:48:23 +05:30
parththakkar106andClaude Opus 5 70d024da62 Check the egress work against the database it actually runs on
Everything measured so far ran on SQLite against a synthetic fixture, so two
claims were still on trust: that migration 38 spells BYTEA correctly for a real
server, and that the byte figures survive psycopg's encodings.

Both hold. Production reads schema_version 41 with embedding_blob bytea and
embedded boolean present, the backfill is complete at 134/134, and the packed
vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a
memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and
against a throwaway Neon database every shape lands within 0.5% of the SQLite
run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of
it.

The harness writes, so it refuses any target whose name does not say stress or
scratch -- pointed at the production database it stops rather than seeding it
with a fake user and 200 fake turns. It also empties a Postgres target before
building, which a fresh SQLite temp file never needed.

Two corrections fall out, both recorded in plan/13. The page-load model has the
wrong shape: real actions are half the fixture's weight but real stories run to
607 actions, not 200, so the worst real page load is 589.5 kB. And the decision
to leave context_snapshot in the database costed egress but never storage --
it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the
ceiling this deploy will hit first.

Measured with counts and octet_length sums only. No user content was read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 12:37:23 +05:30
parththakkar106andClaude Opus 5 b7e53ae581 Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.

Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.

    one turn      3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
    run_post_turn 3,139.1 kB -> 0.7 kB
    Insights      3,223.7 kB -> 117.9 kB
    Memories drawer  ~3.1 MB -> 23.6 kB

A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.

The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.

memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.

Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:28:00 +05:30
parththakkar106andClaude Opus 5 c56864877a Store embeddings as packed float32 instead of a JSON list
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.

It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.

Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.

The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.

Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:17:44 +05:30
parththakkar106andClaude Opus 5 7ee5ceea6c Measure database egress in bytes, from inside the repo
Both egress blowouts this project has had were one query fetching a column
nobody read, and a statement count would have shown nothing wrong in either.
So the meter counts bytes, at the DBAPI cursor -- everything that crosses
that line crossed the wire.

tools/stress_session.py drives a production-shaped adventure through the real
routes with only the network faked. It reproduces both figures measured
directly on production: 426.7 kB for a 200-action page load against 423 KB,
and 3,258.7 kB for one turn against 3,153 kB.

The memory bank is on by default, which is the whole point -- the previous
harness ran without an embedding model, so retrieval returned early and the
heaviest read in a turn never happened. --no-embeddings reproduces that
deliberately, and the gap is 29x.

It also turned up two callers the production SQL could not see: run_post_turn
walks the whole bank again every turn, and Insights pays for it a third time.
A played turn costs ~6.4 MB, not 3.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:11:56 +05:30