f8d401029f12c4a9205c5a0cb90083ce79da74b3
36
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f8d401029f |
Stop a turn locking out its own memory bank, and let the long run notice
The first M01 trial with the memory bank on was 26 turns on a GPU host. It
accepted every turn and reported "complete". It also wrote two memories and
no summary, and logged 180 `database is locked` errors, while derived status
still read `idle`.
The cause was a single uncommitted UPDATE. Retrieval bumped each used
memory's counter before the model call, and the turn commits only after the
reply has streamed. SQLite has one writer, so the turn held the write lock for
the whole reply. Every post-turn memory, summary and status write in that
window waited out the five-second timeout and failed. Recording the failure
needed a write as well, and without a rollback first it raised
PendingRollbackError. The loss therefore reached the log and never reached
the status the Insights panel reads, which F08 forbids. The draco run never
hit this because the bank was off there.
- `retrieve_memories` now only reads. `record_use` writes the counters in the
turn's single commit, so a turn that never lands counts nothing.
- The post-turn task's outer handler rolls back before it records a failure.
The harness could not have caught any of this. It read three prompt sections
under names the builder does not use: `memories` (really `used_memories`),
`story_history` (really `history`/`recent_history`), and a `knowledge` prefix
that matched the fixed instruction section instead of the imported passages.
Memory tokens read 0 whatever the prompt held, and the in-history and
in-memories recall checks could never come out true. The labels are now
constants, pinned by a test against a prompt the real builder assembled.
The harness also stops at the first sign of failed post-turn work. It checks
/derived and new server.log lines after every turn, keeps its log position
across --resume, and waits for background work to settle before its final
checks. A run with no memories or no summaries now ends "failed", not
"complete".
Both new application tests fail on
|
||
|
|
fec46f66bb |
Turn the memory bank on for the long run, and refuse one that cannot use it
The first complete hundred-turn campaign did not exercise M01's "summary/memory activation" step. Memory bank and auto-summarize are per-campaign switches that default to off, and m11_long_run never turned them on: summary_tokens and memory_tokens were 0 on every turn, memories_used was empty, and M04's clue was recalled through narrative state alone. The retrieval path M6 built was never asked, and nothing in the evidence said so except a row of zeros. setup now PATCHes both switches on, reads the campaign back, and stops before the first turn if either did not take. memories_in_bank is recorded on every turn, in the final summary and in the recall, and the recall also says whether a summary exists, so which of the two recall paths succeeded is stated rather than implied. The embedding model is now required. Without one the summary pass still writes memories, but memorybank.retrieve answers "No embedding model configured" and returns none -- the same unexercised path in a fuller bank. The harness refuses before it starts a server or claims --out. tests/test_m11_long_run_memory.py drives setup against the real application in-process: the switches are on afterwards, a server that ignores the PATCH is refused before any state is written, the bank count comes from the application and reads -1 rather than raising when it cannot, and a run with no embedding model is refused. The four that exercise setup and the bank count were run against the previous harness and fail there; the premise test (a fresh campaign has both switches off) passes on both, as it should. The 2026-09-10 run in ~/m11-evidence/m01 therefore does not count as M01. It has to be run again on this harness. Backend 1,382 passed, 18 skipped, 0 failed. The eighteenth skip is test_built_spa_fetches_no_fonts_remotely, which wants a built frontend/dist this worktree does not have; it is an environment condition, not a change here. The frontend is untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XKWHt2DXuvqP83cAk6Zq88 |
||
|
|
ef25b0a876 |
Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB |
||
|
|
144406cd48 |
M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was whether any of the earlier evidence meant what it said. M8 measured a deployment enforcing a 4,096-token input window while the application budgeted 16,384. Every request returned 200. What Ollama does with the excess is drop the oldest tokens, and the oldest tokens here are the system block — the narrator's rules and the campaign canon. A hundred-turn certification against that server would have looked perfect and proved nothing, which is why this milestone could not begin with a hundred turns. So the application asks now. Ollama's window is a property of how a model was loaded rather than of the request — sending num_ctx is accepted, ignored, and worse, reloads the model at the server's own default — so the only honest move is to find out and then tell the truth about it. /api/ps reports what a resident model is being served with, /api/show what an unloaded one will load with, both on the same host inference already uses, through the same endpoint policy and the same TLS trust store. A verified window is a ceiling on the budget; an unverified one leaves the budget alone and is recorded as unverified in the turn's own provenance, so an old turn can be asked afterwards whether it was built against a checked window. There is no third behaviour, and in particular no hard-coded 4,096: a number the server did not say would be right on one machine and wrong on the next. The proof that this is doing something is a campaign whose canon sits at the front of the prompt, 120 turns of history, and a 4,096-token window. The canon is still there afterwards and the oldest history is gone. The same campaign built the old way produces a prompt more than twice the window — the defect, reproduced, so the fix is measured against it rather than asserted. Two defects the validation found on its own, and they are the same defect twice: something was true and nobody was told. A manual state correction of four changes with one bad reference applied three, returned 201, and said nothing — while recording the refusal on the audit row nobody reads. It came to light because the identity diagnostic's own fixture was refused that way and the whole run proceeded on a campaign with no scene, which would have read as a model failure. And the narration-length setting moved no number: brief, medium and long each became one English sentence, while the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three. Both now say what they did. The other two post-M8 findings are closed as well. The tab said AI D&D, which no document had ever claimed it did not; it says Interactive Story now, with the open campaign first, and the name is the owner's decision rather than a find-and-replace to something narrower than the engine. After an Undo the reader could not tell where they had landed; the control row now ends with "Moment 11 · later story ahead", from the server's own answer, in the word the transcript already uses, with none of head, branch or depth anywhere near it. The identity diagnostic exists and the root cause does not. That campaign was destroyed, so no cause can be established — what M11 owes the finding is something that can classify the next occurrence, and a diagnostic that makes only the judgements a program can honestly make: duplicate keys, shared names, protagonist drift, state and context disagreeing. Whether prose misattributed a line is left to a person reading it beside its prompt, because a regex cannot read dialogue and one that pretended to would produce exactly the confident wrong answer this finding is about. Its detectors are proved to fire against a planted second Alice. Two entities may still share a display name. That was checked first, as the finding asked, and left permitted: a mother and a daughter, or a stranger giving a false name, are ordinary fiction, and refusing them to guard against a model mistake would refuse the wrong thing. What was missing was that it happened silently. It is reported now. Evidence, not inference: a hundred accepted turns against a real narrator with genuine process restarts; a real browser against the built SPA; a container with no network at all; a campaign moved into a data directory that never existed. Each was discarded and re-run whenever the product changed under it, and the runs that were thrown away are listed in the report with the reason, along with ten defects in the harnesses themselves — because a harness that has only ever agreed with itself is not evidence, and two of M8's five harness defects were masking real ones. No dependency was added, removed or upgraded. No acceptance test was retired, relaxed or reclassified. M11 is implemented and verified; it is not accepted, and there is no release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B |
||
|
|
1013c94eb1 |
M10: the seam for media, and no media
The media extension contract asks for a scene snapshot a future image or video provider could be handed: location, who is present, what they hold, what must stay true, and where in the story it sits. Building one was the milestone's obvious first task, and it was the wrong one. That snapshot has existed since M5. `narrative_state["scene"]` holds the summary, the location, the cast and the coordinate it was written at; a validated `set_scene` event writes it, every position snapshots it, and every head move restores it. It survives Undo, Redo, Retry, divergence, Save Point restore and a process restart because it is the authoritative state rather than a copy of it. So there is no scenes table here. A second scene store would have been a second answer to "where is the story now", with its own lineage rules to get wrong — and the lineage rules are the expensive part, which is the argument for reusing the ones that already work rather than against it. The Scene Packet is derived on read, and its identity is computed from the campaign and the position rather than allocated: the same position yields the same id in another process, after a restart, and after the packet is thrown away and rebuilt, with no row to keep in step. That is the part of a future media_assets table that would be expensive to retrofit, so it is fixed now even though the table is not built. One table, then: visual_profiles, the only thing the contract's scene list asks for that nothing already stored. Campaign-scoped and not per-position, because a character does not change appearance when the story forks — a reader who diverged would otherwise lose their cast, and the same descriptors would land in every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn campaign to say something that never varies. Keyed by the M5 entity key rather than a new identity namespace, and one table for characters, locations and items alike, because a location is an entity with a type and splitting them would reintroduce the genre shape M5 spent a milestone removing. What the packet leaves out is the more interesting half. Not the transcript, and not imported knowledge — none of it, not merely the sources marked hidden. The rule is what the story established at this position, not everything the narrator was told, and drawing it by class is what makes it hold for a secret nobody thought to mark. A hidden Canon source proves it, with a positive control showing the narrator did receive the sentinel the packet does not carry. Once a validated event puts the observer in the room, the observer is in the packet: that is no longer narrator-only knowledge, and a packet that hid it would be hiding the story from itself. The providers are contracts and nothing else. Protocols for image, video, audio, speech and transcription, an empty registry, no adapter, no dependency, no socket, and no media setting to point anywhere — a setting that exists can be pointed at a cloud by mistake. A future provider endpoint must be loopback, stricter than narration's trusted-LAN allowance, because a picture of a scene carries the scene with it. Transcription returns an editable draft with no commit method, so STT structurally cannot bypass the authoritative path. Nothing here can write the story. Not by convention: no module under media/ imports the code that writes state, no media event type exists in the state vocabulary, and every test in the authority suite compares the authoritative document byte for byte either side of a media operation — including one where a provider insists Alice is in a red coat in a corridor, and the campaign goes on disagreeing. One defect, found by the milestone's own tests. M10 first added a migration creating an index that create_all already builds from the column, so an upgraded database ended up with two indexes and a fresh install with one. Comparing the two schemas is what caught it; neither database examined alone would have. The migration is gone rather than renamed, and the right number of migrations for a new table whose indexes are declared on its columns is zero. Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145 passed. Lint, production build and Docker build clean. No frontend file changed: M10 adds no reader-facing surface, and ordinary play — turns, state, memory, knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media configuration, no warning, no connection attempt and no media row written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B |
||
|
|
44edece67e |
M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B |
||
|
|
d27ee34901 |
Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was authoritative. Phase 0 execution prompts sat beside the specification; four completed milestone reports sat beside the current one; and upstream AI-DnD's own `plan/` build log and `docs/` project site still described a hosted, scripted, multi-user product with accounts — every screenshot in it showed a Scripts tab and a Sign up button, none of which has existed since M2. `planning/archive/` now holds the history and says so in its own README: `phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied. `planning/reports/` holds only the current milestone's report, because that is the one M4 planning has to read; it moves to the archive when M4's replaces it. Deleted rather than archived: the Phase 0B execution prompts and the handoff/status/summary documents, the Phase 0A discovery and triage reports, upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's template boilerplate. All of it is in Git history, and the two recommendation reports carry every conclusion the deleted research reached. Archived documents are kept verbatim. Paths written inside them point at where those files were when the document was written, which is the point: an evidence record that has been quietly edited is no longer evidence. Active documentation is corrected where it pointed at the removed trees or described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not touch" list had gone stale at M2 and claimed QuickJS scripting was still tested; its test count was 604 against an actual 638. `README.md` loses the upstream CI badge, which reported upstream's pipeline rather than this fork's, and a reference to `backend/app/worldstate/engine.py`, a file that does not exist. `planning/README.md` is rewritten as the documentation index. New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the manifest of what belongs in the ChatGPT project's Sources. Source comments referring to the deleted trees are reworded; no behaviour changes. 638 backend tests pass, frontend lints and builds, and a reference scan over all 48 tracked Markdown files reports no unresolved path in active documentation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu |
||
|
|
8c65ae99de |
M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo. |
||
|
|
745a4ea9f3 |
Stop printing the database password
The report opened with the connection string it was about to work on, taken straight from AIDND_DATABASE_URL. On the hosted deploy that string carries the Neon password, so the first line of every run put a live credential into the console — and from there into scrollback, a screenshot, or a pasted bug report. Nothing in the output said it was there to notice. It now prints scheme, user, host and database name. The query string goes whole: sslmode is the only part worth reading, and some drivers accept a password there as well. 631 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
eeab8ef5bb |
Survive the console this will actually be run from
The report prints an arrow between the old and new word counts, and it prints the memories themselves, which are model-written prose. A Windows console defaults to cp1252 and cannot encode either, so the run dies on UnicodeEncodeError partway through — after some memories have been rewritten and committed, which is the worst place to stop. STATUS already carries the same warning for tools/dbmeter.py. stdout is reconfigured to UTF-8 with errors="replace" instead, so the report degrades a character at a time rather than failing. The usage notes were also written for a POSIX shell, and this project is developed in PowerShell, where `VAR=x command` is not a thing. Both docs now give the PowerShell form first, via Read-Host so the Neon URL and the secret stay out of the history file, and name the venv's Python. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
4c94dd0fc0 |
Keep a production run to the adventures you meant
The hosted database is not the local one. It holds other people's stories, and `summary_provider` builds from the adventure owner's Settings, so an unfiltered `--write` against it would spend other people's money rewriting memories they never asked about. `--adventure` could already hold a run down, but only if you knew the ids. `--email` names accounts instead, and every line of the report now says who owns the adventure, so a dry run answers "whose keys would this spend" before anything is written. Guests have no email and stay reachable only by id, which is the right amount of friction for touching a stranger's bank. The other half is written down rather than built: stored API keys are encrypted with AIDND_SECRET_KEY, so a run from a checkout against the Neon database needs the same value the web service holds. With a different one `decrypt_secret` returns "" instead of failing, and every adventure is reported as having no key — a run that looks like it worked and did nothing. The Dockerfile copies backend/app alone, so tools/ is not on the box either way; the recipe is a checkout pointed at AIDND_DATABASE_URL. 629 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
0633cb624e |
Run the new memory prompt back over an old bank
The prompt change only reaches memories written after it. plan/18 decided to leave the existing ones alone and let eviction age them out at memory_bank_capacity, on the grounds that re-summarizing would duplicate whatever was still in the bank because nothing deletes the old rows. That was wrong about the only option. A memory can be rewritten in place. The row carries more than its text — whether it is pinned, how often it has been retrieved, and the node it hangs off, which is what makes a fork inherit the right memories — and rewriting `text` keeps all of it. Deleting the bank and rewinding the cursor would lose that, and would trickle memories back at MAX_MEMORIES_PER_RUN per turn, so an adventure nobody is playing would never recover. tools/rewrite_memories.py does it. Without --write it makes no model calls and only reports the scope; --write rewrites, --embed re-embeds in the run rather than leaving it to the app's post-turn pass. It reads whichever database the app reads, so it works against the hosted Postgres as well as a local file. Two things it needed from the app. `summarize_block` is now the one place a memory prompt is assembled, and the post-turn pass calls it too — a backfill that built its own prompt would be writing memories with a prompt that never shipped, and nothing would report the drift. `source_block` reads a memory's block back out of the story, which nothing has ever had to do: it reads on the lineage of the branch the memory was written on, not the branch being played, because after a fork the same depths hold different actions on each side and a read through the adventure's path would summarize the wrong story silently. It also excludes the discarded attempts at a retried turn, and tolerates a block an action has since been deleted from. Left alone: a memory with no source range, which is hand-written or migrated by 62 and may be the player's own words; a memory whose actions are gone; and an adventure whose owner has no API key, because summarization spends the user's own key by construction and never the shared demo key. --api-key/--model/ --endpoint override that, the last of them aiming a run at claude_shim.py. The vector is cleared for every rewrite, because the stored one describes wording that no longer exists. Re-embedding always uses the owner's own embedding model, never --endpoint: a vector only means anything against the vectors it is ranked beside. 17 tests, 627 green. The fork case is the one that would fail quietly, so the test builds a fork whose depths hold different actions on each side. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
712ef44f57 |
Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and would have been gone with the container. The numbers in plan/18 were therefore assertions nobody could check. plan/18-appendix-memory-ab-run.md now carries the whole transcript: both memories, both summaries, and the thirteen-action story they were written from. backend/tools/memory_ab.py reproduces it. It replaces the throwaway script the first run used, and differs in two ways that matter. It goes through OpenAICompatibleProvider rather than calling a model directly, so a run exercises the provider, the streaming path and complete() instead of a stub. And it reads the control prompt out of git at the commit given to --before, so the thing being compared against cannot drift from what actually shipped. There was already a claude_shim.py serving an OpenAI-compatible endpoint backed by the CLI, which is exactly what the throwaway script had reinvented. memory_ab.py points at it by default, so a run spends a Claude subscription rather than API credit, and --endpoint aims it at the provider the deployed app really uses — which is the one question this whole exercise could not answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
6a5d8e98f9 |
Give the stress fixture back its sibling attempts
`place_action` used to read `depth` from `index`, so the fixture's retry attempts all landed on the turn's own depth and formed one sibling group. Now that depth comes from the head, each attempt moved the head one step and every retry became a turn of its own. The fixture's own check caught it: a coordinate holding a single superseded attempt has no live row, and that turn disappears from the story. The attempts at one turn share that turn's coordinate, so only the first one goes through `place_action` and the rest copy its placement. This is what `attempts.add_attempt` does. The fixture cannot call it directly, because it makes the newest attempt live and the fixture needs a live attempt that is often not the newest. Also drop the `variant_index` and `variant_count` arguments, which raise `TypeError` now, and read the two mark positions from a helper instead of from the adventure columns migrations 72 and 73 removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV |
||
|
|
f1bebe18d0 |
Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`, `variant_count`, `state_before`, and `world_state_before`, plus `adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in SQLite, so migration 71 quotes it. Nothing outside the migrations read these. `models.py`, `schemas.py`, and `ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders by `id`, and `attempts.renumber`, `context.history.max_action_index`, and `nodes.next_index` are deleted. Two changes keep the migration replayable on a `create_all` database: - `_split_variants_into_siblings` wrote through the live ORM table, so it stopped compiling once migration 66 removed five of its columns. It now writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`. - Five data passes read columns these migrations drop. Each now calls `_has_columns` and returns early when the columns are absent. `bootstrap` takes a `through` version so a migration test can stop at the schema it asserts on. 555 tests pass, up from 549. Eight of the new cases assert each column is gone after a real schema-45 database migrates all the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK |
||
|
|
0623dd8b78 |
Split Play.jsx into a directory of twelve files
`frontend/src/pages/Play.jsx` was 2280 lines holding 27 components. It is now `frontend/src/pages/Play/`: `index.jsx` with the page component, five panels, two drawers, the reports, the take pager, the refresh dialog, and the shared formatting helpers. Every line moved verbatim. A coverage check confirms every non-blank line of the original appears exactly once, in order, across the twelve files, and a name-resolution check confirms every identifier each file references is defined or imported there, with no unused imports. The line split stranded a comment at six of the boundaries. A leading comment sits above the section it describes, so each cut left one at the end of the file before it. All six moved to the section they describe, rewritten in the house style. `usePlaySession.js` is not here. The page component still owns all of the session state. Moving eighteen `useState` calls and seven `useEffect` calls is a rewrite rather than a move, and no frontend test would catch a mistake in it today, so it waits for Stage 5. Four stand-in providers under `backend/tools/` now define `last_usage = None`. The turn engine reads that attribute after every call, and the fixtures never defined it, so `tools.tree_fixture` crashed. That break predates this branch. Verified by 549 passing tests, a clean `npm run lint` and `npm run build`, and by driving the Play screen: the story view, all five panels, both drawers including the world-state edit form, the branch map, the refresh dialog, and the take pager stepping onto a take that lives on another branch. No console errors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
2fa812c056 |
Split the adventures router into a package
`backend/app/routers/adventures.py` held 2353 lines and 35 endpoints. It is now a package of 14 modules, the largest 443 lines. The split moves text rather than rewriting it. An AST comparison against the old file confirms all 86 definitions are identical, and the OpenAPI schema still lists the same 35 operations. Names a test replaces now live in `turns.py` only, and other modules reach them as `turns.<name>`. Rebinding a re-exported alias changes the alias and leaves every caller reading the original, so the package root does not re-export them. A patch aimed at the old target raises `AttributeError` instead of passing while doing nothing. Tests and the fixtures in `backend/tools/` say `adventures.turns.<name>`. The same rule keeps the turn lock working. One module owns `_active_turns`, so one lock guards one set. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
ef4da7fda2 |
Keep the local Claude test rig, and record what it found
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by the local `claude` command line tool, so a demo can be played against a real model with no API key. Each request spawns one `claude --print` process, which suits the turn engine: the app assembles the whole prompt every turn and expects a stateless endpoint. Playing the Pokemon demo through it made all five of the world-state changes in `plan/16` work, and the refusal loop ran end to end for the first time: turn 3 clamped to nothing, turn 4's assembled prompt carried the correction verbatim, and the model's next delta was right. Two failures previously blamed on the code were the demo model. One is a schema fault and is still open: Milo's three Pokemon share one `active_hp` stat, so a switch leaves the newcomer at 0 HP. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
ae39514d1e |
Name the Pokemon demo, and land a seed rename on its own row
`seed.py` matches a seed file to its scenario by title, so renaming one inserted a second public scenario and stranded the first. The stranded row stays public forever and has to be deleted by hand on every deployment, which is what happened to "Road to the Champion". A seed file now lists its old titles under `previous_titles`, and the rename updates the existing row. The cover art is a PNG data URI. `app/images.py` accepts raster formats only, because SVG can carry script and the bytes are served from the app's own origin, so an SVG stores but yields an empty `image_url`. `tools/make_pokeball.py` draws the ball with `zlib` alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
40d2555f84 |
Say what the app is now, everywhere it is published
The README, the project page and the engineering guide all describe a linear story. The tree shipped two days ago. Every published surface is a phase behind, and the guide is not merely behind — it is wrong in a way that costs a reader time. Its 2.2 was "Two coordinate systems, and the bug class they create", and it explained the codebase through position_of_index, note_action_removed and settled_story_actions. All three were deleted in SP3. 2.3 explained retry through Action.variants and state_before. Somebody reading either would go looking for machinery that is not there, which is worse than a gap. So 2.2 is now "The story is a tree", written at the depth 1.2 and 1.3 are written at: the seven bugs that turned out to be one bug, the lineage clause and the two properties that make fork count free, why takes group by parent_id rather than by coordinate, cursors becoming anchors, and a closing list of what the design is honest about. 2.3 is rewritten around state_after and takes, and 1.1 and 1.5 follow, because the pipeline no longer snapshots before the call and the memory bank no longer holds an action back. The numbers were simply old: 151 tests where there are 440, 37 migrations where there are 64, twelve phases where there are fourteen. They appear in four places across the README, the project page's stat tiles and the guide's results table. The measured branch cost — 103 B, and 1.007x the page load of the same story flat — is added beside the egress and turn-cost figures it belongs with, since it is the number that answers "what does branching cost me". Three screenshots, on a new tools/shots_fixture.py: the Bandit Camp demo driven through eight written turns with written deltas, three discarded takes forked onto branches of their own, one off a branch so the map has to nest. Same reason tree_fixture.py is committed — the shots have to be reproducible and the frontend still has no test runner. play-world-state.jpg is reshot because it predates the entire tree UI; the map and the branches panel are new. Note for next time: docs/guide.html is hand-written, not generated from the Markdown, so every guide edit is two edits in two vocabularies. Both files were checked for tag balance and both pages rendered locally before this landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
c8e081e9d6 |
Draw the tree the branch list can only spell out
The Branches panel says which lines exist. It cannot say where they parted or how much story each one is, because those are the two numbers `GET /branches` already answers and a list has nowhere to put them. A ⌗ See the tree button opens a map: one horizontal lane per branch, running from the moment it left its parent to the moment it ends, joined to the parent by an elbow at the fork. The horizontal axis is the story's own clock, so two lanes at the same x are at the same moment and a short branch reads as short. This is a branch map, not the per-node map SP7 refused, and that is the whole reason it was cheap. Lanes are bounded by branch count, not node count, so the 600-node windowing problem never arrives. It reads the single request the rail already made and nothing else. `branches.js` holds the tree maths and `BranchMap.jsx` the drawing. Play.jsx gives up its private copies of branchLabel and orderBranches: the panel and the map now label and order a branch through the same functions, so a branch cannot be called two things by the two views. The three operations stay in BranchPanel and are passed down, and `run` answers whether it worked so neither view clears a half-typed name on a refusal. Three things came out of driving it, none of which a test could have seen. `clientWidth` counts the canvas padding the ResizeObserver leaves out, so the first paint drew an svg 24px wider than its box — and the observer's initial observation never arrived here, so dropping the seed left the map never drawing at all. Both are needed and the seed subtracts the padding. Only the name was being clipped, not the meta line under it, so a late fork ran its text off the right edge; both are clipped now, and a lane starting in the right third hangs its labels back over the fork, where its own band guarantees nothing to collide with. And the delete rule lived in the server and in the map but not in the list, which offered Delete on a branch the head was forked from and answered with a toast from the server's 400. `headLineage` is the client's copy of that rule and both views use it; the server stays the authority. `tools/tree_fixture.py` is the counterpart to `tools/branch_fixture.py` — four branches at three fork depths, one forked off a fork. With no frontend test runner it is the whole of the map's coverage, and it exists to be looked at. Nothing in the backend changed; the 440 tests pass unmoved. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
811368d048 |
Refresh what a switch changes, not what a turn changes
A branch switch does not change the length of the story. It changes which story it is. Four panels keyed on actions.length and so could not tell the difference: the Branches panel drew one branch while the reader was already on a second, Insights showed the prompt built for the path just left, the script-state drawer kept the other line's numbers, and the Memory Bank did not notice a deleted branch taking its memories with it. Only the world-state drawer was right, and only because it happened to carry stateKey already. They all key on the pair now, and deleting a branch bumps it too — that is the one operation that changes what is stored without a turn being played and without the story on the current path moving by a single action. tools/branch_fixture.py is the thing that could show it. The stress fixture's world state is empty, so it cannot answer whether a switch puts the scoreboard back, and its story is one branch. This builds a small bootable adventure with a stat schema, a gold script, two takes on one turn that differ by 35 hit points, a fork, and a memory on each side — with both branches the same length on purpose, because equal length is precisely the case a length-based key cannot see. Verified in a browser with the drawer open: hp 60 to 95 and back, the bar redrawn, the story swapped to the other take, Insights carrying the scratch and not the beating. The Memory Bank deliberately does not change on a switch: the drawer is adventure-wide so a memory is always findable to delete, and retrieval is the path-scoped half. 396 tests, build clean, no new lint. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g |
||
|
|
a7bf47a35e |
Let a backup carry a story that went two ways
A bundle had one list and a forked adventure has two stories, so export was emitting every branch's turns interleaved by index — a mangled story rather than lost data, and unreachable only because forking has no UI yet. `ai-dnd-adventure-v2` carries the branches, the depth each one left its parent at, which attempt at every turn is the story, and what each node left behind. That last one is not decoration: the after-snapshots are what a branch switch puts back, and a bundle without them imports a tree nobody can switch inside. `app/bundle.py` owns both formats and nothing else knows either. The v1 reader stays — those files are already on people's disks — and it is now the only place a `variants` array exists anywhere. The rule the module is built on is that a bundle carries what was chosen and never what is derived. The head branch, the fork points, the live flags and the anchors are decisions somebody made. The lineage, the head depth, the legacy `index` and the variant ordinals are computed from those and are rebuilt on the way in, because a bundle is a text file anybody can edit and a derived field shipped beside its source is a chance for the file to disagree with itself where no read would report it. `index` is the one that stops being academic here. It agreed with `depth` until SP5, and this is the first writer that has to fill it for a forked story, where two branches both hold a node at depth 4. It is allocated one per turn instead: siblings share it, no two coordinates do. Everything a hand-edited file can get wrong about the shape of a tree is a 400 raised before the adventure row exists, because a half-applied import is exactly the failure this phase exists to end — a story that goes quiet. A file wrong about which attempt is live is corrected rather than refused; that is an invariant of the database, not of the format. Measured on the 600-action fixture: 587 kB to 911 kB, and all of the increase is the outcomes at 489 B a node — the coordinates themselves save 57.5 B a node against the old turn-and-variants shape. Twenty forks add 660 B. 4.3% of the import body cap. 381 tests green, 16 of them new in test_bundle_v2.py. No migration, no vacuum owed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g |
||
|
|
0a12d9cd47 |
Make a retry a node, not a rewrite
Every attempt at a turn is now its own row at the same (branch, depth), with `live` naming the one the story tells. The JSON repeating group on `actions.variants` is read one last time, by a migration that writes it out as the sibling rows it always described, and then goes unread. The snapshots turn around with it: an action carries the state it left behind rather than the state it started from, because attempts at one turn share a starting position and differ exactly in their outcome. Rolling back is "what the node in front left behind", one lookup on the path, and it is what undo and retry now both read. And the memory holdback goes. It existed because retry rewrote a row under a mark that had already moved past it; a retry writes a sibling now, and replacing what a coordinate says withdraws what was derived from it — the same repair undo and delete already made. The assembled prompt is still stored once per turn: it moves with the live flag, so a superseded attempt keeps only the few hundred bytes that were its own. Measured on the 600-action fixture: 700 rows for the same 600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB and page load 62.7 kB unmoved. 347 tests green. `tests/test_story_tree_baseline.py` and `tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move the baseline for the variant-count semantics and did not need to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
c51531709d |
Mark the story with a node, not with a count
The memory bank and the story summary each kept a cursor: how many story actions they had already covered. A count is a position in a list, and this list moves — delete an action in front of the mark and every later one slides down a slot, so the mark now covers one it has never read. All the cursor bookkeeping existed to patch that up. Both marks are now (branch_id, depth): the node up to and including which the work is done. A depth is a coordinate along a path, not an offset into a list, so nothing in front of it can move it. That deletes rather than rewrites `position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`, `prune_dangling_memories` and the every-pass clamp in `run_post_turn`. A memory hangs off the node its block ends on, so a fork inherits its ancestors' memories without copying any, and retrieval selects through the branch clause over the *whole* lineage — recall is long-range by definition and cannot be windowed. Measured: 1,807 B on a story forked twenty times against 1,823 B on a flat one of the same length. Migrations 53-56 translate the old counts into nodes. They rewrite `adventures` and not `actions`, so this one needs no VACUUM FULL. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
d3756abdaa |
Give every action a branch and a depth
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a `branches` table, `branch_id`/`depth` on actions and memories, a head pointer on adventures, migrations 46-52, and a server-side backfill that re-reads every existing adventure as a tree with one branch. `depth` holds the number `index` already held, gaps included, so no story changes — a linear story *is* a tree with one branch, which is what makes the SP0 baseline passing unmodified the pass condition rather than a hope. The writer had to come with it. No migration will ever visit a row written after it ran, so columns backfilled today and populated next subphase would leave a hole exactly the width of one deploy, and from SP2 on a row without a branch is a row no read can see. `app/tree.py` owns that: one module, because a node written without a branch fails by disappearing rather than by raising. Three things the schema itself insisted on: - `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing both ways makes the two tables a cycle create_all cannot order, and its escape hatch needs an ALTER SQLite does not have. It is a cache, and a head naming a branch that is gone recovers onto the root. - `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`, because Postgres `json` has no equality operator. - SQLite will not drop a column a foreign key names, which is how two existing tests broke: they simulated an old database by rewinding the stamp while leaving the new columns in place. Every ADD COLUMN migration is now idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by rebuilding three tables from frozen DDL so the real ALTERs run. 297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn; page load and index are byte-identical to the recorded figures. The deploy that ships this needs one `VACUUM FULL actions;` on the direct endpoint afterwards — it rewrites every row. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
5c1bcf7305 |
Build a fixture that can witness what a tree changes
The stress fixture is sized from production and exists to weigh bytes, so it leaves every column it does not weigh at its default. Checked against a freshly built one, that is exactly the set phase 14 has to migrate: state_before and world_state_before NULL on all 600 rows, no scenario and so no RPG layer, no adventure scripts, both cursors 0, and a hundred retry histories whose two attempts carry byte-identical text with variant_index pinned at 0. That last one matters most. "Which attempt is live?" is the question SP4's migration answers when it decides which sibling node becomes the head of a turn, and against that fixture the question had no observable answer -- a migration that picked wrong would have looked correct. --rich fills in those columns and no others. A real RPG scenario read from the seed data rather than invented, with a world state played forward so hp declines and flags flip; monotonic per-action state snapshots, so a rollback that does not happen reads as a wrong number instead of as nothing; a gold script, story cards, non-zero cursors, pinned and forgotten memories; retry attempts with distinct texts, counts of two and three, and a live attempt that is often not the last written. Plus a second adventure, because a branch clause that forgot its adventure still looks right on a database holding one. The invariant SP4 reads -- text mirrors variants[variant_index] -- is asserted at build time rather than assumed. The plain fixture is untouched and re-measured unchanged at 1.8 kB, so the egress ceilings stay comparable. All six shapes run clean on --rich, including the turn path with the RPG layer and script pipeline now live. 283 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
0f1e05c808 |
Keep the fixture, so there is something to scroll (#5)
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.
--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:
- create_all() builds the current schema but leaves the version stamp at its
default, and bootstrap() reads a populated-but-unstamped database as
ancient — it replays every migration against a schema that already has the
columns, and fails on the first.
- the fixture's user is a registered one, but local mode looks for the row
with email IS NULL and is_guest false, so without clearing the email the
app opens on an empty library.
--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.
Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ae6e5af6c7 |
Store context_snapshot compressed
One column is 89% of the database and the free tier allows 512 MB. Reads were already solved -- the column is deferred, so a page load never touches it and one screen fetches one row at a time -- but nothing had costed storage, and storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per action, so the ceiling arrives around 5,400 actions and 944 are stored. Postgres already compresses it and only gets 1.7x. pglz is tuned for fast decompression of data a query might filter on, and nothing has ever filtered on an assembled prompt -- it is written once and read whole, rarely, by the Insights viewer. zlib gets 3.5x on the same text for a decompress on a request that already made an LLM call. Done as a TypeDecorator rather than a second column, so every call site still writes a dict and reads a dict back, and deferred/undefer/load_only keep naming the same attribute. Only the storage format moves. Migrations 43-45: add the bytea, convert into it, drop the original, rename. The backfill is the one destructive step in the file -- 44 removes the only other copy -- so it decompresses every row and compares it against what went in, and a row that fails aborts the run. The whole loop is one transaction, so an abort rolls the DROP back and the prompts are still there. Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column came out named context_snapshot, every snapshot compared equal and the one NULL stayed NULL. Postgres does not return the disk by itself: DROP COLUMN only marks the column gone and the backfill leaves a dead tuple per row, so the table peaks near twice its size before settling. The deploy needs one VACUUM FULL to collect it; the migration comment says so. The egress fixture's snapshots are prose now rather than "x" * 20_000, and the prose generator moved to tools/fakeprose.py so the harness and the tests share one definition. A repeated character compresses a thousandfold: against the old fixture a compressed column looked free and the byte ceilings would have been guarding nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
12d57afdac |
Put a number on what an endpoint may fetch, not just a column list
test_egress.py asserted which columns a statement names, which is the shape both of this project's egress blowouts took. It would all still pass if a response grew tenfold within the columns it is allowed to read -- and a story that keeps getting longer does exactly that. Production's longest adventure is 607 actions where the plan assumed 200. So dbmeter, which was built to be importable from tests and was not yet used by any, now backs four byte ceilings: the page load, the action list, and one action's snapshot fetched on demand. Budgets are per action rather than absolute, so they mean the same thing whatever size the fixture is set to, and generous -- 3 kB against a real 994 B. They are there to catch an order of magnitude, not to freeze a byte count. The fourth test is the one that keeps the other three honest. A ceiling proves nothing unless the thing it excludes would breach it, so it undefers the snapshot on purpose and asserts the same twelve rows cost more than ten times the budget. If the fixture ever shrinks below the point where that holds, that test fails rather than the ceilings quietly passing on nothing. Meter grows detach() and a context manager. A script exits and takes the wrapping with it; a test does not, and one test leaving the shared engine metered would charge bytes to a scope nobody opened. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
be66780a26 |
Size the stress fixture from what production actually holds
The old fixture was wrong in both directions at once and happened to land near the right total. Actions were modelled at ~2.1 KB against a real 886 B of text, and adventures at 200 actions against a real 607. Width flattered, length did not, and length is what a page load pays for. Re-sized from the 2026-08-17 measurements: 600 actions, 1700 B of narration alternating with a one-line player input, 232 KB of context_snapshot a row. The page-load shape now reports 606.0 kB against the 589.5 kB measured on production's longest adventure -- 2.8% out, where the old defaults were 28% out on a story a third of the length. Filler text is now generated word by word instead of one sentence repeated. That matters for what comes next: the repeated string compresses 313x and the generated prose 3.7x, so any compression ratio measured against the old fixture would have been fiction, and shrinking context_snapshot is the open question it exists to answer. context_snapshot also gains a flag of its own rather than being hardcoded, and the 74 KB figure in the comment -- inherited from models.py -- is corrected: the real column averages 163 KB a row across the table. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
70d024da62 |
Check the egress work against the database it actually runs on
Everything measured so far ran on SQLite against a synthetic fixture, so two claims were still on trust: that migration 38 spells BYTEA correctly for a real server, and that the byte figures survive psycopg's encodings. Both hold. Production reads schema_version 41 with embedding_blob bytea and embedded boolean present, the backfill is complete at 134/134, and the packed vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and against a throwaway Neon database every shape lands within 0.5% of the SQLite run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of it. The harness writes, so it refuses any target whose name does not say stress or scratch -- pointed at the production database it stops rather than seeding it with a fake user and 200 fake turns. It also empties a Postgres target before building, which a fresh SQLite temp file never needed. Two corrections fall out, both recorded in plan/13. The page-load model has the wrong shape: real actions are half the fixture's weight but real stories run to 607 actions, not 200, so the worst real page load is 589.5 kB. And the decision to leave context_snapshot in the database costed egress but never storage -- it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the ceiling this deploy will hit first. Measured with counts and octet_length sums only. No user content was read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
b7e53ae581 |
Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.
Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.
one turn 3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
run_post_turn 3,139.1 kB -> 0.7 kB
Insights 3,223.7 kB -> 117.9 kB
Memories drawer ~3.1 MB -> 23.6 kB
A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.
The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.
memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.
Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
|
||
|
|
c56864877a |
Store embeddings as packed float32 instead of a JSON list
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.
It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.
Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.
The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.
Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
|
||
|
|
7ee5ceea6c |
Measure database egress in bytes, from inside the repo
Both egress blowouts this project has had were one query fetching a column nobody read, and a statement count would have shown nothing wrong in either. So the meter counts bytes, at the DBAPI cursor -- everything that crosses that line crossed the wire. tools/stress_session.py drives a production-shaped adventure through the real routes with only the network faked. It reproduces both figures measured directly on production: 426.7 kB for a 200-action page load against 423 KB, and 3,258.7 kB for one turn against 3,153 kB. The memory bank is on by default, which is the whole point -- the previous harness ran without an embedding model, so retrieval returned early and the heaviest read in a turn never happened. --no-embeddings reproduces that deliberately, and the gap is 29x. It also turned up two callers the production SQL could not see: run_post_turn walks the whole bank again every turn, and Insights pays for it a third time. A played turn costs ~6.4 MB, not 3.2. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7 |