903fa7a74f886a5530efcfa5998144d599ca9bae
119
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
903fa7a74f |
M3: move the story's head instead of deleting its turns
Undo deleted. It removed the trailing AI action and the player action in front of it, pruned the memories covering them, and let the tip fall back to whatever survived. That made it the one operation in the application that destroyed accepted story, and it was why there was no Redo: the turns to move forward into no longer existed. Phase 0B demonstrated the head-cursor alternative in a disposable spike; this is that concept as production code. backend/app/head.py is the whole of it. Three questions that used to be one — where the story is being read, how far it is retained, and where it opens — are now three functions, and every caller that moves the head or asks about it goes through this module. The spike put the fork check in the write path and left Retry and Add-take on the old one; sharing the rules is what stops that divergence coming back. lineage.Path now caps every entry at the head, so hiding the retained future costs nothing at the call sites: the transcript, the context builder, attempts.preceding and memory retrieval already funnelled through path_of and narrow together. Path.uncapped() is the deliberate exception, and only Redo and the fork check may use it. The memory bank needs no pruning for the same reason — a memory carries the coordinate of the node its block ends on, so one derived past the head falls outside the capped clause and becomes retrievable again on Redo without having been deleted and re-embedded. Undo alone does not fork. Moving the head is not a decision to abandon anything, since the user may be reading or about to Redo; the first write below the head is where the story states which continuation it means. A head already at the tip forks nothing, so a story that is never undone forks exactly as often as it did before and the branch table does not fill up with one branch per turn. Redo follows the lineage rather than choosing among branches, which is what invalidates it after a divergence with no flag to set or clear. Migrations 78 and 79 give a branch superseded_at and superseded_depth. Nothing reads them to decide behaviour — Redo is decided by the lineage, so a stale or hand-edited value here cannot make the story wrong. They exist so the cleanup and discarded-history features left to a later version have something to select on, and so a divergence is observable in a test. Deleting an action no longer drags a moved-back head forward to the recomputed tip, which would have silently redone the story. can_undo and can_redo ride on AdventureOut and ActionPage because the client can work out neither for itself: the campaign opening may be off the top of the loaded window, and the retained future is never sent to it. This is a checkpoint, not the finished milestone. 601 backend tests pass. Five still assert the destructive contract — they count rows after an undo and expect the story to be shorter — and need rewriting against the new one; the world-state assertions inside them already pass. Export and import do not yet carry the head coordinate, so a bundle still reopens at the deepest node and can silently redo an undone story, which is the Phase 0B finding this milestone exists to close. The browser has no Redo control yet. None of the M3 acceptance coverage (D01-D10, E01-E04, I01-I03, I07, L01-L02) is written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u |
||
|
|
8652fe7cd8 |
M2 review: two regressions the green suite hid, and the reports
The post-implementation review of M2, plus the three corrections it took
to make the evidence true. Reports:
planning/reports/M2-BASELINE-REPORT.md 868 lines, the measurements
planning/reports/M2-IMPLEMENTATION-REPORT.md 758 lines, the reading of them
Verdict is PASS, accept with non-blocking debt, proceed to M3. Every M2
requirement is met and the ones that matter were tested by running the
build rather than reading it: a cloud endpoint written straight into
SQLite with sqlite3, behind the API's back, still refused at the wire;
trusted-LAN HTTPS against the real second machine with verification on;
captures showing zero packets outside loopback and the approved host.
Three defects, all found by running the shipped image.
The memory bank was dead. M2 removed Settings.api_key_plain with the API
key, and memorybank's two provider factories still read it. It failed
inside a fire-and-forget task, so no user error, no log anyone would
read, and no test — every memory test stubs those factories. All 604
tests passed with summaries and embeddings silently not happening.
The configurable model timeout never reached the turn engine. Stored,
validated, exposed in the API, rendered in the UI, and not passed to the
provider. M2's own exit criterion was half met: the constant had moved
but the setting did nothing.
And requirements.lock still pinned quickjs, psycopg and cryptography, so
the setup path DEVELOPMENT.md gives a new developer would have
reinstalled all three.
Both code defects now have the test that would have caught them: one
constructs every provider factory from a real Settings row, one drives
the turn endpoint, the chat endpoint and the summariser and asserts the
configured timeout arrives at each. That is the lesson worth keeping from
this milestone — after removing an attribute, build each consumer from a
real object; after adding a setting, prove it lands. Both failures were
in background or plumbing paths, which is exactly where a subtractive
change cannot see itself.
606 tests pass, up from 604. Lint, build and image are clean. Every
runtime result in the baseline report came from an image built after
these fixes; the reports say plainly that commit
|
||
|
|
8c65ae99de |
M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo. |
||
|
|
c1a73b3d77 |
M1: make the first story turn work with no Internet
Phase 0B ran the upstream application on a network with no route out and the first turn died in tiktoken, which downloads its BPE table the first time anything counts a token. The browser separately fetched three font families from Google on every page load. Neither is visible on a machine that has been online once, which is why both now have tests. The tokenizer table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it directly, verifying its SHA-256 against the digest tiktoken itself pins for that URL. No code path in the tokenizer can reach the network any more — not a warm cache, not an environment variable a deployment could forget. The encoding was checked token for token against tiktoken's own. The three font families are self-hosted as variable fonts under frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in frontend/src/styles/fonts.css, and re-vendored by frontend/tools/vendor_fonts.py. Their OFL licences ship beside them. With no remote asset left, the CSP drops both Google hosts and gains object-src, base-uri and form-action; woff2 also gets its real media type, which Python's table lacks on a slim image. A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx verifies against the certifi bundle, so an endpoint whose certificate comes from a CA the user installed on their own machines — a StartOS server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED while curl and the browser on the same host accepted it. app/tlstrust.py builds one context that unions the platform CA store with certifi's, and all four outbound clients use it. A union rather than a swap, so an image with an empty system store cannot start failing on endpoints that worked before. Verification itself is untouched: CERT_REQUIRED, hostname checking on, and no insecure escape hatch. The storyteller listener is now loopback by explicit statement rather than by inheriting uvicorn's default: start.sh, start.ps1, and docker-compose.yml, which publishes to 127.0.0.1 rather than every interface. Reaching an Ollama on another machine is outbound and needs none of that inbound exposure. backend/requirements.lock pins the exact tested closure; requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the same-host and trusted-LAN Ollama configurations, and how to re-run the offline proof. PROVENANCE.md records the upstream commit, the MIT terms, and both vendored assets. Verified, not just compiled. On an --internal Docker network with 1.1.1.1 unreachable and no name resolving, a campaign was created and played for six turns through same-host Ollama, restarted, and resumed. A second run played ten turns through Ollama on a separate physical machine on the LAN over verified HTTPS, summaries and embeddings included, with the storyteller's default route deleted so the LAN was reachable and the Internet was not. Its capture: 893 packets to the approved host, 730 loopback, zero anywhere else, and zero DNS queries. Two induced model failures left the accepted story bit-identical. The inherited SPA was opened in a browser and a campaign read back from it. Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the findings that did not belong in this change. 648 backend tests pass, up from the inherited 632; frontend lint and build are clean; the image builds. No M2 work is included: the hosted, cloud, analytics, Postgres and scripting surfaces are untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL |
||
|
|
d72f7c1bda |
Stop paying twice for a block a retry can still throw away
A memory whose block ends on the newest action is the one memory a player can reach: retry and take-switching both refuse anything else. Each retry of that turn withdrew the memory and wrote it again, and a block closes every six actions while a normal turn writes two, so that was one turn in three. SETTLE_SLACK asks for one action past a block before the block is summarized. The block is still MEMORY_INTERVAL actions; only the moment moves. This is not the pre-SP4 holdback returning: that one was about a retry rewriting text in place, which sibling attempts and forget_node settled, and correctness still rests on the withdrawal rather than on the slack. Undo and delete can carry a summarized node back to the tip, so the withdrawal path stays reachable, just rarely. The slack buys nothing back. The block that just closed is still in the history window in full, so a memory of it says what the model can already read. Tests: the settling suite asserts the new rule and that a retry at the tip finds nothing to withdraw; the rewrite suite builds thirteen actions so both of its blocks settle; the one path test that ended on a block boundary sets the slack to zero, because it is about which actions a block is read from rather than about when a block forms. 632 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NcrxCJjqgDvAkamKeLWdn |
||
|
|
745a4ea9f3 |
Stop printing the database password
The report opened with the connection string it was about to work on, taken straight from AIDND_DATABASE_URL. On the hosted deploy that string carries the Neon password, so the first line of every run put a live credential into the console — and from there into scrollback, a screenshot, or a pasted bug report. Nothing in the output said it was there to notice. It now prints scheme, user, host and database name. The query string goes whole: sslmode is the only part worth reading, and some drivers accept a password there as well. 631 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
eeab8ef5bb |
Survive the console this will actually be run from
The report prints an arrow between the old and new word counts, and it prints the memories themselves, which are model-written prose. A Windows console defaults to cp1252 and cannot encode either, so the run dies on UnicodeEncodeError partway through — after some memories have been rewritten and committed, which is the worst place to stop. STATUS already carries the same warning for tools/dbmeter.py. stdout is reconfigured to UTF-8 with errors="replace" instead, so the report degrades a character at a time rather than failing. The usage notes were also written for a POSIX shell, and this project is developed in PowerShell, where `VAR=x command` is not a thing. Both docs now give the PowerShell form first, via Read-Host so the Neon URL and the secret stay out of the history file, and name the venv's Python. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
4c94dd0fc0 |
Keep a production run to the adventures you meant
The hosted database is not the local one. It holds other people's stories, and `summary_provider` builds from the adventure owner's Settings, so an unfiltered `--write` against it would spend other people's money rewriting memories they never asked about. `--adventure` could already hold a run down, but only if you knew the ids. `--email` names accounts instead, and every line of the report now says who owns the adventure, so a dry run answers "whose keys would this spend" before anything is written. Guests have no email and stay reachable only by id, which is the right amount of friction for touching a stranger's bank. The other half is written down rather than built: stored API keys are encrypted with AIDND_SECRET_KEY, so a run from a checkout against the Neon database needs the same value the web service holds. With a different one `decrypt_secret` returns "" instead of failing, and every adventure is reported as having no key — a run that looks like it worked and did nothing. The Dockerfile copies backend/app alone, so tools/ is not on the box either way; the recipe is a checkout pointed at AIDND_DATABASE_URL. 629 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
0633cb624e |
Run the new memory prompt back over an old bank
The prompt change only reaches memories written after it. plan/18 decided to leave the existing ones alone and let eviction age them out at memory_bank_capacity, on the grounds that re-summarizing would duplicate whatever was still in the bank because nothing deletes the old rows. That was wrong about the only option. A memory can be rewritten in place. The row carries more than its text — whether it is pinned, how often it has been retrieved, and the node it hangs off, which is what makes a fork inherit the right memories — and rewriting `text` keeps all of it. Deleting the bank and rewinding the cursor would lose that, and would trickle memories back at MAX_MEMORIES_PER_RUN per turn, so an adventure nobody is playing would never recover. tools/rewrite_memories.py does it. Without --write it makes no model calls and only reports the scope; --write rewrites, --embed re-embeds in the run rather than leaving it to the app's post-turn pass. It reads whichever database the app reads, so it works against the hosted Postgres as well as a local file. Two things it needed from the app. `summarize_block` is now the one place a memory prompt is assembled, and the post-turn pass calls it too — a backfill that built its own prompt would be writing memories with a prompt that never shipped, and nothing would report the drift. `source_block` reads a memory's block back out of the story, which nothing has ever had to do: it reads on the lineage of the branch the memory was written on, not the branch being played, because after a fork the same depths hold different actions on each side and a read through the adventure's path would summarize the wrong story silently. It also excludes the discarded attempts at a retried turn, and tolerates a block an action has since been deleted from. Left alone: a memory with no source range, which is hand-written or migrated by 62 and may be the player's own words; a memory whose actions are gone; and an adventure whose owner has no API key, because summarization spends the user's own key by construction and never the shared demo key. --api-key/--model/ --endpoint override that, the last of them aiming a run at claude_shim.py. The vector is cleared for every rewrite, because the stored one describes wording that no longer exists. Re-embedding always uses the owner's own embedding model, never --endpoint: a vector only means anything against the vectors it is ranked beside. 17 tests, 627 green. The fork case is the one that would fail quietly, so the test builds a fork whose depths hold different actions on each side. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
712ef44f57 |
Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and would have been gone with the container. The numbers in plan/18 were therefore assertions nobody could check. plan/18-appendix-memory-ab-run.md now carries the whole transcript: both memories, both summaries, and the thirteen-action story they were written from. backend/tools/memory_ab.py reproduces it. It replaces the throwaway script the first run used, and differs in two ways that matter. It goes through OpenAICompatibleProvider rather than calling a model directly, so a run exercises the provider, the streaming path and complete() instead of a stub. And it reads the control prompt out of git at the commit given to --before, so the thing being compared against cannot drift from what actually shipped. There was already a claude_shim.py serving an OpenAI-compatible endpoint backed by the CLI, which is exactly what the throwaway script had reinvented. memory_ab.py points at it by default, so a run spends a Claude subscription rather than API credit, and --endpoint aims it at the provider the deployed app really uses — which is the one question this whole exercise could not answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
07192767d8 |
Give a memory a word ceiling, after measuring one
Ran the memory prompt end to end against a real model as a controlled A/B: one story generated through the app's own build_context a turn at a time, then both prompts run over the same blocks, so the story is held constant and the prompt is the only variable. Fresh process per call, so neither arm sees the other and the model is never told what is being tested. The control is the exact prompt from 9cdcb55. The reported fault reproduced. Two consecutive memories written from one story minutes apart came back in two different persons — "You crept low through the mist" and "The player asked Gwen to". With the cast brief both named Kaelen. The control also inverted who acted on a move whose player text was "grab her wrist and pull her down", which is the failure the brief predicts: with no cast there is nothing to say whose wrist "her wrist" was. It also found something the prompt review had not. "1-2 plain sentences" is not a length, and the same model wrote 34 words for one block and 105 for the next. A 105-word memory is a paragraph, and `memory_top_k` injects five every turn, so the bank's running cost was set by a number nobody had ever stated. MEMORY_MAX_WORDS states it, and the prompt now says which details to keep when trimming: the ones a later scene could turn on. Re-run over the identical story, the same two blocks came back at 32 and 58 words, still named, still third person, still carrying the camp map, the strongbox behind the second tent, and the strap frayed near through. Variance is the real gain — 34..105 became 32..58. Overshooting 50 slightly is expected. Models exceed word budgets, which is why builder.length_hint already carries a buffer for the same reason. What this does not show: the run used a Claude model, and the app talks to an OpenAI-compatible endpoint whose weaker models are why worldstate/parse.py tolerates trailing commas. The prompt is followable and the brief supplies the missing information; a weaker model is not proven to comply as well. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
1b5e4c56cd |
Tell the summarizer who the characters are
`_create_due_memories` sent six actions of second-person prose and nothing else — no protagonist, no cast, no setting, and no instruction about what person to write in. So for `You push the door open. She grabs your arm.` the only honest memory was "You entered a room and she stopped you", which names nobody when it is retrieved forty turns later. The framing wandered too: with no rule, the model picked a person per call, and one bank ended up holding "You entered the crypt", "The player entered the crypt" and "He entered the crypt" for the same kind of event. Both prompts now carry a cast brief and a framing rule: third person, the protagonist by name, other characters named rather than left as bare pronouns. The rule states its reason, because a memory really is read in isolation much later and a model told why complies far more consistently than one handed a bare instruction. The cast comes from the story cards, not from `stat_schema`. Every schema NPC is already turned into a card at adventure creation, deduplicated against the hand-written ones by name, so the cards cover schema NPCs, an author's own cards, and an adventure with no RPG layer at all through one path instead of three. Keyword matching alone was not enough, and finding that out changed the design. Built that way first, the brief for "She grabs your arm" listed the protagonist and nobody else: the block that most needs a cast is exactly the one written in bare pronouns, and Gwen's trigger keys include "her" but the text says "she". So matched cards come first and the remaining slots are filled with the other character cards. Places and items are not topped up — an unmentioned tavern is not who "she" was — though a place that is mentioned still matches normally. The asymmetry with the turn prompt is deliberate: an untriggered card is wrong as lore and right in a roster, because the roster answers "who could these pronouns be" rather than "what is on stage". Fixed descriptions only, never live values. `Gwen: trust 40 (wary)` in the brief would make the same event summarized at two different times come out framed differently, which is the fault this removes. An adventure with no persona still gets the cast and the setting, and the model is told to write "the player". One with nothing to say sends byte for byte the prompt it sent before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
9ee052c51e |
Give the adventure a persona, so the protagonist has a name
The player had stats but no identity. `stat_schema.player` carried hp and
mana beside `npc.gwen.trust`, but where an NPC has a name and a
description the player had neither, so the block rendered as
`You: hp 100/100` and nothing in the prompt said who "you" was.
Three columns on `adventures`: name, pronouns, description. All
user-only, all optional, and an empty name means the app behaves exactly
as it did before — no backfill, no special case for an adventure that
predates the migration.
They are adventure columns rather than part of `stat_schema` for two
reasons. An adventure with no RPG layer still has a protagonist, and
that is the case this was added for. And `worldstate.schema._initials`
treats every dict inside a stat section as a stat definition, so a
persona placed there would be instantiated, rendered in the guide, and
handed an `initial` value as though it were one.
The paths do not change. `player.hp` stays `player.hp`; only the label
moves, to `Kaelen (player): hp 100/100`, the same way NPC lines already
print a display name beside the id. A path carrying the persona's name
would break the moment a player renamed their character, because
`_history_text` replays every past turn's stored delta into the prompt
and those blobs hold literal `player.hp` strings.
The section sits in the system block. Only the user can edit it, so it
never changes mid-story and stays inside the cached prefix. That is what
makes it free, and it is why the AI must not be able to move it — a
delta aimed at `persona.*` is already refused by `_resolve`, and there
is now a test holding that in place.
The modal that used to appear only for scenarios with `${Placeholder}`
tokens now always opens, and is where the character is named. Persona
and placeholders stay independent: a scenario asking for `${Name}` is
asking its own question. No scenario in the repo uses placeholders at
all, so the overlap is hypothetical.
Phase 2, which feeds the persona and the cast to the summarizer, is
written up in plan/18 and not started. That is where the memory-quality
problem actually gets fixed; this change is what gives it a name to use.
Not yet driven in a browser — plan/18 lists what to check by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
|
||
|
|
6a5d8e98f9 |
Give the stress fixture back its sibling attempts
`place_action` used to read `depth` from `index`, so the fixture's retry attempts all landed on the turn's own depth and formed one sibling group. Now that depth comes from the head, each attempt moved the head one step and every retry became a turn of its own. The fixture's own check caught it: a coordinate holding a single superseded attempt has no live row, and that turn disappears from the story. The attempts at one turn share that turn's coordinate, so only the first one goes through `place_action` and the rest copy its placement. This is what `attempts.add_attempt` does. The fixture cannot call it directly, because it makes the newest attempt live and the fixture needs a live attempt that is often not the newest. Also drop the `variant_index` and `variant_count` arguments, which raise `TypeError` now, and read the two mark positions from a helper instead of from the adventure columns migrations 72 and 73 removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV |
||
|
|
949f58a054 |
Sweep the last three uses of the columns SP8 dropped
`seed_demo.py` still passed `index=0` to `models.Action`, which raises `TypeError` now that the attribute is gone. Every call site inside `app/` was updated when the column went, but the seed script sits outside the package and was missed. `bundle.settle` wrote `memory_cursor` and `summary_cursor`, which are no longer mapped, so the assignments only set transient Python attributes. The version 2 branch existed solely to make those assignments, and it ran two queries per cursor to do it, so it goes. `settle` now returns early when the bundle carries anchors. The `db` parameter is unused after that. Also move the comment about deleted memories next to the `db.delete(branch)` it describes. Splitting the router package left it after the `finally`, where it read as attached to nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV |
||
|
|
20f0753277 |
Merge main into the Phase 17 refactor
`main` changed the four files this branch split into packages, so all four came back as modify/delete conflicts. Each change is ported to the module that now holds the code, unchanged in behavior: - `delete_action` moves to `routers/adventures/actions.py`. It reaches the turn lock as `turns.acquire_turn_lock` and `turns._active_turns`, which is the rule this branch set: the package root no longer re-exports the names a test rebinds. - `EMIT_RULE` moves to `worldstate/parse.py`, and `_describe_stat` and `render_reference` to `worldstate/render.py`. - The chip-wrap rules move to `styles/schema-editor.css`, and `StateChangeChips` to `pages/Play/reports.jsx`. `index.css` keeps this branch's import list. `test_delete_state.py` needed four edits to run here. It drops the eight-line database prologue, because `conftest.py` does that once now. It imports the shared `ScriptedProvider` from `fakes.py` instead of carrying a copy. It patches `adventures.turns.OpenAICompatibleProvider` and reaches the lock the same way. And it no longer passes `index=0` when it builds the start action, because migration 71 drops that column. 564 tests pass. `npm run lint` and `npm run build` are clean, and the CSS bundle is 56.71 kB, which is the pre-merge size plus main's new rules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
91907a30bd |
Name every stat in the guide by the path the AI has to write
The referee refused `npc.<id>.status` on a character that has a status stat under another name, and the refusal was right: the model was reaching for a path the guide never gave it. The guide is the fixed list of everything a scenario tracks, and it named stats in prose. A player stat and a world stat both read as a bare name, and an NPC's stats read as the display name plus the stat — "Trainer Milo active_status". The model had to turn that back into `npc.milo.active_status` itself, and in the Pokemon demo five other characters carry a stat called `status`, so the path it built was `npc.milo.status`. The live values do state the paths, but only for the NPCs a scene has mentioned, so an NPC off screen was addressable only by guesswork. Each line now leads with the path: `player.potions`, `npc.milo.active_status`, `flags.sandstorm_active`. An NPC's header is written even when it has no description, because it is the one line that ties a display name to its id. `EMIT_RULE` points at the guide for paths rather than at the live values alone. A free-text stat is marked `(free text)` beside its path, which is the wording `EMIT_RULE` already used to describe it — the guide had been writing "free text" at the end of the line instead. Such a stat now also gets a line when it has no description, where before it was dropped and the model was left to send a number for a stat that holds a string. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST |
||
|
|
41ed2f63b5 |
Put the shared state back when a turn is deleted
Deleting an AI reply removed the text and left everything the turn did to the numbers standing. `script_state` and `world_state` belong to the adventure rather than to the node, so the row going away takes nothing back. Undo, retry, a take and a branch switch all restore them; this endpoint never did. The visible symptom was the cooldown clock, which lives in `world_state._meta.last_changed` and holds a depth. A deleted turn left its depth there, and the turn played in its place is played at that same depth, so the referee refused the change as one that had happened this very turn — on a turn the story no longer contains. Delete the reply because you did not like the stat change it proposed, press Continue, and the same change comes back marked "changed too recently". A script's state stacked the same way: the gold a deleted turn paid out stayed paid, and the replacement turn paid it again. The restore reads the tip's own outcome rather than the deleted node's neighbour, which is what `switch` does. Deleting the newest turn then rewinds to the node in front of it, and deleting one from the middle of the story restores the state the adventure is already in, so the text goes and the numbers stay. The endpoint now takes the turn lock too, for the reason undo takes it: it writes the shared state, and a turn that is still generating is about to write it as well. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST |
||
|
|
f1bebe18d0 |
Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`, `variant_count`, `state_before`, and `world_state_before`, plus `adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in SQLite, so migration 71 quotes it. Nothing outside the migrations read these. `models.py`, `schemas.py`, and `ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders by `id`, and `attempts.renumber`, `context.history.max_action_index`, and `nodes.next_index` are deleted. Two changes keep the migration replayable on a `create_all` database: - `_split_variants_into_siblings` wrote through the live ORM table, so it stopped compiling once migration 66 removed five of its columns. It now writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`. - Five data passes read columns these migrations drop. Each now calls `_has_columns` and returns early when the columns are absent. `bootstrap` takes a `through` version so a migration test can stop at the schema it asserts on. 555 tests pass, up from 549. Eight of the new cases assert each column is gone after a real schema-45 database migrates all the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK |
||
|
|
e0bf2b61d9 |
Add a useDebouncedSave hook and delete Settings.stream
Two items from Stage 2 of `plan/17-refactor.md`. `frontend/src/hooks/useDebouncedSave.js` replaces three copies of the same debounce. `PlotPanel` and `ScenarioEditor` held identical per-key timer maps. `ScriptEditor` held a single shared timer, so editing two fields inside the same 600 ms window canceled the first save. The hook gives every key its own timer, which fixes that. `Settings.stream` was dead state. Nothing read it and every turn streams. This removes the column, both schema fields, and adds migration 65 to drop it. It is item S1 in `docs/self-review.md`. Migration 65 needs a new guard. `_column_already_gone` is the counterpart to `_column_already_there`: `create_all` builds the current schema, which is already missing every dropped column, so a fixture that stamps an old version and replays would fail on a column that is not there. Verified: 549 backend tests pass, lint and build are clean, and migration 65 runs both ways, once against a database that still has the column and once against one that does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
2c57b1ceab |
Remove three kinds of duplication in the backend
Stage 2, items 1, 4, and 5 of `plan/17-refactor.md`. **One path resolver in `worldstate`.** `apply_delta` and `apply_override` routed `flags.<name>`, `milestones.<id>`, `world.<stat>`, `player.<stat>`, and `npc.<id>.<stat>` with parallel code, about 100 lines each. `_resolve` now says what a path points at and returns either a target or the rejection to report. Each function keeps its own write rule, because the rules genuinely differ: an override sets a number rather than adding to it, ignores `cooldown`, `max_delta_per_turn`, and the rule that a counter only counts up, and can un-set a milestone. A differential check ran both implementations over 3960 payloads: twenty paths, fourteen values, three starting states, plus every three-path combination. The results are identical except that 674 rejections from `apply_override` now carry a `fix` string. `apply_delta` already worded those, and the world-state editor renders them, so an override that names an unknown flag now explains itself the way a delta does. **`sse`, `SSE_HEADERS`, and `turn_error` move to `app/sse.py`.** Two routers stream, and `chat.py` had to import from `routers.adventures` to reach them. **`get_adventure_or_404` becomes the `current_adventure` dependency.** All 32 handlers repeated the call as their first statement. The ownership check now reads in the signature and runs before the body. FastAPI caches a dependency for one request, so the handler's `db` is the session the adventure came from. The generated OpenAPI document is byte-identical except on `rename_branch`, where `branch_id` is now listed before `adventure_id`, because that handler no longer names `adventure_id` itself. Parameter order in the document is cosmetic. Six tests in `test_state_revert.py` call `undo_turn` and `retry_action` directly rather than over HTTP. They pass the adventure they already hold instead of an id. 549 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
0623dd8b78 |
Split Play.jsx into a directory of twelve files
`frontend/src/pages/Play.jsx` was 2280 lines holding 27 components. It is now `frontend/src/pages/Play/`: `index.jsx` with the page component, five panels, two drawers, the reports, the take pager, the refresh dialog, and the shared formatting helpers. Every line moved verbatim. A coverage check confirms every non-blank line of the original appears exactly once, in order, across the twelve files, and a name-resolution check confirms every identifier each file references is defined or imported there, with no unused imports. The line split stranded a comment at six of the boundaries. A leading comment sits above the section it describes, so each cut left one at the end of the file before it. All six moved to the section they describe, rewritten in the house style. `usePlaySession.js` is not here. The page component still owns all of the session state. Moving eighteen `useState` calls and seven `useEffect` calls is a rewrite rather than a move, and no frontend test would catch a mistake in it today, so it waits for Stage 5. Four stand-in providers under `backend/tools/` now define `last_usage = None`. The turn engine reads that attribute after every call, and the fixtures never defined it, so `tools.tree_fixture` crashed. That break predates this branch. Verified by 549 passing tests, a clean `npm run lint` and `npm run build`, and by driving the Play screen: the story view, all five panels, both drawers including the world-state edit form, the branch map, the refresh dialog, and the take pager stepping onto a take that lives on another branch. No console errors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
8422ff24f6 |
Split the world-state engine into four modules
`backend/app/worldstate/engine.py` held 918 lines covering four separate jobs: reading a scenario's schema, parsing the block the model writes, applying a change within the schema's limits, and rendering state as prompt text. Each is now its own module, the largest 418 lines. An AST comparison against the old file confirms all 29 definitions are identical. No call site changes, because `worldstate/__init__.py` exports the same names it did before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
2fa812c056 |
Split the adventures router into a package
`backend/app/routers/adventures.py` held 2353 lines and 35 endpoints. It is now a package of 14 modules, the largest 443 lines. The split moves text rather than rewriting it. An AST comparison against the old file confirms all 86 definitions are identical, and the OpenAPI schema still lists the same 35 operations. Names a test replaces now live in `turns.py` only, and other modules reach them as `turns.<name>`. Rebinding a re-exported alias changes the alias and leaves every caller reading the original, so the package root does not re-export them. A patch aimed at the old target raises `AttributeError` instead of passing while doing nothing. Tests and the fixtures in `backend/tools/` say `adventures.turns.<name>`. The same rule keeps the turn lock working. One module owns `_active_turns`, so one lock guards one set. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
32cd7c1077 |
Give the tests one setup instead of thirty-five
Every test module carried the same eight-line prologue redirecting the database to a temp file. Only the first one to be imported ever took effect: `app.database` reads `AIDND_DB_PATH` at import and builds `engine` from it once, so by the time the second module ran the engine already existed. The other 34 copies created a temp file that nothing opened and nothing deleted, and leaked one per module per run. `conftest.py` now does it once, which is early enough because pytest imports conftest before any test module. It also deletes the file when the run ends. The tests still share one database, exactly as they already did: each `client` fixture calls `create_all` on setup and `drop_all` on teardown, so no test sees another test's rows. `tests/fakes.py` holds the one `ScriptedProvider`. Nine modules each had a copy, and the copies had drifted into four feature sets, so a test that needed to raise a provider error had to be written in one of the files whose copy supported that. The shared one is the superset. The two `FakeProvider` copies were the same class with a fixed reply, so they use it too. `test_chat.py` keeps its own, which implements `chat` rather than `generate` and records what it was constructed with. An autouse fixture resets the fake's class state between tests, so a stale reply list can no longer reach the next test. 435 lines out of the suite. 549 tests pass. Verified live by sabotage: breaking the shared fake fails 13 tests across four modules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
b1772c6e21 |
Plan the readability refactor, and clear the tree for it
Phase 17 splits the four files that hold most of the code, finishes the schema migration SP8 left half done, and stops the published guide from drifting away from its Markdown source. `plan/17-refactor.md` carries the plan and the progress table, and `plan/STATUS.md` points at it. Stage 0 is hygiene only. Both abandoned worktrees are gone, which freed about 104 MB. Removing `sp7-tree-ui` needed one extra step: a Vite dev server had been running out of it since 2026-08-18, holding `frontend/.vite` open and owning port 5173, and serving a tree 54 commits behind `main`. The three stale `.db` files are deleted; `data.db` is not. `AIDND_TRUSTED_PROXY_HOPS` is now documented. It was read at `limits.py:55` and named in no `.env.example`, README, or blueprint. It sets how many proxy hops the rate limiter trusts in `X-Forwarded-For`, so a deployment that adds a hop without setting it gets the bucket-rotation bypass back. The 19 squash-landed branches are still there. `git branch -D` is blocked by the permission classifier; the verified command is in the plan file. 549 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
9398c13da5 |
Delete seeded scenarios no seed file claims any more
`previous_titles` stops a rename stranding the row it left behind, but the rows already stranded still had to be deleted by hand on every deployment. The seeder now removes them on the next boot, which retires the stale "Road to the Champion" demo without a database console. Only rows with a NULL owner and `is_public` are considered, and a player's own scenario is neither, so nothing anybody created is reachable. An adventure started from a deleted demo survives: `adventures.scenario_id` is `ON DELETE SET NULL`, and the adventure holds its own copies of the cards and scripts, so it loses only the inherited cover art. Two cases skip the sweep, because neither is an instruction to remove live content: a seed file that fails to parse claims no title, and an empty seed directory reads as a packaging failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
c2d3f0d8b9 |
Show the starter adventure the artwork of the demo it came from
An adventure has no cover art of its own and inherits its scenario's, and a bundle carries no scenario id, because an id means nothing in another database. The starter card therefore fell back to a monogram tile while the demo beside it showed the Pokeball. The starter file names its source under `scenarioTitle`, and the copy is linked to the seeded scenario with that title. If no seed answers to the name, the adventure keeps a NULL `scenario_id`, which is the state every imported bundle is in and costs only the artwork. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
ef4da7fda2 |
Keep the local Claude test rig, and record what it found
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by the local `claude` command line tool, so a demo can be played against a real model with no API key. Each request spawns one `claude --print` process, which suits the turn engine: the app assembles the whole prompt every turn and expects a stateless endpoint. Playing the Pokemon demo through it made all five of the world-state changes in `plan/16` work, and the refusal loop ran end to end for the first time: turn 3 clamped to nothing, turn 4's assembled prompt carried the correction verbatim, and the model's next delta was right. Two failures previously blamed on the code were the demo model. One is a schema fault and is still open: Milo's three Pokemon share one `active_hp` stat, so a switch leaves the newcomer at 0 HP. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
6cbf6d996f |
Give every new guest an adventure that is already played
An empty account gives a visitor nothing to read, and the daily demo turns are limited, so learning what the app does cost one of them. `app/starter.py` now copies a shipped export bundle into each new guest at the point the row is created. The bundle is two exchanges of the Pokemon demo, which ends on a knockout and shows an applied change, a refused one, and a milestone. The guest row is committed before the copy is attempted, so a failure there still leaves them with an account, and the copy runs inside a savepoint. The row building that `POST /adventures/import` did inline moved into `bundle.materialize`, which both callers use. The rate and size checks stayed in the endpoint: the starter writes a file the server ships, so it has no untrusted list to cap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
ae39514d1e |
Name the Pokemon demo, and land a seed rename on its own row
`seed.py` matches a seed file to its scenario by title, so renaming one inserted a second public scenario and stranded the first. The stranded row stays public forever and has to be deleted by hand on every deployment, which is what happened to "Road to the Champion". A seed file now lists its old titles under `previous_titles`, and the rename updates the existing row. The cover art is a PNG data URI. `app/images.py` accepts raster formats only, because SVG can carry script and the bytes are served from the app's own origin, so an SVG stores but yields an empty `image_url`. `tools/make_pokeball.py` draws the ball with `zlib` alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
ab546d6f3c |
Ask the Pokemon demo for changes, not for totals
The model sent a total rather than a change for almost every number: `npc.ivysaur.hp: 96`, `npc.milo.active_hp: 88`, `player.potions: 2`. Because every `hp` starts at its maximum, each one clamped back to where it started, and the potion count rose when the player spent one. `EMIT_RULE` does say "deltas (not new totals)", but the scenario contradicted it at closer range. `milo.active_hp`'s description said "Reset this to the newcomer's full HP", which asks for an absolute and is injected every turn. The bullets said "drop the HP, and raise it when healed", naming a direction but no sign. Five HP descriptions said only whose HP it was. The one line that said "not a delta" covered `player.active_pokemon`, so naming the exception made the rule look optional. `pokemon_fainted` is the control: its description says "add 1 each time", and it is the only number that behaved. Every stat description now states the sign, and the lead-in gives a worked example. `milo.active_hp.max_delta_per_turn` goes 65 to 98, because a switch moves that stat a full bar from 0 and the old cap made the reset unreachable in one turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
1988979aeb |
Store the refusals, so the chips and the model can see them
`world_delta_of` wrote `delta` and `applied` only. Every consumer that tells a refused change from a successful one reads the two lists it dropped: `Action.world_changes` marks a clamped chip from `clamped` and builds its refusal chips from `rejected`, and `worldstate.refusals` reads both. So no chip could report a limit, no rejection chip could appear, and no correction ever reached the next prompt. The three mechanisms merged last session were live in the code and unreachable in production. Found by playing the Pokemon demo. Turn 3's snapshot held two clamped entries with correct `fix` text, both chips came back `clamped: false`, and turn 4's prompt carried no correction, so the model repeated the same mistake. The 21 tests passed because `action()` built the column by hand with every list present. It now fills the column through `world_delta_of`. Removing the two new lines fails 9 of the 23 tests; that was checked by sabotage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
47c7800903 |
Report the world-state changes the engine refuses
`apply_delta` records three outcomes for every change the model sends: `applied`, `clamped`, and `rejected`. Everything downstream read only `applied`. A refused change reached the player as an ordinary chip, and reached the model on the next turn as a change that had succeeded. Five parts: - `Action.world_changes` reads `clamped` and `rejected` beside `applied`. Accepted stats carry a `clamped` flag; refusals become `kind: "rejected"` entries. The `fix` key is present only when the engine wrote one, because this property runs for every action of every list response. - The UI separates the three outcomes. A clamp to a standstill reads `no change - at its limit` on a dashed chip, a partial clamp is marked `(limited)`, and a rejection carries its reason. Dashed and dimmed rather than red: a refused change means the rules are working. - The goals line names the milestone id, as `milestones.<id>`. The ids appeared nowhere in the prompt before, so the model could not send one. - Each rejection, and each clamp that moved nothing, builds a `fix` string from the stat definition at the point of refusal. `render_refusals()` renders them into the next prompt above `EMIT_REMINDER`. - `_history_text` replays `applied_delta()` instead of the sent delta, so a past turn's state block shows only what the engine accepted. A clamp that reduced a change but still moved the value reports nothing. If you tell a model its 80 damage became 30, it can treat the shortfall as a debt and send the remaining 50 next turn, which is the swing `max_delta_per_turn` prevents. In the demo scenario, `pokemon_left` becomes `pokemon_fainted` (`type: counter`, `initial: 0`). Starting at the ceiling turned a wrong-signed delta into a silent no-op; counting up puts the wrong sign on the counter rule, which refuses it out loud. The instructions also now ask for `world.turn`, which sat at 0 for a whole playtest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
2ca2dadc72 |
Fix stale description wording after the flags-to-text-field change
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
2805eb1fb1 |
Nest the player's Pokemon under npcs instead of flattening them
Each teammate is its own named entity with hp/status, the same shape Milo already uses, rather than <name>_hp/<name>_status flattened under player. Also replaces the five mutually-exclusive _active flags with a single player.active_pokemon text field, mirroring npc.milo.active_pokemon, so the AI no longer has to self-enforce "exactly one flag true." Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
cd1ae219ad |
Rework League Championship demo into a real Pokemon battle
Each of 5 team Pokemon now tracks its own HP and status condition, one flag marks who's active, and potions are a depletable resource. The opponent NPC tracks its active Pokemon the same way. Drops the crowd-favor stat, which didn't fit a battle scenario. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
246a355a49 |
Add League Championship demo scenario
Showcases world-state stats and bands, a counter, a text stat, flags, two NPCs with different stat shapes, and milestones in one scene. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
a408c7b6f7 |
Lay the prompt out so the endpoint can cache most of it
Prompt caching bills on a shared prefix: the endpoint reuses the request up to the first byte that differs from last time and no further. The live world-state block sat third from the top of the system message, so every turn re-priced the instructions, the plot essentials and the whole story history underneath it. The retrieved memories and the rewritten summary did it again. Everything fixed is emitted first now, and everything that moves goes after the history, ordered least-volatile first — which is also where recency serves it best, the reasoning that already put the emit reminder last. The three tail sections that are last for their own reasons stay last. The moved sections are still charged to the token budget; only their position changed. Two smaller halves of the same problem. OpenRouter serves a model from whichever upstream is free and each upstream holds its own cache, so a deepseek model now names deepseek as its preferred upstream — a preference, not a restriction, so a turn still runs if that upstream is down. And the endpoint's usage block is read back off the response and kept per attempt, so the hit rate shows up in Insights and the debug log instead of being assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
041f9e25f3 |
Count the visits, and say whether anyone got anywhere
A hosted demo raises a question a local app never does: is anyone using it, and do they reach the part that matters? `/analytics` answers it — visitors, pages, referrers, countries, devices, which shared scenarios get played, turns and demo-key spend, API and turn errors, and a funnel from visited to played a turn to signed up. Not a third-party script, for reasons specific to this one. The CSP allows `script-src 'self'`, so a tracker means loosening it; adblockers eat the popular ones, which silently biases exactly the technical audience this project gets shown to; and none of them can see the measurement that actually matters here, which is a turn, not a pageview. **A visit is a write and never a read.** After the 189x egress fix it would be perverse to add a feature that reads rows per request, so counts accumulate in a process-local dict and flush every 60s as UPSERTs. Storage is a generic `(day, metric, label) -> hits` counter, so measuring something new later costs a constant rather than a migration, plus one row per visitor per day for the funnel flags. Every dashboard query is a GROUP BY returning tens of rows however much traffic sits behind it; a month reads back in a few kilobytes. The buffer's cost is that a hard restart can lose up to a minute — the flusher also runs on shutdown, and a tier that sleeps when idle sleeps on an empty buffer anyway. **The counters are anonymous; the access log beside them is not, on purpose.** A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars — one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined back to `users`, and keyed, so no client can compute one. Story content never reaches that module, and the only content it ever names is a seeded public scenario's title; a player's own titles are theirs. `accesslog.py` is the identifying half and is a separate module writing a separate table so that separation is a property of the code rather than a convention: `access_events` records sessions, sign-ins, registrations and failed attempts with address, email and device, read on a second tab of the same page behind the same gate. Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An unmetered tester is not automatically someone who should see the traffic. The route 404s and the nav link is absent for everyone else, the same treatment AI Chat gets; unset in a hosted deploy means nobody sees it, including me. Three things came out of building it that a test would not have suggested. **A failed turn is an HTTP 200 with a bad ending.** The status-code middleware cannot see one, so a demo whose model had started refusing every request would look perfectly healthy from outside. All five SSE error paths in `_generate_turn` now go through a `turn_error()` helper that counts on the way out. Error buckets elsewhere are labelled by the matched route template rather than the requested path — one bucket per endpoint instead of one per adventure id, and, the reason it isn't merely tidier, an unmatched path is entirely attacker-chosen, so labelling by it would let anyone mint rows. **The funnel counts people, not clicks.** A player who starts six adventures is one person who started an adventure. That is the whole reason the per-visitor-day table exists; its flags only ever turn on, and `is_new` is settled by the first write of a visitor's first day. **The tests run on SQLite and production is Neon.** A flush that raises is caught and logged, so a dialect mistake in the UPSERTs would have stayed invisible until the dashboard quietly never filled. `test_the_upserts_compile_for_postgres` compiles both statements against the Postgres dialect without connecting to one. Two things this leans on elsewhere. `limits._client_ip` is now public `client_ip`: the access log needs the same answer, and two functions both deciding which hop is the caller's is how one of them ends up trusting a header it shouldn't. And the cleanup sweeper now starts if *either* job has work — a deployment can keep every guest forever and still want its visitor-day rows aged out. No migration. Both tables are new and `bootstrap()` calls `create_all` on existing databases too, the route `branches` took in Phase 14, so `LATEST_VERSION` is still 64. 497 tests green, frontend lint and build clean, driven by hand against a synthetic 90-day fixture at 1568px. The narrow-screen layout follows the existing 720px block but is unverified: `resize_window` is ignored on a maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run through the test names themselves, leaving `testclient_ip_*` — still collected by pytest, which is why it passed unnoticed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
3b9e6b3d50 |
Give the length hint a floor, not just a wall
A ceiling alone is a one-sided instruction, and models read it in opposite directions. A verbose one is held back by it; a terse one has nothing to act on except "write only as much as the moment needs -- a typical turn is much shorter" and collapses to two paragraphs. Same prompt, wildly different turn lengths depending on which model is behind it. State a floor as well, so the guidance is a band. The two bounds are deliberately asymmetric -- "must not exceed" for the wall the endpoint enforces, "should not stop short of" for the floor -- so neither reads as a number to hit, which is the property the earlier A/B says decides whether this hint helps or hurts. "Prefer the lower end" inherits the anti-overshoot job the deleted "much shorter" line was doing, but now with a number under it, so a terse model lands on the floor instead of at forty words. Below MIN_LENGTH_FLOOR_WORDS the floor is dropped and the tight-cap wording is left byte-identical: at a tight cap a short turn is the correct turn, and that phrasing is the one measured to keep the state block alive (0/6 truncations at cap 250 against 2/6 unhinted). So this only moves loose caps. MAX_LENGTH_FLOOR_WORDS keeps the share from demanding 555 words minimum at cap 2400 -- a big cap means long turns are allowed, not compulsory. Shipped without an A/B run, deliberately. Two things to watch live: whether a stated range invites landing mid-range on verbose models (drop the share to ~0.25 if so), and whether the state block still survives -- nothing reads finish_reason yet, so truncation is silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
40d2555f84 |
Say what the app is now, everywhere it is published
The README, the project page and the engineering guide all describe a linear story. The tree shipped two days ago. Every published surface is a phase behind, and the guide is not merely behind — it is wrong in a way that costs a reader time. Its 2.2 was "Two coordinate systems, and the bug class they create", and it explained the codebase through position_of_index, note_action_removed and settled_story_actions. All three were deleted in SP3. 2.3 explained retry through Action.variants and state_before. Somebody reading either would go looking for machinery that is not there, which is worse than a gap. So 2.2 is now "The story is a tree", written at the depth 1.2 and 1.3 are written at: the seven bugs that turned out to be one bug, the lineage clause and the two properties that make fork count free, why takes group by parent_id rather than by coordinate, cursors becoming anchors, and a closing list of what the design is honest about. 2.3 is rewritten around state_after and takes, and 1.1 and 1.5 follow, because the pipeline no longer snapshots before the call and the memory bank no longer holds an action back. The numbers were simply old: 151 tests where there are 440, 37 migrations where there are 64, twelve phases where there are fourteen. They appear in four places across the README, the project page's stat tiles and the guide's results table. The measured branch cost — 103 B, and 1.007x the page load of the same story flat — is added beside the egress and turn-cost figures it belongs with, since it is the number that answers "what does branching cost me". Three screenshots, on a new tools/shots_fixture.py: the Bandit Camp demo driven through eight written turns with written deltas, three discarded takes forked onto branches of their own, one off a branch so the map has to nest. Same reason tree_fixture.py is committed — the shots have to be reproducible and the frontend still has no test runner. play-world-state.jpg is reshot because it predates the entire tree UI; the map and the branches panel are new. Note for next time: docs/guide.html is hand-written, not generated from the Markdown, so every guide edit is two edits in two vocabularies. Both files were checked for tag balance and both pages rendered locally before this landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
c8e081e9d6 |
Draw the tree the branch list can only spell out
The Branches panel says which lines exist. It cannot say where they parted or how much story each one is, because those are the two numbers `GET /branches` already answers and a list has nowhere to put them. A ⌗ See the tree button opens a map: one horizontal lane per branch, running from the moment it left its parent to the moment it ends, joined to the parent by an elbow at the fork. The horizontal axis is the story's own clock, so two lanes at the same x are at the same moment and a short branch reads as short. This is a branch map, not the per-node map SP7 refused, and that is the whole reason it was cheap. Lanes are bounded by branch count, not node count, so the 600-node windowing problem never arrives. It reads the single request the rail already made and nothing else. `branches.js` holds the tree maths and `BranchMap.jsx` the drawing. Play.jsx gives up its private copies of branchLabel and orderBranches: the panel and the map now label and order a branch through the same functions, so a branch cannot be called two things by the two views. The three operations stay in BranchPanel and are passed down, and `run` answers whether it worked so neither view clears a half-typed name on a refusal. Three things came out of driving it, none of which a test could have seen. `clientWidth` counts the canvas padding the ResizeObserver leaves out, so the first paint drew an svg 24px wider than its box — and the observer's initial observation never arrived here, so dropping the seed left the map never drawing at all. Both are needed and the seed subtracts the padding. Only the name was being clipped, not the meta line under it, so a late fork ran its text off the right edge; both are clipped now, and a lane starting in the right third hangs its labels back over the fork, where its own band guarantees nothing to collide with. And the delete rule lived in the server and in the map but not in the list, which offered Delete on a branch the head was forked from and answered with a toast from the server's 400. `headLineage` is the client's copy of that rule and both views use it; the server stays the authority. `tools/tree_fixture.py` is the counterpart to `tools/branch_fixture.py` — four branches at three fork depths, one forked off a fork. With no frontend test runner it is the whole of the map's coverage, and it exists to be looked at. Nothing in the backend changed; the 440 tests pass unmoved. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
ff190d2453 |
Edit the take you are reading, not the one it replaced
The pager parks a turn on take 2 of 4 and the row draws that take's words, but the row itself is keyed by the live node — the take the story tells. The edit button seeded from that node and saved back to it, so opening the editor on a 2/4 turn showed 4/4's text and saving overwrote 4/4. The take being read was never reachable. It has an id of its own, carried on the preview, and the edit endpoint takes any row by id whether or not the path runs through it. So an edit opened over a preview carries the take's id, the transcript matches the editor to the row through the preview rather than the node, and the saved text goes back into the preview because there is no row in `actions` to put it in. The pager caches the take list it fetched, so it is told to drop it — otherwise stepping away and back showed the words from before the edit. Leaving the take at all drops a half-typed edit with it. The fork button had the same seed and now takes its text from the screen too. Its id stays the live node's: a take branches just above the turn, and the server only accepts a turn that is on the path. The tests are on the promise the fix leans on — a take is an ordinary row to the edit endpoint, addressed by its own id, and the group listing says so afterwards. That already held; nothing in the backend changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c0cd6fa7ce |
Stop a full memory bank from shutting itself
Eviction ranked on use_count first. A memory written this turn has never been used, so once every survivor in a full bank had been retrieved even once, the newborn was the lowest row in the bank and was marked forgotten by the same post-turn run that wrote it — one pass after the one that embedded it, before retrieval ever saw it. That state never ends, because counts only go up. The bank an adventure happened to hold when it first filled is the bank it keeps for good, and everything the story does afterwards is summarized, evicted and never ranked. A test now plays four turns against a full bank and asserts the survivors are the new memories, not the opening ones; before this it kept the opening three and forgot all four. Order on last-touch instead, with the count as the tiebreak. A new memory carries the newest timestamp there is, so it is the safest row in the bank rather than the most doomed, and it has until something outlives it to prove itself — the "protect the newest" behaviour falls out of the ordering rather than being a count of rows to spare. Little is given up: a memory that is genuinely used stays recently-used by being retrieved, so the two orders only disagree about memories that mattered once and have not been wanted since. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
c4936f75f5 |
Fix two things only a person clicking could find
Driving SP9 by hand found both, and neither was reachable from a test that did not exist yet. **The pager arrived one page load late.** A retry's reply *is* the second take of its turn, so it comes back needing a pager -- but the SSE stream builds its own ActionOut and so never got the annotation. The numbers only appeared once the page was reloaded, which is the one moment nobody reloads. That is the third place this codebase builds an action payload a different way; the other two were patched when `take_count` was added, and this one was missed because nothing read it from the wire. **"> You > You follow the pulse down the corridor."** The editor is seeded from the stored text, which is already the formatted form -- the same text plain edit puts in the box and writes back verbatim. Running it through the formatter again doubles the prefix, so `run_player_turn` learns `preformatted`. And the question the driving raised: does the shared state follow the branch? `test_take_state.py` answers it with a script that adds ten gold a turn, which is the clearest possible witness -- a take that stacks instead of replacing says so in one digit. Three of its five fail with `roll_back_before` removed; the other two pass either way on purpose, one being the premise and one going through `stand_on`, which has restored state since SP5. Confirmed against the running app too, on the HP-script demo: a path crossing three branches carried exactly the damage of the nine AI nodes on it, and none of the 66 points sitting on takes those branches never told. 433 tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
f9b836fc0b |
Put the pager back, and let any turn be played again
SP7 replaced the pager with chips, on the grounds that a chip could also offer "take this path" while a pager could only step. Driving it by hand said otherwise, and the reason is worth keeping: the chip meant two different things depending on where the reader was standing -- a real switch at the tip, a preview needing a second button above it further back. Two meanings in one control is what made the tree unusable. So: one control that does one thing. Stepping reads a take and nothing else, and it tells the server nothing, because reading is not a decision. The transcript below a take that is not live simply ends -- such a take is a leaf by construction, since whatever was played after the turn was played after the take that *is* live. The decision is made by writing, and `after_id` carries it. One step does reach the server and is still not a fork: a take with a story of its own lives on its own branch, so going there is a branch switch and only the server can say what is underneath. `branch_id` on the take is what tells the two apart without asking first. And a fork button on every turn but the opening. On the AI's it regenerates; on your own it opens the text so you can say something else. What the story made of the old take is kept, on the line it was written on. `selectVariant` and `forkFromAttempt` leave the client. Both endpoints stay -- tested, and `stand_on` is shared with the write path -- but the pager needs neither. 426 backend tests; lint and build clean. Not yet driven by hand: the frontend still has no test runner, so this needs the `--keep` fixture and eyes, exactly as SP7 did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
ea7336e5d3 |
Let any turn be played again, not just the newest
`retry` only ever saw the last action: mid-story there was no way to ask for another take at all. Now the take endpoint takes either kind of node, and the kind decides what happens -- an AI turn regenerates, a player turn takes the text supplied. The client asks the same way for both. The tip is the only case that needs no branch, and only for an AI turn, where the takes are still leaves nobody has built on. That is the existing retry, reached by another road. Everywhere else `branch_at` leaves the path just before the turn so the story after it keeps the take it was written for. Both paths roll the shared state back to before the turn ran, which the fork endpoint was already doing and the player-take path was not -- a script would otherwise have stacked this take's output mutations on the one being replaced. 426 tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
c3c9b310eb |
Put the pager's numbers on the page
`2/4` needs the shape of a turn's take group for every message on screen. Asking `attempts.group` per row would put a query behind each one -- the exact cost `variant_count` was cached to avoid, and the reason SP8 could not just drop that column and be done. So one query per page, keyed on the parents the page mentions, and the count and ordinal land on the rows before they are serialised. `take_count` / `take_index` rather than reusing the old pair, because they do not mean the same thing: the old ones cache a coordinate's siblings and say 0 for a turn nobody retook, these count the *turn's* takes across whatever branches they ended up on and say 1/1. SP8 still drops `variant_count` and `variant_index`. Two traps, one of them mine. `parent_id` was not in ACTION_LIST_COLUMNS, so reading it off a windowed row would have been a lazy load per action -- an N+1 hidden behind the thing `load_only` exists to prevent. And the adventure GET does not build `ActionOut` at all: it hands the window to the relationship with `set_committed_value` and lets Pydantic walk it, so patching the three places that do build ActionOut missed the one path every page load takes. The test caught it by reading the wire instead of the helper. 424 tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |