The report prints an arrow between the old and new word counts, and it prints
the memories themselves, which are model-written prose. A Windows console
defaults to cp1252 and cannot encode either, so the run dies on
UnicodeEncodeError partway through — after some memories have been rewritten
and committed, which is the worst place to stop. STATUS already carries the
same warning for tools/dbmeter.py.
stdout is reconfigured to UTF-8 with errors="replace" instead, so the report
degrades a character at a time rather than failing.
The usage notes were also written for a POSIX shell, and this project is
developed in PowerShell, where `VAR=x command` is not a thing. Both docs now
give the PowerShell form first, via Read-Host so the Neon URL and the secret
stay out of the history file, and name the venv's Python.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW
The hosted database is not the local one. It holds other people's stories, and
`summary_provider` builds from the adventure owner's Settings, so an unfiltered
`--write` against it would spend other people's money rewriting memories they
never asked about. `--adventure` could already hold a run down, but only if you
knew the ids.
`--email` names accounts instead, and every line of the report now says who owns
the adventure, so a dry run answers "whose keys would this spend" before
anything is written. Guests have no email and stay reachable only by id, which
is the right amount of friction for touching a stranger's bank.
The other half is written down rather than built: stored API keys are encrypted
with AIDND_SECRET_KEY, so a run from a checkout against the Neon database needs
the same value the web service holds. With a different one `decrypt_secret`
returns "" instead of failing, and every adventure is reported as having no key
— a run that looks like it worked and did nothing. The Dockerfile copies
backend/app alone, so tools/ is not on the box either way; the recipe is a
checkout pointed at AIDND_DATABASE_URL.
629 green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW
The prompt change only reaches memories written after it. plan/18 decided to
leave the existing ones alone and let eviction age them out at
memory_bank_capacity, on the grounds that re-summarizing would duplicate
whatever was still in the bank because nothing deletes the old rows.
That was wrong about the only option. A memory can be rewritten in place. The
row carries more than its text — whether it is pinned, how often it has been
retrieved, and the node it hangs off, which is what makes a fork inherit the
right memories — and rewriting `text` keeps all of it. Deleting the bank and
rewinding the cursor would lose that, and would trickle memories back at
MAX_MEMORIES_PER_RUN per turn, so an adventure nobody is playing would never
recover.
tools/rewrite_memories.py does it. Without --write it makes no model calls and
only reports the scope; --write rewrites, --embed re-embeds in the run rather
than leaving it to the app's post-turn pass. It reads whichever database the
app reads, so it works against the hosted Postgres as well as a local file.
Two things it needed from the app. `summarize_block` is now the one place a
memory prompt is assembled, and the post-turn pass calls it too — a backfill
that built its own prompt would be writing memories with a prompt that never
shipped, and nothing would report the drift. `source_block` reads a memory's
block back out of the story, which nothing has ever had to do: it reads on the
lineage of the branch the memory was written on, not the branch being played,
because after a fork the same depths hold different actions on each side and a
read through the adventure's path would summarize the wrong story silently.
It also excludes the discarded attempts at a retried turn, and tolerates a
block an action has since been deleted from.
Left alone: a memory with no source range, which is hand-written or migrated by
62 and may be the player's own words; a memory whose actions are gone; and an
adventure whose owner has no API key, because summarization spends the user's
own key by construction and never the shared demo key. --api-key/--model/
--endpoint override that, the last of them aiming a run at claude_shim.py.
The vector is cleared for every rewrite, because the stored one describes
wording that no longer exists. Re-embedding always uses the owner's own
embedding model, never --endpoint: a vector only means anything against the
vectors it is ranked beside.
17 tests, 627 green. The fork case is the one that would fail quietly, so the
test builds a fork whose depths hold different actions on each side.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW
The checked-in harness was run end to end through the shim on a newly
generated story, which is both a check that it works and an independent
replication.
The finding the change is for held: across two runs, none of the four
control memories names the protagonist and all four treatment memories
do. The length finding held too — controls at 34, 89, 105 and 107 words,
a four-fold spread with no budget stated anywhere, against 54 and 55
under MEMORY_MAX_WORDS.
One claim did not replicate, and the writeups now say so. Run 1 produced
two control memories in two different persons, which is the reported
complaint exactly, and plan/18 presented that as reproduced. Run 2's
controls were both "the player", consistently. Drifting between second
and third person is therefore something a model sometimes does, observed
once, not something it does every time. The naming gap is the durable
result and the writeups now lead with that instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and
would have been gone with the container. The numbers in plan/18 were
therefore assertions nobody could check.
plan/18-appendix-memory-ab-run.md now carries the whole transcript: both
memories, both summaries, and the thirteen-action story they were
written from.
backend/tools/memory_ab.py reproduces it. It replaces the throwaway
script the first run used, and differs in two ways that matter. It goes
through OpenAICompatibleProvider rather than calling a model directly,
so a run exercises the provider, the streaming path and complete()
instead of a stub. And it reads the control prompt out of git at the
commit given to --before, so the thing being compared against cannot
drift from what actually shipped.
There was already a claude_shim.py serving an OpenAI-compatible endpoint
backed by the CLI, which is exactly what the throwaway script had
reinvented. memory_ab.py points at it by default, so a run spends a
Claude subscription rather than API credit, and --endpoint aims it at
the provider the deployed app really uses — which is the one question
this whole exercise could not answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
Ran the memory prompt end to end against a real model as a controlled
A/B: one story generated through the app's own build_context a turn at
a time, then both prompts run over the same blocks, so the story is held
constant and the prompt is the only variable. Fresh process per call, so
neither arm sees the other and the model is never told what is being
tested. The control is the exact prompt from 9cdcb55.
The reported fault reproduced. Two consecutive memories written from one
story minutes apart came back in two different persons — "You crept low
through the mist" and "The player asked Gwen to". With the cast brief
both named Kaelen. The control also inverted who acted on a move whose
player text was "grab her wrist and pull her down", which is the failure
the brief predicts: with no cast there is nothing to say whose wrist
"her wrist" was.
It also found something the prompt review had not. "1-2 plain sentences"
is not a length, and the same model wrote 34 words for one block and 105
for the next. A 105-word memory is a paragraph, and `memory_top_k`
injects five every turn, so the bank's running cost was set by a number
nobody had ever stated.
MEMORY_MAX_WORDS states it, and the prompt now says which details to
keep when trimming: the ones a later scene could turn on. Re-run over
the identical story, the same two blocks came back at 32 and 58 words,
still named, still third person, still carrying the camp map, the
strongbox behind the second tent, and the strap frayed near through.
Variance is the real gain — 34..105 became 32..58.
Overshooting 50 slightly is expected. Models exceed word budgets, which
is why builder.length_hint already carries a buffer for the same reason.
What this does not show: the run used a Claude model, and the app talks
to an OpenAI-compatible endpoint whose weaker models are why
worldstate/parse.py tolerates trailing commas. The prompt is followable
and the brief supplies the missing information; a weaker model is not
proven to comply as well.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
`_create_due_memories` sent six actions of second-person prose and
nothing else — no protagonist, no cast, no setting, and no instruction
about what person to write in. So for `You push the door open. She grabs
your arm.` the only honest memory was "You entered a room and she
stopped you", which names nobody when it is retrieved forty turns later.
The framing wandered too: with no rule, the model picked a person per
call, and one bank ended up holding "You entered the crypt", "The player
entered the crypt" and "He entered the crypt" for the same kind of
event.
Both prompts now carry a cast brief and a framing rule: third person,
the protagonist by name, other characters named rather than left as bare
pronouns. The rule states its reason, because a memory really is read in
isolation much later and a model told why complies far more
consistently than one handed a bare instruction.
The cast comes from the story cards, not from `stat_schema`. Every
schema NPC is already turned into a card at adventure creation,
deduplicated against the hand-written ones by name, so the cards cover
schema NPCs, an author's own cards, and an adventure with no RPG layer
at all through one path instead of three.
Keyword matching alone was not enough, and finding that out changed the
design. Built that way first, the brief for "She grabs your arm" listed
the protagonist and nobody else: the block that most needs a cast is
exactly the one written in bare pronouns, and Gwen's trigger keys
include "her" but the text says "she". So matched cards come first and
the remaining slots are filled with the other character cards. Places
and items are not topped up — an unmentioned tavern is not who "she"
was — though a place that is mentioned still matches normally. The
asymmetry with the turn prompt is deliberate: an untriggered card is
wrong as lore and right in a roster, because the roster answers "who
could these pronouns be" rather than "what is on stage".
Fixed descriptions only, never live values. `Gwen: trust 40 (wary)` in
the brief would make the same event summarized at two different times
come out framed differently, which is the fault this removes.
An adventure with no persona still gets the cast and the setting, and
the model is told to write "the player". One with nothing to say sends
byte for byte the prompt it sent before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
Chromium against a fresh database with the demo scenarios seeded. No API
key needed: Insights assembles the prompt without calling a model, so the
checks run on the real assembled context rather than a stub.
21/21. The two request failures in the run were the sandbox — Google
Fonts is blocked by the egress policy, and the analytics beacon is
aborted on unload — not the app.
The case worth naming: a blank adventure with no RPG layer produced
sections `['narrator', 'persona', 'length_hint']`. That is what the
persona was added for, and it is the one thing no unit test in this repo
would have caught if the section had been gated behind `has_ws`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
The player had stats but no identity. `stat_schema.player` carried hp and
mana beside `npc.gwen.trust`, but where an NPC has a name and a
description the player had neither, so the block rendered as
`You: hp 100/100` and nothing in the prompt said who "you" was.
Three columns on `adventures`: name, pronouns, description. All
user-only, all optional, and an empty name means the app behaves exactly
as it did before — no backfill, no special case for an adventure that
predates the migration.
They are adventure columns rather than part of `stat_schema` for two
reasons. An adventure with no RPG layer still has a protagonist, and
that is the case this was added for. And `worldstate.schema._initials`
treats every dict inside a stat section as a stat definition, so a
persona placed there would be instantiated, rendered in the guide, and
handed an `initial` value as though it were one.
The paths do not change. `player.hp` stays `player.hp`; only the label
moves, to `Kaelen (player): hp 100/100`, the same way NPC lines already
print a display name beside the id. A path carrying the persona's name
would break the moment a player renamed their character, because
`_history_text` replays every past turn's stored delta into the prompt
and those blobs hold literal `player.hp` strings.
The section sits in the system block. Only the user can edit it, so it
never changes mid-story and stays inside the cached prefix. That is what
makes it free, and it is why the AI must not be able to move it — a
delta aimed at `persona.*` is already refused by `_resolve`, and there
is now a test holding that in place.
The modal that used to appear only for scenarios with `${Placeholder}`
tokens now always opens, and is where the character is named. Persona
and placeholders stay independent: a scenario asking for `${Name}` is
asking its own question. No scenario in the repo uses placeholders at
all, so the overlap is hypothetical.
Phase 2, which feeds the persona and the cast to the summarizer, is
written up in plan/18 and not started. That is where the memory-quality
problem actually gets fixed; this change is what gives it a name to use.
Not yet driven in a browser — plan/18 lists what to check by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
`place_action` used to read `depth` from `index`, so the fixture's retry
attempts all landed on the turn's own depth and formed one sibling group.
Now that depth comes from the head, each attempt moved the head one step
and every retry became a turn of its own. The fixture's own check caught
it: a coordinate holding a single superseded attempt has no live row, and
that turn disappears from the story.
The attempts at one turn share that turn's coordinate, so only the first
one goes through `place_action` and the rest copy its placement. This is
what `attempts.add_attempt` does. The fixture cannot call it directly,
because it makes the newest attempt live and the fixture needs a live
attempt that is often not the newest.
Also drop the `variant_index` and `variant_count` arguments, which raise
`TypeError` now, and read the two mark positions from a helper instead of
from the adventure columns migrations 72 and 73 removed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV
`seed_demo.py` still passed `index=0` to `models.Action`, which raises
`TypeError` now that the attribute is gone. Every call site inside `app/`
was updated when the column went, but the seed script sits outside the
package and was missed.
`bundle.settle` wrote `memory_cursor` and `summary_cursor`, which are no
longer mapped, so the assignments only set transient Python attributes.
The version 2 branch existed solely to make those assignments, and it ran
two queries per cursor to do it, so it goes. `settle` now returns early
when the bundle carries anchors. The `db` parameter is unused after that.
Also move the comment about deleted memories next to the
`db.delete(branch)` it describes. Splitting the router package left it
after the `finally`, where it read as attached to nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV
`main` changed the four files this branch split into packages, so all four
came back as modify/delete conflicts. Each change is ported to the module
that now holds the code, unchanged in behavior:
- `delete_action` moves to `routers/adventures/actions.py`. It reaches the
turn lock as `turns.acquire_turn_lock` and `turns._active_turns`, which is
the rule this branch set: the package root no longer re-exports the names a
test rebinds.
- `EMIT_RULE` moves to `worldstate/parse.py`, and `_describe_stat` and
`render_reference` to `worldstate/render.py`.
- The chip-wrap rules move to `styles/schema-editor.css`, and
`StateChangeChips` to `pages/Play/reports.jsx`. `index.css` keeps this
branch's import list.
`test_delete_state.py` needed four edits to run here. It drops the eight-line
database prologue, because `conftest.py` does that once now. It imports the
shared `ScriptedProvider` from `fakes.py` instead of carrying a copy. It
patches `adventures.turns.OpenAICompatibleProvider` and reaches the lock the
same way. And it no longer passes `index=0` when it builds the start action,
because migration 71 drops that column.
564 tests pass. `npm run lint` and `npm run build` are clean, and the CSS
bundle is 56.71 kB, which is the pre-merge size plus main's new rules.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A chip's value carried `white-space: nowrap`, which was there so a signed
number never breaks between its sign and its digits. It was applied to every
value a chip can hold, and most of them are phrases rather than numbers: a
refusal reason, "no change — at its limit", and the whole new value of a
free-text stat, which is a sentence the AI wrote rather than a word.
An unbreakable phrase sets the chip's minimum width to the width of the
phrase, so `max-width: 100%` capped the box while the text kept painting past
the rounded border and off the screen. The chip row sits inside the story
text, which has nothing to scroll, so the page grew a horizontal scrollbar
instead: a 87-character status value measured 470px against a 320px phone.
The nowrap now applies only to the signed number, under `chg-num`. Every other
value wraps inside its chip, which is what the surrounding `overflow-wrap`
already did for the label.
Measured in Chromium at 320, 360, 390 and 430px wide: that same value paints
between 55px and 165px outside its chip before, and 0px after, with the page
no longer scrolling sideways. The number chip stays on one line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST
The referee refused `npc.<id>.status` on a character that has a status stat
under another name, and the refusal was right: the model was reaching for a
path the guide never gave it.
The guide is the fixed list of everything a scenario tracks, and it named
stats in prose. A player stat and a world stat both read as a bare name, and
an NPC's stats read as the display name plus the stat — "Trainer Milo
active_status". The model had to turn that back into `npc.milo.active_status`
itself, and in the Pokemon demo five other characters carry a stat called
`status`, so the path it built was `npc.milo.status`. The live values do state
the paths, but only for the NPCs a scene has mentioned, so an NPC off screen
was addressable only by guesswork.
Each line now leads with the path: `player.potions`, `npc.milo.active_status`,
`flags.sandstorm_active`. An NPC's header is written even when it has no
description, because it is the one line that ties a display name to its id.
`EMIT_RULE` points at the guide for paths rather than at the live values
alone.
A free-text stat is marked `(free text)` beside its path, which is the wording
`EMIT_RULE` already used to describe it — the guide had been writing "free
text" at the end of the line instead. Such a stat now also gets a line when it
has no description, where before it was dropped and the model was left to send
a number for a stat that holds a string.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST
Deleting an AI reply removed the text and left everything the turn did to the
numbers standing. `script_state` and `world_state` belong to the adventure
rather than to the node, so the row going away takes nothing back. Undo, retry,
a take and a branch switch all restore them; this endpoint never did.
The visible symptom was the cooldown clock, which lives in
`world_state._meta.last_changed` and holds a depth. A deleted turn left its
depth there, and the turn played in its place is played at that same depth, so
the referee refused the change as one that had happened this very turn — on a
turn the story no longer contains. Delete the reply because you did not like
the stat change it proposed, press Continue, and the same change comes back
marked "changed too recently". A script's state stacked the same way: the gold
a deleted turn paid out stayed paid, and the replacement turn paid it again.
The restore reads the tip's own outcome rather than the deleted node's
neighbour, which is what `switch` does. Deleting the newest turn then rewinds
to the node in front of it, and deleting one from the middle of the story
restores the state the adventure is already in, so the text goes and the
numbers stay.
The endpoint now takes the turn lock too, for the reason undo takes it: it
writes the shared state, and a turn that is still generating is about to write
it as well.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`,
`variant_count`, `state_before`, and `world_state_before`, plus
`adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in
SQLite, so migration 71 quotes it.
Nothing outside the migrations read these. `models.py`, `schemas.py`, and
`ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders
by `id`, and `attempts.renumber`, `context.history.max_action_index`, and
`nodes.next_index` are deleted.
Two changes keep the migration replayable on a `create_all` database:
- `_split_variants_into_siblings` wrote through the live ORM table, so it
stopped compiling once migration 66 removed five of its columns. It now
writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`.
- Five data passes read columns these migrations drop. Each now calls
`_has_columns` and returns early when the columns are absent.
`bootstrap` takes a `through` version so a migration test can stop at the
schema it asserts on.
555 tests pass, up from 549. Eight of the new cases assert each column is
gone after a real schema-45 database migrates all the way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK
Two items from Stage 2 of `plan/17-refactor.md`.
`frontend/src/hooks/useDebouncedSave.js` replaces three copies of the same
debounce. `PlotPanel` and `ScenarioEditor` held identical per-key timer maps.
`ScriptEditor` held a single shared timer, so editing two fields inside the
same 600 ms window canceled the first save. The hook gives every key its own
timer, which fixes that.
`Settings.stream` was dead state. Nothing read it and every turn streams. This
removes the column, both schema fields, and adds migration 65 to drop it. It is
item S1 in `docs/self-review.md`.
Migration 65 needs a new guard. `_column_already_gone` is the counterpart to
`_column_already_there`: `create_all` builds the current schema, which is
already missing every dropped column, so a fixture that stamps an old version
and replays would fail on a column that is not there.
Verified: 549 backend tests pass, lint and build are clean, and migration 65
runs both ways, once against a database that still has the column and once
against one that does not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
Stage 2, items 1, 4, and 5 of `plan/17-refactor.md`.
**One path resolver in `worldstate`.** `apply_delta` and `apply_override` routed
`flags.<name>`, `milestones.<id>`, `world.<stat>`, `player.<stat>`, and
`npc.<id>.<stat>` with parallel code, about 100 lines each. `_resolve` now says
what a path points at and returns either a target or the rejection to report.
Each function keeps its own write rule, because the rules genuinely differ: an
override sets a number rather than adding to it, ignores `cooldown`,
`max_delta_per_turn`, and the rule that a counter only counts up, and can un-set
a milestone.
A differential check ran both implementations over 3960 payloads: twenty paths,
fourteen values, three starting states, plus every three-path combination. The
results are identical except that 674 rejections from `apply_override` now carry
a `fix` string. `apply_delta` already worded those, and the world-state editor
renders them, so an override that names an unknown flag now explains itself the
way a delta does.
**`sse`, `SSE_HEADERS`, and `turn_error` move to `app/sse.py`.** Two routers
stream, and `chat.py` had to import from `routers.adventures` to reach them.
**`get_adventure_or_404` becomes the `current_adventure` dependency.** All 32
handlers repeated the call as their first statement. The ownership check now
reads in the signature and runs before the body. FastAPI caches a dependency for
one request, so the handler's `db` is the session the adventure came from.
The generated OpenAPI document is byte-identical except on `rename_branch`,
where `branch_id` is now listed before `adventure_id`, because that handler no
longer names `adventure_id` itself. Parameter order in the document is
cosmetic.
Six tests in `test_state_revert.py` call `undo_turn` and `retry_action`
directly rather than over HTTP. They pass the adventure they already hold
instead of an id.
549 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
`frontend/src/pages/Play.jsx` was 2280 lines holding 27 components. It is now
`frontend/src/pages/Play/`: `index.jsx` with the page component, five panels,
two drawers, the reports, the take pager, the refresh dialog, and the shared
formatting helpers.
Every line moved verbatim. A coverage check confirms every non-blank line of
the original appears exactly once, in order, across the twelve files, and a
name-resolution check confirms every identifier each file references is defined
or imported there, with no unused imports.
The line split stranded a comment at six of the boundaries. A leading comment
sits above the section it describes, so each cut left one at the end of the file
before it. All six moved to the section they describe, rewritten in the house
style.
`usePlaySession.js` is not here. The page component still owns all of the
session state. Moving eighteen `useState` calls and seven `useEffect` calls is a
rewrite rather than a move, and no frontend test would catch a mistake in it
today, so it waits for Stage 5.
Four stand-in providers under `backend/tools/` now define `last_usage = None`.
The turn engine reads that attribute after every call, and the fixtures never
defined it, so `tools.tree_fixture` crashed. That break predates this branch.
Verified by 549 passing tests, a clean `npm run lint` and `npm run build`, and
by driving the Play screen: the story view, all five panels, both drawers
including the world-state edit form, the branch map, the refresh dialog, and the
take pager stepping onto a take that lives on another branch. No console errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
`frontend/src/index.css` held 2865 lines. It is now an import list, and each
section is its own file under `frontend/src/styles/`.
The order is unchanged. Two comments in the file already recorded that the
cascade depends on it: several selectors in the tome section win only because
they come after the base card rules, and the 720px query overrides the whole
desktop design. The built CSS bundle is byte-identical before and after, at
56686 bytes.
The plan asked for each `@media` block to move next to the rules it overrides.
That is not done, and should not be. Moving a rule past a later rule of equal
specificity changes which one wins, and there is no frontend test that would
catch it. The 720px block is now `responsive.css`, imported last.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
`backend/app/worldstate/engine.py` held 918 lines covering four separate jobs:
reading a scenario's schema, parsing the block the model writes, applying a
change within the schema's limits, and rendering state as prompt text. Each is
now its own module, the largest 418 lines.
An AST comparison against the old file confirms all 29 definitions are
identical. No call site changes, because `worldstate/__init__.py` exports the
same names it did before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
`backend/app/routers/adventures.py` held 2353 lines and 35 endpoints. It is now
a package of 14 modules, the largest 443 lines.
The split moves text rather than rewriting it. An AST comparison against the
old file confirms all 86 definitions are identical, and the OpenAPI schema
still lists the same 35 operations.
Names a test replaces now live in `turns.py` only, and other modules reach them
as `turns.<name>`. Rebinding a re-exported alias changes the alias and leaves
every caller reading the original, so the package root does not re-export them.
A patch aimed at the old target raises `AttributeError` instead of passing while
doing nothing. Tests and the fixtures in `backend/tools/` say
`adventures.turns.<name>`.
The same rule keeps the turn lock working. One module owns `_active_turns`, so
one lock guards one set.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
Every test module carried the same eight-line prologue redirecting the
database to a temp file. Only the first one to be imported ever took
effect: `app.database` reads `AIDND_DB_PATH` at import and builds `engine`
from it once, so by the time the second module ran the engine already
existed. The other 34 copies created a temp file that nothing opened and
nothing deleted, and leaked one per module per run.
`conftest.py` now does it once, which is early enough because pytest
imports conftest before any test module. It also deletes the file when the
run ends. The tests still share one database, exactly as they already did:
each `client` fixture calls `create_all` on setup and `drop_all` on
teardown, so no test sees another test's rows.
`tests/fakes.py` holds the one `ScriptedProvider`. Nine modules each had a
copy, and the copies had drifted into four feature sets, so a test that
needed to raise a provider error had to be written in one of the files
whose copy supported that. The shared one is the superset. The two
`FakeProvider` copies were the same class with a fixed reply, so they use
it too. `test_chat.py` keeps its own, which implements `chat` rather than
`generate` and records what it was constructed with.
An autouse fixture resets the fake's class state between tests, so a stale
reply list can no longer reach the next test.
435 lines out of the suite. 549 tests pass. Verified live by sabotage:
breaking the shared fake fails 13 tests across four modules.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
Phase 17 splits the four files that hold most of the code, finishes the
schema migration SP8 left half done, and stops the published guide from
drifting away from its Markdown source. `plan/17-refactor.md` carries the
plan and the progress table, and `plan/STATUS.md` points at it.
Stage 0 is hygiene only. Both abandoned worktrees are gone, which freed
about 104 MB. Removing `sp7-tree-ui` needed one extra step: a Vite dev
server had been running out of it since 2026-08-18, holding
`frontend/.vite` open and owning port 5173, and serving a tree 54 commits
behind `main`. The three stale `.db` files are deleted; `data.db` is not.
`AIDND_TRUSTED_PROXY_HOPS` is now documented. It was read at `limits.py:55`
and named in no `.env.example`, README, or blueprint. It sets how many
proxy hops the rate limiter trusts in `X-Forwarded-For`, so a deployment
that adds a hop without setting it gets the bucket-rotation bypass back.
The 19 squash-landed branches are still there. `git branch -D` is blocked
by the permission classifier; the verified command is in the plan file.
549 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
`previous_titles` stops a rename stranding the row it left behind, but the rows
already stranded still had to be deleted by hand on every deployment. The
seeder now removes them on the next boot, which retires the stale "Road to the
Champion" demo without a database console.
Only rows with a NULL owner and `is_public` are considered, and a player's own
scenario is neither, so nothing anybody created is reachable. An adventure
started from a deleted demo survives: `adventures.scenario_id` is `ON DELETE
SET NULL`, and the adventure holds its own copies of the cards and scripts, so
it loses only the inherited cover art.
Two cases skip the sweep, because neither is an instruction to remove live
content: a seed file that fails to parse claims no title, and an empty seed
directory reads as a packaging failure.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
An adventure has no cover art of its own and inherits its scenario's, and a
bundle carries no scenario id, because an id means nothing in another database.
The starter card therefore fell back to a monogram tile while the demo beside
it showed the Pokeball.
The starter file names its source under `scenarioTitle`, and the copy is linked
to the seeded scenario with that title. If no seed answers to the name, the
adventure keeps a NULL `scenario_id`, which is the state every imported bundle
is in and costs only the artwork.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by
the local `claude` command line tool, so a demo can be played against a real
model with no API key. Each request spawns one `claude --print` process, which
suits the turn engine: the app assembles the whole prompt every turn and
expects a stateless endpoint.
Playing the Pokemon demo through it made all five of the world-state changes in
`plan/16` work, and the refusal loop ran end to end for the first time: turn 3
clamped to nothing, turn 4's assembled prompt carried the correction verbatim,
and the model's next delta was right. Two failures previously blamed on the
code were the demo model. One is a schema fault and is still open: Milo's three
Pokemon share one `active_hp` stat, so a switch leaves the newcomer at 0 HP.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
An empty account gives a visitor nothing to read, and the daily demo turns are
limited, so learning what the app does cost one of them. `app/starter.py` now
copies a shipped export bundle into each new guest at the point the row is
created. The bundle is two exchanges of the Pokemon demo, which ends on a
knockout and shows an applied change, a refused one, and a milestone.
The guest row is committed before the copy is attempted, so a failure there
still leaves them with an account, and the copy runs inside a savepoint.
The row building that `POST /adventures/import` did inline moved into
`bundle.materialize`, which both callers use. The rate and size checks stayed
in the endpoint: the starter writes a file the server ships, so it has no
untrusted list to cap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
`seed.py` matches a seed file to its scenario by title, so renaming one
inserted a second public scenario and stranded the first. The stranded row
stays public forever and has to be deleted by hand on every deployment, which
is what happened to "Road to the Champion". A seed file now lists its old
titles under `previous_titles`, and the rename updates the existing row.
The cover art is a PNG data URI. `app/images.py` accepts raster formats only,
because SVG can carry script and the bytes are served from the app's own
origin, so an SVG stores but yields an empty `image_url`.
`tools/make_pokeball.py` draws the ball with `zlib` alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
The model sent a total rather than a change for almost every number:
`npc.ivysaur.hp: 96`, `npc.milo.active_hp: 88`, `player.potions: 2`. Because
every `hp` starts at its maximum, each one clamped back to where it started, and
the potion count rose when the player spent one.
`EMIT_RULE` does say "deltas (not new totals)", but the scenario contradicted it
at closer range. `milo.active_hp`'s description said "Reset this to the
newcomer's full HP", which asks for an absolute and is injected every turn. The
bullets said "drop the HP, and raise it when healed", naming a direction but no
sign. Five HP descriptions said only whose HP it was. The one line that said
"not a delta" covered `player.active_pokemon`, so naming the exception made the
rule look optional. `pokemon_fainted` is the control: its description says "add
1 each time", and it is the only number that behaved.
Every stat description now states the sign, and the lead-in gives a worked
example. `milo.active_hp.max_delta_per_turn` goes 65 to 98, because a switch
moves that stat a full bar from 0 and the old cap made the reset unreachable in
one turn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
`world_delta_of` wrote `delta` and `applied` only. Every consumer that tells a
refused change from a successful one reads the two lists it dropped:
`Action.world_changes` marks a clamped chip from `clamped` and builds its
refusal chips from `rejected`, and `worldstate.refusals` reads both. So no chip
could report a limit, no rejection chip could appear, and no correction ever
reached the next prompt. The three mechanisms merged last session were live in
the code and unreachable in production.
Found by playing the Pokemon demo. Turn 3's snapshot held two clamped entries
with correct `fix` text, both chips came back `clamped: false`, and turn 4's
prompt carried no correction, so the model repeated the same mistake.
The 21 tests passed because `action()` built the column by hand with every list
present. It now fills the column through `world_delta_of`. Removing the two new
lines fails 9 of the 23 tests; that was checked by sabotage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
`plan/STATUS.md` pointed at SP9 as the next step and called it unmerged and
undeployed. SP9 and SP10 are both on `main`. They landed as squash merges, so
`git branch --no-merged` still lists their branches. The doc now points at SP8
and says to check `actions.parent_id` and commit `c0cd6fa` instead of the
branch list.
`plan/15-pokemon-demo-handover.md` named a cause for each of two bugs and both
causes were wrong. It carries a banner and inline corrections rather than a
rewrite, because the reasoning trap it fell into is worth keeping: the visible
evidence was "the number did not move", which reads as the model never trying.
`plan/16-world-state-refusals.md` records what the engine was actually doing,
the five changes, and a browser checklist. None of the UI work has been driven
in a browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
`apply_delta` records three outcomes for every change the model sends:
`applied`, `clamped`, and `rejected`. Everything downstream read only
`applied`. A refused change reached the player as an ordinary chip, and
reached the model on the next turn as a change that had succeeded.
Five parts:
- `Action.world_changes` reads `clamped` and `rejected` beside `applied`.
Accepted stats carry a `clamped` flag; refusals become `kind: "rejected"`
entries. The `fix` key is present only when the engine wrote one, because
this property runs for every action of every list response.
- The UI separates the three outcomes. A clamp to a standstill reads
`no change - at its limit` on a dashed chip, a partial clamp is marked
`(limited)`, and a rejection carries its reason. Dashed and dimmed rather
than red: a refused change means the rules are working.
- The goals line names the milestone id, as `milestones.<id>`. The ids
appeared nowhere in the prompt before, so the model could not send one.
- Each rejection, and each clamp that moved nothing, builds a `fix` string
from the stat definition at the point of refusal. `render_refusals()`
renders them into the next prompt above `EMIT_REMINDER`.
- `_history_text` replays `applied_delta()` instead of the sent delta, so a
past turn's state block shows only what the engine accepted.
A clamp that reduced a change but still moved the value reports nothing. If
you tell a model its 80 damage became 30, it can treat the shortfall as a
debt and send the remaining 50 next turn, which is the swing
`max_delta_per_turn` prevents.
In the demo scenario, `pokemon_left` becomes `pokemon_fainted`
(`type: counter`, `initial: 0`). Starting at the ceiling turned a wrong-signed
delta into a silent no-op; counting up puts the wrong sign on the counter
rule, which refuses it out loud. The instructions also now ask for
`world.turn`, which sat at 0 for a whole playtest.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
Each teammate is its own named entity with hp/status, the same shape
Milo already uses, rather than <name>_hp/<name>_status flattened
under player. Also replaces the five mutually-exclusive _active
flags with a single player.active_pokemon text field, mirroring
npc.milo.active_pokemon, so the AI no longer has to self-enforce
"exactly one flag true."
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
Each of 5 team Pokemon now tracks its own HP and status condition,
one flag marks who's active, and potions are a depletable resource.
The opponent NPC tracks its active Pokemon the same way. Drops the
crowd-favor stat, which didn't fit a battle scenario.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
Adds docs/architecture.html (a standalone walkthrough of the context
budget, world-state engine, story tree, egress fix, and pen-test
findings) and links it from the landing page. Also clarifies in
GUIDE.md §1.8 that "no branching" describes per-turn control flow,
not the story tree's branching data.
Trim em dashes, convert first-person plural to second person, and
tighten sentences in README.md, docs/GUIDE.md, docs/self-review.md,
and frontend/README.md, matching the style already applied to code
comments. No technical content, numbers, or code blocks changed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7vUFpYkLrJnuSgwcRjqSx
* Rewrite comments in Google developer documentation style
Rewrite the comments and docstrings across the backend core modules so they
read plainly. The previous prose was accurate but dense and figurative, which
made it slow to skim.
Applies the Google developer documentation style guide: short sentences, active
voice, present tense, American spelling, and no metaphors, idioms, or
rhetorical asides. Replaces em-dash chains with separate sentences.
Writing below a take is the one moment the server is told which take is
meant, and it obeys before a token is generated. The transcript only
learned that from the resync after the turn, so clearing the preview at
Send time put the replaced take back on screen for the whole turn — and
left it there for good if the resync never landed: a turn that errored, a
lost connection, a tab closed mid-generation. The story then read one way
and reloading the page read another, which is what a player reported.
The chosen take is pinned instead of dropped. Same text override as a
preview, but it does not truncate the story below it, because the turn
being played goes there; the pager's ordinal follows it so the number
matches the words above it. The pin is released only when a re-read
actually succeeds, so a failed read leaves the correct take on screen
rather than reverting to the one it replaced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vo14iJFj4TZuMLhbUeH2LF
Prompt caching bills on a shared prefix: the endpoint reuses the request up
to the first byte that differs from last time and no further. The live
world-state block sat third from the top of the system message, so every turn
re-priced the instructions, the plot essentials and the whole story history
underneath it. The retrieved memories and the rewritten summary did it again.
Everything fixed is emitted first now, and everything that moves goes after
the history, ordered least-volatile first — which is also where recency serves
it best, the reasoning that already put the emit reminder last. The three tail
sections that are last for their own reasons stay last. The moved sections are
still charged to the token budget; only their position changed.
Two smaller halves of the same problem. OpenRouter serves a model from
whichever upstream is free and each upstream holds its own cache, so a
deepseek model now names deepseek as its preferred upstream — a preference,
not a restriction, so a turn still runs if that upstream is down. And the
endpoint's usage block is read back off the response and kept per attempt, so
the hit rate shows up in Insights and the debug log instead of being assumed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
A change chip like "bandit leader aggression +10" was nowrap inside the
story column, which has nothing to scroll, so it overran narrow screens.
The label may now break as a last resort while the value stays glued
together, and the analytics range buttons wrap onto a second row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
The section was written before the commit and said "not committed, not
deployed", which stopped being true one commit ago. It now names 041f9e2 and
the deploy.
The item that outlives the commit gets its own paragraph:
`AIDND_ANALYTICS_EMAILS` is `sync: false`, so nothing in the repo can set it
and the dashboard stays invisible to everybody until it is filled in by hand
in the Render dashboard. Worth saying plainly that collection runs regardless
— the counters fill either way, so setting it late costs no data, which is the
thing that decides whether this is urgent.
The gate paragraph below it loses its closing sentence, which said the same
thing in the same words.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
A hosted demo raises a question a local app never does: is anyone using it,
and do they reach the part that matters? `/analytics` answers it — visitors,
pages, referrers, countries, devices, which shared scenarios get played, turns
and demo-key spend, API and turn errors, and a funnel from visited to played a
turn to signed up.
Not a third-party script, for reasons specific to this one. The CSP allows
`script-src 'self'`, so a tracker means loosening it; adblockers eat the
popular ones, which silently biases exactly the technical audience this
project gets shown to; and none of them can see the measurement that actually
matters here, which is a turn, not a pageview.
**A visit is a write and never a read.** After the 189x egress fix it would be
perverse to add a feature that reads rows per request, so counts accumulate in
a process-local dict and flush every 60s as UPSERTs. Storage is a generic
`(day, metric, label) -> hits` counter, so measuring something new later costs
a constant rather than a migration, plus one row per visitor per day for the
funnel flags. Every dashboard query is a GROUP BY returning tens of rows
however much traffic sits behind it; a month reads back in a few kilobytes.
The buffer's cost is that a hard restart can lose up to a minute — the flusher
also runs on shutdown, and a tier that sleeps when idle sleeps on an empty
buffer anyway.
**The counters are anonymous; the access log beside them is not, on purpose.**
A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars —
one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined
back to `users`, and keyed, so no client can compute one. Story content never
reaches that module, and the only content it ever names is a seeded public
scenario's title; a player's own titles are theirs. `accesslog.py` is the
identifying half and is a separate module writing a separate table so that
separation is a property of the code rather than a convention: `access_events`
records sessions, sign-ins, registrations and failed attempts with address,
email and device, read on a second tab of the same page behind the same gate.
Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An
unmetered tester is not automatically someone who should see the traffic. The
route 404s and the nav link is absent for everyone else, the same treatment
AI Chat gets; unset in a hosted deploy means nobody sees it, including me.
Three things came out of building it that a test would not have suggested.
**A failed turn is an HTTP 200 with a bad ending.** The status-code middleware
cannot see one, so a demo whose model had started refusing every request would
look perfectly healthy from outside. All five SSE error paths in
`_generate_turn` now go through a `turn_error()` helper that counts on the way
out. Error buckets elsewhere are labelled by the matched route template rather
than the requested path — one bucket per endpoint instead of one per adventure
id, and, the reason it isn't merely tidier, an unmatched path is entirely
attacker-chosen, so labelling by it would let anyone mint rows.
**The funnel counts people, not clicks.** A player who starts six adventures
is one person who started an adventure. That is the whole reason the
per-visitor-day table exists; its flags only ever turn on, and `is_new` is
settled by the first write of a visitor's first day.
**The tests run on SQLite and production is Neon.** A flush that raises is
caught and logged, so a dialect mistake in the UPSERTs would have stayed
invisible until the dashboard quietly never filled.
`test_the_upserts_compile_for_postgres` compiles both statements against the
Postgres dialect without connecting to one.
Two things this leans on elsewhere. `limits._client_ip` is now public
`client_ip`: the access log needs the same answer, and two functions both
deciding which hop is the caller's is how one of them ends up trusting a
header it shouldn't. And the cleanup sweeper now starts if *either* job has
work — a deployment can keep every guest forever and still want its
visitor-day rows aged out.
No migration. Both tables are new and `bootstrap()` calls `create_all` on
existing databases too, the route `branches` took in Phase 14, so
`LATEST_VERSION` is still 64.
497 tests green, frontend lint and build clean, driven by hand against a
synthetic 90-day fixture at 1568px. The narrow-screen layout follows the
existing 720px block but is unverified: `resize_window` is ignored on a
maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized
iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run
through the test names themselves, leaving `testclient_ip_*` — still collected
by pytest, which is why it passed unnoticed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
A ceiling alone is a one-sided instruction, and models read it in opposite
directions. A verbose one is held back by it; a terse one has nothing to act
on except "write only as much as the moment needs -- a typical turn is much
shorter" and collapses to two paragraphs. Same prompt, wildly different turn
lengths depending on which model is behind it.
State a floor as well, so the guidance is a band. The two bounds are
deliberately asymmetric -- "must not exceed" for the wall the endpoint
enforces, "should not stop short of" for the floor -- so neither reads as a
number to hit, which is the property the earlier A/B says decides whether
this hint helps or hurts. "Prefer the lower end" inherits the anti-overshoot
job the deleted "much shorter" line was doing, but now with a number under
it, so a terse model lands on the floor instead of at forty words.
Below MIN_LENGTH_FLOOR_WORDS the floor is dropped and the tight-cap wording
is left byte-identical: at a tight cap a short turn is the correct turn, and
that phrasing is the one measured to keep the state block alive (0/6
truncations at cap 250 against 2/6 unhinted). So this only moves loose caps.
MAX_LENGTH_FLOOR_WORDS keeps the share from demanding 555 words minimum at
cap 2400 -- a big cap means long turns are allowed, not compulsory.
Shipped without an A/B run, deliberately. Two things to watch live: whether a
stated range invites landing mid-range on verbose models (drop the share to
~0.25 if so), and whether the state block still survives -- nothing reads
finish_reason yet, so truncation is silent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
The README, the project page and the engineering guide all describe a linear
story. The tree shipped two days ago. Every published surface is a phase
behind, and the guide is not merely behind — it is wrong in a way that costs a
reader time.
Its 2.2 was "Two coordinate systems, and the bug class they create", and it
explained the codebase through position_of_index, note_action_removed and
settled_story_actions. All three were deleted in SP3. 2.3 explained retry
through Action.variants and state_before. Somebody reading either would go
looking for machinery that is not there, which is worse than a gap.
So 2.2 is now "The story is a tree", written at the depth 1.2 and 1.3 are
written at: the seven bugs that turned out to be one bug, the lineage clause
and the two properties that make fork count free, why takes group by parent_id
rather than by coordinate, cursors becoming anchors, and a closing list of what
the design is honest about. 2.3 is rewritten around state_after and takes, and
1.1 and 1.5 follow, because the pipeline no longer snapshots before the call
and the memory bank no longer holds an action back.
The numbers were simply old: 151 tests where there are 440, 37 migrations where
there are 64, twelve phases where there are fourteen. They appear in four
places across the README, the project page's stat tiles and the guide's results
table. The measured branch cost — 103 B, and 1.007x the page load of the same
story flat — is added beside the egress and turn-cost figures it belongs with,
since it is the number that answers "what does branching cost me".
Three screenshots, on a new tools/shots_fixture.py: the Bandit Camp demo driven
through eight written turns with written deltas, three discarded takes forked
onto branches of their own, one off a branch so the map has to nest. Same
reason tree_fixture.py is committed — the shots have to be reproducible and the
frontend still has no test runner. play-world-state.jpg is reshot because it
predates the entire tree UI; the map and the branches panel are new.
Note for next time: docs/guide.html is hand-written, not generated from the
Markdown, so every guide edit is two edits in two vocabularies. Both files were
checked for tag balance and both pages rendered locally before this landed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY