4c94dd0fc08da9c80d4c2828ea52da6b917e3f90
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4c94dd0fc0 |
Keep a production run to the adventures you meant
The hosted database is not the local one. It holds other people's stories, and `summary_provider` builds from the adventure owner's Settings, so an unfiltered `--write` against it would spend other people's money rewriting memories they never asked about. `--adventure` could already hold a run down, but only if you knew the ids. `--email` names accounts instead, and every line of the report now says who owns the adventure, so a dry run answers "whose keys would this spend" before anything is written. Guests have no email and stay reachable only by id, which is the right amount of friction for touching a stranger's bank. The other half is written down rather than built: stored API keys are encrypted with AIDND_SECRET_KEY, so a run from a checkout against the Neon database needs the same value the web service holds. With a different one `decrypt_secret` returns "" instead of failing, and every adventure is reported as having no key — a run that looks like it worked and did nothing. The Dockerfile copies backend/app alone, so tools/ is not on the box either way; the recipe is a checkout pointed at AIDND_DATABASE_URL. 629 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
0633cb624e |
Run the new memory prompt back over an old bank
The prompt change only reaches memories written after it. plan/18 decided to leave the existing ones alone and let eviction age them out at memory_bank_capacity, on the grounds that re-summarizing would duplicate whatever was still in the bank because nothing deletes the old rows. That was wrong about the only option. A memory can be rewritten in place. The row carries more than its text — whether it is pinned, how often it has been retrieved, and the node it hangs off, which is what makes a fork inherit the right memories — and rewriting `text` keeps all of it. Deleting the bank and rewinding the cursor would lose that, and would trickle memories back at MAX_MEMORIES_PER_RUN per turn, so an adventure nobody is playing would never recover. tools/rewrite_memories.py does it. Without --write it makes no model calls and only reports the scope; --write rewrites, --embed re-embeds in the run rather than leaving it to the app's post-turn pass. It reads whichever database the app reads, so it works against the hosted Postgres as well as a local file. Two things it needed from the app. `summarize_block` is now the one place a memory prompt is assembled, and the post-turn pass calls it too — a backfill that built its own prompt would be writing memories with a prompt that never shipped, and nothing would report the drift. `source_block` reads a memory's block back out of the story, which nothing has ever had to do: it reads on the lineage of the branch the memory was written on, not the branch being played, because after a fork the same depths hold different actions on each side and a read through the adventure's path would summarize the wrong story silently. It also excludes the discarded attempts at a retried turn, and tolerates a block an action has since been deleted from. Left alone: a memory with no source range, which is hand-written or migrated by 62 and may be the player's own words; a memory whose actions are gone; and an adventure whose owner has no API key, because summarization spends the user's own key by construction and never the shared demo key. --api-key/--model/ --endpoint override that, the last of them aiming a run at claude_shim.py. The vector is cleared for every rewrite, because the stored one describes wording that no longer exists. Re-embedding always uses the owner's own embedding model, never --endpoint: a vector only means anything against the vectors it is ranked beside. 17 tests, 627 green. The fork case is the one that would fail quietly, so the test builds a fork whose depths hold different actions on each side. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
71b24b6229 |
Replicate the A/B on a fresh story, and narrow one claim
The checked-in harness was run end to end through the shim on a newly generated story, which is both a check that it works and an independent replication. The finding the change is for held: across two runs, none of the four control memories names the protagonist and all four treatment memories do. The length finding held too — controls at 34, 89, 105 and 107 words, a four-fold spread with no budget stated anywhere, against 54 and 55 under MEMORY_MAX_WORDS. One claim did not replicate, and the writeups now say so. Run 1 produced two control memories in two different persons, which is the reported complaint exactly, and plan/18 presented that as reproduced. Run 2's controls were both "the player", consistently. Drifting between second and third person is therefore something a model sometimes does, observed once, not something it does every time. The naming gap is the durable result and the writeups now lead with that instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
712ef44f57 |
Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and would have been gone with the container. The numbers in plan/18 were therefore assertions nobody could check. plan/18-appendix-memory-ab-run.md now carries the whole transcript: both memories, both summaries, and the thirteen-action story they were written from. backend/tools/memory_ab.py reproduces it. It replaces the throwaway script the first run used, and differs in two ways that matter. It goes through OpenAICompatibleProvider rather than calling a model directly, so a run exercises the provider, the streaming path and complete() instead of a stub. And it reads the control prompt out of git at the commit given to --before, so the thing being compared against cannot drift from what actually shipped. There was already a claude_shim.py serving an OpenAI-compatible endpoint backed by the CLI, which is exactly what the throwaway script had reinvented. memory_ab.py points at it by default, so a run spends a Claude subscription rather than API credit, and --endpoint aims it at the provider the deployed app really uses — which is the one question this whole exercise could not answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
07192767d8 |
Give a memory a word ceiling, after measuring one
Ran the memory prompt end to end against a real model as a controlled A/B: one story generated through the app's own build_context a turn at a time, then both prompts run over the same blocks, so the story is held constant and the prompt is the only variable. Fresh process per call, so neither arm sees the other and the model is never told what is being tested. The control is the exact prompt from 9cdcb55. The reported fault reproduced. Two consecutive memories written from one story minutes apart came back in two different persons — "You crept low through the mist" and "The player asked Gwen to". With the cast brief both named Kaelen. The control also inverted who acted on a move whose player text was "grab her wrist and pull her down", which is the failure the brief predicts: with no cast there is nothing to say whose wrist "her wrist" was. It also found something the prompt review had not. "1-2 plain sentences" is not a length, and the same model wrote 34 words for one block and 105 for the next. A 105-word memory is a paragraph, and `memory_top_k` injects five every turn, so the bank's running cost was set by a number nobody had ever stated. MEMORY_MAX_WORDS states it, and the prompt now says which details to keep when trimming: the ones a later scene could turn on. Re-run over the identical story, the same two blocks came back at 32 and 58 words, still named, still third person, still carrying the camp map, the strongbox behind the second tent, and the strap frayed near through. Variance is the real gain — 34..105 became 32..58. Overshooting 50 slightly is expected. Models exceed word budgets, which is why builder.length_hint already carries a buffer for the same reason. What this does not show: the run used a Claude model, and the app talks to an OpenAI-compatible endpoint whose weaker models are why worldstate/parse.py tolerates trailing commas. The prompt is followable and the brief supplies the missing information; a weaker model is not proven to comply as well. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
1b5e4c56cd |
Tell the summarizer who the characters are
`_create_due_memories` sent six actions of second-person prose and nothing else — no protagonist, no cast, no setting, and no instruction about what person to write in. So for `You push the door open. She grabs your arm.` the only honest memory was "You entered a room and she stopped you", which names nobody when it is retrieved forty turns later. The framing wandered too: with no rule, the model picked a person per call, and one bank ended up holding "You entered the crypt", "The player entered the crypt" and "He entered the crypt" for the same kind of event. Both prompts now carry a cast brief and a framing rule: third person, the protagonist by name, other characters named rather than left as bare pronouns. The rule states its reason, because a memory really is read in isolation much later and a model told why complies far more consistently than one handed a bare instruction. The cast comes from the story cards, not from `stat_schema`. Every schema NPC is already turned into a card at adventure creation, deduplicated against the hand-written ones by name, so the cards cover schema NPCs, an author's own cards, and an adventure with no RPG layer at all through one path instead of three. Keyword matching alone was not enough, and finding that out changed the design. Built that way first, the brief for "She grabs your arm" listed the protagonist and nobody else: the block that most needs a cast is exactly the one written in bare pronouns, and Gwen's trigger keys include "her" but the text says "she". So matched cards come first and the remaining slots are filled with the other character cards. Places and items are not topped up — an unmentioned tavern is not who "she" was — though a place that is mentioned still matches normally. The asymmetry with the turn prompt is deliberate: an untriggered card is wrong as lore and right in a roster, because the roster answers "who could these pronouns be" rather than "what is on stage". Fixed descriptions only, never live values. `Gwen: trust 40 (wary)` in the brief would make the same event summarized at two different times come out framed differently, which is the fault this removes. An adventure with no persona still gets the cast and the setting, and the model is told to write "the player". One with nothing to say sends byte for byte the prompt it sent before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
dfb56d4569 |
Record that the persona was driven in a browser
Chromium against a fresh database with the demo scenarios seeded. No API key needed: Insights assembles the prompt without calling a model, so the checks run on the real assembled context rather than a stub. 21/21. The two request failures in the run were the sandbox — Google Fonts is blocked by the egress policy, and the analytics beacon is aborted on unload — not the app. The case worth naming: a blank adventure with no RPG layer produced sections `['narrator', 'persona', 'length_hint']`. That is what the persona was added for, and it is the one thing no unit test in this repo would have caught if the section had been gated behind `has_ws`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
9ee052c51e |
Give the adventure a persona, so the protagonist has a name
The player had stats but no identity. `stat_schema.player` carried hp and
mana beside `npc.gwen.trust`, but where an NPC has a name and a
description the player had neither, so the block rendered as
`You: hp 100/100` and nothing in the prompt said who "you" was.
Three columns on `adventures`: name, pronouns, description. All
user-only, all optional, and an empty name means the app behaves exactly
as it did before — no backfill, no special case for an adventure that
predates the migration.
They are adventure columns rather than part of `stat_schema` for two
reasons. An adventure with no RPG layer still has a protagonist, and
that is the case this was added for. And `worldstate.schema._initials`
treats every dict inside a stat section as a stat definition, so a
persona placed there would be instantiated, rendered in the guide, and
handed an `initial` value as though it were one.
The paths do not change. `player.hp` stays `player.hp`; only the label
moves, to `Kaelen (player): hp 100/100`, the same way NPC lines already
print a display name beside the id. A path carrying the persona's name
would break the moment a player renamed their character, because
`_history_text` replays every past turn's stored delta into the prompt
and those blobs hold literal `player.hp` strings.
The section sits in the system block. Only the user can edit it, so it
never changes mid-story and stays inside the cached prefix. That is what
makes it free, and it is why the AI must not be able to move it — a
delta aimed at `persona.*` is already refused by `_resolve`, and there
is now a test holding that in place.
The modal that used to appear only for scenarios with `${Placeholder}`
tokens now always opens, and is where the character is named. Persona
and placeholders stay independent: a scenario asking for `${Name}` is
asking its own question. No scenario in the repo uses placeholders at
all, so the overlap is hypothetical.
Phase 2, which feeds the persona and the cast to the summarizer, is
written up in plan/18 and not started. That is where the memory-quality
problem actually gets fixed; this change is what gives it a name to use.
Not yet driven in a browser — plan/18 lists what to check by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
|