Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and would have been gone with the container. The numbers in plan/18 were therefore assertions nobody could check. plan/18-appendix-memory-ab-run.md now carries the whole transcript: both memories, both summaries, and the thirteen-action story they were written from. backend/tools/memory_ab.py reproduces it. It replaces the throwaway script the first run used, and differs in two ways that matter. It goes through OpenAICompatibleProvider rather than calling a model directly, so a run exercises the provider, the streaming path and complete() instead of a stub. And it reads the control prompt out of git at the commit given to --before, so the thing being compared against cannot drift from what actually shipped. There was already a claude_shim.py serving an OpenAI-compatible endpoint backed by the CLI, which is exactly what the throwaway script had reinvented. memory_ab.py points at it by default, so a run spends a Claude subscription rather than API credit, and --endpoint aims it at the provider the deployed app really uses — which is the one question this whole exercise could not answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
This commit is contained in:
+40
-1
@@ -3,7 +3,7 @@
|
||||
Read this first when picking the project back up. Updated at the end of a working
|
||||
session; the per-phase plan files hold the detail, this holds the thread.
|
||||
|
||||
**Last updated: 2026-08-28.**
|
||||
**Last updated: 2026-08-31.**
|
||||
|
||||
---
|
||||
|
||||
@@ -76,6 +76,45 @@ needed; nothing requires reading a row of anyone's story.
|
||||
|
||||
---
|
||||
|
||||
## What happened on 2026-08-31 — the persona, and what the summarizer is told
|
||||
|
||||
**`plan/18-persona-and-memory-quality.md` is the writeup; it is on
|
||||
`claude/ai-dnd-memories-summarization-3muo98`, not yet merged.** Two changes, both green
|
||||
at 610 tests.
|
||||
|
||||
**The protagonist now has a name.** An adventure carries `persona_name`,
|
||||
`persona_pronouns` and `persona_desc` (migrations 74-76), and the player's stat block
|
||||
renders as `Kaelen (player): hp 100/100` instead of `You: hp 100/100`. The paths do not
|
||||
change — a path carrying the persona's name would break the moment a player renamed
|
||||
their character, because `_history_text` replays stored deltas holding literal
|
||||
`player.hp` strings. Empty name means the app behaves exactly as before, so no backfill.
|
||||
Driven in a browser, 21/21 checks.
|
||||
|
||||
**The summarizer used to be told nothing.** It got six actions of second-person prose
|
||||
and no cast, no setting, and no instruction about what person to write in. It now gets a
|
||||
cast brief built from the story cards — which already cover the schema NPCs, because
|
||||
`scenario_card_specs` turns every one of them into a card at adventure creation.
|
||||
|
||||
**Keyword matching alone was not enough, and only running it showed that.** Built to the
|
||||
plan first, the brief for "She grabs your arm" listed the protagonist and nobody else:
|
||||
the block that most needs a cast is exactly the one written in bare pronouns. Matched
|
||||
cards now come first and the rest of the roster is filled with the other `character`
|
||||
cards.
|
||||
|
||||
**Running it against a real model found a second fault.** "1-2 plain sentences" is not a
|
||||
length — the same model wrote 34 words for one block and 105 for the next, and
|
||||
`memory_top_k` injects five every turn. `MEMORY_MAX_WORDS = 50` states it; the same
|
||||
blocks then came back at 32 and 58. The A/B harness is `backend/tools/memory_ab.py`,
|
||||
it drives the real provider through `tools/claude_shim.py`, and the full transcript is
|
||||
in `plan/18-appendix-memory-ab-run.md`.
|
||||
|
||||
**Still unmeasured: whether a weaker model complies.** The run used a Claude model
|
||||
through the shim. The app talks to an OpenAI-compatible endpoint, and
|
||||
`worldstate/parse.py` tolerates trailing commas because free models emit them. Point
|
||||
`memory_ab.py --endpoint` at the real provider to find out.
|
||||
|
||||
---
|
||||
|
||||
## Pick up here
|
||||
|
||||
**`plan/17-refactor.md` is the active phase.** It carries its own progress table, which
|
||||
|
||||
Reference in New Issue
Block a user