Replicate the A/B on a fresh story, and narrow one claim

The checked-in harness was run end to end through the shim on a newly
generated story, which is both a check that it works and an independent
replication.

The finding the change is for held: across two runs, none of the four
control memories names the protagonist and all four treatment memories
do. The length finding held too — controls at 34, 89, 105 and 107 words,
a four-fold spread with no budget stated anywhere, against 54 and 55
under MEMORY_MAX_WORDS.

One claim did not replicate, and the writeups now say so. Run 1 produced
two control memories in two different persons, which is the reported
complaint exactly, and plan/18 presented that as reproduced. Run 2's
controls were both "the player", consistently. Drifting between second
and third person is therefore something a model sometimes does, observed
once, not something it does every time. The naming gap is the durable
result and the writeups now lead with that instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
This commit is contained in:
Claude
2026-08-31 15:03:20 +05:30
committed by Parth
parent 712ef44f57
commit 71b24b6229
3 changed files with 176 additions and 10 deletions
+10 -3
View File
@@ -104,9 +104,16 @@ cards.
**Running it against a real model found a second fault.** "1-2 plain sentences" is not a
length — the same model wrote 34 words for one block and 105 for the next, and
`memory_top_k` injects five every turn. `MEMORY_MAX_WORDS = 50` states it; the same
blocks then came back at 32 and 58. The A/B harness is `backend/tools/memory_ab.py`,
it drives the real provider through `tools/claude_shim.py`, and the full transcript is
in `plan/18-appendix-memory-ab-run.md`.
blocks then came back at 32 and 58, and a second run on a fresh story came back at 54
and 55 against controls of 89 and 107. The A/B harness is `backend/tools/memory_ab.py`,
it drives the real provider through `tools/claude_shim.py`, and both transcripts are in
`plan/18-appendix-memory-ab-run.md` and `-run-2.md`.
**Read one of those findings narrowly.** Run 1 produced two control memories in two
different persons, which is the reported complaint exactly; run 2's controls were both
"the player", so that drift is something a model sometimes does, not always. What holds
across both runs is that none of the four control memories names the protagonist and all
four treatment memories do.
**Still unmeasured: whether a weaker model complies.** The run used a Claude model
through the shim. The app talks to an OpenAI-compatible endpoint, and