Replicate the A/B on a fresh story, and narrow one claim
The checked-in harness was run end to end through the shim on a newly generated story, which is both a check that it works and an independent replication. The finding the change is for held: across two runs, none of the four control memories names the protagonist and all four treatment memories do. The length finding held too — controls at 34, 89, 105 and 107 words, a four-fold spread with no budget stated anywhere, against 54 and 55 under MEMORY_MAX_WORDS. One claim did not replicate, and the writeups now say so. Run 1 produced two control memories in two different persons, which is the reported complaint exactly, and plan/18 presented that as reproduced. Run 2's controls were both "the player", consistently. Drifting between second and third person is therefore something a model sometimes does, observed once, not something it does every time. The naming gap is the durable result and the writeups now lead with that instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
This commit is contained in:
+10
-3
@@ -104,9 +104,16 @@ cards.
|
||||
**Running it against a real model found a second fault.** "1-2 plain sentences" is not a
|
||||
length — the same model wrote 34 words for one block and 105 for the next, and
|
||||
`memory_top_k` injects five every turn. `MEMORY_MAX_WORDS = 50` states it; the same
|
||||
blocks then came back at 32 and 58. The A/B harness is `backend/tools/memory_ab.py`,
|
||||
it drives the real provider through `tools/claude_shim.py`, and the full transcript is
|
||||
in `plan/18-appendix-memory-ab-run.md`.
|
||||
blocks then came back at 32 and 58, and a second run on a fresh story came back at 54
|
||||
and 55 against controls of 89 and 107. The A/B harness is `backend/tools/memory_ab.py`,
|
||||
it drives the real provider through `tools/claude_shim.py`, and both transcripts are in
|
||||
`plan/18-appendix-memory-ab-run.md` and `-run-2.md`.
|
||||
|
||||
**Read one of those findings narrowly.** Run 1 produced two control memories in two
|
||||
different persons, which is the reported complaint exactly; run 2's controls were both
|
||||
"the player", so that drift is something a model sometimes does, not always. What holds
|
||||
across both runs is that none of the four control memories names the protagonist and all
|
||||
four treatment memories do.
|
||||
|
||||
**Still unmeasured: whether a weaker model complies.** The run used a Claude model
|
||||
through the shim. The app talks to an OpenAI-compatible endpoint, and
|
||||
|
||||
Reference in New Issue
Block a user