Keep the A/B, and the harness that produced it

The run that justified MEMORY_MAX_WORDS lived in a scratch directory and
would have been gone with the container. The numbers in plan/18 were
therefore assertions nobody could check.

plan/18-appendix-memory-ab-run.md now carries the whole transcript: both
memories, both summaries, and the thirteen-action story they were
written from.

backend/tools/memory_ab.py reproduces it. It replaces the throwaway
script the first run used, and differs in two ways that matter. It goes
through OpenAICompatibleProvider rather than calling a model directly,
so a run exercises the provider, the streaming path and complete()
instead of a stub. And it reads the control prompt out of git at the
commit given to --before, so the thing being compared against cannot
drift from what actually shipped.

There was already a claude_shim.py serving an OpenAI-compatible endpoint
backed by the CLI, which is exactly what the throwaway script had
reinvented. memory_ab.py points at it by default, so a run spends a
Claude subscription rather than API credit, and --endpoint aims it at
the provider the deployed app really uses — which is the one question
this whole exercise could not answer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
This commit is contained in:
Claude
2026-08-31 15:03:20 +05:30
committed by Parth
parent 07192767d8
commit 712ef44f57
4 changed files with 448 additions and 7 deletions
+17 -6
View File
@@ -460,12 +460,23 @@ is the first thing to try.
## Run with a real model, and what it changed
Run end to end against a real model, using the `claude` CLI as the provider: one
story generated through the app's own `build_context` a turn at a time, then
**both** memory prompts run over the **same** blocks, so the story is held
constant and the prompt is the only variable. Every call is a fresh process, so
no arm can see the other's output and the model is never told what is being
tested. The control is the exact prompt from commit `9cdcb55`, not a paraphrase.
Run end to end against a real model. One story generated through the app's own
`build_context` a turn at a time, then **both** memory prompts run over the
**same** blocks, so the story is held constant and the prompt is the only
variable. Neither arm sees the other's output, and the model is never told what
is being measured. The control is `MEMORY_SYSTEM_PROMPT` as of commit `9cdcb55`,
read out of git rather than pasted, so it cannot drift from what shipped.
The harness is `backend/tools/memory_ab.py`, and it goes through
`OpenAICompatibleProvider` rather than calling a model directly, so the run
exercises the provider, the streaming path and `complete()`. Pointed at
`tools/claude_shim.py` it spends a Claude subscription instead of API credit:
python tools/claude_shim.py --port 8787 &
python tools/memory_ab.py --out ab.md
**The full transcript, with every memory, both summaries and the story they were
written from, is in `plan/18-appendix-memory-ab-run.md`.**
### The reported fault reproduced, and the fix held