Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and would have been gone with the container. The numbers in plan/18 were therefore assertions nobody could check. plan/18-appendix-memory-ab-run.md now carries the whole transcript: both memories, both summaries, and the thirteen-action story they were written from. backend/tools/memory_ab.py reproduces it. It replaces the throwaway script the first run used, and differs in two ways that matter. It goes through OpenAICompatibleProvider rather than calling a model directly, so a run exercises the provider, the streaming path and complete() instead of a stub. And it reads the control prompt out of git at the commit given to --before, so the thing being compared against cannot drift from what actually shipped. There was already a claude_shim.py serving an OpenAI-compatible endpoint backed by the CLI, which is exactly what the throwaway script had reinvented. memory_ab.py points at it by default, so a run spends a Claude subscription rather than API credit, and --endpoint aims it at the provider the deployed app really uses — which is the one question this whole exercise could not answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
This commit is contained in:
@@ -460,12 +460,23 @@ is the first thing to try.
|
||||
|
||||
## Run with a real model, and what it changed
|
||||
|
||||
Run end to end against a real model, using the `claude` CLI as the provider: one
|
||||
story generated through the app's own `build_context` a turn at a time, then
|
||||
**both** memory prompts run over the **same** blocks, so the story is held
|
||||
constant and the prompt is the only variable. Every call is a fresh process, so
|
||||
no arm can see the other's output and the model is never told what is being
|
||||
tested. The control is the exact prompt from commit `9cdcb55`, not a paraphrase.
|
||||
Run end to end against a real model. One story generated through the app's own
|
||||
`build_context` a turn at a time, then **both** memory prompts run over the
|
||||
**same** blocks, so the story is held constant and the prompt is the only
|
||||
variable. Neither arm sees the other's output, and the model is never told what
|
||||
is being measured. The control is `MEMORY_SYSTEM_PROMPT` as of commit `9cdcb55`,
|
||||
read out of git rather than pasted, so it cannot drift from what shipped.
|
||||
|
||||
The harness is `backend/tools/memory_ab.py`, and it goes through
|
||||
`OpenAICompatibleProvider` rather than calling a model directly, so the run
|
||||
exercises the provider, the streaming path and `complete()`. Pointed at
|
||||
`tools/claude_shim.py` it spends a Claude subscription instead of API credit:
|
||||
|
||||
python tools/claude_shim.py --port 8787 &
|
||||
python tools/memory_ab.py --out ab.md
|
||||
|
||||
**The full transcript, with every memory, both summaries and the story they were
|
||||
written from, is in `plan/18-appendix-memory-ab-run.md`.**
|
||||
|
||||
### The reported fault reproduced, and the fix held
|
||||
|
||||
|
||||
Reference in New Issue
Block a user