Files
interactive-story/backend
Claude 712ef44f57 Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and
would have been gone with the container. The numbers in plan/18 were
therefore assertions nobody could check.

plan/18-appendix-memory-ab-run.md now carries the whole transcript: both
memories, both summaries, and the thirteen-action story they were
written from.

backend/tools/memory_ab.py reproduces it. It replaces the throwaway
script the first run used, and differs in two ways that matter. It goes
through OpenAICompatibleProvider rather than calling a model directly,
so a run exercises the provider, the streaming path and complete()
instead of a stub. And it reads the control prompt out of git at the
commit given to --before, so the thing being compared against cannot
drift from what actually shipped.

There was already a claude_shim.py serving an OpenAI-compatible endpoint
backed by the CLI, which is exactly what the throwaway script had
reinvented. memory_ab.py points at it by default, so a run spends a
Claude subscription rather than API credit, and --endpoint aims it at
the provider the deployed app really uses — which is the one question
this whole exercise could not answer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
2026-08-31 15:03:20 +05:30
..
2026-07-07 12:30:06 +05:30