Turn the memory bank on for the long run, and refuse one that cannot use it

The first complete hundred-turn campaign did not exercise M01's
"summary/memory activation" step. Memory bank and auto-summarize are
per-campaign switches that default to off, and m11_long_run never
turned them on: summary_tokens and memory_tokens were 0 on every turn,
memories_used was empty, and M04's clue was recalled through narrative
state alone. The retrieval path M6 built was never asked, and nothing
in the evidence said so except a row of zeros.

setup now PATCHes both switches on, reads the campaign back, and stops
before the first turn if either did not take. memories_in_bank is
recorded on every turn, in the final summary and in the recall, and the
recall also says whether a summary exists, so which of the two recall
paths succeeded is stated rather than implied.

The embedding model is now required. Without one the summary pass still
writes memories, but memorybank.retrieve answers "No embedding model
configured" and returns none -- the same unexercised path in a fuller
bank. The harness refuses before it starts a server or claims --out.

tests/test_m11_long_run_memory.py drives setup against the real
application in-process: the switches are on afterwards, a server that
ignores the PATCH is refused before any state is written, the bank
count comes from the application and reads -1 rather than raising when
it cannot, and a run with no embedding model is refused. The four that
exercise setup and the bank count were run against the previous
harness and fail there; the premise test (a fresh campaign has both
switches off) passes on both, as it should.

The 2026-09-10 run in ~/m11-evidence/m01 therefore does not count as
M01. It has to be run again on this harness.

Backend 1,382 passed, 18 skipped, 0 failed. The eighteenth skip is
test_built_spa_fetches_no_fonts_remotely, which wants a built
frontend/dist this worktree does not have; it is an environment
condition, not a change here. The frontend is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKWHt2DXuvqP83cAk6Zq88
This commit is contained in:
JesseMarkowitz
2026-09-12 21:56:14 -04:00
co-authored by Claude Opus 5
parent ef25b0a876
commit fec46f66bb
2 changed files with 240 additions and 3 deletions
+56 -3
View File
@@ -3,7 +3,7 @@
python -m tools.m11_long_run --turns 100 --out <dir>
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT`, `AIDND_TEST_MODEL` and
optionally `AIDND_TEST_EMBED_MODEL`.
`AIDND_TEST_EMBED_MODEL`, all three required.
## Why this is a script that spawns servers rather than a test
@@ -26,6 +26,17 @@ at planned points. Everything that survives crosses as bytes on disk.
window and then retrieved, so the run measures the window at intervals and the
recall check at the end asks the application what it would actually send.
M01's step list also asks for **"summary/memory activation"**, and both are
per-campaign switches that default to off. An earlier version of this harness
never turned them on, so a hundred turns ran with an empty memory bank and no
summaries: six of M01's seven clauses were exercised and the seventh was
reported by silence. Worse for M04, whose whole question is whether a fact
planted at turn one is still reachable at turn a hundred — with the bank off it
was reachable through narrative state alone, and the retrieval path M6 built was
never asked. `setup` now turns both on and proves it, and `measure` records how
many memories and summaries exist so that "the bank stayed empty" is a number in
the evidence rather than an absence nobody looked for.
## What it records
A JSON line per turn (`timeline.jsonl`) carrying the context measurements M03
@@ -401,6 +412,22 @@ class Run:
self.adv = created["id"]
self.note("campaign", id=self.adv)
# M01 asks for "summary/memory activation". Both are per-campaign and
# both default to False (`models.Adventure`), so a campaign created and
# played without touching them never writes a memory or a summary.
# Enabled here, and read back rather than assumed: a PATCH that silently
# did nothing would leave the same hole this closes.
self.server.call("PATCH", f"/adventures/{self.adv}",
{"memory_bank_enabled": True, "auto_summarize": True})
back = self.server.call("GET", f"/adventures/{self.adv}")
self.note("memory_activated",
memory_bank_enabled=back.get("memory_bank_enabled"),
auto_summarize=back.get("auto_summarize"))
if not (back.get("memory_bank_enabled") and back.get("auto_summarize")):
raise SystemExit(
"the campaign did not accept memory-bank and auto-summarize; "
"M01's summary/memory clause cannot be measured from this run.")
for name, body, kind in (("canon.md", CANON_MD, "canon"),
("reference.md", REFERENCE_MD, "reference"),
("inspiration.md", INSPIRATION_MD, "inspiration")):
@@ -505,6 +532,10 @@ class Run:
# campaign resumed against an older build has neither.
"history_floor_depth": report["history"].get("floor_depth"),
"history_trim_block": report["history"].get("trim_block"),
# M01's summary/memory clause, as a count on every turn. A bank that
# stays at zero is then visible in the evidence while the run is
# still going, instead of being discovered afterwards.
"memories_in_bank": self.bank_size(),
"prompt_tokens": tokens["total"],
"budget": tokens["budget"],
"configured_budget": tokens.get("configured_budget"),
@@ -529,6 +560,15 @@ class Run:
"clue_in_prompt": CLUE_SENTINEL in json.dumps(report["sections"]),
}
def bank_size(self) -> int:
"""How many memories the bank holds. Never raises: this is measurement,
and a turn is not worth failing over a count."""
try:
return len(self.server.call(
"GET", f"/adventures/{self.adv}/memories") or [])
except Exception: # noqa: BLE001
return -1
def state(self) -> dict:
return self.server.call("GET", f"/adventures/{self.adv}/state")
@@ -554,8 +594,12 @@ def main() -> int:
help="stop and write the evidence after this many unaccepted turns")
args = parser.parse_args()
if not (ENDPOINT and MODEL):
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
if not (ENDPOINT and MODEL and EMBED_MODEL):
# The embedding model is not optional. Without one the summary pass
# still writes memories, but `memorybank.retrieve` answers "No embedding
# model configured" and returns none — so the bank fills and M04's
# memory path is never asked, which is the hole `setup` exists to close.
print("set AIDND_TEST_ENDPOINT, AIDND_TEST_MODEL and AIDND_TEST_EMBED_MODEL")
return 2
if not 30 <= args.turn_timeout <= 3600:
# The settings schema's own bound, checked here so a mistyped timeout
@@ -687,6 +731,8 @@ def main() -> int:
"final_measurement": _or_none(run.measure),
"db_bytes": db_path.stat().st_size,
"bundle_bytes": bundle_bytes,
# Whether the seventh clause of M01's step list actually happened.
"memories_in_bank": _or_none(run.bank_size),
}
(out / "summary.json").write_text(json.dumps(summary, indent=2, default=str))
print(json.dumps({k: v for k, v in summary.items()
@@ -939,6 +985,13 @@ def _recall(run: Run) -> dict:
"memories_used": [
m.get("text", "")[:120] for m in (after.get("memories") or {}).get("used", [])
],
# M01's summary/memory clause, stated rather than implied. A recall that
# succeeds only through narrative state, with an empty bank, has proved
# one of the two paths the design has — and the reader of this report
# should be able to see which.
"memories_in_bank": run.bank_size(),
"summary_present": bool(
(server.call("GET", f"/adventures/{adv}") or {}).get("story_summary")),
"prompt_tokens": after["tokens"]["total"],
}