Keep the state section and its proposal out of the story

The first complete M01 run with the memory bank on (f8d4010, 101 turns on
a GPU host) reported "complete". It still did not prove M04. The planted
clue was found at turn 100 only because the narrator had pasted the
narrative-state section into its own prose, and the paste was still in
recent history. No memory and no summary carried the clue.

The narrator is a small local model. It wrote protocol into its stored
narration on 42 of 104 turns, starting at depth 2, in four shapes:

- a copy of the state section: `Scene:`, `Who and what exists:`, `Held:`,
  `Established:`, `Still open:`
- that copy above a correct ```state block, which was stripped while the
  copy stayed
- the copy, then a bare `State` heading, then a `> {"events": ...}`
  proposal quoted like a player turn, sometimes with story after it
- the same block cut off by the output-token limit, on 10 turns

Stored text is replayed verbatim as history, so each leak also put a second,
older account of the state into the next prompt. That is what M5 review
Finding 4 removed from history replay, and every leak gave the model another
example to copy.

The extractor now removes:

- a pasted state section, recognised by at least two of the renderer's own
  headings as whole lines. The headings are constants in `render.py`, so the
  renderer and the extractor cannot drift apart. One heading alone, or a
  `Scene:` line of prose, is left.
- an unfenced proposal that starts a line, quoted or not, when it parses and
  is a proposal. With no fence it becomes the turn's proposal. A `State`
  heading directly above goes with it. Candidates are taken outermost first,
  so a finished event line inside an unfinished block is never taken as a
  proposal by itself.
- an unfinished unfenced proposal at the end that reads as protocol.
- whatever is left at the end: a `State` heading, a bare `>`, a parroted
  reminder or continue hint (closed or not), and a ```json fence cut off
  before it names its events. These are cut repeatedly until nothing more
  comes off.

This also fixes an older bug. `_STATE_FENCE_RE` read "a ```state block"
inside a parroted reminder as a fence opening and cut out the middle of the
reminder. The label must now end its line or run straight into the payload.

A reply whose only removal is a pasted state section records no raw block,
so the turn is not marked unparseable for a block it never started.

Every AI turn in four real runs was replayed through the new extractor:
draco M01, the two 26-turn GPU trials, and this M01 run. 339 turns in all.
No turn the old extractor had left clean changed. Every leak of our own
protocol is gone: 42 of 42 in this M01 run, 5 in trial 2, 3 on draco.
Trial 1 still has model-invented headings ("Identifiers established:",
"Set of events made true:") on 10 turns. They paraphrase the instruction and
are not our renderer's text, so they are left, not guessed at.

The long-run harness now records an explicit M04 verdict, which is never a
recovery while the clue is still in recent history. It also counts the AI
turns in the export that still carry protocol.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-13 21:18:18 -04:00
co-authored by Claude Opus 5
parent f8d401029f
commit 0c7316f951
5 changed files with 615 additions and 25 deletions
+40 -2
View File
@@ -828,11 +828,13 @@ def main() -> int:
# Attempted even for an aborted run: the recovery check and the storage
# numbers are worth having at whatever turn count was reached.
bundle_bytes = None
leaks = None
try:
bundle = server.call("GET", f"/adventures/{run.adv}/export")
(out / "bundle.json").write_text(json.dumps(bundle))
bundle_bytes = len((out / "bundle.json").read_bytes())
run.note("exported", bytes=bundle_bytes)
leaks = _protocol_leaks(bundle)
run.note("exported", bytes=bundle_bytes, protocol_leaks=leaks)
except Exception as exc: # noqa: BLE001
run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300])
@@ -859,6 +861,7 @@ def main() -> int:
"aborted_reason": aborted,
"failed_reason": failed_reason,
"summaries": _or_none(run.summary_count),
"protocol_leaks": leaks,
"accepted_turns": run.accepted,
"turns_requested": args.turns,
"restarts": server.starts - 1,
@@ -1127,7 +1130,7 @@ def _recall(run: Run) -> dict:
document = run.state()["document"]
fact_present = any(
CLUE_SENTINEL in json.dumps(f) for f in document.get("facts") or [])
return {
result = {
"clue_in_recent_history_window": in_history,
"clue_in_prompt": CLUE_SENTINEL in whole_prompt,
"in_state_section": CLUE_SENTINEL in sections.get(STATE_LABEL, ""),
@@ -1151,6 +1154,41 @@ def _recall(run: Run) -> dict:
(server.call("GET", f"/adventures/{adv}") or {}).get("story_summary")),
"prompt_tokens": after["tokens"]["total"],
}
result["m04_verdict"] = _m04_verdict(result)
return result
def _m04_verdict(recall: dict) -> str:
"""What the recall check proved, in one word the report can quote.
The first long run with the bank on found the clue "in the prompt" at turn
100, but only because the narrator had pasted the state section into its
prose and the paste was still in recent history. A clue in recent history is
no evidence about memory, so that case gets its own verdict and never counts
as recovery."""
if recall["clue_in_recent_history_window"]:
return "precondition_not_met"
if recall["in_memories_section"] or recall["in_summary_section"]:
return "recovered_through_memory_or_summary"
if recall["in_state_section"]:
return "recovered_through_state_only"
return "not_recovered"
#: Signs the application stored protocol as story: the state section's own
#: headings, and a proposal's event list. Copied rather than imported from
#: `app.narrative.render`, for the reason `HISTORY_LABELS` is copied.
PROTOCOL_LEAK_MARKERS = ("Who and what exists:", "\nHeld:\n", "\nEstablished:\n",
'"events"')
def _protocol_leaks(bundle: dict) -> dict:
"""How many stored AI turns still carry protocol, on every branch."""
ai = [a for a in bundle.get("actions") or [] if a.get("type") == "ai"]
leaking = [a for a in ai
if any(marker in (a.get("text") or "") for marker in PROTOCOL_LEAK_MARKERS)]
return {"ai_actions": len(ai), "leaking": len(leaking),
"example_ids": [a.get("id") for a in leaking[:10]]}
if __name__ == "__main__":