Keep the state section and its proposal out of the story
The first complete M01 run with the memory bank on (f8d4010, 101 turns on
a GPU host) reported "complete". It still did not prove M04. The planted
clue was found at turn 100 only because the narrator had pasted the
narrative-state section into its own prose, and the paste was still in
recent history. No memory and no summary carried the clue.
The narrator is a small local model. It wrote protocol into its stored
narration on 42 of 104 turns, starting at depth 2, in four shapes:
- a copy of the state section: `Scene:`, `Who and what exists:`, `Held:`,
`Established:`, `Still open:`
- that copy above a correct ```state block, which was stripped while the
copy stayed
- the copy, then a bare `State` heading, then a `> {"events": ...}`
proposal quoted like a player turn, sometimes with story after it
- the same block cut off by the output-token limit, on 10 turns
Stored text is replayed verbatim as history, so each leak also put a second,
older account of the state into the next prompt. That is what M5 review
Finding 4 removed from history replay, and every leak gave the model another
example to copy.
The extractor now removes:
- a pasted state section, recognised by at least two of the renderer's own
headings as whole lines. The headings are constants in `render.py`, so the
renderer and the extractor cannot drift apart. One heading alone, or a
`Scene:` line of prose, is left.
- an unfenced proposal that starts a line, quoted or not, when it parses and
is a proposal. With no fence it becomes the turn's proposal. A `State`
heading directly above goes with it. Candidates are taken outermost first,
so a finished event line inside an unfinished block is never taken as a
proposal by itself.
- an unfinished unfenced proposal at the end that reads as protocol.
- whatever is left at the end: a `State` heading, a bare `>`, a parroted
reminder or continue hint (closed or not), and a ```json fence cut off
before it names its events. These are cut repeatedly until nothing more
comes off.
This also fixes an older bug. `_STATE_FENCE_RE` read "a ```state block"
inside a parroted reminder as a fence opening and cut out the middle of the
reminder. The label must now end its line or run straight into the payload.
A reply whose only removal is a pasted state section records no raw block,
so the turn is not marked unparseable for a block it never started.
Every AI turn in four real runs was replayed through the new extractor:
draco M01, the two 26-turn GPU trials, and this M01 run. 339 turns in all.
No turn the old extractor had left clean changed. Every leak of our own
protocol is gone: 42 of 42 in this M01 run, 5 in trial 2, 3 on draco.
Trial 1 still has model-invented headings ("Identifiers established:",
"Set of events made true:") on 10 turns. They paraphrase the instruction and
are not our renderer's text, so they are left, not guessed at.
The long-run harness now records an explicit M04 verdict, which is never a
recovery while the clue is still in recent history. It also counts the AI
turns in the export that still carry protocol.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
co-authored by
Claude Opus 5
parent
f8d401029f
commit
0c7316f951
@@ -828,11 +828,13 @@ def main() -> int:
|
||||
# Attempted even for an aborted run: the recovery check and the storage
|
||||
# numbers are worth having at whatever turn count was reached.
|
||||
bundle_bytes = None
|
||||
leaks = None
|
||||
try:
|
||||
bundle = server.call("GET", f"/adventures/{run.adv}/export")
|
||||
(out / "bundle.json").write_text(json.dumps(bundle))
|
||||
bundle_bytes = len((out / "bundle.json").read_bytes())
|
||||
run.note("exported", bytes=bundle_bytes)
|
||||
leaks = _protocol_leaks(bundle)
|
||||
run.note("exported", bytes=bundle_bytes, protocol_leaks=leaks)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300])
|
||||
|
||||
@@ -859,6 +861,7 @@ def main() -> int:
|
||||
"aborted_reason": aborted,
|
||||
"failed_reason": failed_reason,
|
||||
"summaries": _or_none(run.summary_count),
|
||||
"protocol_leaks": leaks,
|
||||
"accepted_turns": run.accepted,
|
||||
"turns_requested": args.turns,
|
||||
"restarts": server.starts - 1,
|
||||
@@ -1127,7 +1130,7 @@ def _recall(run: Run) -> dict:
|
||||
document = run.state()["document"]
|
||||
fact_present = any(
|
||||
CLUE_SENTINEL in json.dumps(f) for f in document.get("facts") or [])
|
||||
return {
|
||||
result = {
|
||||
"clue_in_recent_history_window": in_history,
|
||||
"clue_in_prompt": CLUE_SENTINEL in whole_prompt,
|
||||
"in_state_section": CLUE_SENTINEL in sections.get(STATE_LABEL, ""),
|
||||
@@ -1151,6 +1154,41 @@ def _recall(run: Run) -> dict:
|
||||
(server.call("GET", f"/adventures/{adv}") or {}).get("story_summary")),
|
||||
"prompt_tokens": after["tokens"]["total"],
|
||||
}
|
||||
result["m04_verdict"] = _m04_verdict(result)
|
||||
return result
|
||||
|
||||
|
||||
def _m04_verdict(recall: dict) -> str:
|
||||
"""What the recall check proved, in one word the report can quote.
|
||||
|
||||
The first long run with the bank on found the clue "in the prompt" at turn
|
||||
100, but only because the narrator had pasted the state section into its
|
||||
prose and the paste was still in recent history. A clue in recent history is
|
||||
no evidence about memory, so that case gets its own verdict and never counts
|
||||
as recovery."""
|
||||
if recall["clue_in_recent_history_window"]:
|
||||
return "precondition_not_met"
|
||||
if recall["in_memories_section"] or recall["in_summary_section"]:
|
||||
return "recovered_through_memory_or_summary"
|
||||
if recall["in_state_section"]:
|
||||
return "recovered_through_state_only"
|
||||
return "not_recovered"
|
||||
|
||||
|
||||
#: Signs the application stored protocol as story: the state section's own
|
||||
#: headings, and a proposal's event list. Copied rather than imported from
|
||||
#: `app.narrative.render`, for the reason `HISTORY_LABELS` is copied.
|
||||
PROTOCOL_LEAK_MARKERS = ("Who and what exists:", "\nHeld:\n", "\nEstablished:\n",
|
||||
'"events"')
|
||||
|
||||
|
||||
def _protocol_leaks(bundle: dict) -> dict:
|
||||
"""How many stored AI turns still carry protocol, on every branch."""
|
||||
ai = [a for a in bundle.get("actions") or [] if a.get("type") == "ai"]
|
||||
leaking = [a for a in ai
|
||||
if any(marker in (a.get("text") or "") for marker in PROTOCOL_LEAK_MARKERS)]
|
||||
return {"ai_actions": len(ai), "leaking": len(leaking),
|
||||
"example_ids": [a.get("id") for a in leaking[:10]]}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
Reference in New Issue
Block a user