Keep the state section and its proposal out of the story

The first complete M01 run with the memory bank on (f8d4010, 101 turns on
a GPU host) reported "complete". It still did not prove M04. The planted
clue was found at turn 100 only because the narrator had pasted the
narrative-state section into its own prose, and the paste was still in
recent history. No memory and no summary carried the clue.

The narrator is a small local model. It wrote protocol into its stored
narration on 42 of 104 turns, starting at depth 2, in four shapes:

- a copy of the state section: `Scene:`, `Who and what exists:`, `Held:`,
  `Established:`, `Still open:`
- that copy above a correct ```state block, which was stripped while the
  copy stayed
- the copy, then a bare `State` heading, then a `> {"events": ...}`
  proposal quoted like a player turn, sometimes with story after it
- the same block cut off by the output-token limit, on 10 turns

Stored text is replayed verbatim as history, so each leak also put a second,
older account of the state into the next prompt. That is what M5 review
Finding 4 removed from history replay, and every leak gave the model another
example to copy.

The extractor now removes:

- a pasted state section, recognised by at least two of the renderer's own
  headings as whole lines. The headings are constants in `render.py`, so the
  renderer and the extractor cannot drift apart. One heading alone, or a
  `Scene:` line of prose, is left.
- an unfenced proposal that starts a line, quoted or not, when it parses and
  is a proposal. With no fence it becomes the turn's proposal. A `State`
  heading directly above goes with it. Candidates are taken outermost first,
  so a finished event line inside an unfinished block is never taken as a
  proposal by itself.
- an unfinished unfenced proposal at the end that reads as protocol.
- whatever is left at the end: a `State` heading, a bare `>`, a parroted
  reminder or continue hint (closed or not), and a ```json fence cut off
  before it names its events. These are cut repeatedly until nothing more
  comes off.

This also fixes an older bug. `_STATE_FENCE_RE` read "a ```state block"
inside a parroted reminder as a fence opening and cut out the middle of the
reminder. The label must now end its line or run straight into the payload.

A reply whose only removal is a pasted state section records no raw block,
so the turn is not marked unparseable for a block it never started.

Every AI turn in four real runs was replayed through the new extractor:
draco M01, the two 26-turn GPU trials, and this M01 run. 339 turns in all.
No turn the old extractor had left clean changed. Every leak of our own
protocol is gone: 42 of 42 in this M01 run, 5 in trial 2, 3 on draco.
Trial 1 still has model-invented headings ("Identifiers established:",
"Set of events made true:") on 10 turns. They paraphrase the instruction and
are not our renderer's text, so they are left, not guessed at.

The long-run harness now records an explicit M04 verdict, which is never a
recovery while the clue is still in recent history. It also counts the AI
turns in the export that still carry protocol.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-13 21:18:18 -04:00
co-authored by Claude Opus 5
parent f8d401029f
commit 0c7316f951
5 changed files with 615 additions and 25 deletions
+24 -7
View File
@@ -23,6 +23,23 @@ PROMPT_FACTS = 30
PROMPT_RELATIONSHIPS = 20
PROMPT_THREADS = 12
# The headings of `for_prompt`, each a whole line. `extract` recognises a copy of
# this section pasted into a narration by these, so they are named once here and
# the two cannot drift apart. A small local model reproduced the section in its
# prose on 42 of 104 turns in the first M01 run with the memory bank on.
HEADING_SCENE = "Scene:"
HEADING_ENTITIES = "Who and what exists:"
HEADING_HELD = "Held:"
HEADING_FACTS = "Established:"
HEADING_WITHDRAWN = "No longer true — do not treat these as established:"
HEADING_RELATIONSHIPS = "Between them:"
HEADING_THREADS = "Still open:"
#: Every heading except the scene's, which also opens the line it heads.
SECTION_HEADINGS = (
HEADING_ENTITIES, HEADING_HELD, HEADING_FACTS, HEADING_WITHDRAWN,
HEADING_RELATIONSHIPS, HEADING_THREADS,
)
def for_prompt(state) -> str:
"""The current state as the narrator is shown it.
@@ -38,7 +55,7 @@ def for_prompt(state) -> str:
scene = document.get("scene") or {}
if scene.get("summary") or scene.get("location"):
where = scene.get("location")
head = "Scene: " + str(scene.get("summary") or "").strip()
head = f"{HEADING_SCENE} " + str(scene.get("summary") or "").strip()
if where:
head += f" (at {model.entity_name(document, where)})"
lines.append(head.strip())
@@ -46,14 +63,14 @@ def for_prompt(state) -> str:
entities = document["entities"]
if entities:
lines.append("")
lines.append("Who and what exists:")
lines.append(HEADING_ENTITIES)
for key, entity in entities.items():
lines.append(f" {key}: {_entity_line(document, key, entity)}")
possessions = document["possessions"]
if possessions:
lines.append("")
lines.append("Held:")
lines.append(HEADING_HELD)
for item, owner in sorted(possessions.items()):
lines.append(
f" {model.entity_name(document, item)} — "
@@ -63,7 +80,7 @@ def for_prompt(state) -> str:
facts = model.active_facts(document)
if facts:
lines.append("")
lines.append("Established:")
lines.append(HEADING_FACTS)
for fact in facts[-PROMPT_FACTS:]:
lines.append(f" {_fact_line(document, fact)}")
@@ -74,7 +91,7 @@ def for_prompt(state) -> str:
withdrawn = model.withdrawn_facts(document)
if withdrawn:
lines.append("")
lines.append("No longer true — do not treat these as established:")
lines.append(HEADING_WITHDRAWN)
for fact in withdrawn[-PROMPT_FACTS:]:
line = f" {_fact_line(document, fact)}"
reason = fact.get("invalidated_reason")
@@ -85,7 +102,7 @@ def for_prompt(state) -> str:
relationships = model.active_relationships(document)
if relationships:
lines.append("")
lines.append("Between them:")
lines.append(HEADING_RELATIONSHIPS)
for relationship in relationships[-PROMPT_RELATIONSHIPS:]:
lines.append(
f" {model.entity_name(document, relationship['source'])} "
@@ -96,7 +113,7 @@ def for_prompt(state) -> str:
threads = model.open_threads(document)
if threads:
lines.append("")
lines.append("Still open:")
lines.append(HEADING_THREADS)
for key, thread in list(threads.items())[:PROMPT_THREADS]:
lines.append(f" {key}: {thread.get('title', key)}")