Keep the state section and its proposal out of the story

The first complete M01 run with the memory bank on (f8d4010, 101 turns on
a GPU host) reported "complete". It still did not prove M04. The planted
clue was found at turn 100 only because the narrator had pasted the
narrative-state section into its own prose, and the paste was still in
recent history. No memory and no summary carried the clue.

The narrator is a small local model. It wrote protocol into its stored
narration on 42 of 104 turns, starting at depth 2, in four shapes:

- a copy of the state section: `Scene:`, `Who and what exists:`, `Held:`,
  `Established:`, `Still open:`
- that copy above a correct ```state block, which was stripped while the
  copy stayed
- the copy, then a bare `State` heading, then a `> {"events": ...}`
  proposal quoted like a player turn, sometimes with story after it
- the same block cut off by the output-token limit, on 10 turns

Stored text is replayed verbatim as history, so each leak also put a second,
older account of the state into the next prompt. That is what M5 review
Finding 4 removed from history replay, and every leak gave the model another
example to copy.

The extractor now removes:

- a pasted state section, recognised by at least two of the renderer's own
  headings as whole lines. The headings are constants in `render.py`, so the
  renderer and the extractor cannot drift apart. One heading alone, or a
  `Scene:` line of prose, is left.
- an unfenced proposal that starts a line, quoted or not, when it parses and
  is a proposal. With no fence it becomes the turn's proposal. A `State`
  heading directly above goes with it. Candidates are taken outermost first,
  so a finished event line inside an unfinished block is never taken as a
  proposal by itself.
- an unfinished unfenced proposal at the end that reads as protocol.
- whatever is left at the end: a `State` heading, a bare `>`, a parroted
  reminder or continue hint (closed or not), and a ```json fence cut off
  before it names its events. These are cut repeatedly until nothing more
  comes off.

This also fixes an older bug. `_STATE_FENCE_RE` read "a ```state block"
inside a parroted reminder as a fence opening and cut out the middle of the
reminder. The label must now end its line or run straight into the payload.

A reply whose only removal is a pasted state section records no raw block,
so the turn is not marked unparseable for a block it never started.

Every AI turn in four real runs was replayed through the new extractor:
draco M01, the two 26-turn GPU trials, and this M01 run. 339 turns in all.
No turn the old extractor had left clean changed. Every leak of our own
protocol is gone: 42 of 42 in this M01 run, 5 in trial 2, 3 on draco.
Trial 1 still has model-invented headings ("Identifiers established:",
"Set of events made true:") on 10 turns. They paraphrase the instruction and
are not our renderer's text, so they are left, not guessed at.

The long-run harness now records an explicit M04 verdict, which is never a
recovery while the clue is still in recent history. It also counts the AI
turns in the export that still carry protocol.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-13 21:18:18 -04:00
co-authored by Claude Opus 5
parent f8d401029f
commit 0c7316f951
5 changed files with 615 additions and 25 deletions
+281 -8
View File
@@ -24,7 +24,7 @@ from __future__ import annotations
import json import json
import re import re
from . import events from . import events, render
# The block the model is asked to append. Built from the vocabulary rather than # The block the model is asked to append. Built from the vocabulary rather than
# written beside it, so the instruction cannot describe an event the application # written beside it, so the instruction cannot describe an event the application
@@ -67,8 +67,13 @@ EMIT_REMINDER = (
# will take. That block must still leave the prose, and must still be recorded, # will take. That block must still leave the prose, and must still be recorded,
# because an unparseable proposal is exactly the failure the audit exists to # because an unparseable proposal is exactly the failure the audit exists to
# make visible. # make visible.
#
# The label must end the fence line or run straight into the payload. Without
# that, "a ```state block" inside a parroted reminder read as a fence opening,
# and everything up to the next fence was cut out of the middle of the reminder
# (M11 long-run trial).
_STATE_FENCE_RE = re.compile( _STATE_FENCE_RE = re.compile(
r"```state[^\S\n]*\n?(.*?)```", re.DOTALL | re.IGNORECASE r"```state[^\S\n]*(?:\n|(?=[\[{]))(.*?)```", re.DOTALL | re.IGNORECASE
) )
# `json` is *not* our label. Models reach for it anyway, so a ```json fence is # `json` is *not* our label. Models reach for it anyway, so a ```json fence is
# taken only when what it contains is actually a proposal. A character who # taken only when what it contains is actually a proposal. A character who
@@ -91,7 +96,9 @@ _TRAILING_RE = re.compile(r"(\{.*\})\s*$", re.DOTALL)
# Our own label ends the story unconditionally. A dangling ```json fence is # Our own label ends the story unconditionally. A dangling ```json fence is
# judged on what follows it, because an unterminated code block in a story is # judged on what follows it, because an unterminated code block in a story is
# still the author's (M5 review, Finding 6). # still the author's (M5 review, Finding 6).
_DANGLING_STATE_RE = re.compile(r"\n?```state\b.*\Z", re.DOTALL | re.IGNORECASE) _DANGLING_STATE_RE = re.compile(
r"\n?```state[^\S\n]*(?:\n|(?=[\[{])|\Z).*\Z", re.DOTALL | re.IGNORECASE
)
_DANGLING_JSON_RE = re.compile(r"\n?```json\b(.*)\Z", re.DOTALL | re.IGNORECASE) _DANGLING_JSON_RE = re.compile(r"\n?```json\b(.*)\Z", re.DOTALL | re.IGNORECASE)
# The reminder, parroted back. Small local models reproduce the bracketed # The reminder, parroted back. Small local models reproduce the bracketed
@@ -104,6 +111,9 @@ _DANGLING_JSON_RE = re.compile(r"\n?```json\b(.*)\Z", re.DOTALL | re.IGNORECASE)
# the echo is the shape of the instruction itself — the fence token, the word it # the echo is the shape of the instruction itself — the fence token, the word it
# opens with, or the pair of phrases the reminder uses together. # opens with, or the pair of phrases the reminder uses together.
_TRAILING_BRACKET_RE = re.compile(r"\n?\[([^\]]*)\]\s*\Z", re.DOTALL) _TRAILING_BRACKET_RE = re.compile(r"\n?\[([^\]]*)\]\s*\Z", re.DOTALL)
# The same echo cut off before its closing bracket, which a reply that runs
# into the output limit leaves at the end.
_UNCLOSED_BRACKET_RE = re.compile(r"\n?\[([^\]\n]*)\Z")
def _is_echoed_instruction(inner: str) -> bool: def _is_echoed_instruction(inner: str) -> bool:
@@ -113,11 +123,28 @@ def _is_echoed_instruction(inner: str) -> bool:
return True return True
if low.lstrip().startswith("reminder:"): if low.lstrip().startswith("reminder:"):
return True return True
# `CHAT_CONTINUE_HINT` in `providers/openai_compatible.py`, which a model
# also parrots back, observed in the M11 long-run trial. Matched by its
# opening words only, because the echo is often cut off before it ends.
if low.lstrip().startswith("continue the story directly"):
return True
# The reminder names both; prose about the protocol rarely names either the # The reminder names both; prose about the protocol rarely names either the
# way the instruction does, and effectively never both. # way the instruction does, and effectively never both.
return "state block" in low and "events list" in low return "state block" in low and "events list" in low
# A heading the model writes above a block it did not fence: `State`, sometimes
# as `State:`, `**State**` or `### State`. It is removed only in two places:
# directly above a proposal that is removed, and as the last line of the reply.
# A line reading "State" in the middle of a story is left alone.
_STATE_HEADING_RE = re.compile(r"^[ \t>*#_]*state[ \t*_:]*$", re.IGNORECASE)
# An unfenced object that starts a line, optionally quoted with `>`, which small
# models copy from the player-turn convention.
_LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
def _clean(prose: str) -> str: def _clean(prose: str) -> str:
"""Removes protocol the block extraction could not, and nothing else. """Removes protocol the block extraction could not, and nothing else.
@@ -126,18 +153,244 @@ def _clean(prose: str) -> str:
instruction into the narration, and the reader would have been shown it. instruction into the narration, and the reader would have been shown it.
Neither case here is hypothetical — both were observed against a real local Neither case here is hypothetical — both were observed against a real local
model. model.
The M11 long run found two more, on 42 of 104 turns. The model pasted a copy
of the narrative-state section into its prose, and it wrote its proposal
unfenced under a bare `State` heading, sometimes quoted, sometimes with more
story after it. Stored text is replayed as history, so every leak also
showed the next prompt a second, older account of the state, which is what
M5 review Finding 4 removed from replayed history.
""" """
cleaned = prose cleaned, _found = _inline_proposals(prose)
bracket = _TRAILING_BRACKET_RE.search(cleaned) cleaned = _strip_echoed_state(cleaned)
# The end of the reply is cut until nothing more comes off, because one kind
# of leftover can hide another. In a real reply, a `State` heading sat above
# a block the model never finished, and a parroted reminder sat above an
# unclosed fence.
while True:
before = cleaned
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
bracket = pattern.search(cleaned)
if bracket is not None and _is_echoed_instruction(bracket.group(1)): if bracket is not None and _is_echoed_instruction(bracket.group(1)):
cleaned = cleaned[: bracket.start()] cleaned = cleaned[: bracket.start()]
cleaned = _DANGLING_STATE_RE.sub("", cleaned) cleaned = _DANGLING_STATE_RE.sub("", cleaned)
dangling = _DANGLING_JSON_RE.search(cleaned) dangling = _DANGLING_JSON_RE.search(cleaned)
if dangling is not None and _reads_as_protocol(dangling.group(1)): if dangling is not None and (_reads_as_protocol(dangling.group(1))
or _is_opening_of_proposal(dangling.group(1))):
cleaned = cleaned[: dangling.start()] cleaned = cleaned[: dangling.start()]
cleaned = _strip_dangling_object(cleaned)
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
# A bare quote marker, the start of a quoted block that never came.
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
if cleaned == before:
return cleaned.strip() return cleaned.strip()
def _is_state_heading(line: str) -> bool:
return bool(_STATE_HEADING_RE.match(line))
def _strip_trailing_state_heading(text: str) -> str:
lines = text.rstrip().split("\n")
if lines and _is_state_heading(lines[-1]):
return "\n".join(lines[:-1])
return text
def _strip_dangling_object(text: str) -> str:
"""Cuts an unfenced proposal the model never finished, and what follows it.
A reply that runs into the output-token limit mid-block ends inside the
object, often a quoted one. That happened on 10 of 104 turns in the M11 long
run. The object never closes, so `_inline_proposals` cannot take it. The
candidate is the outermost object that stays unclosed, not the last line
that opens one. Its finished event objects open lines too, and they close,
so cutting at the last of them left the list above it in the story. It is
cut when it reads as protocol (`_reads_as_protocol`), the same test a
truncated ```json fence has to pass.
"""
skip_until = 0
for match in _LINE_OBJECT_RE.finditer(text):
line_start = match.start()
if line_start < skip_until:
continue
body = "\n".join(_QUOTE_PREFIX_RE.sub("", line, count=1)
for line in text[line_start:].split("\n"))
closing = _object_end(body, body.find("{"))
if closing is None:
return text[:line_start] if _reads_as_protocol(body) else text
# Quote markers came off `body`, so this position is never past the
# real end of the object. A line inside the object that is examined
# anyway closes inside it, and is passed over too.
skip_until = line_start + closing
return text
def _strip_echoed_state(text: str) -> str:
"""Removes a copy of the narrative-state section pasted into the prose.
Judged by the section's own headings (`render.SECTION_HEADINGS`), as whole
lines. A block qualifies only when it carries two of them, or one and the
scene line directly above it. A story may contain a line reading "Held:",
and one heading on its own is left there. The block runs over the headings,
their indented entries and the blank lines between them, and stops at the
first line of ordinary prose.
"""
lines = text.split("\n")
drop = [False] * len(lines)
index = 0
while index < len(lines):
if lines[index].strip() not in render.SECTION_HEADINGS:
index += 1
continue
start = index
above = index - 1
while above >= 0 and not lines[above].strip():
above -= 1
scene = above >= 0 and (
lines[above].strip() == render.HEADING_SCENE
or lines[above].lstrip().startswith(render.HEADING_SCENE + " ")
)
if scene:
start = above
headings: set[str] = set()
end = index
cursor = index
while cursor < len(lines):
line = lines[cursor]
stripped = line.strip()
if stripped in render.SECTION_HEADINGS:
headings.add(stripped)
end = cursor
elif stripped and line[:1] in (" ", "\t"):
end = cursor
elif stripped:
break
cursor += 1
if len(headings) + (1 if scene else 0) >= 2:
for position in range(start, end + 1):
drop[position] = True
index = end + 1
if not any(drop):
return text
kept = "\n".join(line for line, gone in zip(lines, drop) if not gone)
return re.sub(r"\n{3,}", "\n\n", kept)
def _object_end(text: str, start: int) -> int | None:
"""Where the JSON object opening at `start` closes, strings respected."""
depth, in_string, escaped = 0, False, False
for position in range(start, len(text)):
char = text[position]
if in_string:
if escaped:
escaped = False
elif char == "\\":
escaped = True
elif char == '"':
in_string = False
elif char == '"':
in_string = True
elif char == "{":
depth += 1
elif char == "}":
depth -= 1
if depth == 0:
return position + 1
return None
def _inline_proposals(text: str) -> tuple[str, list[tuple[dict, str]]]:
"""Removes unfenced proposals that start a line, and returns them.
A candidate must parse and must be a proposal (`_looks_like_proposal`), the
same bar as a bare trailing object. JSON a character wrote stays where it
is. A quoted candidate is read with its `>` markers taken off, across the
consecutive quoted lines. A candidate with prose after it on its closing
line is not on its own lines, and is left alone. A bare `State` heading
directly above a removed proposal goes with it.
Returns the text without them, and `(parsed, raw)` for each, oldest first.
"""
found: list[tuple[dict, str]] = []
cuts: list[tuple[int, int]] = []
# Candidates are taken outermost first. A line inside an object already
# examined is part of that object, and a proposal's own event lines open
# objects too, so one of them must never be taken as a proposal by itself.
# An object that never closes runs to the end of the text, so everything
# after it is inside it.
skip_until = 0
for match in _LINE_OBJECT_RE.finditer(text):
line_start = match.start()
if line_start < skip_until:
continue
line_end = text.find("\n", line_start)
line_end = len(text) if line_end == -1 else line_end
if _QUOTE_PREFIX_RE.match(text[line_start:line_end]):
# Gather the quoted run, unquote it, and find the object inside.
spans, cursor = [], line_start
while cursor < len(text):
stop = text.find("\n", cursor)
stop = len(text) if stop == -1 else stop
if not _QUOTE_PREFIX_RE.match(text[cursor:stop]):
break
spans.append((cursor, stop))
cursor = stop + 1
body_lines = [_QUOTE_PREFIX_RE.sub("", text[a:b], count=1) for a, b in spans]
body = "\n".join(body_lines)
opening = body.find("{")
closing = _object_end(body, opening)
if closing is None:
break
consumed = body[:closing].count("\n")
region_end = spans[consumed][1]
skip_until = region_end
if body[closing:].split("\n", 1)[0].strip():
continue
raw = body[opening:closing]
else:
opening = match.end() - 1
closing = _object_end(text, opening)
if closing is None:
break
rest = text.find("\n", closing)
rest = len(text) if rest == -1 else rest
skip_until = rest
if text[closing:rest].strip():
continue
raw = text[opening:closing]
region_end = rest
parsed = _tolerant_load(raw)
if not _looks_like_proposal(parsed):
continue
region_start = line_start
before = text[:line_start].rstrip("\n").rstrip()
heading_start = before.rfind("\n") + 1
if before and _is_state_heading(before[heading_start:]):
region_start = heading_start
cuts.append((region_start, region_end))
found.append((parsed, raw))
if not cuts:
return text, found
pieces, cursor = [], 0
for start, end in cuts:
pieces.append(text[cursor:start])
cursor = end
pieces.append(text[cursor:])
return re.sub(r"\n{3,}", "\n\n", "".join(pieces)), found
def _is_opening_of_proposal(tail: str) -> bool:
"""Whether a truncated fence stopped before it could say what it was.
`{` followed by nothing but the start of `"events"`. The output limit cut
one reply there, before `_reads_as_protocol` had anything to go on. A
story's own code block is not that short, and one that is holds nothing to
lose."""
body = tail.strip()
return body.startswith("{") and '"events"'.startswith(body[1:].strip())
def _reads_as_protocol(tail: str) -> bool: def _reads_as_protocol(tail: str) -> bool:
"""Whether a truncated fence was on its way to being a proposal.""" """Whether a truncated fence was on its way to being a proposal."""
if '"events"' in tail: if '"events"' in tail:
@@ -204,12 +457,32 @@ def split(text: str) -> tuple[str, dict | None, str]:
if _looks_like_proposal(parsed): if _looks_like_proposal(parsed):
return _clean(text[: match.start()]), parsed, raw return _clean(text[: match.start()]), parsed, raw
# An unfenced proposal on its own lines but not at the end: quoted, or
# followed by more story. The last one is the turn's proposal, as with
# fences, and every one leaves the prose.
without, found = _inline_proposals(text)
if found:
parsed, raw = found[-1]
return _clean(without), parsed, raw
# No block at all — but the reply may still carry protocol the model wrote # No block at all — but the reply may still carry protocol the model wrote
# as prose, or a fence it never closed. # as prose, or a fence it never closed.
cleaned = _clean(text) cleaned = _clean(text)
if cleaned != text.strip(): whole = text.strip()
return cleaned, None, text.strip()[len(cleaned):].strip() if cleaned == whole:
return cleaned, None, "" return cleaned, None, ""
# What came off is kept for the audit when it was protocol: an unfinished
# block, a parroted reminder, or a fence. A pasted copy of the state section
# is not a proposal, so a reply with nothing else removed records no block.
# That keeps the turn from being marked unparseable for a block it never
# started.
if whole.startswith(cleaned):
removed = whole[len(cleaned):].strip()
keep = (_reads_as_protocol(removed) or "```" in removed
or removed.startswith("["))
return cleaned, None, removed if keep else ""
# Text also came out of the middle, so what was removed is not one suffix.
return cleaned, None, whole if _reads_as_protocol(whole) else ""
def _looks_like_proposal(parsed) -> bool: def _looks_like_proposal(parsed) -> bool:
+24 -7
View File
@@ -23,6 +23,23 @@ PROMPT_FACTS = 30
PROMPT_RELATIONSHIPS = 20 PROMPT_RELATIONSHIPS = 20
PROMPT_THREADS = 12 PROMPT_THREADS = 12
# The headings of `for_prompt`, each a whole line. `extract` recognises a copy of
# this section pasted into a narration by these, so they are named once here and
# the two cannot drift apart. A small local model reproduced the section in its
# prose on 42 of 104 turns in the first M01 run with the memory bank on.
HEADING_SCENE = "Scene:"
HEADING_ENTITIES = "Who and what exists:"
HEADING_HELD = "Held:"
HEADING_FACTS = "Established:"
HEADING_WITHDRAWN = "No longer true — do not treat these as established:"
HEADING_RELATIONSHIPS = "Between them:"
HEADING_THREADS = "Still open:"
#: Every heading except the scene's, which also opens the line it heads.
SECTION_HEADINGS = (
HEADING_ENTITIES, HEADING_HELD, HEADING_FACTS, HEADING_WITHDRAWN,
HEADING_RELATIONSHIPS, HEADING_THREADS,
)
def for_prompt(state) -> str: def for_prompt(state) -> str:
"""The current state as the narrator is shown it. """The current state as the narrator is shown it.
@@ -38,7 +55,7 @@ def for_prompt(state) -> str:
scene = document.get("scene") or {} scene = document.get("scene") or {}
if scene.get("summary") or scene.get("location"): if scene.get("summary") or scene.get("location"):
where = scene.get("location") where = scene.get("location")
head = "Scene: " + str(scene.get("summary") or "").strip() head = f"{HEADING_SCENE} " + str(scene.get("summary") or "").strip()
if where: if where:
head += f" (at {model.entity_name(document, where)})" head += f" (at {model.entity_name(document, where)})"
lines.append(head.strip()) lines.append(head.strip())
@@ -46,14 +63,14 @@ def for_prompt(state) -> str:
entities = document["entities"] entities = document["entities"]
if entities: if entities:
lines.append("") lines.append("")
lines.append("Who and what exists:") lines.append(HEADING_ENTITIES)
for key, entity in entities.items(): for key, entity in entities.items():
lines.append(f" {key}: {_entity_line(document, key, entity)}") lines.append(f" {key}: {_entity_line(document, key, entity)}")
possessions = document["possessions"] possessions = document["possessions"]
if possessions: if possessions:
lines.append("") lines.append("")
lines.append("Held:") lines.append(HEADING_HELD)
for item, owner in sorted(possessions.items()): for item, owner in sorted(possessions.items()):
lines.append( lines.append(
f" {model.entity_name(document, item)} — " f" {model.entity_name(document, item)} — "
@@ -63,7 +80,7 @@ def for_prompt(state) -> str:
facts = model.active_facts(document) facts = model.active_facts(document)
if facts: if facts:
lines.append("") lines.append("")
lines.append("Established:") lines.append(HEADING_FACTS)
for fact in facts[-PROMPT_FACTS:]: for fact in facts[-PROMPT_FACTS:]:
lines.append(f" {_fact_line(document, fact)}") lines.append(f" {_fact_line(document, fact)}")
@@ -74,7 +91,7 @@ def for_prompt(state) -> str:
withdrawn = model.withdrawn_facts(document) withdrawn = model.withdrawn_facts(document)
if withdrawn: if withdrawn:
lines.append("") lines.append("")
lines.append("No longer true — do not treat these as established:") lines.append(HEADING_WITHDRAWN)
for fact in withdrawn[-PROMPT_FACTS:]: for fact in withdrawn[-PROMPT_FACTS:]:
line = f" {_fact_line(document, fact)}" line = f" {_fact_line(document, fact)}"
reason = fact.get("invalidated_reason") reason = fact.get("invalidated_reason")
@@ -85,7 +102,7 @@ def for_prompt(state) -> str:
relationships = model.active_relationships(document) relationships = model.active_relationships(document)
if relationships: if relationships:
lines.append("") lines.append("")
lines.append("Between them:") lines.append(HEADING_RELATIONSHIPS)
for relationship in relationships[-PROMPT_RELATIONSHIPS:]: for relationship in relationships[-PROMPT_RELATIONSHIPS:]:
lines.append( lines.append(
f" {model.entity_name(document, relationship['source'])} " f" {model.entity_name(document, relationship['source'])} "
@@ -96,7 +113,7 @@ def for_prompt(state) -> str:
threads = model.open_threads(document) threads = model.open_threads(document)
if threads: if threads:
lines.append("") lines.append("")
lines.append("Still open:") lines.append(HEADING_THREADS)
for key, thread in list(threads.items())[:PROMPT_THREADS]: for key, thread in list(threads.items())[:PROMPT_THREADS]:
lines.append(f" {key}: {thread.get('title', key)}") lines.append(f" {key}: {thread.get('title', key)}")
+29
View File
@@ -269,3 +269,32 @@ def test_a_run_with_no_summary_or_no_memory_is_not_complete(run_for, tmp_path):
def _with_adv(run): def _with_adv(run):
run.adv = 1 run.adv = 1
return run return run
# ------------------------------------------------------- what M04 actually proved
def test_the_m04_verdict_never_credits_a_clue_still_in_recent_history():
"""The first long run with the bank on found the clue at turn 100 only
because the narrator had pasted the state into recent history."""
base = {"clue_in_recent_history_window": False, "in_memories_section": False,
"in_summary_section": False, "in_state_section": False}
assert lr._m04_verdict({**base, "clue_in_recent_history_window": True,
"in_memories_section": True}) == "precondition_not_met"
assert lr._m04_verdict({**base, "in_memories_section": True}) == \
"recovered_through_memory_or_summary"
assert lr._m04_verdict({**base, "in_summary_section": True}) == \
"recovered_through_memory_or_summary"
assert lr._m04_verdict({**base, "in_state_section": True}) == \
"recovered_through_state_only"
assert lr._m04_verdict(base) == "not_recovered"
def test_protocol_left_in_stored_narration_is_counted():
bundle = {"actions": [
{"id": 1, "type": "do", "text": '> You say {"events": []}'},
{"id": 2, "type": "ai", "text": "The rain eases."},
{"id": 3, "type": "ai", "text": "Beat.\n\nWho and what exists:\n mara: Mara"},
{"id": 4, "type": "ai", "text": 'Beat.\n\n> {"events": []}'},
]}
assert lr._protocol_leaks(bundle) == {
"ai_actions": 3, "leaking": 2, "example_ids": [3, 4]}
+234 -1
View File
@@ -30,7 +30,7 @@ from app.database import Base, SessionLocal, engine, get_db
from app.main import app from app.main import app
from app.narrative import apply as napply from app.narrative import apply as napply
from app.narrative import events as nevents from app.narrative import events as nevents
from app.narrative import extract, model, store, validate from app.narrative import extract, model, render, store, validate
from app.routers import adventures from app.routers import adventures
from fakes import ScriptedProvider, state_block from fakes import ScriptedProvider, state_block
@@ -1558,3 +1558,236 @@ def test_ordinary_prose_is_never_trimmed(reply):
story to satisfy a regex is a worse failure than leaving a stray bracket.""" story to satisfy a regex is a worse failure than leaving a stray bracket."""
prose, _parsed, _raw = extract.split(reply) prose, _parsed, _raw = extract.split(reply)
assert prose == reply assert prose == reply
# ============================ M11: the state section, pasted into the narration
#
# Found by the first M01 long run with the memory bank on. A small local model
# pasted the narrative-state section into its prose on 42 of 104 turns, and
# wrote its proposal unfenced under a bare `State` heading, sometimes quoted and
# sometimes with more story after it. Stored text is replayed as history, so
# each leak also put a second, older account of the state into the next prompt.
# The replies below are cut down from that run's raw output, not imagined.
PASTED_STATE = (
"Scene: Aldric, Mara, and Edrin in The Crooked Lantern, the silver key in "
"Mara’s possession. (at The Crooked Lantern)\n"
"\n"
"Who and what exists:\n"
" aldric: Aldric (character)\n"
" mara: Mara (character)\n"
" silver_key: the silver key (item)\n"
"\n"
"Held:\n"
" the silver key — Mara\n"
"\n"
"Established:\n"
" Aldric knows The Old Abbey the silver key opens the crypt beneath the Old "
"Abbey (SILVER-KEY-CRYPT-OLD-ABBEY) [corrected by the player]"
)
def test_a_pasted_state_section_and_an_unfenced_trailing_block_leave_the_prose():
reply = (
"Edrin looks at them both. “Do you trust me, Mara?”\n\n"
"> Aldric steps forward, placing the silver key in Mara’s hand. “I do.”\n\n"
+ PASTED_STATE + "\n\nState\n"
'{\n "events": [\n {"type": "set_possession", "item": "silver_key", '
'"owner": "mara"}\n ]\n}'
)
prose, parsed, _raw = extract.split(reply)
assert prose == (
"Edrin looks at them both. “Do you trust me, Mara?”\n\n"
"> Aldric steps forward, placing the silver key in Mara’s hand. “I do.”"
)
assert parsed["events"][0]["owner"] == "mara"
def test_a_quoted_block_in_the_middle_of_the_story_is_taken_and_removed():
reply = (
"Aldric steps forward. “We need to be cautious,” he says.\n\n"
+ PASTED_STATE + "\n\nState\n\n"
'> {"events": [{"type": "set_current_location", "entity": "aldric", '
'"location": "ridge_track"}]}\n\n'
"The forest ahead was dense, and the road lower than expected."
)
prose, parsed, raw = extract.split(reply)
assert prose == (
"Aldric steps forward. “We need to be cautious,” he says.\n\n"
"The forest ahead was dense, and the road lower than expected."
)
assert parsed["events"][0]["location"] == "ridge_track"
assert raw.startswith('{"events"'), "the proposal is kept for the audit"
def test_a_pasted_state_section_above_a_proper_block_is_removed_too():
reply = "The rain eases.\n\n" + PASTED_STATE + '\n\n```state\n{"events": []}\n```'
prose, parsed, _raw = extract.split(reply)
assert prose == "The rain eases."
assert parsed == {"events": []}
def test_a_multi_line_quoted_block_is_read_without_its_markers():
reply = (
'Beat.\n\n> {\n> "events": [\n> {"type": "add_fact", '
'"predicate": "the door opened"}\n> ]\n> }\n\nAfter.'
)
prose, parsed, _raw = extract.split(reply)
assert prose == "Beat.\n\nAfter."
assert parsed["events"][0]["predicate"] == "the door opened"
def test_every_section_the_renderer_writes_is_recognised_when_pasted():
"""The extractor finds a paste by the renderer's own headings. A section
added to `render.for_prompt` without a constant would leak again unnoticed,
so this renders every section and pastes the lot."""
document = model.empty()
document["entities"] = {
"mara": {"type": "character", "name": "Mara"},
"key": {"type": "item", "name": "the key"},
}
document["possessions"] = {"key": "mara"}
document["facts"] = [
{"id": "f1", "subject": "mara", "predicate": "knows",
"value": "the door is locked", "status": "active"},
{"id": "f2", "subject": "mara", "predicate": "believed",
"value": "the door was open", "status": "invalidated"},
]
document["relationships"] = [
{"source": "mara", "type": "guards", "target": "key", "status": "active"}]
document["threads"] = {"door": {"title": "Who locked the door", "status": "open"}}
document["scene"] = {"summary": "Mara at the door.", "location": None,
"present": ["mara"]}
rendered = render.for_prompt(document)
lines = rendered.split("\n")
for heading in render.SECTION_HEADINGS:
assert heading in lines, f"{heading!r} is not a line of the rendered section"
prose, _parsed, _raw = extract.split("Mara listens.\n\n" + rendered)
assert prose == "Mara listens."
def test_a_quoted_block_cut_off_by_the_token_limit_is_not_story():
"""10 of 104 turns in the long run ended inside a quoted object."""
reply = (
"The sky above is a canvas of gray, the rain relentless.\n\n"
+ PASTED_STATE + "\n\nState\n\n"
'> {"events": [{"type": "set_entity_status", "entity":'
)
prose, parsed, raw = extract.split(reply)
assert prose == "The sky above is a canvas of gray, the rain relentless."
assert parsed is None
assert raw, "the unfinished block is kept for the audit"
def test_a_block_missing_only_its_last_brace_is_not_story():
reply = (
"Mara’s eyes widen.\n\nState\n\n"
'> {"events": [{"type": "add_fact", "predicate": "Mara understands."}]'
)
prose, _parsed, _raw = extract.split(reply)
assert prose == "Mara’s eyes widen."
def test_an_unfinished_block_is_cut_at_its_outer_brace_not_an_inner_one():
"""The finished event objects close; the block around them never does.
Cutting at the last line that opens an object left the list in the story."""
reply = (
"They brace themselves for what they must face.\n\nState\n{\n"
' "events": [\n'
' { "type": "set_current_location", "entity": "aldric", "location": "docks" },\n'
' { "type": "set_entity_attribute", "entity": "aldric", "attribute": "mood", '
'"value": "tense"\n'
" ]\n}"
)
prose, parsed, _raw = extract.split(reply)
assert prose == "They brace themselves for what they must face."
assert parsed is None
def test_a_finished_event_inside_an_unfinished_block_is_not_taken_on_its_own():
reply = (
'Beat.\n\nState\n{\n "events": [\n'
' {"type": "add_fact", "predicate": "the door opened"}\n'
)
prose, parsed, _raw = extract.split(reply)
assert prose == "Beat."
assert parsed is None, "one event line was taken as the whole proposal"
def test_a_bare_quote_marker_left_at_the_end_is_not_story():
reply = (
"They are ready.\n\n"
"[Reminder: end your reply with a ```state block listing the events your "
"narration made true, with absolute values. Send an empty events list if "
"nothing changed.]\n\n>"
)
prose, _parsed, _raw = extract.split(reply)
assert prose == "They are ready."
def test_a_parroted_reminder_above_an_unfinished_json_fence_is_all_removed():
"""The reminder names "a ```state block". That phrase once read as a fence
opening, and the middle of the reminder was cut out while the rest of it
and the unfinished JSON stayed in the story."""
reply = (
"He turned back towards the town instead.\n\n"
"[Reminder: end your reply with a ```state block listing the events your "
"narration made true, with absolute values. Send an empty events list if "
"nothing changed.]\n\n"
'```json\n{\n "events": [\n {'
)
prose, parsed, _raw = extract.split(reply)
assert prose == "He turned back towards the town instead."
assert parsed is None
def test_a_fence_cut_off_before_it_names_its_events_is_not_story():
reply = (
"The silver key was his only guide.\n\n"
"[Reminder: end your reply with a ```state block listing the events your "
"narration made true, with absolute values. Send an empty events list if "
"nothing changed.]\n\n"
'```json\n{\n "'
)
prose, _parsed, _raw = extract.split(reply)
assert prose == "The silver key was his only guide."
def test_a_parroted_continue_hint_cut_off_mid_sentence_is_not_story():
reply = (
"Aldric's steps were firm.\n\n"
"[Reminder: end your reply with a ```state block listing the events your "
"narration made true, with absolute values. Send an empty events list if "
"nothing changed.]\n\n"
"[Continue the story directly."
)
prose, _parsed, _raw = extract.split(reply)
assert prose == "Aldric's steps were firm."
def test_a_pasted_state_section_with_no_block_records_no_block():
"""A paste is not a proposal, so the turn is not marked unparseable."""
prose, parsed, raw = extract.split("The rain eases.\n\n" + PASTED_STATE)
assert prose == "The rain eases."
assert parsed is None
assert raw == ""
@pytest.mark.parametrize("reply", [
"The notice on the door read:\n\nHeld:\n nothing, by order of the Watch.",
"He chalked the first mark of a sum on the wall:\n{",
"The sign said only [continue at your own risk",
'She typed it out:\n```json\n{"name":',
"Scene: the docks at dawn.\n\nThe gulls cried over the water.",
'She typed:\n> {"name": "Mara"}\nand pressed enter.',
"Aldric read the old charter.\n\nState\n\nof the realm, it began, is fragile.",
"> You open the door.\n\nThe hinges complain.",
])
def test_story_that_only_resembles_the_protocol_is_kept(reply):
"""One heading, a scene line on its own, JSON that is not a proposal, and a
`State` line with story after it are all somebody's story."""
prose, parsed, _raw = extract.split(reply)
assert prose == reply
assert parsed is None
+40 -2
View File
@@ -828,11 +828,13 @@ def main() -> int:
# Attempted even for an aborted run: the recovery check and the storage # Attempted even for an aborted run: the recovery check and the storage
# numbers are worth having at whatever turn count was reached. # numbers are worth having at whatever turn count was reached.
bundle_bytes = None bundle_bytes = None
leaks = None
try: try:
bundle = server.call("GET", f"/adventures/{run.adv}/export") bundle = server.call("GET", f"/adventures/{run.adv}/export")
(out / "bundle.json").write_text(json.dumps(bundle)) (out / "bundle.json").write_text(json.dumps(bundle))
bundle_bytes = len((out / "bundle.json").read_bytes()) bundle_bytes = len((out / "bundle.json").read_bytes())
run.note("exported", bytes=bundle_bytes) leaks = _protocol_leaks(bundle)
run.note("exported", bytes=bundle_bytes, protocol_leaks=leaks)
except Exception as exc: # noqa: BLE001 except Exception as exc: # noqa: BLE001
run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300]) run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300])
@@ -859,6 +861,7 @@ def main() -> int:
"aborted_reason": aborted, "aborted_reason": aborted,
"failed_reason": failed_reason, "failed_reason": failed_reason,
"summaries": _or_none(run.summary_count), "summaries": _or_none(run.summary_count),
"protocol_leaks": leaks,
"accepted_turns": run.accepted, "accepted_turns": run.accepted,
"turns_requested": args.turns, "turns_requested": args.turns,
"restarts": server.starts - 1, "restarts": server.starts - 1,
@@ -1127,7 +1130,7 @@ def _recall(run: Run) -> dict:
document = run.state()["document"] document = run.state()["document"]
fact_present = any( fact_present = any(
CLUE_SENTINEL in json.dumps(f) for f in document.get("facts") or []) CLUE_SENTINEL in json.dumps(f) for f in document.get("facts") or [])
return { result = {
"clue_in_recent_history_window": in_history, "clue_in_recent_history_window": in_history,
"clue_in_prompt": CLUE_SENTINEL in whole_prompt, "clue_in_prompt": CLUE_SENTINEL in whole_prompt,
"in_state_section": CLUE_SENTINEL in sections.get(STATE_LABEL, ""), "in_state_section": CLUE_SENTINEL in sections.get(STATE_LABEL, ""),
@@ -1151,6 +1154,41 @@ def _recall(run: Run) -> dict:
(server.call("GET", f"/adventures/{adv}") or {}).get("story_summary")), (server.call("GET", f"/adventures/{adv}") or {}).get("story_summary")),
"prompt_tokens": after["tokens"]["total"], "prompt_tokens": after["tokens"]["total"],
} }
result["m04_verdict"] = _m04_verdict(result)
return result
def _m04_verdict(recall: dict) -> str:
"""What the recall check proved, in one word the report can quote.
The first long run with the bank on found the clue "in the prompt" at turn
100, but only because the narrator had pasted the state section into its
prose and the paste was still in recent history. A clue in recent history is
no evidence about memory, so that case gets its own verdict and never counts
as recovery."""
if recall["clue_in_recent_history_window"]:
return "precondition_not_met"
if recall["in_memories_section"] or recall["in_summary_section"]:
return "recovered_through_memory_or_summary"
if recall["in_state_section"]:
return "recovered_through_state_only"
return "not_recovered"
#: Signs the application stored protocol as story: the state section's own
#: headings, and a proposal's event list. Copied rather than imported from
#: `app.narrative.render`, for the reason `HISTORY_LABELS` is copied.
PROTOCOL_LEAK_MARKERS = ("Who and what exists:", "\nHeld:\n", "\nEstablished:\n",
'"events"')
def _protocol_leaks(bundle: dict) -> dict:
"""How many stored AI turns still carry protocol, on every branch."""
ai = [a for a in bundle.get("actions") or [] if a.get("type") == "ai"]
leaking = [a for a in ai
if any(marker in (a.get("text") or "") for marker in PROTOCOL_LEAK_MARKERS)]
return {"ai_actions": len(ai), "leaking": len(leaking),
"example_ids": [a.get("id") for a in leaking[:10]]}
if __name__ == "__main__": if __name__ == "__main__":