Files
interactive-story/backend/app/narrative/extract.py
T
JesseMarkowitzandClaude Opus 5 96c1bf5ded Measure M04 by where the planted turn is, and catch a section of the model's own
The M04 re-run on 0c7316f ran 101 turns with no failures, and the verdict
still came out `precondition_not_met`. That verdict was wrong. The planted
player turn (depth 1) was 65 actions outside the history window, whose floor
was 66. The sentinel's text was in recent history for two other reasons.

- On 5 turns the narrator wrote a section of its own, `## Established:` over
  indented facts, with the planted clue copied into it from the state section.
  The extractor passed it: it was a single heading, and the `## ` meant it did
  not match. The harness leak count passed it the same way, and read 1 where
  5 turns leaked.
- On 4 turns the narrator used the sentinel as a name inside ordinary
  sentences ("the SILVER-KEY-CRYPT-OLD-ABBEY, flickers with latent power"). That
  is story text and cannot be stripped.

So the sentinel's text in history can never be the precondition. M04's
written pass is "Fact/event remains recoverable without entire transcript in
prompt" (V1-ACCEPTANCE-TESTS.md). The owner agreed on 2026-09-13 that the
precondition is positional, and that recovery through authoritative state
counts; memory is not required.

Extractor:
- A state-section heading is recognised with any markdown the model wrapped
  it in (`## Established:`, `**Held:**`, `> Held:`).
- One heading with an indented entry under it now qualifies as protocol.
  Before, a block needed two headings, or one heading and the scene line. A
  heading followed by unindented prose is still story. The earlier guard test
  "Held:" with an indented line is now protocol, and the case was rewritten
  unindented.

Harness:
- It records `planted_depth` when the clue is planted, and carries it across
  --resume. No endpoint reports an action's depth, but a fresh campaign's path
  is the opening, the planted turn and its reply, so the depth is
  `total_actions - 2`. That was checked against the database in three runs.
- `_recall` reports `planted_depth`, `history_floor_depth` and
  `planted_turn_in_history_window`. The verdict is `precondition_not_met` only
  when the planted turn is still in the window, and `precondition_unknown`
  when its depth was never recorded. `clue_in_recent_history_window` stays as a
  fact about the prompt.
- The protocol-leak count matches a state heading with an indented entry under
  it, in any markdown.

Every AI turn in five real runs was replayed through the new extractor: draco
M01, the two 26-turn GPU trials, the f8d4010 M01 run and the M04 re-run. 443
turns in all. No turn the old extractor had left clean changed, and the M04
re-run lost 7 more leaks. The sentinel-as-a-name turns are story and remain.
Trial 1's model-invented headings remain, as before.

The M04 re-run's own evidence, reclassified under the new precondition from
its database and its recall-turn snapshot, reads
`recovered_through_state_only`. The original recall.json is kept unchanged
beside `recall-reclassified.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
2026-09-13 22:10:33 -04:00

551 lines
24 KiB
Python

"""M5: getting a typed proposal out of a narration, and keeping it out of the prose.
The model writes the story and, after it, one fenced block of typed events. This
module holds the instruction it is given, the parser that survives the ways a
model gets a format wrong, and the separation that keeps machine-readable output
from reaching the reader.
Two properties matter more than elegance here:
* **The prose must never carry the protocol.** A reader should not see a JSON
block under their story, and a stored narration should not contain one either,
because everything downstream — memory, summaries, export, the transcript —
treats stored text as the story. The block is removed before the text is
stored, not before it is displayed.
* **An unreadable block must not be a failed turn.** A narration the user watched
arrive is worth keeping even when the state block after it is garbage. Parsing
returns "no events" rather than raising, the turn commits with the state
unchanged, and the proposal record keeps the raw output so the failure is
visible in the audit rather than only in a log.
"""
from __future__ import annotations
import json
import re
from . import events, render
# The block the model is asked to append. Built from the vocabulary rather than
# written beside it, so the instruction cannot describe an event the application
# would then reject (`events.vocabulary_for_prompt`).
EMIT_RULE = (
"After your narration, append a fenced code block labelled `state` containing "
"a JSON object with an \"events\" list, recording what your own narration made "
"true. Treat your narration as authoritative: if you wrote that someone moved, "
"took something, learned something, was hurt, or that a new person or place "
"appeared, record it.\n"
"\n"
"Every value is ABSOLUTE — the new state of things, never a change or a "
"difference. Use only these events:\n"
f"{events.vocabulary_for_prompt()}\n"
"\n"
"Identifiers are short lower-case slugs (mara, silver-key, old-abbey) and must "
"match the ones already in the state you were shown. Introduce a person, place "
"or thing with create_entity before referring to it. If the turn established "
"nothing, send an empty events list.\n"
"Example:\n"
'```state\n'
'{"events": [{"type": "set_possession", "item": "silver-key", "owner": "aldric"},'
' {"type": "set_current_location", "entity": "aldric", "location": "old-abbey"}]}\n'
'```'
)
# Placed last, where recency is strongest, the same way the delta protocol did.
EMIT_REMINDER = (
"[Reminder: end your reply with a ```state block listing the events your "
"narration made true, with absolute values. Send an empty events list if "
"nothing changed.]"
)
# Three patterns, and the difference between them is the whole of this module's
# safety. A story is allowed to contain code, and taking a code block out of
# someone's prose is a worse failure than leaving a stray proposal in it.
#
# `state` is the label the application asks for, so a fence carrying it is ours
# whatever is inside it — including a truncated `{oh no` that no JSON parser
# will take. That block must still leave the prose, and must still be recorded,
# because an unparseable proposal is exactly the failure the audit exists to
# make visible.
#
# The label must end the fence line or run straight into the payload. Without
# that, "a ```state block" inside a parroted reminder read as a fence opening,
# and everything up to the next fence was cut out of the middle of the reminder
# (M11 long-run trial).
_STATE_FENCE_RE = re.compile(
r"```state[^\S\n]*(?:\n|(?=[\[{]))(.*?)```", re.DOTALL | re.IGNORECASE
)
# `json` is *not* our label. Models reach for it anyway, so a ```json fence is
# taken only when what it contains is actually a proposal. A character who
# writes `{"name": "Mara"}` into a terminal keeps their code block (M5 review,
# Finding 6).
_JSON_FENCE_RE = re.compile(
r"```json[^\S\n]*\n?(.*?)```", re.DOTALL | re.IGNORECASE
)
# An *unlabelled* fence is ours on the same terms: it has to be a proposal, not
# merely JSON-shaped.
_BARE_FENCE_RE = re.compile(r"```\s*([\[{].*?[\]}])\s*```", re.DOTALL)
# A bare object hugging the end of the text, for a model that forgets the fence.
_TRAILING_RE = re.compile(r"(\{.*\})\s*$", re.DOTALL)
# An opener with no closing fence. A model that runs out of output tokens
# mid-block leaves one of these, and everything after it is protocol rather than
# story — so the story ends where the opener begins.
#
# Our own label ends the story unconditionally. A dangling ```json fence is
# judged on what follows it, because an unterminated code block in a story is
# still the author's (M5 review, Finding 6).
_DANGLING_STATE_RE = re.compile(
r"\n?```state[^\S\n]*(?:\n|(?=[\[{])|\Z).*\Z", re.DOTALL | re.IGNORECASE
)
_DANGLING_JSON_RE = re.compile(r"\n?```json\b(.*)\Z", re.DOTALL | re.IGNORECASE)
# The reminder, parroted back. Small local models reproduce the bracketed
# instruction they were given, and it arrives as ordinary prose — no fence, so
# nothing above strips it, and the reader is shown a piece of the prompt.
#
# The bracket is *found* broadly and *judged* narrowly. Merely naming the
# protocol is not enough: a story may end on an aside about a state block, and
# deleting that sentence is the worse failure (M5 review, Finding 6). What marks
# the echo is the shape of the instruction itself — the fence token, the word it
# opens with, or the pair of phrases the reminder uses together.
_TRAILING_BRACKET_RE = re.compile(r"\n?\[([^\]]*)\]\s*\Z", re.DOTALL)
# The same echo cut off before its closing bracket, which a reply that runs
# into the output limit leaves at the end.
_UNCLOSED_BRACKET_RE = re.compile(r"\n?\[([^\]\n]*)\Z")
def _is_echoed_instruction(inner: str) -> bool:
"""Whether a trailing bracketed segment is the prompt's own reminder."""
low = inner.lower()
if "```state" in low:
return True
if low.lstrip().startswith("reminder:"):
return True
# `CHAT_CONTINUE_HINT` in `providers/openai_compatible.py`, which a model
# also parrots back, observed in the M11 long-run trial. Matched by its
# opening words only, because the echo is often cut off before it ends.
if low.lstrip().startswith("continue the story directly"):
return True
# The reminder names both; prose about the protocol rarely names either the
# way the instruction does, and effectively never both.
return "state block" in low and "events list" in low
# A heading the model writes above a block it did not fence: `State`, sometimes
# as `State:`, `**State**` or `### State`. It is removed only in two places:
# directly above a proposal that is removed, and as the last line of the reply.
# A line reading "State" in the middle of a story is left alone.
_STATE_HEADING_RE = re.compile(r"^[ \t>*#_]*state[ \t*_:]*$", re.IGNORECASE)
# An unfenced object that starts a line, optionally quoted with `>`, which small
# models copy from the player-turn convention.
_LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
def _clean(prose: str) -> str:
"""Removes protocol the block extraction could not, and nothing else.
Found by the M5 realistic-context run (§12), which is the failure class
Phase 0B warned about: under a full prompt the model echoed its own
instruction into the narration, and the reader would have been shown it.
Neither case here is hypothetical — both were observed against a real local
model.
The M11 long run found two more, on 42 of 104 turns. The model pasted a copy
of the narrative-state section into its prose, and it wrote its proposal
unfenced under a bare `State` heading, sometimes quoted, sometimes with more
story after it. Stored text is replayed as history, so every leak also
showed the next prompt a second, older account of the state, which is what
M5 review Finding 4 removed from replayed history.
"""
cleaned, _found = _inline_proposals(prose)
cleaned = _strip_echoed_state(cleaned)
# The end of the reply is cut until nothing more comes off, because one kind
# of leftover can hide another. In a real reply, a `State` heading sat above
# a block the model never finished, and a parroted reminder sat above an
# unclosed fence.
while True:
before = cleaned
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
bracket = pattern.search(cleaned)
if bracket is not None and _is_echoed_instruction(bracket.group(1)):
cleaned = cleaned[: bracket.start()]
cleaned = _DANGLING_STATE_RE.sub("", cleaned)
dangling = _DANGLING_JSON_RE.search(cleaned)
if dangling is not None and (_reads_as_protocol(dangling.group(1))
or _is_opening_of_proposal(dangling.group(1))):
cleaned = cleaned[: dangling.start()]
cleaned = _strip_dangling_object(cleaned)
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
# A bare quote marker, the start of a quoted block that never came.
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
if cleaned == before:
return cleaned.strip()
def _is_state_heading(line: str) -> bool:
return bool(_STATE_HEADING_RE.match(line))
def _strip_trailing_state_heading(text: str) -> str:
lines = text.rstrip().split("\n")
if lines and _is_state_heading(lines[-1]):
return "\n".join(lines[:-1])
return text
def _strip_dangling_object(text: str) -> str:
"""Cuts an unfenced proposal the model never finished, and what follows it.
A reply that runs into the output-token limit mid-block ends inside the
object, often a quoted one. That happened on 10 of 104 turns in the M11 long
run. The object never closes, so `_inline_proposals` cannot take it. The
candidate is the outermost object that stays unclosed, not the last line
that opens one. Its finished event objects open lines too, and they close,
so cutting at the last of them left the list above it in the story. It is
cut when it reads as protocol (`_reads_as_protocol`), the same test a
truncated ```json fence has to pass.
"""
skip_until = 0
for match in _LINE_OBJECT_RE.finditer(text):
line_start = match.start()
if line_start < skip_until:
continue
body = "\n".join(_QUOTE_PREFIX_RE.sub("", line, count=1)
for line in text[line_start:].split("\n"))
closing = _object_end(body, body.find("{"))
if closing is None:
return text[:line_start] if _reads_as_protocol(body) else text
# Quote markers came off `body`, so this position is never past the
# real end of the object. A line inside the object that is examined
# anyway closes inside it, and is passed over too.
skip_until = line_start + closing
return text
# The markdown a model wraps a heading in: `## Established:`, `**Held:**`,
# `> Held:`.
_HEADING_DECORATION_RE = re.compile(r"^[\s#>*_]+|[\s*_]+$")
def _section_heading(line: str) -> str | None:
"""The state-section heading this line is, markdown aside, or None."""
bare = _HEADING_DECORATION_RE.sub("", line)
return bare if bare in render.SECTION_HEADINGS else None
def _strip_echoed_state(text: str) -> str:
"""Removes a copy of the narrative-state section pasted into the prose.
Judged by the section's own headings (`render.SECTION_HEADINGS`) as whole
lines, with any markdown the model wrapped them in taken off. A block
qualifies when it carries two headings, or one and the scene line directly
above it, or one heading with an indented entry under it. That last case
is the model writing a section of its own: the M04 re-run found
`## Established:` over two indented facts on 5 turns, one of them copying
the planted clue out of the state section. A lone "Held:" with prose after
it is still somebody's story. The block runs over the headings, their
indented entries and the blank lines between them, and stops at the first
line of ordinary prose.
"""
lines = text.split("\n")
drop = [False] * len(lines)
index = 0
while index < len(lines):
if _section_heading(lines[index]) is None:
index += 1
continue
start = index
above = index - 1
while above >= 0 and not lines[above].strip():
above -= 1
scene = above >= 0 and (
lines[above].strip() == render.HEADING_SCENE
or lines[above].lstrip().startswith(render.HEADING_SCENE + " ")
)
if scene:
start = above
headings: set[str] = set()
entries = 0
end = index
cursor = index
while cursor < len(lines):
line = lines[cursor]
stripped = line.strip()
heading = _section_heading(line)
if heading is not None:
headings.add(heading)
end = cursor
elif stripped and line[:1] in (" ", "\t"):
entries += 1
end = cursor
elif stripped:
break
cursor += 1
if len(headings) + (1 if scene else 0) >= 2 or (headings and entries):
for position in range(start, end + 1):
drop[position] = True
index = end + 1
if not any(drop):
return text
kept = "\n".join(line for line, gone in zip(lines, drop) if not gone)
return re.sub(r"\n{3,}", "\n\n", kept)
def _object_end(text: str, start: int) -> int | None:
"""Where the JSON object opening at `start` closes, strings respected."""
depth, in_string, escaped = 0, False, False
for position in range(start, len(text)):
char = text[position]
if in_string:
if escaped:
escaped = False
elif char == "\\":
escaped = True
elif char == '"':
in_string = False
elif char == '"':
in_string = True
elif char == "{":
depth += 1
elif char == "}":
depth -= 1
if depth == 0:
return position + 1
return None
def _inline_proposals(text: str) -> tuple[str, list[tuple[dict, str]]]:
"""Removes unfenced proposals that start a line, and returns them.
A candidate must parse and must be a proposal (`_looks_like_proposal`), the
same bar as a bare trailing object. JSON a character wrote stays where it
is. A quoted candidate is read with its `>` markers taken off, across the
consecutive quoted lines. A candidate with prose after it on its closing
line is not on its own lines, and is left alone. A bare `State` heading
directly above a removed proposal goes with it.
Returns the text without them, and `(parsed, raw)` for each, oldest first.
"""
found: list[tuple[dict, str]] = []
cuts: list[tuple[int, int]] = []
# Candidates are taken outermost first. A line inside an object already
# examined is part of that object, and a proposal's own event lines open
# objects too, so one of them must never be taken as a proposal by itself.
# An object that never closes runs to the end of the text, so everything
# after it is inside it.
skip_until = 0
for match in _LINE_OBJECT_RE.finditer(text):
line_start = match.start()
if line_start < skip_until:
continue
line_end = text.find("\n", line_start)
line_end = len(text) if line_end == -1 else line_end
if _QUOTE_PREFIX_RE.match(text[line_start:line_end]):
# Gather the quoted run, unquote it, and find the object inside.
spans, cursor = [], line_start
while cursor < len(text):
stop = text.find("\n", cursor)
stop = len(text) if stop == -1 else stop
if not _QUOTE_PREFIX_RE.match(text[cursor:stop]):
break
spans.append((cursor, stop))
cursor = stop + 1
body_lines = [_QUOTE_PREFIX_RE.sub("", text[a:b], count=1) for a, b in spans]
body = "\n".join(body_lines)
opening = body.find("{")
closing = _object_end(body, opening)
if closing is None:
break
consumed = body[:closing].count("\n")
region_end = spans[consumed][1]
skip_until = region_end
if body[closing:].split("\n", 1)[0].strip():
continue
raw = body[opening:closing]
else:
opening = match.end() - 1
closing = _object_end(text, opening)
if closing is None:
break
rest = text.find("\n", closing)
rest = len(text) if rest == -1 else rest
skip_until = rest
if text[closing:rest].strip():
continue
raw = text[opening:closing]
region_end = rest
parsed = _tolerant_load(raw)
if not _looks_like_proposal(parsed):
continue
region_start = line_start
before = text[:line_start].rstrip("\n").rstrip()
heading_start = before.rfind("\n") + 1
if before and _is_state_heading(before[heading_start:]):
region_start = heading_start
cuts.append((region_start, region_end))
found.append((parsed, raw))
if not cuts:
return text, found
pieces, cursor = [], 0
for start, end in cuts:
pieces.append(text[cursor:start])
cursor = end
pieces.append(text[cursor:])
return re.sub(r"\n{3,}", "\n\n", "".join(pieces)), found
def _is_opening_of_proposal(tail: str) -> bool:
"""Whether a truncated fence stopped before it could say what it was.
`{` followed by nothing but the start of `"events"`. The output limit cut
one reply there, before `_reads_as_protocol` had anything to go on. A
story's own code block is not that short, and one that is holds nothing to
lose."""
body = tail.strip()
return body.startswith("{") and '"events"'.startswith(body[1:].strip())
def _reads_as_protocol(tail: str) -> bool:
"""Whether a truncated fence was on its way to being a proposal."""
if '"events"' in tail:
return True
return any(f'"{name}"' in tail for name in events.SPECS)
def _tolerant_load(blob: str):
"""Parses a block, forgiving what small local models get wrong.
Trailing commas and a leading `+` on a number are both common and both
rejected by strict JSON. Repairing them is not guessing at meaning — the
intended value is unambiguous — which is the line this function stays on the
right side of. Anything it cannot parse returns None, and the caller treats
that as no proposal rather than as an empty one.
"""
cleaned = re.sub(r",(\s*[}\]])", r"\1", blob)
cleaned = re.sub(r"(:\s*)\+(\d)", r"\1\2", cleaned)
try:
parsed = json.loads(cleaned)
except (json.JSONDecodeError, ValueError):
return None
return parsed
def split(text: str) -> tuple[str, dict | None, str]:
"""Separates a reply into `(prose, proposal, raw_block)`.
`proposal` is None when there is no block or it cannot be parsed at all,
which the caller records as a malformed proposal. `raw_block` is what the
model actually wrote, kept for the audit record even — especially — when it
did not parse.
A bare trailing object is only stripped when it parses *and* looks like a
proposal. Prose that happens to end in a brace is left alone, because
removing a sentence from someone's story to satisfy a regex is a worse
failure than leaving a stray brace in it.
"""
matches = list(_STATE_FENCE_RE.finditer(text))
if matches:
match = matches[-1]
raw = match.group(1).strip()
prose = _clean(text[: match.start()] + text[match.end():])
return prose, _tolerant_load(raw), raw
# A `json` or unlabelled fence is ours only when its contents are this
# protocol. That is judged two ways, and it needs both: a block that parses
# into a proposal, or one that plainly reads as protocol even though it does
# not parse. The second half matters — a small model that mangles its own
# JSON must not have the wreckage shown to the reader, which is what the
# realistic-model run caught during the corrective pass.
for pattern in (_JSON_FENCE_RE, _BARE_FENCE_RE):
for match in reversed(list(pattern.finditer(text))):
raw = match.group(1).strip()
parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
prose = _clean(text[: match.start()] + text[match.end():])
return prose, parsed, raw
match = _TRAILING_RE.search(text)
if match:
raw = match.group(1)
parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed):
return _clean(text[: match.start()]), parsed, raw
# An unfenced proposal on its own lines but not at the end: quoted, or
# followed by more story. The last one is the turn's proposal, as with
# fences, and every one leaves the prose.
without, found = _inline_proposals(text)
if found:
parsed, raw = found[-1]
return _clean(without), parsed, raw
# No block at all — but the reply may still carry protocol the model wrote
# as prose, or a fence it never closed.
cleaned = _clean(text)
whole = text.strip()
if cleaned == whole:
return cleaned, None, ""
# What came off is kept for the audit when it was protocol: an unfinished
# block, a parroted reminder, or a fence. A pasted copy of the state section
# is not a proposal, so a reply with nothing else removed records no block.
# That keeps the turn from being marked unparseable for a block it never
# started.
if whole.startswith(cleaned):
removed = whole[len(cleaned):].strip()
keep = (_reads_as_protocol(removed) or "```" in removed
or removed.startswith("["))
return cleaned, None, removed if keep else ""
# Text also came out of the middle, so what was removed is not one suffix.
return cleaned, None, whole if _reads_as_protocol(whole) else ""
def _looks_like_proposal(parsed) -> bool:
"""Whether a bare trailing object is this protocol rather than prose."""
if not isinstance(parsed, dict):
return False
if isinstance(parsed.get("events"), list):
return True
return isinstance(parsed.get("type"), str) and events.is_allowed(parsed["type"])
def render_block(accepted: list[dict]) -> str:
"""Renders accepted events back into the block the model emitted.
Replayed into the prompt for past turns so the model copies the format it is
being asked for. **Accepted** events rather than proposed ones, for the
reason the delta protocol learned the hard way: showing the model a refused
event standing as though it had worked, contradicted by the state in the
same prompt, teaches it to send the event again.
"""
if not accepted:
return ""
return "```state\n" + json.dumps({"events": accepted}, ensure_ascii=False) + "\n```"
def render_rejections(rejected: list[dict]) -> str:
"""The correction note appended after the most recent AI turn.
Only what was lost. A model that is told what it got wrong can fix it next
turn; a model told nothing repeats it.
"""
if not rejected:
return ""
lines = []
for entry in rejected[:6]:
if not isinstance(entry, dict):
continue
detail = entry.get("detail") or entry.get("reason") or ""
if detail:
lines.append(f"- {detail}")
if not lines:
return ""
body = "\n".join(lines)
return (
"[Part of your last state block was not accepted. Correct it in this "
f"turn's block:\n{body}]"
)