WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
725 lines
31 KiB
Python
725 lines
31 KiB
Python
"""M5: getting a typed proposal out of a narration, and keeping it out of the prose.
|
|
|
|
The model writes the story and, after it, one fenced block of typed events. This
|
|
module holds the instruction it is given, the parser that survives the ways a
|
|
model gets a format wrong, and the separation that keeps machine-readable output
|
|
from reaching the reader.
|
|
|
|
Two properties matter more than elegance here:
|
|
|
|
* **The prose must never carry the protocol.** A reader should not see a JSON
|
|
block under their story, and a stored narration should not contain one either,
|
|
because everything downstream — memory, summaries, export, the transcript —
|
|
treats stored text as the story. The block is removed before the text is
|
|
stored, not before it is displayed.
|
|
* **An unreadable block must not be a failed turn.** A narration the user watched
|
|
arrive is worth keeping even when the state block after it is garbage. Parsing
|
|
returns "no events" rather than raising, the turn commits with the state
|
|
unchanged, and the proposal record keeps the raw output so the failure is
|
|
visible in the audit rather than only in a log.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import re
|
|
|
|
from . import events, render
|
|
|
|
# The block the model is asked to append. Built from the vocabulary rather than
|
|
# written beside it, so the instruction cannot describe an event the application
|
|
# would then reject (`events.vocabulary_for_prompt`).
|
|
EMIT_RULE = (
|
|
"After your narration, append a fenced code block labelled `state` containing "
|
|
"a JSON object with an \"events\" list, recording what your own narration made "
|
|
"true. Treat your narration as authoritative: if you wrote that someone moved, "
|
|
"took something, learned something, was hurt, or that a new person or place "
|
|
"appeared, record it.\n"
|
|
"\n"
|
|
"Every value is ABSOLUTE — the new state of things, never a change or a "
|
|
"difference. Use only these events, in exactly this shape:\n"
|
|
f"{events.vocabulary_for_prompt()}\n"
|
|
"\n"
|
|
"Identifiers are short lower-case slugs and must match the ones already in the "
|
|
"state you were shown; the example's identifiers are placeholders. Introduce a "
|
|
"person, place or thing with create_entity before referring to it. If the turn "
|
|
"established nothing, send an empty events list.\n"
|
|
"Example:\n"
|
|
'```state\n'
|
|
'{"events": [{"type": "set_possession", "item": "item-1", "owner": "character-1"},'
|
|
' {"type": "set_current_location", "entity": "character-1", "location": "location-1"}]}\n'
|
|
'```'
|
|
)
|
|
|
|
# Placed last, where recency is strongest, the same way the delta protocol did.
|
|
EMIT_REMINDER = (
|
|
"[Reminder: end your reply with a ```state block listing the events your "
|
|
"narration made true, with absolute values. Send an empty events list if "
|
|
"nothing changed.]"
|
|
)
|
|
|
|
# v1.1 WP-A2: the length hint's own words, named once. `builder.length_hint`
|
|
# builds the hint from these, and the extractor recognises an echo of it by
|
|
# them, so the two cannot drift apart.
|
|
LENGTH_HINT_OPENING = "[Hard limit:"
|
|
LENGTH_HINT_TAIL = "Finish the narration and append the state block well inside the limit."
|
|
#: The application's wording inside a hint. A 3B narrator reworded the front
|
|
#: ("your next turn") and the end ("This story ends here."), and kept one or the
|
|
#: other of these every time.
|
|
_LENGTH_HINT_PHRASE_RE = re.compile(
|
|
r"append the state block|turn must not exceed \d+ words", re.IGNORECASE
|
|
)
|
|
|
|
#: v1.1 WP-A2: the rules that remove protocol a narrator copied, named so the
|
|
#: replay tool and the report can say which removed what.
|
|
RULE_EVENT_CALL = "event_call_line"
|
|
RULE_LENGTH_HINT = "echoed_length_hint"
|
|
RULE_SCENE_LINE = "rendered_scene_line"
|
|
RULE_EMPTY_FENCE = "empty_dangling_fence"
|
|
RULE_INSTRUCTION_TAIL = "echoed_instruction_tail"
|
|
|
|
#: v1.1 WP-A2 corrective (R5). The sentence `CHAT_CONTINUE_HINT` in
|
|
#: `providers/openai_compatible.py` carries, which a narrator echoed with the rest
|
|
#: of the hint reworded around it. Kept as a copy rather than an import, so the
|
|
#: narrative package does not depend on the provider; a test pins that the
|
|
#: hint still contains it.
|
|
CONTINUE_HINT_PHRASE = "Output only story text"
|
|
|
|
# R1. A whole line opening with a call to an event this protocol has. The names
|
|
# come from the vocabulary, so a call-shaped line naming anything else — a
|
|
# character's `open_door(north)` — is not matched.
|
|
_EVENT_CALL_LINE_RE = re.compile(
|
|
r"^[ \t]*(?:>[ \t]*)?(?:"
|
|
+ "|".join(re.escape(name) for name in events.SPECS)
|
|
+ r")[ \t]*\(",
|
|
re.IGNORECASE,
|
|
)
|
|
# R3. The renderer's scene line carries its location this way.
|
|
_RENDERED_SCENE_LOCATION_RE = re.compile(r"\(at [^()\n]+\)\s*$")
|
|
# R4. An opener with nothing after it.
|
|
_EMPTY_FENCE_LINE_RE = re.compile(r"```(?:json)?[ \t]*", re.IGNORECASE)
|
|
|
|
# Three patterns, and the difference between them is the whole of this module's
|
|
# safety. A story is allowed to contain code, and taking a code block out of
|
|
# someone's prose is a worse failure than leaving a stray proposal in it.
|
|
#
|
|
# `state` is the label the application asks for, so a fence carrying it is ours
|
|
# whatever is inside it — including a truncated `{oh no` that no JSON parser
|
|
# will take. That block must still leave the prose, and must still be recorded,
|
|
# because an unparseable proposal is exactly the failure the audit exists to
|
|
# make visible.
|
|
#
|
|
# The label must end the fence line or run straight into the payload. Without
|
|
# that, "a ```state block" inside a parroted reminder read as a fence opening,
|
|
# and everything up to the next fence was cut out of the middle of the reminder
|
|
# (M11 long-run trial).
|
|
_STATE_FENCE_RE = re.compile(
|
|
r"```state[^\S\n]*(?:\n|(?=[\[{]))(.*?)```", re.DOTALL | re.IGNORECASE
|
|
)
|
|
# `json` is *not* our label. Models reach for it anyway, so a ```json fence is
|
|
# taken only when what it contains is actually a proposal. A character who
|
|
# writes `{"name": "Mara"}` into a terminal keeps their code block (M5 review,
|
|
# Finding 6).
|
|
_JSON_FENCE_RE = re.compile(
|
|
r"```json[^\S\n]*\n?(.*?)```", re.DOTALL | re.IGNORECASE
|
|
)
|
|
# An *unlabelled* fence is ours on the same terms: it has to be a proposal, not
|
|
# merely JSON-shaped.
|
|
_BARE_FENCE_RE = re.compile(r"```\s*([\[{].*?[\]}])\s*```", re.DOTALL)
|
|
|
|
# A bare object hugging the end of the text, for a model that forgets the fence.
|
|
_TRAILING_RE = re.compile(r"(\{.*\})\s*$", re.DOTALL)
|
|
|
|
# An opener with no closing fence. A model that runs out of output tokens
|
|
# mid-block leaves one of these, and everything after it is protocol rather than
|
|
# story — so the story ends where the opener begins.
|
|
#
|
|
# Our own label ends the story unconditionally. A dangling ```json fence is
|
|
# judged on what follows it, because an unterminated code block in a story is
|
|
# still the author's (M5 review, Finding 6).
|
|
_DANGLING_STATE_RE = re.compile(
|
|
r"\n?```state[^\S\n]*(?:\n|(?=[\[{])|\Z).*\Z", re.DOTALL | re.IGNORECASE
|
|
)
|
|
_DANGLING_JSON_RE = re.compile(r"\n?```json\b(.*)\Z", re.DOTALL | re.IGNORECASE)
|
|
|
|
# The reminder, parroted back. Small local models reproduce the bracketed
|
|
# instruction they were given, and it arrives as ordinary prose — no fence, so
|
|
# nothing above strips it, and the reader is shown a piece of the prompt.
|
|
#
|
|
# The bracket is *found* broadly and *judged* narrowly. Merely naming the
|
|
# protocol is not enough: a story may end on an aside about a state block, and
|
|
# deleting that sentence is the worse failure (M5 review, Finding 6). What marks
|
|
# the echo is the shape of the instruction itself — the fence token, the word it
|
|
# opens with, or the pair of phrases the reminder uses together.
|
|
_TRAILING_BRACKET_RE = re.compile(r"\n?\[([^\]]*)\]\s*\Z", re.DOTALL)
|
|
# The same echo cut off before its closing bracket, which a reply that runs
|
|
# into the output limit leaves at the end.
|
|
_UNCLOSED_BRACKET_RE = re.compile(r"\n?\[([^\]\n]*)\Z")
|
|
|
|
|
|
def _is_echoed_instruction(inner: str) -> bool:
|
|
"""Whether a trailing bracketed segment is the prompt's own reminder."""
|
|
low = inner.lower()
|
|
if "```state" in low:
|
|
return True
|
|
if low.lstrip().startswith("reminder:"):
|
|
return True
|
|
# `CHAT_CONTINUE_HINT` in `providers/openai_compatible.py`, which a model
|
|
# also parrots back, observed in the M11 long-run trial. Matched by its
|
|
# opening words only, because the echo is often cut off before it ends.
|
|
if low.lstrip().startswith("continue the story directly"):
|
|
return True
|
|
# v1.1 WP-A2 corrective (R5): the same hint, reworded at the front. The M11
|
|
# closeout-era identity re-run stored "[You don't need to continue; … Continue
|
|
# the story here, directly. Output only story text.]" as the last line of a
|
|
# reply, and because nothing recognised it, nothing above it was trailing.
|
|
if CONTINUE_HINT_PHRASE.lower() in low:
|
|
return True
|
|
# v1.1 WP-A2 (R2): the length hint, which names "state block" but not
|
|
# "events list", so it passed every check above.
|
|
if _is_length_hint(inner):
|
|
return True
|
|
# The reminder names both; prose about the protocol rarely names either the
|
|
# way the instruction does, and effectively never both.
|
|
return "state block" in low and "events list" in low
|
|
|
|
|
|
def _opens_like_length_hint(inner: str) -> bool:
|
|
"""R5. The bracket opens with the length hint's own `Hard limit:`, whatever follows.
|
|
|
|
Never enough on its own: an in-world "[Hard limit: forty days]" opens the same
|
|
way. `_clean` takes it only directly above an echoed instruction it has already
|
|
removed from the end of the same reply.
|
|
"""
|
|
return inner.lstrip().lower().startswith(LENGTH_HINT_OPENING[1:].lower())
|
|
|
|
|
|
def _is_length_hint(inner: str) -> bool:
|
|
"""Whether a bracket's contents are `builder.length_hint`, however reworded.
|
|
|
|
It must open the way the hint opens *and* carry the hint's own wording. An
|
|
in-world "Hard limit: forty days" has the opening and none of the wording.
|
|
"""
|
|
opening = LENGTH_HINT_OPENING[1:].lower()
|
|
return (inner.lstrip().lower().startswith(opening)
|
|
and bool(_LENGTH_HINT_PHRASE_RE.search(inner)))
|
|
|
|
|
|
# A heading the model writes above a block it did not fence: `State`, sometimes
|
|
# as `State:`, `**State**` or `### State`. It is removed only in two places:
|
|
# directly above a proposal that is removed, and as the last line of the reply.
|
|
# A line reading "State" in the middle of a story is left alone.
|
|
_STATE_HEADING_RE = re.compile(r"^[ \t>*#_]*state[ \t*_:]*$", re.IGNORECASE)
|
|
|
|
# An unfenced object that starts a line, optionally quoted with `>`, which small
|
|
# models copy from the player-turn convention.
|
|
_LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
|
|
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
|
|
|
|
|
|
def _clean(prose: str, *, after_block: bool = False) -> str:
|
|
"""Removes protocol the block extraction could not, and nothing else.
|
|
|
|
Found by the M5 realistic-context run (§12), which is the failure class
|
|
Phase 0B warned about: under a full prompt the model echoed its own
|
|
instruction into the narration, and the reader would have been shown it.
|
|
Neither case here is hypothetical — both were observed against a real local
|
|
model.
|
|
|
|
The M11 long run found two more, on 42 of 104 turns. The model pasted a copy
|
|
of the narrative-state section into its prose, and it wrote its proposal
|
|
unfenced under a bare `State` heading, sometimes quoted, sometimes with more
|
|
story after it. Stored text is replayed as history, so every leak also
|
|
showed the next prompt a second, older account of the state, which is what
|
|
M5 review Finding 4 removed from replayed history.
|
|
|
|
v1.1 WP-A2 added four shapes, from the M11 closeout's identity run and the
|
|
v1 corpus, each anchored to something the application owns rather than to
|
|
what prose looks like: a line opening with a vocabulary call (R1), the
|
|
length hint echoed at the end (R2), the renderer's scene line left last
|
|
(R3), and an empty fence opener left last (R4). `after_block` says a
|
|
proposal block was already taken out of this reply, which is what lets R3
|
|
remove a bare scene line that sat above it.
|
|
"""
|
|
cleaned, calls_removed = _strip_event_call_lines(prose)
|
|
cleaned, _found = _inline_proposals(cleaned)
|
|
cleaned = _strip_echoed_state(cleaned)
|
|
protocol_cut = after_block or calls_removed
|
|
# R5: set once an echoed instruction bracket has come off the end. Only then
|
|
# may a bracket that merely opens the way the length hint opens be taken as
|
|
# part of the same echoed tail.
|
|
instruction_cut = False
|
|
# The end of the reply is cut until nothing more comes off, because one kind
|
|
# of leftover can hide another. In a real reply, a `State` heading sat above
|
|
# a block the model never finished, and a parroted reminder sat above an
|
|
# unclosed fence.
|
|
while True:
|
|
before = cleaned
|
|
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
|
|
bracket = pattern.search(cleaned)
|
|
if bracket is None:
|
|
continue
|
|
if _is_echoed_instruction(bracket.group(1)):
|
|
cleaned = cleaned[: bracket.start()]
|
|
instruction_cut = True
|
|
elif instruction_cut and _opens_like_length_hint(bracket.group(1)):
|
|
cleaned = cleaned[: bracket.start()]
|
|
cleaned = _DANGLING_STATE_RE.sub("", cleaned)
|
|
dangling = _DANGLING_JSON_RE.search(cleaned)
|
|
if dangling is not None and (_reads_as_protocol(dangling.group(1))
|
|
or _is_opening_of_proposal(dangling.group(1))):
|
|
cleaned = cleaned[: dangling.start()]
|
|
cleaned = _strip_dangling_object(cleaned)
|
|
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
|
|
# A bare quote marker, the start of a quoted block that never came.
|
|
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
|
|
cleaned = _strip_empty_dangling_fence(cleaned)
|
|
if cleaned.rstrip() != before.rstrip():
|
|
protocol_cut = True
|
|
cleaned = _strip_trailing_scene_line(cleaned, protocol_cut)
|
|
if cleaned == before:
|
|
return cleaned.strip()
|
|
|
|
|
|
def _strip_event_call_lines(text: str) -> tuple[str, bool]:
|
|
"""R1. Removes whole lines that open with a call to a vocabulary event.
|
|
|
|
A line inside a fenced code block is the story's own code and is never
|
|
examined. Returns the text and whether anything was removed.
|
|
"""
|
|
kept: list[str] = []
|
|
in_fence = False
|
|
removed = False
|
|
for line in text.split("\n"):
|
|
if line.lstrip().startswith("```"):
|
|
in_fence = not in_fence
|
|
kept.append(line)
|
|
continue
|
|
if not in_fence and _EVENT_CALL_LINE_RE.match(line):
|
|
removed = True
|
|
continue
|
|
kept.append(line)
|
|
if not removed:
|
|
return text, False
|
|
return re.sub(r"\n{3,}", "\n\n", "\n".join(kept)), True
|
|
|
|
|
|
def _strip_empty_dangling_fence(text: str) -> str:
|
|
"""R4. A ```` ```json ```` or ```` ``` ```` opener as the last line, with nothing after it.
|
|
|
|
Only an *opener*: the fence lines are counted, and an even count means the
|
|
last one closes a story's own code block, which stays.
|
|
"""
|
|
lines = text.rstrip().split("\n")
|
|
if len(lines) < 2 or not _EMPTY_FENCE_LINE_RE.fullmatch(lines[-1].strip()):
|
|
return text
|
|
fences = sum(1 for line in lines if line.lstrip().startswith("```"))
|
|
if fences % 2 == 0:
|
|
return text
|
|
return "\n".join(lines[:-1]).rstrip()
|
|
|
|
|
|
def _strip_trailing_scene_line(text: str, protocol_cut: bool) -> str:
|
|
"""R3. The renderer's scene line, left as the last line of the reply.
|
|
|
|
Taken when it carries the renderer's own `(at <location>)`, or when protocol
|
|
was already cut from this reply, which makes a bare scene line part of the
|
|
same pasted tail. A final screenplay-style "Scene: …" line in a reply with
|
|
no protocol in it stays, and so does any scene line with story after it.
|
|
"""
|
|
lines = text.rstrip().split("\n")
|
|
if len(lines) < 2:
|
|
return text
|
|
last = lines[-1].strip()
|
|
if not last.startswith(render.HEADING_SCENE + " "):
|
|
return text
|
|
if not (_RENDERED_SCENE_LOCATION_RE.search(last) or protocol_cut):
|
|
return text
|
|
return "\n".join(lines[:-1]).rstrip()
|
|
|
|
|
|
def explain_removed_line(line: str) -> str | None:
|
|
"""Which v1.1 rule removes a line of this shape, for the replay report.
|
|
|
|
None means no v1.1 rule explains it, which the replay treats as a failure.
|
|
"""
|
|
stripped = line.strip()
|
|
if _EVENT_CALL_LINE_RE.match(line):
|
|
return RULE_EVENT_CALL
|
|
if stripped.startswith("["):
|
|
inner = stripped[1:]
|
|
inner = inner[:-1] if inner.endswith("]") else inner
|
|
if _is_length_hint(inner):
|
|
return RULE_LENGTH_HINT
|
|
if _is_echoed_instruction(inner) or _opens_like_length_hint(inner):
|
|
return RULE_INSTRUCTION_TAIL
|
|
if stripped.startswith(render.HEADING_SCENE + " "):
|
|
return RULE_SCENE_LINE
|
|
if _EMPTY_FENCE_LINE_RE.fullmatch(stripped):
|
|
return RULE_EMPTY_FENCE
|
|
return None
|
|
|
|
|
|
def _is_state_heading(line: str) -> bool:
|
|
return bool(_STATE_HEADING_RE.match(line))
|
|
|
|
|
|
def _strip_trailing_state_heading(text: str) -> str:
|
|
lines = text.rstrip().split("\n")
|
|
if lines and _is_state_heading(lines[-1]):
|
|
return "\n".join(lines[:-1])
|
|
return text
|
|
|
|
|
|
def _strip_dangling_object(text: str) -> str:
|
|
"""Cuts an unfenced proposal the model never finished, and what follows it.
|
|
|
|
A reply that runs into the output-token limit mid-block ends inside the
|
|
object, often a quoted one. That happened on 10 of 104 turns in the M11 long
|
|
run. The object never closes, so `_inline_proposals` cannot take it. The
|
|
candidate is the outermost object that stays unclosed, not the last line
|
|
that opens one. Its finished event objects open lines too, and they close,
|
|
so cutting at the last of them left the list above it in the story. It is
|
|
cut when it reads as protocol (`_reads_as_protocol`), the same test a
|
|
truncated ```json fence has to pass.
|
|
"""
|
|
skip_until = 0
|
|
for match in _LINE_OBJECT_RE.finditer(text):
|
|
line_start = match.start()
|
|
if line_start < skip_until:
|
|
continue
|
|
body = "\n".join(_QUOTE_PREFIX_RE.sub("", line, count=1)
|
|
for line in text[line_start:].split("\n"))
|
|
closing = _object_end(body, body.find("{"))
|
|
if closing is None:
|
|
return text[:line_start] if _reads_as_protocol(body) else text
|
|
# Quote markers came off `body`, so this position is never past the
|
|
# real end of the object. A line inside the object that is examined
|
|
# anyway closes inside it, and is passed over too.
|
|
skip_until = line_start + closing
|
|
return text
|
|
|
|
|
|
# The markdown a model wraps a heading in: `## Established:`, `**Held:**`,
|
|
# `> Held:`.
|
|
_HEADING_DECORATION_RE = re.compile(r"^[\s#>*_]+|[\s*_]+$")
|
|
|
|
|
|
def _section_heading(line: str) -> str | None:
|
|
"""The state-section heading this line is, markdown aside, or None."""
|
|
bare = _HEADING_DECORATION_RE.sub("", line)
|
|
return bare if bare in render.SECTION_HEADINGS else None
|
|
|
|
|
|
def _strip_echoed_state(text: str) -> str:
|
|
"""Removes a copy of the narrative-state section pasted into the prose.
|
|
|
|
Judged by the section's own headings (`render.SECTION_HEADINGS`) as whole
|
|
lines, with any markdown the model wrapped them in taken off. A block
|
|
qualifies when it carries two headings, or one and the scene line directly
|
|
above it, or one heading with an indented entry under it. That last case
|
|
is the model writing a section of its own: the M04 re-run found
|
|
`## Established:` over two indented facts on 5 turns, one of them copying
|
|
the planted clue out of the state section. A lone "Held:" with prose after
|
|
it is still somebody's story. The block runs over the headings, their
|
|
indented entries and the blank lines between them, and stops at the first
|
|
line of ordinary prose.
|
|
"""
|
|
lines = text.split("\n")
|
|
drop = [False] * len(lines)
|
|
index = 0
|
|
while index < len(lines):
|
|
if _section_heading(lines[index]) is None:
|
|
index += 1
|
|
continue
|
|
start = index
|
|
above = index - 1
|
|
while above >= 0 and not lines[above].strip():
|
|
above -= 1
|
|
scene = above >= 0 and (
|
|
lines[above].strip() == render.HEADING_SCENE
|
|
or lines[above].lstrip().startswith(render.HEADING_SCENE + " ")
|
|
)
|
|
if scene:
|
|
start = above
|
|
headings: set[str] = set()
|
|
entries = 0
|
|
end = index
|
|
cursor = index
|
|
while cursor < len(lines):
|
|
line = lines[cursor]
|
|
stripped = line.strip()
|
|
heading = _section_heading(line)
|
|
if heading is not None:
|
|
headings.add(heading)
|
|
end = cursor
|
|
elif stripped and line[:1] in (" ", "\t"):
|
|
entries += 1
|
|
end = cursor
|
|
elif stripped:
|
|
break
|
|
cursor += 1
|
|
if len(headings) + (1 if scene else 0) >= 2 or (headings and entries):
|
|
for position in range(start, end + 1):
|
|
drop[position] = True
|
|
index = end + 1
|
|
if not any(drop):
|
|
return text
|
|
kept = "\n".join(line for line, gone in zip(lines, drop) if not gone)
|
|
return re.sub(r"\n{3,}", "\n\n", kept)
|
|
|
|
|
|
def _object_end(text: str, start: int) -> int | None:
|
|
"""Where the JSON object opening at `start` closes, strings respected."""
|
|
depth, in_string, escaped = 0, False, False
|
|
for position in range(start, len(text)):
|
|
char = text[position]
|
|
if in_string:
|
|
if escaped:
|
|
escaped = False
|
|
elif char == "\\":
|
|
escaped = True
|
|
elif char == '"':
|
|
in_string = False
|
|
elif char == '"':
|
|
in_string = True
|
|
elif char == "{":
|
|
depth += 1
|
|
elif char == "}":
|
|
depth -= 1
|
|
if depth == 0:
|
|
return position + 1
|
|
return None
|
|
|
|
|
|
def _inline_proposals(text: str) -> tuple[str, list[tuple[dict, str]]]:
|
|
"""Removes unfenced proposals that start a line, and returns them.
|
|
|
|
A candidate must parse and must be a proposal (`_looks_like_proposal`), the
|
|
same bar as a bare trailing object. JSON a character wrote stays where it
|
|
is. A quoted candidate is read with its `>` markers taken off, across the
|
|
consecutive quoted lines. A candidate with prose after it on its closing
|
|
line is not on its own lines, and is left alone. A bare `State` heading
|
|
directly above a removed proposal goes with it.
|
|
|
|
Returns the text without them, and `(parsed, raw)` for each, oldest first.
|
|
"""
|
|
found: list[tuple[dict, str]] = []
|
|
cuts: list[tuple[int, int]] = []
|
|
# Candidates are taken outermost first. A line inside an object already
|
|
# examined is part of that object, and a proposal's own event lines open
|
|
# objects too, so one of them must never be taken as a proposal by itself.
|
|
# An object that never closes runs to the end of the text, so everything
|
|
# after it is inside it.
|
|
skip_until = 0
|
|
for match in _LINE_OBJECT_RE.finditer(text):
|
|
line_start = match.start()
|
|
if line_start < skip_until:
|
|
continue
|
|
line_end = text.find("\n", line_start)
|
|
line_end = len(text) if line_end == -1 else line_end
|
|
if _QUOTE_PREFIX_RE.match(text[line_start:line_end]):
|
|
# Gather the quoted run, unquote it, and find the object inside.
|
|
spans, cursor = [], line_start
|
|
while cursor < len(text):
|
|
stop = text.find("\n", cursor)
|
|
stop = len(text) if stop == -1 else stop
|
|
if not _QUOTE_PREFIX_RE.match(text[cursor:stop]):
|
|
break
|
|
spans.append((cursor, stop))
|
|
cursor = stop + 1
|
|
body_lines = [_QUOTE_PREFIX_RE.sub("", text[a:b], count=1) for a, b in spans]
|
|
body = "\n".join(body_lines)
|
|
opening = body.find("{")
|
|
closing = _object_end(body, opening)
|
|
if closing is None:
|
|
break
|
|
consumed = body[:closing].count("\n")
|
|
region_end = spans[consumed][1]
|
|
skip_until = region_end
|
|
if body[closing:].split("\n", 1)[0].strip():
|
|
continue
|
|
raw = body[opening:closing]
|
|
else:
|
|
opening = match.end() - 1
|
|
closing = _object_end(text, opening)
|
|
if closing is None:
|
|
break
|
|
rest = text.find("\n", closing)
|
|
rest = len(text) if rest == -1 else rest
|
|
skip_until = rest
|
|
if text[closing:rest].strip():
|
|
continue
|
|
raw = text[opening:closing]
|
|
region_end = rest
|
|
parsed = _tolerant_load(raw)
|
|
if not _looks_like_proposal(parsed):
|
|
continue
|
|
region_start = line_start
|
|
before = text[:line_start].rstrip("\n").rstrip()
|
|
heading_start = before.rfind("\n") + 1
|
|
if before and _is_state_heading(before[heading_start:]):
|
|
region_start = heading_start
|
|
cuts.append((region_start, region_end))
|
|
found.append((parsed, raw))
|
|
if not cuts:
|
|
return text, found
|
|
pieces, cursor = [], 0
|
|
for start, end in cuts:
|
|
pieces.append(text[cursor:start])
|
|
cursor = end
|
|
pieces.append(text[cursor:])
|
|
return re.sub(r"\n{3,}", "\n\n", "".join(pieces)), found
|
|
|
|
|
|
def _is_opening_of_proposal(tail: str) -> bool:
|
|
"""Whether a truncated fence stopped before it could say what it was.
|
|
|
|
`{` followed by nothing but the start of `"events"`. The output limit cut
|
|
one reply there, before `_reads_as_protocol` had anything to go on. A
|
|
story's own code block is not that short, and one that is holds nothing to
|
|
lose."""
|
|
body = tail.strip()
|
|
return body.startswith("{") and '"events"'.startswith(body[1:].strip())
|
|
|
|
|
|
def _reads_as_protocol(tail: str) -> bool:
|
|
"""Whether a truncated fence was on its way to being a proposal."""
|
|
if '"events"' in tail:
|
|
return True
|
|
return any(f'"{name}"' in tail for name in events.SPECS)
|
|
|
|
|
|
def _tolerant_load(blob: str):
|
|
"""Parses a block, forgiving what small local models get wrong.
|
|
|
|
Trailing commas and a leading `+` on a number are both common and both
|
|
rejected by strict JSON. Repairing them is not guessing at meaning — the
|
|
intended value is unambiguous — which is the line this function stays on the
|
|
right side of. Anything it cannot parse returns None, and the caller treats
|
|
that as no proposal rather than as an empty one.
|
|
"""
|
|
cleaned = re.sub(r",(\s*[}\]])", r"\1", blob)
|
|
cleaned = re.sub(r"(:\s*)\+(\d)", r"\1\2", cleaned)
|
|
try:
|
|
parsed = json.loads(cleaned)
|
|
except (json.JSONDecodeError, ValueError):
|
|
return None
|
|
return parsed
|
|
|
|
|
|
def split(text: str) -> tuple[str, dict | None, str]:
|
|
"""Separates a reply into `(prose, proposal, raw_block)`.
|
|
|
|
`proposal` is None when there is no block or it cannot be parsed at all,
|
|
which the caller records as a malformed proposal. `raw_block` is what the
|
|
model actually wrote, kept for the audit record even — especially — when it
|
|
did not parse.
|
|
|
|
A bare trailing object is only stripped when it parses *and* looks like a
|
|
proposal. Prose that happens to end in a brace is left alone, because
|
|
removing a sentence from someone's story to satisfy a regex is a worse
|
|
failure than leaving a stray brace in it.
|
|
"""
|
|
matches = list(_STATE_FENCE_RE.finditer(text))
|
|
if matches:
|
|
match = matches[-1]
|
|
raw = match.group(1).strip()
|
|
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
|
return prose, _tolerant_load(raw), raw
|
|
|
|
# A `json` or unlabelled fence is ours only when its contents are this
|
|
# protocol. That is judged two ways, and it needs both: a block that parses
|
|
# into a proposal, or one that plainly reads as protocol even though it does
|
|
# not parse. The second half matters — a small model that mangles its own
|
|
# JSON must not have the wreckage shown to the reader, which is what the
|
|
# realistic-model run caught during the corrective pass.
|
|
for pattern in (_JSON_FENCE_RE, _BARE_FENCE_RE):
|
|
for match in reversed(list(pattern.finditer(text))):
|
|
raw = match.group(1).strip()
|
|
parsed = _tolerant_load(raw)
|
|
if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
|
|
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
|
return prose, parsed, raw
|
|
|
|
match = _TRAILING_RE.search(text)
|
|
if match:
|
|
raw = match.group(1)
|
|
parsed = _tolerant_load(raw)
|
|
if _looks_like_proposal(parsed):
|
|
return _clean(text[: match.start()], after_block=True), parsed, raw
|
|
|
|
# An unfenced proposal on its own lines but not at the end: quoted, or
|
|
# followed by more story. The last one is the turn's proposal, as with
|
|
# fences, and every one leaves the prose.
|
|
without, found = _inline_proposals(text)
|
|
if found:
|
|
parsed, raw = found[-1]
|
|
return _clean(without, after_block=True), parsed, raw
|
|
|
|
# No block at all — but the reply may still carry protocol the model wrote
|
|
# as prose, or a fence it never closed.
|
|
cleaned = _clean(text)
|
|
whole = text.strip()
|
|
if cleaned == whole:
|
|
return cleaned, None, ""
|
|
# What came off is kept for the audit when it was protocol: an unfinished
|
|
# block, a parroted reminder, or a fence. A pasted copy of the state section
|
|
# is not a proposal, so a reply with nothing else removed records no block.
|
|
# That keeps the turn from being marked unparseable for a block it never
|
|
# started.
|
|
if whole.startswith(cleaned):
|
|
removed = whole[len(cleaned):].strip()
|
|
keep = (_reads_as_protocol(removed) or "```" in removed
|
|
or removed.startswith("["))
|
|
return cleaned, None, removed if keep else ""
|
|
# Text also came out of the middle, so what was removed is not one suffix.
|
|
return cleaned, None, whole if _reads_as_protocol(whole) else ""
|
|
|
|
|
|
def _looks_like_proposal(parsed) -> bool:
|
|
"""Whether a bare trailing object is this protocol rather than prose."""
|
|
if not isinstance(parsed, dict):
|
|
return False
|
|
if isinstance(parsed.get("events"), list):
|
|
return True
|
|
return isinstance(parsed.get("type"), str) and events.is_allowed(parsed["type"])
|
|
|
|
|
|
def render_block(accepted: list[dict]) -> str:
|
|
"""Renders accepted events back into the block the model emitted.
|
|
|
|
Replayed into the prompt for past turns so the model copies the format it is
|
|
being asked for. **Accepted** events rather than proposed ones, for the
|
|
reason the delta protocol learned the hard way: showing the model a refused
|
|
event standing as though it had worked, contradicted by the state in the
|
|
same prompt, teaches it to send the event again.
|
|
"""
|
|
if not accepted:
|
|
return ""
|
|
return "```state\n" + json.dumps({"events": accepted}, ensure_ascii=False) + "\n```"
|
|
|
|
|
|
def render_rejections(rejected: list[dict]) -> str:
|
|
"""The correction note appended after the most recent AI turn.
|
|
|
|
Only what was lost. A model that is told what it got wrong can fix it next
|
|
turn; a model told nothing repeats it.
|
|
"""
|
|
if not rejected:
|
|
return ""
|
|
lines = []
|
|
for entry in rejected[:6]:
|
|
if not isinstance(entry, dict):
|
|
continue
|
|
detail = entry.get("detail") or entry.get("reason") or ""
|
|
if detail:
|
|
lines.append(f"- {detail}")
|
|
if not lines:
|
|
return ""
|
|
body = "\n".join(lines)
|
|
return (
|
|
"[Part of your last state block was not accepted. Correct it in this "
|
|
f"turn's block:\n{body}]"
|
|
)
|