Files
interactive-story/backend/tests/test_v11_protocol_echo.py
T
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00

325 lines
14 KiB
Python

"""v1.1 WP-A2: the protocol a narrator copies stays out of the story, and nothing else does.
The M11 closeout's identity run (an office meeting, a 3B narrator, a 4,096
window) stored four turns carrying text the application wrote, not the story:
- `> Create_entity(new_person, "john", …` — the event vocabulary as the prompt
printed it, `name(field, …)`, copied as if it were a call;
- `> set_possession(silver-key, "alice") Adds the silver key to Alice's
possession.` — the call again, naming the fantasy example slug from the fixed
state rule, in a meeting room;
- `Scene: Bill, Alice, … (at The meeting room)` — the renderer's own scene line;
- `[Hard limit: your next turn must not exceed 180 words, … append the state
block well inside the limit.]` — the length hint, reworded at the front and
verbatim at the end.
v1.0.0 removed none of them. The rule this module is held to is unchanged from
M5: **removing story is worse than leaving protocol.** Every removal below is
anchored to a string or a vocabulary the application owns, and every one has
story beside it that must survive.
python -m pytest tests/test_v11_protocol_echo.py -v
"""
import json
import re
import pytest
from app.context import builder
from app.narrative import events, extract, render
# ---------------------------------------------------------------- the prompt
#: The identifiers the v1 state rule taught every campaign, from the fantasy
#: acceptance fixture. None may come back into a fixed instruction.
FANTASY_IDENTIFIERS = ("mara", "silver-key", "silver key", "old-abbey", "abbey",
"aldric", "westhaven", "crypt", "edrin")
#: And nothing from the science-fiction fixture either: neutral means neutral,
#: not "the other genre".
SCIFI_IDENTIFIERS = ("persephone", "imani", "data-crystal", "data crystal", "airlock")
def _fixed_instructions() -> str:
return "\n".join([
extract.EMIT_RULE,
extract.EMIT_REMINDER,
events.vocabulary_for_prompt(),
builder.length_hint(500),
builder.length_hint(500, "brief"),
builder.length_hint(500, "long"),
builder.length_hint(120),
]).lower()
@pytest.mark.parametrize("identifier", FANTASY_IDENTIFIERS + SCIFI_IDENTIFIERS)
def test_the_fixed_state_instructions_name_no_fixture_identifier(identifier):
"""A2-1. The example slug the office run copied cannot come back."""
assert not re.search(rf"\b{re.escape(identifier)}\b", _fixed_instructions())
def test_the_example_uses_neutral_identifiers():
for neutral in ("character-1", "item-1", "location-1"):
assert neutral in extract.EMIT_RULE
def test_the_worked_example_is_a_block_this_extractor_accepts():
"""The example is the wire format, byte for byte, not an illustration of it."""
example = extract.EMIT_RULE[extract.EMIT_RULE.index("```state"):]
prose, parsed, _raw = extract.split("The door opens.\n\n" + example)
assert prose == "The door opens."
assert isinstance(parsed, dict)
assert [e["type"] for e in parsed["events"]] == [
"set_possession", "set_current_location"]
for event in parsed["events"]:
assert events.is_allowed(event["type"])
def test_the_vocabulary_is_not_written_as_function_calls():
"""A2-2. `set_possession(item, owner)` is the notation the narrator copied."""
vocabulary = events.vocabulary_for_prompt()
for name in events.SPECS:
assert not re.search(rf"\b{name}\s*\(", vocabulary), name
def test_every_event_is_described_in_the_shape_the_model_must_send():
lines = events.vocabulary_for_prompt().splitlines()
assert len(lines) == len(events.SPECS)
for name, line in zip(events.SPECS, lines):
shape = line.strip().split(" — ", 1)[0]
obj = json.loads(shape)
assert obj["type"] == name
assert set(obj) - {"type"} == set(events.SPECS[name]["required"])
for optional in events.SPECS[name]["optional"]:
assert optional in line
def test_the_length_hint_carries_the_phrases_the_extractor_recognises():
"""One source for the words, so the builder and the extractor cannot drift."""
for narration_length in ("", "brief", "medium", "long"):
hint = builder.length_hint(500, narration_length)
assert hint.startswith(extract.LENGTH_HINT_OPENING)
assert extract.LENGTH_HINT_TAIL in hint
# ---------------------------------------------------------- observed shapes
STORY = (
"Alice looks at John, the tension in the room palpable.\n\n"
"John nods. \"I'm ready to contribute.\""
)
@pytest.mark.parametrize("leak", [
# Depth 20: the call, the fantasy slug, and a gloss on the same line.
'> set_possession(silver-key, "alice") Adds the silver key to Alice\'s possession.',
# Depth 10: cut off by the output limit mid-call.
'> Create_entity(new_person, "john", "character", "A determined team member", ["john',
# Unquoted, and a vocabulary name in any case.
'SET_CURRENT_LOCATION(bill, office)',
'add_fact(predicate="knows the plan", subject="alice")',
])
def test_an_event_call_line_at_the_end_leaves_the_story(leak):
prose, parsed, _raw = extract.split(f"{STORY}\n\n{leak}")
assert prose == STORY
assert parsed is None
def test_an_event_call_line_in_the_middle_leaves_and_the_story_after_it_stays():
"""Depth 12: the call, then more narration."""
reply = (
f"{STORY}\n\n"
'> Create_entity(new_person, "mike", "character", "A new team member.", ["mike"])\n\n'
"Mike takes the empty chair by the window."
)
prose, _parsed, _raw = extract.split(reply)
assert prose == f"{STORY}\n\nMike takes the empty chair by the window."
def test_the_depth_fourteen_tail_leaves_entirely():
"""A call, a rendered scene line, and a reworded length hint, in that order."""
reply = (
f"{STORY}\n\n"
'> Create_entity(mike, "character", "A new team member.", ["mike"])\n\n'
"Scene: Bill, Alice, Roger, John, and Mike at the table. (at The meeting room)\n\n"
"[Hard limit: your next turn must not exceed 180 words, and it should not stop "
"short of about 70. Prefer the lower end of that range unless the scene genuinely "
"needs more. Finish the narration and append the state block well inside the limit.]"
)
prose, parsed, _raw = extract.split(reply)
assert prose == STORY
assert parsed is None
@pytest.mark.parametrize("hint", [
builder.length_hint(500),
builder.length_hint(500, "brief"),
# Cut off by the output limit before the tail.
"[Hard limit: this turn must not exceed 180 words, and it should not stop short",
# Reworded at the front, as the 3B narrator did.
"[Hard limit: your next turn must not exceed 506 words. Write only as much as the "
"moment needs — a typical turn is much shorter. Finish the narration and append "
"the state block well inside the limit.]",
])
def test_a_parroted_length_hint_at_the_end_leaves_the_story(hint):
prose, _parsed, _raw = extract.split(f"{STORY}\n\n{hint}")
assert prose == STORY
def test_a_rendered_scene_line_at_the_end_leaves_the_story():
prose, _parsed, _raw = extract.split(
f"{STORY}\n\nScene: A tense budget meeting. (at The meeting room)")
assert prose == STORY
def test_a_fenced_block_with_a_call_line_above_it_still_parses_and_applies():
"""A2-6. The proposal is still read when protocol litter surrounds it."""
reply = (
f"{STORY}\n\n"
'> set_current_location(john, office)\n\n'
'```state\n{"events": [{"type": "set_current_location", '
'"entity": "john", "location": "office"}]}\n```'
)
prose, parsed, raw = extract.split(reply)
assert prose == STORY
assert parsed["events"][0]["entity"] == "john"
assert raw.startswith("{")
# ------------------------------------------------ adversarial story that stays
@pytest.mark.parametrize("reply", [
# The owner's cases.
'The engineer writes "set_power(core, 80)" on the whiteboard.',
'She says, "Create_entity is a terrible name for a company."',
'The old manual contains a heading labeled "Scene:"',
'He reads aloud: "[Hard limit: 500 words]" and laughs.',
# A vocabulary name, written into a story, not at the start of a line.
'Nadia squints at the log: the last command was set_possession(badge, guard).',
# Call-shaped, at the start of a line, but not an event this protocol has.
"The terminal scrolls.\n\n> open_door(north)\n\nNothing happens.",
# A vocabulary call inside the story's own code block is the story's code.
"She types:\n\n```python\ncreate_entity(ship)\nset_possession(key, captain)\n```\n\n"
"The console beeps twice.",
# A bracket at the very end, in-world, that is not the application's hint.
"The warning light blinks.\n\n[Hard limit of the reactor: three hours]",
"The contract ends with a clause.\n\n[Hard limit: forty days, no extensions]",
# A scene heading in a screenplay the characters are writing, mid-story.
"Scene: a kitchen, late.\n\nShe crosses it out and starts again.",
# A last line that starts like the renderer's but is not its shape.
"The director calls it.\n\nScene: take two, and nobody moves.",
# A fact restated inside a sentence.
"Alice knew the badge opened the server room, and said nothing.",
"Memory: she remembered the bells.",
])
def test_story_that_resembles_the_new_rules_is_kept(reply):
"""A2-5."""
prose, parsed, _raw = extract.split(reply)
assert prose == reply
assert parsed is None
# ------------------------------------------------------ the replay attribution
@pytest.mark.parametrize("line, rule", [
('> set_possession(silver-key, "alice") Adds the key.', extract.RULE_EVENT_CALL),
("Create_entity(new_person", extract.RULE_EVENT_CALL),
("[Hard limit: this turn must not exceed 90 words. Finish the narration and append "
"the state block well inside the limit.]", extract.RULE_LENGTH_HINT),
("Scene: A meeting. (at The meeting room)", extract.RULE_SCENE_LINE),
("John nods.", None),
('He reads aloud: "[Hard limit: 500 words]" and laughs.', None),
])
def test_a_removed_line_is_attributed_to_the_rule_that_removes_it(line, rule):
assert extract.explain_removed_line(line) == rule
# ------------------------------------- corrective: the depth-16 instruction tail
#: Cut down from the v1.1 identity diagnostic's depth-16 turn, whose stored text
#: was exactly the extractor's output. The two story paragraphs are shortened;
#: the four trailing lines are verbatim.
DEPTH_16_STORY = (
"John's initial ideas are thoughtful and insightful, and the room fills with a "
"sense of optimism.\n\n"
"John's enthusiasm is contagious, and the meeting room is electric with the "
"excitement of a fruitful collaboration ahead."
)
DEPTH_16_TAIL = (
"Scene: Bill, Alice and Roger at the table; John not yet arrived.\n\n"
"[Hard limit: this is now 180 words.]\n\n"
"[Reminder: end your reply with a `state` block listing the events your narration "
"made true, with absolute values.]\n\n"
"[You don't need to continue; your turn must now be about John entering the room. "
"Continue the story here, directly. Output only story text.]"
)
def test_the_depth_sixteen_instruction_tail_leaves_entirely():
"""The corrective's positive regression. v1.1's first A2 left all four lines:
the last bracket was a reworded continue hint nothing recognised, so nothing
above it was ever at the end."""
prose, parsed, _raw = extract.split(f"{DEPTH_16_STORY}\n\n{DEPTH_16_TAIL}")
assert prose == DEPTH_16_STORY
assert parsed is None
def test_the_continue_hint_phrase_is_the_providers_own_sentence():
from app.providers.openai_compatible import CHAT_CONTINUE_HINT
assert extract.CONTINUE_HINT_PHRASE in CHAT_CONTINUE_HINT
def test_an_echoed_continue_hint_alone_at_the_end_leaves():
prose, _p, _r = extract.split(
f"{STORY}\n\n[Keep going. Continue the story here, directly. Output only story text.]")
assert prose == STORY
@pytest.mark.parametrize("reply", [
# A hint-opened bracket with no echoed instruction below it is in-world.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# Nor does a state block below it make it an instruction.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# The phrase in the middle of a story is prose, not a trailing echo.
'She wrote "output only story text" on the card, then crossed it out.\n\nThe rain went on.',
# A trailing in-world bracket that only resembles a continuation.
f"{STORY}\n\n[To be continued]",
])
def test_story_brackets_near_the_corrective_rule_are_kept(reply):
prose, _p, _r = extract.split(reply)
assert prose == reply
def test_a_hint_opened_bracket_above_a_state_block_is_kept():
reply = (f"{STORY}\n\n[Hard limit: forty days, no extensions]\n\n"
'```state\n{"events": []}\n```')
prose, parsed, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[Hard limit: forty days, no extensions]"
assert parsed == {"events": []}
def test_an_in_world_bracket_above_an_echoed_hint_is_kept():
"""Only a bracket opening the way the application's hint opens is taken with
the echo. Any other bracket above it is the story's."""
reply = (f"{STORY}\n\n[The sign on the door reads: Closed]\n\n"
"[Continue the story here, directly. Output only story text.]")
prose, _p, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[The sign on the door reads: Closed]"
@pytest.mark.parametrize("line, rule", [
("[Hard limit: this is now 180 words.]", extract.RULE_INSTRUCTION_TAIL),
("[Reminder: end your reply with a `state` block listing the events.]",
extract.RULE_INSTRUCTION_TAIL),
("[You don't need to continue. Output only story text.]", extract.RULE_INSTRUCTION_TAIL),
])
def test_the_corrective_rule_is_attributed(line, rule):
assert extract.explain_removed_line(line) == rule
def test_the_new_rules_do_not_disturb_the_section_headings_they_share_a_module_with():
"""The renderer's headings are the M11 rules' anchor. A2 adds none."""
assert render.HEADING_SCENE == "Scene:"
assert "Scene:" not in render.SECTION_HEADINGS