Files
interactive-story/backend/app/narrative/events.py
T
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00

216 lines
8.7 KiB
Python

"""M5: the typed event vocabulary, and the allowlist that bounds it.
ADR 010 replaced AI-DnD's relative-delta protocol because the ambiguity was
architectural: a number in a delta field is syntactically legal whether the
model meant "add 50" or "set to 50", and no validator can tell which. Every
event here therefore states its operation in its `type`, and every value it
carries is **absolute**. There is no event whose meaning depends on a prompt
instruction having been followed.
## The allowlist is a security boundary, not a convenience
Model output is untrusted input (`SECURITY-THREAT-MODEL.md`), and this table is
the entire set of things a model may cause to happen. H05's
`{"event_type": "execute_shell", ...}` is refused here — not because "shell" is
recognised and blocked, but because it is not in `SPECS`, and nothing outside
`SPECS` is dispatched. There is no fallback branch, no generic handler and no
name-to-callable lookup that a payload could steer.
Adding an event means adding a spec here and a case in `apply.py`. Nothing else
in the application can widen the vocabulary, which is what keeps
"state extraction" from drifting into "tool execution".
## Shape of a spec
required fields that must be present and non-empty
optional fields that may be present
refs fields naming an entity that must already exist
creates the field naming an entity this event may bring into being
`refs` is what `validate.py` uses for referential integrity, and `creates` is
the deliberate exception: exactly one event type may introduce an entity, so a
typo in any other event surfaces as an unknown reference rather than silently
creating a second, empty Mara.
"""
from __future__ import annotations
import json
# Field types the schema layer enforces. Kept deliberately small: a narrative
# state event carries names, labels and plain values, and nothing here needs a
# nested structure a model could hide something inside.
TEXT = "text"
KEY = "key" # an entity/thread identifier: a slug the campaign chose
VALUE = "value" # a JSON scalar — str, int, float, bool or None
LABELS = "labels" # a list of short strings
#: The whole vocabulary. Nothing outside this mapping is dispatched, ever.
SPECS: dict[str, dict] = {
"create_entity": {
"required": {"entity": KEY, "name": TEXT},
"optional": {"entity_type": TEXT, "description": TEXT, "aliases": LABELS},
"refs": (),
"creates": "entity",
"summary": "brings a person, place, thing or group into the story",
},
"set_entity_status": {
"required": {"entity": KEY, "status": TEXT},
"optional": {},
"refs": ("entity",),
"creates": None,
"summary": "sets whether an entity is active, gone, destroyed …",
},
"set_entity_attribute": {
# The one numeric-capable event, and it is an assignment. ADR 010's
# `set_value`: the operation is in the name, so a value of 50 can only
# mean fifty. An `increment_value` could be added later without
# ambiguity, because it would be a different `type`.
"required": {"entity": KEY, "attribute": TEXT, "value": VALUE},
"optional": {},
"refs": ("entity",),
"creates": None,
"summary": "sets a named value on an entity, absolutely",
},
"set_entity_conditions": {
# Absolute too: the full set replaces the old one. "Add a condition"
# would need the current set to be known by the model, which is exactly
# the assumption that made deltas unreliable.
"required": {"entity": KEY, "conditions": LABELS},
"optional": {},
"refs": ("entity",),
"creates": None,
"summary": "replaces the conditions an entity is under",
},
"set_current_location": {
"required": {"entity": KEY, "location": KEY},
"optional": {},
"refs": ("entity", "location"),
"creates": None,
"summary": "moves an entity to a location",
},
"set_possession": {
"required": {"item": KEY, "owner": KEY},
"optional": {},
"refs": ("item", "owner"),
"creates": None,
"summary": "gives an item to an owner",
},
"clear_possession": {
"required": {"item": KEY},
"optional": {},
"refs": ("item",),
"creates": None,
"summary": "leaves an item held by nobody",
},
"add_fact": {
"required": {"predicate": TEXT},
"optional": {
"subject": KEY, "object": KEY, "value": VALUE, "fact_id": TEXT,
},
# Only the subject is checked as an entity. The *object* of a fact is
# routinely not one — "Mara knows where the key was found" has another
# fact as its object, and C03 needs exactly that — so it is checked
# against entities *and* known facts in `validate._check`. Requiring an
# entity here would make the knowledge distinction C03 asks for
# unrepresentable.
"refs": ("subject",),
"creates": None,
"summary": "asserts something about the world",
},
"invalidate_fact": {
"required": {"fact_id": TEXT},
"optional": {"reason": TEXT},
"refs": (),
"creates": None,
"summary": "withdraws a fact without deleting the record of it",
},
"add_relationship": {
"required": {"source": KEY, "target": KEY, "relationship": TEXT},
"optional": {"description": TEXT},
"refs": ("source", "target"),
"creates": None,
"summary": "ties two entities together",
},
"end_relationship": {
"required": {"source": KEY, "target": KEY, "relationship": TEXT},
"optional": {},
"refs": ("source", "target"),
"creates": None,
"summary": "ends a tie without erasing that it existed",
},
"open_story_thread": {
"required": {"thread": KEY, "title": TEXT},
"optional": {"description": TEXT},
"refs": (),
"creates": None,
"summary": "records narrative business left open",
},
"resolve_story_thread": {
"required": {"thread": KEY},
"optional": {"resolution": TEXT},
"refs": (),
"creates": None,
"summary": "closes narrative business",
},
"set_scene": {
"required": {},
"optional": {"summary": TEXT, "location": KEY, "present": LABELS},
"refs": ("location",),
"creates": None,
"summary": "records the immediate situation",
},
}
#: The allowlist itself, as a set, for the one question that matters most.
ALLOWED = frozenset(SPECS)
def is_allowed(event_type) -> bool:
"""Whether `event_type` names an event this application will ever apply.
A string is required: a dict, a list or None is not a type, and coercing one
with `str()` would turn a malformed payload into a lookup that might
accidentally succeed.
"""
return isinstance(event_type, str) and event_type in ALLOWED
def spec(event_type: str) -> dict | None:
return SPECS.get(event_type)
def vocabulary_for_prompt() -> str:
"""The event list as the narrator prompt describes it.
Generated from `SPECS` rather than written out beside it, so the model can
never be told about an event the application does not implement — the drift
that would produce proposals rejected for reasons nobody could see.
v1.1 WP-A2: each event is shown as the object the model must put in the
`events` list, with its required fields, not as `name(field, …)`. The call
notation was never the wire format, and a 3B narrator copied it into its
prose as `> set_possession(silver-key, "alice")`. An object copied into prose
is a proposal the extractor already recognises and removes; a call is not.
"""
lines = []
for name, definition in SPECS.items():
shape = {"type": name}
for field, kind in definition["required"].items():
shape[field] = _PLACEHOLDER[kind]
body = json.dumps(shape, ensure_ascii=False, separators=(",", ":"))
line = f" {body} — {definition['summary']}"
if definition["optional"]:
line += f" (optional: {', '.join(definition['optional'])})"
lines.append(line)
return "\n".join(lines)
#: What a field of each kind looks like in the prompt's vocabulary. Placeholders,
#: never example identifiers, so the vocabulary names nothing a story could copy.
#: A list field is shown as a list, so the model is told its shape; every other
#: field is an ellipsis. Measured: `"<key>"`-style placeholders with spaced
#: separators cost 456 tokens against v1.0.0's 258; this form costs about 380,
#: and every line is still the object the model must send.
_PLACEHOLDER = {KEY: "…", TEXT: "…", VALUE: "…", LABELS: ["…"]}