M5: genre-neutral authoritative narrative state, with review corrections

Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.

This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:

    visible active transcript position == stored head == authoritative state

Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)

  A narrator edit no longer rewrites a row. It returns to the state before the
  turn, takes the reader's exact text as the accepted narration, re-derives the
  state that text implies, and becomes a new active continuation — while the
  original narration keeps its words, its live flag and its whole future as
  retained history. At the tip the correction is another take; with story below
  it, it forks. No new history machinery: this is the existing fork/take/head
  path with the reader's text in place of a generated reply. The §14A refusal
  is therefore gone for narrator turns, and remains only for player input.

Pre-M5 positions

  Migration 88 backfills the empty narrative document onto every action written
  before M5, and a missing snapshot now restores the empty document instead of
  leaving the previous position's state standing. Restoring to an old Save
  Point no longer leaves a later position's entities and facts on screen.

Narrator context

  Replayed history carries prose only; the machine-readable block is no longer
  reconstructed into past turns, where it contradicted the authoritative state
  in the same prompt. A fact withdrawn by a manual correction is now named as
  no longer true, with the reader's reason, rather than silently dropped.

Also

  - state_changes joins the action-list bulk read, removing one query per row.
  - Extraction takes only the application's own protocol payload: an ordinary
    ```json or ```python block in a story survives, and a mangled proposal
    still does not reach the reader.

Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-05 07:01:50 -04:00
co-authored by Claude Opus 5
parent 62a997f364
commit b7005e6fdd
57 changed files with 7257 additions and 474 deletions
+21 -18
View File
@@ -22,6 +22,7 @@ import re
import pytest
from app import models, worldstate
from app import narrative
from app.context import builder
from app.database import Base, SessionLocal, engine
@@ -75,20 +76,20 @@ def with_schema(db, scenario_id, adventure):
def test_budget_is_the_cap_minus_headroom_and_buffer_in_words():
"""800-token cap → 750 after headroom → ~562 words → 506 after the buffer."""
hint = builder.length_hint(800, has_ws=True)
hint = builder.length_hint(800)
assert "506" in hint
assert "token" not in hint.lower(), "a model cannot count its own tokens"
def test_budget_tracks_the_setting():
small = builder.length_hint(800, has_ws=True)
large = builder.length_hint(2400, has_ws=True)
small = builder.length_hint(800)
large = builder.length_hint(2400)
assert small != large
assert "1586" in large
def asked_words(cap):
return int(re.search(r"(\d+) words", builder.length_hint(cap, has_ws=True)).group(1))
return int(re.search(r"(\d+) words", builder.length_hint(cap)).group(1))
def test_buffer_leaves_room_for_overshoot():
@@ -106,7 +107,7 @@ def test_hint_is_phrased_as_a_ceiling_not_a_budget():
reads as a target to fill. It moved the mean turn from 174 to 246
words, toward the limit it exists to avoid. The limit framing must
survive future prompt edits."""
hint = builder.length_hint(800, has_ws=True)
hint = builder.length_hint(800)
assert "must not exceed" in hint
assert "under about" not in hint
assert "lower end" in hint, "without this the number still reads as a target"
@@ -117,7 +118,7 @@ def test_hint_states_a_floor_as_well_as_a_ceiling():
the "only as much as the moment needs" clause and produces only two
paragraphs. The floor is what makes the same prompt produce a similar
length across models with different tendencies."""
hint = builder.length_hint(800, has_ws=True)
hint = builder.length_hint(800)
assert "506" in hint and "177" in hint
assert "should not stop short of" in hint
# Asymmetric on purpose: the ceiling is a hard limit and the floor is a
@@ -127,7 +128,7 @@ def test_hint_states_a_floor_as_well_as_a_ceiling():
def test_floor_stays_well_under_the_ceiling():
for cap in (400, 800, 1500, 2400):
hint = builder.length_hint(cap, has_ws=True)
hint = builder.length_hint(cap)
ceiling, floor = (int(n) for n in re.findall(r"(\d+)", hint)[:2])
assert floor < ceiling * 0.5
@@ -137,7 +138,7 @@ def test_floor_is_dropped_when_the_cap_is_too_tight_for_one():
wording is the one measured to keep the state block from being
truncated (0/6 truncations at cap 250, against 2/6 unhinted), so it
is left exactly as it was."""
hint = builder.length_hint(250, has_ws=True)
hint = builder.length_hint(250)
assert "should not stop short of" not in hint
assert "much shorter" in hint
@@ -145,26 +146,28 @@ def test_floor_is_dropped_when_the_cap_is_too_tight_for_one():
def test_floor_does_not_grow_without_bound():
"""A big cap means "long turns are allowed", not "every turn must be an essay":
the share alone would demand 555 words minimum at cap 2400."""
hint = builder.length_hint(2400, has_ws=True)
hint = builder.length_hint(2400)
assert str(builder.MAX_LENGTH_FLOOR_WORDS) in hint
def test_no_hint_when_the_cap_is_too_small_to_phrase():
"""Under the floor the hint is noise the model pays for in context."""
assert builder.length_hint(100, has_ws=True) == ""
assert builder.length_hint(builder.LENGTH_HEADROOM, has_ws=True) == ""
assert builder.length_hint(0, has_ws=True) == ""
assert builder.length_hint(100) == ""
assert builder.length_hint(builder.LENGTH_HEADROOM) == ""
assert builder.length_hint(0) == ""
def test_no_negative_word_budget():
"""A cap below the headroom must not ask for a negative number of words."""
for cap in (1, 10, 49, 51):
assert builder.length_hint(cap, has_ws=True) == ""
assert builder.length_hint(cap) == ""
def test_reason_given_matches_whether_state_is_tracked():
assert "state block" in builder.length_hint(800, has_ws=True)
assert "state block" not in builder.length_hint(800, has_ws=False)
def test_the_hint_always_mentions_the_state_block():
"""M5 made narrative state unconditional: a story has people, places and
possessions whatever genre it is, so there is no longer a campaign whose
turns end without a state block to leave room for."""
assert "state block" in builder.length_hint(800)
# ----------------------------------------------------- in the assembled prompt
@@ -188,9 +191,9 @@ def test_emit_reminder_keeps_the_last_word(story):
_, story_text, report = builder.build_context(adventure, settings)
assert story_text.rstrip().endswith(worldstate.EMIT_REMINDER.rstrip())
assert story_text.rstrip().endswith(narrative.extract.EMIT_REMINDER.rstrip())
labels = [s["label"] for s in report["sections"]]
assert labels.index("length_hint") < labels.index("world_state_reminder")
assert labels.index("length_hint") < labels.index("state_reminder")
def test_prompt_stays_inside_the_budget_on_a_long_story(story):