M5: genre-neutral authoritative narrative state, with review corrections

Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.

This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:

    visible active transcript position == stored head == authoritative state

Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)

  A narrator edit no longer rewrites a row. It returns to the state before the
  turn, takes the reader's exact text as the accepted narration, re-derives the
  state that text implies, and becomes a new active continuation — while the
  original narration keeps its words, its live flag and its whole future as
  retained history. At the tip the correction is another take; with story below
  it, it forks. No new history machinery: this is the existing fork/take/head
  path with the reader's text in place of a generated reply. The §14A refusal
  is therefore gone for narrator turns, and remains only for player input.

Pre-M5 positions

  Migration 88 backfills the empty narrative document onto every action written
  before M5, and a missing snapshot now restores the empty document instead of
  leaving the previous position's state standing. Restoring to an old Save
  Point no longer leaves a later position's entities and facts on screen.

Narrator context

  Replayed history carries prose only; the machine-readable block is no longer
  reconstructed into past turns, where it contradicted the authoritative state
  in the same prompt. A fact withdrawn by a manual correction is now named as
  no longer true, with the reader's reason, rather than silently dropped.

Also

  - state_changes joins the action-list bulk read, removing one query per row.
  - Extraction takes only the application's own protocol payload: an ordinary
    ```json or ```python block in a story survives, and a mangled proposal
    still does not reach the reader.

Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-05 07:01:50 -04:00
co-authored by Claude Opus 5
parent 62a997f364
commit b7005e6fdd
57 changed files with 7257 additions and 474 deletions
+82 -46
View File
@@ -1,6 +1,32 @@
"""End-to-end HTTP test for RPG world state (Phase 12): a scenario with a
stat_schema, a turn whose (faked) AI reply carries a state delta block, and
undo rolling the world state back.
"""M5: the RPG world state is demoted, and this file is what holds it there.
This was the end-to-end test for Phase 12's RPG world state: a scenario with a
`stat_schema`, a turn whose reply carried a relative-delta block, the referee
applying it, and undo rolling it back.
**M5 removed that pipeline from the turn path** (ADR 010). Relative deltas were
ambiguous by construction — a number in a delta field is syntactically legal
whether the model meant "add 50" or "set to 50" — and the genre-neutral typed
events in `app/narrative/` replaced them. The engine in `app/worldstate/` is
retained so a pre-M5 database opens unchanged and its numbers stay readable, and
its own logic is still covered by `test_worldstate.py`; what no longer exists is
its authority over new turns.
So this file now tests the demotion, which is worth a test precisely because it
is an absence: nothing else would notice if the delta pipeline were quietly
rewired, and a second authoritative state engine is the specific failure the M5
brief names.
What was deleted, and why, rather than being re-instrumented:
* `test_turn_applies_clamped_delta_and_strips_block`, `test_action_world_changes_summary`,
`test_world_state_endpoint`, `test_undo_reverts_world_state`,
`test_retry_does_not_double_apply` — each asserted that a turn applies a
delta through the referee. That is the removed capability itself, not
instrumentation for something else. The invariants they *also* touched —
undo restoring state, a retry not double-applying — are alive and better
covered in `test_head_cursor.py`, `test_take_state.py` and
`test_narrative_state.py`, now measured through the production state path.
python -m pytest tests/test_worldstate_integration.py -v
"""
@@ -103,47 +129,74 @@ def _play(client, text="attack the goblin"):
return r
def test_turn_applies_clamped_delta_and_strips_block(client):
def test_a_turn_no_longer_applies_a_relative_delta_block(client):
"""The demotion, asserted.
The fixture's reply carries the old protocol. Under Phase 12 this moved hp
from 100 to 70. Under M5 the referee is not in the turn path at all, so the
numbers stay where they were — and the block still leaves the prose, because
a reader must never see the protocol whichever protocol it is.
"""
_play(client)
ws = _world(client.adv_id)
assert ws["player"]["hp"] == 70 # -80 capped to -30
assert ws["npc"]["gwen"]["trust"] == 15
assert ws["flags"]["alarm"] is True
assert ws["milestones"]["win"]["reached"] is True
# The state block is not shown to the player.
assert "```state" not in _last_ai_text(client.adv_id)
assert "goblin's blade" in _last_ai_text(client.adv_id)
# The raw model reply, including the block, is kept for the Insights view.
world = _world(client.adv_id)
assert world["player"]["hp"] == 100, "the delta pipeline is still wired into turns"
assert world.get("milestones", {}) == {}
assert world.get("flags", {}) in ({}, {"alarm": False})
text = _last_ai_text(client.adv_id)
assert "```state" not in text, "the protocol reached the reader"
assert "goblin's blade" in text
# The raw reply is still kept, so the Insights view and the audit can show
# what the model actually sent.
db = SessionLocal()
try:
snap = db.get(models.Adventure, client.adv_id).actions[-1].context_snapshot
snapshot = db.get(models.Adventure, client.adv_id).actions[-1].context_snapshot
finally:
db.close()
assert "```state" in snap["raw_output"]
assert '"player.hp": -80' in snap["raw_output"]
assert '"player.hp": -80' in snapshot["raw_output"]
def test_action_world_changes_summary(client):
def test_an_old_style_block_proposes_no_narrative_events(client):
"""A delta block is not a typed proposal, and must not be read as one.
`{"player.hp": -80}` is a JSON object with no `events` list. It parses, so
it is not malformed; it simply proposes nothing. What matters is that no
part of it is coerced into an event — a state engine that guessed here would
reintroduce exactly the ambiguity ADR 010 removed.
"""
_play(client)
db = SessionLocal()
try:
changes = db.get(models.Adventure, client.adv_id).actions[-1].world_changes
proposal = (
db.query(models.StateProposal)
.filter_by(adventure_id=client.adv_id)
.order_by(models.StateProposal.id.desc())
.first()
)
assert proposal is not None, "no proposal was recorded for the turn"
assert db.query(models.StateEvent).filter_by(
adventure_id=client.adv_id).count() == 0
finally:
db.close()
by_label = {c["label"]: c for c in changes}
assert by_label["hp"]["delta"] == -30 # clamped stat, signed delta
assert by_label["gwen trust"]["delta"] == 15 # npc.<id>.<stat> -> "id stat"
assert by_label["alarm"] == {"kind": "flag", "label": "alarm", "on": True}
assert by_label["win"]["kind"] == "milestone"
def test_world_state_endpoint(client):
_play(client)
r = client.get(f"/api/adventures/{client.adv_id}/world-state")
def test_the_narrative_state_is_what_the_turn_now_establishes(client):
"""And the replacement works on the same campaign, in the same turn shape."""
ScriptedProvider.replies = [
'The blade turns aside.\n```state\n{"events": ['
'{"type": "create_entity", "entity": "gwen", "entity_type": "character",'
' "name": "Gwen"},'
'{"type": "set_entity_status", "entity": "gwen", "status": "wary"}'
']}\n```'
]
_play(client, "parry")
r = client.get(f"/api/adventures/{client.adv_id}/state")
assert r.status_code == 200, r.text
body = r.json()
assert body["schema"]["player"]["hp"]["max"] == 100
assert body["state"]["player"]["hp"] == 70
document = r.json()["document"]
assert document["entities"]["gwen"]["status"] == "wary"
def test_override_world_state_endpoint(client):
@@ -160,20 +213,3 @@ def test_override_world_state_endpoint(client):
# This bypasses max_delta_per_turn (30) because it is a direct correction, not a turn.
r = client.put(f"/api/adventures/{client.adv_id}/world-state", json={"player.hp": 100})
assert r.json()["state"]["player"]["hp"] == 100
def test_undo_reverts_world_state(client):
_play(client)
assert _world(client.adv_id)["player"]["hp"] == 70
r = client.post(f"/api/adventures/{client.adv_id}/undo")
assert r.status_code == 200, r.text
assert _world(client.adv_id)["player"]["hp"] == 100 # back to initial
assert _world(client.adv_id)["milestones"] == {}
def test_retry_does_not_double_apply(client):
_play(client)
assert _world(client.adv_id)["player"]["hp"] == 70
r = client.post(f"/api/adventures/{client.adv_id}/retry")
assert r.status_code == 200, r.text
assert _world(client.adv_id)["player"]["hp"] == 70 # not 40