Files
interactive-story/backend/tests/fakes.py
T
JesseMarkowitzandClaude Opus 5 b7005e6fdd M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.

This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:

    visible active transcript position == stored head == authoritative state

Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)

  A narrator edit no longer rewrites a row. It returns to the state before the
  turn, takes the reader's exact text as the accepted narration, re-derives the
  state that text implies, and becomes a new active continuation — while the
  original narration keeps its words, its live flag and its whole future as
  retained history. At the tip the correction is another take; with story below
  it, it forks. No new history machinery: this is the existing fork/take/head
  path with the reader's text in place of a generated reply. The §14A refusal
  is therefore gone for narrator turns, and remains only for player input.

Pre-M5 positions

  Migration 88 backfills the empty narrative document onto every action written
  before M5, and a missing snapshot now restores the empty document instead of
  leaving the previous position's state standing. Restoring to an old Save
  Point no longer leaves a later position's entities and facts on screen.

Narrator context

  Replayed history carries prose only; the machine-readable block is no longer
  reconstructed into past turns, where it contradicted the authoritative state
  in the same prompt. A fact withdrawn by a manual correction is now named as
  no longer true, with the reader's reason, rather than silently dropped.

Also

  - state_changes joins the action-list bulk read, removing one query per row.
  - Extraction takes only the application's own protocol payload: an ordinary
    ```json or ```python block in a story survives, and a mangled proposal
    still does not reach the reader.

Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-05 07:01:50 -04:00

159 lines
6.6 KiB
Python

"""Stand-ins for the parts of the app a test must not really call.
Import these rather than writing another copy. Nine test modules each carried
their own `ScriptedProvider`, and the copies had drifted into four different
feature sets, so a test that needed to raise a provider error had to be written
in one of the files whose copy supported it.
"""
import json
class ScriptedProvider:
"""Streams canned replies in place of `OpenAICompatibleProvider`.
Set `replies` to the texts the model returns, one per call. The last entry
repeats once the list runs out, so a test that plays more turns than it
scripted still gets text. To drive the provider-error path, put an
`Exception` in the list. It is raised rather than streamed.
State lives on the class, not on the instance, because the turn engine
constructs the provider itself and a test never sees the object. The autouse
`reset_scripted_provider` fixture in `conftest.py` clears it between tests.
`prompts` records every assembled `(system, story)` pair, which is what a
test asserts on to check what the model was shown.
"""
last_usage = None
replies: list = []
calls = 0
prompts: list = []
def __init__(self, *a, **k):
pass
async def generate(self, parts, *, temperature, max_tokens):
index = min(ScriptedProvider.calls, len(ScriptedProvider.replies) - 1)
ScriptedProvider.calls += 1
ScriptedProvider.prompts.append((parts.system, parts.story))
reply = ScriptedProvider.replies[index]
if isinstance(reply, Exception):
raise reply
yield ("text", reply)
# ---------------------------------------------------------------------------
# Deterministic per-turn state instrumentation
# ---------------------------------------------------------------------------
# Several tests need a value that changes by a fixed amount on every turn, so
# that a rollback failure is arithmetic rather than a judgement call: if a take
# stacks instead of replacing, the total is off by exactly one turn's worth.
#
# The instrument has moved twice, and both moves were the same move: it follows
# whatever the production state path is, so the tests exercise real code rather
# than a test hook. It began as a QuickJS `state.gold += 10` (removed with
# scripting in M2), became an RPG world-state delta block (M3/M4), and is now a
# typed narrative-state event (M5).
#
# What the tests using it measure is unchanged, and worth restating because it
# is why they were re-instrumented rather than deleted: the state at a story
# position, rollback, Redo restoration, retry, alternate takes, divergence,
# abandoned-future isolation, and Save Point restore. None of that was ever
# about gold, or about RPG stats.
#
# The M5 instrument is deliberately genre-neutral: a `chronicle` entity — a
# concept, not a character, not an item — carrying one named attribute. Every
# reply sets it to an ABSOLUTE total, which is ADR 010's whole point. A delta
# protocol could not tell "+10" from "= 10"; here the event type says which, so
# `TALLY_PER_TURN * n` after n turns is arithmetic rather than an assumption.
#: The instrument is a fact, not an entity attribute, and deliberately so.
#: `set_entity_attribute` names an entity that must already exist, which is the
#: right rule for the product and the wrong one for an instrument that tests
#: script in isolation — a one-off reply in the middle of a test would be
#: refused for a reference the test never meant to be about. `add_fact` needs no
#: subject, so any reply can state the tally on its own. Entity creation,
#: possession and the referential rule get their own tests in
#: `test_narrative_state.py`, where they are the subject rather than scaffolding.
TALLY_PREDICATE = "tally"
TALLY_PER_TURN = 10
# Kept as an alias so the many tests that speak in these terms keep reading
# naturally. The number is the same; only the protocol underneath changed.
GOLD_PER_TURN = TALLY_PER_TURN
#: A scenario schema is no longer needed for state to work — narrative state is
#: not an opt-in RPG layer. The name survives for fixtures that still pass
#: something, and empty is the honest value: this campaign has no RPG layer, and
#: under M5 it does not need one to have state.
GOLD_SCHEMA: dict = {}
def state_block(events: list) -> str:
"""The fenced block the model is asked to emit, around `events`."""
return "```state\n" + json.dumps({"events": events}, ensure_ascii=False) + "\n```"
def tally_reply(text: str, total: int) -> str:
"""A reply that narrates `text` and records the tally as `total`.
Absolute, always — which is the whole of ADR 010. A delta protocol could not
tell "+10" from "= 10"; here the event says which, so `TALLY_PER_TURN * n`
after n turns is arithmetic rather than an assumption, and a replayed or
duplicated reply cannot silently double it.
Each reply supersedes the last, so the newest active tally fact is the
current one and the document does not grow without bound.
"""
return f"{text}\n" + state_block([{
"type": "add_fact",
"predicate": TALLY_PREDICATE,
"value": total,
"fact_id": f"tally-{total}",
}])
def gold_reply(text: str, amount: int = TALLY_PER_TURN) -> str:
"""One reply banking `amount`, for tests that build a single reply.
The value is absolute underneath, so a caller asking for the default gets
the first turn's total, which is what those call sites mean.
"""
return tally_reply(text, amount)
def tally_replies(prefix: str = "Take", count: int = 40) -> list:
"""`count` numbered replies whose tally runs 10, 20, 30 …"""
return [
tally_reply(f"{prefix} {n}.", n * TALLY_PER_TURN)
for n in range(1, count + 1)
]
#: The historical name, unchanged in meaning for every caller.
gold_replies = tally_replies
def tally_of(state) -> int:
"""Reads the instrument back out of a narrative state document.
The newest active tally fact wins, which is what "absolute assignment"
means when the assignments are appended. Returns 0 when the campaign has
recorded none — what "no turns have been played" means, and what a restore
to before the first turn should produce.
"""
if not isinstance(state, dict):
return 0
facts = state.get("facts")
if not isinstance(facts, list):
return 0
for fact in reversed(facts):
if (
isinstance(fact, dict)
and fact.get("predicate") == TALLY_PREDICATE
and fact.get("status", "active") == "active"
and isinstance(fact.get("value"), (int, float))
):
return fact["value"]
return 0