Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.
This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:
visible active transcript position == stored head == authoritative state
Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)
A narrator edit no longer rewrites a row. It returns to the state before the
turn, takes the reader's exact text as the accepted narration, re-derives the
state that text implies, and becomes a new active continuation — while the
original narration keeps its words, its live flag and its whole future as
retained history. At the tip the correction is another take; with story below
it, it forks. No new history machinery: this is the existing fork/take/head
path with the reader's text in place of a generated reply. The §14A refusal
is therefore gone for narrator turns, and remains only for player input.
Pre-M5 positions
Migration 88 backfills the empty narrative document onto every action written
before M5, and a missing snapshot now restores the empty document instead of
leaving the previous position's state standing. Restoring to an old Save
Point no longer leaves a later position's entities and facts on screen.
Narrator context
Replayed history carries prose only; the machine-readable block is no longer
reconstructed into past turns, where it contradicted the authoritative state
in the same prompt. A fact withdrawn by a manual correction is now named as
no longer true, with the reader's reason, rather than silently dropped.
Also
- state_changes joins the action-list bulk read, removing one query per row.
- Extraction takes only the application's own protocol payload: an ordinary
```json or ```python block in a story survives, and a mangled proposal
still does not reach the reader.
Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
159 lines
6.6 KiB
Python
159 lines
6.6 KiB
Python
"""Stand-ins for the parts of the app a test must not really call.
|
|
|
|
Import these rather than writing another copy. Nine test modules each carried
|
|
their own `ScriptedProvider`, and the copies had drifted into four different
|
|
feature sets, so a test that needed to raise a provider error had to be written
|
|
in one of the files whose copy supported it.
|
|
"""
|
|
import json
|
|
|
|
|
|
class ScriptedProvider:
|
|
"""Streams canned replies in place of `OpenAICompatibleProvider`.
|
|
|
|
Set `replies` to the texts the model returns, one per call. The last entry
|
|
repeats once the list runs out, so a test that plays more turns than it
|
|
scripted still gets text. To drive the provider-error path, put an
|
|
`Exception` in the list. It is raised rather than streamed.
|
|
|
|
State lives on the class, not on the instance, because the turn engine
|
|
constructs the provider itself and a test never sees the object. The autouse
|
|
`reset_scripted_provider` fixture in `conftest.py` clears it between tests.
|
|
|
|
`prompts` records every assembled `(system, story)` pair, which is what a
|
|
test asserts on to check what the model was shown.
|
|
"""
|
|
|
|
last_usage = None
|
|
replies: list = []
|
|
calls = 0
|
|
prompts: list = []
|
|
|
|
def __init__(self, *a, **k):
|
|
pass
|
|
|
|
async def generate(self, parts, *, temperature, max_tokens):
|
|
index = min(ScriptedProvider.calls, len(ScriptedProvider.replies) - 1)
|
|
ScriptedProvider.calls += 1
|
|
ScriptedProvider.prompts.append((parts.system, parts.story))
|
|
reply = ScriptedProvider.replies[index]
|
|
if isinstance(reply, Exception):
|
|
raise reply
|
|
yield ("text", reply)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Deterministic per-turn state instrumentation
|
|
# ---------------------------------------------------------------------------
|
|
# Several tests need a value that changes by a fixed amount on every turn, so
|
|
# that a rollback failure is arithmetic rather than a judgement call: if a take
|
|
# stacks instead of replacing, the total is off by exactly one turn's worth.
|
|
#
|
|
# The instrument has moved twice, and both moves were the same move: it follows
|
|
# whatever the production state path is, so the tests exercise real code rather
|
|
# than a test hook. It began as a QuickJS `state.gold += 10` (removed with
|
|
# scripting in M2), became an RPG world-state delta block (M3/M4), and is now a
|
|
# typed narrative-state event (M5).
|
|
#
|
|
# What the tests using it measure is unchanged, and worth restating because it
|
|
# is why they were re-instrumented rather than deleted: the state at a story
|
|
# position, rollback, Redo restoration, retry, alternate takes, divergence,
|
|
# abandoned-future isolation, and Save Point restore. None of that was ever
|
|
# about gold, or about RPG stats.
|
|
#
|
|
# The M5 instrument is deliberately genre-neutral: a `chronicle` entity — a
|
|
# concept, not a character, not an item — carrying one named attribute. Every
|
|
# reply sets it to an ABSOLUTE total, which is ADR 010's whole point. A delta
|
|
# protocol could not tell "+10" from "= 10"; here the event type says which, so
|
|
# `TALLY_PER_TURN * n` after n turns is arithmetic rather than an assumption.
|
|
|
|
#: The instrument is a fact, not an entity attribute, and deliberately so.
|
|
#: `set_entity_attribute` names an entity that must already exist, which is the
|
|
#: right rule for the product and the wrong one for an instrument that tests
|
|
#: script in isolation — a one-off reply in the middle of a test would be
|
|
#: refused for a reference the test never meant to be about. `add_fact` needs no
|
|
#: subject, so any reply can state the tally on its own. Entity creation,
|
|
#: possession and the referential rule get their own tests in
|
|
#: `test_narrative_state.py`, where they are the subject rather than scaffolding.
|
|
TALLY_PREDICATE = "tally"
|
|
TALLY_PER_TURN = 10
|
|
|
|
# Kept as an alias so the many tests that speak in these terms keep reading
|
|
# naturally. The number is the same; only the protocol underneath changed.
|
|
GOLD_PER_TURN = TALLY_PER_TURN
|
|
|
|
#: A scenario schema is no longer needed for state to work — narrative state is
|
|
#: not an opt-in RPG layer. The name survives for fixtures that still pass
|
|
#: something, and empty is the honest value: this campaign has no RPG layer, and
|
|
#: under M5 it does not need one to have state.
|
|
GOLD_SCHEMA: dict = {}
|
|
|
|
|
|
def state_block(events: list) -> str:
|
|
"""The fenced block the model is asked to emit, around `events`."""
|
|
return "```state\n" + json.dumps({"events": events}, ensure_ascii=False) + "\n```"
|
|
|
|
|
|
def tally_reply(text: str, total: int) -> str:
|
|
"""A reply that narrates `text` and records the tally as `total`.
|
|
|
|
Absolute, always — which is the whole of ADR 010. A delta protocol could not
|
|
tell "+10" from "= 10"; here the event says which, so `TALLY_PER_TURN * n`
|
|
after n turns is arithmetic rather than an assumption, and a replayed or
|
|
duplicated reply cannot silently double it.
|
|
|
|
Each reply supersedes the last, so the newest active tally fact is the
|
|
current one and the document does not grow without bound.
|
|
"""
|
|
return f"{text}\n" + state_block([{
|
|
"type": "add_fact",
|
|
"predicate": TALLY_PREDICATE,
|
|
"value": total,
|
|
"fact_id": f"tally-{total}",
|
|
}])
|
|
|
|
|
|
def gold_reply(text: str, amount: int = TALLY_PER_TURN) -> str:
|
|
"""One reply banking `amount`, for tests that build a single reply.
|
|
|
|
The value is absolute underneath, so a caller asking for the default gets
|
|
the first turn's total, which is what those call sites mean.
|
|
"""
|
|
return tally_reply(text, amount)
|
|
|
|
|
|
def tally_replies(prefix: str = "Take", count: int = 40) -> list:
|
|
"""`count` numbered replies whose tally runs 10, 20, 30 …"""
|
|
return [
|
|
tally_reply(f"{prefix} {n}.", n * TALLY_PER_TURN)
|
|
for n in range(1, count + 1)
|
|
]
|
|
|
|
|
|
#: The historical name, unchanged in meaning for every caller.
|
|
gold_replies = tally_replies
|
|
|
|
|
|
def tally_of(state) -> int:
|
|
"""Reads the instrument back out of a narrative state document.
|
|
|
|
The newest active tally fact wins, which is what "absolute assignment"
|
|
means when the assignments are appended. Returns 0 when the campaign has
|
|
recorded none — what "no turns have been played" means, and what a restore
|
|
to before the first turn should produce.
|
|
"""
|
|
if not isinstance(state, dict):
|
|
return 0
|
|
facts = state.get("facts")
|
|
if not isinstance(facts, list):
|
|
return 0
|
|
for fact in reversed(facts):
|
|
if (
|
|
isinstance(fact, dict)
|
|
and fact.get("predicate") == TALLY_PREDICATE
|
|
and fact.get("status", "active") == "active"
|
|
and isinstance(fact.get("value"), (int, float))
|
|
):
|
|
return fact["value"]
|
|
return 0
|