M9: a campaign you can actually get back

A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
JesseMarkowitz
2026-09-07 01:55:45 -04:00
co-authored by Claude Opus 5
parent 1ce9972760
commit 44edece67e
46 changed files with 9227 additions and 178 deletions
+406
View File
@@ -0,0 +1,406 @@
"""M9: one campaign that exercises every portable data family at once.
`TEST-CAMPAIGN-FIXTURE.md` describes the standard Continuity Test campaign, and
the acceptance suites use it. This is a different thing and does not replace it:
the Continuity Test is shaped to read like a story, and this one is shaped to
break a round trip. Every property M9 promises has a source in this campaign that
would be silently lost by a plausible mistake in the exporter or the importer.
Opening
|
+-- normal turns transcript, state events, snapshots
+-- Retry two takes at one coordinate
+-- knowledge retrieval imported passages in a stored prompt
+-- Save Point S1 a named coordinate on the first line
+-- more turns a future the reader will leave
|
+-- Undo x2 the head steps back
|
+-- divergent continuation a second branch, and a second future
+-- Save Point S2 a named coordinate on the second line
+-- manual state correction an event nothing narrated
+-- Undo x1 the head ends behind the newest row
The shape is chosen so that no single fact identifies a position. The active head
is not the newest row, not the deepest row, not the last row written, and not on
the branch that holds the most story — an importer that guesses any one of those
lands somewhere else.
Two campaigns are built, not one. `build` returns the rich campaign; the fixture
also leaves a neighbour beside it, because a bundle that accidentally exported
another campaign's rows would otherwise export nothing and pass.
The builder speaks HTTP throughout. A fixture that wrote rows directly would
prove the exporter can read what the fixture wrote, which is not the claim.
"""
from __future__ import annotations
import asyncio
from app import memorybank
from fakes import ScriptedProvider, state_block
# --------------------------------------------------------------- source files
# Three imported sources, one per class, plus the two lifecycle states that a
# round trip most easily loses: a source someone switched off, and one only the
# narrator may see.
CANON_MD = """# Westhaven
## The Old Abbey
The abbey above Westhaven has stood since the founding. Its crypt is sealed,
and the seal has never been broken.
## What cannot happen here
The dead do not return. No rite, relic or bargain in Westhaven has ever
returned anyone from death, and none ever will.
"""
REFERENCE_MD = """# The Crooked Lantern
The tavern on Fen Street is timber-framed, low-beamed, and older than the
street it stands on. The hearth is never allowed to go out.
## The keeper
Mara keeps the Crooked Lantern. She was born in Westhaven and has never left
it.
"""
INSPIRATION_MD = """# Weather notes
Rain on shutters. Lantern light through wet glass. The smell of a hearth
banked for the night.
"""
SECRET_MD = """# The seal
The abbey seal was broken once, sixty years ago, and set again by a hand that
is still alive. Nobody in Westhaven knows this.
"""
DISABLED_MD = """# Discarded draft
An earlier draft of the Westhaven material, kept for reference and switched off
so it cannot reach the narrator.
"""
#: The campaign's own rule, so the correction and the canon block have something
#: real to be measured against.
CAMPAIGN_CANON = {"rules": ["The dead do not return."]}
OPENING = "Aldric sits in the Crooked Lantern with Mara, and the rain starts."
# ------------------------------------------------------------------- helpers
def _play(client, adv_id, text, prose, events=None, kind="do"):
ScriptedProvider.replies = [f"{prose}\n{state_block(events or [])}"]
response = client.post(
f"/api/adventures/{adv_id}/actions", json={"type": kind, "text": text}
)
assert response.status_code == 200, response.text[:400]
return response
def _fact(predicate, value, fact_id):
return {"type": "add_fact", "predicate": predicate, "value": value,
"fact_id": fact_id}
def upload(client, adv_id, name, body, classification, **fields):
"""Imports a file the way the browser does: multipart, and no pathname."""
data = {"classification": classification}
data.update({k: str(v).lower() if isinstance(v, bool) else str(v)
for k, v in fields.items()})
response = client.post(
f"/api/adventures/{adv_id}/knowledge",
files={"file": (name, body.encode("utf-8"), "text/markdown")},
data=data,
)
assert response.status_code == 201, response.text[:400]
return response.json()["id"]
def _checkpoint(client, adv_id, name, note=""):
response = client.post(
f"/api/adventures/{adv_id}/checkpoints", json={"name": name, "note": note}
)
assert response.status_code == 201, response.text[:400]
return response.json()
def _undo(client, adv_id, times=1):
for _ in range(times):
response = client.post(f"/api/adventures/{adv_id}/undo")
assert response.status_code == 200, response.text[:400]
def settle_derived(adv_id):
"""Runs the background memory and summary pass to completion.
The turn endpoint fires this as a fire-and-forget task, which a test client
does not wait for. Calling it directly is the same code on the same rows —
what is skipped is the scheduling, not the work — and it is what
`test_context_realistic.py` does for the same reason.
"""
asyncio.run(memorybank.run_post_turn(adv_id))
# --------------------------------------------------------------------- build
def build(client, adv_id) -> dict:
"""Plays the fixture campaign onto `adv_id`, and returns what it built.
The returned dictionary is the assertion source for every round-trip test:
it names the properties that must survive, measured from the campaign as it
stands here rather than restated as constants, so a test compares the copy
against the original instead of against a guess about the original.
"""
# Story memory and the rolling summary on, because a campaign that
# generated neither would let an exporter omit both and still pass. The
# abandoned line below gets long enough to earn its own, which is what E03
# is about after a round trip.
switched_on = client.patch(
f"/api/adventures/{adv_id}",
json={"auto_summarize": True, "memory_bank_enabled": True},
)
assert switched_on.status_code == 200, switched_on.text[:400]
sources = {
"canon": upload(client, adv_id, "canon.md", CANON_MD, "canon",
always_include=True),
"reference": upload(client, adv_id, "reference.md", REFERENCE_MD,
"reference"),
"inspiration": upload(client, adv_id, "inspiration.md", INSPIRATION_MD,
"inspiration"),
"secret": upload(client, adv_id, "secret.md", SECRET_MD, "canon",
visibility="hidden"),
"disabled": upload(client, adv_id, "draft.md", DISABLED_MD, "reference"),
}
disable = client.patch(
f"/api/adventures/{adv_id}/knowledge/{sources['disabled']}",
json={"enabled": False},
)
assert disable.status_code == 200, disable.text[:400]
# ---- the first line of story -----------------------------------------
# Turn 1 asks about the abbey, so the canon source is retrieved and the
# stored prompt for this turn holds an imported passage. That turn is the
# one the provenance tests read back after the round trip.
_play(client, adv_id, "ask Mara about the abbey",
"Mara sets down the cloth. The abbey, she says, is sealed.",
[_fact("tally", 10, "tally-10")])
_play(client, adv_id, "walk up to the abbey",
"The path climbs out of the town and the rain follows.",
[_fact("tally", 20, "tally-20")])
# A retry, so one coordinate holds two takes and the earlier one is
# retained but not selected.
ScriptedProvider.replies = [
"The door is oak, and the seal on it is unbroken.\n"
+ state_block([_fact("tally", 30, "tally-30")])
]
_play(client, adv_id, "try the crypt door",
"The door will not move.", [_fact("tally", 30, "tally-30")])
retry = client.post(f"/api/adventures/{adv_id}/retry")
assert retry.status_code == 200, retry.text[:400]
s1 = _checkpoint(client, adv_id, "At the crypt door",
"Before anything is decided.")
# The future the reader is about to leave behind. It is played out far
# enough to earn derived data of its own — `memorybank.MEMORY_INTERVAL` is
# six actions — because a summary and a memory belonging to an abandoned
# line are what E03 forbids reaching an active prompt, and a round trip is
# a new way to leak one.
_play(client, adv_id, "force the door",
"The seal gives, and the stair below is dark.",
[_fact("tally", 40, "tally-40")])
_play(client, adv_id, "go down",
"The crypt is dry, and the air has not moved in years.",
[_fact("tally", 50, "tally-50")])
_play(client, adv_id, "read the names on the slabs",
"Sixty years of Westhaven dead, and one slab with no name at all.",
[_fact("tally", 60, "tally-60")])
_play(client, adv_id, "touch the nameless slab",
"The stone is warm, which stone in a crypt is not.",
[_fact("tally", 70, "tally-70")])
# Derived data for the line that is about to be abandoned, written while
# the head is still on it. This is the summary and the memory that must
# come back after a round trip and must still be ineligible there.
settle_derived(adv_id)
tip_state = client.get(f"/api/adventures/{adv_id}/state").json()
# ---- step back, and go somewhere else ---------------------------------
_undo(client, adv_id, 4)
_play(client, adv_id, "turn back and return to the tavern",
"The rain has not let up, and the Lantern's windows are lit.",
[_fact("tally", 41, "tally-41")])
s2 = _checkpoint(client, adv_id, "Back at the Lantern", "The other way.")
_play(client, adv_id, "ask Mara what she is not saying",
"She looks at the fire for a while before she answers.",
[_fact("tally", 51, "tally-51")])
_play(client, adv_id, "wait",
"The rain fills the silence, and then she starts talking.",
[_fact("tally", 61, "tally-61")])
# A manual correction: an accepted state change with no narration behind
# it, which is the one kind of state event a replay could never recreate.
correction = client.post(
f"/api/adventures/{adv_id}/state/corrections",
json={
"events": [{
"type": "add_fact",
"predicate": "keeper_of_the_lantern",
"value": "Mara",
"fact_id": "keeper",
}],
"note": "Established in play before the state system saw it.",
},
)
assert correction.status_code == 201, correction.text[:400]
# Derived data for the line the reader stayed on, so the copy has both an
# eligible and an ineligible summary to tell apart. The generated one landed
# on the abandoned line, which is the E03 case; this one is typed at the
# current head, so it is the eligible case beside it. A round trip has to
# keep them on opposite sides of that line.
settle_derived(adv_id)
# One more Undo, so the head finishes behind the retained tip of its own
# branch as well as behind the abandoned line's.
_undo(client, adv_id, 1)
# Typed at the final head, so it is the eligible summary and the generated
# one on the abandoned line is not. A round trip has to keep them on
# opposite sides of that line.
typed = client.patch(
f"/api/adventures/{adv_id}",
json={"story_summary": "Aldric went back to the Lantern instead."},
)
assert typed.status_code == 200, typed.text[:400]
return snapshot_of(client, adv_id, sources=sources, s1=s1, s2=s2,
tip_state=tip_state)
def snapshot_in(action: dict) -> dict | None:
"""The stored prompt in one bundle entry, decoded.
The export compresses it (`bundle._packed`), so a test that reached for a
plain dict would conclude the evidence was missing when it is merely
encoded. Both keys are read, plain first, exactly as the importer does.
"""
from app import bundle
plain = action.get("contextSnapshot")
if isinstance(plain, dict):
return plain
return bundle._unpacked(action.get("contextSnapshotZ"))
def with_snapshot(action: dict, snapshot: dict | None) -> dict:
"""A bundle entry carrying `snapshot`, written in the plain form.
Tests that break a snapshot on purpose write the readable key, because the
importer prefers it and because a test that had to compress its own fixture
would be testing the encoding rather than the thing it edited.
"""
edited = {k: v for k, v in action.items() if k != "contextSnapshotZ"}
if snapshot is None:
edited.pop("contextSnapshot", None)
else:
edited["contextSnapshot"] = snapshot
return edited
# ------------------------------------------------------------------- reading
def snapshot_of(client, adv_id, *, sources=None, s1=None, s2=None,
tip_state=None) -> dict:
"""Everything about a campaign that a round trip has to reproduce.
Read through the API, so the comparison is between what a reader can see in
the source campaign and what a reader can see in the copy. Two campaigns
that agree here agree on everything the product promises about a restored
campaign; nothing below is a database id, because ids are expected to
differ.
"""
head = client.get(f"/api/adventures/{adv_id}").json()
branches = client.get(f"/api/adventures/{adv_id}/branches").json()
checkpoints = client.get(f"/api/adventures/{adv_id}/checkpoints").json()
knowledge = client.get(f"/api/adventures/{adv_id}/knowledge").json()
state = client.get(f"/api/adventures/{adv_id}/state").json()
events = client.get(f"/api/adventures/{adv_id}/state/events?limit=500").json()
memories = client.get(f"/api/adventures/{adv_id}/memories").json()
derived = client.get(f"/api/adventures/{adv_id}/derived").json()
return {
"id": adv_id,
"title": head["title"],
"canon_rules": head.get("canon_rules") or [],
"can_undo": head.get("can_undo"),
"can_redo": head.get("can_redo"),
"transcript": [(a["type"], a["text"]) for a in head["actions"]],
# Every branch's own story, which is the whole retained tree as text.
"branch_count": len(branches),
"checkpoints": sorted(
(c["name"], c["note"]) for c in checkpoints
),
"knowledge": sorted(
(k["title"], k["classification"], k["enabled"], k["visibility"],
k["always_include"], k["content_hash"])
for k in knowledge
),
"state": _comparable_state(state),
"state_events": sorted(
(e["event_type"], e["source"], _payload_key(e["payload"]))
for e in events
),
"memories": sorted(m["text"] for m in memories),
"summaries": sorted(
(s["preview"], s["trigger"], s["eligible"])
for s in derived.get("summaries", [])
),
# Carried through from `build`, for the tests that need the original
# ids or the state at a position the head has since left.
"sources": sources,
"s1": s1,
"s2": s2,
"tip_state": _comparable_state(tip_state) if tip_state else None,
}
def _comparable_state(state: dict) -> dict:
"""The authoritative state, with only what a reader is shown.
Groups arrive from the API as display sections, which is the right shape to
compare: two campaigns whose State panels read identically hold the same
state, whatever ids sit underneath.
"""
groups = state.get("groups") if isinstance(state, dict) else None
if not isinstance(groups, list):
return {}
return {
str(group.get("title")): sorted(
", ".join(f"{k}={group_row[k]}" for k in sorted(group_row))
for group_row in (group.get("rows") or [])
if isinstance(group_row, dict)
)
for group in groups
}
def _payload_key(payload) -> str:
"""A stable identity for an event payload, for set comparison."""
if not isinstance(payload, dict):
return str(payload)
for key in ("fact_id", "entity_id", "thread_id", "id", "predicate"):
if payload.get(key):
return f"{key}={payload[key]}"
return ",".join(f"{k}={payload[k]}" for k in sorted(payload))