M9: a campaign you can actually get back

A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
JesseMarkowitz
2026-09-07 01:55:45 -04:00
co-authored by Claude Opus 5
parent 1ce9972760
commit 44edece67e
46 changed files with 9227 additions and 178 deletions
+406
View File
@@ -0,0 +1,406 @@
"""M9: one campaign that exercises every portable data family at once.
`TEST-CAMPAIGN-FIXTURE.md` describes the standard Continuity Test campaign, and
the acceptance suites use it. This is a different thing and does not replace it:
the Continuity Test is shaped to read like a story, and this one is shaped to
break a round trip. Every property M9 promises has a source in this campaign that
would be silently lost by a plausible mistake in the exporter or the importer.
Opening
|
+-- normal turns transcript, state events, snapshots
+-- Retry two takes at one coordinate
+-- knowledge retrieval imported passages in a stored prompt
+-- Save Point S1 a named coordinate on the first line
+-- more turns a future the reader will leave
|
+-- Undo x2 the head steps back
|
+-- divergent continuation a second branch, and a second future
+-- Save Point S2 a named coordinate on the second line
+-- manual state correction an event nothing narrated
+-- Undo x1 the head ends behind the newest row
The shape is chosen so that no single fact identifies a position. The active head
is not the newest row, not the deepest row, not the last row written, and not on
the branch that holds the most story — an importer that guesses any one of those
lands somewhere else.
Two campaigns are built, not one. `build` returns the rich campaign; the fixture
also leaves a neighbour beside it, because a bundle that accidentally exported
another campaign's rows would otherwise export nothing and pass.
The builder speaks HTTP throughout. A fixture that wrote rows directly would
prove the exporter can read what the fixture wrote, which is not the claim.
"""
from __future__ import annotations
import asyncio
from app import memorybank
from fakes import ScriptedProvider, state_block
# --------------------------------------------------------------- source files
# Three imported sources, one per class, plus the two lifecycle states that a
# round trip most easily loses: a source someone switched off, and one only the
# narrator may see.
CANON_MD = """# Westhaven
## The Old Abbey
The abbey above Westhaven has stood since the founding. Its crypt is sealed,
and the seal has never been broken.
## What cannot happen here
The dead do not return. No rite, relic or bargain in Westhaven has ever
returned anyone from death, and none ever will.
"""
REFERENCE_MD = """# The Crooked Lantern
The tavern on Fen Street is timber-framed, low-beamed, and older than the
street it stands on. The hearth is never allowed to go out.
## The keeper
Mara keeps the Crooked Lantern. She was born in Westhaven and has never left
it.
"""
INSPIRATION_MD = """# Weather notes
Rain on shutters. Lantern light through wet glass. The smell of a hearth
banked for the night.
"""
SECRET_MD = """# The seal
The abbey seal was broken once, sixty years ago, and set again by a hand that
is still alive. Nobody in Westhaven knows this.
"""
DISABLED_MD = """# Discarded draft
An earlier draft of the Westhaven material, kept for reference and switched off
so it cannot reach the narrator.
"""
#: The campaign's own rule, so the correction and the canon block have something
#: real to be measured against.
CAMPAIGN_CANON = {"rules": ["The dead do not return."]}
OPENING = "Aldric sits in the Crooked Lantern with Mara, and the rain starts."
# ------------------------------------------------------------------- helpers
def _play(client, adv_id, text, prose, events=None, kind="do"):
ScriptedProvider.replies = [f"{prose}\n{state_block(events or [])}"]
response = client.post(
f"/api/adventures/{adv_id}/actions", json={"type": kind, "text": text}
)
assert response.status_code == 200, response.text[:400]
return response
def _fact(predicate, value, fact_id):
return {"type": "add_fact", "predicate": predicate, "value": value,
"fact_id": fact_id}
def upload(client, adv_id, name, body, classification, **fields):
"""Imports a file the way the browser does: multipart, and no pathname."""
data = {"classification": classification}
data.update({k: str(v).lower() if isinstance(v, bool) else str(v)
for k, v in fields.items()})
response = client.post(
f"/api/adventures/{adv_id}/knowledge",
files={"file": (name, body.encode("utf-8"), "text/markdown")},
data=data,
)
assert response.status_code == 201, response.text[:400]
return response.json()["id"]
def _checkpoint(client, adv_id, name, note=""):
response = client.post(
f"/api/adventures/{adv_id}/checkpoints", json={"name": name, "note": note}
)
assert response.status_code == 201, response.text[:400]
return response.json()
def _undo(client, adv_id, times=1):
for _ in range(times):
response = client.post(f"/api/adventures/{adv_id}/undo")
assert response.status_code == 200, response.text[:400]
def settle_derived(adv_id):
"""Runs the background memory and summary pass to completion.
The turn endpoint fires this as a fire-and-forget task, which a test client
does not wait for. Calling it directly is the same code on the same rows —
what is skipped is the scheduling, not the work — and it is what
`test_context_realistic.py` does for the same reason.
"""
asyncio.run(memorybank.run_post_turn(adv_id))
# --------------------------------------------------------------------- build
def build(client, adv_id) -> dict:
"""Plays the fixture campaign onto `adv_id`, and returns what it built.
The returned dictionary is the assertion source for every round-trip test:
it names the properties that must survive, measured from the campaign as it
stands here rather than restated as constants, so a test compares the copy
against the original instead of against a guess about the original.
"""
# Story memory and the rolling summary on, because a campaign that
# generated neither would let an exporter omit both and still pass. The
# abandoned line below gets long enough to earn its own, which is what E03
# is about after a round trip.
switched_on = client.patch(
f"/api/adventures/{adv_id}",
json={"auto_summarize": True, "memory_bank_enabled": True},
)
assert switched_on.status_code == 200, switched_on.text[:400]
sources = {
"canon": upload(client, adv_id, "canon.md", CANON_MD, "canon",
always_include=True),
"reference": upload(client, adv_id, "reference.md", REFERENCE_MD,
"reference"),
"inspiration": upload(client, adv_id, "inspiration.md", INSPIRATION_MD,
"inspiration"),
"secret": upload(client, adv_id, "secret.md", SECRET_MD, "canon",
visibility="hidden"),
"disabled": upload(client, adv_id, "draft.md", DISABLED_MD, "reference"),
}
disable = client.patch(
f"/api/adventures/{adv_id}/knowledge/{sources['disabled']}",
json={"enabled": False},
)
assert disable.status_code == 200, disable.text[:400]
# ---- the first line of story -----------------------------------------
# Turn 1 asks about the abbey, so the canon source is retrieved and the
# stored prompt for this turn holds an imported passage. That turn is the
# one the provenance tests read back after the round trip.
_play(client, adv_id, "ask Mara about the abbey",
"Mara sets down the cloth. The abbey, she says, is sealed.",
[_fact("tally", 10, "tally-10")])
_play(client, adv_id, "walk up to the abbey",
"The path climbs out of the town and the rain follows.",
[_fact("tally", 20, "tally-20")])
# A retry, so one coordinate holds two takes and the earlier one is
# retained but not selected.
ScriptedProvider.replies = [
"The door is oak, and the seal on it is unbroken.\n"
+ state_block([_fact("tally", 30, "tally-30")])
]
_play(client, adv_id, "try the crypt door",
"The door will not move.", [_fact("tally", 30, "tally-30")])
retry = client.post(f"/api/adventures/{adv_id}/retry")
assert retry.status_code == 200, retry.text[:400]
s1 = _checkpoint(client, adv_id, "At the crypt door",
"Before anything is decided.")
# The future the reader is about to leave behind. It is played out far
# enough to earn derived data of its own — `memorybank.MEMORY_INTERVAL` is
# six actions — because a summary and a memory belonging to an abandoned
# line are what E03 forbids reaching an active prompt, and a round trip is
# a new way to leak one.
_play(client, adv_id, "force the door",
"The seal gives, and the stair below is dark.",
[_fact("tally", 40, "tally-40")])
_play(client, adv_id, "go down",
"The crypt is dry, and the air has not moved in years.",
[_fact("tally", 50, "tally-50")])
_play(client, adv_id, "read the names on the slabs",
"Sixty years of Westhaven dead, and one slab with no name at all.",
[_fact("tally", 60, "tally-60")])
_play(client, adv_id, "touch the nameless slab",
"The stone is warm, which stone in a crypt is not.",
[_fact("tally", 70, "tally-70")])
# Derived data for the line that is about to be abandoned, written while
# the head is still on it. This is the summary and the memory that must
# come back after a round trip and must still be ineligible there.
settle_derived(adv_id)
tip_state = client.get(f"/api/adventures/{adv_id}/state").json()
# ---- step back, and go somewhere else ---------------------------------
_undo(client, adv_id, 4)
_play(client, adv_id, "turn back and return to the tavern",
"The rain has not let up, and the Lantern's windows are lit.",
[_fact("tally", 41, "tally-41")])
s2 = _checkpoint(client, adv_id, "Back at the Lantern", "The other way.")
_play(client, adv_id, "ask Mara what she is not saying",
"She looks at the fire for a while before she answers.",
[_fact("tally", 51, "tally-51")])
_play(client, adv_id, "wait",
"The rain fills the silence, and then she starts talking.",
[_fact("tally", 61, "tally-61")])
# A manual correction: an accepted state change with no narration behind
# it, which is the one kind of state event a replay could never recreate.
correction = client.post(
f"/api/adventures/{adv_id}/state/corrections",
json={
"events": [{
"type": "add_fact",
"predicate": "keeper_of_the_lantern",
"value": "Mara",
"fact_id": "keeper",
}],
"note": "Established in play before the state system saw it.",
},
)
assert correction.status_code == 201, correction.text[:400]
# Derived data for the line the reader stayed on, so the copy has both an
# eligible and an ineligible summary to tell apart. The generated one landed
# on the abandoned line, which is the E03 case; this one is typed at the
# current head, so it is the eligible case beside it. A round trip has to
# keep them on opposite sides of that line.
settle_derived(adv_id)
# One more Undo, so the head finishes behind the retained tip of its own
# branch as well as behind the abandoned line's.
_undo(client, adv_id, 1)
# Typed at the final head, so it is the eligible summary and the generated
# one on the abandoned line is not. A round trip has to keep them on
# opposite sides of that line.
typed = client.patch(
f"/api/adventures/{adv_id}",
json={"story_summary": "Aldric went back to the Lantern instead."},
)
assert typed.status_code == 200, typed.text[:400]
return snapshot_of(client, adv_id, sources=sources, s1=s1, s2=s2,
tip_state=tip_state)
def snapshot_in(action: dict) -> dict | None:
"""The stored prompt in one bundle entry, decoded.
The export compresses it (`bundle._packed`), so a test that reached for a
plain dict would conclude the evidence was missing when it is merely
encoded. Both keys are read, plain first, exactly as the importer does.
"""
from app import bundle
plain = action.get("contextSnapshot")
if isinstance(plain, dict):
return plain
return bundle._unpacked(action.get("contextSnapshotZ"))
def with_snapshot(action: dict, snapshot: dict | None) -> dict:
"""A bundle entry carrying `snapshot`, written in the plain form.
Tests that break a snapshot on purpose write the readable key, because the
importer prefers it and because a test that had to compress its own fixture
would be testing the encoding rather than the thing it edited.
"""
edited = {k: v for k, v in action.items() if k != "contextSnapshotZ"}
if snapshot is None:
edited.pop("contextSnapshot", None)
else:
edited["contextSnapshot"] = snapshot
return edited
# ------------------------------------------------------------------- reading
def snapshot_of(client, adv_id, *, sources=None, s1=None, s2=None,
tip_state=None) -> dict:
"""Everything about a campaign that a round trip has to reproduce.
Read through the API, so the comparison is between what a reader can see in
the source campaign and what a reader can see in the copy. Two campaigns
that agree here agree on everything the product promises about a restored
campaign; nothing below is a database id, because ids are expected to
differ.
"""
head = client.get(f"/api/adventures/{adv_id}").json()
branches = client.get(f"/api/adventures/{adv_id}/branches").json()
checkpoints = client.get(f"/api/adventures/{adv_id}/checkpoints").json()
knowledge = client.get(f"/api/adventures/{adv_id}/knowledge").json()
state = client.get(f"/api/adventures/{adv_id}/state").json()
events = client.get(f"/api/adventures/{adv_id}/state/events?limit=500").json()
memories = client.get(f"/api/adventures/{adv_id}/memories").json()
derived = client.get(f"/api/adventures/{adv_id}/derived").json()
return {
"id": adv_id,
"title": head["title"],
"canon_rules": head.get("canon_rules") or [],
"can_undo": head.get("can_undo"),
"can_redo": head.get("can_redo"),
"transcript": [(a["type"], a["text"]) for a in head["actions"]],
# Every branch's own story, which is the whole retained tree as text.
"branch_count": len(branches),
"checkpoints": sorted(
(c["name"], c["note"]) for c in checkpoints
),
"knowledge": sorted(
(k["title"], k["classification"], k["enabled"], k["visibility"],
k["always_include"], k["content_hash"])
for k in knowledge
),
"state": _comparable_state(state),
"state_events": sorted(
(e["event_type"], e["source"], _payload_key(e["payload"]))
for e in events
),
"memories": sorted(m["text"] for m in memories),
"summaries": sorted(
(s["preview"], s["trigger"], s["eligible"])
for s in derived.get("summaries", [])
),
# Carried through from `build`, for the tests that need the original
# ids or the state at a position the head has since left.
"sources": sources,
"s1": s1,
"s2": s2,
"tip_state": _comparable_state(tip_state) if tip_state else None,
}
def _comparable_state(state: dict) -> dict:
"""The authoritative state, with only what a reader is shown.
Groups arrive from the API as display sections, which is the right shape to
compare: two campaigns whose State panels read identically hold the same
state, whatever ids sit underneath.
"""
groups = state.get("groups") if isinstance(state, dict) else None
if not isinstance(groups, list):
return {}
return {
str(group.get("title")): sorted(
", ".join(f"{k}={group_row[k]}" for k in sorted(group_row))
for group_row in (group.get("rows") or [])
if isinstance(group_row, dict)
)
for group in groups
}
def _payload_key(payload) -> str:
"""A stable identity for an event payload, for set comparison."""
if not isinstance(payload, dict):
return str(payload)
for key in ("fact_id", "entity_id", "thread_id", "id", "predicate"):
if payload.get(key):
return f"{key}={payload[key]}"
return ",".join(f"{k}={payload[k]}" for k in sorted(payload))
+22 -3
View File
@@ -568,10 +568,29 @@ def test_the_action_cap_counts_the_rows_a_v1_file_expands_into(client, monkeypat
assert _adventure_count() == before, "and nothing was written"
def test_an_unknown_format_is_refused(client):
r = _import(client, {"format": "ai-dnd-adventure-v3", "title": "From the future"})
def test_a_format_from_a_later_build_is_refused(client):
"""A version this build has never heard of is refused, not guessed at.
The placeholder version here has to stay ahead of `bundle.FORMAT`. It was
`v3` until M9 made v3 real, at which point this test started importing a
bundle it meant to reject — the failure mode a hard-coded "next version"
always eventually has, and the reason the message is asserted against
`bundle.FORMAT` rather than against a literal.
"""
r = _import(client, {"format": "ai-dnd-adventure-v99", "title": "From the future"})
assert r.status_code == 400, r.text
assert bundle.FORMAT in r.json()["detail"]
detail = r.json()["detail"]
assert bundle.FORMAT in detail
assert "ai-dnd-adventure-v99" in detail
def test_something_that_is_not_an_export_at_all_is_refused(client):
r = _import(client, {"title": "A file of some other kind"})
assert r.status_code == 400, r.text
# Every version it can read is named, so the reader can tell whether the
# file they have is one of them.
for readable in bundle.READABLE:
assert readable in r.json()["detail"]
# ------------------------------------------------------- the persona (Phase 18)
+38 -7
View File
@@ -1660,7 +1660,22 @@ def test_an_edited_content_hash_is_recomputed_and_reported(client):
def test_historical_prompt_evidence_survives_an_export_round_trip(client):
"""§33: the round trip does not turn provenance into dangling ids."""
"""§33: the round trip does not turn provenance into dangling ids.
Written in M7 and rewritten in M9, and the rewrite is the point of it.
In M7 the bundle carried no context snapshots at all, so this test pinned
the *absence*: there were no ids to dangle because there was no evidence,
and the imported campaign's turns simply had no snapshot. That was recorded
at the time as a limit owned by M9 rather than as a property worth keeping —
`V1-ACCEPTANCE-TESTS.md` I05 said so in as many words, and the M8 report
made it handoff question B.
M9 answered it: the evidence travels. So the assertion inverts, and what it
now pins is the thing M7 was worried about and could not check — that the
provenance arriving on the other side names *this* campaign's sources rather
than the ids it had on the machine that wrote the file.
"""
ids = import_fixture(client)
play(client, "Aldric asks about the Old Abbey and its broken-circle symbol.")
actions = client.get(f"/api/adventures/{client.adv_id}/actions").json()["actions"]
@@ -1673,17 +1688,33 @@ def test_historical_prompt_evidence_survives_an_export_round_trip(client):
bundle = client.get(f"/api/adventures/{client.adv_id}/export").json()
restored = client.post("/api/adventures/import", json=bundle).json()
# The bundle carries no context snapshots at all — it never has, by the rule
# at the top of `bundle.py` — so there are no ids to dangle. The imported
# campaign's turns simply have no snapshot, which is what a pre-M7 bundle
# already did for every other component of the inspector.
new_actions = client.get(
f"/api/adventures/{restored['id']}/actions"
).json()["actions"]
new_ai = next(a for a in reversed(new_actions) if a["type"] == "ai")
assert client.get(
moved = client.get(
f"/api/adventures/{restored['id']}/actions/{new_ai['id']}/context"
).status_code == 404
)
assert moved.status_code == 200, moved.text[:300]
moved = moved.json()
# The evidence itself is identical: the same passages, the same text, the
# same prompt the turn was actually assembled from.
assert [(r["title"], r["text"]) for r in moved["knowledge"]["used"]] == \
[(r["title"], r["text"]) for r in before["knowledge"]["used"]]
assert moved["prompt"] == before["prompt"]
# And the one pointer that is not evidence has been translated, so the
# inspector's "open this source" reaches the restored library rather than
# whatever holds that id here.
theirs = {
source["id"] for source in
client.get(f"/api/adventures/{restored['id']}/knowledge").json()
}
named = {r["source_id"] for r in moved["knowledge"]["used"]
if r["source_id"] is not None}
assert named and named <= theirs
# And the original campaign's evidence is untouched by having been exported.
after = client.get(
f"/api/adventures/{client.adv_id}/actions/{ai_action['id']}/context"
+451
View File
@@ -0,0 +1,451 @@
"""M9: a consistent copy of the whole database, taken while it is being written.
`app/backup.py` explains why a plain file copy is not a backup. This file is the
evidence for the claim, and the shape of it matters: **every test below opens the
backup as its own database and reads what is in it.** A test that only checked a
file appeared, or that the endpoint returned 201, would pass against a `cp` — and
a `cp` is exactly what this replaces.
The load test is the one that separates the two. It writes to the source
database *while* the backup is being taken, from a second thread, and then asks
the copy for a story it can check turn by turn. A page-torn copy would show a
transcript with a hole in it, a campaign whose head points past its own story, or
a `quick_check` failure — and would show none of those on a quiet database, which
is why the quiet case is not the interesting one.
python -m pytest tests/test_m9_backup.py -v
"""
import os
import sqlite3
import tempfile
import threading
import time
from pathlib import Path
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, backup, limits, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import adventures
from fakes import ScriptedProvider, tally_of, tally_reply
@pytest.fixture()
def client(monkeypatch):
"""The app, and a campaign with enough in it to recognise afterwards."""
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="backup@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(user_id=user.id, model="test-model"))
adventure = models.Adventure(user_id=user.id, title="Backed up")
setup.add(adventure)
setup.flush()
setup.add(models.Action(
adventure_id=adventure.id, type="start", text="The story opens.",
))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture()
def elsewhere(tmp_path, monkeypatch):
"""Backups land under a temporary directory, not beside the real database."""
fake_db = tmp_path / "campaign.db"
fake_db.write_bytes(Path(str(engine.url.database)).read_bytes())
return fake_db
def _play(client, text, total):
ScriptedProvider.replies = [tally_reply(f"Beat {total // 10}.", total)]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": text})
assert response.status_code == 200, response.text[:300]
def _open(path) -> sqlite3.Connection:
"""The backup, as its own database, read-only."""
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
connection.row_factory = sqlite3.Row
return connection
# ------------------------------------------------------------------ the copy
def test_the_backup_is_a_database_that_passes_its_own_integrity_check(client):
for turn in range(1, 4):
_play(client, f"turn {turn}", turn * 10)
result = backup.create()
try:
assert result.integrity == "ok"
assert result.pages > 0
assert result.bytes > 0
with _open(result.path) as db:
assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
assert db.execute("PRAGMA foreign_key_check").fetchall() == []
finally:
result.path.unlink(missing_ok=True)
def test_the_backup_holds_the_schema_and_every_family_of_row(client):
"""Not "the file exists": the copy is opened and asked what is in it."""
for turn in range(1, 4):
_play(client, f"turn {turn}", turn * 10)
checkpoint = client.post(f"/api/adventures/{client.adv_id}/checkpoints",
json={"name": "Here", "note": "A position."})
assert checkpoint.status_code == 201
upload = client.post(
f"/api/adventures/{client.adv_id}/knowledge",
files={"file": ("canon.md", b"# Rule\n\nThe dead do not return.\n",
"text/markdown")},
data={"classification": "canon"},
)
assert upload.status_code == 201, upload.text[:300]
result = backup.create()
try:
with _open(result.path) as db:
tables = {
row["name"] for row in
db.execute("SELECT name FROM sqlite_master WHERE type='table'")
}
for expected in ("adventures", "actions", "branches", "checkpoints",
"knowledge_sources", "knowledge_chunks",
"state_events", "summaries", "settings"):
assert expected in tables, f"{expected} is missing from the backup"
campaign = db.execute(
"SELECT * FROM adventures WHERE id = ?", (client.adv_id,)
).fetchone()
assert campaign["title"] == "Backed up"
# The head, which is the thing a restore has to reproduce.
assert campaign["head_depth"] >= 0
assert campaign["head_branch_id"] is not None
texts = [row["text"] for row in db.execute(
"SELECT text FROM actions WHERE adventure_id = ? ORDER BY id",
(client.adv_id,),
)]
assert "The story opens." in texts
assert any("Beat 3." in text for text in texts)
assert db.execute(
"SELECT name FROM checkpoints WHERE adventure_id = ?",
(client.adv_id,),
).fetchone()["name"] == "Here"
assert db.execute(
"SELECT COUNT(*) c FROM knowledge_sources WHERE adventure_id = ?",
(client.adv_id,),
).fetchone()["c"] == 1
assert db.execute(
"SELECT COUNT(*) c FROM state_events WHERE adventure_id = ?",
(client.adv_id,),
).fetchone()["c"] > 0
# And the head names a turn that is actually in the copy.
assert db.execute(
"SELECT COUNT(*) c FROM actions WHERE adventure_id = ? "
"AND branch_id = ? AND depth = ?",
(client.adv_id, campaign["head_branch_id"], campaign["head_depth"]),
).fetchone()["c"] > 0
finally:
result.path.unlink(missing_ok=True)
def test_the_state_in_the_backup_is_the_state_the_campaign_had(client):
"""The authoritative document, read out of the copy and compared."""
for turn in range(1, 5):
_play(client, f"turn {turn}", turn * 10)
live = client.get(f"/api/adventures/{client.adv_id}/state").json()["document"]
result = backup.create()
try:
with _open(result.path) as db:
from app import compression
blob = db.execute(
"SELECT narrative_state FROM adventures WHERE id = ?",
(client.adv_id,),
).fetchone()["narrative_state"]
assert tally_of(compression.unpack(blob)) == tally_of(live) == 40
finally:
result.path.unlink(missing_ok=True)
# ------------------------------------------------------------ while it is live
def test_a_backup_taken_during_writes_is_consistent(client):
"""The claim a plain file copy cannot make.
Turns are played from a second thread throughout the copy. The backup that
comes out is a snapshot of *some* committed point — which point is not
determined, and asserting on a particular one would be asserting on a race —
so what is checked is that it is a coherent one: `quick_check` passes, no
foreign key dangles, the transcript has no gap in it, and the head names a
turn that exists.
"""
stop = threading.Event()
written: list[int] = []
failures: list[Exception] = []
def keep_writing():
turn = 0
while not stop.is_set() and turn < 40:
turn += 1
try:
_play(client, f"concurrent {turn}", turn * 10)
written.append(turn)
except Exception as exc: # noqa: BLE001 - reported to the test
failures.append(exc)
return
time.sleep(0.005)
writer = threading.Thread(target=keep_writing, daemon=True)
writer.start()
# Let a few turns land, so the copy is taken over a database that is moving
# rather than one that has not started.
while len(written) < 3 and writer.is_alive():
time.sleep(0.01)
result = backup.create()
stop.set()
writer.join(timeout=30)
assert not failures, f"the writer failed: {failures[0]}"
assert written, "no turn was written during the backup"
try:
with _open(result.path) as db:
assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
assert db.execute("PRAGMA foreign_key_check").fetchall() == []
rows = db.execute(
"SELECT depth, type FROM actions WHERE adventure_id = ? "
"AND live = 1 ORDER BY depth",
(client.adv_id,),
).fetchall()
depths = [row["depth"] for row in rows]
assert depths == list(range(len(depths))), (
f"the transcript in the backup has a gap: {depths}"
)
campaign = db.execute(
"SELECT head_branch_id, head_depth FROM adventures WHERE id = ?",
(client.adv_id,),
).fetchone()
assert db.execute(
"SELECT COUNT(*) c FROM actions WHERE adventure_id = ? "
"AND branch_id = ? AND depth = ?",
(client.adv_id, campaign["head_branch_id"], campaign["head_depth"]),
).fetchone()["c"] > 0, "the head points past the story in the backup"
finally:
result.path.unlink(missing_ok=True)
def test_the_source_database_is_untouched_by_a_backup(client):
"""Opened read-only, so this is a guarantee rather than an observation."""
_play(client, "one", 10)
source = Path(str(engine.url.database))
before = source.read_bytes()
result = backup.create()
try:
assert source.read_bytes() == before
assert client.get(f"/api/adventures/{client.adv_id}").status_code == 200
finally:
result.path.unlink(missing_ok=True)
# ---------------------------------------------------------------- the rules
def test_an_existing_backup_is_never_overwritten(client):
"""Yesterday's backup surviving today's mistake is most of the point."""
first = backup.create()
second = backup.create()
try:
assert first.path != second.path
assert first.path.exists() and second.path.exists()
finally:
first.path.unlink(missing_ok=True)
second.path.unlink(missing_ok=True)
def test_two_backups_in_the_same_second_do_not_collide(client, monkeypatch):
from datetime import datetime
fixed = datetime(2026, 9, 7, 4, 30, 0)
first = backup.create(now=fixed)
second = backup.create(now=fixed)
try:
assert first.path != second.path
assert first.path.exists() and second.path.exists()
finally:
first.path.unlink(missing_ok=True)
second.path.unlink(missing_ok=True)
def test_a_failed_verification_leaves_nothing_behind(client, monkeypatch):
"""A backup nobody verified is a belief, and one that fails is not kept."""
monkeypatch.setattr(
backup, "_verify",
lambda path: (_ for _ in ()).throw(backup.BackupError("bad pages")),
)
root = backup.directory()
before = set(root.iterdir())
with pytest.raises(backup.BackupError, match="bad pages"):
backup.create()
assert set(root.iterdir()) == before, "a failed backup left a file behind"
def test_a_failed_copy_leaves_nothing_behind_and_reports_the_reason(
client, monkeypatch
):
monkeypatch.setattr(
backup, "_copy",
lambda source, working: (_ for _ in ()).throw(OSError("disk full")),
)
root = backup.directory()
before = set(root.iterdir())
with pytest.raises(backup.BackupError, match="disk full"):
backup.create()
assert set(root.iterdir()) == before
def test_a_missing_source_database_is_reported_rather_than_guessed_at(tmp_path):
with pytest.raises(backup.BackupError, match="no database"):
backup.create(tmp_path / "not-here.db")
def test_the_partial_file_is_never_left_wearing_a_backups_name(client, monkeypatch):
"""The rename is the last step, so an interrupted run is invisible."""
seen: list[Path] = []
real_copy = backup._copy
def watch(source, working):
seen.append(Path(working))
return real_copy(source, working)
monkeypatch.setattr(backup, "_copy", watch)
result = backup.create()
try:
assert seen and seen[0].name.endswith(".partial")
assert not seen[0].exists(), "the temporary file survived"
assert result.path.exists()
assert not result.path.name.endswith(".partial")
finally:
result.path.unlink(missing_ok=True)
# --------------------------------------------------------------- the endpoint
def test_the_endpoint_takes_a_backup_and_says_where_it_went(client):
response = client.post("/api/backups")
assert response.status_code == 201, response.text[:300]
body = response.json()
path = Path(body["directory"]) / body["filename"]
try:
assert body["integrity"] == "ok"
assert body["bytes"] > 0
assert path.exists()
with _open(path) as db:
assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
finally:
path.unlink(missing_ok=True)
def test_the_endpoint_lists_what_is_there_newest_first(client):
"""Ordered by when the backup was taken, which is what its name records.
Both files here are written in the same instant, so their modification times
are indistinguishable and only the stamp in the name says which is which.
That is not a contrived case: copying a backup to another disk or restoring
one from an archive rewrites its mtime, and a list that reordered itself
afterwards would report when the file was last handled rather than when the
backup was taken.
"""
from datetime import datetime
older = backup.create(now=datetime(2026, 9, 1, 10, 0, 0))
newer = backup.create(now=datetime(2026, 9, 6, 10, 0, 0))
try:
listed = client.get("/api/backups")
assert listed.status_code == 200
rows = listed.json()["backups"]
names = [row["filename"] for row in rows]
assert names.index(newer.path.name) < names.index(older.path.name)
by_name = {row["filename"]: row["taken_at"] for row in rows}
assert by_name[newer.path.name].startswith("2026-09-06T10:00")
assert by_name[older.path.name].startswith("2026-09-01T10:00")
finally:
older.path.unlink(missing_ok=True)
newer.path.unlink(missing_ok=True)
def test_a_backup_this_build_did_not_name_still_lists(client):
"""A file in the directory whose name carries no stamp is still shown.
The modification time answers instead. The fallback exists to keep a
hand-renamed or third-party file visible rather than silently absent from
the list a reader uses to find their backups.
"""
stray = backup.directory() / f"{backup.PREFIX}-handwritten.db"
stray.write_bytes(b"SQLite format 3\x00")
try:
rows = client.get("/api/backups").json()["backups"]
listed = {row["filename"]: row for row in rows}
assert stray.name in listed
assert listed[stray.name]["taken_at"]
finally:
stray.unlink(missing_ok=True)
def test_the_endpoint_accepts_no_path_from_the_caller(client):
"""H08. There is no field to attempt a traversal in.
The destination is derived from the database the application already has
open and the name from the clock, so a body is not merely ignored — there is
nothing for one to name.
"""
from app.main import app as application
schema = application.openapi()["paths"]["/api/backups"]["post"]
assert "requestBody" not in schema
assert not schema.get("parameters")
# And sending one anyway changes nothing about where the file lands.
response = client.post("/api/backups", json={"path": "../../../tmp/escape.db"})
assert response.status_code == 201, response.text[:300]
body = response.json()
path = Path(body["directory"]) / body["filename"]
try:
assert path.parent == backup.directory()
assert ".." not in body["filename"]
finally:
path.unlink(missing_ok=True)
def test_a_failure_is_a_clear_error_rather_than_a_silent_success(
client, monkeypatch
):
monkeypatch.setattr(
backup, "create",
lambda *a, **k: (_ for _ in ()).throw(backup.BackupError("no space left")),
)
response = client.post("/api/backups")
assert response.status_code == 500
assert "no space left" in response.json()["detail"]
+460
View File
@@ -0,0 +1,460 @@
"""M9: the campaign moves to a machine that has never seen it.
This is the milestone's Definition of Done, and it is the one claim the rest of
the M9 suite cannot make. `test_m9_portability.py` imports beside the original,
in one process, against one database — which is the right place to check the
*contract* and the wrong place to check *portability*. A shared id space, a
warm cache, a row the exporter forgot to scope, a session still holding the
original: every one of those would pass there and fail here.
So each test below:
1. starts a real server process against database A, and plays a campaign;
2. exports it over HTTP and stops that process;
3. starts a **second** server process against database B, **a file that has
never existed before**, in a different directory;
4. imports the file over HTTP, and asks the second process what it has.
Nothing crosses between them but the bundle. Migrations run on B from nothing,
because it is a new file — so this is also the fresh-install path, and the
"clean data directory" in the Definition of Done is a directory, not a metaphor.
The final test restarts the *importing* server, which is L03 after a move: a
Save Point restored in the third process must reach the same position and the
same state as it did in the second.
python -m pytest tests/test_m9_clean_import.py -v
"""
import json
import os
import shutil
import sqlite3
import subprocess
import sys
import tempfile
import urllib.error
import urllib.request
from pathlib import Path
import pytest
from fakes import TALLY_PER_TURN, tally_of
from test_process_restart import Server, _free_port
HERE = Path(__file__).resolve().parent
@pytest.fixture()
def machines():
"""Two directories, each with its own database, and the servers on them.
Two directories rather than two filenames, because the backup directory and
anything else the application derives from the database's location must land
in the importing machine's own space rather than beside the exporter's.
"""
root = tempfile.mkdtemp(prefix="m9-clean-")
started: list[Server] = []
def start(name: str) -> Server:
directory = os.path.join(root, name)
os.makedirs(directory, exist_ok=True)
server = Server(os.path.join(directory, "campaign.db"), _free_port())
started.append(server)
server.wait_until_ready()
return server
def path_of(name: str) -> str:
return os.path.join(root, name, "campaign.db")
try:
yield start, path_of
finally:
for server in started:
server.stop()
shutil.rmtree(root, ignore_errors=True)
# ------------------------------------------------------------------ building
def _campaign(server: Server) -> int:
"""A campaign with everything a move has to carry, played over HTTP.
Deliberately not `m9_fixture`: that builds through a `TestClient` and this
file exists to avoid one. What it reproduces is the same shape — a retry, a
Save Point, an imported source that a turn actually used, an undone head and
a retained future.
"""
adventure = server.call("POST", "/adventures", {
"title": "Moved between machines",
"canon_rules": ["The dead do not return."],
"opening": "Aldric sits in the Crooked Lantern with Mara.",
}, expect=201)
adv_id = adventure["id"]
_upload(server, adv_id, "canon.md", "canon", (
"# Westhaven\n\n## The Old Abbey\n\nThe abbey above Westhaven has stood "
"since the founding. Its crypt is sealed, its door is oak, and the seal "
"on it has never been broken.\n"
))
_upload(server, adv_id, "secret.md", "canon", (
"# The seal\n\nIt was broken once, sixty years ago.\n"
), visibility="hidden")
disabled = _upload(server, adv_id, "draft.md", "reference", (
"# Discarded draft\n\nAn earlier version, switched off.\n"
))
server.call("PATCH", f"/adventures/{adv_id}/knowledge/{disabled}",
{"enabled": False}, expect=200)
# The spawned narrator writes "Beat N." and nothing else, so every term the
# retrieval has to work with comes from the player's own words. They are
# written to name things the Canon file names.
server.play(adv_id, "ask Mara about the abbey crypt in Westhaven")
server.play(adv_id, "walk up the hill to the abbey")
server.play(adv_id, "try the sealed crypt door of the abbey")
_retry(server, adv_id)
server.call("POST", f"/adventures/{adv_id}/checkpoints",
{"name": "At the door", "note": "Before deciding."}, expect=201)
server.play(adv_id, "force the door")
server.play(adv_id, "go down the stair")
server.call("POST", f"/adventures/{adv_id}/state/corrections", {
"events": [{"type": "add_fact", "predicate": "keeper", "value": "Mara",
"fact_id": "keeper"}],
"note": "Established in play before the state system saw it.",
}, expect=201)
# Two Undos, so the export is taken behind the retained tip.
server.call("POST", f"/adventures/{adv_id}/undo", expect=200)
server.call("POST", f"/adventures/{adv_id}/undo", expect=200)
return adv_id
def _retry(server: Server, adv_id: int) -> None:
"""Retries the newest turn, over the streaming endpoint it actually uses.
`Server.call` parses JSON, and `/retry` answers with an SSE stream as
`/actions` does — so calling it as JSON reads `data: {...}` as a document and
fails on the first character. Draining the stream is what the browser does.
"""
request = urllib.request.Request(
f"http://127.0.0.1:{server.port}/api/adventures/{adv_id}/retry",
data=b"{}", method="POST",
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(request, timeout=120) as response:
body = response.read()
assert b'"type": "error"' not in body, body[:300]
def _upload(server: Server, adv_id: int, name: str, classification: str,
body: str, **fields) -> int:
"""A multipart knowledge upload over real HTTP, without a client library."""
boundary = "----m9cleanimport"
parts = []
for key, value in {"classification": classification, **fields}.items():
parts.append(
f"--{boundary}\r\nContent-Disposition: form-data; name=\"{key}\"\r\n"
f"\r\n{value}\r\n"
)
parts.append(
f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; "
f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n"
)
payload = ("".join(parts) + f"--{boundary}--\r\n").encode()
request = urllib.request.Request(
f"http://127.0.0.1:{server.port}/api/adventures/{adv_id}/knowledge",
data=payload, method="POST",
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"},
)
with urllib.request.urlopen(request, timeout=60) as response:
return json.loads(response.read())["id"]
def _snapshot(server: Server, adv_id: int) -> dict:
"""What a reader can see, read over HTTP through the API they read."""
page = server.call("GET", f"/adventures/{adv_id}", expect=200)
return {
"title": page["title"],
"canon_rules": page["canon_rules"],
"transcript": [(a["type"], a["text"]) for a in page["actions"]],
"can_undo": page["can_undo"],
"can_redo": page["can_redo"],
"state": server.call("GET", f"/adventures/{adv_id}/state", expect=200)["document"],
"checkpoints": sorted(
(c["name"], c["note"], c["depth"])
for c in server.call("GET", f"/adventures/{adv_id}/checkpoints", expect=200)
),
"knowledge": sorted(
(k["title"], k["classification"], k["enabled"], k["visibility"],
k["content_hash"], k["index_state"], k["chunk_count"] > 0)
for k in server.call("GET", f"/adventures/{adv_id}/knowledge", expect=200)
),
"events": sorted(
(e["event_type"], e["source"], json.dumps(e["payload"], sort_keys=True))
for e in server.call("GET", f"/adventures/{adv_id}/state/events?limit=500",
expect=200)
),
"rows": server.total_rows(adv_id),
}
# ------------------------------------------------------------------- the move
@pytest.fixture()
def moved(machines):
"""The campaign, exported from machine A and imported into a clean B."""
start, path_of = machines
source = start("a")
adv_id = _campaign(source)
before = _snapshot(source, adv_id)
# What the source machine retrieves at this position, recorded while it is
# still running. It is the only thing the copy can honestly be compared to.
retrieved = {
record["filename"] for record in
source.call("GET", f"/adventures/{adv_id}/context", expect=200)
["knowledge"]["used"]
}
bundle = source.call("GET", f"/adventures/{adv_id}/export", expect=200)
source.stop()
assert not os.path.exists(path_of("b")), "machine B must not exist yet"
target = start("b")
assert target.call("GET", "/adventures", expect=200) == [], \
"machine B is not empty"
imported = target.call("POST", "/adventures/import", bundle, expect=201)
return {
"bundle": bundle, "before": before, "target": target,
"retrieved": retrieved,
"copy_id": imported["id"], "imported": imported,
"path": path_of, "start": start,
}
def test_the_campaign_arrives_whole_on_a_machine_that_never_had_it(moved):
"""The Definition of Done, in one assertion per family."""
after = _snapshot(moved["target"], moved["copy_id"])
before = moved["before"]
assert after["transcript"] == before["transcript"]
assert after["state"] == before["state"]
assert after["canon_rules"] == before["canon_rules"]
assert after["checkpoints"] == before["checkpoints"]
assert after["knowledge"] == before["knowledge"]
assert after["events"] == before["events"]
assert after["rows"] == before["rows"], "the retained tree is a different size"
def test_it_opens_at_the_exact_head_it_was_exported_at(moved):
"""I07, across the boundary the acceptance test names.
The export was taken two Undos behind the tip, so a machine that opened the
campaign at its newest retained turn would show a story two turns longer
than the one that was saved.
"""
after = _snapshot(moved["target"], moved["copy_id"])
assert after["transcript"] == moved["before"]["transcript"]
assert after["can_redo"] is True, "the retained future is not reachable"
assert moved["imported"]["can_redo"] is True, (
"the response that opens the campaign says Redo is unavailable"
)
assert after["rows"] > len(after["transcript"]), (
"the retained future is not in the database"
)
def test_the_state_audit_arrives_and_still_names_its_author(moved):
"""The manual correction is still a manual correction on the new machine."""
events = moved["target"].call(
f"GET", f"/adventures/{moved['copy_id']}/state/events?limit=500", expect=200
)
manual = [e for e in events if e["source"] == "manual_correction"]
assert len(manual) == 1
assert manual[0]["payload"]["predicate"] == "keeper"
assert any(e["source"] == "accepted_story" for e in events), (
"and the story's own events are there beside it"
)
def test_the_knowledge_works_with_no_access_to_the_original_machine(moved):
"""§11. The exporting machine is stopped; nothing may reach back to it.
Its process is dead and its directory holds a database this server has never
opened. If retrieval works here, it works from the content the file carried.
The comparison is against what the *source* retrieved, recorded before that
process was killed, and the source's own result is asserted first. A test
that only checked the copy retrieved something would pass by accident on a
day the fixture happened to match, and — worse — would report a portability
failure when what had actually happened is that neither side retrieved
anything. That is M8's finding 10: assert your own precondition.
"""
assert moved["retrieved"], (
"the source campaign retrieved nothing, so this proves nothing about "
"the copy"
)
report = moved["target"].call(
"GET", f"/adventures/{moved['copy_id']}/context", expect=200
)
used = {record["filename"] for record in report["knowledge"]["used"]}
assert used == moved["retrieved"], (
f"the copy retrieved {used} where the source retrieved {moved['retrieved']}"
)
assert "draft.md" not in used, "the disabled source was re-enabled by the move"
assert "canon.md" in used
def test_a_historical_turn_still_shows_what_it_was_given(moved):
"""The M8 handoff, across the boundary that made it a handoff.
Inspect Context on an old narrator turn works on a machine that never
assembled that prompt and could not reassemble it — the sources are here but
the state, the head and the canon have all moved on since.
"""
target, copy_id = moved["target"], moved["copy_id"]
page = target.call("GET", f"/adventures/{copy_id}/actions?limit=200", expect=200)
narrator = [a for a in page["actions"] if a["type"] == "ai"]
assert narrator, "the imported campaign has no narrator turn"
inspected = 0
for action in narrator:
response = target.call(
"GET", f"/adventures/{copy_id}/actions/{action['id']}/context"
)
if response is None:
continue
assert response["prompt"]["system"], "a restored prompt is empty"
assert response["sections"], "a restored prompt has no sections"
inspected += 1
assert inspected, "no turn on the new machine can say what it was told"
def test_no_secret_and_no_path_from_the_old_machine_travelled(moved):
"""I06, and the private-detail half of it.
The bundle is checked as text, because that is what actually left the
machine — a field added to a model the exporter walks would reach the file
without any test of a column noticing.
"""
text = json.dumps(moved["bundle"])
assert "api_key" not in text
assert "11434" not in text, "an inference endpoint travelled with the campaign"
assert "/tmp/" not in text and "campaign.db" not in text, (
"a filesystem path from the exporting machine travelled"
)
def test_the_importing_machine_keeps_its_own_settings(moved):
"""§15. A campaign is not a way to reconfigure the destination.
The bundle carries per-turn model provenance, which is a record of what
happened. It does not carry the endpoint, the model or the context budget,
because those describe the machine rather than the campaign — and importing
a campaign must not silently repoint the destination's inference at the
source's.
"""
settings = moved["target"].call("GET", "/settings", expect=200)
assert settings["endpoint_url"] == "http://localhost:11434/v1", (
"the import changed the destination's inference endpoint"
)
assert settings["context_token_budget"] == 16384
def test_a_missing_model_does_not_stop_the_campaign_arriving(moved):
"""§15. The campaign and its data are portable independently of a model.
The importing server has no model configured at all — nothing has ever
written a `model` into its settings — and the import still succeeds, opens,
and shows its state. Play would fail; recovery does not.
"""
settings = moved["target"].call("GET", "/settings", expect=200)
assert settings["model"] == "", "this test needs an unconfigured destination"
after = _snapshot(moved["target"], moved["copy_id"])
assert after["transcript"] == moved["before"]["transcript"]
# ------------------------------------------------- L03, after the campaign moved
def test_l03_a_save_point_restored_on_the_new_machine_survives_its_restart(moved):
"""L03, with the move in front of it.
Restore a Save Point in the second process, record the position and the
state, kill the process, start a **third** against the same file, and ask
again. What crosses is bytes on disk.
"""
target, copy_id = moved["target"], moved["copy_id"]
points = target.call("GET", f"/adventures/{copy_id}/checkpoints", expect=200)
assert points, "the Save Point did not survive the move"
point = points[0]
assert point["resolved"] is True
target.call("POST", f"/adventures/{copy_id}/checkpoints/{point['id']}/restore",
expect=200)
restored = _snapshot(target, copy_id)
rows_before = restored["rows"]
target.stop()
assert not target.is_listening()
third = moved["start"]("b")
again = _snapshot(third, copy_id)
assert again["transcript"] == restored["transcript"]
assert again["state"] == restored["state"]
assert again["rows"] == rows_before, "restoring deleted later history"
# ------------------------------------------------------ the database it wrote
def test_the_importing_machines_database_passes_its_own_integrity_check(moved):
"""A campaign written by an import is a database SQLite is happy with."""
moved["target"].stop()
connection = sqlite3.connect(moved["path"]("b"))
try:
assert connection.execute("PRAGMA quick_check").fetchone()[0] == "ok"
assert connection.execute("PRAGMA foreign_key_check").fetchall() == []
finally:
connection.close()
def test_the_import_left_no_orphan_behind(moved):
"""§17's list, checked against the database rather than against the API.
Every one of these would be invisible from the outside until the moment it
mattered: a Save Point pointing at a turn that is not there, knowledge owned
by a campaign that does not exist, an action on a branch belonging to
something else.
"""
moved["target"].stop()
connection = sqlite3.connect(moved["path"]("b"))
try:
def one(sql):
return connection.execute(sql).fetchone()[0]
assert one("""
SELECT COUNT(*) FROM checkpoints c
LEFT JOIN actions a
ON a.branch_id = c.branch_id AND a.depth = c.depth
AND a.adventure_id = c.adventure_id
WHERE a.id IS NULL
""") == 0, "a Save Point names a position with no turn at it"
assert one("""
SELECT COUNT(*) FROM actions a
LEFT JOIN branches b ON b.id = a.branch_id
WHERE a.branch_id IS NOT NULL
AND (b.id IS NULL OR b.adventure_id <> a.adventure_id)
""") == 0, "an action sits on another campaign's branch"
assert one("""
SELECT COUNT(*) FROM knowledge_sources k
LEFT JOIN adventures adv ON adv.id = k.adventure_id
WHERE adv.id IS NULL
""") == 0, "knowledge owned by no campaign"
assert one("""
SELECT COUNT(*) FROM state_events e
LEFT JOIN actions a ON a.id = e.action_id
WHERE e.action_id IS NOT NULL
AND (a.id IS NULL OR a.adventure_id <> e.adventure_id)
""") == 0, "a state event names a turn in another campaign"
assert one("""
SELECT COUNT(*) FROM adventures adv
LEFT JOIN actions a
ON a.branch_id = adv.head_branch_id AND a.depth = adv.head_depth
AND a.adventure_id = adv.id
WHERE adv.head_depth >= 0 AND a.id IS NULL
""") == 0, "the head points outside the retained story"
finally:
connection.close()
+747
View File
@@ -0,0 +1,747 @@
"""M9: what a broken bundle does, and what it must never do.
A campaign bundle is a file on a disk. It can be truncated by a full volume,
mangled by a text editor, hand-written by somebody curious, or produced by a
build that does not exist yet. Every case below starts from a real export of the
M9 fixture and breaks exactly one thing about it, so what each test measures is
that one break rather than a fixture nobody would recognise.
## The two rules
**Nothing lands.** A refused import leaves no campaign, no branch, no orphan
action, no Save Point pointing at nothing, and no knowledge owned by a campaign
that does not exist. `bundle.plan` has no side effects and runs before a row is
written, and the endpoint commits once, so a refusal is a refusal — checked here
by counting rows before and after rather than by trusting the status code.
**Nothing is fetched, read or run.** A bundle is data. A URL in it is text, a
filename in it is text, and a path in it is text. No test here needs a network
guard to pass, which is the point: there is no code path that would use one.
## Refuse or repair, and why each is which
The two are not interchangeable and the choice is made per field, on one
question — *does a wrong value here make the rest of the campaign wrong?*
refuse the head, the tree, the audit trail
a head past the story misplaces every read of it; a node on a
branch that is not listed is a story with a hole; an audit record
naming a turn that is not there leaves state nobody can explain
repair a knowledge classification that is unreadable, a filename with a
path in it, a live flag nobody set
the value is not load-bearing for anything but itself
drop a Save Point that names no turn, a summary with no coordinate
a bookmark costs a bookmark; refusing the campaign to save it
would lose the story
What none of them ever is: **retarget**. A Save Point whose position is not in
the file does not get moved to a nearby one, because the reader named a position
and no other position is the one they named.
python -m pytest tests/test_m9_corrupt_bundles.py -v
"""
import copy
import json
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, memorybank, models
from app.database import Base, SessionLocal, engine, get_db
from app.knowledge import embeddings
from app.main import app
from app.routers import adventures
import m9_fixture
from fakes import ScriptedProvider
from test_m9_portability import StubDerived
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
embeddings._cache.clear()
setup = SessionLocal()
user = models.User(is_guest=False, email="corrupt@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model="test-model", embedding_model="stub-embed",
context_token_budget=4000, max_output_tokens=400,
))
adventure = models.Adventure(
user_id=user.id, title="Source campaign",
campaign_canon=m9_fixture.CAMPAIGN_CANON,
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
memorybank._vector_cache.clear()
embeddings._cache.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture(scope="module")
def _cache():
"""One place to keep the exported fixture between tests in this module."""
return {}
@pytest.fixture()
def good(client):
"""A real, valid export of the M9 fixture, ready to be broken."""
m9_fixture.build(client, client.adv_id)
response = client.get(f"/api/adventures/{client.adv_id}/export")
assert response.status_code == 200
return response.json()
# ------------------------------------------------------------------ the rules
def _counts() -> dict:
"""Every row that an import can create, per table."""
with SessionLocal() as db:
return {
model.__name__: db.query(model).count()
for model in (
models.Adventure, models.Branch, models.Action, models.Memory,
models.Summary, models.Checkpoint, models.StateEvent,
models.StateProposal, models.KnowledgeSource,
models.KnowledgeChunk, models.StoryCard,
)
}
def refused(client, payload, *, status=(400, 409, 413, 422)) -> str:
"""Imports expecting a refusal, and asserts that nothing at all landed."""
before = _counts()
response = client.post("/api/adventures/import", json=payload)
assert response.status_code in status, (
f"expected a refusal, got {response.status_code}: {response.text[:400]}"
)
assert _counts() == before, (
"a refused import wrote rows: "
f"{ {k: (before[k], v) for k, v in _counts().items() if before[k] != v} }"
)
body = response.json()
return str(body.get("detail", body))
def accepted(client, payload) -> int:
response = client.post("/api/adventures/import", json=payload)
assert response.status_code == 201, response.text[:500]
return response.json()["id"]
def broken(good: dict, **changes) -> dict:
return dict(copy.deepcopy(good), **changes)
# -------------------------------------------------------- format and version
def test_a_payload_that_is_not_an_object_is_refused(client):
for payload in ([], "a string", 7):
response = client.post("/api/adventures/import", json=payload)
assert response.status_code in (400, 422), response.text[:200]
def test_an_empty_object_is_refused(client):
assert "format" in refused(client, {}).lower() or "export" in refused(client, {})
def test_a_missing_format_is_refused(client, good):
payload = copy.deepcopy(good)
del payload["format"]
refused(client, payload)
def test_a_format_of_the_wrong_type_is_refused(client, good):
for wrong in (3, None, ["ai-dnd-adventure-v3"], {"v": 3}):
refused(client, broken(good, format=wrong))
def test_an_unsupported_future_version_is_refused_with_its_name(client, good):
detail = refused(client, broken(good, format="ai-dnd-adventure-v42"))
assert "ai-dnd-adventure-v42" in detail
# ------------------------------------------------------------- the tree graph
def test_an_action_on_a_branch_the_file_does_not_list_is_refused(client, good):
payload = copy.deepcopy(good)
payload["actions"][0]["branch"] = 99
assert "99" in refused(client, payload)
def test_a_branch_forking_from_one_listed_after_it_is_refused(client, good):
"""Which is also how a cycle is made impossible rather than detected.
A branch may only fork from a branch listed before it, so the graph is
acyclic by construction. Without it a lineage walk on a hand-edited file
would not terminate.
"""
payload = copy.deepcopy(good)
payload["branches"][0] = {"parent": 1, "forkDepth": 0}
refused(client, payload)
def test_a_branch_that_forks_from_itself_is_refused(client, good):
payload = copy.deepcopy(good)
payload["branches"][1] = {"parent": 1, "forkDepth": 3}
refused(client, payload)
def test_a_fork_with_no_depth_is_refused(client, good):
payload = copy.deepcopy(good)
payload["branches"][1] = {"parent": 0}
assert "depth" in refused(client, payload)
def test_an_action_with_no_depth_is_refused(client, good):
payload = copy.deepcopy(good)
payload["actions"][1]["depth"] = None
assert "depth" in refused(client, payload)
def test_an_action_with_a_negative_depth_is_refused(client, good):
payload = copy.deepcopy(good)
payload["actions"][1]["depth"] = -4
refused(client, payload)
def test_a_head_past_the_story_is_refused(client, good):
assert "ends at" in refused(client, broken(good, headDepth=10_000))
def test_a_head_depth_of_the_wrong_type_is_refused(client, good):
for wrong in ("3", 3.5, True, [3]):
refused(client, broken(good, headDepth=wrong))
def test_a_head_branch_that_is_not_listed_falls_back_to_the_root(client, good):
"""Repaired rather than refused, and the repair is the safe direction.
The head *depth* is checked against the story and refused when it disagrees,
because a wrong depth silently moves the reader. A head *branch* that names
nothing cannot be read at all, so there is no wrong position to land at —
the root is where a campaign with no chosen branch is read.
"""
payload = copy.deepcopy(good)
payload["headBranch"] = 77
payload.pop("headDepth") # the depth belongs to the branch it names
copy_id = accepted(client, payload)
with SessionLocal() as db:
adventure = db.get(models.Adventure, copy_id)
root = (
db.query(models.Branch)
.filter(models.Branch.adventure_id == copy_id,
models.Branch.parent_branch_id.is_(None))
.first()
)
assert adventure.head_branch_id == root.id
def test_two_actions_claiming_one_identity_are_refused(client, good):
"""Take parentage and the whole audit trail hang off these ids."""
payload = copy.deepcopy(good)
payload["actions"][1]["id"] = payload["actions"][0]["id"]
assert "both call themselves" in refused(client, payload)
def test_a_turn_whose_takes_are_all_dead_still_tells_one(client, good):
"""Repaired, because a turn with no live attempt disappears from the story."""
payload = copy.deepcopy(good)
for action in payload["actions"]:
action["live"] = False
copy_id = accepted(client, payload)
with SessionLocal() as db:
rows = (
db.query(models.Action)
.filter(models.Action.adventure_id == copy_id)
.all()
)
per_turn = {}
for row in rows:
per_turn.setdefault((row.branch_id, row.depth), []).append(row)
for group in per_turn.values():
assert sum(1 for row in group if row.live) == 1
def test_a_parent_naming_a_node_the_file_does_not_hold_is_ignored(client, good):
"""Dropped, not refused: a wrong parent costs a pager, not a campaign."""
payload = copy.deepcopy(good)
for action in payload["actions"]:
if action.get("parentId") is not None:
action["parentId"] = 999_999
copy_id = accepted(client, payload)
story = client.get(f"/api/adventures/{copy_id}").json()
assert story["actions"], "the campaign did not import"
def test_a_node_that_is_its_own_parent_does_not_loop(client, good):
payload = copy.deepcopy(good)
for action in payload["actions"]:
if action.get("id") is not None:
action["parentId"] = action["id"]
copy_id = accepted(client, payload)
with SessionLocal() as db:
assert db.query(models.Action).filter(
models.Action.adventure_id == copy_id,
models.Action.parent_id == models.Action.id,
).count() == 0
# And the pager still resolves rather than recursing.
assert client.get(f"/api/adventures/{copy_id}").status_code == 200
# ----------------------------------------------------------------- save points
def test_a_save_point_beyond_the_retained_story_is_dropped_not_retargeted(
client, good
):
payload = copy.deepcopy(good)
original = payload["checkpoints"][0]["name"]
payload["checkpoints"][0]["depth"] = 5_000
copy_id = accepted(client, payload)
landed = client.get(f"/api/adventures/{copy_id}/checkpoints").json()
assert original not in {point["name"] for point in landed}
assert all(point["depth"] < 5_000 for point in landed)
assert landed, "the good Save Point was lost with the bad one"
def test_a_save_point_on_a_branch_that_is_not_listed_is_dropped(client, good):
payload = copy.deepcopy(good)
payload["checkpoints"][0]["branch"] = 44
copy_id = accepted(client, payload)
landed = client.get(f"/api/adventures/{copy_id}/checkpoints").json()
assert len(landed) == len(good["checkpoints"]) - 1
def test_a_save_point_with_a_blank_name_is_dropped(client, good):
payload = copy.deepcopy(good)
payload["checkpoints"][0]["name"] = " "
copy_id = accepted(client, payload)
assert len(client.get(f"/api/adventures/{copy_id}/checkpoints").json()) == \
len(good["checkpoints"]) - 1
def test_a_checkpoints_section_that_is_not_a_list_costs_the_bookmarks_only(
client, good
):
copy_id = accepted(client, broken(good, checkpoints={"nope": 1}))
assert client.get(f"/api/adventures/{copy_id}/checkpoints").json() == []
assert client.get(f"/api/adventures/{copy_id}").json()["actions"]
# ------------------------------------------------------------ state and audit
def test_a_state_section_that_is_not_a_list_is_refused(client, good):
assert "list" in refused(client, broken(good, stateEvents={"a": 1}))
assert "list" in refused(client, broken(good, stateProposals="events"))
def test_a_state_event_that_is_not_an_object_is_refused(client, good):
payload = copy.deepcopy(good)
payload["stateEvents"][0] = "an event"
refused(client, payload)
def test_a_state_event_with_no_type_is_refused(client, good):
payload = copy.deepcopy(good)
payload["stateEvents"][0]["eventType"] = ""
assert "type" in refused(client, payload)
def test_an_event_naming_a_turn_the_file_does_not_hold_is_refused(client, good):
payload = copy.deepcopy(good)
payload["stateEvents"][0]["action"] = 424_242
assert "424242" in refused(client, payload).replace(",", "")
def test_a_proposal_naming_a_turn_the_file_does_not_hold_is_refused(client, good):
payload = copy.deepcopy(good)
payload["stateProposals"][0]["action"] = 424_242
refused(client, payload)
def test_an_event_naming_a_proposal_that_is_gone_keeps_its_coordinate(client, good):
"""`ON DELETE SET NULL`, as a file. The event is the accepted change.
A proposal can be deleted while the event it produced stands — the schema
says so — so an event whose proposal is not in the file is not a broken
file. It loses the pointer and keeps everything that makes it an audit
record: what changed, where, and who asserted it.
"""
payload = copy.deepcopy(good)
payload["stateProposals"] = []
copy_id = accepted(client, payload)
events = client.get(
f"/api/adventures/{copy_id}/state/events?limit=500"
).json()
assert len(events) == len(good["stateEvents"])
assert any(e["source"] == "manual_correction" for e in events)
with SessionLocal() as db:
assert db.query(models.StateEvent).filter(
models.StateEvent.adventure_id == copy_id,
models.StateEvent.proposal_id.isnot(None),
).count() == 0
def test_a_malformed_narrative_state_costs_the_state_and_not_the_campaign(
client, good
):
"""M5's rule, unchanged: a malformed document is normalised, not fatal.
The story is the valuable thing. A state section that arrives as nonsense
becomes an empty document — which is honest, because nothing in it can be
trusted — and every turn still imports.
"""
copy_id = accepted(client, broken(good, narrativeState={"entities": "wrong"}))
story = client.get(f"/api/adventures/{copy_id}").json()
assert len(story["actions"]) == len(
client.get(f"/api/adventures/{client.adv_id}").json()["actions"]
)
state = client.get(f"/api/adventures/{copy_id}/state").json()
assert state["document"]["entities"] == {}
def test_a_per_position_snapshot_that_is_not_an_object_is_dropped(client, good):
payload = copy.deepcopy(good)
for action in payload["actions"]:
if "narrativeStateAfter" in action:
action["narrativeStateAfter"] = "not a document"
copy_id = accepted(client, payload)
assert client.get(f"/api/adventures/{copy_id}").status_code == 200
# Arriving at such a position gives the empty document rather than a
# later position's state, which is M5's finding 3.
client.post(f"/api/adventures/{copy_id}/undo")
assert client.get(f"/api/adventures/{copy_id}/state").json()["document"]["facts"] == []
# -------------------------------------------------------------- knowledge
def test_a_knowledge_section_that_is_not_a_list_is_refused(client, good):
assert "list" in refused(client, broken(good, knowledge={"a": 1}))
def test_a_source_with_no_content_is_refused(client, good):
"""Refused rather than dropped, and M7 chose that deliberately.
A campaign whose imported Canon quietly did not arrive is a campaign whose
narrator has stopped being told the rules, and the reader has no way to
notice.
"""
payload = copy.deepcopy(good)
payload["knowledge"][0]["content"] = ""
assert "content" in refused(client, payload)
def test_a_source_with_an_unknown_classification_is_refused(client, good):
payload = copy.deepcopy(good)
payload["knowledge"][0]["classification"] = "gospel"
assert "classification" in refused(client, payload)
def test_a_source_that_is_not_an_object_is_refused(client, good):
payload = copy.deepcopy(good)
payload["knowledge"][0] = "canon.md"
refused(client, payload)
def test_an_unreadable_visibility_becomes_normal_rather_than_hidden(client, good):
"""Repaired, and in the direction that reveals rather than conceals.
Visibility is not a permission system — the person who imported the file can
always read it — so a source that should have been narrator-only and lands
as normal costs a spoiler in the prompt framing. The other direction would
silently withhold material the reader expects the narrator to use, with
nothing saying so.
"""
payload = copy.deepcopy(good)
for source in payload["knowledge"]:
source["visibility"] = "invisible"
copy_id = accepted(client, payload)
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
assert all(source["visibility"] == "normal" for source in library)
def test_a_content_hash_that_disagrees_is_recomputed_and_reported(client, good):
"""The one derived value in the file, and the only reason it is there.
The stored hash is recomputed from what actually arrived, so it always
describes the content. The file's own claim is not silently discarded
either: a mismatch means the file was edited after it was written, and the
reader is told on the source itself.
"""
payload = copy.deepcopy(good)
payload["knowledge"][0]["contentHash"] = "0" * 64
copy_id = accepted(client, payload)
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
edited = [s for s in library if s["content_hash"] != "0" * 64]
assert len(edited) == len(library)
detail = client.get(
f"/api/adventures/{copy_id}/knowledge/{library[0]['id']}"
).json()
assert "did not match" in detail["notes"]
def test_more_sources_than_the_cap_is_refused(client, good, monkeypatch):
from app.knowledge import importer
monkeypatch.setattr(importer, "MAX_SOURCES_PER_ADVENTURE", 2)
assert "limit" in refused(client, good)
def test_an_oversized_source_is_refused(client, good, monkeypatch):
from app.knowledge import importer
monkeypatch.setattr(importer, "MAX_SOURCE_BYTES", 32)
assert "larger than" in refused(client, good)
# ------------------------------------------------------------- provenance
def test_a_context_snapshot_that_is_not_an_object_is_dropped(client, good):
"""Evidence is restored verbatim or not at all. It is never guessed at."""
payload = copy.deepcopy(good)
payload["actions"] = [
{k: v for k, v in action.items() if k != "contextSnapshotZ"}
| ({"contextSnapshot": "the prompt was long"}
if m9_fixture.snapshot_in(action) else {})
for action in payload["actions"]
]
copy_id = accepted(client, payload)
page = client.get(f"/api/adventures/{copy_id}").json()
narrator = [a for a in page["actions"] if a["type"] == "ai"]
assert narrator
for action in narrator:
response = client.get(
f"/api/adventures/{copy_id}/actions/{action['id']}/context"
)
assert response.status_code == 404, "a mangled snapshot was restored"
def test_a_snapshot_whose_knowledge_block_is_nonsense_does_not_break_the_import(
client, good
):
payload = copy.deepcopy(good)
rewritten = []
for action in payload["actions"]:
snapshot = m9_fixture.snapshot_in(action)
if isinstance(snapshot, dict) and "knowledge" in snapshot:
snapshot["knowledge"] = ["not", "a", "report"]
rewritten.append(m9_fixture.with_snapshot(action, snapshot))
else:
rewritten.append(action)
payload["actions"] = rewritten
copy_id = accepted(client, payload)
assert client.get(f"/api/adventures/{copy_id}").status_code == 200
def test_a_snapshot_naming_an_impossible_source_is_relinked_to_nothing(
client, good
):
payload = copy.deepcopy(good)
rewritten = []
for action in payload["actions"]:
snapshot = m9_fixture.snapshot_in(action)
if not isinstance(snapshot, dict):
rewritten.append(action)
continue
for record in (snapshot.get("knowledge") or {}).get("used") or []:
record["source_id"] = -1
rewritten.append(m9_fixture.with_snapshot(action, snapshot))
payload["actions"] = rewritten
copy_id = accepted(client, payload)
with SessionLocal() as db:
from sqlalchemy.orm import undefer
for row in (
db.query(models.Action)
.filter(models.Action.adventure_id == copy_id)
.options(undefer(models.Action.context_snapshot))
):
snapshot = row.context_snapshot
if not isinstance(snapshot, dict):
continue
for record in (snapshot.get("knowledge") or {}).get("used") or []:
assert record["source_id"] is None
# ------------------------------------------------------------ summaries
def test_a_summary_with_no_coordinate_is_dropped_not_placed(client, good):
"""Placing it at a guess is how E03's leak would arrive by a new route."""
payload = copy.deepcopy(good)
payload["summaries"][0]["depth"] = None
copy_id = accepted(client, payload)
with SessionLocal() as db:
landed = db.query(models.Summary).filter(
models.Summary.adventure_id == copy_id
).count()
assert landed == len(good["summaries"]) - 1
def test_a_summaries_section_that_is_not_a_list_costs_the_summaries_only(
client, good
):
copy_id = accepted(client, broken(good, summaries="a paragraph"))
assert client.get(f"/api/adventures/{copy_id}").json()["actions"]
with SessionLocal() as db:
assert db.query(models.Summary).filter(
models.Summary.adventure_id == copy_id
).count() == 0
# ------------------------------------------------------------ caps and size
def test_more_actions_than_the_cap_is_refused(client, good, monkeypatch):
monkeypatch.setattr(limits, "MAX_ACTIONS_PER_ADVENTURE", 3)
monkeypatch.setitem(limits._BUNDLE_LIST_CAPS, "actions", 3)
assert "limit" in refused(client, good)
def test_more_branches_than_the_cap_is_refused(client, good, monkeypatch):
monkeypatch.setattr(limits, "MAX_BRANCHES_PER_ADVENTURE", 1)
monkeypatch.setitem(limits._BUNDLE_LIST_CAPS, "branches", 1)
assert "limit" in refused(client, good)
def test_a_body_past_the_import_ceiling_is_refused_before_it_is_parsed(client):
"""413 from the middleware, on the declared length, before any read."""
padding = "x" * (limits.MAX_IMPORT_BODY_BYTES + 1024)
response = client.post(
"/api/adventures/import",
content=json.dumps({"format": "ai-dnd-adventure-v3", "title": padding}),
headers={"Content-Type": "application/json"},
)
assert response.status_code == 413
assert "too large" in response.json()["detail"].lower()
# ------------------------------------------------- the transaction, not the plan
def test_a_failure_deep_inside_the_write_leaves_nothing_behind(
client, good, monkeypatch
):
"""The other half of atomicity, and the half the planner cannot provide.
Every test above is refused by `bundle.plan`, which has no side effects — so
they prove the *planner*, and a passing planner would look identical if the
write phase left debris. This one breaks something the planner has already
approved, half way through writing: the branches, the nodes, their
parentage, the memories, the head and the Save Points are all in the session
by then.
What must survive that is the whole transaction rolling back — every table,
not merely the adventure row. A half-written campaign is the outcome L01
forbids for a turn, and an import is the other place it could happen.
"""
from app import bundle as bundle_module
def explode(*args, **kwargs):
raise RuntimeError("simulated failure deep inside the write")
monkeypatch.setattr(bundle_module, "_write_summaries", explode)
before = _counts()
with pytest.raises(RuntimeError, match="simulated failure"):
client.post("/api/adventures/import", json=good)
assert _counts() == before, (
"a failed write left rows behind: "
f"{ {k: (before[k], v) for k, v in _counts().items() if before[k] != v} }"
)
def test_the_session_is_usable_after_a_failed_import(client, good, monkeypatch):
"""The rollback is explicit, so the next request is not poisoned by it.
Left to the session closing, a failure would leave the request's session in
a state the next caller inherits only by luck of pooling. `bundle_io` rolls
back and re-raises, so the very next import succeeds.
"""
from app import bundle as bundle_module
calls = {"n": 0}
original = bundle_module._write_summaries
def once(*args, **kwargs):
calls["n"] += 1
if calls["n"] == 1:
raise RuntimeError("simulated, once")
return original(*args, **kwargs)
monkeypatch.setattr(bundle_module, "_write_summaries", once)
with pytest.raises(RuntimeError):
client.post("/api/adventures/import", json=good)
copy_id = accepted(client, good)
assert client.get(f"/api/adventures/{copy_id}").json()["actions"]
# --------------------------------------------------------------- inert data
def test_a_url_in_a_bundle_stays_text(client, good):
"""H01/G08 for the import path: nothing in a file is ever fetched.
There is no allowlist to test and no request to intercept, which is the
result rather than a gap — the import has no code that could make one. What
is asserted is that the text arrives as text.
"""
payload = copy.deepcopy(good)
payload["knowledge"][0]["content"] = (
"# Sources\n\nSee https://example.invalid/secret.txt and "
"file:///etc/passwd and ![map](https://example.invalid/map.png)\n"
)
copy_id = accepted(client, payload)
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
detail = client.get(
f"/api/adventures/{copy_id}/knowledge/{library[0]['id']}"
).json()
assert "https://example.invalid/secret.txt" in detail["content"]
def test_a_path_in_a_bundle_never_becomes_a_path(client, good):
"""H08. `originalFilename` is metadata; the import stores no file."""
payload = copy.deepcopy(good)
for hostile in ("../../../etc/passwd", "/etc/shadow", "C:\\Windows\\hosts",
"....//....//etc/passwd"):
payload["knowledge"][0]["originalFilename"] = hostile
copy_id = accepted(client, payload)
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
for source in library:
assert "/" not in source["original_filename"]
assert "\\" not in source["original_filename"]
assert ".." not in source["original_filename"]
def test_a_title_that_looks_like_a_command_is_stored_as_a_title(client, good):
payload = broken(good, title="; rm -rf / #")
copy_id = accepted(client, payload)
assert client.get(f"/api/adventures/{copy_id}").json()["title"] == "; rm -rf / #"
def test_an_over_long_title_is_truncated_rather_than_refused(client, good):
copy_id = accepted(client, broken(good, title="A" * 5_000))
title = client.get(f"/api/adventures/{copy_id}").json()["title"]
assert 0 < len(title) <= 200
+365
View File
@@ -0,0 +1,365 @@
"""M9: every older bundle still imports, and none is reinterpreted.
A backup that stops importing is not a backup, so the importer keeps every
version it has ever written. That is the easy half. The hard half is the rule
`V1-ACCEPTANCE-TESTS.md` I07 states about the head and this file generalises:
> Do not reinterpret missing legacy data using modern assumptions that did not
> exist when the file was written.
An older file is missing things because its **format** could not carry them, not
because the campaign lacked them, and the two demand opposite treatment. A file
written before the head was carried opens at its tip, because tip was the only
position that format could represent — reproducing what it recorded. A file
written before state events existed opens with no state events, because
manufacturing an audit trail from the snapshots it does carry would be this
build's reading of a history it never saw, handed to a reader as the record of
what happened.
Each seam below is built by taking a real v3 export and removing exactly what
the older format could not hold. That is deliberate: a checked-in fixture file
drifts, and a hand-written one tests a shape nothing ever wrote.
python -m pytest tests/test_m9_legacy_bundles.py -v
"""
import copy
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, bundle, limits, memorybank, models
from app.database import Base, SessionLocal, engine, get_db
from app.knowledge import embeddings
from app.main import app
from app.routers import adventures
import m9_fixture
from fakes import ScriptedProvider
from test_m9_portability import StubDerived
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
embeddings._cache.clear()
setup = SessionLocal()
user = models.User(is_guest=False, email="legacy@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model="test-model", embedding_model="stub-embed",
context_token_budget=4000, max_output_tokens=400,
))
adventure = models.Adventure(
user_id=user.id, title="Source",
campaign_canon=m9_fixture.CAMPAIGN_CANON,
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
memorybank._vector_cache.clear()
embeddings._cache.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture()
def current(client):
"""A real v3 export of the M9 fixture, to age backwards from."""
m9_fixture.build(client, client.adv_id)
response = client.get(f"/api/adventures/{client.adv_id}/export")
assert response.status_code == 200
return response.json()
# ---------------------------------------------------- ageing a bundle backwards
def as_of(payload: dict, era: str) -> dict:
"""The same campaign as an export from an earlier era.
Each step removes only what that era's format genuinely could not carry, so
the result is the file a build of that vintage would have produced from this
campaign — not a mutilated modern one.
"""
older = copy.deepcopy(payload)
eras = ("pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head")
assert era in eras, era
reached = eras.index(era)
# M9 (v3): the evidence sections and the node identities.
older["format"] = bundle.TREE_FORMAT
for key in ("stateEvents", "stateProposals", "summaries"):
older.pop(key, None)
for action in older["actions"]:
for key in ("contextSnapshot", "contextSnapshotZ", "id", "parentId"):
action.pop(key, None)
for memory in older.get("memories") or []:
memory.pop("authority", None)
for source in older.get("knowledge") or []:
for key in ("sourceId", "parserVersion", "chunkingVersion"):
source.pop(key, None)
if reached == 0:
return older
# M7: the imported knowledge library.
older.pop("knowledge", None)
if reached == 1:
return older
# M5: the authoritative narrative state, its per-position snapshots, and
# the campaign's own canon.
for key in ("narrativeState", "campaignCanon"):
older.pop(key, None)
for action in older["actions"]:
for key in ("narrativeStateAfter", "stateChanges"):
action.pop(key, None)
if reached == 2:
return older
# M4: named Save Points.
older.pop("checkpoints", None)
if reached == 3:
return older
# M3: the chosen head. Such a file could only ever be read at its tip.
older.pop("headDepth", None)
return older
def bring_back(client, payload) -> int:
response = client.post("/api/adventures/import", json=payload)
assert response.status_code == 201, response.text[:500]
return response.json()["id"]
def _rows(adv_id, model) -> int:
with SessionLocal() as db:
return db.query(model).filter(model.adventure_id == adv_id).count()
def _tree_size(client, adv_id) -> int:
"""Every retained row, which is what "no accepted story was lost" means."""
return len(client.get(f"/api/adventures/{adv_id}/export").json()["actions"])
# --------------------------------------------------------------- every era
@pytest.mark.parametrize("era", [
"pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head",
])
def test_no_accepted_story_is_lost_at_any_seam(client, current, era):
"""The floor under every case below: the turns all arrive.
Counted over the whole retained tree rather than the active path, because
the head moves between eras and a count of what is on screen would move
with it.
"""
copy_id = bring_back(client, as_of(current, era))
assert _tree_size(client, copy_id) == len(current["actions"])
@pytest.mark.parametrize("era", [
"pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head",
])
def test_nothing_is_invented_to_fill_a_gap_the_format_left(client, current, era):
"""Absent means the format could not say. It never means "make one up".
Each era is checked against what that era's files could hold: a pre-M9 file
gets no audit trail and no summaries, a pre-M7 file no knowledge, a pre-M5
file no state, a pre-Save-Point file no Save Points.
"""
copy_id = bring_back(client, as_of(current, era))
reached = ("pre-m9", "pre-m7", "pre-m5", "pre-save-points",
"pre-active-head").index(era)
assert _rows(copy_id, models.StateEvent) == 0
assert _rows(copy_id, models.StateProposal) == 0
assert _rows(copy_id, models.Summary) == 0
if reached >= 1:
assert _rows(copy_id, models.KnowledgeSource) == 0
assert _rows(copy_id, models.KnowledgeChunk) == 0
if reached >= 2:
state = client.get(f"/api/adventures/{copy_id}/state").json()
assert state["document"]["facts"] == []
assert state["document"]["entities"] == {}
assert client.get(f"/api/adventures/{copy_id}").json()["canon_rules"] == []
if reached >= 3:
assert _rows(copy_id, models.Checkpoint) == 0
# ------------------------------------------------------ the head, era by era
def test_a_pre_m9_file_still_opens_at_the_head_it_recorded(client, current):
"""v2 carried the head, so it is honoured exactly as before."""
copy_id = bring_back(client, as_of(current, "pre-m9"))
with SessionLocal() as db:
assert db.get(models.Adventure, copy_id).head_depth == current["headDepth"]
def test_a_pre_active_head_file_opens_at_its_tip(client, current):
"""I07's compatibility clause. Not a degraded path.
Such a file was written when the head could not be anywhere but the tip, so
opening it there reproduces the position it recorded. An import that refused
it, or that guessed some other position, would be the failure.
"""
copy_id = bring_back(client, as_of(current, "pre-active-head"))
with SessionLocal() as db:
adventure = db.get(models.Adventure, copy_id)
tip = max(
row.depth for row in
db.query(models.Action).filter(
models.Action.adventure_id == copy_id,
models.Action.branch_id == adventure.head_branch_id,
)
)
assert adventure.head_depth == tip
assert adventure.head_depth > current["headDepth"], (
"the fixture's head must really be behind its tip, or this proves nothing"
)
def test_a_pre_active_head_file_offers_no_redo_because_it_is_at_the_tip(
client, current
):
copy_id = bring_back(client, as_of(current, "pre-active-head"))
page = client.get(f"/api/adventures/{copy_id}").json()
assert page["can_redo"] is False
assert page["can_undo"] is True
# ----------------------------------------------------- what each era can do
def test_a_pre_m5_campaign_can_be_played_on_and_gains_state_from_there(
client, current
):
"""The M5 rule, applied to an import: no backfill, and no obstacle either.
An old campaign starts with an empty state because its narration was never
read by a state extractor. The next turn fills it in, which is what makes
"no backfill" a decision rather than a loss.
"""
copy_id = bring_back(client, as_of(current, "pre-m5"))
assert client.get(f"/api/adventures/{copy_id}/state").json()["empty"] is True
ScriptedProvider.replies = [
"The door gives at last.\n" + __import__("fakes").state_block([
{"type": "add_fact", "predicate": "tally", "value": 500,
"fact_id": "tally-500"}
])
]
played = client.post(f"/api/adventures/{copy_id}/actions",
json={"type": "do", "text": "push harder"})
assert played.status_code == 200, played.text[:300]
after = client.get(f"/api/adventures/{copy_id}/state").json()
assert after["empty"] is False
assert any(f["predicate"] == "tally" for f in after["document"]["facts"])
def test_a_pre_m7_campaign_needs_no_source_and_can_import_one(client, current):
copy_id = bring_back(client, as_of(current, "pre-m7"))
assert client.get(f"/api/adventures/{copy_id}/knowledge").json() == []
# It plays without one.
assert client.get(f"/api/adventures/{copy_id}/context").status_code == 200
# And gains one.
landed = m9_fixture.upload(
client, copy_id, "canon.md", m9_fixture.CANON_MD, "canon",
)
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
assert [s["id"] for s in library] == [landed]
assert library[0]["index_state"] == "ready"
def test_a_pre_save_point_campaign_can_be_given_one(client, current):
copy_id = bring_back(client, as_of(current, "pre-save-points"))
assert client.get(f"/api/adventures/{copy_id}/checkpoints").json() == []
made = client.post(f"/api/adventures/{copy_id}/checkpoints",
json={"name": "From here", "note": ""})
assert made.status_code == 201, made.text[:300]
assert made.json()["resolved"] is True
def test_a_pre_m9_campaign_re_exports_as_v3_without_gaining_evidence(
client, current
):
"""Re-exporting an old campaign does not turn absence into presence.
The file it writes is a v3 file, because that is what this build writes. Its
evidence sections are empty, because the campaign genuinely has none — and a
later reader can therefore trust a v3 file's empty `stateEvents` to mean
"this campaign has no audit trail" rather than "the file could not say".
"""
copy_id = bring_back(client, as_of(current, "pre-m9"))
again = client.get(f"/api/adventures/{copy_id}/export").json()
assert again["format"] == bundle.FORMAT
assert again["stateEvents"] == []
assert again["stateProposals"] == []
assert again["summaries"] == []
assert not any(a.get("contextSnapshotZ") for a in again["actions"])
# And the story it does have survives a second round trip unchanged.
twice = bring_back(client, again)
assert _tree_size(client, twice) == _tree_size(client, copy_id)
def test_a_v1_file_still_imports_and_reads_in_order(client):
"""The flat format, with its retries as a repeating group."""
copy_id = bring_back(client, {
"format": bundle.LEGACY_FORMAT,
"title": "An old flat file",
"memory": "Kept from before the tree.",
"actions": [
{"index": 0, "type": "start", "text": "It begins."},
{"index": 1, "type": "do", "text": "look around"},
{"index": 2, "type": "ai", "text": "Take two.",
"variants": [{"text": "Take one."}, {"text": "Take two."}],
"variantIndex": 1},
],
})
page = client.get(f"/api/adventures/{copy_id}").json()
assert [a["text"] for a in page["actions"]] == [
"It begins.", "look around", "Take two.",
]
assert page["memory"] == "Kept from before the tree."
# Both attempts arrived; only one is the story.
with SessionLocal() as db:
rows = db.query(models.Action).filter(
models.Action.adventure_id == copy_id, models.Action.type == "ai",
).all()
assert sorted(r.text for r in rows) == ["Take one.", "Take two."]
assert sum(1 for r in rows if r.live) == 1
def test_a_pre_m2_file_with_scripting_still_imports(client, current):
"""M2 removed campaign scripting. Its keys are ignored, not rejected.
The story, the tree and everything else in such a file are still worth
importing, and refusing the campaign over a subsystem that no longer exists
would lose all of it to reject one key.
"""
payload = as_of(current, "pre-m5")
payload["scripts"] = [{"name": "onTurn", "code": "state.gold += 10"}]
payload["scriptState"] = {"gold": 70}
copy_id = bring_back(client, payload)
assert _tree_size(client, copy_id) == len(current["actions"])
File diff suppressed because it is too large Load Diff
+11 -7
View File
@@ -350,13 +350,17 @@ def test_export_and_import_round_trips_variants(client):
assert [a["type"] for a in actions] == ["start", "do", "ai"]
assert actions[-1]["text"] == "Two."
# The pager reads 1/1 on the copy, because the import writes no
# `parent_id` and `annotate_takes` groups on it. The attempts are both
# there, at one coordinate, and `GET .../variants` still lists them. This
# is a gap in the import rather than in the drop: `take_count` has been the
# only number the client reads since SP9, and the import has never set the
# column it is derived from.
assert actions[-1]["take_count"] == 1
# The pager reads 2/2 on the copy, as it does on the original.
#
# It read 1/1 until M9, and this test recorded that as a gap in the import
# rather than in the export: the attempts were both there at one coordinate
# and `GET .../variants` listed them, but the import wrote no `parent_id`,
# so `annotate_takes` grouped on the coordinate instead. That is right for a
# plain retry and wrong the moment two takes of one turn each have takes of
# their own beneath them, which is why M9 carried the parentage rather than
# leaving the pager to a fallback. See `bundle._link_take_parents`.
assert actions[-1]["take_count"] == 2
assert actions[-1]["take_index"] == 1
variants = client.get(
f"/api/adventures/{imported}/actions/{actions[-1]['id']}/variants").json()
assert [v["text"] for v in variants] == ["One.", "Two."]
+12 -2
View File
@@ -50,10 +50,20 @@ def payload() -> dict:
def test_the_shipped_file_is_a_bundle_this_build_can_import():
"""The file is written by an export, so a format change can strand it."""
"""The file is written by an export, so a format change can strand it.
It is checked against every version the importer reads rather than against
the newest one it writes, which is the property that actually matters and
the one the shipped file has to keep. M9 bumped the format to v3 and did not
regenerate this asset: the starter is a linear story with no state events,
no summaries and no stored prompts, so a v3 rewrite of it would differ from
the v2 file in the version string alone — and rewriting a shipped asset to
keep a test's equality holding would be changing the evidence to fit the
test. What it does need is to go on importing, which is asserted below.
"""
data = payload()
version = bundle.check_format(data)
assert version == bundle.FORMAT
assert version in bundle.READABLE
story = bundle.plan(data, version)
assert story["nodes"]
+9 -2
View File
@@ -21,7 +21,7 @@ import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, models
from app import auth, bundle as bundle_module, limits, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import adventures
@@ -425,6 +425,13 @@ def test_export_carries_the_whole_story(client):
array in `test_export_keeps_retry_attempts`. That change reflects the
same fact: a bundle that stores coordinates has no use for a repeating
group. Everything else here still passes unmodified.
M9 changed the same one line again, for the same kind of reason — the
version now says that the file can carry state events and historical
prompts as well as a tree. It is asserted against `bundle.FORMAT` this
time, so the next writer of a new version does not have to find this line:
what the test is about is that an export declares its version, not which
version this build happens to write.
"""
ScriptedProvider.replies = [gold_reply(t) for t in ["One.", "Two."]]
_play(client, "go north")
@@ -433,7 +440,7 @@ def test_export_carries_the_whole_story(client):
r = client.get(f"/api/adventures/{client.adv_id}/export")
assert r.status_code == 200, r.text
bundle = r.json()
assert bundle["format"] == "ai-dnd-adventure-v2"
assert bundle["format"] == bundle_module.FORMAT
assert bundle["title"] == "Cave"
assert [a["text"] for a in bundle["actions"]] == [
OPENING, "> You go north.", "One.", "> You go south.", "Two.",