M9: a campaign you can actually get back

A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
JesseMarkowitz
2026-09-07 01:55:45 -04:00
co-authored by Claude Opus 5
parent 1ce9972760
commit 44edece67e
46 changed files with 9227 additions and 178 deletions
+244 -7
View File
@@ -1800,6 +1800,16 @@ Export standard campaign.
### Pass
Export completes locally and contains enough data to restore story.
### Result — PASS (M9, 2026-09-07)
`ai-dnd-adventure-v3`, checked section by section rather than by file size:
the story and its whole retained tree, the branches and their disposition, the
chosen head, Save Points, the authoritative state and its per-position
snapshots, the state events and proposals, the imported library, the
lineage-anchored summaries, the memories, and a stored prompt for every narrator
turn that has one. `test_m9_portability.py::test_i01_*`, and reproducible with
`python -m tools.m9_portability_report`, which classifies every data family as
PRESERVED, OMITTED or DERIVED/REBUILDABLE.
---
## I02 — Import Exported Campaign
@@ -1814,6 +1824,14 @@ Export completes locally and contains enough data to restore story.
### Pass
Active transcript and state are restored.
### Result — PASS (M9, 2026-09-07)
Across a **genuine clean data directory**: a second server process, in a second
directory, against a database file that has never existed, with the exporting
process stopped. Transcript, authoritative state, canon, Save Points, knowledge,
state events and the size of the retained tree all match the source campaign
(`test_m9_clean_import.py`). The same round trip inside one process is in
`test_m9_portability.py`, and is labelled there as the weaker of the two.
---
## I03 — Branch/Disposable History Export
@@ -1823,6 +1841,19 @@ Active transcript and state are restored.
### Pass
Export preserves retained alternate/disposable history needed for recovery, unless user explicitly chooses a trimmed export.
### Result — PASS (M9, 2026-09-07)
Both futures come back and stay distinguishable: the abandoned line's turns are
present and readable, the branch that was left carries the depth it was left at,
and a superseded take is still at its coordinate with the selected take still
selected. There is no trimmed export — M9 offers no option to drop history, so
the clause about one does not arise.
**M9 also closed a fidelity gap here.** Take parentage was not exported, so every
imported node landed parentless and the pager grouped on the coordinate instead.
That is right for a plain retry and wrong once two takes of one turn each have
takes of their own beneath them; the copy read `5/5` where the source read `2/2`
and `3/3`. v3 carries the parentage.
---
## I04 — Checkpoint Export
@@ -1841,6 +1872,15 @@ campaign. Importing Save Points does **not** move the active head — the head
still comes from the bundle's `headDepth`. Bundles written before M4 carry no
`checkpoints` key, import cleanly, and create none.
### Re-verified — PASS (M9, 2026-09-07)
Unchanged by the format bump, and extended in two directions. Every restored
Save Point resolves, restores to the position it names through M3's head
movement, leaves the retained history it moved back over intact, and the two
in the M9 fixture restore to *different* states. A Save Point whose coordinate
is not in the file is **dropped with the rest of the campaign kept**, never
retargeted to a nearby turn: the reader named a position, and if that position
is not in the file then no other position is the one they named.
---
## I05 — Knowledge Provenance Export
@@ -1870,13 +1910,29 @@ hand-edited knowledge block with an unknown classification or empty content
refuses the import rather than half-landing in it; and an edited content hash is
recomputed from what actually arrived and the discrepancy recorded on the source.
**Limit, unchanged from before M7 and owned by M9.** The bundle carries no
context snapshots at all, so an imported campaign has no historical prompt
provenance — for imported knowledge or for any other component. Nothing M7
creates is turned into a dangling id by a round trip, because no ids are
exported; the evidence simply is not in the file.
`test_historical_prompt_evidence_survives_an_export_round_trip` pins that
behaviour so it cannot regress silently.
**That limit is closed (M9, 2026-09-07).** The paragraph below is kept as
written because it records what was true through M7 and M8, and the M9 decision
is only legible against it.
> *The bundle carries no context snapshots at all, so an imported campaign has no
> historical prompt provenance — for imported knowledge or for any other
> component. Nothing M7 creates is turned into a dangling id by a round trip,
> because no ids are exported; the evidence simply is not in the file.*
M9 carries the snapshots. An old turn in a restored campaign shows the prompt it
was actually assembled from, the passages it was shown and the text each
supplied — after the source has been deleted, the canon edited and the state
moved on. `test_historical_prompt_evidence_survives_an_export_round_trip` was
inverted rather than deleted: it now pins the thing M7 was worried about and
could not check, which is that the provenance arriving on the other side names
*this* campaign's sources rather than the ids they had where the file was
written. Only that pointer is translated; the evidence is restored verbatim, and
a source the file does not carry becomes `null` rather than pointing at a
different file.
Also added in M9: `parserVersion` and `chunkingVersion` per source, recording
what produced the passages a historical retrieval record describes, and a
`sourceId` that exists only so the translation above can be made.
---
@@ -1887,6 +1943,28 @@ behaviour so it cannot regress silently.
### Pass
No external API credentials are embedded in campaign export.
### Result — PASS (M9, 2026-09-07)
Tested rather than assumed, and tested against the file's **text** rather than
against a list of columns — a field added to a model the exporter walks would
otherwise reach the bundle with no test of a column noticing. The inert
`api_key` column is written with a recognisable value first, so its absence is
evidence rather than a tautology.
Nothing matching `api_key`, `apiKey`, the written secret, the inference endpoint
or its port, or an absolute filesystem path appears anywhere in an export —
including inside the compressed snapshots, which the test decodes rather than
skipping. Checked in both suites, so the clean-directory run covers the same
ground across a real process boundary
(`test_m9_portability.py::test_i06_*`, `test_m9_clean_import.py`).
**What deliberately does not travel**, and why it is not an omission: the
inference endpoint, the model name, the context budget and every other row of
`settings`. Those describe the machine, not the campaign, and importing a
campaign must not silently repoint the destination's inference at the source's.
Per-turn model and generation settings *do* travel, inside the historical
snapshot, because there they are a record of what happened rather than a
configuration to apply.
---
## I07 — Export/Import Preserves an Undone Active Head
@@ -1921,6 +1999,28 @@ An export whose stated head lies beyond the story it contains is a file
disagreeing with itself and must be refused rather than opened at a guessed
position.
### Result — PASS (M9, 2026-09-07)
All three clauses, and the main one across a genuine machine boundary.
- **The exact head.** The M9 fixture ends two Undos behind its own branch's
retained tip and behind the abandoned line's, and the last thing it does is an
Undo — so the head is not the newest row written, not the deepest row, not the
tip, and not on the branch holding the most story. An importer guessing any one
of those lands somewhere else. The copy opens exactly where the source was,
the later turns are still in the database as retained future, and Redo is
offered rather than the story having silently been redone. Redo then walks to
the same next turn in both.
- **The legacy clause.** A bundle with its `headDepth` removed opens at the tip
of its head branch, offers no Redo, and offers Undo — which is the position
such a file recorded, because at the time it was written the head could not be
anywhere else. Checked at every seam in `test_m9_legacy_bundles.py`.
- **The self-disagreeing file.** A head past the retained story is refused with
a message naming where the branch actually ends, and nothing is written.
Evidence: `test_m9_clean_import.py::test_it_opens_at_the_exact_head_it_was_exported_at`
(second process, empty directory), plus `test_m9_portability.py::test_i07_*` and
the head cases in `test_m9_corrupt_bundles.py`.
---
# J. Genre Independence
@@ -2082,6 +2182,17 @@ turn.
### Pass
State at each position matches original accepted state.
### Result — PASS (M9, 2026-09-07)
Measured **after a round trip**, which is the M9 form of it: the copy and the
source are walked back three turns and forward three turns in step, and the
authoritative document is compared at every position. They agree throughout.
That this stays a snapshot read rather than a replay is the point. Undo, Redo
and Save Point restore all resolve a coordinate and read the state recorded
there (`TECHNICAL-DESIGN.md` §10.4), so an import that carried the events and
dropped the per-position snapshots would have made every one of them
proportional to campaign length. Both halves of §17's hybrid travel.
---
## L03 — Checkpoint Reconstruction After Restart
@@ -2105,6 +2216,15 @@ exactly that value after the campaign had been advanced past it
`TestClient` restart, which could not distinguish durable state from a live
object.
### Re-verified after a move — PASS (M9, 2026-09-07)
The same claim with a machine boundary in front of it. A campaign is exported
from one server process, imported into a **second process against a database
file that has never existed**, a Save Point is restored there, that process is
killed, and a **third** process against the same file is asked again. The
transcript, the authoritative state and the size of the retained tree all match
what the second process had after restoring
(`test_m9_clean_import.py::test_l03_*`).
---
## L04 — Derived Data Can Be Rebuilt
@@ -2121,6 +2241,34 @@ using a safe test copy.
### Pass
Authoritative campaign history remains intact and derived structures can be recreated.
### Result — PASS (M9, 2026-09-07), and one defect found by running it
On a safe copy — an imported campaign, not the original. Every physically
derived structure is destroyed and rebuilt from the source content the bundle
carried: passages, the FTS rows and the vectors. Afterwards every source is
`ready` with passages again, retrieval works, and the transcript, the
authoritative state, the classifications and the lifecycle flags are identical
either side. Deleting the vectors alone leaves lexical retrieval working, which
is M7's rule that the lexical half is a production path and not a fallback. A
rebuild does not make an abandoned line's summary eligible.
**Running it found a real defect, which is fixed here.** The FTS5 index is a
virtual table, so no foreign key reaches it and no `ON DELETE CASCADE` covers
it: deleting a campaign dropped its passages and left one index row per passage
behind. Nothing read them — every search joins through `knowledge_chunks` — so
the leak was invisible until SQLite handed the freed primary key out again, at
which point the **next source imported into any campaign** failed with an
integrity error. Reindex could not repair it either, because `clear_index` finds
index rows *through* the chunks, and there were none. Both ends are closed: a
campaign's index rows are removed before it is deleted, and the index insert now
replaces a stale row rather than colliding with it — so a database already
carrying the leak repairs itself and needs no migration. See the M9 report's
findings.
The rebuild path M9 depends on is therefore implemented and tested rather than
assumed, which the SHOULD priority above did not require but the milestone did:
knowledge passages and indexes are omitted from the bundle precisely because
they can be rebuilt.
---
# M. Long-Run Test
@@ -2327,3 +2475,92 @@ L01-L03
```
This keeps implementation work tied to observable behavior rather than repository-specific architecture.
---
# P. M11 Test-Design Tasks — Not Yet Acceptance Tests
**Nothing in this section is part of the pass/fail contract.** These are three
behaviours a real play session against accepted M8 showed are worth testing, for
which **the pass criterion is not yet settled**. They are recorded here so the
release tester finds them where they will look, and they are deliberately not
written as A-through-O items: giving them IDs would imply a criterion has been
ratified when it has not, and would destabilise a document whose value is that
every item in it is decidable.
Each names **what must be settled** before it can become an acceptance test.
Full context is in the M9 report's §Y and, durably, in `BUILD-MILESTONES.md`
under the post-M8 playtest findings.
## P1 — The reader can tell where they are after history movement
**From:** a real session in which Undo worked correctly and the reader could not
tell which point in the story they had reached.
**Behaviour to test:** after Undo, Redo, a Save Point restore, or an edit to an
earlier turn, the reader can identify their current position in the visible
story without implementation terminology (`branch`, `head`, `node`, `depth`
remain forbidden at the surface).
**Settle first:** what the indicator *is*. "The reader can tell" is not
decidable as written — it needs an observable artifact, such as a named position
that changes with movement and is present in the DOM. `BROWSER-UX-SPEC.md` §8
now carries the requirement; **no wording is ratified**, and this cannot become
an acceptance test before one is.
## P2 — Narration length has a measurable directional effect
**From:** a reader who chose *2-4 paragraphs* and received substantially longer
replies.
**Behaviour to test:** the narration-length setting produces a **measurable
directional difference** in output length across repeated realistic turns —
brief shorter than standard, standard shorter than detailed — on at least the
reference 3B narrator and one stronger local narrator.
**Settle first:** the numbers, and the mechanism. Today the setting adds one
English sentence to the campaign instructions and changes **no** generation
budget, while a separate numeric hint derived from the global
`max_output_tokens` is identical for every setting (M9 report §Y). Until the
product decides what each setting *means* — and whether it moves the budget —
there is no threshold to test against. A directional test is stateable; an
absolute one is not, and this should not become an acceptance test that asserts
word counts nobody has ratified.
**Do not** turn this into a truncation test: the state block is emitted last and
hard truncation removes it.
## P3 — Multi-character identity continuity
**From:** four people in one scene, and narration that treated one of them as
two different people of the same name. **Root cause unknown** — the playtest
database was destroyed, so no evidence survives.
**Behaviour to test:** across a multi-turn scene with a protagonist and three
supporting characters, the story does not create duplicate characters, does not
duplicate a display name across two entities, does not drift the protagonist's
identity, does not misattribute dialogue, and does not have a character refer to
themself as a separate same-named character — and the authoritative state and
the assembled context do not disagree about who anyone is.
**Settle first:** which of those are **product** guarantees and which are
**model-quality** observations. They are not the same kind of claim and must not
share one verdict:
- *"The state never holds two entities with the same display name"* is
decidable and enforceable, and is a candidate acceptance test today. The
implementation currently permits it and reports nothing (M9 report §Y).
- *"The narrator never confuses two same-named characters"* is not a pass/fail
property of this application — it depends on the model — and belongs in M11's
realistic-model review with a recorded classification, not in this contract.
**A run of this must capture**, on any failure: pre-generation state, the exact
stored prompt snapshot, history, summaries, retrieved memories, imported
knowledge, narrator output, and model settings — then classify as a state,
context-assembly, derived-data, or model failure. M9 made all of that portable,
so a failing campaign can be exported whole and investigated elsewhere.
**Fixture:** the standard Continuity Test does not exercise this — its traps are
knowledge, authority, branch leakage and possession, and its on-stage cast is
effectively two people. A companion fixture is proposed in
`TEST-CAMPAIGN-FIXTURE.md`; the established fixture is deliberately unchanged.