M9: a campaign you can actually get back

A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
JesseMarkowitz
2026-09-07 01:55:45 -04:00
co-authored by Claude Opus 5
parent 1ce9972760
commit 44edece67e
46 changed files with 9227 additions and 178 deletions
+112
View File
@@ -887,6 +887,101 @@ has stopped being told the rules, with nothing to notice.
A bundle written before M7 has no knowledge section and imports with an empty
library, which is what such a campaign had.
### M9: the format became a version, and the evidence started travelling
**M9 bumped the format to `ai-dnd-adventure-v3`**, and the reason is a rule
rather than a preference. Everything M9 added *could* have been an optional key
read with `.get`, the way `persona`, `checkpoints`, `narrativeState` and
`knowledge` each were. That mechanism stops working at exactly this addition:
a v2 file carrying no prompt provenance is **ambiguous** — written before M9,
when no file could carry one, or by M9 from a campaign whose turns predate the
column? Those are different facts about the campaign and a reader has to be able
to tell them apart. It is the same distinction the head rule above draws when it
says a pre-M3 file opens at its tip *because that is the position such a file
recorded*. A version number is how a recovery file states what it was capable of
recording. The reader keeps every older version; only the writer moved.
What each version can be trusted to say:
```text
v1 a linear story, its turns, and its retries as a repeating group
v2 + the tree, the live flags, the after-snapshots, the chosen head,
Save Points, the narrative state document, imported knowledge
v3 + state events and proposals, historical prompt/context provenance,
lineage-anchored summaries, take parentage, memory authority
```
**Three categories, not two.** §31 below distinguishes authoritative from
derived, which was sufficient until M9 had to decide about stored prompts. They
are derived — a machine assembled them — and they must travel anyway, so the
rule the bundle applies has a middle category:
```text
chosen what a person decided: the story, the head, the takes, the
Save Points, the classifications, the canon. travels
evidence what happened, and what the application was told at the time:
the state events and proposals, the per-turn prompt and the
passages it was shown, the model and generation settings that
turn ran under. travels
rebuildable a deterministic function of what travels: knowledge passages,
the FTS index, embeddings, the branch lineage cache.
rebuilt on import
```
The test that separates evidence from rebuildable is **not** "could this be
recomputed" but "would a recomputation answer the same question". Rebuilding the
FTS index answers the same question it answered before. Rebuilding an old turn's
prompt does not — it would say what that turn *would be told now*, from today's
canon, today's sources and today's state, which is the opposite of what the
context inspector is for. Historical evidence is not a cache.
So v3 additionally carries, all restored verbatim:
- **`stateEvents` and `stateProposals`.** §17's hybrid keeps the events for
audit and the snapshots for restore; v2 carried only the snapshots, so a moved
campaign could be read at any position and could no longer say what changed
there, who asserted it, or what the value was before. A **manual correction**
was the worst case: the one state change no narration explains, and with the
events gone nothing distinguished it from something the story established.
Both tables travel, because the inspector reads both — the event says what was
accepted and the proposal says what the model asked for and what was refused.
- **A per-node context snapshot**, which is `SPECIFICATION.md` §6.1's exact
prompt/context record. It carries the assembled prompt section by section, the
passages retrieved with the text each supplied, which summary was eligible,
and the model and generation settings the call ran under — so an old turn can
still say what it was told after the source was deleted, the canon edited, the
chunker changed and the state moved on. Stored once per turn on the live
attempt, so a retried turn is not a multiplier.
- **`summaries`**, with the coordinate that decides eligibility. v2 carried only
the `storySummary` mirror, which has no lineage of its own, so a restored
campaign resumed with no usable long-story continuity — and a summary
belonging to an abandoned line stays ineligible after the move for the same
reason it was before it: eligibility is the coordinate lying on the active
capped lineage, not a stored flag.
- **Take parentage**, so attempts under two different takes of one turn stay two
pagers rather than merging into one.
- **Memory `authority`**, so a heuristic memory is not promoted to accepted
story by being moved (F07).
- **Per-source `parserVersion`/`chunkingVersion`**, recording what produced the
passages a historical retrieval record describes.
**Encoding.** A per-turn prompt contains the story so far, so one per turn is
O(turns²) in campaign length — measured at 20,797 bytes per turn at turn 20 and
55,291 at turn 120, 68% of a 9.7 MB file. The snapshot therefore travels as
`contextSnapshotZ`: the same JSON, zlib-compressed and base64-encoded, using the
same pack/unpack the database column already uses. Nothing is dropped or
summarised; the file is still JSON, and every other section of it is still plain
text. The plain `contextSnapshot` key is still read and takes precedence, so a
hand-edited file keeps importing.
**Two pointers are translated on import, and nothing else is.** Branch numbers
already were. M9 adds the `source_id` inside a restored retrieval record: it
names a row on the machine that wrote the file, so left alone it would point the
inspector's "open this source" at whatever holds that id here. Where the file's
own knowledge section contains the source it is repointed; where it does not — a
source deleted before the export — it becomes `null`, and the record keeps its
text and filename. The evidence is never rewritten; only the pointer is.
## 30. Deletion vs Archival
The system must distinguish:
@@ -916,6 +1011,23 @@ Retry, Undo, Restore, and branch switching must not silently perform permanent d
Derived data should be rebuildable where practical.
**"Where practical" does real work in that sentence, and M9 had to split this
list to act on it** (§29). Embeddings and the lexical and semantic indexes are
deterministic functions of content that travels, so a rebuild answers the same
question and they are not exported. A **summary** is not: it took a model call,
it describes a stretch of story that may since have been abandoned, and
regenerating one on another machine produces different prose about a different
reading — so it is derived, not practically rebuildable, and it travels with the
coordinate that decides whether it still applies.
The same reasoning puts **stored prompt/context snapshots** on the travelling
side, and they are not in either list above because they are neither: they are
not authoritative — nothing decides anything from them — and calling them
derived would invite a rebuild. They are *evidence*: a record of what the
application was told at the time, which a regeneration would not reproduce
because it would use today's canon, today's sources and today's state. §29 states
the three-way rule the export applies.
## 32. Provenance
Important information should answer: