Files
interactive-story/planning/reports/M9-IMPLEMENTATION-REPORT.md
T
JesseMarkowitzandClaude Opus 5 44edece67e M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 01:55:45 -04:00

88 KiB
Raw Blame History

M9 — Export, Backup, Recovery, and Migration Hardening

Implementation report, written for an independent reviewer.

This is a set of claims with the measurements attached. It is not a record of acceptance: M9 is implemented and verified, and acceptance comes after review.

Two conventions carried from M8, because they are what make a report checkable:

  • Where a claim names a boundary, the boundary was crossed. "A clean data directory" is a second process against a database file that has never existed. "A restart" is a different PID. "A backup" is opened as its own database and read. "Migration from M8" is a database built by a server running the signed M8 commit. Where a substitute was used, it is labelled.
  • Every defect found is numbered in §U, including the ones fixed before the final tree. Seven are recorded. Three were found by running M9's own tests rather than by reading them, two more by building the browser harness, and four of the seven predate M9 — they were reachable in M8 and nothing had gone looking.

A. Executive result

PASS.

Every required M9 condition is verified, by the boundary it names. The Definition of Done — a campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact — is satisfied and is measured three separate ways: in the suite, across two real server processes with two data directories, and in a real browser against a real narrator.

Bundle format ai-dnd-adventure-v3. v1 and v2 still import; every seam from pre-active-head onward is tested
Historical prompt provenance carried. The M8 handoff is closed
Story cards compatibility-only, and out of the narrator's prompt
Summaries / memories carried, with the coordinates that decide eligibility
Derived indexes not carried, and rebuilt — with the rebuild path fixed
SQLite backup SQLite online backup API, verified before it is kept
Schema change none, proved against a database M8's own code wrote
Backend 1,102 passed, 14 skipped, 0 failed in 14:54 (M8 baseline: 950)
Frontend 145 passed (was 132), lint and production build clean
Browser 36/36 checks, one frozen build, real narrator, two installations (§Q)
Docker production image builds clean; the backup works inside it, on the mounted volume

Nothing is claimed here that was only compiled. §S maps every acceptance item to the evidence that decides it, and names what is asserted from a synthetic substitute.


B. Repository and provenance state

Branch m9-recovery
Base 1ce99727606124b0f8b3dc3b88842410c874ee14 — M8: the browser becomes the storyteller
Base signature Good signature, RSA key 02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569
Working tree at start clean
HEAD at writing still the M8 base — M9 is staged, not committed (§32)
Upstream ancestry d72f7c1bda0f34fccd84afb7a25c34eb01c901de is an ancestor
LICENSE unchanged; git diff 1ce9972 -- LICENSE is 0 lines
Schema version 92 before, 92 after
Migrations none added. git diff 1ce9972 -- backend/app/migrations.py is 0 lines
Columns none added or removed. 0 mapped_column changes in models.py

Change inventory

backend/app/bundle.py                        928 +-   the v3 format, both directions
backend/app/backup.py                        278      NEW — the online backup
backend/app/routers/backups.py                75      NEW — two endpoints
backend/app/routers/adventures/bundle_io.py   59 +-   the two-phase import
backend/app/context/builder.py                63 +-   story cards out of the prompt
backend/app/knowledge/fts.py                  44 +-   the index leak (finding 1)
backend/app/knowledge/importer.py             22 +    clear_campaign_index
backend/app/schemas.py                        19 +    ImportedAdventureOut
backend/app/routers/adventures/crud.py         8 +    clear the index before delete
backend/app/main.py                            6 +-   register the router
backend/app/starter.py                         4 +-   materialize returns a pair

frontend/src/pages/Settings.jsx              +       the backup control
frontend/src/components.jsx                  +-      the file picker (finding 4)
frontend/src/api.js, styles/library.css      +

backend/tests/m9_fixture.py                  406      NEW — the portability fixture
backend/tests/test_m9_portability.py       1,186      NEW
backend/tests/test_m9_corrupt_bundles.py     689      NEW
backend/tests/test_m9_clean_import.py        460      NEW — two processes
backend/tests/test_m9_backup.py              451      NEW
backend/tests/test_m9_legacy_bundles.py      365      NEW
backend/tools/m9_portability_report.py       354      NEW — the baseline tool
backend/tools/m9_migration_proof.py          331      NEW
backend/tools/m9_scale_report.py             221      NEW
frontend/src/pages/backup.test.jsx           116      NEW
frontend/src/filePicker.test.jsx              85      NEW

Five existing test files changed. Each had pinned a gap M9 closes, and each was reworked to assert the new guarantee rather than deleted — the rule BUILD-MILESTONES.md records from M2 about moving instrumentation rather than removing it. §U-6 lists them.


C. The M8 portability baseline, measured

The brief required the baseline to be measured, not assumed. It was, with a tool a reviewer can rerun on either commit:

cd backend && python -m tools.m9_portability_report          # or --json

It builds the M9 fixture (§E) in a throwaway database, exports it, imports the result, and classifies every data family. Run on the M8 commit and on the M9 tree, the difference is the milestone's claim in reproducible form.

The fixture it measures

Deliberately not the standard Continuity Test: that one is shaped to read like a story, and this one is shaped to break a round trip. 22 actions, 2 branches, 2 Save Points on different branches, 1 superseded take, 12 state events, 12 proposals, 11 stored prompts, 2 memories, 2 summaries (one on the abandoned line, one at the head), 5 imported sources across all three classes including one disabled and one narrator-only, and a manual state correction.

The head ends behind the retained tip of its own branch and behind the abandoned line's, and the last thing the fixture does is an Undo — so the head is not the newest row written, not the deepest row, not the tip, and not on the branch holding the most story. An importer guessing any one of those lands somewhere else.

M8 — ai-dnd-adventure-v2, 29,630 bytes

Verdict Family
PRESERVED campaign identity, transcript, branches, branch disposition, active head, alternate takes, Save Points, narrative state (current), narrative state (per position), imported knowledge, memories, scene metadata, story cards
OMITTED take grouping — SP9 parentage, so imported nodes landed parentless
OMITTED state events — the audit half of §17's hybrid
OMITTED state proposals — what the model asked for, and what was refused
OMITTED manual corrections — indistinguishable from what the story established
OMITTED historical prompt/context — the M8 handoff
OMITTED retrieval provenance — which passages an old turn was shown
OMITTED per-turn model settings — what a historical turn ran under
OMITTED knowledge parser versions
OMITTED summaries — only the lineage-less mirror column travelled
OMITTED memory authority — a heuristic memory imported as accepted story

Two of those were visible from outside, through the API a reader reads: after a round trip the copy's state events and summaries differed from the source's. The rest were invisible until something needed them.

Derived and correctly not carried: knowledge passages, the FTS index, knowledge and memory embeddings, the branch lineage cache, derived status.

M9 — ai-dnd-adventure-v3, 90,117 bytes on the same fixture

Every family PRESERVED. No disagreement on any reader-visible family.

The file is 3.0× larger on this fixture; §R has the measurement on a long campaign, where the ratio is very different and the conclusion is not the obvious one.


D. The final bundle contract

Documented in DATA-MODEL.md §29 and TECHNICAL-DESIGN.md §9.3; the reasoning is repeated in bundle.py's own docstring, which is where the next maintainer will be standing.

Why the version was bumped

Everything M9 adds could have been an optional key read with .get, the way persona, checkpoints, narrativeState and knowledge each were — §9.1 and §9.2 each considered a bump and correctly declined one.

That mechanism stops working here, and the reason is the rule the format already lives by. A v2 file with no prompt provenance is ambiguous: written before M9, when no file could carry one, or by M9 from a campaign whose turns predate the column? Those are different facts and a reader has to be able to tell them apart — the same distinction I07 draws when it says a file written before the head was carried opens at the tip because tip was the only position that format could represent. A version number is how a recovery file states what it was capable of recording.

The family name is unchanged, deliberately: renaming it would break every reader for no gain, and PROVENANCE.md is where the fork is recorded.

v1   a linear story, its turns, and its retries as a repeating group
v2   + the tree, the live flags, the after-snapshots, the chosen head,
       Save Points, the narrative state document, imported knowledge
v3   + state events and proposals, historical prompt/context provenance,
       lineage-anchored summaries, take parentage, memory authority

The importer reads all three. A version it has never heard of is refused with its own name in the message, rather than read as the newest it knows.

Three categories, not two

§31 of DATA-MODEL.md distinguishes authoritative from derived, which was sufficient until M9 had to decide about stored prompts. They are derived — a machine assembled them — and they must travel anyway:

chosen         the story, the head, the takes, the Save Points, the
               classifications, the canon                          travels
evidence       the state events and proposals, the per-turn prompt and the
               passages it was shown, the model and generation settings
               that turn ran under                                 travels
rebuildable    knowledge passages, the FTS index, embeddings, the branch
               lineage cache                              rebuilt on import

The test separating the last two is not "could this be recomputed" but "would a recomputation answer the same question". A rebuilt FTS index answers the same question. A rebuilt prompt does not — it says what the turn would be told now, from today's canon, today's sources and today's state, which is the opposite of what the inspector is for. Historical evidence is not a cache, and M9 regenerates no prompt at any point in the import.

Required, optional, and what an absent key means

Key Required Absent means
format yes refused
actions in practice an empty campaign
branches, headBranch no the root branch
headDepth no the format could not say — open at the tip (I07)
checkpoints no the campaign had none, or predates M4
narrativeState, narrativeStateAfter no predates M5; empty document
knowledge no predates M7; empty library
stateEvents, stateProposals no predates v3 — no audit trail is invented
contextSnapshotZ / contextSnapshot no predates v3, or the turn has none
summaries no predates v3
id / parentId on a node no pre-SP9 rule: attempts.group falls back to the coordinate
authority on a memory no the column default

Integrity and validation

One derived value is in the file, and only because its purpose is to be checked: a knowledge source's contentHash. The import recomputes it from what arrived, stores the recomputed value, and records the discrepancy on the source where a reader can find it. The stated hash is never trusted and never silently discarded.

The two pointers that are translated

Branch numbers already were. M9 adds source_id inside a restored retrieval record: it names a row on the machine that wrote the file, so left alone it points the inspector's "open this source" at whatever holds that id here. It is repointed at the source that landed, or set to null where the file carries no such source — at which point the record still holds the passage's title, filename and text. The evidence is never rewritten; only the pointer is. chunk_id is deliberately untouched: passages are rebuilt and get new ids, so no translation exists, and chunk_index and heading_path still say which passage it was.


E. Story graph, head, and takes

Round-tripped and compared through the API a reader reads, never by row count. test_m9_portability.py, and again across two processes in test_m9_clean_import.py.

Property Evidence
Every accepted action, live and superseded the whole exported tree is the same size either side
Branch and fork relationships both branches, the fork depth, and the names
Active branch and head §I07 below
Retained tip the future past the head is in the database and Redo reaches it
Branch disposition the superseded branch still carries the depth it was left at (6)
Retries / alternate takes the superseded take is still at its coordinate; the selected one still selected
Which take is active exactly one live node per coordinate, decided by the writer, not read from the file
Timestamps createdAt restored per node, per branch, per Save Point, per event

The undo-head case (I07)

Exported with active head < retained tip, imported into a genuinely clean data directory, and the copy opens at exactly the exported head. The later turns are present as retained future, Redo is offered rather than the story having silently been redone, and Redo then walks to the same next turn in both campaigns.

The diverged case (I03)

Both futures survive and stay distinguishable — the abandoned line's turns are readable on their branch, and the branch that was left still records that it was left, at the depth it was left at. There is no trimmed-export option, so I03's clause about one does not arise.

The retry case, and a fidelity gap M9 closed

v2 exported no take parentage, so every imported node landed parentless and attempts.group fell back to the coordinate. That is right for a plain retry and wrong as soon as two takes of one turn each have takes of their own beneath them: those share a branch and a depth, so the copy read 5/5 where the source read 2/2 and 3/3. v3 carries the parentage, and the pager now reads identically in the copy and the original.

test_retry_variants.py had recorded the old behaviour as "a gap in the import"; that test now asserts 2/2 (§U-6).


F. Save Points

Name, note, coordinate, createdAt all restored; compared through the panel's own API
Resolution every restored Save Point reports resolved: true
Restore reaches the position it names, through M3's head movement — no second restore architecture
Later history intact after restoring; the retained tree is the same size
Two Save Points, two branches restore to different states, so they are not accidentally the same pointer
The head unmoved by importing them: a campaign exported at turn 30 with a Save Point at turn 12 opens at turn 30

A Save Point the file cannot satisfy

Dropped, with the rest of the campaign kept — and never retargeted. The decision and its reasoning are in bundle.py and TECHNICAL-DESIGN.md §9.2:

  • not a refusal, because a bookmark costs a bookmark, where refusing the campaign would lose the story to save the bookmark;
  • not a repair, because the reader named a position, and if that position is not in the file then no other position is the one they named.

Three cases are covered: a coordinate beyond the retained story, a branch the file does not list, and a name that is blank once trimmed. In each, the other Save Point in the fixture survives — asserted, because "dropped" must not mean "dropped them all".


G. Narrative state and the audit trail

What travels, and why each

The audit inventory was made from the contract rather than by serialising every table:

Record Travels Why
narrative_state (campaign) yes the authoritative document at the exported head
narrative_state_after (per node) yes the restore path. Without it, head movement becomes replay
state_events yes (new) §17's audit half: what changed, where, who asserted it, what it was before
state_proposals yes (new) the inspector reads both — the event says what was accepted, the proposal says what was asked for and refused
rejected/unparseable proposals yes the record that explains why the state does not say what the narration seems to say
derived_status no describes the last run of a background pass, not the story

Exporting only the accepted half would have kept the answers and lost every question, which is why both tables travel.

Verified after the move

  • Current authoritative state at the exported head: identical, compared as the document.
  • Historical positions: the copy and the source are walked back three turns and forward three turns in step, and the document is compared at every position. They agree throughout (L02).
  • Restoration stays O(1)-style: every position's snapshot travels, so Undo, Redo and Save Point restore remain one row read rather than a replay (TECHNICAL-DESIGN.md §10.4, and the M4 note that made it load-bearing).
  • A manual correction is still identifiable as one. This is the case the omission hurt most: it is the one state change no narration explains, and with the events gone nothing distinguished it from something the story established.
  • The two tables arrive linked: events point at proposals that are in this campaign, and proposals point at this campaign's own turns.

Refuse, repair, drop — decided per field

Rule
Refuse an audit record naming a turn the file does not contain. Unlike a Save Point, an audit record that quietly did not arrive leaves a campaign whose state cannot be explained — and the explanation is what a reader goes looking for exactly when something looks wrong
Accept a proposal with no action: that is a manual correction, which has a coordinate and no narration behind it
Keep the coordinate, drop the pointer an event whose proposal is missing. That is what the schema's ON DELETE SET NULL already says happens
Normalise a malformed state document — M5's rule, unchanged: the story is the valuable thing, and a malformed section should cost the section

H. Historical prompt and retrieval provenance

The M8 handoff (question B) is closed: the evidence travels.

SPECIFICATION.md §6.1 requires every accepted turn to record an exact prompt/context snapshot or reproducible equivalent. The product already stored one; the bundle did not carry it, so that guarantee was local to the machine that played the campaign. It is now portable.

Each node carries contextSnapshotZ: the assembled prompt section by section, the passages retrieved with the text each supplied, which summary was eligible, the token accounting, and the model and generation settings the call ran under. Restored verbatim — nothing is regenerated, and nothing is validated beyond its being an object, because imposing today's expectations on a record an older build wrote would be the retroactive reading the inspector exists to rule out.

Verified to survive each thing that could destroy it

A source disabled later the record is unchanged
A source deleted after import asserted directly: the sources an old turn was shown are deleted in the copy, and the turn still shows the same passages and the same text. The record holds the text, not a pointer to a row that can go away (IMPORTED-KNOWLEDGE-DESIGN.md §49-50)
A different chunking on the importing machine the record is not rebuilt from chunks, so it cannot move
Current state moved on the snapshot is a record, not a re-render
Canon edited later same
Summaries rebuilt same

And across a real machine boundary: on machine B, which never assembled that prompt and could not reassemble it, Inspect Context on an old narrator turn returns a prompt with its sections intact.

Retrieval provenance

Each used record keeps source title, filename, classification, visibility, chunk index, heading path, the scores each path gave it, how it was admitted, and the rendered text. source_id is translated (§D); everything else is evidence and is untouched.

Cost control

The prompt is stored once per turn on the live attempt (attempts.py), so a superseded take carries only its own slices — its reply, its state proposal, its token accounting. Asserted: no superseded take in the bundle carries sections or prompt. A campaign retried twenty times does not carry twenty prompts.


I. Imported knowledge

Preserved title, original filename (metadata only), content, classification, enabled, visibility, always-include, notes, media type, import timestamp, content hash, parser and chunking version, and a sourceId that exists only so historical retrieval records can be relinked
Rebuilt passages, FTS rows, vectors — before the import returns for the first two, so the campaign is searchable immediately
Needs the original path no. Verified with the exporting machine's process stopped
Disabled source still disabled, and still out of retrieval
Narrator-only source still narrator-only
Canon with always-include still both
Content byte-identical, by SHA-256, checked in the browser run too

L04 — the rebuild, and the defect running it found

On a safe copy (the imported campaign, not the original), every physically derived structure is destroyed — passages, FTS rows, vectors — and rebuilt from the content the bundle carried. Afterwards: every source ready with passages, retrieval working, and the transcript, authoritative state, classifications and lifecycle flags identical either side. Deleting the vectors alone leaves lexical retrieval working, which is M7's rule that lexical is a production path and not a fallback. A rebuild does not make an abandoned line's summary eligible.

Running it found finding 1 (§U), a pre-existing defect that made Reindex — the documented repair — unable to repair the state it most needed to.

The M7 calibration rule

Untouched. nomic-embed-text keeps its measured calibration and an uncalibrated embedding model still degrades to lexical-only rather than borrowing the 0.58 threshold; no M9 code path reads or writes the threshold, and the M7 calibration tests pass unchanged.


J. Summaries, memories, and derived data

Summaries — carried, with their lineage

v2 carried only adventures.story_summary, the mirror column with no lineage of its own, so a restored campaign resumed with no usable long-story continuity. v3 carries the rows: text, coordinate, source range, trigger, model, timestamp.

Eligibility is not a stored flag — it is whether the coordinate lies on the active capped lineage — so restoring the coordinate restores the answer, including the "no". The fixture has one generated summary on the line it abandons and one typed at the head it keeps, and after the move:

  • exactly one is eligible, and it is the one at the head;
  • the abandoned line's text does not appear in the prompt the copy would send next;
  • the same holds after a full index rebuild, so M6's leak does not return by the rebuild path either;
  • story_summary is re-derived from the lineage after import rather than trusted from the file, because a mirror disagreeing with the lineage is exactly M6's finding M6-F1.

Memories — and an authority that was being promoted

Text, pinned, forgotten, source range, use count, coordinate, and now authority. Without it every imported memory landed on the column default, which silently promoted a heuristic memory to accepted_story — the one direction F07 forbids. An unreadable value is read as heuristic, which is the safe direction: a demoted memory costs a little ranking weight, a promoted one puts a guess in front of the narrator as fact.

Vectors

Not carried, in either subsystem, and rebuilt against whatever embedding model this machine has. Exporting them would tie a campaign to one machine's model, which is the reason M7 gave and M9 keeps.


K. Story cards — the compatibility decision

Settled: compatibility-only legacy data, and no longer in the narrator's prompt. Recorded in IMPORTED-KNOWLEDGE-DESIGN.md §73.

What was found

Story cards were the "alternate untracked path around the new knowledge authority/provenance rules" that §73 already forbade, and not in principle. A keyword-matched card was injected into the narrator's prompt as World Lore: <entry>, taking up to 40% of what remained after the imported knowledge had been placed, with:

  • no class, so nothing framed how far the narrator could rely on it;
  • no visibility, so no narrator-only distinction existed;
  • no source, no hash, no lifecycle — nothing to disable;
  • no browser surface at all since M8 removed the editor;
  • and no row in the context inspector, which renders knowledge and never rendered cards.

It also competed with imported Canon for one budget, which is the arrangement M7 spent a milestone separating.

What happens now

New export carries them, unchanged, under storyCards
Legacy import accepted unchanged, from every format version
Normal narration not reached. The world_lore section is gone
Re-export carries them again — a round trip destroys nothing

Nothing is deleted: the rows stay, the /api/story-cards endpoints stay, and memorybank.cast_brief still reads them as the summariser's character roster — that names who is on stage so a memory says "Aldric" rather than "he", never reaches the narrator, and every memory written from it is authority-classified by the application afterwards. The cards key stays in the context report and is now always empty for a new turn, because M9 has just made historical snapshots portable and an old turn's record must go on saying that story cards were included. That is asserted.

This is the smallest safe change: it closes the one path that asserted campaign facts to the narrator without any of §73's controls, and touches nothing else.


L. Scene and media metadata

No media tables exist in this build, and none were invented. Two active documents say that is the right answer rather than a gap:

  • K04's second branch — "if media tables are deferred, architecture/types should demonstrate equivalent extension point".
  • MEDIA-EXTENSION-CONTRACT.md §58, on export, in as many words: "For v1, if media is not implemented, preserve schema compatibility."

M9 preserves it in the strongest available form: the format is now versioned, so M10 can add media sections as v4 and every v3 reader keeps working, and every v3 file keeps importing. Inventing an empty media_assets table to satisfy the word "metadata" would have been schema M10 then had to live with.

What does exist is the scene section of the authoritative narrative state document, which is SPECIFICATION.md §14's scene snapshot as this build represents it. It is campaign-scoped and per-position, and it round-trips with the document: asserted both on the campaign-level state and on the per-node snapshots.

Nothing in M9 adds image, video, TTS, STT, a media provider, or a job architecture.


M. Model and context-window portability

The M8 handoff's question D, answered by keeping two things apart.

What travels

Per-turn model and generation settings, inside the historical snapshot — model name, API mode, temperature, max output tokens. There they are a record of what happened, which is exactly §6.1's requirement.

Campaign-owned configuration: title, narrator instructions, persona, canon, memory/summary switches, the campaign's own plot fields.

What does not, and why that is the point

The settings row — endpoint, model, context budget, timeouts, embedding model. Those describe the machine, not the campaign. Asserted directly on the importing machine: after importing a campaign played against a configured endpoint, machine B's endpoint_url is still its own default and its context_token_budget is still 16,384. Importing a campaign is not a way to reconfigure the destination's inference.

Import does not depend on a model

Asserted: machine B has no model configured at all, and the campaign still imports, opens, and shows its state and transcript. Playing on would fail; recovering does not. Campaign restored successfully is not the destination narrator has the same capacity, and M9 does not treat it as such.

The ceiling, and what M9 does about it

No assumption that 16,384 tokens are available is hard-coded anywhere M9 touched, and no native context-window detection architecture was built — that is M11's, with the 100-turn certification.

What M9 adds is the sentence a deployer needs, in DEVELOPMENT.md: a long imported campaign fills a prompt on its first turn, so a machine that has applied neither the derived-model procedure nor a matching budget meets its ceiling immediately rather than gradually. Check /api/ps on the destination before playing on an imported campaign.


N. SQLite backup

backend/app/backup.py, backend/app/routers/backups.py, and a control in Settings → Advanced.

Why not a file copy

A copy taken while the application runs is a copy of a moving target: SQLite writes in pages, and a plain copy can read page 5 before a transaction and page 900 after it, producing a file that opens, reports a schema, and is quietly missing rows. Nothing warns anyone, and it is found later.

So it uses SQLite's online backup API (sqlite3.Connection.backup), which copies under the right locks and yields a transactionally consistent snapshot of a committed point. The application keeps running; no session is closed and no turn is blocked.

The procedure's guarantees, each tested

Source opened read-only via a mode=ro URI, so it is a guarantee and not an observation. The source file is byte-identical after a backup
Temporary destination, finalised by rename the .partial file is gone and the final name never wore it
PRAGMA quick_check before the rename a failed verification leaves nothing behind
Never overwrites two backups get two names; two in the same second get two names
Failure reported a clear error with the reason; the source untouched
No caller-supplied path the endpoint accepts no request body at all — asserted against the OpenAPI schema, and by sending one anyway and checking where the file landed (H08)

Opened independently, and read

Every backup test opens the copy as its own database and asks what is in it — a test that only checked a file appeared would pass against a cp. Verified present in the copy: the schema, campaigns with their titles, story rows, the head branch and depth naming a turn that exists in the copy, Save Points, knowledge sources, state events, and the authoritative state document decoded and compared against the live campaign's.

Taken under load

The test that separates this from a file copy: turns are played from a second thread throughout the copy. The resulting file passes quick_check, has no dangling foreign key, has a transcript with no gap in its depths, and a head that names a turn present in it. Which committed point it caught is not asserted — that would be asserting on a race.

No restore endpoint, and no download

Both are decisions, stated in the code and in the panel. Restore means replacing the file the running process has open, which is how both copies are lost at once; the stop-move-start procedure is in DEVELOPMENT.md. Download would put a copy of every campaign into the browser's download directory and cache, which for an application whose premise is that the story does not leave the machine is a worse default than a path the reader can copy.


O. Import validation and corrupt bundles

test_m9_corrupt_bundles.py. Every case starts from a real export of the M9 fixture and breaks exactly one thing, so what each test measures is that break rather than a hand-written shape nothing ever wrote.

Two rules, checked rather than assumed

Nothing lands. A refused import is checked by counting rows in eleven tables before and after, not by trusting the status code. No campaign, no branch, no orphan action, no Save Point pointing at nothing, no knowledge owned by a campaign that does not exist.

Nothing is fetched, read or run. A URL in a bundle stays text; a path stays text. No test needs a network guard to pass — there is no code path that would use one.

Coverage

Class Cases
Format / version not an object; empty object; missing format; format of four wrong types; an unsupported future version, refused with its own name in the message
History graph action on an unlisted branch; branch forking from one listed after it (which is also how a cycle is made impossible rather than detected); a branch forking from itself; a fork with no depth; an action with no depth; a negative depth; two actions claiming one identity; a turn with no live take; a parent naming a node the file lacks; a node that is its own parent
Head past the retained story (refused, message names where the branch ends); four wrong types; a head branch that is not listed (repaired to the root, and the reasoning for why that one is repaired)
Save Points coordinate beyond the story; branch not listed; blank name; the whole section not a list
State sections not lists; an event that is not an object; an event with no type; an event naming a turn the file lacks (refused); a proposal naming a turn the file lacks; an event whose proposal is gone (keeps its coordinate); a malformed state document; a malformed per-position snapshot
Knowledge section not a list; no content; unknown classification; a source that is not an object; unreadable visibility; content hash mismatch, recomputed and reported; over the source cap; over the size cap
Provenance a snapshot that is not an object (dropped, never guessed at); a nonsense knowledge block; a snapshot naming an impossible source (relinked to null, evidence intact)
Summaries no coordinate (dropped, never placed at a guess — that is E03's leak by a new route); section not a list
Caps actions, branches, and a body past the import ceiling (413 from the middleware, on the declared length, before any parse)
Inert data URLs, file://, remote images; four traversal filenames; a title that looks like a command; an over-long title

The authoritative/derived distinction on failure

Two outcomes a caller can see, and they are distinguishable:

4xx   the authoritative import failed        and no campaign exists
201   the authoritative import succeeded     and the campaign is complete

A rebuildable index failing is neither, and does not roll the campaign back: it is reported on the 201 as import_warnings, visible per source in the Knowledge panel, and repaired by Reindex. The response model is a subclass (ImportedAdventureOut) rather than a field on AdventureOut, because "which of your indexes failed to rebuild" is a fact about one import, not a property of a campaign.

The rollback is explicit rather than left to the session closing, so a test can assert on it rather than on teardown.


P. Migration and backward compatibility

No schema change, proved

git diff 1ce9972 -- backend/app/migrations.py    →  0 lines
mapped_column additions/removals in models.py    →  0
PRAGMA user_version                              →  92 before, 92 after

git diff is necessary and not sufficient — a migration can also be missing, and the failure then is a database that opens and quietly answers wrongly. So the claim is proved the way M8 proved its own:

A campaign is built by a server running the signed M8 commit, from the git worktree at 1ce9972, using M8's own interpreter and M8's own code: story, an alternate take, a Save Point, a manual state correction, memories, a summary, three imported sources including a disabled and a narrator-only one, per-turn context snapshots, and two Undos so the head is behind the tip. M9 then opens that file.

$ python -m tools.m9_migration_proof open <db> --expected <m8-report.json>
{ "problems": [],
  "schema_version_before": 92, "schema_version_after": 92,
  "schema_version_second_open": 92,
  "checked": { "families": 12, "snapshot_rows": 8,
               "counts": { "actions": 16, "checkpoints": 1,
                           "knowledge_chunks": 3, "knowledge_sources": 3,
                           "memories": 2, "state_events": 9,
                           "state_proposals": 9, "summaries": 1 } } }

Twelve families compared and identical; a second open migrates nothing (idempotence); and the M8-built campaign then exports as v3 and round-trips, and Redo still works on it. Neither half of the tool imports the other's code — what crosses is the database file and one JSON report.

Older bundle seams

test_m9_legacy_bundles.py ages a real v3 export backwards, removing only what each era's format genuinely could not carry — a checked-in fixture drifts, and a hand-written one tests a shape nothing ever wrote.

Era No story lost Nothing invented Head
pre-M9 (v2) ✓ no events, no proposals, no summaries, no prompts honoured, as v2 carried it
pre-M7 ✓ + no knowledge sources, no chunks honoured
pre-M5 ✓ + empty state, no canon honoured
pre-Save-Point ✓ + no Save Points honoured
pre-active-head ✓ + no Save Points opens at the tip, offers no Redo
v1 (flat) ✓ both attempts arrive, one live tip
pre-M2 (with scripts) ✓ scripting keys ignored, not rejected as its era

And each era can still be used, not merely read: a pre-M5 campaign gains state from its next turn, a pre-M7 campaign imports a source and it indexes, a pre-Save-Point campaign can be given one.

A pre-M9 campaign re-exports as v3 with its evidence sections empty, because the campaign genuinely has none — which is what lets a later reader trust a v3 file's empty stateEvents to mean "this campaign has no audit trail" rather than "the file could not say".


Q. Browser recovery workflow

36/36 checks, zero failures, on one frozen production build against a real narrator across two installations.

Environment

Browser Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1)
Application the frozen production build dist/assets/index-DcHbz7ga.js, served by the real backend as static files — not a dev server
Backend two uvicorn app.main:app processes on loopback, one per machine, each with its own data directory
Databases machine A: a real campaign. Machine B: a file that had never existed, created by the migrations on first start
Narrator qwen2.5:3b-instruct-16k on the trusted-LAN Ollama, over HTTPS with a privately issued certificate — real inference for every turn
Campaign played over real HTTP: five turns, a retry, a Save Point, three imported sources (one disabled, one narrator-only), a manual state correction, and two Undos so the export is taken behind the retained tip

The §25 workflow, item by item

1. Export campaign PASS. Clicked in the browser; the page produced a 60,160-byte ai-dnd-adventure-v3 bundle
2. Return to Campaigns PASS
3. Import into a clean target PASS. Machine B had never had a database, its library was empty, and it had no narrator configured at all. The import reported no incomplete rebuild
4. Open imported campaign PASS
5. Opens at the exact active position PASS. The transcript matches the source exactly; Redo is offered; the retained future is in the database (11 rows behind 7 on screen); and all 4 narrated turns are found in the rendered page
6. State is correct PASS. The authoritative document matches, the manual correction is in the restored audit trail, with the fact the reader asserted, and the whole trail arrived
7. Save Points present PASS, in the browser's own panel, with the note
8. Knowledge present, with classifications and status PASS. All three sources in the panel; the disabled one still disabled, the narrator-only one still narrator-only, every source reindexed and searchable, and the content byte-identical by hash
9. Historical Inspect Context works PASS. 3 of 3 restored narrator turns show the prompt they were given; 3 name the passages they were shown; and the inspector opens through the browser's own control
10. Undo/Redo coherent PASS. Both offered at the restored head; Redo moves forward into the retained future
Backup control PASS. Taken from Settings; the panel names the file; the file is on disk where it says; it passes its own integrity check opened alone, and holds the restored campaign with all 11 of its rows
I06 on the downloaded file PASS. No credential, no endpoint, no path from the exporting machine

Two substitutions, and both are on the browser's side of the line

Named precisely, because §30 of the brief is about not letting a substitute stand in for the boundary it bypasses.

  1. The export's write to disk. downloadJSON builds a Blob, makes a blob: URL and clicks an anchor. Headless Firefox does not complete that download in this configuration — the click fires, the page reports "Campaign exported.", and no file appears in the download directory, in ~/Downloads, or anywhere under the snap's own tree. So URL.createObjectURL was wrapped to capture the exact bytes the page handed the browser, and those bytes were written to disk and used for everything downstream. Proved: the browser assembled the correct v3 file and offered it. Not proved: Firefox's own download writer, which is browser behaviour.
  2. The import's read of the file. Firefox here is a snap, and its sandbox will not let the content process read a path outside its confinement. WebDriver hands the input the file successfully — the File arrives with the right name and the right byte count, and change fires — and FileReader then fails with NotFoundError. The click and the handover are exercised; the last hop, the POST the page would make with the bytes it read, was made against the same endpoint the page calls.

Neither is papered over, and neither leaves the end-to-end path unproven: test_m9_clean_import.py does the whole of it — export, a real file on disk, a real HTTP import into a second process against a database that never existed — with no browser in the way, 11/11.

There is no unconfined Firefox and no Xvfb on this machine, so a windowed run was not available. That is a property of the machine and is recorded as such.

One observation worth a reviewer's attention

The real narrator emitted no typed state blocks across five turns. qwen2.5:3b-instruct-16k narrated well and never produced a state block, so the campaign's only state event is the manual correction — which is why the audit-trail checks above read "1 events, source had 1".

That is not an M9 defect and not a regression: it is the behaviour ADR 003 and ADR 010 exist for. The application owns state and the model does not, C04's manual correction path is how a reader asserts what the model did not, and the per-position snapshots keep the head, the transcript and the state in agreement regardless. It does mean the browser run's state evidence is thinner than the suites', where a scripted narrator produces twelve events and two proposals — so the rich state case is proved there and the real-model case is proved here, and this report does not let the second stand in for the first.


R. Performance and size

The M9 fixture

M8 (v2) M9 (v3)
Bundle 29,630 B 90,117 B
Actions 22 22
Source database 258 kB 319 kB
Export 0.028 s 0.044 s
Import (write + rebuild) 0.063 s 0.119 s

A long campaign — and the result that changed the plan

python -m tools.m9_scale_report --turns 150 --budget 16384. A real prompt builder, a 16,384-token budget, and prose long enough that a turn is a turn.

 turns  actions     bundle B   snapshots B    B/turn  export s  import s
    25       51      311,346       130,766     5,230     0.066     0.171
    50      101      883,616       279,208     5,584     0.129     0.335
    75      151    1,715,306       439,634     5,861     0.269     0.232
   100      201    2,870,121       675,880     6,758     0.572     0.371
   125      251    4,313,937       951,870     7,614     0.518     0.603
   150      301    6,022,430     1,242,324     8,282     1.012     1.109

A per-turn prompt contains the story so far, so carrying one per turn is O(turns²). That was the risk in the decision, and measuring it changed what was done about it and what is claimed:

  • Uncompressed, the snapshots were 68% of a 9.7 MB file at 120 turns. They now travel compressed inside the bundle — the same JSON through the same compression.pack/unpack the database column already uses, base64-encoded so the file is still JSON. 7.9× smaller, and every other section of the file is still plain readable text.
  • At M11's 100-turn certification target: 2.87 MB, 14% of the cap.
  • Export and import stay under 1.1 s at 150 turns.

Where the bytes are, and what M9 actually cost

The same tool's per-section breakdown, at 150 turns. It exists because "the prompts are 21% of the file" does not answer "what did M9 add" — the events, the proposals and the summaries are v3 additions too, and a cost claim counting only the prompts would understate it.

  per-position state (v2 already)       4,474,294   74.3%
  prompts (contextSnapshotZ)            1,245,144   20.7%
  state proposals                          78,938    1.3%
  state events                             41,870    0.7%
  node ids + parentage                      6,993    0.1%
  summaries                                 3,958    0.1%
  --- everything v3 added               1,376,903   22.9%
  a v2 file of the same campaign         4,645,347
  ceiling with v3 additions:          ~279 turns
  ceiling without them:               ~318 turns

The result is not the one the decision was worried about. The largest thing in a campaign bundle is not the prompts M9 added — it is the per-position narrative state document, at 74% of the file, which v2 already carried. It grows because a state document accumulates facts, so it is quadratic for the same reason the prompts are, and it has been since M5.

Everything M9 added comes to 22.9% of the file, and moves the import ceiling from about 318 turns to about 279 — a cost of roughly 12% of reachable campaign length for closing the provenance gap, not the halving it looked like before it was measured, and not the 11% a prompts-only reading would have claimed.

Export and import stay under 1.2 s at 150 turns.

N+1 and duplication

Checked for, and none found:

  • The export undefers every deferred per-node column in one query, including context_snapshot, and the proposals' detail once for the whole list.
  • Nodes are referenced by their own exported ids, so the audit records need no per-row lookup.
  • Knowledge content appears once per source; the prompt appears once per turn, on the live attempt, with superseded takes carrying only their own slices (asserted).
  • The import's flushes are bounded by the number of knowledge sources, not by the size of the story: one after the nodes, one after the proposals, two in materialize, and one per source — the last because build_index needs the source's id, which is the same one-flush-per-source the live upload path already does. Nothing flushes per action, per branch, per event or per Save Point, which is where a long campaign would have felt it.

Residual

The ceiling above is a residual limit, stated with its measurement. It is far beyond M11's certification target, and the fix if a later milestone needs one is a streaming or chunked import, which is architecture beyond M9. The asymmetry worth naming for a reviewer: a campaign past that length can still be exported and would be refused on import with a 413.


S. Acceptance matrix

Every M9-owned item, with the evidence that decides it. Where a claim names a boundary, the row says which boundary was actually crossed.

I-series

Result Evidence, and the boundary
I01 Export campaign PASS Every section present and non-empty on the M9 fixture, checked family by family rather than by file size. test_i01_*; reproducible with tools/m9_portability_report
I02 Import exported campaign PASS Second process, second directory, database file that never existed, exporting process stopped. Transcript, state, canon, Save Points, knowledge, events and retained-tree size all match. test_m9_clean_import.py. Same round trip in-process in test_m9_portability.py, labelled there as the weaker of the two
I03 Branch / disposable history PASS Both futures present and readable; the abandoned branch still records the depth it was left at; the superseded take still at its coordinate. No trimmed-export option exists, so that clause does not arise
I04 Checkpoint export PASS Names, notes, coordinates; every one resolves; each restores to its own position and its own state; later history intact. A Save Point the file cannot satisfy is dropped, never retargeted
I05 Knowledge provenance PASS Content byte-identical by SHA-256, classification, enabled, visibility, always-include, filename, parser/chunking versions. Verified with the exporting machine stopped, and in the browser
I06 No API secrets PASS Searched against the file's text, with the inert api_key column written first so absence is evidence. Compressed snapshots decoded, not skipped. Checked in the unit suite, the two-process run, and the browser run
I07 Undone active head PASS All three clauses. The exact head across a machine boundary; a head-less file opens at its tip with no Redo; a head past the story is refused with a message naming where the branch ends

L-series

Result Evidence
L01 Atomic turn commit PASS (not regressed) M9 changed no turn-commit path; the inherited tests pass unchanged. M9 adds the import's own version, in two forms: a refused import writes zero rows across eleven tables, checked by counting; and a failure deep inside the write phase — past every planner check, with the branches, nodes, parentage, memories, head and Save Points already in the session — also leaves zero rows, which proves the transaction rather than the planner. The rollback is explicit, so the next import succeeds rather than inheriting a poisoned session
L02 State reconstruction PASS Measured after a round trip: copy and source walked back three turns and forward three turns in step, document compared at every position. Restoration stays a snapshot read, not a replay
L03 Save Point after restart PASS Three processes. Export from A, import into a clean B, restore a Save Point in B, kill B, start a third against the same file, ask again. Transcript, state and retained-tree size all match
L04 Derived data rebuilt PASS Passages, FTS rows and vectors destroyed and rebuilt on a copy; authoritative campaign identical either side; retrieval works again; vectors alone gone still leaves lexical working. Running it found finding 1

M9-specific round trips

Result
Undone-head round trip PASS (I07)
Branched campaign round trip PASS (I03)
Checkpoint round trip PASS (I04)
Knowledge provenance round trip PASS (I05)
Derived indexes can be rebuilt PASS (L04)
Historical prompt/context round trip PASS — survives the source being deleted, in the copy
State audit round trip PASS — events, proposals, and the manual correction still identifiable
Summary lineage round trip PASS — the abandoned line's summary is still ineligible, before and after a rebuild
Take parentage round trip PASS — the pager reads identically in copy and source
Memory authority round trip PASS — a heuristic memory is not promoted
Legacy round trips PASS — v1, v2, and four earlier seams; nothing invented at any of them
Migration from an M8-built database PASS — 12 families, 0 problems, version 92→92→92

Adjacent items re-verified rather than assumed

Result
E03 abandoned summary cannot leak PASS after a move, and after an index rebuild
F07 heuristic memory is not canon PASS — the authority survives the move
C04 manual correction PASS — still identifiable as one in the copy
H08 path traversal PASS — four shapes sanitised; the backup endpoints accept no path
H09 ZIP slip N/A — no archive format was introduced; the bundle is one JSON document
G04 disabled source stays out of retrieval PASS after the move

Synthetic substitutes, labelled

Where Substitute Why it is sound
Every backend suite ScriptedProvider narrator, stub embedder/summariser The property under test is the round trip, not the model. Both derived factories are stubbed, per M6's finding M6-F3
test_m9_clean_import.py the spawned server's deterministic narrator Real processes, real HTTP, real files; only the model text is canned
Legacy-bundle seams v3 exports aged backwards A checked-in fixture drifts and a hand-written one tests a shape nothing wrote
Browser run: the export's write to disk (§Q) the bytes the page produced, captured from URL.createObjectURL Headless Firefox does not complete a blob: download here. Proves the browser assembled and offered the right file; does not prove Firefox's download writer
Browser run: the import's read of the file (§Q) the POST the page would make, made against the same endpoint Firefox is snap-confined and cannot read a WebDriver-uploaded path — the File arrives with the right size and change fires, then FileReader raises NotFoundError. The click and handover are exercised
Not substituted the browser run's narrator and its two installations, the migration proof's M8 build, the backup's SQLite, every process boundary, the file-on-disk import in test_m9_clean_import.py Each is the boundary the claim names. The end-to-end file path the browser could not complete is proved there, with no browser in the way

T. Security and local-only

Internet none. No M9 code path opens a socket. Export, import and backup are file and database operations
Embedded URLs inert. A URL in a bundle is stored and displayed as text
Remote images unchanged; CSP still img-src 'self' data:
Shell none. No subprocess in any M9 application file
Imported content executed never. It is stored and rendered through M8's safe Markdown path
Loopback unchanged. The backup endpoints are on the same loopback API
Inference endpoint policy untouched. endpoints.py unchanged
Cloud storage / remote backup none. The backup writes to a directory beside the database

I06

Tested against the file's text, not against a list of columns — a field added to a model the exporter walks would otherwise reach the bundle with no column test noticing. The inert api_key column is written with a recognisable value first, so its absence is evidence rather than a tautology. Searched for and absent: the written secret, api_key, apiKey, the endpoint and its port, and absolute filesystem paths. The compressed snapshots are decoded before searching rather than skipped.

Checked in three places: the unit suite, the two-process run (against a real endpoint URL), and the browser run. Measured on the M9 fixture, over 90,074 bytes of bundle plus 154,279 bytes of decoded snapshot:

  'SECRET-MUST-NOT-TRAVEL'      present: False     (written into api_key first)
  'api_key' / 'apiKey'          present: False
  the endpoint host and port    present: False
  'endpoint_url'                present: False
  '/home/' , '/tmp/'            present: False
  'secret.key'                  present: False

What is present, and correctly: the per-turn model name inside each historical snapshot. That is a record of what the turn ran under, which §6.1 requires — not a configuration the import applies. §M has the distinction.

File, path and archive safety

No archive format was introduced, so H09 does not arise: the bundle is still one JSON document, and the compression in §R is an encoding of one field inside it, not a container.

  • No bundle-supplied value becomes a read or write path. originalFilename is metadata; four traversal shapes are sanitised on import and asserted absent.
  • Knowledge restores from the content in the bundle, never by reopening the source machine's path — verified with the exporting process stopped.
  • The backup endpoints accept no path at all.

U. Findings

Numbered, including those fixed before the final tree.

1. Deleting a campaign leaked its lexical index, and broke the next import — PRE-EXISTING (M7), FIXED

Severity: high. An ordinary upload in an unrelated campaign failed with a 500.

The FTS5 index is a virtual table, so no foreign key reaches it and no ON DELETE CASCADE covers it. Deleting a campaign cascaded knowledge_sources → knowledge_chunks and stopped there, leaving one index row per passage pointing at a chunk that no longer existed. Nothing read them — every search joins through knowledge_chunks — so the leak was invisible until SQLite handed the freed primary key out again, at which point the next source imported into any campaign collided on INSERT and raised.

Worse, clear_index finds index rows through the chunks, so with the chunks gone the orphans were unreachable: Reindex, the documented repair, could not repair it.

Found by running L04 — it failed on the rebuild with an integrity error, which is the only way this shows itself. Fixed at both ends: importer.clear_campaign_index removes a campaign's index rows before it is deleted, and fts.add uses INSERT OR REPLACE — the rowid is a chunk's primary key, so a row already there is by definition stale. The second half means a database already carrying the leak repairs itself, with no migration. Three regression tests.

2. An imported node with no state snapshot was given the head's state — PRE-EXISTING, FIXED

Severity: medium. M5's review finding 3, arriving through the import.

tree.stamp_outcome runs on every flush and fills a missing narrative_state_after from the campaign's current state. On an import that is the state at the exported head — so every position whose snapshot the file did not carry came back holding the newest position's state, and an Undo to turn 2 showed what the story knew at turn 20.

_write_nodes now writes the empty document explicitly when the file carries none. Legacy behaviour is unchanged: a pre-M5 bundle has no narrativeState either, so the fallback was already writing empty() for those files.

Severity: high, and it would have shipped looking correct.

_relink_snapshots edited the dict the attribute already held and assigned it back. context_snapshot is a plain CompressedJSON column, not a MutableDict, so SQLAlchemy tracks it by assignment: at flush the loaded value and the current value were the same object, the history reported no change, and no UPDATE was emitted. The import looked right in memory and wrote the untranslated ids to disk.

Found because the test asserted the outcome — that the restored records name this campaign's sources — rather than that the function was called. It now builds a new document.

4. Cancelling the import file dialog hung the Import button — PRE-EXISTING, FIXED

Severity: low, and user-visible. pickJSONFile resolved on change only, so closing the picker without choosing anything never settled the promise: the Campaigns screen's await never returned, its finally never ran, and the Import button stayed disabled reading "Importing…" until the page was reloaded.

Found while making the input reachable for the browser suite. The screen's own comment already said "a cancelled file picker is not a failure worth a message" — it had simply never received one. Now oncancel rejects with an empty message, which is exactly what that comment describes.

5. The file input was detached from the document — FIXED

Severity: none for a user; structural for testing. pickJSONFile created an input and clicked it without appending it, so no element existed for a test or for WebDriver to hand a path to — the import workflow could only ever be checked by calling the API underneath it, which is not the workflow. It is now in the document and removed however the promise settles. Six tests.

And a second lesson from fixing it. The first attempt marked the input hidden, which reads correctly and is wrong: a hidden element is non-interactable, and WebDriver will set files on one without dispatching change — the file lands and nothing happens, which is a worse failure than the detached input because it looks like it worked. It now uses the ordinary visually-hidden pattern — off-screen, zero-sized, aria-hidden, out of the tab order — so no reader meets a stray "Choose file" control while the browser's own dialog is what they are looking at. The test asserts hidden === false and says why, so the next person does not "tidy" it back.

6. Five existing tests pinned the gaps M9 closes — REWORKED, NOT DELETED

Each asserted the M8 behaviour that was the debt. Following the rule BUILD-MILESTONES.md records from M2 — move the instrumentation, do not delete the test — each now asserts the new guarantee, and each carries a note saying what it used to assert and why it changed.

Test Was Now
test_historical_prompt_evidence_survives_an_export_round_trip the copy's turn has no snapshot (404) it has the same snapshot, and its source_ids name this campaign
test_export_and_import_round_trips_variants the pager reads 1/1 on the copy, "a gap in the import" it reads 2/2, as the source does
test_an_unknown_format_is_refused used v3 as its "from the future" placeholder — which M9 made real, so it started importing the file it meant to reject uses v99, and asserts against bundle.FORMAT rather than a literal
test_export_carries_the_whole_story format == "ai-dnd-adventure-v2" format == bundle.FORMAT
test_the_shipped_file_is_a_bundle_this_build_can_import version == bundle.FORMAT version in bundle.READABLE — the shipped starter is a v2 file and was not regenerated, because rewriting a shipped asset to keep a test's equality holding is changing the evidence to fit the test

7. The scale tool measured the wrong thing after compression — FIXED IN THE TOOL

It looked for the plain contextSnapshot key and reported 0% once the export started writing contextSnapshotZ — i.e. it would have reported the evidence as omitted when it was merely encoded, which is precisely the mistake the portability report exists to avoid making about anything. Both tools now decode. Recorded because a reviewer reading an intermediate number in this report's history would otherwise be misled.


V. Planning changes

Distinguished by kind, as the brief asks.

Requirement corrections: none

M9 altered no product requirement. SPECIFICATION.md and SECURITY-THREAT-MODEL.md are unchanged: §16 already required the export to preserve the exact active position, §6.1 already required an exact prompt/context snapshot per turn, and M9 implements both rather than redefining either. No acceptance condition was weakened.

Newly settled design decisions

Document Decision
IMPORTED-KNOWLEDGE-DESIGN.md §73 Story Cards are compatibility-only legacy data and no longer enter the narrator's prompt. The brief asked for this to be decided; §K has the evidence that they were the untracked path §73 already forbade
DATA-MODEL.md §29, TECHNICAL-DESIGN.md §9.3 The bundle carries historical evidence, and the two-category rule becomes three
DATA-MODEL.md §29, TECHNICAL-DESIGN.md §9.3 The format is versioned rather than extended, and why §9.1's and §9.2's reasoning does not stretch to cover this addition

Implementation facts recorded

TECHNICAL-DESIGN.md §9.3 (the v3 contract, the encoding, the two-phase transaction) and new §9.4 (the backup); DATA-MODEL.md §29 (the format, the three categories, required vs optional, the pointers translated) and §31, which listed summaries as derived-and-rebuildable and now says what "where practical" excludes, and where stored prompts sit — a reader landing on §31 alone would otherwise conclude the export should be regenerating both; V1-ACCEPTANCE-TESTS.md I01-I07 and L02-L04 results — with I05's M7-era limit marked closed and the original paragraph kept, because the M9 decision is only legible against it; BUILD-MILESTONES.md M9 status; README.md and VERSION.md; DEVELOPMENT.md (a new section on the two recovery tools, taking a backup, and the stop-move-start restore, plus the imported-campaign context note).

M8's report moved to archive/milestone-reports/, by the rotation convention planning/README.md states.


W. Residual risks and the M10-M11 handoff

Residual risks

  1. A long campaign's bundle has a measured ceiling. ~279 turns against the 20 MB import limit; a longer campaign can still be exported and would be refused on import, which is the asymmetry worth naming. Far beyond M11's 100-turn target, which is 14% of the cap. M9's own additions account for 12% of that ceiling — the other 88% is the per-position state document v2 already carried, so lifting the ceiling means addressing that, not the evidence. Risk: low for v1; the fix is a streaming or chunked import, which is architecture beyond M9.
  2. quick_check rather than integrity_check on a backup. It does the structural work without the full index cross-check; a corrupt index that quick_check misses would ride into the copy. Risk: low, and the trade is stated in the code — a backup verified too slowly to be taken is worse.
  3. The backup is not offered on a schedule, and nothing prompts for one. Risk: low, and outside M9's scope. DEVELOPMENT.md shows the curl for cron.
  4. chunk_id in a restored snapshot is stale. It is a React key and a data attribute, not a live pointer, and chunk_index plus heading_path still identify the passage. Risk: very low; documented in the code.
  5. The importing machine's context window may differ from the source's. Documented rather than detected; detection is M11's.
  6. This machine cannot drive a file into or out of the browser. Firefox here is a snap, so its sandbox blocks reading a WebDriver-uploaded path, and its headless download manager does not complete a blob: download. §Q labels both substitutions and the end-to-end file path is proved without a browser in test_m9_clean_import.py. Risk: low for the product, real for future evidence — M11 will meet the same wall on any file-based browser claim, and the cheapest fix is an unconfined Firefox or Xvfb on the certifying machine.

Carried to M10

  • Media tables do not exist. §L records that as not applicable rather than inventing schema, and M10 owns the coordinator, the providers and the jobs.
  • The scene section of the state document is the extension point that exists today, and it round-trips.

Carried to M11

  • The context window, end to end: detection, the 100-turn certification, and whether a Settings warning that reads the real window belongs in v1.
  • The streaming import, if the bundle ceiling above is judged too low.
  • Contrast and focus measurement, tablet tuning, and the broader WCAG audit (inherited from M8's §U, untouched by M9).
  • An unconfined browser on the certifying machine, so that a file-based browser claim can be made end to end rather than in two labelled halves (residual 6).

Recorded here but not M9's: four hands-on playtest findings

A play session against accepted, signed M8 — after M8 was accepted and before M9 closeout — surfaced four product-quality observations. None is an M9 defect, none was caused by M9, and none blocks M9 acceptance. They are written up in full in §Y, and durably in BUILD-MILESTONES.md under M11 so they survive this report's archival:

Owner
A. The browser tab still reads AI D&D M11 release polish
B. After Undo, the reader cannot tell where they are M11 UX/release polish
C. The narration-length setting has no measurable effect M11 realistic-model behaviour
D. Character identity / coreference confusion (root cause unknown) M11 realistic-model / context diagnostic

D is the significant one, and its evidence is gone — the disposable playtest database was destroyed, so no root cause is claimed. §Y records the reproduction and classification plan, and one structural fact worth checking first: the narrative state permits two entities to share a display name and reports nothing.

Still open from earlier milestones, and not M9's

Cross-layer duplication (CONTEXT-AND-MEMORY.md §22); the discarded-history recovery screen (§63); whole-transcript copy and story search (§77, §78); the read-only RPG world state; the inert legacy tables and the dual-dialect migration code awaiting their cleanup migration.


X. Final milestone assessment

Is M9's Definition of Done satisfied?

BUILD-MILESTONES.md states it as: "A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact."

Yes. Measured three ways — in the suite, across two server processes with two data directories and the exporting process stopped, and in a real browser against a real narrator.

The brief's closing questions, answered directly

Question Answer
Can a campaign be moved into a clean data directory and reopened at the exact head? Yes. Into a database file that had never existed, from a second process, with the first stopped. The fixture's head is behind its own branch's tip and the abandoned line's, and is not the newest row written — so an importer guessing the tip, the newest row or the deepest row lands elsewhere. It opens where it was left, the retained future is there, and Redo reaches the same next turn in copy and source
Is all authoritative state intact? Yes. The document at the exported head matches exactly; every historical position matches, walked in step; restoration stays a snapshot read rather than a replay; and the audit that explains the state travels with it — including the manual correction, which is the one change no narration explains
Is historical prompt/context evidence intact? Yes, and this is the M8 handoff closed. An old turn in a restored campaign shows the prompt it was actually given and the passages it was shown, after the source has been deleted in the copy — asserted directly, not argued
Are Save Points intact? Yes. Names, notes and coordinates; each resolves; each restores to its own position and its own state through M3's head movement; later history survives. A Save Point the file cannot satisfy is dropped, never retargeted
Is knowledge intact without original filesystem paths? Yes. Content identical by SHA-256, classifications and lifecycle intact, searchable immediately — verified with the exporting machine's process dead and its directory holding a database the importer never opened
Can derived data be rebuilt? Yes, and running that test found a defect that had made Reindex — the documented repair — unable to repair the state it most needed to. Both ends are fixed, and a database already carrying the damage repairs itself
Can a consistent SQLite backup be created? Yes, through SQLite's online backup API, while the application is being written to, verified with quick_check before it is kept, and opened independently and read in every test. It works in the Docker image and lands on the mounted volume
Are legacy bundles still supported? Yes. v1, v2, and every seam from pre-active-head onward — with nothing invented at any of them. A pre-M9 file gets no manufactured audit trail; a pre-M3 file opens at its tip because that is the position it recorded
Are any required M9 conditions unverified? No. Every required condition has evidence at the boundary it names. What remains is stated as residual risk in §W with its measurement, not as an unverified claim
Is it safe to proceed to M10 after review? Yes, on this evidence. M9 changed no schema, altered no product requirement, and added no media surface. It leaves M10 the scene section of the state document as the extension point that exists today, and §L records the absence of media tables as not applicable rather than inventing schema to satisfy the word "metadata"

What a reviewer should look at hardest

Named deliberately, because a report that only presents its strengths is harder to review than one that points at its own seams:

  1. The version bump (§D). It is the one place M9 chose differently from two earlier milestones that faced a similar decision. The argument is that an absent key is unambiguous for a position and ambiguous for evidence; a reviewer who disagrees should say so, because it is a contract decision and not an implementation detail.
  2. The story-card decision (§K). It removes something from the narrator's prompt, which is a behaviour change in a portability milestone. The brief authorised it and §73 required it; the judgement to check is whether stopping at the prompt — and leaving the summariser's roster alone — is the right line.
  3. The bundle ceiling (§R). Real, measured, and stated rather than hidden. Worth checking that ~279 turns is acceptable for v1 given a 100-turn certification target, and that the export-succeeds/import-refuses asymmetry is tolerable until a streaming import exists.
  4. Finding 3. It would have shipped looking correct in memory and wrong on disk. It was caught only because the test asserted the outcome rather than the call, which is worth generalising.

Not done, and deliberately

No image, video, TTS, STT, media provider or job architecture; no M10 media coordinator; no 100-turn certification; no WCAG audit; no new context-window or provider architecture; no discarded-history recovery browser. Retained history survives correctly so a later screen can use it, which was M9's part of that.


Y. Post-M8 hands-on playtest findings

These are not M9 defects, were not caused by M9, and do not block M9 acceptance. They are recorded here because this is the report a reviewer is holding, and because the M9 report will be archived when M10's replaces it — BUILD-MILESTONES.md carries the durable copy under M11 for that reason.

They come from a real play session against accepted, signed M8 — a real browser, a real trusted-LAN Ollama, narrator qwen2.5:3b-instruct-16k, and a disposable isolated campaign database. That database was deliberately destroyed afterwards, so the stored context snapshot for the turn in finding D no longer exists. Everything below is therefore recorded as an observed symptom, and no root cause is claimed that cannot now be proven.

Nothing in this section changed any application code. Where the text states how the product behaves today, it is from reading the code and running it, and it is labelled as a mechanism rather than as a proven cause of what was seen.

A. The browser still calls the product "AI D&D"

Observed: the browser tab/title reads AI D&D.

Verified: frontend/index.html line 18 is <title>AI D&amp;D</title>. It is the inherited upstream title and M8 did not change it.

Not a false claim by any accepted document. M8's report never asserted the title was changed — its terminology audit covered branch, fork, node, head and depth, and its "does it feel like the intended storyteller" argument lists the navigation, the inspector, the glyph buttons and the removed screens. So this is an uncovered gap, not documentation that needs correcting. No documentation was corrected, and no code was changed.

The naming question is genuinely open, and should not be closed by a find-and-replace. The planning package and README.md call the product Adventure Storyteller, but SPECIFICATION.md is explicit that the engine must stay genre-agnostic — science fiction, mystery, horror, historical, westerns — and Adventure is narrower than the product it names. A repository-wide rename was deliberately not performed in this pass. For planning purposes a neutral working name such as Interactive Story is used, and a browser-tab form such as <Campaign Name> — Interactive Story is a candidate, not a decision.

Owner: M11 release polish. No earlier milestone touches the shell metadata. It is a small, self-contained change whose only hard part is the naming decision, which is the repository owner's.

B. Undo/Redo loses the reader's orientation

Observed: Undo worked, and it was hard to tell which point in the story the reader had moved to.

Not a correctness problem. M3's active-head semantics behaved correctly and M9 re-verified them at every boundary (§E, §S). This is presentation.

What the spec said, and why that was not enough. BROWSER-UX-SPEC.md §8 — "The current endpoint should be clear." — is five words, and the second sentence beside it ("the input box always continues from the currently active story head") was already true while the reader was lost. A requirement that a working implementation satisfies while a real user cannot answer the question is too vague to hold the behaviour. §8 has been strengthened in this pass; no UI text is prescribed, because none is ratified.

The requirement, stated without wording: after Undo, Redo, a Save Point restore, an edit to an earlier turn, or any other movement of the active position, a reader should be able to tell where they now are in the visible story without needing implementation terminology — branch, head, node and depth remain forbidden at the surface (§38 of the spec, and M8's audit). A lightweight indicator such as Moment 8 → Moment 7, optionally noting that later story is still available, is a candidate. Exact wording deliberately not settled here.

Owner: M11 UX/release polish, with a browser regression scenario.

C. Narration length did not feel effective

Observed: setup offered roughly one paragraph / 2-4 paragraphs / longer exposition; the reader chose 2-4 paragraphs and felt replies were substantially longer than that.

Deliberately not recorded as "the model ignored instructions." The evidence does not establish that, and there is a mechanism in the code worth checking first.

The mechanism, verified by reading and running the code. There are two independent length controls, and they do not reference each other:

  1. The setup choice becomes one English sentence appended to the campaign's ai_instructions — "Keep responses to roughly two to four paragraphs." (frontend/src/pages/NewCampaign.jsx, LENGTH_SENTENCE). It changes no generation setting. It sits in the cached system block, near the top of the prompt.
  2. context/builder.py's length_hint() derives a numeric word range from Settings.max_output_tokens — a global application setting, not the campaign's choice — and emits it as a [Hard limit: …] line placed after the history, near the end of the prompt.

At the default max_output_tokens = 800, that second line reads, measured:

[Hard limit: this turn must not exceed 506 words, and it should not stop short
 of about 177. Prefer the lower end of that range unless the scene genuinely
 needs more. Finish the narration and append the state block well inside the
 limit.]

and it is byte-identical whether the reader chose brief, medium or long. A reader asking for two to four paragraphs is simultaneously told, in the more recent and more numerically explicit of the two instructions, not to stop short of about 177 words and that up to 506 are permitted.

This is a mechanism, not a proven cause. Whether it produced what this reader saw needs the reproduction below; the narrator's own instruction following is also in play, and both could contribute.

What a reproduction must measure, rather than judge by eye:

  1. exactly what length instruction(s) enter the stored prompt — both of the above, with their positions;
  2. whether the setup choice changes max_output_tokens or any generation setting (today: it does not);
  3. actual words, tokens and paragraph counts across repeated realistic turns per setting, so a directional effect can be shown or disproved;
  4. at least the reference 3B narrator and a stronger local narrator, since instruction-following differs.

A design candidate, explicitly not ratified: clearer approximate targets — Brief ~100-200 words, Standard ~200-400, Detailed ~400-700 — with the setting actually moving the numeric budget. Do not hard-truncate prose: the state block is emitted last and truncation removes it, which is the failure length_hint's own comments exist to avoid.

Owner: M11 realistic-model behaviour validation, with a realistic-model test rather than only a scripted-provider one.

D. Character identity / coreference confusion — the most important one

Observed: a story established four people in an office — Bill (protagonist), Roger, John, Alice. Later narration treated Alice as though there were two different Alices, in language equivalent to "Alice wondered what Alice was doing."

Root cause: UNKNOWN, and it cannot now be established. The disposable playtest database was deliberately destroyed, so the stored context snapshot for that turn is gone. This report claims neither an inference-model defect nor an application defect. The candidate classes are:

Model failure the stored prompt correctly identifies one Alice and the 3B model still makes a coreference error
State failure the narrative state holds duplicate or conflicting Alice records
Context/derived failure state is correct, but assembled context, a summary, a retrieved memory or the history rendering presents two identities
Combined weakness the context is not contradictory but is insufficiently explicit for a small model, producing an avoidable failure

One structural fact a reproduction should check first, verified by reading the code and stated as a fact about the implementation rather than as a cause:

The narrative state permits two distinct entities to share one display name, and nothing reports it. Entities are keyed by the id the model supplies (state["entities"][event["entity"]], narrative/apply.py), with name a separate display field. validate.py's DUPLICATE_ENTITY rejects re-creating an entity with the same key — key in known — and there is no check anywhere in the narrative layer on the display name. So create_entity(entity="alice", …, name="Alice") followed by create_entity(entity="alice_2", …, name="Alice") both succeed, and the state then holds two entities that both render as "Alice".

That is exactly one of the failure modes this finding describes. It does not establish that it happened here — no evidence survives — and a reproduction may well land on one of the other three classes instead.

Owner: M11 realistic-model / context diagnostic. The test plan is below and in BUILD-MILESTONES.md under M11.

The M11 character-identity diagnostic, as a test plan

Deterministic setup with at least a protagonist and three same-scene supporting characters — Bill (protagonist), Alice, Roger, John — with unambiguous identities and roles established up front. Then a multi-character interaction over enough turns to stress: pronouns; dialogue attribution; people entering and leaving; reference by name; reference by role; and one character speaking about another.

Detect and report, at minimum: duplicate character creation; same-name entity duplication; protagonist identity drift; dialogue attributed to the wrong person; a character referring to themself as a separate same-named character; and state/context disagreement about identity.

On any failure, preserve and report all of: the authoritative state immediately before generation; the exact stored context/prompt snapshot; the recent-history section; summaries; retrieved memories; imported knowledge if any; the narrator output; and the model identifier and settings. M9 makes all of that portable (§H), so a failing campaign can now be exported and handed to whoever investigates it — which is the practical reason this diagnostic is newly worth writing.

Then classify from the evidence:

STATE DEFECT
CONTEXT ASSEMBLY DEFECT
DERIVED MEMORY/SUMMARY DEFECT
MODEL FAILURE WITH CORRECT CONTEXT
AMBIGUOUS / MULTIPLE CONTRIBUTORS

Two rules for whoever runs it: do not "fix" a model failure by changing authoritative story state, and do not blame the model if the prompt already contained the identity error. That distinction should become part of M11's realistic-model review methodology rather than a one-off judgement.

The standard fixture does not cover this failure class

Reviewed, as asked. TEST-CAMPAIGN-FIXTURE.md stresses knowledge boundaries, secrets, authority precedence, branch leakage and possession — its seven deliberate traps are all of those — and it contains no identity trap at all; the word coreference does not appear in it. Its on-stage cast is effectively two people, Aldric and Mara, with Edrin established as missing rather than present.

So the user's suspicion is correct: same-scene multi-character identity continuity is not exercised by the standard fixture.

The established fixture was deliberately not modified. It is the deterministic baseline several milestones' results are compared against, and changing it would invalidate those comparisons. A companion fixture is proposed instead, in a new appendix to that document, named Multi-Character Identity Test and explicitly additive.


Written at the end of implementation, before review. The tree is staged; §32 of the brief governs what happens next.