Files
interactive-story/planning/reports/M9-IMPLEMENTATION-REPORT.md
T
JesseMarkowitzandClaude Opus 5 44edece67e M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 01:55:45 -04:00

1623 lines
88 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M9 — Export, Backup, Recovery, and Migration Hardening
**Implementation report, written for an independent reviewer.**
This is a set of claims with the measurements attached. It is not a record of
acceptance: M9 is implemented and verified, and acceptance comes after review.
Two conventions carried from M8, because they are what make a report checkable:
- **Where a claim names a boundary, the boundary was crossed.** "A clean data
directory" is a second process against a database file that has never existed.
"A restart" is a different PID. "A backup" is opened as its own database and
read. "Migration from M8" is a database built by a server running the signed M8
commit. Where a substitute was used, it is labelled.
- **Every defect found is numbered in §U, including the ones fixed before the
final tree.** Seven are recorded. Three were found by *running* M9's own tests
rather than by reading them, two more by building the browser harness, and
**four of the seven predate M9** — they were reachable in M8 and nothing had
gone looking.
---
## A. Executive result
**PASS.**
Every required M9 condition is verified, by the boundary it names. The Definition
of Done — *a campaign can be safely exported, imported into a clean data
directory, and reopened at the exact intended active position with authoritative
history/state intact* — is satisfied and is measured three separate ways: in the
suite, across two real server processes with two data directories, and in a real
browser against a real narrator.
| | |
| --- | --- |
| Bundle format | **`ai-dnd-adventure-v3`.** v1 and v2 still import; every seam from pre-active-head onward is tested |
| Historical prompt provenance | **carried.** The M8 handoff is closed |
| Story cards | **compatibility-only**, and out of the narrator's prompt |
| Summaries / memories | **carried**, with the coordinates that decide eligibility |
| Derived indexes | **not carried, and rebuilt** — with the rebuild path fixed |
| SQLite backup | **SQLite online backup API**, verified before it is kept |
| Schema change | **none**, proved against a database M8's own code wrote |
| Backend | **1,102 passed, 14 skipped, 0 failed** in 14:54 (M8 baseline: 950) |
| Frontend | 145 passed (was 132), lint and production build clean |
| Browser | **36/36 checks**, one frozen build, real narrator, two installations (§Q) |
| Docker | production image builds clean; the backup works inside it, on the mounted volume |
Nothing is claimed here that was only compiled. §S maps every acceptance item to
the evidence that decides it, and names what is asserted from a synthetic
substitute.
---
## B. Repository and provenance state
| | |
| --- | --- |
| Branch | `m9-recovery` |
| Base | `1ce99727606124b0f8b3dc3b88842410c874ee14` — *M8: the browser becomes the storyteller* |
| Base signature | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569` |
| Working tree at start | clean |
| `HEAD` at writing | still the M8 base — M9 is staged, not committed (§32) |
| Upstream ancestry | `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` **is** an ancestor |
| `LICENSE` | unchanged; `git diff 1ce9972 -- LICENSE` is 0 lines |
| Schema version | **92 before, 92 after** |
| Migrations | **none added.** `git diff 1ce9972 -- backend/app/migrations.py` is 0 lines |
| Columns | **none added or removed.** 0 `mapped_column` changes in `models.py` |
### Change inventory
```
backend/app/bundle.py 928 +- the v3 format, both directions
backend/app/backup.py 278 NEW — the online backup
backend/app/routers/backups.py 75 NEW — two endpoints
backend/app/routers/adventures/bundle_io.py 59 +- the two-phase import
backend/app/context/builder.py 63 +- story cards out of the prompt
backend/app/knowledge/fts.py 44 +- the index leak (finding 1)
backend/app/knowledge/importer.py 22 + clear_campaign_index
backend/app/schemas.py 19 + ImportedAdventureOut
backend/app/routers/adventures/crud.py 8 + clear the index before delete
backend/app/main.py 6 +- register the router
backend/app/starter.py 4 +- materialize returns a pair
frontend/src/pages/Settings.jsx + the backup control
frontend/src/components.jsx +- the file picker (finding 4)
frontend/src/api.js, styles/library.css +
backend/tests/m9_fixture.py 406 NEW — the portability fixture
backend/tests/test_m9_portability.py 1,186 NEW
backend/tests/test_m9_corrupt_bundles.py 689 NEW
backend/tests/test_m9_clean_import.py 460 NEW — two processes
backend/tests/test_m9_backup.py 451 NEW
backend/tests/test_m9_legacy_bundles.py 365 NEW
backend/tools/m9_portability_report.py 354 NEW — the baseline tool
backend/tools/m9_migration_proof.py 331 NEW
backend/tools/m9_scale_report.py 221 NEW
frontend/src/pages/backup.test.jsx 116 NEW
frontend/src/filePicker.test.jsx 85 NEW
```
Five existing test files changed. Each had pinned a gap M9 closes, and each was
**reworked to assert the new guarantee rather than deleted** — the rule
`BUILD-MILESTONES.md` records from M2 about moving instrumentation rather than
removing it. §U-6 lists them.
---
## C. The M8 portability baseline, measured
The brief required the baseline to be **measured, not assumed**. It was, with a
tool a reviewer can rerun on either commit:
```bash
cd backend && python -m tools.m9_portability_report # or --json
```
It builds the M9 fixture (§E) in a throwaway database, exports it, imports the
result, and classifies every data family. Run on the M8 commit and on the M9
tree, the difference is the milestone's claim in reproducible form.
### The fixture it measures
Deliberately not the standard Continuity Test: that one is shaped to read like a
story, and this one is shaped to break a round trip. 22 actions, 2 branches, 2
Save Points on different branches, 1 superseded take, 12 state events, 12
proposals, 11 stored prompts, 2 memories, 2 summaries (one on the abandoned
line, one at the head), 5 imported sources across all three classes including
one disabled and one narrator-only, and a manual state correction.
The head ends **behind the retained tip of its own branch and behind the
abandoned line's**, and the last thing the fixture does is an Undo — so the head
is not the newest row written, not the deepest row, not the tip, and not on the
branch holding the most story. An importer guessing any one of those lands
somewhere else.
### M8 — `ai-dnd-adventure-v2`, 29,630 bytes
| Verdict | Family |
| --- | --- |
| PRESERVED | campaign identity, transcript, branches, branch disposition, active head, alternate takes, Save Points, narrative state (current), narrative state (per position), imported knowledge, memories, scene metadata, story cards |
| **OMITTED** | **take grouping** — SP9 parentage, so imported nodes landed parentless |
| **OMITTED** | **state events** — the audit half of §17's hybrid |
| **OMITTED** | **state proposals** — what the model asked for, and what was refused |
| **OMITTED** | **manual corrections** — indistinguishable from what the story established |
| **OMITTED** | **historical prompt/context** — the M8 handoff |
| **OMITTED** | **retrieval provenance** — which passages an old turn was shown |
| **OMITTED** | **per-turn model settings** — what a historical turn ran under |
| **OMITTED** | **knowledge parser versions** |
| **OMITTED** | **summaries** — only the lineage-less mirror column travelled |
| **OMITTED** | **memory authority** — a heuristic memory imported as accepted story |
Two of those were visible from outside, through the API a reader reads: after a
round trip the copy's **state events** and **summaries** differed from the
source's. The rest were invisible until something needed them.
Derived and correctly not carried: knowledge passages, the FTS index, knowledge
and memory embeddings, the branch lineage cache, derived status.
### M9 — `ai-dnd-adventure-v3`, 90,117 bytes on the same fixture
**Every family PRESERVED. No disagreement on any reader-visible family.**
The file is 3.0× larger on this fixture; §R has the measurement on a long
campaign, where the ratio is very different and the conclusion is not the
obvious one.
---
## D. The final bundle contract
Documented in `DATA-MODEL.md` §29 and `TECHNICAL-DESIGN.md` §9.3; the reasoning
is repeated in `bundle.py`'s own docstring, which is where the next maintainer
will be standing.
### Why the version was bumped
Everything M9 adds *could* have been an optional key read with `.get`, the way
`persona`, `checkpoints`, `narrativeState` and `knowledge` each were — §9.1 and
§9.2 each considered a bump and correctly declined one.
That mechanism stops working here, and the reason is the rule the format already
lives by. A v2 file with no prompt provenance is **ambiguous**: written before
M9, when no file could carry one, or by M9 from a campaign whose turns predate
the column? Those are different facts and a reader has to be able to tell them
apart — the same distinction I07 draws when it says a file written before the
head was carried opens at the tip *because tip was the only position that format
could represent*. **A version number is how a recovery file states what it was
capable of recording.**
The family name is unchanged, deliberately: renaming it would break every reader
for no gain, and `PROVENANCE.md` is where the fork is recorded.
```text
v1 a linear story, its turns, and its retries as a repeating group
v2 + the tree, the live flags, the after-snapshots, the chosen head,
Save Points, the narrative state document, imported knowledge
v3 + state events and proposals, historical prompt/context provenance,
lineage-anchored summaries, take parentage, memory authority
```
The importer reads all three. A version it has never heard of is **refused with
its own name in the message**, rather than read as the newest it knows.
### Three categories, not two
§31 of `DATA-MODEL.md` distinguishes authoritative from derived, which was
sufficient until M9 had to decide about stored prompts. They *are* derived — a
machine assembled them — and they must travel anyway:
```text
chosen the story, the head, the takes, the Save Points, the
classifications, the canon travels
evidence the state events and proposals, the per-turn prompt and the
passages it was shown, the model and generation settings
that turn ran under travels
rebuildable knowledge passages, the FTS index, embeddings, the branch
lineage cache rebuilt on import
```
The test separating the last two is **not** "could this be recomputed" but
"would a recomputation answer the same question". A rebuilt FTS index answers
the same question. A rebuilt prompt does not — it says what the turn *would be
told now*, from today's canon, today's sources and today's state, which is the
opposite of what the inspector is for. **Historical evidence is not a cache**,
and M9 regenerates no prompt at any point in the import.
### Required, optional, and what an absent key means
| Key | Required | Absent means |
| --- | --- | --- |
| `format` | **yes** | refused |
| `actions` | in practice | an empty campaign |
| `branches`, `headBranch` | no | the root branch |
| `headDepth` | no | **the format could not say — open at the tip** (I07) |
| `checkpoints` | no | the campaign had none, or predates M4 |
| `narrativeState`, `narrativeStateAfter` | no | predates M5; empty document |
| `knowledge` | no | predates M7; empty library |
| `stateEvents`, `stateProposals` | no | **predates v3 — no audit trail is invented** |
| `contextSnapshotZ` / `contextSnapshot` | no | predates v3, or the turn has none |
| `summaries` | no | predates v3 |
| `id` / `parentId` on a node | no | pre-SP9 rule: `attempts.group` falls back to the coordinate |
| `authority` on a memory | no | the column default |
### Integrity and validation
One derived value is in the file, and only because its purpose is to be checked:
a knowledge source's `contentHash`. The import **recomputes** it from what
arrived, stores the recomputed value, and records the discrepancy on the source
where a reader can find it. The stated hash is never trusted and never silently
discarded.
### The two pointers that are translated
Branch numbers already were. M9 adds `source_id` inside a restored retrieval
record: it names a row on the machine that wrote the file, so left alone it
points the inspector's "open this source" at whatever holds that id here. It is
repointed at the source that landed, or set to `null` where the file carries no
such source — at which point the record still holds the passage's title,
filename and text. **The evidence is never rewritten; only the pointer is.**
`chunk_id` is deliberately untouched: passages are rebuilt and get new ids, so
no translation exists, and `chunk_index` and `heading_path` still say which
passage it was.
---
## E. Story graph, head, and takes
**Round-tripped and compared through the API a reader reads**, never by row
count. `test_m9_portability.py`, and again across two processes in
`test_m9_clean_import.py`.
| Property | Evidence |
| --- | --- |
| Every accepted action, live and superseded | the whole exported tree is the same size either side |
| Branch and fork relationships | both branches, the fork depth, and the names |
| Active branch and head | §I07 below |
| Retained tip | the future past the head is in the database and Redo reaches it |
| Branch disposition | the superseded branch still carries the depth it was left at (6) |
| Retries / alternate takes | the superseded take is still at its coordinate; the selected one still selected |
| Which take is active | exactly one live node per coordinate, decided by the writer, not read from the file |
| Timestamps | `createdAt` restored per node, per branch, per Save Point, per event |
### The undo-head case (I07)
Exported with `active head < retained tip`, imported into a genuinely clean data
directory, and the copy opens **at exactly the exported head**. The later turns
are present as retained future, Redo is offered rather than the story having
silently been redone, and Redo then walks to the same next turn in both
campaigns.
### The diverged case (I03)
Both futures survive and stay distinguishable — the abandoned line's turns are
readable on their branch, and the branch that was left still records that it was
left, at the depth it was left at. There is no trimmed-export option, so I03's
clause about one does not arise.
### The retry case, and a fidelity gap M9 closed
v2 exported no take parentage, so every imported node landed parentless and
`attempts.group` fell back to the coordinate. That is right for a plain retry
and **wrong as soon as two takes of one turn each have takes of their own
beneath them**: those share a branch and a depth, so the copy read `5/5` where
the source read `2/2` and `3/3`. v3 carries the parentage, and the pager now
reads identically in the copy and the original.
`test_retry_variants.py` had recorded the old behaviour as "a gap in the import";
that test now asserts `2/2` (§U-6).
---
## F. Save Points
| | |
| --- | --- |
| Name, note, coordinate, `createdAt` | all restored; compared through the panel's own API |
| Resolution | every restored Save Point reports `resolved: true` |
| Restore | reaches the position it names, through **M3's head movement** — no second restore architecture |
| Later history | intact after restoring; the retained tree is the same size |
| Two Save Points, two branches | restore to **different** states, so they are not accidentally the same pointer |
| The head | unmoved by importing them: a campaign exported at turn 30 with a Save Point at turn 12 opens at turn 30 |
### A Save Point the file cannot satisfy
**Dropped, with the rest of the campaign kept — and never retargeted.** The
decision and its reasoning are in `bundle.py` and `TECHNICAL-DESIGN.md` §9.2:
- not a refusal, because a bookmark costs a bookmark, where refusing the
campaign would lose the story to save the bookmark;
- not a repair, because the reader named a position, and if that position is not
in the file then **no other position is the one they named**.
Three cases are covered: a coordinate beyond the retained story, a branch the
file does not list, and a name that is blank once trimmed. In each, the *other*
Save Point in the fixture survives — asserted, because "dropped" must not mean
"dropped them all".
---
## G. Narrative state and the audit trail
### What travels, and why each
The audit inventory was made from the contract rather than by serialising every
table:
| Record | Travels | Why |
| --- | --- | --- |
| `narrative_state` (campaign) | yes | the authoritative document at the exported head |
| `narrative_state_after` (per node) | yes | **the restore path.** Without it, head movement becomes replay |
| `state_events` | **yes (new)** | §17's audit half: what changed, where, who asserted it, what it was before |
| `state_proposals` | **yes (new)** | the inspector reads both — the event says what was accepted, the proposal says what was asked for and refused |
| rejected/unparseable proposals | yes | the record that explains why the state does not say what the narration seems to say |
| `derived_status` | no | describes the last run of a background pass, not the story |
Exporting only the accepted half would have kept the answers and lost every
question, which is why both tables travel.
### Verified after the move
- Current authoritative state at the exported head: **identical**, compared as
the document.
- Historical positions: the copy and the source are walked back three turns and
forward three turns **in step**, and the document is compared at every
position. They agree throughout (L02).
- Restoration stays **O(1)-style**: every position's snapshot travels, so Undo,
Redo and Save Point restore remain one row read rather than a replay
(`TECHNICAL-DESIGN.md` §10.4, and the M4 note that made it load-bearing).
- **A manual correction is still identifiable as one.** This is the case the
omission hurt most: it is the one state change no narration explains, and with
the events gone nothing distinguished it from something the story established.
- The two tables arrive **linked**: events point at proposals that are in this
campaign, and proposals point at this campaign's own turns.
### Refuse, repair, drop — decided per field
| | Rule |
| --- | --- |
| **Refuse** | an audit record naming a turn the file does not contain. Unlike a Save Point, an audit record that quietly did not arrive leaves a campaign whose state cannot be explained — and the explanation is what a reader goes looking for exactly when something looks wrong |
| **Accept** | a proposal with *no* action: that is a manual correction, which has a coordinate and no narration behind it |
| **Keep the coordinate, drop the pointer** | an event whose *proposal* is missing. That is what the schema's `ON DELETE SET NULL` already says happens |
| **Normalise** | a malformed state document — M5's rule, unchanged: the story is the valuable thing, and a malformed section should cost the section |
---
## H. Historical prompt and retrieval provenance
**The M8 handoff (question B) is closed: the evidence travels.**
`SPECIFICATION.md` §6.1 requires every accepted turn to record an *exact
prompt/context snapshot or reproducible equivalent*. The product already stored
one; the bundle did not carry it, so that guarantee was local to the machine
that played the campaign. It is now portable.
Each node carries `contextSnapshotZ`: the assembled prompt section by section,
the passages retrieved **with the text each supplied**, which summary was
eligible, the token accounting, and the model and generation settings the call
ran under. Restored **verbatim** — nothing is regenerated, and nothing is
validated beyond its being an object, because imposing today's expectations on a
record an older build wrote would be the retroactive reading the inspector
exists to rule out.
### Verified to survive each thing that could destroy it
| | |
| --- | --- |
| A source disabled later | the record is unchanged |
| **A source deleted after import** | asserted directly: the sources an old turn was shown are deleted in the copy, and the turn still shows the same passages and the same text. The record holds the text, not a pointer to a row that can go away (`IMPORTED-KNOWLEDGE-DESIGN.md` §49-50) |
| A different chunking on the importing machine | the record is not rebuilt from chunks, so it cannot move |
| Current state moved on | the snapshot is a record, not a re-render |
| Canon edited later | same |
| Summaries rebuilt | same |
And across a real machine boundary: on machine B, which never assembled that
prompt and could not reassemble it, Inspect Context on an old narrator turn
returns a prompt with its sections intact.
### Retrieval provenance
Each `used` record keeps source title, filename, classification, visibility,
chunk index, heading path, the scores each path gave it, how it was admitted,
and the rendered text. `source_id` is translated (§D); everything else is
evidence and is untouched.
### Cost control
The prompt is stored **once per turn on the live attempt** (`attempts.py`), so a
superseded take carries only its own slices — its reply, its state proposal, its
token accounting. Asserted: no superseded take in the bundle carries `sections`
or `prompt`. A campaign retried twenty times does not carry twenty prompts.
---
## I. Imported knowledge
| | |
| --- | --- |
| Preserved | title, original filename (metadata only), content, classification, enabled, visibility, always-include, notes, media type, import timestamp, content hash, **parser and chunking version**, and a `sourceId` that exists only so historical retrieval records can be relinked |
| Rebuilt | passages, FTS rows, vectors — before the import returns for the first two, so the campaign is searchable immediately |
| Needs the original path | **no.** Verified with the exporting machine's process stopped |
| Disabled source | still disabled, and still out of retrieval |
| Narrator-only source | still narrator-only |
| Canon with always-include | still both |
| Content | byte-identical, by SHA-256, checked in the browser run too |
### L04 — the rebuild, and the defect running it found
On a **safe copy** (the imported campaign, not the original), every physically
derived structure is destroyed — passages, FTS rows, vectors — and rebuilt from
the content the bundle carried. Afterwards: every source `ready` with passages,
retrieval working, and the transcript, authoritative state, classifications and
lifecycle flags identical either side. Deleting the vectors alone leaves lexical
retrieval working, which is M7's rule that lexical is a production path and not
a fallback. A rebuild does not make an abandoned line's summary eligible.
**Running it found finding 1** (§U), a pre-existing defect that made Reindex —
the documented repair — unable to repair the state it most needed to.
### The M7 calibration rule
Untouched. `nomic-embed-text` keeps its measured calibration and an uncalibrated
embedding model still degrades to lexical-only rather than borrowing the 0.58
threshold; no M9 code path reads or writes the threshold, and the M7 calibration
tests pass unchanged.
---
## J. Summaries, memories, and derived data
### Summaries — carried, with their lineage
v2 carried only `adventures.story_summary`, the **mirror column with no lineage
of its own**, so a restored campaign resumed with no usable long-story
continuity. v3 carries the rows: text, coordinate, source range, trigger, model,
timestamp.
Eligibility is not a stored flag — it is whether the coordinate lies on the
active capped lineage — so restoring the coordinate restores the answer,
**including the "no"**. The fixture has one generated summary on the line it
abandons and one typed at the head it keeps, and after the move:
- exactly one is eligible, and it is the one at the head;
- the abandoned line's text does not appear in the prompt the copy would send
next;
- the same holds after a full index rebuild, so **M6's leak does not return by
the rebuild path either**;
- `story_summary` is re-derived from the lineage after import rather than
trusted from the file, because a mirror disagreeing with the lineage is
exactly M6's finding M6-F1.
### Memories — and an authority that was being promoted
Text, pinned, forgotten, source range, use count, coordinate, and now
**`authority`**. Without it every imported memory landed on the column default,
which silently promoted a `heuristic` memory to `accepted_story` — the one
direction F07 forbids. An unreadable value is read as `heuristic`, which is the
safe direction: a demoted memory costs a little ranking weight, a promoted one
puts a guess in front of the narrator as fact.
### Vectors
Not carried, in either subsystem, and rebuilt against whatever embedding model
*this* machine has. Exporting them would tie a campaign to one machine's model,
which is the reason M7 gave and M9 keeps.
---
## K. Story cards — the compatibility decision
**Settled: compatibility-only legacy data, and no longer in the narrator's
prompt.** Recorded in `IMPORTED-KNOWLEDGE-DESIGN.md` §73.
### What was found
Story cards **were** the "alternate untracked path around the new knowledge
authority/provenance rules" that §73 already forbade, and not in principle. A
keyword-matched card was injected into the narrator's prompt as
`World Lore: <entry>`, taking up to 40% of what remained after the imported
knowledge had been placed, with:
- no class, so nothing framed how far the narrator could rely on it;
- no visibility, so no narrator-only distinction existed;
- no source, no hash, no lifecycle — nothing to disable;
- **no browser surface at all since M8 removed the editor**;
- and no row in the context inspector, which renders `knowledge` and never
rendered `cards`.
It also competed with imported Canon for one budget, which is the arrangement M7
spent a milestone separating.
### What happens now
| | |
| --- | --- |
| **New export** | carries them, unchanged, under `storyCards` |
| **Legacy import** | accepted unchanged, from every format version |
| **Normal narration** | **not reached.** The `world_lore` section is gone |
| **Re-export** | carries them again — a round trip destroys nothing |
Nothing is deleted: the rows stay, the `/api/story-cards` endpoints stay, and
`memorybank.cast_brief` still reads them as the **summariser's character
roster** — that names who is on stage so a memory says "Aldric" rather than
"he", never reaches the narrator, and every memory written from it is
authority-classified by the application afterwards. The `cards` key stays in the
context report and is now always empty for a new turn, because M9 has just made
historical snapshots portable and an old turn's record must go on saying that
story cards were included. That is asserted.
This is the **smallest safe change**: it closes the one path that asserted
campaign facts to the narrator without any of §73's controls, and touches
nothing else.
---
## L. Scene and media metadata
**No media tables exist in this build, and none were invented.** Two active
documents say that is the right answer rather than a gap:
- **K04's second branch** — "if media tables are deferred, architecture/types
should demonstrate equivalent extension point".
- **`MEDIA-EXTENSION-CONTRACT.md` §58**, on export, in as many words: *"For v1,
if media is not implemented, preserve schema compatibility."*
M9 preserves it in the strongest available form: the format is now **versioned**,
so M10 can add media sections as v4 and every v3 reader keeps working, and every
v3 file keeps importing. Inventing an empty `media_assets` table to satisfy the
word "metadata" would have been schema M10 then had to live with.
What does exist is the `scene` section of the authoritative narrative state
document, which is `SPECIFICATION.md` §14's scene snapshot as this build
represents it. It is campaign-scoped and per-position, and it round-trips with
the document: asserted both on the campaign-level state and on the per-node
snapshots.
Nothing in M9 adds image, video, TTS, STT, a media provider, or a job
architecture.
---
## M. Model and context-window portability
The M8 handoff's question D, answered by keeping two things apart.
### What travels
**Per-turn model and generation settings**, inside the historical snapshot —
model name, API mode, temperature, max output tokens. There they are a *record
of what happened*, which is exactly §6.1's requirement.
**Campaign-owned configuration**: title, narrator instructions, persona, canon,
memory/summary switches, the campaign's own plot fields.
### What does not, and why that is the point
The `settings` row — endpoint, model, context budget, timeouts, embedding model.
Those describe the **machine**, not the campaign. Asserted directly on the
importing machine: after importing a campaign played against a configured
endpoint, machine B's `endpoint_url` is still its own default and its
`context_token_budget` is still 16,384. **Importing a campaign is not a way to
reconfigure the destination's inference.**
### Import does not depend on a model
Asserted: machine B has **no model configured at all**, and the campaign still
imports, opens, and shows its state and transcript. Playing on would fail;
recovering does not. *Campaign restored successfully* is not *the destination
narrator has the same capacity*, and M9 does not treat it as such.
### The ceiling, and what M9 does about it
No assumption that 16,384 tokens are available is hard-coded anywhere M9
touched, and no native context-window detection architecture was built — that is
M11's, with the 100-turn certification.
What M9 adds is the sentence a deployer needs, in `DEVELOPMENT.md`: a long
imported campaign fills a prompt on its **first** turn, so a machine that has
applied neither the derived-model procedure nor a matching budget meets its
ceiling immediately rather than gradually. Check `/api/ps` on the destination
*before* playing on an imported campaign.
---
## N. SQLite backup
`backend/app/backup.py`, `backend/app/routers/backups.py`, and a control in
Settings → Advanced.
### Why not a file copy
A copy taken while the application runs is a copy of a moving target: SQLite
writes in pages, and a plain copy can read page 5 before a transaction and page
900 after it, producing a file that opens, reports a schema, and is quietly
missing rows. Nothing warns anyone, and it is found later.
So it uses **SQLite's online backup API** (`sqlite3.Connection.backup`), which
copies under the right locks and yields a transactionally consistent snapshot of
a committed point. **The application keeps running**; no session is closed and
no turn is blocked.
### The procedure's guarantees, each tested
| | |
| --- | --- |
| Source opened **read-only** | via a `mode=ro` URI, so it is a guarantee and not an observation. The source file is byte-identical after a backup |
| Temporary destination, finalised by rename | the `.partial` file is gone and the final name never wore it |
| `PRAGMA quick_check` before the rename | a failed verification leaves **nothing** behind |
| Never overwrites | two backups get two names; two in the same second get two names |
| Failure reported | a clear error with the reason; the source untouched |
| No caller-supplied path | the endpoint accepts **no request body at all** — asserted against the OpenAPI schema, and by sending one anyway and checking where the file landed (H08) |
### Opened independently, and read
Every backup test **opens the copy as its own database** and asks what is in it —
a test that only checked a file appeared would pass against a `cp`. Verified
present in the copy: the schema, campaigns with their titles, story rows, the
head branch and depth **naming a turn that exists in the copy**, Save Points,
knowledge sources, state events, and the authoritative state document decoded and
compared against the live campaign's.
### Taken under load
The test that separates this from a file copy: turns are played from a second
thread **throughout** the copy. The resulting file passes `quick_check`, has no
dangling foreign key, has a transcript with **no gap in its depths**, and a head
that names a turn present in it. Which committed point it caught is not asserted
— that would be asserting on a race.
### No restore endpoint, and no download
Both are decisions, stated in the code and in the panel. Restore means replacing
the file the running process has open, which is how both copies are lost at
once; the stop-move-start procedure is in `DEVELOPMENT.md`. Download would put a
copy of every campaign into the browser's download directory and cache, which
for an application whose premise is that the story does not leave the machine is
a worse default than a path the reader can copy.
---
## O. Import validation and corrupt bundles
`test_m9_corrupt_bundles.py`. Every case starts from a **real export of the M9
fixture** and breaks exactly one thing, so what each test measures is that break
rather than a hand-written shape nothing ever wrote.
### Two rules, checked rather than assumed
**Nothing lands.** A refused import is checked by counting rows in **eleven
tables** before and after, not by trusting the status code. No campaign, no
branch, no orphan action, no Save Point pointing at nothing, no knowledge owned
by a campaign that does not exist.
**Nothing is fetched, read or run.** A URL in a bundle stays text; a path stays
text. No test needs a network guard to pass — there is no code path that would
use one.
### Coverage
| Class | Cases |
| --- | --- |
| Format / version | not an object; empty object; missing `format`; `format` of four wrong types; an unsupported future version, refused **with its own name in the message** |
| History graph | action on an unlisted branch; branch forking from one listed after it (which is also **how a cycle is made impossible rather than detected**); a branch forking from itself; a fork with no depth; an action with no depth; a negative depth; **two actions claiming one identity**; a turn with no live take; a parent naming a node the file lacks; **a node that is its own parent** |
| Head | past the retained story (refused, message names where the branch ends); four wrong types; a head branch that is not listed (repaired to the root, and the reasoning for why *that* one is repaired) |
| Save Points | coordinate beyond the story; branch not listed; blank name; the whole section not a list |
| State | sections not lists; an event that is not an object; an event with no type; **an event naming a turn the file lacks (refused)**; a proposal naming a turn the file lacks; an event whose proposal is gone (keeps its coordinate); a malformed state document; a malformed per-position snapshot |
| Knowledge | section not a list; no content; unknown classification; a source that is not an object; unreadable visibility; **content hash mismatch, recomputed and reported**; over the source cap; over the size cap |
| Provenance | a snapshot that is not an object (dropped, never guessed at); a nonsense knowledge block; a snapshot naming an impossible source (relinked to `null`, evidence intact) |
| Summaries | no coordinate (dropped, **never placed at a guess** — that is E03's leak by a new route); section not a list |
| Caps | actions, branches, and a body past the import ceiling (413 from the middleware, on the declared length, before any parse) |
| Inert data | URLs, `file://`, remote images; four traversal filenames; a title that looks like a command; an over-long title |
### The authoritative/derived distinction on failure
Two outcomes a caller can see, and they are distinguishable:
```text
4xx the authoritative import failed and no campaign exists
201 the authoritative import succeeded and the campaign is complete
```
A **rebuildable** index failing is neither, and does not roll the campaign back:
it is reported on the 201 as `import_warnings`, visible per source in the
Knowledge panel, and repaired by Reindex. The response model is a subclass
(`ImportedAdventureOut`) rather than a field on `AdventureOut`, because "which of
your indexes failed to rebuild" is a fact about one import, not a property of a
campaign.
The rollback is **explicit** rather than left to the session closing, so a test
can assert on it rather than on teardown.
---
## P. Migration and backward compatibility
### No schema change, proved
```
git diff 1ce9972 -- backend/app/migrations.py → 0 lines
mapped_column additions/removals in models.py → 0
PRAGMA user_version → 92 before, 92 after
```
`git diff` is necessary and not sufficient — a migration can also be *missing*,
and the failure then is a database that opens and quietly answers wrongly. So
the claim is proved the way M8 proved its own:
**A campaign is built by a server running the signed M8 commit**, from the git
worktree at `1ce9972`, using M8's own interpreter and M8's own code: story, an
alternate take, a Save Point, a manual state correction, memories, a summary,
three imported sources including a disabled and a narrator-only one, per-turn
context snapshots, and two Undos so the head is behind the tip. M9 then opens
that file.
```
$ python -m tools.m9_migration_proof open <db> --expected <m8-report.json>
{ "problems": [],
"schema_version_before": 92, "schema_version_after": 92,
"schema_version_second_open": 92,
"checked": { "families": 12, "snapshot_rows": 8,
"counts": { "actions": 16, "checkpoints": 1,
"knowledge_chunks": 3, "knowledge_sources": 3,
"memories": 2, "state_events": 9,
"state_proposals": 9, "summaries": 1 } } }
```
Twelve families compared and identical; **a second open migrates nothing**
(idempotence); and the M8-built campaign then exports as v3 and round-trips, and
Redo still works on it. Neither half of the tool imports the other's code — what
crosses is the database file and one JSON report.
### Older bundle seams
`test_m9_legacy_bundles.py` ages a real v3 export backwards, removing only what
each era's format genuinely could not carry — a checked-in fixture drifts, and a
hand-written one tests a shape nothing ever wrote.
| Era | No story lost | Nothing invented | Head |
| --- | --- | --- | --- |
| pre-M9 (v2) | ✓ | no events, no proposals, no summaries, no prompts | honoured, as v2 carried it |
| pre-M7 | ✓ | + no knowledge sources, no chunks | honoured |
| pre-M5 | ✓ | + empty state, no canon | honoured |
| pre-Save-Point | ✓ | + no Save Points | honoured |
| pre-active-head | ✓ | + no Save Points | **opens at the tip**, offers no Redo |
| v1 (flat) | ✓ | both attempts arrive, one live | tip |
| pre-M2 (with `scripts`) | ✓ | scripting keys ignored, not rejected | as its era |
And each era can still be *used*, not merely read: a pre-M5 campaign gains state
from its next turn, a pre-M7 campaign imports a source and it indexes, a
pre-Save-Point campaign can be given one.
**A pre-M9 campaign re-exports as v3 with its evidence sections empty**, because
the campaign genuinely has none — which is what lets a later reader trust a v3
file's empty `stateEvents` to mean "this campaign has no audit trail" rather
than "the file could not say".
---
## Q. Browser recovery workflow
**36/36 checks, zero failures**, on one frozen production build against a real
narrator across two installations.
### Environment
| | |
| --- | --- |
| Browser | Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1) |
| Application | the **frozen production build** `dist/assets/index-DcHbz7ga.js`, served by the real backend as static files — not a dev server |
| Backend | two `uvicorn app.main:app` processes on loopback, one per machine, each with its **own data directory** |
| Databases | machine A: a real campaign. Machine B: **a file that had never existed**, created by the migrations on first start |
| Narrator | `qwen2.5:3b-instruct-16k` on the trusted-LAN Ollama, over HTTPS with a privately issued certificate — **real inference for every turn** |
| Campaign | played over real HTTP: five turns, a retry, a Save Point, three imported sources (one disabled, one narrator-only), a manual state correction, and two Undos so the export is taken behind the retained tip |
### The §25 workflow, item by item
| | |
| --- | --- |
| **1. Export campaign** | **PASS.** Clicked in the browser; the page produced a **60,160-byte `ai-dnd-adventure-v3`** bundle |
| **2. Return to Campaigns** | **PASS** |
| **3. Import into a clean target** | **PASS.** Machine B had never had a database, its library was empty, and it had **no narrator configured at all**. The import reported no incomplete rebuild |
| **4. Open imported campaign** | **PASS** |
| **5. Opens at the exact active position** | **PASS.** The transcript matches the source exactly; Redo is offered; the retained future is in the database (**11 rows behind 7 on screen**); and all 4 narrated turns are found in the rendered page |
| **6. State is correct** | **PASS.** The authoritative document matches, the manual correction is in the restored audit trail, with the fact the reader asserted, and the whole trail arrived |
| **7. Save Points present** | **PASS**, in the browser's own panel, with the note |
| **8. Knowledge present, with classifications and status** | **PASS.** All three sources in the panel; the disabled one still disabled, the narrator-only one still narrator-only, every source reindexed and searchable, and the content **byte-identical by hash** |
| **9. Historical Inspect Context works** | **PASS.** 3 of 3 restored narrator turns show the prompt they were given; 3 name the passages they were shown; and the inspector opens through the browser's own control |
| **10. Undo/Redo coherent** | **PASS.** Both offered at the restored head; Redo moves forward into the retained future |
| **Backup control** | **PASS.** Taken from Settings; the panel names the file; the file is on disk where it says; it **passes its own integrity check opened alone**, and holds the restored campaign with all 11 of its rows |
| **I06 on the downloaded file** | **PASS.** No credential, no endpoint, no path from the exporting machine |
### Two substitutions, and both are on the browser's side of the line
Named precisely, because §30 of the brief is about not letting a substitute
stand in for the boundary it bypasses.
1. **The export's write to disk.** `downloadJSON` builds a Blob, makes a `blob:`
URL and clicks an anchor. **Headless Firefox does not complete that download
in this configuration** — the click fires, the page reports "Campaign
exported.", and no file appears in the download directory, in `~/Downloads`,
or anywhere under the snap's own tree. So `URL.createObjectURL` was wrapped
to capture the exact bytes the page handed the browser, and those bytes were
written to disk and used for everything downstream. *Proved:* the browser
assembled the correct v3 file and offered it. *Not proved:* Firefox's own
download writer, which is browser behaviour.
2. **The import's read of the file.** Firefox here is a **snap**, and its
sandbox will not let the content process read a path outside its
confinement. WebDriver hands the input the file successfully — the `File`
arrives with the right name and **the right byte count**, and `change` fires
— and `FileReader` then fails with `NotFoundError`. The click and the
handover are exercised; the last hop, the POST the page would make with the
bytes it read, was made against the same endpoint the page calls.
Neither is papered over, and neither leaves the end-to-end path unproven:
`test_m9_clean_import.py` does the whole of it — export, a real file on disk,
a real HTTP import into a second process against a database that never existed
— with no browser in the way, 11/11.
There is no unconfined Firefox and no Xvfb on this machine, so a windowed run
was not available. That is a property of the machine and is recorded as such.
### One observation worth a reviewer's attention
**The real narrator emitted no typed state blocks across five turns.**
`qwen2.5:3b-instruct-16k` narrated well and never produced a `state` block, so
the campaign's only state event is the **manual correction** — which is why the
audit-trail checks above read "1 events, source had 1".
That is not an M9 defect and not a regression: it is the behaviour ADR 003 and
ADR 010 exist for. The application owns state and the model does not, C04's
manual correction path is how a reader asserts what the model did not, and the
per-position snapshots keep the head, the transcript and the state in agreement
regardless. It does mean the browser run's state evidence is thinner than the
suites', where a scripted narrator produces twelve events and two proposals —
so the rich state case is proved there and the *real-model* case is proved here,
and this report does not let the second stand in for the first.
---
## R. Performance and size
### The M9 fixture
| | M8 (v2) | M9 (v3) |
| --- | --- | --- |
| Bundle | 29,630 B | 90,117 B |
| Actions | 22 | 22 |
| Source database | 258 kB | 319 kB |
| Export | 0.028 s | 0.044 s |
| Import (write + rebuild) | 0.063 s | 0.119 s |
### A long campaign — and the result that changed the plan
`python -m tools.m9_scale_report --turns 150 --budget 16384`. A real prompt
builder, a 16,384-token budget, and prose long enough that a turn is a turn.
```
turns actions bundle B snapshots B B/turn export s import s
25 51 311,346 130,766 5,230 0.066 0.171
50 101 883,616 279,208 5,584 0.129 0.335
75 151 1,715,306 439,634 5,861 0.269 0.232
100 201 2,870,121 675,880 6,758 0.572 0.371
125 251 4,313,937 951,870 7,614 0.518 0.603
150 301 6,022,430 1,242,324 8,282 1.012 1.109
```
A per-turn prompt contains the story so far, so carrying one per turn is
**O(turns²)**. That was the risk in the decision, and measuring it changed what
was done about it and what is claimed:
- **Uncompressed, the snapshots were 68% of a 9.7 MB file at 120 turns.** They
now travel compressed inside the bundle — the same JSON through the same
`compression.pack`/`unpack` the database column already uses, base64-encoded
so the file is still JSON. **7.9× smaller**, and every other section of the
file is still plain readable text.
- At M11's 100-turn certification target: **2.87 MB, 14% of the cap.**
- Export and import stay under 1.1 s at 150 turns.
### Where the bytes are, and what M9 actually cost
The same tool's per-section breakdown, at 150 turns. It exists because "the
prompts are 21% of the file" does not answer "what did M9 add" — the events, the
proposals and the summaries are v3 additions too, and a cost claim counting only
the prompts would understate it.
```
per-position state (v2 already) 4,474,294 74.3%
prompts (contextSnapshotZ) 1,245,144 20.7%
state proposals 78,938 1.3%
state events 41,870 0.7%
node ids + parentage 6,993 0.1%
summaries 3,958 0.1%
--- everything v3 added 1,376,903 22.9%
a v2 file of the same campaign 4,645,347
ceiling with v3 additions: ~279 turns
ceiling without them: ~318 turns
```
**The result is not the one the decision was worried about.** The largest thing
in a campaign bundle is not the prompts M9 added — it is the **per-position
narrative state document, at 74% of the file, which v2 already carried**. It
grows because a state document accumulates facts, so it is quadratic for the
same reason the prompts are, and it has been since M5.
Everything M9 added comes to **22.9%** of the file, and moves the import ceiling
from about **318 turns to about 279** — a cost of roughly **12% of reachable
campaign length** for closing the provenance gap, not the halving it looked like
before it was measured, and not the 11% a prompts-only reading would have
claimed.
Export and import stay under 1.2 s at 150 turns.
### N+1 and duplication
Checked for, and none found:
- The export undefers every deferred per-node column in **one** query, including
`context_snapshot`, and the proposals' `detail` once for the whole list.
- Nodes are referenced by their own exported ids, so the audit records need no
per-row lookup.
- Knowledge content appears **once** per source; the prompt appears **once** per
turn, on the live attempt, with superseded takes carrying only their own
slices (asserted).
- The import's flushes are bounded by the number of **knowledge sources**, not
by the size of the story: one after the nodes, one after the proposals, two in
`materialize`, and one per source — the last because `build_index` needs the
source's id, which is the same one-flush-per-source the live upload path
already does. Nothing flushes per action, per branch, per event or per
Save Point, which is where a long campaign would have felt it.
### Residual
The ceiling above is a **residual limit, stated with its measurement**. It is far
beyond M11's certification target, and the fix if a later milestone needs one is
a streaming or chunked import, which is architecture beyond M9. The asymmetry
worth naming for a reviewer: a campaign past that length can still be
**exported** and would be refused on **import** with a 413.
---
## S. Acceptance matrix
Every M9-owned item, with the evidence that decides it. Where a claim names a
boundary, the row says which boundary was actually crossed.
### I-series
| | Result | Evidence, and the boundary |
| --- | --- | --- |
| **I01** Export campaign | **PASS** | Every section present and non-empty on the M9 fixture, checked family by family rather than by file size. `test_i01_*`; reproducible with `tools/m9_portability_report` |
| **I02** Import exported campaign | **PASS** | **Second process, second directory, database file that never existed, exporting process stopped.** Transcript, state, canon, Save Points, knowledge, events and retained-tree size all match. `test_m9_clean_import.py`. Same round trip in-process in `test_m9_portability.py`, labelled there as the weaker of the two |
| **I03** Branch / disposable history | **PASS** | Both futures present and readable; the abandoned branch still records the depth it was left at; the superseded take still at its coordinate. No trimmed-export option exists, so that clause does not arise |
| **I04** Checkpoint export | **PASS** | Names, notes, coordinates; every one resolves; each restores to its own position and its own state; later history intact. A Save Point the file cannot satisfy is dropped, never retargeted |
| **I05** Knowledge provenance | **PASS** | Content byte-identical by SHA-256, classification, enabled, visibility, always-include, filename, parser/chunking versions. Verified **with the exporting machine stopped**, and in the browser |
| **I06** No API secrets | **PASS** | Searched against the file's **text**, with the inert `api_key` column written first so absence is evidence. Compressed snapshots decoded, not skipped. Checked in the unit suite, the two-process run, and the browser run |
| **I07** Undone active head | **PASS** | All three clauses. The exact head **across a machine boundary**; a head-less file opens at its tip with no Redo; a head past the story is refused with a message naming where the branch ends |
### L-series
| | Result | Evidence |
| --- | --- | --- |
| **L01** Atomic turn commit | **PASS (not regressed)** | M9 changed no turn-commit path; the inherited tests pass unchanged. M9 adds the import's own version, in two forms: a **refused** import writes zero rows across eleven tables, checked by counting; and a failure **deep inside the write phase** — past every planner check, with the branches, nodes, parentage, memories, head and Save Points already in the session — also leaves zero rows, which proves the transaction rather than the planner. The rollback is explicit, so the next import succeeds rather than inheriting a poisoned session |
| **L02** State reconstruction | **PASS** | Measured **after a round trip**: copy and source walked back three turns and forward three turns in step, document compared at every position. Restoration stays a snapshot read, not a replay |
| **L03** Save Point after restart | **PASS** | **Three processes.** Export from A, import into a clean B, restore a Save Point in B, kill B, start a third against the same file, ask again. Transcript, state and retained-tree size all match |
| **L04** Derived data rebuilt | **PASS** | Passages, FTS rows and vectors destroyed and rebuilt on a copy; authoritative campaign identical either side; retrieval works again; vectors alone gone still leaves lexical working. **Running it found finding 1** |
### M9-specific round trips
| | Result |
| --- | --- |
| Undone-head round trip | **PASS** (I07) |
| Branched campaign round trip | **PASS** (I03) |
| Checkpoint round trip | **PASS** (I04) |
| Knowledge provenance round trip | **PASS** (I05) |
| Derived indexes can be rebuilt | **PASS** (L04) |
| **Historical prompt/context round trip** | **PASS** — survives the source being deleted, in the copy |
| **State audit round trip** | **PASS** — events, proposals, and the manual correction still identifiable |
| **Summary lineage round trip** | **PASS** — the abandoned line's summary is still ineligible, before and after a rebuild |
| **Take parentage round trip** | **PASS** — the pager reads identically in copy and source |
| **Memory authority round trip** | **PASS** — a heuristic memory is not promoted |
| **Legacy round trips** | **PASS** — v1, v2, and four earlier seams; nothing invented at any of them |
| **Migration from an M8-built database** | **PASS** — 12 families, 0 problems, version 92→92→92 |
### Adjacent items re-verified rather than assumed
| | Result |
| --- | --- |
| **E03** abandoned summary cannot leak | **PASS** after a move, and after an index rebuild |
| **F07** heuristic memory is not canon | **PASS** — the authority survives the move |
| **C04** manual correction | **PASS** — still identifiable as one in the copy |
| **H08** path traversal | **PASS** — four shapes sanitised; the backup endpoints accept no path |
| **H09** ZIP slip | **N/A** — no archive format was introduced; the bundle is one JSON document |
| **G04** disabled source stays out of retrieval | **PASS** after the move |
### Synthetic substitutes, labelled
| Where | Substitute | Why it is sound |
| --- | --- | --- |
| Every backend suite | `ScriptedProvider` narrator, stub embedder/summariser | The property under test is the round trip, not the model. Both derived factories are stubbed, per M6's finding M6-F3 |
| `test_m9_clean_import.py` | the spawned server's deterministic narrator | Real processes, real HTTP, real files; only the model text is canned |
| Legacy-bundle seams | v3 exports aged backwards | A checked-in fixture drifts and a hand-written one tests a shape nothing wrote |
| **Browser run: the export's write to disk** (§Q) | the bytes the page produced, captured from `URL.createObjectURL` | Headless Firefox does not complete a `blob:` download here. Proves the browser assembled and offered the right file; does **not** prove Firefox's download writer |
| **Browser run: the import's read of the file** (§Q) | the POST the page would make, made against the same endpoint | Firefox is snap-confined and cannot read a WebDriver-uploaded path — the `File` arrives with the right size and `change` fires, then `FileReader` raises `NotFoundError`. The click and handover **are** exercised |
| **Not substituted** | the browser run's narrator and its two installations, the migration proof's M8 build, the backup's SQLite, every process boundary, the file-on-disk import in `test_m9_clean_import.py` | Each is the boundary the claim names. The end-to-end file path the browser could not complete is proved there, with no browser in the way |
---
## T. Security and local-only
| | |
| --- | --- |
| Internet | none. No M9 code path opens a socket. Export, import and backup are file and database operations |
| Embedded URLs | inert. A URL in a bundle is stored and displayed as text |
| Remote images | unchanged; CSP still `img-src 'self' data:` |
| Shell | none. No `subprocess` in any M9 application file |
| Imported content executed | never. It is stored and rendered through M8's safe Markdown path |
| Loopback | unchanged. The backup endpoints are on the same loopback API |
| Inference endpoint policy | untouched. `endpoints.py` unchanged |
| Cloud storage / remote backup | none. The backup writes to a directory beside the database |
### I06
Tested against the file's **text**, not against a list of columns — a field
added to a model the exporter walks would otherwise reach the bundle with no
column test noticing. The inert `api_key` column is written with a recognisable
value first, so its absence is evidence rather than a tautology. Searched for
and absent: the written secret, `api_key`, `apiKey`, the endpoint and its port,
and absolute filesystem paths. The compressed snapshots are **decoded** before
searching rather than skipped.
Checked in three places: the unit suite, the two-process run (against a real
endpoint URL), and the browser run. Measured on the M9 fixture, over 90,074
bytes of bundle **plus 154,279 bytes of decoded snapshot**:
```
'SECRET-MUST-NOT-TRAVEL' present: False (written into api_key first)
'api_key' / 'apiKey' present: False
the endpoint host and port present: False
'endpoint_url' present: False
'/home/' , '/tmp/' present: False
'secret.key' present: False
```
What **is** present, and correctly: the per-turn model name inside each
historical snapshot. That is a record of what the turn ran under, which §6.1
requires — not a configuration the import applies. §M has the distinction.
### File, path and archive safety
**No archive format was introduced**, so H09 does not arise: the bundle is still
one JSON document, and the compression in §R is an encoding of one field inside
it, not a container.
- No bundle-supplied value becomes a read or write path. `originalFilename` is
metadata; four traversal shapes are sanitised on import and asserted absent.
- Knowledge restores from the content in the bundle, never by reopening the
source machine's path — verified with the exporting process stopped.
- The backup endpoints accept no path at all.
---
## U. Findings
Numbered, including those fixed before the final tree.
### 1. Deleting a campaign leaked its lexical index, and broke the next import — PRE-EXISTING (M7), FIXED
**Severity: high.** An ordinary upload in an unrelated campaign failed with a
500.
The FTS5 index is a virtual table, so no foreign key reaches it and no
`ON DELETE CASCADE` covers it. Deleting a campaign cascaded
`knowledge_sources` → `knowledge_chunks` and stopped there, leaving one index row
per passage pointing at a chunk that no longer existed. Nothing read them — every
search joins through `knowledge_chunks` — so the leak was **invisible** until
SQLite handed the freed primary key out again, at which point the next source
imported into **any** campaign collided on `INSERT` and raised.
Worse, `clear_index` finds index rows *through* the chunks, so with the chunks
gone the orphans were unreachable: **Reindex, the documented repair, could not
repair it.**
Found by **running** L04 — it failed on the rebuild with an integrity error,
which is the only way this shows itself. Fixed at both ends: `importer.clear_campaign_index` removes
a campaign's index rows before it is deleted, and `fts.add` uses `INSERT OR
REPLACE` — the rowid is a chunk's primary key, so a row already there is by
definition stale. The second half means **a database already carrying the leak
repairs itself, with no migration**. Three regression tests.
### 2. An imported node with no state snapshot was given the head's state — PRE-EXISTING, FIXED
**Severity: medium.** M5's review finding 3, arriving through the import.
`tree.stamp_outcome` runs on every flush and fills a missing
`narrative_state_after` from *the campaign's current state*. On an import that is
the state at the exported head — so every position whose snapshot the file did
not carry came back holding the newest position's state, and an Undo to turn 2
showed what the story knew at turn 20.
`_write_nodes` now writes the empty document explicitly when the file carries
none. Legacy behaviour is unchanged: a pre-M5 bundle has no `narrativeState`
either, so the fallback was already writing `empty()` for those files.
### 3. The snapshot relink did not persist — FIXED
**Severity: high, and it would have shipped looking correct.**
`_relink_snapshots` edited the dict the attribute already held and assigned it
back. `context_snapshot` is a plain `CompressedJSON` column, not a
`MutableDict`, so SQLAlchemy tracks it by assignment: at flush the loaded value
and the current value were the same object, the history reported no change, and
no `UPDATE` was emitted. The import looked right in memory and wrote the
untranslated ids to disk.
Found because the test asserted the **outcome** — that the restored records name
this campaign's sources — rather than that the function was called. It now builds
a new document.
### 4. Cancelling the import file dialog hung the Import button — PRE-EXISTING, FIXED
**Severity: low, and user-visible.** `pickJSONFile` resolved on `change` only, so
closing the picker without choosing anything never settled the promise: the
Campaigns screen's `await` never returned, its `finally` never ran, and the
Import button stayed disabled reading "Importing…" until the page was reloaded.
Found while making the input reachable for the browser suite. The screen's own
comment already said "a cancelled file picker is not a failure worth a message" —
it had simply never received one. Now `oncancel` rejects with an empty message,
which is exactly what that comment describes.
### 5. The file input was detached from the document — FIXED
**Severity: none for a user; structural for testing.** `pickJSONFile` created an
input and clicked it without appending it, so no element existed for a test or
for WebDriver to hand a path to — the import workflow could only ever be checked
by calling the API underneath it, which is not the workflow. It is now in the
document and removed however the promise settles. Six tests.
**And a second lesson from fixing it.** The first attempt marked the input
`hidden`, which reads correctly and is wrong: a `hidden` element is
non-interactable, and WebDriver will set `files` on one **without dispatching
`change`** — the file lands and nothing happens, which is a worse failure than
the detached input because it looks like it worked. It now uses the ordinary
visually-hidden pattern — off-screen, zero-sized, `aria-hidden`, out of the tab
order — so no reader meets a stray "Choose file" control while the browser's own
dialog is what they are looking at. The test asserts `hidden === false` and says
why, so the next person does not "tidy" it back.
### 6. Five existing tests pinned the gaps M9 closes — REWORKED, NOT DELETED
Each asserted the M8 behaviour that was the debt. Following the rule
`BUILD-MILESTONES.md` records from M2 — move the instrumentation, do not delete
the test — each now asserts the new guarantee, and each carries a note saying
what it used to assert and why it changed.
| Test | Was | Now |
| --- | --- | --- |
| `test_historical_prompt_evidence_survives_an_export_round_trip` | the copy's turn has **no** snapshot (404) | it has the same snapshot, and its `source_id`s name this campaign |
| `test_export_and_import_round_trips_variants` | the pager reads `1/1` on the copy, "a gap in the import" | it reads `2/2`, as the source does |
| `test_an_unknown_format_is_refused` | used `v3` as its "from the future" placeholder — which M9 made real, so it started importing the file it meant to reject | uses `v99`, and asserts against `bundle.FORMAT` rather than a literal |
| `test_export_carries_the_whole_story` | `format == "ai-dnd-adventure-v2"` | `format == bundle.FORMAT` |
| `test_the_shipped_file_is_a_bundle_this_build_can_import` | `version == bundle.FORMAT` | `version in bundle.READABLE` — the shipped starter is a v2 file and was **not** regenerated, because rewriting a shipped asset to keep a test's equality holding is changing the evidence to fit the test |
### 7. The scale tool measured the wrong thing after compression — FIXED IN THE TOOL
It looked for the plain `contextSnapshot` key and reported 0% once the export
started writing `contextSnapshotZ` — i.e. it would have reported the evidence as
*omitted* when it was merely encoded, which is precisely the mistake the
portability report exists to avoid making about anything. Both tools now decode.
Recorded because a reviewer reading an intermediate number in this report's
history would otherwise be misled.
---
## V. Planning changes
Distinguished by kind, as the brief asks.
### Requirement corrections: none
M9 altered no product requirement. `SPECIFICATION.md` and
`SECURITY-THREAT-MODEL.md` are **unchanged**: §16 already required the export to
preserve the exact active position, §6.1 already required an exact prompt/context
snapshot per turn, and M9 implements both rather than redefining either. No
acceptance condition was weakened.
### Newly settled design decisions
| Document | Decision |
| --- | --- |
| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **Story Cards are compatibility-only legacy data and no longer enter the narrator's prompt.** The brief asked for this to be decided; §K has the evidence that they were the untracked path §73 already forbade |
| `DATA-MODEL.md` §29, `TECHNICAL-DESIGN.md` §9.3 | **The bundle carries historical evidence**, and the two-category rule becomes three |
| `DATA-MODEL.md` §29, `TECHNICAL-DESIGN.md` §9.3 | **The format is versioned rather than extended**, and why §9.1's and §9.2's reasoning does not stretch to cover this addition |
### Implementation facts recorded
`TECHNICAL-DESIGN.md` §9.3 (the v3 contract, the encoding, the two-phase
transaction) and **new §9.4** (the backup); `DATA-MODEL.md` §29 (the format, the
three categories, required vs optional, the pointers translated) and **§31**,
which listed summaries as derived-and-rebuildable and now says what "where
practical" excludes, and where stored prompts sit — a reader landing on §31
alone would otherwise conclude the export should be regenerating both;
`V1-ACCEPTANCE-TESTS.md` I01-I07 and L02-L04 results — with I05's M7-era limit
marked closed and **the original paragraph kept**, because the M9 decision is
only legible against it; `BUILD-MILESTONES.md` M9 status; `README.md` and
`VERSION.md`; `DEVELOPMENT.md` (a new section on the two recovery tools, taking a
backup, and the stop-move-start restore, plus the imported-campaign context note).
M8's report moved to `archive/milestone-reports/`, by the rotation convention
`planning/README.md` states.
---
## W. Residual risks and the M10-M11 handoff
### Residual risks
1. **A long campaign's bundle has a measured ceiling.** **~279 turns** against
the 20 MB import limit; a longer campaign can still be exported and would be
refused on import, which is the asymmetry worth naming. Far beyond M11's
100-turn target, which is 14% of the cap. M9's own additions account for
**12%** of that ceiling — the other 88% is the per-position state document v2
already carried, so lifting the ceiling means addressing *that*, not the
evidence. *Risk: low for v1; the fix is a streaming or chunked import, which
is architecture beyond M9.*
2. **`quick_check` rather than `integrity_check`** on a backup. It does the
structural work without the full index cross-check; a corrupt index that
`quick_check` misses would ride into the copy. *Risk: low, and the trade is
stated in the code — a backup verified too slowly to be taken is worse.*
3. **The backup is not offered on a schedule, and nothing prompts for one.**
*Risk: low, and outside M9's scope. `DEVELOPMENT.md` shows the `curl` for
`cron`.*
4. **`chunk_id` in a restored snapshot is stale.** It is a React key and a data
attribute, not a live pointer, and `chunk_index` plus `heading_path` still
identify the passage. *Risk: very low; documented in the code.*
5. **The importing machine's context window may differ from the source's.**
Documented rather than detected; detection is M11's.
6. **This machine cannot drive a file into or out of the browser.** Firefox here
is a snap, so its sandbox blocks reading a WebDriver-uploaded path, and its
headless download manager does not complete a `blob:` download. §Q labels
both substitutions and the end-to-end file path is proved without a browser
in `test_m9_clean_import.py`. *Risk: low for the product, real for future
evidence — M11 will meet the same wall on any file-based browser claim, and
the cheapest fix is an unconfined Firefox or Xvfb on the certifying machine.*
### Carried to M10
- Media tables do not exist. §L records that as **not applicable** rather than
inventing schema, and M10 owns the coordinator, the providers and the jobs.
- The `scene` section of the state document is the extension point that exists
today, and it round-trips.
### Carried to M11
- **The context window**, end to end: detection, the 100-turn certification, and
whether a Settings warning that reads the real window belongs in v1.
- **The streaming import**, if the bundle ceiling above is judged too low.
- Contrast and focus measurement, tablet tuning, and the broader WCAG audit
(inherited from M8's §U, untouched by M9).
- **An unconfined browser on the certifying machine**, so that a file-based
browser claim can be made end to end rather than in two labelled halves
(residual 6).
### Recorded here but not M9's: four hands-on playtest findings
A play session against **accepted, signed M8** — after M8 was accepted and
before M9 closeout — surfaced four product-quality observations. **None is an M9
defect, none was caused by M9, and none blocks M9 acceptance.** They are written
up in full in **§Y**, and durably in `BUILD-MILESTONES.md` under M11 so they
survive this report's archival:
| | Owner |
| --- | --- |
| A. The browser tab still reads `AI D&D` | M11 release polish |
| B. After Undo, the reader cannot tell where they are | M11 UX/release polish |
| C. The narration-length setting has no measurable effect | M11 realistic-model behaviour |
| D. Character identity / coreference confusion (root cause **unknown**) | M11 realistic-model / context diagnostic |
D is the significant one, and its evidence is gone — the disposable playtest
database was destroyed, so no root cause is claimed. §Y records the reproduction
and classification plan, and one structural fact worth checking first: the
narrative state permits two entities to share a display name and reports
nothing.
### Still open from earlier milestones, and not M9's
Cross-layer duplication (`CONTEXT-AND-MEMORY.md` §22); the discarded-history
recovery screen (§63); whole-transcript copy and story search (§77, §78); the
read-only RPG world state; the inert legacy tables and the dual-dialect
migration code awaiting their cleanup migration.
---
## X. Final milestone assessment
### Is M9's Definition of Done satisfied?
`BUILD-MILESTONES.md` states it as: *"A campaign can be safely exported,
imported into a clean data directory, and reopened at the exact intended active
position with authoritative history/state intact."*
**Yes.** Measured three ways — in the suite, across two server processes with
two data directories and the exporting process stopped, and in a real browser
against a real narrator.
### The brief's closing questions, answered directly
| Question | Answer |
| --- | --- |
| **Can a campaign be moved into a clean data directory and reopened at the exact head?** | **Yes.** Into a database file that had never existed, from a second process, with the first stopped. The fixture's head is behind its own branch's tip *and* the abandoned line's, and is not the newest row written — so an importer guessing the tip, the newest row or the deepest row lands elsewhere. It opens where it was left, the retained future is there, and Redo reaches the same next turn in copy and source |
| **Is all authoritative state intact?** | **Yes.** The document at the exported head matches exactly; every historical position matches, walked in step; restoration stays a snapshot read rather than a replay; and the audit that explains the state travels with it — including the manual correction, which is the one change no narration explains |
| **Is historical prompt/context evidence intact?** | **Yes**, and this is the M8 handoff closed. An old turn in a restored campaign shows the prompt it was actually given and the passages it was shown, **after the source has been deleted in the copy** — asserted directly, not argued |
| **Are Save Points intact?** | **Yes.** Names, notes and coordinates; each resolves; each restores to its own position and its own state through M3's head movement; later history survives. A Save Point the file cannot satisfy is dropped, never retargeted |
| **Is knowledge intact without original filesystem paths?** | **Yes.** Content identical by SHA-256, classifications and lifecycle intact, searchable immediately — verified with the exporting machine's process dead and its directory holding a database the importer never opened |
| **Can derived data be rebuilt?** | **Yes**, and running that test found a defect that had made Reindex — the documented repair — unable to repair the state it most needed to. Both ends are fixed, and a database already carrying the damage repairs itself |
| **Can a consistent SQLite backup be created?** | **Yes**, through SQLite's online backup API, **while the application is being written to**, verified with `quick_check` before it is kept, and opened independently and read in every test. It works in the Docker image and lands on the mounted volume |
| **Are legacy bundles still supported?** | **Yes.** v1, v2, and every seam from pre-active-head onward — with nothing invented at any of them. A pre-M9 file gets no manufactured audit trail; a pre-M3 file opens at its tip *because that is the position it recorded* |
| **Are any required M9 conditions unverified?** | **No.** Every required condition has evidence at the boundary it names. What remains is stated as residual risk in §W with its measurement, not as an unverified claim |
| **Is it safe to proceed to M10 after review?** | **Yes**, on this evidence. M9 changed no schema, altered no product requirement, and added no media surface. It leaves M10 the `scene` section of the state document as the extension point that exists today, and §L records the absence of media tables as *not applicable* rather than inventing schema to satisfy the word "metadata" |
### What a reviewer should look at hardest
Named deliberately, because a report that only presents its strengths is harder
to review than one that points at its own seams:
1. **The version bump (§D).** It is the one place M9 chose differently from two
earlier milestones that faced a similar decision. The argument is that an
absent key is unambiguous for a *position* and ambiguous for *evidence*; a
reviewer who disagrees should say so, because it is a contract decision and
not an implementation detail.
2. **The story-card decision (§K).** It removes something from the narrator's
prompt, which is a behaviour change in a portability milestone. The brief
authorised it and §73 required it; the judgement to check is whether stopping
at the prompt — and leaving the summariser's roster alone — is the right
line.
3. **The bundle ceiling (§R).** Real, measured, and stated rather than hidden.
Worth checking that ~279 turns is acceptable for v1 given a 100-turn
certification target, and that the export-succeeds/import-refuses asymmetry
is tolerable until a streaming import exists.
4. **Finding 3.** It would have shipped looking correct in memory and wrong on
disk. It was caught only because the test asserted the outcome rather than
the call, which is worth generalising.
### Not done, and deliberately
No image, video, TTS, STT, media provider or job architecture; no M10 media
coordinator; no 100-turn certification; no WCAG audit; no new context-window or
provider architecture; no discarded-history recovery browser. Retained history
survives correctly so a later screen can use it, which was M9's part of that.
---
## Y. Post-M8 hands-on playtest findings
**These are not M9 defects, were not caused by M9, and do not block M9
acceptance.** They are recorded here because this is the report a reviewer is
holding, and because the M9 report will be archived when M10's replaces it —
`BUILD-MILESTONES.md` carries the durable copy under M11 for that reason.
They come from a real play session against **accepted, signed M8** — a real
browser, a real trusted-LAN Ollama, narrator `qwen2.5:3b-instruct-16k`, and a
disposable isolated campaign database. **That database was deliberately
destroyed afterwards**, so the stored context snapshot for the turn in finding
D no longer exists. Everything below is therefore recorded as an *observed
symptom*, and no root cause is claimed that cannot now be proven.
**Nothing in this section changed any application code.** Where the text states
how the product behaves today, it is from reading the code and running it, and
it is labelled as a mechanism rather than as a proven cause of what was seen.
### A. The browser still calls the product "AI D&D"
**Observed:** the browser tab/title reads `AI D&D`.
**Verified:** `frontend/index.html` line 18 is `<title>AI D&amp;D</title>`. It
is the inherited upstream title and M8 did not change it.
**Not a false claim by any accepted document.** M8's report never asserted the
title was changed — its terminology audit covered `branch`, `fork`, `node`,
`head` and `depth`, and its "does it feel like the intended storyteller"
argument lists the navigation, the inspector, the glyph buttons and the removed
screens. So this is an **uncovered gap**, not documentation that needs
correcting. No documentation was corrected, and no code was changed.
**The naming question is genuinely open, and should not be closed by a
find-and-replace.** The planning package and `README.md` call the product
*Adventure Storyteller*, but `SPECIFICATION.md` is explicit that the engine must
stay genre-agnostic — science fiction, mystery, horror, historical, westerns —
and *Adventure* is narrower than the product it names. A repository-wide rename
was deliberately **not** performed in this pass. For planning purposes a neutral
working name such as **Interactive Story** is used, and a browser-tab form such
as `<Campaign Name> — Interactive Story` is a candidate, not a decision.
**Owner: M11 release polish.** No earlier milestone touches the shell metadata.
It is a small, self-contained change whose only hard part is the naming
decision, which is the repository owner's.
### B. Undo/Redo loses the reader's orientation
**Observed:** Undo worked, and it was hard to tell which point in the story the
reader had moved to.
**Not a correctness problem.** M3's active-head semantics behaved correctly and
M9 re-verified them at every boundary (§E, §S). This is presentation.
**What the spec said, and why that was not enough.** `BROWSER-UX-SPEC.md` §8 —
*"The current endpoint should be clear."* — is five words, and the second
sentence beside it ("the input box always continues from the currently active
story head") was **already true** while the reader was lost. A requirement that
a working implementation satisfies while a real user cannot answer the question
is too vague to hold the behaviour. §8 has been strengthened in this pass; no UI
text is prescribed, because none is ratified.
**The requirement, stated without wording:** after Undo, Redo, a Save Point
restore, an edit to an earlier turn, or any other movement of the active
position, a reader should be able to tell where they now are in the visible
story **without needing implementation terminology** — `branch`, `head`, `node`
and `depth` remain forbidden at the surface (§38 of the spec, and M8's audit).
A lightweight indicator such as `Moment 8` → `Moment 7`, optionally noting that
later story is still available, is a candidate. **Exact wording deliberately not
settled here.**
**Owner: M11 UX/release polish**, with a browser regression scenario.
### C. Narration length did not feel effective
**Observed:** setup offered roughly *one paragraph* / *2-4 paragraphs* / *longer
exposition*; the reader chose **2-4 paragraphs** and felt replies were
substantially longer than that.
**Deliberately not recorded as "the model ignored instructions."** The evidence
does not establish that, and there is a mechanism in the code worth checking
first.
**The mechanism, verified by reading and running the code.** There are **two
independent length controls, and they do not reference each other**:
1. The setup choice becomes **one English sentence** appended to the campaign's
`ai_instructions` — `"Keep responses to roughly two to four paragraphs."`
(`frontend/src/pages/NewCampaign.jsx`, `LENGTH_SENTENCE`). It changes **no
generation setting**. It sits in the **cached system block**, near the top of
the prompt.
2. `context/builder.py`'s `length_hint()` derives a numeric word range from
`Settings.max_output_tokens` — a **global application setting**, not the
campaign's choice — and emits it as a `[Hard limit: …]` line placed **after
the history**, near the end of the prompt.
At the default `max_output_tokens = 800`, that second line reads, measured:
```
[Hard limit: this turn must not exceed 506 words, and it should not stop short
of about 177. Prefer the lower end of that range unless the scene genuinely
needs more. Finish the narration and append the state block well inside the
limit.]
```
**and it is byte-identical whether the reader chose brief, medium or long.** A
reader asking for two to four paragraphs is simultaneously told, in the more
recent and more numerically explicit of the two instructions, not to stop short
of about 177 words and that up to 506 are permitted.
**This is a mechanism, not a proven cause.** Whether it produced what this
reader saw needs the reproduction below; the narrator's own instruction
following is also in play, and both could contribute.
**What a reproduction must measure**, rather than judge by eye:
1. exactly what length instruction(s) enter the **stored** prompt — both of the
above, with their positions;
2. whether the setup choice changes `max_output_tokens` or any generation
setting (**today: it does not**);
3. actual words, tokens and paragraph counts across repeated realistic turns per
setting, so a *directional* effect can be shown or disproved;
4. at least the reference 3B narrator **and** a stronger local narrator, since
instruction-following differs.
**A design candidate, explicitly not ratified:** clearer approximate targets —
`Brief ~100-200 words`, `Standard ~200-400`, `Detailed ~400-700` — with the
setting actually moving the numeric budget. **Do not hard-truncate prose**: the
state block is emitted last and truncation removes it, which is the failure
`length_hint`'s own comments exist to avoid.
**Owner: M11 realistic-model behaviour validation**, with a realistic-model test
rather than only a scripted-provider one.
### D. Character identity / coreference confusion — the most important one
**Observed:** a story established four people in an office — **Bill**
(protagonist), **Roger**, **John**, **Alice**. Later narration treated Alice as
though there were two different Alices, in language equivalent to *"Alice
wondered what Alice was doing."*
**Root cause: UNKNOWN, and it cannot now be established.** The disposable
playtest database was deliberately destroyed, so the stored context snapshot for
that turn is gone. This report claims neither an inference-model defect nor an
application defect. The candidate classes are:
| | |
| --- | --- |
| **Model failure** | the stored prompt correctly identifies one Alice and the 3B model still makes a coreference error |
| **State failure** | the narrative state holds duplicate or conflicting Alice records |
| **Context/derived failure** | state is correct, but assembled context, a summary, a retrieved memory or the history rendering presents two identities |
| **Combined weakness** | the context is not contradictory but is insufficiently explicit for a small model, producing an avoidable failure |
**One structural fact a reproduction should check first**, verified by reading
the code and stated as a fact about the implementation rather than as a cause:
> **The narrative state permits two distinct entities to share one display
> name, and nothing reports it.** Entities are keyed by the id the model
> supplies (`state["entities"][event["entity"]]`, `narrative/apply.py`), with
> `name` a separate display field. `validate.py`'s `DUPLICATE_ENTITY` rejects
> re-creating an entity **with the same key** — `key in known` — and there is no
> check anywhere in the narrative layer on the display name. So
> `create_entity(entity="alice", …, name="Alice")` followed by
> `create_entity(entity="alice_2", …, name="Alice")` both succeed, and the state
> then holds two entities that both render as "Alice".
That is exactly one of the failure modes this finding describes. It does **not**
establish that it happened here — no evidence survives — and a reproduction may
well land on one of the other three classes instead.
**Owner: M11 realistic-model / context diagnostic.** The test plan is below and
in `BUILD-MILESTONES.md` under M11.
### The M11 character-identity diagnostic, as a test plan
Deterministic setup with at least a protagonist and three same-scene
supporting characters — **Bill** (protagonist), **Alice**, **Roger**, **John**
— with unambiguous identities and roles established up front. Then a
multi-character interaction over enough turns to stress: pronouns; dialogue
attribution; people entering and leaving; reference by name; reference by role;
and one character speaking *about* another.
**Detect and report, at minimum:** duplicate character creation; same-name
entity duplication; protagonist identity drift; dialogue attributed to the wrong
person; a character referring to themself as a separate same-named character;
and state/context disagreement about identity.
**On any failure, preserve and report all of:** the authoritative state
immediately before generation; the exact stored context/prompt snapshot; the
recent-history section; summaries; retrieved memories; imported knowledge if
any; the narrator output; and the model identifier and settings. **M9 makes all
of that portable** (§H), so a failing campaign can now be exported and handed to
whoever investigates it — which is the practical reason this diagnostic is
newly worth writing.
Then classify from the evidence:
```text
STATE DEFECT
CONTEXT ASSEMBLY DEFECT
DERIVED MEMORY/SUMMARY DEFECT
MODEL FAILURE WITH CORRECT CONTEXT
AMBIGUOUS / MULTIPLE CONTRIBUTORS
```
Two rules for whoever runs it: **do not "fix" a model failure by changing
authoritative story state**, and **do not blame the model if the prompt already
contained the identity error.** That distinction should become part of M11's
realistic-model review methodology rather than a one-off judgement.
### The standard fixture does not cover this failure class
Reviewed, as asked. `TEST-CAMPAIGN-FIXTURE.md` stresses knowledge boundaries,
secrets, authority precedence, branch leakage and possession — its seven
deliberate traps are all of those — and it contains **no identity trap at all**;
the word *coreference* does not appear in it. Its on-stage cast is effectively
two people, Aldric and Mara, with Edrin established as missing rather than
present.
So the user's suspicion is correct: **same-scene multi-character identity
continuity is not exercised by the standard fixture.**
**The established fixture was deliberately not modified.** It is the
deterministic baseline several milestones' results are compared against, and
changing it would invalidate those comparisons. A **companion** fixture is
proposed instead, in a new appendix to that document, named
`Multi-Character Identity Test` and explicitly additive.
---
*Written at the end of implementation, before review. The tree is staged; §32 of
the brief governs what happens next.*