The release-validation milestone, and the thing it had to settle first was whether any of the earlier evidence meant what it said. M8 measured a deployment enforcing a 4,096-token input window while the application budgeted 16,384. Every request returned 200. What Ollama does with the excess is drop the oldest tokens, and the oldest tokens here are the system block — the narrator's rules and the campaign canon. A hundred-turn certification against that server would have looked perfect and proved nothing, which is why this milestone could not begin with a hundred turns. So the application asks now. Ollama's window is a property of how a model was loaded rather than of the request — sending num_ctx is accepted, ignored, and worse, reloads the model at the server's own default — so the only honest move is to find out and then tell the truth about it. /api/ps reports what a resident model is being served with, /api/show what an unloaded one will load with, both on the same host inference already uses, through the same endpoint policy and the same TLS trust store. A verified window is a ceiling on the budget; an unverified one leaves the budget alone and is recorded as unverified in the turn's own provenance, so an old turn can be asked afterwards whether it was built against a checked window. There is no third behaviour, and in particular no hard-coded 4,096: a number the server did not say would be right on one machine and wrong on the next. The proof that this is doing something is a campaign whose canon sits at the front of the prompt, 120 turns of history, and a 4,096-token window. The canon is still there afterwards and the oldest history is gone. The same campaign built the old way produces a prompt more than twice the window — the defect, reproduced, so the fix is measured against it rather than asserted. Two defects the validation found on its own, and they are the same defect twice: something was true and nobody was told. A manual state correction of four changes with one bad reference applied three, returned 201, and said nothing — while recording the refusal on the audit row nobody reads. It came to light because the identity diagnostic's own fixture was refused that way and the whole run proceeded on a campaign with no scene, which would have read as a model failure. And the narration-length setting moved no number: brief, medium and long each became one English sentence, while the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three. Both now say what they did. The other two post-M8 findings are closed as well. The tab said AI D&D, which no document had ever claimed it did not; it says Interactive Story now, with the open campaign first, and the name is the owner's decision rather than a find-and-replace to something narrower than the engine. After an Undo the reader could not tell where they had landed; the control row now ends with "Moment 11 · later story ahead", from the server's own answer, in the word the transcript already uses, with none of head, branch or depth anywhere near it. The identity diagnostic exists and the root cause does not. That campaign was destroyed, so no cause can be established — what M11 owes the finding is something that can classify the next occurrence, and a diagnostic that makes only the judgements a program can honestly make: duplicate keys, shared names, protagonist drift, state and context disagreeing. Whether prose misattributed a line is left to a person reading it beside its prompt, because a regex cannot read dialogue and one that pretended to would produce exactly the confident wrong answer this finding is about. Its detectors are proved to fire against a planted second Alice. Two entities may still share a display name. That was checked first, as the finding asked, and left permitted: a mother and a daughter, or a stranger giving a false name, are ordinary fiction, and refusing them to guard against a model mistake would refuse the wrong thing. What was missing was that it happened silently. It is reported now. Evidence, not inference: a hundred accepted turns against a real narrator with genuine process restarts; a real browser against the built SPA; a container with no network at all; a campaign moved into a data directory that never existed. Each was discarded and re-run whenever the product changed under it, and the runs that were thrown away are listed in the report with the reason, along with ten defects in the harnesses themselves — because a harness that has only ever agreed with itself is not evidence, and two of M8's five harness defects were masking real ones. No dependency was added, removed or upgraded. No acceptance test was retired, relaxed or reclassified. M11 is implemented and verified; it is not accepted, and there is no release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
1623 lines
88 KiB
Markdown
1623 lines
88 KiB
Markdown
# M9 — Export, Backup, Recovery, and Migration Hardening
|
||
|
||
**Implementation report, written for an independent reviewer.**
|
||
|
||
This is a set of claims with the measurements attached. It is not a record of
|
||
acceptance: M9 is implemented and verified, and acceptance comes after review.
|
||
|
||
Two conventions carried from M8, because they are what make a report checkable:
|
||
|
||
- **Where a claim names a boundary, the boundary was crossed.** "A clean data
|
||
directory" is a second process against a database file that has never existed.
|
||
"A restart" is a different PID. "A backup" is opened as its own database and
|
||
read. "Migration from M8" is a database built by a server running the signed M8
|
||
commit. Where a substitute was used, it is labelled.
|
||
- **Every defect found is numbered in §U, including the ones fixed before the
|
||
final tree.** Seven are recorded. Three were found by *running* M9's own tests
|
||
rather than by reading them, two more by building the browser harness, and
|
||
**four of the seven predate M9** — they were reachable in M8 and nothing had
|
||
gone looking.
|
||
|
||
---
|
||
|
||
## A. Executive result
|
||
|
||
**PASS.**
|
||
|
||
Every required M9 condition is verified, by the boundary it names. The Definition
|
||
of Done — *a campaign can be safely exported, imported into a clean data
|
||
directory, and reopened at the exact intended active position with authoritative
|
||
history/state intact* — is satisfied and is measured three separate ways: in the
|
||
suite, across two real server processes with two data directories, and in a real
|
||
browser against a real narrator.
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Bundle format | **`ai-dnd-adventure-v3`.** v1 and v2 still import; every seam from pre-active-head onward is tested |
|
||
| Historical prompt provenance | **carried.** The M8 handoff is closed |
|
||
| Story cards | **compatibility-only**, and out of the narrator's prompt |
|
||
| Summaries / memories | **carried**, with the coordinates that decide eligibility |
|
||
| Derived indexes | **not carried, and rebuilt** — with the rebuild path fixed |
|
||
| SQLite backup | **SQLite online backup API**, verified before it is kept |
|
||
| Schema change | **none**, proved against a database M8's own code wrote |
|
||
| Backend | **1,102 passed, 14 skipped, 0 failed** in 14:54 (M8 baseline: 950) |
|
||
| Frontend | 145 passed (was 132), lint and production build clean |
|
||
| Browser | **36/36 checks**, one frozen build, real narrator, two installations (§Q) |
|
||
| Docker | production image builds clean; the backup works inside it, on the mounted volume |
|
||
|
||
Nothing is claimed here that was only compiled. §S maps every acceptance item to
|
||
the evidence that decides it, and names what is asserted from a synthetic
|
||
substitute.
|
||
|
||
---
|
||
|
||
## B. Repository and provenance state
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Branch | `m9-recovery` |
|
||
| Base | `1ce99727606124b0f8b3dc3b88842410c874ee14` — *M8: the browser becomes the storyteller* |
|
||
| Base signature | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569` |
|
||
| Working tree at start | clean |
|
||
| `HEAD` at writing | still the M8 base — M9 is staged, not committed (§32) |
|
||
| Upstream ancestry | `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` **is** an ancestor |
|
||
| `LICENSE` | unchanged; `git diff 1ce9972 -- LICENSE` is 0 lines |
|
||
| Schema version | **92 before, 92 after** |
|
||
| Migrations | **none added.** `git diff 1ce9972 -- backend/app/migrations.py` is 0 lines |
|
||
| Columns | **none added or removed.** 0 `mapped_column` changes in `models.py` |
|
||
|
||
### Change inventory
|
||
|
||
```
|
||
backend/app/bundle.py 928 +- the v3 format, both directions
|
||
backend/app/backup.py 278 NEW — the online backup
|
||
backend/app/routers/backups.py 75 NEW — two endpoints
|
||
backend/app/routers/adventures/bundle_io.py 59 +- the two-phase import
|
||
backend/app/context/builder.py 63 +- story cards out of the prompt
|
||
backend/app/knowledge/fts.py 44 +- the index leak (finding 1)
|
||
backend/app/knowledge/importer.py 22 + clear_campaign_index
|
||
backend/app/schemas.py 19 + ImportedAdventureOut
|
||
backend/app/routers/adventures/crud.py 8 + clear the index before delete
|
||
backend/app/main.py 6 +- register the router
|
||
backend/app/starter.py 4 +- materialize returns a pair
|
||
|
||
frontend/src/pages/Settings.jsx + the backup control
|
||
frontend/src/components.jsx +- the file picker (finding 4)
|
||
frontend/src/api.js, styles/library.css +
|
||
|
||
backend/tests/m9_fixture.py 406 NEW — the portability fixture
|
||
backend/tests/test_m9_portability.py 1,186 NEW
|
||
backend/tests/test_m9_corrupt_bundles.py 689 NEW
|
||
backend/tests/test_m9_clean_import.py 460 NEW — two processes
|
||
backend/tests/test_m9_backup.py 451 NEW
|
||
backend/tests/test_m9_legacy_bundles.py 365 NEW
|
||
backend/tools/m9_portability_report.py 354 NEW — the baseline tool
|
||
backend/tools/m9_migration_proof.py 331 NEW
|
||
backend/tools/m9_scale_report.py 221 NEW
|
||
frontend/src/pages/backup.test.jsx 116 NEW
|
||
frontend/src/filePicker.test.jsx 85 NEW
|
||
```
|
||
|
||
Five existing test files changed. Each had pinned a gap M9 closes, and each was
|
||
**reworked to assert the new guarantee rather than deleted** — the rule
|
||
`BUILD-MILESTONES.md` records from M2 about moving instrumentation rather than
|
||
removing it. §U-6 lists them.
|
||
|
||
---
|
||
|
||
## C. The M8 portability baseline, measured
|
||
|
||
The brief required the baseline to be **measured, not assumed**. It was, with a
|
||
tool a reviewer can rerun on either commit:
|
||
|
||
```bash
|
||
cd backend && python -m tools.m9_portability_report # or --json
|
||
```
|
||
|
||
It builds the M9 fixture (§E) in a throwaway database, exports it, imports the
|
||
result, and classifies every data family. Run on the M8 commit and on the M9
|
||
tree, the difference is the milestone's claim in reproducible form.
|
||
|
||
### The fixture it measures
|
||
|
||
Deliberately not the standard Continuity Test: that one is shaped to read like a
|
||
story, and this one is shaped to break a round trip. 22 actions, 2 branches, 2
|
||
Save Points on different branches, 1 superseded take, 12 state events, 12
|
||
proposals, 11 stored prompts, 2 memories, 2 summaries (one on the abandoned
|
||
line, one at the head), 5 imported sources across all three classes including
|
||
one disabled and one narrator-only, and a manual state correction.
|
||
|
||
The head ends **behind the retained tip of its own branch and behind the
|
||
abandoned line's**, and the last thing the fixture does is an Undo — so the head
|
||
is not the newest row written, not the deepest row, not the tip, and not on the
|
||
branch holding the most story. An importer guessing any one of those lands
|
||
somewhere else.
|
||
|
||
### M8 — `ai-dnd-adventure-v2`, 29,630 bytes
|
||
|
||
| Verdict | Family |
|
||
| --- | --- |
|
||
| PRESERVED | campaign identity, transcript, branches, branch disposition, active head, alternate takes, Save Points, narrative state (current), narrative state (per position), imported knowledge, memories, scene metadata, story cards |
|
||
| **OMITTED** | **take grouping** — SP9 parentage, so imported nodes landed parentless |
|
||
| **OMITTED** | **state events** — the audit half of §17's hybrid |
|
||
| **OMITTED** | **state proposals** — what the model asked for, and what was refused |
|
||
| **OMITTED** | **manual corrections** — indistinguishable from what the story established |
|
||
| **OMITTED** | **historical prompt/context** — the M8 handoff |
|
||
| **OMITTED** | **retrieval provenance** — which passages an old turn was shown |
|
||
| **OMITTED** | **per-turn model settings** — what a historical turn ran under |
|
||
| **OMITTED** | **knowledge parser versions** |
|
||
| **OMITTED** | **summaries** — only the lineage-less mirror column travelled |
|
||
| **OMITTED** | **memory authority** — a heuristic memory imported as accepted story |
|
||
|
||
Two of those were visible from outside, through the API a reader reads: after a
|
||
round trip the copy's **state events** and **summaries** differed from the
|
||
source's. The rest were invisible until something needed them.
|
||
|
||
Derived and correctly not carried: knowledge passages, the FTS index, knowledge
|
||
and memory embeddings, the branch lineage cache, derived status.
|
||
|
||
### M9 — `ai-dnd-adventure-v3`, 90,117 bytes on the same fixture
|
||
|
||
**Every family PRESERVED. No disagreement on any reader-visible family.**
|
||
|
||
The file is 3.0× larger on this fixture; §R has the measurement on a long
|
||
campaign, where the ratio is very different and the conclusion is not the
|
||
obvious one.
|
||
|
||
---
|
||
|
||
## D. The final bundle contract
|
||
|
||
Documented in `DATA-MODEL.md` §29 and `TECHNICAL-DESIGN.md` §9.3; the reasoning
|
||
is repeated in `bundle.py`'s own docstring, which is where the next maintainer
|
||
will be standing.
|
||
|
||
### Why the version was bumped
|
||
|
||
Everything M9 adds *could* have been an optional key read with `.get`, the way
|
||
`persona`, `checkpoints`, `narrativeState` and `knowledge` each were — §9.1 and
|
||
§9.2 each considered a bump and correctly declined one.
|
||
|
||
That mechanism stops working here, and the reason is the rule the format already
|
||
lives by. A v2 file with no prompt provenance is **ambiguous**: written before
|
||
M9, when no file could carry one, or by M9 from a campaign whose turns predate
|
||
the column? Those are different facts and a reader has to be able to tell them
|
||
apart — the same distinction I07 draws when it says a file written before the
|
||
head was carried opens at the tip *because tip was the only position that format
|
||
could represent*. **A version number is how a recovery file states what it was
|
||
capable of recording.**
|
||
|
||
The family name is unchanged, deliberately: renaming it would break every reader
|
||
for no gain, and `PROVENANCE.md` is where the fork is recorded.
|
||
|
||
```text
|
||
v1 a linear story, its turns, and its retries as a repeating group
|
||
v2 + the tree, the live flags, the after-snapshots, the chosen head,
|
||
Save Points, the narrative state document, imported knowledge
|
||
v3 + state events and proposals, historical prompt/context provenance,
|
||
lineage-anchored summaries, take parentage, memory authority
|
||
```
|
||
|
||
The importer reads all three. A version it has never heard of is **refused with
|
||
its own name in the message**, rather than read as the newest it knows.
|
||
|
||
### Three categories, not two
|
||
|
||
§31 of `DATA-MODEL.md` distinguishes authoritative from derived, which was
|
||
sufficient until M9 had to decide about stored prompts. They *are* derived — a
|
||
machine assembled them — and they must travel anyway:
|
||
|
||
```text
|
||
chosen the story, the head, the takes, the Save Points, the
|
||
classifications, the canon travels
|
||
evidence the state events and proposals, the per-turn prompt and the
|
||
passages it was shown, the model and generation settings
|
||
that turn ran under travels
|
||
rebuildable knowledge passages, the FTS index, embeddings, the branch
|
||
lineage cache rebuilt on import
|
||
```
|
||
|
||
The test separating the last two is **not** "could this be recomputed" but
|
||
"would a recomputation answer the same question". A rebuilt FTS index answers
|
||
the same question. A rebuilt prompt does not — it says what the turn *would be
|
||
told now*, from today's canon, today's sources and today's state, which is the
|
||
opposite of what the inspector is for. **Historical evidence is not a cache**,
|
||
and M9 regenerates no prompt at any point in the import.
|
||
|
||
### Required, optional, and what an absent key means
|
||
|
||
| Key | Required | Absent means |
|
||
| --- | --- | --- |
|
||
| `format` | **yes** | refused |
|
||
| `actions` | in practice | an empty campaign |
|
||
| `branches`, `headBranch` | no | the root branch |
|
||
| `headDepth` | no | **the format could not say — open at the tip** (I07) |
|
||
| `checkpoints` | no | the campaign had none, or predates M4 |
|
||
| `narrativeState`, `narrativeStateAfter` | no | predates M5; empty document |
|
||
| `knowledge` | no | predates M7; empty library |
|
||
| `stateEvents`, `stateProposals` | no | **predates v3 — no audit trail is invented** |
|
||
| `contextSnapshotZ` / `contextSnapshot` | no | predates v3, or the turn has none |
|
||
| `summaries` | no | predates v3 |
|
||
| `id` / `parentId` on a node | no | pre-SP9 rule: `attempts.group` falls back to the coordinate |
|
||
| `authority` on a memory | no | the column default |
|
||
|
||
### Integrity and validation
|
||
|
||
One derived value is in the file, and only because its purpose is to be checked:
|
||
a knowledge source's `contentHash`. The import **recomputes** it from what
|
||
arrived, stores the recomputed value, and records the discrepancy on the source
|
||
where a reader can find it. The stated hash is never trusted and never silently
|
||
discarded.
|
||
|
||
### The two pointers that are translated
|
||
|
||
Branch numbers already were. M9 adds `source_id` inside a restored retrieval
|
||
record: it names a row on the machine that wrote the file, so left alone it
|
||
points the inspector's "open this source" at whatever holds that id here. It is
|
||
repointed at the source that landed, or set to `null` where the file carries no
|
||
such source — at which point the record still holds the passage's title,
|
||
filename and text. **The evidence is never rewritten; only the pointer is.**
|
||
`chunk_id` is deliberately untouched: passages are rebuilt and get new ids, so
|
||
no translation exists, and `chunk_index` and `heading_path` still say which
|
||
passage it was.
|
||
|
||
---
|
||
|
||
## E. Story graph, head, and takes
|
||
|
||
**Round-tripped and compared through the API a reader reads**, never by row
|
||
count. `test_m9_portability.py`, and again across two processes in
|
||
`test_m9_clean_import.py`.
|
||
|
||
| Property | Evidence |
|
||
| --- | --- |
|
||
| Every accepted action, live and superseded | the whole exported tree is the same size either side |
|
||
| Branch and fork relationships | both branches, the fork depth, and the names |
|
||
| Active branch and head | §I07 below |
|
||
| Retained tip | the future past the head is in the database and Redo reaches it |
|
||
| Branch disposition | the superseded branch still carries the depth it was left at (6) |
|
||
| Retries / alternate takes | the superseded take is still at its coordinate; the selected one still selected |
|
||
| Which take is active | exactly one live node per coordinate, decided by the writer, not read from the file |
|
||
| Timestamps | `createdAt` restored per node, per branch, per Save Point, per event |
|
||
|
||
### The undo-head case (I07)
|
||
|
||
Exported with `active head < retained tip`, imported into a genuinely clean data
|
||
directory, and the copy opens **at exactly the exported head**. The later turns
|
||
are present as retained future, Redo is offered rather than the story having
|
||
silently been redone, and Redo then walks to the same next turn in both
|
||
campaigns.
|
||
|
||
### The diverged case (I03)
|
||
|
||
Both futures survive and stay distinguishable — the abandoned line's turns are
|
||
readable on their branch, and the branch that was left still records that it was
|
||
left, at the depth it was left at. There is no trimmed-export option, so I03's
|
||
clause about one does not arise.
|
||
|
||
### The retry case, and a fidelity gap M9 closed
|
||
|
||
v2 exported no take parentage, so every imported node landed parentless and
|
||
`attempts.group` fell back to the coordinate. That is right for a plain retry
|
||
and **wrong as soon as two takes of one turn each have takes of their own
|
||
beneath them**: those share a branch and a depth, so the copy read `5/5` where
|
||
the source read `2/2` and `3/3`. v3 carries the parentage, and the pager now
|
||
reads identically in the copy and the original.
|
||
|
||
`test_retry_variants.py` had recorded the old behaviour as "a gap in the import";
|
||
that test now asserts `2/2` (§U-6).
|
||
|
||
---
|
||
|
||
## F. Save Points
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Name, note, coordinate, `createdAt` | all restored; compared through the panel's own API |
|
||
| Resolution | every restored Save Point reports `resolved: true` |
|
||
| Restore | reaches the position it names, through **M3's head movement** — no second restore architecture |
|
||
| Later history | intact after restoring; the retained tree is the same size |
|
||
| Two Save Points, two branches | restore to **different** states, so they are not accidentally the same pointer |
|
||
| The head | unmoved by importing them: a campaign exported at turn 30 with a Save Point at turn 12 opens at turn 30 |
|
||
|
||
### A Save Point the file cannot satisfy
|
||
|
||
**Dropped, with the rest of the campaign kept — and never retargeted.** The
|
||
decision and its reasoning are in `bundle.py` and `TECHNICAL-DESIGN.md` §9.2:
|
||
|
||
- not a refusal, because a bookmark costs a bookmark, where refusing the
|
||
campaign would lose the story to save the bookmark;
|
||
- not a repair, because the reader named a position, and if that position is not
|
||
in the file then **no other position is the one they named**.
|
||
|
||
Three cases are covered: a coordinate beyond the retained story, a branch the
|
||
file does not list, and a name that is blank once trimmed. In each, the *other*
|
||
Save Point in the fixture survives — asserted, because "dropped" must not mean
|
||
"dropped them all".
|
||
|
||
---
|
||
|
||
## G. Narrative state and the audit trail
|
||
|
||
### What travels, and why each
|
||
|
||
The audit inventory was made from the contract rather than by serialising every
|
||
table:
|
||
|
||
| Record | Travels | Why |
|
||
| --- | --- | --- |
|
||
| `narrative_state` (campaign) | yes | the authoritative document at the exported head |
|
||
| `narrative_state_after` (per node) | yes | **the restore path.** Without it, head movement becomes replay |
|
||
| `state_events` | **yes (new)** | §17's audit half: what changed, where, who asserted it, what it was before |
|
||
| `state_proposals` | **yes (new)** | the inspector reads both — the event says what was accepted, the proposal says what was asked for and refused |
|
||
| rejected/unparseable proposals | yes | the record that explains why the state does not say what the narration seems to say |
|
||
| `derived_status` | no | describes the last run of a background pass, not the story |
|
||
|
||
Exporting only the accepted half would have kept the answers and lost every
|
||
question, which is why both tables travel.
|
||
|
||
### Verified after the move
|
||
|
||
- Current authoritative state at the exported head: **identical**, compared as
|
||
the document.
|
||
- Historical positions: the copy and the source are walked back three turns and
|
||
forward three turns **in step**, and the document is compared at every
|
||
position. They agree throughout (L02).
|
||
- Restoration stays **O(1)-style**: every position's snapshot travels, so Undo,
|
||
Redo and Save Point restore remain one row read rather than a replay
|
||
(`TECHNICAL-DESIGN.md` §10.4, and the M4 note that made it load-bearing).
|
||
- **A manual correction is still identifiable as one.** This is the case the
|
||
omission hurt most: it is the one state change no narration explains, and with
|
||
the events gone nothing distinguished it from something the story established.
|
||
- The two tables arrive **linked**: events point at proposals that are in this
|
||
campaign, and proposals point at this campaign's own turns.
|
||
|
||
### Refuse, repair, drop — decided per field
|
||
|
||
| | Rule |
|
||
| --- | --- |
|
||
| **Refuse** | an audit record naming a turn the file does not contain. Unlike a Save Point, an audit record that quietly did not arrive leaves a campaign whose state cannot be explained — and the explanation is what a reader goes looking for exactly when something looks wrong |
|
||
| **Accept** | a proposal with *no* action: that is a manual correction, which has a coordinate and no narration behind it |
|
||
| **Keep the coordinate, drop the pointer** | an event whose *proposal* is missing. That is what the schema's `ON DELETE SET NULL` already says happens |
|
||
| **Normalise** | a malformed state document — M5's rule, unchanged: the story is the valuable thing, and a malformed section should cost the section |
|
||
|
||
---
|
||
|
||
## H. Historical prompt and retrieval provenance
|
||
|
||
**The M8 handoff (question B) is closed: the evidence travels.**
|
||
|
||
`SPECIFICATION.md` §6.1 requires every accepted turn to record an *exact
|
||
prompt/context snapshot or reproducible equivalent*. The product already stored
|
||
one; the bundle did not carry it, so that guarantee was local to the machine
|
||
that played the campaign. It is now portable.
|
||
|
||
Each node carries `contextSnapshotZ`: the assembled prompt section by section,
|
||
the passages retrieved **with the text each supplied**, which summary was
|
||
eligible, the token accounting, and the model and generation settings the call
|
||
ran under. Restored **verbatim** — nothing is regenerated, and nothing is
|
||
validated beyond its being an object, because imposing today's expectations on a
|
||
record an older build wrote would be the retroactive reading the inspector
|
||
exists to rule out.
|
||
|
||
### Verified to survive each thing that could destroy it
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| A source disabled later | the record is unchanged |
|
||
| **A source deleted after import** | asserted directly: the sources an old turn was shown are deleted in the copy, and the turn still shows the same passages and the same text. The record holds the text, not a pointer to a row that can go away (`IMPORTED-KNOWLEDGE-DESIGN.md` §49-50) |
|
||
| A different chunking on the importing machine | the record is not rebuilt from chunks, so it cannot move |
|
||
| Current state moved on | the snapshot is a record, not a re-render |
|
||
| Canon edited later | same |
|
||
| Summaries rebuilt | same |
|
||
|
||
And across a real machine boundary: on machine B, which never assembled that
|
||
prompt and could not reassemble it, Inspect Context on an old narrator turn
|
||
returns a prompt with its sections intact.
|
||
|
||
### Retrieval provenance
|
||
|
||
Each `used` record keeps source title, filename, classification, visibility,
|
||
chunk index, heading path, the scores each path gave it, how it was admitted,
|
||
and the rendered text. `source_id` is translated (§D); everything else is
|
||
evidence and is untouched.
|
||
|
||
### Cost control
|
||
|
||
The prompt is stored **once per turn on the live attempt** (`attempts.py`), so a
|
||
superseded take carries only its own slices — its reply, its state proposal, its
|
||
token accounting. Asserted: no superseded take in the bundle carries `sections`
|
||
or `prompt`. A campaign retried twenty times does not carry twenty prompts.
|
||
|
||
---
|
||
|
||
## I. Imported knowledge
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Preserved | title, original filename (metadata only), content, classification, enabled, visibility, always-include, notes, media type, import timestamp, content hash, **parser and chunking version**, and a `sourceId` that exists only so historical retrieval records can be relinked |
|
||
| Rebuilt | passages, FTS rows, vectors — before the import returns for the first two, so the campaign is searchable immediately |
|
||
| Needs the original path | **no.** Verified with the exporting machine's process stopped |
|
||
| Disabled source | still disabled, and still out of retrieval |
|
||
| Narrator-only source | still narrator-only |
|
||
| Canon with always-include | still both |
|
||
| Content | byte-identical, by SHA-256, checked in the browser run too |
|
||
|
||
### L04 — the rebuild, and the defect running it found
|
||
|
||
On a **safe copy** (the imported campaign, not the original), every physically
|
||
derived structure is destroyed — passages, FTS rows, vectors — and rebuilt from
|
||
the content the bundle carried. Afterwards: every source `ready` with passages,
|
||
retrieval working, and the transcript, authoritative state, classifications and
|
||
lifecycle flags identical either side. Deleting the vectors alone leaves lexical
|
||
retrieval working, which is M7's rule that lexical is a production path and not
|
||
a fallback. A rebuild does not make an abandoned line's summary eligible.
|
||
|
||
**Running it found finding 1** (§U), a pre-existing defect that made Reindex —
|
||
the documented repair — unable to repair the state it most needed to.
|
||
|
||
### The M7 calibration rule
|
||
|
||
Untouched. `nomic-embed-text` keeps its measured calibration and an uncalibrated
|
||
embedding model still degrades to lexical-only rather than borrowing the 0.58
|
||
threshold; no M9 code path reads or writes the threshold, and the M7 calibration
|
||
tests pass unchanged.
|
||
|
||
---
|
||
|
||
## J. Summaries, memories, and derived data
|
||
|
||
### Summaries — carried, with their lineage
|
||
|
||
v2 carried only `adventures.story_summary`, the **mirror column with no lineage
|
||
of its own**, so a restored campaign resumed with no usable long-story
|
||
continuity. v3 carries the rows: text, coordinate, source range, trigger, model,
|
||
timestamp.
|
||
|
||
Eligibility is not a stored flag — it is whether the coordinate lies on the
|
||
active capped lineage — so restoring the coordinate restores the answer,
|
||
**including the "no"**. The fixture has one generated summary on the line it
|
||
abandons and one typed at the head it keeps, and after the move:
|
||
|
||
- exactly one is eligible, and it is the one at the head;
|
||
- the abandoned line's text does not appear in the prompt the copy would send
|
||
next;
|
||
- the same holds after a full index rebuild, so **M6's leak does not return by
|
||
the rebuild path either**;
|
||
- `story_summary` is re-derived from the lineage after import rather than
|
||
trusted from the file, because a mirror disagreeing with the lineage is
|
||
exactly M6's finding M6-F1.
|
||
|
||
### Memories — and an authority that was being promoted
|
||
|
||
Text, pinned, forgotten, source range, use count, coordinate, and now
|
||
**`authority`**. Without it every imported memory landed on the column default,
|
||
which silently promoted a `heuristic` memory to `accepted_story` — the one
|
||
direction F07 forbids. An unreadable value is read as `heuristic`, which is the
|
||
safe direction: a demoted memory costs a little ranking weight, a promoted one
|
||
puts a guess in front of the narrator as fact.
|
||
|
||
### Vectors
|
||
|
||
Not carried, in either subsystem, and rebuilt against whatever embedding model
|
||
*this* machine has. Exporting them would tie a campaign to one machine's model,
|
||
which is the reason M7 gave and M9 keeps.
|
||
|
||
---
|
||
|
||
## K. Story cards — the compatibility decision
|
||
|
||
**Settled: compatibility-only legacy data, and no longer in the narrator's
|
||
prompt.** Recorded in `IMPORTED-KNOWLEDGE-DESIGN.md` §73.
|
||
|
||
### What was found
|
||
|
||
Story cards **were** the "alternate untracked path around the new knowledge
|
||
authority/provenance rules" that §73 already forbade, and not in principle. A
|
||
keyword-matched card was injected into the narrator's prompt as
|
||
`World Lore: <entry>`, taking up to 40% of what remained after the imported
|
||
knowledge had been placed, with:
|
||
|
||
- no class, so nothing framed how far the narrator could rely on it;
|
||
- no visibility, so no narrator-only distinction existed;
|
||
- no source, no hash, no lifecycle — nothing to disable;
|
||
- **no browser surface at all since M8 removed the editor**;
|
||
- and no row in the context inspector, which renders `knowledge` and never
|
||
rendered `cards`.
|
||
|
||
It also competed with imported Canon for one budget, which is the arrangement M7
|
||
spent a milestone separating.
|
||
|
||
### What happens now
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| **New export** | carries them, unchanged, under `storyCards` |
|
||
| **Legacy import** | accepted unchanged, from every format version |
|
||
| **Normal narration** | **not reached.** The `world_lore` section is gone |
|
||
| **Re-export** | carries them again — a round trip destroys nothing |
|
||
|
||
Nothing is deleted: the rows stay, the `/api/story-cards` endpoints stay, and
|
||
`memorybank.cast_brief` still reads them as the **summariser's character
|
||
roster** — that names who is on stage so a memory says "Aldric" rather than
|
||
"he", never reaches the narrator, and every memory written from it is
|
||
authority-classified by the application afterwards. The `cards` key stays in the
|
||
context report and is now always empty for a new turn, because M9 has just made
|
||
historical snapshots portable and an old turn's record must go on saying that
|
||
story cards were included. That is asserted.
|
||
|
||
This is the **smallest safe change**: it closes the one path that asserted
|
||
campaign facts to the narrator without any of §73's controls, and touches
|
||
nothing else.
|
||
|
||
---
|
||
|
||
## L. Scene and media metadata
|
||
|
||
**No media tables exist in this build, and none were invented.** Two active
|
||
documents say that is the right answer rather than a gap:
|
||
|
||
- **K04's second branch** — "if media tables are deferred, architecture/types
|
||
should demonstrate equivalent extension point".
|
||
- **`MEDIA-EXTENSION-CONTRACT.md` §58**, on export, in as many words: *"For v1,
|
||
if media is not implemented, preserve schema compatibility."*
|
||
|
||
M9 preserves it in the strongest available form: the format is now **versioned**,
|
||
so M10 can add media sections as v4 and every v3 reader keeps working, and every
|
||
v3 file keeps importing. Inventing an empty `media_assets` table to satisfy the
|
||
word "metadata" would have been schema M10 then had to live with.
|
||
|
||
What does exist is the `scene` section of the authoritative narrative state
|
||
document, which is `SPECIFICATION.md` §14's scene snapshot as this build
|
||
represents it. It is campaign-scoped and per-position, and it round-trips with
|
||
the document: asserted both on the campaign-level state and on the per-node
|
||
snapshots.
|
||
|
||
Nothing in M9 adds image, video, TTS, STT, a media provider, or a job
|
||
architecture.
|
||
|
||
---
|
||
|
||
## M. Model and context-window portability
|
||
|
||
The M8 handoff's question D, answered by keeping two things apart.
|
||
|
||
### What travels
|
||
|
||
**Per-turn model and generation settings**, inside the historical snapshot —
|
||
model name, API mode, temperature, max output tokens. There they are a *record
|
||
of what happened*, which is exactly §6.1's requirement.
|
||
|
||
**Campaign-owned configuration**: title, narrator instructions, persona, canon,
|
||
memory/summary switches, the campaign's own plot fields.
|
||
|
||
### What does not, and why that is the point
|
||
|
||
The `settings` row — endpoint, model, context budget, timeouts, embedding model.
|
||
Those describe the **machine**, not the campaign. Asserted directly on the
|
||
importing machine: after importing a campaign played against a configured
|
||
endpoint, machine B's `endpoint_url` is still its own default and its
|
||
`context_token_budget` is still 16,384. **Importing a campaign is not a way to
|
||
reconfigure the destination's inference.**
|
||
|
||
### Import does not depend on a model
|
||
|
||
Asserted: machine B has **no model configured at all**, and the campaign still
|
||
imports, opens, and shows its state and transcript. Playing on would fail;
|
||
recovering does not. *Campaign restored successfully* is not *the destination
|
||
narrator has the same capacity*, and M9 does not treat it as such.
|
||
|
||
### The ceiling, and what M9 does about it
|
||
|
||
No assumption that 16,384 tokens are available is hard-coded anywhere M9
|
||
touched, and no native context-window detection architecture was built — that is
|
||
M11's, with the 100-turn certification.
|
||
|
||
What M9 adds is the sentence a deployer needs, in `DEVELOPMENT.md`: a long
|
||
imported campaign fills a prompt on its **first** turn, so a machine that has
|
||
applied neither the derived-model procedure nor a matching budget meets its
|
||
ceiling immediately rather than gradually. Check `/api/ps` on the destination
|
||
*before* playing on an imported campaign.
|
||
|
||
---
|
||
|
||
## N. SQLite backup
|
||
|
||
`backend/app/backup.py`, `backend/app/routers/backups.py`, and a control in
|
||
Settings → Advanced.
|
||
|
||
### Why not a file copy
|
||
|
||
A copy taken while the application runs is a copy of a moving target: SQLite
|
||
writes in pages, and a plain copy can read page 5 before a transaction and page
|
||
900 after it, producing a file that opens, reports a schema, and is quietly
|
||
missing rows. Nothing warns anyone, and it is found later.
|
||
|
||
So it uses **SQLite's online backup API** (`sqlite3.Connection.backup`), which
|
||
copies under the right locks and yields a transactionally consistent snapshot of
|
||
a committed point. **The application keeps running**; no session is closed and
|
||
no turn is blocked.
|
||
|
||
### The procedure's guarantees, each tested
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Source opened **read-only** | via a `mode=ro` URI, so it is a guarantee and not an observation. The source file is byte-identical after a backup |
|
||
| Temporary destination, finalised by rename | the `.partial` file is gone and the final name never wore it |
|
||
| `PRAGMA quick_check` before the rename | a failed verification leaves **nothing** behind |
|
||
| Never overwrites | two backups get two names; two in the same second get two names |
|
||
| Failure reported | a clear error with the reason; the source untouched |
|
||
| No caller-supplied path | the endpoint accepts **no request body at all** — asserted against the OpenAPI schema, and by sending one anyway and checking where the file landed (H08) |
|
||
|
||
### Opened independently, and read
|
||
|
||
Every backup test **opens the copy as its own database** and asks what is in it —
|
||
a test that only checked a file appeared would pass against a `cp`. Verified
|
||
present in the copy: the schema, campaigns with their titles, story rows, the
|
||
head branch and depth **naming a turn that exists in the copy**, Save Points,
|
||
knowledge sources, state events, and the authoritative state document decoded and
|
||
compared against the live campaign's.
|
||
|
||
### Taken under load
|
||
|
||
The test that separates this from a file copy: turns are played from a second
|
||
thread **throughout** the copy. The resulting file passes `quick_check`, has no
|
||
dangling foreign key, has a transcript with **no gap in its depths**, and a head
|
||
that names a turn present in it. Which committed point it caught is not asserted
|
||
— that would be asserting on a race.
|
||
|
||
### No restore endpoint, and no download
|
||
|
||
Both are decisions, stated in the code and in the panel. Restore means replacing
|
||
the file the running process has open, which is how both copies are lost at
|
||
once; the stop-move-start procedure is in `DEVELOPMENT.md`. Download would put a
|
||
copy of every campaign into the browser's download directory and cache, which
|
||
for an application whose premise is that the story does not leave the machine is
|
||
a worse default than a path the reader can copy.
|
||
|
||
---
|
||
|
||
## O. Import validation and corrupt bundles
|
||
|
||
`test_m9_corrupt_bundles.py`. Every case starts from a **real export of the M9
|
||
fixture** and breaks exactly one thing, so what each test measures is that break
|
||
rather than a hand-written shape nothing ever wrote.
|
||
|
||
### Two rules, checked rather than assumed
|
||
|
||
**Nothing lands.** A refused import is checked by counting rows in **eleven
|
||
tables** before and after, not by trusting the status code. No campaign, no
|
||
branch, no orphan action, no Save Point pointing at nothing, no knowledge owned
|
||
by a campaign that does not exist.
|
||
|
||
**Nothing is fetched, read or run.** A URL in a bundle stays text; a path stays
|
||
text. No test needs a network guard to pass — there is no code path that would
|
||
use one.
|
||
|
||
### Coverage
|
||
|
||
| Class | Cases |
|
||
| --- | --- |
|
||
| Format / version | not an object; empty object; missing `format`; `format` of four wrong types; an unsupported future version, refused **with its own name in the message** |
|
||
| History graph | action on an unlisted branch; branch forking from one listed after it (which is also **how a cycle is made impossible rather than detected**); a branch forking from itself; a fork with no depth; an action with no depth; a negative depth; **two actions claiming one identity**; a turn with no live take; a parent naming a node the file lacks; **a node that is its own parent** |
|
||
| Head | past the retained story (refused, message names where the branch ends); four wrong types; a head branch that is not listed (repaired to the root, and the reasoning for why *that* one is repaired) |
|
||
| Save Points | coordinate beyond the story; branch not listed; blank name; the whole section not a list |
|
||
| State | sections not lists; an event that is not an object; an event with no type; **an event naming a turn the file lacks (refused)**; a proposal naming a turn the file lacks; an event whose proposal is gone (keeps its coordinate); a malformed state document; a malformed per-position snapshot |
|
||
| Knowledge | section not a list; no content; unknown classification; a source that is not an object; unreadable visibility; **content hash mismatch, recomputed and reported**; over the source cap; over the size cap |
|
||
| Provenance | a snapshot that is not an object (dropped, never guessed at); a nonsense knowledge block; a snapshot naming an impossible source (relinked to `null`, evidence intact) |
|
||
| Summaries | no coordinate (dropped, **never placed at a guess** — that is E03's leak by a new route); section not a list |
|
||
| Caps | actions, branches, and a body past the import ceiling (413 from the middleware, on the declared length, before any parse) |
|
||
| Inert data | URLs, `file://`, remote images; four traversal filenames; a title that looks like a command; an over-long title |
|
||
|
||
### The authoritative/derived distinction on failure
|
||
|
||
Two outcomes a caller can see, and they are distinguishable:
|
||
|
||
```text
|
||
4xx the authoritative import failed and no campaign exists
|
||
201 the authoritative import succeeded and the campaign is complete
|
||
```
|
||
|
||
A **rebuildable** index failing is neither, and does not roll the campaign back:
|
||
it is reported on the 201 as `import_warnings`, visible per source in the
|
||
Knowledge panel, and repaired by Reindex. The response model is a subclass
|
||
(`ImportedAdventureOut`) rather than a field on `AdventureOut`, because "which of
|
||
your indexes failed to rebuild" is a fact about one import, not a property of a
|
||
campaign.
|
||
|
||
The rollback is **explicit** rather than left to the session closing, so a test
|
||
can assert on it rather than on teardown.
|
||
|
||
---
|
||
|
||
## P. Migration and backward compatibility
|
||
|
||
### No schema change, proved
|
||
|
||
```
|
||
git diff 1ce9972 -- backend/app/migrations.py → 0 lines
|
||
mapped_column additions/removals in models.py → 0
|
||
PRAGMA user_version → 92 before, 92 after
|
||
```
|
||
|
||
`git diff` is necessary and not sufficient — a migration can also be *missing*,
|
||
and the failure then is a database that opens and quietly answers wrongly. So
|
||
the claim is proved the way M8 proved its own:
|
||
|
||
**A campaign is built by a server running the signed M8 commit**, from the git
|
||
worktree at `1ce9972`, using M8's own interpreter and M8's own code: story, an
|
||
alternate take, a Save Point, a manual state correction, memories, a summary,
|
||
three imported sources including a disabled and a narrator-only one, per-turn
|
||
context snapshots, and two Undos so the head is behind the tip. M9 then opens
|
||
that file.
|
||
|
||
```
|
||
$ python -m tools.m9_migration_proof open <db> --expected <m8-report.json>
|
||
{ "problems": [],
|
||
"schema_version_before": 92, "schema_version_after": 92,
|
||
"schema_version_second_open": 92,
|
||
"checked": { "families": 12, "snapshot_rows": 8,
|
||
"counts": { "actions": 16, "checkpoints": 1,
|
||
"knowledge_chunks": 3, "knowledge_sources": 3,
|
||
"memories": 2, "state_events": 9,
|
||
"state_proposals": 9, "summaries": 1 } } }
|
||
```
|
||
|
||
Twelve families compared and identical; **a second open migrates nothing**
|
||
(idempotence); and the M8-built campaign then exports as v3 and round-trips, and
|
||
Redo still works on it. Neither half of the tool imports the other's code — what
|
||
crosses is the database file and one JSON report.
|
||
|
||
### Older bundle seams
|
||
|
||
`test_m9_legacy_bundles.py` ages a real v3 export backwards, removing only what
|
||
each era's format genuinely could not carry — a checked-in fixture drifts, and a
|
||
hand-written one tests a shape nothing ever wrote.
|
||
|
||
| Era | No story lost | Nothing invented | Head |
|
||
| --- | --- | --- | --- |
|
||
| pre-M9 (v2) | ✓ | no events, no proposals, no summaries, no prompts | honoured, as v2 carried it |
|
||
| pre-M7 | ✓ | + no knowledge sources, no chunks | honoured |
|
||
| pre-M5 | ✓ | + empty state, no canon | honoured |
|
||
| pre-Save-Point | ✓ | + no Save Points | honoured |
|
||
| pre-active-head | ✓ | + no Save Points | **opens at the tip**, offers no Redo |
|
||
| v1 (flat) | ✓ | both attempts arrive, one live | tip |
|
||
| pre-M2 (with `scripts`) | ✓ | scripting keys ignored, not rejected | as its era |
|
||
|
||
And each era can still be *used*, not merely read: a pre-M5 campaign gains state
|
||
from its next turn, a pre-M7 campaign imports a source and it indexes, a
|
||
pre-Save-Point campaign can be given one.
|
||
|
||
**A pre-M9 campaign re-exports as v3 with its evidence sections empty**, because
|
||
the campaign genuinely has none — which is what lets a later reader trust a v3
|
||
file's empty `stateEvents` to mean "this campaign has no audit trail" rather
|
||
than "the file could not say".
|
||
|
||
---
|
||
|
||
## Q. Browser recovery workflow
|
||
|
||
**36/36 checks, zero failures**, on one frozen production build against a real
|
||
narrator across two installations.
|
||
|
||
### Environment
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Browser | Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1) |
|
||
| Application | the **frozen production build** `dist/assets/index-DcHbz7ga.js`, served by the real backend as static files — not a dev server |
|
||
| Backend | two `uvicorn app.main:app` processes on loopback, one per machine, each with its **own data directory** |
|
||
| Databases | machine A: a real campaign. Machine B: **a file that had never existed**, created by the migrations on first start |
|
||
| Narrator | `qwen2.5:3b-instruct-16k` on the trusted-LAN Ollama, over HTTPS with a privately issued certificate — **real inference for every turn** |
|
||
| Campaign | played over real HTTP: five turns, a retry, a Save Point, three imported sources (one disabled, one narrator-only), a manual state correction, and two Undos so the export is taken behind the retained tip |
|
||
|
||
### The §25 workflow, item by item
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| **1. Export campaign** | **PASS.** Clicked in the browser; the page produced a **60,160-byte `ai-dnd-adventure-v3`** bundle |
|
||
| **2. Return to Campaigns** | **PASS** |
|
||
| **3. Import into a clean target** | **PASS.** Machine B had never had a database, its library was empty, and it had **no narrator configured at all**. The import reported no incomplete rebuild |
|
||
| **4. Open imported campaign** | **PASS** |
|
||
| **5. Opens at the exact active position** | **PASS.** The transcript matches the source exactly; Redo is offered; the retained future is in the database (**11 rows behind 7 on screen**); and all 4 narrated turns are found in the rendered page |
|
||
| **6. State is correct** | **PASS.** The authoritative document matches, the manual correction is in the restored audit trail, with the fact the reader asserted, and the whole trail arrived |
|
||
| **7. Save Points present** | **PASS**, in the browser's own panel, with the note |
|
||
| **8. Knowledge present, with classifications and status** | **PASS.** All three sources in the panel; the disabled one still disabled, the narrator-only one still narrator-only, every source reindexed and searchable, and the content **byte-identical by hash** |
|
||
| **9. Historical Inspect Context works** | **PASS.** 3 of 3 restored narrator turns show the prompt they were given; 3 name the passages they were shown; and the inspector opens through the browser's own control |
|
||
| **10. Undo/Redo coherent** | **PASS.** Both offered at the restored head; Redo moves forward into the retained future |
|
||
| **Backup control** | **PASS.** Taken from Settings; the panel names the file; the file is on disk where it says; it **passes its own integrity check opened alone**, and holds the restored campaign with all 11 of its rows |
|
||
| **I06 on the downloaded file** | **PASS.** No credential, no endpoint, no path from the exporting machine |
|
||
|
||
### Two substitutions, and both are on the browser's side of the line
|
||
|
||
Named precisely, because §30 of the brief is about not letting a substitute
|
||
stand in for the boundary it bypasses.
|
||
|
||
1. **The export's write to disk.** `downloadJSON` builds a Blob, makes a `blob:`
|
||
URL and clicks an anchor. **Headless Firefox does not complete that download
|
||
in this configuration** — the click fires, the page reports "Campaign
|
||
exported.", and no file appears in the download directory, in `~/Downloads`,
|
||
or anywhere under the snap's own tree. So `URL.createObjectURL` was wrapped
|
||
to capture the exact bytes the page handed the browser, and those bytes were
|
||
written to disk and used for everything downstream. *Proved:* the browser
|
||
assembled the correct v3 file and offered it. *Not proved:* Firefox's own
|
||
download writer, which is browser behaviour.
|
||
2. **The import's read of the file.** Firefox here is a **snap**, and its
|
||
sandbox will not let the content process read a path outside its
|
||
confinement. WebDriver hands the input the file successfully — the `File`
|
||
arrives with the right name and **the right byte count**, and `change` fires
|
||
— and `FileReader` then fails with `NotFoundError`. The click and the
|
||
handover are exercised; the last hop, the POST the page would make with the
|
||
bytes it read, was made against the same endpoint the page calls.
|
||
|
||
Neither is papered over, and neither leaves the end-to-end path unproven:
|
||
`test_m9_clean_import.py` does the whole of it — export, a real file on disk,
|
||
a real HTTP import into a second process against a database that never existed
|
||
— with no browser in the way, 11/11.
|
||
|
||
There is no unconfined Firefox and no Xvfb on this machine, so a windowed run
|
||
was not available. That is a property of the machine and is recorded as such.
|
||
|
||
### One observation worth a reviewer's attention
|
||
|
||
**The real narrator emitted no typed state blocks across five turns.**
|
||
`qwen2.5:3b-instruct-16k` narrated well and never produced a `state` block, so
|
||
the campaign's only state event is the **manual correction** — which is why the
|
||
audit-trail checks above read "1 events, source had 1".
|
||
|
||
That is not an M9 defect and not a regression: it is the behaviour ADR 003 and
|
||
ADR 010 exist for. The application owns state and the model does not, C04's
|
||
manual correction path is how a reader asserts what the model did not, and the
|
||
per-position snapshots keep the head, the transcript and the state in agreement
|
||
regardless. It does mean the browser run's state evidence is thinner than the
|
||
suites', where a scripted narrator produces twelve events and two proposals —
|
||
so the rich state case is proved there and the *real-model* case is proved here,
|
||
and this report does not let the second stand in for the first.
|
||
|
||
---
|
||
|
||
## R. Performance and size
|
||
|
||
### The M9 fixture
|
||
|
||
| | M8 (v2) | M9 (v3) |
|
||
| --- | --- | --- |
|
||
| Bundle | 29,630 B | 90,117 B |
|
||
| Actions | 22 | 22 |
|
||
| Source database | 258 kB | 319 kB |
|
||
| Export | 0.028 s | 0.044 s |
|
||
| Import (write + rebuild) | 0.063 s | 0.119 s |
|
||
|
||
### A long campaign — and the result that changed the plan
|
||
|
||
`python -m tools.m9_scale_report --turns 150 --budget 16384`. A real prompt
|
||
builder, a 16,384-token budget, and prose long enough that a turn is a turn.
|
||
|
||
```
|
||
turns actions bundle B snapshots B B/turn export s import s
|
||
25 51 311,346 130,766 5,230 0.066 0.171
|
||
50 101 883,616 279,208 5,584 0.129 0.335
|
||
75 151 1,715,306 439,634 5,861 0.269 0.232
|
||
100 201 2,870,121 675,880 6,758 0.572 0.371
|
||
125 251 4,313,937 951,870 7,614 0.518 0.603
|
||
150 301 6,022,430 1,242,324 8,282 1.012 1.109
|
||
```
|
||
|
||
A per-turn prompt contains the story so far, so carrying one per turn is
|
||
**O(turns²)**. That was the risk in the decision, and measuring it changed what
|
||
was done about it and what is claimed:
|
||
|
||
- **Uncompressed, the snapshots were 68% of a 9.7 MB file at 120 turns.** They
|
||
now travel compressed inside the bundle — the same JSON through the same
|
||
`compression.pack`/`unpack` the database column already uses, base64-encoded
|
||
so the file is still JSON. **7.9× smaller**, and every other section of the
|
||
file is still plain readable text.
|
||
- At M11's 100-turn certification target: **2.87 MB, 14% of the cap.**
|
||
- Export and import stay under 1.1 s at 150 turns.
|
||
|
||
### Where the bytes are, and what M9 actually cost
|
||
|
||
The same tool's per-section breakdown, at 150 turns. It exists because "the
|
||
prompts are 21% of the file" does not answer "what did M9 add" — the events, the
|
||
proposals and the summaries are v3 additions too, and a cost claim counting only
|
||
the prompts would understate it.
|
||
|
||
```
|
||
per-position state (v2 already) 4,474,294 74.3%
|
||
prompts (contextSnapshotZ) 1,245,144 20.7%
|
||
state proposals 78,938 1.3%
|
||
state events 41,870 0.7%
|
||
node ids + parentage 6,993 0.1%
|
||
summaries 3,958 0.1%
|
||
--- everything v3 added 1,376,903 22.9%
|
||
a v2 file of the same campaign 4,645,347
|
||
ceiling with v3 additions: ~279 turns
|
||
ceiling without them: ~318 turns
|
||
```
|
||
|
||
**The result is not the one the decision was worried about.** The largest thing
|
||
in a campaign bundle is not the prompts M9 added — it is the **per-position
|
||
narrative state document, at 74% of the file, which v2 already carried**. It
|
||
grows because a state document accumulates facts, so it is quadratic for the
|
||
same reason the prompts are, and it has been since M5.
|
||
|
||
Everything M9 added comes to **22.9%** of the file, and moves the import ceiling
|
||
from about **318 turns to about 279** — a cost of roughly **12% of reachable
|
||
campaign length** for closing the provenance gap, not the halving it looked like
|
||
before it was measured, and not the 11% a prompts-only reading would have
|
||
claimed.
|
||
|
||
Export and import stay under 1.2 s at 150 turns.
|
||
|
||
### N+1 and duplication
|
||
|
||
Checked for, and none found:
|
||
|
||
- The export undefers every deferred per-node column in **one** query, including
|
||
`context_snapshot`, and the proposals' `detail` once for the whole list.
|
||
- Nodes are referenced by their own exported ids, so the audit records need no
|
||
per-row lookup.
|
||
- Knowledge content appears **once** per source; the prompt appears **once** per
|
||
turn, on the live attempt, with superseded takes carrying only their own
|
||
slices (asserted).
|
||
- The import's flushes are bounded by the number of **knowledge sources**, not
|
||
by the size of the story: one after the nodes, one after the proposals, two in
|
||
`materialize`, and one per source — the last because `build_index` needs the
|
||
source's id, which is the same one-flush-per-source the live upload path
|
||
already does. Nothing flushes per action, per branch, per event or per
|
||
Save Point, which is where a long campaign would have felt it.
|
||
|
||
### Residual
|
||
|
||
The ceiling above is a **residual limit, stated with its measurement**. It is far
|
||
beyond M11's certification target, and the fix if a later milestone needs one is
|
||
a streaming or chunked import, which is architecture beyond M9. The asymmetry
|
||
worth naming for a reviewer: a campaign past that length can still be
|
||
**exported** and would be refused on **import** with a 413.
|
||
|
||
---
|
||
|
||
## S. Acceptance matrix
|
||
|
||
Every M9-owned item, with the evidence that decides it. Where a claim names a
|
||
boundary, the row says which boundary was actually crossed.
|
||
|
||
### I-series
|
||
|
||
| | Result | Evidence, and the boundary |
|
||
| --- | --- | --- |
|
||
| **I01** Export campaign | **PASS** | Every section present and non-empty on the M9 fixture, checked family by family rather than by file size. `test_i01_*`; reproducible with `tools/m9_portability_report` |
|
||
| **I02** Import exported campaign | **PASS** | **Second process, second directory, database file that never existed, exporting process stopped.** Transcript, state, canon, Save Points, knowledge, events and retained-tree size all match. `test_m9_clean_import.py`. Same round trip in-process in `test_m9_portability.py`, labelled there as the weaker of the two |
|
||
| **I03** Branch / disposable history | **PASS** | Both futures present and readable; the abandoned branch still records the depth it was left at; the superseded take still at its coordinate. No trimmed-export option exists, so that clause does not arise |
|
||
| **I04** Checkpoint export | **PASS** | Names, notes, coordinates; every one resolves; each restores to its own position and its own state; later history intact. A Save Point the file cannot satisfy is dropped, never retargeted |
|
||
| **I05** Knowledge provenance | **PASS** | Content byte-identical by SHA-256, classification, enabled, visibility, always-include, filename, parser/chunking versions. Verified **with the exporting machine stopped**, and in the browser |
|
||
| **I06** No API secrets | **PASS** | Searched against the file's **text**, with the inert `api_key` column written first so absence is evidence. Compressed snapshots decoded, not skipped. Checked in the unit suite, the two-process run, and the browser run |
|
||
| **I07** Undone active head | **PASS** | All three clauses. The exact head **across a machine boundary**; a head-less file opens at its tip with no Redo; a head past the story is refused with a message naming where the branch ends |
|
||
|
||
### L-series
|
||
|
||
| | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| **L01** Atomic turn commit | **PASS (not regressed)** | M9 changed no turn-commit path; the inherited tests pass unchanged. M9 adds the import's own version, in two forms: a **refused** import writes zero rows across eleven tables, checked by counting; and a failure **deep inside the write phase** — past every planner check, with the branches, nodes, parentage, memories, head and Save Points already in the session — also leaves zero rows, which proves the transaction rather than the planner. The rollback is explicit, so the next import succeeds rather than inheriting a poisoned session |
|
||
| **L02** State reconstruction | **PASS** | Measured **after a round trip**: copy and source walked back three turns and forward three turns in step, document compared at every position. Restoration stays a snapshot read, not a replay |
|
||
| **L03** Save Point after restart | **PASS** | **Three processes.** Export from A, import into a clean B, restore a Save Point in B, kill B, start a third against the same file, ask again. Transcript, state and retained-tree size all match |
|
||
| **L04** Derived data rebuilt | **PASS** | Passages, FTS rows and vectors destroyed and rebuilt on a copy; authoritative campaign identical either side; retrieval works again; vectors alone gone still leaves lexical working. **Running it found finding 1** |
|
||
|
||
### M9-specific round trips
|
||
|
||
| | Result |
|
||
| --- | --- |
|
||
| Undone-head round trip | **PASS** (I07) |
|
||
| Branched campaign round trip | **PASS** (I03) |
|
||
| Checkpoint round trip | **PASS** (I04) |
|
||
| Knowledge provenance round trip | **PASS** (I05) |
|
||
| Derived indexes can be rebuilt | **PASS** (L04) |
|
||
| **Historical prompt/context round trip** | **PASS** — survives the source being deleted, in the copy |
|
||
| **State audit round trip** | **PASS** — events, proposals, and the manual correction still identifiable |
|
||
| **Summary lineage round trip** | **PASS** — the abandoned line's summary is still ineligible, before and after a rebuild |
|
||
| **Take parentage round trip** | **PASS** — the pager reads identically in copy and source |
|
||
| **Memory authority round trip** | **PASS** — a heuristic memory is not promoted |
|
||
| **Legacy round trips** | **PASS** — v1, v2, and four earlier seams; nothing invented at any of them |
|
||
| **Migration from an M8-built database** | **PASS** — 12 families, 0 problems, version 92→92→92 |
|
||
|
||
### Adjacent items re-verified rather than assumed
|
||
|
||
| | Result |
|
||
| --- | --- |
|
||
| **E03** abandoned summary cannot leak | **PASS** after a move, and after an index rebuild |
|
||
| **F07** heuristic memory is not canon | **PASS** — the authority survives the move |
|
||
| **C04** manual correction | **PASS** — still identifiable as one in the copy |
|
||
| **H08** path traversal | **PASS** — four shapes sanitised; the backup endpoints accept no path |
|
||
| **H09** ZIP slip | **N/A** — no archive format was introduced; the bundle is one JSON document |
|
||
| **G04** disabled source stays out of retrieval | **PASS** after the move |
|
||
|
||
### Synthetic substitutes, labelled
|
||
|
||
| Where | Substitute | Why it is sound |
|
||
| --- | --- | --- |
|
||
| Every backend suite | `ScriptedProvider` narrator, stub embedder/summariser | The property under test is the round trip, not the model. Both derived factories are stubbed, per M6's finding M6-F3 |
|
||
| `test_m9_clean_import.py` | the spawned server's deterministic narrator | Real processes, real HTTP, real files; only the model text is canned |
|
||
| Legacy-bundle seams | v3 exports aged backwards | A checked-in fixture drifts and a hand-written one tests a shape nothing wrote |
|
||
| **Browser run: the export's write to disk** (§Q) | the bytes the page produced, captured from `URL.createObjectURL` | Headless Firefox does not complete a `blob:` download here. Proves the browser assembled and offered the right file; does **not** prove Firefox's download writer |
|
||
| **Browser run: the import's read of the file** (§Q) | the POST the page would make, made against the same endpoint | Firefox is snap-confined and cannot read a WebDriver-uploaded path — the `File` arrives with the right size and `change` fires, then `FileReader` raises `NotFoundError`. The click and handover **are** exercised |
|
||
| **Not substituted** | the browser run's narrator and its two installations, the migration proof's M8 build, the backup's SQLite, every process boundary, the file-on-disk import in `test_m9_clean_import.py` | Each is the boundary the claim names. The end-to-end file path the browser could not complete is proved there, with no browser in the way |
|
||
|
||
---
|
||
|
||
## T. Security and local-only
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Internet | none. No M9 code path opens a socket. Export, import and backup are file and database operations |
|
||
| Embedded URLs | inert. A URL in a bundle is stored and displayed as text |
|
||
| Remote images | unchanged; CSP still `img-src 'self' data:` |
|
||
| Shell | none. No `subprocess` in any M9 application file |
|
||
| Imported content executed | never. It is stored and rendered through M8's safe Markdown path |
|
||
| Loopback | unchanged. The backup endpoints are on the same loopback API |
|
||
| Inference endpoint policy | untouched. `endpoints.py` unchanged |
|
||
| Cloud storage / remote backup | none. The backup writes to a directory beside the database |
|
||
|
||
### I06
|
||
|
||
Tested against the file's **text**, not against a list of columns — a field
|
||
added to a model the exporter walks would otherwise reach the bundle with no
|
||
column test noticing. The inert `api_key` column is written with a recognisable
|
||
value first, so its absence is evidence rather than a tautology. Searched for
|
||
and absent: the written secret, `api_key`, `apiKey`, the endpoint and its port,
|
||
and absolute filesystem paths. The compressed snapshots are **decoded** before
|
||
searching rather than skipped.
|
||
|
||
Checked in three places: the unit suite, the two-process run (against a real
|
||
endpoint URL), and the browser run. Measured on the M9 fixture, over 90,074
|
||
bytes of bundle **plus 154,279 bytes of decoded snapshot**:
|
||
|
||
```
|
||
'SECRET-MUST-NOT-TRAVEL' present: False (written into api_key first)
|
||
'api_key' / 'apiKey' present: False
|
||
the endpoint host and port present: False
|
||
'endpoint_url' present: False
|
||
'/home/' , '/tmp/' present: False
|
||
'secret.key' present: False
|
||
```
|
||
|
||
What **is** present, and correctly: the per-turn model name inside each
|
||
historical snapshot. That is a record of what the turn ran under, which §6.1
|
||
requires — not a configuration the import applies. §M has the distinction.
|
||
|
||
### File, path and archive safety
|
||
|
||
**No archive format was introduced**, so H09 does not arise: the bundle is still
|
||
one JSON document, and the compression in §R is an encoding of one field inside
|
||
it, not a container.
|
||
|
||
- No bundle-supplied value becomes a read or write path. `originalFilename` is
|
||
metadata; four traversal shapes are sanitised on import and asserted absent.
|
||
- Knowledge restores from the content in the bundle, never by reopening the
|
||
source machine's path — verified with the exporting process stopped.
|
||
- The backup endpoints accept no path at all.
|
||
|
||
---
|
||
|
||
## U. Findings
|
||
|
||
Numbered, including those fixed before the final tree.
|
||
|
||
### 1. Deleting a campaign leaked its lexical index, and broke the next import — PRE-EXISTING (M7), FIXED
|
||
|
||
**Severity: high.** An ordinary upload in an unrelated campaign failed with a
|
||
500.
|
||
|
||
The FTS5 index is a virtual table, so no foreign key reaches it and no
|
||
`ON DELETE CASCADE` covers it. Deleting a campaign cascaded
|
||
`knowledge_sources` → `knowledge_chunks` and stopped there, leaving one index row
|
||
per passage pointing at a chunk that no longer existed. Nothing read them — every
|
||
search joins through `knowledge_chunks` — so the leak was **invisible** until
|
||
SQLite handed the freed primary key out again, at which point the next source
|
||
imported into **any** campaign collided on `INSERT` and raised.
|
||
|
||
Worse, `clear_index` finds index rows *through* the chunks, so with the chunks
|
||
gone the orphans were unreachable: **Reindex, the documented repair, could not
|
||
repair it.**
|
||
|
||
Found by **running** L04 — it failed on the rebuild with an integrity error,
|
||
which is the only way this shows itself. Fixed at both ends: `importer.clear_campaign_index` removes
|
||
a campaign's index rows before it is deleted, and `fts.add` uses `INSERT OR
|
||
REPLACE` — the rowid is a chunk's primary key, so a row already there is by
|
||
definition stale. The second half means **a database already carrying the leak
|
||
repairs itself, with no migration**. Three regression tests.
|
||
|
||
### 2. An imported node with no state snapshot was given the head's state — PRE-EXISTING, FIXED
|
||
|
||
**Severity: medium.** M5's review finding 3, arriving through the import.
|
||
|
||
`tree.stamp_outcome` runs on every flush and fills a missing
|
||
`narrative_state_after` from *the campaign's current state*. On an import that is
|
||
the state at the exported head — so every position whose snapshot the file did
|
||
not carry came back holding the newest position's state, and an Undo to turn 2
|
||
showed what the story knew at turn 20.
|
||
|
||
`_write_nodes` now writes the empty document explicitly when the file carries
|
||
none. Legacy behaviour is unchanged: a pre-M5 bundle has no `narrativeState`
|
||
either, so the fallback was already writing `empty()` for those files.
|
||
|
||
### 3. The snapshot relink did not persist — FIXED
|
||
|
||
**Severity: high, and it would have shipped looking correct.**
|
||
|
||
`_relink_snapshots` edited the dict the attribute already held and assigned it
|
||
back. `context_snapshot` is a plain `CompressedJSON` column, not a
|
||
`MutableDict`, so SQLAlchemy tracks it by assignment: at flush the loaded value
|
||
and the current value were the same object, the history reported no change, and
|
||
no `UPDATE` was emitted. The import looked right in memory and wrote the
|
||
untranslated ids to disk.
|
||
|
||
Found because the test asserted the **outcome** — that the restored records name
|
||
this campaign's sources — rather than that the function was called. It now builds
|
||
a new document.
|
||
|
||
### 4. Cancelling the import file dialog hung the Import button — PRE-EXISTING, FIXED
|
||
|
||
**Severity: low, and user-visible.** `pickJSONFile` resolved on `change` only, so
|
||
closing the picker without choosing anything never settled the promise: the
|
||
Campaigns screen's `await` never returned, its `finally` never ran, and the
|
||
Import button stayed disabled reading "Importing…" until the page was reloaded.
|
||
|
||
Found while making the input reachable for the browser suite. The screen's own
|
||
comment already said "a cancelled file picker is not a failure worth a message" —
|
||
it had simply never received one. Now `oncancel` rejects with an empty message,
|
||
which is exactly what that comment describes.
|
||
|
||
### 5. The file input was detached from the document — FIXED
|
||
|
||
**Severity: none for a user; structural for testing.** `pickJSONFile` created an
|
||
input and clicked it without appending it, so no element existed for a test or
|
||
for WebDriver to hand a path to — the import workflow could only ever be checked
|
||
by calling the API underneath it, which is not the workflow. It is now in the
|
||
document and removed however the promise settles. Six tests.
|
||
|
||
**And a second lesson from fixing it.** The first attempt marked the input
|
||
`hidden`, which reads correctly and is wrong: a `hidden` element is
|
||
non-interactable, and WebDriver will set `files` on one **without dispatching
|
||
`change`** — the file lands and nothing happens, which is a worse failure than
|
||
the detached input because it looks like it worked. It now uses the ordinary
|
||
visually-hidden pattern — off-screen, zero-sized, `aria-hidden`, out of the tab
|
||
order — so no reader meets a stray "Choose file" control while the browser's own
|
||
dialog is what they are looking at. The test asserts `hidden === false` and says
|
||
why, so the next person does not "tidy" it back.
|
||
|
||
### 6. Five existing tests pinned the gaps M9 closes — REWORKED, NOT DELETED
|
||
|
||
Each asserted the M8 behaviour that was the debt. Following the rule
|
||
`BUILD-MILESTONES.md` records from M2 — move the instrumentation, do not delete
|
||
the test — each now asserts the new guarantee, and each carries a note saying
|
||
what it used to assert and why it changed.
|
||
|
||
| Test | Was | Now |
|
||
| --- | --- | --- |
|
||
| `test_historical_prompt_evidence_survives_an_export_round_trip` | the copy's turn has **no** snapshot (404) | it has the same snapshot, and its `source_id`s name this campaign |
|
||
| `test_export_and_import_round_trips_variants` | the pager reads `1/1` on the copy, "a gap in the import" | it reads `2/2`, as the source does |
|
||
| `test_an_unknown_format_is_refused` | used `v3` as its "from the future" placeholder — which M9 made real, so it started importing the file it meant to reject | uses `v99`, and asserts against `bundle.FORMAT` rather than a literal |
|
||
| `test_export_carries_the_whole_story` | `format == "ai-dnd-adventure-v2"` | `format == bundle.FORMAT` |
|
||
| `test_the_shipped_file_is_a_bundle_this_build_can_import` | `version == bundle.FORMAT` | `version in bundle.READABLE` — the shipped starter is a v2 file and was **not** regenerated, because rewriting a shipped asset to keep a test's equality holding is changing the evidence to fit the test |
|
||
|
||
### 7. The scale tool measured the wrong thing after compression — FIXED IN THE TOOL
|
||
|
||
It looked for the plain `contextSnapshot` key and reported 0% once the export
|
||
started writing `contextSnapshotZ` — i.e. it would have reported the evidence as
|
||
*omitted* when it was merely encoded, which is precisely the mistake the
|
||
portability report exists to avoid making about anything. Both tools now decode.
|
||
Recorded because a reviewer reading an intermediate number in this report's
|
||
history would otherwise be misled.
|
||
|
||
---
|
||
|
||
## V. Planning changes
|
||
|
||
Distinguished by kind, as the brief asks.
|
||
|
||
### Requirement corrections: none
|
||
|
||
M9 altered no product requirement. `SPECIFICATION.md` and
|
||
`SECURITY-THREAT-MODEL.md` are **unchanged**: §16 already required the export to
|
||
preserve the exact active position, §6.1 already required an exact prompt/context
|
||
snapshot per turn, and M9 implements both rather than redefining either. No
|
||
acceptance condition was weakened.
|
||
|
||
### Newly settled design decisions
|
||
|
||
| Document | Decision |
|
||
| --- | --- |
|
||
| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **Story Cards are compatibility-only legacy data and no longer enter the narrator's prompt.** The brief asked for this to be decided; §K has the evidence that they were the untracked path §73 already forbade |
|
||
| `DATA-MODEL.md` §29, `TECHNICAL-DESIGN.md` §9.3 | **The bundle carries historical evidence**, and the two-category rule becomes three |
|
||
| `DATA-MODEL.md` §29, `TECHNICAL-DESIGN.md` §9.3 | **The format is versioned rather than extended**, and why §9.1's and §9.2's reasoning does not stretch to cover this addition |
|
||
|
||
### Implementation facts recorded
|
||
|
||
`TECHNICAL-DESIGN.md` §9.3 (the v3 contract, the encoding, the two-phase
|
||
transaction) and **new §9.4** (the backup); `DATA-MODEL.md` §29 (the format, the
|
||
three categories, required vs optional, the pointers translated) and **§31**,
|
||
which listed summaries as derived-and-rebuildable and now says what "where
|
||
practical" excludes, and where stored prompts sit — a reader landing on §31
|
||
alone would otherwise conclude the export should be regenerating both;
|
||
`V1-ACCEPTANCE-TESTS.md` I01-I07 and L02-L04 results — with I05's M7-era limit
|
||
marked closed and **the original paragraph kept**, because the M9 decision is
|
||
only legible against it; `BUILD-MILESTONES.md` M9 status; `README.md` and
|
||
`VERSION.md`; `DEVELOPMENT.md` (a new section on the two recovery tools, taking a
|
||
backup, and the stop-move-start restore, plus the imported-campaign context note).
|
||
|
||
M8's report moved to `archive/milestone-reports/`, by the rotation convention
|
||
`planning/README.md` states.
|
||
|
||
---
|
||
|
||
## W. Residual risks and the M10-M11 handoff
|
||
|
||
### Residual risks
|
||
|
||
1. **A long campaign's bundle has a measured ceiling.** **~279 turns** against
|
||
the 20 MB import limit; a longer campaign can still be exported and would be
|
||
refused on import, which is the asymmetry worth naming. Far beyond M11's
|
||
100-turn target, which is 14% of the cap. M9's own additions account for
|
||
**12%** of that ceiling — the other 88% is the per-position state document v2
|
||
already carried, so lifting the ceiling means addressing *that*, not the
|
||
evidence. *Risk: low for v1; the fix is a streaming or chunked import, which
|
||
is architecture beyond M9.*
|
||
2. **`quick_check` rather than `integrity_check`** on a backup. It does the
|
||
structural work without the full index cross-check; a corrupt index that
|
||
`quick_check` misses would ride into the copy. *Risk: low, and the trade is
|
||
stated in the code — a backup verified too slowly to be taken is worse.*
|
||
3. **The backup is not offered on a schedule, and nothing prompts for one.**
|
||
*Risk: low, and outside M9's scope. `DEVELOPMENT.md` shows the `curl` for
|
||
`cron`.*
|
||
4. **`chunk_id` in a restored snapshot is stale.** It is a React key and a data
|
||
attribute, not a live pointer, and `chunk_index` plus `heading_path` still
|
||
identify the passage. *Risk: very low; documented in the code.*
|
||
5. **The importing machine's context window may differ from the source's.**
|
||
Documented rather than detected; detection is M11's.
|
||
6. **This machine cannot drive a file into or out of the browser.** Firefox here
|
||
is a snap, so its sandbox blocks reading a WebDriver-uploaded path, and its
|
||
headless download manager does not complete a `blob:` download. §Q labels
|
||
both substitutions and the end-to-end file path is proved without a browser
|
||
in `test_m9_clean_import.py`. *Risk: low for the product, real for future
|
||
evidence — M11 will meet the same wall on any file-based browser claim, and
|
||
the cheapest fix is an unconfined Firefox or Xvfb on the certifying machine.*
|
||
|
||
### Carried to M10
|
||
|
||
- Media tables do not exist. §L records that as **not applicable** rather than
|
||
inventing schema, and M10 owns the coordinator, the providers and the jobs.
|
||
- The `scene` section of the state document is the extension point that exists
|
||
today, and it round-trips.
|
||
|
||
### Carried to M11
|
||
|
||
- **The context window**, end to end: detection, the 100-turn certification, and
|
||
whether a Settings warning that reads the real window belongs in v1.
|
||
- **The streaming import**, if the bundle ceiling above is judged too low.
|
||
- Contrast and focus measurement, tablet tuning, and the broader WCAG audit
|
||
(inherited from M8's §U, untouched by M9).
|
||
- **An unconfined browser on the certifying machine**, so that a file-based
|
||
browser claim can be made end to end rather than in two labelled halves
|
||
(residual 6).
|
||
|
||
### Recorded here but not M9's: four hands-on playtest findings
|
||
|
||
A play session against **accepted, signed M8** — after M8 was accepted and
|
||
before M9 closeout — surfaced four product-quality observations. **None is an M9
|
||
defect, none was caused by M9, and none blocks M9 acceptance.** They are written
|
||
up in full in **§Y**, and durably in `BUILD-MILESTONES.md` under M11 so they
|
||
survive this report's archival:
|
||
|
||
| | Owner |
|
||
| --- | --- |
|
||
| A. The browser tab still reads `AI D&D` | M11 release polish |
|
||
| B. After Undo, the reader cannot tell where they are | M11 UX/release polish |
|
||
| C. The narration-length setting has no measurable effect | M11 realistic-model behaviour |
|
||
| D. Character identity / coreference confusion (root cause **unknown**) | M11 realistic-model / context diagnostic |
|
||
|
||
D is the significant one, and its evidence is gone — the disposable playtest
|
||
database was destroyed, so no root cause is claimed. §Y records the reproduction
|
||
and classification plan, and one structural fact worth checking first: the
|
||
narrative state permits two entities to share a display name and reports
|
||
nothing.
|
||
|
||
### Still open from earlier milestones, and not M9's
|
||
|
||
Cross-layer duplication (`CONTEXT-AND-MEMORY.md` §22); the discarded-history
|
||
recovery screen (§63); whole-transcript copy and story search (§77, §78); the
|
||
read-only RPG world state; the inert legacy tables and the dual-dialect
|
||
migration code awaiting their cleanup migration.
|
||
|
||
---
|
||
|
||
## X. Final milestone assessment
|
||
|
||
### Is M9's Definition of Done satisfied?
|
||
|
||
`BUILD-MILESTONES.md` states it as: *"A campaign can be safely exported,
|
||
imported into a clean data directory, and reopened at the exact intended active
|
||
position with authoritative history/state intact."*
|
||
|
||
**Yes.** Measured three ways — in the suite, across two server processes with
|
||
two data directories and the exporting process stopped, and in a real browser
|
||
against a real narrator.
|
||
|
||
### The brief's closing questions, answered directly
|
||
|
||
| Question | Answer |
|
||
| --- | --- |
|
||
| **Can a campaign be moved into a clean data directory and reopened at the exact head?** | **Yes.** Into a database file that had never existed, from a second process, with the first stopped. The fixture's head is behind its own branch's tip *and* the abandoned line's, and is not the newest row written — so an importer guessing the tip, the newest row or the deepest row lands elsewhere. It opens where it was left, the retained future is there, and Redo reaches the same next turn in copy and source |
|
||
| **Is all authoritative state intact?** | **Yes.** The document at the exported head matches exactly; every historical position matches, walked in step; restoration stays a snapshot read rather than a replay; and the audit that explains the state travels with it — including the manual correction, which is the one change no narration explains |
|
||
| **Is historical prompt/context evidence intact?** | **Yes**, and this is the M8 handoff closed. An old turn in a restored campaign shows the prompt it was actually given and the passages it was shown, **after the source has been deleted in the copy** — asserted directly, not argued |
|
||
| **Are Save Points intact?** | **Yes.** Names, notes and coordinates; each resolves; each restores to its own position and its own state through M3's head movement; later history survives. A Save Point the file cannot satisfy is dropped, never retargeted |
|
||
| **Is knowledge intact without original filesystem paths?** | **Yes.** Content identical by SHA-256, classifications and lifecycle intact, searchable immediately — verified with the exporting machine's process dead and its directory holding a database the importer never opened |
|
||
| **Can derived data be rebuilt?** | **Yes**, and running that test found a defect that had made Reindex — the documented repair — unable to repair the state it most needed to. Both ends are fixed, and a database already carrying the damage repairs itself |
|
||
| **Can a consistent SQLite backup be created?** | **Yes**, through SQLite's online backup API, **while the application is being written to**, verified with `quick_check` before it is kept, and opened independently and read in every test. It works in the Docker image and lands on the mounted volume |
|
||
| **Are legacy bundles still supported?** | **Yes.** v1, v2, and every seam from pre-active-head onward — with nothing invented at any of them. A pre-M9 file gets no manufactured audit trail; a pre-M3 file opens at its tip *because that is the position it recorded* |
|
||
| **Are any required M9 conditions unverified?** | **No.** Every required condition has evidence at the boundary it names. What remains is stated as residual risk in §W with its measurement, not as an unverified claim |
|
||
| **Is it safe to proceed to M10 after review?** | **Yes**, on this evidence. M9 changed no schema, altered no product requirement, and added no media surface. It leaves M10 the `scene` section of the state document as the extension point that exists today, and §L records the absence of media tables as *not applicable* rather than inventing schema to satisfy the word "metadata" |
|
||
|
||
### What a reviewer should look at hardest
|
||
|
||
Named deliberately, because a report that only presents its strengths is harder
|
||
to review than one that points at its own seams:
|
||
|
||
1. **The version bump (§D).** It is the one place M9 chose differently from two
|
||
earlier milestones that faced a similar decision. The argument is that an
|
||
absent key is unambiguous for a *position* and ambiguous for *evidence*; a
|
||
reviewer who disagrees should say so, because it is a contract decision and
|
||
not an implementation detail.
|
||
2. **The story-card decision (§K).** It removes something from the narrator's
|
||
prompt, which is a behaviour change in a portability milestone. The brief
|
||
authorised it and §73 required it; the judgement to check is whether stopping
|
||
at the prompt — and leaving the summariser's roster alone — is the right
|
||
line.
|
||
3. **The bundle ceiling (§R).** Real, measured, and stated rather than hidden.
|
||
Worth checking that ~279 turns is acceptable for v1 given a 100-turn
|
||
certification target, and that the export-succeeds/import-refuses asymmetry
|
||
is tolerable until a streaming import exists.
|
||
4. **Finding 3.** It would have shipped looking correct in memory and wrong on
|
||
disk. It was caught only because the test asserted the outcome rather than
|
||
the call, which is worth generalising.
|
||
|
||
### Not done, and deliberately
|
||
|
||
No image, video, TTS, STT, media provider or job architecture; no M10 media
|
||
coordinator; no 100-turn certification; no WCAG audit; no new context-window or
|
||
provider architecture; no discarded-history recovery browser. Retained history
|
||
survives correctly so a later screen can use it, which was M9's part of that.
|
||
|
||
---
|
||
|
||
## Y. Post-M8 hands-on playtest findings
|
||
|
||
**These are not M9 defects, were not caused by M9, and do not block M9
|
||
acceptance.** They are recorded here because this is the report a reviewer is
|
||
holding, and because the M9 report will be archived when M10's replaces it —
|
||
`BUILD-MILESTONES.md` carries the durable copy under M11 for that reason.
|
||
|
||
They come from a real play session against **accepted, signed M8** — a real
|
||
browser, a real trusted-LAN Ollama, narrator `qwen2.5:3b-instruct-16k`, and a
|
||
disposable isolated campaign database. **That database was deliberately
|
||
destroyed afterwards**, so the stored context snapshot for the turn in finding
|
||
D no longer exists. Everything below is therefore recorded as an *observed
|
||
symptom*, and no root cause is claimed that cannot now be proven.
|
||
|
||
**Nothing in this section changed any application code.** Where the text states
|
||
how the product behaves today, it is from reading the code and running it, and
|
||
it is labelled as a mechanism rather than as a proven cause of what was seen.
|
||
|
||
### A. The browser still calls the product "AI D&D"
|
||
|
||
**Observed:** the browser tab/title reads `AI D&D`.
|
||
|
||
**Verified:** `frontend/index.html` line 18 is `<title>AI D&D</title>`. It
|
||
is the inherited upstream title and M8 did not change it.
|
||
|
||
**Not a false claim by any accepted document.** M8's report never asserted the
|
||
title was changed — its terminology audit covered `branch`, `fork`, `node`,
|
||
`head` and `depth`, and its "does it feel like the intended storyteller"
|
||
argument lists the navigation, the inspector, the glyph buttons and the removed
|
||
screens. So this is an **uncovered gap**, not documentation that needs
|
||
correcting. No documentation was corrected, and no code was changed.
|
||
|
||
**The naming question is genuinely open, and should not be closed by a
|
||
find-and-replace.** The planning package and `README.md` call the product
|
||
*Adventure Storyteller*, but `SPECIFICATION.md` is explicit that the engine must
|
||
stay genre-agnostic — science fiction, mystery, horror, historical, westerns —
|
||
and *Adventure* is narrower than the product it names. A repository-wide rename
|
||
was deliberately **not** performed in this pass. For planning purposes a neutral
|
||
working name such as **Interactive Story** is used, and a browser-tab form such
|
||
as `<Campaign Name> — Interactive Story` is a candidate, not a decision.
|
||
|
||
**Owner: M11 release polish.** No earlier milestone touches the shell metadata.
|
||
It is a small, self-contained change whose only hard part is the naming
|
||
decision, which is the repository owner's.
|
||
|
||
### B. Undo/Redo loses the reader's orientation
|
||
|
||
**Observed:** Undo worked, and it was hard to tell which point in the story the
|
||
reader had moved to.
|
||
|
||
**Not a correctness problem.** M3's active-head semantics behaved correctly and
|
||
M9 re-verified them at every boundary (§E, §S). This is presentation.
|
||
|
||
**What the spec said, and why that was not enough.** `BROWSER-UX-SPEC.md` §8 —
|
||
*"The current endpoint should be clear."* — is five words, and the second
|
||
sentence beside it ("the input box always continues from the currently active
|
||
story head") was **already true** while the reader was lost. A requirement that
|
||
a working implementation satisfies while a real user cannot answer the question
|
||
is too vague to hold the behaviour. §8 has been strengthened in this pass; no UI
|
||
text is prescribed, because none is ratified.
|
||
|
||
**The requirement, stated without wording:** after Undo, Redo, a Save Point
|
||
restore, an edit to an earlier turn, or any other movement of the active
|
||
position, a reader should be able to tell where they now are in the visible
|
||
story **without needing implementation terminology** — `branch`, `head`, `node`
|
||
and `depth` remain forbidden at the surface (§38 of the spec, and M8's audit).
|
||
A lightweight indicator such as `Moment 8` → `Moment 7`, optionally noting that
|
||
later story is still available, is a candidate. **Exact wording deliberately not
|
||
settled here.**
|
||
|
||
**Owner: M11 UX/release polish**, with a browser regression scenario.
|
||
|
||
### C. Narration length did not feel effective
|
||
|
||
**Observed:** setup offered roughly *one paragraph* / *2-4 paragraphs* / *longer
|
||
exposition*; the reader chose **2-4 paragraphs** and felt replies were
|
||
substantially longer than that.
|
||
|
||
**Deliberately not recorded as "the model ignored instructions."** The evidence
|
||
does not establish that, and there is a mechanism in the code worth checking
|
||
first.
|
||
|
||
**The mechanism, verified by reading and running the code.** There are **two
|
||
independent length controls, and they do not reference each other**:
|
||
|
||
1. The setup choice becomes **one English sentence** appended to the campaign's
|
||
`ai_instructions` — `"Keep responses to roughly two to four paragraphs."`
|
||
(`frontend/src/pages/NewCampaign.jsx`, `LENGTH_SENTENCE`). It changes **no
|
||
generation setting**. It sits in the **cached system block**, near the top of
|
||
the prompt.
|
||
2. `context/builder.py`'s `length_hint()` derives a numeric word range from
|
||
`Settings.max_output_tokens` — a **global application setting**, not the
|
||
campaign's choice — and emits it as a `[Hard limit: …]` line placed **after
|
||
the history**, near the end of the prompt.
|
||
|
||
At the default `max_output_tokens = 800`, that second line reads, measured:
|
||
|
||
```
|
||
[Hard limit: this turn must not exceed 506 words, and it should not stop short
|
||
of about 177. Prefer the lower end of that range unless the scene genuinely
|
||
needs more. Finish the narration and append the state block well inside the
|
||
limit.]
|
||
```
|
||
|
||
**and it is byte-identical whether the reader chose brief, medium or long.** A
|
||
reader asking for two to four paragraphs is simultaneously told, in the more
|
||
recent and more numerically explicit of the two instructions, not to stop short
|
||
of about 177 words and that up to 506 are permitted.
|
||
|
||
**This is a mechanism, not a proven cause.** Whether it produced what this
|
||
reader saw needs the reproduction below; the narrator's own instruction
|
||
following is also in play, and both could contribute.
|
||
|
||
**What a reproduction must measure**, rather than judge by eye:
|
||
|
||
1. exactly what length instruction(s) enter the **stored** prompt — both of the
|
||
above, with their positions;
|
||
2. whether the setup choice changes `max_output_tokens` or any generation
|
||
setting (**today: it does not**);
|
||
3. actual words, tokens and paragraph counts across repeated realistic turns per
|
||
setting, so a *directional* effect can be shown or disproved;
|
||
4. at least the reference 3B narrator **and** a stronger local narrator, since
|
||
instruction-following differs.
|
||
|
||
**A design candidate, explicitly not ratified:** clearer approximate targets —
|
||
`Brief ~100-200 words`, `Standard ~200-400`, `Detailed ~400-700` — with the
|
||
setting actually moving the numeric budget. **Do not hard-truncate prose**: the
|
||
state block is emitted last and truncation removes it, which is the failure
|
||
`length_hint`'s own comments exist to avoid.
|
||
|
||
**Owner: M11 realistic-model behaviour validation**, with a realistic-model test
|
||
rather than only a scripted-provider one.
|
||
|
||
### D. Character identity / coreference confusion — the most important one
|
||
|
||
**Observed:** a story established four people in an office — **Bill**
|
||
(protagonist), **Roger**, **John**, **Alice**. Later narration treated Alice as
|
||
though there were two different Alices, in language equivalent to *"Alice
|
||
wondered what Alice was doing."*
|
||
|
||
**Root cause: UNKNOWN, and it cannot now be established.** The disposable
|
||
playtest database was deliberately destroyed, so the stored context snapshot for
|
||
that turn is gone. This report claims neither an inference-model defect nor an
|
||
application defect. The candidate classes are:
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| **Model failure** | the stored prompt correctly identifies one Alice and the 3B model still makes a coreference error |
|
||
| **State failure** | the narrative state holds duplicate or conflicting Alice records |
|
||
| **Context/derived failure** | state is correct, but assembled context, a summary, a retrieved memory or the history rendering presents two identities |
|
||
| **Combined weakness** | the context is not contradictory but is insufficiently explicit for a small model, producing an avoidable failure |
|
||
|
||
**One structural fact a reproduction should check first**, verified by reading
|
||
the code and stated as a fact about the implementation rather than as a cause:
|
||
|
||
> **The narrative state permits two distinct entities to share one display
|
||
> name, and nothing reports it.** Entities are keyed by the id the model
|
||
> supplies (`state["entities"][event["entity"]]`, `narrative/apply.py`), with
|
||
> `name` a separate display field. `validate.py`'s `DUPLICATE_ENTITY` rejects
|
||
> re-creating an entity **with the same key** — `key in known` — and there is no
|
||
> check anywhere in the narrative layer on the display name. So
|
||
> `create_entity(entity="alice", …, name="Alice")` followed by
|
||
> `create_entity(entity="alice_2", …, name="Alice")` both succeed, and the state
|
||
> then holds two entities that both render as "Alice".
|
||
|
||
That is exactly one of the failure modes this finding describes. It does **not**
|
||
establish that it happened here — no evidence survives — and a reproduction may
|
||
well land on one of the other three classes instead.
|
||
|
||
**Owner: M11 realistic-model / context diagnostic.** The test plan is below and
|
||
in `BUILD-MILESTONES.md` under M11.
|
||
|
||
### The M11 character-identity diagnostic, as a test plan
|
||
|
||
Deterministic setup with at least a protagonist and three same-scene
|
||
supporting characters — **Bill** (protagonist), **Alice**, **Roger**, **John**
|
||
— with unambiguous identities and roles established up front. Then a
|
||
multi-character interaction over enough turns to stress: pronouns; dialogue
|
||
attribution; people entering and leaving; reference by name; reference by role;
|
||
and one character speaking *about* another.
|
||
|
||
**Detect and report, at minimum:** duplicate character creation; same-name
|
||
entity duplication; protagonist identity drift; dialogue attributed to the wrong
|
||
person; a character referring to themself as a separate same-named character;
|
||
and state/context disagreement about identity.
|
||
|
||
**On any failure, preserve and report all of:** the authoritative state
|
||
immediately before generation; the exact stored context/prompt snapshot; the
|
||
recent-history section; summaries; retrieved memories; imported knowledge if
|
||
any; the narrator output; and the model identifier and settings. **M9 makes all
|
||
of that portable** (§H), so a failing campaign can now be exported and handed to
|
||
whoever investigates it — which is the practical reason this diagnostic is
|
||
newly worth writing.
|
||
|
||
Then classify from the evidence:
|
||
|
||
```text
|
||
STATE DEFECT
|
||
CONTEXT ASSEMBLY DEFECT
|
||
DERIVED MEMORY/SUMMARY DEFECT
|
||
MODEL FAILURE WITH CORRECT CONTEXT
|
||
AMBIGUOUS / MULTIPLE CONTRIBUTORS
|
||
```
|
||
|
||
Two rules for whoever runs it: **do not "fix" a model failure by changing
|
||
authoritative story state**, and **do not blame the model if the prompt already
|
||
contained the identity error.** That distinction should become part of M11's
|
||
realistic-model review methodology rather than a one-off judgement.
|
||
|
||
### The standard fixture does not cover this failure class
|
||
|
||
Reviewed, as asked. `TEST-CAMPAIGN-FIXTURE.md` stresses knowledge boundaries,
|
||
secrets, authority precedence, branch leakage and possession — its seven
|
||
deliberate traps are all of those — and it contains **no identity trap at all**;
|
||
the word *coreference* does not appear in it. Its on-stage cast is effectively
|
||
two people, Aldric and Mara, with Edrin established as missing rather than
|
||
present.
|
||
|
||
So the user's suspicion is correct: **same-scene multi-character identity
|
||
continuity is not exercised by the standard fixture.**
|
||
|
||
**The established fixture was deliberately not modified.** It is the
|
||
deterministic baseline several milestones' results are compared against, and
|
||
changing it would invalidate those comparisons. A **companion** fixture is
|
||
proposed instead, in a new appendix to that document, named
|
||
`Multi-Character Identity Test` and explicitly additive.
|
||
|
||
---
|
||
|
||
*Written at the end of implementation, before review. The tree is staged; §32 of
|
||
the brief governs what happens next.*
|