M9: a campaign you can actually get back

A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
JesseMarkowitz
2026-09-07 01:55:45 -04:00
co-authored by Claude Opus 5
parent 1ce9972760
commit 44edece67e
46 changed files with 9227 additions and 178 deletions
+29
View File
@@ -128,6 +128,35 @@ The current endpoint should be clear.
The input box always continues from the currently active story head.
### 8A. The reader must be able to tell where they are (recorded 2026-09-07)
A hands-on session against accepted M8 found the sentence above **too weak to
hold the behaviour it names**. Undo worked correctly, the input box did continue
from the active head — so both statements above were satisfied — and the reader
still could not tell which point in the story they had moved to.
The requirement, stated so that a working implementation cannot satisfy it while
a reader is lost:
> After Undo, Redo, a Save Point restore, an edit to an earlier turn, or any
> other movement of the active story position, the reader should be able to
> identify **where they now are in the visible story** — and, where it matters,
> whether later story remains available ahead of them.
Two constraints on any solution:
- **No implementation terminology.** `branch`, `fork`, `node`, `head` and
`depth` stay off the reader-facing surface (§38), which is what makes this a
presentation problem rather than a labelling one.
- **It must be observable**, not merely inferable from the transcript scrolling,
so that a release test can decide it.
**No wording is prescribed here, and none is ratified.** A lightweight named
position — `Moment 8` becoming `Moment 7` after an Undo, optionally noting that
later story is available — is one candidate among others. Ownership is M11
release polish; `V1-ACCEPTANCE-TESTS.md` §P1 records what must be settled before
this can become an acceptance test.
## 9. User Turn Presentation
User messages should support:
+193 -1
View File
@@ -1118,6 +1118,187 @@ Make campaigns portable and recoverable without losing lineage, state, knowledge
A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact.
## Status: COMPLETE — 2026-09-07, pending independent review
Implemented on `m9-recovery` from the signed M8 commit `1ce9972`, measured
before and after against the same fixture, and verified in a real browser
against a real narrator. `planning/reports/M9-IMPLEMENTATION-REPORT.md` is the
implementer's account, written for a reviewer.
**What it delivered, beyond the scope list above:**
- **The bundle became a version, and the reason is a rule.** `ai-dnd-adventure-v3`.
Everything M9 adds could have been an optional key, the way four earlier
additions were — and that mechanism fails exactly here, because a v2 file with
no prompt provenance is ambiguous between "written before M9" and "written by
M9 from a campaign with none". A version is how a recovery file states what it
was capable of recording. v1 and v2 are still read, and every seam from
pre-active-head onward is tested.
- **A third category of data.** "Chosen travels, derived is recomputed" was
enough until stored prompts had to be decided. They are derived and must
travel, so the rule is now chosen / **evidence** / rebuildable, and the test
separating the last two is not "could this be recomputed" but "would a
recomputation answer the same question".
- **The M8 handoff on prompt provenance is closed.** An old turn in a restored
campaign shows what it was actually given, after the source has been deleted,
the canon edited and the state moved on.
- **State events and proposals travel**, so a moved campaign can still say why
its state is what it is — and a manual correction is still identifiable as
one, which it was not before.
- **Summaries travel with their coordinates**, so an abandoned line's summary is
still ineligible after the move, and a moved campaign resumes with its
long-story continuity instead of behaving like a new one.
- **A verified SQLite backup**, using the online backup API rather than a file
copy, taken while the application is running, with a browser control in
Settings.
- **Take parentage**, memory authority and parser/chunking versions, each
closing a smaller fidelity loss.
**Three defects found by running the milestone's own tests, and fixed here:**
1. **Deleting a campaign leaked its FTS index rows**, and SQLite then handed the
freed ids to the *next* source imported into *any* campaign, which failed
with an integrity error. Reindex could not repair it either. Pre-existing
since M7; both ends are now closed and an already-damaged database repairs
itself with no migration.
2. **An imported node with no state snapshot was stamped with the campaign's
head state**, so an Undo to turn 2 showed what the story knew at turn 20 —
M5's review finding 3, arriving through the import.
3. **A snapshot's `source_id` was not being translated on import** because the
relink mutated a dict in place, which a non-`MutableDict` column does not
notice. Found by a test asserting the outcome rather than the call.
**One decision the brief asked for, made and recorded:**
**Story cards are compatibility-only legacy data, and no longer enter the
narrator's prompt.** They still travel in both directions, the rows and the API
stay, and `memorybank.cast_brief` still reads them as the summariser's character
roster. What stops is the injection: a keyword-matched card arrived in front of
the narrator as `World Lore: …` with no class, no visibility, no source, no way
to switch it off since M8 removed the editor, and no row in the context
inspector — which is `IMPORTED-KNOWLEDGE-DESIGN.md` §73's "alternate untracked
path around the new knowledge authority/provenance rules" in as many words.
**Debt carried forward, deliberately:**
- **A long campaign's bundle has a measured ceiling: ~279 turns** against the
20 MB import limit. A per-turn prompt contains the story so far, so carrying
one per turn is O(turns²); compressing them inside the file cut that to about
an eighth of what it would have been. **M9's additions account for only 12% of
the ceiling.** The other 88% is the per-position narrative state document,
which is 74% of a bundle and which v2 already carried — so lifting the ceiling
means addressing that, not the evidence. Far beyond M11's 100-turn
certification (14% of the cap), and stated with its measurement rather than
hidden. A streaming or chunked import is the fix if a later milestone needs
one; the asymmetry to know about is that such a campaign can still be exported
and would be refused on import.
- **No discarded-history recovery screen** (§63). M9's job was that retained
history survives correctly so a later screen can use it; it does.
- **No whole-transcript copy and no story search** (§77, §78) — still M11 or
later.
---
# Post-M8 hands-on playtest findings — recorded 2026-09-07, owned by M11
**Not M9 defects, not caused by M9, and they did not block M9 acceptance.**
Recorded here rather than only in the M9 report because a milestone report is
archived when the next one replaces it, and these must not go with it.
They come from a real play session against **accepted, signed M8**: real
browser, real trusted-LAN Ollama, narrator `qwen2.5:3b-instruct-16k`, disposable
isolated campaign database. **That database was deliberately destroyed
afterwards**, so the stored context snapshot for finding D is gone and no root
cause is claimed for it. Full write-up, with the verification behind each
mechanism, is in the M9 report's **§Y** (in `archive/milestone-reports/` once
M10's report replaces it).
**These are M11's, and explicitly not M10's.** M10 is media-readiness
architecture and stays bounded; it inherits them as known carry-forward only.
### A. The browser still calls the product "AI D&D" — M11 release polish
`frontend/index.html` still carries the inherited `<title>AI D&amp;D</title>`.
M8 changed the navigation, the inspector and the screens, and never claimed the
title — so this is an uncovered gap rather than a false claim.
**Do not fix it with a find-and-replace to "Adventure Storyteller".**
`SPECIFICATION.md` requires a genre-agnostic engine, and *Adventure* is narrower
than the product. The naming decision is the repository owner's; a neutral
working name such as **Interactive Story** and a tab form such as
`<Campaign Name> — Interactive Story` are candidates, not decisions.
### B. After Undo, the reader cannot tell where they are — M11 UX polish
Undo behaved correctly (M3 semantics; re-verified throughout M9). The reader
could not tell **which point in the story** they had moved to.
`BROWSER-UX-SPEC.md` §8 said only *"The current endpoint should be clear"*, and
its companion sentence was already true while the reader was lost — so the
requirement could not hold the behaviour. §8 has been strengthened to state the
orientation requirement; **no UI text is prescribed**, because none is ratified.
A `Moment 8` → `Moment 7` style indicator is a candidate. Needs a browser
regression scenario.
### C. Narration length has no measurable effect — M11 realistic-model behaviour
The setup choice becomes **one English sentence** in the campaign's
`ai_instructions` and changes **no generation setting**. Independently,
`length_hint()` derives a numeric word range from the **global**
`Settings.max_output_tokens` and places it after the history — and at the default
800 it reads *"must not exceed 506 words, and it should not stop short of about
177"* **identically for brief, medium and long**.
That is a mechanism, verified by reading and running the code — **not a proven
cause** of what the reader saw; the narrator's instruction following is also in
play. A reproduction must measure what enters the stored prompt, whether the
setting moves any generation budget, and actual word/paragraph counts across
repeated turns, on the reference 3B narrator **and** a stronger local one.
Clearer numeric targets (`Brief ~100-200 words` and so on) are a design
candidate, not ratified. **Do not hard-truncate prose** — the state block is
emitted last and truncation removes it.
### D. Character identity / coreference confusion — M11 diagnostic
Four people in one scene — Bill (protagonist), Roger, John, Alice — and later
narration treated Alice as two different Alices.
**Root cause UNKNOWN and no longer establishable.** Candidates: a model
coreference failure on a correct prompt; duplicate/conflicting state; a
context/summary/memory assembly failure; or a context that is not contradictory
but too implicit for a small model.
**One structural fact to check first**, verified by reading the code: the
narrative state **permits two entities to share a display name and reports
nothing**. Entities are keyed by the model-supplied id; `DUPLICATE_ENTITY`
rejects only a repeated *key*; no check exists on `name`. That is one of this
finding's failure modes, and establishes nothing about what happened.
**M11 must run an explicit diagnostic** with a protagonist and three same-scene
supporting characters, stressing pronouns, dialogue attribution, entrances and
exits, reference by name and by role, and one character speaking about another.
It must detect duplicate creation, same-name duplication, protagonist drift,
misattributed dialogue, self-as-other reference, and state/context disagreement
— and on any failure preserve the pre-generation state, the exact stored prompt
snapshot, history, summaries, memories, imported knowledge, narrator output and
model settings, then classify:
```text
STATE DEFECT / CONTEXT ASSEMBLY DEFECT / DERIVED MEMORY-SUMMARY DEFECT /
MODEL FAILURE WITH CORRECT CONTEXT / AMBIGUOUS
```
**Do not "fix" a model failure by changing authoritative state, and do not blame
the model if the prompt already contained the error.** M9 made all of that
evidence portable, so a failing campaign can be exported and handed over intact.
**The standard fixture does not cover this class.** `TEST-CAMPAIGN-FIXTURE.md`'s
seven traps are knowledge, authority, branch leakage and possession; there is no
identity trap and its on-stage cast is effectively two people. The established
fixture was **not modified** — it is the deterministic baseline earlier results
are compared against. A companion fixture, `Multi-Character Identity Test`, is
proposed in an appendix to that document.
---
# M10 — Future Media Extension Hooks Only
@@ -1126,6 +1307,13 @@ A campaign can be safely exported, imported into a clean data directory, and reo
Preserve the approved future media interfaces without adding a media-generation dependency to v1.
## Note — the post-M8 playtest findings are **not** M10 scope
The four findings recorded above are owned by M11. M10 inherits them as known
carry-forward items only: it should neither implement nor test them, and its
scope below is unchanged by them. They are listed before this milestone rather
than after it only because they were recorded during M9's closeout.
## Scope
- scene snapshots/packets suitable for future providers,
@@ -1178,7 +1366,11 @@ Validate the full product against the release contract after all functional mile
- fantasy and science-fiction fixtures,
- export/import/recovery tests,
- migration tests,
- documentation and packaging.
- documentation and packaging,
- **the four post-M8 hands-on playtest findings above**: the browser product
name, reader orientation after history movement, a narration-length setting
with a measurable effect, and the multi-character identity diagnostic with its
companion fixture.
## Tests / Acceptance
+112
View File
@@ -887,6 +887,101 @@ has stopped being told the rules, with nothing to notice.
A bundle written before M7 has no knowledge section and imports with an empty
library, which is what such a campaign had.
### M9: the format became a version, and the evidence started travelling
**M9 bumped the format to `ai-dnd-adventure-v3`**, and the reason is a rule
rather than a preference. Everything M9 added *could* have been an optional key
read with `.get`, the way `persona`, `checkpoints`, `narrativeState` and
`knowledge` each were. That mechanism stops working at exactly this addition:
a v2 file carrying no prompt provenance is **ambiguous** — written before M9,
when no file could carry one, or by M9 from a campaign whose turns predate the
column? Those are different facts about the campaign and a reader has to be able
to tell them apart. It is the same distinction the head rule above draws when it
says a pre-M3 file opens at its tip *because that is the position such a file
recorded*. A version number is how a recovery file states what it was capable of
recording. The reader keeps every older version; only the writer moved.
What each version can be trusted to say:
```text
v1 a linear story, its turns, and its retries as a repeating group
v2 + the tree, the live flags, the after-snapshots, the chosen head,
Save Points, the narrative state document, imported knowledge
v3 + state events and proposals, historical prompt/context provenance,
lineage-anchored summaries, take parentage, memory authority
```
**Three categories, not two.** §31 below distinguishes authoritative from
derived, which was sufficient until M9 had to decide about stored prompts. They
are derived — a machine assembled them — and they must travel anyway, so the
rule the bundle applies has a middle category:
```text
chosen what a person decided: the story, the head, the takes, the
Save Points, the classifications, the canon. travels
evidence what happened, and what the application was told at the time:
the state events and proposals, the per-turn prompt and the
passages it was shown, the model and generation settings that
turn ran under. travels
rebuildable a deterministic function of what travels: knowledge passages,
the FTS index, embeddings, the branch lineage cache.
rebuilt on import
```
The test that separates evidence from rebuildable is **not** "could this be
recomputed" but "would a recomputation answer the same question". Rebuilding the
FTS index answers the same question it answered before. Rebuilding an old turn's
prompt does not — it would say what that turn *would be told now*, from today's
canon, today's sources and today's state, which is the opposite of what the
context inspector is for. Historical evidence is not a cache.
So v3 additionally carries, all restored verbatim:
- **`stateEvents` and `stateProposals`.** §17's hybrid keeps the events for
audit and the snapshots for restore; v2 carried only the snapshots, so a moved
campaign could be read at any position and could no longer say what changed
there, who asserted it, or what the value was before. A **manual correction**
was the worst case: the one state change no narration explains, and with the
events gone nothing distinguished it from something the story established.
Both tables travel, because the inspector reads both — the event says what was
accepted and the proposal says what the model asked for and what was refused.
- **A per-node context snapshot**, which is `SPECIFICATION.md` §6.1's exact
prompt/context record. It carries the assembled prompt section by section, the
passages retrieved with the text each supplied, which summary was eligible,
and the model and generation settings the call ran under — so an old turn can
still say what it was told after the source was deleted, the canon edited, the
chunker changed and the state moved on. Stored once per turn on the live
attempt, so a retried turn is not a multiplier.
- **`summaries`**, with the coordinate that decides eligibility. v2 carried only
the `storySummary` mirror, which has no lineage of its own, so a restored
campaign resumed with no usable long-story continuity — and a summary
belonging to an abandoned line stays ineligible after the move for the same
reason it was before it: eligibility is the coordinate lying on the active
capped lineage, not a stored flag.
- **Take parentage**, so attempts under two different takes of one turn stay two
pagers rather than merging into one.
- **Memory `authority`**, so a heuristic memory is not promoted to accepted
story by being moved (F07).
- **Per-source `parserVersion`/`chunkingVersion`**, recording what produced the
passages a historical retrieval record describes.
**Encoding.** A per-turn prompt contains the story so far, so one per turn is
O(turns²) in campaign length — measured at 20,797 bytes per turn at turn 20 and
55,291 at turn 120, 68% of a 9.7 MB file. The snapshot therefore travels as
`contextSnapshotZ`: the same JSON, zlib-compressed and base64-encoded, using the
same pack/unpack the database column already uses. Nothing is dropped or
summarised; the file is still JSON, and every other section of it is still plain
text. The plain `contextSnapshot` key is still read and takes precedence, so a
hand-edited file keeps importing.
**Two pointers are translated on import, and nothing else is.** Branch numbers
already were. M9 adds the `source_id` inside a restored retrieval record: it
names a row on the machine that wrote the file, so left alone it would point the
inspector's "open this source" at whatever holds that id here. Where the file's
own knowledge section contains the source it is repointed; where it does not — a
source deleted before the export — it becomes `null`, and the record keeps its
text and filename. The evidence is never rewritten; only the pointer is.
## 30. Deletion vs Archival
The system must distinguish:
@@ -916,6 +1011,23 @@ Retry, Undo, Restore, and branch switching must not silently perform permanent d
Derived data should be rebuildable where practical.
**"Where practical" does real work in that sentence, and M9 had to split this
list to act on it** (§29). Embeddings and the lexical and semantic indexes are
deterministic functions of content that travels, so a rebuild answers the same
question and they are not exported. A **summary** is not: it took a model call,
it describes a stretch of story that may since have been abandoned, and
regenerating one on another machine produces different prose about a different
reading — so it is derived, not practically rebuildable, and it travels with the
coordinate that decides whether it still applies.
The same reasoning puts **stored prompt/context snapshots** on the travelling
side, and they are not in either list above because they are neither: they are
not authoritative — nothing decides anything from them — and calling them
derived would invite a rebuild. They are *evidence*: a record of what the
application was told at the time, which a regeneration would not reproduce
because it would use today's canon, today's sources and today's state. §29 states
the three-way rule the export applies.
## 32. Provenance
Important information should answer:
+41
View File
@@ -1110,6 +1110,47 @@ Normal imported files are campaign-level source material and need not inherit st
Story Cards may remain as an inherited authored-rule/lore primitive during migration if useful, but they must not become an alternate untracked path around the new knowledge authority/provenance rules.
### Settled in M9 (2026-09-07): compatibility-only, and out of the prompt
M8 removed the Story Card browser editor and left the question open; the M9
brief asked for it to be decided. The finding was that story cards **were** the
alternate untracked path the paragraph above forbids, and not in principle: a
keyword-matched card was injected into the narrator's prompt as
`World Lore: <entry>`, taking up to 40% of what was left after the imported
knowledge had been placed, with
- no class, so nothing framed how far the narrator could rely on it;
- no visibility, so no narrator-only distinction existed;
- no source, no hash and no lifecycle, so there was nothing to disable;
- no browser surface after M8, so a reader could neither see nor switch it off;
- no row in the context inspector, which renders `knowledge` and never rendered
`cards`;
and competing with imported Canon for one budget, which is the arrangement M7
spent a milestone separating.
**The decision, and it is the smallest change that closes it:**
| | |
| --- | --- |
| New export | carries them, unchanged, under `storyCards` |
| Legacy import | accepted, unchanged, from every format version |
| Normal narration | **no longer reached.** The `world_lore` section is gone |
| Re-export | carries them again, so a round trip destroys nothing |
Nothing is deleted. The rows stay, the `/api/story-cards` endpoints stay, and
`memorybank.cast_brief` still reads them as the **summariser's character
roster** — that names who is on stage so a memory says "Aldric" rather than
"he", never reaches the narrator, and every memory written from it is
authority-classified by the application afterwards. The `cards` key stays in the
context report and is now always empty for a new turn, because M9 made
historical snapshots portable and an old turn's record must go on saying that
story cards were included.
A campaign that wants the narrator to know something imports it as Canon,
Reference or Inspiration, where it is classified, inspectable, disableable and
attributable — which is what §73 asks for.
### Retrieval implementation direction
Use:
+48 -36
View File
@@ -3,7 +3,8 @@
**This file is the index. Start here.**
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
milestones **M1 through M6 implemented and accepted** (M3 and M4: 2026-09-03;
milestones **M1 through M8 implemented and accepted**, and **M9 implemented and
awaiting review**. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
independent review found a real defect and a corrective pass fixed it.
@@ -19,8 +20,15 @@ carries the closeout: the build-evidence classification in its §P, finding 14's
operational resolution, and the acceptance record in its §V. The M8 tree is
staged and awaits the repository owner's signed commit.
**Next: M9 — Export, Backup, Recovery, and Migration Hardening.** It has not
been started.
**M9 — Export, Backup, Recovery, and Migration Hardening — is implemented and
awaiting independent review** (2026-09-07).
`reports/M9-IMPLEMENTATION-REPORT.md` is the implementer's account, written for
a reviewer: a set of claims with the measurements attached, not yet a record of
acceptance. M8's report has moved to `archive/milestone-reports/`, which is
where a milestone report goes once the next milestone's report replaces it.
**Next: M10 — Future Media Extension Hooks Only.** It has not been started, and
no brief for it exists.
**Package version:** see `VERSION.md`, which records what each revision changed
and why.
@@ -82,8 +90,8 @@ Two standing qualifications:
| Document | What it is for |
| --- | --- |
| `SPECIFICATION.md` | What the product must do. The top of the authority order. |
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M7 built, recorded as fact. |
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the export shape. |
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M9 built, recorded as fact. |
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the v3 export contract. |
| `STORY-BRANCH-SEMANTICS.md` | Undo/Redo/Retry/branch/take behavior, including the M3 ratifications. |
| `CONTEXT-AND-MEMORY.md` | Prompt assembly, summarization, branch-safe memory. |
| `IMPORTED-KNOWLEDGE-DESIGN.md` | Canon / Reference / Inspiration knowledge as a first-class subsystem. |
@@ -110,7 +118,7 @@ Two standing qualifications:
10. `BROWSER-UX-SPEC.md`
11. `V1-ACCEPTANCE-TESTS.md`
12. `DECISIONS/` — all of them; they are short.
13. `reports/M8-IMPLEMENTATION-REPORT.md`, for what the most recent milestone
13. `reports/M9-IMPLEMENTATION-REPORT.md`, for what the most recent milestone
actually left behind — reading it as a claim to check, not a record, until
it is reviewed. Nothing in `planning/archive/` unless sent there.
@@ -143,20 +151,18 @@ work until Phase 0 closes — which Phase 0 satisfied on 2026-09-01. It is in
`reports/` holds the report for the milestone most recently completed, because
that is the one the next milestone's planning has to consult:
- `reports/M8-IMPLEMENTATION-REPORT.md` — the M8 implementation, its baseline
UX measurement, and the browser evidence for every acceptance test it claims.
Written by the implementer for an independent reviewer, and completed at
closeout after that review accepted the milestone: it is a set of claims with
the measurements attached **and** the record of the acceptance. Its §U carries
the M9 handoff — the four questions the next brief has to decide.
- `reports/M9-IMPLEMENTATION-REPORT.md` — the M9 implementation: the measured M8
portability baseline it started from, the final bundle contract, and the
evidence for every acceptance test it claims. Written by the implementer for
an independent reviewer, so it is a set of claims with the measurements
attached and **not** a record of acceptance. Its §W carries the M10-M11
handoff.
**It stays here until M9's report replaces it.** A milestone report is useful
during the immediately following milestone; M8's is not archived merely
because M8 is accepted.
**It stays here until M10's report replaces it.**
Completed earlier milestones are in `archive/milestone-reports/`, which M7's
report joined when M8's was written: a milestone report is useful during the
immediate next milestone and historical afterwards. M1-M7 are all there,
Completed earlier milestones are in `archive/milestone-reports/`, which M8's
report joined when M9's was written: a milestone report is useful during the
immediate next milestone and historical afterwards. M1-M8 are all there,
unedited.
## The decision this package rests on
@@ -282,33 +288,39 @@ Milestone M8 COMPLETE / ACCEPTED (2026-09-06)
story operations review + closeout, in sequence
|
v
Milestone M9 NEXT — not started
export, backup, recovery, see BUILD-MILESTONES.md
migration hardening
Milestone M9 COMPLETE — awaiting review (2026-09-07)
export, backup, recovery, reports/M9-IMPLEMENTATION-REPORT.md
migration hardening bundle format v3; SQLite online backup
|
v
M10-M11, one at a time see BUILD-MILESTONES.md
Milestone M10 NEXT — not started
future media extension hooks see BUILD-MILESTONES.md
|
v
Milestone M11 see BUILD-MILESTONES.md
```
## Stop Rule
**One milestone at a time. Do not begin a milestone before its brief exists.**
**No M9 brief has been prepared.** Writing one is the current action, informed
by the M8 report and by the debt `BUILD-MILESTONES.md` records against M8 — in
particular that the campaign bundle still carries no context snapshots, so an
imported campaign has no historical prompt provenance; that story cards survive
in the backend and the bundle with no browser surface, and M9 should decide
deliberately whether the bundle keeps carrying them; and that a deployment whose
Ollama enforces a small context window truncates an imported long campaign
immediately unless the `DEVELOPMENT.md` procedure or a matching
`context_token_budget` is applied.
**No M10 brief has been prepared**, and M9 is not accepted — it is implemented
and awaiting an independent review. Writing the M10 brief is the action after
that review closes, informed by the M9 report's §W.
**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`: story
cards have no browser editor, the RPG world state is read-only, copy is
per-message only, there is no discarded-history recovery screen, and the tablet
layout is usable but untuned. Each names the milestone that owns it; none is an
open M8 condition.
All three questions the M8 debt raised against M9 are settled and recorded:
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
cards are compatibility-only legacy data and no longer reach the narrator
(`IMPORTED-KNOWLEDGE-DESIGN.md` §73); and the deployment context ceiling is
documented in `DEVELOPMENT.md` with the note that an imported long campaign
meets it on its first turn rather than gradually. The window itself stays M11's.
**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`. Two of
its five items are now closed by M9 — story cards have a settled policy, and the
bundle carries the provenance. The RPG world state is still read-only, copy is
still per-message only, there is still no discarded-history recovery screen, and
the tablet layout is still untuned. Each names the milestone that owns it; none
is an open M8 condition.
The M6 retrieval debt this milestone was warned about is partly addressed and
partly still open. Imported material does **not** compete with story memory for
+83
View File
@@ -608,6 +608,89 @@ stated reason. A misplaced head affects every read in the file; a bookmark
pointing outside the story affects only itself, and rejecting a whole campaign
to protect one bookmark would lose the story to save the pointer.
### 9.3 As implemented in M9 — the version, and the third category
**The format is now `ai-dnd-adventure-v3`, and the bump is the design.** §9.1 and
§9.2 each declined one, correctly: an absent `headDepth` or `checkpoints` key is
unambiguous, because a file either states a position or it does not. That
property fails for what M9 adds. A v2 file with no prompt provenance may have
been written before M9, when no file could carry any, or by M9 from a campaign
whose turns predate the column — different facts about the campaign, and a
reader has to be able to tell them apart. A version number is how a recovery
file states what it was capable of recording, which is exactly the reasoning
§9.1 uses to justify opening a pre-M3 file at its tip. The reader keeps every
version; only the writer moved.
**§9.1's two categories became three.** "Chosen travels, derived is recomputed"
was sufficient until M9 had to decide about stored prompts, which are derived —
a machine assembled them — and must travel anyway:
```text
chosen the story, the head, the takes, the Save Points, the
classifications, the canon travels
evidence the state events and proposals, the per-turn prompt and the
passages it was shown, the model and generation settings that
turn ran under travels
rebuildable knowledge passages, the FTS index, embeddings, the branch
lineage cache rebuilt on import
```
The test separating the last two is not "could this be recomputed" but "would a
recomputation answer the same question". A rebuilt FTS index answers the same
question. A rebuilt prompt does not — it says what the turn *would be told now*,
from today's canon, today's sources and today's state, which is the opposite of
what the inspector is for. Historical evidence is not a cache, so M9 does not
regenerate one on import at any point.
`DATA-MODEL.md` §29 lists what v3 carries. Three implementation facts belong
here rather than there:
1. **The snapshots are encoded, not summarised.** A per-turn prompt contains the
story so far, so one per turn is O(turns²) — measured at 68% of a 9.7 MB file
at 120 turns, against a 20 MB import ceiling. The snapshot therefore travels
as `contextSnapshotZ`, zlib-compressed and base64-encoded through the same
`compression.pack`/`unpack` the database column already uses. The file is
still JSON and every other section of it is still plain text. The readable
`contextSnapshot` key is still accepted and wins when both are present, so a
hand-edited file keeps importing. A residual ceiling remains and is stated in
the M9 report rather than hidden.
2. **One more pointer is translated, and only pointers ever are.** Branch
numbers already were. A restored retrieval record's `source_id` names a row
on the machine that wrote the file, so it is repointed at the source that
landed here, or set to `null` when the file carries no such source. The text
the record holds — the evidence — is never rewritten.
3. **Import stays two-phase inside one transaction.** `plan` refuses everything
a hand-edited file can get wrong before a row exists; `materialize` writes,
and the endpoint commits once and rolls back explicitly otherwise. A
*rebuildable* index failing after that does not roll the campaign back: it is
reported on the response as a warning, shown per source in the Knowledge
panel, and repaired by Reindex. So a caller sees either "the campaign is not
there" or "the campaign is complete", never a third thing.
### 9.4 The database backup, as implemented in M9
A second recovery tool, deliberately not merged with the first. The bundle is a
logical, portable, human-readable copy of **one campaign** and is the supported
way to move a campaign between installations; the backup is a physical copy of
**this machine's whole database** and is what you take before an upgrade.
`backend/app/backup.py` uses SQLite's online backup API rather than a file copy,
because a copy taken while the application runs can read one page before a
transaction and another after it and produce a file that opens, reports a schema
and is quietly missing rows. It writes to a temporary name beside the
destination, runs `PRAGMA quick_check` against the finished file, and only then
renames it into place; it opens the source read-only, never overwrites an
existing backup, and leaves nothing behind on failure.
No path comes from a caller: the destination is derived from the database the
application already has open and the filename from the clock, so the endpoints
accept no body at all (H08).
**There is no restore endpoint, and that is a decision.** Restoring means
replacing the file the running process has open, which is how both copies are
lost at once. The procedure is in `DEVELOPMENT.md` and is a procedure precisely
because each step needs the application stopped.
## 10. Authoritative Narrative State
### 10.1 Do not retain the RPG state protocol as the product model
+114
View File
@@ -996,3 +996,117 @@ actual persistence / memory / canon bugs
The prose may vary.
The expected state, authority, and lineage rules should not.
---
# Appendix A — Proposed companion fixture: Multi-Character Identity Test
**Status: proposed, not built. This appendix changes nothing above it.**
## Why a companion rather than an extension
The Continuity Test above is the deterministic baseline that several milestones'
results are compared against. Adding characters or turns to it would invalidate
those comparisons, so **it is deliberately left exactly as it is**.
## The gap this fills
Reviewed on 2026-09-07 against a hands-on finding. The Continuity Test's seven
deliberate traps (§13) are:
```text
1-3 knowledge boundaries and secrets
4 reference authority
5 canon precedence
6 branch leakage
7 possession
```
There is **no identity trap**, the word *coreference* does not appear, and the
on-stage cast is effectively two people — Aldric and Mara, with Edrin
established as missing rather than present. So **same-scene multi-character
identity continuity is not exercised anywhere in the standard fixture.**
A play session against accepted M8 produced exactly that failure: four people in
one office, and narration that treated one of them as two different people
sharing a name. Root cause is unknown and no longer establishable — the playtest
database was destroyed — which is itself part of why a *deterministic* fixture
for this class is worth having.
## Shape
Four people, all present in one ordinary scene, with no fantasy vocabulary — the
point is identity, not genre:
```yaml
bill: { type: character, role: protagonist, controlled_by: reader }
alice: { type: character, role: coworker }
roger: { type: character, role: coworker }
john: { type: character, role: coworker }
location: { type: location, name: the office }
```
Identities and roles established unambiguously before the first test turn, so
that any later ambiguity is the system's and not the setup's.
## What the sequence must stress
- pronouns with more than one plausible referent in scene;
- dialogue attribution across three speakers;
- characters entering and leaving;
- reference by name **and** by role, for the same person;
- one character speaking *about* another;
- one character speaking about **themself in the third person**, which is the
shape the observed failure took.
## Traps
| | |
| --- | --- |
| **I1 — one Alice** | No turn may produce a second entity whose display name is `Alice`. The state model currently permits this and reports nothing, so the trap is real rather than theoretical |
| **I2 — the protagonist stays the protagonist** | Bill must not drift into being narrated as a third party, or acquire a second entity |
| **I3 — attribution** | A line spoken by Roger must not be attributed to John |
| **I4 — self-reference** | No character may refer to themself as a separate same-named person |
| **I5 — state and context agree** | The authoritative state and the assembled prompt must not disagree about who is present or who anyone is |
| **I6 — exit and return** | A character who leaves and returns is the same entity, not a new one |
## Required evidence on failure
Unlike the fixture above, this one exists to **classify** a failure rather than
only to detect it, because its failure modes are split between the application
and the model. Any failing turn must capture:
```text
authoritative state immediately before generation
the exact stored context/prompt snapshot
the recent-history section
summaries
retrieved memories
imported knowledge, if any
narrator output
model identifier and generation settings
```
then classify:
```text
STATE DEFECT
CONTEXT ASSEMBLY DEFECT
DERIVED MEMORY/SUMMARY DEFECT
MODEL FAILURE WITH CORRECT CONTEXT
AMBIGUOUS / MULTIPLE CONTRIBUTORS
```
Two rules for whoever runs it: **do not "fix" a model failure by editing
authoritative state**, and **do not blame the model when the prompt already
contained the identity error.**
Since M9 the whole of that evidence is portable in one campaign bundle, so a
failing run can be exported intact and investigated elsewhere.
## Ownership
M11, alongside the realistic-model review. `V1-ACCEPTANCE-TESTS.md` §P3 records
which parts of this are candidate **acceptance** criteria — the state-level
traps, which are decidable — and which are model-quality observations that
belong in a recorded review rather than in the pass/fail contract.
+244 -7
View File
@@ -1800,6 +1800,16 @@ Export standard campaign.
### Pass
Export completes locally and contains enough data to restore story.
### Result — PASS (M9, 2026-09-07)
`ai-dnd-adventure-v3`, checked section by section rather than by file size:
the story and its whole retained tree, the branches and their disposition, the
chosen head, Save Points, the authoritative state and its per-position
snapshots, the state events and proposals, the imported library, the
lineage-anchored summaries, the memories, and a stored prompt for every narrator
turn that has one. `test_m9_portability.py::test_i01_*`, and reproducible with
`python -m tools.m9_portability_report`, which classifies every data family as
PRESERVED, OMITTED or DERIVED/REBUILDABLE.
---
## I02 — Import Exported Campaign
@@ -1814,6 +1824,14 @@ Export completes locally and contains enough data to restore story.
### Pass
Active transcript and state are restored.
### Result — PASS (M9, 2026-09-07)
Across a **genuine clean data directory**: a second server process, in a second
directory, against a database file that has never existed, with the exporting
process stopped. Transcript, authoritative state, canon, Save Points, knowledge,
state events and the size of the retained tree all match the source campaign
(`test_m9_clean_import.py`). The same round trip inside one process is in
`test_m9_portability.py`, and is labelled there as the weaker of the two.
---
## I03 — Branch/Disposable History Export
@@ -1823,6 +1841,19 @@ Active transcript and state are restored.
### Pass
Export preserves retained alternate/disposable history needed for recovery, unless user explicitly chooses a trimmed export.
### Result — PASS (M9, 2026-09-07)
Both futures come back and stay distinguishable: the abandoned line's turns are
present and readable, the branch that was left carries the depth it was left at,
and a superseded take is still at its coordinate with the selected take still
selected. There is no trimmed export — M9 offers no option to drop history, so
the clause about one does not arise.
**M9 also closed a fidelity gap here.** Take parentage was not exported, so every
imported node landed parentless and the pager grouped on the coordinate instead.
That is right for a plain retry and wrong once two takes of one turn each have
takes of their own beneath them; the copy read `5/5` where the source read `2/2`
and `3/3`. v3 carries the parentage.
---
## I04 — Checkpoint Export
@@ -1841,6 +1872,15 @@ campaign. Importing Save Points does **not** move the active head — the head
still comes from the bundle's `headDepth`. Bundles written before M4 carry no
`checkpoints` key, import cleanly, and create none.
### Re-verified — PASS (M9, 2026-09-07)
Unchanged by the format bump, and extended in two directions. Every restored
Save Point resolves, restores to the position it names through M3's head
movement, leaves the retained history it moved back over intact, and the two
in the M9 fixture restore to *different* states. A Save Point whose coordinate
is not in the file is **dropped with the rest of the campaign kept**, never
retargeted to a nearby turn: the reader named a position, and if that position
is not in the file then no other position is the one they named.
---
## I05 — Knowledge Provenance Export
@@ -1870,13 +1910,29 @@ hand-edited knowledge block with an unknown classification or empty content
refuses the import rather than half-landing in it; and an edited content hash is
recomputed from what actually arrived and the discrepancy recorded on the source.
**Limit, unchanged from before M7 and owned by M9.** The bundle carries no
context snapshots at all, so an imported campaign has no historical prompt
provenance — for imported knowledge or for any other component. Nothing M7
creates is turned into a dangling id by a round trip, because no ids are
exported; the evidence simply is not in the file.
`test_historical_prompt_evidence_survives_an_export_round_trip` pins that
behaviour so it cannot regress silently.
**That limit is closed (M9, 2026-09-07).** The paragraph below is kept as
written because it records what was true through M7 and M8, and the M9 decision
is only legible against it.
> *The bundle carries no context snapshots at all, so an imported campaign has no
> historical prompt provenance — for imported knowledge or for any other
> component. Nothing M7 creates is turned into a dangling id by a round trip,
> because no ids are exported; the evidence simply is not in the file.*
M9 carries the snapshots. An old turn in a restored campaign shows the prompt it
was actually assembled from, the passages it was shown and the text each
supplied — after the source has been deleted, the canon edited and the state
moved on. `test_historical_prompt_evidence_survives_an_export_round_trip` was
inverted rather than deleted: it now pins the thing M7 was worried about and
could not check, which is that the provenance arriving on the other side names
*this* campaign's sources rather than the ids they had where the file was
written. Only that pointer is translated; the evidence is restored verbatim, and
a source the file does not carry becomes `null` rather than pointing at a
different file.
Also added in M9: `parserVersion` and `chunkingVersion` per source, recording
what produced the passages a historical retrieval record describes, and a
`sourceId` that exists only so the translation above can be made.
---
@@ -1887,6 +1943,28 @@ behaviour so it cannot regress silently.
### Pass
No external API credentials are embedded in campaign export.
### Result — PASS (M9, 2026-09-07)
Tested rather than assumed, and tested against the file's **text** rather than
against a list of columns — a field added to a model the exporter walks would
otherwise reach the bundle with no test of a column noticing. The inert
`api_key` column is written with a recognisable value first, so its absence is
evidence rather than a tautology.
Nothing matching `api_key`, `apiKey`, the written secret, the inference endpoint
or its port, or an absolute filesystem path appears anywhere in an export —
including inside the compressed snapshots, which the test decodes rather than
skipping. Checked in both suites, so the clean-directory run covers the same
ground across a real process boundary
(`test_m9_portability.py::test_i06_*`, `test_m9_clean_import.py`).
**What deliberately does not travel**, and why it is not an omission: the
inference endpoint, the model name, the context budget and every other row of
`settings`. Those describe the machine, not the campaign, and importing a
campaign must not silently repoint the destination's inference at the source's.
Per-turn model and generation settings *do* travel, inside the historical
snapshot, because there they are a record of what happened rather than a
configuration to apply.
---
## I07 — Export/Import Preserves an Undone Active Head
@@ -1921,6 +1999,28 @@ An export whose stated head lies beyond the story it contains is a file
disagreeing with itself and must be refused rather than opened at a guessed
position.
### Result — PASS (M9, 2026-09-07)
All three clauses, and the main one across a genuine machine boundary.
- **The exact head.** The M9 fixture ends two Undos behind its own branch's
retained tip and behind the abandoned line's, and the last thing it does is an
Undo — so the head is not the newest row written, not the deepest row, not the
tip, and not on the branch holding the most story. An importer guessing any one
of those lands somewhere else. The copy opens exactly where the source was,
the later turns are still in the database as retained future, and Redo is
offered rather than the story having silently been redone. Redo then walks to
the same next turn in both.
- **The legacy clause.** A bundle with its `headDepth` removed opens at the tip
of its head branch, offers no Redo, and offers Undo — which is the position
such a file recorded, because at the time it was written the head could not be
anywhere else. Checked at every seam in `test_m9_legacy_bundles.py`.
- **The self-disagreeing file.** A head past the retained story is refused with
a message naming where the branch actually ends, and nothing is written.
Evidence: `test_m9_clean_import.py::test_it_opens_at_the_exact_head_it_was_exported_at`
(second process, empty directory), plus `test_m9_portability.py::test_i07_*` and
the head cases in `test_m9_corrupt_bundles.py`.
---
# J. Genre Independence
@@ -2082,6 +2182,17 @@ turn.
### Pass
State at each position matches original accepted state.
### Result — PASS (M9, 2026-09-07)
Measured **after a round trip**, which is the M9 form of it: the copy and the
source are walked back three turns and forward three turns in step, and the
authoritative document is compared at every position. They agree throughout.
That this stays a snapshot read rather than a replay is the point. Undo, Redo
and Save Point restore all resolve a coordinate and read the state recorded
there (`TECHNICAL-DESIGN.md` §10.4), so an import that carried the events and
dropped the per-position snapshots would have made every one of them
proportional to campaign length. Both halves of §17's hybrid travel.
---
## L03 — Checkpoint Reconstruction After Restart
@@ -2105,6 +2216,15 @@ exactly that value after the campaign had been advanced past it
`TestClient` restart, which could not distinguish durable state from a live
object.
### Re-verified after a move — PASS (M9, 2026-09-07)
The same claim with a machine boundary in front of it. A campaign is exported
from one server process, imported into a **second process against a database
file that has never existed**, a Save Point is restored there, that process is
killed, and a **third** process against the same file is asked again. The
transcript, the authoritative state and the size of the retained tree all match
what the second process had after restoring
(`test_m9_clean_import.py::test_l03_*`).
---
## L04 — Derived Data Can Be Rebuilt
@@ -2121,6 +2241,34 @@ using a safe test copy.
### Pass
Authoritative campaign history remains intact and derived structures can be recreated.
### Result — PASS (M9, 2026-09-07), and one defect found by running it
On a safe copy — an imported campaign, not the original. Every physically
derived structure is destroyed and rebuilt from the source content the bundle
carried: passages, the FTS rows and the vectors. Afterwards every source is
`ready` with passages again, retrieval works, and the transcript, the
authoritative state, the classifications and the lifecycle flags are identical
either side. Deleting the vectors alone leaves lexical retrieval working, which
is M7's rule that the lexical half is a production path and not a fallback. A
rebuild does not make an abandoned line's summary eligible.
**Running it found a real defect, which is fixed here.** The FTS5 index is a
virtual table, so no foreign key reaches it and no `ON DELETE CASCADE` covers
it: deleting a campaign dropped its passages and left one index row per passage
behind. Nothing read them — every search joins through `knowledge_chunks` — so
the leak was invisible until SQLite handed the freed primary key out again, at
which point the **next source imported into any campaign** failed with an
integrity error. Reindex could not repair it either, because `clear_index` finds
index rows *through* the chunks, and there were none. Both ends are closed: a
campaign's index rows are removed before it is deleted, and the index insert now
replaces a stale row rather than colliding with it — so a database already
carrying the leak repairs itself and needs no migration. See the M9 report's
findings.
The rebuild path M9 depends on is therefore implemented and tested rather than
assumed, which the SHOULD priority above did not require but the milestone did:
knowledge passages and indexes are omitted from the bundle precisely because
they can be rebuilt.
---
# M. Long-Run Test
@@ -2327,3 +2475,92 @@ L01-L03
```
This keeps implementation work tied to observable behavior rather than repository-specific architecture.
---
# P. M11 Test-Design Tasks — Not Yet Acceptance Tests
**Nothing in this section is part of the pass/fail contract.** These are three
behaviours a real play session against accepted M8 showed are worth testing, for
which **the pass criterion is not yet settled**. They are recorded here so the
release tester finds them where they will look, and they are deliberately not
written as A-through-O items: giving them IDs would imply a criterion has been
ratified when it has not, and would destabilise a document whose value is that
every item in it is decidable.
Each names **what must be settled** before it can become an acceptance test.
Full context is in the M9 report's §Y and, durably, in `BUILD-MILESTONES.md`
under the post-M8 playtest findings.
## P1 — The reader can tell where they are after history movement
**From:** a real session in which Undo worked correctly and the reader could not
tell which point in the story they had reached.
**Behaviour to test:** after Undo, Redo, a Save Point restore, or an edit to an
earlier turn, the reader can identify their current position in the visible
story without implementation terminology (`branch`, `head`, `node`, `depth`
remain forbidden at the surface).
**Settle first:** what the indicator *is*. "The reader can tell" is not
decidable as written — it needs an observable artifact, such as a named position
that changes with movement and is present in the DOM. `BROWSER-UX-SPEC.md` §8
now carries the requirement; **no wording is ratified**, and this cannot become
an acceptance test before one is.
## P2 — Narration length has a measurable directional effect
**From:** a reader who chose *2-4 paragraphs* and received substantially longer
replies.
**Behaviour to test:** the narration-length setting produces a **measurable
directional difference** in output length across repeated realistic turns —
brief shorter than standard, standard shorter than detailed — on at least the
reference 3B narrator and one stronger local narrator.
**Settle first:** the numbers, and the mechanism. Today the setting adds one
English sentence to the campaign instructions and changes **no** generation
budget, while a separate numeric hint derived from the global
`max_output_tokens` is identical for every setting (M9 report §Y). Until the
product decides what each setting *means* — and whether it moves the budget —
there is no threshold to test against. A directional test is stateable; an
absolute one is not, and this should not become an acceptance test that asserts
word counts nobody has ratified.
**Do not** turn this into a truncation test: the state block is emitted last and
hard truncation removes it.
## P3 — Multi-character identity continuity
**From:** four people in one scene, and narration that treated one of them as
two different people of the same name. **Root cause unknown** — the playtest
database was destroyed, so no evidence survives.
**Behaviour to test:** across a multi-turn scene with a protagonist and three
supporting characters, the story does not create duplicate characters, does not
duplicate a display name across two entities, does not drift the protagonist's
identity, does not misattribute dialogue, and does not have a character refer to
themself as a separate same-named character — and the authoritative state and
the assembled context do not disagree about who anyone is.
**Settle first:** which of those are **product** guarantees and which are
**model-quality** observations. They are not the same kind of claim and must not
share one verdict:
- *"The state never holds two entities with the same display name"* is
decidable and enforceable, and is a candidate acceptance test today. The
implementation currently permits it and reports nothing (M9 report §Y).
- *"The narrator never confuses two same-named characters"* is not a pass/fail
property of this application — it depends on the model — and belongs in M11's
realistic-model review with a recorded classification, not in this contract.
**A run of this must capture**, on any failure: pre-generation state, the exact
stored prompt snapshot, history, summaries, retrieved memories, imported
knowledge, narrator output, and model settings — then classify as a state,
context-assembly, derived-data, or model failure. M9 made all of that portable,
so a failing campaign can be exported whole and investigated elsewhere.
**Fixture:** the standard Continuity Test does not exercise this — its traps are
knowledge, authority, branch leakage and possession, and its on-stage cast is
effectively two people. A companion fixture is proposed in
`TEST-CAMPAIGN-FIXTURE.md`; the established fixture is deliberately unchanged.
+96 -3
View File
@@ -1,8 +1,101 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v3.3
- **Revision date:** 2026-09-06
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted** (M8 closed out 2026-09-06). M9 is next and has not been started.
- **Package:** Adventure Storyteller Planning Package v3.5
- **Revision date:** 2026-09-07
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 implemented and awaiting independent review** (2026-09-07). M10 has not been started.
## v3.5 — Post-M8 hands-on playtest findings recorded (2026-09-07)
**Documentation only. No application code changed, and M9's verified result is
untouched** — see the note at the end of this entry.
A real play session against **accepted, signed M8** (real browser, trusted-LAN
Ollama, `qwen2.5:3b-instruct-16k`, disposable database since destroyed) surfaced
four product-quality observations. **None is an M9 defect, none was caused by
M9, and none blocks M9 acceptance.** They are recorded so they cannot be lost
when the M9 report is archived.
| Document | Change | Kind |
| --- | --- | --- |
| `reports/M9-IMPLEMENTATION-REPORT.md` | **New §Y**, the full write-up with the verification behind each mechanism, plus a pointer from §W. Explicitly labelled as not-M9. | observation record |
| `BUILD-MILESTONES.md` | **New section before M10**, the durable copy owned by M11; a note bounding M10 out of it; the four items added to M11's scope list. | milestone sequencing |
| `V1-ACCEPTANCE-TESTS.md` | **New §P**, three items as **test-design tasks, explicitly not acceptance tests**, each naming what must be settled before it could become one. No existing test changed or weakened. | test design |
| `BROWSER-UX-SPEC.md` | **New §8A.** §8's *"The current endpoint should be clear"* was satisfied while a real reader was lost, so it could not hold the behaviour. States the orientation requirement; **prescribes no wording**. | requirement clarification |
| `TEST-CAMPAIGN-FIXTURE.md` | **New Appendix A** proposing a companion `Multi-Character Identity Test`. The established deterministic fixture is **unchanged** — altering it would invalidate earlier milestones' comparisons. | test design |
**What was verified rather than assumed**, since the playtest campaign no longer
exists and root causes largely cannot be proven:
- The browser title genuinely is `AI D&amp;D` (`frontend/index.html`), and **no
accepted document ever claimed otherwise** — so this is an uncovered gap, not
documentation needing correction. Nothing was corrected and nothing was
renamed: *Adventure Storyteller* is itself narrower than the genre-agnostic
engine `SPECIFICATION.md` requires, and the naming decision is the owner's.
- The narration-length setting adds **one English sentence** and changes **no
generation budget**, while the numeric hint derived from the global
`max_output_tokens` is **identical for every setting** — measured at the
default as *"must not exceed 506 words, and it should not stop short of about
177."* Recorded as a mechanism to check first, **not** as the proven cause.
- The narrative state **permits two entities to share a display name and
reports nothing** — `DUPLICATE_ENTITY` rejects a repeated key only. That is
one of the identity finding's failure modes; it establishes nothing about what
actually happened.
**M9 is unaffected.** No application file changed in this pass, so M9's final
verified result stands exactly as recorded: **1,102 backend passed / 14 skipped
/ 0 failed**, 145 frontend, 36/36 browser. The expensive M9 suites were
deliberately **not** re-run, because only Markdown changed.
## v3.4 — M9 Implementation (2026-09-07)
M9 — Export, Backup, Recovery, and Migration Hardening — is implemented on
`m9-recovery` from the signed M8 commit `1ce9972`. This revision records what the
implementation settled. It is **not** an acceptance: the milestone report is
written for a reviewer and the tree is staged for the repository owner's signed
commit.
**Requirement corrections:** none. M9 altered no product requirement.
`SPECIFICATION.md` and `SECURITY-THREAT-MODEL.md` are unchanged — §16 already
required the export to preserve the exact active position, §6.1 already required
an exact prompt/context snapshot per turn, and M9 implements both rather than
redefining either.
| Document | Change | Kind |
| --- | --- | --- |
| `DATA-MODEL.md` §29 | The v3 format, the three-category rule (chosen / evidence / rebuildable), what each version can be trusted to say, the encoding of the snapshots, and the two pointers the import translates. | implementation fact |
| `TECHNICAL-DESIGN.md` §9.3 | Why the version was bumped when §9.1 and §9.2 each correctly declined one; the third data category; the two-phase transaction and the warning path for a failed derived rebuild. | implementation fact |
| `TECHNICAL-DESIGN.md` §9.4 | **New.** The SQLite backup: the online backup API rather than a file copy, the verify-then-rename order, and why there is no restore endpoint. | implementation fact |
| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **New subsection.** Story Cards settled as compatibility-only legacy data and removed from the narrator's prompt, with the evidence that they were the "alternate untracked path" §73 already forbade. | newly settled design decision |
| `V1-ACCEPTANCE-TESTS.md` I01-I07, L02-L04 | Results recorded. I05's M7-era limit is marked closed with the original paragraph kept, because the M9 decision is only legible against it. L04 records the defect running it found. | implementation fact |
| `BUILD-MILESTONES.md` M9 | Marked complete, with what it delivered, the three defects it found, the story-card decision, and the debt carried forward. | implementation fact |
| `README.md`, `VERSION.md` | Status. | implementation fact |
| `DEVELOPMENT.md` | **New section**: the two recovery tools and when each applies, taking a backup, and the stop-move-start restore procedure. Plus a note that an imported long campaign meets a small context ceiling on its first turn rather than gradually. | implementation fact |
**The four M8 handoff questions, answered**
| | Answer |
| --- | --- |
| **A. Complete campaign portability** | Every family travels and is measured family by family, before and after, by a tool a reviewer can rerun. |
| **B. Historical prompt provenance** | **It belongs in the bundle, and it is in it.** An old turn in a restored campaign shows what it was actually given, after the source has been deleted and the canon edited. |
| **C. Legacy story cards** | Compatibility-only. Carried in both directions; removed from the narrator's prompt; still the summariser's character roster. |
| **D. Context-window portability** | The campaign travels; the machine's model configuration does not. Importing changes no setting of the destination's, and a campaign imports whether or not any model is installed. The window itself remains M11's. |
**What the implementation found rather than assumed**
Three defects, all found by running the milestone's own tests rather than by
reading: an FTS index leak that made an ordinary import fail in an unrelated
campaign and that Reindex could not repair; an imported node with no state
snapshot being stamped with the campaign's *head* state; and a snapshot pointer
that was not being translated because the code mutated a dict in place. The
first two predate M9.
One measurement changed a plan, and then corrected the conclusion drawn from it.
Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured, and after compressing them inside the file, everything
M9 added costs **12%** of reachable campaign length — the import ceiling moves
from about 318 turns to about 279, against a 100-turn certification target. The
dominant cost is not M9's at all: the **per-position narrative state document is
74% of a bundle**, and v2 already carried it.
## v3.3 — M8 Closeout (2026-09-06)
File diff suppressed because it is too large Load Diff