M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
co-authored by
Claude Opus 5
parent
1ce9972760
commit
44edece67e
@@ -128,6 +128,35 @@ The current endpoint should be clear.
|
||||
|
||||
The input box always continues from the currently active story head.
|
||||
|
||||
### 8A. The reader must be able to tell where they are (recorded 2026-09-07)
|
||||
|
||||
A hands-on session against accepted M8 found the sentence above **too weak to
|
||||
hold the behaviour it names**. Undo worked correctly, the input box did continue
|
||||
from the active head — so both statements above were satisfied — and the reader
|
||||
still could not tell which point in the story they had moved to.
|
||||
|
||||
The requirement, stated so that a working implementation cannot satisfy it while
|
||||
a reader is lost:
|
||||
|
||||
> After Undo, Redo, a Save Point restore, an edit to an earlier turn, or any
|
||||
> other movement of the active story position, the reader should be able to
|
||||
> identify **where they now are in the visible story** — and, where it matters,
|
||||
> whether later story remains available ahead of them.
|
||||
|
||||
Two constraints on any solution:
|
||||
|
||||
- **No implementation terminology.** `branch`, `fork`, `node`, `head` and
|
||||
`depth` stay off the reader-facing surface (§38), which is what makes this a
|
||||
presentation problem rather than a labelling one.
|
||||
- **It must be observable**, not merely inferable from the transcript scrolling,
|
||||
so that a release test can decide it.
|
||||
|
||||
**No wording is prescribed here, and none is ratified.** A lightweight named
|
||||
position — `Moment 8` becoming `Moment 7` after an Undo, optionally noting that
|
||||
later story is available — is one candidate among others. Ownership is M11
|
||||
release polish; `V1-ACCEPTANCE-TESTS.md` §P1 records what must be settled before
|
||||
this can become an acceptance test.
|
||||
|
||||
## 9. User Turn Presentation
|
||||
|
||||
User messages should support:
|
||||
|
||||
@@ -1118,6 +1118,187 @@ Make campaigns portable and recoverable without losing lineage, state, knowledge
|
||||
|
||||
A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact.
|
||||
|
||||
## Status: COMPLETE — 2026-09-07, pending independent review
|
||||
|
||||
Implemented on `m9-recovery` from the signed M8 commit `1ce9972`, measured
|
||||
before and after against the same fixture, and verified in a real browser
|
||||
against a real narrator. `planning/reports/M9-IMPLEMENTATION-REPORT.md` is the
|
||||
implementer's account, written for a reviewer.
|
||||
|
||||
**What it delivered, beyond the scope list above:**
|
||||
|
||||
- **The bundle became a version, and the reason is a rule.** `ai-dnd-adventure-v3`.
|
||||
Everything M9 adds could have been an optional key, the way four earlier
|
||||
additions were — and that mechanism fails exactly here, because a v2 file with
|
||||
no prompt provenance is ambiguous between "written before M9" and "written by
|
||||
M9 from a campaign with none". A version is how a recovery file states what it
|
||||
was capable of recording. v1 and v2 are still read, and every seam from
|
||||
pre-active-head onward is tested.
|
||||
- **A third category of data.** "Chosen travels, derived is recomputed" was
|
||||
enough until stored prompts had to be decided. They are derived and must
|
||||
travel, so the rule is now chosen / **evidence** / rebuildable, and the test
|
||||
separating the last two is not "could this be recomputed" but "would a
|
||||
recomputation answer the same question".
|
||||
- **The M8 handoff on prompt provenance is closed.** An old turn in a restored
|
||||
campaign shows what it was actually given, after the source has been deleted,
|
||||
the canon edited and the state moved on.
|
||||
- **State events and proposals travel**, so a moved campaign can still say why
|
||||
its state is what it is — and a manual correction is still identifiable as
|
||||
one, which it was not before.
|
||||
- **Summaries travel with their coordinates**, so an abandoned line's summary is
|
||||
still ineligible after the move, and a moved campaign resumes with its
|
||||
long-story continuity instead of behaving like a new one.
|
||||
- **A verified SQLite backup**, using the online backup API rather than a file
|
||||
copy, taken while the application is running, with a browser control in
|
||||
Settings.
|
||||
- **Take parentage**, memory authority and parser/chunking versions, each
|
||||
closing a smaller fidelity loss.
|
||||
|
||||
**Three defects found by running the milestone's own tests, and fixed here:**
|
||||
|
||||
1. **Deleting a campaign leaked its FTS index rows**, and SQLite then handed the
|
||||
freed ids to the *next* source imported into *any* campaign, which failed
|
||||
with an integrity error. Reindex could not repair it either. Pre-existing
|
||||
since M7; both ends are now closed and an already-damaged database repairs
|
||||
itself with no migration.
|
||||
2. **An imported node with no state snapshot was stamped with the campaign's
|
||||
head state**, so an Undo to turn 2 showed what the story knew at turn 20 —
|
||||
M5's review finding 3, arriving through the import.
|
||||
3. **A snapshot's `source_id` was not being translated on import** because the
|
||||
relink mutated a dict in place, which a non-`MutableDict` column does not
|
||||
notice. Found by a test asserting the outcome rather than the call.
|
||||
|
||||
**One decision the brief asked for, made and recorded:**
|
||||
|
||||
**Story cards are compatibility-only legacy data, and no longer enter the
|
||||
narrator's prompt.** They still travel in both directions, the rows and the API
|
||||
stay, and `memorybank.cast_brief` still reads them as the summariser's character
|
||||
roster. What stops is the injection: a keyword-matched card arrived in front of
|
||||
the narrator as `World Lore: …` with no class, no visibility, no source, no way
|
||||
to switch it off since M8 removed the editor, and no row in the context
|
||||
inspector — which is `IMPORTED-KNOWLEDGE-DESIGN.md` §73's "alternate untracked
|
||||
path around the new knowledge authority/provenance rules" in as many words.
|
||||
|
||||
**Debt carried forward, deliberately:**
|
||||
|
||||
- **A long campaign's bundle has a measured ceiling: ~279 turns** against the
|
||||
20 MB import limit. A per-turn prompt contains the story so far, so carrying
|
||||
one per turn is O(turns²); compressing them inside the file cut that to about
|
||||
an eighth of what it would have been. **M9's additions account for only 12% of
|
||||
the ceiling.** The other 88% is the per-position narrative state document,
|
||||
which is 74% of a bundle and which v2 already carried — so lifting the ceiling
|
||||
means addressing that, not the evidence. Far beyond M11's 100-turn
|
||||
certification (14% of the cap), and stated with its measurement rather than
|
||||
hidden. A streaming or chunked import is the fix if a later milestone needs
|
||||
one; the asymmetry to know about is that such a campaign can still be exported
|
||||
and would be refused on import.
|
||||
- **No discarded-history recovery screen** (§63). M9's job was that retained
|
||||
history survives correctly so a later screen can use it; it does.
|
||||
- **No whole-transcript copy and no story search** (§77, §78) — still M11 or
|
||||
later.
|
||||
|
||||
---
|
||||
|
||||
# Post-M8 hands-on playtest findings — recorded 2026-09-07, owned by M11
|
||||
|
||||
**Not M9 defects, not caused by M9, and they did not block M9 acceptance.**
|
||||
Recorded here rather than only in the M9 report because a milestone report is
|
||||
archived when the next one replaces it, and these must not go with it.
|
||||
|
||||
They come from a real play session against **accepted, signed M8**: real
|
||||
browser, real trusted-LAN Ollama, narrator `qwen2.5:3b-instruct-16k`, disposable
|
||||
isolated campaign database. **That database was deliberately destroyed
|
||||
afterwards**, so the stored context snapshot for finding D is gone and no root
|
||||
cause is claimed for it. Full write-up, with the verification behind each
|
||||
mechanism, is in the M9 report's **§Y** (in `archive/milestone-reports/` once
|
||||
M10's report replaces it).
|
||||
|
||||
**These are M11's, and explicitly not M10's.** M10 is media-readiness
|
||||
architecture and stays bounded; it inherits them as known carry-forward only.
|
||||
|
||||
### A. The browser still calls the product "AI D&D" — M11 release polish
|
||||
|
||||
`frontend/index.html` still carries the inherited `<title>AI D&D</title>`.
|
||||
M8 changed the navigation, the inspector and the screens, and never claimed the
|
||||
title — so this is an uncovered gap rather than a false claim.
|
||||
|
||||
**Do not fix it with a find-and-replace to "Adventure Storyteller".**
|
||||
`SPECIFICATION.md` requires a genre-agnostic engine, and *Adventure* is narrower
|
||||
than the product. The naming decision is the repository owner's; a neutral
|
||||
working name such as **Interactive Story** and a tab form such as
|
||||
`<Campaign Name> — Interactive Story` are candidates, not decisions.
|
||||
|
||||
### B. After Undo, the reader cannot tell where they are — M11 UX polish
|
||||
|
||||
Undo behaved correctly (M3 semantics; re-verified throughout M9). The reader
|
||||
could not tell **which point in the story** they had moved to.
|
||||
|
||||
`BROWSER-UX-SPEC.md` §8 said only *"The current endpoint should be clear"*, and
|
||||
its companion sentence was already true while the reader was lost — so the
|
||||
requirement could not hold the behaviour. §8 has been strengthened to state the
|
||||
orientation requirement; **no UI text is prescribed**, because none is ratified.
|
||||
A `Moment 8` → `Moment 7` style indicator is a candidate. Needs a browser
|
||||
regression scenario.
|
||||
|
||||
### C. Narration length has no measurable effect — M11 realistic-model behaviour
|
||||
|
||||
The setup choice becomes **one English sentence** in the campaign's
|
||||
`ai_instructions` and changes **no generation setting**. Independently,
|
||||
`length_hint()` derives a numeric word range from the **global**
|
||||
`Settings.max_output_tokens` and places it after the history — and at the default
|
||||
800 it reads *"must not exceed 506 words, and it should not stop short of about
|
||||
177"* **identically for brief, medium and long**.
|
||||
|
||||
That is a mechanism, verified by reading and running the code — **not a proven
|
||||
cause** of what the reader saw; the narrator's instruction following is also in
|
||||
play. A reproduction must measure what enters the stored prompt, whether the
|
||||
setting moves any generation budget, and actual word/paragraph counts across
|
||||
repeated turns, on the reference 3B narrator **and** a stronger local one.
|
||||
Clearer numeric targets (`Brief ~100-200 words` and so on) are a design
|
||||
candidate, not ratified. **Do not hard-truncate prose** — the state block is
|
||||
emitted last and truncation removes it.
|
||||
|
||||
### D. Character identity / coreference confusion — M11 diagnostic
|
||||
|
||||
Four people in one scene — Bill (protagonist), Roger, John, Alice — and later
|
||||
narration treated Alice as two different Alices.
|
||||
|
||||
**Root cause UNKNOWN and no longer establishable.** Candidates: a model
|
||||
coreference failure on a correct prompt; duplicate/conflicting state; a
|
||||
context/summary/memory assembly failure; or a context that is not contradictory
|
||||
but too implicit for a small model.
|
||||
|
||||
**One structural fact to check first**, verified by reading the code: the
|
||||
narrative state **permits two entities to share a display name and reports
|
||||
nothing**. Entities are keyed by the model-supplied id; `DUPLICATE_ENTITY`
|
||||
rejects only a repeated *key*; no check exists on `name`. That is one of this
|
||||
finding's failure modes, and establishes nothing about what happened.
|
||||
|
||||
**M11 must run an explicit diagnostic** with a protagonist and three same-scene
|
||||
supporting characters, stressing pronouns, dialogue attribution, entrances and
|
||||
exits, reference by name and by role, and one character speaking about another.
|
||||
It must detect duplicate creation, same-name duplication, protagonist drift,
|
||||
misattributed dialogue, self-as-other reference, and state/context disagreement
|
||||
— and on any failure preserve the pre-generation state, the exact stored prompt
|
||||
snapshot, history, summaries, memories, imported knowledge, narrator output and
|
||||
model settings, then classify:
|
||||
|
||||
```text
|
||||
STATE DEFECT / CONTEXT ASSEMBLY DEFECT / DERIVED MEMORY-SUMMARY DEFECT /
|
||||
MODEL FAILURE WITH CORRECT CONTEXT / AMBIGUOUS
|
||||
```
|
||||
|
||||
**Do not "fix" a model failure by changing authoritative state, and do not blame
|
||||
the model if the prompt already contained the error.** M9 made all of that
|
||||
evidence portable, so a failing campaign can be exported and handed over intact.
|
||||
|
||||
**The standard fixture does not cover this class.** `TEST-CAMPAIGN-FIXTURE.md`'s
|
||||
seven traps are knowledge, authority, branch leakage and possession; there is no
|
||||
identity trap and its on-stage cast is effectively two people. The established
|
||||
fixture was **not modified** — it is the deterministic baseline earlier results
|
||||
are compared against. A companion fixture, `Multi-Character Identity Test`, is
|
||||
proposed in an appendix to that document.
|
||||
|
||||
---
|
||||
|
||||
# M10 — Future Media Extension Hooks Only
|
||||
@@ -1126,6 +1307,13 @@ A campaign can be safely exported, imported into a clean data directory, and reo
|
||||
|
||||
Preserve the approved future media interfaces without adding a media-generation dependency to v1.
|
||||
|
||||
## Note — the post-M8 playtest findings are **not** M10 scope
|
||||
|
||||
The four findings recorded above are owned by M11. M10 inherits them as known
|
||||
carry-forward items only: it should neither implement nor test them, and its
|
||||
scope below is unchanged by them. They are listed before this milestone rather
|
||||
than after it only because they were recorded during M9's closeout.
|
||||
|
||||
## Scope
|
||||
|
||||
- scene snapshots/packets suitable for future providers,
|
||||
@@ -1178,7 +1366,11 @@ Validate the full product against the release contract after all functional mile
|
||||
- fantasy and science-fiction fixtures,
|
||||
- export/import/recovery tests,
|
||||
- migration tests,
|
||||
- documentation and packaging.
|
||||
- documentation and packaging,
|
||||
- **the four post-M8 hands-on playtest findings above**: the browser product
|
||||
name, reader orientation after history movement, a narration-length setting
|
||||
with a measurable effect, and the multi-character identity diagnostic with its
|
||||
companion fixture.
|
||||
|
||||
## Tests / Acceptance
|
||||
|
||||
|
||||
@@ -887,6 +887,101 @@ has stopped being told the rules, with nothing to notice.
|
||||
A bundle written before M7 has no knowledge section and imports with an empty
|
||||
library, which is what such a campaign had.
|
||||
|
||||
### M9: the format became a version, and the evidence started travelling
|
||||
|
||||
**M9 bumped the format to `ai-dnd-adventure-v3`**, and the reason is a rule
|
||||
rather than a preference. Everything M9 added *could* have been an optional key
|
||||
read with `.get`, the way `persona`, `checkpoints`, `narrativeState` and
|
||||
`knowledge` each were. That mechanism stops working at exactly this addition:
|
||||
a v2 file carrying no prompt provenance is **ambiguous** — written before M9,
|
||||
when no file could carry one, or by M9 from a campaign whose turns predate the
|
||||
column? Those are different facts about the campaign and a reader has to be able
|
||||
to tell them apart. It is the same distinction the head rule above draws when it
|
||||
says a pre-M3 file opens at its tip *because that is the position such a file
|
||||
recorded*. A version number is how a recovery file states what it was capable of
|
||||
recording. The reader keeps every older version; only the writer moved.
|
||||
|
||||
What each version can be trusted to say:
|
||||
|
||||
```text
|
||||
v1 a linear story, its turns, and its retries as a repeating group
|
||||
v2 + the tree, the live flags, the after-snapshots, the chosen head,
|
||||
Save Points, the narrative state document, imported knowledge
|
||||
v3 + state events and proposals, historical prompt/context provenance,
|
||||
lineage-anchored summaries, take parentage, memory authority
|
||||
```
|
||||
|
||||
**Three categories, not two.** §31 below distinguishes authoritative from
|
||||
derived, which was sufficient until M9 had to decide about stored prompts. They
|
||||
are derived — a machine assembled them — and they must travel anyway, so the
|
||||
rule the bundle applies has a middle category:
|
||||
|
||||
```text
|
||||
chosen what a person decided: the story, the head, the takes, the
|
||||
Save Points, the classifications, the canon. travels
|
||||
evidence what happened, and what the application was told at the time:
|
||||
the state events and proposals, the per-turn prompt and the
|
||||
passages it was shown, the model and generation settings that
|
||||
turn ran under. travels
|
||||
rebuildable a deterministic function of what travels: knowledge passages,
|
||||
the FTS index, embeddings, the branch lineage cache.
|
||||
rebuilt on import
|
||||
```
|
||||
|
||||
The test that separates evidence from rebuildable is **not** "could this be
|
||||
recomputed" but "would a recomputation answer the same question". Rebuilding the
|
||||
FTS index answers the same question it answered before. Rebuilding an old turn's
|
||||
prompt does not — it would say what that turn *would be told now*, from today's
|
||||
canon, today's sources and today's state, which is the opposite of what the
|
||||
context inspector is for. Historical evidence is not a cache.
|
||||
|
||||
So v3 additionally carries, all restored verbatim:
|
||||
|
||||
- **`stateEvents` and `stateProposals`.** §17's hybrid keeps the events for
|
||||
audit and the snapshots for restore; v2 carried only the snapshots, so a moved
|
||||
campaign could be read at any position and could no longer say what changed
|
||||
there, who asserted it, or what the value was before. A **manual correction**
|
||||
was the worst case: the one state change no narration explains, and with the
|
||||
events gone nothing distinguished it from something the story established.
|
||||
Both tables travel, because the inspector reads both — the event says what was
|
||||
accepted and the proposal says what the model asked for and what was refused.
|
||||
- **A per-node context snapshot**, which is `SPECIFICATION.md` §6.1's exact
|
||||
prompt/context record. It carries the assembled prompt section by section, the
|
||||
passages retrieved with the text each supplied, which summary was eligible,
|
||||
and the model and generation settings the call ran under — so an old turn can
|
||||
still say what it was told after the source was deleted, the canon edited, the
|
||||
chunker changed and the state moved on. Stored once per turn on the live
|
||||
attempt, so a retried turn is not a multiplier.
|
||||
- **`summaries`**, with the coordinate that decides eligibility. v2 carried only
|
||||
the `storySummary` mirror, which has no lineage of its own, so a restored
|
||||
campaign resumed with no usable long-story continuity — and a summary
|
||||
belonging to an abandoned line stays ineligible after the move for the same
|
||||
reason it was before it: eligibility is the coordinate lying on the active
|
||||
capped lineage, not a stored flag.
|
||||
- **Take parentage**, so attempts under two different takes of one turn stay two
|
||||
pagers rather than merging into one.
|
||||
- **Memory `authority`**, so a heuristic memory is not promoted to accepted
|
||||
story by being moved (F07).
|
||||
- **Per-source `parserVersion`/`chunkingVersion`**, recording what produced the
|
||||
passages a historical retrieval record describes.
|
||||
|
||||
**Encoding.** A per-turn prompt contains the story so far, so one per turn is
|
||||
O(turns²) in campaign length — measured at 20,797 bytes per turn at turn 20 and
|
||||
55,291 at turn 120, 68% of a 9.7 MB file. The snapshot therefore travels as
|
||||
`contextSnapshotZ`: the same JSON, zlib-compressed and base64-encoded, using the
|
||||
same pack/unpack the database column already uses. Nothing is dropped or
|
||||
summarised; the file is still JSON, and every other section of it is still plain
|
||||
text. The plain `contextSnapshot` key is still read and takes precedence, so a
|
||||
hand-edited file keeps importing.
|
||||
|
||||
**Two pointers are translated on import, and nothing else is.** Branch numbers
|
||||
already were. M9 adds the `source_id` inside a restored retrieval record: it
|
||||
names a row on the machine that wrote the file, so left alone it would point the
|
||||
inspector's "open this source" at whatever holds that id here. Where the file's
|
||||
own knowledge section contains the source it is repointed; where it does not — a
|
||||
source deleted before the export — it becomes `null`, and the record keeps its
|
||||
text and filename. The evidence is never rewritten; only the pointer is.
|
||||
|
||||
## 30. Deletion vs Archival
|
||||
|
||||
The system must distinguish:
|
||||
@@ -916,6 +1011,23 @@ Retry, Undo, Restore, and branch switching must not silently perform permanent d
|
||||
|
||||
Derived data should be rebuildable where practical.
|
||||
|
||||
**"Where practical" does real work in that sentence, and M9 had to split this
|
||||
list to act on it** (§29). Embeddings and the lexical and semantic indexes are
|
||||
deterministic functions of content that travels, so a rebuild answers the same
|
||||
question and they are not exported. A **summary** is not: it took a model call,
|
||||
it describes a stretch of story that may since have been abandoned, and
|
||||
regenerating one on another machine produces different prose about a different
|
||||
reading — so it is derived, not practically rebuildable, and it travels with the
|
||||
coordinate that decides whether it still applies.
|
||||
|
||||
The same reasoning puts **stored prompt/context snapshots** on the travelling
|
||||
side, and they are not in either list above because they are neither: they are
|
||||
not authoritative — nothing decides anything from them — and calling them
|
||||
derived would invite a rebuild. They are *evidence*: a record of what the
|
||||
application was told at the time, which a regeneration would not reproduce
|
||||
because it would use today's canon, today's sources and today's state. §29 states
|
||||
the three-way rule the export applies.
|
||||
|
||||
## 32. Provenance
|
||||
|
||||
Important information should answer:
|
||||
|
||||
@@ -1110,6 +1110,47 @@ Normal imported files are campaign-level source material and need not inherit st
|
||||
|
||||
Story Cards may remain as an inherited authored-rule/lore primitive during migration if useful, but they must not become an alternate untracked path around the new knowledge authority/provenance rules.
|
||||
|
||||
### Settled in M9 (2026-09-07): compatibility-only, and out of the prompt
|
||||
|
||||
M8 removed the Story Card browser editor and left the question open; the M9
|
||||
brief asked for it to be decided. The finding was that story cards **were** the
|
||||
alternate untracked path the paragraph above forbids, and not in principle: a
|
||||
keyword-matched card was injected into the narrator's prompt as
|
||||
`World Lore: <entry>`, taking up to 40% of what was left after the imported
|
||||
knowledge had been placed, with
|
||||
|
||||
- no class, so nothing framed how far the narrator could rely on it;
|
||||
- no visibility, so no narrator-only distinction existed;
|
||||
- no source, no hash and no lifecycle, so there was nothing to disable;
|
||||
- no browser surface after M8, so a reader could neither see nor switch it off;
|
||||
- no row in the context inspector, which renders `knowledge` and never rendered
|
||||
`cards`;
|
||||
|
||||
and competing with imported Canon for one budget, which is the arrangement M7
|
||||
spent a milestone separating.
|
||||
|
||||
**The decision, and it is the smallest change that closes it:**
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| New export | carries them, unchanged, under `storyCards` |
|
||||
| Legacy import | accepted, unchanged, from every format version |
|
||||
| Normal narration | **no longer reached.** The `world_lore` section is gone |
|
||||
| Re-export | carries them again, so a round trip destroys nothing |
|
||||
|
||||
Nothing is deleted. The rows stay, the `/api/story-cards` endpoints stay, and
|
||||
`memorybank.cast_brief` still reads them as the **summariser's character
|
||||
roster** — that names who is on stage so a memory says "Aldric" rather than
|
||||
"he", never reaches the narrator, and every memory written from it is
|
||||
authority-classified by the application afterwards. The `cards` key stays in the
|
||||
context report and is now always empty for a new turn, because M9 made
|
||||
historical snapshots portable and an old turn's record must go on saying that
|
||||
story cards were included.
|
||||
|
||||
A campaign that wants the narrator to know something imports it as Canon,
|
||||
Reference or Inspiration, where it is classified, inspectable, disableable and
|
||||
attributable — which is what §73 asks for.
|
||||
|
||||
### Retrieval implementation direction
|
||||
|
||||
Use:
|
||||
|
||||
+48
-36
@@ -3,7 +3,8 @@
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
milestones **M1 through M6 implemented and accepted** (M3 and M4: 2026-09-03;
|
||||
milestones **M1 through M8 implemented and accepted**, and **M9 implemented and
|
||||
awaiting review**. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
|
||||
independent review found a real defect and a corrective pass fixed it.
|
||||
|
||||
@@ -19,8 +20,15 @@ carries the closeout: the build-evidence classification in its §P, finding 14's
|
||||
operational resolution, and the acceptance record in its §V. The M8 tree is
|
||||
staged and awaits the repository owner's signed commit.
|
||||
|
||||
**Next: M9 — Export, Backup, Recovery, and Migration Hardening.** It has not
|
||||
been started.
|
||||
**M9 — Export, Backup, Recovery, and Migration Hardening — is implemented and
|
||||
awaiting independent review** (2026-09-07).
|
||||
`reports/M9-IMPLEMENTATION-REPORT.md` is the implementer's account, written for
|
||||
a reviewer: a set of claims with the measurements attached, not yet a record of
|
||||
acceptance. M8's report has moved to `archive/milestone-reports/`, which is
|
||||
where a milestone report goes once the next milestone's report replaces it.
|
||||
|
||||
**Next: M10 — Future Media Extension Hooks Only.** It has not been started, and
|
||||
no brief for it exists.
|
||||
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
and why.
|
||||
@@ -82,8 +90,8 @@ Two standing qualifications:
|
||||
| Document | What it is for |
|
||||
| --- | --- |
|
||||
| `SPECIFICATION.md` | What the product must do. The top of the authority order. |
|
||||
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M7 built, recorded as fact. |
|
||||
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the export shape. |
|
||||
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M9 built, recorded as fact. |
|
||||
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the v3 export contract. |
|
||||
| `STORY-BRANCH-SEMANTICS.md` | Undo/Redo/Retry/branch/take behavior, including the M3 ratifications. |
|
||||
| `CONTEXT-AND-MEMORY.md` | Prompt assembly, summarization, branch-safe memory. |
|
||||
| `IMPORTED-KNOWLEDGE-DESIGN.md` | Canon / Reference / Inspiration knowledge as a first-class subsystem. |
|
||||
@@ -110,7 +118,7 @@ Two standing qualifications:
|
||||
10. `BROWSER-UX-SPEC.md`
|
||||
11. `V1-ACCEPTANCE-TESTS.md`
|
||||
12. `DECISIONS/` — all of them; they are short.
|
||||
13. `reports/M8-IMPLEMENTATION-REPORT.md`, for what the most recent milestone
|
||||
13. `reports/M9-IMPLEMENTATION-REPORT.md`, for what the most recent milestone
|
||||
actually left behind — reading it as a claim to check, not a record, until
|
||||
it is reviewed. Nothing in `planning/archive/` unless sent there.
|
||||
|
||||
@@ -143,20 +151,18 @@ work until Phase 0 closes — which Phase 0 satisfied on 2026-09-01. It is in
|
||||
`reports/` holds the report for the milestone most recently completed, because
|
||||
that is the one the next milestone's planning has to consult:
|
||||
|
||||
- `reports/M8-IMPLEMENTATION-REPORT.md` — the M8 implementation, its baseline
|
||||
UX measurement, and the browser evidence for every acceptance test it claims.
|
||||
Written by the implementer for an independent reviewer, and completed at
|
||||
closeout after that review accepted the milestone: it is a set of claims with
|
||||
the measurements attached **and** the record of the acceptance. Its §U carries
|
||||
the M9 handoff — the four questions the next brief has to decide.
|
||||
- `reports/M9-IMPLEMENTATION-REPORT.md` — the M9 implementation: the measured M8
|
||||
portability baseline it started from, the final bundle contract, and the
|
||||
evidence for every acceptance test it claims. Written by the implementer for
|
||||
an independent reviewer, so it is a set of claims with the measurements
|
||||
attached and **not** a record of acceptance. Its §W carries the M10-M11
|
||||
handoff.
|
||||
|
||||
**It stays here until M9's report replaces it.** A milestone report is useful
|
||||
during the immediately following milestone; M8's is not archived merely
|
||||
because M8 is accepted.
|
||||
**It stays here until M10's report replaces it.**
|
||||
|
||||
Completed earlier milestones are in `archive/milestone-reports/`, which M7's
|
||||
report joined when M8's was written: a milestone report is useful during the
|
||||
immediate next milestone and historical afterwards. M1-M7 are all there,
|
||||
Completed earlier milestones are in `archive/milestone-reports/`, which M8's
|
||||
report joined when M9's was written: a milestone report is useful during the
|
||||
immediate next milestone and historical afterwards. M1-M8 are all there,
|
||||
unedited.
|
||||
|
||||
## The decision this package rests on
|
||||
@@ -282,33 +288,39 @@ Milestone M8 COMPLETE / ACCEPTED (2026-09-06)
|
||||
story operations review + closeout, in sequence
|
||||
|
|
||||
v
|
||||
Milestone M9 NEXT — not started
|
||||
export, backup, recovery, see BUILD-MILESTONES.md
|
||||
migration hardening
|
||||
Milestone M9 COMPLETE — awaiting review (2026-09-07)
|
||||
export, backup, recovery, reports/M9-IMPLEMENTATION-REPORT.md
|
||||
migration hardening bundle format v3; SQLite online backup
|
||||
|
|
||||
v
|
||||
M10-M11, one at a time see BUILD-MILESTONES.md
|
||||
Milestone M10 NEXT — not started
|
||||
future media extension hooks see BUILD-MILESTONES.md
|
||||
|
|
||||
v
|
||||
Milestone M11 see BUILD-MILESTONES.md
|
||||
```
|
||||
|
||||
## Stop Rule
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
|
||||
**No M9 brief has been prepared.** Writing one is the current action, informed
|
||||
by the M8 report and by the debt `BUILD-MILESTONES.md` records against M8 — in
|
||||
particular that the campaign bundle still carries no context snapshots, so an
|
||||
imported campaign has no historical prompt provenance; that story cards survive
|
||||
in the backend and the bundle with no browser surface, and M9 should decide
|
||||
deliberately whether the bundle keeps carrying them; and that a deployment whose
|
||||
Ollama enforces a small context window truncates an imported long campaign
|
||||
immediately unless the `DEVELOPMENT.md` procedure or a matching
|
||||
`context_token_budget` is applied.
|
||||
**No M10 brief has been prepared**, and M9 is not accepted — it is implemented
|
||||
and awaiting an independent review. Writing the M10 brief is the action after
|
||||
that review closes, informed by the M9 report's §W.
|
||||
|
||||
**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`: story
|
||||
cards have no browser editor, the RPG world state is read-only, copy is
|
||||
per-message only, there is no discarded-history recovery screen, and the tablet
|
||||
layout is usable but untuned. Each names the milestone that owns it; none is an
|
||||
open M8 condition.
|
||||
All three questions the M8 debt raised against M9 are settled and recorded:
|
||||
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
|
||||
cards are compatibility-only legacy data and no longer reach the narrator
|
||||
(`IMPORTED-KNOWLEDGE-DESIGN.md` §73); and the deployment context ceiling is
|
||||
documented in `DEVELOPMENT.md` with the note that an imported long campaign
|
||||
meets it on its first turn rather than gradually. The window itself stays M11's.
|
||||
|
||||
**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`. Two of
|
||||
its five items are now closed by M9 — story cards have a settled policy, and the
|
||||
bundle carries the provenance. The RPG world state is still read-only, copy is
|
||||
still per-message only, there is still no discarded-history recovery screen, and
|
||||
the tablet layout is still untuned. Each names the milestone that owns it; none
|
||||
is an open M8 condition.
|
||||
|
||||
The M6 retrieval debt this milestone was warned about is partly addressed and
|
||||
partly still open. Imported material does **not** compete with story memory for
|
||||
|
||||
@@ -608,6 +608,89 @@ stated reason. A misplaced head affects every read in the file; a bookmark
|
||||
pointing outside the story affects only itself, and rejecting a whole campaign
|
||||
to protect one bookmark would lose the story to save the pointer.
|
||||
|
||||
### 9.3 As implemented in M9 — the version, and the third category
|
||||
|
||||
**The format is now `ai-dnd-adventure-v3`, and the bump is the design.** §9.1 and
|
||||
§9.2 each declined one, correctly: an absent `headDepth` or `checkpoints` key is
|
||||
unambiguous, because a file either states a position or it does not. That
|
||||
property fails for what M9 adds. A v2 file with no prompt provenance may have
|
||||
been written before M9, when no file could carry any, or by M9 from a campaign
|
||||
whose turns predate the column — different facts about the campaign, and a
|
||||
reader has to be able to tell them apart. A version number is how a recovery
|
||||
file states what it was capable of recording, which is exactly the reasoning
|
||||
§9.1 uses to justify opening a pre-M3 file at its tip. The reader keeps every
|
||||
version; only the writer moved.
|
||||
|
||||
**§9.1's two categories became three.** "Chosen travels, derived is recomputed"
|
||||
was sufficient until M9 had to decide about stored prompts, which are derived —
|
||||
a machine assembled them — and must travel anyway:
|
||||
|
||||
```text
|
||||
chosen the story, the head, the takes, the Save Points, the
|
||||
classifications, the canon travels
|
||||
evidence the state events and proposals, the per-turn prompt and the
|
||||
passages it was shown, the model and generation settings that
|
||||
turn ran under travels
|
||||
rebuildable knowledge passages, the FTS index, embeddings, the branch
|
||||
lineage cache rebuilt on import
|
||||
```
|
||||
|
||||
The test separating the last two is not "could this be recomputed" but "would a
|
||||
recomputation answer the same question". A rebuilt FTS index answers the same
|
||||
question. A rebuilt prompt does not — it says what the turn *would be told now*,
|
||||
from today's canon, today's sources and today's state, which is the opposite of
|
||||
what the inspector is for. Historical evidence is not a cache, so M9 does not
|
||||
regenerate one on import at any point.
|
||||
|
||||
`DATA-MODEL.md` §29 lists what v3 carries. Three implementation facts belong
|
||||
here rather than there:
|
||||
|
||||
1. **The snapshots are encoded, not summarised.** A per-turn prompt contains the
|
||||
story so far, so one per turn is O(turns²) — measured at 68% of a 9.7 MB file
|
||||
at 120 turns, against a 20 MB import ceiling. The snapshot therefore travels
|
||||
as `contextSnapshotZ`, zlib-compressed and base64-encoded through the same
|
||||
`compression.pack`/`unpack` the database column already uses. The file is
|
||||
still JSON and every other section of it is still plain text. The readable
|
||||
`contextSnapshot` key is still accepted and wins when both are present, so a
|
||||
hand-edited file keeps importing. A residual ceiling remains and is stated in
|
||||
the M9 report rather than hidden.
|
||||
2. **One more pointer is translated, and only pointers ever are.** Branch
|
||||
numbers already were. A restored retrieval record's `source_id` names a row
|
||||
on the machine that wrote the file, so it is repointed at the source that
|
||||
landed here, or set to `null` when the file carries no such source. The text
|
||||
the record holds — the evidence — is never rewritten.
|
||||
3. **Import stays two-phase inside one transaction.** `plan` refuses everything
|
||||
a hand-edited file can get wrong before a row exists; `materialize` writes,
|
||||
and the endpoint commits once and rolls back explicitly otherwise. A
|
||||
*rebuildable* index failing after that does not roll the campaign back: it is
|
||||
reported on the response as a warning, shown per source in the Knowledge
|
||||
panel, and repaired by Reindex. So a caller sees either "the campaign is not
|
||||
there" or "the campaign is complete", never a third thing.
|
||||
|
||||
### 9.4 The database backup, as implemented in M9
|
||||
|
||||
A second recovery tool, deliberately not merged with the first. The bundle is a
|
||||
logical, portable, human-readable copy of **one campaign** and is the supported
|
||||
way to move a campaign between installations; the backup is a physical copy of
|
||||
**this machine's whole database** and is what you take before an upgrade.
|
||||
|
||||
`backend/app/backup.py` uses SQLite's online backup API rather than a file copy,
|
||||
because a copy taken while the application runs can read one page before a
|
||||
transaction and another after it and produce a file that opens, reports a schema
|
||||
and is quietly missing rows. It writes to a temporary name beside the
|
||||
destination, runs `PRAGMA quick_check` against the finished file, and only then
|
||||
renames it into place; it opens the source read-only, never overwrites an
|
||||
existing backup, and leaves nothing behind on failure.
|
||||
|
||||
No path comes from a caller: the destination is derived from the database the
|
||||
application already has open and the filename from the clock, so the endpoints
|
||||
accept no body at all (H08).
|
||||
|
||||
**There is no restore endpoint, and that is a decision.** Restoring means
|
||||
replacing the file the running process has open, which is how both copies are
|
||||
lost at once. The procedure is in `DEVELOPMENT.md` and is a procedure precisely
|
||||
because each step needs the application stopped.
|
||||
|
||||
## 10. Authoritative Narrative State
|
||||
|
||||
### 10.1 Do not retain the RPG state protocol as the product model
|
||||
|
||||
@@ -996,3 +996,117 @@ actual persistence / memory / canon bugs
|
||||
The prose may vary.
|
||||
|
||||
The expected state, authority, and lineage rules should not.
|
||||
|
||||
---
|
||||
|
||||
# Appendix A — Proposed companion fixture: Multi-Character Identity Test
|
||||
|
||||
**Status: proposed, not built. This appendix changes nothing above it.**
|
||||
|
||||
## Why a companion rather than an extension
|
||||
|
||||
The Continuity Test above is the deterministic baseline that several milestones'
|
||||
results are compared against. Adding characters or turns to it would invalidate
|
||||
those comparisons, so **it is deliberately left exactly as it is**.
|
||||
|
||||
## The gap this fills
|
||||
|
||||
Reviewed on 2026-09-07 against a hands-on finding. The Continuity Test's seven
|
||||
deliberate traps (§13) are:
|
||||
|
||||
```text
|
||||
1-3 knowledge boundaries and secrets
|
||||
4 reference authority
|
||||
5 canon precedence
|
||||
6 branch leakage
|
||||
7 possession
|
||||
```
|
||||
|
||||
There is **no identity trap**, the word *coreference* does not appear, and the
|
||||
on-stage cast is effectively two people — Aldric and Mara, with Edrin
|
||||
established as missing rather than present. So **same-scene multi-character
|
||||
identity continuity is not exercised anywhere in the standard fixture.**
|
||||
|
||||
A play session against accepted M8 produced exactly that failure: four people in
|
||||
one office, and narration that treated one of them as two different people
|
||||
sharing a name. Root cause is unknown and no longer establishable — the playtest
|
||||
database was destroyed — which is itself part of why a *deterministic* fixture
|
||||
for this class is worth having.
|
||||
|
||||
## Shape
|
||||
|
||||
Four people, all present in one ordinary scene, with no fantasy vocabulary — the
|
||||
point is identity, not genre:
|
||||
|
||||
```yaml
|
||||
bill: { type: character, role: protagonist, controlled_by: reader }
|
||||
alice: { type: character, role: coworker }
|
||||
roger: { type: character, role: coworker }
|
||||
john: { type: character, role: coworker }
|
||||
location: { type: location, name: the office }
|
||||
```
|
||||
|
||||
Identities and roles established unambiguously before the first test turn, so
|
||||
that any later ambiguity is the system's and not the setup's.
|
||||
|
||||
## What the sequence must stress
|
||||
|
||||
- pronouns with more than one plausible referent in scene;
|
||||
- dialogue attribution across three speakers;
|
||||
- characters entering and leaving;
|
||||
- reference by name **and** by role, for the same person;
|
||||
- one character speaking *about* another;
|
||||
- one character speaking about **themself in the third person**, which is the
|
||||
shape the observed failure took.
|
||||
|
||||
## Traps
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **I1 — one Alice** | No turn may produce a second entity whose display name is `Alice`. The state model currently permits this and reports nothing, so the trap is real rather than theoretical |
|
||||
| **I2 — the protagonist stays the protagonist** | Bill must not drift into being narrated as a third party, or acquire a second entity |
|
||||
| **I3 — attribution** | A line spoken by Roger must not be attributed to John |
|
||||
| **I4 — self-reference** | No character may refer to themself as a separate same-named person |
|
||||
| **I5 — state and context agree** | The authoritative state and the assembled prompt must not disagree about who is present or who anyone is |
|
||||
| **I6 — exit and return** | A character who leaves and returns is the same entity, not a new one |
|
||||
|
||||
## Required evidence on failure
|
||||
|
||||
Unlike the fixture above, this one exists to **classify** a failure rather than
|
||||
only to detect it, because its failure modes are split between the application
|
||||
and the model. Any failing turn must capture:
|
||||
|
||||
```text
|
||||
authoritative state immediately before generation
|
||||
the exact stored context/prompt snapshot
|
||||
the recent-history section
|
||||
summaries
|
||||
retrieved memories
|
||||
imported knowledge, if any
|
||||
narrator output
|
||||
model identifier and generation settings
|
||||
```
|
||||
|
||||
then classify:
|
||||
|
||||
```text
|
||||
STATE DEFECT
|
||||
CONTEXT ASSEMBLY DEFECT
|
||||
DERIVED MEMORY/SUMMARY DEFECT
|
||||
MODEL FAILURE WITH CORRECT CONTEXT
|
||||
AMBIGUOUS / MULTIPLE CONTRIBUTORS
|
||||
```
|
||||
|
||||
Two rules for whoever runs it: **do not "fix" a model failure by editing
|
||||
authoritative state**, and **do not blame the model when the prompt already
|
||||
contained the identity error.**
|
||||
|
||||
Since M9 the whole of that evidence is portable in one campaign bundle, so a
|
||||
failing run can be exported intact and investigated elsewhere.
|
||||
|
||||
## Ownership
|
||||
|
||||
M11, alongside the realistic-model review. `V1-ACCEPTANCE-TESTS.md` §P3 records
|
||||
which parts of this are candidate **acceptance** criteria — the state-level
|
||||
traps, which are decidable — and which are model-quality observations that
|
||||
belong in a recorded review rather than in the pass/fail contract.
|
||||
|
||||
@@ -1800,6 +1800,16 @@ Export standard campaign.
|
||||
### Pass
|
||||
Export completes locally and contains enough data to restore story.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
`ai-dnd-adventure-v3`, checked section by section rather than by file size:
|
||||
the story and its whole retained tree, the branches and their disposition, the
|
||||
chosen head, Save Points, the authoritative state and its per-position
|
||||
snapshots, the state events and proposals, the imported library, the
|
||||
lineage-anchored summaries, the memories, and a stored prompt for every narrator
|
||||
turn that has one. `test_m9_portability.py::test_i01_*`, and reproducible with
|
||||
`python -m tools.m9_portability_report`, which classifies every data family as
|
||||
PRESERVED, OMITTED or DERIVED/REBUILDABLE.
|
||||
|
||||
---
|
||||
|
||||
## I02 — Import Exported Campaign
|
||||
@@ -1814,6 +1824,14 @@ Export completes locally and contains enough data to restore story.
|
||||
### Pass
|
||||
Active transcript and state are restored.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Across a **genuine clean data directory**: a second server process, in a second
|
||||
directory, against a database file that has never existed, with the exporting
|
||||
process stopped. Transcript, authoritative state, canon, Save Points, knowledge,
|
||||
state events and the size of the retained tree all match the source campaign
|
||||
(`test_m9_clean_import.py`). The same round trip inside one process is in
|
||||
`test_m9_portability.py`, and is labelled there as the weaker of the two.
|
||||
|
||||
---
|
||||
|
||||
## I03 — Branch/Disposable History Export
|
||||
@@ -1823,6 +1841,19 @@ Active transcript and state are restored.
|
||||
### Pass
|
||||
Export preserves retained alternate/disposable history needed for recovery, unless user explicitly chooses a trimmed export.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Both futures come back and stay distinguishable: the abandoned line's turns are
|
||||
present and readable, the branch that was left carries the depth it was left at,
|
||||
and a superseded take is still at its coordinate with the selected take still
|
||||
selected. There is no trimmed export — M9 offers no option to drop history, so
|
||||
the clause about one does not arise.
|
||||
|
||||
**M9 also closed a fidelity gap here.** Take parentage was not exported, so every
|
||||
imported node landed parentless and the pager grouped on the coordinate instead.
|
||||
That is right for a plain retry and wrong once two takes of one turn each have
|
||||
takes of their own beneath them; the copy read `5/5` where the source read `2/2`
|
||||
and `3/3`. v3 carries the parentage.
|
||||
|
||||
---
|
||||
|
||||
## I04 — Checkpoint Export
|
||||
@@ -1841,6 +1872,15 @@ campaign. Importing Save Points does **not** move the active head — the head
|
||||
still comes from the bundle's `headDepth`. Bundles written before M4 carry no
|
||||
`checkpoints` key, import cleanly, and create none.
|
||||
|
||||
### Re-verified — PASS (M9, 2026-09-07)
|
||||
Unchanged by the format bump, and extended in two directions. Every restored
|
||||
Save Point resolves, restores to the position it names through M3's head
|
||||
movement, leaves the retained history it moved back over intact, and the two
|
||||
in the M9 fixture restore to *different* states. A Save Point whose coordinate
|
||||
is not in the file is **dropped with the rest of the campaign kept**, never
|
||||
retargeted to a nearby turn: the reader named a position, and if that position
|
||||
is not in the file then no other position is the one they named.
|
||||
|
||||
---
|
||||
|
||||
## I05 — Knowledge Provenance Export
|
||||
@@ -1870,13 +1910,29 @@ hand-edited knowledge block with an unknown classification or empty content
|
||||
refuses the import rather than half-landing in it; and an edited content hash is
|
||||
recomputed from what actually arrived and the discrepancy recorded on the source.
|
||||
|
||||
**Limit, unchanged from before M7 and owned by M9.** The bundle carries no
|
||||
context snapshots at all, so an imported campaign has no historical prompt
|
||||
provenance — for imported knowledge or for any other component. Nothing M7
|
||||
creates is turned into a dangling id by a round trip, because no ids are
|
||||
exported; the evidence simply is not in the file.
|
||||
`test_historical_prompt_evidence_survives_an_export_round_trip` pins that
|
||||
behaviour so it cannot regress silently.
|
||||
**That limit is closed (M9, 2026-09-07).** The paragraph below is kept as
|
||||
written because it records what was true through M7 and M8, and the M9 decision
|
||||
is only legible against it.
|
||||
|
||||
> *The bundle carries no context snapshots at all, so an imported campaign has no
|
||||
> historical prompt provenance — for imported knowledge or for any other
|
||||
> component. Nothing M7 creates is turned into a dangling id by a round trip,
|
||||
> because no ids are exported; the evidence simply is not in the file.*
|
||||
|
||||
M9 carries the snapshots. An old turn in a restored campaign shows the prompt it
|
||||
was actually assembled from, the passages it was shown and the text each
|
||||
supplied — after the source has been deleted, the canon edited and the state
|
||||
moved on. `test_historical_prompt_evidence_survives_an_export_round_trip` was
|
||||
inverted rather than deleted: it now pins the thing M7 was worried about and
|
||||
could not check, which is that the provenance arriving on the other side names
|
||||
*this* campaign's sources rather than the ids they had where the file was
|
||||
written. Only that pointer is translated; the evidence is restored verbatim, and
|
||||
a source the file does not carry becomes `null` rather than pointing at a
|
||||
different file.
|
||||
|
||||
Also added in M9: `parserVersion` and `chunkingVersion` per source, recording
|
||||
what produced the passages a historical retrieval record describes, and a
|
||||
`sourceId` that exists only so the translation above can be made.
|
||||
|
||||
---
|
||||
|
||||
@@ -1887,6 +1943,28 @@ behaviour so it cannot regress silently.
|
||||
### Pass
|
||||
No external API credentials are embedded in campaign export.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Tested rather than assumed, and tested against the file's **text** rather than
|
||||
against a list of columns — a field added to a model the exporter walks would
|
||||
otherwise reach the bundle with no test of a column noticing. The inert
|
||||
`api_key` column is written with a recognisable value first, so its absence is
|
||||
evidence rather than a tautology.
|
||||
|
||||
Nothing matching `api_key`, `apiKey`, the written secret, the inference endpoint
|
||||
or its port, or an absolute filesystem path appears anywhere in an export —
|
||||
including inside the compressed snapshots, which the test decodes rather than
|
||||
skipping. Checked in both suites, so the clean-directory run covers the same
|
||||
ground across a real process boundary
|
||||
(`test_m9_portability.py::test_i06_*`, `test_m9_clean_import.py`).
|
||||
|
||||
**What deliberately does not travel**, and why it is not an omission: the
|
||||
inference endpoint, the model name, the context budget and every other row of
|
||||
`settings`. Those describe the machine, not the campaign, and importing a
|
||||
campaign must not silently repoint the destination's inference at the source's.
|
||||
Per-turn model and generation settings *do* travel, inside the historical
|
||||
snapshot, because there they are a record of what happened rather than a
|
||||
configuration to apply.
|
||||
|
||||
---
|
||||
|
||||
## I07 — Export/Import Preserves an Undone Active Head
|
||||
@@ -1921,6 +1999,28 @@ An export whose stated head lies beyond the story it contains is a file
|
||||
disagreeing with itself and must be refused rather than opened at a guessed
|
||||
position.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
All three clauses, and the main one across a genuine machine boundary.
|
||||
|
||||
- **The exact head.** The M9 fixture ends two Undos behind its own branch's
|
||||
retained tip and behind the abandoned line's, and the last thing it does is an
|
||||
Undo — so the head is not the newest row written, not the deepest row, not the
|
||||
tip, and not on the branch holding the most story. An importer guessing any one
|
||||
of those lands somewhere else. The copy opens exactly where the source was,
|
||||
the later turns are still in the database as retained future, and Redo is
|
||||
offered rather than the story having silently been redone. Redo then walks to
|
||||
the same next turn in both.
|
||||
- **The legacy clause.** A bundle with its `headDepth` removed opens at the tip
|
||||
of its head branch, offers no Redo, and offers Undo — which is the position
|
||||
such a file recorded, because at the time it was written the head could not be
|
||||
anywhere else. Checked at every seam in `test_m9_legacy_bundles.py`.
|
||||
- **The self-disagreeing file.** A head past the retained story is refused with
|
||||
a message naming where the branch actually ends, and nothing is written.
|
||||
|
||||
Evidence: `test_m9_clean_import.py::test_it_opens_at_the_exact_head_it_was_exported_at`
|
||||
(second process, empty directory), plus `test_m9_portability.py::test_i07_*` and
|
||||
the head cases in `test_m9_corrupt_bundles.py`.
|
||||
|
||||
---
|
||||
|
||||
# J. Genre Independence
|
||||
@@ -2082,6 +2182,17 @@ turn.
|
||||
### Pass
|
||||
State at each position matches original accepted state.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Measured **after a round trip**, which is the M9 form of it: the copy and the
|
||||
source are walked back three turns and forward three turns in step, and the
|
||||
authoritative document is compared at every position. They agree throughout.
|
||||
|
||||
That this stays a snapshot read rather than a replay is the point. Undo, Redo
|
||||
and Save Point restore all resolve a coordinate and read the state recorded
|
||||
there (`TECHNICAL-DESIGN.md` §10.4), so an import that carried the events and
|
||||
dropped the per-position snapshots would have made every one of them
|
||||
proportional to campaign length. Both halves of §17's hybrid travel.
|
||||
|
||||
---
|
||||
|
||||
## L03 — Checkpoint Reconstruction After Restart
|
||||
@@ -2105,6 +2216,15 @@ exactly that value after the campaign had been advanced past it
|
||||
`TestClient` restart, which could not distinguish durable state from a live
|
||||
object.
|
||||
|
||||
### Re-verified after a move — PASS (M9, 2026-09-07)
|
||||
The same claim with a machine boundary in front of it. A campaign is exported
|
||||
from one server process, imported into a **second process against a database
|
||||
file that has never existed**, a Save Point is restored there, that process is
|
||||
killed, and a **third** process against the same file is asked again. The
|
||||
transcript, the authoritative state and the size of the retained tree all match
|
||||
what the second process had after restoring
|
||||
(`test_m9_clean_import.py::test_l03_*`).
|
||||
|
||||
---
|
||||
|
||||
## L04 — Derived Data Can Be Rebuilt
|
||||
@@ -2121,6 +2241,34 @@ using a safe test copy.
|
||||
### Pass
|
||||
Authoritative campaign history remains intact and derived structures can be recreated.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07), and one defect found by running it
|
||||
On a safe copy — an imported campaign, not the original. Every physically
|
||||
derived structure is destroyed and rebuilt from the source content the bundle
|
||||
carried: passages, the FTS rows and the vectors. Afterwards every source is
|
||||
`ready` with passages again, retrieval works, and the transcript, the
|
||||
authoritative state, the classifications and the lifecycle flags are identical
|
||||
either side. Deleting the vectors alone leaves lexical retrieval working, which
|
||||
is M7's rule that the lexical half is a production path and not a fallback. A
|
||||
rebuild does not make an abandoned line's summary eligible.
|
||||
|
||||
**Running it found a real defect, which is fixed here.** The FTS5 index is a
|
||||
virtual table, so no foreign key reaches it and no `ON DELETE CASCADE` covers
|
||||
it: deleting a campaign dropped its passages and left one index row per passage
|
||||
behind. Nothing read them — every search joins through `knowledge_chunks` — so
|
||||
the leak was invisible until SQLite handed the freed primary key out again, at
|
||||
which point the **next source imported into any campaign** failed with an
|
||||
integrity error. Reindex could not repair it either, because `clear_index` finds
|
||||
index rows *through* the chunks, and there were none. Both ends are closed: a
|
||||
campaign's index rows are removed before it is deleted, and the index insert now
|
||||
replaces a stale row rather than colliding with it — so a database already
|
||||
carrying the leak repairs itself and needs no migration. See the M9 report's
|
||||
findings.
|
||||
|
||||
The rebuild path M9 depends on is therefore implemented and tested rather than
|
||||
assumed, which the SHOULD priority above did not require but the milestone did:
|
||||
knowledge passages and indexes are omitted from the bundle precisely because
|
||||
they can be rebuilt.
|
||||
|
||||
---
|
||||
|
||||
# M. Long-Run Test
|
||||
@@ -2327,3 +2475,92 @@ L01-L03
|
||||
```
|
||||
|
||||
This keeps implementation work tied to observable behavior rather than repository-specific architecture.
|
||||
|
||||
---
|
||||
|
||||
# P. M11 Test-Design Tasks — Not Yet Acceptance Tests
|
||||
|
||||
**Nothing in this section is part of the pass/fail contract.** These are three
|
||||
behaviours a real play session against accepted M8 showed are worth testing, for
|
||||
which **the pass criterion is not yet settled**. They are recorded here so the
|
||||
release tester finds them where they will look, and they are deliberately not
|
||||
written as A-through-O items: giving them IDs would imply a criterion has been
|
||||
ratified when it has not, and would destabilise a document whose value is that
|
||||
every item in it is decidable.
|
||||
|
||||
Each names **what must be settled** before it can become an acceptance test.
|
||||
Full context is in the M9 report's §Y and, durably, in `BUILD-MILESTONES.md`
|
||||
under the post-M8 playtest findings.
|
||||
|
||||
## P1 — The reader can tell where they are after history movement
|
||||
|
||||
**From:** a real session in which Undo worked correctly and the reader could not
|
||||
tell which point in the story they had reached.
|
||||
|
||||
**Behaviour to test:** after Undo, Redo, a Save Point restore, or an edit to an
|
||||
earlier turn, the reader can identify their current position in the visible
|
||||
story without implementation terminology (`branch`, `head`, `node`, `depth`
|
||||
remain forbidden at the surface).
|
||||
|
||||
**Settle first:** what the indicator *is*. "The reader can tell" is not
|
||||
decidable as written — it needs an observable artifact, such as a named position
|
||||
that changes with movement and is present in the DOM. `BROWSER-UX-SPEC.md` §8
|
||||
now carries the requirement; **no wording is ratified**, and this cannot become
|
||||
an acceptance test before one is.
|
||||
|
||||
## P2 — Narration length has a measurable directional effect
|
||||
|
||||
**From:** a reader who chose *2-4 paragraphs* and received substantially longer
|
||||
replies.
|
||||
|
||||
**Behaviour to test:** the narration-length setting produces a **measurable
|
||||
directional difference** in output length across repeated realistic turns —
|
||||
brief shorter than standard, standard shorter than detailed — on at least the
|
||||
reference 3B narrator and one stronger local narrator.
|
||||
|
||||
**Settle first:** the numbers, and the mechanism. Today the setting adds one
|
||||
English sentence to the campaign instructions and changes **no** generation
|
||||
budget, while a separate numeric hint derived from the global
|
||||
`max_output_tokens` is identical for every setting (M9 report §Y). Until the
|
||||
product decides what each setting *means* — and whether it moves the budget —
|
||||
there is no threshold to test against. A directional test is stateable; an
|
||||
absolute one is not, and this should not become an acceptance test that asserts
|
||||
word counts nobody has ratified.
|
||||
|
||||
**Do not** turn this into a truncation test: the state block is emitted last and
|
||||
hard truncation removes it.
|
||||
|
||||
## P3 — Multi-character identity continuity
|
||||
|
||||
**From:** four people in one scene, and narration that treated one of them as
|
||||
two different people of the same name. **Root cause unknown** — the playtest
|
||||
database was destroyed, so no evidence survives.
|
||||
|
||||
**Behaviour to test:** across a multi-turn scene with a protagonist and three
|
||||
supporting characters, the story does not create duplicate characters, does not
|
||||
duplicate a display name across two entities, does not drift the protagonist's
|
||||
identity, does not misattribute dialogue, and does not have a character refer to
|
||||
themself as a separate same-named character — and the authoritative state and
|
||||
the assembled context do not disagree about who anyone is.
|
||||
|
||||
**Settle first:** which of those are **product** guarantees and which are
|
||||
**model-quality** observations. They are not the same kind of claim and must not
|
||||
share one verdict:
|
||||
|
||||
- *"The state never holds two entities with the same display name"* is
|
||||
decidable and enforceable, and is a candidate acceptance test today. The
|
||||
implementation currently permits it and reports nothing (M9 report §Y).
|
||||
- *"The narrator never confuses two same-named characters"* is not a pass/fail
|
||||
property of this application — it depends on the model — and belongs in M11's
|
||||
realistic-model review with a recorded classification, not in this contract.
|
||||
|
||||
**A run of this must capture**, on any failure: pre-generation state, the exact
|
||||
stored prompt snapshot, history, summaries, retrieved memories, imported
|
||||
knowledge, narrator output, and model settings — then classify as a state,
|
||||
context-assembly, derived-data, or model failure. M9 made all of that portable,
|
||||
so a failing campaign can be exported whole and investigated elsewhere.
|
||||
|
||||
**Fixture:** the standard Continuity Test does not exercise this — its traps are
|
||||
knowledge, authority, branch leakage and possession, and its on-stage cast is
|
||||
effectively two people. A companion fixture is proposed in
|
||||
`TEST-CAMPAIGN-FIXTURE.md`; the established fixture is deliberately unchanged.
|
||||
|
||||
+96
-3
@@ -1,8 +1,101 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v3.3
|
||||
- **Revision date:** 2026-09-06
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted** (M8 closed out 2026-09-06). M9 is next and has not been started.
|
||||
- **Package:** Adventure Storyteller Planning Package v3.5
|
||||
- **Revision date:** 2026-09-07
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 implemented and awaiting independent review** (2026-09-07). M10 has not been started.
|
||||
|
||||
## v3.5 — Post-M8 hands-on playtest findings recorded (2026-09-07)
|
||||
|
||||
**Documentation only. No application code changed, and M9's verified result is
|
||||
untouched** — see the note at the end of this entry.
|
||||
|
||||
A real play session against **accepted, signed M8** (real browser, trusted-LAN
|
||||
Ollama, `qwen2.5:3b-instruct-16k`, disposable database since destroyed) surfaced
|
||||
four product-quality observations. **None is an M9 defect, none was caused by
|
||||
M9, and none blocks M9 acceptance.** They are recorded so they cannot be lost
|
||||
when the M9 report is archived.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `reports/M9-IMPLEMENTATION-REPORT.md` | **New §Y**, the full write-up with the verification behind each mechanism, plus a pointer from §W. Explicitly labelled as not-M9. | observation record |
|
||||
| `BUILD-MILESTONES.md` | **New section before M10**, the durable copy owned by M11; a note bounding M10 out of it; the four items added to M11's scope list. | milestone sequencing |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | **New §P**, three items as **test-design tasks, explicitly not acceptance tests**, each naming what must be settled before it could become one. No existing test changed or weakened. | test design |
|
||||
| `BROWSER-UX-SPEC.md` | **New §8A.** §8's *"The current endpoint should be clear"* was satisfied while a real reader was lost, so it could not hold the behaviour. States the orientation requirement; **prescribes no wording**. | requirement clarification |
|
||||
| `TEST-CAMPAIGN-FIXTURE.md` | **New Appendix A** proposing a companion `Multi-Character Identity Test`. The established deterministic fixture is **unchanged** — altering it would invalidate earlier milestones' comparisons. | test design |
|
||||
|
||||
**What was verified rather than assumed**, since the playtest campaign no longer
|
||||
exists and root causes largely cannot be proven:
|
||||
|
||||
- The browser title genuinely is `AI D&D` (`frontend/index.html`), and **no
|
||||
accepted document ever claimed otherwise** — so this is an uncovered gap, not
|
||||
documentation needing correction. Nothing was corrected and nothing was
|
||||
renamed: *Adventure Storyteller* is itself narrower than the genre-agnostic
|
||||
engine `SPECIFICATION.md` requires, and the naming decision is the owner's.
|
||||
- The narration-length setting adds **one English sentence** and changes **no
|
||||
generation budget**, while the numeric hint derived from the global
|
||||
`max_output_tokens` is **identical for every setting** — measured at the
|
||||
default as *"must not exceed 506 words, and it should not stop short of about
|
||||
177."* Recorded as a mechanism to check first, **not** as the proven cause.
|
||||
- The narrative state **permits two entities to share a display name and
|
||||
reports nothing** — `DUPLICATE_ENTITY` rejects a repeated key only. That is
|
||||
one of the identity finding's failure modes; it establishes nothing about what
|
||||
actually happened.
|
||||
|
||||
**M9 is unaffected.** No application file changed in this pass, so M9's final
|
||||
verified result stands exactly as recorded: **1,102 backend passed / 14 skipped
|
||||
/ 0 failed**, 145 frontend, 36/36 browser. The expensive M9 suites were
|
||||
deliberately **not** re-run, because only Markdown changed.
|
||||
|
||||
## v3.4 — M9 Implementation (2026-09-07)
|
||||
|
||||
M9 — Export, Backup, Recovery, and Migration Hardening — is implemented on
|
||||
`m9-recovery` from the signed M8 commit `1ce9972`. This revision records what the
|
||||
implementation settled. It is **not** an acceptance: the milestone report is
|
||||
written for a reviewer and the tree is staged for the repository owner's signed
|
||||
commit.
|
||||
|
||||
**Requirement corrections:** none. M9 altered no product requirement.
|
||||
`SPECIFICATION.md` and `SECURITY-THREAT-MODEL.md` are unchanged — §16 already
|
||||
required the export to preserve the exact active position, §6.1 already required
|
||||
an exact prompt/context snapshot per turn, and M9 implements both rather than
|
||||
redefining either.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `DATA-MODEL.md` §29 | The v3 format, the three-category rule (chosen / evidence / rebuildable), what each version can be trusted to say, the encoding of the snapshots, and the two pointers the import translates. | implementation fact |
|
||||
| `TECHNICAL-DESIGN.md` §9.3 | Why the version was bumped when §9.1 and §9.2 each correctly declined one; the third data category; the two-phase transaction and the warning path for a failed derived rebuild. | implementation fact |
|
||||
| `TECHNICAL-DESIGN.md` §9.4 | **New.** The SQLite backup: the online backup API rather than a file copy, the verify-then-rename order, and why there is no restore endpoint. | implementation fact |
|
||||
| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **New subsection.** Story Cards settled as compatibility-only legacy data and removed from the narrator's prompt, with the evidence that they were the "alternate untracked path" §73 already forbade. | newly settled design decision |
|
||||
| `V1-ACCEPTANCE-TESTS.md` I01-I07, L02-L04 | Results recorded. I05's M7-era limit is marked closed with the original paragraph kept, because the M9 decision is only legible against it. L04 records the defect running it found. | implementation fact |
|
||||
| `BUILD-MILESTONES.md` M9 | Marked complete, with what it delivered, the three defects it found, the story-card decision, and the debt carried forward. | implementation fact |
|
||||
| `README.md`, `VERSION.md` | Status. | implementation fact |
|
||||
| `DEVELOPMENT.md` | **New section**: the two recovery tools and when each applies, taking a backup, and the stop-move-start restore procedure. Plus a note that an imported long campaign meets a small context ceiling on its first turn rather than gradually. | implementation fact |
|
||||
|
||||
**The four M8 handoff questions, answered**
|
||||
|
||||
| | Answer |
|
||||
| --- | --- |
|
||||
| **A. Complete campaign portability** | Every family travels and is measured family by family, before and after, by a tool a reviewer can rerun. |
|
||||
| **B. Historical prompt provenance** | **It belongs in the bundle, and it is in it.** An old turn in a restored campaign shows what it was actually given, after the source has been deleted and the canon edited. |
|
||||
| **C. Legacy story cards** | Compatibility-only. Carried in both directions; removed from the narrator's prompt; still the summariser's character roster. |
|
||||
| **D. Context-window portability** | The campaign travels; the machine's model configuration does not. Importing changes no setting of the destination's, and a campaign imports whether or not any model is installed. The window itself remains M11's. |
|
||||
|
||||
**What the implementation found rather than assumed**
|
||||
|
||||
Three defects, all found by running the milestone's own tests rather than by
|
||||
reading: an FTS index leak that made an ordinary import fail in an unrelated
|
||||
campaign and that Reindex could not repair; an imported node with no state
|
||||
snapshot being stamped with the campaign's *head* state; and a snapshot pointer
|
||||
that was not being translated because the code mutated a dict in place. The
|
||||
first two predate M9.
|
||||
|
||||
One measurement changed a plan, and then corrected the conclusion drawn from it.
|
||||
Carrying per-turn prompts looked like it would halve the length of campaign that
|
||||
can be restored. Measured, and after compressing them inside the file, everything
|
||||
M9 added costs **12%** of reachable campaign length — the import ceiling moves
|
||||
from about 318 turns to about 279, against a 100-turn certification target. The
|
||||
dominant cost is not M9's at all: the **per-position narrative state document is
|
||||
74% of a bundle**, and v2 already carried it.
|
||||
|
||||
## v3.3 — M8 Closeout (2026-09-06)
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user