M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
co-authored by
Claude Opus 5
parent
1ce9972760
commit
44edece67e
@@ -996,3 +996,117 @@ actual persistence / memory / canon bugs
|
||||
The prose may vary.
|
||||
|
||||
The expected state, authority, and lineage rules should not.
|
||||
|
||||
---
|
||||
|
||||
# Appendix A — Proposed companion fixture: Multi-Character Identity Test
|
||||
|
||||
**Status: proposed, not built. This appendix changes nothing above it.**
|
||||
|
||||
## Why a companion rather than an extension
|
||||
|
||||
The Continuity Test above is the deterministic baseline that several milestones'
|
||||
results are compared against. Adding characters or turns to it would invalidate
|
||||
those comparisons, so **it is deliberately left exactly as it is**.
|
||||
|
||||
## The gap this fills
|
||||
|
||||
Reviewed on 2026-09-07 against a hands-on finding. The Continuity Test's seven
|
||||
deliberate traps (§13) are:
|
||||
|
||||
```text
|
||||
1-3 knowledge boundaries and secrets
|
||||
4 reference authority
|
||||
5 canon precedence
|
||||
6 branch leakage
|
||||
7 possession
|
||||
```
|
||||
|
||||
There is **no identity trap**, the word *coreference* does not appear, and the
|
||||
on-stage cast is effectively two people — Aldric and Mara, with Edrin
|
||||
established as missing rather than present. So **same-scene multi-character
|
||||
identity continuity is not exercised anywhere in the standard fixture.**
|
||||
|
||||
A play session against accepted M8 produced exactly that failure: four people in
|
||||
one office, and narration that treated one of them as two different people
|
||||
sharing a name. Root cause is unknown and no longer establishable — the playtest
|
||||
database was destroyed — which is itself part of why a *deterministic* fixture
|
||||
for this class is worth having.
|
||||
|
||||
## Shape
|
||||
|
||||
Four people, all present in one ordinary scene, with no fantasy vocabulary — the
|
||||
point is identity, not genre:
|
||||
|
||||
```yaml
|
||||
bill: { type: character, role: protagonist, controlled_by: reader }
|
||||
alice: { type: character, role: coworker }
|
||||
roger: { type: character, role: coworker }
|
||||
john: { type: character, role: coworker }
|
||||
location: { type: location, name: the office }
|
||||
```
|
||||
|
||||
Identities and roles established unambiguously before the first test turn, so
|
||||
that any later ambiguity is the system's and not the setup's.
|
||||
|
||||
## What the sequence must stress
|
||||
|
||||
- pronouns with more than one plausible referent in scene;
|
||||
- dialogue attribution across three speakers;
|
||||
- characters entering and leaving;
|
||||
- reference by name **and** by role, for the same person;
|
||||
- one character speaking *about* another;
|
||||
- one character speaking about **themself in the third person**, which is the
|
||||
shape the observed failure took.
|
||||
|
||||
## Traps
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **I1 — one Alice** | No turn may produce a second entity whose display name is `Alice`. The state model currently permits this and reports nothing, so the trap is real rather than theoretical |
|
||||
| **I2 — the protagonist stays the protagonist** | Bill must not drift into being narrated as a third party, or acquire a second entity |
|
||||
| **I3 — attribution** | A line spoken by Roger must not be attributed to John |
|
||||
| **I4 — self-reference** | No character may refer to themself as a separate same-named person |
|
||||
| **I5 — state and context agree** | The authoritative state and the assembled prompt must not disagree about who is present or who anyone is |
|
||||
| **I6 — exit and return** | A character who leaves and returns is the same entity, not a new one |
|
||||
|
||||
## Required evidence on failure
|
||||
|
||||
Unlike the fixture above, this one exists to **classify** a failure rather than
|
||||
only to detect it, because its failure modes are split between the application
|
||||
and the model. Any failing turn must capture:
|
||||
|
||||
```text
|
||||
authoritative state immediately before generation
|
||||
the exact stored context/prompt snapshot
|
||||
the recent-history section
|
||||
summaries
|
||||
retrieved memories
|
||||
imported knowledge, if any
|
||||
narrator output
|
||||
model identifier and generation settings
|
||||
```
|
||||
|
||||
then classify:
|
||||
|
||||
```text
|
||||
STATE DEFECT
|
||||
CONTEXT ASSEMBLY DEFECT
|
||||
DERIVED MEMORY/SUMMARY DEFECT
|
||||
MODEL FAILURE WITH CORRECT CONTEXT
|
||||
AMBIGUOUS / MULTIPLE CONTRIBUTORS
|
||||
```
|
||||
|
||||
Two rules for whoever runs it: **do not "fix" a model failure by editing
|
||||
authoritative state**, and **do not blame the model when the prompt already
|
||||
contained the identity error.**
|
||||
|
||||
Since M9 the whole of that evidence is portable in one campaign bundle, so a
|
||||
failing run can be exported intact and investigated elsewhere.
|
||||
|
||||
## Ownership
|
||||
|
||||
M11, alongside the realistic-model review. `V1-ACCEPTANCE-TESTS.md` §P3 records
|
||||
which parts of this are candidate **acceptance** criteria — the state-level
|
||||
traps, which are decidable — and which are model-quality observations that
|
||||
belong in a recorded review rather than in the pass/fail contract.
|
||||
|
||||
Reference in New Issue
Block a user