# Adventure Storyteller — v1.1 Plan **Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the signed v1.0.0 release commit `432f041`. **WP-A1 and WP-A2** are committed and signed as `d63804f` (`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`). **WP-B.1**, the memory-retention diagnostic, is committed and signed as `beb17ad` (`reports/v1.1/V1.1-WP-B1-REPORT.md`). **WP-B.2**, the retention fix, is complete and staged for the owner's signed commit. It is **accepted with a documented real-model limitation** (owner decision, 2026-09-15): - B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship; - deterministic independent-memory recovery passes; - the one isolation-valid reference-model run failed at memory creation; - the B2.4 prompt experiment did not fix that and was reverted. The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T. WP-C, WP-D and WP-E have not started. No v1.1 version or tag exists. This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1 history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1 specification and acceptance contract are unchanged: every package below improves how an existing requirement is met. No package adds a requirement. --- ## 1. Terminology - **Work package (WP).** One independently reviewable unit of v1.1 work, with its own coding brief, its own report, and its own review. Work packages are lettered. **They are not milestones, and there is no M12.** - **Brief.** The coding prompt for one work package. It holds only the objective, the source documents, the scope and non-scope, the acceptance criteria, and the stop condition. No work package begins before its brief exists. - **Reference narrator.** `qwen2.5:3b-instruct`, with and without `num_ctx` 16,384 baked in, and `nomic-embed-text` for embeddings. These are the models the v1 evidence used (M11 report §E.1). v1.1 real-model evidence uses the same models so that its numbers compare with v1's. - **Reference hosts.** The CPU reference host serves Ollama over HTTPS with a private CA, which is the A06 evidence path. The GPU inference host serves plain HTTP on the LAN and is used for long runs (M11 report §E.1). - **Planning package version** (`VERSION.md`, v4.x) and **product version** (v1.0.0, v1.1.0) are separate numbers. `SECURITY-THREAT-MODEL.md`'s own "Status: v1.1" line is that document's revision label from Phase 0B. It does not refer to this release. ## 2. Baseline | | | | --- | --- | | Release | **v1.0.0**, 2026-09-14 | | Release commit | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, signed by the owner | | Tag | `v1.0.0`, a signed annotated tag on that commit, pushed | | `main` | `432f041`, the same commit | | v1.1 branch | `v1.1-development`, created at `432f041` | | v1 contract | 82 REQUIRED FOR V1 tests: 81 PASS and H09 NOT APPLICABLE (M11 report §T) | | Schema | migrations up to 94 (`LATEST_VERSION` 94) | | Bundle format | `ai-dnd-adventure-v3` | | Provenance | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` | ## 3. v1.1 goals 1. **No silent loss of what the narrator is given.** A prompt must not reach the server's window edge by an accident of tokenizer arithmetic. A turn the server truncated must be reported. 2. **No protocol in the story.** The shapes of the application's own protocol that a narrator copies must not be stored as narration. The fixed instructions must not hand every campaign one genre's nouns to copy. Story prose must not be removed in the process. 3. **Memory that remembers on its own.** An old fact must be recoverable from the memory bank itself, with provenance, when nothing else carries it. This must not weaken lineage safety. 4. **Release evidence without API-only gaps.** Every reader-facing history, state, length, failure and export behaviour must be driven in a real browser. 5. **Recovery that says when it cannot recover.** Backups get a full integrity check, and an export that import would refuse must say so. 6. **Control boundaries that meet WCAG 1.4.11.** And throughout: **every v1.0.0 campaign opens unchanged in v1.1.** ## 4. Non-goals - New product features: media generation, TTS, STT, story search, whole-transcript copy, a discarded-history recovery screen, a restore button, tablet redesign. - Changes to the history, branch, head or Save Point architecture (ADR 005, ADR 012). - Changes to the authoritative state model, its event vocabulary or its validator (ADR 010, ADR 013). Changes to the prompt *wording* that describes the vocabulary are in scope, in WP-A2. - Changes to the knowledge authority classes or their retrieval. - A new bundle format version. - Any new runtime dependency, any network access beyond the configured inference endpoint, and any first-use download, including per-model tokenizers. - Supporting inference servers other than Ollama, beyond what `context_window_override` already allows. - Re-writing stored narration or memories of existing campaigns automatically. - Redesigning identity or state handling on the strength of post-M8 finding D, which has not been reproduced. ## 5. Rules for every work package The v1 milestone rules (`BUILD-MILESTONES.md` §2) apply to every work package, and so do the following: 1. **Compatibility default.** Existing v1.0.0 campaign databases open unchanged. Any migration is forward-only and additive, and is tested by opening a real v1.0.0 database. Fresh-install and upgraded schemas must still compare identical (`test_m11_migration.py`). v1.0.0 bundles still import. 2. **The bundle stays `ai-dnd-adventure-v3`** unless a brief makes the case under M9's semantic test and the owner agrees. Adding a key inside stored evidence is not a format change when an absent key unambiguously means "not recorded". 3. **The v1 contract is the regression floor.** No REQUIRED FOR V1 test is retired, relaxed or reclassified. 4. **Evidence discipline** (M11 report §O). Any product change made after a black-box run was taken invalidates that run for release purposes. Harness defects and product defects are reported separately. A check that cannot fail is not evidence. 5. **Each package has a report** (`planning/reports/V1.1-WP--REPORT.md`). It records what was built, the acceptance results with measurements, what was not verified, and the as-implemented documentation changes. The report rotation for `reports/` is decided when the first one is written (`planning/README.md`). 6. **Real-model runs longer than a few minutes on the GPU host** require the power, link and kernel logging in `DEVELOPMENT.md`, started before the run. 7. **No real hostnames, addresses or people's names** in committed files, fixtures or reports. Evidence stays under `$HOME`, never `/tmp`. 8. **Commits and tags are the owner's.** A package ends with its changes staged for a signed commit. --- ## 6. Backlog triage Severity reflects user risk: **High** means story corruption or lost continuity that the reader is not told about. **Medium** means silent degradation or a real verification gap. **Low** means visible, rare or cosmetic. ### 6.1 The recorded post-v1 backlog | # | Item | Source | Current evidence | Severity | Disposition | | --- | --- | --- | --- | --- | --- | | 1 | Context-window safety margin | M11 §P risk 3, §N; post-v1 backlog | The largest prompts left 23, 34 and 42 real tokens of headroom in the GPU runs. The application's `cl100k_base` count ran 16 tokens below the narrator's on every prompt measured. The only slack is `OUTPUT_SAFETY_MARGIN = 64` (`context/builder.py`), which is fixed and must also absorb the section separators. Ollama 0.34 cut an over-window synthetic prompt to 8,194 tokens with no error. Truncation drops the oldest tokens, which are the narrator's rules and the canon. The server's reported usage is already stored per turn (`snapshot["usage"]`, `routers/adventures/turns.py`) and read by nothing. | **High**: silent, invisible, and reachable by a model whose tokenizer diverges further | **v1.1, WP-A1** | | 2 | Narrator restating prompt and state text | §P risk 4, §G.4 | The state's fact line was restated as a sentence on 74 of 104 turns in the evidence run. Phrases such as "Scene set, continue your adventure." and lines opening "Memory:" are model-invented, and none is an application string (verified by search). The restatements sit inside story sentences. | **Medium**: narration quality, and it is the path by which M04's fact reached memory, which masks item 5 | **v1.1, WP-A2** measures it. Reducing in-sentence restatement is **v1.2**, because removing it means judging prose | | 3 | Protocol leakage: event-call syntax, stray headings, repeated length hints | §P risk 6, §S.6 | On 4 of 10 turns of the closeout identity run (4,096 window) and 0 of 104 in the 100-turn runs. Causes verified in code. `events.vocabulary_for_prompt` shows every event as `name(field, …)`, which is the notation copied. `_is_echoed_instruction` requires both "state block" and "events list", and the length hint's tail says only "state block". A lone trailing `Scene:` is not removed. Stored text is replayed as history (TECHNICAL-DESIGN §15.4). | **Medium-high**: silent and self-reinforcing, but rare at a full window | **v1.1, WP-A2** | | 4 | Genre-specific example in the fixed state rule | §P risk 16; `narrative/extract.py` `EMIT_RULE` | The example names `mara`, `silver-key`, `old-abbey` and `aldric`. In the office-meeting identity run the narrator proposed giving `silver-key` to Alice, and the validator refused it. No state was corrupted. | **Medium**: every non-fantasy campaign gets fantasy nouns to copy | **v1.1, WP-A2** | | 5 | Independent long-term-memory retention | §P risk 5, §G.4; F02 and M04 qualifications | No run showed a planting-era memory carrying the planted fact. Recovery ran from state to the narrator's restatement to memory. Code facts that bear on it, none yet proven to be the cause: a memory is at most 50 words for a 6-action block. The summariser sees the block's *last* 2,000 tokens (`truncate_to_last_tokens`), so an early fact in a long block can be cut. Eviction is least-recently-used above `memory_bank_capacity` (default 80), so a never-retrieved early memory is the first to go once a campaign passes about 480 actions; the 33-memory evidence run never reached that. Ranking is cosine similarity plus a pin (CONTEXT-AND-MEMORY §20). | **High** for long campaigns: lost continuity, seen only as the story forgetting | **v1.1, WP-B** | | 6 | Browser automation gaps: Retry, Save Point, state correction, narration length, failed generation, export download | §S.3, §T qualifications, §P risk 11 | Proved through the API, the component suite and the 100-turn campaign, but never driven in a browser. The export control is exercised only as far as the click, because a `blob:` download does not leave the headless snap Firefox. No defect is known. | **Medium**: a verification gap on the release path | **v1.1, WP-C** | | 7 | Practical export and bundle-size ceiling | §P risk 12, §K; M9 debt | The import body limit is 20 MB (`limits.MAX_IMPORT_BODY_BYTES`). M9's conservative ceiling is about 279 turns. The real 100-turn campaign was 2.74 MB, about 13 kB per action, or roughly 1,600 actions. A larger campaign exports and is then refused on import, and nothing says so at export. | **Medium** impact, low frequency | **v1.1, WP-D**: warn at export. Raising the limit or a streaming import is **v1.2** | | 8 | Identity-confusion diagnostic follow-up | §P risk 8, §S.5 | Not reproduced. The diagnostic cannot see a stale scene, has never run with memory on, and has never run at a 16,384 window. It found no product deficiency. | **Low**: no occurrence since the original, whose campaign is gone | **No v1.1 package.** The diagnostic is re-run as a WP-A2 regression and in the v1.1 release gate, with memory on. A stale-scene detector is **v1.2**. The next real occurrence is classified with `tools/m11_identity.py` | | 9 | WCAG 1.4.11 control-boundary contrast | §P risk 10, §L | Measured at 1.33:1 resting and 1.75:1 on hover (`--border` and `--border-bright` against `--bg-panel`). M11's argument that the label identifies the control is weakest for text inputs, where the boundary shows where to type. | **Low-medium** | **v1.1, WP-E** | | 10a | Backup integrity: `integrity_check` | §P risk 14; M9 | `backup.py` runs `PRAGMA quick_check`, which skips index-content verification. | **Low** | **v1.1, WP-D** | | 10b | Scheduled backups | §P risk 14; M9 | None exist; backups are manual. They need owner policy on interval, retention, location and disk use, and must not contend with a turn's write lock (§O.7). | **Medium** impact, but a new feature | **v1.2** | | 11 | Real media-provider adapters | §P risk 13; M10; K05 and K06 (FUTURE) | The seam has no consumer. Adding one brings dependencies, a stricter endpoint rule (`DEVELOPMENT.md`) and a new UI. | n/a: a feature, not a risk to existing stories | **Future / optional**, not v1.1 | ### 6.2 Other recorded v1 residual risks and debt | # | Item | Source | Disposition | | --- | --- | --- | --- | | 12 | The state lags the narration at 4,096 with the 3B narrator (1 proposal in 10 applied) | §P risk 17, §S.5 | Model behaviour, not a defect. WP-A2's real-model runs record the proposal outcome counts. No package | | 13 | Every realistic observation is a 3B model's | §P risk 2 | No package. v1.1 keeps the reference narrator for comparability. A comparative model recommendation stays open (`planning/README.md`, *Still open*) | | 14 | A campaign that never chose a narration length keeps the pre-M11 hint | §P risk 15 | Deliberate. No change. WP-C drives the control | | 15 | The GPU host dropped off the PCIe bus after the evidence run | §P risk 7, §E.1 | Operational. Rule 6 in §5 applies to every long run | | 16 | Release timings are two hosts' | §P risk 1 | No action. No performance requirement exists or is invented | | 17 | Cross-layer duplication: one fact in state, memory, history and a passage at once | CONTEXT-AND-MEMORY §22; open since M6 and M7 | **v1.2.** It costs tokens (item 1) and is related to item 2. WP-B may touch ranking only where its diagnostic requires | | 18 | Memory-ranking factors beyond similarity and pin not implemented | CONTEXT-AND-MEMORY §20 | Inside WP-B's bounded fix menu, only if WP-B's diagnostic places the failure at ranking | | 19 | M8 carried debt: no whole-transcript copy or story search (§77, §78), no discarded-history recovery screen (§63), tablet untuned, RPG world state read-only | `BUILD-MILESTONES.md` M8 and M9 | **Future backlog.** Features, not reliability | | 20 | M10 debt: `ambience` is an empty shape; a deleted visual profile is unrecoverable in the app; profiles are API-only | M11 §D | **Future**, with media | | 21 | A hostile host on the trusted LAN; the DNS-rebinding interval | `SECURITY-THREAT-MODEL.md` §10A | Accepted for v1. **Future.** No v1.1 package | | 22 | A stale `chunk_id` in a restored snapshot | M9 | No action. It is a React key, not a live pointer | | 23 | Developer-venv residue (`quickjs`, `psycopg`) | §M | No action. It does not ship | | 24 | A Settings warning comparing the real window to the budget | M8 debt, "unowned" | **Closed by M11**: Settings reports the window and warns | | 25 | Root `README.md` has stale facts: "1,191 backend tests" (1,421 at closeout), and the Screenshots paragraph calls screenshots "a job for the UI pass in M8" | found in this review | Documentation debt, not release status, so it is not changed in v4.1. It is fixed in WP-A1's documentation update | --- ## 7. Grouping decisions **The suggested "narrator boundary hardening" package is split into A1 and A2.** They share a theme but not a mechanism or a test method. - **A1**, the context reserve, is budget arithmetic in `context/builder.py` and `contextwindow.py`. It is proved deterministically with a scripted provider that reports token counts. - **A2**, the protocol echo, is the contract between the prompt's wording (`extract.EMIT_RULE`, `events.vocabulary_for_prompt`, the length hint) and the extractor. It is proved with a replay corpus of real narration and real-model runs. Bundling them would make one review carry two unrelated risk profiles. A2's three items belong together. The example slugs and the call notation are the *source* of the copied text, and the extractor is its *sink*. Changing one without the other means proving the change twice. Both also alter the stored prompt, so they invalidate the same evidence. **B stands alone.** It is the only package whose success depends on model output quality. Its test design, a fact that nothing but memory carries, is the hard part. It follows A2, so that it measures the prompt v1.1 will ship. **D is narrowed to recovery honesty**, meaning `integrity_check` and the export warning. Both are small and deterministic, and both are recovery-path changes with no policy question. Scheduled backups need owner decisions and have real design risk, including write-lock contention. Raising or streaming the import limit changes a DoS guard. Both move to v1.2. **E stays focused** on control boundaries. The same measurement found no other accessibility defect (M11 §L). **F is not a package.** Nothing concrete in the product is deficient. The diagnostic's known blind spots are a stale scene, memory off, and one window size. The first two can be addressed by running it differently: the v1.1 gate runs it with memory on. The stale-scene detector is new diagnostic work with no occurrence to justify it yet, so it is v1.2. **G is not in v1.1.** It adds no reliability or quality to existing stories, it brings dependencies and a new endpoint surface, and its tests (K05, K06) are FUTURE. `MEDIA-EXTENSION-CONTRACT.md` and the M10 seam remain authoritative for whenever it is taken up. --- ## 8. Work packages ### WP-A1 — Context-window safety reserve **Objective.** An assembled prompt plus its reply reserve stays at least a documented safety reserve below the effective window, whatever tokenizer the narrator uses. Where the server reports its own prompt count, a turn that exceeded the window, or was evidently truncated, is recorded and shown, never silent. **Rationale.** §15.2 closed "budget larger than the window". It did not close "our count is not the server's count". Measured headroom was 23-42 tokens (item 1), and the fixed 64-token margin absorbs separators as well as tokenizer drift. A narrator whose tokenizer runs a few percent heavier than `cl100k_base` overflows a full 16,384 window. The failure deletes the canon at the front of the prompt with every request returning 200. The data to detect it is already stored with each turn and unread. **Scope.** - A safety reserve that scales with the effective budget and has a floor. It replaces or supplements `OUTPUT_SAFETY_MARGIN`, is defined once, and applies to verified windows, declared overrides and an unverified configured budget alike. The brief states the tokenizer divergence it is sized to absorb. - Reading the server-reported prompt-token count from the turn's usage. It is recorded beside the application's count in the turn's window provenance. Recorded states: - `fits`; - `exceeded`: reported count plus reply reserve is over the window; - `truncation_suspected`: the reported count is materially below the application's count; - `unknown`: no usage was reported. `exceeded` and `truncation_suspected` are surfaced in the context inspector, and in the turn's response so the reader sees a notice. - **Optional, decided in the brief:** calibrating the reserve from observed server-to-application ratios per endpoint and model. It must be bounded and quantised so that it does not re-price the prompt prefix every turn (§15.3). Detection is required either way. - `ContextOverflow` remains the explicit failure when protected context plus both reserves exceeds the budget. - `TECHNICAL-DESIGN.md` §15.2 and `DEVELOPMENT.md`'s context-window section, as implemented. The root `README.md` stale facts from item 25. **Non-scope.** The history block trim mechanism, section order, the knowledge budget share, a per-model tokenizer or any tokenizer download, a hard-coded window, raising any window, a provider abstraction, a Settings redesign, and failing or discarding a turn after the fact. **Likely affected.** `backend/app/context/builder.py`, `backend/app/contextwindow.py`, `backend/app/routers/adventures/turns.py` and `insights.py`, `backend/app/providers/openai_compatible.py` (usage, read-only), the context inspector panel in `frontend/src/pages/Play/`. Tests: `test_m11_context_window.py`, `test_m11_declared_window.py`, `test_history_block_trim.py`. **Acceptance criteria.** 1. **Per configuration.** The application's count plus the output reserve plus the safety reserve is no more than the effective budget, and the safety reserve is at least its documented value. This holds for each of: a verified 4,096 window, a verified 16,384 window, a declared override with no probe, and an unverified configured budget. There is a test per configuration. 2. **Divergence.** A scripted provider reports prompt counts at 1.00×, and at the brief's stated divergence, of the application's count, across a campaign that fills the window. At both, no turn's reported count plus reply reserve exceeds the window. At 1.25×, either no turn exceeds (calibration built) or every exceeding turn is recorded `exceeded`. **An unflagged overflow fails.** 3. **Truncation.** A scripted provider reports 8,194 tokens against an application count near 15,700. The turn is recorded `truncation_suspected` in stored provenance, shown in the inspector, and reported in the turn's response. A provider that reports no usage records `unknown`, never `fits`. 4. **Explicit overflow.** Protected context that fits without the safety reserve but not with it raises `ContextOverflow` with an actionable message. The story and state are unchanged (A05, L01). 5. **Canon survives.** `test_the_canon_at_the_front_survives_a_window_far_too_small` and `test_without_the_cap_the_same_prompt_would_have_overflowed` both pass with the reserve in place. 6. **Cache stability.** At steady state the history floor still moves in blocks (`test_history_block_trim.py`). If calibration is built, a test shows the budget changes only when the observed ratio crosses a stated step. 7. **Real model.** A long run of at least 50 turns runs at a 16,384 window with memory on, using the reference narrator on the GPU host. Its ten largest stored prompts are re-counted by the server, as §N did. Every one leaves at least the documented reserve, and every turn's provenance reads `fits`. The report carries the headroom table beside v1's. **Regression requirements.** The full backend and frontend suites; F03, F04 and M03 coverage; the I-series (a new provenance key must import and export); the offline container; the browser harness's F05 inspector checks. **Dependencies.** None. First. **Compatibility.** | Area | Effect | | --- | --- | | Databases | None expected. Provenance lives in the turn snapshot's JSON, so no migration | | Bundles | None. Snapshots carry new keys, and absence means "not recorded" | | Save Points, branches, state, memories, knowledge | None | | Model settings | None preferred. If a setting is added, it takes a default that reproduces the documented reserve, plus an additive migration | | Behaviour | At the largest prompts the history window holds slightly fewer actions. Nothing stored changes | | Docker and local-only | None | **Test modes.** Deterministic: yes. Real model: yes. Browser: inspector display, via the harness. Offline: regression only. Long run: yes, at least 50 turns. **Risk.** Low implementation risk and high value. Its blast radius is the budget arithmetic of every turn. It changes prompts at full windows, so it invalidates the v1 long-run evidence for v1.1 (§10). --- ### WP-A2 — Protocol echo hardening and genre-neutral instructions **Objective.** Stored narration no longer keeps the application-owned protocol shapes a narrator copies, and the fixed instructions stop supplying genre-specific identifiers. Recognition uses only strings and syntax the application itself owns, so story prose is not removed. **Rationale.** Items 3 and 4. The call notation in the prompt is copied, the length hint's echo escapes `_is_echoed_instruction`, and a lone trailing `Scene:` survives. Each leak is replayed as history, so it teaches the next turn. The example slugs were copied into an office scene. **Scope.** - A genre-neutral `EMIT_RULE` example and slug examples. No identifier in the fixed instructions names a fixture entity or a genre noun. - A decision, by measurement, on whether the vocabulary is shown in a notation that does not look callable. The extractor handles call lines regardless. - The length hint's wording becomes a named constant shared by the builder and the extractor, in the way `render.SECTION_HEADINGS` is shared. The extractor then recognises: - a line consisting only of a call to an event named in `events.SPECS` (case-insensitive), optionally `>`-quoted; - the length hint echoed as a trailing bracket, closed or cut off; - a renderer heading that is the reply's final line with nothing after it. - A replay tool that runs the old and new extractors over a corpus of stored AI turns and writes a per-turn diff report. - Measurement only: per-turn counts of state fact-line restatements, model-invented headings, and proposal outcomes (applied, empty, unparseable, refused), added to `tools/m11_long_run.py`. **Non-scope.** Removing a fact restated inside a sentence; removing model-invented headings that are not the renderer's; the event vocabulary itself, the validator, the fence protocol, or history replay; any model-based cleanup pass; rewriting stored narration of existing campaigns. **Likely affected.** `backend/app/narrative/extract.py`, `events.py` (prompt rendering only), `render.py` (constants), `backend/app/context/builder.py` (length-hint constant), `tests/test_narrative_state.py`, `tests/test_m11_scifi.py` (the J03 vocabulary check), `tools/m11_long_run.py`. Also `TECHNICAL-DESIGN.md` §15.4 and ADR 013's implementation note. **Acceptance criteria.** 1. **Known shapes are removed.** There is a committed case per shape, cut down from the four §S.6 turns and anonymised: - `> set_possession(silver-key, "alice")`; - `> Create_entity(...)`; - the length hint echoed as `[Hard limit: …append the state block well inside the limit.]`, closed and cut off; - a lone trailing `Scene:`. The four stored §S.6 texts, replayed, lose exactly those lines. Every other sentence is intact, asserted as set equality of the remaining sentences. 2. **Adversarial prose is untouched.** Each of these is a test: - dialogue that mentions `set_possession` mid-sentence; - `create_entity(ship)` inside a story's own ```python fence; - a call-shaped line whose name is not in the vocabulary, such as `> open_door(north)`; - an in-world bracket, `[Hard limit of the reactor: three hours]`; - `Scene:` followed by prose; - `Memory: she remembered the bells`; - a fact restated inside a sentence; - a trailing bracket about a state block that lacks the hint's wording. §O.8's existing negative controls all still pass. 3. **Replay corpus.** Every stored AI turn available from the v1 evidence runs is replayed through both extractors: §O.8's 443 and the identity run's 10, kept under `$HOME`, not committed. The package passes when all of these hold: - no turn changes except by removing a criterion-1 shape; - every changed turn appears in the diff report, and the WP report reviews it; - the review classifies zero changes as removed story. The report gives the counts. 4. **False-positive detection.** The replay tool flags any removal that is not the reply's final segment or a whole line matching a criterion-1 shape, and any removal of more than a stated share of a turn's prose. An unreviewed flag fails the package. 5. **Genre-neutral instructions.** A test asserts that `EMIT_RULE`, `EMIT_REMINDER`, the length hints and the vocabulary text contain no identifier from either acceptance fixture (Westhaven, Persephone) and none of J03's genre nouns. 6. **Real model.** - **Identity diagnostic.** The office fixture at 4,096 on the CPU/HTTPS host gives 0 proposals naming an example identifier (baseline 1 of 10), protocol shapes in 0 of 10 stored turns (baseline 4 of 10), and 0 identity signals. Its `--scripted --inject` self-test still fires. - **Long run.** At least 50 turns at 16,384 with memory on stores protocol in 0 turns (baseline 0 of 104). Restatement counts are recorded against the 74 of 104 baseline, as a measurement, not a gate. **Regression requirements.** The full backend suite, including the 25 §O.8 cases; C06 and H05 (refusals are still recorded and shown to the model); J01-J03; the I-series; the browser harness's hostile-narration checks (H04, H06, H07). **Dependencies.** After A1. Both change the prompt, and one real-model run can then cover both. The brief can be written while A1 is under review. **Compatibility.** | Area | Effect | | --- | --- | | Databases | None. No migration | | Stored narration of v1.0.0 campaigns | Not rewritten | | Bundles | None | | Save Points, branches | None | | State | None; the validator is unchanged | | Memories and summaries | Future ones summarise cleaner prose. Existing ones are untouched | | Docker and local-only | None | **Test modes.** Deterministic: yes. Replay corpus: yes. Real model: yes. Browser: regression only. Offline: regression only. Long run: yes, at least 50 turns, and this run may be the same one as A1's criterion 7 if A1 is already merged. **Risk.** Medium. The failure to fear is silent removal of prose. Constants-only recognition, the corpus and the detector are the mitigations. --- ### WP-B — Independent long-term memory retention **Objective.** An important fact planted early is recoverable at depth 100 or more from the memory bank itself, with provenance to a memory whose source range covers the planting turn. This must hold when neither authoritative state, later narration, the summary nor the history window carries the fact, and lineage safety must be unchanged. **Rationale.** Item 5. F02 and M04 pass on state-based recovery, which the owner accepted. Memory's own retention is unproven, and in a long campaign whose facts are not all state-shaped it is the only continuity there is. **Scope.** Diagnosis first, then the smallest sufficient fix. - **B.1, the diagnostic.** A retention harness that reports, for a planted fact, each stage with ids and depths: - **created:** does a memory whose `source_start`..`source_end` covers the planting turn contain the fact? - **retained:** is it not `forgotten`? - **ranked:** where is it in similarity order for the recall query? - **injected:** is it in the memories section? The harness runs in two modes. The deterministic mode uses a scripted narrator, summariser and embedder. The real-model mode uses the reference narrator and embedder. It also adds a `recovered_through_memory_independent` verdict to `tools/m11_long_run.py`, keeping the existing verdicts and their meanings. - **B.2, the fix,** only at the stages B.1 shows failing, from this bounded menu: - how the summariser's excerpt is chosen, instead of keeping only the last 2,000 tokens; - the memory prompt's instruction to keep named facts and objects; - an eviction rule that does not throw out a never-retrieved early memory first; - one additional ranking term (lexical or entity overlap, CONTEXT-AND-MEMORY §20). Anything outside the menu needs the owner's agreement in the brief. **Non-scope.** - Memory becoming authoritative, or outranking state (F07). - Any change to lineage attachment or filtering (`tree.attach_memory`, `forget_node`, E02) or summary lineage (E03). - Merging memory with imported knowledge, or new embedding models or dependencies. - Automatic re-summarisation of existing memories. - Redesigning cross-layer duplication (§22). **Likely affected.** `backend/app/memorybank.py` (creation, eviction, ranking); `backend/app/context/builder.py` (memories section, only if the query changes); `backend/app/models.py`, only if a field is unavoidable. Tools: `tools/memory_ab.py`, `tools/m11_long_run.py`. Tests: `test_context_memory.py`, `test_memory_nodes.py`, `test_m11_leakage.py`, `test_m11_long_run_memory.py`. **Acceptance criteria.** 1. **Deterministic retention test**, committed. Fact F is planted at depth 3 or less through narration only, with no state correction and no knowledge source. These are asserted for the whole run: - F is never in authoritative state; - no turn after the planting block contains F's sentinel tokens; - the summary section never contains F; - the planting turn is outside the history window at recall. At a recall depth of 100 or more, the memories section holds a memory that contains F and whose source range includes the planting depth, and the context report lists its id and similarity. **The test must fail on the `432f041` tree**, and the report must name the stage at which it fails. 2. **Past capacity.** The same test runs with more memories written than `memory_bank_capacity`. The planting-era memory is still active and retrieved, or its eviction follows a documented, tested rule that the WP report justifies. No pinned memory is evicted. The "frozen bank" regression (a new memory evicted at once) still passes. 3. **Lineage.** Fact G is planted only on a line later abandoned by Undo and divergence. G's memory stays on disk and is absent from every active-line prompt and from `memories.used`. All 14 tests in `test_m11_leakage.py` and the E02 tests pass unchanged. 4. **Authority.** A memory contradicting state loses, and memory is still framed as non-canon (F07). 5. **Real model.** A run of at least 100 turns at 16,384 with memory on, on the GPU host with logging. The fact is chosen so that the harness can verify the criterion-1 preconditions on real output. **Pass:** one run meets every precondition and returns `recovered_through_memory_independent`, with the memory's provenance. A run whose precondition fails reports which one, and counts as neither pass nor failure. The report gives the stage-by-stage diagnostic for every run. 6. **Budget.** The memories section stays inside its budget (F03), and the steady-state prompt is no larger than A1's arithmetic allows. 7. `CONTEXT-AND-MEMORY.md` §15, §20 and §21 are updated as implemented. **Regression requirements.** F01-F08, E01-E04, M04 verdicts (the new verdict is added and the old ones are unchanged), the I-series (memories travel in the bundle), L04 (`test_memory_rewrite.py`), §O.7's no-write-lock and failure-recording tests. **Dependencies.** After A2, so that it measures the prompt v1.1 ships and a narrator no longer handed protocol to restate. B.1's deterministic harness may be built earlier. **Compatibility.** | Area | Effect | | --- | --- | | Databases | No migration preferred. If a memory field is unavoidable, it is additive with a default, tested from a real v1.0.0 database, with schema parity | | Bundles | Stay v3. Any new memory field is optional on import, and its absence means "unknown". v1.0.0 bundles import | | Existing memories | Valid and used as they are. A rebuild stays opt-in (`tools/rewrite_memories.py`) | | Branch history, Save Points, state, knowledge | None | | Settings | Any change to capacity semantics keeps v1 defaults | | Docker and local-only | None | **Test modes.** Deterministic: yes. Real model: yes. Browser: no. Offline: regression only. Long run: yes, at least 100 turns, possibly several. **Risk.** High uncertainty, because the outcome depends on the model, and medium implementation risk. The blast radius is the memory subsystem. --- ### WP-C — Browser release coverage **Objective.** Every reader-facing behaviour that v1 proved only through the API or the component suite is driven in a real browser, including an export that actually leaves the browser as a file. **Rationale.** Item 6. §T carries two qualifications that exist only because the harness stops short. **Scope.** New scenarios in `tools/m11_browser.py`, or a sibling that reuses `tools/m11_webdriver.py`: Retry; Save Point create and restore; state correction; narration length; failed generation; export download. Also the Firefox profile preferences that direct a download to a harness-owned directory under `$HOME`. The package is harness-only, unless it finds a product defect or a control with no accessible name. Either is fixed with a regression test and reported as a product change. **Non-scope.** New UI, frontend refactors, Selenium or any new dependency, screenshot diffing, tablet layout, CI. **Likely affected.** `backend/tools/m11_browser.py`, `backend/tools/m11_webdriver.py`, the `DEVELOPMENT.md` harness section. **Acceptance criteria.** The run ends with 0 failed and 0 skipped. It runs against the built SPA served by FastAPI, with turns from the reference narrator over trusted-LAN HTTPS. Every assertion reads the rendered DOM or a file on disk. 1. **Retry.** Retry on the newest turn yields a second take, and the indicator reads 2/2. Stepping to 1/2 shows the original narration unchanged. The takes persist after a reload. 2. **Save Point.** Create a named Save Point through the UI, play two more turns, then restore it through the UI. The transcript ends at the named moment, the position indicator says later story is ahead, and Redo walks into the later turns. After a reload the Save Point is still listed. 3. **State correction.** A correction submitted through the State panel applies and is still shown after a reload. A partly refused correction shows its refusal and reason to the reader. 4. **Narration length.** After changing the control to brief and playing a turn, then to long and playing a turn, each turn's context inspector shows its band's word range. 5. **Failed generation.** With the model set through Settings to a name the server does not serve, a submitted turn shows an error, adds no narration and keeps the typed input. Setting the model back, the next turn succeeds and the earlier story is intact. 6. **Export download.** A real click on Export, from both the campaign library and campaign settings, writes a file with no manual step. The file: - exists and is not empty; - parses as `ai-dnd-adventure-v3`; - imports into a fresh data directory with the same action count, head position and Save Points. If the snap Firefox cannot be made to download, the check runs on a non-snap Firefox under `$HOME`, and `DEVELOPMENT.md` says so. **Exercising only the click does not pass.** 7. The existing 38 checks pass in the same run. **Regression requirements.** The existing browser checks; the frontend suite where a product fix is made. **Dependencies.** None. It can run at any point. Its final run is repeated on the v1.1 release candidate. **Compatibility.** None, unless a product fix is made, and then per §5. **Test modes.** Browser: yes. Real model: yes, over the HTTPS reference host. Deterministic: no. Offline: no. Long run: no. **Risk.** Low product risk. The medium risk is harness flakiness, so every wait is on a DOM condition, never a sleep used as an assertion. --- ### WP-D — Recovery honesty **Objective.** A backup is kept only after a full integrity check. A campaign whose export exceeds what import accepts is exported with a warning saying so, not silently. **Rationale.** Items 7 and 10a. Both are recovery gaps that are known, measured and cheap to close. Neither needs a policy decision. **Scope.** - `backup.py` runs `PRAGMA integrity_check` on the finished copy, and the time it takes is measured. - Export compares the serialised bundle's size with `limits.MAX_IMPORT_BODY_BYTES`. When it is over, the file is still delivered and the reader sees a warning naming the limit and what it means. - The import refusal for an oversized bundle names the limit. - `DEVELOPMENT.md` documents the ceiling as measured: about 13 kB per action on a real campaign, and M9's conservative figure of about 279 turns. **Non-scope.** Scheduled backups; raising the import limit or streaming import; a bundle format change or further compression; a restore button. **Likely affected.** `backend/app/backup.py`, `backend/app/routers/adventures/bundle_io.py`, `frontend/src/pages/Campaigns.jsx`, `frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx`, `frontend/src/pages/backup.test.jsx`, and the backup and bundle tests. **Acceptance criteria.** 1. A backup of a healthy database reports `integrity_check` ok and is kept. 2. A copy with damage that `integrity_check` detects and `quick_check` does not, such as an index inconsistent with its table, is rejected and not kept. Existing backups are still never overwritten. 3. The time for `integrity_check` is recorded on the 100-turn evidence database and on a synthetic database of 100 MB or more, and the backup completes through the UI on both. 4. Exporting a fixture campaign over 20 MB succeeds, delivers the file, and shows a warning naming the import limit, asserted in the API response and in a component test. A campaign under the limit shows no warning. 5. Importing that bundle is refused with a message naming the limit. 6. A normal campaign's export is byte-identical before and after the package, apart from timestamp fields. **Regression requirements.** I01-I07, L01-L04, the backup case in `test_m11_migration.py`, `backup.test.jsx`, and the offline container's export and import. **Dependencies.** None. **Compatibility.** Databases and bundles: no change. Backups: a stricter check, with the same file. **Test modes.** Deterministic: yes. Browser: the warning, optionally in WP-C's harness. Offline: regression only. Real model: no. Long run: no. **Risk.** Low. --- ### WP-E — Control-boundary contrast **Objective.** Control boundaries meet WCAG 1.4.11's 3:1 against their panel, at rest and on hover. **Rationale.** Item 9. **Scope.** - The boundary token values in `frontend/src/styles/tokens.css`, and any component that overrides them. - `tools/contrast_audit.py` treats boundary pairs below 3:1 as failures, not advisories. - The browser harness measures rendered boundary contrast. - Before-and-after screenshots for the owner's approval. **Non-scope.** A palette redesign, typography, layout, tablet work, a screen-reader audit, and other WCAG criteria. A defect the same measurement finds is recorded, not taken on. **Likely affected.** `frontend/src/styles/tokens.css`, `backend/tools/contrast_audit.py`, `backend/tools/m11_browser.py`, `frontend/src/a11y.test.jsx`. **Acceptance criteria.** 1. `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and exits 0 on the package tree. 2. Every text pair still clears 4.5:1 (1.4.3). The rendered text contrasts measured by the harness do not fall below v1's (14.57, 5.48, 13.57 and 5.88:1) without a stated reason. 3. The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1. The focus indicator is still visible. 4. The owner approves the before-and-after screenshots, and the WP report records the approval. **Regression requirements.** The frontend suite and lint; the browser harness's accessibility checks. **Dependencies.** None. Its browser checks go into WP-C's harness if WP-C has landed, and into `m11_browser.py` otherwise. **Compatibility.** None. **Test modes.** Deterministic: yes. Browser: yes. Everything else: no. **Risk.** Low. --- ## 9. Order and dependencies ```text A1 context safety reserve ──► A2 protocol echo + neutral instructions ──► B memory retention (B.1 harness may start early) C browser coverage ── independent D recovery honesty ── independent E boundary contrast ── independent (checks land in C's harness if C is first) │ ▼ v1.1 release validation (§10) ``` **Recommended order: A1, A2, B, C, D, E.** - **A1 first.** It is the only item that can silently remove canon from a prompt. It is deterministic to test, small in blast radius, and it establishes the provenance that A2's and B's real-model runs will read. - **A2 second.** Its leak is silent and compounds through replay. It must precede B, because it changes what the memory pass summarises. - **B third.** It carries the highest continuity value and the most uncertainty. Measuring it before A1 and A2 settle would measure a prompt v1.1 does not ship. - **C, D and E** close verification, recovery and accessibility gaps with no known story risk. They depend on nothing, so the owner may move any of them earlier. While B waits on long runs is a natural slot. Each still has its own brief and review. | Property | A1 | A2 | B | C | D | E | | --- | --- | --- | --- | --- | --- | --- | | Can begin independently | yes | after A1 | after A2 | yes | yes | yes | | Schema or migration | no | no | avoid; additive if unavoidable | no | no | no | | Export format | no (additive evidence key) | no | no; optional field at most | no | no | no | | Changes acceptance tests (`V1-ACCEPTANCE-TESTS.md`) | no | no | no | no | no | no | | Changes the stored prompt | yes | yes | possibly | no | no | no | | Real-model validation | yes | yes | yes | yes, for turns | no | no | | Browser testing | regression | regression | no | yes | optional | yes | | Offline / no-network testing | regression | regression | regression | no | regression | no | | Long-run testing | yes, 50 turns or more | yes, 50 turns or more | yes, 100 turns or more | no | no | no | | Risk | low | medium | high uncertainty | low | low | low | ## 10. Compatibility summary | Area | A1 | A2 | B | C | D | E | | --- | --- | --- | --- | --- | --- | --- | | v1.0.0 campaign databases | open unchanged | unchanged | unchanged; an additive migration only if unavoidable | — | unchanged | — | | Export and import bundles | v3; new evidence key | — | v3; an optional field at most | — | v3; export warns | — | | Save Points | — | — | — | — | — | — | | Branch history | — | — | lineage rules unchanged | — | — | — | | Narrative state | — | validator unchanged | memory never outranks state | — | — | — | | Memories and summaries | — | new ones from cleaner prose | creation, retention and ranking change; existing rows kept | — | — | — | | Knowledge sources | — | — | — | — | — | — | | Model settings | none preferred | — | capacity defaults kept | — | — | — | | Docker and local-only | — | — | — | — | — | — | ## 11. v1.1 release criteria v1.1 is not called v1.1.0 until all of the following hold on one release candidate tree: 1. **Every package in scope is accepted**, each with its report, and with its acceptance criteria passing on the candidate or on a tree whose product code the candidate carries unchanged. 2. **The v1 contract still passes.** All 82 REQUIRED FOR V1 tests hold, with H09 not applicable on the same condition. None is relaxed. 3. **Suites and builds:** the backend suite, the frontend suite and lint, the production build, and `docker build --no-cache`, with the image's SPA identical to the local build. 4. **Offline:** `tools/m11_offline.py` passes all checks with no network and a fresh volume. 5. **Browser:** the existing 38 checks plus WP-C's and WP-E's pass with 0 failed and 0 skipped, on the candidate, over trusted-LAN HTTPS. 6. **One v1.1 long run** on the candidate's product code: 100 turns or more at a 16,384 window with memory on, on the GPU host with logging. It passes M01-M04. A1's headroom table shows the documented reserve on every re-counted prompt, and every turn is `fits`. A2's leak count is 0. B's independent-retention verdict is recorded. 12. **WP-B's memory limitation is reported, not summarised away.** The v1.1 release report states each of these, and never shortens them to "WP-B passed": - deterministic independent-memory recovery: **PASS**; - reference-model independent-memory recovery: **FAIL** on the precondition-valid attempt; - the failing stage: **memory creation**, the summariser's content selection; - the owner's decision to accept that limitation for v1.1. The release long run's independent-retention verdict (item 6) is read against it. A recovery there is reported as evidence, not as a reversal of the limitation, unless it meets every isolation precondition. 13. **Carried residuals are listed with their status:** - the mid-reply narrator instruction echo that A2's trailing cleanup does not remove (WP-B.1 §K); - the doubled full stop in the memory-search scene text (WP-B.2 §R 10). 7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0 protocol shapes in stored narration. 8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all checks. 9. **Upgrade from a real v1.0.0 database.** A database is created by the `v1.0.0` tree and played. It has retained history, an undone head, Save Points, memories, summaries, imported knowledge and a narration-length choice, and uses loopback or placeholder settings with no real hostnames. Opened by the candidate, its transcript, head, Redo availability, state, Save Points, memories, knowledge and settings compare identical, and any migration is forward-only with schema parity. A v1.0.0 export imports into v1.1. Because the format stays v3, a v1.1 export of that campaign is checked for import into v1.0.0, and the result is reported. 10. **Documentation:** `README.md`, `DEVELOPMENT.md`, the as-implemented sections named by each package, and `VERSION.md`. 11. **Owner events**, each separate: the signed release commit, `main`, and a `v1.1.0` tag. ## 12. Scope recommendation **Recommended: Option 1, a focused v1.1.** | Ship in v1.1 | Defer to v1.2 | Future / optional | | --- | --- | --- | | WP-A1 context safety reserve | Scheduled backups (10b) | Real media-provider adapters (11; K05, K06) | | WP-A2 protocol echo and neutral instructions | Raising or streaming the import limit (7) | Whole-transcript copy, story search (§77, §78) | | WP-B independent memory retention | Reducing in-sentence restatement (2) | Discarded-history recovery screen (§63) | | WP-C browser release coverage | Stale-scene and derived-contamination detectors in the identity diagnostic (8) | Tablet layout | | WP-D recovery honesty | Cross-layer duplication suppression (17) | Editable RPG world state | | WP-E control-boundary contrast | | Media debt: `ambience`, visual-profile recovery and UI (20) | | | | Trusted-LAN residual limits: address pinning against DNS rebinding (21) | | | | Comparative narrator-model recommendation (13) | **Why focused.** Four of the six packages close silent failure modes or verification gaps that the v1 evidence itself named. The other two are small and deterministic. Everything deferred is either a new feature, or needs a policy decision the owner has not been asked for, or rests on an occurrence that has not happened. A broader v1.1 that took scheduled backups and media would add the two packages with the most new surface. It would also push the long-run re-validation, which every prompt change already requires, further from the changes it validates. **If B's real-model criterion cannot be met.** Precondition-valid runs may recover nothing even after B.2. In that case the owner chooses between shipping v1.1 with B's deterministic criteria met and the real-model result recorded as a residual risk, or holding v1.1 for B. The plan does not pre-decide this. ## 13. Coding briefs needed next These briefs are not written here, and none is to be executed from this document. 1. **WP-A1 — context-window safety reserve. Write this one first.** It carries the highest user risk, has no dependencies, is deterministically testable, and produces the provenance that later real-model runs read. The brief must settle three things: the tokenizer divergence the reserve is sized for, whether calibration is built or detection alone, and whether any setting is added (recommended: none). 2. WP-A2 — protocol echo hardening and genre-neutral instructions. It can be drafted while A1 is under review. 3. WP-B — independent memory retention: B.1 diagnostic, then B.2 fix. Two briefs are an option if B.1's findings should be reviewed before a fix is chosen. 4. WP-C — browser release coverage. 5. WP-D — recovery honesty. 6. WP-E — control-boundary contrast. 7. v1.1 release validation, written only after every in-scope package is accepted. ## 14. Decisions for the owner - Confirm Option 1, the focused v1.1. - A1: the divergence the reserve must absorb; calibration or detection only; that a detected truncation is flagged, not turned into a failed turn. - B: whether B.1 and B.2 are one brief or two; the choice in §12 if the real-model criterion is not met. - D: that the 20 MB import limit stays in v1.1. - E: approval of the visual change. - Whether v1.1 real-model validation stays on the reference narrator, as this plan recommends. - The report naming and rotation in `planning/reports/` (§5 rule 5).