Files
interactive-story/planning/V1.1-PLAN.md
T
JesseMarkowitzandClaude Opus 5 ac465ed867 Planning v4.1: record the v1.0.0 release, and plan v1.1
Documentation only. No product code, requirement, acceptance test or
schema changes.

Post-release correction. v4.0 was written before the closeout commit was
signed (432f041), main was fast-forwarded to it, and the signed v1.0.0 tag
was pushed. Current-state wording now says so in README.md,
planning/README.md, BUILD-MILESTONES.md and VERSION.md. BUILD-MILESTONES.md's
header had been stale since M8. The M11 report is not edited: its §T
records the state at closeout.

v1.1 plan. planning/V1.1-PLAN.md triages the post-v1 backlog and the other
recorded v1 residual risks, and orders them into work packages, not
milestones:
- A1: a context-window safety reserve, plus reporting a turn the server
  truncated
- A2: removing protocol echoes from stored narration, and a genre-neutral
  state rule
- B: long-term memory retention that holds without help from state
- C: browser coverage of Retry, Save Points, correction, length, failure
  and export download
- D: integrity_check on backups, and a warning when an export exceeds the
  import limit
- E: WCAG 1.4.11 control-boundary contrast

Scheduled backups, the import limit, identity detectors and duplication
suppression move to v1.2; media adapters are future work. The first brief
to write is A1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 08:19:31 -04:00

52 KiB
Raw Blame History

Adventure Storyteller — v1.1 Plan

Status: PLANNING. Written 2026-09-14 on v1.1-development, from the signed v1.0.0 release commit 432f041. No work package has started, and no v1.1 version or tag exists.

This document replaces nothing. BUILD-MILESTONES.md stays as the closed v1 history (M1-M11), and its Post-v1 backlog is this plan's input. The v1 specification and acceptance contract are unchanged: every package below improves how an existing requirement is met. No package adds a requirement.


1. Terminology

  • Work package (WP). One independently reviewable unit of v1.1 work, with its own coding brief, its own report, and its own review. Work packages are lettered. They are not milestones, and there is no M12.
  • Brief. The coding prompt for one work package. It holds only the objective, the source documents, the scope and non-scope, the acceptance criteria, and the stop condition. No work package begins before its brief exists.
  • Reference narrator. qwen2.5:3b-instruct, with and without num_ctx 16,384 baked in, and nomic-embed-text for embeddings. These are the models the v1 evidence used (M11 report §E.1). v1.1 real-model evidence uses the same models so that its numbers compare with v1's.
  • Reference hosts. The CPU reference host serves Ollama over HTTPS with a private CA, which is the A06 evidence path. The GPU inference host serves plain HTTP on the LAN and is used for long runs (M11 report §E.1).
  • Planning package version (VERSION.md, v4.x) and product version (v1.0.0, v1.1.0) are separate numbers. SECURITY-THREAT-MODEL.md's own "Status: v1.1" line is that document's revision label from Phase 0B. It does not refer to this release.

2. Baseline

Release v1.0.0, 2026-09-14
Release commit 432f04100b9a67198bcdc46c6ff8ee0f181e1667, signed by the owner
Tag v1.0.0, a signed annotated tag on that commit, pushed
main 432f041, the same commit
v1.1 branch v1.1-development, created at 432f041
v1 contract 82 REQUIRED FOR V1 tests: 81 PASS and H09 NOT APPLICABLE (M11 report §T)
Schema migrations up to 94 (LATEST_VERSION 94)
Bundle format ai-dnd-adventure-v3
Provenance the AI-DnD fork point d72f7c1b is an ancestor. LICENSE and PROVENANCE.md are unchanged since 1013c94

3. v1.1 goals

  1. No silent loss of what the narrator is given. A prompt must not reach the server's window edge by an accident of tokenizer arithmetic. A turn the server truncated must be reported.
  2. No protocol in the story. The shapes of the application's own protocol that a narrator copies must not be stored as narration. The fixed instructions must not hand every campaign one genre's nouns to copy. Story prose must not be removed in the process.
  3. Memory that remembers on its own. An old fact must be recoverable from the memory bank itself, with provenance, when nothing else carries it. This must not weaken lineage safety.
  4. Release evidence without API-only gaps. Every reader-facing history, state, length, failure and export behaviour must be driven in a real browser.
  5. Recovery that says when it cannot recover. Backups get a full integrity check, and an export that import would refuse must say so.
  6. Control boundaries that meet WCAG 1.4.11.

And throughout: every v1.0.0 campaign opens unchanged in v1.1.

4. Non-goals

  • New product features: media generation, TTS, STT, story search, whole-transcript copy, a discarded-history recovery screen, a restore button, tablet redesign.
  • Changes to the history, branch, head or Save Point architecture (ADR 005, ADR 012).
  • Changes to the authoritative state model, its event vocabulary or its validator (ADR 010, ADR 013). Changes to the prompt wording that describes the vocabulary are in scope, in WP-A2.
  • Changes to the knowledge authority classes or their retrieval.
  • A new bundle format version.
  • Any new runtime dependency, any network access beyond the configured inference endpoint, and any first-use download, including per-model tokenizers.
  • Supporting inference servers other than Ollama, beyond what context_window_override already allows.
  • Re-writing stored narration or memories of existing campaigns automatically.
  • Redesigning identity or state handling on the strength of post-M8 finding D, which has not been reproduced.

5. Rules for every work package

The v1 milestone rules (BUILD-MILESTONES.md §2) apply to every work package, and so do the following:

  1. Compatibility default. Existing v1.0.0 campaign databases open unchanged. Any migration is forward-only and additive, and is tested by opening a real v1.0.0 database. Fresh-install and upgraded schemas must still compare identical (test_m11_migration.py). v1.0.0 bundles still import.
  2. The bundle stays ai-dnd-adventure-v3 unless a brief makes the case under M9's semantic test and the owner agrees. Adding a key inside stored evidence is not a format change when an absent key unambiguously means "not recorded".
  3. The v1 contract is the regression floor. No REQUIRED FOR V1 test is retired, relaxed or reclassified.
  4. Evidence discipline (M11 report §O). Any product change made after a black-box run was taken invalidates that run for release purposes. Harness defects and product defects are reported separately. A check that cannot fail is not evidence.
  5. Each package has a report (planning/reports/V1.1-WP-<id>-REPORT.md). It records what was built, the acceptance results with measurements, what was not verified, and the as-implemented documentation changes. The report rotation for reports/ is decided when the first one is written (planning/README.md).
  6. Real-model runs longer than a few minutes on the GPU host require the power, link and kernel logging in DEVELOPMENT.md, started before the run.
  7. No real hostnames, addresses or people's names in committed files, fixtures or reports. Evidence stays under $HOME, never /tmp.
  8. Commits and tags are the owner's. A package ends with its changes staged for a signed commit.

6. Backlog triage

Severity reflects user risk: High means story corruption or lost continuity that the reader is not told about. Medium means silent degradation or a real verification gap. Low means visible, rare or cosmetic.

6.1 The recorded post-v1 backlog

# Item Source Current evidence Severity Disposition
1 Context-window safety margin M11 §P risk 3, §N; post-v1 backlog The largest prompts left 23, 34 and 42 real tokens of headroom in the GPU runs. The application's cl100k_base count ran 16 tokens below the narrator's on every prompt measured. The only slack is OUTPUT_SAFETY_MARGIN = 64 (context/builder.py), which is fixed and must also absorb the section separators. Ollama 0.34 cut an over-window synthetic prompt to 8,194 tokens with no error. Truncation drops the oldest tokens, which are the narrator's rules and the canon. The server's reported usage is already stored per turn (snapshot["usage"], routers/adventures/turns.py) and read by nothing. High: silent, invisible, and reachable by a model whose tokenizer diverges further v1.1, WP-A1
2 Narrator restating prompt and state text §P risk 4, §G.4 The state's fact line was restated as a sentence on 74 of 104 turns in the evidence run. Phrases such as "Scene set, continue your adventure." and lines opening "Memory:" are model-invented, and none is an application string (verified by search). The restatements sit inside story sentences. Medium: narration quality, and it is the path by which M04's fact reached memory, which masks item 5 v1.1, WP-A2 measures it. Reducing in-sentence restatement is v1.2, because removing it means judging prose
3 Protocol leakage: event-call syntax, stray headings, repeated length hints §P risk 6, §S.6 On 4 of 10 turns of the closeout identity run (4,096 window) and 0 of 104 in the 100-turn runs. Causes verified in code. events.vocabulary_for_prompt shows every event as name(field, …), which is the notation copied. _is_echoed_instruction requires both "state block" and "events list", and the length hint's tail says only "state block". A lone trailing Scene: is not removed. Stored text is replayed as history (TECHNICAL-DESIGN §15.4). Medium-high: silent and self-reinforcing, but rare at a full window v1.1, WP-A2
4 Genre-specific example in the fixed state rule §P risk 16; narrative/extract.py EMIT_RULE The example names mara, silver-key, old-abbey and aldric. In the office-meeting identity run the narrator proposed giving silver-key to Alice, and the validator refused it. No state was corrupted. Medium: every non-fantasy campaign gets fantasy nouns to copy v1.1, WP-A2
5 Independent long-term-memory retention §P risk 5, §G.4; F02 and M04 qualifications No run showed a planting-era memory carrying the planted fact. Recovery ran from state to the narrator's restatement to memory. Code facts that bear on it, none yet proven to be the cause: a memory is at most 50 words for a 6-action block. The summariser sees the block's last 2,000 tokens (truncate_to_last_tokens), so an early fact in a long block can be cut. Eviction is least-recently-used above memory_bank_capacity (default 80), so a never-retrieved early memory is the first to go once a campaign passes about 480 actions; the 33-memory evidence run never reached that. Ranking is cosine similarity plus a pin (CONTEXT-AND-MEMORY §20). High for long campaigns: lost continuity, seen only as the story forgetting v1.1, WP-B
6 Browser automation gaps: Retry, Save Point, state correction, narration length, failed generation, export download §S.3, §T qualifications, §P risk 11 Proved through the API, the component suite and the 100-turn campaign, but never driven in a browser. The export control is exercised only as far as the click, because a blob: download does not leave the headless snap Firefox. No defect is known. Medium: a verification gap on the release path v1.1, WP-C
7 Practical export and bundle-size ceiling §P risk 12, §K; M9 debt The import body limit is 20 MB (limits.MAX_IMPORT_BODY_BYTES). M9's conservative ceiling is about 279 turns. The real 100-turn campaign was 2.74 MB, about 13 kB per action, or roughly 1,600 actions. A larger campaign exports and is then refused on import, and nothing says so at export. Medium impact, low frequency v1.1, WP-D: warn at export. Raising the limit or a streaming import is v1.2
8 Identity-confusion diagnostic follow-up §P risk 8, §S.5 Not reproduced. The diagnostic cannot see a stale scene, has never run with memory on, and has never run at a 16,384 window. It found no product deficiency. Low: no occurrence since the original, whose campaign is gone No v1.1 package. The diagnostic is re-run as a WP-A2 regression and in the v1.1 release gate, with memory on. A stale-scene detector is v1.2. The next real occurrence is classified with tools/m11_identity.py
9 WCAG 1.4.11 control-boundary contrast §P risk 10, §L Measured at 1.33:1 resting and 1.75:1 on hover (--border and --border-bright against --bg-panel). M11's argument that the label identifies the control is weakest for text inputs, where the boundary shows where to type. Low-medium v1.1, WP-E
10a Backup integrity: integrity_check §P risk 14; M9 backup.py runs PRAGMA quick_check, which skips index-content verification. Low v1.1, WP-D
10b Scheduled backups §P risk 14; M9 None exist; backups are manual. They need owner policy on interval, retention, location and disk use, and must not contend with a turn's write lock (§O.7). Medium impact, but a new feature v1.2
11 Real media-provider adapters §P risk 13; M10; K05 and K06 (FUTURE) The seam has no consumer. Adding one brings dependencies, a stricter endpoint rule (DEVELOPMENT.md) and a new UI. n/a: a feature, not a risk to existing stories Future / optional, not v1.1

6.2 Other recorded v1 residual risks and debt

# Item Source Disposition
12 The state lags the narration at 4,096 with the 3B narrator (1 proposal in 10 applied) §P risk 17, §S.5 Model behaviour, not a defect. WP-A2's real-model runs record the proposal outcome counts. No package
13 Every realistic observation is a 3B model's §P risk 2 No package. v1.1 keeps the reference narrator for comparability. A comparative model recommendation stays open (planning/README.md, Still open)
14 A campaign that never chose a narration length keeps the pre-M11 hint §P risk 15 Deliberate. No change. WP-C drives the control
15 The GPU host dropped off the PCIe bus after the evidence run §P risk 7, §E.1 Operational. Rule 6 in §5 applies to every long run
16 Release timings are two hosts' §P risk 1 No action. No performance requirement exists or is invented
17 Cross-layer duplication: one fact in state, memory, history and a passage at once CONTEXT-AND-MEMORY §22; open since M6 and M7 v1.2. It costs tokens (item 1) and is related to item 2. WP-B may touch ranking only where its diagnostic requires
18 Memory-ranking factors beyond similarity and pin not implemented CONTEXT-AND-MEMORY §20 Inside WP-B's bounded fix menu, only if WP-B's diagnostic places the failure at ranking
19 M8 carried debt: no whole-transcript copy or story search (§77, §78), no discarded-history recovery screen (§63), tablet untuned, RPG world state read-only BUILD-MILESTONES.md M8 and M9 Future backlog. Features, not reliability
20 M10 debt: ambience is an empty shape; a deleted visual profile is unrecoverable in the app; profiles are API-only M11 §D Future, with media
21 A hostile host on the trusted LAN; the DNS-rebinding interval SECURITY-THREAT-MODEL.md §10A Accepted for v1. Future. No v1.1 package
22 A stale chunk_id in a restored snapshot M9 No action. It is a React key, not a live pointer
23 Developer-venv residue (quickjs, psycopg) §M No action. It does not ship
24 A Settings warning comparing the real window to the budget M8 debt, "unowned" Closed by M11: Settings reports the window and warns
25 Root README.md has stale facts: "1,191 backend tests" (1,421 at closeout), and the Screenshots paragraph calls screenshots "a job for the UI pass in M8" found in this review Documentation debt, not release status, so it is not changed in v4.1. It is fixed in WP-A1's documentation update

7. Grouping decisions

The suggested "narrator boundary hardening" package is split into A1 and A2. They share a theme but not a mechanism or a test method.

  • A1, the context reserve, is budget arithmetic in context/builder.py and contextwindow.py. It is proved deterministically with a scripted provider that reports token counts.
  • A2, the protocol echo, is the contract between the prompt's wording (extract.EMIT_RULE, events.vocabulary_for_prompt, the length hint) and the extractor. It is proved with a replay corpus of real narration and real-model runs.

Bundling them would make one review carry two unrelated risk profiles.

A2's three items belong together. The example slugs and the call notation are the source of the copied text, and the extractor is its sink. Changing one without the other means proving the change twice. Both also alter the stored prompt, so they invalidate the same evidence.

B stands alone. It is the only package whose success depends on model output quality. Its test design, a fact that nothing but memory carries, is the hard part. It follows A2, so that it measures the prompt v1.1 will ship.

D is narrowed to recovery honesty, meaning integrity_check and the export warning. Both are small and deterministic, and both are recovery-path changes with no policy question. Scheduled backups need owner decisions and have real design risk, including write-lock contention. Raising or streaming the import limit changes a DoS guard. Both move to v1.2.

E stays focused on control boundaries. The same measurement found no other accessibility defect (M11 §L).

F is not a package. Nothing concrete in the product is deficient. The diagnostic's known blind spots are a stale scene, memory off, and one window size. The first two can be addressed by running it differently: the v1.1 gate runs it with memory on. The stale-scene detector is new diagnostic work with no occurrence to justify it yet, so it is v1.2.

G is not in v1.1. It adds no reliability or quality to existing stories, it brings dependencies and a new endpoint surface, and its tests (K05, K06) are FUTURE. MEDIA-EXTENSION-CONTRACT.md and the M10 seam remain authoritative for whenever it is taken up.


8. Work packages

WP-A1 — Context-window safety reserve

Objective. An assembled prompt plus its reply reserve stays at least a documented safety reserve below the effective window, whatever tokenizer the narrator uses. Where the server reports its own prompt count, a turn that exceeded the window, or was evidently truncated, is recorded and shown, never silent.

Rationale. §15.2 closed "budget larger than the window". It did not close "our count is not the server's count". Measured headroom was 23-42 tokens (item 1), and the fixed 64-token margin absorbs separators as well as tokenizer drift. A narrator whose tokenizer runs a few percent heavier than cl100k_base overflows a full 16,384 window. The failure deletes the canon at the front of the prompt with every request returning 200. The data to detect it is already stored with each turn and unread.

Scope.

  • A safety reserve that scales with the effective budget and has a floor. It replaces or supplements OUTPUT_SAFETY_MARGIN, is defined once, and applies to verified windows, declared overrides and an unverified configured budget alike. The brief states the tokenizer divergence it is sized to absorb.

  • Reading the server-reported prompt-token count from the turn's usage. It is recorded beside the application's count in the turn's window provenance. Recorded states:

    • fits;
    • exceeded: reported count plus reply reserve is over the window;
    • truncation_suspected: the reported count is materially below the application's count;
    • unknown: no usage was reported.

    exceeded and truncation_suspected are surfaced in the context inspector, and in the turn's response so the reader sees a notice.

  • Optional, decided in the brief: calibrating the reserve from observed server-to-application ratios per endpoint and model. It must be bounded and quantised so that it does not re-price the prompt prefix every turn (§15.3). Detection is required either way.

  • ContextOverflow remains the explicit failure when protected context plus both reserves exceeds the budget.

  • TECHNICAL-DESIGN.md §15.2 and DEVELOPMENT.md's context-window section, as implemented. The root README.md stale facts from item 25.

Non-scope. The history block trim mechanism, section order, the knowledge budget share, a per-model tokenizer or any tokenizer download, a hard-coded window, raising any window, a provider abstraction, a Settings redesign, and failing or discarding a turn after the fact.

Likely affected. backend/app/context/builder.py, backend/app/contextwindow.py, backend/app/routers/adventures/turns.py and insights.py, backend/app/providers/openai_compatible.py (usage, read-only), the context inspector panel in frontend/src/pages/Play/. Tests: test_m11_context_window.py, test_m11_declared_window.py, test_history_block_trim.py.

Acceptance criteria.

  1. Per configuration. The application's count plus the output reserve plus the safety reserve is no more than the effective budget, and the safety reserve is at least its documented value. This holds for each of: a verified 4,096 window, a verified 16,384 window, a declared override with no probe, and an unverified configured budget. There is a test per configuration.
  2. Divergence. A scripted provider reports prompt counts at 1.00×, and at the brief's stated divergence, of the application's count, across a campaign that fills the window. At both, no turn's reported count plus reply reserve exceeds the window. At 1.25×, either no turn exceeds (calibration built) or every exceeding turn is recorded exceeded. An unflagged overflow fails.
  3. Truncation. A scripted provider reports 8,194 tokens against an application count near 15,700. The turn is recorded truncation_suspected in stored provenance, shown in the inspector, and reported in the turn's response. A provider that reports no usage records unknown, never fits.
  4. Explicit overflow. Protected context that fits without the safety reserve but not with it raises ContextOverflow with an actionable message. The story and state are unchanged (A05, L01).
  5. Canon survives. test_the_canon_at_the_front_survives_a_window_far_too_small and test_without_the_cap_the_same_prompt_would_have_overflowed both pass with the reserve in place.
  6. Cache stability. At steady state the history floor still moves in blocks (test_history_block_trim.py). If calibration is built, a test shows the budget changes only when the observed ratio crosses a stated step.
  7. Real model. A long run of at least 50 turns runs at a 16,384 window with memory on, using the reference narrator on the GPU host. Its ten largest stored prompts are re-counted by the server, as §N did. Every one leaves at least the documented reserve, and every turn's provenance reads fits. The report carries the headroom table beside v1's.

Regression requirements. The full backend and frontend suites; F03, F04 and M03 coverage; the I-series (a new provenance key must import and export); the offline container; the browser harness's F05 inspector checks.

Dependencies. None. First.

Compatibility.

Area Effect
Databases None expected. Provenance lives in the turn snapshot's JSON, so no migration
Bundles None. Snapshots carry new keys, and absence means "not recorded"
Save Points, branches, state, memories, knowledge None
Model settings None preferred. If a setting is added, it takes a default that reproduces the documented reserve, plus an additive migration
Behaviour At the largest prompts the history window holds slightly fewer actions. Nothing stored changes
Docker and local-only None

Test modes. Deterministic: yes. Real model: yes. Browser: inspector display, via the harness. Offline: regression only. Long run: yes, at least 50 turns.

Risk. Low implementation risk and high value. Its blast radius is the budget arithmetic of every turn. It changes prompts at full windows, so it invalidates the v1 long-run evidence for v1.1 (§10).


WP-A2 — Protocol echo hardening and genre-neutral instructions

Objective. Stored narration no longer keeps the application-owned protocol shapes a narrator copies, and the fixed instructions stop supplying genre-specific identifiers. Recognition uses only strings and syntax the application itself owns, so story prose is not removed.

Rationale. Items 3 and 4. The call notation in the prompt is copied, the length hint's echo escapes _is_echoed_instruction, and a lone trailing Scene: survives. Each leak is replayed as history, so it teaches the next turn. The example slugs were copied into an office scene.

Scope.

  • A genre-neutral EMIT_RULE example and slug examples. No identifier in the fixed instructions names a fixture entity or a genre noun.
  • A decision, by measurement, on whether the vocabulary is shown in a notation that does not look callable. The extractor handles call lines regardless.
  • The length hint's wording becomes a named constant shared by the builder and the extractor, in the way render.SECTION_HEADINGS is shared. The extractor then recognises:
    • a line consisting only of a call to an event named in events.SPECS (case-insensitive), optionally >-quoted;
    • the length hint echoed as a trailing bracket, closed or cut off;
    • a renderer heading that is the reply's final line with nothing after it.
  • A replay tool that runs the old and new extractors over a corpus of stored AI turns and writes a per-turn diff report.
  • Measurement only: per-turn counts of state fact-line restatements, model-invented headings, and proposal outcomes (applied, empty, unparseable, refused), added to tools/m11_long_run.py.

Non-scope. Removing a fact restated inside a sentence; removing model-invented headings that are not the renderer's; the event vocabulary itself, the validator, the fence protocol, or history replay; any model-based cleanup pass; rewriting stored narration of existing campaigns.

Likely affected. backend/app/narrative/extract.py, events.py (prompt rendering only), render.py (constants), backend/app/context/builder.py (length-hint constant), tests/test_narrative_state.py, tests/test_m11_scifi.py (the J03 vocabulary check), tools/m11_long_run.py. Also TECHNICAL-DESIGN.md §15.4 and ADR 013's implementation note.

Acceptance criteria.

  1. Known shapes are removed. There is a committed case per shape, cut down from the four §S.6 turns and anonymised:

    • > set_possession(silver-key, "alice");
    • > Create_entity(...);
    • the length hint echoed as [Hard limit: …append the state block well inside the limit.], closed and cut off;
    • a lone trailing Scene:.

    The four stored §S.6 texts, replayed, lose exactly those lines. Every other sentence is intact, asserted as set equality of the remaining sentences.

  2. Adversarial prose is untouched. Each of these is a test:

    • dialogue that mentions set_possession mid-sentence;
    • create_entity(ship) inside a story's own ```python fence;
    • a call-shaped line whose name is not in the vocabulary, such as > open_door(north);
    • an in-world bracket, [Hard limit of the reactor: three hours];
    • Scene: followed by prose;
    • Memory: she remembered the bells;
    • a fact restated inside a sentence;
    • a trailing bracket about a state block that lacks the hint's wording.

    §O.8's existing negative controls all still pass.

  3. Replay corpus. Every stored AI turn available from the v1 evidence runs is replayed through both extractors: §O.8's 443 and the identity run's 10, kept under $HOME, not committed. The package passes when all of these hold:

    • no turn changes except by removing a criterion-1 shape;
    • every changed turn appears in the diff report, and the WP report reviews it;
    • the review classifies zero changes as removed story.

    The report gives the counts.

  4. False-positive detection. The replay tool flags any removal that is not the reply's final segment or a whole line matching a criterion-1 shape, and any removal of more than a stated share of a turn's prose. An unreviewed flag fails the package.

  5. Genre-neutral instructions. A test asserts that EMIT_RULE, EMIT_REMINDER, the length hints and the vocabulary text contain no identifier from either acceptance fixture (Westhaven, Persephone) and none of J03's genre nouns.

  6. Real model.

    • Identity diagnostic. The office fixture at 4,096 on the CPU/HTTPS host gives 0 proposals naming an example identifier (baseline 1 of 10), protocol shapes in 0 of 10 stored turns (baseline 4 of 10), and 0 identity signals. Its --scripted --inject self-test still fires.
    • Long run. At least 50 turns at 16,384 with memory on stores protocol in 0 turns (baseline 0 of 104). Restatement counts are recorded against the 74 of 104 baseline, as a measurement, not a gate.

Regression requirements. The full backend suite, including the 25 §O.8 cases; C06 and H05 (refusals are still recorded and shown to the model); J01-J03; the I-series; the browser harness's hostile-narration checks (H04, H06, H07).

Dependencies. After A1. Both change the prompt, and one real-model run can then cover both. The brief can be written while A1 is under review.

Compatibility.

Area Effect
Databases None. No migration
Stored narration of v1.0.0 campaigns Not rewritten
Bundles None
Save Points, branches None
State None; the validator is unchanged
Memories and summaries Future ones summarise cleaner prose. Existing ones are untouched
Docker and local-only None

Test modes. Deterministic: yes. Replay corpus: yes. Real model: yes. Browser: regression only. Offline: regression only. Long run: yes, at least 50 turns, and this run may be the same one as A1's criterion 7 if A1 is already merged.

Risk. Medium. The failure to fear is silent removal of prose. Constants-only recognition, the corpus and the detector are the mitigations.


WP-B — Independent long-term memory retention

Objective. An important fact planted early is recoverable at depth 100 or more from the memory bank itself, with provenance to a memory whose source range covers the planting turn. This must hold when neither authoritative state, later narration, the summary nor the history window carries the fact, and lineage safety must be unchanged.

Rationale. Item 5. F02 and M04 pass on state-based recovery, which the owner accepted. Memory's own retention is unproven, and in a long campaign whose facts are not all state-shaped it is the only continuity there is.

Scope. Diagnosis first, then the smallest sufficient fix.

  • B.1, the diagnostic. A retention harness that reports, for a planted fact, each stage with ids and depths:

    • created: does a memory whose source_start..source_end covers the planting turn contain the fact?
    • retained: is it not forgotten?
    • ranked: where is it in similarity order for the recall query?
    • injected: is it in the memories section?

    The harness runs in two modes. The deterministic mode uses a scripted narrator, summariser and embedder. The real-model mode uses the reference narrator and embedder. It also adds a recovered_through_memory_independent verdict to tools/m11_long_run.py, keeping the existing verdicts and their meanings.

  • B.2, the fix, only at the stages B.1 shows failing, from this bounded menu:

    • how the summariser's excerpt is chosen, instead of keeping only the last 2,000 tokens;
    • the memory prompt's instruction to keep named facts and objects;
    • an eviction rule that does not throw out a never-retrieved early memory first;
    • one additional ranking term (lexical or entity overlap, CONTEXT-AND-MEMORY §20).

    Anything outside the menu needs the owner's agreement in the brief.

Non-scope.

  • Memory becoming authoritative, or outranking state (F07).
  • Any change to lineage attachment or filtering (tree.attach_memory, forget_node, E02) or summary lineage (E03).
  • Merging memory with imported knowledge, or new embedding models or dependencies.
  • Automatic re-summarisation of existing memories.
  • Redesigning cross-layer duplication (§22).

Likely affected. backend/app/memorybank.py (creation, eviction, ranking); backend/app/context/builder.py (memories section, only if the query changes); backend/app/models.py, only if a field is unavoidable. Tools: tools/memory_ab.py, tools/m11_long_run.py. Tests: test_context_memory.py, test_memory_nodes.py, test_m11_leakage.py, test_m11_long_run_memory.py.

Acceptance criteria.

  1. Deterministic retention test, committed. Fact F is planted at depth 3 or less through narration only, with no state correction and no knowledge source. These are asserted for the whole run:

    • F is never in authoritative state;
    • no turn after the planting block contains F's sentinel tokens;
    • the summary section never contains F;
    • the planting turn is outside the history window at recall.

    At a recall depth of 100 or more, the memories section holds a memory that contains F and whose source range includes the planting depth, and the context report lists its id and similarity. The test must fail on the 432f041 tree, and the report must name the stage at which it fails.

  2. Past capacity. The same test runs with more memories written than memory_bank_capacity. The planting-era memory is still active and retrieved, or its eviction follows a documented, tested rule that the WP report justifies. No pinned memory is evicted. The "frozen bank" regression (a new memory evicted at once) still passes.

  3. Lineage. Fact G is planted only on a line later abandoned by Undo and divergence. G's memory stays on disk and is absent from every active-line prompt and from memories.used. All 14 tests in test_m11_leakage.py and the E02 tests pass unchanged.

  4. Authority. A memory contradicting state loses, and memory is still framed as non-canon (F07).

  5. Real model. A run of at least 100 turns at 16,384 with memory on, on the GPU host with logging. The fact is chosen so that the harness can verify the criterion-1 preconditions on real output. Pass: one run meets every precondition and returns recovered_through_memory_independent, with the memory's provenance. A run whose precondition fails reports which one, and counts as neither pass nor failure. The report gives the stage-by-stage diagnostic for every run.

  6. Budget. The memories section stays inside its budget (F03), and the steady-state prompt is no larger than A1's arithmetic allows.

  7. CONTEXT-AND-MEMORY.md §15, §20 and §21 are updated as implemented.

Regression requirements. F01-F08, E01-E04, M04 verdicts (the new verdict is added and the old ones are unchanged), the I-series (memories travel in the bundle), L04 (test_memory_rewrite.py), §O.7's no-write-lock and failure-recording tests.

Dependencies. After A2, so that it measures the prompt v1.1 ships and a narrator no longer handed protocol to restate. B.1's deterministic harness may be built earlier.

Compatibility.

Area Effect
Databases No migration preferred. If a memory field is unavoidable, it is additive with a default, tested from a real v1.0.0 database, with schema parity
Bundles Stay v3. Any new memory field is optional on import, and its absence means "unknown". v1.0.0 bundles import
Existing memories Valid and used as they are. A rebuild stays opt-in (tools/rewrite_memories.py)
Branch history, Save Points, state, knowledge None
Settings Any change to capacity semantics keeps v1 defaults
Docker and local-only None

Test modes. Deterministic: yes. Real model: yes. Browser: no. Offline: regression only. Long run: yes, at least 100 turns, possibly several.

Risk. High uncertainty, because the outcome depends on the model, and medium implementation risk. The blast radius is the memory subsystem.


WP-C — Browser release coverage

Objective. Every reader-facing behaviour that v1 proved only through the API or the component suite is driven in a real browser, including an export that actually leaves the browser as a file.

Rationale. Item 6. §T carries two qualifications that exist only because the harness stops short.

Scope. New scenarios in tools/m11_browser.py, or a sibling that reuses tools/m11_webdriver.py: Retry; Save Point create and restore; state correction; narration length; failed generation; export download. Also the Firefox profile preferences that direct a download to a harness-owned directory under $HOME. The package is harness-only, unless it finds a product defect or a control with no accessible name. Either is fixed with a regression test and reported as a product change.

Non-scope. New UI, frontend refactors, Selenium or any new dependency, screenshot diffing, tablet layout, CI.

Likely affected. backend/tools/m11_browser.py, backend/tools/m11_webdriver.py, the DEVELOPMENT.md harness section.

Acceptance criteria. The run ends with 0 failed and 0 skipped. It runs against the built SPA served by FastAPI, with turns from the reference narrator over trusted-LAN HTTPS. Every assertion reads the rendered DOM or a file on disk.

  1. Retry. Retry on the newest turn yields a second take, and the indicator reads 2/2. Stepping to 1/2 shows the original narration unchanged. The takes persist after a reload.

  2. Save Point. Create a named Save Point through the UI, play two more turns, then restore it through the UI. The transcript ends at the named moment, the position indicator says later story is ahead, and Redo walks into the later turns. After a reload the Save Point is still listed.

  3. State correction. A correction submitted through the State panel applies and is still shown after a reload. A partly refused correction shows its refusal and reason to the reader.

  4. Narration length. After changing the control to brief and playing a turn, then to long and playing a turn, each turn's context inspector shows its band's word range.

  5. Failed generation. With the model set through Settings to a name the server does not serve, a submitted turn shows an error, adds no narration and keeps the typed input. Setting the model back, the next turn succeeds and the earlier story is intact.

  6. Export download. A real click on Export, from both the campaign library and campaign settings, writes a file with no manual step. The file:

    • exists and is not empty;
    • parses as ai-dnd-adventure-v3;
    • imports into a fresh data directory with the same action count, head position and Save Points.

    If the snap Firefox cannot be made to download, the check runs on a non-snap Firefox under $HOME, and DEVELOPMENT.md says so. Exercising only the click does not pass.

  7. The existing 38 checks pass in the same run.

Regression requirements. The existing browser checks; the frontend suite where a product fix is made.

Dependencies. None. It can run at any point. Its final run is repeated on the v1.1 release candidate.

Compatibility. None, unless a product fix is made, and then per §5.

Test modes. Browser: yes. Real model: yes, over the HTTPS reference host. Deterministic: no. Offline: no. Long run: no.

Risk. Low product risk. The medium risk is harness flakiness, so every wait is on a DOM condition, never a sleep used as an assertion.


WP-D — Recovery honesty

Objective. A backup is kept only after a full integrity check. A campaign whose export exceeds what import accepts is exported with a warning saying so, not silently.

Rationale. Items 7 and 10a. Both are recovery gaps that are known, measured and cheap to close. Neither needs a policy decision.

Scope.

  • backup.py runs PRAGMA integrity_check on the finished copy, and the time it takes is measured.
  • Export compares the serialised bundle's size with limits.MAX_IMPORT_BODY_BYTES. When it is over, the file is still delivered and the reader sees a warning naming the limit and what it means.
  • The import refusal for an oversized bundle names the limit.
  • DEVELOPMENT.md documents the ceiling as measured: about 13 kB per action on a real campaign, and M9's conservative figure of about 279 turns.

Non-scope. Scheduled backups; raising the import limit or streaming import; a bundle format change or further compression; a restore button.

Likely affected. backend/app/backup.py, backend/app/routers/adventures/bundle_io.py, frontend/src/pages/Campaigns.jsx, frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx, frontend/src/pages/backup.test.jsx, and the backup and bundle tests.

Acceptance criteria.

  1. A backup of a healthy database reports integrity_check ok and is kept.
  2. A copy with damage that integrity_check detects and quick_check does not, such as an index inconsistent with its table, is rejected and not kept. Existing backups are still never overwritten.
  3. The time for integrity_check is recorded on the 100-turn evidence database and on a synthetic database of 100 MB or more, and the backup completes through the UI on both.
  4. Exporting a fixture campaign over 20 MB succeeds, delivers the file, and shows a warning naming the import limit, asserted in the API response and in a component test. A campaign under the limit shows no warning.
  5. Importing that bundle is refused with a message naming the limit.
  6. A normal campaign's export is byte-identical before and after the package, apart from timestamp fields.

Regression requirements. I01-I07, L01-L04, the backup case in test_m11_migration.py, backup.test.jsx, and the offline container's export and import.

Dependencies. None.

Compatibility. Databases and bundles: no change. Backups: a stricter check, with the same file.

Test modes. Deterministic: yes. Browser: the warning, optionally in WP-C's harness. Offline: regression only. Real model: no. Long run: no.

Risk. Low.


WP-E — Control-boundary contrast

Objective. Control boundaries meet WCAG 1.4.11's 3:1 against their panel, at rest and on hover.

Rationale. Item 9.

Scope.

  • The boundary token values in frontend/src/styles/tokens.css, and any component that overrides them.
  • tools/contrast_audit.py treats boundary pairs below 3:1 as failures, not advisories.
  • The browser harness measures rendered boundary contrast.
  • Before-and-after screenshots for the owner's approval.

Non-scope. A palette redesign, typography, layout, tablet work, a screen-reader audit, and other WCAG criteria. A defect the same measurement finds is recorded, not taken on.

Likely affected. frontend/src/styles/tokens.css, backend/tools/contrast_audit.py, backend/tools/m11_browser.py, frontend/src/a11y.test.jsx.

Acceptance criteria.

  1. tools.contrast_audit exits non-zero on any boundary pair below 3:1, and exits 0 on the package tree.
  2. Every text pair still clears 4.5:1 (1.4.3). The rendered text contrasts measured by the harness do not fall below v1's (14.57, 5.48, 13.57 and 5.88:1) without a stated reason.
  3. The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1. The focus indicator is still visible.
  4. The owner approves the before-and-after screenshots, and the WP report records the approval.

Regression requirements. The frontend suite and lint; the browser harness's accessibility checks.

Dependencies. None. Its browser checks go into WP-C's harness if WP-C has landed, and into m11_browser.py otherwise.

Compatibility. None.

Test modes. Deterministic: yes. Browser: yes. Everything else: no.

Risk. Low.


9. Order and dependencies

A1  context safety reserve ──► A2  protocol echo + neutral instructions ──► B  memory retention
                                                                              (B.1 harness may start early)
C   browser coverage        ── independent
D   recovery honesty        ── independent
E   boundary contrast       ── independent (checks land in C's harness if C is first)
                                   │
                                   ▼
                        v1.1 release validation (§10)

Recommended order: A1, A2, B, C, D, E.

  • A1 first. It is the only item that can silently remove canon from a prompt. It is deterministic to test, small in blast radius, and it establishes the provenance that A2's and B's real-model runs will read.
  • A2 second. Its leak is silent and compounds through replay. It must precede B, because it changes what the memory pass summarises.
  • B third. It carries the highest continuity value and the most uncertainty. Measuring it before A1 and A2 settle would measure a prompt v1.1 does not ship.
  • C, D and E close verification, recovery and accessibility gaps with no known story risk. They depend on nothing, so the owner may move any of them earlier. While B waits on long runs is a natural slot. Each still has its own brief and review.
Property A1 A2 B C D E
Can begin independently yes after A1 after A2 yes yes yes
Schema or migration no no avoid; additive if unavoidable no no no
Export format no (additive evidence key) no no; optional field at most no no no
Changes acceptance tests (V1-ACCEPTANCE-TESTS.md) no no no no no no
Changes the stored prompt yes yes possibly no no no
Real-model validation yes yes yes yes, for turns no no
Browser testing regression regression no yes optional yes
Offline / no-network testing regression regression regression no regression no
Long-run testing yes, 50 turns or more yes, 50 turns or more yes, 100 turns or more no no no
Risk low medium high uncertainty low low low

10. Compatibility summary

Area A1 A2 B C D E
v1.0.0 campaign databases open unchanged unchanged unchanged; an additive migration only if unavoidable — unchanged —
Export and import bundles v3; new evidence key — v3; an optional field at most — v3; export warns —
Save Points — — — — — —
Branch history — — lineage rules unchanged — — —
Narrative state — validator unchanged memory never outranks state — — —
Memories and summaries — new ones from cleaner prose creation, retention and ranking change; existing rows kept — — —
Knowledge sources — — — — — —
Model settings none preferred — capacity defaults kept — — —
Docker and local-only — — — — — —

11. v1.1 release criteria

v1.1 is not called v1.1.0 until all of the following hold on one release candidate tree:

  1. Every package in scope is accepted, each with its report, and with its acceptance criteria passing on the candidate or on a tree whose product code the candidate carries unchanged.
  2. The v1 contract still passes. All 82 REQUIRED FOR V1 tests hold, with H09 not applicable on the same condition. None is relaxed.
  3. Suites and builds: the backend suite, the frontend suite and lint, the production build, and docker build --no-cache, with the image's SPA identical to the local build.
  4. Offline: tools/m11_offline.py passes all checks with no network and a fresh volume.
  5. Browser: the existing 38 checks plus WP-C's and WP-E's pass with 0 failed and 0 skipped, on the candidate, over trusted-LAN HTTPS.
  6. One v1.1 long run on the candidate's product code: 100 turns or more at a 16,384 window with memory on, on the GPU host with logging. It passes M01-M04. A1's headroom table shows the documented reserve on every re-counted prompt, and every turn is fits. A2's leak count is 0. B's independent-retention verdict is recorded.
  7. Identity diagnostic on the candidate with memory on: 0 signals, and 0 protocol shapes in stored narration.
  8. Recovery: tools/m11_recovery.py on the v1.1 long run's bundle passes all checks.
  9. Upgrade from a real v1.0.0 database. A database is created by the v1.0.0 tree and played. It has retained history, an undone head, Save Points, memories, summaries, imported knowledge and a narration-length choice, and uses loopback or placeholder settings with no real hostnames. Opened by the candidate, its transcript, head, Redo availability, state, Save Points, memories, knowledge and settings compare identical, and any migration is forward-only with schema parity. A v1.0.0 export imports into v1.1. Because the format stays v3, a v1.1 export of that campaign is checked for import into v1.0.0, and the result is reported.
  10. Documentation: README.md, DEVELOPMENT.md, the as-implemented sections named by each package, and VERSION.md.
  11. Owner events, each separate: the signed release commit, main, and a v1.1.0 tag.

12. Scope recommendation

Recommended: Option 1, a focused v1.1.

Ship in v1.1 Defer to v1.2 Future / optional
WP-A1 context safety reserve Scheduled backups (10b) Real media-provider adapters (11; K05, K06)
WP-A2 protocol echo and neutral instructions Raising or streaming the import limit (7) Whole-transcript copy, story search (§77, §78)
WP-B independent memory retention Reducing in-sentence restatement (2) Discarded-history recovery screen (§63)
WP-C browser release coverage Stale-scene and derived-contamination detectors in the identity diagnostic (8) Tablet layout
WP-D recovery honesty Cross-layer duplication suppression (17) Editable RPG world state
WP-E control-boundary contrast Media debt: ambience, visual-profile recovery and UI (20)
Trusted-LAN residual limits: address pinning against DNS rebinding (21)
Comparative narrator-model recommendation (13)

Why focused. Four of the six packages close silent failure modes or verification gaps that the v1 evidence itself named. The other two are small and deterministic. Everything deferred is either a new feature, or needs a policy decision the owner has not been asked for, or rests on an occurrence that has not happened. A broader v1.1 that took scheduled backups and media would add the two packages with the most new surface. It would also push the long-run re-validation, which every prompt change already requires, further from the changes it validates.

If B's real-model criterion cannot be met. Precondition-valid runs may recover nothing even after B.2. In that case the owner chooses between shipping v1.1 with B's deterministic criteria met and the real-model result recorded as a residual risk, or holding v1.1 for B. The plan does not pre-decide this.

13. Coding briefs needed next

These briefs are not written here, and none is to be executed from this document.

  1. WP-A1 — context-window safety reserve. Write this one first. It carries the highest user risk, has no dependencies, is deterministically testable, and produces the provenance that later real-model runs read. The brief must settle three things: the tokenizer divergence the reserve is sized for, whether calibration is built or detection alone, and whether any setting is added (recommended: none).
  2. WP-A2 — protocol echo hardening and genre-neutral instructions. It can be drafted while A1 is under review.
  3. WP-B — independent memory retention: B.1 diagnostic, then B.2 fix. Two briefs are an option if B.1's findings should be reviewed before a fix is chosen.
  4. WP-C — browser release coverage.
  5. WP-D — recovery honesty.
  6. WP-E — control-boundary contrast.
  7. v1.1 release validation, written only after every in-scope package is accepted.

14. Decisions for the owner

  • Confirm Option 1, the focused v1.1.
  • A1: the divergence the reserve must absorb; calibration or detection only; that a detected truncation is flagged, not turned into a failed turn.
  • B: whether B.1 and B.2 are one brief or two; the choice in §12 if the real-model criterion is not met.
  • D: that the 20 MB import limit stays in v1.1.
  • E: approval of the visual change.
  • Whether v1.1 real-model validation stays on the reference narrator, as this plan recommends.
  • The report naming and rotation in planning/reports/ (§5 rule 5).