Files
interactive-story/planning/V1.1-PLAN.md
T
JesseMarkowitzandClaude Opus 5 db7b309e3d
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
v1.1 closeout: accept integrated release validation
Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.

V1.1 RELEASE VALIDATION: PASS

What was run, on this candidate:

- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
  0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
  175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
  to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
  trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
  3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
  keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
  scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
  that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
  identical on all 15 census fields, schema parity at user_version 94, and both
  bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
  verified, public endpoint refused, a real turn, restart, persistence, and
  Firefox rendering the reopened campaign.

Carried residuals, stated rather than summarised away:

- WP-B: deterministic independent-memory recovery PASS; reference-model
  independent-memory recovery FAIL at memory creation — the owner-accepted
  limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
  reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
  release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
  backlog, reproduced and not fixed during validation.

Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.

Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.

Still the owner's to do: sign the release commit, update main, tag v1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-16 07:12:23 -04:00

941 lines
54 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Adventure Storyteller — v1.1 Plan
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
signed v1.0.0 release commit `432f041`.
**WP-A1 and WP-A2** are committed and signed as `d63804f`
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`). **WP-B.1**, the memory-retention
diagnostic, is committed and signed as `beb17ad`
(`reports/v1.1/V1.1-WP-B1-REPORT.md`). **WP-B.2**, the retention fix, is complete and
staged for the owner's signed commit. It is **accepted with a documented
real-model limitation** (owner decision, 2026-09-15):
- B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship;
- deterministic independent-memory recovery passes;
- the one isolation-valid reference-model run failed at memory creation;
- the B2.4 prompt experiment did not fix that and was reverted.
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
**WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release
coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`).
**WP-D** (recovery honesty) and **WP-E**
(control-boundary contrast) are committed and signed as `87a4032`, each with its
own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`.
**Integrated release validation has since run on candidate `87a4032` and
passed** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): the 82 REQUIRED v1 tests hold
(81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a
`--no-cache` image whose SPA is file-for-file identical to the local build,
offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window
run passing M01-M04 with every turn `fits` and an A2 leak count of 0, identity
0 signals / 0 protocol shapes, recovery 16/16, a real v1.0.0 upgrade comparing
identical on all 15 fields with both bundle directions importing, and a release
smoke of 15/15.
WP-B's reference-model memory limitation is carried as an accepted residual, as
are the mid-reply instruction echo and the doubled full stop; K1 is classified as
v1.2 backlog. **No `v1.1.0` tag exists, `main` is unchanged, and no release
commit has been made** — those three remain the owner's separate events.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
specification and acceptance contract are unchanged: every package below
improves how an existing requirement is met. No package adds a requirement.
---
## 1. Terminology
- **Work package (WP).** One independently reviewable unit of v1.1 work, with
its own coding brief, its own report, and its own review. Work packages are
lettered. **They are not milestones, and there is no M12.**
- **Brief.** The coding prompt for one work package. It holds only the
objective, the source documents, the scope and non-scope, the acceptance
criteria, and the stop condition. No work package begins before its brief
exists.
- **Reference narrator.** `qwen2.5:3b-instruct`, with and without `num_ctx`
16,384 baked in, and `nomic-embed-text` for embeddings. These are the models
the v1 evidence used (M11 report §E.1). v1.1 real-model evidence uses the
same models so that its numbers compare with v1's.
- **Reference hosts.** The CPU reference host serves Ollama over HTTPS with a
private CA, which is the A06 evidence path. The GPU inference host serves
plain HTTP on the LAN and is used for long runs (M11 report §E.1).
- **Planning package version** (`VERSION.md`, v4.x) and **product version**
(v1.0.0, v1.1.0) are separate numbers. `SECURITY-THREAT-MODEL.md`'s own
"Status: v1.1" line is that document's revision label from Phase 0B. It does
not refer to this release.
## 2. Baseline
| | |
| --- | --- |
| Release | **v1.0.0**, 2026-09-14 |
| Release commit | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, signed by the owner |
| Tag | `v1.0.0`, a signed annotated tag on that commit, pushed |
| `main` | `432f041`, the same commit |
| v1.1 branch | `v1.1-development`, created at `432f041` |
| v1 contract | 82 REQUIRED FOR V1 tests: 81 PASS and H09 NOT APPLICABLE (M11 report §T) |
| Schema | migrations up to 94 (`LATEST_VERSION` 94) |
| Bundle format | `ai-dnd-adventure-v3` |
| Provenance | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` |
## 3. v1.1 goals
1. **No silent loss of what the narrator is given.** A prompt must not reach the
server's window edge by an accident of tokenizer arithmetic. A turn the
server truncated must be reported.
2. **No protocol in the story.** The shapes of the application's own protocol
that a narrator copies must not be stored as narration. The fixed
instructions must not hand every campaign one genre's nouns to copy. Story
prose must not be removed in the process.
3. **Memory that remembers on its own.** An old fact must be recoverable from
the memory bank itself, with provenance, when nothing else carries it. This
must not weaken lineage safety.
4. **Release evidence without API-only gaps.** Every reader-facing history,
state, length, failure and export behaviour must be driven in a real browser.
5. **Recovery that says when it cannot recover.** Backups get a full integrity
check, and an export that import would refuse must say so.
6. **Control boundaries that meet WCAG 1.4.11.**
And throughout: **every v1.0.0 campaign opens unchanged in v1.1.**
## 4. Non-goals
- New product features: media generation, TTS, STT, story search, whole-transcript
copy, a discarded-history recovery screen, a restore button, tablet redesign.
- Changes to the history, branch, head or Save Point architecture (ADR 005,
ADR 012).
- Changes to the authoritative state model, its event vocabulary or its
validator (ADR 010, ADR 013). Changes to the prompt *wording* that describes
the vocabulary are in scope, in WP-A2.
- Changes to the knowledge authority classes or their retrieval.
- A new bundle format version.
- Any new runtime dependency, any network access beyond the configured
inference endpoint, and any first-use download, including per-model
tokenizers.
- Supporting inference servers other than Ollama, beyond what
`context_window_override` already allows.
- Re-writing stored narration or memories of existing campaigns automatically.
- Redesigning identity or state handling on the strength of post-M8 finding D,
which has not been reproduced.
## 5. Rules for every work package
The v1 milestone rules (`BUILD-MILESTONES.md` §2) apply to every work package,
and so do the following:
1. **Compatibility default.** Existing v1.0.0 campaign databases open
unchanged. Any migration is forward-only and additive, and is tested by
opening a real v1.0.0 database. Fresh-install and upgraded schemas must still
compare identical (`test_m11_migration.py`). v1.0.0 bundles still import.
2. **The bundle stays `ai-dnd-adventure-v3`** unless a brief makes the case under
M9's semantic test and the owner agrees. Adding a key inside stored evidence
is not a format change when an absent key unambiguously means "not
recorded".
3. **The v1 contract is the regression floor.** No REQUIRED FOR V1 test is
retired, relaxed or reclassified.
4. **Evidence discipline** (M11 report §O). Any product change made after a
black-box run was taken invalidates that run for release purposes. Harness
defects and product defects are reported separately. A check that cannot fail
is not evidence.
5. **Each package has a report** (`planning/reports/V1.1-WP-<id>-REPORT.md`).
It records what was built, the acceptance results with measurements, what
was not verified, and the as-implemented documentation changes. The report
rotation for `reports/` is decided when the first one is written
(`planning/README.md`).
6. **Real-model runs longer than a few minutes on the GPU host** require the
power, link and kernel logging in `DEVELOPMENT.md`, started before the run.
7. **No real hostnames, addresses or people's names** in committed files,
fixtures or reports. Evidence stays under `$HOME`, never `/tmp`.
8. **Commits and tags are the owner's.** A package ends with its changes staged
for a signed commit.
---
## 6. Backlog triage
Severity reflects user risk: **High** means story corruption or lost continuity
that the reader is not told about. **Medium** means silent degradation or a
real verification gap. **Low** means visible, rare or cosmetic.
### 6.1 The recorded post-v1 backlog
| # | Item | Source | Current evidence | Severity | Disposition |
| --- | --- | --- | --- | --- | --- |
| 1 | Context-window safety margin | M11 §P risk 3, §N; post-v1 backlog | The largest prompts left 23, 34 and 42 real tokens of headroom in the GPU runs. The application's `cl100k_base` count ran 16 tokens below the narrator's on every prompt measured. The only slack is `OUTPUT_SAFETY_MARGIN = 64` (`context/builder.py`), which is fixed and must also absorb the section separators. Ollama 0.34 cut an over-window synthetic prompt to 8,194 tokens with no error. Truncation drops the oldest tokens, which are the narrator's rules and the canon. The server's reported usage is already stored per turn (`snapshot["usage"]`, `routers/adventures/turns.py`) and read by nothing. | **High**: silent, invisible, and reachable by a model whose tokenizer diverges further | **v1.1, WP-A1** |
| 2 | Narrator restating prompt and state text | §P risk 4, §G.4 | The state's fact line was restated as a sentence on 74 of 104 turns in the evidence run. Phrases such as "Scene set, continue your adventure." and lines opening "Memory:" are model-invented, and none is an application string (verified by search). The restatements sit inside story sentences. | **Medium**: narration quality, and it is the path by which M04's fact reached memory, which masks item 5 | **v1.1, WP-A2** measures it. Reducing in-sentence restatement is **v1.2**, because removing it means judging prose |
| 3 | Protocol leakage: event-call syntax, stray headings, repeated length hints | §P risk 6, §S.6 | On 4 of 10 turns of the closeout identity run (4,096 window) and 0 of 104 in the 100-turn runs. Causes verified in code. `events.vocabulary_for_prompt` shows every event as `name(field, …)`, which is the notation copied. `_is_echoed_instruction` requires both "state block" and "events list", and the length hint's tail says only "state block". A lone trailing `Scene:` is not removed. Stored text is replayed as history (TECHNICAL-DESIGN §15.4). | **Medium-high**: silent and self-reinforcing, but rare at a full window | **v1.1, WP-A2** |
| 4 | Genre-specific example in the fixed state rule | §P risk 16; `narrative/extract.py` `EMIT_RULE` | The example names `mara`, `silver-key`, `old-abbey` and `aldric`. In the office-meeting identity run the narrator proposed giving `silver-key` to Alice, and the validator refused it. No state was corrupted. | **Medium**: every non-fantasy campaign gets fantasy nouns to copy | **v1.1, WP-A2** |
| 5 | Independent long-term-memory retention | §P risk 5, §G.4; F02 and M04 qualifications | No run showed a planting-era memory carrying the planted fact. Recovery ran from state to the narrator's restatement to memory. Code facts that bear on it, none yet proven to be the cause: a memory is at most 50 words for a 6-action block. The summariser sees the block's *last* 2,000 tokens (`truncate_to_last_tokens`), so an early fact in a long block can be cut. Eviction is least-recently-used above `memory_bank_capacity` (default 80), so a never-retrieved early memory is the first to go once a campaign passes about 480 actions; the 33-memory evidence run never reached that. Ranking is cosine similarity plus a pin (CONTEXT-AND-MEMORY §20). | **High** for long campaigns: lost continuity, seen only as the story forgetting | **v1.1, WP-B** |
| 6 | Browser automation gaps: Retry, Save Point, state correction, narration length, failed generation, export download | §S.3, §T qualifications, §P risk 11 | Proved through the API, the component suite and the 100-turn campaign, but never driven in a browser. The export control is exercised only as far as the click, because a `blob:` download does not leave the headless snap Firefox. No defect is known. | **Medium**: a verification gap on the release path | **v1.1, WP-C** |
| 7 | Practical export and bundle-size ceiling | §P risk 12, §K; M9 debt | The import body limit is 20 MB (`limits.MAX_IMPORT_BODY_BYTES`). M9's conservative ceiling is about 279 turns. The real 100-turn campaign was 2.74 MB, about 13 kB per action, or roughly 1,600 actions. A larger campaign exports and is then refused on import, and nothing says so at export. | **Medium** impact, low frequency | **v1.1, WP-D**: warn at export. Raising the limit or a streaming import is **v1.2** |
| 8 | Identity-confusion diagnostic follow-up | §P risk 8, §S.5 | Not reproduced. The diagnostic cannot see a stale scene, has never run with memory on, and has never run at a 16,384 window. It found no product deficiency. | **Low**: no occurrence since the original, whose campaign is gone | **No v1.1 package.** The diagnostic is re-run as a WP-A2 regression and in the v1.1 release gate, with memory on. A stale-scene detector is **v1.2**. The next real occurrence is classified with `tools/m11_identity.py` |
| 9 | WCAG 1.4.11 control-boundary contrast | §P risk 10, §L | Measured at 1.33:1 resting and 1.75:1 on hover (`--border` and `--border-bright` against `--bg-panel`). M11's argument that the label identifies the control is weakest for text inputs, where the boundary shows where to type. | **Low-medium** | **v1.1, WP-E** |
| 10a | Backup integrity: `integrity_check` | §P risk 14; M9 | `backup.py` runs `PRAGMA quick_check`, which skips index-content verification. | **Low** | **v1.1, WP-D** |
| 10b | Scheduled backups | §P risk 14; M9 | None exist; backups are manual. They need owner policy on interval, retention, location and disk use, and must not contend with a turn's write lock (§O.7). | **Medium** impact, but a new feature | **v1.2** |
| 11 | Real media-provider adapters | §P risk 13; M10; K05 and K06 (FUTURE) | The seam has no consumer. Adding one brings dependencies, a stricter endpoint rule (`DEVELOPMENT.md`) and a new UI. | n/a: a feature, not a risk to existing stories | **Future / optional**, not v1.1 |
### 6.2 Other recorded v1 residual risks and debt
| # | Item | Source | Disposition |
| --- | --- | --- | --- |
| 12 | The state lags the narration at 4,096 with the 3B narrator (1 proposal in 10 applied) | §P risk 17, §S.5 | Model behaviour, not a defect. WP-A2's real-model runs record the proposal outcome counts. No package |
| 13 | Every realistic observation is a 3B model's | §P risk 2 | No package. v1.1 keeps the reference narrator for comparability. A comparative model recommendation stays open (`planning/README.md`, *Still open*) |
| 14 | A campaign that never chose a narration length keeps the pre-M11 hint | §P risk 15 | Deliberate. No change. WP-C drives the control |
| 15 | The GPU host dropped off the PCIe bus after the evidence run | §P risk 7, §E.1 | Operational. Rule 6 in §5 applies to every long run |
| 16 | Release timings are two hosts' | §P risk 1 | No action. No performance requirement exists or is invented |
| 17 | Cross-layer duplication: one fact in state, memory, history and a passage at once | CONTEXT-AND-MEMORY §22; open since M6 and M7 | **v1.2.** It costs tokens (item 1) and is related to item 2. WP-B may touch ranking only where its diagnostic requires |
| 18 | Memory-ranking factors beyond similarity and pin not implemented | CONTEXT-AND-MEMORY §20 | Inside WP-B's bounded fix menu, only if WP-B's diagnostic places the failure at ranking |
| 19 | M8 carried debt: no whole-transcript copy or story search (§77, §78), no discarded-history recovery screen (§63), tablet untuned, RPG world state read-only | `BUILD-MILESTONES.md` M8 and M9 | **Future backlog.** Features, not reliability |
| 20 | M10 debt: `ambience` is an empty shape; a deleted visual profile is unrecoverable in the app; profiles are API-only | M11 §D | **Future**, with media |
| 21 | A hostile host on the trusted LAN; the DNS-rebinding interval | `SECURITY-THREAT-MODEL.md` §10A | Accepted for v1. **Future.** No v1.1 package |
| 22 | A stale `chunk_id` in a restored snapshot | M9 | No action. It is a React key, not a live pointer |
| 23 | Developer-venv residue (`quickjs`, `psycopg`) | §M | No action. It does not ship |
| 24 | A Settings warning comparing the real window to the budget | M8 debt, "unowned" | **Closed by M11**: Settings reports the window and warns |
| 25 | Root `README.md` has stale facts: "1,191 backend tests" (1,421 at closeout), and the Screenshots paragraph calls screenshots "a job for the UI pass in M8" | found in this review | Documentation debt, not release status, so it is not changed in v4.1. It is fixed in WP-A1's documentation update |
---
## 7. Grouping decisions
**The suggested "narrator boundary hardening" package is split into A1 and A2.**
They share a theme but not a mechanism or a test method.
- **A1**, the context reserve, is budget arithmetic in `context/builder.py` and
`contextwindow.py`. It is proved deterministically with a scripted provider
that reports token counts.
- **A2**, the protocol echo, is the contract between the prompt's wording
(`extract.EMIT_RULE`, `events.vocabulary_for_prompt`, the length hint) and the
extractor. It is proved with a replay corpus of real narration and
real-model runs.
Bundling them would make one review carry two unrelated risk profiles.
A2's three items belong together. The example slugs and the call notation are
the *source* of the copied text, and the extractor is its *sink*. Changing one
without the other means proving the change twice. Both also alter the stored
prompt, so they invalidate the same evidence.
**B stands alone.** It is the only package whose success depends on model output
quality. Its test design, a fact that nothing but memory carries, is the hard
part. It follows A2, so that it measures the prompt v1.1 will ship.
**D is narrowed to recovery honesty**, meaning `integrity_check` and the export
warning. Both are small and deterministic, and both are recovery-path changes
with no policy question. Scheduled backups need owner decisions and have real
design risk, including write-lock contention. Raising or streaming the import
limit changes a DoS guard. Both move to v1.2.
**E stays focused** on control boundaries. The same measurement found no other
accessibility defect (M11 §L).
**F is not a package.** Nothing concrete in the product is deficient. The
diagnostic's known blind spots are a stale scene, memory off, and one window
size. The first two can be addressed by running it differently: the v1.1 gate
runs it with memory on. The stale-scene detector is new diagnostic work with no
occurrence to justify it yet, so it is v1.2.
**G is not in v1.1.** It adds no reliability or quality to existing stories, it
brings dependencies and a new endpoint surface, and its tests (K05, K06) are
FUTURE. `MEDIA-EXTENSION-CONTRACT.md` and the M10 seam remain authoritative for
whenever it is taken up.
---
## 8. Work packages
### WP-A1 — Context-window safety reserve
**Objective.** An assembled prompt plus its reply reserve stays at least a
documented safety reserve below the effective window, whatever tokenizer the
narrator uses. Where the server reports its own prompt count, a turn that
exceeded the window, or was evidently truncated, is recorded and shown, never
silent.
**Rationale.** §15.2 closed "budget larger than the window". It did not close
"our count is not the server's count". Measured headroom was 23-42 tokens
(item 1), and the fixed 64-token margin absorbs separators as well as tokenizer
drift. A narrator whose tokenizer runs a few percent heavier than
`cl100k_base` overflows a full 16,384 window. The failure deletes the canon at
the front of the prompt with every request returning 200. The data to detect it
is already stored with each turn and unread.
**Scope.**
- A safety reserve that scales with the effective budget and has a floor. It
replaces or supplements `OUTPUT_SAFETY_MARGIN`, is defined once, and applies
to verified windows, declared overrides and an unverified configured budget
alike. The brief states the tokenizer divergence it is sized to absorb.
- Reading the server-reported prompt-token count from the turn's usage. It is
recorded beside the application's count in the turn's window provenance.
Recorded states:
- `fits`;
- `exceeded`: reported count plus reply reserve is over the window;
- `truncation_suspected`: the reported count is materially below the
application's count;
- `unknown`: no usage was reported.
`exceeded` and `truncation_suspected` are surfaced in the context inspector,
and in the turn's response so the reader sees a notice.
- **Optional, decided in the brief:** calibrating the reserve from observed
server-to-application ratios per endpoint and model. It must be bounded and
quantised so that it does not re-price the prompt prefix every turn (§15.3).
Detection is required either way.
- `ContextOverflow` remains the explicit failure when protected context plus
both reserves exceeds the budget.
- `TECHNICAL-DESIGN.md` §15.2 and `DEVELOPMENT.md`'s context-window section, as
implemented. The root `README.md` stale facts from item 25.
**Non-scope.** The history block trim mechanism, section order, the knowledge
budget share, a per-model tokenizer or any tokenizer download, a hard-coded
window, raising any window, a provider abstraction, a Settings redesign, and
failing or discarding a turn after the fact.
**Likely affected.** `backend/app/context/builder.py`,
`backend/app/contextwindow.py`, `backend/app/routers/adventures/turns.py` and
`insights.py`, `backend/app/providers/openai_compatible.py` (usage, read-only),
the context inspector panel in `frontend/src/pages/Play/`. Tests:
`test_m11_context_window.py`, `test_m11_declared_window.py`,
`test_history_block_trim.py`.
**Acceptance criteria.**
1. **Per configuration.** The application's count plus the output reserve plus
the safety reserve is no more than the effective budget, and the safety
reserve is at least its documented value. This holds for each of: a
verified 4,096 window, a verified 16,384 window, a declared override with no
probe, and an unverified configured budget. There is a test per
configuration.
2. **Divergence.** A scripted provider reports prompt counts at 1.00×, and at
the brief's stated divergence, of the application's count, across a campaign
that fills the window. At both, no turn's reported count plus reply reserve
exceeds the window. At 1.25×, either no turn exceeds (calibration built) or
every exceeding turn is recorded `exceeded`. **An unflagged overflow fails.**
3. **Truncation.** A scripted provider reports 8,194 tokens against an
application count near 15,700. The turn is recorded `truncation_suspected`
in stored provenance, shown in the inspector, and reported in the turn's
response. A provider that reports no usage records `unknown`, never `fits`.
4. **Explicit overflow.** Protected context that fits without the safety
reserve but not with it raises `ContextOverflow` with an actionable message.
The story and state are unchanged (A05, L01).
5. **Canon survives.** `test_the_canon_at_the_front_survives_a_window_far_too_small`
and `test_without_the_cap_the_same_prompt_would_have_overflowed` both pass
with the reserve in place.
6. **Cache stability.** At steady state the history floor still moves in blocks
(`test_history_block_trim.py`). If calibration is built, a test shows the
budget changes only when the observed ratio crosses a stated step.
7. **Real model.** A long run of at least 50 turns runs at a 16,384 window with
memory on, using the reference narrator on the GPU host. Its ten largest
stored prompts are re-counted by the server, as §N did. Every one leaves at
least the documented reserve, and every turn's provenance reads `fits`. The
report carries the headroom table beside v1's.
**Regression requirements.** The full backend and frontend suites; F03, F04 and
M03 coverage; the I-series (a new provenance key must import and export); the
offline container; the browser harness's F05 inspector checks.
**Dependencies.** None. First.
**Compatibility.**
| Area | Effect |
| --- | --- |
| Databases | None expected. Provenance lives in the turn snapshot's JSON, so no migration |
| Bundles | None. Snapshots carry new keys, and absence means "not recorded" |
| Save Points, branches, state, memories, knowledge | None |
| Model settings | None preferred. If a setting is added, it takes a default that reproduces the documented reserve, plus an additive migration |
| Behaviour | At the largest prompts the history window holds slightly fewer actions. Nothing stored changes |
| Docker and local-only | None |
**Test modes.** Deterministic: yes. Real model: yes. Browser: inspector display,
via the harness. Offline: regression only. Long run: yes, at least 50 turns.
**Risk.** Low implementation risk and high value. Its blast radius is the
budget arithmetic of every turn. It changes prompts at full windows, so it
invalidates the v1 long-run evidence for v1.1 (§10).
---
### WP-A2 — Protocol echo hardening and genre-neutral instructions
**Objective.** Stored narration no longer keeps the application-owned protocol
shapes a narrator copies, and the fixed instructions stop supplying
genre-specific identifiers. Recognition uses only strings and syntax the
application itself owns, so story prose is not removed.
**Rationale.** Items 3 and 4. The call notation in the prompt is copied, the
length hint's echo escapes `_is_echoed_instruction`, and a lone trailing
`Scene:` survives. Each leak is replayed as history, so it teaches the next turn.
The example slugs were copied into an office scene.
**Scope.**
- A genre-neutral `EMIT_RULE` example and slug examples. No identifier in the
fixed instructions names a fixture entity or a genre noun.
- A decision, by measurement, on whether the vocabulary is shown in a
notation that does not look callable. The extractor handles call lines
regardless.
- The length hint's wording becomes a named constant shared by the builder and
the extractor, in the way `render.SECTION_HEADINGS` is shared. The extractor
then recognises:
- a line consisting only of a call to an event named in `events.SPECS`
(case-insensitive), optionally `>`-quoted;
- the length hint echoed as a trailing bracket, closed or cut off;
- a renderer heading that is the reply's final line with nothing after it.
- A replay tool that runs the old and new extractors over a corpus of stored AI
turns and writes a per-turn diff report.
- Measurement only: per-turn counts of state fact-line restatements,
model-invented headings, and proposal outcomes (applied, empty, unparseable,
refused), added to `tools/m11_long_run.py`.
**Non-scope.** Removing a fact restated inside a sentence; removing
model-invented headings that are not the renderer's; the event vocabulary itself,
the validator, the fence protocol, or history replay; any model-based cleanup
pass; rewriting stored narration of existing campaigns.
**Likely affected.** `backend/app/narrative/extract.py`, `events.py` (prompt
rendering only), `render.py` (constants), `backend/app/context/builder.py`
(length-hint constant), `tests/test_narrative_state.py`, `tests/test_m11_scifi.py`
(the J03 vocabulary check), `tools/m11_long_run.py`. Also
`TECHNICAL-DESIGN.md` §15.4 and ADR 013's implementation note.
**Acceptance criteria.**
1. **Known shapes are removed.** There is a committed case per shape, cut down
from the four §S.6 turns and anonymised:
- `> set_possession(silver-key, "alice")`;
- `> Create_entity(...)`;
- the length hint echoed as `[Hard limit: …append the state block well
inside the limit.]`, closed and cut off;
- a lone trailing `Scene:`.
The four stored §S.6 texts, replayed, lose exactly those lines. Every other
sentence is intact, asserted as set equality of the remaining sentences.
2. **Adversarial prose is untouched.** Each of these is a test:
- dialogue that mentions `set_possession` mid-sentence;
- `create_entity(ship)` inside a story's own ```python fence;
- a call-shaped line whose name is not in the vocabulary, such as
`> open_door(north)`;
- an in-world bracket, `[Hard limit of the reactor: three hours]`;
- `Scene:` followed by prose;
- `Memory: she remembered the bells`;
- a fact restated inside a sentence;
- a trailing bracket about a state block that lacks the hint's wording.
§O.8's existing negative controls all still pass.
3. **Replay corpus.** Every stored AI turn available from the v1 evidence runs
is replayed through both extractors: §O.8's 443 and the identity run's 10,
kept under `$HOME`, not committed. The package passes when all of these
hold:
- no turn changes except by removing a criterion-1 shape;
- every changed turn appears in the diff report, and the WP report reviews
it;
- the review classifies zero changes as removed story.
The report gives the counts.
4. **False-positive detection.** The replay tool flags any removal that is not
the reply's final segment or a whole line matching a criterion-1 shape, and
any removal of more than a stated share of a turn's prose. An unreviewed
flag fails the package.
5. **Genre-neutral instructions.** A test asserts that `EMIT_RULE`,
`EMIT_REMINDER`, the length hints and the vocabulary text contain no
identifier from either acceptance fixture (Westhaven, Persephone) and none of
J03's genre nouns.
6. **Real model.**
- **Identity diagnostic.** The office fixture at 4,096 on the CPU/HTTPS host
gives 0 proposals naming an example identifier (baseline 1 of 10),
protocol shapes in 0 of 10 stored turns (baseline 4 of 10), and 0 identity
signals. Its `--scripted --inject` self-test still fires.
- **Long run.** At least 50 turns at 16,384 with memory on stores protocol in
0 turns (baseline 0 of 104). Restatement counts are recorded against the
74 of 104 baseline, as a measurement, not a gate.
**Regression requirements.** The full backend suite, including the 25 §O.8
cases; C06 and H05 (refusals are still recorded and shown to the model); J01-J03;
the I-series; the browser harness's hostile-narration checks (H04, H06, H07).
**Dependencies.** After A1. Both change the prompt, and one real-model run can
then cover both. The brief can be written while A1 is under review.
**Compatibility.**
| Area | Effect |
| --- | --- |
| Databases | None. No migration |
| Stored narration of v1.0.0 campaigns | Not rewritten |
| Bundles | None |
| Save Points, branches | None |
| State | None; the validator is unchanged |
| Memories and summaries | Future ones summarise cleaner prose. Existing ones are untouched |
| Docker and local-only | None |
**Test modes.** Deterministic: yes. Replay corpus: yes. Real model: yes. Browser:
regression only. Offline: regression only. Long run: yes, at least 50 turns, and
this run may be the same one as A1's criterion 7 if A1 is already merged.
**Risk.** Medium. The failure to fear is silent removal of prose. Constants-only
recognition, the corpus and the detector are the mitigations.
---
### WP-B — Independent long-term memory retention
**Objective.** An important fact planted early is recoverable at depth 100 or
more from the memory bank itself, with provenance to a memory whose source range
covers the planting turn. This must hold when neither authoritative state, later
narration, the summary nor the history window carries the fact, and lineage
safety must be unchanged.
**Rationale.** Item 5. F02 and M04 pass on state-based recovery, which the owner
accepted. Memory's own retention is unproven, and in a long campaign whose facts
are not all state-shaped it is the only continuity there is.
**Scope.** Diagnosis first, then the smallest sufficient fix.
- **B.1, the diagnostic.** A retention harness that reports, for a planted fact,
each stage with ids and depths:
- **created:** does a memory whose `source_start`..`source_end` covers the
planting turn contain the fact?
- **retained:** is it not `forgotten`?
- **ranked:** where is it in similarity order for the recall query?
- **injected:** is it in the memories section?
The harness runs in two modes. The deterministic mode uses a scripted
narrator, summariser and embedder. The real-model mode uses the reference
narrator and embedder. It also adds a `recovered_through_memory_independent`
verdict to `tools/m11_long_run.py`, keeping the existing verdicts and their
meanings.
- **B.2, the fix,** only at the stages B.1 shows failing, from this bounded menu:
- how the summariser's excerpt is chosen, instead of keeping only the last
2,000 tokens;
- the memory prompt's instruction to keep named facts and objects;
- an eviction rule that does not throw out a never-retrieved early memory
first;
- one additional ranking term (lexical or entity overlap, CONTEXT-AND-MEMORY
§20).
Anything outside the menu needs the owner's agreement in the brief.
**Non-scope.**
- Memory becoming authoritative, or outranking state (F07).
- Any change to lineage attachment or filtering (`tree.attach_memory`,
`forget_node`, E02) or summary lineage (E03).
- Merging memory with imported knowledge, or new embedding models or
dependencies.
- Automatic re-summarisation of existing memories.
- Redesigning cross-layer duplication (§22).
**Likely affected.** `backend/app/memorybank.py` (creation, eviction, ranking);
`backend/app/context/builder.py` (memories section, only if the query changes);
`backend/app/models.py`, only if a field is unavoidable. Tools: `tools/memory_ab.py`,
`tools/m11_long_run.py`. Tests: `test_context_memory.py`, `test_memory_nodes.py`,
`test_m11_leakage.py`, `test_m11_long_run_memory.py`.
**Acceptance criteria.**
1. **Deterministic retention test**, committed. Fact F is planted at depth 3 or
less through narration only, with no state correction and no knowledge source.
These are asserted for the whole run:
- F is never in authoritative state;
- no turn after the planting block contains F's sentinel tokens;
- the summary section never contains F;
- the planting turn is outside the history window at recall.
At a recall depth of 100 or more, the memories section holds a memory that
contains F and whose source range includes the planting depth, and the
context report lists its id and similarity. **The test must fail on the
`432f041` tree**, and the report must name the stage at which it fails.
2. **Past capacity.** The same test runs with more memories written than
`memory_bank_capacity`. The planting-era memory is still active and
retrieved, or its eviction follows a documented, tested rule that the WP
report justifies. No pinned memory is evicted. The "frozen bank" regression
(a new memory evicted at once) still passes.
3. **Lineage.** Fact G is planted only on a line later abandoned by Undo and
divergence. G's memory stays on disk and is absent from every active-line
prompt and from `memories.used`. All 14 tests in `test_m11_leakage.py` and the
E02 tests pass unchanged.
4. **Authority.** A memory contradicting state loses, and memory is still framed
as non-canon (F07).
5. **Real model.** A run of at least 100 turns at 16,384 with memory on, on the
GPU host with logging. The fact is chosen so that the harness can verify the
criterion-1 preconditions on real output. **Pass:** one run meets every
precondition and returns `recovered_through_memory_independent`, with the
memory's provenance. A run whose precondition fails reports which one, and
counts as neither pass nor failure. The report gives the stage-by-stage
diagnostic for every run.
6. **Budget.** The memories section stays inside its budget (F03), and the
steady-state prompt is no larger than A1's arithmetic allows.
7. `CONTEXT-AND-MEMORY.md` §15, §20 and §21 are updated as implemented.
**Regression requirements.** F01-F08, E01-E04, M04 verdicts (the new verdict is
added and the old ones are unchanged), the I-series (memories travel in the
bundle), L04 (`test_memory_rewrite.py`), §O.7's no-write-lock and
failure-recording tests.
**Dependencies.** After A2, so that it measures the prompt v1.1 ships and a
narrator no longer handed protocol to restate. B.1's deterministic harness may
be built earlier.
**Compatibility.**
| Area | Effect |
| --- | --- |
| Databases | No migration preferred. If a memory field is unavoidable, it is additive with a default, tested from a real v1.0.0 database, with schema parity |
| Bundles | Stay v3. Any new memory field is optional on import, and its absence means "unknown". v1.0.0 bundles import |
| Existing memories | Valid and used as they are. A rebuild stays opt-in (`tools/rewrite_memories.py`) |
| Branch history, Save Points, state, knowledge | None |
| Settings | Any change to capacity semantics keeps v1 defaults |
| Docker and local-only | None |
**Test modes.** Deterministic: yes. Real model: yes. Browser: no. Offline:
regression only. Long run: yes, at least 100 turns, possibly several.
**Risk.** High uncertainty, because the outcome depends on the model, and medium
implementation risk. The blast radius is the memory subsystem.
---
### WP-C — Browser release coverage
**Objective.** Every reader-facing behaviour that v1 proved only through the API
or the component suite is driven in a real browser, including an export that
actually leaves the browser as a file.
**Rationale.** Item 6. §T carries two qualifications that exist only because the
harness stops short.
**Scope.** New scenarios in `tools/m11_browser.py`, or a sibling that reuses
`tools/m11_webdriver.py`: Retry; Save Point create and restore; state correction;
narration length; failed generation; export download. Also the Firefox profile
preferences that direct a download to a harness-owned directory under `$HOME`.
The package is harness-only, unless it finds a product defect or a control with
no accessible name. Either is fixed with a regression test and reported as a
product change.
**Non-scope.** New UI, frontend refactors, Selenium or any new dependency,
screenshot diffing, tablet layout, CI.
**Likely affected.** `backend/tools/m11_browser.py`, `backend/tools/m11_webdriver.py`,
the `DEVELOPMENT.md` harness section.
**Acceptance criteria.** The run ends with 0 failed and 0 skipped. It runs against
the built SPA served by FastAPI, with turns from the reference narrator over
trusted-LAN HTTPS. Every assertion reads the rendered DOM or a file on disk.
1. **Retry.** Retry on the newest turn yields a second take, and the indicator
reads 2/2. Stepping to 1/2 shows the original narration unchanged. The takes
persist after a reload.
2. **Save Point.** Create a named Save Point through the UI, play two more
turns, then restore it through the UI. The transcript ends at the named
moment, the position indicator says later story is ahead, and Redo walks
into the later turns. After a reload the Save Point is still listed.
3. **State correction.** A correction submitted through the State panel applies
and is still shown after a reload. A partly refused correction shows its
refusal and reason to the reader.
4. **Narration length.** After changing the control to brief and playing a turn,
then to long and playing a turn, each turn's context inspector shows its
band's word range.
5. **Failed generation.** With the model set through Settings to a name the
server does not serve, a submitted turn shows an error, adds no narration and
keeps the typed input. Setting the model back, the next turn succeeds and the
earlier story is intact.
6. **Export download.** A real click on Export, from both the campaign library
and campaign settings, writes a file with no manual step. The file:
- exists and is not empty;
- parses as `ai-dnd-adventure-v3`;
- imports into a fresh data directory with the same action count, head
position and Save Points.
If the snap Firefox cannot be made to download, the check runs on a non-snap
Firefox under `$HOME`, and `DEVELOPMENT.md` says so. **Exercising only the
click does not pass.**
7. The existing 38 checks pass in the same run.
**Regression requirements.** The existing browser checks; the frontend suite
where a product fix is made.
**Dependencies.** None. It can run at any point. Its final run is repeated on the
v1.1 release candidate.
**Compatibility.** None, unless a product fix is made, and then per §5.
**Test modes.** Browser: yes. Real model: yes, over the HTTPS reference host.
Deterministic: no. Offline: no. Long run: no.
**Risk.** Low product risk. The medium risk is harness flakiness, so every wait is
on a DOM condition, never a sleep used as an assertion.
---
### WP-D — Recovery honesty
**Objective.** A backup is kept only after a full integrity check. A campaign
whose export exceeds what import accepts is exported with a warning saying so,
not silently.
**Rationale.** Items 7 and 10a. Both are recovery gaps that are known, measured
and cheap to close. Neither needs a policy decision.
**Scope.**
- `backup.py` runs `PRAGMA integrity_check` on the finished copy, and the time
it takes is measured.
- Export compares the serialised bundle's size with
`limits.MAX_IMPORT_BODY_BYTES`. When it is over, the file is still delivered
and the reader sees a warning naming the limit and what it means.
- The import refusal for an oversized bundle names the limit.
- `DEVELOPMENT.md` documents the ceiling as measured: about 13 kB per action on
a real campaign, and M9's conservative figure of about 279 turns.
**Non-scope.** Scheduled backups; raising the import limit or streaming import; a
bundle format change or further compression; a restore button.
**Likely affected.** `backend/app/backup.py`, `backend/app/routers/adventures/bundle_io.py`,
`frontend/src/pages/Campaigns.jsx`,
`frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx`,
`frontend/src/pages/backup.test.jsx`, and the backup and bundle tests.
**Acceptance criteria.**
1. A backup of a healthy database reports `integrity_check` ok and is kept.
2. A copy with damage that `integrity_check` detects and `quick_check` does not,
such as an index inconsistent with its table, is rejected and not kept.
Existing backups are still never overwritten.
3. The time for `integrity_check` is recorded on the 100-turn evidence database
and on a synthetic database of 100 MB or more, and the backup completes
through the UI on both.
4. Exporting a fixture campaign over 20 MB succeeds, delivers the file, and
shows a warning naming the import limit, asserted in the API response and in
a component test. A campaign under the limit shows no warning.
5. Importing that bundle is refused with a message naming the limit.
6. A normal campaign's export is byte-identical before and after the package,
apart from timestamp fields.
**Regression requirements.** I01-I07, L01-L04, the backup case in
`test_m11_migration.py`, `backup.test.jsx`, and the offline container's
export and import.
**Dependencies.** None.
**Compatibility.** Databases and bundles: no change. Backups: a stricter check,
with the same file.
**Test modes.** Deterministic: yes. Browser: the warning, optionally in WP-C's
harness. Offline: regression only. Real model: no. Long run: no.
**Risk.** Low.
---
### WP-E — Control-boundary contrast
**Objective.** Control boundaries meet WCAG 1.4.11's 3:1 against their panel, at
rest and on hover.
**Rationale.** Item 9.
**Scope.**
- The boundary token values in `frontend/src/styles/tokens.css`, and any
component that overrides them.
- `tools/contrast_audit.py` treats boundary pairs below 3:1 as failures, not
advisories.
- The browser harness measures rendered boundary contrast.
- Before-and-after screenshots for the owner's approval.
**Non-scope.** A palette redesign, typography, layout, tablet work, a
screen-reader audit, and other WCAG criteria. A defect the same measurement finds
is recorded, not taken on.
**Likely affected.** `frontend/src/styles/tokens.css`, `backend/tools/contrast_audit.py`,
`backend/tools/m11_browser.py`, `frontend/src/a11y.test.jsx`.
**Acceptance criteria.**
1. `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and
exits 0 on the package tree.
2. Every text pair still clears 4.5:1 (1.4.3). The rendered text contrasts
measured by the harness do not fall below v1's (14.57, 5.48, 13.57 and
5.88:1) without a stated reason.
3. The rendered boundary of the story input and of a primary control, at rest
and on hover, is at least 3:1. The focus indicator is still visible.
4. The owner approves the before-and-after screenshots, and the WP report records
the approval.
**Regression requirements.** The frontend suite and lint; the browser harness's
accessibility checks.
**Dependencies.** None. Its browser checks go into WP-C's harness if WP-C has
landed, and into `m11_browser.py` otherwise.
**Compatibility.** None.
**Test modes.** Deterministic: yes. Browser: yes. Everything else: no.
**Risk.** Low.
---
## 9. Order and dependencies
```text
A1 context safety reserve ──► A2 protocol echo + neutral instructions ──► B memory retention
(B.1 harness may start early)
C browser coverage ── independent
D recovery honesty ── independent
E boundary contrast ── independent (checks land in C's harness if C is first)
│
▼
v1.1 release validation (§10)
```
**Recommended order: A1, A2, B, C, D, E.**
- **A1 first.** It is the only item that can silently remove canon from a
prompt. It is deterministic to test, small in blast radius, and it
establishes the provenance that A2's and B's real-model runs will read.
- **A2 second.** Its leak is silent and compounds through replay. It must
precede B, because it changes what the memory pass summarises.
- **B third.** It carries the highest continuity value and the most uncertainty.
Measuring it before A1 and A2 settle would measure a prompt v1.1 does not ship.
- **C, D and E** close verification, recovery and accessibility gaps with no
known story risk. They depend on nothing, so the owner may move any of them
earlier. While B waits on long runs is a natural slot. Each still has its own
brief and review.
| Property | A1 | A2 | B | C | D | E |
| --- | --- | --- | --- | --- | --- | --- |
| Can begin independently | yes | after A1 | after A2 | yes | yes | yes |
| Schema or migration | no | no | avoid; additive if unavoidable | no | no | no |
| Export format | no (additive evidence key) | no | no; optional field at most | no | no | no |
| Changes acceptance tests (`V1-ACCEPTANCE-TESTS.md`) | no | no | no | no | no | no |
| Changes the stored prompt | yes | yes | possibly | no | no | no |
| Real-model validation | yes | yes | yes | yes, for turns | no | no |
| Browser testing | regression | regression | no | yes | optional | yes |
| Offline / no-network testing | regression | regression | regression | no | regression | no |
| Long-run testing | yes, 50 turns or more | yes, 50 turns or more | yes, 100 turns or more | no | no | no |
| Risk | low | medium | high uncertainty | low | low | low |
## 10. Compatibility summary
| Area | A1 | A2 | B | C | D | E |
| --- | --- | --- | --- | --- | --- | --- |
| v1.0.0 campaign databases | open unchanged | unchanged | unchanged; an additive migration only if unavoidable | — | unchanged | — |
| Export and import bundles | v3; new evidence key | — | v3; an optional field at most | — | v3; export warns | — |
| Save Points | — | — | — | — | — | — |
| Branch history | — | — | lineage rules unchanged | — | — | — |
| Narrative state | — | validator unchanged | memory never outranks state | — | — | — |
| Memories and summaries | — | new ones from cleaner prose | creation, retention and ranking change; existing rows kept | — | — | — |
| Knowledge sources | — | — | — | — | — | — |
| Model settings | none preferred | — | capacity defaults kept | — | — | — |
| Docker and local-only | — | — | — | — | — | — |
## 11. v1.1 release criteria
v1.1 is not called v1.1.0 until all of the following hold on one release
candidate tree:
1. **Every package in scope is accepted**, each with its report, and with its
acceptance criteria passing on the candidate or on a tree whose product code
the candidate carries unchanged.
2. **The v1 contract still passes.** All 82 REQUIRED FOR V1 tests hold, with H09
not applicable on the same condition. None is relaxed.
3. **Suites and builds:** the backend suite, the frontend suite and lint, the
production build, and `docker build --no-cache`, with the image's SPA
identical to the local build.
4. **Offline:** `tools/m11_offline.py` passes all checks with no network and a
fresh volume.
5. **Browser:** the existing 38 checks plus WP-C's and WP-E's pass with 0 failed
and 0 skipped, on the candidate, over trusted-LAN HTTPS.
6. **One v1.1 long run** on the candidate's product code: 100 turns or more at a
16,384 window with memory on, on the GPU host with logging. It passes M01-M04.
A1's headroom table shows the documented reserve on every re-counted prompt,
and every turn is `fits`. A2's leak count is 0. B's independent-retention
verdict is recorded.
12. **WP-B's memory limitation is reported, not summarised away.** The v1.1
release report states each of these, and never shortens them to "WP-B
passed":
- deterministic independent-memory recovery: **PASS**;
- reference-model independent-memory recovery: **FAIL** on the
precondition-valid attempt;
- the failing stage: **memory creation**, the summariser's content
selection;
- the owner's decision to accept that limitation for v1.1.
The release long run's independent-retention verdict (item 6) is read
against it. A recovery there is reported as evidence, not as a reversal of
the limitation, unless it meets every isolation precondition.
13. **Carried residuals are listed with their status:**
- the mid-reply narrator instruction echo that A2's trailing cleanup does
not remove (WP-B.1 §K);
- the doubled full stop in the memory-search scene text (WP-B.2 §R 10).
7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0
protocol shapes in stored narration.
8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all
checks.
9. **Upgrade from a real v1.0.0 database.** A database is created by the
`v1.0.0` tree and played. It has retained history, an undone head, Save
Points, memories, summaries, imported knowledge and a narration-length
choice, and uses loopback or placeholder settings with no real hostnames.
Opened by the candidate, its transcript, head, Redo availability, state, Save
Points, memories, knowledge and settings compare identical, and any migration
is forward-only with schema parity. A v1.0.0 export imports into v1.1.
Because the format stays v3, a v1.1 export of that campaign is checked for
import into v1.0.0, and the result is reported.
10. **Documentation:** `README.md`, `DEVELOPMENT.md`, the as-implemented sections
named by each package, and `VERSION.md`.
11. **Owner events**, each separate: the signed release commit, `main`, and a
`v1.1.0` tag.
## 12. Scope recommendation
**Recommended: Option 1, a focused v1.1.**
| Ship in v1.1 | Defer to v1.2 | Future / optional |
| --- | --- | --- |
| WP-A1 context safety reserve | Scheduled backups (10b) | Real media-provider adapters (11; K05, K06) |
| WP-A2 protocol echo and neutral instructions | Raising or streaming the import limit (7) | Whole-transcript copy, story search (§77, §78) |
| WP-B independent memory retention | Reducing in-sentence restatement (2) | Discarded-history recovery screen (§63) |
| WP-C browser release coverage | Stale-scene and derived-contamination detectors in the identity diagnostic (8) | Tablet layout |
| WP-D recovery honesty | Cross-layer duplication suppression (17) | Editable RPG world state |
| WP-E control-boundary contrast | | Media debt: `ambience`, visual-profile recovery and UI (20) |
| | | Trusted-LAN residual limits: address pinning against DNS rebinding (21) |
| | | Comparative narrator-model recommendation (13) |
**Why focused.** Four of the six packages close silent failure modes or
verification gaps that the v1 evidence itself named. The other two are small and
deterministic. Everything deferred is either a new feature, or needs a policy
decision the owner has not been asked for, or rests on an occurrence that has not
happened. A broader v1.1 that took scheduled backups and media would add the two
packages with the most new surface. It would also push the long-run
re-validation, which every prompt change already requires, further from the
changes it validates.
**If B's real-model criterion cannot be met.** Precondition-valid runs may recover
nothing even after B.2. In that case the owner chooses between shipping v1.1 with
B's deterministic criteria met and the real-model result recorded as a residual
risk, or holding v1.1 for B. The plan does not pre-decide this.
## 13. Coding briefs needed next
These briefs are not written here, and none is to be executed from this document.
1. **WP-A1 — context-window safety reserve. Write this one first.** It carries
the highest user risk, has no dependencies, is deterministically testable,
and produces the provenance that later real-model runs read. The brief must
settle three things: the tokenizer divergence the reserve is sized for, whether
calibration is built or detection alone, and whether any setting is added
(recommended: none).
2. WP-A2 — protocol echo hardening and genre-neutral instructions. It can be
drafted while A1 is under review.
3. WP-B — independent memory retention: B.1 diagnostic, then B.2 fix. Two briefs
are an option if B.1's findings should be reviewed before a fix is chosen.
4. WP-C — browser release coverage.
5. WP-D — recovery honesty.
6. WP-E — control-boundary contrast.
7. v1.1 release validation, written only after every in-scope package is
accepted.
## 14. Decisions for the owner
- Confirm Option 1, the focused v1.1.
- A1: the divergence the reserve must absorb; calibration or detection only; that
a detected truncation is flagged, not turned into a failed turn.
- B: whether B.1 and B.2 are one brief or two; the choice in §12 if the
real-model criterion is not met.
- D: that the 20 MB import limit stays in v1.1.
- E: approval of the visual change.
- Whether v1.1 real-model validation stays on the reference narrator, as this plan
recommends.
- The report naming and rotation in `planning/reports/` (§5 rule 5).