Planning v4.1: record the v1.0.0 release, and plan v1.1
Documentation only. No product code, requirement, acceptance test or
schema changes.
Post-release correction. v4.0 was written before the closeout commit was
signed (432f041), main was fast-forwarded to it, and the signed v1.0.0 tag
was pushed. Current-state wording now says so in README.md,
planning/README.md, BUILD-MILESTONES.md and VERSION.md. BUILD-MILESTONES.md's
header had been stale since M8. The M11 report is not edited: its §T
records the state at closeout.
v1.1 plan. planning/V1.1-PLAN.md triages the post-v1 backlog and the other
recorded v1 residual risks, and orders them into work packages, not
milestones:
- A1: a context-window safety reserve, plus reporting a turn the server
truncated
- A2: removing protocol echoes from stored narration, and a genre-neutral
state rule
- B: long-term memory retention that holds without help from state
- C: browser coverage of Retry, Save Points, correction, length, failure
and export download
- D: integrity_check on backups, and a warning when an export exceeds the
import limit
- E: WCAG 1.4.11 control-boundary contrast
Scheduled backups, the import limit, identity detectors and duplication
suppression move to v1.2; media adapters are future work. The first brief
to write is A1.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
432f04100b
commit
ac465ed867
@@ -1,6 +1,6 @@
|
||||
# Adventure Storyteller — Production Build Milestones
|
||||
|
||||
**Status:** In implementation. M1-M8 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6, M7 and M8: 2026-09-06, each of the last four after an independent review and a corrective pass). **M9 — Export, Backup, Recovery, and Migration Hardening — is next, and has not been started.**
|
||||
**Status:** **Complete. M1-M11 are closed, and v1.0.0 was released on 2026-09-14** (signed tag `v1.0.0` on signed commit `432f041`, which `main` also points at). This document is the v1 milestone history and is not extended. Post-v1 work is planned as v1.1 work packages in `V1.1-PLAN.md`, not as further milestones.
|
||||
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
||||
|
||||
## 1. Purpose
|
||||
@@ -1466,6 +1466,10 @@ applicable (report §T). The acceptance takes effect with the owner's signed
|
||||
closeout commit. **No release tag exists**: tagging `v1.0.0` is a separate
|
||||
decision, and the tag must point at that signed commit.
|
||||
|
||||
*Post-release note (2026-09-14):* both events have since happened. The closeout
|
||||
commit was signed as `432f041`, `main` was fast-forwarded to it, and the signed
|
||||
tag `v1.0.0` points at it. The paragraph above is kept as it stood at closeout.
|
||||
|
||||
**M1-M11 are all complete. There is no M12.** Post-v1 work is backlog, listed
|
||||
below under *Post-v1 backlog*, and none of it is an unfinished v1 milestone.
|
||||
|
||||
@@ -1550,6 +1554,10 @@ Work recorded for after v1. None of it is a v1 requirement or an unfinished v1
|
||||
milestone, and none of it has a brief. Each item needs one before work begins.
|
||||
Sources are the M11 report's §P and §S.6.
|
||||
|
||||
**Triaged on 2026-09-14 in `V1.1-PLAN.md`**, which orders it into v1.1 work
|
||||
packages, a v1.2 list and future work. The list below is kept as recorded at the
|
||||
M11 closeout; `V1.1-PLAN.md` is where its disposition now lives.
|
||||
|
||||
- **Context-window safety margin.** The largest prompts leave 23-42 real tokens,
|
||||
and Ollama cuts an over-window prompt with no error. Consider a deliberate
|
||||
reserve, or counting with the narrator's own tokenizer.
|
||||
|
||||
+42
-27
@@ -2,11 +2,14 @@
|
||||
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
**milestones M1 through M11 complete**. **M11 was accepted at its closeout
|
||||
(2026-09-14), and the v1 release gate passed** on the release-candidate tree.
|
||||
There is no further planned milestone. The signed closeout commit and any
|
||||
`v1.0.0` tag are separate events and the repository owner's to perform. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 planning has begun.**
|
||||
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
|
||||
through M11 complete and closed**. M11 was accepted at its closeout
|
||||
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
|
||||
owner then signed the closeout commit (`432f041`), fast-forwarded `main` to it,
|
||||
and created and pushed the signed tag `v1.0.0` pointing at it. v1.1 development
|
||||
is on the `v1.1-development` branch, from that commit, and is planned in
|
||||
`V1.1-PLAN.md` as work packages, not milestones. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
|
||||
independent review found a real defect and a corrective pass fixed it.
|
||||
|
||||
@@ -46,13 +49,18 @@ The closeout repeated the browser, offline and identity runs on the exact
|
||||
release-candidate tree (`3652dc6`), and recorded the acceptance (report §S and
|
||||
§T). **The v1 release gate passed.**
|
||||
|
||||
**Not yet happened**, and each a separate event:
|
||||
- the owner signing the closeout commit;
|
||||
- merging to `main`, which still points at M10;
|
||||
- any `v1.0.0` tag.
|
||||
**The release, 2026-09-14.** The three events the closeout left to the owner
|
||||
have all happened:
|
||||
- the closeout commit is signed, as `432f041`;
|
||||
- `main` was fast-forwarded to `432f041`;
|
||||
- the signed tag `v1.0.0` points at `432f041`, and is pushed.
|
||||
|
||||
**There is no further planned milestone and no M12.** Post-v1 work is backlog
|
||||
(`BUILD-MILESTONES.md`, *Post-v1 backlog*).
|
||||
The M11 report's §T still lists the last two as not done. That table records the
|
||||
state at closeout and is deliberately left as written.
|
||||
|
||||
**There is no M12.** Post-v1 work is organised as **v1.1 work packages** in
|
||||
`V1.1-PLAN.md`, which triages the *Post-v1 backlog* recorded in
|
||||
`BUILD-MILESTONES.md`.
|
||||
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
and why.
|
||||
@@ -122,7 +130,8 @@ Two standing qualifications:
|
||||
| `SECURITY-THREAT-MODEL.md` | The trust boundary, and the inference endpoint policy as implemented. |
|
||||
| `MEDIA-EXTENSION-CONTRACT.md` | The contract future image/video/audio/TTS/STT work must fit. |
|
||||
| `BROWSER-UX-SPEC.md` | The browser surface, and what is deliberately not in it. |
|
||||
| `BUILD-MILESTONES.md` | M1-M11, what each delivers, what is done, and the notes each milestone leaves its successors. |
|
||||
| `BUILD-MILESTONES.md` | M1-M11, what each delivers, what is done, and the notes each milestone leaves its successors. Closed history; not extended. |
|
||||
| `V1.1-PLAN.md` | **Post-v1 work.** The triaged backlog, the ordered v1.1 work packages with acceptance criteria, the v1.1 release criteria, and what is deferred. |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | The pass/fail contract v1 is measured against. |
|
||||
| `TEST-CAMPAIGN-FIXTURE.md` | The standard campaign the acceptance tests are run on. |
|
||||
| `VERSION.md` | Package revision history: what each closeout changed. |
|
||||
@@ -133,7 +142,8 @@ Two standing qualifications:
|
||||
1. This file.
|
||||
2. `SPECIFICATION.md`
|
||||
3. `TECHNICAL-DESIGN.md`
|
||||
4. `BUILD-MILESTONES.md` — find the milestone you are being asked to do.
|
||||
4. `V1.1-PLAN.md` — find the work package you are being asked to do.
|
||||
`BUILD-MILESTONES.md` is the v1 history behind it.
|
||||
5. `STORY-BRANCH-SEMANTICS.md`
|
||||
6. `DATA-MODEL.md`
|
||||
7. `CONTEXT-AND-MEMORY.md`
|
||||
@@ -185,7 +195,9 @@ that is the one the next milestone's planning has to consult:
|
||||
It stays in `reports/` after acceptance. The rotation moves a report to the
|
||||
archive when the next milestone's report is written, and after M11 there is no
|
||||
next milestone. Nothing here calls for moving it, so no new convention was
|
||||
invented to do so.
|
||||
invented to do so. `V1.1-PLAN.md` asks each v1.1 work package for a report of
|
||||
its own; when the first is written, it joins this one in `reports/` and the
|
||||
rotation is decided then.
|
||||
|
||||
M9's and M10's reports both moved to `archive/milestone-reports/` when this one
|
||||
was written. M9's had been kept here past its turn because M9 was unaccepted;
|
||||
@@ -337,27 +349,30 @@ Milestone M11 COMPLETE / ACCEPTED (2026-09-14)
|
||||
|
|
||||
v
|
||||
v1 release gate PASSED (2026-09-14)
|
||||
signed release commit and v1.0.0 tag:
|
||||
the repository owner's, not yet done
|
||||
|
|
||||
v
|
||||
Post-v1 backlog only; no milestone planned
|
||||
v1.0.0 RELEASED (2026-09-14)
|
||||
signed tag on signed commit 432f041;
|
||||
main points at the same commit
|
||||
|
|
||||
v
|
||||
v1.1 PLANNING (from 2026-09-14)
|
||||
branch v1.1-development; V1.1-PLAN.md
|
||||
work packages, not milestones
|
||||
```
|
||||
|
||||
## Stop Rule
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
**One work package at a time. Do not begin a work package before its brief
|
||||
exists.**
|
||||
|
||||
**Every planned milestone is complete, and M11 is accepted.** The v1 release gate
|
||||
passed on 2026-09-14 (M11 report §T).
|
||||
**Every v1 milestone is complete, M11 is accepted, and v1.0.0 is released**
|
||||
(2026-09-14, tag `v1.0.0` on `432f041`).
|
||||
|
||||
The next actions are the owner's:
|
||||
1. sign the closeout commit;
|
||||
2. decide on the `v1.0.0` tag, pointed at that signed commit.
|
||||
|
||||
Neither is a milestone. **Do not begin post-v1 work as though it were a v1
|
||||
milestone.** Anything after v1 starts from the *Post-v1 backlog* in
|
||||
`BUILD-MILESTONES.md`, with a brief of its own.
|
||||
**Do not begin post-v1 work as though it were a v1 milestone, and do not create
|
||||
M12.** v1.1 work is the ordered work packages in `V1.1-PLAN.md`. Each starts
|
||||
only from a coding brief of its own, and stops at its own boundary for review.
|
||||
The plan names the brief to write first.
|
||||
|
||||
All three questions the M8 debt raised against M9 are settled and recorded:
|
||||
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
|
||||
|
||||
@@ -0,0 +1,889 @@
|
||||
# Adventure Storyteller — v1.1 Plan
|
||||
|
||||
**Status:** PLANNING. Written 2026-09-14 on `v1.1-development`, from the signed
|
||||
v1.0.0 release commit `432f041`. No work package has started, and no v1.1
|
||||
version or tag exists.
|
||||
|
||||
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
|
||||
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
|
||||
specification and acceptance contract are unchanged: every package below
|
||||
improves how an existing requirement is met. No package adds a requirement.
|
||||
|
||||
---
|
||||
|
||||
## 1. Terminology
|
||||
|
||||
- **Work package (WP).** One independently reviewable unit of v1.1 work, with
|
||||
its own coding brief, its own report, and its own review. Work packages are
|
||||
lettered. **They are not milestones, and there is no M12.**
|
||||
- **Brief.** The coding prompt for one work package. It holds only the
|
||||
objective, the source documents, the scope and non-scope, the acceptance
|
||||
criteria, and the stop condition. No work package begins before its brief
|
||||
exists.
|
||||
- **Reference narrator.** `qwen2.5:3b-instruct`, with and without `num_ctx`
|
||||
16,384 baked in, and `nomic-embed-text` for embeddings. These are the models
|
||||
the v1 evidence used (M11 report §E.1). v1.1 real-model evidence uses the
|
||||
same models so that its numbers compare with v1's.
|
||||
- **Reference hosts.** The CPU reference host serves Ollama over HTTPS with a
|
||||
private CA, which is the A06 evidence path. The GPU inference host serves
|
||||
plain HTTP on the LAN and is used for long runs (M11 report §E.1).
|
||||
- **Planning package version** (`VERSION.md`, v4.x) and **product version**
|
||||
(v1.0.0, v1.1.0) are separate numbers. `SECURITY-THREAT-MODEL.md`'s own
|
||||
"Status: v1.1" line is that document's revision label from Phase 0B. It does
|
||||
not refer to this release.
|
||||
|
||||
## 2. Baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Release | **v1.0.0**, 2026-09-14 |
|
||||
| Release commit | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, signed by the owner |
|
||||
| Tag | `v1.0.0`, a signed annotated tag on that commit, pushed |
|
||||
| `main` | `432f041`, the same commit |
|
||||
| v1.1 branch | `v1.1-development`, created at `432f041` |
|
||||
| v1 contract | 82 REQUIRED FOR V1 tests: 81 PASS and H09 NOT APPLICABLE (M11 report §T) |
|
||||
| Schema | migrations up to 94 (`LATEST_VERSION` 94) |
|
||||
| Bundle format | `ai-dnd-adventure-v3` |
|
||||
| Provenance | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` |
|
||||
|
||||
## 3. v1.1 goals
|
||||
|
||||
1. **No silent loss of what the narrator is given.** A prompt must not reach the
|
||||
server's window edge by an accident of tokenizer arithmetic. A turn the
|
||||
server truncated must be reported.
|
||||
2. **No protocol in the story.** The shapes of the application's own protocol
|
||||
that a narrator copies must not be stored as narration. The fixed
|
||||
instructions must not hand every campaign one genre's nouns to copy. Story
|
||||
prose must not be removed in the process.
|
||||
3. **Memory that remembers on its own.** An old fact must be recoverable from
|
||||
the memory bank itself, with provenance, when nothing else carries it. This
|
||||
must not weaken lineage safety.
|
||||
4. **Release evidence without API-only gaps.** Every reader-facing history,
|
||||
state, length, failure and export behaviour must be driven in a real browser.
|
||||
5. **Recovery that says when it cannot recover.** Backups get a full integrity
|
||||
check, and an export that import would refuse must say so.
|
||||
6. **Control boundaries that meet WCAG 1.4.11.**
|
||||
|
||||
And throughout: **every v1.0.0 campaign opens unchanged in v1.1.**
|
||||
|
||||
## 4. Non-goals
|
||||
|
||||
- New product features: media generation, TTS, STT, story search, whole-transcript
|
||||
copy, a discarded-history recovery screen, a restore button, tablet redesign.
|
||||
- Changes to the history, branch, head or Save Point architecture (ADR 005,
|
||||
ADR 012).
|
||||
- Changes to the authoritative state model, its event vocabulary or its
|
||||
validator (ADR 010, ADR 013). Changes to the prompt *wording* that describes
|
||||
the vocabulary are in scope, in WP-A2.
|
||||
- Changes to the knowledge authority classes or their retrieval.
|
||||
- A new bundle format version.
|
||||
- Any new runtime dependency, any network access beyond the configured
|
||||
inference endpoint, and any first-use download, including per-model
|
||||
tokenizers.
|
||||
- Supporting inference servers other than Ollama, beyond what
|
||||
`context_window_override` already allows.
|
||||
- Re-writing stored narration or memories of existing campaigns automatically.
|
||||
- Redesigning identity or state handling on the strength of post-M8 finding D,
|
||||
which has not been reproduced.
|
||||
|
||||
## 5. Rules for every work package
|
||||
|
||||
The v1 milestone rules (`BUILD-MILESTONES.md` §2) apply to every work package,
|
||||
and so do the following:
|
||||
|
||||
1. **Compatibility default.** Existing v1.0.0 campaign databases open
|
||||
unchanged. Any migration is forward-only and additive, and is tested by
|
||||
opening a real v1.0.0 database. Fresh-install and upgraded schemas must still
|
||||
compare identical (`test_m11_migration.py`). v1.0.0 bundles still import.
|
||||
2. **The bundle stays `ai-dnd-adventure-v3`** unless a brief makes the case under
|
||||
M9's semantic test and the owner agrees. Adding a key inside stored evidence
|
||||
is not a format change when an absent key unambiguously means "not
|
||||
recorded".
|
||||
3. **The v1 contract is the regression floor.** No REQUIRED FOR V1 test is
|
||||
retired, relaxed or reclassified.
|
||||
4. **Evidence discipline** (M11 report §O). Any product change made after a
|
||||
black-box run was taken invalidates that run for release purposes. Harness
|
||||
defects and product defects are reported separately. A check that cannot fail
|
||||
is not evidence.
|
||||
5. **Each package has a report** (`planning/reports/V1.1-WP-<id>-REPORT.md`).
|
||||
It records what was built, the acceptance results with measurements, what
|
||||
was not verified, and the as-implemented documentation changes. The report
|
||||
rotation for `reports/` is decided when the first one is written
|
||||
(`planning/README.md`).
|
||||
6. **Real-model runs longer than a few minutes on the GPU host** require the
|
||||
power, link and kernel logging in `DEVELOPMENT.md`, started before the run.
|
||||
7. **No real hostnames, addresses or people's names** in committed files,
|
||||
fixtures or reports. Evidence stays under `$HOME`, never `/tmp`.
|
||||
8. **Commits and tags are the owner's.** A package ends with its changes staged
|
||||
for a signed commit.
|
||||
|
||||
---
|
||||
|
||||
## 6. Backlog triage
|
||||
|
||||
Severity reflects user risk: **High** means story corruption or lost continuity
|
||||
that the reader is not told about. **Medium** means silent degradation or a
|
||||
real verification gap. **Low** means visible, rare or cosmetic.
|
||||
|
||||
### 6.1 The recorded post-v1 backlog
|
||||
|
||||
| # | Item | Source | Current evidence | Severity | Disposition |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| 1 | Context-window safety margin | M11 §P risk 3, §N; post-v1 backlog | The largest prompts left 23, 34 and 42 real tokens of headroom in the GPU runs. The application's `cl100k_base` count ran 16 tokens below the narrator's on every prompt measured. The only slack is `OUTPUT_SAFETY_MARGIN = 64` (`context/builder.py`), which is fixed and must also absorb the section separators. Ollama 0.34 cut an over-window synthetic prompt to 8,194 tokens with no error. Truncation drops the oldest tokens, which are the narrator's rules and the canon. The server's reported usage is already stored per turn (`snapshot["usage"]`, `routers/adventures/turns.py`) and read by nothing. | **High**: silent, invisible, and reachable by a model whose tokenizer diverges further | **v1.1, WP-A1** |
|
||||
| 2 | Narrator restating prompt and state text | §P risk 4, §G.4 | The state's fact line was restated as a sentence on 74 of 104 turns in the evidence run. Phrases such as "Scene set, continue your adventure." and lines opening "Memory:" are model-invented, and none is an application string (verified by search). The restatements sit inside story sentences. | **Medium**: narration quality, and it is the path by which M04's fact reached memory, which masks item 5 | **v1.1, WP-A2** measures it. Reducing in-sentence restatement is **v1.2**, because removing it means judging prose |
|
||||
| 3 | Protocol leakage: event-call syntax, stray headings, repeated length hints | §P risk 6, §S.6 | On 4 of 10 turns of the closeout identity run (4,096 window) and 0 of 104 in the 100-turn runs. Causes verified in code. `events.vocabulary_for_prompt` shows every event as `name(field, …)`, which is the notation copied. `_is_echoed_instruction` requires both "state block" and "events list", and the length hint's tail says only "state block". A lone trailing `Scene:` is not removed. Stored text is replayed as history (TECHNICAL-DESIGN §15.4). | **Medium-high**: silent and self-reinforcing, but rare at a full window | **v1.1, WP-A2** |
|
||||
| 4 | Genre-specific example in the fixed state rule | §P risk 16; `narrative/extract.py` `EMIT_RULE` | The example names `mara`, `silver-key`, `old-abbey` and `aldric`. In the office-meeting identity run the narrator proposed giving `silver-key` to Alice, and the validator refused it. No state was corrupted. | **Medium**: every non-fantasy campaign gets fantasy nouns to copy | **v1.1, WP-A2** |
|
||||
| 5 | Independent long-term-memory retention | §P risk 5, §G.4; F02 and M04 qualifications | No run showed a planting-era memory carrying the planted fact. Recovery ran from state to the narrator's restatement to memory. Code facts that bear on it, none yet proven to be the cause: a memory is at most 50 words for a 6-action block. The summariser sees the block's *last* 2,000 tokens (`truncate_to_last_tokens`), so an early fact in a long block can be cut. Eviction is least-recently-used above `memory_bank_capacity` (default 80), so a never-retrieved early memory is the first to go once a campaign passes about 480 actions; the 33-memory evidence run never reached that. Ranking is cosine similarity plus a pin (CONTEXT-AND-MEMORY §20). | **High** for long campaigns: lost continuity, seen only as the story forgetting | **v1.1, WP-B** |
|
||||
| 6 | Browser automation gaps: Retry, Save Point, state correction, narration length, failed generation, export download | §S.3, §T qualifications, §P risk 11 | Proved through the API, the component suite and the 100-turn campaign, but never driven in a browser. The export control is exercised only as far as the click, because a `blob:` download does not leave the headless snap Firefox. No defect is known. | **Medium**: a verification gap on the release path | **v1.1, WP-C** |
|
||||
| 7 | Practical export and bundle-size ceiling | §P risk 12, §K; M9 debt | The import body limit is 20 MB (`limits.MAX_IMPORT_BODY_BYTES`). M9's conservative ceiling is about 279 turns. The real 100-turn campaign was 2.74 MB, about 13 kB per action, or roughly 1,600 actions. A larger campaign exports and is then refused on import, and nothing says so at export. | **Medium** impact, low frequency | **v1.1, WP-D**: warn at export. Raising the limit or a streaming import is **v1.2** |
|
||||
| 8 | Identity-confusion diagnostic follow-up | §P risk 8, §S.5 | Not reproduced. The diagnostic cannot see a stale scene, has never run with memory on, and has never run at a 16,384 window. It found no product deficiency. | **Low**: no occurrence since the original, whose campaign is gone | **No v1.1 package.** The diagnostic is re-run as a WP-A2 regression and in the v1.1 release gate, with memory on. A stale-scene detector is **v1.2**. The next real occurrence is classified with `tools/m11_identity.py` |
|
||||
| 9 | WCAG 1.4.11 control-boundary contrast | §P risk 10, §L | Measured at 1.33:1 resting and 1.75:1 on hover (`--border` and `--border-bright` against `--bg-panel`). M11's argument that the label identifies the control is weakest for text inputs, where the boundary shows where to type. | **Low-medium** | **v1.1, WP-E** |
|
||||
| 10a | Backup integrity: `integrity_check` | §P risk 14; M9 | `backup.py` runs `PRAGMA quick_check`, which skips index-content verification. | **Low** | **v1.1, WP-D** |
|
||||
| 10b | Scheduled backups | §P risk 14; M9 | None exist; backups are manual. They need owner policy on interval, retention, location and disk use, and must not contend with a turn's write lock (§O.7). | **Medium** impact, but a new feature | **v1.2** |
|
||||
| 11 | Real media-provider adapters | §P risk 13; M10; K05 and K06 (FUTURE) | The seam has no consumer. Adding one brings dependencies, a stricter endpoint rule (`DEVELOPMENT.md`) and a new UI. | n/a: a feature, not a risk to existing stories | **Future / optional**, not v1.1 |
|
||||
|
||||
### 6.2 Other recorded v1 residual risks and debt
|
||||
|
||||
| # | Item | Source | Disposition |
|
||||
| --- | --- | --- | --- |
|
||||
| 12 | The state lags the narration at 4,096 with the 3B narrator (1 proposal in 10 applied) | §P risk 17, §S.5 | Model behaviour, not a defect. WP-A2's real-model runs record the proposal outcome counts. No package |
|
||||
| 13 | Every realistic observation is a 3B model's | §P risk 2 | No package. v1.1 keeps the reference narrator for comparability. A comparative model recommendation stays open (`planning/README.md`, *Still open*) |
|
||||
| 14 | A campaign that never chose a narration length keeps the pre-M11 hint | §P risk 15 | Deliberate. No change. WP-C drives the control |
|
||||
| 15 | The GPU host dropped off the PCIe bus after the evidence run | §P risk 7, §E.1 | Operational. Rule 6 in §5 applies to every long run |
|
||||
| 16 | Release timings are two hosts' | §P risk 1 | No action. No performance requirement exists or is invented |
|
||||
| 17 | Cross-layer duplication: one fact in state, memory, history and a passage at once | CONTEXT-AND-MEMORY §22; open since M6 and M7 | **v1.2.** It costs tokens (item 1) and is related to item 2. WP-B may touch ranking only where its diagnostic requires |
|
||||
| 18 | Memory-ranking factors beyond similarity and pin not implemented | CONTEXT-AND-MEMORY §20 | Inside WP-B's bounded fix menu, only if WP-B's diagnostic places the failure at ranking |
|
||||
| 19 | M8 carried debt: no whole-transcript copy or story search (§77, §78), no discarded-history recovery screen (§63), tablet untuned, RPG world state read-only | `BUILD-MILESTONES.md` M8 and M9 | **Future backlog.** Features, not reliability |
|
||||
| 20 | M10 debt: `ambience` is an empty shape; a deleted visual profile is unrecoverable in the app; profiles are API-only | M11 §D | **Future**, with media |
|
||||
| 21 | A hostile host on the trusted LAN; the DNS-rebinding interval | `SECURITY-THREAT-MODEL.md` §10A | Accepted for v1. **Future.** No v1.1 package |
|
||||
| 22 | A stale `chunk_id` in a restored snapshot | M9 | No action. It is a React key, not a live pointer |
|
||||
| 23 | Developer-venv residue (`quickjs`, `psycopg`) | §M | No action. It does not ship |
|
||||
| 24 | A Settings warning comparing the real window to the budget | M8 debt, "unowned" | **Closed by M11**: Settings reports the window and warns |
|
||||
| 25 | Root `README.md` has stale facts: "1,191 backend tests" (1,421 at closeout), and the Screenshots paragraph calls screenshots "a job for the UI pass in M8" | found in this review | Documentation debt, not release status, so it is not changed in v4.1. It is fixed in WP-A1's documentation update |
|
||||
|
||||
---
|
||||
|
||||
## 7. Grouping decisions
|
||||
|
||||
**The suggested "narrator boundary hardening" package is split into A1 and A2.**
|
||||
They share a theme but not a mechanism or a test method.
|
||||
|
||||
- **A1**, the context reserve, is budget arithmetic in `context/builder.py` and
|
||||
`contextwindow.py`. It is proved deterministically with a scripted provider
|
||||
that reports token counts.
|
||||
- **A2**, the protocol echo, is the contract between the prompt's wording
|
||||
(`extract.EMIT_RULE`, `events.vocabulary_for_prompt`, the length hint) and the
|
||||
extractor. It is proved with a replay corpus of real narration and
|
||||
real-model runs.
|
||||
|
||||
Bundling them would make one review carry two unrelated risk profiles.
|
||||
|
||||
A2's three items belong together. The example slugs and the call notation are
|
||||
the *source* of the copied text, and the extractor is its *sink*. Changing one
|
||||
without the other means proving the change twice. Both also alter the stored
|
||||
prompt, so they invalidate the same evidence.
|
||||
|
||||
**B stands alone.** It is the only package whose success depends on model output
|
||||
quality. Its test design, a fact that nothing but memory carries, is the hard
|
||||
part. It follows A2, so that it measures the prompt v1.1 will ship.
|
||||
|
||||
**D is narrowed to recovery honesty**, meaning `integrity_check` and the export
|
||||
warning. Both are small and deterministic, and both are recovery-path changes
|
||||
with no policy question. Scheduled backups need owner decisions and have real
|
||||
design risk, including write-lock contention. Raising or streaming the import
|
||||
limit changes a DoS guard. Both move to v1.2.
|
||||
|
||||
**E stays focused** on control boundaries. The same measurement found no other
|
||||
accessibility defect (M11 §L).
|
||||
|
||||
**F is not a package.** Nothing concrete in the product is deficient. The
|
||||
diagnostic's known blind spots are a stale scene, memory off, and one window
|
||||
size. The first two can be addressed by running it differently: the v1.1 gate
|
||||
runs it with memory on. The stale-scene detector is new diagnostic work with no
|
||||
occurrence to justify it yet, so it is v1.2.
|
||||
|
||||
**G is not in v1.1.** It adds no reliability or quality to existing stories, it
|
||||
brings dependencies and a new endpoint surface, and its tests (K05, K06) are
|
||||
FUTURE. `MEDIA-EXTENSION-CONTRACT.md` and the M10 seam remain authoritative for
|
||||
whenever it is taken up.
|
||||
|
||||
---
|
||||
|
||||
## 8. Work packages
|
||||
|
||||
### WP-A1 — Context-window safety reserve
|
||||
|
||||
**Objective.** An assembled prompt plus its reply reserve stays at least a
|
||||
documented safety reserve below the effective window, whatever tokenizer the
|
||||
narrator uses. Where the server reports its own prompt count, a turn that
|
||||
exceeded the window, or was evidently truncated, is recorded and shown, never
|
||||
silent.
|
||||
|
||||
**Rationale.** §15.2 closed "budget larger than the window". It did not close
|
||||
"our count is not the server's count". Measured headroom was 23-42 tokens
|
||||
(item 1), and the fixed 64-token margin absorbs separators as well as tokenizer
|
||||
drift. A narrator whose tokenizer runs a few percent heavier than
|
||||
`cl100k_base` overflows a full 16,384 window. The failure deletes the canon at
|
||||
the front of the prompt with every request returning 200. The data to detect it
|
||||
is already stored with each turn and unread.
|
||||
|
||||
**Scope.**
|
||||
- A safety reserve that scales with the effective budget and has a floor. It
|
||||
replaces or supplements `OUTPUT_SAFETY_MARGIN`, is defined once, and applies
|
||||
to verified windows, declared overrides and an unverified configured budget
|
||||
alike. The brief states the tokenizer divergence it is sized to absorb.
|
||||
- Reading the server-reported prompt-token count from the turn's usage. It is
|
||||
recorded beside the application's count in the turn's window provenance.
|
||||
Recorded states:
|
||||
- `fits`;
|
||||
- `exceeded`: reported count plus reply reserve is over the window;
|
||||
- `truncation_suspected`: the reported count is materially below the
|
||||
application's count;
|
||||
- `unknown`: no usage was reported.
|
||||
|
||||
`exceeded` and `truncation_suspected` are surfaced in the context inspector,
|
||||
and in the turn's response so the reader sees a notice.
|
||||
- **Optional, decided in the brief:** calibrating the reserve from observed
|
||||
server-to-application ratios per endpoint and model. It must be bounded and
|
||||
quantised so that it does not re-price the prompt prefix every turn (§15.3).
|
||||
Detection is required either way.
|
||||
- `ContextOverflow` remains the explicit failure when protected context plus
|
||||
both reserves exceeds the budget.
|
||||
- `TECHNICAL-DESIGN.md` §15.2 and `DEVELOPMENT.md`'s context-window section, as
|
||||
implemented. The root `README.md` stale facts from item 25.
|
||||
|
||||
**Non-scope.** The history block trim mechanism, section order, the knowledge
|
||||
budget share, a per-model tokenizer or any tokenizer download, a hard-coded
|
||||
window, raising any window, a provider abstraction, a Settings redesign, and
|
||||
failing or discarding a turn after the fact.
|
||||
|
||||
**Likely affected.** `backend/app/context/builder.py`,
|
||||
`backend/app/contextwindow.py`, `backend/app/routers/adventures/turns.py` and
|
||||
`insights.py`, `backend/app/providers/openai_compatible.py` (usage, read-only),
|
||||
the context inspector panel in `frontend/src/pages/Play/`. Tests:
|
||||
`test_m11_context_window.py`, `test_m11_declared_window.py`,
|
||||
`test_history_block_trim.py`.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. **Per configuration.** The application's count plus the output reserve plus
|
||||
the safety reserve is no more than the effective budget, and the safety
|
||||
reserve is at least its documented value. This holds for each of: a
|
||||
verified 4,096 window, a verified 16,384 window, a declared override with no
|
||||
probe, and an unverified configured budget. There is a test per
|
||||
configuration.
|
||||
2. **Divergence.** A scripted provider reports prompt counts at 1.00×, and at
|
||||
the brief's stated divergence, of the application's count, across a campaign
|
||||
that fills the window. At both, no turn's reported count plus reply reserve
|
||||
exceeds the window. At 1.25×, either no turn exceeds (calibration built) or
|
||||
every exceeding turn is recorded `exceeded`. **An unflagged overflow fails.**
|
||||
3. **Truncation.** A scripted provider reports 8,194 tokens against an
|
||||
application count near 15,700. The turn is recorded `truncation_suspected`
|
||||
in stored provenance, shown in the inspector, and reported in the turn's
|
||||
response. A provider that reports no usage records `unknown`, never `fits`.
|
||||
4. **Explicit overflow.** Protected context that fits without the safety
|
||||
reserve but not with it raises `ContextOverflow` with an actionable message.
|
||||
The story and state are unchanged (A05, L01).
|
||||
5. **Canon survives.** `test_the_canon_at_the_front_survives_a_window_far_too_small`
|
||||
and `test_without_the_cap_the_same_prompt_would_have_overflowed` both pass
|
||||
with the reserve in place.
|
||||
6. **Cache stability.** At steady state the history floor still moves in blocks
|
||||
(`test_history_block_trim.py`). If calibration is built, a test shows the
|
||||
budget changes only when the observed ratio crosses a stated step.
|
||||
7. **Real model.** A long run of at least 50 turns runs at a 16,384 window with
|
||||
memory on, using the reference narrator on the GPU host. Its ten largest
|
||||
stored prompts are re-counted by the server, as §N did. Every one leaves at
|
||||
least the documented reserve, and every turn's provenance reads `fits`. The
|
||||
report carries the headroom table beside v1's.
|
||||
|
||||
**Regression requirements.** The full backend and frontend suites; F03, F04 and
|
||||
M03 coverage; the I-series (a new provenance key must import and export); the
|
||||
offline container; the browser harness's F05 inspector checks.
|
||||
|
||||
**Dependencies.** None. First.
|
||||
|
||||
**Compatibility.**
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Databases | None expected. Provenance lives in the turn snapshot's JSON, so no migration |
|
||||
| Bundles | None. Snapshots carry new keys, and absence means "not recorded" |
|
||||
| Save Points, branches, state, memories, knowledge | None |
|
||||
| Model settings | None preferred. If a setting is added, it takes a default that reproduces the documented reserve, plus an additive migration |
|
||||
| Behaviour | At the largest prompts the history window holds slightly fewer actions. Nothing stored changes |
|
||||
| Docker and local-only | None |
|
||||
|
||||
**Test modes.** Deterministic: yes. Real model: yes. Browser: inspector display,
|
||||
via the harness. Offline: regression only. Long run: yes, at least 50 turns.
|
||||
|
||||
**Risk.** Low implementation risk and high value. Its blast radius is the
|
||||
budget arithmetic of every turn. It changes prompts at full windows, so it
|
||||
invalidates the v1 long-run evidence for v1.1 (§10).
|
||||
|
||||
---
|
||||
|
||||
### WP-A2 — Protocol echo hardening and genre-neutral instructions
|
||||
|
||||
**Objective.** Stored narration no longer keeps the application-owned protocol
|
||||
shapes a narrator copies, and the fixed instructions stop supplying
|
||||
genre-specific identifiers. Recognition uses only strings and syntax the
|
||||
application itself owns, so story prose is not removed.
|
||||
|
||||
**Rationale.** Items 3 and 4. The call notation in the prompt is copied, the
|
||||
length hint's echo escapes `_is_echoed_instruction`, and a lone trailing
|
||||
`Scene:` survives. Each leak is replayed as history, so it teaches the next turn.
|
||||
The example slugs were copied into an office scene.
|
||||
|
||||
**Scope.**
|
||||
- A genre-neutral `EMIT_RULE` example and slug examples. No identifier in the
|
||||
fixed instructions names a fixture entity or a genre noun.
|
||||
- A decision, by measurement, on whether the vocabulary is shown in a
|
||||
notation that does not look callable. The extractor handles call lines
|
||||
regardless.
|
||||
- The length hint's wording becomes a named constant shared by the builder and
|
||||
the extractor, in the way `render.SECTION_HEADINGS` is shared. The extractor
|
||||
then recognises:
|
||||
- a line consisting only of a call to an event named in `events.SPECS`
|
||||
(case-insensitive), optionally `>`-quoted;
|
||||
- the length hint echoed as a trailing bracket, closed or cut off;
|
||||
- a renderer heading that is the reply's final line with nothing after it.
|
||||
- A replay tool that runs the old and new extractors over a corpus of stored AI
|
||||
turns and writes a per-turn diff report.
|
||||
- Measurement only: per-turn counts of state fact-line restatements,
|
||||
model-invented headings, and proposal outcomes (applied, empty, unparseable,
|
||||
refused), added to `tools/m11_long_run.py`.
|
||||
|
||||
**Non-scope.** Removing a fact restated inside a sentence; removing
|
||||
model-invented headings that are not the renderer's; the event vocabulary itself,
|
||||
the validator, the fence protocol, or history replay; any model-based cleanup
|
||||
pass; rewriting stored narration of existing campaigns.
|
||||
|
||||
**Likely affected.** `backend/app/narrative/extract.py`, `events.py` (prompt
|
||||
rendering only), `render.py` (constants), `backend/app/context/builder.py`
|
||||
(length-hint constant), `tests/test_narrative_state.py`, `tests/test_m11_scifi.py`
|
||||
(the J03 vocabulary check), `tools/m11_long_run.py`. Also
|
||||
`TECHNICAL-DESIGN.md` §15.4 and ADR 013's implementation note.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. **Known shapes are removed.** There is a committed case per shape, cut down
|
||||
from the four §S.6 turns and anonymised:
|
||||
- `> set_possession(silver-key, "alice")`;
|
||||
- `> Create_entity(...)`;
|
||||
- the length hint echoed as `[Hard limit: …append the state block well
|
||||
inside the limit.]`, closed and cut off;
|
||||
- a lone trailing `Scene:`.
|
||||
|
||||
The four stored §S.6 texts, replayed, lose exactly those lines. Every other
|
||||
sentence is intact, asserted as set equality of the remaining sentences.
|
||||
2. **Adversarial prose is untouched.** Each of these is a test:
|
||||
- dialogue that mentions `set_possession` mid-sentence;
|
||||
- `create_entity(ship)` inside a story's own ```python fence;
|
||||
- a call-shaped line whose name is not in the vocabulary, such as
|
||||
`> open_door(north)`;
|
||||
- an in-world bracket, `[Hard limit of the reactor: three hours]`;
|
||||
- `Scene:` followed by prose;
|
||||
- `Memory: she remembered the bells`;
|
||||
- a fact restated inside a sentence;
|
||||
- a trailing bracket about a state block that lacks the hint's wording.
|
||||
|
||||
§O.8's existing negative controls all still pass.
|
||||
3. **Replay corpus.** Every stored AI turn available from the v1 evidence runs
|
||||
is replayed through both extractors: §O.8's 443 and the identity run's 10,
|
||||
kept under `$HOME`, not committed. The package passes when all of these
|
||||
hold:
|
||||
- no turn changes except by removing a criterion-1 shape;
|
||||
- every changed turn appears in the diff report, and the WP report reviews
|
||||
it;
|
||||
- the review classifies zero changes as removed story.
|
||||
|
||||
The report gives the counts.
|
||||
4. **False-positive detection.** The replay tool flags any removal that is not
|
||||
the reply's final segment or a whole line matching a criterion-1 shape, and
|
||||
any removal of more than a stated share of a turn's prose. An unreviewed
|
||||
flag fails the package.
|
||||
5. **Genre-neutral instructions.** A test asserts that `EMIT_RULE`,
|
||||
`EMIT_REMINDER`, the length hints and the vocabulary text contain no
|
||||
identifier from either acceptance fixture (Westhaven, Persephone) and none of
|
||||
J03's genre nouns.
|
||||
6. **Real model.**
|
||||
- **Identity diagnostic.** The office fixture at 4,096 on the CPU/HTTPS host
|
||||
gives 0 proposals naming an example identifier (baseline 1 of 10),
|
||||
protocol shapes in 0 of 10 stored turns (baseline 4 of 10), and 0 identity
|
||||
signals. Its `--scripted --inject` self-test still fires.
|
||||
- **Long run.** At least 50 turns at 16,384 with memory on stores protocol in
|
||||
0 turns (baseline 0 of 104). Restatement counts are recorded against the
|
||||
74 of 104 baseline, as a measurement, not a gate.
|
||||
|
||||
**Regression requirements.** The full backend suite, including the 25 §O.8
|
||||
cases; C06 and H05 (refusals are still recorded and shown to the model); J01-J03;
|
||||
the I-series; the browser harness's hostile-narration checks (H04, H06, H07).
|
||||
|
||||
**Dependencies.** After A1. Both change the prompt, and one real-model run can
|
||||
then cover both. The brief can be written while A1 is under review.
|
||||
|
||||
**Compatibility.**
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Databases | None. No migration |
|
||||
| Stored narration of v1.0.0 campaigns | Not rewritten |
|
||||
| Bundles | None |
|
||||
| Save Points, branches | None |
|
||||
| State | None; the validator is unchanged |
|
||||
| Memories and summaries | Future ones summarise cleaner prose. Existing ones are untouched |
|
||||
| Docker and local-only | None |
|
||||
|
||||
**Test modes.** Deterministic: yes. Replay corpus: yes. Real model: yes. Browser:
|
||||
regression only. Offline: regression only. Long run: yes, at least 50 turns, and
|
||||
this run may be the same one as A1's criterion 7 if A1 is already merged.
|
||||
|
||||
**Risk.** Medium. The failure to fear is silent removal of prose. Constants-only
|
||||
recognition, the corpus and the detector are the mitigations.
|
||||
|
||||
---
|
||||
|
||||
### WP-B — Independent long-term memory retention
|
||||
|
||||
**Objective.** An important fact planted early is recoverable at depth 100 or
|
||||
more from the memory bank itself, with provenance to a memory whose source range
|
||||
covers the planting turn. This must hold when neither authoritative state, later
|
||||
narration, the summary nor the history window carries the fact, and lineage
|
||||
safety must be unchanged.
|
||||
|
||||
**Rationale.** Item 5. F02 and M04 pass on state-based recovery, which the owner
|
||||
accepted. Memory's own retention is unproven, and in a long campaign whose facts
|
||||
are not all state-shaped it is the only continuity there is.
|
||||
|
||||
**Scope.** Diagnosis first, then the smallest sufficient fix.
|
||||
- **B.1, the diagnostic.** A retention harness that reports, for a planted fact,
|
||||
each stage with ids and depths:
|
||||
- **created:** does a memory whose `source_start`..`source_end` covers the
|
||||
planting turn contain the fact?
|
||||
- **retained:** is it not `forgotten`?
|
||||
- **ranked:** where is it in similarity order for the recall query?
|
||||
- **injected:** is it in the memories section?
|
||||
|
||||
The harness runs in two modes. The deterministic mode uses a scripted
|
||||
narrator, summariser and embedder. The real-model mode uses the reference
|
||||
narrator and embedder. It also adds a `recovered_through_memory_independent`
|
||||
verdict to `tools/m11_long_run.py`, keeping the existing verdicts and their
|
||||
meanings.
|
||||
- **B.2, the fix,** only at the stages B.1 shows failing, from this bounded menu:
|
||||
- how the summariser's excerpt is chosen, instead of keeping only the last
|
||||
2,000 tokens;
|
||||
- the memory prompt's instruction to keep named facts and objects;
|
||||
- an eviction rule that does not throw out a never-retrieved early memory
|
||||
first;
|
||||
- one additional ranking term (lexical or entity overlap, CONTEXT-AND-MEMORY
|
||||
§20).
|
||||
|
||||
Anything outside the menu needs the owner's agreement in the brief.
|
||||
|
||||
**Non-scope.**
|
||||
- Memory becoming authoritative, or outranking state (F07).
|
||||
- Any change to lineage attachment or filtering (`tree.attach_memory`,
|
||||
`forget_node`, E02) or summary lineage (E03).
|
||||
- Merging memory with imported knowledge, or new embedding models or
|
||||
dependencies.
|
||||
- Automatic re-summarisation of existing memories.
|
||||
- Redesigning cross-layer duplication (§22).
|
||||
|
||||
**Likely affected.** `backend/app/memorybank.py` (creation, eviction, ranking);
|
||||
`backend/app/context/builder.py` (memories section, only if the query changes);
|
||||
`backend/app/models.py`, only if a field is unavoidable. Tools: `tools/memory_ab.py`,
|
||||
`tools/m11_long_run.py`. Tests: `test_context_memory.py`, `test_memory_nodes.py`,
|
||||
`test_m11_leakage.py`, `test_m11_long_run_memory.py`.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. **Deterministic retention test**, committed. Fact F is planted at depth 3 or
|
||||
less through narration only, with no state correction and no knowledge source.
|
||||
These are asserted for the whole run:
|
||||
- F is never in authoritative state;
|
||||
- no turn after the planting block contains F's sentinel tokens;
|
||||
- the summary section never contains F;
|
||||
- the planting turn is outside the history window at recall.
|
||||
|
||||
At a recall depth of 100 or more, the memories section holds a memory that
|
||||
contains F and whose source range includes the planting depth, and the
|
||||
context report lists its id and similarity. **The test must fail on the
|
||||
`432f041` tree**, and the report must name the stage at which it fails.
|
||||
2. **Past capacity.** The same test runs with more memories written than
|
||||
`memory_bank_capacity`. The planting-era memory is still active and
|
||||
retrieved, or its eviction follows a documented, tested rule that the WP
|
||||
report justifies. No pinned memory is evicted. The "frozen bank" regression
|
||||
(a new memory evicted at once) still passes.
|
||||
3. **Lineage.** Fact G is planted only on a line later abandoned by Undo and
|
||||
divergence. G's memory stays on disk and is absent from every active-line
|
||||
prompt and from `memories.used`. All 14 tests in `test_m11_leakage.py` and the
|
||||
E02 tests pass unchanged.
|
||||
4. **Authority.** A memory contradicting state loses, and memory is still framed
|
||||
as non-canon (F07).
|
||||
5. **Real model.** A run of at least 100 turns at 16,384 with memory on, on the
|
||||
GPU host with logging. The fact is chosen so that the harness can verify the
|
||||
criterion-1 preconditions on real output. **Pass:** one run meets every
|
||||
precondition and returns `recovered_through_memory_independent`, with the
|
||||
memory's provenance. A run whose precondition fails reports which one, and
|
||||
counts as neither pass nor failure. The report gives the stage-by-stage
|
||||
diagnostic for every run.
|
||||
6. **Budget.** The memories section stays inside its budget (F03), and the
|
||||
steady-state prompt is no larger than A1's arithmetic allows.
|
||||
7. `CONTEXT-AND-MEMORY.md` §15, §20 and §21 are updated as implemented.
|
||||
|
||||
**Regression requirements.** F01-F08, E01-E04, M04 verdicts (the new verdict is
|
||||
added and the old ones are unchanged), the I-series (memories travel in the
|
||||
bundle), L04 (`test_memory_rewrite.py`), §O.7's no-write-lock and
|
||||
failure-recording tests.
|
||||
|
||||
**Dependencies.** After A2, so that it measures the prompt v1.1 ships and a
|
||||
narrator no longer handed protocol to restate. B.1's deterministic harness may
|
||||
be built earlier.
|
||||
|
||||
**Compatibility.**
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Databases | No migration preferred. If a memory field is unavoidable, it is additive with a default, tested from a real v1.0.0 database, with schema parity |
|
||||
| Bundles | Stay v3. Any new memory field is optional on import, and its absence means "unknown". v1.0.0 bundles import |
|
||||
| Existing memories | Valid and used as they are. A rebuild stays opt-in (`tools/rewrite_memories.py`) |
|
||||
| Branch history, Save Points, state, knowledge | None |
|
||||
| Settings | Any change to capacity semantics keeps v1 defaults |
|
||||
| Docker and local-only | None |
|
||||
|
||||
**Test modes.** Deterministic: yes. Real model: yes. Browser: no. Offline:
|
||||
regression only. Long run: yes, at least 100 turns, possibly several.
|
||||
|
||||
**Risk.** High uncertainty, because the outcome depends on the model, and medium
|
||||
implementation risk. The blast radius is the memory subsystem.
|
||||
|
||||
---
|
||||
|
||||
### WP-C — Browser release coverage
|
||||
|
||||
**Objective.** Every reader-facing behaviour that v1 proved only through the API
|
||||
or the component suite is driven in a real browser, including an export that
|
||||
actually leaves the browser as a file.
|
||||
|
||||
**Rationale.** Item 6. §T carries two qualifications that exist only because the
|
||||
harness stops short.
|
||||
|
||||
**Scope.** New scenarios in `tools/m11_browser.py`, or a sibling that reuses
|
||||
`tools/m11_webdriver.py`: Retry; Save Point create and restore; state correction;
|
||||
narration length; failed generation; export download. Also the Firefox profile
|
||||
preferences that direct a download to a harness-owned directory under `$HOME`.
|
||||
The package is harness-only, unless it finds a product defect or a control with
|
||||
no accessible name. Either is fixed with a regression test and reported as a
|
||||
product change.
|
||||
|
||||
**Non-scope.** New UI, frontend refactors, Selenium or any new dependency,
|
||||
screenshot diffing, tablet layout, CI.
|
||||
|
||||
**Likely affected.** `backend/tools/m11_browser.py`, `backend/tools/m11_webdriver.py`,
|
||||
the `DEVELOPMENT.md` harness section.
|
||||
|
||||
**Acceptance criteria.** The run ends with 0 failed and 0 skipped. It runs against
|
||||
the built SPA served by FastAPI, with turns from the reference narrator over
|
||||
trusted-LAN HTTPS. Every assertion reads the rendered DOM or a file on disk.
|
||||
1. **Retry.** Retry on the newest turn yields a second take, and the indicator
|
||||
reads 2/2. Stepping to 1/2 shows the original narration unchanged. The takes
|
||||
persist after a reload.
|
||||
2. **Save Point.** Create a named Save Point through the UI, play two more
|
||||
turns, then restore it through the UI. The transcript ends at the named
|
||||
moment, the position indicator says later story is ahead, and Redo walks
|
||||
into the later turns. After a reload the Save Point is still listed.
|
||||
3. **State correction.** A correction submitted through the State panel applies
|
||||
and is still shown after a reload. A partly refused correction shows its
|
||||
refusal and reason to the reader.
|
||||
4. **Narration length.** After changing the control to brief and playing a turn,
|
||||
then to long and playing a turn, each turn's context inspector shows its
|
||||
band's word range.
|
||||
5. **Failed generation.** With the model set through Settings to a name the
|
||||
server does not serve, a submitted turn shows an error, adds no narration and
|
||||
keeps the typed input. Setting the model back, the next turn succeeds and the
|
||||
earlier story is intact.
|
||||
6. **Export download.** A real click on Export, from both the campaign library
|
||||
and campaign settings, writes a file with no manual step. The file:
|
||||
- exists and is not empty;
|
||||
- parses as `ai-dnd-adventure-v3`;
|
||||
- imports into a fresh data directory with the same action count, head
|
||||
position and Save Points.
|
||||
|
||||
If the snap Firefox cannot be made to download, the check runs on a non-snap
|
||||
Firefox under `$HOME`, and `DEVELOPMENT.md` says so. **Exercising only the
|
||||
click does not pass.**
|
||||
7. The existing 38 checks pass in the same run.
|
||||
|
||||
**Regression requirements.** The existing browser checks; the frontend suite
|
||||
where a product fix is made.
|
||||
|
||||
**Dependencies.** None. It can run at any point. Its final run is repeated on the
|
||||
v1.1 release candidate.
|
||||
|
||||
**Compatibility.** None, unless a product fix is made, and then per §5.
|
||||
|
||||
**Test modes.** Browser: yes. Real model: yes, over the HTTPS reference host.
|
||||
Deterministic: no. Offline: no. Long run: no.
|
||||
|
||||
**Risk.** Low product risk. The medium risk is harness flakiness, so every wait is
|
||||
on a DOM condition, never a sleep used as an assertion.
|
||||
|
||||
---
|
||||
|
||||
### WP-D — Recovery honesty
|
||||
|
||||
**Objective.** A backup is kept only after a full integrity check. A campaign
|
||||
whose export exceeds what import accepts is exported with a warning saying so,
|
||||
not silently.
|
||||
|
||||
**Rationale.** Items 7 and 10a. Both are recovery gaps that are known, measured
|
||||
and cheap to close. Neither needs a policy decision.
|
||||
|
||||
**Scope.**
|
||||
- `backup.py` runs `PRAGMA integrity_check` on the finished copy, and the time
|
||||
it takes is measured.
|
||||
- Export compares the serialised bundle's size with
|
||||
`limits.MAX_IMPORT_BODY_BYTES`. When it is over, the file is still delivered
|
||||
and the reader sees a warning naming the limit and what it means.
|
||||
- The import refusal for an oversized bundle names the limit.
|
||||
- `DEVELOPMENT.md` documents the ceiling as measured: about 13 kB per action on
|
||||
a real campaign, and M9's conservative figure of about 279 turns.
|
||||
|
||||
**Non-scope.** Scheduled backups; raising the import limit or streaming import; a
|
||||
bundle format change or further compression; a restore button.
|
||||
|
||||
**Likely affected.** `backend/app/backup.py`, `backend/app/routers/adventures/bundle_io.py`,
|
||||
`frontend/src/pages/Campaigns.jsx`,
|
||||
`frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx`,
|
||||
`frontend/src/pages/backup.test.jsx`, and the backup and bundle tests.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. A backup of a healthy database reports `integrity_check` ok and is kept.
|
||||
2. A copy with damage that `integrity_check` detects and `quick_check` does not,
|
||||
such as an index inconsistent with its table, is rejected and not kept.
|
||||
Existing backups are still never overwritten.
|
||||
3. The time for `integrity_check` is recorded on the 100-turn evidence database
|
||||
and on a synthetic database of 100 MB or more, and the backup completes
|
||||
through the UI on both.
|
||||
4. Exporting a fixture campaign over 20 MB succeeds, delivers the file, and
|
||||
shows a warning naming the import limit, asserted in the API response and in
|
||||
a component test. A campaign under the limit shows no warning.
|
||||
5. Importing that bundle is refused with a message naming the limit.
|
||||
6. A normal campaign's export is byte-identical before and after the package,
|
||||
apart from timestamp fields.
|
||||
|
||||
**Regression requirements.** I01-I07, L01-L04, the backup case in
|
||||
`test_m11_migration.py`, `backup.test.jsx`, and the offline container's
|
||||
export and import.
|
||||
|
||||
**Dependencies.** None.
|
||||
|
||||
**Compatibility.** Databases and bundles: no change. Backups: a stricter check,
|
||||
with the same file.
|
||||
|
||||
**Test modes.** Deterministic: yes. Browser: the warning, optionally in WP-C's
|
||||
harness. Offline: regression only. Real model: no. Long run: no.
|
||||
|
||||
**Risk.** Low.
|
||||
|
||||
---
|
||||
|
||||
### WP-E — Control-boundary contrast
|
||||
|
||||
**Objective.** Control boundaries meet WCAG 1.4.11's 3:1 against their panel, at
|
||||
rest and on hover.
|
||||
|
||||
**Rationale.** Item 9.
|
||||
|
||||
**Scope.**
|
||||
- The boundary token values in `frontend/src/styles/tokens.css`, and any
|
||||
component that overrides them.
|
||||
- `tools/contrast_audit.py` treats boundary pairs below 3:1 as failures, not
|
||||
advisories.
|
||||
- The browser harness measures rendered boundary contrast.
|
||||
- Before-and-after screenshots for the owner's approval.
|
||||
|
||||
**Non-scope.** A palette redesign, typography, layout, tablet work, a
|
||||
screen-reader audit, and other WCAG criteria. A defect the same measurement finds
|
||||
is recorded, not taken on.
|
||||
|
||||
**Likely affected.** `frontend/src/styles/tokens.css`, `backend/tools/contrast_audit.py`,
|
||||
`backend/tools/m11_browser.py`, `frontend/src/a11y.test.jsx`.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and
|
||||
exits 0 on the package tree.
|
||||
2. Every text pair still clears 4.5:1 (1.4.3). The rendered text contrasts
|
||||
measured by the harness do not fall below v1's (14.57, 5.48, 13.57 and
|
||||
5.88:1) without a stated reason.
|
||||
3. The rendered boundary of the story input and of a primary control, at rest
|
||||
and on hover, is at least 3:1. The focus indicator is still visible.
|
||||
4. The owner approves the before-and-after screenshots, and the WP report records
|
||||
the approval.
|
||||
|
||||
**Regression requirements.** The frontend suite and lint; the browser harness's
|
||||
accessibility checks.
|
||||
|
||||
**Dependencies.** None. Its browser checks go into WP-C's harness if WP-C has
|
||||
landed, and into `m11_browser.py` otherwise.
|
||||
|
||||
**Compatibility.** None.
|
||||
|
||||
**Test modes.** Deterministic: yes. Browser: yes. Everything else: no.
|
||||
|
||||
**Risk.** Low.
|
||||
|
||||
---
|
||||
|
||||
## 9. Order and dependencies
|
||||
|
||||
```text
|
||||
A1 context safety reserve ──► A2 protocol echo + neutral instructions ──► B memory retention
|
||||
(B.1 harness may start early)
|
||||
C browser coverage ── independent
|
||||
D recovery honesty ── independent
|
||||
E boundary contrast ── independent (checks land in C's harness if C is first)
|
||||
│
|
||||
▼
|
||||
v1.1 release validation (§10)
|
||||
```
|
||||
|
||||
**Recommended order: A1, A2, B, C, D, E.**
|
||||
|
||||
- **A1 first.** It is the only item that can silently remove canon from a
|
||||
prompt. It is deterministic to test, small in blast radius, and it
|
||||
establishes the provenance that A2's and B's real-model runs will read.
|
||||
- **A2 second.** Its leak is silent and compounds through replay. It must
|
||||
precede B, because it changes what the memory pass summarises.
|
||||
- **B third.** It carries the highest continuity value and the most uncertainty.
|
||||
Measuring it before A1 and A2 settle would measure a prompt v1.1 does not ship.
|
||||
- **C, D and E** close verification, recovery and accessibility gaps with no
|
||||
known story risk. They depend on nothing, so the owner may move any of them
|
||||
earlier. While B waits on long runs is a natural slot. Each still has its own
|
||||
brief and review.
|
||||
|
||||
| Property | A1 | A2 | B | C | D | E |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| Can begin independently | yes | after A1 | after A2 | yes | yes | yes |
|
||||
| Schema or migration | no | no | avoid; additive if unavoidable | no | no | no |
|
||||
| Export format | no (additive evidence key) | no | no; optional field at most | no | no | no |
|
||||
| Changes acceptance tests (`V1-ACCEPTANCE-TESTS.md`) | no | no | no | no | no | no |
|
||||
| Changes the stored prompt | yes | yes | possibly | no | no | no |
|
||||
| Real-model validation | yes | yes | yes | yes, for turns | no | no |
|
||||
| Browser testing | regression | regression | no | yes | optional | yes |
|
||||
| Offline / no-network testing | regression | regression | regression | no | regression | no |
|
||||
| Long-run testing | yes, 50 turns or more | yes, 50 turns or more | yes, 100 turns or more | no | no | no |
|
||||
| Risk | low | medium | high uncertainty | low | low | low |
|
||||
|
||||
## 10. Compatibility summary
|
||||
|
||||
| Area | A1 | A2 | B | C | D | E |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| v1.0.0 campaign databases | open unchanged | unchanged | unchanged; an additive migration only if unavoidable | — | unchanged | — |
|
||||
| Export and import bundles | v3; new evidence key | — | v3; an optional field at most | — | v3; export warns | — |
|
||||
| Save Points | — | — | — | — | — | — |
|
||||
| Branch history | — | — | lineage rules unchanged | — | — | — |
|
||||
| Narrative state | — | validator unchanged | memory never outranks state | — | — | — |
|
||||
| Memories and summaries | — | new ones from cleaner prose | creation, retention and ranking change; existing rows kept | — | — | — |
|
||||
| Knowledge sources | — | — | — | — | — | — |
|
||||
| Model settings | none preferred | — | capacity defaults kept | — | — | — |
|
||||
| Docker and local-only | — | — | — | — | — | — |
|
||||
|
||||
## 11. v1.1 release criteria
|
||||
|
||||
v1.1 is not called v1.1.0 until all of the following hold on one release
|
||||
candidate tree:
|
||||
|
||||
1. **Every package in scope is accepted**, each with its report, and with its
|
||||
acceptance criteria passing on the candidate or on a tree whose product code
|
||||
the candidate carries unchanged.
|
||||
2. **The v1 contract still passes.** All 82 REQUIRED FOR V1 tests hold, with H09
|
||||
not applicable on the same condition. None is relaxed.
|
||||
3. **Suites and builds:** the backend suite, the frontend suite and lint, the
|
||||
production build, and `docker build --no-cache`, with the image's SPA
|
||||
identical to the local build.
|
||||
4. **Offline:** `tools/m11_offline.py` passes all checks with no network and a
|
||||
fresh volume.
|
||||
5. **Browser:** the existing 38 checks plus WP-C's and WP-E's pass with 0 failed
|
||||
and 0 skipped, on the candidate, over trusted-LAN HTTPS.
|
||||
6. **One v1.1 long run** on the candidate's product code: 100 turns or more at a
|
||||
16,384 window with memory on, on the GPU host with logging. It passes M01-M04.
|
||||
A1's headroom table shows the documented reserve on every re-counted prompt,
|
||||
and every turn is `fits`. A2's leak count is 0. B's independent-retention
|
||||
verdict is recorded.
|
||||
7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0
|
||||
protocol shapes in stored narration.
|
||||
8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all
|
||||
checks.
|
||||
9. **Upgrade from a real v1.0.0 database.** A database is created by the
|
||||
`v1.0.0` tree and played. It has retained history, an undone head, Save
|
||||
Points, memories, summaries, imported knowledge and a narration-length
|
||||
choice, and uses loopback or placeholder settings with no real hostnames.
|
||||
Opened by the candidate, its transcript, head, Redo availability, state, Save
|
||||
Points, memories, knowledge and settings compare identical, and any migration
|
||||
is forward-only with schema parity. A v1.0.0 export imports into v1.1.
|
||||
Because the format stays v3, a v1.1 export of that campaign is checked for
|
||||
import into v1.0.0, and the result is reported.
|
||||
10. **Documentation:** `README.md`, `DEVELOPMENT.md`, the as-implemented sections
|
||||
named by each package, and `VERSION.md`.
|
||||
11. **Owner events**, each separate: the signed release commit, `main`, and a
|
||||
`v1.1.0` tag.
|
||||
|
||||
## 12. Scope recommendation
|
||||
|
||||
**Recommended: Option 1, a focused v1.1.**
|
||||
|
||||
| Ship in v1.1 | Defer to v1.2 | Future / optional |
|
||||
| --- | --- | --- |
|
||||
| WP-A1 context safety reserve | Scheduled backups (10b) | Real media-provider adapters (11; K05, K06) |
|
||||
| WP-A2 protocol echo and neutral instructions | Raising or streaming the import limit (7) | Whole-transcript copy, story search (§77, §78) |
|
||||
| WP-B independent memory retention | Reducing in-sentence restatement (2) | Discarded-history recovery screen (§63) |
|
||||
| WP-C browser release coverage | Stale-scene and derived-contamination detectors in the identity diagnostic (8) | Tablet layout |
|
||||
| WP-D recovery honesty | Cross-layer duplication suppression (17) | Editable RPG world state |
|
||||
| WP-E control-boundary contrast | | Media debt: `ambience`, visual-profile recovery and UI (20) |
|
||||
| | | Trusted-LAN residual limits: address pinning against DNS rebinding (21) |
|
||||
| | | Comparative narrator-model recommendation (13) |
|
||||
|
||||
**Why focused.** Four of the six packages close silent failure modes or
|
||||
verification gaps that the v1 evidence itself named. The other two are small and
|
||||
deterministic. Everything deferred is either a new feature, or needs a policy
|
||||
decision the owner has not been asked for, or rests on an occurrence that has not
|
||||
happened. A broader v1.1 that took scheduled backups and media would add the two
|
||||
packages with the most new surface. It would also push the long-run
|
||||
re-validation, which every prompt change already requires, further from the
|
||||
changes it validates.
|
||||
|
||||
**If B's real-model criterion cannot be met.** Precondition-valid runs may recover
|
||||
nothing even after B.2. In that case the owner chooses between shipping v1.1 with
|
||||
B's deterministic criteria met and the real-model result recorded as a residual
|
||||
risk, or holding v1.1 for B. The plan does not pre-decide this.
|
||||
|
||||
## 13. Coding briefs needed next
|
||||
|
||||
These briefs are not written here, and none is to be executed from this document.
|
||||
|
||||
1. **WP-A1 — context-window safety reserve. Write this one first.** It carries
|
||||
the highest user risk, has no dependencies, is deterministically testable,
|
||||
and produces the provenance that later real-model runs read. The brief must
|
||||
settle three things: the tokenizer divergence the reserve is sized for, whether
|
||||
calibration is built or detection alone, and whether any setting is added
|
||||
(recommended: none).
|
||||
2. WP-A2 — protocol echo hardening and genre-neutral instructions. It can be
|
||||
drafted while A1 is under review.
|
||||
3. WP-B — independent memory retention: B.1 diagnostic, then B.2 fix. Two briefs
|
||||
are an option if B.1's findings should be reviewed before a fix is chosen.
|
||||
4. WP-C — browser release coverage.
|
||||
5. WP-D — recovery honesty.
|
||||
6. WP-E — control-boundary contrast.
|
||||
7. v1.1 release validation, written only after every in-scope package is
|
||||
accepted.
|
||||
|
||||
## 14. Decisions for the owner
|
||||
|
||||
- Confirm Option 1, the focused v1.1.
|
||||
- A1: the divergence the reserve must absorb; calibration or detection only; that
|
||||
a detected truncation is flagged, not turned into a failed turn.
|
||||
- B: whether B.1 and B.2 are one brief or two; the choice in §12 if the
|
||||
real-model criterion is not met.
|
||||
- D: that the 20 MB import limit stays in v1.1.
|
||||
- E: approval of the visual change.
|
||||
- Whether v1.1 real-model validation stays on the reference narrator, as this plan
|
||||
recommends.
|
||||
- The report naming and rotation in `planning/reports/` (§5 rule 5).
|
||||
+38
-2
@@ -1,8 +1,44 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v4.0
|
||||
- **Package:** Adventure Storyteller Planning Package v4.1
|
||||
- **Revision date:** 2026-09-14
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M11 complete**; **M11 accepted at its closeout (2026-09-14), and the v1 release gate passed** on the release-candidate tree. All 82 REQUIRED FOR V1 tests pass, with H09 not applicable. There is no further planned milestone. The signed closeout commit and any `v1.0.0` tag are the repository owner's, and neither exists as of this revision.
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 planning has begun** on `v1.1-development`: `V1.1-PLAN.md`. No v1.1 work package has started, and no v1.1 version or tag exists.
|
||||
|
||||
## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14)
|
||||
|
||||
Documentation only. No product code, no requirement, no acceptance test and no
|
||||
schema changed.
|
||||
|
||||
**Part 1 — post-release correction.** v4.0 was written before three owner
|
||||
events that have since happened: the closeout commit was signed (`432f041`),
|
||||
`main` was fast-forwarded to it, and the signed tag `v1.0.0` was created on it
|
||||
and pushed. The current-state wording that said otherwise is corrected. The M11
|
||||
report is not edited: its §T records the state at closeout, which was true when
|
||||
written.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `README.md` | Status: v1.0.0 released; tag and `main` at `432f041`; v1.1 on `v1.1-development`. | developer docs |
|
||||
| `planning/README.md` | Current state, the release events, the map (v1.0.0 released, v1.1 planning), the stop rule restated for work packages, `V1.1-PLAN.md` in the document table and reading order. | index |
|
||||
| `planning/BUILD-MILESTONES.md` | Header status, **stale since M8** ("M9 is next"), now says complete and closed. A dated post-release note under M11's status, leaving the closeout paragraph as it stood. The *Post-v1 backlog* points at its triage. | milestone status |
|
||||
| `planning/VERSION.md` | This header and entry. | package version |
|
||||
|
||||
**Part 2 — the v1.1 plan.**
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `planning/V1.1-PLAN.md` | **New.** Every post-v1 backlog item and every other recorded v1 residual risk, triaged. Six v1.1 work packages in order (A1, A2, B, C, D, E), each with objective, rationale, scope, non-scope, affected subsystems, acceptance criteria, regression requirements, dependencies, compatibility and risk. The v1.1 release criteria. A v1.2 list and future work. | planning |
|
||||
|
||||
**Recommended scope: a focused v1.1.** Context-window safety, protocol-echo
|
||||
hardening, independent memory retention, browser coverage, recovery honesty and
|
||||
control-boundary contrast. Scheduled backups, identity-diagnostic extensions and
|
||||
media adapters are deferred, with the reason for each.
|
||||
|
||||
**First brief to write:** WP-A1, the context-window safety reserve.
|
||||
|
||||
**Requirement changes: zero.** Every v1.1 package improves the implementation of
|
||||
an existing requirement, so `SPECIFICATION.md` and `V1-ACCEPTANCE-TESTS.md` are
|
||||
unchanged.
|
||||
|
||||
## v4.0 — M11 closeout: v1 release validation accepted (2026-09-14)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user