M11 closeout: accept v1 release validation
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s

The browser, offline and identity runs had last been taken on ef25b0a. The
closeout repeated them on the exact release-candidate tree, 3652dc6, whose
product code is identical to 96c1bf5, where the 100-turn evidence was run. No
product code changed, so the long-run evidence stands.

On 3652dc6:
- backend suite: 1421 passed, 17 skipped, 0 failed
- frontend suite: 161 of 161; lint clean
- production build clean
- docker build --no-cache: image SPA byte-identical to the local build
- browser regression: 38 of 38
- no-network container: 23 of 23
- identity diagnostic: 0 signals; the scripted self-test's detectors fire
- release-shaped smoke test from the image: 14 of 14

- M11 report: new S (exact-tree verification, including the identity
  results the report never carried) and T (acceptance record). Corrections:
  the REQUIRED FOR V1 count is 82, not 85, and L's browser narrator was
  qwen2.5:3b-instruct. P gains risks 16 and 17; risk 6 is widened.
- BUILD-MILESTONES.md: M11 COMPLETE / ACCEPTED, and a post-v1 backlog.
- V1-ACCEPTANCE-TESTS.md: the P release gate's result, and the P3
  disposition's run.
- planning/README.md, VERSION.md (v4.0), README.md: status, map, stop rule.
- tools/m11_browser.py: the G01 import wait could not fail, because the
  scenario's campaign is titled "Hidden Knowledge". It now waits for the
  imported source's row.

Found and carried, not fixed. The identity run stored protocol shapes the
extractor leaves, on 4 of 10 turns at a 4,096 window: event-call syntax and a
parroted length hint. The owner chose residual risk. The state rule's example
is fantasy, and the state lagged the narration. None occurs in the 100-turn
evidence.

No requirement changes. No release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CBTT3qbkGoemWD7BRvvxpT
This commit is contained in:
JesseMarkowitz
2026-09-14 06:07:16 -04:00
co-authored by Claude Opus 5
parent 3652dc6fae
commit 432f04100b
7 changed files with 540 additions and 91 deletions
+54 -7
View File
@@ -1449,12 +1449,25 @@ Release gate in `V1-ACCEPTANCE-TESTS.md`:
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07; long-run evidence complete 2026-09-13; awaiting independent review/acceptance
## Status: COMPLETE / ACCEPTED — 2026-09-14 (implemented 2026-09-07; long-run evidence 2026-09-13; release-candidate closeout 2026-09-14)
Implemented on `m11-release-validation` from the signed M10 commit `1013c94`.
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package, written
for a release reviewer. **M11 is not marked accepted here**; that is the
reviewer's to record, and no release tag exists.
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package.
**M11 is accepted, and v1 release validation is complete.** The closeout on
2026-09-14 re-verified the exact release-candidate tree (`3652dc6`, whose product
code is identical to `96c1bf5`, where the 100-turn evidence was run). The backend
suite passed 1,421 with 17 skipped and 0 failed. The frontend suite passed 161 of
161, lint was clean, and the production build succeeded. A `--no-cache` Docker
image built. The browser regression passed 38 of 38, the no-network container
passed 23 of 23, and the identity diagnostic ran clean, with its results now in
the report (§S). Every REQUIRED FOR V1 test passes, 82 in all with H09 not
applicable (report §T). The acceptance takes effect with the owner's signed
closeout commit. **No release tag exists**: tagging `v1.0.0` is a separate
decision, and the tag must point at that signed commit.
**M1-M11 are all complete. There is no M12.** Post-v1 work is backlog, listed
below under *Post-v1 backlog*, and none of it is an unfinished v1 milestone.
**The release blocker it was given, and how it was closed.** M8 measured the
reference deployment enforcing a **4,096**-token input window while the
@@ -1520,9 +1533,43 @@ switched on:
extractor now removes every shape observed. Replaying 443 real turns through it
changed no turn it had previously left clean. (`0c7316f`, `96c1bf5`)
**Left for the reviewer**, in the report's §P: the real-token headroom at the
largest prompts is 23-42 tokens; the narrator restates prompt text in its prose;
and the identity diagnostic's results are not in the report.
**Carried past v1 as residual risk**, in the report's §P:
- the real-token headroom at the largest prompts is 23-42 tokens;
- the narrator restates prompt text in its prose;
- stored narration can keep protocol shapes the extractor does not remove.
The closeout found event-call syntax and a parroted length hint, and the owner
chose to carry this rather than change code (report §S.6).
The identity diagnostic's results are now in the report, §S.5.
---
# Post-v1 backlog — not milestones
Work recorded for after v1. None of it is a v1 requirement or an unfinished v1
milestone, and none of it has a brief. Each item needs one before work begins.
Sources are the M11 report's §P and §S.6.
- **Context-window safety margin.** The largest prompts leave 23-42 real tokens,
and Ollama cuts an over-window prompt with no error. Consider a deliberate
reserve, or counting with the narrator's own tokenizer.
- **The narrator restating prompt and state text** in its prose, including the
protocol shapes the extractor still leaves: event-call syntax, a parroted
length hint, and a lone section heading.
- **A genre-neutral example in the state rule**, in place of `silver-key` and
`aldric`.
- **Independent long-term-memory retention.** No run showed memory keeping a
planted fact without authoritative state.
- **Browser automation of the export download**, and real-browser coverage of
Retry, Save Point, state correction, narration length and failed generation.
- **A practical bundle-size ceiling.** M9 estimated about 279 turns against the
20 MB import limit.
- **Identity follow-up on the next real occurrence.** Classify it with
`tools/m11_identity.py`. Consider a stale-scene check, and a run with memory on
at a full window.
- **WCAG 1.4.11 control-boundary contrast** (1.33:1 resting, 1.75:1 hover).
- **Backup hardening:** `integrity_check`, and scheduled backups.
- **Real media-provider adapters** against `MEDIA-EXTENSION-CONTRACT.md`.
---