M11 closeout: accept v1 release validation
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s

The browser, offline and identity runs had last been taken on ef25b0a. The
closeout repeated them on the exact release-candidate tree, 3652dc6, whose
product code is identical to 96c1bf5, where the 100-turn evidence was run. No
product code changed, so the long-run evidence stands.

On 3652dc6:
- backend suite: 1421 passed, 17 skipped, 0 failed
- frontend suite: 161 of 161; lint clean
- production build clean
- docker build --no-cache: image SPA byte-identical to the local build
- browser regression: 38 of 38
- no-network container: 23 of 23
- identity diagnostic: 0 signals; the scripted self-test's detectors fire
- release-shaped smoke test from the image: 14 of 14

- M11 report: new S (exact-tree verification, including the identity
  results the report never carried) and T (acceptance record). Corrections:
  the REQUIRED FOR V1 count is 82, not 85, and L's browser narrator was
  qwen2.5:3b-instruct. P gains risks 16 and 17; risk 6 is widened.
- BUILD-MILESTONES.md: M11 COMPLETE / ACCEPTED, and a post-v1 backlog.
- V1-ACCEPTANCE-TESTS.md: the P release gate's result, and the P3
  disposition's run.
- planning/README.md, VERSION.md (v4.0), README.md: status, map, stop rule.
- tools/m11_browser.py: the G01 import wait could not fail, because the
  scenario's campaign is titled "Hidden Knowledge". It now waits for the
  imported source's row.

Found and carried, not fixed. The identity run stored protocol shapes the
extractor leaves, on 4 of 10 turns at a 4,096 window: event-call syntax and a
parroted length hint. The owner chose residual risk. The state rule's example
is fantasy, and the state lagged the narration. None occurs in the 100-turn
evidence.

No requirement changes. No release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CBTT3qbkGoemWD7BRvvxpT
This commit is contained in:
JesseMarkowitz
2026-09-14 06:07:16 -04:00
co-authored by Claude Opus 5
parent 3652dc6fae
commit 432f04100b
7 changed files with 540 additions and 91 deletions
+40 -5
View File
@@ -2555,6 +2555,36 @@ The release candidate should not be called v1.0 until:
- branch/memory lineage isolation passes,
- fantasy and science-fiction fixtures both pass.
### Result — gate passed on the release candidate (M11 closeout, 2026-09-14)
Measured on the release-candidate tree `3652dc6`, whose product code is identical
to `96c1bf5`. The per-test matrix is M11 report §F. The closeout's exact-tree
verification is §S, and the acceptance record is §T.
- **All REQUIRED FOR V1 tests pass.** There are 82. 81 pass, and H09 is NOT
APPLICABLE on the condition its own text states: the product extracts no
archives, and a test enforces that.
- **Approved exceptions:** none. No REQUIRED test was waived, relaxed or
reclassified.
- **Security offline test:** 23 of 23, in a container with no network and a
fresh volume, built with `--no-cache` from the candidate tree.
- **100-turn long run:** M01-M04 pass on `96c1bf5` (results under M01-M04
above). The closeout changed no product code, so that evidence stands.
- **Export/import recovery, including an undone head:** I01-I07. The 100-turn
campaign moved to a clean data directory, 16 of 16.
- **Undo/Redo/Retry/checkpoint:** D01-D14, in the long run and in the browser.
- **Narrative-state events at realistic context length:** C06, M01.
- **No first-use runtime download:** H11, in the offline container.
- **Branch and memory lineage isolation:** E01-E04.
- **Fantasy and science-fiction fixtures:** both pass.
**Qualifications that stand:**
- M04 recovered its fact through authoritative state, not independent memory
retention. The owner accepted that on 2026-09-13.
- K04, a SHOULD test, passes on its deferred branch.
- The browser's export *download* is exercised only as far as the click.
## Q. Current Recommendation
Use this document as:
@@ -2700,8 +2730,13 @@ the run reports it), which is the control this kind of tool most often lacks.
**This remains a test-design task and is still not an acceptance test.** The
model-quality half is not a pass/fail property of the application, and M11 does
not make it one. **The M11 report does not record the diagnostic's run
results.** This was corrected on 2026-09-14: the section the report pointed to
was left empty. The report records the fixture defect the diagnostic caught in
itself, and the run's evidence is in the implementer's
`m11-evidence/identity-recheck` directory.
not make it one.
**Run at the M11 closeout (2026-09-14), on the release-candidate tree `3652dc6`.**
The narrator was `qwen2.5:3b-instruct` over trusted-LAN HTTPS, at a verified
4,096 window. The fixture was accepted with nothing refused. 10 of 10 turns were
accepted with **0 signals**, and the scripted self-test's detectors fired. The
state lagged the narration, though: one proposal in ten applied, and a stale
scene is outside the objective checks. The run therefore reproduces nothing
about the original finding and establishes no root cause. M11 report §S.5 has
the results and what they cannot establish.