M11 closeout: accept v1 release validation
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s

The browser, offline and identity runs had last been taken on ef25b0a. The
closeout repeated them on the exact release-candidate tree, 3652dc6, whose
product code is identical to 96c1bf5, where the 100-turn evidence was run. No
product code changed, so the long-run evidence stands.

On 3652dc6:
- backend suite: 1421 passed, 17 skipped, 0 failed
- frontend suite: 161 of 161; lint clean
- production build clean
- docker build --no-cache: image SPA byte-identical to the local build
- browser regression: 38 of 38
- no-network container: 23 of 23
- identity diagnostic: 0 signals; the scripted self-test's detectors fire
- release-shaped smoke test from the image: 14 of 14

- M11 report: new S (exact-tree verification, including the identity
  results the report never carried) and T (acceptance record). Corrections:
  the REQUIRED FOR V1 count is 82, not 85, and L's browser narrator was
  qwen2.5:3b-instruct. P gains risks 16 and 17; risk 6 is widened.
- BUILD-MILESTONES.md: M11 COMPLETE / ACCEPTED, and a post-v1 backlog.
- V1-ACCEPTANCE-TESTS.md: the P release gate's result, and the P3
  disposition's run.
- planning/README.md, VERSION.md (v4.0), README.md: status, map, stop rule.
- tools/m11_browser.py: the G01 import wait could not fail, because the
  scenario's campaign is titled "Hidden Knowledge". It now waits for the
  imported source's row.

Found and carried, not fixed. The identity run stored protocol shapes the
extractor leaves, on 4 of 10 turns at a 4,096 window: event-call syntax and a
parroted length hint. The owner chose residual risk. The state rule's example
is fantasy, and the state lagged the narration. None occurs in the 100-turn
evidence.

No requirement changes. No release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CBTT3qbkGoemWD7BRvvxpT
This commit is contained in:
JesseMarkowitz
2026-09-14 06:07:16 -04:00
co-authored by Claude Opus 5
parent 3652dc6fae
commit 432f04100b
7 changed files with 540 additions and 91 deletions
+348 -40
View File
@@ -4,15 +4,21 @@
Branch `m11-release-validation`, from the signed M10 commit `1013c94`.
Implemented and verified 2026-09-07. The long-run evidence was re-established,
and this revision written, on 2026-09-14. This is the evidence package; it is
not a record of acceptance, and nothing in it says M11 is accepted.
and this revision written, on 2026-09-14.
**Closeout, 2026-09-14.** The browser, offline and identity runs were repeated on
the exact release-candidate tree (§S), and M11 was accepted (§T). Sections A-R
are the implementer's evidence as revised before the closeout, corrected where
the closeout found them wrong. §S and §T are the closeout.
---
## A. Executive result
**PASS, for independent review.** All 85 tests marked REQUIRED FOR V1 pass, with
H09 recorded NOT APPLICABLE on the condition its own text states.
**PASS, and accepted at closeout (§T).** All 82 tests marked REQUIRED FOR V1
pass, with H09 recorded NOT APPLICABLE on the condition its own text states.
*Corrected at closeout:* earlier revisions said 85. That figure counted the four
SHOULD rows in §F and left H09 out.
This revision supersedes the first. That revision left M01 PARTIAL at 41
accepted turns, and then the run it described was lost to a host crash along
@@ -55,10 +61,10 @@ narration-length setting moves a number. Post-M8 findings A and B are closed.
- **Performance.** Timings are two particular hosts', not a product
characteristic (§E.1).
- **Post-M8 finding D's root cause**, which remains unestablished. The identity
diagnostic's run results are not in this report (§D.1).
- **That the browser, offline and identity runs cover the later commits.** They
were re-run on the `ef25b0a` tree. The three product commits since then change
the turn commit, memory retrieval and narration extraction (§B).
diagnostic's closeout run, and what it can and cannot establish, are in §S.5.
- *Closed at closeout:* the browser, offline and identity runs had predated the
last three product commits. All three were repeated on the release-candidate
tree (§S).
**Zero requirement weakenings.** No acceptance test was retired, relaxed or
reclassified. The M04 precondition the harness measures was corrected to the
@@ -72,7 +78,7 @@ acceptance text's own wording, "without entire transcript in prompt" (§O).
| **Base commit** | `1013c94eb1ad283e960114aef04c19c2806b5db7` — *"M10: the seam for media, and no media"* |
| **Signature** | `git verify-commit 1013c94` → **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, "JesseMarkowitz", trust `[ultimate]`. `%G?` = `G`. |
| **Branch** | `m11-release-validation`, created from that commit. The working tree was clean at the start (`git status --porcelain` empty). |
| **HEAD at this revision** | `96c1bf5`. Seven commits follow the base, all signed by the owner (`%G?` = `G`): `144406c` (M11), `fedb714`, `ef25b0a`, `fec46f6`, `f8d4010`, `0c7316f`, `96c1bf5`. This revision of the report is staged for the owner to sign. |
| **HEAD at this revision** | `96c1bf5`. Seven commits follow the base, all signed by the owner (`%G?` = `G`): `144406c` (M11), `fedb714`, `ef25b0a`, `fec46f6`, `f8d4010`, `0c7316f`, `96c1bf5`. This revision of the report is staged for the owner to sign. **At closeout** the candidate is `3652dc6`. `d198806` and `3652dc6` follow `96c1bf5`, both are signed, and both change documentation only: no file under `backend/`, `frontend/`, `Dockerfile`, `docker-compose.yml` or the start scripts differs from `96c1bf5` (§S.1). |
| **Upstream ancestry** | `git merge-base --is-ancestor d72f7c1b HEAD` → true. The AI-DnD fork point is still an ancestor. |
| **LICENSE** | Unchanged — md5 `07fde30437134836e2ee875e82a7cd31`, MIT, "Copyright (c) 2026 Parth Thakkar". `PROVENANCE.md` unchanged. |
@@ -94,7 +100,9 @@ the offline container, the identity diagnostic and the contrast audit** were
re-run on the `ef25b0a` tree, as that commit's message records, and their
evidence is dated 2026-09-10. The three product commits after it (`f8d4010`,
`0c7316f`, `96c1bf5`) change backend files only, and those runs were not
repeated (§P). Runs taken before a product change on their own tree were
repeated at the time. **At closeout all three were repeated on the
release-candidate tree `3652dc6`** (§S). No reported black-box result now comes
from product code other than the candidate's. Runs taken before a product change on their own tree were
discarded and re-run; §O records which and why.
---
@@ -222,11 +230,11 @@ turn's state. Before M11 all three settings produced *the identical sentence*;
`test_the_three_lengths_no_longer_say_the_same_thing` fails against that.
**D — character identity confusion.** The diagnostic exists; the root cause does
not, and cannot. **The diagnostic's run results are not reproduced in this
report.** The first revision pointed here to a section that was left empty.
What is recorded is the defect the diagnostic caught in its own fixture (§O.2),
which is why its first run's evidence was discarded. The later run's evidence is
in the implementer's `m11-evidence/identity-recheck` directory (§P).
not, and cannot. The diagnostic caught a defect in its own fixture (§O.2), and
that is why its first run's evidence was discarded. Earlier revisions of this
report reproduced no run's results. **The closeout run's results are in §S.5**:
clean on every objective check, with the detectors proved to fire, and with what
the run cannot establish stated.
---
## E. The context-window resolution
@@ -442,7 +450,7 @@ evidence, it says so and the verdict is qualified.
| A03 No cloud API key | **PASS** | suite: no `api_key` in settings, none settable through the API, no Authorization header; container: no secret in an export |
| A04 Campaign survives restart | **PASS** | campaign: 3 genuine process restarts (4 process starts), with transcript/head/state/Save Points/knowledge/settings compared before and after each and identical every time (§G.2); container: campaigns survive a container restart |
| A05 Failed model call does not corrupt story | **PASS** | campaign: a real failed call at a controlled point (the model name pointed at one the server does not serve), then play resumed; container: same with no model reachable at all — no narration accepted, state unchanged, the player's words kept, earlier story reachable |
| A06 Trusted-LAN Ollama inference | **PASS** | browser + identity diagnostic + the CPU-host campaigns: Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1) |
| A06 Trusted-LAN Ollama inference | **PASS** | browser + identity diagnostic + the CPU-host campaigns: Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1). **Re-established at closeout on the candidate tree** by the browser regression and the identity diagnostic, on the same HTTPS host (§S) |
### B — Core play
@@ -509,7 +517,7 @@ evidence, it says so and the verdict is qualified.
| ID | Result | Evidence |
| --- | --- | --- |
| G01 Import local text | **PASS** | browser: through the real file input; container: offline |
| G01 Import local text | **PASS** | browser: through the real file input. The harness's own assertion could not fail until the closeout fixed it (§S.3); every run's database holds the imported source. Container: offline |
| G02 Import local Markdown | **PASS** | campaign: three sources imported (canon, reference, inspiration); browser |
| G03 Classification | **PASS** | campaign: all three classes present after the move (§K); suite |
| G04 Disable knowledge source | **PASS** | suite `test_change_visibility.py`, `test_imported_knowledge.py` |
@@ -949,8 +957,11 @@ rather than about M7 once M11 added a migration (§O).
**Firefox 154.0.1**, headless, driven over W3C WebDriver (geckodriver 0.37.1),
against the **built** SPA served by FastAPI — the production path from
`DEVELOPMENT.md`, not a Vite dev server. Narrator: `qwen2.5:3b-instruct-16k` on
the trusted-LAN Ollama. **38 checks, 0 failed, 0 skipped**, 107 seconds.
`DEVELOPMENT.md`, not a Vite dev server. Narrator: `qwen2.5:3b-instruct` on
the trusted-LAN Ollama. *Corrected at closeout:* earlier revisions named the
`-16k` model, and the run's own report records the plain one. **38 checks, 0
failed, 0 skipped**, 107 seconds, on the `ef25b0a` tree. The closeout's runs on
the candidate tree are in §S.3.
The harness is `tools/m11_browser.py` and its WebDriver client is
`tools/m11_webdriver.py`, both in the repository — M8's and M9's browser harness
@@ -1365,6 +1376,7 @@ defects. A harness that has only ever agreed with itself is not evidence.
| **It had no way to see failed post-turn work** *(fixed in `f8d4010`)* | it reported `complete` over 180 `database is locked` errors | it would have certified the run that §O.7 destroyed. It now checks derived status and new `server.log` lines after every turn and stops at the first failure; a run with no memories or no summaries ends `failed` |
| **Its protocol-leak count used plain substrings** *(fixed in `96c1bf5`)* | `## Established:` did not match `\nEstablished:\n`, so the count read 1 where 5 turns leaked | the count now matches any state heading with an indented entry, in any markdown |
| **Its M04 precondition was the clue's text being out of recent history** *(fixed in `96c1bf5`)* | the narrator reuses that text in its own prose, so a run whose planting turn was 65 depths outside the window read `precondition_not_met` | the precondition is now the planting turn's position, recorded at planting and carried across `--resume`, which is M04's own wording |
| **The browser import wait could not fail** *(fixed at closeout)* | it waited for "hidden" in the page text, and the scenario's campaign is titled *Hidden Knowledge*, so the wait held before Import was pressed. The modal check that follows then raced the source row, and skipped once (34/0/1) | "a local file imports through the browser" was recorded without evidence. The harness now waits for the imported source's own row and records its title (§S.3) |
### Discarded evidence runs
@@ -1437,7 +1449,16 @@ accepted for v1, or a decision for the owner at acceptance.
6. **Model-invented headings are not removed.** A model that writes its own
section, such as `Identifiers established:` or `Set of events made true:`,
under a heading that is not the renderer's, keeps it in the story. Seen on 10
turns of one 26-turn trial.
turns of one 26-turn trial. **Widened at closeout.** The identity run found
three more shapes left in stored story, on 4 of 10 turns at a 4,096 window:
- event-call syntax copied from the state rule, such as
`> set_possession(silver-key, "alice")`
- a parroted length hint, `[Hard limit: … append the state block well inside
the limit.]`
- a lone `Scene:` line
None occurs in the 100-turn evidence runs. The owner chose on 2026-09-14 to
carry this as a residual risk rather than change product code (§S.6).
7. **The GPU inference host dropped its GPU after the evidence run** (§E.1). The
cause is not established, and power transients at the uncapped 280 W limit
@@ -1447,17 +1468,15 @@ accepted for v1, or a decision for the owner at acceptance.
8. **Post-M8 finding D's root cause is unestablished and will stay that way.**
The campaign that produced it was destroyed. M11 delivers a diagnostic that
can classify the next occurrence, and the detection the finding asked for.
**The diagnostic's run results are not in this report.** The section the
first revision pointed to was left empty, and the evidence is in the
implementer's `m11-evidence/identity-recheck` directory.
Its closeout run's results, and what that run still cannot establish, are
in §S.5.
9. **The browser, offline and identity runs predate the last three product
commits.** They were re-run on the `ef25b0a` tree: that commit's message
records browser 38/0/0, offline 23/0 and the identity diagnostic clean, and the
evidence is dated 2026-09-10. `f8d4010`, `0c7316f` and `96c1bf5` change the
turn commit, memory retrieval, narration extraction and the state renderer's
headings, all of them backend. The backend and frontend suites cover those
changes; the black-box runs were not repeated.
9. *Closed at closeout.* **The browser, offline and identity runs predated the
last three product commits.** They had been re-run on the `ef25b0a` tree:
browser 38/0/0, offline 23/0, and the identity diagnostic clean, with the
evidence dated 2026-09-10. `f8d4010`, `0c7316f` and `96c1bf5` changed backend
behaviour. All three runs were repeated on the release-candidate tree on
2026-09-14 (§S).
10. **Two control-boundary colour pairs are below WCAG 1.4.11** (1.33:1 resting,
1.75:1 hover). Reported rather than fixed, because the control is identified
@@ -1487,6 +1506,19 @@ accepted for v1, or a decision for the owner at acceptance.
existing campaign does not benefit from finding C's fix until someone sets
the control.
16. **The state rule's example is fantasy.** The fixed instruction every campaign
receives names `mara`, `silver-key`, `old-abbey` and `aldric`. In the
closeout's office-meeting identity run, the 3B narrator proposed giving
`silver-key` to Alice. The validator refused it, which is H05 and C06
behaving correctly. J01-J03 concern the schema, which is unaffected. A
genre-neutral example is post-v1 prompt work.
17. **With the 3B narrator at a 4,096 window, the state lags the narration.** In
the closeout identity run, one proposal in ten applied. The scene was never
updated, and a person the narration introduced never became an entity. Every
refusal and unparseable block was recorded and shown to the model, as
designed. This is risk 2 observed, not a new product defect.
---
## Q. Planning and document changes
@@ -1507,6 +1539,9 @@ accepted for v1, or a decision for the owner at acceptance.
| `planning/archive/milestone-reports/` | M9's and M10's reports moved here. | rotation |
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Revised 2026-09-14**: M01-M04 on the complete evidence run; §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R rewritten; the §G.0 addendum removed. | milestone report |
| `DEVELOPMENT.md` | The pointer to the removed §G.0 replaced; a new section on logging a GPU inference host during a long run. | developer docs |
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Closeout, 2026-09-14**: new §S (exact-tree verification, the identity results) and §T (acceptance record); the 82-test count and §L's narrator corrected; §A, §B, §D.1, §F, §L, §O, §P and §R brought up to date. | milestone report |
| `planning/BUILD-MILESTONES.md`, `VERSION.md` (v4.0), `README.md`, `V1-ACCEPTANCE-TESTS.md` §P and §P3, root `README.md` | M11 accepted and the release gate recorded as passed; a post-v1 backlog. | milestone status |
| `backend/tools/m11_browser.py` | The G01 import wait, which could not fail. | release-test infrastructure |
**Revised after this report, in planning package v3.9 (2026-09-14):**
`planning/V1-ACCEPTANCE-TESTS.md` gained result blocks for M01-M04 and corrected
@@ -1532,7 +1567,7 @@ pinned by a test so a later milestone changes it deliberately.
**1. Does every REQUIRED FOR V1 acceptance test pass?**
**Yes.** All 85 pass, with H09 recorded NOT APPLICABLE on the condition its own
**Yes.** All 82 pass, with H09 recorded NOT APPLICABLE on the condition its own
text states. M01 to M04 pass on the evidence run in §G, and every other REQUIRED
test passes on the evidence in §F.
@@ -1558,8 +1593,9 @@ across the GPU runs (§N, §P).
no route, no DNS, first page load, every asset local, campaign creation, state
extraction, knowledge import and retrieval, prompt assembly, export, import, the
media module inert, and campaigns surviving a container restart. A turn with no
model reachable is reported and corrupts nothing. §J. Re-run on the `ef25b0a`
tree (§P).
model reachable is reported and corrupts nothing. §J. Repeated at closeout on
the release-candidate tree, 23 of 23, from an image built with `--no-cache`
(§S.4).
**5. Does branch/memory/summary/scene isolation pass?**
@@ -1569,7 +1605,7 @@ The evidence run also restored a Save Point across a live memory bank. §G.2.
**6. Does trusted-LAN HTTPS inference still pass?**
**Yes, on the CPU reference host's evidence.** The browser run, the identity
**Yes, and re-established at closeout on the candidate tree** (§S.3, §S.5). The browser run, the identity
diagnostic and the CPU-host campaigns ran against Ollama on a separate physical
machine over HTTPS, with a private CA in the application machine's OS trust store,
verification on and no bypass. The storyteller stayed loopback-bound. **The final
@@ -1598,9 +1634,9 @@ stamp. §K.
**10. Does the browser pass the release workflow?**
**Yes, on the `ef25b0a` tree.** 38 checks, 0 failed, 0 skipped, in Firefox
154.0.1 against the built SPA served by FastAPI. The commits after it change no
frontend file, and the browser run was not repeated. §L, §P.
**Yes, on the release-candidate tree.** 38 of 38 in Firefox 155.0.1 at closeout,
against the SPA built from that tree and served by FastAPI, after a harness fix
(§S.3). Before that, 38 of 38 in Firefox 154.0.1 on `ef25b0a` (§L).
**11. Are there any unresolved blockers to independent v1 acceptance?**
@@ -1610,8 +1646,11 @@ was weakened. A reviewer should weigh five things before signing:
- M04's recovery ran through authoritative state, not memory retention (§G.4)
- the thin real-token headroom (§N)
- the narrator restating prompt text (§P)
- the identity diagnostic's results being absent from this report (§P)
- the black-box runs predating the later commits (§P)
- the extractor shapes and the lagging state the closeout's identity run found
(§S.6; §P risks 6, 16 and 17)
*At closeout:* the identity results are now in §S.5, and the black-box runs no
longer predate the candidate (§S). The acceptance decision is §T.
**12. Is the tree safe to commit as the M11 release candidate?**
@@ -1622,6 +1661,275 @@ was weakened. A reviewer should weigh five things before signing:
the Docker image were built for the first revision and not rebuilt since. The
remaining commit, this report's revision, is the owner's to sign.
*At closeout:* the report's revision is committed as `d198806`, and the planning
revision as `3652dc6`, both signed. The suites, the production build and a
`--no-cache` Docker image were all re-run on `3652dc6` (§S.2). The closeout
commit is staged for the owner's signature.
---
*Written by the implementer. Not an acceptance record.*
## S. Closeout verification on the release candidate (2026-09-14)
The browser, offline and identity runs above were taken on `ef25b0a`. Three later
product commits changed the turn commit, memory retrieval, narration extraction
and the state renderer's headings. The closeout repeated all three runs on the
exact release-candidate tree, and rebuilt and re-tested everything else there.
**The 100-turn campaign was not repeated.** The closeout changed no product code,
and that run was taken on `96c1bf5`, whose product code the candidate carries
unchanged.
### S.1 The candidate
| | |
| --- | --- |
| **Candidate** | `3652dc6fae903cad379da8efd5086709b5b9c6f5` on `m11-release-validation`, signed by the owner (`%G?` = `G`) |
| **Product code** | identical to `96c1bf5`. `d198806` and `3652dc6` change documentation only. No file under `backend/`, `frontend/`, `Dockerfile`, `docker-compose.yml` or the start scripts differs |
| **Working tree** | clean at the start, nothing staged, nothing untracked. The one source change during the closeout is the harness fix in §S.3, to `backend/tools/m11_browser.py`, which is neither part of the application nor in the image. No application file was modified during any run |
| **Signatures** | every commit from `144406c` to `3652dc6` is signed by the owner |
| **Provenance** | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` md5 is `07fde30437134836e2ee875e82a7cd31`, and `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` |
| **Dependencies** | `requirements.txt`, `requirements.lock`, `package.json` and `package-lock.json` are unchanged since `1013c94`. The lock's 35 pins match the environment, with no drift. FastAPI 0.141.1, uvicorn 0.52.4, SQLAlchemy 2.0.52, httpx 0.28.1, pydantic 2.13.5, tiktoken 0.14.0, python-multipart 0.0.32; React and React DOM 19.2.7, React Router 7.18.3; Vite 8.1.3 |
| **Application machine** | as §E.1, now on Ubuntu 24.04.5 LTS, kernel 7.0.0-31-generic, Docker 29.8.0, and Firefox 155.0.1 headless with geckodriver 0.37.1 |
| **Inference** | trusted-LAN: the CPU reference host (§E.1), Ollama 0.33.0 over HTTPS with a private CA in the application machine's OS trust store, verification on, no bypass. `qwen2.5:3b-instruct` (digest `357c53fb659c`), served at a verified 4,096 window. The storyteller listened on loopback only |
| **Evidence** | `$HOME/m11-evidence/closeout-3652dc6/` |
### S.2 Suites and builds
| Check | Command | Result |
| --- | --- | --- |
| Backend suite | `.venv/bin/python -m pytest tests/ -q`, with no `AIDND_TEST_*` set | **1,421 passed, 17 skipped, 0 failed**, 786 s |
| Frontend suite | `npm test` | **161 of 161**, 14 files |
| Lint | `npm run lint` | exit 0; six `only-export-components` warnings, no errors |
| Production build | `npm run build`, after deleting `frontend/dist` | 16 files, including `index-CZ7_g_3j.js` (392.69 kB) and `index-g-TAQQWd.css` (47.96 kB) |
| Docker image | `docker build --no-cache`, run by `tools/m11_offline.py` | `sha256:c9b872c2f4b701c8df32c58af944a3f456d1f4a6dd45105257ba6739f2382ed0`. `npm ci` and `pip install` ran rather than coming from cache. **The image's SPA is byte-identical, file for file, to the local build**, so the browser runs below exercised what ships |
### S.3 Browser regression
`tools/m11_browser.py`, against the SPA from §S.2, served by FastAPI on
loopback, with the narrator over trusted-LAN HTTPS:
| Run | Harness | Result |
| --- | --- | --- |
| 1 | as committed | **34 passed, 0 failed, 1 skipped**, 110 s. Modal focus containment skipped: "no Delete control found" |
| 2 | as committed, unchanged | **38 passed, 0 failed, 0 skipped**, 87 s |
| 3 | with the fix below | **38 passed, 0 failed, 0 skipped**, 69 s. **This is the reported result** |
The skip was a harness defect, not a product one. After Import is pressed, G01
waited for the word "hidden" in the page. The scenario's campaign is titled
*Hidden Knowledge*, and the page shows that title in its heading, so the wait
held at once. "A local file imports through the browser" was therefore recorded
without evidence. The modal check that follows could also run before the
imported source's row, which carries the Delete control, had rendered.
The frontend and the harness are unchanged since `ef25b0a`. Every run's database,
`ef25b0a`'s included, holds the imported source, so the import itself worked
every time. The harness now waits for the source's own row and records its
title (`▸ hidden`). Runs 1 and 2 are kept, and only run 3 used the fixed harness.
Run 3 passed everything §L lists:
- two real turns, and the tab title;
- the position indicator before and after Undo, and Undo and Redo;
- H06, H07, G09 and H04, with hostile narration;
- G01, through the real file input;
- the dialog's focus, accessible name, contents and Escape;
- §38, the context inspector, H10 and the CSP;
- the accessibility measurements, with rendered contrast unchanged at 14.57,
5.48, 13.57 and 5.88 to 1.
**Not driven in a real browser by this harness**, on this tree or before:
- Retry, Save Point creation and restore, state correction, and the
narration-length control;
- a failed generation shown to the reader;
- the export download, which is exercised only as far as the click (§L).
Those behaviours are covered on the candidate tree by the backend and component
suites, and by the 100-turn campaign through the API (§F, §G.2). The harness was
not extended, because the closeout is a regression rerun.
### S.4 Offline
`tools/m11_offline.py`: **23 passed, 0 failed**, 28 s. It ran the §S.2 image with
`--network none` and a fresh volume, and made the same 23 checks as §J:
- **Isolation:** no route, and no DNS.
- **Page and assets:** first page load; no remote origin; a CSP; every
referenced asset local.
- **Local operations:** campaign creation; state extraction; knowledge import;
prompt assembly; retrieval.
- **Failed turn with no model reachable:** reported, with no narration accepted,
the player's words kept, state unchanged, and the earlier story intact.
- **Export and import:** both work, with no secret in the bundle.
- **Media:** the module is inert with no provider.
- **Restart:** campaigns survive a container restart.
### S.5 The identity diagnostic
**Self-test** (`--scripted --inject`, no model): 7 signals, all STATE DEFECT. At
beat 5 the scripted narrator creates a second Alice. `shared_display_name` fires
at beat 5, and `duplicate_character_creation` fires at beats 5-10. The detectors
fire.
**Against a real narrator:**
| | |
| --- | --- |
| **Candidate** | `3652dc6` |
| **Narrator** | `qwen2.5:3b-instruct` on the CPU reference host over HTTPS. The window was verified at 4,096 on every turn, and the 16,384 budget was capped to it. `max_output_tokens` 500. No embedding model, so memory and summaries were off |
| **Fixture** | the companion *Multi-Character Identity Test* (`TEST-CAMPAIGN-FIXTURE.md` Appendix A): Bill, the protagonist, with Alice, Roger and John in one meeting room; ten beats |
| **Fixture accepted** | yes. The opening correction applied all six events and refused none, and the scene named its location and who was present. The harness stops on either failure rather than running a degraded campaign |
| **Turns** | 10 of 10 accepted, in 18 min 46 s |
| **Result** | **0 signals. Verdict: "no objective identity defect detected."** |
| **Evidence** | `identity/stdout.txt`; the final state, prompt, settings and an M9 bundle in `identity/turn-99/` |
**What the state did.** The narrator made ten state proposals:
- one applied (beat 3: John present, in the meeting room);
- five carried no block;
- three were unparseable: two cut off by the output limit, and one a parroted
reminder;
- one was refused: `set_possession` of `silver-key`, which does not exist.
The scene the fixture set was never updated, so Roger's exit went unrecorded.
From beat 6 the narration introduced a fifth person, "Mike", who never became an
entity, because the proposal that would have created him was cut off.
**What the run establishes:**
- the fixture is valid and was not silently refused;
- the authoritative state held four distinct people, with no shared display name
and no drift in the protagonist;
- the prompt's state section agreed with the document on every turn;
- the detectors fire when a confusion is planted;
- what a future occurrence needs is preserved and portable.
**What it cannot establish:**
- **The root cause of post-M8 finding D.** No same-name confusion occurred, so
the finding was not reproduced.
- **Prose-level misattribution.** This is not graded. The narration is kept for a
person to read.
- **A stale scene.** The objective checks cover only the protagonist's presence,
so Roger's unrecorded exit and the invented Mike raise no signal.
- **Derived contamination.** Memory and summaries were off.
- **Behaviour at a 16,384 window, or with another narrator.**
A clean verdict over a state that lags its narration is not evidence that
identities stay straight under a small model. What the diagnostic does show is
that, when identities go wrong, it can say whether the state, the context or the
derived data carried the error.
### S.6 Findings
**Release blockers: none.**
**Non-blocking, found at closeout:**
1. **Protocol shapes the extractor does not remove.** On 4 of 10 stored AI turns
in the identity run (depths 10, 12, 14 and 20), the narration kept event-call
syntax (`> Create_entity(...)`, `> set_possession(silver-key, "alice")`), a
lone `Scene:` line, and a parroted length hint ending "append the state block
well inside the limit.]". Replaying those four texts through the candidate's
`narrative.extract.split` leaves all four unchanged.
- **Why it survives.** The bracket survives because `_is_echoed_instruction`
recognises the reminder by "state block" together with "events list", and
the length hint names only the first.
- **Scope.** The 100-turn runs on `0c7316f` and `96c1bf5` store 0 turns in
these shapes.
- **Why it is not a blocker.** No REQUIRED FOR V1 test names stored-narration
purity, and `TECHNICAL-DESIGN.md` §15.4 holds that removing story is worse
than leaving protocol.
- **Decision.** The owner chose on 2026-09-14 to carry it as a residual risk
(§P risk 6) rather than change product code. A code change would have
invalidated the 100-turn evidence. §O.8's "every shape observed" was true
of the shapes observed when it was written.
2. **The state rule's example is fantasy** (§P risk 16).
3. **The state lagged the narration** at a 4,096 window with the 3B narrator
(§P risk 17).
**Harness defect, fixed:** the G01 import wait (§S.3, §O).
**Corrections to this report:** the REQUIRED FOR V1 count is 82 (§A, §R), and the
browser narrator in §L was `qwen2.5:3b-instruct`.
### S.7 Release-shaped smoke test
This ran after the documentation changes above. It started from the Dockerfile,
not from a development server. The script is not a repository harness, and its
evidence is `smoke-run2/`. The one source difference from the candidate is the
§S.3 harness fix, which is not in the image. **14 passed, 0 failed:**
- **Build and start:** the image builds with `--no-cache`; the container starts
on a fresh volume; it is published on host loopback only
(`127.0.0.1:18080`).
- **Page:** the first page loads, with every asset it names served locally.
- **Endpoint policy:** the approved trusted-LAN endpoint verifies over HTTPS,
with the private CA supplied to the container through the host's trust bundle,
mounted read-only. There is no bypass. A public endpoint is refused with
HTTP 400, and the approved endpoint stays configured.
- **Play:** a campaign is created, and one normal story turn is accepted (24 s,
`qwen2.5:3b-instruct`, window 4,096).
- **Persistence:** after a container restart, the campaign reopens with the same
actions and the same state.
- **Browser:** Firefox 155.0.1 loads the reopened campaign and shows its story,
with the tab reading "Closeout Smoke — Interactive Story".
The first attempt (`smoke/`) passed its first 13 checks. It then stopped on a
defect in the smoke script itself (`browser.title` is a property), before the
browser check was recorded. It is kept, and superseded by the run above.
---
## T. Acceptance record
**Decision: PASS. M11 is accepted at closeout, 2026-09-14.**
> Does this exact build meet the v1 black-box acceptance contract, and can it be
> packaged as the first production release?
**Yes.** Every REQUIRED FOR V1 condition remains satisfied. No test was waived,
relaxed or reclassified.
| Area | Status | Evidence |
| --- | --- | --- |
| All REQUIRED FOR V1 tests | 82: 81 PASS, H09 NOT APPLICABLE | §F |
| M01-M04 | PASS, on `96c1bf5`, whose product code the candidate carries unchanged | §G |
| Offline / local-only operation | PASS, on the candidate | §S.4 |
| Trusted-LAN inference (A06) | PASS, on the candidate: HTTPS, private CA, verification on, storyteller on loopback | §S.3, §S.5 |
| Undo / Redo / Retry / Save Point | PASS: D01-D14 in the long run and the suite; Undo and Redo in a real browser on the candidate | §F, §G.2, §S.3 |
| Branch, memory, summary and scene isolation | PASS: E01-E04, in the candidate's backend suite | §H, §S.2 |
| State atomicity | PASS: L01, in the long run and in the candidate's container | §F, §S.4 |
| Export / import / recovery | PASS: I01-I07; the long campaign onto a clean data directory, 16 of 16 | §K |
| Fresh-install vs upgraded schema parity | PASS: `test_m11_migration.py`, in the candidate's suite | §K, §S.2 |
| Fantasy fixture | PASS: the long-run campaign | §G, §I |
| Science-fiction fixture | PASS: `test_m11_scifi.py`, in the candidate's suite | §I, §S.2 |
| Security and local endpoint policy | PASS: H01-H12 | §F, §J, §S.4 |
| Browser release workflow | PASS: 38 of 38 on the candidate | §S.3 |
**Qualifications that stand:**
- M04's fact was recovered through authoritative state, not independent memory
retention. The owner accepted that on 2026-09-13.
- K04, a SHOULD test, passes on its deferred branch.
- The export download is exercised in the browser only as far as the click.
- Retry, Save Point, state correction, narration length and failed generation are
proved through the API and the component suite, not driven in a real browser.
- The real-token headroom at the largest long-run prompt is 42 tokens (§N).
- A06's HTTPS evidence ran at a 4,096 window on the CPU host, and the long run
used plain HTTP to a LAN GPU host.
- The residual risks in §P remain residual.
**Four separate events:**
| Event | State |
| --- | --- |
| M11 accepted | recorded here, 2026-09-14; takes effect with the owner's signed closeout commit |
| Release candidate verified | done, 2026-09-14, on `3652dc6` (§S) |
| Release commit signed | not yet: the closeout commit is staged for the owner |
| `v1.0.0` tagged | not done. The owner's decision, and the tag must point at the signed closeout commit |
**Where this report stays.** `planning/reports/` holds the most recently
completed milestone's report. That report moves to the archive when the next
milestone's report is written. There is no next milestone, so this report stays
where it is. No new convention was invented for it.
---
*§A-§R written by the implementer. §S and §T added at the M11 closeout,
2026-09-14; §T is the acceptance record.*