M11 closeout: accept v1 release validation
The browser, offline and identity runs had last been taken onef25b0a. The closeout repeated them on the exact release-candidate tree,3652dc6, whose product code is identical to96c1bf5, where the 100-turn evidence was run. No product code changed, so the long-run evidence stands. On3652dc6: - backend suite: 1421 passed, 17 skipped, 0 failed - frontend suite: 161 of 161; lint clean - production build clean - docker build --no-cache: image SPA byte-identical to the local build - browser regression: 38 of 38 - no-network container: 23 of 23 - identity diagnostic: 0 signals; the scripted self-test's detectors fire - release-shaped smoke test from the image: 14 of 14 - M11 report: new S (exact-tree verification, including the identity results the report never carried) and T (acceptance record). Corrections: the REQUIRED FOR V1 count is 82, not 85, and L's browser narrator was qwen2.5:3b-instruct. P gains risks 16 and 17; risk 6 is widened. - BUILD-MILESTONES.md: M11 COMPLETE / ACCEPTED, and a post-v1 backlog. - V1-ACCEPTANCE-TESTS.md: the P release gate's result, and the P3 disposition's run. - planning/README.md, VERSION.md (v4.0), README.md: status, map, stop rule. - tools/m11_browser.py: the G01 import wait could not fail, because the scenario's campaign is titled "Hidden Knowledge". It now waits for the imported source's row. Found and carried, not fixed. The identity run stored protocol shapes the extractor leaves, on 4 of 10 turns at a 4,096 window: event-call syntax and a parroted length hint. The owner chose residual risk. The state rule's example is fantasy, and the state lagged the narration. None occurs in the 100-turn evidence. No requirement changes. No release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CBTT3qbkGoemWD7BRvvxpT
This commit is contained in:
co-authored by
Claude Opus 5
parent
3652dc6fae
commit
432f04100b
@@ -346,6 +346,10 @@ most interesting engineering in the repo.
|
||||
|
||||
## Repo notes
|
||||
|
||||
- **Status:** milestones M1-M11 are complete. The v1 release gate passed on
|
||||
2026-09-14 (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md),
|
||||
§T). Passing the gate is not a release: there is no `v1.0.0` tag yet.
|
||||
|
||||
- `planning/` is this fork's own package: the product specification, the architecture
|
||||
decisions, the milestone plan, the acceptance contract, and a review report for every
|
||||
milestone shipped. Start at [`planning/README.md`](planning/README.md).
|
||||
|
||||
@@ -291,10 +291,16 @@ def check_hidden_knowledge(browser: Browser, site: Site, checks: Checks):
|
||||
checks.record("G01", "Import becomes available once a file is chosen",
|
||||
disabled is False, f"disabled={disabled}")
|
||||
browser.click(submit)
|
||||
browser.wait_until(
|
||||
"document.body.textContent.toLowerCase().includes('hidden')", timeout=90,
|
||||
what="the imported source appears in the library")
|
||||
checks.record("G01", "a local file imports through the browser", True, "hidden.md")
|
||||
# Wait for the source's own row, not for the word "hidden" in the page. This
|
||||
# campaign is titled "Hidden Knowledge", so that condition held before Import
|
||||
# was pressed: the check could not fail, and the modal check below raced the
|
||||
# row and its Delete control (M11 closeout).
|
||||
browser.wait_for(".knowledge-row button.danger", timeout=90)
|
||||
row_title = browser.js(
|
||||
"const t = document.querySelector('.knowledge-row .knowledge-title');"
|
||||
"return t ? t.textContent.trim() : ''")
|
||||
checks.record("G01", "a local file imports through the browser",
|
||||
"hidden" in row_title.lower(), row_title)
|
||||
|
||||
# §21: a real modal, opened from a real control, containing focus.
|
||||
_check_modal_focus(browser, checks)
|
||||
|
||||
@@ -1449,12 +1449,25 @@ Release gate in `V1-ACCEPTANCE-TESTS.md`:
|
||||
|
||||
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
|
||||
|
||||
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07; long-run evidence complete 2026-09-13; awaiting independent review/acceptance
|
||||
## Status: COMPLETE / ACCEPTED — 2026-09-14 (implemented 2026-09-07; long-run evidence 2026-09-13; release-candidate closeout 2026-09-14)
|
||||
|
||||
Implemented on `m11-release-validation` from the signed M10 commit `1013c94`.
|
||||
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package, written
|
||||
for a release reviewer. **M11 is not marked accepted here**; that is the
|
||||
reviewer's to record, and no release tag exists.
|
||||
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package.
|
||||
|
||||
**M11 is accepted, and v1 release validation is complete.** The closeout on
|
||||
2026-09-14 re-verified the exact release-candidate tree (`3652dc6`, whose product
|
||||
code is identical to `96c1bf5`, where the 100-turn evidence was run). The backend
|
||||
suite passed 1,421 with 17 skipped and 0 failed. The frontend suite passed 161 of
|
||||
161, lint was clean, and the production build succeeded. A `--no-cache` Docker
|
||||
image built. The browser regression passed 38 of 38, the no-network container
|
||||
passed 23 of 23, and the identity diagnostic ran clean, with its results now in
|
||||
the report (§S). Every REQUIRED FOR V1 test passes, 82 in all with H09 not
|
||||
applicable (report §T). The acceptance takes effect with the owner's signed
|
||||
closeout commit. **No release tag exists**: tagging `v1.0.0` is a separate
|
||||
decision, and the tag must point at that signed commit.
|
||||
|
||||
**M1-M11 are all complete. There is no M12.** Post-v1 work is backlog, listed
|
||||
below under *Post-v1 backlog*, and none of it is an unfinished v1 milestone.
|
||||
|
||||
**The release blocker it was given, and how it was closed.** M8 measured the
|
||||
reference deployment enforcing a **4,096**-token input window while the
|
||||
@@ -1520,9 +1533,43 @@ switched on:
|
||||
extractor now removes every shape observed. Replaying 443 real turns through it
|
||||
changed no turn it had previously left clean. (`0c7316f`, `96c1bf5`)
|
||||
|
||||
**Left for the reviewer**, in the report's §P: the real-token headroom at the
|
||||
largest prompts is 23-42 tokens; the narrator restates prompt text in its prose;
|
||||
and the identity diagnostic's results are not in the report.
|
||||
**Carried past v1 as residual risk**, in the report's §P:
|
||||
- the real-token headroom at the largest prompts is 23-42 tokens;
|
||||
- the narrator restates prompt text in its prose;
|
||||
- stored narration can keep protocol shapes the extractor does not remove.
|
||||
The closeout found event-call syntax and a parroted length hint, and the owner
|
||||
chose to carry this rather than change code (report §S.6).
|
||||
|
||||
The identity diagnostic's results are now in the report, §S.5.
|
||||
|
||||
---
|
||||
|
||||
# Post-v1 backlog — not milestones
|
||||
|
||||
Work recorded for after v1. None of it is a v1 requirement or an unfinished v1
|
||||
milestone, and none of it has a brief. Each item needs one before work begins.
|
||||
Sources are the M11 report's §P and §S.6.
|
||||
|
||||
- **Context-window safety margin.** The largest prompts leave 23-42 real tokens,
|
||||
and Ollama cuts an over-window prompt with no error. Consider a deliberate
|
||||
reserve, or counting with the narrator's own tokenizer.
|
||||
- **The narrator restating prompt and state text** in its prose, including the
|
||||
protocol shapes the extractor still leaves: event-call syntax, a parroted
|
||||
length hint, and a lone section heading.
|
||||
- **A genre-neutral example in the state rule**, in place of `silver-key` and
|
||||
`aldric`.
|
||||
- **Independent long-term-memory retention.** No run showed memory keeping a
|
||||
planted fact without authoritative state.
|
||||
- **Browser automation of the export download**, and real-browser coverage of
|
||||
Retry, Save Point, state correction, narration length and failed generation.
|
||||
- **A practical bundle-size ceiling.** M9 estimated about 279 turns against the
|
||||
20 MB import limit.
|
||||
- **Identity follow-up on the next real occurrence.** Classify it with
|
||||
`tools/m11_identity.py`. Consider a stale-scene check, and a run with memory on
|
||||
at a full window.
|
||||
- **WCAG 1.4.11 control-boundary contrast** (1.33:1 resting, 1.75:1 hover).
|
||||
- **Backup hardening:** `integrity_check`, and scheduled backups.
|
||||
- **Real media-provider adapters** against `MEDIA-EXTENSION-CONTRACT.md`.
|
||||
|
||||
---
|
||||
|
||||
|
||||
+53
-33
@@ -3,9 +3,10 @@
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
milestones **M1 through M8 implemented and accepted**; **M9 and M10 implemented,
|
||||
committed and signed**; **M11 implemented and verified, awaiting independent
|
||||
review and v1 acceptance**, with its long-run evidence complete (2026-09-13). M11 is the last planned milestone. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
**milestones M1 through M11 complete**. **M11 was accepted at its closeout
|
||||
(2026-09-14), and the v1 release gate passed** on the release-candidate tree.
|
||||
There is no further planned milestone. The signed closeout commit and any
|
||||
`v1.0.0` tag are separate events and the repository owner's to perform. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
|
||||
independent review found a real defect and a corrective pass fixed it.
|
||||
|
||||
@@ -33,21 +34,25 @@ scene packet derived on read, provider contracts with an empty registry, and no
|
||||
dependency added. Its central finding was that the scene snapshot the media
|
||||
contract asks for **already existed**, built by M5.
|
||||
|
||||
**M11 — v1 Security, Long-Run, and Release Validation — is implemented and
|
||||
verified, and awaits independent review** (2026-09-07).
|
||||
`reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. It closed the
|
||||
context-window release blocker M8 found, fixed two defects the validation itself
|
||||
surfaced, and disposed of all four post-M8 playtest findings. **It is not
|
||||
accepted**, and there is no release tag. **M01-M04 now pass on a complete
|
||||
100-turn run** (2026-09-13, commit `96c1bf5`). That run found two more product
|
||||
defects, both fixed: a turn locking out its own memory bank, and the narrator's
|
||||
protocol stored as story. The report was revised to match (`d198806`). All of it
|
||||
is committed and signed on `m11-release-validation`; `main` remains M10, the last
|
||||
*accepted* milestone.
|
||||
**M11 — v1 Security, Long-Run, and Release Validation — is complete and
|
||||
accepted** (implemented 2026-09-07; accepted at closeout 2026-09-14).
|
||||
`reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. M11:
|
||||
- closed the context-window release blocker M8 found;
|
||||
- fixed the defects the validation itself surfaced;
|
||||
- disposed of all four post-M8 playtest findings;
|
||||
- passed M01-M04 on a complete 100-turn run (2026-09-13, commit `96c1bf5`).
|
||||
|
||||
**After M11 there is no further planned milestone.** What follows is
|
||||
independent review and the v1 acceptance decision, which is the repository
|
||||
owner's.
|
||||
The closeout repeated the browser, offline and identity runs on the exact
|
||||
release-candidate tree (`3652dc6`), and recorded the acceptance (report §S and
|
||||
§T). **The v1 release gate passed.**
|
||||
|
||||
**Not yet happened**, and each a separate event:
|
||||
- the owner signing the closeout commit;
|
||||
- merging to `main`, which still points at M10;
|
||||
- any `v1.0.0` tag.
|
||||
|
||||
**There is no further planned milestone and no M12.** Post-v1 work is backlog
|
||||
(`BUILD-MILESTONES.md`, *Post-v1 backlog*).
|
||||
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
and why.
|
||||
@@ -138,8 +143,9 @@ Two standing qualifications:
|
||||
11. `V1-ACCEPTANCE-TESTS.md`
|
||||
12. `DECISIONS/` — all of them; they are short.
|
||||
13. `reports/M11-IMPLEMENTATION-REPORT.md`, for what the release validation
|
||||
actually found — read as claims to check, not a record, until it is
|
||||
reviewed. Nothing in `planning/archive/` unless sent there.
|
||||
actually found. §S is the closeout's verification on the release candidate,
|
||||
and §T is the acceptance record. Nothing in `planning/archive/` unless sent
|
||||
there.
|
||||
|
||||
## Architectural decisions
|
||||
|
||||
@@ -172,9 +178,14 @@ that is the one the next milestone's planning has to consult:
|
||||
|
||||
- `reports/M11-IMPLEMENTATION-REPORT.md` — the release validation: the
|
||||
acceptance matrix run end to end, the 100-turn campaign, the offline and
|
||||
browser evidence, and every defect the run found. Written by the implementer
|
||||
for an independent reviewer, so it is a set of claims with the measurements
|
||||
attached and **not** a record of acceptance.
|
||||
browser evidence, and every defect the run found. §A-§R are the implementer's
|
||||
claims with the measurements attached. §S is the closeout's re-verification on
|
||||
the release candidate, and §T is the acceptance record.
|
||||
|
||||
It stays in `reports/` after acceptance. The rotation moves a report to the
|
||||
archive when the next milestone's report is written, and after M11 there is no
|
||||
next milestone. Nothing here calls for moving it, so no new convention was
|
||||
invented to do so.
|
||||
|
||||
M9's and M10's reports both moved to `archive/milestone-reports/` when this one
|
||||
was written. M9's had been kept here past its turn because M9 was unaccepted;
|
||||
@@ -310,34 +321,43 @@ Milestone M8 COMPLETE / ACCEPTED (2026-09-06)
|
||||
story operations review + closeout, in sequence
|
||||
|
|
||||
v
|
||||
Milestone M9 COMPLETE — awaiting review (2026-09-07)
|
||||
Milestone M9 COMPLETE — committed and signed (44edece)
|
||||
export, backup, recovery, archive/milestone-reports/
|
||||
migration hardening bundle format v3; SQLite online backup
|
||||
|
|
||||
v
|
||||
Milestone M10 COMMITTED AND SIGNED (1013c94)
|
||||
Milestone M10 COMPLETE — committed and signed (1013c94)
|
||||
future media extension hooks archive/milestone-reports/
|
||||
scene packet derived, not stored; no media
|
||||
|
|
||||
v
|
||||
Milestone M11 IMPLEMENTED AND VERIFIED (2026-09-07)
|
||||
Milestone M11 COMPLETE / ACCEPTED (2026-09-14)
|
||||
v1 security, long-run, release reports/M11-IMPLEMENTATION-REPORT.md
|
||||
validation M01-M04 evidence complete (2026-09-13);
|
||||
awaiting independent review; no release tag
|
||||
validation M01-M04 on 96c1bf5; closeout on 3652dc6
|
||||
|
|
||||
v
|
||||
v1 acceptance the repository owner's decision
|
||||
v1 release gate PASSED (2026-09-14)
|
||||
signed release commit and v1.0.0 tag:
|
||||
the repository owner's, not yet done
|
||||
|
|
||||
v
|
||||
Post-v1 backlog only; no milestone planned
|
||||
```
|
||||
|
||||
## Stop Rule
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
|
||||
**Every planned milestone is now implemented.** M11 is verified and committed,
|
||||
its long-run evidence is complete, and the next action is not another milestone: it is an independent review of
|
||||
`reports/M11-IMPLEMENTATION-REPORT.md` against the acceptance contract, and then
|
||||
the owner's v1 acceptance decision. **Do not begin post-v1 work before that
|
||||
decision**, and do not treat M11's own report as the acceptance record.
|
||||
**Every planned milestone is complete, and M11 is accepted.** The v1 release gate
|
||||
passed on 2026-09-14 (M11 report §T).
|
||||
|
||||
The next actions are the owner's:
|
||||
1. sign the closeout commit;
|
||||
2. decide on the `v1.0.0` tag, pointed at that signed commit.
|
||||
|
||||
Neither is a milestone. **Do not begin post-v1 work as though it were a v1
|
||||
milestone.** Anything after v1 starts from the *Post-v1 backlog* in
|
||||
`BUILD-MILESTONES.md`, with a brief of its own.
|
||||
|
||||
All three questions the M8 debt raised against M9 are settled and recorded:
|
||||
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
|
||||
|
||||
@@ -2555,6 +2555,36 @@ The release candidate should not be called v1.0 until:
|
||||
- branch/memory lineage isolation passes,
|
||||
- fantasy and science-fiction fixtures both pass.
|
||||
|
||||
### Result — gate passed on the release candidate (M11 closeout, 2026-09-14)
|
||||
|
||||
Measured on the release-candidate tree `3652dc6`, whose product code is identical
|
||||
to `96c1bf5`. The per-test matrix is M11 report §F. The closeout's exact-tree
|
||||
verification is §S, and the acceptance record is §T.
|
||||
|
||||
- **All REQUIRED FOR V1 tests pass.** There are 82. 81 pass, and H09 is NOT
|
||||
APPLICABLE on the condition its own text states: the product extracts no
|
||||
archives, and a test enforces that.
|
||||
- **Approved exceptions:** none. No REQUIRED test was waived, relaxed or
|
||||
reclassified.
|
||||
- **Security offline test:** 23 of 23, in a container with no network and a
|
||||
fresh volume, built with `--no-cache` from the candidate tree.
|
||||
- **100-turn long run:** M01-M04 pass on `96c1bf5` (results under M01-M04
|
||||
above). The closeout changed no product code, so that evidence stands.
|
||||
- **Export/import recovery, including an undone head:** I01-I07. The 100-turn
|
||||
campaign moved to a clean data directory, 16 of 16.
|
||||
- **Undo/Redo/Retry/checkpoint:** D01-D14, in the long run and in the browser.
|
||||
- **Narrative-state events at realistic context length:** C06, M01.
|
||||
- **No first-use runtime download:** H11, in the offline container.
|
||||
- **Branch and memory lineage isolation:** E01-E04.
|
||||
- **Fantasy and science-fiction fixtures:** both pass.
|
||||
|
||||
**Qualifications that stand:**
|
||||
|
||||
- M04 recovered its fact through authoritative state, not independent memory
|
||||
retention. The owner accepted that on 2026-09-13.
|
||||
- K04, a SHOULD test, passes on its deferred branch.
|
||||
- The browser's export *download* is exercised only as far as the click.
|
||||
|
||||
## Q. Current Recommendation
|
||||
|
||||
Use this document as:
|
||||
@@ -2700,8 +2730,13 @@ the run reports it), which is the control this kind of tool most often lacks.
|
||||
|
||||
**This remains a test-design task and is still not an acceptance test.** The
|
||||
model-quality half is not a pass/fail property of the application, and M11 does
|
||||
not make it one. **The M11 report does not record the diagnostic's run
|
||||
results.** This was corrected on 2026-09-14: the section the report pointed to
|
||||
was left empty. The report records the fixture defect the diagnostic caught in
|
||||
itself, and the run's evidence is in the implementer's
|
||||
`m11-evidence/identity-recheck` directory.
|
||||
not make it one.
|
||||
|
||||
**Run at the M11 closeout (2026-09-14), on the release-candidate tree `3652dc6`.**
|
||||
The narrator was `qwen2.5:3b-instruct` over trusted-LAN HTTPS, at a verified
|
||||
4,096 window. The fixture was accepted with nothing refused. 10 of 10 turns were
|
||||
accepted with **0 signals**, and the scripted self-test's detectors fired. The
|
||||
state lagged the narration, though: one proposal in ten applied, and a stale
|
||||
scene is outside the objective checks. The run therefore reproduces nothing
|
||||
about the original finding and establishes no root cause. M11 report §S.5 has
|
||||
the results and what they cannot establish.
|
||||
|
||||
+31
-2
@@ -1,8 +1,37 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v3.9
|
||||
- **Package:** Adventure Storyteller Planning Package v4.0
|
||||
- **Revision date:** 2026-09-14
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. Its long-run evidence is complete (2026-09-13): all 85 REQUIRED tests pass, with H09 not applicable.
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M11 complete**; **M11 accepted at its closeout (2026-09-14), and the v1 release gate passed** on the release-candidate tree. All 82 REQUIRED FOR V1 tests pass, with H09 not applicable. There is no further planned milestone. The signed closeout commit and any `v1.0.0` tag are the repository owner's, and neither exists as of this revision.
|
||||
|
||||
## v4.0 — M11 closeout: v1 release validation accepted (2026-09-14)
|
||||
|
||||
No requirement change, and no product code change. The closeout re-ran the
|
||||
black-box evidence on the exact release-candidate tree, `3652dc6`. Its product
|
||||
code is identical to `96c1bf5`, where the 100-turn evidence was taken. The
|
||||
browser, offline and identity runs had last been run on `ef25b0a`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `reports/M11-IMPLEMENTATION-REPORT.md` | **New §S**: the closeout's exact-tree verification, including the identity diagnostic's results, which the report had never carried. **New §T**: the acceptance record. **Corrections**: the REQUIRED FOR V1 count is 82, not 85 (the 85 counted four SHOULD rows and left out H09), and §L's browser narrator was `qwen2.5:3b-instruct`, not the `-16k` model. §A, §B, §D.1, §F (A06, G01), §L, §O, §P and §R now reflect the closeout. | milestone report + corrections |
|
||||
| `BUILD-MILESTONES.md` | **M11 marked COMPLETE / ACCEPTED.** A *Post-v1 backlog* section records work that is not a milestone. | milestone status |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | **§P V1 Release Gate gains its result.** The §P3 M11 disposition now records the diagnostic's run. | acceptance evidence |
|
||||
| `planning/README.md` | Status, milestone map, stop rule, and why the M11 report stays in `reports/`. | index |
|
||||
| `README.md` | One status sentence: v1 release validation passed; not yet tagged. | developer docs |
|
||||
| `backend/tools/m11_browser.py` | **Harness fix.** G01's import wait could not fail, because the campaign's own title contains "hidden". It now waits for the imported source's row. | release-test infrastructure |
|
||||
|
||||
**What the closeout found.** No release blocker. One harness defect, fixed. Three
|
||||
non-blocking observations, all from the identity run on a 4,096 window:
|
||||
|
||||
- the extractor leaves event-call syntax and a parroted length hint in stored
|
||||
narration;
|
||||
- the state rule's fixed example carries fantasy slugs, which the model copied;
|
||||
- the state lagged the narration.
|
||||
|
||||
The owner chose on 2026-09-14 to carry the extractor gap as a residual risk
|
||||
rather than change product code. Report §S.4.
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
## v3.9 — M11's long-run evidence, and the two defects it took to get it (2026-09-14)
|
||||
|
||||
|
||||
@@ -4,15 +4,21 @@
|
||||
|
||||
Branch `m11-release-validation`, from the signed M10 commit `1013c94`.
|
||||
Implemented and verified 2026-09-07. The long-run evidence was re-established,
|
||||
and this revision written, on 2026-09-14. This is the evidence package; it is
|
||||
not a record of acceptance, and nothing in it says M11 is accepted.
|
||||
and this revision written, on 2026-09-14.
|
||||
|
||||
**Closeout, 2026-09-14.** The browser, offline and identity runs were repeated on
|
||||
the exact release-candidate tree (§S), and M11 was accepted (§T). Sections A-R
|
||||
are the implementer's evidence as revised before the closeout, corrected where
|
||||
the closeout found them wrong. §S and §T are the closeout.
|
||||
|
||||
---
|
||||
|
||||
## A. Executive result
|
||||
|
||||
**PASS, for independent review.** All 85 tests marked REQUIRED FOR V1 pass, with
|
||||
H09 recorded NOT APPLICABLE on the condition its own text states.
|
||||
**PASS, and accepted at closeout (§T).** All 82 tests marked REQUIRED FOR V1
|
||||
pass, with H09 recorded NOT APPLICABLE on the condition its own text states.
|
||||
*Corrected at closeout:* earlier revisions said 85. That figure counted the four
|
||||
SHOULD rows in §F and left H09 out.
|
||||
|
||||
This revision supersedes the first. That revision left M01 PARTIAL at 41
|
||||
accepted turns, and then the run it described was lost to a host crash along
|
||||
@@ -55,10 +61,10 @@ narration-length setting moves a number. Post-M8 findings A and B are closed.
|
||||
- **Performance.** Timings are two particular hosts', not a product
|
||||
characteristic (§E.1).
|
||||
- **Post-M8 finding D's root cause**, which remains unestablished. The identity
|
||||
diagnostic's run results are not in this report (§D.1).
|
||||
- **That the browser, offline and identity runs cover the later commits.** They
|
||||
were re-run on the `ef25b0a` tree. The three product commits since then change
|
||||
the turn commit, memory retrieval and narration extraction (§B).
|
||||
diagnostic's closeout run, and what it can and cannot establish, are in §S.5.
|
||||
- *Closed at closeout:* the browser, offline and identity runs had predated the
|
||||
last three product commits. All three were repeated on the release-candidate
|
||||
tree (§S).
|
||||
|
||||
**Zero requirement weakenings.** No acceptance test was retired, relaxed or
|
||||
reclassified. The M04 precondition the harness measures was corrected to the
|
||||
@@ -72,7 +78,7 @@ acceptance text's own wording, "without entire transcript in prompt" (§O).
|
||||
| **Base commit** | `1013c94eb1ad283e960114aef04c19c2806b5db7` — *"M10: the seam for media, and no media"* |
|
||||
| **Signature** | `git verify-commit 1013c94` → **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, "JesseMarkowitz", trust `[ultimate]`. `%G?` = `G`. |
|
||||
| **Branch** | `m11-release-validation`, created from that commit. The working tree was clean at the start (`git status --porcelain` empty). |
|
||||
| **HEAD at this revision** | `96c1bf5`. Seven commits follow the base, all signed by the owner (`%G?` = `G`): `144406c` (M11), `fedb714`, `ef25b0a`, `fec46f6`, `f8d4010`, `0c7316f`, `96c1bf5`. This revision of the report is staged for the owner to sign. |
|
||||
| **HEAD at this revision** | `96c1bf5`. Seven commits follow the base, all signed by the owner (`%G?` = `G`): `144406c` (M11), `fedb714`, `ef25b0a`, `fec46f6`, `f8d4010`, `0c7316f`, `96c1bf5`. This revision of the report is staged for the owner to sign. **At closeout** the candidate is `3652dc6`. `d198806` and `3652dc6` follow `96c1bf5`, both are signed, and both change documentation only: no file under `backend/`, `frontend/`, `Dockerfile`, `docker-compose.yml` or the start scripts differs from `96c1bf5` (§S.1). |
|
||||
| **Upstream ancestry** | `git merge-base --is-ancestor d72f7c1b HEAD` → true. The AI-DnD fork point is still an ancestor. |
|
||||
| **LICENSE** | Unchanged — md5 `07fde30437134836e2ee875e82a7cd31`, MIT, "Copyright (c) 2026 Parth Thakkar". `PROVENANCE.md` unchanged. |
|
||||
|
||||
@@ -94,7 +100,9 @@ the offline container, the identity diagnostic and the contrast audit** were
|
||||
re-run on the `ef25b0a` tree, as that commit's message records, and their
|
||||
evidence is dated 2026-09-10. The three product commits after it (`f8d4010`,
|
||||
`0c7316f`, `96c1bf5`) change backend files only, and those runs were not
|
||||
repeated (§P). Runs taken before a product change on their own tree were
|
||||
repeated at the time. **At closeout all three were repeated on the
|
||||
release-candidate tree `3652dc6`** (§S). No reported black-box result now comes
|
||||
from product code other than the candidate's. Runs taken before a product change on their own tree were
|
||||
discarded and re-run; §O records which and why.
|
||||
|
||||
---
|
||||
@@ -222,11 +230,11 @@ turn's state. Before M11 all three settings produced *the identical sentence*;
|
||||
`test_the_three_lengths_no_longer_say_the_same_thing` fails against that.
|
||||
|
||||
**D — character identity confusion.** The diagnostic exists; the root cause does
|
||||
not, and cannot. **The diagnostic's run results are not reproduced in this
|
||||
report.** The first revision pointed here to a section that was left empty.
|
||||
What is recorded is the defect the diagnostic caught in its own fixture (§O.2),
|
||||
which is why its first run's evidence was discarded. The later run's evidence is
|
||||
in the implementer's `m11-evidence/identity-recheck` directory (§P).
|
||||
not, and cannot. The diagnostic caught a defect in its own fixture (§O.2), and
|
||||
that is why its first run's evidence was discarded. Earlier revisions of this
|
||||
report reproduced no run's results. **The closeout run's results are in §S.5**:
|
||||
clean on every objective check, with the detectors proved to fire, and with what
|
||||
the run cannot establish stated.
|
||||
|
||||
---
|
||||
## E. The context-window resolution
|
||||
@@ -442,7 +450,7 @@ evidence, it says so and the verdict is qualified.
|
||||
| A03 No cloud API key | **PASS** | suite: no `api_key` in settings, none settable through the API, no Authorization header; container: no secret in an export |
|
||||
| A04 Campaign survives restart | **PASS** | campaign: 3 genuine process restarts (4 process starts), with transcript/head/state/Save Points/knowledge/settings compared before and after each and identical every time (§G.2); container: campaigns survive a container restart |
|
||||
| A05 Failed model call does not corrupt story | **PASS** | campaign: a real failed call at a controlled point (the model name pointed at one the server does not serve), then play resumed; container: same with no model reachable at all — no narration accepted, state unchanged, the player's words kept, earlier story reachable |
|
||||
| A06 Trusted-LAN Ollama inference | **PASS** | browser + identity diagnostic + the CPU-host campaigns: Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1) |
|
||||
| A06 Trusted-LAN Ollama inference | **PASS** | browser + identity diagnostic + the CPU-host campaigns: Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1). **Re-established at closeout on the candidate tree** by the browser regression and the identity diagnostic, on the same HTTPS host (§S) |
|
||||
|
||||
### B — Core play
|
||||
|
||||
@@ -509,7 +517,7 @@ evidence, it says so and the verdict is qualified.
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| G01 Import local text | **PASS** | browser: through the real file input; container: offline |
|
||||
| G01 Import local text | **PASS** | browser: through the real file input. The harness's own assertion could not fail until the closeout fixed it (§S.3); every run's database holds the imported source. Container: offline |
|
||||
| G02 Import local Markdown | **PASS** | campaign: three sources imported (canon, reference, inspiration); browser |
|
||||
| G03 Classification | **PASS** | campaign: all three classes present after the move (§K); suite |
|
||||
| G04 Disable knowledge source | **PASS** | suite `test_change_visibility.py`, `test_imported_knowledge.py` |
|
||||
@@ -949,8 +957,11 @@ rather than about M7 once M11 added a migration (§O).
|
||||
|
||||
**Firefox 154.0.1**, headless, driven over W3C WebDriver (geckodriver 0.37.1),
|
||||
against the **built** SPA served by FastAPI — the production path from
|
||||
`DEVELOPMENT.md`, not a Vite dev server. Narrator: `qwen2.5:3b-instruct-16k` on
|
||||
the trusted-LAN Ollama. **38 checks, 0 failed, 0 skipped**, 107 seconds.
|
||||
`DEVELOPMENT.md`, not a Vite dev server. Narrator: `qwen2.5:3b-instruct` on
|
||||
the trusted-LAN Ollama. *Corrected at closeout:* earlier revisions named the
|
||||
`-16k` model, and the run's own report records the plain one. **38 checks, 0
|
||||
failed, 0 skipped**, 107 seconds, on the `ef25b0a` tree. The closeout's runs on
|
||||
the candidate tree are in §S.3.
|
||||
|
||||
The harness is `tools/m11_browser.py` and its WebDriver client is
|
||||
`tools/m11_webdriver.py`, both in the repository — M8's and M9's browser harness
|
||||
@@ -1365,6 +1376,7 @@ defects. A harness that has only ever agreed with itself is not evidence.
|
||||
| **It had no way to see failed post-turn work** *(fixed in `f8d4010`)* | it reported `complete` over 180 `database is locked` errors | it would have certified the run that §O.7 destroyed. It now checks derived status and new `server.log` lines after every turn and stops at the first failure; a run with no memories or no summaries ends `failed` |
|
||||
| **Its protocol-leak count used plain substrings** *(fixed in `96c1bf5`)* | `## Established:` did not match `\nEstablished:\n`, so the count read 1 where 5 turns leaked | the count now matches any state heading with an indented entry, in any markdown |
|
||||
| **Its M04 precondition was the clue's text being out of recent history** *(fixed in `96c1bf5`)* | the narrator reuses that text in its own prose, so a run whose planting turn was 65 depths outside the window read `precondition_not_met` | the precondition is now the planting turn's position, recorded at planting and carried across `--resume`, which is M04's own wording |
|
||||
| **The browser import wait could not fail** *(fixed at closeout)* | it waited for "hidden" in the page text, and the scenario's campaign is titled *Hidden Knowledge*, so the wait held before Import was pressed. The modal check that follows then raced the source row, and skipped once (34/0/1) | "a local file imports through the browser" was recorded without evidence. The harness now waits for the imported source's own row and records its title (§S.3) |
|
||||
|
||||
### Discarded evidence runs
|
||||
|
||||
@@ -1437,7 +1449,16 @@ accepted for v1, or a decision for the owner at acceptance.
|
||||
6. **Model-invented headings are not removed.** A model that writes its own
|
||||
section, such as `Identifiers established:` or `Set of events made true:`,
|
||||
under a heading that is not the renderer's, keeps it in the story. Seen on 10
|
||||
turns of one 26-turn trial.
|
||||
turns of one 26-turn trial. **Widened at closeout.** The identity run found
|
||||
three more shapes left in stored story, on 4 of 10 turns at a 4,096 window:
|
||||
- event-call syntax copied from the state rule, such as
|
||||
`> set_possession(silver-key, "alice")`
|
||||
- a parroted length hint, `[Hard limit: … append the state block well inside
|
||||
the limit.]`
|
||||
- a lone `Scene:` line
|
||||
|
||||
None occurs in the 100-turn evidence runs. The owner chose on 2026-09-14 to
|
||||
carry this as a residual risk rather than change product code (§S.6).
|
||||
|
||||
7. **The GPU inference host dropped its GPU after the evidence run** (§E.1). The
|
||||
cause is not established, and power transients at the uncapped 280 W limit
|
||||
@@ -1447,17 +1468,15 @@ accepted for v1, or a decision for the owner at acceptance.
|
||||
8. **Post-M8 finding D's root cause is unestablished and will stay that way.**
|
||||
The campaign that produced it was destroyed. M11 delivers a diagnostic that
|
||||
can classify the next occurrence, and the detection the finding asked for.
|
||||
**The diagnostic's run results are not in this report.** The section the
|
||||
first revision pointed to was left empty, and the evidence is in the
|
||||
implementer's `m11-evidence/identity-recheck` directory.
|
||||
Its closeout run's results, and what that run still cannot establish, are
|
||||
in §S.5.
|
||||
|
||||
9. **The browser, offline and identity runs predate the last three product
|
||||
commits.** They were re-run on the `ef25b0a` tree: that commit's message
|
||||
records browser 38/0/0, offline 23/0 and the identity diagnostic clean, and the
|
||||
evidence is dated 2026-09-10. `f8d4010`, `0c7316f` and `96c1bf5` change the
|
||||
turn commit, memory retrieval, narration extraction and the state renderer's
|
||||
headings, all of them backend. The backend and frontend suites cover those
|
||||
changes; the black-box runs were not repeated.
|
||||
9. *Closed at closeout.* **The browser, offline and identity runs predated the
|
||||
last three product commits.** They had been re-run on the `ef25b0a` tree:
|
||||
browser 38/0/0, offline 23/0, and the identity diagnostic clean, with the
|
||||
evidence dated 2026-09-10. `f8d4010`, `0c7316f` and `96c1bf5` changed backend
|
||||
behaviour. All three runs were repeated on the release-candidate tree on
|
||||
2026-09-14 (§S).
|
||||
|
||||
10. **Two control-boundary colour pairs are below WCAG 1.4.11** (1.33:1 resting,
|
||||
1.75:1 hover). Reported rather than fixed, because the control is identified
|
||||
@@ -1487,6 +1506,19 @@ accepted for v1, or a decision for the owner at acceptance.
|
||||
existing campaign does not benefit from finding C's fix until someone sets
|
||||
the control.
|
||||
|
||||
16. **The state rule's example is fantasy.** The fixed instruction every campaign
|
||||
receives names `mara`, `silver-key`, `old-abbey` and `aldric`. In the
|
||||
closeout's office-meeting identity run, the 3B narrator proposed giving
|
||||
`silver-key` to Alice. The validator refused it, which is H05 and C06
|
||||
behaving correctly. J01-J03 concern the schema, which is unaffected. A
|
||||
genre-neutral example is post-v1 prompt work.
|
||||
|
||||
17. **With the 3B narrator at a 4,096 window, the state lags the narration.** In
|
||||
the closeout identity run, one proposal in ten applied. The scene was never
|
||||
updated, and a person the narration introduced never became an entity. Every
|
||||
refusal and unparseable block was recorded and shown to the model, as
|
||||
designed. This is risk 2 observed, not a new product defect.
|
||||
|
||||
---
|
||||
|
||||
## Q. Planning and document changes
|
||||
@@ -1507,6 +1539,9 @@ accepted for v1, or a decision for the owner at acceptance.
|
||||
| `planning/archive/milestone-reports/` | M9's and M10's reports moved here. | rotation |
|
||||
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Revised 2026-09-14**: M01-M04 on the complete evidence run; §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R rewritten; the §G.0 addendum removed. | milestone report |
|
||||
| `DEVELOPMENT.md` | The pointer to the removed §G.0 replaced; a new section on logging a GPU inference host during a long run. | developer docs |
|
||||
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Closeout, 2026-09-14**: new §S (exact-tree verification, the identity results) and §T (acceptance record); the 82-test count and §L's narrator corrected; §A, §B, §D.1, §F, §L, §O, §P and §R brought up to date. | milestone report |
|
||||
| `planning/BUILD-MILESTONES.md`, `VERSION.md` (v4.0), `README.md`, `V1-ACCEPTANCE-TESTS.md` §P and §P3, root `README.md` | M11 accepted and the release gate recorded as passed; a post-v1 backlog. | milestone status |
|
||||
| `backend/tools/m11_browser.py` | The G01 import wait, which could not fail. | release-test infrastructure |
|
||||
|
||||
**Revised after this report, in planning package v3.9 (2026-09-14):**
|
||||
`planning/V1-ACCEPTANCE-TESTS.md` gained result blocks for M01-M04 and corrected
|
||||
@@ -1532,7 +1567,7 @@ pinned by a test so a later milestone changes it deliberately.
|
||||
|
||||
**1. Does every REQUIRED FOR V1 acceptance test pass?**
|
||||
|
||||
**Yes.** All 85 pass, with H09 recorded NOT APPLICABLE on the condition its own
|
||||
**Yes.** All 82 pass, with H09 recorded NOT APPLICABLE on the condition its own
|
||||
text states. M01 to M04 pass on the evidence run in §G, and every other REQUIRED
|
||||
test passes on the evidence in §F.
|
||||
|
||||
@@ -1558,8 +1593,9 @@ across the GPU runs (§N, §P).
|
||||
no route, no DNS, first page load, every asset local, campaign creation, state
|
||||
extraction, knowledge import and retrieval, prompt assembly, export, import, the
|
||||
media module inert, and campaigns surviving a container restart. A turn with no
|
||||
model reachable is reported and corrupts nothing. §J. Re-run on the `ef25b0a`
|
||||
tree (§P).
|
||||
model reachable is reported and corrupts nothing. §J. Repeated at closeout on
|
||||
the release-candidate tree, 23 of 23, from an image built with `--no-cache`
|
||||
(§S.4).
|
||||
|
||||
**5. Does branch/memory/summary/scene isolation pass?**
|
||||
|
||||
@@ -1569,7 +1605,7 @@ The evidence run also restored a Save Point across a live memory bank. §G.2.
|
||||
|
||||
**6. Does trusted-LAN HTTPS inference still pass?**
|
||||
|
||||
**Yes, on the CPU reference host's evidence.** The browser run, the identity
|
||||
**Yes, and re-established at closeout on the candidate tree** (§S.3, §S.5). The browser run, the identity
|
||||
diagnostic and the CPU-host campaigns ran against Ollama on a separate physical
|
||||
machine over HTTPS, with a private CA in the application machine's OS trust store,
|
||||
verification on and no bypass. The storyteller stayed loopback-bound. **The final
|
||||
@@ -1598,9 +1634,9 @@ stamp. §K.
|
||||
|
||||
**10. Does the browser pass the release workflow?**
|
||||
|
||||
**Yes, on the `ef25b0a` tree.** 38 checks, 0 failed, 0 skipped, in Firefox
|
||||
154.0.1 against the built SPA served by FastAPI. The commits after it change no
|
||||
frontend file, and the browser run was not repeated. §L, §P.
|
||||
**Yes, on the release-candidate tree.** 38 of 38 in Firefox 155.0.1 at closeout,
|
||||
against the SPA built from that tree and served by FastAPI, after a harness fix
|
||||
(§S.3). Before that, 38 of 38 in Firefox 154.0.1 on `ef25b0a` (§L).
|
||||
|
||||
**11. Are there any unresolved blockers to independent v1 acceptance?**
|
||||
|
||||
@@ -1610,8 +1646,11 @@ was weakened. A reviewer should weigh five things before signing:
|
||||
- M04's recovery ran through authoritative state, not memory retention (§G.4)
|
||||
- the thin real-token headroom (§N)
|
||||
- the narrator restating prompt text (§P)
|
||||
- the identity diagnostic's results being absent from this report (§P)
|
||||
- the black-box runs predating the later commits (§P)
|
||||
- the extractor shapes and the lagging state the closeout's identity run found
|
||||
(§S.6; §P risks 6, 16 and 17)
|
||||
|
||||
*At closeout:* the identity results are now in §S.5, and the black-box runs no
|
||||
longer predate the candidate (§S). The acceptance decision is §T.
|
||||
|
||||
**12. Is the tree safe to commit as the M11 release candidate?**
|
||||
|
||||
@@ -1622,6 +1661,275 @@ was weakened. A reviewer should weigh five things before signing:
|
||||
the Docker image were built for the first revision and not rebuilt since. The
|
||||
remaining commit, this report's revision, is the owner's to sign.
|
||||
|
||||
*At closeout:* the report's revision is committed as `d198806`, and the planning
|
||||
revision as `3652dc6`, both signed. The suites, the production build and a
|
||||
`--no-cache` Docker image were all re-run on `3652dc6` (§S.2). The closeout
|
||||
commit is staged for the owner's signature.
|
||||
|
||||
---
|
||||
|
||||
*Written by the implementer. Not an acceptance record.*
|
||||
## S. Closeout verification on the release candidate (2026-09-14)
|
||||
|
||||
The browser, offline and identity runs above were taken on `ef25b0a`. Three later
|
||||
product commits changed the turn commit, memory retrieval, narration extraction
|
||||
and the state renderer's headings. The closeout repeated all three runs on the
|
||||
exact release-candidate tree, and rebuilt and re-tested everything else there.
|
||||
|
||||
**The 100-turn campaign was not repeated.** The closeout changed no product code,
|
||||
and that run was taken on `96c1bf5`, whose product code the candidate carries
|
||||
unchanged.
|
||||
|
||||
### S.1 The candidate
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **Candidate** | `3652dc6fae903cad379da8efd5086709b5b9c6f5` on `m11-release-validation`, signed by the owner (`%G?` = `G`) |
|
||||
| **Product code** | identical to `96c1bf5`. `d198806` and `3652dc6` change documentation only. No file under `backend/`, `frontend/`, `Dockerfile`, `docker-compose.yml` or the start scripts differs |
|
||||
| **Working tree** | clean at the start, nothing staged, nothing untracked. The one source change during the closeout is the harness fix in §S.3, to `backend/tools/m11_browser.py`, which is neither part of the application nor in the image. No application file was modified during any run |
|
||||
| **Signatures** | every commit from `144406c` to `3652dc6` is signed by the owner |
|
||||
| **Provenance** | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` md5 is `07fde30437134836e2ee875e82a7cd31`, and `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` |
|
||||
| **Dependencies** | `requirements.txt`, `requirements.lock`, `package.json` and `package-lock.json` are unchanged since `1013c94`. The lock's 35 pins match the environment, with no drift. FastAPI 0.141.1, uvicorn 0.52.4, SQLAlchemy 2.0.52, httpx 0.28.1, pydantic 2.13.5, tiktoken 0.14.0, python-multipart 0.0.32; React and React DOM 19.2.7, React Router 7.18.3; Vite 8.1.3 |
|
||||
| **Application machine** | as §E.1, now on Ubuntu 24.04.5 LTS, kernel 7.0.0-31-generic, Docker 29.8.0, and Firefox 155.0.1 headless with geckodriver 0.37.1 |
|
||||
| **Inference** | trusted-LAN: the CPU reference host (§E.1), Ollama 0.33.0 over HTTPS with a private CA in the application machine's OS trust store, verification on, no bypass. `qwen2.5:3b-instruct` (digest `357c53fb659c`), served at a verified 4,096 window. The storyteller listened on loopback only |
|
||||
| **Evidence** | `$HOME/m11-evidence/closeout-3652dc6/` |
|
||||
|
||||
### S.2 Suites and builds
|
||||
|
||||
| Check | Command | Result |
|
||||
| --- | --- | --- |
|
||||
| Backend suite | `.venv/bin/python -m pytest tests/ -q`, with no `AIDND_TEST_*` set | **1,421 passed, 17 skipped, 0 failed**, 786 s |
|
||||
| Frontend suite | `npm test` | **161 of 161**, 14 files |
|
||||
| Lint | `npm run lint` | exit 0; six `only-export-components` warnings, no errors |
|
||||
| Production build | `npm run build`, after deleting `frontend/dist` | 16 files, including `index-CZ7_g_3j.js` (392.69 kB) and `index-g-TAQQWd.css` (47.96 kB) |
|
||||
| Docker image | `docker build --no-cache`, run by `tools/m11_offline.py` | `sha256:c9b872c2f4b701c8df32c58af944a3f456d1f4a6dd45105257ba6739f2382ed0`. `npm ci` and `pip install` ran rather than coming from cache. **The image's SPA is byte-identical, file for file, to the local build**, so the browser runs below exercised what ships |
|
||||
|
||||
### S.3 Browser regression
|
||||
|
||||
`tools/m11_browser.py`, against the SPA from §S.2, served by FastAPI on
|
||||
loopback, with the narrator over trusted-LAN HTTPS:
|
||||
|
||||
| Run | Harness | Result |
|
||||
| --- | --- | --- |
|
||||
| 1 | as committed | **34 passed, 0 failed, 1 skipped**, 110 s. Modal focus containment skipped: "no Delete control found" |
|
||||
| 2 | as committed, unchanged | **38 passed, 0 failed, 0 skipped**, 87 s |
|
||||
| 3 | with the fix below | **38 passed, 0 failed, 0 skipped**, 69 s. **This is the reported result** |
|
||||
|
||||
The skip was a harness defect, not a product one. After Import is pressed, G01
|
||||
waited for the word "hidden" in the page. The scenario's campaign is titled
|
||||
*Hidden Knowledge*, and the page shows that title in its heading, so the wait
|
||||
held at once. "A local file imports through the browser" was therefore recorded
|
||||
without evidence. The modal check that follows could also run before the
|
||||
imported source's row, which carries the Delete control, had rendered.
|
||||
|
||||
The frontend and the harness are unchanged since `ef25b0a`. Every run's database,
|
||||
`ef25b0a`'s included, holds the imported source, so the import itself worked
|
||||
every time. The harness now waits for the source's own row and records its
|
||||
title (`▸ hidden`). Runs 1 and 2 are kept, and only run 3 used the fixed harness.
|
||||
|
||||
Run 3 passed everything §L lists:
|
||||
- two real turns, and the tab title;
|
||||
- the position indicator before and after Undo, and Undo and Redo;
|
||||
- H06, H07, G09 and H04, with hostile narration;
|
||||
- G01, through the real file input;
|
||||
- the dialog's focus, accessible name, contents and Escape;
|
||||
- §38, the context inspector, H10 and the CSP;
|
||||
- the accessibility measurements, with rendered contrast unchanged at 14.57,
|
||||
5.48, 13.57 and 5.88 to 1.
|
||||
|
||||
**Not driven in a real browser by this harness**, on this tree or before:
|
||||
- Retry, Save Point creation and restore, state correction, and the
|
||||
narration-length control;
|
||||
- a failed generation shown to the reader;
|
||||
- the export download, which is exercised only as far as the click (§L).
|
||||
|
||||
Those behaviours are covered on the candidate tree by the backend and component
|
||||
suites, and by the 100-turn campaign through the API (§F, §G.2). The harness was
|
||||
not extended, because the closeout is a regression rerun.
|
||||
|
||||
### S.4 Offline
|
||||
|
||||
`tools/m11_offline.py`: **23 passed, 0 failed**, 28 s. It ran the §S.2 image with
|
||||
`--network none` and a fresh volume, and made the same 23 checks as §J:
|
||||
|
||||
- **Isolation:** no route, and no DNS.
|
||||
- **Page and assets:** first page load; no remote origin; a CSP; every
|
||||
referenced asset local.
|
||||
- **Local operations:** campaign creation; state extraction; knowledge import;
|
||||
prompt assembly; retrieval.
|
||||
- **Failed turn with no model reachable:** reported, with no narration accepted,
|
||||
the player's words kept, state unchanged, and the earlier story intact.
|
||||
- **Export and import:** both work, with no secret in the bundle.
|
||||
- **Media:** the module is inert with no provider.
|
||||
- **Restart:** campaigns survive a container restart.
|
||||
|
||||
### S.5 The identity diagnostic
|
||||
|
||||
**Self-test** (`--scripted --inject`, no model): 7 signals, all STATE DEFECT. At
|
||||
beat 5 the scripted narrator creates a second Alice. `shared_display_name` fires
|
||||
at beat 5, and `duplicate_character_creation` fires at beats 5-10. The detectors
|
||||
fire.
|
||||
|
||||
**Against a real narrator:**
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **Candidate** | `3652dc6` |
|
||||
| **Narrator** | `qwen2.5:3b-instruct` on the CPU reference host over HTTPS. The window was verified at 4,096 on every turn, and the 16,384 budget was capped to it. `max_output_tokens` 500. No embedding model, so memory and summaries were off |
|
||||
| **Fixture** | the companion *Multi-Character Identity Test* (`TEST-CAMPAIGN-FIXTURE.md` Appendix A): Bill, the protagonist, with Alice, Roger and John in one meeting room; ten beats |
|
||||
| **Fixture accepted** | yes. The opening correction applied all six events and refused none, and the scene named its location and who was present. The harness stops on either failure rather than running a degraded campaign |
|
||||
| **Turns** | 10 of 10 accepted, in 18 min 46 s |
|
||||
| **Result** | **0 signals. Verdict: "no objective identity defect detected."** |
|
||||
| **Evidence** | `identity/stdout.txt`; the final state, prompt, settings and an M9 bundle in `identity/turn-99/` |
|
||||
|
||||
**What the state did.** The narrator made ten state proposals:
|
||||
- one applied (beat 3: John present, in the meeting room);
|
||||
- five carried no block;
|
||||
- three were unparseable: two cut off by the output limit, and one a parroted
|
||||
reminder;
|
||||
- one was refused: `set_possession` of `silver-key`, which does not exist.
|
||||
|
||||
The scene the fixture set was never updated, so Roger's exit went unrecorded.
|
||||
From beat 6 the narration introduced a fifth person, "Mike", who never became an
|
||||
entity, because the proposal that would have created him was cut off.
|
||||
|
||||
**What the run establishes:**
|
||||
- the fixture is valid and was not silently refused;
|
||||
- the authoritative state held four distinct people, with no shared display name
|
||||
and no drift in the protagonist;
|
||||
- the prompt's state section agreed with the document on every turn;
|
||||
- the detectors fire when a confusion is planted;
|
||||
- what a future occurrence needs is preserved and portable.
|
||||
|
||||
**What it cannot establish:**
|
||||
- **The root cause of post-M8 finding D.** No same-name confusion occurred, so
|
||||
the finding was not reproduced.
|
||||
- **Prose-level misattribution.** This is not graded. The narration is kept for a
|
||||
person to read.
|
||||
- **A stale scene.** The objective checks cover only the protagonist's presence,
|
||||
so Roger's unrecorded exit and the invented Mike raise no signal.
|
||||
- **Derived contamination.** Memory and summaries were off.
|
||||
- **Behaviour at a 16,384 window, or with another narrator.**
|
||||
|
||||
A clean verdict over a state that lags its narration is not evidence that
|
||||
identities stay straight under a small model. What the diagnostic does show is
|
||||
that, when identities go wrong, it can say whether the state, the context or the
|
||||
derived data carried the error.
|
||||
|
||||
### S.6 Findings
|
||||
|
||||
**Release blockers: none.**
|
||||
|
||||
**Non-blocking, found at closeout:**
|
||||
|
||||
1. **Protocol shapes the extractor does not remove.** On 4 of 10 stored AI turns
|
||||
in the identity run (depths 10, 12, 14 and 20), the narration kept event-call
|
||||
syntax (`> Create_entity(...)`, `> set_possession(silver-key, "alice")`), a
|
||||
lone `Scene:` line, and a parroted length hint ending "append the state block
|
||||
well inside the limit.]". Replaying those four texts through the candidate's
|
||||
`narrative.extract.split` leaves all four unchanged.
|
||||
- **Why it survives.** The bracket survives because `_is_echoed_instruction`
|
||||
recognises the reminder by "state block" together with "events list", and
|
||||
the length hint names only the first.
|
||||
- **Scope.** The 100-turn runs on `0c7316f` and `96c1bf5` store 0 turns in
|
||||
these shapes.
|
||||
- **Why it is not a blocker.** No REQUIRED FOR V1 test names stored-narration
|
||||
purity, and `TECHNICAL-DESIGN.md` §15.4 holds that removing story is worse
|
||||
than leaving protocol.
|
||||
- **Decision.** The owner chose on 2026-09-14 to carry it as a residual risk
|
||||
(§P risk 6) rather than change product code. A code change would have
|
||||
invalidated the 100-turn evidence. §O.8's "every shape observed" was true
|
||||
of the shapes observed when it was written.
|
||||
2. **The state rule's example is fantasy** (§P risk 16).
|
||||
3. **The state lagged the narration** at a 4,096 window with the 3B narrator
|
||||
(§P risk 17).
|
||||
|
||||
**Harness defect, fixed:** the G01 import wait (§S.3, §O).
|
||||
|
||||
**Corrections to this report:** the REQUIRED FOR V1 count is 82 (§A, §R), and the
|
||||
browser narrator in §L was `qwen2.5:3b-instruct`.
|
||||
|
||||
### S.7 Release-shaped smoke test
|
||||
|
||||
This ran after the documentation changes above. It started from the Dockerfile,
|
||||
not from a development server. The script is not a repository harness, and its
|
||||
evidence is `smoke-run2/`. The one source difference from the candidate is the
|
||||
§S.3 harness fix, which is not in the image. **14 passed, 0 failed:**
|
||||
|
||||
- **Build and start:** the image builds with `--no-cache`; the container starts
|
||||
on a fresh volume; it is published on host loopback only
|
||||
(`127.0.0.1:18080`).
|
||||
- **Page:** the first page loads, with every asset it names served locally.
|
||||
- **Endpoint policy:** the approved trusted-LAN endpoint verifies over HTTPS,
|
||||
with the private CA supplied to the container through the host's trust bundle,
|
||||
mounted read-only. There is no bypass. A public endpoint is refused with
|
||||
HTTP 400, and the approved endpoint stays configured.
|
||||
- **Play:** a campaign is created, and one normal story turn is accepted (24 s,
|
||||
`qwen2.5:3b-instruct`, window 4,096).
|
||||
- **Persistence:** after a container restart, the campaign reopens with the same
|
||||
actions and the same state.
|
||||
- **Browser:** Firefox 155.0.1 loads the reopened campaign and shows its story,
|
||||
with the tab reading "Closeout Smoke — Interactive Story".
|
||||
|
||||
The first attempt (`smoke/`) passed its first 13 checks. It then stopped on a
|
||||
defect in the smoke script itself (`browser.title` is a property), before the
|
||||
browser check was recorded. It is kept, and superseded by the run above.
|
||||
|
||||
---
|
||||
|
||||
## T. Acceptance record
|
||||
|
||||
**Decision: PASS. M11 is accepted at closeout, 2026-09-14.**
|
||||
|
||||
> Does this exact build meet the v1 black-box acceptance contract, and can it be
|
||||
> packaged as the first production release?
|
||||
|
||||
**Yes.** Every REQUIRED FOR V1 condition remains satisfied. No test was waived,
|
||||
relaxed or reclassified.
|
||||
|
||||
| Area | Status | Evidence |
|
||||
| --- | --- | --- |
|
||||
| All REQUIRED FOR V1 tests | 82: 81 PASS, H09 NOT APPLICABLE | §F |
|
||||
| M01-M04 | PASS, on `96c1bf5`, whose product code the candidate carries unchanged | §G |
|
||||
| Offline / local-only operation | PASS, on the candidate | §S.4 |
|
||||
| Trusted-LAN inference (A06) | PASS, on the candidate: HTTPS, private CA, verification on, storyteller on loopback | §S.3, §S.5 |
|
||||
| Undo / Redo / Retry / Save Point | PASS: D01-D14 in the long run and the suite; Undo and Redo in a real browser on the candidate | §F, §G.2, §S.3 |
|
||||
| Branch, memory, summary and scene isolation | PASS: E01-E04, in the candidate's backend suite | §H, §S.2 |
|
||||
| State atomicity | PASS: L01, in the long run and in the candidate's container | §F, §S.4 |
|
||||
| Export / import / recovery | PASS: I01-I07; the long campaign onto a clean data directory, 16 of 16 | §K |
|
||||
| Fresh-install vs upgraded schema parity | PASS: `test_m11_migration.py`, in the candidate's suite | §K, §S.2 |
|
||||
| Fantasy fixture | PASS: the long-run campaign | §G, §I |
|
||||
| Science-fiction fixture | PASS: `test_m11_scifi.py`, in the candidate's suite | §I, §S.2 |
|
||||
| Security and local endpoint policy | PASS: H01-H12 | §F, §J, §S.4 |
|
||||
| Browser release workflow | PASS: 38 of 38 on the candidate | §S.3 |
|
||||
|
||||
**Qualifications that stand:**
|
||||
|
||||
- M04's fact was recovered through authoritative state, not independent memory
|
||||
retention. The owner accepted that on 2026-09-13.
|
||||
- K04, a SHOULD test, passes on its deferred branch.
|
||||
- The export download is exercised in the browser only as far as the click.
|
||||
- Retry, Save Point, state correction, narration length and failed generation are
|
||||
proved through the API and the component suite, not driven in a real browser.
|
||||
- The real-token headroom at the largest long-run prompt is 42 tokens (§N).
|
||||
- A06's HTTPS evidence ran at a 4,096 window on the CPU host, and the long run
|
||||
used plain HTTP to a LAN GPU host.
|
||||
- The residual risks in §P remain residual.
|
||||
|
||||
**Four separate events:**
|
||||
|
||||
| Event | State |
|
||||
| --- | --- |
|
||||
| M11 accepted | recorded here, 2026-09-14; takes effect with the owner's signed closeout commit |
|
||||
| Release candidate verified | done, 2026-09-14, on `3652dc6` (§S) |
|
||||
| Release commit signed | not yet: the closeout commit is staged for the owner |
|
||||
| `v1.0.0` tagged | not done. The owner's decision, and the tag must point at the signed closeout commit |
|
||||
|
||||
**Where this report stays.** `planning/reports/` holds the most recently
|
||||
completed milestone's report. That report moves to the archive when the next
|
||||
milestone's report is written. There is no next milestone, so this report stays
|
||||
where it is. No new convention was invented for it.
|
||||
|
||||
---
|
||||
|
||||
*§A-§R written by the implementer. §S and §T added at the M11 closeout,
|
||||
2026-09-14; §T is the acceptance record.*
|
||||
|
||||
Reference in New Issue
Block a user