Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.
V1.1 RELEASE VALIDATION: PASS
What was run, on this candidate:
- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
identical on all 15 census fields, schema parity at user_version 94, and both
bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
verified, public endpoint refused, a real turn, restart, persistence, and
Firefox rendering the reopened campaign.
Carried residuals, stated rather than summarised away:
- WP-B: deterministic independent-memory recovery PASS; reference-model
independent-memory recovery FAIL at memory creation — the owner-accepted
limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
backlog, reproduced and not fixed during validation.
Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.
Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.
Still the owner's to do: sign the release commit, update main, tag v1.1.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
882 lines
47 KiB
Markdown
882 lines
47 KiB
Markdown
# Adventure Storyteller v1.1 — Integrated Release Validation
|
||
|
||
**Status:** COMPLETE — **V1.1 RELEASE VALIDATION: PASS**. The decision, and what
|
||
it deliberately does not cover, is in §W.
|
||
|
||
This report answers one question: **does this exact candidate preserve the
|
||
complete v1 contract and satisfy every accepted v1.1 package on one integrated
|
||
release tree?** It is release validation, not a work package. Nothing here adds
|
||
a feature, and no release tag is created by it.
|
||
|
||
---
|
||
|
||
## A. Repository / provenance
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| **Candidate SHA** | **`87a40326a29533c8d52c9f9f41022e7b499b1de7`** |
|
||
| Branch | `v1.1-development`, up to date with `origin/v1.1-development` |
|
||
| Working tree at freeze | **clean** — nothing modified, nothing staged |
|
||
| Commit | *v1.1: harden recovery and control boundaries* (WP-D + WP-E) |
|
||
| **Owner signature** | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, made 2026-09-16 05:37:13 EDT |
|
||
| Tag at HEAD | **none** — no `v1.1.0` tag exists |
|
||
| v1.0.0 baseline | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, **an ancestor** |
|
||
| Package ancestry | `d63804f` (WP-A1/A2), `beb17ad` (WP-B.1), `0c1ba83` (WP-B.2), `59b5ebc` (WP-C) — **all ancestors** |
|
||
| Diff v1.0.0..HEAD | 61 files, +14,833 / −366 |
|
||
| LICENSE / PROVENANCE | **unchanged since v1.0.0** (empty diff) |
|
||
|
||
### A.1 Frozen candidate identity
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Dependency locks | `backend/requirements.txt` `sha256:ed28bc0f8970cf4e…`, `frontend/package-lock.json` `sha256:355cb370837ade01…`, `frontend/package.json` `sha256:2016580ddfa176a9…`, `backend/requirements-dev.txt` `sha256:06d7695816b201e9…` |
|
||
| Schema | `LATEST_VERSION` **94**, 93 migrations (`PRAGMA user_version`) |
|
||
| Bundle format | **`ai-dnd-adventure-v3`** |
|
||
| Import ceiling | 20 MB (`MAX_IMPORT_BODY_BYTES`), unchanged |
|
||
| Frontend build | `dist` built 2026-09-16T05:39:55, 16 files, `sha256(dist) = ea2753ad24f61959fe084f4674911acc` |
|
||
| Docker image | `storyteller:release-87a4032`, `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398`, 312 MB |
|
||
| Firefox / geckodriver | 155.0.1 / 0.37.1 (2026-09-04) |
|
||
| Docker | 29.8.0, build 88096ef |
|
||
| CPU HTTPS reference host | Ollama **0.33.0**, serving `qwen2.5:3b-instruct`, `qwen2.5:3b-instruct-16k`, `nomic-embed-text:latest`; certificate verifies through the machine's CA store with no bypass |
|
||
| GPU inference host | Ollama **0.34.0**; `qwen2.5:3b-instruct-16k` digest `21ff8cc52f375f19`, `nomic-embed-text:latest` digest `0a109f422b47e3a3`; no model resident at start |
|
||
| Evidence root | `$HOME/v11-evidence/release-87a4032/` — never `/tmp`, and no real hostname appears in any committed file |
|
||
|
||
---
|
||
|
||
## B. Package acceptance inventory
|
||
|
||
| Package | Status | Source |
|
||
| --- | --- | --- |
|
||
| **WP-A1** context-window safety reserve | **ACCEPTED** | `V1.1-WP-A1-A2-REPORT.md`, signed `d63804f` |
|
||
| **WP-A2** protocol-echo cleanup, genre-neutral prompting | **ACCEPTED** | same report and commit |
|
||
| **WP-B** independent memory | **ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION** | `V1.1-WP-B1-REPORT.md`, `V1.1-WP-B2-REPORT.md` §S, signed `beb17ad` / `0c1ba83` |
|
||
| **WP-C** browser release coverage | **ACCEPTED** | `V1.1-WP-C-REPORT.md`, signed `59b5ebc` |
|
||
| **WP-D** recovery honesty | **ACCEPTED** | `V1.1-WP-D-REPORT.md`, signed `87a4032` |
|
||
| **WP-E** control-boundary contrast | **ACCEPTED** | `V1.1-WP-E-REPORT.md`, signed `87a4032` |
|
||
|
||
### B.1 WP-B's qualification, carried whole
|
||
|
||
The WP-B disposition is **not** shortened to "WP-B passed" anywhere in this
|
||
report. Its own §S records:
|
||
|
||
```text
|
||
B2.1 RANKING: PASS
|
||
B2.2 EVICTION: PASS
|
||
B2.3 EXCERPT CREATION: PASS
|
||
B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED
|
||
|
||
DETERMINISTIC WP-B: PASS
|
||
REAL-MODEL WP-B: FAIL
|
||
|
||
WP-B OVERALL:
|
||
ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION
|
||
```
|
||
|
||
In the release contract's own words (§11 items 12), that is:
|
||
|
||
```text
|
||
DETERMINISTIC WP-B: PASS
|
||
REFERENCE-MODEL INDEPENDENT MEMORY: FAIL
|
||
OWNER ACCEPTED THE LIMITATION FOR v1.1
|
||
```
|
||
|
||
The failing stage is **memory creation** — the summariser's content selection —
|
||
not ranking, eviction or injection, each of which passes deterministically.
|
||
|
||
### B.2 WP-E screenshot approval — a correction of record
|
||
|
||
The committed WP-E report read `OWNER SCREENSHOT APPROVAL: PENDING`, because it
|
||
was written before the owner reviewed the images. The owner's release-validation
|
||
brief (2026-09-16) states the before/after screenshots are approved and
|
||
instructs this validation to record it. The WP-E report is updated to
|
||
`APPROVED` as part of this closeout (§V), sourced to that brief and dated. No
|
||
visual code changed during release validation, so the approval stands (§Q).
|
||
|
||
---
|
||
|
||
## C. v1 acceptance matrix
|
||
|
||
Every test marked **REQUIRED FOR V1** — there are **82** — against evidence taken
|
||
on **this candidate**. Evidence types follow M11's: `browser` (the 101-check run,
|
||
§G), `campaign` (the 102-turn integrated run, §H), `container` (the offline run
|
||
on the candidate image, §F), `process` (spawned server processes — recovery §M,
|
||
upgrade §N), `suite` (the 1,723-test backend suite, §D). No historical result
|
||
from different product code is used where the contract asks for candidate
|
||
evidence.
|
||
|
||
**Result: 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified.**
|
||
|
||
### A — Local-first operation
|
||
|
||
| ID | Result | Evidence on this candidate |
|
||
| --- | --- | --- |
|
||
| A01 Start application offline | **PASS** | container: first page load, fresh volume, no route and no DNS |
|
||
| A02 Storyteller loopback default | **PASS** | suite; every harness reached it on `127.0.0.1`; `docker-compose.yml` publishes `127.0.0.1:8000:8000` |
|
||
| A03 No cloud API key | **PASS** | suite; container: no secret in an export |
|
||
| A04 Campaign survives restart | **PASS** | campaign: **3 process restarts, 4 process starts**, state compared across each; container: campaigns survive a container restart |
|
||
| A05 Failed model call does not corrupt story | **PASS** | campaign: a real `failed_call` at turn 69 against an unserved model, play resumed; container: same with no model reachable |
|
||
| A06 Trusted-LAN Ollama inference | **PASS** | browser: the whole 101-check run over **trusted-LAN HTTPS with a private CA**, verification on, no bypass, storyteller loopback-bound |
|
||
|
||
### B — Core play
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| B01 Natural language action | **PASS** | browser (real turns through the UI) + campaign (102 accepted) |
|
||
| B02 Dialogue input | **PASS** | campaign: dialogue beats in the turn list |
|
||
| B03 Continue | **PASS** | suite; browser: the Continue control present and enabled |
|
||
|
||
### C — Story authority and state
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| C01 Campaign canon is preserved | **PASS** | campaign: canon present in the prompt on **102 of 102** turns |
|
||
| C02 Possession state | **PASS** | campaign (the silver key) + suite |
|
||
| C03 Character knowledge is not invented | **PASS** | suite |
|
||
| C04 Manual state correction | **PASS** | campaign: **2 state corrections**; browser: C3's accepted and refused corrections; suite |
|
||
| C05 Canon beats reference | **PASS** | suite |
|
||
| C06 Structured state matches accepted narrative consequence | **PASS** | campaign: real extraction across 102 turns, every event validated or refused; suite |
|
||
|
||
### D — Non-destructive history
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| D01 Undo one turn | **PASS** | browser + campaign (`undo`) |
|
||
| D02 Minimum five undos | **PASS** | suite; campaign (`undo_redo`) |
|
||
| D04 Redo | **PASS** | browser + campaign |
|
||
| D05 Redo invalidated by new continuation | **PASS** | campaign: `diverged`, after which Redo is gone |
|
||
| D06 Retry narrator response | **PASS** | campaign: **2 retries** |
|
||
| D07 Select prior retry take | **PASS** | campaign: `take_selected` |
|
||
| D08 Retry does not delete prior take | **PASS** | campaign + suite |
|
||
| D09 Edit earlier user input | **PASS** | suite |
|
||
| D10 Edit narrator output | **PASS** | suite; browser (hostile-Markdown plants through the narrator-edit path) |
|
||
| D11 Named checkpoint | **PASS** | campaign: **2 Save Points**; browser; suite |
|
||
| D12 Restore checkpoint | **PASS** | campaign: `save_point_restored`; browser |
|
||
| D13 Restore does not delete later history | **PASS** | campaign: retained actions after the restore; recovery §M |
|
||
| D14 Delete checkpoint | **PASS** | suite; browser: the delete confirmation dialog |
|
||
|
||
### E — Branch and derived-data isolation
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| E01 Abandoned future cannot affect active state | **PASS** | suite `test_m11_leakage.py`, with a positive control |
|
||
| E02 Abandoned memory cannot leak | **PASS** | as above |
|
||
| E03 Abandoned summary cannot leak | **PASS** | as above |
|
||
| E04 Scene state is lineage-safe | **PASS** | as above |
|
||
|
||
### F — Long-term memory and context
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| F01 Recent turns remain coherent | **PASS** | campaign: history populated every turn, newest always included |
|
||
| F02 Old important event retrieval | **PASS** | campaign §K: the planting turn outside the window and the fact recovered — **through authoritative state**, not independent memory (§K states which) |
|
||
| F03 Prompt remains bounded | **PASS** | campaign: 1,602–14,982 tokens against a 16,384 budget across 102 turns |
|
||
| F04 Output token reserve | **PASS** | campaign: `output_reserve` 500 present and subtracted on every turn |
|
||
| F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt |
|
||
| F06 Retrieval provenance | **PASS** | campaign: knowledge and memory provenance per turn; suite |
|
||
| F07 Heuristic memory is not canon | **PASS** | suite |
|
||
| F08 Memory failure is non-fatal | **PASS** | suite; container: derived work fails with no model and turns still commit; campaign: **0 post-turn failures**, no database-lock errors |
|
||
|
||
### G — Imported knowledge
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| G01 Import local text | **PASS** | browser: through the real file input; container: offline |
|
||
| G02 Import local Markdown | **PASS** | campaign: **3 sources imported**; browser |
|
||
| G03 Classification | **PASS** | campaign: all three classes; recovery §M confirms them after a move |
|
||
| G04 Disable knowledge source | **PASS** | suite |
|
||
| G05 Canon retrieval | **PASS** | campaign: canon passages in stored prompts; suite |
|
||
| G06 Reference retrieval | **PASS** | suite |
|
||
| G07 Inspiration is low authority | **PASS** | suite |
|
||
| G08 No automatic URL fetch | **PASS** | suite `test_egress.py`; container: no network at all, import still works |
|
||
| G09 Remote Markdown image does not auto-load | **PASS** | browser: no remote image src in the rendered story |
|
||
| G10 Prompt injection in source is treated as data | **PASS** | suite; browser: injection text rendered as text |
|
||
|
||
### H — Security
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| H01 No unexpected outbound connections | **PASS** | container (no network at all) + suite `test_egress.py` |
|
||
| H02 No telemetry | **PASS** | suite |
|
||
| H03 No cloud provider required | **PASS** | container: a full campaign offline |
|
||
| H04 Model output cannot execute shell | **PASS** | browser + suite |
|
||
| H05 Invalid state event rejected | **PASS** | suite; campaign: refusals recorded |
|
||
| H06 Stored XSS protection | **PASS** | browser: `onerror` and `<script>` in accepted narration, neither executed |
|
||
| H07 JavaScript URL protection | **PASS** | browser: no `javascript:` href in the DOM |
|
||
| H08 Path traversal import rejected | **PASS** | suite |
|
||
| **H09 ZIP Slip protection** | **NOT APPLICABLE** | the product extracts no archives, and a test enforces it — the same condition v1 recorded |
|
||
| H10 Restrictive CORS and local API behaviour | **PASS** | browser: an unknown API path is a 404 with a non-HTML body; suite |
|
||
| H11 No first-use runtime asset download | **PASS** | container: every asset local with no network; suite: the tokenizer table is vendored |
|
||
| H12 Inference endpoint enforcement | **PASS** | suite; smoke §R: a public endpoint refused by the running image; the container's refusal of an unresolvable name is the same policy (§R) |
|
||
|
||
### I — Export, import and recovery
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| I01 Export campaign | **PASS** | campaign: the 102-turn campaign exported (3,071,683 bytes) |
|
||
| I02 Import exported campaign | **PASS** | process §M: imported into a database and directory that never existed |
|
||
| I03 Branch/disposable history export | **PASS** | §M: **88 actions retained beyond the active line** after the move |
|
||
| I04 Checkpoint export | **PASS** | §M: both Save Points restore after the move |
|
||
| I05 Knowledge provenance export | **PASS** | §M: all three classes with their content |
|
||
| I06 Database/export contains no API secrets | **PASS** | §M: the bundle carries no secret; suite; container |
|
||
| I07 Export/import preserves an undone active head | **PASS** | §N: the v1.0.0 campaign's undone head and `can_redo` survive upgrade and both bundle directions; suite `test_m9_portability.py`. *(The long-run bundle ended head-at-tip, so §M exercises the other case — stated in §M rather than implied)* |
|
||
|
||
### J — Genre neutrality
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| J01 Science-fiction campaign | **PASS** | suite `test_m11_scifi.py` |
|
||
| J02 Generic entity support | **PASS** | as above |
|
||
| J03 Genre profiles are configuration | **PASS** | as above |
|
||
|
||
### K — Future media architecture
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| K01 Scene snapshot exists | **PASS** | suite; container: a scene packet builds offline |
|
||
| K02 Visual character profile | **PASS** | suite |
|
||
| K03 Visual location profile | **PASS** | suite |
|
||
|
||
### L — Data integrity
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| L01 Atomic turn commit | **PASS** | container + campaign: a real induced failure, no narration accepted, no half-written state |
|
||
| L02 State reconstruction | **PASS** | suite; campaign: state compared across 3 restarts |
|
||
| L03 Checkpoint reconstruction after restart | **PASS** | campaign + §M |
|
||
|
||
### M — Long-run
|
||
|
||
| ID | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| **M01** 100-turn campaign | **PASS** | **102 accepted turns**, every scheduled operation exercised, **0 post-turn failures** (§H) |
|
||
| **M02** Restart during long campaign | **PASS** | **3 genuine process restarts** (4 process starts); everything crossed as bytes on disk |
|
||
| **M03** Long-run context stability | **PASS** | the prompt held **13,492–14,982** tokens over the last 70 turns against a 16,384 budget; window verified **102/102**; canon present **102/102** |
|
||
| **M04** Long-run memory recall | **PASS**, qualified | the planted clue was outside the history window (planted depth 1, floor 72) and reached the prompt: `m04_verdict: **recovered_through_state_only**`. Recovery was **through authoritative state**, not independent memory — §K states this distinction and does not relabel it |
|
||
|
||
### SHOULD and FUTURE
|
||
|
||
Not counted as REQUIRED. **SHOULD:** B04, D03, K04, L04 — all still pass on the
|
||
candidate (suite; browser for B04's direction toggle). **FUTURE:** K05, K06 —
|
||
deliberately not run; both need a media provider this release does not build.
|
||
|
||
## D. Backend / frontend suites
|
||
|
||
| Suite | Result |
|
||
| --- | --- |
|
||
| **Backend** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,210.6 s) |
|
||
| **Frontend** (`npm test`) | **175 passed**, 15 files, 0 failed |
|
||
| **Lint** (`npm run lint`, oxlint) | **exit 0 — 0 errors**, 15 warnings |
|
||
| **Production build** (`npm run build`) | succeeded |
|
||
|
||
**Every skip explained — one category, and it is the expected one.** All 17 are
|
||
environment-gated real-model tests, skipped because `AIDND_TEST_*` is
|
||
deliberately unset for the deterministic suite:
|
||
|
||
| File | Skipped | Gate |
|
||
| --- | --- | --- |
|
||
| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
|
||
| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` + `AIDND_TEST_MODEL` |
|
||
| `test_narrative_realistic.py` | 3 | same |
|
||
| `test_m11_real_window.py` | 3 | same; one needs `AIDND_TEST_WIDE_MODEL` |
|
||
| `test_provider_wiring.py` | 1 | same |
|
||
|
||
**0 xfailed.** No deficiency WP-B fixed remains parked as an expected failure —
|
||
B.1's two strict xfails became ordinary passes in B.2 and stayed that way.
|
||
|
||
**Lint warnings are the documented, unchanged set:** the pre-existing
|
||
`only-export-components` and unused-import warnings recorded at WP-C, WP-D and
|
||
WP-E. None is in a file this candidate changed relative to those packages.
|
||
|
||
## E. Production build and Docker image
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Command | the repository's documented production build (`DEVELOPMENT.md`), with `--no-cache` |
|
||
| Image | `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398` (312 MB) |
|
||
| Dependencies | **installed, not reused** — `npm ci` and `pip install --no-cache-dir` both executed in the log |
|
||
| `CACHED` steps | **2**, and both are `WORKDIR` metadata (`/build`, `/app`) — no dependency or source layer was cached |
|
||
| **Image SPA vs local build** | **file-for-file identical**: 16 files each, `diff -r` clean, combined `sha256 = ea2753ad24f61959fe084f4674911acc` on both sides |
|
||
|
||
The image is built from the candidate tree for this validation. No earlier
|
||
work-package image was reused.
|
||
|
||
## F. Offline / no-network
|
||
|
||
`tools/m11_offline.py` against **the candidate image**, `--network none`, fresh
|
||
volume. Evidence: `…/release-87a4032/offline/offline-report.json`.
|
||
|
||
**23 checks, 23 passed, 0 failed.** Including: no route to the public Internet;
|
||
no external DNS; first page load; no remote origin named; CSP served; every
|
||
referenced asset local; campaign creation; state extraction; local file import;
|
||
prompt assembly; knowledge search; a turn with no model reachable reported as a
|
||
failure with no narration accepted, the player's words kept and state unchanged;
|
||
export; import with state; no secret in the export; the media module inert with
|
||
no provider; campaigns surviving a container restart.
|
||
|
||
**No first-use download occurred**, which is what the container's absent network
|
||
makes unfalsifiable rather than merely unobserved.
|
||
|
||
## G. Browser release validation
|
||
|
||
`tools/m11_browser.py` on the candidate, production build served by FastAPI,
|
||
Firefox 155.0.1 / geckodriver 0.37.1, narrator over **trusted-LAN HTTPS with the
|
||
private CA** (`endpoint_class: trusted-LAN HTTPS`), storyteller on loopback.
|
||
`kind: release regression` — not partial, not `--only`. 693 s.
|
||
|
||
| Suite | Passed | Failed | Skipped |
|
||
| --- | --- | --- | --- |
|
||
| **M11** (the v1 release regression) | **38** | **0** | **0** |
|
||
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
|
||
| **WP-E** (control boundaries) | **10** | **0** | **0** |
|
||
| **Total** | **101** | **0** | **0** |
|
||
|
||
**A1 accounting:** 8 narrator turns, **`fits` on every one**; `turns_not_clean`
|
||
**0**; `protocol_shapes_in_narration` **0**. The verified window on this host was
|
||
4,096 (`source: loaded`); the 16,384 evidence is the long run's (§I).
|
||
|
||
WP-C's proofs are all present in the 53: Retry and alternate takes, Save Point
|
||
create/restore/Redo, state correction including a refusal shown as a refusal,
|
||
narration length reaching the prompt, failed generation and recovery, and **real
|
||
export downloads** from both the library and campaign settings, each file landing
|
||
on disk and importing into a fresh application. WP-E's ten are the rendered
|
||
boundary measurements of §Q.
|
||
|
||
## H. Integrated 100+ turn long run
|
||
|
||
One new campaign on the candidate's product code, GPU inference host, with the
|
||
owner's power/link/kernel logging running before the first turn.
|
||
Evidence: `…/release-87a4032/long-run/`.
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Narrator | **`qwen2.5:3b-instruct-16k`**, digest `21ff8cc52f375f19` |
|
||
| Embeddings | **`nomic-embed-text`**, digest `0a109f422b47e3a3` |
|
||
| Window | **16,384**, `window_verified` on **102 of 102** turns |
|
||
| Memory / summaries | **on** — 19 memories in the bank, **12 summaries** written |
|
||
| **Accepted turns** | **102** (target 100) |
|
||
| Restarts | **3** genuine process restarts, 4 process starts |
|
||
| Elapsed | 1,330 s |
|
||
| Export | 3,071,683 bytes; database 2,613,248 bytes |
|
||
| Status | `complete`; `aborted_reason` null, `failed_reason` null |
|
||
|
||
**Not a repeated-turn benchmark.** Every scheduled operation fired and is in the
|
||
timeline: 3 restarts, 1 undo, 1 undo→redo, 2 retries, 1 take selection, 2 Save
|
||
Points, 1 Save Point restore, 1 divergence, 2 state corrections, 3 knowledge
|
||
imports, memory activation, 1 deliberate failed call, 1 export, the planted clue
|
||
and the planted independent fact, and the recall probe.
|
||
|
||
### H.1 M01–M04
|
||
|
||
| | Verdict | What decides it |
|
||
| --- | --- | --- |
|
||
| **M01** 100-turn campaign | **PASS** | 102 accepted turns, each with a committed action and state document |
|
||
| **M02** Restart during long campaign | **PASS** | 3 genuine `uvicorn` restarts; everything that survived crossed as bytes on disk |
|
||
| **M03** Long-run context stability | **PASS** | prompt 13,492–14,982 tokens over the last 70 turns against a 16,384 budget; canon present 102/102; window verified 102/102 |
|
||
| **M04** Long-run memory recall | **PASS**, and qualified | the clue was planted at depth 1, the history floor reached depth 72, and it was **not** in the recent window; it reached the prompt through **authoritative state**. `m04_verdict: recovered_through_state_only`. §K keeps the distinction the criterion was written around |
|
||
|
||
### H.2 State and derived-work integrity
|
||
|
||
**0 post-turn failures across 102 turns, and no `database is locked` error.** The
|
||
only failure-shaped events in the whole timeline are the two the run creates on
|
||
purpose: the scheduled `failed_call` at turn 69 (a model name the server does not
|
||
serve — A05/L01 evidence) and the independent-fact precondition notes (§K).
|
||
Narrative-state proposals were recorded, applied or refused as designed across
|
||
the run, and no accepted narration carried an unresolved protocol block (§J).
|
||
|
||
## I. Context-window / A1 evidence
|
||
|
||
**Every turn `fits`.** Across all 102 accepted turns the accounting status was
|
||
`fits` — **0 `exceeded`, 0 `truncation_suspected`** — and the window was verified
|
||
at 16,384 on every one.
|
||
|
||
**The ten largest stored prompts**, re-counted against what the server itself
|
||
reported:
|
||
|
||
| Turn | App estimate | Server count | Difference | Window | Output reserve | Safety reserve | Observed margin | Status |
|
||
| ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | :--- |
|
||
| 86 | 14,990 | 15,005 | +15 | 16,384 | 500 | 820 | 879 | fits |
|
||
| 83 | 14,978 | 14,993 | +15 | 16,384 | 500 | 820 | 891 | fits |
|
||
| 97 | 14,965 | 14,980 | +15 | 16,384 | 500 | 820 | 904 | fits |
|
||
| 41 | 14,960 | 14,975 | +15 | 16,384 | 500 | 820 | 909 | fits |
|
||
| 66 | 14,951 | 14,966 | +15 | 16,384 | 500 | 820 | 918 | fits |
|
||
| 33 | 14,941 | 14,956 | +15 | 16,384 | 500 | 820 | 928 | fits |
|
||
| 35 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
|
||
| 77 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
|
||
| 38 | 14,927 | 14,942 | +15 | 16,384 | 500 | 820 | 942 | fits |
|
||
| 45 | 14,922 | 14,937 | +15 | 16,384 | 500 | 820 | 947 | fits |
|
||
|
||
**The reserve is preserved on every re-counted prompt.** The estimate runs
|
||
exactly **15 tokens below** the server's own count on all ten — a constant,
|
||
known offset rather than drift — and the smallest observed margin anywhere in the
|
||
run is **879 tokens**, against the documented safety reserve of 820.
|
||
|
||
**Against v1.** M11's closeout recorded a remaining margin of **23–42 tokens**.
|
||
The same measurement on this candidate is **879 at its tightest** — roughly
|
||
twenty to thirty times the headroom, which is what WP-A1 was for.
|
||
|
||
## J. Protocol-leak / A2 evidence
|
||
|
||
**Release-gate result: 0.** Across the run's **105 stored AI actions**,
|
||
`protocol_leaks` reports **0 leaking**, `example_ids` empty.
|
||
|
||
Measured separately, by the detector's own four rules:
|
||
|
||
| Shape | Count in this run |
|
||
| --- | --- |
|
||
| State/section heading with an indented entry | **0** |
|
||
| Event list (`"events"`) | **0** |
|
||
| Event-call syntax (`create_entity(`, `set_scene(`, …) | **0** |
|
||
| Hard-limit / continue-hint echo | **0** |
|
||
|
||
The browser run agrees independently: `protocol_shapes_in_narration` **0** across
|
||
its narrated turns (§G).
|
||
|
||
**The known mid-reply echo did not recur — and is still not fixed.** WP-B.1
|
||
recorded one stored reply (action 153, depth 143) where the narrator echoed the
|
||
length hint mid-reply and then continued the story, which A2's trailing cleanup
|
||
does not remove. Replaying **that stored fixture through this candidate's
|
||
detector** still flags it — 1 of 105 AI actions, matched by the hard-limit/hint
|
||
rule alone. So the residual is live (§T.2); what this release run shows is that
|
||
no equivalent shape occurred in **its** 105 replies. This report does not claim
|
||
protocol leakage is solved.
|
||
|
||
Ordinary fact and state restatement in prose was not counted: it is measurement,
|
||
not application-owned protocol, and no rule treats it as a leak.
|
||
|
||
## K. Memory / WP-B evidence
|
||
|
||
### K.1 At the final recall point
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| **Created** | 19 memories in the bank; 12 summaries |
|
||
| **Retained** | the planting turn's era survived to the end of a 102-turn run |
|
||
| **Ranked** | 4 memories were selected into the prompt at the recall point |
|
||
| **Injected** | the memory section reached the prompt (14,395 tokens that turn) |
|
||
| **Independent-memory verdict** | **not demonstrated** — `precondition_failed: absent_from_later_narration` |
|
||
|
||
### K.2 The independent-fact probe, precondition by precondition
|
||
|
||
| Precondition | Held? |
|
||
| --- | --- |
|
||
| The planted turn is outside the history window (planted depth 3, floor 72) | **yes** |
|
||
| Absent from authoritative state | **yes** |
|
||
| Absent from the summary | **yes** |
|
||
| Absent from imported knowledge | **yes** |
|
||
| Absent from later narration | **no** — 6 violations, the first at turn 6 |
|
||
|
||
The narrator restated the planted fact in later narration, so the probe could not
|
||
isolate memory as the only path. The run therefore **records no independent
|
||
recovery**, and nothing here is relabelled as one.
|
||
|
||
### K.3 The two statements the contract requires, kept apart
|
||
|
||
```text
|
||
deterministic independent-memory recovery: PASS
|
||
reference-model independent-memory limitation: ACCEPTED RESIDUAL
|
||
```
|
||
|
||
- **Deterministic** (`WP-B.2` §I, and the suite on this candidate): the
|
||
`independent_full` scenario fails on v1.0.0 at creation and returns
|
||
`recovered_through_memory_independent` on the candidate, with isolation
|
||
asserted every turn and provenance resolving to the planting turn.
|
||
- **Reference model:** failed on the precondition-valid attempt in WP-B.2, and
|
||
in this release run the attempt was not precondition-valid at all. The failing
|
||
stage remains **memory creation** — the summariser's content selection.
|
||
- **No new regression.** B2.1 ranking, B2.2 eviction and B2.3 excerpt creation
|
||
all pass deterministically in the 1,723-test suite on this candidate, and the
|
||
bank behaved normally through the run (19 memories, ranked and injected). What
|
||
this run shows is the known summariser-quality limitation, not a fault in
|
||
ranking, eviction or injection.
|
||
|
||
## L. Identity diagnostic
|
||
|
||
**Scripted half — complete.** `tools/m11_identity.py --scripted` on the
|
||
candidate: 10 identity stresses, **0 signals**, verdict *no objective identity
|
||
defect detected*.
|
||
|
||
**The detector's negative control fires.** With `--inject`, the same harness on
|
||
the same candidate raises **7 signals** — `shared_display_name` once and
|
||
`duplicate_character_creation` on turns 5–10 — and preserves each turn's
|
||
evidence. A clean run therefore means something: the check is capable of
|
||
failing.
|
||
|
||
**Model-backed half — complete, on the candidate, with memory on.**
|
||
Evidence: `…/release-87a4032/identity-memory/`.
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Model | **`qwen2.5:3b-instruct-16k`** (the reference narrator) |
|
||
| Context window | **16,384** |
|
||
| Memory status | **on** — `memoryBankEnabled: true`, embedding model `nomic-embed-text` |
|
||
| Summary status | **on** — `autoSummarize: true`; **1 summary written** |
|
||
| **Identity signals** | **0** across all 10 stresses |
|
||
| **Stored protocol shapes** | **0** of 10 AI actions, by the release-gate detector |
|
||
| Fixture | accepted successfully — five entities kept distinct (`bill`, `alice`, `roger`, `john`, `office`) |
|
||
| State proposals | recorded and applied with no shared display name and no duplicate creation |
|
||
| Scripted detector self-test | still fires (7 signals under `--inject`) |
|
||
|
||
**A harness correction made during validation, and what it does not invalidate.**
|
||
The diagnostic shipped with `embedding_model=""` and the memory bank switched
|
||
off, so a release run of it would have reported a clean identity result with
|
||
memory never taking part — which is not what the gate asks for. Two lines now
|
||
read `AIDND_TEST_EMBED_MODEL` and enable the bank. **Harness-only: no product
|
||
code changed**, so no product evidence became stale; the corrected harness
|
||
repeated its own check, which is the run reported above.
|
||
|
||
**One honest observation:** with memory enabled the bank still wrote **0
|
||
memories** in this 10-beat campaign — 21 actions is enough to pass
|
||
`MEMORY_START`, but the summary pass is what ran and produced the single
|
||
summary. Memory was configured and active; it was not meaningfully *exercised*
|
||
here. The bank's real exercise is the 102-turn run (§K), which wrote 19.
|
||
|
||
**No claim about the historical root cause.** The post-M8 identity finding's
|
||
campaign was destroyed and its cause cannot be established. A clean run here is
|
||
evidence that the product does not do the things it can be blamed for on this
|
||
fixture — not a discovery of what happened then.
|
||
|
||
## M. Recovery
|
||
|
||
`tools/m11_recovery.py` against **the integrated run's own bundle**
|
||
(3,071,683 bytes), imported into a database file that never existed, in a
|
||
directory that never existed, by a second server process — so migrations ran
|
||
from nothing and this is the fresh-install path as well as the import path.
|
||
|
||
**16 checks, 16 passed, 0 failed.**
|
||
|
||
| Claim | Result |
|
||
| --- | --- |
|
||
| The destination database did not exist beforehand | PASS |
|
||
| The bundle imports into a clean directory | PASS — 209 actions in the file, 121 on the active line |
|
||
| The active transcript is not empty | PASS |
|
||
| Authoritative state came across | PASS — 6 entities, 1 fact |
|
||
| The campaign's own canon came across | PASS |
|
||
| The narration-length choice came across | PASS |
|
||
| Redo availability matches what the file said | PASS |
|
||
| Retained (undone) history came across | PASS — **88 actions retained beyond the active line** |
|
||
| Both Save Points restore | PASS — *On the ridge*, *Before the ridge* |
|
||
| Every imported class came across, with content | PASS — 3 sources |
|
||
| The moved campaign accepts a new change, unrefused | PASS |
|
||
| The bundle carries no secret | PASS |
|
||
| The moved campaign exports again, same story length | PASS |
|
||
|
||
**One thing this does not prove, stated rather than implied.** The long run
|
||
ended with its head at the tip, so the file's head was at the tip and Redo was
|
||
correctly unavailable after import. The **undone-head** case (I07) is proved by
|
||
§N's upgrade campaign, which ends on an Undo with Redo available and survives
|
||
both bundle directions, and by `test_m9_portability.py` in the suite — not by
|
||
this bundle.
|
||
|
||
## N. v1.0.0 upgrade compatibility
|
||
|
||
A campaign **built and played by the `432f041` application** in its own
|
||
worktree, then opened by the candidate. Evidence:
|
||
`…/release-87a4032/upgrade/upgrade-report.json`.
|
||
|
||
**Phase 1 — v1.0.0 builds it.** 20 turns accepted, 39 actions, canon knowledge
|
||
imported, a Save Point taken, one Undo so the head is not at the tip, narration
|
||
length chosen, memory bank and auto-summary on. It contains what §11 item 9
|
||
names: retained history, an undone head with Redo available, a Save Point,
|
||
**6 memories**, **2 summaries**, imported knowledge, a narration-length choice
|
||
and narrative state. Settings hold a **loopback placeholder**
|
||
(`http://127.0.0.1:11434/v1`) — no real hostname is in the evidence database.
|
||
|
||
**Phase 2 — the candidate opens the same file.** Nothing was copied; the
|
||
candidate's migrations ran against it.
|
||
|
||
| Field | Before | After |
|
||
| --- | --- | --- |
|
||
| transcript | 39 actions | **identical** |
|
||
| newest action (the head) | — | **identical** |
|
||
| total / can_undo / can_redo | 39 / true / **true** | **identical** |
|
||
| checkpoints | 1 | **identical** |
|
||
| narrative state | 7 keys | **identical** |
|
||
| memories | **6** | **identical** |
|
||
| summaries | **2** | **identical** |
|
||
| knowledge sources | 1 | **identical** |
|
||
| narration length | `brief` | **identical** |
|
||
| memory_bank_enabled / auto_summarize | true / true | **identical** |
|
||
| settings | loopback placeholder | **identical** |
|
||
| **schema `user_version`** | **94** | **94** |
|
||
|
||
**All 15 census fields compared identical; 0 differ.** Schema parity is exact:
|
||
v1.0.0 and the candidate both stamp `user_version` 94 with 93 migrations, so the
|
||
upgrade required no migration at all, and nothing was rewritten in passing. A
|
||
fresh-install database from the candidate carries the same 94 (§M's import ran
|
||
migrations from nothing).
|
||
|
||
**A first pass at this gate was discarded.** It played 5 turns, which is below
|
||
the memory and summary thresholds, so it compared **0 memories against 0
|
||
memories** and proved nothing about two of the criterion's required contents;
|
||
its settings also carried the live hostname rather than a placeholder. Both were
|
||
corrected and the gate was rerun — the run reported above.
|
||
|
||
## O. Bundle compatibility
|
||
|
||
| Direction | Result |
|
||
| --- | --- |
|
||
| **v1.0.0 export → imported by v1.1** | **PASS** — imported, id 2 |
|
||
| **v1.1 export → offered to v1.0.0** | **ACCEPTED** — v1.0.0 imported it, id 1 |
|
||
|
||
Both directions were **executed**, not inferred from the unchanged format
|
||
string. The format is `ai-dnd-adventure-v3` on both sides, and it did not change
|
||
during release validation.
|
||
|
||
Backward import succeeding means the compatibility question the brief raised —
|
||
whether optional or additive v1.1 evidence data would break a v1.0.0 importer —
|
||
is answered in the negative for this campaign's contents: v1.0.0 accepted the
|
||
candidate's bundle whole. No bundle-format change was made or needed.
|
||
|
||
## P. WP-D regression
|
||
|
||
Reconfirmed on the candidate; no backup-affecting product code changed after
|
||
WP-D, so its 117 MB browser measurement is not repeated (the owner's brief
|
||
permits this).
|
||
|
||
| Claim | Result |
|
||
| --- | --- |
|
||
| The completed backup copy is verified with `PRAGMA integrity_check` | **PASS** — `app/backup.py:193`, docstring at 198–201 records why the full check replaced `quick_check` |
|
||
| The corruption fixture still separates the two pragmas | **PASS** — `quick_check` → `ok`, `integrity_check` → `row 145 missing from index i_t_k` |
|
||
| Oversized export still succeeds and is delivered | **PASS** |
|
||
| Importability metadata names the effective ceiling | **PASS** — "This export is larger than this version's **20 MB** import limit (… bytes). The file was exported successfully, but this version cannot import it." |
|
||
| Normal export unchanged in content | **PASS** |
|
||
| Oversized import still refused | **PASS** — 413 naming the limit |
|
||
| Suite | **12 passed** (`tests/test_v11_d_recovery.py`) |
|
||
|
||
## Q. WP-E regression
|
||
|
||
| Claim | Result |
|
||
| --- | --- |
|
||
| Contrast audit exit code | **0** |
|
||
| All applicable control boundaries ≥ 3:1 | **PASS** — every boundary pair clears 3:1 (1.4.11) |
|
||
| Applicable text contrast still compliant | **PASS** — every text pair clears 4.5:1 (1.4.3); baselines 14.57 / 13.57 / 5.48 / 5.88 unchanged |
|
||
| Browser boundary checks | **PASS** — the 10 WP-E rows in §G |
|
||
| Focus visibility | **PASS** — M11's visible-focus check inside the 38, plus WP-E's focused-edge measurement |
|
||
| Gate tests | **11 passed** (`tests/test_v11_e_contrast.py`), including 2.99:1 failing and 3.00:1 passing |
|
||
|
||
```text
|
||
OWNER SCREENSHOT APPROVAL: APPROVED
|
||
```
|
||
|
||
Approved by the owner in the release-validation brief of 2026-09-16. **No visual
|
||
code changed during release validation**, so that approval remains valid; had any
|
||
changed, it would have been void and new screenshots would have been required.
|
||
|
||
## R. Release-shaped smoke test
|
||
|
||
The **final no-cache candidate image**, a fresh volume, published on loopback,
|
||
with the private CA installed into the container's own trust store. Evidence:
|
||
`…/release-87a4032/smoke/`.
|
||
|
||
**15 checks, 15 passed, 0 failed.**
|
||
|
||
| Claim | Result |
|
||
| --- | --- |
|
||
| The container starts | PASS |
|
||
| The application answers on loopback | PASS |
|
||
| The port is published on **loopback only** | PASS — `8000/tcp -> 127.0.0.1:…` |
|
||
| This machine's **LAN address does not serve** the application | PASS |
|
||
| The first page loads | PASS — HTTP 200 |
|
||
| The shell references no remote origin | PASS — none found |
|
||
| A CSP is served | PASS |
|
||
| **The approved HTTPS narrator verifies through its private CA** | PASS — HTTP 200 through `tlstrust.ssl_context()`, **no bypass** |
|
||
| **A public endpoint is refused** | PASS — HTTP 400 |
|
||
| A campaign is created | PASS |
|
||
| **One real narrator turn is accepted** | PASS |
|
||
| The container restarts and serves again | PASS |
|
||
| The transcript survived the restart | PASS |
|
||
| The narrative state survived the restart | PASS |
|
||
| **Firefox renders the reopened campaign** | PASS — 470 characters of story |
|
||
|
||
**A finding worth recording, and it is not a product defect.** The first attempt
|
||
failed at `PUT /api/settings` with **HTTP 400**. The cause: a `.local` name is
|
||
mDNS, a Docker container has no mDNS resolver, and `endpoints.py` correctly
|
||
refuses an endpoint whose address it cannot classify — the policy behaving
|
||
exactly as designed. The fix is to resolve the name inside the container
|
||
(`--add-host`), **not** to substitute the IP address, because the certificate is
|
||
issued for the hostname and substituting the address would have quietly bypassed
|
||
the hostname verification this test exists to prove. Harness-only; no product
|
||
code changed.
|
||
|
||
This is supplemental evidence, not a substitute for the gates above.
|
||
|
||
## S. Security / local-only review
|
||
|
||
| Claim | Evidence on the candidate |
|
||
| --- | --- |
|
||
| Served on loopback only | every harness reached the application on `127.0.0.1`; the documented container run publishes loopback |
|
||
| Endpoint policy | `app/endpoints.py` admits loopback (v4 and v6), the three RFC1918 ranges, link-local, IPv6 unique-local and CGNAT, and refuses the public Internet; a name resolving to both a private and a public address is refused |
|
||
| Inference actually used | trusted-LAN **HTTPS** with a private CA for the browser gate (§G); plain HTTP to a LAN GPU host for the long run, which `SECURITY-THREAT-MODEL.md` §83 permits and which is not A06 evidence |
|
||
| No secret in exports | offline gate, WP-D tests and the M11 suite |
|
||
| CSP served | `default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none'; base-uri 'none'; form-action 'self'; frame-ancestors 'none'` |
|
||
| Other headers | `x-content-type-options: nosniff`, `referrer-policy: same-origin`, `x-frame-options: DENY` |
|
||
| No remote origin in the shell | offline gate: the page names none, and every asset is local |
|
||
|
||
`'unsafe-inline'` remains on `style-src` only, because React writes inline
|
||
`style` attributes; it is deliberately absent from `script-src`.
|
||
|
||
## T. Known residual risks
|
||
|
||
Each is classified, and none is collapsed into another category.
|
||
|
||
### T.1 WP-B reference-model memory limitation — ACCEPTED RESIDUAL
|
||
|
||
```text
|
||
deterministic independent memory: PASS
|
||
reference-model independent memory: FAIL
|
||
failing stage: memory creation — the summariser's content selection
|
||
owner decision: accepted for v1.1
|
||
```
|
||
|
||
**Status on this candidate:** unchanged, and no broader regression. The release
|
||
long run did not demonstrate independent recovery, but it also could not: its
|
||
probe failed the `absent_from_later_narration` precondition because the narrator
|
||
restated the fact (§K.2). Ranking, eviction and excerpt creation all pass
|
||
deterministically in the 1,723-test suite, and the bank worked normally through
|
||
102 turns (19 memories, ranked and injected). **Accepted residual**, carried
|
||
visibly, not a blocker.
|
||
|
||
### T.2 A2 mid-reply application-instruction echo — ACCEPTED RESIDUAL, still live
|
||
|
||
The known occurrence is WP-B.1's action 153 (depth 143): the narrator echoed the
|
||
length hint **mid-reply** and then continued the story, which A2's trailing
|
||
cleanup does not remove.
|
||
|
||
**Reproduced on this candidate.** Replaying that stored fixture through the
|
||
release-gate detector still flags it — **1 of 105** AI actions, matched by the
|
||
hard-limit/hint rule alone. The extractor still leaves it. This report therefore
|
||
does **not** claim protocol leakage is solved.
|
||
|
||
**But it did not recur in release evidence.** This run's own 105 stored replies
|
||
leak **0** (§J), and the browser run's narrated turns leak 0. The brief's
|
||
stop-and-report condition — *an equivalent shape occurring in this final run* —
|
||
was **not** triggered, so validation continues. No broader sanitizer was written;
|
||
that remains for owner review.
|
||
|
||
### T.3 Doubled full stop in memory-search scene text — ACCEPTED RESIDUAL
|
||
|
||
`"rain outside.."` when the state's scene summary already ends in punctuation.
|
||
Only the embedding query sees it; the effect is one stray token. Not changed
|
||
during release validation, deliberately: cleanliness is not a reason to alter
|
||
product behaviour after evidence is taken.
|
||
|
||
### T.4 K1 — "Correct" on an Important Facts row is always refused
|
||
|
||
**Reproduced, unchanged on this candidate**, deterministically and without a
|
||
browser: Correct on a Characters row (`subject='mara'`) applies, **201**; Correct
|
||
on an Important Facts row (`subject='f1'`) is refused, **400 — "add_fact names
|
||
subject='f1', which does not exist."** The cause is frontend-side: the panel
|
||
sends the row key as `add_fact.subject`, and the validator checks `subject` as an
|
||
entity reference.
|
||
|
||
**Classification: v1.2 backlog, not a release blocker.** It blocks no v1
|
||
REQUIRED test — C04 passes through the working correction paths (§C) — and WP-C's
|
||
State-panel correction coverage passes. It is a narrow bug with an obvious fix
|
||
(offer Correct only against entities, or send facts without a subject), but
|
||
fixing it during release validation would change product code after the evidence
|
||
above was taken, which the brief forbids without a product brief. **Left for the
|
||
owner.**
|
||
|
||
### T.5 Import ceiling and scheduled backups — INTENTIONAL, not unfinished work
|
||
|
||
The 20 MB import ceiling is deliberate and unchanged; WP-D made it honest rather
|
||
than raising it. Scheduled backups remain unbuilt by design. Neither is a
|
||
residual defect.
|
||
|
||
### T.6 Harness corrections made during validation — no product evidence invalidated
|
||
|
||
Three, all harness-only, each named where it occurred: the identity diagnostic
|
||
did not enable memory (§L); the smoke test needed hostname resolution inside the
|
||
container (§R); and the first upgrade campaign was too short to write memories
|
||
and carried a live hostname (§N). **No product code changed at any point during
|
||
release validation**, so no black-box or long-run evidence became stale. Each
|
||
corrected harness repeated its own affected check.
|
||
|
||
## U. Deferred v1.2 / future work
|
||
|
||
| Item | Why it is deferred |
|
||
| --- | --- |
|
||
| Raising the import ceiling, or a streaming import | The plan assigns it to v1.2; WP-D's scope was honesty about the limit, not the limit |
|
||
| Scheduled backups; a restore button | Explicitly out of WP-D's scope |
|
||
| K1's Correct-on-a-fact-row fix (§T.4) | A narrow frontend bug needing a product brief |
|
||
| A broader mid-reply protocol sanitizer (§T.2) | Needs owner review; A2's cleanup is deliberately trailing-only |
|
||
| The doubled full stop (§T.3) | Cosmetic, embedding-query only |
|
||
| K05 generate local image, K06 multi-turn video | FUTURE tests; both need a media provider this release does not build |
|
||
| Reference-model independent memory (§T.1) | Needs a stronger summariser or a different creation strategy — a v1.2 investigation, not a v1.1 fix |
|
||
|
||
## V. Documentation changes
|
||
|
||
Current documents were brought up to date. **No historical milestone report was
|
||
rewritten, and no failed WP-B real-model evidence was turned into success.**
|
||
|
||
| Document | Change |
|
||
| --- | --- |
|
||
| `README.md` | Status now says v1.0.0 **remains** the released version, that all six v1.1 packages are complete and accepted, that release validation passed on candidate `87a4032`, that WP-B ships with a documented limitation, and that **no `v1.1.0` tag exists and `main` is unchanged`**. Also corrected a stale figure: the schema is versioned at **94**, not "92 and counting" |
|
||
| `planning/V1.1-PLAN.md` | WP-D/WP-E recorded as signed `87a4032`; the release-validation outcome summarised with its residuals; the three owner events named as still outstanding |
|
||
| `planning/VERSION.md` | Same status correction, plus a new revision entry for the closeout |
|
||
| `planning/README.md` | Current-state paragraph rewritten for the same facts |
|
||
| `planning/reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL: PENDING` → **`APPROVED`**, with the source (the release-validation brief) and date recorded, and a note that the signed commit predated the review |
|
||
| `planning/reports/v1.1/V1.1-RELEASE-REPORT.md` | **New** — this document |
|
||
| `DEVELOPMENT.md` | Unchanged: its WP-C/WP-D sections already describe the candidate as built |
|
||
|
||
**New tools committed with this closeout** (harness only, no product code):
|
||
`backend/tools/v11_upgrade_check.py` (Gate 9) and
|
||
`backend/tools/v11_release_smoke.py` (§R), plus two narrow corrections to
|
||
`backend/tools/m11_identity.py` (read `AIDND_TEST_EMBED_MODEL`; enable the memory
|
||
bank) so the diagnostic can run with memory on.
|
||
|
||
## W. Final release decision
|
||
|
||
**The question this validation set out to answer:** does candidate `87a4032`
|
||
preserve the complete v1 contract and satisfy every accepted v1.1 package on one
|
||
integrated release tree?
|
||
|
||
| Gate | Result |
|
||
| --- | --- |
|
||
| 1 Package acceptance | **PASS** — six packages accepted; WP-B's qualification carried whole (§B.1) |
|
||
| 2 v1 acceptance contract | **PASS** — 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified (§C) |
|
||
| 3 Suites, lint, build, Docker | **PASS** — 1,723 / 175 / 0 errors; image SPA file-for-file identical (§D, §E) |
|
||
| 4 Offline / no-network | **PASS** — 23/23 on the candidate image (§F) |
|
||
| 5 Browser release run | **PASS** — 101/0/0 over trusted-LAN HTTPS, every turn `fits` (§G) |
|
||
| 6 Integrated long run | **PASS** — 102 turns, M01–M04, 0 post-turn failures (§H) |
|
||
| 7 Identity diagnostic | **PASS** — 0 signals, 0 protocol shapes, memory on (§L) |
|
||
| 8 Recovery | **PASS** — 16/16 on the run's own bundle (§M) |
|
||
| 9 v1.0.0 upgrade | **PASS** — 15/15 identical, schema parity at 94 (§N) |
|
||
| 10 WP-D regression | **PASS** (§P) |
|
||
| 11 WP-E regression | **PASS**, screenshots approved (§Q) |
|
||
| Bundle compatibility | **Both directions execute and import** (§O) |
|
||
| Release smoke | **PASS** — 15/15 from the shipped image (§R) |
|
||
|
||
**A1** holds its reserve on every re-counted prompt, with a smallest margin of
|
||
**879 tokens** against v1's 23–42. **A2**'s release-gate leak count is **0**.
|
||
**WP-B**'s deterministic independent-memory recovery passes and its
|
||
reference-model limitation remains an accepted, documented residual — stated in
|
||
§K.3 in both halves, never shortened to "WP-B passed".
|
||
|
||
No product code was changed at any point during release validation, so no
|
||
evidence was invalidated. Three harness corrections were made and each corrected
|
||
harness repeated its own check (§T.6).
|
||
|
||
```text
|
||
V1.1 RELEASE VALIDATION:
|
||
PASS
|
||
```
|
||
|
||
### What this decision is not
|
||
|
||
These are separate, and only the first is done:
|
||
|
||
```text
|
||
WP-A-E accepted: YES
|
||
release validation passed: YES
|
||
release candidate prepared: YES (87a4032, with this report)
|
||
release commit signed: NO
|
||
main updated to v1.1: NO
|
||
v1.1.0 tagged: NO
|
||
```
|
||
|
||
The closeout changes are **staged and uncommitted**. Nothing was committed,
|
||
pushed, merged or tagged by this validation. The owner's next decision is to
|
||
review this evidence, resolve anything they disagree with, then sign the v1.1
|
||
release commit and publish `v1.1.0`.
|