Files
interactive-story/planning/reports/v1.1/V1.1-RELEASE-REPORT.md
JesseMarkowitzandClaude Opus 5 db7b309e3d
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
v1.1 closeout: accept integrated release validation
Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.

V1.1 RELEASE VALIDATION: PASS

What was run, on this candidate:

- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
  0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
  175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
  to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
  trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
  3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
  keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
  scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
  that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
  identical on all 15 census fields, schema parity at user_version 94, and both
  bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
  verified, public endpoint refused, a real turn, restart, persistence, and
  Firefox rendering the reopened campaign.

Carried residuals, stated rather than summarised away:

- WP-B: deterministic independent-memory recovery PASS; reference-model
  independent-memory recovery FAIL at memory creation — the owner-accepted
  limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
  reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
  release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
  backlog, reproduced and not fixed during validation.

Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.

Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.

Still the owner's to do: sign the release commit, update main, tag v1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-16 07:12:23 -04:00

47 KiB
Raw Permalink Blame History

Adventure Storyteller v1.1 — Integrated Release Validation

Status: COMPLETE — V1.1 RELEASE VALIDATION: PASS. The decision, and what it deliberately does not cover, is in §W.

This report answers one question: does this exact candidate preserve the complete v1 contract and satisfy every accepted v1.1 package on one integrated release tree? It is release validation, not a work package. Nothing here adds a feature, and no release tag is created by it.


A. Repository / provenance

Candidate SHA 87a40326a29533c8d52c9f9f41022e7b499b1de7
Branch v1.1-development, up to date with origin/v1.1-development
Working tree at freeze clean — nothing modified, nothing staged
Commit v1.1: harden recovery and control boundaries (WP-D + WP-E)
Owner signature Good signature, RSA key 02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569, made 2026-09-16 05:37:13 EDT
Tag at HEAD none — no v1.1.0 tag exists
v1.0.0 baseline 432f04100b9a67198bcdc46c6ff8ee0f181e1667, an ancestor
Package ancestry d63804f (WP-A1/A2), beb17ad (WP-B.1), 0c1ba83 (WP-B.2), 59b5ebc (WP-C) — all ancestors
Diff v1.0.0..HEAD 61 files, +14,833 / −366
LICENSE / PROVENANCE unchanged since v1.0.0 (empty diff)

A.1 Frozen candidate identity

Dependency locks backend/requirements.txt sha256:ed28bc0f8970cf4e…, frontend/package-lock.json sha256:355cb370837ade01…, frontend/package.json sha256:2016580ddfa176a9…, backend/requirements-dev.txt sha256:06d7695816b201e9…
Schema LATEST_VERSION 94, 93 migrations (PRAGMA user_version)
Bundle format ai-dnd-adventure-v3
Import ceiling 20 MB (MAX_IMPORT_BODY_BYTES), unchanged
Frontend build dist built 2026-09-16T05:39:55, 16 files, sha256(dist) = ea2753ad24f61959fe084f4674911acc
Docker image storyteller:release-87a4032, sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398, 312 MB
Firefox / geckodriver 155.0.1 / 0.37.1 (2026-09-04)
Docker 29.8.0, build 88096ef
CPU HTTPS reference host Ollama 0.33.0, serving qwen2.5:3b-instruct, qwen2.5:3b-instruct-16k, nomic-embed-text:latest; certificate verifies through the machine's CA store with no bypass
GPU inference host Ollama 0.34.0; qwen2.5:3b-instruct-16k digest 21ff8cc52f375f19, nomic-embed-text:latest digest 0a109f422b47e3a3; no model resident at start
Evidence root $HOME/v11-evidence/release-87a4032/ — never /tmp, and no real hostname appears in any committed file

B. Package acceptance inventory

Package Status Source
WP-A1 context-window safety reserve ACCEPTED V1.1-WP-A1-A2-REPORT.md, signed d63804f
WP-A2 protocol-echo cleanup, genre-neutral prompting ACCEPTED same report and commit
WP-B independent memory ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION V1.1-WP-B1-REPORT.md, V1.1-WP-B2-REPORT.md §S, signed beb17ad / 0c1ba83
WP-C browser release coverage ACCEPTED V1.1-WP-C-REPORT.md, signed 59b5ebc
WP-D recovery honesty ACCEPTED V1.1-WP-D-REPORT.md, signed 87a4032
WP-E control-boundary contrast ACCEPTED V1.1-WP-E-REPORT.md, signed 87a4032

B.1 WP-B's qualification, carried whole

The WP-B disposition is not shortened to "WP-B passed" anywhere in this report. Its own §S records:

B2.1 RANKING: PASS
B2.2 EVICTION: PASS
B2.3 EXCERPT CREATION: PASS
B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED

DETERMINISTIC WP-B: PASS
REAL-MODEL WP-B: FAIL

WP-B OVERALL:
ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION

In the release contract's own words (§11 items 12), that is:

DETERMINISTIC WP-B: PASS
REFERENCE-MODEL INDEPENDENT MEMORY: FAIL
OWNER ACCEPTED THE LIMITATION FOR v1.1

The failing stage is memory creation — the summariser's content selection — not ranking, eviction or injection, each of which passes deterministically.

B.2 WP-E screenshot approval — a correction of record

The committed WP-E report read OWNER SCREENSHOT APPROVAL: PENDING, because it was written before the owner reviewed the images. The owner's release-validation brief (2026-09-16) states the before/after screenshots are approved and instructs this validation to record it. The WP-E report is updated to APPROVED as part of this closeout (§V), sourced to that brief and dated. No visual code changed during release validation, so the approval stands (§Q).


C. v1 acceptance matrix

Every test marked REQUIRED FOR V1 — there are 82 — against evidence taken on this candidate. Evidence types follow M11's: browser (the 101-check run, §G), campaign (the 102-turn integrated run, §H), container (the offline run on the candidate image, §F), process (spawned server processes — recovery §M, upgrade §N), suite (the 1,723-test backend suite, §D). No historical result from different product code is used where the contract asks for candidate evidence.

Result: 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified.

A — Local-first operation

ID Result Evidence on this candidate
A01 Start application offline PASS container: first page load, fresh volume, no route and no DNS
A02 Storyteller loopback default PASS suite; every harness reached it on 127.0.0.1; docker-compose.yml publishes 127.0.0.1:8000:8000
A03 No cloud API key PASS suite; container: no secret in an export
A04 Campaign survives restart PASS campaign: 3 process restarts, 4 process starts, state compared across each; container: campaigns survive a container restart
A05 Failed model call does not corrupt story PASS campaign: a real failed_call at turn 69 against an unserved model, play resumed; container: same with no model reachable
A06 Trusted-LAN Ollama inference PASS browser: the whole 101-check run over trusted-LAN HTTPS with a private CA, verification on, no bypass, storyteller loopback-bound

B — Core play

ID Result Evidence
B01 Natural language action PASS browser (real turns through the UI) + campaign (102 accepted)
B02 Dialogue input PASS campaign: dialogue beats in the turn list
B03 Continue PASS suite; browser: the Continue control present and enabled

C — Story authority and state

ID Result Evidence
C01 Campaign canon is preserved PASS campaign: canon present in the prompt on 102 of 102 turns
C02 Possession state PASS campaign (the silver key) + suite
C03 Character knowledge is not invented PASS suite
C04 Manual state correction PASS campaign: 2 state corrections; browser: C3's accepted and refused corrections; suite
C05 Canon beats reference PASS suite
C06 Structured state matches accepted narrative consequence PASS campaign: real extraction across 102 turns, every event validated or refused; suite

D — Non-destructive history

ID Result Evidence
D01 Undo one turn PASS browser + campaign (undo)
D02 Minimum five undos PASS suite; campaign (undo_redo)
D04 Redo PASS browser + campaign
D05 Redo invalidated by new continuation PASS campaign: diverged, after which Redo is gone
D06 Retry narrator response PASS campaign: 2 retries
D07 Select prior retry take PASS campaign: take_selected
D08 Retry does not delete prior take PASS campaign + suite
D09 Edit earlier user input PASS suite
D10 Edit narrator output PASS suite; browser (hostile-Markdown plants through the narrator-edit path)
D11 Named checkpoint PASS campaign: 2 Save Points; browser; suite
D12 Restore checkpoint PASS campaign: save_point_restored; browser
D13 Restore does not delete later history PASS campaign: retained actions after the restore; recovery §M
D14 Delete checkpoint PASS suite; browser: the delete confirmation dialog

E — Branch and derived-data isolation

ID Result Evidence
E01 Abandoned future cannot affect active state PASS suite test_m11_leakage.py, with a positive control
E02 Abandoned memory cannot leak PASS as above
E03 Abandoned summary cannot leak PASS as above
E04 Scene state is lineage-safe PASS as above

F — Long-term memory and context

ID Result Evidence
F01 Recent turns remain coherent PASS campaign: history populated every turn, newest always included
F02 Old important event retrieval PASS campaign §K: the planting turn outside the window and the fact recovered — through authoritative state, not independent memory (§K states which)
F03 Prompt remains bounded PASS campaign: 1,602–14,982 tokens against a 16,384 budget across 102 turns
F04 Output token reserve PASS campaign: output_reserve 500 present and subtracted on every turn
F05 Prompt inspector PASS browser: the context panel shows the assembled prompt
F06 Retrieval provenance PASS campaign: knowledge and memory provenance per turn; suite
F07 Heuristic memory is not canon PASS suite
F08 Memory failure is non-fatal PASS suite; container: derived work fails with no model and turns still commit; campaign: 0 post-turn failures, no database-lock errors

G — Imported knowledge

ID Result Evidence
G01 Import local text PASS browser: through the real file input; container: offline
G02 Import local Markdown PASS campaign: 3 sources imported; browser
G03 Classification PASS campaign: all three classes; recovery §M confirms them after a move
G04 Disable knowledge source PASS suite
G05 Canon retrieval PASS campaign: canon passages in stored prompts; suite
G06 Reference retrieval PASS suite
G07 Inspiration is low authority PASS suite
G08 No automatic URL fetch PASS suite test_egress.py; container: no network at all, import still works
G09 Remote Markdown image does not auto-load PASS browser: no remote image src in the rendered story
G10 Prompt injection in source is treated as data PASS suite; browser: injection text rendered as text

H — Security

ID Result Evidence
H01 No unexpected outbound connections PASS container (no network at all) + suite test_egress.py
H02 No telemetry PASS suite
H03 No cloud provider required PASS container: a full campaign offline
H04 Model output cannot execute shell PASS browser + suite
H05 Invalid state event rejected PASS suite; campaign: refusals recorded
H06 Stored XSS protection PASS browser: onerror and <script> in accepted narration, neither executed
H07 JavaScript URL protection PASS browser: no javascript: href in the DOM
H08 Path traversal import rejected PASS suite
H09 ZIP Slip protection NOT APPLICABLE the product extracts no archives, and a test enforces it — the same condition v1 recorded
H10 Restrictive CORS and local API behaviour PASS browser: an unknown API path is a 404 with a non-HTML body; suite
H11 No first-use runtime asset download PASS container: every asset local with no network; suite: the tokenizer table is vendored
H12 Inference endpoint enforcement PASS suite; smoke §R: a public endpoint refused by the running image; the container's refusal of an unresolvable name is the same policy (§R)

I — Export, import and recovery

ID Result Evidence
I01 Export campaign PASS campaign: the 102-turn campaign exported (3,071,683 bytes)
I02 Import exported campaign PASS process §M: imported into a database and directory that never existed
I03 Branch/disposable history export PASS §M: 88 actions retained beyond the active line after the move
I04 Checkpoint export PASS §M: both Save Points restore after the move
I05 Knowledge provenance export PASS §M: all three classes with their content
I06 Database/export contains no API secrets PASS §M: the bundle carries no secret; suite; container
I07 Export/import preserves an undone active head PASS §N: the v1.0.0 campaign's undone head and can_redo survive upgrade and both bundle directions; suite test_m9_portability.py. (The long-run bundle ended head-at-tip, so §M exercises the other case — stated in §M rather than implied)

J — Genre neutrality

ID Result Evidence
J01 Science-fiction campaign PASS suite test_m11_scifi.py
J02 Generic entity support PASS as above
J03 Genre profiles are configuration PASS as above

K — Future media architecture

ID Result Evidence
K01 Scene snapshot exists PASS suite; container: a scene packet builds offline
K02 Visual character profile PASS suite
K03 Visual location profile PASS suite

L — Data integrity

ID Result Evidence
L01 Atomic turn commit PASS container + campaign: a real induced failure, no narration accepted, no half-written state
L02 State reconstruction PASS suite; campaign: state compared across 3 restarts
L03 Checkpoint reconstruction after restart PASS campaign + §M

M — Long-run

ID Result Evidence
M01 100-turn campaign PASS 102 accepted turns, every scheduled operation exercised, 0 post-turn failures (§H)
M02 Restart during long campaign PASS 3 genuine process restarts (4 process starts); everything crossed as bytes on disk
M03 Long-run context stability PASS the prompt held 13,492–14,982 tokens over the last 70 turns against a 16,384 budget; window verified 102/102; canon present 102/102
M04 Long-run memory recall PASS, qualified the planted clue was outside the history window (planted depth 1, floor 72) and reached the prompt: m04_verdict: **recovered_through_state_only**. Recovery was through authoritative state, not independent memory — §K states this distinction and does not relabel it

SHOULD and FUTURE

Not counted as REQUIRED. SHOULD: B04, D03, K04, L04 — all still pass on the candidate (suite; browser for B04's direction toggle). FUTURE: K05, K06 — deliberately not run; both need a media provider this release does not build.

D. Backend / frontend suites

Suite Result
Backend (pytest -q, no AIDND_TEST_* set) 1,723 passed, 17 skipped, 0 failed, 0 xfailed (1,210.6 s)
Frontend (npm test) 175 passed, 15 files, 0 failed
Lint (npm run lint, oxlint) exit 0 — 0 errors, 15 warnings
Production build (npm run build) succeeded

Every skip explained — one category, and it is the expected one. All 17 are environment-gated real-model tests, skipped because AIDND_TEST_* is deliberately unset for the deterministic suite:

File Skipped Gate
test_knowledge_real_model.py 7 AIDND_TEST_ENDPOINT (and AIDND_TEST_EMBED_MODEL)
test_context_realistic.py 3 AIDND_TEST_ENDPOINT + AIDND_TEST_MODEL
test_narrative_realistic.py 3 same
test_m11_real_window.py 3 same; one needs AIDND_TEST_WIDE_MODEL
test_provider_wiring.py 1 same

0 xfailed. No deficiency WP-B fixed remains parked as an expected failure — B.1's two strict xfails became ordinary passes in B.2 and stayed that way.

Lint warnings are the documented, unchanged set: the pre-existing only-export-components and unused-import warnings recorded at WP-C, WP-D and WP-E. None is in a file this candidate changed relative to those packages.

E. Production build and Docker image

Command the repository's documented production build (DEVELOPMENT.md), with --no-cache
Image sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398 (312 MB)
Dependencies installed, not reused — npm ci and pip install --no-cache-dir both executed in the log
CACHED steps 2, and both are WORKDIR metadata (/build, /app) — no dependency or source layer was cached
Image SPA vs local build file-for-file identical: 16 files each, diff -r clean, combined sha256 = ea2753ad24f61959fe084f4674911acc on both sides

The image is built from the candidate tree for this validation. No earlier work-package image was reused.

F. Offline / no-network

tools/m11_offline.py against the candidate image, --network none, fresh volume. Evidence: …/release-87a4032/offline/offline-report.json.

23 checks, 23 passed, 0 failed. Including: no route to the public Internet; no external DNS; first page load; no remote origin named; CSP served; every referenced asset local; campaign creation; state extraction; local file import; prompt assembly; knowledge search; a turn with no model reachable reported as a failure with no narration accepted, the player's words kept and state unchanged; export; import with state; no secret in the export; the media module inert with no provider; campaigns surviving a container restart.

No first-use download occurred, which is what the container's absent network makes unfalsifiable rather than merely unobserved.

G. Browser release validation

tools/m11_browser.py on the candidate, production build served by FastAPI, Firefox 155.0.1 / geckodriver 0.37.1, narrator over trusted-LAN HTTPS with the private CA (endpoint_class: trusted-LAN HTTPS), storyteller on loopback. kind: release regression — not partial, not --only. 693 s.

Suite Passed Failed Skipped
M11 (the v1 release regression) 38 0 0
WP-C (browser release coverage) 53 0 0
WP-E (control boundaries) 10 0 0
Total 101 0 0

A1 accounting: 8 narrator turns, fits on every one; turns_not_clean 0; protocol_shapes_in_narration 0. The verified window on this host was 4,096 (source: loaded); the 16,384 evidence is the long run's (§I).

WP-C's proofs are all present in the 53: Retry and alternate takes, Save Point create/restore/Redo, state correction including a refusal shown as a refusal, narration length reaching the prompt, failed generation and recovery, and real export downloads from both the library and campaign settings, each file landing on disk and importing into a fresh application. WP-E's ten are the rendered boundary measurements of §Q.

H. Integrated 100+ turn long run

One new campaign on the candidate's product code, GPU inference host, with the owner's power/link/kernel logging running before the first turn. Evidence: …/release-87a4032/long-run/.

Narrator qwen2.5:3b-instruct-16k, digest 21ff8cc52f375f19
Embeddings nomic-embed-text, digest 0a109f422b47e3a3
Window 16,384, window_verified on 102 of 102 turns
Memory / summaries on — 19 memories in the bank, 12 summaries written
Accepted turns 102 (target 100)
Restarts 3 genuine process restarts, 4 process starts
Elapsed 1,330 s
Export 3,071,683 bytes; database 2,613,248 bytes
Status complete; aborted_reason null, failed_reason null

Not a repeated-turn benchmark. Every scheduled operation fired and is in the timeline: 3 restarts, 1 undo, 1 undo→redo, 2 retries, 1 take selection, 2 Save Points, 1 Save Point restore, 1 divergence, 2 state corrections, 3 knowledge imports, memory activation, 1 deliberate failed call, 1 export, the planted clue and the planted independent fact, and the recall probe.

H.1 M01–M04

Verdict What decides it
M01 100-turn campaign PASS 102 accepted turns, each with a committed action and state document
M02 Restart during long campaign PASS 3 genuine uvicorn restarts; everything that survived crossed as bytes on disk
M03 Long-run context stability PASS prompt 13,492–14,982 tokens over the last 70 turns against a 16,384 budget; canon present 102/102; window verified 102/102
M04 Long-run memory recall PASS, and qualified the clue was planted at depth 1, the history floor reached depth 72, and it was not in the recent window; it reached the prompt through authoritative state. m04_verdict: recovered_through_state_only. §K keeps the distinction the criterion was written around

H.2 State and derived-work integrity

0 post-turn failures across 102 turns, and no database is locked error. The only failure-shaped events in the whole timeline are the two the run creates on purpose: the scheduled failed_call at turn 69 (a model name the server does not serve — A05/L01 evidence) and the independent-fact precondition notes (§K). Narrative-state proposals were recorded, applied or refused as designed across the run, and no accepted narration carried an unresolved protocol block (§J).

I. Context-window / A1 evidence

Every turn fits. Across all 102 accepted turns the accounting status was fits — 0 exceeded, 0 truncation_suspected — and the window was verified at 16,384 on every one.

The ten largest stored prompts, re-counted against what the server itself reported:

Turn App estimate Server count Difference Window Output reserve Safety reserve Observed margin Status
86 14,990 15,005 +15 16,384 500 820 879 fits
83 14,978 14,993 +15 16,384 500 820 891 fits
97 14,965 14,980 +15 16,384 500 820 904 fits
41 14,960 14,975 +15 16,384 500 820 909 fits
66 14,951 14,966 +15 16,384 500 820 918 fits
33 14,941 14,956 +15 16,384 500 820 928 fits
35 14,940 14,955 +15 16,384 500 820 929 fits
77 14,940 14,955 +15 16,384 500 820 929 fits
38 14,927 14,942 +15 16,384 500 820 942 fits
45 14,922 14,937 +15 16,384 500 820 947 fits

The reserve is preserved on every re-counted prompt. The estimate runs exactly 15 tokens below the server's own count on all ten — a constant, known offset rather than drift — and the smallest observed margin anywhere in the run is 879 tokens, against the documented safety reserve of 820.

Against v1. M11's closeout recorded a remaining margin of 23–42 tokens. The same measurement on this candidate is 879 at its tightest — roughly twenty to thirty times the headroom, which is what WP-A1 was for.

J. Protocol-leak / A2 evidence

Release-gate result: 0. Across the run's 105 stored AI actions, protocol_leaks reports 0 leaking, example_ids empty.

Measured separately, by the detector's own four rules:

Shape Count in this run
State/section heading with an indented entry 0
Event list ("events") 0
Event-call syntax (create_entity(, set_scene(, …) 0
Hard-limit / continue-hint echo 0

The browser run agrees independently: protocol_shapes_in_narration 0 across its narrated turns (§G).

The known mid-reply echo did not recur — and is still not fixed. WP-B.1 recorded one stored reply (action 153, depth 143) where the narrator echoed the length hint mid-reply and then continued the story, which A2's trailing cleanup does not remove. Replaying that stored fixture through this candidate's detector still flags it — 1 of 105 AI actions, matched by the hard-limit/hint rule alone. So the residual is live (§T.2); what this release run shows is that no equivalent shape occurred in its 105 replies. This report does not claim protocol leakage is solved.

Ordinary fact and state restatement in prose was not counted: it is measurement, not application-owned protocol, and no rule treats it as a leak.

K. Memory / WP-B evidence

K.1 At the final recall point

Created 19 memories in the bank; 12 summaries
Retained the planting turn's era survived to the end of a 102-turn run
Ranked 4 memories were selected into the prompt at the recall point
Injected the memory section reached the prompt (14,395 tokens that turn)
Independent-memory verdict not demonstrated — precondition_failed: absent_from_later_narration

K.2 The independent-fact probe, precondition by precondition

Precondition Held?
The planted turn is outside the history window (planted depth 3, floor 72) yes
Absent from authoritative state yes
Absent from the summary yes
Absent from imported knowledge yes
Absent from later narration no — 6 violations, the first at turn 6

The narrator restated the planted fact in later narration, so the probe could not isolate memory as the only path. The run therefore records no independent recovery, and nothing here is relabelled as one.

K.3 The two statements the contract requires, kept apart

deterministic independent-memory recovery: PASS
reference-model independent-memory limitation: ACCEPTED RESIDUAL
  • Deterministic (WP-B.2 §I, and the suite on this candidate): the independent_full scenario fails on v1.0.0 at creation and returns recovered_through_memory_independent on the candidate, with isolation asserted every turn and provenance resolving to the planting turn.
  • Reference model: failed on the precondition-valid attempt in WP-B.2, and in this release run the attempt was not precondition-valid at all. The failing stage remains memory creation — the summariser's content selection.
  • No new regression. B2.1 ranking, B2.2 eviction and B2.3 excerpt creation all pass deterministically in the 1,723-test suite on this candidate, and the bank behaved normally through the run (19 memories, ranked and injected). What this run shows is the known summariser-quality limitation, not a fault in ranking, eviction or injection.

L. Identity diagnostic

Scripted half — complete. tools/m11_identity.py --scripted on the candidate: 10 identity stresses, 0 signals, verdict no objective identity defect detected.

The detector's negative control fires. With --inject, the same harness on the same candidate raises 7 signals — shared_display_name once and duplicate_character_creation on turns 5–10 — and preserves each turn's evidence. A clean run therefore means something: the check is capable of failing.

Model-backed half — complete, on the candidate, with memory on. Evidence: …/release-87a4032/identity-memory/.

Model qwen2.5:3b-instruct-16k (the reference narrator)
Context window 16,384
Memory status on — memoryBankEnabled: true, embedding model nomic-embed-text
Summary status on — autoSummarize: true; 1 summary written
Identity signals 0 across all 10 stresses
Stored protocol shapes 0 of 10 AI actions, by the release-gate detector
Fixture accepted successfully — five entities kept distinct (bill, alice, roger, john, office)
State proposals recorded and applied with no shared display name and no duplicate creation
Scripted detector self-test still fires (7 signals under --inject)

A harness correction made during validation, and what it does not invalidate. The diagnostic shipped with embedding_model="" and the memory bank switched off, so a release run of it would have reported a clean identity result with memory never taking part — which is not what the gate asks for. Two lines now read AIDND_TEST_EMBED_MODEL and enable the bank. Harness-only: no product code changed, so no product evidence became stale; the corrected harness repeated its own check, which is the run reported above.

One honest observation: with memory enabled the bank still wrote 0 memories in this 10-beat campaign — 21 actions is enough to pass MEMORY_START, but the summary pass is what ran and produced the single summary. Memory was configured and active; it was not meaningfully exercised here. The bank's real exercise is the 102-turn run (§K), which wrote 19.

No claim about the historical root cause. The post-M8 identity finding's campaign was destroyed and its cause cannot be established. A clean run here is evidence that the product does not do the things it can be blamed for on this fixture — not a discovery of what happened then.

M. Recovery

tools/m11_recovery.py against the integrated run's own bundle (3,071,683 bytes), imported into a database file that never existed, in a directory that never existed, by a second server process — so migrations ran from nothing and this is the fresh-install path as well as the import path.

16 checks, 16 passed, 0 failed.

Claim Result
The destination database did not exist beforehand PASS
The bundle imports into a clean directory PASS — 209 actions in the file, 121 on the active line
The active transcript is not empty PASS
Authoritative state came across PASS — 6 entities, 1 fact
The campaign's own canon came across PASS
The narration-length choice came across PASS
Redo availability matches what the file said PASS
Retained (undone) history came across PASS — 88 actions retained beyond the active line
Both Save Points restore PASS — On the ridge, Before the ridge
Every imported class came across, with content PASS — 3 sources
The moved campaign accepts a new change, unrefused PASS
The bundle carries no secret PASS
The moved campaign exports again, same story length PASS

One thing this does not prove, stated rather than implied. The long run ended with its head at the tip, so the file's head was at the tip and Redo was correctly unavailable after import. The undone-head case (I07) is proved by §N's upgrade campaign, which ends on an Undo with Redo available and survives both bundle directions, and by test_m9_portability.py in the suite — not by this bundle.

N. v1.0.0 upgrade compatibility

A campaign built and played by the 432f041 application in its own worktree, then opened by the candidate. Evidence: …/release-87a4032/upgrade/upgrade-report.json.

Phase 1 — v1.0.0 builds it. 20 turns accepted, 39 actions, canon knowledge imported, a Save Point taken, one Undo so the head is not at the tip, narration length chosen, memory bank and auto-summary on. It contains what §11 item 9 names: retained history, an undone head with Redo available, a Save Point, 6 memories, 2 summaries, imported knowledge, a narration-length choice and narrative state. Settings hold a loopback placeholder (http://127.0.0.1:11434/v1) — no real hostname is in the evidence database.

Phase 2 — the candidate opens the same file. Nothing was copied; the candidate's migrations ran against it.

Field Before After
transcript 39 actions identical
newest action (the head) — identical
total / can_undo / can_redo 39 / true / true identical
checkpoints 1 identical
narrative state 7 keys identical
memories 6 identical
summaries 2 identical
knowledge sources 1 identical
narration length brief identical
memory_bank_enabled / auto_summarize true / true identical
settings loopback placeholder identical
schema user_version 94 94

All 15 census fields compared identical; 0 differ. Schema parity is exact: v1.0.0 and the candidate both stamp user_version 94 with 93 migrations, so the upgrade required no migration at all, and nothing was rewritten in passing. A fresh-install database from the candidate carries the same 94 (§M's import ran migrations from nothing).

A first pass at this gate was discarded. It played 5 turns, which is below the memory and summary thresholds, so it compared 0 memories against 0 memories and proved nothing about two of the criterion's required contents; its settings also carried the live hostname rather than a placeholder. Both were corrected and the gate was rerun — the run reported above.

O. Bundle compatibility

Direction Result
v1.0.0 export → imported by v1.1 PASS — imported, id 2
v1.1 export → offered to v1.0.0 ACCEPTED — v1.0.0 imported it, id 1

Both directions were executed, not inferred from the unchanged format string. The format is ai-dnd-adventure-v3 on both sides, and it did not change during release validation.

Backward import succeeding means the compatibility question the brief raised — whether optional or additive v1.1 evidence data would break a v1.0.0 importer — is answered in the negative for this campaign's contents: v1.0.0 accepted the candidate's bundle whole. No bundle-format change was made or needed.

P. WP-D regression

Reconfirmed on the candidate; no backup-affecting product code changed after WP-D, so its 117 MB browser measurement is not repeated (the owner's brief permits this).

Claim Result
The completed backup copy is verified with PRAGMA integrity_check PASS — app/backup.py:193, docstring at 198–201 records why the full check replaced quick_check
The corruption fixture still separates the two pragmas PASS — quick_check → ok, integrity_check → row 145 missing from index i_t_k
Oversized export still succeeds and is delivered PASS
Importability metadata names the effective ceiling PASS — "This export is larger than this version's 20 MB import limit (… bytes). The file was exported successfully, but this version cannot import it."
Normal export unchanged in content PASS
Oversized import still refused PASS — 413 naming the limit
Suite 12 passed (tests/test_v11_d_recovery.py)

Q. WP-E regression

Claim Result
Contrast audit exit code 0
All applicable control boundaries ≥ 3:1 PASS — every boundary pair clears 3:1 (1.4.11)
Applicable text contrast still compliant PASS — every text pair clears 4.5:1 (1.4.3); baselines 14.57 / 13.57 / 5.48 / 5.88 unchanged
Browser boundary checks PASS — the 10 WP-E rows in §G
Focus visibility PASS — M11's visible-focus check inside the 38, plus WP-E's focused-edge measurement
Gate tests 11 passed (tests/test_v11_e_contrast.py), including 2.99:1 failing and 3.00:1 passing
OWNER SCREENSHOT APPROVAL: APPROVED

Approved by the owner in the release-validation brief of 2026-09-16. No visual code changed during release validation, so that approval remains valid; had any changed, it would have been void and new screenshots would have been required.

R. Release-shaped smoke test

The final no-cache candidate image, a fresh volume, published on loopback, with the private CA installed into the container's own trust store. Evidence: …/release-87a4032/smoke/.

15 checks, 15 passed, 0 failed.

Claim Result
The container starts PASS
The application answers on loopback PASS
The port is published on loopback only PASS — 8000/tcp -> 127.0.0.1:…
This machine's LAN address does not serve the application PASS
The first page loads PASS — HTTP 200
The shell references no remote origin PASS — none found
A CSP is served PASS
The approved HTTPS narrator verifies through its private CA PASS — HTTP 200 through tlstrust.ssl_context(), no bypass
A public endpoint is refused PASS — HTTP 400
A campaign is created PASS
One real narrator turn is accepted PASS
The container restarts and serves again PASS
The transcript survived the restart PASS
The narrative state survived the restart PASS
Firefox renders the reopened campaign PASS — 470 characters of story

A finding worth recording, and it is not a product defect. The first attempt failed at PUT /api/settings with HTTP 400. The cause: a .local name is mDNS, a Docker container has no mDNS resolver, and endpoints.py correctly refuses an endpoint whose address it cannot classify — the policy behaving exactly as designed. The fix is to resolve the name inside the container (--add-host), not to substitute the IP address, because the certificate is issued for the hostname and substituting the address would have quietly bypassed the hostname verification this test exists to prove. Harness-only; no product code changed.

This is supplemental evidence, not a substitute for the gates above.

S. Security / local-only review

Claim Evidence on the candidate
Served on loopback only every harness reached the application on 127.0.0.1; the documented container run publishes loopback
Endpoint policy app/endpoints.py admits loopback (v4 and v6), the three RFC1918 ranges, link-local, IPv6 unique-local and CGNAT, and refuses the public Internet; a name resolving to both a private and a public address is refused
Inference actually used trusted-LAN HTTPS with a private CA for the browser gate (§G); plain HTTP to a LAN GPU host for the long run, which SECURITY-THREAT-MODEL.md §83 permits and which is not A06 evidence
No secret in exports offline gate, WP-D tests and the M11 suite
CSP served default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none'; base-uri 'none'; form-action 'self'; frame-ancestors 'none'
Other headers x-content-type-options: nosniff, referrer-policy: same-origin, x-frame-options: DENY
No remote origin in the shell offline gate: the page names none, and every asset is local

'unsafe-inline' remains on style-src only, because React writes inline style attributes; it is deliberately absent from script-src.

T. Known residual risks

Each is classified, and none is collapsed into another category.

T.1 WP-B reference-model memory limitation — ACCEPTED RESIDUAL

deterministic independent memory: PASS
reference-model independent memory: FAIL
failing stage: memory creation — the summariser's content selection
owner decision: accepted for v1.1

Status on this candidate: unchanged, and no broader regression. The release long run did not demonstrate independent recovery, but it also could not: its probe failed the absent_from_later_narration precondition because the narrator restated the fact (§K.2). Ranking, eviction and excerpt creation all pass deterministically in the 1,723-test suite, and the bank worked normally through 102 turns (19 memories, ranked and injected). Accepted residual, carried visibly, not a blocker.

T.2 A2 mid-reply application-instruction echo — ACCEPTED RESIDUAL, still live

The known occurrence is WP-B.1's action 153 (depth 143): the narrator echoed the length hint mid-reply and then continued the story, which A2's trailing cleanup does not remove.

Reproduced on this candidate. Replaying that stored fixture through the release-gate detector still flags it — 1 of 105 AI actions, matched by the hard-limit/hint rule alone. The extractor still leaves it. This report therefore does not claim protocol leakage is solved.

But it did not recur in release evidence. This run's own 105 stored replies leak 0 (§J), and the browser run's narrated turns leak 0. The brief's stop-and-report condition — an equivalent shape occurring in this final run — was not triggered, so validation continues. No broader sanitizer was written; that remains for owner review.

T.3 Doubled full stop in memory-search scene text — ACCEPTED RESIDUAL

"rain outside.." when the state's scene summary already ends in punctuation. Only the embedding query sees it; the effect is one stray token. Not changed during release validation, deliberately: cleanliness is not a reason to alter product behaviour after evidence is taken.

T.4 K1 — "Correct" on an Important Facts row is always refused

Reproduced, unchanged on this candidate, deterministically and without a browser: Correct on a Characters row (subject='mara') applies, 201; Correct on an Important Facts row (subject='f1') is refused, 400 — "add_fact names subject='f1', which does not exist." The cause is frontend-side: the panel sends the row key as add_fact.subject, and the validator checks subject as an entity reference.

Classification: v1.2 backlog, not a release blocker. It blocks no v1 REQUIRED test — C04 passes through the working correction paths (§C) — and WP-C's State-panel correction coverage passes. It is a narrow bug with an obvious fix (offer Correct only against entities, or send facts without a subject), but fixing it during release validation would change product code after the evidence above was taken, which the brief forbids without a product brief. Left for the owner.

T.5 Import ceiling and scheduled backups — INTENTIONAL, not unfinished work

The 20 MB import ceiling is deliberate and unchanged; WP-D made it honest rather than raising it. Scheduled backups remain unbuilt by design. Neither is a residual defect.

T.6 Harness corrections made during validation — no product evidence invalidated

Three, all harness-only, each named where it occurred: the identity diagnostic did not enable memory (§L); the smoke test needed hostname resolution inside the container (§R); and the first upgrade campaign was too short to write memories and carried a live hostname (§N). No product code changed at any point during release validation, so no black-box or long-run evidence became stale. Each corrected harness repeated its own affected check.

U. Deferred v1.2 / future work

Item Why it is deferred
Raising the import ceiling, or a streaming import The plan assigns it to v1.2; WP-D's scope was honesty about the limit, not the limit
Scheduled backups; a restore button Explicitly out of WP-D's scope
K1's Correct-on-a-fact-row fix (§T.4) A narrow frontend bug needing a product brief
A broader mid-reply protocol sanitizer (§T.2) Needs owner review; A2's cleanup is deliberately trailing-only
The doubled full stop (§T.3) Cosmetic, embedding-query only
K05 generate local image, K06 multi-turn video FUTURE tests; both need a media provider this release does not build
Reference-model independent memory (§T.1) Needs a stronger summariser or a different creation strategy — a v1.2 investigation, not a v1.1 fix

V. Documentation changes

Current documents were brought up to date. No historical milestone report was rewritten, and no failed WP-B real-model evidence was turned into success.

Document Change
README.md Status now says v1.0.0 remains the released version, that all six v1.1 packages are complete and accepted, that release validation passed on candidate 87a4032, that WP-B ships with a documented limitation, and that no v1.1.0 tag exists and main is unchanged`. Also corrected a stale figure: the schema is versioned at 94, not "92 and counting"
planning/V1.1-PLAN.md WP-D/WP-E recorded as signed 87a4032; the release-validation outcome summarised with its residuals; the three owner events named as still outstanding
planning/VERSION.md Same status correction, plus a new revision entry for the closeout
planning/README.md Current-state paragraph rewritten for the same facts
planning/reports/v1.1/V1.1-WP-E-REPORT.md OWNER SCREENSHOT APPROVAL: PENDING → APPROVED, with the source (the release-validation brief) and date recorded, and a note that the signed commit predated the review
planning/reports/v1.1/V1.1-RELEASE-REPORT.md New — this document
DEVELOPMENT.md Unchanged: its WP-C/WP-D sections already describe the candidate as built

New tools committed with this closeout (harness only, no product code): backend/tools/v11_upgrade_check.py (Gate 9) and backend/tools/v11_release_smoke.py (§R), plus two narrow corrections to backend/tools/m11_identity.py (read AIDND_TEST_EMBED_MODEL; enable the memory bank) so the diagnostic can run with memory on.

W. Final release decision

The question this validation set out to answer: does candidate 87a4032 preserve the complete v1 contract and satisfy every accepted v1.1 package on one integrated release tree?

Gate Result
1 Package acceptance PASS — six packages accepted; WP-B's qualification carried whole (§B.1)
2 v1 acceptance contract PASS — 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified (§C)
3 Suites, lint, build, Docker PASS — 1,723 / 175 / 0 errors; image SPA file-for-file identical (§D, §E)
4 Offline / no-network PASS — 23/23 on the candidate image (§F)
5 Browser release run PASS — 101/0/0 over trusted-LAN HTTPS, every turn fits (§G)
6 Integrated long run PASS — 102 turns, M01–M04, 0 post-turn failures (§H)
7 Identity diagnostic PASS — 0 signals, 0 protocol shapes, memory on (§L)
8 Recovery PASS — 16/16 on the run's own bundle (§M)
9 v1.0.0 upgrade PASS — 15/15 identical, schema parity at 94 (§N)
10 WP-D regression PASS (§P)
11 WP-E regression PASS, screenshots approved (§Q)
Bundle compatibility Both directions execute and import (§O)
Release smoke PASS — 15/15 from the shipped image (§R)

A1 holds its reserve on every re-counted prompt, with a smallest margin of 879 tokens against v1's 23–42. A2's release-gate leak count is 0. WP-B's deterministic independent-memory recovery passes and its reference-model limitation remains an accepted, documented residual — stated in §K.3 in both halves, never shortened to "WP-B passed".

No product code was changed at any point during release validation, so no evidence was invalidated. Three harness corrections were made and each corrected harness repeated its own check (§T.6).

V1.1 RELEASE VALIDATION:
PASS

What this decision is not

These are separate, and only the first is done:

WP-A-E accepted:              YES
release validation passed:    YES
release candidate prepared:   YES  (87a4032, with this report)
release commit signed:        NO
main updated to v1.1:         NO
v1.1.0 tagged:                NO

The closeout changes are staged and uncommitted. Nothing was committed, pushed, merged or tagged by this validation. The owner's next decision is to review this evidence, resolve anything they disagree with, then sign the v1.1 release commit and publish v1.1.0.