Files
interactive-story/planning/reports/M11-IMPLEMENTATION-REPORT.md
T
JesseMarkowitzandClaude Opus 5 3652dc6fae Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed
The planning package still described M01 as outstanding. It now records the
evidence run on 96c1bf5 and the two product defects found on the way. It also
corrects three statements that were never true.

- V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as
  recovered through authoritative state, with the owner's acceptance of that on
  2026-09-13 and the positional precondition explained.
  Correction: v3.7 said this file carried M11 results against every REQUIRED
  test. None were written, and the per-test matrix is the M11 report's §F. The
  §P3 M11 disposition said the report records the identity diagnostic's
  findings. It does not, and the disposition now says so.
- BUILD-MILESTONES.md: the M11 status block records the long-run evidence,
  the write-lock and protocol-leak defects, and what is left for the reviewer.
- DATA-MODEL.md §28B: M11 added two columns, not one.
  settings.context_window_override (migration 94, ef25b0a) was never
  recorded.
- TECHNICAL-DESIGN.md: "Background failure observability" gains the rule
  that nothing in a turn writes before the model call, and new §15.4 records
  that stored narration carries story only, with the extractor's rules.
- CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same
  two fixes.
- README.md and VERSION.md: status, milestone map, stop rule, and the v3.9
  entry.
- M11 report §Q: the "not revised" note is replaced by what v3.9 revised.

No requirement changes. No code changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
2026-09-14 03:22:14 -04:00

101 KiB

M11 — v1 Security, Long-Run, and Release Validation

Implementation and verification report, written for independent release review.

Branch m11-release-validation, from the signed M10 commit 1013c94. Implemented and verified 2026-09-07. The long-run evidence was re-established, and this revision written, on 2026-09-14. This is the evidence package; it is not a record of acceptance, and nothing in it says M11 is accepted.


A. Executive result

PASS, for independent review. All 85 tests marked REQUIRED FOR V1 pass, with H09 recorded NOT APPLICABLE on the condition its own text states.

This revision supersedes the first. That revision left M01 PARTIAL at 41 accepted turns, and then the run it described was lost to a host crash along with its evidence. M01 to M04 now pass on a complete 100-turn campaign, run by the committed harness on commit 96c1bf5: 101 accepted turns, three genuine process restarts, every scheduled history operation performed, zero failed post-turn passes, and recovery onto a clean data directory 16 of 16 (§G, §K).

Getting there took six further runs. They found two more product defects, both fixed and both invisible to every earlier piece of evidence because the memory bank had never been switched on during a long run:

  1. A turn locked its own memory bank out (§O.7). Retrieval wrote a use counter before the model call, and the turn committed only after the reply, so SQLite's single write lock was held for the whole reply. Post-turn memory and summary writes timed out behind it, and recording those failures timed out too. A 26-turn run reported "complete" with two memories, no summary and 180 database is locked errors, while derived status read idle.
  2. The narrator's protocol leaked into stored story (§O.8). A small model pasted the narrative-state section and unfenced proposals into its prose, on 42 of 104 turns in one run, and stored text is replayed as history. The extractor now removes every shape observed. Replaying 443 real turns from five runs changed no turn the old extractor had left clean.

The first revision's three corrections stand: the context-window mismatch is resolved (§E), a partly refused manual correction now says so, and the narration-length setting moves a number. Post-M8 findings A and B are closed.

What is not claimed.

  • That memory keeps a planted fact on its own. M04 passes because the fact was recoverable with its planting turn 53 depths outside the history window. The recovery path ran through authoritative state: the narrator restated the state's fact line in its prose, and memory summarised those restatements (§G.4). No memory of the planting era carried the fact in any run. The repository owner accepted state-based recovery for M04 on 2026-09-13.
  • Headroom in the context window. The largest prompts fill 16,342 of 16,384 tokens by the narrator's own count, and the inference server cuts an over-window prompt without an error (§N, §P).
  • Performance. Timings are two particular hosts', not a product characteristic (§E.1).
  • Post-M8 finding D's root cause, which remains unestablished. The identity diagnostic's run results are not in this report (§D.1).
  • That the browser, offline and identity runs cover the later commits. They were re-run on the ef25b0a tree. The three product commits since then change the turn commit, memory retrieval and narration extraction (§B).

Zero requirement weakenings. No acceptance test was retired, relaxed or reclassified. The M04 precondition the harness measures was corrected to the acceptance text's own wording, "without entire transcript in prompt" (§O).


B. Repository and provenance

Base commit 1013c94eb1ad283e960114aef04c19c2806b5db7 — "M10: the seam for media, and no media"
Signature git verify-commit 1013c94 → Good signature, RSA key 02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569, "JesseMarkowitz", trust [ultimate]. %G? = G.
Branch m11-release-validation, created from that commit. The working tree was clean at the start (git status --porcelain empty).
HEAD at this revision 96c1bf5. Seven commits follow the base, all signed by the owner (%G? = G): 144406c (M11), fedb714, ef25b0a, fec46f6, f8d4010, 0c7316f, 96c1bf5. This revision of the report is staged for the owner to sign.
Upstream ancestry git merge-base --is-ancestor d72f7c1b HEAD → true. The AI-DnD fork point is still an ancestor.
LICENSE Unchanged — md5 07fde30437134836e2ee875e82a7cd31, MIT, "Copyright (c) 2026 Parth Thakkar". PROVENANCE.md unchanged.

The four trees, kept distinct, because M11's evidence discipline depends on which one produced a given number:

committed base   1013c94, signed by the owner: M1-M10 as accepted
working tree     the base plus M11's changes; this is what was tested
staged tree      identical to the working tree (§29 lists it)
frozen tree      the working tree at the point each black-box run started,
                 with no source edit during any reported run

The black-box runs come from two trees. The long-run evidence, the 100-turn campaign and its recovery run, was taken on commit 96c1bf5, after the last product change, with no source edit during the run. The browser regression, the offline container, the identity diagnostic and the contrast audit were re-run on the ef25b0a tree, as that commit's message records, and their evidence is dated 2026-09-10. The three product commits after it (f8d4010, 0c7316f, 96c1bf5) change backend files only, and those runs were not repeated (§P). Runs taken before a product change on their own tree were discarded and re-run; §O records which and why.


C. Change inventory

Product code — backend

File Change
app/contextwindow.py New. Discovers the server's real context window and provides the ceiling. The whole of §E.
app/context/builder.py Takes a window; the effective budget is min(configured, verified); the context report carries what was verified and how; length_hint gains the campaign's length band.
app/routers/adventures/turns.py Probes the window before assembling a turn, and stores the verdict in the turn's provenance.
app/routers/adventures/insights.py The same probe, so the inspector shows the prompt the next turn will actually send.
app/routers/settings.py The connection test reports the window, or why it could not be checked; changing endpoint or model clears what was learned.
app/routers/adventures/state.py A correction reports refusals (refused) and shared display names.
app/routers/adventures/crud.py Persists the campaign's narration-length choice.
app/narrative/model.py New duplicate_names — detection, deliberately not a refusal.
app/models.py, app/migrations.py, app/schemas.py, app/bundle.py adventures.narration_length: column, migration 93, API in and out, bundle carriage with an unknown value dropped.

Product code — frontend

File Change
src/documentTitle.js New. The product's name, in one place; the tab follows the open campaign.
index.html, src/pages/Play/index.jsx The inherited AI D&D title replaced and kept in step (finding A).
src/pages/Play/Composer.jsx, styles/story.css The position indicator: Moment 11 · later story ahead (finding B, §8A).
src/pages/NewCampaign.jsx, panels/CampaignSettingsPanel.jsx Narration length sent and editable as data (finding C).
src/pages/Settings.jsx, styles/library.css, styles/tokens.css The context-window report and warning; a --warning token measured at 7.85:1.
panels/StatePanel.jsx, styles/insights.css Refused corrections and shared names, shown to the reader.

Release-test infrastructure — no application code imports any of it

File Purpose
tools/m11_long_run.py The 100-turn campaign: M01-M04.
tools/m11_recovery.py That campaign, moved to a clean data directory: I01-I07.
tools/m11_browser.py, tools/m11_webdriver.py The browser regression and accessibility measurements; a dependency-free W3C WebDriver client.
tools/m11_offline.py A container with no network: §18 and the packaging path.
tools/m11_identity.py The multi-character identity diagnostic (finding D).
tools/contrast_audit.py The palette against WCAG AA.
tests/test_m11_context_window.py 20 tests: the probe, the cap, the truncation sentinel.
tests/test_m11_leakage.py 14 tests: E01-E04 together, in one long campaign.
tests/test_m11_scifi.py 10 tests: J01-J03, the Persephone fixture.
tests/test_m11_migration.py 9 tests: fresh-versus-upgraded schema parity, and the upgrade.
tests/test_m11_security.py 25 tests: the H-series against the assembled product.
tests/test_m11_findings.py 20 tests: findings C and D, and the refused-correction regression.
tests/test_m11_real_window.py 3 tests: the probe against a real Ollama (skipped without one).
tests/schema_rewind.py Migration 93's inverse, so the suite can replay it.

Dependencies: none added, none removed, none upgraded. requirements.txt, requirements.lock and package.json are byte-identical to M10's.

Changes after the first revision. 144406c is the first revision itself. Every later commit is listed here, and all are signed by the owner.

Commit Change Files
fedb714 Release evidence must be written to a durable --out, and the report recorded the lost run DEVELOPMENT.md, this report
ef25b0a The history window trims in blocks, so the server's prompt cache survives. settings.context_window_override (migration 94) covers a server the probe cannot ask. The long run gains --resume, a configurable turn timeout and an abort after consecutive failures. Two checks that could not fail were fixed. Every black-box harness was re-run on this tree app/context/builder.py, app/contextwindow.py, app/models.py, app/migrations.py, app/schemas.py, app/routers/settings.py, app/routers/adventures/turns.py, app/routers/adventures/insights.py, frontend/src/pages/Settings.jsx, tools/m11_long_run.py, tools/m11_browser.py; tests test_history_block_trim.py, test_m11_declared_window.py, test_m11_long_run_resume.py, settingsWindow.test.jsx; planning package v3.8
fec46f6 The long run switches the memory bank and auto-summarise on, and refuses a run with no embedding model tools/m11_long_run.py, tests/test_m11_long_run_memory.py
f8d4010 §O.7: no write lock is held across the model call, and a failure that breaks the session is still recorded. The harness reads the real section labels and stops on failed post-turn work app/memorybank.py, app/routers/adventures/turns.py, app/routers/adventures/insights.py, tools/m11_long_run.py, five test files
0c7316f §O.8: protocol is kept out of stored narration, and the fence-label bug is fixed. The harness gains an M04 verdict and a protocol-leak count app/narrative/extract.py, app/narrative/render.py, tools/m11_long_run.py, tests/test_narrative_state.py, tests/test_m11_long_run_memory.py
96c1bf5 §O.8 extended to a model's own sections under decorated headings. M04's precondition is now positional app/narrative/extract.py, tools/m11_long_run.py, tests/test_narrative_state.py, tests/test_m11_long_run_memory.py

No later commit touches requirements.txt, requirements.lock or package.json.


D. M10 and M8 handoff

Every item handed to M11, and what happened to it. Nothing here is closed by "the automated tests are green".

From M10's §O.1 (its own residual risk)

Item Disposition
K04 satisfied structurally, not physically — no media_jobs/media_assets Unchanged, and verified as such. M11 exercised the deferred-table path rather than building the tables: a dummy provider, a real packet, a fake asset carrying the packet's scene_id, and the story model byte-identical afterwards. Reported as K04 = PASS on the acceptance text's own deferred branch (§F).
ambience is an empty shape Unchanged. Filling it means extending set_scene, which is a prompt-path change and not release validation. Still an empty shape with the right fields.
The seam has no consumer, so it is unexercised by real use Partly answered. The Persephone fixture put a third genre through the packet and a visual profile through a starship hull (test_m11_scifi.py), and the offline container proved the media module imports and stays inert with no network. Still no real adapter, and that remains true until someone writes one.
A visual profile cannot be recovered from within the app Unchanged. It travels in the bundle; there is no undo for deleting one.
Profiles are API-only, with no reader-facing surface Unchanged, deliberately. M11 is release validation; adding a media UI would be new feature work.

From M9's residual risks

Item Disposition
Bundle ceiling ~279 turns Measured against the real 100-turn campaign rather than the fixture — §N gives the actual bundle size and what fraction of the 20 MB import limit it is.
quick_check rather than integrity_check Unchanged; re-exercised on a migrated database (test_m11_migration.py).
No scheduled backup Unchanged, and outside the acceptance contract.
Stale chunk_id in a restored snapshot Unchanged; a React key, not a live pointer.
The importing machine's context window may differ Closed. This was the same defect as M8's, seen from the import side, and §E closes both: the application caps to what the destination server accepts and says when it could not check. DEVELOPMENT.md's import warning is rewritten accordingly.
This machine cannot drive a file into or out of the browser Narrower than recorded, and half of it is closed. The snap Firefox refuses a WebDriver file path under /tmp; a path under $HOME works. Knowledge import is now proved end-to-end in a real browser (§L). The download half — a blob: export leaving the browser — is still not driveable here and is still recorded as a limitation.

From M8 (via M9 and M10)

Item Disposition
The deployment context ceiling Closed — §E. This was M11's stated release blocker.
Contrast and visible focus checked by eye Measured — §L. Palette pairs by calculation (tools/contrast_audit.py), rendered colours in a real browser, plus focus visibility, accessible names, tab order, hover-only controls and modal focus.
The four post-M8 playtest findings All four disposed of — §D.1 below.

D.1 The four post-M8 playtest findings

A — the tab read AI D&D. Fixed. The name is Interactive Story, chosen by the repository owner when M11 asked, and deliberately not "Adventure Storyteller": SPECIFICATION.md requires a genre-agnostic engine and Adventure is narrower than the thing it names. The tab shows the open campaign first (Westhaven — Interactive Story), and one module owns the string.

B — no orientation after Undo. Fixed, to BROWSER-UX-SPEC.md §8A. A status line at the end of the control row reads Moment 11 · later story ahead. The number and the clause are both the server's answers, it changes visibly after Undo, and it uses no implementation vocabulary. Asserted in the component suite and in the browser regression, the second of which is what §8A demanded when it said the requirement must be observable rather than inferable.

C — narration length had no measurable effect. Fixed, at the mechanism the finding identified. The choice is now data on the campaign, and length_hint turns it into a real word band (70-180 / 150-380 / 320-700), bounded by the reply cap. The generation budget is deliberately not touched: capping it per length would make a brief turn likelier to be cut off mid-sentence, and the state block is emitted last, so the first thing a truncated reply loses is the turn's state. Before M11 all three settings produced the identical sentence; test_the_three_lengths_no_longer_say_the_same_thing fails against that.

D — character identity confusion. The diagnostic exists; the root cause does not, and cannot. The diagnostic's run results are not reproduced in this report. The first revision pointed here to a section that was left empty. What is recorded is the defect the diagnostic caught in its own fixture (§O.2), which is why its first run's evidence was discarded. The later run's evidence is in the implementer's m11-evidence/identity-recheck directory (§P).


E. The context-window resolution

The original mismatch

M8 measured the reference deployment enforcing 4,096 input tokens while the application budgeted 16,384. M11 re-measured it, on the same server, before changing anything:

$ curl -sk https://<host>/api/show -d '{"model":"qwen2.5:3b-instruct"}'
  model_info["qwen2.context_length"] = 32768      the architecture's ceiling
  parameters                          = (none)    no num_ctx is baked in

$ (one /v1/chat/completions call to load it, then /api/ps)
  qwen2.5:3b-instruct   context_length = 4096   size_vram = 0

So the mismatch was live on the reference deployment on the day M11 started, and size_vram = 0 is why: with no VRAM Ollama picks a 4,096 default.

Root cause

Two independent facts that only bite together.

  1. Ollama's window is a property of how the model was loaded, not of the request. Its OpenAI-compatible endpoint accepts num_ctx — nested or top-level — returns 200, and ignores it; M8 established that, and it is why the operational fix is a model with num_ctx baked in or OLLAMA_CONTEXT_LENGTH on the server.
  2. The application had no way to know. Settings.context_token_budget was the only number in play, so the builder assembled to it and the server quietly did what it liked with the excess — which is drop the oldest tokens. The oldest tokens here are the system block: the narrator's rules and the campaign canon. The failure therefore looks like a narrator that stops respecting canon deep into a long session, with nothing on screen to explain it, and every acceptance test that reads a 200 as success passing throughout.

The fix

backend/app/contextwindow.py, plus four call sites. The rule:

A verified window is a ceiling on the configured budget. An unverified one leaves the budget standing and is recorded as unverified. There is no third behaviour.

  • Discovery asks the server the application is already talking to, on the native path beside /v1, through endpoints.rejection_reason and the shared TLS trust store — so it can reach exactly what a turn can reach and nothing more. /api/ps gives the window a resident model is actually being served with; /api/show gives the num_ctx an unloaded one will load with, capped by the architecture's own ceiling.
  • Enforcement is one line in the builder: the budget every section is priced against is min(configured, verified). Because the output reserve is subtracted from that budget by the existing arithmetic, the assembled prompt plus the reply reserve fits inside the window by construction.
  • Reporting. The context report and the turn's stored snapshot carry window: {verified, tokens, source, model_max, detail, capped}, so an old turn can be asked afterwards whether it was built against a checked window. The connection test in Settings shows the number or explains why it could not be checked, in three distinct messages, because "could not check", "smaller than your budget" and "fine" need three different things done about them.
  • What it does not do: hard-code 4,096 (right on one machine, wrong on the next), raise anyone's window, guess from a model's name, or add a provider abstraction. An unknown window is reported as unknown.

The numbers, on the reference deployment

Ollama 0.33.0, CPU-only (size_vram = 0)
qwen2.5:3b-instruct 4,096 tokens, source loaded (/api/ps)
qwen2.5:3b-instruct-16k 16,384 tokens, source parameters (num_ctx baked in), architecture ceiling 32,768
Application budget (M01) context_token_budget = 16,384, max_output_tokens = 500
M01's effective budget 16,384 — verified equal to the server's window, so nothing was capped away
Largest assembled prompt in M01 see §N

Measured end to end in tests/test_m11_real_window.py against the real server:

turn 1  window verified=False  (the model is not resident yet)
turn 2  window verified=True  tokens=4096  source=loaded
        budget configured=16384 effective=4096
        prompt 850 tokens + 464 reserved  ->  1314 <= 4096

That is the whole behaviour in six lines: the first turn on a cold model cannot verify and says so; from the second turn the cap is live and the prompt provably fits.

Why silent truncation can no longer invalidate M01

Three separate reasons, and the third is the one that matters:

  1. M01 ran on the 16k model, whose window (16,384) equals the application's budget, so nothing was capped and nothing was near the edge — recorded per turn in timeline.jsonl as window_verified and window_tokens.
  2. Every M01 turn recorded the verification, so the claim is per-turn evidence rather than a statement about the configuration at the start.
  3. The failure mode is now impossible to reach silently. If the window were smaller than the budget, the prompt would be built to the window, and the canon at the front would survive by construction — test_the_canon_at_the_front_survives_a_window_far_too_small plays 120 turns into a 4,096-token window and finds the canon sentinel still present with the oldest history dropped instead. Its companion, test_without_the_cap_the_same_prompt_would_have_overflowed, builds the same campaign with no verified window and measures a prompt more than twice the size — the defect, reproduced, so the fix is shown to be doing something.

E.1 The release environment

Recorded because a timeout without hardware beside it is not a measurement. Two inference hosts produced this report's evidence, and every result below says which.

The application machine

It runs the storyteller, every harness, the browser and the container.

OS Ubuntu 24.04.4 LTS, kernel 7.0.0-30-generic, at the first revision; Ubuntu 24.04.5 LTS, kernel 7.0.0-31-generic, for the 2026-09-13 long runs
CPU AMD Ryzen 7 7840HS, 4 cores available to this VM, 1 thread per core
GPU none: VMware SVGA II
RAM 15 GiB
Python 3.12.3
Node / npm v22.23.1 / 10.9.8
Docker 29.7.2
Browser Firefox 154.0.1, headless, via geckodriver over W3C WebDriver

The CPU reference host

Used from 2026-09-07 to 2026-09-10: the lost run, the memory-off 100-turn run, the browser run and the identity diagnostic.

Machine a separate physical machine on the trusted LAN, no GPU (size_vram = 0)
Ollama 0.33.0, HTTPS with a private CA installed in the application machine's OS trust store
Measured cost 88-203 s per turn at a 4,096 window; 229-291 s at 16,384

The GPU inference host

Used on 2026-09-13 and 2026-09-14: both 26-turn trials and the three 100-turn runs, including the evidence run.

Machine a separate physical machine on the trusted LAN; Ubuntu 24.04.2 LTS; 16 logical CPUs; 14 GiB RAM
GPU NVIDIA GeForce GTX 1080 Ti, 11 GiB (Pascal, compute capability 6.1), driver 580.173.02, in an OCuLink dock: PCIe x4, measured at Gen 3 under load and Gen 1 at idle
Ollama 0.34.0. Its CUDA 13 runner skips a compute-6.1 card and the CUDA 12 runner serves it; both models loaded 100% into VRAM
Transport plain HTTP on the LAN. H12 permits a LAN address; this host is not A06 evidence
Power limit the card's default, 280 W, during every run
Measured cost prefill ~1,350-1,850 tokens/s and generation ~87 tokens/s with synthetic prompts; 4.1-22.8 s per turn in the evidence run

Models and budget

Narrator qwen2.5:3b-instruct-16k: num_ctx 16,384 baked in, architecture ceiling 32,768. Digest 21ff8cc52f37, identical on both hosts
Base narrator qwen2.5:3b-instruct, digest 357c53fb659c, identical on both hosts. With no num_ctx the server gives it 4,096
State/summariser model the narrator; no separate summariser is configured
Embedding model nomic-embed-text, digest 0a109f422b47, identical on both hosts
Application context budget context_token_budget = 16,384, max_output_tokens = 500, reply reserve 564
Effective budget 4,096 in the lost run, where it was capped every turn; 16,384 in every later run, where window and budget were equal

Why the first campaign ran at 4,096. On the CPU host the recommended 16,384 window cost 229-291 s a turn, about eight hours for a hundred. The first revision therefore ran the default window, which doubled as the longest test of §E's cap. That run reached 97 of 100 turns and was lost (§G.6). Every later campaign runs the recommended window.

The GPU host incident

About 30 seconds after the evidence run's last post-turn pass finished, the GPU dropped off the PCIe bus: NVRM: Xid 79, GPU has fallen off the bus, then Xid 154, reset required. Ollama stayed up without a GPU, stopped completing requests, and the host had to be rebooted. The same host had logged one Xid 13 graphics exception from the inference server earlier that day. There was no out-of-memory event and no thermal event recorded. The firmware does not support PCIe AER, so no link-error trail exists.

After the reboot, under a temporary 200 W cap with logging running, the link held Gen 3 x4 under load, peak draw was 207.8 W, and there was no fault. The cause is not established. Power transients at the uncapped 280 W over a long run are the leading candidate. The evidence is unaffected: every record the result rests on was written before the drop (§G.1).

Required for every future long run on a GPU host: log the card's power, temperature, utilisation and PCIe link state, and the kernel's and the inference server's messages, to disk for the whole run. A recurrence can then be tied to power or ruled out. The commands are in DEVELOPMENT.md, "Logging the inference host during a long run".


F. Acceptance matrix

Every test currently marked REQUIRED FOR V1, individually. Evidence type is browser (real Firefox, frozen build), campaign (the 100-turn run against a real narrator), container (no-network Docker run), process (spawned server processes over HTTP), suite (automated tests), or a combination. A REQUIRED test is not marked PASS on source inspection alone; where inspection is the only evidence, it says so and the verdict is qualified.

A — Local-first operation

ID Result Evidence
A01 Start application offline PASS container: first page load from a fresh volume with no route and no DNS; every referenced asset served locally
A02 Storyteller loopback default PASS suite test_local_only_surface.py (start scripts, compose publishes 127.0.0.1:8000:8000); every M11 harness reached it only on loopback
A03 No cloud API key PASS suite: no api_key in settings, none settable through the API, no Authorization header; container: no secret in an export
A04 Campaign survives restart PASS campaign: 3 genuine process restarts (4 process starts), with transcript/head/state/Save Points/knowledge/settings compared before and after each and identical every time (§G.2); container: campaigns survive a container restart
A05 Failed model call does not corrupt story PASS campaign: a real failed call at a controlled point (the model name pointed at one the server does not serve), then play resumed; container: same with no model reachable at all — no narration accepted, state unchanged, the player's words kept, earlier story reachable
A06 Trusted-LAN Ollama inference PASS browser + identity diagnostic + the CPU-host campaigns: Ollama on a separate physical machine over HTTPS with a private CA installed in this machine's OS trust store, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1)

B — Core play

ID Result Evidence
B01 Natural language action PASS browser (two real turns through the UI) + campaign (100+)
B02 Dialogue input PASS campaign: dialogue beats are part of the fixture's turn list
B03 Continue PASS suite test_turn_flow_integration.py; browser: the Continue control is present and enabled
B04 Story direction (SHOULD) PASS suite; browser: the direction toggle and its hint

C — Story authority and state

ID Result Evidence
C01 Campaign canon is preserved PASS campaign: the canon section is in the prompt on 101 of 101 turns (canon_tokens, §G.3); suite
C02 Possession state PASS campaign (the silver key) + suite + sci-fi fixture (the data crystal)
C03 Character knowledge is not invented PASS suite test_worldstate_integration.py, test_narrative_state.py
C04 Manual state correction PASS campaign: two corrections, both accepted, one carrying the planted clue; suite, including the M11 regression that a partly refused correction now says so
C05 Canon beats reference PASS suite test_knowledge_calibration.py, test_imported_knowledge.py
C06 Structured state matches accepted narrative consequence PASS campaign: real state extraction across 100 turns against the reference narrator, with every accepted event validated and every refusal recorded; suite test_narrative_realistic.py against a real model

D — Non-destructive history

ID Result Evidence
D01 Undo one turn PASS browser + campaign
D02 Minimum five undos PASS suite test_head_cursor.py; campaign (undo/redo and undo/diverge sequences)
D03 Unlimited undo (SHOULD) PASS suite: undo to the root and back
D04 Redo PASS browser (returns to the same position) + campaign
D05 Redo invalidated by new continuation PASS campaign: after diverging, redo is no longer available — recorded in the timeline
D06 Retry narrator response PASS campaign: two retries at scheduled points (§G.2)
D07 Select prior retry take PASS campaign: take selection back to index 0
D08 Retry does not delete prior take PASS campaign: take count on the turn after retry; suite
D09 Edit earlier user input PASS suite test_take_edit.py, editRouting.test.jsx
D10 Edit narrator output PASS suite; browser (the hostile-Markdown scenario plants text through the narrator-edit path)
D11 Named checkpoint PASS campaign: two named Save Points; suite; process-restart suite
D12 Restore checkpoint PASS campaign: a restore at a scheduled point; suite
D13 Restore does not delete later history PASS campaign: retained actions after the restore; suite
D14 Delete checkpoint PASS suite test_save_points.py; browser: the delete confirmation dialog

E — Branch and derived-data isolation

ID Result Evidence
E01 Abandoned future cannot affect active state PASS suite test_m11_leakage.py (§H), with a positive control
E02 Abandoned memory cannot leak PASS as above; retained on disk, absent from the prompt
E03 Abandoned summary cannot leak PASS as above — a summary regenerated after the divergence, which is the shape M6's review established
E04 Scene state is lineage-safe PASS as above, including M10's derived Scene Packet

F — Long-term memory and context

ID Result Evidence
F01 Recent turns remain coherent PASS campaign: the history window is populated every turn and the newest turns are always included
F02 Old important event retrieval PASS campaign M04 (§G.4): the planting turn outside the history window and the fact recovered, through authoritative state and the narrator's restatements rather than independent memory retention
F03 Prompt remains bounded PASS campaign: prompt size across 100 turns (§N), plus the cap itself (§E)
F04 Output token reserve PASS campaign: output_reserve present in every turn's measurement and subtracted before history is chosen
F05 Prompt inspector PASS browser: the context panel shows the assembled prompt and its budget
F06 Retrieval provenance PASS campaign: knowledge.used per turn in the stored snapshot; suite
F07 Heuristic memory is not canon PASS suite test_memory_nodes.py (authority)
F08 Memory failure is non-fatal PASS suite test_context_memory.py, including §O.7's regressions that a failure is recorded even when it breaks the session; container: derived work fails with no model and turns still commit; campaign: post-turn work checked after every turn, 0 failures

G — Imported knowledge

ID Result Evidence
G01 Import local text PASS browser: through the real file input; container: offline
G02 Import local Markdown PASS campaign: three sources imported (canon, reference, inspiration); browser
G03 Classification PASS campaign: all three classes present after the move (§K); suite
G04 Disable knowledge source PASS suite test_change_visibility.py, test_imported_knowledge.py
G05 Canon retrieval PASS campaign: canon passages in stored prompts; suite
G06 Reference retrieval PASS sci-fi fixture (spin gravity) + suite
G07 Inspiration is low authority PASS suite test_knowledge_calibration.py
G08 No automatic URL fetch PASS suite test_egress.py; container: no network at all and import still works
G09 Remote Markdown image does not auto-load PASS browser: no http image src in the rendered story
G10 Prompt injection in source is treated as data PASS suite test_imported_knowledge.py; browser: injection text rendered as text

H — Security

ID Result Evidence
H01 No unexpected outbound connections PASS container (no network at all) + the CPU-host campaign (one destination, the configured Ollama; the later runs were not network-monitored, §J) + suite test_egress.py
H02 No telemetry PASS suite + dependency audit (§M)
H03 No cloud provider required PASS container: a full campaign offline; suite
H04 Model output cannot execute shell PASS browser (shell text rendered as text) + suite (no subprocess/eval anywhere in the turn path)
H05 Invalid state event rejected PASS suite: unknown type and unknown reference both refused with the document unchanged
H06 Stored XSS protection PASS browser: onerror and <script> in accepted narration, neither executed
H07 JavaScript URL protection PASS browser: no javascript: href in the DOM
H08 Path traversal import rejected PASS suite: a ../../../etc/cron.d/... filename stored as metadata, no file written
H09 ZIP Slip protection NOT APPLICABLE the product extracts no archives, and a test enforces it (§J)
H10 Restrictive CORS and local API behaviour PASS process: a wildcard origin refuses startup, a named origin does not; browser: an unknown API path is a 404 with a non-HTML body
H11 No first-use runtime asset download PASS container: every referenced asset served locally with no network; suite: the tokenizer table is vendored and no HTTP client is in that module
H12 Inference endpoint enforcement PASS suite: loopback v4 and v6, LAN address, CGNAT allowed; cloud hosts and public addresses refused; a public endpoint written into the database behind the API is refused at request time; the M11 window probe obeys the same policy. A06 covers the LAN-hostname-over-HTTPS case in production

I — Export, import and recovery

ID Result Evidence
I01 Export campaign PASS campaign: the 100-turn campaign exported (§K)
I02 Import exported campaign PASS process: imported into a database that never existed, in a directory that never existed (§K)
I03 Branch/disposable history export PASS §K: retained history present after the move, Redo walks into it
I04 Checkpoint export PASS §K: Save Points restore to the positions they name after the move
I05 Knowledge provenance export PASS §K: all three classes with their content after the move
I06 Database/export contains no API secrets PASS suite + container + §K
I07 Export/import preserves an undone active head PASS §K, and suite test_m9_portability.py for the older-format seams

J — Genre neutrality

ID Result Evidence
J01 Science-fiction campaign PASS suite test_m11_scifi.py (§I)
J02 Generic entity support PASS as above: five entity types in one document, no schema change
J03 Genre profiles are configuration PASS as above, plus the whole-vocabulary check that no event type names a genre noun

K — Future media architecture

ID Result Evidence
K01 Scene snapshot exists PASS suite test_m10_*; sci-fi fixture; campaign (the scene follows the active line through every history operation)
K02 Visual character profile PASS suite; sci-fi fixture
K03 Visual location profile PASS suite; sci-fi fixture (a starship hull)
K04 Attach media asset to scene (SHOULD) PASS on the deferred branch suite test_m10_authority.py: a dummy provider produces an asset carrying the packet's scene_id and the story model is byte-identical afterwards. Media tables remain deliberately unbuilt; a reviewer requiring physical tables should read this as PARTIAL

L — Data integrity

ID Result Evidence
L01 Atomic turn commit PASS container + campaign: real induced failures; no narration accepted, no half-written state, earlier story reachable, play resumes
L02 State reconstruction PASS suite; campaign: state compared across every restart
L03 Checkpoint reconstruction after restart PASS process-restart suite; campaign: Save Points present and restorable after each restart
L04 Derived data can be rebuilt (SHOULD) PASS suite test_knowledge_migration.py (reindex), test_memory_rewrite.py

M — Long-run

ID Result Evidence
M01 100-turn campaign PASS campaign: 101 accepted turns, every scheduled operation, 0 failed post-turn passes (§G)
M02 Restart during long campaign PASS campaign: 3 genuine process restarts, one carrying retained history, all identical (§G.2)
M03 Long-run context stability PASS campaign: the prompt held between 14.6k and 15.7k tokens for 70 turns, the window verified 101/101, the canon present on every turn (§G.3, §N)
M04 Long-run memory recall PASS campaign: the planting turn outside the window, the fact recovered through state and through memory of the narrator's restatements (§G.4)

SHOULD and FUTURE disposition

SHOULD tests run and passing: B04, D03, K04 (on its deferred branch), L04.

FUTURE tests not run, and deliberately: K05 (generate local image) and K06 (multi-turn video request). Both require a media provider, which §5 of the M11 brief forbids adding and M10 deliberately did not build. They are not v1 blockers and no part of this milestone treats them as one.


G. The long-run campaign — M01 to M04

M01 to M04 pass on a complete 100-turn campaign run by the committed harness. This section reports that run (§G.1-§G.5), then every earlier run and why none of them is the evidence (§G.6).

G.1 The evidence run

Tree commit 96c1bf5, signed; tools/m11_long_run.py --turns 100
Inference the GPU inference host (§E.1), qwen2.5:3b-instruct-16k, window 16,384
Evidence $HOME/m11-evidence/m04-final/: timeline.jsonl, summary.json, recall.json, bundle.json, campaign.db, server.log, recovery-report.json
Status complete; no aborted reason, no failed reason
Accepted turns 101: 100 scheduled beats and the recall turn
Genuine process restarts 3, making 4 process starts
Wall clock 1,238 s; a turn took 4.1 s at least, 9.6 s median, 22.8 s at most
Memory bank and auto-summarise switched on at setup and read back as on
Post-turn work checked after every turn against derived status and new server.log lines: 0 failed passes, 0 database is locked, 0 unrecorded failures, 0 tracebacks
Written 33 memories in the database across every branch, 19 reported by the /memories endpoint at the end; 12 summaries
Protocol in stored narration 0 of 104 AI turns

Timing against the GPU fault (§E.1). The last AI action was written at 02:31:19 UTC, the memory, summary and embedding passes finished at 02:31:22, and the GPU dropped off the bus at 02:31:52. The run's final settle and summary followed; nothing they report was pending.

G.2 Timeline of operations

"At turn" is the accepted-turn count when the operation fired.

At turn Operation Outcome
0 memory bank and auto-summarise both read back as on
0 knowledge import canon, reference and inspiration sources
0 state correction 8 events, the opening cast
1 state correction, clue planted the clue accepted as state; verified in state; the planting turn at depth 1
6 Save Point Before the ridge
13 restart 1 identical: true
20 Undo, then Redo 41 → 39 → 41 actions, position restored
27 Retry two takes on the newest turn
34 Save Point On the ridge
41 Undo, then restart 2 with the undone history retained identical: true, Redo available on both sides
48 Retry two takes on the newest turn
56 Undo, then new writing Redo no longer available
62 take selection index 0 of 2 chosen, and it became live
69 failed model call reported: HTTP 404 for a model the server does not serve; state unchanged
76 restore On the ridge active line 148 → 69 actions; the later story retained; Redo available
83 restart 3 identical: true
100 recall check §G.4
101 export 2,743,080 bytes; 0 AI turns carrying protocol

G.3 Context growth

The application's token counts, from timeline.jsonl.

 turn  actions  in prompt   prompt  summary  memory  knowl  state  canon  floor  window
    1        3          3    1,441        0       0    127     93     57      -  16384
   10       21         21    6,422      185     423    330    137     57      -  16384
   20       41         41   11,447      603     544    230    161     57      -  16384
   30       61         55   15,299      275     700    230    165     57      6  16384
   40       81         57   15,300      289     809    230    194     57     24  16384
   50       99         57   15,474      395     946    230    194     57     42  16384
   60      115         55   14,718      374     875    127    194     57     60  16384
   70      136         58   15,549      244   1,000    127    194     57     78  16384
   80       77         59   15,203      289     664    127    164     57     18  16384
   90       97         61   15,504      217     764    227    193     57     36  16384
  100      117         63   15,745      146     757    127    193     57     54  16384
  • budget 16,384 (configured 16,384); reply reserve 564 on every turn
  • window verified on 101 of 101 turns
  • campaign canon in the prompt on 101 of 101 turns; imported knowledge on 101; memories on 94; a summary on 93
  • largest assembled prompt 15,745 tokens by the application's count; the history window was first trimmed at turn 30
  • the drop in actions at turn 80 is the Save Point restore at turn 76

G.4 Recall (M04)

The precondition is positional. M04's pass text is "Fact/event remains recoverable without entire transcript in prompt", so the harness asks whether the planting turn has left the history window. It does not ask whether the clue's text has, because the narrator reuses that text in its own prose (below).

Planting turn's depth 1
History window floor at the recall check depth 54
Planting turn in the history window no
Clue's sentinel text somewhere in recent history yes: the narrator's restatements, below
Clue in the narrative-state section yes
Clue in the memories section yes
Clue in the summary section no
Clue in imported knowledge no
Fact still in authoritative state yes
History after the recall turn 59 of 119 actions, in a 14,638-token prompt
Verdict recovered_through_memory_or_summary

How the fact reached memory. From depth 30 onward the narrator wrote the state section's fact line into its prose as an ordinary sentence, on 74 of 104 AI turns:

Aldric knew the silver key opened the crypt beneath it (SILVER-KEY-CRYPT-OLD-ABBEY).

The memory pass summarises narration, so the 16 memories that carry the clue all summarise depth 54 or later. None comes from the planting era. The chain is authoritative state, then the narrator's restatement, then memory. The fact survived because state carried it. The repository owner accepted state-based recovery as satisfying M04 on 2026-09-13.

A restatement sits inside a line of story, so the extractor rightly leaves it. §P records what it means for narration quality.

G.5 What the numbers say

M01, 100-turn campaign: PASS. 101 accepted turns with none refused. Every one of the thirteen scheduled history operations performed as designed. All three restarts compared transcript, head, Undo/Redo availability, scene, entities, facts, Save Points, imported knowledge and settings and found them identical. The induced failed call changed nothing. The campaign moved to a clean data directory 16 of 16 (§K). The step list's summary/memory activation was exercised: 33 memories and 12 summaries were written with no failed pass.

M02, restart during a long campaign: PASS. Three genuine uvicorn process boundaries, one of them carrying undone history across. Each was identical.

M03, long-run context stability: PASS. The prompt reached about 15.3k tokens at turn 30 and stayed between 14.6k and 15.7k for the next 70 turns, while the active story grew to 136 actions. The canon was present on every turn, and the reply reserve was subtracted on every turn. By the narrator's own tokenizer the largest prompt plus the reserve is 16,342 of 16,384, which is bounded but nearly full (§N).

M04, long-run memory recall: PASS, on §G.4's positional precondition, with the recovery path stated there.

G.6 The runs that led here

Every run's evidence directory is kept.

# Run Tree Host Result Why it is not the evidence
1 first release campaign, 2026-09-07 first revision CPU, 4,096 window reached 97 of 100 turns the host crashed; its evidence was under /tmp and the reboot cleared it. The harness now requires --out and checkpoints for --resume
2 100 turns, 2026-09-10 ef25b0a CPU, 16,384 complete: 101 turns, 3 restarts, 45,656 s the memory bank and auto-summarise were never switched on (harness defect, §O), so summary/memory activation went untested and recall ran through state only
3 26-turn trial, 2026-09-13 fec46f6 GPU reported complete 180 database is locked, 20 failures that could not be recorded, 2 memories, 0 summaries. Found §O.7 and two harness defects
4 26-turn trial f8d4010 GPU 0 lock errors, 7 memories, 2 summaries, 258 s a trial, not a campaign
5 100 turns f8d4010 GPU complete: 101 turns, 1,289 s, 33 memories, 12 summaries the narrator's protocol was stored as story on 42 of 104 turns (§O.8), and a pasted state section kept the clue in recent history, so M04 proved nothing
6 100 turns 0c7316f GPU complete: 101 turns, 1,251 s, 1 leak by the harness count the narrator wrote ## Established: sections of its own on 5 turns, which the extractor and the harness's leak count both missed; verdict precondition_not_met on the sentinel-text precondition. Reclassified on the positional precondition, in recall-reclassified.json beside the original, as recovered_through_state_only. Superseded by run 7 on the corrected tree
7 the evidence run 96c1bf5 GPU §G.1-§G.5

H. Lineage leakage — E01 to E04

tests/test_m11_leakage.py, 14 tests, all passing. What M11 adds to the existing E-series coverage is that all four leaks are exercised together, in one campaign, under long-story conditions — 22 turns on the abandoned line with state, memories, summaries and knowledge all live, then 15 on the new one — because the four share one mechanism and a campaign with only one of them cannot show the mechanism holding for one and failing for another.

Four sentinels, one per class. Every negative control has a positive control that fails loudly if the fixture did not actually establish the thing:

Positive control (path A) Negative control (path B)
E01 state the fact is in the document, and in the prompt absent from the document, absent from the prompt
E02 memory the memory reaches path A's prompt absent from the prompt and from memories.used; still on disk, because the story was left, not erased
E03 summary a summary exists on path A and contains the sentinel a new summary row was generated on path B; it carries nothing from A; the summariser was never offered A's summary; no A turn is on B's lineage; A's row is retained but ineligible
E04 scene the abandoned line moved to the crypt, in state and in M10's packet the current scene is the tavern; the protagonist's location followed the active line; the derived Scene Packet shows the active line and a different scene_id; the abandoned scene is still retained at its own position

E03 is the one with history: M6's review found the first implementation passing while the defect was live, because the test checked only that the old row was ineligible. The shape this file uses is the one that review demanded — the summary is regenerated after the divergence — and the assertion that matters most is that the summariser's input never contained the abandoned prose. A filter over the output would be a different bug.

E04's Scene Packet assertion cannot fail while the state assertion passes, since the packet is derived from the state on read. It is asserted anyway, because the packet is a surface that did not exist when E04 was written, and a later change that gave it a store of its own would fail here.


I. Fantasy and science fiction

The fantasy Continuity Test is the foundation of the 100-turn campaign (§G): Westhaven, the Crooked Lantern, Aldric, Mara, Edrin, the silver key, the sealed abbey crypt, and the canon that the dead do not return.

The Persephone science-fiction fixture — tests/test_m11_scifi.py, 10 tests, all passing — is TEST-CAMPAIGN-FIXTURE.md §31's, with its three hard-technology canon rules and its full cast.

What is actually being checked is not that a science-fiction story can be told, but that no code path knows the difference:

Check Result
Five entity types in one document — character, vehicle, location, item, organization all present, all through the same create_entity
The type list is suggested, not closed SUGGESTED_TYPES — a genre needing a type nobody listed uses one without a migration
Where a thing is lives on the entity the same field puts Aldric in a tavern and Imani aboard a ship
Possession the data crystal is Imani's, through the same set_possession
Canon reaches the prompt as the campaign's highest authority "FTL does not exist" in the campaign_canon section
M10's Scene Packet describes a starship under spin with no field it did not already have
A visual profile holds hull: pitted white composite as readily as a face
Reference retrieval a science-fiction query retrieves the spin-gravity passage
The bundle same ai-dnd-adventure-v3, vehicle type intact on the far side
The event vocabulary contains no genre noun — no spell, no sword, no warp, no airlock

No schema change, no code change, no new event type was required for the science-fiction fixture. J03's claim — genre is configuration — holds in the strong form: the same schema, the same validator, the same builder, the same bundle.


J. Offline and security

The offline run

tools/m11_offline.py — 23 checks, 0 failed. A container built with docker build --no-cache and run with --network none: a loopback interface and nothing else, no resolver, no route, and a fresh volume. The exercise runs inside over docker exec, because with no network there is no published port to reach — that is the only honest way to drive an isolated process.

unshare -rn was the first choice and is unavailable here: Ubuntu 24.04 sets kernel.apparmor_restrict_unprivileged_userns=1. The container gives the same isolation and doubles as §24's packaging evidence.

The isolation is real a TCP connection to 1.1.1.1 fails; getaddrinfo("example.com") fails
First page load, from fresh data succeeds; names no remote origin; a CSP is served
Every asset the shell references served locally — none remote, none missing
Campaign creation, state extraction work
Knowledge import, prompt assembly, retrieval work
A turn with no model reachable reported as a failure; no narration accepted; the player's own words kept (A05); state unchanged; earlier story still there
Export and import work; no secret in the bundle
M10's media module imports; registry empty; a scene packet builds; no provider required
Container restart campaigns survive on the volume

Inference is not exercised offline, and that is stated rather than implied. This deployment's Ollama is on the trusted LAN, which SECURITY-THREAT-MODEL.md §73 permits and which is not an Internet dependency — but it is also unreachable from a container with no network. What the offline run proves about inference is the useful half: with no model reachable the application degrades to a reported error and the campaign stays intact.

Observed outbound destinations

During the first revision's long campaign, on the CPU reference host, the application contacted exactly one host: the configured trusted-LAN Ollama, over HTTPS with a private CA in the OS trust store. It used the /v1 path for inference and the /api/ps and /api/show paths for the window probe. There was no other destination, and no DNS lookup for any other name. In the container run there was no destination at all, because there was no network.

The later long runs were not network-monitored. They were configured with one endpoint, the GPU inference host over plain HTTP. tests/test_egress.py is what continues to assert that nothing else is contacted.

H-series

Full per-test results are in §F. The M11-specific additions:

  • H09 is NOT APPLICABLE, and the condition is now enforced. Its own text makes it conditional on ZIP import/export existing. Nothing in the application opens an archive — the bundle is JSON, an imported source is a single file — and test_h09_the_product_extracts_no_archives scans every module for zipfile, tarfile, unpack_archive, py7zr and rarfile, so the day that stops being true H09 becomes required again.
  • H10's startup refusal is proved by starting a process, not by importing a module: AIDND_CORS_ORIGINS=* makes the application refuse to come up, and a named origin is accepted, which is the control that makes the first assertion about the wildcard rather than about the variable.
  • H12's tampering case writes a cloud endpoint into the settings row behind the API, and the request-time check still refuses it — ADR 011's point being that the check is not only at the front door.
  • The M11 window probe is held to the same policy, verified with no transport installed, so a probe that ignored the policy would attempt a real connection and be caught.
  • Hostile content in the browser (§L): an onerror image, a <script> tag, a javascript: link, a remote image and shell text all reached the real renderer as accepted narration. None executed, none loaded, none became markup.

K. Recovery — export, import, migration

The long campaign, moved to a machine that has never seen it

tools/m11_recovery.py: 16 checks, 0 failed. The input is not a fixture. It is the evidence run's own campaign (§G.1), exported through the API and imported into a database file that did not exist, in a directory that did not exist, opened by a second server process. Migrations ran there from nothing, so this is the fresh-install path as well as the import path.

Bundle 2,743,080 bytes, 13.1% of the 20 MB import limit, at 207 actions
Actions in the bundle / on the active line after import 207 / 119
Retained beyond the active line 88
Entities, facts 6, 2
Save Points 2; On the ridge restores to 69 actions and Before the ridge to 13, the positions they name
Knowledge sources 3: canon, reference and inspiration, with content, not just filenames
Narration-length choice came across
Campaign canon came across
Redo available exactly when the file said so
Secrets in the bundle none
The moved campaign accepts a new correction with nothing refused, and exports again at the same story length, 207 actions

The first revision's recovery run moved a 41-turn campaign at the 4,096 window (654,803 bytes, 83 actions), and that campaign's files were lost. The run above replaces it.

On the bundle ceiling. M9 estimated about 279 turns against the 20 MB limit, using a fixture built to be heavy. This real campaign plays at the recommended 16,384 window, with every stored prompt near that size. It averages about 13 kB per action, which puts the limit near 1,600 actions. M9's number remains the conservative one to quote.

Migration

tests/test_m11_migration.py — 9 tests, all passing, and the first of them is the permanent form of the defect M10 found by accident:

Check Result
A fresh install and an upgraded M10 database produce the same schema identical — every table, every column with type and nullability, every index with its columns and uniqueness, every foreign key, the primary keys, and the version stamp
No table carries two indexes over the same columns none does
Every table the models declare exists including visual_profiles and the new narration_length column
An M10-era database (stamped 92, no visual_profiles) opened by this build gains the table; campaign, action and Save Point intact; foreign_key_check empty; quick_check ok
The new column arrives as "" which means "this campaign never chose", so no existing prompt changes under the upgrade
Opening the database repeatedly index set and version identical after each
A migrated database still plays and still travels played through a real server process, corrected, exported and re-imported
A backup of the migrated database verifies, and carries the new column
A fresh install from nothing creates the database, stamps the current version, and has every expected table

The comparison is written generically rather than about visual_profiles, so a future migration that diverges the two paths fails here whatever it is about.

M11's own migrations are numbers 93 and 94: adventures.narration_length in the first revision and settings.context_window_override in ef25b0a, one column each, no backfill. The knowledge-migration suite's version assertions were corrected from a literal == 92 to == LATEST_VERSION >= M7_VERSION, because they had been asserting that M7's version was the newest — true when written, and a statement about M11 rather than about M7 once M11 added a migration (§O).


L. Browser

Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1), against the built SPA served by FastAPI — the production path from DEVELOPMENT.md, not a Vite dev server. Narrator: qwen2.5:3b-instruct-16k on the trusted-LAN Ollama. 38 checks, 0 failed, 0 skipped, 107 seconds.

The harness is tools/m11_browser.py and its WebDriver client is tools/m11_webdriver.py, both in the repository — M8's and M9's browser harness lived outside it, which made their browser evidence unrepeatable by anyone else. No Selenium: WebDriver is an HTTP protocol and urllib speaks HTTP, so browser evidence adds nothing to the dependency surface.

Area Checks
B01 narration two real turns accepted through the real engine
A/UX (finding A) the tab carries no inherited name; names the product; names the open campaign
B (finding B, §8A) the position is shown; it visibly changes after Undo; it says when later story is available
D01/D04 Undo offered; Redo becomes available after Undo; Redo returns to where the reader was
H06 an onerror attribute never executes; a <script> in narration never executes; markup in the source is not markup in the page
H07 a javascript: URL never becomes an href
G09 a remote Markdown image is not loaded
H04 shell text in narration is text
G01 a local file imports through the real file input; Import is enabled only once a file is chosen
§38 narrator-only text is absent from the DOM, not merely hidden
F05 the context inspector shows the assembled prompt and its budget
H10 an unknown API path is a 404 with a non-HTML body; a page path is the SPA
CSP a policy is served and names no remote origin
A11y every visible control has an accessible name; focus is visible; no positive tabindex; nothing revealed only on hover; the story input takes keyboard focus
A11y modal the dialog takes focus, has an accessible name, contains something focusable, and Escape closes it
A11y contrast measured on rendered colours (below)

Accessibility, measured rather than eyeballed

M8 recorded contrast and visible focus as checked by eye and handed the measurement to M11. Both halves were done.

Rendered contrast, computed in the page from the actual colours after inheritance and layering, with the WCAG 2.1 formula:

story prose   14.57:1 at 18.25px      (needs 4.5:1)
control        5.48:1 at 12.48px      (needs 4.5:1)
input         13.57:1 at 16.81px      (needs 4.5:1)
position       5.88:1 at 12.48px      (needs 4.5:1)

The palette, by calculation (tools/contrast_audit.py): every text pair the design uses clears WCAG AA 1.4.3, the lowest being an error message at 5.02:1.

Two boundary pairs are below 1.4.11's 3:1 — a control's resting edge at 1.33:1 and its hover edge at 1.75:1 — and they are reported, not fixed. 1.4.11 applies to the visual information required to identify a component, and in this design that is the control's text label, which is measured at 5.48:1 and passes. Restyling the palette would be a design change made inside a release-validation milestone to satisfy a threshold the reader is not affected by. It is recorded here so a reviewer can disagree.

What was measured versus inspected. Measured: contrast (both ways), accessible names, focus visibility, tab order, hover-only revelation, modal focus and dismissal, keyboard reachability of the story input. Inspected by reading rather than measured: reading order beyond tabindex, and screen-reader announcement quality. No claim is made about tablet layout beyond what the specification promises.

The limitation that remains

A file can now be driven into the browser (the snap sandbox accepts a path under $HOME, which is what M9's residual risk 6 had recorded as impossible). Driving one out — a blob: download from the export control — still does not complete under this headless snap Firefox. Export is proved end-to-end without a browser (§K) and the browser's own export control is exercised only as far as the click.


M. Dependencies and packaging

The runtime dependency surface

Nothing was added, removed or upgraded by M11. requirements.txt, requirements.lock and package.json are byte-identical to M10's. The browser harness deliberately speaks WebDriver over urllib rather than adding Selenium, because a package added to press buttons would still be a package in the audit surface.

npm audit 0 vulnerabilities, both with dev dependencies and with --omit=dev
Production npm dependencies three: react, react-dom, react-router-dom. Everything else is devDependencies
pip check no broken requirements
Lock consistency 35 pinned entries, none missing from the environment, zero version drift
What ships the image installs from requirements.txt only: 31 packages
Cloud/auth/analytics packages none. No openai, no anthropic, no telemetry SDK, no auth library
Runtime CDN or font references none — the fonts are vendored and test_offline_assets.py fails if a remote origin returns
New network client from M10 or M11 none. httpx was already the only HTTP client; contextwindow.py uses it with the shared TLS context and the shared endpoint policy
Media implementation dependency none — no ComfyUI, diffusers, Whisper, Kokoro, or model download

One observation, not a defect. The developer venv carries three packages the lock does not pin and the image does not install: quickjs (upstream's scripting engine, removed in M2), psycopg/psycopg-binary (the hosted deployment, removed in M2) and cryptography. They are residue in a long-lived local environment, not shipped: the image has none of them, and test_local_only_surface.py already fails if any module imports quickjs. A fresh pip install -r requirements.txt produces the 31-package set.

Unresolved advisories: none reported by either audit.

Packaging

Every production-shaped path the repository claims:

Path Result
uvicorn app.main:app --host 127.0.0.1 --port 8000 the path the 100-turn run and the browser run both used; started 5 times across M01's restarts
SPA served by FastAPI the browser regression ran entirely against the built frontend/dist, served by the backend
docker build --no-cache clean; log in the offline evidence directory
Container startup serves with --network none
Persistence across container restart campaigns survive on the volume
Loopback publication docker-compose.yml publishes 127.0.0.1:8000:8000; test_local_only_surface.py asserts it
Vendored assets fonts and the tokenizer table are in the image; the offline run fetched every asset the shell references from the container itself
No dependency on development source the image contains backend/app and frontend/dist only — no tests, no tools, no node_modules

No release was created and no tag exists.


N. Performance and storage

Measurements, not requirements. The planning package sets no performance target and none is invented here.

Long-run storage

The evidence run, §G.1:

Database after 101 accepted turns: 207 actions across the active line and retained history, 33 memories, 12 summaries, 3 imported sources 2,367,488 bytes
Export of the same campaign 2,743,080 bytes, 13.1% of the 20 MB import limit
Per action, roughly ~11 kB in the database and ~13 kB in the export, dominated by the per-position state snapshot and a stored prompt of up to ~16k tokens

No live backup was taken during this run. The first revision's backup figure (160 pages, integrity: ok) came from the lost campaign.

Prompt size across the run

The application's count:

turn   1    1,441 tokens    3 of   3 actions in the prompt
turn  10    6,422 tokens   21 of  21
turn  20   11,447 tokens   41 of  41
turn  30   15,299 tokens   55 of  61
turn  40   15,300 tokens   57 of  81
turn  50   15,474 tokens   57 of  99
turn  60   14,718 tokens   55 of 115
turn  70   15,549 tokens   58 of 136
turn  80   15,203 tokens   59 of  77   (after the Save Point restore)
turn  90   15,504 tokens   61 of  97
turn 100   15,745 tokens   63 of 117

The prompt rises until the history window is full and then stops. That is F03's claim measured rather than argued. The window's contents keep moving, always to the newest turns, while its size stays put.

Real-token headroom

The application counts tokens with cl100k_base, and the narrator counts with Qwen's own tokenizer. Each run's largest stored prompts were re-sent to the same model, identified by digest, and its prompt_eval_count was read:

Run Largest prompt, narrator's count Plus reply reserve 564 Headroom in 16,384
CPU host, memory off (ef25b0a) 15,728 16,292 92
GPU host, 100 turns (f8d4010) 15,797 16,361 23
GPU host, 100 turns (0c7316f) 15,786 16,350 34
GPU host, evidence run (96c1bf5) 15,778 16,342 42

No prompt in any run exceeded the window. On every prompt re-counted, the application's count was 16 tokens below the narrator's. The evidence run's re-count was taken after the GPU host's reboot, with the model digest verified unchanged.

The margin matters because of how the server fails. Ollama 0.34 cuts a prompt longer than the window down to 8,194 tokens and returns no error. That was observed with synthetic prompts, not in a campaign. §P records it as a risk.

Inference cost, by host

Host Window Prompt at steady state Seconds per turn 100 turns
CPU reference host 4,096 ~3.4k tokens 88-203, median 178 ~4 h projected; that run was lost at 97
CPU reference host 16,384 13-14k tokens 229-291 in the first attempt 45,656 s for 101 turns, with memory off
GPU inference host 16,384 14.6-15.7k tokens 4.1-22.8, median 9.6 1,238 s for 101 turns, with memory on

Prompt processing dominates on the CPU host. On the GPU host a full 16k prompt is processed in about ten seconds (§E.1). Both are these machines' characteristics.

Query behaviour

No new query growth was introduced. M11 adds one HTTP round trip per session per (endpoint, model) pair, the window probe, cached for ten minutes on success and one minute on failure. It adds one derived computation per state read, duplicate_names, a single pass over the entities already in memory. The connection test was changed to use that cache after the browser run showed the model-status badge calling it on every page load.

The §O.7 correction moves one statement rather than adding one: the memory use-counter UPDATE now runs inside the turn's existing commit instead of before the model call.

Nothing pathological was found

No unbounded growth, no per-row query, no repeated snapshot write, no duplicated knowledge content. The M10 measurement tool (tools/m10_media_cost.py) remains valid: a scene packet is four SQL statements at any campaign length.


O. Findings

Product defects first, then defects in the tests and harnesses, which are kept separate because conflating them is how a milestone reports confidence it has not earned.

Product defects — found by M11, fixed in M11

O.1 — The application silently budgeted more input than the server would read Severity: high. Requirement: F03, F04, M03, and the honesty of every long-run claim. Blocker: yes — it was M11's stated blocker. Status: fixed.

Reproduction: configure a model with no num_ctx on a server with no VRAM (/api/ps reports 4,096); leave context_token_budget at its 16,384 default; play a long campaign. Every request returns 200 and the server drops the oldest tokens — the narrator's rules and the campaign canon.

Root cause: the window is a property of the model load, not of the request, and the application had no way to learn it. §E in full.

Correction: app/contextwindow.py discovers it and the builder caps to it, or records the turn as unverified.

Regression evidence: tests/test_m11_context_window.py (20 tests) including the truncation sentinel and its negative control — the same campaign built without the cap, measured at more than twice the window. Plus tests/test_m11_real_window.py against a real Ollama.

O.2 — A partly refused manual state correction reported success Severity: medium. Requirement: C04, and SECURITY-THREAT-MODEL.md §69 auditability. Blocker: no. Status: fixed.

Reproduction: POST /state/corrections with two changes, one naming a location entity that does not exist. Before M11: HTTP 201, the good change applied, the bad one silently dropped, nothing in the response to say so.

Root cause: the handler raised 400 only when nothing was accepted. Partial acceptance is deliberate and correct — validate.py argues that discarding three good changes because of one typo is worse — but the same file states the rule this broke: "what is never allowed is a rejected event mutating anything, or a rejection being silent". The refusal was recorded on the proposal row for the audit trail; the person who wrote it was simply never told.

How it was found: the identity diagnostic's own fixture set a scene naming a location entity it had not created. The event was refused, the 201 said nothing, and the entire first diagnostic run happened on a campaign with no scene and no list of who was in the room — a degraded context that could easily have been read as a model failure. That run's evidence was discarded.

Correction: the correction response carries refused (event, reason, detail), and the State panel shows it. Behaviour is otherwise unchanged.

Regression evidence: test_a_partly_refused_correction_reports_what_did_not_apply, with controls for the fully-applied and wholly-refused cases and a check that the audit record still records partially_accepted.

O.3 — The narration-length setting moved no number (post-M8 finding C) Severity: medium. Requirement: the setting's own promise; B-series narration quality. Blocker: no. Status: fixed.

Reproduction: create three campaigns choosing brief, medium and long; compare the length_hint section of the stored prompts. Before M11 they were identical — at the default cap, "must not exceed 506 words, and it should not stop short of about 177" in all three.

Root cause: the choice became one English sentence in ai_instructions and the numeric hint was derived from the global max_output_tokens.

Correction: the choice is data on the campaign; length_hint maps it to a word band bounded by the reply cap. The generation budget is deliberately untouched.

Regression evidence: tests/test_m11_findings.py, including test_the_three_lengths_no_longer_say_the_same_thing (fails against the old behaviour) and test_the_stored_prompt_carries_the_campaigns_own_range end to end.

Product gaps closed without a defect

O.4 — The tab carried the inherited name (post-M8 finding A). Not a false claim by any document, so a gap rather than a defect. Fixed.

O.5 — No orientation after history movement (post-M8 finding B). The requirement it violated did not exist until §8A was written after the playtest. Fixed to that requirement.

O.6 — Two entities could share a display name silently (post-M8 finding D's structural half). Deliberately not made an error: two people called Alice is ordinary fiction. Made visible instead — duplicate_names in the state API and the State panel.

Product defects — found by the long runs after the first revision, fixed

O.7 — A turn held SQLite's write lock through the model call, and post-turn memory and summary work was lost without a trace Severity: high. Requirement: M01's summary/memory activation, F02, F08, M04. Blocker: yes, for M01. Status: fixed in f8d4010.

Reproduction: turn the memory bank and auto-summarise on, configure an embedding model, and play consecutive turns once a memory exists. On a fast inference host: database is locked in the server log, memories stop accumulating, no summary is written, and derived status reads idle with no failures. The first 26-turn trial (§G.6, run 3) logged 180 such errors, wrote 2 memories and no summary, and reported complete.

Root cause: retrieve_memories(update_stats=True) ran an uncommitted UPDATE of the used memories' counters before the model call. The turn commits once, after the reply has streamed, so the write transaction stayed open for the whole reply. SQLite has one writer. Every post-turn memory, summary and status write in that window waited out the driver's five-second timeout and failed. Recording the failure needs a write as well. The post-turn task's outer handler recorded without rolling back first, so it raised PendingRollbackError and the failure reached only the log. F08 requires a memory failure to be visible, and this one was not. No earlier long run hit it, because none had the memory bank on.

Correction: retrieval only reads. memorybank.record_use writes the counters in the turn's single commit, so a turn that never lands counts nothing. The outer handler rolls back before it records.

Regression evidence: test_no_write_lock_is_held_while_the_narrator_is_talking probes for the lock from a second connection during the model call. test_a_failure_that_breaks_the_session_is_still_recorded covers the recorder. Both fail on fec46f6, with database is locked and idle respectively. Also test_a_failed_turn_counts_no_memory_as_used. The next trial had 0 lock errors, 7 memories and 2 summaries.

O.8 — The narrator's protocol was stored as story Severity: high. Requirement: story authority (M5 review Finding 4), C04, and the validity of M04. Blocker: yes, for M04's evidence. Status: fixed in 0c7316f and 96c1bf5.

Reproduction: play a long campaign against a small local model, then search the stored AI turns for the state section's headings or "events". In the f8d4010 run 42 of 104 turns carried protocol, the first at depth 2, in four shapes:

  • a copy of the narrative-state section: Scene:, Who and what exists:, Held:, Established:, Still open:
  • that copy above a correct ```state block, which was removed while the copy stayed
  • the copy, a bare State heading, and a > {"events": ...} proposal quoted like a player turn, sometimes with story after it
  • the same block cut off by the output-token limit, on 10 turns

In the next run the narrator wrote sections of its own instead, such as ## Established: over indented facts, on 5 turns.

Root cause: the extractor removed fenced blocks, a bare object at the very end, and a parroted bracketed reminder. It did not remove a pasted state section, an unfenced proposal elsewhere in the reply, or an unfinished one. Stored text is replayed verbatim as history (_history_text). Each leak therefore put a second, older account of the state into the next prompt, which is what Finding 4 removed from replay, and it gave the model another example to copy. A second, older bug was in the fence pattern itself. _STATE_FENCE_RE read "a ```state block" inside a parroted reminder as a fence opening and cut out the middle of the reminder.

Correction: the extractor recognises the renderer's own section headings, which are now named constants in render.py, with any markdown wrapped around them. It removes a block carrying two headings, or one heading with an indented entry. It removes an unfenced proposal that starts a line, quoted or not, taking outermost objects first, and uses it as the turn's proposal when there is no fence. It removes an unfinished proposal at the end and whatever is left behind at the end: a State heading, a bare >, a parroted reminder or continue hint, closed or not, and a ```json fence cut off before it names its events. The fence label must now end its line or run straight into the payload. A reply whose only removal is a pasted section records no raw block, so the turn is not marked unparseable for a block it never started.

Regression evidence: 25 new cases, from 18 test functions, in test_narrative_state.py, cut down from the runs' real output, including negative controls. A lone Held: with prose after it, a Scene: line of story, quoted JSON that is not a proposal, and a State line followed by story all stay. A test renders every section the renderer writes and pastes the lot. Beyond the suite, every AI turn in five real runs was replayed through the new extractor, 443 in all, and no turn the old extractor had left clean changed. The evidence run stored 0 of 104 turns with protocol.

Not removed, deliberately: a restatement of a fact inside a line of story (§G.4), and headings a model invents that are not the renderer's (Identifiers established:, on 10 turns of run 3). Removing either means judging prose, and §P records both.

Harness and test defects — found and fixed, no product change

Recorded separately, and at this length, because M8's review found five harness defects against seven product defects and two of the five were masking product defects. A harness that has only ever agreed with itself is not evidence.

What it did Why it mattered
The identity fixture set a scene naming an uncreated entity the scene never existed, and the diagnostic ran on a degraded campaign it would have been read as a model failure. It also surfaced product defect O.2
The long-run harness read the turn endpoint as JSON crashed on the SSE body a harness that read the status code instead would have called every failed turn a success — sse.py says a failed turn is a 200 with an error inside the stream
The same harness called retry as JSON crashed at the fourth scheduled step found by a 14-turn shakeout run rather than 50 turns into the release campaign, which is what the shakeout was for
The offline check compared total action counts after a failed turn reported corruption where the product was behaving as designed A05 deliberately keeps the player's submitted text and the head sits on it. The check now asserts the real contract: no narration accepted, state unchanged, earlier story reachable
The browser import scenario never pressed Import choosing a file only stages it looked exactly like a broken import
The browser modal check used the Save Point control that opens a panel, not a dialog, so the check skipped itself a skip that reports nothing is worse than a failure
A set_scene bounds test asked for 43 present labels (M10, recorded again here) the state model correctly refused the event the test measured the wrong scene
Two substring checks matched inside words "ahead" contains "head"; "immediately" contains "media" both now match whole words or parse imports
A contrast check treated a control boundary as body text would have failed the run on a WCAG clause that does not apply now distinguishes 1.4.3 from 1.4.11 and says which
An audit assertion read a deferred column after the session closed DetachedInstanceError read inside the session
The long-run harness never switched the memory bank or auto-summarise on (fixed in fec46f6) both are per-campaign and default to off, so a complete 100-turn run wrote no memory and no summary M01's summary/memory activation clause was reported by silence, and M04's memory path was never asked. The harness now switches both on, reads them back, and refuses a run with no embedding model
It read three prompt sections under names the builder does not use (fixed in f8d4010) memories (really used_memories), story_history (really history and recent_history), and a knowledge prefix that matched the fixed instruction section instead of the imported passages memory tokens read 0 whatever the prompt held, and the in-history and in-memories recall checks could never be true. The labels are now constants pinned by a test against a prompt the real builder assembled
It had no way to see failed post-turn work (fixed in f8d4010) it reported complete over 180 database is locked errors it would have certified the run that §O.7 destroyed. It now checks derived status and new server.log lines after every turn and stops at the first failure; a run with no memories or no summaries ends failed
Its protocol-leak count used plain substrings (fixed in 96c1bf5) ## Established: did not match \nEstablished:\n, so the count read 1 where 5 turns leaked the count now matches any state heading with an indented entry, in any markdown
Its M04 precondition was the clue's text being out of recent history (fixed in 96c1bf5) the narrator reuses that text in its own prose, so a run whose planting turn was 65 depths outside the window read precondition_not_met the precondition is now the planting turn's position, recorded at planting and carried across --resume, which is M04's own wording

Discarded evidence runs

Per §7, evidence taken before a product change was discarded rather than reported:

  1. The first identity diagnostic run — void, because its own fixture had been refused (O.2). Re-run after the fixture and the product were fixed.
  2. The first browser regression run — 31 passed, 1 harness failure, 1 skip. Discarded and re-run after the harness fixes and the connection-test caching change; the reported run is the second: 38 passed, 0 failed, 0 skipped.
  3. The first offline run — 20 passed, 1 failure that was the harness asserting the wrong contract. Discarded and re-run: 23 passed, 0 failed.
  4. The first 100-turn campaign run — abandoned at 3 turns when the connection-test caching change landed, so that the reported campaign runs entirely on the final tree.
  5. Two harness shakeout runs (6 and 14 turns) — never reported as evidence; their purpose was to find the two SSE defects above.
  6. The first 100-turn release campaign, which reached 97 of 100 turns and was lost with its evidence to a host crash, because it wrote under /tmp. The first revision's §G figures came from it and can no longer be checked. They are replaced, not repeated.
  7. The memory-off 100-turn run (ef25b0a). It is complete and internally consistent, but it never exercised summary/memory activation. It is kept as the CPU-host timing and headroom record (§N).
  8. The 100-turn runs on f8d4010 and 0c7316f. They are complete, and they are superseded because their M04 evidence was contaminated by protocol leaks (§O.8). The 0c7316f run's reclassified verdict is kept beside its original and is not reported as the result.

P. Residual risks

Genuine remaining risk and debt only. There is no M12; everything below is either accepted for v1, or a decision for the owner at acceptance.

  1. The release evidence carries two hosts' speeds. The CPU reference host took minutes a turn, and the GPU host took seconds. Nothing in the planning package sets a performance requirement, and none is invented here. Read §N's timings as these machines', not as a product characteristic.

  2. The narrator is a 3B model. Every realistic-model observation is that model's: state extraction quality, narration length adherence, identity handling, and what it restates. A stronger local model would behave differently, probably better. The application's guarantees are deliberately independent of which, since what is asserted is that the application stays correct whatever the model proposes.

  3. The 16k window is nearly full. By the narrator's own tokenizer the largest prompts plus the reply reserve left 23, 34 and 42 tokens of headroom in the three GPU runs (§N). The application's cl100k_base count ran 16 tokens low on every prompt measured. The inference server cuts an over-window prompt to 8,194 tokens with no error. No prompt overflowed. A model whose tokenizer diverges further from cl100k_base, or a larger reply reserve, could overflow without anyone noticing. A deployment that wants margin can lower context_token_budget below the window.

  4. The narrator restates prompt text in its prose. On 74 of 104 turns of the evidence run it wrote the state's fact line as a sentence, and it also echoed phrases such as "Scene set, continue your adventure." and lines opening "Memory:". These sit inside story, so the extractor leaves them. They cost narration quality, and they are the route by which M04's fact reached memory.

  5. Memory does not keep a planted fact on its own with this summariser. In no run did a memory summarising the planting era carry the fact. Recall rests on authoritative state, which the repository owner accepted for M04 on 2026-09-13. A reviewer who reads F02 or M04 as requiring memory retention in its own right should read them as unproven.

  6. Model-invented headings are not removed. A model that writes its own section, such as Identifiers established: or Set of events made true:, under a heading that is not the renderer's, keeps it in the story. Seen on 10 turns of one 26-turn trial.

  7. The GPU inference host dropped its GPU after the evidence run (§E.1). The cause is not established, and power transients at the uncapped 280 W limit are the leading candidate. The evidence is unaffected. Future long runs must log power, link state and kernel messages (DEVELOPMENT.md).

  8. Post-M8 finding D's root cause is unestablished and will stay that way. The campaign that produced it was destroyed. M11 delivers a diagnostic that can classify the next occurrence, and the detection the finding asked for. The diagnostic's run results are not in this report. The section the first revision pointed to was left empty, and the evidence is in the implementer's m11-evidence/identity-recheck directory.

  9. The browser, offline and identity runs predate the last three product commits. They were re-run on the ef25b0a tree: that commit's message records browser 38/0/0, offline 23/0 and the identity diagnostic clean, and the evidence is dated 2026-09-10. f8d4010, 0c7316f and 96c1bf5 change the turn commit, memory retrieval, narration extraction and the state renderer's headings, all of them backend. The backend and frontend suites cover those changes; the black-box runs were not repeated.

  10. Two control-boundary colour pairs are below WCAG 1.4.11 (1.33:1 resting, 1.75:1 hover). Reported rather than fixed, because the control is identified by its label, measured at 5.48:1, and restyling the palette inside a release-validation milestone would be the wrong kind of change. An owner who disagrees has the measurement.

  11. A file cannot be driven out of this headless snap Firefox. Export is proved end to end without a browser; the browser's own export control is exercised only as far as the click. Import is fully proved.

  12. The bundle ceiling is unchanged. M9 measured about 279 turns against the 20 MB import limit, and §K puts the real 100-turn campaign at 13.1% of it. Beyond the ceiling a campaign can still be exported and would be refused on import, which is the asymmetry worth knowing.

  13. The media seam has no real adapter. This is M10's own residual risk, unchanged: the contracts are shaped by the contract document rather than by an adapter that had to work.

  14. quick_check rather than integrity_check on a backup, and no scheduled backup. Both are M9's, both unchanged, both outside the acceptance contract.

  15. A campaign with no narration-length choice keeps the pre-M11 hint. That is deliberate, because an empty value means the reader never chose. It means an existing campaign does not benefit from finding C's fix until someone sets the control.


Q. Planning and document changes

Document Change Kind
planning/TECHNICAL-DESIGN.md New §15.2 — the inference window as a ceiling: discovery, enforcement, and why there is no hard-coded 4,096. implementation fact
planning/DATA-MODEL.md New §28B — M11's one column, and why the window ceiling, duplicate_names and refused are deliberately not stored. implementation fact
planning/BROWSER-UX-SPEC.md §8A gains "As implemented (M11)" — the position indicator and the three properties that make it answer the requirement. §8A's own text is unchanged. implementation fact
planning/SECURITY-THREAT-MODEL.md New §42B — the probe under §73, the silent partial correction as a §69 gap now closed, H09 not applicable with the condition enforced, and the measured contrast. implementation fact + boundary note
planning/V1-ACCEPTANCE-TESTS.md Results for every REQUIRED test (§F); §P1's duplicate-name question settled (report, do not refuse); §P3 gains an M11 disposition recording that the identity diagnostic exists and remains a test-design task rather than an acceptance test. acceptance evidence + disposition
planning/BUILD-MILESTONES.md M11 status block: the blocker closed, the four post-M8 findings disposed of, the two defects the validation found. milestone status
planning/VERSION.md v3.7 entry. package version
planning/README.md Status, milestone map, reading order; M9's and M10's reports rotated to the archive and the reason the exception ended. index
README.md contextwindow.py in the architecture map; the context-window behaviour as a feature. developer docs
DEVELOPMENT.md The context-window section rewritten around what the application now does; a new section on the six release harnesses. developer docs
planning/reports/M11-IMPLEMENTATION-REPORT.md New. milestone report
planning/archive/milestone-reports/ M9's and M10's reports moved here. rotation
planning/reports/M11-IMPLEMENTATION-REPORT.md Revised 2026-09-14: M01-M04 on the complete evidence run; §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R rewritten; the §G.0 addendum removed. milestone report
DEVELOPMENT.md The pointer to the removed §G.0 replaced; a new section on logging a GPU inference host during a long run. developer docs

Revised after this report, in planning package v3.9 (2026-09-14): planning/V1-ACCEPTANCE-TESTS.md gained result blocks for M01-M04 and corrected its M11 disposition. planning/BUILD-MILESTONES.md, planning/VERSION.md and planning/README.md record the long-run evidence. DATA-MODEL.md §28B now records migration 94, and TECHNICAL-DESIGN.md (new §15.4), CONTEXT-AND-MEMORY.md §51 and ADR 013 record the two defect fixes as implemented.

Implementation facts added: §15.2, §28B, §8A's implementation note, §42B, and the M11 status block. Each records what the code does; none changes what is required.

Requirement changes: zero. No acceptance test was retired, relaxed, reclassified or rewritten to match behaviour. H09 is reported NOT APPLICABLE on the condition its own text states, and that condition is now enforced by a test rather than asserted. §P1's question was settled — the answer being that a shared display name is reported rather than refused — which resolves an open design question rather than weakening a requirement; the permissive behaviour is pinned by a test so a later milestone changes it deliberately.


R. Final release-readiness assessment

1. Does every REQUIRED FOR V1 acceptance test pass?

Yes. All 85 pass, with H09 recorded NOT APPLICABLE on the condition its own text states. M01 to M04 pass on the evidence run in §G, and every other REQUIRED test passes on the evidence in §F.

2. Does M01 pass with 100+ accepted turns?

Yes. 101 accepted turns on commit 96c1bf5, with three genuine process restarts, all thirteen scheduled history operations, an induced failed call that corrupted nothing, zero failed post-turn passes, 33 memories and 12 summaries, and recovery onto a clean data directory 16 of 16. §G.1-§G.5.

3. Does actual model context capacity match the application's release assumptions?

Yes, it is enforced rather than assumed, and the margin is thin. The window was verified on 101 of 101 turns at 16,384, and the application caps its budget to whatever the server reports (§E). By the narrator's own tokenizer the largest prompt plus reserve is 16,342 of 16,384. It fits, with 23-42 tokens of headroom across the GPU runs (§N, §P).

4. Does offline operation pass from a fresh state with no Internet route?

Yes. 23 of 23 checks in a container with --network none and a fresh volume: no route, no DNS, first page load, every asset local, campaign creation, state extraction, knowledge import and retrieval, prompt assembly, export, import, the media module inert, and campaigns surviving a container restart. A turn with no model reachable is reported and corrupts nothing. §J. Re-run on the ef25b0a tree (§P).

5. Does branch/memory/summary/scene isolation pass?

Yes. All four in one long campaign, each with a positive control, including a summary regenerated after the divergence and M10's derived Scene Packet. §H. The evidence run also restored a Save Point across a live memory bank. §G.2.

6. Does trusted-LAN HTTPS inference still pass?

Yes, on the CPU reference host's evidence. The browser run, the identity diagnostic and the CPU-host campaigns ran against Ollama on a separate physical machine over HTTPS, with a private CA in the application machine's OS trust store, verification on and no bypass. The storyteller stayed loopback-bound. The final long runs used plain HTTP to a LAN GPU host. H12 permits that, and it is not A06 evidence. §E.1, §F, §J.

7. Does export/import/recovery pass?

Yes. 16 of 16 checks moving the evidence run's campaign (207 actions, 88 of them retained history, 2.74 MB) into a data directory that never existed, plus M9's own recovery suite and the migration parity suite. §K.

8. Do both fantasy and science-fiction fixtures pass?

Yes. The fantasy Continuity Test is the long-run campaign. The Persephone fixture passes 10 checks, including five entity types in one document, hard canon, possession, reference retrieval, a starship in M10's Scene Packet, and a whole-vocabulary check that no event type names a genre noun. No schema or code change was required for either. §I.

9. Do fresh-install and upgrade schemas agree?

Yes, compared field by field: tables, columns with type and nullability, indexes with columns and uniqueness, foreign keys, primary keys and the version stamp. §K.

10. Does the browser pass the release workflow?

Yes, on the ef25b0a tree. 38 checks, 0 failed, 0 skipped, in Firefox 154.0.1 against the built SPA served by FastAPI. The commits after it change no frontend file, and the browser run was not repeated. §L, §P.

11. Are there any unresolved blockers to independent v1 acceptance?

No known blocker. No product defect is known and unfixed, and no requirement was weakened. A reviewer should weigh five things before signing:

  • M04's recovery ran through authoritative state, not memory retention (§G.4)
  • the thin real-token headroom (§N)
  • the narrator restating prompt text (§P)
  • the identity diagnostic's results being absent from this report (§P)
  • the black-box runs predating the later commits (§P)

12. Is the tree safe to commit as the M11 release candidate?

It is committed. Seven signed commits follow the M10 base, the newest 96c1bf5, and there is no release tag. On that tree the backend suite passed 1,421 with 17 skipped and 0 failed (2026-09-13). The frontend suite passed 161 of 161, and lint exited 0 with warnings only (2026-09-14). The production build and the Docker image were built for the first revision and not rebuilt since. The remaining commit, this report's revision, is the owner's to sign.


Written by the implementer. Not an acceptance record.