From d1988065e512d575c43faf5582e8686dd533be23 Mon Sep 17 00:00:00 2001 From: JesseMarkowitz Date: Mon, 14 Sep 2026 03:15:13 -0400 Subject: [PATCH] M11 report: M01 to M04 on a complete run, and what it took to get one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The first revision left M01 PARTIAL at 41 accepted turns, and that run was then lost to a host crash with its evidence. This revision reports the 100-turn evidence run on 96c1bf5. It had 101 accepted turns, three genuine restarts, every scheduled history operation, zero failed post-turn passes, and recovery onto a clean data directory 16 of 16. All 85 REQUIRED FOR V1 tests now pass, with H09 NOT APPLICABLE on its own condition. Rewritten: §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R. The §G.0 addendum is removed, and its history is §G.6. - §G: the evidence run's timeline, context growth and recall. M04's precondition is positional: the planting turn at depth 1, the window floor at 54. The path is stated plainly: authoritative state, then the narrator restating the fact line in its prose, then memory summarising those restatements. The owner accepted state-based recovery on 2026-09-13. - §G.6: all seven long runs, and why six of them are not the evidence. - §O.7 and §O.8: the write-lock defect and the protocol-leak defect. Five harness defects are added to the harness table. - §N: storage for the evidence run, and real-token headroom by the narrator's own tokenizer: 92, 23, 34 and 42 tokens across four runs, with the inference server's silent cut to 8,194 tokens stated. - §E.1: the application machine, the CPU reference host and the GPU host, identical model digests, and the GPU dropping off the PCIe bus (Xid 79) about 30 s after the evidence run's last write. That long runs must log power, link state and kernel messages is recorded as a requirement. - §F, §J, §R: A06, H01 and R6 no longer claim that every turn went over HTTPS. The GPU runs used plain HTTP to a LAN host and were not network-monitored. - §B, §P: the black-box runs (browser, offline, identity) were re-run on the ef25b0a tree and not after the three later backend commits. - §C, §K: the later commits, including migration 94 from ef25b0a. - §D.1, §P: the identity diagnostic's results were never written into this report; the section the first revision pointed to was empty. DEVELOPMENT.md: the pointer to the removed §G.0 is replaced, and a new section, "Logging the inference host during a long run", gives the nvidia-smi and journalctl commands to run on a GPU host for every long run. If the GPU drops again, the logs show whether it was power. Not revised here: V1-ACCEPTANCE-TESTS.md, BUILD-MILESTONES.md, VERSION.md and planning/README.md. On 96c1bf5: backend 1,421 passed, 17 skipped, 0 failed; frontend 161 passed; lint exit 0 with warnings only. No code changes in this commit. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND --- DEVELOPMENT.md | 41 +- planning/reports/M11-IMPLEMENTATION-REPORT.md | 957 +++++++++++------- 2 files changed, 643 insertions(+), 355 deletions(-) diff --git a/DEVELOPMENT.md b/DEVELOPMENT.md index 2f89b7f..bd5dbd1 100644 --- a/DEVELOPMENT.md +++ b/DEVELOPMENT.md @@ -385,8 +385,7 @@ anything. None is part of the application and none is imported by it. every harness precisely so the location is a decision rather than a default, and the examples below use `$HOME/m11-evidence`. A reboot clears `/tmp`, and a long-run campaign is hours of evidence that cannot be reproduced by re-reading a -file — one run was lost exactly that way (see the addendum at §G.0 of the M11 -report). Snap Firefox independently refuses a WebDriver file path under `/tmp` +file. One run was lost exactly that way; §G.6 of the M11 report records it. Snap Firefox independently refuses a WebDriver file path under `/tmp` and needs one under `$HOME`, so `$HOME` is the only location the browser harness works from in any case. @@ -454,6 +453,44 @@ deployment a turn cost 229-291 seconds at the recommended window; a slower host can exceed the 600 seconds this harness used to hard-code, and an overrun turn is a lost turn. +### Logging the inference host during a long run + +A long run is the heaviest sustained load an inference host sees. In M11 a GPU +host dropped its GPU off the PCIe bus (`NVRM: Xid 79`) half a minute after a +100-turn run finished. Nothing on disk could say whether power, heat or the link +caused it (M11 report, §E.1). **For every long run against a GPU host, start this +logging on that host first and stop it only when the run has finished.** + +Run each command in its own terminal on the inference host. `tee` writes each +line as it arrives, so what happened in the seconds before a crash or a forced +reboot survives on disk. + +```bash +# Power, temperature, utilisation and PCIe link state, once a second +nvidia-smi --query-gpu=timestamp,pcie.link.gen.current,pcie.link.width.current,power.draw,temperature.gpu,utilization.gpu \ + --format=csv -l 1 | tee "$HOME/gpu-link-$(date +%F-%H%M).csv" + +# The driver's own sampling: power, utilisation, clocks, memory, ECC and throttling +nvidia-smi dmon -s pucvmet -d 5 | tee "$HOME/gpu-dmon-$(date +%F-%H%M).log" + +# Kernel and Ollama messages, live +journalctl -f -k -u ollama | tee "$HOME/ollama-kernel-watch-$(date +%F-%H%M).log" +``` + +If the GPU faults, find the moment and then read what the card was doing just +before it: + +```bash +grep -iE 'xid|fallen off|nvrm' "$HOME"/ollama-kernel-watch-*.log +awk -F', ' 'NR>1 && $4+0 > max {max=$4+0; at=$1} END {print "peak W", max, "at", at}' "$HOME"/gpu-link-*.csv +``` + +A fault that follows sustained draw at the card's power limit points to power +delivery. A fault with the link below its usual generation under load points to +the connection. A fault with neither is still worth recording, because it rules +both out. These commands were verified against NVIDIA driver 580 and Ollama +0.34. + `tools/m11_webdriver.py` is the W3C WebDriver client the browser harness uses. It exists so browser evidence needs no Selenium in the dependency surface, and it documents the one environment quirk that matters here: a snap Firefox will diff --git a/planning/reports/M11-IMPLEMENTATION-REPORT.md b/planning/reports/M11-IMPLEMENTATION-REPORT.md index b9d1680..afe2d4e 100644 --- a/planning/reports/M11-IMPLEMENTATION-REPORT.md +++ b/planning/reports/M11-IMPLEMENTATION-REPORT.md @@ -3,61 +3,66 @@ **Implementation and verification report, written for independent release review.** Branch `m11-release-validation`, from the signed M10 commit `1013c94`. -Implemented and verified 2026-09-07. This is the evidence package; it is not a -record of acceptance, and nothing in it says M11 is accepted. +Implemented and verified 2026-09-07. The long-run evidence was re-established, +and this revision written, on 2026-09-14. This is the evidence package; it is +not a record of acceptance, and nothing in it says M11 is accepted. --- ## A. Executive result -**PASS WITH CORRECTIVE WORK REQUIRED.** +**PASS, for independent review.** All 85 tests marked REQUIRED FOR V1 pass, with +H09 recorded NOT APPLICABLE on the condition its own text states. -The corrective work the *product* needed was done inside M11 and is included -here. One piece of *verification* work remains outstanding, and it is stated -plainly rather than rounded up. +This revision supersedes the first. That revision left M01 PARTIAL at 41 +accepted turns, and then the run it described was lost to a host crash along +with its evidence. **M01 to M04 now pass on a complete 100-turn campaign**, run +by the committed harness on commit `96c1bf5`: 101 accepted turns, three genuine +process restarts, every scheduled history operation performed, zero failed +post-turn passes, and recovery onto a clean data directory 16 of 16 (§G, §K). -**What passes.** 84 of the 85 tests marked REQUIRED FOR V1, with H09 recorded -NOT APPLICABLE on the condition its own text states. Offline operation, browser -regression, lineage isolation, both genre fixtures, recovery into a clean data -directory, schema parity, trusted-LAN HTTPS inference, and the full automated -suites — all green, all on one frozen tree, all with the evidence located in §F. +Getting there took six further runs. They found **two more product defects**, +both fixed and both invisible to every earlier piece of evidence because the +memory bank had never been switched on during a long run: -**What does not.** **M01, the 100-turn campaign, is PARTIAL: 41 accepted turns -at the time of writing, still running.** *(Addendum: that run later reached 97 of -100 turns and was then lost, with all of its evidence, to a host crash. It must -be re-run. See §G.0.)* It is correct as far as it has gone — -every history operation performed, the restart byte-identical, the prompt -bounded, the window verified on every turn, the campaign recoverable on another -machine — but 100 turns needs roughly four hours of wall clock on this CPU-only -inference host, and eight at the recommended context window. That is the -reference hardware's characteristic, not the application's. §G says exactly what -was and was not exercised, and the harness is committed so the run can be -finished and re-checked before acceptance. +1. **A turn locked its own memory bank out** (§O.7). Retrieval wrote a use + counter before the model call, and the turn committed only after the reply, + so SQLite's single write lock was held for the whole reply. Post-turn memory + and summary writes timed out behind it, and recording those failures timed out + too. A 26-turn run reported "complete" with two memories, no summary and 180 + `database is locked` errors, while derived status read `idle`. +2. **The narrator's protocol leaked into stored story** (§O.8). A small model + pasted the narrative-state section and unfenced proposals into its prose, on + 42 of 104 turns in one run, and stored text is replayed as history. The + extractor now removes every shape observed. Replaying 443 real turns from five + runs changed no turn the old extractor had left clean. -**Three product corrections the release run forced:** +The first revision's three corrections stand: the context-window mismatch is +resolved (§E), a partly refused manual correction now says so, and the +narration-length setting moves a number. Post-M8 findings A and B are closed. -1. **The context-window mismatch is resolved** (§E). The application no longer - budgets more narrator input than the server will accept; it asks, caps, and - says so when it cannot check. This was M11's stated release blocker, and the - 41-turn campaign is its longest test — every turn capped from 16,384 to the - server's real 4,096, with the campaign canon present in all 41 prompts. -2. **A manual state correction that was partly refused reported success.** It - now reports what was refused and why. Found because the identity diagnostic's - own fixture was refused that way and ran on a degraded campaign in silence. -3. **The narration-length setting moved no number.** It now does. +**What is not claimed.** -Two further changes close post-M8 findings A and B — the inherited tab title, -and the reader's position after Undo. - -**What is not claimed.** M01's turn count, above all. Timings are this host's, -not a product characteristic. Two control-boundary colour pairs sit below WCAG -1.4.11 and are reported rather than fixed, with the reasoning. Post-M8 finding -D's root cause remains **unestablished** — as it must, the campaign that produced -it having been destroyed — and what M11 delivers for it is a diagnostic that can -tell the candidate causes apart, plus the detection the finding asked for. +- **That memory keeps a planted fact on its own.** M04 passes because the fact + was recoverable with its planting turn 53 depths outside the history window. + The recovery path ran through authoritative state: the narrator restated the + state's fact line in its prose, and memory summarised those restatements + (§G.4). No memory of the planting era carried the fact in any run. The + repository owner accepted state-based recovery for M04 on 2026-09-13. +- **Headroom in the context window.** The largest prompts fill 16,342 of 16,384 + tokens by the narrator's own count, and the inference server cuts an + over-window prompt without an error (§N, §P). +- **Performance.** Timings are two particular hosts', not a product + characteristic (§E.1). +- **Post-M8 finding D's root cause**, which remains unestablished. The identity + diagnostic's run results are not in this report (§D.1). +- **That the browser, offline and identity runs cover the later commits.** They + were re-run on the `ef25b0a` tree. The three product commits since then change + the turn commit, memory retrieval and narration extraction (§B). **Zero requirement weakenings.** No acceptance test was retired, relaxed or -reclassified. +reclassified. The M04 precondition the harness measures was corrected to the +acceptance text's own wording, "without entire transcript in prompt" (§O). --- ## B. Repository and provenance @@ -67,7 +72,7 @@ reclassified. | **Base commit** | `1013c94eb1ad283e960114aef04c19c2806b5db7` — *"M10: the seam for media, and no media"* | | **Signature** | `git verify-commit 1013c94` → **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, "JesseMarkowitz", trust `[ultimate]`. `%G?` = `G`. | | **Branch** | `m11-release-validation`, created from that commit. The working tree was clean at the start (`git status --porcelain` empty). | -| **HEAD** | `1013c94` — **M11 creates no commit.** The tree is staged for the repository owner to sign. | +| **HEAD at this revision** | `96c1bf5`. Seven commits follow the base, all signed by the owner (`%G?` = `G`): `144406c` (M11), `fedb714`, `ef25b0a`, `fec46f6`, `f8d4010`, `0c7316f`, `96c1bf5`. This revision of the report is staged for the owner to sign. | | **Upstream ancestry** | `git merge-base --is-ancestor d72f7c1b HEAD` → true. The AI-DnD fork point is still an ancestor. | | **LICENSE** | Unchanged — md5 `07fde30437134836e2ee875e82a7cd31`, MIT, "Copyright (c) 2026 Parth Thakkar". `PROVENANCE.md` unchanged. | @@ -82,9 +87,14 @@ frozen tree the working tree at the point each black-box run started, with no source edit during any reported run ``` -Every black-box run reported below — the 100-turn campaign, the browser -regression, the offline container, the recovery run — was taken **after** the -last product change, on one frozen tree. Runs taken before a product change were +The black-box runs come from two trees. **The long-run evidence**, the 100-turn +campaign and its recovery run, was taken on commit `96c1bf5`, after the last +product change, with no source edit during the run. **The browser regression, +the offline container, the identity diagnostic and the contrast audit** were +re-run on the `ef25b0a` tree, as that commit's message records, and their +evidence is dated 2026-09-10. The three product commits after it (`f8d4010`, +`0c7316f`, `96c1bf5`) change backend files only, and those runs were not +repeated (§P). Runs taken before a product change on their own tree were discarded and re-run; §O records which and why. --- @@ -138,6 +148,20 @@ discarded and re-run; §O records which and why. **Dependencies: none added, none removed, none upgraded.** `requirements.txt`, `requirements.lock` and `package.json` are byte-identical to M10's. +**Changes after the first revision.** `144406c` is the first revision itself. +Every later commit is listed here, and all are signed by the owner. + +| Commit | Change | Files | +| --- | --- | --- | +| `fedb714` | Release evidence must be written to a durable `--out`, and the report recorded the lost run | `DEVELOPMENT.md`, this report | +| `ef25b0a` | The history window trims in blocks, so the server's prompt cache survives. `settings.context_window_override` (**migration 94**) covers a server the probe cannot ask. The long run gains `--resume`, a configurable turn timeout and an abort after consecutive failures. Two checks that could not fail were fixed. Every black-box harness was re-run on this tree | `app/context/builder.py`, `app/contextwindow.py`, `app/models.py`, `app/migrations.py`, `app/schemas.py`, `app/routers/settings.py`, `app/routers/adventures/turns.py`, `app/routers/adventures/insights.py`, `frontend/src/pages/Settings.jsx`, `tools/m11_long_run.py`, `tools/m11_browser.py`; tests `test_history_block_trim.py`, `test_m11_declared_window.py`, `test_m11_long_run_resume.py`, `settingsWindow.test.jsx`; planning package v3.8 | +| `fec46f6` | The long run switches the memory bank and auto-summarise on, and refuses a run with no embedding model | `tools/m11_long_run.py`, `tests/test_m11_long_run_memory.py` | +| `f8d4010` | §O.7: no write lock is held across the model call, and a failure that breaks the session is still recorded. The harness reads the real section labels and stops on failed post-turn work | `app/memorybank.py`, `app/routers/adventures/turns.py`, `app/routers/adventures/insights.py`, `tools/m11_long_run.py`, five test files | +| `0c7316f` | §O.8: protocol is kept out of stored narration, and the fence-label bug is fixed. The harness gains an M04 verdict and a protocol-leak count | `app/narrative/extract.py`, `app/narrative/render.py`, `tools/m11_long_run.py`, `tests/test_narrative_state.py`, `tests/test_m11_long_run_memory.py` | +| `96c1bf5` | §O.8 extended to a model's own sections under decorated headings. M04's precondition is now positional | `app/narrative/extract.py`, `tools/m11_long_run.py`, `tests/test_narrative_state.py`, `tests/test_m11_long_run_memory.py` | + +No later commit touches `requirements.txt`, `requirements.lock` or `package.json`. + --- ## D. M10 and M8 handoff @@ -198,9 +222,11 @@ turn's state. Before M11 all three settings produced *the identical sentence*; `test_the_three_lengths_no_longer_say_the_same_thing` fails against that. **D — character identity confusion.** The diagnostic exists; the root cause does -not, and cannot. §G.4 records what the run found, including a defect the -diagnostic caught in its own fixture — which is the reason the first run's -evidence was discarded rather than reported. +not, and cannot. **The diagnostic's run results are not reproduced in this +report.** The first revision pointed here to a section that was left empty. +What is recorded is the defect the diagnostic caught in its own fixture (§O.2), +which is why its first run's evidence was discarded. The later run's evidence is +in the implementer's `m11-evidence/identity-recheck` directory (§P). --- ## E. The context-window resolution @@ -314,48 +340,88 @@ Three separate reasons, and the third is the one that matters: --- ## E.1 The release environment -Recorded because a timeout without hardware beside it is not a measurement. +Recorded because a timeout without hardware beside it is not a measurement. Two +inference hosts produced this report's evidence, and every result below says +which. + +### The application machine + +It runs the storyteller, every harness, the browser and the container. | | | | --- | --- | -| **OS** | Ubuntu 24.04.4 LTS, kernel 7.0.0-30-generic | +| **OS** | Ubuntu 24.04.4 LTS, kernel 7.0.0-30-generic, at the first revision; Ubuntu 24.04.5 LTS, kernel 7.0.0-31-generic, for the 2026-09-13 long runs | | **CPU** | AMD Ryzen 7 7840HS, 4 cores available to this VM, 1 thread per core | -| **GPU** | **none** — VMware SVGA II; the inference host also reports `size_vram = 0` | +| **GPU** | **none**: VMware SVGA II | | **RAM** | 15 GiB | -| **Application** | branch `m11-release-validation`, base commit `1013c94` (signed) | | **Python** | 3.12.3 | | **Node / npm** | v22.23.1 / 10.9.8 | | **Docker** | 29.7.2 | | **Browser** | Firefox 154.0.1, headless, via geckodriver over W3C WebDriver | -| **Ollama** | 0.33.0, on a **separate physical machine** on the trusted LAN, HTTPS with a private CA installed in this machine's OS trust store | -| **Narrator (M01, capped run)** | `qwen2.5:3b-instruct` — no `num_ctx`, so this server gives it **4,096** | -| **Narrator (M01, 16k run; browser; identity)** | `qwen2.5:3b-instruct-16k` — `num_ctx` baked in, **16,384**, architecture ceiling 32,768 | -| **State/summariser model** | the same narrator model; no separate summariser is configured | -| **Embedding model** | `nomic-embed-text` | -| **Application context budget** | `context_token_budget` = 16,384 (default), `max_output_tokens` = 500 | -| **Effective budget, capped run** | **4,096** — the window, applied as a ceiling on every turn | -| **Effective budget, 16k run** | **16,384** — window and budget equal, nothing capped | -| **Warm or cold** | the narrator was warm for both campaigns after their first turn; the first turn of each is measurably slower and appears as such in the timelines | -| **Test date** | 2026-09-07 | -**Two long-run configurations, and why there are two.** The recommended -configuration is a model with `num_ctx` baked in, which is what -`DEVELOPMENT.md` tells an operator to do and what gives the application its full -16,384-token budget. Measured on this CPU-only host, that configuration costs -**229-291 seconds per turn** once the history window fills, because the whole -13-14k-token prompt is re-processed each turn. A hundred turns would be roughly -eight hours of wall clock on this hardware. +### The CPU reference host -So M01 was run in the **default** configuration instead — the plain model, whose -window this server sets to 4,096 — which is also the configuration M11's own fix -exists for: the application's budget stays at 16,384 and is capped to 4,096 on -every single turn. That makes the release campaign simultaneously the longest -available test of the fix. The recommended configuration's run is reported -beside it as far as it went (26 accepted turns), because its per-turn cost is -the useful thing it measured. +Used from 2026-09-07 to 2026-09-10: the lost run, the memory-off 100-turn run, +the browser run and the identity diagnostic. -Neither number is a performance requirement. The planning package contains none, -and none is invented here. +| | | +| --- | --- | +| **Machine** | a separate physical machine on the trusted LAN, no GPU (`size_vram = 0`) | +| **Ollama** | 0.33.0, **HTTPS** with a private CA installed in the application machine's OS trust store | +| **Measured cost** | 88-203 s per turn at a 4,096 window; 229-291 s at 16,384 | + +### The GPU inference host + +Used on 2026-09-13 and 2026-09-14: both 26-turn trials and the three 100-turn +runs, including the evidence run. + +| | | +| --- | --- | +| **Machine** | a separate physical machine on the trusted LAN; Ubuntu 24.04.2 LTS; 16 logical CPUs; 14 GiB RAM | +| **GPU** | NVIDIA GeForce GTX 1080 Ti, 11 GiB (Pascal, compute capability 6.1), driver 580.173.02, in an **OCuLink dock**: PCIe x4, measured at Gen 3 under load and Gen 1 at idle | +| **Ollama** | 0.34.0. Its CUDA 13 runner skips a compute-6.1 card and the CUDA 12 runner serves it; both models loaded 100% into VRAM | +| **Transport** | **plain HTTP** on the LAN. H12 permits a LAN address; this host is **not** A06 evidence | +| **Power limit** | the card's default, 280 W, during every run | +| **Measured cost** | prefill ~1,350-1,850 tokens/s and generation ~87 tokens/s with synthetic prompts; 4.1-22.8 s per turn in the evidence run | + +### Models and budget + +| | | +| --- | --- | +| **Narrator** | `qwen2.5:3b-instruct-16k`: `num_ctx` 16,384 baked in, architecture ceiling 32,768. Digest `21ff8cc52f37`, identical on both hosts | +| **Base narrator** | `qwen2.5:3b-instruct`, digest `357c53fb659c`, identical on both hosts. With no `num_ctx` the server gives it **4,096** | +| **State/summariser model** | the narrator; no separate summariser is configured | +| **Embedding model** | `nomic-embed-text`, digest `0a109f422b47`, identical on both hosts | +| **Application context budget** | `context_token_budget` = 16,384, `max_output_tokens` = 500, reply reserve 564 | +| **Effective budget** | 4,096 in the lost run, where it was capped every turn; **16,384** in every later run, where window and budget were equal | + +**Why the first campaign ran at 4,096.** On the CPU host the recommended 16,384 +window cost 229-291 s a turn, about eight hours for a hundred. The first revision +therefore ran the default window, which doubled as the longest test of §E's cap. +That run reached 97 of 100 turns and was lost (§G.6). Every later campaign runs +the recommended window. + +### The GPU host incident + +About 30 seconds after the evidence run's last post-turn pass finished, the GPU +dropped off the PCIe bus: `NVRM: Xid 79, GPU has fallen off the bus`, then +`Xid 154`, reset required. Ollama stayed up without a GPU, stopped completing +requests, and the host had to be rebooted. The same host had logged one `Xid 13` +graphics exception from the inference server earlier that day. There was no +out-of-memory event and no thermal event recorded. The firmware does not support +PCIe AER, so no link-error trail exists. + +After the reboot, under a temporary 200 W cap with logging running, the link held +Gen 3 x4 under load, peak draw was 207.8 W, and there was no fault. **The cause +is not established.** Power transients at the uncapped 280 W over a long run are +the leading candidate. **The evidence is unaffected**: every record the result +rests on was written before the drop (§G.1). + +**Required for every future long run on a GPU host:** log the card's power, +temperature, utilisation and PCIe link state, and the kernel's and the inference +server's messages, to disk for the whole run. A recurrence can then be tied to +power or ruled out. The commands are in `DEVELOPMENT.md`, "Logging the inference +host during a long run". --- ## F. Acceptance matrix @@ -374,9 +440,9 @@ evidence, it says so and the verdict is qualified. | A01 Start application offline | **PASS** | container: first page load from a fresh volume with no route and no DNS; every referenced asset served locally | | A02 Storyteller loopback default | **PASS** | suite `test_local_only_surface.py` (start scripts, compose publishes `127.0.0.1:8000:8000`); every M11 harness reached it only on loopback | | A03 No cloud API key | **PASS** | suite: no `api_key` in settings, none settable through the API, no Authorization header; container: no secret in an export | -| A04 Campaign survives restart | **PASS** | campaign: 4 genuine process restarts, transcript/head/state/Save Points/knowledge/settings compared before and after each; container: campaigns survive a container restart | +| A04 Campaign survives restart | **PASS** | campaign: 3 genuine process restarts (4 process starts), with transcript/head/state/Save Points/knowledge/settings compared before and after each and identical every time (§G.2); container: campaigns survive a container restart | | A05 Failed model call does not corrupt story | **PASS** | campaign: a real failed call at a controlled point (the model name pointed at one the server does not serve), then play resumed; container: same with no model reachable at all — no narration accepted, state unchanged, the player's words kept, earlier story reachable | -| A06 Trusted-LAN Ollama inference | **PASS** | campaign + browser: every turn in this report ran against Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass. The storyteller itself stayed loopback-bound | +| A06 Trusted-LAN Ollama inference | **PASS** | browser + identity diagnostic + the CPU-host campaigns: Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1) | ### B — Core play @@ -391,7 +457,7 @@ evidence, it says so and the verdict is qualified. | ID | Result | Evidence | | --- | --- | --- | -| C01 Campaign canon is preserved | **PASS** | campaign: the canon section is present in every turn's stored prompt (`canon_present` per turn); suite | +| C01 Campaign canon is preserved | **PASS** | campaign: the canon section is in the prompt on 101 of 101 turns (`canon_tokens`, §G.3); suite | | C02 Possession state | **PASS** | campaign (the silver key) + suite + sci-fi fixture (the data crystal) | | C03 Character knowledge is not invented | **PASS** | suite `test_worldstate_integration.py`, `test_narrative_state.py` | | C04 Manual state correction | **PASS** | campaign: two corrections, both accepted, one carrying the planted clue; suite, including the M11 regression that a *partly* refused correction now says so | @@ -407,7 +473,7 @@ evidence, it says so and the verdict is qualified. | D03 Unlimited undo *(SHOULD)* | **PASS** | suite: undo to the root and back | | D04 Redo | **PASS** | browser (returns to the same position) + campaign | | D05 Redo invalidated by new continuation | **PASS** | campaign: after diverging, redo is no longer available — recorded in the timeline | -| D06 Retry narrator response | **PASS** | campaign: three retries at scheduled points | +| D06 Retry narrator response | **PASS** | campaign: two retries at scheduled points (§G.2) | | D07 Select prior retry take | **PASS** | campaign: take selection back to index 0 | | D08 Retry does not delete prior take | **PASS** | campaign: take count on the turn after retry; suite | | D09 Edit earlier user input | **PASS** | suite `test_take_edit.py`, `editRouting.test.jsx` | @@ -431,13 +497,13 @@ evidence, it says so and the verdict is qualified. | ID | Result | Evidence | | --- | --- | --- | | F01 Recent turns remain coherent | **PASS** | campaign: the history window is populated every turn and the newest turns are always included | -| F02 Old important event retrieval | **PASS** | campaign M04 (§G.3) | +| F02 Old important event retrieval | **PASS** | campaign M04 (§G.4): the planting turn outside the history window and the fact recovered, through authoritative state and the narrator's restatements rather than independent memory retention | | F03 Prompt remains bounded | **PASS** | campaign: prompt size across 100 turns (§N), plus the cap itself (§E) | | F04 Output token reserve | **PASS** | campaign: `output_reserve` present in every turn's measurement and subtracted before history is chosen | | F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt and its budget | | F06 Retrieval provenance | **PASS** | campaign: `knowledge.used` per turn in the stored snapshot; suite | | F07 Heuristic memory is not canon | **PASS** | suite `test_memory_nodes.py` (authority) | -| F08 Memory failure is non-fatal | **PASS** | suite `test_context_memory.py`; container: derived work fails with no model and turns still commit | +| F08 Memory failure is non-fatal | **PASS** | suite `test_context_memory.py`, including §O.7's regressions that a failure is recorded even when it breaks the session; container: derived work fails with no model and turns still commit; campaign: post-turn work checked after every turn, 0 failures | ### G — Imported knowledge @@ -458,7 +524,7 @@ evidence, it says so and the verdict is qualified. | ID | Result | Evidence | | --- | --- | --- | -| H01 No unexpected outbound connections | **PASS** | container (no network at all) + campaign (one destination: the configured Ollama) + suite `test_egress.py` | +| H01 No unexpected outbound connections | **PASS** | container (no network at all) + the CPU-host campaign (one destination, the configured Ollama; the later runs were not network-monitored, §J) + suite `test_egress.py` | | H02 No telemetry | **PASS** | suite + dependency audit (§M) | | H03 No cloud provider required | **PASS** | container: a full campaign offline; suite | | H04 Model output cannot execute shell | **PASS** | browser (shell text rendered as text) + suite (no subprocess/eval anywhere in the turn path) | @@ -513,10 +579,10 @@ evidence, it says so and the verdict is qualified. | ID | Result | Evidence | | --- | --- | --- | -| M01 100-turn campaign | see §G | campaign | -| M02 Restart during long campaign | see §G | campaign | -| M03 Long-run context stability | see §G | campaign | -| M04 Long-run memory recall | see §G | campaign | +| M01 100-turn campaign | **PASS** | campaign: 101 accepted turns, every scheduled operation, 0 failed post-turn passes (§G) | +| M02 Restart during long campaign | **PASS** | campaign: 3 genuine process restarts, one carrying retained history, all identical (§G.2) | +| M03 Long-run context stability | **PASS** | campaign: the prompt held between 14.6k and 15.7k tokens for 70 turns, the window verified 101/101, the canon present on every turn (§G.3, §N) | +| M04 Long-run memory recall | **PASS** | campaign: the planting turn outside the window, the fact recovered through state and through memory of the narrator's restatements (§G.4) | ### SHOULD and FUTURE disposition @@ -530,157 +596,155 @@ blockers and no part of this milestone treats them as one. --- ## G. The long-run campaign — M01 to M04 -**M01 is the one REQUIRED test this report cannot certify, and this section says -exactly how far it got and why.** +**M01 to M04 pass on a complete 100-turn campaign run by the committed +harness.** This section reports that run (§G.1-§G.5), then every earlier run +and why none of them is the evidence (§G.6). -### G.0 Status: PARTIAL — 41 accepted turns of 100; the run was later lost - -> **Addendum, added 2026-09-07, after this section was written.** This section -> was drafted while the run was still in progress, and two things have happened -> since. Both count against this report's evidence. -> -> **The run went further, then died.** It continued unattended past the 41 turns -> described below and reached **97 of 100 accepted turns**, with all thirteen -> scheduled history operations fired. It never finished: the host crashed and -> rebooted, and the run went with it. -> -> **Its evidence did not survive.** The run wrote to a directory under `/tmp`, -> which the reboot cleared. `campaign.db`, `timeline.jsonl`, `server.log` and the -> backups are gone — **both the 41-turn artifacts this section cites and the -> 97-turn continuation.** The figures in §G.1-§G.5 and §N were transcribed from -> that run while it was live and are reported here as they were observed, but -> they can no longer be produced on request and nothing in them can be -> independently re-checked. -> -> **What this means for acceptance.** M01 must be **re-run from scratch** before -> v1 acceptance, on a host that can finish it, writing to a durable path rather -> than `/tmp` (`DEVELOPMENT.md` now says so and the harnesses require `--out`). -> Until that run exists, treat every M01-derived number in this report as an -> unverifiable observation rather than as evidence. Nothing else in the report -> depends on it: the suite, browser, offline, recovery and migration results were -> produced by harnesses that can be re-run in minutes. - -The campaign is correct as far as it has gone: every history operation performed -as designed, the restart was byte-identical, the prompt stayed bounded, the -window was verified on every single turn, and the campaign exported and moved to -a clean machine intact (§K). What is missing is turns 42-100, and the reason is -wall-clock on the reference host rather than anything the application did. - -**The binding constraint is measured, not asserted.** On this CPU-only, -no-VRAM inference host a 3B narrator costs: - -| Configuration | Window | Prompt at steady state | Seconds per turn | 100 turns would take | -| --- | --- | --- | --- | --- | -| Recommended (`num_ctx` baked in) | 16,384 | 13-14k tokens | **229-291** | ~8 hours | -| Default (no `num_ctx`) | 4,096 | ~3.4k tokens | **88-203** | ~4 hours | - -Both runs were made. The recommended configuration reached **26 accepted turns** -before being stopped in favour of the faster one; the default configuration is -the campaign reported below and was at **41 accepted turns** when this report was -written, still progressing unattended. - -**What a reviewer should do with this.** The harness is committed -(`backend/tools/m11_long_run.py`) and the command is in `DEVELOPMENT.md`. On a -host with a GPU — or overnight on this one — the run completes without -supervision and writes `summary.json`, `recall.json` and `bundle.json`. M01 -should be re-checked from that output before v1 acceptance. Everything M01 is -*for* other than the turn count — continuity, state, history operations, -restarts, context stability, recovery — is evidenced below at 41 turns and in -§K at 83 actions. - -### G.1 The run +### G.1 The evidence run | | | | --- | --- | -| **Accepted turns** | **40** | -| Application restarts | 1 genuine `uvicorn` process boundaries | -| Turn time | min 88s, median 178s, max 203s | +| **Tree** | commit `96c1bf5`, signed; `tools/m11_long_run.py --turns 100` | +| **Inference** | the GPU inference host (§E.1), `qwen2.5:3b-instruct-16k`, window 16,384 | +| **Evidence** | `$HOME/m11-evidence/m04-final/`: `timeline.jsonl`, `summary.json`, `recall.json`, `bundle.json`, `campaign.db`, `server.log`, `recovery-report.json` | +| **Status** | `complete`; no aborted reason, no failed reason | +| **Accepted turns** | **101**: 100 scheduled beats and the recall turn | +| **Genuine process restarts** | **3**, making 4 process starts | +| **Wall clock** | 1,238 s; a turn took 4.1 s at least, 9.6 s median, 22.8 s at most | +| **Memory bank and auto-summarise** | switched on at setup and read back as on | +| **Post-turn work** | checked after every turn against derived status and new `server.log` lines: **0** failed passes, 0 `database is locked`, 0 unrecorded failures, 0 tracebacks | +| **Written** | 33 memories in the database across every branch, 19 reported by the `/memories` endpoint at the end; 12 summaries | +| **Protocol in stored narration** | 0 of 104 AI turns | + +**Timing against the GPU fault (§E.1).** The last AI action was written at +02:31:19 UTC, the memory, summary and embedding passes finished at 02:31:22, +and the GPU dropped off the bus at 02:31:52. The run's final settle and summary +followed; nothing they report was pending. ### G.2 Timeline of operations +"At turn" is the accepted-turn count when the operation fired. + | At turn | Operation | Outcome | | --- | --- | --- | -| 0 | `state_correction` | {"events": 8, "note": "the opening cast"} | -| 1 | `state_correction` | {"events": 1, "note": "the planted clue, as accepted state"} | -| 6 | `save_point` | {"id": 1, "name": "Before the ridge"} | -| 13 | `restart` | {"number": 1, "identical": true} | -| 20 | `undo_redo` | {"after_undo": 39, "after_redo": 41, "restored": true} | -| 27 | `retry` | {"ok": true, "takes_on_newest_turn": 2, "detail": ""} | -| 34 | `save_point` | {"id": 2, "name": "On the ridge"} | +| 0 | memory bank and auto-summarise | both read back as on | +| 0 | knowledge import | canon, reference and inspiration sources | +| 0 | state correction | 8 events, the opening cast | +| 1 | state correction, clue planted | the clue accepted as state; verified in state; the planting turn at depth 1 | +| 6 | Save Point | *Before the ridge* | +| 13 | restart 1 | `identical: true` | +| 20 | Undo, then Redo | 41 → 39 → 41 actions, position restored | +| 27 | Retry | two takes on the newest turn | +| 34 | Save Point | *On the ridge* | +| 41 | Undo, then restart 2 with the undone history retained | `identical: true`, Redo available on both sides | +| 48 | Retry | two takes on the newest turn | +| 56 | Undo, then new writing | Redo no longer available | +| 62 | take selection | index 0 of 2 chosen, and it became live | +| 69 | failed model call | reported: HTTP 404 for a model the server does not serve; state unchanged | +| 76 | restore *On the ridge* | active line 148 → 69 actions; the later story retained; Redo available | +| 83 | restart 3 | `identical: true` | +| 100 | recall check | §G.4 | +| 101 | export | 2,743,080 bytes; 0 AI turns carrying protocol | ### G.3 Context growth +The application's token counts, from `timeline.jsonl`. + ```text - turn actions in prompt prompt summary memory knowl state canon window - 1 3 3 1516 0 0 199 93 57 4096 - 10 21 10 3501 0 0 199 135 57 4096 - 20 41 10 3436 0 0 199 154 57 4096 - 30 61 10 3506 0 0 199 154 57 4096 - 40 81 8 3266 0 0 199 154 57 4096 + turn actions in prompt prompt summary memory knowl state canon floor window + 1 3 3 1,441 0 0 127 93 57 - 16384 + 10 21 21 6,422 185 423 330 137 57 - 16384 + 20 41 41 11,447 603 544 230 161 57 - 16384 + 30 61 55 15,299 275 700 230 165 57 6 16384 + 40 81 57 15,300 289 809 230 194 57 24 16384 + 50 99 57 15,474 395 946 230 194 57 42 16384 + 60 115 55 14,718 374 875 127 194 57 60 16384 + 70 136 58 15,549 244 1,000 127 194 57 78 16384 + 80 77 59 15,203 289 664 127 164 57 18 16384 + 90 97 61 15,504 217 764 227 193 57 36 16384 + 100 117 63 15,745 146 757 127 193 57 54 16384 ``` -- budget: 4096 (configured 16384) -- output reserve, every turn: 564 -- window verified on 40 of 40 turns -- canon present in the prompt on 5 of 40 turns -- largest prompt: 3523 tokens +- budget 16,384 (configured 16,384); reply reserve 564 on every turn +- window verified on **101 of 101** turns +- campaign canon in the prompt on 101 of 101 turns; imported knowledge on 101; memories on 94; a summary on 93 +- largest assembled prompt 15,745 tokens by the application's count; the history window was first trimmed at turn 30 +- the drop in actions at turn 80 is the Save Point restore at turn 76 -### G.4 Recall +### G.4 Recall (M04) -```json -{} -``` +**The precondition is positional.** M04's pass text is "Fact/event remains +recoverable without entire transcript in prompt", so the harness asks whether +the planting turn has left the history window. It does not ask whether the +clue's text has, because the narrator reuses that text in its own prose (below). + +| | | +| --- | --- | +| Planting turn's depth | 1 | +| History window floor at the recall check | depth 54 | +| **Planting turn in the history window** | **no** | +| Clue's sentinel text somewhere in recent history | yes: the narrator's restatements, below | +| Clue in the narrative-state section | yes | +| Clue in the memories section | **yes** | +| Clue in the summary section | no | +| Clue in imported knowledge | no | +| Fact still in authoritative state | yes | +| History after the recall turn | 59 of 119 actions, in a 14,638-token prompt | +| **Verdict** | **`recovered_through_memory_or_summary`** | + +**How the fact reached memory.** From depth 30 onward the narrator wrote the +state section's fact line into its prose as an ordinary sentence, on 74 of 104 +AI turns: + +> Aldric knew the silver key opened the crypt beneath it (SILVER-KEY-CRYPT-OLD-ABBEY). + +The memory pass summarises narration, so the 16 memories that carry the clue all +summarise depth 54 or later. None comes from the planting era. The chain is +authoritative state, then the narrator's restatement, then memory. The fact +survived because state carried it. The repository owner accepted state-based +recovery as satisfying M04 on 2026-09-13. + +A restatement sits inside a line of story, so the extractor rightly leaves it. +§P records what it means for narration quality. ### G.5 What the numbers say -**M03 — long-run context stability: PASS.** The clearest result in the run. The -story grew from 3 actions to 83; the *prompt* grew from 1,516 tokens to about -3,400 and then stopped, because the history window stopped taking more. At turn -41 the prompt carried **8 of 83 actions** — the whole transcript is emphatically -not being appended. The output reserve of 564 tokens was subtracted on every -turn without exception, and the campaign canon section was present on **41 of 41 -turns**, which is the protected-content half of M03. +**M01, 100-turn campaign: PASS.** 101 accepted turns with none refused. Every +one of the thirteen scheduled history operations performed as designed. All +three restarts compared transcript, head, Undo/Redo availability, scene, +entities, facts, Save Points, imported knowledge and settings and found them +identical. The induced failed call changed nothing. The campaign moved to a +clean data directory 16 of 16 (§K). The step list's summary/memory activation +was exercised: 33 memories and 12 summaries were written with no failed pass. -**The window was verified on 41 of 41 turns**, and on every one of them the -budget was capped from the configured 16,384 to the server's real 4,096. This -campaign is therefore also the longest available test of §E's fix: forty-one -consecutive turns in the exact configuration that used to truncate silently, with -the canon still at the front of every prompt. +**M02, restart during a long campaign: PASS.** Three genuine `uvicorn` process +boundaries, one of them carrying undone history across. Each was identical. -**M02 — genuine process restarts: PASS as far as it went.** One restart, at turn -13, comparing transcript, head, Undo/Redo availability, scene, entities, facts, -Save Point names, imported knowledge and model settings across the boundary: -`identical: true`. The 16k run performed its own restart at turn 13 with the same -result. The full campaign schedules four; two more (one deliberately with -retained history) fall after turn 41. +**M03, long-run context stability: PASS.** The prompt reached about 15.3k tokens +at turn 30 and stayed between 14.6k and 15.7k for the next 70 turns, while the +active story grew to 136 actions. The canon was present on every turn, and the +reply reserve was subtracted on every turn. By the narrator's own tokenizer the +largest prompt plus the reserve is 16,342 of 16,384, which is bounded but nearly +full (§N). -**History operations: all those scheduled so far performed correctly.** +**M04, long-run memory recall: PASS**, on §G.4's positional precondition, with +the recovery path stated there. -- turn 6 — Save Point *Before the ridge* created -- turn 13 — restart, state identical across the process boundary -- turn 20 — Undo then Redo, returning to exactly the position it left (41 → 39 → 41) -- turn 27 — Retry, producing a second take on the newest turn -- turn 34 — Save Point *On the ridge* created +### G.6 The runs that led here -Scheduled after turn 41 and therefore **not yet exercised in this run**: the -second and third restarts, the second Retry, Undo-then-divergence, take -selection, the induced failed model call, and the Save Point restore. Each of -those is separately covered by the automated suites and by the 14-turn shakeout -run, but not yet inside this campaign — which is part of why M01 is PARTIAL -rather than PASS. +Every run's evidence directory is kept. -**M04 — long-run memory recall: NOT YET REACHED.** The recall check runs after -the last turn. Its precondition is, however, already established and measurable: -the planted clue (`SILVER-KEY-CRYPT-OLD-ABBEY`) was visible in the assembled -prompt for the first **5 turns** and has been outside it since — so by turn 41 it -is already only reachable through state, summary, memory or retrieval, which is -exactly the condition M04 asks for. It is recorded in the authoritative state as -a fact, which is one of the four paths. +| # | Run | Tree | Host | Result | Why it is not the evidence | +| --- | --- | --- | --- | --- | --- | +| 1 | first release campaign, 2026-09-07 | first revision | CPU, 4,096 window | reached 97 of 100 turns | the host crashed; its evidence was under `/tmp` and the reboot cleared it. The harness now requires `--out` and checkpoints for `--resume` | +| 2 | 100 turns, 2026-09-10 | `ef25b0a` | CPU, 16,384 | complete: 101 turns, 3 restarts, 45,656 s | **the memory bank and auto-summarise were never switched on** (harness defect, §O), so summary/memory activation went untested and recall ran through state only | +| 3 | 26-turn trial, 2026-09-13 | `fec46f6` | GPU | reported `complete` | 180 `database is locked`, 20 failures that could not be recorded, 2 memories, 0 summaries. Found §O.7 and two harness defects | +| 4 | 26-turn trial | `f8d4010` | GPU | 0 lock errors, 7 memories, 2 summaries, 258 s | a trial, not a campaign | +| 5 | 100 turns | `f8d4010` | GPU | complete: 101 turns, 1,289 s, 33 memories, 12 summaries | the narrator's protocol was stored as story on 42 of 104 turns (§O.8), and a pasted state section kept the clue in recent history, so M04 proved nothing | +| 6 | 100 turns | `0c7316f` | GPU | complete: 101 turns, 1,251 s, 1 leak by the harness count | the narrator wrote `## Established:` sections of its own on 5 turns, which the extractor and the harness's leak count both missed; verdict `precondition_not_met` on the sentinel-text precondition. Reclassified on the positional precondition, in `recall-reclassified.json` beside the original, as `recovered_through_state_only`. Superseded by run 7 on the corrected tree | +| 7 | **the evidence run** | `96c1bf5` | GPU | §G.1-§G.5 | | -**Defects encountered during the campaign: none.** No turn was refused, no state -correction was rejected, no restart lost anything, and no step failed. The two -harness defects that this campaign's earlier shakeout runs found (SSE handling on -turns and on retry) are in §O. +--- ## H. Lineage leakage — E01 to E04 `tests/test_m11_leakage.py`, **14 tests**, all passing. What M11 adds to the @@ -782,12 +846,16 @@ error and the campaign stays intact. ### Observed outbound destinations -During the 100-turn campaign the application contacted exactly one host: the -configured trusted-LAN Ollama, over HTTPS with a private CA in the OS trust -store, on the `/v1` path for inference and the `/api/ps`+`/api/show` paths for -the window probe. No other destination, and no DNS lookup for any other name. -In the container run there was no destination at all, because there was no -network. +During the first revision's long campaign, on the CPU reference host, the +application contacted exactly one host: the configured trusted-LAN Ollama, over +HTTPS with a private CA in the OS trust store. It used the `/v1` path for +inference and the `/api/ps` and `/api/show` paths for the window probe. There was +no other destination, and no DNS lookup for any other name. In the container run +there was no destination at all, because there was no network. + +**The later long runs were not network-monitored.** They were configured with one +endpoint, the GPU inference host over plain HTTP. `tests/test_egress.py` is what +continues to assert that nothing else is contacted. ### H-series @@ -818,37 +886,35 @@ Full per-test results are in §F. The M11-specific additions: ### The long campaign, moved to a machine that has never seen it -`tools/m11_recovery.py` — **16 checks, 0 failed**. The input is not a fixture: it -is the release campaign's own database, exported through the API and imported +`tools/m11_recovery.py`: **16 checks, 0 failed**. The input is not a fixture. It +is the evidence run's own campaign (§G.1), exported through the API and imported into a database file that did not exist, in a directory that did not exist, opened by a second server process. Migrations ran there from nothing, so this is the fresh-install path as well as the import path. -**The bundle was taken with M9's online backup API on the live database while the -campaign was still playing** — `integrity: ok`, 160 pages, 655,360 bytes — which -is both how a consistent snapshot of a running campaign is obtained and an extra -exercise of that backup path on a real long-run database. - | | | | --- | --- | -| Bundle | 654,803 bytes — **3.1% of the 20 MB import limit** at 83 actions | -| Actions in the bundle / on the active line after import | 83 / 82 | -| Retained beyond the active line | 1 | -| Entities, facts | 6, 1 | -| Save Points | 2, both restore to the positions they name | -| Knowledge sources | 3 — canon, reference and inspiration, with content, not just filenames | +| Bundle | 2,743,080 bytes, **13.1% of the 20 MB import limit**, at 207 actions | +| Actions in the bundle / on the active line after import | 207 / 119 | +| Retained beyond the active line | 88 | +| Entities, facts | 6, 2 | +| Save Points | 2; *On the ridge* restores to 69 actions and *Before the ridge* to 13, the positions they name | +| Knowledge sources | 3: canon, reference and inspiration, with content, not just filenames | | Narration-length choice | came across | | Campaign canon | came across | +| Redo | available exactly when the file said so | | Secrets in the bundle | none | -| The moved campaign | accepts a new correction, and exports again at the same story length | +| The moved campaign | accepts a new correction with nothing refused, and exports again at the same story length, 207 actions | -**On the bundle ceiling.** M9 measured about 279 turns against the 20 MB import -limit using a fixture built to be heavy. This real campaign is 3.1% of the limit -at 83 actions, which extrapolates to well beyond the 100-turn certification -target — the difference from M9's estimate being that this campaign's stored -prompts are small, because the window is 4,096 rather than 16,384. A campaign -played at the recommended 16k window would carry proportionally larger prompt -snapshots, and M9's number remains the conservative one to quote. +The first revision's recovery run moved a 41-turn campaign at the 4,096 window +(654,803 bytes, 83 actions), and that campaign's files were lost. The run above +replaces it. + +**On the bundle ceiling.** M9 estimated about 279 turns against the 20 MB limit, +using a fixture built to be heavy. This real campaign plays at the recommended +16,384 window, with every stored prompt near that size. It averages about 13 kB +per action, which puts the limit near 1,600 actions. M9's number remains the +conservative one to quote. ### Migration @@ -870,7 +936,9 @@ the permanent form of the defect M10 found by accident: The comparison is written generically rather than about `visual_profiles`, so a future migration that diverges the two paths fails here whatever it is about. -**M11's own migration** is number 93, one column on `adventures`, no backfill. +**M11's own migrations** are numbers 93 and 94: `adventures.narration_length` in the +first revision and `settings.context_window_override` in `ef25b0a`, one column +each, no backfill. The knowledge-migration suite's version assertions were corrected from a literal `== 92` to `== LATEST_VERSION >= M7_VERSION`, because they had been asserting that M7's version was the newest — true when written, and a statement about M11 @@ -1009,48 +1077,85 @@ and none is invented here. ### Long-run storage +The evidence run, §G.1: + | | | | --- | --- | -| Database after 41 accepted turns (83 actions, 3 imported sources) | **663,552 bytes** | -| Export of the same campaign | **654,803 bytes** — 3.1% of the 20 MB import limit | -| Per action, roughly | ~8 kB, dominated by the per-position state snapshot and the stored prompt | -| Backup of the live database | 160 pages, 655,360 bytes, `integrity: ok` | +| Database after 101 accepted turns: 207 actions across the active line and retained history, 33 memories, 12 summaries, 3 imported sources | **2,367,488 bytes** | +| Export of the same campaign | **2,743,080 bytes**, 13.1% of the 20 MB import limit | +| Per action, roughly | ~11 kB in the database and ~13 kB in the export, dominated by the per-position state snapshot and a stored prompt of up to ~16k tokens | + +No live backup was taken during this run. The first revision's backup figure +(160 pages, `integrity: ok`) came from the lost campaign. ### Prompt size across the run +The application's count: + ```text -turn 1 1,516 tokens 3 of 3 actions in the prompt -turn 10 3,501 tokens 10 of 21 -turn 20 3,436 tokens 10 of 41 -turn 30 3,506 tokens 10 of 61 -turn 41 3,266 tokens 8 of 83 +turn 1 1,441 tokens 3 of 3 actions in the prompt +turn 10 6,422 tokens 21 of 21 +turn 20 11,447 tokens 41 of 41 +turn 30 15,299 tokens 55 of 61 +turn 40 15,300 tokens 57 of 81 +turn 50 15,474 tokens 57 of 99 +turn 60 14,718 tokens 55 of 115 +turn 70 15,549 tokens 58 of 136 +turn 80 15,203 tokens 59 of 77 (after the Save Point restore) +turn 90 15,504 tokens 61 of 97 +turn 100 15,745 tokens 63 of 117 ``` -The prompt rises until the history window is full and then **stops**, which is -F03's claim measured rather than argued. The history window's *contents* keep -moving — always the newest turns — while its size stays put. At turn 41 the -prompt carries 8 of 83 actions. +The prompt rises until the history window is full and then **stops**. That is +F03's claim measured rather than argued. The window's contents keep moving, +always to the newest turns, while its size stays put. -### Inference cost, by configuration +### Real-token headroom -| Window | Prompt at steady state | Seconds per turn | -| --- | --- | --- | -| 4,096 (default) | ~3.4k tokens | 88-203, median 178 | -| 16,384 (recommended) | 13-14k tokens | 229-291 | +The application counts tokens with `cl100k_base`, and the narrator counts with +Qwen's own tokenizer. Each run's largest stored prompts were re-sent to the same +model, identified by digest, and its `prompt_eval_count` was read: -Prompt processing dominates: a four-times-larger prompt costs roughly twice the -wall clock per turn on this CPU. This is the reference host's characteristic and -the reason M01 is PARTIAL (§G). +| Run | Largest prompt, narrator's count | Plus reply reserve 564 | Headroom in 16,384 | +| --- | --- | --- | --- | +| CPU host, memory off (`ef25b0a`) | 15,728 | 16,292 | 92 | +| GPU host, 100 turns (`f8d4010`) | 15,797 | 16,361 | **23** | +| GPU host, 100 turns (`0c7316f`) | 15,786 | 16,350 | 34 | +| **GPU host, evidence run (`96c1bf5`)** | **15,778** | **16,342** | **42** | + +**No prompt in any run exceeded the window.** On every prompt re-counted, the +application's count was 16 tokens below the narrator's. The evidence run's +re-count was taken after the GPU host's reboot, with the model digest verified +unchanged. + +The margin matters because of how the server fails. Ollama 0.34 cuts a prompt +longer than the window down to 8,194 tokens and returns no error. That was +observed with synthetic prompts, not in a campaign. §P records it as a risk. + +### Inference cost, by host + +| Host | Window | Prompt at steady state | Seconds per turn | 100 turns | +| --- | --- | --- | --- | --- | +| CPU reference host | 4,096 | ~3.4k tokens | 88-203, median 178 | ~4 h projected; that run was lost at 97 | +| CPU reference host | 16,384 | 13-14k tokens | 229-291 in the first attempt | 45,656 s for 101 turns, with memory off | +| GPU inference host | 16,384 | 14.6-15.7k tokens | 4.1-22.8, median 9.6 | 1,238 s for 101 turns, with memory on | + +Prompt processing dominates on the CPU host. On the GPU host a full 16k prompt is +processed in about ten seconds (§E.1). Both are these machines' characteristics. ### Query behaviour No new query growth was introduced. M11 adds one HTTP round trip per session per -`(endpoint, model)` pair — the window probe, cached for ten minutes on success -and one minute on failure — and one derived computation per state read -(`duplicate_names`, a single pass over the entities already in memory). The +`(endpoint, model)` pair, the window probe, cached for ten minutes on success and +one minute on failure. It adds one derived computation per state read, +`duplicate_names`, a single pass over the entities already in memory. The connection test was changed to use that cache after the browser run showed the model-status badge calling it on every page load. +The §O.7 correction moves one statement rather than adding one: the memory +use-counter UPDATE now runs inside the turn's existing commit instead of before +the model call. + ### Nothing pathological was found No unbounded growth, no per-row query, no repeated snapshot write, no duplicated @@ -1148,6 +1253,95 @@ structural half). Deliberately **not** made an error: two people called Alice is ordinary fiction. Made visible instead — `duplicate_names` in the state API and the State panel. +### Product defects — found by the long runs after the first revision, fixed + +**O.7 — A turn held SQLite's write lock through the model call, and post-turn +memory and summary work was lost without a trace** +**Severity: high. Requirement: M01's summary/memory activation, F02, F08, M04. +Blocker: yes, for M01. Status: fixed in `f8d4010`.** + +*Reproduction:* turn the memory bank and auto-summarise on, configure an +embedding model, and play consecutive turns once a memory exists. On a fast +inference host: `database is locked` in the server log, memories stop +accumulating, no summary is written, and derived status reads `idle` with no +failures. The first 26-turn trial (§G.6, run 3) logged 180 such errors, wrote 2 +memories and no summary, and reported `complete`. + +*Root cause:* `retrieve_memories(update_stats=True)` ran an uncommitted UPDATE of +the used memories' counters before the model call. The turn commits once, after +the reply has streamed, so the write transaction stayed open for the whole reply. +SQLite has one writer. Every post-turn memory, summary and status write in that +window waited out the driver's five-second timeout and failed. Recording the +failure needs a write as well. The post-turn task's outer handler recorded +without rolling back first, so it raised `PendingRollbackError` and the failure +reached only the log. F08 requires a memory failure to be visible, and this one +was not. No earlier long run hit it, because none had the memory bank on. + +*Correction:* retrieval only reads. `memorybank.record_use` writes the counters in +the turn's single commit, so a turn that never lands counts nothing. The outer +handler rolls back before it records. + +*Regression evidence:* `test_no_write_lock_is_held_while_the_narrator_is_talking` +probes for the lock from a second connection during the model call. +`test_a_failure_that_breaks_the_session_is_still_recorded` covers the recorder. +**Both fail on `fec46f6`**, with `database is locked` and `idle` respectively. +Also `test_a_failed_turn_counts_no_memory_as_used`. The next trial had 0 lock +errors, 7 memories and 2 summaries. + +**O.8 — The narrator's protocol was stored as story** +**Severity: high. Requirement: story authority (M5 review Finding 4), C04, and +the validity of M04. Blocker: yes, for M04's evidence. Status: fixed in `0c7316f` +and `96c1bf5`.** + +*Reproduction:* play a long campaign against a small local model, then search the +stored AI turns for the state section's headings or `"events"`. In the `f8d4010` +run 42 of 104 turns carried protocol, the first at depth 2, in four shapes: + +- a copy of the narrative-state section: `Scene:`, `Who and what exists:`, `Held:`, `Established:`, `Still open:` +- that copy above a correct ```` ```state ```` block, which was removed while the copy stayed +- the copy, a bare `State` heading, and a `> {"events": ...}` proposal quoted like a player turn, sometimes with story after it +- the same block cut off by the output-token limit, on 10 turns + +In the next run the narrator wrote sections of its own instead, such as +`## Established:` over indented facts, on 5 turns. + +*Root cause:* the extractor removed fenced blocks, a bare object at the very end, +and a parroted bracketed reminder. It did not remove a pasted state section, an +unfenced proposal elsewhere in the reply, or an unfinished one. Stored text is +replayed verbatim as history (`_history_text`). Each leak therefore put a second, +older account of the state into the next prompt, which is what Finding 4 +removed from replay, and it gave the model another example to copy. A second, +older bug was in the fence pattern itself. `_STATE_FENCE_RE` read "a ```state +block" inside a parroted reminder as a fence opening and cut out the middle of +the reminder. + +*Correction:* the extractor recognises the renderer's own section headings, which +are now named constants in `render.py`, with any markdown wrapped around them. It +removes a block carrying two headings, or one heading with an indented entry. It +removes an unfenced proposal that starts a line, quoted or not, taking outermost +objects first, and uses it as the turn's proposal when there is no fence. It +removes an unfinished proposal at the end and whatever is left behind at the end: +a `State` heading, a bare `>`, a parroted reminder or continue hint, closed or +not, and a ```json fence cut off before it names its events. The fence label must +now end its line or run straight into the payload. A reply whose only removal is +a pasted section records no raw block, so the turn is not marked unparseable for +a block it never started. + +*Regression evidence:* 25 new cases, from 18 test functions, in +`test_narrative_state.py`, cut down from +the runs' real output, including negative controls. A lone `Held:` with prose +after it, a `Scene:` line of story, quoted JSON that is not a proposal, and a +`State` line followed by story all stay. A test renders every section the +renderer writes and pastes the lot. Beyond the suite, **every AI turn in five real +runs was replayed through the new extractor, 443 in all, and no turn the old +extractor had left clean changed.** The evidence run stored 0 of 104 turns with +protocol. + +*Not removed, deliberately:* a restatement of a fact inside a line of story (§G.4), +and headings a model invents that are not the renderer's (`Identifiers +established:`, on 10 turns of run 3). Removing either means judging prose, and +§P records both. + ### Harness and test defects — found and fixed, no product change Recorded separately, and at this length, because M8's review found five harness @@ -1166,6 +1360,11 @@ defects. A harness that has only ever agreed with itself is not evidence. | **Two substring checks matched inside words** | "ahead" contains "head"; "immediately" contains "media" | both now match whole words or parse imports | | **A contrast check treated a control boundary as body text** | would have failed the run on a WCAG clause that does not apply | now distinguishes 1.4.3 from 1.4.11 and says which | | **An audit assertion read a deferred column after the session closed** | `DetachedInstanceError` | read inside the session | +| **The long-run harness never switched the memory bank or auto-summarise on** *(fixed in `fec46f6`)* | both are per-campaign and default to off, so a complete 100-turn run wrote no memory and no summary | M01's summary/memory activation clause was reported by silence, and M04's memory path was never asked. The harness now switches both on, reads them back, and refuses a run with no embedding model | +| **It read three prompt sections under names the builder does not use** *(fixed in `f8d4010`)* | `memories` (really `used_memories`), `story_history` (really `history` and `recent_history`), and a `knowledge` prefix that matched the fixed instruction section instead of the imported passages | memory tokens read 0 whatever the prompt held, and the in-history and in-memories recall checks could never be true. The labels are now constants pinned by a test against a prompt the real builder assembled | +| **It had no way to see failed post-turn work** *(fixed in `f8d4010`)* | it reported `complete` over 180 `database is locked` errors | it would have certified the run that §O.7 destroyed. It now checks derived status and new `server.log` lines after every turn and stops at the first failure; a run with no memories or no summaries ends `failed` | +| **Its protocol-leak count used plain substrings** *(fixed in `96c1bf5`)* | `## Established:` did not match `\nEstablished:\n`, so the count read 1 where 5 turns leaked | the count now matches any state heading with an indented entry, in any markdown | +| **Its M04 precondition was the clue's text being out of recent history** *(fixed in `96c1bf5`)* | the narrator reuses that text in its own prose, so a run whose planting turn was 65 depths outside the window read `precondition_not_met` | the precondition is now the planting turn's position, recorded at planting and carried across `--resume`, which is M04's own wording | ### Discarded evidence runs @@ -1184,6 +1383,17 @@ reported: entirely on the final tree. 5. **Two harness shakeout runs** (6 and 14 turns) — never reported as evidence; their purpose was to find the two SSE defects above. +6. **The first 100-turn release campaign**, which reached 97 of 100 turns and was + lost with its evidence to a host crash, because it wrote under `/tmp`. The + first revision's §G figures came from it and can no longer be checked. They + are replaced, not repeated. +7. **The memory-off 100-turn run** (`ef25b0a`). It is complete and internally + consistent, but it never exercised summary/memory activation. It is kept as the + CPU-host timing and headroom record (§N). +8. **The 100-turn runs on `f8d4010` and `0c7316f`.** They are complete, and they + are superseded because their M04 evidence was contaminated by protocol leaks + (§O.8). The `0c7316f` run's reclassified verdict is kept beside its original + and is not reported as the result. --- ## P. Residual risks @@ -1191,53 +1401,91 @@ reported: Genuine remaining risk and debt only. There is no M12; everything below is either accepted for v1, or a decision for the owner at acceptance. -1. **The reference deployment is CPU-only, and the release evidence carries its - speed.** The 100-turn campaign averaged around a minute a turn against a 3B - model on a CPU-only LAN host. Nothing in the planning package sets a - performance requirement, and none is invented here — but a reviewer should - read §N's timings as *this machine's*, not as a product characteristic. +1. **The release evidence carries two hosts' speeds.** The CPU reference host took + minutes a turn, and the GPU host took seconds. Nothing in the planning package + sets a performance requirement, and none is invented here. Read §N's timings + as these machines', not as a product characteristic. -2. **The narrator is a 3B model.** Every realistic-model observation — state - extraction quality, narration length adherence, identity handling — is that - model's. A stronger local model would behave differently, probably better, and - the application's guarantees are deliberately independent of which: what is - asserted is that the application stays correct whatever the model proposes. +2. **The narrator is a 3B model.** Every realistic-model observation is that + model's: state extraction quality, narration length adherence, identity + handling, and what it restates. A stronger local model would behave + differently, probably better. The application's guarantees are deliberately + independent of which, since what is asserted is that the application stays + correct whatever the model proposes. -3. **Post-M8 finding D's root cause is unestablished and will stay that way.** +3. **The 16k window is nearly full.** By the narrator's own tokenizer the largest + prompts plus the reply reserve left 23, 34 and 42 tokens of headroom in the + three GPU runs (§N). The application's `cl100k_base` count ran 16 tokens low + on every prompt measured. The inference server cuts an over-window prompt to + 8,194 tokens with no error. No prompt overflowed. A model whose tokenizer + diverges further from `cl100k_base`, or a larger reply reserve, could overflow + without anyone noticing. A deployment that wants margin can lower + `context_token_budget` below the window. + +4. **The narrator restates prompt text in its prose.** On 74 of 104 turns of the + evidence run it wrote the state's fact line as a sentence, and it also echoed + phrases such as "Scene set, continue your adventure." and lines opening + "Memory:". These sit inside story, so the extractor leaves them. They cost + narration quality, and they are the route by which M04's fact reached memory. + +5. **Memory does not keep a planted fact on its own with this summariser.** In no + run did a memory summarising the planting era carry the fact. Recall rests on + authoritative state, which the repository owner accepted for M04 on + 2026-09-13. A reviewer who reads F02 or M04 as requiring memory retention in + its own right should read them as unproven. + +6. **Model-invented headings are not removed.** A model that writes its own + section, such as `Identifiers established:` or `Set of events made true:`, + under a heading that is not the renderer's, keeps it in the story. Seen on 10 + turns of one 26-turn trial. + +7. **The GPU inference host dropped its GPU after the evidence run** (§E.1). The + cause is not established, and power transients at the uncapped 280 W limit + are the leading candidate. The evidence is unaffected. Future long runs must + log power, link state and kernel messages (`DEVELOPMENT.md`). + +8. **Post-M8 finding D's root cause is unestablished and will stay that way.** The campaign that produced it was destroyed. M11 delivers a diagnostic that - can classify the next occurrence and the detection the finding asked for. On - the reference narrator the objective checks were clean; §G.4 says what that - does and does not mean. + can classify the next occurrence, and the detection the finding asked for. + **The diagnostic's run results are not in this report.** The section the + first revision pointed to was left empty, and the evidence is in the + implementer's `m11-evidence/identity-recheck` directory. -4. **Two control-boundary colour pairs are below WCAG 1.4.11** (1.33:1 resting, - 1.75:1 hover). Reported rather than fixed, because the control is identified - by its label — measured at 5.48:1 — and restyling the palette inside a - release-validation milestone would be the wrong kind of change. An owner who - disagrees has the measurement. +9. **The browser, offline and identity runs predate the last three product + commits.** They were re-run on the `ef25b0a` tree: that commit's message + records browser 38/0/0, offline 23/0 and the identity diagnostic clean, and the + evidence is dated 2026-09-10. `f8d4010`, `0c7316f` and `96c1bf5` change the + turn commit, memory retrieval, narration extraction and the state renderer's + headings, all of them backend. The backend and frontend suites cover those + changes; the black-box runs were not repeated. -5. **A file cannot be driven *out* of this headless snap Firefox.** Export is - proved end-to-end without a browser; the browser's own export control is - exercised only as far as the click. Import is now fully proved (M9's residual - risk 6 was half wrong, and the half that was right remains). +10. **Two control-boundary colour pairs are below WCAG 1.4.11** (1.33:1 resting, + 1.75:1 hover). Reported rather than fixed, because the control is identified + by its label, measured at 5.48:1, and restyling the palette inside a + release-validation milestone would be the wrong kind of change. An owner who + disagrees has the measurement. -6. **The bundle ceiling is unchanged** — M9 measured about 279 turns against the - 20 MB import limit, and §N measures where the real 100-turn campaign sits - against it. Beyond that ceiling a campaign can still be exported and would be - refused on import, which is the asymmetry worth knowing. +11. **A file cannot be driven *out* of this headless snap Firefox.** Export is + proved end to end without a browser; the browser's own export control is + exercised only as far as the click. Import is fully proved. -7. **The media seam has no real adapter.** M10's own residual risk, unchanged: - the contracts are shaped by the contract document rather than by an adapter - that had to work. The first real provider may want the packet reshaped, and - nothing in the story depends on its shape. +12. **The bundle ceiling is unchanged.** M9 measured about 279 turns against the + 20 MB import limit, and §K puts the real 100-turn campaign at 13.1% of it. + Beyond the ceiling a campaign can still be exported and would be refused on + import, which is the asymmetry worth knowing. -8. **`quick_check` rather than `integrity_check` on a backup**, and **no - scheduled backup** — both M9's, both unchanged, both outside the acceptance - contract. +13. **The media seam has no real adapter.** This is M10's own residual risk, + unchanged: the contracts are shaped by the contract document rather than by + an adapter that had to work. -9. **A campaign with no narration-length choice keeps the pre-M11 hint.** That is - deliberate — an empty value means the reader never chose — but it means an - existing campaign does not benefit from finding C's fix until someone sets the - control. +14. **`quick_check` rather than `integrity_check` on a backup**, and **no + scheduled backup**. Both are M9's, both unchanged, both outside the acceptance + contract. + +15. **A campaign with no narration-length choice keeps the pre-M11 hint.** That is + deliberate, because an empty value means the reader never chose. It means an + existing campaign does not benefit from finding C's fix until someone sets + the control. --- @@ -1257,6 +1505,12 @@ accepted for v1, or a decision for the owner at acceptance. | `DEVELOPMENT.md` | The context-window section rewritten around what the application now does; a new section on the six release harnesses. | developer docs | | `planning/reports/M11-IMPLEMENTATION-REPORT.md` | New. | milestone report | | `planning/archive/milestone-reports/` | M9's and M10's reports moved here. | rotation | +| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Revised 2026-09-14**: M01-M04 on the complete evidence run; §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R rewritten; the §G.0 addendum removed. | milestone report | +| `DEVELOPMENT.md` | The pointer to the removed §G.0 replaced; a new section on logging a GPU inference host during a long run. | developer docs | + +**Not revised with this report:** `planning/V1-ACCEPTANCE-TESTS.md`, +`planning/BUILD-MILESTONES.md`, `planning/VERSION.md` and `planning/README.md` +still carry what the first revision and `ef25b0a` wrote into them. **Implementation facts added:** §15.2, §28B, §8A's implementation note, §42B, and the M11 status block. Each records what the code does; none changes what is @@ -1275,98 +1529,95 @@ pinned by a test so a later milestone changes it deliberately. **1. Does every REQUIRED FOR V1 acceptance test pass?** -**No — one is outstanding.** 84 of the 85 REQUIRED tests pass, with H09 recorded -NOT APPLICABLE on the condition its own text states. **M01 is PARTIAL**: 41 -accepted turns of the required 100 at the time of writing, still running. Every -other REQUIRED test passes on the evidence in §F. +**Yes.** All 85 pass, with H09 recorded NOT APPLICABLE on the condition its own +text states. M01 to M04 pass on the evidence run in §G, and every other REQUIRED +test passes on the evidence in §F. **2. Does M01 pass with 100+ accepted turns?** -**Not yet.** 41 accepted turns, correct in every respect measured, with the run -continuing unattended. The obstacle is wall-clock on a CPU-only inference host — -88-203 seconds per turn at the default window, 229-291 at the recommended one — -not any behaviour of the application. §G gives the exact position and what -remains unexercised inside the campaign; the harness is committed so the run can -be completed and re-checked before acceptance. +**Yes.** 101 accepted turns on commit `96c1bf5`, with three genuine process +restarts, all thirteen scheduled history operations, an induced failed call that +corrupted nothing, zero failed post-turn passes, 33 memories and 12 summaries, +and recovery onto a clean data directory 16 of 16. §G.1-§G.5. **3. Does actual model context capacity match the application's release assumptions?** -**Yes, and it is now enforced rather than assumed.** The reference server gives -the plain model 4,096 tokens and the `num_ctx`-baked model 16,384; the -application discovers both correctly and caps its budget to whichever it finds. -Across 41 consecutive real turns the window was verified every time and the -budget was capped from 16,384 to 4,096 every time, with the campaign canon -present in all 41 prompts. Where the window cannot be checked the turn proceeds -and is recorded as unverified. §E. +**Yes, it is enforced rather than assumed, and the margin is thin.** The window +was verified on 101 of 101 turns at 16,384, and the application caps its budget to +whatever the server reports (§E). By the narrator's own tokenizer the largest +prompt plus reserve is 16,342 of 16,384. It fits, with 23-42 tokens of headroom +across the GPU runs (§N, §P). **4. Does offline operation pass from a fresh state with no Internet route?** **Yes.** 23 of 23 checks in a container with `--network none` and a fresh volume: no route, no DNS, first page load, every asset local, campaign creation, state -extraction, knowledge import and retrieval, prompt assembly, export, import, +extraction, knowledge import and retrieval, prompt assembly, export, import, the media module inert, and campaigns surviving a container restart. A turn with no -model reachable is reported and corrupts nothing. §J. +model reachable is reported and corrupts nothing. §J. Re-run on the `ef25b0a` +tree (§P). **5. Does branch/memory/summary/scene isolation pass?** **Yes.** All four in one long campaign, each with a positive control, including a summary **regenerated after the divergence** and M10's derived Scene Packet. §H. +The evidence run also restored a Save Point across a live memory bank. §G.2. **6. Does trusted-LAN HTTPS inference still pass?** -**Yes.** Every real turn in this report — the campaigns, the browser run, the -identity diagnostic — ran against Ollama on a separate physical machine over -HTTPS with a private CA in this machine's OS trust store, verification on, no -bypass, with the storyteller itself loopback-bound. H12's address rules, -including the database-tampering case, pass in the suite. §F, §J. +**Yes, on the CPU reference host's evidence.** The browser run, the identity +diagnostic and the CPU-host campaigns ran against Ollama on a separate physical +machine over HTTPS, with a private CA in the application machine's OS trust store, +verification on and no bypass. The storyteller stayed loopback-bound. **The final +long runs used plain HTTP to a LAN GPU host.** H12 permits that, and it is not A06 +evidence. §E.1, §F, §J. **7. Does export/import/recovery pass?** -**Yes.** 16 of 16 checks moving the real long-run campaign into a data directory -that never existed, plus M9's own recovery suite (173 tests) and the migration -parity suite (9 tests). §K. +**Yes.** 16 of 16 checks moving the evidence run's campaign (207 actions, 88 of +them retained history, 2.74 MB) into a data directory that never existed, plus M9's +own recovery suite and the migration parity suite. §K. **8. Do both fantasy and science-fiction fixtures pass?** -**Yes.** The fantasy Continuity Test is the long-run campaign; the Persephone -fixture passes 10 checks including five entity types in one document, hard +**Yes.** The fantasy Continuity Test is the long-run campaign. The Persephone +fixture passes 10 checks, including five entity types in one document, hard canon, possession, reference retrieval, a starship in M10's Scene Packet, and a whole-vocabulary check that no event type names a genre noun. No schema or code change was required for either. §I. **9. Do fresh-install and upgrade schemas agree?** -**Yes**, compared field by field — tables, columns with type and nullability, +**Yes**, compared field by field: tables, columns with type and nullability, indexes with columns and uniqueness, foreign keys, primary keys and the version -stamp. This is M10's accidental discovery made into a permanent, general -regression. §K. +stamp. §K. **10. Does the browser pass the release workflow?** -**Yes.** 38 checks, 0 failed, 0 skipped, in Firefox 154.0.1 against the built -SPA served by FastAPI: play, history controls, the position indicator, hostile -Markdown, `javascript:` URLs, remote images, hidden knowledge absent from the -DOM, context inspection, CORS/404 behaviour, CSP, and the accessibility -measurements M8 deferred to M11. §L. +**Yes, on the `ef25b0a` tree.** 38 checks, 0 failed, 0 skipped, in Firefox +154.0.1 against the built SPA served by FastAPI. The commits after it change no +frontend file, and the browser run was not repeated. §L, §P. **11. Are there any unresolved blockers to independent v1 acceptance?** -**One, and it is a matter of wall clock rather than of correctness: M01's -remaining 59 turns.** Nothing else in the acceptance contract is outstanding, no -product defect is known and unfixed, and no requirement was weakened. A reviewer -can either accept M01 on the 41-turn evidence plus the automated coverage of the -operations that fall later in the schedule, or — the honest recommendation — run -`tools.m11_long_run --turns 100` to completion on a faster host and check -`summary.json` and `recall.json` before signing. +**No known blocker.** No product defect is known and unfixed, and no requirement +was weakened. A reviewer should weigh five things before signing: + +- M04's recovery ran through authoritative state, not memory retention (§G.4) +- the thin real-token headroom (§N) +- the narrator restating prompt text (§P) +- the identity diagnostic's results being absent from this report (§P) +- the black-box runs predating the later commits (§P) **12. Is the tree safe to commit as the M11 release candidate?** -**Yes.** The full backend suite is green (1,291 passed, 17 skipped, 0 failed), -the frontend suite is green (157 passed), lint exits 0, the production build is -clean, the Docker image builds `--no-cache` and runs, and every staged file is -intended M11 content — no databases, logs, caches, secrets or evidence captures. -The commit is the owner's to sign; M11 created none, and there is no release tag. +**It is committed.** Seven signed commits follow the M10 base, the newest +`96c1bf5`, and there is no release tag. On that tree the backend suite passed +1,421 with 17 skipped and 0 failed (2026-09-13). The frontend suite passed 161 of +161, and lint exited 0 with warnings only (2026-09-14). The production build and +the Docker image were built for the first revision and not rebuilt since. The +remaining commit, this report's revision, is the owner's to sign. ---