# M11 — v1 Security, Long-Run, and Release Validation **Implementation and verification report, written for independent release review.** Branch `m11-release-validation`, from the signed M10 commit `1013c94`. Implemented and verified 2026-09-07. The long-run evidence was re-established, and this revision written, on 2026-09-14. This is the evidence package; it is not a record of acceptance, and nothing in it says M11 is accepted. --- ## A. Executive result **PASS, for independent review.** All 85 tests marked REQUIRED FOR V1 pass, with H09 recorded NOT APPLICABLE on the condition its own text states. This revision supersedes the first. That revision left M01 PARTIAL at 41 accepted turns, and then the run it described was lost to a host crash along with its evidence. **M01 to M04 now pass on a complete 100-turn campaign**, run by the committed harness on commit `96c1bf5`: 101 accepted turns, three genuine process restarts, every scheduled history operation performed, zero failed post-turn passes, and recovery onto a clean data directory 16 of 16 (§G, §K). Getting there took six further runs. They found **two more product defects**, both fixed and both invisible to every earlier piece of evidence because the memory bank had never been switched on during a long run: 1. **A turn locked its own memory bank out** (§O.7). Retrieval wrote a use counter before the model call, and the turn committed only after the reply, so SQLite's single write lock was held for the whole reply. Post-turn memory and summary writes timed out behind it, and recording those failures timed out too. A 26-turn run reported "complete" with two memories, no summary and 180 `database is locked` errors, while derived status read `idle`. 2. **The narrator's protocol leaked into stored story** (§O.8). A small model pasted the narrative-state section and unfenced proposals into its prose, on 42 of 104 turns in one run, and stored text is replayed as history. The extractor now removes every shape observed. Replaying 443 real turns from five runs changed no turn the old extractor had left clean. The first revision's three corrections stand: the context-window mismatch is resolved (§E), a partly refused manual correction now says so, and the narration-length setting moves a number. Post-M8 findings A and B are closed. **What is not claimed.** - **That memory keeps a planted fact on its own.** M04 passes because the fact was recoverable with its planting turn 53 depths outside the history window. The recovery path ran through authoritative state: the narrator restated the state's fact line in its prose, and memory summarised those restatements (§G.4). No memory of the planting era carried the fact in any run. The repository owner accepted state-based recovery for M04 on 2026-09-13. - **Headroom in the context window.** The largest prompts fill 16,342 of 16,384 tokens by the narrator's own count, and the inference server cuts an over-window prompt without an error (§N, §P). - **Performance.** Timings are two particular hosts', not a product characteristic (§E.1). - **Post-M8 finding D's root cause**, which remains unestablished. The identity diagnostic's run results are not in this report (§D.1). - **That the browser, offline and identity runs cover the later commits.** They were re-run on the `ef25b0a` tree. The three product commits since then change the turn commit, memory retrieval and narration extraction (§B). **Zero requirement weakenings.** No acceptance test was retired, relaxed or reclassified. The M04 precondition the harness measures was corrected to the acceptance text's own wording, "without entire transcript in prompt" (§O). --- ## B. Repository and provenance | | | | --- | --- | | **Base commit** | `1013c94eb1ad283e960114aef04c19c2806b5db7` — *"M10: the seam for media, and no media"* | | **Signature** | `git verify-commit 1013c94` → **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, "JesseMarkowitz", trust `[ultimate]`. `%G?` = `G`. | | **Branch** | `m11-release-validation`, created from that commit. The working tree was clean at the start (`git status --porcelain` empty). | | **HEAD at this revision** | `96c1bf5`. Seven commits follow the base, all signed by the owner (`%G?` = `G`): `144406c` (M11), `fedb714`, `ef25b0a`, `fec46f6`, `f8d4010`, `0c7316f`, `96c1bf5`. This revision of the report is staged for the owner to sign. | | **Upstream ancestry** | `git merge-base --is-ancestor d72f7c1b HEAD` → true. The AI-DnD fork point is still an ancestor. | | **LICENSE** | Unchanged — md5 `07fde30437134836e2ee875e82a7cd31`, MIT, "Copyright (c) 2026 Parth Thakkar". `PROVENANCE.md` unchanged. | **The four trees, kept distinct**, because M11's evidence discipline depends on which one produced a given number: ```text committed base 1013c94, signed by the owner: M1-M10 as accepted working tree the base plus M11's changes; this is what was tested staged tree identical to the working tree (§29 lists it) frozen tree the working tree at the point each black-box run started, with no source edit during any reported run ``` The black-box runs come from two trees. **The long-run evidence**, the 100-turn campaign and its recovery run, was taken on commit `96c1bf5`, after the last product change, with no source edit during the run. **The browser regression, the offline container, the identity diagnostic and the contrast audit** were re-run on the `ef25b0a` tree, as that commit's message records, and their evidence is dated 2026-09-10. The three product commits after it (`f8d4010`, `0c7316f`, `96c1bf5`) change backend files only, and those runs were not repeated (§P). Runs taken before a product change on their own tree were discarded and re-run; §O records which and why. --- ## C. Change inventory **Product code — backend** | File | Change | | --- | --- | | `app/contextwindow.py` | **New.** Discovers the server's real context window and provides the ceiling. The whole of §E. | | `app/context/builder.py` | Takes a `window`; the effective budget is `min(configured, verified)`; the context report carries what was verified and how; `length_hint` gains the campaign's length band. | | `app/routers/adventures/turns.py` | Probes the window before assembling a turn, and stores the verdict in the turn's provenance. | | `app/routers/adventures/insights.py` | The same probe, so the inspector shows the prompt the next turn will actually send. | | `app/routers/settings.py` | The connection test reports the window, or why it could not be checked; changing endpoint or model clears what was learned. | | `app/routers/adventures/state.py` | A correction reports refusals (`refused`) and shared display names. | | `app/routers/adventures/crud.py` | Persists the campaign's narration-length choice. | | `app/narrative/model.py` | **New** `duplicate_names` — detection, deliberately not a refusal. | | `app/models.py`, `app/migrations.py`, `app/schemas.py`, `app/bundle.py` | `adventures.narration_length`: column, migration 93, API in and out, bundle carriage with an unknown value dropped. | **Product code — frontend** | File | Change | | --- | --- | | `src/documentTitle.js` | **New.** The product's name, in one place; the tab follows the open campaign. | | `index.html`, `src/pages/Play/index.jsx` | The inherited `AI D&D` title replaced and kept in step (finding A). | | `src/pages/Play/Composer.jsx`, `styles/story.css` | The position indicator: `Moment 11 · later story ahead` (finding B, §8A). | | `src/pages/NewCampaign.jsx`, `panels/CampaignSettingsPanel.jsx` | Narration length sent and editable as data (finding C). | | `src/pages/Settings.jsx`, `styles/library.css`, `styles/tokens.css` | The context-window report and warning; a `--warning` token measured at 7.85:1. | | `panels/StatePanel.jsx`, `styles/insights.css` | Refused corrections and shared names, shown to the reader. | **Release-test infrastructure — no application code imports any of it** | File | Purpose | | --- | --- | | `tools/m11_long_run.py` | The 100-turn campaign: M01-M04. | | `tools/m11_recovery.py` | That campaign, moved to a clean data directory: I01-I07. | | `tools/m11_browser.py`, `tools/m11_webdriver.py` | The browser regression and accessibility measurements; a dependency-free W3C WebDriver client. | | `tools/m11_offline.py` | A container with no network: §18 and the packaging path. | | `tools/m11_identity.py` | The multi-character identity diagnostic (finding D). | | `tools/contrast_audit.py` | The palette against WCAG AA. | | `tests/test_m11_context_window.py` | 20 tests: the probe, the cap, the truncation sentinel. | | `tests/test_m11_leakage.py` | 14 tests: E01-E04 together, in one long campaign. | | `tests/test_m11_scifi.py` | 10 tests: J01-J03, the Persephone fixture. | | `tests/test_m11_migration.py` | 9 tests: fresh-versus-upgraded schema parity, and the upgrade. | | `tests/test_m11_security.py` | 25 tests: the H-series against the assembled product. | | `tests/test_m11_findings.py` | 20 tests: findings C and D, and the refused-correction regression. | | `tests/test_m11_real_window.py` | 3 tests: the probe against a real Ollama (skipped without one). | | `tests/schema_rewind.py` | Migration 93's inverse, so the suite can replay it. | **Dependencies: none added, none removed, none upgraded.** `requirements.txt`, `requirements.lock` and `package.json` are byte-identical to M10's. **Changes after the first revision.** `144406c` is the first revision itself. Every later commit is listed here, and all are signed by the owner. | Commit | Change | Files | | --- | --- | --- | | `fedb714` | Release evidence must be written to a durable `--out`, and the report recorded the lost run | `DEVELOPMENT.md`, this report | | `ef25b0a` | The history window trims in blocks, so the server's prompt cache survives. `settings.context_window_override` (**migration 94**) covers a server the probe cannot ask. The long run gains `--resume`, a configurable turn timeout and an abort after consecutive failures. Two checks that could not fail were fixed. Every black-box harness was re-run on this tree | `app/context/builder.py`, `app/contextwindow.py`, `app/models.py`, `app/migrations.py`, `app/schemas.py`, `app/routers/settings.py`, `app/routers/adventures/turns.py`, `app/routers/adventures/insights.py`, `frontend/src/pages/Settings.jsx`, `tools/m11_long_run.py`, `tools/m11_browser.py`; tests `test_history_block_trim.py`, `test_m11_declared_window.py`, `test_m11_long_run_resume.py`, `settingsWindow.test.jsx`; planning package v3.8 | | `fec46f6` | The long run switches the memory bank and auto-summarise on, and refuses a run with no embedding model | `tools/m11_long_run.py`, `tests/test_m11_long_run_memory.py` | | `f8d4010` | §O.7: no write lock is held across the model call, and a failure that breaks the session is still recorded. The harness reads the real section labels and stops on failed post-turn work | `app/memorybank.py`, `app/routers/adventures/turns.py`, `app/routers/adventures/insights.py`, `tools/m11_long_run.py`, five test files | | `0c7316f` | §O.8: protocol is kept out of stored narration, and the fence-label bug is fixed. The harness gains an M04 verdict and a protocol-leak count | `app/narrative/extract.py`, `app/narrative/render.py`, `tools/m11_long_run.py`, `tests/test_narrative_state.py`, `tests/test_m11_long_run_memory.py` | | `96c1bf5` | §O.8 extended to a model's own sections under decorated headings. M04's precondition is now positional | `app/narrative/extract.py`, `tools/m11_long_run.py`, `tests/test_narrative_state.py`, `tests/test_m11_long_run_memory.py` | No later commit touches `requirements.txt`, `requirements.lock` or `package.json`. --- ## D. M10 and M8 handoff Every item handed to M11, and what happened to it. Nothing here is closed by "the automated tests are green". ### From M10's §O.1 (its own residual risk) | Item | Disposition | | --- | --- | | **K04 satisfied structurally, not physically** — no `media_jobs`/`media_assets` | **Unchanged, and verified as such.** M11 exercised the deferred-table path rather than building the tables: a dummy provider, a real packet, a fake asset carrying the packet's `scene_id`, and the story model byte-identical afterwards. Reported as K04 = PASS on the acceptance text's own deferred branch (§F). | | **`ambience` is an empty shape** | **Unchanged.** Filling it means extending `set_scene`, which is a prompt-path change and not release validation. Still an empty shape with the right fields. | | **The seam has no consumer, so it is unexercised by real use** | **Partly answered.** The Persephone fixture put a third genre through the packet and a visual profile through a starship hull (`test_m11_scifi.py`), and the offline container proved the media module imports and stays inert with no network. Still no real adapter, and that remains true until someone writes one. | | **A visual profile cannot be recovered from within the app** | **Unchanged.** It travels in the bundle; there is no undo for deleting one. | | **Profiles are API-only, with no reader-facing surface** | **Unchanged, deliberately.** M11 is release validation; adding a media UI would be new feature work. | ### From M9's residual risks | Item | Disposition | | --- | --- | | **Bundle ceiling ~279 turns** | **Measured against the real 100-turn campaign** rather than the fixture — §N gives the actual bundle size and what fraction of the 20 MB import limit it is. | | **`quick_check` rather than `integrity_check`** | Unchanged; re-exercised on a migrated database (`test_m11_migration.py`). | | **No scheduled backup** | Unchanged, and outside the acceptance contract. | | **Stale `chunk_id` in a restored snapshot** | Unchanged; a React key, not a live pointer. | | **The importing machine's context window may differ** | **Closed.** This was the same defect as M8's, seen from the import side, and §E closes both: the application caps to what the destination server accepts and says when it could not check. `DEVELOPMENT.md`'s import warning is rewritten accordingly. | | **This machine cannot drive a file into or out of the browser** | **Narrower than recorded, and half of it is closed.** The snap Firefox refuses a WebDriver file path under `/tmp`; a path under `$HOME` works. Knowledge import is now proved end-to-end in a real browser (§L). The *download* half — a `blob:` export leaving the browser — is still not driveable here and is still recorded as a limitation. | ### From M8 (via M9 and M10) | Item | Disposition | | --- | --- | | **The deployment context ceiling** | **Closed** — §E. This was M11's stated release blocker. | | **Contrast and visible focus checked by eye** | **Measured** — §L. Palette pairs by calculation (`tools/contrast_audit.py`), rendered colours in a real browser, plus focus visibility, accessible names, tab order, hover-only controls and modal focus. | | **The four post-M8 playtest findings** | **All four disposed of** — §D.1 below. | ### D.1 The four post-M8 playtest findings **A — the tab read `AI D&D`.** Fixed. The name is `Interactive Story`, chosen by the repository owner when M11 asked, and deliberately not "Adventure Storyteller": `SPECIFICATION.md` requires a genre-agnostic engine and *Adventure* is narrower than the thing it names. The tab shows the open campaign first (`Westhaven — Interactive Story`), and one module owns the string. **B — no orientation after Undo.** Fixed, to `BROWSER-UX-SPEC.md` §8A. A status line at the end of the control row reads `Moment 11 · later story ahead`. The number and the clause are both the *server's* answers, it changes visibly after Undo, and it uses no implementation vocabulary. Asserted in the component suite and in the browser regression, the second of which is what §8A demanded when it said the requirement must be observable rather than inferable. **C — narration length had no measurable effect.** Fixed, at the mechanism the finding identified. The choice is now data on the campaign, and `length_hint` turns it into a real word band (70-180 / 150-380 / 320-700), bounded by the reply cap. The generation budget is deliberately **not** touched: capping it per length would make a brief turn likelier to be cut off mid-sentence, and the state block is emitted last, so the first thing a truncated reply loses is the turn's state. Before M11 all three settings produced *the identical sentence*; `test_the_three_lengths_no_longer_say_the_same_thing` fails against that. **D — character identity confusion.** The diagnostic exists; the root cause does not, and cannot. **The diagnostic's run results are not reproduced in this report.** The first revision pointed here to a section that was left empty. What is recorded is the defect the diagnostic caught in its own fixture (§O.2), which is why its first run's evidence was discarded. The later run's evidence is in the implementer's `m11-evidence/identity-recheck` directory (§P). --- ## E. The context-window resolution ### The original mismatch M8 measured the reference deployment enforcing **4,096** input tokens while the application budgeted **16,384**. M11 re-measured it, on the same server, before changing anything: ```text $ curl -sk https:///api/show -d '{"model":"qwen2.5:3b-instruct"}' model_info["qwen2.context_length"] = 32768 the architecture's ceiling parameters = (none) no num_ctx is baked in $ (one /v1/chat/completions call to load it, then /api/ps) qwen2.5:3b-instruct context_length = 4096 size_vram = 0 ``` So the mismatch was live on the reference deployment on the day M11 started, and `size_vram = 0` is why: with no VRAM Ollama picks a 4,096 default. ### Root cause Two independent facts that only bite together. 1. **Ollama's window is a property of how the model was loaded**, not of the request. Its OpenAI-compatible endpoint accepts `num_ctx` — nested or top-level — returns 200, and ignores it; M8 established that, and it is why the operational fix is a model with `num_ctx` baked in or `OLLAMA_CONTEXT_LENGTH` on the server. 2. **The application had no way to know.** `Settings.context_token_budget` was the only number in play, so the builder assembled to it and the server quietly did what it liked with the excess — which is drop the **oldest** tokens. The oldest tokens here are the system block: the narrator's rules and the campaign canon. The failure therefore looks like a narrator that stops respecting canon deep into a long session, with nothing on screen to explain it, and every acceptance test that reads a 200 as success passing throughout. ### The fix `backend/app/contextwindow.py`, plus four call sites. The rule: > A **verified** window is a ceiling on the configured budget. An **unverified** > one leaves the budget standing and is recorded as unverified. There is no > third behaviour. - **Discovery** asks the server the application is already talking to, on the native path beside `/v1`, through `endpoints.rejection_reason` and the shared TLS trust store — so it can reach exactly what a turn can reach and nothing more. `/api/ps` gives the window a resident model is *actually* being served with; `/api/show` gives the `num_ctx` an unloaded one will load with, capped by the architecture's own ceiling. - **Enforcement** is one line in the builder: the budget every section is priced against is `min(configured, verified)`. Because the output reserve is subtracted from that budget by the existing arithmetic, the assembled prompt plus the reply reserve fits inside the window by construction. - **Reporting.** The context report and the turn's stored snapshot carry `window: {verified, tokens, source, model_max, detail, capped}`, so an old turn can be asked afterwards whether it was built against a checked window. The connection test in Settings shows the number or explains why it could not be checked, in three distinct messages, because "could not check", "smaller than your budget" and "fine" need three different things done about them. - **What it does not do:** hard-code 4,096 (right on one machine, wrong on the next), raise anyone's window, guess from a model's name, or add a provider abstraction. An unknown window is reported as unknown. ### The numbers, on the reference deployment | | | | --- | --- | | Ollama | 0.33.0, CPU-only (`size_vram = 0`) | | `qwen2.5:3b-instruct` | **4,096** tokens, source `loaded` (`/api/ps`) | | `qwen2.5:3b-instruct-16k` | **16,384** tokens, source `parameters` (`num_ctx` baked in), architecture ceiling 32,768 | | Application budget (M01) | `context_token_budget` = 16,384, `max_output_tokens` = 500 | | M01's effective budget | **16,384** — verified equal to the server's window, so nothing was capped away | | Largest assembled prompt in M01 | see §N | Measured end to end in `tests/test_m11_real_window.py` against the real server: ```text turn 1 window verified=False (the model is not resident yet) turn 2 window verified=True tokens=4096 source=loaded budget configured=16384 effective=4096 prompt 850 tokens + 464 reserved -> 1314 <= 4096 ``` That is the whole behaviour in six lines: the first turn on a cold model cannot verify and says so; from the second turn the cap is live and the prompt provably fits. ### Why silent truncation can no longer invalidate M01 Three separate reasons, and the third is the one that matters: 1. **M01 ran on the 16k model**, whose window (16,384) equals the application's budget, so nothing was capped and nothing was near the edge — recorded per turn in `timeline.jsonl` as `window_verified` and `window_tokens`. 2. **Every M01 turn recorded the verification**, so the claim is per-turn evidence rather than a statement about the configuration at the start. 3. **The failure mode is now impossible to reach silently.** If the window were smaller than the budget, the prompt would be built to the *window*, and the canon at the front would survive by construction — `test_the_canon_at_the_front_survives_a_window_far_too_small` plays 120 turns into a 4,096-token window and finds the canon sentinel still present with the *oldest history* dropped instead. Its companion, `test_without_the_cap_the_same_prompt_would_have_overflowed`, builds the same campaign with no verified window and measures a prompt more than twice the size — the defect, reproduced, so the fix is shown to be doing something. --- ## E.1 The release environment Recorded because a timeout without hardware beside it is not a measurement. Two inference hosts produced this report's evidence, and every result below says which. ### The application machine It runs the storyteller, every harness, the browser and the container. | | | | --- | --- | | **OS** | Ubuntu 24.04.4 LTS, kernel 7.0.0-30-generic, at the first revision; Ubuntu 24.04.5 LTS, kernel 7.0.0-31-generic, for the 2026-09-13 long runs | | **CPU** | AMD Ryzen 7 7840HS, 4 cores available to this VM, 1 thread per core | | **GPU** | **none**: VMware SVGA II | | **RAM** | 15 GiB | | **Python** | 3.12.3 | | **Node / npm** | v22.23.1 / 10.9.8 | | **Docker** | 29.7.2 | | **Browser** | Firefox 154.0.1, headless, via geckodriver over W3C WebDriver | ### The CPU reference host Used from 2026-09-07 to 2026-09-10: the lost run, the memory-off 100-turn run, the browser run and the identity diagnostic. | | | | --- | --- | | **Machine** | a separate physical machine on the trusted LAN, no GPU (`size_vram = 0`) | | **Ollama** | 0.33.0, **HTTPS** with a private CA installed in the application machine's OS trust store | | **Measured cost** | 88-203 s per turn at a 4,096 window; 229-291 s at 16,384 | ### The GPU inference host Used on 2026-09-13 and 2026-09-14: both 26-turn trials and the three 100-turn runs, including the evidence run. | | | | --- | --- | | **Machine** | a separate physical machine on the trusted LAN; Ubuntu 24.04.2 LTS; 16 logical CPUs; 14 GiB RAM | | **GPU** | NVIDIA GeForce GTX 1080 Ti, 11 GiB (Pascal, compute capability 6.1), driver 580.173.02, in an **OCuLink dock**: PCIe x4, measured at Gen 3 under load and Gen 1 at idle | | **Ollama** | 0.34.0. Its CUDA 13 runner skips a compute-6.1 card and the CUDA 12 runner serves it; both models loaded 100% into VRAM | | **Transport** | **plain HTTP** on the LAN. H12 permits a LAN address; this host is **not** A06 evidence | | **Power limit** | the card's default, 280 W, during every run | | **Measured cost** | prefill ~1,350-1,850 tokens/s and generation ~87 tokens/s with synthetic prompts; 4.1-22.8 s per turn in the evidence run | ### Models and budget | | | | --- | --- | | **Narrator** | `qwen2.5:3b-instruct-16k`: `num_ctx` 16,384 baked in, architecture ceiling 32,768. Digest `21ff8cc52f37`, identical on both hosts | | **Base narrator** | `qwen2.5:3b-instruct`, digest `357c53fb659c`, identical on both hosts. With no `num_ctx` the server gives it **4,096** | | **State/summariser model** | the narrator; no separate summariser is configured | | **Embedding model** | `nomic-embed-text`, digest `0a109f422b47`, identical on both hosts | | **Application context budget** | `context_token_budget` = 16,384, `max_output_tokens` = 500, reply reserve 564 | | **Effective budget** | 4,096 in the lost run, where it was capped every turn; **16,384** in every later run, where window and budget were equal | **Why the first campaign ran at 4,096.** On the CPU host the recommended 16,384 window cost 229-291 s a turn, about eight hours for a hundred. The first revision therefore ran the default window, which doubled as the longest test of §E's cap. That run reached 97 of 100 turns and was lost (§G.6). Every later campaign runs the recommended window. ### The GPU host incident About 30 seconds after the evidence run's last post-turn pass finished, the GPU dropped off the PCIe bus: `NVRM: Xid 79, GPU has fallen off the bus`, then `Xid 154`, reset required. Ollama stayed up without a GPU, stopped completing requests, and the host had to be rebooted. The same host had logged one `Xid 13` graphics exception from the inference server earlier that day. There was no out-of-memory event and no thermal event recorded. The firmware does not support PCIe AER, so no link-error trail exists. After the reboot, under a temporary 200 W cap with logging running, the link held Gen 3 x4 under load, peak draw was 207.8 W, and there was no fault. **The cause is not established.** Power transients at the uncapped 280 W over a long run are the leading candidate. **The evidence is unaffected**: every record the result rests on was written before the drop (§G.1). **Required for every future long run on a GPU host:** log the card's power, temperature, utilisation and PCIe link state, and the kernel's and the inference server's messages, to disk for the whole run. A recurrence can then be tied to power or ruled out. The commands are in `DEVELOPMENT.md`, "Logging the inference host during a long run". --- ## F. Acceptance matrix Every test currently marked **REQUIRED FOR V1**, individually. Evidence type is `browser` (real Firefox, frozen build), `campaign` (the 100-turn run against a real narrator), `container` (no-network Docker run), `process` (spawned server processes over HTTP), `suite` (automated tests), or a combination. A REQUIRED test is not marked PASS on source inspection alone; where inspection is the only evidence, it says so and the verdict is qualified. ### A — Local-first operation | ID | Result | Evidence | | --- | --- | --- | | A01 Start application offline | **PASS** | container: first page load from a fresh volume with no route and no DNS; every referenced asset served locally | | A02 Storyteller loopback default | **PASS** | suite `test_local_only_surface.py` (start scripts, compose publishes `127.0.0.1:8000:8000`); every M11 harness reached it only on loopback | | A03 No cloud API key | **PASS** | suite: no `api_key` in settings, none settable through the API, no Authorization header; container: no secret in an export | | A04 Campaign survives restart | **PASS** | campaign: 3 genuine process restarts (4 process starts), with transcript/head/state/Save Points/knowledge/settings compared before and after each and identical every time (§G.2); container: campaigns survive a container restart | | A05 Failed model call does not corrupt story | **PASS** | campaign: a real failed call at a controlled point (the model name pointed at one the server does not serve), then play resumed; container: same with no model reachable at all — no narration accepted, state unchanged, the player's words kept, earlier story reachable | | A06 Trusted-LAN Ollama inference | **PASS** | browser + identity diagnostic + the CPU-host campaigns: Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1) | ### B — Core play | ID | Result | Evidence | | --- | --- | --- | | B01 Natural language action | **PASS** | browser (two real turns through the UI) + campaign (100+) | | B02 Dialogue input | **PASS** | campaign: dialogue beats are part of the fixture's turn list | | B03 Continue | **PASS** | suite `test_turn_flow_integration.py`; browser: the Continue control is present and enabled | | B04 Story direction *(SHOULD)* | **PASS** | suite; browser: the direction toggle and its hint | ### C — Story authority and state | ID | Result | Evidence | | --- | --- | --- | | C01 Campaign canon is preserved | **PASS** | campaign: the canon section is in the prompt on 101 of 101 turns (`canon_tokens`, §G.3); suite | | C02 Possession state | **PASS** | campaign (the silver key) + suite + sci-fi fixture (the data crystal) | | C03 Character knowledge is not invented | **PASS** | suite `test_worldstate_integration.py`, `test_narrative_state.py` | | C04 Manual state correction | **PASS** | campaign: two corrections, both accepted, one carrying the planted clue; suite, including the M11 regression that a *partly* refused correction now says so | | C05 Canon beats reference | **PASS** | suite `test_knowledge_calibration.py`, `test_imported_knowledge.py` | | C06 Structured state matches accepted narrative consequence | **PASS** | campaign: real state extraction across 100 turns against the reference narrator, with every accepted event validated and every refusal recorded; suite `test_narrative_realistic.py` against a real model | ### D — Non-destructive history | ID | Result | Evidence | | --- | --- | --- | | D01 Undo one turn | **PASS** | browser + campaign | | D02 Minimum five undos | **PASS** | suite `test_head_cursor.py`; campaign (undo/redo and undo/diverge sequences) | | D03 Unlimited undo *(SHOULD)* | **PASS** | suite: undo to the root and back | | D04 Redo | **PASS** | browser (returns to the same position) + campaign | | D05 Redo invalidated by new continuation | **PASS** | campaign: after diverging, redo is no longer available — recorded in the timeline | | D06 Retry narrator response | **PASS** | campaign: two retries at scheduled points (§G.2) | | D07 Select prior retry take | **PASS** | campaign: take selection back to index 0 | | D08 Retry does not delete prior take | **PASS** | campaign: take count on the turn after retry; suite | | D09 Edit earlier user input | **PASS** | suite `test_take_edit.py`, `editRouting.test.jsx` | | D10 Edit narrator output | **PASS** | suite; browser (the hostile-Markdown scenario plants text through the narrator-edit path) | | D11 Named checkpoint | **PASS** | campaign: two named Save Points; suite; process-restart suite | | D12 Restore checkpoint | **PASS** | campaign: a restore at a scheduled point; suite | | D13 Restore does not delete later history | **PASS** | campaign: retained actions after the restore; suite | | D14 Delete checkpoint | **PASS** | suite `test_save_points.py`; browser: the delete confirmation dialog | ### E — Branch and derived-data isolation | ID | Result | Evidence | | --- | --- | --- | | E01 Abandoned future cannot affect active state | **PASS** | suite `test_m11_leakage.py` (§H), with a positive control | | E02 Abandoned memory cannot leak | **PASS** | as above; retained on disk, absent from the prompt | | E03 Abandoned summary cannot leak | **PASS** | as above — **a summary regenerated after the divergence**, which is the shape M6's review established | | E04 Scene state is lineage-safe | **PASS** | as above, including M10's derived Scene Packet | ### F — Long-term memory and context | ID | Result | Evidence | | --- | --- | --- | | F01 Recent turns remain coherent | **PASS** | campaign: the history window is populated every turn and the newest turns are always included | | F02 Old important event retrieval | **PASS** | campaign M04 (§G.4): the planting turn outside the history window and the fact recovered, through authoritative state and the narrator's restatements rather than independent memory retention | | F03 Prompt remains bounded | **PASS** | campaign: prompt size across 100 turns (§N), plus the cap itself (§E) | | F04 Output token reserve | **PASS** | campaign: `output_reserve` present in every turn's measurement and subtracted before history is chosen | | F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt and its budget | | F06 Retrieval provenance | **PASS** | campaign: `knowledge.used` per turn in the stored snapshot; suite | | F07 Heuristic memory is not canon | **PASS** | suite `test_memory_nodes.py` (authority) | | F08 Memory failure is non-fatal | **PASS** | suite `test_context_memory.py`, including §O.7's regressions that a failure is recorded even when it breaks the session; container: derived work fails with no model and turns still commit; campaign: post-turn work checked after every turn, 0 failures | ### G — Imported knowledge | ID | Result | Evidence | | --- | --- | --- | | G01 Import local text | **PASS** | browser: through the real file input; container: offline | | G02 Import local Markdown | **PASS** | campaign: three sources imported (canon, reference, inspiration); browser | | G03 Classification | **PASS** | campaign: all three classes present after the move (§K); suite | | G04 Disable knowledge source | **PASS** | suite `test_change_visibility.py`, `test_imported_knowledge.py` | | G05 Canon retrieval | **PASS** | campaign: canon passages in stored prompts; suite | | G06 Reference retrieval | **PASS** | sci-fi fixture (spin gravity) + suite | | G07 Inspiration is low authority | **PASS** | suite `test_knowledge_calibration.py` | | G08 No automatic URL fetch | **PASS** | suite `test_egress.py`; container: no network at all and import still works | | G09 Remote Markdown image does not auto-load | **PASS** | browser: no `http` image src in the rendered story | | G10 Prompt injection in source is treated as data | **PASS** | suite `test_imported_knowledge.py`; browser: injection text rendered as text | ### H — Security | ID | Result | Evidence | | --- | --- | --- | | H01 No unexpected outbound connections | **PASS** | container (no network at all) + the CPU-host campaign (one destination, the configured Ollama; the later runs were not network-monitored, §J) + suite `test_egress.py` | | H02 No telemetry | **PASS** | suite + dependency audit (§M) | | H03 No cloud provider required | **PASS** | container: a full campaign offline; suite | | H04 Model output cannot execute shell | **PASS** | browser (shell text rendered as text) + suite (no subprocess/eval anywhere in the turn path) | | H05 Invalid state event rejected | **PASS** | suite: unknown type and unknown reference both refused with the document unchanged | | H06 Stored XSS protection | **PASS** | browser: `onerror` and `