Files
interactive-story/planning/reports/M11-IMPLEMENTATION-REPORT.md
T
JesseMarkowitzandClaude Opus 5 432f04100b
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
M11 closeout: accept v1 release validation
The browser, offline and identity runs had last been taken on ef25b0a. The
closeout repeated them on the exact release-candidate tree, 3652dc6, whose
product code is identical to 96c1bf5, where the 100-turn evidence was run. No
product code changed, so the long-run evidence stands.

On 3652dc6:
- backend suite: 1421 passed, 17 skipped, 0 failed
- frontend suite: 161 of 161; lint clean
- production build clean
- docker build --no-cache: image SPA byte-identical to the local build
- browser regression: 38 of 38
- no-network container: 23 of 23
- identity diagnostic: 0 signals; the scripted self-test's detectors fire
- release-shaped smoke test from the image: 14 of 14

- M11 report: new S (exact-tree verification, including the identity
  results the report never carried) and T (acceptance record). Corrections:
  the REQUIRED FOR V1 count is 82, not 85, and L's browser narrator was
  qwen2.5:3b-instruct. P gains risks 16 and 17; risk 6 is widened.
- BUILD-MILESTONES.md: M11 COMPLETE / ACCEPTED, and a post-v1 backlog.
- V1-ACCEPTANCE-TESTS.md: the P release gate's result, and the P3
  disposition's run.
- planning/README.md, VERSION.md (v4.0), README.md: status, map, stop rule.
- tools/m11_browser.py: the G01 import wait could not fail, because the
  scenario's campaign is titled "Hidden Knowledge". It now waits for the
  imported source's row.

Found and carried, not fixed. The identity run stored protocol shapes the
extractor leaves, on 4 of 10 turns at a 4,096 window: event-call syntax and a
parroted length hint. The owner chose residual risk. The state rule's example
is fantasy, and the state lagged the narration. None occurs in the 100-turn
evidence.

No requirement changes. No release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CBTT3qbkGoemWD7BRvvxpT
2026-09-14 06:07:16 -04:00

1936 lines
120 KiB
Markdown

# M11 — v1 Security, Long-Run, and Release Validation
**Implementation and verification report, written for independent release review.**
Branch `m11-release-validation`, from the signed M10 commit `1013c94`.
Implemented and verified 2026-09-07. The long-run evidence was re-established,
and this revision written, on 2026-09-14.
**Closeout, 2026-09-14.** The browser, offline and identity runs were repeated on
the exact release-candidate tree (§S), and M11 was accepted (§T). Sections A-R
are the implementer's evidence as revised before the closeout, corrected where
the closeout found them wrong. §S and §T are the closeout.
---
## A. Executive result
**PASS, and accepted at closeout (§T).** All 82 tests marked REQUIRED FOR V1
pass, with H09 recorded NOT APPLICABLE on the condition its own text states.
*Corrected at closeout:* earlier revisions said 85. That figure counted the four
SHOULD rows in §F and left H09 out.
This revision supersedes the first. That revision left M01 PARTIAL at 41
accepted turns, and then the run it described was lost to a host crash along
with its evidence. **M01 to M04 now pass on a complete 100-turn campaign**, run
by the committed harness on commit `96c1bf5`: 101 accepted turns, three genuine
process restarts, every scheduled history operation performed, zero failed
post-turn passes, and recovery onto a clean data directory 16 of 16 (§G, §K).
Getting there took six further runs. They found **two more product defects**,
both fixed and both invisible to every earlier piece of evidence because the
memory bank had never been switched on during a long run:
1. **A turn locked its own memory bank out** (§O.7). Retrieval wrote a use
counter before the model call, and the turn committed only after the reply,
so SQLite's single write lock was held for the whole reply. Post-turn memory
and summary writes timed out behind it, and recording those failures timed out
too. A 26-turn run reported "complete" with two memories, no summary and 180
`database is locked` errors, while derived status read `idle`.
2. **The narrator's protocol leaked into stored story** (§O.8). A small model
pasted the narrative-state section and unfenced proposals into its prose, on
42 of 104 turns in one run, and stored text is replayed as history. The
extractor now removes every shape observed. Replaying 443 real turns from five
runs changed no turn the old extractor had left clean.
The first revision's three corrections stand: the context-window mismatch is
resolved (§E), a partly refused manual correction now says so, and the
narration-length setting moves a number. Post-M8 findings A and B are closed.
**What is not claimed.**
- **That memory keeps a planted fact on its own.** M04 passes because the fact
was recoverable with its planting turn 53 depths outside the history window.
The recovery path ran through authoritative state: the narrator restated the
state's fact line in its prose, and memory summarised those restatements
(§G.4). No memory of the planting era carried the fact in any run. The
repository owner accepted state-based recovery for M04 on 2026-09-13.
- **Headroom in the context window.** The largest prompts fill 16,342 of 16,384
tokens by the narrator's own count, and the inference server cuts an
over-window prompt without an error (§N, §P).
- **Performance.** Timings are two particular hosts', not a product
characteristic (§E.1).
- **Post-M8 finding D's root cause**, which remains unestablished. The identity
diagnostic's closeout run, and what it can and cannot establish, are in §S.5.
- *Closed at closeout:* the browser, offline and identity runs had predated the
last three product commits. All three were repeated on the release-candidate
tree (§S).
**Zero requirement weakenings.** No acceptance test was retired, relaxed or
reclassified. The M04 precondition the harness measures was corrected to the
acceptance text's own wording, "without entire transcript in prompt" (§O).
---
## B. Repository and provenance
| | |
| --- | --- |
| **Base commit** | `1013c94eb1ad283e960114aef04c19c2806b5db7` — *"M10: the seam for media, and no media"* |
| **Signature** | `git verify-commit 1013c94` → **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, "JesseMarkowitz", trust `[ultimate]`. `%G?` = `G`. |
| **Branch** | `m11-release-validation`, created from that commit. The working tree was clean at the start (`git status --porcelain` empty). |
| **HEAD at this revision** | `96c1bf5`. Seven commits follow the base, all signed by the owner (`%G?` = `G`): `144406c` (M11), `fedb714`, `ef25b0a`, `fec46f6`, `f8d4010`, `0c7316f`, `96c1bf5`. This revision of the report is staged for the owner to sign. **At closeout** the candidate is `3652dc6`. `d198806` and `3652dc6` follow `96c1bf5`, both are signed, and both change documentation only: no file under `backend/`, `frontend/`, `Dockerfile`, `docker-compose.yml` or the start scripts differs from `96c1bf5` (§S.1). |
| **Upstream ancestry** | `git merge-base --is-ancestor d72f7c1b HEAD` → true. The AI-DnD fork point is still an ancestor. |
| **LICENSE** | Unchanged — md5 `07fde30437134836e2ee875e82a7cd31`, MIT, "Copyright (c) 2026 Parth Thakkar". `PROVENANCE.md` unchanged. |
**The four trees, kept distinct**, because M11's evidence discipline depends on
which one produced a given number:
```text
committed base 1013c94, signed by the owner: M1-M10 as accepted
working tree the base plus M11's changes; this is what was tested
staged tree identical to the working tree (§29 lists it)
frozen tree the working tree at the point each black-box run started,
with no source edit during any reported run
```
The black-box runs come from two trees. **The long-run evidence**, the 100-turn
campaign and its recovery run, was taken on commit `96c1bf5`, after the last
product change, with no source edit during the run. **The browser regression,
the offline container, the identity diagnostic and the contrast audit** were
re-run on the `ef25b0a` tree, as that commit's message records, and their
evidence is dated 2026-09-10. The three product commits after it (`f8d4010`,
`0c7316f`, `96c1bf5`) change backend files only, and those runs were not
repeated at the time. **At closeout all three were repeated on the
release-candidate tree `3652dc6`** (§S). No reported black-box result now comes
from product code other than the candidate's. Runs taken before a product change on their own tree were
discarded and re-run; §O records which and why.
---
## C. Change inventory
**Product code — backend**
| File | Change |
| --- | --- |
| `app/contextwindow.py` | **New.** Discovers the server's real context window and provides the ceiling. The whole of §E. |
| `app/context/builder.py` | Takes a `window`; the effective budget is `min(configured, verified)`; the context report carries what was verified and how; `length_hint` gains the campaign's length band. |
| `app/routers/adventures/turns.py` | Probes the window before assembling a turn, and stores the verdict in the turn's provenance. |
| `app/routers/adventures/insights.py` | The same probe, so the inspector shows the prompt the next turn will actually send. |
| `app/routers/settings.py` | The connection test reports the window, or why it could not be checked; changing endpoint or model clears what was learned. |
| `app/routers/adventures/state.py` | A correction reports refusals (`refused`) and shared display names. |
| `app/routers/adventures/crud.py` | Persists the campaign's narration-length choice. |
| `app/narrative/model.py` | **New** `duplicate_names` — detection, deliberately not a refusal. |
| `app/models.py`, `app/migrations.py`, `app/schemas.py`, `app/bundle.py` | `adventures.narration_length`: column, migration 93, API in and out, bundle carriage with an unknown value dropped. |
**Product code — frontend**
| File | Change |
| --- | --- |
| `src/documentTitle.js` | **New.** The product's name, in one place; the tab follows the open campaign. |
| `index.html`, `src/pages/Play/index.jsx` | The inherited `AI D&D` title replaced and kept in step (finding A). |
| `src/pages/Play/Composer.jsx`, `styles/story.css` | The position indicator: `Moment 11 · later story ahead` (finding B, §8A). |
| `src/pages/NewCampaign.jsx`, `panels/CampaignSettingsPanel.jsx` | Narration length sent and editable as data (finding C). |
| `src/pages/Settings.jsx`, `styles/library.css`, `styles/tokens.css` | The context-window report and warning; a `--warning` token measured at 7.85:1. |
| `panels/StatePanel.jsx`, `styles/insights.css` | Refused corrections and shared names, shown to the reader. |
**Release-test infrastructure — no application code imports any of it**
| File | Purpose |
| --- | --- |
| `tools/m11_long_run.py` | The 100-turn campaign: M01-M04. |
| `tools/m11_recovery.py` | That campaign, moved to a clean data directory: I01-I07. |
| `tools/m11_browser.py`, `tools/m11_webdriver.py` | The browser regression and accessibility measurements; a dependency-free W3C WebDriver client. |
| `tools/m11_offline.py` | A container with no network: §18 and the packaging path. |
| `tools/m11_identity.py` | The multi-character identity diagnostic (finding D). |
| `tools/contrast_audit.py` | The palette against WCAG AA. |
| `tests/test_m11_context_window.py` | 20 tests: the probe, the cap, the truncation sentinel. |
| `tests/test_m11_leakage.py` | 14 tests: E01-E04 together, in one long campaign. |
| `tests/test_m11_scifi.py` | 10 tests: J01-J03, the Persephone fixture. |
| `tests/test_m11_migration.py` | 9 tests: fresh-versus-upgraded schema parity, and the upgrade. |
| `tests/test_m11_security.py` | 25 tests: the H-series against the assembled product. |
| `tests/test_m11_findings.py` | 20 tests: findings C and D, and the refused-correction regression. |
| `tests/test_m11_real_window.py` | 3 tests: the probe against a real Ollama (skipped without one). |
| `tests/schema_rewind.py` | Migration 93's inverse, so the suite can replay it. |
**Dependencies: none added, none removed, none upgraded.** `requirements.txt`,
`requirements.lock` and `package.json` are byte-identical to M10's.
**Changes after the first revision.** `144406c` is the first revision itself.
Every later commit is listed here, and all are signed by the owner.
| Commit | Change | Files |
| --- | --- | --- |
| `fedb714` | Release evidence must be written to a durable `--out`, and the report recorded the lost run | `DEVELOPMENT.md`, this report |
| `ef25b0a` | The history window trims in blocks, so the server's prompt cache survives. `settings.context_window_override` (**migration 94**) covers a server the probe cannot ask. The long run gains `--resume`, a configurable turn timeout and an abort after consecutive failures. Two checks that could not fail were fixed. Every black-box harness was re-run on this tree | `app/context/builder.py`, `app/contextwindow.py`, `app/models.py`, `app/migrations.py`, `app/schemas.py`, `app/routers/settings.py`, `app/routers/adventures/turns.py`, `app/routers/adventures/insights.py`, `frontend/src/pages/Settings.jsx`, `tools/m11_long_run.py`, `tools/m11_browser.py`; tests `test_history_block_trim.py`, `test_m11_declared_window.py`, `test_m11_long_run_resume.py`, `settingsWindow.test.jsx`; planning package v3.8 |
| `fec46f6` | The long run switches the memory bank and auto-summarise on, and refuses a run with no embedding model | `tools/m11_long_run.py`, `tests/test_m11_long_run_memory.py` |
| `f8d4010` | §O.7: no write lock is held across the model call, and a failure that breaks the session is still recorded. The harness reads the real section labels and stops on failed post-turn work | `app/memorybank.py`, `app/routers/adventures/turns.py`, `app/routers/adventures/insights.py`, `tools/m11_long_run.py`, five test files |
| `0c7316f` | §O.8: protocol is kept out of stored narration, and the fence-label bug is fixed. The harness gains an M04 verdict and a protocol-leak count | `app/narrative/extract.py`, `app/narrative/render.py`, `tools/m11_long_run.py`, `tests/test_narrative_state.py`, `tests/test_m11_long_run_memory.py` |
| `96c1bf5` | §O.8 extended to a model's own sections under decorated headings. M04's precondition is now positional | `app/narrative/extract.py`, `tools/m11_long_run.py`, `tests/test_narrative_state.py`, `tests/test_m11_long_run_memory.py` |
No later commit touches `requirements.txt`, `requirements.lock` or `package.json`.
---
## D. M10 and M8 handoff
Every item handed to M11, and what happened to it. Nothing here is closed by
"the automated tests are green".
### From M10's §O.1 (its own residual risk)
| Item | Disposition |
| --- | --- |
| **K04 satisfied structurally, not physically** — no `media_jobs`/`media_assets` | **Unchanged, and verified as such.** M11 exercised the deferred-table path rather than building the tables: a dummy provider, a real packet, a fake asset carrying the packet's `scene_id`, and the story model byte-identical afterwards. Reported as K04 = PASS on the acceptance text's own deferred branch (§F). |
| **`ambience` is an empty shape** | **Unchanged.** Filling it means extending `set_scene`, which is a prompt-path change and not release validation. Still an empty shape with the right fields. |
| **The seam has no consumer, so it is unexercised by real use** | **Partly answered.** The Persephone fixture put a third genre through the packet and a visual profile through a starship hull (`test_m11_scifi.py`), and the offline container proved the media module imports and stays inert with no network. Still no real adapter, and that remains true until someone writes one. |
| **A visual profile cannot be recovered from within the app** | **Unchanged.** It travels in the bundle; there is no undo for deleting one. |
| **Profiles are API-only, with no reader-facing surface** | **Unchanged, deliberately.** M11 is release validation; adding a media UI would be new feature work. |
### From M9's residual risks
| Item | Disposition |
| --- | --- |
| **Bundle ceiling ~279 turns** | **Measured against the real 100-turn campaign** rather than the fixture — §N gives the actual bundle size and what fraction of the 20 MB import limit it is. |
| **`quick_check` rather than `integrity_check`** | Unchanged; re-exercised on a migrated database (`test_m11_migration.py`). |
| **No scheduled backup** | Unchanged, and outside the acceptance contract. |
| **Stale `chunk_id` in a restored snapshot** | Unchanged; a React key, not a live pointer. |
| **The importing machine's context window may differ** | **Closed.** This was the same defect as M8's, seen from the import side, and §E closes both: the application caps to what the destination server accepts and says when it could not check. `DEVELOPMENT.md`'s import warning is rewritten accordingly. |
| **This machine cannot drive a file into or out of the browser** | **Narrower than recorded, and half of it is closed.** The snap Firefox refuses a WebDriver file path under `/tmp`; a path under `$HOME` works. Knowledge import is now proved end-to-end in a real browser (§L). The *download* half — a `blob:` export leaving the browser — is still not driveable here and is still recorded as a limitation. |
### From M8 (via M9 and M10)
| Item | Disposition |
| --- | --- |
| **The deployment context ceiling** | **Closed** — §E. This was M11's stated release blocker. |
| **Contrast and visible focus checked by eye** | **Measured** — §L. Palette pairs by calculation (`tools/contrast_audit.py`), rendered colours in a real browser, plus focus visibility, accessible names, tab order, hover-only controls and modal focus. |
| **The four post-M8 playtest findings** | **All four disposed of** — §D.1 below. |
### D.1 The four post-M8 playtest findings
**A — the tab read `AI D&D`.** Fixed. The name is `Interactive Story`, chosen by
the repository owner when M11 asked, and deliberately not "Adventure
Storyteller": `SPECIFICATION.md` requires a genre-agnostic engine and *Adventure*
is narrower than the thing it names. The tab shows the open campaign first
(`Westhaven — Interactive Story`), and one module owns the string.
**B — no orientation after Undo.** Fixed, to `BROWSER-UX-SPEC.md` §8A. A status
line at the end of the control row reads `Moment 11 · later story ahead`. The
number and the clause are both the *server's* answers, it changes visibly after
Undo, and it uses no implementation vocabulary. Asserted in the component suite
and in the browser regression, the second of which is what §8A demanded when it
said the requirement must be observable rather than inferable.
**C — narration length had no measurable effect.** Fixed, at the mechanism the
finding identified. The choice is now data on the campaign, and `length_hint`
turns it into a real word band (70-180 / 150-380 / 320-700), bounded by the
reply cap. The generation budget is deliberately **not** touched: capping it per
length would make a brief turn likelier to be cut off mid-sentence, and the
state block is emitted last, so the first thing a truncated reply loses is the
turn's state. Before M11 all three settings produced *the identical sentence*;
`test_the_three_lengths_no_longer_say_the_same_thing` fails against that.
**D — character identity confusion.** The diagnostic exists; the root cause does
not, and cannot. The diagnostic caught a defect in its own fixture (§O.2), and
that is why its first run's evidence was discarded. Earlier revisions of this
report reproduced no run's results. **The closeout run's results are in §S.5**:
clean on every objective check, with the detectors proved to fire, and with what
the run cannot establish stated.
---
## E. The context-window resolution
### The original mismatch
M8 measured the reference deployment enforcing **4,096** input tokens while the
application budgeted **16,384**. M11 re-measured it, on the same server, before
changing anything:
```text
$ curl -sk https://<host>/api/show -d '{"model":"qwen2.5:3b-instruct"}'
model_info["qwen2.context_length"] = 32768 the architecture's ceiling
parameters = (none) no num_ctx is baked in
$ (one /v1/chat/completions call to load it, then /api/ps)
qwen2.5:3b-instruct context_length = 4096 size_vram = 0
```
So the mismatch was live on the reference deployment on the day M11 started, and
`size_vram = 0` is why: with no VRAM Ollama picks a 4,096 default.
### Root cause
Two independent facts that only bite together.
1. **Ollama's window is a property of how the model was loaded**, not of the
request. Its OpenAI-compatible endpoint accepts `num_ctx` — nested or
top-level — returns 200, and ignores it; M8 established that, and it is why
the operational fix is a model with `num_ctx` baked in or
`OLLAMA_CONTEXT_LENGTH` on the server.
2. **The application had no way to know.** `Settings.context_token_budget` was
the only number in play, so the builder assembled to it and the server
quietly did what it liked with the excess — which is drop the **oldest**
tokens. The oldest tokens here are the system block: the narrator's rules and
the campaign canon. The failure therefore looks like a narrator that stops
respecting canon deep into a long session, with nothing on screen to explain
it, and every acceptance test that reads a 200 as success passing throughout.
### The fix
`backend/app/contextwindow.py`, plus four call sites. The rule:
> A **verified** window is a ceiling on the configured budget. An **unverified**
> one leaves the budget standing and is recorded as unverified. There is no
> third behaviour.
- **Discovery** asks the server the application is already talking to, on the
native path beside `/v1`, through `endpoints.rejection_reason` and the shared
TLS trust store — so it can reach exactly what a turn can reach and nothing
more. `/api/ps` gives the window a resident model is *actually* being served
with; `/api/show` gives the `num_ctx` an unloaded one will load with, capped
by the architecture's own ceiling.
- **Enforcement** is one line in the builder: the budget every section is priced
against is `min(configured, verified)`. Because the output reserve is
subtracted from that budget by the existing arithmetic, the assembled prompt
plus the reply reserve fits inside the window by construction.
- **Reporting.** The context report and the turn's stored snapshot carry
`window: {verified, tokens, source, model_max, detail, capped}`, so an old turn
can be asked afterwards whether it was built against a checked window. The
connection test in Settings shows the number or explains why it could not be
checked, in three distinct messages, because "could not check", "smaller than
your budget" and "fine" need three different things done about them.
- **What it does not do:** hard-code 4,096 (right on one machine, wrong on the
next), raise anyone's window, guess from a model's name, or add a provider
abstraction. An unknown window is reported as unknown.
### The numbers, on the reference deployment
| | |
| --- | --- |
| Ollama | 0.33.0, CPU-only (`size_vram = 0`) |
| `qwen2.5:3b-instruct` | **4,096** tokens, source `loaded` (`/api/ps`) |
| `qwen2.5:3b-instruct-16k` | **16,384** tokens, source `parameters` (`num_ctx` baked in), architecture ceiling 32,768 |
| Application budget (M01) | `context_token_budget` = 16,384, `max_output_tokens` = 500 |
| M01's effective budget | **16,384** — verified equal to the server's window, so nothing was capped away |
| Largest assembled prompt in M01 | see §N |
Measured end to end in `tests/test_m11_real_window.py` against the real server:
```text
turn 1 window verified=False (the model is not resident yet)
turn 2 window verified=True tokens=4096 source=loaded
budget configured=16384 effective=4096
prompt 850 tokens + 464 reserved -> 1314 <= 4096
```
That is the whole behaviour in six lines: the first turn on a cold model cannot
verify and says so; from the second turn the cap is live and the prompt provably
fits.
### Why silent truncation can no longer invalidate M01
Three separate reasons, and the third is the one that matters:
1. **M01 ran on the 16k model**, whose window (16,384) equals the application's
budget, so nothing was capped and nothing was near the edge — recorded per
turn in `timeline.jsonl` as `window_verified` and `window_tokens`.
2. **Every M01 turn recorded the verification**, so the claim is per-turn
evidence rather than a statement about the configuration at the start.
3. **The failure mode is now impossible to reach silently.** If the window were
smaller than the budget, the prompt would be built to the *window*, and the
canon at the front would survive by construction —
`test_the_canon_at_the_front_survives_a_window_far_too_small` plays 120 turns
into a 4,096-token window and finds the canon sentinel still present with the
*oldest history* dropped instead. Its companion,
`test_without_the_cap_the_same_prompt_would_have_overflowed`, builds the same
campaign with no verified window and measures a prompt more than twice the
size — the defect, reproduced, so the fix is shown to be doing something.
---
## E.1 The release environment
Recorded because a timeout without hardware beside it is not a measurement. Two
inference hosts produced this report's evidence, and every result below says
which.
### The application machine
It runs the storyteller, every harness, the browser and the container.
| | |
| --- | --- |
| **OS** | Ubuntu 24.04.4 LTS, kernel 7.0.0-30-generic, at the first revision; Ubuntu 24.04.5 LTS, kernel 7.0.0-31-generic, for the 2026-09-13 long runs |
| **CPU** | AMD Ryzen 7 7840HS, 4 cores available to this VM, 1 thread per core |
| **GPU** | **none**: VMware SVGA II |
| **RAM** | 15 GiB |
| **Python** | 3.12.3 |
| **Node / npm** | v22.23.1 / 10.9.8 |
| **Docker** | 29.7.2 |
| **Browser** | Firefox 154.0.1, headless, via geckodriver over W3C WebDriver |
### The CPU reference host
Used from 2026-09-07 to 2026-09-10: the lost run, the memory-off 100-turn run,
the browser run and the identity diagnostic.
| | |
| --- | --- |
| **Machine** | a separate physical machine on the trusted LAN, no GPU (`size_vram = 0`) |
| **Ollama** | 0.33.0, **HTTPS** with a private CA installed in the application machine's OS trust store |
| **Measured cost** | 88-203 s per turn at a 4,096 window; 229-291 s at 16,384 |
### The GPU inference host
Used on 2026-09-13 and 2026-09-14: both 26-turn trials and the three 100-turn
runs, including the evidence run.
| | |
| --- | --- |
| **Machine** | a separate physical machine on the trusted LAN; Ubuntu 24.04.2 LTS; 16 logical CPUs; 14 GiB RAM |
| **GPU** | NVIDIA GeForce GTX 1080 Ti, 11 GiB (Pascal, compute capability 6.1), driver 580.173.02, in an **OCuLink dock**: PCIe x4, measured at Gen 3 under load and Gen 1 at idle |
| **Ollama** | 0.34.0. Its CUDA 13 runner skips a compute-6.1 card and the CUDA 12 runner serves it; both models loaded 100% into VRAM |
| **Transport** | **plain HTTP** on the LAN. H12 permits a LAN address; this host is **not** A06 evidence |
| **Power limit** | the card's default, 280 W, during every run |
| **Measured cost** | prefill ~1,350-1,850 tokens/s and generation ~87 tokens/s with synthetic prompts; 4.1-22.8 s per turn in the evidence run |
### Models and budget
| | |
| --- | --- |
| **Narrator** | `qwen2.5:3b-instruct-16k`: `num_ctx` 16,384 baked in, architecture ceiling 32,768. Digest `21ff8cc52f37`, identical on both hosts |
| **Base narrator** | `qwen2.5:3b-instruct`, digest `357c53fb659c`, identical on both hosts. With no `num_ctx` the server gives it **4,096** |
| **State/summariser model** | the narrator; no separate summariser is configured |
| **Embedding model** | `nomic-embed-text`, digest `0a109f422b47`, identical on both hosts |
| **Application context budget** | `context_token_budget` = 16,384, `max_output_tokens` = 500, reply reserve 564 |
| **Effective budget** | 4,096 in the lost run, where it was capped every turn; **16,384** in every later run, where window and budget were equal |
**Why the first campaign ran at 4,096.** On the CPU host the recommended 16,384
window cost 229-291 s a turn, about eight hours for a hundred. The first revision
therefore ran the default window, which doubled as the longest test of §E's cap.
That run reached 97 of 100 turns and was lost (§G.6). Every later campaign runs
the recommended window.
### The GPU host incident
About 30 seconds after the evidence run's last post-turn pass finished, the GPU
dropped off the PCIe bus: `NVRM: Xid 79, GPU has fallen off the bus`, then
`Xid 154`, reset required. Ollama stayed up without a GPU, stopped completing
requests, and the host had to be rebooted. The same host had logged one `Xid 13`
graphics exception from the inference server earlier that day. There was no
out-of-memory event and no thermal event recorded. The firmware does not support
PCIe AER, so no link-error trail exists.
After the reboot, under a temporary 200 W cap with logging running, the link held
Gen 3 x4 under load, peak draw was 207.8 W, and there was no fault. **The cause
is not established.** Power transients at the uncapped 280 W over a long run are
the leading candidate. **The evidence is unaffected**: every record the result
rests on was written before the drop (§G.1).
**Required for every future long run on a GPU host:** log the card's power,
temperature, utilisation and PCIe link state, and the kernel's and the inference
server's messages, to disk for the whole run. A recurrence can then be tied to
power or ruled out. The commands are in `DEVELOPMENT.md`, "Logging the inference
host during a long run".
---
## F. Acceptance matrix
Every test currently marked **REQUIRED FOR V1**, individually. Evidence type is
`browser` (real Firefox, frozen build), `campaign` (the 100-turn run against a
real narrator), `container` (no-network Docker run), `process` (spawned server
processes over HTTP), `suite` (automated tests), or a combination. A REQUIRED
test is not marked PASS on source inspection alone; where inspection is the only
evidence, it says so and the verdict is qualified.
### A — Local-first operation
| ID | Result | Evidence |
| --- | --- | --- |
| A01 Start application offline | **PASS** | container: first page load from a fresh volume with no route and no DNS; every referenced asset served locally |
| A02 Storyteller loopback default | **PASS** | suite `test_local_only_surface.py` (start scripts, compose publishes `127.0.0.1:8000:8000`); every M11 harness reached it only on loopback |
| A03 No cloud API key | **PASS** | suite: no `api_key` in settings, none settable through the API, no Authorization header; container: no secret in an export |
| A04 Campaign survives restart | **PASS** | campaign: 3 genuine process restarts (4 process starts), with transcript/head/state/Save Points/knowledge/settings compared before and after each and identical every time (§G.2); container: campaigns survive a container restart |
| A05 Failed model call does not corrupt story | **PASS** | campaign: a real failed call at a controlled point (the model name pointed at one the server does not serve), then play resumed; container: same with no model reachable at all — no narration accepted, state unchanged, the player's words kept, earlier story reachable |
| A06 Trusted-LAN Ollama inference | **PASS** | browser + identity diagnostic + the CPU-host campaigns: Ollama on a **separate physical machine** over **HTTPS** with a **private CA installed in this machine's OS trust store**, certificate and hostname verification on, no bypass; the storyteller stayed loopback-bound. The evidence campaign (§G) used plain HTTP to a LAN GPU host, which H12 permits and which is not A06 evidence (§E.1). **Re-established at closeout on the candidate tree** by the browser regression and the identity diagnostic, on the same HTTPS host (§S) |
### B — Core play
| ID | Result | Evidence |
| --- | --- | --- |
| B01 Natural language action | **PASS** | browser (two real turns through the UI) + campaign (100+) |
| B02 Dialogue input | **PASS** | campaign: dialogue beats are part of the fixture's turn list |
| B03 Continue | **PASS** | suite `test_turn_flow_integration.py`; browser: the Continue control is present and enabled |
| B04 Story direction *(SHOULD)* | **PASS** | suite; browser: the direction toggle and its hint |
### C — Story authority and state
| ID | Result | Evidence |
| --- | --- | --- |
| C01 Campaign canon is preserved | **PASS** | campaign: the canon section is in the prompt on 101 of 101 turns (`canon_tokens`, §G.3); suite |
| C02 Possession state | **PASS** | campaign (the silver key) + suite + sci-fi fixture (the data crystal) |
| C03 Character knowledge is not invented | **PASS** | suite `test_worldstate_integration.py`, `test_narrative_state.py` |
| C04 Manual state correction | **PASS** | campaign: two corrections, both accepted, one carrying the planted clue; suite, including the M11 regression that a *partly* refused correction now says so |
| C05 Canon beats reference | **PASS** | suite `test_knowledge_calibration.py`, `test_imported_knowledge.py` |
| C06 Structured state matches accepted narrative consequence | **PASS** | campaign: real state extraction across 100 turns against the reference narrator, with every accepted event validated and every refusal recorded; suite `test_narrative_realistic.py` against a real model |
### D — Non-destructive history
| ID | Result | Evidence |
| --- | --- | --- |
| D01 Undo one turn | **PASS** | browser + campaign |
| D02 Minimum five undos | **PASS** | suite `test_head_cursor.py`; campaign (undo/redo and undo/diverge sequences) |
| D03 Unlimited undo *(SHOULD)* | **PASS** | suite: undo to the root and back |
| D04 Redo | **PASS** | browser (returns to the same position) + campaign |
| D05 Redo invalidated by new continuation | **PASS** | campaign: after diverging, redo is no longer available — recorded in the timeline |
| D06 Retry narrator response | **PASS** | campaign: two retries at scheduled points (§G.2) |
| D07 Select prior retry take | **PASS** | campaign: take selection back to index 0 |
| D08 Retry does not delete prior take | **PASS** | campaign: take count on the turn after retry; suite |
| D09 Edit earlier user input | **PASS** | suite `test_take_edit.py`, `editRouting.test.jsx` |
| D10 Edit narrator output | **PASS** | suite; browser (the hostile-Markdown scenario plants text through the narrator-edit path) |
| D11 Named checkpoint | **PASS** | campaign: two named Save Points; suite; process-restart suite |
| D12 Restore checkpoint | **PASS** | campaign: a restore at a scheduled point; suite |
| D13 Restore does not delete later history | **PASS** | campaign: retained actions after the restore; suite |
| D14 Delete checkpoint | **PASS** | suite `test_save_points.py`; browser: the delete confirmation dialog |
### E — Branch and derived-data isolation
| ID | Result | Evidence |
| --- | --- | --- |
| E01 Abandoned future cannot affect active state | **PASS** | suite `test_m11_leakage.py` (§H), with a positive control |
| E02 Abandoned memory cannot leak | **PASS** | as above; retained on disk, absent from the prompt |
| E03 Abandoned summary cannot leak | **PASS** | as above — **a summary regenerated after the divergence**, which is the shape M6's review established |
| E04 Scene state is lineage-safe | **PASS** | as above, including M10's derived Scene Packet |
### F — Long-term memory and context
| ID | Result | Evidence |
| --- | --- | --- |
| F01 Recent turns remain coherent | **PASS** | campaign: the history window is populated every turn and the newest turns are always included |
| F02 Old important event retrieval | **PASS** | campaign M04 (§G.4): the planting turn outside the history window and the fact recovered, through authoritative state and the narrator's restatements rather than independent memory retention |
| F03 Prompt remains bounded | **PASS** | campaign: prompt size across 100 turns (§N), plus the cap itself (§E) |
| F04 Output token reserve | **PASS** | campaign: `output_reserve` present in every turn's measurement and subtracted before history is chosen |
| F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt and its budget |
| F06 Retrieval provenance | **PASS** | campaign: `knowledge.used` per turn in the stored snapshot; suite |
| F07 Heuristic memory is not canon | **PASS** | suite `test_memory_nodes.py` (authority) |
| F08 Memory failure is non-fatal | **PASS** | suite `test_context_memory.py`, including §O.7's regressions that a failure is recorded even when it breaks the session; container: derived work fails with no model and turns still commit; campaign: post-turn work checked after every turn, 0 failures |
### G — Imported knowledge
| ID | Result | Evidence |
| --- | --- | --- |
| G01 Import local text | **PASS** | browser: through the real file input. The harness's own assertion could not fail until the closeout fixed it (§S.3); every run's database holds the imported source. Container: offline |
| G02 Import local Markdown | **PASS** | campaign: three sources imported (canon, reference, inspiration); browser |
| G03 Classification | **PASS** | campaign: all three classes present after the move (§K); suite |
| G04 Disable knowledge source | **PASS** | suite `test_change_visibility.py`, `test_imported_knowledge.py` |
| G05 Canon retrieval | **PASS** | campaign: canon passages in stored prompts; suite |
| G06 Reference retrieval | **PASS** | sci-fi fixture (spin gravity) + suite |
| G07 Inspiration is low authority | **PASS** | suite `test_knowledge_calibration.py` |
| G08 No automatic URL fetch | **PASS** | suite `test_egress.py`; container: no network at all and import still works |
| G09 Remote Markdown image does not auto-load | **PASS** | browser: no `http` image src in the rendered story |
| G10 Prompt injection in source is treated as data | **PASS** | suite `test_imported_knowledge.py`; browser: injection text rendered as text |
### H — Security
| ID | Result | Evidence |
| --- | --- | --- |
| H01 No unexpected outbound connections | **PASS** | container (no network at all) + the CPU-host campaign (one destination, the configured Ollama; the later runs were not network-monitored, §J) + suite `test_egress.py` |
| H02 No telemetry | **PASS** | suite + dependency audit (§M) |
| H03 No cloud provider required | **PASS** | container: a full campaign offline; suite |
| H04 Model output cannot execute shell | **PASS** | browser (shell text rendered as text) + suite (no subprocess/eval anywhere in the turn path) |
| H05 Invalid state event rejected | **PASS** | suite: unknown type and unknown reference both refused with the document unchanged |
| H06 Stored XSS protection | **PASS** | browser: `onerror` and `<script>` in accepted narration, neither executed |
| H07 JavaScript URL protection | **PASS** | browser: no `javascript:` href in the DOM |
| H08 Path traversal import rejected | **PASS** | suite: a `../../../etc/cron.d/...` filename stored as metadata, no file written |
| H09 ZIP Slip protection | **NOT APPLICABLE** | the product extracts no archives, and a test enforces it (§J) |
| H10 Restrictive CORS and local API behaviour | **PASS** | process: a wildcard origin refuses startup, a named origin does not; browser: an unknown API path is a 404 with a non-HTML body |
| H11 No first-use runtime asset download | **PASS** | container: every referenced asset served locally with no network; suite: the tokenizer table is vendored and no HTTP client is in that module |
| H12 Inference endpoint enforcement | **PASS** | suite: loopback v4 and v6, LAN address, CGNAT allowed; cloud hosts and public addresses refused; **a public endpoint written into the database behind the API is refused at request time**; the M11 window probe obeys the same policy. A06 covers the LAN-hostname-over-HTTPS case in production |
### I — Export, import and recovery
| ID | Result | Evidence |
| --- | --- | --- |
| I01 Export campaign | **PASS** | campaign: the 100-turn campaign exported (§K) |
| I02 Import exported campaign | **PASS** | process: imported into a database that never existed, in a directory that never existed (§K) |
| I03 Branch/disposable history export | **PASS** | §K: retained history present after the move, Redo walks into it |
| I04 Checkpoint export | **PASS** | §K: Save Points restore to the positions they name after the move |
| I05 Knowledge provenance export | **PASS** | §K: all three classes with their content after the move |
| I06 Database/export contains no API secrets | **PASS** | suite + container + §K |
| I07 Export/import preserves an undone active head | **PASS** | §K, and suite `test_m9_portability.py` for the older-format seams |
### J — Genre neutrality
| ID | Result | Evidence |
| --- | --- | --- |
| J01 Science-fiction campaign | **PASS** | suite `test_m11_scifi.py` (§I) |
| J02 Generic entity support | **PASS** | as above: five entity types in one document, no schema change |
| J03 Genre profiles are configuration | **PASS** | as above, plus the whole-vocabulary check that no event type names a genre noun |
### K — Future media architecture
| ID | Result | Evidence |
| --- | --- | --- |
| K01 Scene snapshot exists | **PASS** | suite `test_m10_*`; sci-fi fixture; campaign (the scene follows the active line through every history operation) |
| K02 Visual character profile | **PASS** | suite; sci-fi fixture |
| K03 Visual location profile | **PASS** | suite; sci-fi fixture (a starship hull) |
| K04 Attach media asset to scene *(SHOULD)* | **PASS on the deferred branch** | suite `test_m10_authority.py`: a dummy provider produces an asset carrying the packet's `scene_id` and the story model is byte-identical afterwards. Media tables remain deliberately unbuilt; a reviewer requiring physical tables should read this as PARTIAL |
### L — Data integrity
| ID | Result | Evidence |
| --- | --- | --- |
| L01 Atomic turn commit | **PASS** | container + campaign: real induced failures; no narration accepted, no half-written state, earlier story reachable, play resumes |
| L02 State reconstruction | **PASS** | suite; campaign: state compared across every restart |
| L03 Checkpoint reconstruction after restart | **PASS** | process-restart suite; campaign: Save Points present and restorable after each restart |
| L04 Derived data can be rebuilt *(SHOULD)* | **PASS** | suite `test_knowledge_migration.py` (reindex), `test_memory_rewrite.py` |
### M — Long-run
| ID | Result | Evidence |
| --- | --- | --- |
| M01 100-turn campaign | **PASS** | campaign: 101 accepted turns, every scheduled operation, 0 failed post-turn passes (§G) |
| M02 Restart during long campaign | **PASS** | campaign: 3 genuine process restarts, one carrying retained history, all identical (§G.2) |
| M03 Long-run context stability | **PASS** | campaign: the prompt held between 14.6k and 15.7k tokens for 70 turns, the window verified 101/101, the canon present on every turn (§G.3, §N) |
| M04 Long-run memory recall | **PASS** | campaign: the planting turn outside the window, the fact recovered through state and through memory of the narrator's restatements (§G.4) |
### SHOULD and FUTURE disposition
**SHOULD tests run and passing:** B04, D03, K04 (on its deferred branch), L04.
**FUTURE tests not run, and deliberately:** K05 (generate local image) and K06
(multi-turn video request). Both require a media provider, which §5 of the M11
brief forbids adding and M10 deliberately did not build. They are not v1
blockers and no part of this milestone treats them as one.
---
## G. The long-run campaign — M01 to M04
**M01 to M04 pass on a complete 100-turn campaign run by the committed
harness.** This section reports that run (§G.1-§G.5), then every earlier run
and why none of them is the evidence (§G.6).
### G.1 The evidence run
| | |
| --- | --- |
| **Tree** | commit `96c1bf5`, signed; `tools/m11_long_run.py --turns 100` |
| **Inference** | the GPU inference host (§E.1), `qwen2.5:3b-instruct-16k`, window 16,384 |
| **Evidence** | `$HOME/m11-evidence/m04-final/`: `timeline.jsonl`, `summary.json`, `recall.json`, `bundle.json`, `campaign.db`, `server.log`, `recovery-report.json` |
| **Status** | `complete`; no aborted reason, no failed reason |
| **Accepted turns** | **101**: 100 scheduled beats and the recall turn |
| **Genuine process restarts** | **3**, making 4 process starts |
| **Wall clock** | 1,238 s; a turn took 4.1 s at least, 9.6 s median, 22.8 s at most |
| **Memory bank and auto-summarise** | switched on at setup and read back as on |
| **Post-turn work** | checked after every turn against derived status and new `server.log` lines: **0** failed passes, 0 `database is locked`, 0 unrecorded failures, 0 tracebacks |
| **Written** | 33 memories in the database across every branch, 19 reported by the `/memories` endpoint at the end; 12 summaries |
| **Protocol in stored narration** | 0 of 104 AI turns |
**Timing against the GPU fault (§E.1).** The last AI action was written at
02:31:19 UTC, the memory, summary and embedding passes finished at 02:31:22,
and the GPU dropped off the bus at 02:31:52. The run's final settle and summary
followed; nothing they report was pending.
### G.2 Timeline of operations
"At turn" is the accepted-turn count when the operation fired.
| At turn | Operation | Outcome |
| --- | --- | --- |
| 0 | memory bank and auto-summarise | both read back as on |
| 0 | knowledge import | canon, reference and inspiration sources |
| 0 | state correction | 8 events, the opening cast |
| 1 | state correction, clue planted | the clue accepted as state; verified in state; the planting turn at depth 1 |
| 6 | Save Point | *Before the ridge* |
| 13 | restart 1 | `identical: true` |
| 20 | Undo, then Redo | 41 → 39 → 41 actions, position restored |
| 27 | Retry | two takes on the newest turn |
| 34 | Save Point | *On the ridge* |
| 41 | Undo, then restart 2 with the undone history retained | `identical: true`, Redo available on both sides |
| 48 | Retry | two takes on the newest turn |
| 56 | Undo, then new writing | Redo no longer available |
| 62 | take selection | index 0 of 2 chosen, and it became live |
| 69 | failed model call | reported: HTTP 404 for a model the server does not serve; state unchanged |
| 76 | restore *On the ridge* | active line 148 → 69 actions; the later story retained; Redo available |
| 83 | restart 3 | `identical: true` |
| 100 | recall check | §G.4 |
| 101 | export | 2,743,080 bytes; 0 AI turns carrying protocol |
### G.3 Context growth
The application's token counts, from `timeline.jsonl`.
```text
turn actions in prompt prompt summary memory knowl state canon floor window
1 3 3 1,441 0 0 127 93 57 - 16384
10 21 21 6,422 185 423 330 137 57 - 16384
20 41 41 11,447 603 544 230 161 57 - 16384
30 61 55 15,299 275 700 230 165 57 6 16384
40 81 57 15,300 289 809 230 194 57 24 16384
50 99 57 15,474 395 946 230 194 57 42 16384
60 115 55 14,718 374 875 127 194 57 60 16384
70 136 58 15,549 244 1,000 127 194 57 78 16384
80 77 59 15,203 289 664 127 164 57 18 16384
90 97 61 15,504 217 764 227 193 57 36 16384
100 117 63 15,745 146 757 127 193 57 54 16384
```
- budget 16,384 (configured 16,384); reply reserve 564 on every turn
- window verified on **101 of 101** turns
- campaign canon in the prompt on 101 of 101 turns; imported knowledge on 101; memories on 94; a summary on 93
- largest assembled prompt 15,745 tokens by the application's count; the history window was first trimmed at turn 30
- the drop in actions at turn 80 is the Save Point restore at turn 76
### G.4 Recall (M04)
**The precondition is positional.** M04's pass text is "Fact/event remains
recoverable without entire transcript in prompt", so the harness asks whether
the planting turn has left the history window. It does not ask whether the
clue's text has, because the narrator reuses that text in its own prose (below).
| | |
| --- | --- |
| Planting turn's depth | 1 |
| History window floor at the recall check | depth 54 |
| **Planting turn in the history window** | **no** |
| Clue's sentinel text somewhere in recent history | yes: the narrator's restatements, below |
| Clue in the narrative-state section | yes |
| Clue in the memories section | **yes** |
| Clue in the summary section | no |
| Clue in imported knowledge | no |
| Fact still in authoritative state | yes |
| History after the recall turn | 59 of 119 actions, in a 14,638-token prompt |
| **Verdict** | **`recovered_through_memory_or_summary`** |
**How the fact reached memory.** From depth 30 onward the narrator wrote the
state section's fact line into its prose as an ordinary sentence, on 74 of 104
AI turns:
> Aldric knew the silver key opened the crypt beneath it (SILVER-KEY-CRYPT-OLD-ABBEY).
The memory pass summarises narration, so the 16 memories that carry the clue all
summarise depth 54 or later. None comes from the planting era. The chain is
authoritative state, then the narrator's restatement, then memory. The fact
survived because state carried it. The repository owner accepted state-based
recovery as satisfying M04 on 2026-09-13.
A restatement sits inside a line of story, so the extractor rightly leaves it.
§P records what it means for narration quality.
### G.5 What the numbers say
**M01, 100-turn campaign: PASS.** 101 accepted turns with none refused. Every
one of the thirteen scheduled history operations performed as designed. All
three restarts compared transcript, head, Undo/Redo availability, scene,
entities, facts, Save Points, imported knowledge and settings and found them
identical. The induced failed call changed nothing. The campaign moved to a
clean data directory 16 of 16 (§K). The step list's summary/memory activation
was exercised: 33 memories and 12 summaries were written with no failed pass.
**M02, restart during a long campaign: PASS.** Three genuine `uvicorn` process
boundaries, one of them carrying undone history across. Each was identical.
**M03, long-run context stability: PASS.** The prompt reached about 15.3k tokens
at turn 30 and stayed between 14.6k and 15.7k for the next 70 turns, while the
active story grew to 136 actions. The canon was present on every turn, and the
reply reserve was subtracted on every turn. By the narrator's own tokenizer the
largest prompt plus the reserve is 16,342 of 16,384, which is bounded but nearly
full (§N).
**M04, long-run memory recall: PASS**, on §G.4's positional precondition, with
the recovery path stated there.
### G.6 The runs that led here
Every run's evidence directory is kept.
| # | Run | Tree | Host | Result | Why it is not the evidence |
| --- | --- | --- | --- | --- | --- |
| 1 | first release campaign, 2026-09-07 | first revision | CPU, 4,096 window | reached 97 of 100 turns | the host crashed; its evidence was under `/tmp` and the reboot cleared it. The harness now requires `--out` and checkpoints for `--resume` |
| 2 | 100 turns, 2026-09-10 | `ef25b0a` | CPU, 16,384 | complete: 101 turns, 3 restarts, 45,656 s | **the memory bank and auto-summarise were never switched on** (harness defect, §O), so summary/memory activation went untested and recall ran through state only |
| 3 | 26-turn trial, 2026-09-13 | `fec46f6` | GPU | reported `complete` | 180 `database is locked`, 20 failures that could not be recorded, 2 memories, 0 summaries. Found §O.7 and two harness defects |
| 4 | 26-turn trial | `f8d4010` | GPU | 0 lock errors, 7 memories, 2 summaries, 258 s | a trial, not a campaign |
| 5 | 100 turns | `f8d4010` | GPU | complete: 101 turns, 1,289 s, 33 memories, 12 summaries | the narrator's protocol was stored as story on 42 of 104 turns (§O.8), and a pasted state section kept the clue in recent history, so M04 proved nothing |
| 6 | 100 turns | `0c7316f` | GPU | complete: 101 turns, 1,251 s, 1 leak by the harness count | the narrator wrote `## Established:` sections of its own on 5 turns, which the extractor and the harness's leak count both missed; verdict `precondition_not_met` on the sentinel-text precondition. Reclassified on the positional precondition, in `recall-reclassified.json` beside the original, as `recovered_through_state_only`. Superseded by run 7 on the corrected tree |
| 7 | **the evidence run** | `96c1bf5` | GPU | §G.1-§G.5 | |
---
## H. Lineage leakage — E01 to E04
`tests/test_m11_leakage.py`, **14 tests**, all passing. What M11 adds to the
existing E-series coverage is that all four leaks are exercised **together, in
one campaign, under long-story conditions** — 22 turns on the abandoned line
with state, memories, summaries and knowledge all live, then 15 on the new one —
because the four share one mechanism and a campaign with only one of them cannot
show the mechanism holding for one and failing for another.
Four sentinels, one per class. Every negative control has a positive control
that fails loudly if the fixture did not actually establish the thing:
| | Positive control (path A) | Negative control (path B) |
| --- | --- | --- |
| **E01 state** | the fact is in the document, and in the prompt | absent from the document, absent from the prompt |
| **E02 memory** | the memory reaches path A's prompt | absent from the prompt and from `memories.used`; **still on disk**, because the story was left, not erased |
| **E03 summary** | a summary exists on path A and contains the sentinel | a **new** summary row was generated on path B; it carries nothing from A; the summariser was never *offered* A's summary; no A turn is on B's lineage; A's row is retained but ineligible |
| **E04 scene** | the abandoned line moved to the crypt, in state and in M10's packet | the current scene is the tavern; the protagonist's location followed the active line; the derived Scene Packet shows the active line and a different `scene_id`; the abandoned scene is still retained at its own position |
E03 is the one with history: M6's review found the first implementation passing
while the defect was live, because the test checked only that the old *row* was
ineligible. The shape this file uses is the one that review demanded — **the
summary is regenerated after the divergence** — and the assertion that matters
most is that the summariser's *input* never contained the abandoned prose. A
filter over the output would be a different bug.
E04's Scene Packet assertion cannot fail while the state assertion passes, since
the packet is derived from the state on read. It is asserted anyway, because the
packet is a surface that did not exist when E04 was written, and a later change
that gave it a store of its own would fail here.
---
## I. Fantasy and science fiction
**The fantasy Continuity Test** is the foundation of the 100-turn campaign (§G):
Westhaven, the Crooked Lantern, Aldric, Mara, Edrin, the silver key, the sealed
abbey crypt, and the canon that the dead do not return.
**The Persephone science-fiction fixture** — `tests/test_m11_scifi.py`, **10
tests**, all passing — is `TEST-CAMPAIGN-FIXTURE.md` §31's, with its three
hard-technology canon rules and its full cast.
What is actually being checked is not that a science-fiction story can be told,
but that **no code path knows the difference**:
| Check | Result |
| --- | --- |
| Five entity types in one document — character, vehicle, location, item, organization | all present, all through the same `create_entity` |
| The type list is *suggested*, not closed | `SUGGESTED_TYPES` — a genre needing a type nobody listed uses one without a migration |
| Where a thing is lives on the entity | the same field puts Aldric in a tavern and Imani aboard a ship |
| Possession | the data crystal is Imani's, through the same `set_possession` |
| Canon reaches the prompt as the campaign's highest authority | "FTL does not exist" in the `campaign_canon` section |
| M10's Scene Packet | describes a starship under spin with no field it did not already have |
| A visual profile | holds `hull: pitted white composite` as readily as a face |
| Reference retrieval | a science-fiction query retrieves the spin-gravity passage |
| The bundle | same `ai-dnd-adventure-v3`, vehicle type intact on the far side |
| The event vocabulary | contains no genre noun — no spell, no sword, no warp, no airlock |
**No schema change, no code change, no new event type** was required for the
science-fiction fixture. J03's claim — genre is configuration — holds in the
strong form: the same schema, the same validator, the same builder, the same
bundle.
---
## J. Offline and security
### The offline run
`tools/m11_offline.py` — **23 checks, 0 failed**. A container built with
`docker build --no-cache` and run with `--network none`: a loopback interface and
nothing else, no resolver, no route, and a fresh volume. The exercise runs inside
over `docker exec`, because with no network there is no published port to reach —
that is the only honest way to drive an isolated process.
`unshare -rn` was the first choice and is unavailable here: Ubuntu 24.04 sets
`kernel.apparmor_restrict_unprivileged_userns=1`. The container gives the same
isolation and doubles as §24's packaging evidence.
| | |
| --- | --- |
| **The isolation is real** | a TCP connection to 1.1.1.1 fails; `getaddrinfo("example.com")` fails |
| First page load, from fresh data | succeeds; names no remote origin; a CSP is served |
| Every asset the shell references | served locally — none remote, none missing |
| Campaign creation, state extraction | work |
| Knowledge import, prompt assembly, retrieval | work |
| A turn with **no model reachable** | reported as a failure; **no narration accepted**; the player's own words kept (A05); state unchanged; earlier story still there |
| Export and import | work; no secret in the bundle |
| M10's media module | imports; registry empty; a scene packet builds; no provider required |
| Container restart | campaigns survive on the volume |
**Inference is not exercised offline, and that is stated rather than implied.**
This deployment's Ollama is on the trusted LAN, which `SECURITY-THREAT-MODEL.md`
§73 permits and which is not an Internet dependency — but it is also unreachable
from a container with no network. What the offline run proves about inference is
the useful half: with no model reachable the application degrades to a reported
error and the campaign stays intact.
### Observed outbound destinations
During the first revision's long campaign, on the CPU reference host, the
application contacted exactly one host: the configured trusted-LAN Ollama, over
HTTPS with a private CA in the OS trust store. It used the `/v1` path for
inference and the `/api/ps` and `/api/show` paths for the window probe. There was
no other destination, and no DNS lookup for any other name. In the container run
there was no destination at all, because there was no network.
**The later long runs were not network-monitored.** They were configured with one
endpoint, the GPU inference host over plain HTTP. `tests/test_egress.py` is what
continues to assert that nothing else is contacted.
### H-series
Full per-test results are in §F. The M11-specific additions:
- **H09 is NOT APPLICABLE, and the condition is now enforced.** Its own text
makes it conditional on ZIP import/export existing. Nothing in the application
opens an archive — the bundle is JSON, an imported source is a single file —
and `test_h09_the_product_extracts_no_archives` scans every module for
`zipfile`, `tarfile`, `unpack_archive`, `py7zr` and `rarfile`, so the day that
stops being true H09 becomes required again.
- **H10's startup refusal is proved by starting a process**, not by importing a
module: `AIDND_CORS_ORIGINS=*` makes the application refuse to come up, and a
named origin is accepted, which is the control that makes the first assertion
about the wildcard rather than about the variable.
- **H12's tampering case** writes a cloud endpoint into the settings row behind
the API, and the request-time check still refuses it — ADR 011's point being
that the check is not only at the front door.
- **The M11 window probe is held to the same policy**, verified with no
transport installed, so a probe that ignored the policy would attempt a real
connection and be caught.
- **Hostile content in the browser** (§L): an `onerror` image, a `<script>` tag,
a `javascript:` link, a remote image and shell text all reached the real
renderer as accepted narration. None executed, none loaded, none became markup.
---
## K. Recovery — export, import, migration
### The long campaign, moved to a machine that has never seen it
`tools/m11_recovery.py`: **16 checks, 0 failed**. The input is not a fixture. It
is the evidence run's own campaign (§G.1), exported through the API and imported
into a database file that did not exist, in a directory that did not exist,
opened by a second server process. Migrations ran there from nothing, so this is
the fresh-install path as well as the import path.
| | |
| --- | --- |
| Bundle | 2,743,080 bytes, **13.1% of the 20 MB import limit**, at 207 actions |
| Actions in the bundle / on the active line after import | 207 / 119 |
| Retained beyond the active line | 88 |
| Entities, facts | 6, 2 |
| Save Points | 2; *On the ridge* restores to 69 actions and *Before the ridge* to 13, the positions they name |
| Knowledge sources | 3: canon, reference and inspiration, with content, not just filenames |
| Narration-length choice | came across |
| Campaign canon | came across |
| Redo | available exactly when the file said so |
| Secrets in the bundle | none |
| The moved campaign | accepts a new correction with nothing refused, and exports again at the same story length, 207 actions |
The first revision's recovery run moved a 41-turn campaign at the 4,096 window
(654,803 bytes, 83 actions), and that campaign's files were lost. The run above
replaces it.
**On the bundle ceiling.** M9 estimated about 279 turns against the 20 MB limit,
using a fixture built to be heavy. This real campaign plays at the recommended
16,384 window, with every stored prompt near that size. It averages about 13 kB
per action, which puts the limit near 1,600 actions. M9's number remains the
conservative one to quote.
### Migration
`tests/test_m11_migration.py` — **9 tests**, all passing, and the first of them is
the permanent form of the defect M10 found by accident:
| Check | Result |
| --- | --- |
| **A fresh install and an upgraded M10 database produce the same schema** | identical — every table, every column with type and nullability, every index with its columns and uniqueness, every foreign key, the primary keys, and the version stamp |
| No table carries two indexes over the same columns | none does |
| Every table the models declare exists | including `visual_profiles` and the new `narration_length` column |
| An M10-era database (stamped 92, no `visual_profiles`) opened by this build | gains the table; campaign, action and Save Point intact; `foreign_key_check` empty; `quick_check` ok |
| The new column arrives as `""` | which means "this campaign never chose", so no existing prompt changes under the upgrade |
| Opening the database repeatedly | index set and version identical after each |
| A migrated database still plays and still travels | played through a real server process, corrected, exported and re-imported |
| A backup of the migrated database | verifies, and carries the new column |
| A fresh install from nothing | creates the database, stamps the current version, and has every expected table |
The comparison is written generically rather than about `visual_profiles`, so a
future migration that diverges the two paths fails here whatever it is about.
**M11's own migrations** are numbers 93 and 94: `adventures.narration_length` in the
first revision and `settings.context_window_override` in `ef25b0a`, one column
each, no backfill.
The knowledge-migration suite's version assertions were corrected from a literal
`== 92` to `== LATEST_VERSION >= M7_VERSION`, because they had been asserting
that M7's version was the newest — true when written, and a statement about M11
rather than about M7 once M11 added a migration (§O).
---
## L. Browser
**Firefox 154.0.1**, headless, driven over W3C WebDriver (geckodriver 0.37.1),
against the **built** SPA served by FastAPI — the production path from
`DEVELOPMENT.md`, not a Vite dev server. Narrator: `qwen2.5:3b-instruct` on
the trusted-LAN Ollama. *Corrected at closeout:* earlier revisions named the
`-16k` model, and the run's own report records the plain one. **38 checks, 0
failed, 0 skipped**, 107 seconds, on the `ef25b0a` tree. The closeout's runs on
the candidate tree are in §S.3.
The harness is `tools/m11_browser.py` and its WebDriver client is
`tools/m11_webdriver.py`, both in the repository — M8's and M9's browser harness
lived outside it, which made their browser evidence unrepeatable by anyone else.
No Selenium: WebDriver is an HTTP protocol and `urllib` speaks HTTP, so browser
evidence adds nothing to the dependency surface.
| Area | Checks |
| --- | --- |
| **B01 narration** | two real turns accepted through the real engine |
| **A/UX (finding A)** | the tab carries no inherited name; names the product; names the open campaign |
| **B (finding B, §8A)** | the position is shown; it visibly changes after Undo; it says when later story is available |
| **D01/D04** | Undo offered; Redo becomes available after Undo; Redo returns to where the reader was |
| **H06** | an `onerror` attribute never executes; a `<script>` in narration never executes; markup in the source is not markup in the page |
| **H07** | a `javascript:` URL never becomes an href |
| **G09** | a remote Markdown image is not loaded |
| **H04** | shell text in narration is text |
| **G01** | a local file imports **through the real file input**; Import is enabled only once a file is chosen |
| **§38** | narrator-only text is absent from the DOM, not merely hidden |
| **F05** | the context inspector shows the assembled prompt and its budget |
| **H10** | an unknown API path is a 404 with a non-HTML body; a page path is the SPA |
| **CSP** | a policy is served and names no remote origin |
| **A11y** | every visible control has an accessible name; focus is visible; no positive tabindex; nothing revealed only on hover; the story input takes keyboard focus |
| **A11y modal** | the dialog takes focus, has an accessible name, contains something focusable, and Escape closes it |
| **A11y contrast** | measured on rendered colours (below) |
### Accessibility, measured rather than eyeballed
M8 recorded contrast and visible focus as checked by eye and handed the
measurement to M11. Both halves were done.
**Rendered contrast**, computed in the page from the actual colours after
inheritance and layering, with the WCAG 2.1 formula:
```text
story prose 14.57:1 at 18.25px (needs 4.5:1)
control 5.48:1 at 12.48px (needs 4.5:1)
input 13.57:1 at 16.81px (needs 4.5:1)
position 5.88:1 at 12.48px (needs 4.5:1)
```
**The palette**, by calculation (`tools/contrast_audit.py`): every text pair the
design uses clears WCAG AA 1.4.3, the lowest being an error message at 5.02:1.
**Two boundary pairs are below 1.4.11's 3:1** — a control's resting edge at
1.33:1 and its hover edge at 1.75:1 — and they are **reported, not fixed**.
1.4.11 applies to the visual information *required to identify* a component, and
in this design that is the control's text label, which is measured at 5.48:1 and
passes. Restyling the palette would be a design change made inside a
release-validation milestone to satisfy a threshold the reader is not affected
by. It is recorded here so a reviewer can disagree.
**What was measured versus inspected.** Measured: contrast (both ways),
accessible names, focus visibility, tab order, hover-only revelation, modal focus
and dismissal, keyboard reachability of the story input. Inspected by reading
rather than measured: reading order beyond tabindex, and screen-reader
announcement quality. No claim is made about tablet layout beyond what the
specification promises.
### The limitation that remains
A file can now be driven **into** the browser (the snap sandbox accepts a path
under `$HOME`, which is what M9's residual risk 6 had recorded as impossible).
Driving one **out** — a `blob:` download from the export control — still does not
complete under this headless snap Firefox. Export is proved end-to-end without a
browser (§K) and the browser's own export control is exercised only as far as
the click.
---
## M. Dependencies and packaging
### The runtime dependency surface
**Nothing was added, removed or upgraded by M11.** `requirements.txt`,
`requirements.lock` and `package.json` are byte-identical to M10's. The browser
harness deliberately speaks WebDriver over `urllib` rather than adding Selenium,
because a package added to press buttons would still be a package in the audit
surface.
| | |
| --- | --- |
| **npm audit** | **0 vulnerabilities**, both with dev dependencies and with `--omit=dev` |
| **Production npm dependencies** | three: `react`, `react-dom`, `react-router-dom`. Everything else is `devDependencies` |
| **pip check** | no broken requirements |
| **Lock consistency** | 35 pinned entries, none missing from the environment, **zero version drift** |
| **What ships** | the image installs from `requirements.txt` only: **31 packages** |
| **Cloud/auth/analytics packages** | none. No `openai`, no `anthropic`, no telemetry SDK, no auth library |
| **Runtime CDN or font references** | none — the fonts are vendored and `test_offline_assets.py` fails if a remote origin returns |
| **New network client from M10 or M11** | none. `httpx` was already the only HTTP client; `contextwindow.py` uses it with the shared TLS context and the shared endpoint policy |
| **Media implementation dependency** | none — no ComfyUI, diffusers, Whisper, Kokoro, or model download |
**One observation, not a defect.** The *developer venv* carries three packages
the lock does not pin and the image does not install: `quickjs` (upstream's
scripting engine, removed in M2), `psycopg`/`psycopg-binary` (the hosted
deployment, removed in M2) and `cryptography`. They are residue in a long-lived
local environment, not shipped: the image has none of them, and
`test_local_only_surface.py` already fails if any module imports `quickjs`. A
fresh `pip install -r requirements.txt` produces the 31-package set.
**Unresolved advisories:** none reported by either audit.
### Packaging
Every production-shaped path the repository claims:
| Path | Result |
| --- | --- |
| `uvicorn app.main:app --host 127.0.0.1 --port 8000` | the path the 100-turn run and the browser run both used; started 5 times across M01's restarts |
| SPA served by FastAPI | the browser regression ran entirely against the built `frontend/dist`, served by the backend |
| `docker build --no-cache` | clean; log in the offline evidence directory |
| Container startup | serves with `--network none` |
| Persistence across container restart | campaigns survive on the volume |
| Loopback publication | `docker-compose.yml` publishes `127.0.0.1:8000:8000`; `test_local_only_surface.py` asserts it |
| Vendored assets | fonts and the tokenizer table are in the image; the offline run fetched every asset the shell references from the container itself |
| No dependency on development source | the image contains `backend/app` and `frontend/dist` only — no tests, no tools, no `node_modules` |
**No release was created and no tag exists.**
---
## N. Performance and storage
Measurements, not requirements. The planning package sets no performance target
and none is invented here.
### Long-run storage
The evidence run, §G.1:
| | |
| --- | --- |
| Database after 101 accepted turns: 207 actions across the active line and retained history, 33 memories, 12 summaries, 3 imported sources | **2,367,488 bytes** |
| Export of the same campaign | **2,743,080 bytes**, 13.1% of the 20 MB import limit |
| Per action, roughly | ~11 kB in the database and ~13 kB in the export, dominated by the per-position state snapshot and a stored prompt of up to ~16k tokens |
No live backup was taken during this run. The first revision's backup figure
(160 pages, `integrity: ok`) came from the lost campaign.
### Prompt size across the run
The application's count:
```text
turn 1 1,441 tokens 3 of 3 actions in the prompt
turn 10 6,422 tokens 21 of 21
turn 20 11,447 tokens 41 of 41
turn 30 15,299 tokens 55 of 61
turn 40 15,300 tokens 57 of 81
turn 50 15,474 tokens 57 of 99
turn 60 14,718 tokens 55 of 115
turn 70 15,549 tokens 58 of 136
turn 80 15,203 tokens 59 of 77 (after the Save Point restore)
turn 90 15,504 tokens 61 of 97
turn 100 15,745 tokens 63 of 117
```
The prompt rises until the history window is full and then **stops**. That is
F03's claim measured rather than argued. The window's contents keep moving,
always to the newest turns, while its size stays put.
### Real-token headroom
The application counts tokens with `cl100k_base`, and the narrator counts with
Qwen's own tokenizer. Each run's largest stored prompts were re-sent to the same
model, identified by digest, and its `prompt_eval_count` was read:
| Run | Largest prompt, narrator's count | Plus reply reserve 564 | Headroom in 16,384 |
| --- | --- | --- | --- |
| CPU host, memory off (`ef25b0a`) | 15,728 | 16,292 | 92 |
| GPU host, 100 turns (`f8d4010`) | 15,797 | 16,361 | **23** |
| GPU host, 100 turns (`0c7316f`) | 15,786 | 16,350 | 34 |
| **GPU host, evidence run (`96c1bf5`)** | **15,778** | **16,342** | **42** |
**No prompt in any run exceeded the window.** On every prompt re-counted, the
application's count was 16 tokens below the narrator's. The evidence run's
re-count was taken after the GPU host's reboot, with the model digest verified
unchanged.
The margin matters because of how the server fails. Ollama 0.34 cuts a prompt
longer than the window down to 8,194 tokens and returns no error. That was
observed with synthetic prompts, not in a campaign. §P records it as a risk.
### Inference cost, by host
| Host | Window | Prompt at steady state | Seconds per turn | 100 turns |
| --- | --- | --- | --- | --- |
| CPU reference host | 4,096 | ~3.4k tokens | 88-203, median 178 | ~4 h projected; that run was lost at 97 |
| CPU reference host | 16,384 | 13-14k tokens | 229-291 in the first attempt | 45,656 s for 101 turns, with memory off |
| GPU inference host | 16,384 | 14.6-15.7k tokens | 4.1-22.8, median 9.6 | 1,238 s for 101 turns, with memory on |
Prompt processing dominates on the CPU host. On the GPU host a full 16k prompt is
processed in about ten seconds (§E.1). Both are these machines' characteristics.
### Query behaviour
No new query growth was introduced. M11 adds one HTTP round trip per session per
`(endpoint, model)` pair, the window probe, cached for ten minutes on success and
one minute on failure. It adds one derived computation per state read,
`duplicate_names`, a single pass over the entities already in memory. The
connection test was changed to use that cache after the browser run showed the
model-status badge calling it on every page load.
The §O.7 correction moves one statement rather than adding one: the memory
use-counter UPDATE now runs inside the turn's existing commit instead of before
the model call.
### Nothing pathological was found
No unbounded growth, no per-row query, no repeated snapshot write, no duplicated
knowledge content. The M10 measurement tool (`tools/m10_media_cost.py`) remains
valid: a scene packet is four SQL statements at any campaign length.
---
## O. Findings
Product defects first, then defects in the tests and harnesses, which are kept
separate because conflating them is how a milestone reports confidence it has
not earned.
### Product defects — found by M11, fixed in M11
**O.1 — The application silently budgeted more input than the server would read**
**Severity: high. Requirement: F03, F04, M03, and the honesty of every long-run
claim. Blocker: yes — it was M11's stated blocker. Status: fixed.**
*Reproduction:* configure a model with no `num_ctx` on a server with no VRAM
(`/api/ps` reports 4,096); leave `context_token_budget` at its 16,384 default;
play a long campaign. Every request returns 200 and the server drops the oldest
tokens — the narrator's rules and the campaign canon.
*Root cause:* the window is a property of the model load, not of the request,
and the application had no way to learn it. §E in full.
*Correction:* `app/contextwindow.py` discovers it and the builder caps to it, or
records the turn as unverified.
*Regression evidence:* `tests/test_m11_context_window.py` (20 tests) including
the truncation sentinel and its negative control — the same campaign built
without the cap, measured at more than twice the window. Plus
`tests/test_m11_real_window.py` against a real Ollama.
**O.2 — A partly refused manual state correction reported success**
**Severity: medium. Requirement: C04, and `SECURITY-THREAT-MODEL.md` §69
auditability. Blocker: no. Status: fixed.**
*Reproduction:* `POST /state/corrections` with two changes, one naming a
location entity that does not exist. Before M11: **HTTP 201**, the good change
applied, the bad one silently dropped, nothing in the response to say so.
*Root cause:* the handler raised 400 only when *nothing* was accepted. Partial
acceptance is deliberate and correct — `validate.py` argues that discarding three
good changes because of one typo is worse — but the same file states the rule
this broke: "what is never allowed is a rejected event mutating anything, or **a
rejection being silent**". The refusal was recorded on the proposal row for the
audit trail; the person who wrote it was simply never told.
*How it was found:* the identity diagnostic's own fixture set a scene naming a
location entity it had not created. The event was refused, the 201 said nothing,
and **the entire first diagnostic run happened on a campaign with no scene and no
list of who was in the room** — a degraded context that could easily have been
read as a model failure. That run's evidence was discarded.
*Correction:* the correction response carries `refused` (event, reason, detail),
and the State panel shows it. Behaviour is otherwise unchanged.
*Regression evidence:* `test_a_partly_refused_correction_reports_what_did_not_apply`,
with controls for the fully-applied and wholly-refused cases and a check that the
audit record still records `partially_accepted`.
**O.3 — The narration-length setting moved no number** (post-M8 finding C)
**Severity: medium. Requirement: the setting's own promise; B-series narration
quality. Blocker: no. Status: fixed.**
*Reproduction:* create three campaigns choosing brief, medium and long; compare
the `length_hint` section of the stored prompts. Before M11 they were identical —
at the default cap, "must not exceed 506 words, and it should not stop short of
about 177" in all three.
*Root cause:* the choice became one English sentence in `ai_instructions` and the
numeric hint was derived from the *global* `max_output_tokens`.
*Correction:* the choice is data on the campaign; `length_hint` maps it to a word
band bounded by the reply cap. The generation budget is deliberately untouched.
*Regression evidence:* `tests/test_m11_findings.py`, including
`test_the_three_lengths_no_longer_say_the_same_thing` (fails against the old
behaviour) and `test_the_stored_prompt_carries_the_campaigns_own_range` end to
end.
### Product gaps closed without a defect
**O.4 — The tab carried the inherited name** (post-M8 finding A). Not a false
claim by any document, so a gap rather than a defect. Fixed.
**O.5 — No orientation after history movement** (post-M8 finding B). The
requirement it violated did not exist until §8A was written after the playtest.
Fixed to that requirement.
**O.6 — Two entities could share a display name silently** (post-M8 finding D's
structural half). Deliberately **not** made an error: two people called Alice is
ordinary fiction. Made visible instead — `duplicate_names` in the state API and
the State panel.
### Product defects — found by the long runs after the first revision, fixed
**O.7 — A turn held SQLite's write lock through the model call, and post-turn
memory and summary work was lost without a trace**
**Severity: high. Requirement: M01's summary/memory activation, F02, F08, M04.
Blocker: yes, for M01. Status: fixed in `f8d4010`.**
*Reproduction:* turn the memory bank and auto-summarise on, configure an
embedding model, and play consecutive turns once a memory exists. On a fast
inference host: `database is locked` in the server log, memories stop
accumulating, no summary is written, and derived status reads `idle` with no
failures. The first 26-turn trial (§G.6, run 3) logged 180 such errors, wrote 2
memories and no summary, and reported `complete`.
*Root cause:* `retrieve_memories(update_stats=True)` ran an uncommitted UPDATE of
the used memories' counters before the model call. The turn commits once, after
the reply has streamed, so the write transaction stayed open for the whole reply.
SQLite has one writer. Every post-turn memory, summary and status write in that
window waited out the driver's five-second timeout and failed. Recording the
failure needs a write as well. The post-turn task's outer handler recorded
without rolling back first, so it raised `PendingRollbackError` and the failure
reached only the log. F08 requires a memory failure to be visible, and this one
was not. No earlier long run hit it, because none had the memory bank on.
*Correction:* retrieval only reads. `memorybank.record_use` writes the counters in
the turn's single commit, so a turn that never lands counts nothing. The outer
handler rolls back before it records.
*Regression evidence:* `test_no_write_lock_is_held_while_the_narrator_is_talking`
probes for the lock from a second connection during the model call.
`test_a_failure_that_breaks_the_session_is_still_recorded` covers the recorder.
**Both fail on `fec46f6`**, with `database is locked` and `idle` respectively.
Also `test_a_failed_turn_counts_no_memory_as_used`. The next trial had 0 lock
errors, 7 memories and 2 summaries.
**O.8 — The narrator's protocol was stored as story**
**Severity: high. Requirement: story authority (M5 review Finding 4), C04, and
the validity of M04. Blocker: yes, for M04's evidence. Status: fixed in `0c7316f`
and `96c1bf5`.**
*Reproduction:* play a long campaign against a small local model, then search the
stored AI turns for the state section's headings or `"events"`. In the `f8d4010`
run 42 of 104 turns carried protocol, the first at depth 2, in four shapes:
- a copy of the narrative-state section: `Scene:`, `Who and what exists:`, `Held:`, `Established:`, `Still open:`
- that copy above a correct ```` ```state ```` block, which was removed while the copy stayed
- the copy, a bare `State` heading, and a `> {"events": ...}` proposal quoted like a player turn, sometimes with story after it
- the same block cut off by the output-token limit, on 10 turns
In the next run the narrator wrote sections of its own instead, such as
`## Established:` over indented facts, on 5 turns.
*Root cause:* the extractor removed fenced blocks, a bare object at the very end,
and a parroted bracketed reminder. It did not remove a pasted state section, an
unfenced proposal elsewhere in the reply, or an unfinished one. Stored text is
replayed verbatim as history (`_history_text`). Each leak therefore put a second,
older account of the state into the next prompt, which is what Finding 4
removed from replay, and it gave the model another example to copy. A second,
older bug was in the fence pattern itself. `_STATE_FENCE_RE` read "a ```state
block" inside a parroted reminder as a fence opening and cut out the middle of
the reminder.
*Correction:* the extractor recognises the renderer's own section headings, which
are now named constants in `render.py`, with any markdown wrapped around them. It
removes a block carrying two headings, or one heading with an indented entry. It
removes an unfenced proposal that starts a line, quoted or not, taking outermost
objects first, and uses it as the turn's proposal when there is no fence. It
removes an unfinished proposal at the end and whatever is left behind at the end:
a `State` heading, a bare `>`, a parroted reminder or continue hint, closed or
not, and a ```json fence cut off before it names its events. The fence label must
now end its line or run straight into the payload. A reply whose only removal is
a pasted section records no raw block, so the turn is not marked unparseable for
a block it never started.
*Regression evidence:* 25 new cases, from 18 test functions, in
`test_narrative_state.py`, cut down from
the runs' real output, including negative controls. A lone `Held:` with prose
after it, a `Scene:` line of story, quoted JSON that is not a proposal, and a
`State` line followed by story all stay. A test renders every section the
renderer writes and pastes the lot. Beyond the suite, **every AI turn in five real
runs was replayed through the new extractor, 443 in all, and no turn the old
extractor had left clean changed.** The evidence run stored 0 of 104 turns with
protocol.
*Not removed, deliberately:* a restatement of a fact inside a line of story (§G.4),
and headings a model invents that are not the renderer's (`Identifiers
established:`, on 10 turns of run 3). Removing either means judging prose, and
§P records both.
### Harness and test defects — found and fixed, no product change
Recorded separately, and at this length, because M8's review found five harness
defects against seven product defects and two of the five were *masking* product
defects. A harness that has only ever agreed with itself is not evidence.
| | What it did | Why it mattered |
| --- | --- | --- |
| **The identity fixture set a scene naming an uncreated entity** | the scene never existed, and the diagnostic ran on a degraded campaign | it would have been read as a model failure. It also surfaced product defect O.2 |
| **The long-run harness read the turn endpoint as JSON** | crashed on the SSE body | a harness that read the status code instead would have called every failed turn a success — `sse.py` says a failed turn is a 200 with an error *inside the stream* |
| **The same harness called `retry` as JSON** | crashed at the fourth scheduled step | found by a 14-turn shakeout run rather than 50 turns into the release campaign, which is what the shakeout was for |
| **The offline check compared total action counts after a failed turn** | reported corruption where the product was behaving as designed | A05 deliberately keeps the player's submitted text and the head sits on it. The check now asserts the real contract: no narration accepted, state unchanged, earlier story reachable |
| **The browser import scenario never pressed Import** | choosing a file only stages it | looked exactly like a broken import |
| **The browser modal check used the Save Point control** | that opens a panel, not a dialog, so the check skipped itself | a skip that reports nothing is worse than a failure |
| **A `set_scene` bounds test asked for 43 present labels** *(M10, recorded again here)* | the state model correctly refused the event | the test measured the wrong scene |
| **Two substring checks matched inside words** | "ahead" contains "head"; "immediately" contains "media" | both now match whole words or parse imports |
| **A contrast check treated a control boundary as body text** | would have failed the run on a WCAG clause that does not apply | now distinguishes 1.4.3 from 1.4.11 and says which |
| **An audit assertion read a deferred column after the session closed** | `DetachedInstanceError` | read inside the session |
| **The long-run harness never switched the memory bank or auto-summarise on** *(fixed in `fec46f6`)* | both are per-campaign and default to off, so a complete 100-turn run wrote no memory and no summary | M01's summary/memory activation clause was reported by silence, and M04's memory path was never asked. The harness now switches both on, reads them back, and refuses a run with no embedding model |
| **It read three prompt sections under names the builder does not use** *(fixed in `f8d4010`)* | `memories` (really `used_memories`), `story_history` (really `history` and `recent_history`), and a `knowledge` prefix that matched the fixed instruction section instead of the imported passages | memory tokens read 0 whatever the prompt held, and the in-history and in-memories recall checks could never be true. The labels are now constants pinned by a test against a prompt the real builder assembled |
| **It had no way to see failed post-turn work** *(fixed in `f8d4010`)* | it reported `complete` over 180 `database is locked` errors | it would have certified the run that §O.7 destroyed. It now checks derived status and new `server.log` lines after every turn and stops at the first failure; a run with no memories or no summaries ends `failed` |
| **Its protocol-leak count used plain substrings** *(fixed in `96c1bf5`)* | `## Established:` did not match `\nEstablished:\n`, so the count read 1 where 5 turns leaked | the count now matches any state heading with an indented entry, in any markdown |
| **Its M04 precondition was the clue's text being out of recent history** *(fixed in `96c1bf5`)* | the narrator reuses that text in its own prose, so a run whose planting turn was 65 depths outside the window read `precondition_not_met` | the precondition is now the planting turn's position, recorded at planting and carried across `--resume`, which is M04's own wording |
| **The browser import wait could not fail** *(fixed at closeout)* | it waited for "hidden" in the page text, and the scenario's campaign is titled *Hidden Knowledge*, so the wait held before Import was pressed. The modal check that follows then raced the source row, and skipped once (34/0/1) | "a local file imports through the browser" was recorded without evidence. The harness now waits for the imported source's own row and records its title (§S.3) |
### Discarded evidence runs
Per §7, evidence taken before a product change was discarded rather than
reported:
1. **The first identity diagnostic run** — void, because its own fixture had been
refused (O.2). Re-run after the fixture and the product were fixed.
2. **The first browser regression run** — 31 passed, 1 harness failure, 1 skip.
Discarded and re-run after the harness fixes and the connection-test caching
change; the reported run is the second: 38 passed, 0 failed, 0 skipped.
3. **The first offline run** — 20 passed, 1 failure that was the harness asserting
the wrong contract. Discarded and re-run: 23 passed, 0 failed.
4. **The first 100-turn campaign run** — abandoned at 3 turns when the
connection-test caching change landed, so that the reported campaign runs
entirely on the final tree.
5. **Two harness shakeout runs** (6 and 14 turns) — never reported as evidence;
their purpose was to find the two SSE defects above.
6. **The first 100-turn release campaign**, which reached 97 of 100 turns and was
lost with its evidence to a host crash, because it wrote under `/tmp`. The
first revision's §G figures came from it and can no longer be checked. They
are replaced, not repeated.
7. **The memory-off 100-turn run** (`ef25b0a`). It is complete and internally
consistent, but it never exercised summary/memory activation. It is kept as the
CPU-host timing and headroom record (§N).
8. **The 100-turn runs on `f8d4010` and `0c7316f`.** They are complete, and they
are superseded because their M04 evidence was contaminated by protocol leaks
(§O.8). The `0c7316f` run's reclassified verdict is kept beside its original
and is not reported as the result.
---
## P. Residual risks
Genuine remaining risk and debt only. There is no M12; everything below is either
accepted for v1, or a decision for the owner at acceptance.
1. **The release evidence carries two hosts' speeds.** The CPU reference host took
minutes a turn, and the GPU host took seconds. Nothing in the planning package
sets a performance requirement, and none is invented here. Read §N's timings
as these machines', not as a product characteristic.
2. **The narrator is a 3B model.** Every realistic-model observation is that
model's: state extraction quality, narration length adherence, identity
handling, and what it restates. A stronger local model would behave
differently, probably better. The application's guarantees are deliberately
independent of which, since what is asserted is that the application stays
correct whatever the model proposes.
3. **The 16k window is nearly full.** By the narrator's own tokenizer the largest
prompts plus the reply reserve left 23, 34 and 42 tokens of headroom in the
three GPU runs (§N). The application's `cl100k_base` count ran 16 tokens low
on every prompt measured. The inference server cuts an over-window prompt to
8,194 tokens with no error. No prompt overflowed. A model whose tokenizer
diverges further from `cl100k_base`, or a larger reply reserve, could overflow
without anyone noticing. A deployment that wants margin can lower
`context_token_budget` below the window.
4. **The narrator restates prompt text in its prose.** On 74 of 104 turns of the
evidence run it wrote the state's fact line as a sentence, and it also echoed
phrases such as "Scene set, continue your adventure." and lines opening
"Memory:". These sit inside story, so the extractor leaves them. They cost
narration quality, and they are the route by which M04's fact reached memory.
5. **Memory does not keep a planted fact on its own with this summariser.** In no
run did a memory summarising the planting era carry the fact. Recall rests on
authoritative state, which the repository owner accepted for M04 on
2026-09-13. A reviewer who reads F02 or M04 as requiring memory retention in
its own right should read them as unproven.
6. **Model-invented headings are not removed.** A model that writes its own
section, such as `Identifiers established:` or `Set of events made true:`,
under a heading that is not the renderer's, keeps it in the story. Seen on 10
turns of one 26-turn trial. **Widened at closeout.** The identity run found
three more shapes left in stored story, on 4 of 10 turns at a 4,096 window:
- event-call syntax copied from the state rule, such as
`> set_possession(silver-key, "alice")`
- a parroted length hint, `[Hard limit: … append the state block well inside
the limit.]`
- a lone `Scene:` line
None occurs in the 100-turn evidence runs. The owner chose on 2026-09-14 to
carry this as a residual risk rather than change product code (§S.6).
7. **The GPU inference host dropped its GPU after the evidence run** (§E.1). The
cause is not established, and power transients at the uncapped 280 W limit
are the leading candidate. The evidence is unaffected. Future long runs must
log power, link state and kernel messages (`DEVELOPMENT.md`).
8. **Post-M8 finding D's root cause is unestablished and will stay that way.**
The campaign that produced it was destroyed. M11 delivers a diagnostic that
can classify the next occurrence, and the detection the finding asked for.
Its closeout run's results, and what that run still cannot establish, are
in §S.5.
9. *Closed at closeout.* **The browser, offline and identity runs predated the
last three product commits.** They had been re-run on the `ef25b0a` tree:
browser 38/0/0, offline 23/0, and the identity diagnostic clean, with the
evidence dated 2026-09-10. `f8d4010`, `0c7316f` and `96c1bf5` changed backend
behaviour. All three runs were repeated on the release-candidate tree on
2026-09-14 (§S).
10. **Two control-boundary colour pairs are below WCAG 1.4.11** (1.33:1 resting,
1.75:1 hover). Reported rather than fixed, because the control is identified
by its label, measured at 5.48:1, and restyling the palette inside a
release-validation milestone would be the wrong kind of change. An owner who
disagrees has the measurement.
11. **A file cannot be driven *out* of this headless snap Firefox.** Export is
proved end to end without a browser; the browser's own export control is
exercised only as far as the click. Import is fully proved.
12. **The bundle ceiling is unchanged.** M9 measured about 279 turns against the
20 MB import limit, and §K puts the real 100-turn campaign at 13.1% of it.
Beyond the ceiling a campaign can still be exported and would be refused on
import, which is the asymmetry worth knowing.
13. **The media seam has no real adapter.** This is M10's own residual risk,
unchanged: the contracts are shaped by the contract document rather than by
an adapter that had to work.
14. **`quick_check` rather than `integrity_check` on a backup**, and **no
scheduled backup**. Both are M9's, both unchanged, both outside the acceptance
contract.
15. **A campaign with no narration-length choice keeps the pre-M11 hint.** That is
deliberate, because an empty value means the reader never chose. It means an
existing campaign does not benefit from finding C's fix until someone sets
the control.
16. **The state rule's example is fantasy.** The fixed instruction every campaign
receives names `mara`, `silver-key`, `old-abbey` and `aldric`. In the
closeout's office-meeting identity run, the 3B narrator proposed giving
`silver-key` to Alice. The validator refused it, which is H05 and C06
behaving correctly. J01-J03 concern the schema, which is unaffected. A
genre-neutral example is post-v1 prompt work.
17. **With the 3B narrator at a 4,096 window, the state lags the narration.** In
the closeout identity run, one proposal in ten applied. The scene was never
updated, and a person the narration introduced never became an entity. Every
refusal and unparseable block was recorded and shown to the model, as
designed. This is risk 2 observed, not a new product defect.
---
## Q. Planning and document changes
| Document | Change | Kind |
| --- | --- | --- |
| `planning/TECHNICAL-DESIGN.md` | **New §15.2** — the inference window as a ceiling: discovery, enforcement, and why there is no hard-coded 4,096. | implementation fact |
| `planning/DATA-MODEL.md` | **New §28B** — M11's one column, and why the window ceiling, `duplicate_names` and `refused` are deliberately not stored. | implementation fact |
| `planning/BROWSER-UX-SPEC.md` | **§8A gains "As implemented (M11)"** — the position indicator and the three properties that make it answer the requirement. §8A's own text is unchanged. | implementation fact |
| `planning/SECURITY-THREAT-MODEL.md` | **New §42B** — the probe under §73, the silent partial correction as a §69 gap now closed, H09 not applicable with the condition enforced, and the measured contrast. | implementation fact + boundary note |
| `planning/V1-ACCEPTANCE-TESTS.md` | Results for every REQUIRED test (§F); §P1's duplicate-name question **settled** (report, do not refuse); §P3 gains an M11 disposition recording that the identity diagnostic exists and remains a test-design task rather than an acceptance test. | acceptance evidence + disposition |
| `planning/BUILD-MILESTONES.md` | **M11 status block**: the blocker closed, the four post-M8 findings disposed of, the two defects the validation found. | milestone status |
| `planning/VERSION.md` | **v3.7 entry.** | package version |
| `planning/README.md` | Status, milestone map, reading order; M9's and M10's reports rotated to the archive and the reason the exception ended. | index |
| `README.md` | `contextwindow.py` in the architecture map; the context-window behaviour as a feature. | developer docs |
| `DEVELOPMENT.md` | The context-window section rewritten around what the application now does; a new section on the six release harnesses. | developer docs |
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | New. | milestone report |
| `planning/archive/milestone-reports/` | M9's and M10's reports moved here. | rotation |
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Revised 2026-09-14**: M01-M04 on the complete evidence run; §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R rewritten; the §G.0 addendum removed. | milestone report |
| `DEVELOPMENT.md` | The pointer to the removed §G.0 replaced; a new section on logging a GPU inference host during a long run. | developer docs |
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Closeout, 2026-09-14**: new §S (exact-tree verification, the identity results) and §T (acceptance record); the 82-test count and §L's narrator corrected; §A, §B, §D.1, §F, §L, §O, §P and §R brought up to date. | milestone report |
| `planning/BUILD-MILESTONES.md`, `VERSION.md` (v4.0), `README.md`, `V1-ACCEPTANCE-TESTS.md` §P and §P3, root `README.md` | M11 accepted and the release gate recorded as passed; a post-v1 backlog. | milestone status |
| `backend/tools/m11_browser.py` | The G01 import wait, which could not fail. | release-test infrastructure |
**Revised after this report, in planning package v3.9 (2026-09-14):**
`planning/V1-ACCEPTANCE-TESTS.md` gained result blocks for M01-M04 and corrected
its M11 disposition. `planning/BUILD-MILESTONES.md`, `planning/VERSION.md` and
`planning/README.md` record the long-run evidence. `DATA-MODEL.md` §28B now
records migration 94, and `TECHNICAL-DESIGN.md` (new §15.4), `CONTEXT-AND-MEMORY.md`
§51 and ADR 013 record the two defect fixes as implemented.
**Implementation facts added:** §15.2, §28B, §8A's implementation note, §42B, and
the M11 status block. Each records what the code does; none changes what is
required.
**Requirement changes: zero.** No acceptance test was retired, relaxed,
reclassified or rewritten to match behaviour. H09 is reported NOT APPLICABLE on
the condition its own text states, and that condition is now enforced by a test
rather than asserted. §P1's question was *settled* — the answer being that a
shared display name is reported rather than refused — which resolves an open
design question rather than weakening a requirement; the permissive behaviour is
pinned by a test so a later milestone changes it deliberately.
---
## R. Final release-readiness assessment
**1. Does every REQUIRED FOR V1 acceptance test pass?**
**Yes.** All 82 pass, with H09 recorded NOT APPLICABLE on the condition its own
text states. M01 to M04 pass on the evidence run in §G, and every other REQUIRED
test passes on the evidence in §F.
**2. Does M01 pass with 100+ accepted turns?**
**Yes.** 101 accepted turns on commit `96c1bf5`, with three genuine process
restarts, all thirteen scheduled history operations, an induced failed call that
corrupted nothing, zero failed post-turn passes, 33 memories and 12 summaries,
and recovery onto a clean data directory 16 of 16. §G.1-§G.5.
**3. Does actual model context capacity match the application's release
assumptions?**
**Yes, it is enforced rather than assumed, and the margin is thin.** The window
was verified on 101 of 101 turns at 16,384, and the application caps its budget to
whatever the server reports (§E). By the narrator's own tokenizer the largest
prompt plus reserve is 16,342 of 16,384. It fits, with 23-42 tokens of headroom
across the GPU runs (§N, §P).
**4. Does offline operation pass from a fresh state with no Internet route?**
**Yes.** 23 of 23 checks in a container with `--network none` and a fresh volume:
no route, no DNS, first page load, every asset local, campaign creation, state
extraction, knowledge import and retrieval, prompt assembly, export, import, the
media module inert, and campaigns surviving a container restart. A turn with no
model reachable is reported and corrupts nothing. §J. Repeated at closeout on
the release-candidate tree, 23 of 23, from an image built with `--no-cache`
(§S.4).
**5. Does branch/memory/summary/scene isolation pass?**
**Yes.** All four in one long campaign, each with a positive control, including a
summary **regenerated after the divergence** and M10's derived Scene Packet. §H.
The evidence run also restored a Save Point across a live memory bank. §G.2.
**6. Does trusted-LAN HTTPS inference still pass?**
**Yes, and re-established at closeout on the candidate tree** (§S.3, §S.5). The browser run, the identity
diagnostic and the CPU-host campaigns ran against Ollama on a separate physical
machine over HTTPS, with a private CA in the application machine's OS trust store,
verification on and no bypass. The storyteller stayed loopback-bound. **The final
long runs used plain HTTP to a LAN GPU host.** H12 permits that, and it is not A06
evidence. §E.1, §F, §J.
**7. Does export/import/recovery pass?**
**Yes.** 16 of 16 checks moving the evidence run's campaign (207 actions, 88 of
them retained history, 2.74 MB) into a data directory that never existed, plus M9's
own recovery suite and the migration parity suite. §K.
**8. Do both fantasy and science-fiction fixtures pass?**
**Yes.** The fantasy Continuity Test is the long-run campaign. The Persephone
fixture passes 10 checks, including five entity types in one document, hard
canon, possession, reference retrieval, a starship in M10's Scene Packet, and a
whole-vocabulary check that no event type names a genre noun. No schema or code
change was required for either. §I.
**9. Do fresh-install and upgrade schemas agree?**
**Yes**, compared field by field: tables, columns with type and nullability,
indexes with columns and uniqueness, foreign keys, primary keys and the version
stamp. §K.
**10. Does the browser pass the release workflow?**
**Yes, on the release-candidate tree.** 38 of 38 in Firefox 155.0.1 at closeout,
against the SPA built from that tree and served by FastAPI, after a harness fix
(§S.3). Before that, 38 of 38 in Firefox 154.0.1 on `ef25b0a` (§L).
**11. Are there any unresolved blockers to independent v1 acceptance?**
**No known blocker.** No product defect is known and unfixed, and no requirement
was weakened. A reviewer should weigh five things before signing:
- M04's recovery ran through authoritative state, not memory retention (§G.4)
- the thin real-token headroom (§N)
- the narrator restating prompt text (§P)
- the extractor shapes and the lagging state the closeout's identity run found
(§S.6; §P risks 6, 16 and 17)
*At closeout:* the identity results are now in §S.5, and the black-box runs no
longer predate the candidate (§S). The acceptance decision is §T.
**12. Is the tree safe to commit as the M11 release candidate?**
**It is committed.** Seven signed commits follow the M10 base, the newest
`96c1bf5`, and there is no release tag. On that tree the backend suite passed
1,421 with 17 skipped and 0 failed (2026-09-13). The frontend suite passed 161 of
161, and lint exited 0 with warnings only (2026-09-14). The production build and
the Docker image were built for the first revision and not rebuilt since. The
remaining commit, this report's revision, is the owner's to sign.
*At closeout:* the report's revision is committed as `d198806`, and the planning
revision as `3652dc6`, both signed. The suites, the production build and a
`--no-cache` Docker image were all re-run on `3652dc6` (§S.2). The closeout
commit is staged for the owner's signature.
---
## S. Closeout verification on the release candidate (2026-09-14)
The browser, offline and identity runs above were taken on `ef25b0a`. Three later
product commits changed the turn commit, memory retrieval, narration extraction
and the state renderer's headings. The closeout repeated all three runs on the
exact release-candidate tree, and rebuilt and re-tested everything else there.
**The 100-turn campaign was not repeated.** The closeout changed no product code,
and that run was taken on `96c1bf5`, whose product code the candidate carries
unchanged.
### S.1 The candidate
| | |
| --- | --- |
| **Candidate** | `3652dc6fae903cad379da8efd5086709b5b9c6f5` on `m11-release-validation`, signed by the owner (`%G?` = `G`) |
| **Product code** | identical to `96c1bf5`. `d198806` and `3652dc6` change documentation only. No file under `backend/`, `frontend/`, `Dockerfile`, `docker-compose.yml` or the start scripts differs |
| **Working tree** | clean at the start, nothing staged, nothing untracked. The one source change during the closeout is the harness fix in §S.3, to `backend/tools/m11_browser.py`, which is neither part of the application nor in the image. No application file was modified during any run |
| **Signatures** | every commit from `144406c` to `3652dc6` is signed by the owner |
| **Provenance** | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` md5 is `07fde30437134836e2ee875e82a7cd31`, and `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` |
| **Dependencies** | `requirements.txt`, `requirements.lock`, `package.json` and `package-lock.json` are unchanged since `1013c94`. The lock's 35 pins match the environment, with no drift. FastAPI 0.141.1, uvicorn 0.52.4, SQLAlchemy 2.0.52, httpx 0.28.1, pydantic 2.13.5, tiktoken 0.14.0, python-multipart 0.0.32; React and React DOM 19.2.7, React Router 7.18.3; Vite 8.1.3 |
| **Application machine** | as §E.1, now on Ubuntu 24.04.5 LTS, kernel 7.0.0-31-generic, Docker 29.8.0, and Firefox 155.0.1 headless with geckodriver 0.37.1 |
| **Inference** | trusted-LAN: the CPU reference host (§E.1), Ollama 0.33.0 over HTTPS with a private CA in the application machine's OS trust store, verification on, no bypass. `qwen2.5:3b-instruct` (digest `357c53fb659c`), served at a verified 4,096 window. The storyteller listened on loopback only |
| **Evidence** | `$HOME/m11-evidence/closeout-3652dc6/` |
### S.2 Suites and builds
| Check | Command | Result |
| --- | --- | --- |
| Backend suite | `.venv/bin/python -m pytest tests/ -q`, with no `AIDND_TEST_*` set | **1,421 passed, 17 skipped, 0 failed**, 786 s |
| Frontend suite | `npm test` | **161 of 161**, 14 files |
| Lint | `npm run lint` | exit 0; six `only-export-components` warnings, no errors |
| Production build | `npm run build`, after deleting `frontend/dist` | 16 files, including `index-CZ7_g_3j.js` (392.69 kB) and `index-g-TAQQWd.css` (47.96 kB) |
| Docker image | `docker build --no-cache`, run by `tools/m11_offline.py` | `sha256:c9b872c2f4b701c8df32c58af944a3f456d1f4a6dd45105257ba6739f2382ed0`. `npm ci` and `pip install` ran rather than coming from cache. **The image's SPA is byte-identical, file for file, to the local build**, so the browser runs below exercised what ships |
### S.3 Browser regression
`tools/m11_browser.py`, against the SPA from §S.2, served by FastAPI on
loopback, with the narrator over trusted-LAN HTTPS:
| Run | Harness | Result |
| --- | --- | --- |
| 1 | as committed | **34 passed, 0 failed, 1 skipped**, 110 s. Modal focus containment skipped: "no Delete control found" |
| 2 | as committed, unchanged | **38 passed, 0 failed, 0 skipped**, 87 s |
| 3 | with the fix below | **38 passed, 0 failed, 0 skipped**, 69 s. **This is the reported result** |
The skip was a harness defect, not a product one. After Import is pressed, G01
waited for the word "hidden" in the page. The scenario's campaign is titled
*Hidden Knowledge*, and the page shows that title in its heading, so the wait
held at once. "A local file imports through the browser" was therefore recorded
without evidence. The modal check that follows could also run before the
imported source's row, which carries the Delete control, had rendered.
The frontend and the harness are unchanged since `ef25b0a`. Every run's database,
`ef25b0a`'s included, holds the imported source, so the import itself worked
every time. The harness now waits for the source's own row and records its
title (`▸ hidden`). Runs 1 and 2 are kept, and only run 3 used the fixed harness.
Run 3 passed everything §L lists:
- two real turns, and the tab title;
- the position indicator before and after Undo, and Undo and Redo;
- H06, H07, G09 and H04, with hostile narration;
- G01, through the real file input;
- the dialog's focus, accessible name, contents and Escape;
- §38, the context inspector, H10 and the CSP;
- the accessibility measurements, with rendered contrast unchanged at 14.57,
5.48, 13.57 and 5.88 to 1.
**Not driven in a real browser by this harness**, on this tree or before:
- Retry, Save Point creation and restore, state correction, and the
narration-length control;
- a failed generation shown to the reader;
- the export download, which is exercised only as far as the click (§L).
Those behaviours are covered on the candidate tree by the backend and component
suites, and by the 100-turn campaign through the API (§F, §G.2). The harness was
not extended, because the closeout is a regression rerun.
### S.4 Offline
`tools/m11_offline.py`: **23 passed, 0 failed**, 28 s. It ran the §S.2 image with
`--network none` and a fresh volume, and made the same 23 checks as §J:
- **Isolation:** no route, and no DNS.
- **Page and assets:** first page load; no remote origin; a CSP; every
referenced asset local.
- **Local operations:** campaign creation; state extraction; knowledge import;
prompt assembly; retrieval.
- **Failed turn with no model reachable:** reported, with no narration accepted,
the player's words kept, state unchanged, and the earlier story intact.
- **Export and import:** both work, with no secret in the bundle.
- **Media:** the module is inert with no provider.
- **Restart:** campaigns survive a container restart.
### S.5 The identity diagnostic
**Self-test** (`--scripted --inject`, no model): 7 signals, all STATE DEFECT. At
beat 5 the scripted narrator creates a second Alice. `shared_display_name` fires
at beat 5, and `duplicate_character_creation` fires at beats 5-10. The detectors
fire.
**Against a real narrator:**
| | |
| --- | --- |
| **Candidate** | `3652dc6` |
| **Narrator** | `qwen2.5:3b-instruct` on the CPU reference host over HTTPS. The window was verified at 4,096 on every turn, and the 16,384 budget was capped to it. `max_output_tokens` 500. No embedding model, so memory and summaries were off |
| **Fixture** | the companion *Multi-Character Identity Test* (`TEST-CAMPAIGN-FIXTURE.md` Appendix A): Bill, the protagonist, with Alice, Roger and John in one meeting room; ten beats |
| **Fixture accepted** | yes. The opening correction applied all six events and refused none, and the scene named its location and who was present. The harness stops on either failure rather than running a degraded campaign |
| **Turns** | 10 of 10 accepted, in 18 min 46 s |
| **Result** | **0 signals. Verdict: "no objective identity defect detected."** |
| **Evidence** | `identity/stdout.txt`; the final state, prompt, settings and an M9 bundle in `identity/turn-99/` |
**What the state did.** The narrator made ten state proposals:
- one applied (beat 3: John present, in the meeting room);
- five carried no block;
- three were unparseable: two cut off by the output limit, and one a parroted
reminder;
- one was refused: `set_possession` of `silver-key`, which does not exist.
The scene the fixture set was never updated, so Roger's exit went unrecorded.
From beat 6 the narration introduced a fifth person, "Mike", who never became an
entity, because the proposal that would have created him was cut off.
**What the run establishes:**
- the fixture is valid and was not silently refused;
- the authoritative state held four distinct people, with no shared display name
and no drift in the protagonist;
- the prompt's state section agreed with the document on every turn;
- the detectors fire when a confusion is planted;
- what a future occurrence needs is preserved and portable.
**What it cannot establish:**
- **The root cause of post-M8 finding D.** No same-name confusion occurred, so
the finding was not reproduced.
- **Prose-level misattribution.** This is not graded. The narration is kept for a
person to read.
- **A stale scene.** The objective checks cover only the protagonist's presence,
so Roger's unrecorded exit and the invented Mike raise no signal.
- **Derived contamination.** Memory and summaries were off.
- **Behaviour at a 16,384 window, or with another narrator.**
A clean verdict over a state that lags its narration is not evidence that
identities stay straight under a small model. What the diagnostic does show is
that, when identities go wrong, it can say whether the state, the context or the
derived data carried the error.
### S.6 Findings
**Release blockers: none.**
**Non-blocking, found at closeout:**
1. **Protocol shapes the extractor does not remove.** On 4 of 10 stored AI turns
in the identity run (depths 10, 12, 14 and 20), the narration kept event-call
syntax (`> Create_entity(...)`, `> set_possession(silver-key, "alice")`), a
lone `Scene:` line, and a parroted length hint ending "append the state block
well inside the limit.]". Replaying those four texts through the candidate's
`narrative.extract.split` leaves all four unchanged.
- **Why it survives.** The bracket survives because `_is_echoed_instruction`
recognises the reminder by "state block" together with "events list", and
the length hint names only the first.
- **Scope.** The 100-turn runs on `0c7316f` and `96c1bf5` store 0 turns in
these shapes.
- **Why it is not a blocker.** No REQUIRED FOR V1 test names stored-narration
purity, and `TECHNICAL-DESIGN.md` §15.4 holds that removing story is worse
than leaving protocol.
- **Decision.** The owner chose on 2026-09-14 to carry it as a residual risk
(§P risk 6) rather than change product code. A code change would have
invalidated the 100-turn evidence. §O.8's "every shape observed" was true
of the shapes observed when it was written.
2. **The state rule's example is fantasy** (§P risk 16).
3. **The state lagged the narration** at a 4,096 window with the 3B narrator
(§P risk 17).
**Harness defect, fixed:** the G01 import wait (§S.3, §O).
**Corrections to this report:** the REQUIRED FOR V1 count is 82 (§A, §R), and the
browser narrator in §L was `qwen2.5:3b-instruct`.
### S.7 Release-shaped smoke test
This ran after the documentation changes above. It started from the Dockerfile,
not from a development server. The script is not a repository harness, and its
evidence is `smoke-run2/`. The one source difference from the candidate is the
§S.3 harness fix, which is not in the image. **14 passed, 0 failed:**
- **Build and start:** the image builds with `--no-cache`; the container starts
on a fresh volume; it is published on host loopback only
(`127.0.0.1:18080`).
- **Page:** the first page loads, with every asset it names served locally.
- **Endpoint policy:** the approved trusted-LAN endpoint verifies over HTTPS,
with the private CA supplied to the container through the host's trust bundle,
mounted read-only. There is no bypass. A public endpoint is refused with
HTTP 400, and the approved endpoint stays configured.
- **Play:** a campaign is created, and one normal story turn is accepted (24 s,
`qwen2.5:3b-instruct`, window 4,096).
- **Persistence:** after a container restart, the campaign reopens with the same
actions and the same state.
- **Browser:** Firefox 155.0.1 loads the reopened campaign and shows its story,
with the tab reading "Closeout Smoke — Interactive Story".
The first attempt (`smoke/`) passed its first 13 checks. It then stopped on a
defect in the smoke script itself (`browser.title` is a property), before the
browser check was recorded. It is kept, and superseded by the run above.
---
## T. Acceptance record
**Decision: PASS. M11 is accepted at closeout, 2026-09-14.**
> Does this exact build meet the v1 black-box acceptance contract, and can it be
> packaged as the first production release?
**Yes.** Every REQUIRED FOR V1 condition remains satisfied. No test was waived,
relaxed or reclassified.
| Area | Status | Evidence |
| --- | --- | --- |
| All REQUIRED FOR V1 tests | 82: 81 PASS, H09 NOT APPLICABLE | §F |
| M01-M04 | PASS, on `96c1bf5`, whose product code the candidate carries unchanged | §G |
| Offline / local-only operation | PASS, on the candidate | §S.4 |
| Trusted-LAN inference (A06) | PASS, on the candidate: HTTPS, private CA, verification on, storyteller on loopback | §S.3, §S.5 |
| Undo / Redo / Retry / Save Point | PASS: D01-D14 in the long run and the suite; Undo and Redo in a real browser on the candidate | §F, §G.2, §S.3 |
| Branch, memory, summary and scene isolation | PASS: E01-E04, in the candidate's backend suite | §H, §S.2 |
| State atomicity | PASS: L01, in the long run and in the candidate's container | §F, §S.4 |
| Export / import / recovery | PASS: I01-I07; the long campaign onto a clean data directory, 16 of 16 | §K |
| Fresh-install vs upgraded schema parity | PASS: `test_m11_migration.py`, in the candidate's suite | §K, §S.2 |
| Fantasy fixture | PASS: the long-run campaign | §G, §I |
| Science-fiction fixture | PASS: `test_m11_scifi.py`, in the candidate's suite | §I, §S.2 |
| Security and local endpoint policy | PASS: H01-H12 | §F, §J, §S.4 |
| Browser release workflow | PASS: 38 of 38 on the candidate | §S.3 |
**Qualifications that stand:**
- M04's fact was recovered through authoritative state, not independent memory
retention. The owner accepted that on 2026-09-13.
- K04, a SHOULD test, passes on its deferred branch.
- The export download is exercised in the browser only as far as the click.
- Retry, Save Point, state correction, narration length and failed generation are
proved through the API and the component suite, not driven in a real browser.
- The real-token headroom at the largest long-run prompt is 42 tokens (§N).
- A06's HTTPS evidence ran at a 4,096 window on the CPU host, and the long run
used plain HTTP to a LAN GPU host.
- The residual risks in §P remain residual.
**Four separate events:**
| Event | State |
| --- | --- |
| M11 accepted | recorded here, 2026-09-14; takes effect with the owner's signed closeout commit |
| Release candidate verified | done, 2026-09-14, on `3652dc6` (§S) |
| Release commit signed | not yet: the closeout commit is staged for the owner |
| `v1.0.0` tagged | not done. The owner's decision, and the tag must point at the signed closeout commit |
**Where this report stays.** `planning/reports/` holds the most recently
completed milestone's report. That report moves to the archive when the next
milestone's report is written. There is no next milestone, so this report stays
where it is. No new convention was invented for it.
---
*§A-§R written by the implementer. §S and §T added at the M11 closeout,
2026-09-14; §T is the acceptance record.*