Files
JesseMarkowitzandClaude Opus 5 59b5ebc2d8 v1.1 WP-C: browser release coverage
Drives in a real browser the reader workflows v1 proved only through the API
or the component suite, including an export that leaves the browser as a
file. Final run: 91 checks (the 38 existing M11 checks plus 53 new), 0 failed,
0 skipped, on the production build over trusted-LAN HTTPS.

- tools/m11_browser.py: scenarios for Retry and takes, Save Point create /
  restore / Redo, state correction (accepted, and a refused correction with
  its reason), narration length reaching each turn's prompt, failed
  generation (an unserved model blocked up front; a listed model that cannot
  narrate failing in the open) and recovery, and export download from the
  library and from campaign settings, imported into a fresh application.
  Rows are tagged M11 / WP-C and counted separately; --only for development.
  The M11 checks now wait on conditions instead of sleeping.
- tools/m11_webdriver.py: Firefox download preferences, a $HOME-only
  download folder, a download wait that ignores partial, empty, pre-existing
  and still-growing files, centred real clicks, tabs, and condition waits.
- tests/test_v11_c_browser_helpers.py: the download wait, prefs and $HOME
  guard, without a browser.
- frontend: a correction the story refused was presented as "Generation
  failed" with a Retry offer and a typed-input claim. It is now "That
  correction was not applied", not retryable, with the reason kept
  (errors.js, FailureNotice.jsx; 3 regression tests).
- DEVELOPMENT.md: the harness command, download profile and $HOME rule,
  what counts as a finished download, and the no-sleep rule.
- docs: V1.1-PLAN, VERSION v4.4, planning README,
  reports/v1.1/V1.1-WP-C-REPORT.md.

Open for the owner: "Correct" on an Important Facts row is always refused
(K1), and a partly refused correction is not reachable from the reader UI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 18:29:45 -04:00

476 lines
27 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# v1.1 WP-C — Browser Release Coverage
**Status:** COMPLETE, staged for owner review. The final run passed 91 checks with 0 failed and 0 skipped. The decision is in §R.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD at start | `0c1ba836babe1447ad3693b4d95189a325b7b3b6` — *v1.1 WP-B.2: independent long-term memory retention*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
| Its ancestry | `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `ac465ed` (plan v4.1), `432f041` (v1.0.0) |
| Working tree at start | clean; nothing staged |
| WP-D, WP-E | not started |
---
## B. Existing browser harness
`backend/tools/m11_browser.py` drives Firefox through `backend/tools/m11_webdriver.py`, a
dependency-free W3C WebDriver client. It builds its fixture campaign through the API,
plays two narrator turns through the streaming endpoint, and then asserts on the
rendered DOM. The M11 closeout run (`3652dc6`, run 3) passed **38/38**, 0 failed,
0 skipped, on snap Firefox 155.0.1 with geckodriver 0.37.1 and `qwen2.5:3b-instruct`
over trusted-LAN HTTPS.
What it did not do, and WP-C closes: drive Retry and takes, Save Point create and
restore, state correction, narration length or failed generation through the UI,
and prove an export leaves the browser as a file.
Before WP-C it also slept, for a fixed time, before several assertions: after Undo
and Redo, after planting hostile narration, after choosing a knowledge file, around
the delete dialog, and around opening a panel. §J records that as a harness defect.
---
## C. Browser and download environment
| | |
| --- | --- |
| Firefox | **155.0.1, the snap** (`/snap/bin/firefox`). No second Firefox was installed |
| geckodriver | 0.37.1 (the snap) |
| Headless | yes |
| Frontend | the production build (`vite build`) served by FastAPI on loopback; no Vite dev server |
| Certificates | `acceptInsecureCerts: false`, unchanged |
**The download profile.** `m11_webdriver.firefox_download_prefs` is passed as
`moz:firefoxOptions.prefs`:
- `browser.download.folderList` 2, `browser.download.dir` `<--out>/downloads`,
`browser.download.useDownloadDir` true;
- `browser.download.start_downloads_in_tmp_dir` false;
- `browser.download.always_ask_before_handling_new_types` false;
- `browser.helperApps.neverAsk.saveToDisk` `application/json,application/octet-stream`;
- the download panel suppressed.
**The folder.** `--out/downloads`, which must be under `$HOME`
(`require_under_home`). The harness deletes it at the start of a run and creates it
fresh. Evidence lives under `$HOME/v11-evidence/wp-c/`.
**Does the snap Firefox download?** Yes. Measured first, with a probe
(`$HOME/v11-evidence/wp-c/probe/`): a loopback page runs the product's own download
pattern (a JSON blob, an `<a download>` click, an immediate `revokeObjectURL`).
- Download folder under `~/v11-evidence`: 33-byte file written.
- Download folder under `~/Downloads`: 33-byte file written.
The first probe wrote nothing, and that was a probe defect (J1), not the snap. The
non-snap fallback the plan allows was therefore not needed.
**When a download counts as finished** (`m11_webdriver.wait_for_download`, tested
without a browser in `test_v11_c_browser_helpers.py`, 7 tests). All of these at once:
- a name absent from the listing taken before the click;
- no `*.part` file in the folder;
- more than zero bytes;
- the same size across 3 consecutive polls.
A zero-byte, partial, pre-existing or still-growing file never counts, and neither
does the "Campaign exported." toast.
---
## D. Retry scenario
Real narration: **yes**. All checks use the reader-facing controls on the play page.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| Type in "What you do next", press **Send** | a new narration renders and the page is idle; its exact text is recorded | PASS |
| — | **Retry** is offered (enabled) on the newest narration | PASS |
| Press **Retry** | the take indicator on the newest narration reads **2/2** | PASS |
| — | the second take's text differs from the first (otherwise 1/2 could not be told from 2/2) | PASS |
| Press **‹** (Previous take) | the indicator reads **1/2**, and the narration is identical to the first recorded text | PASS |
| — | the second take's text is not shown anywhere in the transcript | PASS |
| Press **›** (Next take) | **2/2**, showing the second take | PASS |
| Reload the page | the indicator still reads **2/2** on the live take, showing the second take | PASS |
| Press **‹** after the reload | **1/2** still shows the first text, unchanged | PASS |
All text comparisons are of the rendered `.turn-text`. No database was read.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## E. Save Point scenario
Real narration: **yes** (two turns after the Save Point).
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| Press **Save Point**, type a name in "Save this moment", submit | a row with that exact name appears | PASS |
| — | the row's moment is the moment being read ("Moment 7") | PASS |
| **Send** two turns | two new narrations render; position "Moment 11" | PASS |
| In Save Points, press **Restore**, then confirm **Restore** | the position reads "Moment 7 · later story ahead" | PASS |
| — | the transcript ends at the Save Point: its last narration is the one read there, and neither later narration is shown | PASS |
| — | the position says later story is ahead | PASS |
| — | **Redo** is enabled | PASS |
| Press **Redo** until it is disabled (each press waits for the position to change) | the last two narrations are the two later turns, text-identical, at "Moment 11" | PASS |
| Reload the page | the position is still "Moment 11" | PASS |
| — | the named Save Point is still listed | PASS |
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## F. State-correction scenario
**Owner decision (2026-09-15).** The State panel cannot produce a *partly* refused
correction:
- "Save correction" sends exactly one `add_fact`, and "That's wrong" sends exactly
one `invalidate_fact`;
- so a correction is applied whole or refused whole (HTTP 400);
- the route's partial application (`refused` alongside applied changes) is
reachable only through the API.
WP-C drives what the reader can reach: an accepted correction that persists, and a
refused correction whose refusal and reason are visible, with the refused change
not applied. It records the partial refusal as unreachable from the reader UI. No
product change was made for it.
Real narration: **yes** for the refusal (it needs history to step back over); the
accepted correction needs none and also passed in the no-narrator smoke runs.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| In **State**, press **Correct something**, type a fact, press **Save correction** | the form closes and the State panel shows the fact | PASS |
| Reload, open **State** | the fact is still shown | PASS |
| In a second tab press **Undo** (the correction belongs to the moment it was made at); in the first tab, whose panel still shows the fact, press **That's wrong** on it | a failure notice carrying the correction's refusal ("can't be applied") appears | PASS |
| — | it is labelled **"That correction was not applied"**, with no "Try that turn again" and no claim that typed text was kept (§K2) | PASS |
| Open **Show technical details** | the reason is visible: "That correction can't be applied — no fact 'f3' to invalidate." | PASS |
| Second tab **Redo**, close it; reload the first tab at the corrected moment | the fact still stands, with its **That's wrong** control: the refused withdrawal was not applied | PASS |
**Partial refusal.** As the owner decided, a *partly* refused correction is not
reachable from the reader UI, and was not driven. The route's partial application
remains covered by the backend suite.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## G. Narration-length scenario
Real narration: **yes** (one turn per band). The model's actual length is not
asserted, because nothing in the product contract requires it. What is asserted is
that the chosen band reached the turn's own prompt.
The expected sentence comes from the product's own `builder.length_hint` at the
run's 400-token output cap:
- **brief:** "must not exceed 180 words, and it should not stop short of about 70";
- **long:** "must not exceed 236 words, and it should not stop short of about 118".
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| **Settings** panel: choose **Brief** in "Narration length", press **Save changes** | the button reads "Saved" | PASS |
| **Send** a turn; choose **Long**, **Save changes**; **Send** a second turn | both turns render | PASS (both) |
| On the brief turn press **Inspect context** | the prompt sections of "The exact text the narrator was sent" contain brief's range, and do **not** contain that turn's own reply (so this is the turn's record, not the dry run of the next) | PASS |
| On the long turn press **Inspect context** | the same, with long's range | PASS |
| Reload, open **Settings** | "Narration length" still reads **long** | PASS |
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## H. Failed-generation scenario
**Owner decision (2026-09-15).** The literal sequence (save an unserved model, then
submit a turn) cannot be driven. The model check (`modelStatus.jsx`) marks a
configured model absent from the endpoint's list as `missing-model`, and
`blocksPlay` disables Send, Continue and Retry up front. That is M8's intended
behaviour, not a defect. So WP-C drives both of these:
1. **The unserved model** saved through Settings: the reader is told and cannot
send, and the story is unchanged.
2. **A submitted failure:** a model the endpoint lists but that cannot narrate,
`nomic-embed-text:latest`. Send stays enabled, and the turn fails in the open.
Then recovery with the reference model. The report records that path 2 uses a
listed, non-narrating model rather than "a name the server does not serve".
Real narration: **yes**.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| **Settings**: "Type a model name instead", type an unserved name, **Save** | the page says "Saved" | PASS |
| Open the campaign | the header model status is `missing-model` and the setup notice is shown | PASS |
| — | **Send** and **Continue** are disabled | PASS |
| — | the story is unchanged | PASS |
| **Settings**: choose `nomic-embed-text:latest` from the installed-model picker, **Save** | "Saved" | PASS |
| Open the campaign (model status `ready`); type a turn; **Send** | a failure notice is shown: "Generation failed", with the server's reason under the details, `"nomic-embed-text:latest" does not support chat` (HTTP 400) | PASS |
| — | no narration was added | PASS |
| — | the typed text is still in the input box | PASS |
| — | the earlier story is text-identical | PASS |
| Open **State** | the rendered state is identical to before the failure | PASS |
| **Settings**: choose `qwen2.5:3b-instruct`, **Save** | "Saved" | PASS |
| Open the campaign; type a turn; **Send** | exactly one new narration | PASS |
| — | the earlier story is intact | PASS |
| Reload | the successful turn is still the last narration | PASS |
The failed turn's player moment stays in the transcript, as A05 intends
(`player_moment_kept_in_transcript` in the report).
**Deviation, owner-approved.** Path 2 uses a model the endpoint *lists* but that
cannot narrate, not "a name the server does not serve", because an unserved name
is caught before a turn can be submitted. Endpoint policy was not bypassed: both
paths use the same trusted-LAN HTTPS endpoint, and only the model name changed.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## I. Export-download scenario
Real narration: not needed for the download itself. In the final run the exported
campaign holds 19 moments of real narration, two takes and a Save Point.
**C6a — the campaign library**
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| On the library page, press **Export** on the "Release Regression" card | a new file is written to `downloads/`; it is finished (no `.part`, stable size) | PASS |
| — | it is not empty: **136,739 bytes** | PASS |
| Parse the file | `format` is **`ai-dnd-adventure-v3`** | PASS |
| Start a **fresh application** (new database); press **Import campaign**; give its file input the downloaded path | the browser lands on the imported campaign's play page | PASS |
| Compare what the reader sees of the import with what the reader saw of the original (library card, position, Save Points panel) | same title ("Release Regression") | PASS |
| — | same number of moments (19) | PASS |
| — | same position ("Moment 18") | PASS |
| — | same Save Points, which also match the file | PASS |
The file's `headDepth` is 17, which the play page shows as "Moment 18", and its
`actions` count is 19.
**C6b — campaign settings**
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| In the campaign's **Settings** panel, press **Export campaign** | a second, separate file is written and finished | PASS |
| — | not empty: 136,739 bytes | PASS |
| Parse the file | `format` is `ai-dnd-adventure-v3` | PASS |
Both files: `$HOME/v11-evidence/wp-c/final/``downloads/library-export.json` and `downloads/settings-export.json`.
The imported application's log is `import-server.log`.
---
## J. Harness defects found
| # | Defect | How found | Fix |
| --- | --- | --- | --- |
| J1 | The download probe's page wrote `URL.createObjectURL` inside an inline `onclick`, where `URL` is `document.URL`, a string. No blob was made, so it looked exactly like "the snap cannot download" | `gecko.log`: `TypeError: URL.createObjectURL is not a function` | The probe's script uses `window.URL`, the way the product's module does. Both download folders then worked |
| J2 | C3's first refusal withdrew a fact in a second tab and withdrew it again in the stale tab. Withdrawing keeps the fact, marked `invalidated` (C04's audit record), so the second withdrawal was valid and accepted. Nothing was refused, and "the refused change was not applied" passed without meaning anything | Smoke run: two C3 failures; `server.log` shows 201 for every correction; `narrative/apply.py` | The second tab steps the story back past the correction with Undo, so the stale tab withdraws a fact the story at that position does not have. The validator then refuses it: `no fact … to invalidate` |
| J3 | Clicks on a control just under the play page's fixed composer were intercepted ("Show technical details" in a failure notice, a turn's "Inspect context") | Dev run 1: `element click intercepted` | `Browser.click` scrolls the element to the centre of the view, then uses the real WebDriver click |
| J4 | The Settings model field is a text box until the endpoint's model list arrives, then a picker. Choosing before the check finished raced that swap | Dev run 1: `no such element: input#model` | Wait for the header's model status to leave `checking` first |
| J5 | C3's "a refused correction is shown" waited for *any* failure notice, so a notice about something else would have passed | Dev run 1: it passed on a notice titled "Generation failed" (§K2) | It now requires the notice to carry this correction's refusal ("can't be applied"), and asserts how it is labelled |
| J7 | C4 required the band's sentence in the inspector and the turn's own reply to be absent, to tell a turn's record from the dry run. But the inspector also renders what came back (`raw_output`, "What came back, before the state block was removed") inside the same section, so the absence could never hold | Dev run 2: both C4 inspector checks failed. The stored records show the brief turn's `length_hint` section carrying "must not exceed 180 words … about 70", the long turn's carrying "…236 … 118", and every turn's reply present in `raw_output` | The sentence and the absence are read from the prompt sections only, excluding the "What came back" block |
| J6 | M11's checks slept before assertions: after Undo and Redo, after planting hostile narration (1 s), after choosing a knowledge file (0.5 s), around the delete dialog (0.8 s and 0.6 s), and around opening a panel (0.8 s). A sleep is not evidence of what it waited for | Reading the harness against the brief's rule | Each is now a wait on the condition the check needs: the position changed or returned, the planted text rendered, Import enabled, a dialog present or gone, a panel-specific element present. §38's absence check now first waits for the knowledge library to render |
**Checked and not a defect.** In dev run 2, the original take at depth 6 (action 9)
had no `length_hint` in its stored record, while its Retry (action 10) did. The
original take's record holds only the per-attempt fields
(`attempts.ATTEMPT_KEYS`: world state, narrative state, raw output, usage,
accounting), with no sections and no settings. That is how a take that is no
longer live is stored, not a prompt built without the range. Every live turn's
prompt carried its band.
Also added: a panel opens only if it is not already open, because a tab toggles
its panel closed, and it is recognised by an element only that panel renders, not
by its title, which the tab itself already shows.
---
## K. Product defects found
### K1 — "Correct" on an Important Facts row is always refused (not fixed)
The State panel offers **Correct** on every row with a key. On an Important Facts
row the key is the fact's id, and on the scene-summary row it is `"summary"`.
`saveCorrection` sends that key as `add_fact.subject`, and the validator checks
`subject` as an entity reference. So every such correction is refused.
**Reproduced deterministically** against the real application (scratch `TestClient`,
no browser, no model):
| Correction the panel sends | Result |
| --- | --- |
| "Correct" on the Characters row (`subject='mara'`) | **201**, applied |
| "Correct" on the Important Facts row (`subject='f1'`) | **400** "That correction can't be applied — add_fact names subject='f1', which does not exist." |
**Not fixed in WP-C.** It does not prevent any WP-C behaviour: "Correct something"
and entity-row corrections work, and C3 uses them. A fix (offer Correct only
against entities, or send facts without a subject) is a small UX choice, left to
the owner as a v1.1 follow-up.
### K2 — a refused correction was presented as a failed turn (fixed)
**Found in the browser** (dev run 1, §J5). When the story refused a correction, the
failure notice said:
- the title **"Generation failed"**;
- the hint "Nothing was added to your story. You can try that turn again.";
- the button **Try that turn again**;
- the line "what you typed is still in the box below".
None of that is true of a State-panel correction. `classifyError` had no rule for
the server's refusal ("That correction can't be applied — …"), so it fell through
to the generation default. That falsified exactly what C3 checks: that a refusal
is shown to the reader as a refusal.
**Fix** (frontend only, no backend change):
- `errors.js`: one rule, checked first, for "correction can't be applied". It gives
`kind: state`, the title "That correction was not applied", the hint "Nothing in
the story or its state was changed. The reason is in the technical details.",
`retryable: false` and `keptInput: false`.
- `FailureNotice.jsx`: the "what you typed" line is shown unless a failure says
`keptInput: false`. Every other kind still shows it, so M8's A05 contract is
unchanged.
**Regression** (`failurePaths.test.jsx`, 3 tests):
- a refused correction classifies as a state refusal, not retryable, with the
reason kept;
- its notice shows the reason and no "Try that turn again" or typed-input claim;
- a failed turn still claims the typed words were kept.
In the browser, C3 now asserts the label, the absence of "Try that turn again" and
the absence of the typed-input line.
Frontend after the fix: **168/168** tests, lint exit 0 (15 pre-existing warnings, 0
errors, none in changed files), production build passes.
---
## L. Existing 38-check regression
**38/38 passed, 0 failed, 0 skipped** in the same run, tagged `M11`. The names are
identical to the M11 closeout run, so there is a one-to-one mapping and no check was
split, merged or dropped:
- B01 a turn is accepted (×2);
- A/UX: the tab title (×3);
- B position indicator (×3), D01 Undo, D04 Redo (×2);
- H06 (×3), H07, G09, H04;
- G01 knowledge import (×2);
- A11y dialog focus (×4);
- §38 narrator-only text absent;
- F05 context inspector;
- H10 (×2), H11/CSP (×2);
- A11y names, focus, tabindex, hover, contrast (×4), input focus.
What changed in them is how they wait (J6), not what they assert.
## M. New WP-C checks
**53/53 passed, 0 failed, 0 skipped**, tagged `WP-C`:
| Scenario | Checks | Result |
| --- | --- | --- |
| C1 Retry | 8 | 8 PASS |
| C2 Save Point | 9 | 9 PASS |
| C3 State correction | 6 | 6 PASS |
| C4 Narration length | 5 | 5 PASS |
| C5 Failed generation | 14 | 14 PASS |
| C6a Library export and import | 8 | 8 PASS |
| C6b Settings export | 3 | 3 PASS |
```text
existing M11 checks: 38/38
WP-C new checks: 53/53
failed: 0
skipped: 0
```
**Runs that are not the final evidence**, kept under `$HOME/v11-evidence/wp-c/`:
| Run | Where | Result | Why it is not evidence |
| --- | --- | --- | --- |
| `probe/` | loopback page | first: no download (J1); second: both folders written | environment probe |
| `smoke-1` | no narrator | 44 passed, 2 failed (J2), 6 skipped | partial |
| `smoke-2` | no narrator | 43 passed, 0 failed, 7 skipped | partial |
| `dev-gpu-1` | GPU host, plain HTTP | 71 passed, 3 failed (J3, J4; J5 found) | development, and not HTTPS |
| `dev-gpu-2` | GPU host, plain HTTP | 89 passed, 2 failed (J7) | development, and not HTTPS |
| `dev-gpu-3` | GPU host, `--only length` | 7 passed | development, and a subset |
## N. Production build / frontend verification
| | Result |
| --- | --- |
| Frontend suite (`npm test`) | **168/168**, 14 files. It was 165 before; the 3 new tests are K2's regressions |
| Lint (`npm run lint`, oxlint) | **exit 0: 0 errors**, 15 warnings. All are the pre-existing `only-export-components` kind, and none is in a file WP-C changed |
| Production build (`npm run build`) | passes; `dist/index.html` sha256 `62b6ea5eb4ce02a09f23cc1d48c335c2ada36b208b23c237338b39bf63a26cc5`, built 2026-09-15 15:11 from the final WP-C tree |
| What the browser ran | that build, served by FastAPI (`uvicorn app.main:app` on `127.0.0.1`); no Vite server |
| Backend product code | **unchanged**: nothing under `backend/app` is in the diff. The full backend suite was not rerun. The changed harness is tested by `test_v11_c_browser_helpers.py` (7 passed) |
| Offline regression | **23 passed, 0 failed** (`tools/m11_offline.py`, fresh `--no-cache` image, `--network none`, on the final WP-C tree), rerun because the frontend bundle changed (K2). Evidence: `$HOME/v11-evidence/wp-c/offline/` |
## O. Trusted-LAN / security
| | |
| --- | --- |
| Narrator | `qwen2.5:3b-instruct` on the CPU reference host, over **trusted-LAN HTTPS**. The certificate is from the private CA in this machine's trust store, and verified; there is no bypass |
| Window and A1 | every narrator turn (8): window **verified at 4,096**, accounting **`fits`**. None `exceeded` or `truncation_suspected` |
| Protocol echoes (A2) | none of the protocol shapes the harness looks for (a state fence, a hard-limit or reminder bracket, `Events: [`) appeared in any stored narration in this run. This is not a v1.1 protocol-leak result: the release gate owns that, and the mid-reply echo from WP-B.1 remains a separate residual |
| Storyteller | loopback only, both applications (the original and the fresh import) |
| Browser | `acceptInsecureCerts: false`; the CSP checks (H11) passed |
| Endpoint policy | unchanged; C5 changed only the model name, never the endpoint |
| Downloads | written only under `$HOME` (enforced); the harness makes no network request of its own beyond loopback and the configured endpoint |
| New dependencies | none: the harness still uses only `urllib`, no Selenium or Playwright |
| Real identifiers in committed files | none (scanned at staging) |
## P. Compatibility
| Area | Effect |
| --- | --- |
| Database schema, migrations | none |
| Bundle format | none: the downloads are `ai-dnd-adventure-v3` and import unchanged |
| History, Save Points, state semantics, memory, knowledge | none. WP-C drove them and changed nothing in them |
| Backend | no application code changed |
| Frontend | one behaviour change (K2): the refusal of a State-panel correction is labelled as a refusal rather than as a failed turn. Every other failure's classification, retry offer and typed-input claim is unchanged (M8's tests pass) |
| WP-A, WP-B | untouched |
## Q. Residual risks
| # | Risk |
| --- | --- |
| 1 | **K1**: "Correct" on an Important Facts or scene-summary row is always refused. Reproduced, not fixed: a small UX choice for the owner |
| 2 | **Partial refusal is API-only.** The reader UI cannot produce a partly refused correction (owner decision). The display for one exists but is unreachable from the panel's own controls |
| 3 | **C5 path 2 depends on the reference host listing an embedding model.** On a host without one, the submitted-failure path has no model to use |
| 4 | **Model nondeterminism in C1.** "The second take is a different narration" would fail if the model returned identical text for a retry. It did not in any run |
| 5 | **The harness runs on this machine's snap Firefox.** A different Firefox or a Chromium would need the download preferences re-checked |
| 6 | **The heuristic protocol-shape scan is not A2 evidence.** It is recorded only |
| 7 | The WP-B real-model memory limitation and the doubled-full-stop scene text are unchanged, and are not WP-C's |
## R. Final decision
```text
RETRY: PASS
SAVE POINT: PASS
STATE CORRECTION: PASS
NARRATION LENGTH: PASS
FAILED GENERATION: PASS
EXPORT DOWNLOAD — LIBRARY: PASS
EXPORT DOWNLOAD — SETTINGS: PASS
EXISTING BROWSER REGRESSION: PASS
WP-C NEW BROWSER COVERAGE: PASS
WP-C OVERALL:
PASS
```
**PASS**, on these grounds:
- the final run ended with **failed: 0, skipped: 0**;
- both export controls produced a real, finished, non-empty `ai-dnd-adventure-v3`
file on disk;
- the library file imported into a fresh application with the same title, moment
count, position and Save Points.
Two criteria were met in the form the owner approved (2026-09-15):
- **State correction:** an accepted correction, and a refused correction with its
reason. Partial refusal is recorded as unreachable from the reader UI.
- **Failed generation:** the up-front block for an unserved model, plus a submitted
failure with a listed model that cannot narrate.
**Product change:** K2, a mislabelled refusal, fixed narrowly with regression tests.
**Product defect left open:** K1.
Nothing is committed, pushed or tagged. WP-D and WP-E have not started.