Drives in a real browser the reader workflows v1 proved only through the API or the component suite, including an export that leaves the browser as a file. Final run: 91 checks (the 38 existing M11 checks plus 53 new), 0 failed, 0 skipped, on the production build over trusted-LAN HTTPS. - tools/m11_browser.py: scenarios for Retry and takes, Save Point create / restore / Redo, state correction (accepted, and a refused correction with its reason), narration length reaching each turn's prompt, failed generation (an unserved model blocked up front; a listed model that cannot narrate failing in the open) and recovery, and export download from the library and from campaign settings, imported into a fresh application. Rows are tagged M11 / WP-C and counted separately; --only for development. The M11 checks now wait on conditions instead of sleeping. - tools/m11_webdriver.py: Firefox download preferences, a $HOME-only download folder, a download wait that ignores partial, empty, pre-existing and still-growing files, centred real clicks, tabs, and condition waits. - tests/test_v11_c_browser_helpers.py: the download wait, prefs and $HOME guard, without a browser. - frontend: a correction the story refused was presented as "Generation failed" with a Retry offer and a typed-input claim. It is now "That correction was not applied", not retryable, with the reason kept (errors.js, FailureNotice.jsx; 3 regression tests). - DEVELOPMENT.md: the harness command, download profile and $HOME rule, what counts as a finished download, and the no-sleep rule. - docs: V1.1-PLAN, VERSION v4.4, planning README, reports/v1.1/V1.1-WP-C-REPORT.md. Open for the owner: "Correct" on an Important Facts row is always refused (K1), and a partly refused correction is not reachable from the reader UI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
476 lines
27 KiB
Markdown
476 lines
27 KiB
Markdown
# v1.1 WP-C — Browser Release Coverage
|
||
|
||
**Status:** COMPLETE, staged for owner review. The final run passed 91 checks with 0 failed and 0 skipped. The decision is in §R.
|
||
|
||
---
|
||
|
||
## A. Repository baseline
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Branch | `v1.1-development` |
|
||
| HEAD at start | `0c1ba836babe1447ad3693b4d95189a325b7b3b6` — *v1.1 WP-B.2: independent long-term memory retention*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
|
||
| Its ancestry | `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `ac465ed` (plan v4.1), `432f041` (v1.0.0) |
|
||
| Working tree at start | clean; nothing staged |
|
||
| WP-D, WP-E | not started |
|
||
|
||
---
|
||
|
||
## B. Existing browser harness
|
||
|
||
`backend/tools/m11_browser.py` drives Firefox through `backend/tools/m11_webdriver.py`, a
|
||
dependency-free W3C WebDriver client. It builds its fixture campaign through the API,
|
||
plays two narrator turns through the streaming endpoint, and then asserts on the
|
||
rendered DOM. The M11 closeout run (`3652dc6`, run 3) passed **38/38**, 0 failed,
|
||
0 skipped, on snap Firefox 155.0.1 with geckodriver 0.37.1 and `qwen2.5:3b-instruct`
|
||
over trusted-LAN HTTPS.
|
||
|
||
What it did not do, and WP-C closes: drive Retry and takes, Save Point create and
|
||
restore, state correction, narration length or failed generation through the UI,
|
||
and prove an export leaves the browser as a file.
|
||
|
||
Before WP-C it also slept, for a fixed time, before several assertions: after Undo
|
||
and Redo, after planting hostile narration, after choosing a knowledge file, around
|
||
the delete dialog, and around opening a panel. §J records that as a harness defect.
|
||
|
||
---
|
||
|
||
## C. Browser and download environment
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Firefox | **155.0.1, the snap** (`/snap/bin/firefox`). No second Firefox was installed |
|
||
| geckodriver | 0.37.1 (the snap) |
|
||
| Headless | yes |
|
||
| Frontend | the production build (`vite build`) served by FastAPI on loopback; no Vite dev server |
|
||
| Certificates | `acceptInsecureCerts: false`, unchanged |
|
||
|
||
**The download profile.** `m11_webdriver.firefox_download_prefs` is passed as
|
||
`moz:firefoxOptions.prefs`:
|
||
- `browser.download.folderList` 2, `browser.download.dir` `<--out>/downloads`,
|
||
`browser.download.useDownloadDir` true;
|
||
- `browser.download.start_downloads_in_tmp_dir` false;
|
||
- `browser.download.always_ask_before_handling_new_types` false;
|
||
- `browser.helperApps.neverAsk.saveToDisk` `application/json,application/octet-stream`;
|
||
- the download panel suppressed.
|
||
|
||
**The folder.** `--out/downloads`, which must be under `$HOME`
|
||
(`require_under_home`). The harness deletes it at the start of a run and creates it
|
||
fresh. Evidence lives under `$HOME/v11-evidence/wp-c/`.
|
||
|
||
**Does the snap Firefox download?** Yes. Measured first, with a probe
|
||
(`$HOME/v11-evidence/wp-c/probe/`): a loopback page runs the product's own download
|
||
pattern (a JSON blob, an `<a download>` click, an immediate `revokeObjectURL`).
|
||
- Download folder under `~/v11-evidence`: 33-byte file written.
|
||
- Download folder under `~/Downloads`: 33-byte file written.
|
||
|
||
The first probe wrote nothing, and that was a probe defect (J1), not the snap. The
|
||
non-snap fallback the plan allows was therefore not needed.
|
||
|
||
**When a download counts as finished** (`m11_webdriver.wait_for_download`, tested
|
||
without a browser in `test_v11_c_browser_helpers.py`, 7 tests). All of these at once:
|
||
- a name absent from the listing taken before the click;
|
||
- no `*.part` file in the folder;
|
||
- more than zero bytes;
|
||
- the same size across 3 consecutive polls.
|
||
|
||
A zero-byte, partial, pre-existing or still-growing file never counts, and neither
|
||
does the "Campaign exported." toast.
|
||
|
||
---
|
||
|
||
## D. Retry scenario
|
||
|
||
Real narration: **yes**. All checks use the reader-facing controls on the play page.
|
||
|
||
| Browser action | Observable assertion | Result |
|
||
| --- | --- | --- |
|
||
| Type in "What you do next", press **Send** | a new narration renders and the page is idle; its exact text is recorded | PASS |
|
||
| — | **Retry** is offered (enabled) on the newest narration | PASS |
|
||
| Press **Retry** | the take indicator on the newest narration reads **2/2** | PASS |
|
||
| — | the second take's text differs from the first (otherwise 1/2 could not be told from 2/2) | PASS |
|
||
| Press **‹** (Previous take) | the indicator reads **1/2**, and the narration is identical to the first recorded text | PASS |
|
||
| — | the second take's text is not shown anywhere in the transcript | PASS |
|
||
| Press **›** (Next take) | **2/2**, showing the second take | PASS |
|
||
| Reload the page | the indicator still reads **2/2** on the live take, showing the second take | PASS |
|
||
| Press **‹** after the reload | **1/2** still shows the first text, unchanged | PASS |
|
||
|
||
All text comparisons are of the rendered `.turn-text`. No database was read.
|
||
|
||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||
|
||
## E. Save Point scenario
|
||
|
||
Real narration: **yes** (two turns after the Save Point).
|
||
|
||
| Browser action | Observable assertion | Result |
|
||
| --- | --- | --- |
|
||
| Press **Save Point**, type a name in "Save this moment", submit | a row with that exact name appears | PASS |
|
||
| — | the row's moment is the moment being read ("Moment 7") | PASS |
|
||
| **Send** two turns | two new narrations render; position "Moment 11" | PASS |
|
||
| In Save Points, press **Restore**, then confirm **Restore** | the position reads "Moment 7 · later story ahead" | PASS |
|
||
| — | the transcript ends at the Save Point: its last narration is the one read there, and neither later narration is shown | PASS |
|
||
| — | the position says later story is ahead | PASS |
|
||
| — | **Redo** is enabled | PASS |
|
||
| Press **Redo** until it is disabled (each press waits for the position to change) | the last two narrations are the two later turns, text-identical, at "Moment 11" | PASS |
|
||
| Reload the page | the position is still "Moment 11" | PASS |
|
||
| — | the named Save Point is still listed | PASS |
|
||
|
||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||
|
||
## F. State-correction scenario
|
||
|
||
**Owner decision (2026-09-15).** The State panel cannot produce a *partly* refused
|
||
correction:
|
||
- "Save correction" sends exactly one `add_fact`, and "That's wrong" sends exactly
|
||
one `invalidate_fact`;
|
||
- so a correction is applied whole or refused whole (HTTP 400);
|
||
- the route's partial application (`refused` alongside applied changes) is
|
||
reachable only through the API.
|
||
|
||
WP-C drives what the reader can reach: an accepted correction that persists, and a
|
||
refused correction whose refusal and reason are visible, with the refused change
|
||
not applied. It records the partial refusal as unreachable from the reader UI. No
|
||
product change was made for it.
|
||
|
||
Real narration: **yes** for the refusal (it needs history to step back over); the
|
||
accepted correction needs none and also passed in the no-narrator smoke runs.
|
||
|
||
| Browser action | Observable assertion | Result |
|
||
| --- | --- | --- |
|
||
| In **State**, press **Correct something**, type a fact, press **Save correction** | the form closes and the State panel shows the fact | PASS |
|
||
| Reload, open **State** | the fact is still shown | PASS |
|
||
| In a second tab press **Undo** (the correction belongs to the moment it was made at); in the first tab, whose panel still shows the fact, press **That's wrong** on it | a failure notice carrying the correction's refusal ("can't be applied") appears | PASS |
|
||
| — | it is labelled **"That correction was not applied"**, with no "Try that turn again" and no claim that typed text was kept (§K2) | PASS |
|
||
| Open **Show technical details** | the reason is visible: "That correction can't be applied — no fact 'f3' to invalidate." | PASS |
|
||
| Second tab **Redo**, close it; reload the first tab at the corrected moment | the fact still stands, with its **That's wrong** control: the refused withdrawal was not applied | PASS |
|
||
|
||
**Partial refusal.** As the owner decided, a *partly* refused correction is not
|
||
reachable from the reader UI, and was not driven. The route's partial application
|
||
remains covered by the backend suite.
|
||
|
||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||
|
||
## G. Narration-length scenario
|
||
|
||
Real narration: **yes** (one turn per band). The model's actual length is not
|
||
asserted, because nothing in the product contract requires it. What is asserted is
|
||
that the chosen band reached the turn's own prompt.
|
||
|
||
The expected sentence comes from the product's own `builder.length_hint` at the
|
||
run's 400-token output cap:
|
||
- **brief:** "must not exceed 180 words, and it should not stop short of about 70";
|
||
- **long:** "must not exceed 236 words, and it should not stop short of about 118".
|
||
|
||
| Browser action | Observable assertion | Result |
|
||
| --- | --- | --- |
|
||
| **Settings** panel: choose **Brief** in "Narration length", press **Save changes** | the button reads "Saved" | PASS |
|
||
| **Send** a turn; choose **Long**, **Save changes**; **Send** a second turn | both turns render | PASS (both) |
|
||
| On the brief turn press **Inspect context** | the prompt sections of "The exact text the narrator was sent" contain brief's range, and do **not** contain that turn's own reply (so this is the turn's record, not the dry run of the next) | PASS |
|
||
| On the long turn press **Inspect context** | the same, with long's range | PASS |
|
||
| Reload, open **Settings** | "Narration length" still reads **long** | PASS |
|
||
|
||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||
|
||
## H. Failed-generation scenario
|
||
|
||
**Owner decision (2026-09-15).** The literal sequence (save an unserved model, then
|
||
submit a turn) cannot be driven. The model check (`modelStatus.jsx`) marks a
|
||
configured model absent from the endpoint's list as `missing-model`, and
|
||
`blocksPlay` disables Send, Continue and Retry up front. That is M8's intended
|
||
behaviour, not a defect. So WP-C drives both of these:
|
||
1. **The unserved model** saved through Settings: the reader is told and cannot
|
||
send, and the story is unchanged.
|
||
2. **A submitted failure:** a model the endpoint lists but that cannot narrate,
|
||
`nomic-embed-text:latest`. Send stays enabled, and the turn fails in the open.
|
||
|
||
Then recovery with the reference model. The report records that path 2 uses a
|
||
listed, non-narrating model rather than "a name the server does not serve".
|
||
|
||
Real narration: **yes**.
|
||
|
||
| Browser action | Observable assertion | Result |
|
||
| --- | --- | --- |
|
||
| **Settings**: "Type a model name instead", type an unserved name, **Save** | the page says "Saved" | PASS |
|
||
| Open the campaign | the header model status is `missing-model` and the setup notice is shown | PASS |
|
||
| — | **Send** and **Continue** are disabled | PASS |
|
||
| — | the story is unchanged | PASS |
|
||
| **Settings**: choose `nomic-embed-text:latest` from the installed-model picker, **Save** | "Saved" | PASS |
|
||
| Open the campaign (model status `ready`); type a turn; **Send** | a failure notice is shown: "Generation failed", with the server's reason under the details, `"nomic-embed-text:latest" does not support chat` (HTTP 400) | PASS |
|
||
| — | no narration was added | PASS |
|
||
| — | the typed text is still in the input box | PASS |
|
||
| — | the earlier story is text-identical | PASS |
|
||
| Open **State** | the rendered state is identical to before the failure | PASS |
|
||
| **Settings**: choose `qwen2.5:3b-instruct`, **Save** | "Saved" | PASS |
|
||
| Open the campaign; type a turn; **Send** | exactly one new narration | PASS |
|
||
| — | the earlier story is intact | PASS |
|
||
| Reload | the successful turn is still the last narration | PASS |
|
||
|
||
The failed turn's player moment stays in the transcript, as A05 intends
|
||
(`player_moment_kept_in_transcript` in the report).
|
||
|
||
**Deviation, owner-approved.** Path 2 uses a model the endpoint *lists* but that
|
||
cannot narrate, not "a name the server does not serve", because an unserved name
|
||
is caught before a turn can be submitted. Endpoint policy was not bypassed: both
|
||
paths use the same trusted-LAN HTTPS endpoint, and only the model name changed.
|
||
|
||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||
|
||
## I. Export-download scenario
|
||
|
||
Real narration: not needed for the download itself. In the final run the exported
|
||
campaign holds 19 moments of real narration, two takes and a Save Point.
|
||
|
||
**C6a — the campaign library**
|
||
|
||
| Browser action | Observable assertion | Result |
|
||
| --- | --- | --- |
|
||
| On the library page, press **Export** on the "Release Regression" card | a new file is written to `downloads/`; it is finished (no `.part`, stable size) | PASS |
|
||
| — | it is not empty: **136,739 bytes** | PASS |
|
||
| Parse the file | `format` is **`ai-dnd-adventure-v3`** | PASS |
|
||
| Start a **fresh application** (new database); press **Import campaign**; give its file input the downloaded path | the browser lands on the imported campaign's play page | PASS |
|
||
| Compare what the reader sees of the import with what the reader saw of the original (library card, position, Save Points panel) | same title ("Release Regression") | PASS |
|
||
| — | same number of moments (19) | PASS |
|
||
| — | same position ("Moment 18") | PASS |
|
||
| — | same Save Points, which also match the file | PASS |
|
||
|
||
The file's `headDepth` is 17, which the play page shows as "Moment 18", and its
|
||
`actions` count is 19.
|
||
|
||
**C6b — campaign settings**
|
||
|
||
| Browser action | Observable assertion | Result |
|
||
| --- | --- | --- |
|
||
| In the campaign's **Settings** panel, press **Export campaign** | a second, separate file is written and finished | PASS |
|
||
| — | not empty: 136,739 bytes | PASS |
|
||
| Parse the file | `format` is `ai-dnd-adventure-v3` | PASS |
|
||
|
||
Both files: `$HOME/v11-evidence/wp-c/final/``downloads/library-export.json` and `downloads/settings-export.json`.
|
||
The imported application's log is `import-server.log`.
|
||
|
||
---
|
||
|
||
## J. Harness defects found
|
||
|
||
| # | Defect | How found | Fix |
|
||
| --- | --- | --- | --- |
|
||
| J1 | The download probe's page wrote `URL.createObjectURL` inside an inline `onclick`, where `URL` is `document.URL`, a string. No blob was made, so it looked exactly like "the snap cannot download" | `gecko.log`: `TypeError: URL.createObjectURL is not a function` | The probe's script uses `window.URL`, the way the product's module does. Both download folders then worked |
|
||
| J2 | C3's first refusal withdrew a fact in a second tab and withdrew it again in the stale tab. Withdrawing keeps the fact, marked `invalidated` (C04's audit record), so the second withdrawal was valid and accepted. Nothing was refused, and "the refused change was not applied" passed without meaning anything | Smoke run: two C3 failures; `server.log` shows 201 for every correction; `narrative/apply.py` | The second tab steps the story back past the correction with Undo, so the stale tab withdraws a fact the story at that position does not have. The validator then refuses it: `no fact … to invalidate` |
|
||
| J3 | Clicks on a control just under the play page's fixed composer were intercepted ("Show technical details" in a failure notice, a turn's "Inspect context") | Dev run 1: `element click intercepted` | `Browser.click` scrolls the element to the centre of the view, then uses the real WebDriver click |
|
||
| J4 | The Settings model field is a text box until the endpoint's model list arrives, then a picker. Choosing before the check finished raced that swap | Dev run 1: `no such element: input#model` | Wait for the header's model status to leave `checking` first |
|
||
| J5 | C3's "a refused correction is shown" waited for *any* failure notice, so a notice about something else would have passed | Dev run 1: it passed on a notice titled "Generation failed" (§K2) | It now requires the notice to carry this correction's refusal ("can't be applied"), and asserts how it is labelled |
|
||
| J7 | C4 required the band's sentence in the inspector and the turn's own reply to be absent, to tell a turn's record from the dry run. But the inspector also renders what came back (`raw_output`, "What came back, before the state block was removed") inside the same section, so the absence could never hold | Dev run 2: both C4 inspector checks failed. The stored records show the brief turn's `length_hint` section carrying "must not exceed 180 words … about 70", the long turn's carrying "…236 … 118", and every turn's reply present in `raw_output` | The sentence and the absence are read from the prompt sections only, excluding the "What came back" block |
|
||
| J6 | M11's checks slept before assertions: after Undo and Redo, after planting hostile narration (1 s), after choosing a knowledge file (0.5 s), around the delete dialog (0.8 s and 0.6 s), and around opening a panel (0.8 s). A sleep is not evidence of what it waited for | Reading the harness against the brief's rule | Each is now a wait on the condition the check needs: the position changed or returned, the planted text rendered, Import enabled, a dialog present or gone, a panel-specific element present. §38's absence check now first waits for the knowledge library to render |
|
||
|
||
**Checked and not a defect.** In dev run 2, the original take at depth 6 (action 9)
|
||
had no `length_hint` in its stored record, while its Retry (action 10) did. The
|
||
original take's record holds only the per-attempt fields
|
||
(`attempts.ATTEMPT_KEYS`: world state, narrative state, raw output, usage,
|
||
accounting), with no sections and no settings. That is how a take that is no
|
||
longer live is stored, not a prompt built without the range. Every live turn's
|
||
prompt carried its band.
|
||
|
||
Also added: a panel opens only if it is not already open, because a tab toggles
|
||
its panel closed, and it is recognised by an element only that panel renders, not
|
||
by its title, which the tab itself already shows.
|
||
|
||
---
|
||
|
||
## K. Product defects found
|
||
|
||
### K1 — "Correct" on an Important Facts row is always refused (not fixed)
|
||
|
||
The State panel offers **Correct** on every row with a key. On an Important Facts
|
||
row the key is the fact's id, and on the scene-summary row it is `"summary"`.
|
||
`saveCorrection` sends that key as `add_fact.subject`, and the validator checks
|
||
`subject` as an entity reference. So every such correction is refused.
|
||
|
||
**Reproduced deterministically** against the real application (scratch `TestClient`,
|
||
no browser, no model):
|
||
|
||
| Correction the panel sends | Result |
|
||
| --- | --- |
|
||
| "Correct" on the Characters row (`subject='mara'`) | **201**, applied |
|
||
| "Correct" on the Important Facts row (`subject='f1'`) | **400** "That correction can't be applied — add_fact names subject='f1', which does not exist." |
|
||
|
||
**Not fixed in WP-C.** It does not prevent any WP-C behaviour: "Correct something"
|
||
and entity-row corrections work, and C3 uses them. A fix (offer Correct only
|
||
against entities, or send facts without a subject) is a small UX choice, left to
|
||
the owner as a v1.1 follow-up.
|
||
|
||
### K2 — a refused correction was presented as a failed turn (fixed)
|
||
|
||
**Found in the browser** (dev run 1, §J5). When the story refused a correction, the
|
||
failure notice said:
|
||
- the title **"Generation failed"**;
|
||
- the hint "Nothing was added to your story. You can try that turn again.";
|
||
- the button **Try that turn again**;
|
||
- the line "what you typed is still in the box below".
|
||
|
||
None of that is true of a State-panel correction. `classifyError` had no rule for
|
||
the server's refusal ("That correction can't be applied — …"), so it fell through
|
||
to the generation default. That falsified exactly what C3 checks: that a refusal
|
||
is shown to the reader as a refusal.
|
||
|
||
**Fix** (frontend only, no backend change):
|
||
- `errors.js`: one rule, checked first, for "correction can't be applied". It gives
|
||
`kind: state`, the title "That correction was not applied", the hint "Nothing in
|
||
the story or its state was changed. The reason is in the technical details.",
|
||
`retryable: false` and `keptInput: false`.
|
||
- `FailureNotice.jsx`: the "what you typed" line is shown unless a failure says
|
||
`keptInput: false`. Every other kind still shows it, so M8's A05 contract is
|
||
unchanged.
|
||
|
||
**Regression** (`failurePaths.test.jsx`, 3 tests):
|
||
- a refused correction classifies as a state refusal, not retryable, with the
|
||
reason kept;
|
||
- its notice shows the reason and no "Try that turn again" or typed-input claim;
|
||
- a failed turn still claims the typed words were kept.
|
||
|
||
In the browser, C3 now asserts the label, the absence of "Try that turn again" and
|
||
the absence of the typed-input line.
|
||
|
||
Frontend after the fix: **168/168** tests, lint exit 0 (15 pre-existing warnings, 0
|
||
errors, none in changed files), production build passes.
|
||
|
||
---
|
||
|
||
## L. Existing 38-check regression
|
||
|
||
**38/38 passed, 0 failed, 0 skipped** in the same run, tagged `M11`. The names are
|
||
identical to the M11 closeout run, so there is a one-to-one mapping and no check was
|
||
split, merged or dropped:
|
||
- B01 a turn is accepted (×2);
|
||
- A/UX: the tab title (×3);
|
||
- B position indicator (×3), D01 Undo, D04 Redo (×2);
|
||
- H06 (×3), H07, G09, H04;
|
||
- G01 knowledge import (×2);
|
||
- A11y dialog focus (×4);
|
||
- §38 narrator-only text absent;
|
||
- F05 context inspector;
|
||
- H10 (×2), H11/CSP (×2);
|
||
- A11y names, focus, tabindex, hover, contrast (×4), input focus.
|
||
|
||
What changed in them is how they wait (J6), not what they assert.
|
||
|
||
## M. New WP-C checks
|
||
|
||
**53/53 passed, 0 failed, 0 skipped**, tagged `WP-C`:
|
||
|
||
| Scenario | Checks | Result |
|
||
| --- | --- | --- |
|
||
| C1 Retry | 8 | 8 PASS |
|
||
| C2 Save Point | 9 | 9 PASS |
|
||
| C3 State correction | 6 | 6 PASS |
|
||
| C4 Narration length | 5 | 5 PASS |
|
||
| C5 Failed generation | 14 | 14 PASS |
|
||
| C6a Library export and import | 8 | 8 PASS |
|
||
| C6b Settings export | 3 | 3 PASS |
|
||
|
||
```text
|
||
existing M11 checks: 38/38
|
||
WP-C new checks: 53/53
|
||
failed: 0
|
||
skipped: 0
|
||
```
|
||
|
||
**Runs that are not the final evidence**, kept under `$HOME/v11-evidence/wp-c/`:
|
||
|
||
| Run | Where | Result | Why it is not evidence |
|
||
| --- | --- | --- | --- |
|
||
| `probe/` | loopback page | first: no download (J1); second: both folders written | environment probe |
|
||
| `smoke-1` | no narrator | 44 passed, 2 failed (J2), 6 skipped | partial |
|
||
| `smoke-2` | no narrator | 43 passed, 0 failed, 7 skipped | partial |
|
||
| `dev-gpu-1` | GPU host, plain HTTP | 71 passed, 3 failed (J3, J4; J5 found) | development, and not HTTPS |
|
||
| `dev-gpu-2` | GPU host, plain HTTP | 89 passed, 2 failed (J7) | development, and not HTTPS |
|
||
| `dev-gpu-3` | GPU host, `--only length` | 7 passed | development, and a subset |
|
||
|
||
## N. Production build / frontend verification
|
||
|
||
| | Result |
|
||
| --- | --- |
|
||
| Frontend suite (`npm test`) | **168/168**, 14 files. It was 165 before; the 3 new tests are K2's regressions |
|
||
| Lint (`npm run lint`, oxlint) | **exit 0: 0 errors**, 15 warnings. All are the pre-existing `only-export-components` kind, and none is in a file WP-C changed |
|
||
| Production build (`npm run build`) | passes; `dist/index.html` sha256 `62b6ea5eb4ce02a09f23cc1d48c335c2ada36b208b23c237338b39bf63a26cc5`, built 2026-09-15 15:11 from the final WP-C tree |
|
||
| What the browser ran | that build, served by FastAPI (`uvicorn app.main:app` on `127.0.0.1`); no Vite server |
|
||
| Backend product code | **unchanged**: nothing under `backend/app` is in the diff. The full backend suite was not rerun. The changed harness is tested by `test_v11_c_browser_helpers.py` (7 passed) |
|
||
| Offline regression | **23 passed, 0 failed** (`tools/m11_offline.py`, fresh `--no-cache` image, `--network none`, on the final WP-C tree), rerun because the frontend bundle changed (K2). Evidence: `$HOME/v11-evidence/wp-c/offline/` |
|
||
|
||
## O. Trusted-LAN / security
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Narrator | `qwen2.5:3b-instruct` on the CPU reference host, over **trusted-LAN HTTPS**. The certificate is from the private CA in this machine's trust store, and verified; there is no bypass |
|
||
| Window and A1 | every narrator turn (8): window **verified at 4,096**, accounting **`fits`**. None `exceeded` or `truncation_suspected` |
|
||
| Protocol echoes (A2) | none of the protocol shapes the harness looks for (a state fence, a hard-limit or reminder bracket, `Events: [`) appeared in any stored narration in this run. This is not a v1.1 protocol-leak result: the release gate owns that, and the mid-reply echo from WP-B.1 remains a separate residual |
|
||
| Storyteller | loopback only, both applications (the original and the fresh import) |
|
||
| Browser | `acceptInsecureCerts: false`; the CSP checks (H11) passed |
|
||
| Endpoint policy | unchanged; C5 changed only the model name, never the endpoint |
|
||
| Downloads | written only under `$HOME` (enforced); the harness makes no network request of its own beyond loopback and the configured endpoint |
|
||
| New dependencies | none: the harness still uses only `urllib`, no Selenium or Playwright |
|
||
| Real identifiers in committed files | none (scanned at staging) |
|
||
|
||
## P. Compatibility
|
||
|
||
| Area | Effect |
|
||
| --- | --- |
|
||
| Database schema, migrations | none |
|
||
| Bundle format | none: the downloads are `ai-dnd-adventure-v3` and import unchanged |
|
||
| History, Save Points, state semantics, memory, knowledge | none. WP-C drove them and changed nothing in them |
|
||
| Backend | no application code changed |
|
||
| Frontend | one behaviour change (K2): the refusal of a State-panel correction is labelled as a refusal rather than as a failed turn. Every other failure's classification, retry offer and typed-input claim is unchanged (M8's tests pass) |
|
||
| WP-A, WP-B | untouched |
|
||
|
||
## Q. Residual risks
|
||
|
||
| # | Risk |
|
||
| --- | --- |
|
||
| 1 | **K1**: "Correct" on an Important Facts or scene-summary row is always refused. Reproduced, not fixed: a small UX choice for the owner |
|
||
| 2 | **Partial refusal is API-only.** The reader UI cannot produce a partly refused correction (owner decision). The display for one exists but is unreachable from the panel's own controls |
|
||
| 3 | **C5 path 2 depends on the reference host listing an embedding model.** On a host without one, the submitted-failure path has no model to use |
|
||
| 4 | **Model nondeterminism in C1.** "The second take is a different narration" would fail if the model returned identical text for a retry. It did not in any run |
|
||
| 5 | **The harness runs on this machine's snap Firefox.** A different Firefox or a Chromium would need the download preferences re-checked |
|
||
| 6 | **The heuristic protocol-shape scan is not A2 evidence.** It is recorded only |
|
||
| 7 | The WP-B real-model memory limitation and the doubled-full-stop scene text are unchanged, and are not WP-C's |
|
||
|
||
## R. Final decision
|
||
|
||
```text
|
||
RETRY: PASS
|
||
SAVE POINT: PASS
|
||
STATE CORRECTION: PASS
|
||
NARRATION LENGTH: PASS
|
||
FAILED GENERATION: PASS
|
||
EXPORT DOWNLOAD — LIBRARY: PASS
|
||
EXPORT DOWNLOAD — SETTINGS: PASS
|
||
|
||
EXISTING BROWSER REGRESSION: PASS
|
||
WP-C NEW BROWSER COVERAGE: PASS
|
||
|
||
WP-C OVERALL:
|
||
PASS
|
||
```
|
||
|
||
**PASS**, on these grounds:
|
||
- the final run ended with **failed: 0, skipped: 0**;
|
||
- both export controls produced a real, finished, non-empty `ai-dnd-adventure-v3`
|
||
file on disk;
|
||
- the library file imported into a fresh application with the same title, moment
|
||
count, position and Save Points.
|
||
|
||
Two criteria were met in the form the owner approved (2026-09-15):
|
||
- **State correction:** an accepted correction, and a refused correction with its
|
||
reason. Partial refusal is recorded as unreachable from the reader UI.
|
||
- **Failed generation:** the up-front block for an unserved model, plus a submitted
|
||
failure with a listed model that cannot narrate.
|
||
|
||
**Product change:** K2, a mislabelled refusal, fixed narrowly with regression tests.
|
||
**Product defect left open:** K1.
|
||
|
||
Nothing is committed, pushed or tagged. WP-D and WP-E have not started.
|