v1.1 WP-C: browser release coverage

Drives in a real browser the reader workflows v1 proved only through the API
or the component suite, including an export that leaves the browser as a
file. Final run: 91 checks (the 38 existing M11 checks plus 53 new), 0 failed,
0 skipped, on the production build over trusted-LAN HTTPS.

- tools/m11_browser.py: scenarios for Retry and takes, Save Point create /
  restore / Redo, state correction (accepted, and a refused correction with
  its reason), narration length reaching each turn's prompt, failed
  generation (an unserved model blocked up front; a listed model that cannot
  narrate failing in the open) and recovery, and export download from the
  library and from campaign settings, imported into a fresh application.
  Rows are tagged M11 / WP-C and counted separately; --only for development.
  The M11 checks now wait on conditions instead of sleeping.
- tools/m11_webdriver.py: Firefox download preferences, a $HOME-only
  download folder, a download wait that ignores partial, empty, pre-existing
  and still-growing files, centred real clicks, tabs, and condition waits.
- tests/test_v11_c_browser_helpers.py: the download wait, prefs and $HOME
  guard, without a browser.
- frontend: a correction the story refused was presented as "Generation
  failed" with a Retry offer and a typed-input claim. It is now "That
  correction was not applied", not retryable, with the reason kept
  (errors.js, FailureNotice.jsx; 3 regression tests).
- DEVELOPMENT.md: the harness command, download profile and $HOME rule,
  what counts as a finished download, and the no-sleep rule.
- docs: V1.1-PLAN, VERSION v4.4, planning README,
  reports/v1.1/V1.1-WP-C-REPORT.md.

Open for the owner: "Correct" on an Important Facts row is always refused
(K1), and a partly refused correction is not reachable from the reader UI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-15 18:29:45 -04:00
co-authored by Claude Opus 5
parent 0c1ba836ba
commit 59b5ebc2d8
11 changed files with 1711 additions and 157 deletions
+38 -3
View File
@@ -408,9 +408,17 @@ AIDND_TEST_ENDPOINT=... AIDND_TEST_MODEL=... AIDND_TEST_EMBED_MODEL=... \
# What that campaign is worth on a machine that has never seen it (I01-I07). # What that campaign is worth on a machine that has never seen it (I01-I07).
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01" .venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"
# The browser release regression and the accessibility measurements. Needs # The browser release regression and the accessibility measurements: M11's 38
# `frontend/dist` built and geckodriver on PATH. # checks plus v1.1 WP-C's reader workflows (Retry, Save Points, state correction,
.venv/bin/python -m tools.m11_browser --out "$HOME/m11-evidence/browser" # narration length, failed generation, export download). Needs `frontend/dist`
# built, geckodriver on PATH, and --out under $HOME (the downloads land inside
# it). Release evidence needs the narrator over trusted-LAN HTTPS.
AIDND_TEST_ENDPOINT=https://... AIDND_TEST_MODEL=qwen2.5:3b-instruct \
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/browser"
# Without a narrator (a partial smoke run, not evidence), or one scenario while
# developing (--only takes: shell, history, markdown, hidden, context, csp, a11y,
# retry, savepoint, state, length, failure, export).
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/smoke" --no-narrator
# A container with no network at all: the offline run and the packaging path. # A container with no network at all: the offline run and the packaging path.
.venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline" .venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline"
@@ -499,6 +507,33 @@ It exists so browser evidence needs no Selenium in the dependency surface, and
it documents the one environment quirk that matters here: a snap Firefox will it documents the one environment quirk that matters here: a snap Firefox will
not open a file the driver names under `/tmp`, but will under `$HOME`. not open a file the driver names under `/tmp`, but will under `$HOME`.
**Downloads in the browser harness (v1.1 WP-C).** The export checks click the
real Export controls and wait for the file on disk, so the browser has to save
without asking. `m11_webdriver.firefox_download_prefs` gives the WebDriver
session a profile that does that:
- `browser.download.folderList` 2, `browser.download.dir` the run's
`downloads/` folder, `browser.download.useDownloadDir` true;
- no "always ask", and `application/json` saved to disk.
It works on the snap Firefox this machine has (155.0.1, geckodriver 0.37.1), and
no separate Firefox is needed. The same sandbox rule applies as for opening
files: the download folder must be under `$HOME`, and the harness refuses one
that is not.
A download counts as finished only when all of these hold at once
(`m11_webdriver.wait_for_download`):
- a new name has appeared;
- no `*.part` file is left;
- the file is more than zero bytes;
- its size is the same across consecutive polls.
The toast that says "Campaign exported." is not evidence.
**Waiting.** Nothing in the harness sleeps before an assertion. Every wait is on
something the page, the browser or the filesystem shows. A condition that
already holds before the action it waits for does not count as waiting for that
action; the harness defects found in M8, M11 and WP-C were all of that shape.
## Backing up, and getting a campaign back ## Backing up, and getting a campaign back
There are two recovery tools and they answer different questions. Using the There are two recovery tools and they answer different questions. Using the
@@ -0,0 +1,91 @@
"""v1.1 WP-C: the browser harness's download helpers, without a browser.
`tools/m11_webdriver.wait_for_download` is what decides that an export actually
left the browser as a file. It must never call a download finished because a
file appeared, because it is still being written, or because it is empty — each
of those would make "the export works" a claim the evidence does not support.
python -m pytest tests/test_v11_c_browser_helpers.py -v
"""
import threading
import time
from pathlib import Path
import pytest
from tools import m11_webdriver as wd
def later(seconds, action):
timer = threading.Timer(seconds, action)
timer.start()
return timer
def test_a_finished_file_is_returned(tmp_path):
later(0.2, lambda: (tmp_path / "campaign.json").write_text('{"format": "x"}'))
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05)
assert found == tmp_path / "campaign.json"
def test_a_file_that_was_already_there_is_not_the_download(tmp_path):
(tmp_path / "old.json").write_text("{}")
with pytest.raises(wd.WebDriverError):
wd.wait_for_download(tmp_path, {"old.json"}, timeout=0.6, poll=0.05)
def test_an_empty_file_never_counts(tmp_path):
(tmp_path / "empty.json").write_bytes(b"")
with pytest.raises(wd.WebDriverError):
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
def test_nothing_counts_while_firefox_is_still_writing(tmp_path):
"""Firefox writes `<name>.part` beside the final name until it is done."""
(tmp_path / "campaign.json").write_text('{"format": "x"}')
(tmp_path / "campaign.json.part").write_text("")
with pytest.raises(wd.WebDriverError):
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
later(0.1, lambda: (tmp_path / "campaign.json.part").unlink())
assert wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05).name == "campaign.json"
def test_a_file_that_is_still_growing_is_not_finished(tmp_path):
target = tmp_path / "big.json"
target.write_text("{")
stop = threading.Event()
def grow():
for _ in range(8):
if stop.is_set():
return
with target.open("a") as fh:
fh.write("x" * 100)
time.sleep(0.05)
writer = threading.Thread(target=grow)
started = time.monotonic()
writer.start()
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05, stable_polls=3)
writer.join()
# Returned only once the size held still, so after the last write.
assert found == target
assert target.stat().st_size == 1 + 8 * 100
assert time.monotonic() - started >= 0.4
def test_the_prefs_save_downloads_unasked_to_the_folder_given(tmp_path):
prefs = wd.firefox_download_prefs(tmp_path)
assert prefs["browser.download.folderList"] == 2
assert prefs["browser.download.dir"] == str(tmp_path)
assert prefs["browser.download.useDownloadDir"] is True
assert prefs["browser.download.always_ask_before_handling_new_types"] is False
assert "application/json" in prefs["browser.helperApps.neverAsk.saveToDisk"]
def test_a_download_folder_must_be_under_home():
with pytest.raises(wd.WebDriverError):
wd.require_under_home(Path("/tmp/wp-c-downloads"))
inside = Path.home() / "v11-evidence" / "wp-c" / "downloads"
assert wd.require_under_home(inside) == inside.resolve()
File diff suppressed because it is too large Load Diff
+135 -3
View File
@@ -15,6 +15,12 @@ what M9 recorded as "this machine cannot drive a file into the browser". The
narrower and more useful statement is that it refuses `/tmp`: a path under the narrower and more useful statement is that it refuses `/tmp`: a path under the
user's home works. `stage()` exists to put evidence files there, so knowledge user's home works. `stage()` exists to put evidence files there, so knowledge
import can be exercised through the real file input rather than in two halves. import can be exercised through the real file input rather than in two halves.
**Downloads (v1.1 WP-C).** The same snap Firefox saves a download without any
dialog when its profile says where to, and the folder is under `$HOME`.
`firefox_download_prefs` is that profile, `require_under_home` refuses a folder
the sandbox would not let it write, and `wait_for_download` decides when a file
has actually finished arriving — never the click that started it.
""" """
from __future__ import annotations from __future__ import annotations
@@ -33,6 +39,10 @@ GECKODRIVER = shutil.which("geckodriver") or "/snap/bin/geckodriver"
#: Where files the browser must open are staged. Under $HOME because the snap #: Where files the browser must open are staged. Under $HOME because the snap
#: sandbox denies /tmp; see the module docstring. #: sandbox denies /tmp; see the module docstring.
STAGE = Path.home() / "m11-evidence" STAGE = Path.home() / "m11-evidence"
#: The W3C key an element reference is returned under.
ELEMENT_KEY = "element-6066-11e4-a52e-4f735466cecf"
#: What Firefox names a download while it is still arriving.
PARTIAL_SUFFIXES = (".part",)
def stage(name: str, body: str | bytes) -> str: def stage(name: str, body: str | bytes) -> str:
@@ -51,14 +61,86 @@ def free_port() -> int:
return s.getsockname()[1] return s.getsockname()[1]
def geckodriver_version() -> str:
try:
out = subprocess.run([GECKODRIVER, "--version"], capture_output=True, text=True, timeout=30)
return (out.stdout.splitlines() or ["?"])[0].strip()
except (OSError, subprocess.SubprocessError):
return "?"
class WebDriverError(RuntimeError): class WebDriverError(RuntimeError):
pass pass
# ------------------------------------------------------------------ downloads
def require_under_home(path: Path) -> Path:
"""`path`, resolved, if it is inside the user's home; otherwise refuse.
The snap sandbox will not write elsewhere, and a download folder under
`/tmp` would also put evidence where a reboot deletes it.
"""
resolved = Path(path).expanduser().resolve()
home = Path.home().resolve()
if resolved != home and home not in resolved.parents:
raise WebDriverError(f"{resolved} is not under {home}; the browser cannot write there")
return resolved
def firefox_download_prefs(directory: Path) -> dict:
"""Profile preferences that save every download to `directory`, unasked."""
return {
"browser.download.folderList": 2, # 2 = the folder named below
"browser.download.dir": str(directory),
"browser.download.useDownloadDir": True,
"browser.download.start_downloads_in_tmp_dir": False,
"browser.download.always_ask_before_handling_new_types": False,
"browser.helperApps.neverAsk.saveToDisk": "application/json,application/octet-stream",
"browser.download.manager.showWhenStarting": False,
"browser.download.alwaysOpenPanel": False,
"browser.download.panel.shown": True,
}
def wait_for_download(directory: Path, before: set[str], *, timeout: float = 60,
poll: float = 0.2, stable_polls: int = 3) -> Path:
"""The file a download wrote into `directory`, once it has finished.
Finished means all of these, at once:
- a name that was not in `before` (the listing taken before the click);
- no in-progress file (`*.part`) left in the folder;
- more than zero bytes;
- the same size for `stable_polls` consecutive polls.
A first appearance is not a finished download, and a zero-byte or partial
file never counts. Raises `WebDriverError` when nothing finishes in time.
"""
deadline = time.monotonic() + timeout
last: dict[str, int] = {}
steady: dict[str, int] = {}
while time.monotonic() < deadline:
names = {p.name for p in directory.iterdir()} if directory.exists() else set()
partial = any(n.endswith(PARTIAL_SUFFIXES) for n in names)
fresh = sorted(n for n in names - before if not n.endswith(PARTIAL_SUFFIXES))
for name in fresh:
size = (directory / name).stat().st_size
steady[name] = steady.get(name, 0) + 1 if last.get(name) == size else 1
last[name] = size
if not partial and size > 0 and steady[name] >= stable_polls:
return directory / name
time.sleep(poll)
listing = sorted(p.name for p in directory.iterdir()) if directory.exists() else []
raise WebDriverError(f"no finished download in {directory} within {timeout}s; saw {listing}")
# -------------------------------------------------------------------- browser
class Browser: class Browser:
"""One headless Firefox, driven over the wire protocol.""" """One headless Firefox, driven over the wire protocol."""
def __init__(self, *, headless: bool = True, log: Path | None = None): def __init__(self, *, headless: bool = True, log: Path | None = None,
download_dir: Path | None = None):
self.port = free_port() self.port = free_port()
handle = open(log, "ab") if log else subprocess.DEVNULL handle = open(log, "ab") if log else subprocess.DEVNULL
self.proc = subprocess.Popen( self.proc = subprocess.Popen(
@@ -68,9 +150,15 @@ class Browser:
self.base = f"http://127.0.0.1:{self.port}" self.base = f"http://127.0.0.1:{self.port}"
self._wait_for_driver() self._wait_for_driver()
args = ["-headless"] if headless else [] args = ["-headless"] if headless else []
options: dict = {"args": args}
self.download_dir = None
if download_dir is not None:
self.download_dir = require_under_home(download_dir)
self.download_dir.mkdir(parents=True, exist_ok=True)
options["prefs"] = firefox_download_prefs(self.download_dir)
answer = self._call("POST", "/session", {"capabilities": {"alwaysMatch": { answer = self._call("POST", "/session", {"capabilities": {"alwaysMatch": {
"browserName": "firefox", "browserName": "firefox",
"moz:firefoxOptions": {"args": args}, "moz:firefoxOptions": options,
# Never silently accept a bad certificate: the endpoint policy and # Never silently accept a bad certificate: the endpoint policy and
# the TLS trust union are release claims (H12, A06), and a browser # the TLS trust union are release claims (H12, A06), and a browser
# that ignored certificates would hide a failure of either. # that ignored certificates would hide a failure of either.
@@ -78,6 +166,7 @@ class Browser:
}}})["value"] }}})["value"]
self.session = answer["sessionId"] self.session = answer["sessionId"]
self.version = answer["capabilities"].get("browserVersion", "?") self.version = answer["capabilities"].get("browserVersion", "?")
self.capabilities = answer["capabilities"]
# ------------------------------------------------------------- plumbing # ------------------------------------------------------------- plumbing
@@ -123,6 +212,9 @@ class Browser:
def go(self, url: str) -> None: def go(self, url: str) -> None:
self._call("POST", self._s("/url"), {"url": url}) self._call("POST", self._s("/url"), {"url": url})
def reload(self) -> None:
self._call("POST", self._s("/refresh"), {})
@property @property
def url(self) -> str: def url(self) -> str:
return self._call("GET", self._s("/url"))["value"] return self._call("GET", self._s("/url"))["value"]
@@ -138,6 +230,13 @@ class Browser:
return self._call("POST", self._s("/execute/sync"), return self._call("POST", self._s("/execute/sync"),
{"script": script, "args": list(args)})["value"] {"script": script, "args": list(args)})["value"]
def element_by_js(self, script: str, *args):
"""An element a script returns, as a reference `click` can use, or None."""
value = self.js(script, *args)
if isinstance(value, dict) and ELEMENT_KEY in value:
return value[ELEMENT_KEY]
return None
def find(self, css: str, *, required=True): def find(self, css: str, *, required=True):
try: try:
answer = self._call("POST", self._s("/element"), answer = self._call("POST", self._s("/element"),
@@ -163,6 +262,17 @@ class Browser:
return self._call("GET", self._s(f"/element/{element}/property/{name}"))["value"] return self._call("GET", self._s(f"/element/{element}/property/{name}"))["value"]
def click(self, element: str) -> None: def click(self, element: str) -> None:
"""A real click, on an element first scrolled to the middle of the view.
WebDriver scrolls a target only as far as its edge, and the play page's
composer is fixed to the bottom of the window: a control just under it
(a failure notice's details, a turn's Inspect button) is then covered,
and the click is intercepted. A reader scrolls it clear first; so does
this (v1.1 WP-C).
"""
self._call("POST", self._s("/execute/sync"), {
"script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
"args": [{ELEMENT_KEY: element}]})
self._call("POST", self._s(f"/element/{element}/click"), {}) self._call("POST", self._s(f"/element/{element}/click"), {})
def clear(self, element: str) -> None: def clear(self, element: str) -> None:
@@ -183,6 +293,21 @@ class Browser:
answer = self._call("GET", self._s("/element/active")) answer = self._call("GET", self._s("/element/active"))
return list(answer["value"].values())[0] return list(answer["value"].values())[0]
# -------------------------------------------------------------- windows
@property
def window(self) -> str:
return self._call("GET", self._s("/window"))["value"]
def new_tab(self) -> str:
return self._call("POST", self._s("/window/new"), {"type": "tab"})["value"]["handle"]
def switch_to(self, handle: str) -> None:
self._call("POST", self._s("/window"), {"handle": handle})
def close_window(self) -> None:
self._call("DELETE", self._s("/window"))
# ------------------------------------------------------------- waiting # ------------------------------------------------------------- waiting
def wait_for(self, css: str, *, timeout=90, gone=False): def wait_for(self, css: str, *, timeout=90, gone=False):
@@ -196,12 +321,19 @@ class Browser:
f"{'still present' if gone else 'never appeared'}: {css}") f"{'still present' if gone else 'never appeared'}: {css}")
def wait_until(self, script: str, *, timeout=90, what=""): def wait_until(self, script: str, *, timeout=90, what=""):
if self.wait_js(script, timeout=timeout):
return True
raise WebDriverError(f"condition never held: {what or script}")
def wait_js(self, script: str, *, timeout=90) -> bool:
"""Whether `script` became true within `timeout`. For a check to record,
where `wait_until` is for a precondition that must hold."""
deadline = time.monotonic() + timeout deadline = time.monotonic() + timeout
while time.monotonic() < deadline: while time.monotonic() < deadline:
if self.js(f"return ({script})"): if self.js(f"return ({script})"):
return True return True
time.sleep(0.25) time.sleep(0.25)
raise WebDriverError(f"condition never held: {what or script}") return False
class Site: class Site:
+19
View File
@@ -45,6 +45,25 @@ export function classifyError(message) {
const detail = String(message || '').trim() || 'No detail was reported.' const detail = String(message || '').trim() || 'No detail was reported.'
const low = detail.toLowerCase() const low = detail.toLowerCase()
// ---- A correction the story refused ----
//
// v1.1 WP-C, found in the browser. The State panel's refusals ("That
// correction can't be applied — no fact 'f1' to invalidate.") matched none of
// the rules below and fell through to "Generation failed", with a Retry button
// and a line about what you typed. No turn was attempted and nothing was
// typed: the reader corrected the story's state and the rules refused it. So
// it is said as that, and the reason stays under the technical details.
if (low.includes("correction can't be applied")) {
return {
kind: KIND.STATE,
title: 'That correction was not applied',
detail,
hint: 'Nothing in the story or its state was changed. The reason is in the technical details.',
retryable: false,
keptInput: false,
}
}
// ---- Model / endpoint: the story cannot be told at all ---- // ---- Model / endpoint: the story cannot be told at all ----
// "No model configured — set one in Settings." // "No model configured — set one in Settings."
+7 -3
View File
@@ -55,9 +55,13 @@ export function FailureNotice({ failure, partial, onRetry, onDismiss }) {
{failure.hint && <p className="failure-hint">{failure.hint}</p>} {failure.hint && <p className="failure-hint">{failure.hint}</p>}
<p className="failure-kept"> {/* Every failed turn keeps what was typed (A05). A refused correction
Your story is unchanged, and what you typed is still in the box below. had nothing typed in the box, so it does not claim to (v1.1 WP-C). */}
</p> {failure.keptInput !== false && (
<p className="failure-kept">
Your story is unchanged, and what you typed is still in the box below.
</p>
)}
{partial ? ( {partial ? (
<details className="failure-partial"> <details className="failure-partial">
+35 -2
View File
@@ -29,8 +29,8 @@ beforeEach(() => { vi.restoreAllMocks() })
describe('the failure notice states only what is true', () => { describe('the failure notice states only what is true', () => {
it('claims the typed words were kept — so the caller must actually keep them', async () => { it('claims the typed words were kept — so the caller must actually keep them', async () => {
// This assertion is the contract the defect broke. The notice is // This assertion is the contract the defect broke. Every failed turn shows
// unconditional, so every path that shows it owes the reader their text. // it, so every path that shows it owes the reader their text.
await renderWith( await renderWith(
<FailureNotice failure={classifyError('Could not connect to http://127.0.0.1:9/v1')} <FailureNotice failure={classifyError('Could not connect to http://127.0.0.1:9/v1')}
partial={null} onRetry={vi.fn()} onDismiss={vi.fn()} />, partial={null} onRetry={vi.fn()} onDismiss={vi.fn()} />,
@@ -64,6 +64,39 @@ describe('the failure notice states only what is true', () => {
expect(onRetry).toHaveBeenCalled() expect(onRetry).toHaveBeenCalled()
}) })
// v1.1 WP-C, found in the browser: a correction the story refused fell through
// to "Generation failed", offered to try the turn again, and claimed the
// typed words were kept — none of which a State-panel correction involves.
const REFUSAL = "That correction can't be applied — no fact 'f1' to invalidate."
it('classifies a refused correction as a state refusal, not a failed turn', () => {
const refusal = classifyError(REFUSAL)
expect(refusal.kind).toBe(KIND.STATE)
expect(refusal.title).toBe('That correction was not applied')
expect(refusal.retryable).toBe(false)
expect(refusal.keptInput).toBe(false)
expect(refusal.detail).toBe(REFUSAL)
})
it('shows a refused correction with its reason, and no turn to retry', async () => {
await renderWith(
<FailureNotice failure={classifyError(REFUSAL)} partial={null}
onRetry={vi.fn()} onDismiss={vi.fn()} />,
)
expect(screen.getByText('That correction was not applied')).toBeInTheDocument()
expect(screen.getByText(/no fact 'f1' to invalidate/)).toBeInTheDocument()
expect(screen.queryByRole('button', { name: 'Try that turn again' })).toBeNull()
expect(screen.queryByText(/what you typed is still in the box/)).toBeNull()
})
it('still claims the typed words were kept for a failed turn', async () => {
await renderWith(
<FailureNotice failure={classifyError('boom')} partial={null}
onRetry={vi.fn()} onDismiss={vi.fn()} />,
)
expect(screen.getByText(/what you typed is still in the box/)).toBeInTheDocument()
})
it('labels partial prose as not kept rather than showing it as story', async () => { it('labels partial prose as not kept rather than showing it as story', async () => {
await renderWith( await renderWith(
<FailureNotice failure={classifyError('boom')} partial="The tavern door swung" <FailureNotice failure={classifyError('boom')} partial="The tavern door swung"
+4 -5
View File
@@ -2,11 +2,10 @@
**This file is the index. Start here.** **This file is the index. Start here.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: WP-A1 **Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress:
and WP-A2 are committed (`d63804f`), the WP-B.1 memory diagnostic is committed WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`) are committed, and
(`beb17ad`), and WP-B.2, the memory-retention fix, is complete and staged for the owner's WP-C, browser release coverage, is complete and staged for owner review**
signed commit, accepted with a documented reference-model memory limitation** (`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`, (`reports/v1.1/V1.1-WP-C-REPORT.md`).
`reports/v1.1/V1.1-WP-B1-REPORT.md`, `reports/v1.1/V1.1-WP-B2-REPORT.md`).
Phase 0 complete; AI-DnD forked as the production base; **milestones M1 Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the (2026-09-14), the v1 release gate passed on the release-candidate tree, and the
+5 -1
View File
@@ -15,7 +15,11 @@ real-model limitation** (owner decision, 2026-09-15):
- the B2.4 prompt experiment did not fix that and was reverted. - the B2.4 prompt experiment did not fix that and was reverted.
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T. The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
WP-C, WP-D and WP-E have not started. No v1.1 version or tag exists. **WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release
coverage, is complete and staged for owner review. Its final run passed 91 checks
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E
have not started. No v1.1 version or tag exists.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1 This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1 history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
+18 -3
View File
@@ -1,8 +1,23 @@
# Planning Package Version # Planning Package Version
- **Package:** Adventure Storyteller Planning Package v4.2 - **Package:** Adventure Storyteller Planning Package v4.4
- **Revision date:** 2026-09-14 - **Revision date:** 2026-09-15
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1 and WP-A2 are committed (`d63804f`) and WP-B.1 is committed (`beb17ad`). **WP-B.2 is complete and staged for the owner's signed commit, accepted with a documented real-model limitation: B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship; deterministic independent recovery passes; the reference model failed at memory creation on the precondition-valid attempt, and the B2.4 prompt experiment was rejected and reverted** (`reports/v1.1/V1.1-WP-B2-REPORT.md` §S, §T). WP-C, WP-D and WP-E have not started, and no v1.1 version or tag exists. - **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. **WP-C, browser release coverage, is complete and staged for owner review: 91 checks, 0 failed, 0 skipped, over trusted-LAN HTTPS, with real export downloads** (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E have not started, and no v1.1 version or tag exists.
## v4.4 — WP-C browser release coverage (2026-09-15)
Harness work, with one narrow product fix it found. No requirement, acceptance
test, schema or bundle format changed. The evidence is in
`reports/v1.1/V1.1-WP-C-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `reports/v1.1/V1.1-WP-C-REPORT.md` | **New.** The six reader workflows driven in a real browser, the existing 38 checks, the download environment, harness defects J1-J7, and product defects K1 (open) and K2 (fixed) | work-package report |
| `V1.1-PLAN.md` | Status: WP-B.2 committed; WP-C complete and staged | status |
| `planning/README.md` | Current state | index |
| `DEVELOPMENT.md` | The browser harness command and flags; Firefox download preferences and the `$HOME` rule; what counts as a finished download; waiting on conditions, never sleeping | developer docs |
**Requirement changes: zero.**
## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15) ## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15)
+475
View File
@@ -0,0 +1,475 @@
# v1.1 WP-C — Browser Release Coverage
**Status:** COMPLETE, staged for owner review. The final run passed 91 checks with 0 failed and 0 skipped. The decision is in §R.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD at start | `0c1ba836babe1447ad3693b4d95189a325b7b3b6` — *v1.1 WP-B.2: independent long-term memory retention*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
| Its ancestry | `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `ac465ed` (plan v4.1), `432f041` (v1.0.0) |
| Working tree at start | clean; nothing staged |
| WP-D, WP-E | not started |
---
## B. Existing browser harness
`backend/tools/m11_browser.py` drives Firefox through `backend/tools/m11_webdriver.py`, a
dependency-free W3C WebDriver client. It builds its fixture campaign through the API,
plays two narrator turns through the streaming endpoint, and then asserts on the
rendered DOM. The M11 closeout run (`3652dc6`, run 3) passed **38/38**, 0 failed,
0 skipped, on snap Firefox 155.0.1 with geckodriver 0.37.1 and `qwen2.5:3b-instruct`
over trusted-LAN HTTPS.
What it did not do, and WP-C closes: drive Retry and takes, Save Point create and
restore, state correction, narration length or failed generation through the UI,
and prove an export leaves the browser as a file.
Before WP-C it also slept, for a fixed time, before several assertions: after Undo
and Redo, after planting hostile narration, after choosing a knowledge file, around
the delete dialog, and around opening a panel. §J records that as a harness defect.
---
## C. Browser and download environment
| | |
| --- | --- |
| Firefox | **155.0.1, the snap** (`/snap/bin/firefox`). No second Firefox was installed |
| geckodriver | 0.37.1 (the snap) |
| Headless | yes |
| Frontend | the production build (`vite build`) served by FastAPI on loopback; no Vite dev server |
| Certificates | `acceptInsecureCerts: false`, unchanged |
**The download profile.** `m11_webdriver.firefox_download_prefs` is passed as
`moz:firefoxOptions.prefs`:
- `browser.download.folderList` 2, `browser.download.dir` `<--out>/downloads`,
`browser.download.useDownloadDir` true;
- `browser.download.start_downloads_in_tmp_dir` false;
- `browser.download.always_ask_before_handling_new_types` false;
- `browser.helperApps.neverAsk.saveToDisk` `application/json,application/octet-stream`;
- the download panel suppressed.
**The folder.** `--out/downloads`, which must be under `$HOME`
(`require_under_home`). The harness deletes it at the start of a run and creates it
fresh. Evidence lives under `$HOME/v11-evidence/wp-c/`.
**Does the snap Firefox download?** Yes. Measured first, with a probe
(`$HOME/v11-evidence/wp-c/probe/`): a loopback page runs the product's own download
pattern (a JSON blob, an `<a download>` click, an immediate `revokeObjectURL`).
- Download folder under `~/v11-evidence`: 33-byte file written.
- Download folder under `~/Downloads`: 33-byte file written.
The first probe wrote nothing, and that was a probe defect (J1), not the snap. The
non-snap fallback the plan allows was therefore not needed.
**When a download counts as finished** (`m11_webdriver.wait_for_download`, tested
without a browser in `test_v11_c_browser_helpers.py`, 7 tests). All of these at once:
- a name absent from the listing taken before the click;
- no `*.part` file in the folder;
- more than zero bytes;
- the same size across 3 consecutive polls.
A zero-byte, partial, pre-existing or still-growing file never counts, and neither
does the "Campaign exported." toast.
---
## D. Retry scenario
Real narration: **yes**. All checks use the reader-facing controls on the play page.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| Type in "What you do next", press **Send** | a new narration renders and the page is idle; its exact text is recorded | PASS |
| — | **Retry** is offered (enabled) on the newest narration | PASS |
| Press **Retry** | the take indicator on the newest narration reads **2/2** | PASS |
| — | the second take's text differs from the first (otherwise 1/2 could not be told from 2/2) | PASS |
| Press **‹** (Previous take) | the indicator reads **1/2**, and the narration is identical to the first recorded text | PASS |
| — | the second take's text is not shown anywhere in the transcript | PASS |
| Press **›** (Next take) | **2/2**, showing the second take | PASS |
| Reload the page | the indicator still reads **2/2** on the live take, showing the second take | PASS |
| Press **‹** after the reload | **1/2** still shows the first text, unchanged | PASS |
All text comparisons are of the rendered `.turn-text`. No database was read.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## E. Save Point scenario
Real narration: **yes** (two turns after the Save Point).
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| Press **Save Point**, type a name in "Save this moment", submit | a row with that exact name appears | PASS |
| — | the row's moment is the moment being read ("Moment 7") | PASS |
| **Send** two turns | two new narrations render; position "Moment 11" | PASS |
| In Save Points, press **Restore**, then confirm **Restore** | the position reads "Moment 7 · later story ahead" | PASS |
| — | the transcript ends at the Save Point: its last narration is the one read there, and neither later narration is shown | PASS |
| — | the position says later story is ahead | PASS |
| — | **Redo** is enabled | PASS |
| Press **Redo** until it is disabled (each press waits for the position to change) | the last two narrations are the two later turns, text-identical, at "Moment 11" | PASS |
| Reload the page | the position is still "Moment 11" | PASS |
| — | the named Save Point is still listed | PASS |
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## F. State-correction scenario
**Owner decision (2026-09-15).** The State panel cannot produce a *partly* refused
correction:
- "Save correction" sends exactly one `add_fact`, and "That's wrong" sends exactly
one `invalidate_fact`;
- so a correction is applied whole or refused whole (HTTP 400);
- the route's partial application (`refused` alongside applied changes) is
reachable only through the API.
WP-C drives what the reader can reach: an accepted correction that persists, and a
refused correction whose refusal and reason are visible, with the refused change
not applied. It records the partial refusal as unreachable from the reader UI. No
product change was made for it.
Real narration: **yes** for the refusal (it needs history to step back over); the
accepted correction needs none and also passed in the no-narrator smoke runs.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| In **State**, press **Correct something**, type a fact, press **Save correction** | the form closes and the State panel shows the fact | PASS |
| Reload, open **State** | the fact is still shown | PASS |
| In a second tab press **Undo** (the correction belongs to the moment it was made at); in the first tab, whose panel still shows the fact, press **That's wrong** on it | a failure notice carrying the correction's refusal ("can't be applied") appears | PASS |
| — | it is labelled **"That correction was not applied"**, with no "Try that turn again" and no claim that typed text was kept (§K2) | PASS |
| Open **Show technical details** | the reason is visible: "That correction can't be applied — no fact 'f3' to invalidate." | PASS |
| Second tab **Redo**, close it; reload the first tab at the corrected moment | the fact still stands, with its **That's wrong** control: the refused withdrawal was not applied | PASS |
**Partial refusal.** As the owner decided, a *partly* refused correction is not
reachable from the reader UI, and was not driven. The route's partial application
remains covered by the backend suite.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## G. Narration-length scenario
Real narration: **yes** (one turn per band). The model's actual length is not
asserted, because nothing in the product contract requires it. What is asserted is
that the chosen band reached the turn's own prompt.
The expected sentence comes from the product's own `builder.length_hint` at the
run's 400-token output cap:
- **brief:** "must not exceed 180 words, and it should not stop short of about 70";
- **long:** "must not exceed 236 words, and it should not stop short of about 118".
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| **Settings** panel: choose **Brief** in "Narration length", press **Save changes** | the button reads "Saved" | PASS |
| **Send** a turn; choose **Long**, **Save changes**; **Send** a second turn | both turns render | PASS (both) |
| On the brief turn press **Inspect context** | the prompt sections of "The exact text the narrator was sent" contain brief's range, and do **not** contain that turn's own reply (so this is the turn's record, not the dry run of the next) | PASS |
| On the long turn press **Inspect context** | the same, with long's range | PASS |
| Reload, open **Settings** | "Narration length" still reads **long** | PASS |
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## H. Failed-generation scenario
**Owner decision (2026-09-15).** The literal sequence (save an unserved model, then
submit a turn) cannot be driven. The model check (`modelStatus.jsx`) marks a
configured model absent from the endpoint's list as `missing-model`, and
`blocksPlay` disables Send, Continue and Retry up front. That is M8's intended
behaviour, not a defect. So WP-C drives both of these:
1. **The unserved model** saved through Settings: the reader is told and cannot
send, and the story is unchanged.
2. **A submitted failure:** a model the endpoint lists but that cannot narrate,
`nomic-embed-text:latest`. Send stays enabled, and the turn fails in the open.
Then recovery with the reference model. The report records that path 2 uses a
listed, non-narrating model rather than "a name the server does not serve".
Real narration: **yes**.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| **Settings**: "Type a model name instead", type an unserved name, **Save** | the page says "Saved" | PASS |
| Open the campaign | the header model status is `missing-model` and the setup notice is shown | PASS |
| — | **Send** and **Continue** are disabled | PASS |
| — | the story is unchanged | PASS |
| **Settings**: choose `nomic-embed-text:latest` from the installed-model picker, **Save** | "Saved" | PASS |
| Open the campaign (model status `ready`); type a turn; **Send** | a failure notice is shown: "Generation failed", with the server's reason under the details, `"nomic-embed-text:latest" does not support chat` (HTTP 400) | PASS |
| — | no narration was added | PASS |
| — | the typed text is still in the input box | PASS |
| — | the earlier story is text-identical | PASS |
| Open **State** | the rendered state is identical to before the failure | PASS |
| **Settings**: choose `qwen2.5:3b-instruct`, **Save** | "Saved" | PASS |
| Open the campaign; type a turn; **Send** | exactly one new narration | PASS |
| — | the earlier story is intact | PASS |
| Reload | the successful turn is still the last narration | PASS |
The failed turn's player moment stays in the transcript, as A05 intends
(`player_moment_kept_in_transcript` in the report).
**Deviation, owner-approved.** Path 2 uses a model the endpoint *lists* but that
cannot narrate, not "a name the server does not serve", because an unserved name
is caught before a turn can be submitted. Endpoint policy was not bypassed: both
paths use the same trusted-LAN HTTPS endpoint, and only the model name changed.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## I. Export-download scenario
Real narration: not needed for the download itself. In the final run the exported
campaign holds 19 moments of real narration, two takes and a Save Point.
**C6a — the campaign library**
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| On the library page, press **Export** on the "Release Regression" card | a new file is written to `downloads/`; it is finished (no `.part`, stable size) | PASS |
| — | it is not empty: **136,739 bytes** | PASS |
| Parse the file | `format` is **`ai-dnd-adventure-v3`** | PASS |
| Start a **fresh application** (new database); press **Import campaign**; give its file input the downloaded path | the browser lands on the imported campaign's play page | PASS |
| Compare what the reader sees of the import with what the reader saw of the original (library card, position, Save Points panel) | same title ("Release Regression") | PASS |
| — | same number of moments (19) | PASS |
| — | same position ("Moment 18") | PASS |
| — | same Save Points, which also match the file | PASS |
The file's `headDepth` is 17, which the play page shows as "Moment 18", and its
`actions` count is 19.
**C6b — campaign settings**
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| In the campaign's **Settings** panel, press **Export campaign** | a second, separate file is written and finished | PASS |
| — | not empty: 136,739 bytes | PASS |
| Parse the file | `format` is `ai-dnd-adventure-v3` | PASS |
Both files: `$HOME/v11-evidence/wp-c/final/``downloads/library-export.json` and `downloads/settings-export.json`.
The imported application's log is `import-server.log`.
---
## J. Harness defects found
| # | Defect | How found | Fix |
| --- | --- | --- | --- |
| J1 | The download probe's page wrote `URL.createObjectURL` inside an inline `onclick`, where `URL` is `document.URL`, a string. No blob was made, so it looked exactly like "the snap cannot download" | `gecko.log`: `TypeError: URL.createObjectURL is not a function` | The probe's script uses `window.URL`, the way the product's module does. Both download folders then worked |
| J2 | C3's first refusal withdrew a fact in a second tab and withdrew it again in the stale tab. Withdrawing keeps the fact, marked `invalidated` (C04's audit record), so the second withdrawal was valid and accepted. Nothing was refused, and "the refused change was not applied" passed without meaning anything | Smoke run: two C3 failures; `server.log` shows 201 for every correction; `narrative/apply.py` | The second tab steps the story back past the correction with Undo, so the stale tab withdraws a fact the story at that position does not have. The validator then refuses it: `no fact … to invalidate` |
| J3 | Clicks on a control just under the play page's fixed composer were intercepted ("Show technical details" in a failure notice, a turn's "Inspect context") | Dev run 1: `element click intercepted` | `Browser.click` scrolls the element to the centre of the view, then uses the real WebDriver click |
| J4 | The Settings model field is a text box until the endpoint's model list arrives, then a picker. Choosing before the check finished raced that swap | Dev run 1: `no such element: input#model` | Wait for the header's model status to leave `checking` first |
| J5 | C3's "a refused correction is shown" waited for *any* failure notice, so a notice about something else would have passed | Dev run 1: it passed on a notice titled "Generation failed" (§K2) | It now requires the notice to carry this correction's refusal ("can't be applied"), and asserts how it is labelled |
| J7 | C4 required the band's sentence in the inspector and the turn's own reply to be absent, to tell a turn's record from the dry run. But the inspector also renders what came back (`raw_output`, "What came back, before the state block was removed") inside the same section, so the absence could never hold | Dev run 2: both C4 inspector checks failed. The stored records show the brief turn's `length_hint` section carrying "must not exceed 180 words … about 70", the long turn's carrying "…236 … 118", and every turn's reply present in `raw_output` | The sentence and the absence are read from the prompt sections only, excluding the "What came back" block |
| J6 | M11's checks slept before assertions: after Undo and Redo, after planting hostile narration (1 s), after choosing a knowledge file (0.5 s), around the delete dialog (0.8 s and 0.6 s), and around opening a panel (0.8 s). A sleep is not evidence of what it waited for | Reading the harness against the brief's rule | Each is now a wait on the condition the check needs: the position changed or returned, the planted text rendered, Import enabled, a dialog present or gone, a panel-specific element present. §38's absence check now first waits for the knowledge library to render |
**Checked and not a defect.** In dev run 2, the original take at depth 6 (action 9)
had no `length_hint` in its stored record, while its Retry (action 10) did. The
original take's record holds only the per-attempt fields
(`attempts.ATTEMPT_KEYS`: world state, narrative state, raw output, usage,
accounting), with no sections and no settings. That is how a take that is no
longer live is stored, not a prompt built without the range. Every live turn's
prompt carried its band.
Also added: a panel opens only if it is not already open, because a tab toggles
its panel closed, and it is recognised by an element only that panel renders, not
by its title, which the tab itself already shows.
---
## K. Product defects found
### K1 — "Correct" on an Important Facts row is always refused (not fixed)
The State panel offers **Correct** on every row with a key. On an Important Facts
row the key is the fact's id, and on the scene-summary row it is `"summary"`.
`saveCorrection` sends that key as `add_fact.subject`, and the validator checks
`subject` as an entity reference. So every such correction is refused.
**Reproduced deterministically** against the real application (scratch `TestClient`,
no browser, no model):
| Correction the panel sends | Result |
| --- | --- |
| "Correct" on the Characters row (`subject='mara'`) | **201**, applied |
| "Correct" on the Important Facts row (`subject='f1'`) | **400** "That correction can't be applied — add_fact names subject='f1', which does not exist." |
**Not fixed in WP-C.** It does not prevent any WP-C behaviour: "Correct something"
and entity-row corrections work, and C3 uses them. A fix (offer Correct only
against entities, or send facts without a subject) is a small UX choice, left to
the owner as a v1.1 follow-up.
### K2 — a refused correction was presented as a failed turn (fixed)
**Found in the browser** (dev run 1, §J5). When the story refused a correction, the
failure notice said:
- the title **"Generation failed"**;
- the hint "Nothing was added to your story. You can try that turn again.";
- the button **Try that turn again**;
- the line "what you typed is still in the box below".
None of that is true of a State-panel correction. `classifyError` had no rule for
the server's refusal ("That correction can't be applied — …"), so it fell through
to the generation default. That falsified exactly what C3 checks: that a refusal
is shown to the reader as a refusal.
**Fix** (frontend only, no backend change):
- `errors.js`: one rule, checked first, for "correction can't be applied". It gives
`kind: state`, the title "That correction was not applied", the hint "Nothing in
the story or its state was changed. The reason is in the technical details.",
`retryable: false` and `keptInput: false`.
- `FailureNotice.jsx`: the "what you typed" line is shown unless a failure says
`keptInput: false`. Every other kind still shows it, so M8's A05 contract is
unchanged.
**Regression** (`failurePaths.test.jsx`, 3 tests):
- a refused correction classifies as a state refusal, not retryable, with the
reason kept;
- its notice shows the reason and no "Try that turn again" or typed-input claim;
- a failed turn still claims the typed words were kept.
In the browser, C3 now asserts the label, the absence of "Try that turn again" and
the absence of the typed-input line.
Frontend after the fix: **168/168** tests, lint exit 0 (15 pre-existing warnings, 0
errors, none in changed files), production build passes.
---
## L. Existing 38-check regression
**38/38 passed, 0 failed, 0 skipped** in the same run, tagged `M11`. The names are
identical to the M11 closeout run, so there is a one-to-one mapping and no check was
split, merged or dropped:
- B01 a turn is accepted (×2);
- A/UX: the tab title (×3);
- B position indicator (×3), D01 Undo, D04 Redo (×2);
- H06 (×3), H07, G09, H04;
- G01 knowledge import (×2);
- A11y dialog focus (×4);
- §38 narrator-only text absent;
- F05 context inspector;
- H10 (×2), H11/CSP (×2);
- A11y names, focus, tabindex, hover, contrast (×4), input focus.
What changed in them is how they wait (J6), not what they assert.
## M. New WP-C checks
**53/53 passed, 0 failed, 0 skipped**, tagged `WP-C`:
| Scenario | Checks | Result |
| --- | --- | --- |
| C1 Retry | 8 | 8 PASS |
| C2 Save Point | 9 | 9 PASS |
| C3 State correction | 6 | 6 PASS |
| C4 Narration length | 5 | 5 PASS |
| C5 Failed generation | 14 | 14 PASS |
| C6a Library export and import | 8 | 8 PASS |
| C6b Settings export | 3 | 3 PASS |
```text
existing M11 checks: 38/38
WP-C new checks: 53/53
failed: 0
skipped: 0
```
**Runs that are not the final evidence**, kept under `$HOME/v11-evidence/wp-c/`:
| Run | Where | Result | Why it is not evidence |
| --- | --- | --- | --- |
| `probe/` | loopback page | first: no download (J1); second: both folders written | environment probe |
| `smoke-1` | no narrator | 44 passed, 2 failed (J2), 6 skipped | partial |
| `smoke-2` | no narrator | 43 passed, 0 failed, 7 skipped | partial |
| `dev-gpu-1` | GPU host, plain HTTP | 71 passed, 3 failed (J3, J4; J5 found) | development, and not HTTPS |
| `dev-gpu-2` | GPU host, plain HTTP | 89 passed, 2 failed (J7) | development, and not HTTPS |
| `dev-gpu-3` | GPU host, `--only length` | 7 passed | development, and a subset |
## N. Production build / frontend verification
| | Result |
| --- | --- |
| Frontend suite (`npm test`) | **168/168**, 14 files. It was 165 before; the 3 new tests are K2's regressions |
| Lint (`npm run lint`, oxlint) | **exit 0: 0 errors**, 15 warnings. All are the pre-existing `only-export-components` kind, and none is in a file WP-C changed |
| Production build (`npm run build`) | passes; `dist/index.html` sha256 `62b6ea5eb4ce02a09f23cc1d48c335c2ada36b208b23c237338b39bf63a26cc5`, built 2026-09-15 15:11 from the final WP-C tree |
| What the browser ran | that build, served by FastAPI (`uvicorn app.main:app` on `127.0.0.1`); no Vite server |
| Backend product code | **unchanged**: nothing under `backend/app` is in the diff. The full backend suite was not rerun. The changed harness is tested by `test_v11_c_browser_helpers.py` (7 passed) |
| Offline regression | **23 passed, 0 failed** (`tools/m11_offline.py`, fresh `--no-cache` image, `--network none`, on the final WP-C tree), rerun because the frontend bundle changed (K2). Evidence: `$HOME/v11-evidence/wp-c/offline/` |
## O. Trusted-LAN / security
| | |
| --- | --- |
| Narrator | `qwen2.5:3b-instruct` on the CPU reference host, over **trusted-LAN HTTPS**. The certificate is from the private CA in this machine's trust store, and verified; there is no bypass |
| Window and A1 | every narrator turn (8): window **verified at 4,096**, accounting **`fits`**. None `exceeded` or `truncation_suspected` |
| Protocol echoes (A2) | none of the protocol shapes the harness looks for (a state fence, a hard-limit or reminder bracket, `Events: [`) appeared in any stored narration in this run. This is not a v1.1 protocol-leak result: the release gate owns that, and the mid-reply echo from WP-B.1 remains a separate residual |
| Storyteller | loopback only, both applications (the original and the fresh import) |
| Browser | `acceptInsecureCerts: false`; the CSP checks (H11) passed |
| Endpoint policy | unchanged; C5 changed only the model name, never the endpoint |
| Downloads | written only under `$HOME` (enforced); the harness makes no network request of its own beyond loopback and the configured endpoint |
| New dependencies | none: the harness still uses only `urllib`, no Selenium or Playwright |
| Real identifiers in committed files | none (scanned at staging) |
## P. Compatibility
| Area | Effect |
| --- | --- |
| Database schema, migrations | none |
| Bundle format | none: the downloads are `ai-dnd-adventure-v3` and import unchanged |
| History, Save Points, state semantics, memory, knowledge | none. WP-C drove them and changed nothing in them |
| Backend | no application code changed |
| Frontend | one behaviour change (K2): the refusal of a State-panel correction is labelled as a refusal rather than as a failed turn. Every other failure's classification, retry offer and typed-input claim is unchanged (M8's tests pass) |
| WP-A, WP-B | untouched |
## Q. Residual risks
| # | Risk |
| --- | --- |
| 1 | **K1**: "Correct" on an Important Facts or scene-summary row is always refused. Reproduced, not fixed: a small UX choice for the owner |
| 2 | **Partial refusal is API-only.** The reader UI cannot produce a partly refused correction (owner decision). The display for one exists but is unreachable from the panel's own controls |
| 3 | **C5 path 2 depends on the reference host listing an embedding model.** On a host without one, the submitted-failure path has no model to use |
| 4 | **Model nondeterminism in C1.** "The second take is a different narration" would fail if the model returned identical text for a retry. It did not in any run |
| 5 | **The harness runs on this machine's snap Firefox.** A different Firefox or a Chromium would need the download preferences re-checked |
| 6 | **The heuristic protocol-shape scan is not A2 evidence.** It is recorded only |
| 7 | The WP-B real-model memory limitation and the doubled-full-stop scene text are unchanged, and are not WP-C's |
## R. Final decision
```text
RETRY: PASS
SAVE POINT: PASS
STATE CORRECTION: PASS
NARRATION LENGTH: PASS
FAILED GENERATION: PASS
EXPORT DOWNLOAD — LIBRARY: PASS
EXPORT DOWNLOAD — SETTINGS: PASS
EXISTING BROWSER REGRESSION: PASS
WP-C NEW BROWSER COVERAGE: PASS
WP-C OVERALL:
PASS
```
**PASS**, on these grounds:
- the final run ended with **failed: 0, skipped: 0**;
- both export controls produced a real, finished, non-empty `ai-dnd-adventure-v3`
file on disk;
- the library file imported into a fresh application with the same title, moment
count, position and Save Points.
Two criteria were met in the form the owner approved (2026-09-15):
- **State correction:** an accepted correction, and a refused correction with its
reason. Partial refusal is recorded as unreachable from the reader UI.
- **Failed generation:** the up-front block for an unserved model, plus a submitted
failure with a listed model that cannot narrate.
**Product change:** K2, a mislabelled refusal, fixed narrowly with regression tests.
**Product defect left open:** K1.
Nothing is committed, pushed or tagged. WP-D and WP-E have not started.