v1.1 WP-C: browser release coverage
Drives in a real browser the reader workflows v1 proved only through the API or the component suite, including an export that leaves the browser as a file. Final run: 91 checks (the 38 existing M11 checks plus 53 new), 0 failed, 0 skipped, on the production build over trusted-LAN HTTPS. - tools/m11_browser.py: scenarios for Retry and takes, Save Point create / restore / Redo, state correction (accepted, and a refused correction with its reason), narration length reaching each turn's prompt, failed generation (an unserved model blocked up front; a listed model that cannot narrate failing in the open) and recovery, and export download from the library and from campaign settings, imported into a fresh application. Rows are tagged M11 / WP-C and counted separately; --only for development. The M11 checks now wait on conditions instead of sleeping. - tools/m11_webdriver.py: Firefox download preferences, a $HOME-only download folder, a download wait that ignores partial, empty, pre-existing and still-growing files, centred real clicks, tabs, and condition waits. - tests/test_v11_c_browser_helpers.py: the download wait, prefs and $HOME guard, without a browser. - frontend: a correction the story refused was presented as "Generation failed" with a Retry offer and a typed-input claim. It is now "That correction was not applied", not retryable, with the reason kept (errors.js, FailureNotice.jsx; 3 regression tests). - DEVELOPMENT.md: the harness command, download profile and $HOME rule, what counts as a finished download, and the no-sleep rule. - docs: V1.1-PLAN, VERSION v4.4, planning README, reports/v1.1/V1.1-WP-C-REPORT.md. Open for the owner: "Correct" on an Important Facts row is always refused (K1), and a partly refused correction is not reachable from the reader UI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
0c1ba836ba
commit
59b5ebc2d8
+38
-3
@@ -408,9 +408,17 @@ AIDND_TEST_ENDPOINT=... AIDND_TEST_MODEL=... AIDND_TEST_EMBED_MODEL=... \
|
||||
# What that campaign is worth on a machine that has never seen it (I01-I07).
|
||||
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"
|
||||
|
||||
# The browser release regression and the accessibility measurements. Needs
|
||||
# `frontend/dist` built and geckodriver on PATH.
|
||||
.venv/bin/python -m tools.m11_browser --out "$HOME/m11-evidence/browser"
|
||||
# The browser release regression and the accessibility measurements: M11's 38
|
||||
# checks plus v1.1 WP-C's reader workflows (Retry, Save Points, state correction,
|
||||
# narration length, failed generation, export download). Needs `frontend/dist`
|
||||
# built, geckodriver on PATH, and --out under $HOME (the downloads land inside
|
||||
# it). Release evidence needs the narrator over trusted-LAN HTTPS.
|
||||
AIDND_TEST_ENDPOINT=https://... AIDND_TEST_MODEL=qwen2.5:3b-instruct \
|
||||
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/browser"
|
||||
# Without a narrator (a partial smoke run, not evidence), or one scenario while
|
||||
# developing (--only takes: shell, history, markdown, hidden, context, csp, a11y,
|
||||
# retry, savepoint, state, length, failure, export).
|
||||
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/smoke" --no-narrator
|
||||
|
||||
# A container with no network at all: the offline run and the packaging path.
|
||||
.venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline"
|
||||
@@ -499,6 +507,33 @@ It exists so browser evidence needs no Selenium in the dependency surface, and
|
||||
it documents the one environment quirk that matters here: a snap Firefox will
|
||||
not open a file the driver names under `/tmp`, but will under `$HOME`.
|
||||
|
||||
**Downloads in the browser harness (v1.1 WP-C).** The export checks click the
|
||||
real Export controls and wait for the file on disk, so the browser has to save
|
||||
without asking. `m11_webdriver.firefox_download_prefs` gives the WebDriver
|
||||
session a profile that does that:
|
||||
- `browser.download.folderList` 2, `browser.download.dir` the run's
|
||||
`downloads/` folder, `browser.download.useDownloadDir` true;
|
||||
- no "always ask", and `application/json` saved to disk.
|
||||
|
||||
It works on the snap Firefox this machine has (155.0.1, geckodriver 0.37.1), and
|
||||
no separate Firefox is needed. The same sandbox rule applies as for opening
|
||||
files: the download folder must be under `$HOME`, and the harness refuses one
|
||||
that is not.
|
||||
|
||||
A download counts as finished only when all of these hold at once
|
||||
(`m11_webdriver.wait_for_download`):
|
||||
- a new name has appeared;
|
||||
- no `*.part` file is left;
|
||||
- the file is more than zero bytes;
|
||||
- its size is the same across consecutive polls.
|
||||
|
||||
The toast that says "Campaign exported." is not evidence.
|
||||
|
||||
**Waiting.** Nothing in the harness sleeps before an assertion. Every wait is on
|
||||
something the page, the browser or the filesystem shows. A condition that
|
||||
already holds before the action it waits for does not count as waiting for that
|
||||
action; the harness defects found in M8, M11 and WP-C were all of that shape.
|
||||
|
||||
## Backing up, and getting a campaign back
|
||||
|
||||
There are two recovery tools and they answer different questions. Using the
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
"""v1.1 WP-C: the browser harness's download helpers, without a browser.
|
||||
|
||||
`tools/m11_webdriver.wait_for_download` is what decides that an export actually
|
||||
left the browser as a file. It must never call a download finished because a
|
||||
file appeared, because it is still being written, or because it is empty — each
|
||||
of those would make "the export works" a claim the evidence does not support.
|
||||
|
||||
python -m pytest tests/test_v11_c_browser_helpers.py -v
|
||||
"""
|
||||
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from tools import m11_webdriver as wd
|
||||
|
||||
|
||||
def later(seconds, action):
|
||||
timer = threading.Timer(seconds, action)
|
||||
timer.start()
|
||||
return timer
|
||||
|
||||
|
||||
def test_a_finished_file_is_returned(tmp_path):
|
||||
later(0.2, lambda: (tmp_path / "campaign.json").write_text('{"format": "x"}'))
|
||||
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05)
|
||||
assert found == tmp_path / "campaign.json"
|
||||
|
||||
|
||||
def test_a_file_that_was_already_there_is_not_the_download(tmp_path):
|
||||
(tmp_path / "old.json").write_text("{}")
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.wait_for_download(tmp_path, {"old.json"}, timeout=0.6, poll=0.05)
|
||||
|
||||
|
||||
def test_an_empty_file_never_counts(tmp_path):
|
||||
(tmp_path / "empty.json").write_bytes(b"")
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
|
||||
|
||||
|
||||
def test_nothing_counts_while_firefox_is_still_writing(tmp_path):
|
||||
"""Firefox writes `<name>.part` beside the final name until it is done."""
|
||||
(tmp_path / "campaign.json").write_text('{"format": "x"}')
|
||||
(tmp_path / "campaign.json.part").write_text("")
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
|
||||
later(0.1, lambda: (tmp_path / "campaign.json.part").unlink())
|
||||
assert wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05).name == "campaign.json"
|
||||
|
||||
|
||||
def test_a_file_that_is_still_growing_is_not_finished(tmp_path):
|
||||
target = tmp_path / "big.json"
|
||||
target.write_text("{")
|
||||
stop = threading.Event()
|
||||
|
||||
def grow():
|
||||
for _ in range(8):
|
||||
if stop.is_set():
|
||||
return
|
||||
with target.open("a") as fh:
|
||||
fh.write("x" * 100)
|
||||
time.sleep(0.05)
|
||||
|
||||
writer = threading.Thread(target=grow)
|
||||
started = time.monotonic()
|
||||
writer.start()
|
||||
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05, stable_polls=3)
|
||||
writer.join()
|
||||
# Returned only once the size held still, so after the last write.
|
||||
assert found == target
|
||||
assert target.stat().st_size == 1 + 8 * 100
|
||||
assert time.monotonic() - started >= 0.4
|
||||
|
||||
|
||||
def test_the_prefs_save_downloads_unasked_to_the_folder_given(tmp_path):
|
||||
prefs = wd.firefox_download_prefs(tmp_path)
|
||||
assert prefs["browser.download.folderList"] == 2
|
||||
assert prefs["browser.download.dir"] == str(tmp_path)
|
||||
assert prefs["browser.download.useDownloadDir"] is True
|
||||
assert prefs["browser.download.always_ask_before_handling_new_types"] is False
|
||||
assert "application/json" in prefs["browser.helperApps.neverAsk.saveToDisk"]
|
||||
|
||||
|
||||
def test_a_download_folder_must_be_under_home():
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.require_under_home(Path("/tmp/wp-c-downloads"))
|
||||
inside = Path.home() / "v11-evidence" / "wp-c" / "downloads"
|
||||
assert wd.require_under_home(inside) == inside.resolve()
|
||||
+884
-137
File diff suppressed because it is too large
Load Diff
@@ -15,6 +15,12 @@ what M9 recorded as "this machine cannot drive a file into the browser". The
|
||||
narrower and more useful statement is that it refuses `/tmp`: a path under the
|
||||
user's home works. `stage()` exists to put evidence files there, so knowledge
|
||||
import can be exercised through the real file input rather than in two halves.
|
||||
|
||||
**Downloads (v1.1 WP-C).** The same snap Firefox saves a download without any
|
||||
dialog when its profile says where to, and the folder is under `$HOME`.
|
||||
`firefox_download_prefs` is that profile, `require_under_home` refuses a folder
|
||||
the sandbox would not let it write, and `wait_for_download` decides when a file
|
||||
has actually finished arriving — never the click that started it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -33,6 +39,10 @@ GECKODRIVER = shutil.which("geckodriver") or "/snap/bin/geckodriver"
|
||||
#: Where files the browser must open are staged. Under $HOME because the snap
|
||||
#: sandbox denies /tmp; see the module docstring.
|
||||
STAGE = Path.home() / "m11-evidence"
|
||||
#: The W3C key an element reference is returned under.
|
||||
ELEMENT_KEY = "element-6066-11e4-a52e-4f735466cecf"
|
||||
#: What Firefox names a download while it is still arriving.
|
||||
PARTIAL_SUFFIXES = (".part",)
|
||||
|
||||
|
||||
def stage(name: str, body: str | bytes) -> str:
|
||||
@@ -51,14 +61,86 @@ def free_port() -> int:
|
||||
return s.getsockname()[1]
|
||||
|
||||
|
||||
def geckodriver_version() -> str:
|
||||
try:
|
||||
out = subprocess.run([GECKODRIVER, "--version"], capture_output=True, text=True, timeout=30)
|
||||
return (out.stdout.splitlines() or ["?"])[0].strip()
|
||||
except (OSError, subprocess.SubprocessError):
|
||||
return "?"
|
||||
|
||||
|
||||
class WebDriverError(RuntimeError):
|
||||
pass
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ downloads
|
||||
|
||||
def require_under_home(path: Path) -> Path:
|
||||
"""`path`, resolved, if it is inside the user's home; otherwise refuse.
|
||||
|
||||
The snap sandbox will not write elsewhere, and a download folder under
|
||||
`/tmp` would also put evidence where a reboot deletes it.
|
||||
"""
|
||||
resolved = Path(path).expanduser().resolve()
|
||||
home = Path.home().resolve()
|
||||
if resolved != home and home not in resolved.parents:
|
||||
raise WebDriverError(f"{resolved} is not under {home}; the browser cannot write there")
|
||||
return resolved
|
||||
|
||||
|
||||
def firefox_download_prefs(directory: Path) -> dict:
|
||||
"""Profile preferences that save every download to `directory`, unasked."""
|
||||
return {
|
||||
"browser.download.folderList": 2, # 2 = the folder named below
|
||||
"browser.download.dir": str(directory),
|
||||
"browser.download.useDownloadDir": True,
|
||||
"browser.download.start_downloads_in_tmp_dir": False,
|
||||
"browser.download.always_ask_before_handling_new_types": False,
|
||||
"browser.helperApps.neverAsk.saveToDisk": "application/json,application/octet-stream",
|
||||
"browser.download.manager.showWhenStarting": False,
|
||||
"browser.download.alwaysOpenPanel": False,
|
||||
"browser.download.panel.shown": True,
|
||||
}
|
||||
|
||||
|
||||
def wait_for_download(directory: Path, before: set[str], *, timeout: float = 60,
|
||||
poll: float = 0.2, stable_polls: int = 3) -> Path:
|
||||
"""The file a download wrote into `directory`, once it has finished.
|
||||
|
||||
Finished means all of these, at once:
|
||||
- a name that was not in `before` (the listing taken before the click);
|
||||
- no in-progress file (`*.part`) left in the folder;
|
||||
- more than zero bytes;
|
||||
- the same size for `stable_polls` consecutive polls.
|
||||
|
||||
A first appearance is not a finished download, and a zero-byte or partial
|
||||
file never counts. Raises `WebDriverError` when nothing finishes in time.
|
||||
"""
|
||||
deadline = time.monotonic() + timeout
|
||||
last: dict[str, int] = {}
|
||||
steady: dict[str, int] = {}
|
||||
while time.monotonic() < deadline:
|
||||
names = {p.name for p in directory.iterdir()} if directory.exists() else set()
|
||||
partial = any(n.endswith(PARTIAL_SUFFIXES) for n in names)
|
||||
fresh = sorted(n for n in names - before if not n.endswith(PARTIAL_SUFFIXES))
|
||||
for name in fresh:
|
||||
size = (directory / name).stat().st_size
|
||||
steady[name] = steady.get(name, 0) + 1 if last.get(name) == size else 1
|
||||
last[name] = size
|
||||
if not partial and size > 0 and steady[name] >= stable_polls:
|
||||
return directory / name
|
||||
time.sleep(poll)
|
||||
listing = sorted(p.name for p in directory.iterdir()) if directory.exists() else []
|
||||
raise WebDriverError(f"no finished download in {directory} within {timeout}s; saw {listing}")
|
||||
|
||||
|
||||
# -------------------------------------------------------------------- browser
|
||||
|
||||
class Browser:
|
||||
"""One headless Firefox, driven over the wire protocol."""
|
||||
|
||||
def __init__(self, *, headless: bool = True, log: Path | None = None):
|
||||
def __init__(self, *, headless: bool = True, log: Path | None = None,
|
||||
download_dir: Path | None = None):
|
||||
self.port = free_port()
|
||||
handle = open(log, "ab") if log else subprocess.DEVNULL
|
||||
self.proc = subprocess.Popen(
|
||||
@@ -68,9 +150,15 @@ class Browser:
|
||||
self.base = f"http://127.0.0.1:{self.port}"
|
||||
self._wait_for_driver()
|
||||
args = ["-headless"] if headless else []
|
||||
options: dict = {"args": args}
|
||||
self.download_dir = None
|
||||
if download_dir is not None:
|
||||
self.download_dir = require_under_home(download_dir)
|
||||
self.download_dir.mkdir(parents=True, exist_ok=True)
|
||||
options["prefs"] = firefox_download_prefs(self.download_dir)
|
||||
answer = self._call("POST", "/session", {"capabilities": {"alwaysMatch": {
|
||||
"browserName": "firefox",
|
||||
"moz:firefoxOptions": {"args": args},
|
||||
"moz:firefoxOptions": options,
|
||||
# Never silently accept a bad certificate: the endpoint policy and
|
||||
# the TLS trust union are release claims (H12, A06), and a browser
|
||||
# that ignored certificates would hide a failure of either.
|
||||
@@ -78,6 +166,7 @@ class Browser:
|
||||
}}})["value"]
|
||||
self.session = answer["sessionId"]
|
||||
self.version = answer["capabilities"].get("browserVersion", "?")
|
||||
self.capabilities = answer["capabilities"]
|
||||
|
||||
# ------------------------------------------------------------- plumbing
|
||||
|
||||
@@ -123,6 +212,9 @@ class Browser:
|
||||
def go(self, url: str) -> None:
|
||||
self._call("POST", self._s("/url"), {"url": url})
|
||||
|
||||
def reload(self) -> None:
|
||||
self._call("POST", self._s("/refresh"), {})
|
||||
|
||||
@property
|
||||
def url(self) -> str:
|
||||
return self._call("GET", self._s("/url"))["value"]
|
||||
@@ -138,6 +230,13 @@ class Browser:
|
||||
return self._call("POST", self._s("/execute/sync"),
|
||||
{"script": script, "args": list(args)})["value"]
|
||||
|
||||
def element_by_js(self, script: str, *args):
|
||||
"""An element a script returns, as a reference `click` can use, or None."""
|
||||
value = self.js(script, *args)
|
||||
if isinstance(value, dict) and ELEMENT_KEY in value:
|
||||
return value[ELEMENT_KEY]
|
||||
return None
|
||||
|
||||
def find(self, css: str, *, required=True):
|
||||
try:
|
||||
answer = self._call("POST", self._s("/element"),
|
||||
@@ -163,6 +262,17 @@ class Browser:
|
||||
return self._call("GET", self._s(f"/element/{element}/property/{name}"))["value"]
|
||||
|
||||
def click(self, element: str) -> None:
|
||||
"""A real click, on an element first scrolled to the middle of the view.
|
||||
|
||||
WebDriver scrolls a target only as far as its edge, and the play page's
|
||||
composer is fixed to the bottom of the window: a control just under it
|
||||
(a failure notice's details, a turn's Inspect button) is then covered,
|
||||
and the click is intercepted. A reader scrolls it clear first; so does
|
||||
this (v1.1 WP-C).
|
||||
"""
|
||||
self._call("POST", self._s("/execute/sync"), {
|
||||
"script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
|
||||
"args": [{ELEMENT_KEY: element}]})
|
||||
self._call("POST", self._s(f"/element/{element}/click"), {})
|
||||
|
||||
def clear(self, element: str) -> None:
|
||||
@@ -183,6 +293,21 @@ class Browser:
|
||||
answer = self._call("GET", self._s("/element/active"))
|
||||
return list(answer["value"].values())[0]
|
||||
|
||||
# -------------------------------------------------------------- windows
|
||||
|
||||
@property
|
||||
def window(self) -> str:
|
||||
return self._call("GET", self._s("/window"))["value"]
|
||||
|
||||
def new_tab(self) -> str:
|
||||
return self._call("POST", self._s("/window/new"), {"type": "tab"})["value"]["handle"]
|
||||
|
||||
def switch_to(self, handle: str) -> None:
|
||||
self._call("POST", self._s("/window"), {"handle": handle})
|
||||
|
||||
def close_window(self) -> None:
|
||||
self._call("DELETE", self._s("/window"))
|
||||
|
||||
# ------------------------------------------------------------- waiting
|
||||
|
||||
def wait_for(self, css: str, *, timeout=90, gone=False):
|
||||
@@ -196,12 +321,19 @@ class Browser:
|
||||
f"{'still present' if gone else 'never appeared'}: {css}")
|
||||
|
||||
def wait_until(self, script: str, *, timeout=90, what=""):
|
||||
if self.wait_js(script, timeout=timeout):
|
||||
return True
|
||||
raise WebDriverError(f"condition never held: {what or script}")
|
||||
|
||||
def wait_js(self, script: str, *, timeout=90) -> bool:
|
||||
"""Whether `script` became true within `timeout`. For a check to record,
|
||||
where `wait_until` is for a precondition that must hold."""
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
if self.js(f"return ({script})"):
|
||||
return True
|
||||
time.sleep(0.25)
|
||||
raise WebDriverError(f"condition never held: {what or script}")
|
||||
return False
|
||||
|
||||
|
||||
class Site:
|
||||
|
||||
@@ -45,6 +45,25 @@ export function classifyError(message) {
|
||||
const detail = String(message || '').trim() || 'No detail was reported.'
|
||||
const low = detail.toLowerCase()
|
||||
|
||||
// ---- A correction the story refused ----
|
||||
//
|
||||
// v1.1 WP-C, found in the browser. The State panel's refusals ("That
|
||||
// correction can't be applied — no fact 'f1' to invalidate.") matched none of
|
||||
// the rules below and fell through to "Generation failed", with a Retry button
|
||||
// and a line about what you typed. No turn was attempted and nothing was
|
||||
// typed: the reader corrected the story's state and the rules refused it. So
|
||||
// it is said as that, and the reason stays under the technical details.
|
||||
if (low.includes("correction can't be applied")) {
|
||||
return {
|
||||
kind: KIND.STATE,
|
||||
title: 'That correction was not applied',
|
||||
detail,
|
||||
hint: 'Nothing in the story or its state was changed. The reason is in the technical details.',
|
||||
retryable: false,
|
||||
keptInput: false,
|
||||
}
|
||||
}
|
||||
|
||||
// ---- Model / endpoint: the story cannot be told at all ----
|
||||
|
||||
// "No model configured — set one in Settings."
|
||||
|
||||
@@ -55,9 +55,13 @@ export function FailureNotice({ failure, partial, onRetry, onDismiss }) {
|
||||
|
||||
{failure.hint && <p className="failure-hint">{failure.hint}</p>}
|
||||
|
||||
<p className="failure-kept">
|
||||
Your story is unchanged, and what you typed is still in the box below.
|
||||
</p>
|
||||
{/* Every failed turn keeps what was typed (A05). A refused correction
|
||||
had nothing typed in the box, so it does not claim to (v1.1 WP-C). */}
|
||||
{failure.keptInput !== false && (
|
||||
<p className="failure-kept">
|
||||
Your story is unchanged, and what you typed is still in the box below.
|
||||
</p>
|
||||
)}
|
||||
|
||||
{partial ? (
|
||||
<details className="failure-partial">
|
||||
|
||||
@@ -29,8 +29,8 @@ beforeEach(() => { vi.restoreAllMocks() })
|
||||
|
||||
describe('the failure notice states only what is true', () => {
|
||||
it('claims the typed words were kept — so the caller must actually keep them', async () => {
|
||||
// This assertion is the contract the defect broke. The notice is
|
||||
// unconditional, so every path that shows it owes the reader their text.
|
||||
// This assertion is the contract the defect broke. Every failed turn shows
|
||||
// it, so every path that shows it owes the reader their text.
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError('Could not connect to http://127.0.0.1:9/v1')}
|
||||
partial={null} onRetry={vi.fn()} onDismiss={vi.fn()} />,
|
||||
@@ -64,6 +64,39 @@ describe('the failure notice states only what is true', () => {
|
||||
expect(onRetry).toHaveBeenCalled()
|
||||
})
|
||||
|
||||
// v1.1 WP-C, found in the browser: a correction the story refused fell through
|
||||
// to "Generation failed", offered to try the turn again, and claimed the
|
||||
// typed words were kept — none of which a State-panel correction involves.
|
||||
const REFUSAL = "That correction can't be applied — no fact 'f1' to invalidate."
|
||||
|
||||
it('classifies a refused correction as a state refusal, not a failed turn', () => {
|
||||
const refusal = classifyError(REFUSAL)
|
||||
expect(refusal.kind).toBe(KIND.STATE)
|
||||
expect(refusal.title).toBe('That correction was not applied')
|
||||
expect(refusal.retryable).toBe(false)
|
||||
expect(refusal.keptInput).toBe(false)
|
||||
expect(refusal.detail).toBe(REFUSAL)
|
||||
})
|
||||
|
||||
it('shows a refused correction with its reason, and no turn to retry', async () => {
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError(REFUSAL)} partial={null}
|
||||
onRetry={vi.fn()} onDismiss={vi.fn()} />,
|
||||
)
|
||||
expect(screen.getByText('That correction was not applied')).toBeInTheDocument()
|
||||
expect(screen.getByText(/no fact 'f1' to invalidate/)).toBeInTheDocument()
|
||||
expect(screen.queryByRole('button', { name: 'Try that turn again' })).toBeNull()
|
||||
expect(screen.queryByText(/what you typed is still in the box/)).toBeNull()
|
||||
})
|
||||
|
||||
it('still claims the typed words were kept for a failed turn', async () => {
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError('boom')} partial={null}
|
||||
onRetry={vi.fn()} onDismiss={vi.fn()} />,
|
||||
)
|
||||
expect(screen.getByText(/what you typed is still in the box/)).toBeInTheDocument()
|
||||
})
|
||||
|
||||
it('labels partial prose as not kept rather than showing it as story', async () => {
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError('boom')} partial="The tavern door swung"
|
||||
|
||||
+4
-5
@@ -2,11 +2,10 @@
|
||||
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: WP-A1
|
||||
and WP-A2 are committed (`d63804f`), the WP-B.1 memory diagnostic is committed
|
||||
(`beb17ad`), and WP-B.2, the memory-retention fix, is complete and staged for the owner's
|
||||
signed commit, accepted with a documented reference-model memory limitation** (`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`,
|
||||
`reports/v1.1/V1.1-WP-B1-REPORT.md`, `reports/v1.1/V1.1-WP-B2-REPORT.md`).
|
||||
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress:
|
||||
WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`) are committed, and
|
||||
WP-C, browser release coverage, is complete and staged for owner review**
|
||||
(`reports/v1.1/V1.1-WP-C-REPORT.md`).
|
||||
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
|
||||
through M11 complete and closed**. M11 was accepted at its closeout
|
||||
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
|
||||
|
||||
@@ -15,7 +15,11 @@ real-model limitation** (owner decision, 2026-09-15):
|
||||
- the B2.4 prompt experiment did not fix that and was reverted.
|
||||
|
||||
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
|
||||
WP-C, WP-D and WP-E have not started. No v1.1 version or tag exists.
|
||||
**WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release
|
||||
coverage, is complete and staged for owner review. Its final run passed 91 checks
|
||||
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
|
||||
including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E
|
||||
have not started. No v1.1 version or tag exists.
|
||||
|
||||
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
|
||||
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
|
||||
|
||||
+18
-3
@@ -1,8 +1,23 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v4.2
|
||||
- **Revision date:** 2026-09-14
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1 and WP-A2 are committed (`d63804f`) and WP-B.1 is committed (`beb17ad`). **WP-B.2 is complete and staged for the owner's signed commit, accepted with a documented real-model limitation: B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship; deterministic independent recovery passes; the reference model failed at memory creation on the precondition-valid attempt, and the B2.4 prompt experiment was rejected and reverted** (`reports/v1.1/V1.1-WP-B2-REPORT.md` §S, §T). WP-C, WP-D and WP-E have not started, and no v1.1 version or tag exists.
|
||||
- **Package:** Adventure Storyteller Planning Package v4.4
|
||||
- **Revision date:** 2026-09-15
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. **WP-C, browser release coverage, is complete and staged for owner review: 91 checks, 0 failed, 0 skipped, over trusted-LAN HTTPS, with real export downloads** (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E have not started, and no v1.1 version or tag exists.
|
||||
|
||||
## v4.4 — WP-C browser release coverage (2026-09-15)
|
||||
|
||||
Harness work, with one narrow product fix it found. No requirement, acceptance
|
||||
test, schema or bundle format changed. The evidence is in
|
||||
`reports/v1.1/V1.1-WP-C-REPORT.md`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `reports/v1.1/V1.1-WP-C-REPORT.md` | **New.** The six reader workflows driven in a real browser, the existing 38 checks, the download environment, harness defects J1-J7, and product defects K1 (open) and K2 (fixed) | work-package report |
|
||||
| `V1.1-PLAN.md` | Status: WP-B.2 committed; WP-C complete and staged | status |
|
||||
| `planning/README.md` | Current state | index |
|
||||
| `DEVELOPMENT.md` | The browser harness command and flags; Firefox download preferences and the `$HOME` rule; what counts as a finished download; waiting on conditions, never sleeping | developer docs |
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15)
|
||||
|
||||
|
||||
@@ -0,0 +1,475 @@
|
||||
# v1.1 WP-C — Browser Release Coverage
|
||||
|
||||
**Status:** COMPLETE, staged for owner review. The final run passed 91 checks with 0 failed and 0 skipped. The decision is in §R.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `v1.1-development` |
|
||||
| HEAD at start | `0c1ba836babe1447ad3693b4d95189a325b7b3b6` — *v1.1 WP-B.2: independent long-term memory retention*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
|
||||
| Its ancestry | `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `ac465ed` (plan v4.1), `432f041` (v1.0.0) |
|
||||
| Working tree at start | clean; nothing staged |
|
||||
| WP-D, WP-E | not started |
|
||||
|
||||
---
|
||||
|
||||
## B. Existing browser harness
|
||||
|
||||
`backend/tools/m11_browser.py` drives Firefox through `backend/tools/m11_webdriver.py`, a
|
||||
dependency-free W3C WebDriver client. It builds its fixture campaign through the API,
|
||||
plays two narrator turns through the streaming endpoint, and then asserts on the
|
||||
rendered DOM. The M11 closeout run (`3652dc6`, run 3) passed **38/38**, 0 failed,
|
||||
0 skipped, on snap Firefox 155.0.1 with geckodriver 0.37.1 and `qwen2.5:3b-instruct`
|
||||
over trusted-LAN HTTPS.
|
||||
|
||||
What it did not do, and WP-C closes: drive Retry and takes, Save Point create and
|
||||
restore, state correction, narration length or failed generation through the UI,
|
||||
and prove an export leaves the browser as a file.
|
||||
|
||||
Before WP-C it also slept, for a fixed time, before several assertions: after Undo
|
||||
and Redo, after planting hostile narration, after choosing a knowledge file, around
|
||||
the delete dialog, and around opening a panel. §J records that as a harness defect.
|
||||
|
||||
---
|
||||
|
||||
## C. Browser and download environment
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Firefox | **155.0.1, the snap** (`/snap/bin/firefox`). No second Firefox was installed |
|
||||
| geckodriver | 0.37.1 (the snap) |
|
||||
| Headless | yes |
|
||||
| Frontend | the production build (`vite build`) served by FastAPI on loopback; no Vite dev server |
|
||||
| Certificates | `acceptInsecureCerts: false`, unchanged |
|
||||
|
||||
**The download profile.** `m11_webdriver.firefox_download_prefs` is passed as
|
||||
`moz:firefoxOptions.prefs`:
|
||||
- `browser.download.folderList` 2, `browser.download.dir` `<--out>/downloads`,
|
||||
`browser.download.useDownloadDir` true;
|
||||
- `browser.download.start_downloads_in_tmp_dir` false;
|
||||
- `browser.download.always_ask_before_handling_new_types` false;
|
||||
- `browser.helperApps.neverAsk.saveToDisk` `application/json,application/octet-stream`;
|
||||
- the download panel suppressed.
|
||||
|
||||
**The folder.** `--out/downloads`, which must be under `$HOME`
|
||||
(`require_under_home`). The harness deletes it at the start of a run and creates it
|
||||
fresh. Evidence lives under `$HOME/v11-evidence/wp-c/`.
|
||||
|
||||
**Does the snap Firefox download?** Yes. Measured first, with a probe
|
||||
(`$HOME/v11-evidence/wp-c/probe/`): a loopback page runs the product's own download
|
||||
pattern (a JSON blob, an `<a download>` click, an immediate `revokeObjectURL`).
|
||||
- Download folder under `~/v11-evidence`: 33-byte file written.
|
||||
- Download folder under `~/Downloads`: 33-byte file written.
|
||||
|
||||
The first probe wrote nothing, and that was a probe defect (J1), not the snap. The
|
||||
non-snap fallback the plan allows was therefore not needed.
|
||||
|
||||
**When a download counts as finished** (`m11_webdriver.wait_for_download`, tested
|
||||
without a browser in `test_v11_c_browser_helpers.py`, 7 tests). All of these at once:
|
||||
- a name absent from the listing taken before the click;
|
||||
- no `*.part` file in the folder;
|
||||
- more than zero bytes;
|
||||
- the same size across 3 consecutive polls.
|
||||
|
||||
A zero-byte, partial, pre-existing or still-growing file never counts, and neither
|
||||
does the "Campaign exported." toast.
|
||||
|
||||
---
|
||||
|
||||
## D. Retry scenario
|
||||
|
||||
Real narration: **yes**. All checks use the reader-facing controls on the play page.
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| Type in "What you do next", press **Send** | a new narration renders and the page is idle; its exact text is recorded | PASS |
|
||||
| — | **Retry** is offered (enabled) on the newest narration | PASS |
|
||||
| Press **Retry** | the take indicator on the newest narration reads **2/2** | PASS |
|
||||
| — | the second take's text differs from the first (otherwise 1/2 could not be told from 2/2) | PASS |
|
||||
| Press **‹** (Previous take) | the indicator reads **1/2**, and the narration is identical to the first recorded text | PASS |
|
||||
| — | the second take's text is not shown anywhere in the transcript | PASS |
|
||||
| Press **›** (Next take) | **2/2**, showing the second take | PASS |
|
||||
| Reload the page | the indicator still reads **2/2** on the live take, showing the second take | PASS |
|
||||
| Press **‹** after the reload | **1/2** still shows the first text, unchanged | PASS |
|
||||
|
||||
All text comparisons are of the rendered `.turn-text`. No database was read.
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## E. Save Point scenario
|
||||
|
||||
Real narration: **yes** (two turns after the Save Point).
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| Press **Save Point**, type a name in "Save this moment", submit | a row with that exact name appears | PASS |
|
||||
| — | the row's moment is the moment being read ("Moment 7") | PASS |
|
||||
| **Send** two turns | two new narrations render; position "Moment 11" | PASS |
|
||||
| In Save Points, press **Restore**, then confirm **Restore** | the position reads "Moment 7 · later story ahead" | PASS |
|
||||
| — | the transcript ends at the Save Point: its last narration is the one read there, and neither later narration is shown | PASS |
|
||||
| — | the position says later story is ahead | PASS |
|
||||
| — | **Redo** is enabled | PASS |
|
||||
| Press **Redo** until it is disabled (each press waits for the position to change) | the last two narrations are the two later turns, text-identical, at "Moment 11" | PASS |
|
||||
| Reload the page | the position is still "Moment 11" | PASS |
|
||||
| — | the named Save Point is still listed | PASS |
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## F. State-correction scenario
|
||||
|
||||
**Owner decision (2026-09-15).** The State panel cannot produce a *partly* refused
|
||||
correction:
|
||||
- "Save correction" sends exactly one `add_fact`, and "That's wrong" sends exactly
|
||||
one `invalidate_fact`;
|
||||
- so a correction is applied whole or refused whole (HTTP 400);
|
||||
- the route's partial application (`refused` alongside applied changes) is
|
||||
reachable only through the API.
|
||||
|
||||
WP-C drives what the reader can reach: an accepted correction that persists, and a
|
||||
refused correction whose refusal and reason are visible, with the refused change
|
||||
not applied. It records the partial refusal as unreachable from the reader UI. No
|
||||
product change was made for it.
|
||||
|
||||
Real narration: **yes** for the refusal (it needs history to step back over); the
|
||||
accepted correction needs none and also passed in the no-narrator smoke runs.
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| In **State**, press **Correct something**, type a fact, press **Save correction** | the form closes and the State panel shows the fact | PASS |
|
||||
| Reload, open **State** | the fact is still shown | PASS |
|
||||
| In a second tab press **Undo** (the correction belongs to the moment it was made at); in the first tab, whose panel still shows the fact, press **That's wrong** on it | a failure notice carrying the correction's refusal ("can't be applied") appears | PASS |
|
||||
| — | it is labelled **"That correction was not applied"**, with no "Try that turn again" and no claim that typed text was kept (§K2) | PASS |
|
||||
| Open **Show technical details** | the reason is visible: "That correction can't be applied — no fact 'f3' to invalidate." | PASS |
|
||||
| Second tab **Redo**, close it; reload the first tab at the corrected moment | the fact still stands, with its **That's wrong** control: the refused withdrawal was not applied | PASS |
|
||||
|
||||
**Partial refusal.** As the owner decided, a *partly* refused correction is not
|
||||
reachable from the reader UI, and was not driven. The route's partial application
|
||||
remains covered by the backend suite.
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## G. Narration-length scenario
|
||||
|
||||
Real narration: **yes** (one turn per band). The model's actual length is not
|
||||
asserted, because nothing in the product contract requires it. What is asserted is
|
||||
that the chosen band reached the turn's own prompt.
|
||||
|
||||
The expected sentence comes from the product's own `builder.length_hint` at the
|
||||
run's 400-token output cap:
|
||||
- **brief:** "must not exceed 180 words, and it should not stop short of about 70";
|
||||
- **long:** "must not exceed 236 words, and it should not stop short of about 118".
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| **Settings** panel: choose **Brief** in "Narration length", press **Save changes** | the button reads "Saved" | PASS |
|
||||
| **Send** a turn; choose **Long**, **Save changes**; **Send** a second turn | both turns render | PASS (both) |
|
||||
| On the brief turn press **Inspect context** | the prompt sections of "The exact text the narrator was sent" contain brief's range, and do **not** contain that turn's own reply (so this is the turn's record, not the dry run of the next) | PASS |
|
||||
| On the long turn press **Inspect context** | the same, with long's range | PASS |
|
||||
| Reload, open **Settings** | "Narration length" still reads **long** | PASS |
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## H. Failed-generation scenario
|
||||
|
||||
**Owner decision (2026-09-15).** The literal sequence (save an unserved model, then
|
||||
submit a turn) cannot be driven. The model check (`modelStatus.jsx`) marks a
|
||||
configured model absent from the endpoint's list as `missing-model`, and
|
||||
`blocksPlay` disables Send, Continue and Retry up front. That is M8's intended
|
||||
behaviour, not a defect. So WP-C drives both of these:
|
||||
1. **The unserved model** saved through Settings: the reader is told and cannot
|
||||
send, and the story is unchanged.
|
||||
2. **A submitted failure:** a model the endpoint lists but that cannot narrate,
|
||||
`nomic-embed-text:latest`. Send stays enabled, and the turn fails in the open.
|
||||
|
||||
Then recovery with the reference model. The report records that path 2 uses a
|
||||
listed, non-narrating model rather than "a name the server does not serve".
|
||||
|
||||
Real narration: **yes**.
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| **Settings**: "Type a model name instead", type an unserved name, **Save** | the page says "Saved" | PASS |
|
||||
| Open the campaign | the header model status is `missing-model` and the setup notice is shown | PASS |
|
||||
| — | **Send** and **Continue** are disabled | PASS |
|
||||
| — | the story is unchanged | PASS |
|
||||
| **Settings**: choose `nomic-embed-text:latest` from the installed-model picker, **Save** | "Saved" | PASS |
|
||||
| Open the campaign (model status `ready`); type a turn; **Send** | a failure notice is shown: "Generation failed", with the server's reason under the details, `"nomic-embed-text:latest" does not support chat` (HTTP 400) | PASS |
|
||||
| — | no narration was added | PASS |
|
||||
| — | the typed text is still in the input box | PASS |
|
||||
| — | the earlier story is text-identical | PASS |
|
||||
| Open **State** | the rendered state is identical to before the failure | PASS |
|
||||
| **Settings**: choose `qwen2.5:3b-instruct`, **Save** | "Saved" | PASS |
|
||||
| Open the campaign; type a turn; **Send** | exactly one new narration | PASS |
|
||||
| — | the earlier story is intact | PASS |
|
||||
| Reload | the successful turn is still the last narration | PASS |
|
||||
|
||||
The failed turn's player moment stays in the transcript, as A05 intends
|
||||
(`player_moment_kept_in_transcript` in the report).
|
||||
|
||||
**Deviation, owner-approved.** Path 2 uses a model the endpoint *lists* but that
|
||||
cannot narrate, not "a name the server does not serve", because an unserved name
|
||||
is caught before a turn can be submitted. Endpoint policy was not bypassed: both
|
||||
paths use the same trusted-LAN HTTPS endpoint, and only the model name changed.
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## I. Export-download scenario
|
||||
|
||||
Real narration: not needed for the download itself. In the final run the exported
|
||||
campaign holds 19 moments of real narration, two takes and a Save Point.
|
||||
|
||||
**C6a — the campaign library**
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| On the library page, press **Export** on the "Release Regression" card | a new file is written to `downloads/`; it is finished (no `.part`, stable size) | PASS |
|
||||
| — | it is not empty: **136,739 bytes** | PASS |
|
||||
| Parse the file | `format` is **`ai-dnd-adventure-v3`** | PASS |
|
||||
| Start a **fresh application** (new database); press **Import campaign**; give its file input the downloaded path | the browser lands on the imported campaign's play page | PASS |
|
||||
| Compare what the reader sees of the import with what the reader saw of the original (library card, position, Save Points panel) | same title ("Release Regression") | PASS |
|
||||
| — | same number of moments (19) | PASS |
|
||||
| — | same position ("Moment 18") | PASS |
|
||||
| — | same Save Points, which also match the file | PASS |
|
||||
|
||||
The file's `headDepth` is 17, which the play page shows as "Moment 18", and its
|
||||
`actions` count is 19.
|
||||
|
||||
**C6b — campaign settings**
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| In the campaign's **Settings** panel, press **Export campaign** | a second, separate file is written and finished | PASS |
|
||||
| — | not empty: 136,739 bytes | PASS |
|
||||
| Parse the file | `format` is `ai-dnd-adventure-v3` | PASS |
|
||||
|
||||
Both files: `$HOME/v11-evidence/wp-c/final/``downloads/library-export.json` and `downloads/settings-export.json`.
|
||||
The imported application's log is `import-server.log`.
|
||||
|
||||
---
|
||||
|
||||
## J. Harness defects found
|
||||
|
||||
| # | Defect | How found | Fix |
|
||||
| --- | --- | --- | --- |
|
||||
| J1 | The download probe's page wrote `URL.createObjectURL` inside an inline `onclick`, where `URL` is `document.URL`, a string. No blob was made, so it looked exactly like "the snap cannot download" | `gecko.log`: `TypeError: URL.createObjectURL is not a function` | The probe's script uses `window.URL`, the way the product's module does. Both download folders then worked |
|
||||
| J2 | C3's first refusal withdrew a fact in a second tab and withdrew it again in the stale tab. Withdrawing keeps the fact, marked `invalidated` (C04's audit record), so the second withdrawal was valid and accepted. Nothing was refused, and "the refused change was not applied" passed without meaning anything | Smoke run: two C3 failures; `server.log` shows 201 for every correction; `narrative/apply.py` | The second tab steps the story back past the correction with Undo, so the stale tab withdraws a fact the story at that position does not have. The validator then refuses it: `no fact … to invalidate` |
|
||||
| J3 | Clicks on a control just under the play page's fixed composer were intercepted ("Show technical details" in a failure notice, a turn's "Inspect context") | Dev run 1: `element click intercepted` | `Browser.click` scrolls the element to the centre of the view, then uses the real WebDriver click |
|
||||
| J4 | The Settings model field is a text box until the endpoint's model list arrives, then a picker. Choosing before the check finished raced that swap | Dev run 1: `no such element: input#model` | Wait for the header's model status to leave `checking` first |
|
||||
| J5 | C3's "a refused correction is shown" waited for *any* failure notice, so a notice about something else would have passed | Dev run 1: it passed on a notice titled "Generation failed" (§K2) | It now requires the notice to carry this correction's refusal ("can't be applied"), and asserts how it is labelled |
|
||||
| J7 | C4 required the band's sentence in the inspector and the turn's own reply to be absent, to tell a turn's record from the dry run. But the inspector also renders what came back (`raw_output`, "What came back, before the state block was removed") inside the same section, so the absence could never hold | Dev run 2: both C4 inspector checks failed. The stored records show the brief turn's `length_hint` section carrying "must not exceed 180 words … about 70", the long turn's carrying "…236 … 118", and every turn's reply present in `raw_output` | The sentence and the absence are read from the prompt sections only, excluding the "What came back" block |
|
||||
| J6 | M11's checks slept before assertions: after Undo and Redo, after planting hostile narration (1 s), after choosing a knowledge file (0.5 s), around the delete dialog (0.8 s and 0.6 s), and around opening a panel (0.8 s). A sleep is not evidence of what it waited for | Reading the harness against the brief's rule | Each is now a wait on the condition the check needs: the position changed or returned, the planted text rendered, Import enabled, a dialog present or gone, a panel-specific element present. §38's absence check now first waits for the knowledge library to render |
|
||||
|
||||
**Checked and not a defect.** In dev run 2, the original take at depth 6 (action 9)
|
||||
had no `length_hint` in its stored record, while its Retry (action 10) did. The
|
||||
original take's record holds only the per-attempt fields
|
||||
(`attempts.ATTEMPT_KEYS`: world state, narrative state, raw output, usage,
|
||||
accounting), with no sections and no settings. That is how a take that is no
|
||||
longer live is stored, not a prompt built without the range. Every live turn's
|
||||
prompt carried its band.
|
||||
|
||||
Also added: a panel opens only if it is not already open, because a tab toggles
|
||||
its panel closed, and it is recognised by an element only that panel renders, not
|
||||
by its title, which the tab itself already shows.
|
||||
|
||||
---
|
||||
|
||||
## K. Product defects found
|
||||
|
||||
### K1 — "Correct" on an Important Facts row is always refused (not fixed)
|
||||
|
||||
The State panel offers **Correct** on every row with a key. On an Important Facts
|
||||
row the key is the fact's id, and on the scene-summary row it is `"summary"`.
|
||||
`saveCorrection` sends that key as `add_fact.subject`, and the validator checks
|
||||
`subject` as an entity reference. So every such correction is refused.
|
||||
|
||||
**Reproduced deterministically** against the real application (scratch `TestClient`,
|
||||
no browser, no model):
|
||||
|
||||
| Correction the panel sends | Result |
|
||||
| --- | --- |
|
||||
| "Correct" on the Characters row (`subject='mara'`) | **201**, applied |
|
||||
| "Correct" on the Important Facts row (`subject='f1'`) | **400** "That correction can't be applied — add_fact names subject='f1', which does not exist." |
|
||||
|
||||
**Not fixed in WP-C.** It does not prevent any WP-C behaviour: "Correct something"
|
||||
and entity-row corrections work, and C3 uses them. A fix (offer Correct only
|
||||
against entities, or send facts without a subject) is a small UX choice, left to
|
||||
the owner as a v1.1 follow-up.
|
||||
|
||||
### K2 — a refused correction was presented as a failed turn (fixed)
|
||||
|
||||
**Found in the browser** (dev run 1, §J5). When the story refused a correction, the
|
||||
failure notice said:
|
||||
- the title **"Generation failed"**;
|
||||
- the hint "Nothing was added to your story. You can try that turn again.";
|
||||
- the button **Try that turn again**;
|
||||
- the line "what you typed is still in the box below".
|
||||
|
||||
None of that is true of a State-panel correction. `classifyError` had no rule for
|
||||
the server's refusal ("That correction can't be applied — …"), so it fell through
|
||||
to the generation default. That falsified exactly what C3 checks: that a refusal
|
||||
is shown to the reader as a refusal.
|
||||
|
||||
**Fix** (frontend only, no backend change):
|
||||
- `errors.js`: one rule, checked first, for "correction can't be applied". It gives
|
||||
`kind: state`, the title "That correction was not applied", the hint "Nothing in
|
||||
the story or its state was changed. The reason is in the technical details.",
|
||||
`retryable: false` and `keptInput: false`.
|
||||
- `FailureNotice.jsx`: the "what you typed" line is shown unless a failure says
|
||||
`keptInput: false`. Every other kind still shows it, so M8's A05 contract is
|
||||
unchanged.
|
||||
|
||||
**Regression** (`failurePaths.test.jsx`, 3 tests):
|
||||
- a refused correction classifies as a state refusal, not retryable, with the
|
||||
reason kept;
|
||||
- its notice shows the reason and no "Try that turn again" or typed-input claim;
|
||||
- a failed turn still claims the typed words were kept.
|
||||
|
||||
In the browser, C3 now asserts the label, the absence of "Try that turn again" and
|
||||
the absence of the typed-input line.
|
||||
|
||||
Frontend after the fix: **168/168** tests, lint exit 0 (15 pre-existing warnings, 0
|
||||
errors, none in changed files), production build passes.
|
||||
|
||||
---
|
||||
|
||||
## L. Existing 38-check regression
|
||||
|
||||
**38/38 passed, 0 failed, 0 skipped** in the same run, tagged `M11`. The names are
|
||||
identical to the M11 closeout run, so there is a one-to-one mapping and no check was
|
||||
split, merged or dropped:
|
||||
- B01 a turn is accepted (×2);
|
||||
- A/UX: the tab title (×3);
|
||||
- B position indicator (×3), D01 Undo, D04 Redo (×2);
|
||||
- H06 (×3), H07, G09, H04;
|
||||
- G01 knowledge import (×2);
|
||||
- A11y dialog focus (×4);
|
||||
- §38 narrator-only text absent;
|
||||
- F05 context inspector;
|
||||
- H10 (×2), H11/CSP (×2);
|
||||
- A11y names, focus, tabindex, hover, contrast (×4), input focus.
|
||||
|
||||
What changed in them is how they wait (J6), not what they assert.
|
||||
|
||||
## M. New WP-C checks
|
||||
|
||||
**53/53 passed, 0 failed, 0 skipped**, tagged `WP-C`:
|
||||
|
||||
| Scenario | Checks | Result |
|
||||
| --- | --- | --- |
|
||||
| C1 Retry | 8 | 8 PASS |
|
||||
| C2 Save Point | 9 | 9 PASS |
|
||||
| C3 State correction | 6 | 6 PASS |
|
||||
| C4 Narration length | 5 | 5 PASS |
|
||||
| C5 Failed generation | 14 | 14 PASS |
|
||||
| C6a Library export and import | 8 | 8 PASS |
|
||||
| C6b Settings export | 3 | 3 PASS |
|
||||
|
||||
```text
|
||||
existing M11 checks: 38/38
|
||||
WP-C new checks: 53/53
|
||||
failed: 0
|
||||
skipped: 0
|
||||
```
|
||||
|
||||
**Runs that are not the final evidence**, kept under `$HOME/v11-evidence/wp-c/`:
|
||||
|
||||
| Run | Where | Result | Why it is not evidence |
|
||||
| --- | --- | --- | --- |
|
||||
| `probe/` | loopback page | first: no download (J1); second: both folders written | environment probe |
|
||||
| `smoke-1` | no narrator | 44 passed, 2 failed (J2), 6 skipped | partial |
|
||||
| `smoke-2` | no narrator | 43 passed, 0 failed, 7 skipped | partial |
|
||||
| `dev-gpu-1` | GPU host, plain HTTP | 71 passed, 3 failed (J3, J4; J5 found) | development, and not HTTPS |
|
||||
| `dev-gpu-2` | GPU host, plain HTTP | 89 passed, 2 failed (J7) | development, and not HTTPS |
|
||||
| `dev-gpu-3` | GPU host, `--only length` | 7 passed | development, and a subset |
|
||||
|
||||
## N. Production build / frontend verification
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| Frontend suite (`npm test`) | **168/168**, 14 files. It was 165 before; the 3 new tests are K2's regressions |
|
||||
| Lint (`npm run lint`, oxlint) | **exit 0: 0 errors**, 15 warnings. All are the pre-existing `only-export-components` kind, and none is in a file WP-C changed |
|
||||
| Production build (`npm run build`) | passes; `dist/index.html` sha256 `62b6ea5eb4ce02a09f23cc1d48c335c2ada36b208b23c237338b39bf63a26cc5`, built 2026-09-15 15:11 from the final WP-C tree |
|
||||
| What the browser ran | that build, served by FastAPI (`uvicorn app.main:app` on `127.0.0.1`); no Vite server |
|
||||
| Backend product code | **unchanged**: nothing under `backend/app` is in the diff. The full backend suite was not rerun. The changed harness is tested by `test_v11_c_browser_helpers.py` (7 passed) |
|
||||
| Offline regression | **23 passed, 0 failed** (`tools/m11_offline.py`, fresh `--no-cache` image, `--network none`, on the final WP-C tree), rerun because the frontend bundle changed (K2). Evidence: `$HOME/v11-evidence/wp-c/offline/` |
|
||||
|
||||
## O. Trusted-LAN / security
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Narrator | `qwen2.5:3b-instruct` on the CPU reference host, over **trusted-LAN HTTPS**. The certificate is from the private CA in this machine's trust store, and verified; there is no bypass |
|
||||
| Window and A1 | every narrator turn (8): window **verified at 4,096**, accounting **`fits`**. None `exceeded` or `truncation_suspected` |
|
||||
| Protocol echoes (A2) | none of the protocol shapes the harness looks for (a state fence, a hard-limit or reminder bracket, `Events: [`) appeared in any stored narration in this run. This is not a v1.1 protocol-leak result: the release gate owns that, and the mid-reply echo from WP-B.1 remains a separate residual |
|
||||
| Storyteller | loopback only, both applications (the original and the fresh import) |
|
||||
| Browser | `acceptInsecureCerts: false`; the CSP checks (H11) passed |
|
||||
| Endpoint policy | unchanged; C5 changed only the model name, never the endpoint |
|
||||
| Downloads | written only under `$HOME` (enforced); the harness makes no network request of its own beyond loopback and the configured endpoint |
|
||||
| New dependencies | none: the harness still uses only `urllib`, no Selenium or Playwright |
|
||||
| Real identifiers in committed files | none (scanned at staging) |
|
||||
|
||||
## P. Compatibility
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Database schema, migrations | none |
|
||||
| Bundle format | none: the downloads are `ai-dnd-adventure-v3` and import unchanged |
|
||||
| History, Save Points, state semantics, memory, knowledge | none. WP-C drove them and changed nothing in them |
|
||||
| Backend | no application code changed |
|
||||
| Frontend | one behaviour change (K2): the refusal of a State-panel correction is labelled as a refusal rather than as a failed turn. Every other failure's classification, retry offer and typed-input claim is unchanged (M8's tests pass) |
|
||||
| WP-A, WP-B | untouched |
|
||||
|
||||
## Q. Residual risks
|
||||
|
||||
| # | Risk |
|
||||
| --- | --- |
|
||||
| 1 | **K1**: "Correct" on an Important Facts or scene-summary row is always refused. Reproduced, not fixed: a small UX choice for the owner |
|
||||
| 2 | **Partial refusal is API-only.** The reader UI cannot produce a partly refused correction (owner decision). The display for one exists but is unreachable from the panel's own controls |
|
||||
| 3 | **C5 path 2 depends on the reference host listing an embedding model.** On a host without one, the submitted-failure path has no model to use |
|
||||
| 4 | **Model nondeterminism in C1.** "The second take is a different narration" would fail if the model returned identical text for a retry. It did not in any run |
|
||||
| 5 | **The harness runs on this machine's snap Firefox.** A different Firefox or a Chromium would need the download preferences re-checked |
|
||||
| 6 | **The heuristic protocol-shape scan is not A2 evidence.** It is recorded only |
|
||||
| 7 | The WP-B real-model memory limitation and the doubled-full-stop scene text are unchanged, and are not WP-C's |
|
||||
|
||||
## R. Final decision
|
||||
|
||||
```text
|
||||
RETRY: PASS
|
||||
SAVE POINT: PASS
|
||||
STATE CORRECTION: PASS
|
||||
NARRATION LENGTH: PASS
|
||||
FAILED GENERATION: PASS
|
||||
EXPORT DOWNLOAD — LIBRARY: PASS
|
||||
EXPORT DOWNLOAD — SETTINGS: PASS
|
||||
|
||||
EXISTING BROWSER REGRESSION: PASS
|
||||
WP-C NEW BROWSER COVERAGE: PASS
|
||||
|
||||
WP-C OVERALL:
|
||||
PASS
|
||||
```
|
||||
|
||||
**PASS**, on these grounds:
|
||||
- the final run ended with **failed: 0, skipped: 0**;
|
||||
- both export controls produced a real, finished, non-empty `ai-dnd-adventure-v3`
|
||||
file on disk;
|
||||
- the library file imported into a fresh application with the same title, moment
|
||||
count, position and Save Points.
|
||||
|
||||
Two criteria were met in the form the owner approved (2026-09-15):
|
||||
- **State correction:** an accepted correction, and a refused correction with its
|
||||
reason. Partial refusal is recorded as unreachable from the reader UI.
|
||||
- **Failed generation:** the up-front block for an unserved model, plus a submitted
|
||||
failure with a listed model that cannot narrate.
|
||||
|
||||
**Product change:** K2, a mislabelled refusal, fixed narrowly with regression tests.
|
||||
**Product defect left open:** K1.
|
||||
|
||||
Nothing is committed, pushed or tagged. WP-D and WP-E have not started.
|
||||
Reference in New Issue
Block a user