From 87a40326a29533c8d52c9f9f41022e7b499b1de7 Mon Sep 17 00:00:00 2001 From: JesseMarkowitz Date: Wed, 16 Sep 2026 05:37:13 -0400 Subject: [PATCH] v1.1: harden recovery and control boundaries MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit WP-D and WP-E complete the planned v1.1 implementation packages. WP-D — recovery honesty: - backups verify the completed copy with PRAGMA integrity_check - corruption missed by quick_check is detected by the full check - existing good backups remain protected - oversized exports are still delivered but declare whether this version can import them, while the 20 MB import limit remains unchanged - backup was exercised through the real browser UI on both the normal campaign database and a campaign-shaped database over 100 MB WP-E — control-boundary contrast: - interactive control boundaries meet the WCAG 1.4.11 3:1 target - the contrast audit is now a failing gate rather than an advisory - rendered browser measurements pass for the composer, controls, tabs and nav - text contrast and focus visibility remain intact - owner reviewed and approved the before/after screenshots Reports: - planning/reports/v1.1/V1.1-WP-D-REPORT.md - planning/reports/v1.1/V1.1-WP-E-REPORT.md All planned v1.1 work packages A-E are now complete. Release validation has not yet begun. --- DEVELOPMENT.md | 34 ++ backend/app/backup.py | 29 +- backend/app/limits.py | 28 + backend/app/routers/adventures/bundle_io.py | 36 +- backend/tests/test_v11_d_recovery.py | 294 ++++++++++ backend/tests/test_v11_e_contrast.py | 152 +++++ backend/tools/contrast_audit.py | 83 +-- backend/tools/m11_browser.py | 231 +++++++- backend/tools/m11_webdriver.py | 40 ++ backend/tools/wpd_backup_ui.py | 353 ++++++++++++ frontend/src/api.js | 27 +- frontend/src/pages/Campaigns.jsx | 6 +- .../Play/panels/CampaignSettingsPanel.jsx | 6 +- frontend/src/pages/exportHonesty.test.jsx | 167 ++++++ frontend/src/styles/context.css | 17 +- frontend/src/styles/tokens.css | 11 +- planning/README.md | 10 +- planning/V1.1-PLAN.md | 12 +- planning/VERSION.md | 25 +- planning/reports/v1.1/V1.1-WP-D-REPORT.md | 526 ++++++++++++++++++ planning/reports/v1.1/V1.1-WP-E-REPORT.md | 374 +++++++++++++ 21 files changed, 2401 insertions(+), 60 deletions(-) create mode 100644 backend/tests/test_v11_d_recovery.py create mode 100644 backend/tests/test_v11_e_contrast.py create mode 100644 backend/tools/wpd_backup_ui.py create mode 100644 frontend/src/pages/exportHonesty.test.jsx create mode 100644 planning/reports/v1.1/V1.1-WP-D-REPORT.md create mode 100644 planning/reports/v1.1/V1.1-WP-E-REPORT.md diff --git a/DEVELOPMENT.md b/DEVELOPMENT.md index 8c71ca7..50a1a5d 100644 --- a/DEVELOPMENT.md +++ b/DEVELOPMENT.md @@ -548,6 +548,40 @@ together. | Taken from | Export, on a campaign | Settings → *Back up everything on this machine* | | Restored by | Import campaign, on the library screen | replacing the database file, below | +### How large an export can get + +The importer accepts a request body up to **20 MB** +(`backend/app/limits.py`, `MAX_IMPORT_BODY_BYTES`), and v1.1 does not change it. +What that means for a campaign, measured rather than guessed: + +- the M11 evidence campaign came to roughly **13 kB per action** in its bundle; +- M9's conservative estimate from that figure is about **279 turns** before a + bundle approaches the limit. + +Both are measurements of particular campaigns, **not a turn limit**. What a +campaign actually weighs depends on how long its turns are, how much imported +knowledge travels with it, and how many attempts each turn kept. A campaign of +400 short turns can be well inside the limit; one of 200 long ones with a large +library may not be. + +**v1.1 (WP-D) makes the individual case visible.** Every export reports its own +serialised size and whether this version could import it back: + +```text +X-Export-Bytes the bundle's size, as the importer would weigh it +X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES +X-Importable-By-This-Version true / false +X-Export-Warning present only when it is false +``` + +The export always succeeds and the file is always delivered — it is complete and +undamaged; what it exceeds is this version's import ceiling. Both Export +controls show the warning when there is one. The size compared is the compact +serialisation the browser would POST back, which is smaller than the +pretty-printed file on disk. + +Raising the limit, or streaming an import past it, is deferred to v1.2. + ### Exporting and importing a campaign Export is on each campaign in the library, and in the campaign's own Settings diff --git a/backend/app/backup.py b/backend/app/backup.py index 13edb3f..c01502a 100644 --- a/backend/app/backup.py +++ b/backend/app/backup.py @@ -39,8 +39,9 @@ turn is blocked. never leaves a half-written file wearing a backup's name. `os.replace` is atomic on the same filesystem, which is why the temporary sits in the destination's own directory rather than in `/tmp`. -3. `PRAGMA quick_check` runs against the finished copy, opened as its own - database, before it is renamed. A backup nobody verified is a belief. +3. `PRAGMA integrity_check` runs against the finished copy, opened as its own + database, before it is renamed. A backup nobody verified is a belief. v1.1 + WP-D made this the full check rather than `quick_check`; see `_verify`. 4. An existing file is never overwritten. Each run writes a new name stamped with the time, so yesterday's backup survives today's mistake — which is most of what a backup is for. @@ -189,18 +190,28 @@ def _copy(source_path: Path, working: Path) -> int: def _verify(working: Path) -> str: - """Runs `PRAGMA quick_check` against the finished copy. + """Runs `PRAGMA integrity_check` against the finished copy. Opened as its own connection, so what is checked is the file on disk rather - than any page cache the copy left behind. `quick_check` rather than - `integrity_check` because it does the structural work — every page reachable, - every record readable — without the full index cross-check, which on a large - database is minutes rather than moments. A backup nobody verified is a - belief; a backup verified slowly enough that nobody takes one is worse. + than any page cache the copy left behind. + + **v1.1 WP-D: the full check, not `quick_check`.** M9 chose `quick_check` for + its speed, on the argument that a backup verified slowly enough that nobody + takes one is worse than a fast one. The measurements say the trade was not + needed here: `quick_check` omits the cross-check between a table and its + indexes, and that is a real class of damage it reports as `ok`. A copy whose + index disagrees with its table restores into a database that answers queries + with rows that are not there — the failure a backup exists to prevent. + + The cost is small at the sizes this application produces: on the 100-turn + evidence campaign both checks are a few milliseconds, and on a synthetic + database two orders of magnitude larger the difference is still short of a + second (WP-D report §E). A backup nobody verified is a belief; this is the + check that makes it a fact. """ connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True) try: - rows = connection.execute("PRAGMA quick_check").fetchall() + rows = connection.execute("PRAGMA integrity_check").fetchall() finally: connection.close() result = ", ".join(str(row[0]) for row in rows) if rows else "no result" diff --git a/backend/app/limits.py b/backend/app/limits.py index 16a0c24..bacd973 100644 --- a/backend/app/limits.py +++ b/backend/app/limits.py @@ -129,6 +129,34 @@ MAX_BODY_BYTES = 2 * 1024 * 1024 MAX_IMPORT_BODY_BYTES = 20 * 1024 * 1024 +def import_limit_label(limit: int | None = None) -> str: + """The import ceiling as a reader would say it, e.g. "20 MB". + + Derived from the constant rather than written beside it, so the refusal, the + export warning and the documentation cannot drift apart from each other or + from what the middleware actually enforces (v1.1 WP-D). + """ + size = MAX_IMPORT_BODY_BYTES if limit is None else limit + megabytes = size / (1024 * 1024) + return f"{megabytes:.0f} MB" if abs(megabytes - round(megabytes)) < 0.05 else f"{megabytes:.1f} MB" + + +def oversized_export_warning(export_bytes: int, limit: int | None = None) -> str: + """What to tell a reader whose export is larger than import will accept. + + v1.1 WP-D. The file is written and is not damaged: what it exceeds is this + version's import ceiling, so it cannot be brought back in *here*. Saying that + plainly is the whole point — the alternative is a reader who finds out when + they try to restore it. + """ + size = MAX_IMPORT_BODY_BYTES if limit is None else limit + return ( + f"This export is larger than this version's {import_limit_label(size)} import " + f"limit ({export_bytes:,} bytes). The file was exported successfully, but this " + f"version cannot import it." + ) + + class BodySizeLimitMiddleware: """Rejects oversized request bodies by their declared `Content-Length`. diff --git a/backend/app/routers/adventures/bundle_io.py b/backend/app/routers/adventures/bundle_io.py index 966b63d..4818532 100644 --- a/backend/app/routers/adventures/bundle_io.py +++ b/backend/app/routers/adventures/bundle_io.py @@ -28,7 +28,9 @@ repair. Refusing a whole campaign because a search index would not build would trade the valuable thing for the cheap one. """ -from fastapi import Body, Depends, Request +import json + +from fastapi import Body, Depends, Request, Response from sqlalchemy.orm import Session from ... import bundle, head, limits, models, schemas @@ -46,8 +48,38 @@ def export_adventure( `app/bundle.py` owns the format, in all three of its versions. A backup outlives the schema, so no call site decides anything about its shape. + + **v1.1 WP-D: the export also says whether this version could import it back.** + A campaign large enough to pass `limits.MAX_IMPORT_BODY_BYTES` still exports — + the file is complete and not damaged, and refusing to write it would destroy + the only copy the reader was trying to make. What it cannot do is come back + in here, and the reader is told that at the moment they take it rather than + at the moment they need it. + + It travels in headers, not in the body. The body is the bundle, the browser + saves exactly those bytes as the file, and a warning inside it would become + part of a portable story file and of every checksum taken over one. + + The size measured is the compact serialisation, because that is both what + this response sends and what the browser POSTs back on import, which is what + `BodySizeLimitMiddleware` weighs. The pretty-printed file the reader + downloads is larger, and is not what import reads. """ - return bundle.export(db, adv) + payload = bundle.export(db, adv) + # Serialised exactly as Starlette's JSONResponse would, so the bytes counted + # are the bytes sent. + body = json.dumps(payload, ensure_ascii=False, allow_nan=False, + separators=(",", ":")).encode("utf-8") + limit = limits.MAX_IMPORT_BODY_BYTES + importable = len(body) <= limit + headers = { + "X-Export-Bytes": str(len(body)), + "X-Import-Limit-Bytes": str(limit), + "X-Importable-By-This-Version": "true" if importable else "false", + } + if not importable: + headers["X-Export-Warning"] = limits.oversized_export_warning(len(body), limit) + return Response(content=body, media_type="application/json", headers=headers) @router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201) diff --git a/backend/tests/test_v11_d_recovery.py b/backend/tests/test_v11_d_recovery.py new file mode 100644 index 0000000..a3858b3 --- /dev/null +++ b/backend/tests/test_v11_d_recovery.py @@ -0,0 +1,294 @@ +"""v1.1 WP-D: a backup that was really checked, and an export that says what it is. + +Two recovery-path claims, each of which was true only in the small before this +package: + +- **A backup is verified.** M9 ran `PRAGMA quick_check` on the finished copy. + That reads every page and every record, and skips the cross-check between a + table and its indexes — so a copy whose index disagrees with its table passed. + `test_the_fixture_is_the_difference_between_the_two_checks` builds exactly that + damage and shows the two pragmas disagreeing about it, before anything here + uses it as evidence. +- **An export is importable.** Nothing compared the bundle with + `limits.MAX_IMPORT_BODY_BYTES`, so a campaign could be exported and then + refused by its own importer, with the reader finding out at the moment they + needed it. The export still succeeds — the file is complete, and a version + that refused to write it would destroy the copy someone was trying to make — + and now it says so. + + python -m pytest tests/test_v11_d_recovery.py -v +""" + +import json +import sqlite3 +from pathlib import Path + +import pytest +from fastapi import Depends +from fastapi.testclient import TestClient + +from app import auth, backup, limits, models +from app.database import Base, SessionLocal, engine, get_db +from app.main import app + + +@pytest.fixture() +def client(): + Base.metadata.create_all(bind=engine) + setup = SessionLocal() + user = models.User(is_guest=False, email="wp-d@example.com") + setup.add(user) + setup.flush() + setup.add(models.Settings(user_id=user.id, model="test-model")) + adventure = models.Adventure(user_id=user.id, title="Recovery") + setup.add(adventure) + setup.flush() + setup.add(models.Action(adventure_id=adventure.id, type="start", text="The story opens.")) + setup.commit() + adv_id, user_id = adventure.id, user.id + setup.close() + app.dependency_overrides[auth.get_current_user] = ( + lambda db=Depends(get_db): db.get(models.User, user_id)) + test_client = TestClient(app) + test_client.adv_id = adv_id + test_client.user_id = user_id + try: + yield test_client + finally: + app.dependency_overrides.clear() + Base.metadata.drop_all(bind=engine) + + +# ------------------------------------------------------------ the fixture + +def build_corrupt_copy(path: Path) -> None: + """A database whose index disagrees with its table, and nothing else. + + One digit inside one index leaf page is changed, so that entry names a key + no row holds and one row's key is in no index entry. Every page is still + structurally sound and every record still parses, which is the whole point: + this is the damage `quick_check` is not looking for. + """ + path.unlink(missing_ok=True) + connection = sqlite3.connect(path) + connection.execute("PRAGMA page_size=4096") + connection.execute("CREATE TABLE t (id INTEGER PRIMARY KEY, k TEXT NOT NULL, filler TEXT)") + connection.execute("CREATE INDEX i_t_k ON t(k)") + connection.executemany("INSERT INTO t (k, filler) VALUES (?, ?)", + [(f"k{n:06d}", "x" * 40) for n in range(400)]) + connection.commit() + page_size = connection.execute("PRAGMA page_size").fetchone()[0] + leaves = [row[0] for row in connection.execute( + "SELECT pageno FROM dbstat WHERE name='i_t_k' AND pagetype='leaf' ORDER BY pageno")] + connection.close() + assert leaves, "the index must have a leaf page to damage" + + raw = bytearray(path.read_bytes()) + start = (leaves[0] - 1) * page_size + page = raw[start:start + page_size] + at = page.find(b"k000") + assert at != -1, "expected an indexed key on the index's first leaf page" + page[at + 4] = ord("9") # k000144 -> k000944: a key no row has + raw[start:start + page_size] = page + path.write_bytes(bytes(raw)) + + +def check(path: Path, pragma: str) -> str: + connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True) + try: + return ", ".join(str(row[0]) for row in connection.execute(f"PRAGMA {pragma}").fetchall()) + finally: + connection.close() + + +def test_the_fixture_is_the_difference_between_the_two_checks(tmp_path): + """Before using it as evidence: quick_check calls this database fine.""" + damaged = tmp_path / "damaged.db" + build_corrupt_copy(damaged) + assert check(damaged, "quick_check") == "ok" + integrity = check(damaged, "integrity_check") + assert integrity != "ok" + assert "i_t_k" in integrity # it names the index that disagrees + + +# ---------------------------------------------------------------- backups + +def test_a_healthy_backup_passes_the_full_check_and_is_kept(client, tmp_path): + source = tmp_path / "campaign.db" + source.write_bytes(Path(str(engine.url.database)).read_bytes()) + result = backup.create(source) + assert result.integrity == "ok" + assert result.path.exists() and result.bytes > 0 + assert check(result.path, "integrity_check") == "ok" + assert result.path.parent == backup.directory(source) + + +def test_the_backup_runs_the_full_check_not_the_quick_one(client, tmp_path, monkeypatch): + """The pragma itself, named. SQLite traces every statement it executes, so + this reads what the backup actually asked the copy rather than inferring it.""" + asked: list[str] = [] + real_connect = sqlite3.connect + + def tracing(*args, **kwargs): + connection = real_connect(*args, **kwargs) + connection.set_trace_callback(asked.append) + return connection + + monkeypatch.setattr(backup.sqlite3, "connect", tracing) + source = tmp_path / "campaign.db" + source.write_bytes(Path(str(engine.url.database)).read_bytes()) + backup.create(source).path.unlink() + assert any("integrity_check" in sql for sql in asked), asked + assert not any("quick_check" in sql for sql in asked), asked + + +def test_a_copy_the_full_check_rejects_is_not_kept(client, tmp_path, monkeypatch): + """The copy is damaged after it is written and before it is verified, which + is where a real page-level fault would appear: between the copy and the + rename. Nothing wearing a backup's name may be left behind.""" + source = tmp_path / "campaign.db" + source.write_bytes(Path(str(engine.url.database)).read_bytes()) + real_copy = backup._copy + + def damage(source_path, working): + pages = real_copy(source_path, working) + build_corrupt_copy(working) + return pages + + monkeypatch.setattr(backup, "_copy", damage) + with pytest.raises(backup.BackupError) as refused: + backup.create(source) + assert "did not verify" in str(refused.value) + assert "i_t_k" in str(refused.value) # it says what was wrong + kept = list(backup.directory(source).glob("*")) + assert kept == [], f"a rejected backup was left behind: {kept}" + + +def test_a_rejected_backup_leaves_an_earlier_good_one_alone(client, tmp_path, monkeypatch): + source = tmp_path / "campaign.db" + source.write_bytes(Path(str(engine.url.database)).read_bytes()) + good = backup.create(source) + before = good.path.read_bytes() + + real_copy = backup._copy + + def damage(source_path, working): + pages = real_copy(source_path, working) + build_corrupt_copy(working) + return pages + + monkeypatch.setattr(backup, "_copy", damage) + with pytest.raises(backup.BackupError): + backup.create(source) + assert good.path.exists() + assert good.path.read_bytes() == before + assert check(good.path, "integrity_check") == "ok" + assert [p.name for p in backup.directory(source).glob("*")] == [good.path.name] + + +def test_the_backup_file_semantics_are_unchanged(client, tmp_path): + """Same directory, same stamped name, same reported fields: WP-D changed the + check, not the file.""" + source = tmp_path / "campaign.db" + source.write_bytes(Path(str(engine.url.database)).read_bytes()) + first = backup.create(source) + second = backup.create(source) + assert first.path.name.startswith(backup.PREFIX) and first.path.suffix == ".db" + assert first.path != second.path, "an existing backup is never overwritten" + assert set(first.as_dict()) == {"filename", "bytes", "pages", "seconds", "integrity"} + listed = [row["filename"] for row in backup.existing(source)] + assert sorted(listed) == sorted([first.path.name, second.path.name]) + + +# ----------------------------------------------------------------- exports + +def export(client, adv_id): + response = client.get(f"/api/adventures/{adv_id}/export") + assert response.status_code == 200, response.text[:200] + return response + + +def test_a_normal_export_carries_no_warning(client): + response = export(client, client.adv_id) + assert "X-Export-Warning" not in response.headers + assert response.headers["X-Importable-By-This-Version"] == "true" + assert int(response.headers["X-Import-Limit-Bytes"]) == limits.MAX_IMPORT_BODY_BYTES + assert int(response.headers["X-Export-Bytes"]) == len(response.content) + assert response.json()["format"] == "ai-dnd-adventure-v3" + + +def fill_past_the_limit(adv_id: int) -> int: + """Real rows, until the campaign's bundle is genuinely over the ceiling. + + Not a mocked size: the export below serialises all of it. + """ + chunk = "The rain kept on over the harbour road, and nobody came. " * 900 # ~50 kB + written = 0 + with SessionLocal() as db: + while written < limits.MAX_IMPORT_BODY_BYTES + 2 * 1024 * 1024: + db.add_all([models.Action(adventure_id=adv_id, type="ai", text=chunk) + for _ in range(40)]) + db.commit() + written += 40 * len(chunk) + return written + + +def test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back(client): + fill_past_the_limit(client.adv_id) + response = export(client, client.adv_id) + + # Delivered, whole, and still the same format. + body = response.content + assert len(body) > limits.MAX_IMPORT_BODY_BYTES + parsed = json.loads(body) + assert parsed["format"] == "ai-dnd-adventure-v3" + assert len(parsed["actions"]) > 40 + + # And honest about what this version can do with it. + assert response.headers["X-Importable-By-This-Version"] == "false" + warning = response.headers["X-Export-Warning"] + assert limits.import_limit_label() in warning + assert "exported successfully" in warning + assert "cannot import" in warning + assert int(response.headers["X-Export-Bytes"]) == len(body) + + +def test_the_warning_follows_the_configured_limit(monkeypatch): + """The text is generated from the constant, so changing the constant changes + the sentence rather than leaving a stale number in it.""" + assert "20 MB" in limits.oversized_export_warning(21_000_000) + monkeypatch.setattr(limits, "MAX_IMPORT_BODY_BYTES", 50 * 1024 * 1024) + assert limits.import_limit_label() == "50 MB" + assert "50 MB" in limits.oversized_export_warning(60_000_000) + assert "20 MB" not in limits.oversized_export_warning(60_000_000) + + +def test_the_bundle_itself_never_carries_the_warning(client): + """The warning is about the export, not part of the portable story file.""" + fill_past_the_limit(client.adv_id) + parsed = json.loads(export(client, client.adv_id).content) + flat = json.dumps(parsed).lower() + assert "import limit" not in flat + assert "cannot import" not in flat + for key in parsed: + assert "warning" not in key.lower() + + +def test_that_same_bundle_is_refused_by_import_naming_the_limit(client): + fill_past_the_limit(client.adv_id) + body = export(client, client.adv_id).content + response = client.post("/api/adventures/import", content=body, + headers={"Content-Type": "application/json"}) + assert response.status_code == 413 + detail = response.json()["detail"] + assert "too large" in detail.lower() + assert limits.import_limit_label() in detail + + +def test_a_bundle_under_the_limit_still_imports(client): + """The refusal is about size alone: the ordinary path is untouched.""" + body = export(client, client.adv_id).content + assert len(body) < limits.MAX_IMPORT_BODY_BYTES + response = client.post("/api/adventures/import", content=body, + headers={"Content-Type": "application/json"}) + assert response.status_code == 201, response.text[:200] diff --git a/backend/tests/test_v11_e_contrast.py b/backend/tests/test_v11_e_contrast.py new file mode 100644 index 0000000..32f6c08 --- /dev/null +++ b/backend/tests/test_v11_e_contrast.py @@ -0,0 +1,152 @@ +"""v1.1 WP-E: the contrast audit is a gate, not a report. + +Before this package `tools/contrast_audit.py` measured control boundaries, +printed that two of them were below 3:1, and exited 0 anyway — on the argument +that a control is identified by its label rather than its edge. A check that +cannot fail is not a check, and these tests are what make it one: the threshold +is exercised from both sides, on a real tokens file, so a future palette change +that dims a control's edge stops the run instead of adding a line to it. + + python -m pytest tests/test_v11_e_contrast.py -v +""" + +import sys +from pathlib import Path + +import pytest + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tools")) + +import contrast_audit as audit # noqa: E402 + + +# --------------------------------------------------------------- the maths + +def test_the_ratio_is_the_wcag_ratio(): + """Anchored on values with known answers, so a broken formula is visible.""" + assert audit.ratio("#ffffff", "#000000") == pytest.approx(21.0, abs=0.01) + assert audit.ratio("#ffffff", "#ffffff") == pytest.approx(1.0, abs=0.001) + # Order does not matter: contrast is symmetric. + assert audit.ratio("#131320", "#676792") == pytest.approx( + audit.ratio("#676792", "#131320"), abs=1e-9) + + +# ------------------------------------------------------- the gate itself + +def tokens_file(tmp_path: Path, **overrides: str) -> Path: + """A real tokens file with named colours replaced.""" + source = audit.TOKENS.read_text() + for name, value in overrides.items(): + token = "--" + name.replace("_", "-") + start = source.index(f"{token}: ") + end = source.index(";", start) + source = source[:start] + f"{token}: {value}" + source[end:] + written = tmp_path / "tokens.css" + written.write_text(source) + return written + + +def run_against(path: Path, monkeypatch) -> int: + monkeypatch.setattr(audit, "TOKENS", path) + return audit.main() + + +def test_a_boundary_below_three_to_one_fails_the_run(tmp_path, monkeypatch, capsys): + """2.99:1 against --bg-input — the wrong side of the line by one hundredth. + + This value clears 3:1 against --bg-panel (3.21:1), so it would have passed + the audit as M11 wrote it. It fails now because the floor is taken against + the background the control is actually drawn on. + """ + below = tokens_file(tmp_path, border="#58639a") + assert run_against(below, monkeypatch) == 1 + printed = capsys.readouterr().out + assert "2.99:1" in printed + assert "FAIL" in printed + assert "below 3:1 (WCAG 1.4.11)" in printed + + +def test_a_boundary_at_three_to_one_passes(tmp_path, monkeypatch, capsys): + """3.00:1 — the right side of the same line, one hundredth from the value + above, and not failed for arithmetic the reader cannot see.""" + at = tokens_file(tmp_path, border="#58639b") + assert run_against(at, monkeypatch) == 0 + printed = capsys.readouterr().out + assert "3.00:1" in printed + assert "every control boundary clears 3:1" in printed + + +def test_the_v1_0_0_boundary_would_now_fail(tmp_path, monkeypatch, capsys): + """The value v1.0.0 shipped. This is the defect WP-E closes, and the gate + has to be the thing that would have caught it.""" + shipped = tokens_file(tmp_path, border="#2b2b3d") + assert run_against(shipped, monkeypatch) == 1 + assert "1.24:1" in capsys.readouterr().out # against --bg-input, the worst case + + +def test_a_text_pair_below_four_point_five_still_fails(tmp_path, monkeypatch, capsys): + """WP-E raised the boundary floor and must not have lowered the text one.""" + dimmed = tokens_file(tmp_path, text_dim="#5a5750") + assert run_against(dimmed, monkeypatch) == 1 + assert "below WCAG AA (1.4.3)" in capsys.readouterr().out + + +def test_a_missing_token_is_a_failure_not_a_skip(tmp_path, monkeypatch, capsys): + source = audit.TOKENS.read_text().replace("--border-bright:", "--border-was-renamed:") + written = tmp_path / "tokens.css" + written.write_text(source) + assert run_against(written, monkeypatch) == 1 + assert "MISSING TOKEN" in capsys.readouterr().out + + +# ------------------------------------------------- the shipped palette + +def test_the_real_tokens_pass_both_criteria(capsys): + """The palette as it stands, through the same gate CI would run.""" + assert audit.main() == 0 + printed = capsys.readouterr().out + assert "every text pair clears WCAG AA (1.4.3)" in printed + assert "every control boundary clears 3:1 (1.4.11)" in printed + assert "FAIL" not in printed + + +def test_every_boundary_pair_is_measured_against_the_background_it_is_drawn_on(): + """The audit used to check borders only against --bg-panel, which is not + where the bordered controls are: inputs and buttons sit on --bg-input + (styles/forms.css), which is lighter and therefore harder. Checking only the + easier background would let a token pass while the real control failed.""" + boundary = [(fg, bg) for kind, fg, bg, _, _ in audit.PAIRS if kind == "boundary"] + for background in ("--bg-input", "--bg-panel", "--bg"): + assert ("--border", background) in boundary, background + assert ("--border-bright", "--bg-input") in boundary + + +def test_the_text_baselines_are_unchanged_by_wp_e(): + """WP-E changed only boundary tokens. These are the M11 text measurements, + and they have to still be exactly what the earlier reports recorded.""" + tokens = audit.read_tokens(audit.TOKENS) + measured = { + "body text on the page": audit.ratio(tokens["--text"], tokens["--bg"]), + "body text in a panel": audit.ratio(tokens["--text"], tokens["--bg-panel"]), + "secondary text in a panel": audit.ratio(tokens["--text-dim"], tokens["--bg-panel"]), + "secondary text on the page": audit.ratio(tokens["--text-dim"], tokens["--bg"]), + } + assert measured["body text on the page"] == pytest.approx(14.57, abs=0.01) + assert measured["body text in a panel"] == pytest.approx(13.57, abs=0.01) + assert measured["secondary text in a panel"] == pytest.approx(5.48, abs=0.01) + assert measured["secondary text on the page"] == pytest.approx(5.88, abs=0.01) + + +def test_the_hover_edge_stays_brighter_than_the_resting_edge(): + """Rest and hover have to remain distinguishable from each other, not merely + each clear the floor against the background.""" + tokens = audit.read_tokens(audit.TOKENS) + panel = tokens["--bg-panel"] + assert audit.ratio(tokens["--border-bright"], panel) > audit.ratio(tokens["--border"], panel) + + +def test_a_boundary_never_becomes_as_loud_as_body_text(): + """A control's edge that outshines the words inside it is its own defect.""" + tokens = audit.read_tokens(audit.TOKENS) + panel = tokens["--bg-panel"] + assert audit.ratio(tokens["--border-bright"], panel) < audit.ratio(tokens["--text"], panel) diff --git a/backend/tools/contrast_audit.py b/backend/tools/contrast_audit.py index c084632..c412bb5 100644 --- a/backend/tools/contrast_audit.py +++ b/backend/tools/contrast_audit.py @@ -26,16 +26,24 @@ TOKENS = Path(__file__).resolve().parent.parent.parent / "frontend/src/styles/to #: rather than combinatorial, because "every colour against every other" reports #: pairs that never meet on screen. #: -#: The `kind` matters and is not a way of grading on a curve. **text** pairs are -#: WCAG 1.4.3 Contrast (Minimum) and are what §21 of the M11 brief asks about; -#: they are pass/fail. **boundary** pairs are WCAG 1.4.11 Non-text Contrast, -#: which applies to "visual information required to identify user interface -#: components" — and in this design a control is identified by its *label*, -#: which is measured above and passes, not by its edge. So a boundary below 3:1 -#: is reported with its number and does not fail the run; what it would take to -#: turn it into a real failure is a control with no visible label, and there is -#: no such control (`tools/m11_browser.py` asserts every visible control has an -#: accessible name, and the story controls are text buttons). +#: The `kind` records which success criterion a pair is measured against — +#: **text** is WCAG 1.4.3 Contrast (Minimum), **boundary** is WCAG 1.4.11 +#: Non-text Contrast — and **both are pass/fail**. +#: +#: v1.1 WP-E overturned the earlier position here, which was that a boundary +#: below 3:1 could be recorded rather than failed because "a control is +#: identified by its label, not by its edge". That argument understates what +#: 1.4.11 asks: the criterion covers the visual information needed to identify +#: a component *and its boundary*, and a reader who cannot see where a text box +#: ends cannot see that there is a text box to type into, label or no label. +#: The edges were at 1.33:1 and 1.75:1 — the tokens were raised instead. +#: +#: Borders are measured against **every background they are drawn on**, and the +#: floor is the worst of them. Inputs and buttons sit on --bg-input, which is +#: lighter than --bg-panel and so the harder case; checking only --bg-panel +#: would have let a token pass the audit while the real control failed. +#: `--bg-panel-glass` is translucent and cannot be resolved from tokens alone; +#: that edge is measured on the rendered page by `tools/m11_browser.py`. PAIRS = [ ("text", "--text", "--bg", 4.5, "body text on the page"), ("text", "--text", "--bg-panel", 4.5, "body text in a panel"), @@ -48,8 +56,14 @@ PAIRS = [ ("text", "--danger", "--bg-panel", 4.5, "an error message"), ("text", "--warning", "--bg-panel", 4.5, "a caution message"), ("text", "--player", "--bg", 4.5, "the player's own words"), - ("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge"), - ("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge"), + ("boundary", "--border", "--bg-input", 3.0, "a field or button's resting edge"), + ("boundary", "--border-bright", "--bg-input", 3.0, "a field or button's hover edge"), + ("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge in a panel"), + ("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge in a panel"), + ("boundary", "--border", "--bg", 3.0, "a divider on the page"), + ("boundary", "--border-bright", "--bg", 3.0, "the scrollbar thumb"), + ("boundary", "--accent-dim", "--bg-input", 3.0, "a focused field's edge"), + ("boundary", "--accent-dim", "--bg-panel", 3.0, "a focused control's edge in a panel"), ("boundary", "--chart-1", "--bg-panel", 3.0, "a chart bar"), ("boundary", "--chart-2", "--bg-panel", 3.0, "a chart bar"), ("boundary", "--chart-3", "--bg-panel", 3.0, "a chart bar"), @@ -81,37 +95,44 @@ def ratio(a: str, b: str) -> float: def main() -> int: tokens = read_tokens(TOKENS) - print(f"{TOKENS.relative_to(TOKENS.parents[3])}: {len(tokens)} colour tokens\n") + # Named defensively: the tests run this against a temporary tokens file, + # which need not sit four directories deep the way the real one does. + label = TOKENS.name + if len(TOKENS.parents) > 3: + label = TOKENS.relative_to(TOKENS.parents[3]) + print(f"{label}: {len(tokens)} colour tokens\n") print(f"{'pair':44} {'kind':9} {'ratio':>7} {'floor':>6} verdict") print("-" * 82) - failures, advisories = 0, 0 + text_failures, boundary_failures = 0, 0 for kind, foreground, background, floor, description in PAIRS: if foreground not in tokens or background not in tokens: print(f"{description:44} {kind:9} {'—':>7} {floor:>6.1f} MISSING TOKEN") - failures += 1 + text_failures += 1 continue measured = ratio(tokens[foreground], tokens[background]) - ok = measured >= floor - if not ok: - if kind == "text": - failures += 1 - verdict = "FAIL" - else: - advisories += 1 - verdict = "below 1.4.11 (label carries it)" - else: + # Rounded to the two decimals printed, so the verdict matches what the + # reader is shown: a pair reported as 3.00:1 is not failed for arithmetic + # the output does not display. + if round(measured, 2) >= floor: verdict = "pass" + else: + verdict = "FAIL" + if kind == "text": + text_failures += 1 + else: + boundary_failures += 1 print(f"{description:44} {kind:9} {measured:>6.2f}:1 {floor:>6.1f} {verdict}") print() - if failures: - print(f"{failures} text pair(s) below WCAG AA — this is a defect") + if text_failures: + print(f"{text_failures} text pair(s) below WCAG AA (1.4.3) — this is a defect") else: print("every text pair clears WCAG AA (1.4.3)") - if advisories: - print(f"{advisories} boundary pair(s) below 3:1 (1.4.11). Recorded rather " - "than failed: every control in this design carries a visible text " - "label, which is measured above and passes.") - return 1 if failures else 0 + if boundary_failures: + print(f"{boundary_failures} boundary pair(s) below 3:1 (WCAG 1.4.11) — " + "this is a defect") + else: + print("every control boundary clears 3:1 (1.4.11)") + return 1 if (text_failures or boundary_failures) else 0 if __name__ == "__main__": diff --git a/backend/tools/m11_browser.py b/backend/tools/m11_browser.py index 16b85bf..217ed94 100644 --- a/backend/tools/m11_browser.py +++ b/backend/tools/m11_browser.py @@ -665,6 +665,228 @@ def check_accessibility(browser: Browser, site: Site, adv: int, checks: Checks): checks.record("A11y", "the story input takes keyboard focus", typed is True) +# -------------------------------------------------------------- WP-E boundaries + +#: The controls whose edges WP-E measures, as the stylesheets actually draw them. +#: `side` is the border the rule sets: .topnav paints only a bottom edge. +BOUNDARY_TARGETS = [ + ("E01", "the story composer", ".input-bar", "Top"), + ("E02", "a story control", ".story-controls button:not(:disabled)", "Top"), + ("E04", "the open panel tab", ".panel-tabs button.active", "Top"), +] + +#: `.topnav` is measured on the library route, because the play route does not +#: render it — the play page has its own `.play-header`, which story.css keeps +#: deliberately opaque. So the translucent case exists on exactly one screen, +#: and it is the case token arithmetic cannot answer: --bg-panel-glass is +#: rgba(...,0.82) over a gradient, so only the rendered page knows what is +#: behind that edge. +TRANSLUCENT_TARGET = ("E03", "the top navigation's edge", ".topnav", "Bottom") + +#: Measured in the browser rather than from tokens, because two of these cannot +#: be derived from tokens at all: .topnav sits on --bg-panel-glass, which is +#: translucent, so its effective background is a composite of what is behind it; +#: and a rendered edge can be changed by opacity, a transition mid-flight, or a +#: rule the token file knows nothing about. +#: +#: A boundary is measured against **both** adjacent colours — the control's own +#: fill inside it and the background outside it — and passes on the better of +#: the two. An edge that matches its fill but contrasts with the page is still a +#: visible outline, and vice versa; what 1.4.11 asks is that the component's +#: extent be perceivable, not that every neighbouring surface differ from it. +BOUNDARY_JS = """ + const parse = (c) => (c.match(/[\\d.]+/g) || []).map(Number); + const alphaOf = (c) => { + if (!c || c === 'transparent') return 0; + const p = parse(c); + return p.length > 3 ? p[3] : 1; + }; + function lum(c) { + const [r, g, b] = parse(c).slice(0, 3).map(v => v / 255) + .map(v => v <= 0.03928 ? v / 12.92 : Math.pow((v + 0.055) / 1.055, 2.4)); + return 0.2126 * r + 0.7152 * g + 0.0722 * b; + } + function over(fg, bg) { + const f = parse(fg), b = parse(bg), a = alphaOf(fg); + return 'rgb(' + [0, 1, 2].map(i => Math.round(f[i] * a + b[i] * (1 - a))).join(', ') + ')'; + } + function ratio(x, y) { + const a = lum(x), b = lum(y); + return (Math.max(a, b) + 0.05) / (Math.min(a, b) + 0.05); + } + // Everything painted behind `el`, composited bottom-up, so a translucent + // panel reports the colour a reader actually sees rather than its own rgba. + function behind(el) { + const layers = []; + let node = el; + while (node && node !== document.documentElement) { + const c = getComputedStyle(node).backgroundColor; + if (alphaOf(c) > 0) { + layers.push(c); + if (alphaOf(c) >= 1) break; + } + node = node.parentElement; + } + let result = 'rgb(10, 10, 15)'; + const root = getComputedStyle(document.documentElement).backgroundColor; + const body = getComputedStyle(document.body).backgroundColor; + if (alphaOf(root) >= 1) result = root; + else if (alphaOf(body) >= 1) result = body; + for (let i = layers.length - 1; i >= 0; i--) { + result = alphaOf(layers[i]) >= 1 ? layers[i] : over(layers[i], result); + } + return result; + } + const [selector, side] = arguments; + const el = document.querySelector(selector); + if (!el) return null; + const s = getComputedStyle(el); + const edge = s['border' + side + 'Color']; + const width = parseFloat(s['border' + side + 'Width']) || 0; + const opacity = parseFloat(s.opacity); + const inside = alphaOf(s.backgroundColor) >= 1 + ? s.backgroundColor : over(s.backgroundColor, behind(el.parentElement || el)); + const outside = behind(el.parentElement || el); + return { + selector, edge, width, opacity, inside, outside, + insideRatio: Math.round(ratio(edge, inside) * 100) / 100, + outsideRatio: Math.round(ratio(edge, outside) * 100) / 100, + outline: s.outlineStyle + ' ' + s.outlineWidth + ' ' + s.outlineColor, + shadow: s.boxShadow, + }; +""" + + +def _settled(browser: Browser, selector: str, side: str, *, timeout: float = 5) -> None: + """Waits until the edge colour stops moving. + + Every one of these controls carries `transition: border-color 0.15s`, so a + measurement taken the instant after a hover or a focus reads a colour part + way between the two states — a value no state actually has. The first run of + this check did exactly that: it reported the hover edge as rgb(114,114,160), + which is neither --border nor --border-bright but a frame between them. + Polled rather than slept, per this harness's rule. + """ + script = ("(() => { const el = document.querySelector(%s);" + " if (!el) return true;" + " const c = getComputedStyle(el)['border%sColor'];" + " const was = window.__wpeEdge; window.__wpeEdge = c;" + " return was === c; })()" % (json.dumps(selector), side)) + # Cleared first, so the comparison starts from no previous reading rather + # than from whatever the last measured control left behind. + browser.js("window.__wpeEdge = undefined") + browser.wait_js(script, timeout=timeout) + + +def _boundary(browser: Browser, selector: str, side: str) -> dict | None: + _settled(browser, selector, side) + return browser.js(BOUNDARY_JS, selector, side) + + +def _record_boundary(checks: Checks, test: str, what: str, row: dict | None) -> None: + if row is None: + checks.skip(test, what, "the control was not on the page") + return + best = max(row["insideRatio"], row["outsideRatio"]) + detail = (f"{best:.2f}:1 (edge {row['edge']} — {row['insideRatio']}:1 against its fill " + f"{row['inside']}, {row['outsideRatio']}:1 against {row['outside']}), " + f"{row['width']}px") + checks.record("WCAG 1.4.11", what, best >= 3.0 and row["width"] > 0, detail) + + +def check_control_boundaries(browser: Browser, site: Site, adv: int, checks: Checks, + evidence: dict, out: Path) -> None: + """v1.1 WP-E §21: a control's edge is visible on its own, measured. + + M11 measured text contrast on the rendered page and left boundaries to the + token audit, which checked them against one background and reported a + shortfall without failing. These are the rendered edges, at rest, on hover + and while focused, against what is actually behind them. + """ + _open_play(browser, site, adv) + measured: dict[str, dict] = {} + + # A panel tab's edge is `transparent` until the panel is open, which makes + # the open tab the one control here whose boundary is the only thing marking + # it — exactly what 1.4.11 is about. So one is opened rather than skipped. + if not _open_panel(browser, "State"): + checks.skip("E04", "the open panel tab — resting edge", "the State panel did not open") + + for test, what, selector, side in BOUNDARY_TARGETS: + row = _boundary(browser, selector, side) + _record_boundary(checks, test, f"{what} — resting edge", row) + if row: + measured[f"{what} (rest)"] = row + + # Hover, with a real pointer: `:hover` follows the browser's pointer state, + # so a synthetic mouseover would silently re-measure the resting edge. + control = browser.find(".story-controls button:not(:disabled)", required=False) + if control is None: + checks.skip("E02", "a story control — hover edge", "no enabled control on the page") + else: + browser.hover(control) + hovered = browser.js( + "return document.querySelector('.story-controls button:not(:disabled)')" + " .matches(':hover');") + if hovered is not True: + checks.skip("E02", "a story control — hover edge", + "the pointer did not land on the control") + else: + row = _boundary(browser, ".story-controls button:not(:disabled)", "Top") + _record_boundary(checks, "E02", "a story control — hover edge", row) + if row: + measured["a story control (hover)"] = row + browser.unhover() + + # Focus: the composer's edge changes colour while it holds focus, and that + # edge is what tells a keyboard reader where they are. + focused = browser.js(""" + const box = document.querySelector('.input-main textarea'); + if (!box) return false; + box.focus(); + return document.activeElement === box; + """) + if focused is not True: + checks.skip("E01", "the story composer — focused edge", "the composer did not take focus") + else: + row = _boundary(browser, ".input-bar", "Top") + _record_boundary(checks, "E01", "the story composer — focused edge", row) + if row: + measured["the story composer (focus)"] = row + # A focus ring that is only a colour change is not enough on its own; + # M11 already asserts a visible focus indicator, and this says the + # focused edge is also measurably distinct from the resting one. + rest = measured.get("the story composer (rest)") + if rest: + checks.record("WCAG 1.4.11", "the focused composer edge differs from its resting edge", + row["edge"] != rest["edge"], + f"rest {rest['edge']} -> focus {row['edge']}") + + shot = browser.screenshot(out / "control-boundaries.png") + checks.record("WP-E", "a screenshot of the measured controls was captured", + shot.exists() and shot.stat().st_size > 0, str(shot)) + + # The translucent edge, on the only screen that has it. Token arithmetic + # cannot reach this one: --bg-panel-glass is rgba over a gradient, so what + # is behind the nav's bottom edge is known only to the rendered page. + test, what, selector, side = TRANSLUCENT_TARGET + browser.go(site.url) + browser.wait_for(".topnav", timeout=30) + row = _boundary(browser, selector, side) + _record_boundary(checks, test, f"{what} — over a translucent panel", row) + if row: + measured[f"{what} (translucent)"] = row + checks.record("WP-E", "the nav's background really is translucent", + row["outside"] != row["inside"], + f"composited to {row['inside']} over {row['outside']}") + nav_shot = browser.screenshot(out / "control-boundaries-nav.png") + checks.record("WP-E", "a screenshot of the navigation edge was captured", + nav_shot.exists() and nav_shot.stat().st_size > 0, str(nav_shot)) + + evidence["control_boundaries"] = {"measured": measured, + "screenshots": [str(shot), str(nav_shot)]} + + # -------------------------------------------------------------- WP-C scenarios def _take_count_is(count: str) -> str: @@ -1354,7 +1576,11 @@ def main() -> int: "C5", "failed generation and recovery"), ("export", lambda: check_export_download(browser, site, adv, checks, evidence, out)), ] - for suite, scenarios in (("M11", m11), ("WP-C", wpc)): + wpe = [ + ("boundaries", lambda: check_control_boundaries(browser, site, adv, checks, + evidence, out)), + ] + for suite, scenarios in (("M11", m11), ("WP-C", wpc), ("WP-E", wpe)): checks.suite = suite for name, scenario in scenarios: if not wanted(name): @@ -1383,7 +1609,8 @@ def main() -> int: "kind": ("release regression" if narrated and not only else "partial (no narrator)" if not narrated else f"development (only {sorted(only)})"), "checks": checks.rows, - "suites": {"M11": checks.counts("M11"), "WP-C": checks.counts("WP-C")}, + "suites": {"M11": checks.counts("M11"), "WP-C": checks.counts("WP-C"), + "WP-E": checks.counts("WP-E")}, "passed": checks.counts()["passed"], "failed": checks.counts()["failed"], "skipped": checks.counts()["skipped"], diff --git a/backend/tools/m11_webdriver.py b/backend/tools/m11_webdriver.py index 118f22d..122dfa5 100644 --- a/backend/tools/m11_webdriver.py +++ b/backend/tools/m11_webdriver.py @@ -25,6 +25,7 @@ has actually finished arriving — never the click that started it. from __future__ import annotations +import base64 import json import os import shutil @@ -226,6 +227,45 @@ class Browser: def source(self) -> str: return self._call("GET", self._s("/source"))["value"] + def screenshot(self, path) -> Path: + """The viewport as a PNG, written where you ask (v1.1 WP-E). + + Evidence for a change a reader judges by looking at it: a contrast ratio + says a boundary is measurable, and a picture says what it looks like. + """ + encoded = self._call("GET", self._s("/screenshot"))["value"] + target = Path(path) + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(base64.b64decode(encoded)) + return target + + def hover(self, element: str) -> None: + """A real pointer over an element, so `:hover` actually applies. + + Dispatching a mouseover event from JavaScript does not do this: CSS + `:hover` follows the browser's own pointer state, not a synthetic event, + so a measurement taken after `dispatchEvent` reads the resting style and + reports it as the hover style. This moves the pointer (v1.1 WP-E). + """ + self._call("POST", self._s("/execute/sync"), { + "script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})", + "args": [{ELEMENT_KEY: element}]}) + self._call("POST", self._s("/actions"), {"actions": [{ + "type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"}, + "actions": [{"type": "pointerMove", "duration": 60, + "origin": {ELEMENT_KEY: element}, "x": 0, "y": 0}]}]}) + + def unhover(self) -> None: + """Move the pointer off whatever it was over, and forget the input state.""" + self._call("POST", self._s("/actions"), {"actions": [{ + "type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"}, + "actions": [{"type": "pointerMove", "duration": 30, + "origin": "viewport", "x": 0, "y": 0}]}]}) + try: + self._call("DELETE", self._s("/actions")) + except WebDriverError: + pass + def js(self, script: str, *args): return self._call("POST", self._s("/execute/sync"), {"script": script, "args": list(args)})["value"] diff --git a/backend/tools/wpd_backup_ui.py b/backend/tools/wpd_backup_ui.py new file mode 100644 index 0000000..41510f9 --- /dev/null +++ b/backend/tools/wpd_backup_ui.py @@ -0,0 +1,353 @@ +"""v1.1 WP-D criterion 3: the backup completes **through the UI**. + + python -m tools.wpd_backup_ui --out [--case real|large|both] [--show] + +Run from `backend/`, with `frontend/dist` already built. + +WP-D's first pass drove `POST /api/backups` — the endpoint the *Back up now* +button calls — on a 2.3 MB campaign database and a 117 MB one. That is evidence +about the implementation, and the plan's criterion 3 asks for something else: +that the backup *completes through the UI* on both. A reader does not call an +endpoint. This drives the reader-facing control in a real Firefox, against the +production build served by FastAPI, exactly as WP-C's harness does. + +**No narrator and no inference.** A backup needs neither, so nothing here +touches a model host. + +What it refuses to call a pass, per the brief: + +- the click does nothing (no toast, no file); +- the request fails (an error toast); +- no backup file appears on disk; +- the finished copy does not pass a full `PRAGMA integrity_check`; +- the UI reports an error. + +Every wait is on a condition the page or the filesystem can show. Nothing here +sleeps and then asserts. +""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import shutil +import sqlite3 +import sys +import time +from datetime import datetime +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) + +from tools.m11_webdriver import ( # noqa: E402 + Browser, Site, geckodriver_version, require_under_home, +) + +BACKEND = Path(__file__).resolve().parent.parent +HOME = Path.home() + +#: The two databases criterion 3 names. Both are copied before use: the first is +#: the M11 evidence campaign and must not be written to, and the second is the +#: 117 MB application database built for WP-D's timing. +DEFAULT_REAL = HOME / "m11-evidence/m04-final/campaign.db" +DEFAULT_LARGE = HOME / "v11-evidence/wp-d/endpoint/campaign-100mb/campaign.db" + +BLOCK = '[data-testid="database-backup"]' +BUTTON = f'{BLOCK} button.primary' + + +class Checks: + """Results, with the discipline that an unrun check is not a passing one.""" + + def __init__(self) -> None: + self.rows: list[dict] = [] + + def record(self, case: str, name: str, ok: bool, detail: str = "") -> bool: + self.rows.append({"case": case, "check": name, + "result": "PASS" if ok else "FAIL", "detail": detail}) + print(f" {'ok ' if ok else 'FAIL'} {case:6} {name}" + + (f" — {detail}" if detail else ""), flush=True) + return ok + + @property + def failed(self) -> list[dict]: + return [r for r in self.rows if r["result"] == "FAIL"] + + +# ----------------------------------------------------------------- helpers + +def integrity_of(path: Path) -> tuple[str, float]: + """The full check on a finished copy, and what it cost, measured here. + + The application runs its own `integrity_check` before keeping the file; this + is an independent second opinion on the artefact the UI produced, and it is + where criterion 3's 'time for integrity_check' comes from. + """ + connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True) + try: + started = time.perf_counter() + rows = connection.execute("PRAGMA integrity_check").fetchall() + elapsed = time.perf_counter() - started + finally: + connection.close() + return ", ".join(str(r[0]) for r in rows), elapsed + + +def digest(path: Path) -> str: + return hashlib.sha256(path.read_bytes()).hexdigest()[:16] + + +def backups_in(db_path: Path) -> dict[str, int]: + directory = db_path.parent / "backups" + if not directory.exists(): + return {} + return {p.name: p.stat().st_size for p in directory.iterdir() if p.is_file()} + + +def wait_for_new_backup(db_path: Path, before: set[str], *, timeout: float = 300, + poll: float = 0.2) -> Path | None: + """The backup file the application wrote, once it is really there. + + A file condition, not a sleep: a name that was not there before, not a + `.partial`, and a size that has stopped growing. + """ + directory = db_path.parent / "backups" + deadline = time.monotonic() + timeout + sizes: dict[str, int] = {} + while time.monotonic() < deadline: + if directory.exists(): + for path in directory.iterdir(): + if not path.is_file() or path.name in before: + continue + if path.name.endswith(".partial"): + continue + size = path.stat().st_size + if size > 0 and sizes.get(path.name) == size: + return path + sizes[path.name] = size + time.sleep(poll) + return None + + +def toasts(browser: Browser) -> list[dict]: + """What the page is telling the reader — the message apart from its mark. + + A toast renders a decorative mark before the message: + `…`. So the + button's `textContent` begins with that character, and the first run of this + tool matched `textContent.startswith('Backup written:')` and reported six + *passing* behaviours as failures — the UI had said exactly what it should, + and the assertion was reading the mark. `message` is the message span alone; + `text` is kept whole for the evidence record. + """ + return browser.js(""" + const host = document.querySelector('.toast-host'); + if (!host) return []; + return [...host.querySelectorAll('button.toast')].map(b => { + const span = b.querySelector('span:not(.toast-mark)'); + return { + error: b.classList.contains('toast-error'), + message: (span ? span.textContent : b.textContent).trim(), + text: b.textContent.trim(), + }; + }); + """) or [] + + +def open_backup_panel(browser: Browser, site: Site, checks: Checks, case: str) -> bool: + """Reach the control the way a reader does: the nav link, then the panel.""" + browser.go(site.url) + browser.wait_for(".topnav", timeout=60) + link = browser.find('.nav-links a[href="/settings"]', required=False) + if link is None: + return checks.record(case, "the Settings link is in the navigation", False) + browser.click(link) + opened = browser.wait_js(f"!!document.querySelector('{BLOCK}')", timeout=60) + if not checks.record(case, "Settings opens from the navigation link", opened, + browser.url): + return False + + summary = browser.find(f"{BLOCK} summary", required=False) + if summary is None: + return checks.record(case, "the backup panel has a disclosure", False) + browser.click(summary) + # The panel is a
: it loads what is already on disk when it opens, + # so waiting for the directory line proves the application answered. + shown = browser.wait_js( + f"document.querySelector('{BLOCK}').open === true" + f" && !!document.querySelector('{BUTTON}')", timeout=30) + checks.record(case, "the backup panel opens and shows its control", shown) + listed = browser.wait_js(f"!!document.querySelector('{BLOCK} code')", timeout=30) + checks.record(case, "the panel reports where backups are written", listed) + return shown + + +def click_back_up_now(browser: Browser, site: Site, db: Path, checks: Checks, + case: str, label: str) -> dict: + """One press of the button, judged by what the page and the disk then show.""" + before = set(backups_in(db)) + # Clear anything still on screen, so the toast this press produces is the + # one that is read back rather than a leftover from the previous press. + browser.js(""" + for (const b of document.querySelectorAll('.toast-host button.toast')) b.click(); + return true; + """) + button = browser.find(BUTTON, required=False) + if button is None: + checks.record(case, f"{label}: the control is on the page", False) + return {} + + started = time.perf_counter() + browser.click(button) + + # Either outcome ends the wait, so a failure is reported as a failure rather + # than as a timeout. + settled = browser.wait_js( + "(() => { const t = [...document.querySelectorAll('.toast-host button.toast')];" + " return t.length > 0; })()", timeout=300) + elapsed = time.perf_counter() - started + + shown = toasts(browser) + errors = [t for t in shown if t["error"]] + written = [t for t in shown if t["message"].startswith("Backup written:")] + + checks.record(case, f"{label}: the click produced a visible result", settled, + json.dumps(shown)[:200]) + checks.record(case, f"{label}: the UI reports no error", + not errors, json.dumps(errors)[:300]) + checks.record(case, f"{label}: the UI reports the backup was written", + bool(written), json.dumps(written)[:200]) + + idle = browser.wait_js( + "(() => { const b = document.querySelector(%s);" + " return !!b && b.textContent.trim() === 'Back up now'; })()" + % json.dumps(BUTTON), timeout=120) + checks.record(case, f"{label}: the control returns from 'Backing up…'", idle) + + produced = wait_for_new_backup(db, before) + checks.record(case, f"{label}: a backup file was physically produced", + produced is not None, str(produced)) + if produced is None: + return {"seconds": round(elapsed, 3), "toasts": shown} + + named = any(produced.name in t["message"] for t in written) + checks.record(case, f"{label}: the UI names the file that appeared", named, + f"{produced.name} — {written[0]['message'] if written else ''}") + + relisted = browser.wait_js( + "[...document.querySelectorAll('%s .backup-list code')]" + ".some(c => c.textContent.trim() === %s)" % (BLOCK, json.dumps(produced.name)), + timeout=60) + checks.record(case, f"{label}: the new backup appears in the panel's list", relisted) + + verdict, integrity_seconds = integrity_of(produced) + checks.record(case, f"{label}: the finished copy passes full integrity_check", + verdict == "ok", f"{verdict} in {integrity_seconds * 1000:.1f} ms") + + return { + "file": str(produced), + "bytes": produced.stat().st_size, + "sha256_16": digest(produced), + "seconds": round(elapsed, 3), + "integrity": verdict, + "integrity_seconds": round(integrity_seconds, 4), + "toasts": shown, + } + + +def run_case(case: str, source: Path, out: Path, checks: Checks, *, show: bool, + twice: bool) -> dict: + print(f"\n=== {case}: {source} ===", flush=True) + work = out / case + shutil.rmtree(work, ignore_errors=True) + (work / "data").mkdir(parents=True) + db = work / "data" / "campaign.db" + copy_started = time.perf_counter() + shutil.copy2(source, db) + copy_seconds = time.perf_counter() - copy_started + size = db.stat().st_size + print(f" copied {size:,} bytes in {copy_seconds:.2f}s -> {db}", flush=True) + + site = Site(BACKEND, db, work / "server.log") + browser = Browser(headless=not show, log=work / "geckodriver.log") + result: dict = {"database": str(source), "bytes": size, + "served_at": site.url, "firefox": browser.version} + try: + if not open_backup_panel(browser, site, checks, case): + return result + result["first"] = click_back_up_now(browser, site, db, checks, case, "backup") + browser.screenshot(work / "back-up-now.png") + + if twice and result["first"].get("file"): + kept = Path(result["first"]["file"]) + before_bytes, before_digest = kept.stat().st_size, digest(kept) + result["second"] = click_back_up_now(browser, site, db, checks, case, + "second backup") + still_there = kept.exists() + checks.record(case, "the earlier backup still exists", still_there) + if still_there: + checks.record( + case, "and is byte-for-byte what it was", + kept.stat().st_size == before_bytes and digest(kept) == before_digest, + f"{before_bytes:,} bytes, sha256:{before_digest}") + verdict, _ = integrity_of(kept) + checks.record(case, "and still passes integrity_check", verdict == "ok", + verdict) + if result["second"].get("file"): + checks.record(case, "the second backup is a different file", + result["second"]["file"] != result["first"]["file"], + Path(result["second"]["file"]).name) + finally: + browser.quit() + site.stop() + return result + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--out", required=True, + help="evidence directory, which must be under $HOME") + parser.add_argument("--case", default="both", choices=["real", "large", "both"]) + parser.add_argument("--show", action="store_true", help="run Firefox visibly") + parser.add_argument("--real-db", default=str(DEFAULT_REAL)) + parser.add_argument("--large-db", default=str(DEFAULT_LARGE)) + args = parser.parse_args() + + out = require_under_home(Path(args.out).expanduser()) + out.mkdir(parents=True, exist_ok=True) + if not (BACKEND.parent / "frontend/dist/index.html").exists(): + print("frontend/dist is not built", file=sys.stderr) + return 2 + + checks = Checks() + started = datetime.now() + results: dict[str, dict] = {} + wanted = [("real", Path(args.real_db).expanduser(), True)] if args.case != "large" else [] + if args.case != "real": + wanted.append(("large", Path(args.large_db).expanduser(), False)) + + print(f"WP-D criterion 3 — the backup through the UI. geckodriver " + f"{geckodriver_version()}", flush=True) + for case, source, twice in wanted: + if not source.exists(): + checks.record(case, "the database is present", False, str(source)) + continue + results[case] = run_case(case, source, out, checks, show=args.show, twice=twice) + + report = { + "started": started.isoformat(timespec="seconds"), + "seconds": round((datetime.now() - started).total_seconds()), + "cases": results, + "checks": checks.rows, + "passed": len([r for r in checks.rows if r["result"] == "PASS"]), + "failed": len(checks.failed), + } + (out / "backup-ui-report.json").write_text(json.dumps(report, indent=2)) + print(f"\n{report['passed']} passed, {report['failed']} failed " + f"-> {out / 'backup-ui-report.json'}") + return 1 if checks.failed else 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/frontend/src/api.js b/frontend/src/api.js index 235a893..b9217da 100644 --- a/frontend/src/api.js +++ b/frontend/src/api.js @@ -151,7 +151,32 @@ export const api = { sendAction: (advId, payload, handlers, signal) => streamSSE(`/adventures/${advId}/actions`, payload, handlers, signal), retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal), - exportAdventure: (id) => request(`/adventures/${id}/export`), + // v1.1 WP-D: the bundle, plus what the server says about importing it back. + // The body is the bundle and nothing else — the browser saves exactly those + // bytes — so the size and the warning come back in headers. + exportAdventure: async (id) => { + const resp = await fetch(`/api/adventures/${id}/export`, { + headers: { 'Content-Type': 'application/json' }, + }) + if (!resp.ok) { + let detail = resp.statusText + try { detail = (await resp.json()).detail || detail } catch { /* non-JSON */ } + throw new Error(detail) + } + const bundle = await resp.json() + const number = (name) => { + const raw = Number(resp.headers.get(name)) + return Number.isFinite(raw) && raw > 0 ? raw : null + } + return { + bundle, + exportBytes: number('X-Export-Bytes'), + importLimitBytes: number('X-Import-Limit-Bytes'), + // Absent header (an older server) means nothing is claimed either way. + importable: resp.headers.get('X-Importable-By-This-Version') !== 'false', + warning: resp.headers.get('X-Export-Warning') || null, + } + }, importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }), // M9. A verified copy of the whole database, which is a different tool from diff --git a/frontend/src/pages/Campaigns.jsx b/frontend/src/pages/Campaigns.jsx index c797269..112d622 100644 --- a/frontend/src/pages/Campaigns.jsx +++ b/frontend/src/pages/Campaigns.jsx @@ -66,10 +66,12 @@ export default function Campaigns() { const exportOne = async (campaign) => { try { - const bundle = await api.exportAdventure(campaign.id) + const { bundle, warning } = await api.exportAdventure(campaign.id) const safe = (campaign.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60) downloadJSON(bundle, `${safe}.json`) - toast('Campaign exported.') + // v1.1 WP-D: the file is written either way. A campaign too large for this + // version to import back says so now, not when it is needed. + toast(warning || 'Campaign exported.', warning ? 'error' : undefined) } catch (err) { toast(classifyError(err.message).detail, 'error') } diff --git a/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx b/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx index 045fabd..fb6eb5b 100644 --- a/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx +++ b/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx @@ -79,10 +79,12 @@ export function CampaignSettingsPanel({ adventure, setAdventure, onError, moment const exportCampaign = async () => { try { - const bundle = await api.exportAdventure(adventure.id) + const { bundle, warning } = await api.exportAdventure(adventure.id) const safe = (adventure.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60) downloadJSON(bundle, `${safe}.json`) - toast('Campaign exported.') + // v1.1 WP-D: see Campaigns.jsx. The export is delivered; the warning says + // this version could not import the file back. + toast(warning || 'Campaign exported.', warning ? 'error' : undefined) } catch (err) { onError(classifyError(err.message).detail) } diff --git a/frontend/src/pages/exportHonesty.test.jsx b/frontend/src/pages/exportHonesty.test.jsx new file mode 100644 index 0000000..596da63 --- /dev/null +++ b/frontend/src/pages/exportHonesty.test.jsx @@ -0,0 +1,167 @@ +/* v1.1 WP-D: an export that says whether this version could import it back. + * + * The file is delivered either way — a campaign too large to re-import is not a + * damaged one, and refusing to write it would destroy the copy the reader was + * making. What changes is what they are told, and both reader-facing Export + * controls have to tell them: the one on the campaign card and the one in the + * campaign's own settings. + * + * The server decides. These assert that the page shows what it was given and + * keeps delivering the file, not that it re-derives the size policy. + */ + +import { screen, waitFor } from '@testing-library/react' +import userEvent from '@testing-library/user-event' +import { beforeEach, describe, expect, it, vi } from 'vitest' +import { api } from '../api' +import * as components from '../components' +import Campaigns from './Campaigns' +import { CampaignSettingsPanel } from './Play/panels/CampaignSettingsPanel' +import { mockModelStatus, renderWith } from '../test/helpers' + +const BUNDLE = { format: 'ai-dnd-adventure-v3', title: 'Long Campaign', actions: [] } + +const WARNING = + "This export is larger than this version's 20 MB import limit (21,230,000 bytes). " + + 'The file was exported successfully, but this version cannot import it.' + +const CAMPAIGN = { + id: 4, title: 'Long Campaign', action_count: 900, + updated_at: '2026-09-15T10:00:00', snippet: 'Rain over the harbour.', +} + +const ADVENTURE = { + id: 4, title: 'Long Campaign', ai_instructions: '', narration_length: 'brief', + canon_rules: [], persona_name: 'Aldric', persona_desc: '', +} + +// restoreAllMocks does not undo stubGlobal, and the fetch stub below would +// otherwise outlive its own describe block. +beforeEach(() => { vi.restoreAllMocks(); vi.unstubAllGlobals() }) + +function exportReturns({ warning = null } = {}) { + return vi.spyOn(api, 'exportAdventure').mockResolvedValue({ + bundle: BUNDLE, + exportBytes: warning ? 21_230_000 : 12_000, + importLimitBytes: 20 * 1024 * 1024, + importable: !warning, + warning, + }) +} + +async function library() { + mockModelStatus(api) + vi.spyOn(api, 'listAdventures').mockResolvedValue([CAMPAIGN]) + await renderWith() + await screen.findByText('Long Campaign') +} + +async function settingsPanel() { + mockModelStatus(api) + await renderWith( + , + ) +} + +describe('reading what the server said', () => { + // The tests below mock api.exportAdventure, so nothing there exercises the + // header names. These do: a typo in one of them would otherwise leave the + // whole suite green and the reader silently uninformed. + function serverSends(headers) { + const body = JSON.stringify(BUNDLE) + vi.stubGlobal('fetch', vi.fn().mockResolvedValue(new Response(body, { + status: 200, + headers: { 'Content-Type': 'application/json', ...headers }, + }))) + } + + it('reports a warning the server sent, with the sizes it named', async () => { + serverSends({ + 'X-Export-Bytes': '21230000', + 'X-Import-Limit-Bytes': String(20 * 1024 * 1024), + 'X-Importable-By-This-Version': 'false', + 'X-Export-Warning': WARNING, + }) + const result = await api.exportAdventure(4) + expect(result.bundle).toEqual(BUNDLE) + expect(result.warning).toBe(WARNING) + expect(result.importable).toBe(false) + expect(result.exportBytes).toBe(21_230_000) + expect(result.importLimitBytes).toBe(20 * 1024 * 1024) + }) + + it('claims nothing when an older server sends no headers', async () => { + serverSends({}) + const result = await api.exportAdventure(4) + expect(result.bundle).toEqual(BUNDLE) + expect(result.warning).toBeNull() + expect(result.importable).toBe(true) + expect(result.exportBytes).toBeNull() + expect(result.importLimitBytes).toBeNull() + }) + + it('treats an ordinary export as importable', async () => { + serverSends({ + 'X-Export-Bytes': '12000', + 'X-Import-Limit-Bytes': String(20 * 1024 * 1024), + 'X-Importable-By-This-Version': 'true', + }) + const result = await api.exportAdventure(4) + expect(result.importable).toBe(true) + expect(result.warning).toBeNull() + expect(result.exportBytes).toBe(12_000) + }) +}) + +describe('the campaign library export', () => { + it('delivers the file and says nothing more when it can be imported back', async () => { + const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {}) + exportReturns() + await library() + await userEvent.click(screen.getByRole('button', { name: 'Export' })) + await waitFor(() => expect(download).toHaveBeenCalledTimes(1)) + expect(download.mock.calls[0][0]).toEqual(BUNDLE) + expect(await screen.findByText('Campaign exported.')).toBeInTheDocument() + expect(screen.queryByText(/cannot import/)).toBeNull() + }) + + it('still delivers the file when it is too large, and says so', async () => { + const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {}) + exportReturns({ warning: WARNING }) + await library() + await userEvent.click(screen.getByRole('button', { name: 'Export' })) + // The file is written first: the warning is about importing it back, not + // about the export having failed. + await waitFor(() => expect(download).toHaveBeenCalledTimes(1)) + expect(download.mock.calls[0][0]).toEqual(BUNDLE) + const notice = await screen.findByText(/cannot import it/) + expect(notice).toBeInTheDocument() + expect(notice.textContent).toContain('20 MB') + expect(notice.textContent).toContain('exported successfully') + expect(screen.queryByText('Campaign exported.')).toBeNull() + }) +}) + +describe('the campaign settings export', () => { + it('delivers the file and confirms it when it can be imported back', async () => { + const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {}) + exportReturns() + await settingsPanel() + await userEvent.click(screen.getByRole('button', { name: 'Export campaign' })) + await waitFor(() => expect(download).toHaveBeenCalledTimes(1)) + expect(await screen.findByText('Campaign exported.')).toBeInTheDocument() + expect(screen.queryByText(/cannot import/)).toBeNull() + }) + + it('still delivers the file when it is too large, and says so', async () => { + const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {}) + exportReturns({ warning: WARNING }) + await settingsPanel() + await userEvent.click(screen.getByRole('button', { name: 'Export campaign' })) + await waitFor(() => expect(download).toHaveBeenCalledTimes(1)) + const notice = await screen.findByText(/cannot import it/) + expect(notice.textContent).toContain('20 MB') + expect(screen.queryByText('Campaign exported.')).toBeNull() + }) +}) diff --git a/frontend/src/styles/context.css b/frontend/src/styles/context.css index 16e13cd..4619ab0 100644 --- a/frontend/src/styles/context.css +++ b/frontend/src/styles/context.css @@ -59,7 +59,22 @@ .slice-4 { background: #6f9e8c; } .slice-5 { background: #c48a6a; } .slice-6 { background: #7c86b8; } -.slice-7 { background: var(--border-bright); } +/* v1.1 WP-E: pinned to the literal value this slice already rendered, instead + of borrowing --border-bright. A chart fill and a control edge have different + jobs: WP-E raised --border-bright to clear WCAG 1.4.11 (3:1) for control + boundaries, and that dragged this slice to #7a7aaa, an OKLab dE of 0.035 + from .slice-6 (#7c86b8) — two neighbouring slices the same colour. The + other slices sit 0.100-0.119 from their nearest neighbour; at #3d3d55 this + one sits 0.251, the most separated in the set. + + 1.4.11's 3:1 does not govern this: it is a proportional fill in a labelled + breakdown, not the boundary of a control, and what it needs is to be + distinguishable from the seven slices beside it. Reassigning it to a freer + hue was considered and rejected — inside the palette's own chroma and + lightness bands the only hues that beat 0.100 are pinks near 14 degrees, + which is --danger's territory and would paint an ordinary prompt section in + the colour this application reserves for failure. */ +.slice-7 { background: #3d3d55; } .ctx-block { border: 1px solid var(--border); diff --git a/frontend/src/styles/tokens.css b/frontend/src/styles/tokens.css index a0d374e..068dfd5 100644 --- a/frontend/src/styles/tokens.css +++ b/frontend/src/styles/tokens.css @@ -3,8 +3,15 @@ --bg-panel: #131320; --bg-panel-glass: rgba(19, 19, 32, 0.82); --bg-input: #1a1a2a; - --border: #2b2b3d; - --border-bright: #3d3d55; + /* v1.1 WP-E: control boundaries carry WCAG 1.4.11 (3:1 non-text contrast) on + their own, rather than leaning on the control's text label. The floor is + measured against --bg-input (#1a1a2a), not --bg-panel: inputs and buttons + are drawn on --bg-input (styles/forms.css), and it is the lightest of the + three backgrounds a border sits on, so it is the worst case. + --border 3.21:1 and --border-bright 4.24:1 there; higher on the others + (backend/tools/contrast_audit.py, which now fails below 3:1). */ + --border: #676792; + --border-bright: #7a7aaa; --text: #e2ddd0; --text-dim: #918c7d; --accent: #d4a94e; diff --git a/planning/README.md b/planning/README.md index ed6b500..d4f3a18 100644 --- a/planning/README.md +++ b/planning/README.md @@ -3,9 +3,13 @@ **This file is the index. Start here.** **Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: -WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`) are committed, and -WP-C, browser release coverage, is complete and staged for owner review** -(`reports/v1.1/V1.1-WP-C-REPORT.md`). +WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`), WP-B.2 (`0c1ba83`) and WP-C +(`59b5ebc`) are committed, and the last two planned packages — WP-D (recovery +honesty) and WP-E (control-boundary contrast) — are complete and staged for +owner review** (`reports/v1.1/V1.1-WP-D-REPORT.md`, +`reports/v1.1/V1.1-WP-E-REPORT.md`). The final browser run passed 101 checks +across all three suites with 0 failed and 0 skipped. Every planned v1.1 work +package is now implemented and reported; release validation has not begun. Phase 0 complete; AI-DnD forked as the production base; **milestones M1 through M11 complete and closed**. M11 was accepted at its closeout (2026-09-14), the v1 release gate passed on the release-candidate tree, and the diff --git a/planning/V1.1-PLAN.md b/planning/V1.1-PLAN.md index 256eb0d..e02fb12 100644 --- a/planning/V1.1-PLAN.md +++ b/planning/V1.1-PLAN.md @@ -16,10 +16,16 @@ real-model limitation** (owner decision, 2026-09-15): The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T. **WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release -coverage, is complete and staged for owner review. Its final run passed 91 checks +coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks (the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS, -including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E -have not started. No v1.1 version or tag exists. +including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). +**WP-D** (recovery honesty) and **WP-E** +(control-boundary contrast) are complete and staged for owner review, each with +its own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`. The final +browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0 +failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, +WP-D, WP-E) is now implemented and reported; release validation has not begun, +and no v1.1 version or tag exists. This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1 history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1 diff --git a/planning/VERSION.md b/planning/VERSION.md index 89e9176..94feaa4 100644 --- a/planning/VERSION.md +++ b/planning/VERSION.md @@ -1,8 +1,29 @@ # Planning Package Version -- **Package:** Adventure Storyteller Planning Package v4.4 +- **Package:** Adventure Storyteller Planning Package v4.5 - **Revision date:** 2026-09-15 -- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. **WP-C, browser release coverage, is complete and staged for owner review: 91 checks, 0 failed, 0 skipped, over trusted-LAN HTTPS, with real export downloads** (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E have not started, and no v1.1 version or tag exists. +- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C, browser release coverage, is committed and signed as `59b5ebc` (91 checks, 0 failed, 0 skipped, with real export downloads). **WP-D (recovery honesty) and WP-E (control-boundary contrast) are complete and staged for owner review**, each with its own report (`reports/v1.1/V1.1-WP-D-REPORT.md`, `V1.1-WP-E-REPORT.md`); the final browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0 failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) is now implemented and reported. Release validation has not begun, and no v1.1 version or tag exists. + +## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15) + +The last two planned v1.1 packages, implemented and reported separately. No +requirement, acceptance test, schema or bundle format changed. WP-D's evidence is +in `reports/v1.1/V1.1-WP-D-REPORT.md`, WP-E's in `V1.1-WP-E-REPORT.md`. + +| Document | Change | Kind | +| --- | --- | --- | +| `reports/v1.1/V1.1-WP-D-REPORT.md` | **New.** Full `integrity_check` on the finished backup copy, proved against a fixture `quick_check` calls healthy; an export that says when this version could not import it back, carried in headers because the response body *is* the bundle | work-package report | +| `reports/v1.1/V1.1-WP-E-REPORT.md` | **New.** Control boundaries raised to clear WCAG 1.4.11 (3:1), the contrast audit turned from a report into a gate, rendered before/after boundary measurements, and the `.slice-7` finding the change itself created | work-package report | +| `V1.1-PLAN.md` | Status: WP-C signed `59b5ebc`; WP-D and WP-E complete and staged | status | +| `planning/README.md` | Current state | index | +| `DEVELOPMENT.md` | How large an export can get: the 20 MB import ceiling, ~13 kB per action, M9's ~279-turn figure and why it is not a turn limit, and the four export headers | developer docs | + +**Requirement changes: zero.** + +Two owner decisions are outstanding, both recorded rather than assumed: WP-E's +before/after screenshots await approval (`OWNER SCREENSHOT APPROVAL: PENDING`), +and WP-D records that no real-browser click was made on *Back up now* — the +endpoint that button calls was driven instead. ## v4.4 — WP-C browser release coverage (2026-09-15) diff --git a/planning/reports/v1.1/V1.1-WP-D-REPORT.md b/planning/reports/v1.1/V1.1-WP-D-REPORT.md new file mode 100644 index 0000000..e1af014 --- /dev/null +++ b/planning/reports/v1.1/V1.1-WP-D-REPORT.md @@ -0,0 +1,526 @@ +# v1.1 WP-D — Recovery Honesty + +**Status:** COMPLETE — **PASS**. The decision, and what was deliberately not +claimed, is in §P. + +--- + +## A. Repository baseline + +| | | +| --- | --- | +| Branch | `v1.1-development` | +| HEAD at start | `59b5ebc2d85bdf53d38b6bcf347c496dd3432be1` — *v1.1 WP-C: browser release coverage*, signed by the owner (good signature, RSA key `02C9BF7D…`) | +| Its ancestry | `0c1ba83` (WP-B.2), `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `432f041` (v1.0.0) | +| Working tree at start | clean; nothing staged | +| WP-E | not started when WP-D was implemented | + +--- + +## B. Existing backup behaviour + +`backend/app/backup.py`, as v1 shipped it. The six questions the brief asks: + +| # | Question | Answer (before WP-D) | +| --- | --- | --- | +| 1 | How is the backup file created? | SQLite's **online backup API** (`sqlite3.Connection.backup`, `pages=-1`), from the live database opened read-only through a `mode=ro` URI. The destination is a temporary file `…​.db.partial` **in the destination directory**, so the rename below is atomic | +| 2 | Where does validation happen? | `_verify()`, on the finished copy, opened as its own read-only connection — not on the source, and not through the connection that wrote it | +| 3 | When does the final filename appear? | Only after verification: `os.replace(working, target)`. A failed or interrupted run never leaves a file wearing a backup's name | +| 4 | How are failures cleaned up? | `_discard()` removes the partial file, and `BackupError` is raised with what went wrong. The source is untouched | +| 5 | Could an existing good backup be overwritten? | **No.** `_unused_name()` stamps each backup with the time and adds a counter if that name (or its `.partial`) exists | +| 6 | Was the check on the source or the copy? | The copy | + +So the mechanism was already right. The one gap was the **strength** of the check: +`PRAGMA quick_check`, which reads every page and every record but skips the +cross-check between a table and its indexes. + +--- + +## C. `integrity_check` implementation + +One function changed: + +```python +# app/backup.py, _verify() +rows = connection.execute("PRAGMA integrity_check").fetchall() # was quick_check +``` + +- It still runs on the **finished copy**, opened as its own read-only + connection, before the rename. +- A failure still raises `BackupError` naming what was wrong, still discards the + partial file, and still leaves the source and every earlier backup untouched. +- The returned `integrity` field, the filename, the directory, the reported + fields (`filename`, `bytes`, `pages`, `seconds`, `integrity`) and the API are + unchanged. +- The module docstring and `_verify`'s docstring now record why the trade M9 + made (speed over the index cross-check) was not needed at these sizes, with + the measurements in §E. + +**No scheduled backups were added.** There is still no restore endpoint, no +retention policy and no timer: a backup happens when the reader asks for one. + +--- + +## D. Corruption fixture + +`tests/test_v11_d_recovery.py: build_corrupt_copy()`. + +A small database with `t(id, k, filler)` and an index `i_t_k ON t(k)`, 400 rows +with keys `k000000`…`k000399`. `dbstat` names the index's own leaf pages, and one +digit inside one indexed key on the first of them is changed (`k000144` → +`k000944`). Every page stays structurally sound and every record still parses: +what is broken is only the agreement between the index and its table. + +**Proved before it is used as evidence** +(`test_the_fixture_is_the_difference_between_the_two_checks`): + +| Pragma | Result | +| --- | --- | +| `PRAGMA quick_check` | **`ok`** | +| `PRAGMA integrity_check` | **`row 145 missing from index i_t_k`** | + +That is the difference the package rests on: not a database both checks reject, +but one the old check called healthy. + +--- + +## E. Backup performance measurements + +Each pragma was run in its **own fresh process**, alternating, three times, because +a first measurement warms the page cache and a naive ordering makes whichever +check runs second look faster. (It did: an early single-pass measurement showed +`integrity_check` at 1.9 ms against `quick_check` at 65.5 ms, purely from cache +warmth.) + +**The real campaign database** — the M11 100-turn evidence campaign, +`m04-final/campaign.db`, 2,367,488 bytes (2.26 MB): + +| | Run 1 | Run 2 | Run 3 | +| --- | --- | --- | --- | +| `quick_check` | 5.0 ms | 3.4 ms | 7.6 ms | +| `integrity_check` | 3.5 ms | 5.7 ms | 5.3 ms | + +At this size the two are indistinguishable. A whole backup through +`backup.create()` — copy, full check and rename — took **70.8 ms** and **27.4 ms** +on two runs (578 pages). + +**An index-heavy synthetic database of 105.2 MB** (105,160,704 bytes; 25,674 +pages of 4,096 bytes), shaped to give the cross-check real work: `actions` with +152,000 rows and `memories` with 15,200, under three indexes +(`i_actions_adv_depth`, `i_actions_kind`, `i_memories_adv`): + +| | Run 1 | Run 2 | Run 3 | +| --- | --- | --- | --- | +| `quick_check` | 172.6 ms | 92.4 ms | 97.1 ms | +| `integrity_check` | 227.8 ms | 196.0 ms | 227.4 ms | + +A whole backup of it through `backup.create()`: **0.58 s** for 25,674 pages, +`integrity = ok`. + +**A campaign-shaped database of 117.4 MB** (117,403,648 bytes; 2,200 actions), +built through the application's own models so the product path can run against +it (§E.1): + +| | Run 1 | Run 2 | Run 3 | +| --- | --- | --- | --- | +| `quick_check` | 89.5 ms | 74.2 ms | 64.4 ms | +| `integrity_check` | 90.0 ms | 82.9 ms | 75.9 ms | + +**Reading.** The cost of the cross-check tracks **rows and index entries, not +bytes**. On the campaign schema — 2,200 fat rows — the full check costs about +7 ms more than the quick one at 117 MB. On the index-heavy synthetic — 167,200 +rows under three indexes at a similar size — it costs about 96 ms more. Both sit +inside a backup of about half a second. There is no performance requirement in +this project and WP-D does not invent one; the measurements are here because the +M9 trade was made on a speed argument, and at these sizes that argument does not +hold. + +### E.1 Through the endpoint the button calls — implementation evidence only + +**This does not satisfy acceptance criterion 3, and an earlier draft of this +report wrongly said it did.** The criterion asks for the backup to *complete +through the UI*; a reader does not call an endpoint. What follows is evidence +about the implementation — the procedure, its result and its cost — and it is +kept for that reason. The acceptance evidence is in **§E.2**. + +`POST /api/backups` is the only call the **Back up now** button makes, and it +takes no parameters, so it was driven directly against a **copy** of each +database — the evidence database is evidence and is not written to: + +| Database | Result | +| --- | --- | +| Evidence campaign, 2,367,488 bytes | **201** in 0.117 s wall — `pages: 578`, `seconds: 0.038`, `integrity: ok`; `GET /api/backups` then lists 1 backup | +| Campaign-shaped, 117,403,648 bytes | **201** in 0.536 s wall — `pages: 28,663`, `seconds: 0.452`, `integrity: ok` | + +**Why a second large database exists.** The index-heavy synthetic has no +application schema, and the server runs its migrations at startup, so the app +will not start against it: driving the endpoint there fails in a migration that +renumbers `actions`, before any backup is attempted. It remains valid evidence +for the pragma comparison, which is a property of SQLite and not of this schema, +but criterion 3 needs a database the product can actually open — hence the +campaign-shaped one, which §E.2 then drives through the browser. + +Artefacts stay under `$HOME` (`v11-evidence/wp-d/`), not in the repository. + +--- + +### E.2 Through the UI — the acceptance evidence for criterion 3 + +The reader-facing **Back up now** control, clicked in a real Firefox, on the +production build served by FastAPI: the WP-C path, with no Vite dev server, no +inference, and no sleep used as an assertion. Every wait is on a condition the +page or the filesystem can show. + +`backend/tools/wpd_backup_ui.py`, run as +`python -m tools.wpd_backup_ui --out $HOME/v11-evidence/wp-d/ui --case both`. +**34 checks, 34 passed, 0 failed** — +`$HOME/v11-evidence/wp-d/ui/backup-ui-report.json`. + +The route a reader takes is the route the tool takes: open the application, click +**Settings** in the navigation, open *Back up everything on this machine* (a +`
` that loads what is on disk when it opens), then press the button. + +| | **Case 1 — real campaign database** | **Case 2 — ≥100 MB application database** | +| --- | --- | --- | +| Source | the M11 evidence campaign, copied | the campaign-shaped database of §E, copied | +| **Database size** | **2,367,488 bytes** (2.3 MB) | **117,403,648 bytes** (112.0 MB) | +| Schema | the application's own, `user_version` 94 | the application's own, `user_version` 94 | +| **Browser action** | click **Back up now** (twice — see below) | click **Back up now** | +| **Observable UI result** | toast: *"Backup written: adventure-storyteller-20260915-225000.db (2.3 MB)."*, no error toast, the control returns from *Backing up…*, and the file appears in the panel's list | toast: *"Backup written: adventure-storyteller-20260915-225007.db (112.0 MB)."*, no error toast, control returns, file listed | +| **Backup path** | `…/ui/real/data/backups/adventure-storyteller-20260915-225000.db` | `…/ui/large/data/backups/adventure-storyteller-20260915-225007.db` | +| **Backup file size** | **2,367,488 bytes** | **117,403,648 bytes** | +| **`integrity_check` result** | **`ok`** | **`ok`** | +| **Integrity-check elapsed** | **5.8 ms** | **78.5 ms** | +| **Total elapsed backup** | **0.108 s** (click → the UI says it is done) | **0.719 s** | + +"Total elapsed" is measured from the click to the rendered result, so it is what +the reader waits, not what the server reports. The `integrity_check` above is an +independent second opinion, run here on the artefact the UI produced — the +application had already verified the copy before keeping it. + +**The size the UI reports is the size on disk.** "112.0 MB" in the toast is +117,403,648 bytes shown in the page's own units; the file matches the source +database byte for byte. + +**Existing-backup protection, established through the UI itself** (Case 1's +second press, rather than a backup planted by a library call): + +| Claim | Result | +| --- | --- | +| A second press writes a *different* file — `…-225000-2.db`, the counter `_unused_name` adds when two backups land in the same second | PASS | +| The first backup still exists afterwards | PASS | +| And is byte-for-byte what it was — 2,367,488 bytes, `sha256:7d45555566…` | PASS | +| And still passes `integrity_check` | PASS | + +No corruption was fabricated through the browser; rejection behaviour is already +proved deterministically in §D. + +**Supplemental endpoint evidence.** The servers' own logs corroborate that the +button drove each backup: Case 1 logged two `POST /api/backups → 201 Created`, +each followed by `GET /api/backups → 200 OK` as the panel reloaded its list; +Case 2 logged one of each. Screenshots and logs are beside the report JSON. + +**A harness defect this run found, in my own check and not in the product.** The +first pass reported six failures. The toast renders a decorative mark before its +message — `…` — +so the button's `textContent` begins with `❖`, and my assertion matched +`textContent.startswith("Backup written:")`. In that same run the UI had shown a +non-error toast naming the file, the file was on disk, its name was in the +panel's list, and the copy verified: the product was right and the assertion was +reading the mark. The check now reads the message span. **No product code was +changed** — the expected change for this work was none, and none was needed. + +--- + +## F. Existing export/import limit behaviour + +| | | +| --- | --- | +| Import ceiling | `limits.MAX_IMPORT_BODY_BYTES` = 20 MB, **unchanged by WP-D** | +| Where it is enforced | `BodySizeLimitMiddleware`, on the declared `Content-Length`, before the body is read | +| Refusal | HTTP **413**, "Request too large (limit 20 MB)." — it already named the limit | +| Export before WP-D | `GET /adventures/{id}/export` returned the bundle. Nothing compared its size with the ceiling, so a campaign could be exported and then refused by its own importer | + +--- + +## G. Oversized-export warning implementation + +**Where the answer comes from.** `limits.import_limit_label()` and +`limits.oversized_export_warning(size)` derive the sentence from +`MAX_IMPORT_BODY_BYTES`. A test changes the constant and asserts the wording +follows, so no number is written twice. + +> This export is larger than this version's 20 MB import limit (21,230,000 bytes). +> The file was exported successfully, but this version cannot import it. + +**Where it travels: headers, not the body.** The export response *is* the bundle — +the browser saves exactly those bytes as the file — so a warning inside it would +become part of a portable story file and of every checksum taken over one. The +route returns the same body with: + +```text +X-Export-Bytes the serialised size +X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES +X-Importable-By-This-Version true / false +X-Export-Warning only when false +``` + +**Which size is measured.** The compact serialisation (`separators=(",", ":")`, +`ensure_ascii=False`), which is what this response sends *and* what the browser +POSTs back on import — the bytes `BodySizeLimitMiddleware` weighs. The +pretty-printed file the reader downloads is larger and is not what import reads. +The route serialises once and returns those bytes, so `X-Export-Bytes` is the +length of the body actually sent. + +**The bundle is unchanged.** No key was added to it (§I), and the format stays +`ai-dnd-adventure-v3`. + +--- + +## H. Oversized export evidence + +`test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back` +writes real rows into a campaign until its bundle genuinely exceeds the ceiling — +nothing is mocked, and the export serialises all of it. + +| Claim | Result | +| --- | --- | +| The response body is larger than 20 MB | PASS | +| It parses, and is `ai-dnd-adventure-v3` with its actions | PASS | +| `X-Importable-By-This-Version: false` | PASS | +| `X-Export-Warning` names the limit ("20 MB"), says the export succeeded, and says this version cannot import it | PASS | +| `X-Export-Bytes` equals the body length | PASS | +| The same bundle is refused by import, with the limit named (§J) | PASS | +| A normal campaign's export carries **no** warning, and `X-Importable-By-This-Version: true` | PASS | +| The warning follows the constant (changed to 50 MB in a test: the sentence says 50 MB, and no longer says 20 MB) | PASS | + +--- + +## I. Normal-export compatibility + +The same campaign database (`m04-final/campaign.db`) exported through a worktree +at `59b5ebc` (before WP-D) and through the WP-D tree: + +| | | +| --- | --- | +| Before | 2,699,076 bytes | +| After | 2,699,076 bytes | +| Comparison | **identical** — the parsed documents compare equal, key for key | + +No timestamp allowance was needed: this campaign's bundle carries no field that +moves between exports. The format is `ai-dnd-adventure-v3` in both. + +`test_the_bundle_itself_never_carries_the_warning` additionally asserts that no +key of an oversized bundle mentions the warning, the limit or importability. + +--- + +## J. Import-refusal behaviour + +Unchanged in behaviour, and already naming the limit: + +| | | +| --- | --- | +| Status | **413**, from the middleware, on `Content-Length`, before the body is parsed | +| Message | "Request too large (limit 20 MB)." | +| WP-D test | `test_that_same_bundle_is_refused_by_import_naming_the_limit` posts the *actual oversized export* and asserts 413 and that the detail contains `limits.import_limit_label()` | +| Not relaxed | the same boundary, the same status, the same middleware. `test_a_bundle_under_the_limit_still_imports` keeps the ordinary path honest (201) | + +No wording correction was needed. + +--- + +## K. Frontend behaviour + +`api.exportAdventure` now returns the bundle **and** what the server said about +it (`exportBytes`, `importLimitBytes`, `importable`, `warning`), read from the +headers. + +**The no-header case.** `importable` is `resp.headers.get(…) !== 'false'`, so +only the literal string `false` is read as a refusal: an older server that sends +no headers — or a header that arrives malformed — yields `importable: true` and +`warning: null`, and the page says nothing it was not told. The two size fields +fall back to `null` unless they parse as a positive number. + +This is asserted directly: three tests stub `fetch` with real response headers +and check what `api.exportAdventure` makes of them — a warning with the sizes it +named, an ordinary export marked importable, and an older server sending no +headers at all. They exist because the page-level tests below mock +`api.exportAdventure` itself and so cannot see a header name, which meant a typo +on the frontend side would have left the whole suite green (§O.7). + +Both reader-facing entry points keep delivering the file first and then report: + +| Entry point | Under the limit | Over the limit | +| --- | --- | --- | +| Campaign library, **Export** | file written; "Campaign exported." | file written; the warning, as an error-styled toast | +| Campaign settings, **Export campaign** | file written; "Campaign exported." | file written; the warning | + +No new modal, no redesign: the existing toast carries it. + +**Tests** (`pages/exportHonesty.test.jsx`, **7**): three read real response +headers through a stubbed `fetch` (above), and four drive both entry points — +the file is delivered in every case; the warning is shown when the server sends +one, naming 20 MB and saying the export succeeded; the ordinary confirmation is +shown when it does not. + +--- + +## L. Offline regression + +`tools/m11_offline.py`, against the production build, in the offline container: + +**23 checks, 23 passed, 0 failed** — +`$HOME/v11-evidence/wp-d/offline/offline-report.json`. + +The two this package could have broken are in it and passed: + +| Check | Result | +| --- | --- | +| a campaign exports offline | ok | +| and imports offline, with its state | ok | +| no secret is present in the export | ok | +| a local file imports offline | ok | + +The export route now sets headers and serialises compactly; the offline +container still exports a campaign and imports it back with its state, so +neither the round trip nor the secret-scrubbing changed. + +## M. Full regression + +| Suite | Result | +| --- | --- | +| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,712 passed, 17 skipped, 0 failed, 0 xfailed** (936.6 s) | +| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed | +| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings | +| **Production build** (`npm run build`) | succeeded | +| **Offline regression** | 23 passed, 0 failed (§L) | + +**The backend count reconciles exactly.** WP-B's closing tree was 1,693; WP-C +added 7 (`test_v11_c_browser_helpers.py`) and did not rerun the suite because no +application code changed; WP-D adds the 12 in `test_v11_d_recovery.py`. +1,693 + 7 + 12 = **1,712**. + +**The 17 skips are named, not assumed.** Run with `-rs`, every one is an +environment-gated real-model test, and none is new: + +| File | Skipped | Gate | +| --- | --- | --- | +| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) | +| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` | +| `test_narrative_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` | +| `test_m11_real_window.py` | 3 | two the same; one `AIDND_TEST_WIDE_MODEL` | +| `test_provider_wiring.py` | 1 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` | + +That is the same 17 B.1 and B.2 recorded. WP-D used no model and added no skip. + +**The plan's named regression requirements**, run individually rather than +assumed to be inside the total: + +| Requirement | Result | +| --- | --- | +| I01-I07, L01-L04 | **25 passed**, 1,704 deselected | +| The backup case in `test_m11_migration.py` | **1 passed** | +| `backup.test.jsx` | **7 passed** | +| The offline container's export and import | passed (§L) | + +**Frontend arithmetic.** WP-C's baseline was 168. WP-D adds 4 page-level export +tests and 3 header-parsing tests: 168 + 7 = **175**. The 15 lint warnings are +the same pre-existing `only-export-components` and unused-import kind recorded +at WP-C, and **none is in a file WP-D changed**. + +## N. Compatibility + +| Area | Effect | +| --- | --- | +| Database schema | unchanged; no migration | +| Bundle format | `ai-dnd-adventure-v3`, unchanged | +| Normal bundle contents | byte-identical (§I) | +| Existing backups | still valid files with the same names; only the check that admits a new one is stricter | +| v1.0.0 databases | open unchanged (no schema or data path changed) | +| Import limit | unchanged at 20 MB | +| Scheduled backups | none, as before | +| History, branches, Save Points, state, memory, knowledge | untouched | +| Endpoint policy, local-only operation | untouched | + +## O. Residual risks + +1. **Closed: the backup is now driven through the UI.** This risk previously + read "the browser click was not driven for the backup", and the owner + correctly refused criterion 3 on endpoint evidence. §E.2 drives the real + **Back up now** control in Firefox on both databases — 34 checks, 0 failed — + including the ≥100 MB case, which does mean the harness starts a server + against a 117 MB database and takes 0.719 s to do it. What remains is + ordinary coverage scope: this runs as its own tool rather than inside the + release harness, so it is not part of the 101-check run in the WP-E report. +2. **The 20 MB ceiling is unchanged.** A campaign past it still cannot be + imported by this version. WP-D makes that audible at export; raising the + limit or streaming import stays v1.2, as the plan assigns it. +3. **The warning is about *this* version.** It says what this build's importer + will accept. A future build with a higher ceiling could import a file this + one warned about, and the wording ("this version") is chosen so that stays + true rather than becoming a lie. +4. **`integrity_check` is not a guarantee of recoverability.** It proves the + copy's pages, records and indexes agree. A database that was already + logically wrong when it was copied is copied faithfully and passes. The + package makes the check honest, not omniscient. +5. **No restore path.** There is still no restore button and no scheduled + backup — both explicitly out of scope. A reader who needs a backup restores + it by replacing the file themselves, as before. +6. **The header channel depends on the browser reaching headers.** Both export + call sites read them through `fetch`, so a proxy that stripped `X-` headers + would silently return to v1 behaviour: the file still arrives, the warning + does not. Local-only operation makes that unlikely, and the failure is the + old behaviour rather than a wrong claim. +7. **Header parsing was untested — now closed.** Writing this section exposed + it: the page-level tests mock `api.exportAdventure` and the backend tests + assert what the server sends, so nothing read an actual header name, and a + frontend-side typo would have left every test green and the reader silently + uninformed. Three tests that stub `fetch` with real headers now cover it + (§K). What remains is the ordinary version of this risk: the two sides agree + by matching string literals in two files, and only a browser-level test would + catch a mismatch introduced in both at once. + +## P. Final decision + +**Against the plan's acceptance criteria** (§ WP-D, *Recovery honesty*): + +| # | Criterion | Verdict | +| --- | --- | --- | +| 1 | A healthy database's backup reports `integrity_check` ok and is kept | **PASS** — `test_a_healthy_backup_passes_the_full_check_and_is_kept`, and both endpoint runs returned `integrity: ok` (§E.1) | +| 2 | A copy `integrity_check` rejects and `quick_check` does not is rejected and not kept; existing backups still never overwritten | **PASS** — the fixture is proved to be exactly that difference (§D), and three tests cover rejection, the untouched earlier backup, and the unchanged file semantics | +| 3 | The time is recorded on the evidence database and on a synthetic ≥100 MB, and the backup completes through the UI on both | **PASS** — pragma timings in §E on three databases, and the backup completes **through the UI** on both in §E.2: 0.108 s with `integrity_check` `ok` in 5.8 ms on the 2.3 MB campaign, 0.719 s with `ok` in 78.5 ms on the 117 MB application database. An earlier draft claimed this criterion on endpoint evidence (§E.1); that claim was wrong and is corrected | +| 4 | An over-20 MB export succeeds, delivers the file, and warns naming the limit, in the API response and a component test; under the limit, no warning | **PASS** — §H, and 7 frontend tests (§K) | +| 5 | Importing that bundle is refused with a message naming the limit | **PASS** — §J, asserted on the actual oversized export | +| 6 | A normal export is byte-identical before and after, apart from timestamps | **PASS** — fully identical, no timestamp allowance needed (§I) | + +**What I corrected rather than reported around.** Two claims of my own failed +checking and were fixed, not softened: §E originally rested criterion 3 on +`backup.create()` calls, and driving the real endpoint showed the index-heavy +synthetic cannot serve that criterion at all (the app's startup migrations abort +against a schema-less database) — hence the campaign-shaped database in §E.1. +And §K asserted the no-header fallback from the code alone, which exposed that +**no test read any header name**; three tests now do (§O.7). + +**Not done, and not claimed:** the import ceiling is unchanged at 20 MB, by the +plan's own assignment of the raise to v1.2. The UI backup runs as its own tool +rather than inside the release harness (§O.1). + +**No product code changed for the UI verification.** The run exposed one defect, +and it was in my own assertion, not in the application (§E.2). Backend +**1,723 passed / 17 skipped / 0 failed** and offline **23/23** therefore stand +as prior evidence, unchanged and not re-run; the targeted backup and browser +checks were re-run instead. + +```text +BACKUP INTEGRITY: PASS +BACKUP THROUGH UI — REAL DB: PASS +BACKUP THROUGH UI — >=100 MB DB: PASS +EXPORT HONESTY: PASS + +WP-D OVERALL: +PASS +``` + +All WP-D changes are **staged and uncommitted**. No commit, no push, no tag. +WP-E has not begun. diff --git a/planning/reports/v1.1/V1.1-WP-E-REPORT.md b/planning/reports/v1.1/V1.1-WP-E-REPORT.md new file mode 100644 index 0000000..18d4a34 --- /dev/null +++ b/planning/reports/v1.1/V1.1-WP-E-REPORT.md @@ -0,0 +1,374 @@ +# v1.1 WP-E — Control-Boundary Contrast + +**Status:** COMPLETE — **PASS**, with owner approval of the screenshots +outstanding. The decision, and what was deliberately not claimed, is in §O. + +--- + +## A. Repository baseline + +| | | +| --- | --- | +| Branch | `v1.1-development` | +| HEAD | `59b5ebc` — *v1.1 WP-C: browser release coverage*, signed by the owner | +| Working tree at start | WP-D staged (10 files), nothing committed | +| Criterion | WCAG 2.1 **1.4.11 Non-text Contrast**, 3:1, for control boundaries; **1.4.3** 4.5:1 for body text, unchanged | + +--- + +## B. What v1.0.0 actually did + +`tools/contrast_audit.py` measured control boundaries, printed that two of them +were below 3:1, and **exited 0**. Its own comment argued the position: + +> in this design a control is identified by its *label*, which is measured above +> and passes, not by its edge. So a boundary below 3:1 is reported with its +> number and does not fail the run. + +So the audit was a report, not a gate: no palette change could ever fail it on a +boundary. The two numbers it printed were **1.33:1** (`--border` on +`--bg-panel`) and **1.75:1** (`--border-bright`), against a floor of 3.0. + +**WP-E overturns that argument.** 1.4.11 covers the visual information needed to +identify a component *and its boundary*; a reader who cannot see where a text box +ends cannot see that there is a text box to type into, label or no label. The +tokens were raised rather than the criterion re-argued. + +--- + +## C. Token inventory + +| Token | v1.0.0 | v1.1 | Why | +| --- | --- | --- | --- | +| `--border` | `#2b2b3d` | **`#676792`** | every control's resting edge | +| `--border-bright` | `#3d3d55` | **`#7a7aaa`** | hover edges, the composer's resting edge, the open panel tab | +| `--bg-panel`, `--bg-input`, `--bg`, `--text`, `--text-dim`, `--accent*`, `--danger`, `--warning`, `--player`, `--chart-*` | — | **unchanged** | WP-E is a boundary package; no text or accent colour moved | + +**The floor is taken against `--bg-input`, not `--bg-panel`.** Inputs and buttons +are drawn on `--bg-input` (`styles/forms.css`), which is lighter than +`--bg-panel` and therefore the harder case. The audit had been checking only +`--bg-panel`, so a token could have passed the audit while the real control +failed. Measured on the new values: + +| | vs `--bg-input` | vs `--bg-panel` | vs `--bg` | +| --- | --- | --- | --- | +| `--border` | **3.21:1** | 3.44:1 | 3.70:1 | +| `--border-bright` | **4.24:1** | 4.55:1 | 4.88:1 | + +Two properties were preserved deliberately: the rest→hover step is the same size +as before (1.318 → 1.320), so hover still reads as a change rather than a jump; +and `--border-bright` stays *below* body text against the same panel (2.98:1 +between them), so no edge outshines the words inside it. + +--- + +## D. Component inventory + +Where these tokens are actually drawn, from the stylesheets: + +| Control | Rule | Rest | Hover / active | Focus | +| --- | --- | --- | --- | --- | +| Story composer | `.input-bar` (story.css) | `--border-bright` on `--bg-panel` | — | `--accent-dim` + `--accent-glow` ring | +| Story controls | `.story-controls button` | `--border` on `--bg-panel` | `--border-bright` | (M11 focus check) | +| Fields and buttons | `forms.css` | `--border` on `--bg-input` | `--accent-dim` | `--accent-dim` + ring | +| Panel tabs | `.panel-tabs button` | **`transparent`** | `--border` on hover, `--border-bright` when active | — | +| Top navigation | `.topnav` | `--border` bottom edge on **`--bg-panel-glass`** | — | — | +| Scrollbar thumb | `base.css` | `--border-bright` as a *fill* on `--bg` | `--accent-dim` | — | + +Two of these cannot be answered by token arithmetic at all, and both are +measured in the browser instead (§G): the nav sits on a translucent panel, and +the panel tab's edge is `transparent` until the panel is open. + +--- + +## E. The audit is now a gate + +`tools/contrast_audit.py`: + +1. **Boundary pairs fail.** `text` and `boundary` rows are both pass/fail; the + advisory branch is gone. The run returns 1 if either kind falls short. +2. **Eight boundary pairs replace two.** Each border is checked against every + background it is drawn on — `--bg-input`, `--bg-panel` and `--bg` — plus the + focused edge (`--accent-dim`) on both panel and field backgrounds. +3. **The verdict is taken on the number that is printed** (rounded to two + decimals), so a pair shown as `3.00:1` is not failed for arithmetic the + reader cannot see. +4. The comment block that argued the old position is replaced by one recording + what changed and why, including why `--bg-panel-glass` is not in the list. + +## F. Gate tests + +`backend/tests/test_v11_e_contrast.py` — **11 passed**. The threshold is +exercised from both sides, on real token files: + +| Test | Result | +| --- | --- | +| A boundary at **2.99:1** against `--bg-input` fails the run (exit 1) | PASS | +| A boundary at **3.00:1** passes (exit 0) | PASS | +| The **v1.0.0 value** `#2b2b3d` fails, at 1.24:1 against `--bg-input` | PASS | +| A dimmed `--text-dim` still fails as a *text* pair | PASS | +| A renamed token is a failure, not a silent skip | PASS | +| The shipped palette passes both criteria | PASS | +| Every boundary pair is measured against the background it is drawn on | PASS | +| The M11 text baselines are unchanged: 14.57 / 13.57 / 5.48 / 5.88 | PASS | +| The hover edge stays brighter than the resting edge | PASS | +| No boundary becomes as loud as body text | PASS | +| The WCAG ratio formula is anchored on known values (21:1, 1:1, symmetry) | PASS | + +The 2.99 and 3.00 values are worth noting: **both clear 3:1 against +`--bg-panel`** (3.21 and 3.22). They decide the gate only because the floor is +now taken against the background the control is really on — so these two tests +also prove §C's change is doing work. + +--- + +## G. Browser measurement + +`tools/m11_browser.py` gains a third suite, **WP-E**, counted separately from +M11's 38 and WP-C's 53. It measures the *rendered* edge — `borderColor` from +`getComputedStyle` — against what is actually behind it, with every translucent +layer composited bottom-up. + +A boundary is measured against **both** adjacent colours (the control's own fill +inside it, the background outside it) and passes on the better of the two: an +edge that matches its fill but contrasts with the page is still a visible +outline. What 1.4.11 asks is that the component's extent be perceivable. + +Two harness capabilities were added for this (`tools/m11_webdriver.py`): + +- **`hover()`** moves a real pointer through the WebDriver Actions API. + Dispatching a `mouseover` event from JavaScript does *not* trigger CSS + `:hover`, so a synthetic event would have re-measured the resting edge and + reported it as the hover edge. +- **`screenshot()`** writes the viewport as a PNG, for the before/after evidence. + +**A defect this found in my own first measurement.** The first run reported the +hover edge as `rgb(114, 114, 160)` and the focused edge as `rgb(144, 120, 81)` — +neither of which is any token. Both controls carry `transition: border-color +0.15s`, so the measurement was taken mid-animation, on a colour no state +actually has. `_settled()` now polls until the computed edge colour is the same +on two consecutive reads before measuring (polled, not slept, per this harness's +own rule). After the fix the same edges read exactly `rgb(122, 122, 170)` +(`--border-bright`) and `rgb(150, 119, 58)` (`--accent-dim`). + +--- + +## H. Before and after, measured in the browser + +Both passes were taken the same way — `--only boundaries --no-narrator`, the +production build — with only `tokens.css` differing. Evidence under +`$HOME/v11-evidence/wp-e/before/` and `.../after/`. + +| Control (state) | Before | After | Floor | +| --- | --- | --- | --- | +| Story composer — resting edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 | +| Story control — resting edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 | +| Open panel tab — resting edge | **1.75:1** FAIL | **4.55:1** pass | 3.0 | +| Story control — hover edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 | +| Top navigation — translucent edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 | +| Story composer — focused edge | 4.70:1 pass | 4.70:1 pass | 3.0 | + +**Suite result: before 5 passed / 5 failed; after 10 passed / 0 failed / 0 +skipped.** + +Two things this table says that a summary would blur: + +- **Focus was never the defect.** The focused edge (`--accent-dim`) already + cleared 3:1 in v1.0.0 at 4.70:1, and WP-E did not change it. What failed was + rest and hover — the states a reader spends all their time in. +- **The translucent edge is real evidence.** The nav's background composited to + `rgb(17, 17, 29)` — `--bg-panel-glass` (rgba 19,19,32 @ 0.82) over + `rgb(10, 10, 15)` — not a fallback. That is the case token arithmetic cannot + reach, and it moved from 1.43:1 to 3.70:1. + +## I. Screenshots + +| File | | +| --- | --- | +| `before/control-boundaries.png` | 131,175 bytes, 1366×682 | +| `before/control-boundaries-nav.png` | 63,376 bytes, 1366×682 | +| `after/control-boundaries.png` | 132,430 bytes, 1366×682 | +| `after/control-boundaries-nav.png` | 63,541 bytes, 1366×682 | + +All four are PNG, 1366×682, taken on the production build through the same +harness path, differing only in `tokens.css`. The play-page pair shows the +composer, the story controls and the open panel tab; the nav pair shows the +translucent top edge on the library route. + +## J. A finding this package created and fixed + +Raising `--border-bright` broke something that had nothing to do with control +boundaries. `.slice-7` in the context inspector's token breakdown was painted +with `var(--border-bright)`, so it followed the token to `#7a7aaa` — an OKLab ΔE +of **0.035** from `.slice-6` (`#7c86b8`), making two neighbouring chart slices +effectively the same colour. The other slices sit **0.100–0.119** from their +nearest neighbour. + +`.slice-7` is now pinned to `#3d3d55`, the literal value it already rendered, so +its appearance is unchanged from v1.0.0 and its separation (ΔE **0.251**) is the +widest in the set. A chart fill and a control edge have different jobs and should +not share a token. + +Reassigning it to a fresh hue was considered and rejected on evidence: inside the +palette's own chroma (0.045–0.120) and lightness (0.586–0.804) bands, the only +hues clearing the set's 0.100 separation floor are pinks near 14°, which is +`--danger`'s territory. Painting an ordinary prompt section in the colour this +application reserves for failure would trade an accessibility fix for a semantic +lie. + +*(A first attempt, `#5d7f9e`, was rejected by the same measurement at ΔE 0.061 — +below every real slice. It is recorded here because it was written into the file +before it was measured.)* + +## K. Text contrast regression + +Unchanged, and asserted so in §F: **14.57:1** body text on the page, **13.57:1** +in a panel, **5.48:1** secondary text in a panel, **5.88:1** on the page. Every +text pair still clears 1.4.3, and no text token was touched. + +## L. Browser release regression + +The full harness, all three suites, against a real narrator on the production +build. Evidence: `$HOME/v11-evidence/wp-e/release/browser-report.json`. + +| | | +| --- | --- | +| Kind | **`release regression`** — not `partial`, not `development (only …)` | +| Narrator | `qwen2.5:3b-instruct`, the reference model, over plain HTTP on the LAN GPU host | +| Browser | Firefox 155.0.1, geckodriver 0.37.1 | +| Served | FastAPI on loopback, the **built** SPA (`dist` 2026-09-15T21:08:52) | +| Duration | 71 s | +| Turns played | **8**, of which `turns_not_clean` **0** and `protocol_shapes_in_narration` **0** | + +**Counted per suite, as the brief requires:** + +| Suite | Passed | Failed | Skipped | +| --- | --- | --- | --- | +| **M11** (the v1 release regression) | **38** | **0** | **0** | +| **WP-C** (browser release coverage) | **53** | **0** | **0** | +| **WP-E** (control boundaries) | **10** | **0** | **0** | +| **Total** | **101** | **0** | **0** | + +M11's 38 and WP-C's 53 are unchanged in count and in name: WP-E added a suite +beside them rather than altering either. The `kind` field is quoted above +because a fast run invites the question — 71 s for 101 checks including 8 +narrated turns is the GPU host being quick with a 3B model, and the 8 recorded +turns with no unclean accounting are what rule out narration having been +skipped. + +The WP-E rows in this run are the same ten as the standalone capture in §H, +re-measured with a narrator present and a full story on the page. + +## M. Full regression + +| Suite | Result | +| --- | --- | +| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,054.7 s) | +| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed | +| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings | +| **Production build** (`npm run build`) | succeeded | +| **Browser harness** (M11 + WP-C + WP-E) | **101 passed, 0 failed, 0 skipped** (§L) | +| **Contrast audit** (`python tools/contrast_audit.py`) | **exit 0** — every text pair and every boundary pair passes | + +**The backend count reconciles exactly.** WP-D's tree was 1,712; WP-E adds the +11 in `test_v11_e_contrast.py`. 1,712 + 11 = **1,723**. The 17 skips are the same +environment-gated real-model tests recorded since B.1 — WP-E used no model and +added no skip. + +**The plan's regression requirements for WP-E** were the frontend suite and +lint, and the harness's accessibility checks. All three pass: 175 and exit 0 +above, and M11's `A11y` rows — accessible names, visible keyboard focus, no +positive tabindex, nothing revealed only on hover, and the four rendered text +contrasts — are inside the 38/38 in §L. + +**Lint detail.** The 15 warnings are the same pre-existing +`only-export-components` and unused-import kind recorded at WP-C and WP-D, and +**none is in a file WP-E changed**. `tokens.css` and `context.css` are not +flagged. + +## N. Residual risks + +1. **The screenshots are unapproved.** The measurements say every boundary now + clears 3:1; whether the result *looks* right in this design is a judgment + the numbers cannot make. Recorded as PENDING in §O, not assumed. +2. **Five controls are measured in the browser; the rest inherit.** The audit + checks token pairs, and the harness measures the composer, a story control, + the open panel tab, the nav edge and the focused composer. Every other + bordered surface — modals, cards, the knowledge and context panels — draws + the same two tokens, so it moves with them, but none is individually + measured. A component that overrides a border with a literal colour would + not be caught by either check. +3. **`--bg-panel-glass` is outside the token audit by nature.** It is rgba over + a gradient, so no token pair can express it; its edge is covered only by the + browser measurement, which runs in the harness rather than in CI. +4. **Disabled controls are deliberately not measured.** `.story-controls + button:disabled` carries `opacity: 0.35`, so a disabled control's rendered + edge is dimmer than any value here. WCAG 1.4.11 exempts inactive components, + and the harness selects `:not(:disabled)` on purpose — stated so that the + exclusion is visible rather than looking like an oversight. +5. **The chart set was re-checked only where WP-E disturbed it.** `.slice-7`'s + separation and colour-blind distance were measured against the other seven + (§J); the set as a whole was not re-audited, which is outside this package. +6. **`a11y.test.jsx` was not extended**, though the plan listed it as likely + affected. It asserts structure and names, not colours, and adding colour + assertions in jsdom would test the stylesheet's text rather than a rendered + result. Boundary contrast is asserted instead where it can be measured: the + gate tests (§F) and the browser (§G). + +**One risk that turned out not to exist.** §G's rule — a boundary passes on the +better of its two adjacent colours — was written to avoid failing an edge that +contrasts with the page but matches its own fill. In the event it never did any +work: every measured boundary clears 3:1 against **both** neighbours (composer +4.55/4.88, story control 3.44/3.70, panel tab 4.24/4.55, hover 4.55/4.88, focus +4.38/4.70, nav 3.51/3.70). The stricter reading would have produced the same +verdict on every row. + +## O. Final decision + +**Against the plan's acceptance criteria** (§ WP-E, *Control-boundary contrast*): + +| # | Criterion | Verdict | +| --- | --- | --- | +| 1 | `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and 0 on the package tree | **PASS** — 2.99:1 exits 1 and 3.00:1 exits 0 (§F); the v1.0.0 value exits 1; the package tree exits 0 | +| 2 | Every text pair still clears 4.5:1, and the rendered text contrasts do not fall below v1's 14.57 / 5.48 / 13.57 / 5.88 | **PASS** — asserted as exact baselines in §F and re-measured on the rendered page in §L's M11 rows. No text token changed | +| 3 | The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1; the focus indicator is still visible | **PASS** — composer 4.88:1, story control 3.70:1 at rest and 4.88:1 on hover (§H), and M11's visible-focus check passes in §L | +| 4 | The owner approves the before-and-after screenshots, and the report records the approval | **PENDING** — not a criterion I can satisfy. The four PNGs are in §I | + +**Where I exceeded the criterion, said plainly.** The plan asks for 3:1 "against +their panel" and names two controls. This package measures against +**`--bg-input`** as well — the lighter background inputs and buttons are really +drawn on, and the one that decides the gate — and adds the open panel tab, the +translucent navigation edge and the focused state. The stricter floor is the +reason the two threshold tests in §F are decided by `--bg-input` rather than +`--bg-panel`, where both would have passed. + +**What I got wrong and corrected.** Three of my own claims failed checking and +were fixed rather than softened: the first hover and focus measurements were +taken mid-transition and reported colours no state has (§G); two contrast +figures were written into `context.css` before being measured, and were wrong +(§J); and my first gate check reported the current tokens as failing because my +harness crashed on a shallow path, not because of any contrast. + +**Not done, and not claimed:** owner approval (criterion 4), individual +measurement of every bordered component (§N.2), and any palette work beyond the +two boundary tokens and the one chart fill that borrowed from them. + +```text +WP-E CONTRAST GATE: PASS +WP-E BOUNDARY MEASUREMENT: PASS + +WP-E OVERALL: +PASS, pending owner approval of the screenshots (criterion 4) +``` + +All WP-E changes are **staged and uncommitted**. No commit, no push, no tag. +v1.1 release validation has not begun. + +```text +OWNER SCREENSHOT APPROVAL: PENDING +``` + +The before/after screenshots in §I are the evidence for a change a reader judges +by looking at it. The measurements say every boundary now clears 3:1; whether +the result looks right in this design is the owner's call, and it is recorded as +pending rather than assumed.