diff --git a/DEVELOPMENT.md b/DEVELOPMENT.md
index 8c71ca7..50a1a5d 100644
--- a/DEVELOPMENT.md
+++ b/DEVELOPMENT.md
@@ -548,6 +548,40 @@ together.
| Taken from | Export, on a campaign | Settings → *Back up everything on this machine* |
| Restored by | Import campaign, on the library screen | replacing the database file, below |
+### How large an export can get
+
+The importer accepts a request body up to **20 MB**
+(`backend/app/limits.py`, `MAX_IMPORT_BODY_BYTES`), and v1.1 does not change it.
+What that means for a campaign, measured rather than guessed:
+
+- the M11 evidence campaign came to roughly **13 kB per action** in its bundle;
+- M9's conservative estimate from that figure is about **279 turns** before a
+ bundle approaches the limit.
+
+Both are measurements of particular campaigns, **not a turn limit**. What a
+campaign actually weighs depends on how long its turns are, how much imported
+knowledge travels with it, and how many attempts each turn kept. A campaign of
+400 short turns can be well inside the limit; one of 200 long ones with a large
+library may not be.
+
+**v1.1 (WP-D) makes the individual case visible.** Every export reports its own
+serialised size and whether this version could import it back:
+
+```text
+X-Export-Bytes the bundle's size, as the importer would weigh it
+X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
+X-Importable-By-This-Version true / false
+X-Export-Warning present only when it is false
+```
+
+The export always succeeds and the file is always delivered — it is complete and
+undamaged; what it exceeds is this version's import ceiling. Both Export
+controls show the warning when there is one. The size compared is the compact
+serialisation the browser would POST back, which is smaller than the
+pretty-printed file on disk.
+
+Raising the limit, or streaming an import past it, is deferred to v1.2.
+
### Exporting and importing a campaign
Export is on each campaign in the library, and in the campaign's own Settings
diff --git a/backend/app/backup.py b/backend/app/backup.py
index 13edb3f..c01502a 100644
--- a/backend/app/backup.py
+++ b/backend/app/backup.py
@@ -39,8 +39,9 @@ turn is blocked.
never leaves a half-written file wearing a backup's name. `os.replace` is
atomic on the same filesystem, which is why the temporary sits in the
destination's own directory rather than in `/tmp`.
-3. `PRAGMA quick_check` runs against the finished copy, opened as its own
- database, before it is renamed. A backup nobody verified is a belief.
+3. `PRAGMA integrity_check` runs against the finished copy, opened as its own
+ database, before it is renamed. A backup nobody verified is a belief. v1.1
+ WP-D made this the full check rather than `quick_check`; see `_verify`.
4. An existing file is never overwritten. Each run writes a new name stamped
with the time, so yesterday's backup survives today's mistake — which is most
of what a backup is for.
@@ -189,18 +190,28 @@ def _copy(source_path: Path, working: Path) -> int:
def _verify(working: Path) -> str:
- """Runs `PRAGMA quick_check` against the finished copy.
+ """Runs `PRAGMA integrity_check` against the finished copy.
Opened as its own connection, so what is checked is the file on disk rather
- than any page cache the copy left behind. `quick_check` rather than
- `integrity_check` because it does the structural work — every page reachable,
- every record readable — without the full index cross-check, which on a large
- database is minutes rather than moments. A backup nobody verified is a
- belief; a backup verified slowly enough that nobody takes one is worse.
+ than any page cache the copy left behind.
+
+ **v1.1 WP-D: the full check, not `quick_check`.** M9 chose `quick_check` for
+ its speed, on the argument that a backup verified slowly enough that nobody
+ takes one is worse than a fast one. The measurements say the trade was not
+ needed here: `quick_check` omits the cross-check between a table and its
+ indexes, and that is a real class of damage it reports as `ok`. A copy whose
+ index disagrees with its table restores into a database that answers queries
+ with rows that are not there — the failure a backup exists to prevent.
+
+ The cost is small at the sizes this application produces: on the 100-turn
+ evidence campaign both checks are a few milliseconds, and on a synthetic
+ database two orders of magnitude larger the difference is still short of a
+ second (WP-D report §E). A backup nobody verified is a belief; this is the
+ check that makes it a fact.
"""
connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True)
try:
- rows = connection.execute("PRAGMA quick_check").fetchall()
+ rows = connection.execute("PRAGMA integrity_check").fetchall()
finally:
connection.close()
result = ", ".join(str(row[0]) for row in rows) if rows else "no result"
diff --git a/backend/app/limits.py b/backend/app/limits.py
index 16a0c24..bacd973 100644
--- a/backend/app/limits.py
+++ b/backend/app/limits.py
@@ -129,6 +129,34 @@ MAX_BODY_BYTES = 2 * 1024 * 1024
MAX_IMPORT_BODY_BYTES = 20 * 1024 * 1024
+def import_limit_label(limit: int | None = None) -> str:
+ """The import ceiling as a reader would say it, e.g. "20 MB".
+
+ Derived from the constant rather than written beside it, so the refusal, the
+ export warning and the documentation cannot drift apart from each other or
+ from what the middleware actually enforces (v1.1 WP-D).
+ """
+ size = MAX_IMPORT_BODY_BYTES if limit is None else limit
+ megabytes = size / (1024 * 1024)
+ return f"{megabytes:.0f} MB" if abs(megabytes - round(megabytes)) < 0.05 else f"{megabytes:.1f} MB"
+
+
+def oversized_export_warning(export_bytes: int, limit: int | None = None) -> str:
+ """What to tell a reader whose export is larger than import will accept.
+
+ v1.1 WP-D. The file is written and is not damaged: what it exceeds is this
+ version's import ceiling, so it cannot be brought back in *here*. Saying that
+ plainly is the whole point — the alternative is a reader who finds out when
+ they try to restore it.
+ """
+ size = MAX_IMPORT_BODY_BYTES if limit is None else limit
+ return (
+ f"This export is larger than this version's {import_limit_label(size)} import "
+ f"limit ({export_bytes:,} bytes). The file was exported successfully, but this "
+ f"version cannot import it."
+ )
+
+
class BodySizeLimitMiddleware:
"""Rejects oversized request bodies by their declared `Content-Length`.
diff --git a/backend/app/routers/adventures/bundle_io.py b/backend/app/routers/adventures/bundle_io.py
index 966b63d..4818532 100644
--- a/backend/app/routers/adventures/bundle_io.py
+++ b/backend/app/routers/adventures/bundle_io.py
@@ -28,7 +28,9 @@ repair. Refusing a whole campaign because a search index would not build would
trade the valuable thing for the cheap one.
"""
-from fastapi import Body, Depends, Request
+import json
+
+from fastapi import Body, Depends, Request, Response
from sqlalchemy.orm import Session
from ... import bundle, head, limits, models, schemas
@@ -46,8 +48,38 @@ def export_adventure(
`app/bundle.py` owns the format, in all three of its versions. A backup
outlives the schema, so no call site decides anything about its shape.
+
+ **v1.1 WP-D: the export also says whether this version could import it back.**
+ A campaign large enough to pass `limits.MAX_IMPORT_BODY_BYTES` still exports —
+ the file is complete and not damaged, and refusing to write it would destroy
+ the only copy the reader was trying to make. What it cannot do is come back
+ in here, and the reader is told that at the moment they take it rather than
+ at the moment they need it.
+
+ It travels in headers, not in the body. The body is the bundle, the browser
+ saves exactly those bytes as the file, and a warning inside it would become
+ part of a portable story file and of every checksum taken over one.
+
+ The size measured is the compact serialisation, because that is both what
+ this response sends and what the browser POSTs back on import, which is what
+ `BodySizeLimitMiddleware` weighs. The pretty-printed file the reader
+ downloads is larger, and is not what import reads.
"""
- return bundle.export(db, adv)
+ payload = bundle.export(db, adv)
+ # Serialised exactly as Starlette's JSONResponse would, so the bytes counted
+ # are the bytes sent.
+ body = json.dumps(payload, ensure_ascii=False, allow_nan=False,
+ separators=(",", ":")).encode("utf-8")
+ limit = limits.MAX_IMPORT_BODY_BYTES
+ importable = len(body) <= limit
+ headers = {
+ "X-Export-Bytes": str(len(body)),
+ "X-Import-Limit-Bytes": str(limit),
+ "X-Importable-By-This-Version": "true" if importable else "false",
+ }
+ if not importable:
+ headers["X-Export-Warning"] = limits.oversized_export_warning(len(body), limit)
+ return Response(content=body, media_type="application/json", headers=headers)
@router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201)
diff --git a/backend/tests/test_v11_d_recovery.py b/backend/tests/test_v11_d_recovery.py
new file mode 100644
index 0000000..a3858b3
--- /dev/null
+++ b/backend/tests/test_v11_d_recovery.py
@@ -0,0 +1,294 @@
+"""v1.1 WP-D: a backup that was really checked, and an export that says what it is.
+
+Two recovery-path claims, each of which was true only in the small before this
+package:
+
+- **A backup is verified.** M9 ran `PRAGMA quick_check` on the finished copy.
+ That reads every page and every record, and skips the cross-check between a
+ table and its indexes — so a copy whose index disagrees with its table passed.
+ `test_the_fixture_is_the_difference_between_the_two_checks` builds exactly that
+ damage and shows the two pragmas disagreeing about it, before anything here
+ uses it as evidence.
+- **An export is importable.** Nothing compared the bundle with
+ `limits.MAX_IMPORT_BODY_BYTES`, so a campaign could be exported and then
+ refused by its own importer, with the reader finding out at the moment they
+ needed it. The export still succeeds — the file is complete, and a version
+ that refused to write it would destroy the copy someone was trying to make —
+ and now it says so.
+
+ python -m pytest tests/test_v11_d_recovery.py -v
+"""
+
+import json
+import sqlite3
+from pathlib import Path
+
+import pytest
+from fastapi import Depends
+from fastapi.testclient import TestClient
+
+from app import auth, backup, limits, models
+from app.database import Base, SessionLocal, engine, get_db
+from app.main import app
+
+
+@pytest.fixture()
+def client():
+ Base.metadata.create_all(bind=engine)
+ setup = SessionLocal()
+ user = models.User(is_guest=False, email="wp-d@example.com")
+ setup.add(user)
+ setup.flush()
+ setup.add(models.Settings(user_id=user.id, model="test-model"))
+ adventure = models.Adventure(user_id=user.id, title="Recovery")
+ setup.add(adventure)
+ setup.flush()
+ setup.add(models.Action(adventure_id=adventure.id, type="start", text="The story opens."))
+ setup.commit()
+ adv_id, user_id = adventure.id, user.id
+ setup.close()
+ app.dependency_overrides[auth.get_current_user] = (
+ lambda db=Depends(get_db): db.get(models.User, user_id))
+ test_client = TestClient(app)
+ test_client.adv_id = adv_id
+ test_client.user_id = user_id
+ try:
+ yield test_client
+ finally:
+ app.dependency_overrides.clear()
+ Base.metadata.drop_all(bind=engine)
+
+
+# ------------------------------------------------------------ the fixture
+
+def build_corrupt_copy(path: Path) -> None:
+ """A database whose index disagrees with its table, and nothing else.
+
+ One digit inside one index leaf page is changed, so that entry names a key
+ no row holds and one row's key is in no index entry. Every page is still
+ structurally sound and every record still parses, which is the whole point:
+ this is the damage `quick_check` is not looking for.
+ """
+ path.unlink(missing_ok=True)
+ connection = sqlite3.connect(path)
+ connection.execute("PRAGMA page_size=4096")
+ connection.execute("CREATE TABLE t (id INTEGER PRIMARY KEY, k TEXT NOT NULL, filler TEXT)")
+ connection.execute("CREATE INDEX i_t_k ON t(k)")
+ connection.executemany("INSERT INTO t (k, filler) VALUES (?, ?)",
+ [(f"k{n:06d}", "x" * 40) for n in range(400)])
+ connection.commit()
+ page_size = connection.execute("PRAGMA page_size").fetchone()[0]
+ leaves = [row[0] for row in connection.execute(
+ "SELECT pageno FROM dbstat WHERE name='i_t_k' AND pagetype='leaf' ORDER BY pageno")]
+ connection.close()
+ assert leaves, "the index must have a leaf page to damage"
+
+ raw = bytearray(path.read_bytes())
+ start = (leaves[0] - 1) * page_size
+ page = raw[start:start + page_size]
+ at = page.find(b"k000")
+ assert at != -1, "expected an indexed key on the index's first leaf page"
+ page[at + 4] = ord("9") # k000144 -> k000944: a key no row has
+ raw[start:start + page_size] = page
+ path.write_bytes(bytes(raw))
+
+
+def check(path: Path, pragma: str) -> str:
+ connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
+ try:
+ return ", ".join(str(row[0]) for row in connection.execute(f"PRAGMA {pragma}").fetchall())
+ finally:
+ connection.close()
+
+
+def test_the_fixture_is_the_difference_between_the_two_checks(tmp_path):
+ """Before using it as evidence: quick_check calls this database fine."""
+ damaged = tmp_path / "damaged.db"
+ build_corrupt_copy(damaged)
+ assert check(damaged, "quick_check") == "ok"
+ integrity = check(damaged, "integrity_check")
+ assert integrity != "ok"
+ assert "i_t_k" in integrity # it names the index that disagrees
+
+
+# ---------------------------------------------------------------- backups
+
+def test_a_healthy_backup_passes_the_full_check_and_is_kept(client, tmp_path):
+ source = tmp_path / "campaign.db"
+ source.write_bytes(Path(str(engine.url.database)).read_bytes())
+ result = backup.create(source)
+ assert result.integrity == "ok"
+ assert result.path.exists() and result.bytes > 0
+ assert check(result.path, "integrity_check") == "ok"
+ assert result.path.parent == backup.directory(source)
+
+
+def test_the_backup_runs_the_full_check_not_the_quick_one(client, tmp_path, monkeypatch):
+ """The pragma itself, named. SQLite traces every statement it executes, so
+ this reads what the backup actually asked the copy rather than inferring it."""
+ asked: list[str] = []
+ real_connect = sqlite3.connect
+
+ def tracing(*args, **kwargs):
+ connection = real_connect(*args, **kwargs)
+ connection.set_trace_callback(asked.append)
+ return connection
+
+ monkeypatch.setattr(backup.sqlite3, "connect", tracing)
+ source = tmp_path / "campaign.db"
+ source.write_bytes(Path(str(engine.url.database)).read_bytes())
+ backup.create(source).path.unlink()
+ assert any("integrity_check" in sql for sql in asked), asked
+ assert not any("quick_check" in sql for sql in asked), asked
+
+
+def test_a_copy_the_full_check_rejects_is_not_kept(client, tmp_path, monkeypatch):
+ """The copy is damaged after it is written and before it is verified, which
+ is where a real page-level fault would appear: between the copy and the
+ rename. Nothing wearing a backup's name may be left behind."""
+ source = tmp_path / "campaign.db"
+ source.write_bytes(Path(str(engine.url.database)).read_bytes())
+ real_copy = backup._copy
+
+ def damage(source_path, working):
+ pages = real_copy(source_path, working)
+ build_corrupt_copy(working)
+ return pages
+
+ monkeypatch.setattr(backup, "_copy", damage)
+ with pytest.raises(backup.BackupError) as refused:
+ backup.create(source)
+ assert "did not verify" in str(refused.value)
+ assert "i_t_k" in str(refused.value) # it says what was wrong
+ kept = list(backup.directory(source).glob("*"))
+ assert kept == [], f"a rejected backup was left behind: {kept}"
+
+
+def test_a_rejected_backup_leaves_an_earlier_good_one_alone(client, tmp_path, monkeypatch):
+ source = tmp_path / "campaign.db"
+ source.write_bytes(Path(str(engine.url.database)).read_bytes())
+ good = backup.create(source)
+ before = good.path.read_bytes()
+
+ real_copy = backup._copy
+
+ def damage(source_path, working):
+ pages = real_copy(source_path, working)
+ build_corrupt_copy(working)
+ return pages
+
+ monkeypatch.setattr(backup, "_copy", damage)
+ with pytest.raises(backup.BackupError):
+ backup.create(source)
+ assert good.path.exists()
+ assert good.path.read_bytes() == before
+ assert check(good.path, "integrity_check") == "ok"
+ assert [p.name for p in backup.directory(source).glob("*")] == [good.path.name]
+
+
+def test_the_backup_file_semantics_are_unchanged(client, tmp_path):
+ """Same directory, same stamped name, same reported fields: WP-D changed the
+ check, not the file."""
+ source = tmp_path / "campaign.db"
+ source.write_bytes(Path(str(engine.url.database)).read_bytes())
+ first = backup.create(source)
+ second = backup.create(source)
+ assert first.path.name.startswith(backup.PREFIX) and first.path.suffix == ".db"
+ assert first.path != second.path, "an existing backup is never overwritten"
+ assert set(first.as_dict()) == {"filename", "bytes", "pages", "seconds", "integrity"}
+ listed = [row["filename"] for row in backup.existing(source)]
+ assert sorted(listed) == sorted([first.path.name, second.path.name])
+
+
+# ----------------------------------------------------------------- exports
+
+def export(client, adv_id):
+ response = client.get(f"/api/adventures/{adv_id}/export")
+ assert response.status_code == 200, response.text[:200]
+ return response
+
+
+def test_a_normal_export_carries_no_warning(client):
+ response = export(client, client.adv_id)
+ assert "X-Export-Warning" not in response.headers
+ assert response.headers["X-Importable-By-This-Version"] == "true"
+ assert int(response.headers["X-Import-Limit-Bytes"]) == limits.MAX_IMPORT_BODY_BYTES
+ assert int(response.headers["X-Export-Bytes"]) == len(response.content)
+ assert response.json()["format"] == "ai-dnd-adventure-v3"
+
+
+def fill_past_the_limit(adv_id: int) -> int:
+ """Real rows, until the campaign's bundle is genuinely over the ceiling.
+
+ Not a mocked size: the export below serialises all of it.
+ """
+ chunk = "The rain kept on over the harbour road, and nobody came. " * 900 # ~50 kB
+ written = 0
+ with SessionLocal() as db:
+ while written < limits.MAX_IMPORT_BODY_BYTES + 2 * 1024 * 1024:
+ db.add_all([models.Action(adventure_id=adv_id, type="ai", text=chunk)
+ for _ in range(40)])
+ db.commit()
+ written += 40 * len(chunk)
+ return written
+
+
+def test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back(client):
+ fill_past_the_limit(client.adv_id)
+ response = export(client, client.adv_id)
+
+ # Delivered, whole, and still the same format.
+ body = response.content
+ assert len(body) > limits.MAX_IMPORT_BODY_BYTES
+ parsed = json.loads(body)
+ assert parsed["format"] == "ai-dnd-adventure-v3"
+ assert len(parsed["actions"]) > 40
+
+ # And honest about what this version can do with it.
+ assert response.headers["X-Importable-By-This-Version"] == "false"
+ warning = response.headers["X-Export-Warning"]
+ assert limits.import_limit_label() in warning
+ assert "exported successfully" in warning
+ assert "cannot import" in warning
+ assert int(response.headers["X-Export-Bytes"]) == len(body)
+
+
+def test_the_warning_follows_the_configured_limit(monkeypatch):
+ """The text is generated from the constant, so changing the constant changes
+ the sentence rather than leaving a stale number in it."""
+ assert "20 MB" in limits.oversized_export_warning(21_000_000)
+ monkeypatch.setattr(limits, "MAX_IMPORT_BODY_BYTES", 50 * 1024 * 1024)
+ assert limits.import_limit_label() == "50 MB"
+ assert "50 MB" in limits.oversized_export_warning(60_000_000)
+ assert "20 MB" not in limits.oversized_export_warning(60_000_000)
+
+
+def test_the_bundle_itself_never_carries_the_warning(client):
+ """The warning is about the export, not part of the portable story file."""
+ fill_past_the_limit(client.adv_id)
+ parsed = json.loads(export(client, client.adv_id).content)
+ flat = json.dumps(parsed).lower()
+ assert "import limit" not in flat
+ assert "cannot import" not in flat
+ for key in parsed:
+ assert "warning" not in key.lower()
+
+
+def test_that_same_bundle_is_refused_by_import_naming_the_limit(client):
+ fill_past_the_limit(client.adv_id)
+ body = export(client, client.adv_id).content
+ response = client.post("/api/adventures/import", content=body,
+ headers={"Content-Type": "application/json"})
+ assert response.status_code == 413
+ detail = response.json()["detail"]
+ assert "too large" in detail.lower()
+ assert limits.import_limit_label() in detail
+
+
+def test_a_bundle_under_the_limit_still_imports(client):
+ """The refusal is about size alone: the ordinary path is untouched."""
+ body = export(client, client.adv_id).content
+ assert len(body) < limits.MAX_IMPORT_BODY_BYTES
+ response = client.post("/api/adventures/import", content=body,
+ headers={"Content-Type": "application/json"})
+ assert response.status_code == 201, response.text[:200]
diff --git a/backend/tests/test_v11_e_contrast.py b/backend/tests/test_v11_e_contrast.py
new file mode 100644
index 0000000..32f6c08
--- /dev/null
+++ b/backend/tests/test_v11_e_contrast.py
@@ -0,0 +1,152 @@
+"""v1.1 WP-E: the contrast audit is a gate, not a report.
+
+Before this package `tools/contrast_audit.py` measured control boundaries,
+printed that two of them were below 3:1, and exited 0 anyway — on the argument
+that a control is identified by its label rather than its edge. A check that
+cannot fail is not a check, and these tests are what make it one: the threshold
+is exercised from both sides, on a real tokens file, so a future palette change
+that dims a control's edge stops the run instead of adding a line to it.
+
+ python -m pytest tests/test_v11_e_contrast.py -v
+"""
+
+import sys
+from pathlib import Path
+
+import pytest
+
+sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tools"))
+
+import contrast_audit as audit # noqa: E402
+
+
+# --------------------------------------------------------------- the maths
+
+def test_the_ratio_is_the_wcag_ratio():
+ """Anchored on values with known answers, so a broken formula is visible."""
+ assert audit.ratio("#ffffff", "#000000") == pytest.approx(21.0, abs=0.01)
+ assert audit.ratio("#ffffff", "#ffffff") == pytest.approx(1.0, abs=0.001)
+ # Order does not matter: contrast is symmetric.
+ assert audit.ratio("#131320", "#676792") == pytest.approx(
+ audit.ratio("#676792", "#131320"), abs=1e-9)
+
+
+# ------------------------------------------------------- the gate itself
+
+def tokens_file(tmp_path: Path, **overrides: str) -> Path:
+ """A real tokens file with named colours replaced."""
+ source = audit.TOKENS.read_text()
+ for name, value in overrides.items():
+ token = "--" + name.replace("_", "-")
+ start = source.index(f"{token}: ")
+ end = source.index(";", start)
+ source = source[:start] + f"{token}: {value}" + source[end:]
+ written = tmp_path / "tokens.css"
+ written.write_text(source)
+ return written
+
+
+def run_against(path: Path, monkeypatch) -> int:
+ monkeypatch.setattr(audit, "TOKENS", path)
+ return audit.main()
+
+
+def test_a_boundary_below_three_to_one_fails_the_run(tmp_path, monkeypatch, capsys):
+ """2.99:1 against --bg-input — the wrong side of the line by one hundredth.
+
+ This value clears 3:1 against --bg-panel (3.21:1), so it would have passed
+ the audit as M11 wrote it. It fails now because the floor is taken against
+ the background the control is actually drawn on.
+ """
+ below = tokens_file(tmp_path, border="#58639a")
+ assert run_against(below, monkeypatch) == 1
+ printed = capsys.readouterr().out
+ assert "2.99:1" in printed
+ assert "FAIL" in printed
+ assert "below 3:1 (WCAG 1.4.11)" in printed
+
+
+def test_a_boundary_at_three_to_one_passes(tmp_path, monkeypatch, capsys):
+ """3.00:1 — the right side of the same line, one hundredth from the value
+ above, and not failed for arithmetic the reader cannot see."""
+ at = tokens_file(tmp_path, border="#58639b")
+ assert run_against(at, monkeypatch) == 0
+ printed = capsys.readouterr().out
+ assert "3.00:1" in printed
+ assert "every control boundary clears 3:1" in printed
+
+
+def test_the_v1_0_0_boundary_would_now_fail(tmp_path, monkeypatch, capsys):
+ """The value v1.0.0 shipped. This is the defect WP-E closes, and the gate
+ has to be the thing that would have caught it."""
+ shipped = tokens_file(tmp_path, border="#2b2b3d")
+ assert run_against(shipped, monkeypatch) == 1
+ assert "1.24:1" in capsys.readouterr().out # against --bg-input, the worst case
+
+
+def test_a_text_pair_below_four_point_five_still_fails(tmp_path, monkeypatch, capsys):
+ """WP-E raised the boundary floor and must not have lowered the text one."""
+ dimmed = tokens_file(tmp_path, text_dim="#5a5750")
+ assert run_against(dimmed, monkeypatch) == 1
+ assert "below WCAG AA (1.4.3)" in capsys.readouterr().out
+
+
+def test_a_missing_token_is_a_failure_not_a_skip(tmp_path, monkeypatch, capsys):
+ source = audit.TOKENS.read_text().replace("--border-bright:", "--border-was-renamed:")
+ written = tmp_path / "tokens.css"
+ written.write_text(source)
+ assert run_against(written, monkeypatch) == 1
+ assert "MISSING TOKEN" in capsys.readouterr().out
+
+
+# ------------------------------------------------- the shipped palette
+
+def test_the_real_tokens_pass_both_criteria(capsys):
+ """The palette as it stands, through the same gate CI would run."""
+ assert audit.main() == 0
+ printed = capsys.readouterr().out
+ assert "every text pair clears WCAG AA (1.4.3)" in printed
+ assert "every control boundary clears 3:1 (1.4.11)" in printed
+ assert "FAIL" not in printed
+
+
+def test_every_boundary_pair_is_measured_against_the_background_it_is_drawn_on():
+ """The audit used to check borders only against --bg-panel, which is not
+ where the bordered controls are: inputs and buttons sit on --bg-input
+ (styles/forms.css), which is lighter and therefore harder. Checking only the
+ easier background would let a token pass while the real control failed."""
+ boundary = [(fg, bg) for kind, fg, bg, _, _ in audit.PAIRS if kind == "boundary"]
+ for background in ("--bg-input", "--bg-panel", "--bg"):
+ assert ("--border", background) in boundary, background
+ assert ("--border-bright", "--bg-input") in boundary
+
+
+def test_the_text_baselines_are_unchanged_by_wp_e():
+ """WP-E changed only boundary tokens. These are the M11 text measurements,
+ and they have to still be exactly what the earlier reports recorded."""
+ tokens = audit.read_tokens(audit.TOKENS)
+ measured = {
+ "body text on the page": audit.ratio(tokens["--text"], tokens["--bg"]),
+ "body text in a panel": audit.ratio(tokens["--text"], tokens["--bg-panel"]),
+ "secondary text in a panel": audit.ratio(tokens["--text-dim"], tokens["--bg-panel"]),
+ "secondary text on the page": audit.ratio(tokens["--text-dim"], tokens["--bg"]),
+ }
+ assert measured["body text on the page"] == pytest.approx(14.57, abs=0.01)
+ assert measured["body text in a panel"] == pytest.approx(13.57, abs=0.01)
+ assert measured["secondary text in a panel"] == pytest.approx(5.48, abs=0.01)
+ assert measured["secondary text on the page"] == pytest.approx(5.88, abs=0.01)
+
+
+def test_the_hover_edge_stays_brighter_than_the_resting_edge():
+ """Rest and hover have to remain distinguishable from each other, not merely
+ each clear the floor against the background."""
+ tokens = audit.read_tokens(audit.TOKENS)
+ panel = tokens["--bg-panel"]
+ assert audit.ratio(tokens["--border-bright"], panel) > audit.ratio(tokens["--border"], panel)
+
+
+def test_a_boundary_never_becomes_as_loud_as_body_text():
+ """A control's edge that outshines the words inside it is its own defect."""
+ tokens = audit.read_tokens(audit.TOKENS)
+ panel = tokens["--bg-panel"]
+ assert audit.ratio(tokens["--border-bright"], panel) < audit.ratio(tokens["--text"], panel)
diff --git a/backend/tools/contrast_audit.py b/backend/tools/contrast_audit.py
index c084632..c412bb5 100644
--- a/backend/tools/contrast_audit.py
+++ b/backend/tools/contrast_audit.py
@@ -26,16 +26,24 @@ TOKENS = Path(__file__).resolve().parent.parent.parent / "frontend/src/styles/to
#: rather than combinatorial, because "every colour against every other" reports
#: pairs that never meet on screen.
#:
-#: The `kind` matters and is not a way of grading on a curve. **text** pairs are
-#: WCAG 1.4.3 Contrast (Minimum) and are what §21 of the M11 brief asks about;
-#: they are pass/fail. **boundary** pairs are WCAG 1.4.11 Non-text Contrast,
-#: which applies to "visual information required to identify user interface
-#: components" — and in this design a control is identified by its *label*,
-#: which is measured above and passes, not by its edge. So a boundary below 3:1
-#: is reported with its number and does not fail the run; what it would take to
-#: turn it into a real failure is a control with no visible label, and there is
-#: no such control (`tools/m11_browser.py` asserts every visible control has an
-#: accessible name, and the story controls are text buttons).
+#: The `kind` records which success criterion a pair is measured against —
+#: **text** is WCAG 1.4.3 Contrast (Minimum), **boundary** is WCAG 1.4.11
+#: Non-text Contrast — and **both are pass/fail**.
+#:
+#: v1.1 WP-E overturned the earlier position here, which was that a boundary
+#: below 3:1 could be recorded rather than failed because "a control is
+#: identified by its label, not by its edge". That argument understates what
+#: 1.4.11 asks: the criterion covers the visual information needed to identify
+#: a component *and its boundary*, and a reader who cannot see where a text box
+#: ends cannot see that there is a text box to type into, label or no label.
+#: The edges were at 1.33:1 and 1.75:1 — the tokens were raised instead.
+#:
+#: Borders are measured against **every background they are drawn on**, and the
+#: floor is the worst of them. Inputs and buttons sit on --bg-input, which is
+#: lighter than --bg-panel and so the harder case; checking only --bg-panel
+#: would have let a token pass the audit while the real control failed.
+#: `--bg-panel-glass` is translucent and cannot be resolved from tokens alone;
+#: that edge is measured on the rendered page by `tools/m11_browser.py`.
PAIRS = [
("text", "--text", "--bg", 4.5, "body text on the page"),
("text", "--text", "--bg-panel", 4.5, "body text in a panel"),
@@ -48,8 +56,14 @@ PAIRS = [
("text", "--danger", "--bg-panel", 4.5, "an error message"),
("text", "--warning", "--bg-panel", 4.5, "a caution message"),
("text", "--player", "--bg", 4.5, "the player's own words"),
- ("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge"),
- ("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge"),
+ ("boundary", "--border", "--bg-input", 3.0, "a field or button's resting edge"),
+ ("boundary", "--border-bright", "--bg-input", 3.0, "a field or button's hover edge"),
+ ("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge in a panel"),
+ ("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge in a panel"),
+ ("boundary", "--border", "--bg", 3.0, "a divider on the page"),
+ ("boundary", "--border-bright", "--bg", 3.0, "the scrollbar thumb"),
+ ("boundary", "--accent-dim", "--bg-input", 3.0, "a focused field's edge"),
+ ("boundary", "--accent-dim", "--bg-panel", 3.0, "a focused control's edge in a panel"),
("boundary", "--chart-1", "--bg-panel", 3.0, "a chart bar"),
("boundary", "--chart-2", "--bg-panel", 3.0, "a chart bar"),
("boundary", "--chart-3", "--bg-panel", 3.0, "a chart bar"),
@@ -81,37 +95,44 @@ def ratio(a: str, b: str) -> float:
def main() -> int:
tokens = read_tokens(TOKENS)
- print(f"{TOKENS.relative_to(TOKENS.parents[3])}: {len(tokens)} colour tokens\n")
+ # Named defensively: the tests run this against a temporary tokens file,
+ # which need not sit four directories deep the way the real one does.
+ label = TOKENS.name
+ if len(TOKENS.parents) > 3:
+ label = TOKENS.relative_to(TOKENS.parents[3])
+ print(f"{label}: {len(tokens)} colour tokens\n")
print(f"{'pair':44} {'kind':9} {'ratio':>7} {'floor':>6} verdict")
print("-" * 82)
- failures, advisories = 0, 0
+ text_failures, boundary_failures = 0, 0
for kind, foreground, background, floor, description in PAIRS:
if foreground not in tokens or background not in tokens:
print(f"{description:44} {kind:9} {'—':>7} {floor:>6.1f} MISSING TOKEN")
- failures += 1
+ text_failures += 1
continue
measured = ratio(tokens[foreground], tokens[background])
- ok = measured >= floor
- if not ok:
- if kind == "text":
- failures += 1
- verdict = "FAIL"
- else:
- advisories += 1
- verdict = "below 1.4.11 (label carries it)"
- else:
+ # Rounded to the two decimals printed, so the verdict matches what the
+ # reader is shown: a pair reported as 3.00:1 is not failed for arithmetic
+ # the output does not display.
+ if round(measured, 2) >= floor:
verdict = "pass"
+ else:
+ verdict = "FAIL"
+ if kind == "text":
+ text_failures += 1
+ else:
+ boundary_failures += 1
print(f"{description:44} {kind:9} {measured:>6.2f}:1 {floor:>6.1f} {verdict}")
print()
- if failures:
- print(f"{failures} text pair(s) below WCAG AA — this is a defect")
+ if text_failures:
+ print(f"{text_failures} text pair(s) below WCAG AA (1.4.3) — this is a defect")
else:
print("every text pair clears WCAG AA (1.4.3)")
- if advisories:
- print(f"{advisories} boundary pair(s) below 3:1 (1.4.11). Recorded rather "
- "than failed: every control in this design carries a visible text "
- "label, which is measured above and passes.")
- return 1 if failures else 0
+ if boundary_failures:
+ print(f"{boundary_failures} boundary pair(s) below 3:1 (WCAG 1.4.11) — "
+ "this is a defect")
+ else:
+ print("every control boundary clears 3:1 (1.4.11)")
+ return 1 if (text_failures or boundary_failures) else 0
if __name__ == "__main__":
diff --git a/backend/tools/m11_browser.py b/backend/tools/m11_browser.py
index 16b85bf..217ed94 100644
--- a/backend/tools/m11_browser.py
+++ b/backend/tools/m11_browser.py
@@ -665,6 +665,228 @@ def check_accessibility(browser: Browser, site: Site, adv: int, checks: Checks):
checks.record("A11y", "the story input takes keyboard focus", typed is True)
+# -------------------------------------------------------------- WP-E boundaries
+
+#: The controls whose edges WP-E measures, as the stylesheets actually draw them.
+#: `side` is the border the rule sets: .topnav paints only a bottom edge.
+BOUNDARY_TARGETS = [
+ ("E01", "the story composer", ".input-bar", "Top"),
+ ("E02", "a story control", ".story-controls button:not(:disabled)", "Top"),
+ ("E04", "the open panel tab", ".panel-tabs button.active", "Top"),
+]
+
+#: `.topnav` is measured on the library route, because the play route does not
+#: render it — the play page has its own `.play-header`, which story.css keeps
+#: deliberately opaque. So the translucent case exists on exactly one screen,
+#: and it is the case token arithmetic cannot answer: --bg-panel-glass is
+#: rgba(...,0.82) over a gradient, so only the rendered page knows what is
+#: behind that edge.
+TRANSLUCENT_TARGET = ("E03", "the top navigation's edge", ".topnav", "Bottom")
+
+#: Measured in the browser rather than from tokens, because two of these cannot
+#: be derived from tokens at all: .topnav sits on --bg-panel-glass, which is
+#: translucent, so its effective background is a composite of what is behind it;
+#: and a rendered edge can be changed by opacity, a transition mid-flight, or a
+#: rule the token file knows nothing about.
+#:
+#: A boundary is measured against **both** adjacent colours — the control's own
+#: fill inside it and the background outside it — and passes on the better of
+#: the two. An edge that matches its fill but contrasts with the page is still a
+#: visible outline, and vice versa; what 1.4.11 asks is that the component's
+#: extent be perceivable, not that every neighbouring surface differ from it.
+BOUNDARY_JS = """
+ const parse = (c) => (c.match(/[\\d.]+/g) || []).map(Number);
+ const alphaOf = (c) => {
+ if (!c || c === 'transparent') return 0;
+ const p = parse(c);
+ return p.length > 3 ? p[3] : 1;
+ };
+ function lum(c) {
+ const [r, g, b] = parse(c).slice(0, 3).map(v => v / 255)
+ .map(v => v <= 0.03928 ? v / 12.92 : Math.pow((v + 0.055) / 1.055, 2.4));
+ return 0.2126 * r + 0.7152 * g + 0.0722 * b;
+ }
+ function over(fg, bg) {
+ const f = parse(fg), b = parse(bg), a = alphaOf(fg);
+ return 'rgb(' + [0, 1, 2].map(i => Math.round(f[i] * a + b[i] * (1 - a))).join(', ') + ')';
+ }
+ function ratio(x, y) {
+ const a = lum(x), b = lum(y);
+ return (Math.max(a, b) + 0.05) / (Math.min(a, b) + 0.05);
+ }
+ // Everything painted behind `el`, composited bottom-up, so a translucent
+ // panel reports the colour a reader actually sees rather than its own rgba.
+ function behind(el) {
+ const layers = [];
+ let node = el;
+ while (node && node !== document.documentElement) {
+ const c = getComputedStyle(node).backgroundColor;
+ if (alphaOf(c) > 0) {
+ layers.push(c);
+ if (alphaOf(c) >= 1) break;
+ }
+ node = node.parentElement;
+ }
+ let result = 'rgb(10, 10, 15)';
+ const root = getComputedStyle(document.documentElement).backgroundColor;
+ const body = getComputedStyle(document.body).backgroundColor;
+ if (alphaOf(root) >= 1) result = root;
+ else if (alphaOf(body) >= 1) result = body;
+ for (let i = layers.length - 1; i >= 0; i--) {
+ result = alphaOf(layers[i]) >= 1 ? layers[i] : over(layers[i], result);
+ }
+ return result;
+ }
+ const [selector, side] = arguments;
+ const el = document.querySelector(selector);
+ if (!el) return null;
+ const s = getComputedStyle(el);
+ const edge = s['border' + side + 'Color'];
+ const width = parseFloat(s['border' + side + 'Width']) || 0;
+ const opacity = parseFloat(s.opacity);
+ const inside = alphaOf(s.backgroundColor) >= 1
+ ? s.backgroundColor : over(s.backgroundColor, behind(el.parentElement || el));
+ const outside = behind(el.parentElement || el);
+ return {
+ selector, edge, width, opacity, inside, outside,
+ insideRatio: Math.round(ratio(edge, inside) * 100) / 100,
+ outsideRatio: Math.round(ratio(edge, outside) * 100) / 100,
+ outline: s.outlineStyle + ' ' + s.outlineWidth + ' ' + s.outlineColor,
+ shadow: s.boxShadow,
+ };
+"""
+
+
+def _settled(browser: Browser, selector: str, side: str, *, timeout: float = 5) -> None:
+ """Waits until the edge colour stops moving.
+
+ Every one of these controls carries `transition: border-color 0.15s`, so a
+ measurement taken the instant after a hover or a focus reads a colour part
+ way between the two states — a value no state actually has. The first run of
+ this check did exactly that: it reported the hover edge as rgb(114,114,160),
+ which is neither --border nor --border-bright but a frame between them.
+ Polled rather than slept, per this harness's rule.
+ """
+ script = ("(() => { const el = document.querySelector(%s);"
+ " if (!el) return true;"
+ " const c = getComputedStyle(el)['border%sColor'];"
+ " const was = window.__wpeEdge; window.__wpeEdge = c;"
+ " return was === c; })()" % (json.dumps(selector), side))
+ # Cleared first, so the comparison starts from no previous reading rather
+ # than from whatever the last measured control left behind.
+ browser.js("window.__wpeEdge = undefined")
+ browser.wait_js(script, timeout=timeout)
+
+
+def _boundary(browser: Browser, selector: str, side: str) -> dict | None:
+ _settled(browser, selector, side)
+ return browser.js(BOUNDARY_JS, selector, side)
+
+
+def _record_boundary(checks: Checks, test: str, what: str, row: dict | None) -> None:
+ if row is None:
+ checks.skip(test, what, "the control was not on the page")
+ return
+ best = max(row["insideRatio"], row["outsideRatio"])
+ detail = (f"{best:.2f}:1 (edge {row['edge']} — {row['insideRatio']}:1 against its fill "
+ f"{row['inside']}, {row['outsideRatio']}:1 against {row['outside']}), "
+ f"{row['width']}px")
+ checks.record("WCAG 1.4.11", what, best >= 3.0 and row["width"] > 0, detail)
+
+
+def check_control_boundaries(browser: Browser, site: Site, adv: int, checks: Checks,
+ evidence: dict, out: Path) -> None:
+ """v1.1 WP-E §21: a control's edge is visible on its own, measured.
+
+ M11 measured text contrast on the rendered page and left boundaries to the
+ token audit, which checked them against one background and reported a
+ shortfall without failing. These are the rendered edges, at rest, on hover
+ and while focused, against what is actually behind them.
+ """
+ _open_play(browser, site, adv)
+ measured: dict[str, dict] = {}
+
+ # A panel tab's edge is `transparent` until the panel is open, which makes
+ # the open tab the one control here whose boundary is the only thing marking
+ # it — exactly what 1.4.11 is about. So one is opened rather than skipped.
+ if not _open_panel(browser, "State"):
+ checks.skip("E04", "the open panel tab — resting edge", "the State panel did not open")
+
+ for test, what, selector, side in BOUNDARY_TARGETS:
+ row = _boundary(browser, selector, side)
+ _record_boundary(checks, test, f"{what} — resting edge", row)
+ if row:
+ measured[f"{what} (rest)"] = row
+
+ # Hover, with a real pointer: `:hover` follows the browser's pointer state,
+ # so a synthetic mouseover would silently re-measure the resting edge.
+ control = browser.find(".story-controls button:not(:disabled)", required=False)
+ if control is None:
+ checks.skip("E02", "a story control — hover edge", "no enabled control on the page")
+ else:
+ browser.hover(control)
+ hovered = browser.js(
+ "return document.querySelector('.story-controls button:not(:disabled)')"
+ " .matches(':hover');")
+ if hovered is not True:
+ checks.skip("E02", "a story control — hover edge",
+ "the pointer did not land on the control")
+ else:
+ row = _boundary(browser, ".story-controls button:not(:disabled)", "Top")
+ _record_boundary(checks, "E02", "a story control — hover edge", row)
+ if row:
+ measured["a story control (hover)"] = row
+ browser.unhover()
+
+ # Focus: the composer's edge changes colour while it holds focus, and that
+ # edge is what tells a keyboard reader where they are.
+ focused = browser.js("""
+ const box = document.querySelector('.input-main textarea');
+ if (!box) return false;
+ box.focus();
+ return document.activeElement === box;
+ """)
+ if focused is not True:
+ checks.skip("E01", "the story composer — focused edge", "the composer did not take focus")
+ else:
+ row = _boundary(browser, ".input-bar", "Top")
+ _record_boundary(checks, "E01", "the story composer — focused edge", row)
+ if row:
+ measured["the story composer (focus)"] = row
+ # A focus ring that is only a colour change is not enough on its own;
+ # M11 already asserts a visible focus indicator, and this says the
+ # focused edge is also measurably distinct from the resting one.
+ rest = measured.get("the story composer (rest)")
+ if rest:
+ checks.record("WCAG 1.4.11", "the focused composer edge differs from its resting edge",
+ row["edge"] != rest["edge"],
+ f"rest {rest['edge']} -> focus {row['edge']}")
+
+ shot = browser.screenshot(out / "control-boundaries.png")
+ checks.record("WP-E", "a screenshot of the measured controls was captured",
+ shot.exists() and shot.stat().st_size > 0, str(shot))
+
+ # The translucent edge, on the only screen that has it. Token arithmetic
+ # cannot reach this one: --bg-panel-glass is rgba over a gradient, so what
+ # is behind the nav's bottom edge is known only to the rendered page.
+ test, what, selector, side = TRANSLUCENT_TARGET
+ browser.go(site.url)
+ browser.wait_for(".topnav", timeout=30)
+ row = _boundary(browser, selector, side)
+ _record_boundary(checks, test, f"{what} — over a translucent panel", row)
+ if row:
+ measured[f"{what} (translucent)"] = row
+ checks.record("WP-E", "the nav's background really is translucent",
+ row["outside"] != row["inside"],
+ f"composited to {row['inside']} over {row['outside']}")
+ nav_shot = browser.screenshot(out / "control-boundaries-nav.png")
+ checks.record("WP-E", "a screenshot of the navigation edge was captured",
+ nav_shot.exists() and nav_shot.stat().st_size > 0, str(nav_shot))
+
+ evidence["control_boundaries"] = {"measured": measured,
+ "screenshots": [str(shot), str(nav_shot)]}
+
+
# -------------------------------------------------------------- WP-C scenarios
def _take_count_is(count: str) -> str:
@@ -1354,7 +1576,11 @@ def main() -> int:
"C5", "failed generation and recovery"),
("export", lambda: check_export_download(browser, site, adv, checks, evidence, out)),
]
- for suite, scenarios in (("M11", m11), ("WP-C", wpc)):
+ wpe = [
+ ("boundaries", lambda: check_control_boundaries(browser, site, adv, checks,
+ evidence, out)),
+ ]
+ for suite, scenarios in (("M11", m11), ("WP-C", wpc), ("WP-E", wpe)):
checks.suite = suite
for name, scenario in scenarios:
if not wanted(name):
@@ -1383,7 +1609,8 @@ def main() -> int:
"kind": ("release regression" if narrated and not only else
"partial (no narrator)" if not narrated else f"development (only {sorted(only)})"),
"checks": checks.rows,
- "suites": {"M11": checks.counts("M11"), "WP-C": checks.counts("WP-C")},
+ "suites": {"M11": checks.counts("M11"), "WP-C": checks.counts("WP-C"),
+ "WP-E": checks.counts("WP-E")},
"passed": checks.counts()["passed"],
"failed": checks.counts()["failed"],
"skipped": checks.counts()["skipped"],
diff --git a/backend/tools/m11_webdriver.py b/backend/tools/m11_webdriver.py
index 118f22d..122dfa5 100644
--- a/backend/tools/m11_webdriver.py
+++ b/backend/tools/m11_webdriver.py
@@ -25,6 +25,7 @@ has actually finished arriving — never the click that started it.
from __future__ import annotations
+import base64
import json
import os
import shutil
@@ -226,6 +227,45 @@ class Browser:
def source(self) -> str:
return self._call("GET", self._s("/source"))["value"]
+ def screenshot(self, path) -> Path:
+ """The viewport as a PNG, written where you ask (v1.1 WP-E).
+
+ Evidence for a change a reader judges by looking at it: a contrast ratio
+ says a boundary is measurable, and a picture says what it looks like.
+ """
+ encoded = self._call("GET", self._s("/screenshot"))["value"]
+ target = Path(path)
+ target.parent.mkdir(parents=True, exist_ok=True)
+ target.write_bytes(base64.b64decode(encoded))
+ return target
+
+ def hover(self, element: str) -> None:
+ """A real pointer over an element, so `:hover` actually applies.
+
+ Dispatching a mouseover event from JavaScript does not do this: CSS
+ `:hover` follows the browser's own pointer state, not a synthetic event,
+ so a measurement taken after `dispatchEvent` reads the resting style and
+ reports it as the hover style. This moves the pointer (v1.1 WP-E).
+ """
+ self._call("POST", self._s("/execute/sync"), {
+ "script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
+ "args": [{ELEMENT_KEY: element}]})
+ self._call("POST", self._s("/actions"), {"actions": [{
+ "type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
+ "actions": [{"type": "pointerMove", "duration": 60,
+ "origin": {ELEMENT_KEY: element}, "x": 0, "y": 0}]}]})
+
+ def unhover(self) -> None:
+ """Move the pointer off whatever it was over, and forget the input state."""
+ self._call("POST", self._s("/actions"), {"actions": [{
+ "type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
+ "actions": [{"type": "pointerMove", "duration": 30,
+ "origin": "viewport", "x": 0, "y": 0}]}]})
+ try:
+ self._call("DELETE", self._s("/actions"))
+ except WebDriverError:
+ pass
+
def js(self, script: str, *args):
return self._call("POST", self._s("/execute/sync"),
{"script": script, "args": list(args)})["value"]
diff --git a/backend/tools/wpd_backup_ui.py b/backend/tools/wpd_backup_ui.py
new file mode 100644
index 0000000..41510f9
--- /dev/null
+++ b/backend/tools/wpd_backup_ui.py
@@ -0,0 +1,353 @@
+"""v1.1 WP-D criterion 3: the backup completes **through the UI**.
+
+ python -m tools.wpd_backup_ui --out
[--case real|large|both] [--show]
+
+Run from `backend/`, with `frontend/dist` already built.
+
+WP-D's first pass drove `POST /api/backups` — the endpoint the *Back up now*
+button calls — on a 2.3 MB campaign database and a 117 MB one. That is evidence
+about the implementation, and the plan's criterion 3 asks for something else:
+that the backup *completes through the UI* on both. A reader does not call an
+endpoint. This drives the reader-facing control in a real Firefox, against the
+production build served by FastAPI, exactly as WP-C's harness does.
+
+**No narrator and no inference.** A backup needs neither, so nothing here
+touches a model host.
+
+What it refuses to call a pass, per the brief:
+
+- the click does nothing (no toast, no file);
+- the request fails (an error toast);
+- no backup file appears on disk;
+- the finished copy does not pass a full `PRAGMA integrity_check`;
+- the UI reports an error.
+
+Every wait is on a condition the page or the filesystem can show. Nothing here
+sleeps and then asserts.
+"""
+
+from __future__ import annotations
+
+import argparse
+import hashlib
+import json
+import shutil
+import sqlite3
+import sys
+import time
+from datetime import datetime
+from pathlib import Path
+
+sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
+
+from tools.m11_webdriver import ( # noqa: E402
+ Browser, Site, geckodriver_version, require_under_home,
+)
+
+BACKEND = Path(__file__).resolve().parent.parent
+HOME = Path.home()
+
+#: The two databases criterion 3 names. Both are copied before use: the first is
+#: the M11 evidence campaign and must not be written to, and the second is the
+#: 117 MB application database built for WP-D's timing.
+DEFAULT_REAL = HOME / "m11-evidence/m04-final/campaign.db"
+DEFAULT_LARGE = HOME / "v11-evidence/wp-d/endpoint/campaign-100mb/campaign.db"
+
+BLOCK = '[data-testid="database-backup"]'
+BUTTON = f'{BLOCK} button.primary'
+
+
+class Checks:
+ """Results, with the discipline that an unrun check is not a passing one."""
+
+ def __init__(self) -> None:
+ self.rows: list[dict] = []
+
+ def record(self, case: str, name: str, ok: bool, detail: str = "") -> bool:
+ self.rows.append({"case": case, "check": name,
+ "result": "PASS" if ok else "FAIL", "detail": detail})
+ print(f" {'ok ' if ok else 'FAIL'} {case:6} {name}"
+ + (f" — {detail}" if detail else ""), flush=True)
+ return ok
+
+ @property
+ def failed(self) -> list[dict]:
+ return [r for r in self.rows if r["result"] == "FAIL"]
+
+
+# ----------------------------------------------------------------- helpers
+
+def integrity_of(path: Path) -> tuple[str, float]:
+ """The full check on a finished copy, and what it cost, measured here.
+
+ The application runs its own `integrity_check` before keeping the file; this
+ is an independent second opinion on the artefact the UI produced, and it is
+ where criterion 3's 'time for integrity_check' comes from.
+ """
+ connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
+ try:
+ started = time.perf_counter()
+ rows = connection.execute("PRAGMA integrity_check").fetchall()
+ elapsed = time.perf_counter() - started
+ finally:
+ connection.close()
+ return ", ".join(str(r[0]) for r in rows), elapsed
+
+
+def digest(path: Path) -> str:
+ return hashlib.sha256(path.read_bytes()).hexdigest()[:16]
+
+
+def backups_in(db_path: Path) -> dict[str, int]:
+ directory = db_path.parent / "backups"
+ if not directory.exists():
+ return {}
+ return {p.name: p.stat().st_size for p in directory.iterdir() if p.is_file()}
+
+
+def wait_for_new_backup(db_path: Path, before: set[str], *, timeout: float = 300,
+ poll: float = 0.2) -> Path | None:
+ """The backup file the application wrote, once it is really there.
+
+ A file condition, not a sleep: a name that was not there before, not a
+ `.partial`, and a size that has stopped growing.
+ """
+ directory = db_path.parent / "backups"
+ deadline = time.monotonic() + timeout
+ sizes: dict[str, int] = {}
+ while time.monotonic() < deadline:
+ if directory.exists():
+ for path in directory.iterdir():
+ if not path.is_file() or path.name in before:
+ continue
+ if path.name.endswith(".partial"):
+ continue
+ size = path.stat().st_size
+ if size > 0 and sizes.get(path.name) == size:
+ return path
+ sizes[path.name] = size
+ time.sleep(poll)
+ return None
+
+
+def toasts(browser: Browser) -> list[dict]:
+ """What the page is telling the reader — the message apart from its mark.
+
+ A toast renders a decorative mark before the message:
+ `❖…`. So the
+ button's `textContent` begins with that character, and the first run of this
+ tool matched `textContent.startswith('Backup written:')` and reported six
+ *passing* behaviours as failures — the UI had said exactly what it should,
+ and the assertion was reading the mark. `message` is the message span alone;
+ `text` is kept whole for the evidence record.
+ """
+ return browser.js("""
+ const host = document.querySelector('.toast-host');
+ if (!host) return [];
+ return [...host.querySelectorAll('button.toast')].map(b => {
+ const span = b.querySelector('span:not(.toast-mark)');
+ return {
+ error: b.classList.contains('toast-error'),
+ message: (span ? span.textContent : b.textContent).trim(),
+ text: b.textContent.trim(),
+ };
+ });
+ """) or []
+
+
+def open_backup_panel(browser: Browser, site: Site, checks: Checks, case: str) -> bool:
+ """Reach the control the way a reader does: the nav link, then the panel."""
+ browser.go(site.url)
+ browser.wait_for(".topnav", timeout=60)
+ link = browser.find('.nav-links a[href="/settings"]', required=False)
+ if link is None:
+ return checks.record(case, "the Settings link is in the navigation", False)
+ browser.click(link)
+ opened = browser.wait_js(f"!!document.querySelector('{BLOCK}')", timeout=60)
+ if not checks.record(case, "Settings opens from the navigation link", opened,
+ browser.url):
+ return False
+
+ summary = browser.find(f"{BLOCK} summary", required=False)
+ if summary is None:
+ return checks.record(case, "the backup panel has a disclosure", False)
+ browser.click(summary)
+ # The panel is a : it loads what is already on disk when it opens,
+ # so waiting for the directory line proves the application answered.
+ shown = browser.wait_js(
+ f"document.querySelector('{BLOCK}').open === true"
+ f" && !!document.querySelector('{BUTTON}')", timeout=30)
+ checks.record(case, "the backup panel opens and shows its control", shown)
+ listed = browser.wait_js(f"!!document.querySelector('{BLOCK} code')", timeout=30)
+ checks.record(case, "the panel reports where backups are written", listed)
+ return shown
+
+
+def click_back_up_now(browser: Browser, site: Site, db: Path, checks: Checks,
+ case: str, label: str) -> dict:
+ """One press of the button, judged by what the page and the disk then show."""
+ before = set(backups_in(db))
+ # Clear anything still on screen, so the toast this press produces is the
+ # one that is read back rather than a leftover from the previous press.
+ browser.js("""
+ for (const b of document.querySelectorAll('.toast-host button.toast')) b.click();
+ return true;
+ """)
+ button = browser.find(BUTTON, required=False)
+ if button is None:
+ checks.record(case, f"{label}: the control is on the page", False)
+ return {}
+
+ started = time.perf_counter()
+ browser.click(button)
+
+ # Either outcome ends the wait, so a failure is reported as a failure rather
+ # than as a timeout.
+ settled = browser.wait_js(
+ "(() => { const t = [...document.querySelectorAll('.toast-host button.toast')];"
+ " return t.length > 0; })()", timeout=300)
+ elapsed = time.perf_counter() - started
+
+ shown = toasts(browser)
+ errors = [t for t in shown if t["error"]]
+ written = [t for t in shown if t["message"].startswith("Backup written:")]
+
+ checks.record(case, f"{label}: the click produced a visible result", settled,
+ json.dumps(shown)[:200])
+ checks.record(case, f"{label}: the UI reports no error",
+ not errors, json.dumps(errors)[:300])
+ checks.record(case, f"{label}: the UI reports the backup was written",
+ bool(written), json.dumps(written)[:200])
+
+ idle = browser.wait_js(
+ "(() => { const b = document.querySelector(%s);"
+ " return !!b && b.textContent.trim() === 'Back up now'; })()"
+ % json.dumps(BUTTON), timeout=120)
+ checks.record(case, f"{label}: the control returns from 'Backing up…'", idle)
+
+ produced = wait_for_new_backup(db, before)
+ checks.record(case, f"{label}: a backup file was physically produced",
+ produced is not None, str(produced))
+ if produced is None:
+ return {"seconds": round(elapsed, 3), "toasts": shown}
+
+ named = any(produced.name in t["message"] for t in written)
+ checks.record(case, f"{label}: the UI names the file that appeared", named,
+ f"{produced.name} — {written[0]['message'] if written else ''}")
+
+ relisted = browser.wait_js(
+ "[...document.querySelectorAll('%s .backup-list code')]"
+ ".some(c => c.textContent.trim() === %s)" % (BLOCK, json.dumps(produced.name)),
+ timeout=60)
+ checks.record(case, f"{label}: the new backup appears in the panel's list", relisted)
+
+ verdict, integrity_seconds = integrity_of(produced)
+ checks.record(case, f"{label}: the finished copy passes full integrity_check",
+ verdict == "ok", f"{verdict} in {integrity_seconds * 1000:.1f} ms")
+
+ return {
+ "file": str(produced),
+ "bytes": produced.stat().st_size,
+ "sha256_16": digest(produced),
+ "seconds": round(elapsed, 3),
+ "integrity": verdict,
+ "integrity_seconds": round(integrity_seconds, 4),
+ "toasts": shown,
+ }
+
+
+def run_case(case: str, source: Path, out: Path, checks: Checks, *, show: bool,
+ twice: bool) -> dict:
+ print(f"\n=== {case}: {source} ===", flush=True)
+ work = out / case
+ shutil.rmtree(work, ignore_errors=True)
+ (work / "data").mkdir(parents=True)
+ db = work / "data" / "campaign.db"
+ copy_started = time.perf_counter()
+ shutil.copy2(source, db)
+ copy_seconds = time.perf_counter() - copy_started
+ size = db.stat().st_size
+ print(f" copied {size:,} bytes in {copy_seconds:.2f}s -> {db}", flush=True)
+
+ site = Site(BACKEND, db, work / "server.log")
+ browser = Browser(headless=not show, log=work / "geckodriver.log")
+ result: dict = {"database": str(source), "bytes": size,
+ "served_at": site.url, "firefox": browser.version}
+ try:
+ if not open_backup_panel(browser, site, checks, case):
+ return result
+ result["first"] = click_back_up_now(browser, site, db, checks, case, "backup")
+ browser.screenshot(work / "back-up-now.png")
+
+ if twice and result["first"].get("file"):
+ kept = Path(result["first"]["file"])
+ before_bytes, before_digest = kept.stat().st_size, digest(kept)
+ result["second"] = click_back_up_now(browser, site, db, checks, case,
+ "second backup")
+ still_there = kept.exists()
+ checks.record(case, "the earlier backup still exists", still_there)
+ if still_there:
+ checks.record(
+ case, "and is byte-for-byte what it was",
+ kept.stat().st_size == before_bytes and digest(kept) == before_digest,
+ f"{before_bytes:,} bytes, sha256:{before_digest}")
+ verdict, _ = integrity_of(kept)
+ checks.record(case, "and still passes integrity_check", verdict == "ok",
+ verdict)
+ if result["second"].get("file"):
+ checks.record(case, "the second backup is a different file",
+ result["second"]["file"] != result["first"]["file"],
+ Path(result["second"]["file"]).name)
+ finally:
+ browser.quit()
+ site.stop()
+ return result
+
+
+def main() -> int:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--out", required=True,
+ help="evidence directory, which must be under $HOME")
+ parser.add_argument("--case", default="both", choices=["real", "large", "both"])
+ parser.add_argument("--show", action="store_true", help="run Firefox visibly")
+ parser.add_argument("--real-db", default=str(DEFAULT_REAL))
+ parser.add_argument("--large-db", default=str(DEFAULT_LARGE))
+ args = parser.parse_args()
+
+ out = require_under_home(Path(args.out).expanduser())
+ out.mkdir(parents=True, exist_ok=True)
+ if not (BACKEND.parent / "frontend/dist/index.html").exists():
+ print("frontend/dist is not built", file=sys.stderr)
+ return 2
+
+ checks = Checks()
+ started = datetime.now()
+ results: dict[str, dict] = {}
+ wanted = [("real", Path(args.real_db).expanduser(), True)] if args.case != "large" else []
+ if args.case != "real":
+ wanted.append(("large", Path(args.large_db).expanduser(), False))
+
+ print(f"WP-D criterion 3 — the backup through the UI. geckodriver "
+ f"{geckodriver_version()}", flush=True)
+ for case, source, twice in wanted:
+ if not source.exists():
+ checks.record(case, "the database is present", False, str(source))
+ continue
+ results[case] = run_case(case, source, out, checks, show=args.show, twice=twice)
+
+ report = {
+ "started": started.isoformat(timespec="seconds"),
+ "seconds": round((datetime.now() - started).total_seconds()),
+ "cases": results,
+ "checks": checks.rows,
+ "passed": len([r for r in checks.rows if r["result"] == "PASS"]),
+ "failed": len(checks.failed),
+ }
+ (out / "backup-ui-report.json").write_text(json.dumps(report, indent=2))
+ print(f"\n{report['passed']} passed, {report['failed']} failed "
+ f"-> {out / 'backup-ui-report.json'}")
+ return 1 if checks.failed else 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/frontend/src/api.js b/frontend/src/api.js
index 235a893..b9217da 100644
--- a/frontend/src/api.js
+++ b/frontend/src/api.js
@@ -151,7 +151,32 @@ export const api = {
sendAction: (advId, payload, handlers, signal) =>
streamSSE(`/adventures/${advId}/actions`, payload, handlers, signal),
retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal),
- exportAdventure: (id) => request(`/adventures/${id}/export`),
+ // v1.1 WP-D: the bundle, plus what the server says about importing it back.
+ // The body is the bundle and nothing else — the browser saves exactly those
+ // bytes — so the size and the warning come back in headers.
+ exportAdventure: async (id) => {
+ const resp = await fetch(`/api/adventures/${id}/export`, {
+ headers: { 'Content-Type': 'application/json' },
+ })
+ if (!resp.ok) {
+ let detail = resp.statusText
+ try { detail = (await resp.json()).detail || detail } catch { /* non-JSON */ }
+ throw new Error(detail)
+ }
+ const bundle = await resp.json()
+ const number = (name) => {
+ const raw = Number(resp.headers.get(name))
+ return Number.isFinite(raw) && raw > 0 ? raw : null
+ }
+ return {
+ bundle,
+ exportBytes: number('X-Export-Bytes'),
+ importLimitBytes: number('X-Import-Limit-Bytes'),
+ // Absent header (an older server) means nothing is claimed either way.
+ importable: resp.headers.get('X-Importable-By-This-Version') !== 'false',
+ warning: resp.headers.get('X-Export-Warning') || null,
+ }
+ },
importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }),
// M9. A verified copy of the whole database, which is a different tool from
diff --git a/frontend/src/pages/Campaigns.jsx b/frontend/src/pages/Campaigns.jsx
index c797269..112d622 100644
--- a/frontend/src/pages/Campaigns.jsx
+++ b/frontend/src/pages/Campaigns.jsx
@@ -66,10 +66,12 @@ export default function Campaigns() {
const exportOne = async (campaign) => {
try {
- const bundle = await api.exportAdventure(campaign.id)
+ const { bundle, warning } = await api.exportAdventure(campaign.id)
const safe = (campaign.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
downloadJSON(bundle, `${safe}.json`)
- toast('Campaign exported.')
+ // v1.1 WP-D: the file is written either way. A campaign too large for this
+ // version to import back says so now, not when it is needed.
+ toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
} catch (err) {
toast(classifyError(err.message).detail, 'error')
}
diff --git a/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx b/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx
index 045fabd..fb6eb5b 100644
--- a/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx
+++ b/frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx
@@ -79,10 +79,12 @@ export function CampaignSettingsPanel({ adventure, setAdventure, onError, moment
const exportCampaign = async () => {
try {
- const bundle = await api.exportAdventure(adventure.id)
+ const { bundle, warning } = await api.exportAdventure(adventure.id)
const safe = (adventure.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
downloadJSON(bundle, `${safe}.json`)
- toast('Campaign exported.')
+ // v1.1 WP-D: see Campaigns.jsx. The export is delivered; the warning says
+ // this version could not import the file back.
+ toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
} catch (err) {
onError(classifyError(err.message).detail)
}
diff --git a/frontend/src/pages/exportHonesty.test.jsx b/frontend/src/pages/exportHonesty.test.jsx
new file mode 100644
index 0000000..596da63
--- /dev/null
+++ b/frontend/src/pages/exportHonesty.test.jsx
@@ -0,0 +1,167 @@
+/* v1.1 WP-D: an export that says whether this version could import it back.
+ *
+ * The file is delivered either way — a campaign too large to re-import is not a
+ * damaged one, and refusing to write it would destroy the copy the reader was
+ * making. What changes is what they are told, and both reader-facing Export
+ * controls have to tell them: the one on the campaign card and the one in the
+ * campaign's own settings.
+ *
+ * The server decides. These assert that the page shows what it was given and
+ * keeps delivering the file, not that it re-derives the size policy.
+ */
+
+import { screen, waitFor } from '@testing-library/react'
+import userEvent from '@testing-library/user-event'
+import { beforeEach, describe, expect, it, vi } from 'vitest'
+import { api } from '../api'
+import * as components from '../components'
+import Campaigns from './Campaigns'
+import { CampaignSettingsPanel } from './Play/panels/CampaignSettingsPanel'
+import { mockModelStatus, renderWith } from '../test/helpers'
+
+const BUNDLE = { format: 'ai-dnd-adventure-v3', title: 'Long Campaign', actions: [] }
+
+const WARNING =
+ "This export is larger than this version's 20 MB import limit (21,230,000 bytes). "
+ + 'The file was exported successfully, but this version cannot import it.'
+
+const CAMPAIGN = {
+ id: 4, title: 'Long Campaign', action_count: 900,
+ updated_at: '2026-09-15T10:00:00', snippet: 'Rain over the harbour.',
+}
+
+const ADVENTURE = {
+ id: 4, title: 'Long Campaign', ai_instructions: '', narration_length: 'brief',
+ canon_rules: [], persona_name: 'Aldric', persona_desc: '',
+}
+
+// restoreAllMocks does not undo stubGlobal, and the fetch stub below would
+// otherwise outlive its own describe block.
+beforeEach(() => { vi.restoreAllMocks(); vi.unstubAllGlobals() })
+
+function exportReturns({ warning = null } = {}) {
+ return vi.spyOn(api, 'exportAdventure').mockResolvedValue({
+ bundle: BUNDLE,
+ exportBytes: warning ? 21_230_000 : 12_000,
+ importLimitBytes: 20 * 1024 * 1024,
+ importable: !warning,
+ warning,
+ })
+}
+
+async function library() {
+ mockModelStatus(api)
+ vi.spyOn(api, 'listAdventures').mockResolvedValue([CAMPAIGN])
+ await renderWith()
+ await screen.findByText('Long Campaign')
+}
+
+async function settingsPanel() {
+ mockModelStatus(api)
+ await renderWith(
+ ,
+ )
+}
+
+describe('reading what the server said', () => {
+ // The tests below mock api.exportAdventure, so nothing there exercises the
+ // header names. These do: a typo in one of them would otherwise leave the
+ // whole suite green and the reader silently uninformed.
+ function serverSends(headers) {
+ const body = JSON.stringify(BUNDLE)
+ vi.stubGlobal('fetch', vi.fn().mockResolvedValue(new Response(body, {
+ status: 200,
+ headers: { 'Content-Type': 'application/json', ...headers },
+ })))
+ }
+
+ it('reports a warning the server sent, with the sizes it named', async () => {
+ serverSends({
+ 'X-Export-Bytes': '21230000',
+ 'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
+ 'X-Importable-By-This-Version': 'false',
+ 'X-Export-Warning': WARNING,
+ })
+ const result = await api.exportAdventure(4)
+ expect(result.bundle).toEqual(BUNDLE)
+ expect(result.warning).toBe(WARNING)
+ expect(result.importable).toBe(false)
+ expect(result.exportBytes).toBe(21_230_000)
+ expect(result.importLimitBytes).toBe(20 * 1024 * 1024)
+ })
+
+ it('claims nothing when an older server sends no headers', async () => {
+ serverSends({})
+ const result = await api.exportAdventure(4)
+ expect(result.bundle).toEqual(BUNDLE)
+ expect(result.warning).toBeNull()
+ expect(result.importable).toBe(true)
+ expect(result.exportBytes).toBeNull()
+ expect(result.importLimitBytes).toBeNull()
+ })
+
+ it('treats an ordinary export as importable', async () => {
+ serverSends({
+ 'X-Export-Bytes': '12000',
+ 'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
+ 'X-Importable-By-This-Version': 'true',
+ })
+ const result = await api.exportAdventure(4)
+ expect(result.importable).toBe(true)
+ expect(result.warning).toBeNull()
+ expect(result.exportBytes).toBe(12_000)
+ })
+})
+
+describe('the campaign library export', () => {
+ it('delivers the file and says nothing more when it can be imported back', async () => {
+ const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
+ exportReturns()
+ await library()
+ await userEvent.click(screen.getByRole('button', { name: 'Export' }))
+ await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
+ expect(download.mock.calls[0][0]).toEqual(BUNDLE)
+ expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
+ expect(screen.queryByText(/cannot import/)).toBeNull()
+ })
+
+ it('still delivers the file when it is too large, and says so', async () => {
+ const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
+ exportReturns({ warning: WARNING })
+ await library()
+ await userEvent.click(screen.getByRole('button', { name: 'Export' }))
+ // The file is written first: the warning is about importing it back, not
+ // about the export having failed.
+ await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
+ expect(download.mock.calls[0][0]).toEqual(BUNDLE)
+ const notice = await screen.findByText(/cannot import it/)
+ expect(notice).toBeInTheDocument()
+ expect(notice.textContent).toContain('20 MB')
+ expect(notice.textContent).toContain('exported successfully')
+ expect(screen.queryByText('Campaign exported.')).toBeNull()
+ })
+})
+
+describe('the campaign settings export', () => {
+ it('delivers the file and confirms it when it can be imported back', async () => {
+ const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
+ exportReturns()
+ await settingsPanel()
+ await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
+ await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
+ expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
+ expect(screen.queryByText(/cannot import/)).toBeNull()
+ })
+
+ it('still delivers the file when it is too large, and says so', async () => {
+ const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
+ exportReturns({ warning: WARNING })
+ await settingsPanel()
+ await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
+ await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
+ const notice = await screen.findByText(/cannot import it/)
+ expect(notice.textContent).toContain('20 MB')
+ expect(screen.queryByText('Campaign exported.')).toBeNull()
+ })
+})
diff --git a/frontend/src/styles/context.css b/frontend/src/styles/context.css
index 16e13cd..4619ab0 100644
--- a/frontend/src/styles/context.css
+++ b/frontend/src/styles/context.css
@@ -59,7 +59,22 @@
.slice-4 { background: #6f9e8c; }
.slice-5 { background: #c48a6a; }
.slice-6 { background: #7c86b8; }
-.slice-7 { background: var(--border-bright); }
+/* v1.1 WP-E: pinned to the literal value this slice already rendered, instead
+ of borrowing --border-bright. A chart fill and a control edge have different
+ jobs: WP-E raised --border-bright to clear WCAG 1.4.11 (3:1) for control
+ boundaries, and that dragged this slice to #7a7aaa, an OKLab dE of 0.035
+ from .slice-6 (#7c86b8) — two neighbouring slices the same colour. The
+ other slices sit 0.100-0.119 from their nearest neighbour; at #3d3d55 this
+ one sits 0.251, the most separated in the set.
+
+ 1.4.11's 3:1 does not govern this: it is a proportional fill in a labelled
+ breakdown, not the boundary of a control, and what it needs is to be
+ distinguishable from the seven slices beside it. Reassigning it to a freer
+ hue was considered and rejected — inside the palette's own chroma and
+ lightness bands the only hues that beat 0.100 are pinks near 14 degrees,
+ which is --danger's territory and would paint an ordinary prompt section in
+ the colour this application reserves for failure. */
+.slice-7 { background: #3d3d55; }
.ctx-block {
border: 1px solid var(--border);
diff --git a/frontend/src/styles/tokens.css b/frontend/src/styles/tokens.css
index a0d374e..068dfd5 100644
--- a/frontend/src/styles/tokens.css
+++ b/frontend/src/styles/tokens.css
@@ -3,8 +3,15 @@
--bg-panel: #131320;
--bg-panel-glass: rgba(19, 19, 32, 0.82);
--bg-input: #1a1a2a;
- --border: #2b2b3d;
- --border-bright: #3d3d55;
+ /* v1.1 WP-E: control boundaries carry WCAG 1.4.11 (3:1 non-text contrast) on
+ their own, rather than leaning on the control's text label. The floor is
+ measured against --bg-input (#1a1a2a), not --bg-panel: inputs and buttons
+ are drawn on --bg-input (styles/forms.css), and it is the lightest of the
+ three backgrounds a border sits on, so it is the worst case.
+ --border 3.21:1 and --border-bright 4.24:1 there; higher on the others
+ (backend/tools/contrast_audit.py, which now fails below 3:1). */
+ --border: #676792;
+ --border-bright: #7a7aaa;
--text: #e2ddd0;
--text-dim: #918c7d;
--accent: #d4a94e;
diff --git a/planning/README.md b/planning/README.md
index ed6b500..d4f3a18 100644
--- a/planning/README.md
+++ b/planning/README.md
@@ -3,9 +3,13 @@
**This file is the index. Start here.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress:
-WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`) are committed, and
-WP-C, browser release coverage, is complete and staged for owner review**
-(`reports/v1.1/V1.1-WP-C-REPORT.md`).
+WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`), WP-B.2 (`0c1ba83`) and WP-C
+(`59b5ebc`) are committed, and the last two planned packages — WP-D (recovery
+honesty) and WP-E (control-boundary contrast) — are complete and staged for
+owner review** (`reports/v1.1/V1.1-WP-D-REPORT.md`,
+`reports/v1.1/V1.1-WP-E-REPORT.md`). The final browser run passed 101 checks
+across all three suites with 0 failed and 0 skipped. Every planned v1.1 work
+package is now implemented and reported; release validation has not begun.
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
diff --git a/planning/V1.1-PLAN.md b/planning/V1.1-PLAN.md
index 256eb0d..e02fb12 100644
--- a/planning/V1.1-PLAN.md
+++ b/planning/V1.1-PLAN.md
@@ -16,10 +16,16 @@ real-model limitation** (owner decision, 2026-09-15):
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
**WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release
-coverage, is complete and staged for owner review. Its final run passed 91 checks
+coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
-including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E
-have not started. No v1.1 version or tag exists.
+including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`).
+**WP-D** (recovery honesty) and **WP-E**
+(control-boundary contrast) are complete and staged for owner review, each with
+its own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`. The final
+browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0
+failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C,
+WP-D, WP-E) is now implemented and reported; release validation has not begun,
+and no v1.1 version or tag exists.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
diff --git a/planning/VERSION.md b/planning/VERSION.md
index 89e9176..94feaa4 100644
--- a/planning/VERSION.md
+++ b/planning/VERSION.md
@@ -1,8 +1,29 @@
# Planning Package Version
-- **Package:** Adventure Storyteller Planning Package v4.4
+- **Package:** Adventure Storyteller Planning Package v4.5
- **Revision date:** 2026-09-15
-- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. **WP-C, browser release coverage, is complete and staged for owner review: 91 checks, 0 failed, 0 skipped, over trusted-LAN HTTPS, with real export downloads** (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E have not started, and no v1.1 version or tag exists.
+- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C, browser release coverage, is committed and signed as `59b5ebc` (91 checks, 0 failed, 0 skipped, with real export downloads). **WP-D (recovery honesty) and WP-E (control-boundary contrast) are complete and staged for owner review**, each with its own report (`reports/v1.1/V1.1-WP-D-REPORT.md`, `V1.1-WP-E-REPORT.md`); the final browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0 failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) is now implemented and reported. Release validation has not begun, and no v1.1 version or tag exists.
+
+## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15)
+
+The last two planned v1.1 packages, implemented and reported separately. No
+requirement, acceptance test, schema or bundle format changed. WP-D's evidence is
+in `reports/v1.1/V1.1-WP-D-REPORT.md`, WP-E's in `V1.1-WP-E-REPORT.md`.
+
+| Document | Change | Kind |
+| --- | --- | --- |
+| `reports/v1.1/V1.1-WP-D-REPORT.md` | **New.** Full `integrity_check` on the finished backup copy, proved against a fixture `quick_check` calls healthy; an export that says when this version could not import it back, carried in headers because the response body *is* the bundle | work-package report |
+| `reports/v1.1/V1.1-WP-E-REPORT.md` | **New.** Control boundaries raised to clear WCAG 1.4.11 (3:1), the contrast audit turned from a report into a gate, rendered before/after boundary measurements, and the `.slice-7` finding the change itself created | work-package report |
+| `V1.1-PLAN.md` | Status: WP-C signed `59b5ebc`; WP-D and WP-E complete and staged | status |
+| `planning/README.md` | Current state | index |
+| `DEVELOPMENT.md` | How large an export can get: the 20 MB import ceiling, ~13 kB per action, M9's ~279-turn figure and why it is not a turn limit, and the four export headers | developer docs |
+
+**Requirement changes: zero.**
+
+Two owner decisions are outstanding, both recorded rather than assumed: WP-E's
+before/after screenshots await approval (`OWNER SCREENSHOT APPROVAL: PENDING`),
+and WP-D records that no real-browser click was made on *Back up now* — the
+endpoint that button calls was driven instead.
## v4.4 — WP-C browser release coverage (2026-09-15)
diff --git a/planning/reports/v1.1/V1.1-WP-D-REPORT.md b/planning/reports/v1.1/V1.1-WP-D-REPORT.md
new file mode 100644
index 0000000..e1af014
--- /dev/null
+++ b/planning/reports/v1.1/V1.1-WP-D-REPORT.md
@@ -0,0 +1,526 @@
+# v1.1 WP-D — Recovery Honesty
+
+**Status:** COMPLETE — **PASS**. The decision, and what was deliberately not
+claimed, is in §P.
+
+---
+
+## A. Repository baseline
+
+| | |
+| --- | --- |
+| Branch | `v1.1-development` |
+| HEAD at start | `59b5ebc2d85bdf53d38b6bcf347c496dd3432be1` — *v1.1 WP-C: browser release coverage*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
+| Its ancestry | `0c1ba83` (WP-B.2), `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `432f041` (v1.0.0) |
+| Working tree at start | clean; nothing staged |
+| WP-E | not started when WP-D was implemented |
+
+---
+
+## B. Existing backup behaviour
+
+`backend/app/backup.py`, as v1 shipped it. The six questions the brief asks:
+
+| # | Question | Answer (before WP-D) |
+| --- | --- | --- |
+| 1 | How is the backup file created? | SQLite's **online backup API** (`sqlite3.Connection.backup`, `pages=-1`), from the live database opened read-only through a `mode=ro` URI. The destination is a temporary file `….db.partial` **in the destination directory**, so the rename below is atomic |
+| 2 | Where does validation happen? | `_verify()`, on the finished copy, opened as its own read-only connection — not on the source, and not through the connection that wrote it |
+| 3 | When does the final filename appear? | Only after verification: `os.replace(working, target)`. A failed or interrupted run never leaves a file wearing a backup's name |
+| 4 | How are failures cleaned up? | `_discard()` removes the partial file, and `BackupError` is raised with what went wrong. The source is untouched |
+| 5 | Could an existing good backup be overwritten? | **No.** `_unused_name()` stamps each backup with the time and adds a counter if that name (or its `.partial`) exists |
+| 6 | Was the check on the source or the copy? | The copy |
+
+So the mechanism was already right. The one gap was the **strength** of the check:
+`PRAGMA quick_check`, which reads every page and every record but skips the
+cross-check between a table and its indexes.
+
+---
+
+## C. `integrity_check` implementation
+
+One function changed:
+
+```python
+# app/backup.py, _verify()
+rows = connection.execute("PRAGMA integrity_check").fetchall() # was quick_check
+```
+
+- It still runs on the **finished copy**, opened as its own read-only
+ connection, before the rename.
+- A failure still raises `BackupError` naming what was wrong, still discards the
+ partial file, and still leaves the source and every earlier backup untouched.
+- The returned `integrity` field, the filename, the directory, the reported
+ fields (`filename`, `bytes`, `pages`, `seconds`, `integrity`) and the API are
+ unchanged.
+- The module docstring and `_verify`'s docstring now record why the trade M9
+ made (speed over the index cross-check) was not needed at these sizes, with
+ the measurements in §E.
+
+**No scheduled backups were added.** There is still no restore endpoint, no
+retention policy and no timer: a backup happens when the reader asks for one.
+
+---
+
+## D. Corruption fixture
+
+`tests/test_v11_d_recovery.py: build_corrupt_copy()`.
+
+A small database with `t(id, k, filler)` and an index `i_t_k ON t(k)`, 400 rows
+with keys `k000000`…`k000399`. `dbstat` names the index's own leaf pages, and one
+digit inside one indexed key on the first of them is changed (`k000144` →
+`k000944`). Every page stays structurally sound and every record still parses:
+what is broken is only the agreement between the index and its table.
+
+**Proved before it is used as evidence**
+(`test_the_fixture_is_the_difference_between_the_two_checks`):
+
+| Pragma | Result |
+| --- | --- |
+| `PRAGMA quick_check` | **`ok`** |
+| `PRAGMA integrity_check` | **`row 145 missing from index i_t_k`** |
+
+That is the difference the package rests on: not a database both checks reject,
+but one the old check called healthy.
+
+---
+
+## E. Backup performance measurements
+
+Each pragma was run in its **own fresh process**, alternating, three times, because
+a first measurement warms the page cache and a naive ordering makes whichever
+check runs second look faster. (It did: an early single-pass measurement showed
+`integrity_check` at 1.9 ms against `quick_check` at 65.5 ms, purely from cache
+warmth.)
+
+**The real campaign database** — the M11 100-turn evidence campaign,
+`m04-final/campaign.db`, 2,367,488 bytes (2.26 MB):
+
+| | Run 1 | Run 2 | Run 3 |
+| --- | --- | --- | --- |
+| `quick_check` | 5.0 ms | 3.4 ms | 7.6 ms |
+| `integrity_check` | 3.5 ms | 5.7 ms | 5.3 ms |
+
+At this size the two are indistinguishable. A whole backup through
+`backup.create()` — copy, full check and rename — took **70.8 ms** and **27.4 ms**
+on two runs (578 pages).
+
+**An index-heavy synthetic database of 105.2 MB** (105,160,704 bytes; 25,674
+pages of 4,096 bytes), shaped to give the cross-check real work: `actions` with
+152,000 rows and `memories` with 15,200, under three indexes
+(`i_actions_adv_depth`, `i_actions_kind`, `i_memories_adv`):
+
+| | Run 1 | Run 2 | Run 3 |
+| --- | --- | --- | --- |
+| `quick_check` | 172.6 ms | 92.4 ms | 97.1 ms |
+| `integrity_check` | 227.8 ms | 196.0 ms | 227.4 ms |
+
+A whole backup of it through `backup.create()`: **0.58 s** for 25,674 pages,
+`integrity = ok`.
+
+**A campaign-shaped database of 117.4 MB** (117,403,648 bytes; 2,200 actions),
+built through the application's own models so the product path can run against
+it (§E.1):
+
+| | Run 1 | Run 2 | Run 3 |
+| --- | --- | --- | --- |
+| `quick_check` | 89.5 ms | 74.2 ms | 64.4 ms |
+| `integrity_check` | 90.0 ms | 82.9 ms | 75.9 ms |
+
+**Reading.** The cost of the cross-check tracks **rows and index entries, not
+bytes**. On the campaign schema — 2,200 fat rows — the full check costs about
+7 ms more than the quick one at 117 MB. On the index-heavy synthetic — 167,200
+rows under three indexes at a similar size — it costs about 96 ms more. Both sit
+inside a backup of about half a second. There is no performance requirement in
+this project and WP-D does not invent one; the measurements are here because the
+M9 trade was made on a speed argument, and at these sizes that argument does not
+hold.
+
+### E.1 Through the endpoint the button calls — implementation evidence only
+
+**This does not satisfy acceptance criterion 3, and an earlier draft of this
+report wrongly said it did.** The criterion asks for the backup to *complete
+through the UI*; a reader does not call an endpoint. What follows is evidence
+about the implementation — the procedure, its result and its cost — and it is
+kept for that reason. The acceptance evidence is in **§E.2**.
+
+`POST /api/backups` is the only call the **Back up now** button makes, and it
+takes no parameters, so it was driven directly against a **copy** of each
+database — the evidence database is evidence and is not written to:
+
+| Database | Result |
+| --- | --- |
+| Evidence campaign, 2,367,488 bytes | **201** in 0.117 s wall — `pages: 578`, `seconds: 0.038`, `integrity: ok`; `GET /api/backups` then lists 1 backup |
+| Campaign-shaped, 117,403,648 bytes | **201** in 0.536 s wall — `pages: 28,663`, `seconds: 0.452`, `integrity: ok` |
+
+**Why a second large database exists.** The index-heavy synthetic has no
+application schema, and the server runs its migrations at startup, so the app
+will not start against it: driving the endpoint there fails in a migration that
+renumbers `actions`, before any backup is attempted. It remains valid evidence
+for the pragma comparison, which is a property of SQLite and not of this schema,
+but criterion 3 needs a database the product can actually open — hence the
+campaign-shaped one, which §E.2 then drives through the browser.
+
+Artefacts stay under `$HOME` (`v11-evidence/wp-d/`), not in the repository.
+
+---
+
+### E.2 Through the UI — the acceptance evidence for criterion 3
+
+The reader-facing **Back up now** control, clicked in a real Firefox, on the
+production build served by FastAPI: the WP-C path, with no Vite dev server, no
+inference, and no sleep used as an assertion. Every wait is on a condition the
+page or the filesystem can show.
+
+`backend/tools/wpd_backup_ui.py`, run as
+`python -m tools.wpd_backup_ui --out $HOME/v11-evidence/wp-d/ui --case both`.
+**34 checks, 34 passed, 0 failed** —
+`$HOME/v11-evidence/wp-d/ui/backup-ui-report.json`.
+
+The route a reader takes is the route the tool takes: open the application, click
+**Settings** in the navigation, open *Back up everything on this machine* (a
+`` that loads what is on disk when it opens), then press the button.
+
+| | **Case 1 — real campaign database** | **Case 2 — ≥100 MB application database** |
+| --- | --- | --- |
+| Source | the M11 evidence campaign, copied | the campaign-shaped database of §E, copied |
+| **Database size** | **2,367,488 bytes** (2.3 MB) | **117,403,648 bytes** (112.0 MB) |
+| Schema | the application's own, `user_version` 94 | the application's own, `user_version` 94 |
+| **Browser action** | click **Back up now** (twice — see below) | click **Back up now** |
+| **Observable UI result** | toast: *"Backup written: adventure-storyteller-20260915-225000.db (2.3 MB)."*, no error toast, the control returns from *Backing up…*, and the file appears in the panel's list | toast: *"Backup written: adventure-storyteller-20260915-225007.db (112.0 MB)."*, no error toast, control returns, file listed |
+| **Backup path** | `…/ui/real/data/backups/adventure-storyteller-20260915-225000.db` | `…/ui/large/data/backups/adventure-storyteller-20260915-225007.db` |
+| **Backup file size** | **2,367,488 bytes** | **117,403,648 bytes** |
+| **`integrity_check` result** | **`ok`** | **`ok`** |
+| **Integrity-check elapsed** | **5.8 ms** | **78.5 ms** |
+| **Total elapsed backup** | **0.108 s** (click → the UI says it is done) | **0.719 s** |
+
+"Total elapsed" is measured from the click to the rendered result, so it is what
+the reader waits, not what the server reports. The `integrity_check` above is an
+independent second opinion, run here on the artefact the UI produced — the
+application had already verified the copy before keeping it.
+
+**The size the UI reports is the size on disk.** "112.0 MB" in the toast is
+117,403,648 bytes shown in the page's own units; the file matches the source
+database byte for byte.
+
+**Existing-backup protection, established through the UI itself** (Case 1's
+second press, rather than a backup planted by a library call):
+
+| Claim | Result |
+| --- | --- |
+| A second press writes a *different* file — `…-225000-2.db`, the counter `_unused_name` adds when two backups land in the same second | PASS |
+| The first backup still exists afterwards | PASS |
+| And is byte-for-byte what it was — 2,367,488 bytes, `sha256:7d45555566…` | PASS |
+| And still passes `integrity_check` | PASS |
+
+No corruption was fabricated through the browser; rejection behaviour is already
+proved deterministically in §D.
+
+**Supplemental endpoint evidence.** The servers' own logs corroborate that the
+button drove each backup: Case 1 logged two `POST /api/backups → 201 Created`,
+each followed by `GET /api/backups → 200 OK` as the panel reloaded its list;
+Case 2 logged one of each. Screenshots and logs are beside the report JSON.
+
+**A harness defect this run found, in my own check and not in the product.** The
+first pass reported six failures. The toast renders a decorative mark before its
+message — `❖…` —
+so the button's `textContent` begins with `❖`, and my assertion matched
+`textContent.startswith("Backup written:")`. In that same run the UI had shown a
+non-error toast naming the file, the file was on disk, its name was in the
+panel's list, and the copy verified: the product was right and the assertion was
+reading the mark. The check now reads the message span. **No product code was
+changed** — the expected change for this work was none, and none was needed.
+
+---
+
+## F. Existing export/import limit behaviour
+
+| | |
+| --- | --- |
+| Import ceiling | `limits.MAX_IMPORT_BODY_BYTES` = 20 MB, **unchanged by WP-D** |
+| Where it is enforced | `BodySizeLimitMiddleware`, on the declared `Content-Length`, before the body is read |
+| Refusal | HTTP **413**, "Request too large (limit 20 MB)." — it already named the limit |
+| Export before WP-D | `GET /adventures/{id}/export` returned the bundle. Nothing compared its size with the ceiling, so a campaign could be exported and then refused by its own importer |
+
+---
+
+## G. Oversized-export warning implementation
+
+**Where the answer comes from.** `limits.import_limit_label()` and
+`limits.oversized_export_warning(size)` derive the sentence from
+`MAX_IMPORT_BODY_BYTES`. A test changes the constant and asserts the wording
+follows, so no number is written twice.
+
+> This export is larger than this version's 20 MB import limit (21,230,000 bytes).
+> The file was exported successfully, but this version cannot import it.
+
+**Where it travels: headers, not the body.** The export response *is* the bundle —
+the browser saves exactly those bytes as the file — so a warning inside it would
+become part of a portable story file and of every checksum taken over one. The
+route returns the same body with:
+
+```text
+X-Export-Bytes the serialised size
+X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
+X-Importable-By-This-Version true / false
+X-Export-Warning only when false
+```
+
+**Which size is measured.** The compact serialisation (`separators=(",", ":")`,
+`ensure_ascii=False`), which is what this response sends *and* what the browser
+POSTs back on import — the bytes `BodySizeLimitMiddleware` weighs. The
+pretty-printed file the reader downloads is larger and is not what import reads.
+The route serialises once and returns those bytes, so `X-Export-Bytes` is the
+length of the body actually sent.
+
+**The bundle is unchanged.** No key was added to it (§I), and the format stays
+`ai-dnd-adventure-v3`.
+
+---
+
+## H. Oversized export evidence
+
+`test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back`
+writes real rows into a campaign until its bundle genuinely exceeds the ceiling —
+nothing is mocked, and the export serialises all of it.
+
+| Claim | Result |
+| --- | --- |
+| The response body is larger than 20 MB | PASS |
+| It parses, and is `ai-dnd-adventure-v3` with its actions | PASS |
+| `X-Importable-By-This-Version: false` | PASS |
+| `X-Export-Warning` names the limit ("20 MB"), says the export succeeded, and says this version cannot import it | PASS |
+| `X-Export-Bytes` equals the body length | PASS |
+| The same bundle is refused by import, with the limit named (§J) | PASS |
+| A normal campaign's export carries **no** warning, and `X-Importable-By-This-Version: true` | PASS |
+| The warning follows the constant (changed to 50 MB in a test: the sentence says 50 MB, and no longer says 20 MB) | PASS |
+
+---
+
+## I. Normal-export compatibility
+
+The same campaign database (`m04-final/campaign.db`) exported through a worktree
+at `59b5ebc` (before WP-D) and through the WP-D tree:
+
+| | |
+| --- | --- |
+| Before | 2,699,076 bytes |
+| After | 2,699,076 bytes |
+| Comparison | **identical** — the parsed documents compare equal, key for key |
+
+No timestamp allowance was needed: this campaign's bundle carries no field that
+moves between exports. The format is `ai-dnd-adventure-v3` in both.
+
+`test_the_bundle_itself_never_carries_the_warning` additionally asserts that no
+key of an oversized bundle mentions the warning, the limit or importability.
+
+---
+
+## J. Import-refusal behaviour
+
+Unchanged in behaviour, and already naming the limit:
+
+| | |
+| --- | --- |
+| Status | **413**, from the middleware, on `Content-Length`, before the body is parsed |
+| Message | "Request too large (limit 20 MB)." |
+| WP-D test | `test_that_same_bundle_is_refused_by_import_naming_the_limit` posts the *actual oversized export* and asserts 413 and that the detail contains `limits.import_limit_label()` |
+| Not relaxed | the same boundary, the same status, the same middleware. `test_a_bundle_under_the_limit_still_imports` keeps the ordinary path honest (201) |
+
+No wording correction was needed.
+
+---
+
+## K. Frontend behaviour
+
+`api.exportAdventure` now returns the bundle **and** what the server said about
+it (`exportBytes`, `importLimitBytes`, `importable`, `warning`), read from the
+headers.
+
+**The no-header case.** `importable` is `resp.headers.get(…) !== 'false'`, so
+only the literal string `false` is read as a refusal: an older server that sends
+no headers — or a header that arrives malformed — yields `importable: true` and
+`warning: null`, and the page says nothing it was not told. The two size fields
+fall back to `null` unless they parse as a positive number.
+
+This is asserted directly: three tests stub `fetch` with real response headers
+and check what `api.exportAdventure` makes of them — a warning with the sizes it
+named, an ordinary export marked importable, and an older server sending no
+headers at all. They exist because the page-level tests below mock
+`api.exportAdventure` itself and so cannot see a header name, which meant a typo
+on the frontend side would have left the whole suite green (§O.7).
+
+Both reader-facing entry points keep delivering the file first and then report:
+
+| Entry point | Under the limit | Over the limit |
+| --- | --- | --- |
+| Campaign library, **Export** | file written; "Campaign exported." | file written; the warning, as an error-styled toast |
+| Campaign settings, **Export campaign** | file written; "Campaign exported." | file written; the warning |
+
+No new modal, no redesign: the existing toast carries it.
+
+**Tests** (`pages/exportHonesty.test.jsx`, **7**): three read real response
+headers through a stubbed `fetch` (above), and four drive both entry points —
+the file is delivered in every case; the warning is shown when the server sends
+one, naming 20 MB and saying the export succeeded; the ordinary confirmation is
+shown when it does not.
+
+---
+
+## L. Offline regression
+
+`tools/m11_offline.py`, against the production build, in the offline container:
+
+**23 checks, 23 passed, 0 failed** —
+`$HOME/v11-evidence/wp-d/offline/offline-report.json`.
+
+The two this package could have broken are in it and passed:
+
+| Check | Result |
+| --- | --- |
+| a campaign exports offline | ok |
+| and imports offline, with its state | ok |
+| no secret is present in the export | ok |
+| a local file imports offline | ok |
+
+The export route now sets headers and serialises compactly; the offline
+container still exports a campaign and imports it back with its state, so
+neither the round trip nor the secret-scrubbing changed.
+
+## M. Full regression
+
+| Suite | Result |
+| --- | --- |
+| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,712 passed, 17 skipped, 0 failed, 0 xfailed** (936.6 s) |
+| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
+| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
+| **Production build** (`npm run build`) | succeeded |
+| **Offline regression** | 23 passed, 0 failed (§L) |
+
+**The backend count reconciles exactly.** WP-B's closing tree was 1,693; WP-C
+added 7 (`test_v11_c_browser_helpers.py`) and did not rerun the suite because no
+application code changed; WP-D adds the 12 in `test_v11_d_recovery.py`.
+1,693 + 7 + 12 = **1,712**.
+
+**The 17 skips are named, not assumed.** Run with `-rs`, every one is an
+environment-gated real-model test, and none is new:
+
+| File | Skipped | Gate |
+| --- | --- | --- |
+| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
+| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
+| `test_narrative_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
+| `test_m11_real_window.py` | 3 | two the same; one `AIDND_TEST_WIDE_MODEL` |
+| `test_provider_wiring.py` | 1 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
+
+That is the same 17 B.1 and B.2 recorded. WP-D used no model and added no skip.
+
+**The plan's named regression requirements**, run individually rather than
+assumed to be inside the total:
+
+| Requirement | Result |
+| --- | --- |
+| I01-I07, L01-L04 | **25 passed**, 1,704 deselected |
+| The backup case in `test_m11_migration.py` | **1 passed** |
+| `backup.test.jsx` | **7 passed** |
+| The offline container's export and import | passed (§L) |
+
+**Frontend arithmetic.** WP-C's baseline was 168. WP-D adds 4 page-level export
+tests and 3 header-parsing tests: 168 + 7 = **175**. The 15 lint warnings are
+the same pre-existing `only-export-components` and unused-import kind recorded
+at WP-C, and **none is in a file WP-D changed**.
+
+## N. Compatibility
+
+| Area | Effect |
+| --- | --- |
+| Database schema | unchanged; no migration |
+| Bundle format | `ai-dnd-adventure-v3`, unchanged |
+| Normal bundle contents | byte-identical (§I) |
+| Existing backups | still valid files with the same names; only the check that admits a new one is stricter |
+| v1.0.0 databases | open unchanged (no schema or data path changed) |
+| Import limit | unchanged at 20 MB |
+| Scheduled backups | none, as before |
+| History, branches, Save Points, state, memory, knowledge | untouched |
+| Endpoint policy, local-only operation | untouched |
+
+## O. Residual risks
+
+1. **Closed: the backup is now driven through the UI.** This risk previously
+ read "the browser click was not driven for the backup", and the owner
+ correctly refused criterion 3 on endpoint evidence. §E.2 drives the real
+ **Back up now** control in Firefox on both databases — 34 checks, 0 failed —
+ including the ≥100 MB case, which does mean the harness starts a server
+ against a 117 MB database and takes 0.719 s to do it. What remains is
+ ordinary coverage scope: this runs as its own tool rather than inside the
+ release harness, so it is not part of the 101-check run in the WP-E report.
+2. **The 20 MB ceiling is unchanged.** A campaign past it still cannot be
+ imported by this version. WP-D makes that audible at export; raising the
+ limit or streaming import stays v1.2, as the plan assigns it.
+3. **The warning is about *this* version.** It says what this build's importer
+ will accept. A future build with a higher ceiling could import a file this
+ one warned about, and the wording ("this version") is chosen so that stays
+ true rather than becoming a lie.
+4. **`integrity_check` is not a guarantee of recoverability.** It proves the
+ copy's pages, records and indexes agree. A database that was already
+ logically wrong when it was copied is copied faithfully and passes. The
+ package makes the check honest, not omniscient.
+5. **No restore path.** There is still no restore button and no scheduled
+ backup — both explicitly out of scope. A reader who needs a backup restores
+ it by replacing the file themselves, as before.
+6. **The header channel depends on the browser reaching headers.** Both export
+ call sites read them through `fetch`, so a proxy that stripped `X-` headers
+ would silently return to v1 behaviour: the file still arrives, the warning
+ does not. Local-only operation makes that unlikely, and the failure is the
+ old behaviour rather than a wrong claim.
+7. **Header parsing was untested — now closed.** Writing this section exposed
+ it: the page-level tests mock `api.exportAdventure` and the backend tests
+ assert what the server sends, so nothing read an actual header name, and a
+ frontend-side typo would have left every test green and the reader silently
+ uninformed. Three tests that stub `fetch` with real headers now cover it
+ (§K). What remains is the ordinary version of this risk: the two sides agree
+ by matching string literals in two files, and only a browser-level test would
+ catch a mismatch introduced in both at once.
+
+## P. Final decision
+
+**Against the plan's acceptance criteria** (§ WP-D, *Recovery honesty*):
+
+| # | Criterion | Verdict |
+| --- | --- | --- |
+| 1 | A healthy database's backup reports `integrity_check` ok and is kept | **PASS** — `test_a_healthy_backup_passes_the_full_check_and_is_kept`, and both endpoint runs returned `integrity: ok` (§E.1) |
+| 2 | A copy `integrity_check` rejects and `quick_check` does not is rejected and not kept; existing backups still never overwritten | **PASS** — the fixture is proved to be exactly that difference (§D), and three tests cover rejection, the untouched earlier backup, and the unchanged file semantics |
+| 3 | The time is recorded on the evidence database and on a synthetic ≥100 MB, and the backup completes through the UI on both | **PASS** — pragma timings in §E on three databases, and the backup completes **through the UI** on both in §E.2: 0.108 s with `integrity_check` `ok` in 5.8 ms on the 2.3 MB campaign, 0.719 s with `ok` in 78.5 ms on the 117 MB application database. An earlier draft claimed this criterion on endpoint evidence (§E.1); that claim was wrong and is corrected |
+| 4 | An over-20 MB export succeeds, delivers the file, and warns naming the limit, in the API response and a component test; under the limit, no warning | **PASS** — §H, and 7 frontend tests (§K) |
+| 5 | Importing that bundle is refused with a message naming the limit | **PASS** — §J, asserted on the actual oversized export |
+| 6 | A normal export is byte-identical before and after, apart from timestamps | **PASS** — fully identical, no timestamp allowance needed (§I) |
+
+**What I corrected rather than reported around.** Two claims of my own failed
+checking and were fixed, not softened: §E originally rested criterion 3 on
+`backup.create()` calls, and driving the real endpoint showed the index-heavy
+synthetic cannot serve that criterion at all (the app's startup migrations abort
+against a schema-less database) — hence the campaign-shaped database in §E.1.
+And §K asserted the no-header fallback from the code alone, which exposed that
+**no test read any header name**; three tests now do (§O.7).
+
+**Not done, and not claimed:** the import ceiling is unchanged at 20 MB, by the
+plan's own assignment of the raise to v1.2. The UI backup runs as its own tool
+rather than inside the release harness (§O.1).
+
+**No product code changed for the UI verification.** The run exposed one defect,
+and it was in my own assertion, not in the application (§E.2). Backend
+**1,723 passed / 17 skipped / 0 failed** and offline **23/23** therefore stand
+as prior evidence, unchanged and not re-run; the targeted backup and browser
+checks were re-run instead.
+
+```text
+BACKUP INTEGRITY: PASS
+BACKUP THROUGH UI — REAL DB: PASS
+BACKUP THROUGH UI — >=100 MB DB: PASS
+EXPORT HONESTY: PASS
+
+WP-D OVERALL:
+PASS
+```
+
+All WP-D changes are **staged and uncommitted**. No commit, no push, no tag.
+WP-E has not begun.
diff --git a/planning/reports/v1.1/V1.1-WP-E-REPORT.md b/planning/reports/v1.1/V1.1-WP-E-REPORT.md
new file mode 100644
index 0000000..18d4a34
--- /dev/null
+++ b/planning/reports/v1.1/V1.1-WP-E-REPORT.md
@@ -0,0 +1,374 @@
+# v1.1 WP-E — Control-Boundary Contrast
+
+**Status:** COMPLETE — **PASS**, with owner approval of the screenshots
+outstanding. The decision, and what was deliberately not claimed, is in §O.
+
+---
+
+## A. Repository baseline
+
+| | |
+| --- | --- |
+| Branch | `v1.1-development` |
+| HEAD | `59b5ebc` — *v1.1 WP-C: browser release coverage*, signed by the owner |
+| Working tree at start | WP-D staged (10 files), nothing committed |
+| Criterion | WCAG 2.1 **1.4.11 Non-text Contrast**, 3:1, for control boundaries; **1.4.3** 4.5:1 for body text, unchanged |
+
+---
+
+## B. What v1.0.0 actually did
+
+`tools/contrast_audit.py` measured control boundaries, printed that two of them
+were below 3:1, and **exited 0**. Its own comment argued the position:
+
+> in this design a control is identified by its *label*, which is measured above
+> and passes, not by its edge. So a boundary below 3:1 is reported with its
+> number and does not fail the run.
+
+So the audit was a report, not a gate: no palette change could ever fail it on a
+boundary. The two numbers it printed were **1.33:1** (`--border` on
+`--bg-panel`) and **1.75:1** (`--border-bright`), against a floor of 3.0.
+
+**WP-E overturns that argument.** 1.4.11 covers the visual information needed to
+identify a component *and its boundary*; a reader who cannot see where a text box
+ends cannot see that there is a text box to type into, label or no label. The
+tokens were raised rather than the criterion re-argued.
+
+---
+
+## C. Token inventory
+
+| Token | v1.0.0 | v1.1 | Why |
+| --- | --- | --- | --- |
+| `--border` | `#2b2b3d` | **`#676792`** | every control's resting edge |
+| `--border-bright` | `#3d3d55` | **`#7a7aaa`** | hover edges, the composer's resting edge, the open panel tab |
+| `--bg-panel`, `--bg-input`, `--bg`, `--text`, `--text-dim`, `--accent*`, `--danger`, `--warning`, `--player`, `--chart-*` | — | **unchanged** | WP-E is a boundary package; no text or accent colour moved |
+
+**The floor is taken against `--bg-input`, not `--bg-panel`.** Inputs and buttons
+are drawn on `--bg-input` (`styles/forms.css`), which is lighter than
+`--bg-panel` and therefore the harder case. The audit had been checking only
+`--bg-panel`, so a token could have passed the audit while the real control
+failed. Measured on the new values:
+
+| | vs `--bg-input` | vs `--bg-panel` | vs `--bg` |
+| --- | --- | --- | --- |
+| `--border` | **3.21:1** | 3.44:1 | 3.70:1 |
+| `--border-bright` | **4.24:1** | 4.55:1 | 4.88:1 |
+
+Two properties were preserved deliberately: the rest→hover step is the same size
+as before (1.318 → 1.320), so hover still reads as a change rather than a jump;
+and `--border-bright` stays *below* body text against the same panel (2.98:1
+between them), so no edge outshines the words inside it.
+
+---
+
+## D. Component inventory
+
+Where these tokens are actually drawn, from the stylesheets:
+
+| Control | Rule | Rest | Hover / active | Focus |
+| --- | --- | --- | --- | --- |
+| Story composer | `.input-bar` (story.css) | `--border-bright` on `--bg-panel` | — | `--accent-dim` + `--accent-glow` ring |
+| Story controls | `.story-controls button` | `--border` on `--bg-panel` | `--border-bright` | (M11 focus check) |
+| Fields and buttons | `forms.css` | `--border` on `--bg-input` | `--accent-dim` | `--accent-dim` + ring |
+| Panel tabs | `.panel-tabs button` | **`transparent`** | `--border` on hover, `--border-bright` when active | — |
+| Top navigation | `.topnav` | `--border` bottom edge on **`--bg-panel-glass`** | — | — |
+| Scrollbar thumb | `base.css` | `--border-bright` as a *fill* on `--bg` | `--accent-dim` | — |
+
+Two of these cannot be answered by token arithmetic at all, and both are
+measured in the browser instead (§G): the nav sits on a translucent panel, and
+the panel tab's edge is `transparent` until the panel is open.
+
+---
+
+## E. The audit is now a gate
+
+`tools/contrast_audit.py`:
+
+1. **Boundary pairs fail.** `text` and `boundary` rows are both pass/fail; the
+ advisory branch is gone. The run returns 1 if either kind falls short.
+2. **Eight boundary pairs replace two.** Each border is checked against every
+ background it is drawn on — `--bg-input`, `--bg-panel` and `--bg` — plus the
+ focused edge (`--accent-dim`) on both panel and field backgrounds.
+3. **The verdict is taken on the number that is printed** (rounded to two
+ decimals), so a pair shown as `3.00:1` is not failed for arithmetic the
+ reader cannot see.
+4. The comment block that argued the old position is replaced by one recording
+ what changed and why, including why `--bg-panel-glass` is not in the list.
+
+## F. Gate tests
+
+`backend/tests/test_v11_e_contrast.py` — **11 passed**. The threshold is
+exercised from both sides, on real token files:
+
+| Test | Result |
+| --- | --- |
+| A boundary at **2.99:1** against `--bg-input` fails the run (exit 1) | PASS |
+| A boundary at **3.00:1** passes (exit 0) | PASS |
+| The **v1.0.0 value** `#2b2b3d` fails, at 1.24:1 against `--bg-input` | PASS |
+| A dimmed `--text-dim` still fails as a *text* pair | PASS |
+| A renamed token is a failure, not a silent skip | PASS |
+| The shipped palette passes both criteria | PASS |
+| Every boundary pair is measured against the background it is drawn on | PASS |
+| The M11 text baselines are unchanged: 14.57 / 13.57 / 5.48 / 5.88 | PASS |
+| The hover edge stays brighter than the resting edge | PASS |
+| No boundary becomes as loud as body text | PASS |
+| The WCAG ratio formula is anchored on known values (21:1, 1:1, symmetry) | PASS |
+
+The 2.99 and 3.00 values are worth noting: **both clear 3:1 against
+`--bg-panel`** (3.21 and 3.22). They decide the gate only because the floor is
+now taken against the background the control is really on — so these two tests
+also prove §C's change is doing work.
+
+---
+
+## G. Browser measurement
+
+`tools/m11_browser.py` gains a third suite, **WP-E**, counted separately from
+M11's 38 and WP-C's 53. It measures the *rendered* edge — `borderColor` from
+`getComputedStyle` — against what is actually behind it, with every translucent
+layer composited bottom-up.
+
+A boundary is measured against **both** adjacent colours (the control's own fill
+inside it, the background outside it) and passes on the better of the two: an
+edge that matches its fill but contrasts with the page is still a visible
+outline. What 1.4.11 asks is that the component's extent be perceivable.
+
+Two harness capabilities were added for this (`tools/m11_webdriver.py`):
+
+- **`hover()`** moves a real pointer through the WebDriver Actions API.
+ Dispatching a `mouseover` event from JavaScript does *not* trigger CSS
+ `:hover`, so a synthetic event would have re-measured the resting edge and
+ reported it as the hover edge.
+- **`screenshot()`** writes the viewport as a PNG, for the before/after evidence.
+
+**A defect this found in my own first measurement.** The first run reported the
+hover edge as `rgb(114, 114, 160)` and the focused edge as `rgb(144, 120, 81)` —
+neither of which is any token. Both controls carry `transition: border-color
+0.15s`, so the measurement was taken mid-animation, on a colour no state
+actually has. `_settled()` now polls until the computed edge colour is the same
+on two consecutive reads before measuring (polled, not slept, per this harness's
+own rule). After the fix the same edges read exactly `rgb(122, 122, 170)`
+(`--border-bright`) and `rgb(150, 119, 58)` (`--accent-dim`).
+
+---
+
+## H. Before and after, measured in the browser
+
+Both passes were taken the same way — `--only boundaries --no-narrator`, the
+production build — with only `tokens.css` differing. Evidence under
+`$HOME/v11-evidence/wp-e/before/` and `.../after/`.
+
+| Control (state) | Before | After | Floor |
+| --- | --- | --- | --- |
+| Story composer — resting edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
+| Story control — resting edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
+| Open panel tab — resting edge | **1.75:1** FAIL | **4.55:1** pass | 3.0 |
+| Story control — hover edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
+| Top navigation — translucent edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
+| Story composer — focused edge | 4.70:1 pass | 4.70:1 pass | 3.0 |
+
+**Suite result: before 5 passed / 5 failed; after 10 passed / 0 failed / 0
+skipped.**
+
+Two things this table says that a summary would blur:
+
+- **Focus was never the defect.** The focused edge (`--accent-dim`) already
+ cleared 3:1 in v1.0.0 at 4.70:1, and WP-E did not change it. What failed was
+ rest and hover — the states a reader spends all their time in.
+- **The translucent edge is real evidence.** The nav's background composited to
+ `rgb(17, 17, 29)` — `--bg-panel-glass` (rgba 19,19,32 @ 0.82) over
+ `rgb(10, 10, 15)` — not a fallback. That is the case token arithmetic cannot
+ reach, and it moved from 1.43:1 to 3.70:1.
+
+## I. Screenshots
+
+| File | |
+| --- | --- |
+| `before/control-boundaries.png` | 131,175 bytes, 1366×682 |
+| `before/control-boundaries-nav.png` | 63,376 bytes, 1366×682 |
+| `after/control-boundaries.png` | 132,430 bytes, 1366×682 |
+| `after/control-boundaries-nav.png` | 63,541 bytes, 1366×682 |
+
+All four are PNG, 1366×682, taken on the production build through the same
+harness path, differing only in `tokens.css`. The play-page pair shows the
+composer, the story controls and the open panel tab; the nav pair shows the
+translucent top edge on the library route.
+
+## J. A finding this package created and fixed
+
+Raising `--border-bright` broke something that had nothing to do with control
+boundaries. `.slice-7` in the context inspector's token breakdown was painted
+with `var(--border-bright)`, so it followed the token to `#7a7aaa` — an OKLab ΔE
+of **0.035** from `.slice-6` (`#7c86b8`), making two neighbouring chart slices
+effectively the same colour. The other slices sit **0.100–0.119** from their
+nearest neighbour.
+
+`.slice-7` is now pinned to `#3d3d55`, the literal value it already rendered, so
+its appearance is unchanged from v1.0.0 and its separation (ΔE **0.251**) is the
+widest in the set. A chart fill and a control edge have different jobs and should
+not share a token.
+
+Reassigning it to a fresh hue was considered and rejected on evidence: inside the
+palette's own chroma (0.045–0.120) and lightness (0.586–0.804) bands, the only
+hues clearing the set's 0.100 separation floor are pinks near 14°, which is
+`--danger`'s territory. Painting an ordinary prompt section in the colour this
+application reserves for failure would trade an accessibility fix for a semantic
+lie.
+
+*(A first attempt, `#5d7f9e`, was rejected by the same measurement at ΔE 0.061 —
+below every real slice. It is recorded here because it was written into the file
+before it was measured.)*
+
+## K. Text contrast regression
+
+Unchanged, and asserted so in §F: **14.57:1** body text on the page, **13.57:1**
+in a panel, **5.48:1** secondary text in a panel, **5.88:1** on the page. Every
+text pair still clears 1.4.3, and no text token was touched.
+
+## L. Browser release regression
+
+The full harness, all three suites, against a real narrator on the production
+build. Evidence: `$HOME/v11-evidence/wp-e/release/browser-report.json`.
+
+| | |
+| --- | --- |
+| Kind | **`release regression`** — not `partial`, not `development (only …)` |
+| Narrator | `qwen2.5:3b-instruct`, the reference model, over plain HTTP on the LAN GPU host |
+| Browser | Firefox 155.0.1, geckodriver 0.37.1 |
+| Served | FastAPI on loopback, the **built** SPA (`dist` 2026-09-15T21:08:52) |
+| Duration | 71 s |
+| Turns played | **8**, of which `turns_not_clean` **0** and `protocol_shapes_in_narration` **0** |
+
+**Counted per suite, as the brief requires:**
+
+| Suite | Passed | Failed | Skipped |
+| --- | --- | --- | --- |
+| **M11** (the v1 release regression) | **38** | **0** | **0** |
+| **WP-C** (browser release coverage) | **53** | **0** | **0** |
+| **WP-E** (control boundaries) | **10** | **0** | **0** |
+| **Total** | **101** | **0** | **0** |
+
+M11's 38 and WP-C's 53 are unchanged in count and in name: WP-E added a suite
+beside them rather than altering either. The `kind` field is quoted above
+because a fast run invites the question — 71 s for 101 checks including 8
+narrated turns is the GPU host being quick with a 3B model, and the 8 recorded
+turns with no unclean accounting are what rule out narration having been
+skipped.
+
+The WP-E rows in this run are the same ten as the standalone capture in §H,
+re-measured with a narrator present and a full story on the page.
+
+## M. Full regression
+
+| Suite | Result |
+| --- | --- |
+| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,054.7 s) |
+| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
+| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
+| **Production build** (`npm run build`) | succeeded |
+| **Browser harness** (M11 + WP-C + WP-E) | **101 passed, 0 failed, 0 skipped** (§L) |
+| **Contrast audit** (`python tools/contrast_audit.py`) | **exit 0** — every text pair and every boundary pair passes |
+
+**The backend count reconciles exactly.** WP-D's tree was 1,712; WP-E adds the
+11 in `test_v11_e_contrast.py`. 1,712 + 11 = **1,723**. The 17 skips are the same
+environment-gated real-model tests recorded since B.1 — WP-E used no model and
+added no skip.
+
+**The plan's regression requirements for WP-E** were the frontend suite and
+lint, and the harness's accessibility checks. All three pass: 175 and exit 0
+above, and M11's `A11y` rows — accessible names, visible keyboard focus, no
+positive tabindex, nothing revealed only on hover, and the four rendered text
+contrasts — are inside the 38/38 in §L.
+
+**Lint detail.** The 15 warnings are the same pre-existing
+`only-export-components` and unused-import kind recorded at WP-C and WP-D, and
+**none is in a file WP-E changed**. `tokens.css` and `context.css` are not
+flagged.
+
+## N. Residual risks
+
+1. **The screenshots are unapproved.** The measurements say every boundary now
+ clears 3:1; whether the result *looks* right in this design is a judgment
+ the numbers cannot make. Recorded as PENDING in §O, not assumed.
+2. **Five controls are measured in the browser; the rest inherit.** The audit
+ checks token pairs, and the harness measures the composer, a story control,
+ the open panel tab, the nav edge and the focused composer. Every other
+ bordered surface — modals, cards, the knowledge and context panels — draws
+ the same two tokens, so it moves with them, but none is individually
+ measured. A component that overrides a border with a literal colour would
+ not be caught by either check.
+3. **`--bg-panel-glass` is outside the token audit by nature.** It is rgba over
+ a gradient, so no token pair can express it; its edge is covered only by the
+ browser measurement, which runs in the harness rather than in CI.
+4. **Disabled controls are deliberately not measured.** `.story-controls
+ button:disabled` carries `opacity: 0.35`, so a disabled control's rendered
+ edge is dimmer than any value here. WCAG 1.4.11 exempts inactive components,
+ and the harness selects `:not(:disabled)` on purpose — stated so that the
+ exclusion is visible rather than looking like an oversight.
+5. **The chart set was re-checked only where WP-E disturbed it.** `.slice-7`'s
+ separation and colour-blind distance were measured against the other seven
+ (§J); the set as a whole was not re-audited, which is outside this package.
+6. **`a11y.test.jsx` was not extended**, though the plan listed it as likely
+ affected. It asserts structure and names, not colours, and adding colour
+ assertions in jsdom would test the stylesheet's text rather than a rendered
+ result. Boundary contrast is asserted instead where it can be measured: the
+ gate tests (§F) and the browser (§G).
+
+**One risk that turned out not to exist.** §G's rule — a boundary passes on the
+better of its two adjacent colours — was written to avoid failing an edge that
+contrasts with the page but matches its own fill. In the event it never did any
+work: every measured boundary clears 3:1 against **both** neighbours (composer
+4.55/4.88, story control 3.44/3.70, panel tab 4.24/4.55, hover 4.55/4.88, focus
+4.38/4.70, nav 3.51/3.70). The stricter reading would have produced the same
+verdict on every row.
+
+## O. Final decision
+
+**Against the plan's acceptance criteria** (§ WP-E, *Control-boundary contrast*):
+
+| # | Criterion | Verdict |
+| --- | --- | --- |
+| 1 | `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and 0 on the package tree | **PASS** — 2.99:1 exits 1 and 3.00:1 exits 0 (§F); the v1.0.0 value exits 1; the package tree exits 0 |
+| 2 | Every text pair still clears 4.5:1, and the rendered text contrasts do not fall below v1's 14.57 / 5.48 / 13.57 / 5.88 | **PASS** — asserted as exact baselines in §F and re-measured on the rendered page in §L's M11 rows. No text token changed |
+| 3 | The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1; the focus indicator is still visible | **PASS** — composer 4.88:1, story control 3.70:1 at rest and 4.88:1 on hover (§H), and M11's visible-focus check passes in §L |
+| 4 | The owner approves the before-and-after screenshots, and the report records the approval | **PENDING** — not a criterion I can satisfy. The four PNGs are in §I |
+
+**Where I exceeded the criterion, said plainly.** The plan asks for 3:1 "against
+their panel" and names two controls. This package measures against
+**`--bg-input`** as well — the lighter background inputs and buttons are really
+drawn on, and the one that decides the gate — and adds the open panel tab, the
+translucent navigation edge and the focused state. The stricter floor is the
+reason the two threshold tests in §F are decided by `--bg-input` rather than
+`--bg-panel`, where both would have passed.
+
+**What I got wrong and corrected.** Three of my own claims failed checking and
+were fixed rather than softened: the first hover and focus measurements were
+taken mid-transition and reported colours no state has (§G); two contrast
+figures were written into `context.css` before being measured, and were wrong
+(§J); and my first gate check reported the current tokens as failing because my
+harness crashed on a shallow path, not because of any contrast.
+
+**Not done, and not claimed:** owner approval (criterion 4), individual
+measurement of every bordered component (§N.2), and any palette work beyond the
+two boundary tokens and the one chart fill that borrowed from them.
+
+```text
+WP-E CONTRAST GATE: PASS
+WP-E BOUNDARY MEASUREMENT: PASS
+
+WP-E OVERALL:
+PASS, pending owner approval of the screenshots (criterion 4)
+```
+
+All WP-E changes are **staged and uncommitted**. No commit, no push, no tag.
+v1.1 release validation has not begun.
+
+```text
+OWNER SCREENSHOT APPROVAL: PENDING
+```
+
+The before/after screenshots in §I are the evidence for a change a reader judges
+by looking at it. The measurements say every boundary now clears 3:1; whether
+the result looks right in this design is the owner's call, and it is recorded as
+pending rather than assumed.