v1.1: harden recovery and control boundaries

WP-D and WP-E complete the planned v1.1 implementation packages.

WP-D — recovery honesty:
- backups verify the completed copy with PRAGMA integrity_check
- corruption missed by quick_check is detected by the full check
- existing good backups remain protected
- oversized exports are still delivered but declare whether this version can
  import them, while the 20 MB import limit remains unchanged
- backup was exercised through the real browser UI on both the normal campaign
  database and a campaign-shaped database over 100 MB

WP-E — control-boundary contrast:
- interactive control boundaries meet the WCAG 1.4.11 3:1 target
- the contrast audit is now a failing gate rather than an advisory
- rendered browser measurements pass for the composer, controls, tabs and nav
- text contrast and focus visibility remain intact
- owner reviewed and approved the before/after screenshots

Reports:
- planning/reports/v1.1/V1.1-WP-D-REPORT.md
- planning/reports/v1.1/V1.1-WP-E-REPORT.md

All planned v1.1 work packages A-E are now complete. Release validation has not
yet begun.
This commit is contained in:
JesseMarkowitz
2026-09-16 05:37:13 -04:00
parent 59b5ebc2d8
commit 87a40326a2
21 changed files with 2401 additions and 60 deletions
+34
View File
@@ -548,6 +548,40 @@ together.
| Taken from | Export, on a campaign | Settings → *Back up everything on this machine* | | Taken from | Export, on a campaign | Settings → *Back up everything on this machine* |
| Restored by | Import campaign, on the library screen | replacing the database file, below | | Restored by | Import campaign, on the library screen | replacing the database file, below |
### How large an export can get
The importer accepts a request body up to **20 MB**
(`backend/app/limits.py`, `MAX_IMPORT_BODY_BYTES`), and v1.1 does not change it.
What that means for a campaign, measured rather than guessed:
- the M11 evidence campaign came to roughly **13 kB per action** in its bundle;
- M9's conservative estimate from that figure is about **279 turns** before a
bundle approaches the limit.
Both are measurements of particular campaigns, **not a turn limit**. What a
campaign actually weighs depends on how long its turns are, how much imported
knowledge travels with it, and how many attempts each turn kept. A campaign of
400 short turns can be well inside the limit; one of 200 long ones with a large
library may not be.
**v1.1 (WP-D) makes the individual case visible.** Every export reports its own
serialised size and whether this version could import it back:
```text
X-Export-Bytes the bundle's size, as the importer would weigh it
X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
X-Importable-By-This-Version true / false
X-Export-Warning present only when it is false
```
The export always succeeds and the file is always delivered — it is complete and
undamaged; what it exceeds is this version's import ceiling. Both Export
controls show the warning when there is one. The size compared is the compact
serialisation the browser would POST back, which is smaller than the
pretty-printed file on disk.
Raising the limit, or streaming an import past it, is deferred to v1.2.
### Exporting and importing a campaign ### Exporting and importing a campaign
Export is on each campaign in the library, and in the campaign's own Settings Export is on each campaign in the library, and in the campaign's own Settings
+20 -9
View File
@@ -39,8 +39,9 @@ turn is blocked.
never leaves a half-written file wearing a backup's name. `os.replace` is never leaves a half-written file wearing a backup's name. `os.replace` is
atomic on the same filesystem, which is why the temporary sits in the atomic on the same filesystem, which is why the temporary sits in the
destination's own directory rather than in `/tmp`. destination's own directory rather than in `/tmp`.
3. `PRAGMA quick_check` runs against the finished copy, opened as its own 3. `PRAGMA integrity_check` runs against the finished copy, opened as its own
database, before it is renamed. A backup nobody verified is a belief. database, before it is renamed. A backup nobody verified is a belief. v1.1
WP-D made this the full check rather than `quick_check`; see `_verify`.
4. An existing file is never overwritten. Each run writes a new name stamped 4. An existing file is never overwritten. Each run writes a new name stamped
with the time, so yesterday's backup survives today's mistake — which is most with the time, so yesterday's backup survives today's mistake — which is most
of what a backup is for. of what a backup is for.
@@ -189,18 +190,28 @@ def _copy(source_path: Path, working: Path) -> int:
def _verify(working: Path) -> str: def _verify(working: Path) -> str:
"""Runs `PRAGMA quick_check` against the finished copy. """Runs `PRAGMA integrity_check` against the finished copy.
Opened as its own connection, so what is checked is the file on disk rather Opened as its own connection, so what is checked is the file on disk rather
than any page cache the copy left behind. `quick_check` rather than than any page cache the copy left behind.
`integrity_check` because it does the structural work — every page reachable,
every record readable — without the full index cross-check, which on a large **v1.1 WP-D: the full check, not `quick_check`.** M9 chose `quick_check` for
database is minutes rather than moments. A backup nobody verified is a its speed, on the argument that a backup verified slowly enough that nobody
belief; a backup verified slowly enough that nobody takes one is worse. takes one is worse than a fast one. The measurements say the trade was not
needed here: `quick_check` omits the cross-check between a table and its
indexes, and that is a real class of damage it reports as `ok`. A copy whose
index disagrees with its table restores into a database that answers queries
with rows that are not there — the failure a backup exists to prevent.
The cost is small at the sizes this application produces: on the 100-turn
evidence campaign both checks are a few milliseconds, and on a synthetic
database two orders of magnitude larger the difference is still short of a
second (WP-D report §E). A backup nobody verified is a belief; this is the
check that makes it a fact.
""" """
connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True) connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True)
try: try:
rows = connection.execute("PRAGMA quick_check").fetchall() rows = connection.execute("PRAGMA integrity_check").fetchall()
finally: finally:
connection.close() connection.close()
result = ", ".join(str(row[0]) for row in rows) if rows else "no result" result = ", ".join(str(row[0]) for row in rows) if rows else "no result"
+28
View File
@@ -129,6 +129,34 @@ MAX_BODY_BYTES = 2 * 1024 * 1024
MAX_IMPORT_BODY_BYTES = 20 * 1024 * 1024 MAX_IMPORT_BODY_BYTES = 20 * 1024 * 1024
def import_limit_label(limit: int | None = None) -> str:
"""The import ceiling as a reader would say it, e.g. "20 MB".
Derived from the constant rather than written beside it, so the refusal, the
export warning and the documentation cannot drift apart from each other or
from what the middleware actually enforces (v1.1 WP-D).
"""
size = MAX_IMPORT_BODY_BYTES if limit is None else limit
megabytes = size / (1024 * 1024)
return f"{megabytes:.0f} MB" if abs(megabytes - round(megabytes)) < 0.05 else f"{megabytes:.1f} MB"
def oversized_export_warning(export_bytes: int, limit: int | None = None) -> str:
"""What to tell a reader whose export is larger than import will accept.
v1.1 WP-D. The file is written and is not damaged: what it exceeds is this
version's import ceiling, so it cannot be brought back in *here*. Saying that
plainly is the whole point — the alternative is a reader who finds out when
they try to restore it.
"""
size = MAX_IMPORT_BODY_BYTES if limit is None else limit
return (
f"This export is larger than this version's {import_limit_label(size)} import "
f"limit ({export_bytes:,} bytes). The file was exported successfully, but this "
f"version cannot import it."
)
class BodySizeLimitMiddleware: class BodySizeLimitMiddleware:
"""Rejects oversized request bodies by their declared `Content-Length`. """Rejects oversized request bodies by their declared `Content-Length`.
+34 -2
View File
@@ -28,7 +28,9 @@ repair. Refusing a whole campaign because a search index would not build would
trade the valuable thing for the cheap one. trade the valuable thing for the cheap one.
""" """
from fastapi import Body, Depends, Request import json
from fastapi import Body, Depends, Request, Response
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from ... import bundle, head, limits, models, schemas from ... import bundle, head, limits, models, schemas
@@ -46,8 +48,38 @@ def export_adventure(
`app/bundle.py` owns the format, in all three of its versions. A backup `app/bundle.py` owns the format, in all three of its versions. A backup
outlives the schema, so no call site decides anything about its shape. outlives the schema, so no call site decides anything about its shape.
**v1.1 WP-D: the export also says whether this version could import it back.**
A campaign large enough to pass `limits.MAX_IMPORT_BODY_BYTES` still exports —
the file is complete and not damaged, and refusing to write it would destroy
the only copy the reader was trying to make. What it cannot do is come back
in here, and the reader is told that at the moment they take it rather than
at the moment they need it.
It travels in headers, not in the body. The body is the bundle, the browser
saves exactly those bytes as the file, and a warning inside it would become
part of a portable story file and of every checksum taken over one.
The size measured is the compact serialisation, because that is both what
this response sends and what the browser POSTs back on import, which is what
`BodySizeLimitMiddleware` weighs. The pretty-printed file the reader
downloads is larger, and is not what import reads.
""" """
return bundle.export(db, adv) payload = bundle.export(db, adv)
# Serialised exactly as Starlette's JSONResponse would, so the bytes counted
# are the bytes sent.
body = json.dumps(payload, ensure_ascii=False, allow_nan=False,
separators=(",", ":")).encode("utf-8")
limit = limits.MAX_IMPORT_BODY_BYTES
importable = len(body) <= limit
headers = {
"X-Export-Bytes": str(len(body)),
"X-Import-Limit-Bytes": str(limit),
"X-Importable-By-This-Version": "true" if importable else "false",
}
if not importable:
headers["X-Export-Warning"] = limits.oversized_export_warning(len(body), limit)
return Response(content=body, media_type="application/json", headers=headers)
@router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201) @router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201)
+294
View File
@@ -0,0 +1,294 @@
"""v1.1 WP-D: a backup that was really checked, and an export that says what it is.
Two recovery-path claims, each of which was true only in the small before this
package:
- **A backup is verified.** M9 ran `PRAGMA quick_check` on the finished copy.
That reads every page and every record, and skips the cross-check between a
table and its indexes — so a copy whose index disagrees with its table passed.
`test_the_fixture_is_the_difference_between_the_two_checks` builds exactly that
damage and shows the two pragmas disagreeing about it, before anything here
uses it as evidence.
- **An export is importable.** Nothing compared the bundle with
`limits.MAX_IMPORT_BODY_BYTES`, so a campaign could be exported and then
refused by its own importer, with the reader finding out at the moment they
needed it. The export still succeeds — the file is complete, and a version
that refused to write it would destroy the copy someone was trying to make —
and now it says so.
python -m pytest tests/test_v11_d_recovery.py -v
"""
import json
import sqlite3
from pathlib import Path
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, backup, limits, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
@pytest.fixture()
def client():
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="wp-d@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(user_id=user.id, model="test-model"))
adventure = models.Adventure(user_id=user.id, title="Recovery")
setup.add(adventure)
setup.flush()
setup.add(models.Action(adventure_id=adventure.id, type="start", text="The story opens."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id))
test_client = TestClient(app)
test_client.adv_id = adv_id
test_client.user_id = user_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
Base.metadata.drop_all(bind=engine)
# ------------------------------------------------------------ the fixture
def build_corrupt_copy(path: Path) -> None:
"""A database whose index disagrees with its table, and nothing else.
One digit inside one index leaf page is changed, so that entry names a key
no row holds and one row's key is in no index entry. Every page is still
structurally sound and every record still parses, which is the whole point:
this is the damage `quick_check` is not looking for.
"""
path.unlink(missing_ok=True)
connection = sqlite3.connect(path)
connection.execute("PRAGMA page_size=4096")
connection.execute("CREATE TABLE t (id INTEGER PRIMARY KEY, k TEXT NOT NULL, filler TEXT)")
connection.execute("CREATE INDEX i_t_k ON t(k)")
connection.executemany("INSERT INTO t (k, filler) VALUES (?, ?)",
[(f"k{n:06d}", "x" * 40) for n in range(400)])
connection.commit()
page_size = connection.execute("PRAGMA page_size").fetchone()[0]
leaves = [row[0] for row in connection.execute(
"SELECT pageno FROM dbstat WHERE name='i_t_k' AND pagetype='leaf' ORDER BY pageno")]
connection.close()
assert leaves, "the index must have a leaf page to damage"
raw = bytearray(path.read_bytes())
start = (leaves[0] - 1) * page_size
page = raw[start:start + page_size]
at = page.find(b"k000")
assert at != -1, "expected an indexed key on the index's first leaf page"
page[at + 4] = ord("9") # k000144 -> k000944: a key no row has
raw[start:start + page_size] = page
path.write_bytes(bytes(raw))
def check(path: Path, pragma: str) -> str:
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
return ", ".join(str(row[0]) for row in connection.execute(f"PRAGMA {pragma}").fetchall())
finally:
connection.close()
def test_the_fixture_is_the_difference_between_the_two_checks(tmp_path):
"""Before using it as evidence: quick_check calls this database fine."""
damaged = tmp_path / "damaged.db"
build_corrupt_copy(damaged)
assert check(damaged, "quick_check") == "ok"
integrity = check(damaged, "integrity_check")
assert integrity != "ok"
assert "i_t_k" in integrity # it names the index that disagrees
# ---------------------------------------------------------------- backups
def test_a_healthy_backup_passes_the_full_check_and_is_kept(client, tmp_path):
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
result = backup.create(source)
assert result.integrity == "ok"
assert result.path.exists() and result.bytes > 0
assert check(result.path, "integrity_check") == "ok"
assert result.path.parent == backup.directory(source)
def test_the_backup_runs_the_full_check_not_the_quick_one(client, tmp_path, monkeypatch):
"""The pragma itself, named. SQLite traces every statement it executes, so
this reads what the backup actually asked the copy rather than inferring it."""
asked: list[str] = []
real_connect = sqlite3.connect
def tracing(*args, **kwargs):
connection = real_connect(*args, **kwargs)
connection.set_trace_callback(asked.append)
return connection
monkeypatch.setattr(backup.sqlite3, "connect", tracing)
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
backup.create(source).path.unlink()
assert any("integrity_check" in sql for sql in asked), asked
assert not any("quick_check" in sql for sql in asked), asked
def test_a_copy_the_full_check_rejects_is_not_kept(client, tmp_path, monkeypatch):
"""The copy is damaged after it is written and before it is verified, which
is where a real page-level fault would appear: between the copy and the
rename. Nothing wearing a backup's name may be left behind."""
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
real_copy = backup._copy
def damage(source_path, working):
pages = real_copy(source_path, working)
build_corrupt_copy(working)
return pages
monkeypatch.setattr(backup, "_copy", damage)
with pytest.raises(backup.BackupError) as refused:
backup.create(source)
assert "did not verify" in str(refused.value)
assert "i_t_k" in str(refused.value) # it says what was wrong
kept = list(backup.directory(source).glob("*"))
assert kept == [], f"a rejected backup was left behind: {kept}"
def test_a_rejected_backup_leaves_an_earlier_good_one_alone(client, tmp_path, monkeypatch):
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
good = backup.create(source)
before = good.path.read_bytes()
real_copy = backup._copy
def damage(source_path, working):
pages = real_copy(source_path, working)
build_corrupt_copy(working)
return pages
monkeypatch.setattr(backup, "_copy", damage)
with pytest.raises(backup.BackupError):
backup.create(source)
assert good.path.exists()
assert good.path.read_bytes() == before
assert check(good.path, "integrity_check") == "ok"
assert [p.name for p in backup.directory(source).glob("*")] == [good.path.name]
def test_the_backup_file_semantics_are_unchanged(client, tmp_path):
"""Same directory, same stamped name, same reported fields: WP-D changed the
check, not the file."""
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
first = backup.create(source)
second = backup.create(source)
assert first.path.name.startswith(backup.PREFIX) and first.path.suffix == ".db"
assert first.path != second.path, "an existing backup is never overwritten"
assert set(first.as_dict()) == {"filename", "bytes", "pages", "seconds", "integrity"}
listed = [row["filename"] for row in backup.existing(source)]
assert sorted(listed) == sorted([first.path.name, second.path.name])
# ----------------------------------------------------------------- exports
def export(client, adv_id):
response = client.get(f"/api/adventures/{adv_id}/export")
assert response.status_code == 200, response.text[:200]
return response
def test_a_normal_export_carries_no_warning(client):
response = export(client, client.adv_id)
assert "X-Export-Warning" not in response.headers
assert response.headers["X-Importable-By-This-Version"] == "true"
assert int(response.headers["X-Import-Limit-Bytes"]) == limits.MAX_IMPORT_BODY_BYTES
assert int(response.headers["X-Export-Bytes"]) == len(response.content)
assert response.json()["format"] == "ai-dnd-adventure-v3"
def fill_past_the_limit(adv_id: int) -> int:
"""Real rows, until the campaign's bundle is genuinely over the ceiling.
Not a mocked size: the export below serialises all of it.
"""
chunk = "The rain kept on over the harbour road, and nobody came. " * 900 # ~50 kB
written = 0
with SessionLocal() as db:
while written < limits.MAX_IMPORT_BODY_BYTES + 2 * 1024 * 1024:
db.add_all([models.Action(adventure_id=adv_id, type="ai", text=chunk)
for _ in range(40)])
db.commit()
written += 40 * len(chunk)
return written
def test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back(client):
fill_past_the_limit(client.adv_id)
response = export(client, client.adv_id)
# Delivered, whole, and still the same format.
body = response.content
assert len(body) > limits.MAX_IMPORT_BODY_BYTES
parsed = json.loads(body)
assert parsed["format"] == "ai-dnd-adventure-v3"
assert len(parsed["actions"]) > 40
# And honest about what this version can do with it.
assert response.headers["X-Importable-By-This-Version"] == "false"
warning = response.headers["X-Export-Warning"]
assert limits.import_limit_label() in warning
assert "exported successfully" in warning
assert "cannot import" in warning
assert int(response.headers["X-Export-Bytes"]) == len(body)
def test_the_warning_follows_the_configured_limit(monkeypatch):
"""The text is generated from the constant, so changing the constant changes
the sentence rather than leaving a stale number in it."""
assert "20 MB" in limits.oversized_export_warning(21_000_000)
monkeypatch.setattr(limits, "MAX_IMPORT_BODY_BYTES", 50 * 1024 * 1024)
assert limits.import_limit_label() == "50 MB"
assert "50 MB" in limits.oversized_export_warning(60_000_000)
assert "20 MB" not in limits.oversized_export_warning(60_000_000)
def test_the_bundle_itself_never_carries_the_warning(client):
"""The warning is about the export, not part of the portable story file."""
fill_past_the_limit(client.adv_id)
parsed = json.loads(export(client, client.adv_id).content)
flat = json.dumps(parsed).lower()
assert "import limit" not in flat
assert "cannot import" not in flat
for key in parsed:
assert "warning" not in key.lower()
def test_that_same_bundle_is_refused_by_import_naming_the_limit(client):
fill_past_the_limit(client.adv_id)
body = export(client, client.adv_id).content
response = client.post("/api/adventures/import", content=body,
headers={"Content-Type": "application/json"})
assert response.status_code == 413
detail = response.json()["detail"]
assert "too large" in detail.lower()
assert limits.import_limit_label() in detail
def test_a_bundle_under_the_limit_still_imports(client):
"""The refusal is about size alone: the ordinary path is untouched."""
body = export(client, client.adv_id).content
assert len(body) < limits.MAX_IMPORT_BODY_BYTES
response = client.post("/api/adventures/import", content=body,
headers={"Content-Type": "application/json"})
assert response.status_code == 201, response.text[:200]
+152
View File
@@ -0,0 +1,152 @@
"""v1.1 WP-E: the contrast audit is a gate, not a report.
Before this package `tools/contrast_audit.py` measured control boundaries,
printed that two of them were below 3:1, and exited 0 anyway — on the argument
that a control is identified by its label rather than its edge. A check that
cannot fail is not a check, and these tests are what make it one: the threshold
is exercised from both sides, on a real tokens file, so a future palette change
that dims a control's edge stops the run instead of adding a line to it.
python -m pytest tests/test_v11_e_contrast.py -v
"""
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tools"))
import contrast_audit as audit # noqa: E402
# --------------------------------------------------------------- the maths
def test_the_ratio_is_the_wcag_ratio():
"""Anchored on values with known answers, so a broken formula is visible."""
assert audit.ratio("#ffffff", "#000000") == pytest.approx(21.0, abs=0.01)
assert audit.ratio("#ffffff", "#ffffff") == pytest.approx(1.0, abs=0.001)
# Order does not matter: contrast is symmetric.
assert audit.ratio("#131320", "#676792") == pytest.approx(
audit.ratio("#676792", "#131320"), abs=1e-9)
# ------------------------------------------------------- the gate itself
def tokens_file(tmp_path: Path, **overrides: str) -> Path:
"""A real tokens file with named colours replaced."""
source = audit.TOKENS.read_text()
for name, value in overrides.items():
token = "--" + name.replace("_", "-")
start = source.index(f"{token}: ")
end = source.index(";", start)
source = source[:start] + f"{token}: {value}" + source[end:]
written = tmp_path / "tokens.css"
written.write_text(source)
return written
def run_against(path: Path, monkeypatch) -> int:
monkeypatch.setattr(audit, "TOKENS", path)
return audit.main()
def test_a_boundary_below_three_to_one_fails_the_run(tmp_path, monkeypatch, capsys):
"""2.99:1 against --bg-input — the wrong side of the line by one hundredth.
This value clears 3:1 against --bg-panel (3.21:1), so it would have passed
the audit as M11 wrote it. It fails now because the floor is taken against
the background the control is actually drawn on.
"""
below = tokens_file(tmp_path, border="#58639a")
assert run_against(below, monkeypatch) == 1
printed = capsys.readouterr().out
assert "2.99:1" in printed
assert "FAIL" in printed
assert "below 3:1 (WCAG 1.4.11)" in printed
def test_a_boundary_at_three_to_one_passes(tmp_path, monkeypatch, capsys):
"""3.00:1 — the right side of the same line, one hundredth from the value
above, and not failed for arithmetic the reader cannot see."""
at = tokens_file(tmp_path, border="#58639b")
assert run_against(at, monkeypatch) == 0
printed = capsys.readouterr().out
assert "3.00:1" in printed
assert "every control boundary clears 3:1" in printed
def test_the_v1_0_0_boundary_would_now_fail(tmp_path, monkeypatch, capsys):
"""The value v1.0.0 shipped. This is the defect WP-E closes, and the gate
has to be the thing that would have caught it."""
shipped = tokens_file(tmp_path, border="#2b2b3d")
assert run_against(shipped, monkeypatch) == 1
assert "1.24:1" in capsys.readouterr().out # against --bg-input, the worst case
def test_a_text_pair_below_four_point_five_still_fails(tmp_path, monkeypatch, capsys):
"""WP-E raised the boundary floor and must not have lowered the text one."""
dimmed = tokens_file(tmp_path, text_dim="#5a5750")
assert run_against(dimmed, monkeypatch) == 1
assert "below WCAG AA (1.4.3)" in capsys.readouterr().out
def test_a_missing_token_is_a_failure_not_a_skip(tmp_path, monkeypatch, capsys):
source = audit.TOKENS.read_text().replace("--border-bright:", "--border-was-renamed:")
written = tmp_path / "tokens.css"
written.write_text(source)
assert run_against(written, monkeypatch) == 1
assert "MISSING TOKEN" in capsys.readouterr().out
# ------------------------------------------------- the shipped palette
def test_the_real_tokens_pass_both_criteria(capsys):
"""The palette as it stands, through the same gate CI would run."""
assert audit.main() == 0
printed = capsys.readouterr().out
assert "every text pair clears WCAG AA (1.4.3)" in printed
assert "every control boundary clears 3:1 (1.4.11)" in printed
assert "FAIL" not in printed
def test_every_boundary_pair_is_measured_against_the_background_it_is_drawn_on():
"""The audit used to check borders only against --bg-panel, which is not
where the bordered controls are: inputs and buttons sit on --bg-input
(styles/forms.css), which is lighter and therefore harder. Checking only the
easier background would let a token pass while the real control failed."""
boundary = [(fg, bg) for kind, fg, bg, _, _ in audit.PAIRS if kind == "boundary"]
for background in ("--bg-input", "--bg-panel", "--bg"):
assert ("--border", background) in boundary, background
assert ("--border-bright", "--bg-input") in boundary
def test_the_text_baselines_are_unchanged_by_wp_e():
"""WP-E changed only boundary tokens. These are the M11 text measurements,
and they have to still be exactly what the earlier reports recorded."""
tokens = audit.read_tokens(audit.TOKENS)
measured = {
"body text on the page": audit.ratio(tokens["--text"], tokens["--bg"]),
"body text in a panel": audit.ratio(tokens["--text"], tokens["--bg-panel"]),
"secondary text in a panel": audit.ratio(tokens["--text-dim"], tokens["--bg-panel"]),
"secondary text on the page": audit.ratio(tokens["--text-dim"], tokens["--bg"]),
}
assert measured["body text on the page"] == pytest.approx(14.57, abs=0.01)
assert measured["body text in a panel"] == pytest.approx(13.57, abs=0.01)
assert measured["secondary text in a panel"] == pytest.approx(5.48, abs=0.01)
assert measured["secondary text on the page"] == pytest.approx(5.88, abs=0.01)
def test_the_hover_edge_stays_brighter_than_the_resting_edge():
"""Rest and hover have to remain distinguishable from each other, not merely
each clear the floor against the background."""
tokens = audit.read_tokens(audit.TOKENS)
panel = tokens["--bg-panel"]
assert audit.ratio(tokens["--border-bright"], panel) > audit.ratio(tokens["--border"], panel)
def test_a_boundary_never_becomes_as_loud_as_body_text():
"""A control's edge that outshines the words inside it is its own defect."""
tokens = audit.read_tokens(audit.TOKENS)
panel = tokens["--bg-panel"]
assert audit.ratio(tokens["--border-bright"], panel) < audit.ratio(tokens["--text"], panel)
+52 -31
View File
@@ -26,16 +26,24 @@ TOKENS = Path(__file__).resolve().parent.parent.parent / "frontend/src/styles/to
#: rather than combinatorial, because "every colour against every other" reports #: rather than combinatorial, because "every colour against every other" reports
#: pairs that never meet on screen. #: pairs that never meet on screen.
#: #:
#: The `kind` matters and is not a way of grading on a curve. **text** pairs are #: The `kind` records which success criterion a pair is measured against —
#: WCAG 1.4.3 Contrast (Minimum) and are what §21 of the M11 brief asks about; #: **text** is WCAG 1.4.3 Contrast (Minimum), **boundary** is WCAG 1.4.11
#: they are pass/fail. **boundary** pairs are WCAG 1.4.11 Non-text Contrast, #: Non-text Contrast — and **both are pass/fail**.
#: which applies to "visual information required to identify user interface #:
#: components" — and in this design a control is identified by its *label*, #: v1.1 WP-E overturned the earlier position here, which was that a boundary
#: which is measured above and passes, not by its edge. So a boundary below 3:1 #: below 3:1 could be recorded rather than failed because "a control is
#: is reported with its number and does not fail the run; what it would take to #: identified by its label, not by its edge". That argument understates what
#: turn it into a real failure is a control with no visible label, and there is #: 1.4.11 asks: the criterion covers the visual information needed to identify
#: no such control (`tools/m11_browser.py` asserts every visible control has an #: a component *and its boundary*, and a reader who cannot see where a text box
#: accessible name, and the story controls are text buttons). #: ends cannot see that there is a text box to type into, label or no label.
#: The edges were at 1.33:1 and 1.75:1 — the tokens were raised instead.
#:
#: Borders are measured against **every background they are drawn on**, and the
#: floor is the worst of them. Inputs and buttons sit on --bg-input, which is
#: lighter than --bg-panel and so the harder case; checking only --bg-panel
#: would have let a token pass the audit while the real control failed.
#: `--bg-panel-glass` is translucent and cannot be resolved from tokens alone;
#: that edge is measured on the rendered page by `tools/m11_browser.py`.
PAIRS = [ PAIRS = [
("text", "--text", "--bg", 4.5, "body text on the page"), ("text", "--text", "--bg", 4.5, "body text on the page"),
("text", "--text", "--bg-panel", 4.5, "body text in a panel"), ("text", "--text", "--bg-panel", 4.5, "body text in a panel"),
@@ -48,8 +56,14 @@ PAIRS = [
("text", "--danger", "--bg-panel", 4.5, "an error message"), ("text", "--danger", "--bg-panel", 4.5, "an error message"),
("text", "--warning", "--bg-panel", 4.5, "a caution message"), ("text", "--warning", "--bg-panel", 4.5, "a caution message"),
("text", "--player", "--bg", 4.5, "the player's own words"), ("text", "--player", "--bg", 4.5, "the player's own words"),
("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge"), ("boundary", "--border", "--bg-input", 3.0, "a field or button's resting edge"),
("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge"), ("boundary", "--border-bright", "--bg-input", 3.0, "a field or button's hover edge"),
("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge in a panel"),
("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge in a panel"),
("boundary", "--border", "--bg", 3.0, "a divider on the page"),
("boundary", "--border-bright", "--bg", 3.0, "the scrollbar thumb"),
("boundary", "--accent-dim", "--bg-input", 3.0, "a focused field's edge"),
("boundary", "--accent-dim", "--bg-panel", 3.0, "a focused control's edge in a panel"),
("boundary", "--chart-1", "--bg-panel", 3.0, "a chart bar"), ("boundary", "--chart-1", "--bg-panel", 3.0, "a chart bar"),
("boundary", "--chart-2", "--bg-panel", 3.0, "a chart bar"), ("boundary", "--chart-2", "--bg-panel", 3.0, "a chart bar"),
("boundary", "--chart-3", "--bg-panel", 3.0, "a chart bar"), ("boundary", "--chart-3", "--bg-panel", 3.0, "a chart bar"),
@@ -81,37 +95,44 @@ def ratio(a: str, b: str) -> float:
def main() -> int: def main() -> int:
tokens = read_tokens(TOKENS) tokens = read_tokens(TOKENS)
print(f"{TOKENS.relative_to(TOKENS.parents[3])}: {len(tokens)} colour tokens\n") # Named defensively: the tests run this against a temporary tokens file,
# which need not sit four directories deep the way the real one does.
label = TOKENS.name
if len(TOKENS.parents) > 3:
label = TOKENS.relative_to(TOKENS.parents[3])
print(f"{label}: {len(tokens)} colour tokens\n")
print(f"{'pair':44} {'kind':9} {'ratio':>7} {'floor':>6} verdict") print(f"{'pair':44} {'kind':9} {'ratio':>7} {'floor':>6} verdict")
print("-" * 82) print("-" * 82)
failures, advisories = 0, 0 text_failures, boundary_failures = 0, 0
for kind, foreground, background, floor, description in PAIRS: for kind, foreground, background, floor, description in PAIRS:
if foreground not in tokens or background not in tokens: if foreground not in tokens or background not in tokens:
print(f"{description:44} {kind:9} {'—':>7} {floor:>6.1f} MISSING TOKEN") print(f"{description:44} {kind:9} {'—':>7} {floor:>6.1f} MISSING TOKEN")
failures += 1 text_failures += 1
continue continue
measured = ratio(tokens[foreground], tokens[background]) measured = ratio(tokens[foreground], tokens[background])
ok = measured >= floor # Rounded to the two decimals printed, so the verdict matches what the
if not ok: # reader is shown: a pair reported as 3.00:1 is not failed for arithmetic
if kind == "text": # the output does not display.
failures += 1 if round(measured, 2) >= floor:
verdict = "FAIL"
else:
advisories += 1
verdict = "below 1.4.11 (label carries it)"
else:
verdict = "pass" verdict = "pass"
else:
verdict = "FAIL"
if kind == "text":
text_failures += 1
else:
boundary_failures += 1
print(f"{description:44} {kind:9} {measured:>6.2f}:1 {floor:>6.1f} {verdict}") print(f"{description:44} {kind:9} {measured:>6.2f}:1 {floor:>6.1f} {verdict}")
print() print()
if failures: if text_failures:
print(f"{failures} text pair(s) below WCAG AA — this is a defect") print(f"{text_failures} text pair(s) below WCAG AA (1.4.3) — this is a defect")
else: else:
print("every text pair clears WCAG AA (1.4.3)") print("every text pair clears WCAG AA (1.4.3)")
if advisories: if boundary_failures:
print(f"{advisories} boundary pair(s) below 3:1 (1.4.11). Recorded rather " print(f"{boundary_failures} boundary pair(s) below 3:1 (WCAG 1.4.11) — "
"than failed: every control in this design carries a visible text " "this is a defect")
"label, which is measured above and passes.") else:
return 1 if failures else 0 print("every control boundary clears 3:1 (1.4.11)")
return 1 if (text_failures or boundary_failures) else 0
if __name__ == "__main__": if __name__ == "__main__":
+229 -2
View File
@@ -665,6 +665,228 @@ def check_accessibility(browser: Browser, site: Site, adv: int, checks: Checks):
checks.record("A11y", "the story input takes keyboard focus", typed is True) checks.record("A11y", "the story input takes keyboard focus", typed is True)
# -------------------------------------------------------------- WP-E boundaries
#: The controls whose edges WP-E measures, as the stylesheets actually draw them.
#: `side` is the border the rule sets: .topnav paints only a bottom edge.
BOUNDARY_TARGETS = [
("E01", "the story composer", ".input-bar", "Top"),
("E02", "a story control", ".story-controls button:not(:disabled)", "Top"),
("E04", "the open panel tab", ".panel-tabs button.active", "Top"),
]
#: `.topnav` is measured on the library route, because the play route does not
#: render it — the play page has its own `.play-header`, which story.css keeps
#: deliberately opaque. So the translucent case exists on exactly one screen,
#: and it is the case token arithmetic cannot answer: --bg-panel-glass is
#: rgba(...,0.82) over a gradient, so only the rendered page knows what is
#: behind that edge.
TRANSLUCENT_TARGET = ("E03", "the top navigation's edge", ".topnav", "Bottom")
#: Measured in the browser rather than from tokens, because two of these cannot
#: be derived from tokens at all: .topnav sits on --bg-panel-glass, which is
#: translucent, so its effective background is a composite of what is behind it;
#: and a rendered edge can be changed by opacity, a transition mid-flight, or a
#: rule the token file knows nothing about.
#:
#: A boundary is measured against **both** adjacent colours — the control's own
#: fill inside it and the background outside it — and passes on the better of
#: the two. An edge that matches its fill but contrasts with the page is still a
#: visible outline, and vice versa; what 1.4.11 asks is that the component's
#: extent be perceivable, not that every neighbouring surface differ from it.
BOUNDARY_JS = """
const parse = (c) => (c.match(/[\\d.]+/g) || []).map(Number);
const alphaOf = (c) => {
if (!c || c === 'transparent') return 0;
const p = parse(c);
return p.length > 3 ? p[3] : 1;
};
function lum(c) {
const [r, g, b] = parse(c).slice(0, 3).map(v => v / 255)
.map(v => v <= 0.03928 ? v / 12.92 : Math.pow((v + 0.055) / 1.055, 2.4));
return 0.2126 * r + 0.7152 * g + 0.0722 * b;
}
function over(fg, bg) {
const f = parse(fg), b = parse(bg), a = alphaOf(fg);
return 'rgb(' + [0, 1, 2].map(i => Math.round(f[i] * a + b[i] * (1 - a))).join(', ') + ')';
}
function ratio(x, y) {
const a = lum(x), b = lum(y);
return (Math.max(a, b) + 0.05) / (Math.min(a, b) + 0.05);
}
// Everything painted behind `el`, composited bottom-up, so a translucent
// panel reports the colour a reader actually sees rather than its own rgba.
function behind(el) {
const layers = [];
let node = el;
while (node && node !== document.documentElement) {
const c = getComputedStyle(node).backgroundColor;
if (alphaOf(c) > 0) {
layers.push(c);
if (alphaOf(c) >= 1) break;
}
node = node.parentElement;
}
let result = 'rgb(10, 10, 15)';
const root = getComputedStyle(document.documentElement).backgroundColor;
const body = getComputedStyle(document.body).backgroundColor;
if (alphaOf(root) >= 1) result = root;
else if (alphaOf(body) >= 1) result = body;
for (let i = layers.length - 1; i >= 0; i--) {
result = alphaOf(layers[i]) >= 1 ? layers[i] : over(layers[i], result);
}
return result;
}
const [selector, side] = arguments;
const el = document.querySelector(selector);
if (!el) return null;
const s = getComputedStyle(el);
const edge = s['border' + side + 'Color'];
const width = parseFloat(s['border' + side + 'Width']) || 0;
const opacity = parseFloat(s.opacity);
const inside = alphaOf(s.backgroundColor) >= 1
? s.backgroundColor : over(s.backgroundColor, behind(el.parentElement || el));
const outside = behind(el.parentElement || el);
return {
selector, edge, width, opacity, inside, outside,
insideRatio: Math.round(ratio(edge, inside) * 100) / 100,
outsideRatio: Math.round(ratio(edge, outside) * 100) / 100,
outline: s.outlineStyle + ' ' + s.outlineWidth + ' ' + s.outlineColor,
shadow: s.boxShadow,
};
"""
def _settled(browser: Browser, selector: str, side: str, *, timeout: float = 5) -> None:
"""Waits until the edge colour stops moving.
Every one of these controls carries `transition: border-color 0.15s`, so a
measurement taken the instant after a hover or a focus reads a colour part
way between the two states — a value no state actually has. The first run of
this check did exactly that: it reported the hover edge as rgb(114,114,160),
which is neither --border nor --border-bright but a frame between them.
Polled rather than slept, per this harness's rule.
"""
script = ("(() => { const el = document.querySelector(%s);"
" if (!el) return true;"
" const c = getComputedStyle(el)['border%sColor'];"
" const was = window.__wpeEdge; window.__wpeEdge = c;"
" return was === c; })()" % (json.dumps(selector), side))
# Cleared first, so the comparison starts from no previous reading rather
# than from whatever the last measured control left behind.
browser.js("window.__wpeEdge = undefined")
browser.wait_js(script, timeout=timeout)
def _boundary(browser: Browser, selector: str, side: str) -> dict | None:
_settled(browser, selector, side)
return browser.js(BOUNDARY_JS, selector, side)
def _record_boundary(checks: Checks, test: str, what: str, row: dict | None) -> None:
if row is None:
checks.skip(test, what, "the control was not on the page")
return
best = max(row["insideRatio"], row["outsideRatio"])
detail = (f"{best:.2f}:1 (edge {row['edge']} — {row['insideRatio']}:1 against its fill "
f"{row['inside']}, {row['outsideRatio']}:1 against {row['outside']}), "
f"{row['width']}px")
checks.record("WCAG 1.4.11", what, best >= 3.0 and row["width"] > 0, detail)
def check_control_boundaries(browser: Browser, site: Site, adv: int, checks: Checks,
evidence: dict, out: Path) -> None:
"""v1.1 WP-E §21: a control's edge is visible on its own, measured.
M11 measured text contrast on the rendered page and left boundaries to the
token audit, which checked them against one background and reported a
shortfall without failing. These are the rendered edges, at rest, on hover
and while focused, against what is actually behind them.
"""
_open_play(browser, site, adv)
measured: dict[str, dict] = {}
# A panel tab's edge is `transparent` until the panel is open, which makes
# the open tab the one control here whose boundary is the only thing marking
# it — exactly what 1.4.11 is about. So one is opened rather than skipped.
if not _open_panel(browser, "State"):
checks.skip("E04", "the open panel tab — resting edge", "the State panel did not open")
for test, what, selector, side in BOUNDARY_TARGETS:
row = _boundary(browser, selector, side)
_record_boundary(checks, test, f"{what} — resting edge", row)
if row:
measured[f"{what} (rest)"] = row
# Hover, with a real pointer: `:hover` follows the browser's pointer state,
# so a synthetic mouseover would silently re-measure the resting edge.
control = browser.find(".story-controls button:not(:disabled)", required=False)
if control is None:
checks.skip("E02", "a story control — hover edge", "no enabled control on the page")
else:
browser.hover(control)
hovered = browser.js(
"return document.querySelector('.story-controls button:not(:disabled)')"
" .matches(':hover');")
if hovered is not True:
checks.skip("E02", "a story control — hover edge",
"the pointer did not land on the control")
else:
row = _boundary(browser, ".story-controls button:not(:disabled)", "Top")
_record_boundary(checks, "E02", "a story control — hover edge", row)
if row:
measured["a story control (hover)"] = row
browser.unhover()
# Focus: the composer's edge changes colour while it holds focus, and that
# edge is what tells a keyboard reader where they are.
focused = browser.js("""
const box = document.querySelector('.input-main textarea');
if (!box) return false;
box.focus();
return document.activeElement === box;
""")
if focused is not True:
checks.skip("E01", "the story composer — focused edge", "the composer did not take focus")
else:
row = _boundary(browser, ".input-bar", "Top")
_record_boundary(checks, "E01", "the story composer — focused edge", row)
if row:
measured["the story composer (focus)"] = row
# A focus ring that is only a colour change is not enough on its own;
# M11 already asserts a visible focus indicator, and this says the
# focused edge is also measurably distinct from the resting one.
rest = measured.get("the story composer (rest)")
if rest:
checks.record("WCAG 1.4.11", "the focused composer edge differs from its resting edge",
row["edge"] != rest["edge"],
f"rest {rest['edge']} -> focus {row['edge']}")
shot = browser.screenshot(out / "control-boundaries.png")
checks.record("WP-E", "a screenshot of the measured controls was captured",
shot.exists() and shot.stat().st_size > 0, str(shot))
# The translucent edge, on the only screen that has it. Token arithmetic
# cannot reach this one: --bg-panel-glass is rgba over a gradient, so what
# is behind the nav's bottom edge is known only to the rendered page.
test, what, selector, side = TRANSLUCENT_TARGET
browser.go(site.url)
browser.wait_for(".topnav", timeout=30)
row = _boundary(browser, selector, side)
_record_boundary(checks, test, f"{what} — over a translucent panel", row)
if row:
measured[f"{what} (translucent)"] = row
checks.record("WP-E", "the nav's background really is translucent",
row["outside"] != row["inside"],
f"composited to {row['inside']} over {row['outside']}")
nav_shot = browser.screenshot(out / "control-boundaries-nav.png")
checks.record("WP-E", "a screenshot of the navigation edge was captured",
nav_shot.exists() and nav_shot.stat().st_size > 0, str(nav_shot))
evidence["control_boundaries"] = {"measured": measured,
"screenshots": [str(shot), str(nav_shot)]}
# -------------------------------------------------------------- WP-C scenarios # -------------------------------------------------------------- WP-C scenarios
def _take_count_is(count: str) -> str: def _take_count_is(count: str) -> str:
@@ -1354,7 +1576,11 @@ def main() -> int:
"C5", "failed generation and recovery"), "C5", "failed generation and recovery"),
("export", lambda: check_export_download(browser, site, adv, checks, evidence, out)), ("export", lambda: check_export_download(browser, site, adv, checks, evidence, out)),
] ]
for suite, scenarios in (("M11", m11), ("WP-C", wpc)): wpe = [
("boundaries", lambda: check_control_boundaries(browser, site, adv, checks,
evidence, out)),
]
for suite, scenarios in (("M11", m11), ("WP-C", wpc), ("WP-E", wpe)):
checks.suite = suite checks.suite = suite
for name, scenario in scenarios: for name, scenario in scenarios:
if not wanted(name): if not wanted(name):
@@ -1383,7 +1609,8 @@ def main() -> int:
"kind": ("release regression" if narrated and not only else "kind": ("release regression" if narrated and not only else
"partial (no narrator)" if not narrated else f"development (only {sorted(only)})"), "partial (no narrator)" if not narrated else f"development (only {sorted(only)})"),
"checks": checks.rows, "checks": checks.rows,
"suites": {"M11": checks.counts("M11"), "WP-C": checks.counts("WP-C")}, "suites": {"M11": checks.counts("M11"), "WP-C": checks.counts("WP-C"),
"WP-E": checks.counts("WP-E")},
"passed": checks.counts()["passed"], "passed": checks.counts()["passed"],
"failed": checks.counts()["failed"], "failed": checks.counts()["failed"],
"skipped": checks.counts()["skipped"], "skipped": checks.counts()["skipped"],
+40
View File
@@ -25,6 +25,7 @@ has actually finished arriving — never the click that started it.
from __future__ import annotations from __future__ import annotations
import base64
import json import json
import os import os
import shutil import shutil
@@ -226,6 +227,45 @@ class Browser:
def source(self) -> str: def source(self) -> str:
return self._call("GET", self._s("/source"))["value"] return self._call("GET", self._s("/source"))["value"]
def screenshot(self, path) -> Path:
"""The viewport as a PNG, written where you ask (v1.1 WP-E).
Evidence for a change a reader judges by looking at it: a contrast ratio
says a boundary is measurable, and a picture says what it looks like.
"""
encoded = self._call("GET", self._s("/screenshot"))["value"]
target = Path(path)
target.parent.mkdir(parents=True, exist_ok=True)
target.write_bytes(base64.b64decode(encoded))
return target
def hover(self, element: str) -> None:
"""A real pointer over an element, so `:hover` actually applies.
Dispatching a mouseover event from JavaScript does not do this: CSS
`:hover` follows the browser's own pointer state, not a synthetic event,
so a measurement taken after `dispatchEvent` reads the resting style and
reports it as the hover style. This moves the pointer (v1.1 WP-E).
"""
self._call("POST", self._s("/execute/sync"), {
"script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
"args": [{ELEMENT_KEY: element}]})
self._call("POST", self._s("/actions"), {"actions": [{
"type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
"actions": [{"type": "pointerMove", "duration": 60,
"origin": {ELEMENT_KEY: element}, "x": 0, "y": 0}]}]})
def unhover(self) -> None:
"""Move the pointer off whatever it was over, and forget the input state."""
self._call("POST", self._s("/actions"), {"actions": [{
"type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
"actions": [{"type": "pointerMove", "duration": 30,
"origin": "viewport", "x": 0, "y": 0}]}]})
try:
self._call("DELETE", self._s("/actions"))
except WebDriverError:
pass
def js(self, script: str, *args): def js(self, script: str, *args):
return self._call("POST", self._s("/execute/sync"), return self._call("POST", self._s("/execute/sync"),
{"script": script, "args": list(args)})["value"] {"script": script, "args": list(args)})["value"]
+353
View File
@@ -0,0 +1,353 @@
"""v1.1 WP-D criterion 3: the backup completes **through the UI**.
python -m tools.wpd_backup_ui --out <dir under $HOME> [--case real|large|both] [--show]
Run from `backend/`, with `frontend/dist` already built.
WP-D's first pass drove `POST /api/backups` — the endpoint the *Back up now*
button calls — on a 2.3 MB campaign database and a 117 MB one. That is evidence
about the implementation, and the plan's criterion 3 asks for something else:
that the backup *completes through the UI* on both. A reader does not call an
endpoint. This drives the reader-facing control in a real Firefox, against the
production build served by FastAPI, exactly as WP-C's harness does.
**No narrator and no inference.** A backup needs neither, so nothing here
touches a model host.
What it refuses to call a pass, per the brief:
- the click does nothing (no toast, no file);
- the request fails (an error toast);
- no backup file appears on disk;
- the finished copy does not pass a full `PRAGMA integrity_check`;
- the UI reports an error.
Every wait is on a condition the page or the filesystem can show. Nothing here
sleeps and then asserts.
"""
from __future__ import annotations
import argparse
import hashlib
import json
import shutil
import sqlite3
import sys
import time
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.m11_webdriver import ( # noqa: E402
Browser, Site, geckodriver_version, require_under_home,
)
BACKEND = Path(__file__).resolve().parent.parent
HOME = Path.home()
#: The two databases criterion 3 names. Both are copied before use: the first is
#: the M11 evidence campaign and must not be written to, and the second is the
#: 117 MB application database built for WP-D's timing.
DEFAULT_REAL = HOME / "m11-evidence/m04-final/campaign.db"
DEFAULT_LARGE = HOME / "v11-evidence/wp-d/endpoint/campaign-100mb/campaign.db"
BLOCK = '[data-testid="database-backup"]'
BUTTON = f'{BLOCK} button.primary'
class Checks:
"""Results, with the discipline that an unrun check is not a passing one."""
def __init__(self) -> None:
self.rows: list[dict] = []
def record(self, case: str, name: str, ok: bool, detail: str = "") -> bool:
self.rows.append({"case": case, "check": name,
"result": "PASS" if ok else "FAIL", "detail": detail})
print(f" {'ok ' if ok else 'FAIL'} {case:6} {name}"
+ (f" — {detail}" if detail else ""), flush=True)
return ok
@property
def failed(self) -> list[dict]:
return [r for r in self.rows if r["result"] == "FAIL"]
# ----------------------------------------------------------------- helpers
def integrity_of(path: Path) -> tuple[str, float]:
"""The full check on a finished copy, and what it cost, measured here.
The application runs its own `integrity_check` before keeping the file; this
is an independent second opinion on the artefact the UI produced, and it is
where criterion 3's 'time for integrity_check' comes from.
"""
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
started = time.perf_counter()
rows = connection.execute("PRAGMA integrity_check").fetchall()
elapsed = time.perf_counter() - started
finally:
connection.close()
return ", ".join(str(r[0]) for r in rows), elapsed
def digest(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()[:16]
def backups_in(db_path: Path) -> dict[str, int]:
directory = db_path.parent / "backups"
if not directory.exists():
return {}
return {p.name: p.stat().st_size for p in directory.iterdir() if p.is_file()}
def wait_for_new_backup(db_path: Path, before: set[str], *, timeout: float = 300,
poll: float = 0.2) -> Path | None:
"""The backup file the application wrote, once it is really there.
A file condition, not a sleep: a name that was not there before, not a
`.partial`, and a size that has stopped growing.
"""
directory = db_path.parent / "backups"
deadline = time.monotonic() + timeout
sizes: dict[str, int] = {}
while time.monotonic() < deadline:
if directory.exists():
for path in directory.iterdir():
if not path.is_file() or path.name in before:
continue
if path.name.endswith(".partial"):
continue
size = path.stat().st_size
if size > 0 and sizes.get(path.name) == size:
return path
sizes[path.name] = size
time.sleep(poll)
return None
def toasts(browser: Browser) -> list[dict]:
"""What the page is telling the reader — the message apart from its mark.
A toast renders a decorative mark before the message:
`<span class="toast-mark" aria-hidden="true">❖</span><span>…</span>`. So the
button's `textContent` begins with that character, and the first run of this
tool matched `textContent.startswith('Backup written:')` and reported six
*passing* behaviours as failures — the UI had said exactly what it should,
and the assertion was reading the mark. `message` is the message span alone;
`text` is kept whole for the evidence record.
"""
return browser.js("""
const host = document.querySelector('.toast-host');
if (!host) return [];
return [...host.querySelectorAll('button.toast')].map(b => {
const span = b.querySelector('span:not(.toast-mark)');
return {
error: b.classList.contains('toast-error'),
message: (span ? span.textContent : b.textContent).trim(),
text: b.textContent.trim(),
};
});
""") or []
def open_backup_panel(browser: Browser, site: Site, checks: Checks, case: str) -> bool:
"""Reach the control the way a reader does: the nav link, then the panel."""
browser.go(site.url)
browser.wait_for(".topnav", timeout=60)
link = browser.find('.nav-links a[href="/settings"]', required=False)
if link is None:
return checks.record(case, "the Settings link is in the navigation", False)
browser.click(link)
opened = browser.wait_js(f"!!document.querySelector('{BLOCK}')", timeout=60)
if not checks.record(case, "Settings opens from the navigation link", opened,
browser.url):
return False
summary = browser.find(f"{BLOCK} summary", required=False)
if summary is None:
return checks.record(case, "the backup panel has a disclosure", False)
browser.click(summary)
# The panel is a <details>: it loads what is already on disk when it opens,
# so waiting for the directory line proves the application answered.
shown = browser.wait_js(
f"document.querySelector('{BLOCK}').open === true"
f" && !!document.querySelector('{BUTTON}')", timeout=30)
checks.record(case, "the backup panel opens and shows its control", shown)
listed = browser.wait_js(f"!!document.querySelector('{BLOCK} code')", timeout=30)
checks.record(case, "the panel reports where backups are written", listed)
return shown
def click_back_up_now(browser: Browser, site: Site, db: Path, checks: Checks,
case: str, label: str) -> dict:
"""One press of the button, judged by what the page and the disk then show."""
before = set(backups_in(db))
# Clear anything still on screen, so the toast this press produces is the
# one that is read back rather than a leftover from the previous press.
browser.js("""
for (const b of document.querySelectorAll('.toast-host button.toast')) b.click();
return true;
""")
button = browser.find(BUTTON, required=False)
if button is None:
checks.record(case, f"{label}: the control is on the page", False)
return {}
started = time.perf_counter()
browser.click(button)
# Either outcome ends the wait, so a failure is reported as a failure rather
# than as a timeout.
settled = browser.wait_js(
"(() => { const t = [...document.querySelectorAll('.toast-host button.toast')];"
" return t.length > 0; })()", timeout=300)
elapsed = time.perf_counter() - started
shown = toasts(browser)
errors = [t for t in shown if t["error"]]
written = [t for t in shown if t["message"].startswith("Backup written:")]
checks.record(case, f"{label}: the click produced a visible result", settled,
json.dumps(shown)[:200])
checks.record(case, f"{label}: the UI reports no error",
not errors, json.dumps(errors)[:300])
checks.record(case, f"{label}: the UI reports the backup was written",
bool(written), json.dumps(written)[:200])
idle = browser.wait_js(
"(() => { const b = document.querySelector(%s);"
" return !!b && b.textContent.trim() === 'Back up now'; })()"
% json.dumps(BUTTON), timeout=120)
checks.record(case, f"{label}: the control returns from 'Backing up…'", idle)
produced = wait_for_new_backup(db, before)
checks.record(case, f"{label}: a backup file was physically produced",
produced is not None, str(produced))
if produced is None:
return {"seconds": round(elapsed, 3), "toasts": shown}
named = any(produced.name in t["message"] for t in written)
checks.record(case, f"{label}: the UI names the file that appeared", named,
f"{produced.name} — {written[0]['message'] if written else ''}")
relisted = browser.wait_js(
"[...document.querySelectorAll('%s .backup-list code')]"
".some(c => c.textContent.trim() === %s)" % (BLOCK, json.dumps(produced.name)),
timeout=60)
checks.record(case, f"{label}: the new backup appears in the panel's list", relisted)
verdict, integrity_seconds = integrity_of(produced)
checks.record(case, f"{label}: the finished copy passes full integrity_check",
verdict == "ok", f"{verdict} in {integrity_seconds * 1000:.1f} ms")
return {
"file": str(produced),
"bytes": produced.stat().st_size,
"sha256_16": digest(produced),
"seconds": round(elapsed, 3),
"integrity": verdict,
"integrity_seconds": round(integrity_seconds, 4),
"toasts": shown,
}
def run_case(case: str, source: Path, out: Path, checks: Checks, *, show: bool,
twice: bool) -> dict:
print(f"\n=== {case}: {source} ===", flush=True)
work = out / case
shutil.rmtree(work, ignore_errors=True)
(work / "data").mkdir(parents=True)
db = work / "data" / "campaign.db"
copy_started = time.perf_counter()
shutil.copy2(source, db)
copy_seconds = time.perf_counter() - copy_started
size = db.stat().st_size
print(f" copied {size:,} bytes in {copy_seconds:.2f}s -> {db}", flush=True)
site = Site(BACKEND, db, work / "server.log")
browser = Browser(headless=not show, log=work / "geckodriver.log")
result: dict = {"database": str(source), "bytes": size,
"served_at": site.url, "firefox": browser.version}
try:
if not open_backup_panel(browser, site, checks, case):
return result
result["first"] = click_back_up_now(browser, site, db, checks, case, "backup")
browser.screenshot(work / "back-up-now.png")
if twice and result["first"].get("file"):
kept = Path(result["first"]["file"])
before_bytes, before_digest = kept.stat().st_size, digest(kept)
result["second"] = click_back_up_now(browser, site, db, checks, case,
"second backup")
still_there = kept.exists()
checks.record(case, "the earlier backup still exists", still_there)
if still_there:
checks.record(
case, "and is byte-for-byte what it was",
kept.stat().st_size == before_bytes and digest(kept) == before_digest,
f"{before_bytes:,} bytes, sha256:{before_digest}")
verdict, _ = integrity_of(kept)
checks.record(case, "and still passes integrity_check", verdict == "ok",
verdict)
if result["second"].get("file"):
checks.record(case, "the second backup is a different file",
result["second"]["file"] != result["first"]["file"],
Path(result["second"]["file"]).name)
finally:
browser.quit()
site.stop()
return result
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--out", required=True,
help="evidence directory, which must be under $HOME")
parser.add_argument("--case", default="both", choices=["real", "large", "both"])
parser.add_argument("--show", action="store_true", help="run Firefox visibly")
parser.add_argument("--real-db", default=str(DEFAULT_REAL))
parser.add_argument("--large-db", default=str(DEFAULT_LARGE))
args = parser.parse_args()
out = require_under_home(Path(args.out).expanduser())
out.mkdir(parents=True, exist_ok=True)
if not (BACKEND.parent / "frontend/dist/index.html").exists():
print("frontend/dist is not built", file=sys.stderr)
return 2
checks = Checks()
started = datetime.now()
results: dict[str, dict] = {}
wanted = [("real", Path(args.real_db).expanduser(), True)] if args.case != "large" else []
if args.case != "real":
wanted.append(("large", Path(args.large_db).expanduser(), False))
print(f"WP-D criterion 3 — the backup through the UI. geckodriver "
f"{geckodriver_version()}", flush=True)
for case, source, twice in wanted:
if not source.exists():
checks.record(case, "the database is present", False, str(source))
continue
results[case] = run_case(case, source, out, checks, show=args.show, twice=twice)
report = {
"started": started.isoformat(timespec="seconds"),
"seconds": round((datetime.now() - started).total_seconds()),
"cases": results,
"checks": checks.rows,
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
"failed": len(checks.failed),
}
(out / "backup-ui-report.json").write_text(json.dumps(report, indent=2))
print(f"\n{report['passed']} passed, {report['failed']} failed "
f"-> {out / 'backup-ui-report.json'}")
return 1 if checks.failed else 0
if __name__ == "__main__":
raise SystemExit(main())
+26 -1
View File
@@ -151,7 +151,32 @@ export const api = {
sendAction: (advId, payload, handlers, signal) => sendAction: (advId, payload, handlers, signal) =>
streamSSE(`/adventures/${advId}/actions`, payload, handlers, signal), streamSSE(`/adventures/${advId}/actions`, payload, handlers, signal),
retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal), retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal),
exportAdventure: (id) => request(`/adventures/${id}/export`), // v1.1 WP-D: the bundle, plus what the server says about importing it back.
// The body is the bundle and nothing else — the browser saves exactly those
// bytes — so the size and the warning come back in headers.
exportAdventure: async (id) => {
const resp = await fetch(`/api/adventures/${id}/export`, {
headers: { 'Content-Type': 'application/json' },
})
if (!resp.ok) {
let detail = resp.statusText
try { detail = (await resp.json()).detail || detail } catch { /* non-JSON */ }
throw new Error(detail)
}
const bundle = await resp.json()
const number = (name) => {
const raw = Number(resp.headers.get(name))
return Number.isFinite(raw) && raw > 0 ? raw : null
}
return {
bundle,
exportBytes: number('X-Export-Bytes'),
importLimitBytes: number('X-Import-Limit-Bytes'),
// Absent header (an older server) means nothing is claimed either way.
importable: resp.headers.get('X-Importable-By-This-Version') !== 'false',
warning: resp.headers.get('X-Export-Warning') || null,
}
},
importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }), importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }),
// M9. A verified copy of the whole database, which is a different tool from // M9. A verified copy of the whole database, which is a different tool from
+4 -2
View File
@@ -66,10 +66,12 @@ export default function Campaigns() {
const exportOne = async (campaign) => { const exportOne = async (campaign) => {
try { try {
const bundle = await api.exportAdventure(campaign.id) const { bundle, warning } = await api.exportAdventure(campaign.id)
const safe = (campaign.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60) const safe = (campaign.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
downloadJSON(bundle, `${safe}.json`) downloadJSON(bundle, `${safe}.json`)
toast('Campaign exported.') // v1.1 WP-D: the file is written either way. A campaign too large for this
// version to import back says so now, not when it is needed.
toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
} catch (err) { } catch (err) {
toast(classifyError(err.message).detail, 'error') toast(classifyError(err.message).detail, 'error')
} }
@@ -79,10 +79,12 @@ export function CampaignSettingsPanel({ adventure, setAdventure, onError, moment
const exportCampaign = async () => { const exportCampaign = async () => {
try { try {
const bundle = await api.exportAdventure(adventure.id) const { bundle, warning } = await api.exportAdventure(adventure.id)
const safe = (adventure.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60) const safe = (adventure.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
downloadJSON(bundle, `${safe}.json`) downloadJSON(bundle, `${safe}.json`)
toast('Campaign exported.') // v1.1 WP-D: see Campaigns.jsx. The export is delivered; the warning says
// this version could not import the file back.
toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
} catch (err) { } catch (err) {
onError(classifyError(err.message).detail) onError(classifyError(err.message).detail)
} }
+167
View File
@@ -0,0 +1,167 @@
/* v1.1 WP-D: an export that says whether this version could import it back.
*
* The file is delivered either way — a campaign too large to re-import is not a
* damaged one, and refusing to write it would destroy the copy the reader was
* making. What changes is what they are told, and both reader-facing Export
* controls have to tell them: the one on the campaign card and the one in the
* campaign's own settings.
*
* The server decides. These assert that the page shows what it was given and
* keeps delivering the file, not that it re-derives the size policy.
*/
import { screen, waitFor } from '@testing-library/react'
import userEvent from '@testing-library/user-event'
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { api } from '../api'
import * as components from '../components'
import Campaigns from './Campaigns'
import { CampaignSettingsPanel } from './Play/panels/CampaignSettingsPanel'
import { mockModelStatus, renderWith } from '../test/helpers'
const BUNDLE = { format: 'ai-dnd-adventure-v3', title: 'Long Campaign', actions: [] }
const WARNING =
"This export is larger than this version's 20 MB import limit (21,230,000 bytes). "
+ 'The file was exported successfully, but this version cannot import it.'
const CAMPAIGN = {
id: 4, title: 'Long Campaign', action_count: 900,
updated_at: '2026-09-15T10:00:00', snippet: 'Rain over the harbour.',
}
const ADVENTURE = {
id: 4, title: 'Long Campaign', ai_instructions: '', narration_length: 'brief',
canon_rules: [], persona_name: 'Aldric', persona_desc: '',
}
// restoreAllMocks does not undo stubGlobal, and the fetch stub below would
// otherwise outlive its own describe block.
beforeEach(() => { vi.restoreAllMocks(); vi.unstubAllGlobals() })
function exportReturns({ warning = null } = {}) {
return vi.spyOn(api, 'exportAdventure').mockResolvedValue({
bundle: BUNDLE,
exportBytes: warning ? 21_230_000 : 12_000,
importLimitBytes: 20 * 1024 * 1024,
importable: !warning,
warning,
})
}
async function library() {
mockModelStatus(api)
vi.spyOn(api, 'listAdventures').mockResolvedValue([CAMPAIGN])
await renderWith(<Campaigns />)
await screen.findByText('Long Campaign')
}
async function settingsPanel() {
mockModelStatus(api)
await renderWith(
<CampaignSettingsPanel adventure={ADVENTURE} setAdventure={vi.fn()}
onError={vi.fn()} moments={900} />,
)
}
describe('reading what the server said', () => {
// The tests below mock api.exportAdventure, so nothing there exercises the
// header names. These do: a typo in one of them would otherwise leave the
// whole suite green and the reader silently uninformed.
function serverSends(headers) {
const body = JSON.stringify(BUNDLE)
vi.stubGlobal('fetch', vi.fn().mockResolvedValue(new Response(body, {
status: 200,
headers: { 'Content-Type': 'application/json', ...headers },
})))
}
it('reports a warning the server sent, with the sizes it named', async () => {
serverSends({
'X-Export-Bytes': '21230000',
'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
'X-Importable-By-This-Version': 'false',
'X-Export-Warning': WARNING,
})
const result = await api.exportAdventure(4)
expect(result.bundle).toEqual(BUNDLE)
expect(result.warning).toBe(WARNING)
expect(result.importable).toBe(false)
expect(result.exportBytes).toBe(21_230_000)
expect(result.importLimitBytes).toBe(20 * 1024 * 1024)
})
it('claims nothing when an older server sends no headers', async () => {
serverSends({})
const result = await api.exportAdventure(4)
expect(result.bundle).toEqual(BUNDLE)
expect(result.warning).toBeNull()
expect(result.importable).toBe(true)
expect(result.exportBytes).toBeNull()
expect(result.importLimitBytes).toBeNull()
})
it('treats an ordinary export as importable', async () => {
serverSends({
'X-Export-Bytes': '12000',
'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
'X-Importable-By-This-Version': 'true',
})
const result = await api.exportAdventure(4)
expect(result.importable).toBe(true)
expect(result.warning).toBeNull()
expect(result.exportBytes).toBe(12_000)
})
})
describe('the campaign library export', () => {
it('delivers the file and says nothing more when it can be imported back', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns()
await library()
await userEvent.click(screen.getByRole('button', { name: 'Export' }))
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
expect(download.mock.calls[0][0]).toEqual(BUNDLE)
expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
expect(screen.queryByText(/cannot import/)).toBeNull()
})
it('still delivers the file when it is too large, and says so', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns({ warning: WARNING })
await library()
await userEvent.click(screen.getByRole('button', { name: 'Export' }))
// The file is written first: the warning is about importing it back, not
// about the export having failed.
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
expect(download.mock.calls[0][0]).toEqual(BUNDLE)
const notice = await screen.findByText(/cannot import it/)
expect(notice).toBeInTheDocument()
expect(notice.textContent).toContain('20 MB')
expect(notice.textContent).toContain('exported successfully')
expect(screen.queryByText('Campaign exported.')).toBeNull()
})
})
describe('the campaign settings export', () => {
it('delivers the file and confirms it when it can be imported back', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns()
await settingsPanel()
await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
expect(screen.queryByText(/cannot import/)).toBeNull()
})
it('still delivers the file when it is too large, and says so', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns({ warning: WARNING })
await settingsPanel()
await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
const notice = await screen.findByText(/cannot import it/)
expect(notice.textContent).toContain('20 MB')
expect(screen.queryByText('Campaign exported.')).toBeNull()
})
})
+16 -1
View File
@@ -59,7 +59,22 @@
.slice-4 { background: #6f9e8c; } .slice-4 { background: #6f9e8c; }
.slice-5 { background: #c48a6a; } .slice-5 { background: #c48a6a; }
.slice-6 { background: #7c86b8; } .slice-6 { background: #7c86b8; }
.slice-7 { background: var(--border-bright); } /* v1.1 WP-E: pinned to the literal value this slice already rendered, instead
of borrowing --border-bright. A chart fill and a control edge have different
jobs: WP-E raised --border-bright to clear WCAG 1.4.11 (3:1) for control
boundaries, and that dragged this slice to #7a7aaa, an OKLab dE of 0.035
from .slice-6 (#7c86b8) — two neighbouring slices the same colour. The
other slices sit 0.100-0.119 from their nearest neighbour; at #3d3d55 this
one sits 0.251, the most separated in the set.
1.4.11's 3:1 does not govern this: it is a proportional fill in a labelled
breakdown, not the boundary of a control, and what it needs is to be
distinguishable from the seven slices beside it. Reassigning it to a freer
hue was considered and rejected — inside the palette's own chroma and
lightness bands the only hues that beat 0.100 are pinks near 14 degrees,
which is --danger's territory and would paint an ordinary prompt section in
the colour this application reserves for failure. */
.slice-7 { background: #3d3d55; }
.ctx-block { .ctx-block {
border: 1px solid var(--border); border: 1px solid var(--border);
+9 -2
View File
@@ -3,8 +3,15 @@
--bg-panel: #131320; --bg-panel: #131320;
--bg-panel-glass: rgba(19, 19, 32, 0.82); --bg-panel-glass: rgba(19, 19, 32, 0.82);
--bg-input: #1a1a2a; --bg-input: #1a1a2a;
--border: #2b2b3d; /* v1.1 WP-E: control boundaries carry WCAG 1.4.11 (3:1 non-text contrast) on
--border-bright: #3d3d55; their own, rather than leaning on the control's text label. The floor is
measured against --bg-input (#1a1a2a), not --bg-panel: inputs and buttons
are drawn on --bg-input (styles/forms.css), and it is the lightest of the
three backgrounds a border sits on, so it is the worst case.
--border 3.21:1 and --border-bright 4.24:1 there; higher on the others
(backend/tools/contrast_audit.py, which now fails below 3:1). */
--border: #676792;
--border-bright: #7a7aaa;
--text: #e2ddd0; --text: #e2ddd0;
--text-dim: #918c7d; --text-dim: #918c7d;
--accent: #d4a94e; --accent: #d4a94e;
+7 -3
View File
@@ -3,9 +3,13 @@
**This file is the index. Start here.** **This file is the index. Start here.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: **Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress:
WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`) are committed, and WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`), WP-B.2 (`0c1ba83`) and WP-C
WP-C, browser release coverage, is complete and staged for owner review** (`59b5ebc`) are committed, and the last two planned packages — WP-D (recovery
(`reports/v1.1/V1.1-WP-C-REPORT.md`). honesty) and WP-E (control-boundary contrast) — are complete and staged for
owner review** (`reports/v1.1/V1.1-WP-D-REPORT.md`,
`reports/v1.1/V1.1-WP-E-REPORT.md`). The final browser run passed 101 checks
across all three suites with 0 failed and 0 skipped. Every planned v1.1 work
package is now implemented and reported; release validation has not begun.
Phase 0 complete; AI-DnD forked as the production base; **milestones M1 Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the (2026-09-14), the v1 release gate passed on the release-candidate tree, and the
+9 -3
View File
@@ -16,10 +16,16 @@ real-model limitation** (owner decision, 2026-09-15):
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T. The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
**WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release **WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release
coverage, is complete and staged for owner review. Its final run passed 91 checks coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS, (the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`).
have not started. No v1.1 version or tag exists. **WP-D** (recovery honesty) and **WP-E**
(control-boundary contrast) are complete and staged for owner review, each with
its own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`. The final
browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0
failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C,
WP-D, WP-E) is now implemented and reported; release validation has not begun,
and no v1.1 version or tag exists.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1 This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1 history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
+23 -2
View File
@@ -1,8 +1,29 @@
# Planning Package Version # Planning Package Version
- **Package:** Adventure Storyteller Planning Package v4.4 - **Package:** Adventure Storyteller Planning Package v4.5
- **Revision date:** 2026-09-15 - **Revision date:** 2026-09-15
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. **WP-C, browser release coverage, is complete and staged for owner review: 91 checks, 0 failed, 0 skipped, over trusted-LAN HTTPS, with real export downloads** (`reports/v1.1/V1.1-WP-C-REPORT.md`). WP-D and WP-E have not started, and no v1.1 version or tag exists. - **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C, browser release coverage, is committed and signed as `59b5ebc` (91 checks, 0 failed, 0 skipped, with real export downloads). **WP-D (recovery honesty) and WP-E (control-boundary contrast) are complete and staged for owner review**, each with its own report (`reports/v1.1/V1.1-WP-D-REPORT.md`, `V1.1-WP-E-REPORT.md`); the final browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0 failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) is now implemented and reported. Release validation has not begun, and no v1.1 version or tag exists.
## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15)
The last two planned v1.1 packages, implemented and reported separately. No
requirement, acceptance test, schema or bundle format changed. WP-D's evidence is
in `reports/v1.1/V1.1-WP-D-REPORT.md`, WP-E's in `V1.1-WP-E-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `reports/v1.1/V1.1-WP-D-REPORT.md` | **New.** Full `integrity_check` on the finished backup copy, proved against a fixture `quick_check` calls healthy; an export that says when this version could not import it back, carried in headers because the response body *is* the bundle | work-package report |
| `reports/v1.1/V1.1-WP-E-REPORT.md` | **New.** Control boundaries raised to clear WCAG 1.4.11 (3:1), the contrast audit turned from a report into a gate, rendered before/after boundary measurements, and the `.slice-7` finding the change itself created | work-package report |
| `V1.1-PLAN.md` | Status: WP-C signed `59b5ebc`; WP-D and WP-E complete and staged | status |
| `planning/README.md` | Current state | index |
| `DEVELOPMENT.md` | How large an export can get: the 20 MB import ceiling, ~13 kB per action, M9's ~279-turn figure and why it is not a turn limit, and the four export headers | developer docs |
**Requirement changes: zero.**
Two owner decisions are outstanding, both recorded rather than assumed: WP-E's
before/after screenshots await approval (`OWNER SCREENSHOT APPROVAL: PENDING`),
and WP-D records that no real-browser click was made on *Back up now* — the
endpoint that button calls was driven instead.
## v4.4 — WP-C browser release coverage (2026-09-15) ## v4.4 — WP-C browser release coverage (2026-09-15)
+526
View File
@@ -0,0 +1,526 @@
# v1.1 WP-D — Recovery Honesty
**Status:** COMPLETE — **PASS**. The decision, and what was deliberately not
claimed, is in §P.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD at start | `59b5ebc2d85bdf53d38b6bcf347c496dd3432be1` — *v1.1 WP-C: browser release coverage*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
| Its ancestry | `0c1ba83` (WP-B.2), `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `432f041` (v1.0.0) |
| Working tree at start | clean; nothing staged |
| WP-E | not started when WP-D was implemented |
---
## B. Existing backup behaviour
`backend/app/backup.py`, as v1 shipped it. The six questions the brief asks:
| # | Question | Answer (before WP-D) |
| --- | --- | --- |
| 1 | How is the backup file created? | SQLite's **online backup API** (`sqlite3.Connection.backup`, `pages=-1`), from the live database opened read-only through a `mode=ro` URI. The destination is a temporary file `…​.db.partial` **in the destination directory**, so the rename below is atomic |
| 2 | Where does validation happen? | `_verify()`, on the finished copy, opened as its own read-only connection — not on the source, and not through the connection that wrote it |
| 3 | When does the final filename appear? | Only after verification: `os.replace(working, target)`. A failed or interrupted run never leaves a file wearing a backup's name |
| 4 | How are failures cleaned up? | `_discard()` removes the partial file, and `BackupError` is raised with what went wrong. The source is untouched |
| 5 | Could an existing good backup be overwritten? | **No.** `_unused_name()` stamps each backup with the time and adds a counter if that name (or its `.partial`) exists |
| 6 | Was the check on the source or the copy? | The copy |
So the mechanism was already right. The one gap was the **strength** of the check:
`PRAGMA quick_check`, which reads every page and every record but skips the
cross-check between a table and its indexes.
---
## C. `integrity_check` implementation
One function changed:
```python
# app/backup.py, _verify()
rows = connection.execute("PRAGMA integrity_check").fetchall() # was quick_check
```
- It still runs on the **finished copy**, opened as its own read-only
connection, before the rename.
- A failure still raises `BackupError` naming what was wrong, still discards the
partial file, and still leaves the source and every earlier backup untouched.
- The returned `integrity` field, the filename, the directory, the reported
fields (`filename`, `bytes`, `pages`, `seconds`, `integrity`) and the API are
unchanged.
- The module docstring and `_verify`'s docstring now record why the trade M9
made (speed over the index cross-check) was not needed at these sizes, with
the measurements in §E.
**No scheduled backups were added.** There is still no restore endpoint, no
retention policy and no timer: a backup happens when the reader asks for one.
---
## D. Corruption fixture
`tests/test_v11_d_recovery.py: build_corrupt_copy()`.
A small database with `t(id, k, filler)` and an index `i_t_k ON t(k)`, 400 rows
with keys `k000000`…`k000399`. `dbstat` names the index's own leaf pages, and one
digit inside one indexed key on the first of them is changed (`k000144` →
`k000944`). Every page stays structurally sound and every record still parses:
what is broken is only the agreement between the index and its table.
**Proved before it is used as evidence**
(`test_the_fixture_is_the_difference_between_the_two_checks`):
| Pragma | Result |
| --- | --- |
| `PRAGMA quick_check` | **`ok`** |
| `PRAGMA integrity_check` | **`row 145 missing from index i_t_k`** |
That is the difference the package rests on: not a database both checks reject,
but one the old check called healthy.
---
## E. Backup performance measurements
Each pragma was run in its **own fresh process**, alternating, three times, because
a first measurement warms the page cache and a naive ordering makes whichever
check runs second look faster. (It did: an early single-pass measurement showed
`integrity_check` at 1.9 ms against `quick_check` at 65.5 ms, purely from cache
warmth.)
**The real campaign database** — the M11 100-turn evidence campaign,
`m04-final/campaign.db`, 2,367,488 bytes (2.26 MB):
| | Run 1 | Run 2 | Run 3 |
| --- | --- | --- | --- |
| `quick_check` | 5.0 ms | 3.4 ms | 7.6 ms |
| `integrity_check` | 3.5 ms | 5.7 ms | 5.3 ms |
At this size the two are indistinguishable. A whole backup through
`backup.create()` — copy, full check and rename — took **70.8 ms** and **27.4 ms**
on two runs (578 pages).
**An index-heavy synthetic database of 105.2 MB** (105,160,704 bytes; 25,674
pages of 4,096 bytes), shaped to give the cross-check real work: `actions` with
152,000 rows and `memories` with 15,200, under three indexes
(`i_actions_adv_depth`, `i_actions_kind`, `i_memories_adv`):
| | Run 1 | Run 2 | Run 3 |
| --- | --- | --- | --- |
| `quick_check` | 172.6 ms | 92.4 ms | 97.1 ms |
| `integrity_check` | 227.8 ms | 196.0 ms | 227.4 ms |
A whole backup of it through `backup.create()`: **0.58 s** for 25,674 pages,
`integrity = ok`.
**A campaign-shaped database of 117.4 MB** (117,403,648 bytes; 2,200 actions),
built through the application's own models so the product path can run against
it (§E.1):
| | Run 1 | Run 2 | Run 3 |
| --- | --- | --- | --- |
| `quick_check` | 89.5 ms | 74.2 ms | 64.4 ms |
| `integrity_check` | 90.0 ms | 82.9 ms | 75.9 ms |
**Reading.** The cost of the cross-check tracks **rows and index entries, not
bytes**. On the campaign schema — 2,200 fat rows — the full check costs about
7 ms more than the quick one at 117 MB. On the index-heavy synthetic — 167,200
rows under three indexes at a similar size — it costs about 96 ms more. Both sit
inside a backup of about half a second. There is no performance requirement in
this project and WP-D does not invent one; the measurements are here because the
M9 trade was made on a speed argument, and at these sizes that argument does not
hold.
### E.1 Through the endpoint the button calls — implementation evidence only
**This does not satisfy acceptance criterion 3, and an earlier draft of this
report wrongly said it did.** The criterion asks for the backup to *complete
through the UI*; a reader does not call an endpoint. What follows is evidence
about the implementation — the procedure, its result and its cost — and it is
kept for that reason. The acceptance evidence is in **§E.2**.
`POST /api/backups` is the only call the **Back up now** button makes, and it
takes no parameters, so it was driven directly against a **copy** of each
database — the evidence database is evidence and is not written to:
| Database | Result |
| --- | --- |
| Evidence campaign, 2,367,488 bytes | **201** in 0.117 s wall — `pages: 578`, `seconds: 0.038`, `integrity: ok`; `GET /api/backups` then lists 1 backup |
| Campaign-shaped, 117,403,648 bytes | **201** in 0.536 s wall — `pages: 28,663`, `seconds: 0.452`, `integrity: ok` |
**Why a second large database exists.** The index-heavy synthetic has no
application schema, and the server runs its migrations at startup, so the app
will not start against it: driving the endpoint there fails in a migration that
renumbers `actions`, before any backup is attempted. It remains valid evidence
for the pragma comparison, which is a property of SQLite and not of this schema,
but criterion 3 needs a database the product can actually open — hence the
campaign-shaped one, which §E.2 then drives through the browser.
Artefacts stay under `$HOME` (`v11-evidence/wp-d/`), not in the repository.
---
### E.2 Through the UI — the acceptance evidence for criterion 3
The reader-facing **Back up now** control, clicked in a real Firefox, on the
production build served by FastAPI: the WP-C path, with no Vite dev server, no
inference, and no sleep used as an assertion. Every wait is on a condition the
page or the filesystem can show.
`backend/tools/wpd_backup_ui.py`, run as
`python -m tools.wpd_backup_ui --out $HOME/v11-evidence/wp-d/ui --case both`.
**34 checks, 34 passed, 0 failed** —
`$HOME/v11-evidence/wp-d/ui/backup-ui-report.json`.
The route a reader takes is the route the tool takes: open the application, click
**Settings** in the navigation, open *Back up everything on this machine* (a
`<details>` that loads what is on disk when it opens), then press the button.
| | **Case 1 — real campaign database** | **Case 2 — ≥100 MB application database** |
| --- | --- | --- |
| Source | the M11 evidence campaign, copied | the campaign-shaped database of §E, copied |
| **Database size** | **2,367,488 bytes** (2.3 MB) | **117,403,648 bytes** (112.0 MB) |
| Schema | the application's own, `user_version` 94 | the application's own, `user_version` 94 |
| **Browser action** | click **Back up now** (twice — see below) | click **Back up now** |
| **Observable UI result** | toast: *"Backup written: adventure-storyteller-20260915-225000.db (2.3 MB)."*, no error toast, the control returns from *Backing up…*, and the file appears in the panel's list | toast: *"Backup written: adventure-storyteller-20260915-225007.db (112.0 MB)."*, no error toast, control returns, file listed |
| **Backup path** | `…/ui/real/data/backups/adventure-storyteller-20260915-225000.db` | `…/ui/large/data/backups/adventure-storyteller-20260915-225007.db` |
| **Backup file size** | **2,367,488 bytes** | **117,403,648 bytes** |
| **`integrity_check` result** | **`ok`** | **`ok`** |
| **Integrity-check elapsed** | **5.8 ms** | **78.5 ms** |
| **Total elapsed backup** | **0.108 s** (click → the UI says it is done) | **0.719 s** |
"Total elapsed" is measured from the click to the rendered result, so it is what
the reader waits, not what the server reports. The `integrity_check` above is an
independent second opinion, run here on the artefact the UI produced — the
application had already verified the copy before keeping it.
**The size the UI reports is the size on disk.** "112.0 MB" in the toast is
117,403,648 bytes shown in the page's own units; the file matches the source
database byte for byte.
**Existing-backup protection, established through the UI itself** (Case 1's
second press, rather than a backup planted by a library call):
| Claim | Result |
| --- | --- |
| A second press writes a *different* file — `…-225000-2.db`, the counter `_unused_name` adds when two backups land in the same second | PASS |
| The first backup still exists afterwards | PASS |
| And is byte-for-byte what it was — 2,367,488 bytes, `sha256:7d45555566…` | PASS |
| And still passes `integrity_check` | PASS |
No corruption was fabricated through the browser; rejection behaviour is already
proved deterministically in §D.
**Supplemental endpoint evidence.** The servers' own logs corroborate that the
button drove each backup: Case 1 logged two `POST /api/backups → 201 Created`,
each followed by `GET /api/backups → 200 OK` as the panel reloaded its list;
Case 2 logged one of each. Screenshots and logs are beside the report JSON.
**A harness defect this run found, in my own check and not in the product.** The
first pass reported six failures. The toast renders a decorative mark before its
message — `<span class="toast-mark" aria-hidden="true">❖</span><span>…</span>` —
so the button's `textContent` begins with `❖`, and my assertion matched
`textContent.startswith("Backup written:")`. In that same run the UI had shown a
non-error toast naming the file, the file was on disk, its name was in the
panel's list, and the copy verified: the product was right and the assertion was
reading the mark. The check now reads the message span. **No product code was
changed** — the expected change for this work was none, and none was needed.
---
## F. Existing export/import limit behaviour
| | |
| --- | --- |
| Import ceiling | `limits.MAX_IMPORT_BODY_BYTES` = 20 MB, **unchanged by WP-D** |
| Where it is enforced | `BodySizeLimitMiddleware`, on the declared `Content-Length`, before the body is read |
| Refusal | HTTP **413**, "Request too large (limit 20 MB)." — it already named the limit |
| Export before WP-D | `GET /adventures/{id}/export` returned the bundle. Nothing compared its size with the ceiling, so a campaign could be exported and then refused by its own importer |
---
## G. Oversized-export warning implementation
**Where the answer comes from.** `limits.import_limit_label()` and
`limits.oversized_export_warning(size)` derive the sentence from
`MAX_IMPORT_BODY_BYTES`. A test changes the constant and asserts the wording
follows, so no number is written twice.
> This export is larger than this version's 20 MB import limit (21,230,000 bytes).
> The file was exported successfully, but this version cannot import it.
**Where it travels: headers, not the body.** The export response *is* the bundle —
the browser saves exactly those bytes as the file — so a warning inside it would
become part of a portable story file and of every checksum taken over one. The
route returns the same body with:
```text
X-Export-Bytes the serialised size
X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
X-Importable-By-This-Version true / false
X-Export-Warning only when false
```
**Which size is measured.** The compact serialisation (`separators=(",", ":")`,
`ensure_ascii=False`), which is what this response sends *and* what the browser
POSTs back on import — the bytes `BodySizeLimitMiddleware` weighs. The
pretty-printed file the reader downloads is larger and is not what import reads.
The route serialises once and returns those bytes, so `X-Export-Bytes` is the
length of the body actually sent.
**The bundle is unchanged.** No key was added to it (§I), and the format stays
`ai-dnd-adventure-v3`.
---
## H. Oversized export evidence
`test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back`
writes real rows into a campaign until its bundle genuinely exceeds the ceiling —
nothing is mocked, and the export serialises all of it.
| Claim | Result |
| --- | --- |
| The response body is larger than 20 MB | PASS |
| It parses, and is `ai-dnd-adventure-v3` with its actions | PASS |
| `X-Importable-By-This-Version: false` | PASS |
| `X-Export-Warning` names the limit ("20 MB"), says the export succeeded, and says this version cannot import it | PASS |
| `X-Export-Bytes` equals the body length | PASS |
| The same bundle is refused by import, with the limit named (§J) | PASS |
| A normal campaign's export carries **no** warning, and `X-Importable-By-This-Version: true` | PASS |
| The warning follows the constant (changed to 50 MB in a test: the sentence says 50 MB, and no longer says 20 MB) | PASS |
---
## I. Normal-export compatibility
The same campaign database (`m04-final/campaign.db`) exported through a worktree
at `59b5ebc` (before WP-D) and through the WP-D tree:
| | |
| --- | --- |
| Before | 2,699,076 bytes |
| After | 2,699,076 bytes |
| Comparison | **identical** — the parsed documents compare equal, key for key |
No timestamp allowance was needed: this campaign's bundle carries no field that
moves between exports. The format is `ai-dnd-adventure-v3` in both.
`test_the_bundle_itself_never_carries_the_warning` additionally asserts that no
key of an oversized bundle mentions the warning, the limit or importability.
---
## J. Import-refusal behaviour
Unchanged in behaviour, and already naming the limit:
| | |
| --- | --- |
| Status | **413**, from the middleware, on `Content-Length`, before the body is parsed |
| Message | "Request too large (limit 20 MB)." |
| WP-D test | `test_that_same_bundle_is_refused_by_import_naming_the_limit` posts the *actual oversized export* and asserts 413 and that the detail contains `limits.import_limit_label()` |
| Not relaxed | the same boundary, the same status, the same middleware. `test_a_bundle_under_the_limit_still_imports` keeps the ordinary path honest (201) |
No wording correction was needed.
---
## K. Frontend behaviour
`api.exportAdventure` now returns the bundle **and** what the server said about
it (`exportBytes`, `importLimitBytes`, `importable`, `warning`), read from the
headers.
**The no-header case.** `importable` is `resp.headers.get(…) !== 'false'`, so
only the literal string `false` is read as a refusal: an older server that sends
no headers — or a header that arrives malformed — yields `importable: true` and
`warning: null`, and the page says nothing it was not told. The two size fields
fall back to `null` unless they parse as a positive number.
This is asserted directly: three tests stub `fetch` with real response headers
and check what `api.exportAdventure` makes of them — a warning with the sizes it
named, an ordinary export marked importable, and an older server sending no
headers at all. They exist because the page-level tests below mock
`api.exportAdventure` itself and so cannot see a header name, which meant a typo
on the frontend side would have left the whole suite green (§O.7).
Both reader-facing entry points keep delivering the file first and then report:
| Entry point | Under the limit | Over the limit |
| --- | --- | --- |
| Campaign library, **Export** | file written; "Campaign exported." | file written; the warning, as an error-styled toast |
| Campaign settings, **Export campaign** | file written; "Campaign exported." | file written; the warning |
No new modal, no redesign: the existing toast carries it.
**Tests** (`pages/exportHonesty.test.jsx`, **7**): three read real response
headers through a stubbed `fetch` (above), and four drive both entry points —
the file is delivered in every case; the warning is shown when the server sends
one, naming 20 MB and saying the export succeeded; the ordinary confirmation is
shown when it does not.
---
## L. Offline regression
`tools/m11_offline.py`, against the production build, in the offline container:
**23 checks, 23 passed, 0 failed** —
`$HOME/v11-evidence/wp-d/offline/offline-report.json`.
The two this package could have broken are in it and passed:
| Check | Result |
| --- | --- |
| a campaign exports offline | ok |
| and imports offline, with its state | ok |
| no secret is present in the export | ok |
| a local file imports offline | ok |
The export route now sets headers and serialises compactly; the offline
container still exports a campaign and imports it back with its state, so
neither the round trip nor the secret-scrubbing changed.
## M. Full regression
| Suite | Result |
| --- | --- |
| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,712 passed, 17 skipped, 0 failed, 0 xfailed** (936.6 s) |
| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
| **Production build** (`npm run build`) | succeeded |
| **Offline regression** | 23 passed, 0 failed (§L) |
**The backend count reconciles exactly.** WP-B's closing tree was 1,693; WP-C
added 7 (`test_v11_c_browser_helpers.py`) and did not rerun the suite because no
application code changed; WP-D adds the 12 in `test_v11_d_recovery.py`.
1,693 + 7 + 12 = **1,712**.
**The 17 skips are named, not assumed.** Run with `-rs`, every one is an
environment-gated real-model test, and none is new:
| File | Skipped | Gate |
| --- | --- | --- |
| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
| `test_narrative_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
| `test_m11_real_window.py` | 3 | two the same; one `AIDND_TEST_WIDE_MODEL` |
| `test_provider_wiring.py` | 1 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
That is the same 17 B.1 and B.2 recorded. WP-D used no model and added no skip.
**The plan's named regression requirements**, run individually rather than
assumed to be inside the total:
| Requirement | Result |
| --- | --- |
| I01-I07, L01-L04 | **25 passed**, 1,704 deselected |
| The backup case in `test_m11_migration.py` | **1 passed** |
| `backup.test.jsx` | **7 passed** |
| The offline container's export and import | passed (§L) |
**Frontend arithmetic.** WP-C's baseline was 168. WP-D adds 4 page-level export
tests and 3 header-parsing tests: 168 + 7 = **175**. The 15 lint warnings are
the same pre-existing `only-export-components` and unused-import kind recorded
at WP-C, and **none is in a file WP-D changed**.
## N. Compatibility
| Area | Effect |
| --- | --- |
| Database schema | unchanged; no migration |
| Bundle format | `ai-dnd-adventure-v3`, unchanged |
| Normal bundle contents | byte-identical (§I) |
| Existing backups | still valid files with the same names; only the check that admits a new one is stricter |
| v1.0.0 databases | open unchanged (no schema or data path changed) |
| Import limit | unchanged at 20 MB |
| Scheduled backups | none, as before |
| History, branches, Save Points, state, memory, knowledge | untouched |
| Endpoint policy, local-only operation | untouched |
## O. Residual risks
1. **Closed: the backup is now driven through the UI.** This risk previously
read "the browser click was not driven for the backup", and the owner
correctly refused criterion 3 on endpoint evidence. §E.2 drives the real
**Back up now** control in Firefox on both databases — 34 checks, 0 failed —
including the ≥100 MB case, which does mean the harness starts a server
against a 117 MB database and takes 0.719 s to do it. What remains is
ordinary coverage scope: this runs as its own tool rather than inside the
release harness, so it is not part of the 101-check run in the WP-E report.
2. **The 20 MB ceiling is unchanged.** A campaign past it still cannot be
imported by this version. WP-D makes that audible at export; raising the
limit or streaming import stays v1.2, as the plan assigns it.
3. **The warning is about *this* version.** It says what this build's importer
will accept. A future build with a higher ceiling could import a file this
one warned about, and the wording ("this version") is chosen so that stays
true rather than becoming a lie.
4. **`integrity_check` is not a guarantee of recoverability.** It proves the
copy's pages, records and indexes agree. A database that was already
logically wrong when it was copied is copied faithfully and passes. The
package makes the check honest, not omniscient.
5. **No restore path.** There is still no restore button and no scheduled
backup — both explicitly out of scope. A reader who needs a backup restores
it by replacing the file themselves, as before.
6. **The header channel depends on the browser reaching headers.** Both export
call sites read them through `fetch`, so a proxy that stripped `X-` headers
would silently return to v1 behaviour: the file still arrives, the warning
does not. Local-only operation makes that unlikely, and the failure is the
old behaviour rather than a wrong claim.
7. **Header parsing was untested — now closed.** Writing this section exposed
it: the page-level tests mock `api.exportAdventure` and the backend tests
assert what the server sends, so nothing read an actual header name, and a
frontend-side typo would have left every test green and the reader silently
uninformed. Three tests that stub `fetch` with real headers now cover it
(§K). What remains is the ordinary version of this risk: the two sides agree
by matching string literals in two files, and only a browser-level test would
catch a mismatch introduced in both at once.
## P. Final decision
**Against the plan's acceptance criteria** (§ WP-D, *Recovery honesty*):
| # | Criterion | Verdict |
| --- | --- | --- |
| 1 | A healthy database's backup reports `integrity_check` ok and is kept | **PASS** — `test_a_healthy_backup_passes_the_full_check_and_is_kept`, and both endpoint runs returned `integrity: ok` (§E.1) |
| 2 | A copy `integrity_check` rejects and `quick_check` does not is rejected and not kept; existing backups still never overwritten | **PASS** — the fixture is proved to be exactly that difference (§D), and three tests cover rejection, the untouched earlier backup, and the unchanged file semantics |
| 3 | The time is recorded on the evidence database and on a synthetic ≥100 MB, and the backup completes through the UI on both | **PASS** — pragma timings in §E on three databases, and the backup completes **through the UI** on both in §E.2: 0.108 s with `integrity_check` `ok` in 5.8 ms on the 2.3 MB campaign, 0.719 s with `ok` in 78.5 ms on the 117 MB application database. An earlier draft claimed this criterion on endpoint evidence (§E.1); that claim was wrong and is corrected |
| 4 | An over-20 MB export succeeds, delivers the file, and warns naming the limit, in the API response and a component test; under the limit, no warning | **PASS** — §H, and 7 frontend tests (§K) |
| 5 | Importing that bundle is refused with a message naming the limit | **PASS** — §J, asserted on the actual oversized export |
| 6 | A normal export is byte-identical before and after, apart from timestamps | **PASS** — fully identical, no timestamp allowance needed (§I) |
**What I corrected rather than reported around.** Two claims of my own failed
checking and were fixed, not softened: §E originally rested criterion 3 on
`backup.create()` calls, and driving the real endpoint showed the index-heavy
synthetic cannot serve that criterion at all (the app's startup migrations abort
against a schema-less database) — hence the campaign-shaped database in §E.1.
And §K asserted the no-header fallback from the code alone, which exposed that
**no test read any header name**; three tests now do (§O.7).
**Not done, and not claimed:** the import ceiling is unchanged at 20 MB, by the
plan's own assignment of the raise to v1.2. The UI backup runs as its own tool
rather than inside the release harness (§O.1).
**No product code changed for the UI verification.** The run exposed one defect,
and it was in my own assertion, not in the application (§E.2). Backend
**1,723 passed / 17 skipped / 0 failed** and offline **23/23** therefore stand
as prior evidence, unchanged and not re-run; the targeted backup and browser
checks were re-run instead.
```text
BACKUP INTEGRITY: PASS
BACKUP THROUGH UI — REAL DB: PASS
BACKUP THROUGH UI — >=100 MB DB: PASS
EXPORT HONESTY: PASS
WP-D OVERALL:
PASS
```
All WP-D changes are **staged and uncommitted**. No commit, no push, no tag.
WP-E has not begun.
+374
View File
@@ -0,0 +1,374 @@
# v1.1 WP-E — Control-Boundary Contrast
**Status:** COMPLETE — **PASS**, with owner approval of the screenshots
outstanding. The decision, and what was deliberately not claimed, is in §O.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD | `59b5ebc` — *v1.1 WP-C: browser release coverage*, signed by the owner |
| Working tree at start | WP-D staged (10 files), nothing committed |
| Criterion | WCAG 2.1 **1.4.11 Non-text Contrast**, 3:1, for control boundaries; **1.4.3** 4.5:1 for body text, unchanged |
---
## B. What v1.0.0 actually did
`tools/contrast_audit.py` measured control boundaries, printed that two of them
were below 3:1, and **exited 0**. Its own comment argued the position:
> in this design a control is identified by its *label*, which is measured above
> and passes, not by its edge. So a boundary below 3:1 is reported with its
> number and does not fail the run.
So the audit was a report, not a gate: no palette change could ever fail it on a
boundary. The two numbers it printed were **1.33:1** (`--border` on
`--bg-panel`) and **1.75:1** (`--border-bright`), against a floor of 3.0.
**WP-E overturns that argument.** 1.4.11 covers the visual information needed to
identify a component *and its boundary*; a reader who cannot see where a text box
ends cannot see that there is a text box to type into, label or no label. The
tokens were raised rather than the criterion re-argued.
---
## C. Token inventory
| Token | v1.0.0 | v1.1 | Why |
| --- | --- | --- | --- |
| `--border` | `#2b2b3d` | **`#676792`** | every control's resting edge |
| `--border-bright` | `#3d3d55` | **`#7a7aaa`** | hover edges, the composer's resting edge, the open panel tab |
| `--bg-panel`, `--bg-input`, `--bg`, `--text`, `--text-dim`, `--accent*`, `--danger`, `--warning`, `--player`, `--chart-*` | — | **unchanged** | WP-E is a boundary package; no text or accent colour moved |
**The floor is taken against `--bg-input`, not `--bg-panel`.** Inputs and buttons
are drawn on `--bg-input` (`styles/forms.css`), which is lighter than
`--bg-panel` and therefore the harder case. The audit had been checking only
`--bg-panel`, so a token could have passed the audit while the real control
failed. Measured on the new values:
| | vs `--bg-input` | vs `--bg-panel` | vs `--bg` |
| --- | --- | --- | --- |
| `--border` | **3.21:1** | 3.44:1 | 3.70:1 |
| `--border-bright` | **4.24:1** | 4.55:1 | 4.88:1 |
Two properties were preserved deliberately: the rest→hover step is the same size
as before (1.318 → 1.320), so hover still reads as a change rather than a jump;
and `--border-bright` stays *below* body text against the same panel (2.98:1
between them), so no edge outshines the words inside it.
---
## D. Component inventory
Where these tokens are actually drawn, from the stylesheets:
| Control | Rule | Rest | Hover / active | Focus |
| --- | --- | --- | --- | --- |
| Story composer | `.input-bar` (story.css) | `--border-bright` on `--bg-panel` | — | `--accent-dim` + `--accent-glow` ring |
| Story controls | `.story-controls button` | `--border` on `--bg-panel` | `--border-bright` | (M11 focus check) |
| Fields and buttons | `forms.css` | `--border` on `--bg-input` | `--accent-dim` | `--accent-dim` + ring |
| Panel tabs | `.panel-tabs button` | **`transparent`** | `--border` on hover, `--border-bright` when active | — |
| Top navigation | `.topnav` | `--border` bottom edge on **`--bg-panel-glass`** | — | — |
| Scrollbar thumb | `base.css` | `--border-bright` as a *fill* on `--bg` | `--accent-dim` | — |
Two of these cannot be answered by token arithmetic at all, and both are
measured in the browser instead (§G): the nav sits on a translucent panel, and
the panel tab's edge is `transparent` until the panel is open.
---
## E. The audit is now a gate
`tools/contrast_audit.py`:
1. **Boundary pairs fail.** `text` and `boundary` rows are both pass/fail; the
advisory branch is gone. The run returns 1 if either kind falls short.
2. **Eight boundary pairs replace two.** Each border is checked against every
background it is drawn on — `--bg-input`, `--bg-panel` and `--bg` — plus the
focused edge (`--accent-dim`) on both panel and field backgrounds.
3. **The verdict is taken on the number that is printed** (rounded to two
decimals), so a pair shown as `3.00:1` is not failed for arithmetic the
reader cannot see.
4. The comment block that argued the old position is replaced by one recording
what changed and why, including why `--bg-panel-glass` is not in the list.
## F. Gate tests
`backend/tests/test_v11_e_contrast.py` — **11 passed**. The threshold is
exercised from both sides, on real token files:
| Test | Result |
| --- | --- |
| A boundary at **2.99:1** against `--bg-input` fails the run (exit 1) | PASS |
| A boundary at **3.00:1** passes (exit 0) | PASS |
| The **v1.0.0 value** `#2b2b3d` fails, at 1.24:1 against `--bg-input` | PASS |
| A dimmed `--text-dim` still fails as a *text* pair | PASS |
| A renamed token is a failure, not a silent skip | PASS |
| The shipped palette passes both criteria | PASS |
| Every boundary pair is measured against the background it is drawn on | PASS |
| The M11 text baselines are unchanged: 14.57 / 13.57 / 5.48 / 5.88 | PASS |
| The hover edge stays brighter than the resting edge | PASS |
| No boundary becomes as loud as body text | PASS |
| The WCAG ratio formula is anchored on known values (21:1, 1:1, symmetry) | PASS |
The 2.99 and 3.00 values are worth noting: **both clear 3:1 against
`--bg-panel`** (3.21 and 3.22). They decide the gate only because the floor is
now taken against the background the control is really on — so these two tests
also prove §C's change is doing work.
---
## G. Browser measurement
`tools/m11_browser.py` gains a third suite, **WP-E**, counted separately from
M11's 38 and WP-C's 53. It measures the *rendered* edge — `borderColor` from
`getComputedStyle` — against what is actually behind it, with every translucent
layer composited bottom-up.
A boundary is measured against **both** adjacent colours (the control's own fill
inside it, the background outside it) and passes on the better of the two: an
edge that matches its fill but contrasts with the page is still a visible
outline. What 1.4.11 asks is that the component's extent be perceivable.
Two harness capabilities were added for this (`tools/m11_webdriver.py`):
- **`hover()`** moves a real pointer through the WebDriver Actions API.
Dispatching a `mouseover` event from JavaScript does *not* trigger CSS
`:hover`, so a synthetic event would have re-measured the resting edge and
reported it as the hover edge.
- **`screenshot()`** writes the viewport as a PNG, for the before/after evidence.
**A defect this found in my own first measurement.** The first run reported the
hover edge as `rgb(114, 114, 160)` and the focused edge as `rgb(144, 120, 81)` —
neither of which is any token. Both controls carry `transition: border-color
0.15s`, so the measurement was taken mid-animation, on a colour no state
actually has. `_settled()` now polls until the computed edge colour is the same
on two consecutive reads before measuring (polled, not slept, per this harness's
own rule). After the fix the same edges read exactly `rgb(122, 122, 170)`
(`--border-bright`) and `rgb(150, 119, 58)` (`--accent-dim`).
---
## H. Before and after, measured in the browser
Both passes were taken the same way — `--only boundaries --no-narrator`, the
production build — with only `tokens.css` differing. Evidence under
`$HOME/v11-evidence/wp-e/before/` and `.../after/`.
| Control (state) | Before | After | Floor |
| --- | --- | --- | --- |
| Story composer — resting edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
| Story control — resting edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
| Open panel tab — resting edge | **1.75:1** FAIL | **4.55:1** pass | 3.0 |
| Story control — hover edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
| Top navigation — translucent edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
| Story composer — focused edge | 4.70:1 pass | 4.70:1 pass | 3.0 |
**Suite result: before 5 passed / 5 failed; after 10 passed / 0 failed / 0
skipped.**
Two things this table says that a summary would blur:
- **Focus was never the defect.** The focused edge (`--accent-dim`) already
cleared 3:1 in v1.0.0 at 4.70:1, and WP-E did not change it. What failed was
rest and hover — the states a reader spends all their time in.
- **The translucent edge is real evidence.** The nav's background composited to
`rgb(17, 17, 29)` — `--bg-panel-glass` (rgba 19,19,32 @ 0.82) over
`rgb(10, 10, 15)` — not a fallback. That is the case token arithmetic cannot
reach, and it moved from 1.43:1 to 3.70:1.
## I. Screenshots
| File | |
| --- | --- |
| `before/control-boundaries.png` | 131,175 bytes, 1366×682 |
| `before/control-boundaries-nav.png` | 63,376 bytes, 1366×682 |
| `after/control-boundaries.png` | 132,430 bytes, 1366×682 |
| `after/control-boundaries-nav.png` | 63,541 bytes, 1366×682 |
All four are PNG, 1366×682, taken on the production build through the same
harness path, differing only in `tokens.css`. The play-page pair shows the
composer, the story controls and the open panel tab; the nav pair shows the
translucent top edge on the library route.
## J. A finding this package created and fixed
Raising `--border-bright` broke something that had nothing to do with control
boundaries. `.slice-7` in the context inspector's token breakdown was painted
with `var(--border-bright)`, so it followed the token to `#7a7aaa` — an OKLab ΔE
of **0.035** from `.slice-6` (`#7c86b8`), making two neighbouring chart slices
effectively the same colour. The other slices sit **0.100–0.119** from their
nearest neighbour.
`.slice-7` is now pinned to `#3d3d55`, the literal value it already rendered, so
its appearance is unchanged from v1.0.0 and its separation (ΔE **0.251**) is the
widest in the set. A chart fill and a control edge have different jobs and should
not share a token.
Reassigning it to a fresh hue was considered and rejected on evidence: inside the
palette's own chroma (0.045–0.120) and lightness (0.586–0.804) bands, the only
hues clearing the set's 0.100 separation floor are pinks near 14°, which is
`--danger`'s territory. Painting an ordinary prompt section in the colour this
application reserves for failure would trade an accessibility fix for a semantic
lie.
*(A first attempt, `#5d7f9e`, was rejected by the same measurement at ΔE 0.061 —
below every real slice. It is recorded here because it was written into the file
before it was measured.)*
## K. Text contrast regression
Unchanged, and asserted so in §F: **14.57:1** body text on the page, **13.57:1**
in a panel, **5.48:1** secondary text in a panel, **5.88:1** on the page. Every
text pair still clears 1.4.3, and no text token was touched.
## L. Browser release regression
The full harness, all three suites, against a real narrator on the production
build. Evidence: `$HOME/v11-evidence/wp-e/release/browser-report.json`.
| | |
| --- | --- |
| Kind | **`release regression`** — not `partial`, not `development (only …)` |
| Narrator | `qwen2.5:3b-instruct`, the reference model, over plain HTTP on the LAN GPU host |
| Browser | Firefox 155.0.1, geckodriver 0.37.1 |
| Served | FastAPI on loopback, the **built** SPA (`dist` 2026-09-15T21:08:52) |
| Duration | 71 s |
| Turns played | **8**, of which `turns_not_clean` **0** and `protocol_shapes_in_narration` **0** |
**Counted per suite, as the brief requires:**
| Suite | Passed | Failed | Skipped |
| --- | --- | --- | --- |
| **M11** (the v1 release regression) | **38** | **0** | **0** |
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
| **WP-E** (control boundaries) | **10** | **0** | **0** |
| **Total** | **101** | **0** | **0** |
M11's 38 and WP-C's 53 are unchanged in count and in name: WP-E added a suite
beside them rather than altering either. The `kind` field is quoted above
because a fast run invites the question — 71 s for 101 checks including 8
narrated turns is the GPU host being quick with a 3B model, and the 8 recorded
turns with no unclean accounting are what rule out narration having been
skipped.
The WP-E rows in this run are the same ten as the standalone capture in §H,
re-measured with a narrator present and a full story on the page.
## M. Full regression
| Suite | Result |
| --- | --- |
| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,054.7 s) |
| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
| **Production build** (`npm run build`) | succeeded |
| **Browser harness** (M11 + WP-C + WP-E) | **101 passed, 0 failed, 0 skipped** (§L) |
| **Contrast audit** (`python tools/contrast_audit.py`) | **exit 0** — every text pair and every boundary pair passes |
**The backend count reconciles exactly.** WP-D's tree was 1,712; WP-E adds the
11 in `test_v11_e_contrast.py`. 1,712 + 11 = **1,723**. The 17 skips are the same
environment-gated real-model tests recorded since B.1 — WP-E used no model and
added no skip.
**The plan's regression requirements for WP-E** were the frontend suite and
lint, and the harness's accessibility checks. All three pass: 175 and exit 0
above, and M11's `A11y` rows — accessible names, visible keyboard focus, no
positive tabindex, nothing revealed only on hover, and the four rendered text
contrasts — are inside the 38/38 in §L.
**Lint detail.** The 15 warnings are the same pre-existing
`only-export-components` and unused-import kind recorded at WP-C and WP-D, and
**none is in a file WP-E changed**. `tokens.css` and `context.css` are not
flagged.
## N. Residual risks
1. **The screenshots are unapproved.** The measurements say every boundary now
clears 3:1; whether the result *looks* right in this design is a judgment
the numbers cannot make. Recorded as PENDING in §O, not assumed.
2. **Five controls are measured in the browser; the rest inherit.** The audit
checks token pairs, and the harness measures the composer, a story control,
the open panel tab, the nav edge and the focused composer. Every other
bordered surface — modals, cards, the knowledge and context panels — draws
the same two tokens, so it moves with them, but none is individually
measured. A component that overrides a border with a literal colour would
not be caught by either check.
3. **`--bg-panel-glass` is outside the token audit by nature.** It is rgba over
a gradient, so no token pair can express it; its edge is covered only by the
browser measurement, which runs in the harness rather than in CI.
4. **Disabled controls are deliberately not measured.** `.story-controls
button:disabled` carries `opacity: 0.35`, so a disabled control's rendered
edge is dimmer than any value here. WCAG 1.4.11 exempts inactive components,
and the harness selects `:not(:disabled)` on purpose — stated so that the
exclusion is visible rather than looking like an oversight.
5. **The chart set was re-checked only where WP-E disturbed it.** `.slice-7`'s
separation and colour-blind distance were measured against the other seven
(§J); the set as a whole was not re-audited, which is outside this package.
6. **`a11y.test.jsx` was not extended**, though the plan listed it as likely
affected. It asserts structure and names, not colours, and adding colour
assertions in jsdom would test the stylesheet's text rather than a rendered
result. Boundary contrast is asserted instead where it can be measured: the
gate tests (§F) and the browser (§G).
**One risk that turned out not to exist.** §G's rule — a boundary passes on the
better of its two adjacent colours — was written to avoid failing an edge that
contrasts with the page but matches its own fill. In the event it never did any
work: every measured boundary clears 3:1 against **both** neighbours (composer
4.55/4.88, story control 3.44/3.70, panel tab 4.24/4.55, hover 4.55/4.88, focus
4.38/4.70, nav 3.51/3.70). The stricter reading would have produced the same
verdict on every row.
## O. Final decision
**Against the plan's acceptance criteria** (§ WP-E, *Control-boundary contrast*):
| # | Criterion | Verdict |
| --- | --- | --- |
| 1 | `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and 0 on the package tree | **PASS** — 2.99:1 exits 1 and 3.00:1 exits 0 (§F); the v1.0.0 value exits 1; the package tree exits 0 |
| 2 | Every text pair still clears 4.5:1, and the rendered text contrasts do not fall below v1's 14.57 / 5.48 / 13.57 / 5.88 | **PASS** — asserted as exact baselines in §F and re-measured on the rendered page in §L's M11 rows. No text token changed |
| 3 | The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1; the focus indicator is still visible | **PASS** — composer 4.88:1, story control 3.70:1 at rest and 4.88:1 on hover (§H), and M11's visible-focus check passes in §L |
| 4 | The owner approves the before-and-after screenshots, and the report records the approval | **PENDING** — not a criterion I can satisfy. The four PNGs are in §I |
**Where I exceeded the criterion, said plainly.** The plan asks for 3:1 "against
their panel" and names two controls. This package measures against
**`--bg-input`** as well — the lighter background inputs and buttons are really
drawn on, and the one that decides the gate — and adds the open panel tab, the
translucent navigation edge and the focused state. The stricter floor is the
reason the two threshold tests in §F are decided by `--bg-input` rather than
`--bg-panel`, where both would have passed.
**What I got wrong and corrected.** Three of my own claims failed checking and
were fixed rather than softened: the first hover and focus measurements were
taken mid-transition and reported colours no state has (§G); two contrast
figures were written into `context.css` before being measured, and were wrong
(§J); and my first gate check reported the current tokens as failing because my
harness crashed on a shallow path, not because of any contrast.
**Not done, and not claimed:** owner approval (criterion 4), individual
measurement of every bordered component (§N.2), and any palette work beyond the
two boundary tokens and the one chart fill that borrowed from them.
```text
WP-E CONTRAST GATE: PASS
WP-E BOUNDARY MEASUREMENT: PASS
WP-E OVERALL:
PASS, pending owner approval of the screenshots (criterion 4)
```
All WP-E changes are **staged and uncommitted**. No commit, no push, no tag.
v1.1 release validation has not begun.
```text
OWNER SCREENSHOT APPROVAL: PENDING
```
The before/after screenshots in §I are the evidence for a change a reader judges
by looking at it. The measurements say every boundary now clears 3:1; whether
the result looks right in this design is the owner's call, and it is recorded as
pending rather than assumed.