Files
JesseMarkowitzandClaude Opus 5 db7b309e3d
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
v1.1 closeout: accept integrated release validation
Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.

V1.1 RELEASE VALIDATION: PASS

What was run, on this candidate:

- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
  0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
  175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
  to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
  trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
  3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
  keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
  scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
  that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
  identical on all 15 census fields, schema parity at user_version 94, and both
  bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
  verified, public endpoint refused, a real turn, restart, persistence, and
  Firefox rendering the reopened campaign.

Carried residuals, stated rather than summarised away:

- WP-B: deterministic independent-memory recovery PASS; reference-model
  independent-memory recovery FAIL at memory creation — the owner-accepted
  limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
  reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
  release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
  backlog, reproduced and not fixed during validation.

Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.

Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.

Still the owner's to do: sign the release commit, update main, tag v1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-16 07:12:23 -04:00

381 lines
19 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# v1.1 WP-E — Control-Boundary Contrast
**Status:** COMPLETE — **PASS**, with owner approval of the screenshots
outstanding. The decision, and what was deliberately not claimed, is in §O.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD | `59b5ebc` — *v1.1 WP-C: browser release coverage*, signed by the owner |
| Working tree at start | WP-D staged (10 files), nothing committed |
| Criterion | WCAG 2.1 **1.4.11 Non-text Contrast**, 3:1, for control boundaries; **1.4.3** 4.5:1 for body text, unchanged |
---
## B. What v1.0.0 actually did
`tools/contrast_audit.py` measured control boundaries, printed that two of them
were below 3:1, and **exited 0**. Its own comment argued the position:
> in this design a control is identified by its *label*, which is measured above
> and passes, not by its edge. So a boundary below 3:1 is reported with its
> number and does not fail the run.
So the audit was a report, not a gate: no palette change could ever fail it on a
boundary. The two numbers it printed were **1.33:1** (`--border` on
`--bg-panel`) and **1.75:1** (`--border-bright`), against a floor of 3.0.
**WP-E overturns that argument.** 1.4.11 covers the visual information needed to
identify a component *and its boundary*; a reader who cannot see where a text box
ends cannot see that there is a text box to type into, label or no label. The
tokens were raised rather than the criterion re-argued.
---
## C. Token inventory
| Token | v1.0.0 | v1.1 | Why |
| --- | --- | --- | --- |
| `--border` | `#2b2b3d` | **`#676792`** | every control's resting edge |
| `--border-bright` | `#3d3d55` | **`#7a7aaa`** | hover edges, the composer's resting edge, the open panel tab |
| `--bg-panel`, `--bg-input`, `--bg`, `--text`, `--text-dim`, `--accent*`, `--danger`, `--warning`, `--player`, `--chart-*` | — | **unchanged** | WP-E is a boundary package; no text or accent colour moved |
**The floor is taken against `--bg-input`, not `--bg-panel`.** Inputs and buttons
are drawn on `--bg-input` (`styles/forms.css`), which is lighter than
`--bg-panel` and therefore the harder case. The audit had been checking only
`--bg-panel`, so a token could have passed the audit while the real control
failed. Measured on the new values:
| | vs `--bg-input` | vs `--bg-panel` | vs `--bg` |
| --- | --- | --- | --- |
| `--border` | **3.21:1** | 3.44:1 | 3.70:1 |
| `--border-bright` | **4.24:1** | 4.55:1 | 4.88:1 |
Two properties were preserved deliberately: the rest→hover step is the same size
as before (1.318 → 1.320), so hover still reads as a change rather than a jump;
and `--border-bright` stays *below* body text against the same panel (2.98:1
between them), so no edge outshines the words inside it.
---
## D. Component inventory
Where these tokens are actually drawn, from the stylesheets:
| Control | Rule | Rest | Hover / active | Focus |
| --- | --- | --- | --- | --- |
| Story composer | `.input-bar` (story.css) | `--border-bright` on `--bg-panel` | — | `--accent-dim` + `--accent-glow` ring |
| Story controls | `.story-controls button` | `--border` on `--bg-panel` | `--border-bright` | (M11 focus check) |
| Fields and buttons | `forms.css` | `--border` on `--bg-input` | `--accent-dim` | `--accent-dim` + ring |
| Panel tabs | `.panel-tabs button` | **`transparent`** | `--border` on hover, `--border-bright` when active | — |
| Top navigation | `.topnav` | `--border` bottom edge on **`--bg-panel-glass`** | — | — |
| Scrollbar thumb | `base.css` | `--border-bright` as a *fill* on `--bg` | `--accent-dim` | — |
Two of these cannot be answered by token arithmetic at all, and both are
measured in the browser instead (§G): the nav sits on a translucent panel, and
the panel tab's edge is `transparent` until the panel is open.
---
## E. The audit is now a gate
`tools/contrast_audit.py`:
1. **Boundary pairs fail.** `text` and `boundary` rows are both pass/fail; the
advisory branch is gone. The run returns 1 if either kind falls short.
2. **Eight boundary pairs replace two.** Each border is checked against every
background it is drawn on — `--bg-input`, `--bg-panel` and `--bg` — plus the
focused edge (`--accent-dim`) on both panel and field backgrounds.
3. **The verdict is taken on the number that is printed** (rounded to two
decimals), so a pair shown as `3.00:1` is not failed for arithmetic the
reader cannot see.
4. The comment block that argued the old position is replaced by one recording
what changed and why, including why `--bg-panel-glass` is not in the list.
## F. Gate tests
`backend/tests/test_v11_e_contrast.py` — **11 passed**. The threshold is
exercised from both sides, on real token files:
| Test | Result |
| --- | --- |
| A boundary at **2.99:1** against `--bg-input` fails the run (exit 1) | PASS |
| A boundary at **3.00:1** passes (exit 0) | PASS |
| The **v1.0.0 value** `#2b2b3d` fails, at 1.24:1 against `--bg-input` | PASS |
| A dimmed `--text-dim` still fails as a *text* pair | PASS |
| A renamed token is a failure, not a silent skip | PASS |
| The shipped palette passes both criteria | PASS |
| Every boundary pair is measured against the background it is drawn on | PASS |
| The M11 text baselines are unchanged: 14.57 / 13.57 / 5.48 / 5.88 | PASS |
| The hover edge stays brighter than the resting edge | PASS |
| No boundary becomes as loud as body text | PASS |
| The WCAG ratio formula is anchored on known values (21:1, 1:1, symmetry) | PASS |
The 2.99 and 3.00 values are worth noting: **both clear 3:1 against
`--bg-panel`** (3.21 and 3.22). They decide the gate only because the floor is
now taken against the background the control is really on — so these two tests
also prove §C's change is doing work.
---
## G. Browser measurement
`tools/m11_browser.py` gains a third suite, **WP-E**, counted separately from
M11's 38 and WP-C's 53. It measures the *rendered* edge — `borderColor` from
`getComputedStyle` — against what is actually behind it, with every translucent
layer composited bottom-up.
A boundary is measured against **both** adjacent colours (the control's own fill
inside it, the background outside it) and passes on the better of the two: an
edge that matches its fill but contrasts with the page is still a visible
outline. What 1.4.11 asks is that the component's extent be perceivable.
Two harness capabilities were added for this (`tools/m11_webdriver.py`):
- **`hover()`** moves a real pointer through the WebDriver Actions API.
Dispatching a `mouseover` event from JavaScript does *not* trigger CSS
`:hover`, so a synthetic event would have re-measured the resting edge and
reported it as the hover edge.
- **`screenshot()`** writes the viewport as a PNG, for the before/after evidence.
**A defect this found in my own first measurement.** The first run reported the
hover edge as `rgb(114, 114, 160)` and the focused edge as `rgb(144, 120, 81)` —
neither of which is any token. Both controls carry `transition: border-color
0.15s`, so the measurement was taken mid-animation, on a colour no state
actually has. `_settled()` now polls until the computed edge colour is the same
on two consecutive reads before measuring (polled, not slept, per this harness's
own rule). After the fix the same edges read exactly `rgb(122, 122, 170)`
(`--border-bright`) and `rgb(150, 119, 58)` (`--accent-dim`).
---
## H. Before and after, measured in the browser
Both passes were taken the same way — `--only boundaries --no-narrator`, the
production build — with only `tokens.css` differing. Evidence under
`$HOME/v11-evidence/wp-e/before/` and `.../after/`.
| Control (state) | Before | After | Floor |
| --- | --- | --- | --- |
| Story composer — resting edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
| Story control — resting edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
| Open panel tab — resting edge | **1.75:1** FAIL | **4.55:1** pass | 3.0 |
| Story control — hover edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
| Top navigation — translucent edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
| Story composer — focused edge | 4.70:1 pass | 4.70:1 pass | 3.0 |
**Suite result: before 5 passed / 5 failed; after 10 passed / 0 failed / 0
skipped.**
Two things this table says that a summary would blur:
- **Focus was never the defect.** The focused edge (`--accent-dim`) already
cleared 3:1 in v1.0.0 at 4.70:1, and WP-E did not change it. What failed was
rest and hover — the states a reader spends all their time in.
- **The translucent edge is real evidence.** The nav's background composited to
`rgb(17, 17, 29)` — `--bg-panel-glass` (rgba 19,19,32 @ 0.82) over
`rgb(10, 10, 15)` — not a fallback. That is the case token arithmetic cannot
reach, and it moved from 1.43:1 to 3.70:1.
## I. Screenshots
| File | |
| --- | --- |
| `before/control-boundaries.png` | 131,175 bytes, 1366×682 |
| `before/control-boundaries-nav.png` | 63,376 bytes, 1366×682 |
| `after/control-boundaries.png` | 132,430 bytes, 1366×682 |
| `after/control-boundaries-nav.png` | 63,541 bytes, 1366×682 |
All four are PNG, 1366×682, taken on the production build through the same
harness path, differing only in `tokens.css`. The play-page pair shows the
composer, the story controls and the open panel tab; the nav pair shows the
translucent top edge on the library route.
## J. A finding this package created and fixed
Raising `--border-bright` broke something that had nothing to do with control
boundaries. `.slice-7` in the context inspector's token breakdown was painted
with `var(--border-bright)`, so it followed the token to `#7a7aaa` — an OKLab ΔE
of **0.035** from `.slice-6` (`#7c86b8`), making two neighbouring chart slices
effectively the same colour. The other slices sit **0.100–0.119** from their
nearest neighbour.
`.slice-7` is now pinned to `#3d3d55`, the literal value it already rendered, so
its appearance is unchanged from v1.0.0 and its separation (ΔE **0.251**) is the
widest in the set. A chart fill and a control edge have different jobs and should
not share a token.
Reassigning it to a fresh hue was considered and rejected on evidence: inside the
palette's own chroma (0.045–0.120) and lightness (0.586–0.804) bands, the only
hues clearing the set's 0.100 separation floor are pinks near 14°, which is
`--danger`'s territory. Painting an ordinary prompt section in the colour this
application reserves for failure would trade an accessibility fix for a semantic
lie.
*(A first attempt, `#5d7f9e`, was rejected by the same measurement at ΔE 0.061 —
below every real slice. It is recorded here because it was written into the file
before it was measured.)*
## K. Text contrast regression
Unchanged, and asserted so in §F: **14.57:1** body text on the page, **13.57:1**
in a panel, **5.48:1** secondary text in a panel, **5.88:1** on the page. Every
text pair still clears 1.4.3, and no text token was touched.
## L. Browser release regression
The full harness, all three suites, against a real narrator on the production
build. Evidence: `$HOME/v11-evidence/wp-e/release/browser-report.json`.
| | |
| --- | --- |
| Kind | **`release regression`** — not `partial`, not `development (only …)` |
| Narrator | `qwen2.5:3b-instruct`, the reference model, over plain HTTP on the LAN GPU host |
| Browser | Firefox 155.0.1, geckodriver 0.37.1 |
| Served | FastAPI on loopback, the **built** SPA (`dist` 2026-09-15T21:08:52) |
| Duration | 71 s |
| Turns played | **8**, of which `turns_not_clean` **0** and `protocol_shapes_in_narration` **0** |
**Counted per suite, as the brief requires:**
| Suite | Passed | Failed | Skipped |
| --- | --- | --- | --- |
| **M11** (the v1 release regression) | **38** | **0** | **0** |
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
| **WP-E** (control boundaries) | **10** | **0** | **0** |
| **Total** | **101** | **0** | **0** |
M11's 38 and WP-C's 53 are unchanged in count and in name: WP-E added a suite
beside them rather than altering either. The `kind` field is quoted above
because a fast run invites the question — 71 s for 101 checks including 8
narrated turns is the GPU host being quick with a 3B model, and the 8 recorded
turns with no unclean accounting are what rule out narration having been
skipped.
The WP-E rows in this run are the same ten as the standalone capture in §H,
re-measured with a narrator present and a full story on the page.
## M. Full regression
| Suite | Result |
| --- | --- |
| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,054.7 s) |
| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
| **Production build** (`npm run build`) | succeeded |
| **Browser harness** (M11 + WP-C + WP-E) | **101 passed, 0 failed, 0 skipped** (§L) |
| **Contrast audit** (`python tools/contrast_audit.py`) | **exit 0** — every text pair and every boundary pair passes |
**The backend count reconciles exactly.** WP-D's tree was 1,712; WP-E adds the
11 in `test_v11_e_contrast.py`. 1,712 + 11 = **1,723**. The 17 skips are the same
environment-gated real-model tests recorded since B.1 — WP-E used no model and
added no skip.
**The plan's regression requirements for WP-E** were the frontend suite and
lint, and the harness's accessibility checks. All three pass: 175 and exit 0
above, and M11's `A11y` rows — accessible names, visible keyboard focus, no
positive tabindex, nothing revealed only on hover, and the four rendered text
contrasts — are inside the 38/38 in §L.
**Lint detail.** The 15 warnings are the same pre-existing
`only-export-components` and unused-import kind recorded at WP-C and WP-D, and
**none is in a file WP-E changed**. `tokens.css` and `context.css` are not
flagged.
## N. Residual risks
1. **The screenshots are unapproved.** The measurements say every boundary now
clears 3:1; whether the result *looks* right in this design is a judgment
the numbers cannot make. Recorded as PENDING in §O, not assumed.
2. **Five controls are measured in the browser; the rest inherit.** The audit
checks token pairs, and the harness measures the composer, a story control,
the open panel tab, the nav edge and the focused composer. Every other
bordered surface — modals, cards, the knowledge and context panels — draws
the same two tokens, so it moves with them, but none is individually
measured. A component that overrides a border with a literal colour would
not be caught by either check.
3. **`--bg-panel-glass` is outside the token audit by nature.** It is rgba over
a gradient, so no token pair can express it; its edge is covered only by the
browser measurement, which runs in the harness rather than in CI.
4. **Disabled controls are deliberately not measured.** `.story-controls
button:disabled` carries `opacity: 0.35`, so a disabled control's rendered
edge is dimmer than any value here. WCAG 1.4.11 exempts inactive components,
and the harness selects `:not(:disabled)` on purpose — stated so that the
exclusion is visible rather than looking like an oversight.
5. **The chart set was re-checked only where WP-E disturbed it.** `.slice-7`'s
separation and colour-blind distance were measured against the other seven
(§J); the set as a whole was not re-audited, which is outside this package.
6. **`a11y.test.jsx` was not extended**, though the plan listed it as likely
affected. It asserts structure and names, not colours, and adding colour
assertions in jsdom would test the stylesheet's text rather than a rendered
result. Boundary contrast is asserted instead where it can be measured: the
gate tests (§F) and the browser (§G).
**One risk that turned out not to exist.** §G's rule — a boundary passes on the
better of its two adjacent colours — was written to avoid failing an edge that
contrasts with the page but matches its own fill. In the event it never did any
work: every measured boundary clears 3:1 against **both** neighbours (composer
4.55/4.88, story control 3.44/3.70, panel tab 4.24/4.55, hover 4.55/4.88, focus
4.38/4.70, nav 3.51/3.70). The stricter reading would have produced the same
verdict on every row.
## O. Final decision
**Against the plan's acceptance criteria** (§ WP-E, *Control-boundary contrast*):
| # | Criterion | Verdict |
| --- | --- | --- |
| 1 | `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and 0 on the package tree | **PASS** — 2.99:1 exits 1 and 3.00:1 exits 0 (§F); the v1.0.0 value exits 1; the package tree exits 0 |
| 2 | Every text pair still clears 4.5:1, and the rendered text contrasts do not fall below v1's 14.57 / 5.48 / 13.57 / 5.88 | **PASS** — asserted as exact baselines in §F and re-measured on the rendered page in §L's M11 rows. No text token changed |
| 3 | The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1; the focus indicator is still visible | **PASS** — composer 4.88:1, story control 3.70:1 at rest and 4.88:1 on hover (§H), and M11's visible-focus check passes in §L |
| 4 | The owner approves the before-and-after screenshots, and the report records the approval | **PENDING** — not a criterion I can satisfy. The four PNGs are in §I |
**Where I exceeded the criterion, said plainly.** The plan asks for 3:1 "against
their panel" and names two controls. This package measures against
**`--bg-input`** as well — the lighter background inputs and buttons are really
drawn on, and the one that decides the gate — and adds the open panel tab, the
translucent navigation edge and the focused state. The stricter floor is the
reason the two threshold tests in §F are decided by `--bg-input` rather than
`--bg-panel`, where both would have passed.
**What I got wrong and corrected.** Three of my own claims failed checking and
were fixed rather than softened: the first hover and focus measurements were
taken mid-transition and reported colours no state has (§G); two contrast
figures were written into `context.css` before being measured, and were wrong
(§J); and my first gate check reported the current tokens as failing because my
harness crashed on a shallow path, not because of any contrast.
**Not done, and not claimed:** owner approval (criterion 4), individual
measurement of every bordered component (§N.2), and any palette work beyond the
two boundary tokens and the one chart fill that borrowed from them.
```text
WP-E CONTRAST GATE: PASS
WP-E BOUNDARY MEASUREMENT: PASS
WP-E OVERALL:
PASS, pending owner approval of the screenshots (criterion 4)
```
All WP-E changes are **staged and uncommitted**. No commit, no push, no tag.
v1.1 release validation has not begun.
```text
OWNER SCREENSHOT APPROVAL: APPROVED
```
The before/after screenshots in §I are the evidence for a change a reader judges
by looking at it. The measurements say every boundary now clears 3:1; whether the
result looks right in this design was the owner's call.
**Approved by the owner on 2026-09-16**, in the v1.1 release-validation brief,
after reviewing the before/after pair in `$HOME/v11-evidence/wp-e/`. This line
was `PENDING` in the signed commit `87a4032` because the report predated that
review; it is updated here as part of the release closeout, with its source and
date recorded rather than the approval being assumed. No visual code changed
during release validation, so the approval stands (release report §Q).