Say where release evidence must be written, and what the crash took

The M11 harness examples wrote to /tmp, which a reboot clears. One
100-turn campaign was lost that way at 97 turns. The examples now
write under $HOME, which is also the only place the browser harness
works. G.0 records what the run reached, that its evidence is gone,
and that M01 must be re-run before acceptance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015aH3G73fEdh4qTQdZNUwty
This commit is contained in:
JesseMarkowitz
2026-09-07 14:19:27 -04:00
co-authored by Claude Opus 5
parent 144406cd48
commit fedb7144d0
2 changed files with 44 additions and 7 deletions
+15 -5
View File
@@ -351,28 +351,38 @@ behind `planning/reports/M11-IMPLEMENTATION-REPORT.md`, and they live in the
repository so a reviewer can re-run them rather than take the report's word for
anything. None is part of the application and none is imported by it.
**Write their output somewhere durable, never `/tmp`.** `--out` is required on
every harness precisely so the location is a decision rather than a default, and
the examples below use `$HOME/m11-evidence`. A reboot clears `/tmp`, and a
long-run campaign is hours of evidence that cannot be reproduced by re-reading a
file — one run was lost exactly that way (see the addendum at §G.0 of the M11
report). Snap Firefox independently refuses a WebDriver file path under `/tmp`
and needs one under `$HOME`, so `$HOME` is the only location the browser harness
works from in any case.
```bash
cd backend
mkdir -p "$HOME/m11-evidence"
# The 100-turn release campaign (M01-M04): real narrator, genuine process
# restarts, every history operation. Hours, not minutes.
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 \
AIDND_TEST_MODEL=<model> AIDND_TEST_EMBED_MODEL=<embedding model> \
.venv/bin/python -m tools.m11_long_run --turns 100 --out /tmp/m01
.venv/bin/python -m tools.m11_long_run --turns 100 --out "$HOME/m11-evidence/m01"
# What that campaign is worth on a machine that has never seen it (I01-I07).
.venv/bin/python -m tools.m11_recovery --bundle /tmp/m01/bundle.json --out /tmp/m01
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"
# The browser release regression and the accessibility measurements. Needs
# `frontend/dist` built and geckodriver on PATH.
.venv/bin/python -m tools.m11_browser --out /tmp/browser
.venv/bin/python -m tools.m11_browser --out "$HOME/m11-evidence/browser"
# A container with no network at all: the offline run and the packaging path.
.venv/bin/python -m tools.m11_offline --out /tmp/offline
.venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline"
# The multi-character identity diagnostic (post-M8 finding D), and the run that
# proves its detectors fire.
.venv/bin/python -m tools.m11_identity --out /tmp/identity
.venv/bin/python -m tools.m11_identity --out "$HOME/m11-evidence/identity"
.venv/bin/python -m tools.m11_identity --scripted --inject
# The palette, against WCAG AA.
+29 -2
View File
@@ -23,7 +23,9 @@ directory, schema parity, trusted-LAN HTTPS inference, and the full automated
suites — all green, all on one frozen tree, all with the evidence located in §F.
**What does not.** **M01, the 100-turn campaign, is PARTIAL: 41 accepted turns
at the time of writing, still running.** It is correct as far as it has gone —
at the time of writing, still running.** *(Addendum: that run later reached 97 of
100 turns and was then lost, with all of its evidence, to a host crash. It must
be re-run. See §G.0.)* It is correct as far as it has gone —
every history operation performed, the restart byte-identical, the prompt
bounded, the window verified on every turn, the campaign recoverable on another
machine — but 100 turns needs roughly four hours of wall clock on this CPU-only
@@ -531,7 +533,32 @@ blockers and no part of this milestone treats them as one.
**M01 is the one REQUIRED test this report cannot certify, and this section says
exactly how far it got and why.**
### G.0 Status: PARTIAL — 41 accepted turns of 100, and still running
### G.0 Status: PARTIAL — 41 accepted turns of 100; the run was later lost
> **Addendum, added 2026-09-07, after this section was written.** This section
> was drafted while the run was still in progress, and two things have happened
> since. Both count against this report's evidence.
>
> **The run went further, then died.** It continued unattended past the 41 turns
> described below and reached **97 of 100 accepted turns**, with all thirteen
> scheduled history operations fired. It never finished: the host crashed and
> rebooted, and the run went with it.
>
> **Its evidence did not survive.** The run wrote to a directory under `/tmp`,
> which the reboot cleared. `campaign.db`, `timeline.jsonl`, `server.log` and the
> backups are gone — **both the 41-turn artifacts this section cites and the
> 97-turn continuation.** The figures in §G.1-§G.5 and §N were transcribed from
> that run while it was live and are reported here as they were observed, but
> they can no longer be produced on request and nothing in them can be
> independently re-checked.
>
> **What this means for acceptance.** M01 must be **re-run from scratch** before
> v1 acceptance, on a host that can finish it, writing to a durable path rather
> than `/tmp` (`DEVELOPMENT.md` now says so and the harnesses require `--out`).
> Until that run exists, treat every M01-derived number in this report as an
> unverifiable observation rather than as evidence. Nothing else in the report
> depends on it: the suite, browser, offline, recovery and migration results were
> produced by harnesses that can be re-run in minutes.
The campaign is correct as far as it has gone: every history operation performed
as designed, the restart was byte-identical, the prompt stayed bounded, the