diff --git a/DEVELOPMENT.md b/DEVELOPMENT.md index b80fba5..1133f0f 100644 --- a/DEVELOPMENT.md +++ b/DEVELOPMENT.md @@ -351,28 +351,38 @@ behind `planning/reports/M11-IMPLEMENTATION-REPORT.md`, and they live in the repository so a reviewer can re-run them rather than take the report's word for anything. None is part of the application and none is imported by it. +**Write their output somewhere durable, never `/tmp`.** `--out` is required on +every harness precisely so the location is a decision rather than a default, and +the examples below use `$HOME/m11-evidence`. A reboot clears `/tmp`, and a +long-run campaign is hours of evidence that cannot be reproduced by re-reading a +file — one run was lost exactly that way (see the addendum at §G.0 of the M11 +report). Snap Firefox independently refuses a WebDriver file path under `/tmp` +and needs one under `$HOME`, so `$HOME` is the only location the browser harness +works from in any case. + ```bash cd backend +mkdir -p "$HOME/m11-evidence" # The 100-turn release campaign (M01-M04): real narrator, genuine process # restarts, every history operation. Hours, not minutes. AIDND_TEST_ENDPOINT=https://:/v1 \ AIDND_TEST_MODEL= AIDND_TEST_EMBED_MODEL= \ - .venv/bin/python -m tools.m11_long_run --turns 100 --out /tmp/m01 + .venv/bin/python -m tools.m11_long_run --turns 100 --out "$HOME/m11-evidence/m01" # What that campaign is worth on a machine that has never seen it (I01-I07). -.venv/bin/python -m tools.m11_recovery --bundle /tmp/m01/bundle.json --out /tmp/m01 +.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01" # The browser release regression and the accessibility measurements. Needs # `frontend/dist` built and geckodriver on PATH. -.venv/bin/python -m tools.m11_browser --out /tmp/browser +.venv/bin/python -m tools.m11_browser --out "$HOME/m11-evidence/browser" # A container with no network at all: the offline run and the packaging path. -.venv/bin/python -m tools.m11_offline --out /tmp/offline +.venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline" # The multi-character identity diagnostic (post-M8 finding D), and the run that # proves its detectors fire. -.venv/bin/python -m tools.m11_identity --out /tmp/identity +.venv/bin/python -m tools.m11_identity --out "$HOME/m11-evidence/identity" .venv/bin/python -m tools.m11_identity --scripted --inject # The palette, against WCAG AA. diff --git a/planning/reports/M11-IMPLEMENTATION-REPORT.md b/planning/reports/M11-IMPLEMENTATION-REPORT.md index 30178f4..b9d1680 100644 --- a/planning/reports/M11-IMPLEMENTATION-REPORT.md +++ b/planning/reports/M11-IMPLEMENTATION-REPORT.md @@ -23,7 +23,9 @@ directory, schema parity, trusted-LAN HTTPS inference, and the full automated suites — all green, all on one frozen tree, all with the evidence located in §F. **What does not.** **M01, the 100-turn campaign, is PARTIAL: 41 accepted turns -at the time of writing, still running.** It is correct as far as it has gone — +at the time of writing, still running.** *(Addendum: that run later reached 97 of +100 turns and was then lost, with all of its evidence, to a host crash. It must +be re-run. See §G.0.)* It is correct as far as it has gone — every history operation performed, the restart byte-identical, the prompt bounded, the window verified on every turn, the campaign recoverable on another machine — but 100 turns needs roughly four hours of wall clock on this CPU-only @@ -531,7 +533,32 @@ blockers and no part of this milestone treats them as one. **M01 is the one REQUIRED test this report cannot certify, and this section says exactly how far it got and why.** -### G.0 Status: PARTIAL — 41 accepted turns of 100, and still running +### G.0 Status: PARTIAL — 41 accepted turns of 100; the run was later lost + +> **Addendum, added 2026-09-07, after this section was written.** This section +> was drafted while the run was still in progress, and two things have happened +> since. Both count against this report's evidence. +> +> **The run went further, then died.** It continued unattended past the 41 turns +> described below and reached **97 of 100 accepted turns**, with all thirteen +> scheduled history operations fired. It never finished: the host crashed and +> rebooted, and the run went with it. +> +> **Its evidence did not survive.** The run wrote to a directory under `/tmp`, +> which the reboot cleared. `campaign.db`, `timeline.jsonl`, `server.log` and the +> backups are gone — **both the 41-turn artifacts this section cites and the +> 97-turn continuation.** The figures in §G.1-§G.5 and §N were transcribed from +> that run while it was live and are reported here as they were observed, but +> they can no longer be produced on request and nothing in them can be +> independently re-checked. +> +> **What this means for acceptance.** M01 must be **re-run from scratch** before +> v1 acceptance, on a host that can finish it, writing to a durable path rather +> than `/tmp` (`DEVELOPMENT.md` now says so and the harnesses require `--out`). +> Until that run exists, treat every M01-derived number in this report as an +> unverifiable observation rather than as evidence. Nothing else in the report +> depends on it: the suite, browser, offline, recovery and migration results were +> produced by harnesses that can be re-run in minutes. The campaign is correct as far as it has gone: every history operation performed as designed, the restart was byte-identical, the prompt stayed bounded, the