M11 report: M01 to M04 on a complete run, and what it took to get one

The first revision left M01 PARTIAL at 41 accepted turns, and that run was then
lost to a host crash with its evidence. This revision reports the 100-turn
evidence run on 96c1bf5. It had 101 accepted turns, three genuine restarts,
every scheduled history operation, zero failed post-turn passes, and recovery
onto a clean data directory 16 of 16. All 85 REQUIRED FOR V1 tests now pass,
with H09 NOT APPLICABLE on its own condition.

Rewritten: §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R. The §G.0
addendum is removed, and its history is §G.6.

- §G: the evidence run's timeline, context growth and recall. M04's
  precondition is positional: the planting turn at depth 1, the window floor
  at 54. The path is stated plainly: authoritative state, then the narrator
  restating the fact line in its prose, then memory summarising those
  restatements. The owner accepted state-based recovery on 2026-09-13.
- §G.6: all seven long runs, and why six of them are not the evidence.
- §O.7 and §O.8: the write-lock defect and the protocol-leak defect. Five
  harness defects are added to the harness table.
- §N: storage for the evidence run, and real-token headroom by the narrator's
  own tokenizer: 92, 23, 34 and 42 tokens across four runs, with the
  inference server's silent cut to 8,194 tokens stated.
- §E.1: the application machine, the CPU reference host and the GPU host,
  identical model digests, and the GPU dropping off the PCIe bus (Xid 79)
  about 30 s after the evidence run's last write. That long runs must log
  power, link state and kernel messages is recorded as a requirement.
- §F, §J, §R: A06, H01 and R6 no longer claim that every turn went over HTTPS.
  The GPU runs used plain HTTP to a LAN host and were not network-monitored.
- §B, §P: the black-box runs (browser, offline, identity) were re-run on the
  ef25b0a tree and not after the three later backend commits.
- §C, §K: the later commits, including migration 94 from ef25b0a.
- §D.1, §P: the identity diagnostic's results were never written into this
  report; the section the first revision pointed to was empty.

DEVELOPMENT.md: the pointer to the removed §G.0 is replaced, and a new section,
"Logging the inference host during a long run", gives the nvidia-smi and
journalctl commands to run on a GPU host for every long run. If the GPU drops
again, the logs show whether it was power.

Not revised here: V1-ACCEPTANCE-TESTS.md, BUILD-MILESTONES.md, VERSION.md and
planning/README.md.

On 96c1bf5: backend 1,421 passed, 17 skipped, 0 failed; frontend 161 passed;
lint exit 0 with warnings only. No code changes in this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-14 03:15:13 -04:00
co-authored by Claude Opus 5
parent 96c1bf5ded
commit d1988065e5
2 changed files with 643 additions and 355 deletions
+39 -2
View File
@@ -385,8 +385,7 @@ anything. None is part of the application and none is imported by it.
every harness precisely so the location is a decision rather than a default, and
the examples below use `$HOME/m11-evidence`. A reboot clears `/tmp`, and a
long-run campaign is hours of evidence that cannot be reproduced by re-reading a
file — one run was lost exactly that way (see the addendum at §G.0 of the M11
report). Snap Firefox independently refuses a WebDriver file path under `/tmp`
file. One run was lost exactly that way; §G.6 of the M11 report records it. Snap Firefox independently refuses a WebDriver file path under `/tmp`
and needs one under `$HOME`, so `$HOME` is the only location the browser harness
works from in any case.
@@ -454,6 +453,44 @@ deployment a turn cost 229-291 seconds at the recommended window; a slower host
can exceed the 600 seconds this harness used to hard-code, and an overrun turn is
a lost turn.
### Logging the inference host during a long run
A long run is the heaviest sustained load an inference host sees. In M11 a GPU
host dropped its GPU off the PCIe bus (`NVRM: Xid 79`) half a minute after a
100-turn run finished. Nothing on disk could say whether power, heat or the link
caused it (M11 report, §E.1). **For every long run against a GPU host, start this
logging on that host first and stop it only when the run has finished.**
Run each command in its own terminal on the inference host. `tee` writes each
line as it arrives, so what happened in the seconds before a crash or a forced
reboot survives on disk.
```bash
# Power, temperature, utilisation and PCIe link state, once a second
nvidia-smi --query-gpu=timestamp,pcie.link.gen.current,pcie.link.width.current,power.draw,temperature.gpu,utilization.gpu \
--format=csv -l 1 | tee "$HOME/gpu-link-$(date +%F-%H%M).csv"
# The driver's own sampling: power, utilisation, clocks, memory, ECC and throttling
nvidia-smi dmon -s pucvmet -d 5 | tee "$HOME/gpu-dmon-$(date +%F-%H%M).log"
# Kernel and Ollama messages, live
journalctl -f -k -u ollama | tee "$HOME/ollama-kernel-watch-$(date +%F-%H%M).log"
```
If the GPU faults, find the moment and then read what the card was doing just
before it:
```bash
grep -iE 'xid|fallen off|nvrm' "$HOME"/ollama-kernel-watch-*.log
awk -F', ' 'NR>1 && $4+0 > max {max=$4+0; at=$1} END {print "peak W", max, "at", at}' "$HOME"/gpu-link-*.csv
```
A fault that follows sustained draw at the card's power limit points to power
delivery. A fault with the link below its usual generation under load points to
the connection. A fault with neither is still worth recording, because it rules
both out. These commands were verified against NVIDIA driver 580 and Ollama
0.34.
`tools/m11_webdriver.py` is the W3C WebDriver client the browser harness uses.
It exists so browser evidence needs no Selenium in the dependency surface, and
it documents the one environment quirk that matters here: a snap Firefox will