M11 report: M01 to M04 on a complete run, and what it took to get one
The first revision left M01 PARTIAL at 41 accepted turns, and that run was then lost to a host crash with its evidence. This revision reports the 100-turn evidence run on96c1bf5. It had 101 accepted turns, three genuine restarts, every scheduled history operation, zero failed post-turn passes, and recovery onto a clean data directory 16 of 16. All 85 REQUIRED FOR V1 tests now pass, with H09 NOT APPLICABLE on its own condition. Rewritten: §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R. The §G.0 addendum is removed, and its history is §G.6. - §G: the evidence run's timeline, context growth and recall. M04's precondition is positional: the planting turn at depth 1, the window floor at 54. The path is stated plainly: authoritative state, then the narrator restating the fact line in its prose, then memory summarising those restatements. The owner accepted state-based recovery on 2026-09-13. - §G.6: all seven long runs, and why six of them are not the evidence. - §O.7 and §O.8: the write-lock defect and the protocol-leak defect. Five harness defects are added to the harness table. - §N: storage for the evidence run, and real-token headroom by the narrator's own tokenizer: 92, 23, 34 and 42 tokens across four runs, with the inference server's silent cut to 8,194 tokens stated. - §E.1: the application machine, the CPU reference host and the GPU host, identical model digests, and the GPU dropping off the PCIe bus (Xid 79) about 30 s after the evidence run's last write. That long runs must log power, link state and kernel messages is recorded as a requirement. - §F, §J, §R: A06, H01 and R6 no longer claim that every turn went over HTTPS. The GPU runs used plain HTTP to a LAN host and were not network-monitored. - §B, §P: the black-box runs (browser, offline, identity) were re-run on theef25b0atree and not after the three later backend commits. - §C, §K: the later commits, including migration 94 fromef25b0a. - §D.1, §P: the identity diagnostic's results were never written into this report; the section the first revision pointed to was empty. DEVELOPMENT.md: the pointer to the removed §G.0 is replaced, and a new section, "Logging the inference host during a long run", gives the nvidia-smi and journalctl commands to run on a GPU host for every long run. If the GPU drops again, the logs show whether it was power. Not revised here: V1-ACCEPTANCE-TESTS.md, BUILD-MILESTONES.md, VERSION.md and planning/README.md. On96c1bf5: backend 1,421 passed, 17 skipped, 0 failed; frontend 161 passed; lint exit 0 with warnings only. No code changes in this commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
co-authored by
Claude Opus 5
parent
96c1bf5ded
commit
d1988065e5