Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed

The planning package still described M01 as outstanding. It now records the
evidence run on 96c1bf5 and the two product defects found on the way. It also
corrects three statements that were never true.

- V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as
  recovered through authoritative state, with the owner's acceptance of that on
  2026-09-13 and the positional precondition explained.
  Correction: v3.7 said this file carried M11 results against every REQUIRED
  test. None were written, and the per-test matrix is the M11 report's §F. The
  §P3 M11 disposition said the report records the identity diagnostic's
  findings. It does not, and the disposition now says so.
- BUILD-MILESTONES.md: the M11 status block records the long-run evidence,
  the write-lock and protocol-leak defects, and what is left for the reviewer.
- DATA-MODEL.md §28B: M11 added two columns, not one.
  settings.context_window_override (migration 94, ef25b0a) was never
  recorded.
- TECHNICAL-DESIGN.md: "Background failure observability" gains the rule
  that nothing in a turn writes before the model call, and new §15.4 records
  that stored narration carries story only, with the extractor's rules.
- CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same
  two fixes.
- README.md and VERSION.md: status, milestone map, stop rule, and the v3.9
  entry.
- M11 report §Q: the "not revised" note is replaced by what v3.9 revised.

No requirement changes. No code changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-14 03:22:14 -04:00
co-authored by Claude Opus 5
parent d1988065e5
commit 3652dc6fae
9 changed files with 206 additions and 18 deletions
+29 -1
View File
@@ -1449,7 +1449,7 @@ Release gate in `V1-ACCEPTANCE-TESTS.md`:
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07, awaiting independent review/acceptance
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07; long-run evidence complete 2026-09-13; awaiting independent review/acceptance
Implemented on `m11-release-validation` from the signed M10 commit `1013c94`.
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package, written
@@ -1496,6 +1496,34 @@ Firefox refuses a WebDriver file path under `/tmp`, not all paths. Staging under
`$HOME` makes browser file import work, so knowledge import is now proved
end-to-end in a real browser rather than in two labelled halves.
**The long-run evidence (2026-09-13).** The first report left M01 PARTIAL at 41
turns, and that run was then lost to a host crash with its evidence. M01-M04 now
pass on a complete 100-turn run on `96c1bf5`: 101 accepted turns, three genuine
restarts, all thirteen scheduled history operations, zero failed post-turn
passes, and recovery onto a clean data directory 16 of 16. M04's precondition is
positional: the planting turn was outside the history window. The fact was
recovered through authoritative state and reached memory only as the narrator's
restatement of it. The repository owner accepted that on 2026-09-13. The
report's §G tells the whole path, including six runs that are not the evidence.
**Two further product defects**, found only once a long run had the memory bank
switched on:
3. **A turn locked its own memory bank out.** Retrieval wrote a use counter before
the model call, and the turn committed after the reply, so SQLite's one write
lock was held through the reply. Post-turn memory and summary writes timed out,
and so did recording their failure. The counter is now written in the turn's
own commit, and a failure is recorded after a rollback. (`f8d4010`)
4. **The narrator's protocol was stored as story.** The model wrote a pasted copy
of the state section, unfenced and unfinished proposals, and sections of its
own, on up to 42 of 104 turns, and stored text is replayed as history. The
extractor now removes every shape observed. Replaying 443 real turns through it
changed no turn it had previously left clean. (`0c7316f`, `96c1bf5`)
**Left for the reviewer**, in the report's §P: the real-token headroom at the
largest prompts is 23-42 tokens; the narrator restates prompt text in its prose;
and the identity diagnostic's results are not in the report.
---
## 4. Milestone Dependency Summary