Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed

The planning package still described M01 as outstanding. It now records the
evidence run on 96c1bf5 and the two product defects found on the way. It also
corrects three statements that were never true.

- V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as
  recovered through authoritative state, with the owner's acceptance of that on
  2026-09-13 and the positional precondition explained.
  Correction: v3.7 said this file carried M11 results against every REQUIRED
  test. None were written, and the per-test matrix is the M11 report's §F. The
  §P3 M11 disposition said the report records the identity diagnostic's
  findings. It does not, and the disposition now says so.
- BUILD-MILESTONES.md: the M11 status block records the long-run evidence,
  the write-lock and protocol-leak defects, and what is left for the reviewer.
- DATA-MODEL.md §28B: M11 added two columns, not one.
  settings.context_window_override (migration 94, ef25b0a) was never
  recorded.
- TECHNICAL-DESIGN.md: "Background failure observability" gains the rule
  that nothing in a turn writes before the model call, and new §15.4 records
  that stored narration carries story only, with the extractor's rules.
- CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same
  two fixes.
- README.md and VERSION.md: status, milestone map, stop rule, and the v3.9
  entry.
- M11 report §Q: the "not revised" note is replaced by what v3.9 revised.

No requirement changes. No code changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-14 03:22:14 -04:00
co-authored by Claude Opus 5
parent d1988065e5
commit 3652dc6fae
9 changed files with 206 additions and 18 deletions
+45 -2
View File
@@ -2358,6 +2358,17 @@ Run or automate at least 100 accepted turns with:
### Pass
No major continuity/state/history corruption.
### Result — PASS (M11, long-run evidence 2026-09-13)
101 accepted turns on commit `96c1bf5`, against a real narrator
(`qwen2.5:3b-instruct-16k`, its 16,384-token window verified on every turn). The
run covered every step above: the fixture's characters and locations; two Save
Points and a restore; Undo with divergence; two scheduled retries and a take
selection; three imported knowledge sources; and the memory bank and
auto-summarise on, with 33 memories and 12 summaries written and zero failed
post-turn passes. It also included an induced failed model call. No turn was
refused, nothing in state or history was corrupted, and the campaign moved to a
clean data directory 16 of 16. `tools/m11_long_run.py`; M11 report §G.
---
## M02 — Restart During Long Campaign
@@ -2369,6 +2380,13 @@ Restart application at several points during M01.
### Pass
Campaign resumes correctly.
### Result — PASS (M11, 2026-09-13)
Three genuine `uvicorn` process restarts inside the M01 run, at turns 13, 41 and
83. The restart at turn 41 carried undone history across the boundary. Each one
compared transcript, head, Undo/Redo availability, scene, entities, facts, Save
Points, imported knowledge and settings before and after, and found them
identical every time. M11 report §G.2.
---
## M03 — Long-Run Context Stability
@@ -2378,6 +2396,14 @@ Campaign resumes correctly.
### Pass
Prompt size remains bounded as total transcript grows.
### Result — PASS (M11, 2026-09-13)
In the M01 run the prompt reached about 15.3k tokens at turn 30. It then held
between 14.6k and 15.7k for seventy turns while the active story grew to 136
actions. The window was verified at 16,384 on 101 of 101 turns, the campaign
canon was present on every turn, and the reply reserve was subtracted on every
turn. By the narrator's own tokenizer, the largest prompt plus the reserve is
16,342 of 16,384: bounded, with little headroom. M11 report §G.3 and §N.
---
## M04 — Long-Run Memory Recall
@@ -2391,6 +2417,20 @@ Verify relevant recall near Turn 100.
### Pass
Fact/event remains recoverable without entire transcript in prompt.
### Result — PASS (M11, 2026-09-13), recovered through authoritative state
The clue was planted at depth 1. At the recall check near turn 100 it was outside
the history window, which began at depth 54. It was still recoverable: present in
the authoritative state and in the memories section. It reached memory only
because the narrator restated the state's fact line in its prose and memory
summarised that. No memory of the planting era carried the clue in any run. The
repository owner accepted recovery through authoritative state as satisfying this
test on 2026-09-13.
An earlier harness judged the precondition by whether the clue's text was absent
from recent history. The narrator reuses that text in prose, so the precondition
is now the planting turn's position, which is this test's own wording. M11 report
§G.4.
---
# N. Candidate-Specific Phase 0B Tests
@@ -2660,5 +2700,8 @@ the run reports it), which is the control this kind of tool most often lacks.
**This remains a test-design task and is still not an acceptance test.** The
model-quality half is not a pass/fail property of the application, and M11 does
not make it one. The M11 report records what the diagnostic found on the
reference narrator, including a fixture defect it caught in itself.
not make it one. **The M11 report does not record the diagnostic's run
results.** This was corrected on 2026-09-14: the section the report pointed to
was left empty. The report records the fixture defect the diagnostic caught in
itself, and the run's evidence is in the implementer's
`m11-evidence/identity-recheck` directory.