Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed
The planning package still described M01 as outstanding. It now records the evidence run on96c1bf5and the two product defects found on the way. It also corrects three statements that were never true. - V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as recovered through authoritative state, with the owner's acceptance of that on 2026-09-13 and the positional precondition explained. Correction: v3.7 said this file carried M11 results against every REQUIRED test. None were written, and the per-test matrix is the M11 report's §F. The §P3 M11 disposition said the report records the identity diagnostic's findings. It does not, and the disposition now says so. - BUILD-MILESTONES.md: the M11 status block records the long-run evidence, the write-lock and protocol-leak defects, and what is left for the reviewer. - DATA-MODEL.md §28B: M11 added two columns, not one. settings.context_window_override (migration 94,ef25b0a) was never recorded. - TECHNICAL-DESIGN.md: "Background failure observability" gains the rule that nothing in a turn writes before the model call, and new §15.4 records that stored narration carries story only, with the extractor's rules. - CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same two fixes. - README.md and VERSION.md: status, milestone map, stop rule, and the v3.9 entry. - M11 report §Q: the "not revised" note is replaced by what v3.9 revised. No requirement changes. No code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
co-authored by
Claude Opus 5
parent
d1988065e5
commit
3652dc6fae
@@ -2358,6 +2358,17 @@ Run or automate at least 100 accepted turns with:
|
||||
### Pass
|
||||
No major continuity/state/history corruption.
|
||||
|
||||
### Result — PASS (M11, long-run evidence 2026-09-13)
|
||||
101 accepted turns on commit `96c1bf5`, against a real narrator
|
||||
(`qwen2.5:3b-instruct-16k`, its 16,384-token window verified on every turn). The
|
||||
run covered every step above: the fixture's characters and locations; two Save
|
||||
Points and a restore; Undo with divergence; two scheduled retries and a take
|
||||
selection; three imported knowledge sources; and the memory bank and
|
||||
auto-summarise on, with 33 memories and 12 summaries written and zero failed
|
||||
post-turn passes. It also included an induced failed model call. No turn was
|
||||
refused, nothing in state or history was corrupted, and the campaign moved to a
|
||||
clean data directory 16 of 16. `tools/m11_long_run.py`; M11 report §G.
|
||||
|
||||
---
|
||||
|
||||
## M02 — Restart During Long Campaign
|
||||
@@ -2369,6 +2380,13 @@ Restart application at several points during M01.
|
||||
### Pass
|
||||
Campaign resumes correctly.
|
||||
|
||||
### Result — PASS (M11, 2026-09-13)
|
||||
Three genuine `uvicorn` process restarts inside the M01 run, at turns 13, 41 and
|
||||
83. The restart at turn 41 carried undone history across the boundary. Each one
|
||||
compared transcript, head, Undo/Redo availability, scene, entities, facts, Save
|
||||
Points, imported knowledge and settings before and after, and found them
|
||||
identical every time. M11 report §G.2.
|
||||
|
||||
---
|
||||
|
||||
## M03 — Long-Run Context Stability
|
||||
@@ -2378,6 +2396,14 @@ Campaign resumes correctly.
|
||||
### Pass
|
||||
Prompt size remains bounded as total transcript grows.
|
||||
|
||||
### Result — PASS (M11, 2026-09-13)
|
||||
In the M01 run the prompt reached about 15.3k tokens at turn 30. It then held
|
||||
between 14.6k and 15.7k for seventy turns while the active story grew to 136
|
||||
actions. The window was verified at 16,384 on 101 of 101 turns, the campaign
|
||||
canon was present on every turn, and the reply reserve was subtracted on every
|
||||
turn. By the narrator's own tokenizer, the largest prompt plus the reserve is
|
||||
16,342 of 16,384: bounded, with little headroom. M11 report §G.3 and §N.
|
||||
|
||||
---
|
||||
|
||||
## M04 — Long-Run Memory Recall
|
||||
@@ -2391,6 +2417,20 @@ Verify relevant recall near Turn 100.
|
||||
### Pass
|
||||
Fact/event remains recoverable without entire transcript in prompt.
|
||||
|
||||
### Result — PASS (M11, 2026-09-13), recovered through authoritative state
|
||||
The clue was planted at depth 1. At the recall check near turn 100 it was outside
|
||||
the history window, which began at depth 54. It was still recoverable: present in
|
||||
the authoritative state and in the memories section. It reached memory only
|
||||
because the narrator restated the state's fact line in its prose and memory
|
||||
summarised that. No memory of the planting era carried the clue in any run. The
|
||||
repository owner accepted recovery through authoritative state as satisfying this
|
||||
test on 2026-09-13.
|
||||
|
||||
An earlier harness judged the precondition by whether the clue's text was absent
|
||||
from recent history. The narrator reuses that text in prose, so the precondition
|
||||
is now the planting turn's position, which is this test's own wording. M11 report
|
||||
§G.4.
|
||||
|
||||
---
|
||||
|
||||
# N. Candidate-Specific Phase 0B Tests
|
||||
@@ -2660,5 +2700,8 @@ the run reports it), which is the control this kind of tool most often lacks.
|
||||
|
||||
**This remains a test-design task and is still not an acceptance test.** The
|
||||
model-quality half is not a pass/fail property of the application, and M11 does
|
||||
not make it one. The M11 report records what the diagnostic found on the
|
||||
reference narrator, including a fixture defect it caught in itself.
|
||||
not make it one. **The M11 report does not record the diagnostic's run
|
||||
results.** This was corrected on 2026-09-14: the section the report pointed to
|
||||
was left empty. The report records the fixture defect the diagnostic caught in
|
||||
itself, and the run's evidence is in the implementer's
|
||||
`m11-evidence/identity-recheck` directory.
|
||||
|
||||
Reference in New Issue
Block a user