Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed

The planning package still described M01 as outstanding. It now records the
evidence run on 96c1bf5 and the two product defects found on the way. It also
corrects three statements that were never true.

- V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as
  recovered through authoritative state, with the owner's acceptance of that on
  2026-09-13 and the positional precondition explained.
  Correction: v3.7 said this file carried M11 results against every REQUIRED
  test. None were written, and the per-test matrix is the M11 report's §F. The
  §P3 M11 disposition said the report records the identity diagnostic's
  findings. It does not, and the disposition now says so.
- BUILD-MILESTONES.md: the M11 status block records the long-run evidence,
  the write-lock and protocol-leak defects, and what is left for the reviewer.
- DATA-MODEL.md §28B: M11 added two columns, not one.
  settings.context_window_override (migration 94, ef25b0a) was never
  recorded.
- TECHNICAL-DESIGN.md: "Background failure observability" gains the rule
  that nothing in a turn writes before the model call, and new §15.4 records
  that stored narration carries story only, with the extractor's rules.
- CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same
  two fixes.
- README.md and VERSION.md: status, milestone map, stop rule, and the v3.9
  entry.
- M11 report §Q: the "not revised" note is replaced by what v3.9 revised.

No requirement changes. No code changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-14 03:22:14 -04:00
co-authored by Claude Opus 5
parent d1988065e5
commit 3652dc6fae
9 changed files with 206 additions and 18 deletions
+44 -3
View File
@@ -1,8 +1,49 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v3.8
- **Revision date:** 2026-09-10
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. M01, the 100-turn campaign, is the one REQUIRED test still outstanding.
- **Package:** Adventure Storyteller Planning Package v3.9
- **Revision date:** 2026-09-14
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. Its long-run evidence is complete (2026-09-13): all 85 REQUIRED tests pass, with H09 not applicable.
## v3.9 — M11's long-run evidence, and the two defects it took to get it (2026-09-14)
No milestone, no requirement change, and no acceptance claim. M01 was the one
REQUIRED test v3.8 left outstanding. It now has a complete 100-turn run on commit
`96c1bf5`, and M01-M04 pass on it. Getting there found two product defects and
five harness defects. The M11 report's revision (`d198806`) is the evidence.
| Document | Change | Kind |
| --- | --- | --- |
| `V1-ACCEPTANCE-TESTS.md` | **Result blocks for M01-M04**, citing the evidence run. **Correction:** v3.7 said this file carried results for the M11 release run against every REQUIRED test. No such blocks were written. The per-test matrix is the M11 report's §F. The M11 disposition under §P3 also said the report records the identity diagnostic's findings; it does not, and the disposition now says so. | acceptance evidence + correction |
| `BUILD-MILESTONES.md` | **M11 status block** gains the long-run evidence and the two further product defects. | milestone status |
| `TECHNICAL-DESIGN.md` | **"Background failure observability" gains the write-lock rule.** **New §15.4**: stored narration carries story only. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` | **§51 gains what M11 found**: a derived failure has to be recordable, and must not be caused by the turn itself. | as-implemented record |
| `DATA-MODEL.md` | **§28B corrected.** M11 added two columns, not one. `settings.context_window_override` (migration 94, `ef25b0a`) was never recorded here. | correction |
| `DECISIONS/013-authoritative-narrative-state-document.md` | **Implementation note** on what "the protocol block is separated from the prose" now has to cover. | as-implemented note |
| `planning/README.md` | Status, milestone map and stop rule: the long-run evidence is complete. | index |
| `reports/M11-IMPLEMENTATION-REPORT.md` | **Revised** (`d198806`): M01-M04 on the complete run, §G.0 removed. | milestone report |
| `DEVELOPMENT.md` | Logging a GPU inference host during a long run, so a hardware fault can be tied to power or ruled out. | developer docs |
**The long-run evidence.** 101 accepted turns, three genuine restarts, every
scheduled history operation, zero failed post-turn passes, and recovery onto a
clean data directory 16 of 16. M04 passes on a positional precondition: the
planting turn was outside the history window, which is the acceptance text's own
"without entire transcript in prompt". The fact was recovered through
authoritative state, and it reached memory only as the narrator's restatement of
that state. The repository owner accepted state-based recovery on 2026-09-13.
**Two defects the memory bank had been hiding.** No long run had ever had it
switched on.
- **A turn locked out its own post-turn work.** It held SQLite's single write lock
through the model call. Every memory, summary and status write during the reply
timed out, and so did the record of the failure. A run reported `complete` with
two memories and no summary.
- **The narrator's protocol was stored as story.** The model pasted the state
section and unfenced proposals into its prose on up to 42 of 104 turns, and
stored text is replayed as history.
**Requirement changes: zero.** The M04 precondition the harness measures was
corrected to the acceptance text's wording; the test itself is unchanged.
## v3.8 — Two gaps closed under M11's own rules, and the cost of a long turn (2026-09-10)