Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed
The planning package still described M01 as outstanding. It now records the evidence run on96c1bf5and the two product defects found on the way. It also corrects three statements that were never true. - V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as recovered through authoritative state, with the owner's acceptance of that on 2026-09-13 and the positional precondition explained. Correction: v3.7 said this file carried M11 results against every REQUIRED test. None were written, and the per-test matrix is the M11 report's §F. The §P3 M11 disposition said the report records the identity diagnostic's findings. It does not, and the disposition now says so. - BUILD-MILESTONES.md: the M11 status block records the long-run evidence, the write-lock and protocol-leak defects, and what is left for the reviewer. - DATA-MODEL.md §28B: M11 added two columns, not one. settings.context_window_override (migration 94,ef25b0a) was never recorded. - TECHNICAL-DESIGN.md: "Background failure observability" gains the rule that nothing in a turn writes before the model call, and new §15.4 records that stored narration carries story only, with the extractor's rules. - CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same two fixes. - README.md and VERSION.md: status, milestone map, stop rule, and the v3.9 entry. - M11 report §Q: the "not revised" note is replaced by what v3.9 revised. No requirement changes. No code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
co-authored by
Claude Opus 5
parent
d1988065e5
commit
3652dc6fae
@@ -462,6 +462,15 @@ the error and the attempt count. That row is served by
|
||||
the whole memory bank dead and the suite green; this is the mechanism that makes
|
||||
the same failure visible.
|
||||
|
||||
M11's long run found the other half: a failure is only visible if it can be
|
||||
written down. Retrieval issued a use-counter UPDATE, and the turn committed it only
|
||||
after the reply. That held SQLite's write lock through the model call, so every
|
||||
post-turn write during the reply timed out, including the `derived_status` row
|
||||
that would have said so. **Nothing in a turn writes before the model call.**
|
||||
Memory use counters are written in the turn's single commit
|
||||
(`memorybank.record_use`), and a failure is recorded only after its session has
|
||||
been rolled back.
|
||||
|
||||
**Prompt inspection.** The context report carries per-section token counts, the
|
||||
budget, the output reserve, the protected total, the history allowance, the
|
||||
summary's provenance, each retrieved memory's authority and source coordinate,
|
||||
@@ -1259,6 +1268,38 @@ saving growing with the block and therefore with the budget.
|
||||
records that no performance requirement exists and declines to invent one; that
|
||||
still holds. What changed is the cost of a turn, not what a turn must contain.
|
||||
|
||||
### 15.4 Stored narration carries story only (post-M11)
|
||||
|
||||
A turn's reply is split into prose and proposal (`narrative.extract.split`), and
|
||||
the prose is stored as the action's text. Stored text is replayed verbatim as
|
||||
history (`builder._history_text`). Anything protocol-shaped left in it therefore
|
||||
reaches the next prompt as a second, older account of the state, which is the
|
||||
failure M5 review Finding 4 removed from history replay. It also gives the model
|
||||
an example to copy.
|
||||
|
||||
M11's long run, with a small local narrator, found the model writing protocol into
|
||||
its prose on 42 of 104 turns. Beyond the fenced block, the extractor therefore
|
||||
removes:
|
||||
|
||||
- **a copy of the narrative-state section**, recognised by the renderer's own
|
||||
headings with any markdown around them. The headings are
|
||||
`render.SECTION_HEADINGS`, named once so the renderer and the extractor cannot
|
||||
drift apart. A copy means two headings, or one heading with an indented entry.
|
||||
- **an unfenced proposal that starts a line**, quoted or not. It becomes the turn's
|
||||
proposal when there is no fence.
|
||||
- **an unfinished proposal at the end, and what it leaves behind there**: a `State`
|
||||
heading, a bare `>`, a parroted reminder or continue hint.
|
||||
|
||||
It also treats `state` as a fence label only when the label ends its line or runs
|
||||
straight into the payload. That way a reminder mentioning "a ```state block" is not
|
||||
read as a fence.
|
||||
|
||||
The rule the whole module keeps: **removing story is worse than leaving
|
||||
protocol.** A candidate must parse as a proposal or carry the renderer's
|
||||
headings. A lone heading followed by prose stays, and so do JSON a character
|
||||
typed and a fact restated inside a sentence. That last case is how a narrator can
|
||||
still carry authoritative state into its prose (M11 report §G.4 and §P).
|
||||
|
||||
## 16. Database Direction
|
||||
|
||||
SQLite remains the selected v1 authoritative store.
|
||||
|
||||
Reference in New Issue
Block a user