Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed

The planning package still described M01 as outstanding. It now records the
evidence run on 96c1bf5 and the two product defects found on the way. It also
corrects three statements that were never true.

- V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as
  recovered through authoritative state, with the owner's acceptance of that on
  2026-09-13 and the positional precondition explained.
  Correction: v3.7 said this file carried M11 results against every REQUIRED
  test. None were written, and the per-test matrix is the M11 report's §F. The
  §P3 M11 disposition said the report records the identity diagnostic's
  findings. It does not, and the disposition now says so.
- BUILD-MILESTONES.md: the M11 status block records the long-run evidence,
  the write-lock and protocol-leak defects, and what is left for the reviewer.
- DATA-MODEL.md §28B: M11 added two columns, not one.
  settings.context_window_override (migration 94, ef25b0a) was never
  recorded.
- TECHNICAL-DESIGN.md: "Background failure observability" gains the rule
  that nothing in a turn writes before the model call, and new §15.4 records
  that stored narration carries story only, with the extractor's rules.
- CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same
  two fixes.
- README.md and VERSION.md: status, milestone map, stop rule, and the v3.9
  entry.
- M11 report §Q: the "not revised" note is replaced by what v3.9 revised.

No requirement changes. No code changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-14 03:22:14 -04:00
co-authored by Claude Opus 5
parent d1988065e5
commit 3652dc6fae
9 changed files with 206 additions and 18 deletions
+29 -1
View File
@@ -1449,7 +1449,7 @@ Release gate in `V1-ACCEPTANCE-TESTS.md`:
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07, awaiting independent review/acceptance
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07; long-run evidence complete 2026-09-13; awaiting independent review/acceptance
Implemented on `m11-release-validation` from the signed M10 commit `1013c94`.
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package, written
@@ -1496,6 +1496,34 @@ Firefox refuses a WebDriver file path under `/tmp`, not all paths. Staging under
`$HOME` makes browser file import work, so knowledge import is now proved
end-to-end in a real browser rather than in two labelled halves.
**The long-run evidence (2026-09-13).** The first report left M01 PARTIAL at 41
turns, and that run was then lost to a host crash with its evidence. M01-M04 now
pass on a complete 100-turn run on `96c1bf5`: 101 accepted turns, three genuine
restarts, all thirteen scheduled history operations, zero failed post-turn
passes, and recovery onto a clean data directory 16 of 16. M04's precondition is
positional: the planting turn was outside the history window. The fact was
recovered through authoritative state and reached memory only as the narrator's
restatement of it. The repository owner accepted that on 2026-09-13. The
report's §G tells the whole path, including six runs that are not the evidence.
**Two further product defects**, found only once a long run had the memory bank
switched on:
3. **A turn locked its own memory bank out.** Retrieval wrote a use counter before
the model call, and the turn committed after the reply, so SQLite's one write
lock was held through the reply. Post-turn memory and summary writes timed out,
and so did recording their failure. The counter is now written in the turn's
own commit, and a failure is recorded after a rollback. (`f8d4010`)
4. **The narrator's protocol was stored as story.** The model wrote a pasted copy
of the state section, unfenced and unfinished proposals, and sections of its
own, on up to 42 of 104 turns, and stored text is replayed as history. The
extractor now removes every shape observed. Replaying 443 real turns through it
changed no turn it had previously left clean. (`0c7316f`, `96c1bf5`)
**Left for the reviewer**, in the report's §P: the real-token headroom at the
largest prompts is 23-42 tokens; the narrator restates prompt text in its prose;
and the identity diagnostic's results are not in the report.
---
## 4. Milestone Dependency Summary
+11
View File
@@ -1127,6 +1127,17 @@ If embedding generation fails:
Derived-memory failures must not block story persistence.
**As implemented (M11).** Two more rules, both learned from a long run in which
every turn was accepted while the memory bank quietly stopped filling:
- **A failure must be recordable.** The record is itself a write. It is made after
the failed session has been rolled back, so the status the reader sees
(`derived_status`) never reads `idle` over a failure.
- **The turn must not cause it.** A turn that held SQLite's write lock through the
model call locked every post-turn write out for the length of the reply. Nothing
in a turn writes before the model call (`TECHNICAL-DESIGN.md`, "Background
failure observability").
## 52. Offline Operation
All context and memory operations must work locally.
+13 -3
View File
@@ -886,15 +886,14 @@ never becomes canon (`MEDIA-EXTENSION-CONTRACT.md` §35, §37).
## 28B. What M11 Added (as implemented)
One column, and one field on a response. Both exist because something that was
Two columns, and one field on a response. Each exists because something that was
happening silently had to become visible.
```text
adventures.narration_length "" | "brief" | "medium" | "long"
```
The campaign's own narration-length choice, and the *only* schema change M11
makes. Until M11 the choice became one English sentence inside
The campaign's own narration-length choice (migration 93). Until M11 the choice became one English sentence inside
`ai_instructions` and moved no number: the numeric hint the model actually reads
was derived from the global reply cap and said the same thing for all three
settings. Stored as its own field because the prompt builder has to derive a
@@ -904,6 +903,17 @@ it is a campaign that never chose, which is exactly what every campaign created
before M11 did, so the migration needs no backfill and no existing prompt
changes under it.
```text
settings.context_window_override INTEGER | NULL (migration 94, post-M11)
```
The window an operator says the inference server enforces, for a server the
window probe cannot ask, meaning anything that does not speak Ollama's native
API. It is nullable and null by default, because a default number would be the
application guessing at a window, which is the one thing `contextwindow` refuses
to do. It never overrides a window the server reported. It was added in `ef25b0a`
and not recorded here until 2026-09-14; `TECHNICAL-DESIGN.md` §15.2 has the rule.
**No new table.** The context-window ceiling M11 enforces is *not* stored: it is
asked of the server, cached in the process, and recorded in the turn's context
snapshot as provenance. A stored ceiling would be a second copy of a fact the
@@ -62,6 +62,13 @@ Each stage has one job, and the order is deliberate: the allowlist is checked
before any field is read, so an unknown event type is rejected before its
contents are touched.
**Implementation note (post-M11).** "The protocol block is separated from the
prose" covers more than one fenced block. A small local model also pastes the
rendered state into its prose, writes its proposal unfenced or quoted, and runs
out of tokens partway through it. All of that is removed before the prose is
stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4
has the rules.
### Where it lives
- `adventures.narrative_state` — the current authoritative document. This is
+10 -6
View File
@@ -5,7 +5,7 @@
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
milestones **M1 through M8 implemented and accepted**; **M9 and M10 implemented,
committed and signed**; **M11 implemented and verified, awaiting independent
review and v1 acceptance**. M11 is the last planned milestone. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
review and v1 acceptance**, with its long-run evidence complete (2026-09-13). M11 is the last planned milestone. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
independent review found a real defect and a corrective pass fixed it.
@@ -38,9 +38,12 @@ verified, and awaits independent review** (2026-09-07).
`reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. It closed the
context-window release blocker M8 found, fixed two defects the validation itself
surfaced, and disposed of all four post-M8 playtest findings. **It is not
accepted**, and there is no release tag. The work is committed and signed on
`m11-release-validation` (`fedb714`); `main` remains M10, the last *accepted*
milestone.
accepted**, and there is no release tag. **M01-M04 now pass on a complete
100-turn run** (2026-09-13, commit `96c1bf5`). That run found two more product
defects, both fixed: a turn locking out its own memory bank, and the narrator's
protocol stored as story. The report was revised to match (`d198806`). All of it
is committed and signed on `m11-release-validation`; `main` remains M10, the last
*accepted* milestone.
**After M11 there is no further planned milestone.** What follows is
independent review and the v1 acceptance decision, which is the repository
@@ -319,7 +322,8 @@ Milestone M10 COMMITTED AND SIGNED (1013c94)
v
Milestone M11 IMPLEMENTED AND VERIFIED (2026-09-07)
v1 security, long-run, release reports/M11-IMPLEMENTATION-REPORT.md
validation awaiting independent review; no release tag
validation M01-M04 evidence complete (2026-09-13);
awaiting independent review; no release tag
|
v
v1 acceptance the repository owner's decision
@@ -330,7 +334,7 @@ v1 acceptance the repository owner's decision
**One milestone at a time. Do not begin a milestone before its brief exists.**
**Every planned milestone is now implemented.** M11 is verified and committed,
and the next action is not another milestone: it is an independent review of
its long-run evidence is complete, and the next action is not another milestone: it is an independent review of
`reports/M11-IMPLEMENTATION-REPORT.md` against the acceptance contract, and then
the owner's v1 acceptance decision. **Do not begin post-v1 work before that
decision**, and do not treat M11's own report as the acceptance record.
+41
View File
@@ -462,6 +462,15 @@ the error and the attempt count. That row is served by
the whole memory bank dead and the suite green; this is the mechanism that makes
the same failure visible.
M11's long run found the other half: a failure is only visible if it can be
written down. Retrieval issued a use-counter UPDATE, and the turn committed it only
after the reply. That held SQLite's write lock through the model call, so every
post-turn write during the reply timed out, including the `derived_status` row
that would have said so. **Nothing in a turn writes before the model call.**
Memory use counters are written in the turn's single commit
(`memorybank.record_use`), and a failure is recorded only after its session has
been rolled back.
**Prompt inspection.** The context report carries per-section token counts, the
budget, the output reserve, the protected total, the history allowance, the
summary's provenance, each retrieved memory's authority and source coordinate,
@@ -1259,6 +1268,38 @@ saving growing with the block and therefore with the budget.
records that no performance requirement exists and declines to invent one; that
still holds. What changed is the cost of a turn, not what a turn must contain.
### 15.4 Stored narration carries story only (post-M11)
A turn's reply is split into prose and proposal (`narrative.extract.split`), and
the prose is stored as the action's text. Stored text is replayed verbatim as
history (`builder._history_text`). Anything protocol-shaped left in it therefore
reaches the next prompt as a second, older account of the state, which is the
failure M5 review Finding 4 removed from history replay. It also gives the model
an example to copy.
M11's long run, with a small local narrator, found the model writing protocol into
its prose on 42 of 104 turns. Beyond the fenced block, the extractor therefore
removes:
- **a copy of the narrative-state section**, recognised by the renderer's own
headings with any markdown around them. The headings are
`render.SECTION_HEADINGS`, named once so the renderer and the extractor cannot
drift apart. A copy means two headings, or one heading with an indented entry.
- **an unfenced proposal that starts a line**, quoted or not. It becomes the turn's
proposal when there is no fence.
- **an unfinished proposal at the end, and what it leaves behind there**: a `State`
heading, a bare `>`, a parroted reminder or continue hint.
It also treats `state` as a fence label only when the label ends its line or runs
straight into the payload. That way a reminder mentioning "a ```state block" is not
read as a fence.
The rule the whole module keeps: **removing story is worse than leaving
protocol.** A candidate must parse as a proposal or carry the renderer's
headings. A lone heading followed by prose stays, and so do JSON a character
typed and a fact restated inside a sentence. That last case is how a narrator can
still carry authoritative state into its prose (M11 report §G.4 and §P).
## 16. Database Direction
SQLite remains the selected v1 authoritative store.
+45 -2
View File
@@ -2358,6 +2358,17 @@ Run or automate at least 100 accepted turns with:
### Pass
No major continuity/state/history corruption.
### Result — PASS (M11, long-run evidence 2026-09-13)
101 accepted turns on commit `96c1bf5`, against a real narrator
(`qwen2.5:3b-instruct-16k`, its 16,384-token window verified on every turn). The
run covered every step above: the fixture's characters and locations; two Save
Points and a restore; Undo with divergence; two scheduled retries and a take
selection; three imported knowledge sources; and the memory bank and
auto-summarise on, with 33 memories and 12 summaries written and zero failed
post-turn passes. It also included an induced failed model call. No turn was
refused, nothing in state or history was corrupted, and the campaign moved to a
clean data directory 16 of 16. `tools/m11_long_run.py`; M11 report §G.
---
## M02 — Restart During Long Campaign
@@ -2369,6 +2380,13 @@ Restart application at several points during M01.
### Pass
Campaign resumes correctly.
### Result — PASS (M11, 2026-09-13)
Three genuine `uvicorn` process restarts inside the M01 run, at turns 13, 41 and
83. The restart at turn 41 carried undone history across the boundary. Each one
compared transcript, head, Undo/Redo availability, scene, entities, facts, Save
Points, imported knowledge and settings before and after, and found them
identical every time. M11 report §G.2.
---
## M03 — Long-Run Context Stability
@@ -2378,6 +2396,14 @@ Campaign resumes correctly.
### Pass
Prompt size remains bounded as total transcript grows.
### Result — PASS (M11, 2026-09-13)
In the M01 run the prompt reached about 15.3k tokens at turn 30. It then held
between 14.6k and 15.7k for seventy turns while the active story grew to 136
actions. The window was verified at 16,384 on 101 of 101 turns, the campaign
canon was present on every turn, and the reply reserve was subtracted on every
turn. By the narrator's own tokenizer, the largest prompt plus the reserve is
16,342 of 16,384: bounded, with little headroom. M11 report §G.3 and §N.
---
## M04 — Long-Run Memory Recall
@@ -2391,6 +2417,20 @@ Verify relevant recall near Turn 100.
### Pass
Fact/event remains recoverable without entire transcript in prompt.
### Result — PASS (M11, 2026-09-13), recovered through authoritative state
The clue was planted at depth 1. At the recall check near turn 100 it was outside
the history window, which began at depth 54. It was still recoverable: present in
the authoritative state and in the memories section. It reached memory only
because the narrator restated the state's fact line in its prose and memory
summarised that. No memory of the planting era carried the clue in any run. The
repository owner accepted recovery through authoritative state as satisfying this
test on 2026-09-13.
An earlier harness judged the precondition by whether the clue's text was absent
from recent history. The narrator reuses that text in prose, so the precondition
is now the planting turn's position, which is this test's own wording. M11 report
§G.4.
---
# N. Candidate-Specific Phase 0B Tests
@@ -2660,5 +2700,8 @@ the run reports it), which is the control this kind of tool most often lacks.
**This remains a test-design task and is still not an acceptance test.** The
model-quality half is not a pass/fail property of the application, and M11 does
not make it one. The M11 report records what the diagnostic found on the
reference narrator, including a fixture defect it caught in itself.
not make it one. **The M11 report does not record the diagnostic's run
results.** This was corrected on 2026-09-14: the section the report pointed to
was left empty. The report records the fixture defect the diagnostic caught in
itself, and the run's evidence is in the implementer's
`m11-evidence/identity-recheck` directory.
+44 -3
View File
@@ -1,8 +1,49 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v3.8
- **Revision date:** 2026-09-10
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. M01, the 100-turn campaign, is the one REQUIRED test still outstanding.
- **Package:** Adventure Storyteller Planning Package v3.9
- **Revision date:** 2026-09-14
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. Its long-run evidence is complete (2026-09-13): all 85 REQUIRED tests pass, with H09 not applicable.
## v3.9 — M11's long-run evidence, and the two defects it took to get it (2026-09-14)
No milestone, no requirement change, and no acceptance claim. M01 was the one
REQUIRED test v3.8 left outstanding. It now has a complete 100-turn run on commit
`96c1bf5`, and M01-M04 pass on it. Getting there found two product defects and
five harness defects. The M11 report's revision (`d198806`) is the evidence.
| Document | Change | Kind |
| --- | --- | --- |
| `V1-ACCEPTANCE-TESTS.md` | **Result blocks for M01-M04**, citing the evidence run. **Correction:** v3.7 said this file carried results for the M11 release run against every REQUIRED test. No such blocks were written. The per-test matrix is the M11 report's §F. The M11 disposition under §P3 also said the report records the identity diagnostic's findings; it does not, and the disposition now says so. | acceptance evidence + correction |
| `BUILD-MILESTONES.md` | **M11 status block** gains the long-run evidence and the two further product defects. | milestone status |
| `TECHNICAL-DESIGN.md` | **"Background failure observability" gains the write-lock rule.** **New §15.4**: stored narration carries story only. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` | **§51 gains what M11 found**: a derived failure has to be recordable, and must not be caused by the turn itself. | as-implemented record |
| `DATA-MODEL.md` | **§28B corrected.** M11 added two columns, not one. `settings.context_window_override` (migration 94, `ef25b0a`) was never recorded here. | correction |
| `DECISIONS/013-authoritative-narrative-state-document.md` | **Implementation note** on what "the protocol block is separated from the prose" now has to cover. | as-implemented note |
| `planning/README.md` | Status, milestone map and stop rule: the long-run evidence is complete. | index |
| `reports/M11-IMPLEMENTATION-REPORT.md` | **Revised** (`d198806`): M01-M04 on the complete run, §G.0 removed. | milestone report |
| `DEVELOPMENT.md` | Logging a GPU inference host during a long run, so a hardware fault can be tied to power or ruled out. | developer docs |
**The long-run evidence.** 101 accepted turns, three genuine restarts, every
scheduled history operation, zero failed post-turn passes, and recovery onto a
clean data directory 16 of 16. M04 passes on a positional precondition: the
planting turn was outside the history window, which is the acceptance text's own
"without entire transcript in prompt". The fact was recovered through
authoritative state, and it reached memory only as the narrator's restatement of
that state. The repository owner accepted state-based recovery on 2026-09-13.
**Two defects the memory bank had been hiding.** No long run had ever had it
switched on.
- **A turn locked out its own post-turn work.** It held SQLite's single write lock
through the model call. Every memory, summary and status write during the reply
timed out, and so did the record of the failure. A run reported `complete` with
two memories and no summary.
- **The narrator's protocol was stored as story.** The model pasted the state
section and unfenced proposals into its prose on up to 42 of 104 turns, and
stored text is replayed as history.
**Requirement changes: zero.** The M04 precondition the harness measures was
corrected to the acceptance text's wording; the test itself is unchanged.
## v3.8 — Two gaps closed under M11's own rules, and the cost of a long turn (2026-09-10)
@@ -1508,9 +1508,12 @@ accepted for v1, or a decision for the owner at acceptance.
| `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Revised 2026-09-14**: M01-M04 on the complete evidence run; §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R rewritten; the §G.0 addendum removed. | milestone report |
| `DEVELOPMENT.md` | The pointer to the removed §G.0 replaced; a new section on logging a GPU inference host during a long run. | developer docs |
**Not revised with this report:** `planning/V1-ACCEPTANCE-TESTS.md`,
`planning/BUILD-MILESTONES.md`, `planning/VERSION.md` and `planning/README.md`
still carry what the first revision and `ef25b0a` wrote into them.
**Revised after this report, in planning package v3.9 (2026-09-14):**
`planning/V1-ACCEPTANCE-TESTS.md` gained result blocks for M01-M04 and corrected
its M11 disposition. `planning/BUILD-MILESTONES.md`, `planning/VERSION.md` and
`planning/README.md` record the long-run evidence. `DATA-MODEL.md` §28B now
records migration 94, and `TECHNICAL-DESIGN.md` (new §15.4), `CONTEXT-AND-MEMORY.md`
§51 and ADR 013 record the two defect fixes as implemented.
**Implementation facts added:** §15.2, §28B, §8A's implementation note, §42B, and
the M11 status block. Each records what the code does; none changes what is