From 3652dc6fae903cad379da8efd5086709b5b9c6f5 Mon Sep 17 00:00:00 2001 From: JesseMarkowitz Date: Mon, 14 Sep 2026 03:22:14 -0400 Subject: [PATCH] Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The planning package still described M01 as outstanding. It now records the evidence run on 96c1bf5 and the two product defects found on the way. It also corrects three statements that were never true. - V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as recovered through authoritative state, with the owner's acceptance of that on 2026-09-13 and the positional precondition explained. Correction: v3.7 said this file carried M11 results against every REQUIRED test. None were written, and the per-test matrix is the M11 report's §F. The §P3 M11 disposition said the report records the identity diagnostic's findings. It does not, and the disposition now says so. - BUILD-MILESTONES.md: the M11 status block records the long-run evidence, the write-lock and protocol-leak defects, and what is left for the reviewer. - DATA-MODEL.md §28B: M11 added two columns, not one. settings.context_window_override (migration 94, ef25b0a) was never recorded. - TECHNICAL-DESIGN.md: "Background failure observability" gains the rule that nothing in a turn writes before the model call, and new §15.4 records that stored narration carries story only, with the extractor's rules. - CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same two fixes. - README.md and VERSION.md: status, milestone map, stop rule, and the v3.9 entry. - M11 report §Q: the "not revised" note is replaced by what v3.9 revised. No requirement changes. No code changes. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND --- planning/BUILD-MILESTONES.md | 30 +++++++++++- planning/CONTEXT-AND-MEMORY.md | 11 +++++ planning/DATA-MODEL.md | 16 +++++-- ...-authoritative-narrative-state-document.md | 7 +++ planning/README.md | 16 ++++--- planning/TECHNICAL-DESIGN.md | 41 ++++++++++++++++ planning/V1-ACCEPTANCE-TESTS.md | 47 ++++++++++++++++++- planning/VERSION.md | 47 +++++++++++++++++-- planning/reports/M11-IMPLEMENTATION-REPORT.md | 9 ++-- 9 files changed, 206 insertions(+), 18 deletions(-) diff --git a/planning/BUILD-MILESTONES.md b/planning/BUILD-MILESTONES.md index f1a3a9e..c4d625b 100644 --- a/planning/BUILD-MILESTONES.md +++ b/planning/BUILD-MILESTONES.md @@ -1449,7 +1449,7 @@ Release gate in `V1-ACCEPTANCE-TESTS.md`: The build meets the v1 black-box acceptance contract and can be packaged as the first production release. -## Status: IMPLEMENTED AND VERIFIED — 2026-09-07, awaiting independent review/acceptance +## Status: IMPLEMENTED AND VERIFIED — 2026-09-07; long-run evidence complete 2026-09-13; awaiting independent review/acceptance Implemented on `m11-release-validation` from the signed M10 commit `1013c94`. `planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package, written @@ -1496,6 +1496,34 @@ Firefox refuses a WebDriver file path under `/tmp`, not all paths. Staging under `$HOME` makes browser file import work, so knowledge import is now proved end-to-end in a real browser rather than in two labelled halves. +**The long-run evidence (2026-09-13).** The first report left M01 PARTIAL at 41 +turns, and that run was then lost to a host crash with its evidence. M01-M04 now +pass on a complete 100-turn run on `96c1bf5`: 101 accepted turns, three genuine +restarts, all thirteen scheduled history operations, zero failed post-turn +passes, and recovery onto a clean data directory 16 of 16. M04's precondition is +positional: the planting turn was outside the history window. The fact was +recovered through authoritative state and reached memory only as the narrator's +restatement of it. The repository owner accepted that on 2026-09-13. The +report's §G tells the whole path, including six runs that are not the evidence. + +**Two further product defects**, found only once a long run had the memory bank +switched on: + +3. **A turn locked its own memory bank out.** Retrieval wrote a use counter before + the model call, and the turn committed after the reply, so SQLite's one write + lock was held through the reply. Post-turn memory and summary writes timed out, + and so did recording their failure. The counter is now written in the turn's + own commit, and a failure is recorded after a rollback. (`f8d4010`) +4. **The narrator's protocol was stored as story.** The model wrote a pasted copy + of the state section, unfenced and unfinished proposals, and sections of its + own, on up to 42 of 104 turns, and stored text is replayed as history. The + extractor now removes every shape observed. Replaying 443 real turns through it + changed no turn it had previously left clean. (`0c7316f`, `96c1bf5`) + +**Left for the reviewer**, in the report's §P: the real-token headroom at the +largest prompts is 23-42 tokens; the narrator restates prompt text in its prose; +and the identity diagnostic's results are not in the report. + --- ## 4. Milestone Dependency Summary diff --git a/planning/CONTEXT-AND-MEMORY.md b/planning/CONTEXT-AND-MEMORY.md index 182ddc0..365b821 100644 --- a/planning/CONTEXT-AND-MEMORY.md +++ b/planning/CONTEXT-AND-MEMORY.md @@ -1127,6 +1127,17 @@ If embedding generation fails: Derived-memory failures must not block story persistence. +**As implemented (M11).** Two more rules, both learned from a long run in which +every turn was accepted while the memory bank quietly stopped filling: + +- **A failure must be recordable.** The record is itself a write. It is made after + the failed session has been rolled back, so the status the reader sees + (`derived_status`) never reads `idle` over a failure. +- **The turn must not cause it.** A turn that held SQLite's write lock through the + model call locked every post-turn write out for the length of the reply. Nothing + in a turn writes before the model call (`TECHNICAL-DESIGN.md`, "Background + failure observability"). + ## 52. Offline Operation All context and memory operations must work locally. diff --git a/planning/DATA-MODEL.md b/planning/DATA-MODEL.md index 6f0e8c6..280168e 100644 --- a/planning/DATA-MODEL.md +++ b/planning/DATA-MODEL.md @@ -886,15 +886,14 @@ never becomes canon (`MEDIA-EXTENSION-CONTRACT.md` §35, §37). ## 28B. What M11 Added (as implemented) -One column, and one field on a response. Both exist because something that was +Two columns, and one field on a response. Each exists because something that was happening silently had to become visible. ```text adventures.narration_length "" | "brief" | "medium" | "long" ``` -The campaign's own narration-length choice, and the *only* schema change M11 -makes. Until M11 the choice became one English sentence inside +The campaign's own narration-length choice (migration 93). Until M11 the choice became one English sentence inside `ai_instructions` and moved no number: the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three settings. Stored as its own field because the prompt builder has to derive a @@ -904,6 +903,17 @@ it is a campaign that never chose, which is exactly what every campaign created before M11 did, so the migration needs no backfill and no existing prompt changes under it. +```text +settings.context_window_override INTEGER | NULL (migration 94, post-M11) +``` + +The window an operator says the inference server enforces, for a server the +window probe cannot ask, meaning anything that does not speak Ollama's native +API. It is nullable and null by default, because a default number would be the +application guessing at a window, which is the one thing `contextwindow` refuses +to do. It never overrides a window the server reported. It was added in `ef25b0a` +and not recorded here until 2026-09-14; `TECHNICAL-DESIGN.md` §15.2 has the rule. + **No new table.** The context-window ceiling M11 enforces is *not* stored: it is asked of the server, cached in the process, and recorded in the turn's context snapshot as provenance. A stored ceiling would be a second copy of a fact the diff --git a/planning/DECISIONS/013-authoritative-narrative-state-document.md b/planning/DECISIONS/013-authoritative-narrative-state-document.md index 62c862e..534bf35 100644 --- a/planning/DECISIONS/013-authoritative-narrative-state-document.md +++ b/planning/DECISIONS/013-authoritative-narrative-state-document.md @@ -62,6 +62,13 @@ Each stage has one job, and the order is deliberate: the allowlist is checked before any field is read, so an unknown event type is rejected before its contents are touched. +**Implementation note (post-M11).** "The protocol block is separated from the +prose" covers more than one fenced block. A small local model also pastes the +rendered state into its prose, writes its proposal unfenced or quoted, and runs +out of tokens partway through it. All of that is removed before the prose is +stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4 +has the rules. + ### Where it lives - `adventures.narrative_state` — the current authoritative document. This is diff --git a/planning/README.md b/planning/README.md index bdbecfc..326e5ca 100644 --- a/planning/README.md +++ b/planning/README.md @@ -5,7 +5,7 @@ **Current state:** Phase 0 complete; AI-DnD forked as the production base; milestones **M1 through M8 implemented and accepted**; **M9 and M10 implemented, committed and signed**; **M11 implemented and verified, awaiting independent -review and v1 acceptance**. M11 is the last planned milestone. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03; +review and v1 acceptance**, with its long-run evidence complete (2026-09-13). M11 is the last planned milestone. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03; M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an independent review found a real defect and a corrective pass fixed it. @@ -38,9 +38,12 @@ verified, and awaits independent review** (2026-09-07). `reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. It closed the context-window release blocker M8 found, fixed two defects the validation itself surfaced, and disposed of all four post-M8 playtest findings. **It is not -accepted**, and there is no release tag. The work is committed and signed on -`m11-release-validation` (`fedb714`); `main` remains M10, the last *accepted* -milestone. +accepted**, and there is no release tag. **M01-M04 now pass on a complete +100-turn run** (2026-09-13, commit `96c1bf5`). That run found two more product +defects, both fixed: a turn locking out its own memory bank, and the narrator's +protocol stored as story. The report was revised to match (`d198806`). All of it +is committed and signed on `m11-release-validation`; `main` remains M10, the last +*accepted* milestone. **After M11 there is no further planned milestone.** What follows is independent review and the v1 acceptance decision, which is the repository @@ -319,7 +322,8 @@ Milestone M10 COMMITTED AND SIGNED (1013c94) v Milestone M11 IMPLEMENTED AND VERIFIED (2026-09-07) v1 security, long-run, release reports/M11-IMPLEMENTATION-REPORT.md - validation awaiting independent review; no release tag + validation M01-M04 evidence complete (2026-09-13); + awaiting independent review; no release tag | v v1 acceptance the repository owner's decision @@ -330,7 +334,7 @@ v1 acceptance the repository owner's decision **One milestone at a time. Do not begin a milestone before its brief exists.** **Every planned milestone is now implemented.** M11 is verified and committed, -and the next action is not another milestone: it is an independent review of +its long-run evidence is complete, and the next action is not another milestone: it is an independent review of `reports/M11-IMPLEMENTATION-REPORT.md` against the acceptance contract, and then the owner's v1 acceptance decision. **Do not begin post-v1 work before that decision**, and do not treat M11's own report as the acceptance record. diff --git a/planning/TECHNICAL-DESIGN.md b/planning/TECHNICAL-DESIGN.md index 0292f6d..1a9d717 100644 --- a/planning/TECHNICAL-DESIGN.md +++ b/planning/TECHNICAL-DESIGN.md @@ -462,6 +462,15 @@ the error and the attempt count. That row is served by the whole memory bank dead and the suite green; this is the mechanism that makes the same failure visible. +M11's long run found the other half: a failure is only visible if it can be +written down. Retrieval issued a use-counter UPDATE, and the turn committed it only +after the reply. That held SQLite's write lock through the model call, so every +post-turn write during the reply timed out, including the `derived_status` row +that would have said so. **Nothing in a turn writes before the model call.** +Memory use counters are written in the turn's single commit +(`memorybank.record_use`), and a failure is recorded only after its session has +been rolled back. + **Prompt inspection.** The context report carries per-section token counts, the budget, the output reserve, the protected total, the history allowance, the summary's provenance, each retrieved memory's authority and source coordinate, @@ -1259,6 +1268,38 @@ saving growing with the block and therefore with the budget. records that no performance requirement exists and declines to invent one; that still holds. What changed is the cost of a turn, not what a turn must contain. +### 15.4 Stored narration carries story only (post-M11) + +A turn's reply is split into prose and proposal (`narrative.extract.split`), and +the prose is stored as the action's text. Stored text is replayed verbatim as +history (`builder._history_text`). Anything protocol-shaped left in it therefore +reaches the next prompt as a second, older account of the state, which is the +failure M5 review Finding 4 removed from history replay. It also gives the model +an example to copy. + +M11's long run, with a small local narrator, found the model writing protocol into +its prose on 42 of 104 turns. Beyond the fenced block, the extractor therefore +removes: + +- **a copy of the narrative-state section**, recognised by the renderer's own + headings with any markdown around them. The headings are + `render.SECTION_HEADINGS`, named once so the renderer and the extractor cannot + drift apart. A copy means two headings, or one heading with an indented entry. +- **an unfenced proposal that starts a line**, quoted or not. It becomes the turn's + proposal when there is no fence. +- **an unfinished proposal at the end, and what it leaves behind there**: a `State` + heading, a bare `>`, a parroted reminder or continue hint. + +It also treats `state` as a fence label only when the label ends its line or runs +straight into the payload. That way a reminder mentioning "a ```state block" is not +read as a fence. + +The rule the whole module keeps: **removing story is worse than leaving +protocol.** A candidate must parse as a proposal or carry the renderer's +headings. A lone heading followed by prose stays, and so do JSON a character +typed and a fact restated inside a sentence. That last case is how a narrator can +still carry authoritative state into its prose (M11 report §G.4 and §P). + ## 16. Database Direction SQLite remains the selected v1 authoritative store. diff --git a/planning/V1-ACCEPTANCE-TESTS.md b/planning/V1-ACCEPTANCE-TESTS.md index 0c828e3..2da8bfe 100644 --- a/planning/V1-ACCEPTANCE-TESTS.md +++ b/planning/V1-ACCEPTANCE-TESTS.md @@ -2358,6 +2358,17 @@ Run or automate at least 100 accepted turns with: ### Pass No major continuity/state/history corruption. +### Result — PASS (M11, long-run evidence 2026-09-13) +101 accepted turns on commit `96c1bf5`, against a real narrator +(`qwen2.5:3b-instruct-16k`, its 16,384-token window verified on every turn). The +run covered every step above: the fixture's characters and locations; two Save +Points and a restore; Undo with divergence; two scheduled retries and a take +selection; three imported knowledge sources; and the memory bank and +auto-summarise on, with 33 memories and 12 summaries written and zero failed +post-turn passes. It also included an induced failed model call. No turn was +refused, nothing in state or history was corrupted, and the campaign moved to a +clean data directory 16 of 16. `tools/m11_long_run.py`; M11 report §G. + --- ## M02 — Restart During Long Campaign @@ -2369,6 +2380,13 @@ Restart application at several points during M01. ### Pass Campaign resumes correctly. +### Result — PASS (M11, 2026-09-13) +Three genuine `uvicorn` process restarts inside the M01 run, at turns 13, 41 and +83. The restart at turn 41 carried undone history across the boundary. Each one +compared transcript, head, Undo/Redo availability, scene, entities, facts, Save +Points, imported knowledge and settings before and after, and found them +identical every time. M11 report §G.2. + --- ## M03 — Long-Run Context Stability @@ -2378,6 +2396,14 @@ Campaign resumes correctly. ### Pass Prompt size remains bounded as total transcript grows. +### Result — PASS (M11, 2026-09-13) +In the M01 run the prompt reached about 15.3k tokens at turn 30. It then held +between 14.6k and 15.7k for seventy turns while the active story grew to 136 +actions. The window was verified at 16,384 on 101 of 101 turns, the campaign +canon was present on every turn, and the reply reserve was subtracted on every +turn. By the narrator's own tokenizer, the largest prompt plus the reserve is +16,342 of 16,384: bounded, with little headroom. M11 report §G.3 and §N. + --- ## M04 — Long-Run Memory Recall @@ -2391,6 +2417,20 @@ Verify relevant recall near Turn 100. ### Pass Fact/event remains recoverable without entire transcript in prompt. +### Result — PASS (M11, 2026-09-13), recovered through authoritative state +The clue was planted at depth 1. At the recall check near turn 100 it was outside +the history window, which began at depth 54. It was still recoverable: present in +the authoritative state and in the memories section. It reached memory only +because the narrator restated the state's fact line in its prose and memory +summarised that. No memory of the planting era carried the clue in any run. The +repository owner accepted recovery through authoritative state as satisfying this +test on 2026-09-13. + +An earlier harness judged the precondition by whether the clue's text was absent +from recent history. The narrator reuses that text in prose, so the precondition +is now the planting turn's position, which is this test's own wording. M11 report +§G.4. + --- # N. Candidate-Specific Phase 0B Tests @@ -2660,5 +2700,8 @@ the run reports it), which is the control this kind of tool most often lacks. **This remains a test-design task and is still not an acceptance test.** The model-quality half is not a pass/fail property of the application, and M11 does -not make it one. The M11 report records what the diagnostic found on the -reference narrator, including a fixture defect it caught in itself. +not make it one. **The M11 report does not record the diagnostic's run +results.** This was corrected on 2026-09-14: the section the report pointed to +was left empty. The report records the fixture defect the diagnostic caught in +itself, and the run's evidence is in the implementer's +`m11-evidence/identity-recheck` directory. diff --git a/planning/VERSION.md b/planning/VERSION.md index 218d6be..a818aab 100644 --- a/planning/VERSION.md +++ b/planning/VERSION.md @@ -1,8 +1,49 @@ # Planning Package Version -- **Package:** Adventure Storyteller Planning Package v3.8 -- **Revision date:** 2026-09-10 -- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. M01, the 100-turn campaign, is the one REQUIRED test still outstanding. +- **Package:** Adventure Storyteller Planning Package v3.9 +- **Revision date:** 2026-09-14 +- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. Its long-run evidence is complete (2026-09-13): all 85 REQUIRED tests pass, with H09 not applicable. + +## v3.9 — M11's long-run evidence, and the two defects it took to get it (2026-09-14) + +No milestone, no requirement change, and no acceptance claim. M01 was the one +REQUIRED test v3.8 left outstanding. It now has a complete 100-turn run on commit +`96c1bf5`, and M01-M04 pass on it. Getting there found two product defects and +five harness defects. The M11 report's revision (`d198806`) is the evidence. + +| Document | Change | Kind | +| --- | --- | --- | +| `V1-ACCEPTANCE-TESTS.md` | **Result blocks for M01-M04**, citing the evidence run. **Correction:** v3.7 said this file carried results for the M11 release run against every REQUIRED test. No such blocks were written. The per-test matrix is the M11 report's §F. The M11 disposition under §P3 also said the report records the identity diagnostic's findings; it does not, and the disposition now says so. | acceptance evidence + correction | +| `BUILD-MILESTONES.md` | **M11 status block** gains the long-run evidence and the two further product defects. | milestone status | +| `TECHNICAL-DESIGN.md` | **"Background failure observability" gains the write-lock rule.** **New §15.4**: stored narration carries story only. | as-implemented record | +| `CONTEXT-AND-MEMORY.md` | **§51 gains what M11 found**: a derived failure has to be recordable, and must not be caused by the turn itself. | as-implemented record | +| `DATA-MODEL.md` | **§28B corrected.** M11 added two columns, not one. `settings.context_window_override` (migration 94, `ef25b0a`) was never recorded here. | correction | +| `DECISIONS/013-authoritative-narrative-state-document.md` | **Implementation note** on what "the protocol block is separated from the prose" now has to cover. | as-implemented note | +| `planning/README.md` | Status, milestone map and stop rule: the long-run evidence is complete. | index | +| `reports/M11-IMPLEMENTATION-REPORT.md` | **Revised** (`d198806`): M01-M04 on the complete run, §G.0 removed. | milestone report | +| `DEVELOPMENT.md` | Logging a GPU inference host during a long run, so a hardware fault can be tied to power or ruled out. | developer docs | + +**The long-run evidence.** 101 accepted turns, three genuine restarts, every +scheduled history operation, zero failed post-turn passes, and recovery onto a +clean data directory 16 of 16. M04 passes on a positional precondition: the +planting turn was outside the history window, which is the acceptance text's own +"without entire transcript in prompt". The fact was recovered through +authoritative state, and it reached memory only as the narrator's restatement of +that state. The repository owner accepted state-based recovery on 2026-09-13. + +**Two defects the memory bank had been hiding.** No long run had ever had it +switched on. + +- **A turn locked out its own post-turn work.** It held SQLite's single write lock + through the model call. Every memory, summary and status write during the reply + timed out, and so did the record of the failure. A run reported `complete` with + two memories and no summary. +- **The narrator's protocol was stored as story.** The model pasted the state + section and unfenced proposals into its prose on up to 42 of 104 turns, and + stored text is replayed as history. + +**Requirement changes: zero.** The M04 precondition the harness measures was +corrected to the acceptance text's wording; the test itself is unchanged. ## v3.8 — Two gaps closed under M11's own rules, and the cost of a long turn (2026-09-10) diff --git a/planning/reports/M11-IMPLEMENTATION-REPORT.md b/planning/reports/M11-IMPLEMENTATION-REPORT.md index afe2d4e..2aaf3f2 100644 --- a/planning/reports/M11-IMPLEMENTATION-REPORT.md +++ b/planning/reports/M11-IMPLEMENTATION-REPORT.md @@ -1508,9 +1508,12 @@ accepted for v1, or a decision for the owner at acceptance. | `planning/reports/M11-IMPLEMENTATION-REPORT.md` | **Revised 2026-09-14**: M01-M04 on the complete evidence run; §A, §B, §C, §E.1, §F, §G, §J, §K, §N, §O, §P and §R rewritten; the §G.0 addendum removed. | milestone report | | `DEVELOPMENT.md` | The pointer to the removed §G.0 replaced; a new section on logging a GPU inference host during a long run. | developer docs | -**Not revised with this report:** `planning/V1-ACCEPTANCE-TESTS.md`, -`planning/BUILD-MILESTONES.md`, `planning/VERSION.md` and `planning/README.md` -still carry what the first revision and `ef25b0a` wrote into them. +**Revised after this report, in planning package v3.9 (2026-09-14):** +`planning/V1-ACCEPTANCE-TESTS.md` gained result blocks for M01-M04 and corrected +its M11 disposition. `planning/BUILD-MILESTONES.md`, `planning/VERSION.md` and +`planning/README.md` record the long-run evidence. `DATA-MODEL.md` §28B now +records migration 94, and `TECHNICAL-DESIGN.md` (new §15.4), `CONTEXT-AND-MEMORY.md` +§51 and ADR 013 record the two defect fixes as implemented. **Implementation facts added:** §15.2, §28B, §8A's implementation note, §42B, and the M11 status block. Each records what the code does; none changes what is