Files
interactive-story/planning/BUILD-MILESTONES.md
T
JesseMarkowitzandClaude Opus 5 3652dc6fae Planning v3.9: record M11's long-run evidence, and correct what v3.7 claimed
The planning package still described M01 as outstanding. It now records the
evidence run on 96c1bf5 and the two product defects found on the way. It also
corrects three statements that were never true.

- V1-ACCEPTANCE-TESTS.md: result blocks for M01-M04. M04 is recorded as
  recovered through authoritative state, with the owner's acceptance of that on
  2026-09-13 and the positional precondition explained.
  Correction: v3.7 said this file carried M11 results against every REQUIRED
  test. None were written, and the per-test matrix is the M11 report's §F. The
  §P3 M11 disposition said the report records the identity diagnostic's
  findings. It does not, and the disposition now says so.
- BUILD-MILESTONES.md: the M11 status block records the long-run evidence,
  the write-lock and protocol-leak defects, and what is left for the reviewer.
- DATA-MODEL.md §28B: M11 added two columns, not one.
  settings.context_window_override (migration 94, ef25b0a) was never
  recorded.
- TECHNICAL-DESIGN.md: "Background failure observability" gains the rule
  that nothing in a turn writes before the model call, and new §15.4 records
  that stored narration carries story only, with the extractor's rules.
- CONTEXT-AND-MEMORY.md §51 and ADR 013: as-implemented notes for the same
  two fixes.
- README.md and VERSION.md: status, milestone map, stop rule, and the v3.9
  entry.
- M11 report §Q: the "not revised" note is replaced by what v3.9 revised.

No requirement changes. No code changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
2026-09-14 03:22:14 -04:00

79 KiB

Adventure Storyteller — Production Build Milestones

Status: In implementation. M1-M8 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6, M7 and M8: 2026-09-06, each of the last four after an independent review and a corrective pass). M9 — Export, Backup, Recovery, and Migration Hardening — is next, and has not been started.
Base: AI-DnD d72f7c1bda0f34fccd84afb7a25c34eb01c901de

1. Purpose

This document defines the production implementation sequence after Phase 0.

It is intentionally a milestone plan, not a Codex execution prompt. Each milestone should be converted into a separate coding brief only after the previous milestone has been reviewed and accepted.

The milestone order prioritizes load-bearing correctness and offline safety before broad feature work.

2. Global Rules for Every Milestone

Each implementation milestone must:

  • preserve upstream provenance and licensing notices,
  • keep the application runnable whenever practical,
  • add or update automated tests for changed behavior,
  • avoid unrelated refactors,
  • update affected documentation,
  • record schema/migration changes,
  • preserve local-only defaults,
  • stop at the milestone boundary for review,
  • avoid implementing later milestones opportunistically.

The product specification and acceptance tests outrank convenience inherited from AI-DnD.

3. End-User Capability Progression

The application is expected to remain runnable throughout the sequence. M8 does not create the browser UI from scratch; AI-DnD already provides a browser interface from M1. M8 is where that inherited/adapted interface is completed and simplified into the intended v1 storyteller experience.

  • M1 — First playable baseline: The user can open the inherited browser UI, start or resume a story, send prompts to Ollama, receive streamed narration, and persist the story across restart. The interface may still look and behave substantially like AI-DnD, and advanced history/state features are not yet converted to final semantics.
  • M2 — Clean local storyteller baseline: Normal story play still works, but cloud/hosted/account/scripting surfaces are removed and Ollama can be either same-host or on an explicitly configured trusted-LAN machine. From the user's perspective this is still basic play, but the environment now matches the intended privacy/deployment model.
  • M3 — Safe history controls: The user can play normally with production-grade Undo, Redo, Retry, alternate takes, and divergence without deleting old story history. Export/import also preserves the exact current position even when the user has undone several turns.
  • M4 — Save Points: The user can name a point in the story, continue playing, restart, return to that Save Point, and take a different path without losing the later story. This is the first milestone where deliberate long-form experimentation/recovery should feel comfortable.
  • M5 — Reliable story state: Characters, locations, possessions, relationships, facts, threads, and current scene become application-owned genre-neutral state rather than RPG-stat machinery. The user can begin inspecting/correcting what the storyteller believes, although the final polished state UI comes later.
  • M6 — Long-story memory: The storyteller becomes much better suited to long campaigns because summaries and retrieved older memories remain branch-safe and inspectable. The user should be able to return to old people, promises, clues, and events without abandoned paths contaminating the active story.
  • M7 — Local knowledge library: The user can import .txt/.md material as Canon, Reference, or Inspiration and have it retrieved locally with provenance. The functionality is usable, but some management/inspection surfaces may still be utilitarian until M8.
  • M8 — Finished v1 browser experience: The existing browser UI is reorganized and polished around storytelling: streamlined campaign setup, transcript/input, Undo/Redo/Retry, Save Points, state, knowledge, context inspection, and model status. This milestone makes the product feel like the intended storyteller rather than an adapted AI-DnD application.
  • M9 — Portable/recoverable campaigns: The user can reliably export, back up, import, and recover complete campaigns including history, Save Points, state, knowledge, and active position. This is where moving or restoring a campaign becomes a supported user workflow rather than merely an underlying capability.
  • M10 — Media-ready, still text-first: Little or nothing visibly changes for normal play. The story/scene data and provider boundaries are prepared so image/video/audio/TTS/STT can be added later without redesigning the core application.
  • M11 — Release-quality v1: The user experience should be functionally complete; this milestone proves it stays correct through long campaigns, repeated history operations, offline use, failures, export/import, and multiple genres. It is the release-validation milestone rather than a major new feature milestone.

M1 — Establish Production Fork and Offline Baseline

Objective

Create the production fork from the pinned AI-DnD commit and make ordinary local story operation genuinely offline-capable before larger changes.

Scope

  • establish production repository/fork lineage,
  • record upstream commit and license provenance,
  • establish reproducible dev/test environment,
  • vendor/cache/replace the tiktoken first-use encoding dependency so story generation works without Internet,
  • eliminate runtime Google Fonts/remote font dependency,
  • tighten CSP for local runtime assets,
  • verify same-host Ollama story generation with outbound Internet blocked,
  • verify explicitly configured trusted-LAN Ollama story generation while the storyteller UI/API remains loopback-bound,
  • establish baseline regression/test report.

Added during implementation: outbound TLS trust. A trusted-LAN Ollama may be served over HTTPS with a privately issued certificate, which the inherited HTTP client refused because it verified against a bundled public-CA list only. M1 made outbound HTTPS verify against the operating system's CA store as well, with verification and hostname checking fully intact and no bypass option (ADR 002; backend/app/tlstrust.py). This was not in the scope list above but was on M1's critical path: without it the trusted-LAN Definition of Done below is unreachable against a realistic host.

Explicit Non-Scope

  • no history rewrite yet,
  • no RPG-state redesign,
  • no imported knowledge,
  • no UI redesign,
  • no media generation.

Tests / Acceptance

Must demonstrate:

  • A01-A06 as applicable to the inherited base, including the separate-host trusted-LAN inference path,
  • H01-H03 and H11,
  • no unexpected DNS/HTTP requests during ordinary story generation after setup,
  • no remote font request,
  • no tokenizer/BPE download on first production turn,
  • existing relevant AI-DnD tests remain green except documented upstream/environment exceptions.

Definition of Done

A clean production build can start, open the browser UI, generate and persist story turns through either same-host Ollama or an explicitly configured trusted-LAN Ollama host, restart, and resume with outbound Internet blocked. The storyteller UI/API remains loopback-bound by default.

Status: COMPLETE

Accepted 2026-09-02. Evidence: planning/archive/milestone-reports/M1-BASELINE-REPORT.md (run logs and packet captures) and planning/archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md (review report). A01-A06, H01-H03 and H11 all pass on runtime evidence; 648 backend tests pass, including with no route to the Internet.

Capabilities M1 delivered, which later milestones inherit rather than build:

  • offline first turn — the tokenizer table is vendored and digest-checked,
  • no remote runtime assets — fonts self-hosted, CSP names no remote origin,
  • loopback-bound storyteller in every run path, including the published Docker port,
  • working trusted-LAN inference, plain HTTP and HTTPS with a private CA, demonstrated against a second physical machine,
  • a reproducible environment (backend/requirements.lock) and an offline-capable test suite.

M2 — Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface

Objective

Reduce AI-DnD to the intended single-user local product trust boundary without destabilizing the story/tree/memory foundation.

Scope

Remove or isolate as appropriate:

  • multi-user/account/guest/auth flows,
  • demo API-key behavior,
  • hosted analytics,
  • Render/Neon deployment paths,
  • Postgres/psycopg support,
  • cloud providers not required by v1,
  • arbitrary remote model-provider UI/configuration,
  • QuickJS/campaign scripting,
  • AI-Dungeon compatibility code that exists primarily to support scripting/hosted behavior,
  • hosted-only rate-limit/account infrastructure.

Add:

  • explicit Ollama endpoint policy. Trusted-LAN inference already works as of M1 — same-host loopback default, user-configured LAN endpoint, HTTP or HTTPS with a privately issued certificate, verified against the machine's CA store. M2's job is not to invent that capability but to formalize and narrow it: decide and enforce what an endpoint may be in normal v1 configuration, reject or remove what it may not, and keep the endpoint's configuration from affecting the storyteller's own loopback bind. Do not regress the TLS behavior while narrowing the surface — the shared verification context must follow any client the removal work rewrites,
  • clear local-model connection diagnostics,
  • migration/test instrumentation replacements for any tests that depended on JS hooks or removed hosted paths.

Explicit Non-Scope

  • do not redesign story history,
  • do not replace RPG state yet,
  • do not build knowledge/media systems.

Tests / Acceptance

  • local story play still works,
  • existing tree/retry/memory/context behavior remains intact,
  • application requires no cloud API keys,
  • approved trusted-LAN Ollama endpoints work — including an HTTPS endpoint with a privately issued certificate, per A06 — while arbitrary public/Internet provider endpoints are rejected or absent from normal production configuration,
  • removed provider/auth/analytics/scripting paths are no longer reachable from normal production configuration.

Definition of Done

The codebase has a narrow single-user/local-only surface and the inherited story foundation still passes its relevant regression suite.

Status: COMPLETE

Accepted 2026-09-02. Evidence: planning/archive/milestone-reports/M2-BASELINE-REPORT.md (measurements) and planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md (review); verdict accept with non-blocking debt, proceed to M3. Implementation is commit 8c65ae9, and the three defects the review found are commit 8652fe7 — see the closeout note appended to both reports.

Capabilities M2 delivered, which later milestones inherit rather than build:

  • a single-user product with no accounts, sessions or auth — all four /api/auth/* routes are gone, not gated; 52 API routes down to 36,
  • Ollama as the only inference backend, with no cloud provider code and no API key anywhere in the product,
  • an address-based inference endpoint policy, enforced on save and again before every outbound request, that refuses public addresses even when the database is edited behind the settings API (ADR 011),
  • M1's trusted-LAN and TLS behaviour carried through the removal intact — verified HTTPS against a real second machine, no bypass option,
  • SQLite as the only store; Postgres, Neon and the Render deployment path removed,
  • no campaign scripting: the QuickJS engine and /api/scripts are gone,
  • a materially simpler runtime — 10 environment variables down to two, 6 Python and 21 npm packages removed, and a 933 kB bundle down to 395 kB.

Debt carried forward, none of it blocking M3: Settings.model still defaults to "" with nothing prompting for it (M8); inert legacy tables and columns await a cleanup migration once the schema settles, after M3/M5; there are still no frontend tests (M8). docs/*.html, upstream's project site, was listed here as still linking Google Fonts; the whole inherited docs/ tree was deleted in the 2026-09-03 documentation pass, which closes that item. Full table in the implementation report §P.


M3 — Production Non-Destructive History, Redo, and Active-Head Export

Objective

Replace destructive Undo with the demonstrated head-cursor model and make Undo/Redo/divergence/export/import production-safe.

Scope

Promote the Phase 0B spike concept into maintainable production code:

  • active head can move behind retained tip,
  • Undo moves head and deletes zero accepted turns,
  • Redo follows the retained continuation,
  • new write behind tip forks on first write,
  • retry/add-take/edit paths use the same safe fork/head rules,
  • prevent Undo from walking before the campaign opening/root semantics,
  • mark displaced/inactive futures/takes as retained/disposable using implementation-appropriate metadata,
  • keep ordinary Redo invalidated after divergence,
  • add browser Redo control and correct disabled/enabled states,
  • keep branch complexity out of normal UI,
  • export active head coordinate/depth as well as active branch,
  • import honors exported active head,
  • preserve backward compatibility for older bundles that lack the new field.

Explicit Non-Scope

  • named checkpoints belong in M4,
  • no general branch-tree UI,
  • no abandoned-history cleanup feature.

Tests / Acceptance

Must cover:

  • D01-D10,
  • D02 minimum-five Undo,
  • D03 practical unlimited Undo behavior across retained history,
  • D04-D05 Redo semantics,
  • E01-E04 lineage isolation,
  • I01-I03 and I07 (undone-head round trip),
  • L01-L02,
  • branch-scoped memory remains isolated under Undo/Redo/divergence,
  • row counts demonstrate Undo deletes zero accepted turns.

Definition of Done

History operations are non-destructive, Redo works, divergence preserves old futures, and export/import reopens at the exact active head.

Status: COMPLETE

Accepted 2026-09-03. Evidence: planning/archive/milestone-reports/M3-IMPLEMENTATION-REPORT.md, which is M3's primary evidence record — no separate baseline report was produced, so that document carries the raw counts and runtime observations as well as the review. The architecture is recorded in ADR 012.

Capabilities M3 delivered, which later milestones inherit rather than build:

  • Undo that deletes zero accepted turns — measured directly on row identity: 15 rows before five Undos, 15 after,
  • Redo, round-tripping exactly (head 14 → 4 → 14) in transcript and in state,
  • a single head-movement mechanism every position change goes through, and a single capped lineage that bounds every read of the story,
  • divergence decided by the lineage rather than by a flag: the first write below a moved-back head forks, the displaced future keeps its rows, and ordinary Redo stops offering it with nothing to invalidate,
  • branch-scoped memory isolation preserved for free — a memory past the head is unretrievable and becomes eligible again on Redo, with no pruning and no re-embedding,
  • an export that carries the reader's position, so a campaign exported after two Undos imports still undone, with pre-M3 bundles opening at their tip,
  • Undo across fork points to the campaign opening, and refusal to change what a turn says while story descends from it off screen — both ratified in STORY-BRANCH-SEMANTICS.md (§5, §10, §14A).

Closeout condition — CLOSED at M4 closeout (2026-09-03). M3 was accepted with one condition outstanding: the required browser smoke test had not been performed, because no session in which M3 was implemented or reviewed had a browser available, leaving the DOM-level behaviour of the Redo button and its disabled states unverified by observation.

That condition is now discharged. A real Firefox 154.0.1, driven through geckodriver, exercised M3's controls in the rendered application: Undo enabled and Redo disabled at the tip, two Undos moving the transcript back, Redo becoming enabled and returning the original tip exactly, Retry and the take pager, and a divergent write retiring Redo with no stale old-future text on screen. It passed. Evidence: planning/archive/milestone-reports/M4-IMPLEMENTATION-REPORT.md §W.7.

Debt carried forward, none of it blocking M4: full narrator-edit state re-evaluation is deferred to M5 (STORY-BRANCH-SEMANTICS.md §14A records the interim refusal); POST /adventures/import returns every branch's rows rather than a head-capped window (inherited, harmless in the UI); ActionPage is constructed in two places. Full table in the implementation report §S.


M4 — Named Save Points / Checkpoints

Objective

Add durable user-facing Save Points on top of the active-head model.

Scope

  • named checkpoint storage,
  • checkpoint create/list/rename/delete,
  • restore by moving the active head,
  • later history retained rather than deleted,
  • fork only on first new continuation after restore,
  • checkpoint persistence across restart,
  • checkpoint export/import,
  • browser Save Point UX per BROWSER-UX-SPEC.md.

Explicit Non-Scope

  • no branch merge,
  • no automatic discarded-history cleanup,
  • no complex branch explorer required.

Tests / Acceptance

  • D11-D14,
  • E lineage tests after checkpoint restore/divergence,
  • I04,
  • L03.

Definition of Done

The user can create a named Save Point, continue, restart, restore it, and continue differently without losing later history.

Note from M3 — reuse the head machinery, do not build a second one

A Save Point is a durable named pointer to a recoverable story position, and nothing more. M3 made that position a stored coordinate and made moving to one a row lookup plus a state restore, so restoring a Save Point is head movement with a bounds check — not a restore system of its own.

Concretely, M4 should:

  • store the coordinate, the name, and the metadata around them, and no copy of any story;
  • restore by calling M3's head-movement mechanism, so that state, transcript, context and memory eligibility all move together exactly as they do for Undo and Redo, and so that later history is retained rather than deleted (D13 is already satisfied by the mechanism);
  • let the existing fork-on-first-write-below-the-head rule handle divergence after a restore, rather than forking at restore time;
  • validate that a Save Point's coordinate is still on the lineage being read before moving to it.

A second restore path is the specific failure to avoid. The Phase 0B spike put the fork check in the write path and left Retry and add-take on the old one, and M3's cost was reconciling them; a parallel checkpoint mover would recreate that divergence in a place where the two paths would silently disagree about what "restore" means. See ADR 012.

Status: COMPLETE

Accepted 2026-09-03. Evidence: planning/archive/milestone-reports/M4-IMPLEMENTATION-REPORT.md, including its §W closeout addendum. The Definition of Done is met, and — for the first time in this project — verified in a real browser.

What M4 delivered:

  • checkpoints, a table holding a name, an optional note, and a (branch, depth) coordinate — and no copy of any story. create_all builds it, as it did memories and branches; migration 80 adds the index. No backfill: nobody had named a position before M4, and inventing one would be inventing the decision.
  • Create / list / rename / delete / restore under /api/adventures/{id}/checkpoints, campaign-scoped, with a Save Point from another campaign a 404 rather than a restore of the wrong story.
  • Create at the active head, not the retained tip, so a Save Point made after two Undos names the undone position.
  • Restore that delegates, and is the whole of the milestone's architecture: resolve the coordinate, refuse it if it names no live turn, then head.move_to_node — one function whose depth half is M3's head.move_to unchanged, and whose branch half is the single assignment switch_branch makes. Restore forks nothing.
  • A browser Save Point panel — create form, list, Restore, Rename, Delete, with both confirmations saying what is not destroyed — plus a Save Point button beside Undo and Redo, where it belongs. No branch explorer, no discarded-history browser, no merge UI.
  • Export/import of Save Points with no format version bump, and a pre-M4 bundle importing with none.

The one architectural decision M4 had to make, which ADR 012 does not settle: a Save Point can name a position on a line the story has since left, so restore moves the branch half of the head as well — but only when the coordinate is not on the path being read. Doing it unconditionally would quietly hand back an abandoned continuation whenever a Save Point in a shared prefix was restored. No new ADR: this is ADR 012's mechanism applied to both halves of a coordinate ADR 012 already defines, not a new architecture. TECHNICAL-DESIGN.md §8.8 records it.

Tests: 42 in backend/tests/test_save_points.py, covering D11-D14, I04, L03, E-series lineage and memory isolation after restore and divergence, the edge cases in the brief, and the M3-database migration.

Facts M5 inherits, and must not redesign:

  • a Save Point is a durable story coordinate — (branch, depth) — carrying no copy of transcript, state, prompt, memory or summary;
  • restore is M3 head movement, through the same head.move_to Undo and Redo use; there is no second restore path and M5 must not add one;
  • state recovery stays snapshot/cached-position based, never a replay of the campaign (TECHNICAL-DESIGN.md §10.4). This is now load-bearing for Save Points as well as Undo/Redo;
  • the first divergent write after a restore creates the continuation; restore itself never forks;
  • history a Save Point names cannot disappear through an unrelated deletion (§19.1);
  • M3/M4 history and Save Point semantics are infrastructure now. M5 replaces the state model; it does not revisit how the story is positioned.

Acceptance evidence: D11-D14, I04 and L03 all pass. D11 and L03 are discharged by automation that crosses a genuine OS process boundary — one server process writes the campaign, is killed, and a second process reads it back — rather than by recreating a client in one process.

The browser condition is closed, for M4 and retrospectively for M3. A real Firefox 154.0.1, driven through geckodriver over the W3C WebDriver protocol, exercised the rendered DOM end to end: 44/44 checks passed, covering M3's Undo/Redo enable states and transcript movement, Retry and the take pager, M3 divergence and the disappearance of Redo, and every M4 Save Point operation including both confirmations and the branch-delete warning. No console errors. This closes the outstanding M3 condition recorded above and the equivalent M4 one.

The three review findings were fixed during closeout:

  • B-1 — the Save Point list was an N+1 that loaded whole Action rows, narration included. It is now one bulk two-column coordinate query plus one lineage: 53 SELECTs for 25 Save Points became 5, and the query count no longer moves with the length of the list.
  • B-2 — deleting a branch silently deleted the Save Points naming it. Fixed as a behaviour defect, not a wording one: a branch a Save Point names can no longer be deleted at all. The request is refused with the offending Save Points named, the user deletes them explicitly (which deletes no story), and the branch then goes. Both delete controls, in the branch list and in the tree overlay, disable and explain rather than warning about a loss that no longer happens. STORY-BRANCH-SEMANTICS.md §19.1 records the rule; §28 already required a future cleanup feature to retain checkpoint-referenced paths, and this is that requirement applied to the deletion path that exists today.
  • B-3 — the D11/L03 automation now spawns real server processes.

Also fixed: creating a Save Point takes the campaign's turn lock (review §S C-5), so "save where I am" cannot read a head that a turn in flight is about to move. Rename and Delete deliberately do not take it — neither reads nor moves a story position.

Debt M4 carries forward: the Save Point panel has no frontend test, because the project still has no frontend test runner at all (M8) — the browser smoke test above is a closeout procedure, not a suite. POST /adventures/import still returns every branch's rows rather than a head-capped window (inherited, M3).

Note to M5 — the instrumentation is now larger than M3 estimated

M4 added 60 tests that use inherited RPG world-state values as deterministic instrumentation, on top of the ~20 M3 flagged. M5 replaces that state model, and must move the instrumentation while preserving the behavioural assertions: what those tests measure is where the story is being read and what state belongs to that position, which is exactly as true after M5 as before it. Deleting them would delete the evidence for D11-D14, I04, L03 and the E-series.

The existing constraint stands and is now load-bearing for Save Points as well as Undo/Redo: historical state must remain efficiently snapshot/cache recoverable, so that moving to a position never becomes proportional to campaign length (TECHNICAL-DESIGN.md §10.4). Restore, Undo and Redo all pay whatever that costs.


M5 — Genre-Neutral Authoritative Narrative State

Objective

Replace/generalize AI-DnD's RPG-specific relative-delta state system with the approved genre-neutral typed-event model.

Scope

Establish generic authoritative state for:

  • entities,
  • characters/locations/organizations/items/vehicles as descriptive entity categories,
  • facts,
  • relationships,
  • possession/location/status/conditions,
  • story threads,
  • scene state,
  • manual state/canon corrections,
  • state proposal provenance.

Implement:

  • explicit typed state proposal schema,
  • absolute/unambiguous event semantics,
  • event allowlist,
  • validation,
  • atomic accepted-event commit,
  • current/historical snapshot/cache,
  • rollback/reconstruction tied to active lineage,
  • browser current-state inspector foundation.

Remove or demote:

  • D&D-specific stats/bands/cooldowns/mechanics from the core product model,
  • relative-delta semantics as the generic state protocol.

Explicit Non-Scope

  • no optional RPG module,
  • no imported knowledge yet,
  • no final rich state-editing UX if a simpler inspector is enough for this milestone.

Tests / Acceptance

  • C01-C04 and C06,
  • D/E tests verifying state follows head movement,
  • H05 invalid state event rejection,
  • L01-L02,
  • malformed JSON/state proposal handling,
  • realistic-context tests against representative local models,
  • fantasy and science-fiction state fixtures use the same schema.

Definition of Done

Accepted story state is genre-neutral, auditable, reconstructable, and no longer depends on ambiguous relative deltas.

Note from M2 — eight rollback tests use the world-state engine as instrumentation

M2 removed QuickJS. Eight existing rollback/history tests had used a JavaScript counter as deterministic instrumentation — a value they could change on a turn and then assert had been rolled back — and they now use the inherited RPG/world-state delta machinery for the same purpose.

Read those tests correctly before touching them:

  • what they exercise is rollback and state reconstruction across undo, retry, takes and divergence;
  • their use of the world-state engine is test instrumentation, not an endorsement of RPG-shaped state as the target architecture. Nothing about them argues against the typed-event model this milestone installs;
  • when M5 replaces or generalizes the world-state protocol, the instrumentation must move with it to the new narrative-state mechanism. In practice that is one schema entry and one helper in backend/tests/fakes.py;
  • preserve or rework these tests; do not delete them merely because their current instrumentation is RPG-shaped. The behaviour they pin is exactly the behaviour M5 is most likely to break.

test_state_revert and test_delete_state additionally assert destructive undo semantics and are expected to be rewritten by M3; that is separate from this instrumentation point, and they should likewise be rewritten rather than dropped.

Evidence: planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md §K.2, §Q.

Note from M3 — three constraints this milestone must satisfy

1. Keep state efficiently recoverable at a retained position. M3's head movement is a row lookup plus a state restore, which is why Undo, Redo and — from M4 — Save Point restore all cost the same regardless of how far into a campaign the position is. A state model recoverable only by replaying events from the campaign opening would make every one of those operations proportional to campaign length, on exactly the long campaigns this product exists for.

TECHNICAL-DESIGN.md §10.4 already selects a hybrid of validated events plus snapshots/cache. M3 turns the snapshot half from a preference into a requirement: keep per-node snapshots, or an equivalent cache with the same property, while adding the typed event model. See ADR 012.

2. Move the instrumentation, keep the assertions. The M2 note above applies with more force after M3: test_head_cursor.py adds roughly a dozen more tests that express position and rollback through the inherited gold counter. What they measure is positional — that the state belonging to a story position is restored when the head moves to it, in either direction, and that an abandoned line's state does not survive a divergence. Those properties must still hold over whatever carries state after M5. The file's own docstring says so.

3. Complete the narrator edit. STORY-BRANCH-SEMANTICS.md §14-15 requires that a narrator edit become authoritative and that the state it implies be re-evaluated. That requirement is intact and unimplemented: re-evaluating state from prose a user typed needs this milestone's extraction pass. M3 shipped the safe interim behavior only — an in-place edit is refused when story descends from the turn and is off screen (§14A), so retained history cannot be made to disagree with itself unseen. M5 is where §14-15 is finished: return to the state before the edited narration, treat the edited text as accepted output, re-evaluate the implied state, create a new continuation, and retain the original. The refusal in §14A is then replaced by that behavior rather than kept alongside it.

Delivered in the M5 corrective pass. The first M5 implementation did the state half only and kept editing the row in place, which the independent review found broke the head/state invariant. A narrator edit now forks — the original narration and its whole future are retained untouched, and the correction becomes a new active continuation carrying its own re-derived state. The §14A refusal is gone for narrator turns and remains only for a player's own input (§13), which M5 did not change.


M5 — Outcome

Complete and accepted, 2026-09-04, after an independent implementation review (planning/archive/milestone-reports/M5-IMPLEMENTATION-REPORT.md) and the corrective pass recorded in that report's addendum.

Delivered:

  • Genre-neutral typed narrative state per ADR 010, with the authoritative document shape now recorded in ADR 013.
  • D10 complete. A narrator edit returns to the state before the turn, uses the reader's exact text, re-derives the state it implies, becomes a new active continuation, and retains the original narration and its future (STORY-BRANCH-SEMANTICS.md §§14-15).
  • Pre-M5 compatibility. Migration 88 backfills the empty narrative document onto every action written before M5, and a missing snapshot restores the empty document rather than leaving a later position's state standing. Old campaigns, Save Points, branches and transcripts stay usable, and no legacy RPG machinery becomes authoritative again.
  • Manual corrections govern the narrator's context. Replayed history carries prose only, and a withdrawn fact is named as no longer true rather than silently dropped.

Debt carried forward, deliberately:

  • M8 (browser UX): the scenario editor still exposes the legacy RPG stat schema. Its copy no longer claims that schema is how the story is tracked, but the screen itself is M8's to redesign.
  • M9 (export/recovery): state_events and state_proposals are not carried in an export, so an imported campaign keeps a correction's effect but not its audit trail. C04 therefore passes for a live campaign and not across a round trip.
  • §13 (editing player input) is still an in-place edit guarded by the off-screen refusal. Bringing it onto the §§14-15 footing was out of M5's scope.

M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory

Objective

Align inherited AI-DnD memory/context behavior with the final authority and history model.

Scope

  • preserve recent active-lineage history selection,
  • ensure summaries are anchored to source turn ranges/lineage rather than positional list assumptions,
  • preserve branch-scoped memory isolation,
  • classify memory authority (accepted vs heuristic/inferred),
  • maintain explicit token budgets,
  • preserve/extend prompt-context inspection,
  • ensure memory/summary failures do not corrupt accepted story state,
  • test memory behavior after Undo/Redo/checkpoint/divergence,
  • validate realistic-context prompt construction.

Explicit Non-Scope

  • imported document knowledge belongs in M7,
  • no remote embeddings/vector DB.

Tests / Acceptance

  • F01-F08,
  • E lineage safety,
  • no abandoned-future term appears in active prompt after divergence,
  • negative control proves eligible memory returns when the relevant lineage is active,
  • summary lineage survives head movement correctly,
  • prompt inspector identifies included memories/summaries and token costs.

Definition of Done

Long-running story context is lineage-safe, authority-aware, local, inspectable, and bounded.

Note from M2 — background memory failure must be observable

M2 shipped with the memory bank entirely dead, and the full suite stayed green. Summaries and embeddings raised AttributeError inside a fire-and-forget task: no user-visible error, no log a player would read, and no failing test, because every memory test stubs the provider factories out.

M6 therefore additionally requires:

  • memory/summarization background failures must be observable — a fire-and-forget task that dies must leave a record a user or maintainer can actually find, rather than being swallowed;
  • tests must exercise at least one real provider-construction/wiring path, not only mocked factories, so that a moved or removed setting surfaces as a test failure (TECHNICAL-DESIGN.md §18.1);
  • derived-memory failure must not corrupt accepted story state. Already in the scope list above; M2's evidence is why it stays there. A dead memory bank degraded the storyteller quietly and left the transcript correct, which is the right failure direction — but it must also be a visible one.

Evidence: planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md §A.1, §9.1.


M6 — Outcome

Complete and accepted 2026-09-06, after an independent review that found E03 still failing and a corrective pass that fixed it. Report, including the review findings and the corrective addendum: planning/archive/milestone-reports/M6-IMPLEMENTATION-REPORT.md.

M7 is authorized: the corrected M6 evidence passes, including E03 end to end against a real local summariser.

What was inherited and kept, rather than rebuilt:

  • Memory lineage was already correct. Memories carry (branch_id, depth) and retrieval filters them through the capped-path clause. The ten-step negative control was measured passing against the M5 baseline before any M6 change, and is now pinned by tests. M6 added provenance to the retrieval result and an authority classification; it did not rewrite the bank.
  • The lineage chokepoint, the cursors, and the post-turn pass are unchanged in shape.

What M6 changed:

  • Summary lineage. Summaries move from a single adventures.story_summary column to summaries rows carrying a coordinate and a source range, filtered by the same lineage clause as memories. E03 was measured leaking at the M5 baseline and is now closed.
  • Memory authority. accepted_story vs heuristic, classified by the application, marked in the prompt.
  • Context budgeting. The reply is reserved out of the context budget, and an impossible configuration raises ContextOverflow rather than producing a prompt known to overflow. Nothing was reserved before M6.
  • Background failure observability. derived_status rows, an API endpoint and an Insights surface, so the M2 failure — the whole memory bank dead with a green suite — is visible if it recurs.
  • Real provider-wiring tests, which mock no factory, plus two pre-existing test-suite leaks they exposed and which are fixed.

What the independent review found, and what the corrective pass did:

  • E03 was still failing (M6-F1). Summary rows were lineage-anchored, but generation was seeded from adventures.story_summary, a campaign-global column with no lineage — so a summary generated after a divergence inherited the abandoned line's prose inside a correctly anchored row. Fixed by seeding from summaries.current. The lesson is recorded in V1-ACCEPTANCE-TESTS.md: a valid E03 test must regenerate a summary after diverging, not merely check that the old row went ineligible.
  • F02 passed by one slot (M6-F2). With a real embedding model, four near-identical memories crowded out the one that mattered; the clue survived only because the default memory_top_k is 5. Redundancy suppression now runs before the final cut, and the clue is retrieved at top_k 5, 4 and 3.
  • Two test defects (M6-F3, M6-F4). The unit fixture made real network calls to the default endpoint, and the E03 test and browser check shared the blind spot above. Both fixed; the browser suite now has a dedicated E03 scenario that regenerates a summary after divergence.
  • A misleading status (M6-F5). Derived work that had nothing to do reported ok; it now reports idle.

Debt carried forward, deliberately:

  • Retrieval ranking is cosine similarity plus an explicit pin. Importance, entity overlap, recency and thread overlap are contemplated by CONTEXT-AND-MEMORY.md §20 and are not implemented. Redundancy suppression covers the failure mode M6 measured; the richer ranking is open, and matters to M7 because imported material will compete for the same budget.
  • Cross-layer duplication (§22) is unimplemented: the same fact can appear in state, memory and history at once. Bounded and legible, but not the "highest-authority concise representation" the plan asks for.
  • M7: imported knowledge is not implemented, and nothing was built to fill its inspector section.
  • M8: the Insights additions are functional, not designed; the broader UX pass remains M8's.
  • M9: derived rows are not carried in an export bundle, so an imported campaign starts with no summaries or memories and rebuilds them. Authoritative history is unaffected.
  • M11: the long-context evidence here is a bounded fixture, not the M01 100-turn campaign.

M7 — First-Class Imported Knowledge Library

Objective

Implement the separate local knowledge subsystem required by the specification rather than overloading AI-DnD Story Cards.

Scope

  • .txt and .md import,
  • source validation and local copy/storage,
  • Canon / Reference / Inspiration classification,
  • enable/disable/delete,
  • content hash and provenance,
  • heading/paragraph-aware chunking,
  • SQLite FTS5 lexical index,
  • local Ollama embeddings where semantic retrieval is enabled,
  • hybrid retrieval/reranking,
  • authority-aware context insertion,
  • campaign scoping,
  • source/chunk inspector,
  • prompt retrieval provenance,
  • export/import preservation,
  • no automatic URL/image fetch,
  • prompt-injection framing as untrusted data.

Explicit Non-Scope

  • PDF/DOCX/EPUB import,
  • remote URL ingestion,
  • remote vector stores,
  • general plugin/tool framework.

Tests / Acceptance

  • G01-G10,
  • C05 canon beats reference/inspiration,
  • hidden/abandoned story facts cannot leak through lineage-derived knowledge,
  • H06-H09 as applicable to rendering/import/archive handling,
  • I05 knowledge provenance survives export/import.

Definition of Done

A campaign can import local Canon/Reference/Inspiration files, retrieve them locally with provenance, and maintain authority boundaries.

Status: COMPLETE

Accepted 2026-09-06, after an independent review that returned PASS WITH CORRECTIVE WORK REQUIRED, a corrective pass that closed both blocking findings, and a closeout verification that resolved the calibration boundary the corrective pass had left as debt. Report, including the original findings, the corrective closeout and the closeout verification, all preserved in sequence: archive/milestone-reports/M7-IMPLEMENTATION-REPORT.md.

M8 is authorized.

Capabilities M7 delivered, which later milestones inherit rather than build:

  • a first-class imported knowledge library — .txt/.md import, Canon / Reference / Inspiration classification that decides prompt framing, ranking weight and budget rather than labelling a list, enable/disable/delete, content hashing, deterministic heading-aware chunking, SQLite FTS5, local Ollama embeddings, hybrid retrieval, campaign isolation and a source inspector;
  • relevance admission separated from ranking, so retrieval can return nothing (§13.2 of TECHNICAL-DESIGN.md);
  • prompt provenance that survives its source — the rendered text travels in the turn's snapshot, so deleting a source cannot orphan a historical prompt;
  • imported text framed as untrusted data with the authority order stated in words, verified against a real narrator;
  • an import surface that accepts no filesystem path at all, so H08 is satisfied by the absence of the mechanism.

The two blocking findings the review raised, both closed:

  • M7-F1 — retrieval had no effective no-match gate. Relevance was decided by a floor expressed as a share of the best candidate, which the best clears by construction, so a passage was admitted on every turn regardless of the scene. A query about tide tables retrieved all five sources of a fantasy campaign, hidden Canon among them. Corrected by separating relevance admission from ranking: admission now uses raw, candidate-set-independent signals, and retrieval may return nothing. TECHNICAL-DESIGN.md §13.2 records the lesson.
  • M7-F2 — the retrieval suite could not detect F1. Its stub embedder scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, so the broken gate passed. Corrected with a stub that has a deliberate similarity floor, plus a test that fails if the floor is ever removed and one that demonstrates the superseded rule still being fooled by the same fixture.

What was built, beyond the scope list above:

  • The classification is load-bearing, not a label. It decides the framing a passage is given in the prompt, the weight it carries in ranking, and which budget it competes in. classes.py is the single place all three read.
  • Lexical retrieval is a production path. SQLite FTS5 with porter stemming, campaign- and enabled-scoped in SQL, bounded by a LIMIT before any Python ranking runs. The library is fully usable with no embedding model configured, and a dead inference host costs the semantic half and nothing else.
  • Relevance admission is a separate stage from ranking (added by the corrective pass). Admission reads raw signals — the cosine the model returned, and how many distinct meaningful query terms a passage contains — so it can answer "nothing matched". Ranking reads normalized ones, because bm25 has no fixed range and a real embedding model scores any two pieces of English around 0.3-0.6. The semantic floor is measured against the production model and recorded beside the constant.
  • The class multiplies relevance rather than adding to it, which is what makes "relevant Canon outranks equally relevant Reference" and "irrelevant Canon does not win on class alone" both true.
  • Provenance is the rendered text, not a foreign key. A turn's knowledge record carries what the narrator was actually shown, so deleting a source cannot turn historical evidence into dangling ids.
  • The import surface accepts no pathname at all, so H08 is satisfied by the absence of the mechanism rather than by a check that could be bypassed.

Debt carried forward, deliberately:

  • Entity linking, tags, manual priority and scene pinning are not implemented (IMPORTED-KNOWLEDGE-DESIGN.md §33-36). Entity and place names do reach the retrieval query, because it is built partly from the authoritative state, but there is no explicit source-to-entity link.
  • Conflict detection between two Canon sources (§42) is not implemented. Two Canon sources that disagree are both retrieved and both framed as Canon.
  • Source versioning (§14, §43) is not implemented. A duplicate is refused with a conflict, or imported deliberately as a second source; there is no supersession chain.
  • Canon scope metadata — invariant / initial / descriptive / historical (§45) — is not implemented. Current-state precedence stands in its place, which §45 itself permits for v1.
  • The semantic scan is linear over the campaign's vectors, capped at 4,000 passages, with the shortfall reported rather than hidden. There is no approximate-nearest-neighbour index in v1.
  • The semantic admission floor is calibrated for one embedding model, and the product now knows that. nomic-embed-text was measured over 113 production-path pairs. An uncalibrated model does not inherit the number: semantic retrieval is skipped for it and the library degrades to lexical-only with the reason reported (TECHNICAL-DESIGN.md §13.3). What remains open is only the enhancement — calibrating further models, each a measurement rather than a guess. The cost meanwhile is a conceptual-only paraphrase going unretrieved under an uncalibrated model, which is a missing passage rather than an irrelevant one.
  • M8: the knowledge panel and the Insights knowledge rows are functional, not designed. Two labels missing from the Insights section table since M5 were added while M7 was in that file; the rest of the panel's design is M8's. M8 should also consider how a "nothing was relevant enough" result and an uncalibrated-model warning should look — both are surfaced plainly today.

Two defects were found by the browser run and fixed in this pass rather than carried:

  • The Insights panel rendered state_rule and state_reminder as raw keys, because M5's two sections were never added to the label table.
  • The open source inspector refetched the source on every render of the Play screen, which re-renders on every keystroke in the story box — 38 needless requests for a 38-character sentence. The reporter callback was in the effect's dependencies and arrives as a fresh function each render. The browser suite gained a check that types and counts requests; it was verified to fail against the unfixed code before the fix was kept.
  • M9: the bundle carries knowledge sources but still carries no context snapshots, so an imported campaign has no historical prompt provenance for any component — which is what a pre-M7 bundle already did for every other one.

M8 — Browser UX Completion for v1 Story Operations

Objective

Turn the adapted AI-DnD interface into the focused interactive-story workspace defined by BROWSER-UX-SPEC.md.

Scope

  • campaign library/setup streamlined for non-RPG stories,
  • story transcript and streaming polish,
  • one primary natural-language input flow plus story-direction affordance,
  • Undo/Redo/Retry/alternate-take controls,
  • Save Points UI,
  • edit user input/narrator output with safe lineage semantics,
  • current scene/state side panel,
  • knowledge panel,
  • context inspector,
  • local Ollama status/model selection,
  • no cloud-provider UI,
  • reserve microphone/STT affordance without implementing STT.

Explicit Non-Scope

  • no full branch-management UI,
  • no image/video/TTS/STT implementation,
  • no mobile-native app.

Tests / Acceptance

  • B01-B04,
  • D user-facing behavior,
  • relevant UX acceptance checks,
  • browser state remains consistent after restart/Undo/Redo/checkpoint/failed generation.

Definition of Done

Normal story creation and play feels like a focused local storyteller rather than an RPG or developer console.

Status: COMPLETE / ACCEPTED — 2026-09-06

Implemented on m8-browser-ux from the signed M7 commit 480414e, verified in a real browser against a real local narrator, independently reviewed, and accepted at closeout on 2026-09-06. The independent review returned M8 IMPLEMENTATION: PASS subject to evidence and documentation cleanup, which the closeout completed: the build-evidence classification in the report's §P and finding 14's resolution.

planning/archive/milestone-reports/M8-IMPLEMENTATION-REPORT.md records what was built, what was measured and every finding, including the seven product defects verification found, the five harness defects, and the evidence runs that were discarded.

Closeout evidence

  • Real-browser verification against a real local narrator — 157 checks across six suites, zero failures, all from one frozen production build (index-Ii-lARp9.js): acceptance 57/57, hidden-information sentinel 21/21, A05 failed generation 23/23, security/offline 27/27, genuine process restart 12/12, migration against an M7-built database 17/17.
  • A frontend test foundation, which the project had never had — 132 tests across 10 files, running in about six seconds. Backend: 950 passed, 14 skipped, 0 failed. Lint, production build and Docker build clean.
  • Reader-facing terminology audit: 0 hits across 17 normal-play components and both advanced surfaces.
  • No schema change and no migration. Proved by building a database with a server running the M7 commit's own code and opening it with M8's.
  • BROWSER-UX-SPEC.md §38 clarified and ratified — hidden narrator information is withheld at the surface that can actually expose it, and is absent from the DOM rather than collapsed in it. No hidden-state subsystem was invented, and none is required.
  • Carry-forward: the four M9 handoff questions in the report's §U (portability scope, historical context-snapshot provenance, legacy story cards, context-window portability), and finding 14's context-window ceiling, whose long-campaign release validation stays with M11.

M9 is next and has not been started.

What M8 delivered, beyond the scope list above:

  • One entry point, and everything else inside a campaign. The navigation was Home · Adventures · Scenarios · Settings · AI Chat; it is now the campaign library and Settings, with State, Knowledge, Context, Save Points and campaign Settings reachable from the story screen's own panel.
  • A campaign is created from a form, not from a template. The scenario gallery and its editor — a JSON stat-schema form, a story-card table and an art picker — are gone from the browser. Setup asks for a name and offers genre, tone, voice, length, protagonist, opening scene and canon, none of it required and none of it genre-specific.
  • campaign_canon reached the browser. It has been the highest authority in a campaign since M5, read by the prompt builder and the state validator, and had no API at all — a fixture had to write it with SQL.
  • A resolution path for a blank model. Settings.model could be empty with nothing saying so until a turn failed. The header now reports Ollama's state in five cases, and an unconfigured or missing model offers the models actually installed on the endpoint. Nothing is chosen automatically: an endpoint's first model may be an embedding model, which cannot narrate.
  • Failures have §71's taxonomy rather than one toast, and a failed turn leaves the reader's words in the box (A05).
  • Safe Markdown for story prose. Headings, emphasis, lists, blockquotes and code, built as React elements from parsed text — no dangerouslySetInnerHTML anywhere. A remote image is a placeholder, and a javascript: URL never becomes an href.
  • A frontend test suite, which the project had never had.

Two defects found by driving the product, and fixed here:

  • > You I enter the tavern. AI Dungeon's player-input convention prefixes > You , which was right for the Do mode's bare verb phrase and wrong for §12's one natural-language field. It reached the transcript, the replayed history and therefore the narration, where a small model imitated it. Now the prefix is added only when the reader has not already written a subject.
  • A stale .input-bar { display: flex } in play.css overrode the new composer layout, because that sheet is imported after the new one. Found in the first browser pass, not by reading the CSS.

Debt carried forward, deliberately:

  • A deployment's context ceiling may be far below the app's budget — RESOLVED operationally, no code change. Ollama defaults to a 4,096-token input window when it sees no VRAM; Settings.context_token_budget defaults to 16,384. The OpenAI-compatible endpoint the app speaks accepts num_ctx and silently ignores it, and reloads the model at its own default, so priming over the native API does not help either. The remedy, verified end to end through the app's own path, is a derived model carrying PARAMETER num_ctx, created over /api/create and selected in Settings — no shell access on the Ollama host, no application change, and it appears in the model picker automatically. DEVELOPMENT.md carries the procedure under "The context window your Ollama actually enforces". Nothing was truncated in M8 — the largest assembled prompt across every suite was 2,498 tokens. This is now an operational/deployment note, not an open defect: M11 should confirm the window on the deployment it certifies against before its 100-turn run, and M9 should note that an imported long campaign reaches a small ceiling immediately. A Settings-screen warning that reads the real window and compares it to the budget remains an unowned usability improvement.
  • Story cards are no longer editable in the browser. The backend keeps them and the bundle still carries them; the M7 knowledge library supersedes them for v1, and showing both would offer two unrelated systems for "things the narrator should know".
  • The RPG world state is read-only. It still appears in the context inspector for a campaign that has one; the editing drawer is gone (§18).
  • Copy is per-message only. §77's "copy the whole transcript" and §78's story search are not built.
  • No discarded-history recovery screen (§63) — explicitly future work.
  • Tablet is usable, not tuned. Desktop was the target; the narrow-screen sheet is inherited and was not redesigned.

M9 — Export, Backup, Recovery, and Migration Hardening

Objective

Make campaigns portable and recoverable without losing lineage, state, knowledge, or the active head.

Scope

  • finalize documented campaign bundle format/versioning,
  • preserve active branch/head, retained history, retries, checkpoints,
  • state events/snapshots,
  • prompt/retrieval provenance as selected,
  • knowledge source metadata/content/index rebuild information,
  • scene/media metadata placeholders,
  • safe SQLite backup behavior,
  • import validation,
  • backward-compatible migration rules,
  • corruption/failure handling where practical.

Tests / Acceptance

  • I01-I06,
  • L01-L04,
  • undone-head round trip,
  • branched campaign round trip,
  • checkpoint round trip,
  • knowledge provenance round trip,
  • derived indexes can be rebuilt.

Definition of Done

A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact.

Status: COMPLETE — 2026-09-07, pending independent review

Implemented on m9-recovery from the signed M8 commit 1ce9972, measured before and after against the same fixture, and verified in a real browser against a real narrator. planning/archive/milestone-reports/M9-IMPLEMENTATION-REPORT.md is the implementer's account, written for a reviewer.

What it delivered, beyond the scope list above:

  • The bundle became a version, and the reason is a rule. ai-dnd-adventure-v3. Everything M9 adds could have been an optional key, the way four earlier additions were — and that mechanism fails exactly here, because a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign with none". A version is how a recovery file states what it was capable of recording. v1 and v2 are still read, and every seam from pre-active-head onward is tested.
  • A third category of data. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided. They are derived and must travel, so the rule is now chosen / evidence / rebuildable, and the test separating the last two is not "could this be recomputed" but "would a recomputation answer the same question".
  • The M8 handoff on prompt provenance is closed. An old turn in a restored campaign shows what it was actually given, after the source has been deleted, the canon edited and the state moved on.
  • State events and proposals travel, so a moved campaign can still say why its state is what it is — and a manual correction is still identifiable as one, which it was not before.
  • Summaries travel with their coordinates, so an abandoned line's summary is still ineligible after the move, and a moved campaign resumes with its long-story continuity instead of behaving like a new one.
  • A verified SQLite backup, using the online backup API rather than a file copy, taken while the application is running, with a browser control in Settings.
  • Take parentage, memory authority and parser/chunking versions, each closing a smaller fidelity loss.

Three defects found by running the milestone's own tests, and fixed here:

  1. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error. Reindex could not repair it either. Pre-existing since M7; both ends are now closed and an already-damaged database repairs itself with no migration.
  2. An imported node with no state snapshot was stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20 — M5's review finding 3, arriving through the import.
  3. A snapshot's source_id was not being translated on import because the relink mutated a dict in place, which a non-MutableDict column does not notice. Found by a test asserting the outcome rather than the call.

One decision the brief asked for, made and recorded:

Story cards are compatibility-only legacy data, and no longer enter the narrator's prompt. They still travel in both directions, the rows and the API stay, and memorybank.cast_brief still reads them as the summariser's character roster. What stops is the injection: a keyword-matched card arrived in front of the narrator as World Lore: … with no class, no visibility, no source, no way to switch it off since M8 removed the editor, and no row in the context inspector — which is IMPORTED-KNOWLEDGE-DESIGN.md §73's "alternate untracked path around the new knowledge authority/provenance rules" in as many words.

Debt carried forward, deliberately:

  • A long campaign's bundle has a measured ceiling: ~279 turns against the 20 MB import limit. A per-turn prompt contains the story so far, so carrying one per turn is O(turns²); compressing them inside the file cut that to about an eighth of what it would have been. M9's additions account for only 12% of the ceiling. The other 88% is the per-position narrative state document, which is 74% of a bundle and which v2 already carried — so lifting the ceiling means addressing that, not the evidence. Far beyond M11's 100-turn certification (14% of the cap), and stated with its measurement rather than hidden. A streaming or chunked import is the fix if a later milestone needs one; the asymmetry to know about is that such a campaign can still be exported and would be refused on import.
  • No discarded-history recovery screen (§63). M9's job was that retained history survives correctly so a later screen can use it; it does.
  • No whole-transcript copy and no story search (§77, §78) — still M11 or later.

Post-M8 hands-on playtest findings — recorded 2026-09-07, owned by M11

Not M9 defects, not caused by M9, and they did not block M9 acceptance. Recorded here rather than only in the M9 report because a milestone report is archived when the next one replaces it, and these must not go with it.

They come from a real play session against accepted, signed M8: real browser, real trusted-LAN Ollama, narrator qwen2.5:3b-instruct-16k, disposable isolated campaign database. That database was deliberately destroyed afterwards, so the stored context snapshot for finding D is gone and no root cause is claimed for it. Full write-up, with the verification behind each mechanism, is in the M9 report's §Y (in archive/milestone-reports/ once M10's report replaces it).

These are M11's, and explicitly not M10's. M10 is media-readiness architecture and stays bounded; it inherits them as known carry-forward only.

A. The browser still calls the product "AI D&D" — M11 release polish

frontend/index.html still carries the inherited <title>AI D&amp;D</title>. M8 changed the navigation, the inspector and the screens, and never claimed the title — so this is an uncovered gap rather than a false claim.

Do not fix it with a find-and-replace to "Adventure Storyteller". SPECIFICATION.md requires a genre-agnostic engine, and Adventure is narrower than the product. The naming decision is the repository owner's; a neutral working name such as Interactive Story and a tab form such as <Campaign Name> — Interactive Story are candidates, not decisions.

B. After Undo, the reader cannot tell where they are — M11 UX polish

Undo behaved correctly (M3 semantics; re-verified throughout M9). The reader could not tell which point in the story they had moved to.

BROWSER-UX-SPEC.md §8 said only "The current endpoint should be clear", and its companion sentence was already true while the reader was lost — so the requirement could not hold the behaviour. §8 has been strengthened to state the orientation requirement; no UI text is prescribed, because none is ratified. A Moment 8 → Moment 7 style indicator is a candidate. Needs a browser regression scenario.

C. Narration length has no measurable effect — M11 realistic-model behaviour

The setup choice becomes one English sentence in the campaign's ai_instructions and changes no generation setting. Independently, length_hint() derives a numeric word range from the global Settings.max_output_tokens and places it after the history — and at the default 800 it reads "must not exceed 506 words, and it should not stop short of about 177" identically for brief, medium and long.

That is a mechanism, verified by reading and running the code — not a proven cause of what the reader saw; the narrator's instruction following is also in play. A reproduction must measure what enters the stored prompt, whether the setting moves any generation budget, and actual word/paragraph counts across repeated turns, on the reference 3B narrator and a stronger local one. Clearer numeric targets (Brief ~100-200 words and so on) are a design candidate, not ratified. Do not hard-truncate prose — the state block is emitted last and truncation removes it.

D. Character identity / coreference confusion — M11 diagnostic

Four people in one scene — Bill (protagonist), Roger, John, Alice — and later narration treated Alice as two different Alices.

Root cause UNKNOWN and no longer establishable. Candidates: a model coreference failure on a correct prompt; duplicate/conflicting state; a context/summary/memory assembly failure; or a context that is not contradictory but too implicit for a small model.

One structural fact to check first, verified by reading the code: the narrative state permits two entities to share a display name and reports nothing. Entities are keyed by the model-supplied id; DUPLICATE_ENTITY rejects only a repeated key; no check exists on name. That is one of this finding's failure modes, and establishes nothing about what happened.

M11 must run an explicit diagnostic with a protagonist and three same-scene supporting characters, stressing pronouns, dialogue attribution, entrances and exits, reference by name and by role, and one character speaking about another. It must detect duplicate creation, same-name duplication, protagonist drift, misattributed dialogue, self-as-other reference, and state/context disagreement — and on any failure preserve the pre-generation state, the exact stored prompt snapshot, history, summaries, memories, imported knowledge, narrator output and model settings, then classify:

STATE DEFECT / CONTEXT ASSEMBLY DEFECT / DERIVED MEMORY-SUMMARY DEFECT /
MODEL FAILURE WITH CORRECT CONTEXT / AMBIGUOUS

Do not "fix" a model failure by changing authoritative state, and do not blame the model if the prompt already contained the error. M9 made all of that evidence portable, so a failing campaign can be exported and handed over intact.

The standard fixture does not cover this class. TEST-CAMPAIGN-FIXTURE.md's seven traps are knowledge, authority, branch leakage and possession; there is no identity trap and its on-stage cast is effectively two people. The established fixture was not modified — it is the deterministic baseline earlier results are compared against. A companion fixture, Multi-Character Identity Test, is proposed in an appendix to that document.


M10 — Future Media Extension Hooks Only

Objective

Preserve the approved future media interfaces without adding a media-generation dependency to v1.

Note — the post-M8 playtest findings are not M10 scope

The four findings recorded above are owned by M11. M10 inherits them as known carry-forward items only: it should neither implement nor test them, and its scope below is unchanged by them. They are listed before this milestone rather than after it only because they were recorded during M9's closeout.

Scope

  • scene snapshots/packets suitable for future providers,
  • optional visual character/location/item descriptors,
  • provider-neutral media request/job/asset types or reserved schema as justified,
  • branch/lineage association for scenes/assets,
  • local-only endpoint contract for future providers,
  • STT contract: local transcription -> editable draft -> normal submission,
  • no core story-engine dependency on media availability.

Explicit Non-Scope

Do not implement:

  • image generation,
  • video generation,
  • TTS,
  • STT,
  • ambience/audio generation,
  • ComfyUI/FLUX integration.

Tests / Acceptance

  • K01-K04 architecture/media-readiness tests,
  • no v1 story flow requires a media service,
  • scene data follows active lineage after Undo/restore/divergence.

Definition of Done

Future media providers can be added through defined local interfaces without redesigning core story authority/history.

Status: COMPLETE — 2026-09-07, pending independent review

Implemented on m10-media-hooks from the signed M9 commit 44edece. planning/archive/milestone-reports/M10-IMPLEMENTATION-REPORT.md is the implementer's account, written for a reviewer.

The finding that shaped the milestone: the scene snapshot already existed.

The media contract's §5 asks for a persisted or derived scene snapshot, and M5 built one three milestones ago. narrative_state["scene"] holds the summary, the location, who is present and the (branch_id, depth) coordinate; it is written by a validated set_scene event, snapshotted per position, and restored on every head move. That was verified with a probe — a campaign played, diverged, undone and exported — rather than taken from M9's report.

So M10 built no scenes table, and the Scene Packet is derived on read with a computed identity (c<adventure>:b<branch>:<start>-<end>) rather than an allocated one. A second scene store would have been a duplicate representation of the same fact, with its own lineage rules to get wrong; the lineage rules are the hard part, which is precisely the argument for reusing the ones that already work.

What it delivered:

  • One table, visual_profiles — the only field in the contract's scene list that nothing already stored. Campaign-scoped rather than per-position, because a character does not change appearance when the story forks, and because per-position profiles would have cost 245 copies of the same 367 bytes in a 120-turn campaign to say something that never varies.
  • A scene packet built on read, excluding the raw transcript, all imported knowledge, memories and summaries. Excluding imported knowledge as a class is what keeps a hidden Canon source out of a future depiction without a filter anyone has to remember to extend.
  • Provider contracts as typing.Protocol structural types, with an empty registry, no adapter, and no dependency added. Nothing imports a media library because none is installed.
  • The STT asymmetry in the type: a DraftTranscription is editable and has no commit method, so a transcriber structurally cannot bypass the authoritative commit path.
  • A stricter endpoint policy than narration uses — loopback only, reusing endpoints.py's resolved-address check rather than trusting a hostname. No media configuration setting exists, because one that exists can be pointed at a cloud by mistake.
  • Profiles travel in the bundle with no format bump. M9's own semantic test decides it: an absent visualProfiles key is unambiguous, because a campaign with no profiles is the ordinary case. Older v3 files still import.

A defect found by the milestone's own tests, and fixed here:

  • A redundant index migration made two databases disagree. M10 first shipped migration 93 creating ix_visual_profiles_adventure. create_all already builds ix_visual_profiles_adventure_id from the column's index=True, on fresh installs and existing databases alike — so an upgraded database ended up with both indexes and a fresh one with only the second. The comparison of a fresh schema against an upgraded schema is what caught it. M10 adds no migration at all; LATEST_VERSION stays 92.

Carry-forward, unchanged by this milestone: the four post-M8 playtest findings recorded above (browser title, post-Undo orientation, narration length, character identity) remain M11's. M10 neither implemented nor tested them.


M11 — v1 Security, Long-Run, and Release Validation

Objective

Validate the full product against the release contract after all functional milestones are integrated.

Scope

  • full outbound-network blocked run,
  • dependency/runtime audit,
  • restrictive browser/CORS/CSP checks,
  • failed-model-call recovery,
  • long-running 100-turn campaign,
  • repeated Undo/Redo/Retry/checkpoint cycles,
  • branch/memory/summary leakage checks,
  • realistic-context state extraction across selected recommended models,
  • fantasy and science-fiction fixtures,
  • export/import/recovery tests,
  • migration tests,
  • documentation and packaging,
  • the four post-M8 hands-on playtest findings above: the browser product name, reader orientation after history movement, a narration-length setting with a measurable effect, and the multi-character identity diagnostic with its companion fixture.

Tests / Acceptance

Release gate in V1-ACCEPTANCE-TESTS.md:

  • all REQUIRED FOR V1 tests pass,
  • approved exceptions documented in ADRs,
  • 100-turn test passes,
  • offline operation passes,
  • branch/memory isolation passes,
  • export/import recovery passes,
  • both genre fixtures pass.

Definition of Done

The build meets the v1 black-box acceptance contract and can be packaged as the first production release.

Status: IMPLEMENTED AND VERIFIED — 2026-09-07; long-run evidence complete 2026-09-13; awaiting independent review/acceptance

Implemented on m11-release-validation from the signed M10 commit 1013c94. planning/reports/M11-IMPLEMENTATION-REPORT.md is the evidence package, written for a release reviewer. M11 is not marked accepted here; that is the reviewer's to record, and no release tag exists.

The release blocker it was given, and how it was closed. M8 measured the reference deployment enforcing a 4,096-token input window while the application budgeted 16,384 — every request returning 200, and llama.cpp dropping the oldest tokens, which in this design are the narrator's rules and the campaign canon. A 100-turn certification against that server would have looked perfect and proved nothing.

M11's invariant: the application must not silently budget more narrator input than the runtime will accept. It now asks the server — /api/ps for a loaded model, /api/show for one that is not — under the same endpoint policy and TLS trust as inference, and caps the prompt to what it finds, or records the window as unverified in the turn's own provenance. Not a hard-coded 4,096, which would cripple a correctly configured deployment; not a guess from the model's name. Measured on the reference server: the plain model reports 4,096 and the budget caps to it; the num_ctx-baked model reports 16,384 and the full budget stands.

The four post-M8 playtest findings, disposed of:

Disposition
A. The tab read AI D&D Fixed. Interactive Story, with the open campaign first — a name chosen by the repository owner, and deliberately not "Adventure Storyteller", which is narrower than a genre-agnostic engine. One module owns it.
B. No orientation after Undo Fixed. Moment 11 · later story ahead, from the server's own answer, in the transcript's existing vocabulary, with no implementation words. BROWSER-UX-SPEC.md §8A records the implementation.
C. Narration length had no effect Fixed. The choice is data (adventures.narration_length), and the prompt builder turns it into a real word band. The generation budget is deliberately untouched: capping it would truncate prose, and the state block is emitted last.
D. Character identity confusion Diagnostic built; root cause remains unestablished, as it must. The campaign was destroyed. tools/m11_identity.py runs the finding's own scenario, makes only the judgements a program can make honestly, preserves everything on a signal, and proves its detectors fire. The structural fact the finding asked M11 to check first — two entities may share a display name silently — is now reported rather than refused, because two people called Alice is ordinary fiction.

Two product defects found by the release validation itself, both the same family — something true that nobody was told:

  1. A partly refused manual state correction reported success. Four changes, one refused, HTTP 201, nothing said. Found because the identity diagnostic's own fixture was refused that way and ran on a degraded campaign without noticing. The refusal was already on the audit record; the reader was not told. Now returned as refused, and shown in the State panel.
  2. The narration-length setting above, which is finding C.

Also corrected: M9's residual risk 6 was narrower than recorded — the snap Firefox refuses a WebDriver file path under /tmp, not all paths. Staging under $HOME makes browser file import work, so knowledge import is now proved end-to-end in a real browser rather than in two labelled halves.

The long-run evidence (2026-09-13). The first report left M01 PARTIAL at 41 turns, and that run was then lost to a host crash with its evidence. M01-M04 now pass on a complete 100-turn run on 96c1bf5: 101 accepted turns, three genuine restarts, all thirteen scheduled history operations, zero failed post-turn passes, and recovery onto a clean data directory 16 of 16. M04's precondition is positional: the planting turn was outside the history window. The fact was recovered through authoritative state and reached memory only as the narrator's restatement of it. The repository owner accepted that on 2026-09-13. The report's §G tells the whole path, including six runs that are not the evidence.

Two further product defects, found only once a long run had the memory bank switched on:

  1. A turn locked its own memory bank out. Retrieval wrote a use counter before the model call, and the turn committed after the reply, so SQLite's one write lock was held through the reply. Post-turn memory and summary writes timed out, and so did recording their failure. The counter is now written in the turn's own commit, and a failure is recorded after a rollback. (f8d4010)
  2. The narrator's protocol was stored as story. The model wrote a pasted copy of the state section, unfenced and unfinished proposals, and sections of its own, on up to 42 of 104 turns, and stored text is replayed as history. The extractor now removes every shape observed. Replaying 443 real turns through it changed no turn it had previously left clean. (0c7316f, 96c1bf5)

Left for the reviewer, in the report's §P: the real-token headroom at the largest prompts is 23-42 tokens; the narrator restates prompt text in its prose; and the identity diagnostic's results are not in the report.


4. Milestone Dependency Summary

M1 Production/offline foundation
  -> M2 Local-only surface reduction
      -> M3 Non-destructive history + export head
          -> M4 Save Points
              -> M5 Narrative state
                  -> M6 Context/memory
                      -> M7 Imported knowledge
                          -> M8 Browser UX completion
                              -> M9 Export/recovery hardening
                                  -> M10 Future-media hooks
                                      -> M11 v1 validation/release

Some implementation work may overlap internally, but milestone acceptance should remain sequential so architectural regressions are discovered early.

4. Prompting Rule

When implementation begins, prepare one Codex prompt per milestone.

Do not hand the entire plan to Codex as one production task.

Each prompt should contain only:

  • milestone objective,
  • relevant source documents,
  • scope/non-scope,
  • specific acceptance tests,
  • stop condition and required report/diff.

No implementation prompt is included in the current planning revision.