Files
interactive-story/planning/BUILD-MILESTONES.md
JesseMarkowitzandClaude Opus 5 a6e9c7a32b M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-06 03:00:33 -04:00

1013 lines
46 KiB
Markdown

# Adventure Storyteller — Production Build Milestones
**Status:** In implementation. M1-M6 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6: 2026-09-06, each of the last two after an independent review and a corrective pass); M7 — First-Class Imported Knowledge Library — next to brief
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
## 1. Purpose
This document defines the production implementation sequence after Phase 0.
It is intentionally a milestone plan, **not a Codex execution prompt**. Each milestone should be converted into a separate coding brief only after the previous milestone has been reviewed and accepted.
The milestone order prioritizes load-bearing correctness and offline safety before broad feature work.
## 2. Global Rules for Every Milestone
Each implementation milestone must:
- preserve upstream provenance and licensing notices,
- keep the application runnable whenever practical,
- add or update automated tests for changed behavior,
- avoid unrelated refactors,
- update affected documentation,
- record schema/migration changes,
- preserve local-only defaults,
- stop at the milestone boundary for review,
- avoid implementing later milestones opportunistically.
The product specification and acceptance tests outrank convenience inherited from AI-DnD.
## 3. End-User Capability Progression
The application is expected to remain runnable throughout the sequence. **M8 does not create the browser UI from scratch**; AI-DnD already provides a browser interface from M1. M8 is where that inherited/adapted interface is completed and simplified into the intended v1 storyteller experience.
- **M1 — First playable baseline:** The user can open the inherited browser UI, start or resume a story, send prompts to Ollama, receive streamed narration, and persist the story across restart. The interface may still look and behave substantially like AI-DnD, and advanced history/state features are not yet converted to final semantics.
- **M2 — Clean local storyteller baseline:** Normal story play still works, but cloud/hosted/account/scripting surfaces are removed and Ollama can be either same-host or on an explicitly configured trusted-LAN machine. From the user's perspective this is still basic play, but the environment now matches the intended privacy/deployment model.
- **M3 — Safe history controls:** The user can play normally with production-grade Undo, Redo, Retry, alternate takes, and divergence without deleting old story history. Export/import also preserves the exact current position even when the user has undone several turns.
- **M4 — Save Points:** The user can name a point in the story, continue playing, restart, return to that Save Point, and take a different path without losing the later story. This is the first milestone where deliberate long-form experimentation/recovery should feel comfortable.
- **M5 — Reliable story state:** Characters, locations, possessions, relationships, facts, threads, and current scene become application-owned genre-neutral state rather than RPG-stat machinery. The user can begin inspecting/correcting what the storyteller believes, although the final polished state UI comes later.
- **M6 — Long-story memory:** The storyteller becomes much better suited to long campaigns because summaries and retrieved older memories remain branch-safe and inspectable. The user should be able to return to old people, promises, clues, and events without abandoned paths contaminating the active story.
- **M7 — Local knowledge library:** The user can import `.txt`/`.md` material as Canon, Reference, or Inspiration and have it retrieved locally with provenance. The functionality is usable, but some management/inspection surfaces may still be utilitarian until M8.
- **M8 — Finished v1 browser experience:** The existing browser UI is reorganized and polished around storytelling: streamlined campaign setup, transcript/input, Undo/Redo/Retry, Save Points, state, knowledge, context inspection, and model status. This milestone makes the product feel like the intended storyteller rather than an adapted AI-DnD application.
- **M9 — Portable/recoverable campaigns:** The user can reliably export, back up, import, and recover complete campaigns including history, Save Points, state, knowledge, and active position. This is where moving or restoring a campaign becomes a supported user workflow rather than merely an underlying capability.
- **M10 — Media-ready, still text-first:** Little or nothing visibly changes for normal play. The story/scene data and provider boundaries are prepared so image/video/audio/TTS/STT can be added later without redesigning the core application.
- **M11 — Release-quality v1:** The user experience should be functionally complete; this milestone proves it stays correct through long campaigns, repeated history operations, offline use, failures, export/import, and multiple genres. It is the release-validation milestone rather than a major new feature milestone.
---
# M1 — Establish Production Fork and Offline Baseline
## Objective
Create the production fork from the pinned AI-DnD commit and make ordinary local story operation genuinely offline-capable before larger changes.
## Scope
- establish production repository/fork lineage,
- record upstream commit and license provenance,
- establish reproducible dev/test environment,
- vendor/cache/replace the `tiktoken` first-use encoding dependency so story generation works without Internet,
- eliminate runtime Google Fonts/remote font dependency,
- tighten CSP for local runtime assets,
- verify same-host Ollama story generation with outbound Internet blocked,
- verify explicitly configured trusted-LAN Ollama story generation while the storyteller UI/API remains loopback-bound,
- establish baseline regression/test report.
**Added during implementation:** outbound TLS trust. A trusted-LAN Ollama may be
served over HTTPS with a privately issued certificate, which the inherited HTTP
client refused because it verified against a bundled public-CA list only. M1
made outbound HTTPS verify against the operating system's CA store as well,
with verification and hostname checking fully intact and no bypass option
(ADR 002; `backend/app/tlstrust.py`). This was not in the scope list above but
was on M1's critical path: without it the trusted-LAN Definition of Done below
is unreachable against a realistic host.
## Explicit Non-Scope
- no history rewrite yet,
- no RPG-state redesign,
- no imported knowledge,
- no UI redesign,
- no media generation.
## Tests / Acceptance
Must demonstrate:
- A01-A06 as applicable to the inherited base, including the separate-host trusted-LAN inference path,
- H01-H03 and H11,
- no unexpected DNS/HTTP requests during ordinary story generation after setup,
- no remote font request,
- no tokenizer/BPE download on first production turn,
- existing relevant AI-DnD tests remain green except documented upstream/environment exceptions.
## Definition of Done
A clean production build can start, open the browser UI, generate and persist story turns through either same-host Ollama or an explicitly configured trusted-LAN Ollama host, restart, and resume with outbound Internet blocked. The storyteller UI/API remains loopback-bound by default.
## Status: COMPLETE
Accepted 2026-09-02. Evidence: `planning/archive/milestone-reports/M1-BASELINE-REPORT.md` (run
logs and packet captures) and `planning/archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md`
(review report). A01-A06, H01-H03 and H11 all pass on runtime evidence; 648
backend tests pass, including with no route to the Internet.
**Capabilities M1 delivered, which later milestones inherit rather than build:**
- offline first turn — the tokenizer table is vendored and digest-checked,
- no remote runtime assets — fonts self-hosted, CSP names no remote origin,
- loopback-bound storyteller in every run path, including the published Docker port,
- **working trusted-LAN inference, plain HTTP and HTTPS with a private CA**,
demonstrated against a second physical machine,
- a reproducible environment (`backend/requirements.lock`) and an offline-capable test suite.
---
# M2 — Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface
## Objective
Reduce AI-DnD to the intended single-user local product trust boundary without destabilizing the story/tree/memory foundation.
## Scope
Remove or isolate as appropriate:
- multi-user/account/guest/auth flows,
- demo API-key behavior,
- hosted analytics,
- Render/Neon deployment paths,
- Postgres/`psycopg` support,
- cloud providers not required by v1,
- arbitrary remote model-provider UI/configuration,
- QuickJS/campaign scripting,
- AI-Dungeon compatibility code that exists primarily to support scripting/hosted behavior,
- hosted-only rate-limit/account infrastructure.
Add:
- explicit Ollama endpoint policy. **Trusted-LAN inference already works as of
M1** — same-host loopback default, user-configured LAN endpoint, HTTP or
HTTPS with a privately issued certificate, verified against the machine's CA
store. M2's job is not to invent that capability but to **formalize and
narrow** it: decide and enforce what an endpoint may be in normal v1
configuration, reject or remove what it may not, and keep the endpoint's
configuration from affecting the storyteller's own loopback bind. Do not
regress the TLS behavior while narrowing the surface — the shared
verification context must follow any client the removal work rewrites,
- clear local-model connection diagnostics,
- migration/test instrumentation replacements for any tests that depended on JS hooks or removed hosted paths.
## Explicit Non-Scope
- do not redesign story history,
- do not replace RPG state yet,
- do not build knowledge/media systems.
## Tests / Acceptance
- local story play still works,
- existing tree/retry/memory/context behavior remains intact,
- application requires no cloud API keys,
- approved trusted-LAN Ollama endpoints work — including an HTTPS endpoint with a privately issued certificate, per A06 — while arbitrary public/Internet provider endpoints are rejected or absent from normal production configuration,
- removed provider/auth/analytics/scripting paths are no longer reachable from normal production configuration.
## Definition of Done
The codebase has a narrow single-user/local-only surface and the inherited story foundation still passes its relevant regression suite.
## Status: COMPLETE
Accepted 2026-09-02. Evidence: `planning/archive/milestone-reports/M2-BASELINE-REPORT.md`
(measurements) and `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` (review);
verdict *accept with non-blocking debt, proceed to M3*. Implementation is
commit `8c65ae9`, and the three defects the review found are commit `8652fe7`
— see the closeout note appended to both reports.
**Capabilities M2 delivered, which later milestones inherit rather than build:**
- a single-user product with no accounts, sessions or auth — all four
`/api/auth/*` routes are gone, not gated; 52 API routes down to 36,
- Ollama as the only inference backend, with no cloud provider code and no API
key anywhere in the product,
- an **address-based inference endpoint policy**, enforced on save and again
before every outbound request, that refuses public addresses even when the
database is edited behind the settings API (ADR 011),
- M1's trusted-LAN and TLS behaviour carried through the removal intact —
verified HTTPS against a real second machine, no bypass option,
- SQLite as the only store; Postgres, Neon and the Render deployment path removed,
- no campaign scripting: the QuickJS engine and `/api/scripts` are gone,
- a materially simpler runtime — 10 environment variables down to two, 6 Python
and 21 npm packages removed, and a 933 kB bundle down to 395 kB.
**Debt carried forward, none of it blocking M3:** `Settings.model` still
defaults to `""` with nothing prompting for it (M8); inert legacy tables and
columns await a cleanup migration once the schema settles, after M3/M5; there
are still no frontend tests (M8). `docs/*.html`, upstream's project site, was
listed here as still linking Google Fonts; the whole inherited `docs/` tree was
deleted in the 2026-09-03 documentation pass, which closes that item. Full table
in the implementation report §P.
---
# M3 — Production Non-Destructive History, Redo, and Active-Head Export
## Objective
Replace destructive Undo with the demonstrated head-cursor model and make Undo/Redo/divergence/export/import production-safe.
## Scope
Promote the Phase 0B spike concept into maintainable production code:
- active head can move behind retained tip,
- Undo moves head and deletes zero accepted turns,
- Redo follows the retained continuation,
- new write behind tip forks on first write,
- retry/add-take/edit paths use the same safe fork/head rules,
- prevent Undo from walking before the campaign opening/root semantics,
- mark displaced/inactive futures/takes as retained/disposable using implementation-appropriate metadata,
- keep ordinary Redo invalidated after divergence,
- add browser Redo control and correct disabled/enabled states,
- keep branch complexity out of normal UI,
- export active head coordinate/depth as well as active branch,
- import honors exported active head,
- preserve backward compatibility for older bundles that lack the new field.
## Explicit Non-Scope
- named checkpoints belong in M4,
- no general branch-tree UI,
- no abandoned-history cleanup feature.
## Tests / Acceptance
Must cover:
- D01-D10,
- D02 minimum-five Undo,
- D03 practical unlimited Undo behavior across retained history,
- D04-D05 Redo semantics,
- E01-E04 lineage isolation,
- I01-I03 and I07 (undone-head round trip),
- L01-L02,
- branch-scoped memory remains isolated under Undo/Redo/divergence,
- row counts demonstrate Undo deletes zero accepted turns.
## Definition of Done
History operations are non-destructive, Redo works, divergence preserves old futures, and export/import reopens at the exact active head.
## Status: COMPLETE
Accepted 2026-09-03. Evidence: `planning/archive/milestone-reports/M3-IMPLEMENTATION-REPORT.md`,
which is M3's primary evidence record — no separate baseline report was produced,
so that document carries the raw counts and runtime observations as well as the
review. The architecture is recorded in **ADR 012**.
**Capabilities M3 delivered, which later milestones inherit rather than build:**
- **Undo that deletes zero accepted turns** — measured directly on row identity:
15 rows before five Undos, 15 after,
- **Redo**, round-tripping exactly (head 14 → 4 → 14) in transcript and in state,
- a **single head-movement mechanism** every position change goes through, and a
**single capped lineage** that bounds every read of the story,
- **divergence decided by the lineage** rather than by a flag: the first write
below a moved-back head forks, the displaced future keeps its rows, and
ordinary Redo stops offering it with nothing to invalidate,
- **branch-scoped memory isolation preserved for free** — a memory past the head
is unretrievable and becomes eligible again on Redo, with no pruning and no
re-embedding,
- **an export that carries the reader's position**, so a campaign exported after
two Undos imports still undone, with pre-M3 bundles opening at their tip,
- **Undo across fork points** to the campaign opening, and refusal to change what
a turn says while story descends from it off screen — both ratified in
`STORY-BRANCH-SEMANTICS.md` (§5, §10, §14A).
**Closeout condition — CLOSED at M4 closeout (2026-09-03).** M3 was accepted with
one condition outstanding: the required **browser smoke test had not been
performed**, because no session in which M3 was implemented or reviewed had a
browser available, leaving the DOM-level behaviour of the Redo button and its
disabled states unverified by observation.
That condition is now discharged. A real Firefox 154.0.1, driven through
geckodriver, exercised M3's controls in the rendered application: Undo enabled
and Redo disabled at the tip, two Undos moving the transcript back, Redo becoming
enabled and returning the original tip exactly, Retry and the take pager, and a
divergent write retiring Redo with no stale old-future text on screen. It passed.
Evidence: `planning/reports/M4-IMPLEMENTATION-REPORT.md` §W.7.
**Debt carried forward, none of it blocking M4:** full narrator-edit state
re-evaluation is deferred to M5 (`STORY-BRANCH-SEMANTICS.md` §14A records the
interim refusal); `POST /adventures/import` returns every branch's rows rather
than a head-capped window (inherited, harmless in the UI); `ActionPage` is
constructed in two places. Full table in the implementation report §S.
---
# M4 — Named Save Points / Checkpoints
## Objective
Add durable user-facing Save Points on top of the active-head model.
## Scope
- named checkpoint storage,
- checkpoint create/list/rename/delete,
- restore by moving the active head,
- later history retained rather than deleted,
- fork only on first new continuation after restore,
- checkpoint persistence across restart,
- checkpoint export/import,
- browser Save Point UX per `BROWSER-UX-SPEC.md`.
## Explicit Non-Scope
- no branch merge,
- no automatic discarded-history cleanup,
- no complex branch explorer required.
## Tests / Acceptance
- D11-D14,
- E lineage tests after checkpoint restore/divergence,
- I04,
- L03.
## Definition of Done
The user can create a named Save Point, continue, restart, restore it, and continue differently without losing later history.
## Note from M3 — reuse the head machinery, do not build a second one
A Save Point is **a durable named pointer to a recoverable story position**, and
nothing more. M3 made that position a stored coordinate and made moving to one a
row lookup plus a state restore, so restoring a Save Point is head movement with
a bounds check — not a restore system of its own.
Concretely, M4 should:
- store the coordinate, the name, and the metadata around them, and no copy of
any story;
- restore by calling M3's head-movement mechanism, so that state, transcript,
context and memory eligibility all move together exactly as they do for Undo
and Redo, and so that later history is retained rather than deleted (D13 is
already satisfied by the mechanism);
- let the existing fork-on-first-write-below-the-head rule handle divergence
after a restore, rather than forking at restore time;
- validate that a Save Point's coordinate is still on the lineage being read
before moving to it.
A second restore path is the specific failure to avoid. The Phase 0B spike put
the fork check in the write path and left Retry and add-take on the old one, and
M3's cost was reconciling them; a parallel checkpoint mover would recreate that
divergence in a place where the two paths would silently disagree about what
"restore" means. See ADR 012.
## Status: COMPLETE
Accepted 2026-09-03. Evidence: `planning/reports/M4-IMPLEMENTATION-REPORT.md`,
including its §W closeout addendum. The Definition of Done is met, and — for the
first time in this project — **verified in a real browser**.
**What M4 delivered:**
- **`checkpoints`**, a table holding a name, an optional note, and a
`(branch, depth)` coordinate — and no copy of any story. `create_all` builds
it, as it did `memories` and `branches`; migration 80 adds the index. No
backfill: nobody had named a position before M4, and inventing one would be
inventing the decision.
- **Create / list / rename / delete / restore** under
`/api/adventures/{id}/checkpoints`, campaign-scoped, with a Save Point from
another campaign a 404 rather than a restore of the wrong story.
- **Create at the active head, not the retained tip**, so a Save Point made after
two Undos names the undone position.
- **Restore that delegates**, and is the whole of the milestone's architecture:
resolve the coordinate, refuse it if it names no live turn, then
`head.move_to_node` — one function whose depth half is M3's `head.move_to`
unchanged, and whose branch half is the single assignment `switch_branch`
makes. Restore forks nothing.
- **A browser Save Point panel** — create form, list, Restore, Rename, Delete,
with both confirmations saying what is *not* destroyed — plus a Save Point
button beside Undo and Redo, where it belongs. No branch explorer, no
discarded-history browser, no merge UI.
- **Export/import of Save Points** with no format version bump, and a pre-M4
bundle importing with none.
**The one architectural decision M4 had to make**, which ADR 012 does not settle:
a Save Point can name a position on a line the story has since left, so restore
moves the branch half of the head as well — but *only* when the coordinate is not
on the path being read. Doing it unconditionally would quietly hand back an
abandoned continuation whenever a Save Point in a shared prefix was restored. No
new ADR: this is ADR 012's mechanism applied to both halves of a coordinate ADR
012 already defines, not a new architecture. `TECHNICAL-DESIGN.md` §8.8 records
it.
**Tests:** 42 in `backend/tests/test_save_points.py`, covering D11-D14, I04, L03,
E-series lineage and memory isolation after restore and divergence, the edge
cases in the brief, and the M3-database migration.
**Facts M5 inherits, and must not redesign:**
- a Save Point is a **durable story coordinate** — `(branch, depth)` — carrying
no copy of transcript, state, prompt, memory or summary;
- **restore is M3 head movement**, through the same `head.move_to` Undo and Redo
use; there is no second restore path and M5 must not add one;
- **state recovery stays snapshot/cached-position based**, never a replay of the
campaign (`TECHNICAL-DESIGN.md` §10.4). This is now load-bearing for Save
Points as well as Undo/Redo;
- **the first divergent write** after a restore creates the continuation;
restore itself never forks;
- **history a Save Point names cannot disappear** through an unrelated deletion
(§19.1);
- **M3/M4 history and Save Point semantics are infrastructure now.** M5 replaces
the state *model*; it does not revisit how the story is positioned.
**Acceptance evidence:** D11-D14, I04 and L03 all pass. D11 and L03 are
discharged by automation that crosses a **genuine OS process boundary** — one
server process writes the campaign, is killed, and a second process reads it back
— rather than by recreating a client in one process.
**The browser condition is closed, for M4 and retrospectively for M3.** A real
Firefox 154.0.1, driven through geckodriver over the W3C WebDriver protocol,
exercised the rendered DOM end to end: **44/44 checks passed**, covering M3's
Undo/Redo enable states and transcript movement, Retry and the take pager, M3
divergence and the disappearance of Redo, and every M4 Save Point operation
including both confirmations and the branch-delete warning. No console errors.
This closes the outstanding M3 condition recorded above and the equivalent M4
one.
**The three review findings were fixed during closeout:**
- **B-1** — the Save Point list was an N+1 that loaded whole `Action` rows,
narration included. It is now one bulk two-column coordinate query plus one
lineage: **53 SELECTs for 25 Save Points became 5**, and the query count no
longer moves with the length of the list.
- **B-2** — deleting a branch silently deleted the Save Points naming it. Fixed
as a **behaviour** defect, not a wording one: a branch a Save Point names can
no longer be deleted at all. The request is refused with the offending Save
Points named, the user deletes them explicitly (which deletes no story), and
the branch then goes. Both delete controls, in the branch list and in the tree
overlay, disable and explain rather than warning about a loss that no longer
happens. `STORY-BRANCH-SEMANTICS.md` **§19.1** records the rule; §28 already
required a future cleanup feature to retain checkpoint-referenced paths, and
this is that requirement applied to the deletion path that exists today.
- **B-3** — the D11/L03 automation now spawns real server processes.
**Also fixed:** creating a Save Point takes the campaign's turn lock (review §S
C-5), so "save where I am" cannot read a head that a turn in flight is about to
move. Rename and Delete deliberately do not take it — neither reads nor moves a
story position.
**Debt M4 carries forward:** the Save Point panel has no frontend test, because
the project still has no frontend test runner at all (M8) — the browser smoke
test above is a closeout procedure, not a suite. `POST /adventures/import` still
returns every branch's rows rather than a head-capped window (inherited, M3).
## Note to M5 — the instrumentation is now larger than M3 estimated
M4 added **60 tests that use inherited RPG world-state values as deterministic
instrumentation**, on top of the ~20 M3 flagged. M5 replaces that state model,
and must **move the instrumentation while preserving the behavioural
assertions**: what those tests measure is where the story is being read and what
state belongs to that position, which is exactly as true after M5 as before it.
Deleting them would delete the evidence for D11-D14, I04, L03 and the E-series.
The existing constraint stands and is now load-bearing for Save Points as well as
Undo/Redo: **historical state must remain efficiently snapshot/cache
recoverable**, so that moving to a position never becomes proportional to
campaign length (`TECHNICAL-DESIGN.md` §10.4). Restore, Undo and Redo all pay
whatever that costs.
---
# M5 — Genre-Neutral Authoritative Narrative State
## Objective
Replace/generalize AI-DnD's RPG-specific relative-delta state system with the approved genre-neutral typed-event model.
## Scope
Establish generic authoritative state for:
- entities,
- characters/locations/organizations/items/vehicles as descriptive entity categories,
- facts,
- relationships,
- possession/location/status/conditions,
- story threads,
- scene state,
- manual state/canon corrections,
- state proposal provenance.
Implement:
- explicit typed state proposal schema,
- absolute/unambiguous event semantics,
- event allowlist,
- validation,
- atomic accepted-event commit,
- current/historical snapshot/cache,
- rollback/reconstruction tied to active lineage,
- browser current-state inspector foundation.
Remove or demote:
- D&D-specific stats/bands/cooldowns/mechanics from the core product model,
- relative-delta semantics as the generic state protocol.
## Explicit Non-Scope
- no optional RPG module,
- no imported knowledge yet,
- no final rich state-editing UX if a simpler inspector is enough for this milestone.
## Tests / Acceptance
- C01-C04 and C06,
- D/E tests verifying state follows head movement,
- H05 invalid state event rejection,
- L01-L02,
- malformed JSON/state proposal handling,
- realistic-context tests against representative local models,
- fantasy and science-fiction state fixtures use the same schema.
## Definition of Done
Accepted story state is genre-neutral, auditable, reconstructable, and no longer depends on ambiguous relative deltas.
## Note from M2 — eight rollback tests use the world-state engine as instrumentation
M2 removed QuickJS. Eight existing rollback/history tests had used a JavaScript
counter as deterministic instrumentation — a value they could change on a turn
and then assert had been rolled back — and they now use the inherited
RPG/world-state delta machinery for the same purpose.
Read those tests correctly before touching them:
- **what they exercise** is rollback and state reconstruction across undo,
retry, takes and divergence;
- **their use of the world-state engine is test instrumentation, not an
endorsement** of RPG-shaped state as the target architecture. Nothing about
them argues against the typed-event model this milestone installs;
- **when M5 replaces or generalizes the world-state protocol, the
instrumentation must move with it** to the new narrative-state mechanism. In
practice that is one schema entry and one helper in `backend/tests/fakes.py`;
- **preserve or rework these tests; do not delete them** merely because their
current instrumentation is RPG-shaped. The behaviour they pin is exactly the
behaviour M5 is most likely to break.
`test_state_revert` and `test_delete_state` additionally assert *destructive*
undo semantics and are expected to be rewritten by M3; that is separate from
this instrumentation point, and they should likewise be rewritten rather than
dropped.
Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §K.2, §Q.
## Note from M3 — three constraints this milestone must satisfy
**1. Keep state efficiently recoverable at a retained position.**
M3's head movement is a row lookup plus a state restore, which is why Undo, Redo
and — from M4 — Save Point restore all cost the same regardless of how far into a
campaign the position is. A state model recoverable only by replaying events from
the campaign opening would make every one of those operations proportional to
campaign length, on exactly the long campaigns this product exists for.
`TECHNICAL-DESIGN.md` §10.4 already selects a hybrid of validated events plus
snapshots/cache. M3 turns the snapshot half from a preference into a requirement:
keep per-node snapshots, or an equivalent cache with the same property, while
adding the typed event model. See ADR 012.
**2. Move the instrumentation, keep the assertions.**
The M2 note above applies with more force after M3: `test_head_cursor.py` adds
roughly a dozen more tests that express position and rollback through the
inherited gold counter. What they measure is *positional* — that the state
belonging to a story position is restored when the head moves to it, in either
direction, and that an abandoned line's state does not survive a divergence.
Those properties must still hold over whatever carries state after M5. The file's
own docstring says so.
**3. Complete the narrator edit.**
`STORY-BRANCH-SEMANTICS.md` §14-15 requires that a narrator edit become
authoritative and that the state it implies be re-evaluated. That requirement is
intact and unimplemented: re-evaluating state from prose a user typed needs this
milestone's extraction pass. M3 shipped the safe interim behavior only — an
in-place edit is refused when story descends from the turn and is off screen
(§14A), so retained history cannot be made to disagree with itself unseen.
M5 is where §14-15 is finished: return to the state before the edited narration,
treat the edited text as accepted output, re-evaluate the implied state, create a
new continuation, and retain the original. The refusal in §14A is then replaced
by that behavior rather than kept alongside it.
**Delivered in the M5 corrective pass.** The first M5 implementation did the
state half only and kept editing the row in place, which the independent review
found broke the head/state invariant. A narrator edit now forks — the original
narration and its whole future are retained untouched, and the correction
becomes a new active continuation carrying its own re-derived state. The §14A
refusal is gone for narrator turns and remains only for a player's own input
(§13), which M5 did not change.
---
## M5 — Outcome
**Complete and accepted, 2026-09-04**, after an independent implementation
review (`planning/reports/M5-IMPLEMENTATION-REPORT.md`) and the corrective pass
recorded in that report's addendum.
Delivered:
- Genre-neutral typed narrative state per ADR 010, with the authoritative
document shape now recorded in
[ADR 013](DECISIONS/013-authoritative-narrative-state-document.md).
- **D10 complete.** A narrator edit returns to the state before the turn, uses
the reader's exact text, re-derives the state it implies, becomes a new active
continuation, and retains the original narration and its future
(`STORY-BRANCH-SEMANTICS.md` §§14-15).
- **Pre-M5 compatibility.** Migration 88 backfills the empty narrative document
onto every action written before M5, and a missing snapshot restores the empty
document rather than leaving a later position's state standing. Old campaigns,
Save Points, branches and transcripts stay usable, and no legacy RPG machinery
becomes authoritative again.
- **Manual corrections govern the narrator's context.** Replayed history carries
prose only, and a withdrawn fact is named as no longer true rather than
silently dropped.
Debt carried forward, deliberately:
- **M8 (browser UX):** the scenario editor still exposes the legacy RPG stat
schema. Its copy no longer claims that schema is how the story is tracked, but
the screen itself is M8's to redesign.
- **M9 (export/recovery):** `state_events` and `state_proposals` are not carried
in an export, so an imported campaign keeps a correction's effect but not its
audit trail. C04 therefore passes for a live campaign and not across a round
trip.
- **§13 (editing player input)** is still an in-place edit guarded by the
off-screen refusal. Bringing it onto the §§14-15 footing was out of M5's scope.
---
# M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory
## Objective
Align inherited AI-DnD memory/context behavior with the final authority and history model.
## Scope
- preserve recent active-lineage history selection,
- ensure summaries are anchored to source turn ranges/lineage rather than positional list assumptions,
- preserve branch-scoped memory isolation,
- classify memory authority (accepted vs heuristic/inferred),
- maintain explicit token budgets,
- preserve/extend prompt-context inspection,
- ensure memory/summary failures do not corrupt accepted story state,
- test memory behavior after Undo/Redo/checkpoint/divergence,
- validate realistic-context prompt construction.
## Explicit Non-Scope
- imported document knowledge belongs in M7,
- no remote embeddings/vector DB.
## Tests / Acceptance
- F01-F08,
- E lineage safety,
- no abandoned-future term appears in active prompt after divergence,
- negative control proves eligible memory returns when the relevant lineage is active,
- summary lineage survives head movement correctly,
- prompt inspector identifies included memories/summaries and token costs.
## Definition of Done
Long-running story context is lineage-safe, authority-aware, local, inspectable, and bounded.
## Note from M2 — background memory failure must be observable
M2 shipped with the memory bank entirely dead, and the full suite stayed green.
Summaries and embeddings raised `AttributeError` inside a fire-and-forget task:
no user-visible error, no log a player would read, and no failing test, because
every memory test stubs the provider factories out.
M6 therefore additionally requires:
- **memory/summarization background failures must be observable** — a
fire-and-forget task that dies must leave a record a user or maintainer can
actually find, rather than being swallowed;
- **tests must exercise at least one real provider-construction/wiring path**,
not only mocked factories, so that a moved or removed setting surfaces as a
test failure (`TECHNICAL-DESIGN.md` §18.1);
- **derived-memory failure must not corrupt accepted story state.** Already in
the scope list above; M2's evidence is why it stays there. A dead memory bank
degraded the storyteller quietly and left the transcript correct, which is the
right failure direction — but it must also be a *visible* one.
Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §A.1, §9.1.
---
## M6 — Outcome
**Complete and accepted 2026-09-06**, after an independent review that found E03
still failing and a corrective pass that fixed it. Report, including the review
findings and the corrective addendum:
`planning/reports/M6-IMPLEMENTATION-REPORT.md`.
**M7 is authorized**: the corrected M6 evidence passes, including E03 end to end
against a real local summariser.
What was inherited and kept, rather than rebuilt:
- **Memory lineage was already correct.** Memories carry `(branch_id, depth)`
and retrieval filters them through the capped-path clause. The ten-step
negative control was measured passing against the M5 baseline *before* any M6
change, and is now pinned by tests. M6 added provenance to the retrieval
result and an authority classification; it did not rewrite the bank.
- The lineage chokepoint, the cursors, and the post-turn pass are unchanged in
shape.
What M6 changed:
- **Summary lineage.** Summaries move from a single `adventures.story_summary`
column to `summaries` rows carrying a coordinate and a source range, filtered
by the same lineage clause as memories. E03 was measured leaking at the M5
baseline and is now closed.
- **Memory authority.** `accepted_story` vs `heuristic`, classified by the
application, marked in the prompt.
- **Context budgeting.** The reply is reserved out of the context budget, and an
impossible configuration raises `ContextOverflow` rather than producing a
prompt known to overflow. Nothing was reserved before M6.
- **Background failure observability.** `derived_status` rows, an API endpoint
and an Insights surface, so the M2 failure — the whole memory bank dead with a
green suite — is visible if it recurs.
- **Real provider-wiring tests**, which mock no factory, plus two pre-existing
test-suite leaks they exposed and which are fixed.
What the independent review found, and what the corrective pass did:
- **E03 was still failing (M6-F1).** Summary *rows* were lineage-anchored, but
generation was seeded from `adventures.story_summary`, a campaign-global
column with no lineage — so a summary generated after a divergence inherited
the abandoned line's prose inside a correctly anchored row. Fixed by seeding
from `summaries.current`. The lesson is recorded in `V1-ACCEPTANCE-TESTS.md`:
a valid E03 test must **regenerate** a summary after diverging, not merely
check that the old row went ineligible.
- **F02 passed by one slot (M6-F2).** With a real embedding model, four
near-identical memories crowded out the one that mattered; the clue survived
only because the default `memory_top_k` is 5. Redundancy suppression now runs
before the final cut, and the clue is retrieved at top_k 5, 4 and 3.
- **Two test defects (M6-F3, M6-F4).** The unit fixture made real network calls
to the default endpoint, and the E03 test and browser check shared the blind
spot above. Both fixed; the browser suite now has a dedicated E03 scenario
that regenerates a summary after divergence.
- **A misleading status (M6-F5).** Derived work that had nothing to do reported
`ok`; it now reports `idle`.
Debt carried forward, deliberately:
- **Retrieval ranking** is cosine similarity plus an explicit pin. Importance,
entity overlap, recency and thread overlap are contemplated by
`CONTEXT-AND-MEMORY.md` §20 and are not implemented. Redundancy suppression
covers the failure mode M6 measured; the richer ranking is open, and matters
to M7 because imported material will compete for the same budget.
- **Cross-layer duplication** (§22) is unimplemented: the same fact can appear
in state, memory and history at once. Bounded and legible, but not the
"highest-authority concise representation" the plan asks for.
- **M7:** imported knowledge is not implemented, and nothing was built to fill
its inspector section.
- **M8:** the Insights additions are functional, not designed; the broader UX
pass remains M8's.
- **M9:** derived rows are not carried in an export bundle, so an imported
campaign starts with no summaries or memories and rebuilds them. Authoritative
history is unaffected.
- **M11:** the long-context evidence here is a bounded fixture, not the M01
100-turn campaign.
---
# M7 — First-Class Imported Knowledge Library
## Objective
Implement the separate local knowledge subsystem required by the specification rather than overloading AI-DnD Story Cards.
## Scope
- `.txt` and `.md` import,
- source validation and local copy/storage,
- Canon / Reference / Inspiration classification,
- enable/disable/delete,
- content hash and provenance,
- heading/paragraph-aware chunking,
- SQLite FTS5 lexical index,
- local Ollama embeddings where semantic retrieval is enabled,
- hybrid retrieval/reranking,
- authority-aware context insertion,
- campaign scoping,
- source/chunk inspector,
- prompt retrieval provenance,
- export/import preservation,
- no automatic URL/image fetch,
- prompt-injection framing as untrusted data.
## Explicit Non-Scope
- PDF/DOCX/EPUB import,
- remote URL ingestion,
- remote vector stores,
- general plugin/tool framework.
## Tests / Acceptance
- G01-G10,
- C05 canon beats reference/inspiration,
- hidden/abandoned story facts cannot leak through lineage-derived knowledge,
- H06-H09 as applicable to rendering/import/archive handling,
- I05 knowledge provenance survives export/import.
## Definition of Done
A campaign can import local Canon/Reference/Inspiration files, retrieve them locally with provenance, and maintain authority boundaries.
---
# M8 — Browser UX Completion for v1 Story Operations
## Objective
Turn the adapted AI-DnD interface into the focused interactive-story workspace defined by `BROWSER-UX-SPEC.md`.
## Scope
- campaign library/setup streamlined for non-RPG stories,
- story transcript and streaming polish,
- one primary natural-language input flow plus story-direction affordance,
- Undo/Redo/Retry/alternate-take controls,
- Save Points UI,
- edit user input/narrator output with safe lineage semantics,
- current scene/state side panel,
- knowledge panel,
- context inspector,
- local Ollama status/model selection,
- no cloud-provider UI,
- reserve microphone/STT affordance without implementing STT.
## Explicit Non-Scope
- no full branch-management UI,
- no image/video/TTS/STT implementation,
- no mobile-native app.
## Tests / Acceptance
- B01-B04,
- D user-facing behavior,
- relevant UX acceptance checks,
- browser state remains consistent after restart/Undo/Redo/checkpoint/failed generation.
## Definition of Done
Normal story creation and play feels like a focused local storyteller rather than an RPG or developer console.
---
# M9 — Export, Backup, Recovery, and Migration Hardening
## Objective
Make campaigns portable and recoverable without losing lineage, state, knowledge, or the active head.
## Scope
- finalize documented campaign bundle format/versioning,
- preserve active branch/head, retained history, retries, checkpoints,
- state events/snapshots,
- prompt/retrieval provenance as selected,
- knowledge source metadata/content/index rebuild information,
- scene/media metadata placeholders,
- safe SQLite backup behavior,
- import validation,
- backward-compatible migration rules,
- corruption/failure handling where practical.
## Tests / Acceptance
- I01-I06,
- L01-L04,
- undone-head round trip,
- branched campaign round trip,
- checkpoint round trip,
- knowledge provenance round trip,
- derived indexes can be rebuilt.
## Definition of Done
A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact.
---
# M10 — Future Media Extension Hooks Only
## Objective
Preserve the approved future media interfaces without adding a media-generation dependency to v1.
## Scope
- scene snapshots/packets suitable for future providers,
- optional visual character/location/item descriptors,
- provider-neutral media request/job/asset types or reserved schema as justified,
- branch/lineage association for scenes/assets,
- local-only endpoint contract for future providers,
- STT contract: local transcription -> editable draft -> normal submission,
- no core story-engine dependency on media availability.
## Explicit Non-Scope
Do not implement:
- image generation,
- video generation,
- TTS,
- STT,
- ambience/audio generation,
- ComfyUI/FLUX integration.
## Tests / Acceptance
- K01-K04 architecture/media-readiness tests,
- no v1 story flow requires a media service,
- scene data follows active lineage after Undo/restore/divergence.
## Definition of Done
Future media providers can be added through defined local interfaces without redesigning core story authority/history.
---
# M11 — v1 Security, Long-Run, and Release Validation
## Objective
Validate the full product against the release contract after all functional milestones are integrated.
## Scope
- full outbound-network blocked run,
- dependency/runtime audit,
- restrictive browser/CORS/CSP checks,
- failed-model-call recovery,
- long-running 100-turn campaign,
- repeated Undo/Redo/Retry/checkpoint cycles,
- branch/memory/summary leakage checks,
- realistic-context state extraction across selected recommended models,
- fantasy and science-fiction fixtures,
- export/import/recovery tests,
- migration tests,
- documentation and packaging.
## Tests / Acceptance
Release gate in `V1-ACCEPTANCE-TESTS.md`:
- all REQUIRED FOR V1 tests pass,
- approved exceptions documented in ADRs,
- 100-turn test passes,
- offline operation passes,
- branch/memory isolation passes,
- export/import recovery passes,
- both genre fixtures pass.
## Definition of Done
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
---
## 4. Milestone Dependency Summary
```text
M1 Production/offline foundation
-> M2 Local-only surface reduction
-> M3 Non-destructive history + export head
-> M4 Save Points
-> M5 Narrative state
-> M6 Context/memory
-> M7 Imported knowledge
-> M8 Browser UX completion
-> M9 Export/recovery hardening
-> M10 Future-media hooks
-> M11 v1 validation/release
```
Some implementation work may overlap internally, but milestone acceptance should remain sequential so architectural regressions are discovered early.
## 4. Prompting Rule
When implementation begins, prepare **one Codex prompt per milestone**.
Do not hand the entire plan to Codex as one production task.
Each prompt should contain only:
- milestone objective,
- relevant source documents,
- scope/non-scope,
- specific acceptance tests,
- stop condition and required report/diff.
No implementation prompt is included in the current planning revision.