M2's review reported six planning recommendations rather than applying them, three marked before M3. All six are applied here, plus three additions drawn from the same evidence. No implementation file is touched. The endpoint policy was the gap that mattered. It is the most consequential setting in the application — the storyteller sends the player's prose, the context, the memories and the embedding inputs to whatever address it names — and it existed only as a module docstring. It is now ADR 011 and a new §10A in the threat model, which also retires the assumption in §71A that the inherited guard was a starting point. It was not: AI-DnD's SSRF guard blocked private addresses to stop a hosted server reaching its own internal network, which is the exact opposite of what a local storyteller needs. It was removed, not adapted. Both documents state the rule as implemented — an allowlist of explicit local-network CIDRs, every resolved address checked, enforced on save and again before every outbound request, TLS never traded against it — and both state the two residual limits plainly rather than implying they are covered: a hostile host already on the trusted LAN is inside the permitted boundary, and a rebinding interval exists between the policy's resolution and the client's connection. Accepted risks, not M3 work. The CIDRs are spelled out rather than derived from is_private/is_reserved, and the ADR records why: is_private is true of the documentation ranges and 0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it refuses an ordinary same-host Ollama on [::1]. TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening items. A new §5.2 records the M1/M2 architecture as fact rather than intention, so later milestones inherit what the code does. A new §18.1 carries the lesson of M2's two regressions: when removing a setting, test a real consumer construction path; when adding one, prove it reaches the component that uses it. Both defects hid behind a green suite because the tests at that boundary were mocks. BUILD-MILESTONES records M2 complete, with the capabilities later milestones inherit and the debt carried forward. Two notes go to milestones that would otherwise misread what M2 left them. M5 is told that eight rollback tests now use the world-state engine as instrumentation and not as endorsement — the instrumentation moves when the protocol does, and those tests are reworked rather than deleted. M6 is told that the memory bank died silently under a green suite, so background failure must be observable and at least one real provider-construction path must be tested. The security contract gains what M2 demonstrated. H10 now names the two conditions that were defects during M2: a wildcard origin must be refused at startup, and an unknown /api path must 404 rather than returning the SPA with 200. New H12 covers endpoint enforcement, and its fourth pass condition is the one that matters — a public endpoint written into the database behind the settings API must still be refused at the wire. A build passing the first three and failing that one has configuration validation only. SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement; it removed capability the specification never asked for. The two M2 reports gain appended closeout notes rather than edits. Their original wording about an uncommitted working tree was true when written, and the note records what happened afterwards: the six-file correction is8652fe7,8c65ae9remains the implementation commit, and the two were never squashed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
666 lines
26 KiB
Markdown
666 lines
26 KiB
Markdown
# Adventure Storyteller — Production Build Milestones
|
|
|
|
**Status:** In implementation. M1 and M2 complete (2026-09-02); M3 next
|
|
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
|
|
|
## 1. Purpose
|
|
|
|
This document defines the production implementation sequence after Phase 0.
|
|
|
|
It is intentionally a milestone plan, **not a Codex execution prompt**. Each milestone should be converted into a separate coding brief only after the previous milestone has been reviewed and accepted.
|
|
|
|
The milestone order prioritizes load-bearing correctness and offline safety before broad feature work.
|
|
|
|
## 2. Global Rules for Every Milestone
|
|
|
|
Each implementation milestone must:
|
|
|
|
- preserve upstream provenance and licensing notices,
|
|
- keep the application runnable whenever practical,
|
|
- add or update automated tests for changed behavior,
|
|
- avoid unrelated refactors,
|
|
- update affected documentation,
|
|
- record schema/migration changes,
|
|
- preserve local-only defaults,
|
|
- stop at the milestone boundary for review,
|
|
- avoid implementing later milestones opportunistically.
|
|
|
|
The product specification and acceptance tests outrank convenience inherited from AI-DnD.
|
|
|
|
## 3. End-User Capability Progression
|
|
|
|
The application is expected to remain runnable throughout the sequence. **M8 does not create the browser UI from scratch**; AI-DnD already provides a browser interface from M1. M8 is where that inherited/adapted interface is completed and simplified into the intended v1 storyteller experience.
|
|
|
|
- **M1 — First playable baseline:** The user can open the inherited browser UI, start or resume a story, send prompts to Ollama, receive streamed narration, and persist the story across restart. The interface may still look and behave substantially like AI-DnD, and advanced history/state features are not yet converted to final semantics.
|
|
- **M2 — Clean local storyteller baseline:** Normal story play still works, but cloud/hosted/account/scripting surfaces are removed and Ollama can be either same-host or on an explicitly configured trusted-LAN machine. From the user's perspective this is still basic play, but the environment now matches the intended privacy/deployment model.
|
|
- **M3 — Safe history controls:** The user can play normally with production-grade Undo, Redo, Retry, alternate takes, and divergence without deleting old story history. Export/import also preserves the exact current position even when the user has undone several turns.
|
|
- **M4 — Save Points:** The user can name a point in the story, continue playing, restart, return to that Save Point, and take a different path without losing the later story. This is the first milestone where deliberate long-form experimentation/recovery should feel comfortable.
|
|
- **M5 — Reliable story state:** Characters, locations, possessions, relationships, facts, threads, and current scene become application-owned genre-neutral state rather than RPG-stat machinery. The user can begin inspecting/correcting what the storyteller believes, although the final polished state UI comes later.
|
|
- **M6 — Long-story memory:** The storyteller becomes much better suited to long campaigns because summaries and retrieved older memories remain branch-safe and inspectable. The user should be able to return to old people, promises, clues, and events without abandoned paths contaminating the active story.
|
|
- **M7 — Local knowledge library:** The user can import `.txt`/`.md` material as Canon, Reference, or Inspiration and have it retrieved locally with provenance. The functionality is usable, but some management/inspection surfaces may still be utilitarian until M8.
|
|
- **M8 — Finished v1 browser experience:** The existing browser UI is reorganized and polished around storytelling: streamlined campaign setup, transcript/input, Undo/Redo/Retry, Save Points, state, knowledge, context inspection, and model status. This milestone makes the product feel like the intended storyteller rather than an adapted AI-DnD application.
|
|
- **M9 — Portable/recoverable campaigns:** The user can reliably export, back up, import, and recover complete campaigns including history, Save Points, state, knowledge, and active position. This is where moving or restoring a campaign becomes a supported user workflow rather than merely an underlying capability.
|
|
- **M10 — Media-ready, still text-first:** Little or nothing visibly changes for normal play. The story/scene data and provider boundaries are prepared so image/video/audio/TTS/STT can be added later without redesigning the core application.
|
|
- **M11 — Release-quality v1:** The user experience should be functionally complete; this milestone proves it stays correct through long campaigns, repeated history operations, offline use, failures, export/import, and multiple genres. It is the release-validation milestone rather than a major new feature milestone.
|
|
|
|
---
|
|
|
|
# M1 — Establish Production Fork and Offline Baseline
|
|
|
|
## Objective
|
|
|
|
Create the production fork from the pinned AI-DnD commit and make ordinary local story operation genuinely offline-capable before larger changes.
|
|
|
|
## Scope
|
|
|
|
- establish production repository/fork lineage,
|
|
- record upstream commit and license provenance,
|
|
- establish reproducible dev/test environment,
|
|
- vendor/cache/replace the `tiktoken` first-use encoding dependency so story generation works without Internet,
|
|
- eliminate runtime Google Fonts/remote font dependency,
|
|
- tighten CSP for local runtime assets,
|
|
- verify same-host Ollama story generation with outbound Internet blocked,
|
|
- verify explicitly configured trusted-LAN Ollama story generation while the storyteller UI/API remains loopback-bound,
|
|
- establish baseline regression/test report.
|
|
|
|
**Added during implementation:** outbound TLS trust. A trusted-LAN Ollama may be
|
|
served over HTTPS with a privately issued certificate, which the inherited HTTP
|
|
client refused because it verified against a bundled public-CA list only. M1
|
|
made outbound HTTPS verify against the operating system's CA store as well,
|
|
with verification and hostname checking fully intact and no bypass option
|
|
(ADR 002; `backend/app/tlstrust.py`). This was not in the scope list above but
|
|
was on M1's critical path: without it the trusted-LAN Definition of Done below
|
|
is unreachable against a realistic host.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- no history rewrite yet,
|
|
- no RPG-state redesign,
|
|
- no imported knowledge,
|
|
- no UI redesign,
|
|
- no media generation.
|
|
|
|
## Tests / Acceptance
|
|
|
|
Must demonstrate:
|
|
|
|
- A01-A06 as applicable to the inherited base, including the separate-host trusted-LAN inference path,
|
|
- H01-H03 and H11,
|
|
- no unexpected DNS/HTTP requests during ordinary story generation after setup,
|
|
- no remote font request,
|
|
- no tokenizer/BPE download on first production turn,
|
|
- existing relevant AI-DnD tests remain green except documented upstream/environment exceptions.
|
|
|
|
## Definition of Done
|
|
|
|
A clean production build can start, open the browser UI, generate and persist story turns through either same-host Ollama or an explicitly configured trusted-LAN Ollama host, restart, and resume with outbound Internet blocked. The storyteller UI/API remains loopback-bound by default.
|
|
|
|
## Status: COMPLETE
|
|
|
|
Accepted 2026-09-02. Evidence: `planning/reports/M1-BASELINE-REPORT.md` (run
|
|
logs and packet captures) and `planning/reports/M1-IMPLEMENTATION-REPORT.md`
|
|
(review report). A01-A06, H01-H03 and H11 all pass on runtime evidence; 648
|
|
backend tests pass, including with no route to the Internet.
|
|
|
|
**Capabilities M1 delivered, which later milestones inherit rather than build:**
|
|
|
|
- offline first turn — the tokenizer table is vendored and digest-checked,
|
|
- no remote runtime assets — fonts self-hosted, CSP names no remote origin,
|
|
- loopback-bound storyteller in every run path, including the published Docker port,
|
|
- **working trusted-LAN inference, plain HTTP and HTTPS with a private CA**,
|
|
demonstrated against a second physical machine,
|
|
- a reproducible environment (`backend/requirements.lock`) and an offline-capable test suite.
|
|
|
|
---
|
|
|
|
# M2 — Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface
|
|
|
|
## Objective
|
|
|
|
Reduce AI-DnD to the intended single-user local product trust boundary without destabilizing the story/tree/memory foundation.
|
|
|
|
## Scope
|
|
|
|
Remove or isolate as appropriate:
|
|
|
|
- multi-user/account/guest/auth flows,
|
|
- demo API-key behavior,
|
|
- hosted analytics,
|
|
- Render/Neon deployment paths,
|
|
- Postgres/`psycopg` support,
|
|
- cloud providers not required by v1,
|
|
- arbitrary remote model-provider UI/configuration,
|
|
- QuickJS/campaign scripting,
|
|
- AI-Dungeon compatibility code that exists primarily to support scripting/hosted behavior,
|
|
- hosted-only rate-limit/account infrastructure.
|
|
|
|
Add:
|
|
|
|
- explicit Ollama endpoint policy. **Trusted-LAN inference already works as of
|
|
M1** — same-host loopback default, user-configured LAN endpoint, HTTP or
|
|
HTTPS with a privately issued certificate, verified against the machine's CA
|
|
store. M2's job is not to invent that capability but to **formalize and
|
|
narrow** it: decide and enforce what an endpoint may be in normal v1
|
|
configuration, reject or remove what it may not, and keep the endpoint's
|
|
configuration from affecting the storyteller's own loopback bind. Do not
|
|
regress the TLS behavior while narrowing the surface — the shared
|
|
verification context must follow any client the removal work rewrites,
|
|
- clear local-model connection diagnostics,
|
|
- migration/test instrumentation replacements for any tests that depended on JS hooks or removed hosted paths.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- do not redesign story history,
|
|
- do not replace RPG state yet,
|
|
- do not build knowledge/media systems.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- local story play still works,
|
|
- existing tree/retry/memory/context behavior remains intact,
|
|
- application requires no cloud API keys,
|
|
- approved trusted-LAN Ollama endpoints work — including an HTTPS endpoint with a privately issued certificate, per A06 — while arbitrary public/Internet provider endpoints are rejected or absent from normal production configuration,
|
|
- removed provider/auth/analytics/scripting paths are no longer reachable from normal production configuration.
|
|
|
|
## Definition of Done
|
|
|
|
The codebase has a narrow single-user/local-only surface and the inherited story foundation still passes its relevant regression suite.
|
|
|
|
## Status: COMPLETE
|
|
|
|
Accepted 2026-09-02. Evidence: `planning/reports/M2-BASELINE-REPORT.md`
|
|
(measurements) and `planning/reports/M2-IMPLEMENTATION-REPORT.md` (review);
|
|
verdict *accept with non-blocking debt, proceed to M3*. Implementation is
|
|
commit `8c65ae9`, and the three defects the review found are commit `8652fe7`
|
|
— see the closeout note appended to both reports.
|
|
|
|
**Capabilities M2 delivered, which later milestones inherit rather than build:**
|
|
|
|
- a single-user product with no accounts, sessions or auth — all four
|
|
`/api/auth/*` routes are gone, not gated; 52 API routes down to 36,
|
|
- Ollama as the only inference backend, with no cloud provider code and no API
|
|
key anywhere in the product,
|
|
- an **address-based inference endpoint policy**, enforced on save and again
|
|
before every outbound request, that refuses public addresses even when the
|
|
database is edited behind the settings API (ADR 011),
|
|
- M1's trusted-LAN and TLS behaviour carried through the removal intact —
|
|
verified HTTPS against a real second machine, no bypass option,
|
|
- SQLite as the only store; Postgres, Neon and the Render deployment path removed,
|
|
- no campaign scripting: the QuickJS engine and `/api/scripts` are gone,
|
|
- a materially simpler runtime — 10 environment variables down to two, 6 Python
|
|
and 21 npm packages removed, and a 933 kB bundle down to 395 kB.
|
|
|
|
**Debt carried forward, none of it blocking M3:** `Settings.model` still
|
|
defaults to `""` with nothing prompting for it (M8); inert legacy tables and
|
|
columns await a cleanup migration once the schema settles, after M3/M5; there
|
|
are still no frontend tests (M8); `docs/*.html`, upstream's project site and not
|
|
served by the app, still links Google Fonts. Full table in the implementation
|
|
report §P.
|
|
|
|
---
|
|
|
|
# M3 — Production Non-Destructive History, Redo, and Active-Head Export
|
|
|
|
## Objective
|
|
|
|
Replace destructive Undo with the demonstrated head-cursor model and make Undo/Redo/divergence/export/import production-safe.
|
|
|
|
## Scope
|
|
|
|
Promote the Phase 0B spike concept into maintainable production code:
|
|
|
|
- active head can move behind retained tip,
|
|
- Undo moves head and deletes zero accepted turns,
|
|
- Redo follows the retained continuation,
|
|
- new write behind tip forks on first write,
|
|
- retry/add-take/edit paths use the same safe fork/head rules,
|
|
- prevent Undo from walking before the campaign opening/root semantics,
|
|
- mark displaced/inactive futures/takes as retained/disposable using implementation-appropriate metadata,
|
|
- keep ordinary Redo invalidated after divergence,
|
|
- add browser Redo control and correct disabled/enabled states,
|
|
- keep branch complexity out of normal UI,
|
|
- export active head coordinate/depth as well as active branch,
|
|
- import honors exported active head,
|
|
- preserve backward compatibility for older bundles that lack the new field.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- named checkpoints belong in M4,
|
|
- no general branch-tree UI,
|
|
- no abandoned-history cleanup feature.
|
|
|
|
## Tests / Acceptance
|
|
|
|
Must cover:
|
|
|
|
- D01-D10,
|
|
- D02 minimum-five Undo,
|
|
- D03 practical unlimited Undo behavior across retained history,
|
|
- D04-D05 Redo semantics,
|
|
- E01-E04 lineage isolation,
|
|
- I01-I03 and I07 (undone-head round trip),
|
|
- L01-L02,
|
|
- branch-scoped memory remains isolated under Undo/Redo/divergence,
|
|
- row counts demonstrate Undo deletes zero accepted turns.
|
|
|
|
## Definition of Done
|
|
|
|
History operations are non-destructive, Redo works, divergence preserves old futures, and export/import reopens at the exact active head.
|
|
|
|
---
|
|
|
|
# M4 — Named Save Points / Checkpoints
|
|
|
|
## Objective
|
|
|
|
Add durable user-facing Save Points on top of the active-head model.
|
|
|
|
## Scope
|
|
|
|
- named checkpoint storage,
|
|
- checkpoint create/list/rename/delete,
|
|
- restore by moving the active head,
|
|
- later history retained rather than deleted,
|
|
- fork only on first new continuation after restore,
|
|
- checkpoint persistence across restart,
|
|
- checkpoint export/import,
|
|
- browser Save Point UX per `BROWSER-UX-SPEC.md`.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- no branch merge,
|
|
- no automatic discarded-history cleanup,
|
|
- no complex branch explorer required.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- D11-D14,
|
|
- E lineage tests after checkpoint restore/divergence,
|
|
- I04,
|
|
- L03.
|
|
|
|
## Definition of Done
|
|
|
|
The user can create a named Save Point, continue, restart, restore it, and continue differently without losing later history.
|
|
|
|
---
|
|
|
|
# M5 — Genre-Neutral Authoritative Narrative State
|
|
|
|
## Objective
|
|
|
|
Replace/generalize AI-DnD's RPG-specific relative-delta state system with the approved genre-neutral typed-event model.
|
|
|
|
## Scope
|
|
|
|
Establish generic authoritative state for:
|
|
|
|
- entities,
|
|
- characters/locations/organizations/items/vehicles as descriptive entity categories,
|
|
- facts,
|
|
- relationships,
|
|
- possession/location/status/conditions,
|
|
- story threads,
|
|
- scene state,
|
|
- manual state/canon corrections,
|
|
- state proposal provenance.
|
|
|
|
Implement:
|
|
|
|
- explicit typed state proposal schema,
|
|
- absolute/unambiguous event semantics,
|
|
- event allowlist,
|
|
- validation,
|
|
- atomic accepted-event commit,
|
|
- current/historical snapshot/cache,
|
|
- rollback/reconstruction tied to active lineage,
|
|
- browser current-state inspector foundation.
|
|
|
|
Remove or demote:
|
|
|
|
- D&D-specific stats/bands/cooldowns/mechanics from the core product model,
|
|
- relative-delta semantics as the generic state protocol.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- no optional RPG module,
|
|
- no imported knowledge yet,
|
|
- no final rich state-editing UX if a simpler inspector is enough for this milestone.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- C01-C04 and C06,
|
|
- D/E tests verifying state follows head movement,
|
|
- H05 invalid state event rejection,
|
|
- L01-L02,
|
|
- malformed JSON/state proposal handling,
|
|
- realistic-context tests against representative local models,
|
|
- fantasy and science-fiction state fixtures use the same schema.
|
|
|
|
## Definition of Done
|
|
|
|
Accepted story state is genre-neutral, auditable, reconstructable, and no longer depends on ambiguous relative deltas.
|
|
|
|
## Note from M2 — eight rollback tests use the world-state engine as instrumentation
|
|
|
|
M2 removed QuickJS. Eight existing rollback/history tests had used a JavaScript
|
|
counter as deterministic instrumentation — a value they could change on a turn
|
|
and then assert had been rolled back — and they now use the inherited
|
|
RPG/world-state delta machinery for the same purpose.
|
|
|
|
Read those tests correctly before touching them:
|
|
|
|
- **what they exercise** is rollback and state reconstruction across undo,
|
|
retry, takes and divergence;
|
|
- **their use of the world-state engine is test instrumentation, not an
|
|
endorsement** of RPG-shaped state as the target architecture. Nothing about
|
|
them argues against the typed-event model this milestone installs;
|
|
- **when M5 replaces or generalizes the world-state protocol, the
|
|
instrumentation must move with it** to the new narrative-state mechanism. In
|
|
practice that is one schema entry and one helper in `backend/tests/fakes.py`;
|
|
- **preserve or rework these tests; do not delete them** merely because their
|
|
current instrumentation is RPG-shaped. The behaviour they pin is exactly the
|
|
behaviour M5 is most likely to break.
|
|
|
|
`test_state_revert` and `test_delete_state` additionally assert *destructive*
|
|
undo semantics and are expected to be rewritten by M3; that is separate from
|
|
this instrumentation point, and they should likewise be rewritten rather than
|
|
dropped.
|
|
|
|
Evidence: `planning/reports/M2-IMPLEMENTATION-REPORT.md` §K.2, §Q.
|
|
|
|
---
|
|
|
|
# M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory
|
|
|
|
## Objective
|
|
|
|
Align inherited AI-DnD memory/context behavior with the final authority and history model.
|
|
|
|
## Scope
|
|
|
|
- preserve recent active-lineage history selection,
|
|
- ensure summaries are anchored to source turn ranges/lineage rather than positional list assumptions,
|
|
- preserve branch-scoped memory isolation,
|
|
- classify memory authority (accepted vs heuristic/inferred),
|
|
- maintain explicit token budgets,
|
|
- preserve/extend prompt-context inspection,
|
|
- ensure memory/summary failures do not corrupt accepted story state,
|
|
- test memory behavior after Undo/Redo/checkpoint/divergence,
|
|
- validate realistic-context prompt construction.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- imported document knowledge belongs in M7,
|
|
- no remote embeddings/vector DB.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- F01-F08,
|
|
- E lineage safety,
|
|
- no abandoned-future term appears in active prompt after divergence,
|
|
- negative control proves eligible memory returns when the relevant lineage is active,
|
|
- summary lineage survives head movement correctly,
|
|
- prompt inspector identifies included memories/summaries and token costs.
|
|
|
|
## Definition of Done
|
|
|
|
Long-running story context is lineage-safe, authority-aware, local, inspectable, and bounded.
|
|
|
|
## Note from M2 — background memory failure must be observable
|
|
|
|
M2 shipped with the memory bank entirely dead, and the full suite stayed green.
|
|
Summaries and embeddings raised `AttributeError` inside a fire-and-forget task:
|
|
no user-visible error, no log a player would read, and no failing test, because
|
|
every memory test stubs the provider factories out.
|
|
|
|
M6 therefore additionally requires:
|
|
|
|
- **memory/summarization background failures must be observable** — a
|
|
fire-and-forget task that dies must leave a record a user or maintainer can
|
|
actually find, rather than being swallowed;
|
|
- **tests must exercise at least one real provider-construction/wiring path**,
|
|
not only mocked factories, so that a moved or removed setting surfaces as a
|
|
test failure (`TECHNICAL-DESIGN.md` §18.1);
|
|
- **derived-memory failure must not corrupt accepted story state.** Already in
|
|
the scope list above; M2's evidence is why it stays there. A dead memory bank
|
|
degraded the storyteller quietly and left the transcript correct, which is the
|
|
right failure direction — but it must also be a *visible* one.
|
|
|
|
Evidence: `planning/reports/M2-IMPLEMENTATION-REPORT.md` §A.1, §9.1.
|
|
|
|
---
|
|
|
|
# M7 — First-Class Imported Knowledge Library
|
|
|
|
## Objective
|
|
|
|
Implement the separate local knowledge subsystem required by the specification rather than overloading AI-DnD Story Cards.
|
|
|
|
## Scope
|
|
|
|
- `.txt` and `.md` import,
|
|
- source validation and local copy/storage,
|
|
- Canon / Reference / Inspiration classification,
|
|
- enable/disable/delete,
|
|
- content hash and provenance,
|
|
- heading/paragraph-aware chunking,
|
|
- SQLite FTS5 lexical index,
|
|
- local Ollama embeddings where semantic retrieval is enabled,
|
|
- hybrid retrieval/reranking,
|
|
- authority-aware context insertion,
|
|
- campaign scoping,
|
|
- source/chunk inspector,
|
|
- prompt retrieval provenance,
|
|
- export/import preservation,
|
|
- no automatic URL/image fetch,
|
|
- prompt-injection framing as untrusted data.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- PDF/DOCX/EPUB import,
|
|
- remote URL ingestion,
|
|
- remote vector stores,
|
|
- general plugin/tool framework.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- G01-G10,
|
|
- C05 canon beats reference/inspiration,
|
|
- hidden/abandoned story facts cannot leak through lineage-derived knowledge,
|
|
- H06-H09 as applicable to rendering/import/archive handling,
|
|
- I05 knowledge provenance survives export/import.
|
|
|
|
## Definition of Done
|
|
|
|
A campaign can import local Canon/Reference/Inspiration files, retrieve them locally with provenance, and maintain authority boundaries.
|
|
|
|
---
|
|
|
|
# M8 — Browser UX Completion for v1 Story Operations
|
|
|
|
## Objective
|
|
|
|
Turn the adapted AI-DnD interface into the focused interactive-story workspace defined by `BROWSER-UX-SPEC.md`.
|
|
|
|
## Scope
|
|
|
|
- campaign library/setup streamlined for non-RPG stories,
|
|
- story transcript and streaming polish,
|
|
- one primary natural-language input flow plus story-direction affordance,
|
|
- Undo/Redo/Retry/alternate-take controls,
|
|
- Save Points UI,
|
|
- edit user input/narrator output with safe lineage semantics,
|
|
- current scene/state side panel,
|
|
- knowledge panel,
|
|
- context inspector,
|
|
- local Ollama status/model selection,
|
|
- no cloud-provider UI,
|
|
- reserve microphone/STT affordance without implementing STT.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
- no full branch-management UI,
|
|
- no image/video/TTS/STT implementation,
|
|
- no mobile-native app.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- B01-B04,
|
|
- D user-facing behavior,
|
|
- relevant UX acceptance checks,
|
|
- browser state remains consistent after restart/Undo/Redo/checkpoint/failed generation.
|
|
|
|
## Definition of Done
|
|
|
|
Normal story creation and play feels like a focused local storyteller rather than an RPG or developer console.
|
|
|
|
---
|
|
|
|
# M9 — Export, Backup, Recovery, and Migration Hardening
|
|
|
|
## Objective
|
|
|
|
Make campaigns portable and recoverable without losing lineage, state, knowledge, or the active head.
|
|
|
|
## Scope
|
|
|
|
- finalize documented campaign bundle format/versioning,
|
|
- preserve active branch/head, retained history, retries, checkpoints,
|
|
- state events/snapshots,
|
|
- prompt/retrieval provenance as selected,
|
|
- knowledge source metadata/content/index rebuild information,
|
|
- scene/media metadata placeholders,
|
|
- safe SQLite backup behavior,
|
|
- import validation,
|
|
- backward-compatible migration rules,
|
|
- corruption/failure handling where practical.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- I01-I06,
|
|
- L01-L04,
|
|
- undone-head round trip,
|
|
- branched campaign round trip,
|
|
- checkpoint round trip,
|
|
- knowledge provenance round trip,
|
|
- derived indexes can be rebuilt.
|
|
|
|
## Definition of Done
|
|
|
|
A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact.
|
|
|
|
---
|
|
|
|
# M10 — Future Media Extension Hooks Only
|
|
|
|
## Objective
|
|
|
|
Preserve the approved future media interfaces without adding a media-generation dependency to v1.
|
|
|
|
## Scope
|
|
|
|
- scene snapshots/packets suitable for future providers,
|
|
- optional visual character/location/item descriptors,
|
|
- provider-neutral media request/job/asset types or reserved schema as justified,
|
|
- branch/lineage association for scenes/assets,
|
|
- local-only endpoint contract for future providers,
|
|
- STT contract: local transcription -> editable draft -> normal submission,
|
|
- no core story-engine dependency on media availability.
|
|
|
|
## Explicit Non-Scope
|
|
|
|
Do not implement:
|
|
|
|
- image generation,
|
|
- video generation,
|
|
- TTS,
|
|
- STT,
|
|
- ambience/audio generation,
|
|
- ComfyUI/FLUX integration.
|
|
|
|
## Tests / Acceptance
|
|
|
|
- K01-K04 architecture/media-readiness tests,
|
|
- no v1 story flow requires a media service,
|
|
- scene data follows active lineage after Undo/restore/divergence.
|
|
|
|
## Definition of Done
|
|
|
|
Future media providers can be added through defined local interfaces without redesigning core story authority/history.
|
|
|
|
---
|
|
|
|
# M11 — v1 Security, Long-Run, and Release Validation
|
|
|
|
## Objective
|
|
|
|
Validate the full product against the release contract after all functional milestones are integrated.
|
|
|
|
## Scope
|
|
|
|
- full outbound-network blocked run,
|
|
- dependency/runtime audit,
|
|
- restrictive browser/CORS/CSP checks,
|
|
- failed-model-call recovery,
|
|
- long-running 100-turn campaign,
|
|
- repeated Undo/Redo/Retry/checkpoint cycles,
|
|
- branch/memory/summary leakage checks,
|
|
- realistic-context state extraction across selected recommended models,
|
|
- fantasy and science-fiction fixtures,
|
|
- export/import/recovery tests,
|
|
- migration tests,
|
|
- documentation and packaging.
|
|
|
|
## Tests / Acceptance
|
|
|
|
Release gate in `V1-ACCEPTANCE-TESTS.md`:
|
|
|
|
- all REQUIRED FOR V1 tests pass,
|
|
- approved exceptions documented in ADRs,
|
|
- 100-turn test passes,
|
|
- offline operation passes,
|
|
- branch/memory isolation passes,
|
|
- export/import recovery passes,
|
|
- both genre fixtures pass.
|
|
|
|
## Definition of Done
|
|
|
|
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
|
|
|
|
---
|
|
|
|
## 4. Milestone Dependency Summary
|
|
|
|
```text
|
|
M1 Production/offline foundation
|
|
-> M2 Local-only surface reduction
|
|
-> M3 Non-destructive history + export head
|
|
-> M4 Save Points
|
|
-> M5 Narrative state
|
|
-> M6 Context/memory
|
|
-> M7 Imported knowledge
|
|
-> M8 Browser UX completion
|
|
-> M9 Export/recovery hardening
|
|
-> M10 Future-media hooks
|
|
-> M11 v1 validation/release
|
|
```
|
|
|
|
Some implementation work may overlap internally, but milestone acceptance should remain sequential so architectural regressions are discovered early.
|
|
|
|
## 4. Prompting Rule
|
|
|
|
When implementation begins, prepare **one Codex prompt per milestone**.
|
|
|
|
Do not hand the entire plan to Codex as one production task.
|
|
|
|
Each prompt should contain only:
|
|
|
|
- milestone objective,
|
|
- relevant source documents,
|
|
- scope/non-scope,
|
|
- specific acceptance tests,
|
|
- stop condition and required report/diff.
|
|
|
|
No implementation prompt is included in the current planning revision.
|