Files
interactive-story/planning/TECHNICAL-DESIGN.md
T

654 lines
23 KiB
Markdown

# Adventure Storyteller — Technical Design
**Status:** v1.0 — architecture selected after Phase 0B
**Production base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
## 1. Design Objective
Build a local-first, browser-based interactive storytelling application in which:
- Ollama provides inference on user-controlled local infrastructure, either same-host or on an explicitly approved trusted-LAN machine,
- the application owns authoritative story state,
- complete story history is retained,
- user-facing Undo/Redo/Retry/Save Point behavior is simple,
- internal history is non-destructive and lineage-aware,
- long-running context is reconstructed from state, summaries, retrieval, recent turns, and local knowledge,
- imported knowledge remains local and authority-classified,
- prompt/context provenance is inspectable,
- future image/video/audio/TTS/STT providers can be added without coupling them to the story engine.
The application is an interactive-story system, not a D&D rules engine.
## 2. Production Base Decision
AI-DnD is the production fork/base.
Retain its high-value foundation:
- React/Vite browser application,
- FastAPI backend,
- SQLite persistence/migrations,
- story-tree and lineage machinery,
- alternate takes,
- branch switching,
- per-node state snapshot pattern,
- local Ollama/OpenAI-compatible integration where appropriate,
- branch-scoped memory concepts,
- prompt/context snapshots and Insights concepts,
- SSE streaming,
- export/import framework,
- automated test foundation.
Do not inherit candidate behavior merely because it exists upstream. The product specification and detailed behavior documents remain authoritative.
## 3. Reference Projects and Their Role
### ai-adventure
Primary implementation reference for:
- explicit typed state events,
- model-proposes/application-validates discipline,
- atomic commit behavior,
- head movement instead of destructive Undo,
- named checkpoints,
- replay/state reconstruction,
- narrow local-only endpoint handling,
- deterministic lexical lore concepts.
It is **not** the written specification and does not override this design.
### Open Dungeon
Reference for:
- focused story-reading UX,
- inline generated media presentation,
- character visual continuity,
- local media-service boundaries.
It is not a production-fork candidate.
### Other references
- Chronicler: memory authority/trust tiers.
- Interactive Fiction Framework: canon/application-owned-state concepts.
- Gamentic: provider-neutral asynchronous media boundary.
## 4. Selected High-Level Architecture
```text
Local Browser
|
v
+-------------------+
| React/Vite UI |
+---------+---------+
|
v
+-------------------+
| FastAPI Story |
| Director/API |
+---+-----------+---+
| |
| +-----------------------+
v v
+----------------------+ +----------------------+
| SQLite Authoritative | | Context / Retrieval |
| Story Store | | local only |
+----------+-----------+ +----------+-----------+
| |
| +----------+-----------+
| | FTS + local semantic |
| | retrieval |
| +----------+-----------+
| |
+--------------------+---------------+
|
v
+----------------------+
| Ollama |
| localhost by default |
| or approved LAN host |
+----------------------+
```
Future optional extension:
```text
Accepted Story / Scene State
|
v
+------------------+
| Media Coordinator|
+--+---+---+---+---+
| | | | |
v v v v v
Image Video Audio TTS STT
local providers only by default
```
## 5. Local-Only Runtime Boundary
`Local-only` means user-controlled local infrastructure with no required Internet/cloud dependency; it does not require every process to share one host.
The architecture has **two boundaries, and they are not the same boundary**. The
storyteller is loopback-only, always. Inference may be same-host or on a
specifically configured trusted-LAN machine. Wording that describes Ollama as
simply "loopback/local" collapses the two and understates the intended
deployment:
```text
Browser ──loopback──> storyteller (FastAPI + SPA), bound to 127.0.0.1
│
├── same-host Ollama on 127.0.0.1:11434 (default)
│
└── OR an explicitly configured trusted-LAN Ollama
on another user-controlled machine
http://<host>:11434/v1 or https://<host>/v1
```
Read that as three separate rules:
1. **The storyteller's own listener is loopback, in every run path.** Dev
server, production server, and container alike. Nothing about the inference
choice changes it. Where a container must listen on `0.0.0.0` because a
published port cannot reach anything else, the port is published to the
host's loopback only.
2. **The inference endpoint is an outbound connection, chosen by the user.**
Same-host loopback is the default. A trusted-LAN host is a first-class,
supported v1 configuration — not a workaround and not a development-only
convenience.
3. **The two are independent.** Reaching a LAN Ollama never requires, and must
never cause, LAN exposure of the storyteller UI/API. There is no supported
v1 configuration in which the storyteller itself is reachable from the LAN.
### Transport to a trusted-LAN endpoint
A LAN inference host is often reached over **HTTPS with a certificate issued by
a private or local CA**, and may offer no cleartext port at all. This is
ordinary for a self-hosted server, so v1 must handle it rather than assume the
same plain HTTP that same-host loopback uses:
- outbound HTTPS verifies against the **operating system's trusted CA store** in
addition to any bundled certificate list, so a CA the user installed on their
own machine is honoured here as it is by `curl` and their browser;
- certificate **and hostname** verification stay fully enabled;
- there is **no "ignore TLS errors" option** anywhere — not in the UI, not in
configuration, not as an environment variable;
- the endpoint field therefore accepts `https://` on any port.
Prompts, story text, retrieved knowledge and embedding inputs all travel to
whichever inference host is configured, which is why it must be one the user
controls on a network they trust — and why the endpoint is always explicitly
configured, never discovered.
See ADR 002 (*Transport for a Trusted-LAN Endpoint*) and, for the demonstrated
deployment, `planning/reports/M1-IMPLEMENTATION-REPORT.md` §F.
Allowed future local paths may also include explicitly configured local media services.
Production defaults must not require:
- cloud model providers,
- hosted authentication,
- analytics/telemetry,
- remote database services,
- runtime CDNs/fonts/assets,
- automatic web retrieval,
- external embeddings/vector stores,
- general plugins/MCP/shell execution.
### 5.1 Known AI-DnD hardening work
Phase 0B identified concrete inherited violations, and M1 added a fifth. Items
1, 2 and 5 are **resolved**; items 3 and 4 remain open and belong to M2.
1. ~~`tiktoken` attempts to download the `cl100k_base` encoding on first use.~~
**Done in M1.** The encoding table is vendored in the tree and loaded
directly, with its SHA-256 verified against the digest `tiktoken` pins, so no
code path in the tokenizer can reach the network.
2. ~~the SPA requests Google Fonts at runtime.~~
**Done in M1.** All three families are self-hosted, and the CSP names no
remote origin at all.
3. hosted/multi-user/auth/demo/analytics/Postgres/cloud-provider/QuickJS paths are unnecessary.
- remove them rather than merely hide them where practical.
- **Open — M2.** M1 removed nothing, so this surface is unchanged from the
fork point.
4. endpoint validation must reflect this product's threat model.
- same-host loopback Ollama is the default; an explicitly configured trusted-LAN Ollama endpoint is supported; arbitrary public/Internet model endpoints must be rejected or kept outside normal v1 configuration.
- inference endpoint configuration must not change the storyteller's own loopback bind behavior.
- **Open — M2.** The trusted-LAN path itself works as of M1; what remains is
deciding and enforcing which endpoints normal v1 configuration may name.
5. ~~outbound TLS verified only against a bundled public-CA list, so a LAN host
with a privately issued certificate was refused.~~
**Found and fixed in M1.** Not visible to Phase 0B: every run up to that
point used plain HTTP over loopback, where certificate verification never
happens. See *Transport to a trusted-LAN endpoint* above.
## 6. Browser UI Boundary
The browser remains a presentation/control layer, not the owner of story authority.
Primary areas:
- campaign library/setup,
- active story transcript,
- input composer,
- Undo / Redo / Retry,
- alternate-take selector,
- Save Points,
- current story state inspector,
- imported-knowledge manager,
- prompt/context inspector,
- local-model status/settings,
- export/import,
- future media gallery/actions.
Normal storytelling should not expose branch IDs, database rows, embeddings, or event logs unless the user opens advanced diagnostics.
## 7. Story Director Boundary
The FastAPI service owns the turn lifecycle:
1. resolve campaign, active branch, and active head,
2. load authoritative state at that head,
3. assemble bounded lineage-safe context,
4. retain prompt/retrieval provenance,
5. invoke local Ollama narrator,
6. stream provisional narration,
7. obtain/parse a structured state proposal,
8. validate the proposal,
9. atomically accept the turn plus validated state consequences,
10. update derived memory/summary/index data without allowing derived failures to corrupt the accepted turn,
11. expose the new current head to the browser.
A failed generation must not partially advance authoritative story state.
## 8. Non-Destructive History Model
### 8.1 Core rule
Undo moves the active head. It does not delete accepted history.
Phase 0B demonstrated that AI-DnD already contains the architectural chokepoints needed for this approach:
- stored head depth/position,
- lineage reads through a common path abstraction,
- branch-at-depth behavior.
The disposable spike is evidence, not production code to merge blindly.
### 8.2 Active head versus retained tip
The design distinguishes:
- **retained tip:** newest retained turn on a continuation,
- **active head:** the story position from which the user is currently reading/continuing.
After Undo, the active head may sit behind a retained tip.
Redo moves the head forward along the previous active continuation while no divergent write has occurred.
### 8.3 Divergence after moving backward
If the user writes/retries/edits from a head behind the retained tip:
- the new continuation forks on first write,
- the previous future remains retained,
- ordinary Redo into that old future is invalidated,
- the displaced future is marked abandoned/disposable,
- lineage-sensitive state, summary, and memory selection follows only the new active path.
### 8.4 Retry
Retry preserves alternate narrator takes for the same user input.
Production implementation must ensure retry/add-take while behind the current tip uses the same safe fork/head semantics as other writes.
### 8.5 Checkpoints / Save Points
A named checkpoint is a durable pointer to a recoverable story position.
Conceptually:
```text
campaign_id
branch_id
turn_id or equivalent head coordinate
name
notes
created_at
```
Restoring a checkpoint moves the active head to that position. The existing later future remains retained. A new branch is created only if/when the user creates a different continuation.
### 8.6 Abandoned history
No automatic cleanup policy is required in v1.
Abandoned history must:
- remain retained,
- be marked disposable/inactive through implementation-appropriate metadata,
- stop influencing current state/context/memory/summary,
- remain available for future recovery/cleanup features.
## 9. Export / Import and Head Position
AI-DnD's current export carries branch information but reconstructs the imported head at the branch tip.
That is invalid after non-destructive Undo because an exported campaign can intentionally have:
```text
active head < retained tip
```
The production export format must preserve:
- active branch,
- active head turn/depth/coordinate,
- retained alternate/disposable history,
- checkpoints,
- state/history/provenance required for recovery.
For compatibility with earlier bundles, import may fall back to the retained tip only when no explicit active-head field exists.
Export/import regression tests must include an undone campaign and verify the imported story reopens at the exact exported head rather than silently redoing later turns.
## 10. Authoritative Narrative State
### 10.1 Do not retain the RPG state protocol as the product model
AI-DnD's world-state machinery is useful evidence that state snapshots and rollback are structurally separable from RPG presentation, but the production state model must be genre-neutral.
Core concepts include:
- entities,
- facts,
- relationships,
- locations,
- possessions,
- conditions,
- organizations,
- story threads,
- scene state,
- chronology where needed.
### 10.2 Explicit typed state events
Use the ADR 010 model:
```text
model proposes explicit typed operation
-> schema validation
-> semantic/referential validation
-> accepted event(s)
-> state snapshot/cache
```
Prefer explicit absolute semantics for mutable values.
Examples:
```text
set_current_location
set_entity_status
set_possession
add_fact
invalidate_fact
add_relationship
end_relationship
open_story_thread
resolve_story_thread
set_scene
```
Avoid one generic relative-delta protocol whose numeric meaning depends primarily on prompt compliance.
### 10.3 Validation limitations
Typed events remove the delta/absolute ambiguity but do not guarantee semantic truth.
Validation should include deterministic checks where possible:
- event type allowlist,
- schema/type validation,
- entity/reference existence,
- impossible transitions where explicitly modeled,
- authority constraints,
- conflict handling,
- transaction integrity.
The accepted transcript remains available even if derived state extraction must be retried/repaired according to the final turn-acceptance workflow.
### 10.4 Hybrid storage
Selected direction:
> **validated state events + efficient current/historical snapshots/cache**
Events provide audit/reconstruction value. Snapshots/cache make normal reads, Undo/Redo, and context construction fast.
## 11. Context and Memory
Retain AI-DnD's useful lineage-aware memory foundation, but align it with the product authority model.
Narrator context is assembled in explicit layers:
```text
Narrator/system rules
Campaign profile
Global/explicit canon
Current authoritative state
Lineage-safe summary
Relevant older story memories
Relevant imported knowledge
Recent active-lineage turns
Current user input
```
Requirements:
- no abandoned future may appear in active recent history,
- no memory derived solely from an abandoned future may be retrieved,
- summaries are anchored to source lineage/turn ranges,
- derived memory/summary never becomes more authoritative than accepted state/canon,
- prompt snapshot records what was actually supplied,
- token budgets remain explicit and inspectable.
Phase 0B verified AI-DnD branch-scoped memory isolation against real local embeddings with a negative control. Preserve that property through the history rewrite.
## 12. Realistic-Context Model Testing
The Phase 0B referee failure appeared under full application context even though the same model followed the state protocol correctly in an isolated probe.
Therefore structured-output/state tests must include:
- realistic narrator/context length,
- representative state complexity,
- actual local models likely to be used,
- repeated runs rather than one clean prompt,
- malformed/incorrect semantic proposals,
- validation and recovery behavior.
Model capability recommendations are deferred until these measurements exist; this does not block the architecture.
## 13. Imported Knowledge Subsystem
Do not turn AI-DnD Story Cards into the production imported-knowledge store.
Story Cards may remain a useful reference or authored-rule mechanism, but the imported-knowledge requirements need a separate first-class subsystem with:
- source records,
- `.txt` / `.md` import,
- Canon / Reference / Inspiration classification,
- enable/disable/delete,
- source/version/hash provenance,
- chunk records,
- campaign scoping,
- local lexical index (prefer SQLite FTS5),
- local Ollama embeddings/semantic index where enabled,
- authority-aware hybrid retrieval,
- retrieval provenance,
- export/import support,
- no automatic URL/image fetching,
- imported content treated as data, never executable instructions.
If a future knowledge item is derived from story history rather than imported as global campaign material, it must carry lineage/source-turn information sufficient to avoid abandoned-path leakage.
## 14. Prompt and Provenance Inspection
Preserve and extend AI-DnD's Insights/context-snapshot capability.
For each narrator turn the system should be able to explain:
- narrator/system rules used,
- campaign/canon context,
- current authoritative state included,
- summary included,
- story memories retrieved,
- knowledge chunks retrieved,
- recent history included,
- user input,
- model/settings,
- state proposal,
- validation result,
- accepted events,
- source IDs/turn ranges where applicable.
## 15. Scene and Future Media Boundary
v1 does not require media generation.
It does require preserving scene/entity information so future providers do not have to infer continuity from the entire raw transcript.
Persist or derive a scene snapshot containing relevant fields such as:
- location,
- participants,
- significant objects,
- current actions,
- time/lighting/environment,
- mood,
- visual character/location profiles,
- continuity constraints,
- source turn range and lineage.
Future media coordinator consumes a normalized scene packet and records local asset provenance.
The story engine must remain fully functional with media disabled.
STT specifically follows:
```text
microphone/audio -> local STT -> editable draft -> normal user submission
```
STT never bypasses the ordinary authoritative story commit path.
## 16. Database Direction
SQLite remains the selected v1 authoritative store.
Reasons:
- already present in the selected base,
- local and single-user friendly,
- transactional,
- portable,
- supports FTS5,
- compatible with backup/export tooling,
- no external service required.
Remove Postgres/Neon support from the production fork unless a later explicit requirement reverses this decision.
The exact physical schema may evolve through migrations; the conceptual model is in `DATA-MODEL.md`.
## 17. Transaction Boundaries
Where practical, one accepted turn should atomically establish:
- accepted user input/narration relationship,
- turn/lineage identity,
- active-head advancement,
- validated authoritative state events,
- resulting state snapshot/cache,
- core prompt/model provenance needed for recovery/audit.
Derived work such as embeddings, memory extraction, summary generation, and future media jobs may occur separately, but failure must not corrupt the authoritative commit.
## 18. Testing Strategy
Use AI-DnD's inherited tests as a foundation, then rewrite/add tests around product semantics.
Required categories:
- offline startup/story use,
- no runtime remote assets/tokenizer fetch,
- same-host and trusted-LAN Ollama endpoint handling,
- turn persistence/restart,
- failed generation atomicity,
- non-destructive Undo/Redo,
- divergence after Undo,
- retry/take retention,
- named checkpoint restore,
- branch/lineage state reconstruction,
- abandoned-history memory/summary isolation,
- active-head export/import round trip,
- generic narrative state event validation,
- realistic-context structured state extraction,
- imported knowledge authority/provenance/isolation,
- prompt/context inspection,
- fantasy + science-fiction genre neutrality,
- 100-turn/long-run acceptance,
- future-media schema compatibility.
Acceptance-test IDs in `V1-ACCEPTANCE-TESTS.md` are the black-box release contract.
## 19. Removal / Migration Strategy From Upstream
Production migration should be incremental and test-gated rather than a broad rewrite.
Remove or replace in controlled milestones:
1. runtime external dependency leaks,
2. hosted/multi-user/auth/demo/analytics/cloud/Postgres/QuickJS surfaces,
3. destructive Undo/no-Redo behavior,
4. RPG-specific state/referee protocol and UI assumptions,
5. Story Card assumptions where they conflict with the new knowledge subsystem.
Preserve upstream provenance and license notices.
Do not mechanically merge the ai-adventure or Open Dungeon repositories into the fork.
## 20. Deferred Questions That Do Not Block v1 Architecture
The following are implementation/release measurements, not unresolved foundational choices:
- which narrator/state models should be recommended to users,
- exact embedding model recommendation,
- performance of multi-hour stories,
- concurrency beyond the single-user turn lock,
- which future local image/video/TTS/STT provider is selected,
- abandoned-history cleanup policy/UI after v1.
## 21. Technical Design v1.0 Exit Status
Phase 0 has resolved the foundational choices required for v1.0:
- base repository selected,
- browser/backend stack selected,
- SQLite selected,
- non-destructive history/head model demonstrated,
- typed narrative-state-event direction selected,
- Ollama validated on local infrastructure, including the required trusted-LAN deployment mode,
- memory lineage behavior validated,
- imported-knowledge architecture selected,
- local-only hardening scope identified,
- export/import active-head defect understood,
- future media boundary retained,
- production milestone sequence defined in `BUILD-MILESTONES.md`.
This technical design is therefore the implementation baseline unless revised by a later ADR.