Files
interactive-story/planning/TECHNICAL-DESIGN.md
T

23 KiB

Adventure Storyteller — Technical Design

Status: v1.0 — architecture selected after Phase 0B
Production base: AI-DnD d72f7c1bda0f34fccd84afb7a25c34eb01c901de

1. Design Objective

Build a local-first, browser-based interactive storytelling application in which:

  • Ollama provides inference on user-controlled local infrastructure, either same-host or on an explicitly approved trusted-LAN machine,
  • the application owns authoritative story state,
  • complete story history is retained,
  • user-facing Undo/Redo/Retry/Save Point behavior is simple,
  • internal history is non-destructive and lineage-aware,
  • long-running context is reconstructed from state, summaries, retrieval, recent turns, and local knowledge,
  • imported knowledge remains local and authority-classified,
  • prompt/context provenance is inspectable,
  • future image/video/audio/TTS/STT providers can be added without coupling them to the story engine.

The application is an interactive-story system, not a D&D rules engine.

2. Production Base Decision

AI-DnD is the production fork/base.

Retain its high-value foundation:

  • React/Vite browser application,
  • FastAPI backend,
  • SQLite persistence/migrations,
  • story-tree and lineage machinery,
  • alternate takes,
  • branch switching,
  • per-node state snapshot pattern,
  • local Ollama/OpenAI-compatible integration where appropriate,
  • branch-scoped memory concepts,
  • prompt/context snapshots and Insights concepts,
  • SSE streaming,
  • export/import framework,
  • automated test foundation.

Do not inherit candidate behavior merely because it exists upstream. The product specification and detailed behavior documents remain authoritative.

3. Reference Projects and Their Role

ai-adventure

Primary implementation reference for:

  • explicit typed state events,
  • model-proposes/application-validates discipline,
  • atomic commit behavior,
  • head movement instead of destructive Undo,
  • named checkpoints,
  • replay/state reconstruction,
  • narrow local-only endpoint handling,
  • deterministic lexical lore concepts.

It is not the written specification and does not override this design.

Open Dungeon

Reference for:

  • focused story-reading UX,
  • inline generated media presentation,
  • character visual continuity,
  • local media-service boundaries.

It is not a production-fork candidate.

Other references

  • Chronicler: memory authority/trust tiers.
  • Interactive Fiction Framework: canon/application-owned-state concepts.
  • Gamentic: provider-neutral asynchronous media boundary.

4. Selected High-Level Architecture

                   Local Browser
                        |
                        v
              +-------------------+
              | React/Vite UI     |
              +---------+---------+
                        |
                        v
              +-------------------+
              | FastAPI Story     |
              | Director/API      |
              +---+-----------+---+
                  |           |
                  |           +-----------------------+
                  v                                   v
       +----------------------+             +----------------------+
       | SQLite Authoritative |             | Context / Retrieval  |
       | Story Store          |             | local only           |
       +----------+-----------+             +----------+-----------+
                  |                                    |
                  |                         +----------+-----------+
                  |                         | FTS + local semantic |
                  |                         | retrieval            |
                  |                         +----------+-----------+
                  |                                    |
                  +--------------------+---------------+
                                       |
                                       v
                               +----------------------+
                               | Ollama               |
                               | localhost by default |
                               | or approved LAN host |
                               +----------------------+

Future optional extension:

Accepted Story / Scene State
          |
          v
  +------------------+
  | Media Coordinator|
  +--+---+---+---+---+
     |   |   |   |   |
     v   v   v   v   v
   Image Video Audio TTS STT
   local providers only by default

5. Local-Only Runtime Boundary

Local-only means user-controlled local infrastructure with no required Internet/cloud dependency; it does not require every process to share one host.

The architecture has two boundaries, and they are not the same boundary. The storyteller is loopback-only, always. Inference may be same-host or on a specifically configured trusted-LAN machine. Wording that describes Ollama as simply "loopback/local" collapses the two and understates the intended deployment:

Browser ──loopback──> storyteller (FastAPI + SPA), bound to 127.0.0.1
                          │
                          ├── same-host Ollama on 127.0.0.1:11434        (default)
                          │
                          └── OR an explicitly configured trusted-LAN Ollama
                              on another user-controlled machine
                              http://<host>:11434/v1  or  https://<host>/v1

Read that as three separate rules:

  1. The storyteller's own listener is loopback, in every run path. Dev server, production server, and container alike. Nothing about the inference choice changes it. Where a container must listen on 0.0.0.0 because a published port cannot reach anything else, the port is published to the host's loopback only.
  2. The inference endpoint is an outbound connection, chosen by the user. Same-host loopback is the default. A trusted-LAN host is a first-class, supported v1 configuration — not a workaround and not a development-only convenience.
  3. The two are independent. Reaching a LAN Ollama never requires, and must never cause, LAN exposure of the storyteller UI/API. There is no supported v1 configuration in which the storyteller itself is reachable from the LAN.

Transport to a trusted-LAN endpoint

A LAN inference host is often reached over HTTPS with a certificate issued by a private or local CA, and may offer no cleartext port at all. This is ordinary for a self-hosted server, so v1 must handle it rather than assume the same plain HTTP that same-host loopback uses:

  • outbound HTTPS verifies against the operating system's trusted CA store in addition to any bundled certificate list, so a CA the user installed on their own machine is honoured here as it is by curl and their browser;
  • certificate and hostname verification stay fully enabled;
  • there is no "ignore TLS errors" option anywhere — not in the UI, not in configuration, not as an environment variable;
  • the endpoint field therefore accepts https:// on any port.

Prompts, story text, retrieved knowledge and embedding inputs all travel to whichever inference host is configured, which is why it must be one the user controls on a network they trust — and why the endpoint is always explicitly configured, never discovered.

See ADR 002 (Transport for a Trusted-LAN Endpoint) and, for the demonstrated deployment, planning/reports/M1-IMPLEMENTATION-REPORT.md §F.

Allowed future local paths may also include explicitly configured local media services.

Production defaults must not require:

  • cloud model providers,
  • hosted authentication,
  • analytics/telemetry,
  • remote database services,
  • runtime CDNs/fonts/assets,
  • automatic web retrieval,
  • external embeddings/vector stores,
  • general plugins/MCP/shell execution.

5.1 Known AI-DnD hardening work

Phase 0B identified concrete inherited violations, and M1 added a fifth. Items 1, 2 and 5 are resolved; items 3 and 4 remain open and belong to M2.

  1. tiktoken attempts to download the cl100k_base encoding on first use. Done in M1. The encoding table is vendored in the tree and loaded directly, with its SHA-256 verified against the digest tiktoken pins, so no code path in the tokenizer can reach the network.
  2. the SPA requests Google Fonts at runtime. Done in M1. All three families are self-hosted, and the CSP names no remote origin at all.
  3. hosted/multi-user/auth/demo/analytics/Postgres/cloud-provider/QuickJS paths are unnecessary.
    • remove them rather than merely hide them where practical.
    • Open — M2. M1 removed nothing, so this surface is unchanged from the fork point.
  4. endpoint validation must reflect this product's threat model.
    • same-host loopback Ollama is the default; an explicitly configured trusted-LAN Ollama endpoint is supported; arbitrary public/Internet model endpoints must be rejected or kept outside normal v1 configuration.
    • inference endpoint configuration must not change the storyteller's own loopback bind behavior.
    • Open — M2. The trusted-LAN path itself works as of M1; what remains is deciding and enforcing which endpoints normal v1 configuration may name.
  5. outbound TLS verified only against a bundled public-CA list, so a LAN host with a privately issued certificate was refused. Found and fixed in M1. Not visible to Phase 0B: every run up to that point used plain HTTP over loopback, where certificate verification never happens. See Transport to a trusted-LAN endpoint above.

6. Browser UI Boundary

The browser remains a presentation/control layer, not the owner of story authority.

Primary areas:

  • campaign library/setup,
  • active story transcript,
  • input composer,
  • Undo / Redo / Retry,
  • alternate-take selector,
  • Save Points,
  • current story state inspector,
  • imported-knowledge manager,
  • prompt/context inspector,
  • local-model status/settings,
  • export/import,
  • future media gallery/actions.

Normal storytelling should not expose branch IDs, database rows, embeddings, or event logs unless the user opens advanced diagnostics.

7. Story Director Boundary

The FastAPI service owns the turn lifecycle:

  1. resolve campaign, active branch, and active head,
  2. load authoritative state at that head,
  3. assemble bounded lineage-safe context,
  4. retain prompt/retrieval provenance,
  5. invoke local Ollama narrator,
  6. stream provisional narration,
  7. obtain/parse a structured state proposal,
  8. validate the proposal,
  9. atomically accept the turn plus validated state consequences,
  10. update derived memory/summary/index data without allowing derived failures to corrupt the accepted turn,
  11. expose the new current head to the browser.

A failed generation must not partially advance authoritative story state.

8. Non-Destructive History Model

8.1 Core rule

Undo moves the active head. It does not delete accepted history.

Phase 0B demonstrated that AI-DnD already contains the architectural chokepoints needed for this approach:

  • stored head depth/position,
  • lineage reads through a common path abstraction,
  • branch-at-depth behavior.

The disposable spike is evidence, not production code to merge blindly.

8.2 Active head versus retained tip

The design distinguishes:

  • retained tip: newest retained turn on a continuation,
  • active head: the story position from which the user is currently reading/continuing.

After Undo, the active head may sit behind a retained tip.

Redo moves the head forward along the previous active continuation while no divergent write has occurred.

8.3 Divergence after moving backward

If the user writes/retries/edits from a head behind the retained tip:

  • the new continuation forks on first write,
  • the previous future remains retained,
  • ordinary Redo into that old future is invalidated,
  • the displaced future is marked abandoned/disposable,
  • lineage-sensitive state, summary, and memory selection follows only the new active path.

8.4 Retry

Retry preserves alternate narrator takes for the same user input.

Production implementation must ensure retry/add-take while behind the current tip uses the same safe fork/head semantics as other writes.

8.5 Checkpoints / Save Points

A named checkpoint is a durable pointer to a recoverable story position.

Conceptually:

campaign_id
branch_id
turn_id or equivalent head coordinate
name
notes
created_at

Restoring a checkpoint moves the active head to that position. The existing later future remains retained. A new branch is created only if/when the user creates a different continuation.

8.6 Abandoned history

No automatic cleanup policy is required in v1.

Abandoned history must:

  • remain retained,
  • be marked disposable/inactive through implementation-appropriate metadata,
  • stop influencing current state/context/memory/summary,
  • remain available for future recovery/cleanup features.

9. Export / Import and Head Position

AI-DnD's current export carries branch information but reconstructs the imported head at the branch tip.

That is invalid after non-destructive Undo because an exported campaign can intentionally have:

active head < retained tip

The production export format must preserve:

  • active branch,
  • active head turn/depth/coordinate,
  • retained alternate/disposable history,
  • checkpoints,
  • state/history/provenance required for recovery.

For compatibility with earlier bundles, import may fall back to the retained tip only when no explicit active-head field exists.

Export/import regression tests must include an undone campaign and verify the imported story reopens at the exact exported head rather than silently redoing later turns.

10. Authoritative Narrative State

10.1 Do not retain the RPG state protocol as the product model

AI-DnD's world-state machinery is useful evidence that state snapshots and rollback are structurally separable from RPG presentation, but the production state model must be genre-neutral.

Core concepts include:

  • entities,
  • facts,
  • relationships,
  • locations,
  • possessions,
  • conditions,
  • organizations,
  • story threads,
  • scene state,
  • chronology where needed.

10.2 Explicit typed state events

Use the ADR 010 model:

model proposes explicit typed operation
        -> schema validation
        -> semantic/referential validation
        -> accepted event(s)
        -> state snapshot/cache

Prefer explicit absolute semantics for mutable values.

Examples:

set_current_location
set_entity_status
set_possession
add_fact
invalidate_fact
add_relationship
end_relationship
open_story_thread
resolve_story_thread
set_scene

Avoid one generic relative-delta protocol whose numeric meaning depends primarily on prompt compliance.

10.3 Validation limitations

Typed events remove the delta/absolute ambiguity but do not guarantee semantic truth.

Validation should include deterministic checks where possible:

  • event type allowlist,
  • schema/type validation,
  • entity/reference existence,
  • impossible transitions where explicitly modeled,
  • authority constraints,
  • conflict handling,
  • transaction integrity.

The accepted transcript remains available even if derived state extraction must be retried/repaired according to the final turn-acceptance workflow.

10.4 Hybrid storage

Selected direction:

validated state events + efficient current/historical snapshots/cache

Events provide audit/reconstruction value. Snapshots/cache make normal reads, Undo/Redo, and context construction fast.

11. Context and Memory

Retain AI-DnD's useful lineage-aware memory foundation, but align it with the product authority model.

Narrator context is assembled in explicit layers:

Narrator/system rules
Campaign profile
Global/explicit canon
Current authoritative state
Lineage-safe summary
Relevant older story memories
Relevant imported knowledge
Recent active-lineage turns
Current user input

Requirements:

  • no abandoned future may appear in active recent history,
  • no memory derived solely from an abandoned future may be retrieved,
  • summaries are anchored to source lineage/turn ranges,
  • derived memory/summary never becomes more authoritative than accepted state/canon,
  • prompt snapshot records what was actually supplied,
  • token budgets remain explicit and inspectable.

Phase 0B verified AI-DnD branch-scoped memory isolation against real local embeddings with a negative control. Preserve that property through the history rewrite.

12. Realistic-Context Model Testing

The Phase 0B referee failure appeared under full application context even though the same model followed the state protocol correctly in an isolated probe.

Therefore structured-output/state tests must include:

  • realistic narrator/context length,
  • representative state complexity,
  • actual local models likely to be used,
  • repeated runs rather than one clean prompt,
  • malformed/incorrect semantic proposals,
  • validation and recovery behavior.

Model capability recommendations are deferred until these measurements exist; this does not block the architecture.

13. Imported Knowledge Subsystem

Do not turn AI-DnD Story Cards into the production imported-knowledge store.

Story Cards may remain a useful reference or authored-rule mechanism, but the imported-knowledge requirements need a separate first-class subsystem with:

  • source records,
  • .txt / .md import,
  • Canon / Reference / Inspiration classification,
  • enable/disable/delete,
  • source/version/hash provenance,
  • chunk records,
  • campaign scoping,
  • local lexical index (prefer SQLite FTS5),
  • local Ollama embeddings/semantic index where enabled,
  • authority-aware hybrid retrieval,
  • retrieval provenance,
  • export/import support,
  • no automatic URL/image fetching,
  • imported content treated as data, never executable instructions.

If a future knowledge item is derived from story history rather than imported as global campaign material, it must carry lineage/source-turn information sufficient to avoid abandoned-path leakage.

14. Prompt and Provenance Inspection

Preserve and extend AI-DnD's Insights/context-snapshot capability.

For each narrator turn the system should be able to explain:

  • narrator/system rules used,
  • campaign/canon context,
  • current authoritative state included,
  • summary included,
  • story memories retrieved,
  • knowledge chunks retrieved,
  • recent history included,
  • user input,
  • model/settings,
  • state proposal,
  • validation result,
  • accepted events,
  • source IDs/turn ranges where applicable.

15. Scene and Future Media Boundary

v1 does not require media generation.

It does require preserving scene/entity information so future providers do not have to infer continuity from the entire raw transcript.

Persist or derive a scene snapshot containing relevant fields such as:

  • location,
  • participants,
  • significant objects,
  • current actions,
  • time/lighting/environment,
  • mood,
  • visual character/location profiles,
  • continuity constraints,
  • source turn range and lineage.

Future media coordinator consumes a normalized scene packet and records local asset provenance.

The story engine must remain fully functional with media disabled.

STT specifically follows:

microphone/audio -> local STT -> editable draft -> normal user submission

STT never bypasses the ordinary authoritative story commit path.

16. Database Direction

SQLite remains the selected v1 authoritative store.

Reasons:

  • already present in the selected base,
  • local and single-user friendly,
  • transactional,
  • portable,
  • supports FTS5,
  • compatible with backup/export tooling,
  • no external service required.

Remove Postgres/Neon support from the production fork unless a later explicit requirement reverses this decision.

The exact physical schema may evolve through migrations; the conceptual model is in DATA-MODEL.md.

17. Transaction Boundaries

Where practical, one accepted turn should atomically establish:

  • accepted user input/narration relationship,
  • turn/lineage identity,
  • active-head advancement,
  • validated authoritative state events,
  • resulting state snapshot/cache,
  • core prompt/model provenance needed for recovery/audit.

Derived work such as embeddings, memory extraction, summary generation, and future media jobs may occur separately, but failure must not corrupt the authoritative commit.

18. Testing Strategy

Use AI-DnD's inherited tests as a foundation, then rewrite/add tests around product semantics.

Required categories:

  • offline startup/story use,
  • no runtime remote assets/tokenizer fetch,
  • same-host and trusted-LAN Ollama endpoint handling,
  • turn persistence/restart,
  • failed generation atomicity,
  • non-destructive Undo/Redo,
  • divergence after Undo,
  • retry/take retention,
  • named checkpoint restore,
  • branch/lineage state reconstruction,
  • abandoned-history memory/summary isolation,
  • active-head export/import round trip,
  • generic narrative state event validation,
  • realistic-context structured state extraction,
  • imported knowledge authority/provenance/isolation,
  • prompt/context inspection,
  • fantasy + science-fiction genre neutrality,
  • 100-turn/long-run acceptance,
  • future-media schema compatibility.

Acceptance-test IDs in V1-ACCEPTANCE-TESTS.md are the black-box release contract.

19. Removal / Migration Strategy From Upstream

Production migration should be incremental and test-gated rather than a broad rewrite.

Remove or replace in controlled milestones:

  1. runtime external dependency leaks,
  2. hosted/multi-user/auth/demo/analytics/cloud/Postgres/QuickJS surfaces,
  3. destructive Undo/no-Redo behavior,
  4. RPG-specific state/referee protocol and UI assumptions,
  5. Story Card assumptions where they conflict with the new knowledge subsystem.

Preserve upstream provenance and license notices.

Do not mechanically merge the ai-adventure or Open Dungeon repositories into the fork.

20. Deferred Questions That Do Not Block v1 Architecture

The following are implementation/release measurements, not unresolved foundational choices:

  • which narrator/state models should be recommended to users,
  • exact embedding model recommendation,
  • performance of multi-hour stories,
  • concurrency beyond the single-user turn lock,
  • which future local image/video/TTS/STT provider is selected,
  • abandoned-history cleanup policy/UI after v1.

21. Technical Design v1.0 Exit Status

Phase 0 has resolved the foundational choices required for v1.0:

  • base repository selected,
  • browser/backend stack selected,
  • SQLite selected,
  • non-destructive history/head model demonstrated,
  • typed narrative-state-event direction selected,
  • Ollama validated on local infrastructure, including the required trusted-LAN deployment mode,
  • memory lineage behavior validated,
  • imported-knowledge architecture selected,
  • local-only hardening scope identified,
  • export/import active-head defect understood,
  • future media boundary retained,
  • production milestone sequence defined in BUILD-MILESTONES.md.

This technical design is therefore the implementation baseline unless revised by a later ADR.