16 KiB
Adventure Storyteller — Technical Design
Status: Provisional v0.1
Important: This document describes the current preferred architecture. Phase 0 research is expected to confirm, revise, or replace portions of it.
1. Design Objective
Implement a local-first, browser-based interactive storytelling system in which:
- Ollama provides local AI inference,
- the application owns authoritative story state,
- complete story history is persistent,
- long-running context is reconstructed from state, summaries, retrieval, and recent turns,
- users can checkpoint, restore, retry, and branch,
- imported knowledge remains local,
- future media generation can be added without redesigning the story engine.
2. Current Provisional Architecture
Local Browser
|
v
+------------------+
| Browser UI |
+--------+---------+
|
v
+------------------+
| Story Director |
| API / Service |
+---+----------+---+
| |
+--------+ +----------------+
v v
+----------------------+ +----------------------+
| Authoritative Store | | Context / Retrieval |
| SQLite (provisional) | | local only |
+----------+-----------+ +----------+-----------+
| |
| v
| +----------------------+
| | Local embeddings / |
| | lexical retrieval |
| +----------+-----------+
| |
+-------------------+------------------+
|
v
+---------------+
| Ollama |
| localhost |
+-------+-------+
|
v
Local narrator model
Future:
Story / Scene State
|
v
+-------------------+
| Media Coordinator |
+----+---------+----+
| |
v v
Image Video
Provider Provider
3. Base Repository Strategy
Status: UNDECIDED
Phase 0 will determine whether to:
- fork Open Dungeon and add stronger state/memory/branching,
- fork AI-DnD and remove RPG/cloud complexity,
- use
CaoRuiming/ai-adventureas the core and add browser/Ollama layers, - build a thin new application using selected reusable components,
- choose another candidate discovered during research.
The chosen strategy must be justified with code-level evidence rather than README feature comparison alone.
4. Component Boundaries
4.1 Browser UI
Responsibilities:
- campaign selection,
- campaign creation/editing,
- transcript display,
- streaming narrator output,
- user input,
- branch/checkpoint navigation,
- state inspection/editing,
- library/source management,
- prompt/context inspection,
- settings,
- future media controls/gallery.
The browser UI must not directly own story authority.
4.2 Story Director
Responsibilities:
- accept user turns,
- load authoritative campaign/branch state,
- construct model context,
- invoke Ollama,
- validate and commit accepted outputs,
- trigger summarization/state extraction as needed,
- maintain branch/tree relationships,
- create scene snapshots,
- record provenance/debug metadata,
- expose state/history APIs to the browser.
4.3 Model Adapter
v1 target: Ollama.
Responsibilities:
- enumerate allowed local models,
- invoke chat/generation,
- support streaming,
- invoke local embedding model if selected,
- expose model metadata,
- reject unsupported remote/cloud providers.
Provisional endpoint default:
http://127.0.0.1:11434
4.4 Authoritative Store
Provisional choice: SQLite.
Reasons:
- local,
- transactional,
- portable,
- easy backup,
- strong fit for structured story/state data,
- can support FTS,
- no external service required.
Phase 0 must validate whether the selected fork already has a suitable schema and migration system.
4.5 Knowledge Store
Provisional options:
- SQLite FTS,
- embeddings stored in SQLite,
- local vector library,
- hybrid lexical/vector retrieval.
Remote vector databases are out of scope for v1.
4.6 Media Coordinator
Not required for v1 implementation, but interface boundaries should be reserved.
Responsibilities later:
- accept scene/character/turn-range generation requests,
- transform story state into media-generation packets,
- call pluggable local media providers,
- record asset provenance,
- attach assets to campaigns/scenes/turns.
5. Authoritative Data Model
The exact schema is provisional.
5.1 Campaign
Possible fields:
idtitleprofiletonestylenarrator_rulescreated_atupdated_atactive_branch_idmodel_config_id
5.2 Branch
Possible fields:
idcampaign_idnameroot_turn_idhead_turn_idcreated_from_branch_idcreated_at
5.3 Turn
Possible fields:
idcampaign_idbranch_idparent_turn_idsequence_hintuser_inputassistant_outputcreated_atmodel_idgeneration_configprompt_snapshot_idstate_version_idscene_snapshot_id
The graph/tree relationship should come from parentage, not merely sequential row order.
5.4 Checkpoint
Possible fields:
idcampaign_idturn_idnamenotescreated_at
Every accepted turn is implicitly recoverable even when not given a name.
5.5 Narrative Entity
A generic entity model should avoid genre-specific database design.
Possible categories:
- character,
- location,
- organization,
- item,
- vehicle,
- object,
- concept,
- other.
Possible fields:
idcampaign_idtypenamecanonical_descriptionvisual_descriptionstatusmetadata_json
Separate normalized tables may replace a generic entity table if research shows that is cleaner.
5.6 Fact
Potential representation:
- subject,
- predicate,
- object/value,
- source turn,
- canonical status,
- validity interval/version,
- confidence/review state.
The design should distinguish accepted canon from merely proposed model content.
5.7 Relationship
Potential examples:
- character-to-character,
- character-to-organization,
- entity-to-location,
- ownership,
- allegiance,
- trust/hostility,
- family/friendship.
5.8 Story Thread
Possible fields:
- title,
- description,
- status,
- opened_turn_id,
- resolved_turn_id,
- importance,
- related entities.
5.9 Summary
Potential levels:
- full campaign summary,
- arc/chapter summary,
- branch summary,
- rolling compressed memory.
Each summary should record what source turns it represents.
5.10 Scene Snapshot
Provisional fields:
idcampaign_idbranch_idsource_turn_startsource_turn_endlocation_entity_idtime_descriptionmoodparticipants_jsonenvironment_jsonvisual_notesaction_beats_jsoncontinuity_notes_json
5.11 Asset / Asset Job
May exist in schema before implementation.
Potential asset fields:
idcampaign_idscene_idtypeprovidermodelpromptsettings_jsonfile_pathcreated_atsource_turn_range
No v1 dependency should require these tables to contain data.
6. Turn Processing
Provisional turn pipeline:
1. Receive user input
2. Resolve campaign + active branch + parent turn
3. Load authoritative state
4. Retrieve recent turns
5. Retrieve relevant older story memory
6. Retrieve relevant local canon/reference/inspiration
7. Build narrator context
8. Save prompt/context provenance
9. Invoke Ollama narrator model
10. Stream response to UI
11. Validate completion
12. Extract proposed state changes
13. Validate proposed state changes
14. Commit turn + state + scene atomically
15. Update summaries/indexes when thresholds require it
16. Expose new recoverable branch head
The final implementation may combine or reorder steps depending on selected repository architecture.
7. State Extraction
The system may use a second local model call to convert narration into proposed structured changes.
Example:
{
"new_facts": [],
"changed_entities": [],
"opened_threads": [],
"resolved_threads": [],
"scene_changes": []
}
Rules:
- model-produced state changes are proposals,
- authoritative updates must pass application validation,
- invalid structured output must not corrupt the campaign,
- narration should remain preserved even if extraction must be retried or repaired.
A deterministic/non-LLM extraction layer may supplement this later.
8. Context Construction
Target conceptual structure:
Narrator/system rules
+
Campaign profile
+
Authoritative canon/world rules
+
Current narrative state
+
Campaign/arc summary
+
Relevant older memories
+
Relevant imported local material
+
Recent turns
+
Current user input
Context must be bounded by configurable token budget.
Priority ordering should be explicit.
9. Memory Architecture
The application should distinguish:
-
Authoritative transcript
- never pruned from storage.
-
Recent context
- direct recent turns.
-
Summaries
- compressed representation of older ranges.
-
Structured state
- current accepted facts/entities/threads.
-
Retrievable memories
- indexed older story events.
-
Imported knowledge
- local canon/reference/inspiration.
The exact retrieval implementation is a Phase 0 decision.
10. Knowledge Ingestion
Initial ingestion pipeline:
Local file
|
v
Parse as data
|
v
Classify source:
Canon / Reference / Inspiration
|
v
Chunk
|
+--> lexical index
|
+--> optional local embeddings
|
v
Store provenance
Requirements:
- imported content is not executable,
- no macros/plugins/scripts from imported content,
- no automatic URL fetching,
- original source metadata preserved,
- campaign association explicit,
- re-indexing repeatable.
11. Branching and Restore Model
Preferred behavior:
- accepted turns are immutable historical events,
- editing or retrying creates a new continuation unless the implementation provides an equally auditable versioning model,
- restore changes active branch/head rather than deleting historical rows,
- named checkpoints point to turn/state identities,
- abandoned branches remain navigable,
- explicit delete may be supported later.
Phase 0 should compare existing candidate implementations against this model.
12. Prompt and Provenance Inspection
For each narrator turn, retain enough information to answer:
- what instructions were sent,
- what story state was included,
- what old memories were retrieved,
- what knowledge chunks were retrieved,
- which model/settings were used,
- what structured updates were proposed,
- what was accepted/rejected.
Storage may use normalized tables or compressed prompt snapshots.
13. Security Design
13.1 Network
Default production mode:
- browser connects to local application,
- application connects to local Ollama,
- no required outbound Internet access.
Phase 0 must inventory all network behavior inherited from any fork.
13.2 Remote dependency removal
Candidate fork review must identify:
- analytics SDKs,
- telemetry,
- crash reporting,
- hosted fonts,
- CDNs,
- remote image assets,
- update checks,
- cloud auth,
- remote databases,
- cloud model providers,
- web scraping/fetch features.
Any retained remote behavior must be explicitly justified and configurable; preferred v1 state is none.
13.3 Imported content
Treat imports as untrusted data.
Do not:
- execute HTML/JS from imported material,
- execute scripts,
- execute plugin code,
- follow embedded URLs automatically,
- pass filesystem paths to the model unnecessarily.
13.4 Ollama endpoint
Default loopback.
Potential future LAN support should require explicit configuration and documented security implications.
14. Browser Architecture
The final frontend framework should depend partly on fork selection.
Candidate inherited stacks may include React/Next.js or other browser frameworks.
Required UI capabilities:
- streaming text,
- responsive transcript,
- branch navigation,
- state/library inspectors,
- local settings,
- future image/video display,
- no hard dependency on remote CDN resources at runtime.
15. Future Media Architecture
15.1 Scene packet
The story engine should be able to transform authoritative state into a neutral media packet.
Example:
scene_id: scene-128
source_turns: [128, 129]
location: Crooked Lantern tavern
mood: tense
characters:
- id: aldric
visual_reference: ...
- id: mara
visual_reference: ...
important_actions:
- Aldric enters
- Mara signals from the rear table
visual_continuity:
- same green cloak as prior scene
15.2 Provider abstraction
Future conceptual interface:
generate_scene_image(scene_id, options)
generate_character_portrait(character_id, options)
generate_video(turn_start, turn_end, options)
No story-engine component should depend on a specific media model.
15.3 Asset provenance
Store:
- model/provider,
- prompt,
- settings,
- scene/turn sources,
- generation date,
- local file location,
- optional seed/workflow metadata.
16. Genre Profiles
Profiles should be configuration.
Example hard-SF profile:
genre: hard_scifi
rules:
- Respect established technology limits.
- Do not introduce supernatural events unless canon allows them.
- Preserve travel-time and distance continuity.
Example fantasy profile:
genre: low_fantasy
rules:
- Magic exists only as established in campaign canon.
- Avoid modern technology.
- Preserve setting-specific social and technological constraints.
The storage and story engine should not change between these profiles.
17. Testing Strategy
Phase 0 should determine inherited test quality.
v1 should ultimately cover:
- story turn persistence,
- branch creation,
- checkpoint restore,
- state rollback,
- failed model call recovery,
- malformed structured extraction,
- context budgeting,
- knowledge retrieval,
- source provenance,
- export/import,
- local-only network assumptions,
- schema migration,
- media schema backward compatibility.
18. Open Technical Questions for Phase 0
- Which repository should be the base?
- Which existing story-tree implementation is safest to reuse?
- Should state be event-sourced, snapshot-based, or hybrid?
- Is SQLite alone sufficient for embeddings?
- Which Ollama embedding model is appropriate?
- How should retrieved story memories differ from imported lore?
- How much state extraction can be deterministic?
- Should the app use one model for narration and another for summarization/extraction?
- How should edit/retry semantics map to branches?
- What exact campaign export format should v1 use?
- Which schema pieces should be introduced now solely for future media?
- Which dependencies in candidate forks violate local-only requirements?
19. Technical Design v1.0 Exit Criteria
This document becomes v1.0 only after Phase 0 has:
- selected the base architecture,
- validated the candidate application locally,
- selected storage/versioning strategy,
- selected memory/retrieval design,
- selected Ollama integration approach,
- completed dependency/network/privacy review,
- documented migration/reuse strategy,
- resolved licensing questions,
- defined v1 API/component boundaries,
- defined implementation milestones.