711 lines
16 KiB
Markdown
711 lines
16 KiB
Markdown
# Adventure Storyteller — Technical Design
|
|
|
|
**Status:** Provisional v0.1
|
|
**Important:** This document describes the current preferred architecture. Phase 0 research is expected to confirm, revise, or replace portions of it.
|
|
|
|
## 1. Design Objective
|
|
|
|
Implement a local-first, browser-based interactive storytelling system in which:
|
|
|
|
- Ollama provides local AI inference,
|
|
- the application owns authoritative story state,
|
|
- complete story history is persistent,
|
|
- long-running context is reconstructed from state, summaries, retrieval, and recent turns,
|
|
- users can checkpoint, restore, retry, and branch,
|
|
- imported knowledge remains local,
|
|
- future media generation can be added without redesigning the story engine.
|
|
|
|
## 2. Current Provisional Architecture
|
|
|
|
```text
|
|
Local Browser
|
|
|
|
|
v
|
|
+------------------+
|
|
| Browser UI |
|
|
+--------+---------+
|
|
|
|
|
v
|
|
+------------------+
|
|
| Story Director |
|
|
| API / Service |
|
|
+---+----------+---+
|
|
| |
|
|
+--------+ +----------------+
|
|
v v
|
|
+----------------------+ +----------------------+
|
|
| Authoritative Store | | Context / Retrieval |
|
|
| SQLite (provisional) | | local only |
|
|
+----------+-----------+ +----------+-----------+
|
|
| |
|
|
| v
|
|
| +----------------------+
|
|
| | Local embeddings / |
|
|
| | lexical retrieval |
|
|
| +----------+-----------+
|
|
| |
|
|
+-------------------+------------------+
|
|
|
|
|
v
|
|
+---------------+
|
|
| Ollama |
|
|
| localhost |
|
|
+-------+-------+
|
|
|
|
|
v
|
|
Local narrator model
|
|
```
|
|
|
|
Future:
|
|
|
|
```text
|
|
Story / Scene State
|
|
|
|
|
v
|
|
+-------------------+
|
|
| Media Coordinator |
|
|
+----+---------+----+
|
|
| |
|
|
v v
|
|
Image Video
|
|
Provider Provider
|
|
```
|
|
|
|
## 3. Base Repository Strategy
|
|
|
|
**Status: UNDECIDED**
|
|
|
|
Phase 0 will determine whether to:
|
|
|
|
1. fork Open Dungeon and add stronger state/memory/branching,
|
|
2. fork AI-DnD and remove RPG/cloud complexity,
|
|
3. use `CaoRuiming/ai-adventure` as the core and add browser/Ollama layers,
|
|
4. build a thin new application using selected reusable components,
|
|
5. choose another candidate discovered during research.
|
|
|
|
The chosen strategy must be justified with code-level evidence rather than README feature comparison alone.
|
|
|
|
## 4. Component Boundaries
|
|
|
|
### 4.1 Browser UI
|
|
|
|
Responsibilities:
|
|
|
|
- campaign selection,
|
|
- campaign creation/editing,
|
|
- transcript display,
|
|
- streaming narrator output,
|
|
- user input,
|
|
- branch/checkpoint navigation,
|
|
- state inspection/editing,
|
|
- library/source management,
|
|
- prompt/context inspection,
|
|
- settings,
|
|
- future media controls/gallery.
|
|
|
|
The browser UI must not directly own story authority.
|
|
|
|
### 4.2 Story Director
|
|
|
|
Responsibilities:
|
|
|
|
- accept user turns,
|
|
- load authoritative campaign/branch state,
|
|
- construct model context,
|
|
- invoke Ollama,
|
|
- validate and commit accepted outputs,
|
|
- trigger summarization/state extraction as needed,
|
|
- maintain branch/tree relationships,
|
|
- create scene snapshots,
|
|
- record provenance/debug metadata,
|
|
- expose state/history APIs to the browser.
|
|
|
|
### 4.3 Model Adapter
|
|
|
|
v1 target: Ollama.
|
|
|
|
Responsibilities:
|
|
|
|
- enumerate allowed local models,
|
|
- invoke chat/generation,
|
|
- support streaming,
|
|
- invoke local embedding model if selected,
|
|
- expose model metadata,
|
|
- reject unsupported remote/cloud providers.
|
|
|
|
Provisional endpoint default:
|
|
|
|
```text
|
|
http://127.0.0.1:11434
|
|
```
|
|
|
|
### 4.4 Authoritative Store
|
|
|
|
**Provisional choice:** SQLite.
|
|
|
|
Reasons:
|
|
|
|
- local,
|
|
- transactional,
|
|
- portable,
|
|
- easy backup,
|
|
- strong fit for structured story/state data,
|
|
- can support FTS,
|
|
- no external service required.
|
|
|
|
Phase 0 must validate whether the selected fork already has a suitable schema and migration system.
|
|
|
|
### 4.5 Knowledge Store
|
|
|
|
Provisional options:
|
|
|
|
- SQLite FTS,
|
|
- embeddings stored in SQLite,
|
|
- local vector library,
|
|
- hybrid lexical/vector retrieval.
|
|
|
|
Remote vector databases are out of scope for v1.
|
|
|
|
### 4.6 Media Coordinator
|
|
|
|
Not required for v1 implementation, but interface boundaries should be reserved.
|
|
|
|
Responsibilities later:
|
|
|
|
- accept scene/character/turn-range generation requests,
|
|
- transform story state into media-generation packets,
|
|
- call pluggable local media providers,
|
|
- record asset provenance,
|
|
- attach assets to campaigns/scenes/turns.
|
|
|
|
## 5. Authoritative Data Model
|
|
|
|
The exact schema is provisional.
|
|
|
|
### 5.1 Campaign
|
|
|
|
Possible fields:
|
|
|
|
- `id`
|
|
- `title`
|
|
- `profile`
|
|
- `tone`
|
|
- `style`
|
|
- `narrator_rules`
|
|
- `created_at`
|
|
- `updated_at`
|
|
- `active_branch_id`
|
|
- `model_config_id`
|
|
|
|
### 5.2 Branch
|
|
|
|
Possible fields:
|
|
|
|
- `id`
|
|
- `campaign_id`
|
|
- `name`
|
|
- `root_turn_id`
|
|
- `head_turn_id`
|
|
- `created_from_branch_id`
|
|
- `created_at`
|
|
|
|
### 5.3 Turn
|
|
|
|
Possible fields:
|
|
|
|
- `id`
|
|
- `campaign_id`
|
|
- `branch_id`
|
|
- `parent_turn_id`
|
|
- `sequence_hint`
|
|
- `user_input`
|
|
- `assistant_output`
|
|
- `created_at`
|
|
- `model_id`
|
|
- `generation_config`
|
|
- `prompt_snapshot_id`
|
|
- `state_version_id`
|
|
- `scene_snapshot_id`
|
|
|
|
The graph/tree relationship should come from parentage, not merely sequential row order.
|
|
|
|
### 5.4 Checkpoint
|
|
|
|
Possible fields:
|
|
|
|
- `id`
|
|
- `campaign_id`
|
|
- `turn_id`
|
|
- `name`
|
|
- `notes`
|
|
- `created_at`
|
|
|
|
Every accepted turn is implicitly recoverable even when not given a name.
|
|
|
|
### 5.5 Narrative Entity
|
|
|
|
A generic entity model should avoid genre-specific database design.
|
|
|
|
Possible categories:
|
|
|
|
- character,
|
|
- location,
|
|
- organization,
|
|
- item,
|
|
- vehicle,
|
|
- object,
|
|
- concept,
|
|
- other.
|
|
|
|
Possible fields:
|
|
|
|
- `id`
|
|
- `campaign_id`
|
|
- `type`
|
|
- `name`
|
|
- `canonical_description`
|
|
- `visual_description`
|
|
- `status`
|
|
- `metadata_json`
|
|
|
|
Separate normalized tables may replace a generic entity table if research shows that is cleaner.
|
|
|
|
### 5.6 Fact
|
|
|
|
Potential representation:
|
|
|
|
- subject,
|
|
- predicate,
|
|
- object/value,
|
|
- source turn,
|
|
- canonical status,
|
|
- validity interval/version,
|
|
- confidence/review state.
|
|
|
|
The design should distinguish accepted canon from merely proposed model content.
|
|
|
|
### 5.7 Relationship
|
|
|
|
Potential examples:
|
|
|
|
- character-to-character,
|
|
- character-to-organization,
|
|
- entity-to-location,
|
|
- ownership,
|
|
- allegiance,
|
|
- trust/hostility,
|
|
- family/friendship.
|
|
|
|
### 5.8 Story Thread
|
|
|
|
Possible fields:
|
|
|
|
- title,
|
|
- description,
|
|
- status,
|
|
- opened_turn_id,
|
|
- resolved_turn_id,
|
|
- importance,
|
|
- related entities.
|
|
|
|
### 5.9 Summary
|
|
|
|
Potential levels:
|
|
|
|
- full campaign summary,
|
|
- arc/chapter summary,
|
|
- branch summary,
|
|
- rolling compressed memory.
|
|
|
|
Each summary should record what source turns it represents.
|
|
|
|
### 5.10 Scene Snapshot
|
|
|
|
Provisional fields:
|
|
|
|
- `id`
|
|
- `campaign_id`
|
|
- `branch_id`
|
|
- `source_turn_start`
|
|
- `source_turn_end`
|
|
- `location_entity_id`
|
|
- `time_description`
|
|
- `mood`
|
|
- `participants_json`
|
|
- `environment_json`
|
|
- `visual_notes`
|
|
- `action_beats_json`
|
|
- `continuity_notes_json`
|
|
|
|
### 5.11 Asset / Asset Job
|
|
|
|
May exist in schema before implementation.
|
|
|
|
Potential asset fields:
|
|
|
|
- `id`
|
|
- `campaign_id`
|
|
- `scene_id`
|
|
- `type`
|
|
- `provider`
|
|
- `model`
|
|
- `prompt`
|
|
- `settings_json`
|
|
- `file_path`
|
|
- `created_at`
|
|
- `source_turn_range`
|
|
|
|
No v1 dependency should require these tables to contain data.
|
|
|
|
## 6. Turn Processing
|
|
|
|
Provisional turn pipeline:
|
|
|
|
```text
|
|
1. Receive user input
|
|
2. Resolve campaign + active branch + parent turn
|
|
3. Load authoritative state
|
|
4. Retrieve recent turns
|
|
5. Retrieve relevant older story memory
|
|
6. Retrieve relevant local canon/reference/inspiration
|
|
7. Build narrator context
|
|
8. Save prompt/context provenance
|
|
9. Invoke Ollama narrator model
|
|
10. Stream response to UI
|
|
11. Validate completion
|
|
12. Extract proposed state changes
|
|
13. Validate proposed state changes
|
|
14. Commit turn + state + scene atomically
|
|
15. Update summaries/indexes when thresholds require it
|
|
16. Expose new recoverable branch head
|
|
```
|
|
|
|
The final implementation may combine or reorder steps depending on selected repository architecture.
|
|
|
|
## 7. State Extraction
|
|
|
|
The system may use a second local model call to convert narration into proposed structured changes.
|
|
|
|
Example:
|
|
|
|
```json
|
|
{
|
|
"new_facts": [],
|
|
"changed_entities": [],
|
|
"opened_threads": [],
|
|
"resolved_threads": [],
|
|
"scene_changes": []
|
|
}
|
|
```
|
|
|
|
Rules:
|
|
|
|
- model-produced state changes are proposals,
|
|
- authoritative updates must pass application validation,
|
|
- invalid structured output must not corrupt the campaign,
|
|
- narration should remain preserved even if extraction must be retried or repaired.
|
|
|
|
A deterministic/non-LLM extraction layer may supplement this later.
|
|
|
|
## 8. Context Construction
|
|
|
|
Target conceptual structure:
|
|
|
|
```text
|
|
Narrator/system rules
|
|
+
|
|
Campaign profile
|
|
+
|
|
Authoritative canon/world rules
|
|
+
|
|
Current narrative state
|
|
+
|
|
Campaign/arc summary
|
|
+
|
|
Relevant older memories
|
|
+
|
|
Relevant imported local material
|
|
+
|
|
Recent turns
|
|
+
|
|
Current user input
|
|
```
|
|
|
|
Context must be bounded by configurable token budget.
|
|
|
|
Priority ordering should be explicit.
|
|
|
|
## 9. Memory Architecture
|
|
|
|
The application should distinguish:
|
|
|
|
1. **Authoritative transcript**
|
|
- never pruned from storage.
|
|
|
|
2. **Recent context**
|
|
- direct recent turns.
|
|
|
|
3. **Summaries**
|
|
- compressed representation of older ranges.
|
|
|
|
4. **Structured state**
|
|
- current accepted facts/entities/threads.
|
|
|
|
5. **Retrievable memories**
|
|
- indexed older story events.
|
|
|
|
6. **Imported knowledge**
|
|
- local canon/reference/inspiration.
|
|
|
|
The exact retrieval implementation is a Phase 0 decision.
|
|
|
|
## 10. Knowledge Ingestion
|
|
|
|
Initial ingestion pipeline:
|
|
|
|
```text
|
|
Local file
|
|
|
|
|
v
|
|
Parse as data
|
|
|
|
|
v
|
|
Classify source:
|
|
Canon / Reference / Inspiration
|
|
|
|
|
v
|
|
Chunk
|
|
|
|
|
+--> lexical index
|
|
|
|
|
+--> optional local embeddings
|
|
|
|
|
v
|
|
Store provenance
|
|
```
|
|
|
|
Requirements:
|
|
|
|
- imported content is not executable,
|
|
- no macros/plugins/scripts from imported content,
|
|
- no automatic URL fetching,
|
|
- original source metadata preserved,
|
|
- campaign association explicit,
|
|
- re-indexing repeatable.
|
|
|
|
## 11. Branching and Restore Model
|
|
|
|
Preferred behavior:
|
|
|
|
- accepted turns are immutable historical events,
|
|
- editing or retrying creates a new continuation unless the implementation provides an equally auditable versioning model,
|
|
- restore changes active branch/head rather than deleting historical rows,
|
|
- named checkpoints point to turn/state identities,
|
|
- abandoned branches remain navigable,
|
|
- explicit delete may be supported later.
|
|
|
|
Phase 0 should compare existing candidate implementations against this model.
|
|
|
|
## 12. Prompt and Provenance Inspection
|
|
|
|
For each narrator turn, retain enough information to answer:
|
|
|
|
- what instructions were sent,
|
|
- what story state was included,
|
|
- what old memories were retrieved,
|
|
- what knowledge chunks were retrieved,
|
|
- which model/settings were used,
|
|
- what structured updates were proposed,
|
|
- what was accepted/rejected.
|
|
|
|
Storage may use normalized tables or compressed prompt snapshots.
|
|
|
|
## 13. Security Design
|
|
|
|
### 13.1 Network
|
|
|
|
Default production mode:
|
|
|
|
- browser connects to local application,
|
|
- application connects to local Ollama,
|
|
- no required outbound Internet access.
|
|
|
|
Phase 0 must inventory all network behavior inherited from any fork.
|
|
|
|
### 13.2 Remote dependency removal
|
|
|
|
Candidate fork review must identify:
|
|
|
|
- analytics SDKs,
|
|
- telemetry,
|
|
- crash reporting,
|
|
- hosted fonts,
|
|
- CDNs,
|
|
- remote image assets,
|
|
- update checks,
|
|
- cloud auth,
|
|
- remote databases,
|
|
- cloud model providers,
|
|
- web scraping/fetch features.
|
|
|
|
Any retained remote behavior must be explicitly justified and configurable; preferred v1 state is none.
|
|
|
|
### 13.3 Imported content
|
|
|
|
Treat imports as untrusted data.
|
|
|
|
Do not:
|
|
|
|
- execute HTML/JS from imported material,
|
|
- execute scripts,
|
|
- execute plugin code,
|
|
- follow embedded URLs automatically,
|
|
- pass filesystem paths to the model unnecessarily.
|
|
|
|
### 13.4 Ollama endpoint
|
|
|
|
Default loopback.
|
|
|
|
Potential future LAN support should require explicit configuration and documented security implications.
|
|
|
|
## 14. Browser Architecture
|
|
|
|
The final frontend framework should depend partly on fork selection.
|
|
|
|
Candidate inherited stacks may include React/Next.js or other browser frameworks.
|
|
|
|
Required UI capabilities:
|
|
|
|
- streaming text,
|
|
- responsive transcript,
|
|
- branch navigation,
|
|
- state/library inspectors,
|
|
- local settings,
|
|
- future image/video display,
|
|
- no hard dependency on remote CDN resources at runtime.
|
|
|
|
## 15. Future Media Architecture
|
|
|
|
### 15.1 Scene packet
|
|
|
|
The story engine should be able to transform authoritative state into a neutral media packet.
|
|
|
|
Example:
|
|
|
|
```yaml
|
|
scene_id: scene-128
|
|
source_turns: [128, 129]
|
|
location: Crooked Lantern tavern
|
|
mood: tense
|
|
characters:
|
|
- id: aldric
|
|
visual_reference: ...
|
|
- id: mara
|
|
visual_reference: ...
|
|
important_actions:
|
|
- Aldric enters
|
|
- Mara signals from the rear table
|
|
visual_continuity:
|
|
- same green cloak as prior scene
|
|
```
|
|
|
|
### 15.2 Provider abstraction
|
|
|
|
Future conceptual interface:
|
|
|
|
```text
|
|
generate_scene_image(scene_id, options)
|
|
generate_character_portrait(character_id, options)
|
|
generate_video(turn_start, turn_end, options)
|
|
```
|
|
|
|
No story-engine component should depend on a specific media model.
|
|
|
|
### 15.3 Asset provenance
|
|
|
|
Store:
|
|
|
|
- model/provider,
|
|
- prompt,
|
|
- settings,
|
|
- scene/turn sources,
|
|
- generation date,
|
|
- local file location,
|
|
- optional seed/workflow metadata.
|
|
|
|
## 16. Genre Profiles
|
|
|
|
Profiles should be configuration.
|
|
|
|
Example hard-SF profile:
|
|
|
|
```yaml
|
|
genre: hard_scifi
|
|
rules:
|
|
- Respect established technology limits.
|
|
- Do not introduce supernatural events unless canon allows them.
|
|
- Preserve travel-time and distance continuity.
|
|
```
|
|
|
|
Example fantasy profile:
|
|
|
|
```yaml
|
|
genre: low_fantasy
|
|
rules:
|
|
- Magic exists only as established in campaign canon.
|
|
- Avoid modern technology.
|
|
- Preserve setting-specific social and technological constraints.
|
|
```
|
|
|
|
The storage and story engine should not change between these profiles.
|
|
|
|
## 17. Testing Strategy
|
|
|
|
Phase 0 should determine inherited test quality.
|
|
|
|
v1 should ultimately cover:
|
|
|
|
- story turn persistence,
|
|
- branch creation,
|
|
- checkpoint restore,
|
|
- state rollback,
|
|
- failed model call recovery,
|
|
- malformed structured extraction,
|
|
- context budgeting,
|
|
- knowledge retrieval,
|
|
- source provenance,
|
|
- export/import,
|
|
- local-only network assumptions,
|
|
- schema migration,
|
|
- media schema backward compatibility.
|
|
|
|
## 18. Open Technical Questions for Phase 0
|
|
|
|
1. Which repository should be the base?
|
|
2. Which existing story-tree implementation is safest to reuse?
|
|
3. Should state be event-sourced, snapshot-based, or hybrid?
|
|
4. Is SQLite alone sufficient for embeddings?
|
|
5. Which Ollama embedding model is appropriate?
|
|
6. How should retrieved story memories differ from imported lore?
|
|
7. How much state extraction can be deterministic?
|
|
8. Should the app use one model for narration and another for summarization/extraction?
|
|
9. How should edit/retry semantics map to branches?
|
|
10. What exact campaign export format should v1 use?
|
|
11. Which schema pieces should be introduced now solely for future media?
|
|
12. Which dependencies in candidate forks violate local-only requirements?
|
|
|
|
## 19. Technical Design v1.0 Exit Criteria
|
|
|
|
This document becomes v1.0 only after Phase 0 has:
|
|
|
|
- selected the base architecture,
|
|
- validated the candidate application locally,
|
|
- selected storage/versioning strategy,
|
|
- selected memory/retrieval design,
|
|
- selected Ollama integration approach,
|
|
- completed dependency/network/privacy review,
|
|
- documented migration/reuse strategy,
|
|
- resolved licensing questions,
|
|
- defined v1 API/component boundaries,
|
|
- defined implementation milestones.
|