24 KiB
Adventure Storyteller — Media Extension Contract
Status: Draft v0.1
Purpose: Define the stable interfaces and data boundaries needed to add local image, video, and audio generation later without coupling media generation to the core story engine.
1. Design Goal
The core storyteller must not depend on any specific media generator.
The story system should remain fully functional with:
media_enabled = false
Media is an optional extension layer.
The core rule is:
Story state is authoritative. Media is derived from story state and scene history.
Generated media must never become the database of record for story facts.
2. Scope
This document covers future:
- still-image generation,
- video generation,
- audio ambience,
- narration/TTS,
- character voice,
- speech-to-text (STT) for player input/dictation,
- scene illustration,
- multi-turn scene recap clips.
It does not require any of these to ship in v1.
The v1 requirement is architectural compatibility.
3. Architectural Separation
Recommended architecture:
Story Engine
|
v
Scene Extraction
|
v
Scene Packet
|
v
Media Coordinator
|
+--> Image Provider
+--> Video Provider
+--> Audio Provider
+--> TTS Provider
The Story Engine should not directly call:
- ComfyUI,
- Stable Diffusion,
- video pipelines,
- TTS engines,
- third-party media APIs.
4. Core Responsibilities
Story Engine
Responsible for:
- accepted transcript,
- canon,
- state,
- entities,
- relationships,
- scenes,
- lineage,
- checkpoints,
- prompt provenance.
Scene Extractor
Responsible for turning accepted story state into a structured media-ready scene description.
Media Coordinator
Responsible for:
- provider selection,
- request normalization,
- job creation,
- status tracking,
- retry,
- cancellation,
- output registration,
- provenance.
Media Provider
Responsible for:
- translating normalized request into provider-specific format,
- invoking local generator,
- returning output metadata.
5. Scene Snapshot Requirement
The story system should persist a structured scene snapshot for relevant turns.
Conceptual fields:
scene_id:
campaign_id:
branch_id:
start_turn_id:
end_turn_id:
location_id:
characters_present:
objects_present:
time_of_day:
weather:
lighting:
mood:
visual_notes:
action_summary:
dialogue_summary:
camera_hint:
created_at:
Not every field is required for every scene.
6. Scene Snapshot Authority
A scene snapshot is derived from accepted story state.
If it conflicts with canonical state:
- canonical state wins,
- scene snapshot should be regenerated or corrected.
Scene snapshots from abandoned history remain associated with that abandoned lineage.
7. Visual Character Profiles
Characters may have optional stable visual descriptors.
Example:
character_id: mara
apparent_age: early 40s
build: sturdy
hair: dark auburn
eyes: gray
clothing_baseline: practical innkeeper clothing
distinctive_features:
- small burn scar on right forearm
style_notes:
- grounded realism
These descriptors should support continuity across generated images.
8. Visual Location Profiles
Locations may have optional stable descriptors.
Example:
location_id: crooked-lantern
architecture: timber-framed roadside tavern
interior:
- stone hearth
- dark beams
- shared wooden tables
lighting:
- candles
- oil lamps
visual_identity:
- warm but worn
9. Visual Item Profiles
Important recurring items may also have visual descriptors.
Example:
item_id: silver-key
material: silver
shape: small old-fashioned key
marking: broken-circle symbol
This is optional but useful for continuity.
10. Scene Packet
The Media Coordinator should consume a normalized Scene Packet rather than raw transcript text.
Example:
scene_id: scene-142
campaign_id: continuity-test
turn_range:
start: 140
end: 143
location:
name: Crooked Lantern Tavern
visual_profile: ...
characters:
- name: Aldric
visual_profile: ...
current_condition: ...
- name: Mara
visual_profile: ...
objects:
- Silver Key
action_summary: >
Aldric places the silver key on the table while Mara studies
the broken-circle symbol.
mood: tense curiosity
lighting: dim oil-lamp light
time_of_day: night
continuity_constraints:
- Aldric still owns the key
- Mara has not yet entered the cellar
- no modern objects
11. Scene Packet Purpose
The Scene Packet provides:
- stable provider-independent input,
- continuity,
- exact lineage,
- provenance,
- repeatability,
- easier testing,
- future provider switching.
12. Raw Transcript Access
Providers should not normally receive the entire campaign transcript.
Preferred:
- structured scene packet,
- selected supporting recent text,
- only the minimum needed.
This improves:
- privacy,
- consistency,
- prompt size,
- provider portability.
13. Media Request
Conceptual request:
media_request_id:
campaign_id:
scene_id:
media_type:
provider_id:
requested_by:
created_at:
settings:
prompt_override:
reference_assets:
Media types:
image
video
audio
tts
stt
14. Media Job
Each generation attempt should create a job record.
Conceptual fields:
job_id:
media_request_id:
provider_id:
status:
started_at:
completed_at:
error:
provider_settings:
seed:
model:
model_version:
input_hash:
Statuses:
queued
running
completed
failed
cancelled
15. Media Asset
Successful jobs produce media assets.
Conceptual fields:
asset_id:
campaign_id:
scene_id:
job_id:
media_type:
local_path:
mime_type:
width:
height:
duration:
file_size:
content_hash:
created_at:
Optional:
- thumbnail,
- codec,
- frame rate,
- audio channels.
16. Asset Provenance
Every asset should retain:
- source scene,
- source turn range,
- active branch/lineage at generation,
- provider,
- model,
- model version,
- generation settings,
- seed if available,
- normalized scene packet,
- final rendered prompt if applicable.
This allows later explanation:
What story state produced this image?
17. Branch Awareness
Media must be branch-aware.
Example:
Path A:
Mara enters the cellar.
Path B:
Mara remains upstairs.
An image generated for Path A must not appear as the current scene illustration on Path B.
The asset may remain stored.
It becomes inactive/disposable with the abandoned scene lineage.
18. Undo and Restore
When the story is undone/restored:
- story state changes immediately,
- media is not authoritative,
- old media remains attached to old scene/lineage,
- UI should stop treating old media as current.
Do not delete media automatically.
19. Media Cleanup
Media consumes far more storage than text.
A future cleanup feature should consider:
- abandoned lineage,
- age,
- asset size,
- user favorites,
- checkpoint references,
- regeneration ability.
Potential actions:
Delete abandoned media
Delete unstarred generations
Keep final selections only
No automatic cleanup required initially.
20. User Selection
When multiple generations exist:
Image A
Image B
Image C
The user may select one as:
preferred_asset = true
Other assets remain stored unless deleted.
21. Image Provider Interface
Conceptual interface:
generate_image(scene_packet, settings) -> media_result
Provider capabilities should declare:
supports_seed:
supports_negative_prompt:
supports_reference_images:
supports_character_reference:
supports_controlnet:
supports_inpainting:
supports_upscale:
22. Video Provider Interface
Conceptual interface:
generate_video(scene_packet, settings) -> media_result
Potential capabilities:
supports_text_to_video:
supports_image_to_video:
supports_reference_frames:
supports_audio:
supports_seed:
max_duration_seconds:
23. Audio Provider Interface
Conceptual:
generate_audio(scene_packet, settings) -> media_result
Potential uses:
- tavern ambience,
- rain,
- machinery,
- battle sounds.
24. TTS Provider Interface
Conceptual:
synthesize_speech(text, voice_profile, settings) -> media_result
Potential future uses:
- narrator voice,
- NPC voices,
- replaying dialogue.
24A. Speech-to-Text Provider Interface
Conceptual:
transcribe_audio(audio_input, settings) -> transcription_result
Potential future uses:
- dictate player actions,
- dictate dialogue,
- hands-free story input,
- accessibility.
Required semantic rule:
STT output is draft user input, not an accepted story event.
Recommended workflow:
Microphone / local audio
|
v
Local STT Provider
|
v
Draft transcription
|
v
User review/edit
|
v
Normal story input submission
The user should be able to edit the transcription before it enters the authoritative transcript.
The STT provider should remain local by default and should not require a cloud transcription API.
25. Provider Capability Discovery
Providers should expose capabilities.
Example:
provider_id: local-comfyui
media_types:
- image
supports:
seed: true
reference_images: true
inpainting: true
The UI should adapt based on capability.
26. Provider Configuration
Provider config should be separate from campaign data where possible.
Example:
provider_id: local-comfyui
endpoint: http://127.0.0.1:8188
enabled: true
27. Local-Only Requirement
Preferred v1/future default:
Media provider endpoints must be loopback/local.
No cloud generation should be required.
If remote media providers are ever added:
- they must be explicit,
- off by default,
- clearly marked as data-leaving-machine behavior.
28. Provider Endpoint Validation
Default allowed endpoints:
127.0.0.1
localhost
Future advanced setting may allow:
- LAN,
- Tailscale,
- trusted workstation.
Not required initially.
29. Image Generation Example
Story:
Aldric enters the Crooked Lantern during a storm.
Mara stands behind the bar.
Scene Packet:
Location:
old timber tavern
Characters:
Aldric
Mara
Weather:
heavy rain outside
Lighting:
warm oil lamps
Mood:
tense arrival
The Image Provider turns this into provider-specific prompt/controls.
30. Video Generation Use Case
User wants to convert:
Turns 210-215
into a short battle clip.
Workflow:
Select turn range
|
v
Build multi-turn scene packet
|
v
Extract action beats
|
v
Create shot/sequence plan
|
v
Video Provider
31. Multi-Turn Scene Packet
For video, include ordered beats.
Example:
beats:
- Aldric draws sword
- guard lunges
- Aldric blocks
- lantern falls
- room catches partial fire
Each beat should preserve:
- characters,
- location,
- object state,
- continuity.
32. Shot Planning
A future Video Coordinator may generate:
shots:
- wide establishing shot
- medium combat shot
- close-up of key falling
- final wide shot
This is not a Story Engine responsibility.
33. Video Duration
The video request should explicitly define:
target_duration
rather than assuming one turn equals one fixed duration.
34. Scene Compression
Long turn ranges may need compression into:
- action beats,
- visual summary,
- omitted dialogue.
This derived representation should retain source turn IDs.
35. Media Does Not Change Canon
A generated image may be wrong.
Example:
- wrong hair color,
- extra sword,
- wrong number of characters.
The image does not alter story state.
Correction options:
- regenerate,
- edit prompt,
- inpaint,
- reject asset.
36. Media Feedback
The user may mark an asset:
accepted_visual
This means:
- preferred depiction,
- useful visual continuity reference.
It still should not automatically override explicit canonical data.
37. Visual Canon Promotion
Potential future feature:
User explicitly selects:
Promote visual detail to canon
Example:
- scar shape,
- clothing color,
- vehicle appearance.
This must be explicit.
Media output should never auto-promote visual details.
38. Reference Images
Future image/video providers may accept:
- character portrait,
- location image,
- item image.
These assets should have:
- local IDs,
- provenance,
- explicit role.
Example:
reference_role: character_identity
39. Character Consistency
Media coordinator may use:
- visual profile,
- prior accepted portrait,
- reference image,
- seed,
- adapter/LoRA where supported.
The provider-specific method should remain outside Story Engine.
40. Location Consistency
Same principle:
- stable visual profile,
- accepted reference image,
- provider-specific conditioning.
41. Style Profiles
Campaign may define:
visual_style:
realism: grounded cinematic
palette: muted
era_accuracy: high
Style profile is campaign configuration.
It should not alter story canon.
42. Prompt Templates
Provider-specific prompt templates belong in the media subsystem.
Example:
[visual style]
[location]
[characters]
[action]
[lighting]
[continuity constraints]
Story engine should not contain Stable Diffusion syntax.
43. Negative Prompts
If supported, Media Coordinator may generate:
- no modern objects,
- no duplicate characters,
- no text overlays.
This is provider-specific optional metadata.
44. Determinism
Where provider supports seeds:
Store seed.
This enables:
- recreation,
- variations,
- debugging.
Determinism is helpful but not required across all providers.
45. Media Retries
Retrying media generation should create a new job/asset.
Do not overwrite prior asset.
Example:
scene-42
image-a
image-b
image-c
46. Media Edits
Future:
- inpaint,
- upscale,
- image-to-image,
- video refinement.
Each edit should preserve parent asset lineage.
Conceptually:
asset B derived_from asset A
47. Asset Lineage
Potential fields:
parent_asset_id:
operation:
- regenerate
- inpaint
- upscale
- animate
48. Story-to-Media Provenance
Every media asset should be traceable:
Campaign
-> Branch
-> Turn range
-> Scene snapshot
-> Media request
-> Job
-> Asset
49. Media-to-Story Separation
The reverse must not happen automatically:
Asset
-X-> Story state
unless the user explicitly promotes information.
50. Failure Handling
If media generation fails:
- story remains unaffected,
- scene remains valid,
- job records failure,
- user may retry,
- no partial story mutation.
51. Provider Timeout
Media jobs may be long.
Coordinator should support:
- status,
- timeout,
- cancellation,
- retry.
No need to block story interaction while media generates.
52. Asynchronous Design
The architecture should assume media generation can happen asynchronously relative to story interaction.
The user may continue the story while an image/video job runs.
When complete:
- asset attaches to source scene,
- it should not become current merely because story has advanced.
53. Job Queue
A simple local job queue may be needed.
Conceptual:
pending
running
completed
failed
Implementation may be:
- database-backed,
- in-process,
- worker process.
Do not introduce distributed infrastructure for v1.
54. Restart Recovery
If app restarts during media job:
Preferred:
- mark interrupted jobs,
- allow retry,
- preserve completed files.
Provider-specific resume is optional.
55. Storage Layout
Recommended:
data/
campaigns/
<campaign-id>/
media/
images/
video/
audio/
thumbnails/
Exact layout is implementation-specific.
56. File Naming
Do not use raw user text as filenames.
Use:
- IDs,
- hashes,
- safe extensions.
Original labels can exist in metadata.
57. Media Hashing
Compute content hash.
Useful for:
- duplicate detection,
- integrity,
- export verification.
58. Export
Campaign export should include:
- selected media assets,
- metadata,
- provenance.
Potential export modes:
Full
No Media
Selected Media Only
For v1, if media is not implemented, preserve schema compatibility.
59. Import
Campaign import should restore:
- asset metadata,
- local file associations,
- source scene references.
Missing media files should not break story history.
60. Thumbnail Generation
Thumbnails are derived cache.
They may be recreated.
Do not treat them as canonical assets.
61. Browser Delivery
Media should be served through scoped local routes.
Do not expose arbitrary filesystem paths.
62. Security
Generated files are still untrusted browser content.
Use:
- correct MIME types,
- safe content disposition,
- no arbitrary executable serving,
- no remote media loading by default.
63. Prompt Privacy
Media prompts may contain:
- hidden characters,
- plot secrets,
- campaign state.
Therefore local-only media providers are preferred.
64. Hidden Information
A scene illustration should not accidentally reveal narrator-only hidden canon unless the scene logically exposes it.
Example:
- hidden trap behind wall,
- secret identity,
- concealed character.
Scene extractor should distinguish:
- visible facts,
- narrator-only facts.
65. Visible Scene State
Media should generally receive only visually observable information plus necessary visual continuity data.
It should not receive unrelated hidden plot details.
66. Audio Privacy
TTS text may contain full dialogue.
Keep speech generation local by default.
67. Voice Profiles
Future:
voice_profile_id:
character_id:
provider:
voice_name:
settings:
Voice profile is presentation metadata.
It is not story canon.
67A. Speech-to-Text Privacy and Input Semantics
STT may receive live microphone audio or a local recorded clip.
Security/privacy requirements:
- audio remains local by default,
- microphone access requires normal browser/user permission,
- no automatic background recording,
- recording state must be visibly indicated,
- transcript remains editable before submission,
- failed/partial transcription must not create a story turn,
- raw audio retention should be optional and off by default unless needed for debugging or user-requested history.
STT should not bypass the normal story-input validation and commit path.
68. Music
Future ambient/music generation should be:
- optional,
- local,
- scene-linked.
Do not make it part of core story context.
69. User Controls
Potential media UI:
Generate Image
Generate Video
Generate Audio
Read Aloud
Dictate
Regenerate
Use as Preferred
Delete
Show Provenance
Output-media controls should appear as optional scene actions. Dictate belongs near the story input field because STT produces draft user input.
70. Default Media Behavior
Recommended default:
No automatic media generation.
Reason:
- GPU cost,
- storage,
- latency,
- user control.
User explicitly requests media.
71. Optional Auto-Illustration
Future setting:
Auto-generate one image at scene changes
Off by default.
72. Scene Change Detection
Future media automation may detect:
- new location,
- major character entrance,
- major action event,
- chapter boundary.
This remains optional.
73. Resource Coordination
Local LLM and image/video models may compete for GPU memory.
Media coordinator should eventually support:
- queueing,
- model unload/reload,
- provider limits.
Do not assume simultaneous execution is always possible.
74. Hardware Awareness
Providers may expose:
- VRAM requirement,
- model availability,
- estimated capability.
The core story engine should not need this information.
75. Provider Errors
Normalize provider errors.
Example:
error_type:
model_missing
out_of_memory
invalid_request
provider_unreachable
cancelled
76. Model Discovery
Future media UI may list local models available from provider.
Do not auto-download them.
77. Model Downloads
Same privacy rule as Ollama:
- installation may use Internet,
- normal generation should not require Internet,
- no silent downloads.
78. Media Provider Registry
Conceptual registry:
providers:
- id: comfyui-local
types: [image, video]
- id: kokoro-local
types: [tts]
- id: local-stt
types: [stt]
Exact products are not committed.
79. Open Dungeon Reuse
Open Dungeon is especially relevant for:
- local image generation,
- character visual continuity,
- image workflow UX.
During Phase 0B, inspect:
- how scene prompts are built,
- how images are attached to story,
- provider coupling,
- whether media can be separated from destructive history model.
Do not copy its persistence limitations into the core story architecture.
80. Gamentic Reuse
Gamentic is relevant for:
- provider abstraction,
- asynchronous generation,
- media job concepts,
- local multimodal architecture.
Use as a pattern reference, not a merged codebase.
81. Corvus Story Core Reuse
Corvus may be useful for:
- visual scene extraction,
- TTS/ComfyUI integration patterns.
Again, use concepts selectively.
82. V1 Physical Schema Decision
Open question:
Should media tables physically exist in v1?
Recommended:
Include minimal media-ready schema/interfaces if inexpensive, but do not build provider implementation solely to justify them.
Minimum useful v1 fields:
- scene snapshot,
- stable visual profiles,
- media provider interface/types,
- optional media asset table.
Phase 0B should determine cost.
83. V1 Required Media Readiness
Even without generation, v1 should preserve:
- scene identity,
- scene turn range,
- character visual profiles,
- location visual profiles,
- branch-aware scene snapshots.
This is sufficient to avoid architectural dead ends.
84. Future Image Acceptance Test
Given an accepted scene:
Aldric and Mara examine the Silver Key in the tavern.
Generate image.
Pass if:
- only local provider used,
- asset attached to correct scene,
- provenance recorded,
- story state unchanged,
- retry creates new asset rather than overwriting.
85. Future Branch Media Acceptance Test
Generate image on Path A.
Restore checkpoint and create Path B.
Pass if:
- Path A image remains stored,
- Path A image is not shown as current Path B media,
- no automatic deletion occurs.
86. Future Video Acceptance Test
Select 4-5 combat turns.
Generate short local video.
Pass if:
- source turn range preserved,
- action sequence matches accepted branch,
- abandoned-branch events are not included,
- media generation does not mutate story.
87. Phase 0B Validation Questions
Codex should answer:
- How does Open Dungeon attach generated images to story messages/scenes?
- Is image generation coupled to linear/destructive message history?
- Can its visual character continuity data be reused independently?
- What provider assumptions are hardcoded?
- Can ComfyUI/local provider calls run fully offline?
- Does AI-DnD already have scene-like structured state suitable for media extraction?
- Can scene snapshots be added without RPG schema coupling?
- What parts of Gamentic's provider abstraction are worth reimplementing?
- Should media asset tables exist in v1 or be deferred?
- Can media generation remain entirely optional without special-casing core story logic?
88. Acceptance Criteria for Architecture
The architecture passes if:
- story engine functions with no media provider,
- media requests derive from accepted story state,
- assets are branch-aware,
- assets preserve source scene/turn provenance,
- media failure cannot corrupt story state,
- provider-specific syntax stays outside Story Engine,
- local providers can be substituted,
- future image/video/audio/TTS types fit the output job/asset model,
- future STT fits the same provider architecture while feeding editable draft input rather than story state,
- generated media never automatically becomes canon,
- abandoned-history media remains recoverable but inactive.
89. Current Recommendation
Use this conceptual contract:
Authoritative Story
|
v
Scene Snapshot
|
v
Scene Packet
|
v
Media Coordinator
|
+--> Local Image Provider
+--> Local Video Provider
+--> Local Audio Provider
+--> Local TTS Provider
|
v
Media Asset + Provenance
The media subsystem should depend on the story engine.
The story engine should not depend on the media subsystem.
That one-way dependency is the most important architectural requirement in this document.