# Adventure Storyteller — Media Extension Contract **Status:** Draft v0.1 **Purpose:** Define the stable interfaces and data boundaries needed to add local image, video, and audio generation later without coupling media generation to the core story engine. ## 1. Design Goal The core storyteller must not depend on any specific media generator. The story system should remain fully functional with: ```text media_enabled = false ``` Media is an optional extension layer. The core rule is: > Story state is authoritative. Media is derived from story state and scene history. Generated media must never become the database of record for story facts. ## 2. Scope This document covers future: - still-image generation, - video generation, - audio ambience, - narration/TTS, - character voice, - speech-to-text (STT) for player input/dictation, - scene illustration, - multi-turn scene recap clips. It does not require any of these to ship in v1. The v1 requirement is architectural compatibility. ## 3. Architectural Separation Recommended architecture: ```text Story Engine | v Scene Extraction | v Scene Packet | v Media Coordinator | +--> Image Provider +--> Video Provider +--> Audio Provider +--> TTS Provider ``` The Story Engine should not directly call: - ComfyUI, - Stable Diffusion, - video pipelines, - TTS engines, - third-party media APIs. ## 4. Core Responsibilities ### Story Engine Responsible for: - accepted transcript, - canon, - state, - entities, - relationships, - scenes, - lineage, - checkpoints, - prompt provenance. ### Scene Extractor Responsible for turning accepted story state into a structured media-ready scene description. ### Media Coordinator Responsible for: - provider selection, - request normalization, - job creation, - status tracking, - retry, - cancellation, - output registration, - provenance. ### Media Provider Responsible for: - translating normalized request into provider-specific format, - invoking local generator, - returning output metadata. ## 5. Scene Snapshot Requirement The story system should persist a structured scene snapshot for relevant turns. Conceptual fields: ```yaml scene_id: campaign_id: branch_id: start_turn_id: end_turn_id: location_id: characters_present: objects_present: time_of_day: weather: lighting: mood: visual_notes: action_summary: dialogue_summary: camera_hint: created_at: ``` Not every field is required for every scene. ## 6. Scene Snapshot Authority A scene snapshot is derived from accepted story state. If it conflicts with canonical state: - canonical state wins, - scene snapshot should be regenerated or corrected. Scene snapshots from abandoned history remain associated with that abandoned lineage. ## 7. Visual Character Profiles Characters may have optional stable visual descriptors. Example: ```yaml character_id: mara apparent_age: early 40s build: sturdy hair: dark auburn eyes: gray clothing_baseline: practical innkeeper clothing distinctive_features: - small burn scar on right forearm style_notes: - grounded realism ``` These descriptors should support continuity across generated images. ## 8. Visual Location Profiles Locations may have optional stable descriptors. Example: ```yaml location_id: crooked-lantern architecture: timber-framed roadside tavern interior: - stone hearth - dark beams - shared wooden tables lighting: - candles - oil lamps visual_identity: - warm but worn ``` ## 9. Visual Item Profiles Important recurring items may also have visual descriptors. Example: ```yaml item_id: silver-key material: silver shape: small old-fashioned key marking: broken-circle symbol ``` This is optional but useful for continuity. ## 10. Scene Packet The Media Coordinator should consume a normalized Scene Packet rather than raw transcript text. Example: ```yaml scene_id: scene-142 campaign_id: continuity-test turn_range: start: 140 end: 143 location: name: Crooked Lantern Tavern visual_profile: ... characters: - name: Aldric visual_profile: ... current_condition: ... - name: Mara visual_profile: ... objects: - Silver Key action_summary: > Aldric places the silver key on the table while Mara studies the broken-circle symbol. mood: tense curiosity lighting: dim oil-lamp light time_of_day: night continuity_constraints: - Aldric still owns the key - Mara has not yet entered the cellar - no modern objects ``` ## 11. Scene Packet Purpose The Scene Packet provides: - stable provider-independent input, - continuity, - exact lineage, - provenance, - repeatability, - easier testing, - future provider switching. ## 12. Raw Transcript Access Providers should not normally receive the entire campaign transcript. Preferred: - structured scene packet, - selected supporting recent text, - only the minimum needed. This improves: - privacy, - consistency, - prompt size, - provider portability. ## 13. Media Request Conceptual request: ```yaml media_request_id: campaign_id: scene_id: media_type: provider_id: requested_by: created_at: settings: prompt_override: reference_assets: ``` Media types: ```text image video audio tts stt ``` ## 14. Media Job Each generation attempt should create a job record. Conceptual fields: ```yaml job_id: media_request_id: provider_id: status: started_at: completed_at: error: provider_settings: seed: model: model_version: input_hash: ``` Statuses: ```text queued running completed failed cancelled ``` ## 15. Media Asset Successful jobs produce media assets. Conceptual fields: ```yaml asset_id: campaign_id: scene_id: job_id: media_type: local_path: mime_type: width: height: duration: file_size: content_hash: created_at: ``` Optional: - thumbnail, - codec, - frame rate, - audio channels. ## 16. Asset Provenance Every asset should retain: - source scene, - source turn range, - active branch/lineage at generation, - provider, - model, - model version, - generation settings, - seed if available, - normalized scene packet, - final rendered prompt if applicable. This allows later explanation: ```text What story state produced this image? ``` ## 17. Branch Awareness Media must be branch-aware. Example: Path A: ```text Mara enters the cellar. ``` Path B: ```text Mara remains upstairs. ``` An image generated for Path A must not appear as the current scene illustration on Path B. The asset may remain stored. It becomes inactive/disposable with the abandoned scene lineage. ## 18. Undo and Restore When the story is undone/restored: - story state changes immediately, - media is not authoritative, - old media remains attached to old scene/lineage, - UI should stop treating old media as current. Do not delete media automatically. ## 19. Media Cleanup Media consumes far more storage than text. A future cleanup feature should consider: - abandoned lineage, - age, - asset size, - user favorites, - checkpoint references, - regeneration ability. Potential actions: ```text Delete abandoned media Delete unstarred generations Keep final selections only ``` No automatic cleanup required initially. ## 20. User Selection When multiple generations exist: ```text Image A Image B Image C ``` The user may select one as: ```text preferred_asset = true ``` Other assets remain stored unless deleted. ## 21. Image Provider Interface Conceptual interface: ```text generate_image(scene_packet, settings) -> media_result ``` Provider capabilities should declare: ```yaml supports_seed: supports_negative_prompt: supports_reference_images: supports_character_reference: supports_controlnet: supports_inpainting: supports_upscale: ``` ## 22. Video Provider Interface Conceptual interface: ```text generate_video(scene_packet, settings) -> media_result ``` Potential capabilities: ```yaml supports_text_to_video: supports_image_to_video: supports_reference_frames: supports_audio: supports_seed: max_duration_seconds: ``` ## 23. Audio Provider Interface Conceptual: ```text generate_audio(scene_packet, settings) -> media_result ``` Potential uses: - tavern ambience, - rain, - machinery, - battle sounds. ## 24. TTS Provider Interface Conceptual: ```text synthesize_speech(text, voice_profile, settings) -> media_result ``` Potential future uses: - narrator voice, - NPC voices, - replaying dialogue. ## 24A. Speech-to-Text Provider Interface Conceptual: ```text transcribe_audio(audio_input, settings) -> transcription_result ``` Potential future uses: - dictate player actions, - dictate dialogue, - hands-free story input, - accessibility. Required semantic rule: > STT output is draft user input, not an accepted story event. Recommended workflow: ```text Microphone / local audio | v Local STT Provider | v Draft transcription | v User review/edit | v Normal story input submission ``` The user should be able to edit the transcription before it enters the authoritative transcript. The STT provider should remain local by default and should not require a cloud transcription API. ## 25. Provider Capability Discovery Providers should expose capabilities. Example: ```yaml provider_id: local-comfyui media_types: - image supports: seed: true reference_images: true inpainting: true ``` The UI should adapt based on capability. ## 26. Provider Configuration Provider config should be separate from campaign data where possible. Example: ```yaml provider_id: local-comfyui endpoint: http://127.0.0.1:8188 enabled: true ``` ## 27. Local-Only Requirement Preferred v1/future default: ```text Media provider endpoints must be loopback/local. ``` No cloud generation should be required. If remote media providers are ever added: - they must be explicit, - off by default, - clearly marked as data-leaving-machine behavior. ## 28. Provider Endpoint Validation Default allowed endpoints: ```text 127.0.0.1 localhost ``` Future advanced setting may allow: - LAN, - Tailscale, - trusted workstation. Not required initially. ## 29. Image Generation Example Story: ```text Aldric enters the Crooked Lantern during a storm. Mara stands behind the bar. ``` Scene Packet: ```text Location: old timber tavern Characters: Aldric Mara Weather: heavy rain outside Lighting: warm oil lamps Mood: tense arrival ``` The Image Provider turns this into provider-specific prompt/controls. ## 30. Video Generation Use Case User wants to convert: ```text Turns 210-215 ``` into a short battle clip. Workflow: ```text Select turn range | v Build multi-turn scene packet | v Extract action beats | v Create shot/sequence plan | v Video Provider ``` ## 31. Multi-Turn Scene Packet For video, include ordered beats. Example: ```yaml beats: - Aldric draws sword - guard lunges - Aldric blocks - lantern falls - room catches partial fire ``` Each beat should preserve: - characters, - location, - object state, - continuity. ## 32. Shot Planning A future Video Coordinator may generate: ```yaml shots: - wide establishing shot - medium combat shot - close-up of key falling - final wide shot ``` This is not a Story Engine responsibility. ## 33. Video Duration The video request should explicitly define: ```text target_duration ``` rather than assuming one turn equals one fixed duration. ## 34. Scene Compression Long turn ranges may need compression into: - action beats, - visual summary, - omitted dialogue. This derived representation should retain source turn IDs. ## 35. Media Does Not Change Canon A generated image may be wrong. Example: - wrong hair color, - extra sword, - wrong number of characters. The image does not alter story state. Correction options: - regenerate, - edit prompt, - inpaint, - reject asset. ## 36. Media Feedback The user may mark an asset: ```text accepted_visual ``` This means: - preferred depiction, - useful visual continuity reference. It still should not automatically override explicit canonical data. ## 37. Visual Canon Promotion Potential future feature: User explicitly selects: ```text Promote visual detail to canon ``` Example: - scar shape, - clothing color, - vehicle appearance. This must be explicit. Media output should never auto-promote visual details. ## 38. Reference Images Future image/video providers may accept: - character portrait, - location image, - item image. These assets should have: - local IDs, - provenance, - explicit role. Example: ```yaml reference_role: character_identity ``` ## 39. Character Consistency Media coordinator may use: - visual profile, - prior accepted portrait, - reference image, - seed, - adapter/LoRA where supported. The provider-specific method should remain outside Story Engine. ## 40. Location Consistency Same principle: - stable visual profile, - accepted reference image, - provider-specific conditioning. ## 41. Style Profiles Campaign may define: ```yaml visual_style: realism: grounded cinematic palette: muted era_accuracy: high ``` Style profile is campaign configuration. It should not alter story canon. ## 42. Prompt Templates Provider-specific prompt templates belong in the media subsystem. Example: ```text [visual style] [location] [characters] [action] [lighting] [continuity constraints] ``` Story engine should not contain Stable Diffusion syntax. ## 43. Negative Prompts If supported, Media Coordinator may generate: - no modern objects, - no duplicate characters, - no text overlays. This is provider-specific optional metadata. ## 44. Determinism Where provider supports seeds: Store seed. This enables: - recreation, - variations, - debugging. Determinism is helpful but not required across all providers. ## 45. Media Retries Retrying media generation should create a new job/asset. Do not overwrite prior asset. Example: ```text scene-42 image-a image-b image-c ``` ## 46. Media Edits Future: - inpaint, - upscale, - image-to-image, - video refinement. Each edit should preserve parent asset lineage. Conceptually: ```text asset B derived_from asset A ``` ## 47. Asset Lineage Potential fields: ```yaml parent_asset_id: operation: - regenerate - inpaint - upscale - animate ``` ## 48. Story-to-Media Provenance Every media asset should be traceable: ```text Campaign -> Branch -> Turn range -> Scene snapshot -> Media request -> Job -> Asset ``` ## 49. Media-to-Story Separation The reverse must not happen automatically: ```text Asset -X-> Story state ``` unless the user explicitly promotes information. ## 50. Failure Handling If media generation fails: - story remains unaffected, - scene remains valid, - job records failure, - user may retry, - no partial story mutation. ## 51. Provider Timeout Media jobs may be long. Coordinator should support: - status, - timeout, - cancellation, - retry. No need to block story interaction while media generates. ## 52. Asynchronous Design The architecture should assume media generation can happen asynchronously relative to story interaction. The user may continue the story while an image/video job runs. When complete: - asset attaches to source scene, - it should not become current merely because story has advanced. ## 53. Job Queue A simple local job queue may be needed. Conceptual: ```text pending running completed failed ``` Implementation may be: - database-backed, - in-process, - worker process. Do not introduce distributed infrastructure for v1. ## 54. Restart Recovery If app restarts during media job: Preferred: - mark interrupted jobs, - allow retry, - preserve completed files. Provider-specific resume is optional. ## 55. Storage Layout Recommended: ```text data/ campaigns/ / media/ images/ video/ audio/ thumbnails/ ``` Exact layout is implementation-specific. ## 56. File Naming Do not use raw user text as filenames. Use: - IDs, - hashes, - safe extensions. Original labels can exist in metadata. ## 57. Media Hashing Compute content hash. Useful for: - duplicate detection, - integrity, - export verification. ## 58. Export Campaign export should include: - selected media assets, - metadata, - provenance. Potential export modes: ```text Full No Media Selected Media Only ``` For v1, if media is not implemented, preserve schema compatibility. ## 59. Import Campaign import should restore: - asset metadata, - local file associations, - source scene references. Missing media files should not break story history. ## 60. Thumbnail Generation Thumbnails are derived cache. They may be recreated. Do not treat them as canonical assets. ## 61. Browser Delivery Media should be served through scoped local routes. Do not expose arbitrary filesystem paths. ## 62. Security Generated files are still untrusted browser content. Use: - correct MIME types, - safe content disposition, - no arbitrary executable serving, - no remote media loading by default. ## 63. Prompt Privacy Media prompts may contain: - hidden characters, - plot secrets, - campaign state. Therefore local-only media providers are preferred. ## 64. Hidden Information A scene illustration should not accidentally reveal narrator-only hidden canon unless the scene logically exposes it. Example: - hidden trap behind wall, - secret identity, - concealed character. Scene extractor should distinguish: - visible facts, - narrator-only facts. ## 65. Visible Scene State Media should generally receive only visually observable information plus necessary visual continuity data. It should not receive unrelated hidden plot details. ## 66. Audio Privacy TTS text may contain full dialogue. Keep speech generation local by default. ## 67. Voice Profiles Future: ```yaml voice_profile_id: character_id: provider: voice_name: settings: ``` Voice profile is presentation metadata. It is not story canon. ## 67A. Speech-to-Text Privacy and Input Semantics STT may receive live microphone audio or a local recorded clip. Security/privacy requirements: - audio remains local by default, - microphone access requires normal browser/user permission, - no automatic background recording, - recording state must be visibly indicated, - transcript remains editable before submission, - failed/partial transcription must not create a story turn, - raw audio retention should be optional and off by default unless needed for debugging or user-requested history. STT should not bypass the normal story-input validation and commit path. ## 68. Music Future ambient/music generation should be: - optional, - local, - scene-linked. Do not make it part of core story context. ## 69. User Controls Potential media UI: ```text Generate Image Generate Video Generate Audio Read Aloud Dictate Regenerate Use as Preferred Delete Show Provenance ``` Output-media controls should appear as optional scene actions. `Dictate` belongs near the story input field because STT produces draft user input. ## 70. Default Media Behavior Recommended default: ```text No automatic media generation. ``` Reason: - GPU cost, - storage, - latency, - user control. User explicitly requests media. ## 71. Optional Auto-Illustration Future setting: ```text Auto-generate one image at scene changes ``` Off by default. ## 72. Scene Change Detection Future media automation may detect: - new location, - major character entrance, - major action event, - chapter boundary. This remains optional. ## 73. Resource Coordination Local LLM and image/video models may compete for GPU memory. Media coordinator should eventually support: - queueing, - model unload/reload, - provider limits. Do not assume simultaneous execution is always possible. ## 74. Hardware Awareness Providers may expose: - VRAM requirement, - model availability, - estimated capability. The core story engine should not need this information. ## 75. Provider Errors Normalize provider errors. Example: ```yaml error_type: model_missing out_of_memory invalid_request provider_unreachable cancelled ``` ## 76. Model Discovery Future media UI may list local models available from provider. Do not auto-download them. ## 77. Model Downloads Same privacy rule as Ollama: - installation may use Internet, - normal generation should not require Internet, - no silent downloads. ## 78. Media Provider Registry Conceptual registry: ```yaml providers: - id: comfyui-local types: [image, video] - id: kokoro-local types: [tts] - id: local-stt types: [stt] ``` Exact products are not committed. ## 79. Open Dungeon Reuse Open Dungeon is especially relevant for: - local image generation, - character visual continuity, - image workflow UX. During Phase 0B, inspect: - how scene prompts are built, - how images are attached to story, - provider coupling, - whether media can be separated from destructive history model. Do not copy its persistence limitations into the core story architecture. ## 80. Gamentic Reuse Gamentic is relevant for: - provider abstraction, - asynchronous generation, - media job concepts, - local multimodal architecture. Use as a pattern reference, not a merged codebase. ## 81. Corvus Story Core Reuse Corvus may be useful for: - visual scene extraction, - TTS/ComfyUI integration patterns. Again, use concepts selectively. ## 82. V1 Physical Schema Decision Open question: Should media tables physically exist in v1? Recommended: > Include minimal media-ready schema/interfaces if inexpensive, but do not build provider implementation solely to justify them. Minimum useful v1 fields: - scene snapshot, - stable visual profiles, - media provider interface/types, - optional media asset table. Phase 0B should determine cost. ## 83. V1 Required Media Readiness Even without generation, v1 should preserve: - scene identity, - scene turn range, - character visual profiles, - location visual profiles, - branch-aware scene snapshots. This is sufficient to avoid architectural dead ends. ## 84. Future Image Acceptance Test Given an accepted scene: ```text Aldric and Mara examine the Silver Key in the tavern. ``` Generate image. Pass if: - only local provider used, - asset attached to correct scene, - provenance recorded, - story state unchanged, - retry creates new asset rather than overwriting. ## 85. Future Branch Media Acceptance Test Generate image on Path A. Restore checkpoint and create Path B. Pass if: - Path A image remains stored, - Path A image is not shown as current Path B media, - no automatic deletion occurs. ## 86. Future Video Acceptance Test Select 4-5 combat turns. Generate short local video. Pass if: - source turn range preserved, - action sequence matches accepted branch, - abandoned-branch events are not included, - media generation does not mutate story. ## 87. Phase 0B Validation Questions Codex should answer: 1. How does Open Dungeon attach generated images to story messages/scenes? 2. Is image generation coupled to linear/destructive message history? 3. Can its visual character continuity data be reused independently? 4. What provider assumptions are hardcoded? 5. Can ComfyUI/local provider calls run fully offline? 6. Does AI-DnD already have scene-like structured state suitable for media extraction? 7. Can scene snapshots be added without RPG schema coupling? 8. What parts of Gamentic's provider abstraction are worth reimplementing? 9. Should media asset tables exist in v1 or be deferred? 10. Can media generation remain entirely optional without special-casing core story logic? ## 88. Acceptance Criteria for Architecture The architecture passes if: - story engine functions with no media provider, - media requests derive from accepted story state, - assets are branch-aware, - assets preserve source scene/turn provenance, - media failure cannot corrupt story state, - provider-specific syntax stays outside Story Engine, - local providers can be substituted, - future image/video/audio/TTS types fit the output job/asset model, - future STT fits the same provider architecture while feeding editable draft input rather than story state, - generated media never automatically becomes canon, - abandoned-history media remains recoverable but inactive. ## 89. Current Recommendation Use this conceptual contract: ```text Authoritative Story | v Scene Snapshot | v Scene Packet | v Media Coordinator | +--> Local Image Provider +--> Local Video Provider +--> Local Audio Provider +--> Local TTS Provider | v Media Asset + Provenance ``` The media subsystem should depend on the story engine. The story engine should not depend on the media subsystem. That one-way dependency is the most important architectural requirement in this document.