# Adventure Storyteller — Media Extension Contract **Status:** v1.0 architecture contract — providers remain future/optional **Purpose:** Define the stable interfaces and data boundaries needed to add local image, video, and audio generation later without coupling media generation to the core story engine. ## 1. Design Goal The core storyteller must not depend on any specific media generator. The story system should remain fully functional with: ```text media_enabled = false ``` Media is an optional extension layer. The core rule is: > Story state is authoritative. Media is derived from story state and scene history. Generated media must never become the database of record for story facts. ## 2. Scope This document covers future: - still-image generation, - video generation, - audio ambience, - narration/TTS, - character voice, - speech-to-text (STT) for player input/dictation, - scene illustration, - multi-turn scene recap clips. It does not require any of these to ship in v1. The v1 requirement is architectural compatibility. ## 3. Architectural Separation Recommended architecture: ```text Story Engine | v Scene Extraction | v Scene Packet | v Media Coordinator | +--> Image Provider +--> Video Provider +--> Audio Provider +--> TTS Provider ``` The Story Engine should not directly call: - ComfyUI, - Stable Diffusion, - video pipelines, - TTS engines, - third-party media APIs. ## 4. Core Responsibilities ### Story Engine Responsible for: - accepted transcript, - canon, - state, - entities, - relationships, - scenes, - lineage, - checkpoints, - prompt provenance. ### Scene Extractor Responsible for turning accepted story state into a structured media-ready scene description. ### Media Coordinator Responsible for: - provider selection, - request normalization, - job creation, - status tracking, - retry, - cancellation, - output registration, - provenance. ### Media Provider Responsible for: - translating normalized request into provider-specific format, - invoking local generator, - returning output metadata. ## 5. Scene Snapshot Requirement The story system should persist a structured scene snapshot for relevant turns. Conceptual fields: ```yaml scene_id: campaign_id: branch_id: start_turn_id: end_turn_id: location_id: characters_present: objects_present: time_of_day: weather: lighting: mood: visual_notes: action_summary: dialogue_summary: camera_hint: created_at: ``` Not every field is required for every scene. ## 6. Scene Snapshot Authority A scene snapshot is derived from accepted story state. If it conflicts with canonical state: - canonical state wins, - scene snapshot should be regenerated or corrected. Scene snapshots from abandoned history remain associated with that abandoned lineage. ## 7. Visual Character Profiles Characters may have optional stable visual descriptors. Example: ```yaml character_id: mara apparent_age: early 40s build: sturdy hair: dark auburn eyes: gray clothing_baseline: practical innkeeper clothing distinctive_features: - small burn scar on right forearm style_notes: - grounded realism ``` These descriptors should support continuity across generated images. ## 8. Visual Location Profiles Locations may have optional stable descriptors. Example: ```yaml location_id: crooked-lantern architecture: timber-framed roadside tavern interior: - stone hearth - dark beams - shared wooden tables lighting: - candles - oil lamps visual_identity: - warm but worn ``` ## 9. Visual Item Profiles Important recurring items may also have visual descriptors. Example: ```yaml item_id: silver-key material: silver shape: small old-fashioned key marking: broken-circle symbol ``` This is optional but useful for continuity. ## 10. Scene Packet The Media Coordinator should consume a normalized Scene Packet rather than raw transcript text. Example: ```yaml scene_id: scene-142 campaign_id: continuity-test turn_range: start: 140 end: 143 location: name: Crooked Lantern Tavern visual_profile: ... characters: - name: Aldric visual_profile: ... current_condition: ... - name: Mara visual_profile: ... objects: - Silver Key action_summary: > Aldric places the silver key on the table while Mara studies the broken-circle symbol. mood: tense curiosity lighting: dim oil-lamp light time_of_day: night continuity_constraints: - Aldric still owns the key - Mara has not yet entered the cellar - no modern objects ``` ## 11. Scene Packet Purpose The Scene Packet provides: - stable provider-independent input, - continuity, - exact lineage, - provenance, - repeatability, - easier testing, - future provider switching. ## 12. Raw Transcript Access Providers should not normally receive the entire campaign transcript. Preferred: - structured scene packet, - selected supporting recent text, - only the minimum needed. This improves: - privacy, - consistency, - prompt size, - provider portability. ## 13. Media Request Conceptual request: ```yaml media_request_id: campaign_id: scene_id: media_type: provider_id: requested_by: created_at: settings: prompt_override: reference_assets: ``` Media types: ```text image video audio tts stt ``` ## 14. Media Job Each generation attempt should create a job record. Conceptual fields: ```yaml job_id: media_request_id: provider_id: status: started_at: completed_at: error: provider_settings: seed: model: model_version: input_hash: ``` Statuses: ```text queued running completed failed cancelled ``` ## 15. Media Asset Successful jobs produce media assets. Conceptual fields: ```yaml asset_id: campaign_id: scene_id: job_id: media_type: local_path: mime_type: width: height: duration: file_size: content_hash: created_at: ``` Optional: - thumbnail, - codec, - frame rate, - audio channels. ## 16. Asset Provenance Every asset should retain: - source scene, - source turn range, - active branch/lineage at generation, - provider, - model, - model version, - generation settings, - seed if available, - normalized scene packet, - final rendered prompt if applicable. This allows later explanation: ```text What story state produced this image? ``` ## 17. Branch Awareness Media must be branch-aware. Example: Path A: ```text Mara enters the cellar. ``` Path B: ```text Mara remains upstairs. ``` An image generated for Path A must not appear as the current scene illustration on Path B. The asset may remain stored. It becomes inactive/disposable with the abandoned scene lineage. ## 18. Undo and Restore When the story is undone/restored: - story state changes immediately, - media is not authoritative, - old media remains attached to old scene/lineage, - UI should stop treating old media as current. Do not delete media automatically. ## 19. Media Cleanup Media consumes far more storage than text. A future cleanup feature should consider: - abandoned lineage, - age, - asset size, - user favorites, - checkpoint references, - regeneration ability. Potential actions: ```text Delete abandoned media Delete unstarred generations Keep final selections only ``` No automatic cleanup required initially. ## 20. User Selection When multiple generations exist: ```text Image A Image B Image C ``` The user may select one as: ```text preferred_asset = true ``` Other assets remain stored unless deleted. ## 21. Image Provider Interface Conceptual interface: ```text generate_image(scene_packet, settings) -> media_result ``` Provider capabilities should declare: ```yaml supports_seed: supports_negative_prompt: supports_reference_images: supports_character_reference: supports_controlnet: supports_inpainting: supports_upscale: ``` ## 22. Video Provider Interface Conceptual interface: ```text generate_video(scene_packet, settings) -> media_result ``` Potential capabilities: ```yaml supports_text_to_video: supports_image_to_video: supports_reference_frames: supports_audio: supports_seed: max_duration_seconds: ``` ## 23. Audio Provider Interface Conceptual: ```text generate_audio(scene_packet, settings) -> media_result ``` Potential uses: - tavern ambience, - rain, - machinery, - battle sounds. ## 24. TTS Provider Interface Conceptual: ```text synthesize_speech(text, voice_profile, settings) -> media_result ``` Potential future uses: - narrator voice, - NPC voices, - replaying dialogue. ## 24A. Speech-to-Text Provider Interface Conceptual: ```text transcribe_audio(audio_input, settings) -> transcription_result ``` Potential future uses: - dictate player actions, - dictate dialogue, - hands-free story input, - accessibility. Required semantic rule: > STT output is draft user input, not an accepted story event. Recommended workflow: ```text Microphone / local audio | v Local STT Provider | v Draft transcription | v User review/edit | v Normal story input submission ``` The user should be able to edit the transcription before it enters the authoritative transcript. The STT provider should remain local by default and should not require a cloud transcription API. ## 25. Provider Capability Discovery Providers should expose capabilities. Example: ```yaml provider_id: local-comfyui media_types: - image supports: seed: true reference_images: true inpainting: true ``` The UI should adapt based on capability. ## 26. Provider Configuration Provider config should be separate from campaign data where possible. Example: ```yaml provider_id: local-comfyui endpoint: http://127.0.0.1:8188 enabled: true ``` ## 27. Local-Only Requirement Preferred v1/future default: ```text Media provider endpoints must be loopback/local. ``` No cloud generation should be required. If remote media providers are ever added: - they must be explicit, - off by default, - clearly marked as data-leaving-machine behavior. ## 28. Provider Endpoint Validation Default allowed endpoints: ```text 127.0.0.1 localhost ``` Future advanced setting may allow: - LAN, - Tailscale, - trusted workstation. Not required initially. ## 29. Image Generation Example Story: ```text Aldric enters the Crooked Lantern during a storm. Mara stands behind the bar. ``` Scene Packet: ```text Location: old timber tavern Characters: Aldric Mara Weather: heavy rain outside Lighting: warm oil lamps Mood: tense arrival ``` The Image Provider turns this into provider-specific prompt/controls. ## 30. Video Generation Use Case User wants to convert: ```text Turns 210-215 ``` into a short battle clip. Workflow: ```text Select turn range | v Build multi-turn scene packet | v Extract action beats | v Create shot/sequence plan | v Video Provider ``` ## 31. Multi-Turn Scene Packet For video, include ordered beats. Example: ```yaml beats: - Aldric draws sword - guard lunges - Aldric blocks - lantern falls - room catches partial fire ``` Each beat should preserve: - characters, - location, - object state, - continuity. ## 32. Shot Planning A future Video Coordinator may generate: ```yaml shots: - wide establishing shot - medium combat shot - close-up of key falling - final wide shot ``` This is not a Story Engine responsibility. ## 33. Video Duration The video request should explicitly define: ```text target_duration ``` rather than assuming one turn equals one fixed duration. ## 34. Scene Compression Long turn ranges may need compression into: - action beats, - visual summary, - omitted dialogue. This derived representation should retain source turn IDs. ## 35. Media Does Not Change Canon A generated image may be wrong. Example: - wrong hair color, - extra sword, - wrong number of characters. The image does not alter story state. Correction options: - regenerate, - edit prompt, - inpaint, - reject asset. ## 36. Media Feedback The user may mark an asset: ```text accepted_visual ``` This means: - preferred depiction, - useful visual continuity reference. It still should not automatically override explicit canonical data. ## 37. Visual Canon Promotion Potential future feature: User explicitly selects: ```text Promote visual detail to canon ``` Example: - scar shape, - clothing color, - vehicle appearance. This must be explicit. Media output should never auto-promote visual details. ## 38. Reference Images Future image/video providers may accept: - character portrait, - location image, - item image. These assets should have: - local IDs, - provenance, - explicit role. Example: ```yaml reference_role: character_identity ``` ## 39. Character Consistency Media coordinator may use: - visual profile, - prior accepted portrait, - reference image, - seed, - adapter/LoRA where supported. The provider-specific method should remain outside Story Engine. ## 40. Location Consistency Same principle: - stable visual profile, - accepted reference image, - provider-specific conditioning. ## 41. Style Profiles Campaign may define: ```yaml visual_style: realism: grounded cinematic palette: muted era_accuracy: high ``` Style profile is campaign configuration. It should not alter story canon. ## 42. Prompt Templates Provider-specific prompt templates belong in the media subsystem. Example: ```text [visual style] [location] [characters] [action] [lighting] [continuity constraints] ``` Story engine should not contain Stable Diffusion syntax. ## 43. Negative Prompts If supported, Media Coordinator may generate: - no modern objects, - no duplicate characters, - no text overlays. This is provider-specific optional metadata. ## 44. Determinism Where provider supports seeds: Store seed. This enables: - recreation, - variations, - debugging. Determinism is helpful but not required across all providers. ## 45. Media Retries Retrying media generation should create a new job/asset. Do not overwrite prior asset. Example: ```text scene-42 image-a image-b image-c ``` ## 46. Media Edits Future: - inpaint, - upscale, - image-to-image, - video refinement. Each edit should preserve parent asset lineage. Conceptually: ```text asset B derived_from asset A ``` ## 47. Asset Lineage Potential fields: ```yaml parent_asset_id: operation: - regenerate - inpaint - upscale - animate ``` ## 48. Story-to-Media Provenance Every media asset should be traceable: ```text Campaign -> Branch -> Turn range -> Scene snapshot -> Media request -> Job -> Asset ``` ## 49. Media-to-Story Separation The reverse must not happen automatically: ```text Asset -X-> Story state ``` unless the user explicitly promotes information. ## 50. Failure Handling If media generation fails: - story remains unaffected, - scene remains valid, - job records failure, - user may retry, - no partial story mutation. ## 51. Provider Timeout Media jobs may be long. Coordinator should support: - status, - timeout, - cancellation, - retry. No need to block story interaction while media generates. ## 52. Asynchronous Design The architecture should assume media generation can happen asynchronously relative to story interaction. The user may continue the story while an image/video job runs. When complete: - asset attaches to source scene, - it should not become current merely because story has advanced. ## 53. Job Queue A simple local job queue may be needed. Conceptual: ```text pending running completed failed ``` Implementation may be: - database-backed, - in-process, - worker process. Do not introduce distributed infrastructure for v1. ## 54. Restart Recovery If app restarts during media job: Preferred: - mark interrupted jobs, - allow retry, - preserve completed files. Provider-specific resume is optional. ## 55. Storage Layout Recommended: ```text data/ campaigns/ / media/ images/ video/ audio/ thumbnails/ ``` Exact layout is implementation-specific. ## 56. File Naming Do not use raw user text as filenames. Use: - IDs, - hashes, - safe extensions. Original labels can exist in metadata. ## 57. Media Hashing Compute content hash. Useful for: - duplicate detection, - integrity, - export verification. ## 58. Export Campaign export should include: - selected media assets, - metadata, - provenance. Potential export modes: ```text Full No Media Selected Media Only ``` For v1, if media is not implemented, preserve schema compatibility. ## 59. Import Campaign import should restore: - asset metadata, - local file associations, - source scene references. Missing media files should not break story history. ## 60. Thumbnail Generation Thumbnails are derived cache. They may be recreated. Do not treat them as canonical assets. ## 61. Browser Delivery Media should be served through scoped local routes. Do not expose arbitrary filesystem paths. ## 62. Security Generated files are still untrusted browser content. Use: - correct MIME types, - safe content disposition, - no arbitrary executable serving, - no remote media loading by default. ## 63. Prompt Privacy Media prompts may contain: - hidden characters, - plot secrets, - campaign state. Therefore local-only media providers are preferred. ## 64. Hidden Information A scene illustration should not accidentally reveal narrator-only hidden canon unless the scene logically exposes it. Example: - hidden trap behind wall, - secret identity, - concealed character. Scene extractor should distinguish: - visible facts, - narrator-only facts. ## 65. Visible Scene State Media should generally receive only visually observable information plus necessary visual continuity data. It should not receive unrelated hidden plot details. ## 66. Audio Privacy TTS text may contain full dialogue. Keep speech generation local by default. ## 67. Voice Profiles Future: ```yaml voice_profile_id: character_id: provider: voice_name: settings: ``` Voice profile is presentation metadata. It is not story canon. ## 67A. Speech-to-Text Privacy and Input Semantics STT may receive live microphone audio or a local recorded clip. Security/privacy requirements: - audio remains local by default, - microphone access requires normal browser/user permission, - no automatic background recording, - recording state must be visibly indicated, - transcript remains editable before submission, - failed/partial transcription must not create a story turn, - raw audio retention should be optional and off by default unless needed for debugging or user-requested history. STT should not bypass the normal story-input validation and commit path. ## 68. Music Future ambient/music generation should be: - optional, - local, - scene-linked. Do not make it part of core story context. ## 69. User Controls Potential media UI: ```text Generate Image Generate Video Generate Audio Read Aloud Dictate Regenerate Use as Preferred Delete Show Provenance ``` Output-media controls should appear as optional scene actions. `Dictate` belongs near the story input field because STT produces draft user input. ## 70. Default Media Behavior Recommended default: ```text No automatic media generation. ``` Reason: - GPU cost, - storage, - latency, - user control. User explicitly requests media. ## 71. Optional Auto-Illustration Future setting: ```text Auto-generate one image at scene changes ``` Off by default. ## 72. Scene Change Detection Future media automation may detect: - new location, - major character entrance, - major action event, - chapter boundary. This remains optional. ## 73. Resource Coordination Local LLM and image/video models may compete for GPU memory. Media coordinator should eventually support: - queueing, - model unload/reload, - provider limits. Do not assume simultaneous execution is always possible. ## 74. Hardware Awareness Providers may expose: - VRAM requirement, - model availability, - estimated capability. The core story engine should not need this information. ## 75. Provider Errors Normalize provider errors. Example: ```yaml error_type: model_missing out_of_memory invalid_request provider_unreachable cancelled ``` ## 76. Model Discovery Future media UI may list local models available from provider. Do not auto-download them. ## 77. Model Downloads Same privacy rule as Ollama: - installation may use Internet, - normal generation should not require Internet, - no silent downloads. ## 78. Media Provider Registry Conceptual registry: ```yaml providers: - id: comfyui-local types: [image, video] - id: kokoro-local types: [tts] - id: local-stt types: [stt] ``` Exact products are not committed. ## 79. Open Dungeon Reuse Open Dungeon is especially relevant for: - local image generation, - character visual continuity, - image workflow UX. Phase 0B confirmed Open Dungeon should remain a media/UX reference rather than the production base. Study its local image worker protocol, scene/image attachment, and character visual-continuity ideas during the future media implementation stage if useful. Do not copy its linear/destructive persistence assumptions into the core story architecture. ## 80. Gamentic Reuse Gamentic is relevant for: - provider abstraction, - asynchronous generation, - media job concepts, - local multimodal architecture. Use as a pattern reference, not a merged codebase. ## 81. Corvus Story Core Reuse Corvus may be useful for: - visual scene extraction, - TTS/ComfyUI integration patterns. Again, use concepts selectively. ## 82. V1 Physical Schema Direction The v1 requirement is architectural compatibility, not media generation. Required now: - scene snapshots, - stable optional visual profiles, - provider-neutral request/result types or equivalent interface contract. A physical media job/asset table may be introduced in the dedicated future-media-hooks milestone if it is inexpensive and useful for schema stability. Its physical presence is an implementation detail, not a prerequisite for story functionality. Do not implement a provider merely to justify a table. ## 83. V1 Required Media Readiness Even without generation, v1 should preserve: - scene identity, - scene turn range, - character visual profiles, - location visual profiles, - branch-aware scene snapshots. This is sufficient to avoid architectural dead ends. ## 84. Future Image Acceptance Test Given an accepted scene: ```text Aldric and Mara examine the Silver Key in the tavern. ``` Generate image. Pass if: - only local provider used, - asset attached to correct scene, - provenance recorded, - story state unchanged, - retry creates new asset rather than overwriting. ## 85. Future Branch Media Acceptance Test Generate image on Path A. Restore checkpoint and create Path B. Pass if: - Path A image remains stored, - Path A image is not shown as current Path B media, - no automatic deletion occurs. ## 86. Future Video Acceptance Test Select 4-5 combat turns. Generate short local video. Pass if: - source turn range preserved, - action sequence matches accepted branch, - abandoned-branch events are not included, - media generation does not mutate story. ## 87. Phase 0B Findings Applied - Open Dungeon has useful local image-generation and visual-continuity concepts, but they are not a reason to use it as the production fork. - The production base has no required media subsystem today; this is acceptable for v1. - Media must remain derived from accepted scene/story state and branch-aware. - Provider portability for Open Dungeon's image worker is deferred until image generation is actually scheduled. - Scene snapshots and visual continuity fields are sufficient near-term architecture commitments. - TTS/STT/video remain future providers behind the same local optional boundary. ## 88. Acceptance Criteria for Architecture The architecture passes if: - story engine functions with no media provider, - media requests derive from accepted story state, - assets are branch-aware, - assets preserve source scene/turn provenance, - media failure cannot corrupt story state, - provider-specific syntax stays outside Story Engine, - local providers can be substituted, - future image/video/audio/TTS types fit the output job/asset model, - future STT fits the provider architecture while feeding editable draft input rather than story state, - generated media never automatically becomes canon, - abandoned-history media remains recoverable but inactive. ## 89. Selected Contract Use this conceptual contract: ```text Authoritative Story | v Scene Snapshot / Scene Packet | v Optional Media Coordinator | +--> Image Provider +--> Video Provider +--> Audio Provider +--> TTS Provider +--> STT Provider (draft input path) ``` No production media provider is required for v1.