Files
interactive-story/planning/MEDIA-EXTENSION-CONTRACT.md

24 KiB

Adventure Storyteller — Media Extension Contract

Status: v1.0 architecture contract — providers remain future/optional
Purpose: Define the stable interfaces and data boundaries needed to add local image, video, and audio generation later without coupling media generation to the core story engine.

1. Design Goal

The core storyteller must not depend on any specific media generator.

The story system should remain fully functional with:

media_enabled = false

Media is an optional extension layer.

The core rule is:

Story state is authoritative. Media is derived from story state and scene history.

Generated media must never become the database of record for story facts.

2. Scope

This document covers future:

  • still-image generation,
  • video generation,
  • audio ambience,
  • narration/TTS,
  • character voice,
  • speech-to-text (STT) for player input/dictation,
  • scene illustration,
  • multi-turn scene recap clips.

It does not require any of these to ship in v1.

The v1 requirement is architectural compatibility.

3. Architectural Separation

Recommended architecture:

Story Engine
    |
    v
Scene Extraction
    |
    v
Scene Packet
    |
    v
Media Coordinator
    |
    +--> Image Provider
    +--> Video Provider
    +--> Audio Provider
    +--> TTS Provider

The Story Engine should not directly call:

  • ComfyUI,
  • Stable Diffusion,
  • video pipelines,
  • TTS engines,
  • third-party media APIs.

4. Core Responsibilities

Story Engine

Responsible for:

  • accepted transcript,
  • canon,
  • state,
  • entities,
  • relationships,
  • scenes,
  • lineage,
  • checkpoints,
  • prompt provenance.

Scene Extractor

Responsible for turning accepted story state into a structured media-ready scene description.

Media Coordinator

Responsible for:

  • provider selection,
  • request normalization,
  • job creation,
  • status tracking,
  • retry,
  • cancellation,
  • output registration,
  • provenance.

Media Provider

Responsible for:

  • translating normalized request into provider-specific format,
  • invoking local generator,
  • returning output metadata.

5. Scene Snapshot Requirement

The story system should persist a structured scene snapshot for relevant turns.

Conceptual fields:

scene_id:
campaign_id:
branch_id:
start_turn_id:
end_turn_id:
location_id:
characters_present:
objects_present:
time_of_day:
weather:
lighting:
mood:
visual_notes:
action_summary:
dialogue_summary:
camera_hint:
created_at:

Not every field is required for every scene.

6. Scene Snapshot Authority

A scene snapshot is derived from accepted story state.

If it conflicts with canonical state:

  • canonical state wins,
  • scene snapshot should be regenerated or corrected.

Scene snapshots from abandoned history remain associated with that abandoned lineage.

7. Visual Character Profiles

Characters may have optional stable visual descriptors.

Example:

character_id: mara
apparent_age: early 40s
build: sturdy
hair: dark auburn
eyes: gray
clothing_baseline: practical innkeeper clothing
distinctive_features:
  - small burn scar on right forearm
style_notes:
  - grounded realism

These descriptors should support continuity across generated images.

8. Visual Location Profiles

Locations may have optional stable descriptors.

Example:

location_id: crooked-lantern
architecture: timber-framed roadside tavern
interior:
  - stone hearth
  - dark beams
  - shared wooden tables
lighting:
  - candles
  - oil lamps
visual_identity:
  - warm but worn

9. Visual Item Profiles

Important recurring items may also have visual descriptors.

Example:

item_id: silver-key
material: silver
shape: small old-fashioned key
marking: broken-circle symbol

This is optional but useful for continuity.

10. Scene Packet

The Media Coordinator should consume a normalized Scene Packet rather than raw transcript text.

Example:

scene_id: scene-142
campaign_id: continuity-test
turn_range:
  start: 140
  end: 143

location:
  name: Crooked Lantern Tavern
  visual_profile: ...

characters:
  - name: Aldric
    visual_profile: ...
    current_condition: ...
  - name: Mara
    visual_profile: ...

objects:
  - Silver Key

action_summary: >
  Aldric places the silver key on the table while Mara studies
  the broken-circle symbol.

mood: tense curiosity
lighting: dim oil-lamp light
time_of_day: night

continuity_constraints:
  - Aldric still owns the key
  - Mara has not yet entered the cellar
  - no modern objects

11. Scene Packet Purpose

The Scene Packet provides:

  • stable provider-independent input,
  • continuity,
  • exact lineage,
  • provenance,
  • repeatability,
  • easier testing,
  • future provider switching.

12. Raw Transcript Access

Providers should not normally receive the entire campaign transcript.

Preferred:

  • structured scene packet,
  • selected supporting recent text,
  • only the minimum needed.

This improves:

  • privacy,
  • consistency,
  • prompt size,
  • provider portability.

13. Media Request

Conceptual request:

media_request_id:
campaign_id:
scene_id:
media_type:
provider_id:
requested_by:
created_at:
settings:
prompt_override:
reference_assets:

Media types:

image
video
audio
tts
stt

14. Media Job

Each generation attempt should create a job record.

Conceptual fields:

job_id:
media_request_id:
provider_id:
status:
started_at:
completed_at:
error:
provider_settings:
seed:
model:
model_version:
input_hash:

Statuses:

queued
running
completed
failed
cancelled

15. Media Asset

Successful jobs produce media assets.

Conceptual fields:

asset_id:
campaign_id:
scene_id:
job_id:
media_type:
local_path:
mime_type:
width:
height:
duration:
file_size:
content_hash:
created_at:

Optional:

  • thumbnail,
  • codec,
  • frame rate,
  • audio channels.

16. Asset Provenance

Every asset should retain:

  • source scene,
  • source turn range,
  • active branch/lineage at generation,
  • provider,
  • model,
  • model version,
  • generation settings,
  • seed if available,
  • normalized scene packet,
  • final rendered prompt if applicable.

This allows later explanation:

What story state produced this image?

17. Branch Awareness

Media must be branch-aware.

Example:

Path A:

Mara enters the cellar.

Path B:

Mara remains upstairs.

An image generated for Path A must not appear as the current scene illustration on Path B.

The asset may remain stored.

It becomes inactive/disposable with the abandoned scene lineage.

18. Undo and Restore

When the story is undone/restored:

  • story state changes immediately,
  • media is not authoritative,
  • old media remains attached to old scene/lineage,
  • UI should stop treating old media as current.

Do not delete media automatically.

19. Media Cleanup

Media consumes far more storage than text.

A future cleanup feature should consider:

  • abandoned lineage,
  • age,
  • asset size,
  • user favorites,
  • checkpoint references,
  • regeneration ability.

Potential actions:

Delete abandoned media
Delete unstarred generations
Keep final selections only

No automatic cleanup required initially.

20. User Selection

When multiple generations exist:

Image A
Image B
Image C

The user may select one as:

preferred_asset = true

Other assets remain stored unless deleted.

21. Image Provider Interface

Conceptual interface:

generate_image(scene_packet, settings) -> media_result

Provider capabilities should declare:

supports_seed:
supports_negative_prompt:
supports_reference_images:
supports_character_reference:
supports_controlnet:
supports_inpainting:
supports_upscale:

22. Video Provider Interface

Conceptual interface:

generate_video(scene_packet, settings) -> media_result

Potential capabilities:

supports_text_to_video:
supports_image_to_video:
supports_reference_frames:
supports_audio:
supports_seed:
max_duration_seconds:

23. Audio Provider Interface

Conceptual:

generate_audio(scene_packet, settings) -> media_result

Potential uses:

  • tavern ambience,
  • rain,
  • machinery,
  • battle sounds.

24. TTS Provider Interface

Conceptual:

synthesize_speech(text, voice_profile, settings) -> media_result

Potential future uses:

  • narrator voice,
  • NPC voices,
  • replaying dialogue.

24A. Speech-to-Text Provider Interface

Conceptual:

transcribe_audio(audio_input, settings) -> transcription_result

Potential future uses:

  • dictate player actions,
  • dictate dialogue,
  • hands-free story input,
  • accessibility.

Required semantic rule:

STT output is draft user input, not an accepted story event.

Recommended workflow:

Microphone / local audio
    |
    v
Local STT Provider
    |
    v
Draft transcription
    |
    v
User review/edit
    |
    v
Normal story input submission

The user should be able to edit the transcription before it enters the authoritative transcript.

The STT provider should remain local by default and should not require a cloud transcription API.

25. Provider Capability Discovery

Providers should expose capabilities.

Example:

provider_id: local-comfyui
media_types:
  - image
supports:
  seed: true
  reference_images: true
  inpainting: true

The UI should adapt based on capability.

26. Provider Configuration

Provider config should be separate from campaign data where possible.

Example:

provider_id: local-comfyui
endpoint: http://127.0.0.1:8188
enabled: true

27. Local-Only Requirement

Preferred v1/future default:

Media provider endpoints must be loopback/local.

No cloud generation should be required.

If remote media providers are ever added:

  • they must be explicit,
  • off by default,
  • clearly marked as data-leaving-machine behavior.

28. Provider Endpoint Validation

Default allowed endpoints:

127.0.0.1
localhost

Future advanced setting may allow:

  • LAN,
  • Tailscale,
  • trusted workstation.

Not required initially.

29. Image Generation Example

Story:

Aldric enters the Crooked Lantern during a storm.
Mara stands behind the bar.

Scene Packet:

Location:
old timber tavern

Characters:
Aldric
Mara

Weather:
heavy rain outside

Lighting:
warm oil lamps

Mood:
tense arrival

The Image Provider turns this into provider-specific prompt/controls.

30. Video Generation Use Case

User wants to convert:

Turns 210-215

into a short battle clip.

Workflow:

Select turn range
   |
   v
Build multi-turn scene packet
   |
   v
Extract action beats
   |
   v
Create shot/sequence plan
   |
   v
Video Provider

31. Multi-Turn Scene Packet

For video, include ordered beats.

Example:

beats:
  - Aldric draws sword
  - guard lunges
  - Aldric blocks
  - lantern falls
  - room catches partial fire

Each beat should preserve:

  • characters,
  • location,
  • object state,
  • continuity.

32. Shot Planning

A future Video Coordinator may generate:

shots:
  - wide establishing shot
  - medium combat shot
  - close-up of key falling
  - final wide shot

This is not a Story Engine responsibility.

33. Video Duration

The video request should explicitly define:

target_duration

rather than assuming one turn equals one fixed duration.

34. Scene Compression

Long turn ranges may need compression into:

  • action beats,
  • visual summary,
  • omitted dialogue.

This derived representation should retain source turn IDs.

35. Media Does Not Change Canon

A generated image may be wrong.

Example:

  • wrong hair color,
  • extra sword,
  • wrong number of characters.

The image does not alter story state.

Correction options:

  • regenerate,
  • edit prompt,
  • inpaint,
  • reject asset.

36. Media Feedback

The user may mark an asset:

accepted_visual

This means:

  • preferred depiction,
  • useful visual continuity reference.

It still should not automatically override explicit canonical data.

37. Visual Canon Promotion

Potential future feature:

User explicitly selects:

Promote visual detail to canon

Example:

  • scar shape,
  • clothing color,
  • vehicle appearance.

This must be explicit.

Media output should never auto-promote visual details.

38. Reference Images

Future image/video providers may accept:

  • character portrait,
  • location image,
  • item image.

These assets should have:

  • local IDs,
  • provenance,
  • explicit role.

Example:

reference_role: character_identity

39. Character Consistency

Media coordinator may use:

  • visual profile,
  • prior accepted portrait,
  • reference image,
  • seed,
  • adapter/LoRA where supported.

The provider-specific method should remain outside Story Engine.

40. Location Consistency

Same principle:

  • stable visual profile,
  • accepted reference image,
  • provider-specific conditioning.

41. Style Profiles

Campaign may define:

visual_style:
  realism: grounded cinematic
  palette: muted
  era_accuracy: high

Style profile is campaign configuration.

It should not alter story canon.

42. Prompt Templates

Provider-specific prompt templates belong in the media subsystem.

Example:

[visual style]
[location]
[characters]
[action]
[lighting]
[continuity constraints]

Story engine should not contain Stable Diffusion syntax.

43. Negative Prompts

If supported, Media Coordinator may generate:

  • no modern objects,
  • no duplicate characters,
  • no text overlays.

This is provider-specific optional metadata.

44. Determinism

Where provider supports seeds:

Store seed.

This enables:

  • recreation,
  • variations,
  • debugging.

Determinism is helpful but not required across all providers.

45. Media Retries

Retrying media generation should create a new job/asset.

Do not overwrite prior asset.

Example:

scene-42
  image-a
  image-b
  image-c

46. Media Edits

Future:

  • inpaint,
  • upscale,
  • image-to-image,
  • video refinement.

Each edit should preserve parent asset lineage.

Conceptually:

asset B derived_from asset A

47. Asset Lineage

Potential fields:

parent_asset_id:
operation:
  - regenerate
  - inpaint
  - upscale
  - animate

48. Story-to-Media Provenance

Every media asset should be traceable:

Campaign
  -> Branch
  -> Turn range
  -> Scene snapshot
  -> Media request
  -> Job
  -> Asset

49. Media-to-Story Separation

The reverse must not happen automatically:

Asset
  -X-> Story state

unless the user explicitly promotes information.

50. Failure Handling

If media generation fails:

  • story remains unaffected,
  • scene remains valid,
  • job records failure,
  • user may retry,
  • no partial story mutation.

51. Provider Timeout

Media jobs may be long.

Coordinator should support:

  • status,
  • timeout,
  • cancellation,
  • retry.

No need to block story interaction while media generates.

52. Asynchronous Design

The architecture should assume media generation can happen asynchronously relative to story interaction.

The user may continue the story while an image/video job runs.

When complete:

  • asset attaches to source scene,
  • it should not become current merely because story has advanced.

53. Job Queue

A simple local job queue may be needed.

Conceptual:

pending
running
completed
failed

Implementation may be:

  • database-backed,
  • in-process,
  • worker process.

Do not introduce distributed infrastructure for v1.

54. Restart Recovery

If app restarts during media job:

Preferred:

  • mark interrupted jobs,
  • allow retry,
  • preserve completed files.

Provider-specific resume is optional.

55. Storage Layout

Recommended:

data/
  campaigns/
    <campaign-id>/
      media/
        images/
        video/
        audio/
        thumbnails/

Exact layout is implementation-specific.

56. File Naming

Do not use raw user text as filenames.

Use:

  • IDs,
  • hashes,
  • safe extensions.

Original labels can exist in metadata.

57. Media Hashing

Compute content hash.

Useful for:

  • duplicate detection,
  • integrity,
  • export verification.

58. Export

Campaign export should include:

  • selected media assets,
  • metadata,
  • provenance.

Potential export modes:

Full
No Media
Selected Media Only

For v1, if media is not implemented, preserve schema compatibility.

59. Import

Campaign import should restore:

  • asset metadata,
  • local file associations,
  • source scene references.

Missing media files should not break story history.

60. Thumbnail Generation

Thumbnails are derived cache.

They may be recreated.

Do not treat them as canonical assets.

61. Browser Delivery

Media should be served through scoped local routes.

Do not expose arbitrary filesystem paths.

62. Security

Generated files are still untrusted browser content.

Use:

  • correct MIME types,
  • safe content disposition,
  • no arbitrary executable serving,
  • no remote media loading by default.

63. Prompt Privacy

Media prompts may contain:

  • hidden characters,
  • plot secrets,
  • campaign state.

Therefore local-only media providers are preferred.

64. Hidden Information

A scene illustration should not accidentally reveal narrator-only hidden canon unless the scene logically exposes it.

Example:

  • hidden trap behind wall,
  • secret identity,
  • concealed character.

Scene extractor should distinguish:

  • visible facts,
  • narrator-only facts.

65. Visible Scene State

Media should generally receive only visually observable information plus necessary visual continuity data.

It should not receive unrelated hidden plot details.

66. Audio Privacy

TTS text may contain full dialogue.

Keep speech generation local by default.

67. Voice Profiles

Future:

voice_profile_id:
character_id:
provider:
voice_name:
settings:

Voice profile is presentation metadata.

It is not story canon.

67A. Speech-to-Text Privacy and Input Semantics

STT may receive live microphone audio or a local recorded clip.

Security/privacy requirements:

  • audio remains local by default,
  • microphone access requires normal browser/user permission,
  • no automatic background recording,
  • recording state must be visibly indicated,
  • transcript remains editable before submission,
  • failed/partial transcription must not create a story turn,
  • raw audio retention should be optional and off by default unless needed for debugging or user-requested history.

STT should not bypass the normal story-input validation and commit path.

68. Music

Future ambient/music generation should be:

  • optional,
  • local,
  • scene-linked.

Do not make it part of core story context.

69. User Controls

Potential media UI:

Generate Image
Generate Video
Generate Audio
Read Aloud
Dictate
Regenerate
Use as Preferred
Delete
Show Provenance

Output-media controls should appear as optional scene actions. Dictate belongs near the story input field because STT produces draft user input.

70. Default Media Behavior

Recommended default:

No automatic media generation.

Reason:

  • GPU cost,
  • storage,
  • latency,
  • user control.

User explicitly requests media.

71. Optional Auto-Illustration

Future setting:

Auto-generate one image at scene changes

Off by default.

72. Scene Change Detection

Future media automation may detect:

  • new location,
  • major character entrance,
  • major action event,
  • chapter boundary.

This remains optional.

73. Resource Coordination

Local LLM and image/video models may compete for GPU memory.

Media coordinator should eventually support:

  • queueing,
  • model unload/reload,
  • provider limits.

Do not assume simultaneous execution is always possible.

74. Hardware Awareness

Providers may expose:

  • VRAM requirement,
  • model availability,
  • estimated capability.

The core story engine should not need this information.

75. Provider Errors

Normalize provider errors.

Example:

error_type:
  model_missing
  out_of_memory
  invalid_request
  provider_unreachable
  cancelled

76. Model Discovery

Future media UI may list local models available from provider.

Do not auto-download them.

77. Model Downloads

Same privacy rule as Ollama:

  • installation may use Internet,
  • normal generation should not require Internet,
  • no silent downloads.

78. Media Provider Registry

Conceptual registry:

providers:
  - id: comfyui-local
    types: [image, video]
  - id: kokoro-local
    types: [tts]
  - id: local-stt
    types: [stt]

Exact products are not committed.

79. Open Dungeon Reuse

Open Dungeon is especially relevant for:

  • local image generation,
  • character visual continuity,
  • image workflow UX.

Phase 0B confirmed Open Dungeon should remain a media/UX reference rather than the production base. Study its local image worker protocol, scene/image attachment, and character visual-continuity ideas during the future media implementation stage if useful.

Do not copy its linear/destructive persistence assumptions into the core story architecture.

80. Gamentic Reuse

Gamentic is relevant for:

  • provider abstraction,
  • asynchronous generation,
  • media job concepts,
  • local multimodal architecture.

Use as a pattern reference, not a merged codebase.

81. Corvus Story Core Reuse

Corvus may be useful for:

  • visual scene extraction,
  • TTS/ComfyUI integration patterns.

Again, use concepts selectively.

82. V1 Physical Schema Direction

The v1 requirement is architectural compatibility, not media generation.

Required now:

  • scene snapshots,
  • stable optional visual profiles,
  • provider-neutral request/result types or equivalent interface contract.

A physical media job/asset table may be introduced in the dedicated future-media-hooks milestone if it is inexpensive and useful for schema stability. Its physical presence is an implementation detail, not a prerequisite for story functionality.

Do not implement a provider merely to justify a table.

83. V1 Required Media Readiness

Even without generation, v1 should preserve:

  • scene identity,
  • scene turn range,
  • character visual profiles,
  • location visual profiles,
  • branch-aware scene snapshots.

This is sufficient to avoid architectural dead ends.

84. Future Image Acceptance Test

Given an accepted scene:

Aldric and Mara examine the Silver Key in the tavern.

Generate image.

Pass if:

  • only local provider used,
  • asset attached to correct scene,
  • provenance recorded,
  • story state unchanged,
  • retry creates new asset rather than overwriting.

85. Future Branch Media Acceptance Test

Generate image on Path A.

Restore checkpoint and create Path B.

Pass if:

  • Path A image remains stored,
  • Path A image is not shown as current Path B media,
  • no automatic deletion occurs.

86. Future Video Acceptance Test

Select 4-5 combat turns.

Generate short local video.

Pass if:

  • source turn range preserved,
  • action sequence matches accepted branch,
  • abandoned-branch events are not included,
  • media generation does not mutate story.

87. Phase 0B Findings Applied

  • Open Dungeon has useful local image-generation and visual-continuity concepts, but they are not a reason to use it as the production fork.
  • The production base has no required media subsystem today; this is acceptable for v1.
  • Media must remain derived from accepted scene/story state and branch-aware.
  • Provider portability for Open Dungeon's image worker is deferred until image generation is actually scheduled.
  • Scene snapshots and visual continuity fields are sufficient near-term architecture commitments.
  • TTS/STT/video remain future providers behind the same local optional boundary.

88. Acceptance Criteria for Architecture

The architecture passes if:

  • story engine functions with no media provider,
  • media requests derive from accepted story state,
  • assets are branch-aware,
  • assets preserve source scene/turn provenance,
  • media failure cannot corrupt story state,
  • provider-specific syntax stays outside Story Engine,
  • local providers can be substituted,
  • future image/video/audio/TTS types fit the output job/asset model,
  • future STT fits the provider architecture while feeding editable draft input rather than story state,
  • generated media never automatically becomes canon,
  • abandoned-history media remains recoverable but inactive.

89. Selected Contract

Use this conceptual contract:

Authoritative Story
       |
       v
Scene Snapshot / Scene Packet
       |
       v
Optional Media Coordinator
       |
       +--> Image Provider
       +--> Video Provider
       +--> Audio Provider
       +--> TTS Provider
       +--> STT Provider (draft input path)

No production media provider is required for v1.