Files
JesseMarkowitzandClaude Opus 5 1013c94eb1
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
M10: the seam for media, and no media
The media extension contract asks for a scene snapshot a future image or video
provider could be handed: location, who is present, what they hold, what must
stay true, and where in the story it sits. Building one was the milestone's
obvious first task, and it was the wrong one. That snapshot has existed since
M5. `narrative_state["scene"]` holds the summary, the location, the cast and the
coordinate it was written at; a validated `set_scene` event writes it, every
position snapshots it, and every head move restores it. It survives Undo, Redo,
Retry, divergence, Save Point restore and a process restart because it is the
authoritative state rather than a copy of it.

So there is no scenes table here. A second scene store would have been a second
answer to "where is the story now", with its own lineage rules to get wrong —
and the lineage rules are the expensive part, which is the argument for reusing
the ones that already work rather than against it. The Scene Packet is derived
on read, and its identity is computed from the campaign and the position rather
than allocated: the same position yields the same id in another process, after a
restart, and after the packet is thrown away and rebuilt, with no row to keep in
step. That is the part of a future media_assets table that would be expensive to
retrofit, so it is fixed now even though the table is not built.

One table, then: visual_profiles, the only thing the contract's scene list asks
for that nothing already stored. Campaign-scoped and not per-position, because a
character does not change appearance when the story forks — a reader who
diverged would otherwise lose their cast, and the same descriptors would land in
every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn
campaign to say something that never varies. Keyed by the M5 entity key rather
than a new identity namespace, and one table for characters, locations and items
alike, because a location is an entity with a type and splitting them would
reintroduce the genre shape M5 spent a milestone removing.

What the packet leaves out is the more interesting half. Not the transcript, and
not imported knowledge — none of it, not merely the sources marked hidden. The
rule is what the story established at this position, not everything the narrator
was told, and drawing it by class is what makes it hold for a secret nobody
thought to mark. A hidden Canon source proves it, with a positive control
showing the narrator did receive the sentinel the packet does not carry. Once a
validated event puts the observer in the room, the observer is in the packet:
that is no longer narrator-only knowledge, and a packet that hid it would be
hiding the story from itself.

The providers are contracts and nothing else. Protocols for image, video, audio,
speech and transcription, an empty registry, no adapter, no dependency, no
socket, and no media setting to point anywhere — a setting that exists can be
pointed at a cloud by mistake. A future provider endpoint must be loopback,
stricter than narration's trusted-LAN allowance, because a picture of a scene
carries the scene with it. Transcription returns an editable draft with no
commit method, so STT structurally cannot bypass the authoritative path.

Nothing here can write the story. Not by convention: no module under media/
imports the code that writes state, no media event type exists in the state
vocabulary, and every test in the authority suite compares the authoritative
document byte for byte either side of a media operation — including one where a
provider insists Alice is in a red coat in a corridor, and the campaign goes on
disagreeing.

One defect, found by the milestone's own tests. M10 first added a migration
creating an index that create_all already builds from the column, so an upgraded
database ended up with two indexes and a fresh install with one. Comparing the
two schemas is what caught it; neither database examined alone would have. The
migration is gone rather than renamed, and the right number of migrations for a
new table whose indexes are declared on its columns is zero.

Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145
passed. Lint, production build and Docker build clean. No frontend file changed:
M10 adds no reader-facing surface, and ordinary play — turns, state, memory,
knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media
configuration, no warning, no connection attempt and no media row written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 03:41:04 -04:00

329 lines
14 KiB
Python

"""M10: the Scene Packet — one accepted scene, bounded, for a future provider.
`MEDIA-EXTENSION-CONTRACT.md` §10-12 asks for a normalised, provider-independent
description of a scene, and asks explicitly that a provider **not** normally
receive the campaign transcript. This module builds that description.
## It is constructed, never stored
A packet is a pure function of things that are already persisted: the
authoritative state document at a position, the entity records inside it, and
the campaign's visual profiles. Storing one would create a second copy of all of
that, which could then disagree with the first — and the packet has no field the
source of truth does not already hold.
So there is no `scene_packets` table, nothing to migrate, nothing to keep in
step with the head, and nothing to carry in a bundle. Rebuilding it costs one
state read and one profile query. That is the same reasoning M9 applied to the
FTS index and the knowledge passages, applied to a smaller thing.
## Scene identity, without a scenes table
`MEDIA-EXTENSION-CONTRACT.md` §10 shows a `scene_id`, and the M10 brief asks
that a future asset be able to name unambiguously:
campaign -> lineage/story position -> source turn or turn range -> scene
That is a **coordinate**, and the application already has one. So the identity
is derived rather than allocated:
c<adventure>:b<branch>:<start>-<end>
Two properties follow, and both matter more than a surrogate key would have:
* it is **stable** — the same scene yields the same id on any machine, before
and after an export, without a row having to travel;
* it is **resolvable** — a future asset holding this string can be turned back
into the exact accepted position it depicts, with no lookup table.
A surrogate `scene_id` would have needed a table, a lineage column, a restore
path and bundle carriage, all to name something the coordinate already names.
## Ranges, because a video is not a turn
`build` takes a range, not a position. §30-31 of the contract describe a video
covering several accepted turns, and the M10 brief is explicit that neither
"one turn == one scene" nor "one scene == one asset" may be assumed.
So `start` and `end` are depths on one branch, the identity carries both, and a
single-turn image is the case where they are equal rather than a different kind
of request. Several future assets may name the same identity; nothing here
allocates or records them, so nothing constrains how many there are.
## What is deliberately not in a packet
**The transcript.** Not a summarised version of it either. The packet carries
the scene's own summary — the one sentence the story itself accepted through
`set_scene` — and the entities present. A provider that needs to depict a room
does not need to have read the campaign.
**Imported knowledge, of any class.** Not canon, not reference, not
inspiration, and emphatically not a narrator-only source. This is the hidden
information boundary and it is drawn structurally: this module never reads
`knowledge_sources`, so there is no filter to get wrong and no marker to
overlook. A secret reaches a packet only if the *story* put it into accepted
state through a validated event — which is the correct rule, because at that
point it is something that happened rather than something the narrator knows.
**Memories and summaries.** Derived narrative text about the campaign's past,
which is not what depicting a present moment needs.
**Facts, relationships and threads.** These are the campaign's reasoning about
itself. A `continuity_constraints` list carries the few that bear on depiction —
what a character is holding, where they are — and nothing else.
The result is that the honest answer to "what could leak through a packet" is
"what the accepted scene contains", which is what a picture of that scene would
show anyway.
"""
from __future__ import annotations
from sqlalchemy.orm import Session
from .. import models
from ..context import lineage
from ..narrative import model as narrative_model
from ..narrative import store as narrative_store
from . import profiles as visual_profiles
#: How many entities one packet will describe. A scene is a moment with people
#: in it; a request naming two hundred is a runaway state document rather than a
#: picture, and the bound keeps a future provider's prompt finite.
MAX_CHARACTERS = 24
MAX_OBJECTS = 24
MAX_CONSTRAINTS = 24
def scene_id(adventure_id: int, branch_id: int | None, start: int, end: int) -> str:
"""The derived, stable identity for one scene. See the module docstring."""
branch = branch_id if branch_id is not None else 0
return f"c{adventure_id}:b{branch}:{start}-{end}"
def parse_scene_id(value: str) -> dict | None:
"""Turns a scene identity back into the coordinate it names, or `None`.
The half that makes the derived identity worth having: a future asset
holding this string can be resolved to an accepted position without a table.
"""
try:
campaign, branch, span = str(value).split(":")
start, end = span.split("-")
return {
"adventure_id": int(campaign.lstrip("c")),
"branch_id": int(branch.lstrip("b")),
"start": int(start),
"end": int(end),
}
except (ValueError, AttributeError):
return None
def build(
db: Session,
adventure: models.Adventure,
*,
start: int | None = None,
end: int | None = None,
) -> dict:
"""The Scene Packet for a range of accepted story on the active branch.
Defaults to the scene at the active head, which is the ordinary case: an
image of what is happening now. `start` and `end` are depths on the active
branch; passing both describes a stretch, which is what a future video
would ask for.
Reads. Writes nothing, and cannot: this module imports no writer, emits no
event and does not touch the head. `test_m10_authority.py` asserts the
authoritative document is byte-identical either side of a build.
"""
state = narrative_store.current(adventure)
scene = state.get("scene") if isinstance(state.get("scene"), dict) else {}
branch_id = adventure.head_branch_id
head_depth = adventure.head_depth
# The scene's own coordinate is the position `set_scene` last ran at, which
# is where the depiction belongs. It can sit behind the head — the story may
# have moved on without re-establishing the scene — and that is correct: the
# picture is of the moment the scene was set, not of a later turn that did
# not change it.
at = scene.get("at") if isinstance(scene.get("at"), dict) else {}
scene_branch = at.get("branch_id") if at.get("branch_id") is not None else branch_id
scene_depth = at.get("depth") if _is_int(at.get("depth")) else head_depth
first = start if _is_int(start) else scene_depth
last = end if _is_int(end) else max(first, scene_depth)
if last < first:
first, last = last, first
profiles = visual_profiles.by_key(db, adventure)
location_key = scene.get("location") if isinstance(scene.get("location"), str) else None
present = [k for k in (scene.get("present") or []) if isinstance(k, str)]
return {
"scene_id": scene_id(adventure.id, scene_branch, first, last),
"campaign": {"id": adventure.id, "title": adventure.title},
# Where in the story this is, in the vocabulary the application already
# uses internally. A future provider does not read these; a future
# coordinator resolving an asset back to its source does.
"turn_range": {"branch_id": scene_branch, "start": first, "end": last},
"lineage": _lineage_of(db, adventure),
"location": _entity_view(state, profiles, location_key),
"characters": [
view for key in present[:MAX_CHARACTERS]
if (view := _entity_view(state, profiles, key)) is not None
],
"objects": _objects(state, profiles, present, location_key),
"action_summary": str(scene.get("summary") or ""),
"continuity_constraints": _constraints(state, present, location_key),
# Present, empty, and deliberately so — see `_ambience`.
"ambience": _ambience(scene),
"source": {
# What produced this, so a future asset's provenance can say which
# build's rules bounded the packet it was made from.
"packet_version": PACKET_VERSION,
"head_depth": head_depth,
},
}
#: The packet's own shape version. A future provider adapter can branch on it if
#: the packet gains fields; nothing in the story engine reads it.
PACKET_VERSION = 1
def _lineage_of(db: Session, adventure: models.Adventure) -> list[dict]:
"""The capped lineage this scene sits on, as provenance.
Read through `lineage.path_of`, the same helper every story read uses, so a
packet cannot describe a position the story could not. M10 builds no media
head: there is one head, and this follows it.
"""
try:
path = lineage.path_of(db, adventure)
except Exception: # noqa: BLE001 - a packet is a read; it does not raise
return []
entries = getattr(path, "entries", None)
if not entries:
return []
return [
{"branch_id": branch_id, "through_depth": cap}
for branch_id, cap in entries
]
def _entity_view(state: dict, profiles: dict, key: str | None) -> dict | None:
"""One entity as a packet describes it: what it is, plus how it looks."""
if not key:
return None
found = narrative_model.entity(state, key)
if found is None:
return None
return {
"key": key,
"name": narrative_model.entity_name(state, key),
"type": found.get("type") or "other",
"status": found.get("status") or "active",
"description": found.get("description") or "",
# `None` rather than an empty profile, so a provider can tell "nobody
# said how this looks" from "somebody said it looks like nothing".
"visual_profile": profiles.get(key),
}
def _objects(
state: dict, profiles: dict, present: list[str], location_key: str | None
) -> list[dict]:
"""The things visibly in the scene, from what the present entities hold.
Possession is the only relation in the state document that says an object is
*somewhere*, so it is the honest source for "what would be in the picture".
An item nobody in the scene is carrying is not depicted, which is the same
rule a reader would apply looking at the room.
"""
possessions = state.get("possessions")
if not isinstance(possessions, dict):
return []
holders = set(present) | ({location_key} if location_key else set())
out: list[dict] = []
for item_key, holder in possessions.items():
if holder not in holders or not isinstance(item_key, str):
continue
view = _entity_view(state, profiles, item_key)
if view is None:
continue
view["held_by"] = holder
out.append(view)
if len(out) >= MAX_OBJECTS:
break
return out
def _constraints(
state: dict, present: list[str], location_key: str | None
) -> list[str]:
"""The few facts that bear on depicting *this* scene, as sentences.
Deliberately narrow. The state document's `facts` list is the campaign's
reasoning about itself and most of it has nothing to do with a picture;
forwarding all of it would make the packet a state dump with a different
name, and would be the route by which something the scene has not exposed
reached a provider.
So only two kinds are carried: where the present entities are, and what they
are holding. Both are already visible in the scene by construction.
"""
out: list[str] = []
for key in present:
found = narrative_model.entity(state, key)
if found is None:
continue
name = narrative_model.entity_name(state, key)
status = found.get("status")
if status and status != "active":
out.append(f"{name} is {status}.")
if len(out) >= MAX_CONSTRAINTS:
return out
possessions = state.get("possessions")
if isinstance(possessions, dict):
for item_key, holder in possessions.items():
if holder not in present:
continue
out.append(
f"{narrative_model.entity_name(state, holder)} is carrying "
f"{narrative_model.entity_name(state, item_key)}."
)
if len(out) >= MAX_CONSTRAINTS:
break
return out
def _ambience(scene: dict) -> dict:
"""Time of day, lighting and mood — present in the shape, empty in v1.
`MEDIA-EXTENSION-CONTRACT.md` §5 lists these among a scene snapshot's
conceptual fields, and M10 **does not** add them to the `set_scene` event
that would establish them.
That is a deliberate deferral rather than an oversight. Adding them would
mean extending M5's typed-event vocabulary, which means teaching the
narrator to emit them, which means changing the prompt — and M10's central
acceptance condition is that ordinary story flow is *unchanged*. Buying
three optional fields at the price of touching every narration was the wrong
trade for a milestone whose deliverable is a seam.
So the keys are here and are `None`, read from the scene document if a later
milestone starts recording them. A provider adapter written today against
this shape keeps working when they arrive.
"""
return {
"time_of_day": scene.get("time_of_day") or None,
"lighting": scene.get("lighting") or None,
"mood": scene.get("mood") or None,
}
def _is_int(value) -> bool:
return isinstance(value, int) and not isinstance(value, bool)