The release-validation milestone, and the thing it had to settle first was whether any of the earlier evidence meant what it said. M8 measured a deployment enforcing a 4,096-token input window while the application budgeted 16,384. Every request returned 200. What Ollama does with the excess is drop the oldest tokens, and the oldest tokens here are the system block — the narrator's rules and the campaign canon. A hundred-turn certification against that server would have looked perfect and proved nothing, which is why this milestone could not begin with a hundred turns. So the application asks now. Ollama's window is a property of how a model was loaded rather than of the request — sending num_ctx is accepted, ignored, and worse, reloads the model at the server's own default — so the only honest move is to find out and then tell the truth about it. /api/ps reports what a resident model is being served with, /api/show what an unloaded one will load with, both on the same host inference already uses, through the same endpoint policy and the same TLS trust store. A verified window is a ceiling on the budget; an unverified one leaves the budget alone and is recorded as unverified in the turn's own provenance, so an old turn can be asked afterwards whether it was built against a checked window. There is no third behaviour, and in particular no hard-coded 4,096: a number the server did not say would be right on one machine and wrong on the next. The proof that this is doing something is a campaign whose canon sits at the front of the prompt, 120 turns of history, and a 4,096-token window. The canon is still there afterwards and the oldest history is gone. The same campaign built the old way produces a prompt more than twice the window — the defect, reproduced, so the fix is measured against it rather than asserted. Two defects the validation found on its own, and they are the same defect twice: something was true and nobody was told. A manual state correction of four changes with one bad reference applied three, returned 201, and said nothing — while recording the refusal on the audit row nobody reads. It came to light because the identity diagnostic's own fixture was refused that way and the whole run proceeded on a campaign with no scene, which would have read as a model failure. And the narration-length setting moved no number: brief, medium and long each became one English sentence, while the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three. Both now say what they did. The other two post-M8 findings are closed as well. The tab said AI D&D, which no document had ever claimed it did not; it says Interactive Story now, with the open campaign first, and the name is the owner's decision rather than a find-and-replace to something narrower than the engine. After an Undo the reader could not tell where they had landed; the control row now ends with "Moment 11 · later story ahead", from the server's own answer, in the word the transcript already uses, with none of head, branch or depth anywhere near it. The identity diagnostic exists and the root cause does not. That campaign was destroyed, so no cause can be established — what M11 owes the finding is something that can classify the next occurrence, and a diagnostic that makes only the judgements a program can honestly make: duplicate keys, shared names, protagonist drift, state and context disagreeing. Whether prose misattributed a line is left to a person reading it beside its prompt, because a regex cannot read dialogue and one that pretended to would produce exactly the confident wrong answer this finding is about. Its detectors are proved to fire against a planted second Alice. Two entities may still share a display name. That was checked first, as the finding asked, and left permitted: a mother and a daughter, or a stranger giving a false name, are ordinary fiction, and refusing them to guard against a model mistake would refuse the wrong thing. What was missing was that it happened silently. It is reported now. Evidence, not inference: a hundred accepted turns against a real narrator with genuine process restarts; a real browser against the built SPA; a container with no network at all; a campaign moved into a data directory that never existed. Each was discarded and re-run whenever the product changed under it, and the runs that were thrown away are listed in the report with the reason, along with ten defects in the harnesses themselves — because a harness that has only ever agreed with itself is not evidence, and two of M8's five harness defects were masking real ones. No dependency was added, removed or upgraded. No acceptance test was retired, relaxed or reclassified. M11 is implemented and verified; it is not accepted, and there is no release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
47 KiB
M10 — Future Media Extension Hooks Only
Implementation report, written for an independent reviewer.
Branch m10-media-hooks, from the signed M9 commit 44edece. Implemented
2026-09-07. This is a set of claims with the evidence attached; it is not a
record of acceptance.
The one-sentence version: the scene snapshot the media contract asks for already existed, built by M5, so M10 built the seam around it and no media — one table, a packet derived on read, provider contracts with an empty registry, no dependency, no network, and no change to a single reader-facing surface.
A. Repository baseline
| M9 base commit | 44edece67e7f65bacf78010ef56f98c3c8864073 — "M9: a campaign you can actually get back" |
| Signature | git verify-commit 44edece → Good signature, RSA key 02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569, "JesseMarkowitz", trust [ultimate]. %G? = G. |
| Branch | m10-media-hooks, created from that commit. The working tree was clean at the start. |
| Upstream ancestry | upstream = https://github.com/parththakkar106/AI-DnD.git. git merge-base --is-ancestor d72f7c1b HEAD → true: the fork point is still an ancestor, so this remains a fork rather than a rewrite. |
| License / provenance | LICENSE unchanged — md5 07fde30437134836e2ee875e82a7cd31, still MIT, still "Copyright (c) 2026 Parth Thakkar". PROVENANCE.md unchanged: M10 adds only new files under backend/app/media/, one router, one model and tests. No dependency was added to requirements.txt or package.json. |
B. Architecture implemented
B.1 The finding that decided the shape of the milestone
MEDIA-EXTENSION-CONTRACT.md §5 asks for a persisted or derived scene snapshot
carrying location, participants, objects, actions, ambience, profiles,
continuity constraints, and a turn range with lineage.
Most of it already existed, and had since M5. narrative_state["scene"]
holds:
{"summary": "...", "location": "office", "present": ["bill", "alice", "roger"],
"at": {"branch_id": 3, "depth": 4}}
written by the validated set_scene typed event, snapshotted per position in
actions.narrative_state_after, restored by attempts.restore_state on every
head move, and carried in the M9 v3 bundle.
That was verified rather than inherited from M9's report: a probe played a
campaign to a scene ("Mara enters the cellar" at branch 1, depth 4), undid it and
confirmed the scene cleared to {}, diverged to a second continuation ("Mara
remains upstairs" at branch 3, depth 4), confirmed both were retained and
distinguishable, and confirmed a bundle carried both.
So M10 created no scenes table. A second scene store would have been a duplicate representation of the same fact, with its own lineage rules to get wrong — and the lineage rules are the hard part, which is the argument for reusing the ones that already work rather than against it.
B.2 What exists now
backend/app/media/__init__.py 66 lines the finding, recorded where a
future implementer will hit it
backend/app/media/packet.py 328 lines the Scene Packet, built on read
backend/app/media/profiles.py 220 lines visual profiles: the one thing
§5 asks for that nothing stored
backend/app/media/providers.py 332 lines contracts + endpoint policy
backend/app/routers/adventures/visuals.py 131 four profile endpoints + the packet
backend/app/models.py +1 model VisualProfile
Scene representation — derived, not stored. packet.build(db, adventure, start=, end=) reads narrative_store.current(adventure), which is the
authoritative document at the active head, and returns:
scene_id derived identity, below
campaign {id, title}
turn_range {branch_id, start, end}
lineage the head-capped branch lineage, as coordinates
location entity view + its visual profile, or null
characters present entities, each with its profile or null
objects significant items, held or in the location
action_summary the scene's own summary text
continuity_constraints what must stay true in a depiction
ambience {time_of_day, lighting, mood} — present, empty, §B.5
source {packet_version: 1, head_depth}
Scene identity — c<adventure>:b<branch>:<start>-<end>, e.g.
c7:b3:4-4. Derived, not allocated. The same position yields the same id in
any process, after any restart, and after the packet is thrown away and rebuilt,
with no row to keep in step. parse_scene_id() reads it back. This is the part
of a future media_assets table that would be expensive to retrofit, so it is
guaranteed now even though the table is not built.
Turn-range representation — depths on the branch the scene was set on.
Defaults to the scene's own single position; a caller passing start and end
describes a stretch, which is what a future video provider would ask for
(§31 of the contract, multi-turn packets).
Lineage association — the packet carries turn_range.branch_id and the
head-capped lineage array. The scene's coordinate comes from the scene's own
at, not from the head, and the difference is deliberate: a story can move on
without re-establishing the scene, and a picture belongs to the moment the scene
was set rather than to a later turn that did not change it.
Visual profiles — visual_profiles, keyed (adventure_id, entity_key),
with an open descriptors map, a features list and free style_notes. Four
endpoints under /api/adventures/{id}/visual-profiles — list, read, write,
delete. Campaign-scoped, not
per-position (§D, §C.3).
Provider-neutral interfaces — typing.Protocol structural types, so a
future adapter satisfies them by shape and imports nothing from here:
MediaProvider capabilities() -> ProviderCapabilities
generate(MediaRequest) -> MediaResult
SpeechProvider speak(...) -> MediaResult
TranscriptionProvider transcribe(...) -> DraftTranscription
with MediaRequest, MediaResult, ProviderCapabilities, DraftTranscription
and MediaProviderError beside them, and a registry (register, unregister,
registered, for_kind) that is empty and ships empty.
STT draft-input contract — DraftTranscription carries
editable: bool = True and has no commit method. A transcriber can produce a
draft and structurally cannot submit one; the ordinary authoritative commit path
is the only way in. §24A's rule is enforced by the shape of the type rather than
by a caller remembering it.
Request/job/asset persistence — does not exist. MediaRequest and
MediaResult are contracts; there are no media_jobs or media_assets tables.
See §F/K04 for the reasoning and how it is reported.
B.3 Nothing calls any of it
No turn, prompt, context section, health check or startup path touches
app/media/. Asserted structurally: test_m10_no_media.py parses the import
statements of turns.py, context/builder.py, narrative/apply.py,
narrative/store.py, tree.py, head.py and memorybank.py and requires that
none imports the media package.
B.4 What the packet excludes, which is the more interesting half
Excluded: the raw transcript, all imported knowledge, memories, summaries, and the state document's facts, relationships and threads.
The rule is what the story established at this position, not everything the narrator was told. Excluding imported knowledge as a class rather than filtering marked secrets is what makes §H hold for a secret nobody thought to mark: there is no filter to forget to extend.
B.5 One deliberate gap
ambience returns {time_of_day: null, lighting: null, mood: null}. The fields
are in the shape because a provider adapter should not have to branch on their
absence; they are empty because filling them would mean extending the set_scene
event, which is on the prompt path — a change to what the narrator is asked for,
which is not M10's to make. Stated here rather than left to be discovered.
C. Authority analysis
AUTHORITATIVE DERIVED
┌────────────────────────────────┐ ┌──────────────────────────┐
│ actions (the transcript) │ │ Scene Packet │
│ state_events / state_proposals │──►│ built on read │
│ narrative_state │ │ stored nowhere │
│ narrative_state_after (per pos)│ │ identity computed │
│ branches, head_branch/depth │ │ │
│ checkpoints │ │ (future: media assets) │
└────────────────────────────────┘ └──────────────────────────┘
▲ │
└────────── NO PATH ◄──────────┘
visual_profiles ── presentation metadata, campaign-scoped.
Written by the reader, read by the packet, never by the story.
C.1 The reverse path does not exist, proved three ways
By structure. test_m10_authority.py::test_the_media_package_imports_nothing_that_writes_state
requires that no file under app/media/ references narrative.apply,
set_current, head.move_to or tree.place_action. narrative.model and
narrative_store.current are reads and are used.
By vocabulary. test_no_state_event_type_was_added_for_media requires that
events.ALLOWED contains no name beginning media or containing visual or
asset. There is no way for the media layer to speak in the story's language,
so there is nothing for the validator to accept.
By behaviour, which is the one that would catch a mistake nobody predicted.
Every test in test_m10_authority.py records the authoritative document —
narrative_state, head_branch_id, head_depth, and the counts of state
events, proposals and actions — before and after a media operation, and requires
them to be identical:
| Operation | Result |
|---|---|
| Write a visual profile | authoritative document unchanged |
| Update Alice's profile to "blue coat" | unchanged, and "blue coat" appears in no fact and in no entity |
| Delete a profile | unchanged |
| Build five scene packets | unchanged |
| Build a packet with an explicit range | head does not move |
| A dummy provider returns "Alice in a red coat in a corridor" | unchanged; neither string is anywhere in the state |
A provider raises MediaProviderError |
unchanged; head does not advance |
Packet derivation raises inside build |
unchanged, and the campaign still plays |
| Delete every profile | campaign intact; packet still builds with visual_profile: null |
| A profile naming an entity that does not exist | inert; packet unaffected; play continues |
C.2 The claim in the other direction
§35 and §37 of the contract — a depiction never becomes canon, and promoting a
visual detail into canon must be a deliberate act by the reader. M10 makes that
structural: there is no code path from a MediaResult to a state event, because
there is no code that consumes a MediaResult at all.
C.3 Why a profile carries no branch coordinate
Every other derived record in the schema carries (branch_id, depth) because it
describes a moment. A profile describes none: a character does not change
appearance because the story forked. Per-position profiles would have been wrong
twice — a reader who diverged would lose their cast's appearance, which is the
opposite of the continuity a profile exists for, and a descriptor document would
land in every per-position snapshot (measured: 245 copies of the same 367 bytes
in a 120-turn campaign, §K).
D. Schema and migration
Tables added: one.
visual_profiles
id INTEGER PRIMARY KEY
adventure_id INTEGER NOT NULL FK adventures(id) ON DELETE CASCADE, indexed
entity_key VARCHAR(200) NOT NULL
descriptors JSON
features JSON
style_notes TEXT
created_at DATETIME
updated_at DATETIME
UNIQUE (adventure_id, entity_key) -- uq_visual_entity
INDEX ix_visual_profiles_adventure_id
Columns added to existing tables: none. Columns changed or dropped: none. Story tables touched: none.
Migrations added: none. LATEST_VERSION is 92, exactly as M9 left it.
create_all builds a new table on every path — fresh install, existing database,
test setup — as it did for memories, branches, checkpoints, summaries and
the M7 knowledge tables, and it builds the index too, because the index is
declared on the column rather than in __table_args__. Migration 92's own
comment states this rule for the M7 tables; M10 follows it.
Proof that none is needed, and that the two paths converge
(test_m10_bundle.py, §15 group):
| Check | Result |
|---|---|
An M9-era database (schema at 92, no visual_profiles, a campaign already in it) opened by this build |
table present, empty, ix_visual_profiles_adventure_id present, version still 92 |
| The campaign that was already there | title, action text and PRAGMA foreign_key_check unchanged |
| Opening the same database three times | index set identical after each; no error; no accumulation |
| Fresh install vs upgraded M9 file | sqlite_master DDL for visual_profiles identical; index sets identical |
backup.create() on the upgraded file |
integrity == "ok", quick_check ok, foreign_key_check empty, opens independently, carries the new table and the pre-M10 campaign, user_version 92 |
| A profile written after the upgrade, then backed up | present in the backup with its descriptors |
A defect this found — see §M.1. M10 first shipped migration 93 creating
ix_visual_profiles_adventure. Because create_all had already built
ix_visual_profiles_adventure_id, an upgraded database ended up with both and
a fresh install with one. The fresh-versus-upgraded comparison caught it; the
migration was removed rather than renamed, because the right number of
migrations here is zero.
E. Bundle implications
Did M9's v3 bundle change? Yes — it gained one optional key:
"visualProfiles": [
{"entityKey": "alice",
"descriptors": {"build": "tall", "hair": "short black", "clothing": "grey blazer"},
"features": ["tortoiseshell glasses"],
"styleNotes": "photographic, natural light",
"createdAt": "2026-09-07T…"}
]
Did the format version change? No. It stays ai-dnd-adventure-v3.
Why. M9 introduced a version because a v2 file with no prompt provenance was
ambiguous between "written before M9" and "written by M9 from a campaign that
had none". The test is therefore not "did the format gain a key" but does
omission create ambiguity about what an older file could have recorded. It does
not: a campaign with no visual profiles is the ordinary case — appearance is
something a reader adds, not something a campaign has by default — so an absent
key unambiguously means "none", exactly as checkpoints did before M4 and
knowledge before M7. Bumping to v4 for a key whose absence is unambiguous would
spend the mechanism M9 built and make it mean less next time.
Legacy import behaviour. Unchanged, and re-checked:
| File | Result |
|---|---|
A v3 file with visualProfiles deleted (an M9-written file) |
imports; campaign intact; zero profiles |
| v2, v1 | unchanged — M9's readers are untouched |
| A malformed profile in an otherwise good file | dropped, campaign still imports. A story that would not import because a description of somebody's coat is malformed would be the wrong trade |
| A profile for an entity that no longer exists | imported and inert |
Round trip. Export → import → export produces the same profiles, in the same
shape (test_a_profile_survives_a_second_round_trip_unchanged). The copy's
profiles are its own rows — editing the copy does not reach the original — and
the copy's scene packet is populated from them, which is the point of carrying
them at all. A neighbouring campaign's profiles do not travel. The planner
(bundle.plan, M9's before-anything-is-written checkpoint) checks profiles
there rather than partway through a write.
Clean-directory round trip, with profiles, across two real processes.
test_profiles_reach_a_clean_data_directory_on_another_machine follows M9's
shape — two directories, two databases, two server processes, nothing crossing
but the file, and machine B's database a file that never existed before, so its
migrations run from nothing. Machine A profiles Alice and the office and exports;
machine B imports and builds a scene packet whose action_summary matches,
whose Alice carries her descriptors, whose office carries its lighting, and whose
Roger is still visual_profile: null. Only the campaign id differs, which is
what a new machine's id space means.
This exists because M9's own test_m9_clean_import.py predates visual profiles
and carries none — it passes unchanged (§L), but it could not have caught a
profile that failed to cross. M10 added no cross-machine coupling: entity_key
is a key inside the campaign's own state document, which travels in the same
file, so unlike a branch number or a knowledge source id it needs no translation
on import.
F. Acceptance matrix
| Test | Verdict | Evidence |
|---|---|---|
| K01 — Scene Snapshot Exists | PASS (and was already passing) | narrative_state["scene"] since M5; normalized packet at GET /api/adventures/{id}/scene-packet. test_m10_media_hooks.py (contents, bounds, identity), test_m10_lineage.py (position correctness through every history operation and two process restarts). |
| K02 — Visual Character Profile | PASS | visual_profiles + four endpoints (list, read, write, delete). Optional is tested, not just stated: Roger is deliberately unprofiled and the packet reports visual_profile: null rather than an empty profile. Stability is tested per operation — Undo and divergence, Redo and Save Point restore, a genuine process restart, and a two-process move to a clean data directory. |
| K03 — Visual Location Profile | PASS | Same table and same code path — a location is an entity with a type. the office carries a profile; the packet's location.visual_profile returns it. |
| K04 — Attach Media Asset to Scene | PASS on the deferred branch | Media tables are deliberately deferred, so this is reported against the acceptance text's own second clause ("if media tables are deferred: architecture/types should demonstrate equivalent extension point"). A test registers a dummy provider, builds a packet, generates a fake PNG carrying the packet's scene_id as provenance, and shows the story model byte-for-byte unchanged. It is not PASS on the first clause, and a reviewer who requires physical media tables in v1 should read this as PARTIAL. |
K04's exact status, stated plainly. Physically implementing media_jobs and
media_assets now would mean designing a queue with no producer and no consumer,
whose shape would be decided by a provider nobody has chosen; the codebase
declined the same thing once already (M6's derived_status, commented "not a job
queue"). What is guaranteed instead is the part that would be expensive to
retrofit: a scene identity that is derived from campaign and position, so a
future asset can reference a scene without a scenes table existing to reference.
G. History and lineage evidence
test_m10_lineage.py — 8 tests, all passing, including the brief's own §4
example and its §17 sequence, and two genuine spawned-process restarts
(reusing test_process_restart.Server, so the process really goes away).
| Operation | What was checked | Result |
|---|---|---|
| Set a scene, continue | packet describes the scene's position, not the head's | pass |
| Undo | packet follows the state back; a scene set after the undone point is gone from it | pass |
| Redo | packet returns to the later scene, with the same scene_id it had before |
pass |
| Retry | the take that is live decides the scene; the superseded take's scene does not leak | pass |
| Divergence | Path A's scene and Path B's scene are different packets with different ids; both retained; the abandoned one is not current | pass |
| Save Point restore | packet matches the position the Save Point names | pass |
| Redo, and a Save Point restore | the profile is untouched by either; the scene set after the Save Point is correctly gone from the packet while the profile remains — which is the difference between story state and presentation metadata | pass |
| Restart (real process) | the same position yields the same scene_id and the same packet contents in a new process |
pass |
| Restart after divergence (real process) | the campaign reopens on the branch it was left on, and the packet is that branch's | pass |
The mechanism behind all of it is M5's, not M10's: the head move restores the whole state document and the scene is part of it. What M10 adds is the test that pins it for the media seam, plus the derived identity that makes the restart comparison meaningful — a stored id would have been trivially stable and would have proved nothing.
H. Hidden-information evidence
The test uses a hidden M7 knowledge source, because that is the product's real narrator-only mechanism, rather than an invented marker. Each check carries a positive control, so a pass cannot be a campaign where the secret was never established.
The sentinel. ZARQUON-CONCEALED-OBSERVER-7731, in a hidden Canon source
describing a concealed observer behind the office's north wall.
| Check | Result |
|---|---|
| Control: does the narrator actually receive it? | Yes — the sentinel appears in the assembled prompt for a turn about the north wall panelling |
| Does it reach the Scene Packet? | No — neither the sentinel nor "concealed observer" is anywhere in the packet |
| Control: does a visible reference source reach the narrator? | Yes — the handbook appears in the context report's used-knowledge list |
| Does that visible source reach the packet? | No — imported knowledge is excluded as a class, which is what makes the rule hold for a secret nobody thought to mark |
| Do memories and summaries reach the packet? | No — after eight turns and a summary pass, the recurring "printer incident" text is absent |
| Does something the story established reach the packet? | Yes, and it should — once a validated set_scene event puts the observer in the room, the observer is in the packet. It is no longer narrator-only knowledge; it is something that happened. The sentinel is still absent, because the story never said it. |
That last row is the reason the boundary is drawn where it is: a packet that hid established story from a depiction would be hiding the story from itself.
I. No-media operation
test_m10_no_media.py — 12 tests, all passing. §20's list, run in one campaign
with an empty provider registry:
several story turns ✓ Undo ✓
state extraction ✓ Redo ✓
memory + summary activity ✓ Retry ✓
knowledge retrieval ✓ Save Point restore ✓
restart ✓ *
* In this suite the restart is a fresh session reading what was written, not a
new process. The genuine spawned-process restarts are in
test_m10_lineage.py (§G), where they carry more weight: they run with
profiles written and packets built, and check the derived scene_id is the same
in a process that never saw the first one.
with no media warning, no media connection attempt, no missing-provider error, and no media schema requirement reaching narration.
Also checked:
- The prompt is unchanged. No context section labelled
media,scene_packetorvisual_profile*exists; the stringsvisual_profileandscene_idappear nowhere in the assembled system or story prompt. - Ordinary play writes no media row.
- No media setting exists — neither in the settings API response nor as a
column on
Settings. A setting that exists is a setting that can be pointed at a cloud by mistake. - The turn path cannot reach the media package, checked by parsing imports rather than by grepping text.
- The app serves, plays and passes health checks with
providers.registered() == {}.
J. Security and locality
Outbound destinations added: none. No module under app/media/ references
httpx, requests, urllib.request, socket, aiohttp or subprocess, and
that is a test, not an inspection. No provider adapter ships, so there is nothing
to connect to; the registry is empty at import and stays empty.
Endpoint architecture — stricter than narration.
providers.endpoint_rejection_reason(url) applies endpoints.rejection_reason
first (the shared policy: every address the hostname resolves to must be
loopback, RFC1918, link-local, IPv6 ULA or CGNAT; known cloud inference hosts are
refused by name; the check is on the resolved address, so
localhost.evil.example does not pass) and then requires loopback in
addition. Verified:
| Endpoint | Verdict |
|---|---|
http://127.0.0.1:8188/ |
allowed |
http://192.168.x.x:8188/ (trusted LAN — allowed for narration) |
refused for media |
https://api.openai.com/v1, https://replicate.com, http://8.8.8.8:8188, http://example.com |
refused |
This is deliberately narrower than SECURITY-THREAT-MODEL.md §73 permits, and
§42A now records the discrepancy rather than leaving it to be found. The
reasoning: a picture of a scene carries the scene with it, and a GPU rendering
someone's campaign is a machine that person is sitting at.
No TLS verification bypass was introduced. There is no verify=False, no
-k, and no new HTTP client at all; tlstrust.py is untouched and
test_tls_trust.py passes.
Filesystem/media exposure: none. No asset is stored, no directory is served, no path comes from a caller. M10 adds no file-serving route.
CSP/CORS: unchanged. No frontend file was modified, no new origin is
contacted, and test_offline_assets.py (which reads the built SPA and the CSP)
passes.
Cloud dependency: none. Nothing was installed to demonstrate an interface —
no ComfyUI, no diffusers, no Whisper, no Kokoro, no model download. The
requirements.txt diff is empty.
One new disclosure boundary, and it is not a network one. The Scene Packet is the input a future provider would receive, so its contents are a disclosure decision — covered in §H.
K. Performance and storage
Measured with backend/tools/m10_media_cost.py, a 120-turn campaign played
through the real turn engine and the real state pipeline, with SQL statements
counted by tools/dbmeter.py.
120 turns, 245 action rows
scene records M10 wrote
visual_profiles rows 2 one per profiled entity, written once
scene rows 0 M10 adds no scenes table
M5 per-position state snapshots 245 already there; the scene lives here
bytes added to the database
empty database 188416 B
after the fixture campaign 196608 B
after 120 more turns 1056768 B
visual profile content 367 B 0.035% of the database
profile duplication: campaign-scoped against per-position
as stored, once per entity 367 B
if snapshotted per position 89915 B x245
packet: persisted or constructed
rows written while building one 0
packet rows in any table 0 built on read, never stored
build time, 2 turns 16.6 ms
build time, 122 turns 12.6 ms
current-scene query behaviour
statements, 2 turns 5
statements, 122 turns 4
does not grow with the campaign
The packet's four statements at turn 122 are: the adventure, its visual profiles, the current user, and the head's branch. Nothing walks the transcript, which is the property that matters — a scene derivation that scanned actions would have made every future depiction O(turns).
The x245 row is the measurement behind §C.3: per-position profiles would have
stored the same 367 bytes 245 times in this campaign to say something that never
varies.
Storage added to an ordinary campaign that uses no profiles: one empty table.
L. Regression counts
All runs on this branch, after every M10 change, on 2026-09-07.
| Suite | Result | Time |
|---|---|---|
| Backend, full | 1,191 passed, 14 skipped, 0 failed — 1,205 collected, exit 0 | 862.9 s |
| M10-specific | 89 passed (39 hooks + 8 lineage + 12 authority + 12 no-media + 18 bundle/migration) | 56.4 s |
| Frontend component suite | 145 passed, 12 files, 0 failed | 8.17 s |
Lint (oxlint) |
0 errors, 15 warnings, exit 0 | — |
Production build (vite build) |
clean — index-DcHbz7ga.js 389.66 kB (gzip 119.25 kB), index-B55Q8MSM.css 47.50 kB |
0.58 s |
Docker (--no-cache) |
exit 0; image runs and imports app.media with registered() == {} |
25.2 s |
| Browser | not run — M10 adds no reader-facing surface. See §L.2 |
The 14 skips are M9's and are unchanged — confirmed by running the five files that carry a skip condition on their own (26 passed, 14 skipped): seven need a second machine or an environment the suite cannot create; the rest need a real local model.
The 15 lint warnings are pre-existing (no-unused-vars in two test files, and
react/only-export-components in components that export a constant beside a
component). M10 modified no frontend file, so the count is M9's, unchanged.
L.1 By group
Each group was run on its own, so a reviewer can check a claim without running the whole suite. Every group is green.
| Group | Files | Result |
|---|---|---|
| M10 | test_m10_media_hooks (39), _lineage (8), _authority (12), _no_media (12), _bundle (18) |
89 passed — 56.4 s |
| M9 recovery | test_m9_backup, _clean_import, _corrupt_bundles, _legacy_bundles, _portability, test_bundle_v2 |
173 passed — 371.3 s |
| Migration | test_tree_migration, test_knowledge_migration, test_pre_m5_compatibility, test_snapshot_compression |
50 passed — 27.8 s |
| M5/M6/M7 state, memory, knowledge | test_narrative_state, test_worldstate_integration, test_memory_nodes, test_memory_retrieval, test_context_memory, test_imported_knowledge, test_knowledge_retrieval_quality |
216 passed — 113.7 s |
| M3/M4 history | test_story_tree_baseline, test_branch_forking, test_head_cursor, test_save_points, test_take_state, test_process_restart |
137 passed — 139.6 s |
| Security / offline | test_egress, test_endpoint_policy, test_local_only_surface, test_offline_assets, test_tls_trust |
95 passed — 18.4 s |
The security group is the one that carries §J's claims: no outbound route, the endpoint policy, the local-only API surface, the offline asset and CSP checks, and the TLS trust union. All were green before M10 and are green now.
L.2 On the browser run
M10 adds no reader-facing surface: no page, no control, no copy, no route in
the SPA. git status shows no file under frontend/ modified. The five new
endpoints are backend-only and nothing in the browser calls them. A real-browser
regression pass would therefore be re-verifying M8/M9's surfaces against a build
identical to theirs, and its evidence would be M9's evidence. The frontend
component suite and the production build were run anyway, and are green.
This is stated as a decision, not an omission: if the reviewer wants a browser pass as a matter of process, it has not been done.
M. Findings
M.1 A redundant index migration made two databases disagree — introduced by M10, fixed here
Severity: low in effect, moderate in kind. Blocker: no — fixed. Owner: M10 (closed).
M10 first added migration 93, CREATE INDEX IF NOT EXISTS ix_visual_profiles_adventure ON visual_profiles (adventure_id). But VisualProfile.adventure_id declares
index=True, so create_all already builds ix_visual_profiles_adventure_id —
on a fresh install and on an existing database, since create_all runs before
the migration loop. The result:
upgraded from 92: ix_visual_profiles_adventure, ix_visual_profiles_adventure_id
fresh install: ix_visual_profiles_adventure_id
Two schemas differing by which path the file took, which is the thing a migration exists to prevent, plus a redundant index on every upgraded database.
Found by test_a_fresh_database_arrives_at_the_same_place, which compares a
fresh schema against an upgraded one. Neither database examined on its own would
have shown it. Fixed by removing the migration, not by renaming the index:
migration 92's own comment already records the rule for the M7 tables — when the
index is declared on the column there is nothing left for a CREATE INDEX to do.
LATEST_VERSION returns to 92.
M.2 The state model refuses an over-large scene, and a test asked for one — not a defect
While writing the packet's bounds test I sent 43 entries in set_scene's
present, exceeding validate.MAX_LABELS = 40. The event was correctly refused
and the previous scene stayed, so the test measured the wrong scene and failed.
Recorded because the diagnosis matters: the product was right and the test was
wrong. The test now uses 33 and asserts its own precondition, so it cannot
silently measure a scene it did not set.
M.3 The Story Engine vocabulary grep hit the file that forbids the vocabulary — test defect, fixed
§9's rule is that no provider vocabulary (ComfyUI, Whisper, num_inference_steps,
LoRA) appears in the Story Engine. The first version of the test grepped the
whole backend and hit providers.py, whose docstrings name those things
precisely in order to exclude them. Fixed by scoping the grep to the story
engine, and by adding a complementary test that checks the seam by behaviour:
no provider registered, and no networking import anywhere under app/media/.
M.3a A text search for "media" matched "immediately" — test defect, fixed
The first version of the no-media import check read each turn-path module and
required the string media to be absent. narrative/store.py contains the word
immediately, so the test failed on a module that imports nothing. It now parses
the file and inspects its import statements, which is what the claim was
always about. Recorded because the failure looked briefly like a real coupling
and was not, and because the fixed version is the stronger test: a module could
have imported the package while never spelling the word in prose.
M.4 ambience is present and empty — known gap, deliberate
Severity: low. Blocker: no. Owner: whichever milestone builds a coordinator.
The packet's ambience object has the right shape and no content, because
filling it would mean extending the set_scene event — a change on the prompt
path, asking the narrator for something new, which is outside M10's scope. A
future provider gets a stable shape today and content when someone decides the
narrator should be asked.
M.5 Deleting profiles is not recoverable from within the app — known limit, stated
Severity: low. Blocker: no.
M9's rebuildable data can be regenerated; a visual profile cannot, because it is something a reader wrote. It travels in the bundle, so a backup or an export recovers it, and deletion is per-entity and explicit. There is no undo for it, and none was invented — that would be a second history model beside the story's.
No pre-existing defect was found in M9's or earlier work during this milestone. The full backend suite was green before M10 began and is green now.
N. Planning changes
| Document | Change | Why |
|---|---|---|
planning/DATA-MODEL.md |
New §28A — media extension points as implemented | §20/§27/§28 describe a scene table and job/asset tables. Only one of the three exists, and a reader of the conceptual model needs to know which, and why the scene is derived from §20's own data rather than stored beside it. |
planning/TECHNICAL-DESIGN.md |
New §15.1 under Scene and Future Media Boundary | §15 said "persist or derive". The answer is derive, and the reason (M5 already persisted it) is the milestone's central fact. Also records the packet's exclusions, the STT asymmetry and the endpoint policy. |
planning/MEDIA-EXTENSION-CONTRACT.md |
New §90, appended | The contract is Phase 0B design and stays readable as such. §90 records what was built, the three places implementation answered an open question (§5 already satisfied, §12 drawn wider, §7-9 collapsed into one table), and what is deliberately unbuilt. |
planning/BUILD-MILESTONES.md |
M10 status block | Milestone status, the shaping finding, what shipped, and the defect its own tests caught. The four post-M8 playtest findings above it are untouched and still M11's. |
planning/V1-ACCEPTANCE-TESTS.md |
K01-K04 results | Acceptance evidence. K01 records that it was already passing; K04 records which of its two clauses it passes on. |
planning/SECURITY-THREAT-MODEL.md |
New §42A | The trust boundary did not widen, but in one place the implementation is deliberately narrower than §73 permits. A stricter implementation than the model describes is still a discrepancy, and an undocumented one becomes an accidental relaxation later. |
planning/VERSION.md |
v3.6 entry | Records this milestone's documentation changes and the two decisions (no format bump, no migration). |
planning/README.md |
Status, milestone map, reading order, report rotation | M10 is implemented; M11 is next. Records why M9's report stays in reports/ against the usual rotation: M9 is not accepted, and M10's baseline is M9's. |
README.md |
media/ in the architecture map; VisualProfile in the model list; test count 920 → 1,191 |
The map is the first thing a new reader reads. |
DEVELOPMENT.md |
The stricter future media endpoint rule; a note that the suite takes ~15 minutes | Whoever adds the first provider should find the rule before writing the adapter. |
No planning document was rewritten, and no earlier milestone's evidence was edited.
O. M11 handoff
O.1 M10's residual risk
- K04 is satisfied structurally, not physically. No
media_jobsormedia_assetstable exists. A future coordinator will design them, and the contracts here constrain that design only loosely. Risk: low — the expensive part (a stable scene identity) is fixed; the cheap part (two tables) is not. ambienceis an empty shape (§M.4). Filling it means extendingset_scene, which changes what the narrator is asked for. Risk: low; a provider adapter written today would find the fields and no values.- The seam has no consumer, so it is unexercised by real use. Every test here uses a dummy provider. The contracts are shaped by the contract document and by what the state model can supply, not by an adapter that had to work against a real generator. Risk: moderate for the interfaces' ergonomics, nil for the story engine — the first real adapter may want the packet reshaped, and nothing in the story depends on its shape.
- A visual profile cannot be recovered from within the app (§M.5).
- Profiles are not surfaced to the reader at all. They are API-only. Whoever builds a media UI owns the browser surface, and no reader-facing vocabulary for them has been invented — deliberately, since M10 was told not to introduce reader-facing branding or surfaces.
O.2 M9 carry-forward still relevant
All six of M9's residual risks are unchanged by M10 — none was addressed and none was made worse:
| M9 residual | Status after M10 |
|---|---|
| Bundle ceiling ~279 turns | unchanged. M10 adds ~370 bytes per campaign to a file whose ceiling is set by per-position state; it does not move the number. |
quick_check rather than integrity_check |
unchanged; re-exercised on a migrated database (§D) |
| No scheduled backup | unchanged |
Stale chunk_id in a restored snapshot |
unchanged |
| Importing machine's context window may differ | unchanged; still M11's |
| This machine cannot drive a file into/out of the browser | unchanged; still M11's, and still the cheapest fix is an unconfined Firefox or Xvfb |
M9's two items "carried to M10" are both closed: media tables were owned and
the decision is recorded (§F, K04); the scene section of the state document was
indeed the extension point, and is what the packet is built from.
O.3 The post-M8 hands-on findings — still M11's, untouched
M10 neither implemented nor tested any of them, as its brief required. They are
listed here so they cannot be lost when M9's report is eventually archived; the
durable copy is in BUILD-MILESTONES.md.
| Finding | Owner |
|---|---|
A. The browser tab still reads AI D&D |
M11 release polish |
| B. After Undo, the reader cannot tell where they are | M11 UX/release polish |
| C. The narration-length setting has no measurable effect | M11 realistic-model behaviour |
| D. Character identity / coreference confusion — root cause unknown, and the playtest database was destroyed | M11 realistic-model / context diagnostic |
On D specifically: M10's fixture is deliberately the same shape — a
protagonist, two more characters in the room, and a fourth who is not — but
nothing in M10 asserts anything about coreference, and the fixture's
resemblance is not evidence about the finding. M11 still owes the explicit
diagnostic. What M10 does contribute, incidentally, is that a visual profile
attaches to the canonical entity_key rather than to a display name, so whatever
M11 concludes about identity, profiles are keyed to the thing the state model
considers one character.
O.4 The context-window issue
Unchanged and still M11's: a deployment's enforced context window may be far
below Settings.context_token_budget (Ollama defaults to 4,096 when it sees no
VRAM). Documented in DEVELOPMENT.md; detection, the Settings warning question,
and the 100-turn certification all remain open. M10 sends nothing to a model and
does not touch it.
O.5 Realistic-model / 100-turn work
Untouched by M10, and M10 adds no new realistic-model obligation: the media seam has no model in it. The 100-turn certification, the contrast/focus measurement, the WCAG audit and the streaming-import question all carry forward unchanged.
P. Final verdict
1. Is M10's Definition of Done satisfied?
Future media providers can be added through defined local interfaces without redesigning core story authority/history.
Yes. A provider is added by implementing a Protocol and registering it;
it receives a Scene Packet built from the authoritative state at a position, and
returns a MediaResult. Nothing in story authority or history changes to
accommodate it — proved by the fact that nothing in app/media/ can even import
the code that writes state, and that every media operation leaves the
authoritative document byte-identical.
The honest qualification: no real adapter has been written against these interfaces, so their ergonomics are untested (§O.1.3). The story engine's independence, which is what the Definition of Done is actually about, is tested.
2. Are K01-K03 all PASS? Yes — all three, with evidence in §F and §G. K01 additionally records that it was already passing before M10 began.
3. What is K04's exact status? PASS on the acceptance text's deferred branch ("if media tables are deferred: architecture/types should demonstrate equivalent extension point"), demonstrated by a dummy provider producing an asset against a real packet with the story model unchanged. Not PASS on the first branch, since no media table is physically implemented. A reviewer who requires physical tables in v1 should read K04 as PARTIAL; the decision and its reasoning are in §F.
4. Can all ordinary story operation run with zero media provider? Yes — §I. Turns, state extraction, memory, summaries, knowledge retrieval, Undo, Redo, Retry and Save Point restore, with an empty registry, no media setting in existence, no warning, no connection attempt and no media row written. Restart is covered as a genuine spawned-process restart in the lineage suite (§G); the no-media suite's own restart step is a fresh session read, and §I says so.
5. Can a future image provider be added without modifying story authority or
history? Yes. It implements MediaProvider, is handed a packet, and
returns a result. No story table, event type, or history operation changes.
6. Can a future video provider consume a multi-turn scene representation
without redesigning history? Yes. packet.build(..., start=, end=) takes a
depth range on the branch and the resulting scene_id encodes it
(c7:b3:4-9). The range is expressed in the coordinates history already uses, so
a multi-turn packet is a read of existing structure rather than a new one.
7. Does future STT feed editable draft input rather than authoritative state?
Yes, structurally. TranscriptionProvider.transcribe returns a
DraftTranscription with editable=True and no commit method. A transcriber
cannot submit; the ordinary authoritative path is the only way in.
8. Is scene/media data branch-safe? Yes — §G. The scene follows the active lineage through Undo, Redo, Retry, divergence, Save Point restore and two genuine process restarts, because it is the authoritative state rather than a copy of it. Profiles are campaign-scoped by design and are stable across all of those, which is the correct behaviour for appearance and is argued in §C.3.
9. Did M10 create any network dependency? No. No dependency added, no HTTP client, no socket, no subprocess, no provider adapter, no model download, no media endpoint setting. The endpoint policy that a future provider will meet is stricter than the one narration uses: loopback only.
10. Is there any blocker before M11? No blocker. Two things a reviewer should decide rather than inherit:
- whether K04 on the deferred branch is acceptable for v1, or whether media tables must be physically present (§F);
- whether no browser regression pass is acceptable given that M10 modified no frontend file (§L.2).
Neither is a defect; both are decisions that belong to the reviewer.