Files
interactive-story/planning/archive/milestone-reports/M10-IMPLEMENTATION-REPORT.md
T
JesseMarkowitzandClaude Opus 5 144406cd48 M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was
whether any of the earlier evidence meant what it said. M8 measured a deployment
enforcing a 4,096-token input window while the application budgeted 16,384.
Every request returned 200. What Ollama does with the excess is drop the oldest
tokens, and the oldest tokens here are the system block — the narrator's rules
and the campaign canon. A hundred-turn certification against that server would
have looked perfect and proved nothing, which is why this milestone could not
begin with a hundred turns.

So the application asks now. Ollama's window is a property of how a model was
loaded rather than of the request — sending num_ctx is accepted, ignored, and
worse, reloads the model at the server's own default — so the only honest move
is to find out and then tell the truth about it. /api/ps reports what a resident
model is being served with, /api/show what an unloaded one will load with, both
on the same host inference already uses, through the same endpoint policy and
the same TLS trust store. A verified window is a ceiling on the budget; an
unverified one leaves the budget alone and is recorded as unverified in the
turn's own provenance, so an old turn can be asked afterwards whether it was
built against a checked window. There is no third behaviour, and in particular
no hard-coded 4,096: a number the server did not say would be right on one
machine and wrong on the next.

The proof that this is doing something is a campaign whose canon sits at the
front of the prompt, 120 turns of history, and a 4,096-token window. The canon
is still there afterwards and the oldest history is gone. The same campaign
built the old way produces a prompt more than twice the window — the defect,
reproduced, so the fix is measured against it rather than asserted.

Two defects the validation found on its own, and they are the same defect twice:
something was true and nobody was told. A manual state correction of four
changes with one bad reference applied three, returned 201, and said nothing —
while recording the refusal on the audit row nobody reads. It came to light
because the identity diagnostic's own fixture was refused that way and the whole
run proceeded on a campaign with no scene, which would have read as a model
failure. And the narration-length setting moved no number: brief, medium and
long each became one English sentence, while the numeric hint the model actually
reads was derived from the global reply cap and said the same thing for all
three. Both now say what they did.

The other two post-M8 findings are closed as well. The tab said AI D&D, which no
document had ever claimed it did not; it says Interactive Story now, with the
open campaign first, and the name is the owner's decision rather than a
find-and-replace to something narrower than the engine. After an Undo the reader
could not tell where they had landed; the control row now ends with
"Moment 11 · later story ahead", from the server's own answer, in the word the
transcript already uses, with none of head, branch or depth anywhere near it.

The identity diagnostic exists and the root cause does not. That campaign was
destroyed, so no cause can be established — what M11 owes the finding is
something that can classify the next occurrence, and a diagnostic that makes only
the judgements a program can honestly make: duplicate keys, shared names,
protagonist drift, state and context disagreeing. Whether prose misattributed a
line is left to a person reading it beside its prompt, because a regex cannot
read dialogue and one that pretended to would produce exactly the confident wrong
answer this finding is about. Its detectors are proved to fire against a planted
second Alice.

Two entities may still share a display name. That was checked first, as the
finding asked, and left permitted: a mother and a daughter, or a stranger giving
a false name, are ordinary fiction, and refusing them to guard against a model
mistake would refuse the wrong thing. What was missing was that it happened
silently. It is reported now.

Evidence, not inference: a hundred accepted turns against a real narrator with
genuine process restarts; a real browser against the built SPA; a container with
no network at all; a campaign moved into a data directory that never existed.
Each was discarded and re-run whenever the product changed under it, and the runs
that were thrown away are listed in the report with the reason, along with ten
defects in the harnesses themselves — because a harness that has only ever
agreed with itself is not evidence, and two of M8's five harness defects were
masking real ones.

No dependency was added, removed or upgraded. No acceptance test was retired,
relaxed or reclassified. M11 is implemented and verified; it is not accepted, and
there is no release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 14:01:20 -04:00

856 lines
47 KiB
Markdown

# M10 — Future Media Extension Hooks Only
**Implementation report, written for an independent reviewer.**
Branch `m10-media-hooks`, from the signed M9 commit `44edece`. Implemented
2026-09-07. This is a set of claims with the evidence attached; it is not a
record of acceptance.
**The one-sentence version:** the scene snapshot the media contract asks for
already existed, built by M5, so M10 built the seam around it and no media —
one table, a packet derived on read, provider contracts with an empty registry,
no dependency, no network, and no change to a single reader-facing surface.
---
## A. Repository baseline
| | |
| --- | --- |
| **M9 base commit** | `44edece67e7f65bacf78010ef56f98c3c8864073` — *"M9: a campaign you can actually get back"* |
| **Signature** | `git verify-commit 44edece` → **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, "JesseMarkowitz", trust `[ultimate]`. `%G?` = `G`. |
| **Branch** | `m10-media-hooks`, created from that commit. The working tree was clean at the start. |
| **Upstream ancestry** | `upstream` = `https://github.com/parththakkar106/AI-DnD.git`. `git merge-base --is-ancestor d72f7c1b HEAD` → true: the fork point is still an ancestor, so this remains a fork rather than a rewrite. |
| **License / provenance** | `LICENSE` unchanged — md5 `07fde30437134836e2ee875e82a7cd31`, still MIT, still "Copyright (c) 2026 Parth Thakkar". `PROVENANCE.md` unchanged: M10 adds only new files under `backend/app/media/`, one router, one model and tests. **No dependency was added** to `requirements.txt` or `package.json`. |
---
## B. Architecture implemented
### B.1 The finding that decided the shape of the milestone
`MEDIA-EXTENSION-CONTRACT.md` §5 asks for a persisted or derived scene snapshot
carrying location, participants, objects, actions, ambience, profiles,
continuity constraints, and a turn range with lineage.
**Most of it already existed, and had since M5.** `narrative_state["scene"]`
holds:
```json
{"summary": "...", "location": "office", "present": ["bill", "alice", "roger"],
"at": {"branch_id": 3, "depth": 4}}
```
written by the validated `set_scene` typed event, snapshotted per position in
`actions.narrative_state_after`, restored by `attempts.restore_state` on every
head move, and carried in the M9 v3 bundle.
That was **verified rather than inherited from M9's report**: a probe played a
campaign to a scene ("Mara enters the cellar" at branch 1, depth 4), undid it and
confirmed the scene cleared to `{}`, diverged to a second continuation ("Mara
remains upstairs" at branch 3, depth 4), confirmed both were retained and
distinguishable, and confirmed a bundle carried both.
So **M10 created no scenes table.** A second scene store would have been a
duplicate representation of the same fact, with its own lineage rules to get
wrong — and the lineage rules are the hard part, which is the argument for
reusing the ones that already work rather than against it.
### B.2 What exists now
```text
backend/app/media/__init__.py 66 lines the finding, recorded where a
future implementer will hit it
backend/app/media/packet.py 328 lines the Scene Packet, built on read
backend/app/media/profiles.py 220 lines visual profiles: the one thing
§5 asks for that nothing stored
backend/app/media/providers.py 332 lines contracts + endpoint policy
backend/app/routers/adventures/visuals.py 131 four profile endpoints + the packet
backend/app/models.py +1 model VisualProfile
```
**Scene representation** — derived, not stored. `packet.build(db, adventure,
start=, end=)` reads `narrative_store.current(adventure)`, which is the
authoritative document at the active head, and returns:
```text
scene_id derived identity, below
campaign {id, title}
turn_range {branch_id, start, end}
lineage the head-capped branch lineage, as coordinates
location entity view + its visual profile, or null
characters present entities, each with its profile or null
objects significant items, held or in the location
action_summary the scene's own summary text
continuity_constraints what must stay true in a depiction
ambience {time_of_day, lighting, mood} — present, empty, §B.5
source {packet_version: 1, head_depth}
```
**Scene identity** — `c<adventure>:b<branch>:<start>-<end>`, e.g.
`c7:b3:4-4`. **Derived, not allocated.** The same position yields the same id in
any process, after any restart, and after the packet is thrown away and rebuilt,
with no row to keep in step. `parse_scene_id()` reads it back. This is the part
of a future `media_assets` table that would be expensive to retrofit, so it is
guaranteed now even though the table is not built.
**Turn-range representation** — depths on the branch the scene was set on.
Defaults to the scene's own single position; a caller passing `start` and `end`
describes a stretch, which is what a future video provider would ask for
(§31 of the contract, multi-turn packets).
**Lineage association** — the packet carries `turn_range.branch_id` and the
head-capped `lineage` array. The scene's coordinate comes from the scene's own
`at`, not from the head, and the difference is deliberate: a story can move on
without re-establishing the scene, and a picture belongs to the moment the scene
was set rather than to a later turn that did not change it.
**Visual profiles** — `visual_profiles`, keyed `(adventure_id, entity_key)`,
with an open `descriptors` map, a `features` list and free `style_notes`. Four
endpoints under `/api/adventures/{id}/visual-profiles` — list, read, write,
delete. Campaign-scoped, not
per-position (§D, §C.3).
**Provider-neutral interfaces** — `typing.Protocol` structural types, so a
future adapter satisfies them by shape and imports nothing from here:
```text
MediaProvider capabilities() -> ProviderCapabilities
generate(MediaRequest) -> MediaResult
SpeechProvider speak(...) -> MediaResult
TranscriptionProvider transcribe(...) -> DraftTranscription
```
with `MediaRequest`, `MediaResult`, `ProviderCapabilities`, `DraftTranscription`
and `MediaProviderError` beside them, and a registry (`register`, `unregister`,
`registered`, `for_kind`) that **is empty and ships empty**.
**STT draft-input contract** — `DraftTranscription` carries
`editable: bool = True` and **has no commit method**. A transcriber can produce a
draft and structurally cannot submit one; the ordinary authoritative commit path
is the only way in. §24A's rule is enforced by the shape of the type rather than
by a caller remembering it.
**Request/job/asset persistence** — **does not exist.** `MediaRequest` and
`MediaResult` are contracts; there are no `media_jobs` or `media_assets` tables.
See §F/K04 for the reasoning and how it is reported.
### B.3 Nothing calls any of it
No turn, prompt, context section, health check or startup path touches
`app/media/`. Asserted structurally: `test_m10_no_media.py` parses the import
statements of `turns.py`, `context/builder.py`, `narrative/apply.py`,
`narrative/store.py`, `tree.py`, `head.py` and `memorybank.py` and requires that
none imports the media package.
### B.4 What the packet excludes, which is the more interesting half
Excluded: the raw transcript, **all imported knowledge**, memories, summaries,
and the state document's facts, relationships and threads.
The rule is *what the story established at this position*, not *everything the
narrator was told*. Excluding imported knowledge **as a class** rather than
filtering marked secrets is what makes §H hold for a secret nobody thought to
mark: there is no filter to forget to extend.
### B.5 One deliberate gap
`ambience` returns `{time_of_day: null, lighting: null, mood: null}`. The fields
are in the shape because a provider adapter should not have to branch on their
absence; they are empty because filling them would mean extending the `set_scene`
event, which is on the prompt path — a change to what the narrator is asked for,
which is not M10's to make. Stated here rather than left to be discovered.
---
## C. Authority analysis
```text
AUTHORITATIVE DERIVED
┌────────────────────────────────┐ ┌──────────────────────────┐
│ actions (the transcript) │ │ Scene Packet │
│ state_events / state_proposals │──►│ built on read │
│ narrative_state │ │ stored nowhere │
│ narrative_state_after (per pos)│ │ identity computed │
│ branches, head_branch/depth │ │ │
│ checkpoints │ │ (future: media assets) │
└────────────────────────────────┘ └──────────────────────────┘
▲ │
└────────── NO PATH ◄──────────┘
visual_profiles ── presentation metadata, campaign-scoped.
Written by the reader, read by the packet, never by the story.
```
### C.1 The reverse path does not exist, proved three ways
**By structure.** `test_m10_authority.py::test_the_media_package_imports_nothing_that_writes_state`
requires that no file under `app/media/` references `narrative.apply`,
`set_current`, `head.move_to` or `tree.place_action`. `narrative.model` and
`narrative_store.current` are reads and are used.
**By vocabulary.** `test_no_state_event_type_was_added_for_media` requires that
`events.ALLOWED` contains no name beginning `media` or containing `visual` or
`asset`. There is no way for the media layer to speak in the story's language,
so there is nothing for the validator to accept.
**By behaviour, which is the one that would catch a mistake nobody predicted.**
Every test in `test_m10_authority.py` records the authoritative document —
`narrative_state`, `head_branch_id`, `head_depth`, and the counts of state
events, proposals and actions — before and after a media operation, and requires
them to be **identical**:
| Operation | Result |
| --- | --- |
| Write a visual profile | authoritative document unchanged |
| Update Alice's profile to "blue coat" | unchanged, and `"blue coat"` appears in no fact and in no entity |
| Delete a profile | unchanged |
| Build five scene packets | unchanged |
| Build a packet with an explicit range | head does not move |
| A dummy provider returns "Alice in a red coat in a corridor" | unchanged; neither string is anywhere in the state |
| A provider raises `MediaProviderError` | unchanged; head does not advance |
| Packet derivation raises inside `build` | unchanged, and the campaign still plays |
| Delete every profile | campaign intact; packet still builds with `visual_profile: null` |
| A profile naming an entity that does not exist | inert; packet unaffected; play continues |
### C.2 The claim in the other direction
§35 and §37 of the contract — a depiction never becomes canon, and promoting a
visual detail into canon must be a deliberate act by the reader. M10 makes that
structural: there is no code path from a `MediaResult` to a state event, because
there is no code that consumes a `MediaResult` at all.
### C.3 Why a profile carries no branch coordinate
Every other derived record in the schema carries `(branch_id, depth)` because it
describes a *moment*. A profile describes none: a character does not change
appearance because the story forked. Per-position profiles would have been wrong
twice — a reader who diverged would lose their cast's appearance, which is the
opposite of the continuity a profile exists for, and a descriptor document would
land in every per-position snapshot (measured: 245 copies of the same 367 bytes
in a 120-turn campaign, §K).
---
## D. Schema and migration
**Tables added:** one.
```text
visual_profiles
id INTEGER PRIMARY KEY
adventure_id INTEGER NOT NULL FK adventures(id) ON DELETE CASCADE, indexed
entity_key VARCHAR(200) NOT NULL
descriptors JSON
features JSON
style_notes TEXT
created_at DATETIME
updated_at DATETIME
UNIQUE (adventure_id, entity_key) -- uq_visual_entity
INDEX ix_visual_profiles_adventure_id
```
**Columns added to existing tables:** none.
**Columns changed or dropped:** none.
**Story tables touched:** none.
**Migrations added: none.** `LATEST_VERSION` is **92**, exactly as M9 left it.
`create_all` builds a new table on every path — fresh install, existing database,
test setup — as it did for `memories`, `branches`, `checkpoints`, `summaries` and
the M7 knowledge tables, and it builds the index too, because the index is
declared on the column rather than in `__table_args__`. Migration 92's own
comment states this rule for the M7 tables; M10 follows it.
**Proof that none is needed**, and that the two paths converge
(`test_m10_bundle.py`, §15 group):
| Check | Result |
| --- | --- |
| An M9-era database (schema at 92, no `visual_profiles`, a campaign already in it) opened by this build | table present, empty, `ix_visual_profiles_adventure_id` present, version still 92 |
| The campaign that was already there | title, action text and `PRAGMA foreign_key_check` unchanged |
| Opening the same database three times | index set identical after each; no error; no accumulation |
| Fresh install vs upgraded M9 file | `sqlite_master` DDL for `visual_profiles` **identical**; index sets **identical** |
| `backup.create()` on the upgraded file | `integrity == "ok"`, `quick_check` ok, `foreign_key_check` empty, opens independently, carries the new table and the pre-M10 campaign, `user_version` 92 |
| A profile written after the upgrade, then backed up | present in the backup with its descriptors |
**A defect this found — see §M.1.** M10 first shipped migration 93 creating
`ix_visual_profiles_adventure`. Because `create_all` had already built
`ix_visual_profiles_adventure_id`, an *upgraded* database ended up with both and
a fresh install with one. The fresh-versus-upgraded comparison caught it; the
migration was removed rather than renamed, because the right number of
migrations here is zero.
---
## E. Bundle implications
**Did M9's v3 bundle change?** Yes — it gained one optional key:
```json
"visualProfiles": [
{"entityKey": "alice",
"descriptors": {"build": "tall", "hair": "short black", "clothing": "grey blazer"},
"features": ["tortoiseshell glasses"],
"styleNotes": "photographic, natural light",
"createdAt": "2026-09-07T…"}
]
```
**Did the format version change?** **No. It stays `ai-dnd-adventure-v3`.**
**Why.** M9 introduced a version because a v2 file with no prompt provenance was
ambiguous between "written before M9" and "written by M9 from a campaign that
had none". The test is therefore not "did the format gain a key" but *does
omission create ambiguity about what an older file could have recorded*. It does
not: a campaign with no visual profiles is the ordinary case — appearance is
something a reader adds, not something a campaign has by default — so an absent
key unambiguously means "none", exactly as `checkpoints` did before M4 and
`knowledge` before M7. Bumping to v4 for a key whose absence is unambiguous would
spend the mechanism M9 built and make it mean less next time.
**Legacy import behaviour.** Unchanged, and re-checked:
| File | Result |
| --- | --- |
| A v3 file with `visualProfiles` deleted (an M9-written file) | imports; campaign intact; zero profiles |
| v2, v1 | unchanged — M9's readers are untouched |
| A malformed profile in an otherwise good file | **dropped, campaign still imports.** A story that would not import because a description of somebody's coat is malformed would be the wrong trade |
| A profile for an entity that no longer exists | imported and inert |
**Round trip.** Export → import → export produces the same profiles, in the same
shape (`test_a_profile_survives_a_second_round_trip_unchanged`). The copy's
profiles are its own rows — editing the copy does not reach the original — and
the copy's scene packet is populated from them, which is the point of carrying
them at all. A neighbouring campaign's profiles do not travel. The planner
(`bundle.plan`, M9's before-anything-is-written checkpoint) checks profiles
there rather than partway through a write.
**Clean-directory round trip, with profiles, across two real processes.**
`test_profiles_reach_a_clean_data_directory_on_another_machine` follows M9's
shape — two directories, two databases, two server processes, nothing crossing
but the file, and machine B's database a file that never existed before, so its
migrations run from nothing. Machine A profiles Alice and the office and exports;
machine B imports and builds a scene packet whose `action_summary` matches,
whose Alice carries her descriptors, whose office carries its lighting, and whose
Roger is still `visual_profile: null`. Only the campaign id differs, which is
what a new machine's id space means.
This exists because M9's own `test_m9_clean_import.py` predates visual profiles
and carries none — it passes unchanged (§L), but it could not have caught a
profile that failed to cross. M10 added no cross-machine coupling: `entity_key`
is a key inside the campaign's own state document, which travels in the same
file, so unlike a branch number or a knowledge source id it needs no translation
on import.
---
## F. Acceptance matrix
| Test | Verdict | Evidence |
| --- | --- | --- |
| **K01 — Scene Snapshot Exists** | **PASS** *(and was already passing)* | `narrative_state["scene"]` since M5; normalized packet at `GET /api/adventures/{id}/scene-packet`. `test_m10_media_hooks.py` (contents, bounds, identity), `test_m10_lineage.py` (position correctness through every history operation and two process restarts). |
| **K02 — Visual Character Profile** | **PASS** | `visual_profiles` + four endpoints (list, read, write, delete). Optional is tested, not just stated: Roger is deliberately unprofiled and the packet reports `visual_profile: null` rather than an empty profile. Stability is tested per operation — Undo and divergence, Redo and Save Point restore, a genuine process restart, and a two-process move to a clean data directory. |
| **K03 — Visual Location Profile** | **PASS** | Same table and same code path — a location is an entity with a `type`. `the office` carries a profile; the packet's `location.visual_profile` returns it. |
| **K04 — Attach Media Asset to Scene** | **PASS on the deferred branch** | Media tables are deliberately deferred, so this is reported against the acceptance text's own second clause ("if media tables are deferred: architecture/types should demonstrate equivalent extension point"). A test registers a dummy provider, builds a packet, generates a fake PNG carrying the packet's `scene_id` as provenance, and shows the story model byte-for-byte unchanged. **It is not PASS on the first clause**, and a reviewer who requires physical media tables in v1 should read this as PARTIAL. |
**K04's exact status, stated plainly.** Physically implementing `media_jobs` and
`media_assets` now would mean designing a queue with no producer and no consumer,
whose shape would be decided by a provider nobody has chosen; the codebase
declined the same thing once already (M6's `derived_status`, commented "not a job
queue"). What is guaranteed instead is the part that would be expensive to
retrofit: a scene identity that is *derived* from campaign and position, so a
future asset can reference a scene without a scenes table existing to reference.
---
## G. History and lineage evidence
`test_m10_lineage.py` — 8 tests, all passing, including the brief's own §4
example and its §17 sequence, and two **genuine spawned-process restarts**
(reusing `test_process_restart.Server`, so the process really goes away).
| Operation | What was checked | Result |
| --- | --- | --- |
| Set a scene, continue | packet describes the scene's position, not the head's | pass |
| **Undo** | packet follows the state back; a scene set after the undone point is gone from it | pass |
| **Redo** | packet returns to the later scene, with the same `scene_id` it had before | pass |
| **Retry** | the take that is live decides the scene; the superseded take's scene does not leak | pass |
| **Divergence** | Path A's scene and Path B's scene are different packets with different ids; both retained; the abandoned one is not current | pass |
| **Save Point restore** | packet matches the position the Save Point names | pass |
| **Redo, and a Save Point restore** | the profile is untouched by either; the scene set after the Save Point is correctly gone from the packet while the profile remains — which is the difference between story state and presentation metadata | pass |
| **Restart (real process)** | the same position yields the same `scene_id` and the same packet contents in a new process | pass |
| **Restart after divergence (real process)** | the campaign reopens on the branch it was left on, and the packet is that branch's | pass |
The mechanism behind all of it is M5's, not M10's: the head move restores the
whole state document and the scene is part of it. What M10 adds is the test that
pins it for the media seam, plus the derived identity that makes the restart
comparison meaningful — a stored id would have been trivially stable and would
have proved nothing.
---
## H. Hidden-information evidence
The test uses a **hidden M7 knowledge source**, because that is the product's
real narrator-only mechanism, rather than an invented marker. Each check carries
a **positive control**, so a pass cannot be a campaign where the secret was never
established.
**The sentinel.** `ZARQUON-CONCEALED-OBSERVER-7731`, in a hidden Canon source
describing a concealed observer behind the office's north wall.
| Check | Result |
| --- | --- |
| **Control:** does the narrator actually receive it? | **Yes** — the sentinel appears in the assembled prompt for a turn about the north wall panelling |
| Does it reach the Scene Packet? | **No** — neither the sentinel nor "concealed observer" is anywhere in the packet |
| **Control:** does a *visible* reference source reach the narrator? | **Yes** — the handbook appears in the context report's used-knowledge list |
| Does that visible source reach the packet? | **No** — imported knowledge is excluded as a class, which is what makes the rule hold for a secret nobody thought to mark |
| Do memories and summaries reach the packet? | **No** — after eight turns and a summary pass, the recurring "printer incident" text is absent |
| Does something the *story* established reach the packet? | **Yes**, and it should — once a validated `set_scene` event puts the observer in the room, the observer is in the packet. It is no longer narrator-only knowledge; it is something that happened. The sentinel is still absent, because the story never said it. |
That last row is the reason the boundary is drawn where it is: a packet that hid
established story from a depiction would be hiding the story from itself.
---
## I. No-media operation
`test_m10_no_media.py` — 12 tests, all passing. §20's list, run in one campaign
with an empty provider registry:
```text
several story turns ✓ Undo ✓
state extraction ✓ Redo ✓
memory + summary activity ✓ Retry ✓
knowledge retrieval ✓ Save Point restore ✓
restart ✓ *
```
\* In this suite the restart is a fresh session reading what was written, not a
new process. The **genuine spawned-process restarts are in
`test_m10_lineage.py`** (§G), where they carry more weight: they run with
profiles written and packets built, and check the derived `scene_id` is the same
in a process that never saw the first one.
with **no** media warning, **no** media connection attempt, **no**
missing-provider error, and no media schema requirement reaching narration.
Also checked:
- **The prompt is unchanged.** No context section labelled `media`,
`scene_packet` or `visual_profile*` exists; the strings `visual_profile` and
`scene_id` appear nowhere in the assembled system or story prompt.
- **Ordinary play writes no media row.**
- **No media setting exists** — neither in the settings API response nor as a
column on `Settings`. A setting that exists is a setting that can be pointed at
a cloud by mistake.
- **The turn path cannot reach the media package**, checked by parsing imports
rather than by grepping text.
- The app serves, plays and passes health checks with `providers.registered() == {}`.
---
## J. Security and locality
**Outbound destinations added: none.** No module under `app/media/` references
`httpx`, `requests`, `urllib.request`, `socket`, `aiohttp` or `subprocess`, and
that is a test, not an inspection. No provider adapter ships, so there is nothing
to connect *to*; the registry is empty at import and stays empty.
**Endpoint architecture — stricter than narration.**
`providers.endpoint_rejection_reason(url)` applies `endpoints.rejection_reason`
first (the shared policy: every address the hostname resolves to must be
loopback, RFC1918, link-local, IPv6 ULA or CGNAT; known cloud inference hosts are
refused by name; the check is on the **resolved address**, so
`localhost.evil.example` does not pass) and **then requires loopback in
addition**. Verified:
| Endpoint | Verdict |
| --- | --- |
| `http://127.0.0.1:8188/` | allowed |
| `http://192.168.x.x:8188/` (trusted LAN — allowed for *narration*) | **refused** for media |
| `https://api.openai.com/v1`, `https://replicate.com`, `http://8.8.8.8:8188`, `http://example.com` | refused |
This is deliberately narrower than `SECURITY-THREAT-MODEL.md` §73 permits, and
§42A now records the discrepancy rather than leaving it to be found. The
reasoning: a picture of a scene carries the scene with it, and a GPU rendering
someone's campaign is a machine that person is sitting at.
**No TLS verification bypass** was introduced. There is no `verify=False`, no
`-k`, and no new HTTP client at all; `tlstrust.py` is untouched and
`test_tls_trust.py` passes.
**Filesystem/media exposure: none.** No asset is stored, no directory is served,
no path comes from a caller. M10 adds no file-serving route.
**CSP/CORS: unchanged.** No frontend file was modified, no new origin is
contacted, and `test_offline_assets.py` (which reads the built SPA and the CSP)
passes.
**Cloud dependency: none.** Nothing was installed to demonstrate an interface —
no ComfyUI, no diffusers, no Whisper, no Kokoro, no model download. The
`requirements.txt` diff is empty.
**One new disclosure boundary, and it is not a network one.** The Scene Packet is
the input a future provider would receive, so its contents are a disclosure
decision — covered in §H.
---
## K. Performance and storage
Measured with `backend/tools/m10_media_cost.py`, a 120-turn campaign played
through the real turn engine and the real state pipeline, with SQL statements
counted by `tools/dbmeter.py`.
```text
120 turns, 245 action rows
scene records M10 wrote
visual_profiles rows 2 one per profiled entity, written once
scene rows 0 M10 adds no scenes table
M5 per-position state snapshots 245 already there; the scene lives here
bytes added to the database
empty database 188416 B
after the fixture campaign 196608 B
after 120 more turns 1056768 B
visual profile content 367 B 0.035% of the database
profile duplication: campaign-scoped against per-position
as stored, once per entity 367 B
if snapshotted per position 89915 B x245
packet: persisted or constructed
rows written while building one 0
packet rows in any table 0 built on read, never stored
build time, 2 turns 16.6 ms
build time, 122 turns 12.6 ms
current-scene query behaviour
statements, 2 turns 5
statements, 122 turns 4
does not grow with the campaign
```
The packet's four statements at turn 122 are: the adventure, its visual
profiles, the current user, and the head's branch. **Nothing walks the
transcript**, which is the property that matters — a scene derivation that
scanned actions would have made every future depiction O(turns).
The `x245` row is the measurement behind §C.3: per-position profiles would have
stored the same 367 bytes 245 times in this campaign to say something that never
varies.
**Storage added to an ordinary campaign that uses no profiles: one empty table.**
---
## L. Regression counts
All runs on this branch, after every M10 change, on 2026-09-07.
| Suite | Result | Time |
| --- | --- | --- |
| **Backend, full** | **1,191 passed, 14 skipped, 0 failed** — 1,205 collected, exit 0 | 862.9 s |
| **M10-specific** | **89 passed** (39 hooks + 8 lineage + 12 authority + 12 no-media + 18 bundle/migration) | 56.4 s |
| **Frontend component suite** | **145 passed**, 12 files, 0 failed | 8.17 s |
| **Lint** (`oxlint`) | **0 errors**, 15 warnings, exit 0 | — |
| **Production build** (`vite build`) | clean — `index-DcHbz7ga.js` 389.66 kB (gzip 119.25 kB), `index-B55Q8MSM.css` 47.50 kB | 0.58 s |
| **Docker** (`--no-cache`) | **exit 0**; image runs and imports `app.media` with `registered() == {}` | 25.2 s |
| **Browser** | **not run — M10 adds no reader-facing surface.** See §L.2 |
The 14 skips are M9's and are unchanged — confirmed by running the five files
that carry a skip condition on their own (26 passed, 14 skipped): seven need a
second machine or an environment the suite cannot create; the rest need a real
local model.
The 15 lint warnings are pre-existing (`no-unused-vars` in two test files, and
`react/only-export-components` in components that export a constant beside a
component). **M10 modified no frontend file**, so the count is M9's, unchanged.
### L.1 By group
Each group was run on its own, so a reviewer can check a claim without running
the whole suite. Every group is green.
| Group | Files | Result |
| --- | --- | --- |
| **M10** | `test_m10_media_hooks` (39), `_lineage` (8), `_authority` (12), `_no_media` (12), `_bundle` (18) | **89 passed** — 56.4 s |
| **M9 recovery** | `test_m9_backup`, `_clean_import`, `_corrupt_bundles`, `_legacy_bundles`, `_portability`, `test_bundle_v2` | **173 passed** — 371.3 s |
| **Migration** | `test_tree_migration`, `test_knowledge_migration`, `test_pre_m5_compatibility`, `test_snapshot_compression` | **50 passed** — 27.8 s |
| **M5/M6/M7 state, memory, knowledge** | `test_narrative_state`, `test_worldstate_integration`, `test_memory_nodes`, `test_memory_retrieval`, `test_context_memory`, `test_imported_knowledge`, `test_knowledge_retrieval_quality` | **216 passed** — 113.7 s |
| **M3/M4 history** | `test_story_tree_baseline`, `test_branch_forking`, `test_head_cursor`, `test_save_points`, `test_take_state`, `test_process_restart` | **137 passed** — 139.6 s |
| **Security / offline** | `test_egress`, `test_endpoint_policy`, `test_local_only_surface`, `test_offline_assets`, `test_tls_trust` | **95 passed** — 18.4 s |
The security group is the one that carries §J's claims: no outbound route, the
endpoint policy, the local-only API surface, the offline asset and CSP checks,
and the TLS trust union. All were green before M10 and are green now.
### L.2 On the browser run
M10 adds **no reader-facing surface**: no page, no control, no copy, no route in
the SPA. `git status` shows no file under `frontend/` modified. The five new
endpoints are backend-only and nothing in the browser calls them. A real-browser
regression pass would therefore be re-verifying M8/M9's surfaces against a build
identical to theirs, and its evidence would be M9's evidence. The frontend
component suite and the production build were run anyway, and are green.
This is stated as a decision, not an omission: **if the reviewer wants a browser
pass as a matter of process, it has not been done.**
---
## M. Findings
### M.1 A redundant index migration made two databases disagree — **introduced by M10, fixed here**
**Severity:** low in effect, moderate in kind. **Blocker:** no — fixed.
**Owner:** M10 (closed).
M10 first added migration 93, `CREATE INDEX IF NOT EXISTS ix_visual_profiles_adventure
ON visual_profiles (adventure_id)`. But `VisualProfile.adventure_id` declares
`index=True`, so `create_all` already builds `ix_visual_profiles_adventure_id` —
on a fresh install *and* on an existing database, since `create_all` runs before
the migration loop. The result:
```text
upgraded from 92: ix_visual_profiles_adventure, ix_visual_profiles_adventure_id
fresh install: ix_visual_profiles_adventure_id
```
Two schemas differing by which path the file took, which is the thing a migration
exists to prevent, plus a redundant index on every upgraded database.
**Found by** `test_a_fresh_database_arrives_at_the_same_place`, which compares a
fresh schema against an upgraded one. Neither database examined on its own would
have shown it. **Fixed by removing the migration**, not by renaming the index:
migration 92's own comment already records the rule for the M7 tables — when the
index is declared on the column there is nothing left for a `CREATE INDEX` to do.
`LATEST_VERSION` returns to 92.
### M.2 The state model refuses an over-large scene, and a test asked for one — **not a defect**
While writing the packet's bounds test I sent 43 entries in `set_scene`'s
`present`, exceeding `validate.MAX_LABELS = 40`. The event was correctly refused
and the previous scene stayed, so the test measured the wrong scene and failed.
Recorded because the diagnosis matters: the product was right and the test was
wrong. The test now uses 33 and asserts its own precondition, so it cannot
silently measure a scene it did not set.
### M.3 The Story Engine vocabulary grep hit the file that forbids the vocabulary — **test defect, fixed**
§9's rule is that no provider vocabulary (ComfyUI, Whisper, `num_inference_steps`,
LoRA) appears in the Story Engine. The first version of the test grepped the
whole backend and hit `providers.py`, whose docstrings *name* those things
precisely in order to exclude them. Fixed by scoping the grep to the story
engine, and by adding a complementary test that checks the seam **by behaviour**:
no provider registered, and no networking import anywhere under `app/media/`.
### M.3a A text search for "media" matched "im**media**tely" — **test defect, fixed**
The first version of the no-media import check read each turn-path module and
required the string `media` to be absent. `narrative/store.py` contains the word
*immediately*, so the test failed on a module that imports nothing. It now parses
the file and inspects its **import statements**, which is what the claim was
always about. Recorded because the failure looked briefly like a real coupling
and was not, and because the fixed version is the stronger test: a module could
have imported the package while never spelling the word in prose.
### M.4 `ambience` is present and empty — **known gap, deliberate**
**Severity:** low. **Blocker:** no. **Owner:** whichever milestone builds a
coordinator.
The packet's `ambience` object has the right shape and no content, because
filling it would mean extending the `set_scene` event — a change on the prompt
path, asking the narrator for something new, which is outside M10's scope. A
future provider gets a stable shape today and content when someone decides the
narrator should be asked.
### M.5 Deleting profiles is not recoverable from within the app — **known limit, stated**
**Severity:** low. **Blocker:** no.
M9's rebuildable data can be regenerated; a visual profile cannot, because it is
something a reader wrote. It travels in the bundle, so a backup or an export
recovers it, and deletion is per-entity and explicit. There is no undo for it,
and none was invented — that would be a second history model beside the story's.
**No pre-existing defect was found in M9's or earlier work during this
milestone.** The full backend suite was green before M10 began and is green now.
---
## N. Planning changes
| Document | Change | Why |
| --- | --- | --- |
| `planning/DATA-MODEL.md` | **New §28A** — media extension points as implemented | §20/§27/§28 describe a scene table and job/asset tables. Only one of the three exists, and a reader of the conceptual model needs to know which, and why the scene is derived from §20's own data rather than stored beside it. |
| `planning/TECHNICAL-DESIGN.md` | **New §15.1** under Scene and Future Media Boundary | §15 said "persist or derive". The answer is *derive*, and the reason (M5 already persisted it) is the milestone's central fact. Also records the packet's exclusions, the STT asymmetry and the endpoint policy. |
| `planning/MEDIA-EXTENSION-CONTRACT.md` | **New §90**, appended | The contract is Phase 0B design and stays readable as such. §90 records what was built, the three places implementation answered an open question (§5 already satisfied, §12 drawn wider, §7-9 collapsed into one table), and what is deliberately unbuilt. |
| `planning/BUILD-MILESTONES.md` | **M10 status block** | Milestone status, the shaping finding, what shipped, and the defect its own tests caught. The four post-M8 playtest findings above it are untouched and still M11's. |
| `planning/V1-ACCEPTANCE-TESTS.md` | **K01-K04 results** | Acceptance evidence. K01 records that it was already passing; K04 records which of its two clauses it passes on. |
| `planning/SECURITY-THREAT-MODEL.md` | **New §42A** | The trust boundary did not widen, but in one place the implementation is deliberately **narrower** than §73 permits. A stricter implementation than the model describes is still a discrepancy, and an undocumented one becomes an accidental relaxation later. |
| `planning/VERSION.md` | **v3.6 entry** | Records this milestone's documentation changes and the two decisions (no format bump, no migration). |
| `planning/README.md` | Status, milestone map, reading order, report rotation | M10 is implemented; M11 is next. Records **why M9's report stays in `reports/`** against the usual rotation: M9 is not accepted, and M10's baseline is M9's. |
| `README.md` | `media/` in the architecture map; `VisualProfile` in the model list; test count 920 → 1,191 | The map is the first thing a new reader reads. |
| `DEVELOPMENT.md` | The stricter future media endpoint rule; a note that the suite takes ~15 minutes | Whoever adds the first provider should find the rule before writing the adapter. |
No planning document was rewritten, and no earlier milestone's evidence was
edited.
---
## O. M11 handoff
### O.1 M10's residual risk
1. **K04 is satisfied structurally, not physically.** No `media_jobs` or
`media_assets` table exists. A future coordinator will design them, and the
contracts here constrain that design only loosely. *Risk: low — the expensive
part (a stable scene identity) is fixed; the cheap part (two tables) is not.*
2. **`ambience` is an empty shape** (§M.4). Filling it means extending
`set_scene`, which changes what the narrator is asked for. *Risk: low; a
provider adapter written today would find the fields and no values.*
3. **The seam has no consumer, so it is unexercised by real use.** Every test
here uses a dummy provider. The contracts are shaped by the contract document
and by what the state model can supply, not by an adapter that had to work
against a real generator. *Risk: moderate for the interfaces' ergonomics, nil
for the story engine — the first real adapter may want the packet reshaped,
and nothing in the story depends on its shape.*
4. **A visual profile cannot be recovered from within the app** (§M.5).
5. **Profiles are not surfaced to the reader at all.** They are API-only. Whoever
builds a media UI owns the browser surface, and no reader-facing vocabulary
for them has been invented — deliberately, since M10 was told not to introduce
reader-facing branding or surfaces.
### O.2 M9 carry-forward still relevant
All six of M9's residual risks are unchanged by M10 — none was addressed and none
was made worse:
| M9 residual | Status after M10 |
| --- | --- |
| Bundle ceiling ~279 turns | unchanged. M10 adds ~370 bytes per campaign to a file whose ceiling is set by per-position state; it does not move the number. |
| `quick_check` rather than `integrity_check` | unchanged; re-exercised on a migrated database (§D) |
| No scheduled backup | unchanged |
| Stale `chunk_id` in a restored snapshot | unchanged |
| Importing machine's context window may differ | unchanged; still M11's |
| This machine cannot drive a file into/out of the browser | unchanged; still M11's, and still the cheapest fix is an unconfined Firefox or Xvfb |
M9's two items "carried to M10" are both **closed**: media tables were owned and
the decision is recorded (§F, K04); the `scene` section of the state document was
indeed the extension point, and is what the packet is built from.
### O.3 The post-M8 hands-on findings — still M11's, untouched
M10 neither implemented nor tested any of them, as its brief required. They are
listed here so they cannot be lost when M9's report is eventually archived; the
durable copy is in `BUILD-MILESTONES.md`.
| Finding | Owner |
| --- | --- |
| **A.** The browser tab still reads `AI D&D` | M11 release polish |
| **B.** After Undo, the reader cannot tell where they are | M11 UX/release polish |
| **C.** The narration-length setting has no measurable effect | M11 realistic-model behaviour |
| **D.** Character identity / coreference confusion — root cause **unknown**, and the playtest database was destroyed | M11 realistic-model / context diagnostic |
**On D specifically:** M10's fixture is deliberately the same shape — a
protagonist, two more characters in the room, and a fourth who is not — but
**nothing in M10 asserts anything about coreference**, and the fixture's
resemblance is not evidence about the finding. M11 still owes the explicit
diagnostic. What M10 does contribute, incidentally, is that a visual profile
attaches to the canonical `entity_key` rather than to a display name, so whatever
M11 concludes about identity, profiles are keyed to the thing the state model
considers one character.
### O.4 The context-window issue
Unchanged and still M11's: a deployment's enforced context window may be far
below `Settings.context_token_budget` (Ollama defaults to 4,096 when it sees no
VRAM). Documented in `DEVELOPMENT.md`; detection, the Settings warning question,
and the 100-turn certification all remain open. M10 sends nothing to a model and
does not touch it.
### O.5 Realistic-model / 100-turn work
Untouched by M10, and M10 adds no new realistic-model obligation: the media seam
has no model in it. The 100-turn certification, the contrast/focus measurement,
the WCAG audit and the streaming-import question all carry forward unchanged.
---
## P. Final verdict
**1. Is M10's Definition of Done satisfied?**
> *Future media providers can be added through defined local interfaces without
> redesigning core story authority/history.*
**Yes.** A provider is added by implementing a `Protocol` and registering it;
it receives a Scene Packet built from the authoritative state at a position, and
returns a `MediaResult`. Nothing in story authority or history changes to
accommodate it — proved by the fact that nothing in `app/media/` can even import
the code that writes state, and that every media operation leaves the
authoritative document byte-identical.
The honest qualification: **no real adapter has been written against these
interfaces**, so their ergonomics are untested (§O.1.3). The story engine's
independence, which is what the Definition of Done is actually about, is tested.
**2. Are K01-K03 all PASS?** **Yes** — all three, with evidence in §F and §G.
K01 additionally records that it was already passing before M10 began.
**3. What is K04's exact status?** **PASS on the acceptance text's deferred
branch** ("if media tables are deferred: architecture/types should demonstrate
equivalent extension point"), demonstrated by a dummy provider producing an asset
against a real packet with the story model unchanged. **Not PASS on the first
branch**, since no media table is physically implemented. A reviewer who requires
physical tables in v1 should read K04 as **PARTIAL**; the decision and its
reasoning are in §F.
**4. Can all ordinary story operation run with zero media provider?** **Yes** —
§I. Turns, state extraction, memory, summaries, knowledge retrieval, Undo, Redo,
Retry and Save Point restore, with an empty registry, no media setting in
existence, no warning, no connection attempt and no media row written. Restart is
covered as a genuine spawned-process restart in the lineage suite (§G); the
no-media suite's own restart step is a fresh session read, and §I says so.
**5. Can a future image provider be added without modifying story authority or
history?** **Yes.** It implements `MediaProvider`, is handed a packet, and
returns a result. No story table, event type, or history operation changes.
**6. Can a future video provider consume a multi-turn scene representation
without redesigning history?** **Yes.** `packet.build(..., start=, end=)` takes a
depth range on the branch and the resulting `scene_id` encodes it
(`c7:b3:4-9`). The range is expressed in the coordinates history already uses, so
a multi-turn packet is a read of existing structure rather than a new one.
**7. Does future STT feed editable draft input rather than authoritative state?**
**Yes, structurally.** `TranscriptionProvider.transcribe` returns a
`DraftTranscription` with `editable=True` and no commit method. A transcriber
cannot submit; the ordinary authoritative path is the only way in.
**8. Is scene/media data branch-safe?** **Yes** — §G. The scene follows the
active lineage through Undo, Redo, Retry, divergence, Save Point restore and two
genuine process restarts, because it *is* the authoritative state rather than a
copy of it. Profiles are campaign-scoped by design and are stable across all of
those, which is the correct behaviour for appearance and is argued in §C.3.
**9. Did M10 create any network dependency?** **No.** No dependency added, no
HTTP client, no socket, no subprocess, no provider adapter, no model download, no
media endpoint setting. The endpoint *policy* that a future provider will meet is
stricter than the one narration uses: loopback only.
**10. Is there any blocker before M11?** **No blocker.** Two things a reviewer
should decide rather than inherit:
- whether **K04 on the deferred branch** is acceptable for v1, or whether media
tables must be physically present (§F);
- whether **no browser regression pass** is acceptable given that M10 modified no
frontend file (§L.2).
Neither is a defect; both are decisions that belong to the reviewer.