Files
JesseMarkowitzandClaude Opus 5 1013c94eb1
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
M10: the seam for media, and no media
The media extension contract asks for a scene snapshot a future image or video
provider could be handed: location, who is present, what they hold, what must
stay true, and where in the story it sits. Building one was the milestone's
obvious first task, and it was the wrong one. That snapshot has existed since
M5. `narrative_state["scene"]` holds the summary, the location, the cast and the
coordinate it was written at; a validated `set_scene` event writes it, every
position snapshots it, and every head move restores it. It survives Undo, Redo,
Retry, divergence, Save Point restore and a process restart because it is the
authoritative state rather than a copy of it.

So there is no scenes table here. A second scene store would have been a second
answer to "where is the story now", with its own lineage rules to get wrong —
and the lineage rules are the expensive part, which is the argument for reusing
the ones that already work rather than against it. The Scene Packet is derived
on read, and its identity is computed from the campaign and the position rather
than allocated: the same position yields the same id in another process, after a
restart, and after the packet is thrown away and rebuilt, with no row to keep in
step. That is the part of a future media_assets table that would be expensive to
retrofit, so it is fixed now even though the table is not built.

One table, then: visual_profiles, the only thing the contract's scene list asks
for that nothing already stored. Campaign-scoped and not per-position, because a
character does not change appearance when the story forks — a reader who
diverged would otherwise lose their cast, and the same descriptors would land in
every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn
campaign to say something that never varies. Keyed by the M5 entity key rather
than a new identity namespace, and one table for characters, locations and items
alike, because a location is an entity with a type and splitting them would
reintroduce the genre shape M5 spent a milestone removing.

What the packet leaves out is the more interesting half. Not the transcript, and
not imported knowledge — none of it, not merely the sources marked hidden. The
rule is what the story established at this position, not everything the narrator
was told, and drawing it by class is what makes it hold for a secret nobody
thought to mark. A hidden Canon source proves it, with a positive control
showing the narrator did receive the sentinel the packet does not carry. Once a
validated event puts the observer in the room, the observer is in the packet:
that is no longer narrator-only knowledge, and a packet that hid it would be
hiding the story from itself.

The providers are contracts and nothing else. Protocols for image, video, audio,
speech and transcription, an empty registry, no adapter, no dependency, no
socket, and no media setting to point anywhere — a setting that exists can be
pointed at a cloud by mistake. A future provider endpoint must be loopback,
stricter than narration's trusted-LAN allowance, because a picture of a scene
carries the scene with it. Transcription returns an editable draft with no
commit method, so STT structurally cannot bypass the authoritative path.

Nothing here can write the story. Not by convention: no module under media/
imports the code that writes state, no media event type exists in the state
vocabulary, and every test in the authority suite compares the authoritative
document byte for byte either side of a media operation — including one where a
provider insists Alice is in a red coat in a corridor, and the campaign goes on
disagreeing.

One defect, found by the milestone's own tests. M10 first added a migration
creating an index that create_all already builds from the column, so an upgraded
database ended up with two indexes and a fresh install with one. Comparing the
two schemas is what caught it; neither database examined alone would have. The
migration is gone rather than renamed, and the right number of migrations for a
new table whose indexes are declared on its columns is zero.

Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145
passed. Lint, production build and Docker build clean. No frontend file changed:
M10 adds no reader-facing surface, and ordinary play — turns, state, memory,
knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media
configuration, no warning, no connection attempt and no media row written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 03:41:04 -04:00

226 lines
9.0 KiB
Python

"""What M10 costs a campaign, measured rather than argued.
python -m tools.m10_media_cost [--turns 60]
Run from `backend/`. Plays a campaign of `--turns` turns with the real prompt
builder and the real state pipeline, then reports the five numbers §21 of the
M10 brief asks for.
Four of them are expected to be zero or near it, and that is the point: M10's
central design decision was that **the scene snapshot already exists**, so the
milestone persists nothing per scene and nothing per turn. A design claim like
that is cheap to make and easy to get wrong by one accidental write, so it is
measured here against a campaign long enough for a per-turn cost to show.
scene records written by M10 expected 0, and the scenes that do
exist are M5's, counted for contrast
bytes added to the database one row per profiled entity, once
profile duplication what per-position profiles would have
cost, against what campaign-scoped
profiles do cost
packet: persisted or constructed rows written while building one
current-scene query behaviour statements per packet, at 10 turns and
at N turns — a number that grows with
the campaign is a scan
"""
from __future__ import annotations
import argparse
import json
import os
import sys
import tempfile
import time
from pathlib import Path
_HERE = Path(__file__).resolve().parent
sys.path.insert(0, str(_HERE.parent / "tests"))
_DB = tempfile.NamedTemporaryFile(suffix="-m10-cost.db", delete=False)
_DB.close()
os.environ["AIDND_DB_PATH"] = _DB.name
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from fastapi import Depends # noqa: E402
from fastapi.testclient import TestClient # noqa: E402
import m10_fixture # noqa: E402
from app import auth, limits, memorybank, models # noqa: E402
from app.database import Base, SessionLocal, engine, get_db # noqa: E402
from app.main import app # noqa: E402
from app.routers import adventures # noqa: E402
from fakes import ScriptedProvider, state_block # noqa: E402
from tools import dbmeter # noqa: E402
PROSE = (
"Roger pulled the whiteboard marker apart while he talked, which was how "
"everyone knew the meeting had stopped being about the agenda. Alice wrote "
"nothing down. Outside the glass, somebody wheeled a trolley of monitors "
"past the door and did not look in."
)
class _Stub:
async def complete(self, system, prompt, **kwargs):
return "The meeting went on for some time."
async def embed(self, texts):
return [[1.0, 0.5, 0.25, 0.125] for _ in texts]
def _setup() -> tuple[TestClient, int]:
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
memorybank.embedding_provider = lambda s: _Stub()
memorybank.summary_provider = lambda s: _Stub()
limits.check_row_cap = lambda *a, **k: None
Base.metadata.create_all(bind=engine)
with SessionLocal() as db:
user = models.User(is_guest=False, email="m10cost@example.com")
db.add(user)
db.flush()
db.add(models.Settings(
user_id=user.id, model="cost-model", embedding_model="stub",
context_token_budget=8192, max_output_tokens=600,
))
adventure = models.Adventure(user_id=user.id, title="Cost",
auto_summarize=True, memory_bank_enabled=True)
db.add(adventure)
db.flush()
db.add(models.Action(adventure_id=adventure.id, type="start",
text="Bill badges in on a Tuesday morning."))
db.commit()
adv_id, user_id = adventure.id, user.id
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
return TestClient(app), adv_id
def _db_bytes() -> int:
return Path(_DB.name).stat().st_size
def _counts(adv_id: int) -> dict:
with SessionLocal() as db:
return {
"actions": db.query(models.Action).filter(
models.Action.adventure_id == adv_id).count(),
"visual profiles": db.query(models.VisualProfile).filter(
models.VisualProfile.adventure_id == adv_id).count(),
"M5 per-position state snapshots": db.query(models.Action).filter(
models.Action.adventure_id == adv_id,
models.Action.narrative_state_after.isnot(None)).count(),
}
def _profile_bytes(adv_id: int) -> int:
with SessionLocal() as db:
rows = db.query(models.VisualProfile).filter(
models.VisualProfile.adventure_id == adv_id).all()
return sum(
len(json.dumps({"entity_key": r.entity_key,
"descriptors": r.descriptors,
"features": r.features,
"style_notes": r.style_notes}).encode("utf-8"))
for r in rows
)
def _packet_statements(client, adv_id: int, meter: dbmeter.Meter, label: str):
with meter.scope(label) as scope:
started = time.perf_counter()
response = client.get(f"/api/adventures/{adv_id}/scene-packet")
seconds = time.perf_counter() - started
response.raise_for_status()
return scope, seconds
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--turns", type=int, default=60)
args = parser.parse_args()
client, adv_id = _setup()
empty_bytes = _db_bytes()
m10_fixture.build(client, adv_id)
meter = dbmeter.Meter()
meter.attach(engine)
try:
early_scope, early_seconds = _packet_statements(
client, adv_id, meter, "packet at 2 turns")
before_play = _db_bytes()
for turn in range(1, args.turns + 1):
ScriptedProvider.replies = [
f"{PROSE} [{turn}]\n" + state_block([
{"type": "set_scene",
"summary": f"The meeting reaches item {turn}.",
"location": "office",
"present": ["bill", "alice", "roger"]},
])
]
client.post(f"/api/adventures/{adv_id}/actions",
json={"type": "do", "text": f"item {turn}"}
).raise_for_status()
rows_before = _counts(adv_id)
late_scope, late_seconds = _packet_statements(
client, adv_id, meter, f"packet at {args.turns + 2} turns")
rows_after = _counts(adv_id)
finally:
meter.detach()
played_bytes = _db_bytes()
profile_bytes = _profile_bytes(adv_id)
print(f"\n{args.turns} turns, {rows_before['actions']} action rows\n")
print("scene records M10 wrote")
print(f" visual_profiles rows {rows_before['visual profiles']:>8}"
" (one per profiled entity, written once)")
print(" scene rows 0"
" M10 adds no scenes table")
print(f" M5 per-position state snapshots "
f"{rows_before['M5 per-position state snapshots']:>8}"
" already there since M5; the scene lives here")
print("\nbytes added to the database")
print(f" empty database {empty_bytes:>8} B")
print(f" after the fixture campaign {before_play:>8} B")
print(f" after {args.turns} more turns".ljust(36)
+ f"{played_bytes:>8} B")
print(f" visual profile content {profile_bytes:>8} B"
f" {100 * profile_bytes / max(played_bytes, 1):.3f}% of the database")
per_position = profile_bytes * rows_before["M5 per-position state snapshots"]
print("\nprofile duplication: campaign-scoped against per-position")
print(f" as stored, once per entity {profile_bytes:>8} B")
print(f" if snapshotted per position {per_position:>8} B"
f" x{per_position / max(profile_bytes, 1):.0f}")
print("\npacket: persisted or constructed")
print(f" rows written while building one "
f"{rows_after['visual profiles'] - rows_before['visual profiles']:>8}")
print(" packet rows in any table 0 built on read, never stored")
print(f" build time, 2 turns {early_seconds * 1000:>8.1f} ms")
print(f" build time, {args.turns + 2} turns".ljust(36)
+ f"{late_seconds * 1000:>8.1f} ms")
print("\ncurrent-scene query behaviour")
print(f" statements, 2 turns {early_scope.total.statements:>8}")
print(f" statements, {args.turns + 2} turns".ljust(36)
+ f"{late_scope.total.statements:>8}")
verdict = ("does not grow with the campaign"
if late_scope.total.statements <= early_scope.total.statements
else "GROWS — the scene is being scanned, not read")
print(f" {verdict}")
print("\n" + dbmeter.render_scope(late_scope, statements=6))
return 0
if __name__ == "__main__":
raise SystemExit(main())