The media extension contract asks for a scene snapshot a future image or video provider could be handed: location, who is present, what they hold, what must stay true, and where in the story it sits. Building one was the milestone's obvious first task, and it was the wrong one. That snapshot has existed since M5. `narrative_state["scene"]` holds the summary, the location, the cast and the coordinate it was written at; a validated `set_scene` event writes it, every position snapshots it, and every head move restores it. It survives Undo, Redo, Retry, divergence, Save Point restore and a process restart because it is the authoritative state rather than a copy of it. So there is no scenes table here. A second scene store would have been a second answer to "where is the story now", with its own lineage rules to get wrong — and the lineage rules are the expensive part, which is the argument for reusing the ones that already work rather than against it. The Scene Packet is derived on read, and its identity is computed from the campaign and the position rather than allocated: the same position yields the same id in another process, after a restart, and after the packet is thrown away and rebuilt, with no row to keep in step. That is the part of a future media_assets table that would be expensive to retrofit, so it is fixed now even though the table is not built. One table, then: visual_profiles, the only thing the contract's scene list asks for that nothing already stored. Campaign-scoped and not per-position, because a character does not change appearance when the story forks — a reader who diverged would otherwise lose their cast, and the same descriptors would land in every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn campaign to say something that never varies. Keyed by the M5 entity key rather than a new identity namespace, and one table for characters, locations and items alike, because a location is an entity with a type and splitting them would reintroduce the genre shape M5 spent a milestone removing. What the packet leaves out is the more interesting half. Not the transcript, and not imported knowledge — none of it, not merely the sources marked hidden. The rule is what the story established at this position, not everything the narrator was told, and drawing it by class is what makes it hold for a secret nobody thought to mark. A hidden Canon source proves it, with a positive control showing the narrator did receive the sentinel the packet does not carry. Once a validated event puts the observer in the room, the observer is in the packet: that is no longer narrator-only knowledge, and a packet that hid it would be hiding the story from itself. The providers are contracts and nothing else. Protocols for image, video, audio, speech and transcription, an empty registry, no adapter, no dependency, no socket, and no media setting to point anywhere — a setting that exists can be pointed at a cloud by mistake. A future provider endpoint must be loopback, stricter than narration's trusted-LAN allowance, because a picture of a scene carries the scene with it. Transcription returns an editable draft with no commit method, so STT structurally cannot bypass the authoritative path. Nothing here can write the story. Not by convention: no module under media/ imports the code that writes state, no media event type exists in the state vocabulary, and every test in the authority suite compares the authoritative document byte for byte either side of a media operation — including one where a provider insists Alice is in a red coat in a corridor, and the campaign goes on disagreeing. One defect, found by the milestone's own tests. M10 first added a migration creating an index that create_all already builds from the column, so an upgraded database ended up with two indexes and a fresh install with one. Comparing the two schemas is what caught it; neither database examined alone would have. The migration is gone rather than renamed, and the right number of migrations for a new table whose indexes are declared on its columns is zero. Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145 passed. Lint, production build and Docker build clean. No frontend file changed: M10 adds no reader-facing surface, and ordinary play — turns, state, memory, knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media configuration, no warning, no connection attempt and no media row written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
226 lines
9.0 KiB
Python
226 lines
9.0 KiB
Python
"""What M10 costs a campaign, measured rather than argued.
|
|
|
|
python -m tools.m10_media_cost [--turns 60]
|
|
|
|
Run from `backend/`. Plays a campaign of `--turns` turns with the real prompt
|
|
builder and the real state pipeline, then reports the five numbers §21 of the
|
|
M10 brief asks for.
|
|
|
|
Four of them are expected to be zero or near it, and that is the point: M10's
|
|
central design decision was that **the scene snapshot already exists**, so the
|
|
milestone persists nothing per scene and nothing per turn. A design claim like
|
|
that is cheap to make and easy to get wrong by one accidental write, so it is
|
|
measured here against a campaign long enough for a per-turn cost to show.
|
|
|
|
scene records written by M10 expected 0, and the scenes that do
|
|
exist are M5's, counted for contrast
|
|
bytes added to the database one row per profiled entity, once
|
|
profile duplication what per-position profiles would have
|
|
cost, against what campaign-scoped
|
|
profiles do cost
|
|
packet: persisted or constructed rows written while building one
|
|
current-scene query behaviour statements per packet, at 10 turns and
|
|
at N turns — a number that grows with
|
|
the campaign is a scan
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import json
|
|
import os
|
|
import sys
|
|
import tempfile
|
|
import time
|
|
from pathlib import Path
|
|
|
|
_HERE = Path(__file__).resolve().parent
|
|
sys.path.insert(0, str(_HERE.parent / "tests"))
|
|
|
|
_DB = tempfile.NamedTemporaryFile(suffix="-m10-cost.db", delete=False)
|
|
_DB.close()
|
|
os.environ["AIDND_DB_PATH"] = _DB.name
|
|
os.environ.pop("AIDND_DATABASE_URL", None)
|
|
os.environ.pop("DATABASE_URL", None)
|
|
|
|
from fastapi import Depends # noqa: E402
|
|
from fastapi.testclient import TestClient # noqa: E402
|
|
|
|
import m10_fixture # noqa: E402
|
|
from app import auth, limits, memorybank, models # noqa: E402
|
|
from app.database import Base, SessionLocal, engine, get_db # noqa: E402
|
|
from app.main import app # noqa: E402
|
|
from app.routers import adventures # noqa: E402
|
|
from fakes import ScriptedProvider, state_block # noqa: E402
|
|
from tools import dbmeter # noqa: E402
|
|
|
|
PROSE = (
|
|
"Roger pulled the whiteboard marker apart while he talked, which was how "
|
|
"everyone knew the meeting had stopped being about the agenda. Alice wrote "
|
|
"nothing down. Outside the glass, somebody wheeled a trolley of monitors "
|
|
"past the door and did not look in."
|
|
)
|
|
|
|
|
|
class _Stub:
|
|
async def complete(self, system, prompt, **kwargs):
|
|
return "The meeting went on for some time."
|
|
|
|
async def embed(self, texts):
|
|
return [[1.0, 0.5, 0.25, 0.125] for _ in texts]
|
|
|
|
|
|
def _setup() -> tuple[TestClient, int]:
|
|
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
|
memorybank.embedding_provider = lambda s: _Stub()
|
|
memorybank.summary_provider = lambda s: _Stub()
|
|
limits.check_row_cap = lambda *a, **k: None
|
|
Base.metadata.create_all(bind=engine)
|
|
with SessionLocal() as db:
|
|
user = models.User(is_guest=False, email="m10cost@example.com")
|
|
db.add(user)
|
|
db.flush()
|
|
db.add(models.Settings(
|
|
user_id=user.id, model="cost-model", embedding_model="stub",
|
|
context_token_budget=8192, max_output_tokens=600,
|
|
))
|
|
adventure = models.Adventure(user_id=user.id, title="Cost",
|
|
auto_summarize=True, memory_bank_enabled=True)
|
|
db.add(adventure)
|
|
db.flush()
|
|
db.add(models.Action(adventure_id=adventure.id, type="start",
|
|
text="Bill badges in on a Tuesday morning."))
|
|
db.commit()
|
|
adv_id, user_id = adventure.id, user.id
|
|
app.dependency_overrides[auth.get_current_user] = (
|
|
lambda db=Depends(get_db): db.get(models.User, user_id)
|
|
)
|
|
return TestClient(app), adv_id
|
|
|
|
|
|
def _db_bytes() -> int:
|
|
return Path(_DB.name).stat().st_size
|
|
|
|
|
|
def _counts(adv_id: int) -> dict:
|
|
with SessionLocal() as db:
|
|
return {
|
|
"actions": db.query(models.Action).filter(
|
|
models.Action.adventure_id == adv_id).count(),
|
|
"visual profiles": db.query(models.VisualProfile).filter(
|
|
models.VisualProfile.adventure_id == adv_id).count(),
|
|
"M5 per-position state snapshots": db.query(models.Action).filter(
|
|
models.Action.adventure_id == adv_id,
|
|
models.Action.narrative_state_after.isnot(None)).count(),
|
|
}
|
|
|
|
|
|
def _profile_bytes(adv_id: int) -> int:
|
|
with SessionLocal() as db:
|
|
rows = db.query(models.VisualProfile).filter(
|
|
models.VisualProfile.adventure_id == adv_id).all()
|
|
return sum(
|
|
len(json.dumps({"entity_key": r.entity_key,
|
|
"descriptors": r.descriptors,
|
|
"features": r.features,
|
|
"style_notes": r.style_notes}).encode("utf-8"))
|
|
for r in rows
|
|
)
|
|
|
|
|
|
def _packet_statements(client, adv_id: int, meter: dbmeter.Meter, label: str):
|
|
with meter.scope(label) as scope:
|
|
started = time.perf_counter()
|
|
response = client.get(f"/api/adventures/{adv_id}/scene-packet")
|
|
seconds = time.perf_counter() - started
|
|
response.raise_for_status()
|
|
return scope, seconds
|
|
|
|
|
|
def main() -> int:
|
|
parser = argparse.ArgumentParser(description=__doc__)
|
|
parser.add_argument("--turns", type=int, default=60)
|
|
args = parser.parse_args()
|
|
|
|
client, adv_id = _setup()
|
|
empty_bytes = _db_bytes()
|
|
m10_fixture.build(client, adv_id)
|
|
|
|
meter = dbmeter.Meter()
|
|
meter.attach(engine)
|
|
try:
|
|
early_scope, early_seconds = _packet_statements(
|
|
client, adv_id, meter, "packet at 2 turns")
|
|
|
|
before_play = _db_bytes()
|
|
for turn in range(1, args.turns + 1):
|
|
ScriptedProvider.replies = [
|
|
f"{PROSE} [{turn}]\n" + state_block([
|
|
{"type": "set_scene",
|
|
"summary": f"The meeting reaches item {turn}.",
|
|
"location": "office",
|
|
"present": ["bill", "alice", "roger"]},
|
|
])
|
|
]
|
|
client.post(f"/api/adventures/{adv_id}/actions",
|
|
json={"type": "do", "text": f"item {turn}"}
|
|
).raise_for_status()
|
|
|
|
rows_before = _counts(adv_id)
|
|
late_scope, late_seconds = _packet_statements(
|
|
client, adv_id, meter, f"packet at {args.turns + 2} turns")
|
|
rows_after = _counts(adv_id)
|
|
finally:
|
|
meter.detach()
|
|
|
|
played_bytes = _db_bytes()
|
|
profile_bytes = _profile_bytes(adv_id)
|
|
|
|
print(f"\n{args.turns} turns, {rows_before['actions']} action rows\n")
|
|
|
|
print("scene records M10 wrote")
|
|
print(f" visual_profiles rows {rows_before['visual profiles']:>8}"
|
|
" (one per profiled entity, written once)")
|
|
print(" scene rows 0"
|
|
" M10 adds no scenes table")
|
|
print(f" M5 per-position state snapshots "
|
|
f"{rows_before['M5 per-position state snapshots']:>8}"
|
|
" already there since M5; the scene lives here")
|
|
|
|
print("\nbytes added to the database")
|
|
print(f" empty database {empty_bytes:>8} B")
|
|
print(f" after the fixture campaign {before_play:>8} B")
|
|
print(f" after {args.turns} more turns".ljust(36)
|
|
+ f"{played_bytes:>8} B")
|
|
print(f" visual profile content {profile_bytes:>8} B"
|
|
f" {100 * profile_bytes / max(played_bytes, 1):.3f}% of the database")
|
|
|
|
per_position = profile_bytes * rows_before["M5 per-position state snapshots"]
|
|
print("\nprofile duplication: campaign-scoped against per-position")
|
|
print(f" as stored, once per entity {profile_bytes:>8} B")
|
|
print(f" if snapshotted per position {per_position:>8} B"
|
|
f" x{per_position / max(profile_bytes, 1):.0f}")
|
|
|
|
print("\npacket: persisted or constructed")
|
|
print(f" rows written while building one "
|
|
f"{rows_after['visual profiles'] - rows_before['visual profiles']:>8}")
|
|
print(" packet rows in any table 0 built on read, never stored")
|
|
print(f" build time, 2 turns {early_seconds * 1000:>8.1f} ms")
|
|
print(f" build time, {args.turns + 2} turns".ljust(36)
|
|
+ f"{late_seconds * 1000:>8.1f} ms")
|
|
|
|
print("\ncurrent-scene query behaviour")
|
|
print(f" statements, 2 turns {early_scope.total.statements:>8}")
|
|
print(f" statements, {args.turns + 2} turns".ljust(36)
|
|
+ f"{late_scope.total.statements:>8}")
|
|
verdict = ("does not grow with the campaign"
|
|
if late_scope.total.statements <= early_scope.total.statements
|
|
else "GROWS — the scene is being scanned, not read")
|
|
print(f" {verdict}")
|
|
print("\n" + dbmeter.render_scope(late_scope, statements=6))
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
raise SystemExit(main())
|