M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
co-authored by
Claude Opus 5
parent
1ce9972760
commit
44edece67e
@@ -0,0 +1,331 @@
|
||||
"""M9's migration claim, proved against a database M8's own code wrote.
|
||||
|
||||
# from the M8 worktree, using M8's interpreter:
|
||||
python -m tools.m9_migration_proof build <db_path>
|
||||
|
||||
# from the M9 tree, using M9's interpreter:
|
||||
python -m tools.m9_migration_proof open <db_path>
|
||||
|
||||
M9 claims to add no schema change. `git diff` proves that nothing in
|
||||
`migrations.py` or `models.py` moved, which is necessary and not sufficient: a
|
||||
migration can also be *missing*, and the failure then is a database that opens
|
||||
and quietly answers wrongly. The M8 report set the standard here — a database
|
||||
created by today's code and read by today's code proves nothing — so the
|
||||
campaign below is built by a server running the signed M8 commit, from a git
|
||||
worktree, and read back by M9.
|
||||
|
||||
`build` writes a campaign that touches every family M9 changed the handling of:
|
||||
story with an alternate take, a Save Point, a manual state correction, memories
|
||||
and a summary, imported knowledge including a disabled and a narrator-only
|
||||
source, and per-turn context snapshots. It prints what it wrote, as JSON.
|
||||
|
||||
`open` opens that file with the current code, runs the migration path, and
|
||||
checks every one of those against what `build` reported. It also asserts the
|
||||
schema version did not move and that a second open is a no-op, which is what
|
||||
"no migration" means in practice: the stamp is the same number before and after.
|
||||
|
||||
Neither half imports anything from the other. What crosses is the database file
|
||||
and one JSON report on stdout, which is the only way the two builds can be made
|
||||
to talk without one of them importing the other's code.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tests"))
|
||||
|
||||
|
||||
def _app(db_path: str):
|
||||
"""Imports the application against `db_path`. Must run before any app import."""
|
||||
os.environ["AIDND_DB_PATH"] = db_path
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
from fakes import ScriptedProvider, state_block
|
||||
|
||||
class Stub:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "A memory of what had happened by then."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25] for _ in texts]
|
||||
|
||||
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
||||
memorybank.embedding_provider = lambda s: Stub()
|
||||
memorybank.summary_provider = lambda s: Stub()
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
return {
|
||||
"Base": Base, "SessionLocal": SessionLocal, "engine": engine,
|
||||
"models": models, "app": app, "auth": auth, "get_db": get_db,
|
||||
"Depends": Depends, "TestClient": TestClient,
|
||||
"ScriptedProvider": ScriptedProvider, "state_block": state_block,
|
||||
"memorybank": memorybank,
|
||||
}
|
||||
|
||||
|
||||
def _client(ctx, user_id: int):
|
||||
ctx["app"].dependency_overrides[ctx["auth"].get_current_user] = (
|
||||
lambda db=ctx["Depends"](ctx["get_db"]): db.get(ctx["models"].User, user_id)
|
||||
)
|
||||
return ctx["TestClient"](ctx["app"])
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- building
|
||||
|
||||
def build(db_path: str) -> dict:
|
||||
ctx = _app(db_path)
|
||||
ctx["Base"].metadata.create_all(bind=ctx["engine"])
|
||||
models, SessionLocal = ctx["models"], ctx["SessionLocal"]
|
||||
|
||||
db = SessionLocal()
|
||||
try:
|
||||
user = models.User(is_guest=False, email="m9mig@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="m8-model", embedding_model="stub",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Built by M8", auto_summarize=True,
|
||||
memory_bank_enabled=True,
|
||||
campaign_canon={"rules": ["The dead do not return."]},
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=adventure.id, type="start",
|
||||
text="Aldric sits in the Crooked Lantern with Mara.",
|
||||
))
|
||||
db.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
finally:
|
||||
db.close()
|
||||
|
||||
client = _client(ctx, user_id)
|
||||
upload = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": ("canon.md", (
|
||||
"# Westhaven\n\n## The Old Abbey\n\nThe abbey crypt is sealed.\n"
|
||||
).encode(), "text/markdown")},
|
||||
data={"classification": "canon", "always_include": "true"},
|
||||
)
|
||||
assert upload.status_code == 201, upload.text[:300]
|
||||
hidden = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": ("secret.md", (
|
||||
"# The seal\n\nIt was broken once, sixty years ago.\n"
|
||||
).encode(), "text/markdown")},
|
||||
data={"classification": "canon", "visibility": "hidden"},
|
||||
)
|
||||
assert hidden.status_code == 201, hidden.text[:300]
|
||||
disabled = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": ("draft.md", b"# Draft\n\nAn earlier version.\n",
|
||||
"text/markdown")},
|
||||
data={"classification": "reference"},
|
||||
)
|
||||
assert disabled.status_code == 201, disabled.text[:300]
|
||||
assert client.patch(
|
||||
f"/api/adventures/{adv_id}/knowledge/{disabled.json()['id']}",
|
||||
json={"enabled": False},
|
||||
).status_code == 200
|
||||
|
||||
state_block = ctx["state_block"]
|
||||
for turn in range(1, 8):
|
||||
ctx["ScriptedProvider"].replies = [
|
||||
f"The rain keeps on, and Mara says nothing for a while. [{turn}]\n"
|
||||
+ state_block([{"type": "add_fact", "predicate": "tally",
|
||||
"value": turn * 10, "fact_id": f"tally-{turn * 10}"}])
|
||||
]
|
||||
response = client.post(f"/api/adventures/{adv_id}/actions",
|
||||
json={"type": "do", "text": f"ask about turn {turn}"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
if turn == 3:
|
||||
assert client.post(f"/api/adventures/{adv_id}/retry").status_code == 200
|
||||
point = client.post(f"/api/adventures/{adv_id}/checkpoints",
|
||||
json={"name": "Third turn", "note": "A position."})
|
||||
assert point.status_code == 201, point.text[:300]
|
||||
|
||||
correction = client.post(f"/api/adventures/{adv_id}/state/corrections", json={
|
||||
"events": [{"type": "add_fact", "predicate": "keeper", "value": "Mara",
|
||||
"fact_id": "keeper"}],
|
||||
"note": "Established in play.",
|
||||
})
|
||||
assert correction.status_code == 201, correction.text[:300]
|
||||
|
||||
import asyncio
|
||||
asyncio.run(ctx["memorybank"].run_post_turn(adv_id))
|
||||
|
||||
# Two Undos, so the head is behind the retained tip when M9 opens it.
|
||||
for _ in range(2):
|
||||
assert client.post(f"/api/adventures/{adv_id}/undo").status_code == 200
|
||||
|
||||
report = _describe(ctx, client, adv_id)
|
||||
ctx["app"].dependency_overrides.clear()
|
||||
return report
|
||||
|
||||
|
||||
# -------------------------------------------------------------------- reading
|
||||
|
||||
def _describe(ctx, client, adv_id: int) -> dict:
|
||||
"""Everything the other build has to agree with, read through the API."""
|
||||
models, SessionLocal = ctx["models"], ctx["SessionLocal"]
|
||||
page = client.get(f"/api/adventures/{adv_id}").json()
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
version = db.execute(_pragma()).scalar()
|
||||
counts = {
|
||||
name: db.query(model).filter(model.adventure_id == adv_id).count()
|
||||
for name, model in (
|
||||
("actions", models.Action), ("memories", models.Memory),
|
||||
("summaries", models.Summary), ("checkpoints", models.Checkpoint),
|
||||
("state_events", models.StateEvent),
|
||||
("state_proposals", models.StateProposal),
|
||||
("knowledge_sources", models.KnowledgeSource),
|
||||
("knowledge_chunks", models.KnowledgeChunk),
|
||||
)
|
||||
}
|
||||
head = {"branch_id": adventure.head_branch_id, "depth": adventure.head_depth}
|
||||
snapshots = db.query(models.Action).filter(
|
||||
models.Action.adventure_id == adv_id,
|
||||
models.Action.context_snapshot.isnot(None),
|
||||
).count()
|
||||
return {
|
||||
"adventure_id": adv_id,
|
||||
"schema_version": version,
|
||||
"title": page["title"],
|
||||
"canon_rules": page["canon_rules"],
|
||||
"transcript": [a["text"] for a in page["actions"]],
|
||||
"can_undo": page["can_undo"],
|
||||
"can_redo": page["can_redo"],
|
||||
"head": head,
|
||||
"counts": counts,
|
||||
"snapshot_rows": snapshots,
|
||||
"state": client.get(f"/api/adventures/{adv_id}/state").json()["document"],
|
||||
"checkpoints": sorted(
|
||||
(c["name"], c["depth"])
|
||||
for c in client.get(f"/api/adventures/{adv_id}/checkpoints").json()
|
||||
),
|
||||
"knowledge": sorted(
|
||||
(k["original_filename"], k["classification"], k["enabled"],
|
||||
k["visibility"], k["always_include"], k["content_hash"],
|
||||
k["index_state"])
|
||||
for k in client.get(f"/api/adventures/{adv_id}/knowledge").json()
|
||||
),
|
||||
"events": sorted(
|
||||
(e["event_type"], e["source"], json.dumps(e["payload"], sort_keys=True))
|
||||
for e in client.get(
|
||||
f"/api/adventures/{adv_id}/state/events?limit=500").json()
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def _comparable(value):
|
||||
"""`value` as it survives a JSON round trip, so the two builds compare like."""
|
||||
return json.loads(json.dumps(value, sort_keys=True, default=str))
|
||||
|
||||
|
||||
def _pragma():
|
||||
from sqlalchemy import text
|
||||
|
||||
return text("PRAGMA user_version")
|
||||
|
||||
|
||||
def open_it(db_path: str, expected: dict) -> dict:
|
||||
"""Opens an existing database with this build, and checks it against `expected`."""
|
||||
ctx = _app(db_path)
|
||||
from app import migrations
|
||||
|
||||
# This is the migration run. `main` already called `bootstrap` at import.
|
||||
with ctx["engine"].begin() as conn:
|
||||
after_first = conn.execute(_pragma()).scalar()
|
||||
# And again, to prove idempotence: a second run must change nothing.
|
||||
migrations.bootstrap(ctx["engine"])
|
||||
with ctx["engine"].begin() as conn:
|
||||
after_second = conn.execute(_pragma()).scalar()
|
||||
|
||||
adv_id = expected["adventure_id"]
|
||||
with ctx["SessionLocal"]() as db:
|
||||
user = db.query(ctx["models"].User).first()
|
||||
user_id = user.id
|
||||
client = _client(ctx, user_id)
|
||||
actual = _describe(ctx, client, adv_id)
|
||||
|
||||
problems = []
|
||||
for key in ("title", "canon_rules", "transcript", "head", "counts",
|
||||
"snapshot_rows", "state", "checkpoints", "knowledge", "events",
|
||||
"can_undo", "can_redo"):
|
||||
# Compared through JSON, because that is how the other build's answer
|
||||
# arrived: a tuple written by `_describe` comes back as a list, and a
|
||||
# comparison that called that a difference would report ten differences
|
||||
# in a database nothing had changed.
|
||||
if _comparable(actual[key]) != _comparable(expected[key]):
|
||||
problems.append(f"{key}: expected {expected[key]!r}, got {actual[key]!r}")
|
||||
if expected["schema_version"] != after_first:
|
||||
problems.append(
|
||||
f"the schema version moved: {expected['schema_version']} -> {after_first}"
|
||||
)
|
||||
if after_first != after_second:
|
||||
problems.append(
|
||||
f"a second open migrated again: {after_first} -> {after_second}"
|
||||
)
|
||||
|
||||
# And the campaign still works, rather than merely reading correctly.
|
||||
exported = client.get(f"/api/adventures/{adv_id}/export")
|
||||
if exported.status_code != 200:
|
||||
problems.append(f"export failed: {exported.status_code}")
|
||||
else:
|
||||
imported = client.post("/api/adventures/import", json=exported.json())
|
||||
if imported.status_code != 201:
|
||||
problems.append(f"round trip failed: {imported.text[:300]}")
|
||||
elif exported.json()["format"] != "ai-dnd-adventure-v3":
|
||||
problems.append("the M8 database did not export as v3")
|
||||
redo = client.post(f"/api/adventures/{adv_id}/redo")
|
||||
if redo.status_code != 200:
|
||||
problems.append(f"Redo failed on the migrated campaign: {redo.status_code}")
|
||||
|
||||
ctx["app"].dependency_overrides.clear()
|
||||
return {
|
||||
"schema_version_before": expected["schema_version"],
|
||||
"schema_version_after": after_first,
|
||||
"schema_version_second_open": after_second,
|
||||
"problems": problems,
|
||||
"checked": {
|
||||
"families": 12, "snapshot_rows": actual["snapshot_rows"],
|
||||
"counts": actual["counts"],
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("mode", choices=("build", "open"))
|
||||
parser.add_argument("db_path")
|
||||
parser.add_argument("--expected", help="the JSON `build` printed (open only)")
|
||||
args = parser.parse_args()
|
||||
|
||||
if args.mode == "build":
|
||||
print(json.dumps(build(args.db_path), sort_keys=True))
|
||||
return 0
|
||||
|
||||
expected = json.loads(Path(args.expected).read_text())
|
||||
result = open_it(args.db_path, expected)
|
||||
print(json.dumps(result, indent=2, sort_keys=True))
|
||||
return 1 if result["problems"] else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,354 @@
|
||||
"""Measures what a campaign bundle preserves, omits and rebuilds.
|
||||
|
||||
python -m tools.m9_portability_report # human-readable
|
||||
python -m tools.m9_portability_report --json # machine-readable
|
||||
|
||||
Run from `backend/`, with the virtualenv on the path. The script builds the M9
|
||||
portability fixture in a throwaway database, exports it, imports it into a second
|
||||
throwaway database, and then compares the two campaigns family by family.
|
||||
|
||||
It exists because the M9 brief asks for the baseline to be **measured** rather
|
||||
than assumed. Running it on the M8 commit produces the inventory M9 started from;
|
||||
running it on the M9 tree produces the one M9 finished with, and the difference
|
||||
between the two files is the milestone's portability claim in a form a reviewer
|
||||
can reproduce rather than take on trust.
|
||||
|
||||
The comparison is by data family rather than by row count. "12 actions in, 12
|
||||
actions out" is the check that misses a bundle carrying every turn and none of
|
||||
its state, so each family below reports what a reader could still see afterwards.
|
||||
|
||||
Nothing here touches the developer's own database: two temporary files are
|
||||
created and removed, and no network call is made — the narrator, the summariser
|
||||
and the embedder are all local fakes.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
# The test harness owns the fixture and the fakes. Both live under `tests/`,
|
||||
# which is not a package, so the path is extended rather than imported from.
|
||||
_HERE = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(_HERE.parent / "tests"))
|
||||
|
||||
# `app.database` reads this at import and builds the engine once, exactly as
|
||||
# `tests/conftest.py` explains. It has to be set before the first `app` import.
|
||||
_SOURCE_DB = tempfile.NamedTemporaryFile(suffix="-m9-source.db", delete=False)
|
||||
_SOURCE_DB.close()
|
||||
os.environ["AIDND_DB_PATH"] = _SOURCE_DB.name
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends # noqa: E402
|
||||
from fastapi.testclient import TestClient # noqa: E402
|
||||
|
||||
import m9_fixture # noqa: E402
|
||||
from app import auth, limits, memorybank, models # noqa: E402
|
||||
from app.database import Base, SessionLocal, engine, get_db # noqa: E402
|
||||
from app.main import app # noqa: E402
|
||||
from app.routers import adventures # noqa: E402
|
||||
from fakes import ScriptedProvider # noqa: E402
|
||||
|
||||
|
||||
class _StubDerivedProvider:
|
||||
"""Deterministic vectors and prose, so the report needs no model at all.
|
||||
|
||||
One object serves as both the embedder and the summariser, because the
|
||||
memory pass builds each from the same factory and stubbing only one of them
|
||||
is the M6 finding M6-F3 mistake: the unstubbed factory opens a socket
|
||||
against the default endpoint on every turn.
|
||||
"""
|
||||
|
||||
_written = 0
|
||||
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
_StubDerivedProvider._written += 1
|
||||
return (
|
||||
f"Memory {_StubDerivedProvider._written}: what the story had "
|
||||
f"established by this point."
|
||||
)
|
||||
|
||||
async def embed(self, texts):
|
||||
out = []
|
||||
for text in texts:
|
||||
lowered = text.lower()
|
||||
out.append([
|
||||
1.0,
|
||||
1.0 if "abbey" in lowered or "crypt" in lowered else 0.0,
|
||||
1.0 if "tavern" in lowered or "lantern" in lowered else 0.0,
|
||||
1.0 if "rain" in lowered else 0.0,
|
||||
])
|
||||
return out
|
||||
|
||||
|
||||
def _install_fakes() -> None:
|
||||
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
||||
memorybank.embedding_provider = lambda s: _StubDerivedProvider()
|
||||
memorybank.summary_provider = lambda s: _StubDerivedProvider()
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
|
||||
|
||||
def _new_user_and_campaign(title: str) -> tuple[int, int]:
|
||||
db = SessionLocal()
|
||||
try:
|
||||
user = models.User(is_guest=False, email=f"m9-{title}@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="report-model", embedding_model="stub-embed",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title=title,
|
||||
campaign_canon=m9_fixture.CAMPAIGN_CANON,
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
|
||||
))
|
||||
# A neighbour, so a bundle that reached past its own campaign would
|
||||
# bring back rows this report can see.
|
||||
neighbour = models.Adventure(user_id=user.id, title="Neighbour")
|
||||
db.add(neighbour)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=neighbour.id, type="start", text="A different story.",
|
||||
))
|
||||
db.commit()
|
||||
return adventure.id, user.id
|
||||
finally:
|
||||
db.close()
|
||||
|
||||
|
||||
def _client(user_id: int) -> TestClient:
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
return TestClient(app)
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- the families
|
||||
# One entry per data family the M9 brief asks the baseline to classify. Each
|
||||
# `present` function answers "did this survive into the file?" from the bundle
|
||||
# alone, because that is the question the classification is about.
|
||||
|
||||
def _actions(bundle: dict) -> list[dict]:
|
||||
return [a for a in (bundle.get("actions") or []) if isinstance(a, dict)]
|
||||
|
||||
|
||||
def _snapshots(bundle: dict) -> list[dict]:
|
||||
"""Every stored prompt in the file, decoded.
|
||||
|
||||
The export compresses them (`bundle._packed`), so a report that looked for a
|
||||
plain dict would say the evidence was omitted when it is merely encoded —
|
||||
which is the mistake this whole tool exists to avoid making about anything.
|
||||
"""
|
||||
from app import bundle as bundle_module
|
||||
|
||||
out = []
|
||||
for action in _actions(bundle):
|
||||
snapshot = (
|
||||
action.get("contextSnapshot")
|
||||
if isinstance(action.get("contextSnapshot"), dict)
|
||||
else bundle_module._unpacked(action.get("contextSnapshotZ"))
|
||||
)
|
||||
if isinstance(snapshot, dict):
|
||||
out.append(snapshot)
|
||||
return out
|
||||
|
||||
|
||||
FAMILIES: list[tuple[str, str, callable]] = [
|
||||
("campaign identity",
|
||||
"title, instructions, persona, canon, the campaign's own settings",
|
||||
lambda b: bool(b.get("title"))),
|
||||
("transcript",
|
||||
"every accepted player and narrator action, live and superseded",
|
||||
lambda b: bool(_actions(b))),
|
||||
("branches",
|
||||
"the retained tree, its fork points and its names",
|
||||
lambda b: bool(b.get("branches"))),
|
||||
("branch disposition",
|
||||
"which lines the story left behind, and where",
|
||||
lambda b: any("supersededAt" in x for x in (b.get("branches") or []))),
|
||||
("active head",
|
||||
"the branch and depth the campaign is being read at",
|
||||
lambda b: b.get("headDepth") is not None),
|
||||
("alternate takes",
|
||||
"every attempt at a turn, and which one is the story",
|
||||
lambda b: any(not a.get("live", True) for a in _actions(b))),
|
||||
("take grouping",
|
||||
"which attempts belong to the same turn across a fork (SP9 parentage)",
|
||||
lambda b: any("parentId" in a for a in _actions(b))),
|
||||
("save points",
|
||||
"named coordinates, their notes and their positions",
|
||||
lambda b: bool(b.get("checkpoints"))),
|
||||
("narrative state (current)",
|
||||
"the authoritative document at the exported head",
|
||||
lambda b: b.get("narrativeState") is not None),
|
||||
("narrative state (per position)",
|
||||
"the snapshot every position restores from",
|
||||
lambda b: any("narrativeStateAfter" in a for a in _actions(b))),
|
||||
("state events",
|
||||
"the accepted typed events: the audit half of the hybrid",
|
||||
lambda b: bool(b.get("stateEvents"))),
|
||||
("state proposals",
|
||||
"what the model proposed and what the application did about it",
|
||||
lambda b: bool(b.get("stateProposals"))),
|
||||
("manual corrections",
|
||||
"state the user asserted, distinguishable from state the story did",
|
||||
lambda b: any(e.get("source") == "manual_correction"
|
||||
for e in (b.get("stateEvents") or []))),
|
||||
("historical prompt/context",
|
||||
"the exact prompt each turn was given",
|
||||
lambda b: bool(_snapshots(b))),
|
||||
("retrieval provenance",
|
||||
"which passages a historical turn was shown, and their text",
|
||||
lambda b: any((s.get("knowledge") or {}).get("used") for s in _snapshots(b))),
|
||||
("per-turn model settings",
|
||||
"the model and generation settings a historical turn ran under",
|
||||
lambda b: any(s.get("settings") for s in _snapshots(b))),
|
||||
("imported knowledge",
|
||||
"source content, class, lifecycle, visibility and hash",
|
||||
lambda b: bool(b.get("knowledge"))),
|
||||
("knowledge parser versions",
|
||||
"what produced the chunks the source last had",
|
||||
lambda b: any("parserVersion" in k for k in (b.get("knowledge") or []))),
|
||||
("summaries",
|
||||
"the generated rolling summaries and the story they cover",
|
||||
lambda b: bool(b.get("summaries"))),
|
||||
("memories",
|
||||
"long-term memories and the coordinate each hangs off",
|
||||
lambda b: bool(b.get("memories"))),
|
||||
("memory authority",
|
||||
"whether a memory is accepted story or a heuristic reading of it",
|
||||
lambda b: any("authority" in m for m in (b.get("memories") or []))),
|
||||
("scene metadata",
|
||||
"the scene section of the authoritative state document",
|
||||
lambda b: isinstance(b.get("narrativeState"), dict)
|
||||
and "scene" in b["narrativeState"]),
|
||||
("story cards (legacy)",
|
||||
"the inherited lore primitive, which has no v1 browser surface",
|
||||
lambda b: "storyCards" in b),
|
||||
]
|
||||
|
||||
#: Families that are deliberately rebuilt rather than carried, with the reason.
|
||||
REBUILDABLE = {
|
||||
"knowledge passages": "a deterministic function of the source content",
|
||||
"lexical (FTS) index": "rebuilt from the passages on import",
|
||||
"knowledge embeddings": "belong to the importing machine's embedding model",
|
||||
"memory embeddings": "the same, for the memory bank",
|
||||
"branch lineage cache": "computed from parent plus fork depth",
|
||||
"derived status": "describes the last run of a background pass, not the story",
|
||||
}
|
||||
|
||||
|
||||
def measure(json_out: bool) -> dict:
|
||||
_install_fakes()
|
||||
Base.metadata.create_all(bind=engine)
|
||||
adv_id, user_id = _new_user_and_campaign("M9 Portability Fixture")
|
||||
client = _client(user_id)
|
||||
|
||||
built = time.perf_counter()
|
||||
source = m9_fixture.build(client, adv_id)
|
||||
build_seconds = time.perf_counter() - built
|
||||
|
||||
started = time.perf_counter()
|
||||
response = client.get(f"/api/adventures/{adv_id}/export")
|
||||
export_seconds = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
bundle = response.json()
|
||||
encoded = json.dumps(bundle, ensure_ascii=False).encode("utf-8")
|
||||
|
||||
started = time.perf_counter()
|
||||
imported = client.post("/api/adventures/import", json=bundle)
|
||||
import_seconds = time.perf_counter() - started
|
||||
import_status = imported.status_code
|
||||
copy = (
|
||||
m9_fixture.snapshot_of(client, imported.json()["id"])
|
||||
if import_status == 201 else None
|
||||
)
|
||||
|
||||
source_db_bytes = Path(_SOURCE_DB.name).stat().st_size
|
||||
app.dependency_overrides.clear()
|
||||
|
||||
families = [
|
||||
{"family": name, "what": what,
|
||||
"verdict": "PRESERVED" if present(bundle) else "OMITTED"}
|
||||
for name, what, present in FAMILIES
|
||||
]
|
||||
report = {
|
||||
"format": bundle.get("format"),
|
||||
"families": families,
|
||||
"rebuildable": REBUILDABLE,
|
||||
"sizes": {
|
||||
"source_database_bytes": source_db_bytes,
|
||||
"bundle_bytes": len(encoded),
|
||||
"bundle_actions": len(_actions(bundle)),
|
||||
"bundle_keys": sorted(bundle),
|
||||
},
|
||||
"timings_seconds": {
|
||||
"fixture_build": round(build_seconds, 3),
|
||||
"export": round(export_seconds, 3),
|
||||
"import": round(import_seconds, 3),
|
||||
},
|
||||
"round_trip": {
|
||||
"import_status": import_status,
|
||||
"agrees": _agreement(source, copy) if copy else None,
|
||||
},
|
||||
}
|
||||
return report
|
||||
|
||||
|
||||
def _agreement(source: dict, copy: dict) -> dict:
|
||||
"""Which of the reader-visible families match between original and copy."""
|
||||
keys = ("title", "canon_rules", "transcript", "branch_count", "checkpoints",
|
||||
"knowledge", "state", "state_events", "memories", "summaries",
|
||||
"can_undo", "can_redo")
|
||||
return {key: source.get(key) == copy.get(key) for key in keys}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--json", action="store_true",
|
||||
help="print the report as JSON")
|
||||
args = parser.parse_args()
|
||||
try:
|
||||
report = measure(args.json)
|
||||
finally:
|
||||
for path in (_SOURCE_DB.name,):
|
||||
try:
|
||||
os.unlink(path)
|
||||
except OSError:
|
||||
pass
|
||||
if args.json:
|
||||
print(json.dumps(report, indent=2, sort_keys=True))
|
||||
return 0
|
||||
print(f"bundle format: {report['format']}")
|
||||
print(f"bundle size: {report['sizes']['bundle_bytes']:,} bytes "
|
||||
f"across {report['sizes']['bundle_actions']} actions")
|
||||
print(f"source db: {report['sizes']['source_database_bytes']:,} bytes")
|
||||
print(f"timings: {report['timings_seconds']}")
|
||||
print()
|
||||
width = max(len(name) for name, _, _ in FAMILIES)
|
||||
for row in report["families"]:
|
||||
print(f" {row['verdict']:<10} {row['family']:<{width}} {row['what']}")
|
||||
print()
|
||||
print(" DERIVED/REBUILDABLE (deliberately not carried)")
|
||||
for name, why in REBUILDABLE.items():
|
||||
print(f" {name:<24} {why}")
|
||||
print()
|
||||
print(f"round trip: HTTP {report['round_trip']['import_status']}")
|
||||
for key, agreed in (report["round_trip"]["agrees"] or {}).items():
|
||||
print(f" {'same' if agreed else 'DIFFERS':<8} {key}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,264 @@
|
||||
"""How a campaign bundle grows with the campaign, measured rather than reasoned.
|
||||
|
||||
python -m tools.m9_scale_report [--turns 120] [--budget 16384]
|
||||
|
||||
Run from `backend/`. Plays a campaign of `--turns` turns against a scripted
|
||||
narrator with the real prompt builder and a realistic context budget, then
|
||||
exports it and reports where the bytes are.
|
||||
|
||||
The question it exists to answer is the one M9's decision to carry historical
|
||||
prompts raises: **a per-turn prompt contains the story so far, so storing one per
|
||||
turn is quadratic in campaign length.** That is already true of the database —
|
||||
`compression.py` records the column as 89% of production storage — and M9 makes
|
||||
it true of the export as well. Reasoning about it gives the wrong number, because
|
||||
the prompt is bounded by the context budget rather than by the transcript: once
|
||||
the history window is full, each turn's snapshot stops growing and the total
|
||||
becomes linear again. Where that knee falls is a measurement.
|
||||
|
||||
It also watches for the accidental costs §26 names: a query per row, a
|
||||
duplicated body of knowledge content, or a snapshot written more than once per
|
||||
turn.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
_HERE = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(_HERE.parent / "tests"))
|
||||
|
||||
_DB = tempfile.NamedTemporaryFile(suffix="-m9-scale.db", delete=False)
|
||||
_DB.close()
|
||||
os.environ["AIDND_DB_PATH"] = _DB.name
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends # noqa: E402
|
||||
from fastapi.testclient import TestClient # noqa: E402
|
||||
|
||||
import m9_fixture # noqa: E402
|
||||
from app import auth, limits, memorybank, models # noqa: E402
|
||||
from app.database import Base, SessionLocal, engine, get_db # noqa: E402
|
||||
from app.main import app # noqa: E402
|
||||
from app.routers import adventures # noqa: E402
|
||||
from fakes import ScriptedProvider, state_block # noqa: E402
|
||||
|
||||
#: Prose long enough that a turn is a turn rather than a sentence. The history
|
||||
#: window is what fills the prompt, so a fixture of three-word replies would
|
||||
#: measure a campaign nobody plays.
|
||||
PROSE = (
|
||||
"The rain came harder off the fen and the lantern light shivered on the wet "
|
||||
"boards. Mara set down the cloth she had been folding and looked at him for "
|
||||
"a while without saying anything, the way she did when the answer was going "
|
||||
"to cost her something. Outside, somebody crossed the yard and did not stop."
|
||||
)
|
||||
|
||||
|
||||
class _Stub:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "The story had established a good deal by this point."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25, 0.125] for _ in texts]
|
||||
|
||||
|
||||
def _setup(budget: int) -> tuple[TestClient, int]:
|
||||
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
||||
memorybank.embedding_provider = lambda s: _Stub()
|
||||
memorybank.summary_provider = lambda s: _Stub()
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
Base.metadata.create_all(bind=engine)
|
||||
db = SessionLocal()
|
||||
try:
|
||||
user = models.User(is_guest=False, email="scale@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="scale-model", embedding_model="stub",
|
||||
context_token_budget=budget, max_output_tokens=800,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Scale", auto_summarize=True,
|
||||
memory_bank_enabled=True,
|
||||
campaign_canon=m9_fixture.CAMPAIGN_CANON,
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
|
||||
))
|
||||
db.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
finally:
|
||||
db.close()
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
return TestClient(app), adv_id
|
||||
|
||||
|
||||
def _bundle_bytes(client, adv_id) -> tuple[int, dict, float]:
|
||||
started = time.perf_counter()
|
||||
response = client.get(f"/api/adventures/{adv_id}/export")
|
||||
seconds = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
payload = response.json()
|
||||
return len(json.dumps(payload).encode("utf-8")), payload, seconds
|
||||
|
||||
|
||||
def _snapshot_bytes(payload: dict) -> int:
|
||||
"""What the stored prompts cost **in the file**, which is the encoded size.
|
||||
|
||||
Measured as they appear rather than decoded first: the question this report
|
||||
answers is how large the file gets and how close it comes to the import
|
||||
ceiling, so what counts is the bytes that actually travel.
|
||||
"""
|
||||
return sum(
|
||||
len(json.dumps(action[key]).encode("utf-8"))
|
||||
for action in payload["actions"]
|
||||
for key in ("contextSnapshotZ", "contextSnapshot")
|
||||
if action.get(key)
|
||||
)
|
||||
|
||||
|
||||
def _decoded_snapshot_bytes(payload: dict) -> int:
|
||||
"""What the same prompts would cost uncompressed, for the ratio."""
|
||||
from app import bundle as bundle_module
|
||||
|
||||
total = 0
|
||||
for action in payload["actions"]:
|
||||
snapshot = (
|
||||
action.get("contextSnapshot")
|
||||
if isinstance(action.get("contextSnapshot"), dict)
|
||||
else bundle_module._unpacked(action.get("contextSnapshotZ"))
|
||||
)
|
||||
if isinstance(snapshot, dict):
|
||||
total += len(json.dumps(snapshot).encode("utf-8"))
|
||||
return total
|
||||
|
||||
|
||||
def _section_bytes(payload: dict) -> dict:
|
||||
"""What each v3 addition costs in the file, separately.
|
||||
|
||||
Needed because "the snapshots are 21% of the file" does not answer "what did
|
||||
M9 add": the state events, the proposals and the summaries are v3 additions
|
||||
too, and a claim about M9's cost that counted only the prompts would be
|
||||
understating it.
|
||||
"""
|
||||
def size(value) -> int:
|
||||
return len(json.dumps(value).encode("utf-8"))
|
||||
|
||||
per_node = {"contextSnapshotZ": 0, "id": 0, "parentId": 0}
|
||||
for action in payload["actions"]:
|
||||
for key in per_node:
|
||||
if key in action:
|
||||
per_node[key] += size(action[key]) + len(key) + 4
|
||||
return {
|
||||
"prompts (contextSnapshotZ)": per_node["contextSnapshotZ"],
|
||||
"state events": size(payload.get("stateEvents") or []),
|
||||
"state proposals": size(payload.get("stateProposals") or []),
|
||||
"summaries": size(payload.get("summaries") or []),
|
||||
"node ids + parentage": per_node["id"] + per_node["parentId"],
|
||||
"per-position state (v2 already)": sum(
|
||||
size(a["narrativeStateAfter"]) for a in payload["actions"]
|
||||
if "narrativeStateAfter" in a
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--turns", type=int, default=120)
|
||||
parser.add_argument("--budget", type=int, default=16384)
|
||||
parser.add_argument("--every", type=int, default=20,
|
||||
help="report the running size every N turns")
|
||||
args = parser.parse_args()
|
||||
|
||||
client, adv_id = _setup(args.budget)
|
||||
for name, body, kind in (
|
||||
("canon.md", m9_fixture.CANON_MD, "canon"),
|
||||
("reference.md", m9_fixture.REFERENCE_MD, "reference"),
|
||||
("secret.md", m9_fixture.SECRET_MD, "canon"),
|
||||
):
|
||||
m9_fixture.upload(client, adv_id, name, body, kind)
|
||||
|
||||
print(f"budget {args.budget} tokens, {args.turns} turns\n")
|
||||
print(f"{'turns':>6} {'actions':>8} {'bundle B':>12} {'snapshots B':>13} "
|
||||
f"{'B/turn':>9} {'export s':>9} {'import s':>9}")
|
||||
rows = []
|
||||
sections: dict[int, dict] = {}
|
||||
for turn in range(1, args.turns + 1):
|
||||
ScriptedProvider.replies = [
|
||||
f"{PROSE} [{turn}]\n"
|
||||
+ state_block([{"type": "add_fact", "predicate": "tally",
|
||||
"value": turn * 10, "fact_id": f"tally-{turn}"}])
|
||||
]
|
||||
response = client.post(f"/api/adventures/{adv_id}/actions",
|
||||
json={"type": "do", "text": f"press on, {turn}"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
if turn % args.every and turn != args.turns:
|
||||
continue
|
||||
m9_fixture.settle_derived(adv_id)
|
||||
size, payload, export_seconds = _bundle_bytes(client, adv_id)
|
||||
started = time.perf_counter()
|
||||
imported = client.post("/api/adventures/import", json=payload)
|
||||
import_seconds = time.perf_counter() - started
|
||||
assert imported.status_code == 201, imported.text[:300]
|
||||
client.delete(f"/api/adventures/{imported.json()['id']}")
|
||||
snapshots = _snapshot_bytes(payload)
|
||||
plain = _decoded_snapshot_bytes(payload)
|
||||
rows.append((turn, size, snapshots, plain))
|
||||
sections[turn] = _section_bytes(payload)
|
||||
print(f"{turn:>6} {len(payload['actions']):>8} {size:>12,} "
|
||||
f"{snapshots:>13,} {snapshots // turn:>9,} "
|
||||
f"{export_seconds:>9.3f} {import_seconds:>9.3f}")
|
||||
|
||||
db_bytes = Path(_DB.name).stat().st_size
|
||||
last_turn, last_size, last_snapshots, last_plain = rows[-1]
|
||||
per_turn = last_snapshots // last_turn
|
||||
cap = limits.MAX_IMPORT_BODY_BYTES
|
||||
print()
|
||||
print(f"database on disk: {db_bytes:,} bytes")
|
||||
print(f"snapshot share of file: {100 * last_snapshots // last_size}%")
|
||||
print(f"stored uncompressed: {last_plain:,} bytes "
|
||||
f"({last_plain / max(last_snapshots, 1):.1f}x the encoded size)")
|
||||
print(f"import body cap: {cap:,} bytes")
|
||||
print(f"turns before the cap: ~{cap // max(per_turn, 1):,} "
|
||||
f"at the marginal rate above")
|
||||
# Growth between the last two samples says whether the per-turn cost has
|
||||
# settled. It should: once the history window fills the budget, a prompt
|
||||
# stops growing with the transcript and the total becomes linear.
|
||||
if len(rows) >= 2:
|
||||
(t0, _, s0, _p0), (t1, _, s1, _p1) = rows[-2], rows[-1]
|
||||
print(f"marginal cost, last {t1 - t0} turns: "
|
||||
f"{(s1 - s0) // max(t1 - t0, 1):,} bytes/turn")
|
||||
print()
|
||||
print("where the bytes are, at the last sample:")
|
||||
last = sections[last_turn]
|
||||
added = sum(v for k, v in last.items() if not k.endswith("(v2 already)"))
|
||||
for name, value in sorted(last.items(), key=lambda kv: -kv[1]):
|
||||
print(f" {name:34} {value:>12,} {100 * value / last_size:5.1f}%")
|
||||
print(f" {'--- everything v3 added':34} {added:>12,} "
|
||||
f"{100 * added / last_size:5.1f}%")
|
||||
without = last_size - added
|
||||
print(f" a v2 file of the same campaign {without:>12,}")
|
||||
print(f" ceiling with v3 additions: ~{int((cap / (last_size / last_turn**2)) ** 0.5):,} turns")
|
||||
print(f" ceiling without them: ~{int((cap / (without / last_turn**2)) ** 0.5):,} turns")
|
||||
app.dependency_overrides.clear()
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
try:
|
||||
raise SystemExit(main())
|
||||
finally:
|
||||
try:
|
||||
os.unlink(_DB.name)
|
||||
except OSError:
|
||||
pass
|
||||
Reference in New Issue
Block a user