Files
interactive-story/backend/tests/test_m10_bundle.py
JesseMarkowitzandClaude Opus 5 144406cd48 M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was
whether any of the earlier evidence meant what it said. M8 measured a deployment
enforcing a 4,096-token input window while the application budgeted 16,384.
Every request returned 200. What Ollama does with the excess is drop the oldest
tokens, and the oldest tokens here are the system block — the narrator's rules
and the campaign canon. A hundred-turn certification against that server would
have looked perfect and proved nothing, which is why this milestone could not
begin with a hundred turns.

So the application asks now. Ollama's window is a property of how a model was
loaded rather than of the request — sending num_ctx is accepted, ignored, and
worse, reloads the model at the server's own default — so the only honest move
is to find out and then tell the truth about it. /api/ps reports what a resident
model is being served with, /api/show what an unloaded one will load with, both
on the same host inference already uses, through the same endpoint policy and
the same TLS trust store. A verified window is a ceiling on the budget; an
unverified one leaves the budget alone and is recorded as unverified in the
turn's own provenance, so an old turn can be asked afterwards whether it was
built against a checked window. There is no third behaviour, and in particular
no hard-coded 4,096: a number the server did not say would be right on one
machine and wrong on the next.

The proof that this is doing something is a campaign whose canon sits at the
front of the prompt, 120 turns of history, and a 4,096-token window. The canon
is still there afterwards and the oldest history is gone. The same campaign
built the old way produces a prompt more than twice the window — the defect,
reproduced, so the fix is measured against it rather than asserted.

Two defects the validation found on its own, and they are the same defect twice:
something was true and nobody was told. A manual state correction of four
changes with one bad reference applied three, returned 201, and said nothing —
while recording the refusal on the audit row nobody reads. It came to light
because the identity diagnostic's own fixture was refused that way and the whole
run proceeded on a campaign with no scene, which would have read as a model
failure. And the narration-length setting moved no number: brief, medium and
long each became one English sentence, while the numeric hint the model actually
reads was derived from the global reply cap and said the same thing for all
three. Both now say what they did.

The other two post-M8 findings are closed as well. The tab said AI D&D, which no
document had ever claimed it did not; it says Interactive Story now, with the
open campaign first, and the name is the owner's decision rather than a
find-and-replace to something narrower than the engine. After an Undo the reader
could not tell where they had landed; the control row now ends with
"Moment 11 · later story ahead", from the server's own answer, in the word the
transcript already uses, with none of head, branch or depth anywhere near it.

The identity diagnostic exists and the root cause does not. That campaign was
destroyed, so no cause can be established — what M11 owes the finding is
something that can classify the next occurrence, and a diagnostic that makes only
the judgements a program can honestly make: duplicate keys, shared names,
protagonist drift, state and context disagreeing. Whether prose misattributed a
line is left to a person reading it beside its prompt, because a regex cannot
read dialogue and one that pretended to would produce exactly the confident wrong
answer this finding is about. Its detectors are proved to fire against a planted
second Alice.

Two entities may still share a display name. That was checked first, as the
finding asked, and left permitted: a mother and a daughter, or a stranger giving
a false name, are ordinary fiction, and refusing them to guard against a model
mistake would refuse the wrong thing. What was missing was that it happened
silently. It is reported now.

Evidence, not inference: a hundred accepted turns against a real narrator with
genuine process restarts; a real browser against the built SPA; a container with
no network at all; a campaign moved into a data directory that never existed.
Each was discarded and re-run whenever the product changed under it, and the runs
that were thrown away are listed in the report with the reason, along with ten
defects in the harnesses themselves — because a harness that has only ever
agreed with itself is not evidence, and two of M8's five harness defects were
masking real ones.

No dependency was added, removed or upgraded. No acceptance test was retired,
relaxed or reclassified. M11 is implemented and verified; it is not accepted, and
there is no release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 14:01:20 -04:00

501 lines
21 KiB
Python

"""M10 §14 and §15: the profiles travel, and an M9 database opens.
Two questions, and they are the ones a reader would ask if they knew what M10
had done to their machine:
* **§14 — does a campaign still move?** A visual profile is part of the campaign
the reader built, so it belongs in the bundle. It is also *new*, which is the
risk: an exporter that carries it and an importer that drops it both pass a
test that only checks the campaign still opens.
* **§15 — does the database I already have still work?** M10 adds one table and
nothing else. An existing campaign must survive opening under the new build
untouched, opening must not care how many times it happens, the schema an M9
file reaches must be the schema a fresh install has, and M9's backup must keep
working on the result.
The upgrade needs **no migration**: `create_all` builds a new table and the
indexes declared on its columns on every path. A `CREATE INDEX` migration was
written here first and `test_a_fresh_database_arrives_at_the_same_place` is what
found it wrong — it left an upgraded database holding an index a fresh install
did not have. That test is the one to keep pointed at any future schema change.
The bundle format stays `ai-dnd-adventure-v3`. M9's own test for a version bump
is whether omission creates ambiguity about what an older file *could* have
recorded, and it does not: a campaign with no visual profiles is the ordinary
case, so an absent key means "none" rather than "unknown". The tests below hold
that decision to its consequence — an M9-written v3 file must still import, and
the M10 exporter must still produce a file an M9 build would recognise.
python -m pytest tests/test_m10_bundle.py -v
"""
import copy
import os
import shutil
import sqlite3
import tempfile
from pathlib import Path
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy import create_engine, text
from sqlalchemy.orm import sessionmaker
from app import auth, backup, limits, migrations, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import adventures
import m10_fixture
from fakes import ScriptedProvider
from test_process_restart import Server, _free_port
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="m10bundle@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(user_id=user.id, model="test-model",
embedding_model=""))
adventure = models.Adventure(user_id=user.id, title="Portable office")
setup.add(adventure)
setup.flush()
setup.add(models.Action(adventure_id=adventure.id, type="start",
text="Bill badges in."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
def export(client, adv_id=None) -> dict:
response = client.get(f"/api/adventures/{adv_id or client.adv_id}/export")
assert response.status_code == 200, response.text[:400]
return response.json()
def bring_back(client, payload) -> int:
response = client.post("/api/adventures/import", json=payload)
assert response.status_code == 201, response.text[:600]
return response.json()["id"]
def profiles_of(client, adv_id) -> dict:
body = client.get(f"/api/adventures/{adv_id}/visual-profiles").json()
return {p["entity_key"]: p for p in body["profiles"]}
@pytest.fixture()
def moved(client):
"""The office campaign, its bundle, and the copy the bundle produced."""
m10_fixture.build(client, client.adv_id)
payload = export(client)
return {"bundle": payload, "copy_id": bring_back(client, payload)}
# --------------------------------------------------------------- §14 the file
def test_the_format_version_is_unchanged(moved):
"""The decision, recorded as a test so a later bump is deliberate."""
assert moved["bundle"]["format"] == "ai-dnd-adventure-v3"
def test_the_bundle_carries_the_profiles_that_exist(moved):
exported = {p["entityKey"]: p for p in moved["bundle"]["visualProfiles"]}
assert set(exported) == {"alice", "office"}
assert exported["alice"]["descriptors"]["hair"] == "short black"
assert exported["alice"]["features"] == ["tortoiseshell glasses"]
assert exported["alice"]["styleNotes"] == "photographic, natural light"
def test_an_unprofiled_character_exports_no_empty_profile(moved):
"""Roger has no profile, and the file must say that by omission.
An exporter that wrote a blank row for every entity would lose the
distinction a provider needs: "nobody decided what Roger looks like" is not
"Roger looks like nothing".
"""
keys = [p["entityKey"] for p in moved["bundle"]["visualProfiles"]]
assert "roger" not in keys and "bill" not in keys
def test_the_copy_holds_the_same_profiles(client, moved):
original = profiles_of(client, client.adv_id)
copied = profiles_of(client, moved["copy_id"])
assert set(copied) == set(original)
for key in original:
assert copied[key]["descriptors"] == original[key]["descriptors"]
assert copied[key]["features"] == original[key]["features"]
assert copied[key]["style_notes"] == original[key]["style_notes"]
def test_the_copys_profiles_are_its_own_rows(client, moved):
"""Editing the copy must not reach back into the original."""
client.put(f"/api/adventures/{moved['copy_id']}/visual-profiles/alice",
json={"descriptors": {"hair": "bleached"}})
assert profiles_of(client, client.adv_id)["alice"][
"descriptors"]["hair"] == "short black"
def test_the_copys_scene_packet_is_populated_from_the_imported_profiles(
client, moved):
"""The point of carrying them: the copy can be depicted without redoing work."""
packet = client.get(
f"/api/adventures/{moved['copy_id']}/scene-packet").json()
by_name = {c["name"]: c for c in packet["characters"]}
assert by_name["Alice"]["visual_profile"]["descriptors"]["build"] == "tall"
assert by_name["Roger"]["visual_profile"] is None
assert packet["location"]["visual_profile"]["descriptors"][
"lighting"] == "flat fluorescent"
def test_an_m9_era_file_still_imports_and_simply_has_no_profiles(client, moved):
"""A v3 file written before M10 existed: the key is absent, not empty."""
older = copy.deepcopy(moved["bundle"])
del older["visualProfiles"]
copy_id = bring_back(client, older)
assert profiles_of(client, copy_id) == {}
# And the campaign itself arrived intact.
assert client.get(f"/api/adventures/{copy_id}/scene-packet").json()[
"characters"]
def test_a_malformed_profile_is_dropped_rather_than_refusing_the_campaign(
client, moved):
"""§14's proportionality rule, in the one place M10 could get it wrong.
A story that will not import because a description of somebody's coat is
malformed would be the wrong trade. The campaign arrives; the bad profile
does not; the good one does.
"""
damaged = copy.deepcopy(moved["bundle"])
damaged["visualProfiles"].append(
{"entity_key": "", "descriptors": "not an object"})
damaged["visualProfiles"].append({"descriptors": {"a": "b"}})
copy_id = bring_back(client, damaged)
assert set(profiles_of(client, copy_id)) == {"alice", "office"}
def test_a_profile_survives_a_second_round_trip_unchanged(client, moved):
"""Export, import, export again: the file is a fixed point."""
again = export(client, moved["copy_id"])
first = sorted(moved["bundle"]["visualProfiles"], key=lambda p: p["entityKey"])
second = sorted(again["visualProfiles"], key=lambda p: p["entityKey"])
assert [p["entityKey"] for p in first] == [p["entityKey"] for p in second]
for a, b in zip(first, second):
assert a["descriptors"] == b["descriptors"]
assert a["features"] == b["features"]
assert a["styleNotes"] == b["styleNotes"]
def test_a_neighbouring_campaigns_profiles_do_not_travel(client, moved):
"""Scoping: the exporter must filter by campaign, not by table."""
with SessionLocal() as db:
neighbour = models.Adventure(user_id=None, title="Someone else's")
db.add(neighbour)
db.flush()
db.add(models.VisualProfile(
adventure_id=neighbour.id, entity_key="intruder",
descriptors={"hair": "should not travel"}, features=[],
style_notes=""))
db.commit()
keys = [p["entityKey"] for p in export(client)["visualProfiles"]]
assert "intruder" not in keys
def test_the_planner_checks_the_profiles_before_a_row_is_written(moved):
"""M9's atomicity rule: everything is checked before anything is written.
`bundle.plan` is that checkpoint — it has no side effects and is what the
importer runs first — so a profile that would fail must fail there rather
than halfway through writing a campaign. There is no HTTP preview endpoint;
the planner is called directly for the same reason the importer calls it.
"""
from app import bundle as bundle_module
planned = bundle_module.plan(moved["bundle"], "ai-dnd-adventure-v3")
assert {p["entity_key"] for p in planned["visualProfiles"]} == {
"alice", "office"}
# ---------------------------------------------------------- §15 the migration
@pytest.fixture()
def m9_database():
"""A database as an M9 build left it, with a campaign already in it.
M10's only schema change is the `visual_profiles` table, so an M9-era file
is exactly this: the current schema without that table, stamped at 92 — the
version M9 ended on, and the version an M10-era file still carries, because
M10 added no migration of its own. Opening it brings it to whatever the
current version is; M11 later added 93, which is why these tests compare
against `LATEST_VERSION` rather than a literal. The campaign rows are
written before the upgrade, because the claim under test is that they are
still there afterwards.
"""
directory = tempfile.mkdtemp(prefix="m10-migrate-")
path = Path(directory) / "campaign.db"
older = create_engine(f"sqlite:///{path}")
Base.metadata.create_all(bind=older)
# Written through the ORM, so the campaign in the file is shaped the way the
# application writes one rather than the way a test guessed at.
with sessionmaker(bind=older)() as db:
adventure = models.Adventure(title="An M9 campaign")
db.add(adventure)
db.flush()
db.add(models.Action(adventure_id=adventure.id, type="start",
text="The story opened before M10."))
db.commit()
adv_id = adventure.id
with older.begin() as conn:
conn.execute(text("DROP TABLE visual_profiles"))
conn.execute(text("PRAGMA user_version = 92"))
older.dispose()
yield path, create_engine(f"sqlite:///{path}"), adv_id
def _indexes(engine_) -> set:
with engine_.begin() as conn:
return {row[0] for row in conn.execute(text(
"SELECT name FROM sqlite_master WHERE type = 'index'"))}
def _version(engine_) -> int:
with engine_.begin() as conn:
return conn.execute(text("PRAGMA user_version")).scalar()
def test_an_m9_database_gains_the_new_table_when_it_is_opened(m9_database):
path, older, adv_id = m9_database
assert _version(older) == 92
migrations.bootstrap(older)
assert _version(older) == migrations.LATEST_VERSION
with older.begin() as conn:
assert conn.execute(text("SELECT COUNT(*) FROM visual_profiles")).scalar() == 0
assert "ix_visual_profiles_adventure_id" in _indexes(older)
def test_no_migration_mentions_the_table_m10_added(m9_database):
"""M10's actual claim, stated so a later migration cannot invalidate it.
The first version of this file expressed "M10 adds no migration" as
`LATEST_VERSION == 92`, which stopped being true the moment M11 added a
column to another table — a fact about M11 that says nothing about M10. The
durable claim is that `visual_profiles` arrives through `create_all` and
that no migration anywhere touches it.
"""
for _, sql in migrations.MIGRATIONS:
body = sql if isinstance(sql, str) else " ".join(sql.values())
assert "visual_profiles" not in body, body[:120]
def test_the_campaign_that_was_already_there_is_untouched(m9_database):
path, older, adv_id = m9_database
migrations.bootstrap(older)
with older.begin() as conn:
assert conn.execute(text("SELECT title FROM adventures")).scalar() == (
"An M9 campaign")
assert conn.execute(text("SELECT text FROM actions")).scalar() == (
"The story opened before M10.")
assert conn.execute(text("PRAGMA foreign_key_check")).fetchall() == []
def test_opening_the_database_repeatedly_is_a_no_op(m9_database):
"""Three starts in a row. Nothing accumulates and nothing errors.
This is the idempotence §15 asks about. It is stated as "open it again"
rather than "run the migration again" because opening is what the
application does, and M10 has no migration of its own to rerun.
"""
path, older, adv_id = m9_database
migrations.bootstrap(older)
after_first = _indexes(older)
first_version = _version(older)
for _ in range(2):
migrations.bootstrap(older)
assert _version(older) == first_version == migrations.LATEST_VERSION
assert _indexes(older) == after_first
with older.begin() as conn:
assert conn.execute(text("SELECT COUNT(*) FROM adventures")).scalar() == 1
def test_a_fresh_database_arrives_at_the_same_place(m9_database):
"""An upgraded M9 file and a new install must not differ.
Two schemas that disagree is the failure this catches, and it is the one a
version stamp alone would hide.
"""
path, older, adv_id = m9_database
migrations.bootstrap(older)
fresh_path = path.with_name("fresh.db")
fresh = create_engine(f"sqlite:///{fresh_path}")
migrations.bootstrap(fresh)
assert _version(fresh) == _version(older)
def shape(e):
with e.begin() as conn:
return conn.execute(text(
"SELECT sql FROM sqlite_master WHERE name = 'visual_profiles'"
)).scalar()
assert shape(fresh) == shape(older)
# Including the indexes. This comparison is what caught the redundant
# `CREATE INDEX` migration M10 first shipped: the upgraded file had an index
# the fresh one did not, which is a difference no test of either database on
# its own would have shown.
assert _indexes(fresh) == _indexes(older)
fresh.dispose()
def test_a_backup_of_the_upgraded_database_still_works(m9_database):
"""M9's backup keeps its guarantees on a file M10 added a table to."""
path, older, adv_id = m9_database
migrations.bootstrap(older)
older.dispose()
result = backup.create(path)
try:
assert result.integrity == "ok"
assert result.pages > 0
with sqlite3.connect(f"file:{result.path}?mode=ro", uri=True) as copy_db:
assert copy_db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
assert copy_db.execute("PRAGMA foreign_key_check").fetchall() == []
# It opens independently: the new table is in it, and so is the
# campaign that predates the migration.
assert copy_db.execute(
"SELECT COUNT(*) FROM visual_profiles").fetchone()[0] == 0
assert copy_db.execute(
"SELECT title FROM adventures").fetchone()[0] == "An M9 campaign"
assert copy_db.execute("PRAGMA user_version").fetchone()[0] == (
migrations.LATEST_VERSION)
finally:
result.path.unlink(missing_ok=True)
def test_a_backup_carries_the_profiles_written_after_the_upgrade(m9_database):
path, older, adv_id = m9_database
migrations.bootstrap(older)
with older.begin() as conn:
conn.execute(text(
"INSERT INTO visual_profiles "
"(adventure_id, entity_key, descriptors, features, style_notes, "
" created_at, updated_at) "
"VALUES (:adv, 'bill', '{\"build\": \"heavyset\"}', '[]', '', "
" datetime('now'), datetime('now'))"), {"adv": adv_id})
older.dispose()
result = backup.create(path)
try:
with sqlite3.connect(f"file:{result.path}?mode=ro", uri=True) as copy_db:
row = copy_db.execute(
"SELECT entity_key, descriptors FROM visual_profiles").fetchone()
assert row[0] == "bill" and "heavyset" in row[1]
finally:
result.path.unlink(missing_ok=True)
# ------------------------------------- §14 the move to a machine that never saw it
@pytest.fixture()
def machines():
"""Two directories, each with its own database, and a server on each.
The same shape as `test_m9_clean_import.py`, for the same reason: a shared
id space, a warm cache or a session still holding the original would let an
in-process import pass while a real move failed. M9's version of this test
predates visual profiles and carries none, so this is the profile-carrying
half of the same claim rather than a duplicate of it.
"""
root = tempfile.mkdtemp(prefix="m10-clean-")
started: list[Server] = []
def start(name: str) -> Server:
directory = os.path.join(root, name)
os.makedirs(directory, exist_ok=True)
server = Server(os.path.join(directory, "campaign.db"), _free_port())
started.append(server)
server.wait_until_ready()
return server
try:
yield start
finally:
for server in started:
server.stop()
shutil.rmtree(root, ignore_errors=True)
def test_profiles_reach_a_clean_data_directory_on_another_machine(machines):
"""§14's Definition-of-Done clause, run across two real processes.
Machine A plays a campaign, profiles two entities and exports. Machine B is
a database file that has never existed before, in a different directory, in
a different process — migrations run there from nothing. Nothing crosses but
the bundle.
"""
a = machines("machine-a")
campaign = a.call("POST", "/adventures",
{"title": "Moving day", "opening": "The office is quiet."},
expect=201)
adv = campaign["id"]
a.call("POST", f"/adventures/{adv}/state/corrections", {
"events": [
{"type": "create_entity", "entity": "alice",
"entity_type": "character", "name": "Alice"},
{"type": "create_entity", "entity": "roger",
"entity_type": "character", "name": "Roger"},
{"type": "create_entity", "entity": "office",
"entity_type": "location", "name": "The office"},
{"type": "set_scene", "summary": "Alice and Roger wait in the office.",
"location": "office", "present": ["alice", "roger"]},
],
"note": "setting the scene",
}, expect=201)
a.call("PUT", f"/adventures/{adv}/visual-profiles/alice",
{"descriptors": {"build": "tall", "hair": "short black"},
"features": ["tortoiseshell glasses"],
"style_notes": "photographic, natural light"}, expect=200)
a.call("PUT", f"/adventures/{adv}/visual-profiles/office",
{"descriptors": {"lighting": "flat fluorescent"}}, expect=200)
payload = a.call("GET", f"/adventures/{adv}/export", expect=200)
source_packet = a.call("GET", f"/adventures/{adv}/scene-packet", expect=200)
a.stop()
assert not a.is_listening()
b = machines("machine-b")
moved = b.call("POST", "/adventures/import", payload, expect=201)["id"]
profiles = {p["entity_key"]: p for p in b.call(
"GET", f"/adventures/{moved}/visual-profiles", expect=200)["profiles"]}
assert set(profiles) == {"alice", "office"}
assert profiles["alice"]["features"] == ["tortoiseshell glasses"]
assert profiles["alice"]["style_notes"] == "photographic, natural light"
# The packet the copy builds describes the same scene, with the same
# profiles attached and Roger still deliberately unprofiled. Only the
# campaign id differs, which is what a new machine's id space means.
moved_packet = b.call("GET", f"/adventures/{moved}/scene-packet", expect=200)
assert moved_packet["action_summary"] == source_packet["action_summary"]
by_name = {c["name"]: c for c in moved_packet["characters"]}
assert by_name["Alice"]["visual_profile"]["descriptors"]["hair"] == "short black"
assert by_name["Roger"]["visual_profile"] is None
assert moved_packet["location"]["visual_profile"]["descriptors"][
"lighting"] == "flat fluorescent"