The release-validation milestone, and the thing it had to settle first was whether any of the earlier evidence meant what it said. M8 measured a deployment enforcing a 4,096-token input window while the application budgeted 16,384. Every request returned 200. What Ollama does with the excess is drop the oldest tokens, and the oldest tokens here are the system block — the narrator's rules and the campaign canon. A hundred-turn certification against that server would have looked perfect and proved nothing, which is why this milestone could not begin with a hundred turns. So the application asks now. Ollama's window is a property of how a model was loaded rather than of the request — sending num_ctx is accepted, ignored, and worse, reloads the model at the server's own default — so the only honest move is to find out and then tell the truth about it. /api/ps reports what a resident model is being served with, /api/show what an unloaded one will load with, both on the same host inference already uses, through the same endpoint policy and the same TLS trust store. A verified window is a ceiling on the budget; an unverified one leaves the budget alone and is recorded as unverified in the turn's own provenance, so an old turn can be asked afterwards whether it was built against a checked window. There is no third behaviour, and in particular no hard-coded 4,096: a number the server did not say would be right on one machine and wrong on the next. The proof that this is doing something is a campaign whose canon sits at the front of the prompt, 120 turns of history, and a 4,096-token window. The canon is still there afterwards and the oldest history is gone. The same campaign built the old way produces a prompt more than twice the window — the defect, reproduced, so the fix is measured against it rather than asserted. Two defects the validation found on its own, and they are the same defect twice: something was true and nobody was told. A manual state correction of four changes with one bad reference applied three, returned 201, and said nothing — while recording the refusal on the audit row nobody reads. It came to light because the identity diagnostic's own fixture was refused that way and the whole run proceeded on a campaign with no scene, which would have read as a model failure. And the narration-length setting moved no number: brief, medium and long each became one English sentence, while the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three. Both now say what they did. The other two post-M8 findings are closed as well. The tab said AI D&D, which no document had ever claimed it did not; it says Interactive Story now, with the open campaign first, and the name is the owner's decision rather than a find-and-replace to something narrower than the engine. After an Undo the reader could not tell where they had landed; the control row now ends with "Moment 11 · later story ahead", from the server's own answer, in the word the transcript already uses, with none of head, branch or depth anywhere near it. The identity diagnostic exists and the root cause does not. That campaign was destroyed, so no cause can be established — what M11 owes the finding is something that can classify the next occurrence, and a diagnostic that makes only the judgements a program can honestly make: duplicate keys, shared names, protagonist drift, state and context disagreeing. Whether prose misattributed a line is left to a person reading it beside its prompt, because a regex cannot read dialogue and one that pretended to would produce exactly the confident wrong answer this finding is about. Its detectors are proved to fire against a planted second Alice. Two entities may still share a display name. That was checked first, as the finding asked, and left permitted: a mother and a daughter, or a stranger giving a false name, are ordinary fiction, and refusing them to guard against a model mistake would refuse the wrong thing. What was missing was that it happened silently. It is reported now. Evidence, not inference: a hundred accepted turns against a real narrator with genuine process restarts; a real browser against the built SPA; a container with no network at all; a campaign moved into a data directory that never existed. Each was discarded and re-run whenever the product changed under it, and the runs that were thrown away are listed in the report with the reason, along with ten defects in the harnesses themselves — because a harness that has only ever agreed with itself is not evidence, and two of M8's five harness defects were masking real ones. No dependency was added, removed or upgraded. No acceptance test was retired, relaxed or reclassified. M11 is implemented and verified; it is not accepted, and there is no release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
418 lines
17 KiB
Python
418 lines
17 KiB
Python
"""M11 §13: E01-E04, all four at once, in one long campaign.
|
|
|
|
The E-series already has tests, and good ones — M6's corrective pass rewrote E03
|
|
after an independent review found the first version passing while the defect was
|
|
live. What none of them does is what the M11 brief asks for: exercise all four
|
|
**together, in a single campaign, under realistic long-story conditions**, with
|
|
state, memories, summaries, imported knowledge and a scene all live at once.
|
|
|
|
That matters because the four leaks share one mechanism — a head that moves and
|
|
a lineage that decides what is still true — and a campaign that has only one of
|
|
them cannot show the mechanism failing for one and holding for another. It also
|
|
adds the dimension none of the earlier tests could have: M10's Scene Packet, the
|
|
thing a future depiction would be built from, which has to answer for the active
|
|
line exactly as the state does.
|
|
|
|
Four sentinels, one per class, each with a positive control on path A and a
|
|
negative control on path B:
|
|
|
|
state a fact and a location established on the abandoned line
|
|
memory a distinctive memory extracted from abandoned turns
|
|
summary a summary **regenerated after the divergence** (M6-F1's shape)
|
|
scene the location the abandoned line moved to, in the state *and* in
|
|
the derived packet
|
|
|
|
python -m pytest tests/test_m11_leakage.py -v
|
|
"""
|
|
|
|
import asyncio
|
|
|
|
import pytest
|
|
from fastapi import Depends
|
|
from fastapi.testclient import TestClient
|
|
|
|
from app import auth, limits, memorybank, models, summaries
|
|
from app.context import lineage
|
|
from app.database import Base, SessionLocal, engine, get_db
|
|
from app.knowledge import embeddings
|
|
from app.main import app
|
|
from app.routers import adventures
|
|
|
|
from fakes import ScriptedProvider, state_block
|
|
|
|
#: One per leak class, so a failure names which boundary broke.
|
|
STATE_SENTINEL = "the-abbey-seal-was-broken"
|
|
MEMORY_SENTINEL = "GRIMWALD-CONFESSED-8821"
|
|
SUMMARY_SENTINEL = "ABANDONED-OATH-SWORN-4416"
|
|
SCENE_SENTINEL = "old_abbey_crypt"
|
|
|
|
|
|
class Summariser:
|
|
"""Carries the summary forward and folds in new events, as a real one does.
|
|
|
|
Copied in behaviour from `test_context_memory.CarryingSummariser` — the M6
|
|
corrective pass established that a summariser which *discards* its seed
|
|
cannot show the E03 defect, because the defect is in what the seed contains.
|
|
"""
|
|
|
|
def __init__(self):
|
|
self.seeds: list[str] = []
|
|
|
|
async def complete(self, system, user, *, max_tokens=600):
|
|
if "Current story summary:" not in user:
|
|
found = [s for s in (MEMORY_SENTINEL, SUMMARY_SENTINEL) if s in user]
|
|
if found:
|
|
return "MEM[" + " ".join(found) + "]"
|
|
return "MEM[the road, and nothing sworn]"
|
|
current = user.split("Current story summary:\n", 1)[1].split("\n\nNew events")[0]
|
|
events = user.split("New events since the last update:\n", 1)[1].split(
|
|
"\n\nUpdated summary:")[0]
|
|
self.seeds.append(current.strip())
|
|
carried = "" if current.strip() == "(none yet)" else current.strip() + " "
|
|
return (carried + events.strip().replace("\n", " "))[:2000]
|
|
|
|
async def embed(self, texts):
|
|
out = []
|
|
for text in texts:
|
|
out.append([
|
|
1.0,
|
|
1.0 if MEMORY_SENTINEL in text or "confess" in text.lower() else 0.0,
|
|
1.0 if "road" in text.lower() else 0.0,
|
|
])
|
|
return out
|
|
|
|
|
|
@pytest.fixture()
|
|
def summariser(monkeypatch):
|
|
made = Summariser()
|
|
monkeypatch.setattr(memorybank, "summary_provider", lambda s: made)
|
|
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: made)
|
|
return made
|
|
|
|
|
|
@pytest.fixture()
|
|
def client(monkeypatch, summariser):
|
|
Base.metadata.create_all(bind=engine)
|
|
memorybank._vector_cache.clear()
|
|
embeddings._cache.clear()
|
|
setup = SessionLocal()
|
|
user = models.User(is_guest=False, email="m11leak@example.com")
|
|
setup.add(user)
|
|
setup.flush()
|
|
setup.add(models.Settings(
|
|
user_id=user.id, model="test-model", embedding_model="embed-test",
|
|
context_token_budget=6000, max_output_tokens=400, memory_top_k=4,
|
|
))
|
|
adventure = models.Adventure(
|
|
user_id=user.id, title="Continuity", memory_bank_enabled=True,
|
|
auto_summarize=True,
|
|
campaign_canon={"rules": ["The dead do not return."]},
|
|
)
|
|
setup.add(adventure)
|
|
setup.flush()
|
|
setup.add(models.Action(adventure_id=adventure.id, type="start",
|
|
text="Rain over Westhaven, and the abbey bell tolling."))
|
|
setup.commit()
|
|
adv_id, user_id = adventure.id, user.id
|
|
setup.close()
|
|
|
|
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
|
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
|
app.dependency_overrides[auth.get_current_user] = (
|
|
lambda db=Depends(get_db): db.get(models.User, user_id)
|
|
)
|
|
test_client = TestClient(app)
|
|
test_client.adv_id = adv_id
|
|
try:
|
|
yield test_client
|
|
finally:
|
|
app.dependency_overrides.clear()
|
|
adventures.turns._active_turns.clear()
|
|
memorybank._vector_cache.clear()
|
|
embeddings._cache.clear()
|
|
Base.metadata.drop_all(bind=engine)
|
|
|
|
|
|
# ----------------------------------------------------------------- helpers
|
|
|
|
def play(client, text, prose="The road bends on past the treeline.", events=None):
|
|
ScriptedProvider.replies = [f"{prose}\n{state_block(events or [])}"]
|
|
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
|
json={"type": "do", "text": text})
|
|
assert response.status_code == 200, response.text[:300]
|
|
assert '"error"' not in response.text, response.text[:300]
|
|
|
|
|
|
def report(client) -> dict:
|
|
response = client.get(f"/api/adventures/{client.adv_id}/context")
|
|
assert response.status_code == 200, response.text[:300]
|
|
return response.json()
|
|
|
|
|
|
def prompt_of(report_: dict) -> str:
|
|
return "\n".join(section["text"] for section in report_["sections"])
|
|
|
|
|
|
def state_of(client) -> dict:
|
|
return client.get(f"/api/adventures/{client.adv_id}/state").json()["document"]
|
|
|
|
|
|
def packet_of(client) -> dict:
|
|
response = client.get(f"/api/adventures/{client.adv_id}/scene-packet")
|
|
assert response.status_code == 200, response.text[:300]
|
|
return response.json()
|
|
|
|
|
|
def head_of(client):
|
|
with SessionLocal() as db:
|
|
adventure = db.get(models.Adventure, client.adv_id)
|
|
return adventure.head_branch_id, adventure.head_depth
|
|
|
|
|
|
def settle(client, rounds=8):
|
|
"""Runs the derived pass until it has caught up, as a played campaign would.
|
|
|
|
`MAX_MEMORIES_PER_RUN` is 5, so one call settles at most five blocks — a cap
|
|
that exists so an imported campaign does not do all its catch-up inside one
|
|
turn. A test that calls it once and then asserts on the summary is asserting
|
|
against a half-settled campaign, which is how the first version of this file
|
|
failed: path A's later turns, the ones carrying the summary sentinel, had
|
|
not been summarised yet.
|
|
"""
|
|
for _ in range(rounds):
|
|
before = _settled_marks(client)
|
|
asyncio.run(memorybank.run_post_turn(client.adv_id))
|
|
if _settled_marks(client) == before:
|
|
return
|
|
|
|
|
|
def _settled_marks(client):
|
|
with SessionLocal() as db:
|
|
return (
|
|
db.query(models.Memory).filter(
|
|
models.Memory.adventure_id == client.adv_id).count(),
|
|
db.query(models.Summary).filter(
|
|
models.Summary.adventure_id == client.adv_id).count(),
|
|
)
|
|
|
|
|
|
@pytest.fixture()
|
|
def diverged(client, summariser):
|
|
"""One campaign: a long path A holding all four sentinels, then a path B.
|
|
|
|
Returns what the positive controls established on A, so the negative
|
|
controls on B can be asserted against something rather than against nothing.
|
|
"""
|
|
# ---- Path A. Long enough that summaries and memories are real. ----
|
|
play(client, "arrive", events=[
|
|
{"type": "create_entity", "entity": "aldric", "entity_type": "character",
|
|
"name": "Aldric"},
|
|
{"type": "create_entity", "entity": "grimwald", "entity_type": "character",
|
|
"name": "Grimwald"},
|
|
{"type": "create_entity", "entity": "tavern", "entity_type": "location",
|
|
"name": "The Crooked Lantern"},
|
|
{"type": "create_entity", "entity": SCENE_SENTINEL,
|
|
"entity_type": "location", "name": "The abbey crypt"},
|
|
{"type": "set_scene", "summary": "Aldric and Grimwald take the corner table.",
|
|
"location": "tavern", "present": ["aldric", "grimwald"]},
|
|
])
|
|
for i in range(10):
|
|
play(client, f"a{i}", prose=f"They talk on into the evening. [{i}]")
|
|
|
|
# The four sentinels, established together on the line that will be left.
|
|
play(client, "the confession", prose=(
|
|
f"Grimwald says it plainly: {MEMORY_SENTINEL}. They swear the "
|
|
f"{SUMMARY_SENTINEL} on it."
|
|
), events=[
|
|
{"type": "add_fact", "subject": "grimwald", "predicate": "confessed",
|
|
"object": "aldric", "fact_id": STATE_SENTINEL},
|
|
{"type": "set_current_location", "entity": "aldric",
|
|
"location": SCENE_SENTINEL},
|
|
{"type": "set_scene",
|
|
"summary": "Aldric stands in the abbey crypt, the seal broken.",
|
|
"location": SCENE_SENTINEL, "present": ["aldric"]},
|
|
])
|
|
for i in range(10):
|
|
play(client, f"a2{i}", prose=(
|
|
f"The crypt is cold, and the {SUMMARY_SENTINEL} still stands. [{i}]"))
|
|
settle(client)
|
|
|
|
before = {
|
|
"report": report(client),
|
|
"state": state_of(client),
|
|
"packet": packet_of(client),
|
|
"head": head_of(client),
|
|
}
|
|
with SessionLocal() as db:
|
|
adventure = db.get(models.Adventure, client.adv_id)
|
|
row = summaries.current(db, adventure)
|
|
before["summary_id"] = row.id if row else None
|
|
before["summary_text"] = row.text if row else ""
|
|
|
|
# ---- Move the head below every sentinel, then diverge. ----
|
|
while head_of(client)[1] > 11:
|
|
assert client.post(f"/api/adventures/{client.adv_id}/undo").status_code == 200
|
|
summariser.seeds.clear()
|
|
|
|
# ---- Path B. Far enough that a NEW summary is generated (M6-F1). ----
|
|
play(client, "b-turn", prose="Aldric leaves the table and takes the dry road.",
|
|
events=[
|
|
{"type": "set_current_location", "entity": "aldric", "location": "tavern"},
|
|
{"type": "set_scene", "summary": "Aldric alone on the road out of town.",
|
|
"location": "tavern", "present": ["aldric"]},
|
|
])
|
|
for i in range(14):
|
|
play(client, f"b{i}", prose=f"A dry road, nothing sworn, nothing confessed. [{i}]")
|
|
settle(client)
|
|
|
|
return {"before": before, "after": {
|
|
"report": report(client),
|
|
"state": state_of(client),
|
|
"packet": packet_of(client),
|
|
"head": head_of(client),
|
|
}}
|
|
|
|
|
|
# ------------------------------------------------------- positive controls
|
|
|
|
def test_path_a_really_established_all_four(diverged):
|
|
"""Without this, every assertion below proves only that nothing happened."""
|
|
before = diverged["before"]
|
|
prompt = prompt_of(before["report"])
|
|
|
|
facts = [f.get("id") for f in before["state"].get("facts", [])]
|
|
assert STATE_SENTINEL in facts, "the state sentinel was never established"
|
|
assert before["state"]["scene"]["location"] == SCENE_SENTINEL
|
|
assert before["packet"]["location"]["key"] == SCENE_SENTINEL
|
|
assert before["summary_id"] is not None, "no summary was generated on path A"
|
|
assert SUMMARY_SENTINEL in before["summary_text"], (
|
|
"the fixture did not get the sentinel into path A's summary")
|
|
assert SUMMARY_SENTINEL in prompt, "path A's prompt did not carry its own summary"
|
|
assert MEMORY_SENTINEL in prompt or any(
|
|
MEMORY_SENTINEL in (m.get("text") or "")
|
|
for m in (before["report"].get("memories") or {}).get("used", [])
|
|
), "the memory sentinel never reached path A's prompt"
|
|
|
|
|
|
# ------------------------------------------------------- E01: state
|
|
|
|
def test_e01_the_abandoned_fact_is_not_in_the_active_state(diverged):
|
|
facts = [f.get("id") for f in diverged["after"]["state"].get("facts", [])]
|
|
assert STATE_SENTINEL not in facts
|
|
|
|
|
|
def test_e01_the_abandoned_fact_is_not_in_the_active_prompt(diverged):
|
|
assert STATE_SENTINEL not in prompt_of(diverged["after"]["report"])
|
|
|
|
|
|
# ------------------------------------------------------- E02: memory
|
|
|
|
def test_e02_the_abandoned_memory_does_not_enter_the_active_prompt(diverged):
|
|
after = diverged["after"]["report"]
|
|
assert MEMORY_SENTINEL not in prompt_of(after)
|
|
used = (after.get("memories") or {}).get("used", [])
|
|
assert not any(MEMORY_SENTINEL in (m.get("text") or "") for m in used)
|
|
|
|
|
|
def test_e02_the_abandoned_memory_is_still_on_disk(client, diverged):
|
|
"""Retained, not deleted — the story was left, not erased (ADR 012)."""
|
|
with SessionLocal() as db:
|
|
stored = db.query(models.Memory).filter(
|
|
models.Memory.adventure_id == client.adv_id,
|
|
models.Memory.text.like(f"%{MEMORY_SENTINEL}%"),
|
|
).count()
|
|
assert stored > 0, "the abandoned memory was destroyed rather than retained"
|
|
|
|
|
|
# ------------------------------------------------------- E03: summary
|
|
|
|
def test_e03_a_new_summary_was_generated_on_the_new_line(client, diverged):
|
|
"""M6-F1's shape: the test is worthless unless a regeneration happened."""
|
|
with SessionLocal() as db:
|
|
adventure = db.get(models.Adventure, client.adv_id)
|
|
row = summaries.current(db, adventure)
|
|
assert row is not None, "no summary is eligible on path B"
|
|
assert row.id != diverged["before"]["summary_id"], (
|
|
"path B reused path A's summary row rather than generating one")
|
|
|
|
|
|
def test_e03_the_regenerated_summary_carries_no_abandoned_content(client, diverged):
|
|
with SessionLocal() as db:
|
|
adventure = db.get(models.Adventure, client.adv_id)
|
|
row = summaries.current(db, adventure)
|
|
assert SUMMARY_SENTINEL not in (row.text or "")
|
|
assert MEMORY_SENTINEL not in (row.text or "")
|
|
|
|
|
|
def test_e03_the_summariser_was_never_offered_the_abandoned_summary(summariser, diverged):
|
|
"""The fix is at the input. A filter over the output would be a different bug."""
|
|
assert summariser.seeds, "no summary was generated on path B"
|
|
assert not any(SUMMARY_SENTINEL in seed for seed in summariser.seeds)
|
|
|
|
|
|
def test_e03_no_abandoned_turn_is_on_the_active_lineage(client, diverged):
|
|
with SessionLocal() as db:
|
|
adventure = db.get(models.Adventure, client.adv_id)
|
|
leaked = db.query(models.Action).filter(
|
|
models.Action.adventure_id == client.adv_id,
|
|
lineage.path_of(db, adventure).clause(models.Action),
|
|
models.Action.text.like(f"%{SUMMARY_SENTINEL}%"),
|
|
).count()
|
|
assert leaked == 0, "the fixture left path-A story on path B's lineage"
|
|
|
|
|
|
def test_e03_the_abandoned_summary_row_is_retained(client, diverged):
|
|
with SessionLocal() as db:
|
|
kept = db.query(models.Summary).filter(
|
|
models.Summary.adventure_id == client.adv_id,
|
|
models.Summary.text.like(f"%{SUMMARY_SENTINEL}%"),
|
|
).count()
|
|
assert kept > 0, "the abandoned summary was deleted rather than retired"
|
|
|
|
|
|
# ------------------------------------------------------- E04: scene
|
|
|
|
def test_e04_the_current_scene_is_the_active_lines_scene(diverged):
|
|
"""The acceptance scenario, exactly: the discarded future moved to the abbey."""
|
|
scene = diverged["after"]["state"]["scene"]
|
|
assert scene["location"] == "tavern"
|
|
assert scene["location"] != SCENE_SENTINEL
|
|
|
|
|
|
def test_e04_the_protagonists_location_followed_the_active_line(diverged):
|
|
entities = diverged["after"]["state"].get("entities") or {}
|
|
assert (entities.get("aldric") or {}).get("location") != SCENE_SENTINEL
|
|
|
|
|
|
def test_e04_the_derived_scene_packet_shows_the_active_line_only(diverged):
|
|
"""M10's packet, which is what a future depiction would be built from.
|
|
|
|
The packet is derived from the authoritative state on read, so this cannot
|
|
fail while the state above passes — which is the point. It is asserted
|
|
anyway because the packet is a *new* surface since E04 was written, and a
|
|
later change that gave it a store of its own would fail here.
|
|
"""
|
|
packet = diverged["after"]["packet"]
|
|
assert packet["location"]["key"] == "tavern"
|
|
assert SCENE_SENTINEL not in repr(packet)
|
|
assert packet["scene_id"] != diverged["before"]["packet"]["scene_id"]
|
|
|
|
|
|
def test_e04_the_abandoned_scene_is_still_retained_at_its_own_position(client, diverged):
|
|
"""Retained history keeps its scene; it simply is not current."""
|
|
from sqlalchemy.orm import undefer
|
|
|
|
with SessionLocal() as db:
|
|
rows = (
|
|
db.query(models.Action)
|
|
.filter(models.Action.adventure_id == client.adv_id)
|
|
.options(undefer(models.Action.narrative_state_after))
|
|
.all()
|
|
)
|
|
kept = [
|
|
r for r in rows
|
|
if ((r.narrative_state_after or {}).get("scene") or {}).get("location")
|
|
== SCENE_SENTINEL
|
|
]
|
|
assert kept, "the abandoned line's scene was destroyed rather than retained"
|