Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
co-authored by
Claude Opus 5
parent
fedb7144d0
commit
ef25b0a876
@@ -0,0 +1,298 @@
|
||||
"""The history window moves in blocks, so the prompt's prefix holds still.
|
||||
|
||||
Inference servers cache a prompt by its **prefix**. While a story only grows at
|
||||
the end, every turn re-uses that cache and pays for its own new tokens alone. The
|
||||
builder's old window took whatever fit, which meant that once the budget was full
|
||||
it dropped the *oldest* action every turn — a change near the front of the prompt
|
||||
— and everything after it had to be processed again.
|
||||
|
||||
Measured on the reference deployment, at 7.7k prompt tokens against a 3B model:
|
||||
|
||||
window slid by one turn 343-350 s
|
||||
prefix preserved 6.0 s
|
||||
|
||||
These tests do not measure time. They pin the property the measurement is
|
||||
downstream of: **the oldest included action is the same across consecutive
|
||||
turns**, except on the turns where the window deliberately steps.
|
||||
"""
|
||||
import pytest
|
||||
import pytest as _pytest
|
||||
|
||||
from app.context import builder
|
||||
|
||||
|
||||
def costs_of(n, each=100):
|
||||
return [each] * n
|
||||
|
||||
|
||||
def depths(n, start=1):
|
||||
return list(range(start, start + n))
|
||||
|
||||
|
||||
# ------------------------------------------------------------- the block size
|
||||
|
||||
def test_the_block_is_a_share_of_what_fits():
|
||||
"""Derived from `TRIM_FRACTION` rather than asserting the number it is
|
||||
currently set to, so retuning the dial does not fail a test that was never
|
||||
about the dial's value."""
|
||||
# 16 actions of 500 fit in 8,000, and the block is that share of them.
|
||||
assert builder.trim_block(8000, 500) == 16 // builder.TRIM_FRACTION
|
||||
|
||||
|
||||
def test_trim_fraction_is_the_dial_between_history_and_speed():
|
||||
"""`TRIM_FRACTION` is meant to be retuned, so this pins what retuning does.
|
||||
|
||||
Lower it and the window gives up more at once: bigger blocks, fewer re-reads,
|
||||
less recent history retained. Raise it and the reverse. Nothing else in the
|
||||
builder has to change for that to hold, which is the property worth having a
|
||||
test for.
|
||||
"""
|
||||
budget, per_action = 8000, 500
|
||||
fits = budget // per_action
|
||||
|
||||
def block_at(fraction, monkeypatch):
|
||||
monkeypatch.setattr(builder, "TRIM_FRACTION", fraction)
|
||||
return builder.trim_block(budget, per_action)
|
||||
|
||||
with _pytest.MonkeyPatch.context() as mp:
|
||||
greedier = block_at(2, mp)
|
||||
assert greedier == fits // 2
|
||||
with _pytest.MonkeyPatch.context() as mp:
|
||||
gentler = block_at(8, mp)
|
||||
assert gentler == max(builder.MIN_TRIM_BLOCK, fits // 8)
|
||||
assert greedier > gentler, "a lower fraction must give up more at once"
|
||||
|
||||
# And the floor to the whole thing survives any setting.
|
||||
with _pytest.MonkeyPatch.context() as mp:
|
||||
mp.setattr(builder, "TRIM_FRACTION", 1000)
|
||||
assert builder.trim_block(budget, per_action) >= builder.MIN_TRIM_BLOCK
|
||||
|
||||
|
||||
def test_the_block_never_slides_by_one():
|
||||
"""A block of one is the old behaviour wearing a hat."""
|
||||
assert builder.trim_block(100, 500) >= builder.MIN_TRIM_BLOCK
|
||||
assert builder.trim_block(0, 500) >= builder.MIN_TRIM_BLOCK
|
||||
|
||||
|
||||
def test_the_block_comes_from_settings_not_from_the_story():
|
||||
"""It has to be the same on two consecutive turns, so it cannot be measured
|
||||
from actions whose sizes vary."""
|
||||
assert builder.trim_block(8000, 500) == builder.trim_block(8000, 500)
|
||||
# Bigger budget, bigger step; the ratio is what is fixed.
|
||||
assert builder.trim_block(16000, 500) > builder.trim_block(8000, 500)
|
||||
|
||||
|
||||
# ------------------------------------------------------- nothing to trim yet
|
||||
|
||||
def test_a_story_that_fits_whole_is_not_trimmed():
|
||||
"""Also the append-only regime: every turn is a prefix extension already."""
|
||||
assert builder.history_floor(depths(5), costs_of(5), budget=10_000, block=4) is None
|
||||
|
||||
|
||||
def test_a_short_story_keeps_its_opening():
|
||||
"""Snapping here would drop the start of the story for no reason at all."""
|
||||
assert builder.history_floor(depths(3), costs_of(3), budget=10_000, block=8) is None
|
||||
|
||||
|
||||
def test_an_action_larger_than_the_budget_is_left_to_the_caller():
|
||||
assert builder.history_floor([1], [5000], budget=100, block=4) is None
|
||||
|
||||
|
||||
def test_rows_without_a_depth_are_not_trimmed():
|
||||
"""Legacy rows have no stable coordinate, so behave exactly as before."""
|
||||
assert builder.history_floor([None, None], costs_of(2), 100, 4) is None
|
||||
assert builder.history_floor([], [], 100, 4) is None
|
||||
|
||||
|
||||
# ------------------------------------------------------------ the whole point
|
||||
|
||||
def test_the_floor_holds_still_while_the_story_grows():
|
||||
"""The property the 57x measurement rests on.
|
||||
|
||||
Ten consecutive turns against a full budget. The floor must take a small
|
||||
number of steps, not ten.
|
||||
"""
|
||||
block, budget, each = 4, 1000, 100 # 10 actions fit
|
||||
seen = []
|
||||
for extra in range(10): # the story grows by one action
|
||||
n = 20 + extra
|
||||
seen.append(builder.history_floor(depths(n), costs_of(n, each), budget, block))
|
||||
|
||||
steps = sum(1 for a, b in zip(seen, seen[1:]) if a != b)
|
||||
assert steps <= 3, f"the floor moved {steps} times in 10 turns: {seen}"
|
||||
assert len(set(seen)) > 1, "it never moved at all, so the budget is not binding"
|
||||
|
||||
|
||||
def test_every_floor_sits_on_a_block_boundary():
|
||||
block, budget = 4, 1000
|
||||
for n in range(20, 40):
|
||||
floor = builder.history_floor(depths(n), costs_of(n), budget, block)
|
||||
assert floor is not None
|
||||
assert floor % block == 0, f"{floor} is not a multiple of {block}"
|
||||
|
||||
|
||||
def test_the_floor_only_ever_moves_forward():
|
||||
block, budget = 4, 1000
|
||||
floors = [builder.history_floor(depths(n), costs_of(n), budget, block)
|
||||
for n in range(20, 45)]
|
||||
assert floors == sorted(floors)
|
||||
|
||||
|
||||
# --------------------------------------------------- and still inside budget
|
||||
|
||||
@pytest.mark.parametrize("n", range(20, 40))
|
||||
def test_the_kept_window_never_exceeds_the_budget(n):
|
||||
"""M03's bound is not weakened. Trimming only ever drops more, never less."""
|
||||
block, budget, each = 4, 1000, 100
|
||||
ds, cs = depths(n), costs_of(n, each)
|
||||
floor = builder.history_floor(ds, cs, budget, block)
|
||||
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
|
||||
assert kept <= budget
|
||||
|
||||
|
||||
@pytest.mark.parametrize("n", range(20, 40))
|
||||
def test_the_kept_window_is_not_gutted(n):
|
||||
"""The cost of holding still is bounded: a trim gives up `block` actions,
|
||||
never most of the window."""
|
||||
block, budget, each = 4, 1000, 100
|
||||
ds, cs = depths(n), costs_of(n, each)
|
||||
floor = builder.history_floor(ds, cs, budget, block)
|
||||
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
|
||||
assert kept >= budget - block * each
|
||||
|
||||
|
||||
# ------------------------------------------- the real builder, end to end
|
||||
|
||||
import pytest as _pytest # noqa: E402 (grouped with the fixtures it serves)
|
||||
|
||||
from app import models # noqa: E402
|
||||
from app.context import builder as _builder # noqa: E402
|
||||
from app.database import Base, SessionLocal, engine # noqa: E402
|
||||
|
||||
NARRATION = ("The rain came down over Westhaven in long grey sheets and the "
|
||||
"gutters ran full from the ridge to the waterfront. ") * 6
|
||||
|
||||
|
||||
@_pytest.fixture()
|
||||
def saturated():
|
||||
"""A campaign whose history is longer than its budget, with real depths.
|
||||
|
||||
`depth` is what the floor is expressed in, and every action written through
|
||||
the application has one (`tree.place_action`). The older fixtures in
|
||||
`test_history_window.py` predate the tree and leave it null, which is why
|
||||
trimming does not engage there and those tests still describe the old
|
||||
behaviour exactly.
|
||||
"""
|
||||
Base.metadata.create_all(bind=engine)
|
||||
db = SessionLocal()
|
||||
user = models.User(is_guest=False, email="blocktrim@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
settings = models.Settings(user_id=user.id, model="m", embedding_model="",
|
||||
context_token_budget=2048, max_output_tokens=200)
|
||||
db.add(settings)
|
||||
adventure = models.Adventure(user_id=user.id, title="Long", script_state={})
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
for i in range(60):
|
||||
db.add(models.Action(adventure_id=adventure.id,
|
||||
type="ai" if i % 2 else "do",
|
||||
text=f"[{i}] {NARRATION}", branch_id=None, depth=i))
|
||||
db.commit()
|
||||
db.expire_all()
|
||||
adventure = db.get(models.Adventure, adventure.id)
|
||||
settings = db.get(models.Settings, settings.id)
|
||||
try:
|
||||
yield db, adventure, settings
|
||||
finally:
|
||||
db.close()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _play_one_more(db, adventure, at_depth):
|
||||
db.add(models.Action(adventure_id=adventure.id, type="do",
|
||||
text=f"[{at_depth}] {NARRATION}", depth=at_depth))
|
||||
# `at_depth` may be None: the control below plays a turn into a story whose
|
||||
# rows predate the tree, which is the ungoverned window this replaced.
|
||||
db.commit()
|
||||
db.expire_all()
|
||||
|
||||
|
||||
def _shared_prefix(before: str, after: str) -> float:
|
||||
"""How much of the old prompt the new one still opens with, 0.0 to 1.0.
|
||||
|
||||
This is the quantity the inference server's cache is keyed on, so it is the
|
||||
quantity worth asserting. It is not 1.0 even in the best case: the prompt
|
||||
ends with the turn's length-hint and state-block instructions, which sit
|
||||
*after* the history, so appending a turn always rewrites that tail.
|
||||
"""
|
||||
shared = 0
|
||||
for x, y in zip(before, after):
|
||||
if x != y:
|
||||
break
|
||||
shared += 1
|
||||
return shared / max(1, len(before))
|
||||
|
||||
|
||||
def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
|
||||
"""The property the whole change exists for.
|
||||
|
||||
Not a timing test — it asserts what the timing follows from. The story text
|
||||
a turn sends still opens with almost all of what the previous turn sent, so
|
||||
the server's prompt cache covers that part and only the tail is processed.
|
||||
"""
|
||||
db, adventure, settings = saturated
|
||||
_, before, report_before = _builder.build_context(adventure, settings)
|
||||
assert report_before["history"]["floor_depth"] is not None, (
|
||||
"this fixture is meant to be over budget; trimming never engaged")
|
||||
|
||||
_play_one_more(db, adventure, 60)
|
||||
_, after, report_after = _builder.build_context(adventure, settings)
|
||||
|
||||
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"]
|
||||
assert _shared_prefix(before, after) > 0.85
|
||||
|
||||
|
||||
def test_without_a_stable_floor_the_prefix_collapses(saturated):
|
||||
"""The control, and the behaviour this replaced.
|
||||
|
||||
Rows with no `depth` cannot be placed on the tree, so the floor cannot be
|
||||
computed and the window takes whatever fits — sliding by one action every
|
||||
turn. The new prompt then starts with a *different* action, the shared
|
||||
prefix collapses, and the server reprocesses essentially the whole thing.
|
||||
That is the 343s case in this module's docstring.
|
||||
"""
|
||||
db, adventure, settings = saturated
|
||||
for action in db.query(models.Action).all():
|
||||
action.depth = None
|
||||
db.commit()
|
||||
db.expire_all()
|
||||
|
||||
_, before, report = _builder.build_context(adventure, settings)
|
||||
assert report["history"]["floor_depth"] is None
|
||||
_play_one_more(db, adventure, None)
|
||||
_, after, _ = _builder.build_context(adventure, settings)
|
||||
|
||||
assert _shared_prefix(before, after) < 0.1
|
||||
|
||||
|
||||
def test_the_window_does_step_eventually(saturated):
|
||||
"""It holds still, but it must not hold still for ever — the budget is a
|
||||
bound, and a window that never moved would break it."""
|
||||
db, adventure, settings = saturated
|
||||
first = _builder.build_context(adventure, settings)[2]["history"]["floor_depth"]
|
||||
|
||||
seen = {first}
|
||||
for depth in range(60, 90):
|
||||
_play_one_more(db, adventure, depth)
|
||||
seen.add(_builder.build_context(adventure, settings)[2]["history"]["floor_depth"])
|
||||
assert len(seen) > 1, "the floor never moved across 30 turns"
|
||||
|
||||
|
||||
def test_the_prompt_stays_inside_the_budget_as_the_window_steps(saturated):
|
||||
"""M03's bound, across the step. Trimming only ever drops more history."""
|
||||
db, adventure, settings = saturated
|
||||
for depth in range(60, 85):
|
||||
_play_one_more(db, adventure, depth)
|
||||
report = _builder.build_context(adventure, settings)[2]
|
||||
assert report["tokens"]["total"] <= report["tokens"]["budget"]
|
||||
@@ -369,7 +369,7 @@ def test_a_turn_records_the_window_it_was_built_against(client, monkeypatch):
|
||||
the campaign whether it was built against a checked window, rather than
|
||||
inferring it from what the settings say today.
|
||||
"""
|
||||
async def verified(endpoint, model, use_cache=True):
|
||||
async def verified(endpoint, model, declared=None, use_cache=True):
|
||||
return contextwindow.Window(4096, contextwindow.LOADED, 32768, "fake")
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
|
||||
|
||||
@@ -0,0 +1,165 @@
|
||||
"""The window an operator declares, for a server that cannot be asked for one.
|
||||
|
||||
`contextwindow`'s discovery speaks Ollama's native API. Nothing restricts
|
||||
`endpoint_url` to Ollama, so on vLLM, llama.cpp's own server, or anything else
|
||||
serving an OpenAI-compatible `/v1`, `/api/ps` and `/api/show` are not there:
|
||||
discovery fails as designed, the window is unknown, and the budget is left
|
||||
uncapped at whatever is configured. That is M11's own failure mode reached by a
|
||||
different route — the server drops the oldest tokens, which here are the
|
||||
narrator's rules and the campaign canon.
|
||||
|
||||
`Settings.context_window_override` closes it. These tests pin the two properties
|
||||
that make it safe rather than merely useful:
|
||||
|
||||
1. it is used **only** where discovery left a hole, so it can never talk the
|
||||
application into a longer prompt than a server actually reported, and
|
||||
2. it does not make `verified` true, because `verified` means the server
|
||||
answered and a declaration is a person's claim about a server.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, contextwindow, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
|
||||
|
||||
UNREACHABLE = "http://127.0.0.1:1/v1"
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_window_cache():
|
||||
contextwindow.cache_clear()
|
||||
yield
|
||||
contextwindow.cache_clear()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m11dw@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="some-model", endpoint_url=UNREACHABLE,
|
||||
embedding_model="", context_token_budget=16384, max_output_tokens=800,
|
||||
))
|
||||
setup.commit()
|
||||
user_id = user.id
|
||||
setup.close()
|
||||
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
try:
|
||||
yield TestClient(app)
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def probe(endpoint=UNREACHABLE, model="some-model", declared=None):
|
||||
contextwindow.cache_clear()
|
||||
return asyncio.run(
|
||||
contextwindow.probe(endpoint, model, declared=declared, use_cache=False))
|
||||
|
||||
|
||||
# ------------------------------------------------------- filling the hole
|
||||
|
||||
def test_without_a_declaration_an_unaskable_server_leaves_the_window_unknown():
|
||||
window = probe()
|
||||
assert window.tokens is None
|
||||
assert not window.verified
|
||||
assert not window.enforceable
|
||||
assert window.source == contextwindow.UNKNOWN
|
||||
|
||||
|
||||
def test_a_declaration_becomes_the_ceiling_when_the_server_cannot_be_asked():
|
||||
window = probe(declared=8192)
|
||||
assert window.tokens == 8192
|
||||
assert window.source == contextwindow.DECLARED
|
||||
assert window.enforceable
|
||||
assert contextwindow.effective_budget(16384, window) == 8192
|
||||
|
||||
|
||||
def test_a_declaration_does_not_claim_the_server_was_verified():
|
||||
"""`window_verified` travels in every turn's provenance and the M11 report
|
||||
counts it. A declaration must not inflate that count."""
|
||||
window = probe(declared=8192)
|
||||
assert window.enforceable
|
||||
assert not window.verified
|
||||
|
||||
|
||||
def test_the_detail_says_the_number_came_from_settings():
|
||||
assert "declared in settings" in probe(declared=8192).detail
|
||||
|
||||
|
||||
@pytest.mark.parametrize("declared", [None, 0, -1])
|
||||
def test_a_missing_or_meaningless_declaration_changes_nothing(declared):
|
||||
window = probe(declared=declared)
|
||||
assert window.tokens is None
|
||||
assert window.source == contextwindow.UNKNOWN
|
||||
|
||||
|
||||
def test_a_declaration_still_applies_when_nothing_is_configured():
|
||||
window = asyncio.run(contextwindow.probe("", "", declared=4096))
|
||||
assert window.tokens == 4096
|
||||
assert window.source == contextwindow.DECLARED
|
||||
|
||||
|
||||
def test_a_declaration_applies_to_a_refused_endpoint_without_reaching_it():
|
||||
"""A refused address is a discovery failure like any other (ADR 011, H12).
|
||||
The declaration caps the prompt; it does not make the endpoint usable, and
|
||||
the turn is still refused where endpoints are enforced."""
|
||||
window = probe(endpoint="http://169.254.169.254/v1", declared=4096)
|
||||
assert window.source == contextwindow.DECLARED
|
||||
assert window.tokens == 4096
|
||||
|
||||
|
||||
# ------------------------------------------- a verified answer always wins
|
||||
|
||||
def test_a_verified_window_is_not_overridden(monkeypatch):
|
||||
"""The safety property. An operator may lower an unknown ceiling into
|
||||
existence; they may never raise one the server reported."""
|
||||
async def reported(endpoint, model):
|
||||
return contextwindow.Window(4096, contextwindow.LOADED, 32768, "real")
|
||||
|
||||
monkeypatch.setattr(contextwindow, "_ask", reported)
|
||||
window = probe(declared=32768)
|
||||
assert window.tokens == 4096
|
||||
assert window.source == contextwindow.LOADED
|
||||
assert window.verified
|
||||
assert contextwindow.effective_budget(16384, window) == 4096
|
||||
|
||||
|
||||
def test_a_declared_window_larger_than_the_budget_does_not_raise_it():
|
||||
window = probe(declared=200_000)
|
||||
assert contextwindow.effective_budget(16384, window) == 16384
|
||||
|
||||
|
||||
# --------------------------------------------------------------- plumbing
|
||||
|
||||
def test_the_override_is_readable_and_settable_through_the_api(client):
|
||||
assert client.get("/api/settings").json()["context_window_override"] is None
|
||||
|
||||
body = client.put("/api/settings",
|
||||
json={"context_window_override": 8192}).json()
|
||||
assert body["context_window_override"] == 8192
|
||||
|
||||
# And can be taken back off, which `exclude_unset` makes a real distinction:
|
||||
# sending null clears it, sending nothing leaves it alone.
|
||||
body = client.put("/api/settings", json={"temperature": 0.5}).json()
|
||||
assert body["context_window_override"] == 8192
|
||||
body = client.put("/api/settings",
|
||||
json={"context_window_override": None}).json()
|
||||
assert body["context_window_override"] is None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("bad", [255, 200_001])
|
||||
def test_the_override_is_bounded_like_the_budget_it_caps(client, bad):
|
||||
assert client.put("/api/settings",
|
||||
json={"context_window_override": bad}).status_code == 422
|
||||
@@ -0,0 +1,225 @@
|
||||
"""The long run's resume checkpoint, and the refusals that protect its evidence.
|
||||
|
||||
M11's release campaign was lost twice over: once to a host crash at turn 97, and
|
||||
again to the fact that starting the harness a second time began a new campaign
|
||||
rather than continuing the old one. `tools/m11_long_run.py` now checkpoints
|
||||
`resume.json` and can be pointed back at it.
|
||||
|
||||
These tests exercise that logic without a narrator, a server or a database,
|
||||
because none of it needs one: the checkpoint is a file, and the decisions made
|
||||
around it are decisions about files. What they cannot prove is that a resumed
|
||||
campaign continues correctly against a real application — that is what the run
|
||||
itself proves, and §G of the M11 report is where it is reported.
|
||||
"""
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from tools import m11_long_run as lr
|
||||
|
||||
|
||||
class FakeServer:
|
||||
"""Enough of `Storyteller` for the checkpoint: it records process starts."""
|
||||
|
||||
def __init__(self, starts=1):
|
||||
self.starts = starts
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def run(tmp_path):
|
||||
made = lr.Run(FakeServer(), tmp_path, turns_target=100)
|
||||
yield made
|
||||
made.timeline.close()
|
||||
|
||||
|
||||
# --------------------------------------------------------------- checkpoint
|
||||
|
||||
def test_a_checkpoint_carries_what_a_resume_needs(run, tmp_path):
|
||||
run.adv = 7
|
||||
run.accepted = 41
|
||||
run.beat = 44
|
||||
run.completed_steps = {7, 14, 21}
|
||||
run.save_resume()
|
||||
|
||||
saved = json.loads((tmp_path / lr.RESUME_FILE).read_text())
|
||||
assert saved["adventure"] == 7
|
||||
assert saved["accepted"] == 41
|
||||
assert saved["beat"] == 44
|
||||
assert saved["completed_steps"] == [7, 14, 21]
|
||||
assert saved["server_starts"] == 1
|
||||
assert saved["turns_target"] == 100
|
||||
|
||||
|
||||
def test_the_checkpoint_is_replaced_rather_than_appended(run, tmp_path):
|
||||
run.adv = 7
|
||||
run.accepted = 1
|
||||
run.save_resume()
|
||||
run.accepted = 2
|
||||
run.save_resume()
|
||||
|
||||
assert json.loads((tmp_path / lr.RESUME_FILE).read_text())["accepted"] == 2
|
||||
# The temporary name it is written under must not survive the rename.
|
||||
assert not (tmp_path / (lr.RESUME_FILE + ".tmp")).exists()
|
||||
|
||||
|
||||
def test_a_checkpoint_round_trips_into_a_later_session(run, tmp_path):
|
||||
run.adv = 7
|
||||
run.accepted = 41
|
||||
run.beat = 44
|
||||
run.completed_steps = {7, 14}
|
||||
run.elapsed_before = 100
|
||||
run.save_resume()
|
||||
saved = json.loads((tmp_path / lr.RESUME_FILE).read_text())
|
||||
|
||||
later = lr.Run(FakeServer(starts=3), tmp_path, turns_target=100)
|
||||
try:
|
||||
later.adopt(saved)
|
||||
assert later.adv == 7
|
||||
assert later.accepted == 41
|
||||
assert later.beat == 44
|
||||
assert later.completed_steps == {7, 14}
|
||||
assert later.resumed is True
|
||||
# Run time accumulates across sessions rather than restarting.
|
||||
assert later.elapsed_before >= 100
|
||||
assert later.elapsed() >= 100
|
||||
finally:
|
||||
later.timeline.close()
|
||||
|
||||
|
||||
def test_a_resumed_session_appends_to_the_existing_timeline(run, tmp_path):
|
||||
run.adv = 7
|
||||
run.note("turn", text="the first session")
|
||||
run.timeline.close()
|
||||
|
||||
later = lr.Run(FakeServer(), tmp_path, turns_target=100)
|
||||
try:
|
||||
later.note("resumed", adventure=7)
|
||||
finally:
|
||||
later.timeline.close()
|
||||
|
||||
lines = (tmp_path / "timeline.jsonl").read_text().strip().splitlines()
|
||||
assert [json.loads(line)["kind"] for line in lines] == ["turn", "resumed"]
|
||||
|
||||
|
||||
# ------------------------------------------------------------- the decision
|
||||
|
||||
def test_a_clean_directory_starts_a_run(tmp_path):
|
||||
assert lr._resume_state(
|
||||
tmp_path / lr.RESUME_FILE, tmp_path / "timeline.jsonl", False) is None
|
||||
|
||||
|
||||
def test_an_unfinished_run_is_not_overwritten(tmp_path):
|
||||
resume_path = tmp_path / lr.RESUME_FILE
|
||||
resume_path.write_text(json.dumps({"adventure": 7, "accepted": 41}))
|
||||
|
||||
refusal = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", False)
|
||||
assert isinstance(refusal, str)
|
||||
assert "--resume" in refusal
|
||||
|
||||
|
||||
def test_a_recorded_run_with_no_checkpoint_is_not_reused(tmp_path):
|
||||
"""A run that recorded something and then died before its first checkpoint.
|
||||
Starting here would put a second campaign in the same timeline."""
|
||||
timeline = tmp_path / "timeline.jsonl"
|
||||
timeline.write_text(json.dumps({"kind": "settings"}) + "\n")
|
||||
|
||||
refusal = lr._resume_state(tmp_path / lr.RESUME_FILE, timeline, False)
|
||||
assert isinstance(refusal, str)
|
||||
assert "second campaign" in refusal
|
||||
|
||||
|
||||
def test_a_run_that_recorded_nothing_leaves_the_directory_usable(tmp_path):
|
||||
"""A server that never came up opens the timeline and writes no line to it.
|
||||
Nothing was written that a fresh run could collide with."""
|
||||
(tmp_path / "timeline.jsonl").write_text("")
|
||||
|
||||
assert lr._resume_state(
|
||||
tmp_path / lr.RESUME_FILE, tmp_path / "timeline.jsonl", False) is None
|
||||
|
||||
|
||||
def test_resuming_returns_the_checkpoint(tmp_path):
|
||||
resume_path = tmp_path / lr.RESUME_FILE
|
||||
resume_path.write_text(json.dumps({"adventure": 7, "accepted": 41}))
|
||||
|
||||
prior = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", True)
|
||||
assert prior["adventure"] == 7
|
||||
assert prior["accepted"] == 41
|
||||
|
||||
|
||||
def test_resuming_nothing_is_refused_rather_than_started_fresh(tmp_path):
|
||||
refusal = lr._resume_state(
|
||||
tmp_path / lr.RESUME_FILE, tmp_path / "timeline.jsonl", True)
|
||||
assert isinstance(refusal, str)
|
||||
assert "no resume.json" in refusal
|
||||
|
||||
|
||||
def test_an_unreadable_checkpoint_is_refused(tmp_path):
|
||||
resume_path = tmp_path / lr.RESUME_FILE
|
||||
resume_path.write_text("{not json")
|
||||
|
||||
refusal = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", True)
|
||||
assert isinstance(refusal, str)
|
||||
assert "cannot read" in refusal
|
||||
|
||||
|
||||
def test_a_checkpoint_naming_no_campaign_is_refused(tmp_path):
|
||||
resume_path = tmp_path / lr.RESUME_FILE
|
||||
resume_path.write_text(json.dumps({"accepted": 41}))
|
||||
|
||||
refusal = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", True)
|
||||
assert isinstance(refusal, str)
|
||||
assert "names no campaign" in refusal
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- timeouts
|
||||
|
||||
def test_the_harness_waits_longer_than_the_application_does(tmp_path):
|
||||
"""Otherwise the socket closes before the application can report the failure
|
||||
inside the stream, and a real error is recorded as a transport one."""
|
||||
server = lr.Storyteller(tmp_path / "campaign.db", tmp_path / "server.log",
|
||||
turn_timeout=1800)
|
||||
assert server.stream_timeout > server.turn_timeout
|
||||
|
||||
|
||||
def test_the_default_timeout_is_inside_the_settings_bound():
|
||||
"""`app/schemas.py` bounds model_timeout_seconds at 30..3600."""
|
||||
assert 30 <= lr.DEFAULT_TURN_TIMEOUT <= 3600
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- schedule
|
||||
|
||||
def test_every_scheduled_operation_has_its_own_turn(tmp_path):
|
||||
"""The steps are keyed by turn number, which is what lets a completed one be
|
||||
remembered across a resume: two are called `retry` and two `restart`, so a
|
||||
name does not identify one."""
|
||||
plan = lr._schedule(100)
|
||||
assert len(plan) == len(set(plan)) == 12
|
||||
assert sorted(plan)[0] >= 1
|
||||
assert max(plan) < 100
|
||||
names = list(plan.values())
|
||||
assert names.count("restart") == 2
|
||||
assert names.count("retry") == 2
|
||||
|
||||
|
||||
# --------------------------------------------------------------- the clue
|
||||
|
||||
def test_the_planted_clue_uses_a_field_add_fact_actually_carries():
|
||||
"""M04's state half turns on the clue text reaching the stored fact.
|
||||
|
||||
`add_fact` requires `predicate` and accepts `subject`, `object`, `value` and
|
||||
`fact_id`. A key it does not define is dropped, and the correction still
|
||||
succeeds — so a clue planted into the wrong key leaves a fact asserting
|
||||
nothing, and `_recall` reports a recall failure the application did not
|
||||
cause. This test fails against the `detail` key that used to be sent.
|
||||
"""
|
||||
from app.narrative.events import SPECS
|
||||
|
||||
spec = SPECS["add_fact"]
|
||||
allowed = {"type"} | set(spec["required"]) | set(spec["optional"])
|
||||
assert set(lr.CLUE_FACT) <= allowed, (
|
||||
f"{set(lr.CLUE_FACT) - allowed} is not carried by add_fact")
|
||||
|
||||
|
||||
def test_the_planted_clue_carries_the_sentinel_recall_looks_for():
|
||||
assert lr.CLUE_SENTINEL in lr.CLUE_FACT["value"]
|
||||
assert lr.CLUE in lr.CLUE_FACT["value"]
|
||||
Reference in New Issue
Block a user