Files
interactive-story/backend/tests/test_history_block_trim.py
T
JesseMarkowitzandClaude Opus 5 ef25b0a876 Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
2026-09-10 06:13:55 -04:00

299 lines
12 KiB
Python

"""The history window moves in blocks, so the prompt's prefix holds still.
Inference servers cache a prompt by its **prefix**. While a story only grows at
the end, every turn re-uses that cache and pays for its own new tokens alone. The
builder's old window took whatever fit, which meant that once the budget was full
it dropped the *oldest* action every turn — a change near the front of the prompt
— and everything after it had to be processed again.
Measured on the reference deployment, at 7.7k prompt tokens against a 3B model:
window slid by one turn 343-350 s
prefix preserved 6.0 s
These tests do not measure time. They pin the property the measurement is
downstream of: **the oldest included action is the same across consecutive
turns**, except on the turns where the window deliberately steps.
"""
import pytest
import pytest as _pytest
from app.context import builder
def costs_of(n, each=100):
return [each] * n
def depths(n, start=1):
return list(range(start, start + n))
# ------------------------------------------------------------- the block size
def test_the_block_is_a_share_of_what_fits():
"""Derived from `TRIM_FRACTION` rather than asserting the number it is
currently set to, so retuning the dial does not fail a test that was never
about the dial's value."""
# 16 actions of 500 fit in 8,000, and the block is that share of them.
assert builder.trim_block(8000, 500) == 16 // builder.TRIM_FRACTION
def test_trim_fraction_is_the_dial_between_history_and_speed():
"""`TRIM_FRACTION` is meant to be retuned, so this pins what retuning does.
Lower it and the window gives up more at once: bigger blocks, fewer re-reads,
less recent history retained. Raise it and the reverse. Nothing else in the
builder has to change for that to hold, which is the property worth having a
test for.
"""
budget, per_action = 8000, 500
fits = budget // per_action
def block_at(fraction, monkeypatch):
monkeypatch.setattr(builder, "TRIM_FRACTION", fraction)
return builder.trim_block(budget, per_action)
with _pytest.MonkeyPatch.context() as mp:
greedier = block_at(2, mp)
assert greedier == fits // 2
with _pytest.MonkeyPatch.context() as mp:
gentler = block_at(8, mp)
assert gentler == max(builder.MIN_TRIM_BLOCK, fits // 8)
assert greedier > gentler, "a lower fraction must give up more at once"
# And the floor to the whole thing survives any setting.
with _pytest.MonkeyPatch.context() as mp:
mp.setattr(builder, "TRIM_FRACTION", 1000)
assert builder.trim_block(budget, per_action) >= builder.MIN_TRIM_BLOCK
def test_the_block_never_slides_by_one():
"""A block of one is the old behaviour wearing a hat."""
assert builder.trim_block(100, 500) >= builder.MIN_TRIM_BLOCK
assert builder.trim_block(0, 500) >= builder.MIN_TRIM_BLOCK
def test_the_block_comes_from_settings_not_from_the_story():
"""It has to be the same on two consecutive turns, so it cannot be measured
from actions whose sizes vary."""
assert builder.trim_block(8000, 500) == builder.trim_block(8000, 500)
# Bigger budget, bigger step; the ratio is what is fixed.
assert builder.trim_block(16000, 500) > builder.trim_block(8000, 500)
# ------------------------------------------------------- nothing to trim yet
def test_a_story_that_fits_whole_is_not_trimmed():
"""Also the append-only regime: every turn is a prefix extension already."""
assert builder.history_floor(depths(5), costs_of(5), budget=10_000, block=4) is None
def test_a_short_story_keeps_its_opening():
"""Snapping here would drop the start of the story for no reason at all."""
assert builder.history_floor(depths(3), costs_of(3), budget=10_000, block=8) is None
def test_an_action_larger_than_the_budget_is_left_to_the_caller():
assert builder.history_floor([1], [5000], budget=100, block=4) is None
def test_rows_without_a_depth_are_not_trimmed():
"""Legacy rows have no stable coordinate, so behave exactly as before."""
assert builder.history_floor([None, None], costs_of(2), 100, 4) is None
assert builder.history_floor([], [], 100, 4) is None
# ------------------------------------------------------------ the whole point
def test_the_floor_holds_still_while_the_story_grows():
"""The property the 57x measurement rests on.
Ten consecutive turns against a full budget. The floor must take a small
number of steps, not ten.
"""
block, budget, each = 4, 1000, 100 # 10 actions fit
seen = []
for extra in range(10): # the story grows by one action
n = 20 + extra
seen.append(builder.history_floor(depths(n), costs_of(n, each), budget, block))
steps = sum(1 for a, b in zip(seen, seen[1:]) if a != b)
assert steps <= 3, f"the floor moved {steps} times in 10 turns: {seen}"
assert len(set(seen)) > 1, "it never moved at all, so the budget is not binding"
def test_every_floor_sits_on_a_block_boundary():
block, budget = 4, 1000
for n in range(20, 40):
floor = builder.history_floor(depths(n), costs_of(n), budget, block)
assert floor is not None
assert floor % block == 0, f"{floor} is not a multiple of {block}"
def test_the_floor_only_ever_moves_forward():
block, budget = 4, 1000
floors = [builder.history_floor(depths(n), costs_of(n), budget, block)
for n in range(20, 45)]
assert floors == sorted(floors)
# --------------------------------------------------- and still inside budget
@pytest.mark.parametrize("n", range(20, 40))
def test_the_kept_window_never_exceeds_the_budget(n):
"""M03's bound is not weakened. Trimming only ever drops more, never less."""
block, budget, each = 4, 1000, 100
ds, cs = depths(n), costs_of(n, each)
floor = builder.history_floor(ds, cs, budget, block)
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
assert kept <= budget
@pytest.mark.parametrize("n", range(20, 40))
def test_the_kept_window_is_not_gutted(n):
"""The cost of holding still is bounded: a trim gives up `block` actions,
never most of the window."""
block, budget, each = 4, 1000, 100
ds, cs = depths(n), costs_of(n, each)
floor = builder.history_floor(ds, cs, budget, block)
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
assert kept >= budget - block * each
# ------------------------------------------- the real builder, end to end
import pytest as _pytest # noqa: E402 (grouped with the fixtures it serves)
from app import models # noqa: E402
from app.context import builder as _builder # noqa: E402
from app.database import Base, SessionLocal, engine # noqa: E402
NARRATION = ("The rain came down over Westhaven in long grey sheets and the "
"gutters ran full from the ridge to the waterfront. ") * 6
@_pytest.fixture()
def saturated():
"""A campaign whose history is longer than its budget, with real depths.
`depth` is what the floor is expressed in, and every action written through
the application has one (`tree.place_action`). The older fixtures in
`test_history_window.py` predate the tree and leave it null, which is why
trimming does not engage there and those tests still describe the old
behaviour exactly.
"""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
user = models.User(is_guest=False, email="blocktrim@example.com")
db.add(user)
db.flush()
settings = models.Settings(user_id=user.id, model="m", embedding_model="",
context_token_budget=2048, max_output_tokens=200)
db.add(settings)
adventure = models.Adventure(user_id=user.id, title="Long", script_state={})
db.add(adventure)
db.flush()
for i in range(60):
db.add(models.Action(adventure_id=adventure.id,
type="ai" if i % 2 else "do",
text=f"[{i}] {NARRATION}", branch_id=None, depth=i))
db.commit()
db.expire_all()
adventure = db.get(models.Adventure, adventure.id)
settings = db.get(models.Settings, settings.id)
try:
yield db, adventure, settings
finally:
db.close()
Base.metadata.drop_all(bind=engine)
def _play_one_more(db, adventure, at_depth):
db.add(models.Action(adventure_id=adventure.id, type="do",
text=f"[{at_depth}] {NARRATION}", depth=at_depth))
# `at_depth` may be None: the control below plays a turn into a story whose
# rows predate the tree, which is the ungoverned window this replaced.
db.commit()
db.expire_all()
def _shared_prefix(before: str, after: str) -> float:
"""How much of the old prompt the new one still opens with, 0.0 to 1.0.
This is the quantity the inference server's cache is keyed on, so it is the
quantity worth asserting. It is not 1.0 even in the best case: the prompt
ends with the turn's length-hint and state-block instructions, which sit
*after* the history, so appending a turn always rewrites that tail.
"""
shared = 0
for x, y in zip(before, after):
if x != y:
break
shared += 1
return shared / max(1, len(before))
def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
"""The property the whole change exists for.
Not a timing test — it asserts what the timing follows from. The story text
a turn sends still opens with almost all of what the previous turn sent, so
the server's prompt cache covers that part and only the tail is processed.
"""
db, adventure, settings = saturated
_, before, report_before = _builder.build_context(adventure, settings)
assert report_before["history"]["floor_depth"] is not None, (
"this fixture is meant to be over budget; trimming never engaged")
_play_one_more(db, adventure, 60)
_, after, report_after = _builder.build_context(adventure, settings)
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"]
assert _shared_prefix(before, after) > 0.85
def test_without_a_stable_floor_the_prefix_collapses(saturated):
"""The control, and the behaviour this replaced.
Rows with no `depth` cannot be placed on the tree, so the floor cannot be
computed and the window takes whatever fits — sliding by one action every
turn. The new prompt then starts with a *different* action, the shared
prefix collapses, and the server reprocesses essentially the whole thing.
That is the 343s case in this module's docstring.
"""
db, adventure, settings = saturated
for action in db.query(models.Action).all():
action.depth = None
db.commit()
db.expire_all()
_, before, report = _builder.build_context(adventure, settings)
assert report["history"]["floor_depth"] is None
_play_one_more(db, adventure, None)
_, after, _ = _builder.build_context(adventure, settings)
assert _shared_prefix(before, after) < 0.1
def test_the_window_does_step_eventually(saturated):
"""It holds still, but it must not hold still for ever — the budget is a
bound, and a window that never moved would break it."""
db, adventure, settings = saturated
first = _builder.build_context(adventure, settings)[2]["history"]["floor_depth"]
seen = {first}
for depth in range(60, 90):
_play_one_more(db, adventure, depth)
seen.add(_builder.build_context(adventure, settings)[2]["history"]["floor_depth"])
assert len(seen) > 1, "the floor never moved across 30 turns"
def test_the_prompt_stays_inside_the_budget_as_the_window_steps(saturated):
"""M03's bound, across the step. Trimming only ever drops more history."""
db, adventure, settings = saturated
for depth in range(60, 85):
_play_one_more(db, adventure, depth)
report = _builder.build_context(adventure, settings)[2]
assert report["tokens"]["total"] <= report["tokens"]["budget"]