WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
316 lines
13 KiB
Python
316 lines
13 KiB
Python
"""The history window moves in blocks, so the prompt's prefix holds still.
|
|
|
|
Inference servers cache a prompt by its **prefix**. While a story only grows at
|
|
the end, every turn re-uses that cache and pays for its own new tokens alone. The
|
|
builder's old window took whatever fit, which meant that once the budget was full
|
|
it dropped the *oldest* action every turn — a change near the front of the prompt
|
|
— and everything after it had to be processed again.
|
|
|
|
Measured on the reference deployment, at 7.7k prompt tokens against a 3B model:
|
|
|
|
window slid by one turn 343-350 s
|
|
prefix preserved 6.0 s
|
|
|
|
These tests do not measure time. They pin the property the measurement is
|
|
downstream of: **the oldest included action is the same across consecutive
|
|
turns**, except on the turns where the window deliberately steps.
|
|
"""
|
|
import pytest
|
|
import pytest as _pytest
|
|
|
|
from app.context import builder
|
|
|
|
|
|
def costs_of(n, each=100):
|
|
return [each] * n
|
|
|
|
|
|
def depths(n, start=1):
|
|
return list(range(start, start + n))
|
|
|
|
|
|
# ------------------------------------------------------------- the block size
|
|
|
|
def test_the_block_is_a_share_of_what_fits():
|
|
"""Derived from `TRIM_FRACTION` rather than asserting the number it is
|
|
currently set to, so retuning the dial does not fail a test that was never
|
|
about the dial's value."""
|
|
# 16 actions of 500 fit in 8,000, and the block is that share of them.
|
|
assert builder.trim_block(8000, 500) == 16 // builder.TRIM_FRACTION
|
|
|
|
|
|
def test_trim_fraction_is_the_dial_between_history_and_speed():
|
|
"""`TRIM_FRACTION` is meant to be retuned, so this pins what retuning does.
|
|
|
|
Lower it and the window gives up more at once: bigger blocks, fewer re-reads,
|
|
less recent history retained. Raise it and the reverse. Nothing else in the
|
|
builder has to change for that to hold, which is the property worth having a
|
|
test for.
|
|
"""
|
|
budget, per_action = 8000, 500
|
|
fits = budget // per_action
|
|
|
|
def block_at(fraction, monkeypatch):
|
|
monkeypatch.setattr(builder, "TRIM_FRACTION", fraction)
|
|
return builder.trim_block(budget, per_action)
|
|
|
|
with _pytest.MonkeyPatch.context() as mp:
|
|
greedier = block_at(2, mp)
|
|
assert greedier == fits // 2
|
|
with _pytest.MonkeyPatch.context() as mp:
|
|
gentler = block_at(8, mp)
|
|
assert gentler == max(builder.MIN_TRIM_BLOCK, fits // 8)
|
|
assert greedier > gentler, "a lower fraction must give up more at once"
|
|
|
|
# And the floor to the whole thing survives any setting.
|
|
with _pytest.MonkeyPatch.context() as mp:
|
|
mp.setattr(builder, "TRIM_FRACTION", 1000)
|
|
assert builder.trim_block(budget, per_action) >= builder.MIN_TRIM_BLOCK
|
|
|
|
|
|
def test_the_block_never_slides_by_one():
|
|
"""A block of one is the old behaviour wearing a hat."""
|
|
assert builder.trim_block(100, 500) >= builder.MIN_TRIM_BLOCK
|
|
assert builder.trim_block(0, 500) >= builder.MIN_TRIM_BLOCK
|
|
|
|
|
|
def test_the_block_comes_from_settings_not_from_the_story():
|
|
"""It has to be the same on two consecutive turns, so it cannot be measured
|
|
from actions whose sizes vary."""
|
|
assert builder.trim_block(8000, 500) == builder.trim_block(8000, 500)
|
|
# Bigger budget, bigger step; the ratio is what is fixed.
|
|
assert builder.trim_block(16000, 500) > builder.trim_block(8000, 500)
|
|
|
|
|
|
# ------------------------------------------------------- nothing to trim yet
|
|
|
|
def test_a_story_that_fits_whole_is_not_trimmed():
|
|
"""Also the append-only regime: every turn is a prefix extension already."""
|
|
assert builder.history_floor(depths(5), costs_of(5), budget=10_000, block=4) is None
|
|
|
|
|
|
def test_a_short_story_keeps_its_opening():
|
|
"""Snapping here would drop the start of the story for no reason at all."""
|
|
assert builder.history_floor(depths(3), costs_of(3), budget=10_000, block=8) is None
|
|
|
|
|
|
def test_an_action_larger_than_the_budget_is_left_to_the_caller():
|
|
assert builder.history_floor([1], [5000], budget=100, block=4) is None
|
|
|
|
|
|
def test_rows_without_a_depth_are_not_trimmed():
|
|
"""Legacy rows have no stable coordinate, so behave exactly as before."""
|
|
assert builder.history_floor([None, None], costs_of(2), 100, 4) is None
|
|
assert builder.history_floor([], [], 100, 4) is None
|
|
|
|
|
|
# ------------------------------------------------------------ the whole point
|
|
|
|
def test_the_floor_holds_still_while_the_story_grows():
|
|
"""The property the 57x measurement rests on.
|
|
|
|
Ten consecutive turns against a full budget. The floor must take a small
|
|
number of steps, not ten.
|
|
"""
|
|
block, budget, each = 4, 1000, 100 # 10 actions fit
|
|
seen = []
|
|
for extra in range(10): # the story grows by one action
|
|
n = 20 + extra
|
|
seen.append(builder.history_floor(depths(n), costs_of(n, each), budget, block))
|
|
|
|
steps = sum(1 for a, b in zip(seen, seen[1:]) if a != b)
|
|
assert steps <= 3, f"the floor moved {steps} times in 10 turns: {seen}"
|
|
assert len(set(seen)) > 1, "it never moved at all, so the budget is not binding"
|
|
|
|
|
|
def test_every_floor_sits_on_a_block_boundary():
|
|
block, budget = 4, 1000
|
|
for n in range(20, 40):
|
|
floor = builder.history_floor(depths(n), costs_of(n), budget, block)
|
|
assert floor is not None
|
|
assert floor % block == 0, f"{floor} is not a multiple of {block}"
|
|
|
|
|
|
def test_the_floor_only_ever_moves_forward():
|
|
block, budget = 4, 1000
|
|
floors = [builder.history_floor(depths(n), costs_of(n), budget, block)
|
|
for n in range(20, 45)]
|
|
assert floors == sorted(floors)
|
|
|
|
|
|
# --------------------------------------------------- and still inside budget
|
|
|
|
@pytest.mark.parametrize("n", range(20, 40))
|
|
def test_the_kept_window_never_exceeds_the_budget(n):
|
|
"""M03's bound is not weakened. Trimming only ever drops more, never less."""
|
|
block, budget, each = 4, 1000, 100
|
|
ds, cs = depths(n), costs_of(n, each)
|
|
floor = builder.history_floor(ds, cs, budget, block)
|
|
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
|
|
assert kept <= budget
|
|
|
|
|
|
@pytest.mark.parametrize("n", range(20, 40))
|
|
def test_the_kept_window_is_not_gutted(n):
|
|
"""The cost of holding still is bounded: a trim gives up `block` actions,
|
|
never most of the window."""
|
|
block, budget, each = 4, 1000, 100
|
|
ds, cs = depths(n), costs_of(n, each)
|
|
floor = builder.history_floor(ds, cs, budget, block)
|
|
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
|
|
assert kept >= budget - block * each
|
|
|
|
|
|
# ------------------------------------------- the real builder, end to end
|
|
|
|
import pytest as _pytest # noqa: E402 (grouped with the fixtures it serves)
|
|
|
|
from app import models # noqa: E402
|
|
from app.context import builder as _builder # noqa: E402
|
|
from app.database import Base, SessionLocal, engine # noqa: E402
|
|
|
|
NARRATION = ("The rain came down over Westhaven in long grey sheets and the "
|
|
"gutters ran full from the ridge to the waterfront. ") * 6
|
|
|
|
|
|
@_pytest.fixture()
|
|
def saturated():
|
|
"""A campaign whose history is longer than its budget, with real depths.
|
|
|
|
`depth` is what the floor is expressed in, and every action written through
|
|
the application has one (`tree.place_action`). The older fixtures in
|
|
`test_history_window.py` predate the tree and leave it null, which is why
|
|
trimming does not engage there and those tests still describe the old
|
|
behaviour exactly.
|
|
"""
|
|
Base.metadata.create_all(bind=engine)
|
|
db = SessionLocal()
|
|
user = models.User(is_guest=False, email="blocktrim@example.com")
|
|
db.add(user)
|
|
db.flush()
|
|
settings = models.Settings(user_id=user.id, model="m", embedding_model="",
|
|
context_token_budget=2048, max_output_tokens=200)
|
|
db.add(settings)
|
|
adventure = models.Adventure(user_id=user.id, title="Long", script_state={})
|
|
db.add(adventure)
|
|
db.flush()
|
|
for i in range(60):
|
|
db.add(models.Action(adventure_id=adventure.id,
|
|
type="ai" if i % 2 else "do",
|
|
text=f"[{i}] {NARRATION}", branch_id=None, depth=i))
|
|
db.commit()
|
|
db.expire_all()
|
|
adventure = db.get(models.Adventure, adventure.id)
|
|
settings = db.get(models.Settings, settings.id)
|
|
try:
|
|
yield db, adventure, settings
|
|
finally:
|
|
db.close()
|
|
Base.metadata.drop_all(bind=engine)
|
|
|
|
|
|
def _play_one_more(db, adventure, at_depth):
|
|
db.add(models.Action(adventure_id=adventure.id, type="do",
|
|
text=f"[{at_depth}] {NARRATION}", depth=at_depth))
|
|
# `at_depth` may be None: the control below plays a turn into a story whose
|
|
# rows predate the tree, which is the ungoverned window this replaced.
|
|
db.commit()
|
|
db.expire_all()
|
|
|
|
|
|
def _shared_prefix(before: str, after: str) -> float:
|
|
"""How much of the old prompt the new one still opens with, 0.0 to 1.0.
|
|
|
|
This is the quantity the inference server's cache is keyed on, so it is the
|
|
quantity worth asserting. It is not 1.0 even in the best case: the prompt
|
|
ends with the turn's length-hint and state-block instructions, which sit
|
|
*after* the history, so appending a turn always rewrites that tail.
|
|
"""
|
|
shared = 0
|
|
for x, y in zip(before, after):
|
|
if x != y:
|
|
break
|
|
shared += 1
|
|
return shared / max(1, len(before))
|
|
|
|
|
|
def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
|
|
"""The property the whole change exists for.
|
|
|
|
Not a timing test — it asserts what the timing follows from. The story text
|
|
a turn sends still opens with almost all of what the previous turn sent, so
|
|
the server's prompt cache covers that part and only the tail is processed.
|
|
"""
|
|
db, adventure, settings = saturated
|
|
_, before, report_before = _builder.build_context(adventure, settings)
|
|
assert report_before["history"]["floor_depth"] is not None, (
|
|
"this fixture is meant to be over budget; trimming never engaged")
|
|
|
|
# v1.1 WP-A1: the fixture used to be positioned so that the very next turn
|
|
# held the floor. The safety reserve takes 256 tokens of this 2,048 budget,
|
|
# the block is now the minimum of two, and the next turn is a step. So walk
|
|
# forward until a turn holds, requiring every move on the way to be exactly
|
|
# one block: a window that slides by one action every turn fails either way.
|
|
held = None
|
|
depth = 60
|
|
for _ in range(4):
|
|
_play_one_more(db, adventure, depth)
|
|
depth += 1
|
|
_, after, report_after = _builder.build_context(adventure, settings)
|
|
floor_before = report_before["history"]["floor_depth"]
|
|
floor_after = report_after["history"]["floor_depth"]
|
|
block = report_after["history"]["trim_block"]
|
|
assert floor_after - floor_before in (0, block), (floor_before, floor_after, block)
|
|
if floor_after == floor_before:
|
|
held = (before, after)
|
|
break
|
|
before, report_before = after, report_after
|
|
|
|
assert held is not None, "the floor never held across a turn"
|
|
assert _shared_prefix(*held) > 0.85
|
|
|
|
|
|
def test_without_a_stable_floor_the_prefix_collapses(saturated):
|
|
"""The control, and the behaviour this replaced.
|
|
|
|
Rows with no `depth` cannot be placed on the tree, so the floor cannot be
|
|
computed and the window takes whatever fits — sliding by one action every
|
|
turn. The new prompt then starts with a *different* action, the shared
|
|
prefix collapses, and the server reprocesses essentially the whole thing.
|
|
That is the 343s case in this module's docstring.
|
|
"""
|
|
db, adventure, settings = saturated
|
|
for action in db.query(models.Action).all():
|
|
action.depth = None
|
|
db.commit()
|
|
db.expire_all()
|
|
|
|
_, before, report = _builder.build_context(adventure, settings)
|
|
assert report["history"]["floor_depth"] is None
|
|
_play_one_more(db, adventure, None)
|
|
_, after, _ = _builder.build_context(adventure, settings)
|
|
|
|
assert _shared_prefix(before, after) < 0.1
|
|
|
|
|
|
def test_the_window_does_step_eventually(saturated):
|
|
"""It holds still, but it must not hold still for ever — the budget is a
|
|
bound, and a window that never moved would break it."""
|
|
db, adventure, settings = saturated
|
|
first = _builder.build_context(adventure, settings)[2]["history"]["floor_depth"]
|
|
|
|
seen = {first}
|
|
for depth in range(60, 90):
|
|
_play_one_more(db, adventure, depth)
|
|
seen.add(_builder.build_context(adventure, settings)[2]["history"]["floor_depth"])
|
|
assert len(seen) > 1, "the floor never moved across 30 turns"
|
|
|
|
|
|
def test_the_prompt_stays_inside_the_budget_as_the_window_steps(saturated):
|
|
"""M03's bound, across the step. Trimming only ever drops more history."""
|
|
db, adventure, settings = saturated
|
|
for depth in range(60, 85):
|
|
_play_one_more(db, adventure, depth)
|
|
report = _builder.build_context(adventure, settings)[2]
|
|
assert report["tokens"]["total"] <= report["tokens"]["budget"]
|