Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
co-authored by
Claude Opus 5
parent
fedb7144d0
commit
ef25b0a876
@@ -35,6 +35,21 @@ AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
|
||||
NPC_WINDOW = 6 # actions of story searched for NPC trigger words ("in scene")
|
||||
SEPARATOR = "\n\n"
|
||||
|
||||
#: How much of the history window one trim gives up, as one-over-this. A
|
||||
#: quarter: large enough that the window then holds still for several turns,
|
||||
#: small enough that the narrator never loses most of its recent history at once.
|
||||
#:
|
||||
#: **This is the dial.** Lower it for bigger blocks — fewer prompt re-reads and
|
||||
#: faster long campaigns, at the cost of retaining less recent history. Raise it
|
||||
#: for the reverse. Nothing else has to change: `trim_block` is the only reader,
|
||||
#: and `test_trim_fraction_is_the_dial_between_history_and_speed` pins that.
|
||||
#: Measured at 4, on an 8,192-token budget: 124.0s per turn against 362.4s with
|
||||
#: trimming off.
|
||||
TRIM_FRACTION = 4
|
||||
#: Never trim less than this, or the window slides by one action again and the
|
||||
#: whole point is lost.
|
||||
MIN_TRIM_BLOCK = 2
|
||||
|
||||
# Output-length guidance. The endpoint enforces `max_output_tokens` as a hard
|
||||
# limit, and it truncates the reply mid-sentence when the model reaches it. The
|
||||
# state block is emitted last, so truncation removes it. Asking the model to
|
||||
@@ -301,6 +316,85 @@ def _canon_section(adventure: models.Adventure) -> str:
|
||||
return f"Campaign canon (these are true and may not be contradicted):\n{body}"
|
||||
|
||||
|
||||
def trim_block(history_budget: int, max_output_tokens: int) -> int:
|
||||
"""How many `depth` steps of history one trim gives up.
|
||||
|
||||
Derived from **configuration**, never from the story, because the answer has
|
||||
to be the same on two consecutive turns. A block size that moved with the
|
||||
measured size of recent actions would move the boundary it defines, and a
|
||||
boundary that moves is precisely what this exists to stop.
|
||||
|
||||
An AI action is bounded by `max_output_tokens` and a player action is small
|
||||
beside it, so `max_output_tokens` is the scale of one row of history — a
|
||||
setting, rather than a guess about the data.
|
||||
"""
|
||||
per_action = max(1, max_output_tokens)
|
||||
fits = max(1, history_budget // per_action)
|
||||
return max(MIN_TRIM_BLOCK, fits // TRIM_FRACTION)
|
||||
|
||||
|
||||
def history_floor(depths: list[int | None], costs: list[int], budget: int,
|
||||
block: int) -> int | None:
|
||||
"""The depth of the oldest action to include, snapped to a block boundary.
|
||||
|
||||
## Why this is not just "whatever fits"
|
||||
|
||||
Taking whatever fits is what the builder did, and it is correct. It is also
|
||||
the reason a long campaign costs a full prompt re-read every turn.
|
||||
|
||||
Inference servers cache the prompt they have already processed, keyed on the
|
||||
**prefix**. While the story only grows at the end, each turn re-uses that
|
||||
cache and pays for its own new tokens alone. As soon as the budget is full,
|
||||
"whatever fits" drops the *oldest* action every turn — a change near the
|
||||
front of the prompt — and everything after it has to be processed again.
|
||||
|
||||
So the floor is snapped forward to a multiple of `block` and then held. It
|
||||
moves in steps: several cheap turns that re-use the cache, then one turn that
|
||||
pays to re-read, rather than every turn paying. The cost is history depth —
|
||||
right after a step the window holds up to `block` actions fewer than the
|
||||
budget would allow, which is what `TRIM_FRACTION` bounds.
|
||||
|
||||
Measured against the reference deployment, on prompts this builder produced,
|
||||
at an 8,192 budget where `block` is 3:
|
||||
|
||||
floor held, story grew by one action 14-20 s
|
||||
floor stepped, prompt re-read 333-338 s
|
||||
mean over two whole cycles 124.0 s
|
||||
floor disabled, every turn re-read 362.4 s (361, 361, 365, 361)
|
||||
|
||||
2.9x, and the shape is the point rather than the ratio: the saving grows with
|
||||
`block`, which grows with the budget, so the configuration that hurt most
|
||||
before benefits most now.
|
||||
|
||||
Returns None when nothing needs trimming, which covers two cases that must
|
||||
both stay as they were: a story short enough to fit whole (the window is a
|
||||
growing prefix already, and snapping would drop its opening for no reason),
|
||||
and an action so large that not even the newest one fits, which the caller
|
||||
truncates.
|
||||
"""
|
||||
if not depths or any(depth is None for depth in depths):
|
||||
# Legacy rows, or a path this cannot place on the tree. Trimming needs a
|
||||
# stable coordinate; without one, behave exactly as before.
|
||||
return None
|
||||
|
||||
spent = 0
|
||||
oldest_fitting: int | None = None
|
||||
for depth, cost in zip(reversed(depths), reversed(costs)):
|
||||
if spent + cost > budget:
|
||||
break
|
||||
spent += cost
|
||||
oldest_fitting = depth
|
||||
if oldest_fitting is None:
|
||||
return None
|
||||
if oldest_fitting == depths[0]:
|
||||
# Everything offered fits. There is nothing to drop, and snapping here
|
||||
# would throw away the start of a short story to no purpose.
|
||||
return None
|
||||
|
||||
block = max(1, block)
|
||||
return -(-oldest_fitting // block) * block
|
||||
|
||||
|
||||
def _visible_npcs(actions: list[models.Action], stat_schema: dict) -> dict[str, str]:
|
||||
"""Returns the NPCs whose trigger words appear in the recent story.
|
||||
|
||||
@@ -633,10 +727,31 @@ def build_context(
|
||||
|
||||
# ----- Story history: newest first until the remaining budget is spent -----
|
||||
history_budget = available_after_knowledge - used
|
||||
|
||||
# Where the window starts, snapped to a block so it holds still for several
|
||||
# turns instead of sliding by one action every turn. `history_floor` says
|
||||
# why that matters and what it costs. None means trim nothing, and then
|
||||
# everything below is exactly what it was before.
|
||||
costs = [count_tokens(_history_text(a)) + count_tokens(SEPARATOR)
|
||||
for a in actions]
|
||||
block = trim_block(history_budget, settings.max_output_tokens)
|
||||
floor_depth = history_floor([a.depth for a in actions], costs,
|
||||
history_budget, block)
|
||||
windowed = actions
|
||||
if floor_depth is not None:
|
||||
kept = [a for a in actions if a.depth is not None and a.depth >= floor_depth]
|
||||
# A floor that leaves nothing is a floor worth ignoring: the loop below
|
||||
# still has to produce a turn, and its own truncation path is the honest
|
||||
# way to handle a single action larger than the whole budget.
|
||||
if kept:
|
||||
windowed = kept
|
||||
else:
|
||||
floor_depth = None
|
||||
|
||||
included_actions: list[models.Action] = []
|
||||
spent = 0
|
||||
oldest_truncated = False
|
||||
for action in reversed(actions):
|
||||
for action in reversed(windowed):
|
||||
# Budget against the text as it appears in the prompt, which includes
|
||||
# the state block when this adventure tracks world state.
|
||||
rendered = _history_text(action)
|
||||
@@ -741,9 +856,13 @@ def build_context(
|
||||
"source": (window.source if window is not None else contextwindow.UNKNOWN),
|
||||
"model_max": (window.model_max if window is not None else None),
|
||||
"detail": (window.detail if window is not None else "not checked"),
|
||||
# `enforceable`, not `verified`: an operator-declared window caps
|
||||
# the prompt exactly as a server-reported one does, and a turn built
|
||||
# against it *was* capped. `verified` and `source` above still say
|
||||
# which kind of answer produced the number.
|
||||
"capped": (
|
||||
window is not None
|
||||
and window.verified
|
||||
and window.enforceable
|
||||
and window.tokens < settings.context_token_budget
|
||||
),
|
||||
},
|
||||
@@ -775,6 +894,13 @@ def build_context(
|
||||
# included, so this number must be the real total.
|
||||
"total": history.count(adventure, exclude_action_id),
|
||||
"oldest_truncated": oldest_truncated,
|
||||
# Where the window was cut, and how big a step it takes when it
|
||||
# moves. Both are in `depth` units. `floor_depth` is null while the
|
||||
# story still fits whole, which is also while every turn is a pure
|
||||
# prefix extension of the last one. A reader comparing two turns can
|
||||
# tell from these whether the prompt's prefix was preserved.
|
||||
"floor_depth": floor_depth,
|
||||
"trim_block": block,
|
||||
},
|
||||
"settings": {
|
||||
"model": settings.model,
|
||||
|
||||
Reference in New Issue
Block a user