Stop re-reading the whole prompt every turn, and let a lost run carry on

M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
JesseMarkowitz
2026-09-10 06:13:55 -04:00
co-authored by Claude Opus 5
parent fedb7144d0
commit ef25b0a876
22 changed files with 1654 additions and 75 deletions
+128 -2
View File
@@ -35,6 +35,21 @@ AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
NPC_WINDOW = 6 # actions of story searched for NPC trigger words ("in scene")
SEPARATOR = "\n\n"
#: How much of the history window one trim gives up, as one-over-this. A
#: quarter: large enough that the window then holds still for several turns,
#: small enough that the narrator never loses most of its recent history at once.
#:
#: **This is the dial.** Lower it for bigger blocks — fewer prompt re-reads and
#: faster long campaigns, at the cost of retaining less recent history. Raise it
#: for the reverse. Nothing else has to change: `trim_block` is the only reader,
#: and `test_trim_fraction_is_the_dial_between_history_and_speed` pins that.
#: Measured at 4, on an 8,192-token budget: 124.0s per turn against 362.4s with
#: trimming off.
TRIM_FRACTION = 4
#: Never trim less than this, or the window slides by one action again and the
#: whole point is lost.
MIN_TRIM_BLOCK = 2
# Output-length guidance. The endpoint enforces `max_output_tokens` as a hard
# limit, and it truncates the reply mid-sentence when the model reaches it. The
# state block is emitted last, so truncation removes it. Asking the model to
@@ -301,6 +316,85 @@ def _canon_section(adventure: models.Adventure) -> str:
return f"Campaign canon (these are true and may not be contradicted):\n{body}"
def trim_block(history_budget: int, max_output_tokens: int) -> int:
"""How many `depth` steps of history one trim gives up.
Derived from **configuration**, never from the story, because the answer has
to be the same on two consecutive turns. A block size that moved with the
measured size of recent actions would move the boundary it defines, and a
boundary that moves is precisely what this exists to stop.
An AI action is bounded by `max_output_tokens` and a player action is small
beside it, so `max_output_tokens` is the scale of one row of history — a
setting, rather than a guess about the data.
"""
per_action = max(1, max_output_tokens)
fits = max(1, history_budget // per_action)
return max(MIN_TRIM_BLOCK, fits // TRIM_FRACTION)
def history_floor(depths: list[int | None], costs: list[int], budget: int,
block: int) -> int | None:
"""The depth of the oldest action to include, snapped to a block boundary.
## Why this is not just "whatever fits"
Taking whatever fits is what the builder did, and it is correct. It is also
the reason a long campaign costs a full prompt re-read every turn.
Inference servers cache the prompt they have already processed, keyed on the
**prefix**. While the story only grows at the end, each turn re-uses that
cache and pays for its own new tokens alone. As soon as the budget is full,
"whatever fits" drops the *oldest* action every turn — a change near the
front of the prompt — and everything after it has to be processed again.
So the floor is snapped forward to a multiple of `block` and then held. It
moves in steps: several cheap turns that re-use the cache, then one turn that
pays to re-read, rather than every turn paying. The cost is history depth —
right after a step the window holds up to `block` actions fewer than the
budget would allow, which is what `TRIM_FRACTION` bounds.
Measured against the reference deployment, on prompts this builder produced,
at an 8,192 budget where `block` is 3:
floor held, story grew by one action 14-20 s
floor stepped, prompt re-read 333-338 s
mean over two whole cycles 124.0 s
floor disabled, every turn re-read 362.4 s (361, 361, 365, 361)
2.9x, and the shape is the point rather than the ratio: the saving grows with
`block`, which grows with the budget, so the configuration that hurt most
before benefits most now.
Returns None when nothing needs trimming, which covers two cases that must
both stay as they were: a story short enough to fit whole (the window is a
growing prefix already, and snapping would drop its opening for no reason),
and an action so large that not even the newest one fits, which the caller
truncates.
"""
if not depths or any(depth is None for depth in depths):
# Legacy rows, or a path this cannot place on the tree. Trimming needs a
# stable coordinate; without one, behave exactly as before.
return None
spent = 0
oldest_fitting: int | None = None
for depth, cost in zip(reversed(depths), reversed(costs)):
if spent + cost > budget:
break
spent += cost
oldest_fitting = depth
if oldest_fitting is None:
return None
if oldest_fitting == depths[0]:
# Everything offered fits. There is nothing to drop, and snapping here
# would throw away the start of a short story to no purpose.
return None
block = max(1, block)
return -(-oldest_fitting // block) * block
def _visible_npcs(actions: list[models.Action], stat_schema: dict) -> dict[str, str]:
"""Returns the NPCs whose trigger words appear in the recent story.
@@ -633,10 +727,31 @@ def build_context(
# ----- Story history: newest first until the remaining budget is spent -----
history_budget = available_after_knowledge - used
# Where the window starts, snapped to a block so it holds still for several
# turns instead of sliding by one action every turn. `history_floor` says
# why that matters and what it costs. None means trim nothing, and then
# everything below is exactly what it was before.
costs = [count_tokens(_history_text(a)) + count_tokens(SEPARATOR)
for a in actions]
block = trim_block(history_budget, settings.max_output_tokens)
floor_depth = history_floor([a.depth for a in actions], costs,
history_budget, block)
windowed = actions
if floor_depth is not None:
kept = [a for a in actions if a.depth is not None and a.depth >= floor_depth]
# A floor that leaves nothing is a floor worth ignoring: the loop below
# still has to produce a turn, and its own truncation path is the honest
# way to handle a single action larger than the whole budget.
if kept:
windowed = kept
else:
floor_depth = None
included_actions: list[models.Action] = []
spent = 0
oldest_truncated = False
for action in reversed(actions):
for action in reversed(windowed):
# Budget against the text as it appears in the prompt, which includes
# the state block when this adventure tracks world state.
rendered = _history_text(action)
@@ -741,9 +856,13 @@ def build_context(
"source": (window.source if window is not None else contextwindow.UNKNOWN),
"model_max": (window.model_max if window is not None else None),
"detail": (window.detail if window is not None else "not checked"),
# `enforceable`, not `verified`: an operator-declared window caps
# the prompt exactly as a server-reported one does, and a turn built
# against it *was* capped. `verified` and `source` above still say
# which kind of answer produced the number.
"capped": (
window is not None
and window.verified
and window.enforceable
and window.tokens < settings.context_token_budget
),
},
@@ -775,6 +894,13 @@ def build_context(
# included, so this number must be the real total.
"total": history.count(adventure, exclude_action_id),
"oldest_truncated": oldest_truncated,
# Where the window was cut, and how big a step it takes when it
# moves. Both are in `depth` units. `floor_depth` is null while the
# story still fits whole, which is also while every turn is a pure
# prefix extension of the last one. A reader comparing two turns can
# tell from these whether the prompt's prefix was preserved.
"floor_depth": floor_depth,
"trim_block": block,
},
"settings": {
"model": settings.model,
+90 -10
View File
@@ -56,6 +56,31 @@ It does not hard-code 4,096, which would cripple a correctly configured
deployment; it does not raise the budget, which is the operator's decision; it
does not fall back to a cloud probe, a bundled table of model sizes, or a guess
from the model's name. An unknown window is reported as unknown.
## The server that cannot be asked
Discovery above is Ollama's native API. Nothing restricts `endpoint_url` to
Ollama — any allowed address serving an OpenAI-compatible `/v1` is accepted —
and on vLLM, llama.cpp's own server, or anything else, `/api/ps` and `/api/show`
are simply not there. Discovery then fails exactly as designed and the window is
reported unknown, which is honest but leaves the invariant at the top of this
file unenforced: the budget stands at whatever is configured, and if that server
enforces a smaller window it drops the oldest tokens again.
`context_window_override` is the operator's answer to that. It is a number the
operator states because they know how the server was launched, and it is used
**only when the server could not be asked**:
verified window -> always wins; a declaration cannot raise it
no verified window -> the declaration becomes the ceiling, source DECLARED
neither -> unknown, exactly as before
This does not weaken what `verified` claims. `verified` still means the server
itself answered, so `window_verified` in a turn's provenance keeps the meaning
the M11 report gives it, and a declared window is identifiable as a declaration
wherever it appears. What the declaration buys is enforcement: the prompt is
capped, so the failure mode is a shorter prompt rather than a silently truncated
one.
"""
from __future__ import annotations
@@ -86,8 +111,12 @@ NEGATIVE_TTL = 60.0
#: Sources, in the order of how much they prove.
LOADED = "loaded" # /api/ps: what the runtime is enforcing now
PARAMETERS = "parameters" # /api/show: what the model will load with
DECLARED = "declared" # the operator said so; the server could not be asked
UNKNOWN = "unknown"
#: Sources that mean *the server answered*, as opposed to somebody asserting.
FROM_SERVER = (LOADED, PARAMETERS)
@dataclass(frozen=True)
class Window:
@@ -106,6 +135,18 @@ class Window:
@property
def verified(self) -> bool:
"""The **server** answered. An operator's declaration is not this.
Kept narrow on purpose. `window_verified` travels in every turn's stored
provenance and the M11 report counts on it meaning one thing: that the
runtime was asked and replied. A declaration is a person's claim about a
server, which is worth acting on and is not the same evidence.
"""
return self.tokens is not None and self.source in FROM_SERVER
@property
def enforceable(self) -> bool:
"""There is a number to cap the prompt to, whoever supplied it."""
return self.tokens is not None
@@ -128,9 +169,10 @@ def native_base(endpoint_url: str) -> str:
def effective_budget(configured: int, window: Window | int | None) -> int:
"""The budget the prompt may actually use.
The whole enforcement, in one line: a verified window is a ceiling. The
configured budget still wins when it is *smaller*, because a reader who has
deliberately asked for a shorter prompt should get one.
The whole enforcement, in one line: a known window is a ceiling — whether
the server reported it or the operator declared it. The configured budget
still wins when it is *smaller*, because a reader who has deliberately asked
for a shorter prompt should get one.
"""
tokens = window.tokens if isinstance(window, Window) else window
if tokens is None or tokens <= 0:
@@ -143,17 +185,55 @@ def cache_clear() -> None:
_cache.clear()
async def probe(endpoint_url: str, model: str, *, use_cache: bool = True) -> Window:
"""Asks the server what window `model` gets. Never raises.
async def probe(endpoint_url: str, model: str, *,
declared: int | None = None, use_cache: bool = True) -> Window:
"""What window `model` gets, asked of the server and only then declared.
Returns `UNVERIFIED` for every failure — refused endpoint, unreachable
server, TLS failure, a non-Ollama endpoint, an unparseable answer. The caller
cannot act differently on those and the reader is told the same thing either
way: the window could not be checked.
Returns `UNVERIFIED` for every discovery failure — refused endpoint,
unreachable server, TLS failure, a server with no Ollama-native API, an
unparseable answer — unless `declared` supplies a number to fall back on.
The caller cannot act differently on those failures and the reader is told
the same thing either way: the window could not be checked.
`declared` is `Settings.context_window_override`. It never overrides a
verified answer, so an operator cannot talk the application into a bigger
prompt than the runtime will read; it only fills a gap discovery left.
"""
if not endpoint_url or not model:
return Window(None, UNKNOWN, detail="no endpoint or model configured")
return _declared_or(declared,
Window(None, UNKNOWN,
detail="no endpoint or model configured"))
discovered = await _discover(endpoint_url, model, use_cache=use_cache)
return _declared_or(declared, discovered)
def _declared_or(declared: int | None, discovered: Window) -> Window:
"""The operator's number, but only where the server left a hole.
A verified window always wins. That ordering is the whole safety property:
a declaration can lower an unknown ceiling into existence, never raise a
known one.
"""
if discovered.verified:
return discovered
if not declared or declared <= 0:
return discovered
return Window(
declared, DECLARED, discovered.model_max,
f"{declared:,} tokens, declared in settings — the server was not able "
f"to say ({discovered.detail})",
)
async def _discover(endpoint_url: str, model: str, *,
use_cache: bool = True) -> Window:
"""The server's own answer, cached. Knows nothing about declarations.
The cache holds only what was discovered, so changing the declared override
takes effect on the next turn without having to clear anything: the
declaration is layered on afterwards, in `_declared_or`.
"""
key = (endpoint_url, model)
now = time.monotonic()
if use_cache:
+4
View File
@@ -475,6 +475,10 @@ MIGRATIONS: list[tuple[int, str | dict[str, str]]] = [
# an existing campaign's prompts do not change under it.
(93, "ALTER TABLE adventures ADD COLUMN narration_length VARCHAR(20) "
"NOT NULL DEFAULT ''"),
# Nullable, and null by default: an override that defaulted to a number
# would be the application guessing at a window again, which is the one
# thing `contextwindow` refuses to do. Null means "nobody has said".
(94, "ALTER TABLE settings ADD COLUMN context_window_override INTEGER"),
]
LATEST_VERSION = max((v for v, _ in MIGRATIONS), default=1)
+8
View File
@@ -1224,6 +1224,14 @@ class Settings(Base):
# while the same turn takes seconds once the model is resident. See
# `providers.openai_compatible.DEFAULT_READ_TIMEOUT`.
model_timeout_seconds: Mapped[int] = mapped_column(Integer, default=300)
# What window the inference server enforces, when the server cannot be asked
# for it. Discovery (`contextwindow`) speaks Ollama's native API; a server
# that does not serve one — vLLM, llama.cpp's own server — leaves the window
# unknown and the budget uncapped. This is the operator saying how they
# launched it. It never overrides a window the server did report, and null
# means nobody has said, because a default here would be a guess.
context_window_override: Mapped[int | None] = mapped_column(
Integer, nullable=True, default=None)
narrator_prompt: Mapped[str] = mapped_column(
Text,
default=(
+2 -1
View File
@@ -33,7 +33,8 @@ async def dry_run_context(
# M11: and by the same probe the turn makes, for the same reason — a panel
# that showed a 16,384-token budget while the next turn will be capped to
# 4,096 would be showing a prompt that is not the one about to be sent.
window = await contextwindow.probe(settings.endpoint_url, settings.model)
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
try:
_, _, report = build_context(
adventure, settings, memories, knowledge=knowledge, window=window
+2 -1
View File
@@ -208,7 +208,8 @@ async def _generate_turn(
# network calls — and cached per endpoint and model, so it costs one short
# request per session rather than one per turn. An unverified window does
# not block the turn; it is recorded as unverified in the snapshot below.
window = await contextwindow.probe(settings.endpoint_url, settings.model)
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
try:
system_text, story_text, snapshot = build_context(
adventure,
+25 -5
View File
@@ -177,19 +177,38 @@ async def list_endpoint_models(endpoint_url: str) -> dict:
def _window_warning(window: contextwindow.Window, settings: models.Settings) -> str | None:
"""What to tell the reader about the window, or None when nothing is wrong.
Three cases, and they need three different things done about them, so they
say three different things (the same reasoning as the connection test's own
Four cases, and they need four different things done about them, so they
say four different things (the same reasoning as the connection test's own
four failure kinds).
"""
budget = settings.context_token_budget
if window.source == contextwindow.DECLARED:
# Enforced, but on the operator's word rather than the server's. Worth
# saying plainly: nothing here has checked the number, so a declaration
# that is too large is the silent-truncation failure all over again.
over = (
" It is larger than the story budget, so it changes nothing today."
if window.tokens >= budget else
f" Prompts are being built to {window.tokens:,} rather than "
f"{budget:,}."
)
return (
f"The context window for '{settings.model}' is set in settings to "
f"{window.tokens:,} tokens, because this server cannot be asked for it "
f"— {window.detail}.{over} Nothing has verified that number against "
"the server; if it is larger than the window the server really "
"enforces, the oldest part of the prompt is still being dropped."
)
if not window.verified:
return (
f"The context window this server will give '{settings.model}' could not "
f"be checked — {window.detail}. The story budget is {budget:,} tokens; "
"if the server's window is smaller than that it silently drops the "
"oldest part of the prompt, which here is the narrator's rules and the "
"campaign canon. See DEVELOPMENT.md, 'The context window your Ollama "
"actually enforces'."
"campaign canon. If this server has no Ollama-native API to ask — "
"vLLM, llama.cpp's own server — set the context window in settings so "
"the prompt is capped to it. See DEVELOPMENT.md, 'The context window "
"your Ollama actually enforces'."
)
if window.tokens < budget:
ceiling = (
@@ -226,7 +245,8 @@ async def test_connection(
# an operator reloads a model. Changing the endpoint or the model clears
# the cache (`update_settings`), which covers the case a reader can
# actually cause; the detail line always says where the number came from.
window = await contextwindow.probe(settings.endpoint_url, settings.model)
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
result = result | {"window": {
"verified": window.verified,
"tokens": window.tokens,
+6
View File
@@ -687,6 +687,7 @@ class SettingsOut(ORMModel):
max_output_tokens: int
context_token_budget: int
model_timeout_seconds: int
context_window_override: int | None
narrator_prompt: str
summary_model: str
embedding_model: str
@@ -730,6 +731,11 @@ class SettingsUpdate(BaseModel):
# turn cannot trip it; the ceiling exists so that "wait longer" stays a
# number rather than becoming "wait forever".
model_timeout_seconds: Annotated[int, Field(ge=30, le=3600)] | None = None
# The window an inference server enforces, for servers that cannot be asked.
# Bounded like the budget it caps. It is never a way to *raise* the prompt
# past a window the server did report — `contextwindow._declared_or` — so
# the ceiling here only bounds what an operator can usefully claim.
context_window_override: Annotated[int, Field(ge=256, le=200_000)] | None = None
narrator_prompt: Prose | None = None
summary_model: Name | None = None
embedding_model: Name | None = None