Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
co-authored by
Claude Opus 5
parent
fedb7144d0
commit
ef25b0a876
@@ -177,19 +177,38 @@ async def list_endpoint_models(endpoint_url: str) -> dict:
|
||||
def _window_warning(window: contextwindow.Window, settings: models.Settings) -> str | None:
|
||||
"""What to tell the reader about the window, or None when nothing is wrong.
|
||||
|
||||
Three cases, and they need three different things done about them, so they
|
||||
say three different things (the same reasoning as the connection test's own
|
||||
Four cases, and they need four different things done about them, so they
|
||||
say four different things (the same reasoning as the connection test's own
|
||||
four failure kinds).
|
||||
"""
|
||||
budget = settings.context_token_budget
|
||||
if window.source == contextwindow.DECLARED:
|
||||
# Enforced, but on the operator's word rather than the server's. Worth
|
||||
# saying plainly: nothing here has checked the number, so a declaration
|
||||
# that is too large is the silent-truncation failure all over again.
|
||||
over = (
|
||||
" It is larger than the story budget, so it changes nothing today."
|
||||
if window.tokens >= budget else
|
||||
f" Prompts are being built to {window.tokens:,} rather than "
|
||||
f"{budget:,}."
|
||||
)
|
||||
return (
|
||||
f"The context window for '{settings.model}' is set in settings to "
|
||||
f"{window.tokens:,} tokens, because this server cannot be asked for it "
|
||||
f"— {window.detail}.{over} Nothing has verified that number against "
|
||||
"the server; if it is larger than the window the server really "
|
||||
"enforces, the oldest part of the prompt is still being dropped."
|
||||
)
|
||||
if not window.verified:
|
||||
return (
|
||||
f"The context window this server will give '{settings.model}' could not "
|
||||
f"be checked — {window.detail}. The story budget is {budget:,} tokens; "
|
||||
"if the server's window is smaller than that it silently drops the "
|
||||
"oldest part of the prompt, which here is the narrator's rules and the "
|
||||
"campaign canon. See DEVELOPMENT.md, 'The context window your Ollama "
|
||||
"actually enforces'."
|
||||
"campaign canon. If this server has no Ollama-native API to ask — "
|
||||
"vLLM, llama.cpp's own server — set the context window in settings so "
|
||||
"the prompt is capped to it. See DEVELOPMENT.md, 'The context window "
|
||||
"your Ollama actually enforces'."
|
||||
)
|
||||
if window.tokens < budget:
|
||||
ceiling = (
|
||||
@@ -226,7 +245,8 @@ async def test_connection(
|
||||
# an operator reloads a model. Changing the endpoint or the model clears
|
||||
# the cache (`update_settings`), which covers the case a reader can
|
||||
# actually cause; the detail line always says where the number came from.
|
||||
window = await contextwindow.probe(settings.endpoint_url, settings.model)
|
||||
window = await contextwindow.probe(settings.endpoint_url, settings.model,
|
||||
declared=settings.context_window_override)
|
||||
result = result | {"window": {
|
||||
"verified": window.verified,
|
||||
"tokens": window.tokens,
|
||||
|
||||
Reference in New Issue
Block a user