Stop re-reading the whole prompt every turn, and let a lost run carry on

M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
JesseMarkowitz
2026-09-10 06:13:55 -04:00
co-authored by Claude Opus 5
parent fedb7144d0
commit ef25b0a876
22 changed files with 1654 additions and 75 deletions
+165
View File
@@ -0,0 +1,165 @@
"""The window an operator declares, for a server that cannot be asked for one.
`contextwindow`'s discovery speaks Ollama's native API. Nothing restricts
`endpoint_url` to Ollama, so on vLLM, llama.cpp's own server, or anything else
serving an OpenAI-compatible `/v1`, `/api/ps` and `/api/show` are not there:
discovery fails as designed, the window is unknown, and the budget is left
uncapped at whatever is configured. That is M11's own failure mode reached by a
different route — the server drops the oldest tokens, which here are the
narrator's rules and the campaign canon.
`Settings.context_window_override` closes it. These tests pin the two properties
that make it safe rather than merely useful:
1. it is used **only** where discovery left a hole, so it can never talk the
application into a longer prompt than a server actually reported, and
2. it does not make `verified` true, because `verified` means the server
answered and a declaration is a person's claim about a server.
"""
import asyncio
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, contextwindow, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
UNREACHABLE = "http://127.0.0.1:1/v1"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
@pytest.fixture()
def client():
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="m11dw@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model="some-model", endpoint_url=UNREACHABLE,
embedding_model="", context_token_budget=16384, max_output_tokens=800,
))
setup.commit()
user_id = user.id
setup.close()
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
try:
yield TestClient(app)
finally:
app.dependency_overrides.clear()
Base.metadata.drop_all(bind=engine)
def probe(endpoint=UNREACHABLE, model="some-model", declared=None):
contextwindow.cache_clear()
return asyncio.run(
contextwindow.probe(endpoint, model, declared=declared, use_cache=False))
# ------------------------------------------------------- filling the hole
def test_without_a_declaration_an_unaskable_server_leaves_the_window_unknown():
window = probe()
assert window.tokens is None
assert not window.verified
assert not window.enforceable
assert window.source == contextwindow.UNKNOWN
def test_a_declaration_becomes_the_ceiling_when_the_server_cannot_be_asked():
window = probe(declared=8192)
assert window.tokens == 8192
assert window.source == contextwindow.DECLARED
assert window.enforceable
assert contextwindow.effective_budget(16384, window) == 8192
def test_a_declaration_does_not_claim_the_server_was_verified():
"""`window_verified` travels in every turn's provenance and the M11 report
counts it. A declaration must not inflate that count."""
window = probe(declared=8192)
assert window.enforceable
assert not window.verified
def test_the_detail_says_the_number_came_from_settings():
assert "declared in settings" in probe(declared=8192).detail
@pytest.mark.parametrize("declared", [None, 0, -1])
def test_a_missing_or_meaningless_declaration_changes_nothing(declared):
window = probe(declared=declared)
assert window.tokens is None
assert window.source == contextwindow.UNKNOWN
def test_a_declaration_still_applies_when_nothing_is_configured():
window = asyncio.run(contextwindow.probe("", "", declared=4096))
assert window.tokens == 4096
assert window.source == contextwindow.DECLARED
def test_a_declaration_applies_to_a_refused_endpoint_without_reaching_it():
"""A refused address is a discovery failure like any other (ADR 011, H12).
The declaration caps the prompt; it does not make the endpoint usable, and
the turn is still refused where endpoints are enforced."""
window = probe(endpoint="http://169.254.169.254/v1", declared=4096)
assert window.source == contextwindow.DECLARED
assert window.tokens == 4096
# ------------------------------------------- a verified answer always wins
def test_a_verified_window_is_not_overridden(monkeypatch):
"""The safety property. An operator may lower an unknown ceiling into
existence; they may never raise one the server reported."""
async def reported(endpoint, model):
return contextwindow.Window(4096, contextwindow.LOADED, 32768, "real")
monkeypatch.setattr(contextwindow, "_ask", reported)
window = probe(declared=32768)
assert window.tokens == 4096
assert window.source == contextwindow.LOADED
assert window.verified
assert contextwindow.effective_budget(16384, window) == 4096
def test_a_declared_window_larger_than_the_budget_does_not_raise_it():
window = probe(declared=200_000)
assert contextwindow.effective_budget(16384, window) == 16384
# --------------------------------------------------------------- plumbing
def test_the_override_is_readable_and_settable_through_the_api(client):
assert client.get("/api/settings").json()["context_window_override"] is None
body = client.put("/api/settings",
json={"context_window_override": 8192}).json()
assert body["context_window_override"] == 8192
# And can be taken back off, which `exclude_unset` makes a real distinction:
# sending null clears it, sending nothing leaves it alone.
body = client.put("/api/settings", json={"temperature": 0.5}).json()
assert body["context_window_override"] == 8192
body = client.put("/api/settings",
json={"context_window_override": None}).json()
assert body["context_window_override"] is None
@pytest.mark.parametrize("bad", [255, 200_001])
def test_the_override_is_bounded_like_the_budget_it_caps(client, bad):
assert client.put("/api/settings",
json={"context_window_override": bad}).status_code == 422