M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
166 lines
6.1 KiB
Python
166 lines
6.1 KiB
Python
"""The window an operator declares, for a server that cannot be asked for one.
|
|
|
|
`contextwindow`'s discovery speaks Ollama's native API. Nothing restricts
|
|
`endpoint_url` to Ollama, so on vLLM, llama.cpp's own server, or anything else
|
|
serving an OpenAI-compatible `/v1`, `/api/ps` and `/api/show` are not there:
|
|
discovery fails as designed, the window is unknown, and the budget is left
|
|
uncapped at whatever is configured. That is M11's own failure mode reached by a
|
|
different route — the server drops the oldest tokens, which here are the
|
|
narrator's rules and the campaign canon.
|
|
|
|
`Settings.context_window_override` closes it. These tests pin the two properties
|
|
that make it safe rather than merely useful:
|
|
|
|
1. it is used **only** where discovery left a hole, so it can never talk the
|
|
application into a longer prompt than a server actually reported, and
|
|
2. it does not make `verified` true, because `verified` means the server
|
|
answered and a declaration is a person's claim about a server.
|
|
"""
|
|
import asyncio
|
|
|
|
import pytest
|
|
from fastapi import Depends
|
|
from fastapi.testclient import TestClient
|
|
|
|
from app import auth, contextwindow, models
|
|
from app.database import Base, SessionLocal, engine, get_db
|
|
from app.main import app
|
|
|
|
|
|
UNREACHABLE = "http://127.0.0.1:1/v1"
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _clear_window_cache():
|
|
contextwindow.cache_clear()
|
|
yield
|
|
contextwindow.cache_clear()
|
|
|
|
|
|
@pytest.fixture()
|
|
def client():
|
|
Base.metadata.create_all(bind=engine)
|
|
setup = SessionLocal()
|
|
user = models.User(is_guest=False, email="m11dw@example.com")
|
|
setup.add(user)
|
|
setup.flush()
|
|
setup.add(models.Settings(
|
|
user_id=user.id, model="some-model", endpoint_url=UNREACHABLE,
|
|
embedding_model="", context_token_budget=16384, max_output_tokens=800,
|
|
))
|
|
setup.commit()
|
|
user_id = user.id
|
|
setup.close()
|
|
|
|
app.dependency_overrides[auth.get_current_user] = (
|
|
lambda db=Depends(get_db): db.get(models.User, user_id)
|
|
)
|
|
try:
|
|
yield TestClient(app)
|
|
finally:
|
|
app.dependency_overrides.clear()
|
|
Base.metadata.drop_all(bind=engine)
|
|
|
|
|
|
def probe(endpoint=UNREACHABLE, model="some-model", declared=None):
|
|
contextwindow.cache_clear()
|
|
return asyncio.run(
|
|
contextwindow.probe(endpoint, model, declared=declared, use_cache=False))
|
|
|
|
|
|
# ------------------------------------------------------- filling the hole
|
|
|
|
def test_without_a_declaration_an_unaskable_server_leaves_the_window_unknown():
|
|
window = probe()
|
|
assert window.tokens is None
|
|
assert not window.verified
|
|
assert not window.enforceable
|
|
assert window.source == contextwindow.UNKNOWN
|
|
|
|
|
|
def test_a_declaration_becomes_the_ceiling_when_the_server_cannot_be_asked():
|
|
window = probe(declared=8192)
|
|
assert window.tokens == 8192
|
|
assert window.source == contextwindow.DECLARED
|
|
assert window.enforceable
|
|
assert contextwindow.effective_budget(16384, window) == 8192
|
|
|
|
|
|
def test_a_declaration_does_not_claim_the_server_was_verified():
|
|
"""`window_verified` travels in every turn's provenance and the M11 report
|
|
counts it. A declaration must not inflate that count."""
|
|
window = probe(declared=8192)
|
|
assert window.enforceable
|
|
assert not window.verified
|
|
|
|
|
|
def test_the_detail_says_the_number_came_from_settings():
|
|
assert "declared in settings" in probe(declared=8192).detail
|
|
|
|
|
|
@pytest.mark.parametrize("declared", [None, 0, -1])
|
|
def test_a_missing_or_meaningless_declaration_changes_nothing(declared):
|
|
window = probe(declared=declared)
|
|
assert window.tokens is None
|
|
assert window.source == contextwindow.UNKNOWN
|
|
|
|
|
|
def test_a_declaration_still_applies_when_nothing_is_configured():
|
|
window = asyncio.run(contextwindow.probe("", "", declared=4096))
|
|
assert window.tokens == 4096
|
|
assert window.source == contextwindow.DECLARED
|
|
|
|
|
|
def test_a_declaration_applies_to_a_refused_endpoint_without_reaching_it():
|
|
"""A refused address is a discovery failure like any other (ADR 011, H12).
|
|
The declaration caps the prompt; it does not make the endpoint usable, and
|
|
the turn is still refused where endpoints are enforced."""
|
|
window = probe(endpoint="http://169.254.169.254/v1", declared=4096)
|
|
assert window.source == contextwindow.DECLARED
|
|
assert window.tokens == 4096
|
|
|
|
|
|
# ------------------------------------------- a verified answer always wins
|
|
|
|
def test_a_verified_window_is_not_overridden(monkeypatch):
|
|
"""The safety property. An operator may lower an unknown ceiling into
|
|
existence; they may never raise one the server reported."""
|
|
async def reported(endpoint, model):
|
|
return contextwindow.Window(4096, contextwindow.LOADED, 32768, "real")
|
|
|
|
monkeypatch.setattr(contextwindow, "_ask", reported)
|
|
window = probe(declared=32768)
|
|
assert window.tokens == 4096
|
|
assert window.source == contextwindow.LOADED
|
|
assert window.verified
|
|
assert contextwindow.effective_budget(16384, window) == 4096
|
|
|
|
|
|
def test_a_declared_window_larger_than_the_budget_does_not_raise_it():
|
|
window = probe(declared=200_000)
|
|
assert contextwindow.effective_budget(16384, window) == 16384
|
|
|
|
|
|
# --------------------------------------------------------------- plumbing
|
|
|
|
def test_the_override_is_readable_and_settable_through_the_api(client):
|
|
assert client.get("/api/settings").json()["context_window_override"] is None
|
|
|
|
body = client.put("/api/settings",
|
|
json={"context_window_override": 8192}).json()
|
|
assert body["context_window_override"] == 8192
|
|
|
|
# And can be taken back off, which `exclude_unset` makes a real distinction:
|
|
# sending null clears it, sending nothing leaves it alone.
|
|
body = client.put("/api/settings", json={"temperature": 0.5}).json()
|
|
assert body["context_window_override"] == 8192
|
|
body = client.put("/api/settings",
|
|
json={"context_window_override": None}).json()
|
|
assert body["context_window_override"] is None
|
|
|
|
|
|
@pytest.mark.parametrize("bad", [255, 200_001])
|
|
def test_the_override_is_bounded_like_the_budget_it_caps(client, bad):
|
|
assert client.put("/api/settings",
|
|
json={"context_window_override": bad}).status_code == 422
|