Stop re-reading the whole prompt every turn, and let a lost run carry on

M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
JesseMarkowitz
2026-09-10 06:13:55 -04:00
co-authored by Claude Opus 5
parent fedb7144d0
commit ef25b0a876
22 changed files with 1654 additions and 75 deletions
+84
View File
@@ -344,6 +344,36 @@ subsystem comes back as a route, if an API key becomes settable again, if the
model timeout stops being configurable or becomes unbounded, or if a supported
start path stops binding loopback.
## Why a long campaign is not slow in proportion to its length
An inference server caches the prompt it has already processed, keyed on the
**prefix**. While a story only grows at the end, each turn re-uses that cache and
pays for its own new tokens alone. Once the context budget is full, though, the
history window has to give something up — and a window that gives up its *oldest*
action every turn changes the prompt near the front, which throws the cache away
and makes the server re-read almost the whole thing, every turn.
So the window moves in blocks. `context/builder.py` snaps the oldest included
action to a boundary and holds it there for several turns, then steps. Measured
against the reference deployment on real builder output, at an 8,192-token budget:
| | Per turn |
| --- | --- |
| Window held, story grew by one action | 14-20 s |
| Window stepped (one turn in three) | 333-338 s |
| **Mean over whole cycles** | **124.0 s** |
| Window sliding every turn, as before | 362.4 s |
The cost is history depth: right after a step the window holds up to a block
fewer actions than the budget would allow. `TRIM_FRACTION` bounds that at a
quarter of the window, and it is the one number to change if you would rather
trade recent history for speed, or the reverse.
The saving grows with the block, and the block grows with the budget — so the
larger the context window, the more this is worth. `history["floor_depth"]` and
`history["trim_block"]` are in every context report, and a `floor_depth` that is
the same on two consecutive turns is the prompt's prefix having been preserved.
## The release-validation harnesses
M11 added six runnable harnesses under `backend/tools/`. They are the evidence
@@ -370,6 +400,12 @@ AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 \
AIDND_TEST_MODEL=<model> AIDND_TEST_EMBED_MODEL=<embedding model> \
.venv/bin/python -m tools.m11_long_run --turns 100 --out "$HOME/m11-evidence/m01"
# The same campaign, carried on after a crash, a reboot or a Ctrl-C. It picks up
# the adventure the checkpoint names, keeps its place in the beat cycle, and does
# not fire a scheduled operation that already fired.
AIDND_TEST_ENDPOINT=... AIDND_TEST_MODEL=... AIDND_TEST_EMBED_MODEL=... \
.venv/bin/python -m tools.m11_long_run --turns 100 --resume --out "$HOME/m11-evidence/m01"
# What that campaign is worth on a machine that has never seen it (I01-I07).
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"
@@ -389,6 +425,35 @@ AIDND_TEST_MODEL=<model> AIDND_TEST_EMBED_MODEL=<embedding model> \
.venv/bin/python -m tools.contrast_audit
```
### Resuming the long run, and timing it out
A hundred turns is hours of wall clock, and the first release attempt lost one at
turn 97 to a host crash. The harness now checkpoints `resume.json` into `--out`
after the prologue, after every scheduled operation and after every turn, and
`--resume` continues from it. The file is written under a temporary name and
renamed, so a crash during the write cannot leave a half-parsed one.
`resume.json` is operational state rather than evidence: `timeline.jsonl` stays
the append-only record, a resumed session appends to it, and a finished run
deletes its `resume.json`. That makes the file's presence mean exactly one
thing — there is an unfinished run in this directory — and the harness refuses
to start a fresh campaign on top of one, because two campaigns interleaved in a
single timeline and database are worse evidence than none. It refuses a
directory holding a `campaign.db` with no checkpoint for the same reason.
How long a turn takes is the inference host's characteristic, not the
application's, so the timeout is an option rather than a constant:
| Flag | Default | What it does |
| --- | --- | --- |
| `--turn-timeout` | 1800 | Seconds the application waits for one narrator reply — it becomes `model_timeout_seconds`, so the settings schema's 30..3600 bound applies. The harness waits 300s longer, so the application's own error arrives inside the stream rather than being cut off at the socket. |
| `--max-consecutive-failures` | 5 | Unaccepted turns in a row before the run stops, writes `summary.json` with `status: aborted`, and leaves a `resume.json` that `--resume` can carry on. |
Measure your host before lowering `--turn-timeout`. On the M11 reference
deployment a turn cost 229-291 seconds at the recommended window; a slower host
can exceed the 600 seconds this harness used to hard-code, and an overrun turn is
a lost turn.
`tools/m11_webdriver.py` is the W3C WebDriver client the browser harness uses.
It exists so browser evidence needs no Selenium in the dependency surface, and
it documents the one environment quirk that matters here: a snap Firefox will
@@ -521,6 +586,25 @@ reports the window it found or says plainly that it could not check.
That does not make the window *bigger*, and the rest of this section is still
how you do that.
**On a server that is not Ollama, tell the application the window yourself.**
The check above uses Ollama's *native* API, which vLLM, llama.cpp's own server
and the rest do not serve — so the window comes back unverified and the budget
is left at whatever is configured. Set **`context_window_override`** in settings
to the window you launched that server with:
```bash
curl -X PUT http://127.0.0.1:8000/api/settings \
-H 'Content-Type: application/json' -d '{"context_window_override": 8192}'
```
Prompts are then capped to it. It is used *only* when the server could not be
asked — a window the server did report always wins, so this can never be a way
to over-budget an Ollama that answered — and it does not count as verification:
the turn's provenance still records that nothing checked the number. Send
`null` to remove it. Nothing here validates the figure against the server, so an
override larger than the real window puts you back to silent truncation; take it
from how you started the server, not from the model card.
**Setting it per request does not work from this application.** Ollama's
OpenAI-compatible endpoint accepts `num_ctx` — nested in `options` or at the top
level — returns HTTP 200 and ignores it. Worse, it *reloads the model at its own
+4 -1
View File
@@ -75,7 +75,10 @@ that isn't the live one starts a new branch.
refused, it is silently trimmed from the *oldest* end, which here is the narrator's rules and
your campaign's canon. The application asks the server what window your model gets and caps
the prompt to it, so what a small window costs is history rather than the canon at the front
(`backend/app/contextwindow.py`). If it cannot check, it says so instead of assuming.
(`backend/app/contextwindow.py`). If it cannot check, it says so instead of assuming — and
on a server it cannot ask, which is any server that is not Ollama, `context_window_override`
in settings lets you state the window so the prompt is still capped. A window the server
itself reported always wins over that, and a declared one is never reported as verified.
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
+128 -2
View File
@@ -35,6 +35,21 @@ AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
NPC_WINDOW = 6 # actions of story searched for NPC trigger words ("in scene")
SEPARATOR = "\n\n"
#: How much of the history window one trim gives up, as one-over-this. A
#: quarter: large enough that the window then holds still for several turns,
#: small enough that the narrator never loses most of its recent history at once.
#:
#: **This is the dial.** Lower it for bigger blocks — fewer prompt re-reads and
#: faster long campaigns, at the cost of retaining less recent history. Raise it
#: for the reverse. Nothing else has to change: `trim_block` is the only reader,
#: and `test_trim_fraction_is_the_dial_between_history_and_speed` pins that.
#: Measured at 4, on an 8,192-token budget: 124.0s per turn against 362.4s with
#: trimming off.
TRIM_FRACTION = 4
#: Never trim less than this, or the window slides by one action again and the
#: whole point is lost.
MIN_TRIM_BLOCK = 2
# Output-length guidance. The endpoint enforces `max_output_tokens` as a hard
# limit, and it truncates the reply mid-sentence when the model reaches it. The
# state block is emitted last, so truncation removes it. Asking the model to
@@ -301,6 +316,85 @@ def _canon_section(adventure: models.Adventure) -> str:
return f"Campaign canon (these are true and may not be contradicted):\n{body}"
def trim_block(history_budget: int, max_output_tokens: int) -> int:
"""How many `depth` steps of history one trim gives up.
Derived from **configuration**, never from the story, because the answer has
to be the same on two consecutive turns. A block size that moved with the
measured size of recent actions would move the boundary it defines, and a
boundary that moves is precisely what this exists to stop.
An AI action is bounded by `max_output_tokens` and a player action is small
beside it, so `max_output_tokens` is the scale of one row of history — a
setting, rather than a guess about the data.
"""
per_action = max(1, max_output_tokens)
fits = max(1, history_budget // per_action)
return max(MIN_TRIM_BLOCK, fits // TRIM_FRACTION)
def history_floor(depths: list[int | None], costs: list[int], budget: int,
block: int) -> int | None:
"""The depth of the oldest action to include, snapped to a block boundary.
## Why this is not just "whatever fits"
Taking whatever fits is what the builder did, and it is correct. It is also
the reason a long campaign costs a full prompt re-read every turn.
Inference servers cache the prompt they have already processed, keyed on the
**prefix**. While the story only grows at the end, each turn re-uses that
cache and pays for its own new tokens alone. As soon as the budget is full,
"whatever fits" drops the *oldest* action every turn — a change near the
front of the prompt — and everything after it has to be processed again.
So the floor is snapped forward to a multiple of `block` and then held. It
moves in steps: several cheap turns that re-use the cache, then one turn that
pays to re-read, rather than every turn paying. The cost is history depth —
right after a step the window holds up to `block` actions fewer than the
budget would allow, which is what `TRIM_FRACTION` bounds.
Measured against the reference deployment, on prompts this builder produced,
at an 8,192 budget where `block` is 3:
floor held, story grew by one action 14-20 s
floor stepped, prompt re-read 333-338 s
mean over two whole cycles 124.0 s
floor disabled, every turn re-read 362.4 s (361, 361, 365, 361)
2.9x, and the shape is the point rather than the ratio: the saving grows with
`block`, which grows with the budget, so the configuration that hurt most
before benefits most now.
Returns None when nothing needs trimming, which covers two cases that must
both stay as they were: a story short enough to fit whole (the window is a
growing prefix already, and snapping would drop its opening for no reason),
and an action so large that not even the newest one fits, which the caller
truncates.
"""
if not depths or any(depth is None for depth in depths):
# Legacy rows, or a path this cannot place on the tree. Trimming needs a
# stable coordinate; without one, behave exactly as before.
return None
spent = 0
oldest_fitting: int | None = None
for depth, cost in zip(reversed(depths), reversed(costs)):
if spent + cost > budget:
break
spent += cost
oldest_fitting = depth
if oldest_fitting is None:
return None
if oldest_fitting == depths[0]:
# Everything offered fits. There is nothing to drop, and snapping here
# would throw away the start of a short story to no purpose.
return None
block = max(1, block)
return -(-oldest_fitting // block) * block
def _visible_npcs(actions: list[models.Action], stat_schema: dict) -> dict[str, str]:
"""Returns the NPCs whose trigger words appear in the recent story.
@@ -633,10 +727,31 @@ def build_context(
# ----- Story history: newest first until the remaining budget is spent -----
history_budget = available_after_knowledge - used
# Where the window starts, snapped to a block so it holds still for several
# turns instead of sliding by one action every turn. `history_floor` says
# why that matters and what it costs. None means trim nothing, and then
# everything below is exactly what it was before.
costs = [count_tokens(_history_text(a)) + count_tokens(SEPARATOR)
for a in actions]
block = trim_block(history_budget, settings.max_output_tokens)
floor_depth = history_floor([a.depth for a in actions], costs,
history_budget, block)
windowed = actions
if floor_depth is not None:
kept = [a for a in actions if a.depth is not None and a.depth >= floor_depth]
# A floor that leaves nothing is a floor worth ignoring: the loop below
# still has to produce a turn, and its own truncation path is the honest
# way to handle a single action larger than the whole budget.
if kept:
windowed = kept
else:
floor_depth = None
included_actions: list[models.Action] = []
spent = 0
oldest_truncated = False
for action in reversed(actions):
for action in reversed(windowed):
# Budget against the text as it appears in the prompt, which includes
# the state block when this adventure tracks world state.
rendered = _history_text(action)
@@ -741,9 +856,13 @@ def build_context(
"source": (window.source if window is not None else contextwindow.UNKNOWN),
"model_max": (window.model_max if window is not None else None),
"detail": (window.detail if window is not None else "not checked"),
# `enforceable`, not `verified`: an operator-declared window caps
# the prompt exactly as a server-reported one does, and a turn built
# against it *was* capped. `verified` and `source` above still say
# which kind of answer produced the number.
"capped": (
window is not None
and window.verified
and window.enforceable
and window.tokens < settings.context_token_budget
),
},
@@ -775,6 +894,13 @@ def build_context(
# included, so this number must be the real total.
"total": history.count(adventure, exclude_action_id),
"oldest_truncated": oldest_truncated,
# Where the window was cut, and how big a step it takes when it
# moves. Both are in `depth` units. `floor_depth` is null while the
# story still fits whole, which is also while every turn is a pure
# prefix extension of the last one. A reader comparing two turns can
# tell from these whether the prompt's prefix was preserved.
"floor_depth": floor_depth,
"trim_block": block,
},
"settings": {
"model": settings.model,
+90 -10
View File
@@ -56,6 +56,31 @@ It does not hard-code 4,096, which would cripple a correctly configured
deployment; it does not raise the budget, which is the operator's decision; it
does not fall back to a cloud probe, a bundled table of model sizes, or a guess
from the model's name. An unknown window is reported as unknown.
## The server that cannot be asked
Discovery above is Ollama's native API. Nothing restricts `endpoint_url` to
Ollama — any allowed address serving an OpenAI-compatible `/v1` is accepted —
and on vLLM, llama.cpp's own server, or anything else, `/api/ps` and `/api/show`
are simply not there. Discovery then fails exactly as designed and the window is
reported unknown, which is honest but leaves the invariant at the top of this
file unenforced: the budget stands at whatever is configured, and if that server
enforces a smaller window it drops the oldest tokens again.
`context_window_override` is the operator's answer to that. It is a number the
operator states because they know how the server was launched, and it is used
**only when the server could not be asked**:
verified window -> always wins; a declaration cannot raise it
no verified window -> the declaration becomes the ceiling, source DECLARED
neither -> unknown, exactly as before
This does not weaken what `verified` claims. `verified` still means the server
itself answered, so `window_verified` in a turn's provenance keeps the meaning
the M11 report gives it, and a declared window is identifiable as a declaration
wherever it appears. What the declaration buys is enforcement: the prompt is
capped, so the failure mode is a shorter prompt rather than a silently truncated
one.
"""
from __future__ import annotations
@@ -86,8 +111,12 @@ NEGATIVE_TTL = 60.0
#: Sources, in the order of how much they prove.
LOADED = "loaded" # /api/ps: what the runtime is enforcing now
PARAMETERS = "parameters" # /api/show: what the model will load with
DECLARED = "declared" # the operator said so; the server could not be asked
UNKNOWN = "unknown"
#: Sources that mean *the server answered*, as opposed to somebody asserting.
FROM_SERVER = (LOADED, PARAMETERS)
@dataclass(frozen=True)
class Window:
@@ -106,6 +135,18 @@ class Window:
@property
def verified(self) -> bool:
"""The **server** answered. An operator's declaration is not this.
Kept narrow on purpose. `window_verified` travels in every turn's stored
provenance and the M11 report counts on it meaning one thing: that the
runtime was asked and replied. A declaration is a person's claim about a
server, which is worth acting on and is not the same evidence.
"""
return self.tokens is not None and self.source in FROM_SERVER
@property
def enforceable(self) -> bool:
"""There is a number to cap the prompt to, whoever supplied it."""
return self.tokens is not None
@@ -128,9 +169,10 @@ def native_base(endpoint_url: str) -> str:
def effective_budget(configured: int, window: Window | int | None) -> int:
"""The budget the prompt may actually use.
The whole enforcement, in one line: a verified window is a ceiling. The
configured budget still wins when it is *smaller*, because a reader who has
deliberately asked for a shorter prompt should get one.
The whole enforcement, in one line: a known window is a ceiling — whether
the server reported it or the operator declared it. The configured budget
still wins when it is *smaller*, because a reader who has deliberately asked
for a shorter prompt should get one.
"""
tokens = window.tokens if isinstance(window, Window) else window
if tokens is None or tokens <= 0:
@@ -143,17 +185,55 @@ def cache_clear() -> None:
_cache.clear()
async def probe(endpoint_url: str, model: str, *, use_cache: bool = True) -> Window:
"""Asks the server what window `model` gets. Never raises.
async def probe(endpoint_url: str, model: str, *,
declared: int | None = None, use_cache: bool = True) -> Window:
"""What window `model` gets, asked of the server and only then declared.
Returns `UNVERIFIED` for every failure — refused endpoint, unreachable
server, TLS failure, a non-Ollama endpoint, an unparseable answer. The caller
cannot act differently on those and the reader is told the same thing either
way: the window could not be checked.
Returns `UNVERIFIED` for every discovery failure — refused endpoint,
unreachable server, TLS failure, a server with no Ollama-native API, an
unparseable answer — unless `declared` supplies a number to fall back on.
The caller cannot act differently on those failures and the reader is told
the same thing either way: the window could not be checked.
`declared` is `Settings.context_window_override`. It never overrides a
verified answer, so an operator cannot talk the application into a bigger
prompt than the runtime will read; it only fills a gap discovery left.
"""
if not endpoint_url or not model:
return Window(None, UNKNOWN, detail="no endpoint or model configured")
return _declared_or(declared,
Window(None, UNKNOWN,
detail="no endpoint or model configured"))
discovered = await _discover(endpoint_url, model, use_cache=use_cache)
return _declared_or(declared, discovered)
def _declared_or(declared: int | None, discovered: Window) -> Window:
"""The operator's number, but only where the server left a hole.
A verified window always wins. That ordering is the whole safety property:
a declaration can lower an unknown ceiling into existence, never raise a
known one.
"""
if discovered.verified:
return discovered
if not declared or declared <= 0:
return discovered
return Window(
declared, DECLARED, discovered.model_max,
f"{declared:,} tokens, declared in settings — the server was not able "
f"to say ({discovered.detail})",
)
async def _discover(endpoint_url: str, model: str, *,
use_cache: bool = True) -> Window:
"""The server's own answer, cached. Knows nothing about declarations.
The cache holds only what was discovered, so changing the declared override
takes effect on the next turn without having to clear anything: the
declaration is layered on afterwards, in `_declared_or`.
"""
key = (endpoint_url, model)
now = time.monotonic()
if use_cache:
+4
View File
@@ -475,6 +475,10 @@ MIGRATIONS: list[tuple[int, str | dict[str, str]]] = [
# an existing campaign's prompts do not change under it.
(93, "ALTER TABLE adventures ADD COLUMN narration_length VARCHAR(20) "
"NOT NULL DEFAULT ''"),
# Nullable, and null by default: an override that defaulted to a number
# would be the application guessing at a window again, which is the one
# thing `contextwindow` refuses to do. Null means "nobody has said".
(94, "ALTER TABLE settings ADD COLUMN context_window_override INTEGER"),
]
LATEST_VERSION = max((v for v, _ in MIGRATIONS), default=1)
+8
View File
@@ -1224,6 +1224,14 @@ class Settings(Base):
# while the same turn takes seconds once the model is resident. See
# `providers.openai_compatible.DEFAULT_READ_TIMEOUT`.
model_timeout_seconds: Mapped[int] = mapped_column(Integer, default=300)
# What window the inference server enforces, when the server cannot be asked
# for it. Discovery (`contextwindow`) speaks Ollama's native API; a server
# that does not serve one — vLLM, llama.cpp's own server — leaves the window
# unknown and the budget uncapped. This is the operator saying how they
# launched it. It never overrides a window the server did report, and null
# means nobody has said, because a default here would be a guess.
context_window_override: Mapped[int | None] = mapped_column(
Integer, nullable=True, default=None)
narrator_prompt: Mapped[str] = mapped_column(
Text,
default=(
+2 -1
View File
@@ -33,7 +33,8 @@ async def dry_run_context(
# M11: and by the same probe the turn makes, for the same reason — a panel
# that showed a 16,384-token budget while the next turn will be capped to
# 4,096 would be showing a prompt that is not the one about to be sent.
window = await contextwindow.probe(settings.endpoint_url, settings.model)
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
try:
_, _, report = build_context(
adventure, settings, memories, knowledge=knowledge, window=window
+2 -1
View File
@@ -208,7 +208,8 @@ async def _generate_turn(
# network calls — and cached per endpoint and model, so it costs one short
# request per session rather than one per turn. An unverified window does
# not block the turn; it is recorded as unverified in the snapshot below.
window = await contextwindow.probe(settings.endpoint_url, settings.model)
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
try:
system_text, story_text, snapshot = build_context(
adventure,
+25 -5
View File
@@ -177,19 +177,38 @@ async def list_endpoint_models(endpoint_url: str) -> dict:
def _window_warning(window: contextwindow.Window, settings: models.Settings) -> str | None:
"""What to tell the reader about the window, or None when nothing is wrong.
Three cases, and they need three different things done about them, so they
say three different things (the same reasoning as the connection test's own
Four cases, and they need four different things done about them, so they
say four different things (the same reasoning as the connection test's own
four failure kinds).
"""
budget = settings.context_token_budget
if window.source == contextwindow.DECLARED:
# Enforced, but on the operator's word rather than the server's. Worth
# saying plainly: nothing here has checked the number, so a declaration
# that is too large is the silent-truncation failure all over again.
over = (
" It is larger than the story budget, so it changes nothing today."
if window.tokens >= budget else
f" Prompts are being built to {window.tokens:,} rather than "
f"{budget:,}."
)
return (
f"The context window for '{settings.model}' is set in settings to "
f"{window.tokens:,} tokens, because this server cannot be asked for it "
f"— {window.detail}.{over} Nothing has verified that number against "
"the server; if it is larger than the window the server really "
"enforces, the oldest part of the prompt is still being dropped."
)
if not window.verified:
return (
f"The context window this server will give '{settings.model}' could not "
f"be checked — {window.detail}. The story budget is {budget:,} tokens; "
"if the server's window is smaller than that it silently drops the "
"oldest part of the prompt, which here is the narrator's rules and the "
"campaign canon. See DEVELOPMENT.md, 'The context window your Ollama "
"actually enforces'."
"campaign canon. If this server has no Ollama-native API to ask — "
"vLLM, llama.cpp's own server — set the context window in settings so "
"the prompt is capped to it. See DEVELOPMENT.md, 'The context window "
"your Ollama actually enforces'."
)
if window.tokens < budget:
ceiling = (
@@ -226,7 +245,8 @@ async def test_connection(
# an operator reloads a model. Changing the endpoint or the model clears
# the cache (`update_settings`), which covers the case a reader can
# actually cause; the detail line always says where the number came from.
window = await contextwindow.probe(settings.endpoint_url, settings.model)
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
result = result | {"window": {
"verified": window.verified,
"tokens": window.tokens,
+6
View File
@@ -687,6 +687,7 @@ class SettingsOut(ORMModel):
max_output_tokens: int
context_token_budget: int
model_timeout_seconds: int
context_window_override: int | None
narrator_prompt: str
summary_model: str
embedding_model: str
@@ -730,6 +731,11 @@ class SettingsUpdate(BaseModel):
# turn cannot trip it; the ceiling exists so that "wait longer" stays a
# number rather than becoming "wait forever".
model_timeout_seconds: Annotated[int, Field(ge=30, le=3600)] | None = None
# The window an inference server enforces, for servers that cannot be asked.
# Bounded like the budget it caps. It is never a way to *raise* the prompt
# past a window the server did report — `contextwindow._declared_or` — so
# the ceiling here only bounds what an operator can usefully claim.
context_window_override: Annotated[int, Field(ge=256, le=200_000)] | None = None
narrator_prompt: Prose | None = None
summary_model: Name | None = None
embedding_model: Name | None = None
+298
View File
@@ -0,0 +1,298 @@
"""The history window moves in blocks, so the prompt's prefix holds still.
Inference servers cache a prompt by its **prefix**. While a story only grows at
the end, every turn re-uses that cache and pays for its own new tokens alone. The
builder's old window took whatever fit, which meant that once the budget was full
it dropped the *oldest* action every turn — a change near the front of the prompt
— and everything after it had to be processed again.
Measured on the reference deployment, at 7.7k prompt tokens against a 3B model:
window slid by one turn 343-350 s
prefix preserved 6.0 s
These tests do not measure time. They pin the property the measurement is
downstream of: **the oldest included action is the same across consecutive
turns**, except on the turns where the window deliberately steps.
"""
import pytest
import pytest as _pytest
from app.context import builder
def costs_of(n, each=100):
return [each] * n
def depths(n, start=1):
return list(range(start, start + n))
# ------------------------------------------------------------- the block size
def test_the_block_is_a_share_of_what_fits():
"""Derived from `TRIM_FRACTION` rather than asserting the number it is
currently set to, so retuning the dial does not fail a test that was never
about the dial's value."""
# 16 actions of 500 fit in 8,000, and the block is that share of them.
assert builder.trim_block(8000, 500) == 16 // builder.TRIM_FRACTION
def test_trim_fraction_is_the_dial_between_history_and_speed():
"""`TRIM_FRACTION` is meant to be retuned, so this pins what retuning does.
Lower it and the window gives up more at once: bigger blocks, fewer re-reads,
less recent history retained. Raise it and the reverse. Nothing else in the
builder has to change for that to hold, which is the property worth having a
test for.
"""
budget, per_action = 8000, 500
fits = budget // per_action
def block_at(fraction, monkeypatch):
monkeypatch.setattr(builder, "TRIM_FRACTION", fraction)
return builder.trim_block(budget, per_action)
with _pytest.MonkeyPatch.context() as mp:
greedier = block_at(2, mp)
assert greedier == fits // 2
with _pytest.MonkeyPatch.context() as mp:
gentler = block_at(8, mp)
assert gentler == max(builder.MIN_TRIM_BLOCK, fits // 8)
assert greedier > gentler, "a lower fraction must give up more at once"
# And the floor to the whole thing survives any setting.
with _pytest.MonkeyPatch.context() as mp:
mp.setattr(builder, "TRIM_FRACTION", 1000)
assert builder.trim_block(budget, per_action) >= builder.MIN_TRIM_BLOCK
def test_the_block_never_slides_by_one():
"""A block of one is the old behaviour wearing a hat."""
assert builder.trim_block(100, 500) >= builder.MIN_TRIM_BLOCK
assert builder.trim_block(0, 500) >= builder.MIN_TRIM_BLOCK
def test_the_block_comes_from_settings_not_from_the_story():
"""It has to be the same on two consecutive turns, so it cannot be measured
from actions whose sizes vary."""
assert builder.trim_block(8000, 500) == builder.trim_block(8000, 500)
# Bigger budget, bigger step; the ratio is what is fixed.
assert builder.trim_block(16000, 500) > builder.trim_block(8000, 500)
# ------------------------------------------------------- nothing to trim yet
def test_a_story_that_fits_whole_is_not_trimmed():
"""Also the append-only regime: every turn is a prefix extension already."""
assert builder.history_floor(depths(5), costs_of(5), budget=10_000, block=4) is None
def test_a_short_story_keeps_its_opening():
"""Snapping here would drop the start of the story for no reason at all."""
assert builder.history_floor(depths(3), costs_of(3), budget=10_000, block=8) is None
def test_an_action_larger_than_the_budget_is_left_to_the_caller():
assert builder.history_floor([1], [5000], budget=100, block=4) is None
def test_rows_without_a_depth_are_not_trimmed():
"""Legacy rows have no stable coordinate, so behave exactly as before."""
assert builder.history_floor([None, None], costs_of(2), 100, 4) is None
assert builder.history_floor([], [], 100, 4) is None
# ------------------------------------------------------------ the whole point
def test_the_floor_holds_still_while_the_story_grows():
"""The property the 57x measurement rests on.
Ten consecutive turns against a full budget. The floor must take a small
number of steps, not ten.
"""
block, budget, each = 4, 1000, 100 # 10 actions fit
seen = []
for extra in range(10): # the story grows by one action
n = 20 + extra
seen.append(builder.history_floor(depths(n), costs_of(n, each), budget, block))
steps = sum(1 for a, b in zip(seen, seen[1:]) if a != b)
assert steps <= 3, f"the floor moved {steps} times in 10 turns: {seen}"
assert len(set(seen)) > 1, "it never moved at all, so the budget is not binding"
def test_every_floor_sits_on_a_block_boundary():
block, budget = 4, 1000
for n in range(20, 40):
floor = builder.history_floor(depths(n), costs_of(n), budget, block)
assert floor is not None
assert floor % block == 0, f"{floor} is not a multiple of {block}"
def test_the_floor_only_ever_moves_forward():
block, budget = 4, 1000
floors = [builder.history_floor(depths(n), costs_of(n), budget, block)
for n in range(20, 45)]
assert floors == sorted(floors)
# --------------------------------------------------- and still inside budget
@pytest.mark.parametrize("n", range(20, 40))
def test_the_kept_window_never_exceeds_the_budget(n):
"""M03's bound is not weakened. Trimming only ever drops more, never less."""
block, budget, each = 4, 1000, 100
ds, cs = depths(n), costs_of(n, each)
floor = builder.history_floor(ds, cs, budget, block)
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
assert kept <= budget
@pytest.mark.parametrize("n", range(20, 40))
def test_the_kept_window_is_not_gutted(n):
"""The cost of holding still is bounded: a trim gives up `block` actions,
never most of the window."""
block, budget, each = 4, 1000, 100
ds, cs = depths(n), costs_of(n, each)
floor = builder.history_floor(ds, cs, budget, block)
kept = sum(c for d, c in zip(ds, cs) if d >= floor)
assert kept >= budget - block * each
# ------------------------------------------- the real builder, end to end
import pytest as _pytest # noqa: E402 (grouped with the fixtures it serves)
from app import models # noqa: E402
from app.context import builder as _builder # noqa: E402
from app.database import Base, SessionLocal, engine # noqa: E402
NARRATION = ("The rain came down over Westhaven in long grey sheets and the "
"gutters ran full from the ridge to the waterfront. ") * 6
@_pytest.fixture()
def saturated():
"""A campaign whose history is longer than its budget, with real depths.
`depth` is what the floor is expressed in, and every action written through
the application has one (`tree.place_action`). The older fixtures in
`test_history_window.py` predate the tree and leave it null, which is why
trimming does not engage there and those tests still describe the old
behaviour exactly.
"""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
user = models.User(is_guest=False, email="blocktrim@example.com")
db.add(user)
db.flush()
settings = models.Settings(user_id=user.id, model="m", embedding_model="",
context_token_budget=2048, max_output_tokens=200)
db.add(settings)
adventure = models.Adventure(user_id=user.id, title="Long", script_state={})
db.add(adventure)
db.flush()
for i in range(60):
db.add(models.Action(adventure_id=adventure.id,
type="ai" if i % 2 else "do",
text=f"[{i}] {NARRATION}", branch_id=None, depth=i))
db.commit()
db.expire_all()
adventure = db.get(models.Adventure, adventure.id)
settings = db.get(models.Settings, settings.id)
try:
yield db, adventure, settings
finally:
db.close()
Base.metadata.drop_all(bind=engine)
def _play_one_more(db, adventure, at_depth):
db.add(models.Action(adventure_id=adventure.id, type="do",
text=f"[{at_depth}] {NARRATION}", depth=at_depth))
# `at_depth` may be None: the control below plays a turn into a story whose
# rows predate the tree, which is the ungoverned window this replaced.
db.commit()
db.expire_all()
def _shared_prefix(before: str, after: str) -> float:
"""How much of the old prompt the new one still opens with, 0.0 to 1.0.
This is the quantity the inference server's cache is keyed on, so it is the
quantity worth asserting. It is not 1.0 even in the best case: the prompt
ends with the turn's length-hint and state-block instructions, which sit
*after* the history, so appending a turn always rewrites that tail.
"""
shared = 0
for x, y in zip(before, after):
if x != y:
break
shared += 1
return shared / max(1, len(before))
def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
"""The property the whole change exists for.
Not a timing test — it asserts what the timing follows from. The story text
a turn sends still opens with almost all of what the previous turn sent, so
the server's prompt cache covers that part and only the tail is processed.
"""
db, adventure, settings = saturated
_, before, report_before = _builder.build_context(adventure, settings)
assert report_before["history"]["floor_depth"] is not None, (
"this fixture is meant to be over budget; trimming never engaged")
_play_one_more(db, adventure, 60)
_, after, report_after = _builder.build_context(adventure, settings)
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"]
assert _shared_prefix(before, after) > 0.85
def test_without_a_stable_floor_the_prefix_collapses(saturated):
"""The control, and the behaviour this replaced.
Rows with no `depth` cannot be placed on the tree, so the floor cannot be
computed and the window takes whatever fits — sliding by one action every
turn. The new prompt then starts with a *different* action, the shared
prefix collapses, and the server reprocesses essentially the whole thing.
That is the 343s case in this module's docstring.
"""
db, adventure, settings = saturated
for action in db.query(models.Action).all():
action.depth = None
db.commit()
db.expire_all()
_, before, report = _builder.build_context(adventure, settings)
assert report["history"]["floor_depth"] is None
_play_one_more(db, adventure, None)
_, after, _ = _builder.build_context(adventure, settings)
assert _shared_prefix(before, after) < 0.1
def test_the_window_does_step_eventually(saturated):
"""It holds still, but it must not hold still for ever — the budget is a
bound, and a window that never moved would break it."""
db, adventure, settings = saturated
first = _builder.build_context(adventure, settings)[2]["history"]["floor_depth"]
seen = {first}
for depth in range(60, 90):
_play_one_more(db, adventure, depth)
seen.add(_builder.build_context(adventure, settings)[2]["history"]["floor_depth"])
assert len(seen) > 1, "the floor never moved across 30 turns"
def test_the_prompt_stays_inside_the_budget_as_the_window_steps(saturated):
"""M03's bound, across the step. Trimming only ever drops more history."""
db, adventure, settings = saturated
for depth in range(60, 85):
_play_one_more(db, adventure, depth)
report = _builder.build_context(adventure, settings)[2]
assert report["tokens"]["total"] <= report["tokens"]["budget"]
+1 -1
View File
@@ -369,7 +369,7 @@ def test_a_turn_records_the_window_it_was_built_against(client, monkeypatch):
the campaign whether it was built against a checked window, rather than
inferring it from what the settings say today.
"""
async def verified(endpoint, model, use_cache=True):
async def verified(endpoint, model, declared=None, use_cache=True):
return contextwindow.Window(4096, contextwindow.LOADED, 32768, "fake")
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
+165
View File
@@ -0,0 +1,165 @@
"""The window an operator declares, for a server that cannot be asked for one.
`contextwindow`'s discovery speaks Ollama's native API. Nothing restricts
`endpoint_url` to Ollama, so on vLLM, llama.cpp's own server, or anything else
serving an OpenAI-compatible `/v1`, `/api/ps` and `/api/show` are not there:
discovery fails as designed, the window is unknown, and the budget is left
uncapped at whatever is configured. That is M11's own failure mode reached by a
different route — the server drops the oldest tokens, which here are the
narrator's rules and the campaign canon.
`Settings.context_window_override` closes it. These tests pin the two properties
that make it safe rather than merely useful:
1. it is used **only** where discovery left a hole, so it can never talk the
application into a longer prompt than a server actually reported, and
2. it does not make `verified` true, because `verified` means the server
answered and a declaration is a person's claim about a server.
"""
import asyncio
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, contextwindow, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
UNREACHABLE = "http://127.0.0.1:1/v1"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
@pytest.fixture()
def client():
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="m11dw@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model="some-model", endpoint_url=UNREACHABLE,
embedding_model="", context_token_budget=16384, max_output_tokens=800,
))
setup.commit()
user_id = user.id
setup.close()
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
try:
yield TestClient(app)
finally:
app.dependency_overrides.clear()
Base.metadata.drop_all(bind=engine)
def probe(endpoint=UNREACHABLE, model="some-model", declared=None):
contextwindow.cache_clear()
return asyncio.run(
contextwindow.probe(endpoint, model, declared=declared, use_cache=False))
# ------------------------------------------------------- filling the hole
def test_without_a_declaration_an_unaskable_server_leaves_the_window_unknown():
window = probe()
assert window.tokens is None
assert not window.verified
assert not window.enforceable
assert window.source == contextwindow.UNKNOWN
def test_a_declaration_becomes_the_ceiling_when_the_server_cannot_be_asked():
window = probe(declared=8192)
assert window.tokens == 8192
assert window.source == contextwindow.DECLARED
assert window.enforceable
assert contextwindow.effective_budget(16384, window) == 8192
def test_a_declaration_does_not_claim_the_server_was_verified():
"""`window_verified` travels in every turn's provenance and the M11 report
counts it. A declaration must not inflate that count."""
window = probe(declared=8192)
assert window.enforceable
assert not window.verified
def test_the_detail_says_the_number_came_from_settings():
assert "declared in settings" in probe(declared=8192).detail
@pytest.mark.parametrize("declared", [None, 0, -1])
def test_a_missing_or_meaningless_declaration_changes_nothing(declared):
window = probe(declared=declared)
assert window.tokens is None
assert window.source == contextwindow.UNKNOWN
def test_a_declaration_still_applies_when_nothing_is_configured():
window = asyncio.run(contextwindow.probe("", "", declared=4096))
assert window.tokens == 4096
assert window.source == contextwindow.DECLARED
def test_a_declaration_applies_to_a_refused_endpoint_without_reaching_it():
"""A refused address is a discovery failure like any other (ADR 011, H12).
The declaration caps the prompt; it does not make the endpoint usable, and
the turn is still refused where endpoints are enforced."""
window = probe(endpoint="http://169.254.169.254/v1", declared=4096)
assert window.source == contextwindow.DECLARED
assert window.tokens == 4096
# ------------------------------------------- a verified answer always wins
def test_a_verified_window_is_not_overridden(monkeypatch):
"""The safety property. An operator may lower an unknown ceiling into
existence; they may never raise one the server reported."""
async def reported(endpoint, model):
return contextwindow.Window(4096, contextwindow.LOADED, 32768, "real")
monkeypatch.setattr(contextwindow, "_ask", reported)
window = probe(declared=32768)
assert window.tokens == 4096
assert window.source == contextwindow.LOADED
assert window.verified
assert contextwindow.effective_budget(16384, window) == 4096
def test_a_declared_window_larger_than_the_budget_does_not_raise_it():
window = probe(declared=200_000)
assert contextwindow.effective_budget(16384, window) == 16384
# --------------------------------------------------------------- plumbing
def test_the_override_is_readable_and_settable_through_the_api(client):
assert client.get("/api/settings").json()["context_window_override"] is None
body = client.put("/api/settings",
json={"context_window_override": 8192}).json()
assert body["context_window_override"] == 8192
# And can be taken back off, which `exclude_unset` makes a real distinction:
# sending null clears it, sending nothing leaves it alone.
body = client.put("/api/settings", json={"temperature": 0.5}).json()
assert body["context_window_override"] == 8192
body = client.put("/api/settings",
json={"context_window_override": None}).json()
assert body["context_window_override"] is None
@pytest.mark.parametrize("bad", [255, 200_001])
def test_the_override_is_bounded_like_the_budget_it_caps(client, bad):
assert client.put("/api/settings",
json={"context_window_override": bad}).status_code == 422
+225
View File
@@ -0,0 +1,225 @@
"""The long run's resume checkpoint, and the refusals that protect its evidence.
M11's release campaign was lost twice over: once to a host crash at turn 97, and
again to the fact that starting the harness a second time began a new campaign
rather than continuing the old one. `tools/m11_long_run.py` now checkpoints
`resume.json` and can be pointed back at it.
These tests exercise that logic without a narrator, a server or a database,
because none of it needs one: the checkpoint is a file, and the decisions made
around it are decisions about files. What they cannot prove is that a resumed
campaign continues correctly against a real application — that is what the run
itself proves, and §G of the M11 report is where it is reported.
"""
import json
import pytest
from tools import m11_long_run as lr
class FakeServer:
"""Enough of `Storyteller` for the checkpoint: it records process starts."""
def __init__(self, starts=1):
self.starts = starts
@pytest.fixture
def run(tmp_path):
made = lr.Run(FakeServer(), tmp_path, turns_target=100)
yield made
made.timeline.close()
# --------------------------------------------------------------- checkpoint
def test_a_checkpoint_carries_what_a_resume_needs(run, tmp_path):
run.adv = 7
run.accepted = 41
run.beat = 44
run.completed_steps = {7, 14, 21}
run.save_resume()
saved = json.loads((tmp_path / lr.RESUME_FILE).read_text())
assert saved["adventure"] == 7
assert saved["accepted"] == 41
assert saved["beat"] == 44
assert saved["completed_steps"] == [7, 14, 21]
assert saved["server_starts"] == 1
assert saved["turns_target"] == 100
def test_the_checkpoint_is_replaced_rather_than_appended(run, tmp_path):
run.adv = 7
run.accepted = 1
run.save_resume()
run.accepted = 2
run.save_resume()
assert json.loads((tmp_path / lr.RESUME_FILE).read_text())["accepted"] == 2
# The temporary name it is written under must not survive the rename.
assert not (tmp_path / (lr.RESUME_FILE + ".tmp")).exists()
def test_a_checkpoint_round_trips_into_a_later_session(run, tmp_path):
run.adv = 7
run.accepted = 41
run.beat = 44
run.completed_steps = {7, 14}
run.elapsed_before = 100
run.save_resume()
saved = json.loads((tmp_path / lr.RESUME_FILE).read_text())
later = lr.Run(FakeServer(starts=3), tmp_path, turns_target=100)
try:
later.adopt(saved)
assert later.adv == 7
assert later.accepted == 41
assert later.beat == 44
assert later.completed_steps == {7, 14}
assert later.resumed is True
# Run time accumulates across sessions rather than restarting.
assert later.elapsed_before >= 100
assert later.elapsed() >= 100
finally:
later.timeline.close()
def test_a_resumed_session_appends_to_the_existing_timeline(run, tmp_path):
run.adv = 7
run.note("turn", text="the first session")
run.timeline.close()
later = lr.Run(FakeServer(), tmp_path, turns_target=100)
try:
later.note("resumed", adventure=7)
finally:
later.timeline.close()
lines = (tmp_path / "timeline.jsonl").read_text().strip().splitlines()
assert [json.loads(line)["kind"] for line in lines] == ["turn", "resumed"]
# ------------------------------------------------------------- the decision
def test_a_clean_directory_starts_a_run(tmp_path):
assert lr._resume_state(
tmp_path / lr.RESUME_FILE, tmp_path / "timeline.jsonl", False) is None
def test_an_unfinished_run_is_not_overwritten(tmp_path):
resume_path = tmp_path / lr.RESUME_FILE
resume_path.write_text(json.dumps({"adventure": 7, "accepted": 41}))
refusal = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", False)
assert isinstance(refusal, str)
assert "--resume" in refusal
def test_a_recorded_run_with_no_checkpoint_is_not_reused(tmp_path):
"""A run that recorded something and then died before its first checkpoint.
Starting here would put a second campaign in the same timeline."""
timeline = tmp_path / "timeline.jsonl"
timeline.write_text(json.dumps({"kind": "settings"}) + "\n")
refusal = lr._resume_state(tmp_path / lr.RESUME_FILE, timeline, False)
assert isinstance(refusal, str)
assert "second campaign" in refusal
def test_a_run_that_recorded_nothing_leaves_the_directory_usable(tmp_path):
"""A server that never came up opens the timeline and writes no line to it.
Nothing was written that a fresh run could collide with."""
(tmp_path / "timeline.jsonl").write_text("")
assert lr._resume_state(
tmp_path / lr.RESUME_FILE, tmp_path / "timeline.jsonl", False) is None
def test_resuming_returns_the_checkpoint(tmp_path):
resume_path = tmp_path / lr.RESUME_FILE
resume_path.write_text(json.dumps({"adventure": 7, "accepted": 41}))
prior = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", True)
assert prior["adventure"] == 7
assert prior["accepted"] == 41
def test_resuming_nothing_is_refused_rather_than_started_fresh(tmp_path):
refusal = lr._resume_state(
tmp_path / lr.RESUME_FILE, tmp_path / "timeline.jsonl", True)
assert isinstance(refusal, str)
assert "no resume.json" in refusal
def test_an_unreadable_checkpoint_is_refused(tmp_path):
resume_path = tmp_path / lr.RESUME_FILE
resume_path.write_text("{not json")
refusal = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", True)
assert isinstance(refusal, str)
assert "cannot read" in refusal
def test_a_checkpoint_naming_no_campaign_is_refused(tmp_path):
resume_path = tmp_path / lr.RESUME_FILE
resume_path.write_text(json.dumps({"accepted": 41}))
refusal = lr._resume_state(resume_path, tmp_path / "timeline.jsonl", True)
assert isinstance(refusal, str)
assert "names no campaign" in refusal
# ----------------------------------------------------------------- timeouts
def test_the_harness_waits_longer_than_the_application_does(tmp_path):
"""Otherwise the socket closes before the application can report the failure
inside the stream, and a real error is recorded as a transport one."""
server = lr.Storyteller(tmp_path / "campaign.db", tmp_path / "server.log",
turn_timeout=1800)
assert server.stream_timeout > server.turn_timeout
def test_the_default_timeout_is_inside_the_settings_bound():
"""`app/schemas.py` bounds model_timeout_seconds at 30..3600."""
assert 30 <= lr.DEFAULT_TURN_TIMEOUT <= 3600
# ----------------------------------------------------------------- schedule
def test_every_scheduled_operation_has_its_own_turn(tmp_path):
"""The steps are keyed by turn number, which is what lets a completed one be
remembered across a resume: two are called `retry` and two `restart`, so a
name does not identify one."""
plan = lr._schedule(100)
assert len(plan) == len(set(plan)) == 12
assert sorted(plan)[0] >= 1
assert max(plan) < 100
names = list(plan.values())
assert names.count("restart") == 2
assert names.count("retry") == 2
# --------------------------------------------------------------- the clue
def test_the_planted_clue_uses_a_field_add_fact_actually_carries():
"""M04's state half turns on the clue text reaching the stored fact.
`add_fact` requires `predicate` and accepts `subject`, `object`, `value` and
`fact_id`. A key it does not define is dropped, and the correction still
succeeds — so a clue planted into the wrong key leaves a fact asserting
nothing, and `_recall` reports a recall failure the application did not
cause. This test fails against the `detail` key that used to be sent.
"""
from app.narrative.events import SPECS
spec = SPECS["add_fact"]
allowed = {"type"} | set(spec["required"]) | set(spec["optional"])
assert set(lr.CLUE_FACT) <= allowed, (
f"{set(lr.CLUE_FACT) - allowed} is not carried by add_fact")
def test_the_planted_clue_carries_the_sentinel_recall_looks_for():
assert lr.CLUE_SENTINEL in lr.CLUE_FACT["value"]
assert lr.CLUE in lr.CLUE_FACT["value"]
+49 -8
View File
@@ -2,9 +2,20 @@
python -m tools.m11_browser --out <dir> [--show]
Run from `backend/`, with `frontend/dist` already built. Uses a real narrator
when `AIDND_TEST_ENDPOINT`/`AIDND_TEST_MODEL` are set; the checks that do not
need narration run either way and say which they are.
Run from `backend/`, with `frontend/dist` already built.
**A narrator is required**, from `AIDND_TEST_ENDPOINT`/`AIDND_TEST_MODEL`, and
the run refuses to start without one. `--no-narrator` opts out explicitly, and
then every check that needs a played turn *skips* and the run is marked
`partial` — it is a smoke test of the deterministic checks, not release evidence.
That is deliberate, and it is the second harness defect of this shape M11 has
found. Without a narrator no turn is ever played, so the campaign has no history:
Undo is correctly disabled, `D01` fails, the click that follows it throws, and a
run reports two failures that look exactly like a product regression and are not.
An unset environment variable must not be able to produce that. `--out` is
required on every harness here for the same reason — so the decision is made on
purpose rather than by omission.
## What this is and is not
@@ -542,8 +553,22 @@ def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--out", required=True)
parser.add_argument("--show", action="store_true", help="not headless")
parser.add_argument(
"--no-narrator", action="store_true",
help=("run only the checks that need no narration, and mark the run "
"partial. Not release evidence."))
args = parser.parse_args()
narrated = bool(ENDPOINT and MODEL) and not args.no_narrator
if not narrated and not args.no_narrator:
print("AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL are not set.\n"
"A browser release regression needs a narrator: without one no turn "
"is played, the campaign has no history, and the history controls "
"fail for a reason that is not the product's.\n"
"Set both, or pass --no-narrator to run the deterministic checks "
"alone and get a run marked partial.")
return 2
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
dist = BACKEND.parent / "frontend" / "dist" / "index.html"
@@ -560,12 +585,12 @@ def main() -> int:
print(f"build: {dist.stat().st_mtime} served at {site.url}\n")
try:
if ENDPOINT and MODEL:
if narrated:
site.api("PUT", "/settings", {
"endpoint_url": ENDPOINT, "model": MODEL,
"max_output_tokens": 400, "model_timeout_seconds": 600})
adv = campaign_with_story(site, checks)
if ENDPOINT and MODEL:
if narrated:
for text in ("I ask Mara what she has heard.",
"I show her the silver key."):
events = play_a_turn(site, adv, text)
@@ -574,11 +599,21 @@ def main() -> int:
not errors, errors[0].get("detail", "")[:120] if errors else "")
else:
checks.skip("B01", "narration through a real model",
"AIDND_TEST_ENDPOINT/MODEL not set")
"--no-narrator")
# The history controls read a story that has turns in it, and only a
# narrator puts turns there. Skipped rather than run against an empty
# campaign, because Undo being correctly disabled is not a D01 failure.
history_controls = (
(lambda: check_history_controls(browser, site, adv, checks))
if narrated else
(lambda: checks.skip("D01-D14", "history controls in the browser",
"--no-narrator: the campaign has no turns"))
)
for scenario in (
lambda: check_shell_and_title(browser, site, adv, checks),
lambda: check_history_controls(browser, site, adv, checks),
history_controls,
lambda: check_markdown_safety(browser, site, checks),
lambda: check_hidden_knowledge(browser, site, checks),
lambda: check_context_inspection(browser, site, adv, checks),
@@ -599,7 +634,10 @@ def main() -> int:
"browser": f"Firefox {browser.version}",
"started": started.isoformat(timespec="seconds"),
"seconds": round((datetime.now() - started).total_seconds()),
"narrator": MODEL or "none (deterministic checks only)",
"narrator": MODEL if narrated else "none (--no-narrator)",
# Release evidence, or a smoke test. A reader should not have to infer
# which from the skip count.
"kind": "release regression" if narrated else "partial (no narrator)",
"checks": checks.rows,
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
"failed": len(checks.failed),
@@ -608,6 +646,9 @@ def main() -> int:
(out / "browser-report.json").write_text(json.dumps(report, indent=2))
print(f"\n{report['passed']} passed, {report['failed']} failed, "
f"{report['skipped']} skipped -> {out / 'browser-report.json'}")
if not narrated:
print("PARTIAL: no narrator, so the narration and history checks did "
"not run. This is not release evidence.")
return 1 if checks.failed else 0
+350 -39
View File
@@ -28,10 +28,47 @@ recall check at the end asks the application what it would actually send.
## What it records
A JSON line per turn (`timeline.jsonl`) with the context measurements M03 wants,
a `measurements.json` of the sampled checkpoints, the recall evidence for M04,
and the final bundle. Everything is written as it happens, so a run that dies at
turn 80 still leaves 80 turns of evidence rather than nothing.
A JSON line per turn (`timeline.jsonl`) carrying the context measurements M03
wants, the recall evidence for M04 (`recall.json`), the exported campaign
(`bundle.json`) and `summary.json`. Everything is written as it happens, so a
run that dies at turn 80 still leaves 80 turns of evidence rather than nothing.
## Resuming
A hundred turns is hours of wall clock, and the first release attempt lost one
at turn 97 to a host crash. `timeline.jsonl` survived that; the run did not,
because starting the harness again began a new campaign at turn 1.
So a run now checkpoints `resume.json` beside its evidence — after the prologue,
after every scheduled operation, and after every turn — and `--resume` picks the
campaign back up where it stopped: same adventure row, same accepted count, same
place in the beat cycle, and the scheduled operations that already fired are not
fired again. The file is written under a temporary name and renamed, because the
failure it exists to survive is the host dying mid-write.
`resume.json` is operational state, never evidence. `timeline.jsonl` remains the
append-only record and nothing here rewrites it. A finished run deletes its
`resume.json`, which makes the file's presence mean exactly one thing: there is
an unfinished run in this directory. The harness refuses to start a fresh
campaign in a directory that already holds one, because two campaigns
interleaved in one timeline are worse evidence than none.
## Timeouts, and why they are an option rather than a constant
How long a turn takes belongs to the inference host, not to the application. On
the reference host a turn cost 229-291 seconds at the recommended window, which
fits comfortably inside the 600-second model timeout this harness used to
hard-code. A slower host does not, and the consequence was not a slow run: a
turn that overran the timeout raised out of the loop and ended the run with a
traceback and no summary.
`--turn-timeout` sets what the application will wait for one narrator reply, and
the harness waits longer still, so that the application's own error arrives
inside the stream rather than being cut off at the socket. The default is
deliberately generous; measure your host before lowering it.
`--max-consecutive-failures` ends a run that has stopped producing turns, with
its evidence and a summary written and `--resume` still able to continue it,
instead of spinning against a narrator that is not answering.
"""
from __future__ import annotations
@@ -55,11 +92,52 @@ ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
#: What the application will wait for one narrator reply, in seconds. The
#: settings schema bounds this at 30..3600 (`app/schemas.py`), and this default
#: sits high inside that range deliberately: an overrun turn is a lost turn, and
#: over a hundred of them the cost of waiting is far below the cost of a restart.
DEFAULT_TURN_TIMEOUT = 1800
#: How much longer the harness waits than the application does. The application
#: has to be the thing that times out, because it reports the failure inside the
#: stream the harness is reading; a harness that gave up first would record a
#: socket error and throw away what the application was about to say.
HARNESS_TIMEOUT_MARGIN = 300
#: Turns in a row not accepted before the run stops and writes what it has. One
#: rejected turn is ordinary — `failed_call` causes one on purpose. A run of
#: them means the narrator is gone, and every further attempt costs a timeout.
DEFAULT_MAX_CONSECUTIVE_FAILURES = 5
#: Operational state, not evidence. Its presence means an unfinished run.
RESUME_FILE = "resume.json"
#: The planted clue. Distinctive enough that its presence anywhere is
#: unambiguous, and phrased as something a story would actually establish.
CLUE = "the silver key opens the crypt beneath the Old Abbey"
CLUE_SENTINEL = "SILVER-KEY-CRYPT-OLD-ABBEY"
#: The clue as accepted state, which is the half of M04 that does not depend on
#: the narrator remembering anything.
#:
#: The content goes in `value`. `add_fact` requires `predicate` and accepts
#: `subject`, `object`, `value` and `fact_id` — and nothing else
#: (`app/narrative/events.py` SPECS). An earlier version of this harness put the
#: clue in a `detail` key, which that event does not define: the correction was
#: accepted, the fact was created, and the clue text went nowhere. The stored
#: fact said only that Aldric knows of the abbey, so `_recall`'s
#: `fact_still_in_state` could not answer True however well the application
#: behaved. A check that can only fail is worse than no check, and this is the
#: second harness defect of that shape M11 has found.
CLUE_FACT = {
"type": "add_fact",
"subject": "aldric",
"predicate": "knows",
"object": "abbey",
"value": f"{CLUE} ({CLUE_SENTINEL})",
"fact_id": "silver-key-opens-crypt",
}
CANON = [
"The dead do not return. No rite, relic or bargain has ever returned anyone.",
"The abbey crypt has been sealed since the founding.",
@@ -115,12 +193,15 @@ def free_port() -> int:
class Storyteller:
"""The real application, started the way `DEVELOPMENT.md` says to start it."""
def __init__(self, db_path: Path, log: Path):
def __init__(self, db_path: Path, log: Path, *,
turn_timeout: int = DEFAULT_TURN_TIMEOUT):
self.db_path = db_path
self.port = free_port()
self.log_path = log
self.proc = None
self.starts = 0
self.turn_timeout = turn_timeout
self.stream_timeout = turn_timeout + HARNESS_TIMEOUT_MARGIN
def start(self) -> None:
self.starts += 1
@@ -176,7 +257,7 @@ class Storyteller:
body = response.read().decode()
return json.loads(body) if body else None
def stream(self, path: str, payload, timeout=900) -> list[dict]:
def stream(self, path: str, payload, timeout=None) -> list[dict]:
"""A turn. The reply is SSE, and a failed turn is an event, not a status.
`app/sse.py`: "A failed turn is still an HTTP 200 response, because the
@@ -184,6 +265,8 @@ class Storyteller:
harness that read the status code would call every failure a success —
which is precisely the class of harness defect M8's review warned about.
"""
if timeout is None:
timeout = self.stream_timeout
request = urllib.request.Request(
f"http://127.0.0.1:{self.port}/api{path}",
data=json.dumps(payload).encode(), method="POST",
@@ -204,13 +287,26 @@ class Storyteller:
class Run:
"""One long campaign, and everything measured about it."""
def __init__(self, server: Storyteller, out: Path):
def __init__(self, server: Storyteller, out: Path, *, turns_target: int,
turn_timeout: int = DEFAULT_TURN_TIMEOUT):
self.server = server
self.out = out
self.timeline = (out / "timeline.jsonl").open("a")
self.adv = 0
self.accepted = 0
self.events: list[dict] = []
self.turns_target = turns_target
self.turn_timeout = turn_timeout
#: Where the beat cycle stands. Carried across a resume, so a continued
#: campaign keeps moving rather than replaying its first ten beats.
self.beat = 0
#: Plan keys whose scheduled operation has already fired, keyed by turn
#: number rather than by step name: two of the steps are called `retry`
#: and two `restart`, so a name does not identify one.
self.completed_steps: set[int] = set()
self.resumed = False
self.elapsed_before = 0.0
self.session_started = time.monotonic()
# ------------------------------------------------------------ recording
@@ -221,17 +317,76 @@ class Run:
self.timeline.write(json.dumps(entry, default=str) + "\n")
self.timeline.flush()
# ------------------------------------------------------------- resuming
def elapsed(self) -> int:
"""Seconds of run time, across every session this campaign has had."""
return round(self.elapsed_before + (time.monotonic() - self.session_started))
def save_resume(self) -> None:
"""Checkpoint enough to pick this campaign up again, atomically."""
payload = {
"adventure": self.adv,
"accepted": self.accepted,
"beat": self.beat,
"completed_steps": sorted(self.completed_steps),
"server_starts": self.server.starts,
"elapsed_seconds": self.elapsed(),
"turns_target": self.turns_target,
"written": datetime.now().isoformat(timespec="seconds"),
}
tmp = self.out / (RESUME_FILE + ".tmp")
tmp.write_text(json.dumps(payload, indent=2))
tmp.replace(self.out / RESUME_FILE)
def adopt(self, prior: dict) -> None:
"""Take on the state a previous session checkpointed.
A checkpoint can be at most one turn behind the database, because a turn
is committed by the application before this file is written. Erring that
way costs one extra turn on a campaign that wants *at least* a hundred,
which is the harmless direction.
"""
self.adv = prior["adventure"]
self.accepted = prior["accepted"]
self.beat = prior.get("beat", 0)
self.completed_steps = set(prior.get("completed_steps") or [])
self.elapsed_before = prior.get("elapsed_seconds", 0)
self.resumed = True
def reattach(self) -> None:
"""Prove the campaign is still there, and restate what a run needs.
Settings are re-applied rather than trusted. They live in the database,
and `failed_call` deliberately points the model at a name the server
does not serve before putting it back; a host that died inside that
window left the campaign configured to fail every turn it is given.
"""
page = self.server.call("GET", f"/adventures/{self.adv}/actions?limit=1")
self.apply_settings()
self.note("resumed", adventure=self.adv, accepted=self.accepted,
beat=self.beat, actions_in_db=page["total"],
completed_steps=sorted(self.completed_steps),
process_starts=self.server.starts,
elapsed_before_seconds=round(self.elapsed_before))
# ------------------------------------------------------------- campaign
def setup(self) -> None:
def apply_settings(self) -> dict:
settings = self.server.call("PUT", "/settings", {
"endpoint_url": ENDPOINT, "model": MODEL,
"embedding_model": EMBED_MODEL, "context_token_budget": 16384,
"max_output_tokens": 500, "model_timeout_seconds": 600,
"max_output_tokens": 500,
"model_timeout_seconds": self.turn_timeout,
"memory_top_k": 4,
})
self.note("settings", model=settings["model"],
budget=settings["context_token_budget"])
budget=settings["context_token_budget"],
model_timeout_seconds=settings["model_timeout_seconds"])
return settings
def setup(self) -> None:
self.apply_settings()
created = self.server.call("POST", "/adventures", {
"title": "Continuity Test (M11 long run)",
@@ -308,6 +463,13 @@ class Run:
self.note("turn_failed", text=text, status=exc.code,
detail=exc.read().decode()[:300])
return {"accepted": False}
except Exception as exc: # noqa: BLE001
# A socket timeout, or a connection dropped mid-stream. This used to
# propagate out of the loop and end the run with a traceback and no
# summary, which on a slow host is the likeliest way to lose one.
self.note("turn_failed", text=text,
detail=f"{type(exc).__name__}: {exc}"[:300])
return {"accepted": False}
errors = [e for e in events if e.get("type") == "error"]
if errors:
self.note("turn_error", text=text, detail=errors[0].get("detail", "")[:300])
@@ -334,6 +496,15 @@ class Run:
return {
"total_actions": report["history"]["total"],
"history_included": report["history"]["included"],
# Where the history window was cut, and the step it takes when it
# moves. A `floor_depth` that is the same on two consecutive turns
# is the prompt's prefix having been preserved, which is the whole
# of what trimming the window in blocks buys; a run that recorded
# neither could not say whether it engaged, held or stepped, and a
# continuity finding could not be attributed. `.get` because a
# campaign resumed against an older build has neither.
"history_floor_depth": report["history"].get("floor_depth"),
"history_trim_block": report["history"].get("trim_block"),
"prompt_tokens": tokens["total"],
"budget": tokens["budget"],
"configured_budget": tokens.get("configured_budget"),
@@ -370,77 +541,217 @@ def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--turns", type=int, default=100)
parser.add_argument("--out", required=True)
parser.add_argument(
"--resume", action="store_true",
help="continue the unfinished run in --out rather than starting one")
parser.add_argument(
"--turn-timeout", type=int, default=DEFAULT_TURN_TIMEOUT,
help=("seconds the application waits for one narrator reply "
f"(30-3600, default {DEFAULT_TURN_TIMEOUT})"))
parser.add_argument(
"--max-consecutive-failures", type=int,
default=DEFAULT_MAX_CONSECUTIVE_FAILURES,
help="stop and write the evidence after this many unaccepted turns")
args = parser.parse_args()
if not (ENDPOINT and MODEL):
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
return 2
if not 30 <= args.turn_timeout <= 3600:
# The settings schema's own bound, checked here so a mistyped timeout
# fails in the first second rather than on the first PUT.
print("--turn-timeout must be between 30 and 3600 seconds")
return 2
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
db_path = out / "campaign.db"
server = Storyteller(db_path, out / "server.log")
resume_path = out / RESUME_FILE
prior = _resume_state(resume_path, out / "timeline.jsonl", args.resume)
if isinstance(prior, str):
print(prior)
return 2
server = Storyteller(db_path, out / "server.log",
turn_timeout=args.turn_timeout)
if prior:
server.starts = prior.get("server_starts", 1)
server.start()
run = Run(server, out)
started = datetime.now()
run = Run(server, out, turns_target=args.turns,
turn_timeout=args.turn_timeout)
if prior:
run.adopt(prior)
try:
run.setup()
if run.resumed:
run.reattach()
else:
run.setup()
# ---- The planted clue, at the very beginning. ----
run.turn(f"I tell Mara quietly that {CLUE} — {CLUE_SENTINEL}.")
run.correct([
{"type": "add_fact", "subject": "aldric", "predicate": "knows",
"object": "abbey", "detail": f"{CLUE} ({CLUE_SENTINEL})",
"fact_id": "silver-key-opens-crypt"},
], note="the planted clue, as accepted state")
run.note("clue_planted", sentinel=CLUE_SENTINEL)
# ---- The planted clue, at the very beginning. ----
run.turn(f"I tell Mara quietly that {CLUE} — {CLUE_SENTINEL}.")
run.correct([CLUE_FACT], note="the planted clue, as accepted state")
# Proved here, at turn one, where it costs a single request. M04
# asks whether the clue is still recoverable a hundred turns later,
# and that question is meaningless if it was never stored — so a
# run whose clue did not land is stopped rather than spending hours
# measuring nothing. See CLUE_FACT for how this went wrong before.
planted = any(CLUE_SENTINEL in json.dumps(fact) for fact
in run.state()["document"].get("facts") or [])
run.note("clue_planted", sentinel=CLUE_SENTINEL,
verified_in_state=planted)
if not planted:
raise SystemExit(
"the planted clue is not in accepted state, so M04 cannot "
"be measured from this run. Stopping before the campaign "
"starts rather than reporting a recall failure later.")
# The first checkpoint, and the point from which --resume works: the
# campaign exists and its clue is planted.
run.save_resume()
# ---- The long middle. ----
plan = _schedule(args.turns)
beat = 0
consecutive_failures = 0
aborted = None
while run.accepted < args.turns:
step = plan.get(run.accepted + 1)
if step:
at = run.accepted + 1
step = plan.get(at)
if step is not None and at not in run.completed_steps:
try:
_do_step(run, server, step)
except Exception as exc: # noqa: BLE001
# A step that fails is a finding, not a reason to lose the
# other ninety turns. It is recorded loudly and the campaign
# goes on, because an abandoned run proves nothing at all.
run.note("step_failed", step=step,
run.note("step_failed", step=step, at=at,
error=f"{type(exc).__name__}: {exc}"[:300])
run.turn(BEATS[beat % len(BEATS)])
beat += 1
# Fired, however it went. Before this the step was keyed only by
# the accepted count, so a refused turn afterwards ran the whole
# operation again — and a restart or a retry performed twice is
# not the test the schedule describes.
run.completed_steps.add(at)
run.save_resume()
result = run.turn(BEATS[run.beat % len(BEATS)])
run.beat += 1
if result.get("accepted"):
consecutive_failures = 0
else:
consecutive_failures += 1
if consecutive_failures >= args.max_consecutive_failures:
aborted = (
f"{consecutive_failures} turns in a row were not accepted; "
"the narrator is not answering. Stopping with the evidence "
"written, so --resume can carry this campaign on."
)
run.note("run_aborted", reason=aborted,
accepted_turns=run.accepted)
run.save_resume()
break
run.save_resume()
# ---- M04: the recall check, with controls. ----
run.note("recall_begin")
recall = _recall(run)
(out / "recall.json").write_text(json.dumps(recall, indent=2))
# Skipped on an aborted run: it asks the narrator a question, and the
# reason the run stopped is that the narrator does not answer.
recall = None
if aborted is None:
run.note("recall_begin")
recall = _recall(run)
(out / "recall.json").write_text(json.dumps(recall, indent=2))
# ---- Export the whole thing, for the recovery evidence. ----
bundle = server.call("GET", f"/adventures/{run.adv}/export")
(out / "bundle.json").write_text(json.dumps(bundle))
run.note("exported", bytes=len((out / "bundle.json").read_bytes()))
# ---- Export whatever exists, for the recovery evidence. ----
# Attempted even for an aborted run: the recovery check and the storage
# numbers are worth having at whatever turn count was reached.
bundle_bytes = None
try:
bundle = server.call("GET", f"/adventures/{run.adv}/export")
(out / "bundle.json").write_text(json.dumps(bundle))
bundle_bytes = len((out / "bundle.json").read_bytes())
run.note("exported", bytes=bundle_bytes)
except Exception as exc: # noqa: BLE001
run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300])
summary = {
"status": "aborted" if aborted else "complete",
"aborted_reason": aborted,
"accepted_turns": run.accepted,
"turns_requested": args.turns,
"restarts": server.starts - 1,
"elapsed_seconds": round((datetime.now() - started).total_seconds()),
"process_starts": server.starts,
"resumed": run.resumed,
"elapsed_seconds": run.elapsed(),
"turn_timeout_seconds": args.turn_timeout,
"recall": recall,
"final_state": run.state()["document"],
"final_measurement": run.measure(),
"final_state": _or_none(lambda: run.state()["document"]),
"final_measurement": _or_none(run.measure),
"db_bytes": db_path.stat().st_size,
"bundle_bytes": bundle_bytes,
}
(out / "summary.json").write_text(json.dumps(summary, indent=2, default=str))
print(json.dumps({k: v for k, v in summary.items()
if k not in ("final_state",)}, indent=2, default=str)[:2000])
return 0
if aborted is None:
# A finished run has nothing to resume, and the file's absence is
# what lets a later run use this directory.
resume_path.unlink(missing_ok=True)
return 0
return 1
finally:
server.stop()
run.timeline.close()
def _or_none(read):
"""A summary field worth having when it can be read, and worth skipping when
it cannot. An aborted run still reports the fields that do answer."""
try:
return read()
except Exception as exc: # noqa: BLE001
return {"unavailable": f"{type(exc).__name__}: {exc}"[:200]}
def _resume_state(resume_path: Path, timeline_path: Path, resuming: bool):
"""The previous session's checkpoint, or a message saying why there is none.
Returns the parsed checkpoint, `None` for a clean start, or a string to
print before exiting. The refusals matter as much as the resume: two
campaigns interleaved in one `timeline.jsonl` and one `campaign.db` are
worse evidence than none, and that is what starting fresh on top of an
existing run produces.
A non-empty `timeline.jsonl` is what says a previous attempt got far enough
to record something, and it is the harness's own artifact rather than the
application's. Testing it rather than `campaign.db` means a run that died
before it recorded anything — a server that never came up, a missing
endpoint — leaves the directory usable, because nothing was written into it
to collide with.
"""
if resuming:
if not resume_path.exists():
return (f"--resume: there is no {resume_path.name} in "
f"{resume_path.parent}. Either the run never reached its "
"first checkpoint, or this is the wrong directory.")
try:
prior = json.loads(resume_path.read_text())
except (OSError, json.JSONDecodeError) as exc:
return f"--resume: cannot read {resume_path}: {exc}"
if not prior.get("adventure"):
return f"--resume: {resume_path} names no campaign."
return prior
if resume_path.exists():
return (f"{resume_path} exists, so this directory holds an unfinished "
"run. Pass --resume to carry it on, or choose a new --out.")
if timeline_path.exists() and timeline_path.stat().st_size > 0:
return (f"{timeline_path} already has entries but there is no "
f"{resume_path.name}, so a previous run recorded something here "
"and then died before its first checkpoint. Choose a new --out: "
"starting here would put a second campaign in the same timeline "
"and the same database.")
return None
def _schedule(turns: int) -> dict:
"""Where each required history operation happens. Spread, not clustered."""
unit = max(1, turns // 13)
+20
View File
@@ -331,6 +331,26 @@ export default function Settings() {
/>
<p className="field-hint">In tokens. Must fit your model’s window.</p>
</div>
<div className="field">
<label htmlFor="ctx-window">The window your server enforces</label>
<input
id="ctx-window" type="number" min="256" max="200000"
placeholder="ask the server"
value={settings.context_window_override ?? ''}
onChange={(e) =>
setField(
'context_window_override',
e.target.value === '' ? null : Number(e.target.value),
)}
/>
<p className="field-hint">
Leave this empty and the window is read from the server. Ollama can
be asked; others — vLLM, llama.cpp’s own server — cannot, and then
nothing caps the budget above. Set it here and prompts are capped to
it. A window the server does report always wins over this, and a
number you type is never treated as verified.
</p>
</div>
</div>
<div className="field">
+72
View File
@@ -0,0 +1,72 @@
/* The context window an operator declares, in the Settings screen.
*
* `contextwindow.py` asks the inference server what window the model gets, over
* Ollama's *native* API. Nothing restricts the endpoint to Ollama: against vLLM
* or llama.cpp's own server there is nothing to ask, the window comes back
* unverified, and the story budget is capped to nothing at all. This field is
* where the operator states what they launched the server with.
*
* The one thing worth a test rather than an eye: **empty must mean unset**.
* `Number('')` is 0, and 0 would be rejected by the schema's floor of 256 while
* reading, to anyone looking at the row, like a real declaration of zero.
*/
import { fireEvent, screen, waitFor, within } from '@testing-library/react'
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { api } from './api'
import Settings from './pages/Settings.jsx'
import { SETTINGS, mockModelStatus, renderWith } from './test/helpers'
const FIELD = /window your server enforces/i
async function renderSettings(overrides = {}) {
const row = { ...SETTINGS, ...overrides }
mockModelStatus(api, { settings: row })
vi.spyOn(api, 'updateSettings').mockResolvedValue(row)
await renderWith(<Settings />)
return await screen.findByLabelText(FIELD)
}
/** The Save that belongs to this field's own section: the page has several. */
function saveFor(field) {
return within(field.closest('.settings-section'))
.getByRole('button', { name: /^save$/i })
}
beforeEach(() => {
vi.restoreAllMocks()
})
describe('the declared context window', () => {
it('is empty when nobody has declared one', async () => {
const field = await renderSettings({ context_window_override: null })
expect(field.value).toBe('')
})
it('shows a declaration that has been made', async () => {
const field = await renderSettings({ context_window_override: 8192 })
expect(field.value).toBe('8192')
})
it('sends the number that was typed', async () => {
const field = await renderSettings({ context_window_override: null })
fireEvent.change(field, { target: { value: '8192' } })
fireEvent.click(saveFor(field))
await waitFor(() =>
expect(api.updateSettings).toHaveBeenCalledWith(
expect.objectContaining({ context_window_override: 8192 }),
))
})
it('sends null when cleared, never zero', async () => {
const field = await renderSettings({ context_window_override: 8192 })
fireEvent.change(field, { target: { value: '' } })
fireEvent.click(saveFor(field))
await waitFor(() =>
expect(api.updateSettings).toHaveBeenCalledWith(
expect.objectContaining({ context_window_override: null }),
))
const sent = api.updateSettings.mock.calls.at(-1)[0]
expect(sent.context_window_override).not.toBe(0)
})
})
+3
View File
@@ -23,6 +23,9 @@ export const SETTINGS = {
max_output_tokens: 400,
context_token_budget: 8000,
model_timeout_seconds: 600,
// Null is the shipped default: nobody has declared a window, so the server is
// asked for one. See `contextwindow.py`.
context_window_override: null,
summary_model: '',
embedding_model: 'nomic-embed-text',
memory_bank_capacity: 200,
+5 -4
View File
@@ -38,8 +38,9 @@ verified, and awaits independent review** (2026-09-07).
`reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. It closed the
context-window release blocker M8 found, fixed two defects the validation itself
surfaced, and disposed of all four post-M8 playtest findings. **It is not
accepted**, there is no release tag, and the tree is staged rather than
committed.
accepted**, and there is no release tag. The work is committed and signed on
`m11-release-validation` (`fedb714`); `main` remains M10, the last *accepted*
milestone.
**After M11 there is no further planned milestone.** What follows is
independent review and the v1 acceptance decision, which is the repository
@@ -328,8 +329,8 @@ v1 acceptance the repository owner's decision
**One milestone at a time. Do not begin a milestone before its brief exists.**
**Every planned milestone is now implemented.** M11 is verified and staged, and
the next action is not another milestone: it is an independent review of
**Every planned milestone is now implemented.** M11 is verified and committed,
and the next action is not another milestone: it is an independent review of
`reports/M11-IMPLEMENTATION-REPORT.md` against the acceptance contract, and then
the owner's v1 acceptance decision. **Do not begin post-v1 work before that
decision**, and do not treat M11's own report as the acceptance record.
+69
View File
@@ -1190,6 +1190,75 @@ What this does not do is change the window. That is an operator action — a mod
with `num_ctx` baked in, or `OLLAMA_CONTEXT_LENGTH` — and `DEVELOPMENT.md` says
how. What the application owes the reader is not to lie about it.
**As implemented: the server that cannot be asked.** Discovery above speaks
Ollama's *native* API, and nothing restricts `endpoint_url` to Ollama — any
allowed address serving an OpenAI-compatible `/v1` is accepted. On vLLM, on
llama.cpp's own server, on anything else, `/api/ps` and `/api/show` are not
there. Discovery fails as designed and the window is unverified, which is honest
and leaves the rule above unenforced: the budget stands, and a server with a
smaller window drops the oldest tokens exactly as before. The rule names Ollama
because Ollama is what M8 measured; the failure it forbids is not Ollama's.
`Settings.context_window_override` is the operator stating the window because
they know how they launched the server. It is consulted **only where discovery
left a hole**, in this order:
| Discovery | Override | Result |
| --- | --- | --- |
| verified | any | the verified window; a declaration cannot raise it |
| unverified | set | the declared number, source `declared`, and the prompt is capped |
| unverified | unset | unknown, and the configured budget stands |
The first row is the safety property: an operator may lower an unknown ceiling
into existence and may never raise a known one, so the override cannot become a
route back to over-budgeting a server that already answered.
`verified` keeps its narrow meaning — *the server answered* — so
`window_verified` in a turn's provenance still counts what §15.2 says it counts
and a declaration cannot inflate it. `enforceable` is the separate question of
whether there is a number to cap to at all, and that is what the builder and the
`capped` flag use. A declared window is therefore enforced and identifiable as a
declaration everywhere it appears, and the connection test says plainly that
nothing has checked it against the server.
### 15.3 The history window moves in blocks (post-M11)
§15.2 makes the window a ceiling. This is about what happens at that ceiling.
An inference server caches a prompt by its **prefix**. A history window that
gives up its oldest action every turn changes the prompt near the front, which
discards the cache and makes the server re-read nearly all of it every turn. The
builder's window did exactly that, and it is why a long campaign cost roughly the
same per turn however little had changed since the last one.
`builder.history_floor` snaps the oldest included action's `depth` forward to a
multiple of `builder.trim_block` and holds it there. The window then steps: a
run of turns that re-use the cache, then one that pays to re-read.
Two properties make it safe rather than merely fast, and both are pinned by
tests:
- **It only ever drops more.** The kept window is a suffix of what "whatever
fits" would have kept, so §15.2's ceiling and M03's bound are unweakened.
- **The block comes from configuration, not from the story.** `trim_block` is
derived from the history budget and `max_output_tokens`. A block size measured
from the sizes of recent actions would move the boundary it defines, and a
boundary that moves is the thing this exists to stop.
`depth` is the coordinate because it is stable per action and already branch-
scoped. Rows without one — anything predating the story tree — are not trimmed,
and behave exactly as they did.
The cost is history depth: right after a step the window holds up to a block
fewer actions than the budget allows. `TRIM_FRACTION` bounds that at a quarter of
the window and is the dial between recent history and speed. Measured at an
8,192-token budget: 124.0s per turn against 362.4s with the floor disabled, the
saving growing with the block and therefore with the budget.
**This is a performance change and nothing above it is a requirement.** M11 §P.1
records that no performance requirement exists and declines to invent one; that
still holds. What changed is the cost of a turn, not what a turn must contain.
## 16. Database Direction
SQLite remains the selected v1 authoritative store.
+44 -3
View File
@@ -1,8 +1,49 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v3.7
- **Revision date:** 2026-09-07
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance.
- **Package:** Adventure Storyteller Planning Package v3.8
- **Revision date:** 2026-09-10
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. M01, the 100-turn campaign, is the one REQUIRED test still outstanding.
## v3.8 — Two gaps closed under M11's own rules, and the cost of a long turn (2026-09-10)
No milestone, no requirement change, and no acceptance claim. Three pieces of
work done while M01 was still outstanding: one hole in §15.2's guarantee, one
stale statement of fact, and the reason a long campaign cost the same per turn
however little had changed.
| Document | Change | Kind |
| --- | --- | --- |
| `TECHNICAL-DESIGN.md` | **§15.2 gains "the server that cannot be asked"** — discovery speaks Ollama's native API, nothing restricts the endpoint to Ollama, and on any other server the window goes unverified and the budget uncapped. `Settings.context_window_override` lets an operator state it; a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. | as-implemented record |
| `TECHNICAL-DESIGN.md` | **New §15.3** — the history window moves in blocks. What happens *at* §15.2's ceiling: a window that gave up its oldest action every turn changed the prompt near the front and cost a full re-read every turn. Records the two safety properties and that M11 §P.1's "no performance requirement" still stands. | as-implemented record |
| `planning/README.md` | **Corrected**: said M11's tree was "staged rather than committed" in two places. It was committed and signed on `m11-release-validation` (`fedb714`); `main` is still M10. | correction |
| `README.md`, `DEVELOPMENT.md` | The window override and how to set it; why a long campaign is not slow in proportion to its length, with the measurements. | developer docs |
**The window a server will not tell you.** §15.2 makes the application ask the
inference server what it will accept and cap itself to the answer. It asks over
Ollama's *native* API — and nothing restricts `endpoint_url` to Ollama. Against
vLLM, llama.cpp's own server, or anything else serving an OpenAI-compatible
`/v1`, `/api/ps` and `/api/show` are simply absent, discovery fails as designed,
and the budget stands uncapped at whatever is configured. §15.2's rule names
Ollama because Ollama is what M8 measured; the failure it forbids is not
Ollama's. The override closes that without weakening what `verified` claims:
`verified` still means the server answered, so `window_verified` in a turn's
provenance counts what M11's report says it counts, and a declared window is
identifiable as a declaration everywhere it appears.
**What a long turn was paying for.** An inference server caches a prompt by its
prefix. The history window gave up its oldest action every turn, which changed
the prompt near the front and discarded that cache, so nearly the whole prompt
was reprocessed every turn regardless of how little had changed. The window now
snaps to a block and holds, stepping every few turns. Measured against the
reference deployment on real builder output, at an 8,192-token budget: **124.0s
per turn against 362.4s** with the floor disabled. The cost is history depth —
up to a block fewer actions right after a step — bounded by `TRIM_FRACTION` at a
quarter of the window, which is the dial between recent history and speed.
**Requirement changes: zero.** Nothing was retired, relaxed or reclassified.
§15.3 records explicitly that M11 §P.1 declines to set a performance requirement
and that this does not invent one: what changed is the cost of a turn, not what
a turn must contain.
## v3.7 — M11 implemented: v1 security, long-run and release validation (2026-09-07)