Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
co-authored by
Claude Opus 5
parent
fedb7144d0
commit
ef25b0a876
@@ -344,6 +344,36 @@ subsystem comes back as a route, if an API key becomes settable again, if the
|
||||
model timeout stops being configurable or becomes unbounded, or if a supported
|
||||
start path stops binding loopback.
|
||||
|
||||
## Why a long campaign is not slow in proportion to its length
|
||||
|
||||
An inference server caches the prompt it has already processed, keyed on the
|
||||
**prefix**. While a story only grows at the end, each turn re-uses that cache and
|
||||
pays for its own new tokens alone. Once the context budget is full, though, the
|
||||
history window has to give something up — and a window that gives up its *oldest*
|
||||
action every turn changes the prompt near the front, which throws the cache away
|
||||
and makes the server re-read almost the whole thing, every turn.
|
||||
|
||||
So the window moves in blocks. `context/builder.py` snaps the oldest included
|
||||
action to a boundary and holds it there for several turns, then steps. Measured
|
||||
against the reference deployment on real builder output, at an 8,192-token budget:
|
||||
|
||||
| | Per turn |
|
||||
| --- | --- |
|
||||
| Window held, story grew by one action | 14-20 s |
|
||||
| Window stepped (one turn in three) | 333-338 s |
|
||||
| **Mean over whole cycles** | **124.0 s** |
|
||||
| Window sliding every turn, as before | 362.4 s |
|
||||
|
||||
The cost is history depth: right after a step the window holds up to a block
|
||||
fewer actions than the budget would allow. `TRIM_FRACTION` bounds that at a
|
||||
quarter of the window, and it is the one number to change if you would rather
|
||||
trade recent history for speed, or the reverse.
|
||||
|
||||
The saving grows with the block, and the block grows with the budget — so the
|
||||
larger the context window, the more this is worth. `history["floor_depth"]` and
|
||||
`history["trim_block"]` are in every context report, and a `floor_depth` that is
|
||||
the same on two consecutive turns is the prompt's prefix having been preserved.
|
||||
|
||||
## The release-validation harnesses
|
||||
|
||||
M11 added six runnable harnesses under `backend/tools/`. They are the evidence
|
||||
@@ -370,6 +400,12 @@ AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 \
|
||||
AIDND_TEST_MODEL=<model> AIDND_TEST_EMBED_MODEL=<embedding model> \
|
||||
.venv/bin/python -m tools.m11_long_run --turns 100 --out "$HOME/m11-evidence/m01"
|
||||
|
||||
# The same campaign, carried on after a crash, a reboot or a Ctrl-C. It picks up
|
||||
# the adventure the checkpoint names, keeps its place in the beat cycle, and does
|
||||
# not fire a scheduled operation that already fired.
|
||||
AIDND_TEST_ENDPOINT=... AIDND_TEST_MODEL=... AIDND_TEST_EMBED_MODEL=... \
|
||||
.venv/bin/python -m tools.m11_long_run --turns 100 --resume --out "$HOME/m11-evidence/m01"
|
||||
|
||||
# What that campaign is worth on a machine that has never seen it (I01-I07).
|
||||
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"
|
||||
|
||||
@@ -389,6 +425,35 @@ AIDND_TEST_MODEL=<model> AIDND_TEST_EMBED_MODEL=<embedding model> \
|
||||
.venv/bin/python -m tools.contrast_audit
|
||||
```
|
||||
|
||||
### Resuming the long run, and timing it out
|
||||
|
||||
A hundred turns is hours of wall clock, and the first release attempt lost one at
|
||||
turn 97 to a host crash. The harness now checkpoints `resume.json` into `--out`
|
||||
after the prologue, after every scheduled operation and after every turn, and
|
||||
`--resume` continues from it. The file is written under a temporary name and
|
||||
renamed, so a crash during the write cannot leave a half-parsed one.
|
||||
|
||||
`resume.json` is operational state rather than evidence: `timeline.jsonl` stays
|
||||
the append-only record, a resumed session appends to it, and a finished run
|
||||
deletes its `resume.json`. That makes the file's presence mean exactly one
|
||||
thing — there is an unfinished run in this directory — and the harness refuses
|
||||
to start a fresh campaign on top of one, because two campaigns interleaved in a
|
||||
single timeline and database are worse evidence than none. It refuses a
|
||||
directory holding a `campaign.db` with no checkpoint for the same reason.
|
||||
|
||||
How long a turn takes is the inference host's characteristic, not the
|
||||
application's, so the timeout is an option rather than a constant:
|
||||
|
||||
| Flag | Default | What it does |
|
||||
| --- | --- | --- |
|
||||
| `--turn-timeout` | 1800 | Seconds the application waits for one narrator reply — it becomes `model_timeout_seconds`, so the settings schema's 30..3600 bound applies. The harness waits 300s longer, so the application's own error arrives inside the stream rather than being cut off at the socket. |
|
||||
| `--max-consecutive-failures` | 5 | Unaccepted turns in a row before the run stops, writes `summary.json` with `status: aborted`, and leaves a `resume.json` that `--resume` can carry on. |
|
||||
|
||||
Measure your host before lowering `--turn-timeout`. On the M11 reference
|
||||
deployment a turn cost 229-291 seconds at the recommended window; a slower host
|
||||
can exceed the 600 seconds this harness used to hard-code, and an overrun turn is
|
||||
a lost turn.
|
||||
|
||||
`tools/m11_webdriver.py` is the W3C WebDriver client the browser harness uses.
|
||||
It exists so browser evidence needs no Selenium in the dependency surface, and
|
||||
it documents the one environment quirk that matters here: a snap Firefox will
|
||||
@@ -521,6 +586,25 @@ reports the window it found or says plainly that it could not check.
|
||||
That does not make the window *bigger*, and the rest of this section is still
|
||||
how you do that.
|
||||
|
||||
**On a server that is not Ollama, tell the application the window yourself.**
|
||||
The check above uses Ollama's *native* API, which vLLM, llama.cpp's own server
|
||||
and the rest do not serve — so the window comes back unverified and the budget
|
||||
is left at whatever is configured. Set **`context_window_override`** in settings
|
||||
to the window you launched that server with:
|
||||
|
||||
```bash
|
||||
curl -X PUT http://127.0.0.1:8000/api/settings \
|
||||
-H 'Content-Type: application/json' -d '{"context_window_override": 8192}'
|
||||
```
|
||||
|
||||
Prompts are then capped to it. It is used *only* when the server could not be
|
||||
asked — a window the server did report always wins, so this can never be a way
|
||||
to over-budget an Ollama that answered — and it does not count as verification:
|
||||
the turn's provenance still records that nothing checked the number. Send
|
||||
`null` to remove it. Nothing here validates the figure against the server, so an
|
||||
override larger than the real window puts you back to silent truncation; take it
|
||||
from how you started the server, not from the model card.
|
||||
|
||||
**Setting it per request does not work from this application.** Ollama's
|
||||
OpenAI-compatible endpoint accepts `num_ctx` — nested in `options` or at the top
|
||||
level — returns HTTP 200 and ignores it. Worse, it *reloads the model at its own
|
||||
|
||||
Reference in New Issue
Block a user