Stop re-reading the whole prompt every turn, and let a lost run carry on

M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
JesseMarkowitz
2026-09-10 06:13:55 -04:00
co-authored by Claude Opus 5
parent fedb7144d0
commit ef25b0a876
22 changed files with 1654 additions and 75 deletions
+20
View File
@@ -331,6 +331,26 @@ export default function Settings() {
/>
<p className="field-hint">In tokens. Must fit your model’s window.</p>
</div>
<div className="field">
<label htmlFor="ctx-window">The window your server enforces</label>
<input
id="ctx-window" type="number" min="256" max="200000"
placeholder="ask the server"
value={settings.context_window_override ?? ''}
onChange={(e) =>
setField(
'context_window_override',
e.target.value === '' ? null : Number(e.target.value),
)}
/>
<p className="field-hint">
Leave this empty and the window is read from the server. Ollama can
be asked; others — vLLM, llama.cpp’s own server — cannot, and then
nothing caps the budget above. Set it here and prompts are capped to
it. A window the server does report always wins over this, and a
number you type is never treated as verified.
</p>
</div>
</div>
<div className="field">
+72
View File
@@ -0,0 +1,72 @@
/* The context window an operator declares, in the Settings screen.
*
* `contextwindow.py` asks the inference server what window the model gets, over
* Ollama's *native* API. Nothing restricts the endpoint to Ollama: against vLLM
* or llama.cpp's own server there is nothing to ask, the window comes back
* unverified, and the story budget is capped to nothing at all. This field is
* where the operator states what they launched the server with.
*
* The one thing worth a test rather than an eye: **empty must mean unset**.
* `Number('')` is 0, and 0 would be rejected by the schema's floor of 256 while
* reading, to anyone looking at the row, like a real declaration of zero.
*/
import { fireEvent, screen, waitFor, within } from '@testing-library/react'
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { api } from './api'
import Settings from './pages/Settings.jsx'
import { SETTINGS, mockModelStatus, renderWith } from './test/helpers'
const FIELD = /window your server enforces/i
async function renderSettings(overrides = {}) {
const row = { ...SETTINGS, ...overrides }
mockModelStatus(api, { settings: row })
vi.spyOn(api, 'updateSettings').mockResolvedValue(row)
await renderWith(<Settings />)
return await screen.findByLabelText(FIELD)
}
/** The Save that belongs to this field's own section: the page has several. */
function saveFor(field) {
return within(field.closest('.settings-section'))
.getByRole('button', { name: /^save$/i })
}
beforeEach(() => {
vi.restoreAllMocks()
})
describe('the declared context window', () => {
it('is empty when nobody has declared one', async () => {
const field = await renderSettings({ context_window_override: null })
expect(field.value).toBe('')
})
it('shows a declaration that has been made', async () => {
const field = await renderSettings({ context_window_override: 8192 })
expect(field.value).toBe('8192')
})
it('sends the number that was typed', async () => {
const field = await renderSettings({ context_window_override: null })
fireEvent.change(field, { target: { value: '8192' } })
fireEvent.click(saveFor(field))
await waitFor(() =>
expect(api.updateSettings).toHaveBeenCalledWith(
expect.objectContaining({ context_window_override: 8192 }),
))
})
it('sends null when cleared, never zero', async () => {
const field = await renderSettings({ context_window_override: 8192 })
fireEvent.change(field, { target: { value: '' } })
fireEvent.click(saveFor(field))
await waitFor(() =>
expect(api.updateSettings).toHaveBeenCalledWith(
expect.objectContaining({ context_window_override: null }),
))
const sent = api.updateSettings.mock.calls.at(-1)[0]
expect(sent.context_window_override).not.toBe(0)
})
})
+3
View File
@@ -23,6 +23,9 @@ export const SETTINGS = {
max_output_tokens: 400,
context_token_budget: 8000,
model_timeout_seconds: 600,
// Null is the shipped default: nobody has declared a window, so the server is
// asked for one. See `contextwindow.py`.
context_window_override: null,
summary_model: '',
embedding_model: 'nomic-embed-text',
memory_bank_capacity: 200,