Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
co-authored by
Claude Opus 5
parent
fedb7144d0
commit
ef25b0a876
+5
-4
@@ -38,8 +38,9 @@ verified, and awaits independent review** (2026-09-07).
|
||||
`reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. It closed the
|
||||
context-window release blocker M8 found, fixed two defects the validation itself
|
||||
surfaced, and disposed of all four post-M8 playtest findings. **It is not
|
||||
accepted**, there is no release tag, and the tree is staged rather than
|
||||
committed.
|
||||
accepted**, and there is no release tag. The work is committed and signed on
|
||||
`m11-release-validation` (`fedb714`); `main` remains M10, the last *accepted*
|
||||
milestone.
|
||||
|
||||
**After M11 there is no further planned milestone.** What follows is
|
||||
independent review and the v1 acceptance decision, which is the repository
|
||||
@@ -328,8 +329,8 @@ v1 acceptance the repository owner's decision
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
|
||||
**Every planned milestone is now implemented.** M11 is verified and staged, and
|
||||
the next action is not another milestone: it is an independent review of
|
||||
**Every planned milestone is now implemented.** M11 is verified and committed,
|
||||
and the next action is not another milestone: it is an independent review of
|
||||
`reports/M11-IMPLEMENTATION-REPORT.md` against the acceptance contract, and then
|
||||
the owner's v1 acceptance decision. **Do not begin post-v1 work before that
|
||||
decision**, and do not treat M11's own report as the acceptance record.
|
||||
|
||||
@@ -1190,6 +1190,75 @@ What this does not do is change the window. That is an operator action — a mod
|
||||
with `num_ctx` baked in, or `OLLAMA_CONTEXT_LENGTH` — and `DEVELOPMENT.md` says
|
||||
how. What the application owes the reader is not to lie about it.
|
||||
|
||||
**As implemented: the server that cannot be asked.** Discovery above speaks
|
||||
Ollama's *native* API, and nothing restricts `endpoint_url` to Ollama — any
|
||||
allowed address serving an OpenAI-compatible `/v1` is accepted. On vLLM, on
|
||||
llama.cpp's own server, on anything else, `/api/ps` and `/api/show` are not
|
||||
there. Discovery fails as designed and the window is unverified, which is honest
|
||||
and leaves the rule above unenforced: the budget stands, and a server with a
|
||||
smaller window drops the oldest tokens exactly as before. The rule names Ollama
|
||||
because Ollama is what M8 measured; the failure it forbids is not Ollama's.
|
||||
|
||||
`Settings.context_window_override` is the operator stating the window because
|
||||
they know how they launched the server. It is consulted **only where discovery
|
||||
left a hole**, in this order:
|
||||
|
||||
| Discovery | Override | Result |
|
||||
| --- | --- | --- |
|
||||
| verified | any | the verified window; a declaration cannot raise it |
|
||||
| unverified | set | the declared number, source `declared`, and the prompt is capped |
|
||||
| unverified | unset | unknown, and the configured budget stands |
|
||||
|
||||
The first row is the safety property: an operator may lower an unknown ceiling
|
||||
into existence and may never raise a known one, so the override cannot become a
|
||||
route back to over-budgeting a server that already answered.
|
||||
|
||||
`verified` keeps its narrow meaning — *the server answered* — so
|
||||
`window_verified` in a turn's provenance still counts what §15.2 says it counts
|
||||
and a declaration cannot inflate it. `enforceable` is the separate question of
|
||||
whether there is a number to cap to at all, and that is what the builder and the
|
||||
`capped` flag use. A declared window is therefore enforced and identifiable as a
|
||||
declaration everywhere it appears, and the connection test says plainly that
|
||||
nothing has checked it against the server.
|
||||
|
||||
### 15.3 The history window moves in blocks (post-M11)
|
||||
|
||||
§15.2 makes the window a ceiling. This is about what happens at that ceiling.
|
||||
|
||||
An inference server caches a prompt by its **prefix**. A history window that
|
||||
gives up its oldest action every turn changes the prompt near the front, which
|
||||
discards the cache and makes the server re-read nearly all of it every turn. The
|
||||
builder's window did exactly that, and it is why a long campaign cost roughly the
|
||||
same per turn however little had changed since the last one.
|
||||
|
||||
`builder.history_floor` snaps the oldest included action's `depth` forward to a
|
||||
multiple of `builder.trim_block` and holds it there. The window then steps: a
|
||||
run of turns that re-use the cache, then one that pays to re-read.
|
||||
|
||||
Two properties make it safe rather than merely fast, and both are pinned by
|
||||
tests:
|
||||
|
||||
- **It only ever drops more.** The kept window is a suffix of what "whatever
|
||||
fits" would have kept, so §15.2's ceiling and M03's bound are unweakened.
|
||||
- **The block comes from configuration, not from the story.** `trim_block` is
|
||||
derived from the history budget and `max_output_tokens`. A block size measured
|
||||
from the sizes of recent actions would move the boundary it defines, and a
|
||||
boundary that moves is the thing this exists to stop.
|
||||
|
||||
`depth` is the coordinate because it is stable per action and already branch-
|
||||
scoped. Rows without one — anything predating the story tree — are not trimmed,
|
||||
and behave exactly as they did.
|
||||
|
||||
The cost is history depth: right after a step the window holds up to a block
|
||||
fewer actions than the budget allows. `TRIM_FRACTION` bounds that at a quarter of
|
||||
the window and is the dial between recent history and speed. Measured at an
|
||||
8,192-token budget: 124.0s per turn against 362.4s with the floor disabled, the
|
||||
saving growing with the block and therefore with the budget.
|
||||
|
||||
**This is a performance change and nothing above it is a requirement.** M11 §P.1
|
||||
records that no performance requirement exists and declines to invent one; that
|
||||
still holds. What changed is the cost of a turn, not what a turn must contain.
|
||||
|
||||
## 16. Database Direction
|
||||
|
||||
SQLite remains the selected v1 authoritative store.
|
||||
|
||||
+44
-3
@@ -1,8 +1,49 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v3.7
|
||||
- **Revision date:** 2026-09-07
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance.
|
||||
- **Package:** Adventure Storyteller Planning Package v3.8
|
||||
- **Revision date:** 2026-09-10
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance. M01, the 100-turn campaign, is the one REQUIRED test still outstanding.
|
||||
|
||||
## v3.8 — Two gaps closed under M11's own rules, and the cost of a long turn (2026-09-10)
|
||||
|
||||
No milestone, no requirement change, and no acceptance claim. Three pieces of
|
||||
work done while M01 was still outstanding: one hole in §15.2's guarantee, one
|
||||
stale statement of fact, and the reason a long campaign cost the same per turn
|
||||
however little had changed.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `TECHNICAL-DESIGN.md` | **§15.2 gains "the server that cannot be asked"** — discovery speaks Ollama's native API, nothing restricts the endpoint to Ollama, and on any other server the window goes unverified and the budget uncapped. `Settings.context_window_override` lets an operator state it; a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. | as-implemented record |
|
||||
| `TECHNICAL-DESIGN.md` | **New §15.3** — the history window moves in blocks. What happens *at* §15.2's ceiling: a window that gave up its oldest action every turn changed the prompt near the front and cost a full re-read every turn. Records the two safety properties and that M11 §P.1's "no performance requirement" still stands. | as-implemented record |
|
||||
| `planning/README.md` | **Corrected**: said M11's tree was "staged rather than committed" in two places. It was committed and signed on `m11-release-validation` (`fedb714`); `main` is still M10. | correction |
|
||||
| `README.md`, `DEVELOPMENT.md` | The window override and how to set it; why a long campaign is not slow in proportion to its length, with the measurements. | developer docs |
|
||||
|
||||
**The window a server will not tell you.** §15.2 makes the application ask the
|
||||
inference server what it will accept and cap itself to the answer. It asks over
|
||||
Ollama's *native* API — and nothing restricts `endpoint_url` to Ollama. Against
|
||||
vLLM, llama.cpp's own server, or anything else serving an OpenAI-compatible
|
||||
`/v1`, `/api/ps` and `/api/show` are simply absent, discovery fails as designed,
|
||||
and the budget stands uncapped at whatever is configured. §15.2's rule names
|
||||
Ollama because Ollama is what M8 measured; the failure it forbids is not
|
||||
Ollama's. The override closes that without weakening what `verified` claims:
|
||||
`verified` still means the server answered, so `window_verified` in a turn's
|
||||
provenance counts what M11's report says it counts, and a declared window is
|
||||
identifiable as a declaration everywhere it appears.
|
||||
|
||||
**What a long turn was paying for.** An inference server caches a prompt by its
|
||||
prefix. The history window gave up its oldest action every turn, which changed
|
||||
the prompt near the front and discarded that cache, so nearly the whole prompt
|
||||
was reprocessed every turn regardless of how little had changed. The window now
|
||||
snaps to a block and holds, stepping every few turns. Measured against the
|
||||
reference deployment on real builder output, at an 8,192-token budget: **124.0s
|
||||
per turn against 362.4s** with the floor disabled. The cost is history depth —
|
||||
up to a block fewer actions right after a step — bounded by `TRIM_FRACTION` at a
|
||||
quarter of the window, which is the dial between recent history and speed.
|
||||
|
||||
**Requirement changes: zero.** Nothing was retired, relaxed or reclassified.
|
||||
§15.3 records explicitly that M11 §P.1 declines to set a performance requirement
|
||||
and that this does not invent one: what changed is the cost of a turn, not what
|
||||
a turn must contain.
|
||||
|
||||
## v3.7 — M11 implemented: v1 security, long-run and release validation (2026-09-07)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user