v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
ac465ed867
commit
d63804f22e
@@ -80,6 +80,14 @@ that isn't the live one starts a new branch.
|
||||
in settings lets you state the window so the prompt is still capped. A window the server
|
||||
itself reported always wins over that, and a declared one is never reported as verified.
|
||||
|
||||
Since v1.1 the prompt also stops short of that window on purpose. It leaves
|
||||
`max(256, 5% of the window)` tokens free, because your model counts tokens differently from
|
||||
the application, and the v1 evidence came within 23 tokens of the edge. After each turn,
|
||||
the server's own count of what it read is compared with what was sent. A turn the server
|
||||
appears to have truncated is kept, flagged and shown in the context inspector, not left to
|
||||
pass silently. A model that isn't loaded yet, and so cannot report its window, is loaded
|
||||
once before the turn is built, so the first turn of a session gets the real window too.
|
||||
|
||||
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
|
||||
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
|
||||
card used to arrive in front of it as a world fact with no class, no visibility, no source and
|
||||
@@ -179,9 +187,9 @@ that isn't the live one starts a new branch.
|
||||
|
||||
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
|
||||
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
|
||||
removed rather than left standing as a picture of a product that no longer exists. The M4
|
||||
closeout drove the real application in a real browser, so the screens exist and work; taking
|
||||
presentable screenshots of them is a job for the UI pass in M8.
|
||||
removed rather than left standing as a picture of a product that no longer exists. The
|
||||
screens exist and are driven in a real browser by the release harness
|
||||
(`backend/tools/m11_browser.py`). Presentable screenshots of them have not been taken.
|
||||
|
||||
## Quick start
|
||||
|
||||
@@ -310,7 +318,8 @@ development, Vite proxies `/api` to FastAPI.
|
||||
|
||||
## Tests
|
||||
|
||||
1,191 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||||
1,524 backend tests (1,507 run everywhere, 17 need a real local model and skip without one): unit
|
||||
tests plus full HTTP integration through the real turn engine, with
|
||||
the model provider mocked. They run with no route to the Internet, which is a requirement
|
||||
rather than a convenience — an offline claim proved on a machine that has been online once
|
||||
proves nothing. A further handful need a real local model and skip without one; they exist
|
||||
|
||||
Reference in New Issue
Block a user