v1.1: harden context window and narrator protocol boundary

WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-14 16:35:05 -04:00
co-authored by Claude Opus 5
parent ac465ed867
commit d63804f22e
26 changed files with 3903 additions and 61 deletions
@@ -69,6 +69,18 @@ out of tokens partway through it. All of that is removed before the prose is
stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4
has the rules.
**Implementation note (v1.1 WP-A2).** A small model also copies the protocol's
*instructions*: the vocabulary written as calls, the length hint, and the scene
line. The fix is on both sides:
- **Prompt:** the vocabulary is shown in the wire format, and the fixed example
uses genre-neutral placeholders.
- **Extractor:** it recognises those echoes only by strings and names the
application owns.
No event type, field, validation rule or proposal record changed.
`TECHNICAL-DESIGN.md` §15.4 lists the four rules.
### Where it lives
- `adventures.narrative_state` — the current authoritative document. This is
+3 -1
View File
@@ -2,7 +2,9 @@
**This file is the index. Start here.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 planning has begun.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: WP-A1
and WP-A2 are implemented and staged for owner review**
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`).
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
+116
View File
@@ -1230,6 +1230,75 @@ whether there is a number to cap to at all, and that is what the builder and the
declaration everywhere it appears, and the connection test says plainly that
nothing has checked it against the server.
**As implemented (v1.1 WP-A1): a safety reserve, and the server's own count.**
A ceiling in the application's tokens is not a ceiling in the narrator's. The
builder counts with `cl100k_base`, and the v1 evidence left the largest prompts
23-42 real tokens from the edge of a 16,384 window. Past the edge Ollama does not
refuse. Measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt returned
200 with `prompt_tokens` 2,050.
- **The reserve.** `contextwindow.safety_reserve(budget)` is
`max(256, ceil(5% of the effective budget))`, computed with integer rounding
up: 256 at 4,096, 410 at 8,192, 820 at 16,384. It is taken before any history
is chosen. The effective budget is the verified or declared window when there
is one, and the configured budget otherwise. It is fixed and documented, is not
a setting, and is not calibrated per model.
- **What replaced the 64-token margin.** M6's `OUTPUT_SAFETY_MARGIN` absorbed two
unrelated things.
- The application's own text added after pricing: separators between
sections, and the chat hint the provider appends to every request. This is
now priced exactly as `transport`.
- Tokenizer drift. This is now the reserve.
The reply allocation is exactly `max_output_tokens`. Protected context is
`sections + transport + reply + reserve`, and `ContextOverflow` is raised
before the model call when that does not fit.
- **The server's count.** Streaming requests set `stream_options.include_usage`.
Without it Ollama sends no usage, and none of the 514 AI turns in the v1
evidence has one. After the reply, `contextwindow.classify_usage` compares the
server's `prompt_tokens` with `tokens.estimate`: the assembled text plus what
the provider adds.
- **Accounting states.** The turn's snapshot records `accounting`, whose status
is one of the following, checked in this order:
| Status | Meaning |
| --- | --- |
| `unknown` | No positive integer count was reported. It is never read as `fits`. |
| `truncation_suspected` | The server read fewer tokens than the estimate by more than the reserve. |
| `exceeded` | The server's count plus the reply allocation is over the budget. |
| `fits` | Otherwise. |
The record also carries the server's count, the difference, the reserve and
the observed margin (`budget - reply - server count`).
- **Surfacing.** The record is returned on the turn's `done` event, logged as a
warning when it is `exceeded` or `truncation_suspected`, and shown in the
context inspector, where those two statuses are an alert.
- **The turn is kept.** A discrepancy found after the reply is recorded, never
enforced. The narration has already streamed to the reader, and the accepted
turn is not discarded.
- **Accounting belongs to one attempt.** It sits in `attempts.ATTEMPT_KEYS`
beside `usage`. When a retry or a take selection moves the shared prompt,
each take keeps the accounting for its own call.
- **A cold model is loaded before its turn is built (v1.1 A1 corrective).** A
model that is not resident cannot report its window. The A1 evidence caught
exactly that: 13,875 tokens were sent to a server that read 2,050.
`contextwindow.ensure_window` works like this:
- it probes;
- if the window is unverified and the server answered, it makes one bounded
`POST /api/generate` naming only the model, with no prompt. Ollama loads
the model and generates nothing ("done_reason": "load");
- it probes again, bypassing the cache;
- the turn is built to whatever that second probe says.
A failed load, or a window still unverified afterwards, changes nothing: the
configured budget stands and the accounting still catches a cut. The request
goes to the configured endpoint only, under the same policy and TLS path. It
writes nothing, and it is recorded as `window.preflight` in the turn's
snapshot. The context dry run never loads a model.
No schema change: the accounting lives in the snapshot JSON, and a turn from
v1.0.0 simply has none. The bundle format is unchanged, for the same reason.
### 15.3 The history window moves in blocks (post-M11)
§15.2 makes the window a ceiling. This is about what happens at that ceiling.
@@ -1300,6 +1369,53 @@ headings. A lone heading followed by prose stays, and so do JSON a character
typed and a fact restated inside a sentence. That last case is how a narrator can
still carry authoritative state into its prose (M11 report §G.4 and §P).
**As implemented (v1.1 WP-A2): the source and the sink together.** The M11
closeout's identity run stored four shapes the extractor left. All of them were
application text. v1.1 changes both the prompt that taught them and the
extractor that missed them. Every new removal is anchored to something the
application owns, never to what prose looks like.
- **The prompt.**
- `events.vocabulary_for_prompt` shows each event as the object the model
must send (`{"type": "set_possession", "item": "<key>", "owner": "<key>"}`),
not as `set_possession(item, owner)`. The call notation was never the wire
format, and the narrator copied it.
- `EMIT_RULE`'s example uses the placeholders `character-1`, `item-1` and
`location-1`, not the fantasy fixture's `mara`, `silver-key`, `old-abbey` and
`aldric`. The narrator had proposed `silver-key` in an office meeting.
- The length hint's opening and closing words are named constants shared by
the builder and the extractor.
- **The extractor**, rules R1-R4 (`RULE_*` in `narrative/extract.py`):
- **R1:** a whole line that begins with a call to an event in `events.SPECS`,
optionally `>`-quoted. Not inside a fenced code block, not mid-sentence, and
not for a call-shaped name the vocabulary lacks.
- **R2:** a trailing bracket that opens `Hard limit:` and carries the hint's
own wording ("append the state block", or "turn must not exceed *N* words").
- **R3:** the renderer's scene line left as the reply's last line. It is
removed when it ends in the renderer's `(at <location>)`, or when protocol
was already cut from the same reply.
- **R4:** a ```` ```json ```` or bare ```` ``` ```` opener left as the last
line with nothing after it, counted as an opener rather than a closer.
- **Proven on real narration.** Every stored real reply in the v1 evidence was
replayed through the v1.0.0 and v1.1 extractors (`tools/v11_replay_extractor.py`).
Every changed line is attributed to one of the four rules, and a person reviewed
every change. The results are in the WP-A1/A2 report.
- **Deliberately still left:** a fact restated inside a sentence; model-invented
headings; and a bracket that starts `Hard limit:` but carries none of the
application's wording.
- **R5, the echoed instruction tail (v1.1 A2 corrective).** A v1.1 identity turn
ended in a reworded continue hint, "[… Continue the story here, directly.
Output only story text.]". Because nothing recognised it, nothing above it was
ever trailing, and the reminder, a reworded length hint and a scene line all
stayed. R5 makes two changes:
- **The continue hint is recognised by its own sentence.** A trailing bracket
containing "Output only story text" is an echoed instruction.
`CONTINUE_HINT_PHRASE` is pinned by a test to `CHAT_CONTINUE_HINT`.
- **A bracket opening with the length hint's own `Hard limit:` is removed only
directly above an echoed instruction already cut from the same reply's end.**
An in-world "[Hard limit: forty days]" stays when it is the last line, and
when a state block follows it. Any other bracket above an echo stays.
## 16. Database Direction
SQLite remains the selected v1 authoritative store.
+7 -3
View File
@@ -1,8 +1,12 @@
# Adventure Storyteller — v1.1 Plan
**Status:** PLANNING. Written 2026-09-14 on `v1.1-development`, from the signed
v1.0.0 release commit `432f041`. No work package has started, and no v1.1
version or tag exists.
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
signed v1.0.0 release commit `432f041`.
**WP-A1 and WP-A2 are implemented and staged for owner review** (2026-09-14). They
are reported together, and kept separate, in
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`. No other work package has started. No
v1.1 version or tag exists.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
+28 -2
View File
@@ -1,8 +1,34 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v4.1
- **Package:** Adventure Storyteller Planning Package v4.2
- **Revision date:** 2026-09-14
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 planning has begun** on `v1.1-development`: `V1.1-PLAN.md`. No v1.1 work package has started, and no v1.1 version or tag exists.
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): **WP-A1 and WP-A2 are implemented and staged for owner review.** No other work package has started, and no v1.1 version or tag exists.
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
Two v1.1 work packages, implemented in sequence. No requirement or acceptance
test changed, and no schema or bundle format changed. Evidence is in
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `TECHNICAL-DESIGN.md` §15.2 | **As implemented (v1.1 WP-A1)**, covering: the safety reserve, `max(256, ceil(5%))`; what replaced M6's 64-token margin; the server's own count; the four accounting states; and keeping the turn. | as-implemented record |
| `TECHNICAL-DESIGN.md` §15.4 | **As implemented (v1.1 WP-A2)**: the vocabulary shown in the wire format, genre-neutral placeholders, and extractor rules R1-R4, each anchored to application-owned text. | as-implemented record |
| `DECISIONS/013-authoritative-narrative-state-document.md` | An implementation note for v1.1. No event type, field, validation rule or proposal record changed. | as-implemented note |
| `V1.1-PLAN.md` | Status: A1 and A2 implemented and staged. | status |
| `planning/README.md` | Current state. | index |
| `reports/v1.1/V1.1-WP-A1-A2-REPORT.md` | **New.** The combined review package, with A1 and A2 kept separate. | work-package report |
| `README.md`, `DEVELOPMENT.md` | The reserve and the accounting. The stale backend test count and the Screenshots paragraph are corrected. | developer docs |
**Corrective work before commit (owner review, 2026-09-14).** Report §R.
| Document | Change |
| --- | --- |
| `TECHNICAL-DESIGN.md` §15.2 | A cold model is loaded once before its turn is built (`contextwindow.ensure_window`). |
| `TECHNICAL-DESIGN.md` §15.4 | R5: the echoed continue hint is recognised by its own sentence, and the application-opened tail above it is removed. |
| `DEVELOPMENT.md`, `README.md` | The cold-model load, in operator terms. |
**Requirement changes: zero.**
## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14)
File diff suppressed because it is too large Load Diff