v1.1: harden context window and narrator protocol boundary

WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-14 16:35:05 -04:00
co-authored by Claude Opus 5
parent ac465ed867
commit d63804f22e
26 changed files with 3903 additions and 61 deletions
+51 -20
View File
@@ -25,6 +25,7 @@ from sqlalchemy.orm import object_session
from .. import contextwindow, derived, models, narrative, summaries, worldstate
from ..knowledge import inject as knowledge_inject
from ..providers.openai_compatible import CHAT_CONTINUE_HINT
from ..knowledge import records as knowledge_records
from . import encoding, history
@@ -106,11 +107,22 @@ BAND_FLOOR_SHARE = 0.5
# Built from the table vendored in `encoding.py`, not fetched: the upstream
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
# is called on every turn.
# M6: added to the configured reply budget when reserving output space. It
# absorbs the section separators added after budgeting and the drift between
# this tokenizer and the serving model's. Fixed rather than proportional: what
# it covers does not grow with the size of the budget.
OUTPUT_SAFETY_MARGIN = 64
#
# v1.1 WP-A1: `OUTPUT_SAFETY_MARGIN = 64` was here. M6 added it to the reply
# budget to absorb two unrelated things, and v1.1 separates them:
#
# * **Text the application adds after pricing.** The separators between
# sections, and `CHAT_CONTINUE_HINT`, which the provider appends to every chat
# request and nothing counted. That is not drift, it is our own text, so it is
# now priced exactly (`transport` below).
# * **The drift between this tokenizer and the narrator's.** That is what the
# 64 tokens were really for, and the v1 evidence showed it was too small. It
# is now `contextwindow.safety_reserve`, sized to the window.
#
#: Story sections that can be joined by `SEPARATOR` after pricing: history,
#: author's note, recent history, summary, lore, memories, state, front memory,
#: length hint, refusals, reminder. Knowledge and history rows price their own.
STORY_SECTION_SLOTS = 11
class ContextOverflow(RuntimeError):
@@ -180,19 +192,19 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
words = min(words, band_ceiling)
floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
tail = (
" Finish the narration and append the state block well inside the limit."
" " + narrative.extract.LENGTH_HINT_TAIL
)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
tail = " Finish the narration and append the state block well inside the limit."
tail = " " + narrative.extract.LENGTH_HINT_TAIL
# State the number as a ceiling, never as a budget. In measurements, the
# wording "keep this turn under about N words" read to the model as a target
@@ -204,7 +216,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
# Both numbers are bounds, and the wording is deliberately asymmetric. The
@@ -216,7 +228,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
# so a terse model reading the same clause stops at the floor rather than at
# forty words.
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
@@ -625,12 +637,19 @@ def build_context(
# truncated turn on a model whose window is the budget
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
#
# The margin covers what is added after this arithmetic — the separators
# between sections, and the difference between our tokenizer's count and the
# serving model's. It is small and fixed rather than proportional, because
# what it absorbs does not scale with the budget.
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
protected = reserved + output_reserve
# v1.1 WP-A1: the reply allocation is exactly the reply cap. The text this
# application adds after pricing — separators, and the chat hint the
# provider appends — is counted as `transport`. What neither can know, the
# narrator's tokenizer disagreeing with `cl100k_base`, is the safety reserve,
# which is sized to the window and taken before any history is chosen.
output_reserve = max(0, settings.max_output_tokens)
separator_tokens = count_tokens(SEPARATOR)
transport = (
separator_tokens * (len(system_sections) + STORY_SECTION_SLOTS)
+ count_tokens(CHAT_CONTINUE_HINT)
)
safety = contextwindow.safety_reserve(budget)
protected = reserved + transport + output_reserve + safety
if protected >= budget:
# Failing here is the point. The alternative — carrying on with a token
# or two of history — builds a prompt that is known to overflow, and
@@ -638,8 +657,9 @@ def build_context(
# gracefully if protected context alone is too large."
raise ContextOverflow(
f"The protected context needs {protected} tokens "
f"({reserved} of prompt plus {output_reserve} reserved for the "
f"reply) but the context budget is {budget}. "
f"({reserved} of prompt, {transport} of formatting, {output_reserve} "
f"reserved for the reply and a {safety}-token safety margin) but the "
f"context budget is {budget}. "
+ (
"That budget is what this server was found to accept, so raising "
"the setting alone will not help — load the model with a larger "
@@ -827,6 +847,7 @@ def build_context(
story_text = SEPARATOR.join(s.text for s in story_sections)
all_sections = [s for s in system_sections if s.text] + story_sections
total_tokens = count_tokens(system_text) + count_tokens(story_text)
report = {
"sections": [
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
@@ -837,13 +858,23 @@ def build_context(
# what the history was actually allowed to spend after everything
# protected was subtracted.
"tokens": {
"total": count_tokens(system_text) + count_tokens(story_text),
"total": total_tokens,
"budget": budget,
"configured_budget": settings.context_token_budget,
"output_reserve": output_reserve,
"protected": reserved,
"available_for_history": available,
"history_spent": spent,
# v1.1 WP-A1. `transport` is the formatting priced in above;
# `estimate` is what this application believes it actually sent,
# the assembled text plus what the provider adds to it, and is what
# the server's own count is compared against after the reply.
"transport": transport,
"safety_reserve": safety,
"estimate": total_tokens + (
count_tokens(CHAT_CONTINUE_HINT) if settings.api_mode != "completion"
else separator_tokens
),
},
# M11: what the server was found to accept, and how. `verified` false
# means nobody could check — the prompt was built to the configured