v1.1: harden context window and narrator protocol boundary

WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-14 16:35:05 -04:00
co-authored by Claude Opus 5
parent ac465ed867
commit d63804f22e
26 changed files with 3903 additions and 61 deletions
+8 -1
View File
@@ -46,7 +46,14 @@ from .narrative import model as narrative_model
# token accounting. Each attempt is its own API call, and a retry is the call
# most likely to read the prompt back out of cache. Everything else in a snapshot
# is the prompt, which is assembled once per turn.
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage")
#
# v1.1 WP-A1: `accounting` is one attempt's too. It compares the server's count
# for *that* call with the turn's estimate. Left out of this tuple, it was
# treated as part of the shared prompt, so moving the live flag handed the
# superseded attempt's accounting to the new live one and threw the new one's
# away. Found by the A2 long run: two retries and one take selection left three
# attempts reporting no accounting, or another attempt's.
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage", "accounting")
# ------------------------------------------------------------------ reading
+51 -20
View File
@@ -25,6 +25,7 @@ from sqlalchemy.orm import object_session
from .. import contextwindow, derived, models, narrative, summaries, worldstate
from ..knowledge import inject as knowledge_inject
from ..providers.openai_compatible import CHAT_CONTINUE_HINT
from ..knowledge import records as knowledge_records
from . import encoding, history
@@ -106,11 +107,22 @@ BAND_FLOOR_SHARE = 0.5
# Built from the table vendored in `encoding.py`, not fetched: the upstream
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
# is called on every turn.
# M6: added to the configured reply budget when reserving output space. It
# absorbs the section separators added after budgeting and the drift between
# this tokenizer and the serving model's. Fixed rather than proportional: what
# it covers does not grow with the size of the budget.
OUTPUT_SAFETY_MARGIN = 64
#
# v1.1 WP-A1: `OUTPUT_SAFETY_MARGIN = 64` was here. M6 added it to the reply
# budget to absorb two unrelated things, and v1.1 separates them:
#
# * **Text the application adds after pricing.** The separators between
# sections, and `CHAT_CONTINUE_HINT`, which the provider appends to every chat
# request and nothing counted. That is not drift, it is our own text, so it is
# now priced exactly (`transport` below).
# * **The drift between this tokenizer and the narrator's.** That is what the
# 64 tokens were really for, and the v1 evidence showed it was too small. It
# is now `contextwindow.safety_reserve`, sized to the window.
#
#: Story sections that can be joined by `SEPARATOR` after pricing: history,
#: author's note, recent history, summary, lore, memories, state, front memory,
#: length hint, refusals, reminder. Knowledge and history rows price their own.
STORY_SECTION_SLOTS = 11
class ContextOverflow(RuntimeError):
@@ -180,19 +192,19 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
words = min(words, band_ceiling)
floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
tail = (
" Finish the narration and append the state block well inside the limit."
" " + narrative.extract.LENGTH_HINT_TAIL
)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
tail = " Finish the narration and append the state block well inside the limit."
tail = " " + narrative.extract.LENGTH_HINT_TAIL
# State the number as a ceiling, never as a budget. In measurements, the
# wording "keep this turn under about N words" read to the model as a target
@@ -204,7 +216,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
# Both numbers are bounds, and the wording is deliberately asymmetric. The
@@ -216,7 +228,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
# so a terse model reading the same clause stops at the floor rather than at
# forty words.
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
@@ -625,12 +637,19 @@ def build_context(
# truncated turn on a model whose window is the budget
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
#
# The margin covers what is added after this arithmetic — the separators
# between sections, and the difference between our tokenizer's count and the
# serving model's. It is small and fixed rather than proportional, because
# what it absorbs does not scale with the budget.
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
protected = reserved + output_reserve
# v1.1 WP-A1: the reply allocation is exactly the reply cap. The text this
# application adds after pricing — separators, and the chat hint the
# provider appends — is counted as `transport`. What neither can know, the
# narrator's tokenizer disagreeing with `cl100k_base`, is the safety reserve,
# which is sized to the window and taken before any history is chosen.
output_reserve = max(0, settings.max_output_tokens)
separator_tokens = count_tokens(SEPARATOR)
transport = (
separator_tokens * (len(system_sections) + STORY_SECTION_SLOTS)
+ count_tokens(CHAT_CONTINUE_HINT)
)
safety = contextwindow.safety_reserve(budget)
protected = reserved + transport + output_reserve + safety
if protected >= budget:
# Failing here is the point. The alternative — carrying on with a token
# or two of history — builds a prompt that is known to overflow, and
@@ -638,8 +657,9 @@ def build_context(
# gracefully if protected context alone is too large."
raise ContextOverflow(
f"The protected context needs {protected} tokens "
f"({reserved} of prompt plus {output_reserve} reserved for the "
f"reply) but the context budget is {budget}. "
f"({reserved} of prompt, {transport} of formatting, {output_reserve} "
f"reserved for the reply and a {safety}-token safety margin) but the "
f"context budget is {budget}. "
+ (
"That budget is what this server was found to accept, so raising "
"the setting alone will not help — load the model with a larger "
@@ -827,6 +847,7 @@ def build_context(
story_text = SEPARATOR.join(s.text for s in story_sections)
all_sections = [s for s in system_sections if s.text] + story_sections
total_tokens = count_tokens(system_text) + count_tokens(story_text)
report = {
"sections": [
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
@@ -837,13 +858,23 @@ def build_context(
# what the history was actually allowed to spend after everything
# protected was subtracted.
"tokens": {
"total": count_tokens(system_text) + count_tokens(story_text),
"total": total_tokens,
"budget": budget,
"configured_budget": settings.context_token_budget,
"output_reserve": output_reserve,
"protected": reserved,
"available_for_history": available,
"history_spent": spent,
# v1.1 WP-A1. `transport` is the formatting priced in above;
# `estimate` is what this application believes it actually sent,
# the assembled text plus what the provider adds to it, and is what
# the server's own count is compared against after the reply.
"transport": transport,
"safety_reserve": safety,
"estimate": total_tokens + (
count_tokens(CHAT_CONTINUE_HINT) if settings.api_mode != "completion"
else separator_tokens
),
},
# M11: what the server was found to accept, and how. `verified` false
# means nobody could check — the prompt was built to the configured
+208
View File
@@ -86,6 +86,7 @@ one.
from __future__ import annotations
import logging
import math
import re
import time
from dataclasses import dataclass
@@ -132,6 +133,11 @@ class Window:
model_max: int | None = None
#: Why the window is unknown, or how it was found. Shown to the user.
detail: str = ""
#: v1.1: the server answered a discovery request at all, whatever it said.
#: A server that answered but could not report a window may simply not have
#: the model loaded yet, which `ensure_window` can fix; one that did not
#: answer cannot be helped by asking it to load anything.
reachable: bool = False
@property
def verified(self) -> bool:
@@ -180,6 +186,116 @@ def effective_budget(configured: int, window: Window | int | None) -> int:
return min(configured, tokens)
#: v1.1 WP-A1: the tokens kept free below the effective window, beyond the reply.
#:
#: The builder counts with `cl100k_base`; the narrator counts with its own
#: tokenizer. The v1 evidence put the largest prompts 23-42 real tokens from the
#: edge of a 16,384 window, and Ollama does not refuse a prompt past the edge —
#: measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt came back 200
#: with `prompt_tokens` 2,050. So the reserve is deliberate and sized to the
#: window: the larger of a floor and a share, **rounded up to a whole token**.
#:
#: 4,096 -> 256 8,192 -> 410 16,384 -> 820
#:
#: A fixed, documented tolerance, owner-chosen for v1.1. It is not a setting and
#: it is not calibrated per model.
SAFETY_RESERVE_FLOOR = 256
SAFETY_RESERVE_PERCENT = 5
def safety_reserve(effective_window: int) -> int:
"""`max(256, ceil(5% of the effective window))`, in tokens.
The effective window is the budget the prompt is actually built to — the
verified or declared window when there is one, the configured budget
otherwise — so a 16,384 setting against a 4,096 server reserves 256, not 820.
Integer arithmetic, so the rounding is exact rather than a float's.
"""
share = math.ceil(max(0, effective_window) * SAFETY_RESERVE_PERCENT / 100)
return max(SAFETY_RESERVE_FLOOR, share)
#: v1.1 WP-A1: what the server's own count says about a turn that was sent.
FITS = "fits"
EXCEEDED = "exceeded"
TRUNCATION_SUSPECTED = "truncation_suspected"
#: `UNKNOWN` above: the server reported no usable count.
def classify_usage(usage: dict | None, *, estimate: int, budget: int,
max_output_tokens: int, window_verified: bool) -> dict:
"""Sets the server's reported prompt count against what the application sent.
The order of the checks is the order of what they prove:
``unknown``
No positive integer `prompt_tokens`. Nothing can be said, and nothing
is claimed: an absent count is never read as a prompt that fitted.
``truncation_suspected``
The server read fewer tokens than were sent by more than the safety
reserve. A tokenizer thriftier than `cl100k_base` may honestly count a
little less; a shortfall larger than the tolerance the application keeps
for drift is the signature of a server that cut the prompt — the real
shape was 6,316 sent and 2,050 read.
``exceeded``
The server's count plus the reply allocation is more than the window
the prompt was built for. The drift was larger than the whole reserve,
so the reply may have been cut short.
``fits``
Otherwise.
`observed_margin` is what was left beside the reply by the server's count:
`budget - max_output_tokens - server_prompt_tokens`. The safety reserve is
the tolerance, so a margin between 0 and the reserve is still `fits`.
A discrepancy is recorded, never acted on: the reply has already streamed
to the reader and is accepted story.
"""
prompt = usage.get("prompt_tokens") if isinstance(usage, dict) else None
reserve = safety_reserve(budget)
verified_note = "" if window_verified else (
" The window itself was not verified for this turn.")
record = {
"status": UNKNOWN,
"server_prompt_tokens": None,
"estimate": estimate,
"difference": None,
"budget": budget,
"max_output_tokens": max_output_tokens,
"safety_reserve": reserve,
"observed_margin": None,
"window_verified": bool(window_verified),
"detail": "",
}
if type(prompt) is not int or prompt <= 0:
record["detail"] = ("The server reported no prompt token count, so nothing "
"confirms the whole prompt was read." + verified_note)
return record
record["server_prompt_tokens"] = prompt
record["difference"] = prompt - estimate
record["observed_margin"] = budget - max_output_tokens - prompt
if prompt + reserve < estimate:
record["status"] = TRUNCATION_SUSPECTED
record["detail"] = (
f"The server read {prompt:,} prompt tokens of the {estimate:,} sent, a "
f"shortfall larger than the {reserve:,}-token safety reserve. A server "
"that cuts an over-window prompt reports exactly this, and what it cuts "
"is the start: the narrator's rules and the canon." + verified_note)
elif prompt + max_output_tokens > budget:
record["status"] = EXCEEDED
record["detail"] = (
f"The server counted {prompt:,} prompt tokens; with {max_output_tokens:,} "
f"for the reply that is more than the {budget:,}-token window the prompt "
"was built for, so the reply may have been cut short." + verified_note)
else:
record["status"] = FITS
record["detail"] = (
f"The server read {prompt:,} prompt tokens, leaving "
f"{record['observed_margin']:,} beside the reply." + verified_note)
return record
def cache_clear() -> None:
"""Forgets what was learned. Called when the endpoint or model changes."""
_cache.clear()
@@ -223,6 +339,7 @@ def _declared_or(declared: int | None, discovered: Window) -> Window:
declared, DECLARED, discovered.model_max,
f"{declared:,} tokens, declared in settings — the server was not able "
f"to say ({discovered.detail})",
reachable=discovered.reachable,
)
@@ -266,6 +383,7 @@ async def _ask(endpoint_url: str, model: str) -> Window:
return Window(
tokens, LOADED, ceiling,
f"{tokens:,} tokens, reported by the running model",
reachable=True,
)
return await _declared_window(client, base, model)
except (httpx.HTTPError, ValueError, TypeError, KeyError) as exc:
@@ -293,6 +411,7 @@ async def _declared_window(client, base: str, model: str) -> Window:
return Window(
None, UNKNOWN,
detail=f"the server did not describe the model (HTTP {resp.status_code})",
reachable=True,
)
body = resp.json() or {}
ceiling = _architecture_ceiling(body.get("model_info") or {})
@@ -304,14 +423,103 @@ async def _declared_window(client, base: str, model: str) -> Window:
"the model sets no num_ctx, so the server will load it at its own "
"default — which is 4,096 where there is no VRAM"
),
reachable=True,
)
tokens = min(declared, ceiling) if ceiling else declared
return Window(
tokens, PARAMETERS, ceiling,
f"{tokens:,} tokens, from the model's own num_ctx",
reachable=True,
)
#: v1.1 WP-A1 corrective: loading the configured model so its window can be read.
#:
#: The first real turn of the A1 evidence found a cold model: `/api/ps` knew
#: nothing, `/api/show` found no `num_ctx`, so the window was unverified and the
#: prompt was built to the configured 16,384. Ollama loaded the model at its own
#: 4,096 default, kept 2,050 of 13,875 tokens and answered 200. That case is
#: preventable, because the window becomes readable the moment the model is
#: resident. Ollama's native `POST /api/generate` with a model and **no prompt**
#: loads the model and generates nothing — measured on Ollama 0.33: HTTP 200,
#: `"response": ""`, `"done_reason": "load"`, and `/api/ps` then reported the
#: window. The OpenAI-compatible request that followed did not reload it.
WARM_PATH = "/api/generate"
async def warm(endpoint_url: str, model: str, *, timeout: float) -> tuple[bool, str]:
"""Asks the configured server to load `model`. One request, no story text.
Held to the same endpoint policy and TLS trust as inference and the probe, and
sent to the same host the probe asks. The body names the model and nothing
else: no prompt, so nothing is generated, and no `options` or `keep_alive`, so
the model loads the way the server would load it for the turn itself.
Returns `(loaded, detail)`. Every failure is `(False, why)` and never raises:
a server that will not load the model on request will fail the turn's own
call the ordinary way, which is where that failure belongs.
"""
reason = endpoints.rejection_reason(endpoint_url)
if reason is not None:
return False, f"endpoint not allowed — {reason}"
base = native_base(endpoint_url)
try:
async with httpx.AsyncClient(
timeout=httpx.Timeout(timeout, connect=CONNECT_TIMEOUT),
verify=tlstrust.ssl_context(),
) as client:
resp = await client.post(f"{base}{WARM_PATH}", json={"model": model})
except httpx.HTTPError as exc:
log.debug("model warm-up failed for %s: %s", base, exc)
return False, f"could not ask the server to load the model ({type(exc).__name__})"
if resp.status_code != 200:
return False, f"the server did not load the model (HTTP {resp.status_code})"
try:
body = resp.json() or {}
except ValueError:
return False, "the server answered the load request with something that was not JSON"
return True, f"the server loaded the model ({body.get('done_reason') or 'done'})"
async def ensure_window(endpoint_url: str, model: str, *, declared: int | None = None,
warm_timeout: float = 300.0) -> tuple[Window, dict]:
"""The window for a turn about to be generated, loading the model once if that is what it takes.
1. Probe as before.
2. If the window is not verified, the server answered, and there is a model to
load: one bounded `warm` request.
3. If the model loaded, probe again, bypassing the cache that still holds the
unverified answer.
Whatever the second probe says is the answer. There is no retry loop, no
guessed window, and no hard-coded 4,096: a window still unverified leaves the
configured budget standing, exactly as before, and the turn's accounting
still catches a server that cut the prompt.
Returns the window and a `preflight` record for the turn's provenance.
Not used by the context dry run: loading a model is a side effect, and
opening a panel should not cause one.
"""
window = await probe(endpoint_url, model, declared=declared)
preflight = {"attempted": False, "loaded": None, "verified_before": window.verified,
"verified_after": window.verified, "detail": ""}
if window.verified:
preflight["detail"] = "the window was already verified"
return window, preflight
if not (endpoint_url and model):
preflight["detail"] = "no endpoint or model configured"
return window, preflight
if not window.reachable:
preflight["detail"] = "the server did not answer, so no model was loaded"
return window, preflight
loaded, detail = await warm(endpoint_url, model, timeout=warm_timeout)
preflight.update(attempted=True, loaded=loaded, detail=detail)
if loaded:
window = await probe(endpoint_url, model, declared=declared, use_cache=False)
preflight["verified_after"] = window.verified
return window, preflight
def _num_ctx(parameters) -> int | None:
"""Reads `num_ctx` out of the plain-text parameter block Ollama returns."""
if not isinstance(parameters, str):
+25 -4
View File
@@ -35,6 +35,8 @@ creating a second, empty Mara.
from __future__ import annotations
import json
# Field types the schema layer enforces. Kept deliberately small: a narrative
# state event carries names, labels and plain values, and nothing here needs a
# nested structure a model could hide something inside.
@@ -184,11 +186,30 @@ def vocabulary_for_prompt() -> str:
Generated from `SPECS` rather than written out beside it, so the model can
never be told about an event the application does not implement — the drift
that would produce proposals rejected for reasons nobody could see.
v1.1 WP-A2: each event is shown as the object the model must put in the
`events` list, with its required fields, not as `name(field, …)`. The call
notation was never the wire format, and a 3B narrator copied it into its
prose as `> set_possession(silver-key, "alice")`. An object copied into prose
is a proposal the extractor already recognises and removes; a call is not.
"""
lines = []
for name, definition in SPECS.items():
fields = list(definition["required"]) + [
f"{field}?" for field in definition["optional"]
]
lines.append(f' {name}({", ".join(fields)}) — {definition["summary"]}')
shape = {"type": name}
for field, kind in definition["required"].items():
shape[field] = _PLACEHOLDER[kind]
body = json.dumps(shape, ensure_ascii=False, separators=(",", ":"))
line = f" {body} — {definition['summary']}"
if definition["optional"]:
line += f" (optional: {', '.join(definition['optional'])})"
lines.append(line)
return "\n".join(lines)
#: What a field of each kind looks like in the prompt's vocabulary. Placeholders,
#: never example identifiers, so the vocabulary names nothing a story could copy.
#: A list field is shown as a list, so the model is told its shape; every other
#: field is an ellipsis. Measured: `"<key>"`-style placeholders with spaced
#: separators cost 456 tokens against v1.0.0's 258; this form costs about 380,
#: and every line is still the object the model must send.
_PLACEHOLDER = {KEY: "…", TEXT: "…", VALUE: "…", LABELS: ["…"]}
+188 -14
View File
@@ -37,17 +37,17 @@ EMIT_RULE = (
"appeared, record it.\n"
"\n"
"Every value is ABSOLUTE — the new state of things, never a change or a "
"difference. Use only these events:\n"
"difference. Use only these events, in exactly this shape:\n"
f"{events.vocabulary_for_prompt()}\n"
"\n"
"Identifiers are short lower-case slugs (mara, silver-key, old-abbey) and must "
"match the ones already in the state you were shown. Introduce a person, place "
"or thing with create_entity before referring to it. If the turn established "
"nothing, send an empty events list.\n"
"Identifiers are short lower-case slugs and must match the ones already in the "
"state you were shown; the example's identifiers are placeholders. Introduce a "
"person, place or thing with create_entity before referring to it. If the turn "
"established nothing, send an empty events list.\n"
"Example:\n"
'```state\n'
'{"events": [{"type": "set_possession", "item": "silver-key", "owner": "aldric"},'
' {"type": "set_current_location", "entity": "aldric", "location": "old-abbey"}]}\n'
'{"events": [{"type": "set_possession", "item": "item-1", "owner": "character-1"},'
' {"type": "set_current_location", "entity": "character-1", "location": "location-1"}]}\n'
'```'
)
@@ -58,6 +58,47 @@ EMIT_REMINDER = (
"nothing changed.]"
)
# v1.1 WP-A2: the length hint's own words, named once. `builder.length_hint`
# builds the hint from these, and the extractor recognises an echo of it by
# them, so the two cannot drift apart.
LENGTH_HINT_OPENING = "[Hard limit:"
LENGTH_HINT_TAIL = "Finish the narration and append the state block well inside the limit."
#: The application's wording inside a hint. A 3B narrator reworded the front
#: ("your next turn") and the end ("This story ends here."), and kept one or the
#: other of these every time.
_LENGTH_HINT_PHRASE_RE = re.compile(
r"append the state block|turn must not exceed \d+ words", re.IGNORECASE
)
#: v1.1 WP-A2: the rules that remove protocol a narrator copied, named so the
#: replay tool and the report can say which removed what.
RULE_EVENT_CALL = "event_call_line"
RULE_LENGTH_HINT = "echoed_length_hint"
RULE_SCENE_LINE = "rendered_scene_line"
RULE_EMPTY_FENCE = "empty_dangling_fence"
RULE_INSTRUCTION_TAIL = "echoed_instruction_tail"
#: v1.1 WP-A2 corrective (R5). The sentence `CHAT_CONTINUE_HINT` in
#: `providers/openai_compatible.py` carries, which a narrator echoed with the rest
#: of the hint reworded around it. Kept as a copy rather than an import, so the
#: narrative package does not depend on the provider; a test pins that the
#: hint still contains it.
CONTINUE_HINT_PHRASE = "Output only story text"
# R1. A whole line opening with a call to an event this protocol has. The names
# come from the vocabulary, so a call-shaped line naming anything else — a
# character's `open_door(north)` — is not matched.
_EVENT_CALL_LINE_RE = re.compile(
r"^[ \t]*(?:>[ \t]*)?(?:"
+ "|".join(re.escape(name) for name in events.SPECS)
+ r")[ \t]*\(",
re.IGNORECASE,
)
# R3. The renderer's scene line carries its location this way.
_RENDERED_SCENE_LOCATION_RE = re.compile(r"\(at [^()\n]+\)\s*$")
# R4. An opener with nothing after it.
_EMPTY_FENCE_LINE_RE = re.compile(r"```(?:json)?[ \t]*", re.IGNORECASE)
# Three patterns, and the difference between them is the whole of this module's
# safety. A story is allowed to contain code, and taking a code block out of
# someone's prose is a worse failure than leaving a stray proposal in it.
@@ -128,11 +169,42 @@ def _is_echoed_instruction(inner: str) -> bool:
# opening words only, because the echo is often cut off before it ends.
if low.lstrip().startswith("continue the story directly"):
return True
# v1.1 WP-A2 corrective (R5): the same hint, reworded at the front. The M11
# closeout-era identity re-run stored "[You don't need to continue; … Continue
# the story here, directly. Output only story text.]" as the last line of a
# reply, and because nothing recognised it, nothing above it was trailing.
if CONTINUE_HINT_PHRASE.lower() in low:
return True
# v1.1 WP-A2 (R2): the length hint, which names "state block" but not
# "events list", so it passed every check above.
if _is_length_hint(inner):
return True
# The reminder names both; prose about the protocol rarely names either the
# way the instruction does, and effectively never both.
return "state block" in low and "events list" in low
def _opens_like_length_hint(inner: str) -> bool:
"""R5. The bracket opens with the length hint's own `Hard limit:`, whatever follows.
Never enough on its own: an in-world "[Hard limit: forty days]" opens the same
way. `_clean` takes it only directly above an echoed instruction it has already
removed from the end of the same reply.
"""
return inner.lstrip().lower().startswith(LENGTH_HINT_OPENING[1:].lower())
def _is_length_hint(inner: str) -> bool:
"""Whether a bracket's contents are `builder.length_hint`, however reworded.
It must open the way the hint opens *and* carry the hint's own wording. An
in-world "Hard limit: forty days" has the opening and none of the wording.
"""
opening = LENGTH_HINT_OPENING[1:].lower()
return (inner.lstrip().lower().startswith(opening)
and bool(_LENGTH_HINT_PHRASE_RE.search(inner)))
# A heading the model writes above a block it did not fence: `State`, sometimes
# as `State:`, `**State**` or `### State`. It is removed only in two places:
# directly above a proposal that is removed, and as the last line of the reply.
@@ -145,7 +217,7 @@ _LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
def _clean(prose: str) -> str:
def _clean(prose: str, *, after_block: bool = False) -> str:
"""Removes protocol the block extraction could not, and nothing else.
Found by the M5 realistic-context run (§12), which is the failure class
@@ -160,9 +232,23 @@ def _clean(prose: str) -> str:
story after it. Stored text is replayed as history, so every leak also
showed the next prompt a second, older account of the state, which is what
M5 review Finding 4 removed from replayed history.
v1.1 WP-A2 added four shapes, from the M11 closeout's identity run and the
v1 corpus, each anchored to something the application owns rather than to
what prose looks like: a line opening with a vocabulary call (R1), the
length hint echoed at the end (R2), the renderer's scene line left last
(R3), and an empty fence opener left last (R4). `after_block` says a
proposal block was already taken out of this reply, which is what lets R3
remove a bare scene line that sat above it.
"""
cleaned, _found = _inline_proposals(prose)
cleaned, calls_removed = _strip_event_call_lines(prose)
cleaned, _found = _inline_proposals(cleaned)
cleaned = _strip_echoed_state(cleaned)
protocol_cut = after_block or calls_removed
# R5: set once an echoed instruction bracket has come off the end. Only then
# may a bracket that merely opens the way the length hint opens be taken as
# part of the same echoed tail.
instruction_cut = False
# The end of the reply is cut until nothing more comes off, because one kind
# of leftover can hide another. In a real reply, a `State` heading sat above
# a block the model never finished, and a parroted reminder sat above an
@@ -171,7 +257,12 @@ def _clean(prose: str) -> str:
before = cleaned
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
bracket = pattern.search(cleaned)
if bracket is not None and _is_echoed_instruction(bracket.group(1)):
if bracket is None:
continue
if _is_echoed_instruction(bracket.group(1)):
cleaned = cleaned[: bracket.start()]
instruction_cut = True
elif instruction_cut and _opens_like_length_hint(bracket.group(1)):
cleaned = cleaned[: bracket.start()]
cleaned = _DANGLING_STATE_RE.sub("", cleaned)
dangling = _DANGLING_JSON_RE.search(cleaned)
@@ -182,10 +273,93 @@ def _clean(prose: str) -> str:
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
# A bare quote marker, the start of a quoted block that never came.
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
cleaned = _strip_empty_dangling_fence(cleaned)
if cleaned.rstrip() != before.rstrip():
protocol_cut = True
cleaned = _strip_trailing_scene_line(cleaned, protocol_cut)
if cleaned == before:
return cleaned.strip()
def _strip_event_call_lines(text: str) -> tuple[str, bool]:
"""R1. Removes whole lines that open with a call to a vocabulary event.
A line inside a fenced code block is the story's own code and is never
examined. Returns the text and whether anything was removed.
"""
kept: list[str] = []
in_fence = False
removed = False
for line in text.split("\n"):
if line.lstrip().startswith("```"):
in_fence = not in_fence
kept.append(line)
continue
if not in_fence and _EVENT_CALL_LINE_RE.match(line):
removed = True
continue
kept.append(line)
if not removed:
return text, False
return re.sub(r"\n{3,}", "\n\n", "\n".join(kept)), True
def _strip_empty_dangling_fence(text: str) -> str:
"""R4. A ```` ```json ```` or ```` ``` ```` opener as the last line, with nothing after it.
Only an *opener*: the fence lines are counted, and an even count means the
last one closes a story's own code block, which stays.
"""
lines = text.rstrip().split("\n")
if len(lines) < 2 or not _EMPTY_FENCE_LINE_RE.fullmatch(lines[-1].strip()):
return text
fences = sum(1 for line in lines if line.lstrip().startswith("```"))
if fences % 2 == 0:
return text
return "\n".join(lines[:-1]).rstrip()
def _strip_trailing_scene_line(text: str, protocol_cut: bool) -> str:
"""R3. The renderer's scene line, left as the last line of the reply.
Taken when it carries the renderer's own `(at <location>)`, or when protocol
was already cut from this reply, which makes a bare scene line part of the
same pasted tail. A final screenplay-style "Scene: …" line in a reply with
no protocol in it stays, and so does any scene line with story after it.
"""
lines = text.rstrip().split("\n")
if len(lines) < 2:
return text
last = lines[-1].strip()
if not last.startswith(render.HEADING_SCENE + " "):
return text
if not (_RENDERED_SCENE_LOCATION_RE.search(last) or protocol_cut):
return text
return "\n".join(lines[:-1]).rstrip()
def explain_removed_line(line: str) -> str | None:
"""Which v1.1 rule removes a line of this shape, for the replay report.
None means no v1.1 rule explains it, which the replay treats as a failure.
"""
stripped = line.strip()
if _EVENT_CALL_LINE_RE.match(line):
return RULE_EVENT_CALL
if stripped.startswith("["):
inner = stripped[1:]
inner = inner[:-1] if inner.endswith("]") else inner
if _is_length_hint(inner):
return RULE_LENGTH_HINT
if _is_echoed_instruction(inner) or _opens_like_length_hint(inner):
return RULE_INSTRUCTION_TAIL
if stripped.startswith(render.HEADING_SCENE + " "):
return RULE_SCENE_LINE
if _EMPTY_FENCE_LINE_RE.fullmatch(stripped):
return RULE_EMPTY_FENCE
return None
def _is_state_heading(line: str) -> bool:
return bool(_STATE_HEADING_RE.match(line))
@@ -451,7 +625,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
if matches:
match = matches[-1]
raw = match.group(1).strip()
prose = _clean(text[: match.start()] + text[match.end():])
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
return prose, _tolerant_load(raw), raw
# A `json` or unlabelled fence is ours only when its contents are this
@@ -465,7 +639,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
raw = match.group(1).strip()
parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
prose = _clean(text[: match.start()] + text[match.end():])
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
return prose, parsed, raw
match = _TRAILING_RE.search(text)
@@ -473,7 +647,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
raw = match.group(1)
parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed):
return _clean(text[: match.start()]), parsed, raw
return _clean(text[: match.start()], after_block=True), parsed, raw
# An unfenced proposal on its own lines but not at the end: quoted, or
# followed by more story. The last one is the turn's proposal, as with
@@ -481,7 +655,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
without, found = _inline_proposals(text)
if found:
parsed, raw = found[-1]
return _clean(without), parsed, raw
return _clean(without, after_block=True), parsed, raw
# No block at all — but the reply may still carry protocol the model wrote
# as prose, or a fence it never closed.
+16 -4
View File
@@ -26,6 +26,10 @@ EMBED_READ_TIMEOUT = 60.0
#: v1.1 WP-A1: ask a stream to report its token usage. Without it Ollama sends
#: none, and a prompt the server cut cannot be told from one it read whole.
STREAM_OPTIONS = {"include_usage": True}
# Completion endpoints have no roles, so a chat has to be flattened into one
# labeled transcript that ends on "Assistant:" for the model to continue.
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
@@ -84,10 +88,14 @@ class OpenAICompatibleProvider(Provider):
def _record_usage(self, payload: dict) -> None:
"""Records the endpoint's own token accounting, if it reported any.
OpenRouter now always reports usage, and `usage: {include: true}` and
`stream_options` are deprecated and do nothing. In a stream the usage
arrives on a final chunk that carries no choices, which is why this is
read separately from the text extraction.
In a stream the usage arrives on a final chunk that carries no choices,
which is why this is read separately from the text extraction.
v1.1 WP-A1: Ollama sends that chunk only when asked. Measured on Ollama
0.33: a stream with no `stream_options` carried no usage at all, and not
one of the 514 AI turns in the v1 evidence had a count stored. Every
streaming body therefore sets `stream_options.include_usage`
(`STREAM_OPTIONS`), and the turn compares the count with what it sent.
"""
usage = payload.get("usage")
if isinstance(usage, dict) and usage:
@@ -102,6 +110,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
else:
url = f"{self.base_url}/chat/completions"
@@ -114,6 +123,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
return url, body
@@ -183,6 +193,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
else:
url = f"{self.base_url}/chat/completions"
@@ -192,6 +203,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
async for event in self._stream(url, body):
yield event
+38 -3
View File
@@ -6,6 +6,7 @@ lock guards one set only while one module owns it. And a test that replaces
`OpenAICompatibleProvider` or `generate_turn` patches this module, which every
caller reads through.
"""
import logging
import threading
from fastapi import Depends, HTTPException, Request
@@ -28,6 +29,8 @@ from .deps import CurrentUser, current_adventure, router
from .nodes import _move_to_after, next_depth
from .paging import annotate_takes
log = logging.getLogger(__name__)
def world_delta_of(snapshot: dict | None) -> dict | None:
"""Returns the bulk-read slice of a context snapshot, for `Action.world_delta`.
@@ -212,8 +215,18 @@ async def _generate_turn(
# network calls — and cached per endpoint and model, so it costs one short
# request per session rather than one per turn. An unverified window does
# not block the turn; it is recorded as unverified in the snapshot below.
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
#
# v1.1 WP-A1 corrective: a model that is not resident cannot report its window,
# and a turn built to the configured budget against it was silently cut in the
# A1 evidence (13,875 tokens sent, 2,050 read). So an unverified window gets
# one bounded attempt to load the model, and one more probe, before the
# prompt is assembled. No story text is generated by it and nothing is
# written. A window still unverified afterwards changes nothing below.
window, preflight = await contextwindow.ensure_window(
settings.endpoint_url, settings.model,
declared=settings.context_window_override,
warm_timeout=float(settings.model_timeout_seconds or 300),
)
try:
system_text, story_text, snapshot = build_context(
adventure,
@@ -232,6 +245,9 @@ async def _generate_turn(
yield turn_error(str(exc))
return
if isinstance(snapshot.get("window"), dict):
snapshot["window"]["preflight"] = preflight
parts = PromptParts(system=system_text, story=story_text)
provider = OpenAICompatibleProvider(
@@ -318,6 +334,24 @@ async def _generate_turn(
# prompt came from cache rather than being billed in full. This is recorded
# per attempt, next to the prompt it priced.
snapshot["usage"] = provider.last_usage
# v1.1 WP-A1: what the server says it read, against what was sent. Recorded
# and shown, never acted on: the narration has already streamed to the
# reader, and discarding an accepted turn over an accounting discrepancy
# would lose story to hide a problem. A server that cut the prompt answers
# 200 either way, so this record is the only place the cut is visible.
tokens = snapshot.get("tokens") or {}
accounting = contextwindow.classify_usage(
provider.last_usage,
estimate=tokens.get("estimate") or tokens.get("total") or 0,
budget=tokens.get("budget") or settings.context_token_budget,
max_output_tokens=settings.max_output_tokens,
window_verified=bool((snapshot.get("window") or {}).get("verified")),
)
snapshot["accounting"] = accounting
if accounting["status"] in (contextwindow.EXCEEDED,
contextwindow.TRUNCATION_SUSPECTED):
log.warning("turn accounting for adventure %s: %s — %s",
adventure.id, accounting["status"], accounting["detail"])
reasoning = "".join(reasoning_chunks).strip() or None
ai_action = models.Action(
@@ -381,7 +415,8 @@ async def _generate_turn(
db.commit()
db.refresh(ai_action)
yield _SAVED
yield sse({"type": "done", "action": action_json(ai_action, db)})
yield sse({"type": "done", "action": action_json(ai_action, db),
"accounting": accounting})
# Phase 6: schedule summarization and embedding without waiting for them.
# The task opens its own database session.
memorybank.schedule_post_turn(adventure)