v1.1: harden context window and narrator protocol boundary

WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-14 16:35:05 -04:00
co-authored by Claude Opus 5
parent ac465ed867
commit d63804f22e
26 changed files with 3903 additions and 61 deletions
+29
View File
@@ -620,6 +620,35 @@ campaign gets less history than the setting asks for, which is a visible,
explicable loss rather than a silent one, and Settings' **Test connection** explicable loss rather than a silent one, and Settings' **Test connection**
reports the window it found or says plainly that it could not check. reports the window it found or says plainly that it could not check.
**It also keeps a margin, and checks the server's own count (v1.1).** The
application counts tokens with `cl100k_base`, and your model counts them with
its own tokenizer. The two disagree slightly, so the prompt is built to leave
`max(256, 5% of the window)` tokens free on top of the reply: 256 at 4,096, and
820 at 16,384. After each turn the server's reported prompt-token count is
compared with what was sent. The context inspector shows the result for any
past turn:
- **The server read the whole prompt:** the ordinary case.
- **The server did not say how much it read:** the server reported no usage.
Nothing is wrong, and nothing is confirmed either.
- **The server may have cut the start of the prompt:** it read far fewer tokens
than were sent. Ollama does this, silently, to a prompt larger than the window
the model was loaded with. The turn is kept. Check the window with the
commands above.
- **The prompt was larger than the server allowed for:** its count and the reply
together exceed the window. The reply may have been cut short. The turn is
kept.
The last two also appear in the server log as a warning.
**A model that is not loaded yet is loaded first.** Before a turn, if the
application cannot read the window because your model isn't in memory, it asks
the same Ollama to load it once. That is a `POST /api/generate` naming only the
model, which generates no text. It then reads the window again, so the first
turn of a session is built to the window the model really has rather than to
your setting. If loading fails, or the window still can't be read, the turn goes
ahead exactly as before, unverified, and the check above still applies.
That does not make the window *bigger*, and the rest of this section is still That does not make the window *bigger*, and the rest of this section is still
how you do that. how you do that.
+13 -4
View File
@@ -80,6 +80,14 @@ that isn't the live one starts a new branch.
in settings lets you state the window so the prompt is still capped. A window the server in settings lets you state the window so the prompt is still capped. A window the server
itself reported always wins over that, and a declared one is never reported as verified. itself reported always wins over that, and a declared one is never reported as verified.
Since v1.1 the prompt also stops short of that window on purpose. It leaves
`max(256, 5% of the window)` tokens free, because your model counts tokens differently from
the application, and the v1 evidence came within 23 tokens of the edge. After each turn,
the server's own count of what it read is compared with what was sent. A turn the server
appears to have truncated is kept, flagged and shown in the context inspector, not left to
pass silently. A model that isn't loaded yet, and so cannot report its window, is loaded
once before the turn is built, so the first turn of a session gets the real window too.
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as **Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
card used to arrive in front of it as a world fact with no class, no visibility, no source and card used to arrive in front of it as a world fact with no class, no visibility, no source and
@@ -179,9 +187,9 @@ that isn't the live one starts a new branch.
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up, None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
removed rather than left standing as a picture of a product that no longer exists. The M4 removed rather than left standing as a picture of a product that no longer exists. The
closeout drove the real application in a real browser, so the screens exist and work; taking screens exist and are driven in a real browser by the release harness
presentable screenshots of them is a job for the UI pass in M8. (`backend/tools/m11_browser.py`). Presentable screenshots of them have not been taken.
## Quick start ## Quick start
@@ -310,7 +318,8 @@ development, Vite proxies `/api` to FastAPI.
## Tests ## Tests
1,191 backend tests: unit tests plus full HTTP integration through the real turn engine, with 1,524 backend tests (1,507 run everywhere, 17 need a real local model and skip without one): unit
tests plus full HTTP integration through the real turn engine, with
the model provider mocked. They run with no route to the Internet, which is a requirement the model provider mocked. They run with no route to the Internet, which is a requirement
rather than a convenience — an offline claim proved on a machine that has been online once rather than a convenience — an offline claim proved on a machine that has been online once
proves nothing. A further handful need a real local model and skip without one; they exist proves nothing. A further handful need a real local model and skip without one; they exist
+8 -1
View File
@@ -46,7 +46,14 @@ from .narrative import model as narrative_model
# token accounting. Each attempt is its own API call, and a retry is the call # token accounting. Each attempt is its own API call, and a retry is the call
# most likely to read the prompt back out of cache. Everything else in a snapshot # most likely to read the prompt back out of cache. Everything else in a snapshot
# is the prompt, which is assembled once per turn. # is the prompt, which is assembled once per turn.
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage") #
# v1.1 WP-A1: `accounting` is one attempt's too. It compares the server's count
# for *that* call with the turn's estimate. Left out of this tuple, it was
# treated as part of the shared prompt, so moving the live flag handed the
# superseded attempt's accounting to the new live one and threw the new one's
# away. Found by the A2 long run: two retries and one take selection left three
# attempts reporting no accounting, or another attempt's.
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage", "accounting")
# ------------------------------------------------------------------ reading # ------------------------------------------------------------------ reading
+51 -20
View File
@@ -25,6 +25,7 @@ from sqlalchemy.orm import object_session
from .. import contextwindow, derived, models, narrative, summaries, worldstate from .. import contextwindow, derived, models, narrative, summaries, worldstate
from ..knowledge import inject as knowledge_inject from ..knowledge import inject as knowledge_inject
from ..providers.openai_compatible import CHAT_CONTINUE_HINT
from ..knowledge import records as knowledge_records from ..knowledge import records as knowledge_records
from . import encoding, history from . import encoding, history
@@ -106,11 +107,22 @@ BAND_FLOOR_SHARE = 0.5
# Built from the table vendored in `encoding.py`, not fetched: the upstream # Built from the table vendored in `encoding.py`, not fetched: the upstream
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this # `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
# is called on every turn. # is called on every turn.
# M6: added to the configured reply budget when reserving output space. It #
# absorbs the section separators added after budgeting and the drift between # v1.1 WP-A1: `OUTPUT_SAFETY_MARGIN = 64` was here. M6 added it to the reply
# this tokenizer and the serving model's. Fixed rather than proportional: what # budget to absorb two unrelated things, and v1.1 separates them:
# it covers does not grow with the size of the budget. #
OUTPUT_SAFETY_MARGIN = 64 # * **Text the application adds after pricing.** The separators between
# sections, and `CHAT_CONTINUE_HINT`, which the provider appends to every chat
# request and nothing counted. That is not drift, it is our own text, so it is
# now priced exactly (`transport` below).
# * **The drift between this tokenizer and the narrator's.** That is what the
# 64 tokens were really for, and the v1 evidence showed it was too small. It
# is now `contextwindow.safety_reserve`, sized to the window.
#
#: Story sections that can be joined by `SEPARATOR` after pricing: history,
#: author's note, recent history, summary, lore, memories, state, front memory,
#: length hint, refusals, reminder. Knowledge and history rows price their own.
STORY_SECTION_SLOTS = 11
class ContextOverflow(RuntimeError): class ContextOverflow(RuntimeError):
@@ -180,19 +192,19 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
words = min(words, band_ceiling) words = min(words, band_ceiling)
floor = min(band_floor, int(words * BAND_FLOOR_SHARE)) floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
tail = ( tail = (
" Finish the narration and append the state block well inside the limit." " " + narrative.extract.LENGTH_HINT_TAIL
) )
if floor < MIN_LENGTH_FLOOR_WORDS: if floor < MIN_LENGTH_FLOOR_WORDS:
return ( return (
f"[Hard limit: this turn must not exceed {words} words. Write only as " f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]" f"much as the moment needs — a typical turn is much shorter.{tail}]"
) )
return ( return (
f"[Hard limit: this turn must not exceed {words} words, and it should not " f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless " f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]" f"the scene genuinely needs more.{tail}]"
) )
tail = " Finish the narration and append the state block well inside the limit." tail = " " + narrative.extract.LENGTH_HINT_TAIL
# State the number as a ceiling, never as a budget. In measurements, the # State the number as a ceiling, never as a budget. In measurements, the
# wording "keep this turn under about N words" read to the model as a target # wording "keep this turn under about N words" read to the model as a target
@@ -204,7 +216,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS) floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
if floor < MIN_LENGTH_FLOOR_WORDS: if floor < MIN_LENGTH_FLOOR_WORDS:
return ( return (
f"[Hard limit: this turn must not exceed {words} words. Write only as " f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]" f"much as the moment needs — a typical turn is much shorter.{tail}]"
) )
# Both numbers are bounds, and the wording is deliberately asymmetric. The # Both numbers are bounds, and the wording is deliberately asymmetric. The
@@ -216,7 +228,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
# so a terse model reading the same clause stops at the floor rather than at # so a terse model reading the same clause stops at the floor rather than at
# forty words. # forty words.
return ( return (
f"[Hard limit: this turn must not exceed {words} words, and it should not " f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless " f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]" f"the scene genuinely needs more.{tail}]"
) )
@@ -625,12 +637,19 @@ def build_context(
# truncated turn on a model whose window is the budget # truncated turn on a model whose window is the budget
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04). # (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
# #
# The margin covers what is added after this arithmetic — the separators # v1.1 WP-A1: the reply allocation is exactly the reply cap. The text this
# between sections, and the difference between our tokenizer's count and the # application adds after pricing — separators, and the chat hint the
# serving model's. It is small and fixed rather than proportional, because # provider appends — is counted as `transport`. What neither can know, the
# what it absorbs does not scale with the budget. # narrator's tokenizer disagreeing with `cl100k_base`, is the safety reserve,
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN # which is sized to the window and taken before any history is chosen.
protected = reserved + output_reserve output_reserve = max(0, settings.max_output_tokens)
separator_tokens = count_tokens(SEPARATOR)
transport = (
separator_tokens * (len(system_sections) + STORY_SECTION_SLOTS)
+ count_tokens(CHAT_CONTINUE_HINT)
)
safety = contextwindow.safety_reserve(budget)
protected = reserved + transport + output_reserve + safety
if protected >= budget: if protected >= budget:
# Failing here is the point. The alternative — carrying on with a token # Failing here is the point. The alternative — carrying on with a token
# or two of history — builds a prompt that is known to overflow, and # or two of history — builds a prompt that is known to overflow, and
@@ -638,8 +657,9 @@ def build_context(
# gracefully if protected context alone is too large." # gracefully if protected context alone is too large."
raise ContextOverflow( raise ContextOverflow(
f"The protected context needs {protected} tokens " f"The protected context needs {protected} tokens "
f"({reserved} of prompt plus {output_reserve} reserved for the " f"({reserved} of prompt, {transport} of formatting, {output_reserve} "
f"reply) but the context budget is {budget}. " f"reserved for the reply and a {safety}-token safety margin) but the "
f"context budget is {budget}. "
+ ( + (
"That budget is what this server was found to accept, so raising " "That budget is what this server was found to accept, so raising "
"the setting alone will not help — load the model with a larger " "the setting alone will not help — load the model with a larger "
@@ -827,6 +847,7 @@ def build_context(
story_text = SEPARATOR.join(s.text for s in story_sections) story_text = SEPARATOR.join(s.text for s in story_sections)
all_sections = [s for s in system_sections if s.text] + story_sections all_sections = [s for s in system_sections if s.text] + story_sections
total_tokens = count_tokens(system_text) + count_tokens(story_text)
report = { report = {
"sections": [ "sections": [
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections {"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
@@ -837,13 +858,23 @@ def build_context(
# what the history was actually allowed to spend after everything # what the history was actually allowed to spend after everything
# protected was subtracted. # protected was subtracted.
"tokens": { "tokens": {
"total": count_tokens(system_text) + count_tokens(story_text), "total": total_tokens,
"budget": budget, "budget": budget,
"configured_budget": settings.context_token_budget, "configured_budget": settings.context_token_budget,
"output_reserve": output_reserve, "output_reserve": output_reserve,
"protected": reserved, "protected": reserved,
"available_for_history": available, "available_for_history": available,
"history_spent": spent, "history_spent": spent,
# v1.1 WP-A1. `transport` is the formatting priced in above;
# `estimate` is what this application believes it actually sent,
# the assembled text plus what the provider adds to it, and is what
# the server's own count is compared against after the reply.
"transport": transport,
"safety_reserve": safety,
"estimate": total_tokens + (
count_tokens(CHAT_CONTINUE_HINT) if settings.api_mode != "completion"
else separator_tokens
),
}, },
# M11: what the server was found to accept, and how. `verified` false # M11: what the server was found to accept, and how. `verified` false
# means nobody could check — the prompt was built to the configured # means nobody could check — the prompt was built to the configured
+208
View File
@@ -86,6 +86,7 @@ one.
from __future__ import annotations from __future__ import annotations
import logging import logging
import math
import re import re
import time import time
from dataclasses import dataclass from dataclasses import dataclass
@@ -132,6 +133,11 @@ class Window:
model_max: int | None = None model_max: int | None = None
#: Why the window is unknown, or how it was found. Shown to the user. #: Why the window is unknown, or how it was found. Shown to the user.
detail: str = "" detail: str = ""
#: v1.1: the server answered a discovery request at all, whatever it said.
#: A server that answered but could not report a window may simply not have
#: the model loaded yet, which `ensure_window` can fix; one that did not
#: answer cannot be helped by asking it to load anything.
reachable: bool = False
@property @property
def verified(self) -> bool: def verified(self) -> bool:
@@ -180,6 +186,116 @@ def effective_budget(configured: int, window: Window | int | None) -> int:
return min(configured, tokens) return min(configured, tokens)
#: v1.1 WP-A1: the tokens kept free below the effective window, beyond the reply.
#:
#: The builder counts with `cl100k_base`; the narrator counts with its own
#: tokenizer. The v1 evidence put the largest prompts 23-42 real tokens from the
#: edge of a 16,384 window, and Ollama does not refuse a prompt past the edge —
#: measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt came back 200
#: with `prompt_tokens` 2,050. So the reserve is deliberate and sized to the
#: window: the larger of a floor and a share, **rounded up to a whole token**.
#:
#: 4,096 -> 256 8,192 -> 410 16,384 -> 820
#:
#: A fixed, documented tolerance, owner-chosen for v1.1. It is not a setting and
#: it is not calibrated per model.
SAFETY_RESERVE_FLOOR = 256
SAFETY_RESERVE_PERCENT = 5
def safety_reserve(effective_window: int) -> int:
"""`max(256, ceil(5% of the effective window))`, in tokens.
The effective window is the budget the prompt is actually built to — the
verified or declared window when there is one, the configured budget
otherwise — so a 16,384 setting against a 4,096 server reserves 256, not 820.
Integer arithmetic, so the rounding is exact rather than a float's.
"""
share = math.ceil(max(0, effective_window) * SAFETY_RESERVE_PERCENT / 100)
return max(SAFETY_RESERVE_FLOOR, share)
#: v1.1 WP-A1: what the server's own count says about a turn that was sent.
FITS = "fits"
EXCEEDED = "exceeded"
TRUNCATION_SUSPECTED = "truncation_suspected"
#: `UNKNOWN` above: the server reported no usable count.
def classify_usage(usage: dict | None, *, estimate: int, budget: int,
max_output_tokens: int, window_verified: bool) -> dict:
"""Sets the server's reported prompt count against what the application sent.
The order of the checks is the order of what they prove:
``unknown``
No positive integer `prompt_tokens`. Nothing can be said, and nothing
is claimed: an absent count is never read as a prompt that fitted.
``truncation_suspected``
The server read fewer tokens than were sent by more than the safety
reserve. A tokenizer thriftier than `cl100k_base` may honestly count a
little less; a shortfall larger than the tolerance the application keeps
for drift is the signature of a server that cut the prompt — the real
shape was 6,316 sent and 2,050 read.
``exceeded``
The server's count plus the reply allocation is more than the window
the prompt was built for. The drift was larger than the whole reserve,
so the reply may have been cut short.
``fits``
Otherwise.
`observed_margin` is what was left beside the reply by the server's count:
`budget - max_output_tokens - server_prompt_tokens`. The safety reserve is
the tolerance, so a margin between 0 and the reserve is still `fits`.
A discrepancy is recorded, never acted on: the reply has already streamed
to the reader and is accepted story.
"""
prompt = usage.get("prompt_tokens") if isinstance(usage, dict) else None
reserve = safety_reserve(budget)
verified_note = "" if window_verified else (
" The window itself was not verified for this turn.")
record = {
"status": UNKNOWN,
"server_prompt_tokens": None,
"estimate": estimate,
"difference": None,
"budget": budget,
"max_output_tokens": max_output_tokens,
"safety_reserve": reserve,
"observed_margin": None,
"window_verified": bool(window_verified),
"detail": "",
}
if type(prompt) is not int or prompt <= 0:
record["detail"] = ("The server reported no prompt token count, so nothing "
"confirms the whole prompt was read." + verified_note)
return record
record["server_prompt_tokens"] = prompt
record["difference"] = prompt - estimate
record["observed_margin"] = budget - max_output_tokens - prompt
if prompt + reserve < estimate:
record["status"] = TRUNCATION_SUSPECTED
record["detail"] = (
f"The server read {prompt:,} prompt tokens of the {estimate:,} sent, a "
f"shortfall larger than the {reserve:,}-token safety reserve. A server "
"that cuts an over-window prompt reports exactly this, and what it cuts "
"is the start: the narrator's rules and the canon." + verified_note)
elif prompt + max_output_tokens > budget:
record["status"] = EXCEEDED
record["detail"] = (
f"The server counted {prompt:,} prompt tokens; with {max_output_tokens:,} "
f"for the reply that is more than the {budget:,}-token window the prompt "
"was built for, so the reply may have been cut short." + verified_note)
else:
record["status"] = FITS
record["detail"] = (
f"The server read {prompt:,} prompt tokens, leaving "
f"{record['observed_margin']:,} beside the reply." + verified_note)
return record
def cache_clear() -> None: def cache_clear() -> None:
"""Forgets what was learned. Called when the endpoint or model changes.""" """Forgets what was learned. Called when the endpoint or model changes."""
_cache.clear() _cache.clear()
@@ -223,6 +339,7 @@ def _declared_or(declared: int | None, discovered: Window) -> Window:
declared, DECLARED, discovered.model_max, declared, DECLARED, discovered.model_max,
f"{declared:,} tokens, declared in settings — the server was not able " f"{declared:,} tokens, declared in settings — the server was not able "
f"to say ({discovered.detail})", f"to say ({discovered.detail})",
reachable=discovered.reachable,
) )
@@ -266,6 +383,7 @@ async def _ask(endpoint_url: str, model: str) -> Window:
return Window( return Window(
tokens, LOADED, ceiling, tokens, LOADED, ceiling,
f"{tokens:,} tokens, reported by the running model", f"{tokens:,} tokens, reported by the running model",
reachable=True,
) )
return await _declared_window(client, base, model) return await _declared_window(client, base, model)
except (httpx.HTTPError, ValueError, TypeError, KeyError) as exc: except (httpx.HTTPError, ValueError, TypeError, KeyError) as exc:
@@ -293,6 +411,7 @@ async def _declared_window(client, base: str, model: str) -> Window:
return Window( return Window(
None, UNKNOWN, None, UNKNOWN,
detail=f"the server did not describe the model (HTTP {resp.status_code})", detail=f"the server did not describe the model (HTTP {resp.status_code})",
reachable=True,
) )
body = resp.json() or {} body = resp.json() or {}
ceiling = _architecture_ceiling(body.get("model_info") or {}) ceiling = _architecture_ceiling(body.get("model_info") or {})
@@ -304,14 +423,103 @@ async def _declared_window(client, base: str, model: str) -> Window:
"the model sets no num_ctx, so the server will load it at its own " "the model sets no num_ctx, so the server will load it at its own "
"default — which is 4,096 where there is no VRAM" "default — which is 4,096 where there is no VRAM"
), ),
reachable=True,
) )
tokens = min(declared, ceiling) if ceiling else declared tokens = min(declared, ceiling) if ceiling else declared
return Window( return Window(
tokens, PARAMETERS, ceiling, tokens, PARAMETERS, ceiling,
f"{tokens:,} tokens, from the model's own num_ctx", f"{tokens:,} tokens, from the model's own num_ctx",
reachable=True,
) )
#: v1.1 WP-A1 corrective: loading the configured model so its window can be read.
#:
#: The first real turn of the A1 evidence found a cold model: `/api/ps` knew
#: nothing, `/api/show` found no `num_ctx`, so the window was unverified and the
#: prompt was built to the configured 16,384. Ollama loaded the model at its own
#: 4,096 default, kept 2,050 of 13,875 tokens and answered 200. That case is
#: preventable, because the window becomes readable the moment the model is
#: resident. Ollama's native `POST /api/generate` with a model and **no prompt**
#: loads the model and generates nothing — measured on Ollama 0.33: HTTP 200,
#: `"response": ""`, `"done_reason": "load"`, and `/api/ps` then reported the
#: window. The OpenAI-compatible request that followed did not reload it.
WARM_PATH = "/api/generate"
async def warm(endpoint_url: str, model: str, *, timeout: float) -> tuple[bool, str]:
"""Asks the configured server to load `model`. One request, no story text.
Held to the same endpoint policy and TLS trust as inference and the probe, and
sent to the same host the probe asks. The body names the model and nothing
else: no prompt, so nothing is generated, and no `options` or `keep_alive`, so
the model loads the way the server would load it for the turn itself.
Returns `(loaded, detail)`. Every failure is `(False, why)` and never raises:
a server that will not load the model on request will fail the turn's own
call the ordinary way, which is where that failure belongs.
"""
reason = endpoints.rejection_reason(endpoint_url)
if reason is not None:
return False, f"endpoint not allowed — {reason}"
base = native_base(endpoint_url)
try:
async with httpx.AsyncClient(
timeout=httpx.Timeout(timeout, connect=CONNECT_TIMEOUT),
verify=tlstrust.ssl_context(),
) as client:
resp = await client.post(f"{base}{WARM_PATH}", json={"model": model})
except httpx.HTTPError as exc:
log.debug("model warm-up failed for %s: %s", base, exc)
return False, f"could not ask the server to load the model ({type(exc).__name__})"
if resp.status_code != 200:
return False, f"the server did not load the model (HTTP {resp.status_code})"
try:
body = resp.json() or {}
except ValueError:
return False, "the server answered the load request with something that was not JSON"
return True, f"the server loaded the model ({body.get('done_reason') or 'done'})"
async def ensure_window(endpoint_url: str, model: str, *, declared: int | None = None,
warm_timeout: float = 300.0) -> tuple[Window, dict]:
"""The window for a turn about to be generated, loading the model once if that is what it takes.
1. Probe as before.
2. If the window is not verified, the server answered, and there is a model to
load: one bounded `warm` request.
3. If the model loaded, probe again, bypassing the cache that still holds the
unverified answer.
Whatever the second probe says is the answer. There is no retry loop, no
guessed window, and no hard-coded 4,096: a window still unverified leaves the
configured budget standing, exactly as before, and the turn's accounting
still catches a server that cut the prompt.
Returns the window and a `preflight` record for the turn's provenance.
Not used by the context dry run: loading a model is a side effect, and
opening a panel should not cause one.
"""
window = await probe(endpoint_url, model, declared=declared)
preflight = {"attempted": False, "loaded": None, "verified_before": window.verified,
"verified_after": window.verified, "detail": ""}
if window.verified:
preflight["detail"] = "the window was already verified"
return window, preflight
if not (endpoint_url and model):
preflight["detail"] = "no endpoint or model configured"
return window, preflight
if not window.reachable:
preflight["detail"] = "the server did not answer, so no model was loaded"
return window, preflight
loaded, detail = await warm(endpoint_url, model, timeout=warm_timeout)
preflight.update(attempted=True, loaded=loaded, detail=detail)
if loaded:
window = await probe(endpoint_url, model, declared=declared, use_cache=False)
preflight["verified_after"] = window.verified
return window, preflight
def _num_ctx(parameters) -> int | None: def _num_ctx(parameters) -> int | None:
"""Reads `num_ctx` out of the plain-text parameter block Ollama returns.""" """Reads `num_ctx` out of the plain-text parameter block Ollama returns."""
if not isinstance(parameters, str): if not isinstance(parameters, str):
+25 -4
View File
@@ -35,6 +35,8 @@ creating a second, empty Mara.
from __future__ import annotations from __future__ import annotations
import json
# Field types the schema layer enforces. Kept deliberately small: a narrative # Field types the schema layer enforces. Kept deliberately small: a narrative
# state event carries names, labels and plain values, and nothing here needs a # state event carries names, labels and plain values, and nothing here needs a
# nested structure a model could hide something inside. # nested structure a model could hide something inside.
@@ -184,11 +186,30 @@ def vocabulary_for_prompt() -> str:
Generated from `SPECS` rather than written out beside it, so the model can Generated from `SPECS` rather than written out beside it, so the model can
never be told about an event the application does not implement — the drift never be told about an event the application does not implement — the drift
that would produce proposals rejected for reasons nobody could see. that would produce proposals rejected for reasons nobody could see.
v1.1 WP-A2: each event is shown as the object the model must put in the
`events` list, with its required fields, not as `name(field, …)`. The call
notation was never the wire format, and a 3B narrator copied it into its
prose as `> set_possession(silver-key, "alice")`. An object copied into prose
is a proposal the extractor already recognises and removes; a call is not.
""" """
lines = [] lines = []
for name, definition in SPECS.items(): for name, definition in SPECS.items():
fields = list(definition["required"]) + [ shape = {"type": name}
f"{field}?" for field in definition["optional"] for field, kind in definition["required"].items():
] shape[field] = _PLACEHOLDER[kind]
lines.append(f' {name}({", ".join(fields)}) — {definition["summary"]}') body = json.dumps(shape, ensure_ascii=False, separators=(",", ":"))
line = f" {body} — {definition['summary']}"
if definition["optional"]:
line += f" (optional: {', '.join(definition['optional'])})"
lines.append(line)
return "\n".join(lines) return "\n".join(lines)
#: What a field of each kind looks like in the prompt's vocabulary. Placeholders,
#: never example identifiers, so the vocabulary names nothing a story could copy.
#: A list field is shown as a list, so the model is told its shape; every other
#: field is an ellipsis. Measured: `"<key>"`-style placeholders with spaced
#: separators cost 456 tokens against v1.0.0's 258; this form costs about 380,
#: and every line is still the object the model must send.
_PLACEHOLDER = {KEY: "…", TEXT: "…", VALUE: "…", LABELS: ["…"]}
+188 -14
View File
@@ -37,17 +37,17 @@ EMIT_RULE = (
"appeared, record it.\n" "appeared, record it.\n"
"\n" "\n"
"Every value is ABSOLUTE — the new state of things, never a change or a " "Every value is ABSOLUTE — the new state of things, never a change or a "
"difference. Use only these events:\n" "difference. Use only these events, in exactly this shape:\n"
f"{events.vocabulary_for_prompt()}\n" f"{events.vocabulary_for_prompt()}\n"
"\n" "\n"
"Identifiers are short lower-case slugs (mara, silver-key, old-abbey) and must " "Identifiers are short lower-case slugs and must match the ones already in the "
"match the ones already in the state you were shown. Introduce a person, place " "state you were shown; the example's identifiers are placeholders. Introduce a "
"or thing with create_entity before referring to it. If the turn established " "person, place or thing with create_entity before referring to it. If the turn "
"nothing, send an empty events list.\n" "established nothing, send an empty events list.\n"
"Example:\n" "Example:\n"
'```state\n' '```state\n'
'{"events": [{"type": "set_possession", "item": "silver-key", "owner": "aldric"},' '{"events": [{"type": "set_possession", "item": "item-1", "owner": "character-1"},'
' {"type": "set_current_location", "entity": "aldric", "location": "old-abbey"}]}\n' ' {"type": "set_current_location", "entity": "character-1", "location": "location-1"}]}\n'
'```' '```'
) )
@@ -58,6 +58,47 @@ EMIT_REMINDER = (
"nothing changed.]" "nothing changed.]"
) )
# v1.1 WP-A2: the length hint's own words, named once. `builder.length_hint`
# builds the hint from these, and the extractor recognises an echo of it by
# them, so the two cannot drift apart.
LENGTH_HINT_OPENING = "[Hard limit:"
LENGTH_HINT_TAIL = "Finish the narration and append the state block well inside the limit."
#: The application's wording inside a hint. A 3B narrator reworded the front
#: ("your next turn") and the end ("This story ends here."), and kept one or the
#: other of these every time.
_LENGTH_HINT_PHRASE_RE = re.compile(
r"append the state block|turn must not exceed \d+ words", re.IGNORECASE
)
#: v1.1 WP-A2: the rules that remove protocol a narrator copied, named so the
#: replay tool and the report can say which removed what.
RULE_EVENT_CALL = "event_call_line"
RULE_LENGTH_HINT = "echoed_length_hint"
RULE_SCENE_LINE = "rendered_scene_line"
RULE_EMPTY_FENCE = "empty_dangling_fence"
RULE_INSTRUCTION_TAIL = "echoed_instruction_tail"
#: v1.1 WP-A2 corrective (R5). The sentence `CHAT_CONTINUE_HINT` in
#: `providers/openai_compatible.py` carries, which a narrator echoed with the rest
#: of the hint reworded around it. Kept as a copy rather than an import, so the
#: narrative package does not depend on the provider; a test pins that the
#: hint still contains it.
CONTINUE_HINT_PHRASE = "Output only story text"
# R1. A whole line opening with a call to an event this protocol has. The names
# come from the vocabulary, so a call-shaped line naming anything else — a
# character's `open_door(north)` — is not matched.
_EVENT_CALL_LINE_RE = re.compile(
r"^[ \t]*(?:>[ \t]*)?(?:"
+ "|".join(re.escape(name) for name in events.SPECS)
+ r")[ \t]*\(",
re.IGNORECASE,
)
# R3. The renderer's scene line carries its location this way.
_RENDERED_SCENE_LOCATION_RE = re.compile(r"\(at [^()\n]+\)\s*$")
# R4. An opener with nothing after it.
_EMPTY_FENCE_LINE_RE = re.compile(r"```(?:json)?[ \t]*", re.IGNORECASE)
# Three patterns, and the difference between them is the whole of this module's # Three patterns, and the difference between them is the whole of this module's
# safety. A story is allowed to contain code, and taking a code block out of # safety. A story is allowed to contain code, and taking a code block out of
# someone's prose is a worse failure than leaving a stray proposal in it. # someone's prose is a worse failure than leaving a stray proposal in it.
@@ -128,11 +169,42 @@ def _is_echoed_instruction(inner: str) -> bool:
# opening words only, because the echo is often cut off before it ends. # opening words only, because the echo is often cut off before it ends.
if low.lstrip().startswith("continue the story directly"): if low.lstrip().startswith("continue the story directly"):
return True return True
# v1.1 WP-A2 corrective (R5): the same hint, reworded at the front. The M11
# closeout-era identity re-run stored "[You don't need to continue; … Continue
# the story here, directly. Output only story text.]" as the last line of a
# reply, and because nothing recognised it, nothing above it was trailing.
if CONTINUE_HINT_PHRASE.lower() in low:
return True
# v1.1 WP-A2 (R2): the length hint, which names "state block" but not
# "events list", so it passed every check above.
if _is_length_hint(inner):
return True
# The reminder names both; prose about the protocol rarely names either the # The reminder names both; prose about the protocol rarely names either the
# way the instruction does, and effectively never both. # way the instruction does, and effectively never both.
return "state block" in low and "events list" in low return "state block" in low and "events list" in low
def _opens_like_length_hint(inner: str) -> bool:
"""R5. The bracket opens with the length hint's own `Hard limit:`, whatever follows.
Never enough on its own: an in-world "[Hard limit: forty days]" opens the same
way. `_clean` takes it only directly above an echoed instruction it has already
removed from the end of the same reply.
"""
return inner.lstrip().lower().startswith(LENGTH_HINT_OPENING[1:].lower())
def _is_length_hint(inner: str) -> bool:
"""Whether a bracket's contents are `builder.length_hint`, however reworded.
It must open the way the hint opens *and* carry the hint's own wording. An
in-world "Hard limit: forty days" has the opening and none of the wording.
"""
opening = LENGTH_HINT_OPENING[1:].lower()
return (inner.lstrip().lower().startswith(opening)
and bool(_LENGTH_HINT_PHRASE_RE.search(inner)))
# A heading the model writes above a block it did not fence: `State`, sometimes # A heading the model writes above a block it did not fence: `State`, sometimes
# as `State:`, `**State**` or `### State`. It is removed only in two places: # as `State:`, `**State**` or `### State`. It is removed only in two places:
# directly above a proposal that is removed, and as the last line of the reply. # directly above a proposal that is removed, and as the last line of the reply.
@@ -145,7 +217,7 @@ _LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?") _QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
def _clean(prose: str) -> str: def _clean(prose: str, *, after_block: bool = False) -> str:
"""Removes protocol the block extraction could not, and nothing else. """Removes protocol the block extraction could not, and nothing else.
Found by the M5 realistic-context run (§12), which is the failure class Found by the M5 realistic-context run (§12), which is the failure class
@@ -160,9 +232,23 @@ def _clean(prose: str) -> str:
story after it. Stored text is replayed as history, so every leak also story after it. Stored text is replayed as history, so every leak also
showed the next prompt a second, older account of the state, which is what showed the next prompt a second, older account of the state, which is what
M5 review Finding 4 removed from replayed history. M5 review Finding 4 removed from replayed history.
v1.1 WP-A2 added four shapes, from the M11 closeout's identity run and the
v1 corpus, each anchored to something the application owns rather than to
what prose looks like: a line opening with a vocabulary call (R1), the
length hint echoed at the end (R2), the renderer's scene line left last
(R3), and an empty fence opener left last (R4). `after_block` says a
proposal block was already taken out of this reply, which is what lets R3
remove a bare scene line that sat above it.
""" """
cleaned, _found = _inline_proposals(prose) cleaned, calls_removed = _strip_event_call_lines(prose)
cleaned, _found = _inline_proposals(cleaned)
cleaned = _strip_echoed_state(cleaned) cleaned = _strip_echoed_state(cleaned)
protocol_cut = after_block or calls_removed
# R5: set once an echoed instruction bracket has come off the end. Only then
# may a bracket that merely opens the way the length hint opens be taken as
# part of the same echoed tail.
instruction_cut = False
# The end of the reply is cut until nothing more comes off, because one kind # The end of the reply is cut until nothing more comes off, because one kind
# of leftover can hide another. In a real reply, a `State` heading sat above # of leftover can hide another. In a real reply, a `State` heading sat above
# a block the model never finished, and a parroted reminder sat above an # a block the model never finished, and a parroted reminder sat above an
@@ -171,7 +257,12 @@ def _clean(prose: str) -> str:
before = cleaned before = cleaned
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE): for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
bracket = pattern.search(cleaned) bracket = pattern.search(cleaned)
if bracket is not None and _is_echoed_instruction(bracket.group(1)): if bracket is None:
continue
if _is_echoed_instruction(bracket.group(1)):
cleaned = cleaned[: bracket.start()]
instruction_cut = True
elif instruction_cut and _opens_like_length_hint(bracket.group(1)):
cleaned = cleaned[: bracket.start()] cleaned = cleaned[: bracket.start()]
cleaned = _DANGLING_STATE_RE.sub("", cleaned) cleaned = _DANGLING_STATE_RE.sub("", cleaned)
dangling = _DANGLING_JSON_RE.search(cleaned) dangling = _DANGLING_JSON_RE.search(cleaned)
@@ -182,10 +273,93 @@ def _clean(prose: str) -> str:
cleaned = _strip_trailing_state_heading(cleaned).rstrip() cleaned = _strip_trailing_state_heading(cleaned).rstrip()
# A bare quote marker, the start of a quoted block that never came. # A bare quote marker, the start of a quoted block that never came.
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned) cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
cleaned = _strip_empty_dangling_fence(cleaned)
if cleaned.rstrip() != before.rstrip():
protocol_cut = True
cleaned = _strip_trailing_scene_line(cleaned, protocol_cut)
if cleaned == before: if cleaned == before:
return cleaned.strip() return cleaned.strip()
def _strip_event_call_lines(text: str) -> tuple[str, bool]:
"""R1. Removes whole lines that open with a call to a vocabulary event.
A line inside a fenced code block is the story's own code and is never
examined. Returns the text and whether anything was removed.
"""
kept: list[str] = []
in_fence = False
removed = False
for line in text.split("\n"):
if line.lstrip().startswith("```"):
in_fence = not in_fence
kept.append(line)
continue
if not in_fence and _EVENT_CALL_LINE_RE.match(line):
removed = True
continue
kept.append(line)
if not removed:
return text, False
return re.sub(r"\n{3,}", "\n\n", "\n".join(kept)), True
def _strip_empty_dangling_fence(text: str) -> str:
"""R4. A ```` ```json ```` or ```` ``` ```` opener as the last line, with nothing after it.
Only an *opener*: the fence lines are counted, and an even count means the
last one closes a story's own code block, which stays.
"""
lines = text.rstrip().split("\n")
if len(lines) < 2 or not _EMPTY_FENCE_LINE_RE.fullmatch(lines[-1].strip()):
return text
fences = sum(1 for line in lines if line.lstrip().startswith("```"))
if fences % 2 == 0:
return text
return "\n".join(lines[:-1]).rstrip()
def _strip_trailing_scene_line(text: str, protocol_cut: bool) -> str:
"""R3. The renderer's scene line, left as the last line of the reply.
Taken when it carries the renderer's own `(at <location>)`, or when protocol
was already cut from this reply, which makes a bare scene line part of the
same pasted tail. A final screenplay-style "Scene: …" line in a reply with
no protocol in it stays, and so does any scene line with story after it.
"""
lines = text.rstrip().split("\n")
if len(lines) < 2:
return text
last = lines[-1].strip()
if not last.startswith(render.HEADING_SCENE + " "):
return text
if not (_RENDERED_SCENE_LOCATION_RE.search(last) or protocol_cut):
return text
return "\n".join(lines[:-1]).rstrip()
def explain_removed_line(line: str) -> str | None:
"""Which v1.1 rule removes a line of this shape, for the replay report.
None means no v1.1 rule explains it, which the replay treats as a failure.
"""
stripped = line.strip()
if _EVENT_CALL_LINE_RE.match(line):
return RULE_EVENT_CALL
if stripped.startswith("["):
inner = stripped[1:]
inner = inner[:-1] if inner.endswith("]") else inner
if _is_length_hint(inner):
return RULE_LENGTH_HINT
if _is_echoed_instruction(inner) or _opens_like_length_hint(inner):
return RULE_INSTRUCTION_TAIL
if stripped.startswith(render.HEADING_SCENE + " "):
return RULE_SCENE_LINE
if _EMPTY_FENCE_LINE_RE.fullmatch(stripped):
return RULE_EMPTY_FENCE
return None
def _is_state_heading(line: str) -> bool: def _is_state_heading(line: str) -> bool:
return bool(_STATE_HEADING_RE.match(line)) return bool(_STATE_HEADING_RE.match(line))
@@ -451,7 +625,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
if matches: if matches:
match = matches[-1] match = matches[-1]
raw = match.group(1).strip() raw = match.group(1).strip()
prose = _clean(text[: match.start()] + text[match.end():]) prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
return prose, _tolerant_load(raw), raw return prose, _tolerant_load(raw), raw
# A `json` or unlabelled fence is ours only when its contents are this # A `json` or unlabelled fence is ours only when its contents are this
@@ -465,7 +639,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
raw = match.group(1).strip() raw = match.group(1).strip()
parsed = _tolerant_load(raw) parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed) or _reads_as_protocol(raw): if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
prose = _clean(text[: match.start()] + text[match.end():]) prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
return prose, parsed, raw return prose, parsed, raw
match = _TRAILING_RE.search(text) match = _TRAILING_RE.search(text)
@@ -473,7 +647,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
raw = match.group(1) raw = match.group(1)
parsed = _tolerant_load(raw) parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed): if _looks_like_proposal(parsed):
return _clean(text[: match.start()]), parsed, raw return _clean(text[: match.start()], after_block=True), parsed, raw
# An unfenced proposal on its own lines but not at the end: quoted, or # An unfenced proposal on its own lines but not at the end: quoted, or
# followed by more story. The last one is the turn's proposal, as with # followed by more story. The last one is the turn's proposal, as with
@@ -481,7 +655,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
without, found = _inline_proposals(text) without, found = _inline_proposals(text)
if found: if found:
parsed, raw = found[-1] parsed, raw = found[-1]
return _clean(without), parsed, raw return _clean(without, after_block=True), parsed, raw
# No block at all — but the reply may still carry protocol the model wrote # No block at all — but the reply may still carry protocol the model wrote
# as prose, or a fence it never closed. # as prose, or a fence it never closed.
+16 -4
View File
@@ -26,6 +26,10 @@ EMBED_READ_TIMEOUT = 60.0
#: v1.1 WP-A1: ask a stream to report its token usage. Without it Ollama sends
#: none, and a prompt the server cut cannot be told from one it read whole.
STREAM_OPTIONS = {"include_usage": True}
# Completion endpoints have no roles, so a chat has to be flattened into one # Completion endpoints have no roles, so a chat has to be flattened into one
# labeled transcript that ends on "Assistant:" for the model to continue. # labeled transcript that ends on "Assistant:" for the model to continue.
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"} _ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
@@ -84,10 +88,14 @@ class OpenAICompatibleProvider(Provider):
def _record_usage(self, payload: dict) -> None: def _record_usage(self, payload: dict) -> None:
"""Records the endpoint's own token accounting, if it reported any. """Records the endpoint's own token accounting, if it reported any.
OpenRouter now always reports usage, and `usage: {include: true}` and In a stream the usage arrives on a final chunk that carries no choices,
`stream_options` are deprecated and do nothing. In a stream the usage which is why this is read separately from the text extraction.
arrives on a final chunk that carries no choices, which is why this is
read separately from the text extraction. v1.1 WP-A1: Ollama sends that chunk only when asked. Measured on Ollama
0.33: a stream with no `stream_options` carried no usage at all, and not
one of the 514 AI turns in the v1 evidence had a count stored. Every
streaming body therefore sets `stream_options.include_usage`
(`STREAM_OPTIONS`), and the turn compares the count with what it sent.
""" """
usage = payload.get("usage") usage = payload.get("usage")
if isinstance(usage, dict) and usage: if isinstance(usage, dict) and usage:
@@ -102,6 +110,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature, "temperature": temperature,
"max_tokens": max_tokens, "max_tokens": max_tokens,
"stream": True, "stream": True,
"stream_options": dict(STREAM_OPTIONS),
} }
else: else:
url = f"{self.base_url}/chat/completions" url = f"{self.base_url}/chat/completions"
@@ -114,6 +123,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature, "temperature": temperature,
"max_tokens": max_tokens, "max_tokens": max_tokens,
"stream": True, "stream": True,
"stream_options": dict(STREAM_OPTIONS),
} }
return url, body return url, body
@@ -183,6 +193,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature, "temperature": temperature,
"max_tokens": max_tokens, "max_tokens": max_tokens,
"stream": True, "stream": True,
"stream_options": dict(STREAM_OPTIONS),
} }
else: else:
url = f"{self.base_url}/chat/completions" url = f"{self.base_url}/chat/completions"
@@ -192,6 +203,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature, "temperature": temperature,
"max_tokens": max_tokens, "max_tokens": max_tokens,
"stream": True, "stream": True,
"stream_options": dict(STREAM_OPTIONS),
} }
async for event in self._stream(url, body): async for event in self._stream(url, body):
yield event yield event
+38 -3
View File
@@ -6,6 +6,7 @@ lock guards one set only while one module owns it. And a test that replaces
`OpenAICompatibleProvider` or `generate_turn` patches this module, which every `OpenAICompatibleProvider` or `generate_turn` patches this module, which every
caller reads through. caller reads through.
""" """
import logging
import threading import threading
from fastapi import Depends, HTTPException, Request from fastapi import Depends, HTTPException, Request
@@ -28,6 +29,8 @@ from .deps import CurrentUser, current_adventure, router
from .nodes import _move_to_after, next_depth from .nodes import _move_to_after, next_depth
from .paging import annotate_takes from .paging import annotate_takes
log = logging.getLogger(__name__)
def world_delta_of(snapshot: dict | None) -> dict | None: def world_delta_of(snapshot: dict | None) -> dict | None:
"""Returns the bulk-read slice of a context snapshot, for `Action.world_delta`. """Returns the bulk-read slice of a context snapshot, for `Action.world_delta`.
@@ -212,8 +215,18 @@ async def _generate_turn(
# network calls — and cached per endpoint and model, so it costs one short # network calls — and cached per endpoint and model, so it costs one short
# request per session rather than one per turn. An unverified window does # request per session rather than one per turn. An unverified window does
# not block the turn; it is recorded as unverified in the snapshot below. # not block the turn; it is recorded as unverified in the snapshot below.
window = await contextwindow.probe(settings.endpoint_url, settings.model, #
declared=settings.context_window_override) # v1.1 WP-A1 corrective: a model that is not resident cannot report its window,
# and a turn built to the configured budget against it was silently cut in the
# A1 evidence (13,875 tokens sent, 2,050 read). So an unverified window gets
# one bounded attempt to load the model, and one more probe, before the
# prompt is assembled. No story text is generated by it and nothing is
# written. A window still unverified afterwards changes nothing below.
window, preflight = await contextwindow.ensure_window(
settings.endpoint_url, settings.model,
declared=settings.context_window_override,
warm_timeout=float(settings.model_timeout_seconds or 300),
)
try: try:
system_text, story_text, snapshot = build_context( system_text, story_text, snapshot = build_context(
adventure, adventure,
@@ -232,6 +245,9 @@ async def _generate_turn(
yield turn_error(str(exc)) yield turn_error(str(exc))
return return
if isinstance(snapshot.get("window"), dict):
snapshot["window"]["preflight"] = preflight
parts = PromptParts(system=system_text, story=story_text) parts = PromptParts(system=system_text, story=story_text)
provider = OpenAICompatibleProvider( provider = OpenAICompatibleProvider(
@@ -318,6 +334,24 @@ async def _generate_turn(
# prompt came from cache rather than being billed in full. This is recorded # prompt came from cache rather than being billed in full. This is recorded
# per attempt, next to the prompt it priced. # per attempt, next to the prompt it priced.
snapshot["usage"] = provider.last_usage snapshot["usage"] = provider.last_usage
# v1.1 WP-A1: what the server says it read, against what was sent. Recorded
# and shown, never acted on: the narration has already streamed to the
# reader, and discarding an accepted turn over an accounting discrepancy
# would lose story to hide a problem. A server that cut the prompt answers
# 200 either way, so this record is the only place the cut is visible.
tokens = snapshot.get("tokens") or {}
accounting = contextwindow.classify_usage(
provider.last_usage,
estimate=tokens.get("estimate") or tokens.get("total") or 0,
budget=tokens.get("budget") or settings.context_token_budget,
max_output_tokens=settings.max_output_tokens,
window_verified=bool((snapshot.get("window") or {}).get("verified")),
)
snapshot["accounting"] = accounting
if accounting["status"] in (contextwindow.EXCEEDED,
contextwindow.TRUNCATION_SUSPECTED):
log.warning("turn accounting for adventure %s: %s — %s",
adventure.id, accounting["status"], accounting["detail"])
reasoning = "".join(reasoning_chunks).strip() or None reasoning = "".join(reasoning_chunks).strip() or None
ai_action = models.Action( ai_action = models.Action(
@@ -381,7 +415,8 @@ async def _generate_turn(
db.commit() db.commit()
db.refresh(ai_action) db.refresh(ai_action)
yield _SAVED yield _SAVED
yield sse({"type": "done", "action": action_json(ai_action, db)}) yield sse({"type": "done", "action": action_json(ai_action, db),
"accounting": accounting})
# Phase 6: schedule summarization and embedding without waiting for them. # Phase 6: schedule summarization and embedding without waiting for them.
# The task opens its own database session. # The task opens its own database session.
memorybank.schedule_post_turn(adventure) memorybank.schedule_post_turn(adventure)
+21 -4
View File
@@ -246,11 +246,28 @@ def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
assert report_before["history"]["floor_depth"] is not None, ( assert report_before["history"]["floor_depth"] is not None, (
"this fixture is meant to be over budget; trimming never engaged") "this fixture is meant to be over budget; trimming never engaged")
_play_one_more(db, adventure, 60) # v1.1 WP-A1: the fixture used to be positioned so that the very next turn
_, after, report_after = _builder.build_context(adventure, settings) # held the floor. The safety reserve takes 256 tokens of this 2,048 budget,
# the block is now the minimum of two, and the next turn is a step. So walk
# forward until a turn holds, requiring every move on the way to be exactly
# one block: a window that slides by one action every turn fails either way.
held = None
depth = 60
for _ in range(4):
_play_one_more(db, adventure, depth)
depth += 1
_, after, report_after = _builder.build_context(adventure, settings)
floor_before = report_before["history"]["floor_depth"]
floor_after = report_after["history"]["floor_depth"]
block = report_after["history"]["trim_block"]
assert floor_after - floor_before in (0, block), (floor_before, floor_after, block)
if floor_after == floor_before:
held = (before, after)
break
before, report_before = after, report_after
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"] assert held is not None, "the floor never held across a turn"
assert _shared_prefix(before, after) > 0.85 assert _shared_prefix(*held) > 0.85
def test_without_a_stable_floor_the_prefix_collapses(saturated): def test_without_a_stable_floor_the_prefix_collapses(saturated):
+348
View File
@@ -0,0 +1,348 @@
"""v1.1 WP-A1 corrective: a cold model is loaded, not guessed about.
A1's accounting caught a real cold-model turn: `/api/ps` knew nothing because the
model was not resident, `/api/show` found no `num_ctx`, the window was therefore
unverified, and the prompt was built to the configured 16,384. Ollama loaded the
model at its own 4,096 default, read 2,050 of the 13,875 tokens and answered 200.
Detection was right. The case is also preventable: once the model is loaded its
window is readable. So before an unverified turn is assembled, the application
asks the configured server, once, to load the model (`POST /api/generate` with a
model and no prompt, which Ollama answers with `"done_reason": "load"` and no
text), probes again, and builds the turn to whatever that probe says. A window
still unverified afterwards changes nothing: the configured budget stands and
the post-response accounting still watches for a cut prompt.
python -m pytest tests/test_v11_cold_window.py -v
"""
import asyncio
import httpx
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy import text
from sqlalchemy.orm import undefer
from app import auth, contextwindow, limits, models
from app.context import builder
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.providers.base import ProviderError
from app.routers import adventures
from fakes import ScriptedProvider
ENDPOINT = "http://127.0.0.1:11434/v1"
MODEL = "qwen2.5:3b-instruct"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
class ColdOllama:
"""The shapes a real Ollama 0.33 returned, with a model that starts cold.
`/api/ps` lists only loaded models. `/api/show` carries no `num_ctx`.
`/api/generate` with no prompt loads the model at `load_window`, exactly as
the real server answered: HTTP 200, `"response": ""`, `"done_reason": "load"`.
"""
def __init__(self, *, loaded=None, load_window=4096, generate_status=200,
report_after_load=True):
self.loaded = dict(loaded or {})
self.load_window = load_window
self.generate_status = generate_status
self.report_after_load = report_after_load
self.requests: list[tuple[str, str, dict | None]] = []
def handler(self, request: httpx.Request) -> httpx.Response:
body = None
if request.content:
import json
body = json.loads(request.content)
self.requests.append((request.method, str(request.url), body))
path = request.url.path
if path == "/api/ps":
return httpx.Response(200, json={"models": [
{"name": name, "model": name, "context_length": tokens}
for name, tokens in self.loaded.items()
]})
if path == "/api/show":
return httpx.Response(200, json={
"model_info": {"qwen2.context_length": 32768}, "parameters": ""})
if path == "/api/generate":
if self.generate_status != 200:
return httpx.Response(self.generate_status, json={"error": "model not found"})
if self.report_after_load:
self.loaded[body["model"]] = self.load_window
return httpx.Response(200, json={
"model": body["model"], "response": "", "done": True, "done_reason": "load"})
return httpx.Response(404)
def paths(self):
return [httpx.URL(url).path for _method, url, _body in self.requests]
@pytest.fixture()
def server(monkeypatch):
def install(fake: ColdOllama):
original = httpx.AsyncClient
def build(*args, **kwargs):
kwargs.pop("verify", None)
return original(*args, transport=httpx.MockTransport(fake.handler), **kwargs)
monkeypatch.setattr(contextwindow.httpx, "AsyncClient", build)
return fake
return install
# ------------------------------------------------------------ ensure_window
def test_a_cold_model_is_loaded_once_and_its_window_verified(server):
fake = server(ColdOllama(loaded={}, load_window=4096))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert (window.tokens, window.source, window.verified) == (4096, contextwindow.LOADED, True)
assert preflight == {"attempted": True, "loaded": True, "verified_before": False,
"verified_after": True,
"detail": "the server loaded the model (load)"}
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate", "/api/ps"]
# One load request, naming the model and nothing else: no prompt, so no text.
warms = [body for _m, url, body in fake.requests if url.endswith("/api/generate")]
assert warms == [{"model": MODEL}]
def test_an_already_loaded_model_is_not_warmed(server):
fake = server(ColdOllama(loaded={MODEL: 16384}))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert window.verified and window.tokens == 16384
assert preflight["attempted"] is False
assert "/api/generate" not in fake.paths()
def test_a_model_that_loads_but_still_cannot_be_read_stays_unverified(server):
"""A server that loads the model but whose `/api/ps` still cannot say. The
existing unknown path stands: no guessed window, the configured budget kept."""
fake = server(ColdOllama(loaded={}, report_after_load=False))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert not window.verified and window.tokens is None
assert preflight["attempted"] is True and preflight["loaded"] is True
assert preflight["verified_after"] is False
assert fake.paths().count("/api/generate") == 1
assert contextwindow.effective_budget(16384, window) == 16384
@pytest.mark.parametrize("status", [404, 500])
def test_a_failed_load_is_recorded_and_leaves_the_window_unverified(server, status):
fake = server(ColdOllama(loaded={}, generate_status=status))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert not window.verified
assert preflight["attempted"] is True and preflight["loaded"] is False
assert f"HTTP {status}" in preflight["detail"]
# Bounded: one attempt, and no second probe after a failed load.
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate"]
def test_a_declared_window_does_not_stop_the_server_being_asked(server):
"""A declaration fills a hole the server leaves. Loading the model can close
the hole, and a verified answer always wins over a declaration."""
server(ColdOllama(loaded={}, load_window=4096))
window, _preflight = asyncio.run(
contextwindow.ensure_window(ENDPOINT, MODEL, declared=8192))
assert (window.tokens, window.source) == (4096, contextwindow.LOADED)
def test_an_unreachable_server_is_not_asked_to_load_anything():
"""Nothing listens here. No load is attempted against a server that did not
answer the probe, so an offline turn costs no second timeout."""
window, preflight = asyncio.run(
contextwindow.ensure_window("http://127.0.0.1:1/v1", MODEL))
assert not window.verified
assert preflight["attempted"] is False
assert "did not answer" in preflight["detail"]
def test_the_load_request_obeys_the_endpoint_policy():
"""ADR 011. No transport is installed: a load that ignored the policy would
try to reach a public address for real."""
for url in ("https://api.openai.com/v1", "http://8.8.8.8:11434/v1"):
loaded, detail = asyncio.run(contextwindow.warm(url, MODEL, timeout=2))
assert loaded is False
assert "not allowed" in detail
def test_the_load_request_goes_only_to_the_configured_host(server):
fake = server(ColdOllama(loaded={}))
asyncio.run(contextwindow.ensure_window("http://192.168.0.50:11434/v1", MODEL))
hosts = {httpx.URL(url).host for _m, url, _b in fake.requests}
ports = {httpx.URL(url).port for _m, url, _b in fake.requests}
assert hosts == {"192.168.0.50"} and ports == {11434}
# ------------------------------------------------------------- end to end
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="v11cold@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
context_token_budget=16384, max_output_tokens=500,
))
adventure = models.Adventure(
user_id=user.id, title="Cold",
campaign_canon={"rules": ["The sealed crypt is named CANON-SENTINEL-COLD-2050."]},
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(adventure_id=adventure.id, type="start", text="Rain."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
def _long_story(adv_id, turns=120):
from app import tree
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv_id)
for i in range(turns):
for kind, body in (
("do", f"I search the {i}th chamber of the undercroft."),
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
):
action = models.Action(adventure_id=adv_id, type=kind, text=body)
db.add(action)
db.flush()
tree.place_action(db, adventure, action)
db.commit()
def _counts():
with SessionLocal() as db:
return {table: db.execute(text(f"SELECT COUNT(*) FROM {table}")).scalar()
for table in ("actions", "state_events", "state_proposals", "memories",
"summaries")}
def _latest_ai(adv_id):
with SessionLocal() as db:
return (db.query(models.Action)
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
def test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports(client, server):
"""The observed failure, prevented. Without the load this turn would be built
to the configured 16,384 against a 4,096 server."""
_long_story(client.adv_id)
fake = server(ColdOllama(loaded={}, load_window=4096))
ScriptedProvider.replies = ["The seal holds."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is True
assert snapshot["tokens"]["budget"] == 4096
assert snapshot["window"]["preflight"]["attempted"] is True
assert snapshot["window"]["preflight"]["verified_after"] is True
system, story = ScriptedProvider.prompts[-1]
sent = builder.count_tokens(system) + builder.count_tokens(story)
assert sent + snapshot["tokens"]["transport"] + 500 + 256 <= 4096
assert "CANON-SENTINEL-COLD-2050" in system
assert fake.paths().count("/api/generate") == 1
def test_without_the_load_the_same_cold_turn_would_have_been_built_too_large(client, server,
monkeypatch):
"""The negative control: v1.0.0 and the first A1 tree probed only."""
_long_story(client.adv_id)
server(ColdOllama(loaded={}, load_window=4096))
async def probe_only(endpoint_url, model, *, declared=None, warm_timeout=300.0):
window = await contextwindow.probe(endpoint_url, model, declared=declared)
return window, {"attempted": False}
monkeypatch.setattr(adventures.turns.contextwindow, "ensure_window", probe_only)
ScriptedProvider.replies = ["The seal holds."]
client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is False
assert snapshot["tokens"]["budget"] == 16384
system, story = ScriptedProvider.prompts[-1]
assert builder.count_tokens(system) + builder.count_tokens(story) > 4096 * 2
def test_the_load_itself_writes_nothing(client, server):
"""No action, narration, state event, proposal, memory or summary comes from
the preflight: it is a request to the server and nothing else."""
fake = server(ColdOllama(loaded={}))
before = _counts()
asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert _counts() == before
assert fake.paths().count("/api/generate") == 1
def test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe(client, server):
"""The ordinary failure semantics: the error is reported, no narration is
accepted, and nothing about the state changes."""
server(ColdOllama(loaded={}, generate_status=404))
before = _counts()
ScriptedProvider.replies = [ProviderError("Endpoint or model not found (HTTP 404).")]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the door"})
assert response.status_code == 200
assert '"type": "error"' in response.text or '"error"' in response.text
after = _counts()
assert after["state_events"] == before["state_events"]
assert after["state_proposals"] == before["state_proposals"]
with SessionLocal() as db:
assert db.query(models.Action).filter_by(adventure_id=client.adv_id,
type="ai").count() == 0
def test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer(client, server):
server(ColdOllama(loaded={}, generate_status=500))
ScriptedProvider.replies = ["The door opens."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the door"})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is False
assert snapshot["window"]["preflight"]["loaded"] is False
assert snapshot["tokens"]["budget"] == 16384
assert snapshot["accounting"]["status"] == contextwindow.UNKNOWN
def test_the_context_dry_run_never_loads_a_model(client, server):
fake = server(ColdOllama(loaded={}))
response = client.get(f"/api/adventures/{client.adv_id}/context")
assert response.status_code == 200
assert "/api/generate" not in fake.paths()
+495
View File
@@ -0,0 +1,495 @@
"""v1.1 WP-A1: a deliberate safety reserve, and a turn the server cut is not silent.
M11 made the verified window a ceiling. It did not make the application's count
the server's count. The application counts with `cl100k_base`, the narrator with
its own tokenizer, and the v1 evidence left 23-42 real tokens between the largest
prompt and the edge of a 16,384 window. Past that edge Ollama does not refuse.
Measured against the reference CPU host (Ollama 0.33, a 4,096 window), a
6,316-token prompt came back 200 with `prompt_tokens` 2,050: the front of the
prompt, which in this design is the narrator's rules and the canon, was gone.
So the tests below are in three halves.
**The reserve.** `max(256, ceil(5% of the effective window))`, taken from the
budget before any history is chosen, on top of an exact reply allocation.
**The arithmetic.** The assembled prompt, plus the application text the provider
adds to every request, plus the reply allocation, plus the reserve, fits the
effective window. Protected context that cannot fit that way fails before the
model is called.
**The accounting.** Where the server reports how many prompt tokens it read, the
turn records `fits`, `exceeded` or `truncation_suspected`. Where it reports
nothing, the turn says `unknown`, never `fits`. A discrepancy found after the
reply is recorded and shown; it never costs the reader an accepted turn.
python -m pytest tests/test_v11_context_reserve.py -v
"""
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, contextwindow, limits, models
from app.context import builder
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.providers.openai_compatible import CHAT_CONTINUE_HINT, OpenAICompatibleProvider
from app.routers import adventures
from fakes import ScriptedProvider
ENDPOINT = "http://127.0.0.1:11434/v1"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
# ------------------------------------------------------------- the reserve
@pytest.mark.parametrize("window, reserve", [
(1024, 256),
(4096, 256), # 5% is 204.8, so the floor holds
(5120, 256), # exactly 5% is the floor
(5121, 257), # 256.05 rounds up
(8192, 410), # 409.6 rounds up
(16384, 820), # 819.2 rounds up
(32768, 1639), # 1638.4 rounds up
])
def test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up(window, reserve):
assert contextwindow.safety_reserve(window) == reserve
def test_the_reserve_is_far_larger_than_the_v1_margin_at_the_evidence_window():
"""The v1 evidence left 23-42 tokens at 16,384. 64 tokens of slack was all
the arithmetic kept for drift and separators together."""
assert contextwindow.safety_reserve(16384) >= 10 * 64
# ---------------------------------------------------------- the arithmetic
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="v11reserve@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model="qwen2.5:3b-instruct", endpoint_url=ENDPOINT,
embedding_model="", context_token_budget=16384, max_output_tokens=500,
))
adventure = models.Adventure(
user_id=user.id, title="Reserved",
campaign_canon={"rules": [
"The abbey seal has never been broken.",
"The sealed crypt is named CANON-SENTINEL-RESERVE-5120.",
]},
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(
adventure_id=adventure.id, type="start", text="Rain over Westhaven."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
def _long_story(adv_id, turns=120):
from app import tree
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv_id)
for i in range(turns):
for kind, text in (
("do", f"I search the {i}th chamber of the undercroft."),
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
):
action = models.Action(adventure_id=adv_id, type=kind, text=text)
db.add(action)
db.flush()
tree.place_action(db, adventure, action)
db.commit()
def _settings(**changes):
with SessionLocal() as db:
settings = db.query(models.Settings).first()
for key, value in changes.items():
setattr(settings, key, value)
db.commit()
def _build(client, window):
with SessionLocal() as db:
adventure = db.get(models.Adventure, client.adv_id)
settings = db.query(models.Settings).first()
return builder.build_context(adventure, settings, window=window)
def _sent(system, story) -> int:
"""What the provider actually sends in chat mode, by the application's count."""
return (builder.count_tokens(system) + builder.count_tokens(story)
+ builder.count_tokens(CHAT_CONTINUE_HINT))
CONFIGURATIONS = {
"verified 4,096": (dict(context_token_budget=16384),
contextwindow.Window(4096, contextwindow.LOADED)),
"verified 8,192": (dict(context_token_budget=16384),
contextwindow.Window(8192, contextwindow.PARAMETERS)),
"verified 16,384": (dict(context_token_budget=16384),
contextwindow.Window(16384, contextwindow.LOADED)),
"declared 6,000": (dict(context_token_budget=16384),
contextwindow.Window(6000, contextwindow.DECLARED)),
"unverified, configured 12,000": (dict(context_token_budget=12000),
contextwindow.UNVERIFIED),
}
@pytest.mark.parametrize("name", list(CONFIGURATIONS))
def test_the_prompt_leaves_the_reply_and_the_reserve_free(client, name):
"""A1-2, on the assembled text rather than the builder's own arithmetic."""
changes, window = CONFIGURATIONS[name]
_settings(**changes)
_long_story(client.adv_id, turns=120)
system, story, report = _build(client, window)
tokens = report["tokens"]
budget = tokens["budget"]
assert tokens["safety_reserve"] == contextwindow.safety_reserve(budget)
assert tokens["output_reserve"] == 500
sent = _sent(system, story)
assert sent + tokens["output_reserve"] + tokens["safety_reserve"] <= budget, (
name, sent, tokens)
# The history is what gave way, not the canon.
assert "CANON-SENTINEL-RESERVE-5120" in system
assert report["history"]["included"] < report["history"]["total"]
def test_the_report_prices_the_text_the_provider_adds(client):
"""The chat hint rides on every request and was never counted."""
_, _, report = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
tokens = report["tokens"]
assert tokens["transport"] >= builder.count_tokens(CHAT_CONTINUE_HINT)
assert tokens["estimate"] == tokens["total"] + builder.count_tokens(CHAT_CONTINUE_HINT)
def test_the_reserve_follows_the_effective_window_not_the_setting(client):
"""5% of a 4,096 server, not 5% of a 16,384 setting it will never read."""
_, _, capped = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
_, _, full = _build(client, contextwindow.Window(16384, contextwindow.LOADED))
assert capped["tokens"]["safety_reserve"] == 256
assert full["tokens"]["safety_reserve"] == 820
def test_protected_context_that_only_fits_without_the_reserve_fails_explicitly(client):
"""A1-3 in the builder. Before v1.1 this prompt would have been built.
The canon is sized so that protected text plus the reply fits a 4,096 window
with room to spare, and does not fit once the 256-token reserve is taken.
"""
small = contextwindow.Window(4096, contextwindow.LOADED)
# Measured with a window large enough never to overflow, because repeated
# text merges tokens at its seams and cannot be priced by multiplication.
roomy = contextwindow.Window(32768, contextwindow.LOADED)
rules = None
with SessionLocal() as db:
adventure = db.get(models.Adventure, client.adv_id)
settings = db.query(models.Settings).first()
base_rules = list(adventure.campaign_canon["rules"])
filler = "The bell tolls once for every name in the ledger."
copies = 1
while True:
candidate = base_rules + [" ".join([filler] * copies)]
adventure.campaign_canon = {"rules": candidate}
_, _, measured = builder.build_context(adventure, settings, window=roomy)
t = measured["tokens"]
# What protected context costs at 4,096, without the reserve.
without_reserve = t["protected"] + t["transport"] + t["output_reserve"]
if without_reserve + 64 >= 4096 - 60:
break
copies += 1
rules = candidate
adventure.campaign_canon = {"rules": rules}
db.commit()
# The case this test is about: v1's arithmetic, with its 64-token margin,
# would have built this prompt. v1.1's reserve does not fit.
assert without_reserve + 64 < 4096
assert without_reserve + contextwindow.safety_reserve(4096) >= 4096
with pytest.raises(builder.ContextOverflow) as caught:
builder.build_context(adventure, settings, window=small)
message = str(caught.value)
assert "safety" in message
assert "load the model with a larger window" in message
def test_an_overflowing_turn_never_reaches_the_model(client, monkeypatch):
"""A1-3 end to end: the refusal happens before the provider is called."""
async def verified(endpoint, model, declared=None, use_cache=True):
return contextwindow.Window(1024, contextwindow.LOADED)
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
ScriptedProvider.replies = ["This must never be generated."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the crypt"})
assert response.status_code == 200
assert "safety" in response.text
assert ScriptedProvider.calls == 0
with SessionLocal() as db:
assert db.query(models.Action).filter_by(
adventure_id=client.adv_id, type="ai").count() == 0
# ---------------------------------------------------------- the accounting
def _classify(prompt_tokens=None, *, usage=None, estimate=3500, budget=4096,
output=500, verified=True):
if usage is None and prompt_tokens is not None:
usage = {"prompt_tokens": prompt_tokens, "completion_tokens": 40}
return contextwindow.classify_usage(
usage, estimate=estimate, budget=budget, max_output_tokens=output,
window_verified=verified,
)
def test_a_prompt_the_server_read_in_full_fits():
# The 13-token chat-template overhead measured against the real server.
result = _classify(3513)
assert result["status"] == contextwindow.FITS
assert result["server_prompt_tokens"] == 3513
assert result["difference"] == 13
assert result["safety_reserve"] == 256
assert result["observed_margin"] == 4096 - 500 - 3513
def test_a_server_that_counts_more_than_the_reserve_allows_is_exceeded():
"""The prompt plus the reply allocation no longer fits the window."""
result = _classify(3700)
assert result["status"] == contextwindow.EXCEEDED
assert result["observed_margin"] < 0
def test_a_server_that_read_far_less_than_was_sent_is_suspected_of_truncating():
"""The real shape: 6,316 sent, 2,050 read, HTTP 200, no error."""
result = _classify(2050, estimate=6316)
assert result["status"] == contextwindow.TRUNCATION_SUSPECTED
assert result["difference"] == 2050 - 6316
def test_a_small_undercount_is_tokenizer_drift_not_truncation():
"""A tokenizer thriftier than `cl100k_base` reads fewer tokens honestly. Only
a shortfall larger than the reserve is called truncation."""
assert _classify(3500 - 255)["status"] == contextwindow.FITS
assert _classify(3500 - 257)["status"] == contextwindow.TRUNCATION_SUSPECTED
@pytest.mark.parametrize("usage", [
None,
{},
{"completion_tokens": 40},
{"prompt_tokens": 0},
{"prompt_tokens": "3500"},
{"prompt_tokens": -1},
])
def test_no_usable_count_is_unknown_never_fits(usage):
result = _classify(usage=usage)
assert result["status"] == contextwindow.UNKNOWN
assert result["server_prompt_tokens"] is None
assert result["observed_margin"] is None
def test_the_accounting_says_when_the_window_itself_was_not_verified():
result = _classify(3513, verified=False)
assert result["status"] == contextwindow.FITS
assert result["window_verified"] is False
assert "not verified" in result["detail"]
def test_the_stream_asks_the_server_to_report_its_usage():
"""Measured: Ollama 0.33 sends no usage in a stream unless asked."""
provider = OpenAICompatibleProvider(ENDPOINT, "m")
from app.providers.base import PromptParts
for mode in ("chat", "completion"):
provider.api_mode = mode
_url, body = provider._request(PromptParts(system="s", story="t"), 0.7, 50)
assert body["stream"] is True
assert body["stream_options"] == {"include_usage": True}
def _latest_ai(adv_id):
with SessionLocal() as db:
return (
db.query(models.Action)
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first()
)
def _play(client, monkeypatch, usage, window=4096, reply="The crypt is still sealed."):
async def verified(endpoint, model, declared=None, use_cache=True):
return contextwindow.Window(window, contextwindow.LOADED, 32768, "fake")
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
monkeypatch.setattr(ScriptedProvider, "last_usage", usage)
ScriptedProvider.replies = [reply]
return client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
def test_a_turn_records_what_the_server_read(client, monkeypatch):
response = _play(client, monkeypatch, None)
assert response.status_code == 200, response.text[:300]
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
response = _play(client, monkeypatch,
{"prompt_tokens": estimate + 13, "completion_tokens": 9})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
accounting = snapshot["accounting"]
assert accounting["status"] == contextwindow.FITS
assert accounting["server_prompt_tokens"] == estimate + 13
assert accounting["estimate"] == snapshot["tokens"]["estimate"]
assert '"accounting"' in response.text
assert contextwindow.FITS in response.text
def test_a_turn_with_no_reported_usage_is_unknown(client, monkeypatch):
response = _play(client, monkeypatch, None)
assert response.status_code == 200
assert _latest_ai(client.adv_id).context_snapshot["accounting"]["status"] == (
contextwindow.UNKNOWN)
def test_a_suspected_truncation_keeps_the_turn_and_says_so(client, monkeypatch, caplog):
"""A1-7 and A1-8. The reader watched the narration arrive; it stays."""
response = _play(client, monkeypatch, {"prompt_tokens": 12, "completion_tokens": 9},
reply="The seal holds, and the rain goes on.")
assert response.status_code == 200, response.text[:300]
action = _latest_ai(client.adv_id)
assert action is not None
assert action.text == "The seal holds, and the rain goes on."
accounting = action.context_snapshot["accounting"]
assert accounting["status"] == contextwindow.TRUNCATION_SUSPECTED
assert contextwindow.TRUNCATION_SUSPECTED in response.text
assert any(contextwindow.TRUNCATION_SUSPECTED in r.getMessage() for r in caplog.records)
# Inspectable afterwards through the same route the context panel reads.
context = client.get(
f"/api/adventures/{client.adv_id}/actions/{action.id}/context")
assert context.status_code == 200
assert context.json()["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
def test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves():
"""Found by the A2 long run. Accounting belongs to one API call, not to the
turn's shared prompt. A retry demotes the old attempt, and a take selection
hands the prompt from one attempt to another. Neither may drop an attempt's
accounting or give it another attempt's."""
from app import attempts
class Node:
def __init__(self, snapshot):
self.context_snapshot = snapshot
shared = {"tokens": {"estimate": 3000}, "sections": [], "window": {"verified": True}}
first = Node(shared | {"raw_output": "one", "usage": {"prompt_tokens": 3015},
"accounting": {"status": contextwindow.FITS, "server_prompt_tokens": 3015}})
second = Node({"raw_output": "two", "usage": {"prompt_tokens": 12},
"accounting": {"status": contextwindow.TRUNCATION_SUSPECTED,
"server_prompt_tokens": 12}})
# Superseded by a retry: the old attempt keeps only its own slices.
attempts.keep_own_slices(Node(dict(first.context_snapshot)))
demoted = Node(dict(first.context_snapshot))
attempts.keep_own_slices(demoted)
assert demoted.context_snapshot["accounting"]["server_prompt_tokens"] == 3015
assert "tokens" not in demoted.context_snapshot
# The prompt moves to the second attempt; each keeps its own accounting.
attempts.hand_over_the_prompt(first, second)
assert second.context_snapshot["tokens"] == {"estimate": 3000}
assert second.context_snapshot["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
assert second.context_snapshot["accounting"]["server_prompt_tokens"] == 12
assert first.context_snapshot["accounting"]["status"] == contextwindow.FITS
assert "tokens" not in first.context_snapshot
def test_a_retry_leaves_each_take_with_its_own_accounting(client, monkeypatch):
"""End to end, through the real retry route. Before the fix the live take
inherited the superseded take's accounting, so the inspector could show one
call's server count as another's."""
response = _play(client, monkeypatch, None, reply="The first take.")
assert response.status_code == 200, response.text[:300]
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
# Replay the first take with a real count, so it has accounting of its own.
with SessionLocal() as db:
first = (db.query(models.Action)
.filter(models.Action.adventure_id == client.adv_id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot)).first())
snapshot = dict(first.context_snapshot)
snapshot["accounting"] = contextwindow.classify_usage(
{"prompt_tokens": estimate + 15}, estimate=estimate, budget=4096,
max_output_tokens=500, window_verified=True)
first.context_snapshot = snapshot
db.commit()
first_id = first.id
monkeypatch.setattr(ScriptedProvider, "last_usage",
{"prompt_tokens": 12, "completion_tokens": 9})
ScriptedProvider.replies = ["The second take."]
retried = client.post(f"/api/adventures/{client.adv_id}/retry")
assert retried.status_code == 200, retried.text[:300]
with SessionLocal() as db:
rows = (db.query(models.Action)
.filter(models.Action.adventure_id == client.adv_id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id).all())
by_id = {row.id: row for row in rows}
old = by_id[first_id]
new = [row for row in rows if row.id != first_id][-1]
assert new.text == "The second take."
assert new.live and not old.live
# The superseded take keeps its own accounting and gives up the prompt.
assert old.context_snapshot["accounting"]["status"] == contextwindow.FITS
assert old.context_snapshot["accounting"]["server_prompt_tokens"] == estimate + 15
assert "tokens" not in old.context_snapshot
# The live take carries the prompt and its own accounting, not the old one's.
assert "tokens" in new.context_snapshot
assert new.context_snapshot["accounting"]["status"] == (
contextwindow.TRUNCATION_SUSPECTED)
assert new.context_snapshot["accounting"]["server_prompt_tokens"] == 12
def test_an_exceeded_turn_is_also_kept(client, monkeypatch):
response = _play(client, monkeypatch, {"prompt_tokens": 3900, "completion_tokens": 9})
assert response.status_code == 200
action = _latest_ai(client.adv_id)
assert action.text == "The crypt is still sealed."
assert action.context_snapshot["accounting"]["status"] == contextwindow.EXCEEDED
+324
View File
@@ -0,0 +1,324 @@
"""v1.1 WP-A2: the protocol a narrator copies stays out of the story, and nothing else does.
The M11 closeout's identity run (an office meeting, a 3B narrator, a 4,096
window) stored four turns carrying text the application wrote, not the story:
- `> Create_entity(new_person, "john", …` — the event vocabulary as the prompt
printed it, `name(field, …)`, copied as if it were a call;
- `> set_possession(silver-key, "alice") Adds the silver key to Alice's
possession.` — the call again, naming the fantasy example slug from the fixed
state rule, in a meeting room;
- `Scene: Bill, Alice, … (at The meeting room)` — the renderer's own scene line;
- `[Hard limit: your next turn must not exceed 180 words, … append the state
block well inside the limit.]` — the length hint, reworded at the front and
verbatim at the end.
v1.0.0 removed none of them. The rule this module is held to is unchanged from
M5: **removing story is worse than leaving protocol.** Every removal below is
anchored to a string or a vocabulary the application owns, and every one has
story beside it that must survive.
python -m pytest tests/test_v11_protocol_echo.py -v
"""
import json
import re
import pytest
from app.context import builder
from app.narrative import events, extract, render
# ---------------------------------------------------------------- the prompt
#: The identifiers the v1 state rule taught every campaign, from the fantasy
#: acceptance fixture. None may come back into a fixed instruction.
FANTASY_IDENTIFIERS = ("mara", "silver-key", "silver key", "old-abbey", "abbey",
"aldric", "westhaven", "crypt", "edrin")
#: And nothing from the science-fiction fixture either: neutral means neutral,
#: not "the other genre".
SCIFI_IDENTIFIERS = ("persephone", "imani", "data-crystal", "data crystal", "airlock")
def _fixed_instructions() -> str:
return "\n".join([
extract.EMIT_RULE,
extract.EMIT_REMINDER,
events.vocabulary_for_prompt(),
builder.length_hint(500),
builder.length_hint(500, "brief"),
builder.length_hint(500, "long"),
builder.length_hint(120),
]).lower()
@pytest.mark.parametrize("identifier", FANTASY_IDENTIFIERS + SCIFI_IDENTIFIERS)
def test_the_fixed_state_instructions_name_no_fixture_identifier(identifier):
"""A2-1. The example slug the office run copied cannot come back."""
assert not re.search(rf"\b{re.escape(identifier)}\b", _fixed_instructions())
def test_the_example_uses_neutral_identifiers():
for neutral in ("character-1", "item-1", "location-1"):
assert neutral in extract.EMIT_RULE
def test_the_worked_example_is_a_block_this_extractor_accepts():
"""The example is the wire format, byte for byte, not an illustration of it."""
example = extract.EMIT_RULE[extract.EMIT_RULE.index("```state"):]
prose, parsed, _raw = extract.split("The door opens.\n\n" + example)
assert prose == "The door opens."
assert isinstance(parsed, dict)
assert [e["type"] for e in parsed["events"]] == [
"set_possession", "set_current_location"]
for event in parsed["events"]:
assert events.is_allowed(event["type"])
def test_the_vocabulary_is_not_written_as_function_calls():
"""A2-2. `set_possession(item, owner)` is the notation the narrator copied."""
vocabulary = events.vocabulary_for_prompt()
for name in events.SPECS:
assert not re.search(rf"\b{name}\s*\(", vocabulary), name
def test_every_event_is_described_in_the_shape_the_model_must_send():
lines = events.vocabulary_for_prompt().splitlines()
assert len(lines) == len(events.SPECS)
for name, line in zip(events.SPECS, lines):
shape = line.strip().split(" — ", 1)[0]
obj = json.loads(shape)
assert obj["type"] == name
assert set(obj) - {"type"} == set(events.SPECS[name]["required"])
for optional in events.SPECS[name]["optional"]:
assert optional in line
def test_the_length_hint_carries_the_phrases_the_extractor_recognises():
"""One source for the words, so the builder and the extractor cannot drift."""
for narration_length in ("", "brief", "medium", "long"):
hint = builder.length_hint(500, narration_length)
assert hint.startswith(extract.LENGTH_HINT_OPENING)
assert extract.LENGTH_HINT_TAIL in hint
# ---------------------------------------------------------- observed shapes
STORY = (
"Alice looks at John, the tension in the room palpable.\n\n"
"John nods. \"I'm ready to contribute.\""
)
@pytest.mark.parametrize("leak", [
# Depth 20: the call, the fantasy slug, and a gloss on the same line.
'> set_possession(silver-key, "alice") Adds the silver key to Alice\'s possession.',
# Depth 10: cut off by the output limit mid-call.
'> Create_entity(new_person, "john", "character", "A determined team member", ["john',
# Unquoted, and a vocabulary name in any case.
'SET_CURRENT_LOCATION(bill, office)',
'add_fact(predicate="knows the plan", subject="alice")',
])
def test_an_event_call_line_at_the_end_leaves_the_story(leak):
prose, parsed, _raw = extract.split(f"{STORY}\n\n{leak}")
assert prose == STORY
assert parsed is None
def test_an_event_call_line_in_the_middle_leaves_and_the_story_after_it_stays():
"""Depth 12: the call, then more narration."""
reply = (
f"{STORY}\n\n"
'> Create_entity(new_person, "mike", "character", "A new team member.", ["mike"])\n\n'
"Mike takes the empty chair by the window."
)
prose, _parsed, _raw = extract.split(reply)
assert prose == f"{STORY}\n\nMike takes the empty chair by the window."
def test_the_depth_fourteen_tail_leaves_entirely():
"""A call, a rendered scene line, and a reworded length hint, in that order."""
reply = (
f"{STORY}\n\n"
'> Create_entity(mike, "character", "A new team member.", ["mike"])\n\n'
"Scene: Bill, Alice, Roger, John, and Mike at the table. (at The meeting room)\n\n"
"[Hard limit: your next turn must not exceed 180 words, and it should not stop "
"short of about 70. Prefer the lower end of that range unless the scene genuinely "
"needs more. Finish the narration and append the state block well inside the limit.]"
)
prose, parsed, _raw = extract.split(reply)
assert prose == STORY
assert parsed is None
@pytest.mark.parametrize("hint", [
builder.length_hint(500),
builder.length_hint(500, "brief"),
# Cut off by the output limit before the tail.
"[Hard limit: this turn must not exceed 180 words, and it should not stop short",
# Reworded at the front, as the 3B narrator did.
"[Hard limit: your next turn must not exceed 506 words. Write only as much as the "
"moment needs — a typical turn is much shorter. Finish the narration and append "
"the state block well inside the limit.]",
])
def test_a_parroted_length_hint_at_the_end_leaves_the_story(hint):
prose, _parsed, _raw = extract.split(f"{STORY}\n\n{hint}")
assert prose == STORY
def test_a_rendered_scene_line_at_the_end_leaves_the_story():
prose, _parsed, _raw = extract.split(
f"{STORY}\n\nScene: A tense budget meeting. (at The meeting room)")
assert prose == STORY
def test_a_fenced_block_with_a_call_line_above_it_still_parses_and_applies():
"""A2-6. The proposal is still read when protocol litter surrounds it."""
reply = (
f"{STORY}\n\n"
'> set_current_location(john, office)\n\n'
'```state\n{"events": [{"type": "set_current_location", '
'"entity": "john", "location": "office"}]}\n```'
)
prose, parsed, raw = extract.split(reply)
assert prose == STORY
assert parsed["events"][0]["entity"] == "john"
assert raw.startswith("{")
# ------------------------------------------------ adversarial story that stays
@pytest.mark.parametrize("reply", [
# The owner's cases.
'The engineer writes "set_power(core, 80)" on the whiteboard.',
'She says, "Create_entity is a terrible name for a company."',
'The old manual contains a heading labeled "Scene:"',
'He reads aloud: "[Hard limit: 500 words]" and laughs.',
# A vocabulary name, written into a story, not at the start of a line.
'Nadia squints at the log: the last command was set_possession(badge, guard).',
# Call-shaped, at the start of a line, but not an event this protocol has.
"The terminal scrolls.\n\n> open_door(north)\n\nNothing happens.",
# A vocabulary call inside the story's own code block is the story's code.
"She types:\n\n```python\ncreate_entity(ship)\nset_possession(key, captain)\n```\n\n"
"The console beeps twice.",
# A bracket at the very end, in-world, that is not the application's hint.
"The warning light blinks.\n\n[Hard limit of the reactor: three hours]",
"The contract ends with a clause.\n\n[Hard limit: forty days, no extensions]",
# A scene heading in a screenplay the characters are writing, mid-story.
"Scene: a kitchen, late.\n\nShe crosses it out and starts again.",
# A last line that starts like the renderer's but is not its shape.
"The director calls it.\n\nScene: take two, and nobody moves.",
# A fact restated inside a sentence.
"Alice knew the badge opened the server room, and said nothing.",
"Memory: she remembered the bells.",
])
def test_story_that_resembles_the_new_rules_is_kept(reply):
"""A2-5."""
prose, parsed, _raw = extract.split(reply)
assert prose == reply
assert parsed is None
# ------------------------------------------------------ the replay attribution
@pytest.mark.parametrize("line, rule", [
('> set_possession(silver-key, "alice") Adds the key.', extract.RULE_EVENT_CALL),
("Create_entity(new_person", extract.RULE_EVENT_CALL),
("[Hard limit: this turn must not exceed 90 words. Finish the narration and append "
"the state block well inside the limit.]", extract.RULE_LENGTH_HINT),
("Scene: A meeting. (at The meeting room)", extract.RULE_SCENE_LINE),
("John nods.", None),
('He reads aloud: "[Hard limit: 500 words]" and laughs.', None),
])
def test_a_removed_line_is_attributed_to_the_rule_that_removes_it(line, rule):
assert extract.explain_removed_line(line) == rule
# ------------------------------------- corrective: the depth-16 instruction tail
#: Cut down from the v1.1 identity diagnostic's depth-16 turn, whose stored text
#: was exactly the extractor's output. The two story paragraphs are shortened;
#: the four trailing lines are verbatim.
DEPTH_16_STORY = (
"John's initial ideas are thoughtful and insightful, and the room fills with a "
"sense of optimism.\n\n"
"John's enthusiasm is contagious, and the meeting room is electric with the "
"excitement of a fruitful collaboration ahead."
)
DEPTH_16_TAIL = (
"Scene: Bill, Alice and Roger at the table; John not yet arrived.\n\n"
"[Hard limit: this is now 180 words.]\n\n"
"[Reminder: end your reply with a `state` block listing the events your narration "
"made true, with absolute values.]\n\n"
"[You don't need to continue; your turn must now be about John entering the room. "
"Continue the story here, directly. Output only story text.]"
)
def test_the_depth_sixteen_instruction_tail_leaves_entirely():
"""The corrective's positive regression. v1.1's first A2 left all four lines:
the last bracket was a reworded continue hint nothing recognised, so nothing
above it was ever at the end."""
prose, parsed, _raw = extract.split(f"{DEPTH_16_STORY}\n\n{DEPTH_16_TAIL}")
assert prose == DEPTH_16_STORY
assert parsed is None
def test_the_continue_hint_phrase_is_the_providers_own_sentence():
from app.providers.openai_compatible import CHAT_CONTINUE_HINT
assert extract.CONTINUE_HINT_PHRASE in CHAT_CONTINUE_HINT
def test_an_echoed_continue_hint_alone_at_the_end_leaves():
prose, _p, _r = extract.split(
f"{STORY}\n\n[Keep going. Continue the story here, directly. Output only story text.]")
assert prose == STORY
@pytest.mark.parametrize("reply", [
# A hint-opened bracket with no echoed instruction below it is in-world.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# Nor does a state block below it make it an instruction.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# The phrase in the middle of a story is prose, not a trailing echo.
'She wrote "output only story text" on the card, then crossed it out.\n\nThe rain went on.',
# A trailing in-world bracket that only resembles a continuation.
f"{STORY}\n\n[To be continued]",
])
def test_story_brackets_near_the_corrective_rule_are_kept(reply):
prose, _p, _r = extract.split(reply)
assert prose == reply
def test_a_hint_opened_bracket_above_a_state_block_is_kept():
reply = (f"{STORY}\n\n[Hard limit: forty days, no extensions]\n\n"
'```state\n{"events": []}\n```')
prose, parsed, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[Hard limit: forty days, no extensions]"
assert parsed == {"events": []}
def test_an_in_world_bracket_above_an_echoed_hint_is_kept():
"""Only a bracket opening the way the application's hint opens is taken with
the echo. Any other bracket above it is the story's."""
reply = (f"{STORY}\n\n[The sign on the door reads: Closed]\n\n"
"[Continue the story here, directly. Output only story text.]")
prose, _p, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[The sign on the door reads: Closed]"
@pytest.mark.parametrize("line, rule", [
("[Hard limit: this is now 180 words.]", extract.RULE_INSTRUCTION_TAIL),
("[Reminder: end your reply with a `state` block listing the events.]",
extract.RULE_INSTRUCTION_TAIL),
("[You don't need to continue. Output only story text.]", extract.RULE_INSTRUCTION_TAIL),
])
def test_the_corrective_rule_is_attributed(line, rule):
assert extract.explain_removed_line(line) == rule
def test_the_new_rules_do_not_disturb_the_section_headings_they_share_a_module_with():
"""The renderer's headings are the M11 rules' anchor. A2 adds none."""
assert render.HEADING_SCENE == "Scene:"
assert "Scene:" not in render.SECTION_HEADINGS
+28 -1
View File
@@ -559,6 +559,17 @@ class Run:
self.accepted += 1 self.accepted += 1
seconds = time.monotonic() - started seconds = time.monotonic() - started
sample = self.measure() sample = self.measure()
# v1.1 WP-A1: what the server said it read for the turn just played, from
# the `done` event. `.get` because a build before v1.1 sends none.
done = next((e for e in events if e.get("type") == "done"), {})
accounting = done.get("accounting") or {}
sample.update({
"accounting_status": accounting.get("status"),
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
"app_prompt_estimate": accounting.get("estimate"),
"observed_margin": accounting.get("observed_margin"),
"safety_reserve": accounting.get("safety_reserve"),
})
self.note("turn", text=text, seconds=round(seconds, 1), **sample) self.note("turn", text=text, seconds=round(seconds, 1), **sample)
return {"accepted": True, "seconds": seconds, **sample} return {"accepted": True, "seconds": seconds, **sample}
@@ -1225,10 +1236,26 @@ PROTOCOL_LEAK_HEADING_RE = re.compile(
re.MULTILINE, re.MULTILINE,
) )
PROTOCOL_LEAK_EVENTS = '"events"' PROTOCOL_LEAK_EVENTS = '"events"'
#: v1.1 WP-A2: the two shapes the M11 closeout's identity run stored that the
#: two signs above cannot see — a line opening with a call to an event, and the
#: length hint echoed with the application's own wording. Copied, as above.
PROTOCOL_LEAK_CALL_RE = re.compile(
r"^[ \t]*(?:>[ \t]*)?(?:create_entity|set_entity_status|set_entity_attribute"
r"|set_entity_conditions|set_current_location|set_possession|clear_possession"
r"|add_fact|invalidate_fact|add_relationship|end_relationship"
r"|open_story_thread|resolve_story_thread|set_scene)[ \t]*\(",
re.IGNORECASE | re.MULTILINE,
)
PROTOCOL_LEAK_HINT_RE = re.compile(
r"\[Hard limit:[^\]]*(?:append the state block|turn must not exceed \d+ words)",
re.IGNORECASE,
)
def _leaks_protocol(text: str) -> bool: def _leaks_protocol(text: str) -> bool:
return bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text return (bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
or bool(PROTOCOL_LEAK_CALL_RE.search(text))
or bool(PROTOCOL_LEAK_HINT_RE.search(text)))
def _protocol_leaks(bundle: dict) -> dict: def _protocol_leaks(bundle: dict) -> dict:
+187
View File
@@ -0,0 +1,187 @@
"""v1.1: does a real v1.0.0 database open unchanged?
# the same database, opened by each tree, snapshotted read-only
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
--tree <v1.0.0 worktree>/backend --label v100 --out "$HOME/v11-evidence/compat"
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
--label v11 --exercise --out "$HOME/v11-evidence/compat"
Run from `backend/`. The source database is never opened. It is copied into
`--out` first, and the copy is what the application opens.
A read-only snapshot is taken through the API, the same way a reader sees the
campaign:
- the export bundle, which carries the whole tree, the head, the Save Points,
state, events, summaries, memories and knowledge, and has no timestamp of its
own;
- the narrative state and its events;
- the Save Points, the imported knowledge, the memories, the derived status and
the settings;
- the database schema and `PRAGMA user_version`, before and after the
application opened it.
Two snapshots of the same database from two trees are then compared. Identical
means v1.1 read it exactly as v1.0.0 did, and a matching schema and version mean
nothing migrated.
`--exercise` then uses the v1.1 copy: undo, redo, a Save Point restore, a
context dry run (knowledge retrieval), an export, and an import of that export.
It first points the copy's endpoint at a loopback port that refuses, so nothing
here reaches an inference server.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sqlite3
import sys
from pathlib import Path
def _schema(path: Path) -> dict:
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
version = connection.execute("PRAGMA user_version").fetchone()[0]
rows = connection.execute(
"SELECT type, name, sql FROM sqlite_master WHERE name NOT LIKE 'sqlite_%' "
"ORDER BY type, name").fetchall()
finally:
connection.close()
return {"user_version": version, "objects": [list(r) for r in rows]}
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--db", required=True)
parser.add_argument("--tree", default="", help="a backend/ directory to import the app from")
parser.add_argument("--label", required=True)
parser.add_argument("--exercise", action="store_true")
parser.add_argument("--out", required=True)
args = parser.parse_args()
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
copy = out / f"{args.label}.db"
if copy.exists():
print(f"{copy} exists; choose a new --label or --out")
return 2
shutil.copy2(args.db, copy)
schema_before = _schema(copy)
if args.tree:
sys.path.insert(0, str(Path(args.tree).resolve()))
os.environ["AIDND_DB_PATH"] = str(copy)
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, models
from app.database import SessionLocal, get_db
from app.main import app
print(f"app imported from {Path(sys.modules['app'].__file__).parent}")
limits.check_row_cap = lambda *a, **k: None
with SessionLocal() as db:
owner = db.query(models.Adventure.user_id).order_by(models.Adventure.id).first()
user_id = owner[0] if owner else db.query(models.User.id).first()[0]
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
def call(client, method, url, body=None, expect=200):
response = client.request(method, f"/api{url}", json=body)
if response.status_code != expect:
raise SystemExit(f"{method} {url}: HTTP {response.status_code} {response.text[:300]}")
return response.json() if response.content else None
report: dict = {"label": args.label, "schema_before": schema_before}
with TestClient(app) as client:
with SessionLocal() as db:
adventure_ids = [a for (a,) in db.query(models.Adventure.id)
.filter(models.Adventure.user_id == user_id)
.order_by(models.Adventure.id)]
snapshot = {"settings": call(client, "GET", "/settings"), "adventures": {}}
for adv in adventure_ids:
snapshot["adventures"][str(adv)] = {
"export": call(client, "GET", f"/adventures/{adv}/export"),
"state": call(client, "GET", f"/adventures/{adv}/state"),
"state_events": call(client, "GET", f"/adventures/{adv}/state/events"),
"checkpoints": call(client, "GET", f"/adventures/{adv}/checkpoints"),
"knowledge": call(client, "GET", f"/adventures/{adv}/knowledge"),
"memories": call(client, "GET", f"/adventures/{adv}/memories"),
"derived": call(client, "GET", f"/adventures/{adv}/derived"),
"newest_actions": call(client, "GET", f"/adventures/{adv}/actions?limit=5"),
}
report["snapshot"] = snapshot
if args.exercise and adventure_ids:
adv = adventure_ids[0]
ex: dict = {}
call(client, "PUT", "/settings", {"endpoint_url": "http://127.0.0.1:9/v1",
"embedding_model": ""})
before = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
ex["before"] = {k: before[k] for k in ("total", "can_undo", "can_redo")}
undone = call(client, "POST", f"/adventures/{adv}/undo")
ex["after_undo"] = {k: undone[k] for k in ("total", "can_undo", "can_redo")}
redone = call(client, "POST", f"/adventures/{adv}/redo")
ex["after_redo"] = {k: redone[k] for k in ("total", "can_undo", "can_redo")}
points = call(client, "GET", f"/adventures/{adv}/checkpoints")
if points:
point = points[0]
restored = call(client, "POST",
f"/adventures/{adv}/checkpoints/{point['id']}/restore")
page = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
ex["restore"] = {"save_point": point["name"], "total": page["total"],
"can_redo": page["can_redo"],
"response_keys": sorted(restored or {})}
context = client.get(f"/api/adventures/{adv}/context")
body = context.json()
ex["context"] = {
"status": context.status_code,
"knowledge_used": len(((body.get("knowledge") or {}).get("used")) or []),
"canon_section": any(s["label"] == "campaign_canon"
for s in body.get("sections") or []),
"state_section": any(s["label"] == "narrative_state"
for s in body.get("sections") or []),
"summary": body.get("summary"),
"tokens": body.get("tokens"),
"window": body.get("window"),
}
bundle = call(client, "GET", f"/adventures/{adv}/export")
imported = call(client, "POST", "/adventures/import", bundle, expect=201)
new_id = imported["id"]
reimport = call(client, "GET", f"/adventures/{new_id}/export")
ex["import"] = {
"new_id": new_id,
"actions_in_bundle": len(bundle.get("actions") or []),
"actions_after_import": len(reimport.get("actions") or []),
"head_same": (bundle.get("headBranch") is not None
and bundle.get("headDepth") == reimport.get("headDepth")),
"checkpoints": [len(bundle.get("checkpoints") or []),
len(reimport.get("checkpoints") or [])],
"memories": [len(bundle.get("memories") or []),
len(reimport.get("memories") or [])],
"narrative_state_same": bundle.get("narrativeState") == reimport.get("narrativeState"),
}
report["exercise"] = ex
app.dependency_overrides.clear()
report["schema_after"] = _schema(copy)
(out / f"{args.label}.json").write_text(json.dumps(report, indent=2, sort_keys=True, default=str))
same_schema = report["schema_before"] == report["schema_after"]
print(f"schema unchanged by opening: {same_schema} "
f"(user_version {report['schema_before']['user_version']} -> "
f"{report['schema_after']['user_version']})")
if "exercise" in report:
print(json.dumps(report["exercise"], indent=2, default=str)[:3000])
return 0
if __name__ == "__main__":
sys.exit(main())
+245
View File
@@ -0,0 +1,245 @@
"""v1.1 WP-A2: replay real stored narration through the v1.0.0 and current extractors.
The A2 extractor changes remove more text from a narrator's reply than v1.0.0
did. Removing story is worse than leaving protocol (`TECHNICAL-DESIGN.md`
§15.4), so every change is shown to a person rather than summarised. This tool
takes every real reply the evidence kept, runs it through both extractors, and
writes each turn whose prose differs with:
- the v1.0.0 prose and the current prose;
- every line removed, and the rule that explains it;
- any removal no rule explains, which fails the replay.
**Input.** A reply is read from the turn's stored `raw_output` where the
evidence database kept one: that is exactly what the narrator sent. A bundle
carries no raw reply, so a bundle's turns are replayed from their stored text,
which is v1.0.0's output already. For those the old prose is the input itself,
and the comparison is still exact.
**The v1.0.0 extractor** is read from the release tag with `git show`, not
copied, so this tool compares against what shipped. It shares `events` and
`render` with the current tree. A2 does not change `events.SPECS` or
`render.SECTION_HEADINGS`, and the report verifies that with `git diff`.
Evidence stays outside the repository. Replayed text is fiction from the
acceptance fixtures, but it is still somebody's run.
.venv/bin/python -m tools.v11_replay_extractor \\
--db "$HOME/m11-evidence/**/*.db" \\
--bundle "$HOME/m11-evidence/closeout-3652dc6/identity/turn-99/bundle.json" \\
--out "$HOME/v11-evidence/a2-replay"
"""
from __future__ import annotations
import argparse
import difflib
import glob
import hashlib
import importlib.util
import json
import sqlite3
import subprocess
import sys
import types
import zlib
from pathlib import Path
RELEASE = "v1.0.0"
EXTRACT_PATH = "backend/app/narrative/extract.py"
#: A removal larger than this share of the v1.0.0 prose is flagged for review
#: even when every line is explained, because a rule that eats most of a reply
#: is the shape a false positive takes.
LARGE_REMOVAL_SHARE = 0.25
def load_release_extractor(repo: Path) -> types.ModuleType:
"""`app.narrative.extract` as it was at the release tag."""
source = subprocess.run(
["git", "-C", str(repo), "show", f"{RELEASE}:{EXTRACT_PATH}"],
check=True, capture_output=True, text=True,
).stdout
import app.narrative # noqa: F401 the package the relative import needs
spec = importlib.util.spec_from_loader("app.narrative._extract_release", loader=None)
module = importlib.util.module_from_spec(spec)
module.__package__ = "app.narrative"
exec(compile(source, f"{RELEASE}:{EXTRACT_PATH}", "exec"), module.__dict__)
return module
def _unpack(blob):
if blob is None:
return None
try:
return json.loads(zlib.decompress(bytes(blob)).decode("utf-8"))
except (zlib.error, ValueError, UnicodeDecodeError):
return None
def turns_from_db(path: str):
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
rows = connection.execute(
"SELECT id, depth, text, context_snapshot FROM actions WHERE type = 'ai'"
).fetchall()
finally:
connection.close()
for action_id, depth, text, blob in rows:
snapshot = _unpack(blob) or {}
raw = snapshot.get("raw_output")
if isinstance(raw, str) and raw.strip():
yield {"source": path, "id": action_id, "depth": depth,
"input": raw, "input_kind": "raw_output"}
elif text:
yield {"source": path, "id": action_id, "depth": depth,
"input": text, "input_kind": "stored_text"}
def turns_from_bundle(path: str):
bundle = json.loads(Path(path).read_text())
for action in bundle.get("actions") or []:
if action.get("type") == "ai" and action.get("text"):
yield {"source": path, "id": action.get("id"), "depth": action.get("depth"),
"input": action["text"], "input_kind": "stored_text"}
def removed_lines(before: str, after: str) -> list[str]:
"""Lines present in `before` and gone from `after`, in order.
Compared with trailing whitespace ignored. The extractor strips the end of
every reply it cuts, so a story line that becomes the last line loses a
trailing space. A first version of this tool reported that space as a
rewritten line of story, which it is not.
"""
old = [line.rstrip() for line in before.split("\n")]
new = [line.rstrip() for line in after.split("\n")]
matcher = difflib.SequenceMatcher(a=old, b=new, autojunk=False)
gone: list[str] = []
for tag, a0, a1, _b0, _b1 in matcher.get_opcodes():
if tag in ("delete", "replace"):
gone.extend(old[a0:a1])
return gone
def replay(inputs, old, new) -> dict:
seen: set[str] = set()
unchanged = 0
changed: list[dict] = []
duplicates = 0
for turn in inputs:
digest = hashlib.sha256(turn["input"].encode()).hexdigest()
if digest in seen:
duplicates += 1
continue
seen.add(digest)
old_prose, _old_parsed, _old_raw = old.split(turn["input"])
new_prose, _new_parsed, _new_raw = new.split(turn["input"])
if old_prose == new_prose:
unchanged += 1
continue
lines = []
unexplained = 0
for line in removed_lines(old_prose, new_prose):
if not line.strip():
continue
rule = new.explain_removed_line(line)
if rule is None:
unexplained += 1
lines.append({"line": line, "rule": rule})
added = [line for line in removed_lines(new_prose, old_prose) if line.strip()]
share = 1 - len(new_prose) / max(1, len(old_prose))
flags = []
if unexplained:
flags.append("unexplained_removal")
if added:
flags.append("text_added_or_rewritten")
if share > LARGE_REMOVAL_SHARE:
flags.append("large_removal")
changed.append({
**{k: turn[k] for k in ("source", "id", "depth", "input_kind")},
"sha256": digest,
"old_prose": old_prose,
"new_prose": new_prose,
"removed": lines,
"added_or_rewritten": added,
"removed_chars": len(old_prose) - len(new_prose),
"removed_share": round(share, 4),
"flags": flags,
})
return {
"replayed": unchanged + len(changed),
"duplicates_skipped": duplicates,
"unchanged": unchanged,
"changed": len(changed),
"flagged": sum(1 for c in changed if c["flags"]),
"turns": changed,
}
def write_markdown(result: dict, path: Path) -> None:
out = [
"# A2 extractor replay",
"",
f"- replayed (unique replies): **{result['replayed']}**",
f"- duplicates skipped: {result['duplicates_skipped']}",
f"- unchanged: {result['unchanged']}",
f"- changed: **{result['changed']}**",
f"- flagged: **{result['flagged']}**",
"",
]
for index, turn in enumerate(result["turns"], 1):
out += [
f"## {index}. {Path(turn['source']).parent.name}/{Path(turn['source']).name}"
f" action {turn['id']} depth {turn['depth']} ({turn['input_kind']})",
"",
f"- removed chars: {turn['removed_chars']} ({turn['removed_share']:.1%})",
f"- flags: {', '.join(turn['flags']) or 'none'}",
"",
"Removed lines:",
"",
]
for item in turn["removed"]:
out.append(f"- `{item['rule'] or 'UNEXPLAINED'}` — {item['line']!r}")
out += ["", "<details><summary>v1.0.0 prose</summary>", "", "```text",
turn["old_prose"], "```", "</details>", "",
"<details><summary>current prose</summary>", "", "```text",
turn["new_prose"], "```", "</details>", ""]
path.write_text("\n".join(out))
def main(argv=None) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--db", action="append", default=[],
help="an evidence database, or a glob of them")
parser.add_argument("--bundle", action="append", default=[],
help="an exported bundle whose campaign has no database here")
parser.add_argument("--out", required=True)
args = parser.parse_args(argv)
repo = Path(__file__).resolve().parents[2]
from app.narrative import extract as current
old = load_release_extractor(repo)
dbs = sorted({p for pattern in args.db for p in glob.glob(pattern, recursive=True)})
def inputs():
for path in dbs:
yield from turns_from_db(path)
for path in args.bundle:
yield from turns_from_bundle(path)
result = replay(inputs(), old, current)
result["databases"] = dbs
result["bundles"] = args.bundle
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
(out / "replay.json").write_text(json.dumps(result, indent=2, ensure_ascii=False))
write_markdown(result, out / "replay.md")
print(json.dumps({k: result[k] for k in
("replayed", "duplicates_skipped", "unchanged", "changed", "flagged")}))
return 1 if result["flagged"] else 0
if __name__ == "__main__":
sys.exit(main())
+211
View File
@@ -0,0 +1,211 @@
"""v1.1 WP-A1: real turns, and what the server said it read.
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 AIDND_TEST_MODEL=<model> \\
.venv/bin/python -m tools.v11_window_accounting \\
--bundle "$HOME/m11-evidence/m04-final/bundle.json" --turns 4 \\
--out "$HOME/v11-evidence/a1-accounting/<label>"
Run from `backend/`. The evidence that matters is at the edge of the window, and
a new campaign takes dozens of turns to reach it. So this imports a long
campaign, by default the v1 evidence run's 207-action bundle, and every turn is
assembled against a full window from the first. A real narrator is then asked
for `--turns` turns, and each one is written out with:
configured budget, verified window, and where the window came from
the application's estimate of what it sent (its count, plus the text the
provider adds)
the server's own prompt-token count, from the usage it reported
the reply allocation and the safety reserve
the observed margin: window - reply allocation - the server's count
the accounting status: fits, exceeded, truncation_suspected or unknown
The first turn on a cold model cannot verify the window: `/api/ps` knows nothing
until the model is loaded. That is M11's behaviour, and the row says so rather
than hiding the turn.
The database lives in `--out`, not in `/tmp`, so the stored snapshots behind
every row can be read again. The endpoint is read from the environment and is
never written into a committed file.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
import time
from pathlib import Path
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
TURNS = [
"I look around carefully and take stock of where I am.",
"I ask the nearest person what has happened since I was last here.",
"I check what I am carrying.",
"I move on towards the place I meant to reach.",
"I wait and listen.",
"I say, \"Tell me the part you left out.\"",
]
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--bundle", default="",
help="a campaign bundle to import, so turns start at a full window")
parser.add_argument("--turns", type=int, default=4)
parser.add_argument("--budget", type=int, default=16384)
parser.add_argument("--max-output", type=int, default=500)
parser.add_argument("--timeout", type=int, default=900)
parser.add_argument("--out", required=True)
parser.add_argument("--unload-first", action="store_true",
help="ask the configured server to unload the model before turn 1, "
"so the first turn starts cold (the A1 corrective test)")
args = parser.parse_args()
if not (ENDPOINT and MODEL):
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
return 2
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
db_path = out / "accounting.db"
if db_path.exists():
print(f"{db_path} exists; choose a new --out")
return 2
os.environ["AIDND_DB_PATH"] = str(db_path)
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, limits, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
limits.check_row_cap = lambda *a, **k: None
Base.metadata.create_all(bind=engine)
with SessionLocal() as db:
user = models.User(is_guest=False, email="accounting@example.com")
db.add(user)
db.flush()
db.add(models.Settings(
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
context_token_budget=args.budget, max_output_tokens=args.max_output,
model_timeout_seconds=args.timeout,
))
db.commit()
user_id = user.id
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
client = TestClient(app)
if args.bundle:
bundle = json.loads(Path(args.bundle).read_text())
# The evidence campaign had its memory bank and auto-summarise on. Here
# they would only add post-turn model calls between the measured turns,
# on the same host, and fail noisily as the in-process client closes.
# This tool measures the turn's own prompt, which neither changes.
bundle["memoryBankEnabled"] = False
bundle["autoSummarize"] = False
imported = client.post("/api/adventures/import", json=bundle)
imported.raise_for_status()
adv = imported.json()["id"]
else:
created = client.post("/api/adventures", json={
"title": "Window accounting", "opening": "A quiet road at dusk."})
created.raise_for_status()
adv = created.json()["id"]
if args.unload_first:
# The configured endpoint only, under the same policy and TLS trust as a turn.
import asyncio
import httpx
from app import contextwindow, endpoints, tlstrust
reason = endpoints.rejection_reason(ENDPOINT)
if reason:
print(f"endpoint refused: {reason}")
return 2
base = contextwindow.native_base(ENDPOINT)
with httpx.Client(verify=tlstrust.ssl_context(), timeout=120) as http:
unloaded = http.post(f"{base}/api/generate", json={"model": MODEL, "keep_alive": 0})
resident = [m.get("name") for m in http.get(f"{base}/api/ps").json().get("models", [])]
contextwindow.cache_clear()
print(f"unload: HTTP {unloaded.status_code} {unloaded.text[:120]} | resident now: {resident}")
rows: list[dict] = []
timeline = (out / "turns.jsonl").open("a")
print(f"model {MODEL}, budget {args.budget}, reply {args.max_output}")
print(f"{'#':>2} {'status':22} {'window':>13} {'estimate':>8} {'server':>7} "
f"{'reserve':>7} {'margin':>7} {'sec':>5}")
for index in range(args.turns):
text = TURNS[index % len(TURNS)]
started = time.monotonic()
response = client.post(f"/api/adventures/{adv}/actions",
json={"type": "do", "text": text})
seconds = round(time.monotonic() - started, 1)
error = None
if response.status_code != 200 or '"type": "error"' in response.text:
error = response.text[-400:]
with SessionLocal() as db:
action = (
db.query(models.Action)
.filter(models.Action.adventure_id == adv, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first()
)
snapshot = (action.context_snapshot or {}) if action else {}
tokens = snapshot.get("tokens") or {}
window = snapshot.get("window") or {}
accounting = snapshot.get("accounting") or {}
row = {
"turn": index + 1,
"action_id": action.id if action else None,
"seconds": seconds,
"error": error,
"configured_budget": tokens.get("configured_budget"),
"effective_budget": tokens.get("budget"),
"window_verified": window.get("verified"),
"window_tokens": window.get("tokens"),
"window_source": window.get("source"),
"preflight_attempted": (window.get("preflight") or {}).get("attempted"),
"preflight_loaded": (window.get("preflight") or {}).get("loaded"),
"preflight_verified_before": (window.get("preflight") or {}).get("verified_before"),
"preflight_verified_after": (window.get("preflight") or {}).get("verified_after"),
"preflight_detail": (window.get("preflight") or {}).get("detail"),
"app_prompt_tokens": tokens.get("total"),
"transport_tokens": tokens.get("transport"),
"app_estimate": tokens.get("estimate"),
"output_reserve": tokens.get("output_reserve"),
"safety_reserve": tokens.get("safety_reserve"),
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
"difference": accounting.get("difference"),
"observed_margin": accounting.get("observed_margin"),
"status": accounting.get("status"),
"history_included": (snapshot.get("history") or {}).get("included"),
"history_total": (snapshot.get("history") or {}).get("total"),
}
rows.append(row)
timeline.write(json.dumps(row) + "\n")
timeline.flush()
print(f"{row['turn']:>2} {str(row['status'] if not error else 'ERROR'):22} "
f"{str(row['window_tokens'])+('v' if row['window_verified'] else '?'):>13} "
f"{str(row['app_estimate']):>8} {str(row['server_prompt_tokens']):>7} "
f"{str(row['safety_reserve']):>7} {str(row['observed_margin']):>7} {seconds:>5}")
if error:
print(f" error: {error[:200]}")
timeline.close()
(out / "summary.json").write_text(json.dumps({
"model": MODEL, "budget": args.budget, "max_output_tokens": args.max_output,
"bundle": args.bundle, "rows": rows,
}, indent=2))
app.dependency_overrides.clear()
return 0 if all(r["error"] is None for r in rows) else 1
if __name__ == "__main__":
sys.exit(main())
@@ -47,6 +47,63 @@ function Section({ title, count, children, open = false, testId }) {
) )
} }
/* v1.1 WP-A1: what the server said it read, set against what was sent.
*
* Only a turn that was actually sent has this — the next-turn view has not
* been sent yet. The two conditions that mean something went wrong are shown
* as alerts, because the failure they describe is otherwise silent: Ollama
* answers 200 whether or not it cut the front of the prompt off. */
const ACCOUNTING = {
fits: {
title: 'The server read the whole prompt',
body: 'Its count stayed inside the room kept for the reply and the safety margin.',
},
exceeded: {
title: 'The prompt was larger than the server allowed for',
body: 'The server counted more tokens than the safety margin covers, so the '
+ 'reply may have been cut short. The turn is kept.',
alert: true,
},
truncation_suspected: {
title: 'The server may have cut the start of the prompt',
body: 'It read far fewer tokens than were sent, which is what happens when a '
+ 'prompt is larger than the window the model was loaded with. The '
+ 'narrator’s rules and the campaign canon are at the start. The turn is kept.',
alert: true,
},
unknown: {
title: 'The server did not say how much it read',
body: 'Nothing here can confirm whether the whole prompt was used.',
},
}
function AccountingReport({ accounting }) {
if (!accounting) return null
const copy = ACCOUNTING[accounting.status] || ACCOUNTING.unknown
return (
<div
className={copy.alert ? 'notice error' : 'ctx-accounting'}
role={copy.alert ? 'alert' : undefined}
data-testid="ctx-accounting"
data-status={accounting.status}
>
<strong>{copy.title}</strong>
<p>{copy.body}</p>
{accounting.server_prompt_tokens != null && (
<p className="notice-detail">
Sent {accounting.estimate?.toLocaleString()} by this app’s count; the
server read {accounting.server_prompt_tokens.toLocaleString()}.
{accounting.observed_margin != null
&& ` ${accounting.observed_margin.toLocaleString()} tokens were left beside the reply.`}
</p>
)}
{accounting.window_verified === false && (
<p className="notice-detail">The model’s window was not verified for this turn.</p>
)}
</div>
)
}
function TokenBar({ sections, total }) { function TokenBar({ sections, total }) {
if (!sections.length || total <= 0) return null if (!sections.length || total <= 0) return null
return ( return (
@@ -122,6 +179,12 @@ export function ContextPanel({
<span>{tokens.output_reserve.toLocaleString()}</span> <span>{tokens.output_reserve.toLocaleString()}</span>
</div> </div>
)} )}
{tokens.safety_reserve > 0 && (
<div className="ctx-token-line dim">
<span>Kept free as a safety margin</span>
<span>{tokens.safety_reserve.toLocaleString()}</span>
</div>
)}
<div className="ctx-token-line dim"> <div className="ctx-token-line dim">
<span>What the model can hold</span> <span>What the model can hold</span>
<span>{tokens.budget.toLocaleString()}</span> <span>{tokens.budget.toLocaleString()}</span>
@@ -135,6 +198,8 @@ export function ContextPanel({
)} )}
</div> </div>
<AccountingReport accounting={report.accounting} />
{failing.length > 0 && ( {failing.length > 0 && (
<div className="notice error" role="alert" data-testid="ctx-derived-failing"> <div className="notice error" role="alert" data-testid="ctx-derived-failing">
<strong>Background work is failing</strong> <strong>Background work is failing</strong>
@@ -220,6 +220,50 @@ describe('context inspector (§23, §55)', () => {
expect(tokens).toHaveTextContent('8,000') expect(tokens).toHaveTextContent('8,000')
}) })
it('shows the safety margin kept free beside the reply (v1.1 A1)', async () => {
api.getAdventureContext.mockResolvedValue({
...REPORT, tokens: { ...REPORT.tokens, safety_reserve: 410 },
})
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
expect(screen.getByTestId('ctx-tokens')).toHaveTextContent(/safety margin\s*410/)
})
it('says nothing about accounting for a turn that has not been sent', async () => {
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
expect(screen.queryByTestId('ctx-accounting')).toBeNull()
})
it('alerts when the server may have cut the start of a sent prompt (v1.1 A1)', async () => {
vi.spyOn(api, 'getActionContext').mockResolvedValue({
...REPORT,
accounting: {
status: 'truncation_suspected', estimate: 6316, server_prompt_tokens: 2050,
observed_margin: 1546, window_verified: true,
},
})
await renderWith(
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
const el = screen.getByTestId('ctx-accounting')
expect(el).toHaveAttribute('data-status', 'truncation_suspected')
expect(el).toHaveAttribute('role', 'alert')
expect(el).toHaveTextContent('6,316')
expect(el).toHaveTextContent('2,050')
expect(el).toHaveTextContent(/turn is kept/)
})
it('reports an unknown count plainly, without claiming the prompt fit', async () => {
vi.spyOn(api, 'getActionContext').mockResolvedValue({
...REPORT, accounting: { status: 'unknown', server_prompt_tokens: null },
})
await renderWith(
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
const el = screen.getByTestId('ctx-accounting')
expect(el).toHaveAttribute('data-status', 'unknown')
expect(el).not.toHaveAttribute('role')
expect(el).toHaveTextContent(/did not say/)
expect(el.textContent).not.toMatch(/read the whole prompt/)
})
it('shows a retrieved passage with its source, class, heading and score', async () => { it('shows a retrieved passage with its source, class, heading and score', async () => {
await renderWith(<ContextPanel advId="1" refreshKey="x" />) await renderWith(<ContextPanel advId="1" refreshKey="x" />)
const row = document.querySelector('[data-chunk-id="11"]') const row = document.querySelector('[data-chunk-id="11"]')
+10
View File
@@ -32,6 +32,16 @@
font-variant-numeric: tabular-nums; font-variant-numeric: tabular-nums;
} }
.ctx-warn { margin: 8px 0 0; color: var(--danger); font-size: 0.76rem; } .ctx-warn { margin: 8px 0 0; color: var(--danger); font-size: 0.76rem; }
/* v1.1 A1: what the server said it read. The two failure states render as a
`.notice.error` alert instead; this is the quiet form for `fits` and
`unknown`. */
.ctx-accounting {
padding: 9px 13px;
border: 1px solid var(--border);
border-radius: 7px;
color: var(--text-dim);
}
.ctx-accounting p { margin: 4px 0 0; }
.token-bar { .token-bar {
display: flex; display: flex;
@@ -69,6 +69,18 @@ out of tokens partway through it. All of that is removed before the prose is
stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4 stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4
has the rules. has the rules.
**Implementation note (v1.1 WP-A2).** A small model also copies the protocol's
*instructions*: the vocabulary written as calls, the length hint, and the scene
line. The fix is on both sides:
- **Prompt:** the vocabulary is shown in the wire format, and the fixed example
uses genre-neutral placeholders.
- **Extractor:** it recognises those echoes only by strings and names the
application owns.
No event type, field, validation rule or proposal record changed.
`TECHNICAL-DESIGN.md` §15.4 lists the four rules.
### Where it lives ### Where it lives
- `adventures.narrative_state` — the current authoritative document. This is - `adventures.narrative_state` — the current authoritative document. This is
+3 -1
View File
@@ -2,7 +2,9 @@
**This file is the index. Start here.** **This file is the index. Start here.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 planning has begun.** **Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: WP-A1
and WP-A2 are implemented and staged for owner review**
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`).
Phase 0 complete; AI-DnD forked as the production base; **milestones M1 Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the (2026-09-14), the v1 release gate passed on the release-candidate tree, and the
+116
View File
@@ -1230,6 +1230,75 @@ whether there is a number to cap to at all, and that is what the builder and the
declaration everywhere it appears, and the connection test says plainly that declaration everywhere it appears, and the connection test says plainly that
nothing has checked it against the server. nothing has checked it against the server.
**As implemented (v1.1 WP-A1): a safety reserve, and the server's own count.**
A ceiling in the application's tokens is not a ceiling in the narrator's. The
builder counts with `cl100k_base`, and the v1 evidence left the largest prompts
23-42 real tokens from the edge of a 16,384 window. Past the edge Ollama does not
refuse. Measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt returned
200 with `prompt_tokens` 2,050.
- **The reserve.** `contextwindow.safety_reserve(budget)` is
`max(256, ceil(5% of the effective budget))`, computed with integer rounding
up: 256 at 4,096, 410 at 8,192, 820 at 16,384. It is taken before any history
is chosen. The effective budget is the verified or declared window when there
is one, and the configured budget otherwise. It is fixed and documented, is not
a setting, and is not calibrated per model.
- **What replaced the 64-token margin.** M6's `OUTPUT_SAFETY_MARGIN` absorbed two
unrelated things.
- The application's own text added after pricing: separators between
sections, and the chat hint the provider appends to every request. This is
now priced exactly as `transport`.
- Tokenizer drift. This is now the reserve.
The reply allocation is exactly `max_output_tokens`. Protected context is
`sections + transport + reply + reserve`, and `ContextOverflow` is raised
before the model call when that does not fit.
- **The server's count.** Streaming requests set `stream_options.include_usage`.
Without it Ollama sends no usage, and none of the 514 AI turns in the v1
evidence has one. After the reply, `contextwindow.classify_usage` compares the
server's `prompt_tokens` with `tokens.estimate`: the assembled text plus what
the provider adds.
- **Accounting states.** The turn's snapshot records `accounting`, whose status
is one of the following, checked in this order:
| Status | Meaning |
| --- | --- |
| `unknown` | No positive integer count was reported. It is never read as `fits`. |
| `truncation_suspected` | The server read fewer tokens than the estimate by more than the reserve. |
| `exceeded` | The server's count plus the reply allocation is over the budget. |
| `fits` | Otherwise. |
The record also carries the server's count, the difference, the reserve and
the observed margin (`budget - reply - server count`).
- **Surfacing.** The record is returned on the turn's `done` event, logged as a
warning when it is `exceeded` or `truncation_suspected`, and shown in the
context inspector, where those two statuses are an alert.
- **The turn is kept.** A discrepancy found after the reply is recorded, never
enforced. The narration has already streamed to the reader, and the accepted
turn is not discarded.
- **Accounting belongs to one attempt.** It sits in `attempts.ATTEMPT_KEYS`
beside `usage`. When a retry or a take selection moves the shared prompt,
each take keeps the accounting for its own call.
- **A cold model is loaded before its turn is built (v1.1 A1 corrective).** A
model that is not resident cannot report its window. The A1 evidence caught
exactly that: 13,875 tokens were sent to a server that read 2,050.
`contextwindow.ensure_window` works like this:
- it probes;
- if the window is unverified and the server answered, it makes one bounded
`POST /api/generate` naming only the model, with no prompt. Ollama loads
the model and generates nothing ("done_reason": "load");
- it probes again, bypassing the cache;
- the turn is built to whatever that second probe says.
A failed load, or a window still unverified afterwards, changes nothing: the
configured budget stands and the accounting still catches a cut. The request
goes to the configured endpoint only, under the same policy and TLS path. It
writes nothing, and it is recorded as `window.preflight` in the turn's
snapshot. The context dry run never loads a model.
No schema change: the accounting lives in the snapshot JSON, and a turn from
v1.0.0 simply has none. The bundle format is unchanged, for the same reason.
### 15.3 The history window moves in blocks (post-M11) ### 15.3 The history window moves in blocks (post-M11)
§15.2 makes the window a ceiling. This is about what happens at that ceiling. §15.2 makes the window a ceiling. This is about what happens at that ceiling.
@@ -1300,6 +1369,53 @@ headings. A lone heading followed by prose stays, and so do JSON a character
typed and a fact restated inside a sentence. That last case is how a narrator can typed and a fact restated inside a sentence. That last case is how a narrator can
still carry authoritative state into its prose (M11 report §G.4 and §P). still carry authoritative state into its prose (M11 report §G.4 and §P).
**As implemented (v1.1 WP-A2): the source and the sink together.** The M11
closeout's identity run stored four shapes the extractor left. All of them were
application text. v1.1 changes both the prompt that taught them and the
extractor that missed them. Every new removal is anchored to something the
application owns, never to what prose looks like.
- **The prompt.**
- `events.vocabulary_for_prompt` shows each event as the object the model
must send (`{"type": "set_possession", "item": "<key>", "owner": "<key>"}`),
not as `set_possession(item, owner)`. The call notation was never the wire
format, and the narrator copied it.
- `EMIT_RULE`'s example uses the placeholders `character-1`, `item-1` and
`location-1`, not the fantasy fixture's `mara`, `silver-key`, `old-abbey` and
`aldric`. The narrator had proposed `silver-key` in an office meeting.
- The length hint's opening and closing words are named constants shared by
the builder and the extractor.
- **The extractor**, rules R1-R4 (`RULE_*` in `narrative/extract.py`):
- **R1:** a whole line that begins with a call to an event in `events.SPECS`,
optionally `>`-quoted. Not inside a fenced code block, not mid-sentence, and
not for a call-shaped name the vocabulary lacks.
- **R2:** a trailing bracket that opens `Hard limit:` and carries the hint's
own wording ("append the state block", or "turn must not exceed *N* words").
- **R3:** the renderer's scene line left as the reply's last line. It is
removed when it ends in the renderer's `(at <location>)`, or when protocol
was already cut from the same reply.
- **R4:** a ```` ```json ```` or bare ```` ``` ```` opener left as the last
line with nothing after it, counted as an opener rather than a closer.
- **Proven on real narration.** Every stored real reply in the v1 evidence was
replayed through the v1.0.0 and v1.1 extractors (`tools/v11_replay_extractor.py`).
Every changed line is attributed to one of the four rules, and a person reviewed
every change. The results are in the WP-A1/A2 report.
- **Deliberately still left:** a fact restated inside a sentence; model-invented
headings; and a bracket that starts `Hard limit:` but carries none of the
application's wording.
- **R5, the echoed instruction tail (v1.1 A2 corrective).** A v1.1 identity turn
ended in a reworded continue hint, "[… Continue the story here, directly.
Output only story text.]". Because nothing recognised it, nothing above it was
ever trailing, and the reminder, a reworded length hint and a scene line all
stayed. R5 makes two changes:
- **The continue hint is recognised by its own sentence.** A trailing bracket
containing "Output only story text" is an echoed instruction.
`CONTINUE_HINT_PHRASE` is pinned by a test to `CHAT_CONTINUE_HINT`.
- **A bracket opening with the length hint's own `Hard limit:` is removed only
directly above an echoed instruction already cut from the same reply's end.**
An in-world "[Hard limit: forty days]" stays when it is the last line, and
when a state block follows it. Any other bracket above an echo stays.
## 16. Database Direction ## 16. Database Direction
SQLite remains the selected v1 authoritative store. SQLite remains the selected v1 authoritative store.
+7 -3
View File
@@ -1,8 +1,12 @@
# Adventure Storyteller — v1.1 Plan # Adventure Storyteller — v1.1 Plan
**Status:** PLANNING. Written 2026-09-14 on `v1.1-development`, from the signed **Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
v1.0.0 release commit `432f041`. No work package has started, and no v1.1 signed v1.0.0 release commit `432f041`.
version or tag exists.
**WP-A1 and WP-A2 are implemented and staged for owner review** (2026-09-14). They
are reported together, and kept separate, in
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`. No other work package has started. No
v1.1 version or tag exists.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1 This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1 history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
+28 -2
View File
@@ -1,8 +1,34 @@
# Planning Package Version # Planning Package Version
- **Package:** Adventure Storyteller Planning Package v4.1 - **Package:** Adventure Storyteller Planning Package v4.2
- **Revision date:** 2026-09-14 - **Revision date:** 2026-09-14
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 planning has begun** on `v1.1-development`: `V1.1-PLAN.md`. No v1.1 work package has started, and no v1.1 version or tag exists. - **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): **WP-A1 and WP-A2 are implemented and staged for owner review.** No other work package has started, and no v1.1 version or tag exists.
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
Two v1.1 work packages, implemented in sequence. No requirement or acceptance
test changed, and no schema or bundle format changed. Evidence is in
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `TECHNICAL-DESIGN.md` §15.2 | **As implemented (v1.1 WP-A1)**, covering: the safety reserve, `max(256, ceil(5%))`; what replaced M6's 64-token margin; the server's own count; the four accounting states; and keeping the turn. | as-implemented record |
| `TECHNICAL-DESIGN.md` §15.4 | **As implemented (v1.1 WP-A2)**: the vocabulary shown in the wire format, genre-neutral placeholders, and extractor rules R1-R4, each anchored to application-owned text. | as-implemented record |
| `DECISIONS/013-authoritative-narrative-state-document.md` | An implementation note for v1.1. No event type, field, validation rule or proposal record changed. | as-implemented note |
| `V1.1-PLAN.md` | Status: A1 and A2 implemented and staged. | status |
| `planning/README.md` | Current state. | index |
| `reports/v1.1/V1.1-WP-A1-A2-REPORT.md` | **New.** The combined review package, with A1 and A2 kept separate. | work-package report |
| `README.md`, `DEVELOPMENT.md` | The reserve and the accounting. The stale backend test count and the Screenshots paragraph are corrected. | developer docs |
**Corrective work before commit (owner review, 2026-09-14).** Report §R.
| Document | Change |
| --- | --- |
| `TECHNICAL-DESIGN.md` §15.2 | A cold model is loaded once before its turn is built (`contextwindow.ensure_window`). |
| `TECHNICAL-DESIGN.md` §15.4 | R5: the echoed continue hint is recognised by its own sentence, and the application-opened tail above it is removed. |
| `DEVELOPMENT.md`, `README.md` | The cold-model load, in operator terms. |
**Requirement changes: zero.**
## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14) ## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14)
File diff suppressed because it is too large Load Diff