v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
ac465ed867
commit
d63804f22e
@@ -46,7 +46,14 @@ from .narrative import model as narrative_model
|
||||
# token accounting. Each attempt is its own API call, and a retry is the call
|
||||
# most likely to read the prompt back out of cache. Everything else in a snapshot
|
||||
# is the prompt, which is assembled once per turn.
|
||||
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage")
|
||||
#
|
||||
# v1.1 WP-A1: `accounting` is one attempt's too. It compares the server's count
|
||||
# for *that* call with the turn's estimate. Left out of this tuple, it was
|
||||
# treated as part of the shared prompt, so moving the live flag handed the
|
||||
# superseded attempt's accounting to the new live one and threw the new one's
|
||||
# away. Found by the A2 long run: two retries and one take selection left three
|
||||
# attempts reporting no accounting, or another attempt's.
|
||||
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage", "accounting")
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ reading
|
||||
|
||||
@@ -25,6 +25,7 @@ from sqlalchemy.orm import object_session
|
||||
|
||||
from .. import contextwindow, derived, models, narrative, summaries, worldstate
|
||||
from ..knowledge import inject as knowledge_inject
|
||||
from ..providers.openai_compatible import CHAT_CONTINUE_HINT
|
||||
from ..knowledge import records as knowledge_records
|
||||
from . import encoding, history
|
||||
|
||||
@@ -106,11 +107,22 @@ BAND_FLOOR_SHARE = 0.5
|
||||
# Built from the table vendored in `encoding.py`, not fetched: the upstream
|
||||
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
|
||||
# is called on every turn.
|
||||
# M6: added to the configured reply budget when reserving output space. It
|
||||
# absorbs the section separators added after budgeting and the drift between
|
||||
# this tokenizer and the serving model's. Fixed rather than proportional: what
|
||||
# it covers does not grow with the size of the budget.
|
||||
OUTPUT_SAFETY_MARGIN = 64
|
||||
#
|
||||
# v1.1 WP-A1: `OUTPUT_SAFETY_MARGIN = 64` was here. M6 added it to the reply
|
||||
# budget to absorb two unrelated things, and v1.1 separates them:
|
||||
#
|
||||
# * **Text the application adds after pricing.** The separators between
|
||||
# sections, and `CHAT_CONTINUE_HINT`, which the provider appends to every chat
|
||||
# request and nothing counted. That is not drift, it is our own text, so it is
|
||||
# now priced exactly (`transport` below).
|
||||
# * **The drift between this tokenizer and the narrator's.** That is what the
|
||||
# 64 tokens were really for, and the v1 evidence showed it was too small. It
|
||||
# is now `contextwindow.safety_reserve`, sized to the window.
|
||||
#
|
||||
#: Story sections that can be joined by `SEPARATOR` after pricing: history,
|
||||
#: author's note, recent history, summary, lore, memories, state, front memory,
|
||||
#: length hint, refusals, reminder. Knowledge and history rows price their own.
|
||||
STORY_SECTION_SLOTS = 11
|
||||
|
||||
|
||||
class ContextOverflow(RuntimeError):
|
||||
@@ -180,19 +192,19 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
words = min(words, band_ceiling)
|
||||
floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
|
||||
tail = (
|
||||
" Finish the narration and append the state block well inside the limit."
|
||||
" " + narrative.extract.LENGTH_HINT_TAIL
|
||||
)
|
||||
if floor < MIN_LENGTH_FLOOR_WORDS:
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words. Write only as "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
|
||||
f"much as the moment needs — a typical turn is much shorter.{tail}]"
|
||||
)
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words, and it should not "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
|
||||
f"stop short of about {floor}. Prefer the lower end of that range unless "
|
||||
f"the scene genuinely needs more.{tail}]"
|
||||
)
|
||||
tail = " Finish the narration and append the state block well inside the limit."
|
||||
tail = " " + narrative.extract.LENGTH_HINT_TAIL
|
||||
|
||||
# State the number as a ceiling, never as a budget. In measurements, the
|
||||
# wording "keep this turn under about N words" read to the model as a target
|
||||
@@ -204,7 +216,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
|
||||
if floor < MIN_LENGTH_FLOOR_WORDS:
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words. Write only as "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
|
||||
f"much as the moment needs — a typical turn is much shorter.{tail}]"
|
||||
)
|
||||
# Both numbers are bounds, and the wording is deliberately asymmetric. The
|
||||
@@ -216,7 +228,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
# so a terse model reading the same clause stops at the floor rather than at
|
||||
# forty words.
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words, and it should not "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
|
||||
f"stop short of about {floor}. Prefer the lower end of that range unless "
|
||||
f"the scene genuinely needs more.{tail}]"
|
||||
)
|
||||
@@ -625,12 +637,19 @@ def build_context(
|
||||
# truncated turn on a model whose window is the budget
|
||||
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
|
||||
#
|
||||
# The margin covers what is added after this arithmetic — the separators
|
||||
# between sections, and the difference between our tokenizer's count and the
|
||||
# serving model's. It is small and fixed rather than proportional, because
|
||||
# what it absorbs does not scale with the budget.
|
||||
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
|
||||
protected = reserved + output_reserve
|
||||
# v1.1 WP-A1: the reply allocation is exactly the reply cap. The text this
|
||||
# application adds after pricing — separators, and the chat hint the
|
||||
# provider appends — is counted as `transport`. What neither can know, the
|
||||
# narrator's tokenizer disagreeing with `cl100k_base`, is the safety reserve,
|
||||
# which is sized to the window and taken before any history is chosen.
|
||||
output_reserve = max(0, settings.max_output_tokens)
|
||||
separator_tokens = count_tokens(SEPARATOR)
|
||||
transport = (
|
||||
separator_tokens * (len(system_sections) + STORY_SECTION_SLOTS)
|
||||
+ count_tokens(CHAT_CONTINUE_HINT)
|
||||
)
|
||||
safety = contextwindow.safety_reserve(budget)
|
||||
protected = reserved + transport + output_reserve + safety
|
||||
if protected >= budget:
|
||||
# Failing here is the point. The alternative — carrying on with a token
|
||||
# or two of history — builds a prompt that is known to overflow, and
|
||||
@@ -638,8 +657,9 @@ def build_context(
|
||||
# gracefully if protected context alone is too large."
|
||||
raise ContextOverflow(
|
||||
f"The protected context needs {protected} tokens "
|
||||
f"({reserved} of prompt plus {output_reserve} reserved for the "
|
||||
f"reply) but the context budget is {budget}. "
|
||||
f"({reserved} of prompt, {transport} of formatting, {output_reserve} "
|
||||
f"reserved for the reply and a {safety}-token safety margin) but the "
|
||||
f"context budget is {budget}. "
|
||||
+ (
|
||||
"That budget is what this server was found to accept, so raising "
|
||||
"the setting alone will not help — load the model with a larger "
|
||||
@@ -827,6 +847,7 @@ def build_context(
|
||||
story_text = SEPARATOR.join(s.text for s in story_sections)
|
||||
|
||||
all_sections = [s for s in system_sections if s.text] + story_sections
|
||||
total_tokens = count_tokens(system_text) + count_tokens(story_text)
|
||||
report = {
|
||||
"sections": [
|
||||
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
|
||||
@@ -837,13 +858,23 @@ def build_context(
|
||||
# what the history was actually allowed to spend after everything
|
||||
# protected was subtracted.
|
||||
"tokens": {
|
||||
"total": count_tokens(system_text) + count_tokens(story_text),
|
||||
"total": total_tokens,
|
||||
"budget": budget,
|
||||
"configured_budget": settings.context_token_budget,
|
||||
"output_reserve": output_reserve,
|
||||
"protected": reserved,
|
||||
"available_for_history": available,
|
||||
"history_spent": spent,
|
||||
# v1.1 WP-A1. `transport` is the formatting priced in above;
|
||||
# `estimate` is what this application believes it actually sent,
|
||||
# the assembled text plus what the provider adds to it, and is what
|
||||
# the server's own count is compared against after the reply.
|
||||
"transport": transport,
|
||||
"safety_reserve": safety,
|
||||
"estimate": total_tokens + (
|
||||
count_tokens(CHAT_CONTINUE_HINT) if settings.api_mode != "completion"
|
||||
else separator_tokens
|
||||
),
|
||||
},
|
||||
# M11: what the server was found to accept, and how. `verified` false
|
||||
# means nobody could check — the prompt was built to the configured
|
||||
|
||||
@@ -86,6 +86,7 @@ one.
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import math
|
||||
import re
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
@@ -132,6 +133,11 @@ class Window:
|
||||
model_max: int | None = None
|
||||
#: Why the window is unknown, or how it was found. Shown to the user.
|
||||
detail: str = ""
|
||||
#: v1.1: the server answered a discovery request at all, whatever it said.
|
||||
#: A server that answered but could not report a window may simply not have
|
||||
#: the model loaded yet, which `ensure_window` can fix; one that did not
|
||||
#: answer cannot be helped by asking it to load anything.
|
||||
reachable: bool = False
|
||||
|
||||
@property
|
||||
def verified(self) -> bool:
|
||||
@@ -180,6 +186,116 @@ def effective_budget(configured: int, window: Window | int | None) -> int:
|
||||
return min(configured, tokens)
|
||||
|
||||
|
||||
#: v1.1 WP-A1: the tokens kept free below the effective window, beyond the reply.
|
||||
#:
|
||||
#: The builder counts with `cl100k_base`; the narrator counts with its own
|
||||
#: tokenizer. The v1 evidence put the largest prompts 23-42 real tokens from the
|
||||
#: edge of a 16,384 window, and Ollama does not refuse a prompt past the edge —
|
||||
#: measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt came back 200
|
||||
#: with `prompt_tokens` 2,050. So the reserve is deliberate and sized to the
|
||||
#: window: the larger of a floor and a share, **rounded up to a whole token**.
|
||||
#:
|
||||
#: 4,096 -> 256 8,192 -> 410 16,384 -> 820
|
||||
#:
|
||||
#: A fixed, documented tolerance, owner-chosen for v1.1. It is not a setting and
|
||||
#: it is not calibrated per model.
|
||||
SAFETY_RESERVE_FLOOR = 256
|
||||
SAFETY_RESERVE_PERCENT = 5
|
||||
|
||||
|
||||
def safety_reserve(effective_window: int) -> int:
|
||||
"""`max(256, ceil(5% of the effective window))`, in tokens.
|
||||
|
||||
The effective window is the budget the prompt is actually built to — the
|
||||
verified or declared window when there is one, the configured budget
|
||||
otherwise — so a 16,384 setting against a 4,096 server reserves 256, not 820.
|
||||
Integer arithmetic, so the rounding is exact rather than a float's.
|
||||
"""
|
||||
share = math.ceil(max(0, effective_window) * SAFETY_RESERVE_PERCENT / 100)
|
||||
return max(SAFETY_RESERVE_FLOOR, share)
|
||||
|
||||
|
||||
#: v1.1 WP-A1: what the server's own count says about a turn that was sent.
|
||||
FITS = "fits"
|
||||
EXCEEDED = "exceeded"
|
||||
TRUNCATION_SUSPECTED = "truncation_suspected"
|
||||
#: `UNKNOWN` above: the server reported no usable count.
|
||||
|
||||
|
||||
def classify_usage(usage: dict | None, *, estimate: int, budget: int,
|
||||
max_output_tokens: int, window_verified: bool) -> dict:
|
||||
"""Sets the server's reported prompt count against what the application sent.
|
||||
|
||||
The order of the checks is the order of what they prove:
|
||||
|
||||
``unknown``
|
||||
No positive integer `prompt_tokens`. Nothing can be said, and nothing
|
||||
is claimed: an absent count is never read as a prompt that fitted.
|
||||
``truncation_suspected``
|
||||
The server read fewer tokens than were sent by more than the safety
|
||||
reserve. A tokenizer thriftier than `cl100k_base` may honestly count a
|
||||
little less; a shortfall larger than the tolerance the application keeps
|
||||
for drift is the signature of a server that cut the prompt — the real
|
||||
shape was 6,316 sent and 2,050 read.
|
||||
``exceeded``
|
||||
The server's count plus the reply allocation is more than the window
|
||||
the prompt was built for. The drift was larger than the whole reserve,
|
||||
so the reply may have been cut short.
|
||||
``fits``
|
||||
Otherwise.
|
||||
|
||||
`observed_margin` is what was left beside the reply by the server's count:
|
||||
`budget - max_output_tokens - server_prompt_tokens`. The safety reserve is
|
||||
the tolerance, so a margin between 0 and the reserve is still `fits`.
|
||||
|
||||
A discrepancy is recorded, never acted on: the reply has already streamed
|
||||
to the reader and is accepted story.
|
||||
"""
|
||||
prompt = usage.get("prompt_tokens") if isinstance(usage, dict) else None
|
||||
reserve = safety_reserve(budget)
|
||||
verified_note = "" if window_verified else (
|
||||
" The window itself was not verified for this turn.")
|
||||
record = {
|
||||
"status": UNKNOWN,
|
||||
"server_prompt_tokens": None,
|
||||
"estimate": estimate,
|
||||
"difference": None,
|
||||
"budget": budget,
|
||||
"max_output_tokens": max_output_tokens,
|
||||
"safety_reserve": reserve,
|
||||
"observed_margin": None,
|
||||
"window_verified": bool(window_verified),
|
||||
"detail": "",
|
||||
}
|
||||
if type(prompt) is not int or prompt <= 0:
|
||||
record["detail"] = ("The server reported no prompt token count, so nothing "
|
||||
"confirms the whole prompt was read." + verified_note)
|
||||
return record
|
||||
|
||||
record["server_prompt_tokens"] = prompt
|
||||
record["difference"] = prompt - estimate
|
||||
record["observed_margin"] = budget - max_output_tokens - prompt
|
||||
if prompt + reserve < estimate:
|
||||
record["status"] = TRUNCATION_SUSPECTED
|
||||
record["detail"] = (
|
||||
f"The server read {prompt:,} prompt tokens of the {estimate:,} sent, a "
|
||||
f"shortfall larger than the {reserve:,}-token safety reserve. A server "
|
||||
"that cuts an over-window prompt reports exactly this, and what it cuts "
|
||||
"is the start: the narrator's rules and the canon." + verified_note)
|
||||
elif prompt + max_output_tokens > budget:
|
||||
record["status"] = EXCEEDED
|
||||
record["detail"] = (
|
||||
f"The server counted {prompt:,} prompt tokens; with {max_output_tokens:,} "
|
||||
f"for the reply that is more than the {budget:,}-token window the prompt "
|
||||
"was built for, so the reply may have been cut short." + verified_note)
|
||||
else:
|
||||
record["status"] = FITS
|
||||
record["detail"] = (
|
||||
f"The server read {prompt:,} prompt tokens, leaving "
|
||||
f"{record['observed_margin']:,} beside the reply." + verified_note)
|
||||
return record
|
||||
|
||||
|
||||
def cache_clear() -> None:
|
||||
"""Forgets what was learned. Called when the endpoint or model changes."""
|
||||
_cache.clear()
|
||||
@@ -223,6 +339,7 @@ def _declared_or(declared: int | None, discovered: Window) -> Window:
|
||||
declared, DECLARED, discovered.model_max,
|
||||
f"{declared:,} tokens, declared in settings — the server was not able "
|
||||
f"to say ({discovered.detail})",
|
||||
reachable=discovered.reachable,
|
||||
)
|
||||
|
||||
|
||||
@@ -266,6 +383,7 @@ async def _ask(endpoint_url: str, model: str) -> Window:
|
||||
return Window(
|
||||
tokens, LOADED, ceiling,
|
||||
f"{tokens:,} tokens, reported by the running model",
|
||||
reachable=True,
|
||||
)
|
||||
return await _declared_window(client, base, model)
|
||||
except (httpx.HTTPError, ValueError, TypeError, KeyError) as exc:
|
||||
@@ -293,6 +411,7 @@ async def _declared_window(client, base: str, model: str) -> Window:
|
||||
return Window(
|
||||
None, UNKNOWN,
|
||||
detail=f"the server did not describe the model (HTTP {resp.status_code})",
|
||||
reachable=True,
|
||||
)
|
||||
body = resp.json() or {}
|
||||
ceiling = _architecture_ceiling(body.get("model_info") or {})
|
||||
@@ -304,14 +423,103 @@ async def _declared_window(client, base: str, model: str) -> Window:
|
||||
"the model sets no num_ctx, so the server will load it at its own "
|
||||
"default — which is 4,096 where there is no VRAM"
|
||||
),
|
||||
reachable=True,
|
||||
)
|
||||
tokens = min(declared, ceiling) if ceiling else declared
|
||||
return Window(
|
||||
tokens, PARAMETERS, ceiling,
|
||||
f"{tokens:,} tokens, from the model's own num_ctx",
|
||||
reachable=True,
|
||||
)
|
||||
|
||||
|
||||
#: v1.1 WP-A1 corrective: loading the configured model so its window can be read.
|
||||
#:
|
||||
#: The first real turn of the A1 evidence found a cold model: `/api/ps` knew
|
||||
#: nothing, `/api/show` found no `num_ctx`, so the window was unverified and the
|
||||
#: prompt was built to the configured 16,384. Ollama loaded the model at its own
|
||||
#: 4,096 default, kept 2,050 of 13,875 tokens and answered 200. That case is
|
||||
#: preventable, because the window becomes readable the moment the model is
|
||||
#: resident. Ollama's native `POST /api/generate` with a model and **no prompt**
|
||||
#: loads the model and generates nothing — measured on Ollama 0.33: HTTP 200,
|
||||
#: `"response": ""`, `"done_reason": "load"`, and `/api/ps` then reported the
|
||||
#: window. The OpenAI-compatible request that followed did not reload it.
|
||||
WARM_PATH = "/api/generate"
|
||||
|
||||
|
||||
async def warm(endpoint_url: str, model: str, *, timeout: float) -> tuple[bool, str]:
|
||||
"""Asks the configured server to load `model`. One request, no story text.
|
||||
|
||||
Held to the same endpoint policy and TLS trust as inference and the probe, and
|
||||
sent to the same host the probe asks. The body names the model and nothing
|
||||
else: no prompt, so nothing is generated, and no `options` or `keep_alive`, so
|
||||
the model loads the way the server would load it for the turn itself.
|
||||
|
||||
Returns `(loaded, detail)`. Every failure is `(False, why)` and never raises:
|
||||
a server that will not load the model on request will fail the turn's own
|
||||
call the ordinary way, which is where that failure belongs.
|
||||
"""
|
||||
reason = endpoints.rejection_reason(endpoint_url)
|
||||
if reason is not None:
|
||||
return False, f"endpoint not allowed — {reason}"
|
||||
base = native_base(endpoint_url)
|
||||
try:
|
||||
async with httpx.AsyncClient(
|
||||
timeout=httpx.Timeout(timeout, connect=CONNECT_TIMEOUT),
|
||||
verify=tlstrust.ssl_context(),
|
||||
) as client:
|
||||
resp = await client.post(f"{base}{WARM_PATH}", json={"model": model})
|
||||
except httpx.HTTPError as exc:
|
||||
log.debug("model warm-up failed for %s: %s", base, exc)
|
||||
return False, f"could not ask the server to load the model ({type(exc).__name__})"
|
||||
if resp.status_code != 200:
|
||||
return False, f"the server did not load the model (HTTP {resp.status_code})"
|
||||
try:
|
||||
body = resp.json() or {}
|
||||
except ValueError:
|
||||
return False, "the server answered the load request with something that was not JSON"
|
||||
return True, f"the server loaded the model ({body.get('done_reason') or 'done'})"
|
||||
|
||||
|
||||
async def ensure_window(endpoint_url: str, model: str, *, declared: int | None = None,
|
||||
warm_timeout: float = 300.0) -> tuple[Window, dict]:
|
||||
"""The window for a turn about to be generated, loading the model once if that is what it takes.
|
||||
|
||||
1. Probe as before.
|
||||
2. If the window is not verified, the server answered, and there is a model to
|
||||
load: one bounded `warm` request.
|
||||
3. If the model loaded, probe again, bypassing the cache that still holds the
|
||||
unverified answer.
|
||||
|
||||
Whatever the second probe says is the answer. There is no retry loop, no
|
||||
guessed window, and no hard-coded 4,096: a window still unverified leaves the
|
||||
configured budget standing, exactly as before, and the turn's accounting
|
||||
still catches a server that cut the prompt.
|
||||
|
||||
Returns the window and a `preflight` record for the turn's provenance.
|
||||
Not used by the context dry run: loading a model is a side effect, and
|
||||
opening a panel should not cause one.
|
||||
"""
|
||||
window = await probe(endpoint_url, model, declared=declared)
|
||||
preflight = {"attempted": False, "loaded": None, "verified_before": window.verified,
|
||||
"verified_after": window.verified, "detail": ""}
|
||||
if window.verified:
|
||||
preflight["detail"] = "the window was already verified"
|
||||
return window, preflight
|
||||
if not (endpoint_url and model):
|
||||
preflight["detail"] = "no endpoint or model configured"
|
||||
return window, preflight
|
||||
if not window.reachable:
|
||||
preflight["detail"] = "the server did not answer, so no model was loaded"
|
||||
return window, preflight
|
||||
loaded, detail = await warm(endpoint_url, model, timeout=warm_timeout)
|
||||
preflight.update(attempted=True, loaded=loaded, detail=detail)
|
||||
if loaded:
|
||||
window = await probe(endpoint_url, model, declared=declared, use_cache=False)
|
||||
preflight["verified_after"] = window.verified
|
||||
return window, preflight
|
||||
|
||||
|
||||
def _num_ctx(parameters) -> int | None:
|
||||
"""Reads `num_ctx` out of the plain-text parameter block Ollama returns."""
|
||||
if not isinstance(parameters, str):
|
||||
|
||||
@@ -35,6 +35,8 @@ creating a second, empty Mara.
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
|
||||
# Field types the schema layer enforces. Kept deliberately small: a narrative
|
||||
# state event carries names, labels and plain values, and nothing here needs a
|
||||
# nested structure a model could hide something inside.
|
||||
@@ -184,11 +186,30 @@ def vocabulary_for_prompt() -> str:
|
||||
Generated from `SPECS` rather than written out beside it, so the model can
|
||||
never be told about an event the application does not implement — the drift
|
||||
that would produce proposals rejected for reasons nobody could see.
|
||||
|
||||
v1.1 WP-A2: each event is shown as the object the model must put in the
|
||||
`events` list, with its required fields, not as `name(field, …)`. The call
|
||||
notation was never the wire format, and a 3B narrator copied it into its
|
||||
prose as `> set_possession(silver-key, "alice")`. An object copied into prose
|
||||
is a proposal the extractor already recognises and removes; a call is not.
|
||||
"""
|
||||
lines = []
|
||||
for name, definition in SPECS.items():
|
||||
fields = list(definition["required"]) + [
|
||||
f"{field}?" for field in definition["optional"]
|
||||
]
|
||||
lines.append(f' {name}({", ".join(fields)}) — {definition["summary"]}')
|
||||
shape = {"type": name}
|
||||
for field, kind in definition["required"].items():
|
||||
shape[field] = _PLACEHOLDER[kind]
|
||||
body = json.dumps(shape, ensure_ascii=False, separators=(",", ":"))
|
||||
line = f" {body} — {definition['summary']}"
|
||||
if definition["optional"]:
|
||||
line += f" (optional: {', '.join(definition['optional'])})"
|
||||
lines.append(line)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
#: What a field of each kind looks like in the prompt's vocabulary. Placeholders,
|
||||
#: never example identifiers, so the vocabulary names nothing a story could copy.
|
||||
#: A list field is shown as a list, so the model is told its shape; every other
|
||||
#: field is an ellipsis. Measured: `"<key>"`-style placeholders with spaced
|
||||
#: separators cost 456 tokens against v1.0.0's 258; this form costs about 380,
|
||||
#: and every line is still the object the model must send.
|
||||
_PLACEHOLDER = {KEY: "…", TEXT: "…", VALUE: "…", LABELS: ["…"]}
|
||||
|
||||
@@ -37,17 +37,17 @@ EMIT_RULE = (
|
||||
"appeared, record it.\n"
|
||||
"\n"
|
||||
"Every value is ABSOLUTE — the new state of things, never a change or a "
|
||||
"difference. Use only these events:\n"
|
||||
"difference. Use only these events, in exactly this shape:\n"
|
||||
f"{events.vocabulary_for_prompt()}\n"
|
||||
"\n"
|
||||
"Identifiers are short lower-case slugs (mara, silver-key, old-abbey) and must "
|
||||
"match the ones already in the state you were shown. Introduce a person, place "
|
||||
"or thing with create_entity before referring to it. If the turn established "
|
||||
"nothing, send an empty events list.\n"
|
||||
"Identifiers are short lower-case slugs and must match the ones already in the "
|
||||
"state you were shown; the example's identifiers are placeholders. Introduce a "
|
||||
"person, place or thing with create_entity before referring to it. If the turn "
|
||||
"established nothing, send an empty events list.\n"
|
||||
"Example:\n"
|
||||
'```state\n'
|
||||
'{"events": [{"type": "set_possession", "item": "silver-key", "owner": "aldric"},'
|
||||
' {"type": "set_current_location", "entity": "aldric", "location": "old-abbey"}]}\n'
|
||||
'{"events": [{"type": "set_possession", "item": "item-1", "owner": "character-1"},'
|
||||
' {"type": "set_current_location", "entity": "character-1", "location": "location-1"}]}\n'
|
||||
'```'
|
||||
)
|
||||
|
||||
@@ -58,6 +58,47 @@ EMIT_REMINDER = (
|
||||
"nothing changed.]"
|
||||
)
|
||||
|
||||
# v1.1 WP-A2: the length hint's own words, named once. `builder.length_hint`
|
||||
# builds the hint from these, and the extractor recognises an echo of it by
|
||||
# them, so the two cannot drift apart.
|
||||
LENGTH_HINT_OPENING = "[Hard limit:"
|
||||
LENGTH_HINT_TAIL = "Finish the narration and append the state block well inside the limit."
|
||||
#: The application's wording inside a hint. A 3B narrator reworded the front
|
||||
#: ("your next turn") and the end ("This story ends here."), and kept one or the
|
||||
#: other of these every time.
|
||||
_LENGTH_HINT_PHRASE_RE = re.compile(
|
||||
r"append the state block|turn must not exceed \d+ words", re.IGNORECASE
|
||||
)
|
||||
|
||||
#: v1.1 WP-A2: the rules that remove protocol a narrator copied, named so the
|
||||
#: replay tool and the report can say which removed what.
|
||||
RULE_EVENT_CALL = "event_call_line"
|
||||
RULE_LENGTH_HINT = "echoed_length_hint"
|
||||
RULE_SCENE_LINE = "rendered_scene_line"
|
||||
RULE_EMPTY_FENCE = "empty_dangling_fence"
|
||||
RULE_INSTRUCTION_TAIL = "echoed_instruction_tail"
|
||||
|
||||
#: v1.1 WP-A2 corrective (R5). The sentence `CHAT_CONTINUE_HINT` in
|
||||
#: `providers/openai_compatible.py` carries, which a narrator echoed with the rest
|
||||
#: of the hint reworded around it. Kept as a copy rather than an import, so the
|
||||
#: narrative package does not depend on the provider; a test pins that the
|
||||
#: hint still contains it.
|
||||
CONTINUE_HINT_PHRASE = "Output only story text"
|
||||
|
||||
# R1. A whole line opening with a call to an event this protocol has. The names
|
||||
# come from the vocabulary, so a call-shaped line naming anything else — a
|
||||
# character's `open_door(north)` — is not matched.
|
||||
_EVENT_CALL_LINE_RE = re.compile(
|
||||
r"^[ \t]*(?:>[ \t]*)?(?:"
|
||||
+ "|".join(re.escape(name) for name in events.SPECS)
|
||||
+ r")[ \t]*\(",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# R3. The renderer's scene line carries its location this way.
|
||||
_RENDERED_SCENE_LOCATION_RE = re.compile(r"\(at [^()\n]+\)\s*$")
|
||||
# R4. An opener with nothing after it.
|
||||
_EMPTY_FENCE_LINE_RE = re.compile(r"```(?:json)?[ \t]*", re.IGNORECASE)
|
||||
|
||||
# Three patterns, and the difference between them is the whole of this module's
|
||||
# safety. A story is allowed to contain code, and taking a code block out of
|
||||
# someone's prose is a worse failure than leaving a stray proposal in it.
|
||||
@@ -128,11 +169,42 @@ def _is_echoed_instruction(inner: str) -> bool:
|
||||
# opening words only, because the echo is often cut off before it ends.
|
||||
if low.lstrip().startswith("continue the story directly"):
|
||||
return True
|
||||
# v1.1 WP-A2 corrective (R5): the same hint, reworded at the front. The M11
|
||||
# closeout-era identity re-run stored "[You don't need to continue; … Continue
|
||||
# the story here, directly. Output only story text.]" as the last line of a
|
||||
# reply, and because nothing recognised it, nothing above it was trailing.
|
||||
if CONTINUE_HINT_PHRASE.lower() in low:
|
||||
return True
|
||||
# v1.1 WP-A2 (R2): the length hint, which names "state block" but not
|
||||
# "events list", so it passed every check above.
|
||||
if _is_length_hint(inner):
|
||||
return True
|
||||
# The reminder names both; prose about the protocol rarely names either the
|
||||
# way the instruction does, and effectively never both.
|
||||
return "state block" in low and "events list" in low
|
||||
|
||||
|
||||
def _opens_like_length_hint(inner: str) -> bool:
|
||||
"""R5. The bracket opens with the length hint's own `Hard limit:`, whatever follows.
|
||||
|
||||
Never enough on its own: an in-world "[Hard limit: forty days]" opens the same
|
||||
way. `_clean` takes it only directly above an echoed instruction it has already
|
||||
removed from the end of the same reply.
|
||||
"""
|
||||
return inner.lstrip().lower().startswith(LENGTH_HINT_OPENING[1:].lower())
|
||||
|
||||
|
||||
def _is_length_hint(inner: str) -> bool:
|
||||
"""Whether a bracket's contents are `builder.length_hint`, however reworded.
|
||||
|
||||
It must open the way the hint opens *and* carry the hint's own wording. An
|
||||
in-world "Hard limit: forty days" has the opening and none of the wording.
|
||||
"""
|
||||
opening = LENGTH_HINT_OPENING[1:].lower()
|
||||
return (inner.lstrip().lower().startswith(opening)
|
||||
and bool(_LENGTH_HINT_PHRASE_RE.search(inner)))
|
||||
|
||||
|
||||
# A heading the model writes above a block it did not fence: `State`, sometimes
|
||||
# as `State:`, `**State**` or `### State`. It is removed only in two places:
|
||||
# directly above a proposal that is removed, and as the last line of the reply.
|
||||
@@ -145,7 +217,7 @@ _LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
|
||||
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
|
||||
|
||||
|
||||
def _clean(prose: str) -> str:
|
||||
def _clean(prose: str, *, after_block: bool = False) -> str:
|
||||
"""Removes protocol the block extraction could not, and nothing else.
|
||||
|
||||
Found by the M5 realistic-context run (§12), which is the failure class
|
||||
@@ -160,9 +232,23 @@ def _clean(prose: str) -> str:
|
||||
story after it. Stored text is replayed as history, so every leak also
|
||||
showed the next prompt a second, older account of the state, which is what
|
||||
M5 review Finding 4 removed from replayed history.
|
||||
|
||||
v1.1 WP-A2 added four shapes, from the M11 closeout's identity run and the
|
||||
v1 corpus, each anchored to something the application owns rather than to
|
||||
what prose looks like: a line opening with a vocabulary call (R1), the
|
||||
length hint echoed at the end (R2), the renderer's scene line left last
|
||||
(R3), and an empty fence opener left last (R4). `after_block` says a
|
||||
proposal block was already taken out of this reply, which is what lets R3
|
||||
remove a bare scene line that sat above it.
|
||||
"""
|
||||
cleaned, _found = _inline_proposals(prose)
|
||||
cleaned, calls_removed = _strip_event_call_lines(prose)
|
||||
cleaned, _found = _inline_proposals(cleaned)
|
||||
cleaned = _strip_echoed_state(cleaned)
|
||||
protocol_cut = after_block or calls_removed
|
||||
# R5: set once an echoed instruction bracket has come off the end. Only then
|
||||
# may a bracket that merely opens the way the length hint opens be taken as
|
||||
# part of the same echoed tail.
|
||||
instruction_cut = False
|
||||
# The end of the reply is cut until nothing more comes off, because one kind
|
||||
# of leftover can hide another. In a real reply, a `State` heading sat above
|
||||
# a block the model never finished, and a parroted reminder sat above an
|
||||
@@ -171,7 +257,12 @@ def _clean(prose: str) -> str:
|
||||
before = cleaned
|
||||
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
|
||||
bracket = pattern.search(cleaned)
|
||||
if bracket is not None and _is_echoed_instruction(bracket.group(1)):
|
||||
if bracket is None:
|
||||
continue
|
||||
if _is_echoed_instruction(bracket.group(1)):
|
||||
cleaned = cleaned[: bracket.start()]
|
||||
instruction_cut = True
|
||||
elif instruction_cut and _opens_like_length_hint(bracket.group(1)):
|
||||
cleaned = cleaned[: bracket.start()]
|
||||
cleaned = _DANGLING_STATE_RE.sub("", cleaned)
|
||||
dangling = _DANGLING_JSON_RE.search(cleaned)
|
||||
@@ -182,10 +273,93 @@ def _clean(prose: str) -> str:
|
||||
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
|
||||
# A bare quote marker, the start of a quoted block that never came.
|
||||
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
|
||||
cleaned = _strip_empty_dangling_fence(cleaned)
|
||||
if cleaned.rstrip() != before.rstrip():
|
||||
protocol_cut = True
|
||||
cleaned = _strip_trailing_scene_line(cleaned, protocol_cut)
|
||||
if cleaned == before:
|
||||
return cleaned.strip()
|
||||
|
||||
|
||||
def _strip_event_call_lines(text: str) -> tuple[str, bool]:
|
||||
"""R1. Removes whole lines that open with a call to a vocabulary event.
|
||||
|
||||
A line inside a fenced code block is the story's own code and is never
|
||||
examined. Returns the text and whether anything was removed.
|
||||
"""
|
||||
kept: list[str] = []
|
||||
in_fence = False
|
||||
removed = False
|
||||
for line in text.split("\n"):
|
||||
if line.lstrip().startswith("```"):
|
||||
in_fence = not in_fence
|
||||
kept.append(line)
|
||||
continue
|
||||
if not in_fence and _EVENT_CALL_LINE_RE.match(line):
|
||||
removed = True
|
||||
continue
|
||||
kept.append(line)
|
||||
if not removed:
|
||||
return text, False
|
||||
return re.sub(r"\n{3,}", "\n\n", "\n".join(kept)), True
|
||||
|
||||
|
||||
def _strip_empty_dangling_fence(text: str) -> str:
|
||||
"""R4. A ```` ```json ```` or ```` ``` ```` opener as the last line, with nothing after it.
|
||||
|
||||
Only an *opener*: the fence lines are counted, and an even count means the
|
||||
last one closes a story's own code block, which stays.
|
||||
"""
|
||||
lines = text.rstrip().split("\n")
|
||||
if len(lines) < 2 or not _EMPTY_FENCE_LINE_RE.fullmatch(lines[-1].strip()):
|
||||
return text
|
||||
fences = sum(1 for line in lines if line.lstrip().startswith("```"))
|
||||
if fences % 2 == 0:
|
||||
return text
|
||||
return "\n".join(lines[:-1]).rstrip()
|
||||
|
||||
|
||||
def _strip_trailing_scene_line(text: str, protocol_cut: bool) -> str:
|
||||
"""R3. The renderer's scene line, left as the last line of the reply.
|
||||
|
||||
Taken when it carries the renderer's own `(at <location>)`, or when protocol
|
||||
was already cut from this reply, which makes a bare scene line part of the
|
||||
same pasted tail. A final screenplay-style "Scene: …" line in a reply with
|
||||
no protocol in it stays, and so does any scene line with story after it.
|
||||
"""
|
||||
lines = text.rstrip().split("\n")
|
||||
if len(lines) < 2:
|
||||
return text
|
||||
last = lines[-1].strip()
|
||||
if not last.startswith(render.HEADING_SCENE + " "):
|
||||
return text
|
||||
if not (_RENDERED_SCENE_LOCATION_RE.search(last) or protocol_cut):
|
||||
return text
|
||||
return "\n".join(lines[:-1]).rstrip()
|
||||
|
||||
|
||||
def explain_removed_line(line: str) -> str | None:
|
||||
"""Which v1.1 rule removes a line of this shape, for the replay report.
|
||||
|
||||
None means no v1.1 rule explains it, which the replay treats as a failure.
|
||||
"""
|
||||
stripped = line.strip()
|
||||
if _EVENT_CALL_LINE_RE.match(line):
|
||||
return RULE_EVENT_CALL
|
||||
if stripped.startswith("["):
|
||||
inner = stripped[1:]
|
||||
inner = inner[:-1] if inner.endswith("]") else inner
|
||||
if _is_length_hint(inner):
|
||||
return RULE_LENGTH_HINT
|
||||
if _is_echoed_instruction(inner) or _opens_like_length_hint(inner):
|
||||
return RULE_INSTRUCTION_TAIL
|
||||
if stripped.startswith(render.HEADING_SCENE + " "):
|
||||
return RULE_SCENE_LINE
|
||||
if _EMPTY_FENCE_LINE_RE.fullmatch(stripped):
|
||||
return RULE_EMPTY_FENCE
|
||||
return None
|
||||
|
||||
|
||||
def _is_state_heading(line: str) -> bool:
|
||||
return bool(_STATE_HEADING_RE.match(line))
|
||||
|
||||
@@ -451,7 +625,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
if matches:
|
||||
match = matches[-1]
|
||||
raw = match.group(1).strip()
|
||||
prose = _clean(text[: match.start()] + text[match.end():])
|
||||
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
||||
return prose, _tolerant_load(raw), raw
|
||||
|
||||
# A `json` or unlabelled fence is ours only when its contents are this
|
||||
@@ -465,7 +639,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
raw = match.group(1).strip()
|
||||
parsed = _tolerant_load(raw)
|
||||
if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
|
||||
prose = _clean(text[: match.start()] + text[match.end():])
|
||||
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
||||
return prose, parsed, raw
|
||||
|
||||
match = _TRAILING_RE.search(text)
|
||||
@@ -473,7 +647,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
raw = match.group(1)
|
||||
parsed = _tolerant_load(raw)
|
||||
if _looks_like_proposal(parsed):
|
||||
return _clean(text[: match.start()]), parsed, raw
|
||||
return _clean(text[: match.start()], after_block=True), parsed, raw
|
||||
|
||||
# An unfenced proposal on its own lines but not at the end: quoted, or
|
||||
# followed by more story. The last one is the turn's proposal, as with
|
||||
@@ -481,7 +655,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
without, found = _inline_proposals(text)
|
||||
if found:
|
||||
parsed, raw = found[-1]
|
||||
return _clean(without), parsed, raw
|
||||
return _clean(without, after_block=True), parsed, raw
|
||||
|
||||
# No block at all — but the reply may still carry protocol the model wrote
|
||||
# as prose, or a fence it never closed.
|
||||
|
||||
@@ -26,6 +26,10 @@ EMBED_READ_TIMEOUT = 60.0
|
||||
|
||||
|
||||
|
||||
#: v1.1 WP-A1: ask a stream to report its token usage. Without it Ollama sends
|
||||
#: none, and a prompt the server cut cannot be told from one it read whole.
|
||||
STREAM_OPTIONS = {"include_usage": True}
|
||||
|
||||
# Completion endpoints have no roles, so a chat has to be flattened into one
|
||||
# labeled transcript that ends on "Assistant:" for the model to continue.
|
||||
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
|
||||
@@ -84,10 +88,14 @@ class OpenAICompatibleProvider(Provider):
|
||||
def _record_usage(self, payload: dict) -> None:
|
||||
"""Records the endpoint's own token accounting, if it reported any.
|
||||
|
||||
OpenRouter now always reports usage, and `usage: {include: true}` and
|
||||
`stream_options` are deprecated and do nothing. In a stream the usage
|
||||
arrives on a final chunk that carries no choices, which is why this is
|
||||
read separately from the text extraction.
|
||||
In a stream the usage arrives on a final chunk that carries no choices,
|
||||
which is why this is read separately from the text extraction.
|
||||
|
||||
v1.1 WP-A1: Ollama sends that chunk only when asked. Measured on Ollama
|
||||
0.33: a stream with no `stream_options` carried no usage at all, and not
|
||||
one of the 514 AI turns in the v1 evidence had a count stored. Every
|
||||
streaming body therefore sets `stream_options.include_usage`
|
||||
(`STREAM_OPTIONS`), and the turn compares the count with what it sent.
|
||||
"""
|
||||
usage = payload.get("usage")
|
||||
if isinstance(usage, dict) and usage:
|
||||
@@ -102,6 +110,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
else:
|
||||
url = f"{self.base_url}/chat/completions"
|
||||
@@ -114,6 +123,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
return url, body
|
||||
|
||||
@@ -183,6 +193,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
else:
|
||||
url = f"{self.base_url}/chat/completions"
|
||||
@@ -192,6 +203,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
async for event in self._stream(url, body):
|
||||
yield event
|
||||
|
||||
@@ -6,6 +6,7 @@ lock guards one set only while one module owns it. And a test that replaces
|
||||
`OpenAICompatibleProvider` or `generate_turn` patches this module, which every
|
||||
caller reads through.
|
||||
"""
|
||||
import logging
|
||||
import threading
|
||||
|
||||
from fastapi import Depends, HTTPException, Request
|
||||
@@ -28,6 +29,8 @@ from .deps import CurrentUser, current_adventure, router
|
||||
from .nodes import _move_to_after, next_depth
|
||||
from .paging import annotate_takes
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def world_delta_of(snapshot: dict | None) -> dict | None:
|
||||
"""Returns the bulk-read slice of a context snapshot, for `Action.world_delta`.
|
||||
@@ -212,8 +215,18 @@ async def _generate_turn(
|
||||
# network calls — and cached per endpoint and model, so it costs one short
|
||||
# request per session rather than one per turn. An unverified window does
|
||||
# not block the turn; it is recorded as unverified in the snapshot below.
|
||||
window = await contextwindow.probe(settings.endpoint_url, settings.model,
|
||||
declared=settings.context_window_override)
|
||||
#
|
||||
# v1.1 WP-A1 corrective: a model that is not resident cannot report its window,
|
||||
# and a turn built to the configured budget against it was silently cut in the
|
||||
# A1 evidence (13,875 tokens sent, 2,050 read). So an unverified window gets
|
||||
# one bounded attempt to load the model, and one more probe, before the
|
||||
# prompt is assembled. No story text is generated by it and nothing is
|
||||
# written. A window still unverified afterwards changes nothing below.
|
||||
window, preflight = await contextwindow.ensure_window(
|
||||
settings.endpoint_url, settings.model,
|
||||
declared=settings.context_window_override,
|
||||
warm_timeout=float(settings.model_timeout_seconds or 300),
|
||||
)
|
||||
try:
|
||||
system_text, story_text, snapshot = build_context(
|
||||
adventure,
|
||||
@@ -232,6 +245,9 @@ async def _generate_turn(
|
||||
yield turn_error(str(exc))
|
||||
return
|
||||
|
||||
if isinstance(snapshot.get("window"), dict):
|
||||
snapshot["window"]["preflight"] = preflight
|
||||
|
||||
parts = PromptParts(system=system_text, story=story_text)
|
||||
|
||||
provider = OpenAICompatibleProvider(
|
||||
@@ -318,6 +334,24 @@ async def _generate_turn(
|
||||
# prompt came from cache rather than being billed in full. This is recorded
|
||||
# per attempt, next to the prompt it priced.
|
||||
snapshot["usage"] = provider.last_usage
|
||||
# v1.1 WP-A1: what the server says it read, against what was sent. Recorded
|
||||
# and shown, never acted on: the narration has already streamed to the
|
||||
# reader, and discarding an accepted turn over an accounting discrepancy
|
||||
# would lose story to hide a problem. A server that cut the prompt answers
|
||||
# 200 either way, so this record is the only place the cut is visible.
|
||||
tokens = snapshot.get("tokens") or {}
|
||||
accounting = contextwindow.classify_usage(
|
||||
provider.last_usage,
|
||||
estimate=tokens.get("estimate") or tokens.get("total") or 0,
|
||||
budget=tokens.get("budget") or settings.context_token_budget,
|
||||
max_output_tokens=settings.max_output_tokens,
|
||||
window_verified=bool((snapshot.get("window") or {}).get("verified")),
|
||||
)
|
||||
snapshot["accounting"] = accounting
|
||||
if accounting["status"] in (contextwindow.EXCEEDED,
|
||||
contextwindow.TRUNCATION_SUSPECTED):
|
||||
log.warning("turn accounting for adventure %s: %s — %s",
|
||||
adventure.id, accounting["status"], accounting["detail"])
|
||||
|
||||
reasoning = "".join(reasoning_chunks).strip() or None
|
||||
ai_action = models.Action(
|
||||
@@ -381,7 +415,8 @@ async def _generate_turn(
|
||||
db.commit()
|
||||
db.refresh(ai_action)
|
||||
yield _SAVED
|
||||
yield sse({"type": "done", "action": action_json(ai_action, db)})
|
||||
yield sse({"type": "done", "action": action_json(ai_action, db),
|
||||
"accounting": accounting})
|
||||
# Phase 6: schedule summarization and embedding without waiting for them.
|
||||
# The task opens its own database session.
|
||||
memorybank.schedule_post_turn(adventure)
|
||||
|
||||
Reference in New Issue
Block a user