v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
ac465ed867
commit
d63804f22e
@@ -620,6 +620,35 @@ campaign gets less history than the setting asks for, which is a visible,
|
||||
explicable loss rather than a silent one, and Settings' **Test connection**
|
||||
reports the window it found or says plainly that it could not check.
|
||||
|
||||
**It also keeps a margin, and checks the server's own count (v1.1).** The
|
||||
application counts tokens with `cl100k_base`, and your model counts them with
|
||||
its own tokenizer. The two disagree slightly, so the prompt is built to leave
|
||||
`max(256, 5% of the window)` tokens free on top of the reply: 256 at 4,096, and
|
||||
820 at 16,384. After each turn the server's reported prompt-token count is
|
||||
compared with what was sent. The context inspector shows the result for any
|
||||
past turn:
|
||||
|
||||
- **The server read the whole prompt:** the ordinary case.
|
||||
- **The server did not say how much it read:** the server reported no usage.
|
||||
Nothing is wrong, and nothing is confirmed either.
|
||||
- **The server may have cut the start of the prompt:** it read far fewer tokens
|
||||
than were sent. Ollama does this, silently, to a prompt larger than the window
|
||||
the model was loaded with. The turn is kept. Check the window with the
|
||||
commands above.
|
||||
- **The prompt was larger than the server allowed for:** its count and the reply
|
||||
together exceed the window. The reply may have been cut short. The turn is
|
||||
kept.
|
||||
|
||||
The last two also appear in the server log as a warning.
|
||||
|
||||
**A model that is not loaded yet is loaded first.** Before a turn, if the
|
||||
application cannot read the window because your model isn't in memory, it asks
|
||||
the same Ollama to load it once. That is a `POST /api/generate` naming only the
|
||||
model, which generates no text. It then reads the window again, so the first
|
||||
turn of a session is built to the window the model really has rather than to
|
||||
your setting. If loading fails, or the window still can't be read, the turn goes
|
||||
ahead exactly as before, unverified, and the check above still applies.
|
||||
|
||||
That does not make the window *bigger*, and the rest of this section is still
|
||||
how you do that.
|
||||
|
||||
|
||||
@@ -80,6 +80,14 @@ that isn't the live one starts a new branch.
|
||||
in settings lets you state the window so the prompt is still capped. A window the server
|
||||
itself reported always wins over that, and a declared one is never reported as verified.
|
||||
|
||||
Since v1.1 the prompt also stops short of that window on purpose. It leaves
|
||||
`max(256, 5% of the window)` tokens free, because your model counts tokens differently from
|
||||
the application, and the v1 evidence came within 23 tokens of the edge. After each turn,
|
||||
the server's own count of what it read is compared with what was sent. A turn the server
|
||||
appears to have truncated is kept, flagged and shown in the context inspector, not left to
|
||||
pass silently. A model that isn't loaded yet, and so cannot report its window, is loaded
|
||||
once before the turn is built, so the first turn of a session gets the real window too.
|
||||
|
||||
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
|
||||
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
|
||||
card used to arrive in front of it as a world fact with no class, no visibility, no source and
|
||||
@@ -179,9 +187,9 @@ that isn't the live one starts a new branch.
|
||||
|
||||
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
|
||||
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
|
||||
removed rather than left standing as a picture of a product that no longer exists. The M4
|
||||
closeout drove the real application in a real browser, so the screens exist and work; taking
|
||||
presentable screenshots of them is a job for the UI pass in M8.
|
||||
removed rather than left standing as a picture of a product that no longer exists. The
|
||||
screens exist and are driven in a real browser by the release harness
|
||||
(`backend/tools/m11_browser.py`). Presentable screenshots of them have not been taken.
|
||||
|
||||
## Quick start
|
||||
|
||||
@@ -310,7 +318,8 @@ development, Vite proxies `/api` to FastAPI.
|
||||
|
||||
## Tests
|
||||
|
||||
1,191 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||||
1,524 backend tests (1,507 run everywhere, 17 need a real local model and skip without one): unit
|
||||
tests plus full HTTP integration through the real turn engine, with
|
||||
the model provider mocked. They run with no route to the Internet, which is a requirement
|
||||
rather than a convenience — an offline claim proved on a machine that has been online once
|
||||
proves nothing. A further handful need a real local model and skip without one; they exist
|
||||
|
||||
@@ -46,7 +46,14 @@ from .narrative import model as narrative_model
|
||||
# token accounting. Each attempt is its own API call, and a retry is the call
|
||||
# most likely to read the prompt back out of cache. Everything else in a snapshot
|
||||
# is the prompt, which is assembled once per turn.
|
||||
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage")
|
||||
#
|
||||
# v1.1 WP-A1: `accounting` is one attempt's too. It compares the server's count
|
||||
# for *that* call with the turn's estimate. Left out of this tuple, it was
|
||||
# treated as part of the shared prompt, so moving the live flag handed the
|
||||
# superseded attempt's accounting to the new live one and threw the new one's
|
||||
# away. Found by the A2 long run: two retries and one take selection left three
|
||||
# attempts reporting no accounting, or another attempt's.
|
||||
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage", "accounting")
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ reading
|
||||
|
||||
@@ -25,6 +25,7 @@ from sqlalchemy.orm import object_session
|
||||
|
||||
from .. import contextwindow, derived, models, narrative, summaries, worldstate
|
||||
from ..knowledge import inject as knowledge_inject
|
||||
from ..providers.openai_compatible import CHAT_CONTINUE_HINT
|
||||
from ..knowledge import records as knowledge_records
|
||||
from . import encoding, history
|
||||
|
||||
@@ -106,11 +107,22 @@ BAND_FLOOR_SHARE = 0.5
|
||||
# Built from the table vendored in `encoding.py`, not fetched: the upstream
|
||||
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
|
||||
# is called on every turn.
|
||||
# M6: added to the configured reply budget when reserving output space. It
|
||||
# absorbs the section separators added after budgeting and the drift between
|
||||
# this tokenizer and the serving model's. Fixed rather than proportional: what
|
||||
# it covers does not grow with the size of the budget.
|
||||
OUTPUT_SAFETY_MARGIN = 64
|
||||
#
|
||||
# v1.1 WP-A1: `OUTPUT_SAFETY_MARGIN = 64` was here. M6 added it to the reply
|
||||
# budget to absorb two unrelated things, and v1.1 separates them:
|
||||
#
|
||||
# * **Text the application adds after pricing.** The separators between
|
||||
# sections, and `CHAT_CONTINUE_HINT`, which the provider appends to every chat
|
||||
# request and nothing counted. That is not drift, it is our own text, so it is
|
||||
# now priced exactly (`transport` below).
|
||||
# * **The drift between this tokenizer and the narrator's.** That is what the
|
||||
# 64 tokens were really for, and the v1 evidence showed it was too small. It
|
||||
# is now `contextwindow.safety_reserve`, sized to the window.
|
||||
#
|
||||
#: Story sections that can be joined by `SEPARATOR` after pricing: history,
|
||||
#: author's note, recent history, summary, lore, memories, state, front memory,
|
||||
#: length hint, refusals, reminder. Knowledge and history rows price their own.
|
||||
STORY_SECTION_SLOTS = 11
|
||||
|
||||
|
||||
class ContextOverflow(RuntimeError):
|
||||
@@ -180,19 +192,19 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
words = min(words, band_ceiling)
|
||||
floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
|
||||
tail = (
|
||||
" Finish the narration and append the state block well inside the limit."
|
||||
" " + narrative.extract.LENGTH_HINT_TAIL
|
||||
)
|
||||
if floor < MIN_LENGTH_FLOOR_WORDS:
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words. Write only as "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
|
||||
f"much as the moment needs — a typical turn is much shorter.{tail}]"
|
||||
)
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words, and it should not "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
|
||||
f"stop short of about {floor}. Prefer the lower end of that range unless "
|
||||
f"the scene genuinely needs more.{tail}]"
|
||||
)
|
||||
tail = " Finish the narration and append the state block well inside the limit."
|
||||
tail = " " + narrative.extract.LENGTH_HINT_TAIL
|
||||
|
||||
# State the number as a ceiling, never as a budget. In measurements, the
|
||||
# wording "keep this turn under about N words" read to the model as a target
|
||||
@@ -204,7 +216,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
|
||||
if floor < MIN_LENGTH_FLOOR_WORDS:
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words. Write only as "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
|
||||
f"much as the moment needs — a typical turn is much shorter.{tail}]"
|
||||
)
|
||||
# Both numbers are bounds, and the wording is deliberately asymmetric. The
|
||||
@@ -216,7 +228,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
# so a terse model reading the same clause stops at the floor rather than at
|
||||
# forty words.
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words, and it should not "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
|
||||
f"stop short of about {floor}. Prefer the lower end of that range unless "
|
||||
f"the scene genuinely needs more.{tail}]"
|
||||
)
|
||||
@@ -625,12 +637,19 @@ def build_context(
|
||||
# truncated turn on a model whose window is the budget
|
||||
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
|
||||
#
|
||||
# The margin covers what is added after this arithmetic — the separators
|
||||
# between sections, and the difference between our tokenizer's count and the
|
||||
# serving model's. It is small and fixed rather than proportional, because
|
||||
# what it absorbs does not scale with the budget.
|
||||
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
|
||||
protected = reserved + output_reserve
|
||||
# v1.1 WP-A1: the reply allocation is exactly the reply cap. The text this
|
||||
# application adds after pricing — separators, and the chat hint the
|
||||
# provider appends — is counted as `transport`. What neither can know, the
|
||||
# narrator's tokenizer disagreeing with `cl100k_base`, is the safety reserve,
|
||||
# which is sized to the window and taken before any history is chosen.
|
||||
output_reserve = max(0, settings.max_output_tokens)
|
||||
separator_tokens = count_tokens(SEPARATOR)
|
||||
transport = (
|
||||
separator_tokens * (len(system_sections) + STORY_SECTION_SLOTS)
|
||||
+ count_tokens(CHAT_CONTINUE_HINT)
|
||||
)
|
||||
safety = contextwindow.safety_reserve(budget)
|
||||
protected = reserved + transport + output_reserve + safety
|
||||
if protected >= budget:
|
||||
# Failing here is the point. The alternative — carrying on with a token
|
||||
# or two of history — builds a prompt that is known to overflow, and
|
||||
@@ -638,8 +657,9 @@ def build_context(
|
||||
# gracefully if protected context alone is too large."
|
||||
raise ContextOverflow(
|
||||
f"The protected context needs {protected} tokens "
|
||||
f"({reserved} of prompt plus {output_reserve} reserved for the "
|
||||
f"reply) but the context budget is {budget}. "
|
||||
f"({reserved} of prompt, {transport} of formatting, {output_reserve} "
|
||||
f"reserved for the reply and a {safety}-token safety margin) but the "
|
||||
f"context budget is {budget}. "
|
||||
+ (
|
||||
"That budget is what this server was found to accept, so raising "
|
||||
"the setting alone will not help — load the model with a larger "
|
||||
@@ -827,6 +847,7 @@ def build_context(
|
||||
story_text = SEPARATOR.join(s.text for s in story_sections)
|
||||
|
||||
all_sections = [s for s in system_sections if s.text] + story_sections
|
||||
total_tokens = count_tokens(system_text) + count_tokens(story_text)
|
||||
report = {
|
||||
"sections": [
|
||||
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
|
||||
@@ -837,13 +858,23 @@ def build_context(
|
||||
# what the history was actually allowed to spend after everything
|
||||
# protected was subtracted.
|
||||
"tokens": {
|
||||
"total": count_tokens(system_text) + count_tokens(story_text),
|
||||
"total": total_tokens,
|
||||
"budget": budget,
|
||||
"configured_budget": settings.context_token_budget,
|
||||
"output_reserve": output_reserve,
|
||||
"protected": reserved,
|
||||
"available_for_history": available,
|
||||
"history_spent": spent,
|
||||
# v1.1 WP-A1. `transport` is the formatting priced in above;
|
||||
# `estimate` is what this application believes it actually sent,
|
||||
# the assembled text plus what the provider adds to it, and is what
|
||||
# the server's own count is compared against after the reply.
|
||||
"transport": transport,
|
||||
"safety_reserve": safety,
|
||||
"estimate": total_tokens + (
|
||||
count_tokens(CHAT_CONTINUE_HINT) if settings.api_mode != "completion"
|
||||
else separator_tokens
|
||||
),
|
||||
},
|
||||
# M11: what the server was found to accept, and how. `verified` false
|
||||
# means nobody could check — the prompt was built to the configured
|
||||
|
||||
@@ -86,6 +86,7 @@ one.
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import math
|
||||
import re
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
@@ -132,6 +133,11 @@ class Window:
|
||||
model_max: int | None = None
|
||||
#: Why the window is unknown, or how it was found. Shown to the user.
|
||||
detail: str = ""
|
||||
#: v1.1: the server answered a discovery request at all, whatever it said.
|
||||
#: A server that answered but could not report a window may simply not have
|
||||
#: the model loaded yet, which `ensure_window` can fix; one that did not
|
||||
#: answer cannot be helped by asking it to load anything.
|
||||
reachable: bool = False
|
||||
|
||||
@property
|
||||
def verified(self) -> bool:
|
||||
@@ -180,6 +186,116 @@ def effective_budget(configured: int, window: Window | int | None) -> int:
|
||||
return min(configured, tokens)
|
||||
|
||||
|
||||
#: v1.1 WP-A1: the tokens kept free below the effective window, beyond the reply.
|
||||
#:
|
||||
#: The builder counts with `cl100k_base`; the narrator counts with its own
|
||||
#: tokenizer. The v1 evidence put the largest prompts 23-42 real tokens from the
|
||||
#: edge of a 16,384 window, and Ollama does not refuse a prompt past the edge —
|
||||
#: measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt came back 200
|
||||
#: with `prompt_tokens` 2,050. So the reserve is deliberate and sized to the
|
||||
#: window: the larger of a floor and a share, **rounded up to a whole token**.
|
||||
#:
|
||||
#: 4,096 -> 256 8,192 -> 410 16,384 -> 820
|
||||
#:
|
||||
#: A fixed, documented tolerance, owner-chosen for v1.1. It is not a setting and
|
||||
#: it is not calibrated per model.
|
||||
SAFETY_RESERVE_FLOOR = 256
|
||||
SAFETY_RESERVE_PERCENT = 5
|
||||
|
||||
|
||||
def safety_reserve(effective_window: int) -> int:
|
||||
"""`max(256, ceil(5% of the effective window))`, in tokens.
|
||||
|
||||
The effective window is the budget the prompt is actually built to — the
|
||||
verified or declared window when there is one, the configured budget
|
||||
otherwise — so a 16,384 setting against a 4,096 server reserves 256, not 820.
|
||||
Integer arithmetic, so the rounding is exact rather than a float's.
|
||||
"""
|
||||
share = math.ceil(max(0, effective_window) * SAFETY_RESERVE_PERCENT / 100)
|
||||
return max(SAFETY_RESERVE_FLOOR, share)
|
||||
|
||||
|
||||
#: v1.1 WP-A1: what the server's own count says about a turn that was sent.
|
||||
FITS = "fits"
|
||||
EXCEEDED = "exceeded"
|
||||
TRUNCATION_SUSPECTED = "truncation_suspected"
|
||||
#: `UNKNOWN` above: the server reported no usable count.
|
||||
|
||||
|
||||
def classify_usage(usage: dict | None, *, estimate: int, budget: int,
|
||||
max_output_tokens: int, window_verified: bool) -> dict:
|
||||
"""Sets the server's reported prompt count against what the application sent.
|
||||
|
||||
The order of the checks is the order of what they prove:
|
||||
|
||||
``unknown``
|
||||
No positive integer `prompt_tokens`. Nothing can be said, and nothing
|
||||
is claimed: an absent count is never read as a prompt that fitted.
|
||||
``truncation_suspected``
|
||||
The server read fewer tokens than were sent by more than the safety
|
||||
reserve. A tokenizer thriftier than `cl100k_base` may honestly count a
|
||||
little less; a shortfall larger than the tolerance the application keeps
|
||||
for drift is the signature of a server that cut the prompt — the real
|
||||
shape was 6,316 sent and 2,050 read.
|
||||
``exceeded``
|
||||
The server's count plus the reply allocation is more than the window
|
||||
the prompt was built for. The drift was larger than the whole reserve,
|
||||
so the reply may have been cut short.
|
||||
``fits``
|
||||
Otherwise.
|
||||
|
||||
`observed_margin` is what was left beside the reply by the server's count:
|
||||
`budget - max_output_tokens - server_prompt_tokens`. The safety reserve is
|
||||
the tolerance, so a margin between 0 and the reserve is still `fits`.
|
||||
|
||||
A discrepancy is recorded, never acted on: the reply has already streamed
|
||||
to the reader and is accepted story.
|
||||
"""
|
||||
prompt = usage.get("prompt_tokens") if isinstance(usage, dict) else None
|
||||
reserve = safety_reserve(budget)
|
||||
verified_note = "" if window_verified else (
|
||||
" The window itself was not verified for this turn.")
|
||||
record = {
|
||||
"status": UNKNOWN,
|
||||
"server_prompt_tokens": None,
|
||||
"estimate": estimate,
|
||||
"difference": None,
|
||||
"budget": budget,
|
||||
"max_output_tokens": max_output_tokens,
|
||||
"safety_reserve": reserve,
|
||||
"observed_margin": None,
|
||||
"window_verified": bool(window_verified),
|
||||
"detail": "",
|
||||
}
|
||||
if type(prompt) is not int or prompt <= 0:
|
||||
record["detail"] = ("The server reported no prompt token count, so nothing "
|
||||
"confirms the whole prompt was read." + verified_note)
|
||||
return record
|
||||
|
||||
record["server_prompt_tokens"] = prompt
|
||||
record["difference"] = prompt - estimate
|
||||
record["observed_margin"] = budget - max_output_tokens - prompt
|
||||
if prompt + reserve < estimate:
|
||||
record["status"] = TRUNCATION_SUSPECTED
|
||||
record["detail"] = (
|
||||
f"The server read {prompt:,} prompt tokens of the {estimate:,} sent, a "
|
||||
f"shortfall larger than the {reserve:,}-token safety reserve. A server "
|
||||
"that cuts an over-window prompt reports exactly this, and what it cuts "
|
||||
"is the start: the narrator's rules and the canon." + verified_note)
|
||||
elif prompt + max_output_tokens > budget:
|
||||
record["status"] = EXCEEDED
|
||||
record["detail"] = (
|
||||
f"The server counted {prompt:,} prompt tokens; with {max_output_tokens:,} "
|
||||
f"for the reply that is more than the {budget:,}-token window the prompt "
|
||||
"was built for, so the reply may have been cut short." + verified_note)
|
||||
else:
|
||||
record["status"] = FITS
|
||||
record["detail"] = (
|
||||
f"The server read {prompt:,} prompt tokens, leaving "
|
||||
f"{record['observed_margin']:,} beside the reply." + verified_note)
|
||||
return record
|
||||
|
||||
|
||||
def cache_clear() -> None:
|
||||
"""Forgets what was learned. Called when the endpoint or model changes."""
|
||||
_cache.clear()
|
||||
@@ -223,6 +339,7 @@ def _declared_or(declared: int | None, discovered: Window) -> Window:
|
||||
declared, DECLARED, discovered.model_max,
|
||||
f"{declared:,} tokens, declared in settings — the server was not able "
|
||||
f"to say ({discovered.detail})",
|
||||
reachable=discovered.reachable,
|
||||
)
|
||||
|
||||
|
||||
@@ -266,6 +383,7 @@ async def _ask(endpoint_url: str, model: str) -> Window:
|
||||
return Window(
|
||||
tokens, LOADED, ceiling,
|
||||
f"{tokens:,} tokens, reported by the running model",
|
||||
reachable=True,
|
||||
)
|
||||
return await _declared_window(client, base, model)
|
||||
except (httpx.HTTPError, ValueError, TypeError, KeyError) as exc:
|
||||
@@ -293,6 +411,7 @@ async def _declared_window(client, base: str, model: str) -> Window:
|
||||
return Window(
|
||||
None, UNKNOWN,
|
||||
detail=f"the server did not describe the model (HTTP {resp.status_code})",
|
||||
reachable=True,
|
||||
)
|
||||
body = resp.json() or {}
|
||||
ceiling = _architecture_ceiling(body.get("model_info") or {})
|
||||
@@ -304,14 +423,103 @@ async def _declared_window(client, base: str, model: str) -> Window:
|
||||
"the model sets no num_ctx, so the server will load it at its own "
|
||||
"default — which is 4,096 where there is no VRAM"
|
||||
),
|
||||
reachable=True,
|
||||
)
|
||||
tokens = min(declared, ceiling) if ceiling else declared
|
||||
return Window(
|
||||
tokens, PARAMETERS, ceiling,
|
||||
f"{tokens:,} tokens, from the model's own num_ctx",
|
||||
reachable=True,
|
||||
)
|
||||
|
||||
|
||||
#: v1.1 WP-A1 corrective: loading the configured model so its window can be read.
|
||||
#:
|
||||
#: The first real turn of the A1 evidence found a cold model: `/api/ps` knew
|
||||
#: nothing, `/api/show` found no `num_ctx`, so the window was unverified and the
|
||||
#: prompt was built to the configured 16,384. Ollama loaded the model at its own
|
||||
#: 4,096 default, kept 2,050 of 13,875 tokens and answered 200. That case is
|
||||
#: preventable, because the window becomes readable the moment the model is
|
||||
#: resident. Ollama's native `POST /api/generate` with a model and **no prompt**
|
||||
#: loads the model and generates nothing — measured on Ollama 0.33: HTTP 200,
|
||||
#: `"response": ""`, `"done_reason": "load"`, and `/api/ps` then reported the
|
||||
#: window. The OpenAI-compatible request that followed did not reload it.
|
||||
WARM_PATH = "/api/generate"
|
||||
|
||||
|
||||
async def warm(endpoint_url: str, model: str, *, timeout: float) -> tuple[bool, str]:
|
||||
"""Asks the configured server to load `model`. One request, no story text.
|
||||
|
||||
Held to the same endpoint policy and TLS trust as inference and the probe, and
|
||||
sent to the same host the probe asks. The body names the model and nothing
|
||||
else: no prompt, so nothing is generated, and no `options` or `keep_alive`, so
|
||||
the model loads the way the server would load it for the turn itself.
|
||||
|
||||
Returns `(loaded, detail)`. Every failure is `(False, why)` and never raises:
|
||||
a server that will not load the model on request will fail the turn's own
|
||||
call the ordinary way, which is where that failure belongs.
|
||||
"""
|
||||
reason = endpoints.rejection_reason(endpoint_url)
|
||||
if reason is not None:
|
||||
return False, f"endpoint not allowed — {reason}"
|
||||
base = native_base(endpoint_url)
|
||||
try:
|
||||
async with httpx.AsyncClient(
|
||||
timeout=httpx.Timeout(timeout, connect=CONNECT_TIMEOUT),
|
||||
verify=tlstrust.ssl_context(),
|
||||
) as client:
|
||||
resp = await client.post(f"{base}{WARM_PATH}", json={"model": model})
|
||||
except httpx.HTTPError as exc:
|
||||
log.debug("model warm-up failed for %s: %s", base, exc)
|
||||
return False, f"could not ask the server to load the model ({type(exc).__name__})"
|
||||
if resp.status_code != 200:
|
||||
return False, f"the server did not load the model (HTTP {resp.status_code})"
|
||||
try:
|
||||
body = resp.json() or {}
|
||||
except ValueError:
|
||||
return False, "the server answered the load request with something that was not JSON"
|
||||
return True, f"the server loaded the model ({body.get('done_reason') or 'done'})"
|
||||
|
||||
|
||||
async def ensure_window(endpoint_url: str, model: str, *, declared: int | None = None,
|
||||
warm_timeout: float = 300.0) -> tuple[Window, dict]:
|
||||
"""The window for a turn about to be generated, loading the model once if that is what it takes.
|
||||
|
||||
1. Probe as before.
|
||||
2. If the window is not verified, the server answered, and there is a model to
|
||||
load: one bounded `warm` request.
|
||||
3. If the model loaded, probe again, bypassing the cache that still holds the
|
||||
unverified answer.
|
||||
|
||||
Whatever the second probe says is the answer. There is no retry loop, no
|
||||
guessed window, and no hard-coded 4,096: a window still unverified leaves the
|
||||
configured budget standing, exactly as before, and the turn's accounting
|
||||
still catches a server that cut the prompt.
|
||||
|
||||
Returns the window and a `preflight` record for the turn's provenance.
|
||||
Not used by the context dry run: loading a model is a side effect, and
|
||||
opening a panel should not cause one.
|
||||
"""
|
||||
window = await probe(endpoint_url, model, declared=declared)
|
||||
preflight = {"attempted": False, "loaded": None, "verified_before": window.verified,
|
||||
"verified_after": window.verified, "detail": ""}
|
||||
if window.verified:
|
||||
preflight["detail"] = "the window was already verified"
|
||||
return window, preflight
|
||||
if not (endpoint_url and model):
|
||||
preflight["detail"] = "no endpoint or model configured"
|
||||
return window, preflight
|
||||
if not window.reachable:
|
||||
preflight["detail"] = "the server did not answer, so no model was loaded"
|
||||
return window, preflight
|
||||
loaded, detail = await warm(endpoint_url, model, timeout=warm_timeout)
|
||||
preflight.update(attempted=True, loaded=loaded, detail=detail)
|
||||
if loaded:
|
||||
window = await probe(endpoint_url, model, declared=declared, use_cache=False)
|
||||
preflight["verified_after"] = window.verified
|
||||
return window, preflight
|
||||
|
||||
|
||||
def _num_ctx(parameters) -> int | None:
|
||||
"""Reads `num_ctx` out of the plain-text parameter block Ollama returns."""
|
||||
if not isinstance(parameters, str):
|
||||
|
||||
@@ -35,6 +35,8 @@ creating a second, empty Mara.
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
|
||||
# Field types the schema layer enforces. Kept deliberately small: a narrative
|
||||
# state event carries names, labels and plain values, and nothing here needs a
|
||||
# nested structure a model could hide something inside.
|
||||
@@ -184,11 +186,30 @@ def vocabulary_for_prompt() -> str:
|
||||
Generated from `SPECS` rather than written out beside it, so the model can
|
||||
never be told about an event the application does not implement — the drift
|
||||
that would produce proposals rejected for reasons nobody could see.
|
||||
|
||||
v1.1 WP-A2: each event is shown as the object the model must put in the
|
||||
`events` list, with its required fields, not as `name(field, …)`. The call
|
||||
notation was never the wire format, and a 3B narrator copied it into its
|
||||
prose as `> set_possession(silver-key, "alice")`. An object copied into prose
|
||||
is a proposal the extractor already recognises and removes; a call is not.
|
||||
"""
|
||||
lines = []
|
||||
for name, definition in SPECS.items():
|
||||
fields = list(definition["required"]) + [
|
||||
f"{field}?" for field in definition["optional"]
|
||||
]
|
||||
lines.append(f' {name}({", ".join(fields)}) — {definition["summary"]}')
|
||||
shape = {"type": name}
|
||||
for field, kind in definition["required"].items():
|
||||
shape[field] = _PLACEHOLDER[kind]
|
||||
body = json.dumps(shape, ensure_ascii=False, separators=(",", ":"))
|
||||
line = f" {body} — {definition['summary']}"
|
||||
if definition["optional"]:
|
||||
line += f" (optional: {', '.join(definition['optional'])})"
|
||||
lines.append(line)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
#: What a field of each kind looks like in the prompt's vocabulary. Placeholders,
|
||||
#: never example identifiers, so the vocabulary names nothing a story could copy.
|
||||
#: A list field is shown as a list, so the model is told its shape; every other
|
||||
#: field is an ellipsis. Measured: `"<key>"`-style placeholders with spaced
|
||||
#: separators cost 456 tokens against v1.0.0's 258; this form costs about 380,
|
||||
#: and every line is still the object the model must send.
|
||||
_PLACEHOLDER = {KEY: "…", TEXT: "…", VALUE: "…", LABELS: ["…"]}
|
||||
|
||||
@@ -37,17 +37,17 @@ EMIT_RULE = (
|
||||
"appeared, record it.\n"
|
||||
"\n"
|
||||
"Every value is ABSOLUTE — the new state of things, never a change or a "
|
||||
"difference. Use only these events:\n"
|
||||
"difference. Use only these events, in exactly this shape:\n"
|
||||
f"{events.vocabulary_for_prompt()}\n"
|
||||
"\n"
|
||||
"Identifiers are short lower-case slugs (mara, silver-key, old-abbey) and must "
|
||||
"match the ones already in the state you were shown. Introduce a person, place "
|
||||
"or thing with create_entity before referring to it. If the turn established "
|
||||
"nothing, send an empty events list.\n"
|
||||
"Identifiers are short lower-case slugs and must match the ones already in the "
|
||||
"state you were shown; the example's identifiers are placeholders. Introduce a "
|
||||
"person, place or thing with create_entity before referring to it. If the turn "
|
||||
"established nothing, send an empty events list.\n"
|
||||
"Example:\n"
|
||||
'```state\n'
|
||||
'{"events": [{"type": "set_possession", "item": "silver-key", "owner": "aldric"},'
|
||||
' {"type": "set_current_location", "entity": "aldric", "location": "old-abbey"}]}\n'
|
||||
'{"events": [{"type": "set_possession", "item": "item-1", "owner": "character-1"},'
|
||||
' {"type": "set_current_location", "entity": "character-1", "location": "location-1"}]}\n'
|
||||
'```'
|
||||
)
|
||||
|
||||
@@ -58,6 +58,47 @@ EMIT_REMINDER = (
|
||||
"nothing changed.]"
|
||||
)
|
||||
|
||||
# v1.1 WP-A2: the length hint's own words, named once. `builder.length_hint`
|
||||
# builds the hint from these, and the extractor recognises an echo of it by
|
||||
# them, so the two cannot drift apart.
|
||||
LENGTH_HINT_OPENING = "[Hard limit:"
|
||||
LENGTH_HINT_TAIL = "Finish the narration and append the state block well inside the limit."
|
||||
#: The application's wording inside a hint. A 3B narrator reworded the front
|
||||
#: ("your next turn") and the end ("This story ends here."), and kept one or the
|
||||
#: other of these every time.
|
||||
_LENGTH_HINT_PHRASE_RE = re.compile(
|
||||
r"append the state block|turn must not exceed \d+ words", re.IGNORECASE
|
||||
)
|
||||
|
||||
#: v1.1 WP-A2: the rules that remove protocol a narrator copied, named so the
|
||||
#: replay tool and the report can say which removed what.
|
||||
RULE_EVENT_CALL = "event_call_line"
|
||||
RULE_LENGTH_HINT = "echoed_length_hint"
|
||||
RULE_SCENE_LINE = "rendered_scene_line"
|
||||
RULE_EMPTY_FENCE = "empty_dangling_fence"
|
||||
RULE_INSTRUCTION_TAIL = "echoed_instruction_tail"
|
||||
|
||||
#: v1.1 WP-A2 corrective (R5). The sentence `CHAT_CONTINUE_HINT` in
|
||||
#: `providers/openai_compatible.py` carries, which a narrator echoed with the rest
|
||||
#: of the hint reworded around it. Kept as a copy rather than an import, so the
|
||||
#: narrative package does not depend on the provider; a test pins that the
|
||||
#: hint still contains it.
|
||||
CONTINUE_HINT_PHRASE = "Output only story text"
|
||||
|
||||
# R1. A whole line opening with a call to an event this protocol has. The names
|
||||
# come from the vocabulary, so a call-shaped line naming anything else — a
|
||||
# character's `open_door(north)` — is not matched.
|
||||
_EVENT_CALL_LINE_RE = re.compile(
|
||||
r"^[ \t]*(?:>[ \t]*)?(?:"
|
||||
+ "|".join(re.escape(name) for name in events.SPECS)
|
||||
+ r")[ \t]*\(",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# R3. The renderer's scene line carries its location this way.
|
||||
_RENDERED_SCENE_LOCATION_RE = re.compile(r"\(at [^()\n]+\)\s*$")
|
||||
# R4. An opener with nothing after it.
|
||||
_EMPTY_FENCE_LINE_RE = re.compile(r"```(?:json)?[ \t]*", re.IGNORECASE)
|
||||
|
||||
# Three patterns, and the difference between them is the whole of this module's
|
||||
# safety. A story is allowed to contain code, and taking a code block out of
|
||||
# someone's prose is a worse failure than leaving a stray proposal in it.
|
||||
@@ -128,11 +169,42 @@ def _is_echoed_instruction(inner: str) -> bool:
|
||||
# opening words only, because the echo is often cut off before it ends.
|
||||
if low.lstrip().startswith("continue the story directly"):
|
||||
return True
|
||||
# v1.1 WP-A2 corrective (R5): the same hint, reworded at the front. The M11
|
||||
# closeout-era identity re-run stored "[You don't need to continue; … Continue
|
||||
# the story here, directly. Output only story text.]" as the last line of a
|
||||
# reply, and because nothing recognised it, nothing above it was trailing.
|
||||
if CONTINUE_HINT_PHRASE.lower() in low:
|
||||
return True
|
||||
# v1.1 WP-A2 (R2): the length hint, which names "state block" but not
|
||||
# "events list", so it passed every check above.
|
||||
if _is_length_hint(inner):
|
||||
return True
|
||||
# The reminder names both; prose about the protocol rarely names either the
|
||||
# way the instruction does, and effectively never both.
|
||||
return "state block" in low and "events list" in low
|
||||
|
||||
|
||||
def _opens_like_length_hint(inner: str) -> bool:
|
||||
"""R5. The bracket opens with the length hint's own `Hard limit:`, whatever follows.
|
||||
|
||||
Never enough on its own: an in-world "[Hard limit: forty days]" opens the same
|
||||
way. `_clean` takes it only directly above an echoed instruction it has already
|
||||
removed from the end of the same reply.
|
||||
"""
|
||||
return inner.lstrip().lower().startswith(LENGTH_HINT_OPENING[1:].lower())
|
||||
|
||||
|
||||
def _is_length_hint(inner: str) -> bool:
|
||||
"""Whether a bracket's contents are `builder.length_hint`, however reworded.
|
||||
|
||||
It must open the way the hint opens *and* carry the hint's own wording. An
|
||||
in-world "Hard limit: forty days" has the opening and none of the wording.
|
||||
"""
|
||||
opening = LENGTH_HINT_OPENING[1:].lower()
|
||||
return (inner.lstrip().lower().startswith(opening)
|
||||
and bool(_LENGTH_HINT_PHRASE_RE.search(inner)))
|
||||
|
||||
|
||||
# A heading the model writes above a block it did not fence: `State`, sometimes
|
||||
# as `State:`, `**State**` or `### State`. It is removed only in two places:
|
||||
# directly above a proposal that is removed, and as the last line of the reply.
|
||||
@@ -145,7 +217,7 @@ _LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
|
||||
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
|
||||
|
||||
|
||||
def _clean(prose: str) -> str:
|
||||
def _clean(prose: str, *, after_block: bool = False) -> str:
|
||||
"""Removes protocol the block extraction could not, and nothing else.
|
||||
|
||||
Found by the M5 realistic-context run (§12), which is the failure class
|
||||
@@ -160,9 +232,23 @@ def _clean(prose: str) -> str:
|
||||
story after it. Stored text is replayed as history, so every leak also
|
||||
showed the next prompt a second, older account of the state, which is what
|
||||
M5 review Finding 4 removed from replayed history.
|
||||
|
||||
v1.1 WP-A2 added four shapes, from the M11 closeout's identity run and the
|
||||
v1 corpus, each anchored to something the application owns rather than to
|
||||
what prose looks like: a line opening with a vocabulary call (R1), the
|
||||
length hint echoed at the end (R2), the renderer's scene line left last
|
||||
(R3), and an empty fence opener left last (R4). `after_block` says a
|
||||
proposal block was already taken out of this reply, which is what lets R3
|
||||
remove a bare scene line that sat above it.
|
||||
"""
|
||||
cleaned, _found = _inline_proposals(prose)
|
||||
cleaned, calls_removed = _strip_event_call_lines(prose)
|
||||
cleaned, _found = _inline_proposals(cleaned)
|
||||
cleaned = _strip_echoed_state(cleaned)
|
||||
protocol_cut = after_block or calls_removed
|
||||
# R5: set once an echoed instruction bracket has come off the end. Only then
|
||||
# may a bracket that merely opens the way the length hint opens be taken as
|
||||
# part of the same echoed tail.
|
||||
instruction_cut = False
|
||||
# The end of the reply is cut until nothing more comes off, because one kind
|
||||
# of leftover can hide another. In a real reply, a `State` heading sat above
|
||||
# a block the model never finished, and a parroted reminder sat above an
|
||||
@@ -171,7 +257,12 @@ def _clean(prose: str) -> str:
|
||||
before = cleaned
|
||||
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
|
||||
bracket = pattern.search(cleaned)
|
||||
if bracket is not None and _is_echoed_instruction(bracket.group(1)):
|
||||
if bracket is None:
|
||||
continue
|
||||
if _is_echoed_instruction(bracket.group(1)):
|
||||
cleaned = cleaned[: bracket.start()]
|
||||
instruction_cut = True
|
||||
elif instruction_cut and _opens_like_length_hint(bracket.group(1)):
|
||||
cleaned = cleaned[: bracket.start()]
|
||||
cleaned = _DANGLING_STATE_RE.sub("", cleaned)
|
||||
dangling = _DANGLING_JSON_RE.search(cleaned)
|
||||
@@ -182,10 +273,93 @@ def _clean(prose: str) -> str:
|
||||
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
|
||||
# A bare quote marker, the start of a quoted block that never came.
|
||||
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
|
||||
cleaned = _strip_empty_dangling_fence(cleaned)
|
||||
if cleaned.rstrip() != before.rstrip():
|
||||
protocol_cut = True
|
||||
cleaned = _strip_trailing_scene_line(cleaned, protocol_cut)
|
||||
if cleaned == before:
|
||||
return cleaned.strip()
|
||||
|
||||
|
||||
def _strip_event_call_lines(text: str) -> tuple[str, bool]:
|
||||
"""R1. Removes whole lines that open with a call to a vocabulary event.
|
||||
|
||||
A line inside a fenced code block is the story's own code and is never
|
||||
examined. Returns the text and whether anything was removed.
|
||||
"""
|
||||
kept: list[str] = []
|
||||
in_fence = False
|
||||
removed = False
|
||||
for line in text.split("\n"):
|
||||
if line.lstrip().startswith("```"):
|
||||
in_fence = not in_fence
|
||||
kept.append(line)
|
||||
continue
|
||||
if not in_fence and _EVENT_CALL_LINE_RE.match(line):
|
||||
removed = True
|
||||
continue
|
||||
kept.append(line)
|
||||
if not removed:
|
||||
return text, False
|
||||
return re.sub(r"\n{3,}", "\n\n", "\n".join(kept)), True
|
||||
|
||||
|
||||
def _strip_empty_dangling_fence(text: str) -> str:
|
||||
"""R4. A ```` ```json ```` or ```` ``` ```` opener as the last line, with nothing after it.
|
||||
|
||||
Only an *opener*: the fence lines are counted, and an even count means the
|
||||
last one closes a story's own code block, which stays.
|
||||
"""
|
||||
lines = text.rstrip().split("\n")
|
||||
if len(lines) < 2 or not _EMPTY_FENCE_LINE_RE.fullmatch(lines[-1].strip()):
|
||||
return text
|
||||
fences = sum(1 for line in lines if line.lstrip().startswith("```"))
|
||||
if fences % 2 == 0:
|
||||
return text
|
||||
return "\n".join(lines[:-1]).rstrip()
|
||||
|
||||
|
||||
def _strip_trailing_scene_line(text: str, protocol_cut: bool) -> str:
|
||||
"""R3. The renderer's scene line, left as the last line of the reply.
|
||||
|
||||
Taken when it carries the renderer's own `(at <location>)`, or when protocol
|
||||
was already cut from this reply, which makes a bare scene line part of the
|
||||
same pasted tail. A final screenplay-style "Scene: …" line in a reply with
|
||||
no protocol in it stays, and so does any scene line with story after it.
|
||||
"""
|
||||
lines = text.rstrip().split("\n")
|
||||
if len(lines) < 2:
|
||||
return text
|
||||
last = lines[-1].strip()
|
||||
if not last.startswith(render.HEADING_SCENE + " "):
|
||||
return text
|
||||
if not (_RENDERED_SCENE_LOCATION_RE.search(last) or protocol_cut):
|
||||
return text
|
||||
return "\n".join(lines[:-1]).rstrip()
|
||||
|
||||
|
||||
def explain_removed_line(line: str) -> str | None:
|
||||
"""Which v1.1 rule removes a line of this shape, for the replay report.
|
||||
|
||||
None means no v1.1 rule explains it, which the replay treats as a failure.
|
||||
"""
|
||||
stripped = line.strip()
|
||||
if _EVENT_CALL_LINE_RE.match(line):
|
||||
return RULE_EVENT_CALL
|
||||
if stripped.startswith("["):
|
||||
inner = stripped[1:]
|
||||
inner = inner[:-1] if inner.endswith("]") else inner
|
||||
if _is_length_hint(inner):
|
||||
return RULE_LENGTH_HINT
|
||||
if _is_echoed_instruction(inner) or _opens_like_length_hint(inner):
|
||||
return RULE_INSTRUCTION_TAIL
|
||||
if stripped.startswith(render.HEADING_SCENE + " "):
|
||||
return RULE_SCENE_LINE
|
||||
if _EMPTY_FENCE_LINE_RE.fullmatch(stripped):
|
||||
return RULE_EMPTY_FENCE
|
||||
return None
|
||||
|
||||
|
||||
def _is_state_heading(line: str) -> bool:
|
||||
return bool(_STATE_HEADING_RE.match(line))
|
||||
|
||||
@@ -451,7 +625,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
if matches:
|
||||
match = matches[-1]
|
||||
raw = match.group(1).strip()
|
||||
prose = _clean(text[: match.start()] + text[match.end():])
|
||||
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
||||
return prose, _tolerant_load(raw), raw
|
||||
|
||||
# A `json` or unlabelled fence is ours only when its contents are this
|
||||
@@ -465,7 +639,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
raw = match.group(1).strip()
|
||||
parsed = _tolerant_load(raw)
|
||||
if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
|
||||
prose = _clean(text[: match.start()] + text[match.end():])
|
||||
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
||||
return prose, parsed, raw
|
||||
|
||||
match = _TRAILING_RE.search(text)
|
||||
@@ -473,7 +647,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
raw = match.group(1)
|
||||
parsed = _tolerant_load(raw)
|
||||
if _looks_like_proposal(parsed):
|
||||
return _clean(text[: match.start()]), parsed, raw
|
||||
return _clean(text[: match.start()], after_block=True), parsed, raw
|
||||
|
||||
# An unfenced proposal on its own lines but not at the end: quoted, or
|
||||
# followed by more story. The last one is the turn's proposal, as with
|
||||
@@ -481,7 +655,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
without, found = _inline_proposals(text)
|
||||
if found:
|
||||
parsed, raw = found[-1]
|
||||
return _clean(without), parsed, raw
|
||||
return _clean(without, after_block=True), parsed, raw
|
||||
|
||||
# No block at all — but the reply may still carry protocol the model wrote
|
||||
# as prose, or a fence it never closed.
|
||||
|
||||
@@ -26,6 +26,10 @@ EMBED_READ_TIMEOUT = 60.0
|
||||
|
||||
|
||||
|
||||
#: v1.1 WP-A1: ask a stream to report its token usage. Without it Ollama sends
|
||||
#: none, and a prompt the server cut cannot be told from one it read whole.
|
||||
STREAM_OPTIONS = {"include_usage": True}
|
||||
|
||||
# Completion endpoints have no roles, so a chat has to be flattened into one
|
||||
# labeled transcript that ends on "Assistant:" for the model to continue.
|
||||
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
|
||||
@@ -84,10 +88,14 @@ class OpenAICompatibleProvider(Provider):
|
||||
def _record_usage(self, payload: dict) -> None:
|
||||
"""Records the endpoint's own token accounting, if it reported any.
|
||||
|
||||
OpenRouter now always reports usage, and `usage: {include: true}` and
|
||||
`stream_options` are deprecated and do nothing. In a stream the usage
|
||||
arrives on a final chunk that carries no choices, which is why this is
|
||||
read separately from the text extraction.
|
||||
In a stream the usage arrives on a final chunk that carries no choices,
|
||||
which is why this is read separately from the text extraction.
|
||||
|
||||
v1.1 WP-A1: Ollama sends that chunk only when asked. Measured on Ollama
|
||||
0.33: a stream with no `stream_options` carried no usage at all, and not
|
||||
one of the 514 AI turns in the v1 evidence had a count stored. Every
|
||||
streaming body therefore sets `stream_options.include_usage`
|
||||
(`STREAM_OPTIONS`), and the turn compares the count with what it sent.
|
||||
"""
|
||||
usage = payload.get("usage")
|
||||
if isinstance(usage, dict) and usage:
|
||||
@@ -102,6 +110,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
else:
|
||||
url = f"{self.base_url}/chat/completions"
|
||||
@@ -114,6 +123,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
return url, body
|
||||
|
||||
@@ -183,6 +193,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
else:
|
||||
url = f"{self.base_url}/chat/completions"
|
||||
@@ -192,6 +203,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
async for event in self._stream(url, body):
|
||||
yield event
|
||||
|
||||
@@ -6,6 +6,7 @@ lock guards one set only while one module owns it. And a test that replaces
|
||||
`OpenAICompatibleProvider` or `generate_turn` patches this module, which every
|
||||
caller reads through.
|
||||
"""
|
||||
import logging
|
||||
import threading
|
||||
|
||||
from fastapi import Depends, HTTPException, Request
|
||||
@@ -28,6 +29,8 @@ from .deps import CurrentUser, current_adventure, router
|
||||
from .nodes import _move_to_after, next_depth
|
||||
from .paging import annotate_takes
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def world_delta_of(snapshot: dict | None) -> dict | None:
|
||||
"""Returns the bulk-read slice of a context snapshot, for `Action.world_delta`.
|
||||
@@ -212,8 +215,18 @@ async def _generate_turn(
|
||||
# network calls — and cached per endpoint and model, so it costs one short
|
||||
# request per session rather than one per turn. An unverified window does
|
||||
# not block the turn; it is recorded as unverified in the snapshot below.
|
||||
window = await contextwindow.probe(settings.endpoint_url, settings.model,
|
||||
declared=settings.context_window_override)
|
||||
#
|
||||
# v1.1 WP-A1 corrective: a model that is not resident cannot report its window,
|
||||
# and a turn built to the configured budget against it was silently cut in the
|
||||
# A1 evidence (13,875 tokens sent, 2,050 read). So an unverified window gets
|
||||
# one bounded attempt to load the model, and one more probe, before the
|
||||
# prompt is assembled. No story text is generated by it and nothing is
|
||||
# written. A window still unverified afterwards changes nothing below.
|
||||
window, preflight = await contextwindow.ensure_window(
|
||||
settings.endpoint_url, settings.model,
|
||||
declared=settings.context_window_override,
|
||||
warm_timeout=float(settings.model_timeout_seconds or 300),
|
||||
)
|
||||
try:
|
||||
system_text, story_text, snapshot = build_context(
|
||||
adventure,
|
||||
@@ -232,6 +245,9 @@ async def _generate_turn(
|
||||
yield turn_error(str(exc))
|
||||
return
|
||||
|
||||
if isinstance(snapshot.get("window"), dict):
|
||||
snapshot["window"]["preflight"] = preflight
|
||||
|
||||
parts = PromptParts(system=system_text, story=story_text)
|
||||
|
||||
provider = OpenAICompatibleProvider(
|
||||
@@ -318,6 +334,24 @@ async def _generate_turn(
|
||||
# prompt came from cache rather than being billed in full. This is recorded
|
||||
# per attempt, next to the prompt it priced.
|
||||
snapshot["usage"] = provider.last_usage
|
||||
# v1.1 WP-A1: what the server says it read, against what was sent. Recorded
|
||||
# and shown, never acted on: the narration has already streamed to the
|
||||
# reader, and discarding an accepted turn over an accounting discrepancy
|
||||
# would lose story to hide a problem. A server that cut the prompt answers
|
||||
# 200 either way, so this record is the only place the cut is visible.
|
||||
tokens = snapshot.get("tokens") or {}
|
||||
accounting = contextwindow.classify_usage(
|
||||
provider.last_usage,
|
||||
estimate=tokens.get("estimate") or tokens.get("total") or 0,
|
||||
budget=tokens.get("budget") or settings.context_token_budget,
|
||||
max_output_tokens=settings.max_output_tokens,
|
||||
window_verified=bool((snapshot.get("window") or {}).get("verified")),
|
||||
)
|
||||
snapshot["accounting"] = accounting
|
||||
if accounting["status"] in (contextwindow.EXCEEDED,
|
||||
contextwindow.TRUNCATION_SUSPECTED):
|
||||
log.warning("turn accounting for adventure %s: %s — %s",
|
||||
adventure.id, accounting["status"], accounting["detail"])
|
||||
|
||||
reasoning = "".join(reasoning_chunks).strip() or None
|
||||
ai_action = models.Action(
|
||||
@@ -381,7 +415,8 @@ async def _generate_turn(
|
||||
db.commit()
|
||||
db.refresh(ai_action)
|
||||
yield _SAVED
|
||||
yield sse({"type": "done", "action": action_json(ai_action, db)})
|
||||
yield sse({"type": "done", "action": action_json(ai_action, db),
|
||||
"accounting": accounting})
|
||||
# Phase 6: schedule summarization and embedding without waiting for them.
|
||||
# The task opens its own database session.
|
||||
memorybank.schedule_post_turn(adventure)
|
||||
|
||||
@@ -246,11 +246,28 @@ def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
|
||||
assert report_before["history"]["floor_depth"] is not None, (
|
||||
"this fixture is meant to be over budget; trimming never engaged")
|
||||
|
||||
_play_one_more(db, adventure, 60)
|
||||
_, after, report_after = _builder.build_context(adventure, settings)
|
||||
# v1.1 WP-A1: the fixture used to be positioned so that the very next turn
|
||||
# held the floor. The safety reserve takes 256 tokens of this 2,048 budget,
|
||||
# the block is now the minimum of two, and the next turn is a step. So walk
|
||||
# forward until a turn holds, requiring every move on the way to be exactly
|
||||
# one block: a window that slides by one action every turn fails either way.
|
||||
held = None
|
||||
depth = 60
|
||||
for _ in range(4):
|
||||
_play_one_more(db, adventure, depth)
|
||||
depth += 1
|
||||
_, after, report_after = _builder.build_context(adventure, settings)
|
||||
floor_before = report_before["history"]["floor_depth"]
|
||||
floor_after = report_after["history"]["floor_depth"]
|
||||
block = report_after["history"]["trim_block"]
|
||||
assert floor_after - floor_before in (0, block), (floor_before, floor_after, block)
|
||||
if floor_after == floor_before:
|
||||
held = (before, after)
|
||||
break
|
||||
before, report_before = after, report_after
|
||||
|
||||
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"]
|
||||
assert _shared_prefix(before, after) > 0.85
|
||||
assert held is not None, "the floor never held across a turn"
|
||||
assert _shared_prefix(*held) > 0.85
|
||||
|
||||
|
||||
def test_without_a_stable_floor_the_prefix_collapses(saturated):
|
||||
|
||||
@@ -0,0 +1,348 @@
|
||||
"""v1.1 WP-A1 corrective: a cold model is loaded, not guessed about.
|
||||
|
||||
A1's accounting caught a real cold-model turn: `/api/ps` knew nothing because the
|
||||
model was not resident, `/api/show` found no `num_ctx`, the window was therefore
|
||||
unverified, and the prompt was built to the configured 16,384. Ollama loaded the
|
||||
model at its own 4,096 default, read 2,050 of the 13,875 tokens and answered 200.
|
||||
|
||||
Detection was right. The case is also preventable: once the model is loaded its
|
||||
window is readable. So before an unverified turn is assembled, the application
|
||||
asks the configured server, once, to load the model (`POST /api/generate` with a
|
||||
model and no prompt, which Ollama answers with `"done_reason": "load"` and no
|
||||
text), probes again, and builds the turn to whatever that probe says. A window
|
||||
still unverified afterwards changes nothing: the configured budget stands and
|
||||
the post-response accounting still watches for a cut prompt.
|
||||
|
||||
python -m pytest tests/test_v11_cold_window.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy import text
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, contextwindow, limits, models
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.providers.base import ProviderError
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
ENDPOINT = "http://127.0.0.1:11434/v1"
|
||||
MODEL = "qwen2.5:3b-instruct"
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_window_cache():
|
||||
contextwindow.cache_clear()
|
||||
yield
|
||||
contextwindow.cache_clear()
|
||||
|
||||
|
||||
class ColdOllama:
|
||||
"""The shapes a real Ollama 0.33 returned, with a model that starts cold.
|
||||
|
||||
`/api/ps` lists only loaded models. `/api/show` carries no `num_ctx`.
|
||||
`/api/generate` with no prompt loads the model at `load_window`, exactly as
|
||||
the real server answered: HTTP 200, `"response": ""`, `"done_reason": "load"`.
|
||||
"""
|
||||
|
||||
def __init__(self, *, loaded=None, load_window=4096, generate_status=200,
|
||||
report_after_load=True):
|
||||
self.loaded = dict(loaded or {})
|
||||
self.load_window = load_window
|
||||
self.generate_status = generate_status
|
||||
self.report_after_load = report_after_load
|
||||
self.requests: list[tuple[str, str, dict | None]] = []
|
||||
|
||||
def handler(self, request: httpx.Request) -> httpx.Response:
|
||||
body = None
|
||||
if request.content:
|
||||
import json
|
||||
body = json.loads(request.content)
|
||||
self.requests.append((request.method, str(request.url), body))
|
||||
path = request.url.path
|
||||
if path == "/api/ps":
|
||||
return httpx.Response(200, json={"models": [
|
||||
{"name": name, "model": name, "context_length": tokens}
|
||||
for name, tokens in self.loaded.items()
|
||||
]})
|
||||
if path == "/api/show":
|
||||
return httpx.Response(200, json={
|
||||
"model_info": {"qwen2.context_length": 32768}, "parameters": ""})
|
||||
if path == "/api/generate":
|
||||
if self.generate_status != 200:
|
||||
return httpx.Response(self.generate_status, json={"error": "model not found"})
|
||||
if self.report_after_load:
|
||||
self.loaded[body["model"]] = self.load_window
|
||||
return httpx.Response(200, json={
|
||||
"model": body["model"], "response": "", "done": True, "done_reason": "load"})
|
||||
return httpx.Response(404)
|
||||
|
||||
def paths(self):
|
||||
return [httpx.URL(url).path for _method, url, _body in self.requests]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def server(monkeypatch):
|
||||
def install(fake: ColdOllama):
|
||||
original = httpx.AsyncClient
|
||||
|
||||
def build(*args, **kwargs):
|
||||
kwargs.pop("verify", None)
|
||||
return original(*args, transport=httpx.MockTransport(fake.handler), **kwargs)
|
||||
|
||||
monkeypatch.setattr(contextwindow.httpx, "AsyncClient", build)
|
||||
return fake
|
||||
|
||||
return install
|
||||
|
||||
|
||||
# ------------------------------------------------------------ ensure_window
|
||||
|
||||
def test_a_cold_model_is_loaded_once_and_its_window_verified(server):
|
||||
fake = server(ColdOllama(loaded={}, load_window=4096))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert (window.tokens, window.source, window.verified) == (4096, contextwindow.LOADED, True)
|
||||
assert preflight == {"attempted": True, "loaded": True, "verified_before": False,
|
||||
"verified_after": True,
|
||||
"detail": "the server loaded the model (load)"}
|
||||
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate", "/api/ps"]
|
||||
# One load request, naming the model and nothing else: no prompt, so no text.
|
||||
warms = [body for _m, url, body in fake.requests if url.endswith("/api/generate")]
|
||||
assert warms == [{"model": MODEL}]
|
||||
|
||||
|
||||
def test_an_already_loaded_model_is_not_warmed(server):
|
||||
fake = server(ColdOllama(loaded={MODEL: 16384}))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert window.verified and window.tokens == 16384
|
||||
assert preflight["attempted"] is False
|
||||
assert "/api/generate" not in fake.paths()
|
||||
|
||||
|
||||
def test_a_model_that_loads_but_still_cannot_be_read_stays_unverified(server):
|
||||
"""A server that loads the model but whose `/api/ps` still cannot say. The
|
||||
existing unknown path stands: no guessed window, the configured budget kept."""
|
||||
fake = server(ColdOllama(loaded={}, report_after_load=False))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert not window.verified and window.tokens is None
|
||||
assert preflight["attempted"] is True and preflight["loaded"] is True
|
||||
assert preflight["verified_after"] is False
|
||||
assert fake.paths().count("/api/generate") == 1
|
||||
assert contextwindow.effective_budget(16384, window) == 16384
|
||||
|
||||
|
||||
@pytest.mark.parametrize("status", [404, 500])
|
||||
def test_a_failed_load_is_recorded_and_leaves_the_window_unverified(server, status):
|
||||
fake = server(ColdOllama(loaded={}, generate_status=status))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert not window.verified
|
||||
assert preflight["attempted"] is True and preflight["loaded"] is False
|
||||
assert f"HTTP {status}" in preflight["detail"]
|
||||
# Bounded: one attempt, and no second probe after a failed load.
|
||||
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate"]
|
||||
|
||||
|
||||
def test_a_declared_window_does_not_stop_the_server_being_asked(server):
|
||||
"""A declaration fills a hole the server leaves. Loading the model can close
|
||||
the hole, and a verified answer always wins over a declaration."""
|
||||
server(ColdOllama(loaded={}, load_window=4096))
|
||||
window, _preflight = asyncio.run(
|
||||
contextwindow.ensure_window(ENDPOINT, MODEL, declared=8192))
|
||||
assert (window.tokens, window.source) == (4096, contextwindow.LOADED)
|
||||
|
||||
|
||||
def test_an_unreachable_server_is_not_asked_to_load_anything():
|
||||
"""Nothing listens here. No load is attempted against a server that did not
|
||||
answer the probe, so an offline turn costs no second timeout."""
|
||||
window, preflight = asyncio.run(
|
||||
contextwindow.ensure_window("http://127.0.0.1:1/v1", MODEL))
|
||||
assert not window.verified
|
||||
assert preflight["attempted"] is False
|
||||
assert "did not answer" in preflight["detail"]
|
||||
|
||||
|
||||
def test_the_load_request_obeys_the_endpoint_policy():
|
||||
"""ADR 011. No transport is installed: a load that ignored the policy would
|
||||
try to reach a public address for real."""
|
||||
for url in ("https://api.openai.com/v1", "http://8.8.8.8:11434/v1"):
|
||||
loaded, detail = asyncio.run(contextwindow.warm(url, MODEL, timeout=2))
|
||||
assert loaded is False
|
||||
assert "not allowed" in detail
|
||||
|
||||
|
||||
def test_the_load_request_goes_only_to_the_configured_host(server):
|
||||
fake = server(ColdOllama(loaded={}))
|
||||
asyncio.run(contextwindow.ensure_window("http://192.168.0.50:11434/v1", MODEL))
|
||||
hosts = {httpx.URL(url).host for _m, url, _b in fake.requests}
|
||||
ports = {httpx.URL(url).port for _m, url, _b in fake.requests}
|
||||
assert hosts == {"192.168.0.50"} and ports == {11434}
|
||||
|
||||
|
||||
# ------------------------------------------------------------- end to end
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="v11cold@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
|
||||
context_token_budget=16384, max_output_tokens=500,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Cold",
|
||||
campaign_canon={"rules": ["The sealed crypt is named CANON-SENTINEL-COLD-2050."]},
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(adventure_id=adventure.id, type="start", text="Rain."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _long_story(adv_id, turns=120):
|
||||
from app import tree
|
||||
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
for i in range(turns):
|
||||
for kind, body in (
|
||||
("do", f"I search the {i}th chamber of the undercroft."),
|
||||
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
|
||||
):
|
||||
action = models.Action(adventure_id=adv_id, type=kind, text=body)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
tree.place_action(db, adventure, action)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _counts():
|
||||
with SessionLocal() as db:
|
||||
return {table: db.execute(text(f"SELECT COUNT(*) FROM {table}")).scalar()
|
||||
for table in ("actions", "state_events", "state_proposals", "memories",
|
||||
"summaries")}
|
||||
|
||||
|
||||
def _latest_ai(adv_id):
|
||||
with SessionLocal() as db:
|
||||
return (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
|
||||
|
||||
def test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports(client, server):
|
||||
"""The observed failure, prevented. Without the load this turn would be built
|
||||
to the configured 16,384 against a 4,096 server."""
|
||||
_long_story(client.adv_id)
|
||||
fake = server(ColdOllama(loaded={}, load_window=4096))
|
||||
ScriptedProvider.replies = ["The seal holds."]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "look at the seal"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
assert snapshot["window"]["verified"] is True
|
||||
assert snapshot["tokens"]["budget"] == 4096
|
||||
assert snapshot["window"]["preflight"]["attempted"] is True
|
||||
assert snapshot["window"]["preflight"]["verified_after"] is True
|
||||
system, story = ScriptedProvider.prompts[-1]
|
||||
sent = builder.count_tokens(system) + builder.count_tokens(story)
|
||||
assert sent + snapshot["tokens"]["transport"] + 500 + 256 <= 4096
|
||||
assert "CANON-SENTINEL-COLD-2050" in system
|
||||
assert fake.paths().count("/api/generate") == 1
|
||||
|
||||
|
||||
def test_without_the_load_the_same_cold_turn_would_have_been_built_too_large(client, server,
|
||||
monkeypatch):
|
||||
"""The negative control: v1.0.0 and the first A1 tree probed only."""
|
||||
_long_story(client.adv_id)
|
||||
server(ColdOllama(loaded={}, load_window=4096))
|
||||
|
||||
async def probe_only(endpoint_url, model, *, declared=None, warm_timeout=300.0):
|
||||
window = await contextwindow.probe(endpoint_url, model, declared=declared)
|
||||
return window, {"attempted": False}
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "ensure_window", probe_only)
|
||||
ScriptedProvider.replies = ["The seal holds."]
|
||||
client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "look at the seal"})
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
assert snapshot["window"]["verified"] is False
|
||||
assert snapshot["tokens"]["budget"] == 16384
|
||||
system, story = ScriptedProvider.prompts[-1]
|
||||
assert builder.count_tokens(system) + builder.count_tokens(story) > 4096 * 2
|
||||
|
||||
|
||||
def test_the_load_itself_writes_nothing(client, server):
|
||||
"""No action, narration, state event, proposal, memory or summary comes from
|
||||
the preflight: it is a request to the server and nothing else."""
|
||||
fake = server(ColdOllama(loaded={}))
|
||||
before = _counts()
|
||||
asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert _counts() == before
|
||||
assert fake.paths().count("/api/generate") == 1
|
||||
|
||||
|
||||
def test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe(client, server):
|
||||
"""The ordinary failure semantics: the error is reported, no narration is
|
||||
accepted, and nothing about the state changes."""
|
||||
server(ColdOllama(loaded={}, generate_status=404))
|
||||
before = _counts()
|
||||
ScriptedProvider.replies = [ProviderError("Endpoint or model not found (HTTP 404).")]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "open the door"})
|
||||
assert response.status_code == 200
|
||||
assert '"type": "error"' in response.text or '"error"' in response.text
|
||||
after = _counts()
|
||||
assert after["state_events"] == before["state_events"]
|
||||
assert after["state_proposals"] == before["state_proposals"]
|
||||
with SessionLocal() as db:
|
||||
assert db.query(models.Action).filter_by(adventure_id=client.adv_id,
|
||||
type="ai").count() == 0
|
||||
|
||||
|
||||
def test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer(client, server):
|
||||
server(ColdOllama(loaded={}, generate_status=500))
|
||||
ScriptedProvider.replies = ["The door opens."]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "open the door"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
assert snapshot["window"]["verified"] is False
|
||||
assert snapshot["window"]["preflight"]["loaded"] is False
|
||||
assert snapshot["tokens"]["budget"] == 16384
|
||||
assert snapshot["accounting"]["status"] == contextwindow.UNKNOWN
|
||||
|
||||
|
||||
def test_the_context_dry_run_never_loads_a_model(client, server):
|
||||
fake = server(ColdOllama(loaded={}))
|
||||
response = client.get(f"/api/adventures/{client.adv_id}/context")
|
||||
assert response.status_code == 200
|
||||
assert "/api/generate" not in fake.paths()
|
||||
@@ -0,0 +1,495 @@
|
||||
"""v1.1 WP-A1: a deliberate safety reserve, and a turn the server cut is not silent.
|
||||
|
||||
M11 made the verified window a ceiling. It did not make the application's count
|
||||
the server's count. The application counts with `cl100k_base`, the narrator with
|
||||
its own tokenizer, and the v1 evidence left 23-42 real tokens between the largest
|
||||
prompt and the edge of a 16,384 window. Past that edge Ollama does not refuse.
|
||||
Measured against the reference CPU host (Ollama 0.33, a 4,096 window), a
|
||||
6,316-token prompt came back 200 with `prompt_tokens` 2,050: the front of the
|
||||
prompt, which in this design is the narrator's rules and the canon, was gone.
|
||||
|
||||
So the tests below are in three halves.
|
||||
|
||||
**The reserve.** `max(256, ceil(5% of the effective window))`, taken from the
|
||||
budget before any history is chosen, on top of an exact reply allocation.
|
||||
|
||||
**The arithmetic.** The assembled prompt, plus the application text the provider
|
||||
adds to every request, plus the reply allocation, plus the reserve, fits the
|
||||
effective window. Protected context that cannot fit that way fails before the
|
||||
model is called.
|
||||
|
||||
**The accounting.** Where the server reports how many prompt tokens it read, the
|
||||
turn records `fits`, `exceeded` or `truncation_suspected`. Where it reports
|
||||
nothing, the turn says `unknown`, never `fits`. A discrepancy found after the
|
||||
reply is recorded and shown; it never costs the reader an accepted turn.
|
||||
|
||||
python -m pytest tests/test_v11_context_reserve.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, contextwindow, limits, models
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.providers.openai_compatible import CHAT_CONTINUE_HINT, OpenAICompatibleProvider
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
ENDPOINT = "http://127.0.0.1:11434/v1"
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_window_cache():
|
||||
contextwindow.cache_clear()
|
||||
yield
|
||||
contextwindow.cache_clear()
|
||||
|
||||
|
||||
# ------------------------------------------------------------- the reserve
|
||||
|
||||
@pytest.mark.parametrize("window, reserve", [
|
||||
(1024, 256),
|
||||
(4096, 256), # 5% is 204.8, so the floor holds
|
||||
(5120, 256), # exactly 5% is the floor
|
||||
(5121, 257), # 256.05 rounds up
|
||||
(8192, 410), # 409.6 rounds up
|
||||
(16384, 820), # 819.2 rounds up
|
||||
(32768, 1639), # 1638.4 rounds up
|
||||
])
|
||||
def test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up(window, reserve):
|
||||
assert contextwindow.safety_reserve(window) == reserve
|
||||
|
||||
|
||||
def test_the_reserve_is_far_larger_than_the_v1_margin_at_the_evidence_window():
|
||||
"""The v1 evidence left 23-42 tokens at 16,384. 64 tokens of slack was all
|
||||
the arithmetic kept for drift and separators together."""
|
||||
assert contextwindow.safety_reserve(16384) >= 10 * 64
|
||||
|
||||
|
||||
# ---------------------------------------------------------- the arithmetic
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="v11reserve@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="qwen2.5:3b-instruct", endpoint_url=ENDPOINT,
|
||||
embedding_model="", context_token_budget=16384, max_output_tokens=500,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Reserved",
|
||||
campaign_canon={"rules": [
|
||||
"The abbey seal has never been broken.",
|
||||
"The sealed crypt is named CANON-SENTINEL-RESERVE-5120.",
|
||||
]},
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text="Rain over Westhaven."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _long_story(adv_id, turns=120):
|
||||
from app import tree
|
||||
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
for i in range(turns):
|
||||
for kind, text in (
|
||||
("do", f"I search the {i}th chamber of the undercroft."),
|
||||
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
|
||||
):
|
||||
action = models.Action(adventure_id=adv_id, type=kind, text=text)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
tree.place_action(db, adventure, action)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _settings(**changes):
|
||||
with SessionLocal() as db:
|
||||
settings = db.query(models.Settings).first()
|
||||
for key, value in changes.items():
|
||||
setattr(settings, key, value)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _build(client, window):
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
settings = db.query(models.Settings).first()
|
||||
return builder.build_context(adventure, settings, window=window)
|
||||
|
||||
|
||||
def _sent(system, story) -> int:
|
||||
"""What the provider actually sends in chat mode, by the application's count."""
|
||||
return (builder.count_tokens(system) + builder.count_tokens(story)
|
||||
+ builder.count_tokens(CHAT_CONTINUE_HINT))
|
||||
|
||||
|
||||
CONFIGURATIONS = {
|
||||
"verified 4,096": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(4096, contextwindow.LOADED)),
|
||||
"verified 8,192": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(8192, contextwindow.PARAMETERS)),
|
||||
"verified 16,384": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(16384, contextwindow.LOADED)),
|
||||
"declared 6,000": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(6000, contextwindow.DECLARED)),
|
||||
"unverified, configured 12,000": (dict(context_token_budget=12000),
|
||||
contextwindow.UNVERIFIED),
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", list(CONFIGURATIONS))
|
||||
def test_the_prompt_leaves_the_reply_and_the_reserve_free(client, name):
|
||||
"""A1-2, on the assembled text rather than the builder's own arithmetic."""
|
||||
changes, window = CONFIGURATIONS[name]
|
||||
_settings(**changes)
|
||||
_long_story(client.adv_id, turns=120)
|
||||
system, story, report = _build(client, window)
|
||||
tokens = report["tokens"]
|
||||
budget = tokens["budget"]
|
||||
|
||||
assert tokens["safety_reserve"] == contextwindow.safety_reserve(budget)
|
||||
assert tokens["output_reserve"] == 500
|
||||
sent = _sent(system, story)
|
||||
assert sent + tokens["output_reserve"] + tokens["safety_reserve"] <= budget, (
|
||||
name, sent, tokens)
|
||||
# The history is what gave way, not the canon.
|
||||
assert "CANON-SENTINEL-RESERVE-5120" in system
|
||||
assert report["history"]["included"] < report["history"]["total"]
|
||||
|
||||
|
||||
def test_the_report_prices_the_text_the_provider_adds(client):
|
||||
"""The chat hint rides on every request and was never counted."""
|
||||
_, _, report = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
|
||||
tokens = report["tokens"]
|
||||
assert tokens["transport"] >= builder.count_tokens(CHAT_CONTINUE_HINT)
|
||||
assert tokens["estimate"] == tokens["total"] + builder.count_tokens(CHAT_CONTINUE_HINT)
|
||||
|
||||
|
||||
def test_the_reserve_follows_the_effective_window_not_the_setting(client):
|
||||
"""5% of a 4,096 server, not 5% of a 16,384 setting it will never read."""
|
||||
_, _, capped = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
|
||||
_, _, full = _build(client, contextwindow.Window(16384, contextwindow.LOADED))
|
||||
assert capped["tokens"]["safety_reserve"] == 256
|
||||
assert full["tokens"]["safety_reserve"] == 820
|
||||
|
||||
|
||||
def test_protected_context_that_only_fits_without_the_reserve_fails_explicitly(client):
|
||||
"""A1-3 in the builder. Before v1.1 this prompt would have been built.
|
||||
|
||||
The canon is sized so that protected text plus the reply fits a 4,096 window
|
||||
with room to spare, and does not fit once the 256-token reserve is taken.
|
||||
"""
|
||||
small = contextwindow.Window(4096, contextwindow.LOADED)
|
||||
# Measured with a window large enough never to overflow, because repeated
|
||||
# text merges tokens at its seams and cannot be priced by multiplication.
|
||||
roomy = contextwindow.Window(32768, contextwindow.LOADED)
|
||||
rules = None
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
settings = db.query(models.Settings).first()
|
||||
base_rules = list(adventure.campaign_canon["rules"])
|
||||
filler = "The bell tolls once for every name in the ledger."
|
||||
copies = 1
|
||||
while True:
|
||||
candidate = base_rules + [" ".join([filler] * copies)]
|
||||
adventure.campaign_canon = {"rules": candidate}
|
||||
_, _, measured = builder.build_context(adventure, settings, window=roomy)
|
||||
t = measured["tokens"]
|
||||
# What protected context costs at 4,096, without the reserve.
|
||||
without_reserve = t["protected"] + t["transport"] + t["output_reserve"]
|
||||
if without_reserve + 64 >= 4096 - 60:
|
||||
break
|
||||
copies += 1
|
||||
rules = candidate
|
||||
adventure.campaign_canon = {"rules": rules}
|
||||
db.commit()
|
||||
|
||||
# The case this test is about: v1's arithmetic, with its 64-token margin,
|
||||
# would have built this prompt. v1.1's reserve does not fit.
|
||||
assert without_reserve + 64 < 4096
|
||||
assert without_reserve + contextwindow.safety_reserve(4096) >= 4096
|
||||
|
||||
with pytest.raises(builder.ContextOverflow) as caught:
|
||||
builder.build_context(adventure, settings, window=small)
|
||||
message = str(caught.value)
|
||||
assert "safety" in message
|
||||
assert "load the model with a larger window" in message
|
||||
|
||||
|
||||
def test_an_overflowing_turn_never_reaches_the_model(client, monkeypatch):
|
||||
"""A1-3 end to end: the refusal happens before the provider is called."""
|
||||
async def verified(endpoint, model, declared=None, use_cache=True):
|
||||
return contextwindow.Window(1024, contextwindow.LOADED)
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
|
||||
ScriptedProvider.replies = ["This must never be generated."]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "open the crypt"})
|
||||
assert response.status_code == 200
|
||||
assert "safety" in response.text
|
||||
assert ScriptedProvider.calls == 0
|
||||
with SessionLocal() as db:
|
||||
assert db.query(models.Action).filter_by(
|
||||
adventure_id=client.adv_id, type="ai").count() == 0
|
||||
|
||||
|
||||
# ---------------------------------------------------------- the accounting
|
||||
|
||||
def _classify(prompt_tokens=None, *, usage=None, estimate=3500, budget=4096,
|
||||
output=500, verified=True):
|
||||
if usage is None and prompt_tokens is not None:
|
||||
usage = {"prompt_tokens": prompt_tokens, "completion_tokens": 40}
|
||||
return contextwindow.classify_usage(
|
||||
usage, estimate=estimate, budget=budget, max_output_tokens=output,
|
||||
window_verified=verified,
|
||||
)
|
||||
|
||||
|
||||
def test_a_prompt_the_server_read_in_full_fits():
|
||||
# The 13-token chat-template overhead measured against the real server.
|
||||
result = _classify(3513)
|
||||
assert result["status"] == contextwindow.FITS
|
||||
assert result["server_prompt_tokens"] == 3513
|
||||
assert result["difference"] == 13
|
||||
assert result["safety_reserve"] == 256
|
||||
assert result["observed_margin"] == 4096 - 500 - 3513
|
||||
|
||||
|
||||
def test_a_server_that_counts_more_than_the_reserve_allows_is_exceeded():
|
||||
"""The prompt plus the reply allocation no longer fits the window."""
|
||||
result = _classify(3700)
|
||||
assert result["status"] == contextwindow.EXCEEDED
|
||||
assert result["observed_margin"] < 0
|
||||
|
||||
|
||||
def test_a_server_that_read_far_less_than_was_sent_is_suspected_of_truncating():
|
||||
"""The real shape: 6,316 sent, 2,050 read, HTTP 200, no error."""
|
||||
result = _classify(2050, estimate=6316)
|
||||
assert result["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
assert result["difference"] == 2050 - 6316
|
||||
|
||||
|
||||
def test_a_small_undercount_is_tokenizer_drift_not_truncation():
|
||||
"""A tokenizer thriftier than `cl100k_base` reads fewer tokens honestly. Only
|
||||
a shortfall larger than the reserve is called truncation."""
|
||||
assert _classify(3500 - 255)["status"] == contextwindow.FITS
|
||||
assert _classify(3500 - 257)["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
|
||||
|
||||
@pytest.mark.parametrize("usage", [
|
||||
None,
|
||||
{},
|
||||
{"completion_tokens": 40},
|
||||
{"prompt_tokens": 0},
|
||||
{"prompt_tokens": "3500"},
|
||||
{"prompt_tokens": -1},
|
||||
])
|
||||
def test_no_usable_count_is_unknown_never_fits(usage):
|
||||
result = _classify(usage=usage)
|
||||
assert result["status"] == contextwindow.UNKNOWN
|
||||
assert result["server_prompt_tokens"] is None
|
||||
assert result["observed_margin"] is None
|
||||
|
||||
|
||||
def test_the_accounting_says_when_the_window_itself_was_not_verified():
|
||||
result = _classify(3513, verified=False)
|
||||
assert result["status"] == contextwindow.FITS
|
||||
assert result["window_verified"] is False
|
||||
assert "not verified" in result["detail"]
|
||||
|
||||
|
||||
def test_the_stream_asks_the_server_to_report_its_usage():
|
||||
"""Measured: Ollama 0.33 sends no usage in a stream unless asked."""
|
||||
provider = OpenAICompatibleProvider(ENDPOINT, "m")
|
||||
from app.providers.base import PromptParts
|
||||
|
||||
for mode in ("chat", "completion"):
|
||||
provider.api_mode = mode
|
||||
_url, body = provider._request(PromptParts(system="s", story="t"), 0.7, 50)
|
||||
assert body["stream"] is True
|
||||
assert body["stream_options"] == {"include_usage": True}
|
||||
|
||||
|
||||
def _latest_ai(adv_id):
|
||||
with SessionLocal() as db:
|
||||
return (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
|
||||
|
||||
def _play(client, monkeypatch, usage, window=4096, reply="The crypt is still sealed."):
|
||||
async def verified(endpoint, model, declared=None, use_cache=True):
|
||||
return contextwindow.Window(window, contextwindow.LOADED, 32768, "fake")
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
|
||||
monkeypatch.setattr(ScriptedProvider, "last_usage", usage)
|
||||
ScriptedProvider.replies = [reply]
|
||||
return client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "look at the seal"})
|
||||
|
||||
|
||||
def test_a_turn_records_what_the_server_read(client, monkeypatch):
|
||||
response = _play(client, monkeypatch, None)
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
|
||||
|
||||
response = _play(client, monkeypatch,
|
||||
{"prompt_tokens": estimate + 13, "completion_tokens": 9})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
accounting = snapshot["accounting"]
|
||||
assert accounting["status"] == contextwindow.FITS
|
||||
assert accounting["server_prompt_tokens"] == estimate + 13
|
||||
assert accounting["estimate"] == snapshot["tokens"]["estimate"]
|
||||
assert '"accounting"' in response.text
|
||||
assert contextwindow.FITS in response.text
|
||||
|
||||
|
||||
def test_a_turn_with_no_reported_usage_is_unknown(client, monkeypatch):
|
||||
response = _play(client, monkeypatch, None)
|
||||
assert response.status_code == 200
|
||||
assert _latest_ai(client.adv_id).context_snapshot["accounting"]["status"] == (
|
||||
contextwindow.UNKNOWN)
|
||||
|
||||
|
||||
def test_a_suspected_truncation_keeps_the_turn_and_says_so(client, monkeypatch, caplog):
|
||||
"""A1-7 and A1-8. The reader watched the narration arrive; it stays."""
|
||||
response = _play(client, monkeypatch, {"prompt_tokens": 12, "completion_tokens": 9},
|
||||
reply="The seal holds, and the rain goes on.")
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
action = _latest_ai(client.adv_id)
|
||||
assert action is not None
|
||||
assert action.text == "The seal holds, and the rain goes on."
|
||||
accounting = action.context_snapshot["accounting"]
|
||||
assert accounting["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
assert contextwindow.TRUNCATION_SUSPECTED in response.text
|
||||
assert any(contextwindow.TRUNCATION_SUSPECTED in r.getMessage() for r in caplog.records)
|
||||
|
||||
# Inspectable afterwards through the same route the context panel reads.
|
||||
context = client.get(
|
||||
f"/api/adventures/{client.adv_id}/actions/{action.id}/context")
|
||||
assert context.status_code == 200
|
||||
assert context.json()["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
|
||||
|
||||
def test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves():
|
||||
"""Found by the A2 long run. Accounting belongs to one API call, not to the
|
||||
turn's shared prompt. A retry demotes the old attempt, and a take selection
|
||||
hands the prompt from one attempt to another. Neither may drop an attempt's
|
||||
accounting or give it another attempt's."""
|
||||
from app import attempts
|
||||
|
||||
class Node:
|
||||
def __init__(self, snapshot):
|
||||
self.context_snapshot = snapshot
|
||||
|
||||
shared = {"tokens": {"estimate": 3000}, "sections": [], "window": {"verified": True}}
|
||||
first = Node(shared | {"raw_output": "one", "usage": {"prompt_tokens": 3015},
|
||||
"accounting": {"status": contextwindow.FITS, "server_prompt_tokens": 3015}})
|
||||
second = Node({"raw_output": "two", "usage": {"prompt_tokens": 12},
|
||||
"accounting": {"status": contextwindow.TRUNCATION_SUSPECTED,
|
||||
"server_prompt_tokens": 12}})
|
||||
|
||||
# Superseded by a retry: the old attempt keeps only its own slices.
|
||||
attempts.keep_own_slices(Node(dict(first.context_snapshot)))
|
||||
demoted = Node(dict(first.context_snapshot))
|
||||
attempts.keep_own_slices(demoted)
|
||||
assert demoted.context_snapshot["accounting"]["server_prompt_tokens"] == 3015
|
||||
assert "tokens" not in demoted.context_snapshot
|
||||
|
||||
# The prompt moves to the second attempt; each keeps its own accounting.
|
||||
attempts.hand_over_the_prompt(first, second)
|
||||
assert second.context_snapshot["tokens"] == {"estimate": 3000}
|
||||
assert second.context_snapshot["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
assert second.context_snapshot["accounting"]["server_prompt_tokens"] == 12
|
||||
assert first.context_snapshot["accounting"]["status"] == contextwindow.FITS
|
||||
assert "tokens" not in first.context_snapshot
|
||||
|
||||
|
||||
def test_a_retry_leaves_each_take_with_its_own_accounting(client, monkeypatch):
|
||||
"""End to end, through the real retry route. Before the fix the live take
|
||||
inherited the superseded take's accounting, so the inspector could show one
|
||||
call's server count as another's."""
|
||||
response = _play(client, monkeypatch, None, reply="The first take.")
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
|
||||
# Replay the first take with a real count, so it has accounting of its own.
|
||||
with SessionLocal() as db:
|
||||
first = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == client.adv_id,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot)).first())
|
||||
snapshot = dict(first.context_snapshot)
|
||||
snapshot["accounting"] = contextwindow.classify_usage(
|
||||
{"prompt_tokens": estimate + 15}, estimate=estimate, budget=4096,
|
||||
max_output_tokens=500, window_verified=True)
|
||||
first.context_snapshot = snapshot
|
||||
db.commit()
|
||||
first_id = first.id
|
||||
|
||||
monkeypatch.setattr(ScriptedProvider, "last_usage",
|
||||
{"prompt_tokens": 12, "completion_tokens": 9})
|
||||
ScriptedProvider.replies = ["The second take."]
|
||||
retried = client.post(f"/api/adventures/{client.adv_id}/retry")
|
||||
assert retried.status_code == 200, retried.text[:300]
|
||||
|
||||
with SessionLocal() as db:
|
||||
rows = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == client.adv_id,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id).all())
|
||||
by_id = {row.id: row for row in rows}
|
||||
old = by_id[first_id]
|
||||
new = [row for row in rows if row.id != first_id][-1]
|
||||
assert new.text == "The second take."
|
||||
assert new.live and not old.live
|
||||
# The superseded take keeps its own accounting and gives up the prompt.
|
||||
assert old.context_snapshot["accounting"]["status"] == contextwindow.FITS
|
||||
assert old.context_snapshot["accounting"]["server_prompt_tokens"] == estimate + 15
|
||||
assert "tokens" not in old.context_snapshot
|
||||
# The live take carries the prompt and its own accounting, not the old one's.
|
||||
assert "tokens" in new.context_snapshot
|
||||
assert new.context_snapshot["accounting"]["status"] == (
|
||||
contextwindow.TRUNCATION_SUSPECTED)
|
||||
assert new.context_snapshot["accounting"]["server_prompt_tokens"] == 12
|
||||
|
||||
|
||||
def test_an_exceeded_turn_is_also_kept(client, monkeypatch):
|
||||
response = _play(client, monkeypatch, {"prompt_tokens": 3900, "completion_tokens": 9})
|
||||
assert response.status_code == 200
|
||||
action = _latest_ai(client.adv_id)
|
||||
assert action.text == "The crypt is still sealed."
|
||||
assert action.context_snapshot["accounting"]["status"] == contextwindow.EXCEEDED
|
||||
@@ -0,0 +1,324 @@
|
||||
"""v1.1 WP-A2: the protocol a narrator copies stays out of the story, and nothing else does.
|
||||
|
||||
The M11 closeout's identity run (an office meeting, a 3B narrator, a 4,096
|
||||
window) stored four turns carrying text the application wrote, not the story:
|
||||
|
||||
- `> Create_entity(new_person, "john", …` — the event vocabulary as the prompt
|
||||
printed it, `name(field, …)`, copied as if it were a call;
|
||||
- `> set_possession(silver-key, "alice") Adds the silver key to Alice's
|
||||
possession.` — the call again, naming the fantasy example slug from the fixed
|
||||
state rule, in a meeting room;
|
||||
- `Scene: Bill, Alice, … (at The meeting room)` — the renderer's own scene line;
|
||||
- `[Hard limit: your next turn must not exceed 180 words, … append the state
|
||||
block well inside the limit.]` — the length hint, reworded at the front and
|
||||
verbatim at the end.
|
||||
|
||||
v1.0.0 removed none of them. The rule this module is held to is unchanged from
|
||||
M5: **removing story is worse than leaving protocol.** Every removal below is
|
||||
anchored to a string or a vocabulary the application owns, and every one has
|
||||
story beside it that must survive.
|
||||
|
||||
python -m pytest tests/test_v11_protocol_echo.py -v
|
||||
"""
|
||||
|
||||
import json
|
||||
import re
|
||||
|
||||
import pytest
|
||||
|
||||
from app.context import builder
|
||||
from app.narrative import events, extract, render
|
||||
|
||||
# ---------------------------------------------------------------- the prompt
|
||||
|
||||
#: The identifiers the v1 state rule taught every campaign, from the fantasy
|
||||
#: acceptance fixture. None may come back into a fixed instruction.
|
||||
FANTASY_IDENTIFIERS = ("mara", "silver-key", "silver key", "old-abbey", "abbey",
|
||||
"aldric", "westhaven", "crypt", "edrin")
|
||||
#: And nothing from the science-fiction fixture either: neutral means neutral,
|
||||
#: not "the other genre".
|
||||
SCIFI_IDENTIFIERS = ("persephone", "imani", "data-crystal", "data crystal", "airlock")
|
||||
|
||||
|
||||
def _fixed_instructions() -> str:
|
||||
return "\n".join([
|
||||
extract.EMIT_RULE,
|
||||
extract.EMIT_REMINDER,
|
||||
events.vocabulary_for_prompt(),
|
||||
builder.length_hint(500),
|
||||
builder.length_hint(500, "brief"),
|
||||
builder.length_hint(500, "long"),
|
||||
builder.length_hint(120),
|
||||
]).lower()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("identifier", FANTASY_IDENTIFIERS + SCIFI_IDENTIFIERS)
|
||||
def test_the_fixed_state_instructions_name_no_fixture_identifier(identifier):
|
||||
"""A2-1. The example slug the office run copied cannot come back."""
|
||||
assert not re.search(rf"\b{re.escape(identifier)}\b", _fixed_instructions())
|
||||
|
||||
|
||||
def test_the_example_uses_neutral_identifiers():
|
||||
for neutral in ("character-1", "item-1", "location-1"):
|
||||
assert neutral in extract.EMIT_RULE
|
||||
|
||||
|
||||
def test_the_worked_example_is_a_block_this_extractor_accepts():
|
||||
"""The example is the wire format, byte for byte, not an illustration of it."""
|
||||
example = extract.EMIT_RULE[extract.EMIT_RULE.index("```state"):]
|
||||
prose, parsed, _raw = extract.split("The door opens.\n\n" + example)
|
||||
assert prose == "The door opens."
|
||||
assert isinstance(parsed, dict)
|
||||
assert [e["type"] for e in parsed["events"]] == [
|
||||
"set_possession", "set_current_location"]
|
||||
for event in parsed["events"]:
|
||||
assert events.is_allowed(event["type"])
|
||||
|
||||
|
||||
def test_the_vocabulary_is_not_written_as_function_calls():
|
||||
"""A2-2. `set_possession(item, owner)` is the notation the narrator copied."""
|
||||
vocabulary = events.vocabulary_for_prompt()
|
||||
for name in events.SPECS:
|
||||
assert not re.search(rf"\b{name}\s*\(", vocabulary), name
|
||||
|
||||
|
||||
def test_every_event_is_described_in_the_shape_the_model_must_send():
|
||||
lines = events.vocabulary_for_prompt().splitlines()
|
||||
assert len(lines) == len(events.SPECS)
|
||||
for name, line in zip(events.SPECS, lines):
|
||||
shape = line.strip().split(" — ", 1)[0]
|
||||
obj = json.loads(shape)
|
||||
assert obj["type"] == name
|
||||
assert set(obj) - {"type"} == set(events.SPECS[name]["required"])
|
||||
for optional in events.SPECS[name]["optional"]:
|
||||
assert optional in line
|
||||
|
||||
|
||||
def test_the_length_hint_carries_the_phrases_the_extractor_recognises():
|
||||
"""One source for the words, so the builder and the extractor cannot drift."""
|
||||
for narration_length in ("", "brief", "medium", "long"):
|
||||
hint = builder.length_hint(500, narration_length)
|
||||
assert hint.startswith(extract.LENGTH_HINT_OPENING)
|
||||
assert extract.LENGTH_HINT_TAIL in hint
|
||||
|
||||
|
||||
# ---------------------------------------------------------- observed shapes
|
||||
|
||||
STORY = (
|
||||
"Alice looks at John, the tension in the room palpable.\n\n"
|
||||
"John nods. \"I'm ready to contribute.\""
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("leak", [
|
||||
# Depth 20: the call, the fantasy slug, and a gloss on the same line.
|
||||
'> set_possession(silver-key, "alice") Adds the silver key to Alice\'s possession.',
|
||||
# Depth 10: cut off by the output limit mid-call.
|
||||
'> Create_entity(new_person, "john", "character", "A determined team member", ["john',
|
||||
# Unquoted, and a vocabulary name in any case.
|
||||
'SET_CURRENT_LOCATION(bill, office)',
|
||||
'add_fact(predicate="knows the plan", subject="alice")',
|
||||
])
|
||||
def test_an_event_call_line_at_the_end_leaves_the_story(leak):
|
||||
prose, parsed, _raw = extract.split(f"{STORY}\n\n{leak}")
|
||||
assert prose == STORY
|
||||
assert parsed is None
|
||||
|
||||
|
||||
def test_an_event_call_line_in_the_middle_leaves_and_the_story_after_it_stays():
|
||||
"""Depth 12: the call, then more narration."""
|
||||
reply = (
|
||||
f"{STORY}\n\n"
|
||||
'> Create_entity(new_person, "mike", "character", "A new team member.", ["mike"])\n\n'
|
||||
"Mike takes the empty chair by the window."
|
||||
)
|
||||
prose, _parsed, _raw = extract.split(reply)
|
||||
assert prose == f"{STORY}\n\nMike takes the empty chair by the window."
|
||||
|
||||
|
||||
def test_the_depth_fourteen_tail_leaves_entirely():
|
||||
"""A call, a rendered scene line, and a reworded length hint, in that order."""
|
||||
reply = (
|
||||
f"{STORY}\n\n"
|
||||
'> Create_entity(mike, "character", "A new team member.", ["mike"])\n\n'
|
||||
"Scene: Bill, Alice, Roger, John, and Mike at the table. (at The meeting room)\n\n"
|
||||
"[Hard limit: your next turn must not exceed 180 words, and it should not stop "
|
||||
"short of about 70. Prefer the lower end of that range unless the scene genuinely "
|
||||
"needs more. Finish the narration and append the state block well inside the limit.]"
|
||||
)
|
||||
prose, parsed, _raw = extract.split(reply)
|
||||
assert prose == STORY
|
||||
assert parsed is None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("hint", [
|
||||
builder.length_hint(500),
|
||||
builder.length_hint(500, "brief"),
|
||||
# Cut off by the output limit before the tail.
|
||||
"[Hard limit: this turn must not exceed 180 words, and it should not stop short",
|
||||
# Reworded at the front, as the 3B narrator did.
|
||||
"[Hard limit: your next turn must not exceed 506 words. Write only as much as the "
|
||||
"moment needs — a typical turn is much shorter. Finish the narration and append "
|
||||
"the state block well inside the limit.]",
|
||||
])
|
||||
def test_a_parroted_length_hint_at_the_end_leaves_the_story(hint):
|
||||
prose, _parsed, _raw = extract.split(f"{STORY}\n\n{hint}")
|
||||
assert prose == STORY
|
||||
|
||||
|
||||
def test_a_rendered_scene_line_at_the_end_leaves_the_story():
|
||||
prose, _parsed, _raw = extract.split(
|
||||
f"{STORY}\n\nScene: A tense budget meeting. (at The meeting room)")
|
||||
assert prose == STORY
|
||||
|
||||
|
||||
def test_a_fenced_block_with_a_call_line_above_it_still_parses_and_applies():
|
||||
"""A2-6. The proposal is still read when protocol litter surrounds it."""
|
||||
reply = (
|
||||
f"{STORY}\n\n"
|
||||
'> set_current_location(john, office)\n\n'
|
||||
'```state\n{"events": [{"type": "set_current_location", '
|
||||
'"entity": "john", "location": "office"}]}\n```'
|
||||
)
|
||||
prose, parsed, raw = extract.split(reply)
|
||||
assert prose == STORY
|
||||
assert parsed["events"][0]["entity"] == "john"
|
||||
assert raw.startswith("{")
|
||||
|
||||
|
||||
# ------------------------------------------------ adversarial story that stays
|
||||
|
||||
@pytest.mark.parametrize("reply", [
|
||||
# The owner's cases.
|
||||
'The engineer writes "set_power(core, 80)" on the whiteboard.',
|
||||
'She says, "Create_entity is a terrible name for a company."',
|
||||
'The old manual contains a heading labeled "Scene:"',
|
||||
'He reads aloud: "[Hard limit: 500 words]" and laughs.',
|
||||
# A vocabulary name, written into a story, not at the start of a line.
|
||||
'Nadia squints at the log: the last command was set_possession(badge, guard).',
|
||||
# Call-shaped, at the start of a line, but not an event this protocol has.
|
||||
"The terminal scrolls.\n\n> open_door(north)\n\nNothing happens.",
|
||||
# A vocabulary call inside the story's own code block is the story's code.
|
||||
"She types:\n\n```python\ncreate_entity(ship)\nset_possession(key, captain)\n```\n\n"
|
||||
"The console beeps twice.",
|
||||
# A bracket at the very end, in-world, that is not the application's hint.
|
||||
"The warning light blinks.\n\n[Hard limit of the reactor: three hours]",
|
||||
"The contract ends with a clause.\n\n[Hard limit: forty days, no extensions]",
|
||||
# A scene heading in a screenplay the characters are writing, mid-story.
|
||||
"Scene: a kitchen, late.\n\nShe crosses it out and starts again.",
|
||||
# A last line that starts like the renderer's but is not its shape.
|
||||
"The director calls it.\n\nScene: take two, and nobody moves.",
|
||||
# A fact restated inside a sentence.
|
||||
"Alice knew the badge opened the server room, and said nothing.",
|
||||
"Memory: she remembered the bells.",
|
||||
])
|
||||
def test_story_that_resembles_the_new_rules_is_kept(reply):
|
||||
"""A2-5."""
|
||||
prose, parsed, _raw = extract.split(reply)
|
||||
assert prose == reply
|
||||
assert parsed is None
|
||||
|
||||
|
||||
# ------------------------------------------------------ the replay attribution
|
||||
|
||||
@pytest.mark.parametrize("line, rule", [
|
||||
('> set_possession(silver-key, "alice") Adds the key.', extract.RULE_EVENT_CALL),
|
||||
("Create_entity(new_person", extract.RULE_EVENT_CALL),
|
||||
("[Hard limit: this turn must not exceed 90 words. Finish the narration and append "
|
||||
"the state block well inside the limit.]", extract.RULE_LENGTH_HINT),
|
||||
("Scene: A meeting. (at The meeting room)", extract.RULE_SCENE_LINE),
|
||||
("John nods.", None),
|
||||
('He reads aloud: "[Hard limit: 500 words]" and laughs.', None),
|
||||
])
|
||||
def test_a_removed_line_is_attributed_to_the_rule_that_removes_it(line, rule):
|
||||
assert extract.explain_removed_line(line) == rule
|
||||
|
||||
|
||||
# ------------------------------------- corrective: the depth-16 instruction tail
|
||||
|
||||
#: Cut down from the v1.1 identity diagnostic's depth-16 turn, whose stored text
|
||||
#: was exactly the extractor's output. The two story paragraphs are shortened;
|
||||
#: the four trailing lines are verbatim.
|
||||
DEPTH_16_STORY = (
|
||||
"John's initial ideas are thoughtful and insightful, and the room fills with a "
|
||||
"sense of optimism.\n\n"
|
||||
"John's enthusiasm is contagious, and the meeting room is electric with the "
|
||||
"excitement of a fruitful collaboration ahead."
|
||||
)
|
||||
DEPTH_16_TAIL = (
|
||||
"Scene: Bill, Alice and Roger at the table; John not yet arrived.\n\n"
|
||||
"[Hard limit: this is now 180 words.]\n\n"
|
||||
"[Reminder: end your reply with a `state` block listing the events your narration "
|
||||
"made true, with absolute values.]\n\n"
|
||||
"[You don't need to continue; your turn must now be about John entering the room. "
|
||||
"Continue the story here, directly. Output only story text.]"
|
||||
)
|
||||
|
||||
|
||||
def test_the_depth_sixteen_instruction_tail_leaves_entirely():
|
||||
"""The corrective's positive regression. v1.1's first A2 left all four lines:
|
||||
the last bracket was a reworded continue hint nothing recognised, so nothing
|
||||
above it was ever at the end."""
|
||||
prose, parsed, _raw = extract.split(f"{DEPTH_16_STORY}\n\n{DEPTH_16_TAIL}")
|
||||
assert prose == DEPTH_16_STORY
|
||||
assert parsed is None
|
||||
|
||||
|
||||
def test_the_continue_hint_phrase_is_the_providers_own_sentence():
|
||||
from app.providers.openai_compatible import CHAT_CONTINUE_HINT
|
||||
|
||||
assert extract.CONTINUE_HINT_PHRASE in CHAT_CONTINUE_HINT
|
||||
|
||||
|
||||
def test_an_echoed_continue_hint_alone_at_the_end_leaves():
|
||||
prose, _p, _r = extract.split(
|
||||
f"{STORY}\n\n[Keep going. Continue the story here, directly. Output only story text.]")
|
||||
assert prose == STORY
|
||||
|
||||
|
||||
@pytest.mark.parametrize("reply", [
|
||||
# A hint-opened bracket with no echoed instruction below it is in-world.
|
||||
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
|
||||
# Nor does a state block below it make it an instruction.
|
||||
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
|
||||
# The phrase in the middle of a story is prose, not a trailing echo.
|
||||
'She wrote "output only story text" on the card, then crossed it out.\n\nThe rain went on.',
|
||||
# A trailing in-world bracket that only resembles a continuation.
|
||||
f"{STORY}\n\n[To be continued]",
|
||||
])
|
||||
def test_story_brackets_near_the_corrective_rule_are_kept(reply):
|
||||
prose, _p, _r = extract.split(reply)
|
||||
assert prose == reply
|
||||
|
||||
|
||||
def test_a_hint_opened_bracket_above_a_state_block_is_kept():
|
||||
reply = (f"{STORY}\n\n[Hard limit: forty days, no extensions]\n\n"
|
||||
'```state\n{"events": []}\n```')
|
||||
prose, parsed, _r = extract.split(reply)
|
||||
assert prose == f"{STORY}\n\n[Hard limit: forty days, no extensions]"
|
||||
assert parsed == {"events": []}
|
||||
|
||||
|
||||
def test_an_in_world_bracket_above_an_echoed_hint_is_kept():
|
||||
"""Only a bracket opening the way the application's hint opens is taken with
|
||||
the echo. Any other bracket above it is the story's."""
|
||||
reply = (f"{STORY}\n\n[The sign on the door reads: Closed]\n\n"
|
||||
"[Continue the story here, directly. Output only story text.]")
|
||||
prose, _p, _r = extract.split(reply)
|
||||
assert prose == f"{STORY}\n\n[The sign on the door reads: Closed]"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("line, rule", [
|
||||
("[Hard limit: this is now 180 words.]", extract.RULE_INSTRUCTION_TAIL),
|
||||
("[Reminder: end your reply with a `state` block listing the events.]",
|
||||
extract.RULE_INSTRUCTION_TAIL),
|
||||
("[You don't need to continue. Output only story text.]", extract.RULE_INSTRUCTION_TAIL),
|
||||
])
|
||||
def test_the_corrective_rule_is_attributed(line, rule):
|
||||
assert extract.explain_removed_line(line) == rule
|
||||
|
||||
|
||||
def test_the_new_rules_do_not_disturb_the_section_headings_they_share_a_module_with():
|
||||
"""The renderer's headings are the M11 rules' anchor. A2 adds none."""
|
||||
assert render.HEADING_SCENE == "Scene:"
|
||||
assert "Scene:" not in render.SECTION_HEADINGS
|
||||
@@ -559,6 +559,17 @@ class Run:
|
||||
self.accepted += 1
|
||||
seconds = time.monotonic() - started
|
||||
sample = self.measure()
|
||||
# v1.1 WP-A1: what the server said it read for the turn just played, from
|
||||
# the `done` event. `.get` because a build before v1.1 sends none.
|
||||
done = next((e for e in events if e.get("type") == "done"), {})
|
||||
accounting = done.get("accounting") or {}
|
||||
sample.update({
|
||||
"accounting_status": accounting.get("status"),
|
||||
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
|
||||
"app_prompt_estimate": accounting.get("estimate"),
|
||||
"observed_margin": accounting.get("observed_margin"),
|
||||
"safety_reserve": accounting.get("safety_reserve"),
|
||||
})
|
||||
self.note("turn", text=text, seconds=round(seconds, 1), **sample)
|
||||
return {"accepted": True, "seconds": seconds, **sample}
|
||||
|
||||
@@ -1225,10 +1236,26 @@ PROTOCOL_LEAK_HEADING_RE = re.compile(
|
||||
re.MULTILINE,
|
||||
)
|
||||
PROTOCOL_LEAK_EVENTS = '"events"'
|
||||
#: v1.1 WP-A2: the two shapes the M11 closeout's identity run stored that the
|
||||
#: two signs above cannot see — a line opening with a call to an event, and the
|
||||
#: length hint echoed with the application's own wording. Copied, as above.
|
||||
PROTOCOL_LEAK_CALL_RE = re.compile(
|
||||
r"^[ \t]*(?:>[ \t]*)?(?:create_entity|set_entity_status|set_entity_attribute"
|
||||
r"|set_entity_conditions|set_current_location|set_possession|clear_possession"
|
||||
r"|add_fact|invalidate_fact|add_relationship|end_relationship"
|
||||
r"|open_story_thread|resolve_story_thread|set_scene)[ \t]*\(",
|
||||
re.IGNORECASE | re.MULTILINE,
|
||||
)
|
||||
PROTOCOL_LEAK_HINT_RE = re.compile(
|
||||
r"\[Hard limit:[^\]]*(?:append the state block|turn must not exceed \d+ words)",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
|
||||
|
||||
def _leaks_protocol(text: str) -> bool:
|
||||
return bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
|
||||
return (bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
|
||||
or bool(PROTOCOL_LEAK_CALL_RE.search(text))
|
||||
or bool(PROTOCOL_LEAK_HINT_RE.search(text)))
|
||||
|
||||
|
||||
def _protocol_leaks(bundle: dict) -> dict:
|
||||
|
||||
@@ -0,0 +1,187 @@
|
||||
"""v1.1: does a real v1.0.0 database open unchanged?
|
||||
|
||||
# the same database, opened by each tree, snapshotted read-only
|
||||
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
|
||||
--tree <v1.0.0 worktree>/backend --label v100 --out "$HOME/v11-evidence/compat"
|
||||
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
|
||||
--label v11 --exercise --out "$HOME/v11-evidence/compat"
|
||||
|
||||
Run from `backend/`. The source database is never opened. It is copied into
|
||||
`--out` first, and the copy is what the application opens.
|
||||
|
||||
A read-only snapshot is taken through the API, the same way a reader sees the
|
||||
campaign:
|
||||
|
||||
- the export bundle, which carries the whole tree, the head, the Save Points,
|
||||
state, events, summaries, memories and knowledge, and has no timestamp of its
|
||||
own;
|
||||
- the narrative state and its events;
|
||||
- the Save Points, the imported knowledge, the memories, the derived status and
|
||||
the settings;
|
||||
- the database schema and `PRAGMA user_version`, before and after the
|
||||
application opened it.
|
||||
|
||||
Two snapshots of the same database from two trees are then compared. Identical
|
||||
means v1.1 read it exactly as v1.0.0 did, and a matching schema and version mean
|
||||
nothing migrated.
|
||||
|
||||
`--exercise` then uses the v1.1 copy: undo, redo, a Save Point restore, a
|
||||
context dry run (knowledge retrieval), an export, and an import of that export.
|
||||
It first points the copy's endpoint at a loopback port that refuses, so nothing
|
||||
here reaches an inference server.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def _schema(path: Path) -> dict:
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
version = connection.execute("PRAGMA user_version").fetchone()[0]
|
||||
rows = connection.execute(
|
||||
"SELECT type, name, sql FROM sqlite_master WHERE name NOT LIKE 'sqlite_%' "
|
||||
"ORDER BY type, name").fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
return {"user_version": version, "objects": [list(r) for r in rows]}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--db", required=True)
|
||||
parser.add_argument("--tree", default="", help="a backend/ directory to import the app from")
|
||||
parser.add_argument("--label", required=True)
|
||||
parser.add_argument("--exercise", action="store_true")
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
copy = out / f"{args.label}.db"
|
||||
if copy.exists():
|
||||
print(f"{copy} exists; choose a new --label or --out")
|
||||
return 2
|
||||
shutil.copy2(args.db, copy)
|
||||
schema_before = _schema(copy)
|
||||
|
||||
if args.tree:
|
||||
sys.path.insert(0, str(Path(args.tree).resolve()))
|
||||
os.environ["AIDND_DB_PATH"] = str(copy)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, models
|
||||
from app.database import SessionLocal, get_db
|
||||
from app.main import app
|
||||
|
||||
print(f"app imported from {Path(sys.modules['app'].__file__).parent}")
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
with SessionLocal() as db:
|
||||
owner = db.query(models.Adventure.user_id).order_by(models.Adventure.id).first()
|
||||
user_id = owner[0] if owner else db.query(models.User.id).first()[0]
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
|
||||
def call(client, method, url, body=None, expect=200):
|
||||
response = client.request(method, f"/api{url}", json=body)
|
||||
if response.status_code != expect:
|
||||
raise SystemExit(f"{method} {url}: HTTP {response.status_code} {response.text[:300]}")
|
||||
return response.json() if response.content else None
|
||||
|
||||
report: dict = {"label": args.label, "schema_before": schema_before}
|
||||
with TestClient(app) as client:
|
||||
with SessionLocal() as db:
|
||||
adventure_ids = [a for (a,) in db.query(models.Adventure.id)
|
||||
.filter(models.Adventure.user_id == user_id)
|
||||
.order_by(models.Adventure.id)]
|
||||
snapshot = {"settings": call(client, "GET", "/settings"), "adventures": {}}
|
||||
for adv in adventure_ids:
|
||||
snapshot["adventures"][str(adv)] = {
|
||||
"export": call(client, "GET", f"/adventures/{adv}/export"),
|
||||
"state": call(client, "GET", f"/adventures/{adv}/state"),
|
||||
"state_events": call(client, "GET", f"/adventures/{adv}/state/events"),
|
||||
"checkpoints": call(client, "GET", f"/adventures/{adv}/checkpoints"),
|
||||
"knowledge": call(client, "GET", f"/adventures/{adv}/knowledge"),
|
||||
"memories": call(client, "GET", f"/adventures/{adv}/memories"),
|
||||
"derived": call(client, "GET", f"/adventures/{adv}/derived"),
|
||||
"newest_actions": call(client, "GET", f"/adventures/{adv}/actions?limit=5"),
|
||||
}
|
||||
report["snapshot"] = snapshot
|
||||
|
||||
if args.exercise and adventure_ids:
|
||||
adv = adventure_ids[0]
|
||||
ex: dict = {}
|
||||
call(client, "PUT", "/settings", {"endpoint_url": "http://127.0.0.1:9/v1",
|
||||
"embedding_model": ""})
|
||||
before = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
|
||||
ex["before"] = {k: before[k] for k in ("total", "can_undo", "can_redo")}
|
||||
undone = call(client, "POST", f"/adventures/{adv}/undo")
|
||||
ex["after_undo"] = {k: undone[k] for k in ("total", "can_undo", "can_redo")}
|
||||
redone = call(client, "POST", f"/adventures/{adv}/redo")
|
||||
ex["after_redo"] = {k: redone[k] for k in ("total", "can_undo", "can_redo")}
|
||||
points = call(client, "GET", f"/adventures/{adv}/checkpoints")
|
||||
if points:
|
||||
point = points[0]
|
||||
restored = call(client, "POST",
|
||||
f"/adventures/{adv}/checkpoints/{point['id']}/restore")
|
||||
page = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
|
||||
ex["restore"] = {"save_point": point["name"], "total": page["total"],
|
||||
"can_redo": page["can_redo"],
|
||||
"response_keys": sorted(restored or {})}
|
||||
context = client.get(f"/api/adventures/{adv}/context")
|
||||
body = context.json()
|
||||
ex["context"] = {
|
||||
"status": context.status_code,
|
||||
"knowledge_used": len(((body.get("knowledge") or {}).get("used")) or []),
|
||||
"canon_section": any(s["label"] == "campaign_canon"
|
||||
for s in body.get("sections") or []),
|
||||
"state_section": any(s["label"] == "narrative_state"
|
||||
for s in body.get("sections") or []),
|
||||
"summary": body.get("summary"),
|
||||
"tokens": body.get("tokens"),
|
||||
"window": body.get("window"),
|
||||
}
|
||||
bundle = call(client, "GET", f"/adventures/{adv}/export")
|
||||
imported = call(client, "POST", "/adventures/import", bundle, expect=201)
|
||||
new_id = imported["id"]
|
||||
reimport = call(client, "GET", f"/adventures/{new_id}/export")
|
||||
ex["import"] = {
|
||||
"new_id": new_id,
|
||||
"actions_in_bundle": len(bundle.get("actions") or []),
|
||||
"actions_after_import": len(reimport.get("actions") or []),
|
||||
"head_same": (bundle.get("headBranch") is not None
|
||||
and bundle.get("headDepth") == reimport.get("headDepth")),
|
||||
"checkpoints": [len(bundle.get("checkpoints") or []),
|
||||
len(reimport.get("checkpoints") or [])],
|
||||
"memories": [len(bundle.get("memories") or []),
|
||||
len(reimport.get("memories") or [])],
|
||||
"narrative_state_same": bundle.get("narrativeState") == reimport.get("narrativeState"),
|
||||
}
|
||||
report["exercise"] = ex
|
||||
|
||||
app.dependency_overrides.clear()
|
||||
report["schema_after"] = _schema(copy)
|
||||
(out / f"{args.label}.json").write_text(json.dumps(report, indent=2, sort_keys=True, default=str))
|
||||
same_schema = report["schema_before"] == report["schema_after"]
|
||||
print(f"schema unchanged by opening: {same_schema} "
|
||||
f"(user_version {report['schema_before']['user_version']} -> "
|
||||
f"{report['schema_after']['user_version']})")
|
||||
if "exercise" in report:
|
||||
print(json.dumps(report["exercise"], indent=2, default=str)[:3000])
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,245 @@
|
||||
"""v1.1 WP-A2: replay real stored narration through the v1.0.0 and current extractors.
|
||||
|
||||
The A2 extractor changes remove more text from a narrator's reply than v1.0.0
|
||||
did. Removing story is worse than leaving protocol (`TECHNICAL-DESIGN.md`
|
||||
§15.4), so every change is shown to a person rather than summarised. This tool
|
||||
takes every real reply the evidence kept, runs it through both extractors, and
|
||||
writes each turn whose prose differs with:
|
||||
|
||||
- the v1.0.0 prose and the current prose;
|
||||
- every line removed, and the rule that explains it;
|
||||
- any removal no rule explains, which fails the replay.
|
||||
|
||||
**Input.** A reply is read from the turn's stored `raw_output` where the
|
||||
evidence database kept one: that is exactly what the narrator sent. A bundle
|
||||
carries no raw reply, so a bundle's turns are replayed from their stored text,
|
||||
which is v1.0.0's output already. For those the old prose is the input itself,
|
||||
and the comparison is still exact.
|
||||
|
||||
**The v1.0.0 extractor** is read from the release tag with `git show`, not
|
||||
copied, so this tool compares against what shipped. It shares `events` and
|
||||
`render` with the current tree. A2 does not change `events.SPECS` or
|
||||
`render.SECTION_HEADINGS`, and the report verifies that with `git diff`.
|
||||
|
||||
Evidence stays outside the repository. Replayed text is fiction from the
|
||||
acceptance fixtures, but it is still somebody's run.
|
||||
|
||||
.venv/bin/python -m tools.v11_replay_extractor \\
|
||||
--db "$HOME/m11-evidence/**/*.db" \\
|
||||
--bundle "$HOME/m11-evidence/closeout-3652dc6/identity/turn-99/bundle.json" \\
|
||||
--out "$HOME/v11-evidence/a2-replay"
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import difflib
|
||||
import glob
|
||||
import hashlib
|
||||
import importlib.util
|
||||
import json
|
||||
import sqlite3
|
||||
import subprocess
|
||||
import sys
|
||||
import types
|
||||
import zlib
|
||||
from pathlib import Path
|
||||
|
||||
RELEASE = "v1.0.0"
|
||||
EXTRACT_PATH = "backend/app/narrative/extract.py"
|
||||
#: A removal larger than this share of the v1.0.0 prose is flagged for review
|
||||
#: even when every line is explained, because a rule that eats most of a reply
|
||||
#: is the shape a false positive takes.
|
||||
LARGE_REMOVAL_SHARE = 0.25
|
||||
|
||||
|
||||
def load_release_extractor(repo: Path) -> types.ModuleType:
|
||||
"""`app.narrative.extract` as it was at the release tag."""
|
||||
source = subprocess.run(
|
||||
["git", "-C", str(repo), "show", f"{RELEASE}:{EXTRACT_PATH}"],
|
||||
check=True, capture_output=True, text=True,
|
||||
).stdout
|
||||
import app.narrative # noqa: F401 the package the relative import needs
|
||||
|
||||
spec = importlib.util.spec_from_loader("app.narrative._extract_release", loader=None)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
module.__package__ = "app.narrative"
|
||||
exec(compile(source, f"{RELEASE}:{EXTRACT_PATH}", "exec"), module.__dict__)
|
||||
return module
|
||||
|
||||
|
||||
def _unpack(blob):
|
||||
if blob is None:
|
||||
return None
|
||||
try:
|
||||
return json.loads(zlib.decompress(bytes(blob)).decode("utf-8"))
|
||||
except (zlib.error, ValueError, UnicodeDecodeError):
|
||||
return None
|
||||
|
||||
|
||||
def turns_from_db(path: str):
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
rows = connection.execute(
|
||||
"SELECT id, depth, text, context_snapshot FROM actions WHERE type = 'ai'"
|
||||
).fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
for action_id, depth, text, blob in rows:
|
||||
snapshot = _unpack(blob) or {}
|
||||
raw = snapshot.get("raw_output")
|
||||
if isinstance(raw, str) and raw.strip():
|
||||
yield {"source": path, "id": action_id, "depth": depth,
|
||||
"input": raw, "input_kind": "raw_output"}
|
||||
elif text:
|
||||
yield {"source": path, "id": action_id, "depth": depth,
|
||||
"input": text, "input_kind": "stored_text"}
|
||||
|
||||
|
||||
def turns_from_bundle(path: str):
|
||||
bundle = json.loads(Path(path).read_text())
|
||||
for action in bundle.get("actions") or []:
|
||||
if action.get("type") == "ai" and action.get("text"):
|
||||
yield {"source": path, "id": action.get("id"), "depth": action.get("depth"),
|
||||
"input": action["text"], "input_kind": "stored_text"}
|
||||
|
||||
|
||||
def removed_lines(before: str, after: str) -> list[str]:
|
||||
"""Lines present in `before` and gone from `after`, in order.
|
||||
|
||||
Compared with trailing whitespace ignored. The extractor strips the end of
|
||||
every reply it cuts, so a story line that becomes the last line loses a
|
||||
trailing space. A first version of this tool reported that space as a
|
||||
rewritten line of story, which it is not.
|
||||
"""
|
||||
old = [line.rstrip() for line in before.split("\n")]
|
||||
new = [line.rstrip() for line in after.split("\n")]
|
||||
matcher = difflib.SequenceMatcher(a=old, b=new, autojunk=False)
|
||||
gone: list[str] = []
|
||||
for tag, a0, a1, _b0, _b1 in matcher.get_opcodes():
|
||||
if tag in ("delete", "replace"):
|
||||
gone.extend(old[a0:a1])
|
||||
return gone
|
||||
|
||||
|
||||
def replay(inputs, old, new) -> dict:
|
||||
seen: set[str] = set()
|
||||
unchanged = 0
|
||||
changed: list[dict] = []
|
||||
duplicates = 0
|
||||
for turn in inputs:
|
||||
digest = hashlib.sha256(turn["input"].encode()).hexdigest()
|
||||
if digest in seen:
|
||||
duplicates += 1
|
||||
continue
|
||||
seen.add(digest)
|
||||
old_prose, _old_parsed, _old_raw = old.split(turn["input"])
|
||||
new_prose, _new_parsed, _new_raw = new.split(turn["input"])
|
||||
if old_prose == new_prose:
|
||||
unchanged += 1
|
||||
continue
|
||||
lines = []
|
||||
unexplained = 0
|
||||
for line in removed_lines(old_prose, new_prose):
|
||||
if not line.strip():
|
||||
continue
|
||||
rule = new.explain_removed_line(line)
|
||||
if rule is None:
|
||||
unexplained += 1
|
||||
lines.append({"line": line, "rule": rule})
|
||||
added = [line for line in removed_lines(new_prose, old_prose) if line.strip()]
|
||||
share = 1 - len(new_prose) / max(1, len(old_prose))
|
||||
flags = []
|
||||
if unexplained:
|
||||
flags.append("unexplained_removal")
|
||||
if added:
|
||||
flags.append("text_added_or_rewritten")
|
||||
if share > LARGE_REMOVAL_SHARE:
|
||||
flags.append("large_removal")
|
||||
changed.append({
|
||||
**{k: turn[k] for k in ("source", "id", "depth", "input_kind")},
|
||||
"sha256": digest,
|
||||
"old_prose": old_prose,
|
||||
"new_prose": new_prose,
|
||||
"removed": lines,
|
||||
"added_or_rewritten": added,
|
||||
"removed_chars": len(old_prose) - len(new_prose),
|
||||
"removed_share": round(share, 4),
|
||||
"flags": flags,
|
||||
})
|
||||
return {
|
||||
"replayed": unchanged + len(changed),
|
||||
"duplicates_skipped": duplicates,
|
||||
"unchanged": unchanged,
|
||||
"changed": len(changed),
|
||||
"flagged": sum(1 for c in changed if c["flags"]),
|
||||
"turns": changed,
|
||||
}
|
||||
|
||||
|
||||
def write_markdown(result: dict, path: Path) -> None:
|
||||
out = [
|
||||
"# A2 extractor replay",
|
||||
"",
|
||||
f"- replayed (unique replies): **{result['replayed']}**",
|
||||
f"- duplicates skipped: {result['duplicates_skipped']}",
|
||||
f"- unchanged: {result['unchanged']}",
|
||||
f"- changed: **{result['changed']}**",
|
||||
f"- flagged: **{result['flagged']}**",
|
||||
"",
|
||||
]
|
||||
for index, turn in enumerate(result["turns"], 1):
|
||||
out += [
|
||||
f"## {index}. {Path(turn['source']).parent.name}/{Path(turn['source']).name}"
|
||||
f" action {turn['id']} depth {turn['depth']} ({turn['input_kind']})",
|
||||
"",
|
||||
f"- removed chars: {turn['removed_chars']} ({turn['removed_share']:.1%})",
|
||||
f"- flags: {', '.join(turn['flags']) or 'none'}",
|
||||
"",
|
||||
"Removed lines:",
|
||||
"",
|
||||
]
|
||||
for item in turn["removed"]:
|
||||
out.append(f"- `{item['rule'] or 'UNEXPLAINED'}` — {item['line']!r}")
|
||||
out += ["", "<details><summary>v1.0.0 prose</summary>", "", "```text",
|
||||
turn["old_prose"], "```", "</details>", "",
|
||||
"<details><summary>current prose</summary>", "", "```text",
|
||||
turn["new_prose"], "```", "</details>", ""]
|
||||
path.write_text("\n".join(out))
|
||||
|
||||
|
||||
def main(argv=None) -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--db", action="append", default=[],
|
||||
help="an evidence database, or a glob of them")
|
||||
parser.add_argument("--bundle", action="append", default=[],
|
||||
help="an exported bundle whose campaign has no database here")
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args(argv)
|
||||
|
||||
repo = Path(__file__).resolve().parents[2]
|
||||
from app.narrative import extract as current
|
||||
|
||||
old = load_release_extractor(repo)
|
||||
dbs = sorted({p for pattern in args.db for p in glob.glob(pattern, recursive=True)})
|
||||
|
||||
def inputs():
|
||||
for path in dbs:
|
||||
yield from turns_from_db(path)
|
||||
for path in args.bundle:
|
||||
yield from turns_from_bundle(path)
|
||||
|
||||
result = replay(inputs(), old, current)
|
||||
result["databases"] = dbs
|
||||
result["bundles"] = args.bundle
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
(out / "replay.json").write_text(json.dumps(result, indent=2, ensure_ascii=False))
|
||||
write_markdown(result, out / "replay.md")
|
||||
print(json.dumps({k: result[k] for k in
|
||||
("replayed", "duplicates_skipped", "unchanged", "changed", "flagged")}))
|
||||
return 1 if result["flagged"] else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,211 @@
|
||||
"""v1.1 WP-A1: real turns, and what the server said it read.
|
||||
|
||||
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 AIDND_TEST_MODEL=<model> \\
|
||||
.venv/bin/python -m tools.v11_window_accounting \\
|
||||
--bundle "$HOME/m11-evidence/m04-final/bundle.json" --turns 4 \\
|
||||
--out "$HOME/v11-evidence/a1-accounting/<label>"
|
||||
|
||||
Run from `backend/`. The evidence that matters is at the edge of the window, and
|
||||
a new campaign takes dozens of turns to reach it. So this imports a long
|
||||
campaign, by default the v1 evidence run's 207-action bundle, and every turn is
|
||||
assembled against a full window from the first. A real narrator is then asked
|
||||
for `--turns` turns, and each one is written out with:
|
||||
|
||||
configured budget, verified window, and where the window came from
|
||||
the application's estimate of what it sent (its count, plus the text the
|
||||
provider adds)
|
||||
the server's own prompt-token count, from the usage it reported
|
||||
the reply allocation and the safety reserve
|
||||
the observed margin: window - reply allocation - the server's count
|
||||
the accounting status: fits, exceeded, truncation_suspected or unknown
|
||||
|
||||
The first turn on a cold model cannot verify the window: `/api/ps` knows nothing
|
||||
until the model is loaded. That is M11's behaviour, and the row says so rather
|
||||
than hiding the turn.
|
||||
|
||||
The database lives in `--out`, not in `/tmp`, so the stored snapshots behind
|
||||
every row can be read again. The endpoint is read from the environment and is
|
||||
never written into a committed file.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
|
||||
TURNS = [
|
||||
"I look around carefully and take stock of where I am.",
|
||||
"I ask the nearest person what has happened since I was last here.",
|
||||
"I check what I am carrying.",
|
||||
"I move on towards the place I meant to reach.",
|
||||
"I wait and listen.",
|
||||
"I say, \"Tell me the part you left out.\"",
|
||||
]
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--bundle", default="",
|
||||
help="a campaign bundle to import, so turns start at a full window")
|
||||
parser.add_argument("--turns", type=int, default=4)
|
||||
parser.add_argument("--budget", type=int, default=16384)
|
||||
parser.add_argument("--max-output", type=int, default=500)
|
||||
parser.add_argument("--timeout", type=int, default=900)
|
||||
parser.add_argument("--out", required=True)
|
||||
parser.add_argument("--unload-first", action="store_true",
|
||||
help="ask the configured server to unload the model before turn 1, "
|
||||
"so the first turn starts cold (the A1 corrective test)")
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL):
|
||||
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
|
||||
return 2
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
db_path = out / "accounting.db"
|
||||
if db_path.exists():
|
||||
print(f"{db_path} exists; choose a new --out")
|
||||
return 2
|
||||
os.environ["AIDND_DB_PATH"] = str(db_path)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
Base.metadata.create_all(bind=engine)
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email="accounting@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
|
||||
context_token_budget=args.budget, max_output_tokens=args.max_output,
|
||||
model_timeout_seconds=args.timeout,
|
||||
))
|
||||
db.commit()
|
||||
user_id = user.id
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
client = TestClient(app)
|
||||
|
||||
if args.bundle:
|
||||
bundle = json.loads(Path(args.bundle).read_text())
|
||||
# The evidence campaign had its memory bank and auto-summarise on. Here
|
||||
# they would only add post-turn model calls between the measured turns,
|
||||
# on the same host, and fail noisily as the in-process client closes.
|
||||
# This tool measures the turn's own prompt, which neither changes.
|
||||
bundle["memoryBankEnabled"] = False
|
||||
bundle["autoSummarize"] = False
|
||||
imported = client.post("/api/adventures/import", json=bundle)
|
||||
imported.raise_for_status()
|
||||
adv = imported.json()["id"]
|
||||
else:
|
||||
created = client.post("/api/adventures", json={
|
||||
"title": "Window accounting", "opening": "A quiet road at dusk."})
|
||||
created.raise_for_status()
|
||||
adv = created.json()["id"]
|
||||
|
||||
if args.unload_first:
|
||||
# The configured endpoint only, under the same policy and TLS trust as a turn.
|
||||
import asyncio
|
||||
import httpx
|
||||
from app import contextwindow, endpoints, tlstrust
|
||||
reason = endpoints.rejection_reason(ENDPOINT)
|
||||
if reason:
|
||||
print(f"endpoint refused: {reason}")
|
||||
return 2
|
||||
base = contextwindow.native_base(ENDPOINT)
|
||||
with httpx.Client(verify=tlstrust.ssl_context(), timeout=120) as http:
|
||||
unloaded = http.post(f"{base}/api/generate", json={"model": MODEL, "keep_alive": 0})
|
||||
resident = [m.get("name") for m in http.get(f"{base}/api/ps").json().get("models", [])]
|
||||
contextwindow.cache_clear()
|
||||
print(f"unload: HTTP {unloaded.status_code} {unloaded.text[:120]} | resident now: {resident}")
|
||||
|
||||
rows: list[dict] = []
|
||||
timeline = (out / "turns.jsonl").open("a")
|
||||
print(f"model {MODEL}, budget {args.budget}, reply {args.max_output}")
|
||||
print(f"{'#':>2} {'status':22} {'window':>13} {'estimate':>8} {'server':>7} "
|
||||
f"{'reserve':>7} {'margin':>7} {'sec':>5}")
|
||||
for index in range(args.turns):
|
||||
text = TURNS[index % len(TURNS)]
|
||||
started = time.monotonic()
|
||||
response = client.post(f"/api/adventures/{adv}/actions",
|
||||
json={"type": "do", "text": text})
|
||||
seconds = round(time.monotonic() - started, 1)
|
||||
error = None
|
||||
if response.status_code != 200 or '"type": "error"' in response.text:
|
||||
error = response.text[-400:]
|
||||
with SessionLocal() as db:
|
||||
action = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
snapshot = (action.context_snapshot or {}) if action else {}
|
||||
tokens = snapshot.get("tokens") or {}
|
||||
window = snapshot.get("window") or {}
|
||||
accounting = snapshot.get("accounting") or {}
|
||||
row = {
|
||||
"turn": index + 1,
|
||||
"action_id": action.id if action else None,
|
||||
"seconds": seconds,
|
||||
"error": error,
|
||||
"configured_budget": tokens.get("configured_budget"),
|
||||
"effective_budget": tokens.get("budget"),
|
||||
"window_verified": window.get("verified"),
|
||||
"window_tokens": window.get("tokens"),
|
||||
"window_source": window.get("source"),
|
||||
"preflight_attempted": (window.get("preflight") or {}).get("attempted"),
|
||||
"preflight_loaded": (window.get("preflight") or {}).get("loaded"),
|
||||
"preflight_verified_before": (window.get("preflight") or {}).get("verified_before"),
|
||||
"preflight_verified_after": (window.get("preflight") or {}).get("verified_after"),
|
||||
"preflight_detail": (window.get("preflight") or {}).get("detail"),
|
||||
"app_prompt_tokens": tokens.get("total"),
|
||||
"transport_tokens": tokens.get("transport"),
|
||||
"app_estimate": tokens.get("estimate"),
|
||||
"output_reserve": tokens.get("output_reserve"),
|
||||
"safety_reserve": tokens.get("safety_reserve"),
|
||||
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
|
||||
"difference": accounting.get("difference"),
|
||||
"observed_margin": accounting.get("observed_margin"),
|
||||
"status": accounting.get("status"),
|
||||
"history_included": (snapshot.get("history") or {}).get("included"),
|
||||
"history_total": (snapshot.get("history") or {}).get("total"),
|
||||
}
|
||||
rows.append(row)
|
||||
timeline.write(json.dumps(row) + "\n")
|
||||
timeline.flush()
|
||||
print(f"{row['turn']:>2} {str(row['status'] if not error else 'ERROR'):22} "
|
||||
f"{str(row['window_tokens'])+('v' if row['window_verified'] else '?'):>13} "
|
||||
f"{str(row['app_estimate']):>8} {str(row['server_prompt_tokens']):>7} "
|
||||
f"{str(row['safety_reserve']):>7} {str(row['observed_margin']):>7} {seconds:>5}")
|
||||
if error:
|
||||
print(f" error: {error[:200]}")
|
||||
timeline.close()
|
||||
(out / "summary.json").write_text(json.dumps({
|
||||
"model": MODEL, "budget": args.budget, "max_output_tokens": args.max_output,
|
||||
"bundle": args.bundle, "rows": rows,
|
||||
}, indent=2))
|
||||
app.dependency_overrides.clear()
|
||||
return 0 if all(r["error"] is None for r in rows) else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -47,6 +47,63 @@ function Section({ title, count, children, open = false, testId }) {
|
||||
)
|
||||
}
|
||||
|
||||
/* v1.1 WP-A1: what the server said it read, set against what was sent.
|
||||
*
|
||||
* Only a turn that was actually sent has this — the next-turn view has not
|
||||
* been sent yet. The two conditions that mean something went wrong are shown
|
||||
* as alerts, because the failure they describe is otherwise silent: Ollama
|
||||
* answers 200 whether or not it cut the front of the prompt off. */
|
||||
const ACCOUNTING = {
|
||||
fits: {
|
||||
title: 'The server read the whole prompt',
|
||||
body: 'Its count stayed inside the room kept for the reply and the safety margin.',
|
||||
},
|
||||
exceeded: {
|
||||
title: 'The prompt was larger than the server allowed for',
|
||||
body: 'The server counted more tokens than the safety margin covers, so the '
|
||||
+ 'reply may have been cut short. The turn is kept.',
|
||||
alert: true,
|
||||
},
|
||||
truncation_suspected: {
|
||||
title: 'The server may have cut the start of the prompt',
|
||||
body: 'It read far fewer tokens than were sent, which is what happens when a '
|
||||
+ 'prompt is larger than the window the model was loaded with. The '
|
||||
+ 'narrator’s rules and the campaign canon are at the start. The turn is kept.',
|
||||
alert: true,
|
||||
},
|
||||
unknown: {
|
||||
title: 'The server did not say how much it read',
|
||||
body: 'Nothing here can confirm whether the whole prompt was used.',
|
||||
},
|
||||
}
|
||||
|
||||
function AccountingReport({ accounting }) {
|
||||
if (!accounting) return null
|
||||
const copy = ACCOUNTING[accounting.status] || ACCOUNTING.unknown
|
||||
return (
|
||||
<div
|
||||
className={copy.alert ? 'notice error' : 'ctx-accounting'}
|
||||
role={copy.alert ? 'alert' : undefined}
|
||||
data-testid="ctx-accounting"
|
||||
data-status={accounting.status}
|
||||
>
|
||||
<strong>{copy.title}</strong>
|
||||
<p>{copy.body}</p>
|
||||
{accounting.server_prompt_tokens != null && (
|
||||
<p className="notice-detail">
|
||||
Sent {accounting.estimate?.toLocaleString()} by this app’s count; the
|
||||
server read {accounting.server_prompt_tokens.toLocaleString()}.
|
||||
{accounting.observed_margin != null
|
||||
&& ` ${accounting.observed_margin.toLocaleString()} tokens were left beside the reply.`}
|
||||
</p>
|
||||
)}
|
||||
{accounting.window_verified === false && (
|
||||
<p className="notice-detail">The model’s window was not verified for this turn.</p>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
function TokenBar({ sections, total }) {
|
||||
if (!sections.length || total <= 0) return null
|
||||
return (
|
||||
@@ -122,6 +179,12 @@ export function ContextPanel({
|
||||
<span>{tokens.output_reserve.toLocaleString()}</span>
|
||||
</div>
|
||||
)}
|
||||
{tokens.safety_reserve > 0 && (
|
||||
<div className="ctx-token-line dim">
|
||||
<span>Kept free as a safety margin</span>
|
||||
<span>{tokens.safety_reserve.toLocaleString()}</span>
|
||||
</div>
|
||||
)}
|
||||
<div className="ctx-token-line dim">
|
||||
<span>What the model can hold</span>
|
||||
<span>{tokens.budget.toLocaleString()}</span>
|
||||
@@ -135,6 +198,8 @@ export function ContextPanel({
|
||||
)}
|
||||
</div>
|
||||
|
||||
<AccountingReport accounting={report.accounting} />
|
||||
|
||||
{failing.length > 0 && (
|
||||
<div className="notice error" role="alert" data-testid="ctx-derived-failing">
|
||||
<strong>Background work is failing</strong>
|
||||
|
||||
@@ -220,6 +220,50 @@ describe('context inspector (§23, §55)', () => {
|
||||
expect(tokens).toHaveTextContent('8,000')
|
||||
})
|
||||
|
||||
it('shows the safety margin kept free beside the reply (v1.1 A1)', async () => {
|
||||
api.getAdventureContext.mockResolvedValue({
|
||||
...REPORT, tokens: { ...REPORT.tokens, safety_reserve: 410 },
|
||||
})
|
||||
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
|
||||
expect(screen.getByTestId('ctx-tokens')).toHaveTextContent(/safety margin\s*410/)
|
||||
})
|
||||
|
||||
it('says nothing about accounting for a turn that has not been sent', async () => {
|
||||
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
|
||||
expect(screen.queryByTestId('ctx-accounting')).toBeNull()
|
||||
})
|
||||
|
||||
it('alerts when the server may have cut the start of a sent prompt (v1.1 A1)', async () => {
|
||||
vi.spyOn(api, 'getActionContext').mockResolvedValue({
|
||||
...REPORT,
|
||||
accounting: {
|
||||
status: 'truncation_suspected', estimate: 6316, server_prompt_tokens: 2050,
|
||||
observed_margin: 1546, window_verified: true,
|
||||
},
|
||||
})
|
||||
await renderWith(
|
||||
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
|
||||
const el = screen.getByTestId('ctx-accounting')
|
||||
expect(el).toHaveAttribute('data-status', 'truncation_suspected')
|
||||
expect(el).toHaveAttribute('role', 'alert')
|
||||
expect(el).toHaveTextContent('6,316')
|
||||
expect(el).toHaveTextContent('2,050')
|
||||
expect(el).toHaveTextContent(/turn is kept/)
|
||||
})
|
||||
|
||||
it('reports an unknown count plainly, without claiming the prompt fit', async () => {
|
||||
vi.spyOn(api, 'getActionContext').mockResolvedValue({
|
||||
...REPORT, accounting: { status: 'unknown', server_prompt_tokens: null },
|
||||
})
|
||||
await renderWith(
|
||||
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
|
||||
const el = screen.getByTestId('ctx-accounting')
|
||||
expect(el).toHaveAttribute('data-status', 'unknown')
|
||||
expect(el).not.toHaveAttribute('role')
|
||||
expect(el).toHaveTextContent(/did not say/)
|
||||
expect(el.textContent).not.toMatch(/read the whole prompt/)
|
||||
})
|
||||
|
||||
it('shows a retrieved passage with its source, class, heading and score', async () => {
|
||||
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
|
||||
const row = document.querySelector('[data-chunk-id="11"]')
|
||||
|
||||
@@ -32,6 +32,16 @@
|
||||
font-variant-numeric: tabular-nums;
|
||||
}
|
||||
.ctx-warn { margin: 8px 0 0; color: var(--danger); font-size: 0.76rem; }
|
||||
/* v1.1 A1: what the server said it read. The two failure states render as a
|
||||
`.notice.error` alert instead; this is the quiet form for `fits` and
|
||||
`unknown`. */
|
||||
.ctx-accounting {
|
||||
padding: 9px 13px;
|
||||
border: 1px solid var(--border);
|
||||
border-radius: 7px;
|
||||
color: var(--text-dim);
|
||||
}
|
||||
.ctx-accounting p { margin: 4px 0 0; }
|
||||
|
||||
.token-bar {
|
||||
display: flex;
|
||||
|
||||
@@ -69,6 +69,18 @@ out of tokens partway through it. All of that is removed before the prose is
|
||||
stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4
|
||||
has the rules.
|
||||
|
||||
**Implementation note (v1.1 WP-A2).** A small model also copies the protocol's
|
||||
*instructions*: the vocabulary written as calls, the length hint, and the scene
|
||||
line. The fix is on both sides:
|
||||
|
||||
- **Prompt:** the vocabulary is shown in the wire format, and the fixed example
|
||||
uses genre-neutral placeholders.
|
||||
- **Extractor:** it recognises those echoes only by strings and names the
|
||||
application owns.
|
||||
|
||||
No event type, field, validation rule or proposal record changed.
|
||||
`TECHNICAL-DESIGN.md` §15.4 lists the four rules.
|
||||
|
||||
### Where it lives
|
||||
|
||||
- `adventures.narrative_state` — the current authoritative document. This is
|
||||
|
||||
+3
-1
@@ -2,7 +2,9 @@
|
||||
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 planning has begun.**
|
||||
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: WP-A1
|
||||
and WP-A2 are implemented and staged for owner review**
|
||||
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`).
|
||||
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
|
||||
through M11 complete and closed**. M11 was accepted at its closeout
|
||||
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
|
||||
|
||||
@@ -1230,6 +1230,75 @@ whether there is a number to cap to at all, and that is what the builder and the
|
||||
declaration everywhere it appears, and the connection test says plainly that
|
||||
nothing has checked it against the server.
|
||||
|
||||
**As implemented (v1.1 WP-A1): a safety reserve, and the server's own count.**
|
||||
A ceiling in the application's tokens is not a ceiling in the narrator's. The
|
||||
builder counts with `cl100k_base`, and the v1 evidence left the largest prompts
|
||||
23-42 real tokens from the edge of a 16,384 window. Past the edge Ollama does not
|
||||
refuse. Measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt returned
|
||||
200 with `prompt_tokens` 2,050.
|
||||
|
||||
- **The reserve.** `contextwindow.safety_reserve(budget)` is
|
||||
`max(256, ceil(5% of the effective budget))`, computed with integer rounding
|
||||
up: 256 at 4,096, 410 at 8,192, 820 at 16,384. It is taken before any history
|
||||
is chosen. The effective budget is the verified or declared window when there
|
||||
is one, and the configured budget otherwise. It is fixed and documented, is not
|
||||
a setting, and is not calibrated per model.
|
||||
- **What replaced the 64-token margin.** M6's `OUTPUT_SAFETY_MARGIN` absorbed two
|
||||
unrelated things.
|
||||
- The application's own text added after pricing: separators between
|
||||
sections, and the chat hint the provider appends to every request. This is
|
||||
now priced exactly as `transport`.
|
||||
- Tokenizer drift. This is now the reserve.
|
||||
|
||||
The reply allocation is exactly `max_output_tokens`. Protected context is
|
||||
`sections + transport + reply + reserve`, and `ContextOverflow` is raised
|
||||
before the model call when that does not fit.
|
||||
- **The server's count.** Streaming requests set `stream_options.include_usage`.
|
||||
Without it Ollama sends no usage, and none of the 514 AI turns in the v1
|
||||
evidence has one. After the reply, `contextwindow.classify_usage` compares the
|
||||
server's `prompt_tokens` with `tokens.estimate`: the assembled text plus what
|
||||
the provider adds.
|
||||
- **Accounting states.** The turn's snapshot records `accounting`, whose status
|
||||
is one of the following, checked in this order:
|
||||
|
||||
| Status | Meaning |
|
||||
| --- | --- |
|
||||
| `unknown` | No positive integer count was reported. It is never read as `fits`. |
|
||||
| `truncation_suspected` | The server read fewer tokens than the estimate by more than the reserve. |
|
||||
| `exceeded` | The server's count plus the reply allocation is over the budget. |
|
||||
| `fits` | Otherwise. |
|
||||
|
||||
The record also carries the server's count, the difference, the reserve and
|
||||
the observed margin (`budget - reply - server count`).
|
||||
- **Surfacing.** The record is returned on the turn's `done` event, logged as a
|
||||
warning when it is `exceeded` or `truncation_suspected`, and shown in the
|
||||
context inspector, where those two statuses are an alert.
|
||||
- **The turn is kept.** A discrepancy found after the reply is recorded, never
|
||||
enforced. The narration has already streamed to the reader, and the accepted
|
||||
turn is not discarded.
|
||||
- **Accounting belongs to one attempt.** It sits in `attempts.ATTEMPT_KEYS`
|
||||
beside `usage`. When a retry or a take selection moves the shared prompt,
|
||||
each take keeps the accounting for its own call.
|
||||
- **A cold model is loaded before its turn is built (v1.1 A1 corrective).** A
|
||||
model that is not resident cannot report its window. The A1 evidence caught
|
||||
exactly that: 13,875 tokens were sent to a server that read 2,050.
|
||||
`contextwindow.ensure_window` works like this:
|
||||
- it probes;
|
||||
- if the window is unverified and the server answered, it makes one bounded
|
||||
`POST /api/generate` naming only the model, with no prompt. Ollama loads
|
||||
the model and generates nothing ("done_reason": "load");
|
||||
- it probes again, bypassing the cache;
|
||||
- the turn is built to whatever that second probe says.
|
||||
|
||||
A failed load, or a window still unverified afterwards, changes nothing: the
|
||||
configured budget stands and the accounting still catches a cut. The request
|
||||
goes to the configured endpoint only, under the same policy and TLS path. It
|
||||
writes nothing, and it is recorded as `window.preflight` in the turn's
|
||||
snapshot. The context dry run never loads a model.
|
||||
|
||||
No schema change: the accounting lives in the snapshot JSON, and a turn from
|
||||
v1.0.0 simply has none. The bundle format is unchanged, for the same reason.
|
||||
|
||||
### 15.3 The history window moves in blocks (post-M11)
|
||||
|
||||
§15.2 makes the window a ceiling. This is about what happens at that ceiling.
|
||||
@@ -1300,6 +1369,53 @@ headings. A lone heading followed by prose stays, and so do JSON a character
|
||||
typed and a fact restated inside a sentence. That last case is how a narrator can
|
||||
still carry authoritative state into its prose (M11 report §G.4 and §P).
|
||||
|
||||
**As implemented (v1.1 WP-A2): the source and the sink together.** The M11
|
||||
closeout's identity run stored four shapes the extractor left. All of them were
|
||||
application text. v1.1 changes both the prompt that taught them and the
|
||||
extractor that missed them. Every new removal is anchored to something the
|
||||
application owns, never to what prose looks like.
|
||||
|
||||
- **The prompt.**
|
||||
- `events.vocabulary_for_prompt` shows each event as the object the model
|
||||
must send (`{"type": "set_possession", "item": "<key>", "owner": "<key>"}`),
|
||||
not as `set_possession(item, owner)`. The call notation was never the wire
|
||||
format, and the narrator copied it.
|
||||
- `EMIT_RULE`'s example uses the placeholders `character-1`, `item-1` and
|
||||
`location-1`, not the fantasy fixture's `mara`, `silver-key`, `old-abbey` and
|
||||
`aldric`. The narrator had proposed `silver-key` in an office meeting.
|
||||
- The length hint's opening and closing words are named constants shared by
|
||||
the builder and the extractor.
|
||||
- **The extractor**, rules R1-R4 (`RULE_*` in `narrative/extract.py`):
|
||||
- **R1:** a whole line that begins with a call to an event in `events.SPECS`,
|
||||
optionally `>`-quoted. Not inside a fenced code block, not mid-sentence, and
|
||||
not for a call-shaped name the vocabulary lacks.
|
||||
- **R2:** a trailing bracket that opens `Hard limit:` and carries the hint's
|
||||
own wording ("append the state block", or "turn must not exceed *N* words").
|
||||
- **R3:** the renderer's scene line left as the reply's last line. It is
|
||||
removed when it ends in the renderer's `(at <location>)`, or when protocol
|
||||
was already cut from the same reply.
|
||||
- **R4:** a ```` ```json ```` or bare ```` ``` ```` opener left as the last
|
||||
line with nothing after it, counted as an opener rather than a closer.
|
||||
- **Proven on real narration.** Every stored real reply in the v1 evidence was
|
||||
replayed through the v1.0.0 and v1.1 extractors (`tools/v11_replay_extractor.py`).
|
||||
Every changed line is attributed to one of the four rules, and a person reviewed
|
||||
every change. The results are in the WP-A1/A2 report.
|
||||
- **Deliberately still left:** a fact restated inside a sentence; model-invented
|
||||
headings; and a bracket that starts `Hard limit:` but carries none of the
|
||||
application's wording.
|
||||
- **R5, the echoed instruction tail (v1.1 A2 corrective).** A v1.1 identity turn
|
||||
ended in a reworded continue hint, "[… Continue the story here, directly.
|
||||
Output only story text.]". Because nothing recognised it, nothing above it was
|
||||
ever trailing, and the reminder, a reworded length hint and a scene line all
|
||||
stayed. R5 makes two changes:
|
||||
- **The continue hint is recognised by its own sentence.** A trailing bracket
|
||||
containing "Output only story text" is an echoed instruction.
|
||||
`CONTINUE_HINT_PHRASE` is pinned by a test to `CHAT_CONTINUE_HINT`.
|
||||
- **A bracket opening with the length hint's own `Hard limit:` is removed only
|
||||
directly above an echoed instruction already cut from the same reply's end.**
|
||||
An in-world "[Hard limit: forty days]" stays when it is the last line, and
|
||||
when a state block follows it. Any other bracket above an echo stays.
|
||||
|
||||
## 16. Database Direction
|
||||
|
||||
SQLite remains the selected v1 authoritative store.
|
||||
|
||||
@@ -1,8 +1,12 @@
|
||||
# Adventure Storyteller — v1.1 Plan
|
||||
|
||||
**Status:** PLANNING. Written 2026-09-14 on `v1.1-development`, from the signed
|
||||
v1.0.0 release commit `432f041`. No work package has started, and no v1.1
|
||||
version or tag exists.
|
||||
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
|
||||
signed v1.0.0 release commit `432f041`.
|
||||
|
||||
**WP-A1 and WP-A2 are implemented and staged for owner review** (2026-09-14). They
|
||||
are reported together, and kept separate, in
|
||||
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`. No other work package has started. No
|
||||
v1.1 version or tag exists.
|
||||
|
||||
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
|
||||
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
|
||||
|
||||
+28
-2
@@ -1,8 +1,34 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v4.1
|
||||
- **Package:** Adventure Storyteller Planning Package v4.2
|
||||
- **Revision date:** 2026-09-14
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 planning has begun** on `v1.1-development`: `V1.1-PLAN.md`. No v1.1 work package has started, and no v1.1 version or tag exists.
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): **WP-A1 and WP-A2 are implemented and staged for owner review.** No other work package has started, and no v1.1 version or tag exists.
|
||||
|
||||
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
|
||||
|
||||
Two v1.1 work packages, implemented in sequence. No requirement or acceptance
|
||||
test changed, and no schema or bundle format changed. Evidence is in
|
||||
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `TECHNICAL-DESIGN.md` §15.2 | **As implemented (v1.1 WP-A1)**, covering: the safety reserve, `max(256, ceil(5%))`; what replaced M6's 64-token margin; the server's own count; the four accounting states; and keeping the turn. | as-implemented record |
|
||||
| `TECHNICAL-DESIGN.md` §15.4 | **As implemented (v1.1 WP-A2)**: the vocabulary shown in the wire format, genre-neutral placeholders, and extractor rules R1-R4, each anchored to application-owned text. | as-implemented record |
|
||||
| `DECISIONS/013-authoritative-narrative-state-document.md` | An implementation note for v1.1. No event type, field, validation rule or proposal record changed. | as-implemented note |
|
||||
| `V1.1-PLAN.md` | Status: A1 and A2 implemented and staged. | status |
|
||||
| `planning/README.md` | Current state. | index |
|
||||
| `reports/v1.1/V1.1-WP-A1-A2-REPORT.md` | **New.** The combined review package, with A1 and A2 kept separate. | work-package report |
|
||||
| `README.md`, `DEVELOPMENT.md` | The reserve and the accounting. The stale backend test count and the Screenshots paragraph are corrected. | developer docs |
|
||||
|
||||
**Corrective work before commit (owner review, 2026-09-14).** Report §R.
|
||||
|
||||
| Document | Change |
|
||||
| --- | --- |
|
||||
| `TECHNICAL-DESIGN.md` §15.2 | A cold model is loaded once before its turn is built (`contextwindow.ensure_window`). |
|
||||
| `TECHNICAL-DESIGN.md` §15.4 | R5: the echoed continue hint is recognised by its own sentence, and the application-opened tail above it is removed. |
|
||||
| `DEVELOPMENT.md`, `README.md` | The cold-model load, in operator terms. |
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14)
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user