v1.1: harden context window and narrator protocol boundary

WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-14 16:35:05 -04:00
co-authored by Claude Opus 5
parent ac465ed867
commit d63804f22e
26 changed files with 3903 additions and 61 deletions
+21 -4
View File
@@ -246,11 +246,28 @@ def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
assert report_before["history"]["floor_depth"] is not None, (
"this fixture is meant to be over budget; trimming never engaged")
_play_one_more(db, adventure, 60)
_, after, report_after = _builder.build_context(adventure, settings)
# v1.1 WP-A1: the fixture used to be positioned so that the very next turn
# held the floor. The safety reserve takes 256 tokens of this 2,048 budget,
# the block is now the minimum of two, and the next turn is a step. So walk
# forward until a turn holds, requiring every move on the way to be exactly
# one block: a window that slides by one action every turn fails either way.
held = None
depth = 60
for _ in range(4):
_play_one_more(db, adventure, depth)
depth += 1
_, after, report_after = _builder.build_context(adventure, settings)
floor_before = report_before["history"]["floor_depth"]
floor_after = report_after["history"]["floor_depth"]
block = report_after["history"]["trim_block"]
assert floor_after - floor_before in (0, block), (floor_before, floor_after, block)
if floor_after == floor_before:
held = (before, after)
break
before, report_before = after, report_after
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"]
assert _shared_prefix(before, after) > 0.85
assert held is not None, "the floor never held across a turn"
assert _shared_prefix(*held) > 0.85
def test_without_a_stable_floor_the_prefix_collapses(saturated):
+348
View File
@@ -0,0 +1,348 @@
"""v1.1 WP-A1 corrective: a cold model is loaded, not guessed about.
A1's accounting caught a real cold-model turn: `/api/ps` knew nothing because the
model was not resident, `/api/show` found no `num_ctx`, the window was therefore
unverified, and the prompt was built to the configured 16,384. Ollama loaded the
model at its own 4,096 default, read 2,050 of the 13,875 tokens and answered 200.
Detection was right. The case is also preventable: once the model is loaded its
window is readable. So before an unverified turn is assembled, the application
asks the configured server, once, to load the model (`POST /api/generate` with a
model and no prompt, which Ollama answers with `"done_reason": "load"` and no
text), probes again, and builds the turn to whatever that probe says. A window
still unverified afterwards changes nothing: the configured budget stands and
the post-response accounting still watches for a cut prompt.
python -m pytest tests/test_v11_cold_window.py -v
"""
import asyncio
import httpx
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy import text
from sqlalchemy.orm import undefer
from app import auth, contextwindow, limits, models
from app.context import builder
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.providers.base import ProviderError
from app.routers import adventures
from fakes import ScriptedProvider
ENDPOINT = "http://127.0.0.1:11434/v1"
MODEL = "qwen2.5:3b-instruct"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
class ColdOllama:
"""The shapes a real Ollama 0.33 returned, with a model that starts cold.
`/api/ps` lists only loaded models. `/api/show` carries no `num_ctx`.
`/api/generate` with no prompt loads the model at `load_window`, exactly as
the real server answered: HTTP 200, `"response": ""`, `"done_reason": "load"`.
"""
def __init__(self, *, loaded=None, load_window=4096, generate_status=200,
report_after_load=True):
self.loaded = dict(loaded or {})
self.load_window = load_window
self.generate_status = generate_status
self.report_after_load = report_after_load
self.requests: list[tuple[str, str, dict | None]] = []
def handler(self, request: httpx.Request) -> httpx.Response:
body = None
if request.content:
import json
body = json.loads(request.content)
self.requests.append((request.method, str(request.url), body))
path = request.url.path
if path == "/api/ps":
return httpx.Response(200, json={"models": [
{"name": name, "model": name, "context_length": tokens}
for name, tokens in self.loaded.items()
]})
if path == "/api/show":
return httpx.Response(200, json={
"model_info": {"qwen2.context_length": 32768}, "parameters": ""})
if path == "/api/generate":
if self.generate_status != 200:
return httpx.Response(self.generate_status, json={"error": "model not found"})
if self.report_after_load:
self.loaded[body["model"]] = self.load_window
return httpx.Response(200, json={
"model": body["model"], "response": "", "done": True, "done_reason": "load"})
return httpx.Response(404)
def paths(self):
return [httpx.URL(url).path for _method, url, _body in self.requests]
@pytest.fixture()
def server(monkeypatch):
def install(fake: ColdOllama):
original = httpx.AsyncClient
def build(*args, **kwargs):
kwargs.pop("verify", None)
return original(*args, transport=httpx.MockTransport(fake.handler), **kwargs)
monkeypatch.setattr(contextwindow.httpx, "AsyncClient", build)
return fake
return install
# ------------------------------------------------------------ ensure_window
def test_a_cold_model_is_loaded_once_and_its_window_verified(server):
fake = server(ColdOllama(loaded={}, load_window=4096))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert (window.tokens, window.source, window.verified) == (4096, contextwindow.LOADED, True)
assert preflight == {"attempted": True, "loaded": True, "verified_before": False,
"verified_after": True,
"detail": "the server loaded the model (load)"}
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate", "/api/ps"]
# One load request, naming the model and nothing else: no prompt, so no text.
warms = [body for _m, url, body in fake.requests if url.endswith("/api/generate")]
assert warms == [{"model": MODEL}]
def test_an_already_loaded_model_is_not_warmed(server):
fake = server(ColdOllama(loaded={MODEL: 16384}))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert window.verified and window.tokens == 16384
assert preflight["attempted"] is False
assert "/api/generate" not in fake.paths()
def test_a_model_that_loads_but_still_cannot_be_read_stays_unverified(server):
"""A server that loads the model but whose `/api/ps` still cannot say. The
existing unknown path stands: no guessed window, the configured budget kept."""
fake = server(ColdOllama(loaded={}, report_after_load=False))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert not window.verified and window.tokens is None
assert preflight["attempted"] is True and preflight["loaded"] is True
assert preflight["verified_after"] is False
assert fake.paths().count("/api/generate") == 1
assert contextwindow.effective_budget(16384, window) == 16384
@pytest.mark.parametrize("status", [404, 500])
def test_a_failed_load_is_recorded_and_leaves_the_window_unverified(server, status):
fake = server(ColdOllama(loaded={}, generate_status=status))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert not window.verified
assert preflight["attempted"] is True and preflight["loaded"] is False
assert f"HTTP {status}" in preflight["detail"]
# Bounded: one attempt, and no second probe after a failed load.
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate"]
def test_a_declared_window_does_not_stop_the_server_being_asked(server):
"""A declaration fills a hole the server leaves. Loading the model can close
the hole, and a verified answer always wins over a declaration."""
server(ColdOllama(loaded={}, load_window=4096))
window, _preflight = asyncio.run(
contextwindow.ensure_window(ENDPOINT, MODEL, declared=8192))
assert (window.tokens, window.source) == (4096, contextwindow.LOADED)
def test_an_unreachable_server_is_not_asked_to_load_anything():
"""Nothing listens here. No load is attempted against a server that did not
answer the probe, so an offline turn costs no second timeout."""
window, preflight = asyncio.run(
contextwindow.ensure_window("http://127.0.0.1:1/v1", MODEL))
assert not window.verified
assert preflight["attempted"] is False
assert "did not answer" in preflight["detail"]
def test_the_load_request_obeys_the_endpoint_policy():
"""ADR 011. No transport is installed: a load that ignored the policy would
try to reach a public address for real."""
for url in ("https://api.openai.com/v1", "http://8.8.8.8:11434/v1"):
loaded, detail = asyncio.run(contextwindow.warm(url, MODEL, timeout=2))
assert loaded is False
assert "not allowed" in detail
def test_the_load_request_goes_only_to_the_configured_host(server):
fake = server(ColdOllama(loaded={}))
asyncio.run(contextwindow.ensure_window("http://192.168.0.50:11434/v1", MODEL))
hosts = {httpx.URL(url).host for _m, url, _b in fake.requests}
ports = {httpx.URL(url).port for _m, url, _b in fake.requests}
assert hosts == {"192.168.0.50"} and ports == {11434}
# ------------------------------------------------------------- end to end
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="v11cold@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
context_token_budget=16384, max_output_tokens=500,
))
adventure = models.Adventure(
user_id=user.id, title="Cold",
campaign_canon={"rules": ["The sealed crypt is named CANON-SENTINEL-COLD-2050."]},
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(adventure_id=adventure.id, type="start", text="Rain."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
def _long_story(adv_id, turns=120):
from app import tree
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv_id)
for i in range(turns):
for kind, body in (
("do", f"I search the {i}th chamber of the undercroft."),
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
):
action = models.Action(adventure_id=adv_id, type=kind, text=body)
db.add(action)
db.flush()
tree.place_action(db, adventure, action)
db.commit()
def _counts():
with SessionLocal() as db:
return {table: db.execute(text(f"SELECT COUNT(*) FROM {table}")).scalar()
for table in ("actions", "state_events", "state_proposals", "memories",
"summaries")}
def _latest_ai(adv_id):
with SessionLocal() as db:
return (db.query(models.Action)
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
def test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports(client, server):
"""The observed failure, prevented. Without the load this turn would be built
to the configured 16,384 against a 4,096 server."""
_long_story(client.adv_id)
fake = server(ColdOllama(loaded={}, load_window=4096))
ScriptedProvider.replies = ["The seal holds."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is True
assert snapshot["tokens"]["budget"] == 4096
assert snapshot["window"]["preflight"]["attempted"] is True
assert snapshot["window"]["preflight"]["verified_after"] is True
system, story = ScriptedProvider.prompts[-1]
sent = builder.count_tokens(system) + builder.count_tokens(story)
assert sent + snapshot["tokens"]["transport"] + 500 + 256 <= 4096
assert "CANON-SENTINEL-COLD-2050" in system
assert fake.paths().count("/api/generate") == 1
def test_without_the_load_the_same_cold_turn_would_have_been_built_too_large(client, server,
monkeypatch):
"""The negative control: v1.0.0 and the first A1 tree probed only."""
_long_story(client.adv_id)
server(ColdOllama(loaded={}, load_window=4096))
async def probe_only(endpoint_url, model, *, declared=None, warm_timeout=300.0):
window = await contextwindow.probe(endpoint_url, model, declared=declared)
return window, {"attempted": False}
monkeypatch.setattr(adventures.turns.contextwindow, "ensure_window", probe_only)
ScriptedProvider.replies = ["The seal holds."]
client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is False
assert snapshot["tokens"]["budget"] == 16384
system, story = ScriptedProvider.prompts[-1]
assert builder.count_tokens(system) + builder.count_tokens(story) > 4096 * 2
def test_the_load_itself_writes_nothing(client, server):
"""No action, narration, state event, proposal, memory or summary comes from
the preflight: it is a request to the server and nothing else."""
fake = server(ColdOllama(loaded={}))
before = _counts()
asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert _counts() == before
assert fake.paths().count("/api/generate") == 1
def test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe(client, server):
"""The ordinary failure semantics: the error is reported, no narration is
accepted, and nothing about the state changes."""
server(ColdOllama(loaded={}, generate_status=404))
before = _counts()
ScriptedProvider.replies = [ProviderError("Endpoint or model not found (HTTP 404).")]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the door"})
assert response.status_code == 200
assert '"type": "error"' in response.text or '"error"' in response.text
after = _counts()
assert after["state_events"] == before["state_events"]
assert after["state_proposals"] == before["state_proposals"]
with SessionLocal() as db:
assert db.query(models.Action).filter_by(adventure_id=client.adv_id,
type="ai").count() == 0
def test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer(client, server):
server(ColdOllama(loaded={}, generate_status=500))
ScriptedProvider.replies = ["The door opens."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the door"})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is False
assert snapshot["window"]["preflight"]["loaded"] is False
assert snapshot["tokens"]["budget"] == 16384
assert snapshot["accounting"]["status"] == contextwindow.UNKNOWN
def test_the_context_dry_run_never_loads_a_model(client, server):
fake = server(ColdOllama(loaded={}))
response = client.get(f"/api/adventures/{client.adv_id}/context")
assert response.status_code == 200
assert "/api/generate" not in fake.paths()
+495
View File
@@ -0,0 +1,495 @@
"""v1.1 WP-A1: a deliberate safety reserve, and a turn the server cut is not silent.
M11 made the verified window a ceiling. It did not make the application's count
the server's count. The application counts with `cl100k_base`, the narrator with
its own tokenizer, and the v1 evidence left 23-42 real tokens between the largest
prompt and the edge of a 16,384 window. Past that edge Ollama does not refuse.
Measured against the reference CPU host (Ollama 0.33, a 4,096 window), a
6,316-token prompt came back 200 with `prompt_tokens` 2,050: the front of the
prompt, which in this design is the narrator's rules and the canon, was gone.
So the tests below are in three halves.
**The reserve.** `max(256, ceil(5% of the effective window))`, taken from the
budget before any history is chosen, on top of an exact reply allocation.
**The arithmetic.** The assembled prompt, plus the application text the provider
adds to every request, plus the reply allocation, plus the reserve, fits the
effective window. Protected context that cannot fit that way fails before the
model is called.
**The accounting.** Where the server reports how many prompt tokens it read, the
turn records `fits`, `exceeded` or `truncation_suspected`. Where it reports
nothing, the turn says `unknown`, never `fits`. A discrepancy found after the
reply is recorded and shown; it never costs the reader an accepted turn.
python -m pytest tests/test_v11_context_reserve.py -v
"""
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, contextwindow, limits, models
from app.context import builder
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.providers.openai_compatible import CHAT_CONTINUE_HINT, OpenAICompatibleProvider
from app.routers import adventures
from fakes import ScriptedProvider
ENDPOINT = "http://127.0.0.1:11434/v1"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
# ------------------------------------------------------------- the reserve
@pytest.mark.parametrize("window, reserve", [
(1024, 256),
(4096, 256), # 5% is 204.8, so the floor holds
(5120, 256), # exactly 5% is the floor
(5121, 257), # 256.05 rounds up
(8192, 410), # 409.6 rounds up
(16384, 820), # 819.2 rounds up
(32768, 1639), # 1638.4 rounds up
])
def test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up(window, reserve):
assert contextwindow.safety_reserve(window) == reserve
def test_the_reserve_is_far_larger_than_the_v1_margin_at_the_evidence_window():
"""The v1 evidence left 23-42 tokens at 16,384. 64 tokens of slack was all
the arithmetic kept for drift and separators together."""
assert contextwindow.safety_reserve(16384) >= 10 * 64
# ---------------------------------------------------------- the arithmetic
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="v11reserve@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model="qwen2.5:3b-instruct", endpoint_url=ENDPOINT,
embedding_model="", context_token_budget=16384, max_output_tokens=500,
))
adventure = models.Adventure(
user_id=user.id, title="Reserved",
campaign_canon={"rules": [
"The abbey seal has never been broken.",
"The sealed crypt is named CANON-SENTINEL-RESERVE-5120.",
]},
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(
adventure_id=adventure.id, type="start", text="Rain over Westhaven."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
def _long_story(adv_id, turns=120):
from app import tree
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv_id)
for i in range(turns):
for kind, text in (
("do", f"I search the {i}th chamber of the undercroft."),
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
):
action = models.Action(adventure_id=adv_id, type=kind, text=text)
db.add(action)
db.flush()
tree.place_action(db, adventure, action)
db.commit()
def _settings(**changes):
with SessionLocal() as db:
settings = db.query(models.Settings).first()
for key, value in changes.items():
setattr(settings, key, value)
db.commit()
def _build(client, window):
with SessionLocal() as db:
adventure = db.get(models.Adventure, client.adv_id)
settings = db.query(models.Settings).first()
return builder.build_context(adventure, settings, window=window)
def _sent(system, story) -> int:
"""What the provider actually sends in chat mode, by the application's count."""
return (builder.count_tokens(system) + builder.count_tokens(story)
+ builder.count_tokens(CHAT_CONTINUE_HINT))
CONFIGURATIONS = {
"verified 4,096": (dict(context_token_budget=16384),
contextwindow.Window(4096, contextwindow.LOADED)),
"verified 8,192": (dict(context_token_budget=16384),
contextwindow.Window(8192, contextwindow.PARAMETERS)),
"verified 16,384": (dict(context_token_budget=16384),
contextwindow.Window(16384, contextwindow.LOADED)),
"declared 6,000": (dict(context_token_budget=16384),
contextwindow.Window(6000, contextwindow.DECLARED)),
"unverified, configured 12,000": (dict(context_token_budget=12000),
contextwindow.UNVERIFIED),
}
@pytest.mark.parametrize("name", list(CONFIGURATIONS))
def test_the_prompt_leaves_the_reply_and_the_reserve_free(client, name):
"""A1-2, on the assembled text rather than the builder's own arithmetic."""
changes, window = CONFIGURATIONS[name]
_settings(**changes)
_long_story(client.adv_id, turns=120)
system, story, report = _build(client, window)
tokens = report["tokens"]
budget = tokens["budget"]
assert tokens["safety_reserve"] == contextwindow.safety_reserve(budget)
assert tokens["output_reserve"] == 500
sent = _sent(system, story)
assert sent + tokens["output_reserve"] + tokens["safety_reserve"] <= budget, (
name, sent, tokens)
# The history is what gave way, not the canon.
assert "CANON-SENTINEL-RESERVE-5120" in system
assert report["history"]["included"] < report["history"]["total"]
def test_the_report_prices_the_text_the_provider_adds(client):
"""The chat hint rides on every request and was never counted."""
_, _, report = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
tokens = report["tokens"]
assert tokens["transport"] >= builder.count_tokens(CHAT_CONTINUE_HINT)
assert tokens["estimate"] == tokens["total"] + builder.count_tokens(CHAT_CONTINUE_HINT)
def test_the_reserve_follows_the_effective_window_not_the_setting(client):
"""5% of a 4,096 server, not 5% of a 16,384 setting it will never read."""
_, _, capped = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
_, _, full = _build(client, contextwindow.Window(16384, contextwindow.LOADED))
assert capped["tokens"]["safety_reserve"] == 256
assert full["tokens"]["safety_reserve"] == 820
def test_protected_context_that_only_fits_without_the_reserve_fails_explicitly(client):
"""A1-3 in the builder. Before v1.1 this prompt would have been built.
The canon is sized so that protected text plus the reply fits a 4,096 window
with room to spare, and does not fit once the 256-token reserve is taken.
"""
small = contextwindow.Window(4096, contextwindow.LOADED)
# Measured with a window large enough never to overflow, because repeated
# text merges tokens at its seams and cannot be priced by multiplication.
roomy = contextwindow.Window(32768, contextwindow.LOADED)
rules = None
with SessionLocal() as db:
adventure = db.get(models.Adventure, client.adv_id)
settings = db.query(models.Settings).first()
base_rules = list(adventure.campaign_canon["rules"])
filler = "The bell tolls once for every name in the ledger."
copies = 1
while True:
candidate = base_rules + [" ".join([filler] * copies)]
adventure.campaign_canon = {"rules": candidate}
_, _, measured = builder.build_context(adventure, settings, window=roomy)
t = measured["tokens"]
# What protected context costs at 4,096, without the reserve.
without_reserve = t["protected"] + t["transport"] + t["output_reserve"]
if without_reserve + 64 >= 4096 - 60:
break
copies += 1
rules = candidate
adventure.campaign_canon = {"rules": rules}
db.commit()
# The case this test is about: v1's arithmetic, with its 64-token margin,
# would have built this prompt. v1.1's reserve does not fit.
assert without_reserve + 64 < 4096
assert without_reserve + contextwindow.safety_reserve(4096) >= 4096
with pytest.raises(builder.ContextOverflow) as caught:
builder.build_context(adventure, settings, window=small)
message = str(caught.value)
assert "safety" in message
assert "load the model with a larger window" in message
def test_an_overflowing_turn_never_reaches_the_model(client, monkeypatch):
"""A1-3 end to end: the refusal happens before the provider is called."""
async def verified(endpoint, model, declared=None, use_cache=True):
return contextwindow.Window(1024, contextwindow.LOADED)
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
ScriptedProvider.replies = ["This must never be generated."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the crypt"})
assert response.status_code == 200
assert "safety" in response.text
assert ScriptedProvider.calls == 0
with SessionLocal() as db:
assert db.query(models.Action).filter_by(
adventure_id=client.adv_id, type="ai").count() == 0
# ---------------------------------------------------------- the accounting
def _classify(prompt_tokens=None, *, usage=None, estimate=3500, budget=4096,
output=500, verified=True):
if usage is None and prompt_tokens is not None:
usage = {"prompt_tokens": prompt_tokens, "completion_tokens": 40}
return contextwindow.classify_usage(
usage, estimate=estimate, budget=budget, max_output_tokens=output,
window_verified=verified,
)
def test_a_prompt_the_server_read_in_full_fits():
# The 13-token chat-template overhead measured against the real server.
result = _classify(3513)
assert result["status"] == contextwindow.FITS
assert result["server_prompt_tokens"] == 3513
assert result["difference"] == 13
assert result["safety_reserve"] == 256
assert result["observed_margin"] == 4096 - 500 - 3513
def test_a_server_that_counts_more_than_the_reserve_allows_is_exceeded():
"""The prompt plus the reply allocation no longer fits the window."""
result = _classify(3700)
assert result["status"] == contextwindow.EXCEEDED
assert result["observed_margin"] < 0
def test_a_server_that_read_far_less_than_was_sent_is_suspected_of_truncating():
"""The real shape: 6,316 sent, 2,050 read, HTTP 200, no error."""
result = _classify(2050, estimate=6316)
assert result["status"] == contextwindow.TRUNCATION_SUSPECTED
assert result["difference"] == 2050 - 6316
def test_a_small_undercount_is_tokenizer_drift_not_truncation():
"""A tokenizer thriftier than `cl100k_base` reads fewer tokens honestly. Only
a shortfall larger than the reserve is called truncation."""
assert _classify(3500 - 255)["status"] == contextwindow.FITS
assert _classify(3500 - 257)["status"] == contextwindow.TRUNCATION_SUSPECTED
@pytest.mark.parametrize("usage", [
None,
{},
{"completion_tokens": 40},
{"prompt_tokens": 0},
{"prompt_tokens": "3500"},
{"prompt_tokens": -1},
])
def test_no_usable_count_is_unknown_never_fits(usage):
result = _classify(usage=usage)
assert result["status"] == contextwindow.UNKNOWN
assert result["server_prompt_tokens"] is None
assert result["observed_margin"] is None
def test_the_accounting_says_when_the_window_itself_was_not_verified():
result = _classify(3513, verified=False)
assert result["status"] == contextwindow.FITS
assert result["window_verified"] is False
assert "not verified" in result["detail"]
def test_the_stream_asks_the_server_to_report_its_usage():
"""Measured: Ollama 0.33 sends no usage in a stream unless asked."""
provider = OpenAICompatibleProvider(ENDPOINT, "m")
from app.providers.base import PromptParts
for mode in ("chat", "completion"):
provider.api_mode = mode
_url, body = provider._request(PromptParts(system="s", story="t"), 0.7, 50)
assert body["stream"] is True
assert body["stream_options"] == {"include_usage": True}
def _latest_ai(adv_id):
with SessionLocal() as db:
return (
db.query(models.Action)
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first()
)
def _play(client, monkeypatch, usage, window=4096, reply="The crypt is still sealed."):
async def verified(endpoint, model, declared=None, use_cache=True):
return contextwindow.Window(window, contextwindow.LOADED, 32768, "fake")
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
monkeypatch.setattr(ScriptedProvider, "last_usage", usage)
ScriptedProvider.replies = [reply]
return client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
def test_a_turn_records_what_the_server_read(client, monkeypatch):
response = _play(client, monkeypatch, None)
assert response.status_code == 200, response.text[:300]
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
response = _play(client, monkeypatch,
{"prompt_tokens": estimate + 13, "completion_tokens": 9})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
accounting = snapshot["accounting"]
assert accounting["status"] == contextwindow.FITS
assert accounting["server_prompt_tokens"] == estimate + 13
assert accounting["estimate"] == snapshot["tokens"]["estimate"]
assert '"accounting"' in response.text
assert contextwindow.FITS in response.text
def test_a_turn_with_no_reported_usage_is_unknown(client, monkeypatch):
response = _play(client, monkeypatch, None)
assert response.status_code == 200
assert _latest_ai(client.adv_id).context_snapshot["accounting"]["status"] == (
contextwindow.UNKNOWN)
def test_a_suspected_truncation_keeps_the_turn_and_says_so(client, monkeypatch, caplog):
"""A1-7 and A1-8. The reader watched the narration arrive; it stays."""
response = _play(client, monkeypatch, {"prompt_tokens": 12, "completion_tokens": 9},
reply="The seal holds, and the rain goes on.")
assert response.status_code == 200, response.text[:300]
action = _latest_ai(client.adv_id)
assert action is not None
assert action.text == "The seal holds, and the rain goes on."
accounting = action.context_snapshot["accounting"]
assert accounting["status"] == contextwindow.TRUNCATION_SUSPECTED
assert contextwindow.TRUNCATION_SUSPECTED in response.text
assert any(contextwindow.TRUNCATION_SUSPECTED in r.getMessage() for r in caplog.records)
# Inspectable afterwards through the same route the context panel reads.
context = client.get(
f"/api/adventures/{client.adv_id}/actions/{action.id}/context")
assert context.status_code == 200
assert context.json()["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
def test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves():
"""Found by the A2 long run. Accounting belongs to one API call, not to the
turn's shared prompt. A retry demotes the old attempt, and a take selection
hands the prompt from one attempt to another. Neither may drop an attempt's
accounting or give it another attempt's."""
from app import attempts
class Node:
def __init__(self, snapshot):
self.context_snapshot = snapshot
shared = {"tokens": {"estimate": 3000}, "sections": [], "window": {"verified": True}}
first = Node(shared | {"raw_output": "one", "usage": {"prompt_tokens": 3015},
"accounting": {"status": contextwindow.FITS, "server_prompt_tokens": 3015}})
second = Node({"raw_output": "two", "usage": {"prompt_tokens": 12},
"accounting": {"status": contextwindow.TRUNCATION_SUSPECTED,
"server_prompt_tokens": 12}})
# Superseded by a retry: the old attempt keeps only its own slices.
attempts.keep_own_slices(Node(dict(first.context_snapshot)))
demoted = Node(dict(first.context_snapshot))
attempts.keep_own_slices(demoted)
assert demoted.context_snapshot["accounting"]["server_prompt_tokens"] == 3015
assert "tokens" not in demoted.context_snapshot
# The prompt moves to the second attempt; each keeps its own accounting.
attempts.hand_over_the_prompt(first, second)
assert second.context_snapshot["tokens"] == {"estimate": 3000}
assert second.context_snapshot["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
assert second.context_snapshot["accounting"]["server_prompt_tokens"] == 12
assert first.context_snapshot["accounting"]["status"] == contextwindow.FITS
assert "tokens" not in first.context_snapshot
def test_a_retry_leaves_each_take_with_its_own_accounting(client, monkeypatch):
"""End to end, through the real retry route. Before the fix the live take
inherited the superseded take's accounting, so the inspector could show one
call's server count as another's."""
response = _play(client, monkeypatch, None, reply="The first take.")
assert response.status_code == 200, response.text[:300]
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
# Replay the first take with a real count, so it has accounting of its own.
with SessionLocal() as db:
first = (db.query(models.Action)
.filter(models.Action.adventure_id == client.adv_id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot)).first())
snapshot = dict(first.context_snapshot)
snapshot["accounting"] = contextwindow.classify_usage(
{"prompt_tokens": estimate + 15}, estimate=estimate, budget=4096,
max_output_tokens=500, window_verified=True)
first.context_snapshot = snapshot
db.commit()
first_id = first.id
monkeypatch.setattr(ScriptedProvider, "last_usage",
{"prompt_tokens": 12, "completion_tokens": 9})
ScriptedProvider.replies = ["The second take."]
retried = client.post(f"/api/adventures/{client.adv_id}/retry")
assert retried.status_code == 200, retried.text[:300]
with SessionLocal() as db:
rows = (db.query(models.Action)
.filter(models.Action.adventure_id == client.adv_id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id).all())
by_id = {row.id: row for row in rows}
old = by_id[first_id]
new = [row for row in rows if row.id != first_id][-1]
assert new.text == "The second take."
assert new.live and not old.live
# The superseded take keeps its own accounting and gives up the prompt.
assert old.context_snapshot["accounting"]["status"] == contextwindow.FITS
assert old.context_snapshot["accounting"]["server_prompt_tokens"] == estimate + 15
assert "tokens" not in old.context_snapshot
# The live take carries the prompt and its own accounting, not the old one's.
assert "tokens" in new.context_snapshot
assert new.context_snapshot["accounting"]["status"] == (
contextwindow.TRUNCATION_SUSPECTED)
assert new.context_snapshot["accounting"]["server_prompt_tokens"] == 12
def test_an_exceeded_turn_is_also_kept(client, monkeypatch):
response = _play(client, monkeypatch, {"prompt_tokens": 3900, "completion_tokens": 9})
assert response.status_code == 200
action = _latest_ai(client.adv_id)
assert action.text == "The crypt is still sealed."
assert action.context_snapshot["accounting"]["status"] == contextwindow.EXCEEDED
+324
View File
@@ -0,0 +1,324 @@
"""v1.1 WP-A2: the protocol a narrator copies stays out of the story, and nothing else does.
The M11 closeout's identity run (an office meeting, a 3B narrator, a 4,096
window) stored four turns carrying text the application wrote, not the story:
- `> Create_entity(new_person, "john", …` — the event vocabulary as the prompt
printed it, `name(field, …)`, copied as if it were a call;
- `> set_possession(silver-key, "alice") Adds the silver key to Alice's
possession.` — the call again, naming the fantasy example slug from the fixed
state rule, in a meeting room;
- `Scene: Bill, Alice, … (at The meeting room)` — the renderer's own scene line;
- `[Hard limit: your next turn must not exceed 180 words, … append the state
block well inside the limit.]` — the length hint, reworded at the front and
verbatim at the end.
v1.0.0 removed none of them. The rule this module is held to is unchanged from
M5: **removing story is worse than leaving protocol.** Every removal below is
anchored to a string or a vocabulary the application owns, and every one has
story beside it that must survive.
python -m pytest tests/test_v11_protocol_echo.py -v
"""
import json
import re
import pytest
from app.context import builder
from app.narrative import events, extract, render
# ---------------------------------------------------------------- the prompt
#: The identifiers the v1 state rule taught every campaign, from the fantasy
#: acceptance fixture. None may come back into a fixed instruction.
FANTASY_IDENTIFIERS = ("mara", "silver-key", "silver key", "old-abbey", "abbey",
"aldric", "westhaven", "crypt", "edrin")
#: And nothing from the science-fiction fixture either: neutral means neutral,
#: not "the other genre".
SCIFI_IDENTIFIERS = ("persephone", "imani", "data-crystal", "data crystal", "airlock")
def _fixed_instructions() -> str:
return "\n".join([
extract.EMIT_RULE,
extract.EMIT_REMINDER,
events.vocabulary_for_prompt(),
builder.length_hint(500),
builder.length_hint(500, "brief"),
builder.length_hint(500, "long"),
builder.length_hint(120),
]).lower()
@pytest.mark.parametrize("identifier", FANTASY_IDENTIFIERS + SCIFI_IDENTIFIERS)
def test_the_fixed_state_instructions_name_no_fixture_identifier(identifier):
"""A2-1. The example slug the office run copied cannot come back."""
assert not re.search(rf"\b{re.escape(identifier)}\b", _fixed_instructions())
def test_the_example_uses_neutral_identifiers():
for neutral in ("character-1", "item-1", "location-1"):
assert neutral in extract.EMIT_RULE
def test_the_worked_example_is_a_block_this_extractor_accepts():
"""The example is the wire format, byte for byte, not an illustration of it."""
example = extract.EMIT_RULE[extract.EMIT_RULE.index("```state"):]
prose, parsed, _raw = extract.split("The door opens.\n\n" + example)
assert prose == "The door opens."
assert isinstance(parsed, dict)
assert [e["type"] for e in parsed["events"]] == [
"set_possession", "set_current_location"]
for event in parsed["events"]:
assert events.is_allowed(event["type"])
def test_the_vocabulary_is_not_written_as_function_calls():
"""A2-2. `set_possession(item, owner)` is the notation the narrator copied."""
vocabulary = events.vocabulary_for_prompt()
for name in events.SPECS:
assert not re.search(rf"\b{name}\s*\(", vocabulary), name
def test_every_event_is_described_in_the_shape_the_model_must_send():
lines = events.vocabulary_for_prompt().splitlines()
assert len(lines) == len(events.SPECS)
for name, line in zip(events.SPECS, lines):
shape = line.strip().split(" — ", 1)[0]
obj = json.loads(shape)
assert obj["type"] == name
assert set(obj) - {"type"} == set(events.SPECS[name]["required"])
for optional in events.SPECS[name]["optional"]:
assert optional in line
def test_the_length_hint_carries_the_phrases_the_extractor_recognises():
"""One source for the words, so the builder and the extractor cannot drift."""
for narration_length in ("", "brief", "medium", "long"):
hint = builder.length_hint(500, narration_length)
assert hint.startswith(extract.LENGTH_HINT_OPENING)
assert extract.LENGTH_HINT_TAIL in hint
# ---------------------------------------------------------- observed shapes
STORY = (
"Alice looks at John, the tension in the room palpable.\n\n"
"John nods. \"I'm ready to contribute.\""
)
@pytest.mark.parametrize("leak", [
# Depth 20: the call, the fantasy slug, and a gloss on the same line.
'> set_possession(silver-key, "alice") Adds the silver key to Alice\'s possession.',
# Depth 10: cut off by the output limit mid-call.
'> Create_entity(new_person, "john", "character", "A determined team member", ["john',
# Unquoted, and a vocabulary name in any case.
'SET_CURRENT_LOCATION(bill, office)',
'add_fact(predicate="knows the plan", subject="alice")',
])
def test_an_event_call_line_at_the_end_leaves_the_story(leak):
prose, parsed, _raw = extract.split(f"{STORY}\n\n{leak}")
assert prose == STORY
assert parsed is None
def test_an_event_call_line_in_the_middle_leaves_and_the_story_after_it_stays():
"""Depth 12: the call, then more narration."""
reply = (
f"{STORY}\n\n"
'> Create_entity(new_person, "mike", "character", "A new team member.", ["mike"])\n\n'
"Mike takes the empty chair by the window."
)
prose, _parsed, _raw = extract.split(reply)
assert prose == f"{STORY}\n\nMike takes the empty chair by the window."
def test_the_depth_fourteen_tail_leaves_entirely():
"""A call, a rendered scene line, and a reworded length hint, in that order."""
reply = (
f"{STORY}\n\n"
'> Create_entity(mike, "character", "A new team member.", ["mike"])\n\n'
"Scene: Bill, Alice, Roger, John, and Mike at the table. (at The meeting room)\n\n"
"[Hard limit: your next turn must not exceed 180 words, and it should not stop "
"short of about 70. Prefer the lower end of that range unless the scene genuinely "
"needs more. Finish the narration and append the state block well inside the limit.]"
)
prose, parsed, _raw = extract.split(reply)
assert prose == STORY
assert parsed is None
@pytest.mark.parametrize("hint", [
builder.length_hint(500),
builder.length_hint(500, "brief"),
# Cut off by the output limit before the tail.
"[Hard limit: this turn must not exceed 180 words, and it should not stop short",
# Reworded at the front, as the 3B narrator did.
"[Hard limit: your next turn must not exceed 506 words. Write only as much as the "
"moment needs — a typical turn is much shorter. Finish the narration and append "
"the state block well inside the limit.]",
])
def test_a_parroted_length_hint_at_the_end_leaves_the_story(hint):
prose, _parsed, _raw = extract.split(f"{STORY}\n\n{hint}")
assert prose == STORY
def test_a_rendered_scene_line_at_the_end_leaves_the_story():
prose, _parsed, _raw = extract.split(
f"{STORY}\n\nScene: A tense budget meeting. (at The meeting room)")
assert prose == STORY
def test_a_fenced_block_with_a_call_line_above_it_still_parses_and_applies():
"""A2-6. The proposal is still read when protocol litter surrounds it."""
reply = (
f"{STORY}\n\n"
'> set_current_location(john, office)\n\n'
'```state\n{"events": [{"type": "set_current_location", '
'"entity": "john", "location": "office"}]}\n```'
)
prose, parsed, raw = extract.split(reply)
assert prose == STORY
assert parsed["events"][0]["entity"] == "john"
assert raw.startswith("{")
# ------------------------------------------------ adversarial story that stays
@pytest.mark.parametrize("reply", [
# The owner's cases.
'The engineer writes "set_power(core, 80)" on the whiteboard.',
'She says, "Create_entity is a terrible name for a company."',
'The old manual contains a heading labeled "Scene:"',
'He reads aloud: "[Hard limit: 500 words]" and laughs.',
# A vocabulary name, written into a story, not at the start of a line.
'Nadia squints at the log: the last command was set_possession(badge, guard).',
# Call-shaped, at the start of a line, but not an event this protocol has.
"The terminal scrolls.\n\n> open_door(north)\n\nNothing happens.",
# A vocabulary call inside the story's own code block is the story's code.
"She types:\n\n```python\ncreate_entity(ship)\nset_possession(key, captain)\n```\n\n"
"The console beeps twice.",
# A bracket at the very end, in-world, that is not the application's hint.
"The warning light blinks.\n\n[Hard limit of the reactor: three hours]",
"The contract ends with a clause.\n\n[Hard limit: forty days, no extensions]",
# A scene heading in a screenplay the characters are writing, mid-story.
"Scene: a kitchen, late.\n\nShe crosses it out and starts again.",
# A last line that starts like the renderer's but is not its shape.
"The director calls it.\n\nScene: take two, and nobody moves.",
# A fact restated inside a sentence.
"Alice knew the badge opened the server room, and said nothing.",
"Memory: she remembered the bells.",
])
def test_story_that_resembles_the_new_rules_is_kept(reply):
"""A2-5."""
prose, parsed, _raw = extract.split(reply)
assert prose == reply
assert parsed is None
# ------------------------------------------------------ the replay attribution
@pytest.mark.parametrize("line, rule", [
('> set_possession(silver-key, "alice") Adds the key.', extract.RULE_EVENT_CALL),
("Create_entity(new_person", extract.RULE_EVENT_CALL),
("[Hard limit: this turn must not exceed 90 words. Finish the narration and append "
"the state block well inside the limit.]", extract.RULE_LENGTH_HINT),
("Scene: A meeting. (at The meeting room)", extract.RULE_SCENE_LINE),
("John nods.", None),
('He reads aloud: "[Hard limit: 500 words]" and laughs.', None),
])
def test_a_removed_line_is_attributed_to_the_rule_that_removes_it(line, rule):
assert extract.explain_removed_line(line) == rule
# ------------------------------------- corrective: the depth-16 instruction tail
#: Cut down from the v1.1 identity diagnostic's depth-16 turn, whose stored text
#: was exactly the extractor's output. The two story paragraphs are shortened;
#: the four trailing lines are verbatim.
DEPTH_16_STORY = (
"John's initial ideas are thoughtful and insightful, and the room fills with a "
"sense of optimism.\n\n"
"John's enthusiasm is contagious, and the meeting room is electric with the "
"excitement of a fruitful collaboration ahead."
)
DEPTH_16_TAIL = (
"Scene: Bill, Alice and Roger at the table; John not yet arrived.\n\n"
"[Hard limit: this is now 180 words.]\n\n"
"[Reminder: end your reply with a `state` block listing the events your narration "
"made true, with absolute values.]\n\n"
"[You don't need to continue; your turn must now be about John entering the room. "
"Continue the story here, directly. Output only story text.]"
)
def test_the_depth_sixteen_instruction_tail_leaves_entirely():
"""The corrective's positive regression. v1.1's first A2 left all four lines:
the last bracket was a reworded continue hint nothing recognised, so nothing
above it was ever at the end."""
prose, parsed, _raw = extract.split(f"{DEPTH_16_STORY}\n\n{DEPTH_16_TAIL}")
assert prose == DEPTH_16_STORY
assert parsed is None
def test_the_continue_hint_phrase_is_the_providers_own_sentence():
from app.providers.openai_compatible import CHAT_CONTINUE_HINT
assert extract.CONTINUE_HINT_PHRASE in CHAT_CONTINUE_HINT
def test_an_echoed_continue_hint_alone_at_the_end_leaves():
prose, _p, _r = extract.split(
f"{STORY}\n\n[Keep going. Continue the story here, directly. Output only story text.]")
assert prose == STORY
@pytest.mark.parametrize("reply", [
# A hint-opened bracket with no echoed instruction below it is in-world.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# Nor does a state block below it make it an instruction.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# The phrase in the middle of a story is prose, not a trailing echo.
'She wrote "output only story text" on the card, then crossed it out.\n\nThe rain went on.',
# A trailing in-world bracket that only resembles a continuation.
f"{STORY}\n\n[To be continued]",
])
def test_story_brackets_near_the_corrective_rule_are_kept(reply):
prose, _p, _r = extract.split(reply)
assert prose == reply
def test_a_hint_opened_bracket_above_a_state_block_is_kept():
reply = (f"{STORY}\n\n[Hard limit: forty days, no extensions]\n\n"
'```state\n{"events": []}\n```')
prose, parsed, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[Hard limit: forty days, no extensions]"
assert parsed == {"events": []}
def test_an_in_world_bracket_above_an_echoed_hint_is_kept():
"""Only a bracket opening the way the application's hint opens is taken with
the echo. Any other bracket above it is the story's."""
reply = (f"{STORY}\n\n[The sign on the door reads: Closed]\n\n"
"[Continue the story here, directly. Output only story text.]")
prose, _p, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[The sign on the door reads: Closed]"
@pytest.mark.parametrize("line, rule", [
("[Hard limit: this is now 180 words.]", extract.RULE_INSTRUCTION_TAIL),
("[Reminder: end your reply with a `state` block listing the events.]",
extract.RULE_INSTRUCTION_TAIL),
("[You don't need to continue. Output only story text.]", extract.RULE_INSTRUCTION_TAIL),
])
def test_the_corrective_rule_is_attributed(line, rule):
assert extract.explain_removed_line(line) == rule
def test_the_new_rules_do_not_disturb_the_section_headings_they_share_a_module_with():
"""The renderer's headings are the M11 rules' anchor. A2 adds none."""
assert render.HEADING_SCENE == "Scene:"
assert "Scene:" not in render.SECTION_HEADINGS