Files
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00

188 lines
8.7 KiB
Python

"""v1.1: does a real v1.0.0 database open unchanged?
# the same database, opened by each tree, snapshotted read-only
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
--tree <v1.0.0 worktree>/backend --label v100 --out "$HOME/v11-evidence/compat"
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
--label v11 --exercise --out "$HOME/v11-evidence/compat"
Run from `backend/`. The source database is never opened. It is copied into
`--out` first, and the copy is what the application opens.
A read-only snapshot is taken through the API, the same way a reader sees the
campaign:
- the export bundle, which carries the whole tree, the head, the Save Points,
state, events, summaries, memories and knowledge, and has no timestamp of its
own;
- the narrative state and its events;
- the Save Points, the imported knowledge, the memories, the derived status and
the settings;
- the database schema and `PRAGMA user_version`, before and after the
application opened it.
Two snapshots of the same database from two trees are then compared. Identical
means v1.1 read it exactly as v1.0.0 did, and a matching schema and version mean
nothing migrated.
`--exercise` then uses the v1.1 copy: undo, redo, a Save Point restore, a
context dry run (knowledge retrieval), an export, and an import of that export.
It first points the copy's endpoint at a loopback port that refuses, so nothing
here reaches an inference server.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sqlite3
import sys
from pathlib import Path
def _schema(path: Path) -> dict:
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
version = connection.execute("PRAGMA user_version").fetchone()[0]
rows = connection.execute(
"SELECT type, name, sql FROM sqlite_master WHERE name NOT LIKE 'sqlite_%' "
"ORDER BY type, name").fetchall()
finally:
connection.close()
return {"user_version": version, "objects": [list(r) for r in rows]}
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--db", required=True)
parser.add_argument("--tree", default="", help="a backend/ directory to import the app from")
parser.add_argument("--label", required=True)
parser.add_argument("--exercise", action="store_true")
parser.add_argument("--out", required=True)
args = parser.parse_args()
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
copy = out / f"{args.label}.db"
if copy.exists():
print(f"{copy} exists; choose a new --label or --out")
return 2
shutil.copy2(args.db, copy)
schema_before = _schema(copy)
if args.tree:
sys.path.insert(0, str(Path(args.tree).resolve()))
os.environ["AIDND_DB_PATH"] = str(copy)
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, models
from app.database import SessionLocal, get_db
from app.main import app
print(f"app imported from {Path(sys.modules['app'].__file__).parent}")
limits.check_row_cap = lambda *a, **k: None
with SessionLocal() as db:
owner = db.query(models.Adventure.user_id).order_by(models.Adventure.id).first()
user_id = owner[0] if owner else db.query(models.User.id).first()[0]
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
def call(client, method, url, body=None, expect=200):
response = client.request(method, f"/api{url}", json=body)
if response.status_code != expect:
raise SystemExit(f"{method} {url}: HTTP {response.status_code} {response.text[:300]}")
return response.json() if response.content else None
report: dict = {"label": args.label, "schema_before": schema_before}
with TestClient(app) as client:
with SessionLocal() as db:
adventure_ids = [a for (a,) in db.query(models.Adventure.id)
.filter(models.Adventure.user_id == user_id)
.order_by(models.Adventure.id)]
snapshot = {"settings": call(client, "GET", "/settings"), "adventures": {}}
for adv in adventure_ids:
snapshot["adventures"][str(adv)] = {
"export": call(client, "GET", f"/adventures/{adv}/export"),
"state": call(client, "GET", f"/adventures/{adv}/state"),
"state_events": call(client, "GET", f"/adventures/{adv}/state/events"),
"checkpoints": call(client, "GET", f"/adventures/{adv}/checkpoints"),
"knowledge": call(client, "GET", f"/adventures/{adv}/knowledge"),
"memories": call(client, "GET", f"/adventures/{adv}/memories"),
"derived": call(client, "GET", f"/adventures/{adv}/derived"),
"newest_actions": call(client, "GET", f"/adventures/{adv}/actions?limit=5"),
}
report["snapshot"] = snapshot
if args.exercise and adventure_ids:
adv = adventure_ids[0]
ex: dict = {}
call(client, "PUT", "/settings", {"endpoint_url": "http://127.0.0.1:9/v1",
"embedding_model": ""})
before = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
ex["before"] = {k: before[k] for k in ("total", "can_undo", "can_redo")}
undone = call(client, "POST", f"/adventures/{adv}/undo")
ex["after_undo"] = {k: undone[k] for k in ("total", "can_undo", "can_redo")}
redone = call(client, "POST", f"/adventures/{adv}/redo")
ex["after_redo"] = {k: redone[k] for k in ("total", "can_undo", "can_redo")}
points = call(client, "GET", f"/adventures/{adv}/checkpoints")
if points:
point = points[0]
restored = call(client, "POST",
f"/adventures/{adv}/checkpoints/{point['id']}/restore")
page = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
ex["restore"] = {"save_point": point["name"], "total": page["total"],
"can_redo": page["can_redo"],
"response_keys": sorted(restored or {})}
context = client.get(f"/api/adventures/{adv}/context")
body = context.json()
ex["context"] = {
"status": context.status_code,
"knowledge_used": len(((body.get("knowledge") or {}).get("used")) or []),
"canon_section": any(s["label"] == "campaign_canon"
for s in body.get("sections") or []),
"state_section": any(s["label"] == "narrative_state"
for s in body.get("sections") or []),
"summary": body.get("summary"),
"tokens": body.get("tokens"),
"window": body.get("window"),
}
bundle = call(client, "GET", f"/adventures/{adv}/export")
imported = call(client, "POST", "/adventures/import", bundle, expect=201)
new_id = imported["id"]
reimport = call(client, "GET", f"/adventures/{new_id}/export")
ex["import"] = {
"new_id": new_id,
"actions_in_bundle": len(bundle.get("actions") or []),
"actions_after_import": len(reimport.get("actions") or []),
"head_same": (bundle.get("headBranch") is not None
and bundle.get("headDepth") == reimport.get("headDepth")),
"checkpoints": [len(bundle.get("checkpoints") or []),
len(reimport.get("checkpoints") or [])],
"memories": [len(bundle.get("memories") or []),
len(reimport.get("memories") or [])],
"narrative_state_same": bundle.get("narrativeState") == reimport.get("narrativeState"),
}
report["exercise"] = ex
app.dependency_overrides.clear()
report["schema_after"] = _schema(copy)
(out / f"{args.label}.json").write_text(json.dumps(report, indent=2, sort_keys=True, default=str))
same_schema = report["schema_before"] == report["schema_after"]
print(f"schema unchanged by opening: {same_schema} "
f"(user_version {report['schema_before']['user_version']} -> "
f"{report['schema_after']['user_version']})")
if "exercise" in report:
print(json.dumps(report["exercise"], indent=2, default=str)[:3000])
return 0
if __name__ == "__main__":
sys.exit(main())