v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
ac465ed867
commit
d63804f22e
@@ -559,6 +559,17 @@ class Run:
|
||||
self.accepted += 1
|
||||
seconds = time.monotonic() - started
|
||||
sample = self.measure()
|
||||
# v1.1 WP-A1: what the server said it read for the turn just played, from
|
||||
# the `done` event. `.get` because a build before v1.1 sends none.
|
||||
done = next((e for e in events if e.get("type") == "done"), {})
|
||||
accounting = done.get("accounting") or {}
|
||||
sample.update({
|
||||
"accounting_status": accounting.get("status"),
|
||||
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
|
||||
"app_prompt_estimate": accounting.get("estimate"),
|
||||
"observed_margin": accounting.get("observed_margin"),
|
||||
"safety_reserve": accounting.get("safety_reserve"),
|
||||
})
|
||||
self.note("turn", text=text, seconds=round(seconds, 1), **sample)
|
||||
return {"accepted": True, "seconds": seconds, **sample}
|
||||
|
||||
@@ -1225,10 +1236,26 @@ PROTOCOL_LEAK_HEADING_RE = re.compile(
|
||||
re.MULTILINE,
|
||||
)
|
||||
PROTOCOL_LEAK_EVENTS = '"events"'
|
||||
#: v1.1 WP-A2: the two shapes the M11 closeout's identity run stored that the
|
||||
#: two signs above cannot see — a line opening with a call to an event, and the
|
||||
#: length hint echoed with the application's own wording. Copied, as above.
|
||||
PROTOCOL_LEAK_CALL_RE = re.compile(
|
||||
r"^[ \t]*(?:>[ \t]*)?(?:create_entity|set_entity_status|set_entity_attribute"
|
||||
r"|set_entity_conditions|set_current_location|set_possession|clear_possession"
|
||||
r"|add_fact|invalidate_fact|add_relationship|end_relationship"
|
||||
r"|open_story_thread|resolve_story_thread|set_scene)[ \t]*\(",
|
||||
re.IGNORECASE | re.MULTILINE,
|
||||
)
|
||||
PROTOCOL_LEAK_HINT_RE = re.compile(
|
||||
r"\[Hard limit:[^\]]*(?:append the state block|turn must not exceed \d+ words)",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
|
||||
|
||||
def _leaks_protocol(text: str) -> bool:
|
||||
return bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
|
||||
return (bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
|
||||
or bool(PROTOCOL_LEAK_CALL_RE.search(text))
|
||||
or bool(PROTOCOL_LEAK_HINT_RE.search(text)))
|
||||
|
||||
|
||||
def _protocol_leaks(bundle: dict) -> dict:
|
||||
|
||||
@@ -0,0 +1,187 @@
|
||||
"""v1.1: does a real v1.0.0 database open unchanged?
|
||||
|
||||
# the same database, opened by each tree, snapshotted read-only
|
||||
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
|
||||
--tree <v1.0.0 worktree>/backend --label v100 --out "$HOME/v11-evidence/compat"
|
||||
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
|
||||
--label v11 --exercise --out "$HOME/v11-evidence/compat"
|
||||
|
||||
Run from `backend/`. The source database is never opened. It is copied into
|
||||
`--out` first, and the copy is what the application opens.
|
||||
|
||||
A read-only snapshot is taken through the API, the same way a reader sees the
|
||||
campaign:
|
||||
|
||||
- the export bundle, which carries the whole tree, the head, the Save Points,
|
||||
state, events, summaries, memories and knowledge, and has no timestamp of its
|
||||
own;
|
||||
- the narrative state and its events;
|
||||
- the Save Points, the imported knowledge, the memories, the derived status and
|
||||
the settings;
|
||||
- the database schema and `PRAGMA user_version`, before and after the
|
||||
application opened it.
|
||||
|
||||
Two snapshots of the same database from two trees are then compared. Identical
|
||||
means v1.1 read it exactly as v1.0.0 did, and a matching schema and version mean
|
||||
nothing migrated.
|
||||
|
||||
`--exercise` then uses the v1.1 copy: undo, redo, a Save Point restore, a
|
||||
context dry run (knowledge retrieval), an export, and an import of that export.
|
||||
It first points the copy's endpoint at a loopback port that refuses, so nothing
|
||||
here reaches an inference server.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def _schema(path: Path) -> dict:
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
version = connection.execute("PRAGMA user_version").fetchone()[0]
|
||||
rows = connection.execute(
|
||||
"SELECT type, name, sql FROM sqlite_master WHERE name NOT LIKE 'sqlite_%' "
|
||||
"ORDER BY type, name").fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
return {"user_version": version, "objects": [list(r) for r in rows]}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--db", required=True)
|
||||
parser.add_argument("--tree", default="", help="a backend/ directory to import the app from")
|
||||
parser.add_argument("--label", required=True)
|
||||
parser.add_argument("--exercise", action="store_true")
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
copy = out / f"{args.label}.db"
|
||||
if copy.exists():
|
||||
print(f"{copy} exists; choose a new --label or --out")
|
||||
return 2
|
||||
shutil.copy2(args.db, copy)
|
||||
schema_before = _schema(copy)
|
||||
|
||||
if args.tree:
|
||||
sys.path.insert(0, str(Path(args.tree).resolve()))
|
||||
os.environ["AIDND_DB_PATH"] = str(copy)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, models
|
||||
from app.database import SessionLocal, get_db
|
||||
from app.main import app
|
||||
|
||||
print(f"app imported from {Path(sys.modules['app'].__file__).parent}")
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
with SessionLocal() as db:
|
||||
owner = db.query(models.Adventure.user_id).order_by(models.Adventure.id).first()
|
||||
user_id = owner[0] if owner else db.query(models.User.id).first()[0]
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
|
||||
def call(client, method, url, body=None, expect=200):
|
||||
response = client.request(method, f"/api{url}", json=body)
|
||||
if response.status_code != expect:
|
||||
raise SystemExit(f"{method} {url}: HTTP {response.status_code} {response.text[:300]}")
|
||||
return response.json() if response.content else None
|
||||
|
||||
report: dict = {"label": args.label, "schema_before": schema_before}
|
||||
with TestClient(app) as client:
|
||||
with SessionLocal() as db:
|
||||
adventure_ids = [a for (a,) in db.query(models.Adventure.id)
|
||||
.filter(models.Adventure.user_id == user_id)
|
||||
.order_by(models.Adventure.id)]
|
||||
snapshot = {"settings": call(client, "GET", "/settings"), "adventures": {}}
|
||||
for adv in adventure_ids:
|
||||
snapshot["adventures"][str(adv)] = {
|
||||
"export": call(client, "GET", f"/adventures/{adv}/export"),
|
||||
"state": call(client, "GET", f"/adventures/{adv}/state"),
|
||||
"state_events": call(client, "GET", f"/adventures/{adv}/state/events"),
|
||||
"checkpoints": call(client, "GET", f"/adventures/{adv}/checkpoints"),
|
||||
"knowledge": call(client, "GET", f"/adventures/{adv}/knowledge"),
|
||||
"memories": call(client, "GET", f"/adventures/{adv}/memories"),
|
||||
"derived": call(client, "GET", f"/adventures/{adv}/derived"),
|
||||
"newest_actions": call(client, "GET", f"/adventures/{adv}/actions?limit=5"),
|
||||
}
|
||||
report["snapshot"] = snapshot
|
||||
|
||||
if args.exercise and adventure_ids:
|
||||
adv = adventure_ids[0]
|
||||
ex: dict = {}
|
||||
call(client, "PUT", "/settings", {"endpoint_url": "http://127.0.0.1:9/v1",
|
||||
"embedding_model": ""})
|
||||
before = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
|
||||
ex["before"] = {k: before[k] for k in ("total", "can_undo", "can_redo")}
|
||||
undone = call(client, "POST", f"/adventures/{adv}/undo")
|
||||
ex["after_undo"] = {k: undone[k] for k in ("total", "can_undo", "can_redo")}
|
||||
redone = call(client, "POST", f"/adventures/{adv}/redo")
|
||||
ex["after_redo"] = {k: redone[k] for k in ("total", "can_undo", "can_redo")}
|
||||
points = call(client, "GET", f"/adventures/{adv}/checkpoints")
|
||||
if points:
|
||||
point = points[0]
|
||||
restored = call(client, "POST",
|
||||
f"/adventures/{adv}/checkpoints/{point['id']}/restore")
|
||||
page = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
|
||||
ex["restore"] = {"save_point": point["name"], "total": page["total"],
|
||||
"can_redo": page["can_redo"],
|
||||
"response_keys": sorted(restored or {})}
|
||||
context = client.get(f"/api/adventures/{adv}/context")
|
||||
body = context.json()
|
||||
ex["context"] = {
|
||||
"status": context.status_code,
|
||||
"knowledge_used": len(((body.get("knowledge") or {}).get("used")) or []),
|
||||
"canon_section": any(s["label"] == "campaign_canon"
|
||||
for s in body.get("sections") or []),
|
||||
"state_section": any(s["label"] == "narrative_state"
|
||||
for s in body.get("sections") or []),
|
||||
"summary": body.get("summary"),
|
||||
"tokens": body.get("tokens"),
|
||||
"window": body.get("window"),
|
||||
}
|
||||
bundle = call(client, "GET", f"/adventures/{adv}/export")
|
||||
imported = call(client, "POST", "/adventures/import", bundle, expect=201)
|
||||
new_id = imported["id"]
|
||||
reimport = call(client, "GET", f"/adventures/{new_id}/export")
|
||||
ex["import"] = {
|
||||
"new_id": new_id,
|
||||
"actions_in_bundle": len(bundle.get("actions") or []),
|
||||
"actions_after_import": len(reimport.get("actions") or []),
|
||||
"head_same": (bundle.get("headBranch") is not None
|
||||
and bundle.get("headDepth") == reimport.get("headDepth")),
|
||||
"checkpoints": [len(bundle.get("checkpoints") or []),
|
||||
len(reimport.get("checkpoints") or [])],
|
||||
"memories": [len(bundle.get("memories") or []),
|
||||
len(reimport.get("memories") or [])],
|
||||
"narrative_state_same": bundle.get("narrativeState") == reimport.get("narrativeState"),
|
||||
}
|
||||
report["exercise"] = ex
|
||||
|
||||
app.dependency_overrides.clear()
|
||||
report["schema_after"] = _schema(copy)
|
||||
(out / f"{args.label}.json").write_text(json.dumps(report, indent=2, sort_keys=True, default=str))
|
||||
same_schema = report["schema_before"] == report["schema_after"]
|
||||
print(f"schema unchanged by opening: {same_schema} "
|
||||
f"(user_version {report['schema_before']['user_version']} -> "
|
||||
f"{report['schema_after']['user_version']})")
|
||||
if "exercise" in report:
|
||||
print(json.dumps(report["exercise"], indent=2, default=str)[:3000])
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,245 @@
|
||||
"""v1.1 WP-A2: replay real stored narration through the v1.0.0 and current extractors.
|
||||
|
||||
The A2 extractor changes remove more text from a narrator's reply than v1.0.0
|
||||
did. Removing story is worse than leaving protocol (`TECHNICAL-DESIGN.md`
|
||||
§15.4), so every change is shown to a person rather than summarised. This tool
|
||||
takes every real reply the evidence kept, runs it through both extractors, and
|
||||
writes each turn whose prose differs with:
|
||||
|
||||
- the v1.0.0 prose and the current prose;
|
||||
- every line removed, and the rule that explains it;
|
||||
- any removal no rule explains, which fails the replay.
|
||||
|
||||
**Input.** A reply is read from the turn's stored `raw_output` where the
|
||||
evidence database kept one: that is exactly what the narrator sent. A bundle
|
||||
carries no raw reply, so a bundle's turns are replayed from their stored text,
|
||||
which is v1.0.0's output already. For those the old prose is the input itself,
|
||||
and the comparison is still exact.
|
||||
|
||||
**The v1.0.0 extractor** is read from the release tag with `git show`, not
|
||||
copied, so this tool compares against what shipped. It shares `events` and
|
||||
`render` with the current tree. A2 does not change `events.SPECS` or
|
||||
`render.SECTION_HEADINGS`, and the report verifies that with `git diff`.
|
||||
|
||||
Evidence stays outside the repository. Replayed text is fiction from the
|
||||
acceptance fixtures, but it is still somebody's run.
|
||||
|
||||
.venv/bin/python -m tools.v11_replay_extractor \\
|
||||
--db "$HOME/m11-evidence/**/*.db" \\
|
||||
--bundle "$HOME/m11-evidence/closeout-3652dc6/identity/turn-99/bundle.json" \\
|
||||
--out "$HOME/v11-evidence/a2-replay"
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import difflib
|
||||
import glob
|
||||
import hashlib
|
||||
import importlib.util
|
||||
import json
|
||||
import sqlite3
|
||||
import subprocess
|
||||
import sys
|
||||
import types
|
||||
import zlib
|
||||
from pathlib import Path
|
||||
|
||||
RELEASE = "v1.0.0"
|
||||
EXTRACT_PATH = "backend/app/narrative/extract.py"
|
||||
#: A removal larger than this share of the v1.0.0 prose is flagged for review
|
||||
#: even when every line is explained, because a rule that eats most of a reply
|
||||
#: is the shape a false positive takes.
|
||||
LARGE_REMOVAL_SHARE = 0.25
|
||||
|
||||
|
||||
def load_release_extractor(repo: Path) -> types.ModuleType:
|
||||
"""`app.narrative.extract` as it was at the release tag."""
|
||||
source = subprocess.run(
|
||||
["git", "-C", str(repo), "show", f"{RELEASE}:{EXTRACT_PATH}"],
|
||||
check=True, capture_output=True, text=True,
|
||||
).stdout
|
||||
import app.narrative # noqa: F401 the package the relative import needs
|
||||
|
||||
spec = importlib.util.spec_from_loader("app.narrative._extract_release", loader=None)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
module.__package__ = "app.narrative"
|
||||
exec(compile(source, f"{RELEASE}:{EXTRACT_PATH}", "exec"), module.__dict__)
|
||||
return module
|
||||
|
||||
|
||||
def _unpack(blob):
|
||||
if blob is None:
|
||||
return None
|
||||
try:
|
||||
return json.loads(zlib.decompress(bytes(blob)).decode("utf-8"))
|
||||
except (zlib.error, ValueError, UnicodeDecodeError):
|
||||
return None
|
||||
|
||||
|
||||
def turns_from_db(path: str):
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
rows = connection.execute(
|
||||
"SELECT id, depth, text, context_snapshot FROM actions WHERE type = 'ai'"
|
||||
).fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
for action_id, depth, text, blob in rows:
|
||||
snapshot = _unpack(blob) or {}
|
||||
raw = snapshot.get("raw_output")
|
||||
if isinstance(raw, str) and raw.strip():
|
||||
yield {"source": path, "id": action_id, "depth": depth,
|
||||
"input": raw, "input_kind": "raw_output"}
|
||||
elif text:
|
||||
yield {"source": path, "id": action_id, "depth": depth,
|
||||
"input": text, "input_kind": "stored_text"}
|
||||
|
||||
|
||||
def turns_from_bundle(path: str):
|
||||
bundle = json.loads(Path(path).read_text())
|
||||
for action in bundle.get("actions") or []:
|
||||
if action.get("type") == "ai" and action.get("text"):
|
||||
yield {"source": path, "id": action.get("id"), "depth": action.get("depth"),
|
||||
"input": action["text"], "input_kind": "stored_text"}
|
||||
|
||||
|
||||
def removed_lines(before: str, after: str) -> list[str]:
|
||||
"""Lines present in `before` and gone from `after`, in order.
|
||||
|
||||
Compared with trailing whitespace ignored. The extractor strips the end of
|
||||
every reply it cuts, so a story line that becomes the last line loses a
|
||||
trailing space. A first version of this tool reported that space as a
|
||||
rewritten line of story, which it is not.
|
||||
"""
|
||||
old = [line.rstrip() for line in before.split("\n")]
|
||||
new = [line.rstrip() for line in after.split("\n")]
|
||||
matcher = difflib.SequenceMatcher(a=old, b=new, autojunk=False)
|
||||
gone: list[str] = []
|
||||
for tag, a0, a1, _b0, _b1 in matcher.get_opcodes():
|
||||
if tag in ("delete", "replace"):
|
||||
gone.extend(old[a0:a1])
|
||||
return gone
|
||||
|
||||
|
||||
def replay(inputs, old, new) -> dict:
|
||||
seen: set[str] = set()
|
||||
unchanged = 0
|
||||
changed: list[dict] = []
|
||||
duplicates = 0
|
||||
for turn in inputs:
|
||||
digest = hashlib.sha256(turn["input"].encode()).hexdigest()
|
||||
if digest in seen:
|
||||
duplicates += 1
|
||||
continue
|
||||
seen.add(digest)
|
||||
old_prose, _old_parsed, _old_raw = old.split(turn["input"])
|
||||
new_prose, _new_parsed, _new_raw = new.split(turn["input"])
|
||||
if old_prose == new_prose:
|
||||
unchanged += 1
|
||||
continue
|
||||
lines = []
|
||||
unexplained = 0
|
||||
for line in removed_lines(old_prose, new_prose):
|
||||
if not line.strip():
|
||||
continue
|
||||
rule = new.explain_removed_line(line)
|
||||
if rule is None:
|
||||
unexplained += 1
|
||||
lines.append({"line": line, "rule": rule})
|
||||
added = [line for line in removed_lines(new_prose, old_prose) if line.strip()]
|
||||
share = 1 - len(new_prose) / max(1, len(old_prose))
|
||||
flags = []
|
||||
if unexplained:
|
||||
flags.append("unexplained_removal")
|
||||
if added:
|
||||
flags.append("text_added_or_rewritten")
|
||||
if share > LARGE_REMOVAL_SHARE:
|
||||
flags.append("large_removal")
|
||||
changed.append({
|
||||
**{k: turn[k] for k in ("source", "id", "depth", "input_kind")},
|
||||
"sha256": digest,
|
||||
"old_prose": old_prose,
|
||||
"new_prose": new_prose,
|
||||
"removed": lines,
|
||||
"added_or_rewritten": added,
|
||||
"removed_chars": len(old_prose) - len(new_prose),
|
||||
"removed_share": round(share, 4),
|
||||
"flags": flags,
|
||||
})
|
||||
return {
|
||||
"replayed": unchanged + len(changed),
|
||||
"duplicates_skipped": duplicates,
|
||||
"unchanged": unchanged,
|
||||
"changed": len(changed),
|
||||
"flagged": sum(1 for c in changed if c["flags"]),
|
||||
"turns": changed,
|
||||
}
|
||||
|
||||
|
||||
def write_markdown(result: dict, path: Path) -> None:
|
||||
out = [
|
||||
"# A2 extractor replay",
|
||||
"",
|
||||
f"- replayed (unique replies): **{result['replayed']}**",
|
||||
f"- duplicates skipped: {result['duplicates_skipped']}",
|
||||
f"- unchanged: {result['unchanged']}",
|
||||
f"- changed: **{result['changed']}**",
|
||||
f"- flagged: **{result['flagged']}**",
|
||||
"",
|
||||
]
|
||||
for index, turn in enumerate(result["turns"], 1):
|
||||
out += [
|
||||
f"## {index}. {Path(turn['source']).parent.name}/{Path(turn['source']).name}"
|
||||
f" action {turn['id']} depth {turn['depth']} ({turn['input_kind']})",
|
||||
"",
|
||||
f"- removed chars: {turn['removed_chars']} ({turn['removed_share']:.1%})",
|
||||
f"- flags: {', '.join(turn['flags']) or 'none'}",
|
||||
"",
|
||||
"Removed lines:",
|
||||
"",
|
||||
]
|
||||
for item in turn["removed"]:
|
||||
out.append(f"- `{item['rule'] or 'UNEXPLAINED'}` — {item['line']!r}")
|
||||
out += ["", "<details><summary>v1.0.0 prose</summary>", "", "```text",
|
||||
turn["old_prose"], "```", "</details>", "",
|
||||
"<details><summary>current prose</summary>", "", "```text",
|
||||
turn["new_prose"], "```", "</details>", ""]
|
||||
path.write_text("\n".join(out))
|
||||
|
||||
|
||||
def main(argv=None) -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--db", action="append", default=[],
|
||||
help="an evidence database, or a glob of them")
|
||||
parser.add_argument("--bundle", action="append", default=[],
|
||||
help="an exported bundle whose campaign has no database here")
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args(argv)
|
||||
|
||||
repo = Path(__file__).resolve().parents[2]
|
||||
from app.narrative import extract as current
|
||||
|
||||
old = load_release_extractor(repo)
|
||||
dbs = sorted({p for pattern in args.db for p in glob.glob(pattern, recursive=True)})
|
||||
|
||||
def inputs():
|
||||
for path in dbs:
|
||||
yield from turns_from_db(path)
|
||||
for path in args.bundle:
|
||||
yield from turns_from_bundle(path)
|
||||
|
||||
result = replay(inputs(), old, current)
|
||||
result["databases"] = dbs
|
||||
result["bundles"] = args.bundle
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
(out / "replay.json").write_text(json.dumps(result, indent=2, ensure_ascii=False))
|
||||
write_markdown(result, out / "replay.md")
|
||||
print(json.dumps({k: result[k] for k in
|
||||
("replayed", "duplicates_skipped", "unchanged", "changed", "flagged")}))
|
||||
return 1 if result["flagged"] else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,211 @@
|
||||
"""v1.1 WP-A1: real turns, and what the server said it read.
|
||||
|
||||
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 AIDND_TEST_MODEL=<model> \\
|
||||
.venv/bin/python -m tools.v11_window_accounting \\
|
||||
--bundle "$HOME/m11-evidence/m04-final/bundle.json" --turns 4 \\
|
||||
--out "$HOME/v11-evidence/a1-accounting/<label>"
|
||||
|
||||
Run from `backend/`. The evidence that matters is at the edge of the window, and
|
||||
a new campaign takes dozens of turns to reach it. So this imports a long
|
||||
campaign, by default the v1 evidence run's 207-action bundle, and every turn is
|
||||
assembled against a full window from the first. A real narrator is then asked
|
||||
for `--turns` turns, and each one is written out with:
|
||||
|
||||
configured budget, verified window, and where the window came from
|
||||
the application's estimate of what it sent (its count, plus the text the
|
||||
provider adds)
|
||||
the server's own prompt-token count, from the usage it reported
|
||||
the reply allocation and the safety reserve
|
||||
the observed margin: window - reply allocation - the server's count
|
||||
the accounting status: fits, exceeded, truncation_suspected or unknown
|
||||
|
||||
The first turn on a cold model cannot verify the window: `/api/ps` knows nothing
|
||||
until the model is loaded. That is M11's behaviour, and the row says so rather
|
||||
than hiding the turn.
|
||||
|
||||
The database lives in `--out`, not in `/tmp`, so the stored snapshots behind
|
||||
every row can be read again. The endpoint is read from the environment and is
|
||||
never written into a committed file.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
|
||||
TURNS = [
|
||||
"I look around carefully and take stock of where I am.",
|
||||
"I ask the nearest person what has happened since I was last here.",
|
||||
"I check what I am carrying.",
|
||||
"I move on towards the place I meant to reach.",
|
||||
"I wait and listen.",
|
||||
"I say, \"Tell me the part you left out.\"",
|
||||
]
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--bundle", default="",
|
||||
help="a campaign bundle to import, so turns start at a full window")
|
||||
parser.add_argument("--turns", type=int, default=4)
|
||||
parser.add_argument("--budget", type=int, default=16384)
|
||||
parser.add_argument("--max-output", type=int, default=500)
|
||||
parser.add_argument("--timeout", type=int, default=900)
|
||||
parser.add_argument("--out", required=True)
|
||||
parser.add_argument("--unload-first", action="store_true",
|
||||
help="ask the configured server to unload the model before turn 1, "
|
||||
"so the first turn starts cold (the A1 corrective test)")
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL):
|
||||
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
|
||||
return 2
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
db_path = out / "accounting.db"
|
||||
if db_path.exists():
|
||||
print(f"{db_path} exists; choose a new --out")
|
||||
return 2
|
||||
os.environ["AIDND_DB_PATH"] = str(db_path)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
Base.metadata.create_all(bind=engine)
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email="accounting@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
|
||||
context_token_budget=args.budget, max_output_tokens=args.max_output,
|
||||
model_timeout_seconds=args.timeout,
|
||||
))
|
||||
db.commit()
|
||||
user_id = user.id
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
client = TestClient(app)
|
||||
|
||||
if args.bundle:
|
||||
bundle = json.loads(Path(args.bundle).read_text())
|
||||
# The evidence campaign had its memory bank and auto-summarise on. Here
|
||||
# they would only add post-turn model calls between the measured turns,
|
||||
# on the same host, and fail noisily as the in-process client closes.
|
||||
# This tool measures the turn's own prompt, which neither changes.
|
||||
bundle["memoryBankEnabled"] = False
|
||||
bundle["autoSummarize"] = False
|
||||
imported = client.post("/api/adventures/import", json=bundle)
|
||||
imported.raise_for_status()
|
||||
adv = imported.json()["id"]
|
||||
else:
|
||||
created = client.post("/api/adventures", json={
|
||||
"title": "Window accounting", "opening": "A quiet road at dusk."})
|
||||
created.raise_for_status()
|
||||
adv = created.json()["id"]
|
||||
|
||||
if args.unload_first:
|
||||
# The configured endpoint only, under the same policy and TLS trust as a turn.
|
||||
import asyncio
|
||||
import httpx
|
||||
from app import contextwindow, endpoints, tlstrust
|
||||
reason = endpoints.rejection_reason(ENDPOINT)
|
||||
if reason:
|
||||
print(f"endpoint refused: {reason}")
|
||||
return 2
|
||||
base = contextwindow.native_base(ENDPOINT)
|
||||
with httpx.Client(verify=tlstrust.ssl_context(), timeout=120) as http:
|
||||
unloaded = http.post(f"{base}/api/generate", json={"model": MODEL, "keep_alive": 0})
|
||||
resident = [m.get("name") for m in http.get(f"{base}/api/ps").json().get("models", [])]
|
||||
contextwindow.cache_clear()
|
||||
print(f"unload: HTTP {unloaded.status_code} {unloaded.text[:120]} | resident now: {resident}")
|
||||
|
||||
rows: list[dict] = []
|
||||
timeline = (out / "turns.jsonl").open("a")
|
||||
print(f"model {MODEL}, budget {args.budget}, reply {args.max_output}")
|
||||
print(f"{'#':>2} {'status':22} {'window':>13} {'estimate':>8} {'server':>7} "
|
||||
f"{'reserve':>7} {'margin':>7} {'sec':>5}")
|
||||
for index in range(args.turns):
|
||||
text = TURNS[index % len(TURNS)]
|
||||
started = time.monotonic()
|
||||
response = client.post(f"/api/adventures/{adv}/actions",
|
||||
json={"type": "do", "text": text})
|
||||
seconds = round(time.monotonic() - started, 1)
|
||||
error = None
|
||||
if response.status_code != 200 or '"type": "error"' in response.text:
|
||||
error = response.text[-400:]
|
||||
with SessionLocal() as db:
|
||||
action = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
snapshot = (action.context_snapshot or {}) if action else {}
|
||||
tokens = snapshot.get("tokens") or {}
|
||||
window = snapshot.get("window") or {}
|
||||
accounting = snapshot.get("accounting") or {}
|
||||
row = {
|
||||
"turn": index + 1,
|
||||
"action_id": action.id if action else None,
|
||||
"seconds": seconds,
|
||||
"error": error,
|
||||
"configured_budget": tokens.get("configured_budget"),
|
||||
"effective_budget": tokens.get("budget"),
|
||||
"window_verified": window.get("verified"),
|
||||
"window_tokens": window.get("tokens"),
|
||||
"window_source": window.get("source"),
|
||||
"preflight_attempted": (window.get("preflight") or {}).get("attempted"),
|
||||
"preflight_loaded": (window.get("preflight") or {}).get("loaded"),
|
||||
"preflight_verified_before": (window.get("preflight") or {}).get("verified_before"),
|
||||
"preflight_verified_after": (window.get("preflight") or {}).get("verified_after"),
|
||||
"preflight_detail": (window.get("preflight") or {}).get("detail"),
|
||||
"app_prompt_tokens": tokens.get("total"),
|
||||
"transport_tokens": tokens.get("transport"),
|
||||
"app_estimate": tokens.get("estimate"),
|
||||
"output_reserve": tokens.get("output_reserve"),
|
||||
"safety_reserve": tokens.get("safety_reserve"),
|
||||
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
|
||||
"difference": accounting.get("difference"),
|
||||
"observed_margin": accounting.get("observed_margin"),
|
||||
"status": accounting.get("status"),
|
||||
"history_included": (snapshot.get("history") or {}).get("included"),
|
||||
"history_total": (snapshot.get("history") or {}).get("total"),
|
||||
}
|
||||
rows.append(row)
|
||||
timeline.write(json.dumps(row) + "\n")
|
||||
timeline.flush()
|
||||
print(f"{row['turn']:>2} {str(row['status'] if not error else 'ERROR'):22} "
|
||||
f"{str(row['window_tokens'])+('v' if row['window_verified'] else '?'):>13} "
|
||||
f"{str(row['app_estimate']):>8} {str(row['server_prompt_tokens']):>7} "
|
||||
f"{str(row['safety_reserve']):>7} {str(row['observed_margin']):>7} {seconds:>5}")
|
||||
if error:
|
||||
print(f" error: {error[:200]}")
|
||||
timeline.close()
|
||||
(out / "summary.json").write_text(json.dumps({
|
||||
"model": MODEL, "budget": args.budget, "max_output_tokens": args.max_output,
|
||||
"bundle": args.bundle, "rows": rows,
|
||||
}, indent=2))
|
||||
app.dependency_overrides.clear()
|
||||
return 0 if all(r["error"] is None for r in rows) else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Reference in New Issue
Block a user