From db7b309e3d4998129b2c6d08a0f71b5bca1918e2 Mon Sep 17 00:00:00 2001 From: JesseMarkowitz Date: Wed, 16 Sep 2026 07:12:23 -0400 Subject: [PATCH] v1.1 closeout: accept integrated release validation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Release validation of candidate 87a4032, not a work package. No product code changed, no requirement or acceptance test changed, no schema or bundle format changed, and nothing is tagged or merged by it. V1.1 RELEASE VALIDATION: PASS What was run, on this candidate: - v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified. - Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend 175 passed; lint 0 errors (15 documented warnings); production build clean. - Docker: docker build --no-cache; the image's SPA is file-for-file identical to the local build (16 files, same combined sha256). - Offline: 23/23 against the candidate image with no network and a fresh volume. - Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over trusted-LAN HTTPS with a private CA; every narrator turn "fits". - Long run: 102 accepted turns at a verified 16,384 window with memory on, 3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks. - A1: every turn "fits"; the ten largest prompts re-counted against the server keep the documented reserve, smallest margin 879 tokens against v1's 23-42. - A2: release-gate leak count 0 across 105 stored replies. - Identity: 0 signals and 0 stored protocol shapes, with memory on; the scripted detector still fires on an injected defect. - Recovery: 16/16 on the long run's own bundle, into a database and directory that never existed. - Upgrade: a campaign built and played by the v1.0.0 application compares identical on all 15 census fields, schema parity at user_version 94, and both bundle directions import. - Release smoke: 15/15 from the shipped image — loopback only, private CA verified, public endpoint refused, a real turn, restart, persistence, and Firefox rendering the reopened campaign. Carried residuals, stated rather than summarised away: - WP-B: deterministic independent-memory recovery PASS; reference-model independent-memory recovery FAIL at memory creation — the owner-accepted limitation, unchanged and not a new regression. - The mid-reply instruction echo A2's trailing cleanup does not remove is still reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in release evidence. - The doubled full stop in the memory-search scene text. - K1 ("Correct" on an Important Facts row is refused) is classified v1.2 backlog, reproduced and not fixed during validation. Three harness corrections were made during validation — the identity diagnostic did not enable memory, the smoke test needed hostname resolution inside the container, and the first upgrade campaign was too short to write memories. All harness-only; each corrected harness repeated its own check, and no product evidence became stale. Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the released version, that v1.1 is implemented and validated, and that no v1.1.0 tag exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py. Still the owner's to do: sign the release commit, update main, tag v1.1.0. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY --- README.md | 17 +- backend/tools/m11_identity.py | 13 +- backend/tools/v11_release_smoke.py | 307 +++++++ backend/tools/v11_upgrade_check.py | 381 ++++++++ planning/README.md | 17 +- planning/V1.1-PLAN.md | 23 +- planning/VERSION.md | 29 +- planning/reports/v1.1/V1.1-RELEASE-REPORT.md | 881 +++++++++++++++++++ planning/reports/v1.1/V1.1-WP-E-REPORT.md | 14 +- 9 files changed, 1656 insertions(+), 26 deletions(-) create mode 100644 backend/tools/v11_release_smoke.py create mode 100644 backend/tools/v11_upgrade_check.py create mode 100644 planning/reports/v1.1/V1.1-RELEASE-REPORT.md diff --git a/README.md b/README.md index 5f50a48..1660370 100644 --- a/README.md +++ b/README.md @@ -293,7 +293,7 @@ player input frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI ├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug ├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile - ├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting) + ├─ migrations.py hand-rolled, versioned via PRAGMA user_version (94 and counting) ├─ endpoints.py the inference-endpoint address policy ├─ contextwindow.py what the server will actually accept, and the cap ├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's @@ -355,11 +355,20 @@ most interesting engineering in the repo. ## Repo notes -- **Status:** **v1.0.0 was released on 2026-09-14.** Milestones M1-M11 are +- **Status:** **v1.0.0 remains the released version.** Milestones M1-M11 are complete, and the v1 release gate passed (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md), §T). The signed tag `v1.0.0` and `main` both point at the signed release - commit `432f041`. v1.1 development has begun on the `v1.1-development` branch; - its plan is [`planning/V1.1-PLAN.md`](planning/V1.1-PLAN.md). + commit `432f041`. + + **v1.1 is implemented and validated, but not yet released.** All six work + packages (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) are complete and accepted on + the `v1.1-development` branch, and integrated release validation passed on + candidate `87a4032` — see + [`planning/reports/v1.1/V1.1-RELEASE-REPORT.md`](planning/reports/v1.1/V1.1-RELEASE-REPORT.md). + WP-B ships with a documented reference-model memory limitation, recorded in + that report. **No `v1.1.0` tag exists and `main` is unchanged**; the release + commit, `main` and the tag are the owner's to make. The plan is + [`planning/V1.1-PLAN.md`](planning/V1.1-PLAN.md). - `planning/` is this fork's own package: the product specification, the architecture decisions, the milestone plan, the acceptance contract, and a review report for every diff --git a/backend/tools/m11_identity.py b/backend/tools/m11_identity.py index 259e1e0..1b7272b 100644 --- a/backend/tools/m11_identity.py +++ b/backend/tools/m11_identity.py @@ -90,6 +90,11 @@ from app.routers import adventures # noqa: E402 ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "") MODEL = os.environ.get("AIDND_TEST_MODEL", "") +#: v1.1 release Gate 7 asks for this diagnostic with **memory on**. It shipped +#: with no embedding model and the bank switched off, so a release run of it +#: would have reported a clean identity result without memory ever taking part. +#: Empty keeps the old behaviour, which is what `--scripted` wants. +EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "") #: The cast the finding describes: a protagonist and three others, all on stage. CAST = [ @@ -140,7 +145,7 @@ def _setup(scripted: bool): db.add(models.Settings( user_id=user.id, model=MODEL or "scripted", endpoint_url=ENDPOINT or "http://127.0.0.1:11434/v1", - embedding_model="", context_token_budget=16384, max_output_tokens=500, + embedding_model=EMBED_MODEL, context_token_budget=16384, max_output_tokens=500, model_timeout_seconds=300, )) db.commit() @@ -167,6 +172,12 @@ def _campaign(client) -> int: }) created.raise_for_status() adv = created.json()["id"] + # Memory and summaries are per-campaign switches defaulting to off. Gate 7 + # asks for this diagnostic with memory on, and the ten beats below write + # twenty actions — past `MEMORY_START` — so the bank has something to do. + client.patch(f"/api/adventures/{adv}", + json={"memory_bank_enabled": True, "auto_summarize": True} + ).raise_for_status() answer = client.post(f"/api/adventures/{adv}/state/corrections", json={ "events": [ {"type": "create_entity", "entity": key, "entity_type": kind, "name": name} diff --git a/backend/tools/v11_release_smoke.py b/backend/tools/v11_release_smoke.py new file mode 100644 index 0000000..cf43612 --- /dev/null +++ b/backend/tools/v11_release_smoke.py @@ -0,0 +1,307 @@ +"""v1.1 release smoke test: the shipped image, as a reader would meet it. + + python -m tools.v11_release_smoke --image --out + +Run from `backend/`. Reads `AIDND_TEST_ENDPOINT` (an **HTTPS** Ollama on the +trusted LAN) and `AIDND_TEST_MODEL`. `--ca` names the private CA to install +inside the container, defaulting to this machine's own. + +Supplemental release evidence, not a replacement for the gates: it asks whether +the artefact that ships actually runs, reaches its approved narrator, refuses an +unapproved one, and keeps a campaign across a container restart. + +## The two things this is careful about + +**The CA is installed, not bypassed.** `app/tlstrust.ssl_context()` is +`ssl.create_default_context()` — the platform's own store — unioned with +certifi's. So the private CA is mounted into +`/usr/local/share/ca-certificates/` and registered with +`update-ca-certificates`, and verification is then ordinary. Nothing sets +`verify=False`, and a check inside the container proves the handshake succeeds +through that store. + +**Loopback means the published port.** The process inside the container listens +on `0.0.0.0` because that is the only address a published port can reach +(`docker-compose.yml` says so). What must be loopback-only is the *publish*, so +the container is started with `-p 127.0.0.1::8000` and the check is that +the host's LAN address refuses the same port. +""" + +from __future__ import annotations + +import argparse +import json +import os +import socket +import subprocess +import sys +import time +import urllib.error +import urllib.request +from datetime import datetime +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) + +from tools.m11_webdriver import Browser, free_port, require_under_home # noqa: E402 + +ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "") +MODEL = os.environ.get("AIDND_TEST_MODEL", "") +NAME = "v11-release-smoke" +VOLUME = "v11-release-smoke-data" +#: An endpoint the policy must refuse whatever else is true: a public host. +PUBLIC_ENDPOINT = "https://api.openai.com/v1" + + +class Checks: + def __init__(self) -> None: + self.rows: list[dict] = [] + + def record(self, name: str, ok: bool, detail: str = "") -> bool: + self.rows.append({"check": name, "result": "PASS" if ok else "FAIL", + "detail": detail}) + print(f" {'ok ' if ok else 'FAIL'} {name}" + (f" — {detail}" if detail else ""), + flush=True) + return ok + + @property + def failed(self) -> list[dict]: + return [r for r in self.rows if r["result"] == "FAIL"] + + +def run(*args: str, **kwargs) -> subprocess.CompletedProcess: + return subprocess.run(args, capture_output=True, text=True, **kwargs) + + +def api(base: str, method: str, path: str, payload=None, timeout=900): + data = json.dumps(payload).encode() if payload is not None else None + request = urllib.request.Request( + f"{base}/api{path}", data=data, method=method, + headers={"Content-Type": "application/json"} if data else {}) + with urllib.request.urlopen(request, timeout=timeout) as response: + body = response.read().decode() + return json.loads(body) if body else None + + +def stream_turn(base: str, adv: int, text: str) -> list[dict]: + request = urllib.request.Request( + f"{base}/api/adventures/{adv}/actions", + data=json.dumps({"type": "do", "text": text}).encode(), + method="POST", headers={"Content-Type": "application/json"}) + events: list[dict] = [] + with urllib.request.urlopen(request, timeout=900) as response: + for raw in response: + line = raw.decode(errors="replace").strip() + if line.startswith("data:"): + try: + events.append(json.loads(line[5:].strip())) + except json.JSONDecodeError: + pass + return events + + +def lan_address() -> str | None: + """This machine's own LAN address, for the loopback-only check.""" + probe = socket.socket(socket.AF_INET, socket.SOCK_DGRAM) + try: + probe.connect(("192.0.2.1", 9)) # TEST-NET-1: routed nowhere, sends nothing + return probe.getsockname()[0] + except OSError: + return None + finally: + probe.close() + + +def wait_ready(base: str, *, timeout: float = 180) -> bool: + deadline = time.monotonic() + timeout + while time.monotonic() < deadline: + try: + urllib.request.urlopen(f"{base}/api/settings", timeout=3) + return True + except Exception: + time.sleep(1) + return False + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--image", required=True) + parser.add_argument("--out", required=True) + parser.add_argument("--ca", default="/usr/local/share/ca-certificates/draco.crt") + parser.add_argument( + "--add-host", default="", metavar="NAME:ADDRESS", + help=("resolve the narrator's hostname inside the container. A `.local` " + "name is mDNS, and a container has no mDNS resolver, so the " + "endpoint policy refuses an address it cannot classify and " + "`PUT /api/settings` answers 400. Mapping the name — rather than " + "using the address — keeps the hostname the certificate is issued " + "for, which is the thing this test verifies.")) + args = parser.parse_args() + + if not (ENDPOINT and MODEL): + print("set AIDND_TEST_ENDPOINT (https://…) and AIDND_TEST_MODEL") + return 2 + if not ENDPOINT.startswith("https://"): + print("the smoke test needs an HTTPS endpoint: that is what it verifies") + return 2 + ca = Path(args.ca) + if not ca.exists(): + print(f"no CA at {ca}") + return 2 + + out = require_under_home(Path(args.out).expanduser()) + out.mkdir(parents=True, exist_ok=True) + checks = Checks() + port = free_port() + base = f"http://127.0.0.1:{port}" + started = datetime.now() + + run("docker", "rm", "-f", NAME) + run("docker", "volume", "rm", VOLUME) + run("docker", "volume", "create", VOLUME) + + print(f"starting {args.image} on 127.0.0.1:{port} with a fresh volume …") + start = run( + "docker", "run", "-d", "--name", NAME, + "-p", f"127.0.0.1:{port}:8000", + "-v", f"{VOLUME}:/data", + "-v", f"{ca}:/usr/local/share/ca-certificates/{ca.name}:ro", + *(("--add-host", args.add_host) if args.add_host else ()), + args.image, + "sh", "-c", + "update-ca-certificates >/dev/null 2>&1; " + "exec uvicorn app.main:app --host 0.0.0.0 --port 8000", + ) + if start.returncode != 0: + print(start.stderr[:400]) + return 1 + container = start.stdout.strip()[:12] + + try: + checks.record("the container starts", True, container) + ready = wait_ready(base) + if not checks.record("the application answers on loopback", ready, base): + logs = run("docker", "logs", NAME) + (out / "container.log").write_text(logs.stdout + logs.stderr) + return 1 + + published = run("docker", "port", NAME).stdout.strip() + checks.record("the port is published on loopback only", + "127.0.0.1" in published and "0.0.0.0" not in published, published) + + lan = lan_address() + if lan: + try: + urllib.request.urlopen(f"http://{lan}:{port}/api/settings", timeout=4) + reachable = True + except Exception: + reachable = False + checks.record("the LAN address does not serve the application", not reachable, + f"port {port} on this machine's LAN address") + + page = urllib.request.urlopen(base + "/", timeout=30) + html = page.read().decode(errors="replace") + checks.record("the first page loads", page.status == 200 and "
= 2, + errors[0].get("detail", "")[:160] if errors else + f"{page_after.get('total')} actions") + before = [(a.get("type"), (a.get("text") or "")[:120]) + for a in (page_after.get("actions") or [])] + state_before = api(base, "GET", f"/adventures/{adv}/state") or {} + + print("restarting the container …") + run("docker", "restart", NAME) + ready = wait_ready(base) + checks.record("the container restarts and serves again", ready) + + page_reopened = api(base, "GET", f"/adventures/{adv}/actions?limit=50") or {} + after = [(a.get("type"), (a.get("text") or "")[:120]) + for a in (page_reopened.get("actions") or [])] + checks.record("the transcript survived the restart", after == before, + f"{len(before)} -> {len(after)} actions") + state_after = api(base, "GET", f"/adventures/{adv}/state") or {} + checks.record("the narrative state survived the restart", + state_after == state_before) + + browser = Browser(headless=True, log=out / "geckodriver.log") + try: + browser.go(f"{base}/play/{adv}") + browser.wait_for(".story-controls", timeout=60) + story = browser.js( + "const el = document.querySelector('.story');" + " return el ? el.textContent.trim().length : 0;") + checks.record("Firefox renders the reopened campaign", + isinstance(story, int) and story > 0, f"{story} characters of story") + browser.screenshot(out / "reopened-campaign.png") + finally: + browser.quit() + finally: + logs = run("docker", "logs", NAME) + (out / "container.log").write_text(logs.stdout + logs.stderr) + run("docker", "rm", "-f", NAME) + run("docker", "volume", "rm", VOLUME) + + report = { + "image": args.image, + "started": started.isoformat(timespec="seconds"), + "seconds": round((datetime.now() - started).total_seconds()), + "endpoint_class": "trusted-LAN HTTPS with a private CA", + "checks": checks.rows, + "passed": len([r for r in checks.rows if r["result"] == "PASS"]), + "failed": len(checks.failed), + } + (out / "smoke-report.json").write_text(json.dumps(report, indent=2)) + print(f"\n{report['passed']} passed, {report['failed']} failed " + f"-> {out / 'smoke-report.json'}") + return 1 if checks.failed else 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/backend/tools/v11_upgrade_check.py b/backend/tools/v11_upgrade_check.py new file mode 100644 index 0000000..768f1b2 --- /dev/null +++ b/backend/tools/v11_upgrade_check.py @@ -0,0 +1,381 @@ +"""v1.1 release Gate 9: a real v1.0.0 campaign, opened by the candidate. + + python -m tools.v11_upgrade_check --v100 --out + +Run from `backend/`. Reads `AIDND_TEST_ENDPOINT`, `AIDND_TEST_MODEL` and +`AIDND_TEST_EMBED_MODEL`: the campaign has to be *played*, because memories, +summaries and narrative state are things a narrator produces. A schema-only +fixture would prove nothing about an upgrade, which is why §11 item 9 asks for a +database the v1.0.0 application built. + +Three phases, each its own server process, so everything that survives crosses +as bytes on disk: + +1. **v1.0.0 builds and plays.** The `432f041` tree serves the application: a + campaign is created, canon knowledge imported, turns played, a Save Point + taken, a turn undone so the head is not at the tip and Redo is available, and + a narration length chosen. Then that server stops, and a census is taken. +2. **The candidate opens the same file.** Nothing is copied; the candidate's own + migrations run against it. The census is taken again and compared field by + field. +3. **Bundles cross both ways.** The v1.0.0 export is imported by the candidate. + The candidate's export is offered back to v1.0.0, and whatever happens is + reported — the format string is unchanged, which is a reason to test backward + import, not a reason to assume it. + +Settings hold only what the caller's environment names, and the evidence +directory lives under `$HOME`. + +## Shapes this had to be written against, not guessed + +- A turn is **SSE**: `POST /adventures/{id}/actions`, and a failed turn is an + `error` *event* inside an HTTP 200. Reading the status code would call every + failure a success. +- History is `GET /{id}/actions?limit=N` -> `{actions, total, has_more, + can_undo, can_redo}`. There is **no head-id field**, so the active head is + compared as the newest action plus the two flags. +- Save Points are **checkpoints**. Creating one after an Undo names the undone + position, deliberately. +- There is **no summaries route**; summaries and post-turn health both come from + `GET /{id}/derived`. +- Campaign switches are `PATCH /adventures/{id}` with `memory_bank_enabled` and + `auto_summarize` — not the names a reader would guess. +- Knowledge import is **multipart**, as `m11_long_run` and `m11_offline` build it. +""" + +from __future__ import annotations + +import argparse +import json +import os +import shutil +import sqlite3 +import subprocess +import sys +import time +import urllib.error +import urllib.request +from datetime import datetime +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) + +from tools.m11_webdriver import free_port, require_under_home # noqa: E402 + +ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "") +MODEL = os.environ.get("AIDND_TEST_MODEL", "") +EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "") +TURN_TIMEOUT = 900 +#: What the evidence database is left holding. The campaign has to be played +#: against a real narrator, but nothing about the *upgrade* depends on which +#: host that was, and §11 item 9 asks for no real hostname in this database. +PLACEHOLDER_ENDPOINT = "http://127.0.0.1:11434/v1" + +#: What the upgrade must preserve, compared exactly on both sides. `settings` is +#: included because a migration that silently rewrote an endpoint would be a real +#: defect; `schema_version` is read from the file rather than the API. +CENSUS = ("transcript", "newest_action", "total", "can_undo", "can_redo", + "checkpoints", "state", "memories", "summaries", "knowledge", + "narration_length", "memory_bank_enabled", "auto_summarize", + "settings", "schema_version") + +CANON_MD = """# Westhaven + +The abbey bell is rung only for a death. Mara keeps the harbour ledger. +Aldric carries a silver key he will not explain. +""" + + +class App: + """One application process, from whichever tree it is given.""" + + def __init__(self, tree: Path, db: Path, log: Path, label: str): + self.tree, self.db, self.label = tree, db, label + self.port = free_port() + self.log = log + handle = open(log, "ab") + self.proc = subprocess.Popen( + [str(Path(__file__).resolve().parent.parent / ".venv/bin/uvicorn"), + "app.main:app", "--host", "127.0.0.1", "--port", str(self.port)], + cwd=str(tree / "backend"), stdout=handle, stderr=subprocess.STDOUT, + env={**os.environ, "AIDND_DB_PATH": str(db), + "AIDND_DATABASE_URL": "", "DATABASE_URL": ""}, + ) + self.url = f"http://127.0.0.1:{self.port}" + deadline = time.monotonic() + 120 + while time.monotonic() < deadline: + if self.proc.poll() is not None: + raise RuntimeError(f"{label} exited early; see {log}") + try: + urllib.request.urlopen(self.url + "/api/settings", timeout=2) + print(f" {label} serving {db.name} on {self.url}", flush=True) + return + except Exception: + time.sleep(0.2) + raise RuntimeError(f"{label} never became ready; see {log}") + + def call(self, method: str, path: str, payload=None, timeout=120): + data = json.dumps(payload).encode() if payload is not None else None + request = urllib.request.Request( + f"{self.url}/api{path}", data=data, method=method, + headers={"Content-Type": "application/json"} if data else {}) + with urllib.request.urlopen(request, timeout=timeout) as response: + body = response.read().decode() + return json.loads(body) if body else None + + def stream(self, path: str, payload) -> list[dict]: + """A turn. A failed turn is an event in the stream, not a status code.""" + request = urllib.request.Request( + f"{self.url}/api{path}", data=json.dumps(payload).encode(), + method="POST", headers={"Content-Type": "application/json"}) + events: list[dict] = [] + with urllib.request.urlopen(request, timeout=TURN_TIMEOUT) as response: + for raw in response: + line = raw.decode(errors="replace").strip() + if line.startswith("data:"): + try: + events.append(json.loads(line[5:].strip())) + except json.JSONDecodeError: + pass + return events + + def upload(self, adv: int, name: str, body: str, classification: str) -> dict: + boundary = "----v11upgrade" + parts = ( + f"--{boundary}\r\nContent-Disposition: form-data; name=\"classification\"" + f"\r\n\r\n{classification}\r\n" + f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; " + f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n" + f"--{boundary}--\r\n" + ).encode() + request = urllib.request.Request( + f"{self.url}/api/adventures/{adv}/knowledge", data=parts, method="POST", + headers={"Content-Type": f"multipart/form-data; boundary={boundary}"}) + with urllib.request.urlopen(request, timeout=120) as response: + return json.loads(response.read().decode()) + + def stop(self) -> None: + if self.proc.poll() is None: + self.proc.terminate() + try: + self.proc.wait(timeout=30) + except subprocess.TimeoutExpired: + self.proc.kill() + + +def schema_version(db: Path) -> int: + connection = sqlite3.connect(f"file:{db}?mode=ro", uri=True) + try: + return connection.execute("PRAGMA user_version").fetchone()[0] + finally: + connection.close() + + +def census(app: App, adv: int, db: Path) -> dict: + page = app.call("GET", f"/adventures/{adv}/actions?limit=500") or {} + actions = page.get("actions") or [] + derived = app.call("GET", f"/adventures/{adv}/derived") or {} + adventure = app.call("GET", f"/adventures/{adv}") or {} + settings = app.call("GET", "/settings") or {} + checkpoints = app.call("GET", f"/adventures/{adv}/checkpoints") or [] + memories = app.call("GET", f"/adventures/{adv}/memories") or [] + state = app.call("GET", f"/adventures/{adv}/state") or {} + knowledge = app.call("GET", f"/adventures/{adv}/knowledge") or [] + newest = actions[-1] if actions else {} + return { + "transcript": [(a.get("type"), (a.get("text") or "")[:300]) for a in actions], + # No head id is exposed; the head is the newest action on the read line + # plus the two flags the page carries. + "newest_action": ((newest.get("type"), (newest.get("text") or "")[:300]) + if newest else None), + "total": page.get("total"), + "can_undo": page.get("can_undo"), + "can_redo": page.get("can_redo"), + "checkpoints": sorted((c.get("name"), c.get("depth"), c.get("branch_id"), + c.get("on_path"), c.get("resolved")) + for c in checkpoints), + "state": state.get("document") if isinstance(state, dict) else state, + "memories": sorted((m.get("text") or "")[:200] for m in memories), + "summaries": sorted((s.get("text") or "")[:200] + for s in (derived.get("summaries") or [])), + "knowledge": sorted((k.get("filename") or k.get("title"), + k.get("classification")) for k in knowledge), + "narration_length": adventure.get("narration_length"), + "memory_bank_enabled": adventure.get("memory_bank_enabled"), + "auto_summarize": adventure.get("auto_summarize"), + "settings": {k: settings.get(k) + for k in ("endpoint_url", "model", "max_output_tokens")}, + "schema_version": schema_version(db), + } + + +def play(app: App, adv: int, text: str) -> bool: + events = app.stream(f"/adventures/{adv}/actions", {"type": "do", "text": text}) + errors = [e for e in events if e.get("type") == "error"] + if errors: + print(f" turn refused: {errors[0].get('detail', '')[:150]}", flush=True) + return False + return True + + +def build_v100_campaign(app: App) -> int: + created = app.call("POST", "/adventures", { + "title": "Upgrade Evidence", + "opening": "Rain over Westhaven, and the abbey bell tolling.", + "canon_rules": ["The dead do not return."], + "persona_name": "Aldric", + }) + adv = created["id"] + app.call("PUT", "/settings", { + "endpoint_url": ENDPOINT, "model": MODEL, + "embedding_model": EMBED_MODEL, + "max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT}) + # Memory and summaries are per-campaign switches defaulting to off, so the + # census would otherwise have nothing to compare. + app.call("PATCH", f"/adventures/{adv}", { + "narration_length": "brief", "memory_bank_enabled": True, + "auto_summarize": True}) + app.upload(adv, "westhaven-canon.md", CANON_MD, "canon") + + # Enough turns that memories and summaries actually exist. A memory needs + # MEMORY_INTERVAL (6) actions plus SETTLE_SLACK (1) settled past the anchor, + # and a summary needs SUMMARY_INTERVAL (15) uncovered actions — both counted + # in *actions*, and a turn writes two. A first pass at this gate played five + # turns, wrote neither, and compared 0 against 0, which proves nothing about + # whether the upgrade preserves them. + beats = [ + "I ask Mara what the bell means.", + "I show her the silver key.", + "I follow her to the harbour ledger.", + "I ask who else knows about the key.", + "I read the ledger's last page aloud.", + "I ask the ferryman about the fen road.", + "I wait out the rain and watch the harbour.", + "I ask Mara about the abbey's sealed crypt.", + "I count the entries against the tide table.", + "I ask who signed for the last shipment.", + "I walk the quay to the chandler's door.", + "I ask the chandler what he remembers of that night.", + "I show the chandler the key.", + "I return to Mara with what he said.", + "I ask Mara what she means to do now.", + "I agree to meet her at first light.", + "I take the long way back along the ridge.", + "I check whether anyone followed me.", + "I write down what I have learned so far.", + "I sleep, and wake before the bell.", + ] + played = 0 + for text in beats: + if play(app, adv, text): + played += 1 + print(f" {played} turns accepted by v1.0.0", flush=True) + + app.call("POST", f"/adventures/{adv}/checkpoints", {"name": "before the ledger"}) + # One Undo, so the head is not at the retained tip and Redo is available. + app.call("POST", f"/adventures/{adv}/undo", {}) + + # §11 item 9: the database must carry **loopback or placeholder settings with + # no real hostnames**. The turns above needed a real narrator, so the + # endpoint is reset to loopback once the story exists — before the census is + # taken and before either bundle is exported. + app.call("PUT", "/settings", { + "endpoint_url": PLACEHOLDER_ENDPOINT, "model": MODEL, + "embedding_model": EMBED_MODEL, + "max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT}) + return adv + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--v100", required=True, help="the v1.0.0 worktree") + parser.add_argument("--out", required=True) + args = parser.parse_args() + + if not (ENDPOINT and MODEL and EMBED_MODEL): + print("set AIDND_TEST_ENDPOINT, AIDND_TEST_MODEL and AIDND_TEST_EMBED_MODEL") + return 2 + + out = require_under_home(Path(args.out).expanduser()) + shutil.rmtree(out, ignore_errors=True) + out.mkdir(parents=True) + v100_tree = Path(args.v100).expanduser().resolve() + candidate_tree = Path(__file__).resolve().parent.parent.parent + db = out / "campaign.db" + results: dict = {"started": datetime.now().isoformat(timespec="seconds"), + "v100_tree": str(v100_tree), "candidate": str(candidate_tree)} + failures: list[str] = [] + + print("phase 1 — v1.0.0 builds and plays the campaign") + app = App(v100_tree, db, out / "v100-server.log", "v1.0.0") + try: + adv = build_v100_campaign(app) + before = census(app, adv, db) + v100_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600) + (out / "v100-export.json").write_text(json.dumps(v100_bundle)) + finally: + app.stop() + results["adventure"], results["before"] = adv, before + print(f" {before['total']} actions, redo={before['can_redo']}, " + f"checkpoints={len(before['checkpoints'])}, memories={len(before['memories'])}, " + f"summaries={len(before['summaries'])}, knowledge={len(before['knowledge'])}, " + f"schema={before['schema_version']}") + + print("\nphase 2 — the candidate opens that same database file") + app = App(candidate_tree, db, out / "candidate-server.log", "candidate") + try: + after = census(app, adv, db) + v11_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600) + (out / "v11-export.json").write_text(json.dumps(v11_bundle)) + try: + imported = app.call("POST", "/adventures/import", v100_bundle, timeout=600) + results["v100_bundle_into_v11"] = {"status": "imported", + "id": (imported or {}).get("id")} + print(" the v1.0.0 bundle imported into the candidate") + except urllib.error.HTTPError as exc: + results["v100_bundle_into_v11"] = { + "status": "refused", "code": exc.code, + "detail": exc.read().decode()[:300]} + failures.append("v100_bundle_into_v11") + print(f" the candidate REFUSED the v1.0.0 bundle: {exc.code}") + finally: + app.stop() + results["after"] = after + + print(" comparing the census, field by field:") + for field in CENSUS: + same = before.get(field) == after.get(field) + print(f" {'ok ' if same else 'DIFF'} {field}") + if not same: + failures.append(field) + results.setdefault("differences", {})[field] = { + "before": before.get(field), "after": after.get(field)} + + print("\nphase 3 — the candidate's bundle offered back to v1.0.0") + app = App(v100_tree, out / "backward.db", out / "v100-backward.log", "v1.0.0") + try: + try: + back = app.call("POST", "/adventures/import", v11_bundle, timeout=600) + results["v11_bundle_into_v100"] = {"status": "imported", + "id": (back or {}).get("id")} + print(" v1.0.0 ACCEPTED the v1.1 bundle") + except urllib.error.HTTPError as exc: + detail = exc.read().decode()[:400] + results["v11_bundle_into_v100"] = {"status": "refused", "code": exc.code, + "detail": detail} + # Reported, not failed: §11 item 9 asks for the result, and the + # owner's brief asks whether a refusal breaks the compatibility + # promise — a judgement, not an assertion this script may make. + print(f" v1.0.0 REFUSED the v1.1 bundle: {exc.code} {detail[:160]}") + finally: + app.stop() + + results["failures"] = failures + (out / "upgrade-report.json").write_text(json.dumps(results, indent=2, default=str)) + print(f"\n{'PASS' if not failures else 'FAIL'}: {len(failures)} field(s) differ " + f"-> {out / 'upgrade-report.json'}") + return 1 if failures else 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/planning/README.md b/planning/README.md index d4f3a18..4681287 100644 --- a/planning/README.md +++ b/planning/README.md @@ -2,14 +2,15 @@ **This file is the index. Start here.** -**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: -WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`), WP-B.2 (`0c1ba83`) and WP-C -(`59b5ebc`) are committed, and the last two planned packages — WP-D (recovery -honesty) and WP-E (control-boundary contrast) — are complete and staged for -owner review** (`reports/v1.1/V1.1-WP-D-REPORT.md`, -`reports/v1.1/V1.1-WP-E-REPORT.md`). The final browser run passed 101 checks -across all three suites with 0 failed and 0 skipped. Every planned v1.1 work -package is now implemented and reported; release validation has not begun. +**Current state:** **v1.0.0 remains the released version. Every v1.1 work +package is implemented, accepted and signed** — WP-A1/A2 (`d63804f`), WP-B.1 +(`beb17ad`), WP-B.2 (`0c1ba83`), WP-C (`59b5ebc`), WP-D and WP-E (`87a4032`). +**Integrated release validation passed on candidate `87a4032`** +(`reports/v1.1/V1.1-RELEASE-REPORT.md`): the v1 contract holds (81 PASS, H09 NOT +APPLICABLE), every suite and build passes, and the browser, offline, long-run, +identity, recovery, upgrade and smoke gates are clean. WP-B ships with a +documented reference-model memory limitation. **No release commit, no `main` +update and no `v1.1.0` tag exist yet** — those are the owner's separate events. Phase 0 complete; AI-DnD forked as the production base; **milestones M1 through M11 complete and closed**. M11 was accepted at its closeout (2026-09-14), the v1 release gate passed on the release-candidate tree, and the diff --git a/planning/V1.1-PLAN.md b/planning/V1.1-PLAN.md index e02fb12..071aefd 100644 --- a/planning/V1.1-PLAN.md +++ b/planning/V1.1-PLAN.md @@ -20,12 +20,23 @@ coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks (the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS, including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). **WP-D** (recovery honesty) and **WP-E** -(control-boundary contrast) are complete and staged for owner review, each with -its own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`. The final -browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0 -failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, -WP-D, WP-E) is now implemented and reported; release validation has not begun, -and no v1.1 version or tag exists. +(control-boundary contrast) are committed and signed as `87a4032`, each with its +own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`. + +**Integrated release validation has since run on candidate `87a4032` and +passed** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): the 82 REQUIRED v1 tests hold +(81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a +`--no-cache` image whose SPA is file-for-file identical to the local build, +offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window +run passing M01-M04 with every turn `fits` and an A2 leak count of 0, identity +0 signals / 0 protocol shapes, recovery 16/16, a real v1.0.0 upgrade comparing +identical on all 15 fields with both bundle directions importing, and a release +smoke of 15/15. + +WP-B's reference-model memory limitation is carried as an accepted residual, as +are the mid-reply instruction echo and the doubled full stop; K1 is classified as +v1.2 backlog. **No `v1.1.0` tag exists, `main` is unchanged, and no release +commit has been made** — those three remain the owner's separate events. This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1 history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1 diff --git a/planning/VERSION.md b/planning/VERSION.md index 94feaa4..30343e3 100644 --- a/planning/VERSION.md +++ b/planning/VERSION.md @@ -1,8 +1,31 @@ # Planning Package Version -- **Package:** Adventure Storyteller Planning Package v4.5 -- **Revision date:** 2026-09-15 -- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C, browser release coverage, is committed and signed as `59b5ebc` (91 checks, 0 failed, 0 skipped, with real export downloads). **WP-D (recovery honesty) and WP-E (control-boundary contrast) are complete and staged for owner review**, each with its own report (`reports/v1.1/V1.1-WP-D-REPORT.md`, `V1.1-WP-E-REPORT.md`); the final browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0 failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) is now implemented and reported. Release validation has not begun, and no v1.1 version or tag exists. +- **Package:** Adventure Storyteller Planning Package v4.6 +- **Revision date:** 2026-09-16 +- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C is signed as `59b5ebc`, and **WP-D (recovery honesty) and WP-E (control-boundary contrast) are signed as `87a4032`**. **Integrated v1.1 release validation has run on candidate `87a4032` and PASSED** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): 82 REQUIRED v1 tests hold (81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a `--no-cache` image whose SPA is file-for-file identical to the local build, offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window run passing M01-M04 with every turn `fits` and 0 protocol leaks, identity 0/0, recovery 16/16, a real v1.0.0 upgrade identical on all 15 fields with both bundle directions importing, and a release smoke of 15/15. WP-B's reference-model memory limitation remains an accepted, documented residual. **v1.0.0 is still the released version: no release commit, no `main` update and no `v1.1.0` tag exist** — those are the owner's events. + +## v4.6 — v1.1 integrated release validation (2026-09-16) + +Release validation of candidate `87a4032`, not a work package: no requirement, +acceptance test, schema, bundle format or product code changed. The evidence is +`reports/v1.1/V1.1-RELEASE-REPORT.md`, sections A-W. + +| Document | Change | Kind | +| --- | --- | --- | +| `reports/v1.1/V1.1-RELEASE-REPORT.md` | **New.** The frozen candidate, the 82-row v1 acceptance matrix, every gate's result, the A1 headroom table, the A2 leak count, the WP-B verdict kept in both halves, the residual classification, and the final decision | release report | +| `reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL` **PENDING → APPROVED**, sourced and dated to the owner's release-validation brief; the signed commit predated the review | correction of record | +| `V1.1-PLAN.md`, `planning/README.md`, `VERSION.md` | Status: WP-D/WP-E signed `87a4032`; validation passed; the three owner events still outstanding | status | +| `README.md` | v1.0.0 **remains** released; v1.1 implemented and validated but untagged; schema figure corrected to 94 | product docs | +| `backend/tools/v11_upgrade_check.py`, `backend/tools/v11_release_smoke.py` | **New**, harness only: the real-v1.0.0 upgrade gate and the release-shaped smoke test | tooling | +| `backend/tools/m11_identity.py` | Reads `AIDND_TEST_EMBED_MODEL` and enables the memory bank, so the diagnostic can run with memory on as the gate requires | tooling | + +**Requirement changes: zero. Product-code changes: zero.** + +**Outcome:** `V1.1 RELEASE VALIDATION: PASS`. Carried residuals: WP-B's +reference-model memory limitation, the mid-reply instruction echo (still +reproducible on the stored fixture, absent from release evidence), and the +doubled full stop. K1 is classified v1.2 backlog. **No release commit, no `main` +update, no `v1.1.0` tag.** ## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15) diff --git a/planning/reports/v1.1/V1.1-RELEASE-REPORT.md b/planning/reports/v1.1/V1.1-RELEASE-REPORT.md new file mode 100644 index 0000000..496ef77 --- /dev/null +++ b/planning/reports/v1.1/V1.1-RELEASE-REPORT.md @@ -0,0 +1,881 @@ +# Adventure Storyteller v1.1 — Integrated Release Validation + +**Status:** COMPLETE — **V1.1 RELEASE VALIDATION: PASS**. The decision, and what +it deliberately does not cover, is in §W. + +This report answers one question: **does this exact candidate preserve the +complete v1 contract and satisfy every accepted v1.1 package on one integrated +release tree?** It is release validation, not a work package. Nothing here adds +a feature, and no release tag is created by it. + +--- + +## A. Repository / provenance + +| | | +| --- | --- | +| **Candidate SHA** | **`87a40326a29533c8d52c9f9f41022e7b499b1de7`** | +| Branch | `v1.1-development`, up to date with `origin/v1.1-development` | +| Working tree at freeze | **clean** — nothing modified, nothing staged | +| Commit | *v1.1: harden recovery and control boundaries* (WP-D + WP-E) | +| **Owner signature** | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, made 2026-09-16 05:37:13 EDT | +| Tag at HEAD | **none** — no `v1.1.0` tag exists | +| v1.0.0 baseline | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, **an ancestor** | +| Package ancestry | `d63804f` (WP-A1/A2), `beb17ad` (WP-B.1), `0c1ba83` (WP-B.2), `59b5ebc` (WP-C) — **all ancestors** | +| Diff v1.0.0..HEAD | 61 files, +14,833 / −366 | +| LICENSE / PROVENANCE | **unchanged since v1.0.0** (empty diff) | + +### A.1 Frozen candidate identity + +| | | +| --- | --- | +| Dependency locks | `backend/requirements.txt` `sha256:ed28bc0f8970cf4e…`, `frontend/package-lock.json` `sha256:355cb370837ade01…`, `frontend/package.json` `sha256:2016580ddfa176a9…`, `backend/requirements-dev.txt` `sha256:06d7695816b201e9…` | +| Schema | `LATEST_VERSION` **94**, 93 migrations (`PRAGMA user_version`) | +| Bundle format | **`ai-dnd-adventure-v3`** | +| Import ceiling | 20 MB (`MAX_IMPORT_BODY_BYTES`), unchanged | +| Frontend build | `dist` built 2026-09-16T05:39:55, 16 files, `sha256(dist) = ea2753ad24f61959fe084f4674911acc` | +| Docker image | `storyteller:release-87a4032`, `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398`, 312 MB | +| Firefox / geckodriver | 155.0.1 / 0.37.1 (2026-09-04) | +| Docker | 29.8.0, build 88096ef | +| CPU HTTPS reference host | Ollama **0.33.0**, serving `qwen2.5:3b-instruct`, `qwen2.5:3b-instruct-16k`, `nomic-embed-text:latest`; certificate verifies through the machine's CA store with no bypass | +| GPU inference host | Ollama **0.34.0**; `qwen2.5:3b-instruct-16k` digest `21ff8cc52f375f19`, `nomic-embed-text:latest` digest `0a109f422b47e3a3`; no model resident at start | +| Evidence root | `$HOME/v11-evidence/release-87a4032/` — never `/tmp`, and no real hostname appears in any committed file | + +--- + +## B. Package acceptance inventory + +| Package | Status | Source | +| --- | --- | --- | +| **WP-A1** context-window safety reserve | **ACCEPTED** | `V1.1-WP-A1-A2-REPORT.md`, signed `d63804f` | +| **WP-A2** protocol-echo cleanup, genre-neutral prompting | **ACCEPTED** | same report and commit | +| **WP-B** independent memory | **ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION** | `V1.1-WP-B1-REPORT.md`, `V1.1-WP-B2-REPORT.md` §S, signed `beb17ad` / `0c1ba83` | +| **WP-C** browser release coverage | **ACCEPTED** | `V1.1-WP-C-REPORT.md`, signed `59b5ebc` | +| **WP-D** recovery honesty | **ACCEPTED** | `V1.1-WP-D-REPORT.md`, signed `87a4032` | +| **WP-E** control-boundary contrast | **ACCEPTED** | `V1.1-WP-E-REPORT.md`, signed `87a4032` | + +### B.1 WP-B's qualification, carried whole + +The WP-B disposition is **not** shortened to "WP-B passed" anywhere in this +report. Its own §S records: + +```text +B2.1 RANKING: PASS +B2.2 EVICTION: PASS +B2.3 EXCERPT CREATION: PASS +B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED + +DETERMINISTIC WP-B: PASS +REAL-MODEL WP-B: FAIL + +WP-B OVERALL: +ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION +``` + +In the release contract's own words (§11 items 12), that is: + +```text +DETERMINISTIC WP-B: PASS +REFERENCE-MODEL INDEPENDENT MEMORY: FAIL +OWNER ACCEPTED THE LIMITATION FOR v1.1 +``` + +The failing stage is **memory creation** — the summariser's content selection — +not ranking, eviction or injection, each of which passes deterministically. + +### B.2 WP-E screenshot approval — a correction of record + +The committed WP-E report read `OWNER SCREENSHOT APPROVAL: PENDING`, because it +was written before the owner reviewed the images. The owner's release-validation +brief (2026-09-16) states the before/after screenshots are approved and +instructs this validation to record it. The WP-E report is updated to +`APPROVED` as part of this closeout (§V), sourced to that brief and dated. No +visual code changed during release validation, so the approval stands (§Q). + +--- + +## C. v1 acceptance matrix + +Every test marked **REQUIRED FOR V1** — there are **82** — against evidence taken +on **this candidate**. Evidence types follow M11's: `browser` (the 101-check run, +§G), `campaign` (the 102-turn integrated run, §H), `container` (the offline run +on the candidate image, §F), `process` (spawned server processes — recovery §M, +upgrade §N), `suite` (the 1,723-test backend suite, §D). No historical result +from different product code is used where the contract asks for candidate +evidence. + +**Result: 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified.** + +### A — Local-first operation + +| ID | Result | Evidence on this candidate | +| --- | --- | --- | +| A01 Start application offline | **PASS** | container: first page load, fresh volume, no route and no DNS | +| A02 Storyteller loopback default | **PASS** | suite; every harness reached it on `127.0.0.1`; `docker-compose.yml` publishes `127.0.0.1:8000:8000` | +| A03 No cloud API key | **PASS** | suite; container: no secret in an export | +| A04 Campaign survives restart | **PASS** | campaign: **3 process restarts, 4 process starts**, state compared across each; container: campaigns survive a container restart | +| A05 Failed model call does not corrupt story | **PASS** | campaign: a real `failed_call` at turn 69 against an unserved model, play resumed; container: same with no model reachable | +| A06 Trusted-LAN Ollama inference | **PASS** | browser: the whole 101-check run over **trusted-LAN HTTPS with a private CA**, verification on, no bypass, storyteller loopback-bound | + +### B — Core play + +| ID | Result | Evidence | +| --- | --- | --- | +| B01 Natural language action | **PASS** | browser (real turns through the UI) + campaign (102 accepted) | +| B02 Dialogue input | **PASS** | campaign: dialogue beats in the turn list | +| B03 Continue | **PASS** | suite; browser: the Continue control present and enabled | + +### C — Story authority and state + +| ID | Result | Evidence | +| --- | --- | --- | +| C01 Campaign canon is preserved | **PASS** | campaign: canon present in the prompt on **102 of 102** turns | +| C02 Possession state | **PASS** | campaign (the silver key) + suite | +| C03 Character knowledge is not invented | **PASS** | suite | +| C04 Manual state correction | **PASS** | campaign: **2 state corrections**; browser: C3's accepted and refused corrections; suite | +| C05 Canon beats reference | **PASS** | suite | +| C06 Structured state matches accepted narrative consequence | **PASS** | campaign: real extraction across 102 turns, every event validated or refused; suite | + +### D — Non-destructive history + +| ID | Result | Evidence | +| --- | --- | --- | +| D01 Undo one turn | **PASS** | browser + campaign (`undo`) | +| D02 Minimum five undos | **PASS** | suite; campaign (`undo_redo`) | +| D04 Redo | **PASS** | browser + campaign | +| D05 Redo invalidated by new continuation | **PASS** | campaign: `diverged`, after which Redo is gone | +| D06 Retry narrator response | **PASS** | campaign: **2 retries** | +| D07 Select prior retry take | **PASS** | campaign: `take_selected` | +| D08 Retry does not delete prior take | **PASS** | campaign + suite | +| D09 Edit earlier user input | **PASS** | suite | +| D10 Edit narrator output | **PASS** | suite; browser (hostile-Markdown plants through the narrator-edit path) | +| D11 Named checkpoint | **PASS** | campaign: **2 Save Points**; browser; suite | +| D12 Restore checkpoint | **PASS** | campaign: `save_point_restored`; browser | +| D13 Restore does not delete later history | **PASS** | campaign: retained actions after the restore; recovery §M | +| D14 Delete checkpoint | **PASS** | suite; browser: the delete confirmation dialog | + +### E — Branch and derived-data isolation + +| ID | Result | Evidence | +| --- | --- | --- | +| E01 Abandoned future cannot affect active state | **PASS** | suite `test_m11_leakage.py`, with a positive control | +| E02 Abandoned memory cannot leak | **PASS** | as above | +| E03 Abandoned summary cannot leak | **PASS** | as above | +| E04 Scene state is lineage-safe | **PASS** | as above | + +### F — Long-term memory and context + +| ID | Result | Evidence | +| --- | --- | --- | +| F01 Recent turns remain coherent | **PASS** | campaign: history populated every turn, newest always included | +| F02 Old important event retrieval | **PASS** | campaign §K: the planting turn outside the window and the fact recovered — **through authoritative state**, not independent memory (§K states which) | +| F03 Prompt remains bounded | **PASS** | campaign: 1,602–14,982 tokens against a 16,384 budget across 102 turns | +| F04 Output token reserve | **PASS** | campaign: `output_reserve` 500 present and subtracted on every turn | +| F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt | +| F06 Retrieval provenance | **PASS** | campaign: knowledge and memory provenance per turn; suite | +| F07 Heuristic memory is not canon | **PASS** | suite | +| F08 Memory failure is non-fatal | **PASS** | suite; container: derived work fails with no model and turns still commit; campaign: **0 post-turn failures**, no database-lock errors | + +### G — Imported knowledge + +| ID | Result | Evidence | +| --- | --- | --- | +| G01 Import local text | **PASS** | browser: through the real file input; container: offline | +| G02 Import local Markdown | **PASS** | campaign: **3 sources imported**; browser | +| G03 Classification | **PASS** | campaign: all three classes; recovery §M confirms them after a move | +| G04 Disable knowledge source | **PASS** | suite | +| G05 Canon retrieval | **PASS** | campaign: canon passages in stored prompts; suite | +| G06 Reference retrieval | **PASS** | suite | +| G07 Inspiration is low authority | **PASS** | suite | +| G08 No automatic URL fetch | **PASS** | suite `test_egress.py`; container: no network at all, import still works | +| G09 Remote Markdown image does not auto-load | **PASS** | browser: no remote image src in the rendered story | +| G10 Prompt injection in source is treated as data | **PASS** | suite; browser: injection text rendered as text | + +### H — Security + +| ID | Result | Evidence | +| --- | --- | --- | +| H01 No unexpected outbound connections | **PASS** | container (no network at all) + suite `test_egress.py` | +| H02 No telemetry | **PASS** | suite | +| H03 No cloud provider required | **PASS** | container: a full campaign offline | +| H04 Model output cannot execute shell | **PASS** | browser + suite | +| H05 Invalid state event rejected | **PASS** | suite; campaign: refusals recorded | +| H06 Stored XSS protection | **PASS** | browser: `onerror` and `