v1.1 closeout: accept integrated release validation

Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.

V1.1 RELEASE VALIDATION: PASS

What was run, on this candidate:

- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
  0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
  175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
  to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
  trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
  3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
  keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
  scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
  that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
  identical on all 15 census fields, schema parity at user_version 94, and both
  bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
  verified, public endpoint refused, a real turn, restart, persistence, and
  Firefox rendering the reopened campaign.

Carried residuals, stated rather than summarised away:

- WP-B: deterministic independent-memory recovery PASS; reference-model
  independent-memory recovery FAIL at memory creation — the owner-accepted
  limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
  reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
  release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
  backlog, reproduced and not fixed during validation.

Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.

Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.

Still the owner's to do: sign the release commit, update main, tag v1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-16 07:12:23 -04:00
co-authored by Claude Opus 5
parent 87a40326a2
commit bdbb66d921
9 changed files with 1656 additions and 26 deletions
+13 -4
View File
@@ -293,7 +293,7 @@ player input
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug ├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile ├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting) ├─ migrations.py hand-rolled, versioned via PRAGMA user_version (94 and counting)
├─ endpoints.py the inference-endpoint address policy ├─ endpoints.py the inference-endpoint address policy
├─ contextwindow.py what the server will actually accept, and the cap ├─ contextwindow.py what the server will actually accept, and the cap
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's ├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
@@ -355,11 +355,20 @@ most interesting engineering in the repo.
## Repo notes ## Repo notes
- **Status:** **v1.0.0 was released on 2026-09-14.** Milestones M1-M11 are - **Status:** **v1.0.0 remains the released version.** Milestones M1-M11 are
complete, and the v1 release gate passed (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md), complete, and the v1 release gate passed (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md),
§T). The signed tag `v1.0.0` and `main` both point at the signed release §T). The signed tag `v1.0.0` and `main` both point at the signed release
commit `432f041`. v1.1 development has begun on the `v1.1-development` branch; commit `432f041`.
its plan is [`planning/V1.1-PLAN.md`](planning/V1.1-PLAN.md).
**v1.1 is implemented and validated, but not yet released.** All six work
packages (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) are complete and accepted on
the `v1.1-development` branch, and integrated release validation passed on
candidate `87a4032` — see
[`planning/reports/v1.1/V1.1-RELEASE-REPORT.md`](planning/reports/v1.1/V1.1-RELEASE-REPORT.md).
WP-B ships with a documented reference-model memory limitation, recorded in
that report. **No `v1.1.0` tag exists and `main` is unchanged**; the release
commit, `main` and the tag are the owner's to make. The plan is
[`planning/V1.1-PLAN.md`](planning/V1.1-PLAN.md).
- `planning/` is this fork's own package: the product specification, the architecture - `planning/` is this fork's own package: the product specification, the architecture
decisions, the milestone plan, the acceptance contract, and a review report for every decisions, the milestone plan, the acceptance contract, and a review report for every
+12 -1
View File
@@ -90,6 +90,11 @@ from app.routers import adventures # noqa: E402
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "") ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "") MODEL = os.environ.get("AIDND_TEST_MODEL", "")
#: v1.1 release Gate 7 asks for this diagnostic with **memory on**. It shipped
#: with no embedding model and the bank switched off, so a release run of it
#: would have reported a clean identity result without memory ever taking part.
#: Empty keeps the old behaviour, which is what `--scripted` wants.
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
#: The cast the finding describes: a protagonist and three others, all on stage. #: The cast the finding describes: a protagonist and three others, all on stage.
CAST = [ CAST = [
@@ -140,7 +145,7 @@ def _setup(scripted: bool):
db.add(models.Settings( db.add(models.Settings(
user_id=user.id, user_id=user.id,
model=MODEL or "scripted", endpoint_url=ENDPOINT or "http://127.0.0.1:11434/v1", model=MODEL or "scripted", endpoint_url=ENDPOINT or "http://127.0.0.1:11434/v1",
embedding_model="", context_token_budget=16384, max_output_tokens=500, embedding_model=EMBED_MODEL, context_token_budget=16384, max_output_tokens=500,
model_timeout_seconds=300, model_timeout_seconds=300,
)) ))
db.commit() db.commit()
@@ -167,6 +172,12 @@ def _campaign(client) -> int:
}) })
created.raise_for_status() created.raise_for_status()
adv = created.json()["id"] adv = created.json()["id"]
# Memory and summaries are per-campaign switches defaulting to off. Gate 7
# asks for this diagnostic with memory on, and the ten beats below write
# twenty actions — past `MEMORY_START` — so the bank has something to do.
client.patch(f"/api/adventures/{adv}",
json={"memory_bank_enabled": True, "auto_summarize": True}
).raise_for_status()
answer = client.post(f"/api/adventures/{adv}/state/corrections", json={ answer = client.post(f"/api/adventures/{adv}/state/corrections", json={
"events": [ "events": [
{"type": "create_entity", "entity": key, "entity_type": kind, "name": name} {"type": "create_entity", "entity": key, "entity_type": kind, "name": name}
+307
View File
@@ -0,0 +1,307 @@
"""v1.1 release smoke test: the shipped image, as a reader would meet it.
python -m tools.v11_release_smoke --image <tag> --out <dir under $HOME>
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT` (an **HTTPS** Ollama on the
trusted LAN) and `AIDND_TEST_MODEL`. `--ca` names the private CA to install
inside the container, defaulting to this machine's own.
Supplemental release evidence, not a replacement for the gates: it asks whether
the artefact that ships actually runs, reaches its approved narrator, refuses an
unapproved one, and keeps a campaign across a container restart.
## The two things this is careful about
**The CA is installed, not bypassed.** `app/tlstrust.ssl_context()` is
`ssl.create_default_context()` — the platform's own store — unioned with
certifi's. So the private CA is mounted into
`/usr/local/share/ca-certificates/` and registered with
`update-ca-certificates`, and verification is then ordinary. Nothing sets
`verify=False`, and a check inside the container proves the handshake succeeds
through that store.
**Loopback means the published port.** The process inside the container listens
on `0.0.0.0` because that is the only address a published port can reach
(`docker-compose.yml` says so). What must be loopback-only is the *publish*, so
the container is started with `-p 127.0.0.1:<port>:8000` and the check is that
the host's LAN address refuses the same port.
"""
from __future__ import annotations
import argparse
import json
import os
import socket
import subprocess
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.m11_webdriver import Browser, free_port, require_under_home # noqa: E402
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
NAME = "v11-release-smoke"
VOLUME = "v11-release-smoke-data"
#: An endpoint the policy must refuse whatever else is true: a public host.
PUBLIC_ENDPOINT = "https://api.openai.com/v1"
class Checks:
def __init__(self) -> None:
self.rows: list[dict] = []
def record(self, name: str, ok: bool, detail: str = "") -> bool:
self.rows.append({"check": name, "result": "PASS" if ok else "FAIL",
"detail": detail})
print(f" {'ok ' if ok else 'FAIL'} {name}" + (f" — {detail}" if detail else ""),
flush=True)
return ok
@property
def failed(self) -> list[dict]:
return [r for r in self.rows if r["result"] == "FAIL"]
def run(*args: str, **kwargs) -> subprocess.CompletedProcess:
return subprocess.run(args, capture_output=True, text=True, **kwargs)
def api(base: str, method: str, path: str, payload=None, timeout=900):
data = json.dumps(payload).encode() if payload is not None else None
request = urllib.request.Request(
f"{base}/api{path}", data=data, method=method,
headers={"Content-Type": "application/json"} if data else {})
with urllib.request.urlopen(request, timeout=timeout) as response:
body = response.read().decode()
return json.loads(body) if body else None
def stream_turn(base: str, adv: int, text: str) -> list[dict]:
request = urllib.request.Request(
f"{base}/api/adventures/{adv}/actions",
data=json.dumps({"type": "do", "text": text}).encode(),
method="POST", headers={"Content-Type": "application/json"})
events: list[dict] = []
with urllib.request.urlopen(request, timeout=900) as response:
for raw in response:
line = raw.decode(errors="replace").strip()
if line.startswith("data:"):
try:
events.append(json.loads(line[5:].strip()))
except json.JSONDecodeError:
pass
return events
def lan_address() -> str | None:
"""This machine's own LAN address, for the loopback-only check."""
probe = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
try:
probe.connect(("192.0.2.1", 9)) # TEST-NET-1: routed nowhere, sends nothing
return probe.getsockname()[0]
except OSError:
return None
finally:
probe.close()
def wait_ready(base: str, *, timeout: float = 180) -> bool:
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
try:
urllib.request.urlopen(f"{base}/api/settings", timeout=3)
return True
except Exception:
time.sleep(1)
return False
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--image", required=True)
parser.add_argument("--out", required=True)
parser.add_argument("--ca", default="/usr/local/share/ca-certificates/draco.crt")
parser.add_argument(
"--add-host", default="", metavar="NAME:ADDRESS",
help=("resolve the narrator's hostname inside the container. A `.local` "
"name is mDNS, and a container has no mDNS resolver, so the "
"endpoint policy refuses an address it cannot classify and "
"`PUT /api/settings` answers 400. Mapping the name — rather than "
"using the address — keeps the hostname the certificate is issued "
"for, which is the thing this test verifies."))
args = parser.parse_args()
if not (ENDPOINT and MODEL):
print("set AIDND_TEST_ENDPOINT (https://…) and AIDND_TEST_MODEL")
return 2
if not ENDPOINT.startswith("https://"):
print("the smoke test needs an HTTPS endpoint: that is what it verifies")
return 2
ca = Path(args.ca)
if not ca.exists():
print(f"no CA at {ca}")
return 2
out = require_under_home(Path(args.out).expanduser())
out.mkdir(parents=True, exist_ok=True)
checks = Checks()
port = free_port()
base = f"http://127.0.0.1:{port}"
started = datetime.now()
run("docker", "rm", "-f", NAME)
run("docker", "volume", "rm", VOLUME)
run("docker", "volume", "create", VOLUME)
print(f"starting {args.image} on 127.0.0.1:{port} with a fresh volume …")
start = run(
"docker", "run", "-d", "--name", NAME,
"-p", f"127.0.0.1:{port}:8000",
"-v", f"{VOLUME}:/data",
"-v", f"{ca}:/usr/local/share/ca-certificates/{ca.name}:ro",
*(("--add-host", args.add_host) if args.add_host else ()),
args.image,
"sh", "-c",
"update-ca-certificates >/dev/null 2>&1; "
"exec uvicorn app.main:app --host 0.0.0.0 --port 8000",
)
if start.returncode != 0:
print(start.stderr[:400])
return 1
container = start.stdout.strip()[:12]
try:
checks.record("the container starts", True, container)
ready = wait_ready(base)
if not checks.record("the application answers on loopback", ready, base):
logs = run("docker", "logs", NAME)
(out / "container.log").write_text(logs.stdout + logs.stderr)
return 1
published = run("docker", "port", NAME).stdout.strip()
checks.record("the port is published on loopback only",
"127.0.0.1" in published and "0.0.0.0" not in published, published)
lan = lan_address()
if lan:
try:
urllib.request.urlopen(f"http://{lan}:{port}/api/settings", timeout=4)
reachable = True
except Exception:
reachable = False
checks.record("the LAN address does not serve the application", not reachable,
f"port {port} on this machine's LAN address")
page = urllib.request.urlopen(base + "/", timeout=30)
html = page.read().decode(errors="replace")
checks.record("the first page loads", page.status == 200 and "<div id=\"root\"" in html,
f"HTTP {page.status}, {len(html)} bytes")
remote = [chunk for chunk in html.split('"')
if chunk.startswith("http://") or chunk.startswith("https://")]
checks.record("the shell references no remote origin", not remote, str(remote[:3]))
csp = page.headers.get("content-security-policy") or ""
checks.record("a CSP is served", bool(csp), csp[:80])
# The approved endpoint, verified through the private CA *inside* the
# container, with the application's own trust context and no bypass.
probe = run("docker", "exec", NAME, "python", "-c",
"import json,urllib.request,ssl,sys;"
"sys.path.insert(0,'/app/backend');"
"from app.tlstrust import ssl_context;"
f"r=urllib.request.urlopen('{ENDPOINT}/models',"
" timeout=20, context=ssl_context());"
"print(r.status)")
checks.record("the approved HTTPS narrator verifies through the private CA",
probe.returncode == 0 and "200" in probe.stdout,
(probe.stdout + probe.stderr).strip()[:160])
api(base, "PUT", "/settings", {"endpoint_url": ENDPOINT, "model": MODEL,
"max_output_tokens": 300,
"model_timeout_seconds": 600})
try:
api(base, "PUT", "/settings", {"endpoint_url": PUBLIC_ENDPOINT, "model": MODEL})
refused = False
detail = "accepted"
except urllib.error.HTTPError as exc:
refused = 400 <= exc.code < 500
detail = f"HTTP {exc.code}"
checks.record("a public endpoint is refused", refused, detail)
# Put the approved one back, whatever happened above.
api(base, "PUT", "/settings", {"endpoint_url": ENDPOINT, "model": MODEL,
"max_output_tokens": 300,
"model_timeout_seconds": 600})
created = api(base, "POST", "/adventures", {
"title": "Release Smoke",
"opening": "Rain over Westhaven, and the abbey bell tolling.",
"persona_name": "Aldric"})
adv = created["id"]
checks.record("a campaign is created", bool(adv), f"id {adv}")
events = stream_turn(base, adv, "I ask Mara what the bell means.")
errors = [e for e in events if e.get("type") == "error"]
page_after = api(base, "GET", f"/adventures/{adv}/actions?limit=50") or {}
checks.record("one real narrator turn is accepted",
not errors and (page_after.get("total") or 0) >= 2,
errors[0].get("detail", "")[:160] if errors else
f"{page_after.get('total')} actions")
before = [(a.get("type"), (a.get("text") or "")[:120])
for a in (page_after.get("actions") or [])]
state_before = api(base, "GET", f"/adventures/{adv}/state") or {}
print("restarting the container …")
run("docker", "restart", NAME)
ready = wait_ready(base)
checks.record("the container restarts and serves again", ready)
page_reopened = api(base, "GET", f"/adventures/{adv}/actions?limit=50") or {}
after = [(a.get("type"), (a.get("text") or "")[:120])
for a in (page_reopened.get("actions") or [])]
checks.record("the transcript survived the restart", after == before,
f"{len(before)} -> {len(after)} actions")
state_after = api(base, "GET", f"/adventures/{adv}/state") or {}
checks.record("the narrative state survived the restart",
state_after == state_before)
browser = Browser(headless=True, log=out / "geckodriver.log")
try:
browser.go(f"{base}/play/{adv}")
browser.wait_for(".story-controls", timeout=60)
story = browser.js(
"const el = document.querySelector('.story');"
" return el ? el.textContent.trim().length : 0;")
checks.record("Firefox renders the reopened campaign",
isinstance(story, int) and story > 0, f"{story} characters of story")
browser.screenshot(out / "reopened-campaign.png")
finally:
browser.quit()
finally:
logs = run("docker", "logs", NAME)
(out / "container.log").write_text(logs.stdout + logs.stderr)
run("docker", "rm", "-f", NAME)
run("docker", "volume", "rm", VOLUME)
report = {
"image": args.image,
"started": started.isoformat(timespec="seconds"),
"seconds": round((datetime.now() - started).total_seconds()),
"endpoint_class": "trusted-LAN HTTPS with a private CA",
"checks": checks.rows,
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
"failed": len(checks.failed),
}
(out / "smoke-report.json").write_text(json.dumps(report, indent=2))
print(f"\n{report['passed']} passed, {report['failed']} failed "
f"-> {out / 'smoke-report.json'}")
return 1 if checks.failed else 0
if __name__ == "__main__":
raise SystemExit(main())
+381
View File
@@ -0,0 +1,381 @@
"""v1.1 release Gate 9: a real v1.0.0 campaign, opened by the candidate.
python -m tools.v11_upgrade_check --v100 <worktree> --out <dir under $HOME>
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT`, `AIDND_TEST_MODEL` and
`AIDND_TEST_EMBED_MODEL`: the campaign has to be *played*, because memories,
summaries and narrative state are things a narrator produces. A schema-only
fixture would prove nothing about an upgrade, which is why §11 item 9 asks for a
database the v1.0.0 application built.
Three phases, each its own server process, so everything that survives crosses
as bytes on disk:
1. **v1.0.0 builds and plays.** The `432f041` tree serves the application: a
campaign is created, canon knowledge imported, turns played, a Save Point
taken, a turn undone so the head is not at the tip and Redo is available, and
a narration length chosen. Then that server stops, and a census is taken.
2. **The candidate opens the same file.** Nothing is copied; the candidate's own
migrations run against it. The census is taken again and compared field by
field.
3. **Bundles cross both ways.** The v1.0.0 export is imported by the candidate.
The candidate's export is offered back to v1.0.0, and whatever happens is
reported — the format string is unchanged, which is a reason to test backward
import, not a reason to assume it.
Settings hold only what the caller's environment names, and the evidence
directory lives under `$HOME`.
## Shapes this had to be written against, not guessed
- A turn is **SSE**: `POST /adventures/{id}/actions`, and a failed turn is an
`error` *event* inside an HTTP 200. Reading the status code would call every
failure a success.
- History is `GET /{id}/actions?limit=N` -> `{actions, total, has_more,
can_undo, can_redo}`. There is **no head-id field**, so the active head is
compared as the newest action plus the two flags.
- Save Points are **checkpoints**. Creating one after an Undo names the undone
position, deliberately.
- There is **no summaries route**; summaries and post-turn health both come from
`GET /{id}/derived`.
- Campaign switches are `PATCH /adventures/{id}` with `memory_bank_enabled` and
`auto_summarize` — not the names a reader would guess.
- Knowledge import is **multipart**, as `m11_long_run` and `m11_offline` build it.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sqlite3
import subprocess
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.m11_webdriver import free_port, require_under_home # noqa: E402
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
TURN_TIMEOUT = 900
#: What the evidence database is left holding. The campaign has to be played
#: against a real narrator, but nothing about the *upgrade* depends on which
#: host that was, and §11 item 9 asks for no real hostname in this database.
PLACEHOLDER_ENDPOINT = "http://127.0.0.1:11434/v1"
#: What the upgrade must preserve, compared exactly on both sides. `settings` is
#: included because a migration that silently rewrote an endpoint would be a real
#: defect; `schema_version` is read from the file rather than the API.
CENSUS = ("transcript", "newest_action", "total", "can_undo", "can_redo",
"checkpoints", "state", "memories", "summaries", "knowledge",
"narration_length", "memory_bank_enabled", "auto_summarize",
"settings", "schema_version")
CANON_MD = """# Westhaven
The abbey bell is rung only for a death. Mara keeps the harbour ledger.
Aldric carries a silver key he will not explain.
"""
class App:
"""One application process, from whichever tree it is given."""
def __init__(self, tree: Path, db: Path, log: Path, label: str):
self.tree, self.db, self.label = tree, db, label
self.port = free_port()
self.log = log
handle = open(log, "ab")
self.proc = subprocess.Popen(
[str(Path(__file__).resolve().parent.parent / ".venv/bin/uvicorn"),
"app.main:app", "--host", "127.0.0.1", "--port", str(self.port)],
cwd=str(tree / "backend"), stdout=handle, stderr=subprocess.STDOUT,
env={**os.environ, "AIDND_DB_PATH": str(db),
"AIDND_DATABASE_URL": "", "DATABASE_URL": ""},
)
self.url = f"http://127.0.0.1:{self.port}"
deadline = time.monotonic() + 120
while time.monotonic() < deadline:
if self.proc.poll() is not None:
raise RuntimeError(f"{label} exited early; see {log}")
try:
urllib.request.urlopen(self.url + "/api/settings", timeout=2)
print(f" {label} serving {db.name} on {self.url}", flush=True)
return
except Exception:
time.sleep(0.2)
raise RuntimeError(f"{label} never became ready; see {log}")
def call(self, method: str, path: str, payload=None, timeout=120):
data = json.dumps(payload).encode() if payload is not None else None
request = urllib.request.Request(
f"{self.url}/api{path}", data=data, method=method,
headers={"Content-Type": "application/json"} if data else {})
with urllib.request.urlopen(request, timeout=timeout) as response:
body = response.read().decode()
return json.loads(body) if body else None
def stream(self, path: str, payload) -> list[dict]:
"""A turn. A failed turn is an event in the stream, not a status code."""
request = urllib.request.Request(
f"{self.url}/api{path}", data=json.dumps(payload).encode(),
method="POST", headers={"Content-Type": "application/json"})
events: list[dict] = []
with urllib.request.urlopen(request, timeout=TURN_TIMEOUT) as response:
for raw in response:
line = raw.decode(errors="replace").strip()
if line.startswith("data:"):
try:
events.append(json.loads(line[5:].strip()))
except json.JSONDecodeError:
pass
return events
def upload(self, adv: int, name: str, body: str, classification: str) -> dict:
boundary = "----v11upgrade"
parts = (
f"--{boundary}\r\nContent-Disposition: form-data; name=\"classification\""
f"\r\n\r\n{classification}\r\n"
f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; "
f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n"
f"--{boundary}--\r\n"
).encode()
request = urllib.request.Request(
f"{self.url}/api/adventures/{adv}/knowledge", data=parts, method="POST",
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"})
with urllib.request.urlopen(request, timeout=120) as response:
return json.loads(response.read().decode())
def stop(self) -> None:
if self.proc.poll() is None:
self.proc.terminate()
try:
self.proc.wait(timeout=30)
except subprocess.TimeoutExpired:
self.proc.kill()
def schema_version(db: Path) -> int:
connection = sqlite3.connect(f"file:{db}?mode=ro", uri=True)
try:
return connection.execute("PRAGMA user_version").fetchone()[0]
finally:
connection.close()
def census(app: App, adv: int, db: Path) -> dict:
page = app.call("GET", f"/adventures/{adv}/actions?limit=500") or {}
actions = page.get("actions") or []
derived = app.call("GET", f"/adventures/{adv}/derived") or {}
adventure = app.call("GET", f"/adventures/{adv}") or {}
settings = app.call("GET", "/settings") or {}
checkpoints = app.call("GET", f"/adventures/{adv}/checkpoints") or []
memories = app.call("GET", f"/adventures/{adv}/memories") or []
state = app.call("GET", f"/adventures/{adv}/state") or {}
knowledge = app.call("GET", f"/adventures/{adv}/knowledge") or []
newest = actions[-1] if actions else {}
return {
"transcript": [(a.get("type"), (a.get("text") or "")[:300]) for a in actions],
# No head id is exposed; the head is the newest action on the read line
# plus the two flags the page carries.
"newest_action": ((newest.get("type"), (newest.get("text") or "")[:300])
if newest else None),
"total": page.get("total"),
"can_undo": page.get("can_undo"),
"can_redo": page.get("can_redo"),
"checkpoints": sorted((c.get("name"), c.get("depth"), c.get("branch_id"),
c.get("on_path"), c.get("resolved"))
for c in checkpoints),
"state": state.get("document") if isinstance(state, dict) else state,
"memories": sorted((m.get("text") or "")[:200] for m in memories),
"summaries": sorted((s.get("text") or "")[:200]
for s in (derived.get("summaries") or [])),
"knowledge": sorted((k.get("filename") or k.get("title"),
k.get("classification")) for k in knowledge),
"narration_length": adventure.get("narration_length"),
"memory_bank_enabled": adventure.get("memory_bank_enabled"),
"auto_summarize": adventure.get("auto_summarize"),
"settings": {k: settings.get(k)
for k in ("endpoint_url", "model", "max_output_tokens")},
"schema_version": schema_version(db),
}
def play(app: App, adv: int, text: str) -> bool:
events = app.stream(f"/adventures/{adv}/actions", {"type": "do", "text": text})
errors = [e for e in events if e.get("type") == "error"]
if errors:
print(f" turn refused: {errors[0].get('detail', '')[:150]}", flush=True)
return False
return True
def build_v100_campaign(app: App) -> int:
created = app.call("POST", "/adventures", {
"title": "Upgrade Evidence",
"opening": "Rain over Westhaven, and the abbey bell tolling.",
"canon_rules": ["The dead do not return."],
"persona_name": "Aldric",
})
adv = created["id"]
app.call("PUT", "/settings", {
"endpoint_url": ENDPOINT, "model": MODEL,
"embedding_model": EMBED_MODEL,
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
# Memory and summaries are per-campaign switches defaulting to off, so the
# census would otherwise have nothing to compare.
app.call("PATCH", f"/adventures/{adv}", {
"narration_length": "brief", "memory_bank_enabled": True,
"auto_summarize": True})
app.upload(adv, "westhaven-canon.md", CANON_MD, "canon")
# Enough turns that memories and summaries actually exist. A memory needs
# MEMORY_INTERVAL (6) actions plus SETTLE_SLACK (1) settled past the anchor,
# and a summary needs SUMMARY_INTERVAL (15) uncovered actions — both counted
# in *actions*, and a turn writes two. A first pass at this gate played five
# turns, wrote neither, and compared 0 against 0, which proves nothing about
# whether the upgrade preserves them.
beats = [
"I ask Mara what the bell means.",
"I show her the silver key.",
"I follow her to the harbour ledger.",
"I ask who else knows about the key.",
"I read the ledger's last page aloud.",
"I ask the ferryman about the fen road.",
"I wait out the rain and watch the harbour.",
"I ask Mara about the abbey's sealed crypt.",
"I count the entries against the tide table.",
"I ask who signed for the last shipment.",
"I walk the quay to the chandler's door.",
"I ask the chandler what he remembers of that night.",
"I show the chandler the key.",
"I return to Mara with what he said.",
"I ask Mara what she means to do now.",
"I agree to meet her at first light.",
"I take the long way back along the ridge.",
"I check whether anyone followed me.",
"I write down what I have learned so far.",
"I sleep, and wake before the bell.",
]
played = 0
for text in beats:
if play(app, adv, text):
played += 1
print(f" {played} turns accepted by v1.0.0", flush=True)
app.call("POST", f"/adventures/{adv}/checkpoints", {"name": "before the ledger"})
# One Undo, so the head is not at the retained tip and Redo is available.
app.call("POST", f"/adventures/{adv}/undo", {})
# §11 item 9: the database must carry **loopback or placeholder settings with
# no real hostnames**. The turns above needed a real narrator, so the
# endpoint is reset to loopback once the story exists — before the census is
# taken and before either bundle is exported.
app.call("PUT", "/settings", {
"endpoint_url": PLACEHOLDER_ENDPOINT, "model": MODEL,
"embedding_model": EMBED_MODEL,
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
return adv
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--v100", required=True, help="the v1.0.0 worktree")
parser.add_argument("--out", required=True)
args = parser.parse_args()
if not (ENDPOINT and MODEL and EMBED_MODEL):
print("set AIDND_TEST_ENDPOINT, AIDND_TEST_MODEL and AIDND_TEST_EMBED_MODEL")
return 2
out = require_under_home(Path(args.out).expanduser())
shutil.rmtree(out, ignore_errors=True)
out.mkdir(parents=True)
v100_tree = Path(args.v100).expanduser().resolve()
candidate_tree = Path(__file__).resolve().parent.parent.parent
db = out / "campaign.db"
results: dict = {"started": datetime.now().isoformat(timespec="seconds"),
"v100_tree": str(v100_tree), "candidate": str(candidate_tree)}
failures: list[str] = []
print("phase 1 — v1.0.0 builds and plays the campaign")
app = App(v100_tree, db, out / "v100-server.log", "v1.0.0")
try:
adv = build_v100_campaign(app)
before = census(app, adv, db)
v100_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
(out / "v100-export.json").write_text(json.dumps(v100_bundle))
finally:
app.stop()
results["adventure"], results["before"] = adv, before
print(f" {before['total']} actions, redo={before['can_redo']}, "
f"checkpoints={len(before['checkpoints'])}, memories={len(before['memories'])}, "
f"summaries={len(before['summaries'])}, knowledge={len(before['knowledge'])}, "
f"schema={before['schema_version']}")
print("\nphase 2 — the candidate opens that same database file")
app = App(candidate_tree, db, out / "candidate-server.log", "candidate")
try:
after = census(app, adv, db)
v11_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
(out / "v11-export.json").write_text(json.dumps(v11_bundle))
try:
imported = app.call("POST", "/adventures/import", v100_bundle, timeout=600)
results["v100_bundle_into_v11"] = {"status": "imported",
"id": (imported or {}).get("id")}
print(" the v1.0.0 bundle imported into the candidate")
except urllib.error.HTTPError as exc:
results["v100_bundle_into_v11"] = {
"status": "refused", "code": exc.code,
"detail": exc.read().decode()[:300]}
failures.append("v100_bundle_into_v11")
print(f" the candidate REFUSED the v1.0.0 bundle: {exc.code}")
finally:
app.stop()
results["after"] = after
print(" comparing the census, field by field:")
for field in CENSUS:
same = before.get(field) == after.get(field)
print(f" {'ok ' if same else 'DIFF'} {field}")
if not same:
failures.append(field)
results.setdefault("differences", {})[field] = {
"before": before.get(field), "after": after.get(field)}
print("\nphase 3 — the candidate's bundle offered back to v1.0.0")
app = App(v100_tree, out / "backward.db", out / "v100-backward.log", "v1.0.0")
try:
try:
back = app.call("POST", "/adventures/import", v11_bundle, timeout=600)
results["v11_bundle_into_v100"] = {"status": "imported",
"id": (back or {}).get("id")}
print(" v1.0.0 ACCEPTED the v1.1 bundle")
except urllib.error.HTTPError as exc:
detail = exc.read().decode()[:400]
results["v11_bundle_into_v100"] = {"status": "refused", "code": exc.code,
"detail": detail}
# Reported, not failed: §11 item 9 asks for the result, and the
# owner's brief asks whether a refusal breaks the compatibility
# promise — a judgement, not an assertion this script may make.
print(f" v1.0.0 REFUSED the v1.1 bundle: {exc.code} {detail[:160]}")
finally:
app.stop()
results["failures"] = failures
(out / "upgrade-report.json").write_text(json.dumps(results, indent=2, default=str))
print(f"\n{'PASS' if not failures else 'FAIL'}: {len(failures)} field(s) differ "
f"-> {out / 'upgrade-report.json'}")
return 1 if failures else 0
if __name__ == "__main__":
raise SystemExit(main())
+9 -8
View File
@@ -2,14 +2,15 @@
**This file is the index. Start here.** **This file is the index. Start here.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: **Current state:** **v1.0.0 remains the released version. Every v1.1 work
WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`), WP-B.2 (`0c1ba83`) and WP-C package is implemented, accepted and signed** — WP-A1/A2 (`d63804f`), WP-B.1
(`59b5ebc`) are committed, and the last two planned packages — WP-D (recovery (`beb17ad`), WP-B.2 (`0c1ba83`), WP-C (`59b5ebc`), WP-D and WP-E (`87a4032`).
honesty) and WP-E (control-boundary contrast) — are complete and staged for **Integrated release validation passed on candidate `87a4032`**
owner review** (`reports/v1.1/V1.1-WP-D-REPORT.md`, (`reports/v1.1/V1.1-RELEASE-REPORT.md`): the v1 contract holds (81 PASS, H09 NOT
`reports/v1.1/V1.1-WP-E-REPORT.md`). The final browser run passed 101 checks APPLICABLE), every suite and build passes, and the browser, offline, long-run,
across all three suites with 0 failed and 0 skipped. Every planned v1.1 work identity, recovery, upgrade and smoke gates are clean. WP-B ships with a
package is now implemented and reported; release validation has not begun. documented reference-model memory limitation. **No release commit, no `main`
update and no `v1.1.0` tag exist yet** — those are the owner's separate events.
Phase 0 complete; AI-DnD forked as the production base; **milestones M1 Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the (2026-09-14), the v1 release gate passed on the release-candidate tree, and the
+17 -6
View File
@@ -20,12 +20,23 @@ coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS, (the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`). including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`).
**WP-D** (recovery honesty) and **WP-E** **WP-D** (recovery honesty) and **WP-E**
(control-boundary contrast) are complete and staged for owner review, each with (control-boundary contrast) are committed and signed as `87a4032`, each with its
its own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`. The final own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`.
browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0
failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, **Integrated release validation has since run on candidate `87a4032` and
WP-D, WP-E) is now implemented and reported; release validation has not begun, passed** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): the 82 REQUIRED v1 tests hold
and no v1.1 version or tag exists. (81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a
`--no-cache` image whose SPA is file-for-file identical to the local build,
offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window
run passing M01-M04 with every turn `fits` and an A2 leak count of 0, identity
0 signals / 0 protocol shapes, recovery 16/16, a real v1.0.0 upgrade comparing
identical on all 15 fields with both bundle directions importing, and a release
smoke of 15/15.
WP-B's reference-model memory limitation is carried as an accepted residual, as
are the mid-reply instruction echo and the doubled full stop; K1 is classified as
v1.2 backlog. **No `v1.1.0` tag exists, `main` is unchanged, and no release
commit has been made** — those three remain the owner's separate events.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1 This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1 history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
+26 -3
View File
@@ -1,8 +1,31 @@
# Planning Package Version # Planning Package Version
- **Package:** Adventure Storyteller Planning Package v4.5 - **Package:** Adventure Storyteller Planning Package v4.6
- **Revision date:** 2026-09-15 - **Revision date:** 2026-09-16
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C, browser release coverage, is committed and signed as `59b5ebc` (91 checks, 0 failed, 0 skipped, with real export downloads). **WP-D (recovery honesty) and WP-E (control-boundary contrast) are complete and staged for owner review**, each with its own report (`reports/v1.1/V1.1-WP-D-REPORT.md`, `V1.1-WP-E-REPORT.md`); the final browser run covers all three suites — M11 38, WP-C 53, WP-E 10: **101 passed, 0 failed, 0 skipped**. Every planned v1.1 work package (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) is now implemented and reported. Release validation has not begun, and no v1.1 version or tag exists. - **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C is signed as `59b5ebc`, and **WP-D (recovery honesty) and WP-E (control-boundary contrast) are signed as `87a4032`**. **Integrated v1.1 release validation has run on candidate `87a4032` and PASSED** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): 82 REQUIRED v1 tests hold (81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a `--no-cache` image whose SPA is file-for-file identical to the local build, offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window run passing M01-M04 with every turn `fits` and 0 protocol leaks, identity 0/0, recovery 16/16, a real v1.0.0 upgrade identical on all 15 fields with both bundle directions importing, and a release smoke of 15/15. WP-B's reference-model memory limitation remains an accepted, documented residual. **v1.0.0 is still the released version: no release commit, no `main` update and no `v1.1.0` tag exist** — those are the owner's events.
## v4.6 — v1.1 integrated release validation (2026-09-16)
Release validation of candidate `87a4032`, not a work package: no requirement,
acceptance test, schema, bundle format or product code changed. The evidence is
`reports/v1.1/V1.1-RELEASE-REPORT.md`, sections A-W.
| Document | Change | Kind |
| --- | --- | --- |
| `reports/v1.1/V1.1-RELEASE-REPORT.md` | **New.** The frozen candidate, the 82-row v1 acceptance matrix, every gate's result, the A1 headroom table, the A2 leak count, the WP-B verdict kept in both halves, the residual classification, and the final decision | release report |
| `reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL` **PENDING → APPROVED**, sourced and dated to the owner's release-validation brief; the signed commit predated the review | correction of record |
| `V1.1-PLAN.md`, `planning/README.md`, `VERSION.md` | Status: WP-D/WP-E signed `87a4032`; validation passed; the three owner events still outstanding | status |
| `README.md` | v1.0.0 **remains** released; v1.1 implemented and validated but untagged; schema figure corrected to 94 | product docs |
| `backend/tools/v11_upgrade_check.py`, `backend/tools/v11_release_smoke.py` | **New**, harness only: the real-v1.0.0 upgrade gate and the release-shaped smoke test | tooling |
| `backend/tools/m11_identity.py` | Reads `AIDND_TEST_EMBED_MODEL` and enables the memory bank, so the diagnostic can run with memory on as the gate requires | tooling |
**Requirement changes: zero. Product-code changes: zero.**
**Outcome:** `V1.1 RELEASE VALIDATION: PASS`. Carried residuals: WP-B's
reference-model memory limitation, the mid-reply instruction echo (still
reproducible on the stored fixture, absent from release evidence), and the
doubled full stop. K1 is classified v1.2 backlog. **No release commit, no `main`
update, no `v1.1.0` tag.**
## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15) ## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15)
@@ -0,0 +1,881 @@
# Adventure Storyteller v1.1 — Integrated Release Validation
**Status:** COMPLETE — **V1.1 RELEASE VALIDATION: PASS**. The decision, and what
it deliberately does not cover, is in §W.
This report answers one question: **does this exact candidate preserve the
complete v1 contract and satisfy every accepted v1.1 package on one integrated
release tree?** It is release validation, not a work package. Nothing here adds
a feature, and no release tag is created by it.
---
## A. Repository / provenance
| | |
| --- | --- |
| **Candidate SHA** | **`87a40326a29533c8d52c9f9f41022e7b499b1de7`** |
| Branch | `v1.1-development`, up to date with `origin/v1.1-development` |
| Working tree at freeze | **clean** — nothing modified, nothing staged |
| Commit | *v1.1: harden recovery and control boundaries* (WP-D + WP-E) |
| **Owner signature** | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, made 2026-09-16 05:37:13 EDT |
| Tag at HEAD | **none** — no `v1.1.0` tag exists |
| v1.0.0 baseline | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, **an ancestor** |
| Package ancestry | `d63804f` (WP-A1/A2), `beb17ad` (WP-B.1), `0c1ba83` (WP-B.2), `59b5ebc` (WP-C) — **all ancestors** |
| Diff v1.0.0..HEAD | 61 files, +14,833 / −366 |
| LICENSE / PROVENANCE | **unchanged since v1.0.0** (empty diff) |
### A.1 Frozen candidate identity
| | |
| --- | --- |
| Dependency locks | `backend/requirements.txt` `sha256:ed28bc0f8970cf4e…`, `frontend/package-lock.json` `sha256:355cb370837ade01…`, `frontend/package.json` `sha256:2016580ddfa176a9…`, `backend/requirements-dev.txt` `sha256:06d7695816b201e9…` |
| Schema | `LATEST_VERSION` **94**, 93 migrations (`PRAGMA user_version`) |
| Bundle format | **`ai-dnd-adventure-v3`** |
| Import ceiling | 20 MB (`MAX_IMPORT_BODY_BYTES`), unchanged |
| Frontend build | `dist` built 2026-09-16T05:39:55, 16 files, `sha256(dist) = ea2753ad24f61959fe084f4674911acc` |
| Docker image | `storyteller:release-87a4032`, `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398`, 312 MB |
| Firefox / geckodriver | 155.0.1 / 0.37.1 (2026-09-04) |
| Docker | 29.8.0, build 88096ef |
| CPU HTTPS reference host | Ollama **0.33.0**, serving `qwen2.5:3b-instruct`, `qwen2.5:3b-instruct-16k`, `nomic-embed-text:latest`; certificate verifies through the machine's CA store with no bypass |
| GPU inference host | Ollama **0.34.0**; `qwen2.5:3b-instruct-16k` digest `21ff8cc52f375f19`, `nomic-embed-text:latest` digest `0a109f422b47e3a3`; no model resident at start |
| Evidence root | `$HOME/v11-evidence/release-87a4032/` — never `/tmp`, and no real hostname appears in any committed file |
---
## B. Package acceptance inventory
| Package | Status | Source |
| --- | --- | --- |
| **WP-A1** context-window safety reserve | **ACCEPTED** | `V1.1-WP-A1-A2-REPORT.md`, signed `d63804f` |
| **WP-A2** protocol-echo cleanup, genre-neutral prompting | **ACCEPTED** | same report and commit |
| **WP-B** independent memory | **ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION** | `V1.1-WP-B1-REPORT.md`, `V1.1-WP-B2-REPORT.md` §S, signed `beb17ad` / `0c1ba83` |
| **WP-C** browser release coverage | **ACCEPTED** | `V1.1-WP-C-REPORT.md`, signed `59b5ebc` |
| **WP-D** recovery honesty | **ACCEPTED** | `V1.1-WP-D-REPORT.md`, signed `87a4032` |
| **WP-E** control-boundary contrast | **ACCEPTED** | `V1.1-WP-E-REPORT.md`, signed `87a4032` |
### B.1 WP-B's qualification, carried whole
The WP-B disposition is **not** shortened to "WP-B passed" anywhere in this
report. Its own §S records:
```text
B2.1 RANKING: PASS
B2.2 EVICTION: PASS
B2.3 EXCERPT CREATION: PASS
B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED
DETERMINISTIC WP-B: PASS
REAL-MODEL WP-B: FAIL
WP-B OVERALL:
ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION
```
In the release contract's own words (§11 items 12), that is:
```text
DETERMINISTIC WP-B: PASS
REFERENCE-MODEL INDEPENDENT MEMORY: FAIL
OWNER ACCEPTED THE LIMITATION FOR v1.1
```
The failing stage is **memory creation** — the summariser's content selection —
not ranking, eviction or injection, each of which passes deterministically.
### B.2 WP-E screenshot approval — a correction of record
The committed WP-E report read `OWNER SCREENSHOT APPROVAL: PENDING`, because it
was written before the owner reviewed the images. The owner's release-validation
brief (2026-09-16) states the before/after screenshots are approved and
instructs this validation to record it. The WP-E report is updated to
`APPROVED` as part of this closeout (§V), sourced to that brief and dated. No
visual code changed during release validation, so the approval stands (§Q).
---
## C. v1 acceptance matrix
Every test marked **REQUIRED FOR V1** — there are **82** — against evidence taken
on **this candidate**. Evidence types follow M11's: `browser` (the 101-check run,
§G), `campaign` (the 102-turn integrated run, §H), `container` (the offline run
on the candidate image, §F), `process` (spawned server processes — recovery §M,
upgrade §N), `suite` (the 1,723-test backend suite, §D). No historical result
from different product code is used where the contract asks for candidate
evidence.
**Result: 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified.**
### A — Local-first operation
| ID | Result | Evidence on this candidate |
| --- | --- | --- |
| A01 Start application offline | **PASS** | container: first page load, fresh volume, no route and no DNS |
| A02 Storyteller loopback default | **PASS** | suite; every harness reached it on `127.0.0.1`; `docker-compose.yml` publishes `127.0.0.1:8000:8000` |
| A03 No cloud API key | **PASS** | suite; container: no secret in an export |
| A04 Campaign survives restart | **PASS** | campaign: **3 process restarts, 4 process starts**, state compared across each; container: campaigns survive a container restart |
| A05 Failed model call does not corrupt story | **PASS** | campaign: a real `failed_call` at turn 69 against an unserved model, play resumed; container: same with no model reachable |
| A06 Trusted-LAN Ollama inference | **PASS** | browser: the whole 101-check run over **trusted-LAN HTTPS with a private CA**, verification on, no bypass, storyteller loopback-bound |
### B — Core play
| ID | Result | Evidence |
| --- | --- | --- |
| B01 Natural language action | **PASS** | browser (real turns through the UI) + campaign (102 accepted) |
| B02 Dialogue input | **PASS** | campaign: dialogue beats in the turn list |
| B03 Continue | **PASS** | suite; browser: the Continue control present and enabled |
### C — Story authority and state
| ID | Result | Evidence |
| --- | --- | --- |
| C01 Campaign canon is preserved | **PASS** | campaign: canon present in the prompt on **102 of 102** turns |
| C02 Possession state | **PASS** | campaign (the silver key) + suite |
| C03 Character knowledge is not invented | **PASS** | suite |
| C04 Manual state correction | **PASS** | campaign: **2 state corrections**; browser: C3's accepted and refused corrections; suite |
| C05 Canon beats reference | **PASS** | suite |
| C06 Structured state matches accepted narrative consequence | **PASS** | campaign: real extraction across 102 turns, every event validated or refused; suite |
### D — Non-destructive history
| ID | Result | Evidence |
| --- | --- | --- |
| D01 Undo one turn | **PASS** | browser + campaign (`undo`) |
| D02 Minimum five undos | **PASS** | suite; campaign (`undo_redo`) |
| D04 Redo | **PASS** | browser + campaign |
| D05 Redo invalidated by new continuation | **PASS** | campaign: `diverged`, after which Redo is gone |
| D06 Retry narrator response | **PASS** | campaign: **2 retries** |
| D07 Select prior retry take | **PASS** | campaign: `take_selected` |
| D08 Retry does not delete prior take | **PASS** | campaign + suite |
| D09 Edit earlier user input | **PASS** | suite |
| D10 Edit narrator output | **PASS** | suite; browser (hostile-Markdown plants through the narrator-edit path) |
| D11 Named checkpoint | **PASS** | campaign: **2 Save Points**; browser; suite |
| D12 Restore checkpoint | **PASS** | campaign: `save_point_restored`; browser |
| D13 Restore does not delete later history | **PASS** | campaign: retained actions after the restore; recovery §M |
| D14 Delete checkpoint | **PASS** | suite; browser: the delete confirmation dialog |
### E — Branch and derived-data isolation
| ID | Result | Evidence |
| --- | --- | --- |
| E01 Abandoned future cannot affect active state | **PASS** | suite `test_m11_leakage.py`, with a positive control |
| E02 Abandoned memory cannot leak | **PASS** | as above |
| E03 Abandoned summary cannot leak | **PASS** | as above |
| E04 Scene state is lineage-safe | **PASS** | as above |
### F — Long-term memory and context
| ID | Result | Evidence |
| --- | --- | --- |
| F01 Recent turns remain coherent | **PASS** | campaign: history populated every turn, newest always included |
| F02 Old important event retrieval | **PASS** | campaign §K: the planting turn outside the window and the fact recovered — **through authoritative state**, not independent memory (§K states which) |
| F03 Prompt remains bounded | **PASS** | campaign: 1,602–14,982 tokens against a 16,384 budget across 102 turns |
| F04 Output token reserve | **PASS** | campaign: `output_reserve` 500 present and subtracted on every turn |
| F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt |
| F06 Retrieval provenance | **PASS** | campaign: knowledge and memory provenance per turn; suite |
| F07 Heuristic memory is not canon | **PASS** | suite |
| F08 Memory failure is non-fatal | **PASS** | suite; container: derived work fails with no model and turns still commit; campaign: **0 post-turn failures**, no database-lock errors |
### G — Imported knowledge
| ID | Result | Evidence |
| --- | --- | --- |
| G01 Import local text | **PASS** | browser: through the real file input; container: offline |
| G02 Import local Markdown | **PASS** | campaign: **3 sources imported**; browser |
| G03 Classification | **PASS** | campaign: all three classes; recovery §M confirms them after a move |
| G04 Disable knowledge source | **PASS** | suite |
| G05 Canon retrieval | **PASS** | campaign: canon passages in stored prompts; suite |
| G06 Reference retrieval | **PASS** | suite |
| G07 Inspiration is low authority | **PASS** | suite |
| G08 No automatic URL fetch | **PASS** | suite `test_egress.py`; container: no network at all, import still works |
| G09 Remote Markdown image does not auto-load | **PASS** | browser: no remote image src in the rendered story |
| G10 Prompt injection in source is treated as data | **PASS** | suite; browser: injection text rendered as text |
### H — Security
| ID | Result | Evidence |
| --- | --- | --- |
| H01 No unexpected outbound connections | **PASS** | container (no network at all) + suite `test_egress.py` |
| H02 No telemetry | **PASS** | suite |
| H03 No cloud provider required | **PASS** | container: a full campaign offline |
| H04 Model output cannot execute shell | **PASS** | browser + suite |
| H05 Invalid state event rejected | **PASS** | suite; campaign: refusals recorded |
| H06 Stored XSS protection | **PASS** | browser: `onerror` and `<script>` in accepted narration, neither executed |
| H07 JavaScript URL protection | **PASS** | browser: no `javascript:` href in the DOM |
| H08 Path traversal import rejected | **PASS** | suite |
| **H09 ZIP Slip protection** | **NOT APPLICABLE** | the product extracts no archives, and a test enforces it — the same condition v1 recorded |
| H10 Restrictive CORS and local API behaviour | **PASS** | browser: an unknown API path is a 404 with a non-HTML body; suite |
| H11 No first-use runtime asset download | **PASS** | container: every asset local with no network; suite: the tokenizer table is vendored |
| H12 Inference endpoint enforcement | **PASS** | suite; smoke §R: a public endpoint refused by the running image; the container's refusal of an unresolvable name is the same policy (§R) |
### I — Export, import and recovery
| ID | Result | Evidence |
| --- | --- | --- |
| I01 Export campaign | **PASS** | campaign: the 102-turn campaign exported (3,071,683 bytes) |
| I02 Import exported campaign | **PASS** | process §M: imported into a database and directory that never existed |
| I03 Branch/disposable history export | **PASS** | §M: **88 actions retained beyond the active line** after the move |
| I04 Checkpoint export | **PASS** | §M: both Save Points restore after the move |
| I05 Knowledge provenance export | **PASS** | §M: all three classes with their content |
| I06 Database/export contains no API secrets | **PASS** | §M: the bundle carries no secret; suite; container |
| I07 Export/import preserves an undone active head | **PASS** | §N: the v1.0.0 campaign's undone head and `can_redo` survive upgrade and both bundle directions; suite `test_m9_portability.py`. *(The long-run bundle ended head-at-tip, so §M exercises the other case — stated in §M rather than implied)* |
### J — Genre neutrality
| ID | Result | Evidence |
| --- | --- | --- |
| J01 Science-fiction campaign | **PASS** | suite `test_m11_scifi.py` |
| J02 Generic entity support | **PASS** | as above |
| J03 Genre profiles are configuration | **PASS** | as above |
### K — Future media architecture
| ID | Result | Evidence |
| --- | --- | --- |
| K01 Scene snapshot exists | **PASS** | suite; container: a scene packet builds offline |
| K02 Visual character profile | **PASS** | suite |
| K03 Visual location profile | **PASS** | suite |
### L — Data integrity
| ID | Result | Evidence |
| --- | --- | --- |
| L01 Atomic turn commit | **PASS** | container + campaign: a real induced failure, no narration accepted, no half-written state |
| L02 State reconstruction | **PASS** | suite; campaign: state compared across 3 restarts |
| L03 Checkpoint reconstruction after restart | **PASS** | campaign + §M |
### M — Long-run
| ID | Result | Evidence |
| --- | --- | --- |
| **M01** 100-turn campaign | **PASS** | **102 accepted turns**, every scheduled operation exercised, **0 post-turn failures** (§H) |
| **M02** Restart during long campaign | **PASS** | **3 genuine process restarts** (4 process starts); everything crossed as bytes on disk |
| **M03** Long-run context stability | **PASS** | the prompt held **13,492–14,982** tokens over the last 70 turns against a 16,384 budget; window verified **102/102**; canon present **102/102** |
| **M04** Long-run memory recall | **PASS**, qualified | the planted clue was outside the history window (planted depth 1, floor 72) and reached the prompt: `m04_verdict: **recovered_through_state_only**`. Recovery was **through authoritative state**, not independent memory — §K states this distinction and does not relabel it |
### SHOULD and FUTURE
Not counted as REQUIRED. **SHOULD:** B04, D03, K04, L04 — all still pass on the
candidate (suite; browser for B04's direction toggle). **FUTURE:** K05, K06 —
deliberately not run; both need a media provider this release does not build.
## D. Backend / frontend suites
| Suite | Result |
| --- | --- |
| **Backend** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,210.6 s) |
| **Frontend** (`npm test`) | **175 passed**, 15 files, 0 failed |
| **Lint** (`npm run lint`, oxlint) | **exit 0 — 0 errors**, 15 warnings |
| **Production build** (`npm run build`) | succeeded |
**Every skip explained — one category, and it is the expected one.** All 17 are
environment-gated real-model tests, skipped because `AIDND_TEST_*` is
deliberately unset for the deterministic suite:
| File | Skipped | Gate |
| --- | --- | --- |
| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` + `AIDND_TEST_MODEL` |
| `test_narrative_realistic.py` | 3 | same |
| `test_m11_real_window.py` | 3 | same; one needs `AIDND_TEST_WIDE_MODEL` |
| `test_provider_wiring.py` | 1 | same |
**0 xfailed.** No deficiency WP-B fixed remains parked as an expected failure —
B.1's two strict xfails became ordinary passes in B.2 and stayed that way.
**Lint warnings are the documented, unchanged set:** the pre-existing
`only-export-components` and unused-import warnings recorded at WP-C, WP-D and
WP-E. None is in a file this candidate changed relative to those packages.
## E. Production build and Docker image
| | |
| --- | --- |
| Command | the repository's documented production build (`DEVELOPMENT.md`), with `--no-cache` |
| Image | `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398` (312 MB) |
| Dependencies | **installed, not reused** — `npm ci` and `pip install --no-cache-dir` both executed in the log |
| `CACHED` steps | **2**, and both are `WORKDIR` metadata (`/build`, `/app`) — no dependency or source layer was cached |
| **Image SPA vs local build** | **file-for-file identical**: 16 files each, `diff -r` clean, combined `sha256 = ea2753ad24f61959fe084f4674911acc` on both sides |
The image is built from the candidate tree for this validation. No earlier
work-package image was reused.
## F. Offline / no-network
`tools/m11_offline.py` against **the candidate image**, `--network none`, fresh
volume. Evidence: `…/release-87a4032/offline/offline-report.json`.
**23 checks, 23 passed, 0 failed.** Including: no route to the public Internet;
no external DNS; first page load; no remote origin named; CSP served; every
referenced asset local; campaign creation; state extraction; local file import;
prompt assembly; knowledge search; a turn with no model reachable reported as a
failure with no narration accepted, the player's words kept and state unchanged;
export; import with state; no secret in the export; the media module inert with
no provider; campaigns surviving a container restart.
**No first-use download occurred**, which is what the container's absent network
makes unfalsifiable rather than merely unobserved.
## G. Browser release validation
`tools/m11_browser.py` on the candidate, production build served by FastAPI,
Firefox 155.0.1 / geckodriver 0.37.1, narrator over **trusted-LAN HTTPS with the
private CA** (`endpoint_class: trusted-LAN HTTPS`), storyteller on loopback.
`kind: release regression` — not partial, not `--only`. 693 s.
| Suite | Passed | Failed | Skipped |
| --- | --- | --- | --- |
| **M11** (the v1 release regression) | **38** | **0** | **0** |
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
| **WP-E** (control boundaries) | **10** | **0** | **0** |
| **Total** | **101** | **0** | **0** |
**A1 accounting:** 8 narrator turns, **`fits` on every one**; `turns_not_clean`
**0**; `protocol_shapes_in_narration` **0**. The verified window on this host was
4,096 (`source: loaded`); the 16,384 evidence is the long run's (§I).
WP-C's proofs are all present in the 53: Retry and alternate takes, Save Point
create/restore/Redo, state correction including a refusal shown as a refusal,
narration length reaching the prompt, failed generation and recovery, and **real
export downloads** from both the library and campaign settings, each file landing
on disk and importing into a fresh application. WP-E's ten are the rendered
boundary measurements of §Q.
## H. Integrated 100+ turn long run
One new campaign on the candidate's product code, GPU inference host, with the
owner's power/link/kernel logging running before the first turn.
Evidence: `…/release-87a4032/long-run/`.
| | |
| --- | --- |
| Narrator | **`qwen2.5:3b-instruct-16k`**, digest `21ff8cc52f375f19` |
| Embeddings | **`nomic-embed-text`**, digest `0a109f422b47e3a3` |
| Window | **16,384**, `window_verified` on **102 of 102** turns |
| Memory / summaries | **on** — 19 memories in the bank, **12 summaries** written |
| **Accepted turns** | **102** (target 100) |
| Restarts | **3** genuine process restarts, 4 process starts |
| Elapsed | 1,330 s |
| Export | 3,071,683 bytes; database 2,613,248 bytes |
| Status | `complete`; `aborted_reason` null, `failed_reason` null |
**Not a repeated-turn benchmark.** Every scheduled operation fired and is in the
timeline: 3 restarts, 1 undo, 1 undo→redo, 2 retries, 1 take selection, 2 Save
Points, 1 Save Point restore, 1 divergence, 2 state corrections, 3 knowledge
imports, memory activation, 1 deliberate failed call, 1 export, the planted clue
and the planted independent fact, and the recall probe.
### H.1 M01–M04
| | Verdict | What decides it |
| --- | --- | --- |
| **M01** 100-turn campaign | **PASS** | 102 accepted turns, each with a committed action and state document |
| **M02** Restart during long campaign | **PASS** | 3 genuine `uvicorn` restarts; everything that survived crossed as bytes on disk |
| **M03** Long-run context stability | **PASS** | prompt 13,492–14,982 tokens over the last 70 turns against a 16,384 budget; canon present 102/102; window verified 102/102 |
| **M04** Long-run memory recall | **PASS**, and qualified | the clue was planted at depth 1, the history floor reached depth 72, and it was **not** in the recent window; it reached the prompt through **authoritative state**. `m04_verdict: recovered_through_state_only`. §K keeps the distinction the criterion was written around |
### H.2 State and derived-work integrity
**0 post-turn failures across 102 turns, and no `database is locked` error.** The
only failure-shaped events in the whole timeline are the two the run creates on
purpose: the scheduled `failed_call` at turn 69 (a model name the server does not
serve — A05/L01 evidence) and the independent-fact precondition notes (§K).
Narrative-state proposals were recorded, applied or refused as designed across
the run, and no accepted narration carried an unresolved protocol block (§J).
## I. Context-window / A1 evidence
**Every turn `fits`.** Across all 102 accepted turns the accounting status was
`fits` — **0 `exceeded`, 0 `truncation_suspected`** — and the window was verified
at 16,384 on every one.
**The ten largest stored prompts**, re-counted against what the server itself
reported:
| Turn | App estimate | Server count | Difference | Window | Output reserve | Safety reserve | Observed margin | Status |
| ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | :--- |
| 86 | 14,990 | 15,005 | +15 | 16,384 | 500 | 820 | 879 | fits |
| 83 | 14,978 | 14,993 | +15 | 16,384 | 500 | 820 | 891 | fits |
| 97 | 14,965 | 14,980 | +15 | 16,384 | 500 | 820 | 904 | fits |
| 41 | 14,960 | 14,975 | +15 | 16,384 | 500 | 820 | 909 | fits |
| 66 | 14,951 | 14,966 | +15 | 16,384 | 500 | 820 | 918 | fits |
| 33 | 14,941 | 14,956 | +15 | 16,384 | 500 | 820 | 928 | fits |
| 35 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
| 77 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
| 38 | 14,927 | 14,942 | +15 | 16,384 | 500 | 820 | 942 | fits |
| 45 | 14,922 | 14,937 | +15 | 16,384 | 500 | 820 | 947 | fits |
**The reserve is preserved on every re-counted prompt.** The estimate runs
exactly **15 tokens below** the server's own count on all ten — a constant,
known offset rather than drift — and the smallest observed margin anywhere in the
run is **879 tokens**, against the documented safety reserve of 820.
**Against v1.** M11's closeout recorded a remaining margin of **23–42 tokens**.
The same measurement on this candidate is **879 at its tightest** — roughly
twenty to thirty times the headroom, which is what WP-A1 was for.
## J. Protocol-leak / A2 evidence
**Release-gate result: 0.** Across the run's **105 stored AI actions**,
`protocol_leaks` reports **0 leaking**, `example_ids` empty.
Measured separately, by the detector's own four rules:
| Shape | Count in this run |
| --- | --- |
| State/section heading with an indented entry | **0** |
| Event list (`"events"`) | **0** |
| Event-call syntax (`create_entity(`, `set_scene(`, …) | **0** |
| Hard-limit / continue-hint echo | **0** |
The browser run agrees independently: `protocol_shapes_in_narration` **0** across
its narrated turns (§G).
**The known mid-reply echo did not recur — and is still not fixed.** WP-B.1
recorded one stored reply (action 153, depth 143) where the narrator echoed the
length hint mid-reply and then continued the story, which A2's trailing cleanup
does not remove. Replaying **that stored fixture through this candidate's
detector** still flags it — 1 of 105 AI actions, matched by the hard-limit/hint
rule alone. So the residual is live (§T.2); what this release run shows is that
no equivalent shape occurred in **its** 105 replies. This report does not claim
protocol leakage is solved.
Ordinary fact and state restatement in prose was not counted: it is measurement,
not application-owned protocol, and no rule treats it as a leak.
## K. Memory / WP-B evidence
### K.1 At the final recall point
| | |
| --- | --- |
| **Created** | 19 memories in the bank; 12 summaries |
| **Retained** | the planting turn's era survived to the end of a 102-turn run |
| **Ranked** | 4 memories were selected into the prompt at the recall point |
| **Injected** | the memory section reached the prompt (14,395 tokens that turn) |
| **Independent-memory verdict** | **not demonstrated** — `precondition_failed: absent_from_later_narration` |
### K.2 The independent-fact probe, precondition by precondition
| Precondition | Held? |
| --- | --- |
| The planted turn is outside the history window (planted depth 3, floor 72) | **yes** |
| Absent from authoritative state | **yes** |
| Absent from the summary | **yes** |
| Absent from imported knowledge | **yes** |
| Absent from later narration | **no** — 6 violations, the first at turn 6 |
The narrator restated the planted fact in later narration, so the probe could not
isolate memory as the only path. The run therefore **records no independent
recovery**, and nothing here is relabelled as one.
### K.3 The two statements the contract requires, kept apart
```text
deterministic independent-memory recovery: PASS
reference-model independent-memory limitation: ACCEPTED RESIDUAL
```
- **Deterministic** (`WP-B.2` §I, and the suite on this candidate): the
`independent_full` scenario fails on v1.0.0 at creation and returns
`recovered_through_memory_independent` on the candidate, with isolation
asserted every turn and provenance resolving to the planting turn.
- **Reference model:** failed on the precondition-valid attempt in WP-B.2, and
in this release run the attempt was not precondition-valid at all. The failing
stage remains **memory creation** — the summariser's content selection.
- **No new regression.** B2.1 ranking, B2.2 eviction and B2.3 excerpt creation
all pass deterministically in the 1,723-test suite on this candidate, and the
bank behaved normally through the run (19 memories, ranked and injected). What
this run shows is the known summariser-quality limitation, not a fault in
ranking, eviction or injection.
## L. Identity diagnostic
**Scripted half — complete.** `tools/m11_identity.py --scripted` on the
candidate: 10 identity stresses, **0 signals**, verdict *no objective identity
defect detected*.
**The detector's negative control fires.** With `--inject`, the same harness on
the same candidate raises **7 signals** — `shared_display_name` once and
`duplicate_character_creation` on turns 5–10 — and preserves each turn's
evidence. A clean run therefore means something: the check is capable of
failing.
**Model-backed half — complete, on the candidate, with memory on.**
Evidence: `…/release-87a4032/identity-memory/`.
| | |
| --- | --- |
| Model | **`qwen2.5:3b-instruct-16k`** (the reference narrator) |
| Context window | **16,384** |
| Memory status | **on** — `memoryBankEnabled: true`, embedding model `nomic-embed-text` |
| Summary status | **on** — `autoSummarize: true`; **1 summary written** |
| **Identity signals** | **0** across all 10 stresses |
| **Stored protocol shapes** | **0** of 10 AI actions, by the release-gate detector |
| Fixture | accepted successfully — five entities kept distinct (`bill`, `alice`, `roger`, `john`, `office`) |
| State proposals | recorded and applied with no shared display name and no duplicate creation |
| Scripted detector self-test | still fires (7 signals under `--inject`) |
**A harness correction made during validation, and what it does not invalidate.**
The diagnostic shipped with `embedding_model=""` and the memory bank switched
off, so a release run of it would have reported a clean identity result with
memory never taking part — which is not what the gate asks for. Two lines now
read `AIDND_TEST_EMBED_MODEL` and enable the bank. **Harness-only: no product
code changed**, so no product evidence became stale; the corrected harness
repeated its own check, which is the run reported above.
**One honest observation:** with memory enabled the bank still wrote **0
memories** in this 10-beat campaign — 21 actions is enough to pass
`MEMORY_START`, but the summary pass is what ran and produced the single
summary. Memory was configured and active; it was not meaningfully *exercised*
here. The bank's real exercise is the 102-turn run (§K), which wrote 19.
**No claim about the historical root cause.** The post-M8 identity finding's
campaign was destroyed and its cause cannot be established. A clean run here is
evidence that the product does not do the things it can be blamed for on this
fixture — not a discovery of what happened then.
## M. Recovery
`tools/m11_recovery.py` against **the integrated run's own bundle**
(3,071,683 bytes), imported into a database file that never existed, in a
directory that never existed, by a second server process — so migrations ran
from nothing and this is the fresh-install path as well as the import path.
**16 checks, 16 passed, 0 failed.**
| Claim | Result |
| --- | --- |
| The destination database did not exist beforehand | PASS |
| The bundle imports into a clean directory | PASS — 209 actions in the file, 121 on the active line |
| The active transcript is not empty | PASS |
| Authoritative state came across | PASS — 6 entities, 1 fact |
| The campaign's own canon came across | PASS |
| The narration-length choice came across | PASS |
| Redo availability matches what the file said | PASS |
| Retained (undone) history came across | PASS — **88 actions retained beyond the active line** |
| Both Save Points restore | PASS — *On the ridge*, *Before the ridge* |
| Every imported class came across, with content | PASS — 3 sources |
| The moved campaign accepts a new change, unrefused | PASS |
| The bundle carries no secret | PASS |
| The moved campaign exports again, same story length | PASS |
**One thing this does not prove, stated rather than implied.** The long run
ended with its head at the tip, so the file's head was at the tip and Redo was
correctly unavailable after import. The **undone-head** case (I07) is proved by
§N's upgrade campaign, which ends on an Undo with Redo available and survives
both bundle directions, and by `test_m9_portability.py` in the suite — not by
this bundle.
## N. v1.0.0 upgrade compatibility
A campaign **built and played by the `432f041` application** in its own
worktree, then opened by the candidate. Evidence:
`…/release-87a4032/upgrade/upgrade-report.json`.
**Phase 1 — v1.0.0 builds it.** 20 turns accepted, 39 actions, canon knowledge
imported, a Save Point taken, one Undo so the head is not at the tip, narration
length chosen, memory bank and auto-summary on. It contains what §11 item 9
names: retained history, an undone head with Redo available, a Save Point,
**6 memories**, **2 summaries**, imported knowledge, a narration-length choice
and narrative state. Settings hold a **loopback placeholder**
(`http://127.0.0.1:11434/v1`) — no real hostname is in the evidence database.
**Phase 2 — the candidate opens the same file.** Nothing was copied; the
candidate's migrations ran against it.
| Field | Before | After |
| --- | --- | --- |
| transcript | 39 actions | **identical** |
| newest action (the head) | — | **identical** |
| total / can_undo / can_redo | 39 / true / **true** | **identical** |
| checkpoints | 1 | **identical** |
| narrative state | 7 keys | **identical** |
| memories | **6** | **identical** |
| summaries | **2** | **identical** |
| knowledge sources | 1 | **identical** |
| narration length | `brief` | **identical** |
| memory_bank_enabled / auto_summarize | true / true | **identical** |
| settings | loopback placeholder | **identical** |
| **schema `user_version`** | **94** | **94** |
**All 15 census fields compared identical; 0 differ.** Schema parity is exact:
v1.0.0 and the candidate both stamp `user_version` 94 with 93 migrations, so the
upgrade required no migration at all, and nothing was rewritten in passing. A
fresh-install database from the candidate carries the same 94 (§M's import ran
migrations from nothing).
**A first pass at this gate was discarded.** It played 5 turns, which is below
the memory and summary thresholds, so it compared **0 memories against 0
memories** and proved nothing about two of the criterion's required contents;
its settings also carried the live hostname rather than a placeholder. Both were
corrected and the gate was rerun — the run reported above.
## O. Bundle compatibility
| Direction | Result |
| --- | --- |
| **v1.0.0 export → imported by v1.1** | **PASS** — imported, id 2 |
| **v1.1 export → offered to v1.0.0** | **ACCEPTED** — v1.0.0 imported it, id 1 |
Both directions were **executed**, not inferred from the unchanged format
string. The format is `ai-dnd-adventure-v3` on both sides, and it did not change
during release validation.
Backward import succeeding means the compatibility question the brief raised —
whether optional or additive v1.1 evidence data would break a v1.0.0 importer —
is answered in the negative for this campaign's contents: v1.0.0 accepted the
candidate's bundle whole. No bundle-format change was made or needed.
## P. WP-D regression
Reconfirmed on the candidate; no backup-affecting product code changed after
WP-D, so its 117 MB browser measurement is not repeated (the owner's brief
permits this).
| Claim | Result |
| --- | --- |
| The completed backup copy is verified with `PRAGMA integrity_check` | **PASS** — `app/backup.py:193`, docstring at 198–201 records why the full check replaced `quick_check` |
| The corruption fixture still separates the two pragmas | **PASS** — `quick_check` → `ok`, `integrity_check` → `row 145 missing from index i_t_k` |
| Oversized export still succeeds and is delivered | **PASS** |
| Importability metadata names the effective ceiling | **PASS** — "This export is larger than this version's **20 MB** import limit (… bytes). The file was exported successfully, but this version cannot import it." |
| Normal export unchanged in content | **PASS** |
| Oversized import still refused | **PASS** — 413 naming the limit |
| Suite | **12 passed** (`tests/test_v11_d_recovery.py`) |
## Q. WP-E regression
| Claim | Result |
| --- | --- |
| Contrast audit exit code | **0** |
| All applicable control boundaries ≥ 3:1 | **PASS** — every boundary pair clears 3:1 (1.4.11) |
| Applicable text contrast still compliant | **PASS** — every text pair clears 4.5:1 (1.4.3); baselines 14.57 / 13.57 / 5.48 / 5.88 unchanged |
| Browser boundary checks | **PASS** — the 10 WP-E rows in §G |
| Focus visibility | **PASS** — M11's visible-focus check inside the 38, plus WP-E's focused-edge measurement |
| Gate tests | **11 passed** (`tests/test_v11_e_contrast.py`), including 2.99:1 failing and 3.00:1 passing |
```text
OWNER SCREENSHOT APPROVAL: APPROVED
```
Approved by the owner in the release-validation brief of 2026-09-16. **No visual
code changed during release validation**, so that approval remains valid; had any
changed, it would have been void and new screenshots would have been required.
## R. Release-shaped smoke test
The **final no-cache candidate image**, a fresh volume, published on loopback,
with the private CA installed into the container's own trust store. Evidence:
`…/release-87a4032/smoke/`.
**15 checks, 15 passed, 0 failed.**
| Claim | Result |
| --- | --- |
| The container starts | PASS |
| The application answers on loopback | PASS |
| The port is published on **loopback only** | PASS — `8000/tcp -> 127.0.0.1:…` |
| This machine's **LAN address does not serve** the application | PASS |
| The first page loads | PASS — HTTP 200 |
| The shell references no remote origin | PASS — none found |
| A CSP is served | PASS |
| **The approved HTTPS narrator verifies through its private CA** | PASS — HTTP 200 through `tlstrust.ssl_context()`, **no bypass** |
| **A public endpoint is refused** | PASS — HTTP 400 |
| A campaign is created | PASS |
| **One real narrator turn is accepted** | PASS |
| The container restarts and serves again | PASS |
| The transcript survived the restart | PASS |
| The narrative state survived the restart | PASS |
| **Firefox renders the reopened campaign** | PASS — 470 characters of story |
**A finding worth recording, and it is not a product defect.** The first attempt
failed at `PUT /api/settings` with **HTTP 400**. The cause: a `.local` name is
mDNS, a Docker container has no mDNS resolver, and `endpoints.py` correctly
refuses an endpoint whose address it cannot classify — the policy behaving
exactly as designed. The fix is to resolve the name inside the container
(`--add-host`), **not** to substitute the IP address, because the certificate is
issued for the hostname and substituting the address would have quietly bypassed
the hostname verification this test exists to prove. Harness-only; no product
code changed.
This is supplemental evidence, not a substitute for the gates above.
## S. Security / local-only review
| Claim | Evidence on the candidate |
| --- | --- |
| Served on loopback only | every harness reached the application on `127.0.0.1`; the documented container run publishes loopback |
| Endpoint policy | `app/endpoints.py` admits loopback (v4 and v6), the three RFC1918 ranges, link-local, IPv6 unique-local and CGNAT, and refuses the public Internet; a name resolving to both a private and a public address is refused |
| Inference actually used | trusted-LAN **HTTPS** with a private CA for the browser gate (§G); plain HTTP to a LAN GPU host for the long run, which `SECURITY-THREAT-MODEL.md` §83 permits and which is not A06 evidence |
| No secret in exports | offline gate, WP-D tests and the M11 suite |
| CSP served | `default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none'; base-uri 'none'; form-action 'self'; frame-ancestors 'none'` |
| Other headers | `x-content-type-options: nosniff`, `referrer-policy: same-origin`, `x-frame-options: DENY` |
| No remote origin in the shell | offline gate: the page names none, and every asset is local |
`'unsafe-inline'` remains on `style-src` only, because React writes inline
`style` attributes; it is deliberately absent from `script-src`.
## T. Known residual risks
Each is classified, and none is collapsed into another category.
### T.1 WP-B reference-model memory limitation — ACCEPTED RESIDUAL
```text
deterministic independent memory: PASS
reference-model independent memory: FAIL
failing stage: memory creation — the summariser's content selection
owner decision: accepted for v1.1
```
**Status on this candidate:** unchanged, and no broader regression. The release
long run did not demonstrate independent recovery, but it also could not: its
probe failed the `absent_from_later_narration` precondition because the narrator
restated the fact (§K.2). Ranking, eviction and excerpt creation all pass
deterministically in the 1,723-test suite, and the bank worked normally through
102 turns (19 memories, ranked and injected). **Accepted residual**, carried
visibly, not a blocker.
### T.2 A2 mid-reply application-instruction echo — ACCEPTED RESIDUAL, still live
The known occurrence is WP-B.1's action 153 (depth 143): the narrator echoed the
length hint **mid-reply** and then continued the story, which A2's trailing
cleanup does not remove.
**Reproduced on this candidate.** Replaying that stored fixture through the
release-gate detector still flags it — **1 of 105** AI actions, matched by the
hard-limit/hint rule alone. The extractor still leaves it. This report therefore
does **not** claim protocol leakage is solved.
**But it did not recur in release evidence.** This run's own 105 stored replies
leak **0** (§J), and the browser run's narrated turns leak 0. The brief's
stop-and-report condition — *an equivalent shape occurring in this final run* —
was **not** triggered, so validation continues. No broader sanitizer was written;
that remains for owner review.
### T.3 Doubled full stop in memory-search scene text — ACCEPTED RESIDUAL
`"rain outside.."` when the state's scene summary already ends in punctuation.
Only the embedding query sees it; the effect is one stray token. Not changed
during release validation, deliberately: cleanliness is not a reason to alter
product behaviour after evidence is taken.
### T.4 K1 — "Correct" on an Important Facts row is always refused
**Reproduced, unchanged on this candidate**, deterministically and without a
browser: Correct on a Characters row (`subject='mara'`) applies, **201**; Correct
on an Important Facts row (`subject='f1'`) is refused, **400 — "add_fact names
subject='f1', which does not exist."** The cause is frontend-side: the panel
sends the row key as `add_fact.subject`, and the validator checks `subject` as an
entity reference.
**Classification: v1.2 backlog, not a release blocker.** It blocks no v1
REQUIRED test — C04 passes through the working correction paths (§C) — and WP-C's
State-panel correction coverage passes. It is a narrow bug with an obvious fix
(offer Correct only against entities, or send facts without a subject), but
fixing it during release validation would change product code after the evidence
above was taken, which the brief forbids without a product brief. **Left for the
owner.**
### T.5 Import ceiling and scheduled backups — INTENTIONAL, not unfinished work
The 20 MB import ceiling is deliberate and unchanged; WP-D made it honest rather
than raising it. Scheduled backups remain unbuilt by design. Neither is a
residual defect.
### T.6 Harness corrections made during validation — no product evidence invalidated
Three, all harness-only, each named where it occurred: the identity diagnostic
did not enable memory (§L); the smoke test needed hostname resolution inside the
container (§R); and the first upgrade campaign was too short to write memories
and carried a live hostname (§N). **No product code changed at any point during
release validation**, so no black-box or long-run evidence became stale. Each
corrected harness repeated its own affected check.
## U. Deferred v1.2 / future work
| Item | Why it is deferred |
| --- | --- |
| Raising the import ceiling, or a streaming import | The plan assigns it to v1.2; WP-D's scope was honesty about the limit, not the limit |
| Scheduled backups; a restore button | Explicitly out of WP-D's scope |
| K1's Correct-on-a-fact-row fix (§T.4) | A narrow frontend bug needing a product brief |
| A broader mid-reply protocol sanitizer (§T.2) | Needs owner review; A2's cleanup is deliberately trailing-only |
| The doubled full stop (§T.3) | Cosmetic, embedding-query only |
| K05 generate local image, K06 multi-turn video | FUTURE tests; both need a media provider this release does not build |
| Reference-model independent memory (§T.1) | Needs a stronger summariser or a different creation strategy — a v1.2 investigation, not a v1.1 fix |
## V. Documentation changes
Current documents were brought up to date. **No historical milestone report was
rewritten, and no failed WP-B real-model evidence was turned into success.**
| Document | Change |
| --- | --- |
| `README.md` | Status now says v1.0.0 **remains** the released version, that all six v1.1 packages are complete and accepted, that release validation passed on candidate `87a4032`, that WP-B ships with a documented limitation, and that **no `v1.1.0` tag exists and `main` is unchanged`**. Also corrected a stale figure: the schema is versioned at **94**, not "92 and counting" |
| `planning/V1.1-PLAN.md` | WP-D/WP-E recorded as signed `87a4032`; the release-validation outcome summarised with its residuals; the three owner events named as still outstanding |
| `planning/VERSION.md` | Same status correction, plus a new revision entry for the closeout |
| `planning/README.md` | Current-state paragraph rewritten for the same facts |
| `planning/reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL: PENDING` → **`APPROVED`**, with the source (the release-validation brief) and date recorded, and a note that the signed commit predated the review |
| `planning/reports/v1.1/V1.1-RELEASE-REPORT.md` | **New** — this document |
| `DEVELOPMENT.md` | Unchanged: its WP-C/WP-D sections already describe the candidate as built |
**New tools committed with this closeout** (harness only, no product code):
`backend/tools/v11_upgrade_check.py` (Gate 9) and
`backend/tools/v11_release_smoke.py` (§R), plus two narrow corrections to
`backend/tools/m11_identity.py` (read `AIDND_TEST_EMBED_MODEL`; enable the memory
bank) so the diagnostic can run with memory on.
## W. Final release decision
**The question this validation set out to answer:** does candidate `87a4032`
preserve the complete v1 contract and satisfy every accepted v1.1 package on one
integrated release tree?
| Gate | Result |
| --- | --- |
| 1 Package acceptance | **PASS** — six packages accepted; WP-B's qualification carried whole (§B.1) |
| 2 v1 acceptance contract | **PASS** — 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified (§C) |
| 3 Suites, lint, build, Docker | **PASS** — 1,723 / 175 / 0 errors; image SPA file-for-file identical (§D, §E) |
| 4 Offline / no-network | **PASS** — 23/23 on the candidate image (§F) |
| 5 Browser release run | **PASS** — 101/0/0 over trusted-LAN HTTPS, every turn `fits` (§G) |
| 6 Integrated long run | **PASS** — 102 turns, M01–M04, 0 post-turn failures (§H) |
| 7 Identity diagnostic | **PASS** — 0 signals, 0 protocol shapes, memory on (§L) |
| 8 Recovery | **PASS** — 16/16 on the run's own bundle (§M) |
| 9 v1.0.0 upgrade | **PASS** — 15/15 identical, schema parity at 94 (§N) |
| 10 WP-D regression | **PASS** (§P) |
| 11 WP-E regression | **PASS**, screenshots approved (§Q) |
| Bundle compatibility | **Both directions execute and import** (§O) |
| Release smoke | **PASS** — 15/15 from the shipped image (§R) |
**A1** holds its reserve on every re-counted prompt, with a smallest margin of
**879 tokens** against v1's 23–42. **A2**'s release-gate leak count is **0**.
**WP-B**'s deterministic independent-memory recovery passes and its
reference-model limitation remains an accepted, documented residual — stated in
§K.3 in both halves, never shortened to "WP-B passed".
No product code was changed at any point during release validation, so no
evidence was invalidated. Three harness corrections were made and each corrected
harness repeated its own check (§T.6).
```text
V1.1 RELEASE VALIDATION:
PASS
```
### What this decision is not
These are separate, and only the first is done:
```text
WP-A-E accepted: YES
release validation passed: YES
release candidate prepared: YES (87a4032, with this report)
release commit signed: NO
main updated to v1.1: NO
v1.1.0 tagged: NO
```
The closeout changes are **staged and uncommitted**. Nothing was committed,
pushed, merged or tagged by this validation. The owner's next decision is to
review this evidence, resolve anything they disagree with, then sign the v1.1
release commit and publish `v1.1.0`.
+10 -4
View File
@@ -365,10 +365,16 @@ All WP-E changes are **staged and uncommitted**. No commit, no push, no tag.
v1.1 release validation has not begun. v1.1 release validation has not begun.
```text ```text
OWNER SCREENSHOT APPROVAL: PENDING OWNER SCREENSHOT APPROVAL: APPROVED
``` ```
The before/after screenshots in §I are the evidence for a change a reader judges The before/after screenshots in §I are the evidence for a change a reader judges
by looking at it. The measurements say every boundary now clears 3:1; whether by looking at it. The measurements say every boundary now clears 3:1; whether the
the result looks right in this design is the owner's call, and it is recorded as result looks right in this design was the owner's call.
pending rather than assumed.
**Approved by the owner on 2026-09-16**, in the v1.1 release-validation brief,
after reviewing the before/after pair in `$HOME/v11-evidence/wp-e/`. This line
was `PENDING` in the signed commit `87a4032` because the report predated that
review; it is updated here as part of the release closeout, with its source and
date recorded rather than the approval being assumed. No visual code changed
during release validation, so the approval stands (release report §Q).