Files
JesseMarkowitzandClaude Opus 5 db7b309e3d
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
v1.1 closeout: accept integrated release validation
Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.

V1.1 RELEASE VALIDATION: PASS

What was run, on this candidate:

- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
  0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
  175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
  to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
  trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
  3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
  keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
  scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
  that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
  identical on all 15 census fields, schema parity at user_version 94, and both
  bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
  verified, public endpoint refused, a real turn, restart, persistence, and
  Firefox rendering the reopened campaign.

Carried residuals, stated rather than summarised away:

- WP-B: deterministic independent-memory recovery PASS; reference-model
  independent-memory recovery FAIL at memory creation — the owner-accepted
  limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
  reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
  release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
  backlog, reproduced and not fixed during validation.

Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.

Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.

Still the owner's to do: sign the release commit, update main, tag v1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-16 07:12:23 -04:00

382 lines
17 KiB
Python

"""v1.1 release Gate 9: a real v1.0.0 campaign, opened by the candidate.
python -m tools.v11_upgrade_check --v100 <worktree> --out <dir under $HOME>
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT`, `AIDND_TEST_MODEL` and
`AIDND_TEST_EMBED_MODEL`: the campaign has to be *played*, because memories,
summaries and narrative state are things a narrator produces. A schema-only
fixture would prove nothing about an upgrade, which is why §11 item 9 asks for a
database the v1.0.0 application built.
Three phases, each its own server process, so everything that survives crosses
as bytes on disk:
1. **v1.0.0 builds and plays.** The `432f041` tree serves the application: a
campaign is created, canon knowledge imported, turns played, a Save Point
taken, a turn undone so the head is not at the tip and Redo is available, and
a narration length chosen. Then that server stops, and a census is taken.
2. **The candidate opens the same file.** Nothing is copied; the candidate's own
migrations run against it. The census is taken again and compared field by
field.
3. **Bundles cross both ways.** The v1.0.0 export is imported by the candidate.
The candidate's export is offered back to v1.0.0, and whatever happens is
reported — the format string is unchanged, which is a reason to test backward
import, not a reason to assume it.
Settings hold only what the caller's environment names, and the evidence
directory lives under `$HOME`.
## Shapes this had to be written against, not guessed
- A turn is **SSE**: `POST /adventures/{id}/actions`, and a failed turn is an
`error` *event* inside an HTTP 200. Reading the status code would call every
failure a success.
- History is `GET /{id}/actions?limit=N` -> `{actions, total, has_more,
can_undo, can_redo}`. There is **no head-id field**, so the active head is
compared as the newest action plus the two flags.
- Save Points are **checkpoints**. Creating one after an Undo names the undone
position, deliberately.
- There is **no summaries route**; summaries and post-turn health both come from
`GET /{id}/derived`.
- Campaign switches are `PATCH /adventures/{id}` with `memory_bank_enabled` and
`auto_summarize` — not the names a reader would guess.
- Knowledge import is **multipart**, as `m11_long_run` and `m11_offline` build it.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sqlite3
import subprocess
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.m11_webdriver import free_port, require_under_home # noqa: E402
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
TURN_TIMEOUT = 900
#: What the evidence database is left holding. The campaign has to be played
#: against a real narrator, but nothing about the *upgrade* depends on which
#: host that was, and §11 item 9 asks for no real hostname in this database.
PLACEHOLDER_ENDPOINT = "http://127.0.0.1:11434/v1"
#: What the upgrade must preserve, compared exactly on both sides. `settings` is
#: included because a migration that silently rewrote an endpoint would be a real
#: defect; `schema_version` is read from the file rather than the API.
CENSUS = ("transcript", "newest_action", "total", "can_undo", "can_redo",
"checkpoints", "state", "memories", "summaries", "knowledge",
"narration_length", "memory_bank_enabled", "auto_summarize",
"settings", "schema_version")
CANON_MD = """# Westhaven
The abbey bell is rung only for a death. Mara keeps the harbour ledger.
Aldric carries a silver key he will not explain.
"""
class App:
"""One application process, from whichever tree it is given."""
def __init__(self, tree: Path, db: Path, log: Path, label: str):
self.tree, self.db, self.label = tree, db, label
self.port = free_port()
self.log = log
handle = open(log, "ab")
self.proc = subprocess.Popen(
[str(Path(__file__).resolve().parent.parent / ".venv/bin/uvicorn"),
"app.main:app", "--host", "127.0.0.1", "--port", str(self.port)],
cwd=str(tree / "backend"), stdout=handle, stderr=subprocess.STDOUT,
env={**os.environ, "AIDND_DB_PATH": str(db),
"AIDND_DATABASE_URL": "", "DATABASE_URL": ""},
)
self.url = f"http://127.0.0.1:{self.port}"
deadline = time.monotonic() + 120
while time.monotonic() < deadline:
if self.proc.poll() is not None:
raise RuntimeError(f"{label} exited early; see {log}")
try:
urllib.request.urlopen(self.url + "/api/settings", timeout=2)
print(f" {label} serving {db.name} on {self.url}", flush=True)
return
except Exception:
time.sleep(0.2)
raise RuntimeError(f"{label} never became ready; see {log}")
def call(self, method: str, path: str, payload=None, timeout=120):
data = json.dumps(payload).encode() if payload is not None else None
request = urllib.request.Request(
f"{self.url}/api{path}", data=data, method=method,
headers={"Content-Type": "application/json"} if data else {})
with urllib.request.urlopen(request, timeout=timeout) as response:
body = response.read().decode()
return json.loads(body) if body else None
def stream(self, path: str, payload) -> list[dict]:
"""A turn. A failed turn is an event in the stream, not a status code."""
request = urllib.request.Request(
f"{self.url}/api{path}", data=json.dumps(payload).encode(),
method="POST", headers={"Content-Type": "application/json"})
events: list[dict] = []
with urllib.request.urlopen(request, timeout=TURN_TIMEOUT) as response:
for raw in response:
line = raw.decode(errors="replace").strip()
if line.startswith("data:"):
try:
events.append(json.loads(line[5:].strip()))
except json.JSONDecodeError:
pass
return events
def upload(self, adv: int, name: str, body: str, classification: str) -> dict:
boundary = "----v11upgrade"
parts = (
f"--{boundary}\r\nContent-Disposition: form-data; name=\"classification\""
f"\r\n\r\n{classification}\r\n"
f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; "
f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n"
f"--{boundary}--\r\n"
).encode()
request = urllib.request.Request(
f"{self.url}/api/adventures/{adv}/knowledge", data=parts, method="POST",
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"})
with urllib.request.urlopen(request, timeout=120) as response:
return json.loads(response.read().decode())
def stop(self) -> None:
if self.proc.poll() is None:
self.proc.terminate()
try:
self.proc.wait(timeout=30)
except subprocess.TimeoutExpired:
self.proc.kill()
def schema_version(db: Path) -> int:
connection = sqlite3.connect(f"file:{db}?mode=ro", uri=True)
try:
return connection.execute("PRAGMA user_version").fetchone()[0]
finally:
connection.close()
def census(app: App, adv: int, db: Path) -> dict:
page = app.call("GET", f"/adventures/{adv}/actions?limit=500") or {}
actions = page.get("actions") or []
derived = app.call("GET", f"/adventures/{adv}/derived") or {}
adventure = app.call("GET", f"/adventures/{adv}") or {}
settings = app.call("GET", "/settings") or {}
checkpoints = app.call("GET", f"/adventures/{adv}/checkpoints") or []
memories = app.call("GET", f"/adventures/{adv}/memories") or []
state = app.call("GET", f"/adventures/{adv}/state") or {}
knowledge = app.call("GET", f"/adventures/{adv}/knowledge") or []
newest = actions[-1] if actions else {}
return {
"transcript": [(a.get("type"), (a.get("text") or "")[:300]) for a in actions],
# No head id is exposed; the head is the newest action on the read line
# plus the two flags the page carries.
"newest_action": ((newest.get("type"), (newest.get("text") or "")[:300])
if newest else None),
"total": page.get("total"),
"can_undo": page.get("can_undo"),
"can_redo": page.get("can_redo"),
"checkpoints": sorted((c.get("name"), c.get("depth"), c.get("branch_id"),
c.get("on_path"), c.get("resolved"))
for c in checkpoints),
"state": state.get("document") if isinstance(state, dict) else state,
"memories": sorted((m.get("text") or "")[:200] for m in memories),
"summaries": sorted((s.get("text") or "")[:200]
for s in (derived.get("summaries") or [])),
"knowledge": sorted((k.get("filename") or k.get("title"),
k.get("classification")) for k in knowledge),
"narration_length": adventure.get("narration_length"),
"memory_bank_enabled": adventure.get("memory_bank_enabled"),
"auto_summarize": adventure.get("auto_summarize"),
"settings": {k: settings.get(k)
for k in ("endpoint_url", "model", "max_output_tokens")},
"schema_version": schema_version(db),
}
def play(app: App, adv: int, text: str) -> bool:
events = app.stream(f"/adventures/{adv}/actions", {"type": "do", "text": text})
errors = [e for e in events if e.get("type") == "error"]
if errors:
print(f" turn refused: {errors[0].get('detail', '')[:150]}", flush=True)
return False
return True
def build_v100_campaign(app: App) -> int:
created = app.call("POST", "/adventures", {
"title": "Upgrade Evidence",
"opening": "Rain over Westhaven, and the abbey bell tolling.",
"canon_rules": ["The dead do not return."],
"persona_name": "Aldric",
})
adv = created["id"]
app.call("PUT", "/settings", {
"endpoint_url": ENDPOINT, "model": MODEL,
"embedding_model": EMBED_MODEL,
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
# Memory and summaries are per-campaign switches defaulting to off, so the
# census would otherwise have nothing to compare.
app.call("PATCH", f"/adventures/{adv}", {
"narration_length": "brief", "memory_bank_enabled": True,
"auto_summarize": True})
app.upload(adv, "westhaven-canon.md", CANON_MD, "canon")
# Enough turns that memories and summaries actually exist. A memory needs
# MEMORY_INTERVAL (6) actions plus SETTLE_SLACK (1) settled past the anchor,
# and a summary needs SUMMARY_INTERVAL (15) uncovered actions — both counted
# in *actions*, and a turn writes two. A first pass at this gate played five
# turns, wrote neither, and compared 0 against 0, which proves nothing about
# whether the upgrade preserves them.
beats = [
"I ask Mara what the bell means.",
"I show her the silver key.",
"I follow her to the harbour ledger.",
"I ask who else knows about the key.",
"I read the ledger's last page aloud.",
"I ask the ferryman about the fen road.",
"I wait out the rain and watch the harbour.",
"I ask Mara about the abbey's sealed crypt.",
"I count the entries against the tide table.",
"I ask who signed for the last shipment.",
"I walk the quay to the chandler's door.",
"I ask the chandler what he remembers of that night.",
"I show the chandler the key.",
"I return to Mara with what he said.",
"I ask Mara what she means to do now.",
"I agree to meet her at first light.",
"I take the long way back along the ridge.",
"I check whether anyone followed me.",
"I write down what I have learned so far.",
"I sleep, and wake before the bell.",
]
played = 0
for text in beats:
if play(app, adv, text):
played += 1
print(f" {played} turns accepted by v1.0.0", flush=True)
app.call("POST", f"/adventures/{adv}/checkpoints", {"name": "before the ledger"})
# One Undo, so the head is not at the retained tip and Redo is available.
app.call("POST", f"/adventures/{adv}/undo", {})
# §11 item 9: the database must carry **loopback or placeholder settings with
# no real hostnames**. The turns above needed a real narrator, so the
# endpoint is reset to loopback once the story exists — before the census is
# taken and before either bundle is exported.
app.call("PUT", "/settings", {
"endpoint_url": PLACEHOLDER_ENDPOINT, "model": MODEL,
"embedding_model": EMBED_MODEL,
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
return adv
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--v100", required=True, help="the v1.0.0 worktree")
parser.add_argument("--out", required=True)
args = parser.parse_args()
if not (ENDPOINT and MODEL and EMBED_MODEL):
print("set AIDND_TEST_ENDPOINT, AIDND_TEST_MODEL and AIDND_TEST_EMBED_MODEL")
return 2
out = require_under_home(Path(args.out).expanduser())
shutil.rmtree(out, ignore_errors=True)
out.mkdir(parents=True)
v100_tree = Path(args.v100).expanduser().resolve()
candidate_tree = Path(__file__).resolve().parent.parent.parent
db = out / "campaign.db"
results: dict = {"started": datetime.now().isoformat(timespec="seconds"),
"v100_tree": str(v100_tree), "candidate": str(candidate_tree)}
failures: list[str] = []
print("phase 1 — v1.0.0 builds and plays the campaign")
app = App(v100_tree, db, out / "v100-server.log", "v1.0.0")
try:
adv = build_v100_campaign(app)
before = census(app, adv, db)
v100_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
(out / "v100-export.json").write_text(json.dumps(v100_bundle))
finally:
app.stop()
results["adventure"], results["before"] = adv, before
print(f" {before['total']} actions, redo={before['can_redo']}, "
f"checkpoints={len(before['checkpoints'])}, memories={len(before['memories'])}, "
f"summaries={len(before['summaries'])}, knowledge={len(before['knowledge'])}, "
f"schema={before['schema_version']}")
print("\nphase 2 — the candidate opens that same database file")
app = App(candidate_tree, db, out / "candidate-server.log", "candidate")
try:
after = census(app, adv, db)
v11_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
(out / "v11-export.json").write_text(json.dumps(v11_bundle))
try:
imported = app.call("POST", "/adventures/import", v100_bundle, timeout=600)
results["v100_bundle_into_v11"] = {"status": "imported",
"id": (imported or {}).get("id")}
print(" the v1.0.0 bundle imported into the candidate")
except urllib.error.HTTPError as exc:
results["v100_bundle_into_v11"] = {
"status": "refused", "code": exc.code,
"detail": exc.read().decode()[:300]}
failures.append("v100_bundle_into_v11")
print(f" the candidate REFUSED the v1.0.0 bundle: {exc.code}")
finally:
app.stop()
results["after"] = after
print(" comparing the census, field by field:")
for field in CENSUS:
same = before.get(field) == after.get(field)
print(f" {'ok ' if same else 'DIFF'} {field}")
if not same:
failures.append(field)
results.setdefault("differences", {})[field] = {
"before": before.get(field), "after": after.get(field)}
print("\nphase 3 — the candidate's bundle offered back to v1.0.0")
app = App(v100_tree, out / "backward.db", out / "v100-backward.log", "v1.0.0")
try:
try:
back = app.call("POST", "/adventures/import", v11_bundle, timeout=600)
results["v11_bundle_into_v100"] = {"status": "imported",
"id": (back or {}).get("id")}
print(" v1.0.0 ACCEPTED the v1.1 bundle")
except urllib.error.HTTPError as exc:
detail = exc.read().decode()[:400]
results["v11_bundle_into_v100"] = {"status": "refused", "code": exc.code,
"detail": detail}
# Reported, not failed: §11 item 9 asks for the result, and the
# owner's brief asks whether a refusal breaks the compatibility
# promise — a judgement, not an assertion this script may make.
print(f" v1.0.0 REFUSED the v1.1 bundle: {exc.code} {detail[:160]}")
finally:
app.stop()
results["failures"] = failures
(out / "upgrade-report.json").write_text(json.dumps(results, indent=2, default=str))
print(f"\n{'PASS' if not failures else 'FAIL'}: {len(failures)} field(s) differ "
f"-> {out / 'upgrade-report.json'}")
return 1 if failures else 0
if __name__ == "__main__":
raise SystemExit(main())