Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.
V1.1 RELEASE VALIDATION: PASS
What was run, on this candidate:
- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
identical on all 15 census fields, schema parity at user_version 94, and both
bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
verified, public endpoint refused, a real turn, restart, persistence, and
Firefox rendering the reopened campaign.
Carried residuals, stated rather than summarised away:
- WP-B: deterministic independent-memory recovery PASS; reference-model
independent-memory recovery FAIL at memory creation — the owner-accepted
limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
backlog, reproduced and not fixed during validation.
Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.
Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.
Still the owner's to do: sign the release commit, update main, tag v1.1.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
382 lines
17 KiB
Python
382 lines
17 KiB
Python
"""v1.1 release Gate 9: a real v1.0.0 campaign, opened by the candidate.
|
|
|
|
python -m tools.v11_upgrade_check --v100 <worktree> --out <dir under $HOME>
|
|
|
|
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT`, `AIDND_TEST_MODEL` and
|
|
`AIDND_TEST_EMBED_MODEL`: the campaign has to be *played*, because memories,
|
|
summaries and narrative state are things a narrator produces. A schema-only
|
|
fixture would prove nothing about an upgrade, which is why §11 item 9 asks for a
|
|
database the v1.0.0 application built.
|
|
|
|
Three phases, each its own server process, so everything that survives crosses
|
|
as bytes on disk:
|
|
|
|
1. **v1.0.0 builds and plays.** The `432f041` tree serves the application: a
|
|
campaign is created, canon knowledge imported, turns played, a Save Point
|
|
taken, a turn undone so the head is not at the tip and Redo is available, and
|
|
a narration length chosen. Then that server stops, and a census is taken.
|
|
2. **The candidate opens the same file.** Nothing is copied; the candidate's own
|
|
migrations run against it. The census is taken again and compared field by
|
|
field.
|
|
3. **Bundles cross both ways.** The v1.0.0 export is imported by the candidate.
|
|
The candidate's export is offered back to v1.0.0, and whatever happens is
|
|
reported — the format string is unchanged, which is a reason to test backward
|
|
import, not a reason to assume it.
|
|
|
|
Settings hold only what the caller's environment names, and the evidence
|
|
directory lives under `$HOME`.
|
|
|
|
## Shapes this had to be written against, not guessed
|
|
|
|
- A turn is **SSE**: `POST /adventures/{id}/actions`, and a failed turn is an
|
|
`error` *event* inside an HTTP 200. Reading the status code would call every
|
|
failure a success.
|
|
- History is `GET /{id}/actions?limit=N` -> `{actions, total, has_more,
|
|
can_undo, can_redo}`. There is **no head-id field**, so the active head is
|
|
compared as the newest action plus the two flags.
|
|
- Save Points are **checkpoints**. Creating one after an Undo names the undone
|
|
position, deliberately.
|
|
- There is **no summaries route**; summaries and post-turn health both come from
|
|
`GET /{id}/derived`.
|
|
- Campaign switches are `PATCH /adventures/{id}` with `memory_bank_enabled` and
|
|
`auto_summarize` — not the names a reader would guess.
|
|
- Knowledge import is **multipart**, as `m11_long_run` and `m11_offline` build it.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import json
|
|
import os
|
|
import shutil
|
|
import sqlite3
|
|
import subprocess
|
|
import sys
|
|
import time
|
|
import urllib.error
|
|
import urllib.request
|
|
from datetime import datetime
|
|
from pathlib import Path
|
|
|
|
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
|
|
from tools.m11_webdriver import free_port, require_under_home # noqa: E402
|
|
|
|
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
|
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
|
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
|
|
TURN_TIMEOUT = 900
|
|
#: What the evidence database is left holding. The campaign has to be played
|
|
#: against a real narrator, but nothing about the *upgrade* depends on which
|
|
#: host that was, and §11 item 9 asks for no real hostname in this database.
|
|
PLACEHOLDER_ENDPOINT = "http://127.0.0.1:11434/v1"
|
|
|
|
#: What the upgrade must preserve, compared exactly on both sides. `settings` is
|
|
#: included because a migration that silently rewrote an endpoint would be a real
|
|
#: defect; `schema_version` is read from the file rather than the API.
|
|
CENSUS = ("transcript", "newest_action", "total", "can_undo", "can_redo",
|
|
"checkpoints", "state", "memories", "summaries", "knowledge",
|
|
"narration_length", "memory_bank_enabled", "auto_summarize",
|
|
"settings", "schema_version")
|
|
|
|
CANON_MD = """# Westhaven
|
|
|
|
The abbey bell is rung only for a death. Mara keeps the harbour ledger.
|
|
Aldric carries a silver key he will not explain.
|
|
"""
|
|
|
|
|
|
class App:
|
|
"""One application process, from whichever tree it is given."""
|
|
|
|
def __init__(self, tree: Path, db: Path, log: Path, label: str):
|
|
self.tree, self.db, self.label = tree, db, label
|
|
self.port = free_port()
|
|
self.log = log
|
|
handle = open(log, "ab")
|
|
self.proc = subprocess.Popen(
|
|
[str(Path(__file__).resolve().parent.parent / ".venv/bin/uvicorn"),
|
|
"app.main:app", "--host", "127.0.0.1", "--port", str(self.port)],
|
|
cwd=str(tree / "backend"), stdout=handle, stderr=subprocess.STDOUT,
|
|
env={**os.environ, "AIDND_DB_PATH": str(db),
|
|
"AIDND_DATABASE_URL": "", "DATABASE_URL": ""},
|
|
)
|
|
self.url = f"http://127.0.0.1:{self.port}"
|
|
deadline = time.monotonic() + 120
|
|
while time.monotonic() < deadline:
|
|
if self.proc.poll() is not None:
|
|
raise RuntimeError(f"{label} exited early; see {log}")
|
|
try:
|
|
urllib.request.urlopen(self.url + "/api/settings", timeout=2)
|
|
print(f" {label} serving {db.name} on {self.url}", flush=True)
|
|
return
|
|
except Exception:
|
|
time.sleep(0.2)
|
|
raise RuntimeError(f"{label} never became ready; see {log}")
|
|
|
|
def call(self, method: str, path: str, payload=None, timeout=120):
|
|
data = json.dumps(payload).encode() if payload is not None else None
|
|
request = urllib.request.Request(
|
|
f"{self.url}/api{path}", data=data, method=method,
|
|
headers={"Content-Type": "application/json"} if data else {})
|
|
with urllib.request.urlopen(request, timeout=timeout) as response:
|
|
body = response.read().decode()
|
|
return json.loads(body) if body else None
|
|
|
|
def stream(self, path: str, payload) -> list[dict]:
|
|
"""A turn. A failed turn is an event in the stream, not a status code."""
|
|
request = urllib.request.Request(
|
|
f"{self.url}/api{path}", data=json.dumps(payload).encode(),
|
|
method="POST", headers={"Content-Type": "application/json"})
|
|
events: list[dict] = []
|
|
with urllib.request.urlopen(request, timeout=TURN_TIMEOUT) as response:
|
|
for raw in response:
|
|
line = raw.decode(errors="replace").strip()
|
|
if line.startswith("data:"):
|
|
try:
|
|
events.append(json.loads(line[5:].strip()))
|
|
except json.JSONDecodeError:
|
|
pass
|
|
return events
|
|
|
|
def upload(self, adv: int, name: str, body: str, classification: str) -> dict:
|
|
boundary = "----v11upgrade"
|
|
parts = (
|
|
f"--{boundary}\r\nContent-Disposition: form-data; name=\"classification\""
|
|
f"\r\n\r\n{classification}\r\n"
|
|
f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; "
|
|
f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n"
|
|
f"--{boundary}--\r\n"
|
|
).encode()
|
|
request = urllib.request.Request(
|
|
f"{self.url}/api/adventures/{adv}/knowledge", data=parts, method="POST",
|
|
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"})
|
|
with urllib.request.urlopen(request, timeout=120) as response:
|
|
return json.loads(response.read().decode())
|
|
|
|
def stop(self) -> None:
|
|
if self.proc.poll() is None:
|
|
self.proc.terminate()
|
|
try:
|
|
self.proc.wait(timeout=30)
|
|
except subprocess.TimeoutExpired:
|
|
self.proc.kill()
|
|
|
|
|
|
def schema_version(db: Path) -> int:
|
|
connection = sqlite3.connect(f"file:{db}?mode=ro", uri=True)
|
|
try:
|
|
return connection.execute("PRAGMA user_version").fetchone()[0]
|
|
finally:
|
|
connection.close()
|
|
|
|
|
|
def census(app: App, adv: int, db: Path) -> dict:
|
|
page = app.call("GET", f"/adventures/{adv}/actions?limit=500") or {}
|
|
actions = page.get("actions") or []
|
|
derived = app.call("GET", f"/adventures/{adv}/derived") or {}
|
|
adventure = app.call("GET", f"/adventures/{adv}") or {}
|
|
settings = app.call("GET", "/settings") or {}
|
|
checkpoints = app.call("GET", f"/adventures/{adv}/checkpoints") or []
|
|
memories = app.call("GET", f"/adventures/{adv}/memories") or []
|
|
state = app.call("GET", f"/adventures/{adv}/state") or {}
|
|
knowledge = app.call("GET", f"/adventures/{adv}/knowledge") or []
|
|
newest = actions[-1] if actions else {}
|
|
return {
|
|
"transcript": [(a.get("type"), (a.get("text") or "")[:300]) for a in actions],
|
|
# No head id is exposed; the head is the newest action on the read line
|
|
# plus the two flags the page carries.
|
|
"newest_action": ((newest.get("type"), (newest.get("text") or "")[:300])
|
|
if newest else None),
|
|
"total": page.get("total"),
|
|
"can_undo": page.get("can_undo"),
|
|
"can_redo": page.get("can_redo"),
|
|
"checkpoints": sorted((c.get("name"), c.get("depth"), c.get("branch_id"),
|
|
c.get("on_path"), c.get("resolved"))
|
|
for c in checkpoints),
|
|
"state": state.get("document") if isinstance(state, dict) else state,
|
|
"memories": sorted((m.get("text") or "")[:200] for m in memories),
|
|
"summaries": sorted((s.get("text") or "")[:200]
|
|
for s in (derived.get("summaries") or [])),
|
|
"knowledge": sorted((k.get("filename") or k.get("title"),
|
|
k.get("classification")) for k in knowledge),
|
|
"narration_length": adventure.get("narration_length"),
|
|
"memory_bank_enabled": adventure.get("memory_bank_enabled"),
|
|
"auto_summarize": adventure.get("auto_summarize"),
|
|
"settings": {k: settings.get(k)
|
|
for k in ("endpoint_url", "model", "max_output_tokens")},
|
|
"schema_version": schema_version(db),
|
|
}
|
|
|
|
|
|
def play(app: App, adv: int, text: str) -> bool:
|
|
events = app.stream(f"/adventures/{adv}/actions", {"type": "do", "text": text})
|
|
errors = [e for e in events if e.get("type") == "error"]
|
|
if errors:
|
|
print(f" turn refused: {errors[0].get('detail', '')[:150]}", flush=True)
|
|
return False
|
|
return True
|
|
|
|
|
|
def build_v100_campaign(app: App) -> int:
|
|
created = app.call("POST", "/adventures", {
|
|
"title": "Upgrade Evidence",
|
|
"opening": "Rain over Westhaven, and the abbey bell tolling.",
|
|
"canon_rules": ["The dead do not return."],
|
|
"persona_name": "Aldric",
|
|
})
|
|
adv = created["id"]
|
|
app.call("PUT", "/settings", {
|
|
"endpoint_url": ENDPOINT, "model": MODEL,
|
|
"embedding_model": EMBED_MODEL,
|
|
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
|
|
# Memory and summaries are per-campaign switches defaulting to off, so the
|
|
# census would otherwise have nothing to compare.
|
|
app.call("PATCH", f"/adventures/{adv}", {
|
|
"narration_length": "brief", "memory_bank_enabled": True,
|
|
"auto_summarize": True})
|
|
app.upload(adv, "westhaven-canon.md", CANON_MD, "canon")
|
|
|
|
# Enough turns that memories and summaries actually exist. A memory needs
|
|
# MEMORY_INTERVAL (6) actions plus SETTLE_SLACK (1) settled past the anchor,
|
|
# and a summary needs SUMMARY_INTERVAL (15) uncovered actions — both counted
|
|
# in *actions*, and a turn writes two. A first pass at this gate played five
|
|
# turns, wrote neither, and compared 0 against 0, which proves nothing about
|
|
# whether the upgrade preserves them.
|
|
beats = [
|
|
"I ask Mara what the bell means.",
|
|
"I show her the silver key.",
|
|
"I follow her to the harbour ledger.",
|
|
"I ask who else knows about the key.",
|
|
"I read the ledger's last page aloud.",
|
|
"I ask the ferryman about the fen road.",
|
|
"I wait out the rain and watch the harbour.",
|
|
"I ask Mara about the abbey's sealed crypt.",
|
|
"I count the entries against the tide table.",
|
|
"I ask who signed for the last shipment.",
|
|
"I walk the quay to the chandler's door.",
|
|
"I ask the chandler what he remembers of that night.",
|
|
"I show the chandler the key.",
|
|
"I return to Mara with what he said.",
|
|
"I ask Mara what she means to do now.",
|
|
"I agree to meet her at first light.",
|
|
"I take the long way back along the ridge.",
|
|
"I check whether anyone followed me.",
|
|
"I write down what I have learned so far.",
|
|
"I sleep, and wake before the bell.",
|
|
]
|
|
played = 0
|
|
for text in beats:
|
|
if play(app, adv, text):
|
|
played += 1
|
|
print(f" {played} turns accepted by v1.0.0", flush=True)
|
|
|
|
app.call("POST", f"/adventures/{adv}/checkpoints", {"name": "before the ledger"})
|
|
# One Undo, so the head is not at the retained tip and Redo is available.
|
|
app.call("POST", f"/adventures/{adv}/undo", {})
|
|
|
|
# §11 item 9: the database must carry **loopback or placeholder settings with
|
|
# no real hostnames**. The turns above needed a real narrator, so the
|
|
# endpoint is reset to loopback once the story exists — before the census is
|
|
# taken and before either bundle is exported.
|
|
app.call("PUT", "/settings", {
|
|
"endpoint_url": PLACEHOLDER_ENDPOINT, "model": MODEL,
|
|
"embedding_model": EMBED_MODEL,
|
|
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
|
|
return adv
|
|
|
|
|
|
def main() -> int:
|
|
parser = argparse.ArgumentParser(description=__doc__)
|
|
parser.add_argument("--v100", required=True, help="the v1.0.0 worktree")
|
|
parser.add_argument("--out", required=True)
|
|
args = parser.parse_args()
|
|
|
|
if not (ENDPOINT and MODEL and EMBED_MODEL):
|
|
print("set AIDND_TEST_ENDPOINT, AIDND_TEST_MODEL and AIDND_TEST_EMBED_MODEL")
|
|
return 2
|
|
|
|
out = require_under_home(Path(args.out).expanduser())
|
|
shutil.rmtree(out, ignore_errors=True)
|
|
out.mkdir(parents=True)
|
|
v100_tree = Path(args.v100).expanduser().resolve()
|
|
candidate_tree = Path(__file__).resolve().parent.parent.parent
|
|
db = out / "campaign.db"
|
|
results: dict = {"started": datetime.now().isoformat(timespec="seconds"),
|
|
"v100_tree": str(v100_tree), "candidate": str(candidate_tree)}
|
|
failures: list[str] = []
|
|
|
|
print("phase 1 — v1.0.0 builds and plays the campaign")
|
|
app = App(v100_tree, db, out / "v100-server.log", "v1.0.0")
|
|
try:
|
|
adv = build_v100_campaign(app)
|
|
before = census(app, adv, db)
|
|
v100_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
|
|
(out / "v100-export.json").write_text(json.dumps(v100_bundle))
|
|
finally:
|
|
app.stop()
|
|
results["adventure"], results["before"] = adv, before
|
|
print(f" {before['total']} actions, redo={before['can_redo']}, "
|
|
f"checkpoints={len(before['checkpoints'])}, memories={len(before['memories'])}, "
|
|
f"summaries={len(before['summaries'])}, knowledge={len(before['knowledge'])}, "
|
|
f"schema={before['schema_version']}")
|
|
|
|
print("\nphase 2 — the candidate opens that same database file")
|
|
app = App(candidate_tree, db, out / "candidate-server.log", "candidate")
|
|
try:
|
|
after = census(app, adv, db)
|
|
v11_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
|
|
(out / "v11-export.json").write_text(json.dumps(v11_bundle))
|
|
try:
|
|
imported = app.call("POST", "/adventures/import", v100_bundle, timeout=600)
|
|
results["v100_bundle_into_v11"] = {"status": "imported",
|
|
"id": (imported or {}).get("id")}
|
|
print(" the v1.0.0 bundle imported into the candidate")
|
|
except urllib.error.HTTPError as exc:
|
|
results["v100_bundle_into_v11"] = {
|
|
"status": "refused", "code": exc.code,
|
|
"detail": exc.read().decode()[:300]}
|
|
failures.append("v100_bundle_into_v11")
|
|
print(f" the candidate REFUSED the v1.0.0 bundle: {exc.code}")
|
|
finally:
|
|
app.stop()
|
|
results["after"] = after
|
|
|
|
print(" comparing the census, field by field:")
|
|
for field in CENSUS:
|
|
same = before.get(field) == after.get(field)
|
|
print(f" {'ok ' if same else 'DIFF'} {field}")
|
|
if not same:
|
|
failures.append(field)
|
|
results.setdefault("differences", {})[field] = {
|
|
"before": before.get(field), "after": after.get(field)}
|
|
|
|
print("\nphase 3 — the candidate's bundle offered back to v1.0.0")
|
|
app = App(v100_tree, out / "backward.db", out / "v100-backward.log", "v1.0.0")
|
|
try:
|
|
try:
|
|
back = app.call("POST", "/adventures/import", v11_bundle, timeout=600)
|
|
results["v11_bundle_into_v100"] = {"status": "imported",
|
|
"id": (back or {}).get("id")}
|
|
print(" v1.0.0 ACCEPTED the v1.1 bundle")
|
|
except urllib.error.HTTPError as exc:
|
|
detail = exc.read().decode()[:400]
|
|
results["v11_bundle_into_v100"] = {"status": "refused", "code": exc.code,
|
|
"detail": detail}
|
|
# Reported, not failed: §11 item 9 asks for the result, and the
|
|
# owner's brief asks whether a refusal breaks the compatibility
|
|
# promise — a judgement, not an assertion this script may make.
|
|
print(f" v1.0.0 REFUSED the v1.1 bundle: {exc.code} {detail[:160]}")
|
|
finally:
|
|
app.stop()
|
|
|
|
results["failures"] = failures
|
|
(out / "upgrade-report.json").write_text(json.dumps(results, indent=2, default=str))
|
|
print(f"\n{'PASS' if not failures else 'FAIL'}: {len(failures)} field(s) differ "
|
|
f"-> {out / 'upgrade-report.json'}")
|
|
return 1 if failures else 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
raise SystemExit(main())
|