v1.1 closeout: accept integrated release validation
Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.
V1.1 RELEASE VALIDATION: PASS
What was run, on this candidate:
- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
identical on all 15 census fields, schema parity at user_version 94, and both
bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
verified, public endpoint refused, a real turn, restart, persistence, and
Firefox rendering the reopened campaign.
Carried residuals, stated rather than summarised away:
- WP-B: deterministic independent-memory recovery PASS; reference-model
independent-memory recovery FAIL at memory creation — the owner-accepted
limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
backlog, reproduced and not fixed during validation.
Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.
Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.
Still the owner's to do: sign the release commit, update main, tag v1.1.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
87a40326a2
commit
db7b309e3d
@@ -90,6 +90,11 @@ from app.routers import adventures # noqa: E402
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
#: v1.1 release Gate 7 asks for this diagnostic with **memory on**. It shipped
|
||||
#: with no embedding model and the bank switched off, so a release run of it
|
||||
#: would have reported a clean identity result without memory ever taking part.
|
||||
#: Empty keeps the old behaviour, which is what `--scripted` wants.
|
||||
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
|
||||
|
||||
#: The cast the finding describes: a protagonist and three others, all on stage.
|
||||
CAST = [
|
||||
@@ -140,7 +145,7 @@ def _setup(scripted: bool):
|
||||
db.add(models.Settings(
|
||||
user_id=user.id,
|
||||
model=MODEL or "scripted", endpoint_url=ENDPOINT or "http://127.0.0.1:11434/v1",
|
||||
embedding_model="", context_token_budget=16384, max_output_tokens=500,
|
||||
embedding_model=EMBED_MODEL, context_token_budget=16384, max_output_tokens=500,
|
||||
model_timeout_seconds=300,
|
||||
))
|
||||
db.commit()
|
||||
@@ -167,6 +172,12 @@ def _campaign(client) -> int:
|
||||
})
|
||||
created.raise_for_status()
|
||||
adv = created.json()["id"]
|
||||
# Memory and summaries are per-campaign switches defaulting to off. Gate 7
|
||||
# asks for this diagnostic with memory on, and the ten beats below write
|
||||
# twenty actions — past `MEMORY_START` — so the bank has something to do.
|
||||
client.patch(f"/api/adventures/{adv}",
|
||||
json={"memory_bank_enabled": True, "auto_summarize": True}
|
||||
).raise_for_status()
|
||||
answer = client.post(f"/api/adventures/{adv}/state/corrections", json={
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": key, "entity_type": kind, "name": name}
|
||||
|
||||
Reference in New Issue
Block a user