Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
co-authored by
Claude Opus 5
parent
fedb7144d0
commit
ef25b0a876
@@ -2,9 +2,20 @@
|
||||
|
||||
python -m tools.m11_browser --out <dir> [--show]
|
||||
|
||||
Run from `backend/`, with `frontend/dist` already built. Uses a real narrator
|
||||
when `AIDND_TEST_ENDPOINT`/`AIDND_TEST_MODEL` are set; the checks that do not
|
||||
need narration run either way and say which they are.
|
||||
Run from `backend/`, with `frontend/dist` already built.
|
||||
|
||||
**A narrator is required**, from `AIDND_TEST_ENDPOINT`/`AIDND_TEST_MODEL`, and
|
||||
the run refuses to start without one. `--no-narrator` opts out explicitly, and
|
||||
then every check that needs a played turn *skips* and the run is marked
|
||||
`partial` — it is a smoke test of the deterministic checks, not release evidence.
|
||||
|
||||
That is deliberate, and it is the second harness defect of this shape M11 has
|
||||
found. Without a narrator no turn is ever played, so the campaign has no history:
|
||||
Undo is correctly disabled, `D01` fails, the click that follows it throws, and a
|
||||
run reports two failures that look exactly like a product regression and are not.
|
||||
An unset environment variable must not be able to produce that. `--out` is
|
||||
required on every harness here for the same reason — so the decision is made on
|
||||
purpose rather than by omission.
|
||||
|
||||
## What this is and is not
|
||||
|
||||
@@ -542,8 +553,22 @@ def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--out", required=True)
|
||||
parser.add_argument("--show", action="store_true", help="not headless")
|
||||
parser.add_argument(
|
||||
"--no-narrator", action="store_true",
|
||||
help=("run only the checks that need no narration, and mark the run "
|
||||
"partial. Not release evidence."))
|
||||
args = parser.parse_args()
|
||||
|
||||
narrated = bool(ENDPOINT and MODEL) and not args.no_narrator
|
||||
if not narrated and not args.no_narrator:
|
||||
print("AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL are not set.\n"
|
||||
"A browser release regression needs a narrator: without one no turn "
|
||||
"is played, the campaign has no history, and the history controls "
|
||||
"fail for a reason that is not the product's.\n"
|
||||
"Set both, or pass --no-narrator to run the deterministic checks "
|
||||
"alone and get a run marked partial.")
|
||||
return 2
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
dist = BACKEND.parent / "frontend" / "dist" / "index.html"
|
||||
@@ -560,12 +585,12 @@ def main() -> int:
|
||||
print(f"build: {dist.stat().st_mtime} served at {site.url}\n")
|
||||
|
||||
try:
|
||||
if ENDPOINT and MODEL:
|
||||
if narrated:
|
||||
site.api("PUT", "/settings", {
|
||||
"endpoint_url": ENDPOINT, "model": MODEL,
|
||||
"max_output_tokens": 400, "model_timeout_seconds": 600})
|
||||
adv = campaign_with_story(site, checks)
|
||||
if ENDPOINT and MODEL:
|
||||
if narrated:
|
||||
for text in ("I ask Mara what she has heard.",
|
||||
"I show her the silver key."):
|
||||
events = play_a_turn(site, adv, text)
|
||||
@@ -574,11 +599,21 @@ def main() -> int:
|
||||
not errors, errors[0].get("detail", "")[:120] if errors else "")
|
||||
else:
|
||||
checks.skip("B01", "narration through a real model",
|
||||
"AIDND_TEST_ENDPOINT/MODEL not set")
|
||||
"--no-narrator")
|
||||
|
||||
# The history controls read a story that has turns in it, and only a
|
||||
# narrator puts turns there. Skipped rather than run against an empty
|
||||
# campaign, because Undo being correctly disabled is not a D01 failure.
|
||||
history_controls = (
|
||||
(lambda: check_history_controls(browser, site, adv, checks))
|
||||
if narrated else
|
||||
(lambda: checks.skip("D01-D14", "history controls in the browser",
|
||||
"--no-narrator: the campaign has no turns"))
|
||||
)
|
||||
|
||||
for scenario in (
|
||||
lambda: check_shell_and_title(browser, site, adv, checks),
|
||||
lambda: check_history_controls(browser, site, adv, checks),
|
||||
history_controls,
|
||||
lambda: check_markdown_safety(browser, site, checks),
|
||||
lambda: check_hidden_knowledge(browser, site, checks),
|
||||
lambda: check_context_inspection(browser, site, adv, checks),
|
||||
@@ -599,7 +634,10 @@ def main() -> int:
|
||||
"browser": f"Firefox {browser.version}",
|
||||
"started": started.isoformat(timespec="seconds"),
|
||||
"seconds": round((datetime.now() - started).total_seconds()),
|
||||
"narrator": MODEL or "none (deterministic checks only)",
|
||||
"narrator": MODEL if narrated else "none (--no-narrator)",
|
||||
# Release evidence, or a smoke test. A reader should not have to infer
|
||||
# which from the skip count.
|
||||
"kind": "release regression" if narrated else "partial (no narrator)",
|
||||
"checks": checks.rows,
|
||||
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
|
||||
"failed": len(checks.failed),
|
||||
@@ -608,6 +646,9 @@ def main() -> int:
|
||||
(out / "browser-report.json").write_text(json.dumps(report, indent=2))
|
||||
print(f"\n{report['passed']} passed, {report['failed']} failed, "
|
||||
f"{report['skipped']} skipped -> {out / 'browser-report.json'}")
|
||||
if not narrated:
|
||||
print("PARTIAL: no narrator, so the narration and history checks did "
|
||||
"not run. This is not release evidence.")
|
||||
return 1 if checks.failed else 0
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user