Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
This commit is contained in:
co-authored by
Claude Opus 5
parent
fedb7144d0
commit
ef25b0a876
@@ -2,9 +2,20 @@
|
||||
|
||||
python -m tools.m11_browser --out <dir> [--show]
|
||||
|
||||
Run from `backend/`, with `frontend/dist` already built. Uses a real narrator
|
||||
when `AIDND_TEST_ENDPOINT`/`AIDND_TEST_MODEL` are set; the checks that do not
|
||||
need narration run either way and say which they are.
|
||||
Run from `backend/`, with `frontend/dist` already built.
|
||||
|
||||
**A narrator is required**, from `AIDND_TEST_ENDPOINT`/`AIDND_TEST_MODEL`, and
|
||||
the run refuses to start without one. `--no-narrator` opts out explicitly, and
|
||||
then every check that needs a played turn *skips* and the run is marked
|
||||
`partial` — it is a smoke test of the deterministic checks, not release evidence.
|
||||
|
||||
That is deliberate, and it is the second harness defect of this shape M11 has
|
||||
found. Without a narrator no turn is ever played, so the campaign has no history:
|
||||
Undo is correctly disabled, `D01` fails, the click that follows it throws, and a
|
||||
run reports two failures that look exactly like a product regression and are not.
|
||||
An unset environment variable must not be able to produce that. `--out` is
|
||||
required on every harness here for the same reason — so the decision is made on
|
||||
purpose rather than by omission.
|
||||
|
||||
## What this is and is not
|
||||
|
||||
@@ -542,8 +553,22 @@ def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--out", required=True)
|
||||
parser.add_argument("--show", action="store_true", help="not headless")
|
||||
parser.add_argument(
|
||||
"--no-narrator", action="store_true",
|
||||
help=("run only the checks that need no narration, and mark the run "
|
||||
"partial. Not release evidence."))
|
||||
args = parser.parse_args()
|
||||
|
||||
narrated = bool(ENDPOINT and MODEL) and not args.no_narrator
|
||||
if not narrated and not args.no_narrator:
|
||||
print("AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL are not set.\n"
|
||||
"A browser release regression needs a narrator: without one no turn "
|
||||
"is played, the campaign has no history, and the history controls "
|
||||
"fail for a reason that is not the product's.\n"
|
||||
"Set both, or pass --no-narrator to run the deterministic checks "
|
||||
"alone and get a run marked partial.")
|
||||
return 2
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
dist = BACKEND.parent / "frontend" / "dist" / "index.html"
|
||||
@@ -560,12 +585,12 @@ def main() -> int:
|
||||
print(f"build: {dist.stat().st_mtime} served at {site.url}\n")
|
||||
|
||||
try:
|
||||
if ENDPOINT and MODEL:
|
||||
if narrated:
|
||||
site.api("PUT", "/settings", {
|
||||
"endpoint_url": ENDPOINT, "model": MODEL,
|
||||
"max_output_tokens": 400, "model_timeout_seconds": 600})
|
||||
adv = campaign_with_story(site, checks)
|
||||
if ENDPOINT and MODEL:
|
||||
if narrated:
|
||||
for text in ("I ask Mara what she has heard.",
|
||||
"I show her the silver key."):
|
||||
events = play_a_turn(site, adv, text)
|
||||
@@ -574,11 +599,21 @@ def main() -> int:
|
||||
not errors, errors[0].get("detail", "")[:120] if errors else "")
|
||||
else:
|
||||
checks.skip("B01", "narration through a real model",
|
||||
"AIDND_TEST_ENDPOINT/MODEL not set")
|
||||
"--no-narrator")
|
||||
|
||||
# The history controls read a story that has turns in it, and only a
|
||||
# narrator puts turns there. Skipped rather than run against an empty
|
||||
# campaign, because Undo being correctly disabled is not a D01 failure.
|
||||
history_controls = (
|
||||
(lambda: check_history_controls(browser, site, adv, checks))
|
||||
if narrated else
|
||||
(lambda: checks.skip("D01-D14", "history controls in the browser",
|
||||
"--no-narrator: the campaign has no turns"))
|
||||
)
|
||||
|
||||
for scenario in (
|
||||
lambda: check_shell_and_title(browser, site, adv, checks),
|
||||
lambda: check_history_controls(browser, site, adv, checks),
|
||||
history_controls,
|
||||
lambda: check_markdown_safety(browser, site, checks),
|
||||
lambda: check_hidden_knowledge(browser, site, checks),
|
||||
lambda: check_context_inspection(browser, site, adv, checks),
|
||||
@@ -599,7 +634,10 @@ def main() -> int:
|
||||
"browser": f"Firefox {browser.version}",
|
||||
"started": started.isoformat(timespec="seconds"),
|
||||
"seconds": round((datetime.now() - started).total_seconds()),
|
||||
"narrator": MODEL or "none (deterministic checks only)",
|
||||
"narrator": MODEL if narrated else "none (--no-narrator)",
|
||||
# Release evidence, or a smoke test. A reader should not have to infer
|
||||
# which from the skip count.
|
||||
"kind": "release regression" if narrated else "partial (no narrator)",
|
||||
"checks": checks.rows,
|
||||
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
|
||||
"failed": len(checks.failed),
|
||||
@@ -608,6 +646,9 @@ def main() -> int:
|
||||
(out / "browser-report.json").write_text(json.dumps(report, indent=2))
|
||||
print(f"\n{report['passed']} passed, {report['failed']} failed, "
|
||||
f"{report['skipped']} skipped -> {out / 'browser-report.json'}")
|
||||
if not narrated:
|
||||
print("PARTIAL: no narrator, so the narration and history checks did "
|
||||
"not run. This is not release evidence.")
|
||||
return 1 if checks.failed else 0
|
||||
|
||||
|
||||
|
||||
+350
-39
@@ -28,10 +28,47 @@ recall check at the end asks the application what it would actually send.
|
||||
|
||||
## What it records
|
||||
|
||||
A JSON line per turn (`timeline.jsonl`) with the context measurements M03 wants,
|
||||
a `measurements.json` of the sampled checkpoints, the recall evidence for M04,
|
||||
and the final bundle. Everything is written as it happens, so a run that dies at
|
||||
turn 80 still leaves 80 turns of evidence rather than nothing.
|
||||
A JSON line per turn (`timeline.jsonl`) carrying the context measurements M03
|
||||
wants, the recall evidence for M04 (`recall.json`), the exported campaign
|
||||
(`bundle.json`) and `summary.json`. Everything is written as it happens, so a
|
||||
run that dies at turn 80 still leaves 80 turns of evidence rather than nothing.
|
||||
|
||||
## Resuming
|
||||
|
||||
A hundred turns is hours of wall clock, and the first release attempt lost one
|
||||
at turn 97 to a host crash. `timeline.jsonl` survived that; the run did not,
|
||||
because starting the harness again began a new campaign at turn 1.
|
||||
|
||||
So a run now checkpoints `resume.json` beside its evidence — after the prologue,
|
||||
after every scheduled operation, and after every turn — and `--resume` picks the
|
||||
campaign back up where it stopped: same adventure row, same accepted count, same
|
||||
place in the beat cycle, and the scheduled operations that already fired are not
|
||||
fired again. The file is written under a temporary name and renamed, because the
|
||||
failure it exists to survive is the host dying mid-write.
|
||||
|
||||
`resume.json` is operational state, never evidence. `timeline.jsonl` remains the
|
||||
append-only record and nothing here rewrites it. A finished run deletes its
|
||||
`resume.json`, which makes the file's presence mean exactly one thing: there is
|
||||
an unfinished run in this directory. The harness refuses to start a fresh
|
||||
campaign in a directory that already holds one, because two campaigns
|
||||
interleaved in one timeline are worse evidence than none.
|
||||
|
||||
## Timeouts, and why they are an option rather than a constant
|
||||
|
||||
How long a turn takes belongs to the inference host, not to the application. On
|
||||
the reference host a turn cost 229-291 seconds at the recommended window, which
|
||||
fits comfortably inside the 600-second model timeout this harness used to
|
||||
hard-code. A slower host does not, and the consequence was not a slow run: a
|
||||
turn that overran the timeout raised out of the loop and ended the run with a
|
||||
traceback and no summary.
|
||||
|
||||
`--turn-timeout` sets what the application will wait for one narrator reply, and
|
||||
the harness waits longer still, so that the application's own error arrives
|
||||
inside the stream rather than being cut off at the socket. The default is
|
||||
deliberately generous; measure your host before lowering it.
|
||||
`--max-consecutive-failures` ends a run that has stopped producing turns, with
|
||||
its evidence and a summary written and `--resume` still able to continue it,
|
||||
instead of spinning against a narrator that is not answering.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -55,11 +92,52 @@ ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
|
||||
|
||||
#: What the application will wait for one narrator reply, in seconds. The
|
||||
#: settings schema bounds this at 30..3600 (`app/schemas.py`), and this default
|
||||
#: sits high inside that range deliberately: an overrun turn is a lost turn, and
|
||||
#: over a hundred of them the cost of waiting is far below the cost of a restart.
|
||||
DEFAULT_TURN_TIMEOUT = 1800
|
||||
|
||||
#: How much longer the harness waits than the application does. The application
|
||||
#: has to be the thing that times out, because it reports the failure inside the
|
||||
#: stream the harness is reading; a harness that gave up first would record a
|
||||
#: socket error and throw away what the application was about to say.
|
||||
HARNESS_TIMEOUT_MARGIN = 300
|
||||
|
||||
#: Turns in a row not accepted before the run stops and writes what it has. One
|
||||
#: rejected turn is ordinary — `failed_call` causes one on purpose. A run of
|
||||
#: them means the narrator is gone, and every further attempt costs a timeout.
|
||||
DEFAULT_MAX_CONSECUTIVE_FAILURES = 5
|
||||
|
||||
#: Operational state, not evidence. Its presence means an unfinished run.
|
||||
RESUME_FILE = "resume.json"
|
||||
|
||||
#: The planted clue. Distinctive enough that its presence anywhere is
|
||||
#: unambiguous, and phrased as something a story would actually establish.
|
||||
CLUE = "the silver key opens the crypt beneath the Old Abbey"
|
||||
CLUE_SENTINEL = "SILVER-KEY-CRYPT-OLD-ABBEY"
|
||||
|
||||
#: The clue as accepted state, which is the half of M04 that does not depend on
|
||||
#: the narrator remembering anything.
|
||||
#:
|
||||
#: The content goes in `value`. `add_fact` requires `predicate` and accepts
|
||||
#: `subject`, `object`, `value` and `fact_id` — and nothing else
|
||||
#: (`app/narrative/events.py` SPECS). An earlier version of this harness put the
|
||||
#: clue in a `detail` key, which that event does not define: the correction was
|
||||
#: accepted, the fact was created, and the clue text went nowhere. The stored
|
||||
#: fact said only that Aldric knows of the abbey, so `_recall`'s
|
||||
#: `fact_still_in_state` could not answer True however well the application
|
||||
#: behaved. A check that can only fail is worse than no check, and this is the
|
||||
#: second harness defect of that shape M11 has found.
|
||||
CLUE_FACT = {
|
||||
"type": "add_fact",
|
||||
"subject": "aldric",
|
||||
"predicate": "knows",
|
||||
"object": "abbey",
|
||||
"value": f"{CLUE} ({CLUE_SENTINEL})",
|
||||
"fact_id": "silver-key-opens-crypt",
|
||||
}
|
||||
|
||||
CANON = [
|
||||
"The dead do not return. No rite, relic or bargain has ever returned anyone.",
|
||||
"The abbey crypt has been sealed since the founding.",
|
||||
@@ -115,12 +193,15 @@ def free_port() -> int:
|
||||
class Storyteller:
|
||||
"""The real application, started the way `DEVELOPMENT.md` says to start it."""
|
||||
|
||||
def __init__(self, db_path: Path, log: Path):
|
||||
def __init__(self, db_path: Path, log: Path, *,
|
||||
turn_timeout: int = DEFAULT_TURN_TIMEOUT):
|
||||
self.db_path = db_path
|
||||
self.port = free_port()
|
||||
self.log_path = log
|
||||
self.proc = None
|
||||
self.starts = 0
|
||||
self.turn_timeout = turn_timeout
|
||||
self.stream_timeout = turn_timeout + HARNESS_TIMEOUT_MARGIN
|
||||
|
||||
def start(self) -> None:
|
||||
self.starts += 1
|
||||
@@ -176,7 +257,7 @@ class Storyteller:
|
||||
body = response.read().decode()
|
||||
return json.loads(body) if body else None
|
||||
|
||||
def stream(self, path: str, payload, timeout=900) -> list[dict]:
|
||||
def stream(self, path: str, payload, timeout=None) -> list[dict]:
|
||||
"""A turn. The reply is SSE, and a failed turn is an event, not a status.
|
||||
|
||||
`app/sse.py`: "A failed turn is still an HTTP 200 response, because the
|
||||
@@ -184,6 +265,8 @@ class Storyteller:
|
||||
harness that read the status code would call every failure a success —
|
||||
which is precisely the class of harness defect M8's review warned about.
|
||||
"""
|
||||
if timeout is None:
|
||||
timeout = self.stream_timeout
|
||||
request = urllib.request.Request(
|
||||
f"http://127.0.0.1:{self.port}/api{path}",
|
||||
data=json.dumps(payload).encode(), method="POST",
|
||||
@@ -204,13 +287,26 @@ class Storyteller:
|
||||
class Run:
|
||||
"""One long campaign, and everything measured about it."""
|
||||
|
||||
def __init__(self, server: Storyteller, out: Path):
|
||||
def __init__(self, server: Storyteller, out: Path, *, turns_target: int,
|
||||
turn_timeout: int = DEFAULT_TURN_TIMEOUT):
|
||||
self.server = server
|
||||
self.out = out
|
||||
self.timeline = (out / "timeline.jsonl").open("a")
|
||||
self.adv = 0
|
||||
self.accepted = 0
|
||||
self.events: list[dict] = []
|
||||
self.turns_target = turns_target
|
||||
self.turn_timeout = turn_timeout
|
||||
#: Where the beat cycle stands. Carried across a resume, so a continued
|
||||
#: campaign keeps moving rather than replaying its first ten beats.
|
||||
self.beat = 0
|
||||
#: Plan keys whose scheduled operation has already fired, keyed by turn
|
||||
#: number rather than by step name: two of the steps are called `retry`
|
||||
#: and two `restart`, so a name does not identify one.
|
||||
self.completed_steps: set[int] = set()
|
||||
self.resumed = False
|
||||
self.elapsed_before = 0.0
|
||||
self.session_started = time.monotonic()
|
||||
|
||||
# ------------------------------------------------------------ recording
|
||||
|
||||
@@ -221,17 +317,76 @@ class Run:
|
||||
self.timeline.write(json.dumps(entry, default=str) + "\n")
|
||||
self.timeline.flush()
|
||||
|
||||
# ------------------------------------------------------------- resuming
|
||||
|
||||
def elapsed(self) -> int:
|
||||
"""Seconds of run time, across every session this campaign has had."""
|
||||
return round(self.elapsed_before + (time.monotonic() - self.session_started))
|
||||
|
||||
def save_resume(self) -> None:
|
||||
"""Checkpoint enough to pick this campaign up again, atomically."""
|
||||
payload = {
|
||||
"adventure": self.adv,
|
||||
"accepted": self.accepted,
|
||||
"beat": self.beat,
|
||||
"completed_steps": sorted(self.completed_steps),
|
||||
"server_starts": self.server.starts,
|
||||
"elapsed_seconds": self.elapsed(),
|
||||
"turns_target": self.turns_target,
|
||||
"written": datetime.now().isoformat(timespec="seconds"),
|
||||
}
|
||||
tmp = self.out / (RESUME_FILE + ".tmp")
|
||||
tmp.write_text(json.dumps(payload, indent=2))
|
||||
tmp.replace(self.out / RESUME_FILE)
|
||||
|
||||
def adopt(self, prior: dict) -> None:
|
||||
"""Take on the state a previous session checkpointed.
|
||||
|
||||
A checkpoint can be at most one turn behind the database, because a turn
|
||||
is committed by the application before this file is written. Erring that
|
||||
way costs one extra turn on a campaign that wants *at least* a hundred,
|
||||
which is the harmless direction.
|
||||
"""
|
||||
self.adv = prior["adventure"]
|
||||
self.accepted = prior["accepted"]
|
||||
self.beat = prior.get("beat", 0)
|
||||
self.completed_steps = set(prior.get("completed_steps") or [])
|
||||
self.elapsed_before = prior.get("elapsed_seconds", 0)
|
||||
self.resumed = True
|
||||
|
||||
def reattach(self) -> None:
|
||||
"""Prove the campaign is still there, and restate what a run needs.
|
||||
|
||||
Settings are re-applied rather than trusted. They live in the database,
|
||||
and `failed_call` deliberately points the model at a name the server
|
||||
does not serve before putting it back; a host that died inside that
|
||||
window left the campaign configured to fail every turn it is given.
|
||||
"""
|
||||
page = self.server.call("GET", f"/adventures/{self.adv}/actions?limit=1")
|
||||
self.apply_settings()
|
||||
self.note("resumed", adventure=self.adv, accepted=self.accepted,
|
||||
beat=self.beat, actions_in_db=page["total"],
|
||||
completed_steps=sorted(self.completed_steps),
|
||||
process_starts=self.server.starts,
|
||||
elapsed_before_seconds=round(self.elapsed_before))
|
||||
|
||||
# ------------------------------------------------------------- campaign
|
||||
|
||||
def setup(self) -> None:
|
||||
def apply_settings(self) -> dict:
|
||||
settings = self.server.call("PUT", "/settings", {
|
||||
"endpoint_url": ENDPOINT, "model": MODEL,
|
||||
"embedding_model": EMBED_MODEL, "context_token_budget": 16384,
|
||||
"max_output_tokens": 500, "model_timeout_seconds": 600,
|
||||
"max_output_tokens": 500,
|
||||
"model_timeout_seconds": self.turn_timeout,
|
||||
"memory_top_k": 4,
|
||||
})
|
||||
self.note("settings", model=settings["model"],
|
||||
budget=settings["context_token_budget"])
|
||||
budget=settings["context_token_budget"],
|
||||
model_timeout_seconds=settings["model_timeout_seconds"])
|
||||
return settings
|
||||
|
||||
def setup(self) -> None:
|
||||
self.apply_settings()
|
||||
|
||||
created = self.server.call("POST", "/adventures", {
|
||||
"title": "Continuity Test (M11 long run)",
|
||||
@@ -308,6 +463,13 @@ class Run:
|
||||
self.note("turn_failed", text=text, status=exc.code,
|
||||
detail=exc.read().decode()[:300])
|
||||
return {"accepted": False}
|
||||
except Exception as exc: # noqa: BLE001
|
||||
# A socket timeout, or a connection dropped mid-stream. This used to
|
||||
# propagate out of the loop and end the run with a traceback and no
|
||||
# summary, which on a slow host is the likeliest way to lose one.
|
||||
self.note("turn_failed", text=text,
|
||||
detail=f"{type(exc).__name__}: {exc}"[:300])
|
||||
return {"accepted": False}
|
||||
errors = [e for e in events if e.get("type") == "error"]
|
||||
if errors:
|
||||
self.note("turn_error", text=text, detail=errors[0].get("detail", "")[:300])
|
||||
@@ -334,6 +496,15 @@ class Run:
|
||||
return {
|
||||
"total_actions": report["history"]["total"],
|
||||
"history_included": report["history"]["included"],
|
||||
# Where the history window was cut, and the step it takes when it
|
||||
# moves. A `floor_depth` that is the same on two consecutive turns
|
||||
# is the prompt's prefix having been preserved, which is the whole
|
||||
# of what trimming the window in blocks buys; a run that recorded
|
||||
# neither could not say whether it engaged, held or stepped, and a
|
||||
# continuity finding could not be attributed. `.get` because a
|
||||
# campaign resumed against an older build has neither.
|
||||
"history_floor_depth": report["history"].get("floor_depth"),
|
||||
"history_trim_block": report["history"].get("trim_block"),
|
||||
"prompt_tokens": tokens["total"],
|
||||
"budget": tokens["budget"],
|
||||
"configured_budget": tokens.get("configured_budget"),
|
||||
@@ -370,77 +541,217 @@ def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--turns", type=int, default=100)
|
||||
parser.add_argument("--out", required=True)
|
||||
parser.add_argument(
|
||||
"--resume", action="store_true",
|
||||
help="continue the unfinished run in --out rather than starting one")
|
||||
parser.add_argument(
|
||||
"--turn-timeout", type=int, default=DEFAULT_TURN_TIMEOUT,
|
||||
help=("seconds the application waits for one narrator reply "
|
||||
f"(30-3600, default {DEFAULT_TURN_TIMEOUT})"))
|
||||
parser.add_argument(
|
||||
"--max-consecutive-failures", type=int,
|
||||
default=DEFAULT_MAX_CONSECUTIVE_FAILURES,
|
||||
help="stop and write the evidence after this many unaccepted turns")
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL):
|
||||
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
|
||||
return 2
|
||||
if not 30 <= args.turn_timeout <= 3600:
|
||||
# The settings schema's own bound, checked here so a mistyped timeout
|
||||
# fails in the first second rather than on the first PUT.
|
||||
print("--turn-timeout must be between 30 and 3600 seconds")
|
||||
return 2
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
db_path = out / "campaign.db"
|
||||
server = Storyteller(db_path, out / "server.log")
|
||||
resume_path = out / RESUME_FILE
|
||||
|
||||
prior = _resume_state(resume_path, out / "timeline.jsonl", args.resume)
|
||||
if isinstance(prior, str):
|
||||
print(prior)
|
||||
return 2
|
||||
|
||||
server = Storyteller(db_path, out / "server.log",
|
||||
turn_timeout=args.turn_timeout)
|
||||
if prior:
|
||||
server.starts = prior.get("server_starts", 1)
|
||||
server.start()
|
||||
run = Run(server, out)
|
||||
started = datetime.now()
|
||||
run = Run(server, out, turns_target=args.turns,
|
||||
turn_timeout=args.turn_timeout)
|
||||
if prior:
|
||||
run.adopt(prior)
|
||||
|
||||
try:
|
||||
run.setup()
|
||||
if run.resumed:
|
||||
run.reattach()
|
||||
else:
|
||||
run.setup()
|
||||
|
||||
# ---- The planted clue, at the very beginning. ----
|
||||
run.turn(f"I tell Mara quietly that {CLUE} — {CLUE_SENTINEL}.")
|
||||
run.correct([
|
||||
{"type": "add_fact", "subject": "aldric", "predicate": "knows",
|
||||
"object": "abbey", "detail": f"{CLUE} ({CLUE_SENTINEL})",
|
||||
"fact_id": "silver-key-opens-crypt"},
|
||||
], note="the planted clue, as accepted state")
|
||||
run.note("clue_planted", sentinel=CLUE_SENTINEL)
|
||||
# ---- The planted clue, at the very beginning. ----
|
||||
run.turn(f"I tell Mara quietly that {CLUE} — {CLUE_SENTINEL}.")
|
||||
run.correct([CLUE_FACT], note="the planted clue, as accepted state")
|
||||
# Proved here, at turn one, where it costs a single request. M04
|
||||
# asks whether the clue is still recoverable a hundred turns later,
|
||||
# and that question is meaningless if it was never stored — so a
|
||||
# run whose clue did not land is stopped rather than spending hours
|
||||
# measuring nothing. See CLUE_FACT for how this went wrong before.
|
||||
planted = any(CLUE_SENTINEL in json.dumps(fact) for fact
|
||||
in run.state()["document"].get("facts") or [])
|
||||
run.note("clue_planted", sentinel=CLUE_SENTINEL,
|
||||
verified_in_state=planted)
|
||||
if not planted:
|
||||
raise SystemExit(
|
||||
"the planted clue is not in accepted state, so M04 cannot "
|
||||
"be measured from this run. Stopping before the campaign "
|
||||
"starts rather than reporting a recall failure later.")
|
||||
# The first checkpoint, and the point from which --resume works: the
|
||||
# campaign exists and its clue is planted.
|
||||
run.save_resume()
|
||||
|
||||
# ---- The long middle. ----
|
||||
plan = _schedule(args.turns)
|
||||
beat = 0
|
||||
consecutive_failures = 0
|
||||
aborted = None
|
||||
while run.accepted < args.turns:
|
||||
step = plan.get(run.accepted + 1)
|
||||
if step:
|
||||
at = run.accepted + 1
|
||||
step = plan.get(at)
|
||||
if step is not None and at not in run.completed_steps:
|
||||
try:
|
||||
_do_step(run, server, step)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
# A step that fails is a finding, not a reason to lose the
|
||||
# other ninety turns. It is recorded loudly and the campaign
|
||||
# goes on, because an abandoned run proves nothing at all.
|
||||
run.note("step_failed", step=step,
|
||||
run.note("step_failed", step=step, at=at,
|
||||
error=f"{type(exc).__name__}: {exc}"[:300])
|
||||
run.turn(BEATS[beat % len(BEATS)])
|
||||
beat += 1
|
||||
# Fired, however it went. Before this the step was keyed only by
|
||||
# the accepted count, so a refused turn afterwards ran the whole
|
||||
# operation again — and a restart or a retry performed twice is
|
||||
# not the test the schedule describes.
|
||||
run.completed_steps.add(at)
|
||||
run.save_resume()
|
||||
|
||||
result = run.turn(BEATS[run.beat % len(BEATS)])
|
||||
run.beat += 1
|
||||
if result.get("accepted"):
|
||||
consecutive_failures = 0
|
||||
else:
|
||||
consecutive_failures += 1
|
||||
if consecutive_failures >= args.max_consecutive_failures:
|
||||
aborted = (
|
||||
f"{consecutive_failures} turns in a row were not accepted; "
|
||||
"the narrator is not answering. Stopping with the evidence "
|
||||
"written, so --resume can carry this campaign on."
|
||||
)
|
||||
run.note("run_aborted", reason=aborted,
|
||||
accepted_turns=run.accepted)
|
||||
run.save_resume()
|
||||
break
|
||||
run.save_resume()
|
||||
|
||||
# ---- M04: the recall check, with controls. ----
|
||||
run.note("recall_begin")
|
||||
recall = _recall(run)
|
||||
(out / "recall.json").write_text(json.dumps(recall, indent=2))
|
||||
# Skipped on an aborted run: it asks the narrator a question, and the
|
||||
# reason the run stopped is that the narrator does not answer.
|
||||
recall = None
|
||||
if aborted is None:
|
||||
run.note("recall_begin")
|
||||
recall = _recall(run)
|
||||
(out / "recall.json").write_text(json.dumps(recall, indent=2))
|
||||
|
||||
# ---- Export the whole thing, for the recovery evidence. ----
|
||||
bundle = server.call("GET", f"/adventures/{run.adv}/export")
|
||||
(out / "bundle.json").write_text(json.dumps(bundle))
|
||||
run.note("exported", bytes=len((out / "bundle.json").read_bytes()))
|
||||
# ---- Export whatever exists, for the recovery evidence. ----
|
||||
# Attempted even for an aborted run: the recovery check and the storage
|
||||
# numbers are worth having at whatever turn count was reached.
|
||||
bundle_bytes = None
|
||||
try:
|
||||
bundle = server.call("GET", f"/adventures/{run.adv}/export")
|
||||
(out / "bundle.json").write_text(json.dumps(bundle))
|
||||
bundle_bytes = len((out / "bundle.json").read_bytes())
|
||||
run.note("exported", bytes=bundle_bytes)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300])
|
||||
|
||||
summary = {
|
||||
"status": "aborted" if aborted else "complete",
|
||||
"aborted_reason": aborted,
|
||||
"accepted_turns": run.accepted,
|
||||
"turns_requested": args.turns,
|
||||
"restarts": server.starts - 1,
|
||||
"elapsed_seconds": round((datetime.now() - started).total_seconds()),
|
||||
"process_starts": server.starts,
|
||||
"resumed": run.resumed,
|
||||
"elapsed_seconds": run.elapsed(),
|
||||
"turn_timeout_seconds": args.turn_timeout,
|
||||
"recall": recall,
|
||||
"final_state": run.state()["document"],
|
||||
"final_measurement": run.measure(),
|
||||
"final_state": _or_none(lambda: run.state()["document"]),
|
||||
"final_measurement": _or_none(run.measure),
|
||||
"db_bytes": db_path.stat().st_size,
|
||||
"bundle_bytes": bundle_bytes,
|
||||
}
|
||||
(out / "summary.json").write_text(json.dumps(summary, indent=2, default=str))
|
||||
print(json.dumps({k: v for k, v in summary.items()
|
||||
if k not in ("final_state",)}, indent=2, default=str)[:2000])
|
||||
return 0
|
||||
|
||||
if aborted is None:
|
||||
# A finished run has nothing to resume, and the file's absence is
|
||||
# what lets a later run use this directory.
|
||||
resume_path.unlink(missing_ok=True)
|
||||
return 0
|
||||
return 1
|
||||
finally:
|
||||
server.stop()
|
||||
run.timeline.close()
|
||||
|
||||
|
||||
def _or_none(read):
|
||||
"""A summary field worth having when it can be read, and worth skipping when
|
||||
it cannot. An aborted run still reports the fields that do answer."""
|
||||
try:
|
||||
return read()
|
||||
except Exception as exc: # noqa: BLE001
|
||||
return {"unavailable": f"{type(exc).__name__}: {exc}"[:200]}
|
||||
|
||||
|
||||
def _resume_state(resume_path: Path, timeline_path: Path, resuming: bool):
|
||||
"""The previous session's checkpoint, or a message saying why there is none.
|
||||
|
||||
Returns the parsed checkpoint, `None` for a clean start, or a string to
|
||||
print before exiting. The refusals matter as much as the resume: two
|
||||
campaigns interleaved in one `timeline.jsonl` and one `campaign.db` are
|
||||
worse evidence than none, and that is what starting fresh on top of an
|
||||
existing run produces.
|
||||
|
||||
A non-empty `timeline.jsonl` is what says a previous attempt got far enough
|
||||
to record something, and it is the harness's own artifact rather than the
|
||||
application's. Testing it rather than `campaign.db` means a run that died
|
||||
before it recorded anything — a server that never came up, a missing
|
||||
endpoint — leaves the directory usable, because nothing was written into it
|
||||
to collide with.
|
||||
"""
|
||||
if resuming:
|
||||
if not resume_path.exists():
|
||||
return (f"--resume: there is no {resume_path.name} in "
|
||||
f"{resume_path.parent}. Either the run never reached its "
|
||||
"first checkpoint, or this is the wrong directory.")
|
||||
try:
|
||||
prior = json.loads(resume_path.read_text())
|
||||
except (OSError, json.JSONDecodeError) as exc:
|
||||
return f"--resume: cannot read {resume_path}: {exc}"
|
||||
if not prior.get("adventure"):
|
||||
return f"--resume: {resume_path} names no campaign."
|
||||
return prior
|
||||
if resume_path.exists():
|
||||
return (f"{resume_path} exists, so this directory holds an unfinished "
|
||||
"run. Pass --resume to carry it on, or choose a new --out.")
|
||||
if timeline_path.exists() and timeline_path.stat().st_size > 0:
|
||||
return (f"{timeline_path} already has entries but there is no "
|
||||
f"{resume_path.name}, so a previous run recorded something here "
|
||||
"and then died before its first checkpoint. Choose a new --out: "
|
||||
"starting here would put a second campaign in the same timeline "
|
||||
"and the same database.")
|
||||
return None
|
||||
|
||||
|
||||
def _schedule(turns: int) -> dict:
|
||||
"""Where each required history operation happens. Spread, not clustered."""
|
||||
unit = max(1, turns // 13)
|
||||
|
||||
Reference in New Issue
Block a user