Stop a turn locking out its own memory bank, and let the long run notice
The first M01 trial with the memory bank on was 26 turns on a GPU host. It
accepted every turn and reported "complete". It also wrote two memories and
no summary, and logged 180 `database is locked` errors, while derived status
still read `idle`.
The cause was a single uncommitted UPDATE. Retrieval bumped each used
memory's counter before the model call, and the turn commits only after the
reply has streamed. SQLite has one writer, so the turn held the write lock for
the whole reply. Every post-turn memory, summary and status write in that
window waited out the five-second timeout and failed. Recording the failure
needed a write as well, and without a rollback first it raised
PendingRollbackError. The loss therefore reached the log and never reached
the status the Insights panel reads, which F08 forbids. The draco run never
hit this because the bank was off there.
- `retrieve_memories` now only reads. `record_use` writes the counters in the
turn's single commit, so a turn that never lands counts nothing.
- The post-turn task's outer handler rolls back before it records a failure.
The harness could not have caught any of this. It read three prompt sections
under names the builder does not use: `memories` (really `used_memories`),
`story_history` (really `history`/`recent_history`), and a `knowledge` prefix
that matched the fixed instruction section instead of the imported passages.
Memory tokens read 0 whatever the prompt held, and the in-history and
in-memories recall checks could never come out true. The labels are now
constants, pinned by a test against a prompt the real builder assembled.
The harness also stops at the first sign of failed post-turn work. It checks
/derived and new server.log lines after every turn, keeps its log position
across --resume, and waits for background work to settle before its final
checks. A run with no memories or no summaries now ends "failed", not
"complete".
Both new application tests fail on fec46f6: the lock probe sees
`database is locked`, and memory status stays `idle`. The full backend suite
passes (1392 passed, 17 skipped). A 26-turn re-run against the same host had
0 lock errors, wrote 7 memories and 2 summaries, and used them in the prompt
from turn 8.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
co-authored by
Claude Opus 5
parent
fec46f66bb
commit
f8d401029f
+170
-13
@@ -37,6 +37,17 @@ never asked. `setup` now turns both on and proves it, and `measure` records how
|
||||
many memories and summaries exist so that "the bank stayed empty" is a number in
|
||||
the evidence rather than an absence nobody looked for.
|
||||
|
||||
Turning both switches on did not make them work. The first run with them on
|
||||
was a 26-turn trial on a GPU host. It accepted every turn and reported
|
||||
"complete" with two memories, no summary and 180 `database is locked` errors in
|
||||
`server.log`. Two defects hid that result. The application lost its own
|
||||
post-turn writes and then could not record the loss. The harness read three
|
||||
prompt sections under names the builder does not use, so it measured 0 memory
|
||||
tokens whatever was in the prompt. The harness now checks derived status and
|
||||
`server.log` after every turn and stops at the first sign of failed background
|
||||
work. A run with no memories or no summaries at the end is reported as
|
||||
`failed`, not `complete`.
|
||||
|
||||
## What it records
|
||||
|
||||
A JSON line per turn (`timeline.jsonl`) carrying the context measurements M03
|
||||
@@ -123,6 +134,35 @@ DEFAULT_MAX_CONSECUTIVE_FAILURES = 5
|
||||
#: Operational state, not evidence. Its presence means an unfinished run.
|
||||
RESUME_FILE = "resume.json"
|
||||
|
||||
#: The prompt sections this harness reads, spelled the way
|
||||
#: `app/context/builder.py` and `app/knowledge/classes.py` spell them. A wrong
|
||||
#: label raises no error. `.get` returns 0 and a membership check returns
|
||||
#: False, so each misspelling becomes a check that can never pass. Three of
|
||||
#: them did: `memories`, `story_history`, and a `knowledge` prefix that matched
|
||||
#: the fixed instruction section instead of the retrieved passages.
|
||||
#: `test_context_memory.py` pins these names against a prompt the real builder
|
||||
#: assembled. They are copied rather than imported, because importing `app`
|
||||
#: here would build a database engine in the harness process.
|
||||
HISTORY_LABELS = ("history", "recent_history")
|
||||
SUMMARY_LABEL = "story_summary"
|
||||
MEMORIES_LABEL = "used_memories"
|
||||
STATE_LABEL = "narrative_state"
|
||||
CANON_LABEL = "campaign_canon"
|
||||
IMPORTED_KNOWLEDGE_LABELS = (
|
||||
"imported_canon_always", "imported_canon", "imported_reference",
|
||||
"imported_inspiration",
|
||||
)
|
||||
|
||||
#: Lines `app/derived.py` and `app/memorybank.py` log when post-turn work fails.
|
||||
#: The second one means the failure could not be written to derived status at
|
||||
#: all. Status alone therefore cannot prove the work is healthy: in the first
|
||||
#: 26-turn GPU trial, 20 failures were logged this way and the status still read
|
||||
#: `idle`.
|
||||
LOG_FAILURE_MARKERS = (
|
||||
"work failed for adventure",
|
||||
"could not record derived-work failure",
|
||||
)
|
||||
|
||||
#: The planted clue. Distinctive enough that its presence anywhere is
|
||||
#: unambiguous, and phrased as something a story would actually establish.
|
||||
CLUE = "the silver key opens the crypt beneath the Old Abbey"
|
||||
@@ -318,6 +358,9 @@ class Run:
|
||||
self.resumed = False
|
||||
self.elapsed_before = 0.0
|
||||
self.session_started = time.monotonic()
|
||||
#: How far `background_failures` has read `server.log`. Carried across a
|
||||
#: resume, so a failure is reported once, not again every session.
|
||||
self.log_offset = 0
|
||||
|
||||
# ------------------------------------------------------------ recording
|
||||
|
||||
@@ -344,6 +387,7 @@ class Run:
|
||||
"server_starts": self.server.starts,
|
||||
"elapsed_seconds": self.elapsed(),
|
||||
"turns_target": self.turns_target,
|
||||
"log_offset": self.log_offset,
|
||||
"written": datetime.now().isoformat(timespec="seconds"),
|
||||
}
|
||||
tmp = self.out / (RESUME_FILE + ".tmp")
|
||||
@@ -363,6 +407,7 @@ class Run:
|
||||
self.beat = prior.get("beat", 0)
|
||||
self.completed_steps = set(prior.get("completed_steps") or [])
|
||||
self.elapsed_before = prior.get("elapsed_seconds", 0)
|
||||
self.log_offset = prior.get("log_offset", 0)
|
||||
self.resumed = True
|
||||
|
||||
def reattach(self) -> None:
|
||||
@@ -542,12 +587,12 @@ class Run:
|
||||
"output_reserve": tokens["output_reserve"],
|
||||
"protected": tokens["protected"],
|
||||
"available_for_history": tokens["available_for_history"],
|
||||
"summary_tokens": sections.get("story_summary", 0),
|
||||
"memory_tokens": sections.get("memories", 0),
|
||||
"summary_tokens": sections.get(SUMMARY_LABEL, 0),
|
||||
"memory_tokens": sections.get(MEMORIES_LABEL, 0),
|
||||
"knowledge_tokens": sum(
|
||||
v for k, v in sections.items() if k.startswith("knowledge")),
|
||||
"state_tokens": sections.get("narrative_state", 0),
|
||||
"canon_tokens": sections.get("campaign_canon", 0),
|
||||
sections.get(label, 0) for label in IMPORTED_KNOWLEDGE_LABELS),
|
||||
"state_tokens": sections.get(STATE_LABEL, 0),
|
||||
"canon_tokens": sections.get(CANON_LABEL, 0),
|
||||
"window_verified": window.get("verified"),
|
||||
"window_tokens": window.get("tokens"),
|
||||
# Whether the *planted clue* is still visible anywhere in the
|
||||
@@ -569,6 +614,66 @@ class Run:
|
||||
except Exception: # noqa: BLE001
|
||||
return -1
|
||||
|
||||
def summary_count(self) -> int:
|
||||
"""How many summaries the campaign has written, on any branch. Never
|
||||
raises, and -1 means unreadable, as for `bank_size`."""
|
||||
try:
|
||||
derived = self.server.call("GET", f"/adventures/{self.adv}/derived") or {}
|
||||
return len(derived.get("summaries") or [])
|
||||
except Exception: # noqa: BLE001
|
||||
return -1
|
||||
|
||||
def background_failures(self) -> list[str]:
|
||||
"""Every sign since the last call that post-turn work failed.
|
||||
|
||||
Two sources, because neither is enough alone. Derived status is what the
|
||||
application says. The server log also catches a failure the application
|
||||
could not record, which is the failure that made the first GPU trial
|
||||
report `idle` over 180 `database is locked` errors. Log lines are read
|
||||
from where the last call stopped, so each failure is reported once.
|
||||
"""
|
||||
found: list[str] = []
|
||||
try:
|
||||
derived = self.server.call("GET", f"/adventures/{self.adv}/derived") or {}
|
||||
for row in derived.get("status") or []:
|
||||
if row.get("status") == "failed":
|
||||
found.append(f"{row.get('kind')}: {(row.get('detail') or '')[:200]}")
|
||||
except Exception as exc: # noqa: BLE001 - an unreadable status is noted, not a failure
|
||||
self.note("derived_unreadable", error=f"{type(exc).__name__}: {exc}"[:200])
|
||||
log = getattr(self.server, "log_path", None)
|
||||
if log is not None and Path(log).exists():
|
||||
with open(log, "rb") as handle:
|
||||
handle.seek(self.log_offset)
|
||||
fresh = handle.read()
|
||||
self.log_offset = handle.tell()
|
||||
for line in fresh.decode(errors="replace").splitlines():
|
||||
if any(marker in line for marker in LOG_FAILURE_MARKERS):
|
||||
found.append(f"server.log: {line.strip()[:200]}")
|
||||
return found
|
||||
|
||||
def settle(self, *, quiet_seconds: int = 30, limit_seconds: int = 900) -> None:
|
||||
"""Waits for post-turn work to stop changing derived status.
|
||||
|
||||
Used before the final checks. The last turn's memory and summary passes
|
||||
run after that turn returns, and a check taken before they finish can
|
||||
miss their failure or their output. No endpoint reports running work, so
|
||||
"settled" means derived status unchanged for `quiet_seconds`."""
|
||||
deadline = time.monotonic() + limit_seconds
|
||||
last, since = None, time.monotonic()
|
||||
while time.monotonic() < deadline:
|
||||
try:
|
||||
now = json.dumps(self.server.call(
|
||||
"GET", f"/adventures/{self.adv}/derived"), sort_keys=True)
|
||||
except Exception: # noqa: BLE001
|
||||
now = None
|
||||
if now is not None and now == last:
|
||||
if time.monotonic() - since >= quiet_seconds:
|
||||
return
|
||||
else:
|
||||
last, since = now, time.monotonic()
|
||||
time.sleep(2)
|
||||
self.note("settle_timeout", limit_seconds=limit_seconds)
|
||||
|
||||
def state(self) -> dict:
|
||||
return self.server.call("GET", f"/adventures/{self.adv}/state")
|
||||
|
||||
@@ -693,6 +798,21 @@ def main() -> int:
|
||||
accepted_turns=run.accepted)
|
||||
run.save_resume()
|
||||
break
|
||||
# Checked on every turn, not only at the end. A memory or summary
|
||||
# pass that fails is lost M01 evidence from that turn onward. A run
|
||||
# that carries on would report "complete" over a bank that stopped
|
||||
# filling, which is what the first GPU trial did.
|
||||
failures = run.background_failures()
|
||||
if failures:
|
||||
aborted = (
|
||||
f"post-turn memory/summary work failed ({len(failures)} "
|
||||
"signs, first: " + failures[0] + "). Stopping with the "
|
||||
"evidence written; see server.log."
|
||||
)
|
||||
run.note("run_aborted", reason=aborted, failures=failures[:20],
|
||||
accepted_turns=run.accepted)
|
||||
run.save_resume()
|
||||
break
|
||||
run.save_resume()
|
||||
|
||||
# ---- M04: the recall check, with controls. ----
|
||||
@@ -716,9 +836,29 @@ def main() -> int:
|
||||
except Exception as exc: # noqa: BLE001
|
||||
run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300])
|
||||
|
||||
# Every turn was accepted, but that does not complete M01. Its
|
||||
# summary/memory clause needs a bank that filled and a summary that was
|
||||
# written, and the recall turn's own post-turn work has to have
|
||||
# finished without failing. Otherwise the run is "failed", not
|
||||
# "complete".
|
||||
failed_reason = None
|
||||
if aborted is None:
|
||||
run.settle()
|
||||
late = run.background_failures()
|
||||
if late:
|
||||
failed_reason = (f"post-turn work failed after the last turn "
|
||||
f"({len(late)} signs, first: {late[0]})")
|
||||
else:
|
||||
failed_reason = _activation_shortfall(run)
|
||||
if failed_reason:
|
||||
run.note("run_failed", reason=failed_reason)
|
||||
|
||||
summary = {
|
||||
"status": "aborted" if aborted else "complete",
|
||||
"status": ("aborted" if aborted
|
||||
else "failed" if failed_reason else "complete"),
|
||||
"aborted_reason": aborted,
|
||||
"failed_reason": failed_reason,
|
||||
"summaries": _or_none(run.summary_count),
|
||||
"accepted_turns": run.accepted,
|
||||
"turns_requested": args.turns,
|
||||
"restarts": server.starts - 1,
|
||||
@@ -742,13 +882,30 @@ def main() -> int:
|
||||
# A finished run has nothing to resume, and the file's absence is
|
||||
# what lets a later run use this directory.
|
||||
resume_path.unlink(missing_ok=True)
|
||||
return 0
|
||||
return 1 if failed_reason else 0
|
||||
return 1
|
||||
finally:
|
||||
server.stop()
|
||||
run.timeline.close()
|
||||
|
||||
|
||||
def _activation_shortfall(run) -> str | None:
|
||||
"""Why M01's summary/memory clause was not exercised, or None if it was.
|
||||
|
||||
A count of -1 means the count could not be read. It counts as a shortfall,
|
||||
because an unreadable bank does not show that the bank filled."""
|
||||
memories, summaries = run.bank_size(), run.summary_count()
|
||||
missing = []
|
||||
if memories <= 0:
|
||||
missing.append(f"memories_in_bank={memories}")
|
||||
if summaries <= 0:
|
||||
missing.append(f"summaries={summaries}")
|
||||
if not missing:
|
||||
return None
|
||||
return ("M01's summary/memory clause was not exercised: "
|
||||
+ ", ".join(missing))
|
||||
|
||||
|
||||
def _or_none(read):
|
||||
"""A summary field worth having when it can be read, and worth skipping when
|
||||
it cannot. An aborted run still reports the fields that do answer."""
|
||||
@@ -957,7 +1114,7 @@ def _recall(run: Run) -> dict:
|
||||
# 1. Is the clue outside the recent-history window? (Precondition, not result.)
|
||||
report = server.call("GET", f"/adventures/{adv}/context")
|
||||
history_text = " ".join(
|
||||
s["text"] for s in report["sections"] if s["label"] == "story_history")
|
||||
s["text"] for s in report["sections"] if s["label"] in HISTORY_LABELS)
|
||||
in_history = CLUE_SENTINEL in history_text
|
||||
|
||||
# 2. Ask about the subject, and see what the application assembles.
|
||||
@@ -973,12 +1130,12 @@ def _recall(run: Run) -> dict:
|
||||
return {
|
||||
"clue_in_recent_history_window": in_history,
|
||||
"clue_in_prompt": CLUE_SENTINEL in whole_prompt,
|
||||
"in_state_section": CLUE_SENTINEL in sections.get("narrative_state", ""),
|
||||
"in_summary_section": CLUE_SENTINEL in sections.get("story_summary", ""),
|
||||
"in_memories_section": CLUE_SENTINEL in sections.get("memories", ""),
|
||||
"in_state_section": CLUE_SENTINEL in sections.get(STATE_LABEL, ""),
|
||||
"in_summary_section": CLUE_SENTINEL in sections.get(SUMMARY_LABEL, ""),
|
||||
"in_memories_section": CLUE_SENTINEL in sections.get(MEMORIES_LABEL, ""),
|
||||
"in_knowledge_sections": any(
|
||||
CLUE_SENTINEL in text for label, text in sections.items()
|
||||
if label.startswith("knowledge")),
|
||||
CLUE_SENTINEL in sections.get(label, "")
|
||||
for label in IMPORTED_KNOWLEDGE_LABELS),
|
||||
"fact_still_in_state": fact_present,
|
||||
"history_included": after["history"]["included"],
|
||||
"history_total": after["history"]["total"],
|
||||
|
||||
Reference in New Issue
Block a user