Stop a turn locking out its own memory bank, and let the long run notice

The first M01 trial with the memory bank on was 26 turns on a GPU host. It
accepted every turn and reported "complete". It also wrote two memories and
no summary, and logged 180 `database is locked` errors, while derived status
still read `idle`.

The cause was a single uncommitted UPDATE. Retrieval bumped each used
memory's counter before the model call, and the turn commits only after the
reply has streamed. SQLite has one writer, so the turn held the write lock for
the whole reply. Every post-turn memory, summary and status write in that
window waited out the five-second timeout and failed. Recording the failure
needed a write as well, and without a rollback first it raised
PendingRollbackError. The loss therefore reached the log and never reached
the status the Insights panel reads, which F08 forbids. The draco run never
hit this because the bank was off there.

- `retrieve_memories` now only reads. `record_use` writes the counters in the
  turn's single commit, so a turn that never lands counts nothing.
- The post-turn task's outer handler rolls back before it records a failure.

The harness could not have caught any of this. It read three prompt sections
under names the builder does not use: `memories` (really `used_memories`),
`story_history` (really `history`/`recent_history`), and a `knowledge` prefix
that matched the fixed instruction section instead of the imported passages.
Memory tokens read 0 whatever the prompt held, and the in-history and
in-memories recall checks could never come out true. The labels are now
constants, pinned by a test against a prompt the real builder assembled.

The harness also stops at the first sign of failed post-turn work. It checks
/derived and new server.log lines after every turn, keeps its log position
across --resume, and waits for background work to settle before its final
checks. A run with no memories or no summaries now ends "failed", not
"complete".

Both new application tests fail on fec46f6: the lock probe sees
`database is locked`, and memory status stays `idle`. The full backend suite
passes (1392 passed, 17 skipped). A 26-turn re-run against the same host had
0 lock errors, wrote 7 memories and 2 summaries, and used them in the prompt
from turn 8.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
This commit is contained in:
JesseMarkowitz
2026-09-13 20:21:21 -04:00
co-authored by Claude Opus 5
parent fec46f66bb
commit f8d401029f
9 changed files with 431 additions and 37 deletions
+170 -13
View File
@@ -37,6 +37,17 @@ never asked. `setup` now turns both on and proves it, and `measure` records how
many memories and summaries exist so that "the bank stayed empty" is a number in
the evidence rather than an absence nobody looked for.
Turning both switches on did not make them work. The first run with them on
was a 26-turn trial on a GPU host. It accepted every turn and reported
"complete" with two memories, no summary and 180 `database is locked` errors in
`server.log`. Two defects hid that result. The application lost its own
post-turn writes and then could not record the loss. The harness read three
prompt sections under names the builder does not use, so it measured 0 memory
tokens whatever was in the prompt. The harness now checks derived status and
`server.log` after every turn and stops at the first sign of failed background
work. A run with no memories or no summaries at the end is reported as
`failed`, not `complete`.
## What it records
A JSON line per turn (`timeline.jsonl`) carrying the context measurements M03
@@ -123,6 +134,35 @@ DEFAULT_MAX_CONSECUTIVE_FAILURES = 5
#: Operational state, not evidence. Its presence means an unfinished run.
RESUME_FILE = "resume.json"
#: The prompt sections this harness reads, spelled the way
#: `app/context/builder.py` and `app/knowledge/classes.py` spell them. A wrong
#: label raises no error. `.get` returns 0 and a membership check returns
#: False, so each misspelling becomes a check that can never pass. Three of
#: them did: `memories`, `story_history`, and a `knowledge` prefix that matched
#: the fixed instruction section instead of the retrieved passages.
#: `test_context_memory.py` pins these names against a prompt the real builder
#: assembled. They are copied rather than imported, because importing `app`
#: here would build a database engine in the harness process.
HISTORY_LABELS = ("history", "recent_history")
SUMMARY_LABEL = "story_summary"
MEMORIES_LABEL = "used_memories"
STATE_LABEL = "narrative_state"
CANON_LABEL = "campaign_canon"
IMPORTED_KNOWLEDGE_LABELS = (
"imported_canon_always", "imported_canon", "imported_reference",
"imported_inspiration",
)
#: Lines `app/derived.py` and `app/memorybank.py` log when post-turn work fails.
#: The second one means the failure could not be written to derived status at
#: all. Status alone therefore cannot prove the work is healthy: in the first
#: 26-turn GPU trial, 20 failures were logged this way and the status still read
#: `idle`.
LOG_FAILURE_MARKERS = (
"work failed for adventure",
"could not record derived-work failure",
)
#: The planted clue. Distinctive enough that its presence anywhere is
#: unambiguous, and phrased as something a story would actually establish.
CLUE = "the silver key opens the crypt beneath the Old Abbey"
@@ -318,6 +358,9 @@ class Run:
self.resumed = False
self.elapsed_before = 0.0
self.session_started = time.monotonic()
#: How far `background_failures` has read `server.log`. Carried across a
#: resume, so a failure is reported once, not again every session.
self.log_offset = 0
# ------------------------------------------------------------ recording
@@ -344,6 +387,7 @@ class Run:
"server_starts": self.server.starts,
"elapsed_seconds": self.elapsed(),
"turns_target": self.turns_target,
"log_offset": self.log_offset,
"written": datetime.now().isoformat(timespec="seconds"),
}
tmp = self.out / (RESUME_FILE + ".tmp")
@@ -363,6 +407,7 @@ class Run:
self.beat = prior.get("beat", 0)
self.completed_steps = set(prior.get("completed_steps") or [])
self.elapsed_before = prior.get("elapsed_seconds", 0)
self.log_offset = prior.get("log_offset", 0)
self.resumed = True
def reattach(self) -> None:
@@ -542,12 +587,12 @@ class Run:
"output_reserve": tokens["output_reserve"],
"protected": tokens["protected"],
"available_for_history": tokens["available_for_history"],
"summary_tokens": sections.get("story_summary", 0),
"memory_tokens": sections.get("memories", 0),
"summary_tokens": sections.get(SUMMARY_LABEL, 0),
"memory_tokens": sections.get(MEMORIES_LABEL, 0),
"knowledge_tokens": sum(
v for k, v in sections.items() if k.startswith("knowledge")),
"state_tokens": sections.get("narrative_state", 0),
"canon_tokens": sections.get("campaign_canon", 0),
sections.get(label, 0) for label in IMPORTED_KNOWLEDGE_LABELS),
"state_tokens": sections.get(STATE_LABEL, 0),
"canon_tokens": sections.get(CANON_LABEL, 0),
"window_verified": window.get("verified"),
"window_tokens": window.get("tokens"),
# Whether the *planted clue* is still visible anywhere in the
@@ -569,6 +614,66 @@ class Run:
except Exception: # noqa: BLE001
return -1
def summary_count(self) -> int:
"""How many summaries the campaign has written, on any branch. Never
raises, and -1 means unreadable, as for `bank_size`."""
try:
derived = self.server.call("GET", f"/adventures/{self.adv}/derived") or {}
return len(derived.get("summaries") or [])
except Exception: # noqa: BLE001
return -1
def background_failures(self) -> list[str]:
"""Every sign since the last call that post-turn work failed.
Two sources, because neither is enough alone. Derived status is what the
application says. The server log also catches a failure the application
could not record, which is the failure that made the first GPU trial
report `idle` over 180 `database is locked` errors. Log lines are read
from where the last call stopped, so each failure is reported once.
"""
found: list[str] = []
try:
derived = self.server.call("GET", f"/adventures/{self.adv}/derived") or {}
for row in derived.get("status") or []:
if row.get("status") == "failed":
found.append(f"{row.get('kind')}: {(row.get('detail') or '')[:200]}")
except Exception as exc: # noqa: BLE001 - an unreadable status is noted, not a failure
self.note("derived_unreadable", error=f"{type(exc).__name__}: {exc}"[:200])
log = getattr(self.server, "log_path", None)
if log is not None and Path(log).exists():
with open(log, "rb") as handle:
handle.seek(self.log_offset)
fresh = handle.read()
self.log_offset = handle.tell()
for line in fresh.decode(errors="replace").splitlines():
if any(marker in line for marker in LOG_FAILURE_MARKERS):
found.append(f"server.log: {line.strip()[:200]}")
return found
def settle(self, *, quiet_seconds: int = 30, limit_seconds: int = 900) -> None:
"""Waits for post-turn work to stop changing derived status.
Used before the final checks. The last turn's memory and summary passes
run after that turn returns, and a check taken before they finish can
miss their failure or their output. No endpoint reports running work, so
"settled" means derived status unchanged for `quiet_seconds`."""
deadline = time.monotonic() + limit_seconds
last, since = None, time.monotonic()
while time.monotonic() < deadline:
try:
now = json.dumps(self.server.call(
"GET", f"/adventures/{self.adv}/derived"), sort_keys=True)
except Exception: # noqa: BLE001
now = None
if now is not None and now == last:
if time.monotonic() - since >= quiet_seconds:
return
else:
last, since = now, time.monotonic()
time.sleep(2)
self.note("settle_timeout", limit_seconds=limit_seconds)
def state(self) -> dict:
return self.server.call("GET", f"/adventures/{self.adv}/state")
@@ -693,6 +798,21 @@ def main() -> int:
accepted_turns=run.accepted)
run.save_resume()
break
# Checked on every turn, not only at the end. A memory or summary
# pass that fails is lost M01 evidence from that turn onward. A run
# that carries on would report "complete" over a bank that stopped
# filling, which is what the first GPU trial did.
failures = run.background_failures()
if failures:
aborted = (
f"post-turn memory/summary work failed ({len(failures)} "
"signs, first: " + failures[0] + "). Stopping with the "
"evidence written; see server.log."
)
run.note("run_aborted", reason=aborted, failures=failures[:20],
accepted_turns=run.accepted)
run.save_resume()
break
run.save_resume()
# ---- M04: the recall check, with controls. ----
@@ -716,9 +836,29 @@ def main() -> int:
except Exception as exc: # noqa: BLE001
run.note("export_failed", error=f"{type(exc).__name__}: {exc}"[:300])
# Every turn was accepted, but that does not complete M01. Its
# summary/memory clause needs a bank that filled and a summary that was
# written, and the recall turn's own post-turn work has to have
# finished without failing. Otherwise the run is "failed", not
# "complete".
failed_reason = None
if aborted is None:
run.settle()
late = run.background_failures()
if late:
failed_reason = (f"post-turn work failed after the last turn "
f"({len(late)} signs, first: {late[0]})")
else:
failed_reason = _activation_shortfall(run)
if failed_reason:
run.note("run_failed", reason=failed_reason)
summary = {
"status": "aborted" if aborted else "complete",
"status": ("aborted" if aborted
else "failed" if failed_reason else "complete"),
"aborted_reason": aborted,
"failed_reason": failed_reason,
"summaries": _or_none(run.summary_count),
"accepted_turns": run.accepted,
"turns_requested": args.turns,
"restarts": server.starts - 1,
@@ -742,13 +882,30 @@ def main() -> int:
# A finished run has nothing to resume, and the file's absence is
# what lets a later run use this directory.
resume_path.unlink(missing_ok=True)
return 0
return 1 if failed_reason else 0
return 1
finally:
server.stop()
run.timeline.close()
def _activation_shortfall(run) -> str | None:
"""Why M01's summary/memory clause was not exercised, or None if it was.
A count of -1 means the count could not be read. It counts as a shortfall,
because an unreadable bank does not show that the bank filled."""
memories, summaries = run.bank_size(), run.summary_count()
missing = []
if memories <= 0:
missing.append(f"memories_in_bank={memories}")
if summaries <= 0:
missing.append(f"summaries={summaries}")
if not missing:
return None
return ("M01's summary/memory clause was not exercised: "
+ ", ".join(missing))
def _or_none(read):
"""A summary field worth having when it can be read, and worth skipping when
it cannot. An aborted run still reports the fields that do answer."""
@@ -957,7 +1114,7 @@ def _recall(run: Run) -> dict:
# 1. Is the clue outside the recent-history window? (Precondition, not result.)
report = server.call("GET", f"/adventures/{adv}/context")
history_text = " ".join(
s["text"] for s in report["sections"] if s["label"] == "story_history")
s["text"] for s in report["sections"] if s["label"] in HISTORY_LABELS)
in_history = CLUE_SENTINEL in history_text
# 2. Ask about the subject, and see what the application assembles.
@@ -973,12 +1130,12 @@ def _recall(run: Run) -> dict:
return {
"clue_in_recent_history_window": in_history,
"clue_in_prompt": CLUE_SENTINEL in whole_prompt,
"in_state_section": CLUE_SENTINEL in sections.get("narrative_state", ""),
"in_summary_section": CLUE_SENTINEL in sections.get("story_summary", ""),
"in_memories_section": CLUE_SENTINEL in sections.get("memories", ""),
"in_state_section": CLUE_SENTINEL in sections.get(STATE_LABEL, ""),
"in_summary_section": CLUE_SENTINEL in sections.get(SUMMARY_LABEL, ""),
"in_memories_section": CLUE_SENTINEL in sections.get(MEMORIES_LABEL, ""),
"in_knowledge_sections": any(
CLUE_SENTINEL in text for label, text in sections.items()
if label.startswith("knowledge")),
CLUE_SENTINEL in sections.get(label, "")
for label in IMPORTED_KNOWLEDGE_LABELS),
"fact_still_in_state": fact_present,
"history_included": after["history"]["included"],
"history_total": after["history"]["total"],