0c7316f95132f01d080098472d86952d97ee4383
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0c7316f951 |
Keep the state section and its proposal out of the story
The first complete M01 run with the memory bank on (
|
||
|
|
f8d401029f |
Stop a turn locking out its own memory bank, and let the long run notice
The first M01 trial with the memory bank on was 26 turns on a GPU host. It
accepted every turn and reported "complete". It also wrote two memories and
no summary, and logged 180 `database is locked` errors, while derived status
still read `idle`.
The cause was a single uncommitted UPDATE. Retrieval bumped each used
memory's counter before the model call, and the turn commits only after the
reply has streamed. SQLite has one writer, so the turn held the write lock for
the whole reply. Every post-turn memory, summary and status write in that
window waited out the five-second timeout and failed. Recording the failure
needed a write as well, and without a rollback first it raised
PendingRollbackError. The loss therefore reached the log and never reached
the status the Insights panel reads, which F08 forbids. The draco run never
hit this because the bank was off there.
- `retrieve_memories` now only reads. `record_use` writes the counters in the
turn's single commit, so a turn that never lands counts nothing.
- The post-turn task's outer handler rolls back before it records a failure.
The harness could not have caught any of this. It read three prompt sections
under names the builder does not use: `memories` (really `used_memories`),
`story_history` (really `history`/`recent_history`), and a `knowledge` prefix
that matched the fixed instruction section instead of the imported passages.
Memory tokens read 0 whatever the prompt held, and the in-history and
in-memories recall checks could never come out true. The labels are now
constants, pinned by a test against a prompt the real builder assembled.
The harness also stops at the first sign of failed post-turn work. It checks
/derived and new server.log lines after every turn, keeps its log position
across --resume, and waits for background work to settle before its final
checks. A run with no memories or no summaries now ends "failed", not
"complete".
Both new application tests fail on
|
||
|
|
fec46f66bb |
Turn the memory bank on for the long run, and refuse one that cannot use it
The first complete hundred-turn campaign did not exercise M01's "summary/memory activation" step. Memory bank and auto-summarize are per-campaign switches that default to off, and m11_long_run never turned them on: summary_tokens and memory_tokens were 0 on every turn, memories_used was empty, and M04's clue was recalled through narrative state alone. The retrieval path M6 built was never asked, and nothing in the evidence said so except a row of zeros. setup now PATCHes both switches on, reads the campaign back, and stops before the first turn if either did not take. memories_in_bank is recorded on every turn, in the final summary and in the recall, and the recall also says whether a summary exists, so which of the two recall paths succeeded is stated rather than implied. The embedding model is now required. Without one the summary pass still writes memories, but memorybank.retrieve answers "No embedding model configured" and returns none -- the same unexercised path in a fuller bank. The harness refuses before it starts a server or claims --out. tests/test_m11_long_run_memory.py drives setup against the real application in-process: the switches are on afterwards, a server that ignores the PATCH is refused before any state is written, the bank count comes from the application and reads -1 rather than raising when it cannot, and a run with no embedding model is refused. The four that exercise setup and the bank count were run against the previous harness and fail there; the premise test (a fresh campaign has both switches off) passes on both, as it should. The 2026-09-10 run in ~/m11-evidence/m01 therefore does not count as M01. It has to be run again on this harness. Backend 1,382 passed, 18 skipped, 0 failed. The eighteenth skip is test_built_spa_fetches_no_fonts_remotely, which wants a built frontend/dist this worktree does not have; it is an environment condition, not a change here. The frontend is untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XKWHt2DXuvqP83cAk6Zq88 |