Diagnostic only; no memory behaviour changes. - tools/memory_diagnostic.py: planted-fact isolation checks, the four-stage diagnosis (created / retained / ranked / injected) with a verdict, a production-ranking replica, deterministic summariser/embedder/narrator stubs and seven scenarios (default, past capacity, pinned, low top_k, long-block early/late, lineage control) - tools/v11_b1_memory.py: CLI for the scenarios and for diagnosing a copy of a finished real campaign - tools/m11_long_run.py: opt-in --independent-fact mode with per-turn isolation tracking and the recovered_through_memory_independent verdict; M04 verdicts unchanged - tests: diagnostic stages, eviction, creation window, ranking, lineage and authority controls; two strict xfails record the diagnosed retention and creation defects for WP-B.2 to flip - planning/reports/v1.1/V1.1-WP-B1-REPORT.md First failing stage: ranking (real model); retention past capacity and creation for early facts in long blocks (deterministic, same on v1.0.0). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
38 KiB
v1.1 WP-B.1 — Independent Long-Term Memory Retention Diagnostic
Status: COMPLETE. Diagnostic only: no memory behaviour was changed. The final decision is in §S.
A. Repository baseline
| Branch | v1.1-development |
| HEAD at start | d63804f22ecbaed80741241a154cdaef82f7b2ed — v1.1: harden context window and narrator protocol boundary (WP-A1/A2), signed by the owner (good signature, RSA key 02C9BF7D…) |
| Its parents | ac465ed (planning v4.1, signed) and 432f041 (the signed v1.0.0 release commit; tag v1.0.0; main) |
| Working tree at start | clean |
git diff --stat v1.0.0..HEAD |
27 files, +4,880 / −88: the planning commit and WP-A1/A2 |
| Comparison baseline | 432f041 (v1.0.0), run from a throwaway worktree (§D) |
app/memorybank.py, app/summaries.py, app/context/lineage.py, app/tree.py and
app/vectors.py are unchanged between v1.0.0 and HEAD. A1 and A2 did not touch
the memory pipeline, apart from adding accounting to attempts.ATTEMPT_KEYS.
B. Existing memory pipeline
Answered from the code at HEAD. Nothing was changed.
| # | Question | Answer |
|---|---|---|
| 1 | How are memories generated? | memorybank.run_post_turn, a fire-and-forget task after each accepted turn (schedule_post_turn), runs _create_due_memories when the campaign has auto_summarize. It writes one memory per block of MEMORY_INTERVAL = 6 story actions past the memory cursor. It starts once the story has MEMORY_START = 12 actions, and only when SETTLE_SLACK = 1 action sits past the block. At most MAX_MEMORIES_PER_RUN = 5 memories are written per run. |
| 2 | What range does a memory cover? | The block's first and last action depths: Memory.source_start, Memory.source_end. |
| 3 | How are source depth and lineage stored? | tree.attach_memory sets Memory.branch_id and Memory.depth from the block's last node, so a memory is visible exactly on paths that contain that node. |
| 4 | Memory text length limit | Prompt-only: MEMORY_MAX_WORDS = 50 in MEMORY_SYSTEM_PROMPT ("1-2 plain sentences"). Nothing truncates the stored text. |
| 5 | What input does the summariser receive? | summarize_block: a cast brief (cast_brief, which reads story cards, persona and plot essentials), then "Story excerpt:\n\n{excerpt}\n\nMemory:". The excerpt is the block's action texts joined by blank lines. |
| 6 | Where does the 2,000-token truncation happen? | summarize_block: excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS), with MEMORY_EXCERPT_TOKENS = 2,000. It keeps the block's last 2,000 cl100k_base tokens. The cast brief is matched against the untruncated block. |
| 7 | How does capacity and eviction work? | _evict_over_capacity, at the end of every run_post_turn. It counts non-forgotten memories in the whole adventure (not lineage-scoped). Above settings.memory_bank_capacity (default 80), it marks forgotten = True on the overflow unpinned memories, ordered by coalesce(last_used_at, created_at) ascending, then use_count. Forgotten rows are kept. |
| 8 | How is last_accessed updated? |
The field is Memory.last_used_at. record_use sets it, and increments use_count, for every memory in the turn's memories.used. That write is part of the turn's single commit (§O.7 of the M11 report). Dry runs never count. |
| 9 | How are memories ranked? | retrieve_memories. The query is the text of the newest RETRIEVAL_WINDOW_ACTIONS = 4 actions, truncated to the last RETRIEVAL_WINDOW_TOKENS = 600 tokens. It is embedded with the configured embedding model. Every eligible memory is scored by vectors.cosine. Pinned memories are taken first, the rest fill memory_top_k (default 5) in score order, and _drop_redundant skips a candidate at cosine ≥ 0.93 to one already chosen, within the same authority. There is no lexical term, recency term, importance term or similarity floor (CONTEXT-AND-MEMORY §20, as implemented). |
| 10 | How do pins affect ranking and eviction? | Ranking: a pinned memory is always selected, and counts toward memory_top_k. Eviction: pinned memories are never evicted, and if every active memory is pinned, capacity is exceeded. |
| 11 | How does the active-lineage clause filter memories? | lineage.path_of(db, adventure).clause(models.Memory) matches (branch_id, depth) against the head's path entries, capped at each fork depth. Retrieval also requires forgotten = False and embedded = True. |
| 12 | How do retrieved memories enter build_context? |
Through the memory_bank argument. The builder renders "Memories from earlier in the story. Lines marked [inferred] are interpretation, not established fact — do not treat them as settled truth:" plus one - [inferred]? text line per memory, as section used_memories. It is a live section priced into protected context, placed after history and before narrative_state. |
| 13 | Where is the provenance recorded? | context_snapshot["memories"]: the whole retrieval result, meaning used (id, text, similarity, pinned, authority, source range), considered and suppressed. It is stored per turn, so a past turn's selection is inspectable. |
| 14 | How does memory survive export and import? | bundle._exported_memory carries text, pinned, forgotten, sourceStart, sourceEnd, useCount, authority, branch and depth. It does not carry the vector, last_used_at or created_at. On import the memories are re-embedded by the post-turn pass, and their recency restarts from the import. |
A finding from the inspection itself. The v1 long-run harness (tools/m11_long_run.py)
planted M04's clue as an accepted state correction (CLUE_FACT, via
add_fact). The player turn mentioning it uses the sentinel code, but the fact the
recall looks for was established in state. Memories are written from story
text. So every v1 M04 recovery could only ever run through state, or through the
narrator restating state, and none could have tested memory on its own. That is why
§P risk 5 of the M11 report could say "no run showed memory keeping a planted fact"
without any run having given memory the chance.
C. Diagnostic design
C.1 What is built
| File | Kind | Purpose |
|---|---|---|
backend/tools/memory_diagnostic.py |
tool, new | See below: the fact spec, isolation checks, four-stage diagnosis, deterministic stubs and scenario runner. |
backend/tools/v11_b1_memory.py |
tool, new | CLI. scenarios runs the deterministic campaigns against an isolated database and writes JSON. diagnose runs the four stages against a copy of a finished real campaign's database. |
backend/tools/m11_long_run.py |
tool, extended | --independent-fact: see §C.3. |
backend/tests/test_v11_b1_memory_diagnostic.py |
tests, new | The deterministic diagnostic. |
backend/tests/test_v11_b1_long_run_verdict.py |
tests, new | The new long-run verdict. |
What tools/memory_diagnostic.py holds:
Fact: a planted fact, with whole-word carry and leak matching.isolation(): every non-memory layer checked.diagnose(): the four stages and the verdict.rank_bank(): production ranking recomputed for every eligible memory.- The stubs:
BestCaseSummariser,ConceptEmbedderandScriptNarrator. run_scenario(): a campaign played through the real turn route.
No application file is changed. No column, table, migration or setting is added. The diagnostic's extra fields are computed at report time.
C.2 How each stage is judged
| Stage | Judged from | Output (real field names where they exist) |
|---|---|---|
| created | Memories on the active lineage with source_start ≤ plant_depth ≤ source_end whose text carries F (both "sundial" and "teapot"). For every covering memory, the block is re-read (memorybank.source_block) and cut exactly as summarize_block does, to report whether F was in the block and whether it was in the excerpt the summariser saw. |
memory_id, source_start, source_end, memory_text, covering_memories[].{block_tokens, fact_in_block, fact_in_summariser_excerpt} |
| retained | That memory's row, plus the eviction order production would use (the same ORDER BY, read-only) |
forgotten, pinned, embedded, on_active_lineage, use_count, last_used_at, created_at, active_memories, memory_bank_capacity, eviction_position, reason |
| ranked | rank_bank: the same catalogue clause, cosine, pin rule and memorybank._drop_redundant, over every eligible memory, with the recall turn's own query (history.tail(4, exclude=recall AI node), cut to 600 tokens). It is checked against the recall turn's stored memories.used (replica_matches_stored_selection). |
semantic_score (similarity), lexical_score (always None: no such term exists), final_score, rank, of, top_k_cutoff, selected, suppressed_as_duplicate_of, query |
| injected | The recall turn's stored context_snapshot: memories.used names the memory, and its text is in the used_memories section |
context_component, in_stored_memories_used, text_in_section, token_count |
Verdicts, in order: not_created, created_but_evicted, retained_but_not_ranked,
ranked_but_not_selected, selected_but_not_injected, injected. The fifth is
added to the brief's list, so that "the retrieval picked it" and "the narrator was
shown it" stay distinguishable.
C.3 Isolation (precondition) checks
A result counts only if every check holds.
| Check | How |
|---|---|
state_document |
adventure.narrative_state: entities, facts, relationships, threads, scene and possessions, plus the whole document |
state_snapshots |
narrative_state_after of every node on the active lineage |
later_narration |
every AI turn deeper than the planting block's end, and before the recall turn |
summary |
summaries.current and the recall prompt's story_summary section |
knowledge |
every KnowledgeSource.content, and the recall prompt's imported-knowledge sections |
recent_history |
the recall prompt's history.floor_depth is greater than the planting depth, and F is not in the history or recent_history sections |
state_section |
the recall prompt's narrative_state section |
The negative control for the precondition itself is
test_the_isolation_check_fails_when_another_layer_carries_the_fact.
C.4 The deterministic stubs, and what they model
BestCaseSummariser. An ideal memory writer. A memory keeps every sentence of its excerpt that carries a planted fact, plus one sentence naming the block's own place so that memories differ. Summary updates never mention a planted fact. A creation failure under this stub is the application's, not a model's.ConceptEmbedder. A 96-dimension deterministic embedding. Words in a small concept table ("sundial", "dial", "hour", "clock", …) share a dimension, other words are hashed, and the vector is normalised. It models a paraphrase landing near the original. It says nothing aboutnomic-embed-text.ScriptNarrator. Narration that names only filler places and never a planted fact, with an empty state block, so state never records F.
Scenarios are played through the real POST /actions route. Automatic post-turn
scheduling is replaced by an explicit run_post_turn settle after every turn, so
eviction happens at a known turn. Depths: the opening is 0, turn n's player
action is 2n−1 and its reply 2n. The planted fact is a story action.
C.4.1 Scenarios
| Scenario | Turns | Capacity | top_k | History budget | Prose per reply | Planted at |
|---|---|---|---|---|---|---|
independent_default |
52 (recall at depth 106) | 80 | 5 | 4,096 | ~60 words | depth 1 |
past_capacity |
52 | 6 | 5 | 4,096 | ~60 words | depth 1 |
past_capacity_pinned |
52 | 6 | 5 | 4,096 | ~60 words | depth 1; first other memory pinned |
past_capacity_low_top_k |
52 | 8 | 2 | 4,096 | ~60 words | depth 1 |
long_block_fact_early |
10 | 80 | 5 | 16,384 | ~850 words | depth 1, early in a long block |
long_block_fact_late |
10 | 80 | 5 | 16,384 | ~850 words | depth 5, late in the same-sized block |
lineage_control |
52 | 80 | 5 | 4,096 | ~60 words | F at depth 1; G on line A, then Undo × 9 and divergence at turn 30; Save Points before and after G |
C.5 The long-run verdict
tools/m11_long_run.py --independent-fact plants a second, story-only fact
at depth 3, right after M04's own plant, and never corrects it into state. Every
existing M04 behaviour and verdict is unchanged.
Every accepted turn records the first failure of:
absent_from_state;absent_from_summary;absent_from_later_narration, meaning narration deeper than the planting depth plus 6.
The plant and those first failures survive --resume.
At recall a dedicated question is played. _independent_recall reads the
recall turn's stored context and the campaign database (read-only), and
_independent_memory_verdict returns one of:
recovered_through_memory_independent, only when every one ofplanted_turn_outside_history,absent_from_state,absent_from_summary,absent_from_knowledgeandabsent_from_later_narrationholds, and a memory covering the planting turn carries the fact and was injected;precondition_failed:<name>orprecondition_unknown:<name>, never a recovery;not_recovered:not_created,not_recovered:evictedornot_recovered:not_injected.
Ranking is recomputed afterwards by tools/v11_b1_memory.py diagnose, against a
copy of the run's database.
D. v1.0.0 baseline
How it was run.
- A throwaway worktree was checked out at
432f041(git describe:v1.0.0). - Only the three B.1 files were copied in:
tools/memory_diagnostic.py,tools/v11_b1_memory.pyandtests/test_v11_b1_memory_diagnostic.py. - The imported
appwas confirmed to come from the worktree. - The worktree was removed afterwards, and the
v1.0.0tag and commit were not touched. app/memorybank.pyis byte-identical betweenv1.0.0and HEAD.- Evidence:
$HOME/v11-evidence/b1/v100/, with HEAD's in$HOME/v11-evidence/b1/head/.
v1.0.0 (432f041) |
HEAD (d63804f + B.1 files) |
|
|---|---|---|
test_v11_b1_memory_diagnostic.py |
24 passed, 2 xfailed (strict) | 24 passed, 2 xfailed (strict) |
independent_default |
injected |
injected |
past_capacity (capacity 6, top_k 5) |
created_but_evicted |
created_but_evicted |
past_capacity_pinned |
created_but_evicted (the pinned memory kept) |
same |
past_capacity_low_top_k (capacity 8, top_k 2) |
created_but_evicted |
created_but_evicted |
long_block_fact_early |
not_created |
not_created |
long_block_fact_late |
injected |
injected |
lineage_control |
injected; G never injected after the divergence |
same |
The two independent-retention acceptance criteria fail on v1.0.0, and each names the stage.
criterion: an early fact is recalled from memory past capacity
created: yes
retained: no
FAILURE STAGE: retention (capacity eviction)
criterion: a fact early in a long block is remembered
created: no (the fact was in the block, not in the summariser's excerpt)
FAILURE STAGE: creation (input truncation)
Under the best-case summariser, with blocks shorter than 2,000 tokens and a bank under capacity, v1.0.0 carries the fact all the way to injection. The two failure stages above are therefore application mechanisms, reached under conditions a long campaign meets:
- a block of long narration;
- more memories than
memory_bank_capacity.
Which of them a real campaign meets first is §K's question.
E. Creation results
long_block_fact_early and long_block_fact_late use the same block geometry: the
opening (depth 0) and turns 1-3. Each reply is about 850 words, and the block
source_start 0 … source_end 5 is 2,079 tokens, 79 over
MEMORY_EXCERPT_TOKENS.
| Shape | Planted at | Fact in block | Fact in summariser excerpt | Memory written | Verdict |
|---|---|---|---|---|---|
| F early in the block | depth 1 | yes | no | "The travellers spent time at the ferry landing." | not_created |
| F late in the same-sized block | depth 5 | yes | yes | "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." | injected |
summarize_block keeps the last 2,000 tokens. An early-block fact is cut off
before the summariser reads it, even by an overflow of only 79 tokens. No summariser
quality can recover what it was never given.
independent_default: 60-word replies, a 4,096 budget, and a block under 2,000
tokens. Memory 1 covers depths 0-5, and its text carries F. Both the block and the
excerpt contain F.
F. Retention / capacity results
Default: the bank holds 17 active memories at depth 106, against capacity 80. F's
memory is retained, has use_count 16, and sits at eviction position 14 of 17.
Past capacity (the brief's capacity test). 17 memories are written over 52 turns.
| Scenario | F's uses before eviction | F's last retrieval | First eviction | F evicted | F first evicted? | Created and evicted in the same pass | Pinned kept |
|---|---|---|---|---|---|---|---|
past_capacity (6 / top_k 5) |
12 | turn 18 | turn 21, memory 1 | turn 21 | yes | none | — |
past_capacity_pinned (6 / top_k 5, memory 2 pinned) |
12 | turn 18 | turn 21, memory 1 | turn 21 | yes | none | memory 2 never evicted |
past_capacity_low_top_k (8 / top_k 2) |
4 | turn 24 | turn 27, memory 3 | turn 36 (4th eviction) | no | none | — |
What the trace shows about the rule.
- "Never retrieved" is not the mechanism. F was retrieved, 12 or 4 times, while it still ranked in the top-k for the recent-narration query.
- Eviction orders by
coalesce(last_used_at, created_at). An early fact that recent narration never mentions stops being retrieved once newer memories fill the top-k. It then ages out:- at capacity 6, it is gone 3 turns after its last use, as the first eviction;
- at capacity 8 with
top_k2, 12 turns after its last use, as the fourth.
- Recall itself cannot rescue it. Retrieval is driven by recent narration, and a fact nobody mentions is exactly the one that loses recency.
- Pinned memories stay protected (
past_capacity_pinned). - No memory is evicted by the pass that created it. The frozen-bank regression fix still holds in all three scenarios.
G. Ranking results
The independent_default recall turn is at depth 106. Its query is the newest 4
actions, cut to 600 tokens: three narration turns and the one-line question.
| Value | |
|---|---|
F's similarity (semantic score) |
0.241 |
| Lexical score | none: memory ranking has no lexical term |
| Pin effect | none (not pinned) |
| Final rank | 2 of 17 |
memory_top_k cutoff |
5 |
| Selected | yes |
Replica agrees with the turn's stored memories.used |
yes |
The same memory against three stand-alone queries, ConceptEmbedder:
| Query | Rank | Similarity | Selected |
|---|---|---|---|
| direct: "I ask Mara where she hid the amber sundial." | 1 of 17 | 0.708 | yes |
| paraphrase: "… the little brass dial that tells the hour." | 1 of 17 | 0.636 | yes |
| unrelated: "… what rope costs at the landing this season." | 5 of 17 | 0.064 | yes |
Two diagnostic observations. Neither is a failure in this scenario.
- The production query dilutes the question. The question alone scores 0.708;
inside the four-action window it scores 0.241. F survives at rank 2 of 17. In a
bank where more memories share the recent narration's vocabulary, the same
dilution would push it past
top_k. - There is no relevance floor. With more memories than
memory_top_k, five are injected whatever their similarity. An unrelated query still injects F at 0.064. This matters to F only in the other direction: it can ride along even when it is not relevant.
H. Injection results
independent_default injects F: memories.used in the recall turn's stored snapshot
names memory 1, its text is in the used_memories section, and that section is 97
tokens. The same holds in long_block_fact_late and lineage_control.
selected_but_not_injected never occurred: whatever retrieval selected, the
builder rendered.
I. Lineage negative control
lineage_control:
- G is planted on line A at depth 41, turn 21.
- G's memory 7 is written.
- A Save Point is placed on line A after G.
- Undo ×9, then divergent writing at turn 30. The last action before the divergence has id 59.
| Assertion | Result |
|---|---|
| G's memory stays stored | yes (memory 7 present) |
| G is not eligible on the active lineage | yes (the path clause returns nothing) |
| G is never injected after the divergence | yes (no turn with id > 59 names it or carries its text) |
memories.used does not report it after the divergence |
yes |
| Returning to line A (Save Point restore) makes it eligible again | yes (memory 7 eligible) |
| F on the active line is unaffected | injected, isolation holds |
Before the divergence, G's memory was legitimately used on line A (turn 22). The first version of this check counted that as a leak. That was a defect in the diagnostic, and the scan now starts after the divergence.
No lineage code was touched. test_m11_leakage.py, 14 tests, passes unchanged (§P).
J. Authority negative control
test_a_memory_that_contradicts_state_loses_and_changes_nothing sets up the
conflict like this:
- a state correction adds the fact "the tavern lamp is lit";
- a hand-written memory says "The tavern lamp was never lit that night.";
- the memory is pinned, so it is injected;
- a turn is played with an empty proposal.
| Assertion | Result |
|---|---|
| The narrative state document is unchanged by retrieval and the turn | yes (identical before and after) |
The state fact is in the prompt's narrative_state section |
yes |
The memory is in used_memories, under "Memories from earlier in the story …" |
yes, framed as historical and non-canon context |
narrative_state comes after used_memories, so state is read last and settles the conflict |
yes |
F07 semantics are unchanged: memory never writes state.
K. Real-model attempts
Setup.
- Command:
tools/m11_long_run.py --turns 100 --independent-fact. - Host and models: the GPU inference host (Ollama 0.34.0),
qwen2.5:3b-instruct-16kat a verified 16,384 window, embeddings bynomic-embed-text:latest. - Memory: the memory bank and auto-summarise on.
- Harness settings:
memory_top_k4 andcontext_token_budget16,384. - Logging: the owner's power, link and kernel logging was running on the host before the run started.
- Permission: inference was used only with the owner's explicit approval.
Evidence:
$HOME/v11-evidence/b1/real-1/:summary.json,recall-independent.json,timeline.jsonlandcampaign.db;$HOME/v11-evidence/b1/real-1-diagnosis/diagnosis.json: the four stages, recomputed on a copy of the database with the same embedding model.
K.1 Attempt 1 — PRECONDITION FAILED
| Run status | complete: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 active memories (19 eligible on the recall line), 1,308 s elapsed |
| Window / accounting | verified 16,384 on every turn; fits |
| M04 verdict (unchanged) | recovered_through_state_only |
| Independent fact | planted at depth 3 by the player turn "I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one." It was never corrected into state |
| Verdict | precondition_failed:absent_from_summary (the verdict names the first failed precondition in its fixed order) |
| Precondition | Result | First failure |
|---|---|---|
| planted turn outside recent history | held: the history window started at depth 66 | — |
| absent from authoritative state | held: the document, every node's snapshot and the recall prompt's state section | — |
| absent from imported knowledge | held: 3 sources | — |
| absent from later narration | failed: the narrator mentions the fact at depths 6, 8, 10, 14, 16, 18, 20, 22, 24, 26 and later | accepted turn 5 |
| absent from the active summary | failed | accepted turn 9 |
| fact text in the recall prompt's history sections | present (restated narration inside the window) | — |
Why it did not qualify. The 3B narrator took the planted detail up as a motif and restated it for the rest of the campaign ("Mara's amber sundial flickered softly, a silent reminder of their shared history"). The summariser, whose prompt asks it to "preserve important established facts", folded it into the running summary. Neither is a defect in the harness. They are the other layers doing what they do with a salient fact, which is exactly what makes a clean memory-only measurement hard to obtain with a real narrator.
The four stages, diagnosed anyway. They are not evidence for independent retention, but they are evidence for the mechanism.
| Stage | Result |
|---|---|
| created | yes. Memory 1 covers depths 0-5. The block is 688 tokens, so there was no truncation: the fact was in the block and in the summariser's excerpt. The memory text: "Aldric taps the silver key against his chest … Mara slipped the amber sundial into the cracked teapot, her fingers tightening on the silver key. …" |
| retained | yes. Not forgotten, use_count 53, 33 active memories against capacity 80, eviction position 28 |
| ranked | no. Eligible and embedded. For the recall turn's own query it scored similarity 0.866 and ranked 10 of 19, against top_k_cutoff 4. It was not selected and was not suppressed as a duplicate. The replica matched the stored selection |
| injected | not reached. The recall turn's memories.used = [30, 29, 31, 5] |
What the narrator was given instead. Five later memories also carry the fact, all written from narration that restated it: memories 12 (depths 66-71), 28 (81-86), 30 (93-98), 17 (96-101) and 18 (102-107). Memory 30 was injected at recall. The fact reached the narrator through memory, but through a restatement's memory, not the planting-era one.
The same memory 1 against reference queries (nomic-embed-text):
| Query | Rank | Similarity | Selected |
|---|---|---|---|
| the recall turn's production query | 10 / 19 | 0.866 | no |
| paraphrase: "… the little brass dial that tells the hour" | 6 / 19 | 0.593 | no |
| unrelated: "… what rope costs at the landing this season" | 12 / 19 | 0.435 | no |
Two properties of the real embedder matter to B.2:
- A high floor. An unrelated query still scores 0.44.
- A crowded top. The bank is full of near-identical "Aldric and Mara step out, the silver key's weight in his pocket" memories, so 0.866 was not enough to reach the top 4.
A side finding outside B.1's scope. The export's protocol-leak counter flags
1 of 105 stored turns: action 153, depth 143. Mid-reply, the narrator echoed the
length hint and the state reminder, with an Events: [...] line, and then continued
the story. WP-A2's cleanup rules act only at the end of a reply, so an instruction
echo with story after it stays. It is recorded here for the v1.1 backlog. B.1 did not
touch it.
L. First failing stage
Deterministic, with a best-case summariser and a concept embedder, identical on v1.0.0 and HEAD:
| Condition | First failing stage |
|---|---|
| a bank under capacity, blocks under 2,000 tokens | none: created, retained, ranked (2 of 17) and injected |
a bank past memory_bank_capacity |
retention: evicted by least-recently-used order after recent narration stops retrieving it (§F) |
| the fact early in a block over 2,000 tokens | creation: the fact is in the block, but not in the summariser's excerpt (§E) |
Real model, attempt 1 (isolation not met, mechanism only): created yes, retained
yes, ranking failed first. The planting-era memory scored 0.866 but ranked 10th of
19 behind later, near-identical memories, and outside top_k 4.
Across all the evidence, the first stage an early fact fails at a real campaign's length is ranking. In both the deterministic default and the real run, the memory exists and is retained at 100 turns, where the bank is below capacity. What decides whether the narrator is shown it is its rank against a recent-narration query in a bank of similar memories.
- Retention (eviction) is a second, later failure. It is certain once a campaign outgrows capacity: about 480 actions at the defaults.
- Creation (truncation) is a third, conditional one. It needs blocks longer than 2,000 tokens, which the 16,384 window with 500-token replies did not produce (688 tokens).
M. Evidence for likely root cause
- The retrieval query is recent narration, not the question (§B.9, §G). The query is the last 4 actions cut to 600 tokens, so a one-line recall question is outweighed by three turns of prose. The effect is deterministic: the question's own similarity of 0.708 fell to 0.241 in the production query. In the real run, recent prose about the same tavern, key and people made every memory look similar, and ten ranked above the planting-era one.
- Ranking has no term that favours the planting-era record (§B.9; CONTEXT-AND-MEMORY §20). There is cosine only: no lexical match on the question's rare terms ("sundial", "teapot"), no importance, and no preference for the earliest or a coverage-distinct source. Later restatement memories carry the same words in more familiar company, and outrank the original.
- Eviction is purely least-recently-used (§B.7, §F). Retrieval is driven by recent narration, so exactly the facts nothing recent mentions lose recency and go first. Recall itself cannot rescue them, because they are no longer retrieved.
- The summariser reads only the last 2,000 tokens of a block (§B.6, §E). This is proven deterministically. It did not bite at the real run's block sizes.
- v1's M04 never tested memory (§B finding). The clue was planted as state, so the long-standing "memory does not keep the fact" observation was never a measurement of memory.
N. What B.2 is allowed to change
B.2 is allowed only the smallest changes the evidence supports, one mechanism at a time, each with a failing test first. In order of the evidence:
-
Ranking (the first failing stage). Candidates:
- build the retrieval query so the player's newest input is not drowned out, for example by giving the newest player action its own weight or its own query;
- and/or add one inspectable ranking term from CONTEXT-AND-MEMORY §20, most directly a lexical match on the query's rare terms.
Either must keep
replica_matches_stored_selectionmeaningful: a diagnostic-visible score, recorded inmemories.used. -
Eviction (the certain second failure). Stop least-recently-used eviction from discarding a never-again-retrieved early memory first. For example, weight eviction by coverage, keeping the only memory of a story range, or by age, instead of recency alone. The frozen-bank protection must be kept.
-
Creation (conditional). Choose the summariser's excerpt so that a fact early in a long block is not cut. For example, the head and tail, or the whole block up to a larger bound.
Each change turns one of B.1's diagnostics into a passing result:
past_capacityandlong_block_fact_earlyflip their strict xfails;- a real-model re-run shows
ranked: yesfor the planting-era memory.
O. What B.2 must not change
- Lineage safety.
tree.attach_memory, the path clause,forget_nodeand E02 stay as they are. An abandoned line's memory stays stored and ineligible (§I). - Summary lineage (E03) and summary content policy.
- Authority. Memory never writes state and is never framed as canon (F07, §J).
- Imported-knowledge authority and retrieval.
- Pins. Pinned memories stay always-selected and never evicted.
- The frozen-bank fix. A memory is never evicted by the pass that created it.
- The single-commit use counter (M11 §O.7).
- F01-F08, E01-E04, the M04 verdicts, the bundle format and the schema, unless a migration is separately justified.
- The deterministic diagnostic itself. B.2 flips the strict xfails. It does not weaken the scenarios or the isolation checks.
P. Tests / regression
| Run | Result |
|---|---|
test_v11_b1_memory_diagnostic.py on HEAD |
24 passed, 2 xfailed (strict) |
test_v11_b1_memory_diagnostic.py on v1.0.0 (§D) |
24 passed, 2 xfailed (strict), identical |
test_v11_b1_long_run_verdict.py, test_m11_long_run_memory.py, test_m11_long_run_resume.py |
52 passed |
| Full backend suite, HEAD plus the B.1 files | 1,578 passed, 17 skipped, 2 xfailed, 0 failed (1,323 s). The 17 skips are the tests that need a real model, the same 17 as before. The 2 strict xfails are the two diagnosed retention criteria (§D). |
| Full backend suite, final re-run at staging (after the real-model attempt; the staged tree) | 1,578 passed, 17 skipped, 2 xfailed, 0 failed (1,912 s), identical |
No frontend file was changed, so the frontend suite, lint and build are not affected.
Q. Compatibility
| Result | |
|---|---|
| Application code changed | none. git diff --stat HEAD -- backend/app frontend is empty |
| Schema migration | none |
| Bundle format | unchanged |
| Stored campaign behaviour | unchanged |
| Memory behaviour | unchanged. Creation, eviction, ranking, pins and lineage are all as in v1.0.0. memorybank.py is byte-identical to the tag |
| What changed | Diagnostic tooling (tools/memory_diagnostic.py, tools/v11_b1_memory.py), an opt-in harness mode (m11_long_run.py --independent-fact, with every existing behaviour and M04 verdict unchanged when the flag is off) and tests |
| New strict xfails | 2 (§D). They document the two diagnosed defects, and the suite stays green. B.2 must remove them deliberately when it fixes the mechanisms |
R. Security / local-only
Checked against the B.1 diff: the m11_long_run.py changes plus the four new files.
| Result | |
|---|---|
| New network client, endpoint, URL or TLS setting | none. The only address in the new code is http://127.0.0.1:9/v1, a refused loopback port the deterministic scenarios configure so nothing is contacted |
| External embedding service or remote vector store | none. The deterministic runs use ConceptEmbedder in-process. The real-model run uses the configured Ollama embedding model through the application's existing provider |
Endpoint policy (endpoints.py, ADR 011) and TLS (tlstrust.py) |
unchanged; not in the diff |
| Network calls in the real-model run | only the configured Ollama host on the trusted LAN, through the application's own provider and probe paths |
| Real identifiers in committed files | none. Checked again at staging |
| Offline container regression | not run. The harness builds and runs a Docker container, and this session's permission policy refused it. B.1 changes no runtime code and nothing in the image (the image carries backend/app and frontend/dist only), so the container would be byte-identical to the A1/A2 image. That image passed 23/23 on the corrective tree (V1.1-WP-A1-A2-REPORT.md §R.3) |
S. Final decision
What was established. Every B.1 requirement was carried out except the clean real-model run, which was attempted:
- the current mechanism (§B);
- a deterministic, isolated harness with four stage outputs and a verdict at recall depth ≥ 100 (§C);
- the v1.0.0 baseline (§D);
- capacity and eviction (§F), the creation window (§E) and ranking (§G);
- the lineage and authority negative controls (§I, §J);
- the
recovered_through_memory_independentlong-run verdict, with the M04 verdicts unchanged; - tests and regression (§P), compatibility (§Q) and security (§R).
The real-model gap. The one real-model attempt ran cleanly, but failed isolation
(precondition_failed:absent_from_summary). The narrator and summariser restated the
fact, so a clean memory-only result on a real model was not obtained. The owner
decided not to make a second attempt, because the same narrator behaviour would very
likely repeat. The attempt is reported in full, and its stage diagnosis is used as
mechanism evidence only (§K).
B.1 DIAGNOSTIC: COMPLETE
FIRST FAILING STAGE: RANKING. In a real 100-turn campaign the early fact's memory
was created (it carries the fact) and retained (active, 33 of 80), but ranked 10th of
19 (similarity 0.866) against top_k 4. Later memories with near-identical wording
outranked it, under a query made of recent narration (§K, §L). Two further failures
are proven deterministically, identically on v1.0.0:
- retention, past
memory_bank_capacity: least-recently-used eviction removes an early memory first; - creation, for a fact early in a block over 2,000 tokens: the summariser's last-2,000-token excerpt drops it.
WP-B.2 RECOMMENDED CHANGE: retrieval ranking first. Build the retrieval query so
the newest player input is not diluted by three turns of recent narration, and add one
inspectable lexical term for the query's rare words to cosine ranking. The score must
be recorded in memories.used. Acceptance: a real-model re-run shows ranked: yes
for the planting-era memory.
Then, in separate test-first steps:
- make eviction coverage-aware instead of purely least-recently-used, so the only
memory of an early range is not discarded first (flips
test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity); - choose the summariser excerpt so a fact early in a long block is kept (flips
test_acceptance_a_fact_early_in_a_long_block_is_remembered).
Everything in §O stays unchanged. B.2 has not been started.