# v1.1 WP-B.2 — Independent Long-Term Memory Retention **Status:** COMPLETE — accepted by the owner with a documented real-model limitation (2026-09-15, option 1(a)). Staged, uncommitted, for the owner's signed commit. - **Shipped:** B2.1 ranking, B2.2 eviction and B2.3 excerpt creation. - **Not shipped:** B2.4, a memory-prompt experiment. It did not correct the reference model's creation failure, and its prompt was reverted (§T). - **The disposition** is in §S. --- ## A. Repository baseline | | | | --- | --- | | Branch | `v1.1-development` | | HEAD at start | `beb17ad` — *v1.1 WP-B.1: diagnose independent long-term memory retention*, signed by the owner (good signature, RSA key `02C9BF7D…`) | | Its parents | `d63804f` (WP-A1/A2, signed), `ac465ed` (plan v4.1), `432f041` (v1.0.0, tag `v1.0.0`) | | Working tree at start | clean | | Comparison baseline | `432f041` (v1.0.0), from a throwaway worktree. The B.1 diagnostic tool, which is not in that tree, was copied into it from `beb17ad`, with only the new scenario definitions and the coverage trace added. `app/` in the worktree is v1.0.0's, unmodified | --- ## B. B.1 findings being corrected WP-B.1 (`V1.1-WP-B1-REPORT.md` §L, §M) placed three deficiencies. They are the only memory behaviour this package changes. | # | Stage | B.1 evidence | Corrected in | | --- | --- | --- | --- | | 1 | **Ranking** | Real 100-turn run: the planting-era memory was created, active and carried F, and ranked **10th of 19** (cosine 0.866) against `memory_top_k` 4. The query was the newest four actions cut to 600 tokens, so the one-line question was diluted by narration (deterministically, 0.708 alone fell to 0.241 in the query). That run failed isolation, so it was diagnostic evidence only | B2.1 | | 2 | **Retention** | Past `memory_bank_capacity`, least-recently-used eviction removed the planting-era memory first (turns 21, 21, 36 in the capacity fixtures). Strict xfail `test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity` | B2.2 | | 3 | **Creation** | The summariser saw only a block's last 2,000 tokens; a fact at the start of a 2,079-token block never reached it (`not_created`). Strict xfail `test_acceptance_a_fact_early_in_a_long_block_is_remembered` | B2.3 | A fourth B.1 observation, a mid-reply protocol echo, is **not** WP-B's and is left alone (§N). --- ## C. B2.1 ranking design ### C.1 The retrieval query v1.0.0 embedded one text: the newest four story actions, cut to their last 600 tokens. It is replaced by two short texts embedded in **one** call (`memorybank.retrieval_query`): | Component | Source | Bound | | --- | --- | --- | | `input` | the newest action, when it is a player action with text (`do`, `say`, `story`) | last 200 tokens (`QUERY_INPUT_TOKENS`) | | `context` | the scene from the authoritative state — `scene.summary`, the location's entity name, the names of `scene.present` (at most 8) — then the end of the newest narration | 60 tokens (`QUERY_SCENE_TOKENS`) + 120 tokens (`QUERY_NARRATION_TOKENS`) | - **A continue turn**, and the Insights dry run, have no player input: the context alone is searched, and the lexical score is 0. - **A retry** excludes the attempt being replaced, as before; the query is the retried input and the narration before it. - **Why the context stays.** `CONTEXT-AND-MEMORY.md` §18 requires more than raw input, and a question often means nothing without its scene ("I ask her what she keeps up there"). It is bounded so it can resolve a reference but cannot outweigh the question by length. The full entity list, threads and older narration are left out on purpose. - **Provenance.** The query is recorded in each turn's snapshot as `memories.query`: `input`, `context`, `input_terms`, `input_weight`, `lexical_weight`. ### C.2 The ranking formula ```text semantic_score = INPUT_WEIGHT * cos(input, m) + (1 - INPUT_WEIGHT) * cos(context, m) INPUT_WEIGHT = 0.6; either cosine alone if the other text is empty lexical_score = Σ w(t) for t in input_terms ∩ terms(m) / Σ w(t) for t in input_terms w(t) = ln((N + 1) / (df(t) + 1)) N eligible candidates, df holding t final_score = semantic_score + LEXICAL_WEIGHT * lexical_score, LEXICAL_WEIGHT = 0.15 ``` - **Terms** (`memorybank.lexical_terms`): imported knowledge's tokenizer and stop list (`knowledge.fts.terms`), a possessive `'s` dropped, and a trailing `s` folded from words of five letters or more not ending in `ss`. No stemmer, no dependency. - **Rarity** is computed per call over the eligible candidate set only. Nothing is indexed or stored. A word every candidate holds weighs exactly 0; a word none holds weighs most, so the unmatched part of a question dilutes every candidate equally. - **Range.** `lexical_score` is in [0, 1], so wording can move a memory by at most 0.15 (`test_a_lexical_match_moves_a_memory_at_most_the_lexical_weight`). - **Only the player's input** is matched lexically, never the context. - **Pins** are exactly as before: always used, counted toward `memory_top_k`. - **Ties** on the final score are broken by memory id, so row order never decides. - **Redundancy suppression** (§22, cosine ≥ 0.93, never across authority) is unchanged and runs over the final order. - **Reported per used memory:** `semantic_score`, `lexical_score`, `final_score`. `similarity` keeps its v1.0.0 meaning, the semantic score, so the inspector's "closeness" needs no frontend change. **Egress.** Memory text is read for ranking only on a turn whose input has terms, only for memories not already held, and held beside the vectors in a cache with the same two rules: `set_vector` drops an entry (an edit clears the vector, so a changed text is re-read), and the next read discards memories no longer in the catalogue. A continue turn reads no memory text beyond the top-k detail read, and a second turn in a row reads none again (`test_memory_text_is_read_once_and_then_held`). ### C.3 Choosing the lexical weight Swept over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5 on three deterministic fixtures, each ranked two ways: with the new query, and with the **v1.0.0 query** plus the lexical term (the lexical term's own contribution, with the crowding left in). Rank of the planting-era memory F, `memory_top_k` 4: | Weight | 0 | 0.05 | 0.1 | **0.15** | 0.2 | 0.3 | 0.5 | | --- | --- | --- | --- | --- | --- | --- | --- | | `ranking_crowded`, new query | 1 | 1 | 1 | **1** | 1 | 1 | 1 | | `ranking_crowded`, v1.0.0 query + lexical | 7 | 6 | 6 | **1** | 1 | 1 | 1 | | `ranking_crowded`, unrelated question: F selected? | no (9) | no (9) | no (6) | **no (6)** | no (6) | no (5) | **yes (3)** | | negative control (paraphrase + a decoy's rare place): decoy above F? | no | no | no | **no** | no | no | no | - **0.15** is the smallest weight at which the lexical term alone rescues F from the crowded v1.0.0 query. - **0.5** already lets an incidental shared word ("landing") select F for an unrelated question, which is the failure a larger weight buys. - The **negative control** never put the decoy above F at any weight with the new query. - **Not tuned to one phrase:** the direct question, a paraphrase sharing only "Mara", a context-dependent question sharing no word at all, a decoy word from a different memory, and an unrelated question all ran at every weight. **One normalisation was rejected along the way.** Taking the share over only the input words some candidate holds let a lone incidental match score 1.0, and in the context fixture a decoy holding "fish stalls" then outranked F from weight 0.15. The share is now over all the input's words. --- ## D. B2.1 before/after evidence **Fixtures** (`tools/memory_diagnostic.py`, best-case summariser, concept embedder, isolation valid, below capacity, recall at depth 106): - `ranking_crowded`: 150-word narration, `memory_top_k` 4 — B.1's real-run failure made deterministic. - `ranking_context_dependent`: the same, with the last narration putting Mara at the tavern's top shelf by an old kettle, and the question "I ask her what she keeps up there." Neither "kettle" nor "shelf" is one of F's leak terms. | Fixture | v1.0.0 (`432f041`) | B2.1 | | --- | --- | --- | | `ranking_crowded` | `retained_but_not_ranked`: rank **7** of 17, similarity 0.1653 | `injected`: rank **1** of 17; semantic 0.4671, lexical 0.5328, final 0.5470; selected [1, 7, 11, 12]; 87 tokens | | `ranking_context_dependent` | `retained_but_not_ranked`: rank **6** of 17, similarity 0.1577 | `injected`: rank **2** of 17; semantic 0.1866, lexical **0.0**, final 0.1866; selected [1, 5, 6, 9]; 89 tokens | | `independent_default` (top-k 5) | `injected`: rank 2, similarity 0.241 | `injected`: rank 1; semantic 0.4739, lexical 0.5328, final 0.5538 | In every row the diagnostic's replica of the ranking (production's own `score_candidates` and `select_memories`) matched the recall turn's stored selection. **The required ranking controls:** | Control | Result | | --- | --- | | Direct question | F rank 1 in every fixture (above) | | Semantic paraphrase ("the little brass dial that tells the hour") | rank 1, lexical 0.1125 (only "Mara" shared) against 0.5328 for the direct question: found by meaning | | Context-dependent wording | rank 2 with the scene; **rank 17 of 17** with the context removed; lexical 0 | | Negative control (paraphrase + "out by the fish stalls", held only by a filler memory) | the decoy's lexical score (0.18) exceeds F's (0.09), and F still ranks 1, the decoy 2-5 | | Common words ("The travellers spent time.") | lexical 0.0 for every memory | | Pins | always used, counted toward top-k (`test_a_pin_is_still_always_used_and_counts_toward_top_k`, and v1.0.0's `test_pinned_memories_are_always_used`, unchanged) | | No relevance floor (unchanged) | an unrelated question still selects a full top-k; F is no longer carried along by it (rank 6 of 17, not selected) | **Tests at the checkpoint:** `test_v11_b2_memory_ranking.py` 18 passed; `test_v11_b1_memory_diagnostic.py` 31 passed and 2 xfailed (the eviction and creation criteria, not yet fixed); the memory regression group (13 files) 226 passed. **B2.1 RANKING: PASS** --- ## E. B2.2 eviction design `memorybank.eviction_order(rows, limit)`, run by `_evict_over_capacity` after each post-turn pass. It is a pure function of seven columns per active memory and reads neither text nor vectors (`test_eviction_reads_no_vectors`, unchanged). | Part | Rule | | --- | --- | | **Coverage signal** | Memories with a source range are ordered by position. A memory is judged by the hole its removal would leave: the start of the next memory, less the furthest end before it, less one. Smallest hole first, so the bank thins where it is densest. A memory whose `source_start` another active memory shares leaves no hole (0) | | **Boundaries** | The earliest and the latest memory by position are not coverage candidates: they alone describe the opening and the newest stretch | | **Recency signal** | Among equal holes: `coalesce(last_used_at, created_at)` oldest first | | **Tie-breaking** | then `use_count` lowest first, then `id` lowest first | | **Pinned rows** | never taken; they count as coverage for their neighbours; if every active memory is pinned the bank may exceed capacity, as before | | **Fallback** | once no coverage candidate remains — memories typed by the player or migrated have no range, and a bank can be all boundaries — the rest are taken by the v1.0.0 order (recency, use count) plus id | | **Recomputation** | after every pick, since removing a memory widens its neighbours' holes | - **Unchanged:** the capacity setting and default (80), the count over the whole adventure (not one lineage), marking `forgotten` rather than deleting, running once per post-turn pass. - **Frozen bank.** A memory written this turn is the latest boundary, so it is not a coverage candidate, and in the fallback it has the newest timestamp. It can still go first legitimately when it only repeats a stretch another, more recently used memory describes (`test_a_newborn_can_be_the_legitimate_first_to_go`). - **Lineage.** The rule decides `forgotten` and nothing else; eligibility is the retrieval clause, untouched. A sibling line's memory at the same depth shares a start, so it is the first kind thinned, and the least recently used of the two goes — after a divergence, that is the abandoned line's. - **No new field.** Existing `source_start`, `source_end`, `last_used_at`, `created_at` and `use_count` only. --- ## F. B2.2 capacity evidence **Deterministic fixtures** (v1.0.0 from the worktree, with a coverage trace added to its copy of the tool): | Fixture | v1.0.0 | B2.2 | | --- | --- | --- | | `past_capacity` (capacity 6, top-k 5, 17 memories written) | `created_but_evicted`: F evicted at turn 21, the first eviction. Final bank covers depths 36-101 | `injected`: F retained, rank 1 of 6. Bank covers 0-101, largest uncovered stretch 18 | | `past_capacity_pinned` | evicted at turn 21; bank 6-101 | `injected`; the pinned memory kept; 0-101, gap 18 | | `past_capacity_low_top_k` (capacity 8, top-k 2) | evicted at turn 36 (first eviction 27); bank 18-101 | `injected`: rank 1 of 8; 0-101, gap 12 | In all three: active memories never exceeded capacity after a pass, and no memory was evicted by the pass that created it. **F survives these fixtures as the earliest boundary.** That is the general rule, not a special case — any campaign's opening memory is kept while interior memories remain — but it is also the easiest case, so the rule was tested on interior memories in general. Memories arrive a block at a time, nothing is ever retrieved, and the bank is held at capacity. Largest uncovered stretch, counting from depth 0: | Bank | Coverage-first: first memory kept / largest gap | v1.0.0 order: first memory kept | Average spacing (span / capacity) | | --- | --- | --- | --- | | 17 blocks, capacity 6 | 0 / 18 | depth 66 | 17.0 | | 60 blocks, capacity 10 | 0 / 42 | depth 300 | 36.0 | | 500 blocks, capacity 80 | 0 / 42 | depth 2,520 | 37.5 | | 500 irregular blocks (4-9 actions), capacity 80 | 0 / 47 | depth 2,591 | 38.4 | The largest gap stays within about 1.3 times the average spacing. The v1.0.0 order keeps an unbroken run of the newest blocks and nothing before it. `test_a_long_bank_keeps_describing_the_whole_story` asserts the opening kept and a gap of at most twice the average spacing. **Invariants tested** (`test_v11_b2_memory_eviction.py`, 24 tests): | Invariant | Test | | --- | --- | | Active count ≤ capacity after maintenance (or = the pins, if more) | `test_capacity_holds_and_pins_survive_for_any_bank` (8 random banks) | | Pinned never evicted | same, and `test_pins_are_never_taken_but_still_count_as_coverage` | | Abandoned-line memories not made eligible; lineage separate | `test_eviction_does_not_make_an_abandoned_lines_memory_eligible`, `test_the_pass_changes_nothing_but_forgotten` | | Frozen bank still works | `test_the_newest_memory_is_not_evicted_by_the_bank_it_joins`; v1.0.0's `test_a_newborn_is_not_evicted_by_the_bank_it_joins` and `test_a_full_bank_still_turns_over`, unchanged | | A newborn may still go first when legitimate | `test_a_newborn_can_be_the_legitimate_first_to_go` | | Independent of row order | `test_the_order_does_not_depend_on_row_order`, `test_a_tie_on_every_signal_is_broken_by_id` | | Fallback is v1.0.0's order | `test_memories_without_a_range_take_the_least_recently_used_fallback`; v1.0.0's `test_eviction_marks_the_least_recently_used` and `test_eviction_breaks_ties_on_use_count`, unchanged | The B.1 strict xfail for capacity is now an ordinary test that passes for all three capacity fixtures. B.1's two tests that asserted F's eviction (a diagnosis of v1.0.0) now assert its retention. **B2.2 EVICTION: PASS** --- ## G. B2.3 creation design `memorybank.memory_excerpt(raw, budget=MEMORY_EXCERPT_TOKENS)`, used by `summarize_block` (and so also by the opt-in `tools/rewrite_memories.py`). | Block | Summariser input | | --- | --- | | ≤ 2,000 tokens | the whole block, unchanged. The prompt is byte-identical to v1.0.0's (`test_a_short_block_is_prompted_exactly_as_before`) | | > 2,000 tokens | the block's first tokens, then `\n\n[… the middle of this stretch of story is left out here …]\n\n`, then its last tokens | - **Accounting** (`excerpt_split`). The marker with its blank lines costs 15 tokens and is paid first. The remaining 1,985 are halved, the odd token to the end: **992 + 993 + 15 = 2,000**. - **Seams.** Two decoded runs rejoined can tokenise differently where they meet, so the result is measured and the opening gives up tokens until the whole is within 2,000. Measured alone, a cut part can come to one token over its run; the whole never exceeds the budget, including for mixed-script and emoji text. - **Order and marker.** The opening comes first, and the marker tells the summariser the two parts are not adjacent. The header (`Story excerpt:`) is unchanged. - **The marker is never stored.** `summarize_block` removes it from the model's reply. - **The budget did not grow.** No model context, `MEMORY_MAX_WORDS` or prompt instruction changed. Existing memories were not rewritten. - **Limit.** A fact in the middle of a very long block is still left out (`test_a_fact_in_the_middle_of_a_very_long_block_is_still_omitted`). --- ## H. B2.3 long-block evidence | Fixture | v1.0.0 | B2.3 | | --- | --- | --- | | `long_block_fact_early` (F at depth 1 of a 2,079-token block) | fact in block: yes; in the summariser's excerpt: **no**; `not_created` | in excerpt: **yes**; memory 1 (depths 0-5) carries F; `injected` | | `long_block_fact_late` (F at depth 5 of the same-sized block) | in excerpt: yes; created | in excerpt: yes; created; `injected` | These two fixtures test creation. They recall at depth 22, so the planting turn is still inside the history window and their isolation check is not expected to pass; independence is §I's. **Tests** (`test_v11_b2_memory_excerpt.py`, 17): | Required | Test | | --- | --- | | F early in a > 2,000-token block | `test_a_long_blocks_memory_keeps_its_source_provenance` (the memory holds it); B.1's `test_acceptance_a_fact_early_in_a_long_block_is_remembered`, now ordinary | | F late in the same-sized block | B.1's `test_the_same_fact_late_in_the_same_sized_block_does` | | A block ≤ 2,000 tokens unchanged | `test_a_block_that_fits_is_sent_whole_and_unchanged`, `test_a_block_of_exactly_the_budget_is_unchanged`, `test_a_short_block_is_prompted_exactly_as_before` | | Head/tail accounting bounded | `test_the_split_is_even_and_documented`, `test_the_excerpt_never_exceeds_the_budget` (5 sizes up to 10× the budget), `test_the_budget_holds_for_awkward_text` | | The marker is not stored as story fact | `test_the_marker_is_never_stored_as_part_of_a_memory` (a summariser that echoes its whole input) | | Valid source provenance | `test_a_long_blocks_memory_keeps_its_source_provenance`: `source_start`/`source_end` are the block's depths, `branch_id`/`depth` its last node's | The B.1 strict xfail for creation is now an ordinary test. **B2.3 CREATION: PASS** --- ## I. Full deterministic independent-memory test `independent_full`: all three B.1 failures at once. Every narrator reply is about 850 words, so every memory block (2,079 tokens) is longer than the excerpt; narration crowds the query at `memory_top_k` 4; capacity is 8, and 17 memories are written. F is planted at depth 1 by a `story` action; recall is at depth 106. `context_token_budget` 16,384. | | v1.0.0 | B.2 | | --- | --- | --- | | Verdict | **`not_created`** (the fact never reached the summariser) | **`injected`** | **Isolation, asserted on every one of the 52 turns and at recall:** | Check | Result | | --- | --- | | F absent from authoritative state | every turn's document; every node's `narrative_state_after`; the recall prompt's state section | | F absent from the summary | the active summary on every turn, and the recall prompt's summary section | | F absent from imported knowledge | none imported; no knowledge section mentions it | | F absent from recent history | the history window at recall starts at depth 84; planted at 1; the fact is not in the history sections | | F absent from later narration | no narrator turn after the planting block (ends depth 5) mentions it | **The four stages:** | Stage | Result | | --- | --- | | created | memory 1, source depths 0-5, text "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." The block was 2,079 tokens and the fact was in the head+tail excerpt | | retained | active; bank 8 of 17 written, capacity 8; the bank covers depths 0-101 with a largest uncovered stretch of 12; nothing evicted by its own pass | | ranked | rank **1** of 8, top-k 4; semantic 0.4671, lexical 0.5066, final 0.5430; selected [1, 5, 7, 12]; replica matches the stored selection | | injected | in the recall turn's `memories.used` and its `used_memories` section, 88 tokens | **Provenance.** The recall snapshot's entry for memory 1 records source `{branch_id: 1, depth: 5, source_start: 0, source_end: 5}` and authority `accepted_story`. It equals the row, the range includes the planting depth, and `memorybank.source_block` on that row returns depths 0-5, which hold the planting action. **Long-run verdict.** The same measurements, given to `m11_long_run._independent_memory_verdict`, return **`recovered_through_memory_independent`** (`test_acceptance_full_is_the_long_run_independent_memory_verdict`). All nine deterministic scenarios, final tree: | Scenario | Verdict | Isolation | F rank / top-k | Notes | | --- | --- | --- | --- | --- | | `independent_default` | injected | ok | 1 / 5 | | | `past_capacity` | injected | ok | 1 / 5 | capacity 6 | | `past_capacity_pinned` | injected | ok | 1 / 5 | pin kept | | `past_capacity_low_top_k` | injected | ok | 1 / 2 | capacity 8 | | `long_block_fact_early` | injected | not expected (recall depth 22) | 1 / 5 | fact in excerpt | | `long_block_fact_late` | injected | not expected (recall depth 22) | 1 / 5 | | | `lineage_control` | injected | ok | 1 / 5 | §J | | `ranking_crowded` | injected | ok | 1 / 4 | | | `ranking_context_dependent` | injected | ok | 2 / 4 | lexical 0 | | `independent_full` | injected | ok, every turn | 1 / 4 | all three at once | --- ## J. Lineage `lineage_control` (B.1's fixture, rerun on the final tree). Fact G is planted at turn 21 behind a Save Point; a second Save Point marks line A at turn 30; nine Undos and a different action abandon line A. | Requirement | Result | | --- | --- | | The abandoned memory stays stored | G's memory (id 7) is on disk | | Absent on the active divergent line | not eligible under the active lineage clause | | Absent from `memories.used` | no turn after the divergence names it, and G's text is in no memories section | | Eligible again only when branch semantics permit | restoring the line-A Save Point makes it eligible again | | F on the active line | still `injected` | Eviction cannot change eligibility: it writes only `forgotten` (`test_the_pass_changes_nothing_but_forgotten`, `test_eviction_does_not_make_an_abandoned_lines_memory_eligible`). `tree.attach_memory`, `forget_node`, the lineage clause, E02 and summary lineage (E03) are not in the diff. `test_m11_leakage.py` passes unchanged (§Q). --- ## K. Authority B.1's authority control (`test_a_memory_that_contradicts_state_loses_and_changes_nothing`), rerun unchanged on the final tree: a state correction establishes "the tavern lamp is lit"; a pinned memory says "The tavern lamp was never lit that night." | Requirement | Result | | --- | --- | | State wins | the fact is in the narrative state section, which is read after the memories section | | Memory remains non-authoritative | it is rendered under "Memories from earlier in the story…", as before | | No state mutation from retrieval | the state document is identical before and after the turn | Retrieval still only reads, and a used memory's `authority` is still recorded. F07 and `classify_authority` are not in the diff. --- ## L. Real-model attempts **Setup, identical for both attempts.** - **Command:** `tools/m11_long_run.py --turns 100 --independent-fact`, unchanged since B.1: the same planted fact, prompts and harness. - **Host and models:** the GPU inference host (Ollama 0.34.0); `qwen2.5:3b-instruct-16k` narrating and summarising, `nomic-embed-text:latest` embedding. - **Memory:** bank and summaries on; harness settings `memory_top_k` 4 and `context_token_budget` 16,384. - **Order of events:** - the full deterministic suite had passed (§Q); - the owner confirmed the power, link and kernel logging was running on the host before attempt 1; - attempt 2 was run because attempt 1 failed isolation. - **Cap:** two deliberate attempts at most, as the brief set. Stages come from `tools/v11_b1_memory.py diagnose` on a copy of each run's database, ranked with the same embedding model. ### L.1 Attempt 1 — PRECONDITION FAILED Evidence: `$HOME/v11-evidence/b2/real-1/`, `$HOME/v11-evidence/b2/real-1-diagnosis/`. | | | | --- | --- | | Run | `complete`: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 memories written (18 eligible at recall), 1,352 s, median 10.9 s a turn | | A1 accounting | **`fits` on all 102 turns**; window verified at 16,384 on every turn; largest prompt 14,997 tokens (server count); smallest observed margin 887 tokens | | A2 leak counter | **0 of 105** stored narrator replies | | M04 (unchanged) | `recovered_through_state_only` | | **Independent verdict** | **`precondition_failed:absent_from_summary`** | | Precondition | Result | First failure | | --- | --- | --- | | planting turn outside recent history | held: the history window started at depth 72 | — | | absent from authoritative state | held: the document, every snapshot, the recall state section | — | | absent from imported knowledge | held (3 sources) | — | | **absent from later narration** | **failed**: the narrator names the fact at depths 10, 14, 16, 18, 20, 22, 24, 26, 38, 44 and later | accepted turn **5** | | **absent from the active summary** | **failed** | accepted turn **26** | The history window at recall held narration restating it too. The narrator took up the planted detail as a motif, exactly as in B.1, and the summariser folded it in. The run is **neither a pass nor a failure** for independent retention. **Stages** (mechanism evidence only, not acceptance evidence): | Stage | Result | B.1 attempt, same harness, v1.0.0 memory | | --- | --- | --- | | created | **yes**: memory 1, depths 0-5; 493-token block, fact in the excerpt | yes (memory 1, 688-token block) | | retained | **yes**: active, used 31 times; 33 active of capacity 80 | yes | | ranked | **yes, rank 1 of 19**: semantic 0.7754, lexical 0.2194, final 0.8083; top-k 4 | **no**: rank 10 of 19, cosine 0.866 | | injected | **yes**: memory 1 is in the recall turn's `memories.used` [1, 10, 28, 5] and its text is in the memories section | no (a later restatement's memory carried it) | The recall turn stored the query the diagnostic recomputes: - **Input:** "> I ask Mara quietly where she hid the amber sundial." - **Context:** the scene (Aldric and Mara in the Crooked Lantern) and the end of the last narration. - **Scores:** memory 1 was also first for the paraphrase (0.7509) and 9th, not selected, for the unrelated question (0.6423). - **Duplicates:** five later memories restating the fact (7, 9, 11, 26, 29) were suppressed as duplicates of memory 1, so the planting-era statement is the one whose provenance survived. **One discrepancy, explained.** The replica selected [1, 10, 28, 33] against the stored [1, 10, 28, 5]. The queries are identical. Memory 33 (depths 111-116) was written by the previous turn's post-turn pass at 02:22:27, after the recall turn's retrieval and 9 s before that turn was saved. The turn therefore considered 18 memories, and the diagnosis, run on the finished database, ranks 19. The rows common to both match to rounding (memory 1: 0.7754 / 0.219 / 0.8082 stored). This is the bank changing between retrieval and diagnosis, not a scoring difference. **Found in passing** (not a defect of the result): when the state's scene summary already ends in a full stop, the scene text reads "rain outside..". Only the embedder sees it (§R). ### L.2 Attempt 2 — ISOLATION VALID, NOT RECOVERED (creation) Evidence: `$HOME/v11-evidence/b2/real-2/`, `$HOME/v11-evidence/b2/real-2-diagnosis/`. | | | | --- | --- | | Run | `complete`: 102 accepted turns, 3 restarts (4 process starts), 12 summaries, 18 eligible memories at recall, 1,284 s, median 10.4 s a turn | | A1 accounting | **`fits` on all 102 turns**; window verified at 16,384 on every turn; largest prompt 14,992 tokens; smallest observed margin 892 tokens | | A2 leak counter | **0 of 105** stored narrator replies | | M04 (unchanged) | `recovered_through_state_only` | | **Independent verdict** | **`not_recovered:not_created`** | **Every independent-memory precondition held, on every turn:** | Precondition | Result | | --- | --- | | planting turn outside recent history | held: the window at recall started at depth 72; planted at 3; no fact text in the history sections | | absent from authoritative state | held: the document, every node's snapshot, the recall state section | | absent from the active summary | held, on every turn the harness checked and at recall | | absent from imported knowledge | held (3 sources) | | absent from later narration | held: no narrator turn after depth 5 names the fact | This is the first real run, in B.1 or B.2, in which memory was the only layer that could have carried the fact. It is therefore **acceptance evidence**, and it did not recover the fact. **Stages:** | Stage | Result | | --- | --- | | **created** | **no.** Memory 1 covers the planting block, depths 0-5, 623 tokens. The block fitted the budget, so the summariser was sent all of it (`memory_excerpt` returned the block unchanged), and the fact was in it at depth 3 as the player's own action: "> I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one." **The memory the model wrote does not name the sundial or the teapot.** No other memory in the bank does either | | retained | memory 1 active (used 62 times) — but it does not carry F | | ranked | memory 1 ranked into the top 4 at recall (4th: semantic 0.6934, lexical 0.1463 from "Mara", final 0.7153) — but it does not carry F | | injected | memory 1 was injected (`memories.used` [30, 32, 9, 1]) — without F. The recall reply does not name the fact | **What the summariser wrote instead** (memory 1, verbatim, 102 words against the prompt's 50-word ceiling): > Aldric sits in the Crooked Lantern with a silver key in his pocket, no-one to > give it to. Mara observed him quietly, her eyes unreadable. Edrin finished his > ale and stood, asking if they should see what the crypt beneath the Old Abbey > holds. Aldric decided to go, promising not to tell anyone. They walked to the > Old Abbey's grounds, the crypt door ajar, inviting and foreboding. Inside, the > air was musty and cold, with the smell of damp and decay. They advanced > cautiously, each aware of potential dangers. The silver key, the key to the > crypt, felt heavy in Aldric's hands. Three things are visible in it: 1. **The fact was dropped, but a fragment of its sentence survived:** "promising not to tell anyone", now given to Aldric rather than asked by Mara. *Correction (B2.4 review, §T.3):* in the source Aldric is the one who promises; Mara asks him to. The memory's fault is what the promise is attached to. It ties the promise to Aldric's decision to go to the crypt, and drops Mara and the object it was about. The phrase may also echo Edrin's own "I'll tell no one what we found" in the depth-4 reply. 2. **The block's narration was followed and the player's line was not.** The narrator's reply at depth 4 ignored the sundial and moved to the crypt, and the memory follows the narration: walking to the abbey, the crypt door, entering the crypt. *Correction (B2.4 review, §T.3):* an earlier version of this item said those events were past the block. They are not; the depth-4 reply narrates them. The memory invented nothing. It chose the narration over the player's line. 3. **The length ceiling was ignored**, which the prompt states but nothing enforces (B.1 §B.4). **First failing stage in the qualifying run: creation**, caused by what the model chose to write, not by the excerpt (the whole block was sent), eviction or ranking. ### L.3 Both attempts | Attempt | Preconditions | created | retained | ranked | injected | Verdict | | --- | --- | --- | --- | --- | --- | --- | | 1 | **failed** (later narration from turn 5, summary from turn 26) | yes (memory 1) | yes | yes, rank 1 of 19 | yes | `precondition_failed:absent_from_summary` — neither pass nor fail | | 2 | **all held** | **no** (memory 1 omits F) | (memory 1 yes) | (memory 1 4th of the 4 used) | (memory 1 yes, without F) | **`not_recovered:not_created`** — a failure | The two-attempt limit is reached, and no third attempt was made. The planted fact, prompts and harness were not changed between attempts. Inference use ended with the attempt-2 diagnosis. --- ## M. Context-budget / A1 interaction - **The memories section is unchanged in shape and pricing.** It is the same live section, built from the same `memory_top_k` rows with the same line format, and priced into protected context as before (F03). B.2 changes which memories fill it, not how many or how they are rendered. Deterministically it held 67-99 tokens across the scenarios. - **The query costs no prompt tokens.** It is embedded, never sent to the narrator. It is recorded in the snapshot as provenance only. - **A1 is untouched.** The safety reserve, the transport cost, the accounting states and the cold-model load are not in the diff. The summariser's input is still bounded at 2,000 tokens of story plus the cast brief, so its request is no larger than before. - **Embedding requests.** One call per turn, as before; it now carries two short texts (at most about 200 and 180 tokens) instead of one of up to 600. - **On the real model**, every counted turn of both attempts was `fits`, with the window verified at 16,384 and prompts of at most 14,997 tokens (smallest observed margin 887). The memories section peaked at **716 tokens** (attempt 1) and **1,246 tokens** (attempt 2), against 67-99 deterministically. The difference is the real summariser's memory length: attempt 2's planting memory alone was 102 words, twice the prompt's ceiling (§R 12). It stayed inside the section's budget and the prompt inside the window. --- ## N. Protocol-leak / A2 observations - **WP-B does not modify A2.** `narrative/extract.py`, the extractor rules and the protocol cleanup are not in the diff. - **The B.1 observation stands, unchanged.** In B.1's real run, one stored narrator reply of 105 (action 153, depth 143) echoed application instructions **mid-reply**, with story after them, outside A2's trailing cleanup region. The evidence is preserved in `$HOME/v11-evidence/b1/real-1/campaign.db` (`V1.1-WP-B1-REPORT.md` §K). It is a separate v1.1 follow-up item, not WP-B's. - **This report does not claim the leak count is zero.** The counts from B.2's real-model attempts are in §L. The integrated v1.1 release gate still needs its own protocol-leak result. --- ## O. Compatibility | Area | Result | | --- | --- | | Schema migration | **none**; no model or migration file in the diff | | Bundle format | **v3, unchanged**; no bundle code in the diff | | Existing memory rows | **unchanged** on open: the same 33 rows, row for row | | Existing summaries | **unchanged**: 12 rows, row for row | | Existing history | **unchanged**: 207 actions, 4 branches, 32 state events, 2 Save Points, row for row | | Existing memories regenerated | no. New retrieval and eviction apply only when v1.1 next plays the campaign | | Settings | `memory_top_k`, `memory_bank_capacity` and their defaults unchanged | **A real v1.0.0 database.** `tools/v11_compat_check.py` on `m04-final/campaign.db` (written by `96c1bf5`, product code identical to v1.0.0; `user_version` 94; SHA-256 `c74e8798…4f18`, unchanged by the runs). Each tree opened its own copy: | Comparison | Result | | --- | --- | | Schema before and after opening, v1.0.0 and B.2 trees | identical; `user_version` 94 → 94 on both | | API snapshot (the export bundle, narrative state and events, Save Points, knowledge, memories, derived status, settings) | **0 differences** | | Tables after opening (memories, summaries, actions, state events, Save Points, branches, knowledge sources, adventures) | identical, row for row | `--exercise` on a third copy, endpoint refused: Undo 119 → 117; Redo back to 119; restore *On the ridge* → 69 with Redo available; context dry run HTTP 200 with canon, state and the applicable summary; export and import 207 → 207 actions, the same head, Save Points 2 → 2, memories 33 → 33, identical state. The dry run's memory retrieval reports the unreachable embedding endpoint and uses no memories, as before. Evidence: `$HOME/v11-evidence/b2/compat/`. --- ## P. Offline / security ### P.1 Offline and packaging regression `tools/m11_offline.py --out $HOME/v11-evidence/b2/offline`, on the complete B.2 tree: - the image was built with `--no-cache`; - the container was started with `--network none`: a loopback interface, no resolver, no route; - the exercise was driven inside it over `docker exec`. **23 passed, 0 failed:** - no route to the public Internet; - no external DNS; - first page load offline, and the page names no remote origin; - a CSP is served, and every shell asset is served locally; - a campaign is created, and state extraction works; - a local file imports, and the prompt assembles; - imported knowledge is searchable; - a turn with no model reachable is reported as a failure, with no narration accepted, the player's words kept, state unchanged and the earlier story intact; - export and import work offline, with state, and no secret in the export; - the media module imports, a scene packet builds, and no media provider is required; - campaigns survive a container restart. Evidence: `offline-report.json`, `docker-build.log`, `container.log` and `inside-stdout.txt` in that folder. As in M11, inference is not exercised offline: the model host is on the trusted LAN, which a container with no network cannot reach. ### P.2 What the B.2 diff adds | Check | Result | | --- | --- | | New network client, URL, host, IP, socket or subprocess | **none**: no added line in `backend/app` or `backend/tools` names one | | External vector database | **none**. Vectors stay in SQLite `embedding_blob`; the lexical term is computed in process per turn, with no index | | Cloud embedding | **none**. Embeddings go through the existing `embedding_provider` to the configured local endpoint, one call per turn as before | | New dependency | **none**. The imports added to product code are `math` (standard library) and three internal modules: `context.builder._encoding` (the vendored tokenizer), `knowledge.fts` (its tokenizer and stop list) and `narrative.model` (`entity_name`) | | Endpoint policy (`endpoints.py`, ADR 011), TLS trust (`tlstrust.py`), providers | **unchanged**; not in the diff | | Frontend | **unchanged**; not in the diff | | New data leaving the machine | none. The embedding request carries the player's input and a short scene text instead of 600 tokens of narration, to the same local endpoint | | Real identifiers in committed files | none (scanned at staging, §Q) | Local-only operation is unchanged. --- ## Q. Full regression | Run | Result | | --- | --- | | **Full backend suite**, complete B.2 tree | **1,652 passed, 17 skipped, 0 failed, 0 xfailed** (1,603 s). The 17 skips are the tests that need a real model, the same 17 as in B.1. B.1's 2 strict xfails are now ordinary passes | | `test_v11_b2_memory_ranking.py` (new) | 18 passed | | `test_v11_b2_memory_eviction.py` (new) | 24 passed | | `test_v11_b2_memory_excerpt.py` (new) | 17 passed | | `test_v11_b1_memory_diagnostic.py` (B.1, updated for B.2) | all pass; no xfail marker remains | | Memory regression group at the B2.2 checkpoint: `test_memory_retrieval`, `test_memory_nodes`, `test_context_memory` (F01-F08, E01-E04, write-lock and derived-failure recording), `test_m11_leakage` (all 14), `test_m11_long_run_memory`, `test_memory_rewrite` (L04), `test_cast_brief`, `test_memory_settling`, `test_embedding_model_switch`, `test_bundle_v2`, `test_m10_bundle`, `test_m11_findings`, `test_v11_b1_long_run_verdict` (M04 classifications) | 244 passed | | I-series bundle and memory tests (`test_m9_*`, `test_bundle_v2`, `test_m10_bundle`, `test_process_restart`, `test_save_points`) | in the full suite, all pass | **Frontend: unchanged.** No file under `frontend/` is in the diff, so the frontend tests, lint and production build are unaffected and were not rerun. `similarity` keeps its meaning, so the inspector's display is unchanged. --- ## R. Residual risks | # | Risk | Why it remains | Where it would show | | --- | --- | --- | --- | | 1 | **No relevance floor.** A full `memory_top_k` is still used whenever the bank holds that many, however unrelated | Out of B.2's authorised menu. It costs tokens, not correctness, and a floor is model-dependent | Memories section filled with weak matches on a scene unlike anything remembered | | 2 | **Weights chosen on a deterministic embedder.** `INPUT_WEIGHT` 0.6 and `LEXICAL_WEIGHT` 0.15 were swept with the concept embedder, not `nomic-embed-text` | The deterministic fixtures are isolation-valid; a real run is not guaranteed to be (§L). Real cosines sit higher and closer together (B.1: an unrelated query still scored 0.44), so the lexical term may matter more there, or less | Real-model `ranked` stage; the snapshot now records all three scores to measure it | | 3 | **A fact in the middle of a very long block** is still not shown to the summariser | The excerpt is bounded by design. Blocks over 2,000 tokens need about 330-token actions; B.1's real run had 688-token blocks | `created: no` with `fact_in_block: yes` on a campaign with very long replies | | 4 | **Eviction keeps coverage, not any particular fact.** An interior memory holding an important fact can still be thinned when its neighbours are close | The rule protects the opening and newest stretch, and spreads what remains. It has no notion of importance, which the plan left out of scope (no new field) | A campaign past capacity whose key fact sits in a dense middle stretch | | 5 | **Coverage ignores branches.** A sibling line's memory at the same depth makes an active-line memory look redundant | Eviction is adventure-wide by design (v1.0.0 too). Of two memories sharing a start, the less recently used goes, which after a divergence is normally the abandoned one; but an active-line memory unused since before the divergence could lose to a sibling used later | Heavily branched campaigns past capacity | | 6 | **Real-model isolation is hard to obtain.** A narrator told a salient fact tends to restate it, and a summariser keeps "important established facts" | Model behaviour, not a memory mechanism. B.2 did not change prompts or summary behaviour to manufacture a qualifying run | §L | | 7 | **The summariser may still paraphrase the marker** in a way the exact-string removal does not catch | Only the exact marker is removed; a model rewording it ("part of the story is missing") would store that wording | Memory text mentioning an omission; not seen deterministically | | 8 | **The mid-reply protocol echo** (B.1, action 153) | Not WP-B's; A2 was not modified | §N; the v1.1 release gate's protocol-leak result | | 9 | **Existing campaigns** keep memories written from last-2,000-token excerpts | Existing memories are not regenerated (plan non-scope). `tools/rewrite_memories.py` is the opt-in repair | Old long blocks in v1.0.0 campaigns | | 11 | **The real summariser drops facts stated by the player.** In the one isolation-valid real run, the memory of the planting block omitted the planted fact, kept a fragment of its sentence attributed to the wrong character, and followed the narrator's reply instead | Not one of the three deficiencies this brief authorised. The plan's bounded menu includes "the memory prompt's instruction to keep named facts and objects", but B.2 was told not to change prompts or summary behaviour. **This is the open WP-B failure (§S)** | §L.2; any campaign whose narrator does not echo a player-stated detail | | 12 | **The 50-word memory ceiling is not enforced.** Attempt 2's planting memory was 102 words, and described events past its own block | Prompt-only since v1.0.0 (B.1 §B.4). A longer memory costs section tokens (the section stayed inside budget, §M) and can mix in later events | Memory text lengths in the bank | | 10 | **Cosmetic: a doubled full stop in the scene text** ("rain outside..") when the state's scene summary already ends in one | Found in attempt 1's stored query. Only the embedder reads it; the effect is one stray token. Not changed after the real-model evidence was taken, so that evidence describes the tree as staged | `memories.query.context` in any snapshot whose scene summary ends in punctuation | --- ## S. Final WP-B disposition *This section records the owner's final disposition (2026-09-15). It replaces two earlier decision texts that stood before the owner decided. Their evidence is unchanged in §L and §T.* ```text B2.1 RANKING: PASS B2.2 EVICTION: PASS B2.3 EXCERPT CREATION: PASS B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED DETERMINISTIC WP-B: PASS REAL-MODEL WP-B: FAIL WP-B OVERALL: ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION ``` ```text independent deterministic recovery: PASS reference-model independent recovery: NOT DEMONSTRATED / FAILED ON THE PRECONDITION-VALID ATTEMPT ``` **What ships, and what each part proved:** - **B2.1 ranking** (§C, §D): - the planting-era memory moves into the selected top-k (rank 7 and 6 → 1 and 2); - the player's current input materially affects retrieval; - semantic retrieval stays active, since a paraphrase with no shared word ranks first; - the lexical score is inspectable, recorded per used memory. - **B2.2 coverage-aware eviction** (§E, §F): - the early memory survives every deterministic over-capacity case; - the bank keeps coverage from the opening to the newest stretch; - pins and lineage stay safe, and the bank stays bounded. - **B2.3 bounded head+tail excerpt** (§G, §H): - a fact at the start of a long block reaches the summariser; - late facts stay visible, and the input stays within 2,000 tokens; - short blocks are unchanged. - **Together** (§I): `independent_full`, which fails on `v1.0.0` at creation, returns `recovered_through_memory_independent`. Isolation is asserted on every turn, and provenance resolves to the planting turn. With the B2.4 prompt reverted, this still passes (§T.17). **The real-model limitation, accepted for v1.1.** The reference memory summariser, `qwen2.5:3b-instruct-16k`, can be given the complete relevant source block and still fail to keep a distinctive player-established fact. The failure can be: - omitting the fact; - omitting its specific objects; - attributing it to the wrong character; - preferring the generic narration that follows it. A precondition-valid real attempt failed exactly this way (§L.2). **Real-model independent-memory recovery is therefore not guaranteed**, even though creation-window, retention, ranking and injection now work under deterministic isolation. This is a failure, not an inconclusive result. **B2.4, rejected.** - **Why it was tried:** the precondition-valid attempt failed at memory creation. - **What it was:** a prompt experiment (§T.4). - **Measured on the reference model** (§T.8): - target-fact retention on the original failing block: **0 of 5 under both the old and the experimental prompt**; - it reduced wrong-character attribution (8 → 3 of 40) and invention (2 → 0); - it introduced frequent "Memory:" prefixes (29 of 40) and frequent second-person "you" (18 of 40); - it produced more over-50-word memories (27 against 13); - on one fixture it lost a promise the old prompt always kept (5/5 → 0/5). - **Conclusion:** no reliable net improvement for the target failure. The production prompt is v1.0.0's again (§T.17). B2.4 did not pass. **Scope held.** No further prompt variation was tried. Not introduced: - a larger summariser model; - a second extraction pass or structured fact extraction; - a fact database; - another LLM call, model routing, or automatic re-summarisation. Those are future options (§T.18). **Release-gate visibility.** `V1.1-PLAN.md` §11 now requires the v1.1 release report to state the deterministic PASS, the reference-model FAIL, the failing stage (memory creation, content selection) and the owner's acceptance. It must not reduce these to "WP-B passed" (criterion 12). It must also carry the mid-reply A2 echo and the doubled full stop in the scene text as separate residuals (criterion 13). Nothing is committed, pushed or tagged. WP-C has not started. --- ## T. B2.4 — Memory summariser fidelity for named facts and objects ### T.1 Authorisation The owner's B2.4 brief (2026-09-15), under `V1.1-PLAN.md` §8 WP-B's bounded menu: "the memory prompt's instruction to keep named facts and objects". No new architectural scope. The trigger is attempt 2 (§L.2): precondition-valid, whole block sent, fact omitted. **Repository at start:** - HEAD `beb17ad`, signed (good signature, `02C9BF7D…`); - B2.1-B2.3 staged: 12 files, +2,525/−183; - nothing unstaged; WP-C not started. ### T.2 The prompt before B2.4 `MEMORY_SYSTEM_PROMPT` as shipped in v1.0.0 (184 tokens, verbatim in `tools/memory_fidelity.py` as `BASELINE_MEMORY_PROMPT`, checked against git by a test): | # | Question | Answer | | --- | --- | --- | | 1 | What does it tell the model to preserve? | "the concrete facts and events (names, places, items, promises, injuries)"; "the details a later scene could turn on — a name, a promise, an injury, where something is" | | 2 | What does it prioritise? | Nothing explicitly. "Drop the ones it could not [turn on]" is the only rule of precedence, and it does not say what gives way to what | | 3 | Attribution? | Only through pronouns: name the characters rather than "he", "she" or "they". Nothing on keeping an action, promise or possession with its owner | | 4 | People, objects, places, promises, clues? | Names, places, items, promises and injuries are listed; "where something is" is named. Clues, discoveries, possession and knowledge are not | | 5 | Invention? | Nothing | | 6 | Length? | "1-2 plain sentences", "at most 50 words" | | 7 | Does it implicitly favour narrator prose? | **Yes.** It says "The excerpt is written in the second person: 'you' is the protagonist". The protagonist's own lines are marked `>` and are often first person ("> I watch …", `format_player_input` leaves "I" alone), and nothing says what they establish is story. In the failed block, that line was 26 of about 460 words; the narration around it was 386 | | 8 | Does its wording explain the failure? | Plausibly. "Facts and events" with a two-sentence cap favours the event sequence, and the depth-4 reply was the most event-dense text in the block (the decision, the walk to the abbey, entering the crypt). The static, player-stated fact (where the sundial is) had no stated precedence over it | ### T.3 The observed failure, stated precisely From attempt 2's stored block and memory (§L.2): - The fact was in the block, at depth 3, as the player's line. - The memory omitted the sundial and the teapot. - It was 102 words against a 50-word target. - It followed the narrator's depth-4 reply. **Two corrections to §L.2's first analysis**, now marked there: - **The crypt events are not past the block.** The depth-4 reply narrates them, so the memory invented nothing. - **"promising not to tell anyone" has the right promiser.** In the source, Aldric promises Mara. The memory's error is what the promise is attached to: it ties the promise to going to the crypt, and drops Mara and the object. The phrase may also echo Edrin's own "I'll tell no one what we found". Failure stage: creation. Cause class: the summariser's content selection and fidelity, not the excerpt (the whole 623-token block was sent). ### T.4 The prompt correction One constant changed (`memorybank.MEMORY_SYSTEM_PROMPT`, 184 → 294 tokens), plus a no-behaviour-change helper, `memory_user_prompt(brief, excerpt)`, factored out of `summarize_block` so an evaluation sends a model the identical user message. Still one call, one prompt, the same 50-word target, no example, no genre vocabulary. It states: - keep concrete facts a later scene could turn on: specific people, objects and places; where something is; who has, hid, found, knows, saw or promised what; injuries, clues and commitments; - a distinctive fact comes before mood, scenery, routine movement and small talk: drop those first, and never let a later passage crowd out an earlier fact; - keep each fact with the person it belongs to; never move an action, promise, possession, statement or piece of knowledge between characters; never add a fact the excerpt does not state; - the narration calls the protagonist "you", and lines beginning `>` are the protagonist's own actions and words ("I" or "You"). What they establish is story just as the narration is. This is equal standing, not a preference for player text; - the framing rule is kept: third person, the protagonist by name, no bare pronouns, no preamble. No truncation was added. A memory over the target is stored as written (`test_an_over_long_memory_is_stored_as_written_never_cut`). *Owner disposition: this prompt was rejected and reverted (§T.17). The production prompt is v1.0.0's. The text above records the experiment.* ### T.5 Deterministic fixtures `tools/memory_fidelity.py` holds eight fixtures, each a short block with texture and a fact, declaring the facts to keep, their owners, and forbidden inventions. | Fixture | Requirement | Genre | | --- | --- | --- | | `object_place_office` | B2.4-1 object and place | office | | `player_fact_station` | B2.4-2 player-established fact | science-fiction-neutral | | `promise_contemporary` | B2.4-3 promise | contemporary | | `attribution_station` | B2.4-4 attribution (two characters, a code and a badge) | science-fiction-neutral | | `clutter_office` | B2.4-5 clutter pressure | office | | `no_invention_office` | B2.4-6 no invention (an unclaimed briefcase) | office | | `multiple_facts_station` | B2.4-7 several facts, three owners | science-fiction-neutral | | `regression_attempt_2` | Part 4, the actual failure | the harness campaign | **The checker** (`evaluate`) is deterministic and a stated heuristic: - a fact is kept when one sentence names all its parts; - attribution is the nearest named character before the fact's verb; - inventions are patterns anchored on the wrong character as subject; - words are counted, not enforced. `test_v11_b2_summarizer_fidelity.py`, **49 passed:** - 12 prompt-contract checks, the 50-word target, genre neutrality with no example or fixture name, framing rules kept, and the baseline equal to the shipped prompt; - coverage of every requirement in three genres; - every fixture's facts reaching the summariser whole; - the checker passing each faithful memory and failing each failure shape (13 unfaithful memories: omission, misattribution, invention); - the application sending the corrected prompt with the whole planting block; - no truncation; - the corrected regression memory ranking first under B2.1 scoring. B2.4 adds no xfail. ### T.6 The actual failed-run regression `regression_attempt_2` is attempt 2's planting block verbatim. It is the harness's own fixture campaign, with no hostname, person or other identifier. | Required after B2.4 | Deterministic | Reference model, B2.4 prompt (§T.8) | | --- | --- | --- | | input contains planted fact | yes | yes | | summariser receives planted fact | yes (the whole block; test) | yes | | generated memory contains planted fact | checker passes the corrected memory and fails the stored one (test) | **no: 0 of 5** | | attribution correct | checked | not reached (fact absent) | | distinctive objects retained | checked | **no** | | unsupported facts invented | no | no (0 of 5) | **The pre-B2.4 prompt fails the fixture legitimately.** The stored memory is the real output, and the old prompt also scored 0 of 5 in §T.8; nothing was fabricated. The B2.4 prompt fails it equally. ### T.7 Length and attribution behaviour | 40 memories per arm (8 fixtures × 5) | Baseline prompt | B2.4 prompt | | --- | --- | --- | | median / p90 / max words | 48 / 72 / 82 | 68 / 83 / 105 | | over the 50-word target | 13 | **27** | | misattributed | 8 | **3** | | invented | 2 | **0** | Worst length: 105 words (`object_place_office`, B2.4, a run-on of four "Memory:" clauses). The length target was measured, not enforced; nothing was cut. ### T.8 Real-model comparison `tools/memory_fidelity.py`, 2026-09-15 09:26-09:27: - **Model:** `qwen2.5:3b-instruct-16k` on the GPU inference host, through the application's provider at production temperature (0.3). - **Samples:** 5 per fixture per prompt. - **Logging:** running (§T.11). - **Evidence:** `$HOME/v11-evidence/b2/fidelity-b24/fidelity.json`, every memory verbatim. | Fixture | Baseline passed | B2.4 passed | What the memories show | | --- | --- | --- | --- | | `object_place_office` | 5/5 | 2/5 | B2.4 runs began "Memory:"; two lost the drawer and cabinet ("in the archive room"); one run-on of 105 words, whose "Priya … sliding" the checker mis-scored (`sliding` is not a `slide` form it matches) | | `player_fact_station` | 2/5 | 4/5 | baseline misattributed and invented; B2.4 kept the sample and its owner | | `promise_contemporary` | **5/5** | **0/5** | every B2.4 run began "Memory:" and listed only scenery (floorboards, radiator, dog); four dropped the promise, one mentioned the lease without who promised; two used "You". The baseline kept the promise every time | | `attribution_station` | 1/5 | 4/5 | baseline gave the code or badge to the wrong person | | `clutter_office` | 4/5 | 5/5 | | | `no_invention_office` | 3/5 | 3/5 | | | `multiple_facts_station` | 1/5 | 5/5 | | | **`regression_attempt_2`** | **0/5** | **0/5** | every memory under both prompts retells the key, the decision and the crypt; none names the sundial or teapot | | **total** | 21/40 | 23/40 | B2.4: "Memory:" prefix 29/40, "you" 18/40 (baseline 0 and 0) | Reading it plainly: - **Attribution and invention improved.** - **The target failure did not move.** - **The longer prompt made this 3B model echo the "Memory:" cue from the user message,** drift into second person, write longer, and on one fixture turn the "drop scenery" instruction into a scenery list. - **This is a model-capability limit at this prompt size,** measured, not guessed. ### T.9 B2.1, B2.2 and B2.3 with B2.4 in the tree | Suite | Result | | --- | --- | | `test_v11_b1_memory_diagnostic.py`: all B.1 scenarios, `independent_full` (`recovered_through_memory_independent`, isolation every turn, provenance), lineage (G), authority | pass | | `test_v11_b2_memory_ranking.py`: B2.1, negative controls, pins | pass | | `test_v11_b2_memory_eviction.py`: B2.2, bounds, pins, fallback, lineage separation | pass | | (the three files together) | **81 passed** | | `test_v11_b2_memory_excerpt.py`, `test_cast_brief.py`, `test_memory_rewrite.py` | **53 passed**; short-block prompts byte-identical, excerpt split unchanged | The deterministic scenarios use the best-case scripted summariser, so they are independent of the prompt. That they still pass shows B2.4 disturbed no mechanism; it is not evidence about the prompt. No ranking weight, eviction rule or excerpt split changed. ### T.10 Full deterministic WP-B Unchanged from §I on the B2.4 tree: - `independent_full`: created, retained, ranked and injected; - `recovered_through_memory_independent`, with isolation asserted on every one of 52 turns; - provenance to depths 0-5. Lineage (§J) and authority (§K) pass unchanged. ### T.11 GPU-host evidence review For the WP-B runs so far: B.1 on 2026-09-14 18:16-18:37, and B.2 attempts at 22:00-22:22 and 22:23-22:44. Read-only, from the host's logs and its persistent journal. | Check | Finding | | --- | --- | | Power | peak 294.5 W (B.2 window), 304.6 W (earlier set), 307 W (`dmon`). The power limit now reads 200 W (default 280 W), so it was set after these runs. The owner notes spikes above the limit occur, and the dock is rated to 450 W | | Power-limit events | `dmon` power-violation samples: 17 of 5,804 in the B.2 set. `dmon` has no timestamps, so they cannot be placed within the runs | | Temperature / throttling | peak 68 °C; thermal-violation samples 0 | | PCIe link | Gen3 x4 under load in 2,462 of 2,463 busy samples (one Gen2 sample); PCIe error counter 0 | | Kernel | 4 kernel lines in the whole window; no Xid, NVRM, AER, reset or fallen-off-the-bus | | Ollama | no crash, exit or panic; models loaded once at 22:00:36-38 and stayed loaded through attempt 2 | | HTTP errors | 9 × 404: the harness's deliberate `failed_call` probe (`no-such-model-m11`). 4 × 500: background calls cut off by the harness's planned server restarts (22:18:13 at restart 3; 18:33:13 in B.1). 21:29:54: other traffic outside both windows | | **Contamination** | **None.** No host fault could have altered the WP-B evidence; attempt 2's failure is model output | **A logging defect, found and fixed.** The documented watch command, `journalctl -f -k -u ollama`, matches nothing: `-k` and `-u` are different fields, which journalctl ANDs. Over the B.2 window it returned 0 lines, where the OR form (`_TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service`) returned 30,229. Every such log was therefore empty, and events were read from the persistent journal instead. - `DEVELOPMENT.md` now documents the OR form. - For the B2.4 comparison the owner's watch still ran the old command, and its file is again empty. The journal holds 2,091 Ollama lines for 09:25-09:28. ### T.12 Fresh real-model attempts **0.** The comparison (§T.8) ran first and showed the B2.4 prompt failing on the exact planting block the long run plants. Under the brief ("If the summarizer still fails … then stop. Do not create B2.5 automatically"), the long attempts were not run: they could only have confirmed a known creation failure, or failed isolation. ### T.13 Context-budget interaction - **Memory prompt:** 184 → 294 tokens (+110) in the background memory call only. The narrator's prompt, A1's reserve, the transport cost and the accounting are unchanged. - **Typical memory length** on the fixtures: median 48 → 68 words; worst 105. - **Real-model turns after B2.4:** none, so no turn was `exceeded` or `truncation_suspected`, and A1's protected budget was never at issue. - **Before B2.4:** every counted turn of both B.2 attempts was `fits`, with the memories section at most 1,246 tokens (§M). - **If the B2.4 prompt were kept,** its longer memories would grow the memories section roughly in proportion (median +40%). It is still priced as protected context, and A1 would refuse rather than overflow. ### T.14 Offline regression `tools/m11_offline.py --out $HOME/v11-evidence/b2/offline-b24`, on the B2.4 tree: a fresh `--no-cache` image in a `--network none` container, driven over `docker exec`. **23 passed, 0 failed**, the same checks as §P.1: - no route and no external DNS; - the page loads offline, with a CSP and local assets; - a campaign is created, state is extracted, and a file imports; - the prompt assembles, and knowledge is searchable; - a turn with no model is a reported failure that loses nothing; - export and import work, with no secret in the export; - the media module stays inert; - campaigns survive a restart. **What B2.4 changes on the network:** - **Nothing.** The diff is a prompt string, a pure helper, an evaluation tool, tests and documentation. - **No new destination, dependency or external service.** `tools/memory_fidelity.py` calls only the `--endpoint` it is given, through the application's own provider, and it is a developer tool that production never imports. - **Unchanged:** the endpoint policy, the trusted-LAN policy, embeddings (local), and the vector store (SQLite). ### T.15 v1.0.0 compatibility Rerun on the B2.4 tree against `m04-final/campaign.db` (SHA-256 prefix `c74e8798b68d2e22`, unchanged by the runs), from a fresh `432f041` worktree: | | Result | | --- | --- | | Schema before and after opening | identical to v1.0.0; `user_version` 94 → 94 | | API snapshot | identical | | memories (33), summaries (12), actions (207), state events (32), Save Points (2), branches (4), knowledge sources (3), adventures (1) | identical, row for row | | Exercise | export and import 207 → 207 actions, same head, same narrative state | | Bundle format | v3, unchanged; no memory regenerated | ### T.15a Full backend regression **Full backend suite, B2.4 tree: 1,701 passed, 17 skipped, 0 failed, 0 xfailed** (1,101 s). - The count is B2.1-B2.3's 1,652 plus B2.4's 49 fidelity tests. - The 17 skips are the same real-model tests as before. - The suite covers: all WP-B tests, all B.1 diagnostic tests, B2.1-B2.4, F01-F08 and E01-E04, the M04 classifications, `test_context_memory`, `test_memory_nodes`, `test_memory_retrieval`, `test_m11_leakage`, `test_m11_long_run_memory`, `test_memory_rewrite`, the I-series bundle and memory tests, derived-work failure recording and write-lock regressions. **Frontend: unchanged.** No file under `frontend/` is in the WP-B diff, so the frontend tests, lint and build were not rerun. ### T.16 Residual risks added by B2.4 | # | Risk | | --- | --- | | 13 | **The reference summariser does not keep a player-stated fact against event-dense narration**, under either prompt (0 of 10 on the attempt-2 block). This is the open WP-B limitation | | 14 | ~~The B2.4 prompt regresses framing on a 3B model~~ — **resolved by reverting it** (§T.17). Kept as evidence of why it was rejected | | 15 | **The fidelity checker is a heuristic.** It missed one correct attribution (`sliding`), and may miss paraphrases. Every memory is kept verbatim so a reader can check it | | 16 | **Operator logging:** the documented kernel/Ollama watch recorded nothing on every run before the fix. The journal is the source of those events | | 11-12 (earlier) | still open | | 8 (A2 mid-reply echo) and 10 (doubled full stop in the scene text) | unchanged, and still out of scope | ### T.17 Owner disposition, and the revert **Owner decision, 2026-09-15: option 1(a).** - B2.1, B2.2 and B2.3 are accepted. - The B2.4 prompt does not ship. - Deterministic WP-B is accepted. - The real-model failure is accepted as a documented v1.1 residual limitation. | Reverted | Kept | | --- | --- | | `memorybank.MEMORY_SYSTEM_PROMPT`, restored byte-for-byte to `beb17ad`'s (184 tokens), and the B2.4 comment above it | B2.1 retrieval and ranking, B2.2 eviction, B2.3 head+tail excerpt | | the tests that required the experimental prompt's wording | `memorybank.memory_user_prompt`: a behaviour-neutral helper `summarize_block` calls, so a measurement sends a model the application's exact message | | the B2.4 note in `CONTEXT-AND-MEMORY.md` §15, replaced by the shipped behaviour and the limitation | B.1 and B.2 diagnostics; `tools/memory_fidelity.py`, now diagnostic-only, comparing the shipped prompt with `B24_EXPERIMENT_PROMPT`, which it keeps so the experiment can be repeated | | | this section's evidence, the fixtures, and every measured memory | **The tests after the revert** (`test_v11_b2_summarizer_fidelity.py`) are split in two. - **Acceptance** gates the tree: - the shipped prompt equals v1.0.0's, checked against git; - the experiment is not what ships; - every fixture reaches the summariser whole; - the application sends the shipped prompt with the whole planting block; - an over-long memory is never cut; - a faithful regression memory ranks first under B2.1. - **Diagnostic measurement** proves only the instrument, on hand-written memories with known answers: - fact retention, attribution and invention; - word count, a leading "Memory:" and second-person "you"; - promise retention. No reference-model score is a test gate. Verification after the revert: - The runtime constant equals `beb17ad`'s value. - The only prompt-related lines left in the diff against `beb17ad` are the `summarize_block` call through `memory_user_prompt`. - The fidelity, cast-brief, rewrite and excerpt tests: **94 passed**. **Final regression, offline and compatibility, on the final staged tree** (the B2.4 prompt reverted): | Check | Result | | --- | --- | | **Full backend suite** | **1,693 passed, 17 skipped, 0 failed, 0 xfailed** (1,227 s). This is B2.1-B2.3's 1,652 plus 41 summariser tests; the 8 that required the experimental prompt's wording went with it. The 17 skips are the same real-model tests | | Deterministic WP-B within it | B2.1 ranking, B2.2 eviction, B2.3 excerpt; `independent_full` still `recovered_through_memory_independent`, with isolation every turn and provenance; lineage (G); authority; the B.1 diagnostics; F01-F08, E01-E04, M04, the I-series, derived-work failure recording and write-lock regressions: **all pass** | | **Offline** (`tools/m11_offline.py`, fresh `--no-cache` image, `--network none`) | **23 passed, 0 failed**. Rerun because the runtime differs from §T.14's tree | | **v1.0.0 compatibility** (`m04-final/campaign.db`, compared with the `432f041` snapshot taken the same day) | API snapshot and schema identical; `user_version` 94 → 94; memories (33), summaries (12), actions (207), state events (32), Save Points (2), branches (4), knowledge sources (3) and the adventure identical, row for row; source database unchanged | | Frontend | unchanged; not in the diff | ```text deterministic WP-B: PASS offline regression: PASS (23/23) v1.0.0 compatibility: unchanged ``` ### T.18 Future memory-quality options (not implemented) Each of these is new scope for a future package, and none is part of v1.1: - **A stronger dedicated summariser model,** evaluated with `tools/memory_fidelity.py` against the same fixtures. - **Structured fact extraction** alongside the prose memory: who has, hid, knows or promised what. - **Separate factual and narrative memory,** each retrieved on its own terms. - **Model-specific summariser recommendations** in the operator documentation, once measured.