Files
interactive-story/planning/reports/v1.1/V1.1-WP-B2-REPORT.md
T
JesseMarkowitzandClaude Opus 5 0c1ba836ba v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 11:21:53 -04:00

70 KiB
Raw Blame History

v1.1 WP-B.2 — Independent Long-Term Memory Retention

Status: COMPLETE — accepted by the owner with a documented real-model limitation (2026-09-15, option 1(a)). Staged, uncommitted, for the owner's signed commit.

  • Shipped: B2.1 ranking, B2.2 eviction and B2.3 excerpt creation.
  • Not shipped: B2.4, a memory-prompt experiment. It did not correct the reference model's creation failure, and its prompt was reverted (§T).
  • The disposition is in §S.

A. Repository baseline

Branch v1.1-development
HEAD at start beb17ad — v1.1 WP-B.1: diagnose independent long-term memory retention, signed by the owner (good signature, RSA key 02C9BF7D…)
Its parents d63804f (WP-A1/A2, signed), ac465ed (plan v4.1), 432f041 (v1.0.0, tag v1.0.0)
Working tree at start clean
Comparison baseline 432f041 (v1.0.0), from a throwaway worktree. The B.1 diagnostic tool, which is not in that tree, was copied into it from beb17ad, with only the new scenario definitions and the coverage trace added. app/ in the worktree is v1.0.0's, unmodified

B. B.1 findings being corrected

WP-B.1 (V1.1-WP-B1-REPORT.md §L, §M) placed three deficiencies. They are the only memory behaviour this package changes.

# Stage B.1 evidence Corrected in
1 Ranking Real 100-turn run: the planting-era memory was created, active and carried F, and ranked 10th of 19 (cosine 0.866) against memory_top_k 4. The query was the newest four actions cut to 600 tokens, so the one-line question was diluted by narration (deterministically, 0.708 alone fell to 0.241 in the query). That run failed isolation, so it was diagnostic evidence only B2.1
2 Retention Past memory_bank_capacity, least-recently-used eviction removed the planting-era memory first (turns 21, 21, 36 in the capacity fixtures). Strict xfail test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity B2.2
3 Creation The summariser saw only a block's last 2,000 tokens; a fact at the start of a 2,079-token block never reached it (not_created). Strict xfail test_acceptance_a_fact_early_in_a_long_block_is_remembered B2.3

A fourth B.1 observation, a mid-reply protocol echo, is not WP-B's and is left alone (§N).


C. B2.1 ranking design

C.1 The retrieval query

v1.0.0 embedded one text: the newest four story actions, cut to their last 600 tokens. It is replaced by two short texts embedded in one call (memorybank.retrieval_query):

Component Source Bound
input the newest action, when it is a player action with text (do, say, story) last 200 tokens (QUERY_INPUT_TOKENS)
context the scene from the authoritative state — scene.summary, the location's entity name, the names of scene.present (at most 8) — then the end of the newest narration 60 tokens (QUERY_SCENE_TOKENS) + 120 tokens (QUERY_NARRATION_TOKENS)
  • A continue turn, and the Insights dry run, have no player input: the context alone is searched, and the lexical score is 0.
  • A retry excludes the attempt being replaced, as before; the query is the retried input and the narration before it.
  • Why the context stays. CONTEXT-AND-MEMORY.md §18 requires more than raw input, and a question often means nothing without its scene ("I ask her what she keeps up there"). It is bounded so it can resolve a reference but cannot outweigh the question by length. The full entity list, threads and older narration are left out on purpose.
  • Provenance. The query is recorded in each turn's snapshot as memories.query: input, context, input_terms, input_weight, lexical_weight.

C.2 The ranking formula

semantic_score = INPUT_WEIGHT * cos(input, m) + (1 - INPUT_WEIGHT) * cos(context, m)
                 INPUT_WEIGHT = 0.6; either cosine alone if the other text is empty
lexical_score  = Σ w(t) for t in input_terms ∩ terms(m)  /  Σ w(t) for t in input_terms
w(t)           = ln((N + 1) / (df(t) + 1))       N eligible candidates, df holding t
final_score    = semantic_score + LEXICAL_WEIGHT * lexical_score,   LEXICAL_WEIGHT = 0.15
  • Terms (memorybank.lexical_terms): imported knowledge's tokenizer and stop list (knowledge.fts.terms), a possessive 's dropped, and a trailing s folded from words of five letters or more not ending in ss. No stemmer, no dependency.
  • Rarity is computed per call over the eligible candidate set only. Nothing is indexed or stored. A word every candidate holds weighs exactly 0; a word none holds weighs most, so the unmatched part of a question dilutes every candidate equally.
  • Range. lexical_score is in [0, 1], so wording can move a memory by at most 0.15 (test_a_lexical_match_moves_a_memory_at_most_the_lexical_weight).
  • Only the player's input is matched lexically, never the context.
  • Pins are exactly as before: always used, counted toward memory_top_k.
  • Ties on the final score are broken by memory id, so row order never decides.
  • Redundancy suppression (§22, cosine ≥ 0.93, never across authority) is unchanged and runs over the final order.
  • Reported per used memory: semantic_score, lexical_score, final_score. similarity keeps its v1.0.0 meaning, the semantic score, so the inspector's "closeness" needs no frontend change.

Egress. Memory text is read for ranking only on a turn whose input has terms, only for memories not already held, and held beside the vectors in a cache with the same two rules: set_vector drops an entry (an edit clears the vector, so a changed text is re-read), and the next read discards memories no longer in the catalogue. A continue turn reads no memory text beyond the top-k detail read, and a second turn in a row reads none again (test_memory_text_is_read_once_and_then_held).

C.3 Choosing the lexical weight

Swept over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5 on three deterministic fixtures, each ranked two ways: with the new query, and with the v1.0.0 query plus the lexical term (the lexical term's own contribution, with the crowding left in). Rank of the planting-era memory F, memory_top_k 4:

Weight 0 0.05 0.1 0.15 0.2 0.3 0.5
ranking_crowded, new query 1 1 1 1 1 1 1
ranking_crowded, v1.0.0 query + lexical 7 6 6 1 1 1 1
ranking_crowded, unrelated question: F selected? no (9) no (9) no (6) no (6) no (6) no (5) yes (3)
negative control (paraphrase + a decoy's rare place): decoy above F? no no no no no no no
  • 0.15 is the smallest weight at which the lexical term alone rescues F from the crowded v1.0.0 query.
  • 0.5 already lets an incidental shared word ("landing") select F for an unrelated question, which is the failure a larger weight buys.
  • The negative control never put the decoy above F at any weight with the new query.
  • Not tuned to one phrase: the direct question, a paraphrase sharing only "Mara", a context-dependent question sharing no word at all, a decoy word from a different memory, and an unrelated question all ran at every weight.

One normalisation was rejected along the way. Taking the share over only the input words some candidate holds let a lone incidental match score 1.0, and in the context fixture a decoy holding "fish stalls" then outranked F from weight 0.15. The share is now over all the input's words.


D. B2.1 before/after evidence

Fixtures (tools/memory_diagnostic.py, best-case summariser, concept embedder, isolation valid, below capacity, recall at depth 106):

  • ranking_crowded: 150-word narration, memory_top_k 4 — B.1's real-run failure made deterministic.
  • ranking_context_dependent: the same, with the last narration putting Mara at the tavern's top shelf by an old kettle, and the question "I ask her what she keeps up there." Neither "kettle" nor "shelf" is one of F's leak terms.
Fixture v1.0.0 (432f041) B2.1
ranking_crowded retained_but_not_ranked: rank 7 of 17, similarity 0.1653 injected: rank 1 of 17; semantic 0.4671, lexical 0.5328, final 0.5470; selected [1, 7, 11, 12]; 87 tokens
ranking_context_dependent retained_but_not_ranked: rank 6 of 17, similarity 0.1577 injected: rank 2 of 17; semantic 0.1866, lexical 0.0, final 0.1866; selected [1, 5, 6, 9]; 89 tokens
independent_default (top-k 5) injected: rank 2, similarity 0.241 injected: rank 1; semantic 0.4739, lexical 0.5328, final 0.5538

In every row the diagnostic's replica of the ranking (production's own score_candidates and select_memories) matched the recall turn's stored selection.

The required ranking controls:

Control Result
Direct question F rank 1 in every fixture (above)
Semantic paraphrase ("the little brass dial that tells the hour") rank 1, lexical 0.1125 (only "Mara" shared) against 0.5328 for the direct question: found by meaning
Context-dependent wording rank 2 with the scene; rank 17 of 17 with the context removed; lexical 0
Negative control (paraphrase + "out by the fish stalls", held only by a filler memory) the decoy's lexical score (0.18) exceeds F's (0.09), and F still ranks 1, the decoy 2-5
Common words ("The travellers spent time.") lexical 0.0 for every memory
Pins always used, counted toward top-k (test_a_pin_is_still_always_used_and_counts_toward_top_k, and v1.0.0's test_pinned_memories_are_always_used, unchanged)
No relevance floor (unchanged) an unrelated question still selects a full top-k; F is no longer carried along by it (rank 6 of 17, not selected)

Tests at the checkpoint: test_v11_b2_memory_ranking.py 18 passed; test_v11_b1_memory_diagnostic.py 31 passed and 2 xfailed (the eviction and creation criteria, not yet fixed); the memory regression group (13 files) 226 passed.

B2.1 RANKING: PASS


E. B2.2 eviction design

memorybank.eviction_order(rows, limit), run by _evict_over_capacity after each post-turn pass. It is a pure function of seven columns per active memory and reads neither text nor vectors (test_eviction_reads_no_vectors, unchanged).

Part Rule
Coverage signal Memories with a source range are ordered by position. A memory is judged by the hole its removal would leave: the start of the next memory, less the furthest end before it, less one. Smallest hole first, so the bank thins where it is densest. A memory whose source_start another active memory shares leaves no hole (0)
Boundaries The earliest and the latest memory by position are not coverage candidates: they alone describe the opening and the newest stretch
Recency signal Among equal holes: coalesce(last_used_at, created_at) oldest first
Tie-breaking then use_count lowest first, then id lowest first
Pinned rows never taken; they count as coverage for their neighbours; if every active memory is pinned the bank may exceed capacity, as before
Fallback once no coverage candidate remains — memories typed by the player or migrated have no range, and a bank can be all boundaries — the rest are taken by the v1.0.0 order (recency, use count) plus id
Recomputation after every pick, since removing a memory widens its neighbours' holes
  • Unchanged: the capacity setting and default (80), the count over the whole adventure (not one lineage), marking forgotten rather than deleting, running once per post-turn pass.
  • Frozen bank. A memory written this turn is the latest boundary, so it is not a coverage candidate, and in the fallback it has the newest timestamp. It can still go first legitimately when it only repeats a stretch another, more recently used memory describes (test_a_newborn_can_be_the_legitimate_first_to_go).
  • Lineage. The rule decides forgotten and nothing else; eligibility is the retrieval clause, untouched. A sibling line's memory at the same depth shares a start, so it is the first kind thinned, and the least recently used of the two goes — after a divergence, that is the abandoned line's.
  • No new field. Existing source_start, source_end, last_used_at, created_at and use_count only.

F. B2.2 capacity evidence

Deterministic fixtures (v1.0.0 from the worktree, with a coverage trace added to its copy of the tool):

Fixture v1.0.0 B2.2
past_capacity (capacity 6, top-k 5, 17 memories written) created_but_evicted: F evicted at turn 21, the first eviction. Final bank covers depths 36-101 injected: F retained, rank 1 of 6. Bank covers 0-101, largest uncovered stretch 18
past_capacity_pinned evicted at turn 21; bank 6-101 injected; the pinned memory kept; 0-101, gap 18
past_capacity_low_top_k (capacity 8, top-k 2) evicted at turn 36 (first eviction 27); bank 18-101 injected: rank 1 of 8; 0-101, gap 12

In all three: active memories never exceeded capacity after a pass, and no memory was evicted by the pass that created it.

F survives these fixtures as the earliest boundary. That is the general rule, not a special case — any campaign's opening memory is kept while interior memories remain — but it is also the easiest case, so the rule was tested on interior memories in general. Memories arrive a block at a time, nothing is ever retrieved, and the bank is held at capacity. Largest uncovered stretch, counting from depth 0:

Bank Coverage-first: first memory kept / largest gap v1.0.0 order: first memory kept Average spacing (span / capacity)
17 blocks, capacity 6 0 / 18 depth 66 17.0
60 blocks, capacity 10 0 / 42 depth 300 36.0
500 blocks, capacity 80 0 / 42 depth 2,520 37.5
500 irregular blocks (4-9 actions), capacity 80 0 / 47 depth 2,591 38.4

The largest gap stays within about 1.3 times the average spacing. The v1.0.0 order keeps an unbroken run of the newest blocks and nothing before it. test_a_long_bank_keeps_describing_the_whole_story asserts the opening kept and a gap of at most twice the average spacing.

Invariants tested (test_v11_b2_memory_eviction.py, 24 tests):

Invariant Test
Active count ≤ capacity after maintenance (or = the pins, if more) test_capacity_holds_and_pins_survive_for_any_bank (8 random banks)
Pinned never evicted same, and test_pins_are_never_taken_but_still_count_as_coverage
Abandoned-line memories not made eligible; lineage separate test_eviction_does_not_make_an_abandoned_lines_memory_eligible, test_the_pass_changes_nothing_but_forgotten
Frozen bank still works test_the_newest_memory_is_not_evicted_by_the_bank_it_joins; v1.0.0's test_a_newborn_is_not_evicted_by_the_bank_it_joins and test_a_full_bank_still_turns_over, unchanged
A newborn may still go first when legitimate test_a_newborn_can_be_the_legitimate_first_to_go
Independent of row order test_the_order_does_not_depend_on_row_order, test_a_tie_on_every_signal_is_broken_by_id
Fallback is v1.0.0's order test_memories_without_a_range_take_the_least_recently_used_fallback; v1.0.0's test_eviction_marks_the_least_recently_used and test_eviction_breaks_ties_on_use_count, unchanged

The B.1 strict xfail for capacity is now an ordinary test that passes for all three capacity fixtures. B.1's two tests that asserted F's eviction (a diagnosis of v1.0.0) now assert its retention.

B2.2 EVICTION: PASS


G. B2.3 creation design

memorybank.memory_excerpt(raw, budget=MEMORY_EXCERPT_TOKENS), used by summarize_block (and so also by the opt-in tools/rewrite_memories.py).

Block Summariser input
≤ 2,000 tokens the whole block, unchanged. The prompt is byte-identical to v1.0.0's (test_a_short_block_is_prompted_exactly_as_before)
> 2,000 tokens the block's first tokens, then \n\n[… the middle of this stretch of story is left out here …]\n\n, then its last tokens
  • Accounting (excerpt_split). The marker with its blank lines costs 15 tokens and is paid first. The remaining 1,985 are halved, the odd token to the end: 992 + 993 + 15 = 2,000.
  • Seams. Two decoded runs rejoined can tokenise differently where they meet, so the result is measured and the opening gives up tokens until the whole is within 2,000. Measured alone, a cut part can come to one token over its run; the whole never exceeds the budget, including for mixed-script and emoji text.
  • Order and marker. The opening comes first, and the marker tells the summariser the two parts are not adjacent. The header (Story excerpt:) is unchanged.
  • The marker is never stored. summarize_block removes it from the model's reply.
  • The budget did not grow. No model context, MEMORY_MAX_WORDS or prompt instruction changed. Existing memories were not rewritten.
  • Limit. A fact in the middle of a very long block is still left out (test_a_fact_in_the_middle_of_a_very_long_block_is_still_omitted).

H. B2.3 long-block evidence

Fixture v1.0.0 B2.3
long_block_fact_early (F at depth 1 of a 2,079-token block) fact in block: yes; in the summariser's excerpt: no; not_created in excerpt: yes; memory 1 (depths 0-5) carries F; injected
long_block_fact_late (F at depth 5 of the same-sized block) in excerpt: yes; created in excerpt: yes; created; injected

These two fixtures test creation. They recall at depth 22, so the planting turn is still inside the history window and their isolation check is not expected to pass; independence is §I's.

Tests (test_v11_b2_memory_excerpt.py, 17):

Required Test
F early in a > 2,000-token block test_a_long_blocks_memory_keeps_its_source_provenance (the memory holds it); B.1's test_acceptance_a_fact_early_in_a_long_block_is_remembered, now ordinary
F late in the same-sized block B.1's test_the_same_fact_late_in_the_same_sized_block_does
A block ≤ 2,000 tokens unchanged test_a_block_that_fits_is_sent_whole_and_unchanged, test_a_block_of_exactly_the_budget_is_unchanged, test_a_short_block_is_prompted_exactly_as_before
Head/tail accounting bounded test_the_split_is_even_and_documented, test_the_excerpt_never_exceeds_the_budget (5 sizes up to 10× the budget), test_the_budget_holds_for_awkward_text
The marker is not stored as story fact test_the_marker_is_never_stored_as_part_of_a_memory (a summariser that echoes its whole input)
Valid source provenance test_a_long_blocks_memory_keeps_its_source_provenance: source_start/source_end are the block's depths, branch_id/depth its last node's

The B.1 strict xfail for creation is now an ordinary test.

B2.3 CREATION: PASS


I. Full deterministic independent-memory test

independent_full: all three B.1 failures at once. Every narrator reply is about 850 words, so every memory block (2,079 tokens) is longer than the excerpt; narration crowds the query at memory_top_k 4; capacity is 8, and 17 memories are written. F is planted at depth 1 by a story action; recall is at depth 106. context_token_budget 16,384.

v1.0.0 B.2
Verdict not_created (the fact never reached the summariser) injected

Isolation, asserted on every one of the 52 turns and at recall:

Check Result
F absent from authoritative state every turn's document; every node's narrative_state_after; the recall prompt's state section
F absent from the summary the active summary on every turn, and the recall prompt's summary section
F absent from imported knowledge none imported; no knowledge section mentions it
F absent from recent history the history window at recall starts at depth 84; planted at 1; the fact is not in the history sections
F absent from later narration no narrator turn after the planting block (ends depth 5) mentions it

The four stages:

Stage Result
created memory 1, source depths 0-5, text "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." The block was 2,079 tokens and the fact was in the head+tail excerpt
retained active; bank 8 of 17 written, capacity 8; the bank covers depths 0-101 with a largest uncovered stretch of 12; nothing evicted by its own pass
ranked rank 1 of 8, top-k 4; semantic 0.4671, lexical 0.5066, final 0.5430; selected [1, 5, 7, 12]; replica matches the stored selection
injected in the recall turn's memories.used and its used_memories section, 88 tokens

Provenance. The recall snapshot's entry for memory 1 records source {branch_id: 1, depth: 5, source_start: 0, source_end: 5} and authority accepted_story. It equals the row, the range includes the planting depth, and memorybank.source_block on that row returns depths 0-5, which hold the planting action.

Long-run verdict. The same measurements, given to m11_long_run._independent_memory_verdict, return recovered_through_memory_independent (test_acceptance_full_is_the_long_run_independent_memory_verdict).

All nine deterministic scenarios, final tree:

Scenario Verdict Isolation F rank / top-k Notes
independent_default injected ok 1 / 5
past_capacity injected ok 1 / 5 capacity 6
past_capacity_pinned injected ok 1 / 5 pin kept
past_capacity_low_top_k injected ok 1 / 2 capacity 8
long_block_fact_early injected not expected (recall depth 22) 1 / 5 fact in excerpt
long_block_fact_late injected not expected (recall depth 22) 1 / 5
lineage_control injected ok 1 / 5 §J
ranking_crowded injected ok 1 / 4
ranking_context_dependent injected ok 2 / 4 lexical 0
independent_full injected ok, every turn 1 / 4 all three at once

J. Lineage

lineage_control (B.1's fixture, rerun on the final tree). Fact G is planted at turn 21 behind a Save Point; a second Save Point marks line A at turn 30; nine Undos and a different action abandon line A.

Requirement Result
The abandoned memory stays stored G's memory (id 7) is on disk
Absent on the active divergent line not eligible under the active lineage clause
Absent from memories.used no turn after the divergence names it, and G's text is in no memories section
Eligible again only when branch semantics permit restoring the line-A Save Point makes it eligible again
F on the active line still injected

Eviction cannot change eligibility: it writes only forgotten (test_the_pass_changes_nothing_but_forgotten, test_eviction_does_not_make_an_abandoned_lines_memory_eligible).

tree.attach_memory, forget_node, the lineage clause, E02 and summary lineage (E03) are not in the diff. test_m11_leakage.py passes unchanged (§Q).


K. Authority

B.1's authority control (test_a_memory_that_contradicts_state_loses_and_changes_nothing), rerun unchanged on the final tree: a state correction establishes "the tavern lamp is lit"; a pinned memory says "The tavern lamp was never lit that night."

Requirement Result
State wins the fact is in the narrative state section, which is read after the memories section
Memory remains non-authoritative it is rendered under "Memories from earlier in the story…", as before
No state mutation from retrieval the state document is identical before and after the turn

Retrieval still only reads, and a used memory's authority is still recorded. F07 and classify_authority are not in the diff.


L. Real-model attempts

Setup, identical for both attempts.

  • Command: tools/m11_long_run.py --turns 100 --independent-fact, unchanged since B.1: the same planted fact, prompts and harness.
  • Host and models: the GPU inference host (Ollama 0.34.0); qwen2.5:3b-instruct-16k narrating and summarising, nomic-embed-text:latest embedding.
  • Memory: bank and summaries on; harness settings memory_top_k 4 and context_token_budget 16,384.
  • Order of events:
    • the full deterministic suite had passed (§Q);
    • the owner confirmed the power, link and kernel logging was running on the host before attempt 1;
    • attempt 2 was run because attempt 1 failed isolation.
  • Cap: two deliberate attempts at most, as the brief set.

Stages come from tools/v11_b1_memory.py diagnose on a copy of each run's database, ranked with the same embedding model.

L.1 Attempt 1 — PRECONDITION FAILED

Evidence: $HOME/v11-evidence/b2/real-1/, $HOME/v11-evidence/b2/real-1-diagnosis/.

Run complete: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 memories written (18 eligible at recall), 1,352 s, median 10.9 s a turn
A1 accounting fits on all 102 turns; window verified at 16,384 on every turn; largest prompt 14,997 tokens (server count); smallest observed margin 887 tokens
A2 leak counter 0 of 105 stored narrator replies
M04 (unchanged) recovered_through_state_only
Independent verdict precondition_failed:absent_from_summary
Precondition Result First failure
planting turn outside recent history held: the history window started at depth 72 —
absent from authoritative state held: the document, every snapshot, the recall state section —
absent from imported knowledge held (3 sources) —
absent from later narration failed: the narrator names the fact at depths 10, 14, 16, 18, 20, 22, 24, 26, 38, 44 and later accepted turn 5
absent from the active summary failed accepted turn 26

The history window at recall held narration restating it too. The narrator took up the planted detail as a motif, exactly as in B.1, and the summariser folded it in. The run is neither a pass nor a failure for independent retention.

Stages (mechanism evidence only, not acceptance evidence):

Stage Result B.1 attempt, same harness, v1.0.0 memory
created yes: memory 1, depths 0-5; 493-token block, fact in the excerpt yes (memory 1, 688-token block)
retained yes: active, used 31 times; 33 active of capacity 80 yes
ranked yes, rank 1 of 19: semantic 0.7754, lexical 0.2194, final 0.8083; top-k 4 no: rank 10 of 19, cosine 0.866
injected yes: memory 1 is in the recall turn's memories.used [1, 10, 28, 5] and its text is in the memories section no (a later restatement's memory carried it)

The recall turn stored the query the diagnostic recomputes:

  • Input: "> I ask Mara quietly where she hid the amber sundial."
  • Context: the scene (Aldric and Mara in the Crooked Lantern) and the end of the last narration.
  • Scores: memory 1 was also first for the paraphrase (0.7509) and 9th, not selected, for the unrelated question (0.6423).
  • Duplicates: five later memories restating the fact (7, 9, 11, 26, 29) were suppressed as duplicates of memory 1, so the planting-era statement is the one whose provenance survived.

One discrepancy, explained. The replica selected [1, 10, 28, 33] against the stored [1, 10, 28, 5]. The queries are identical. Memory 33 (depths 111-116) was written by the previous turn's post-turn pass at 02:22:27, after the recall turn's retrieval and 9 s before that turn was saved. The turn therefore considered 18 memories, and the diagnosis, run on the finished database, ranks 19. The rows common to both match to rounding (memory 1: 0.7754 / 0.219 / 0.8082 stored). This is the bank changing between retrieval and diagnosis, not a scoring difference.

Found in passing (not a defect of the result): when the state's scene summary already ends in a full stop, the scene text reads "rain outside..". Only the embedder sees it (§R).

L.2 Attempt 2 — ISOLATION VALID, NOT RECOVERED (creation)

Evidence: $HOME/v11-evidence/b2/real-2/, $HOME/v11-evidence/b2/real-2-diagnosis/.

Run complete: 102 accepted turns, 3 restarts (4 process starts), 12 summaries, 18 eligible memories at recall, 1,284 s, median 10.4 s a turn
A1 accounting fits on all 102 turns; window verified at 16,384 on every turn; largest prompt 14,992 tokens; smallest observed margin 892 tokens
A2 leak counter 0 of 105 stored narrator replies
M04 (unchanged) recovered_through_state_only
Independent verdict not_recovered:not_created

Every independent-memory precondition held, on every turn:

Precondition Result
planting turn outside recent history held: the window at recall started at depth 72; planted at 3; no fact text in the history sections
absent from authoritative state held: the document, every node's snapshot, the recall state section
absent from the active summary held, on every turn the harness checked and at recall
absent from imported knowledge held (3 sources)
absent from later narration held: no narrator turn after depth 5 names the fact

This is the first real run, in B.1 or B.2, in which memory was the only layer that could have carried the fact. It is therefore acceptance evidence, and it did not recover the fact.

Stages:

Stage Result
created no. Memory 1 covers the planting block, depths 0-5, 623 tokens. The block fitted the budget, so the summariser was sent all of it (memory_excerpt returned the block unchanged), and the fact was in it at depth 3 as the player's own action: "> I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one." The memory the model wrote does not name the sundial or the teapot. No other memory in the bank does either
retained memory 1 active (used 62 times) — but it does not carry F
ranked memory 1 ranked into the top 4 at recall (4th: semantic 0.6934, lexical 0.1463 from "Mara", final 0.7153) — but it does not carry F
injected memory 1 was injected (memories.used [30, 32, 9, 1]) — without F. The recall reply does not name the fact

What the summariser wrote instead (memory 1, verbatim, 102 words against the prompt's 50-word ceiling):

Aldric sits in the Crooked Lantern with a silver key in his pocket, no-one to give it to. Mara observed him quietly, her eyes unreadable. Edrin finished his ale and stood, asking if they should see what the crypt beneath the Old Abbey holds. Aldric decided to go, promising not to tell anyone. They walked to the Old Abbey's grounds, the crypt door ajar, inviting and foreboding. Inside, the air was musty and cold, with the smell of damp and decay. They advanced cautiously, each aware of potential dangers. The silver key, the key to the crypt, felt heavy in Aldric's hands.

Three things are visible in it:

  1. The fact was dropped, but a fragment of its sentence survived: "promising not to tell anyone", now given to Aldric rather than asked by Mara.

    Correction (B2.4 review, §T.3): in the source Aldric is the one who promises; Mara asks him to. The memory's fault is what the promise is attached to. It ties the promise to Aldric's decision to go to the crypt, and drops Mara and the object it was about. The phrase may also echo Edrin's own "I'll tell no one what we found" in the depth-4 reply.

  2. The block's narration was followed and the player's line was not. The narrator's reply at depth 4 ignored the sundial and moved to the crypt, and the memory follows the narration: walking to the abbey, the crypt door, entering the crypt.

    Correction (B2.4 review, §T.3): an earlier version of this item said those events were past the block. They are not; the depth-4 reply narrates them. The memory invented nothing. It chose the narration over the player's line.

  3. The length ceiling was ignored, which the prompt states but nothing enforces (B.1 §B.4).

First failing stage in the qualifying run: creation, caused by what the model chose to write, not by the excerpt (the whole block was sent), eviction or ranking.

L.3 Both attempts

Attempt Preconditions created retained ranked injected Verdict
1 failed (later narration from turn 5, summary from turn 26) yes (memory 1) yes yes, rank 1 of 19 yes precondition_failed:absent_from_summary — neither pass nor fail
2 all held no (memory 1 omits F) (memory 1 yes) (memory 1 4th of the 4 used) (memory 1 yes, without F) not_recovered:not_created — a failure

The two-attempt limit is reached, and no third attempt was made. The planted fact, prompts and harness were not changed between attempts.

Inference use ended with the attempt-2 diagnosis.


M. Context-budget / A1 interaction

  • The memories section is unchanged in shape and pricing. It is the same live section, built from the same memory_top_k rows with the same line format, and priced into protected context as before (F03). B.2 changes which memories fill it, not how many or how they are rendered. Deterministically it held 67-99 tokens across the scenarios.
  • The query costs no prompt tokens. It is embedded, never sent to the narrator. It is recorded in the snapshot as provenance only.
  • A1 is untouched. The safety reserve, the transport cost, the accounting states and the cold-model load are not in the diff. The summariser's input is still bounded at 2,000 tokens of story plus the cast brief, so its request is no larger than before.
  • Embedding requests. One call per turn, as before; it now carries two short texts (at most about 200 and 180 tokens) instead of one of up to 600.
  • On the real model, every counted turn of both attempts was fits, with the window verified at 16,384 and prompts of at most 14,997 tokens (smallest observed margin 887). The memories section peaked at 716 tokens (attempt
    1. and 1,246 tokens (attempt 2), against 67-99 deterministically. The difference is the real summariser's memory length: attempt 2's planting memory alone was 102 words, twice the prompt's ceiling (§R 12). It stayed inside the section's budget and the prompt inside the window.

N. Protocol-leak / A2 observations

  • WP-B does not modify A2. narrative/extract.py, the extractor rules and the protocol cleanup are not in the diff.
  • The B.1 observation stands, unchanged. In B.1's real run, one stored narrator reply of 105 (action 153, depth 143) echoed application instructions mid-reply, with story after them, outside A2's trailing cleanup region. The evidence is preserved in $HOME/v11-evidence/b1/real-1/campaign.db (V1.1-WP-B1-REPORT.md §K). It is a separate v1.1 follow-up item, not WP-B's.
  • This report does not claim the leak count is zero. The counts from B.2's real-model attempts are in §L. The integrated v1.1 release gate still needs its own protocol-leak result.

O. Compatibility

Area Result
Schema migration none; no model or migration file in the diff
Bundle format v3, unchanged; no bundle code in the diff
Existing memory rows unchanged on open: the same 33 rows, row for row
Existing summaries unchanged: 12 rows, row for row
Existing history unchanged: 207 actions, 4 branches, 32 state events, 2 Save Points, row for row
Existing memories regenerated no. New retrieval and eviction apply only when v1.1 next plays the campaign
Settings memory_top_k, memory_bank_capacity and their defaults unchanged

A real v1.0.0 database. tools/v11_compat_check.py on m04-final/campaign.db (written by 96c1bf5, product code identical to v1.0.0; user_version 94; SHA-256 c74e8798…4f18, unchanged by the runs). Each tree opened its own copy:

Comparison Result
Schema before and after opening, v1.0.0 and B.2 trees identical; user_version 94 → 94 on both
API snapshot (the export bundle, narrative state and events, Save Points, knowledge, memories, derived status, settings) 0 differences
Tables after opening (memories, summaries, actions, state events, Save Points, branches, knowledge sources, adventures) identical, row for row

--exercise on a third copy, endpoint refused: Undo 119 → 117; Redo back to 119; restore On the ridge → 69 with Redo available; context dry run HTTP 200 with canon, state and the applicable summary; export and import 207 → 207 actions, the same head, Save Points 2 → 2, memories 33 → 33, identical state. The dry run's memory retrieval reports the unreachable embedding endpoint and uses no memories, as before.

Evidence: $HOME/v11-evidence/b2/compat/.


P. Offline / security

P.1 Offline and packaging regression

tools/m11_offline.py --out $HOME/v11-evidence/b2/offline, on the complete B.2 tree:

  • the image was built with --no-cache;
  • the container was started with --network none: a loopback interface, no resolver, no route;
  • the exercise was driven inside it over docker exec.

23 passed, 0 failed:

  • no route to the public Internet;
  • no external DNS;
  • first page load offline, and the page names no remote origin;
  • a CSP is served, and every shell asset is served locally;
  • a campaign is created, and state extraction works;
  • a local file imports, and the prompt assembles;
  • imported knowledge is searchable;
  • a turn with no model reachable is reported as a failure, with no narration accepted, the player's words kept, state unchanged and the earlier story intact;
  • export and import work offline, with state, and no secret in the export;
  • the media module imports, a scene packet builds, and no media provider is required;
  • campaigns survive a container restart.

Evidence: offline-report.json, docker-build.log, container.log and inside-stdout.txt in that folder.

As in M11, inference is not exercised offline: the model host is on the trusted LAN, which a container with no network cannot reach.

P.2 What the B.2 diff adds

Check Result
New network client, URL, host, IP, socket or subprocess none: no added line in backend/app or backend/tools names one
External vector database none. Vectors stay in SQLite embedding_blob; the lexical term is computed in process per turn, with no index
Cloud embedding none. Embeddings go through the existing embedding_provider to the configured local endpoint, one call per turn as before
New dependency none. The imports added to product code are math (standard library) and three internal modules: context.builder._encoding (the vendored tokenizer), knowledge.fts (its tokenizer and stop list) and narrative.model (entity_name)
Endpoint policy (endpoints.py, ADR 011), TLS trust (tlstrust.py), providers unchanged; not in the diff
Frontend unchanged; not in the diff
New data leaving the machine none. The embedding request carries the player's input and a short scene text instead of 600 tokens of narration, to the same local endpoint
Real identifiers in committed files none (scanned at staging, §Q)

Local-only operation is unchanged.


Q. Full regression

Run Result
Full backend suite, complete B.2 tree 1,652 passed, 17 skipped, 0 failed, 0 xfailed (1,603 s). The 17 skips are the tests that need a real model, the same 17 as in B.1. B.1's 2 strict xfails are now ordinary passes
test_v11_b2_memory_ranking.py (new) 18 passed
test_v11_b2_memory_eviction.py (new) 24 passed
test_v11_b2_memory_excerpt.py (new) 17 passed
test_v11_b1_memory_diagnostic.py (B.1, updated for B.2) all pass; no xfail marker remains
Memory regression group at the B2.2 checkpoint: test_memory_retrieval, test_memory_nodes, test_context_memory (F01-F08, E01-E04, write-lock and derived-failure recording), test_m11_leakage (all 14), test_m11_long_run_memory, test_memory_rewrite (L04), test_cast_brief, test_memory_settling, test_embedding_model_switch, test_bundle_v2, test_m10_bundle, test_m11_findings, test_v11_b1_long_run_verdict (M04 classifications) 244 passed
I-series bundle and memory tests (test_m9_*, test_bundle_v2, test_m10_bundle, test_process_restart, test_save_points) in the full suite, all pass

Frontend: unchanged. No file under frontend/ is in the diff, so the frontend tests, lint and production build are unaffected and were not rerun. similarity keeps its meaning, so the inspector's display is unchanged.


R. Residual risks

# Risk Why it remains Where it would show
1 No relevance floor. A full memory_top_k is still used whenever the bank holds that many, however unrelated Out of B.2's authorised menu. It costs tokens, not correctness, and a floor is model-dependent Memories section filled with weak matches on a scene unlike anything remembered
2 Weights chosen on a deterministic embedder. INPUT_WEIGHT 0.6 and LEXICAL_WEIGHT 0.15 were swept with the concept embedder, not nomic-embed-text The deterministic fixtures are isolation-valid; a real run is not guaranteed to be (§L). Real cosines sit higher and closer together (B.1: an unrelated query still scored 0.44), so the lexical term may matter more there, or less Real-model ranked stage; the snapshot now records all three scores to measure it
3 A fact in the middle of a very long block is still not shown to the summariser The excerpt is bounded by design. Blocks over 2,000 tokens need about 330-token actions; B.1's real run had 688-token blocks created: no with fact_in_block: yes on a campaign with very long replies
4 Eviction keeps coverage, not any particular fact. An interior memory holding an important fact can still be thinned when its neighbours are close The rule protects the opening and newest stretch, and spreads what remains. It has no notion of importance, which the plan left out of scope (no new field) A campaign past capacity whose key fact sits in a dense middle stretch
5 Coverage ignores branches. A sibling line's memory at the same depth makes an active-line memory look redundant Eviction is adventure-wide by design (v1.0.0 too). Of two memories sharing a start, the less recently used goes, which after a divergence is normally the abandoned one; but an active-line memory unused since before the divergence could lose to a sibling used later Heavily branched campaigns past capacity
6 Real-model isolation is hard to obtain. A narrator told a salient fact tends to restate it, and a summariser keeps "important established facts" Model behaviour, not a memory mechanism. B.2 did not change prompts or summary behaviour to manufacture a qualifying run §L
7 The summariser may still paraphrase the marker in a way the exact-string removal does not catch Only the exact marker is removed; a model rewording it ("part of the story is missing") would store that wording Memory text mentioning an omission; not seen deterministically
8 The mid-reply protocol echo (B.1, action 153) Not WP-B's; A2 was not modified §N; the v1.1 release gate's protocol-leak result
9 Existing campaigns keep memories written from last-2,000-token excerpts Existing memories are not regenerated (plan non-scope). tools/rewrite_memories.py is the opt-in repair Old long blocks in v1.0.0 campaigns
11 The real summariser drops facts stated by the player. In the one isolation-valid real run, the memory of the planting block omitted the planted fact, kept a fragment of its sentence attributed to the wrong character, and followed the narrator's reply instead Not one of the three deficiencies this brief authorised. The plan's bounded menu includes "the memory prompt's instruction to keep named facts and objects", but B.2 was told not to change prompts or summary behaviour. This is the open WP-B failure (§S) §L.2; any campaign whose narrator does not echo a player-stated detail
12 The 50-word memory ceiling is not enforced. Attempt 2's planting memory was 102 words, and described events past its own block Prompt-only since v1.0.0 (B.1 §B.4). A longer memory costs section tokens (the section stayed inside budget, §M) and can mix in later events Memory text lengths in the bank
10 Cosmetic: a doubled full stop in the scene text ("rain outside..") when the state's scene summary already ends in one Found in attempt 1's stored query. Only the embedder reads it; the effect is one stray token. Not changed after the real-model evidence was taken, so that evidence describes the tree as staged memories.query.context in any snapshot whose scene summary ends in punctuation

S. Final WP-B disposition

This section records the owner's final disposition (2026-09-15). It replaces two earlier decision texts that stood before the owner decided. Their evidence is unchanged in §L and §T.

B2.1 RANKING: PASS
B2.2 EVICTION: PASS
B2.3 EXCERPT CREATION: PASS
B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED

DETERMINISTIC WP-B: PASS
REAL-MODEL WP-B: FAIL

WP-B OVERALL:
ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION
independent deterministic recovery:
PASS

reference-model independent recovery:
NOT DEMONSTRATED / FAILED ON THE PRECONDITION-VALID ATTEMPT

What ships, and what each part proved:

  • B2.1 ranking (§C, §D):
    • the planting-era memory moves into the selected top-k (rank 7 and 6 → 1 and 2);
    • the player's current input materially affects retrieval;
    • semantic retrieval stays active, since a paraphrase with no shared word ranks first;
    • the lexical score is inspectable, recorded per used memory.
  • B2.2 coverage-aware eviction (§E, §F):
    • the early memory survives every deterministic over-capacity case;
    • the bank keeps coverage from the opening to the newest stretch;
    • pins and lineage stay safe, and the bank stays bounded.
  • B2.3 bounded head+tail excerpt (§G, §H):
    • a fact at the start of a long block reaches the summariser;
    • late facts stay visible, and the input stays within 2,000 tokens;
    • short blocks are unchanged.
  • Together (§I): independent_full, which fails on v1.0.0 at creation, returns recovered_through_memory_independent. Isolation is asserted on every turn, and provenance resolves to the planting turn. With the B2.4 prompt reverted, this still passes (§T.17).

The real-model limitation, accepted for v1.1. The reference memory summariser, qwen2.5:3b-instruct-16k, can be given the complete relevant source block and still fail to keep a distinctive player-established fact. The failure can be:

  • omitting the fact;
  • omitting its specific objects;
  • attributing it to the wrong character;
  • preferring the generic narration that follows it.

A precondition-valid real attempt failed exactly this way (§L.2). Real-model independent-memory recovery is therefore not guaranteed, even though creation-window, retention, ranking and injection now work under deterministic isolation. This is a failure, not an inconclusive result.

B2.4, rejected.

  • Why it was tried: the precondition-valid attempt failed at memory creation.
  • What it was: a prompt experiment (§T.4).
  • Measured on the reference model (§T.8):
    • target-fact retention on the original failing block: 0 of 5 under both the old and the experimental prompt;
    • it reduced wrong-character attribution (8 → 3 of 40) and invention (2 → 0);
    • it introduced frequent "Memory:" prefixes (29 of 40) and frequent second-person "you" (18 of 40);
    • it produced more over-50-word memories (27 against 13);
    • on one fixture it lost a promise the old prompt always kept (5/5 → 0/5).
  • Conclusion: no reliable net improvement for the target failure. The production prompt is v1.0.0's again (§T.17). B2.4 did not pass.

Scope held. No further prompt variation was tried. Not introduced:

  • a larger summariser model;
  • a second extraction pass or structured fact extraction;
  • a fact database;
  • another LLM call, model routing, or automatic re-summarisation.

Those are future options (§T.18).

Release-gate visibility. V1.1-PLAN.md §11 now requires the v1.1 release report to state the deterministic PASS, the reference-model FAIL, the failing stage (memory creation, content selection) and the owner's acceptance. It must not reduce these to "WP-B passed" (criterion 12). It must also carry the mid-reply A2 echo and the doubled full stop in the scene text as separate residuals (criterion 13).

Nothing is committed, pushed or tagged. WP-C has not started.


T. B2.4 — Memory summariser fidelity for named facts and objects

T.1 Authorisation

The owner's B2.4 brief (2026-09-15), under V1.1-PLAN.md §8 WP-B's bounded menu: "the memory prompt's instruction to keep named facts and objects". No new architectural scope. The trigger is attempt 2 (§L.2): precondition-valid, whole block sent, fact omitted.

Repository at start:

  • HEAD beb17ad, signed (good signature, 02C9BF7D…);
  • B2.1-B2.3 staged: 12 files, +2,525/−183;
  • nothing unstaged; WP-C not started.

T.2 The prompt before B2.4

MEMORY_SYSTEM_PROMPT as shipped in v1.0.0 (184 tokens, verbatim in tools/memory_fidelity.py as BASELINE_MEMORY_PROMPT, checked against git by a test):

# Question Answer
1 What does it tell the model to preserve? "the concrete facts and events (names, places, items, promises, injuries)"; "the details a later scene could turn on — a name, a promise, an injury, where something is"
2 What does it prioritise? Nothing explicitly. "Drop the ones it could not [turn on]" is the only rule of precedence, and it does not say what gives way to what
3 Attribution? Only through pronouns: name the characters rather than "he", "she" or "they". Nothing on keeping an action, promise or possession with its owner
4 People, objects, places, promises, clues? Names, places, items, promises and injuries are listed; "where something is" is named. Clues, discoveries, possession and knowledge are not
5 Invention? Nothing
6 Length? "1-2 plain sentences", "at most 50 words"
7 Does it implicitly favour narrator prose? Yes. It says "The excerpt is written in the second person: 'you' is the protagonist". The protagonist's own lines are marked > and are often first person ("> I watch …", format_player_input leaves "I" alone), and nothing says what they establish is story. In the failed block, that line was 26 of about 460 words; the narration around it was 386
8 Does its wording explain the failure? Plausibly. "Facts and events" with a two-sentence cap favours the event sequence, and the depth-4 reply was the most event-dense text in the block (the decision, the walk to the abbey, entering the crypt). The static, player-stated fact (where the sundial is) had no stated precedence over it

T.3 The observed failure, stated precisely

From attempt 2's stored block and memory (§L.2):

  • The fact was in the block, at depth 3, as the player's line.
  • The memory omitted the sundial and the teapot.
  • It was 102 words against a 50-word target.
  • It followed the narrator's depth-4 reply.

Two corrections to §L.2's first analysis, now marked there:

  • The crypt events are not past the block. The depth-4 reply narrates them, so the memory invented nothing.
  • "promising not to tell anyone" has the right promiser. In the source, Aldric promises Mara. The memory's error is what the promise is attached to: it ties the promise to going to the crypt, and drops Mara and the object. The phrase may also echo Edrin's own "I'll tell no one what we found".

Failure stage: creation. Cause class: the summariser's content selection and fidelity, not the excerpt (the whole 623-token block was sent).

T.4 The prompt correction

One constant changed (memorybank.MEMORY_SYSTEM_PROMPT, 184 → 294 tokens), plus a no-behaviour-change helper, memory_user_prompt(brief, excerpt), factored out of summarize_block so an evaluation sends a model the identical user message. Still one call, one prompt, the same 50-word target, no example, no genre vocabulary. It states:

  • keep concrete facts a later scene could turn on: specific people, objects and places; where something is; who has, hid, found, knows, saw or promised what; injuries, clues and commitments;
  • a distinctive fact comes before mood, scenery, routine movement and small talk: drop those first, and never let a later passage crowd out an earlier fact;
  • keep each fact with the person it belongs to; never move an action, promise, possession, statement or piece of knowledge between characters; never add a fact the excerpt does not state;
  • the narration calls the protagonist "you", and lines beginning > are the protagonist's own actions and words ("I" or "You"). What they establish is story just as the narration is. This is equal standing, not a preference for player text;
  • the framing rule is kept: third person, the protagonist by name, no bare pronouns, no preamble.

No truncation was added. A memory over the target is stored as written (test_an_over_long_memory_is_stored_as_written_never_cut).

Owner disposition: this prompt was rejected and reverted (§T.17). The production prompt is v1.0.0's. The text above records the experiment.

T.5 Deterministic fixtures

tools/memory_fidelity.py holds eight fixtures, each a short block with texture and a fact, declaring the facts to keep, their owners, and forbidden inventions.

Fixture Requirement Genre
object_place_office B2.4-1 object and place office
player_fact_station B2.4-2 player-established fact science-fiction-neutral
promise_contemporary B2.4-3 promise contemporary
attribution_station B2.4-4 attribution (two characters, a code and a badge) science-fiction-neutral
clutter_office B2.4-5 clutter pressure office
no_invention_office B2.4-6 no invention (an unclaimed briefcase) office
multiple_facts_station B2.4-7 several facts, three owners science-fiction-neutral
regression_attempt_2 Part 4, the actual failure the harness campaign

The checker (evaluate) is deterministic and a stated heuristic:

  • a fact is kept when one sentence names all its parts;
  • attribution is the nearest named character before the fact's verb;
  • inventions are patterns anchored on the wrong character as subject;
  • words are counted, not enforced.

test_v11_b2_summarizer_fidelity.py, 49 passed:

  • 12 prompt-contract checks, the 50-word target, genre neutrality with no example or fixture name, framing rules kept, and the baseline equal to the shipped prompt;
  • coverage of every requirement in three genres;
  • every fixture's facts reaching the summariser whole;
  • the checker passing each faithful memory and failing each failure shape (13 unfaithful memories: omission, misattribution, invention);
  • the application sending the corrected prompt with the whole planting block;
  • no truncation;
  • the corrected regression memory ranking first under B2.1 scoring.

B2.4 adds no xfail.

T.6 The actual failed-run regression

regression_attempt_2 is attempt 2's planting block verbatim. It is the harness's own fixture campaign, with no hostname, person or other identifier.

Required after B2.4 Deterministic Reference model, B2.4 prompt (§T.8)
input contains planted fact yes yes
summariser receives planted fact yes (the whole block; test) yes
generated memory contains planted fact checker passes the corrected memory and fails the stored one (test) no: 0 of 5
attribution correct checked not reached (fact absent)
distinctive objects retained checked no
unsupported facts invented no no (0 of 5)

The pre-B2.4 prompt fails the fixture legitimately. The stored memory is the real output, and the old prompt also scored 0 of 5 in §T.8; nothing was fabricated. The B2.4 prompt fails it equally.

T.7 Length and attribution behaviour

40 memories per arm (8 fixtures × 5) Baseline prompt B2.4 prompt
median / p90 / max words 48 / 72 / 82 68 / 83 / 105
over the 50-word target 13 27
misattributed 8 3
invented 2 0

Worst length: 105 words (object_place_office, B2.4, a run-on of four "Memory:" clauses). The length target was measured, not enforced; nothing was cut.

T.8 Real-model comparison

tools/memory_fidelity.py, 2026-09-15 09:26-09:27:

  • Model: qwen2.5:3b-instruct-16k on the GPU inference host, through the application's provider at production temperature (0.3).
  • Samples: 5 per fixture per prompt.
  • Logging: running (§T.11).
  • Evidence: $HOME/v11-evidence/b2/fidelity-b24/fidelity.json, every memory verbatim.
Fixture Baseline passed B2.4 passed What the memories show
object_place_office 5/5 2/5 B2.4 runs began "Memory:"; two lost the drawer and cabinet ("in the archive room"); one run-on of 105 words, whose "Priya … sliding" the checker mis-scored (sliding is not a slide form it matches)
player_fact_station 2/5 4/5 baseline misattributed and invented; B2.4 kept the sample and its owner
promise_contemporary 5/5 0/5 every B2.4 run began "Memory:" and listed only scenery (floorboards, radiator, dog); four dropped the promise, one mentioned the lease without who promised; two used "You". The baseline kept the promise every time
attribution_station 1/5 4/5 baseline gave the code or badge to the wrong person
clutter_office 4/5 5/5
no_invention_office 3/5 3/5
multiple_facts_station 1/5 5/5
regression_attempt_2 0/5 0/5 every memory under both prompts retells the key, the decision and the crypt; none names the sundial or teapot
total 21/40 23/40 B2.4: "Memory:" prefix 29/40, "you" 18/40 (baseline 0 and 0)

Reading it plainly:

  • Attribution and invention improved.
  • The target failure did not move.
  • The longer prompt made this 3B model echo the "Memory:" cue from the user message, drift into second person, write longer, and on one fixture turn the "drop scenery" instruction into a scenery list.
  • This is a model-capability limit at this prompt size, measured, not guessed.

T.9 B2.1, B2.2 and B2.3 with B2.4 in the tree

Suite Result
test_v11_b1_memory_diagnostic.py: all B.1 scenarios, independent_full (recovered_through_memory_independent, isolation every turn, provenance), lineage (G), authority pass
test_v11_b2_memory_ranking.py: B2.1, negative controls, pins pass
test_v11_b2_memory_eviction.py: B2.2, bounds, pins, fallback, lineage separation pass
(the three files together) 81 passed
test_v11_b2_memory_excerpt.py, test_cast_brief.py, test_memory_rewrite.py 53 passed; short-block prompts byte-identical, excerpt split unchanged

The deterministic scenarios use the best-case scripted summariser, so they are independent of the prompt. That they still pass shows B2.4 disturbed no mechanism; it is not evidence about the prompt. No ranking weight, eviction rule or excerpt split changed.

T.10 Full deterministic WP-B

Unchanged from §I on the B2.4 tree:

  • independent_full: created, retained, ranked and injected;
  • recovered_through_memory_independent, with isolation asserted on every one of 52 turns;
  • provenance to depths 0-5.

Lineage (§J) and authority (§K) pass unchanged.

T.11 GPU-host evidence review

For the WP-B runs so far: B.1 on 2026-09-14 18:16-18:37, and B.2 attempts at 22:00-22:22 and 22:23-22:44. Read-only, from the host's logs and its persistent journal.

Check Finding
Power peak 294.5 W (B.2 window), 304.6 W (earlier set), 307 W (dmon). The power limit now reads 200 W (default 280 W), so it was set after these runs. The owner notes spikes above the limit occur, and the dock is rated to 450 W
Power-limit events dmon power-violation samples: 17 of 5,804 in the B.2 set. dmon has no timestamps, so they cannot be placed within the runs
Temperature / throttling peak 68 °C; thermal-violation samples 0
PCIe link Gen3 x4 under load in 2,462 of 2,463 busy samples (one Gen2 sample); PCIe error counter 0
Kernel 4 kernel lines in the whole window; no Xid, NVRM, AER, reset or fallen-off-the-bus
Ollama no crash, exit or panic; models loaded once at 22:00:36-38 and stayed loaded through attempt 2
HTTP errors 9 × 404: the harness's deliberate failed_call probe (no-such-model-m11). 4 × 500: background calls cut off by the harness's planned server restarts (22:18:13 at restart 3; 18:33:13 in B.1). 21:29:54: other traffic outside both windows
Contamination None. No host fault could have altered the WP-B evidence; attempt 2's failure is model output

A logging defect, found and fixed. The documented watch command, journalctl -f -k -u ollama, matches nothing: -k and -u are different fields, which journalctl ANDs. Over the B.2 window it returned 0 lines, where the OR form (_TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service) returned 30,229. Every such log was therefore empty, and events were read from the persistent journal instead.

  • DEVELOPMENT.md now documents the OR form.
  • For the B2.4 comparison the owner's watch still ran the old command, and its file is again empty. The journal holds 2,091 Ollama lines for 09:25-09:28.

T.12 Fresh real-model attempts

0. The comparison (§T.8) ran first and showed the B2.4 prompt failing on the exact planting block the long run plants. Under the brief ("If the summarizer still fails … then stop. Do not create B2.5 automatically"), the long attempts were not run: they could only have confirmed a known creation failure, or failed isolation.

T.13 Context-budget interaction

  • Memory prompt: 184 → 294 tokens (+110) in the background memory call only. The narrator's prompt, A1's reserve, the transport cost and the accounting are unchanged.
  • Typical memory length on the fixtures: median 48 → 68 words; worst 105.
  • Real-model turns after B2.4: none, so no turn was exceeded or truncation_suspected, and A1's protected budget was never at issue.
  • Before B2.4: every counted turn of both B.2 attempts was fits, with the memories section at most 1,246 tokens (§M).
  • If the B2.4 prompt were kept, its longer memories would grow the memories section roughly in proportion (median +40%). It is still priced as protected context, and A1 would refuse rather than overflow.

T.14 Offline regression

tools/m11_offline.py --out $HOME/v11-evidence/b2/offline-b24, on the B2.4 tree: a fresh --no-cache image in a --network none container, driven over docker exec. 23 passed, 0 failed, the same checks as §P.1:

  • no route and no external DNS;
  • the page loads offline, with a CSP and local assets;
  • a campaign is created, state is extracted, and a file imports;
  • the prompt assembles, and knowledge is searchable;
  • a turn with no model is a reported failure that loses nothing;
  • export and import work, with no secret in the export;
  • the media module stays inert;
  • campaigns survive a restart.

What B2.4 changes on the network:

  • Nothing. The diff is a prompt string, a pure helper, an evaluation tool, tests and documentation.
  • No new destination, dependency or external service. tools/memory_fidelity.py calls only the --endpoint it is given, through the application's own provider, and it is a developer tool that production never imports.
  • Unchanged: the endpoint policy, the trusted-LAN policy, embeddings (local), and the vector store (SQLite).

T.15 v1.0.0 compatibility

Rerun on the B2.4 tree against m04-final/campaign.db (SHA-256 prefix c74e8798b68d2e22, unchanged by the runs), from a fresh 432f041 worktree:

Result
Schema before and after opening identical to v1.0.0; user_version 94 → 94
API snapshot identical
memories (33), summaries (12), actions (207), state events (32), Save Points (2), branches (4), knowledge sources (3), adventures (1) identical, row for row
Exercise export and import 207 → 207 actions, same head, same narrative state
Bundle format v3, unchanged; no memory regenerated

T.15a Full backend regression

Full backend suite, B2.4 tree: 1,701 passed, 17 skipped, 0 failed, 0 xfailed (1,101 s).

  • The count is B2.1-B2.3's 1,652 plus B2.4's 49 fidelity tests.
  • The 17 skips are the same real-model tests as before.
  • The suite covers: all WP-B tests, all B.1 diagnostic tests, B2.1-B2.4, F01-F08 and E01-E04, the M04 classifications, test_context_memory, test_memory_nodes, test_memory_retrieval, test_m11_leakage, test_m11_long_run_memory, test_memory_rewrite, the I-series bundle and memory tests, derived-work failure recording and write-lock regressions.

Frontend: unchanged. No file under frontend/ is in the WP-B diff, so the frontend tests, lint and build were not rerun.

T.16 Residual risks added by B2.4

# Risk
13 The reference summariser does not keep a player-stated fact against event-dense narration, under either prompt (0 of 10 on the attempt-2 block). This is the open WP-B limitation
14 The B2.4 prompt regresses framing on a 3B model — resolved by reverting it (§T.17). Kept as evidence of why it was rejected
15 The fidelity checker is a heuristic. It missed one correct attribution (sliding), and may miss paraphrases. Every memory is kept verbatim so a reader can check it
16 Operator logging: the documented kernel/Ollama watch recorded nothing on every run before the fix. The journal is the source of those events
11-12 (earlier) still open
8 (A2 mid-reply echo) and 10 (doubled full stop in the scene text) unchanged, and still out of scope

T.17 Owner disposition, and the revert

Owner decision, 2026-09-15: option 1(a).

  • B2.1, B2.2 and B2.3 are accepted.
  • The B2.4 prompt does not ship.
  • Deterministic WP-B is accepted.
  • The real-model failure is accepted as a documented v1.1 residual limitation.
Reverted Kept
memorybank.MEMORY_SYSTEM_PROMPT, restored byte-for-byte to beb17ad's (184 tokens), and the B2.4 comment above it B2.1 retrieval and ranking, B2.2 eviction, B2.3 head+tail excerpt
the tests that required the experimental prompt's wording memorybank.memory_user_prompt: a behaviour-neutral helper summarize_block calls, so a measurement sends a model the application's exact message
the B2.4 note in CONTEXT-AND-MEMORY.md §15, replaced by the shipped behaviour and the limitation B.1 and B.2 diagnostics; tools/memory_fidelity.py, now diagnostic-only, comparing the shipped prompt with B24_EXPERIMENT_PROMPT, which it keeps so the experiment can be repeated
this section's evidence, the fixtures, and every measured memory

The tests after the revert (test_v11_b2_summarizer_fidelity.py) are split in two.

  • Acceptance gates the tree:
    • the shipped prompt equals v1.0.0's, checked against git;
    • the experiment is not what ships;
    • every fixture reaches the summariser whole;
    • the application sends the shipped prompt with the whole planting block;
    • an over-long memory is never cut;
    • a faithful regression memory ranks first under B2.1.
  • Diagnostic measurement proves only the instrument, on hand-written memories with known answers:
    • fact retention, attribution and invention;
    • word count, a leading "Memory:" and second-person "you";
    • promise retention.

No reference-model score is a test gate.

Verification after the revert:

  • The runtime constant equals beb17ad's value.
  • The only prompt-related lines left in the diff against beb17ad are the summarize_block call through memory_user_prompt.
  • The fidelity, cast-brief, rewrite and excerpt tests: 94 passed.

Final regression, offline and compatibility, on the final staged tree (the B2.4 prompt reverted):

Check Result
Full backend suite 1,693 passed, 17 skipped, 0 failed, 0 xfailed (1,227 s). This is B2.1-B2.3's 1,652 plus 41 summariser tests; the 8 that required the experimental prompt's wording went with it. The 17 skips are the same real-model tests
Deterministic WP-B within it B2.1 ranking, B2.2 eviction, B2.3 excerpt; independent_full still recovered_through_memory_independent, with isolation every turn and provenance; lineage (G); authority; the B.1 diagnostics; F01-F08, E01-E04, M04, the I-series, derived-work failure recording and write-lock regressions: all pass
Offline (tools/m11_offline.py, fresh --no-cache image, --network none) 23 passed, 0 failed. Rerun because the runtime differs from §T.14's tree
v1.0.0 compatibility (m04-final/campaign.db, compared with the 432f041 snapshot taken the same day) API snapshot and schema identical; user_version 94 → 94; memories (33), summaries (12), actions (207), state events (32), Save Points (2), branches (4), knowledge sources (3) and the adventure identical, row for row; source database unchanged
Frontend unchanged; not in the diff
deterministic WP-B: PASS
offline regression: PASS (23/23)
v1.0.0 compatibility: unchanged

T.18 Future memory-quality options (not implemented)

Each of these is new scope for a future package, and none is part of v1.1:

  • A stronger dedicated summariser model, evaluated with tools/memory_fidelity.py against the same fixtures.
  • Structured fact extraction alongside the prose memory: who has, hid, knows or promised what.
  • Separate factual and narrative memory, each retrieved on its own terms.
  • Model-specific summariser recommendations in the operator documentation, once measured.