Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
70 KiB
v1.1 WP-B.2 — Independent Long-Term Memory Retention
Status: COMPLETE — accepted by the owner with a documented real-model limitation (2026-09-15, option 1(a)). Staged, uncommitted, for the owner's signed commit.
- Shipped: B2.1 ranking, B2.2 eviction and B2.3 excerpt creation.
- Not shipped: B2.4, a memory-prompt experiment. It did not correct the reference model's creation failure, and its prompt was reverted (§T).
- The disposition is in §S.
A. Repository baseline
| Branch | v1.1-development |
| HEAD at start | beb17ad — v1.1 WP-B.1: diagnose independent long-term memory retention, signed by the owner (good signature, RSA key 02C9BF7D…) |
| Its parents | d63804f (WP-A1/A2, signed), ac465ed (plan v4.1), 432f041 (v1.0.0, tag v1.0.0) |
| Working tree at start | clean |
| Comparison baseline | 432f041 (v1.0.0), from a throwaway worktree. The B.1 diagnostic tool, which is not in that tree, was copied into it from beb17ad, with only the new scenario definitions and the coverage trace added. app/ in the worktree is v1.0.0's, unmodified |
B. B.1 findings being corrected
WP-B.1 (V1.1-WP-B1-REPORT.md §L, §M) placed three deficiencies. They are the
only memory behaviour this package changes.
| # | Stage | B.1 evidence | Corrected in |
|---|---|---|---|
| 1 | Ranking | Real 100-turn run: the planting-era memory was created, active and carried F, and ranked 10th of 19 (cosine 0.866) against memory_top_k 4. The query was the newest four actions cut to 600 tokens, so the one-line question was diluted by narration (deterministically, 0.708 alone fell to 0.241 in the query). That run failed isolation, so it was diagnostic evidence only |
B2.1 |
| 2 | Retention | Past memory_bank_capacity, least-recently-used eviction removed the planting-era memory first (turns 21, 21, 36 in the capacity fixtures). Strict xfail test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity |
B2.2 |
| 3 | Creation | The summariser saw only a block's last 2,000 tokens; a fact at the start of a 2,079-token block never reached it (not_created). Strict xfail test_acceptance_a_fact_early_in_a_long_block_is_remembered |
B2.3 |
A fourth B.1 observation, a mid-reply protocol echo, is not WP-B's and is left alone (§N).
C. B2.1 ranking design
C.1 The retrieval query
v1.0.0 embedded one text: the newest four story actions, cut to their last 600
tokens. It is replaced by two short texts embedded in one call
(memorybank.retrieval_query):
| Component | Source | Bound |
|---|---|---|
input |
the newest action, when it is a player action with text (do, say, story) |
last 200 tokens (QUERY_INPUT_TOKENS) |
context |
the scene from the authoritative state — scene.summary, the location's entity name, the names of scene.present (at most 8) — then the end of the newest narration |
60 tokens (QUERY_SCENE_TOKENS) + 120 tokens (QUERY_NARRATION_TOKENS) |
- A continue turn, and the Insights dry run, have no player input: the context alone is searched, and the lexical score is 0.
- A retry excludes the attempt being replaced, as before; the query is the retried input and the narration before it.
- Why the context stays.
CONTEXT-AND-MEMORY.md§18 requires more than raw input, and a question often means nothing without its scene ("I ask her what she keeps up there"). It is bounded so it can resolve a reference but cannot outweigh the question by length. The full entity list, threads and older narration are left out on purpose. - Provenance. The query is recorded in each turn's snapshot as
memories.query:input,context,input_terms,input_weight,lexical_weight.
C.2 The ranking formula
semantic_score = INPUT_WEIGHT * cos(input, m) + (1 - INPUT_WEIGHT) * cos(context, m)
INPUT_WEIGHT = 0.6; either cosine alone if the other text is empty
lexical_score = Σ w(t) for t in input_terms ∩ terms(m) / Σ w(t) for t in input_terms
w(t) = ln((N + 1) / (df(t) + 1)) N eligible candidates, df holding t
final_score = semantic_score + LEXICAL_WEIGHT * lexical_score, LEXICAL_WEIGHT = 0.15
- Terms (
memorybank.lexical_terms): imported knowledge's tokenizer and stop list (knowledge.fts.terms), a possessive'sdropped, and a trailingsfolded from words of five letters or more not ending inss. No stemmer, no dependency. - Rarity is computed per call over the eligible candidate set only. Nothing is indexed or stored. A word every candidate holds weighs exactly 0; a word none holds weighs most, so the unmatched part of a question dilutes every candidate equally.
- Range.
lexical_scoreis in [0, 1], so wording can move a memory by at most 0.15 (test_a_lexical_match_moves_a_memory_at_most_the_lexical_weight). - Only the player's input is matched lexically, never the context.
- Pins are exactly as before: always used, counted toward
memory_top_k. - Ties on the final score are broken by memory id, so row order never decides.
- Redundancy suppression (§22, cosine ≥ 0.93, never across authority) is unchanged and runs over the final order.
- Reported per used memory:
semantic_score,lexical_score,final_score.similaritykeeps its v1.0.0 meaning, the semantic score, so the inspector's "closeness" needs no frontend change.
Egress. Memory text is read for ranking only on a turn whose input has
terms, only for memories not already held, and held beside the vectors in a
cache with the same two rules: set_vector drops an entry (an edit clears the
vector, so a changed text is re-read), and the next read discards memories no
longer in the catalogue. A continue turn reads no memory text beyond the top-k
detail read, and a second turn in a row reads none again
(test_memory_text_is_read_once_and_then_held).
C.3 Choosing the lexical weight
Swept over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5 on three deterministic fixtures,
each ranked two ways: with the new query, and with the v1.0.0 query plus the
lexical term (the lexical term's own contribution, with the crowding left in).
Rank of the planting-era memory F, memory_top_k 4:
| Weight | 0 | 0.05 | 0.1 | 0.15 | 0.2 | 0.3 | 0.5 |
|---|---|---|---|---|---|---|---|
ranking_crowded, new query |
1 | 1 | 1 | 1 | 1 | 1 | 1 |
ranking_crowded, v1.0.0 query + lexical |
7 | 6 | 6 | 1 | 1 | 1 | 1 |
ranking_crowded, unrelated question: F selected? |
no (9) | no (9) | no (6) | no (6) | no (6) | no (5) | yes (3) |
| negative control (paraphrase + a decoy's rare place): decoy above F? | no | no | no | no | no | no | no |
- 0.15 is the smallest weight at which the lexical term alone rescues F from the crowded v1.0.0 query.
- 0.5 already lets an incidental shared word ("landing") select F for an unrelated question, which is the failure a larger weight buys.
- The negative control never put the decoy above F at any weight with the new query.
- Not tuned to one phrase: the direct question, a paraphrase sharing only "Mara", a context-dependent question sharing no word at all, a decoy word from a different memory, and an unrelated question all ran at every weight.
One normalisation was rejected along the way. Taking the share over only the input words some candidate holds let a lone incidental match score 1.0, and in the context fixture a decoy holding "fish stalls" then outranked F from weight 0.15. The share is now over all the input's words.
D. B2.1 before/after evidence
Fixtures (tools/memory_diagnostic.py, best-case summariser, concept
embedder, isolation valid, below capacity, recall at depth 106):
ranking_crowded: 150-word narration,memory_top_k4 — B.1's real-run failure made deterministic.ranking_context_dependent: the same, with the last narration putting Mara at the tavern's top shelf by an old kettle, and the question "I ask her what she keeps up there." Neither "kettle" nor "shelf" is one of F's leak terms.
| Fixture | v1.0.0 (432f041) |
B2.1 |
|---|---|---|
ranking_crowded |
retained_but_not_ranked: rank 7 of 17, similarity 0.1653 |
injected: rank 1 of 17; semantic 0.4671, lexical 0.5328, final 0.5470; selected [1, 7, 11, 12]; 87 tokens |
ranking_context_dependent |
retained_but_not_ranked: rank 6 of 17, similarity 0.1577 |
injected: rank 2 of 17; semantic 0.1866, lexical 0.0, final 0.1866; selected [1, 5, 6, 9]; 89 tokens |
independent_default (top-k 5) |
injected: rank 2, similarity 0.241 |
injected: rank 1; semantic 0.4739, lexical 0.5328, final 0.5538 |
In every row the diagnostic's replica of the ranking (production's own
score_candidates and select_memories) matched the recall turn's stored
selection.
The required ranking controls:
| Control | Result |
|---|---|
| Direct question | F rank 1 in every fixture (above) |
| Semantic paraphrase ("the little brass dial that tells the hour") | rank 1, lexical 0.1125 (only "Mara" shared) against 0.5328 for the direct question: found by meaning |
| Context-dependent wording | rank 2 with the scene; rank 17 of 17 with the context removed; lexical 0 |
| Negative control (paraphrase + "out by the fish stalls", held only by a filler memory) | the decoy's lexical score (0.18) exceeds F's (0.09), and F still ranks 1, the decoy 2-5 |
| Common words ("The travellers spent time.") | lexical 0.0 for every memory |
| Pins | always used, counted toward top-k (test_a_pin_is_still_always_used_and_counts_toward_top_k, and v1.0.0's test_pinned_memories_are_always_used, unchanged) |
| No relevance floor (unchanged) | an unrelated question still selects a full top-k; F is no longer carried along by it (rank 6 of 17, not selected) |
Tests at the checkpoint: test_v11_b2_memory_ranking.py 18 passed;
test_v11_b1_memory_diagnostic.py 31 passed and 2 xfailed (the eviction and
creation criteria, not yet fixed); the memory regression group (13 files) 226
passed.
B2.1 RANKING: PASS
E. B2.2 eviction design
memorybank.eviction_order(rows, limit), run by _evict_over_capacity after
each post-turn pass. It is a pure function of seven columns per active memory
and reads neither text nor vectors (test_eviction_reads_no_vectors, unchanged).
| Part | Rule |
|---|---|
| Coverage signal | Memories with a source range are ordered by position. A memory is judged by the hole its removal would leave: the start of the next memory, less the furthest end before it, less one. Smallest hole first, so the bank thins where it is densest. A memory whose source_start another active memory shares leaves no hole (0) |
| Boundaries | The earliest and the latest memory by position are not coverage candidates: they alone describe the opening and the newest stretch |
| Recency signal | Among equal holes: coalesce(last_used_at, created_at) oldest first |
| Tie-breaking | then use_count lowest first, then id lowest first |
| Pinned rows | never taken; they count as coverage for their neighbours; if every active memory is pinned the bank may exceed capacity, as before |
| Fallback | once no coverage candidate remains — memories typed by the player or migrated have no range, and a bank can be all boundaries — the rest are taken by the v1.0.0 order (recency, use count) plus id |
| Recomputation | after every pick, since removing a memory widens its neighbours' holes |
- Unchanged: the capacity setting and default (80), the count over the whole
adventure (not one lineage), marking
forgottenrather than deleting, running once per post-turn pass. - Frozen bank. A memory written this turn is the latest boundary, so it is
not a coverage candidate, and in the fallback it has the newest timestamp.
It can still go first legitimately when it only repeats a stretch another,
more recently used memory describes
(
test_a_newborn_can_be_the_legitimate_first_to_go). - Lineage. The rule decides
forgottenand nothing else; eligibility is the retrieval clause, untouched. A sibling line's memory at the same depth shares a start, so it is the first kind thinned, and the least recently used of the two goes — after a divergence, that is the abandoned line's. - No new field. Existing
source_start,source_end,last_used_at,created_atanduse_countonly.
F. B2.2 capacity evidence
Deterministic fixtures (v1.0.0 from the worktree, with a coverage trace added to its copy of the tool):
| Fixture | v1.0.0 | B2.2 |
|---|---|---|
past_capacity (capacity 6, top-k 5, 17 memories written) |
created_but_evicted: F evicted at turn 21, the first eviction. Final bank covers depths 36-101 |
injected: F retained, rank 1 of 6. Bank covers 0-101, largest uncovered stretch 18 |
past_capacity_pinned |
evicted at turn 21; bank 6-101 | injected; the pinned memory kept; 0-101, gap 18 |
past_capacity_low_top_k (capacity 8, top-k 2) |
evicted at turn 36 (first eviction 27); bank 18-101 | injected: rank 1 of 8; 0-101, gap 12 |
In all three: active memories never exceeded capacity after a pass, and no memory was evicted by the pass that created it.
F survives these fixtures as the earliest boundary. That is the general rule, not a special case — any campaign's opening memory is kept while interior memories remain — but it is also the easiest case, so the rule was tested on interior memories in general. Memories arrive a block at a time, nothing is ever retrieved, and the bank is held at capacity. Largest uncovered stretch, counting from depth 0:
| Bank | Coverage-first: first memory kept / largest gap | v1.0.0 order: first memory kept | Average spacing (span / capacity) |
|---|---|---|---|
| 17 blocks, capacity 6 | 0 / 18 | depth 66 | 17.0 |
| 60 blocks, capacity 10 | 0 / 42 | depth 300 | 36.0 |
| 500 blocks, capacity 80 | 0 / 42 | depth 2,520 | 37.5 |
| 500 irregular blocks (4-9 actions), capacity 80 | 0 / 47 | depth 2,591 | 38.4 |
The largest gap stays within about 1.3 times the average spacing. The v1.0.0
order keeps an unbroken run of the newest blocks and nothing before it.
test_a_long_bank_keeps_describing_the_whole_story asserts the opening kept and
a gap of at most twice the average spacing.
Invariants tested (test_v11_b2_memory_eviction.py, 24 tests):
| Invariant | Test |
|---|---|
| Active count ≤ capacity after maintenance (or = the pins, if more) | test_capacity_holds_and_pins_survive_for_any_bank (8 random banks) |
| Pinned never evicted | same, and test_pins_are_never_taken_but_still_count_as_coverage |
| Abandoned-line memories not made eligible; lineage separate | test_eviction_does_not_make_an_abandoned_lines_memory_eligible, test_the_pass_changes_nothing_but_forgotten |
| Frozen bank still works | test_the_newest_memory_is_not_evicted_by_the_bank_it_joins; v1.0.0's test_a_newborn_is_not_evicted_by_the_bank_it_joins and test_a_full_bank_still_turns_over, unchanged |
| A newborn may still go first when legitimate | test_a_newborn_can_be_the_legitimate_first_to_go |
| Independent of row order | test_the_order_does_not_depend_on_row_order, test_a_tie_on_every_signal_is_broken_by_id |
| Fallback is v1.0.0's order | test_memories_without_a_range_take_the_least_recently_used_fallback; v1.0.0's test_eviction_marks_the_least_recently_used and test_eviction_breaks_ties_on_use_count, unchanged |
The B.1 strict xfail for capacity is now an ordinary test that passes for all three capacity fixtures. B.1's two tests that asserted F's eviction (a diagnosis of v1.0.0) now assert its retention.
B2.2 EVICTION: PASS
G. B2.3 creation design
memorybank.memory_excerpt(raw, budget=MEMORY_EXCERPT_TOKENS), used by
summarize_block (and so also by the opt-in tools/rewrite_memories.py).
| Block | Summariser input |
|---|---|
| ≤ 2,000 tokens | the whole block, unchanged. The prompt is byte-identical to v1.0.0's (test_a_short_block_is_prompted_exactly_as_before) |
| > 2,000 tokens | the block's first tokens, then \n\n[… the middle of this stretch of story is left out here …]\n\n, then its last tokens |
- Accounting (
excerpt_split). The marker with its blank lines costs 15 tokens and is paid first. The remaining 1,985 are halved, the odd token to the end: 992 + 993 + 15 = 2,000. - Seams. Two decoded runs rejoined can tokenise differently where they meet, so the result is measured and the opening gives up tokens until the whole is within 2,000. Measured alone, a cut part can come to one token over its run; the whole never exceeds the budget, including for mixed-script and emoji text.
- Order and marker. The opening comes first, and the marker tells the
summariser the two parts are not adjacent. The header (
Story excerpt:) is unchanged. - The marker is never stored.
summarize_blockremoves it from the model's reply. - The budget did not grow. No model context,
MEMORY_MAX_WORDSor prompt instruction changed. Existing memories were not rewritten. - Limit. A fact in the middle of a very long block is still left out
(
test_a_fact_in_the_middle_of_a_very_long_block_is_still_omitted).
H. B2.3 long-block evidence
| Fixture | v1.0.0 | B2.3 |
|---|---|---|
long_block_fact_early (F at depth 1 of a 2,079-token block) |
fact in block: yes; in the summariser's excerpt: no; not_created |
in excerpt: yes; memory 1 (depths 0-5) carries F; injected |
long_block_fact_late (F at depth 5 of the same-sized block) |
in excerpt: yes; created | in excerpt: yes; created; injected |
These two fixtures test creation. They recall at depth 22, so the planting turn is still inside the history window and their isolation check is not expected to pass; independence is §I's.
Tests (test_v11_b2_memory_excerpt.py, 17):
| Required | Test |
|---|---|
| F early in a > 2,000-token block | test_a_long_blocks_memory_keeps_its_source_provenance (the memory holds it); B.1's test_acceptance_a_fact_early_in_a_long_block_is_remembered, now ordinary |
| F late in the same-sized block | B.1's test_the_same_fact_late_in_the_same_sized_block_does |
| A block ≤ 2,000 tokens unchanged | test_a_block_that_fits_is_sent_whole_and_unchanged, test_a_block_of_exactly_the_budget_is_unchanged, test_a_short_block_is_prompted_exactly_as_before |
| Head/tail accounting bounded | test_the_split_is_even_and_documented, test_the_excerpt_never_exceeds_the_budget (5 sizes up to 10× the budget), test_the_budget_holds_for_awkward_text |
| The marker is not stored as story fact | test_the_marker_is_never_stored_as_part_of_a_memory (a summariser that echoes its whole input) |
| Valid source provenance | test_a_long_blocks_memory_keeps_its_source_provenance: source_start/source_end are the block's depths, branch_id/depth its last node's |
The B.1 strict xfail for creation is now an ordinary test.
B2.3 CREATION: PASS
I. Full deterministic independent-memory test
independent_full: all three B.1 failures at once. Every narrator reply is
about 850 words, so every memory block (2,079 tokens) is longer than the
excerpt; narration crowds the query at memory_top_k 4; capacity is 8, and 17
memories are written. F is planted at depth 1 by a story action; recall is at
depth 106. context_token_budget 16,384.
| v1.0.0 | B.2 | |
|---|---|---|
| Verdict | not_created (the fact never reached the summariser) |
injected |
Isolation, asserted on every one of the 52 turns and at recall:
| Check | Result |
|---|---|
| F absent from authoritative state | every turn's document; every node's narrative_state_after; the recall prompt's state section |
| F absent from the summary | the active summary on every turn, and the recall prompt's summary section |
| F absent from imported knowledge | none imported; no knowledge section mentions it |
| F absent from recent history | the history window at recall starts at depth 84; planted at 1; the fact is not in the history sections |
| F absent from later narration | no narrator turn after the planting block (ends depth 5) mentions it |
The four stages:
| Stage | Result |
|---|---|
| created | memory 1, source depths 0-5, text "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." The block was 2,079 tokens and the fact was in the head+tail excerpt |
| retained | active; bank 8 of 17 written, capacity 8; the bank covers depths 0-101 with a largest uncovered stretch of 12; nothing evicted by its own pass |
| ranked | rank 1 of 8, top-k 4; semantic 0.4671, lexical 0.5066, final 0.5430; selected [1, 5, 7, 12]; replica matches the stored selection |
| injected | in the recall turn's memories.used and its used_memories section, 88 tokens |
Provenance. The recall snapshot's entry for memory 1 records source
{branch_id: 1, depth: 5, source_start: 0, source_end: 5} and authority
accepted_story. It equals the row, the range includes the planting depth, and
memorybank.source_block on that row returns depths 0-5, which hold the
planting action.
Long-run verdict. The same measurements, given to
m11_long_run._independent_memory_verdict, return
recovered_through_memory_independent
(test_acceptance_full_is_the_long_run_independent_memory_verdict).
All nine deterministic scenarios, final tree:
| Scenario | Verdict | Isolation | F rank / top-k | Notes |
|---|---|---|---|---|
independent_default |
injected | ok | 1 / 5 | |
past_capacity |
injected | ok | 1 / 5 | capacity 6 |
past_capacity_pinned |
injected | ok | 1 / 5 | pin kept |
past_capacity_low_top_k |
injected | ok | 1 / 2 | capacity 8 |
long_block_fact_early |
injected | not expected (recall depth 22) | 1 / 5 | fact in excerpt |
long_block_fact_late |
injected | not expected (recall depth 22) | 1 / 5 | |
lineage_control |
injected | ok | 1 / 5 | §J |
ranking_crowded |
injected | ok | 1 / 4 | |
ranking_context_dependent |
injected | ok | 2 / 4 | lexical 0 |
independent_full |
injected | ok, every turn | 1 / 4 | all three at once |
J. Lineage
lineage_control (B.1's fixture, rerun on the final tree). Fact G is planted at
turn 21 behind a Save Point; a second Save Point marks line A at turn 30; nine
Undos and a different action abandon line A.
| Requirement | Result |
|---|---|
| The abandoned memory stays stored | G's memory (id 7) is on disk |
| Absent on the active divergent line | not eligible under the active lineage clause |
Absent from memories.used |
no turn after the divergence names it, and G's text is in no memories section |
| Eligible again only when branch semantics permit | restoring the line-A Save Point makes it eligible again |
| F on the active line | still injected |
Eviction cannot change eligibility: it writes only forgotten
(test_the_pass_changes_nothing_but_forgotten,
test_eviction_does_not_make_an_abandoned_lines_memory_eligible).
tree.attach_memory, forget_node, the lineage clause, E02 and summary lineage
(E03) are not in the diff. test_m11_leakage.py passes unchanged (§Q).
K. Authority
B.1's authority control (test_a_memory_that_contradicts_state_loses_and_changes_nothing),
rerun unchanged on the final tree: a state correction establishes "the tavern
lamp is lit"; a pinned memory says "The tavern lamp was never lit that night."
| Requirement | Result |
|---|---|
| State wins | the fact is in the narrative state section, which is read after the memories section |
| Memory remains non-authoritative | it is rendered under "Memories from earlier in the story…", as before |
| No state mutation from retrieval | the state document is identical before and after the turn |
Retrieval still only reads, and a used memory's authority is still recorded.
F07 and classify_authority are not in the diff.
L. Real-model attempts
Setup, identical for both attempts.
- Command:
tools/m11_long_run.py --turns 100 --independent-fact, unchanged since B.1: the same planted fact, prompts and harness. - Host and models: the GPU inference host (Ollama 0.34.0);
qwen2.5:3b-instruct-16knarrating and summarising,nomic-embed-text:latestembedding. - Memory: bank and summaries on; harness settings
memory_top_k4 andcontext_token_budget16,384. - Order of events:
- the full deterministic suite had passed (§Q);
- the owner confirmed the power, link and kernel logging was running on the host before attempt 1;
- attempt 2 was run because attempt 1 failed isolation.
- Cap: two deliberate attempts at most, as the brief set.
Stages come from tools/v11_b1_memory.py diagnose on a copy of each run's
database, ranked with the same embedding model.
L.1 Attempt 1 — PRECONDITION FAILED
Evidence: $HOME/v11-evidence/b2/real-1/, $HOME/v11-evidence/b2/real-1-diagnosis/.
| Run | complete: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 memories written (18 eligible at recall), 1,352 s, median 10.9 s a turn |
| A1 accounting | fits on all 102 turns; window verified at 16,384 on every turn; largest prompt 14,997 tokens (server count); smallest observed margin 887 tokens |
| A2 leak counter | 0 of 105 stored narrator replies |
| M04 (unchanged) | recovered_through_state_only |
| Independent verdict | precondition_failed:absent_from_summary |
| Precondition | Result | First failure |
|---|---|---|
| planting turn outside recent history | held: the history window started at depth 72 | — |
| absent from authoritative state | held: the document, every snapshot, the recall state section | — |
| absent from imported knowledge | held (3 sources) | — |
| absent from later narration | failed: the narrator names the fact at depths 10, 14, 16, 18, 20, 22, 24, 26, 38, 44 and later | accepted turn 5 |
| absent from the active summary | failed | accepted turn 26 |
The history window at recall held narration restating it too. The narrator took up the planted detail as a motif, exactly as in B.1, and the summariser folded it in. The run is neither a pass nor a failure for independent retention.
Stages (mechanism evidence only, not acceptance evidence):
| Stage | Result | B.1 attempt, same harness, v1.0.0 memory |
|---|---|---|
| created | yes: memory 1, depths 0-5; 493-token block, fact in the excerpt | yes (memory 1, 688-token block) |
| retained | yes: active, used 31 times; 33 active of capacity 80 | yes |
| ranked | yes, rank 1 of 19: semantic 0.7754, lexical 0.2194, final 0.8083; top-k 4 | no: rank 10 of 19, cosine 0.866 |
| injected | yes: memory 1 is in the recall turn's memories.used [1, 10, 28, 5] and its text is in the memories section |
no (a later restatement's memory carried it) |
The recall turn stored the query the diagnostic recomputes:
- Input: "> I ask Mara quietly where she hid the amber sundial."
- Context: the scene (Aldric and Mara in the Crooked Lantern) and the end of the last narration.
- Scores: memory 1 was also first for the paraphrase (0.7509) and 9th, not selected, for the unrelated question (0.6423).
- Duplicates: five later memories restating the fact (7, 9, 11, 26, 29) were suppressed as duplicates of memory 1, so the planting-era statement is the one whose provenance survived.
One discrepancy, explained. The replica selected [1, 10, 28, 33] against the stored [1, 10, 28, 5]. The queries are identical. Memory 33 (depths 111-116) was written by the previous turn's post-turn pass at 02:22:27, after the recall turn's retrieval and 9 s before that turn was saved. The turn therefore considered 18 memories, and the diagnosis, run on the finished database, ranks 19. The rows common to both match to rounding (memory 1: 0.7754 / 0.219 / 0.8082 stored). This is the bank changing between retrieval and diagnosis, not a scoring difference.
Found in passing (not a defect of the result): when the state's scene summary already ends in a full stop, the scene text reads "rain outside..". Only the embedder sees it (§R).
L.2 Attempt 2 — ISOLATION VALID, NOT RECOVERED (creation)
Evidence: $HOME/v11-evidence/b2/real-2/, $HOME/v11-evidence/b2/real-2-diagnosis/.
| Run | complete: 102 accepted turns, 3 restarts (4 process starts), 12 summaries, 18 eligible memories at recall, 1,284 s, median 10.4 s a turn |
| A1 accounting | fits on all 102 turns; window verified at 16,384 on every turn; largest prompt 14,992 tokens; smallest observed margin 892 tokens |
| A2 leak counter | 0 of 105 stored narrator replies |
| M04 (unchanged) | recovered_through_state_only |
| Independent verdict | not_recovered:not_created |
Every independent-memory precondition held, on every turn:
| Precondition | Result |
|---|---|
| planting turn outside recent history | held: the window at recall started at depth 72; planted at 3; no fact text in the history sections |
| absent from authoritative state | held: the document, every node's snapshot, the recall state section |
| absent from the active summary | held, on every turn the harness checked and at recall |
| absent from imported knowledge | held (3 sources) |
| absent from later narration | held: no narrator turn after depth 5 names the fact |
This is the first real run, in B.1 or B.2, in which memory was the only layer that could have carried the fact. It is therefore acceptance evidence, and it did not recover the fact.
Stages:
| Stage | Result |
|---|---|
| created | no. Memory 1 covers the planting block, depths 0-5, 623 tokens. The block fitted the budget, so the summariser was sent all of it (memory_excerpt returned the block unchanged), and the fact was in it at depth 3 as the player's own action: "> I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one." The memory the model wrote does not name the sundial or the teapot. No other memory in the bank does either |
| retained | memory 1 active (used 62 times) — but it does not carry F |
| ranked | memory 1 ranked into the top 4 at recall (4th: semantic 0.6934, lexical 0.1463 from "Mara", final 0.7153) — but it does not carry F |
| injected | memory 1 was injected (memories.used [30, 32, 9, 1]) — without F. The recall reply does not name the fact |
What the summariser wrote instead (memory 1, verbatim, 102 words against the prompt's 50-word ceiling):
Aldric sits in the Crooked Lantern with a silver key in his pocket, no-one to give it to. Mara observed him quietly, her eyes unreadable. Edrin finished his ale and stood, asking if they should see what the crypt beneath the Old Abbey holds. Aldric decided to go, promising not to tell anyone. They walked to the Old Abbey's grounds, the crypt door ajar, inviting and foreboding. Inside, the air was musty and cold, with the smell of damp and decay. They advanced cautiously, each aware of potential dangers. The silver key, the key to the crypt, felt heavy in Aldric's hands.
Three things are visible in it:
-
The fact was dropped, but a fragment of its sentence survived: "promising not to tell anyone", now given to Aldric rather than asked by Mara.
Correction (B2.4 review, §T.3): in the source Aldric is the one who promises; Mara asks him to. The memory's fault is what the promise is attached to. It ties the promise to Aldric's decision to go to the crypt, and drops Mara and the object it was about. The phrase may also echo Edrin's own "I'll tell no one what we found" in the depth-4 reply.
-
The block's narration was followed and the player's line was not. The narrator's reply at depth 4 ignored the sundial and moved to the crypt, and the memory follows the narration: walking to the abbey, the crypt door, entering the crypt.
Correction (B2.4 review, §T.3): an earlier version of this item said those events were past the block. They are not; the depth-4 reply narrates them. The memory invented nothing. It chose the narration over the player's line.
-
The length ceiling was ignored, which the prompt states but nothing enforces (B.1 §B.4).
First failing stage in the qualifying run: creation, caused by what the model chose to write, not by the excerpt (the whole block was sent), eviction or ranking.
L.3 Both attempts
| Attempt | Preconditions | created | retained | ranked | injected | Verdict |
|---|---|---|---|---|---|---|
| 1 | failed (later narration from turn 5, summary from turn 26) | yes (memory 1) | yes | yes, rank 1 of 19 | yes | precondition_failed:absent_from_summary — neither pass nor fail |
| 2 | all held | no (memory 1 omits F) | (memory 1 yes) | (memory 1 4th of the 4 used) | (memory 1 yes, without F) | not_recovered:not_created — a failure |
The two-attempt limit is reached, and no third attempt was made. The planted fact, prompts and harness were not changed between attempts.
Inference use ended with the attempt-2 diagnosis.
M. Context-budget / A1 interaction
- The memories section is unchanged in shape and pricing. It is the same
live section, built from the same
memory_top_krows with the same line format, and priced into protected context as before (F03). B.2 changes which memories fill it, not how many or how they are rendered. Deterministically it held 67-99 tokens across the scenarios. - The query costs no prompt tokens. It is embedded, never sent to the narrator. It is recorded in the snapshot as provenance only.
- A1 is untouched. The safety reserve, the transport cost, the accounting states and the cold-model load are not in the diff. The summariser's input is still bounded at 2,000 tokens of story plus the cast brief, so its request is no larger than before.
- Embedding requests. One call per turn, as before; it now carries two short texts (at most about 200 and 180 tokens) instead of one of up to 600.
- On the real model, every counted turn of both attempts was
fits, with the window verified at 16,384 and prompts of at most 14,997 tokens (smallest observed margin 887). The memories section peaked at 716 tokens (attempt- and 1,246 tokens (attempt 2), against 67-99 deterministically. The difference is the real summariser's memory length: attempt 2's planting memory alone was 102 words, twice the prompt's ceiling (§R 12). It stayed inside the section's budget and the prompt inside the window.
N. Protocol-leak / A2 observations
- WP-B does not modify A2.
narrative/extract.py, the extractor rules and the protocol cleanup are not in the diff. - The B.1 observation stands, unchanged. In B.1's real run, one stored
narrator reply of 105 (action 153, depth 143) echoed application instructions
mid-reply, with story after them, outside A2's trailing cleanup region.
The evidence is preserved in
$HOME/v11-evidence/b1/real-1/campaign.db(V1.1-WP-B1-REPORT.md§K). It is a separate v1.1 follow-up item, not WP-B's. - This report does not claim the leak count is zero. The counts from B.2's real-model attempts are in §L. The integrated v1.1 release gate still needs its own protocol-leak result.
O. Compatibility
| Area | Result |
|---|---|
| Schema migration | none; no model or migration file in the diff |
| Bundle format | v3, unchanged; no bundle code in the diff |
| Existing memory rows | unchanged on open: the same 33 rows, row for row |
| Existing summaries | unchanged: 12 rows, row for row |
| Existing history | unchanged: 207 actions, 4 branches, 32 state events, 2 Save Points, row for row |
| Existing memories regenerated | no. New retrieval and eviction apply only when v1.1 next plays the campaign |
| Settings | memory_top_k, memory_bank_capacity and their defaults unchanged |
A real v1.0.0 database. tools/v11_compat_check.py on
m04-final/campaign.db (written by 96c1bf5, product code identical to
v1.0.0; user_version 94; SHA-256 c74e8798…4f18, unchanged by the runs). Each
tree opened its own copy:
| Comparison | Result |
|---|---|
| Schema before and after opening, v1.0.0 and B.2 trees | identical; user_version 94 → 94 on both |
| API snapshot (the export bundle, narrative state and events, Save Points, knowledge, memories, derived status, settings) | 0 differences |
| Tables after opening (memories, summaries, actions, state events, Save Points, branches, knowledge sources, adventures) | identical, row for row |
--exercise on a third copy, endpoint refused: Undo 119 → 117; Redo back to
119; restore On the ridge → 69 with Redo available; context dry run HTTP 200
with canon, state and the applicable summary; export and import 207 → 207
actions, the same head, Save Points 2 → 2, memories 33 → 33, identical state.
The dry run's memory retrieval reports the unreachable embedding endpoint and
uses no memories, as before.
Evidence: $HOME/v11-evidence/b2/compat/.
P. Offline / security
P.1 Offline and packaging regression
tools/m11_offline.py --out $HOME/v11-evidence/b2/offline, on the complete B.2
tree:
- the image was built with
--no-cache; - the container was started with
--network none: a loopback interface, no resolver, no route; - the exercise was driven inside it over
docker exec.
23 passed, 0 failed:
- no route to the public Internet;
- no external DNS;
- first page load offline, and the page names no remote origin;
- a CSP is served, and every shell asset is served locally;
- a campaign is created, and state extraction works;
- a local file imports, and the prompt assembles;
- imported knowledge is searchable;
- a turn with no model reachable is reported as a failure, with no narration accepted, the player's words kept, state unchanged and the earlier story intact;
- export and import work offline, with state, and no secret in the export;
- the media module imports, a scene packet builds, and no media provider is required;
- campaigns survive a container restart.
Evidence: offline-report.json, docker-build.log, container.log and
inside-stdout.txt in that folder.
As in M11, inference is not exercised offline: the model host is on the trusted LAN, which a container with no network cannot reach.
P.2 What the B.2 diff adds
| Check | Result |
|---|---|
| New network client, URL, host, IP, socket or subprocess | none: no added line in backend/app or backend/tools names one |
| External vector database | none. Vectors stay in SQLite embedding_blob; the lexical term is computed in process per turn, with no index |
| Cloud embedding | none. Embeddings go through the existing embedding_provider to the configured local endpoint, one call per turn as before |
| New dependency | none. The imports added to product code are math (standard library) and three internal modules: context.builder._encoding (the vendored tokenizer), knowledge.fts (its tokenizer and stop list) and narrative.model (entity_name) |
Endpoint policy (endpoints.py, ADR 011), TLS trust (tlstrust.py), providers |
unchanged; not in the diff |
| Frontend | unchanged; not in the diff |
| New data leaving the machine | none. The embedding request carries the player's input and a short scene text instead of 600 tokens of narration, to the same local endpoint |
| Real identifiers in committed files | none (scanned at staging, §Q) |
Local-only operation is unchanged.
Q. Full regression
| Run | Result |
|---|---|
| Full backend suite, complete B.2 tree | 1,652 passed, 17 skipped, 0 failed, 0 xfailed (1,603 s). The 17 skips are the tests that need a real model, the same 17 as in B.1. B.1's 2 strict xfails are now ordinary passes |
test_v11_b2_memory_ranking.py (new) |
18 passed |
test_v11_b2_memory_eviction.py (new) |
24 passed |
test_v11_b2_memory_excerpt.py (new) |
17 passed |
test_v11_b1_memory_diagnostic.py (B.1, updated for B.2) |
all pass; no xfail marker remains |
Memory regression group at the B2.2 checkpoint: test_memory_retrieval, test_memory_nodes, test_context_memory (F01-F08, E01-E04, write-lock and derived-failure recording), test_m11_leakage (all 14), test_m11_long_run_memory, test_memory_rewrite (L04), test_cast_brief, test_memory_settling, test_embedding_model_switch, test_bundle_v2, test_m10_bundle, test_m11_findings, test_v11_b1_long_run_verdict (M04 classifications) |
244 passed |
I-series bundle and memory tests (test_m9_*, test_bundle_v2, test_m10_bundle, test_process_restart, test_save_points) |
in the full suite, all pass |
Frontend: unchanged. No file under frontend/ is in the diff, so the
frontend tests, lint and production build are unaffected and were not rerun.
similarity keeps its meaning, so the inspector's display is unchanged.
R. Residual risks
| # | Risk | Why it remains | Where it would show |
|---|---|---|---|
| 1 | No relevance floor. A full memory_top_k is still used whenever the bank holds that many, however unrelated |
Out of B.2's authorised menu. It costs tokens, not correctness, and a floor is model-dependent | Memories section filled with weak matches on a scene unlike anything remembered |
| 2 | Weights chosen on a deterministic embedder. INPUT_WEIGHT 0.6 and LEXICAL_WEIGHT 0.15 were swept with the concept embedder, not nomic-embed-text |
The deterministic fixtures are isolation-valid; a real run is not guaranteed to be (§L). Real cosines sit higher and closer together (B.1: an unrelated query still scored 0.44), so the lexical term may matter more there, or less | Real-model ranked stage; the snapshot now records all three scores to measure it |
| 3 | A fact in the middle of a very long block is still not shown to the summariser | The excerpt is bounded by design. Blocks over 2,000 tokens need about 330-token actions; B.1's real run had 688-token blocks | created: no with fact_in_block: yes on a campaign with very long replies |
| 4 | Eviction keeps coverage, not any particular fact. An interior memory holding an important fact can still be thinned when its neighbours are close | The rule protects the opening and newest stretch, and spreads what remains. It has no notion of importance, which the plan left out of scope (no new field) | A campaign past capacity whose key fact sits in a dense middle stretch |
| 5 | Coverage ignores branches. A sibling line's memory at the same depth makes an active-line memory look redundant | Eviction is adventure-wide by design (v1.0.0 too). Of two memories sharing a start, the less recently used goes, which after a divergence is normally the abandoned one; but an active-line memory unused since before the divergence could lose to a sibling used later | Heavily branched campaigns past capacity |
| 6 | Real-model isolation is hard to obtain. A narrator told a salient fact tends to restate it, and a summariser keeps "important established facts" | Model behaviour, not a memory mechanism. B.2 did not change prompts or summary behaviour to manufacture a qualifying run | §L |
| 7 | The summariser may still paraphrase the marker in a way the exact-string removal does not catch | Only the exact marker is removed; a model rewording it ("part of the story is missing") would store that wording | Memory text mentioning an omission; not seen deterministically |
| 8 | The mid-reply protocol echo (B.1, action 153) | Not WP-B's; A2 was not modified | §N; the v1.1 release gate's protocol-leak result |
| 9 | Existing campaigns keep memories written from last-2,000-token excerpts | Existing memories are not regenerated (plan non-scope). tools/rewrite_memories.py is the opt-in repair |
Old long blocks in v1.0.0 campaigns |
| 11 | The real summariser drops facts stated by the player. In the one isolation-valid real run, the memory of the planting block omitted the planted fact, kept a fragment of its sentence attributed to the wrong character, and followed the narrator's reply instead | Not one of the three deficiencies this brief authorised. The plan's bounded menu includes "the memory prompt's instruction to keep named facts and objects", but B.2 was told not to change prompts or summary behaviour. This is the open WP-B failure (§S) | §L.2; any campaign whose narrator does not echo a player-stated detail |
| 12 | The 50-word memory ceiling is not enforced. Attempt 2's planting memory was 102 words, and described events past its own block | Prompt-only since v1.0.0 (B.1 §B.4). A longer memory costs section tokens (the section stayed inside budget, §M) and can mix in later events | Memory text lengths in the bank |
| 10 | Cosmetic: a doubled full stop in the scene text ("rain outside..") when the state's scene summary already ends in one | Found in attempt 1's stored query. Only the embedder reads it; the effect is one stray token. Not changed after the real-model evidence was taken, so that evidence describes the tree as staged | memories.query.context in any snapshot whose scene summary ends in punctuation |
S. Final WP-B disposition
This section records the owner's final disposition (2026-09-15). It replaces two earlier decision texts that stood before the owner decided. Their evidence is unchanged in §L and §T.
B2.1 RANKING: PASS
B2.2 EVICTION: PASS
B2.3 EXCERPT CREATION: PASS
B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED
DETERMINISTIC WP-B: PASS
REAL-MODEL WP-B: FAIL
WP-B OVERALL:
ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION
independent deterministic recovery:
PASS
reference-model independent recovery:
NOT DEMONSTRATED / FAILED ON THE PRECONDITION-VALID ATTEMPT
What ships, and what each part proved:
- B2.1 ranking (§C, §D):
- the planting-era memory moves into the selected top-k (rank 7 and 6 → 1 and 2);
- the player's current input materially affects retrieval;
- semantic retrieval stays active, since a paraphrase with no shared word ranks first;
- the lexical score is inspectable, recorded per used memory.
- B2.2 coverage-aware eviction (§E, §F):
- the early memory survives every deterministic over-capacity case;
- the bank keeps coverage from the opening to the newest stretch;
- pins and lineage stay safe, and the bank stays bounded.
- B2.3 bounded head+tail excerpt (§G, §H):
- a fact at the start of a long block reaches the summariser;
- late facts stay visible, and the input stays within 2,000 tokens;
- short blocks are unchanged.
- Together (§I):
independent_full, which fails onv1.0.0at creation, returnsrecovered_through_memory_independent. Isolation is asserted on every turn, and provenance resolves to the planting turn. With the B2.4 prompt reverted, this still passes (§T.17).
The real-model limitation, accepted for v1.1. The reference memory
summariser, qwen2.5:3b-instruct-16k, can be given the complete relevant
source block and still fail to keep a distinctive player-established fact. The
failure can be:
- omitting the fact;
- omitting its specific objects;
- attributing it to the wrong character;
- preferring the generic narration that follows it.
A precondition-valid real attempt failed exactly this way (§L.2). Real-model independent-memory recovery is therefore not guaranteed, even though creation-window, retention, ranking and injection now work under deterministic isolation. This is a failure, not an inconclusive result.
B2.4, rejected.
- Why it was tried: the precondition-valid attempt failed at memory creation.
- What it was: a prompt experiment (§T.4).
- Measured on the reference model (§T.8):
- target-fact retention on the original failing block: 0 of 5 under both the old and the experimental prompt;
- it reduced wrong-character attribution (8 → 3 of 40) and invention (2 → 0);
- it introduced frequent "Memory:" prefixes (29 of 40) and frequent second-person "you" (18 of 40);
- it produced more over-50-word memories (27 against 13);
- on one fixture it lost a promise the old prompt always kept (5/5 → 0/5).
- Conclusion: no reliable net improvement for the target failure. The production prompt is v1.0.0's again (§T.17). B2.4 did not pass.
Scope held. No further prompt variation was tried. Not introduced:
- a larger summariser model;
- a second extraction pass or structured fact extraction;
- a fact database;
- another LLM call, model routing, or automatic re-summarisation.
Those are future options (§T.18).
Release-gate visibility. V1.1-PLAN.md §11 now requires the v1.1 release
report to state the deterministic PASS, the reference-model FAIL, the failing
stage (memory creation, content selection) and the owner's acceptance. It must
not reduce these to "WP-B passed" (criterion 12). It must also carry the
mid-reply A2 echo and the doubled full stop in the scene text as separate
residuals (criterion 13).
Nothing is committed, pushed or tagged. WP-C has not started.
T. B2.4 — Memory summariser fidelity for named facts and objects
T.1 Authorisation
The owner's B2.4 brief (2026-09-15), under V1.1-PLAN.md §8 WP-B's bounded
menu: "the memory prompt's instruction to keep named facts and objects". No new
architectural scope. The trigger is attempt 2 (§L.2): precondition-valid, whole
block sent, fact omitted.
Repository at start:
- HEAD
beb17ad, signed (good signature,02C9BF7D…); - B2.1-B2.3 staged: 12 files, +2,525/−183;
- nothing unstaged; WP-C not started.
T.2 The prompt before B2.4
MEMORY_SYSTEM_PROMPT as shipped in v1.0.0 (184 tokens, verbatim in
tools/memory_fidelity.py as BASELINE_MEMORY_PROMPT, checked against git by a
test):
| # | Question | Answer |
|---|---|---|
| 1 | What does it tell the model to preserve? | "the concrete facts and events (names, places, items, promises, injuries)"; "the details a later scene could turn on — a name, a promise, an injury, where something is" |
| 2 | What does it prioritise? | Nothing explicitly. "Drop the ones it could not [turn on]" is the only rule of precedence, and it does not say what gives way to what |
| 3 | Attribution? | Only through pronouns: name the characters rather than "he", "she" or "they". Nothing on keeping an action, promise or possession with its owner |
| 4 | People, objects, places, promises, clues? | Names, places, items, promises and injuries are listed; "where something is" is named. Clues, discoveries, possession and knowledge are not |
| 5 | Invention? | Nothing |
| 6 | Length? | "1-2 plain sentences", "at most 50 words" |
| 7 | Does it implicitly favour narrator prose? | Yes. It says "The excerpt is written in the second person: 'you' is the protagonist". The protagonist's own lines are marked > and are often first person ("> I watch …", format_player_input leaves "I" alone), and nothing says what they establish is story. In the failed block, that line was 26 of about 460 words; the narration around it was 386 |
| 8 | Does its wording explain the failure? | Plausibly. "Facts and events" with a two-sentence cap favours the event sequence, and the depth-4 reply was the most event-dense text in the block (the decision, the walk to the abbey, entering the crypt). The static, player-stated fact (where the sundial is) had no stated precedence over it |
T.3 The observed failure, stated precisely
From attempt 2's stored block and memory (§L.2):
- The fact was in the block, at depth 3, as the player's line.
- The memory omitted the sundial and the teapot.
- It was 102 words against a 50-word target.
- It followed the narrator's depth-4 reply.
Two corrections to §L.2's first analysis, now marked there:
- The crypt events are not past the block. The depth-4 reply narrates them, so the memory invented nothing.
- "promising not to tell anyone" has the right promiser. In the source, Aldric promises Mara. The memory's error is what the promise is attached to: it ties the promise to going to the crypt, and drops Mara and the object. The phrase may also echo Edrin's own "I'll tell no one what we found".
Failure stage: creation. Cause class: the summariser's content selection and fidelity, not the excerpt (the whole 623-token block was sent).
T.4 The prompt correction
One constant changed (memorybank.MEMORY_SYSTEM_PROMPT, 184 → 294 tokens), plus
a no-behaviour-change helper, memory_user_prompt(brief, excerpt), factored out
of summarize_block so an evaluation sends a model the identical user message.
Still one call, one prompt, the same 50-word target, no example, no genre
vocabulary. It states:
- keep concrete facts a later scene could turn on: specific people, objects and places; where something is; who has, hid, found, knows, saw or promised what; injuries, clues and commitments;
- a distinctive fact comes before mood, scenery, routine movement and small talk: drop those first, and never let a later passage crowd out an earlier fact;
- keep each fact with the person it belongs to; never move an action, promise, possession, statement or piece of knowledge between characters; never add a fact the excerpt does not state;
- the narration calls the protagonist "you", and lines beginning
>are the protagonist's own actions and words ("I" or "You"). What they establish is story just as the narration is. This is equal standing, not a preference for player text; - the framing rule is kept: third person, the protagonist by name, no bare pronouns, no preamble.
No truncation was added. A memory over the target is stored as written
(test_an_over_long_memory_is_stored_as_written_never_cut).
Owner disposition: this prompt was rejected and reverted (§T.17). The production prompt is v1.0.0's. The text above records the experiment.
T.5 Deterministic fixtures
tools/memory_fidelity.py holds eight fixtures, each a short block with texture
and a fact, declaring the facts to keep, their owners, and forbidden inventions.
| Fixture | Requirement | Genre |
|---|---|---|
object_place_office |
B2.4-1 object and place | office |
player_fact_station |
B2.4-2 player-established fact | science-fiction-neutral |
promise_contemporary |
B2.4-3 promise | contemporary |
attribution_station |
B2.4-4 attribution (two characters, a code and a badge) | science-fiction-neutral |
clutter_office |
B2.4-5 clutter pressure | office |
no_invention_office |
B2.4-6 no invention (an unclaimed briefcase) | office |
multiple_facts_station |
B2.4-7 several facts, three owners | science-fiction-neutral |
regression_attempt_2 |
Part 4, the actual failure | the harness campaign |
The checker (evaluate) is deterministic and a stated heuristic:
- a fact is kept when one sentence names all its parts;
- attribution is the nearest named character before the fact's verb;
- inventions are patterns anchored on the wrong character as subject;
- words are counted, not enforced.
test_v11_b2_summarizer_fidelity.py, 49 passed:
- 12 prompt-contract checks, the 50-word target, genre neutrality with no example or fixture name, framing rules kept, and the baseline equal to the shipped prompt;
- coverage of every requirement in three genres;
- every fixture's facts reaching the summariser whole;
- the checker passing each faithful memory and failing each failure shape (13 unfaithful memories: omission, misattribution, invention);
- the application sending the corrected prompt with the whole planting block;
- no truncation;
- the corrected regression memory ranking first under B2.1 scoring.
B2.4 adds no xfail.
T.6 The actual failed-run regression
regression_attempt_2 is attempt 2's planting block verbatim. It is the
harness's own fixture campaign, with no hostname, person or other identifier.
| Required after B2.4 | Deterministic | Reference model, B2.4 prompt (§T.8) |
|---|---|---|
| input contains planted fact | yes | yes |
| summariser receives planted fact | yes (the whole block; test) | yes |
| generated memory contains planted fact | checker passes the corrected memory and fails the stored one (test) | no: 0 of 5 |
| attribution correct | checked | not reached (fact absent) |
| distinctive objects retained | checked | no |
| unsupported facts invented | no | no (0 of 5) |
The pre-B2.4 prompt fails the fixture legitimately. The stored memory is the real output, and the old prompt also scored 0 of 5 in §T.8; nothing was fabricated. The B2.4 prompt fails it equally.
T.7 Length and attribution behaviour
| 40 memories per arm (8 fixtures × 5) | Baseline prompt | B2.4 prompt |
|---|---|---|
| median / p90 / max words | 48 / 72 / 82 | 68 / 83 / 105 |
| over the 50-word target | 13 | 27 |
| misattributed | 8 | 3 |
| invented | 2 | 0 |
Worst length: 105 words (object_place_office, B2.4, a run-on of four
"Memory:" clauses). The length target was measured, not enforced; nothing was
cut.
T.8 Real-model comparison
tools/memory_fidelity.py, 2026-09-15 09:26-09:27:
- Model:
qwen2.5:3b-instruct-16kon the GPU inference host, through the application's provider at production temperature (0.3). - Samples: 5 per fixture per prompt.
- Logging: running (§T.11).
- Evidence:
$HOME/v11-evidence/b2/fidelity-b24/fidelity.json, every memory verbatim.
| Fixture | Baseline passed | B2.4 passed | What the memories show |
|---|---|---|---|
object_place_office |
5/5 | 2/5 | B2.4 runs began "Memory:"; two lost the drawer and cabinet ("in the archive room"); one run-on of 105 words, whose "Priya … sliding" the checker mis-scored (sliding is not a slide form it matches) |
player_fact_station |
2/5 | 4/5 | baseline misattributed and invented; B2.4 kept the sample and its owner |
promise_contemporary |
5/5 | 0/5 | every B2.4 run began "Memory:" and listed only scenery (floorboards, radiator, dog); four dropped the promise, one mentioned the lease without who promised; two used "You". The baseline kept the promise every time |
attribution_station |
1/5 | 4/5 | baseline gave the code or badge to the wrong person |
clutter_office |
4/5 | 5/5 | |
no_invention_office |
3/5 | 3/5 | |
multiple_facts_station |
1/5 | 5/5 | |
regression_attempt_2 |
0/5 | 0/5 | every memory under both prompts retells the key, the decision and the crypt; none names the sundial or teapot |
| total | 21/40 | 23/40 | B2.4: "Memory:" prefix 29/40, "you" 18/40 (baseline 0 and 0) |
Reading it plainly:
- Attribution and invention improved.
- The target failure did not move.
- The longer prompt made this 3B model echo the "Memory:" cue from the user message, drift into second person, write longer, and on one fixture turn the "drop scenery" instruction into a scenery list.
- This is a model-capability limit at this prompt size, measured, not guessed.
T.9 B2.1, B2.2 and B2.3 with B2.4 in the tree
| Suite | Result |
|---|---|
test_v11_b1_memory_diagnostic.py: all B.1 scenarios, independent_full (recovered_through_memory_independent, isolation every turn, provenance), lineage (G), authority |
pass |
test_v11_b2_memory_ranking.py: B2.1, negative controls, pins |
pass |
test_v11_b2_memory_eviction.py: B2.2, bounds, pins, fallback, lineage separation |
pass |
| (the three files together) | 81 passed |
test_v11_b2_memory_excerpt.py, test_cast_brief.py, test_memory_rewrite.py |
53 passed; short-block prompts byte-identical, excerpt split unchanged |
The deterministic scenarios use the best-case scripted summariser, so they are independent of the prompt. That they still pass shows B2.4 disturbed no mechanism; it is not evidence about the prompt. No ranking weight, eviction rule or excerpt split changed.
T.10 Full deterministic WP-B
Unchanged from §I on the B2.4 tree:
independent_full: created, retained, ranked and injected;recovered_through_memory_independent, with isolation asserted on every one of 52 turns;- provenance to depths 0-5.
Lineage (§J) and authority (§K) pass unchanged.
T.11 GPU-host evidence review
For the WP-B runs so far: B.1 on 2026-09-14 18:16-18:37, and B.2 attempts at 22:00-22:22 and 22:23-22:44. Read-only, from the host's logs and its persistent journal.
| Check | Finding |
|---|---|
| Power | peak 294.5 W (B.2 window), 304.6 W (earlier set), 307 W (dmon). The power limit now reads 200 W (default 280 W), so it was set after these runs. The owner notes spikes above the limit occur, and the dock is rated to 450 W |
| Power-limit events | dmon power-violation samples: 17 of 5,804 in the B.2 set. dmon has no timestamps, so they cannot be placed within the runs |
| Temperature / throttling | peak 68 °C; thermal-violation samples 0 |
| PCIe link | Gen3 x4 under load in 2,462 of 2,463 busy samples (one Gen2 sample); PCIe error counter 0 |
| Kernel | 4 kernel lines in the whole window; no Xid, NVRM, AER, reset or fallen-off-the-bus |
| Ollama | no crash, exit or panic; models loaded once at 22:00:36-38 and stayed loaded through attempt 2 |
| HTTP errors | 9 × 404: the harness's deliberate failed_call probe (no-such-model-m11). 4 × 500: background calls cut off by the harness's planned server restarts (22:18:13 at restart 3; 18:33:13 in B.1). 21:29:54: other traffic outside both windows |
| Contamination | None. No host fault could have altered the WP-B evidence; attempt 2's failure is model output |
A logging defect, found and fixed. The documented watch command,
journalctl -f -k -u ollama, matches nothing: -k and -u are different
fields, which journalctl ANDs. Over the B.2 window it returned 0 lines, where
the OR form (_TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service) returned
30,229. Every such log was therefore empty, and events were read from the
persistent journal instead.
DEVELOPMENT.mdnow documents the OR form.- For the B2.4 comparison the owner's watch still ran the old command, and its file is again empty. The journal holds 2,091 Ollama lines for 09:25-09:28.
T.12 Fresh real-model attempts
0. The comparison (§T.8) ran first and showed the B2.4 prompt failing on the exact planting block the long run plants. Under the brief ("If the summarizer still fails … then stop. Do not create B2.5 automatically"), the long attempts were not run: they could only have confirmed a known creation failure, or failed isolation.
T.13 Context-budget interaction
- Memory prompt: 184 → 294 tokens (+110) in the background memory call only. The narrator's prompt, A1's reserve, the transport cost and the accounting are unchanged.
- Typical memory length on the fixtures: median 48 → 68 words; worst 105.
- Real-model turns after B2.4: none, so no turn was
exceededortruncation_suspected, and A1's protected budget was never at issue. - Before B2.4: every counted turn of both B.2 attempts was
fits, with the memories section at most 1,246 tokens (§M). - If the B2.4 prompt were kept, its longer memories would grow the memories section roughly in proportion (median +40%). It is still priced as protected context, and A1 would refuse rather than overflow.
T.14 Offline regression
tools/m11_offline.py --out $HOME/v11-evidence/b2/offline-b24, on the B2.4
tree: a fresh --no-cache image in a --network none container, driven over
docker exec. 23 passed, 0 failed, the same checks as §P.1:
- no route and no external DNS;
- the page loads offline, with a CSP and local assets;
- a campaign is created, state is extracted, and a file imports;
- the prompt assembles, and knowledge is searchable;
- a turn with no model is a reported failure that loses nothing;
- export and import work, with no secret in the export;
- the media module stays inert;
- campaigns survive a restart.
What B2.4 changes on the network:
- Nothing. The diff is a prompt string, a pure helper, an evaluation tool, tests and documentation.
- No new destination, dependency or external service.
tools/memory_fidelity.pycalls only the--endpointit is given, through the application's own provider, and it is a developer tool that production never imports. - Unchanged: the endpoint policy, the trusted-LAN policy, embeddings (local), and the vector store (SQLite).
T.15 v1.0.0 compatibility
Rerun on the B2.4 tree against m04-final/campaign.db (SHA-256 prefix
c74e8798b68d2e22, unchanged by the runs), from a fresh 432f041 worktree:
| Result | |
|---|---|
| Schema before and after opening | identical to v1.0.0; user_version 94 → 94 |
| API snapshot | identical |
| memories (33), summaries (12), actions (207), state events (32), Save Points (2), branches (4), knowledge sources (3), adventures (1) | identical, row for row |
| Exercise | export and import 207 → 207 actions, same head, same narrative state |
| Bundle format | v3, unchanged; no memory regenerated |
T.15a Full backend regression
Full backend suite, B2.4 tree: 1,701 passed, 17 skipped, 0 failed, 0 xfailed (1,101 s).
- The count is B2.1-B2.3's 1,652 plus B2.4's 49 fidelity tests.
- The 17 skips are the same real-model tests as before.
- The suite covers: all WP-B tests, all B.1 diagnostic tests, B2.1-B2.4, F01-F08
and E01-E04, the M04 classifications,
test_context_memory,test_memory_nodes,test_memory_retrieval,test_m11_leakage,test_m11_long_run_memory,test_memory_rewrite, the I-series bundle and memory tests, derived-work failure recording and write-lock regressions.
Frontend: unchanged. No file under frontend/ is in the WP-B diff, so the
frontend tests, lint and build were not rerun.
T.16 Residual risks added by B2.4
| # | Risk |
|---|---|
| 13 | The reference summariser does not keep a player-stated fact against event-dense narration, under either prompt (0 of 10 on the attempt-2 block). This is the open WP-B limitation |
| 14 | |
| 15 | The fidelity checker is a heuristic. It missed one correct attribution (sliding), and may miss paraphrases. Every memory is kept verbatim so a reader can check it |
| 16 | Operator logging: the documented kernel/Ollama watch recorded nothing on every run before the fix. The journal is the source of those events |
| 11-12 (earlier) | still open |
| 8 (A2 mid-reply echo) and 10 (doubled full stop in the scene text) | unchanged, and still out of scope |
T.17 Owner disposition, and the revert
Owner decision, 2026-09-15: option 1(a).
- B2.1, B2.2 and B2.3 are accepted.
- The B2.4 prompt does not ship.
- Deterministic WP-B is accepted.
- The real-model failure is accepted as a documented v1.1 residual limitation.
| Reverted | Kept |
|---|---|
memorybank.MEMORY_SYSTEM_PROMPT, restored byte-for-byte to beb17ad's (184 tokens), and the B2.4 comment above it |
B2.1 retrieval and ranking, B2.2 eviction, B2.3 head+tail excerpt |
| the tests that required the experimental prompt's wording | memorybank.memory_user_prompt: a behaviour-neutral helper summarize_block calls, so a measurement sends a model the application's exact message |
the B2.4 note in CONTEXT-AND-MEMORY.md §15, replaced by the shipped behaviour and the limitation |
B.1 and B.2 diagnostics; tools/memory_fidelity.py, now diagnostic-only, comparing the shipped prompt with B24_EXPERIMENT_PROMPT, which it keeps so the experiment can be repeated |
| this section's evidence, the fixtures, and every measured memory |
The tests after the revert (test_v11_b2_summarizer_fidelity.py) are split
in two.
- Acceptance gates the tree:
- the shipped prompt equals v1.0.0's, checked against git;
- the experiment is not what ships;
- every fixture reaches the summariser whole;
- the application sends the shipped prompt with the whole planting block;
- an over-long memory is never cut;
- a faithful regression memory ranks first under B2.1.
- Diagnostic measurement proves only the instrument, on hand-written
memories with known answers:
- fact retention, attribution and invention;
- word count, a leading "Memory:" and second-person "you";
- promise retention.
No reference-model score is a test gate.
Verification after the revert:
- The runtime constant equals
beb17ad's value. - The only prompt-related lines left in the diff against
beb17adare thesummarize_blockcall throughmemory_user_prompt. - The fidelity, cast-brief, rewrite and excerpt tests: 94 passed.
Final regression, offline and compatibility, on the final staged tree (the B2.4 prompt reverted):
| Check | Result |
|---|---|
| Full backend suite | 1,693 passed, 17 skipped, 0 failed, 0 xfailed (1,227 s). This is B2.1-B2.3's 1,652 plus 41 summariser tests; the 8 that required the experimental prompt's wording went with it. The 17 skips are the same real-model tests |
| Deterministic WP-B within it | B2.1 ranking, B2.2 eviction, B2.3 excerpt; independent_full still recovered_through_memory_independent, with isolation every turn and provenance; lineage (G); authority; the B.1 diagnostics; F01-F08, E01-E04, M04, the I-series, derived-work failure recording and write-lock regressions: all pass |
Offline (tools/m11_offline.py, fresh --no-cache image, --network none) |
23 passed, 0 failed. Rerun because the runtime differs from §T.14's tree |
v1.0.0 compatibility (m04-final/campaign.db, compared with the 432f041 snapshot taken the same day) |
API snapshot and schema identical; user_version 94 → 94; memories (33), summaries (12), actions (207), state events (32), Save Points (2), branches (4), knowledge sources (3) and the adventure identical, row for row; source database unchanged |
| Frontend | unchanged; not in the diff |
deterministic WP-B: PASS
offline regression: PASS (23/23)
v1.0.0 compatibility: unchanged
T.18 Future memory-quality options (not implemented)
Each of these is new scope for a future package, and none is part of v1.1:
- A stronger dedicated summariser model, evaluated with
tools/memory_fidelity.pyagainst the same fixtures. - Structured fact extraction alongside the prose memory: who has, hid, knows or promised what.
- Separate factual and narrative memory, each retrieved on its own terms.
- Model-specific summariser recommendations in the operator documentation, once measured.