v1.1 WP-B.2: independent long-term memory retention

Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-15 11:21:53 -04:00
co-authored by Claude Opus 5
parent beb17ada10
commit 0c1ba836ba
15 changed files with 3782 additions and 186 deletions
@@ -0,0 +1,277 @@
"""v1.1 WP-B.2 (B2.2): which memory a full bank lets go of.
v1.0.0 evicted the least recently used memory. B.1 showed that this discards
the only memory of an early stretch first, because retrieval follows the present
scene and nothing recent resembles it. `memorybank.eviction_order` now thins the
bank where it is densest and keeps the opening and the newest stretch, with
recency as the tie-break and least-recently-used as the fallback.
The scenario-level tests (the planted fact kept past capacity) are in
`test_v11_b1_memory_diagnostic.py`; the v1.0.0 eviction tests in
`test_memory_retrieval.py` still pass unchanged, because their memories carry
no source range and take the fallback.
python -m pytest tests/test_v11_b2_memory_eviction.py -v
"""
import random
from collections import namedtuple
from datetime import datetime, timedelta
import pytest
from sqlalchemy import select
from app import memorybank, models, tree
from app.context import lineage
from app.database import Base, SessionLocal, engine
T0 = datetime(2026, 1, 1, 12, 0, 0)
Row = namedtuple("Row", "id pinned source_start source_end last_used_at created_at use_count")
def row(id, start, end=None, *, pinned=False, used=None, created=None, uses=0):
"""A memory as eviction sees it. Times are minutes after T0."""
return Row(id, pinned, start, (start + 5) if end is None and start is not None else end,
None if used is None else T0 + timedelta(minutes=used),
T0 + timedelta(minutes=id if created is None else created), uses)
def blocks(n, *, first_id=1):
return [row(first_id + i, 6 * i) for i in range(n)]
def largest_gap_from_opening(rows):
"""The longest uncovered run of depths from depth 0 to the last memory."""
ordered = sorted((r.source_start, r.source_end) for r in rows)
gaps = [ordered[0][0]]
reach = ordered[0][1]
for start, end in ordered[1:]:
gaps.append(max(0, start - reach - 1))
reach = max(reach, end)
return max(gaps)
# ------------------------------------------------------------ the pure order
def test_the_opening_and_the_newest_memory_are_kept():
bank = blocks(7)
doomed = memorybank.eviction_order(bank, 5)
assert bank[0].id not in doomed and bank[-1].id not in doomed
assert len(doomed) == 5
def test_the_densest_stretch_is_thinned_first():
# Memories every 6 depths to 30, then a sparse stretch. Removing one of the
# dense ones leaves a 6-depth hole; removing a sparse one leaves far more.
bank = [row(1, 0), row(2, 6), row(3, 12), row(4, 18), row(5, 60), row(6, 120), row(7, 180)]
assert memorybank.eviction_order(bank, 1)[0] in {2, 3, 4}
assert set(memorybank.eviction_order(bank, 2)) <= {2, 3, 4}
def test_a_stretch_two_memories_describe_loses_one_of_them_first():
"""A shared start (a re-played stretch, or a sibling line) leaves no hole.
Of the two, the less recently used goes, even though a unique memory
elsewhere is older and less used than both."""
bank = [row(1, 0), row(2, 6, used=5), row(3, 12, used=50), row(4, 12, used=40), row(5, 18)]
assert memorybank.eviction_order(bank, 1) == [4]
def test_equal_holes_fall_to_the_least_recently_used():
bank = [row(1, 0), row(2, 6, used=30), row(3, 12, used=10), row(4, 18, used=20), row(5, 24)]
assert memorybank.eviction_order(bank, 1) == [3]
def test_a_newborn_can_be_the_legitimate_first_to_go():
"""The frozen bank is about a newborn losing to a count it cannot have yet.
A newborn that only repeats a stretch another memory describes, one used
after it was written, is legitimately the first to go."""
bank = [row(1, 0), row(2, 6, used=100), row(3, 12), row(4, 6, created=90)]
assert memorybank.eviction_order(bank, 1) == [4]
def test_the_newest_memory_is_not_evicted_by_the_bank_it_joins():
"""The frozen-bank regression under the new rule: every older memory has
been used, the newborn never has, and it still stays."""
bank = [row(i, 6 * (i - 1), used=200 + i, uses=3) for i in range(1, 6)]
newborn = row(6, 30, created=300)
assert newborn.id not in memorybank.eviction_order(bank + [newborn], 1)
def test_pins_are_never_taken_but_still_count_as_coverage():
bank = [row(1, 0), row(2, 6, pinned=True), row(3, 12), row(4, 18, pinned=True), row(5, 24)]
doomed = memorybank.eviction_order(bank, 10)
assert not {2, 4} & set(doomed)
# With the pins covering 6 and 18, memory 3's hole is only its own block.
assert doomed[0] == 3
def test_memories_without_a_range_take_the_least_recently_used_fallback():
hand_written = [row(1, None, None, used=30), row(2, None, None, used=10),
row(3, None, None, used=20)]
assert memorybank.eviction_order(hand_written, 3) == [2, 3, 1]
def test_the_fallback_is_used_only_once_no_interior_memory_remains():
bank = [row(1, 0, used=1), row(2, 6, used=90), row(3, 12, used=2),
row(10, None, None, used=0)]
order = memorybank.eviction_order(bank, 4)
assert order[0] == 2 # the interior memory, although recently used
assert order[1:] == [10, 1, 3] # then least recently used
def test_the_order_does_not_depend_on_row_order():
bank = [row(i, 6 * (i - 1), used=(i * 37) % 11, uses=i % 3) for i in range(1, 30)]
bank += [row(40, 12), row(41, 12)] # a shared start with identical timestamps
expected = memorybank.eviction_order(bank, 20)
for seed in range(5):
shuffled = bank[:]
random.Random(seed).shuffle(shuffled)
assert memorybank.eviction_order(shuffled, 20) == expected
def test_a_tie_on_every_signal_is_broken_by_id():
bank = [row(1, 0), row(9, 6, created=0), row(4, 12, created=0), row(20, 18)]
assert memorybank.eviction_order(bank, 1) == [4]
@pytest.mark.parametrize("seed", range(8))
def test_capacity_holds_and_pins_survive_for_any_bank(seed):
rng = random.Random(seed)
bank = []
for i in range(1, rng.randint(2, 60)):
start = None if rng.random() < 0.15 else rng.randrange(0, 400)
bank.append(row(i, start, None if start is None else start + rng.choice([3, 5, 8]),
pinned=rng.random() < 0.1, used=rng.choice([None, rng.randrange(500)]),
uses=rng.randrange(4)))
capacity = rng.randint(1, 30)
overflow = len(bank) - capacity
doomed = memorybank.eviction_order(bank, max(0, overflow))
pinned = {r.id for r in bank if r.pinned}
assert not pinned & set(doomed)
assert len(set(doomed)) == len(doomed)
remaining = len(bank) - len(doomed)
assert remaining == max(capacity, len(pinned)) if overflow > 0 else remaining == len(bank)
@pytest.mark.parametrize("n, capacity, irregular", [(60, 10, False), (500, 80, False), (500, 80, True)])
def test_a_long_bank_keeps_describing_the_whole_story(n, capacity, irregular):
"""The general property, with nothing ever retrieved: memories arrive one
block at a time and the bank is kept at capacity. The opening stays, and no
stretch goes undescribed for more than twice the average spacing. Least
recently used order, on the same arrivals, keeps only the newest stretch."""
rng = random.Random(n)
kept, lru = [], []
depth = 0
for i in range(1, n + 1):
size = rng.choice([4, 6, 6, 9]) if irregular else 6
memory = row(i, depth, depth + size - 1)
depth += size
kept.append(memory)
lru.append(memory)
if len(kept) > capacity:
doomed = set(memorybank.eviction_order(kept, len(kept) - capacity))
kept = [m for m in kept if m.id not in doomed]
lru = sorted(lru, key=lambda m: (m.created_at, m.id))[len(lru) - capacity:]
assert len(kept) == capacity
assert min(m.source_start for m in kept) == 0
assert max(m.id for m in kept) == n
assert largest_gap_from_opening(kept) <= 2 * depth / capacity
assert largest_gap_from_opening(lru) > depth / 2 # v1.0.0 order: the opening is gone
# --------------------------------------------------------- on real rows
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
session = SessionLocal()
try:
yield session
finally:
session.close()
memorybank._vector_cache.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture()
def adventure(db):
user = models.User(is_guest=False, email="b2-evict@example.com")
db.add(user)
db.flush()
settings = models.Settings(user_id=user.id, model="m", embedding_model="e",
memory_bank_capacity=3)
adv = models.Adventure(user_id=user.id, title="Evict", script_state={}, memory_bank_enabled=True)
db.add_all([settings, adv])
db.commit()
adv.settings_row = settings
return adv
def test_the_pass_changes_nothing_but_forgotten(db, adventure):
for i in range(6):
memory = models.Memory(adventure_id=adventure.id, text=f"block {i}",
source_start=6 * i, source_end=6 * i + 5, branch_id=None, depth=6 * i + 5)
db.add(memory)
db.commit()
columns = (models.Memory.id, models.Memory.text, models.Memory.source_start,
models.Memory.source_end, models.Memory.branch_id, models.Memory.depth,
models.Memory.pinned, models.Memory.use_count, models.Memory.last_used_at)
before = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
after = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
assert before == after
active = db.execute(select(models.Memory.id).where(models.Memory.forgotten.is_(False))).scalars().all()
assert len(active) == 3
assert min(active) == min(before) and max(active) == max(before) # the boundaries
def test_eviction_does_not_make_an_abandoned_lines_memory_eligible(db, adventure):
"""Eviction and lineage are separate: the pass decides only `forgotten`, so
a memory on a line the story left is exactly as ineligible afterwards."""
trunk = []
for i in range(4):
action = models.Action(adventure_id=adventure.id, type="ai", text=f"trunk {i}")
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
trunk.append(action)
abandoned_node = models.Action(adventure_id=adventure.id, type="ai", text="the abandoned line")
tree.place_action(db, adventure, abandoned_node)
db.add(abandoned_node)
db.flush()
abandoned = models.Memory(adventure_id=adventure.id, text="on the abandoned line",
source_start=4, source_end=4)
tree.attach_memory(abandoned, abandoned_node)
db.add(abandoned)
db.commit()
# Move the head back and diverge, so the abandoned node is off the path.
adventure.head_depth = trunk[-1].depth
db.commit()
from app import head
head.fork_if_behind_head(db, adventure)
divergent = models.Action(adventure_id=adventure.id, type="ai", text="the new line")
tree.place_action(db, adventure, divergent)
db.add(divergent)
db.flush()
for i, node in enumerate(trunk + [divergent]):
memory = models.Memory(adventure_id=adventure.id, text=f"active {i}",
source_start=node.depth, source_end=node.depth)
tree.attach_memory(memory, node)
db.add(memory)
db.commit()
def eligible():
return set(db.execute(select(models.Memory.id).where(
models.Memory.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Memory),
models.Memory.forgotten.is_(False))).scalars().all())
assert abandoned.id not in eligible()
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
db.expire_all()
assert abandoned.id not in eligible()
assert len(db.execute(select(models.Memory.id).where(
models.Memory.forgotten.is_(False))).scalars().all()) == 3