v1.1 WP-B.2: independent long-term memory retention

Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-15 11:21:53 -04:00
co-authored by Claude Opus 5
parent beb17ada10
commit 0c1ba836ba
15 changed files with 3782 additions and 186 deletions
+180 -54
View File
@@ -11,8 +11,9 @@ diagnostic in `tools/memory_diagnostic.py`:
2. **What it finds on this tree.** The scenarios run with a best-case summariser,
one that keeps a fact if and only if the fact reached it. Any failure is
therefore the application's mechanism, not a model's writing.
- The criteria the current code does not meet are marked `xfail(strict=True)`,
so B.2 has to flip them deliberately.
- WP-B.1 marked the criteria v1.0.0 did not meet `xfail(strict=True)`. WP-B.2
fixed ranking, eviction and the creation excerpt, and those tests are now
ordinary passes; the v1.0.0 results are recorded in the WP-B.2 report.
- The same file is run unchanged against v1.0.0 for the baseline.
python -m pytest tests/test_v11_b1_memory_diagnostic.py -v
@@ -102,14 +103,14 @@ def test_retention_is_reported_with_the_bank_and_its_eviction_order():
def test_ranking_is_production_ranking_and_agrees_with_the_stored_selection():
ranked = scenario("independent_default")["diagnosis"]["ranked"]
assert ranked["replica_matches_stored_selection"] is True
assert ranked["lexical_score"] is None # memory ranking has no lexical term
assert ranked["top_k_cutoff"] == 5
assert ranked["yes"] is True and ranked["selected"] is True
assert 1 <= ranked["rank"] <= ranked["top_k_cutoff"]
# The production query is the newest four actions, cut to 600 tokens, and the
# one-line question is diluted by the narration around it.
variants = scenario("independent_default")["ranking_variants"]
assert ranked["semantic_score"] < variants["direct"]["similarity"]
# v1.1 WP-B.2: every part of the score is reported, and they add up.
assert 0.0 <= ranked["lexical_score"] <= 1.0
assert ranked["final_score"] == pytest.approx(
ranked["semantic_score"] + memorybank.LEXICAL_WEIGHT * ranked["lexical_score"], abs=2e-4)
assert ranked["query"]["input"].endswith(md.SCENARIOS["independent_default"].recall_text)
def test_injection_is_read_from_the_recall_turns_own_context():
@@ -124,53 +125,118 @@ def test_ranking_variants_direct_paraphrase_and_unrelated():
variants = scenario("independent_default")["ranking_variants"]
assert variants["direct"]["rank"] == 1 and variants["direct"]["selected"]
assert variants["paraphrase"]["rank"] == 1 and variants["paraphrase"]["selected"]
assert (variants["direct"]["similarity"] > variants["paraphrase"]["similarity"]
> 5 * variants["unrelated"]["similarity"])
assert (variants["direct"]["final_score"] > variants["paraphrase"]["final_score"]
> variants["unrelated"]["final_score"])
def test_retrieval_fills_top_k_whatever_the_similarity():
"""Diagnosis: there is no relevance floor. With more memories than
`memory_top_k`, an unrelated query still selects five, and the early fact
rides along at a similarity near zero."""
def test_retrieval_still_fills_top_k_whatever_the_similarity():
"""There is still no relevance floor: an unrelated question selects a full
`memory_top_k`. B.2 changed which memories those are, not how many — the
early fact is no longer carried along by an unrelated question."""
variants = scenario("independent_default")["ranking_variants"]
assert variants["unrelated"]["similarity"] < 0.1
assert variants["unrelated"]["selected"] is True
assert variants["unrelated"]["selected_count"] == 5
assert variants["unrelated"]["selected"] is False
assert variants["unrelated"]["rank"] > 5
# ------------------------------------------------- WP-B.2 ranking acceptance
@pytest.mark.parametrize("name", ["ranking_crowded", "ranking_context_dependent"])
def test_acceptance_the_early_memory_is_ranked_and_injected_below_capacity(name):
"""The B.1 ranking failure, made deterministic. On v1.0.0 both fixtures are
`retained_but_not_ranked` (ranks 7 and 6 of 17 against a top-k of 4)."""
result = scenario(name)
assert result["isolation"]["ok"], result["isolation"]
diagnosis = result["diagnosis"]
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
assert diagnosis["retained"]["active_memories"] <= diagnosis["retained"]["memory_bank_capacity"]
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["rank"] <= 4
assert diagnosis["ranked"]["replica_matches_stored_selection"] is True
assert diagnosis["injected"]["yes"] is True
assert diagnosis["verdict"] == "injected"
def test_the_crowded_fixture_is_won_by_the_players_question():
ranked = scenario("ranking_crowded")["diagnosis"]["ranked"]
assert ranked["rank"] == 1
assert ranked["lexical_score"] > 0 # "sundial" and "amber" are in the question
def test_a_paraphrase_is_found_by_meaning_not_by_shared_words():
"""Lexical matching must not replace semantic retrieval. The paraphrase
shares none of F's distinctive words, yet ranks first."""
for name in ("ranking_crowded", "independent_default"):
paraphrase = scenario(name)["ranking_variants"]["paraphrase"]
assert paraphrase["rank"] == 1 and paraphrase["selected"]
# Only "Mara" is shared, which is far less than the direct question holds.
direct = scenario(name)["ranking_variants"]["direct"]
assert paraphrase["lexical_score"] < direct["lexical_score"] / 2
def test_a_context_dependent_question_needs_the_scene():
""""I ask her what she keeps up there" names nothing F's memory holds. The
scene the last narration set up (Mara, the top shelf, a kettle) is what
finds it; without that context it ranks last."""
result = scenario("ranking_context_dependent")
ranked = result["diagnosis"]["ranked"]
assert ranked["lexical_score"] == 0.0
assert ranked["rank"] <= 4
assert "top shelf" in ranked["query"]["context"]
assert result["ranking_variants"]["input_only"]["rank"] > 4
def test_an_unrelated_rare_word_does_not_outrank_the_relevant_memory():
"""Negative control: the paraphrase plus a place only one other memory
holds. The decoy gains lexical score, and still ranks below F."""
for name in ("ranking_crowded", "independent_default"):
control = scenario(name)["ranking_variants"]["rare_word_with_paraphrase"]
assert control["decoy_lexical_score"] > control["lexical_score"]
assert control["rank"] == 1
assert control["decoy_rank"] > control["rank"]
def test_common_words_contribute_nothing():
common = scenario("independent_default")["ranking_variants"]["common_words"]
assert common["lexical_score"] == 0.0
# ---------------------------------------------------------- capacity/eviction
def test_past_capacity_the_early_memory_is_evicted_and_the_stage_says_so():
"""Diagnosis, not a requirement: what the current eviction rule does to F."""
result = scenario("past_capacity")
assert result["diagnosis"]["created"]["yes"] is True
assert result["diagnosis"]["verdict"] == "created_but_evicted"
eviction = result["eviction"]
assert eviction["f_evicted_at_turn"] is not None
# It was retrieved while the bank was small, stopped being retrieved once
# recent narration filled the top-k, and was then the least recently used.
assert eviction["f_use_count_when_evicted"] > 0
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
assert eviction["f_memory_was_first_evicted"] is True
@pytest.mark.parametrize("name", ["past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"])
def test_past_capacity_the_early_memory_is_retained(name):
"""v1.1 WP-B.2. On v1.0.0 all three are `created_but_evicted`: F was the
least recently used row once recent narration stopped retrieving it, and
went first (turns 21, 21 and 36). Coverage-first eviction keeps the only
memory of the opening, so it stays active and is recalled at depth 106."""
result = scenario(name)
assert result["isolation"]["ok"], result["isolation"]
diagnosis = result["diagnosis"]
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
assert result["eviction"]["f_evicted_at_turn"] is None
# The bank really was past capacity, and stayed bounded.
assert result["eviction"]["first_eviction_turn"] is not None
assert all(t["active"] <= result["scenario"]["capacity"] + (1 if result["scenario"]["pin_first_memory"] else 0)
for t in result["trace"])
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
def test_at_a_lower_top_k_the_early_memory_ages_out_after_it_stops_being_retrieved():
"""Diagnosis with most of the bank unretrieved on any turn, nearer the
shipped 5-in-80 ratio. F is not simply the oldest row: it is evicted some
turns after recent narration stopped pulling it into the top-k, which is
what ordering by last use does to a fact nothing recent mentions."""
result = scenario("past_capacity_low_top_k")
eviction = result["eviction"]
assert result["diagnosis"]["created"]["yes"] is True
assert result["diagnosis"]["verdict"] == "created_but_evicted"
assert eviction["f_use_count_when_evicted"] > 0
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
assert eviction["first_eviction_turn"] <= eviction["f_evicted_at_turn"]
assert eviction["created_and_evicted_same_turn"] == []
def test_past_capacity_the_bank_still_describes_the_whole_story():
"""What the rule buys in general, not only for F: the active bank reaches
from the opening to the newest block, and no stretch between them goes
undescribed for more than twice the average spacing a bank of this capacity
can afford (story span / capacity). On v1.0.0 these banks began at depths 36
and 18: the opening was simply gone."""
for name in ("past_capacity", "past_capacity_low_top_k"):
result = scenario(name)
cover = result["trace"][-1]["coverage"]
assert cover["first_start"] == 0
assert cover["last_end"] >= result["recall_depth"] - 2 * memorybank.MEMORY_INTERVAL
assert cover["largest_gap"] <= 2 * (cover["last_end"] + 1) / result["scenario"]["capacity"]
def test_no_memory_is_evicted_by_the_same_pass_that_created_it():
"""The frozen-bank regression the current rule fixed, still holding."""
for name in ("past_capacity", "past_capacity_pinned"):
"""The frozen-bank regression the v1.0.0 rule fixed, still holding."""
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
assert scenario(name)["eviction"]["created_and_evicted_same_turn"] == []
@@ -180,25 +246,28 @@ def test_a_pinned_memory_survives_capacity():
assert eviction["pinned_memory_forgotten"] is False
@pytest.mark.xfail(strict=True, reason=(
"WP-B.1 diagnosis on this tree: past memory_bank_capacity the planting-era "
"memory is evicted first, because it was never retrieved and eviction orders "
"by last use, then creation. B.2 must flip this deliberately."))
def test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity():
assert scenario("past_capacity")["diagnosis"]["verdict"] == "injected"
"""Was `xfail(strict=True)` in WP-B.1; B.2 fixed the eviction rule."""
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
assert scenario(name)["diagnosis"]["verdict"] == "injected"
# ---------------------------------------------------------- creation window
def test_a_fact_early_in_a_long_block_never_reaches_the_summariser():
def test_a_fact_early_in_a_long_block_now_reaches_the_summariser():
"""v1.1 WP-B.2. On v1.0.0 this block (2,079 tokens) was cut to its last
2,000, the fact at its start was never seen, and the stage was
`not_created`. The excerpt is now the block's opening and end."""
result = scenario("long_block_fact_early")
created = result["diagnosis"]["created"]
covering = created["covering_memories"]
assert covering, "the long block must have been summarised"
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
assert covering[0]["fact_in_block"] is True
assert covering[0]["fact_in_summariser_excerpt"] is False
assert result["diagnosis"]["verdict"] == "not_created"
assert covering[0]["fact_in_summariser_excerpt"] is True
assert created["yes"] is True
assert created["source_start"] <= result["plant_depth"] <= created["source_end"]
assert memorybank.EXCERPT_OMISSION_MARKER not in created["memory_text"]
def test_the_same_fact_late_in_the_same_sized_block_does():
@@ -209,12 +278,69 @@ def test_the_same_fact_late_in_the_same_sized_block_does():
assert result["diagnosis"]["created"]["yes"] is True
@pytest.mark.xfail(strict=True, reason=(
"WP-B.1 diagnosis on this tree: the summariser reads only the last "
f"{memorybank.MEMORY_EXCERPT_TOKENS} tokens of a block, so a fact early in a "
"long block is never seen. B.2 must flip this deliberately."))
def test_acceptance_a_fact_early_in_a_long_block_is_remembered():
"""Was `xfail(strict=True)` in WP-B.1; B.2 changed the excerpt."""
assert scenario("long_block_fact_early")["diagnosis"]["created"]["yes"] is True
assert scenario("long_block_fact_early")["diagnosis"]["verdict"] == "injected"
# ------------------------------------------ WP-B.2 full deterministic acceptance
def test_acceptance_full_isolation_holds_on_every_turn():
"""`independent_full`: long blocks, a crowded query, a bank past capacity.
F must be carried by memory alone for the whole run, not only at recall."""
result = scenario("independent_full")
assert result["plant_depth"] <= 3 and result["recall_depth"] >= 100
assert not any(t["f_in_state"] for t in result["trace"])
assert not any(t["f_in_summary"] for t in result["trace"])
assert result["isolation"]["ok"], result["isolation"]
for check in ("state_document", "state_snapshots", "later_narration", "summary",
"knowledge", "recent_history", "state_section"):
assert result["isolation"]["checks"][check]["ok"], check
def test_acceptance_full_every_stage_passes_past_capacity_with_long_blocks():
"""On v1.0.0 this fixture fails at creation: every block is over 2,000
tokens, and the fact at the start of the first one is never summarised."""
result = scenario("independent_full")
diagnosis = result["diagnosis"]
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
assert diagnosis["created"]["covering_memories"][0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["replica_matches_stored_selection"]
assert diagnosis["injected"]["yes"]
assert diagnosis["verdict"] == "injected"
def test_acceptance_full_provenance_resolves_to_the_planting_turn():
provenance = scenario("independent_full")["provenance"]
assert provenance["recorded"] is not None
assert provenance["range_covers_plant"] and provenance["matches_row"]
assert provenance["source_block_holds_planting"] is True
assert provenance["recorded"]["authority"] == memorybank.ACCEPTED_STORY
def test_acceptance_full_is_the_long_run_independent_memory_verdict():
"""The same measurements, judged by the long-run tool's own verdict."""
from tools import m11_long_run as lr
result = scenario("independent_full")
checks = result["isolation"]["checks"]
diagnosis = result["diagnosis"]
verdict = lr._independent_memory_verdict({
"independent_planted_depth": result["plant_depth"],
"planted_turn_outside_history": checks["recent_history"]["ok"],
"absent_from_state": checks["state_document"]["ok"] and checks["state_snapshots"]["ok"]
and not any(t["f_in_state"] for t in result["trace"]),
"absent_from_summary": checks["summary"]["ok"]
and not any(t["f_in_summary"] for t in result["trace"]),
"absent_from_knowledge": checks["knowledge"]["ok"],
"absent_from_later_narration": checks["later_narration"]["ok"],
"memory_covering_planting_carries_fact": diagnosis["created"]["yes"],
"memory_forgotten": not diagnosis["retained"]["yes"],
"memory_injected": diagnosis["injected"]["yes"],
})
assert verdict == "recovered_through_memory_independent"
# ------------------------------------------------------- lineage control (G)
@@ -0,0 +1,277 @@
"""v1.1 WP-B.2 (B2.2): which memory a full bank lets go of.
v1.0.0 evicted the least recently used memory. B.1 showed that this discards
the only memory of an early stretch first, because retrieval follows the present
scene and nothing recent resembles it. `memorybank.eviction_order` now thins the
bank where it is densest and keeps the opening and the newest stretch, with
recency as the tie-break and least-recently-used as the fallback.
The scenario-level tests (the planted fact kept past capacity) are in
`test_v11_b1_memory_diagnostic.py`; the v1.0.0 eviction tests in
`test_memory_retrieval.py` still pass unchanged, because their memories carry
no source range and take the fallback.
python -m pytest tests/test_v11_b2_memory_eviction.py -v
"""
import random
from collections import namedtuple
from datetime import datetime, timedelta
import pytest
from sqlalchemy import select
from app import memorybank, models, tree
from app.context import lineage
from app.database import Base, SessionLocal, engine
T0 = datetime(2026, 1, 1, 12, 0, 0)
Row = namedtuple("Row", "id pinned source_start source_end last_used_at created_at use_count")
def row(id, start, end=None, *, pinned=False, used=None, created=None, uses=0):
"""A memory as eviction sees it. Times are minutes after T0."""
return Row(id, pinned, start, (start + 5) if end is None and start is not None else end,
None if used is None else T0 + timedelta(minutes=used),
T0 + timedelta(minutes=id if created is None else created), uses)
def blocks(n, *, first_id=1):
return [row(first_id + i, 6 * i) for i in range(n)]
def largest_gap_from_opening(rows):
"""The longest uncovered run of depths from depth 0 to the last memory."""
ordered = sorted((r.source_start, r.source_end) for r in rows)
gaps = [ordered[0][0]]
reach = ordered[0][1]
for start, end in ordered[1:]:
gaps.append(max(0, start - reach - 1))
reach = max(reach, end)
return max(gaps)
# ------------------------------------------------------------ the pure order
def test_the_opening_and_the_newest_memory_are_kept():
bank = blocks(7)
doomed = memorybank.eviction_order(bank, 5)
assert bank[0].id not in doomed and bank[-1].id not in doomed
assert len(doomed) == 5
def test_the_densest_stretch_is_thinned_first():
# Memories every 6 depths to 30, then a sparse stretch. Removing one of the
# dense ones leaves a 6-depth hole; removing a sparse one leaves far more.
bank = [row(1, 0), row(2, 6), row(3, 12), row(4, 18), row(5, 60), row(6, 120), row(7, 180)]
assert memorybank.eviction_order(bank, 1)[0] in {2, 3, 4}
assert set(memorybank.eviction_order(bank, 2)) <= {2, 3, 4}
def test_a_stretch_two_memories_describe_loses_one_of_them_first():
"""A shared start (a re-played stretch, or a sibling line) leaves no hole.
Of the two, the less recently used goes, even though a unique memory
elsewhere is older and less used than both."""
bank = [row(1, 0), row(2, 6, used=5), row(3, 12, used=50), row(4, 12, used=40), row(5, 18)]
assert memorybank.eviction_order(bank, 1) == [4]
def test_equal_holes_fall_to_the_least_recently_used():
bank = [row(1, 0), row(2, 6, used=30), row(3, 12, used=10), row(4, 18, used=20), row(5, 24)]
assert memorybank.eviction_order(bank, 1) == [3]
def test_a_newborn_can_be_the_legitimate_first_to_go():
"""The frozen bank is about a newborn losing to a count it cannot have yet.
A newborn that only repeats a stretch another memory describes, one used
after it was written, is legitimately the first to go."""
bank = [row(1, 0), row(2, 6, used=100), row(3, 12), row(4, 6, created=90)]
assert memorybank.eviction_order(bank, 1) == [4]
def test_the_newest_memory_is_not_evicted_by_the_bank_it_joins():
"""The frozen-bank regression under the new rule: every older memory has
been used, the newborn never has, and it still stays."""
bank = [row(i, 6 * (i - 1), used=200 + i, uses=3) for i in range(1, 6)]
newborn = row(6, 30, created=300)
assert newborn.id not in memorybank.eviction_order(bank + [newborn], 1)
def test_pins_are_never_taken_but_still_count_as_coverage():
bank = [row(1, 0), row(2, 6, pinned=True), row(3, 12), row(4, 18, pinned=True), row(5, 24)]
doomed = memorybank.eviction_order(bank, 10)
assert not {2, 4} & set(doomed)
# With the pins covering 6 and 18, memory 3's hole is only its own block.
assert doomed[0] == 3
def test_memories_without_a_range_take_the_least_recently_used_fallback():
hand_written = [row(1, None, None, used=30), row(2, None, None, used=10),
row(3, None, None, used=20)]
assert memorybank.eviction_order(hand_written, 3) == [2, 3, 1]
def test_the_fallback_is_used_only_once_no_interior_memory_remains():
bank = [row(1, 0, used=1), row(2, 6, used=90), row(3, 12, used=2),
row(10, None, None, used=0)]
order = memorybank.eviction_order(bank, 4)
assert order[0] == 2 # the interior memory, although recently used
assert order[1:] == [10, 1, 3] # then least recently used
def test_the_order_does_not_depend_on_row_order():
bank = [row(i, 6 * (i - 1), used=(i * 37) % 11, uses=i % 3) for i in range(1, 30)]
bank += [row(40, 12), row(41, 12)] # a shared start with identical timestamps
expected = memorybank.eviction_order(bank, 20)
for seed in range(5):
shuffled = bank[:]
random.Random(seed).shuffle(shuffled)
assert memorybank.eviction_order(shuffled, 20) == expected
def test_a_tie_on_every_signal_is_broken_by_id():
bank = [row(1, 0), row(9, 6, created=0), row(4, 12, created=0), row(20, 18)]
assert memorybank.eviction_order(bank, 1) == [4]
@pytest.mark.parametrize("seed", range(8))
def test_capacity_holds_and_pins_survive_for_any_bank(seed):
rng = random.Random(seed)
bank = []
for i in range(1, rng.randint(2, 60)):
start = None if rng.random() < 0.15 else rng.randrange(0, 400)
bank.append(row(i, start, None if start is None else start + rng.choice([3, 5, 8]),
pinned=rng.random() < 0.1, used=rng.choice([None, rng.randrange(500)]),
uses=rng.randrange(4)))
capacity = rng.randint(1, 30)
overflow = len(bank) - capacity
doomed = memorybank.eviction_order(bank, max(0, overflow))
pinned = {r.id for r in bank if r.pinned}
assert not pinned & set(doomed)
assert len(set(doomed)) == len(doomed)
remaining = len(bank) - len(doomed)
assert remaining == max(capacity, len(pinned)) if overflow > 0 else remaining == len(bank)
@pytest.mark.parametrize("n, capacity, irregular", [(60, 10, False), (500, 80, False), (500, 80, True)])
def test_a_long_bank_keeps_describing_the_whole_story(n, capacity, irregular):
"""The general property, with nothing ever retrieved: memories arrive one
block at a time and the bank is kept at capacity. The opening stays, and no
stretch goes undescribed for more than twice the average spacing. Least
recently used order, on the same arrivals, keeps only the newest stretch."""
rng = random.Random(n)
kept, lru = [], []
depth = 0
for i in range(1, n + 1):
size = rng.choice([4, 6, 6, 9]) if irregular else 6
memory = row(i, depth, depth + size - 1)
depth += size
kept.append(memory)
lru.append(memory)
if len(kept) > capacity:
doomed = set(memorybank.eviction_order(kept, len(kept) - capacity))
kept = [m for m in kept if m.id not in doomed]
lru = sorted(lru, key=lambda m: (m.created_at, m.id))[len(lru) - capacity:]
assert len(kept) == capacity
assert min(m.source_start for m in kept) == 0
assert max(m.id for m in kept) == n
assert largest_gap_from_opening(kept) <= 2 * depth / capacity
assert largest_gap_from_opening(lru) > depth / 2 # v1.0.0 order: the opening is gone
# --------------------------------------------------------- on real rows
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
session = SessionLocal()
try:
yield session
finally:
session.close()
memorybank._vector_cache.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture()
def adventure(db):
user = models.User(is_guest=False, email="b2-evict@example.com")
db.add(user)
db.flush()
settings = models.Settings(user_id=user.id, model="m", embedding_model="e",
memory_bank_capacity=3)
adv = models.Adventure(user_id=user.id, title="Evict", script_state={}, memory_bank_enabled=True)
db.add_all([settings, adv])
db.commit()
adv.settings_row = settings
return adv
def test_the_pass_changes_nothing_but_forgotten(db, adventure):
for i in range(6):
memory = models.Memory(adventure_id=adventure.id, text=f"block {i}",
source_start=6 * i, source_end=6 * i + 5, branch_id=None, depth=6 * i + 5)
db.add(memory)
db.commit()
columns = (models.Memory.id, models.Memory.text, models.Memory.source_start,
models.Memory.source_end, models.Memory.branch_id, models.Memory.depth,
models.Memory.pinned, models.Memory.use_count, models.Memory.last_used_at)
before = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
after = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
assert before == after
active = db.execute(select(models.Memory.id).where(models.Memory.forgotten.is_(False))).scalars().all()
assert len(active) == 3
assert min(active) == min(before) and max(active) == max(before) # the boundaries
def test_eviction_does_not_make_an_abandoned_lines_memory_eligible(db, adventure):
"""Eviction and lineage are separate: the pass decides only `forgotten`, so
a memory on a line the story left is exactly as ineligible afterwards."""
trunk = []
for i in range(4):
action = models.Action(adventure_id=adventure.id, type="ai", text=f"trunk {i}")
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
trunk.append(action)
abandoned_node = models.Action(adventure_id=adventure.id, type="ai", text="the abandoned line")
tree.place_action(db, adventure, abandoned_node)
db.add(abandoned_node)
db.flush()
abandoned = models.Memory(adventure_id=adventure.id, text="on the abandoned line",
source_start=4, source_end=4)
tree.attach_memory(abandoned, abandoned_node)
db.add(abandoned)
db.commit()
# Move the head back and diverge, so the abandoned node is off the path.
adventure.head_depth = trunk[-1].depth
db.commit()
from app import head
head.fork_if_behind_head(db, adventure)
divergent = models.Action(adventure_id=adventure.id, type="ai", text="the new line")
tree.place_action(db, adventure, divergent)
db.add(divergent)
db.flush()
for i, node in enumerate(trunk + [divergent]):
memory = models.Memory(adventure_id=adventure.id, text=f"active {i}",
source_start=node.depth, source_end=node.depth)
tree.attach_memory(memory, node)
db.add(memory)
db.commit()
def eligible():
return set(db.execute(select(models.Memory.id).where(
models.Memory.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Memory),
models.Memory.forgotten.is_(False))).scalars().all())
assert abandoned.id not in eligible()
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
db.expire_all()
assert abandoned.id not in eligible()
assert len(db.execute(select(models.Memory.id).where(
models.Memory.forgotten.is_(False))).scalars().all()) == 3
+189
View File
@@ -0,0 +1,189 @@
"""v1.1 WP-B.2 (B2.3): what the memory summariser is shown of a long block.
v1.0.0 sent the last 2,000 tokens of a block, so a fact early in a longer block
never reached the summariser (B.1 §E). A block that fits is still sent whole. A
longer one is now sent as its opening and its end, with a marker between them,
inside the same 2,000-token budget.
The scenario-level test (the planted fact early in a long block, remembered) is
in `test_v11_b1_memory_diagnostic.py`.
python -m pytest tests/test_v11_b2_memory_excerpt.py -v
"""
import asyncio
import random
import pytest
from app import memorybank, models, tree
from app.context import builder
from app.database import Base, SessionLocal, engine
BUDGET = memorybank.MEMORY_EXCERPT_TOKENS
MARKER = memorybank.EXCERPT_OMISSION_MARKER
FILLER = "The travellers walked the long grey road north past the salt market and the reed beds. "
def words_to_tokens(tokens: int) -> str:
"""Filler at least `tokens` long."""
text = FILLER
while builder.count_tokens(text) < tokens:
text += FILLER
return text
# --------------------------------------------------------------- the excerpt
def test_a_block_that_fits_is_sent_whole_and_unchanged():
raw = words_to_tokens(BUDGET - 200)
assert builder.count_tokens(raw) <= BUDGET
assert memorybank.memory_excerpt(raw) == raw
def test_a_block_of_exactly_the_budget_is_unchanged():
raw = words_to_tokens(BUDGET)
tokens = memorybank._excerpt_encoding().encode(raw)[:BUDGET]
exact = memorybank._excerpt_encoding().decode(tokens)
if builder.count_tokens(exact) == BUDGET:
assert memorybank.memory_excerpt(exact) == exact
def test_a_long_block_keeps_its_opening_and_its_end_in_order():
opening = "Mara slipped the amber sundial inside the cracked teapot. "
ending = "Aldric finally reached the north gate at dawn."
raw = opening + words_to_tokens(3 * BUDGET) + ending
excerpt = memorybank.memory_excerpt(raw)
assert excerpt.startswith(opening)
assert excerpt.endswith(ending)
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
assert tail, "the marker must sit between the two parts"
assert excerpt.index(opening) < excerpt.index(MARKER) < excerpt.index(ending)
def test_the_split_is_even_and_documented():
raw = words_to_tokens(4 * BUDGET)
head_budget, tail_budget = memorybank.excerpt_split(BUDGET)
marker_tokens = builder.count_tokens(f"\n\n{MARKER}\n\n")
assert head_budget + tail_budget + marker_tokens == BUDGET
assert abs(head_budget - tail_budget) <= 1
excerpt = memorybank.memory_excerpt(raw)
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
# Each part is cut as a run of `head_budget` / `tail_budget` tokens. Measured
# on its own, a cut run can come to one token more, because the text either
# side of the cut tokenises differently once it is separated; the hard limit
# is the whole excerpt, tested below.
assert builder.count_tokens(head) <= head_budget + 1
assert builder.count_tokens(tail) <= tail_budget + 1
assert builder.count_tokens(excerpt) <= BUDGET
@pytest.mark.parametrize("extra", [1, 7, 500, BUDGET, 9 * BUDGET])
def test_the_excerpt_never_exceeds_the_budget(extra):
raw = words_to_tokens(BUDGET + extra)
excerpt = memorybank.memory_excerpt(raw)
assert builder.count_tokens(excerpt) <= BUDGET
assert memorybank.memory_excerpt(raw) == excerpt # deterministic
@pytest.mark.parametrize("seed", range(4))
def test_the_budget_holds_for_awkward_text(seed):
"""Token boundaries can merge differently once the parts are rejoined, and
text that is not plain English tokenises unevenly. The budget still holds."""
rng = random.Random(seed)
alphabet = "abcdefghij ÄÖÜ ßé漢字かな 🙂🐉 \n\t.,;:—'\""
raw = "".join(rng.choice(alphabet) for _ in range(12000))
assert builder.count_tokens(memorybank.memory_excerpt(raw)) <= BUDGET
def test_a_fact_in_the_middle_of_a_very_long_block_is_still_omitted():
"""The documented limit of a bounded excerpt: head and tail, not everything."""
half = words_to_tokens(3 * BUDGET)
raw = half + "Mara slipped the amber sundial inside the cracked teapot. " + half
assert "sundial" not in memorybank.memory_excerpt(raw)
# ------------------------------------------------------------ memory creation
class EchoSummariser:
"""Returns the whole excerpt it was given as the memory: the worst case for
a marker leaking into stored text."""
def __init__(self):
self.users: list[str] = []
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
self.users.append(user)
return user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
session = SessionLocal()
try:
yield session
finally:
session.close()
Base.metadata.drop_all(bind=engine)
def campaign(db, texts):
user = models.User(is_guest=False, email="b2-excerpt@example.com")
db.add(user)
db.flush()
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
adventure = models.Adventure(user_id=user.id, title="Excerpt", script_state={},
auto_summarize=True, memory_bank_enabled=True)
db.add(adventure)
db.flush()
nodes = []
for i, text in enumerate(texts):
action = models.Action(adventure_id=adventure.id, type="ai" if i % 2 else "do", text=text)
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
nodes.append(action)
db.commit()
return adventure, nodes
def write_memory(db, adventure, monkeypatch, summariser):
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
monkeypatch.setattr(memorybank, "summary_provider", lambda s: summariser)
settings = db.query(models.Settings).first()
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
def test_the_marker_is_never_stored_as_part_of_a_memory(db, monkeypatch):
long = words_to_tokens(800)
adventure, _ = campaign(db, [long] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK))
summariser = EchoSummariser()
[memory] = write_memory(db, adventure, monkeypatch, summariser)
assert MARKER in summariser.users[0] # the summariser was told
assert MARKER not in memory.text # and the memory does not repeat it
assert "[…" not in memory.text and "omitted" not in memory.text
def test_a_short_block_is_prompted_exactly_as_before(db, monkeypatch):
texts = [f"Short action {i}." for i in range(memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK)]
adventure, _ = campaign(db, texts)
summariser = EchoSummariser()
write_memory(db, adventure, monkeypatch, summariser)
block = "\n\n".join(texts[:memorybank.MEMORY_INTERVAL])
assert summariser.users[0] == f"Story excerpt:\n\n{block}\n\nMemory:"
def test_a_long_blocks_memory_keeps_its_source_provenance(db, monkeypatch):
early = "Mara slipped the amber sundial inside the cracked teapot. " + words_to_tokens(900)
texts = [early] + [words_to_tokens(900)] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK - 1)
adventure, nodes = campaign(db, texts)
[memory] = write_memory(db, adventure, monkeypatch, EchoSummariser())
block = nodes[:memorybank.MEMORY_INTERVAL]
assert "sundial" in memory.text # the early fact reached the summariser
assert (memory.source_start, memory.source_end) == (block[0].depth, block[-1].depth)
assert (memory.branch_id, memory.depth) == (block[-1].branch_id, block[-1].depth)
+326
View File
@@ -0,0 +1,326 @@
"""v1.1 WP-B.2 (B2.1): what memory retrieval searches for, and how it scores.
B.1 found the planting-era memory created and retained but ranked out of
`memory_top_k`, because the query was three turns of narration with the player's
question at the end. The query is now the player's input plus a short scene
context, and the score adds one transparent lexical term over the input.
These tests pin the pieces. The end-to-end fixture tests (crowded bank,
context-dependent question, negative controls) are in
`test_v11_b1_memory_diagnostic.py`, beside the diagnostic they use.
python -m pytest tests/test_v11_b2_memory_ranking.py -v
"""
import asyncio
import math
import pytest
from sqlalchemy import event
from app import memorybank, models, tree
from app.context import builder
from app.database import Base, SessionLocal, engine
# ------------------------------------------------------------------ lexical
def test_terms_are_folded_but_not_stemmed():
terms = memorybank.lexical_terms("The tavern's teapots, the glass and the SUNDIAL")
assert {"tavern", "teapot", "glass", "sundial"} <= terms
assert "the" not in terms # the knowledge path's stop list
assert "glas" not in terms # a word ending in "ss" is not a plural
def test_a_word_every_candidate_holds_weighs_nothing():
scores = memorybank.lexical_scores(
frozenset({"travellers"}), {1: frozenset({"travellers", "road"}), 2: frozenset({"travellers"})})
assert scores == {1: 0.0, 2: 0.0}
def test_a_rarer_word_weighs_more_than_a_common_one():
scores = memorybank.lexical_scores(
frozenset({"sundial", "road"}),
{1: frozenset({"sundial"}), 2: frozenset({"road"}), 3: frozenset({"road"}),
4: frozenset({"gate"})})
assert scores[1] > scores[2] == scores[3] > scores[4] == 0.0
def test_a_single_rare_word_of_a_longer_question_is_only_its_share():
"""The share is over the whole question, so one incidental word match
cannot score like a memory that answers it."""
question = frozenset({"where", "amber", "sundial", "fish"})
scores = memorybank.lexical_scores(
question, {1: frozenset({"fish"}), 2: frozenset({"amber", "sundial"}), 3: frozenset({"road"})})
assert 0.0 < scores[1] < scores[2] <= 1.0
assert scores[1] < 0.5
def test_scores_are_bounded_and_empty_inputs_score_zero():
candidates = {1: frozenset({"a1", "b2"}), 2: frozenset({"a1"})}
assert all(0.0 <= v <= 1.0 for v in memorybank.lexical_scores(frozenset({"a1", "b2"}), candidates).values())
assert memorybank.lexical_scores(frozenset(), candidates) == {1: 0.0, 2: 0.0}
assert memorybank.lexical_scores(frozenset({"a1"}), {}) == {}
# ------------------------------------------------------------------ scoring
def _unit(angle):
return [math.cos(angle), math.sin(angle)]
def test_ties_are_broken_by_id_not_by_row_order():
held = {7: [1.0, 0.0], 3: [1.0, 0.0], 5: [1.0, 0.0]}
rows = memorybank.score_candidates([7, 3, 5], held, {}, [1.0, 0.0], None, [])
assert [row[1] for row in rows] == [3, 5, 7]
def test_the_semantic_score_mixes_input_and_context_by_the_fixed_weight():
held = {1: [1.0, 0.0]}
[(final, _, semantic, lexical)] = memorybank.score_candidates(
[1], held, {}, [1.0, 0.0], [0.0, 1.0], [])
assert semantic == pytest.approx(memorybank.INPUT_WEIGHT)
assert final == semantic and lexical == 0.0
# Either part alone is used as it is.
[(_, _, only_context, _)] = memorybank.score_candidates([1], held, {}, None, [0.0, 1.0], [])
assert only_context == pytest.approx(0.0)
@pytest.mark.parametrize("margin, relevant_first", [(0.01, True), (-0.01, False)])
def test_a_lexical_match_moves_a_memory_at_most_the_lexical_weight(margin, relevant_first):
"""The bound that keeps rarity from overruling meaning: a memory more than
`LEXICAL_WEIGHT` behind semantically cannot pass one ahead of it, however
rare the word it shares."""
decoy_cos = 1.0 - memorybank.LEXICAL_WEIGHT - margin
held = {1: [1.0, 0.0], 2: _unit(math.acos(decoy_cos))}
terms = {1: frozenset(), 2: frozenset({"zeppelin"})}
rows = memorybank.score_candidates([1, 2], held, terms, [1.0, 0.0], None, ["zeppelin"])
order = [row[1] for row in rows]
assert (order[0] == 1) is relevant_first
# ---------------------------------------------------- the query, on real rows
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
memorybank._terms_cache.clear()
session = SessionLocal()
try:
yield session
finally:
session.close()
memorybank._vector_cache.clear()
memorybank._terms_cache.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture(autouse=True)
def restore_embedding_provider():
real = memorybank.embedding_provider
try:
yield
finally:
memorybank.embedding_provider = real
@pytest.fixture()
def settings(db):
user = models.User(is_guest=False, email="b2-rank@example.com")
db.add(user)
db.flush()
row = models.Settings(user_id=user.id, model="m", embedding_model="stub-embed",
memory_top_k=2, memory_bank_capacity=80)
db.add(row)
db.commit()
return row
SCENE_STATE = {
"entities": {"mara": {"type": "character", "name": "Mara"},
"tavern": {"type": "location", "name": "The Crooked Lantern"}},
"scene": {"summary": "Closing time", "location": "tavern", "present": ["mara"]},
}
def make_adventure(db, settings, texts, state=None):
"""`texts` is `[(type, text)]`, oldest first, each placed on the tree."""
adventure = models.Adventure(user_id=settings.user_id, title="Rank", script_state={},
memory_bank_enabled=True, narrative_state=state or {})
db.add(adventure)
db.flush()
placed = []
for kind, text in texts:
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
placed.append(action)
db.commit()
return adventure, placed
def test_the_query_is_the_players_input_and_the_scene(db, settings):
adventure, _ = make_adventure(db, settings, [
("start", "Rain over the harbour."),
("ai", "Mara wipes down the counter and glances up at the shelf."),
("do", "> You ask Mara about the brass dial."),
], state=SCENE_STATE)
query = memorybank.retrieval_query(adventure)
assert query["input"] == "> You ask Mara about the brass dial."
assert "The Crooked Lantern" in query["context"] and "Mara" in query["context"]
assert "glances up at the shelf" in query["context"]
assert "brass" in query["input_terms"] and "dial" in query["input_terms"]
def test_a_continue_turn_has_no_input_and_searches_by_the_scene(db, settings):
adventure, _ = make_adventure(db, settings, [
("do", "> You sit down."),
("ai", "The fire burns low in the grate."),
])
query = memorybank.retrieval_query(adventure)
assert query["input"] == "" and query["input_terms"] == []
assert "fire burns low" in query["context"]
def test_a_retry_searches_with_the_input_it_is_retrying(db, settings):
adventure, placed = make_adventure(db, settings, [
("ai", "The market is quiet."),
("do", "> You ask about the sundial."),
("ai", "A discarded attempt about lanterns."),
])
query = memorybank.retrieval_query(adventure, exclude_action_id=placed[-1].id)
assert query["input"] == "> You ask about the sundial."
assert "lanterns" not in query["context"]
assert "market is quiet" in query["context"]
def test_the_query_is_bounded_however_long_the_story(db, settings):
long = "The travellers walked the long grey road north past the salt market. " * 400
adventure, _ = make_adventure(db, settings, [
("ai", long), ("story", long)], state=SCENE_STATE)
query = memorybank.retrieval_query(adventure)
assert builder.count_tokens(query["input"]) <= memorybank.QUERY_INPUT_TOKENS
assert builder.count_tokens(query["context"]) <= (
memorybank.QUERY_SCENE_TOKENS + memorybank.QUERY_NARRATION_TOKENS + 2)
# ------------------------------------------------ retrieval, end to end
class SameVector:
"""Every text embeds the same, so only the lexical term separates memories."""
async def embed(self, texts):
return [[1.0, 0.0, 0.0] for _ in texts]
def add_memory(db, adventure, text, **kwargs):
memory = models.Memory(adventure_id=adventure.id, text=text, **kwargs)
db.add(memory)
db.flush()
memorybank.set_vector(memory, [1.0, 0.0, 0.0])
db.commit()
return memory
def retrieve(adventure, settings, **kwargs):
memorybank.embedding_provider = lambda s: SameVector()
return asyncio.run(memorybank.retrieve_memories(adventure, settings, **kwargs))
@pytest.fixture()
def played(db, settings):
adventure, _ = make_adventure(db, settings, [
("ai", "The tavern is warm."),
("do", "> You ask Mara where the amber sundial went."),
], state=SCENE_STATE)
bank = {
"road": add_memory(db, adventure, "Aldric walked the north road."),
"sundial": add_memory(db, adventure, "Mara hid the amber sundial in the teapot."),
"gate": add_memory(db, adventure, "The gate guard asked for a toll."),
}
return adventure, bank
def test_every_used_memory_reports_the_parts_of_its_score(db, settings, played):
adventure, bank = played
result = retrieve(adventure, settings)
first = result["used"][0]
assert first["id"] == bank["sundial"].id
assert first["similarity"] == first["semantic_score"]
assert first["lexical_score"] > 0
assert first["final_score"] == pytest.approx(
first["semantic_score"] + memorybank.LEXICAL_WEIGHT * first["lexical_score"], abs=2e-4)
assert result["query"]["input"] == "> You ask Mara where the amber sundial went."
assert result["query"]["lexical_weight"] == memorybank.LEXICAL_WEIGHT
assert result["query"]["input_weight"] == memorybank.INPUT_WEIGHT
def test_a_pin_is_still_always_used_and_counts_toward_top_k(db, settings, played):
adventure, bank = played
settings.memory_top_k = 1
bank["gate"].pinned = True
db.commit()
used = retrieve(adventure, settings)["used"]
assert [m["id"] for m in used] == [bank["gate"].id]
assert used[0]["pinned"] is True
def memory_text_reads(statements):
return [s for s in statements
if s.lstrip().upper().startswith("SELECT") and "FROM memories" in s
and "memories.text" in s]
@pytest.fixture()
def sql_log():
statements: list[str] = []
def record(conn, cursor, statement, parameters, context, executemany):
statements.append(statement)
event.listen(engine, "before_cursor_execute", record)
try:
yield statements
finally:
event.remove(engine, "before_cursor_execute", record)
def test_memory_text_is_read_once_and_then_held(db, settings, played, sql_log):
adventure, _ = played
retrieve(adventure, settings)
sql_log.clear()
result = retrieve(adventure, settings)
reads = memory_text_reads(sql_log)
# Only the detail read of the memories chosen remains. (Every memory here
# embeds identically, so redundancy suppression keeps just one of them.)
assert len(reads) == 1 and reads[0].count("?") == len(result["used"])
def test_a_continue_turn_reads_no_memory_text_to_rank(db, settings, sql_log):
adventure, _ = make_adventure(db, settings, [("do", "> You wait."), ("ai", "Night falls.")])
for text in ("one", "two", "three"):
add_memory(db, adventure, f"memory {text}")
sql_log.clear()
result = retrieve(adventure, settings)
assert all(m["lexical_score"] == 0.0 for m in result["used"])
assert len(memory_text_reads(sql_log)) == 1 # the top-k detail read only
def test_an_edited_memory_is_matched_on_its_new_text(db, settings, played):
adventure, bank = played
assert retrieve(adventure, settings)["used"][0]["id"] == bank["sundial"].id
# An edit clears the vector (the route calls set_vector(None)); re-embedding
# sets it again. Both go through set_vector, which drops the held terms.
bank["road"].text = "The amber sundial was traded for the road toll."
memorybank.set_vector(bank["road"], None)
memorybank.set_vector(bank["road"], [1.0, 0.0, 0.0])
bank["sundial"].text = "Mara hid a bottle in the cellar."
memorybank.set_vector(bank["sundial"], None)
memorybank.set_vector(bank["sundial"], [1.0, 0.0, 0.0])
db.commit()
assert retrieve(adventure, settings)["used"][0]["id"] == bank["road"].id
@@ -0,0 +1,225 @@
"""v1.1 WP-B.2: the memory summariser, after the rejected B2.4 prompt experiment.
B2.4 tried a memory prompt instructing the model to keep named facts and objects.
Measured against the reference model, it did not correct the creation failure it
was for, and it was not shipped (`V1.1-WP-B2-REPORT.md` §T). The shipped prompt
is v1.0.0's.
This file keeps two kinds of test apart.
**Acceptance tests** gate the tree:
- the shipped memory prompt is exactly v1.0.0's, so the experiment is gone;
- every fidelity fixture reaches the summariser whole, through the application's
own prompt assembly;
- a long memory is stored as the model wrote it, never cut;
- the memory the attempt-2 block should have produced ranks first under B2.1.
**Diagnostic-measurement tests** check only that `tools/memory_fidelity.py`
measures correctly: fact retention, attribution, invention, word count, a leading
"Memory:", second person and promise retention, on hand-written memories whose
answers are known. What a real model scores on those measurements is
nondeterministic, is taken with inference, and is reported. It is never a gate
here.
python -m pytest tests/test_v11_b2_summarizer_fidelity.py -v
"""
import asyncio
import re
import subprocess
import pytest
from app import memorybank, models, tree
from app.database import Base, SessionLocal, engine
from tools import memory_diagnostic as md
from tools import memory_fidelity as mf
# ==================================================================== acceptance
def test_the_shipped_memory_prompt_is_v1_0_0s():
"""The B2.4 experiment is reverted: production sends the prompt v1.0.0 and
WP-B.1 shipped, unchanged."""
try:
source = subprocess.run(["git", "show", "beb17ad:backend/app/memorybank.py"],
capture_output=True, text=True, check=True).stdout
except (OSError, subprocess.CalledProcessError):
pytest.skip("git history not available")
block = re.search(r"^MEMORY_SYSTEM_PROMPT = \((.*?)^\)$", source, re.S | re.M).group(1)
shipped = eval(f"({block})", {"MEMORY_MAX_WORDS": 50}) # noqa: S307 - our own source
assert memorybank.MEMORY_SYSTEM_PROMPT == shipped
assert memorybank.MEMORY_MAX_WORDS == 50
def test_the_rejected_experiment_is_not_what_ships():
assert mf.B24_EXPERIMENT_PROMPT != memorybank.MEMORY_SYSTEM_PROMPT
assert "Keep each fact with the person it belongs to" not in memorybank.MEMORY_SYSTEM_PROMPT
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
def test_every_fixture_reaches_the_summariser_whole(fixture):
"""Creation can only fail at the model if the fact was sent. Each fixture
fits the excerpt budget, so the whole block is the excerpt."""
user = mf.user_prompt_for(fixture)
assert memorybank.count_tokens(fixture.raw) <= memorybank.MEMORY_EXCERPT_TOKENS
assert f"Story excerpt:\n\n{fixture.raw}\n\nMemory:" in user
assert user.startswith("Cast:\n- " + fixture.protagonist + " — the protagonist.")
class Scripted:
def __init__(self, reply):
self.reply = reply
self.calls: list[tuple[str, str]] = []
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
self.calls.append((system, user))
return self.reply
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
session = SessionLocal()
try:
yield session
finally:
session.close()
Base.metadata.drop_all(bind=engine)
def campaign(db, fixture):
user = models.User(is_guest=False, email="b2-fidelity@example.com")
db.add(user)
db.flush()
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
adventure = models.Adventure(user_id=user.id, title="Fidelity", script_state={}, auto_summarize=True,
persona_name=fixture.protagonist)
db.add(adventure)
db.flush()
nodes = []
for kind, text in fixture.actions + (("ai", "The story moves on."),):
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
nodes.append(action)
db.commit()
return adventure, nodes
def write_memory(db, adventure, monkeypatch, provider):
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
monkeypatch.setattr(memorybank, "summary_provider", lambda s: provider)
settings = db.query(models.Settings).first()
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
def test_the_application_sends_the_shipped_prompt_and_the_whole_planting_block(db, monkeypatch):
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
adventure, nodes = campaign(db, fixture)
provider = Scripted(fixture.faithful)
[memory] = write_memory(db, adventure, monkeypatch, provider)
system, user = provider.calls[0]
assert system == memorybank.MEMORY_SYSTEM_PROMPT
assert "> I watch Mara slip the amber sundial inside the cracked teapot" in user
assert (memory.source_start, memory.source_end) == (nodes[0].depth, nodes[5].depth)
def test_an_over_long_memory_is_stored_as_written_never_cut(db, monkeypatch):
"""The word target is an instruction, not a truncation: cutting a memory
after the fact can split or drop exactly the fact it was written to keep."""
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
adventure, _ = campaign(db, fixture)
long_reply = fixture.faithful + " " + " ".join(["They advanced cautiously through the dark."] * 12)
[memory] = write_memory(db, adventure, monkeypatch, Scripted(long_reply))
assert memory.text == long_reply
assert len(memory.text.split()) > 2 * memorybank.MEMORY_MAX_WORDS
def test_a_faithful_regression_memory_ranks_first_for_its_question():
"""If the summariser keeps the fact, B2.1 finds it: the memory the attempt-2
block should have produced, among the memories its bank really held for that
stretch, under production scoring."""
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
stored, _ = fixture.unfaithful[0]
bank = {
1: fixture.faithful,
2: stored,
3: "Aldric, Mara and Edrin advanced through the cold crypt, the silver key heavy in Aldric's hands.",
4: "Aldric told Mara the silver key opens the crypt beneath the Old Abbey.",
5: "Rain kept falling on Westhaven as the travellers walked toward the abbey grounds.",
}
embed = md.ConceptEmbedder.vector
query = {"input": "> I ask Mara quietly where she hid the amber sundial.",
"context": "Aldric and Mara in the Crooked Lantern, rain outside."}
held = {i: embed(t) for i, t in bank.items()}
terms = {i: memorybank.lexical_terms(t) for i, t in bank.items()}
rows = memorybank.score_candidates(list(bank), held, terms, embed(query["input"]),
embed(query["context"]),
sorted(memorybank.lexical_terms(query["input"])))
assert rows[0][1] == 1
assert rows[0][3] > 0
# ======================================================= diagnostic measurements
# These prove the measuring instrument. They say nothing about any model.
def test_the_fixtures_cover_each_measurement_in_more_than_one_genre():
requirements = {f.requirement for f in mf.FIXTURES}
assert {"distinctive object and place", "player-established concrete fact", "promise / commitment",
"attribution", "clutter pressure", "no invention", "multiple concrete facts",
"the actual failed-run block"} <= requirements
assert {"office", "contemporary", "science-fiction-neutral"} <= {f.genre for f in mf.FIXTURES}
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
def test_the_checker_passes_a_faithful_memory(fixture):
result = mf.evaluate(fixture, fixture.faithful)
assert result["passed"], result
assert not result["over_target"] and not result["memory_prefix"] and not result["second_person"]
@pytest.mark.parametrize("fixture, memory, reason", [
(f, memory, reason) for f in mf.FIXTURES for memory, reason in f.unfaithful
], ids=lambda v: v.fixture_id if isinstance(v, mf.Fixture) else None)
def test_the_checker_fails_each_failure_shape(fixture, memory, reason):
result = mf.evaluate(fixture, memory)
assert not result["passed"], result
if reason == "not retained":
assert not result["retained"]
elif reason == "misattributed":
assert result["misattributed"]
elif reason == "invented":
assert result["inventions"]
def test_the_checker_reads_the_stored_attempt_2_memory_as_the_real_failure():
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
stored, _ = fixture.unfaithful[0]
result = mf.evaluate(fixture, stored)
assert result["retained"] is False and result["words"] == 102 and result["over_target"]
@pytest.mark.parametrize("memory, prefix, you", [
("Memory: Dana promised Marcus the lease by Friday.", True, False),
(" memory: Dana promised the lease.", True, False),
("You thanked Marcus and left.", False, True),
("Dana thanked Marcus; your lease is due.", False, True),
("Dana promised Marcus she would bring the signed lease by Friday.", False, False),
])
def test_the_checker_measures_framing(memory, prefix, you):
result = mf.evaluate(mf.FIXTURES_BY_ID["promise_contemporary"], memory)
assert result["memory_prefix"] is prefix
assert result["second_person"] is you
def test_the_checker_measures_promise_retention():
fixture = mf.FIXTURES_BY_ID["promise_contemporary"]
kept = mf.evaluate(fixture, "Dana promised to bring Marcus the signed lease by Friday.")
scenery = mf.evaluate(fixture, "Memory: Dana looked around the empty living room while a dog barked.")
assert kept["facts"]["lease by Friday"]["kept"] and kept["passed"]
assert not scenery["facts"]["lease by Friday"]["kept"] and scenery["memory_prefix"]