v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
beb17ada10
commit
0c1ba836ba
@@ -11,8 +11,9 @@ diagnostic in `tools/memory_diagnostic.py`:
|
||||
2. **What it finds on this tree.** The scenarios run with a best-case summariser,
|
||||
one that keeps a fact if and only if the fact reached it. Any failure is
|
||||
therefore the application's mechanism, not a model's writing.
|
||||
- The criteria the current code does not meet are marked `xfail(strict=True)`,
|
||||
so B.2 has to flip them deliberately.
|
||||
- WP-B.1 marked the criteria v1.0.0 did not meet `xfail(strict=True)`. WP-B.2
|
||||
fixed ranking, eviction and the creation excerpt, and those tests are now
|
||||
ordinary passes; the v1.0.0 results are recorded in the WP-B.2 report.
|
||||
- The same file is run unchanged against v1.0.0 for the baseline.
|
||||
|
||||
python -m pytest tests/test_v11_b1_memory_diagnostic.py -v
|
||||
@@ -102,14 +103,14 @@ def test_retention_is_reported_with_the_bank_and_its_eviction_order():
|
||||
def test_ranking_is_production_ranking_and_agrees_with_the_stored_selection():
|
||||
ranked = scenario("independent_default")["diagnosis"]["ranked"]
|
||||
assert ranked["replica_matches_stored_selection"] is True
|
||||
assert ranked["lexical_score"] is None # memory ranking has no lexical term
|
||||
assert ranked["top_k_cutoff"] == 5
|
||||
assert ranked["yes"] is True and ranked["selected"] is True
|
||||
assert 1 <= ranked["rank"] <= ranked["top_k_cutoff"]
|
||||
# The production query is the newest four actions, cut to 600 tokens, and the
|
||||
# one-line question is diluted by the narration around it.
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert ranked["semantic_score"] < variants["direct"]["similarity"]
|
||||
# v1.1 WP-B.2: every part of the score is reported, and they add up.
|
||||
assert 0.0 <= ranked["lexical_score"] <= 1.0
|
||||
assert ranked["final_score"] == pytest.approx(
|
||||
ranked["semantic_score"] + memorybank.LEXICAL_WEIGHT * ranked["lexical_score"], abs=2e-4)
|
||||
assert ranked["query"]["input"].endswith(md.SCENARIOS["independent_default"].recall_text)
|
||||
|
||||
|
||||
def test_injection_is_read_from_the_recall_turns_own_context():
|
||||
@@ -124,53 +125,118 @@ def test_ranking_variants_direct_paraphrase_and_unrelated():
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["direct"]["rank"] == 1 and variants["direct"]["selected"]
|
||||
assert variants["paraphrase"]["rank"] == 1 and variants["paraphrase"]["selected"]
|
||||
assert (variants["direct"]["similarity"] > variants["paraphrase"]["similarity"]
|
||||
> 5 * variants["unrelated"]["similarity"])
|
||||
assert (variants["direct"]["final_score"] > variants["paraphrase"]["final_score"]
|
||||
> variants["unrelated"]["final_score"])
|
||||
|
||||
|
||||
def test_retrieval_fills_top_k_whatever_the_similarity():
|
||||
"""Diagnosis: there is no relevance floor. With more memories than
|
||||
`memory_top_k`, an unrelated query still selects five, and the early fact
|
||||
rides along at a similarity near zero."""
|
||||
def test_retrieval_still_fills_top_k_whatever_the_similarity():
|
||||
"""There is still no relevance floor: an unrelated question selects a full
|
||||
`memory_top_k`. B.2 changed which memories those are, not how many — the
|
||||
early fact is no longer carried along by an unrelated question."""
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["unrelated"]["similarity"] < 0.1
|
||||
assert variants["unrelated"]["selected"] is True
|
||||
assert variants["unrelated"]["selected_count"] == 5
|
||||
assert variants["unrelated"]["selected"] is False
|
||||
assert variants["unrelated"]["rank"] > 5
|
||||
|
||||
|
||||
# ------------------------------------------------- WP-B.2 ranking acceptance
|
||||
|
||||
@pytest.mark.parametrize("name", ["ranking_crowded", "ranking_context_dependent"])
|
||||
def test_acceptance_the_early_memory_is_ranked_and_injected_below_capacity(name):
|
||||
"""The B.1 ranking failure, made deterministic. On v1.0.0 both fixtures are
|
||||
`retained_but_not_ranked` (ranks 7 and 6 of 17 against a top-k of 4)."""
|
||||
result = scenario(name)
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
diagnosis = result["diagnosis"]
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert diagnosis["retained"]["active_memories"] <= diagnosis["retained"]["memory_bank_capacity"]
|
||||
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["rank"] <= 4
|
||||
assert diagnosis["ranked"]["replica_matches_stored_selection"] is True
|
||||
assert diagnosis["injected"]["yes"] is True
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_the_crowded_fixture_is_won_by_the_players_question():
|
||||
ranked = scenario("ranking_crowded")["diagnosis"]["ranked"]
|
||||
assert ranked["rank"] == 1
|
||||
assert ranked["lexical_score"] > 0 # "sundial" and "amber" are in the question
|
||||
|
||||
|
||||
def test_a_paraphrase_is_found_by_meaning_not_by_shared_words():
|
||||
"""Lexical matching must not replace semantic retrieval. The paraphrase
|
||||
shares none of F's distinctive words, yet ranks first."""
|
||||
for name in ("ranking_crowded", "independent_default"):
|
||||
paraphrase = scenario(name)["ranking_variants"]["paraphrase"]
|
||||
assert paraphrase["rank"] == 1 and paraphrase["selected"]
|
||||
# Only "Mara" is shared, which is far less than the direct question holds.
|
||||
direct = scenario(name)["ranking_variants"]["direct"]
|
||||
assert paraphrase["lexical_score"] < direct["lexical_score"] / 2
|
||||
|
||||
|
||||
def test_a_context_dependent_question_needs_the_scene():
|
||||
""""I ask her what she keeps up there" names nothing F's memory holds. The
|
||||
scene the last narration set up (Mara, the top shelf, a kettle) is what
|
||||
finds it; without that context it ranks last."""
|
||||
result = scenario("ranking_context_dependent")
|
||||
ranked = result["diagnosis"]["ranked"]
|
||||
assert ranked["lexical_score"] == 0.0
|
||||
assert ranked["rank"] <= 4
|
||||
assert "top shelf" in ranked["query"]["context"]
|
||||
assert result["ranking_variants"]["input_only"]["rank"] > 4
|
||||
|
||||
|
||||
def test_an_unrelated_rare_word_does_not_outrank_the_relevant_memory():
|
||||
"""Negative control: the paraphrase plus a place only one other memory
|
||||
holds. The decoy gains lexical score, and still ranks below F."""
|
||||
for name in ("ranking_crowded", "independent_default"):
|
||||
control = scenario(name)["ranking_variants"]["rare_word_with_paraphrase"]
|
||||
assert control["decoy_lexical_score"] > control["lexical_score"]
|
||||
assert control["rank"] == 1
|
||||
assert control["decoy_rank"] > control["rank"]
|
||||
|
||||
|
||||
def test_common_words_contribute_nothing():
|
||||
common = scenario("independent_default")["ranking_variants"]["common_words"]
|
||||
assert common["lexical_score"] == 0.0
|
||||
|
||||
|
||||
# ---------------------------------------------------------- capacity/eviction
|
||||
|
||||
def test_past_capacity_the_early_memory_is_evicted_and_the_stage_says_so():
|
||||
"""Diagnosis, not a requirement: what the current eviction rule does to F."""
|
||||
result = scenario("past_capacity")
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
assert result["diagnosis"]["verdict"] == "created_but_evicted"
|
||||
eviction = result["eviction"]
|
||||
assert eviction["f_evicted_at_turn"] is not None
|
||||
# It was retrieved while the bank was small, stopped being retrieved once
|
||||
# recent narration filled the top-k, and was then the least recently used.
|
||||
assert eviction["f_use_count_when_evicted"] > 0
|
||||
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
|
||||
assert eviction["f_memory_was_first_evicted"] is True
|
||||
@pytest.mark.parametrize("name", ["past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"])
|
||||
def test_past_capacity_the_early_memory_is_retained(name):
|
||||
"""v1.1 WP-B.2. On v1.0.0 all three are `created_but_evicted`: F was the
|
||||
least recently used row once recent narration stopped retrieving it, and
|
||||
went first (turns 21, 21 and 36). Coverage-first eviction keeps the only
|
||||
memory of the opening, so it stays active and is recalled at depth 106."""
|
||||
result = scenario(name)
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
diagnosis = result["diagnosis"]
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert result["eviction"]["f_evicted_at_turn"] is None
|
||||
# The bank really was past capacity, and stayed bounded.
|
||||
assert result["eviction"]["first_eviction_turn"] is not None
|
||||
assert all(t["active"] <= result["scenario"]["capacity"] + (1 if result["scenario"]["pin_first_memory"] else 0)
|
||||
for t in result["trace"])
|
||||
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
|
||||
|
||||
|
||||
def test_at_a_lower_top_k_the_early_memory_ages_out_after_it_stops_being_retrieved():
|
||||
"""Diagnosis with most of the bank unretrieved on any turn, nearer the
|
||||
shipped 5-in-80 ratio. F is not simply the oldest row: it is evicted some
|
||||
turns after recent narration stopped pulling it into the top-k, which is
|
||||
what ordering by last use does to a fact nothing recent mentions."""
|
||||
result = scenario("past_capacity_low_top_k")
|
||||
eviction = result["eviction"]
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
assert result["diagnosis"]["verdict"] == "created_but_evicted"
|
||||
assert eviction["f_use_count_when_evicted"] > 0
|
||||
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
|
||||
assert eviction["first_eviction_turn"] <= eviction["f_evicted_at_turn"]
|
||||
assert eviction["created_and_evicted_same_turn"] == []
|
||||
def test_past_capacity_the_bank_still_describes_the_whole_story():
|
||||
"""What the rule buys in general, not only for F: the active bank reaches
|
||||
from the opening to the newest block, and no stretch between them goes
|
||||
undescribed for more than twice the average spacing a bank of this capacity
|
||||
can afford (story span / capacity). On v1.0.0 these banks began at depths 36
|
||||
and 18: the opening was simply gone."""
|
||||
for name in ("past_capacity", "past_capacity_low_top_k"):
|
||||
result = scenario(name)
|
||||
cover = result["trace"][-1]["coverage"]
|
||||
assert cover["first_start"] == 0
|
||||
assert cover["last_end"] >= result["recall_depth"] - 2 * memorybank.MEMORY_INTERVAL
|
||||
assert cover["largest_gap"] <= 2 * (cover["last_end"] + 1) / result["scenario"]["capacity"]
|
||||
|
||||
|
||||
def test_no_memory_is_evicted_by_the_same_pass_that_created_it():
|
||||
"""The frozen-bank regression the current rule fixed, still holding."""
|
||||
for name in ("past_capacity", "past_capacity_pinned"):
|
||||
"""The frozen-bank regression the v1.0.0 rule fixed, still holding."""
|
||||
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
|
||||
assert scenario(name)["eviction"]["created_and_evicted_same_turn"] == []
|
||||
|
||||
|
||||
@@ -180,25 +246,28 @@ def test_a_pinned_memory_survives_capacity():
|
||||
assert eviction["pinned_memory_forgotten"] is False
|
||||
|
||||
|
||||
@pytest.mark.xfail(strict=True, reason=(
|
||||
"WP-B.1 diagnosis on this tree: past memory_bank_capacity the planting-era "
|
||||
"memory is evicted first, because it was never retrieved and eviction orders "
|
||||
"by last use, then creation. B.2 must flip this deliberately."))
|
||||
def test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity():
|
||||
assert scenario("past_capacity")["diagnosis"]["verdict"] == "injected"
|
||||
"""Was `xfail(strict=True)` in WP-B.1; B.2 fixed the eviction rule."""
|
||||
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
|
||||
assert scenario(name)["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ---------------------------------------------------------- creation window
|
||||
|
||||
def test_a_fact_early_in_a_long_block_never_reaches_the_summariser():
|
||||
def test_a_fact_early_in_a_long_block_now_reaches_the_summariser():
|
||||
"""v1.1 WP-B.2. On v1.0.0 this block (2,079 tokens) was cut to its last
|
||||
2,000, the fact at its start was never seen, and the stage was
|
||||
`not_created`. The excerpt is now the block's opening and end."""
|
||||
result = scenario("long_block_fact_early")
|
||||
created = result["diagnosis"]["created"]
|
||||
covering = created["covering_memories"]
|
||||
assert covering, "the long block must have been summarised"
|
||||
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert covering[0]["fact_in_block"] is True
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is False
|
||||
assert result["diagnosis"]["verdict"] == "not_created"
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is True
|
||||
assert created["yes"] is True
|
||||
assert created["source_start"] <= result["plant_depth"] <= created["source_end"]
|
||||
assert memorybank.EXCERPT_OMISSION_MARKER not in created["memory_text"]
|
||||
|
||||
|
||||
def test_the_same_fact_late_in_the_same_sized_block_does():
|
||||
@@ -209,12 +278,69 @@ def test_the_same_fact_late_in_the_same_sized_block_does():
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
|
||||
|
||||
@pytest.mark.xfail(strict=True, reason=(
|
||||
"WP-B.1 diagnosis on this tree: the summariser reads only the last "
|
||||
f"{memorybank.MEMORY_EXCERPT_TOKENS} tokens of a block, so a fact early in a "
|
||||
"long block is never seen. B.2 must flip this deliberately."))
|
||||
def test_acceptance_a_fact_early_in_a_long_block_is_remembered():
|
||||
"""Was `xfail(strict=True)` in WP-B.1; B.2 changed the excerpt."""
|
||||
assert scenario("long_block_fact_early")["diagnosis"]["created"]["yes"] is True
|
||||
assert scenario("long_block_fact_early")["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ------------------------------------------ WP-B.2 full deterministic acceptance
|
||||
|
||||
def test_acceptance_full_isolation_holds_on_every_turn():
|
||||
"""`independent_full`: long blocks, a crowded query, a bank past capacity.
|
||||
F must be carried by memory alone for the whole run, not only at recall."""
|
||||
result = scenario("independent_full")
|
||||
assert result["plant_depth"] <= 3 and result["recall_depth"] >= 100
|
||||
assert not any(t["f_in_state"] for t in result["trace"])
|
||||
assert not any(t["f_in_summary"] for t in result["trace"])
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
for check in ("state_document", "state_snapshots", "later_narration", "summary",
|
||||
"knowledge", "recent_history", "state_section"):
|
||||
assert result["isolation"]["checks"][check]["ok"], check
|
||||
|
||||
|
||||
def test_acceptance_full_every_stage_passes_past_capacity_with_long_blocks():
|
||||
"""On v1.0.0 this fixture fails at creation: every block is over 2,000
|
||||
tokens, and the fact at the start of the first one is never summarised."""
|
||||
result = scenario("independent_full")
|
||||
diagnosis = result["diagnosis"]
|
||||
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
|
||||
assert diagnosis["created"]["covering_memories"][0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["replica_matches_stored_selection"]
|
||||
assert diagnosis["injected"]["yes"]
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_acceptance_full_provenance_resolves_to_the_planting_turn():
|
||||
provenance = scenario("independent_full")["provenance"]
|
||||
assert provenance["recorded"] is not None
|
||||
assert provenance["range_covers_plant"] and provenance["matches_row"]
|
||||
assert provenance["source_block_holds_planting"] is True
|
||||
assert provenance["recorded"]["authority"] == memorybank.ACCEPTED_STORY
|
||||
|
||||
|
||||
def test_acceptance_full_is_the_long_run_independent_memory_verdict():
|
||||
"""The same measurements, judged by the long-run tool's own verdict."""
|
||||
from tools import m11_long_run as lr
|
||||
|
||||
result = scenario("independent_full")
|
||||
checks = result["isolation"]["checks"]
|
||||
diagnosis = result["diagnosis"]
|
||||
verdict = lr._independent_memory_verdict({
|
||||
"independent_planted_depth": result["plant_depth"],
|
||||
"planted_turn_outside_history": checks["recent_history"]["ok"],
|
||||
"absent_from_state": checks["state_document"]["ok"] and checks["state_snapshots"]["ok"]
|
||||
and not any(t["f_in_state"] for t in result["trace"]),
|
||||
"absent_from_summary": checks["summary"]["ok"]
|
||||
and not any(t["f_in_summary"] for t in result["trace"]),
|
||||
"absent_from_knowledge": checks["knowledge"]["ok"],
|
||||
"absent_from_later_narration": checks["later_narration"]["ok"],
|
||||
"memory_covering_planting_carries_fact": diagnosis["created"]["yes"],
|
||||
"memory_forgotten": not diagnosis["retained"]["yes"],
|
||||
"memory_injected": diagnosis["injected"]["yes"],
|
||||
})
|
||||
assert verdict == "recovered_through_memory_independent"
|
||||
|
||||
|
||||
# ------------------------------------------------------- lineage control (G)
|
||||
|
||||
@@ -0,0 +1,277 @@
|
||||
"""v1.1 WP-B.2 (B2.2): which memory a full bank lets go of.
|
||||
|
||||
v1.0.0 evicted the least recently used memory. B.1 showed that this discards
|
||||
the only memory of an early stretch first, because retrieval follows the present
|
||||
scene and nothing recent resembles it. `memorybank.eviction_order` now thins the
|
||||
bank where it is densest and keeps the opening and the newest stretch, with
|
||||
recency as the tie-break and least-recently-used as the fallback.
|
||||
|
||||
The scenario-level tests (the planted fact kept past capacity) are in
|
||||
`test_v11_b1_memory_diagnostic.py`; the v1.0.0 eviction tests in
|
||||
`test_memory_retrieval.py` still pass unchanged, because their memories carry
|
||||
no source range and take the fallback.
|
||||
|
||||
python -m pytest tests/test_v11_b2_memory_eviction.py -v
|
||||
"""
|
||||
|
||||
import random
|
||||
from collections import namedtuple
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import select
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.context import lineage
|
||||
from app.database import Base, SessionLocal, engine
|
||||
|
||||
T0 = datetime(2026, 1, 1, 12, 0, 0)
|
||||
Row = namedtuple("Row", "id pinned source_start source_end last_used_at created_at use_count")
|
||||
|
||||
|
||||
def row(id, start, end=None, *, pinned=False, used=None, created=None, uses=0):
|
||||
"""A memory as eviction sees it. Times are minutes after T0."""
|
||||
return Row(id, pinned, start, (start + 5) if end is None and start is not None else end,
|
||||
None if used is None else T0 + timedelta(minutes=used),
|
||||
T0 + timedelta(minutes=id if created is None else created), uses)
|
||||
|
||||
|
||||
def blocks(n, *, first_id=1):
|
||||
return [row(first_id + i, 6 * i) for i in range(n)]
|
||||
|
||||
|
||||
def largest_gap_from_opening(rows):
|
||||
"""The longest uncovered run of depths from depth 0 to the last memory."""
|
||||
ordered = sorted((r.source_start, r.source_end) for r in rows)
|
||||
gaps = [ordered[0][0]]
|
||||
reach = ordered[0][1]
|
||||
for start, end in ordered[1:]:
|
||||
gaps.append(max(0, start - reach - 1))
|
||||
reach = max(reach, end)
|
||||
return max(gaps)
|
||||
|
||||
|
||||
# ------------------------------------------------------------ the pure order
|
||||
|
||||
|
||||
def test_the_opening_and_the_newest_memory_are_kept():
|
||||
bank = blocks(7)
|
||||
doomed = memorybank.eviction_order(bank, 5)
|
||||
assert bank[0].id not in doomed and bank[-1].id not in doomed
|
||||
assert len(doomed) == 5
|
||||
|
||||
|
||||
def test_the_densest_stretch_is_thinned_first():
|
||||
# Memories every 6 depths to 30, then a sparse stretch. Removing one of the
|
||||
# dense ones leaves a 6-depth hole; removing a sparse one leaves far more.
|
||||
bank = [row(1, 0), row(2, 6), row(3, 12), row(4, 18), row(5, 60), row(6, 120), row(7, 180)]
|
||||
assert memorybank.eviction_order(bank, 1)[0] in {2, 3, 4}
|
||||
assert set(memorybank.eviction_order(bank, 2)) <= {2, 3, 4}
|
||||
|
||||
|
||||
def test_a_stretch_two_memories_describe_loses_one_of_them_first():
|
||||
"""A shared start (a re-played stretch, or a sibling line) leaves no hole.
|
||||
Of the two, the less recently used goes, even though a unique memory
|
||||
elsewhere is older and less used than both."""
|
||||
bank = [row(1, 0), row(2, 6, used=5), row(3, 12, used=50), row(4, 12, used=40), row(5, 18)]
|
||||
assert memorybank.eviction_order(bank, 1) == [4]
|
||||
|
||||
|
||||
def test_equal_holes_fall_to_the_least_recently_used():
|
||||
bank = [row(1, 0), row(2, 6, used=30), row(3, 12, used=10), row(4, 18, used=20), row(5, 24)]
|
||||
assert memorybank.eviction_order(bank, 1) == [3]
|
||||
|
||||
|
||||
def test_a_newborn_can_be_the_legitimate_first_to_go():
|
||||
"""The frozen bank is about a newborn losing to a count it cannot have yet.
|
||||
A newborn that only repeats a stretch another memory describes, one used
|
||||
after it was written, is legitimately the first to go."""
|
||||
bank = [row(1, 0), row(2, 6, used=100), row(3, 12), row(4, 6, created=90)]
|
||||
assert memorybank.eviction_order(bank, 1) == [4]
|
||||
|
||||
|
||||
def test_the_newest_memory_is_not_evicted_by_the_bank_it_joins():
|
||||
"""The frozen-bank regression under the new rule: every older memory has
|
||||
been used, the newborn never has, and it still stays."""
|
||||
bank = [row(i, 6 * (i - 1), used=200 + i, uses=3) for i in range(1, 6)]
|
||||
newborn = row(6, 30, created=300)
|
||||
assert newborn.id not in memorybank.eviction_order(bank + [newborn], 1)
|
||||
|
||||
|
||||
def test_pins_are_never_taken_but_still_count_as_coverage():
|
||||
bank = [row(1, 0), row(2, 6, pinned=True), row(3, 12), row(4, 18, pinned=True), row(5, 24)]
|
||||
doomed = memorybank.eviction_order(bank, 10)
|
||||
assert not {2, 4} & set(doomed)
|
||||
# With the pins covering 6 and 18, memory 3's hole is only its own block.
|
||||
assert doomed[0] == 3
|
||||
|
||||
|
||||
def test_memories_without_a_range_take_the_least_recently_used_fallback():
|
||||
hand_written = [row(1, None, None, used=30), row(2, None, None, used=10),
|
||||
row(3, None, None, used=20)]
|
||||
assert memorybank.eviction_order(hand_written, 3) == [2, 3, 1]
|
||||
|
||||
|
||||
def test_the_fallback_is_used_only_once_no_interior_memory_remains():
|
||||
bank = [row(1, 0, used=1), row(2, 6, used=90), row(3, 12, used=2),
|
||||
row(10, None, None, used=0)]
|
||||
order = memorybank.eviction_order(bank, 4)
|
||||
assert order[0] == 2 # the interior memory, although recently used
|
||||
assert order[1:] == [10, 1, 3] # then least recently used
|
||||
|
||||
|
||||
def test_the_order_does_not_depend_on_row_order():
|
||||
bank = [row(i, 6 * (i - 1), used=(i * 37) % 11, uses=i % 3) for i in range(1, 30)]
|
||||
bank += [row(40, 12), row(41, 12)] # a shared start with identical timestamps
|
||||
expected = memorybank.eviction_order(bank, 20)
|
||||
for seed in range(5):
|
||||
shuffled = bank[:]
|
||||
random.Random(seed).shuffle(shuffled)
|
||||
assert memorybank.eviction_order(shuffled, 20) == expected
|
||||
|
||||
|
||||
def test_a_tie_on_every_signal_is_broken_by_id():
|
||||
bank = [row(1, 0), row(9, 6, created=0), row(4, 12, created=0), row(20, 18)]
|
||||
assert memorybank.eviction_order(bank, 1) == [4]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("seed", range(8))
|
||||
def test_capacity_holds_and_pins_survive_for_any_bank(seed):
|
||||
rng = random.Random(seed)
|
||||
bank = []
|
||||
for i in range(1, rng.randint(2, 60)):
|
||||
start = None if rng.random() < 0.15 else rng.randrange(0, 400)
|
||||
bank.append(row(i, start, None if start is None else start + rng.choice([3, 5, 8]),
|
||||
pinned=rng.random() < 0.1, used=rng.choice([None, rng.randrange(500)]),
|
||||
uses=rng.randrange(4)))
|
||||
capacity = rng.randint(1, 30)
|
||||
overflow = len(bank) - capacity
|
||||
doomed = memorybank.eviction_order(bank, max(0, overflow))
|
||||
pinned = {r.id for r in bank if r.pinned}
|
||||
assert not pinned & set(doomed)
|
||||
assert len(set(doomed)) == len(doomed)
|
||||
remaining = len(bank) - len(doomed)
|
||||
assert remaining == max(capacity, len(pinned)) if overflow > 0 else remaining == len(bank)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("n, capacity, irregular", [(60, 10, False), (500, 80, False), (500, 80, True)])
|
||||
def test_a_long_bank_keeps_describing_the_whole_story(n, capacity, irregular):
|
||||
"""The general property, with nothing ever retrieved: memories arrive one
|
||||
block at a time and the bank is kept at capacity. The opening stays, and no
|
||||
stretch goes undescribed for more than twice the average spacing. Least
|
||||
recently used order, on the same arrivals, keeps only the newest stretch."""
|
||||
rng = random.Random(n)
|
||||
kept, lru = [], []
|
||||
depth = 0
|
||||
for i in range(1, n + 1):
|
||||
size = rng.choice([4, 6, 6, 9]) if irregular else 6
|
||||
memory = row(i, depth, depth + size - 1)
|
||||
depth += size
|
||||
kept.append(memory)
|
||||
lru.append(memory)
|
||||
if len(kept) > capacity:
|
||||
doomed = set(memorybank.eviction_order(kept, len(kept) - capacity))
|
||||
kept = [m for m in kept if m.id not in doomed]
|
||||
lru = sorted(lru, key=lambda m: (m.created_at, m.id))[len(lru) - capacity:]
|
||||
assert len(kept) == capacity
|
||||
assert min(m.source_start for m in kept) == 0
|
||||
assert max(m.id for m in kept) == n
|
||||
assert largest_gap_from_opening(kept) <= 2 * depth / capacity
|
||||
assert largest_gap_from_opening(lru) > depth / 2 # v1.0.0 order: the opening is gone
|
||||
|
||||
|
||||
# --------------------------------------------------------- on real rows
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
memorybank._vector_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def adventure(db):
|
||||
user = models.User(is_guest=False, email="b2-evict@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
settings = models.Settings(user_id=user.id, model="m", embedding_model="e",
|
||||
memory_bank_capacity=3)
|
||||
adv = models.Adventure(user_id=user.id, title="Evict", script_state={}, memory_bank_enabled=True)
|
||||
db.add_all([settings, adv])
|
||||
db.commit()
|
||||
adv.settings_row = settings
|
||||
return adv
|
||||
|
||||
|
||||
def test_the_pass_changes_nothing_but_forgotten(db, adventure):
|
||||
for i in range(6):
|
||||
memory = models.Memory(adventure_id=adventure.id, text=f"block {i}",
|
||||
source_start=6 * i, source_end=6 * i + 5, branch_id=None, depth=6 * i + 5)
|
||||
db.add(memory)
|
||||
db.commit()
|
||||
columns = (models.Memory.id, models.Memory.text, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.branch_id, models.Memory.depth,
|
||||
models.Memory.pinned, models.Memory.use_count, models.Memory.last_used_at)
|
||||
before = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
|
||||
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
|
||||
after = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
|
||||
assert before == after
|
||||
active = db.execute(select(models.Memory.id).where(models.Memory.forgotten.is_(False))).scalars().all()
|
||||
assert len(active) == 3
|
||||
assert min(active) == min(before) and max(active) == max(before) # the boundaries
|
||||
|
||||
|
||||
def test_eviction_does_not_make_an_abandoned_lines_memory_eligible(db, adventure):
|
||||
"""Eviction and lineage are separate: the pass decides only `forgotten`, so
|
||||
a memory on a line the story left is exactly as ineligible afterwards."""
|
||||
trunk = []
|
||||
for i in range(4):
|
||||
action = models.Action(adventure_id=adventure.id, type="ai", text=f"trunk {i}")
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
trunk.append(action)
|
||||
abandoned_node = models.Action(adventure_id=adventure.id, type="ai", text="the abandoned line")
|
||||
tree.place_action(db, adventure, abandoned_node)
|
||||
db.add(abandoned_node)
|
||||
db.flush()
|
||||
abandoned = models.Memory(adventure_id=adventure.id, text="on the abandoned line",
|
||||
source_start=4, source_end=4)
|
||||
tree.attach_memory(abandoned, abandoned_node)
|
||||
db.add(abandoned)
|
||||
db.commit()
|
||||
# Move the head back and diverge, so the abandoned node is off the path.
|
||||
adventure.head_depth = trunk[-1].depth
|
||||
db.commit()
|
||||
from app import head
|
||||
head.fork_if_behind_head(db, adventure)
|
||||
divergent = models.Action(adventure_id=adventure.id, type="ai", text="the new line")
|
||||
tree.place_action(db, adventure, divergent)
|
||||
db.add(divergent)
|
||||
db.flush()
|
||||
for i, node in enumerate(trunk + [divergent]):
|
||||
memory = models.Memory(adventure_id=adventure.id, text=f"active {i}",
|
||||
source_start=node.depth, source_end=node.depth)
|
||||
tree.attach_memory(memory, node)
|
||||
db.add(memory)
|
||||
db.commit()
|
||||
|
||||
def eligible():
|
||||
return set(db.execute(select(models.Memory.id).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory),
|
||||
models.Memory.forgotten.is_(False))).scalars().all())
|
||||
|
||||
assert abandoned.id not in eligible()
|
||||
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
|
||||
db.expire_all()
|
||||
assert abandoned.id not in eligible()
|
||||
assert len(db.execute(select(models.Memory.id).where(
|
||||
models.Memory.forgotten.is_(False))).scalars().all()) == 3
|
||||
@@ -0,0 +1,189 @@
|
||||
"""v1.1 WP-B.2 (B2.3): what the memory summariser is shown of a long block.
|
||||
|
||||
v1.0.0 sent the last 2,000 tokens of a block, so a fact early in a longer block
|
||||
never reached the summariser (B.1 §E). A block that fits is still sent whole. A
|
||||
longer one is now sent as its opening and its end, with a marker between them,
|
||||
inside the same 2,000-token budget.
|
||||
|
||||
The scenario-level test (the planted fact early in a long block, remembered) is
|
||||
in `test_v11_b1_memory_diagnostic.py`.
|
||||
|
||||
python -m pytest tests/test_v11_b2_memory_excerpt.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import random
|
||||
|
||||
import pytest
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine
|
||||
|
||||
BUDGET = memorybank.MEMORY_EXCERPT_TOKENS
|
||||
MARKER = memorybank.EXCERPT_OMISSION_MARKER
|
||||
FILLER = "The travellers walked the long grey road north past the salt market and the reed beds. "
|
||||
|
||||
|
||||
def words_to_tokens(tokens: int) -> str:
|
||||
"""Filler at least `tokens` long."""
|
||||
text = FILLER
|
||||
while builder.count_tokens(text) < tokens:
|
||||
text += FILLER
|
||||
return text
|
||||
|
||||
|
||||
# --------------------------------------------------------------- the excerpt
|
||||
|
||||
|
||||
def test_a_block_that_fits_is_sent_whole_and_unchanged():
|
||||
raw = words_to_tokens(BUDGET - 200)
|
||||
assert builder.count_tokens(raw) <= BUDGET
|
||||
assert memorybank.memory_excerpt(raw) == raw
|
||||
|
||||
|
||||
def test_a_block_of_exactly_the_budget_is_unchanged():
|
||||
raw = words_to_tokens(BUDGET)
|
||||
tokens = memorybank._excerpt_encoding().encode(raw)[:BUDGET]
|
||||
exact = memorybank._excerpt_encoding().decode(tokens)
|
||||
if builder.count_tokens(exact) == BUDGET:
|
||||
assert memorybank.memory_excerpt(exact) == exact
|
||||
|
||||
|
||||
def test_a_long_block_keeps_its_opening_and_its_end_in_order():
|
||||
opening = "Mara slipped the amber sundial inside the cracked teapot. "
|
||||
ending = "Aldric finally reached the north gate at dawn."
|
||||
raw = opening + words_to_tokens(3 * BUDGET) + ending
|
||||
excerpt = memorybank.memory_excerpt(raw)
|
||||
assert excerpt.startswith(opening)
|
||||
assert excerpt.endswith(ending)
|
||||
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
|
||||
assert tail, "the marker must sit between the two parts"
|
||||
assert excerpt.index(opening) < excerpt.index(MARKER) < excerpt.index(ending)
|
||||
|
||||
|
||||
def test_the_split_is_even_and_documented():
|
||||
raw = words_to_tokens(4 * BUDGET)
|
||||
head_budget, tail_budget = memorybank.excerpt_split(BUDGET)
|
||||
marker_tokens = builder.count_tokens(f"\n\n{MARKER}\n\n")
|
||||
assert head_budget + tail_budget + marker_tokens == BUDGET
|
||||
assert abs(head_budget - tail_budget) <= 1
|
||||
excerpt = memorybank.memory_excerpt(raw)
|
||||
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
|
||||
# Each part is cut as a run of `head_budget` / `tail_budget` tokens. Measured
|
||||
# on its own, a cut run can come to one token more, because the text either
|
||||
# side of the cut tokenises differently once it is separated; the hard limit
|
||||
# is the whole excerpt, tested below.
|
||||
assert builder.count_tokens(head) <= head_budget + 1
|
||||
assert builder.count_tokens(tail) <= tail_budget + 1
|
||||
assert builder.count_tokens(excerpt) <= BUDGET
|
||||
|
||||
|
||||
@pytest.mark.parametrize("extra", [1, 7, 500, BUDGET, 9 * BUDGET])
|
||||
def test_the_excerpt_never_exceeds_the_budget(extra):
|
||||
raw = words_to_tokens(BUDGET + extra)
|
||||
excerpt = memorybank.memory_excerpt(raw)
|
||||
assert builder.count_tokens(excerpt) <= BUDGET
|
||||
assert memorybank.memory_excerpt(raw) == excerpt # deterministic
|
||||
|
||||
|
||||
@pytest.mark.parametrize("seed", range(4))
|
||||
def test_the_budget_holds_for_awkward_text(seed):
|
||||
"""Token boundaries can merge differently once the parts are rejoined, and
|
||||
text that is not plain English tokenises unevenly. The budget still holds."""
|
||||
rng = random.Random(seed)
|
||||
alphabet = "abcdefghij ÄÖÜ ßé漢字かな 🙂🐉 \n\t.,;:—'\""
|
||||
raw = "".join(rng.choice(alphabet) for _ in range(12000))
|
||||
assert builder.count_tokens(memorybank.memory_excerpt(raw)) <= BUDGET
|
||||
|
||||
|
||||
def test_a_fact_in_the_middle_of_a_very_long_block_is_still_omitted():
|
||||
"""The documented limit of a bounded excerpt: head and tail, not everything."""
|
||||
half = words_to_tokens(3 * BUDGET)
|
||||
raw = half + "Mara slipped the amber sundial inside the cracked teapot. " + half
|
||||
assert "sundial" not in memorybank.memory_excerpt(raw)
|
||||
|
||||
|
||||
# ------------------------------------------------------------ memory creation
|
||||
|
||||
|
||||
class EchoSummariser:
|
||||
"""Returns the whole excerpt it was given as the memory: the worst case for
|
||||
a marker leaking into stored text."""
|
||||
|
||||
def __init__(self):
|
||||
self.users: list[str] = []
|
||||
|
||||
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
|
||||
self.users.append(user)
|
||||
return user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def campaign(db, texts):
|
||||
user = models.User(is_guest=False, email="b2-excerpt@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
|
||||
adventure = models.Adventure(user_id=user.id, title="Excerpt", script_state={},
|
||||
auto_summarize=True, memory_bank_enabled=True)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
nodes = []
|
||||
for i, text in enumerate(texts):
|
||||
action = models.Action(adventure_id=adventure.id, type="ai" if i % 2 else "do", text=text)
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
nodes.append(action)
|
||||
db.commit()
|
||||
return adventure, nodes
|
||||
|
||||
|
||||
def write_memory(db, adventure, monkeypatch, summariser):
|
||||
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
|
||||
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: summariser)
|
||||
settings = db.query(models.Settings).first()
|
||||
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
|
||||
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
|
||||
|
||||
|
||||
def test_the_marker_is_never_stored_as_part_of_a_memory(db, monkeypatch):
|
||||
long = words_to_tokens(800)
|
||||
adventure, _ = campaign(db, [long] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK))
|
||||
summariser = EchoSummariser()
|
||||
[memory] = write_memory(db, adventure, monkeypatch, summariser)
|
||||
assert MARKER in summariser.users[0] # the summariser was told
|
||||
assert MARKER not in memory.text # and the memory does not repeat it
|
||||
assert "[…" not in memory.text and "omitted" not in memory.text
|
||||
|
||||
|
||||
def test_a_short_block_is_prompted_exactly_as_before(db, monkeypatch):
|
||||
texts = [f"Short action {i}." for i in range(memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK)]
|
||||
adventure, _ = campaign(db, texts)
|
||||
summariser = EchoSummariser()
|
||||
write_memory(db, adventure, monkeypatch, summariser)
|
||||
block = "\n\n".join(texts[:memorybank.MEMORY_INTERVAL])
|
||||
assert summariser.users[0] == f"Story excerpt:\n\n{block}\n\nMemory:"
|
||||
|
||||
|
||||
def test_a_long_blocks_memory_keeps_its_source_provenance(db, monkeypatch):
|
||||
early = "Mara slipped the amber sundial inside the cracked teapot. " + words_to_tokens(900)
|
||||
texts = [early] + [words_to_tokens(900)] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK - 1)
|
||||
adventure, nodes = campaign(db, texts)
|
||||
[memory] = write_memory(db, adventure, monkeypatch, EchoSummariser())
|
||||
block = nodes[:memorybank.MEMORY_INTERVAL]
|
||||
assert "sundial" in memory.text # the early fact reached the summariser
|
||||
assert (memory.source_start, memory.source_end) == (block[0].depth, block[-1].depth)
|
||||
assert (memory.branch_id, memory.depth) == (block[-1].branch_id, block[-1].depth)
|
||||
@@ -0,0 +1,326 @@
|
||||
"""v1.1 WP-B.2 (B2.1): what memory retrieval searches for, and how it scores.
|
||||
|
||||
B.1 found the planting-era memory created and retained but ranked out of
|
||||
`memory_top_k`, because the query was three turns of narration with the player's
|
||||
question at the end. The query is now the player's input plus a short scene
|
||||
context, and the score adds one transparent lexical term over the input.
|
||||
|
||||
These tests pin the pieces. The end-to-end fixture tests (crowded bank,
|
||||
context-dependent question, negative controls) are in
|
||||
`test_v11_b1_memory_diagnostic.py`, beside the diagnostic they use.
|
||||
|
||||
python -m pytest tests/test_v11_b2_memory_ranking.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import math
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import event
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine
|
||||
|
||||
# ------------------------------------------------------------------ lexical
|
||||
|
||||
|
||||
def test_terms_are_folded_but_not_stemmed():
|
||||
terms = memorybank.lexical_terms("The tavern's teapots, the glass and the SUNDIAL")
|
||||
assert {"tavern", "teapot", "glass", "sundial"} <= terms
|
||||
assert "the" not in terms # the knowledge path's stop list
|
||||
assert "glas" not in terms # a word ending in "ss" is not a plural
|
||||
|
||||
|
||||
def test_a_word_every_candidate_holds_weighs_nothing():
|
||||
scores = memorybank.lexical_scores(
|
||||
frozenset({"travellers"}), {1: frozenset({"travellers", "road"}), 2: frozenset({"travellers"})})
|
||||
assert scores == {1: 0.0, 2: 0.0}
|
||||
|
||||
|
||||
def test_a_rarer_word_weighs_more_than_a_common_one():
|
||||
scores = memorybank.lexical_scores(
|
||||
frozenset({"sundial", "road"}),
|
||||
{1: frozenset({"sundial"}), 2: frozenset({"road"}), 3: frozenset({"road"}),
|
||||
4: frozenset({"gate"})})
|
||||
assert scores[1] > scores[2] == scores[3] > scores[4] == 0.0
|
||||
|
||||
|
||||
def test_a_single_rare_word_of_a_longer_question_is_only_its_share():
|
||||
"""The share is over the whole question, so one incidental word match
|
||||
cannot score like a memory that answers it."""
|
||||
question = frozenset({"where", "amber", "sundial", "fish"})
|
||||
scores = memorybank.lexical_scores(
|
||||
question, {1: frozenset({"fish"}), 2: frozenset({"amber", "sundial"}), 3: frozenset({"road"})})
|
||||
assert 0.0 < scores[1] < scores[2] <= 1.0
|
||||
assert scores[1] < 0.5
|
||||
|
||||
|
||||
def test_scores_are_bounded_and_empty_inputs_score_zero():
|
||||
candidates = {1: frozenset({"a1", "b2"}), 2: frozenset({"a1"})}
|
||||
assert all(0.0 <= v <= 1.0 for v in memorybank.lexical_scores(frozenset({"a1", "b2"}), candidates).values())
|
||||
assert memorybank.lexical_scores(frozenset(), candidates) == {1: 0.0, 2: 0.0}
|
||||
assert memorybank.lexical_scores(frozenset({"a1"}), {}) == {}
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ scoring
|
||||
|
||||
|
||||
def _unit(angle):
|
||||
return [math.cos(angle), math.sin(angle)]
|
||||
|
||||
|
||||
def test_ties_are_broken_by_id_not_by_row_order():
|
||||
held = {7: [1.0, 0.0], 3: [1.0, 0.0], 5: [1.0, 0.0]}
|
||||
rows = memorybank.score_candidates([7, 3, 5], held, {}, [1.0, 0.0], None, [])
|
||||
assert [row[1] for row in rows] == [3, 5, 7]
|
||||
|
||||
|
||||
def test_the_semantic_score_mixes_input_and_context_by_the_fixed_weight():
|
||||
held = {1: [1.0, 0.0]}
|
||||
[(final, _, semantic, lexical)] = memorybank.score_candidates(
|
||||
[1], held, {}, [1.0, 0.0], [0.0, 1.0], [])
|
||||
assert semantic == pytest.approx(memorybank.INPUT_WEIGHT)
|
||||
assert final == semantic and lexical == 0.0
|
||||
# Either part alone is used as it is.
|
||||
[(_, _, only_context, _)] = memorybank.score_candidates([1], held, {}, None, [0.0, 1.0], [])
|
||||
assert only_context == pytest.approx(0.0)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("margin, relevant_first", [(0.01, True), (-0.01, False)])
|
||||
def test_a_lexical_match_moves_a_memory_at_most_the_lexical_weight(margin, relevant_first):
|
||||
"""The bound that keeps rarity from overruling meaning: a memory more than
|
||||
`LEXICAL_WEIGHT` behind semantically cannot pass one ahead of it, however
|
||||
rare the word it shares."""
|
||||
decoy_cos = 1.0 - memorybank.LEXICAL_WEIGHT - margin
|
||||
held = {1: [1.0, 0.0], 2: _unit(math.acos(decoy_cos))}
|
||||
terms = {1: frozenset(), 2: frozenset({"zeppelin"})}
|
||||
rows = memorybank.score_candidates([1, 2], held, terms, [1.0, 0.0], None, ["zeppelin"])
|
||||
order = [row[1] for row in rows]
|
||||
assert (order[0] == 1) is relevant_first
|
||||
|
||||
|
||||
# ---------------------------------------------------- the query, on real rows
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
memorybank._terms_cache.clear()
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
memorybank._vector_cache.clear()
|
||||
memorybank._terms_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def restore_embedding_provider():
|
||||
real = memorybank.embedding_provider
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
memorybank.embedding_provider = real
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def settings(db):
|
||||
user = models.User(is_guest=False, email="b2-rank@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
row = models.Settings(user_id=user.id, model="m", embedding_model="stub-embed",
|
||||
memory_top_k=2, memory_bank_capacity=80)
|
||||
db.add(row)
|
||||
db.commit()
|
||||
return row
|
||||
|
||||
|
||||
SCENE_STATE = {
|
||||
"entities": {"mara": {"type": "character", "name": "Mara"},
|
||||
"tavern": {"type": "location", "name": "The Crooked Lantern"}},
|
||||
"scene": {"summary": "Closing time", "location": "tavern", "present": ["mara"]},
|
||||
}
|
||||
|
||||
|
||||
def make_adventure(db, settings, texts, state=None):
|
||||
"""`texts` is `[(type, text)]`, oldest first, each placed on the tree."""
|
||||
adventure = models.Adventure(user_id=settings.user_id, title="Rank", script_state={},
|
||||
memory_bank_enabled=True, narrative_state=state or {})
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
placed = []
|
||||
for kind, text in texts:
|
||||
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
placed.append(action)
|
||||
db.commit()
|
||||
return adventure, placed
|
||||
|
||||
|
||||
def test_the_query_is_the_players_input_and_the_scene(db, settings):
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("start", "Rain over the harbour."),
|
||||
("ai", "Mara wipes down the counter and glances up at the shelf."),
|
||||
("do", "> You ask Mara about the brass dial."),
|
||||
], state=SCENE_STATE)
|
||||
query = memorybank.retrieval_query(adventure)
|
||||
assert query["input"] == "> You ask Mara about the brass dial."
|
||||
assert "The Crooked Lantern" in query["context"] and "Mara" in query["context"]
|
||||
assert "glances up at the shelf" in query["context"]
|
||||
assert "brass" in query["input_terms"] and "dial" in query["input_terms"]
|
||||
|
||||
|
||||
def test_a_continue_turn_has_no_input_and_searches_by_the_scene(db, settings):
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("do", "> You sit down."),
|
||||
("ai", "The fire burns low in the grate."),
|
||||
])
|
||||
query = memorybank.retrieval_query(adventure)
|
||||
assert query["input"] == "" and query["input_terms"] == []
|
||||
assert "fire burns low" in query["context"]
|
||||
|
||||
|
||||
def test_a_retry_searches_with_the_input_it_is_retrying(db, settings):
|
||||
adventure, placed = make_adventure(db, settings, [
|
||||
("ai", "The market is quiet."),
|
||||
("do", "> You ask about the sundial."),
|
||||
("ai", "A discarded attempt about lanterns."),
|
||||
])
|
||||
query = memorybank.retrieval_query(adventure, exclude_action_id=placed[-1].id)
|
||||
assert query["input"] == "> You ask about the sundial."
|
||||
assert "lanterns" not in query["context"]
|
||||
assert "market is quiet" in query["context"]
|
||||
|
||||
|
||||
def test_the_query_is_bounded_however_long_the_story(db, settings):
|
||||
long = "The travellers walked the long grey road north past the salt market. " * 400
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("ai", long), ("story", long)], state=SCENE_STATE)
|
||||
query = memorybank.retrieval_query(adventure)
|
||||
assert builder.count_tokens(query["input"]) <= memorybank.QUERY_INPUT_TOKENS
|
||||
assert builder.count_tokens(query["context"]) <= (
|
||||
memorybank.QUERY_SCENE_TOKENS + memorybank.QUERY_NARRATION_TOKENS + 2)
|
||||
|
||||
|
||||
# ------------------------------------------------ retrieval, end to end
|
||||
|
||||
|
||||
class SameVector:
|
||||
"""Every text embeds the same, so only the lexical term separates memories."""
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.0, 0.0] for _ in texts]
|
||||
|
||||
|
||||
def add_memory(db, adventure, text, **kwargs):
|
||||
memory = models.Memory(adventure_id=adventure.id, text=text, **kwargs)
|
||||
db.add(memory)
|
||||
db.flush()
|
||||
memorybank.set_vector(memory, [1.0, 0.0, 0.0])
|
||||
db.commit()
|
||||
return memory
|
||||
|
||||
|
||||
def retrieve(adventure, settings, **kwargs):
|
||||
memorybank.embedding_provider = lambda s: SameVector()
|
||||
return asyncio.run(memorybank.retrieve_memories(adventure, settings, **kwargs))
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def played(db, settings):
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("ai", "The tavern is warm."),
|
||||
("do", "> You ask Mara where the amber sundial went."),
|
||||
], state=SCENE_STATE)
|
||||
bank = {
|
||||
"road": add_memory(db, adventure, "Aldric walked the north road."),
|
||||
"sundial": add_memory(db, adventure, "Mara hid the amber sundial in the teapot."),
|
||||
"gate": add_memory(db, adventure, "The gate guard asked for a toll."),
|
||||
}
|
||||
return adventure, bank
|
||||
|
||||
|
||||
def test_every_used_memory_reports_the_parts_of_its_score(db, settings, played):
|
||||
adventure, bank = played
|
||||
result = retrieve(adventure, settings)
|
||||
first = result["used"][0]
|
||||
assert first["id"] == bank["sundial"].id
|
||||
assert first["similarity"] == first["semantic_score"]
|
||||
assert first["lexical_score"] > 0
|
||||
assert first["final_score"] == pytest.approx(
|
||||
first["semantic_score"] + memorybank.LEXICAL_WEIGHT * first["lexical_score"], abs=2e-4)
|
||||
assert result["query"]["input"] == "> You ask Mara where the amber sundial went."
|
||||
assert result["query"]["lexical_weight"] == memorybank.LEXICAL_WEIGHT
|
||||
assert result["query"]["input_weight"] == memorybank.INPUT_WEIGHT
|
||||
|
||||
|
||||
def test_a_pin_is_still_always_used_and_counts_toward_top_k(db, settings, played):
|
||||
adventure, bank = played
|
||||
settings.memory_top_k = 1
|
||||
bank["gate"].pinned = True
|
||||
db.commit()
|
||||
used = retrieve(adventure, settings)["used"]
|
||||
assert [m["id"] for m in used] == [bank["gate"].id]
|
||||
assert used[0]["pinned"] is True
|
||||
|
||||
|
||||
def memory_text_reads(statements):
|
||||
return [s for s in statements
|
||||
if s.lstrip().upper().startswith("SELECT") and "FROM memories" in s
|
||||
and "memories.text" in s]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def sql_log():
|
||||
statements: list[str] = []
|
||||
|
||||
def record(conn, cursor, statement, parameters, context, executemany):
|
||||
statements.append(statement)
|
||||
|
||||
event.listen(engine, "before_cursor_execute", record)
|
||||
try:
|
||||
yield statements
|
||||
finally:
|
||||
event.remove(engine, "before_cursor_execute", record)
|
||||
|
||||
|
||||
def test_memory_text_is_read_once_and_then_held(db, settings, played, sql_log):
|
||||
adventure, _ = played
|
||||
retrieve(adventure, settings)
|
||||
sql_log.clear()
|
||||
result = retrieve(adventure, settings)
|
||||
reads = memory_text_reads(sql_log)
|
||||
# Only the detail read of the memories chosen remains. (Every memory here
|
||||
# embeds identically, so redundancy suppression keeps just one of them.)
|
||||
assert len(reads) == 1 and reads[0].count("?") == len(result["used"])
|
||||
|
||||
|
||||
def test_a_continue_turn_reads_no_memory_text_to_rank(db, settings, sql_log):
|
||||
adventure, _ = make_adventure(db, settings, [("do", "> You wait."), ("ai", "Night falls.")])
|
||||
for text in ("one", "two", "three"):
|
||||
add_memory(db, adventure, f"memory {text}")
|
||||
sql_log.clear()
|
||||
result = retrieve(adventure, settings)
|
||||
assert all(m["lexical_score"] == 0.0 for m in result["used"])
|
||||
assert len(memory_text_reads(sql_log)) == 1 # the top-k detail read only
|
||||
|
||||
|
||||
def test_an_edited_memory_is_matched_on_its_new_text(db, settings, played):
|
||||
adventure, bank = played
|
||||
assert retrieve(adventure, settings)["used"][0]["id"] == bank["sundial"].id
|
||||
# An edit clears the vector (the route calls set_vector(None)); re-embedding
|
||||
# sets it again. Both go through set_vector, which drops the held terms.
|
||||
bank["road"].text = "The amber sundial was traded for the road toll."
|
||||
memorybank.set_vector(bank["road"], None)
|
||||
memorybank.set_vector(bank["road"], [1.0, 0.0, 0.0])
|
||||
bank["sundial"].text = "Mara hid a bottle in the cellar."
|
||||
memorybank.set_vector(bank["sundial"], None)
|
||||
memorybank.set_vector(bank["sundial"], [1.0, 0.0, 0.0])
|
||||
db.commit()
|
||||
assert retrieve(adventure, settings)["used"][0]["id"] == bank["road"].id
|
||||
@@ -0,0 +1,225 @@
|
||||
"""v1.1 WP-B.2: the memory summariser, after the rejected B2.4 prompt experiment.
|
||||
|
||||
B2.4 tried a memory prompt instructing the model to keep named facts and objects.
|
||||
Measured against the reference model, it did not correct the creation failure it
|
||||
was for, and it was not shipped (`V1.1-WP-B2-REPORT.md` §T). The shipped prompt
|
||||
is v1.0.0's.
|
||||
|
||||
This file keeps two kinds of test apart.
|
||||
|
||||
**Acceptance tests** gate the tree:
|
||||
- the shipped memory prompt is exactly v1.0.0's, so the experiment is gone;
|
||||
- every fidelity fixture reaches the summariser whole, through the application's
|
||||
own prompt assembly;
|
||||
- a long memory is stored as the model wrote it, never cut;
|
||||
- the memory the attempt-2 block should have produced ranks first under B2.1.
|
||||
|
||||
**Diagnostic-measurement tests** check only that `tools/memory_fidelity.py`
|
||||
measures correctly: fact retention, attribution, invention, word count, a leading
|
||||
"Memory:", second person and promise retention, on hand-written memories whose
|
||||
answers are known. What a real model scores on those measurements is
|
||||
nondeterministic, is taken with inference, and is reported. It is never a gate
|
||||
here.
|
||||
|
||||
python -m pytest tests/test_v11_b2_summarizer_fidelity.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import re
|
||||
import subprocess
|
||||
|
||||
import pytest
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.database import Base, SessionLocal, engine
|
||||
from tools import memory_diagnostic as md
|
||||
from tools import memory_fidelity as mf
|
||||
|
||||
# ==================================================================== acceptance
|
||||
|
||||
|
||||
def test_the_shipped_memory_prompt_is_v1_0_0s():
|
||||
"""The B2.4 experiment is reverted: production sends the prompt v1.0.0 and
|
||||
WP-B.1 shipped, unchanged."""
|
||||
try:
|
||||
source = subprocess.run(["git", "show", "beb17ad:backend/app/memorybank.py"],
|
||||
capture_output=True, text=True, check=True).stdout
|
||||
except (OSError, subprocess.CalledProcessError):
|
||||
pytest.skip("git history not available")
|
||||
block = re.search(r"^MEMORY_SYSTEM_PROMPT = \((.*?)^\)$", source, re.S | re.M).group(1)
|
||||
shipped = eval(f"({block})", {"MEMORY_MAX_WORDS": 50}) # noqa: S307 - our own source
|
||||
assert memorybank.MEMORY_SYSTEM_PROMPT == shipped
|
||||
assert memorybank.MEMORY_MAX_WORDS == 50
|
||||
|
||||
|
||||
def test_the_rejected_experiment_is_not_what_ships():
|
||||
assert mf.B24_EXPERIMENT_PROMPT != memorybank.MEMORY_SYSTEM_PROMPT
|
||||
assert "Keep each fact with the person it belongs to" not in memorybank.MEMORY_SYSTEM_PROMPT
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
|
||||
def test_every_fixture_reaches_the_summariser_whole(fixture):
|
||||
"""Creation can only fail at the model if the fact was sent. Each fixture
|
||||
fits the excerpt budget, so the whole block is the excerpt."""
|
||||
user = mf.user_prompt_for(fixture)
|
||||
assert memorybank.count_tokens(fixture.raw) <= memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert f"Story excerpt:\n\n{fixture.raw}\n\nMemory:" in user
|
||||
assert user.startswith("Cast:\n- " + fixture.protagonist + " — the protagonist.")
|
||||
|
||||
|
||||
class Scripted:
|
||||
def __init__(self, reply):
|
||||
self.reply = reply
|
||||
self.calls: list[tuple[str, str]] = []
|
||||
|
||||
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
|
||||
self.calls.append((system, user))
|
||||
return self.reply
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def campaign(db, fixture):
|
||||
user = models.User(is_guest=False, email="b2-fidelity@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
|
||||
adventure = models.Adventure(user_id=user.id, title="Fidelity", script_state={}, auto_summarize=True,
|
||||
persona_name=fixture.protagonist)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
nodes = []
|
||||
for kind, text in fixture.actions + (("ai", "The story moves on."),):
|
||||
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
nodes.append(action)
|
||||
db.commit()
|
||||
return adventure, nodes
|
||||
|
||||
|
||||
def write_memory(db, adventure, monkeypatch, provider):
|
||||
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
|
||||
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: provider)
|
||||
settings = db.query(models.Settings).first()
|
||||
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
|
||||
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
|
||||
|
||||
|
||||
def test_the_application_sends_the_shipped_prompt_and_the_whole_planting_block(db, monkeypatch):
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
adventure, nodes = campaign(db, fixture)
|
||||
provider = Scripted(fixture.faithful)
|
||||
[memory] = write_memory(db, adventure, monkeypatch, provider)
|
||||
system, user = provider.calls[0]
|
||||
assert system == memorybank.MEMORY_SYSTEM_PROMPT
|
||||
assert "> I watch Mara slip the amber sundial inside the cracked teapot" in user
|
||||
assert (memory.source_start, memory.source_end) == (nodes[0].depth, nodes[5].depth)
|
||||
|
||||
|
||||
def test_an_over_long_memory_is_stored_as_written_never_cut(db, monkeypatch):
|
||||
"""The word target is an instruction, not a truncation: cutting a memory
|
||||
after the fact can split or drop exactly the fact it was written to keep."""
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
adventure, _ = campaign(db, fixture)
|
||||
long_reply = fixture.faithful + " " + " ".join(["They advanced cautiously through the dark."] * 12)
|
||||
[memory] = write_memory(db, adventure, monkeypatch, Scripted(long_reply))
|
||||
assert memory.text == long_reply
|
||||
assert len(memory.text.split()) > 2 * memorybank.MEMORY_MAX_WORDS
|
||||
|
||||
|
||||
def test_a_faithful_regression_memory_ranks_first_for_its_question():
|
||||
"""If the summariser keeps the fact, B2.1 finds it: the memory the attempt-2
|
||||
block should have produced, among the memories its bank really held for that
|
||||
stretch, under production scoring."""
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
stored, _ = fixture.unfaithful[0]
|
||||
bank = {
|
||||
1: fixture.faithful,
|
||||
2: stored,
|
||||
3: "Aldric, Mara and Edrin advanced through the cold crypt, the silver key heavy in Aldric's hands.",
|
||||
4: "Aldric told Mara the silver key opens the crypt beneath the Old Abbey.",
|
||||
5: "Rain kept falling on Westhaven as the travellers walked toward the abbey grounds.",
|
||||
}
|
||||
embed = md.ConceptEmbedder.vector
|
||||
query = {"input": "> I ask Mara quietly where she hid the amber sundial.",
|
||||
"context": "Aldric and Mara in the Crooked Lantern, rain outside."}
|
||||
held = {i: embed(t) for i, t in bank.items()}
|
||||
terms = {i: memorybank.lexical_terms(t) for i, t in bank.items()}
|
||||
rows = memorybank.score_candidates(list(bank), held, terms, embed(query["input"]),
|
||||
embed(query["context"]),
|
||||
sorted(memorybank.lexical_terms(query["input"])))
|
||||
assert rows[0][1] == 1
|
||||
assert rows[0][3] > 0
|
||||
|
||||
|
||||
# ======================================================= diagnostic measurements
|
||||
# These prove the measuring instrument. They say nothing about any model.
|
||||
|
||||
|
||||
def test_the_fixtures_cover_each_measurement_in_more_than_one_genre():
|
||||
requirements = {f.requirement for f in mf.FIXTURES}
|
||||
assert {"distinctive object and place", "player-established concrete fact", "promise / commitment",
|
||||
"attribution", "clutter pressure", "no invention", "multiple concrete facts",
|
||||
"the actual failed-run block"} <= requirements
|
||||
assert {"office", "contemporary", "science-fiction-neutral"} <= {f.genre for f in mf.FIXTURES}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
|
||||
def test_the_checker_passes_a_faithful_memory(fixture):
|
||||
result = mf.evaluate(fixture, fixture.faithful)
|
||||
assert result["passed"], result
|
||||
assert not result["over_target"] and not result["memory_prefix"] and not result["second_person"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fixture, memory, reason", [
|
||||
(f, memory, reason) for f in mf.FIXTURES for memory, reason in f.unfaithful
|
||||
], ids=lambda v: v.fixture_id if isinstance(v, mf.Fixture) else None)
|
||||
def test_the_checker_fails_each_failure_shape(fixture, memory, reason):
|
||||
result = mf.evaluate(fixture, memory)
|
||||
assert not result["passed"], result
|
||||
if reason == "not retained":
|
||||
assert not result["retained"]
|
||||
elif reason == "misattributed":
|
||||
assert result["misattributed"]
|
||||
elif reason == "invented":
|
||||
assert result["inventions"]
|
||||
|
||||
|
||||
def test_the_checker_reads_the_stored_attempt_2_memory_as_the_real_failure():
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
stored, _ = fixture.unfaithful[0]
|
||||
result = mf.evaluate(fixture, stored)
|
||||
assert result["retained"] is False and result["words"] == 102 and result["over_target"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("memory, prefix, you", [
|
||||
("Memory: Dana promised Marcus the lease by Friday.", True, False),
|
||||
(" memory: Dana promised the lease.", True, False),
|
||||
("You thanked Marcus and left.", False, True),
|
||||
("Dana thanked Marcus; your lease is due.", False, True),
|
||||
("Dana promised Marcus she would bring the signed lease by Friday.", False, False),
|
||||
])
|
||||
def test_the_checker_measures_framing(memory, prefix, you):
|
||||
result = mf.evaluate(mf.FIXTURES_BY_ID["promise_contemporary"], memory)
|
||||
assert result["memory_prefix"] is prefix
|
||||
assert result["second_person"] is you
|
||||
|
||||
|
||||
def test_the_checker_measures_promise_retention():
|
||||
fixture = mf.FIXTURES_BY_ID["promise_contemporary"]
|
||||
kept = mf.evaluate(fixture, "Dana promised to bring Marcus the signed lease by Friday.")
|
||||
scenery = mf.evaluate(fixture, "Memory: Dana looked around the empty living room while a dog barked.")
|
||||
assert kept["facts"]["lease by Friday"]["kept"] and kept["passed"]
|
||||
assert not scenery["facts"]["lease by Friday"]["kept"] and scenery["memory_prefix"]
|
||||
Reference in New Issue
Block a user