v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
beb17ada10
commit
0c1ba836ba
@@ -11,8 +11,9 @@ diagnostic in `tools/memory_diagnostic.py`:
|
||||
2. **What it finds on this tree.** The scenarios run with a best-case summariser,
|
||||
one that keeps a fact if and only if the fact reached it. Any failure is
|
||||
therefore the application's mechanism, not a model's writing.
|
||||
- The criteria the current code does not meet are marked `xfail(strict=True)`,
|
||||
so B.2 has to flip them deliberately.
|
||||
- WP-B.1 marked the criteria v1.0.0 did not meet `xfail(strict=True)`. WP-B.2
|
||||
fixed ranking, eviction and the creation excerpt, and those tests are now
|
||||
ordinary passes; the v1.0.0 results are recorded in the WP-B.2 report.
|
||||
- The same file is run unchanged against v1.0.0 for the baseline.
|
||||
|
||||
python -m pytest tests/test_v11_b1_memory_diagnostic.py -v
|
||||
@@ -102,14 +103,14 @@ def test_retention_is_reported_with_the_bank_and_its_eviction_order():
|
||||
def test_ranking_is_production_ranking_and_agrees_with_the_stored_selection():
|
||||
ranked = scenario("independent_default")["diagnosis"]["ranked"]
|
||||
assert ranked["replica_matches_stored_selection"] is True
|
||||
assert ranked["lexical_score"] is None # memory ranking has no lexical term
|
||||
assert ranked["top_k_cutoff"] == 5
|
||||
assert ranked["yes"] is True and ranked["selected"] is True
|
||||
assert 1 <= ranked["rank"] <= ranked["top_k_cutoff"]
|
||||
# The production query is the newest four actions, cut to 600 tokens, and the
|
||||
# one-line question is diluted by the narration around it.
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert ranked["semantic_score"] < variants["direct"]["similarity"]
|
||||
# v1.1 WP-B.2: every part of the score is reported, and they add up.
|
||||
assert 0.0 <= ranked["lexical_score"] <= 1.0
|
||||
assert ranked["final_score"] == pytest.approx(
|
||||
ranked["semantic_score"] + memorybank.LEXICAL_WEIGHT * ranked["lexical_score"], abs=2e-4)
|
||||
assert ranked["query"]["input"].endswith(md.SCENARIOS["independent_default"].recall_text)
|
||||
|
||||
|
||||
def test_injection_is_read_from_the_recall_turns_own_context():
|
||||
@@ -124,53 +125,118 @@ def test_ranking_variants_direct_paraphrase_and_unrelated():
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["direct"]["rank"] == 1 and variants["direct"]["selected"]
|
||||
assert variants["paraphrase"]["rank"] == 1 and variants["paraphrase"]["selected"]
|
||||
assert (variants["direct"]["similarity"] > variants["paraphrase"]["similarity"]
|
||||
> 5 * variants["unrelated"]["similarity"])
|
||||
assert (variants["direct"]["final_score"] > variants["paraphrase"]["final_score"]
|
||||
> variants["unrelated"]["final_score"])
|
||||
|
||||
|
||||
def test_retrieval_fills_top_k_whatever_the_similarity():
|
||||
"""Diagnosis: there is no relevance floor. With more memories than
|
||||
`memory_top_k`, an unrelated query still selects five, and the early fact
|
||||
rides along at a similarity near zero."""
|
||||
def test_retrieval_still_fills_top_k_whatever_the_similarity():
|
||||
"""There is still no relevance floor: an unrelated question selects a full
|
||||
`memory_top_k`. B.2 changed which memories those are, not how many — the
|
||||
early fact is no longer carried along by an unrelated question."""
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["unrelated"]["similarity"] < 0.1
|
||||
assert variants["unrelated"]["selected"] is True
|
||||
assert variants["unrelated"]["selected_count"] == 5
|
||||
assert variants["unrelated"]["selected"] is False
|
||||
assert variants["unrelated"]["rank"] > 5
|
||||
|
||||
|
||||
# ------------------------------------------------- WP-B.2 ranking acceptance
|
||||
|
||||
@pytest.mark.parametrize("name", ["ranking_crowded", "ranking_context_dependent"])
|
||||
def test_acceptance_the_early_memory_is_ranked_and_injected_below_capacity(name):
|
||||
"""The B.1 ranking failure, made deterministic. On v1.0.0 both fixtures are
|
||||
`retained_but_not_ranked` (ranks 7 and 6 of 17 against a top-k of 4)."""
|
||||
result = scenario(name)
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
diagnosis = result["diagnosis"]
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert diagnosis["retained"]["active_memories"] <= diagnosis["retained"]["memory_bank_capacity"]
|
||||
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["rank"] <= 4
|
||||
assert diagnosis["ranked"]["replica_matches_stored_selection"] is True
|
||||
assert diagnosis["injected"]["yes"] is True
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_the_crowded_fixture_is_won_by_the_players_question():
|
||||
ranked = scenario("ranking_crowded")["diagnosis"]["ranked"]
|
||||
assert ranked["rank"] == 1
|
||||
assert ranked["lexical_score"] > 0 # "sundial" and "amber" are in the question
|
||||
|
||||
|
||||
def test_a_paraphrase_is_found_by_meaning_not_by_shared_words():
|
||||
"""Lexical matching must not replace semantic retrieval. The paraphrase
|
||||
shares none of F's distinctive words, yet ranks first."""
|
||||
for name in ("ranking_crowded", "independent_default"):
|
||||
paraphrase = scenario(name)["ranking_variants"]["paraphrase"]
|
||||
assert paraphrase["rank"] == 1 and paraphrase["selected"]
|
||||
# Only "Mara" is shared, which is far less than the direct question holds.
|
||||
direct = scenario(name)["ranking_variants"]["direct"]
|
||||
assert paraphrase["lexical_score"] < direct["lexical_score"] / 2
|
||||
|
||||
|
||||
def test_a_context_dependent_question_needs_the_scene():
|
||||
""""I ask her what she keeps up there" names nothing F's memory holds. The
|
||||
scene the last narration set up (Mara, the top shelf, a kettle) is what
|
||||
finds it; without that context it ranks last."""
|
||||
result = scenario("ranking_context_dependent")
|
||||
ranked = result["diagnosis"]["ranked"]
|
||||
assert ranked["lexical_score"] == 0.0
|
||||
assert ranked["rank"] <= 4
|
||||
assert "top shelf" in ranked["query"]["context"]
|
||||
assert result["ranking_variants"]["input_only"]["rank"] > 4
|
||||
|
||||
|
||||
def test_an_unrelated_rare_word_does_not_outrank_the_relevant_memory():
|
||||
"""Negative control: the paraphrase plus a place only one other memory
|
||||
holds. The decoy gains lexical score, and still ranks below F."""
|
||||
for name in ("ranking_crowded", "independent_default"):
|
||||
control = scenario(name)["ranking_variants"]["rare_word_with_paraphrase"]
|
||||
assert control["decoy_lexical_score"] > control["lexical_score"]
|
||||
assert control["rank"] == 1
|
||||
assert control["decoy_rank"] > control["rank"]
|
||||
|
||||
|
||||
def test_common_words_contribute_nothing():
|
||||
common = scenario("independent_default")["ranking_variants"]["common_words"]
|
||||
assert common["lexical_score"] == 0.0
|
||||
|
||||
|
||||
# ---------------------------------------------------------- capacity/eviction
|
||||
|
||||
def test_past_capacity_the_early_memory_is_evicted_and_the_stage_says_so():
|
||||
"""Diagnosis, not a requirement: what the current eviction rule does to F."""
|
||||
result = scenario("past_capacity")
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
assert result["diagnosis"]["verdict"] == "created_but_evicted"
|
||||
eviction = result["eviction"]
|
||||
assert eviction["f_evicted_at_turn"] is not None
|
||||
# It was retrieved while the bank was small, stopped being retrieved once
|
||||
# recent narration filled the top-k, and was then the least recently used.
|
||||
assert eviction["f_use_count_when_evicted"] > 0
|
||||
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
|
||||
assert eviction["f_memory_was_first_evicted"] is True
|
||||
@pytest.mark.parametrize("name", ["past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"])
|
||||
def test_past_capacity_the_early_memory_is_retained(name):
|
||||
"""v1.1 WP-B.2. On v1.0.0 all three are `created_but_evicted`: F was the
|
||||
least recently used row once recent narration stopped retrieving it, and
|
||||
went first (turns 21, 21 and 36). Coverage-first eviction keeps the only
|
||||
memory of the opening, so it stays active and is recalled at depth 106."""
|
||||
result = scenario(name)
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
diagnosis = result["diagnosis"]
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert result["eviction"]["f_evicted_at_turn"] is None
|
||||
# The bank really was past capacity, and stayed bounded.
|
||||
assert result["eviction"]["first_eviction_turn"] is not None
|
||||
assert all(t["active"] <= result["scenario"]["capacity"] + (1 if result["scenario"]["pin_first_memory"] else 0)
|
||||
for t in result["trace"])
|
||||
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
|
||||
|
||||
|
||||
def test_at_a_lower_top_k_the_early_memory_ages_out_after_it_stops_being_retrieved():
|
||||
"""Diagnosis with most of the bank unretrieved on any turn, nearer the
|
||||
shipped 5-in-80 ratio. F is not simply the oldest row: it is evicted some
|
||||
turns after recent narration stopped pulling it into the top-k, which is
|
||||
what ordering by last use does to a fact nothing recent mentions."""
|
||||
result = scenario("past_capacity_low_top_k")
|
||||
eviction = result["eviction"]
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
assert result["diagnosis"]["verdict"] == "created_but_evicted"
|
||||
assert eviction["f_use_count_when_evicted"] > 0
|
||||
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
|
||||
assert eviction["first_eviction_turn"] <= eviction["f_evicted_at_turn"]
|
||||
assert eviction["created_and_evicted_same_turn"] == []
|
||||
def test_past_capacity_the_bank_still_describes_the_whole_story():
|
||||
"""What the rule buys in general, not only for F: the active bank reaches
|
||||
from the opening to the newest block, and no stretch between them goes
|
||||
undescribed for more than twice the average spacing a bank of this capacity
|
||||
can afford (story span / capacity). On v1.0.0 these banks began at depths 36
|
||||
and 18: the opening was simply gone."""
|
||||
for name in ("past_capacity", "past_capacity_low_top_k"):
|
||||
result = scenario(name)
|
||||
cover = result["trace"][-1]["coverage"]
|
||||
assert cover["first_start"] == 0
|
||||
assert cover["last_end"] >= result["recall_depth"] - 2 * memorybank.MEMORY_INTERVAL
|
||||
assert cover["largest_gap"] <= 2 * (cover["last_end"] + 1) / result["scenario"]["capacity"]
|
||||
|
||||
|
||||
def test_no_memory_is_evicted_by_the_same_pass_that_created_it():
|
||||
"""The frozen-bank regression the current rule fixed, still holding."""
|
||||
for name in ("past_capacity", "past_capacity_pinned"):
|
||||
"""The frozen-bank regression the v1.0.0 rule fixed, still holding."""
|
||||
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
|
||||
assert scenario(name)["eviction"]["created_and_evicted_same_turn"] == []
|
||||
|
||||
|
||||
@@ -180,25 +246,28 @@ def test_a_pinned_memory_survives_capacity():
|
||||
assert eviction["pinned_memory_forgotten"] is False
|
||||
|
||||
|
||||
@pytest.mark.xfail(strict=True, reason=(
|
||||
"WP-B.1 diagnosis on this tree: past memory_bank_capacity the planting-era "
|
||||
"memory is evicted first, because it was never retrieved and eviction orders "
|
||||
"by last use, then creation. B.2 must flip this deliberately."))
|
||||
def test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity():
|
||||
assert scenario("past_capacity")["diagnosis"]["verdict"] == "injected"
|
||||
"""Was `xfail(strict=True)` in WP-B.1; B.2 fixed the eviction rule."""
|
||||
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
|
||||
assert scenario(name)["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ---------------------------------------------------------- creation window
|
||||
|
||||
def test_a_fact_early_in_a_long_block_never_reaches_the_summariser():
|
||||
def test_a_fact_early_in_a_long_block_now_reaches_the_summariser():
|
||||
"""v1.1 WP-B.2. On v1.0.0 this block (2,079 tokens) was cut to its last
|
||||
2,000, the fact at its start was never seen, and the stage was
|
||||
`not_created`. The excerpt is now the block's opening and end."""
|
||||
result = scenario("long_block_fact_early")
|
||||
created = result["diagnosis"]["created"]
|
||||
covering = created["covering_memories"]
|
||||
assert covering, "the long block must have been summarised"
|
||||
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert covering[0]["fact_in_block"] is True
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is False
|
||||
assert result["diagnosis"]["verdict"] == "not_created"
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is True
|
||||
assert created["yes"] is True
|
||||
assert created["source_start"] <= result["plant_depth"] <= created["source_end"]
|
||||
assert memorybank.EXCERPT_OMISSION_MARKER not in created["memory_text"]
|
||||
|
||||
|
||||
def test_the_same_fact_late_in_the_same_sized_block_does():
|
||||
@@ -209,12 +278,69 @@ def test_the_same_fact_late_in_the_same_sized_block_does():
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
|
||||
|
||||
@pytest.mark.xfail(strict=True, reason=(
|
||||
"WP-B.1 diagnosis on this tree: the summariser reads only the last "
|
||||
f"{memorybank.MEMORY_EXCERPT_TOKENS} tokens of a block, so a fact early in a "
|
||||
"long block is never seen. B.2 must flip this deliberately."))
|
||||
def test_acceptance_a_fact_early_in_a_long_block_is_remembered():
|
||||
"""Was `xfail(strict=True)` in WP-B.1; B.2 changed the excerpt."""
|
||||
assert scenario("long_block_fact_early")["diagnosis"]["created"]["yes"] is True
|
||||
assert scenario("long_block_fact_early")["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ------------------------------------------ WP-B.2 full deterministic acceptance
|
||||
|
||||
def test_acceptance_full_isolation_holds_on_every_turn():
|
||||
"""`independent_full`: long blocks, a crowded query, a bank past capacity.
|
||||
F must be carried by memory alone for the whole run, not only at recall."""
|
||||
result = scenario("independent_full")
|
||||
assert result["plant_depth"] <= 3 and result["recall_depth"] >= 100
|
||||
assert not any(t["f_in_state"] for t in result["trace"])
|
||||
assert not any(t["f_in_summary"] for t in result["trace"])
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
for check in ("state_document", "state_snapshots", "later_narration", "summary",
|
||||
"knowledge", "recent_history", "state_section"):
|
||||
assert result["isolation"]["checks"][check]["ok"], check
|
||||
|
||||
|
||||
def test_acceptance_full_every_stage_passes_past_capacity_with_long_blocks():
|
||||
"""On v1.0.0 this fixture fails at creation: every block is over 2,000
|
||||
tokens, and the fact at the start of the first one is never summarised."""
|
||||
result = scenario("independent_full")
|
||||
diagnosis = result["diagnosis"]
|
||||
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
|
||||
assert diagnosis["created"]["covering_memories"][0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["replica_matches_stored_selection"]
|
||||
assert diagnosis["injected"]["yes"]
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_acceptance_full_provenance_resolves_to_the_planting_turn():
|
||||
provenance = scenario("independent_full")["provenance"]
|
||||
assert provenance["recorded"] is not None
|
||||
assert provenance["range_covers_plant"] and provenance["matches_row"]
|
||||
assert provenance["source_block_holds_planting"] is True
|
||||
assert provenance["recorded"]["authority"] == memorybank.ACCEPTED_STORY
|
||||
|
||||
|
||||
def test_acceptance_full_is_the_long_run_independent_memory_verdict():
|
||||
"""The same measurements, judged by the long-run tool's own verdict."""
|
||||
from tools import m11_long_run as lr
|
||||
|
||||
result = scenario("independent_full")
|
||||
checks = result["isolation"]["checks"]
|
||||
diagnosis = result["diagnosis"]
|
||||
verdict = lr._independent_memory_verdict({
|
||||
"independent_planted_depth": result["plant_depth"],
|
||||
"planted_turn_outside_history": checks["recent_history"]["ok"],
|
||||
"absent_from_state": checks["state_document"]["ok"] and checks["state_snapshots"]["ok"]
|
||||
and not any(t["f_in_state"] for t in result["trace"]),
|
||||
"absent_from_summary": checks["summary"]["ok"]
|
||||
and not any(t["f_in_summary"] for t in result["trace"]),
|
||||
"absent_from_knowledge": checks["knowledge"]["ok"],
|
||||
"absent_from_later_narration": checks["later_narration"]["ok"],
|
||||
"memory_covering_planting_carries_fact": diagnosis["created"]["yes"],
|
||||
"memory_forgotten": not diagnosis["retained"]["yes"],
|
||||
"memory_injected": diagnosis["injected"]["yes"],
|
||||
})
|
||||
assert verdict == "recovered_through_memory_independent"
|
||||
|
||||
|
||||
# ------------------------------------------------------- lineage control (G)
|
||||
|
||||
Reference in New Issue
Block a user