v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
beb17ada10
commit
0c1ba836ba
@@ -255,70 +255,101 @@ def _section(snapshot: dict, label: str) -> str:
|
||||
|
||||
# ------------------------------------------------------------------- stages
|
||||
|
||||
async def rank_bank(db, adventure, settings, query: str, embed) -> dict:
|
||||
async def rank_bank(db, adventure, settings, query: dict, embed) -> dict:
|
||||
"""Production's ranking, recomputed for `query`, for every eligible memory.
|
||||
|
||||
The same catalogue clause, the same cosine, the same pin rule, the same
|
||||
redundancy suppression helper. Returns every scored row, not just the top-k,
|
||||
because "where did F rank" is the question.
|
||||
`query` is a `memorybank.retrieval_query` dict. The catalogue clause, the
|
||||
scoring (`memorybank.score_candidates`) and the selection with its pins and
|
||||
redundancy rule (`memorybank.select_memories`) are production's own
|
||||
functions, so this is production's ranking, not a second opinion. Returns
|
||||
every scored row, not just the top-k, because "where did F rank" is the
|
||||
question.
|
||||
"""
|
||||
catalogue = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.authority,
|
||||
models.Memory.embedding_blob).where(
|
||||
models.Memory.embedding_blob, models.Memory.text).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory),
|
||||
models.Memory.forgotten.is_(False),
|
||||
models.Memory.embedded.is_(True),
|
||||
)
|
||||
).all()
|
||||
if not catalogue or not query.strip():
|
||||
return {"query": query, "scored": [], "selected": [], "top_k": settings.memory_top_k}
|
||||
[query_vec] = await embed([query])
|
||||
held = {row.id: vectors.unpack(row.embedding_blob) for row in catalogue if row.embedding_blob}
|
||||
authority_of = {row.id: row.authority for row in catalogue}
|
||||
scored = sorted(
|
||||
((vectors.cosine(query_vec, held[row.id]), row.id, row.pinned)
|
||||
for row in catalogue if row.id in held),
|
||||
key=lambda r: r[0], reverse=True,
|
||||
)
|
||||
top_k = max(1, settings.memory_top_k)
|
||||
used = [r for r in scored if r[2]]
|
||||
remaining = max(0, top_k - len(used))
|
||||
candidates = [r for r in scored if not r[2]]
|
||||
kept, suppressed = memorybank._drop_redundant(candidates, held, authority_of, remaining)
|
||||
selected = {r[1] for r in used + kept}
|
||||
texts = [t for t in (query["input"], query["context"]) if t.strip()]
|
||||
if not catalogue or not texts:
|
||||
return {"query": query, "scored": [], "selected": [], "top_k": top_k}
|
||||
vectors_by_text = dict(zip(texts, await embed(texts)))
|
||||
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
|
||||
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
|
||||
held = {row.id: vectors.unpack(row.embedding_blob) for row in catalogue if row.embedding_blob}
|
||||
terms_of = ({row.id: memorybank.lexical_terms(row.text or "") for row in catalogue}
|
||||
if query["input_terms"] else {})
|
||||
authority_of = {row.id: row.authority for row in catalogue}
|
||||
pinned_of = {row.id: row.pinned for row in catalogue}
|
||||
scored = memorybank.score_candidates(
|
||||
[row.id for row in catalogue if row.id in held], held, terms_of,
|
||||
input_vec, context_vec, query["input_terms"])
|
||||
used, suppressed = memorybank.select_memories(scored, pinned_of, held, authority_of, top_k)
|
||||
selected = {row[1] for row in used}
|
||||
suppressed_by = dict(suppressed)
|
||||
return {
|
||||
"query": query,
|
||||
"top_k": top_k,
|
||||
"scored": [
|
||||
{"rank": i + 1, "memory_id": memory_id, "similarity": round(score, 4),
|
||||
"pinned": pinned, "selected": memory_id in selected,
|
||||
{"rank": i + 1, "memory_id": memory_id, "similarity": round(semantic, 4),
|
||||
"semantic_score": round(semantic, 4), "lexical_score": round(lexical, 4),
|
||||
"final_score": round(final, 4),
|
||||
"pinned": pinned_of[memory_id], "selected": memory_id in selected,
|
||||
"suppressed_as_duplicate_of": suppressed_by.get(memory_id)}
|
||||
for i, (score, memory_id, pinned) in enumerate(scored)
|
||||
for i, (final, memory_id, semantic, lexical) in enumerate(scored)
|
||||
],
|
||||
"selected": sorted(selected),
|
||||
}
|
||||
|
||||
|
||||
def production_query(adventure, exclude_action_id: int | None) -> str:
|
||||
"""The retrieval query a turn used: its newest actions, as `retrieve_memories` builds it."""
|
||||
recent = history.tail(adventure, memorybank.RETRIEVAL_WINDOW_ACTIONS, exclude_action_id)
|
||||
return builder.truncate_to_last_tokens(
|
||||
"\n\n".join(a.text for a in recent), memorybank.RETRIEVAL_WINDOW_TOKENS)
|
||||
def production_query(adventure, exclude_action_id: int | None) -> dict:
|
||||
"""The retrieval query a turn used, built by production's own `retrieval_query`."""
|
||||
return memorybank.retrieval_query(adventure, exclude_action_id)
|
||||
|
||||
|
||||
def variant_query(base: dict, player_input: str) -> dict:
|
||||
"""`base` with a different player input: "what if the player had asked this
|
||||
here", with the scene and narration context the recall turn really had."""
|
||||
return {"input": player_input, "context": base["context"],
|
||||
"input_terms": sorted(memorybank.lexical_terms(player_input))}
|
||||
|
||||
|
||||
def eviction_order(db, adventure) -> list[int]:
|
||||
"""The order `_evict_over_capacity` would take unpinned active memories in."""
|
||||
from sqlalchemy import func
|
||||
return db.execute(
|
||||
select(models.Memory.id).where(
|
||||
"""The order `_evict_over_capacity` would take unpinned active memories in.
|
||||
|
||||
Production's own `memorybank.eviction_order`, run to the end of the bank."""
|
||||
rows = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.last_used_at,
|
||||
models.Memory.created_at, models.Memory.use_count).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.forgotten.is_(False),
|
||||
models.Memory.pinned.is_(False),
|
||||
).order_by(func.coalesce(models.Memory.last_used_at, models.Memory.created_at),
|
||||
models.Memory.use_count)
|
||||
).scalars().all()
|
||||
)
|
||||
).all()
|
||||
return memorybank.eviction_order(rows, len(rows))
|
||||
|
||||
|
||||
def coverage(ranges: list[tuple[int, int]], tip: int | None) -> dict:
|
||||
"""How much of the story `ranges` (active memories' source ranges) describe.
|
||||
|
||||
`largest_gap` is the longest run of depths, between the first memory's start
|
||||
and `tip`, that no memory covers. It is how the eviction rule is judged in
|
||||
general, not only for the planted fact."""
|
||||
if not ranges:
|
||||
return {"first_start": None, "last_end": None, "largest_gap": None}
|
||||
ordered = sorted(ranges)
|
||||
gaps = []
|
||||
reach = ordered[0][1]
|
||||
for start, end in ordered[1:]:
|
||||
gaps.append(max(0, start - reach - 1))
|
||||
reach = max(reach, end)
|
||||
return {"first_start": ordered[0][0], "last_end": reach,
|
||||
"largest_gap": max(gaps, default=0)}
|
||||
|
||||
|
||||
async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
|
||||
@@ -342,7 +373,7 @@ async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
|
||||
for memory in covering:
|
||||
block = memorybank.source_block(db, memory)
|
||||
raw = "\n\n".join(a.text for a in block)
|
||||
excerpt = builder.truncate_to_last_tokens(raw, memorybank.MEMORY_EXCERPT_TOKENS)
|
||||
excerpt = memorybank.memory_excerpt(raw) # what `summarize_block` sends
|
||||
creation_input.append({
|
||||
"memory_id": memory.id, "source_start": memory.source_start,
|
||||
"source_end": memory.source_end, "block_tokens": builder.count_tokens(raw),
|
||||
@@ -394,6 +425,8 @@ async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
|
||||
out["verdict"] = "created_but_evicted"
|
||||
return out
|
||||
|
||||
# The recall turn's AI node is excluded, so the newest action is the recall
|
||||
# player action, exactly as the turn saw it before it wrote its reply.
|
||||
query = production_query(adventure, recall_action.id)
|
||||
ranking = await rank_bank(db, adventure, settings, query, embed)
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory.id), None)
|
||||
@@ -401,9 +434,10 @@ async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
|
||||
out["ranked"] = {
|
||||
"yes": row is not None and row["rank"] <= ranking["top_k"],
|
||||
"eligible": row is not None,
|
||||
"lexical_score": None, # memory ranking has no lexical term (CONTEXT-AND-MEMORY §20)
|
||||
"semantic_score": row["similarity"] if row else None,
|
||||
"final_score": row["similarity"] if row else None,
|
||||
"lexical_score": row["lexical_score"] if row else None,
|
||||
"semantic_score": row["semantic_score"] if row else None,
|
||||
"final_score": row["final_score"] if row else None,
|
||||
"selected_top_k": ranking["selected"],
|
||||
"pin_effect": "always selected" if memory.pinned else "none",
|
||||
"rank": row["rank"] if row else None,
|
||||
"of": len(ranking["scored"]),
|
||||
@@ -512,6 +546,9 @@ PLACES = (
|
||||
)
|
||||
PARAPHRASE_QUERY = "I ask Mara where she tucked the little brass dial that tells the hour."
|
||||
UNRELATED_QUERY = "I ask the ferryman what rope costs at the landing this season."
|
||||
#: Built only from words every fixture memory holds ("travellers", "spent",
|
||||
#: "time"), so its rarity weight is zero everywhere.
|
||||
COMMON_WORDS_QUERY = "The travellers spent time."
|
||||
|
||||
|
||||
def filler_prose(index: int, words: int) -> str:
|
||||
@@ -539,6 +576,10 @@ class Scenario:
|
||||
pin_first_memory: bool = False
|
||||
lineage_control: bool = False
|
||||
diagnose_recall: bool = True
|
||||
#: The narration of the last turn before recall, when a fixture needs the
|
||||
#: scene to say something (WP-B.2's context-dependent question). Must not
|
||||
#: name a planted fact.
|
||||
pre_recall_reply: str = ""
|
||||
|
||||
|
||||
SCENARIOS = {
|
||||
@@ -553,6 +594,23 @@ SCENARIOS = {
|
||||
"long_block_fact_late": Scenario("long_block_fact_late", turns=10, prose_words=850,
|
||||
plant_turn=3, budget=16384),
|
||||
"lineage_control": Scenario("lineage_control", lineage_control=True),
|
||||
# v1.1 WP-B.2: the ranking failure B.1 saw on the real model, made
|
||||
# deterministic. Longer narration fills the v1.0.0 query, and `memory_top_k`
|
||||
# is the real run's 4. Below capacity, isolation valid.
|
||||
"ranking_crowded": Scenario("ranking_crowded", prose_words=150, top_k=4),
|
||||
# v1.1 WP-B.2: all three B.1 failures at once. Every block is longer than the
|
||||
# summariser's excerpt, narration crowds the query at `memory_top_k` 4, and
|
||||
# the bank passes a capacity of 8 long before recall at depth 106.
|
||||
"independent_full": Scenario("independent_full", prose_words=850, top_k=4,
|
||||
capacity=8, budget=16384),
|
||||
# v1.1 WP-B.2: a question that names neither the sundial nor the teapot and
|
||||
# cannot be answered without the scene. The last narration puts Mara at the
|
||||
# tavern's top shelf; the player asks "her" what she put "up there".
|
||||
"ranking_context_dependent": Scenario(
|
||||
"ranking_context_dependent", prose_words=150, top_k=4,
|
||||
recall_text="I ask her what she keeps up there.",
|
||||
pre_recall_reply=("Mara stands on a stool at the tavern's top shelf, running a cloth "
|
||||
"around the old kettle up there, and she will not meet your eye.")),
|
||||
}
|
||||
|
||||
|
||||
@@ -662,6 +720,16 @@ def run_scenario(scenario: Scenario) -> dict:
|
||||
models.Memory.created_at).where(models.Memory.adventure_id == adv)
|
||||
.order_by(models.Memory.id)).all()]
|
||||
|
||||
def per_turn_isolation(adventure_id):
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adventure_id)
|
||||
active = summaries.current(db, adventure)
|
||||
return {
|
||||
"f_in_state": FACT_F.mentioned_by(json.dumps(adventure.narrative_state or {},
|
||||
default=str)),
|
||||
"f_in_summary": FACT_F.mentioned_by(active.text if active is not None else ""),
|
||||
}
|
||||
|
||||
def turn(kind, text, reply):
|
||||
ScriptNarrator.next_reply = f"{reply}\n```state\n{{\"events\": []}}\n```"
|
||||
response = client.post(f"/api/adventures/{adv}/actions", json={"type": kind, "text": text})
|
||||
@@ -703,6 +771,8 @@ def run_scenario(scenario: Scenario) -> dict:
|
||||
turn("do", f"I turn away from the mill and walk to the {PLACES[n % len(PLACES)]}.",
|
||||
filler_prose(n + 100, scenario.prose_words))
|
||||
g["diverged_at_turn"] = n
|
||||
elif n == scenario.turns and scenario.pre_recall_reply:
|
||||
turn("do", "I head back to the tavern.", scenario.pre_recall_reply)
|
||||
else:
|
||||
turn("do", f"I walk on to the {PLACES[n % len(PLACES)]}.",
|
||||
filler_prose(n, scenario.prose_words))
|
||||
@@ -735,6 +805,11 @@ def run_scenario(scenario: Scenario) -> dict:
|
||||
"f_memory_id": f_memory_id,
|
||||
"f_forgotten": bool(f_row and f_row["forgotten"]),
|
||||
"f_use_count": f_row["use_count"] if f_row else None,
|
||||
"coverage": coverage([(r["source_start"], r["source_end"]) for r in rows
|
||||
if not r["forgotten"] and r["source_start"] is not None],
|
||||
None),
|
||||
# Isolation on every turn, not only at recall (WP-B.2 acceptance).
|
||||
**per_turn_isolation(adv),
|
||||
})
|
||||
|
||||
turn("do", scenario.recall_text, filler_prose(999, scenario.prose_words))
|
||||
@@ -757,18 +832,64 @@ def run_scenario(scenario: Scenario) -> dict:
|
||||
db, adventure, settings, FACT_F, plant_depth,
|
||||
recall_action=recall_action, embed=embedder.embed))
|
||||
result["summariser_excerpts"] = len(summariser.excerpts)
|
||||
# Provenance as the recall turn recorded it, resolved back to rows.
|
||||
used = (recall_action.context_snapshot.get("memories") or {}).get("used") or []
|
||||
f_entry = next((m for m in used
|
||||
if m.get("id") == result["diagnosis"]["created"]["memory_id"]), None)
|
||||
f_memory = db.get(models.Memory, f_entry["id"]) if f_entry else None
|
||||
block = memorybank.source_block(db, f_memory) if f_memory is not None else []
|
||||
result["provenance"] = {
|
||||
"recorded": f_entry and {k: f_entry.get(k) for k in
|
||||
("id", "source", "semantic_score", "lexical_score",
|
||||
"final_score", "authority")},
|
||||
"range_covers_plant": bool(f_entry and f_entry["source"]["source_start"]
|
||||
<= plant_depth <= f_entry["source"]["source_end"]),
|
||||
"matches_row": bool(f_memory is not None and f_entry["source"] == {
|
||||
"branch_id": f_memory.branch_id, "depth": f_memory.depth,
|
||||
"source_start": f_memory.source_start, "source_end": f_memory.source_end}),
|
||||
"source_block_depths": [a.depth for a in block],
|
||||
"source_block_holds_planting": any(a.text == FACT_F.sentence for a in block),
|
||||
}
|
||||
|
||||
memory_id = result["diagnosis"]["created"]["memory_id"]
|
||||
if memory_id is not None and not result["diagnosis"]["retained"]["forgotten"]:
|
||||
variants = {}
|
||||
for label, query in (("direct", scenario.recall_text),
|
||||
("paraphrase", PARAPHRASE_QUERY),
|
||||
("unrelated", UNRELATED_QUERY)):
|
||||
base = production_query(adventure, recall_action.id)
|
||||
# WP-B.2's negative control needs a decoy: another memory that
|
||||
# holds a word the question adds, and nothing about F.
|
||||
decoy = next(((m.id, found.group(1)) for m in db.execute(
|
||||
select(models.Memory).where(models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.forgotten.is_(False),
|
||||
models.Memory.id != memory_id)
|
||||
.order_by(models.Memory.id)).scalars()
|
||||
if (found := re.search(r"at the ([a-z]+ [a-z]+)\.", m.text or ""))), None)
|
||||
queries = [
|
||||
("direct", variant_query(base, scenario.recall_text)),
|
||||
("paraphrase", variant_query(base, PARAPHRASE_QUERY)),
|
||||
("unrelated", variant_query(base, UNRELATED_QUERY)),
|
||||
# The player's words with the context taken away.
|
||||
("input_only", {**variant_query(base, scenario.recall_text), "context": ""}),
|
||||
# Words every memory in these fixtures holds, and nothing else.
|
||||
("common_words", variant_query(base, COMMON_WORDS_QUERY)),
|
||||
]
|
||||
if decoy is not None:
|
||||
queries.append(("rare_word_with_paraphrase", variant_query(
|
||||
base, PARAPHRASE_QUERY[:-1] + f", out by the {decoy[1]}.")))
|
||||
for label, query in queries:
|
||||
ranking = asyncio.run(rank_bank(db, adventure, settings, query, embedder.embed))
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
|
||||
variants[label] = {"query": query, "rank": row and row["rank"],
|
||||
decoy_row = next((r for r in ranking["scored"]
|
||||
if decoy is not None and r["memory_id"] == decoy[0]), None)
|
||||
text = query["input"]
|
||||
variants[label] = {"query": text, "rank": row and row["rank"],
|
||||
"selected_count": len(ranking["selected"]),
|
||||
"decoy_memory_id": decoy and decoy[0],
|
||||
"decoy_rank": decoy_row and decoy_row["rank"],
|
||||
"decoy_lexical_score": decoy_row and decoy_row["lexical_score"],
|
||||
"of": len(ranking["scored"]),
|
||||
"similarity": row and row["similarity"],
|
||||
"lexical_score": row and row["lexical_score"],
|
||||
"final_score": row and row["final_score"],
|
||||
"selected": bool(row and row["selected"]),
|
||||
"top_k": ranking["top_k"]}
|
||||
result["ranking_variants"] = variants
|
||||
|
||||
Reference in New Issue
Block a user