v1.1 WP-B.1: diagnose independent long-term memory retention

Diagnostic only; no memory behaviour changes.

- tools/memory_diagnostic.py: planted-fact isolation checks, the four-stage
  diagnosis (created / retained / ranked / injected) with a verdict, a
  production-ranking replica, deterministic summariser/embedder/narrator
  stubs and seven scenarios (default, past capacity, pinned, low top_k,
  long-block early/late, lineage control)
- tools/v11_b1_memory.py: CLI for the scenarios and for diagnosing a copy of
  a finished real campaign
- tools/m11_long_run.py: opt-in --independent-fact mode with per-turn
  isolation tracking and the recovered_through_memory_independent verdict;
  M04 verdicts unchanged
- tests: diagnostic stages, eviction, creation window, ranking, lineage and
  authority controls; two strict xfails record the diagnosed retention and
  creation defects for WP-B.2 to flip
- planning/reports/v1.1/V1.1-WP-B1-REPORT.md

First failing stage: ranking (real model); retention past capacity and
creation for early facts in long blocks (deterministic, same on v1.0.0).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-14 20:50:05 -04:00
co-authored by Claude Opus 5
parent d63804f22e
commit beb17ada10
6 changed files with 2184 additions and 0 deletions
@@ -0,0 +1,86 @@
"""v1.1 WP-B.1: the long run's `recovered_through_memory_independent` verdict.
The new verdict must never be reported when anything other than memory could
have carried the fact. Each precondition is named when it fails. The existing M04
verdicts keep their meaning exactly.
python -m pytest tests/test_v11_b1_long_run_verdict.py -v
"""
import pytest
from tools import m11_long_run as lr
GOOD = {
"independent_planted_depth": 3,
"planted_turn_outside_history": True,
"absent_from_state": True,
"absent_from_summary": True,
"absent_from_knowledge": True,
"absent_from_later_narration": True,
"memory_covering_planting_carries_fact": True,
"memory_forgotten": False,
"memory_injected": True,
}
def test_every_precondition_and_an_injected_memory_is_the_new_verdict():
assert lr._independent_memory_verdict(GOOD) == "recovered_through_memory_independent"
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
def test_a_failed_precondition_is_named_and_never_a_recovery(name):
assert lr._independent_memory_verdict({**GOOD, name: False}) == f"precondition_failed:{name}"
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
def test_an_unmeasured_precondition_is_unknown_not_a_pass(name):
assert lr._independent_memory_verdict({**GOOD, name: None}) == f"precondition_unknown:{name}"
def test_no_planted_depth_is_unknown():
assert lr._independent_memory_verdict({**GOOD, "independent_planted_depth": None}) == \
"precondition_unknown:planted_depth"
@pytest.mark.parametrize("change, verdict", [
({"memory_covering_planting_carries_fact": False}, "not_recovered:not_created"),
({"memory_forgotten": True}, "not_recovered:evicted"),
({"memory_injected": False}, "not_recovered:not_injected"),
])
def test_the_failing_memory_stage_is_named(change, verdict):
assert lr._independent_memory_verdict({**GOOD, **change}) == verdict
def test_preconditions_are_judged_before_memory():
"""A carried fact disqualifies the run even when memory also failed."""
both = {**GOOD, "absent_from_state": False, "memory_covering_planting_carries_fact": False}
assert lr._independent_memory_verdict(both) == "precondition_failed:absent_from_state"
def test_the_fact_is_matched_as_whole_words():
assert lr._mentions_fact("She hid the amber Sundial.")
assert lr._mentions_fact("a cracked TEAPOT on the shelf")
assert not lr._mentions_fact("teapots") # a different word, not the fact's
assert not lr._mentions_fact("the sun dialled down")
def test_the_m04_verdicts_are_unchanged():
base = {"planted_turn_in_history_window": False, "in_memories_section": False,
"in_summary_section": False, "in_state_section": False}
assert lr._m04_verdict(base) == "not_recovered"
assert lr._m04_verdict({**base, "in_state_section": True}) == "recovered_through_state_only"
assert lr._m04_verdict({**base, "in_memories_section": True}) == \
"recovered_through_memory_or_summary"
assert lr._m04_verdict({**base, "planted_turn_in_history_window": True}) == \
"precondition_not_met"
def test_the_independent_fact_is_not_in_any_imported_knowledge_file():
for text in (lr.CANON_MD, lr.REFERENCE_MD, lr.INSPIRATION_MD, *lr.BEATS):
assert not lr._mentions_fact(text)
def test_the_planting_text_and_recall_carry_the_fact():
assert lr._mentions_fact(lr.INDEPENDENT_FACT_TEXT)
assert lr._mentions_fact(lr.INDEPENDENT_RECALL_TEXT)
@@ -0,0 +1,308 @@
"""v1.1 WP-B.1: the memory-retention diagnostic, deterministically.
B.1 changes no memory behaviour. These tests prove two things about the
diagnostic in `tools/memory_diagnostic.py`:
1. **It measures what it claims.**
- The fixture keeps the planted fact out of every layer except memory.
- Each stage (created, retained, ranked, injected) is reported from the rows
and the recall turn's own stored context.
- Its ranking agrees with the selection production stored.
2. **What it finds on this tree.** The scenarios run with a best-case summariser,
one that keeps a fact if and only if the fact reached it. Any failure is
therefore the application's mechanism, not a model's writing.
- The criteria the current code does not meet are marked `xfail(strict=True)`,
so B.2 has to flip them deliberately.
- The same file is run unchanged against v1.0.0 for the baseline.
python -m pytest tests/test_v11_b1_memory_diagnostic.py -v
"""
import asyncio
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, limits, memorybank, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import adventures
from tools import memory_diagnostic as md
_results: dict = {}
def scenario(name: str) -> dict:
"""Runs a named scenario once per session and keeps the result."""
if name not in _results:
_results[name] = md.run_scenario(md.SCENARIOS[name])
return _results[name]
# ------------------------------------------------------- fixture preconditions
def test_the_fact_is_planted_early_and_recalled_past_depth_one_hundred():
result = scenario("independent_default")
assert result["plant_depth"] is not None and result["plant_depth"] <= 3
assert result["recall_depth"] >= 100
@pytest.mark.parametrize("check", ["state_document", "state_snapshots", "later_narration",
"summary", "knowledge", "recent_history", "state_section"])
def test_no_layer_but_memory_carries_the_fact(check):
"""A test where another layer carries F is not evidence about memory."""
isolation = scenario("independent_default")["isolation"]
assert isolation["checks"][check]["ok"], isolation["checks"][check]
assert isolation["ok"]
def test_the_isolation_check_fails_when_another_layer_carries_the_fact():
"""The negative control for the precondition itself: a state fact naming F."""
fact = md.FACT_F
with SessionLocal() as db:
Base.metadata.create_all(bind=engine)
try:
user = models.User(is_guest=False, email="b1-iso@example.com")
db.add(user)
db.flush()
adventure = models.Adventure(user_id=user.id, title="iso")
adventure.narrative_state = {"facts": [{"id": "x", "predicate": "hidden",
"value": "the amber sundial is in the teapot"}]}
db.add(adventure)
db.commit()
result = md.isolation(db, adventure, fact, 1)
assert result["ok"] is False
assert result["checks"]["state_document"]["ok"] is False
finally:
db.close()
Base.metadata.drop_all(bind=engine)
# ------------------------------------------------------------------- stages
def test_creation_is_reported_with_the_covering_memory_and_what_the_summariser_saw():
created = scenario("independent_default")["diagnosis"]["created"]
assert created["yes"] is True
assert created["source_start"] <= scenario("independent_default")["plant_depth"] <= created["source_end"]
assert md.FACT_F.carried_by(created["memory_text"])
covering = [c for c in created["covering_memories"] if c["memory_id"] == created["memory_id"]]
assert covering and covering[0]["fact_in_block"] and covering[0]["fact_in_summariser_excerpt"]
def test_retention_is_reported_with_the_bank_and_its_eviction_order():
retained = scenario("independent_default")["diagnosis"]["retained"]
assert retained["yes"] is True and retained["forgotten"] is False
assert retained["on_active_lineage"] is True
assert retained["active_memories"] <= retained["memory_bank_capacity"]
assert retained["eviction_position"] is not None
def test_ranking_is_production_ranking_and_agrees_with_the_stored_selection():
ranked = scenario("independent_default")["diagnosis"]["ranked"]
assert ranked["replica_matches_stored_selection"] is True
assert ranked["lexical_score"] is None # memory ranking has no lexical term
assert ranked["top_k_cutoff"] == 5
assert ranked["yes"] is True and ranked["selected"] is True
assert 1 <= ranked["rank"] <= ranked["top_k_cutoff"]
# The production query is the newest four actions, cut to 600 tokens, and the
# one-line question is diluted by the narration around it.
variants = scenario("independent_default")["ranking_variants"]
assert ranked["semantic_score"] < variants["direct"]["similarity"]
def test_injection_is_read_from_the_recall_turns_own_context():
diagnosis = scenario("independent_default")["diagnosis"]
assert diagnosis["injected"]["yes"] is True
assert diagnosis["injected"]["context_component"] == md.MEMORIES_LABEL
assert diagnosis["injected"]["token_count"] > 0
assert diagnosis["verdict"] == "injected"
def test_ranking_variants_direct_paraphrase_and_unrelated():
variants = scenario("independent_default")["ranking_variants"]
assert variants["direct"]["rank"] == 1 and variants["direct"]["selected"]
assert variants["paraphrase"]["rank"] == 1 and variants["paraphrase"]["selected"]
assert (variants["direct"]["similarity"] > variants["paraphrase"]["similarity"]
> 5 * variants["unrelated"]["similarity"])
def test_retrieval_fills_top_k_whatever_the_similarity():
"""Diagnosis: there is no relevance floor. With more memories than
`memory_top_k`, an unrelated query still selects five, and the early fact
rides along at a similarity near zero."""
variants = scenario("independent_default")["ranking_variants"]
assert variants["unrelated"]["similarity"] < 0.1
assert variants["unrelated"]["selected"] is True
# ---------------------------------------------------------- capacity/eviction
def test_past_capacity_the_early_memory_is_evicted_and_the_stage_says_so():
"""Diagnosis, not a requirement: what the current eviction rule does to F."""
result = scenario("past_capacity")
assert result["diagnosis"]["created"]["yes"] is True
assert result["diagnosis"]["verdict"] == "created_but_evicted"
eviction = result["eviction"]
assert eviction["f_evicted_at_turn"] is not None
# It was retrieved while the bank was small, stopped being retrieved once
# recent narration filled the top-k, and was then the least recently used.
assert eviction["f_use_count_when_evicted"] > 0
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
assert eviction["f_memory_was_first_evicted"] is True
def test_at_a_lower_top_k_the_early_memory_ages_out_after_it_stops_being_retrieved():
"""Diagnosis with most of the bank unretrieved on any turn, nearer the
shipped 5-in-80 ratio. F is not simply the oldest row: it is evicted some
turns after recent narration stopped pulling it into the top-k, which is
what ordering by last use does to a fact nothing recent mentions."""
result = scenario("past_capacity_low_top_k")
eviction = result["eviction"]
assert result["diagnosis"]["created"]["yes"] is True
assert result["diagnosis"]["verdict"] == "created_but_evicted"
assert eviction["f_use_count_when_evicted"] > 0
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
assert eviction["first_eviction_turn"] <= eviction["f_evicted_at_turn"]
assert eviction["created_and_evicted_same_turn"] == []
def test_no_memory_is_evicted_by_the_same_pass_that_created_it():
"""The frozen-bank regression the current rule fixed, still holding."""
for name in ("past_capacity", "past_capacity_pinned"):
assert scenario(name)["eviction"]["created_and_evicted_same_turn"] == []
def test_a_pinned_memory_survives_capacity():
eviction = scenario("past_capacity_pinned")["eviction"]
assert eviction["pinned_memory_id"] is not None
assert eviction["pinned_memory_forgotten"] is False
@pytest.mark.xfail(strict=True, reason=(
"WP-B.1 diagnosis on this tree: past memory_bank_capacity the planting-era "
"memory is evicted first, because it was never retrieved and eviction orders "
"by last use, then creation. B.2 must flip this deliberately."))
def test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity():
assert scenario("past_capacity")["diagnosis"]["verdict"] == "injected"
# ---------------------------------------------------------- creation window
def test_a_fact_early_in_a_long_block_never_reaches_the_summariser():
result = scenario("long_block_fact_early")
created = result["diagnosis"]["created"]
covering = created["covering_memories"]
assert covering, "the long block must have been summarised"
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
assert covering[0]["fact_in_block"] is True
assert covering[0]["fact_in_summariser_excerpt"] is False
assert result["diagnosis"]["verdict"] == "not_created"
def test_the_same_fact_late_in_the_same_sized_block_does():
result = scenario("long_block_fact_late")
covering = result["diagnosis"]["created"]["covering_memories"]
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
assert covering[0]["fact_in_summariser_excerpt"] is True
assert result["diagnosis"]["created"]["yes"] is True
@pytest.mark.xfail(strict=True, reason=(
"WP-B.1 diagnosis on this tree: the summariser reads only the last "
f"{memorybank.MEMORY_EXCERPT_TOKENS} tokens of a block, so a fact early in a "
"long block is never seen. B.2 must flip this deliberately."))
def test_acceptance_a_fact_early_in_a_long_block_is_remembered():
assert scenario("long_block_fact_early")["diagnosis"]["created"]["yes"] is True
# ------------------------------------------------------- lineage control (G)
def test_an_abandoned_lines_memory_is_stored_but_never_eligible_or_injected():
g = scenario("lineage_control")["lineage_control"]
assert g["memory_ids"], "G's memory must exist on line A before it is abandoned"
assert sorted(g["stored"]) == sorted(g["memory_ids"])
assert g["eligible_on_active_line"] == []
assert g["g_text_ever_in_used_memories"] is False
# Any turn that did name G's memory was on line A, before the divergence.
assert g["eligible_after_returning_to_line_a"] == g["memory_ids"]
def test_the_lineage_scenario_still_diagnoses_f_on_the_active_line():
result = scenario("lineage_control")
assert result["isolation"]["ok"], result["isolation"]
assert result["diagnosis"]["verdict"] == "injected"
# ----------------------------------------------------- authority control
@pytest.fixture()
def authority_client(monkeypatch):
embedder = md.ConceptEmbedder()
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
with SessionLocal() as db:
user = models.User(is_guest=False, email="b1-auth@example.com")
db.add(user)
db.flush()
db.add(models.Settings(user_id=user.id, model="script",
endpoint_url="http://127.0.0.1:9/v1",
embedding_model="concept-embed", memory_top_k=5))
adventure = models.Adventure(user_id=user.id, title="auth", memory_bank_enabled=True,
auto_summarize=True)
db.add(adventure)
db.flush()
db.add(models.Action(adventure_id=adventure.id, type="start", text="The tavern at dusk."))
db.commit()
adv, user_id = adventure.id, user.id
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", md.ScriptNarrator)
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: embedder)
monkeypatch.setattr(memorybank, "summary_provider", lambda s: md.BestCaseSummariser())
monkeypatch.setattr(memorybank, "schedule_post_turn", lambda a: None)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id))
client = TestClient(app)
client.adv = adv
try:
yield client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
memorybank._vector_cache.clear()
Base.metadata.drop_all(bind=engine)
def test_a_memory_that_contradicts_state_loses_and_changes_nothing(authority_client):
client, adv = authority_client, authority_client.adv
corrected = client.post(f"/api/adventures/{adv}/state/corrections", json={"events": [
{"type": "add_fact", "predicate": "the tavern lamp is lit", "fact_id": "lamp-lit"}]})
assert corrected.status_code in (200, 201), corrected.text[:300]
made = client.post(f"/api/adventures/{adv}/memories",
json={"text": "The tavern lamp was never lit that night."})
assert made.status_code == 201, made.text[:300]
client.patch(f"/api/adventures/{adv}/memories/{made.json()['id']}", json={"pinned": True})
asyncio.run(memorybank.run_post_turn(adv)) # embed it
before = client.get(f"/api/adventures/{adv}/state").json()["document"]
md.ScriptNarrator.next_reply = 'The fire crackles.\n```state\n{"events": []}\n```'
played = client.post(f"/api/adventures/{adv}/actions",
json={"type": "do", "text": "I look at the lamp."})
assert played.status_code == 200 and '"type": "error"' not in played.text
after = client.get(f"/api/adventures/{adv}/state").json()["document"]
assert after == before # retrieval mutated no state
with SessionLocal() as db:
action = (db.query(models.Action).filter_by(adventure_id=adv, type="ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
snapshot = action.context_snapshot
state_text = md._section(snapshot, md.STATE_LABEL)
memory_text = md._section(snapshot, md.MEMORIES_LABEL)
assert "the tavern lamp is lit" in state_text
assert "never lit" in memory_text
assert memory_text.startswith("Memories from earlier in the story")
labels = [s["label"] for s in snapshot["sections"]]
# State is read last of the live sections: it settles the conflict.
assert labels.index(md.STATE_LABEL) > labels.index(md.MEMORIES_LABEL)
+206
View File
@@ -190,6 +190,31 @@ CLUE_FACT = {
"fact_id": "silver-key-opens-crypt", "fact_id": "silver-key-opens-crypt",
} }
#: v1.1 WP-B.1: a second planted fact, established in the **story only**.
#:
#: The M04 clue above is planted as accepted state, and memories are written
#: from story text, so no memory could ever carry it on its own. That is why
#: every M04 recovery so far ran through state. This fact is told to the reader
#: in narration and never corrected into state, so memory is the only layer that
#: is meant to carry it. `--independent-fact` plants it and reports
#: `recovered_through_memory_independent` only when every other layer is proven
#: not to carry it. The words are copied from `tools/memory_diagnostic.FACT_F`,
#: for the reason `HISTORY_LABELS` is copied.
INDEPENDENT_FACT_TEXT = ("I watch Mara slip the amber sundial inside the cracked teapot on "
"the tavern's top shelf, and she makes me promise to tell no one.")
INDEPENDENT_FACT_TERMS = ("sundial", "teapot")
INDEPENDENT_RECALL_TEXT = "I ask Mara quietly where she hid the amber sundial."
#: How far past the planting turn its memory block can reach. Narration inside
#: that block may repeat the fact; narration after it may not.
INDEPENDENT_BLOCK_SLACK = 6
INDEPENDENT_PRECONDITIONS = (
"planted_turn_outside_history",
"absent_from_state",
"absent_from_summary",
"absent_from_knowledge",
"absent_from_later_narration",
)
CANON = [ CANON = [
"The dead do not return. No rite, relic or bargain has ever returned anyone.", "The dead do not return. No rite, relic or bargain has ever returned anyone.",
"The abbey crypt has been sealed since the founding.", "The abbey crypt has been sealed since the founding.",
@@ -365,6 +390,14 @@ class Run:
#: The depth of the player turn that planted the clue. M04's #: The depth of the player turn that planted the clue. M04's
#: precondition is that this turn has left the history window. #: precondition is that this turn has left the history window.
self.planted_depth: int | None = None self.planted_depth: int | None = None
#: v1.1 WP-B.1, with --independent-fact: where the story-only fact was
#: planted, and the accepted-turn count at which each isolation
#: precondition first failed.
self.independent_fact = False
self.independent_depth: int | None = None
self.independent_violations: dict[str, int] = {}
self.last_done: dict = {}
self.last_report: dict = {}
# ------------------------------------------------------------ recording # ------------------------------------------------------------ recording
@@ -393,6 +426,8 @@ class Run:
"turns_target": self.turns_target, "turns_target": self.turns_target,
"log_offset": self.log_offset, "log_offset": self.log_offset,
"planted_depth": self.planted_depth, "planted_depth": self.planted_depth,
"independent_depth": self.independent_depth,
"independent_violations": self.independent_violations,
"written": datetime.now().isoformat(timespec="seconds"), "written": datetime.now().isoformat(timespec="seconds"),
} }
tmp = self.out / (RESUME_FILE + ".tmp") tmp = self.out / (RESUME_FILE + ".tmp")
@@ -414,6 +449,8 @@ class Run:
self.elapsed_before = prior.get("elapsed_seconds", 0) self.elapsed_before = prior.get("elapsed_seconds", 0)
self.log_offset = prior.get("log_offset", 0) self.log_offset = prior.get("log_offset", 0)
self.planted_depth = prior.get("planted_depth") self.planted_depth = prior.get("planted_depth")
self.independent_depth = prior.get("independent_depth")
self.independent_violations = dict(prior.get("independent_violations") or {})
self.resumed = True self.resumed = True
def reattach(self) -> None: def reattach(self) -> None:
@@ -570,15 +607,44 @@ class Run:
"observed_margin": accounting.get("observed_margin"), "observed_margin": accounting.get("observed_margin"),
"safety_reserve": accounting.get("safety_reserve"), "safety_reserve": accounting.get("safety_reserve"),
}) })
self.last_done = done
if self.independent_fact and self.independent_depth is not None:
self._check_independent_isolation(done, sample)
self.note("turn", text=text, seconds=round(seconds, 1), **sample) self.note("turn", text=text, seconds=round(seconds, 1), **sample)
return {"accepted": True, "seconds": seconds, **sample} return {"accepted": True, "seconds": seconds, **sample}
def _check_independent_isolation(self, done: dict, sample: dict) -> None:
"""v1.1 WP-B.1: does anything but memory carry the story-only fact yet?
Checked on every accepted turn, so a run knows the first turn at which
the experiment stopped being about memory, instead of finding out at
recall. Each precondition records only its first failure.
"""
depth = sample.get("total_actions", 0) - 1
text = (done.get("action") or {}).get("text") or ""
found = {}
if depth > self.independent_depth + INDEPENDENT_BLOCK_SLACK and _mentions_fact(text):
found["absent_from_later_narration"] = f"narration at depth {depth}"
document = self.state().get("document") or {}
if _mentions_fact(json.dumps(document)):
found["absent_from_state"] = "the narrative state names the fact"
summary = next((sec.get("text", "") for sec in (self.last_report.get("sections") or [])
if sec.get("label") == SUMMARY_LABEL), "")
if _mentions_fact(summary):
found["absent_from_summary"] = "the active summary names the fact"
for name, detail in found.items():
if name not in self.independent_violations:
self.independent_violations[name] = self.accepted
self.note("independent_precondition_failed", precondition=name, detail=detail)
sample["independent_violations"] = dict(self.independent_violations)
def count_actions(self) -> int: def count_actions(self) -> int:
return self.server.call("GET", f"/adventures/{self.adv}/actions?limit=1")["total"] return self.server.call("GET", f"/adventures/{self.adv}/actions?limit=1")["total"]
def measure(self) -> dict: def measure(self) -> dict:
"""M03's numbers, read from the prompt the app would send right now.""" """M03's numbers, read from the prompt the app would send right now."""
report = self.server.call("GET", f"/adventures/{self.adv}/context") report = self.server.call("GET", f"/adventures/{self.adv}/context")
self.last_report = report
tokens = report["tokens"] tokens = report["tokens"]
sections = {s["label"]: s["tokens"] for s in report["sections"]} sections = {s["label"]: s["tokens"] for s in report["sections"]}
window = report.get("window") or {} window = report.get("window") or {}
@@ -714,6 +780,10 @@ def main() -> int:
"--max-consecutive-failures", type=int, "--max-consecutive-failures", type=int,
default=DEFAULT_MAX_CONSECUTIVE_FAILURES, default=DEFAULT_MAX_CONSECUTIVE_FAILURES,
help="stop and write the evidence after this many unaccepted turns") help="stop and write the evidence after this many unaccepted turns")
parser.add_argument(
"--independent-fact", action="store_true",
help=("v1.1 WP-B.1: also plant a story-only fact at depth 3 and report "
"whether memory alone recovers it"))
args = parser.parse_args() args = parser.parse_args()
if not (ENDPOINT and MODEL and EMBED_MODEL): if not (ENDPOINT and MODEL and EMBED_MODEL):
@@ -746,6 +816,7 @@ def main() -> int:
server.start() server.start()
run = Run(server, out, turns_target=args.turns, run = Run(server, out, turns_target=args.turns,
turn_timeout=args.turn_timeout) turn_timeout=args.turn_timeout)
run.independent_fact = args.independent_fact
if prior: if prior:
run.adopt(prior) run.adopt(prior)
@@ -782,6 +853,18 @@ def main() -> int:
"the planted clue is not in accepted state, so M04 cannot " "the planted clue is not in accepted state, so M04 cannot "
"be measured from this run. Stopping before the campaign " "be measured from this run. Stopping before the campaign "
"starts rather than reporting a recall failure later.") "starts rather than reporting a recall failure later.")
if args.independent_fact:
# v1.1 WP-B.1: the story-only fact, told in the next turn and
# never corrected into state. Depth 3: the opening, the clue turn
# and its reply come first.
if any(_mentions_fact(md) for md in (CANON_MD, REFERENCE_MD, INSPIRATION_MD)):
raise SystemExit("the imported knowledge names the independent fact")
planting_f = run.turn(INDEPENDENT_FACT_TEXT)
if not planting_f.get("accepted"):
raise SystemExit("the turn that plants the independent fact was not accepted")
run.independent_depth = planting_f["total_actions"] - 2
run.note("independent_fact_planted", depth=run.independent_depth,
terms=list(INDEPENDENT_FACT_TERMS))
# The first checkpoint, and the point from which --resume works: the # The first checkpoint, and the point from which --resume works: the
# campaign exists and its clue is planted. # campaign exists and its clue is planted.
run.save_resume() run.save_resume()
@@ -846,10 +929,15 @@ def main() -> int:
# Skipped on an aborted run: it asks the narrator a question, and the # Skipped on an aborted run: it asks the narrator a question, and the
# reason the run stopped is that the narrator does not answer. # reason the run stopped is that the narrator does not answer.
recall = None recall = None
independent = None
if aborted is None: if aborted is None:
run.note("recall_begin") run.note("recall_begin")
recall = _recall(run) recall = _recall(run)
(out / "recall.json").write_text(json.dumps(recall, indent=2)) (out / "recall.json").write_text(json.dumps(recall, indent=2))
if args.independent_fact and run.independent_depth is not None:
independent = _independent_recall(run, out / "campaign.db")
(out / "recall-independent.json").write_text(json.dumps(independent, indent=2))
run.note("independent_recall", verdict=independent["verdict"])
# ---- Export whatever exists, for the recovery evidence. ---- # ---- Export whatever exists, for the recovery evidence. ----
# Attempted even for an aborted run: the recovery check and the storage # Attempted even for an aborted run: the recovery check and the storage
@@ -897,6 +985,7 @@ def main() -> int:
"elapsed_seconds": run.elapsed(), "elapsed_seconds": run.elapsed(),
"turn_timeout_seconds": args.turn_timeout, "turn_timeout_seconds": args.turn_timeout,
"recall": recall, "recall": recall,
"independent_recall": independent,
"final_state": _or_none(lambda: run.state()["document"]), "final_state": _or_none(lambda: run.state()["document"]),
"final_measurement": _or_none(run.measure), "final_measurement": _or_none(run.measure),
"db_bytes": db_path.stat().st_size, "db_bytes": db_path.stat().st_size,
@@ -1224,6 +1313,123 @@ def _m04_verdict(recall: dict) -> str:
return "not_recovered" return "not_recovered"
def _mentions_fact(text: str | None) -> bool:
"""v1.1 WP-B.1: whether `text` names the independent fact, as a whole word."""
low = (text or "").lower()
return any(re.search(rf"(?<![a-z]){term}(?![a-z])", low) for term in INDEPENDENT_FACT_TERMS)
def _independent_memory_verdict(check: dict) -> str:
"""v1.1 WP-B.1: whether memory alone recovered the story-only fact.
`recovered_through_memory_independent` requires every precondition, so no
other layer could have carried the fact. It also requires that a memory
covering the planting turn carries the fact and was injected into the recall
turn. A failed precondition is named and is never a recovery, and the M04
verdicts above are untouched.
"""
if check.get("independent_planted_depth") is None:
return "precondition_unknown:planted_depth"
for name in INDEPENDENT_PRECONDITIONS:
value = check.get(name)
if value is None:
return f"precondition_unknown:{name}"
if not value:
return f"precondition_failed:{name}"
if not check.get("memory_covering_planting_carries_fact"):
return "not_recovered:not_created"
if check.get("memory_forgotten"):
return "not_recovered:evicted"
if not check.get("memory_injected"):
return "not_recovered:not_injected"
return "recovered_through_memory_independent"
def _independent_recall(run: "Run", db_path: Path) -> dict:
"""v1.1 WP-B.1: ask for the story-only fact, and find out which layer answered.
The prompt-level facts come from the recall turn's own stored context. The
memory rows come from the campaign database, read-only. Ranking is not
recomputed here, because that needs the embedding model;
`tools/v11_b1_memory.py diagnose` does it afterwards against a copy of the
database.
"""
import sqlite3
import zlib
result = run.turn(INDEPENDENT_RECALL_TEXT)
action_id = (run.last_done.get("action") or {}).get("id")
snapshot = (run.server.call("GET", f"/adventures/{run.adv}/actions/{action_id}/context")
if result.get("accepted") and action_id else {}) or {}
sections = {}
for sec in snapshot.get("sections") or []:
sections.setdefault(sec.get("label"), []).append(sec.get("text", ""))
text_of = {label: "\n".join(parts) for label, parts in sections.items()}
floor = (snapshot.get("history") or {}).get("floor_depth")
depth = run.independent_depth
used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
covering = []
connection = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True)
try:
rows = connection.execute(
"SELECT id, text, source_start, source_end, forgotten, pinned, use_count, "
"branch_id, depth FROM memories WHERE adventure_id = ? AND source_start <= ? "
"AND source_end >= ? ORDER BY id", (run.adv, depth, depth)).fetchall()
blob = connection.execute(
"SELECT context_snapshot FROM actions WHERE id = ?", (action_id or -1,)).fetchone()
finally:
connection.close()
for row in rows:
memory_id, text, start, end, forgotten, pinned, use_count, branch_id, node_depth = row
covering.append({
"memory_id": memory_id, "text": text, "source_start": start, "source_end": end,
"forgotten": bool(forgotten), "pinned": bool(pinned), "use_count": use_count,
"branch_id": branch_id, "depth": node_depth,
"carries_fact": all(re.search(rf"(?<![a-z]){t}(?![a-z])", (text or "").lower())
for t in INDEPENDENT_FACT_TERMS),
"injected": memory_id in used,
})
carrying = [c for c in covering if c["carries_fact"]]
best = next((c for c in carrying if c["injected"]), carrying[0] if carrying else None)
stored_snapshot_readable = blob is not None and blob[0] is not None
if stored_snapshot_readable:
try:
json.loads(zlib.decompress(blob[0]))
except Exception: # noqa: BLE001
stored_snapshot_readable = False
document = run.state().get("document") or {}
violations = dict(run.independent_violations)
check = {
"independent_planted_depth": depth,
"recall_accepted": bool(result.get("accepted")),
"history_floor_depth": floor,
"planted_turn_outside_history": (None if not snapshot else
floor is not None and depth < floor),
"absent_from_state": ("absent_from_state" not in violations
and not _mentions_fact(json.dumps(document))
and not _mentions_fact(text_of.get(STATE_LABEL))),
"absent_from_summary": ("absent_from_summary" not in violations
and not _mentions_fact(text_of.get(SUMMARY_LABEL))),
"absent_from_knowledge": not any(
_mentions_fact(text_of.get(label)) for label in IMPORTED_KNOWLEDGE_LABELS),
"absent_from_later_narration": "absent_from_later_narration" not in violations,
"violations_first_turn": violations,
"covering_memories": covering,
"memory_covering_planting_carries_fact": bool(carrying),
"memory_forgotten": bool(best and best["forgotten"]),
"memory_injected": bool(best and best["injected"]),
"memory_text_in_memories_section": bool(
best and best["text"] and best["text"] in (text_of.get(MEMORIES_LABEL) or "")),
"memory_ids_used": used,
"recall_action_id": action_id,
"stored_snapshot_readable": stored_snapshot_readable,
}
check["verdict"] = _independent_memory_verdict(check)
return check
#: Signs the application stored protocol as story. The first is a state-section #: Signs the application stored protocol as story. The first is a state-section
#: heading with an indented entry under it, in any markdown, because #: heading with an indented entry under it, in any markdown, because
#: `## Established:` got past a plain substring match and the count read 1 #: `## Established:` got past a plain substring match and the count read 1
+839
View File
@@ -0,0 +1,839 @@
"""v1.1 WP-B.1: where an early story fact is lost on its way to the narrator.
One planted fact **F** has four stages to survive before the narrator can use
it from memory, and this module reports each one separately:
created a memory whose `source_start`..`source_end` covers the planting
depth carries F
retained that memory is not `forgotten`
ranked it is eligible on the active lineage and embedded, and where it
scores for the recall query against `memory_top_k`
injected the recall turn's own stored `memories.used` names it, and its text
is in that turn's `used_memories` section
A fact is only evidence about memory if memory is the **only** thing carrying it.
`isolation()` checks every other layer: the authoritative document, per-node state
snapshots, the active summary, imported knowledge, the narration after the
planting block, and the recent-history window. A run where any of those carries F
is reported as a failed precondition, never as a memory result.
**Nothing here changes behaviour.**
- It reads rows.
- It reuses production's own pure helpers (`memorybank._drop_redundant`,
`memorybank.classify_authority`, `vectors.cosine`, `lineage.path_of`), so its
ranking is production's ranking, not a second opinion.
- It checks itself against what the recall turn actually recorded.
- The only computed fields are ephemeral report data. No column or table is
added.
The deterministic stubs at the bottom stand in for the models when a test needs a
fixed answer. **Read what they model before reading any result they produce:**
- `BestCaseSummariser` keeps F if and only if F is in the excerpt it is given.
It is the ideal summariser, so a creation failure under it is the
application's, not the model's.
- `ConceptEmbedder` maps words to a small concept table, so that "the brass dial
that tells the hour" lands near "sundial". It models what an embedding is
supposed to do. It says nothing about how well `nomic-embed-text` does it,
which is what the real-model run is for.
"""
from __future__ import annotations
import hashlib
import json
import math
import re
from dataclasses import dataclass, field
from sqlalchemy import select
from app import memorybank, models, summaries, vectors
from app.context import builder, history, lineage
from app.knowledge import classes as knowledge_classes
VERDICTS = (
"not_created",
"created_but_evicted",
"retained_but_not_ranked",
"ranked_but_not_selected",
"selected_but_not_injected",
"injected",
)
#: Section labels in a stored context snapshot. Copied from the builder's
#: vocabulary so a renamed section fails loudly here.
HISTORY_LABELS = ("history", "recent_history")
SUMMARY_LABEL = "story_summary"
MEMORIES_LABEL = "used_memories"
STATE_LABEL = "narrative_state"
KNOWLEDGE_LABELS = (
knowledge_classes.SECTION_CANON,
knowledge_classes.SECTION_REFERENCE,
knowledge_classes.SECTION_INSPIRATION,
)
@dataclass(frozen=True)
class Fact:
"""A planted fact, and how to recognise it in a text.
`carry_groups`: a text carries the fact when every group matches, where a
group matches when any one of its terms appears as a whole word. A memory has
to name both the thing and where it is to carry "where the thing is".
`leak_terms`: any one of these in another layer means that layer carries the
fact. This is deliberately looser than `carry_groups`. For isolation, a
mention is enough to disqualify.
"""
fact_id: str
sentence: str
carry_groups: tuple[tuple[str, ...], ...]
leak_terms: tuple[str, ...]
def carried_by(self, text: str | None) -> bool:
low = (text or "").lower()
return all(any(_has_word(low, term) for term in group) for group in self.carry_groups)
def mentioned_by(self, text: str | None) -> bool:
low = (text or "").lower()
return any(_has_word(low, term) for term in self.leak_terms)
def _has_word(low: str, term: str) -> bool:
return re.search(rf"(?<![a-z]){re.escape(term.lower())}(?![a-z])", low) is not None
#: The fixture's planted fact. Chosen to be natural in a tavern scene and absent
#: from every existing fixture: no "sundial" or "teapot" appears anywhere in the
#: Westhaven campaign, its knowledge files or its beats.
FACT_F = Fact(
fact_id="F-amber-sundial",
sentence="Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf.",
carry_groups=(("sundial",), ("teapot",)),
leak_terms=("sundial", "teapot"),
)
#: The abandoned-line control fact.
FACT_G = Fact(
fact_id="G-iron-weathervane",
sentence="Edrin buried the iron weathervane beneath the mill's broken waterwheel.",
carry_groups=(("weathervane",), ("waterwheel",)),
leak_terms=("weathervane", "waterwheel"),
)
# ------------------------------------------------------------------ reading
def _lineage_actions(db, adventure):
path = lineage.path_of(db, adventure)
return (
db.query(models.Action)
.filter(models.Action.adventure_id == adventure.id, path.clause(models.Action))
.order_by(models.Action.depth, models.Action.id)
.all()
)
def covering_memories(db, adventure, depth: int, *, any_branch: bool = False):
"""Memories whose source range covers `depth`, oldest first."""
query = select(models.Memory).where(
models.Memory.adventure_id == adventure.id,
models.Memory.source_start <= depth,
models.Memory.source_end >= depth,
)
if not any_branch:
query = query.where(lineage.path_of(db, adventure).clause(models.Memory))
return db.execute(query.order_by(models.Memory.id)).scalars().all()
def planting_block_end(db, adventure, plant_depth: int) -> int:
"""The last depth of the memory block holding the planted turn.
Taken from the memory that covers it where one exists. Before one exists it
is the furthest a block could reach, so a later-narration check never counts
a turn inside the planting block as a repetition.
"""
rows = covering_memories(db, adventure, plant_depth)
if rows:
return max(row.source_end for row in rows)
return plant_depth + memorybank.MEMORY_INTERVAL
# ---------------------------------------------------------------- isolation
def isolation(db, adventure, fact: Fact, plant_depth: int, *,
recall_snapshot: dict | None = None,
recall_depth: int | None = None) -> dict:
"""Every layer other than memory that could carry F, checked.
Returns `{check: {"ok": bool, "detail": str}}` and `ok` over all of them.
With `recall_snapshot`, the recall turn's stored context, the prompt-level
checks (history window, summary section, knowledge sections) are made
against what the narrator was actually given.
"""
checks: dict[str, dict] = {}
document = adventure.narrative_state or {}
hits = [key for key in ("entities", "facts", "relationships", "threads", "scene",
"possessions")
if fact.mentioned_by(json.dumps(document.get(key), default=str))]
checks["state_document"] = {
"ok": not hits and not fact.mentioned_by(json.dumps(document, default=str)),
"detail": f"mentioned in {hits}" if hits else "absent",
}
snapshot_hits = []
later_hits = []
block_end = planting_block_end(db, adventure, plant_depth)
for action in _lineage_actions(db, adventure):
if fact.mentioned_by(json.dumps(action.narrative_state_after, default=str)):
snapshot_hits.append(action.depth)
if (action.type == "ai" and action.depth is not None and action.depth > block_end
and (recall_depth is None or action.depth < recall_depth)
and fact.mentioned_by(action.text)):
later_hits.append(action.depth)
checks["state_snapshots"] = {
"ok": not snapshot_hits,
"detail": f"mentioned in snapshots at depths {snapshot_hits[:10]}" if snapshot_hits
else "absent from every node's narrative_state_after on the active lineage",
}
checks["later_narration"] = {
"ok": not later_hits,
"detail": (f"narration after the planting block (ends at depth {block_end}) "
f"mentions the fact at depths {later_hits[:10]}") if later_hits
else f"no narrator turn after depth {block_end} mentions the fact",
}
active = summaries.current(db, adventure)
summary_text = active.text if active is not None else ""
if recall_snapshot is not None:
summary_text += "\n" + _section(recall_snapshot, SUMMARY_LABEL)
checks["summary"] = {
"ok": not fact.mentioned_by(summary_text),
"detail": "the active summary mentions the fact" if fact.mentioned_by(summary_text)
else ("absent from the active summary" if active is not None else "no summary yet"),
}
sources = db.execute(
select(models.KnowledgeSource.content).where(
models.KnowledgeSource.adventure_id == adventure.id)
).scalars().all()
knowledge_text = "\n".join(s or "" for s in sources)
if recall_snapshot is not None:
knowledge_text += "\n" + "\n".join(_section(recall_snapshot, l) for l in KNOWLEDGE_LABELS)
checks["knowledge"] = {
"ok": not fact.mentioned_by(knowledge_text),
"detail": "imported knowledge mentions the fact" if fact.mentioned_by(knowledge_text)
else f"absent from {len(sources)} imported source(s)",
}
if recall_snapshot is not None:
hist = recall_snapshot.get("history") or {}
floor = hist.get("floor_depth")
history_text = "\n".join(_section(recall_snapshot, l) for l in HISTORY_LABELS)
outside = floor is not None and plant_depth < floor
checks["recent_history"] = {
"ok": outside and not fact.carried_by(history_text),
"detail": (f"history window starts at depth {floor}; planted at {plant_depth}; "
f"fact text in history sections: {fact.carried_by(history_text)}"),
}
checks["state_section"] = {
"ok": not fact.mentioned_by(_section(recall_snapshot, STATE_LABEL)),
"detail": "the recall prompt's narrative_state section "
+ ("mentions the fact" if fact.mentioned_by(_section(recall_snapshot, STATE_LABEL))
else "does not mention the fact"),
}
return {"ok": all(c["ok"] for c in checks.values()), "checks": checks}
def _section(snapshot: dict, label: str) -> str:
return "\n".join(s.get("text", "") for s in (snapshot.get("sections") or [])
if s.get("label") == label)
# ------------------------------------------------------------------- stages
async def rank_bank(db, adventure, settings, query: str, embed) -> dict:
"""Production's ranking, recomputed for `query`, for every eligible memory.
The same catalogue clause, the same cosine, the same pin rule, the same
redundancy suppression helper. Returns every scored row, not just the top-k,
because "where did F rank" is the question.
"""
catalogue = db.execute(
select(models.Memory.id, models.Memory.pinned, models.Memory.authority,
models.Memory.embedding_blob).where(
models.Memory.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Memory),
models.Memory.forgotten.is_(False),
models.Memory.embedded.is_(True),
)
).all()
if not catalogue or not query.strip():
return {"query": query, "scored": [], "selected": [], "top_k": settings.memory_top_k}
[query_vec] = await embed([query])
held = {row.id: vectors.unpack(row.embedding_blob) for row in catalogue if row.embedding_blob}
authority_of = {row.id: row.authority for row in catalogue}
scored = sorted(
((vectors.cosine(query_vec, held[row.id]), row.id, row.pinned)
for row in catalogue if row.id in held),
key=lambda r: r[0], reverse=True,
)
top_k = max(1, settings.memory_top_k)
used = [r for r in scored if r[2]]
remaining = max(0, top_k - len(used))
candidates = [r for r in scored if not r[2]]
kept, suppressed = memorybank._drop_redundant(candidates, held, authority_of, remaining)
selected = {r[1] for r in used + kept}
suppressed_by = dict(suppressed)
return {
"query": query,
"top_k": top_k,
"scored": [
{"rank": i + 1, "memory_id": memory_id, "similarity": round(score, 4),
"pinned": pinned, "selected": memory_id in selected,
"suppressed_as_duplicate_of": suppressed_by.get(memory_id)}
for i, (score, memory_id, pinned) in enumerate(scored)
],
"selected": sorted(selected),
}
def production_query(adventure, exclude_action_id: int | None) -> str:
"""The retrieval query a turn used: its newest actions, as `retrieve_memories` builds it."""
recent = history.tail(adventure, memorybank.RETRIEVAL_WINDOW_ACTIONS, exclude_action_id)
return builder.truncate_to_last_tokens(
"\n\n".join(a.text for a in recent), memorybank.RETRIEVAL_WINDOW_TOKENS)
def eviction_order(db, adventure) -> list[int]:
"""The order `_evict_over_capacity` would take unpinned active memories in."""
from sqlalchemy import func
return db.execute(
select(models.Memory.id).where(
models.Memory.adventure_id == adventure.id,
models.Memory.forgotten.is_(False),
models.Memory.pinned.is_(False),
).order_by(func.coalesce(models.Memory.last_used_at, models.Memory.created_at),
models.Memory.use_count)
).scalars().all()
async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
recall_action: models.Action, embed) -> dict:
"""The four stages for `fact`, judged at `recall_action`, the recall turn's AI node.
Ranking is recomputed with the query that turn used, and checked against the
turn's own stored `memories.used`. Injection is read from that snapshot, so
it reports what the narrator was actually given, not a re-run.
"""
snapshot = recall_action.context_snapshot or {}
out: dict = {"fact_id": fact.fact_id, "plant_depth": plant_depth,
"recall_depth": recall_action.depth}
covering = covering_memories(db, adventure, plant_depth)
carrying = [m for m in covering if fact.carried_by(m.text)]
elsewhere = [m for m in db.execute(select(models.Memory).where(
models.Memory.adventure_id == adventure.id)).scalars().all()
if fact.carried_by(m.text) and m not in carrying]
creation_input = []
for memory in covering:
block = memorybank.source_block(db, memory)
raw = "\n\n".join(a.text for a in block)
excerpt = builder.truncate_to_last_tokens(raw, memorybank.MEMORY_EXCERPT_TOKENS)
creation_input.append({
"memory_id": memory.id, "source_start": memory.source_start,
"source_end": memory.source_end, "block_tokens": builder.count_tokens(raw),
"fact_in_block": fact.carried_by(raw),
"fact_in_summariser_excerpt": fact.carried_by(excerpt),
"memory_text": memory.text,
})
memory = carrying[0] if carrying else None
out["created"] = {
"yes": memory is not None,
"memory_id": getattr(memory, "id", None),
"source_start": getattr(memory, "source_start", None),
"source_end": getattr(memory, "source_end", None),
"memory_text": getattr(memory, "text", None),
"covering_memories": creation_input,
"no_covering_memory": not covering,
"carried_by_other_memories": [
{"memory_id": m.id, "source_start": m.source_start, "source_end": m.source_end}
for m in elsewhere],
}
if memory is None:
out["verdict"] = "not_created"
return out
order = eviction_order(db, adventure)
active = db.execute(select(models.Memory.id).where(
models.Memory.adventure_id == adventure.id,
models.Memory.forgotten.is_(False))).scalars().all()
on_lineage = db.execute(select(models.Memory.id).where(
models.Memory.id == memory.id,
lineage.path_of(db, adventure).clause(models.Memory))).scalar() is not None
out["retained"] = {
"yes": not memory.forgotten,
"forgotten": memory.forgotten,
"pinned": memory.pinned,
"embedded": memory.embedded,
"on_active_lineage": on_lineage,
"use_count": memory.use_count,
"last_used_at": str(memory.last_used_at) if memory.last_used_at else None,
"created_at": str(memory.created_at),
"active_memories": len(active),
"memory_bank_capacity": settings.memory_bank_capacity,
"eviction_position": (order.index(memory.id) + 1) if memory.id in order else None,
"reason": ("evicted: marked forgotten by capacity eviction" if memory.forgotten
else "active"),
}
if memory.forgotten:
out["verdict"] = "created_but_evicted"
return out
query = production_query(adventure, recall_action.id)
ranking = await rank_bank(db, adventure, settings, query, embed)
row = next((r for r in ranking["scored"] if r["memory_id"] == memory.id), None)
stored_used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
out["ranked"] = {
"yes": row is not None and row["rank"] <= ranking["top_k"],
"eligible": row is not None,
"lexical_score": None, # memory ranking has no lexical term (CONTEXT-AND-MEMORY §20)
"semantic_score": row["similarity"] if row else None,
"final_score": row["similarity"] if row else None,
"pin_effect": "always selected" if memory.pinned else "none",
"rank": row["rank"] if row else None,
"of": len(ranking["scored"]),
"top_k_cutoff": ranking["top_k"],
"selected": bool(row and row["selected"]),
"suppressed_as_duplicate_of": row["suppressed_as_duplicate_of"] if row else None,
"query": query,
"replica_matches_stored_selection": sorted(stored_used) == ranking["selected"],
}
if row is None or row["rank"] > ranking["top_k"] and not row["selected"]:
out["verdict"] = "retained_but_not_ranked"
return out
if not row["selected"]:
out["verdict"] = "ranked_but_not_selected"
return out
section = _section(snapshot, MEMORIES_LABEL)
injected = memory.id in stored_used and memory.text in section
out["injected"] = {
"yes": injected,
"context_component": MEMORIES_LABEL,
"in_stored_memories_used": memory.id in stored_used,
"text_in_section": memory.text in section,
"token_count": builder.count_tokens(section) if section else 0,
}
out["verdict"] = "injected" if injected else "selected_but_not_injected"
return out
# ---------------------------------------------------------------- the stubs
@dataclass
class BestCaseSummariser:
"""The ideal memory writer: F survives if, and only if, F reached it.
A memory keeps every sentence of the excerpt that carries a planted fact, and
adds one sentence naming the block's own distinct detail so memories differ.
Summary updates never repeat a planted fact, so the summary layer stays out of
the experiment. Every excerpt it was given is kept, for the creation-window
diagnostic.
"""
facts: tuple[Fact, ...] = (FACT_F, FACT_G)
excerpts: list = field(default_factory=list)
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
if "Current story summary:" in user:
return "The travellers kept moving through the country around Westhaven."
excerpt = user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
self.excerpts.append(excerpt)
kept = [s.strip() for s in re.split(r"(?<=[.!?])\s+", excerpt)
if any(f.carried_by(s) for f in self.facts)]
detail = re.findall(r"\bat the ([a-z]+ [a-z]+)\b", excerpt.lower())
tail = f"The travellers spent time at the {detail[-1]}." if detail else \
"The travellers pressed on."
return " ".join(dict.fromkeys(kept + [tail]))
#: Words that mean the same thing to `ConceptEmbedder`. The point is only that a
#: paraphrase lands near the original; the table is the model of that.
CONCEPTS = {
"timepiece": ("sundial", "dial", "hour", "hours", "clock", "timepiece"),
"vessel": ("teapot", "pot", "kettle", "tea", "jar"),
"hid": ("hid", "hide", "hidden", "slipped", "tucked", "put", "stashed"),
"weathervane": ("weathervane", "vane"),
"waterwheel": ("waterwheel", "wheel", "mill"),
}
_WORD_TO_CONCEPT = {w: c for c, words in CONCEPTS.items() for w in words}
DIMENSIONS = 96
@dataclass
class ConceptEmbedder:
"""A deterministic embedding: concepts in fixed dimensions, other words hashed."""
calls: int = 0
async def embed(self, texts):
self.calls += 1
return [self.vector(t) for t in texts]
@staticmethod
def vector(text: str) -> list[float]:
v = [0.0] * DIMENSIONS
v[0] = 0.2 # every text shares a little, as real embeddings do
concept_names = list(CONCEPTS)
for word in re.findall(r"[a-z]+", text.lower()):
concept = _WORD_TO_CONCEPT.get(word)
if concept is not None:
v[1 + concept_names.index(concept)] += 3.0
elif len(word) > 3:
bucket = int(hashlib.sha256(word.encode()).hexdigest(), 16)
v[1 + len(concept_names) + bucket % (DIMENSIONS - 1 - len(concept_names))] += 1.0
norm = math.sqrt(sum(x * x for x in v)) or 1.0
return [x / norm for x in v]
# ---------------------------------------------------------------- scenarios
#: Filler places. No word here is in `CONCEPTS`, and none names a planted fact.
PLACES = (
"north gate", "salt market", "ferry landing", "chapel steps", "rope walk",
"fish stalls", "old bridge", "tanner yard", "lamp street", "weir path",
"grain store", "boat yard", "watch house", "cloth hall", "eel traps",
"sheep fold", "smith forge", "stone quay", "reed beds", "toll booth",
)
PARAPHRASE_QUERY = "I ask Mara where she tucked the little brass dial that tells the hour."
UNRELATED_QUERY = "I ask the ferryman what rope costs at the landing this season."
def filler_prose(index: int, words: int) -> str:
"""Narration that moves on and never touches a planted fact."""
place = PLACES[index % len(PLACES)]
sentence = (f"At the {place} the travellers stopped, listened to the gulls over the "
f"grey water, and talked about the long road north.")
reps = max(1, round(words / len(sentence.split())))
return " ".join([sentence] * reps)
@dataclass
class Scenario:
"""One deterministic campaign. Depths: the opening is 0, turn *n*'s player
action is 2n-1 and its reply 2n."""
name: str
turns: int = 52
capacity: int = 80
top_k: int = 5
budget: int = 4096
prose_words: int = 60
plant_turn: int = 1
recall_text: str = "I ask Mara where she hid the amber sundial."
pin_first_memory: bool = False
lineage_control: bool = False
diagnose_recall: bool = True
SCENARIOS = {
"independent_default": Scenario("independent_default"),
"past_capacity": Scenario("past_capacity", capacity=6),
"past_capacity_pinned": Scenario("past_capacity_pinned", capacity=6, pin_first_memory=True),
# Closer to the shipped ratio (memory_top_k 5 against capacity 80): most of
# the bank is not retrieved on a given turn.
"past_capacity_low_top_k": Scenario("past_capacity_low_top_k", capacity=8, top_k=2),
"long_block_fact_early": Scenario("long_block_fact_early", turns=10, prose_words=850,
plant_turn=1, budget=16384),
"long_block_fact_late": Scenario("long_block_fact_late", turns=10, prose_words=850,
plant_turn=3, budget=16384),
"lineage_control": Scenario("lineage_control", lineage_control=True),
}
class ScriptNarrator:
"""Stands in for the narrator: returns `next_reply`, with an empty state block."""
next_reply = ""
last_usage = None
prompts: list = []
def __init__(self, *a, **k):
pass
async def generate(self, parts, *, temperature, max_tokens):
ScriptNarrator.prompts.append((parts.system, parts.story))
yield ("text", ScriptNarrator.next_reply)
def run_scenario(scenario: Scenario) -> dict:
"""Plays `scenario` through the real turn route and returns everything measured.
Uses the database `app.database` is already bound to, creating and dropping
its tables, the way the suite's fixtures do. Patches are applied here and
removed before returning, so this runs the same under pytest and from the CLI.
"""
import asyncio
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, limits
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import adventures as adventure_routes
summariser = BestCaseSummariser()
embedder = ConceptEmbedder()
patches = [
(memorybank, "summary_provider", lambda s: summariser),
(memorybank, "embedding_provider", lambda s: embedder),
# Post-turn work is settled explicitly after each turn, so eviction
# happens at a known point rather than whenever a background task runs.
(memorybank, "schedule_post_turn", lambda adventure: None),
(adventure_routes.turns, "OpenAICompatibleProvider", ScriptNarrator),
(limits, "check_row_cap", lambda *a, **k: None),
]
saved = [(obj, name, getattr(obj, name)) for obj, name, _ in patches]
for obj, name, value in patches:
setattr(obj, name, value)
ScriptNarrator.prompts = []
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
with SessionLocal() as db:
user = models.User(is_guest=False, email=f"b1-{scenario.name}@example.com")
db.add(user)
db.flush()
db.add(models.Settings(
user_id=user.id, model="script", endpoint_url="http://127.0.0.1:9/v1",
embedding_model="concept-embed", context_token_budget=scenario.budget,
max_output_tokens=500, memory_bank_capacity=scenario.capacity,
memory_top_k=scenario.top_k,
))
adventure = models.Adventure(
user_id=user.id, title=f"B.1 {scenario.name}", memory_bank_enabled=True,
auto_summarize=True, persona_name="Aldric",
)
db.add(adventure)
db.flush()
db.add(models.Action(adventure_id=adventure.id, type="start",
text="Rain over Westhaven, and the tavern door banging in the wind."))
db.commit()
adv, user_id = adventure.id, user.id
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
client = TestClient(app)
result: dict = {"scenario": scenario.__dict__.copy(), "trace": []}
def call(method, path, body=None, expect=200):
response = client.request(method, f"/api/adventures/{adv}{path}", json=body)
assert response.status_code == expect, (path, response.status_code, response.text[:300])
return response.json() if response.content else None
def marks():
with SessionLocal() as db:
rows = db.execute(select(models.Memory.id, models.Memory.forgotten,
models.Memory.embedded).where(
models.Memory.adventure_id == adv)).all()
summaries_n = db.query(models.Summary).filter_by(adventure_id=adv).count()
return tuple(sorted(rows)), summaries_n
def settle():
for _ in range(12):
before = marks()
asyncio.run(memorybank.run_post_turn(adv))
if marks() == before:
return
def memories():
with SessionLocal() as db:
return [dict(row._mapping) for row in db.execute(select(
models.Memory.id, models.Memory.text, models.Memory.source_start,
models.Memory.source_end, models.Memory.forgotten, models.Memory.pinned,
models.Memory.use_count, models.Memory.last_used_at, models.Memory.branch_id,
models.Memory.created_at).where(models.Memory.adventure_id == adv)
.order_by(models.Memory.id)).all()]
def turn(kind, text, reply):
ScriptNarrator.next_reply = f"{reply}\n```state\n{{\"events\": []}}\n```"
response = client.post(f"/api/adventures/{adv}/actions", json={"type": kind, "text": text})
assert response.status_code == 200, response.text[:300]
assert '"type": "error"' not in response.text, response.text[-300:]
plant_depth = None
f_memory_id = None
pinned_id = None
known: dict[int, dict] = {}
g: dict = {}
try:
for n in range(1, scenario.turns + 1):
if n == scenario.plant_turn:
turn("story", FACT_F.sentence, filler_prose(n, scenario.prose_words))
with SessionLocal() as db:
plant_depth = db.query(models.Action.depth).filter_by(
adventure_id=adv, text=FACT_F.sentence).scalar()
elif scenario.lineage_control and n == 21:
call("POST", "/checkpoints", {"name": "before the mill"}, expect=201)
turn("story", FACT_G.sentence, filler_prose(n, scenario.prose_words))
with SessionLocal() as db:
g["plant_depth"] = db.query(models.Action.depth).filter_by(
adventure_id=adv, text=FACT_G.sentence).scalar()
elif scenario.lineage_control and n == 30:
# Line A carries G's memory. Mark it, then abandon it: Undo back
# to before G was planted and write something else.
g["line_a"] = call("POST", "/checkpoints", {"name": "line A, after the mill"},
expect=201)["id"]
with SessionLocal() as db:
g_rows = [m for m in db.execute(select(models.Memory).where(
models.Memory.adventure_id == adv)).scalars() if FACT_G.carried_by(m.text)]
g["memory_ids"] = [m.id for m in g_rows]
with SessionLocal() as db:
g["last_action_id_before_divergence"] = db.query(models.Action.id).filter_by(
adventure_id=adv).order_by(models.Action.id.desc()).limit(1).scalar()
for _ in range(9):
call("POST", "/undo")
turn("do", f"I turn away from the mill and walk to the {PLACES[n % len(PLACES)]}.",
filler_prose(n + 100, scenario.prose_words))
g["diverged_at_turn"] = n
else:
turn("do", f"I walk on to the {PLACES[n % len(PLACES)]}.",
filler_prose(n, scenario.prose_words))
settle()
rows = memories()
created = [r["id"] for r in rows if r["id"] not in known]
newly_forgotten = [r["id"] for r in rows
if r["forgotten"] and not known.get(r["id"], {}).get("forgotten")]
for r in rows:
known[r["id"]] = r
if f_memory_id is None and plant_depth is not None:
for r in rows:
if (r["source_start"] is not None and r["source_start"] <= plant_depth
<= r["source_end"] and FACT_F.carried_by(r["text"])):
f_memory_id = r["id"]
if scenario.pin_first_memory and pinned_id is None:
candidate = next((r for r in rows if r["id"] != f_memory_id), None)
if candidate is not None:
call("PATCH", f"/memories/{candidate['id']}", {"pinned": True})
pinned_id = candidate["id"]
f_row = known.get(f_memory_id) if f_memory_id else None
result["trace"].append({
"turn": n,
"active": sum(1 for r in rows if not r["forgotten"]),
"total": len(rows),
"created": created,
"evicted": newly_forgotten,
"created_and_evicted_same_turn": sorted(set(created) & set(newly_forgotten)),
"f_memory_id": f_memory_id,
"f_forgotten": bool(f_row and f_row["forgotten"]),
"f_use_count": f_row["use_count"] if f_row else None,
})
turn("do", scenario.recall_text, filler_prose(999, scenario.prose_words))
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv)
settings = db.query(models.Settings).filter_by(user_id=user_id).first()
recall_action = (db.query(models.Action)
.filter(models.Action.adventure_id == adv,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
result["plant_depth"] = plant_depth
result["recall_depth"] = recall_action.depth
result["isolation"] = isolation(
db, adventure, FACT_F, plant_depth,
recall_snapshot=recall_action.context_snapshot,
recall_depth=recall_action.depth)
result["diagnosis"] = asyncio.run(diagnose(
db, adventure, settings, FACT_F, plant_depth,
recall_action=recall_action, embed=embedder.embed))
result["summariser_excerpts"] = len(summariser.excerpts)
memory_id = result["diagnosis"]["created"]["memory_id"]
if memory_id is not None and not result["diagnosis"]["retained"]["forgotten"]:
variants = {}
for label, query in (("direct", scenario.recall_text),
("paraphrase", PARAPHRASE_QUERY),
("unrelated", UNRELATED_QUERY)):
ranking = asyncio.run(rank_bank(db, adventure, settings, query, embedder.embed))
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
variants[label] = {"query": query, "rank": row and row["rank"],
"of": len(ranking["scored"]),
"similarity": row and row["similarity"],
"selected": bool(row and row["selected"]),
"top_k": ranking["top_k"]}
result["ranking_variants"] = variants
if memory_id is not None:
result["f_first_used_turn"] = next(
(t["turn"] for t in result["trace"] if (t["f_use_count"] or 0) > 0), None)
result["f_last_use_increase_turn"] = max(
(b["turn"] for a, b in zip(result["trace"], result["trace"][1:])
if (b["f_use_count"] or 0) > (a["f_use_count"] or 0)), default=None)
evicted_turn = next((t["turn"] for t in result["trace"] if t["f_forgotten"]), None)
first_evictions = next((t["evicted"] for t in result["trace"] if t["evicted"]), [])
result["eviction"] = {
"capacity": scenario.capacity,
"f_evicted_at_turn": evicted_turn,
"f_use_count_when_evicted": next(
(t["f_use_count"] for t in result["trace"] if t["f_forgotten"]), None),
"first_eviction_turn": next(
(t["turn"] for t in result["trace"] if t["evicted"]), None),
"first_evicted_ids": first_evictions,
"f_memory_was_first_evicted": bool(f_memory_id and f_memory_id in first_evictions),
"created_and_evicted_same_turn": sorted(
{i for t in result["trace"] for i in t["created_and_evicted_same_turn"]}),
"pinned_memory_id": pinned_id,
"pinned_memory_forgotten": bool(pinned_id and known[pinned_id]["forgotten"]),
}
if scenario.lineage_control:
path_clause = lineage.path_of(db, adventure).clause(models.Memory)
stored = db.execute(select(models.Memory.id).where(
models.Memory.id.in_(g.get("memory_ids") or [-1]))).scalars().all()
eligible = db.execute(select(models.Memory.id).where(
models.Memory.id.in_(g.get("memory_ids") or [-1]), path_clause)).scalars().all()
used_after = set()
injected_text = False
# Only turns played after the divergence. Before it, G was on the
# active line, and a memory of it being used then is correct.
for action in (db.query(models.Action)
.filter(models.Action.adventure_id == adv,
models.Action.type == "ai",
models.Action.id > g["last_action_id_before_divergence"])
.options(undefer(models.Action.context_snapshot))):
snap = action.context_snapshot or {}
for m in (snap.get("memories") or {}).get("used") or []:
if m.get("id") in (g.get("memory_ids") or []):
used_after.add(action.id)
if FACT_G.mentioned_by(_section(snap, MEMORIES_LABEL)):
injected_text = True
g.update(stored=stored, eligible_on_active_line=eligible,
turns_whose_memories_used_named_g=sorted(used_after),
g_text_ever_in_used_memories=injected_text)
if scenario.lineage_control:
call("POST", f"/checkpoints/{g['line_a']}/restore")
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv)
eligible = db.execute(select(models.Memory.id).where(
models.Memory.id.in_(g.get("memory_ids") or [-1]),
lineage.path_of(db, adventure).clause(models.Memory))).scalars().all()
g["eligible_after_returning_to_line_a"] = eligible
result["lineage_control"] = g
return result
finally:
for obj, name, value in saved:
setattr(obj, name, value)
app.dependency_overrides.clear()
adventure_routes.turns._active_turns.clear()
memorybank._vector_cache.clear()
Base.metadata.drop_all(bind=engine)
+122
View File
@@ -0,0 +1,122 @@
"""v1.1 WP-B.1: run the deterministic memory-retention scenarios, or diagnose a real campaign.
# the deterministic scenarios, against an isolated database in --out
.venv/bin/python -m tools.v11_b1_memory scenarios --out "$HOME/v11-evidence/b1/<label>"
# the four stages for a finished real campaign (reads its database; embeds
# the recall query with the campaign's own configured embedding model)
AIDND_TEST_ENDPOINT=... AIDND_TEST_EMBED_MODEL=nomic-embed-text:latest \\
.venv/bin/python -m tools.v11_b1_memory diagnose --db <campaign.db> \\
--plant-depth 3 --out "$HOME/v11-evidence/b1/<label>"
Run from `backend/`. Nothing here changes memory behaviour; see
`tools/memory_diagnostic.py` for what is measured and what the stubs model.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sys
from pathlib import Path
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
sub = parser.add_subparsers(dest="command", required=True)
scen = sub.add_parser("scenarios")
scen.add_argument("--out", required=True)
scen.add_argument("--only", action="append", default=[])
diag = sub.add_parser("diagnose")
diag.add_argument("--db", required=True)
diag.add_argument("--plant-depth", type=int, required=True)
diag.add_argument("--out", required=True)
args = parser.parse_args()
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
if args.command == "scenarios":
db_path = out / "scenarios.db"
if db_path.exists():
db_path.unlink()
os.environ["AIDND_DB_PATH"] = str(db_path)
else:
# A copy, so diagnosis never writes to the evidence database.
copy = out / "diagnosed-copy.db"
shutil.copy2(args.db, copy)
os.environ["AIDND_DB_PATH"] = str(copy)
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from tools import memory_diagnostic as md # after the database is chosen
if args.command == "scenarios":
names = args.only or list(md.SCENARIOS)
summary = {}
for name in names:
result = md.run_scenario(md.SCENARIOS[name])
(out / f"{name}.json").write_text(json.dumps(result, indent=2, default=str))
d = result.get("diagnosis") or {}
summary[name] = {
"verdict": d.get("verdict"),
"isolation_ok": (result.get("isolation") or {}).get("ok"),
"plant_depth": result.get("plant_depth"),
"recall_depth": result.get("recall_depth"),
"f_evicted_at_turn": (result.get("eviction") or {}).get("f_evicted_at_turn"),
}
print(f"{name:26} verdict={d.get('verdict')!s:26} "
f"isolation_ok={summary[name]['isolation_ok']} "
f"plant={result.get('plant_depth')} recall={result.get('recall_depth')}")
(out / "summary.json").write_text(json.dumps(summary, indent=2))
return 0
import asyncio
from sqlalchemy.orm import undefer
from app import memorybank, models
from app.database import SessionLocal
endpoint = os.environ.get("AIDND_TEST_ENDPOINT", "")
embed_model = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
with SessionLocal() as db:
adventure = db.query(models.Adventure).order_by(models.Adventure.id).first()
settings = db.query(models.Settings).filter_by(user_id=adventure.user_id).first()
if endpoint:
settings.endpoint_url = endpoint
if embed_model:
settings.embedding_model = embed_model
recall_action = (db.query(models.Action)
.filter(models.Action.adventure_id == adventure.id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
embed = memorybank.embedding_provider(settings).embed
iso = md.isolation(db, adventure, md.FACT_F, args.plant_depth,
recall_snapshot=recall_action.context_snapshot,
recall_depth=recall_action.depth)
diagnosis = asyncio.run(md.diagnose(db, adventure, settings, md.FACT_F, args.plant_depth,
recall_action=recall_action, embed=embed))
variants = {}
memory_id = diagnosis["created"]["memory_id"]
if memory_id is not None and not diagnosis.get("retained", {}).get("forgotten"):
for label, query in (("recall_turn", diagnosis.get("ranked", {}).get("query", "")),
("paraphrase", md.PARAPHRASE_QUERY),
("unrelated", md.UNRELATED_QUERY)):
ranking = asyncio.run(md.rank_bank(db, adventure, settings, query, embed))
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
variants[label] = {"rank": row and row["rank"], "of": len(ranking["scored"]),
"similarity": row and row["similarity"],
"selected": bool(row and row["selected"])}
db.rollback()
report = {"isolation": iso, "diagnosis": diagnosis, "ranking_variants": variants}
(out / "diagnosis.json").write_text(json.dumps(report, indent=2, default=str))
print(json.dumps({"isolation_ok": iso["ok"], "verdict": diagnosis["verdict"]}, indent=2))
return 0
if __name__ == "__main__":
sys.exit(main())
+623
View File
@@ -0,0 +1,623 @@
# v1.1 WP-B.1 — Independent Long-Term Memory Retention Diagnostic
**Status:** COMPLETE. Diagnostic only: no memory behaviour was changed. The final decision is in §S.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD at start | `d63804f22ecbaed80741241a154cdaef82f7b2ed` — *v1.1: harden context window and narrator protocol boundary* (WP-A1/A2), signed by the owner (good signature, RSA key `02C9BF7D…`) |
| Its parents | `ac465ed` (planning v4.1, signed) and `432f041` (the signed v1.0.0 release commit; tag `v1.0.0`; `main`) |
| Working tree at start | clean |
| `git diff --stat v1.0.0..HEAD` | 27 files, +4,880 / −88: the planning commit and WP-A1/A2 |
| Comparison baseline | `432f041` (v1.0.0), run from a throwaway worktree (§D) |
`app/memorybank.py`, `app/summaries.py`, `app/context/lineage.py`, `app/tree.py` and
`app/vectors.py` are unchanged between `v1.0.0` and HEAD. A1 and A2 did not touch
the memory pipeline, apart from adding `accounting` to `attempts.ATTEMPT_KEYS`.
---
## B. Existing memory pipeline
Answered from the code at HEAD. Nothing was changed.
| # | Question | Answer |
| --- | --- | --- |
| 1 | How are memories generated? | `memorybank.run_post_turn`, a fire-and-forget task after each accepted turn (`schedule_post_turn`), runs `_create_due_memories` when the campaign has `auto_summarize`. It writes one memory per block of `MEMORY_INTERVAL` = 6 story actions past the memory cursor. It starts once the story has `MEMORY_START` = 12 actions, and only when `SETTLE_SLACK` = 1 action sits past the block. At most `MAX_MEMORIES_PER_RUN` = 5 memories are written per run. |
| 2 | What range does a memory cover? | The block's first and last action depths: `Memory.source_start`, `Memory.source_end`. |
| 3 | How are source depth and lineage stored? | `tree.attach_memory` sets `Memory.branch_id` and `Memory.depth` from the block's **last** node, so a memory is visible exactly on paths that contain that node. |
| 4 | Memory text length limit | Prompt-only: `MEMORY_MAX_WORDS` = 50 in `MEMORY_SYSTEM_PROMPT` ("1-2 plain sentences"). Nothing truncates the stored text. |
| 5 | What input does the summariser receive? | `summarize_block`: a cast brief (`cast_brief`, which reads story cards, persona and plot essentials), then `"Story excerpt:\n\n{excerpt}\n\nMemory:"`. The excerpt is the block's action texts joined by blank lines. |
| 6 | Where does the 2,000-token truncation happen? | `summarize_block`: `excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)`, with `MEMORY_EXCERPT_TOKENS` = 2,000. It keeps the block's **last** 2,000 `cl100k_base` tokens. The cast brief is matched against the untruncated block. |
| 7 | How does capacity and eviction work? | `_evict_over_capacity`, at the end of every `run_post_turn`. It counts non-forgotten memories in the whole adventure (not lineage-scoped). Above `settings.memory_bank_capacity` (default 80), it marks `forgotten = True` on the overflow unpinned memories, ordered by `coalesce(last_used_at, created_at)` ascending, then `use_count`. Forgotten rows are kept. |
| 8 | How is `last_accessed` updated? | The field is `Memory.last_used_at`. `record_use` sets it, and increments `use_count`, for every memory in the turn's `memories.used`. That write is part of the turn's single commit (§O.7 of the M11 report). Dry runs never count. |
| 9 | How are memories ranked? | `retrieve_memories`. The query is the text of the newest `RETRIEVAL_WINDOW_ACTIONS` = 4 actions, truncated to the last `RETRIEVAL_WINDOW_TOKENS` = 600 tokens. It is embedded with the configured embedding model. Every eligible memory is scored by `vectors.cosine`. Pinned memories are taken first, the rest fill `memory_top_k` (default 5) in score order, and `_drop_redundant` skips a candidate at cosine ≥ 0.93 to one already chosen, within the same authority. **There is no lexical term, recency term, importance term or similarity floor** (CONTEXT-AND-MEMORY §20, as implemented). |
| 10 | How do pins affect ranking and eviction? | Ranking: a pinned memory is always selected, and counts toward `memory_top_k`. Eviction: pinned memories are never evicted, and if every active memory is pinned, capacity is exceeded. |
| 11 | How does the active-lineage clause filter memories? | `lineage.path_of(db, adventure).clause(models.Memory)` matches `(branch_id, depth)` against the head's path entries, capped at each fork depth. Retrieval also requires `forgotten = False` and `embedded = True`. |
| 12 | How do retrieved memories enter `build_context`? | Through the `memory_bank` argument. The builder renders `"Memories from earlier in the story. Lines marked [inferred] are interpretation, not established fact — do not treat them as settled truth:"` plus one `- [inferred]? text` line per memory, as section `used_memories`. It is a live section priced into protected context, placed after history and before `narrative_state`. |
| 13 | Where is the provenance recorded? | `context_snapshot["memories"]`: the whole retrieval result, meaning `used` (id, text, similarity, pinned, authority, source range), `considered` and `suppressed`. It is stored per turn, so a past turn's selection is inspectable. |
| 14 | How does memory survive export and import? | `bundle._exported_memory` carries text, pinned, forgotten, `sourceStart`, `sourceEnd`, `useCount`, authority, branch and depth. **It does not carry the vector, `last_used_at` or `created_at`.** On import the memories are re-embedded by the post-turn pass, and their recency restarts from the import. |
**A finding from the inspection itself.** The v1 long-run harness (`tools/m11_long_run.py`)
planted M04's clue as an accepted **state correction** (`CLUE_FACT`, via
`add_fact`). The player turn mentioning it uses the sentinel code, but the fact the
recall looks for was established in state. Memories are written from **story
text**. So every v1 M04 recovery could only ever run through state, or through the
narrator restating state, and none could have tested memory on its own. That is why
§P risk 5 of the M11 report could say "no run showed memory keeping a planted fact"
without any run having given memory the chance.
---
## C. Diagnostic design
### C.1 What is built
| File | Kind | Purpose |
| --- | --- | --- |
| `backend/tools/memory_diagnostic.py` | tool, new | See below: the fact spec, isolation checks, four-stage diagnosis, deterministic stubs and scenario runner. |
| `backend/tools/v11_b1_memory.py` | tool, new | CLI. `scenarios` runs the deterministic campaigns against an isolated database and writes JSON. `diagnose` runs the four stages against a **copy** of a finished real campaign's database. |
| `backend/tools/m11_long_run.py` | tool, extended | `--independent-fact`: see §C.3. |
| `backend/tests/test_v11_b1_memory_diagnostic.py` | tests, new | The deterministic diagnostic. |
| `backend/tests/test_v11_b1_long_run_verdict.py` | tests, new | The new long-run verdict. |
What `tools/memory_diagnostic.py` holds:
- **`Fact`:** a planted fact, with whole-word carry and leak matching.
- **`isolation()`:** every non-memory layer checked.
- **`diagnose()`:** the four stages and the verdict.
- **`rank_bank()`:** production ranking recomputed for every eligible memory.
- **The stubs:** `BestCaseSummariser`, `ConceptEmbedder` and `ScriptNarrator`.
- **`run_scenario()`:** a campaign played through the real turn route.
**No application file is changed.** No column, table, migration or setting is
added. The diagnostic's extra fields are computed at report time.
### C.2 How each stage is judged
| Stage | Judged from | Output (real field names where they exist) |
| --- | --- | --- |
| **created** | Memories on the active lineage with `source_start ≤ plant_depth ≤ source_end` whose text carries F (both "sundial" and "teapot"). For every covering memory, the block is re-read (`memorybank.source_block`) and cut exactly as `summarize_block` does, to report whether F was in the block and whether it was in **the excerpt the summariser saw**. | `memory_id`, `source_start`, `source_end`, `memory_text`, `covering_memories[].{block_tokens, fact_in_block, fact_in_summariser_excerpt}` |
| **retained** | That memory's row, plus the eviction order production would use (the same `ORDER BY`, read-only) | `forgotten`, `pinned`, `embedded`, `on_active_lineage`, `use_count`, `last_used_at`, `created_at`, `active_memories`, `memory_bank_capacity`, `eviction_position`, `reason` |
| **ranked** | `rank_bank`: the same catalogue clause, cosine, pin rule and `memorybank._drop_redundant`, over **every** eligible memory, with the recall turn's own query (`history.tail(4, exclude=recall AI node)`, cut to 600 tokens). It is checked against the recall turn's stored `memories.used` (`replica_matches_stored_selection`). | `semantic_score` (`similarity`), `lexical_score` (always `None`: no such term exists), `final_score`, `rank`, `of`, `top_k_cutoff`, `selected`, `suppressed_as_duplicate_of`, `query` |
| **injected** | The recall turn's stored `context_snapshot`: `memories.used` names the memory, and its text is in the `used_memories` section | `context_component`, `in_stored_memories_used`, `text_in_section`, `token_count` |
Verdicts, in order: `not_created`, `created_but_evicted`, `retained_but_not_ranked`,
`ranked_but_not_selected`, `selected_but_not_injected`, `injected`. The fifth is
added to the brief's list, so that "the retrieval picked it" and "the narrator was
shown it" stay distinguishable.
### C.3 Isolation (precondition) checks
A result counts only if every check holds.
| Check | How |
| --- | --- |
| `state_document` | `adventure.narrative_state`: entities, facts, relationships, threads, scene and possessions, plus the whole document |
| `state_snapshots` | `narrative_state_after` of every node on the active lineage |
| `later_narration` | every AI turn deeper than the planting block's end, and before the recall turn |
| `summary` | `summaries.current` and the recall prompt's `story_summary` section |
| `knowledge` | every `KnowledgeSource.content`, and the recall prompt's imported-knowledge sections |
| `recent_history` | the recall prompt's `history.floor_depth` is greater than the planting depth, and F is not in the `history` or `recent_history` sections |
| `state_section` | the recall prompt's `narrative_state` section |
The negative control for the precondition itself is
`test_the_isolation_check_fails_when_another_layer_carries_the_fact`.
### C.4 The deterministic stubs, and what they model
- **`BestCaseSummariser`.** An ideal memory writer. A memory keeps every sentence of
its excerpt that carries a planted fact, plus one sentence naming the block's own
place so that memories differ. Summary updates never mention a planted fact. **A
creation failure under this stub is the application's, not a model's.**
- **`ConceptEmbedder`.** A 96-dimension deterministic embedding. Words in a small
concept table ("sundial", "dial", "hour", "clock", …) share a dimension, other
words are hashed, and the vector is normalised. It models a paraphrase landing
near the original. **It says nothing about `nomic-embed-text`.**
- **`ScriptNarrator`.** Narration that names only filler places and never a planted
fact, with an empty state block, so state never records F.
Scenarios are played through the real `POST /actions` route. Automatic post-turn
scheduling is replaced by an explicit `run_post_turn` settle after every turn, so
eviction happens at a known turn. Depths: the opening is 0, turn *n*'s player
action is 2n−1 and its reply 2n. The planted fact is a `story` action.
### C.4.1 Scenarios
| Scenario | Turns | Capacity | top_k | History budget | Prose per reply | Planted at |
| --- | --- | --- | --- | --- | --- | --- |
| `independent_default` | 52 (recall at depth 106) | 80 | 5 | 4,096 | ~60 words | depth 1 |
| `past_capacity` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1 |
| `past_capacity_pinned` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1; first other memory pinned |
| `past_capacity_low_top_k` | 52 | 8 | 2 | 4,096 | ~60 words | depth 1 |
| `long_block_fact_early` | 10 | 80 | 5 | 16,384 | ~850 words | depth 1, early in a long block |
| `long_block_fact_late` | 10 | 80 | 5 | 16,384 | ~850 words | depth 5, late in the same-sized block |
| `lineage_control` | 52 | 80 | 5 | 4,096 | ~60 words | F at depth 1; G on line A, then Undo × 9 and divergence at turn 30; Save Points before and after G |
### C.5 The long-run verdict
`tools/m11_long_run.py --independent-fact` plants a second, **story-only** fact
at depth 3, right after M04's own plant, and never corrects it into state. Every
existing M04 behaviour and verdict is unchanged.
**Every accepted turn** records the first failure of:
- `absent_from_state`;
- `absent_from_summary`;
- `absent_from_later_narration`, meaning narration deeper than the planting depth
plus 6.
The plant and those first failures survive `--resume`.
**At recall** a dedicated question is played. `_independent_recall` reads the
recall turn's stored context and the campaign database (read-only), and
`_independent_memory_verdict` returns one of:
- `recovered_through_memory_independent`, only when every one of
`planted_turn_outside_history`, `absent_from_state`, `absent_from_summary`,
`absent_from_knowledge` and `absent_from_later_narration` holds, and a memory
covering the planting turn carries the fact and was injected;
- `precondition_failed:<name>` or `precondition_unknown:<name>`, never a recovery;
- `not_recovered:not_created`, `not_recovered:evicted` or
`not_recovered:not_injected`.
Ranking is recomputed afterwards by `tools/v11_b1_memory.py diagnose`, against a
copy of the run's database.
---
## D. v1.0.0 baseline
**How it was run.**
- A throwaway worktree was checked out at `432f041` (`git describe`: `v1.0.0`).
- Only the three B.1 files were copied in: `tools/memory_diagnostic.py`,
`tools/v11_b1_memory.py` and `tests/test_v11_b1_memory_diagnostic.py`.
- The imported `app` was confirmed to come from the worktree.
- The worktree was removed afterwards, and the `v1.0.0` tag and commit were not
touched.
- `app/memorybank.py` is byte-identical between `v1.0.0` and HEAD.
- Evidence: `$HOME/v11-evidence/b1/v100/`, with HEAD's in
`$HOME/v11-evidence/b1/head/`.
| | v1.0.0 (`432f041`) | HEAD (`d63804f` + B.1 files) |
| --- | --- | --- |
| `test_v11_b1_memory_diagnostic.py` | **24 passed, 2 xfailed (strict)** | 24 passed, 2 xfailed (strict) |
| `independent_default` | `injected` | `injected` |
| `past_capacity` (capacity 6, top_k 5) | **`created_but_evicted`** | `created_but_evicted` |
| `past_capacity_pinned` | `created_but_evicted` (the pinned memory kept) | same |
| `past_capacity_low_top_k` (capacity 8, top_k 2) | **`created_but_evicted`** | `created_but_evicted` |
| `long_block_fact_early` | **`not_created`** | `not_created` |
| `long_block_fact_late` | `injected` | `injected` |
| `lineage_control` | `injected`; G never injected after the divergence | same |
**The two independent-retention acceptance criteria fail on v1.0.0, and each names
the stage.**
```text
criterion: an early fact is recalled from memory past capacity
created: yes
retained: no
FAILURE STAGE: retention (capacity eviction)
criterion: a fact early in a long block is remembered
created: no (the fact was in the block, not in the summariser's excerpt)
FAILURE STAGE: creation (input truncation)
```
Under the best-case summariser, with blocks shorter than 2,000 tokens and a bank
under capacity, v1.0.0 carries the fact all the way to injection. The two failure
stages above are therefore **application mechanisms**, reached under conditions
a long campaign meets:
- a block of long narration;
- more memories than `memory_bank_capacity`.
Which of them a real campaign meets first is §K's question.
---
## E. Creation results
`long_block_fact_early` and `long_block_fact_late` use the same block geometry: the
opening (depth 0) and turns 1-3. Each reply is about 850 words, and the block
`source_start` 0 … `source_end` 5 is **2,079 tokens**, 79 over
`MEMORY_EXCERPT_TOKENS`.
| Shape | Planted at | Fact in block | Fact in summariser excerpt | Memory written | Verdict |
| --- | --- | --- | --- | --- | --- |
| F early in the block | depth 1 | yes | **no** | "The travellers spent time at the ferry landing." | **`not_created`** |
| F late in the same-sized block | depth 5 | yes | yes | "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." | `injected` |
`summarize_block` keeps the **last** 2,000 tokens. An early-block fact is cut off
before the summariser reads it, even by an overflow of only 79 tokens. No summariser
quality can recover what it was never given.
`independent_default`: 60-word replies, a 4,096 budget, and a block under 2,000
tokens. Memory 1 covers depths 0-5, and its text carries F. Both the block and the
excerpt contain F.
---
## F. Retention / capacity results
Default: the bank holds 17 active memories at depth 106, against capacity 80. F's
memory is retained, has `use_count` 16, and sits at eviction position 14 of 17.
**Past capacity** (the brief's capacity test). 17 memories are written over 52 turns.
| Scenario | F's uses before eviction | F's last retrieval | First eviction | F evicted | F first evicted? | Created and evicted in the same pass | Pinned kept |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `past_capacity` (6 / top_k 5) | 12 | turn 18 | turn 21, memory 1 | **turn 21** | **yes** | none | — |
| `past_capacity_pinned` (6 / top_k 5, memory 2 pinned) | 12 | turn 18 | turn 21, memory 1 | turn 21 | yes | none | **memory 2 never evicted** |
| `past_capacity_low_top_k` (8 / top_k 2) | 4 | turn 24 | turn 27, memory 3 | **turn 36** (4th eviction) | no | none | — |
**What the trace shows about the rule.**
- **"Never retrieved" is not the mechanism.** F was retrieved, 12 or 4 times, while
it still ranked in the top-k for the recent-narration query.
- **Eviction orders by `coalesce(last_used_at, created_at)`.** An early fact that
recent narration never mentions stops being retrieved once newer memories fill
the top-k. It then ages out:
- at capacity 6, it is gone 3 turns after its last use, as the first eviction;
- at capacity 8 with `top_k` 2, 12 turns after its last use, as the fourth.
- **Recall itself cannot rescue it.** Retrieval is driven by recent narration, and
a fact nobody mentions is exactly the one that loses recency.
- **Pinned memories stay protected** (`past_capacity_pinned`).
- **No memory is evicted by the pass that created it.** The frozen-bank regression
fix still holds in all three scenarios.
---
## G. Ranking results
The `independent_default` recall turn is at depth 106. Its query is the newest 4
actions, cut to 600 tokens: three narration turns and the one-line question.
| | Value |
| --- | --- |
| F's `similarity` (semantic score) | **0.241** |
| Lexical score | none: memory ranking has no lexical term |
| Pin effect | none (not pinned) |
| Final rank | **2 of 17** |
| `memory_top_k` cutoff | 5 |
| Selected | yes |
| Replica agrees with the turn's stored `memories.used` | **yes** |
The same memory against three stand-alone queries, `ConceptEmbedder`:
| Query | Rank | Similarity | Selected |
| --- | --- | --- | --- |
| direct: "I ask Mara where she hid the amber sundial." | 1 of 17 | 0.708 | yes |
| paraphrase: "… the little brass dial that tells the hour." | 1 of 17 | 0.636 | yes |
| unrelated: "… what rope costs at the landing this season." | 5 of 17 | **0.064** | **yes** |
Two diagnostic observations. Neither is a failure in this scenario.
1. **The production query dilutes the question.** The question alone scores 0.708;
inside the four-action window it scores 0.241. F survives at rank 2 of 17. In a
bank where more memories share the recent narration's vocabulary, the same
dilution would push it past `top_k`.
2. **There is no relevance floor.** With more memories than `memory_top_k`, five are
injected whatever their similarity. An unrelated query still injects F at 0.064.
This matters to F only in the other direction: it can ride along even when it is
not relevant.
---
## H. Injection results
`independent_default` injects F: `memories.used` in the recall turn's stored snapshot
names memory 1, its text is in the `used_memories` section, and that section is 97
tokens. The same holds in `long_block_fact_late` and `lineage_control`.
**`selected_but_not_injected` never occurred**: whatever retrieval selected, the
builder rendered.
---
## I. Lineage negative control
`lineage_control`:
- G is planted on line A at depth 41, turn 21.
- G's memory 7 is written.
- A Save Point is placed on line A after G.
- Undo ×9, then divergent writing at turn 30. The last action before the
divergence has id 59.
| Assertion | Result |
| --- | --- |
| G's memory stays stored | **yes** (memory 7 present) |
| G is not eligible on the active lineage | **yes** (the path clause returns nothing) |
| G is never injected after the divergence | **yes** (no turn with id > 59 names it or carries its text) |
| `memories.used` does not report it after the divergence | **yes** |
| Returning to line A (Save Point restore) makes it eligible again | **yes** (memory 7 eligible) |
| F on the active line is unaffected | `injected`, isolation holds |
Before the divergence, G's memory was legitimately used on line A (turn 22). The
first version of this check counted that as a leak. That was a defect in the
diagnostic, and the scan now starts after the divergence.
No lineage code was touched. `test_m11_leakage.py`, 14 tests, passes unchanged (§P).
---
## J. Authority negative control
`test_a_memory_that_contradicts_state_loses_and_changes_nothing` sets up the
conflict like this:
- a state correction adds the fact "the tavern lamp is lit";
- a hand-written memory says "The tavern lamp was never lit that night.";
- the memory is pinned, so it is injected;
- a turn is played with an empty proposal.
| Assertion | Result |
| --- | --- |
| The narrative state document is unchanged by retrieval and the turn | **yes** (identical before and after) |
| The state fact is in the prompt's `narrative_state` section | yes |
| The memory is in `used_memories`, under "Memories from earlier in the story …" | yes, framed as historical and non-canon context |
| `narrative_state` comes after `used_memories`, so state is read last and settles the conflict | yes |
F07 semantics are unchanged: memory never writes state.
---
## K. Real-model attempts
**Setup.**
- **Command:** `tools/m11_long_run.py --turns 100 --independent-fact`.
- **Host and models:** the GPU inference host (Ollama 0.34.0), `qwen2.5:3b-instruct-16k` at a verified 16,384 window, embeddings by `nomic-embed-text:latest`.
- **Memory:** the memory bank and auto-summarise on.
- **Harness settings:** `memory_top_k` 4 and `context_token_budget` 16,384.
- **Logging:** the owner's power, link and kernel logging was running on the host before the run started.
- **Permission:** inference was used only with the owner's explicit approval.
**Evidence:**
- `$HOME/v11-evidence/b1/real-1/`: `summary.json`, `recall-independent.json`, `timeline.jsonl` and `campaign.db`;
- `$HOME/v11-evidence/b1/real-1-diagnosis/diagnosis.json`: the four stages, recomputed on a copy of the database with the same embedding model.
### K.1 Attempt 1 — PRECONDITION FAILED
| | |
| --- | --- |
| Run status | `complete`: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 active memories (19 eligible on the recall line), 1,308 s elapsed |
| Window / accounting | verified 16,384 on every turn; `fits` |
| M04 verdict (unchanged) | `recovered_through_state_only` |
| Independent fact | planted at depth 3 by the player turn *"I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one."* It was never corrected into state |
| **Verdict** | **`precondition_failed:absent_from_summary`** (the verdict names the first failed precondition in its fixed order) |
| Precondition | Result | First failure |
| --- | --- | --- |
| planted turn outside recent history | **held**: the history window started at depth 66 | — |
| absent from authoritative state | **held**: the document, every node's snapshot and the recall prompt's state section | — |
| absent from imported knowledge | **held**: 3 sources | — |
| **absent from later narration** | **failed**: the narrator mentions the fact at depths 6, 8, 10, 14, 16, 18, 20, 22, 24, 26 and later | accepted turn **5** |
| **absent from the active summary** | **failed** | accepted turn **9** |
| fact text in the recall prompt's history sections | **present** (restated narration inside the window) | — |
**Why it did not qualify.** The 3B narrator took the planted detail up as a motif
and restated it for the rest of the campaign ("Mara's amber sundial flickered
softly, a silent reminder of their shared history"). The summariser, whose prompt
asks it to "preserve important established facts", folded it into the running
summary. Neither is a defect in the harness. They are the other layers doing what
they do with a salient fact, which is exactly what makes a clean memory-only
measurement hard to obtain with a real narrator.
**The four stages, diagnosed anyway.** They are not evidence for independent
retention, but they are evidence for the mechanism.
| Stage | Result |
| --- | --- |
| **created** | **yes.** Memory 1 covers depths 0-5. The block is 688 tokens, so there was no truncation: the fact was in the block and in the summariser's excerpt. The memory text: *"Aldric taps the silver key against his chest … Mara slipped the amber sundial into the cracked teapot, her fingers tightening on the silver key. …"* |
| **retained** | **yes.** Not forgotten, `use_count` 53, 33 active memories against capacity 80, eviction position 28 |
| **ranked** | **no.** Eligible and embedded. For the recall turn's own query it scored `similarity` **0.866** and ranked **10 of 19**, against `top_k_cutoff` 4. It was not selected and was not suppressed as a duplicate. The replica matched the stored selection |
| **injected** | not reached. The recall turn's `memories.used` = [30, 29, 31, 5] |
**What the narrator was given instead.** Five **later** memories also carry the
fact, all written from narration that restated it: memories 12 (depths 66-71), 28
(81-86), 30 (93-98), 17 (96-101) and 18 (102-107). Memory 30 was injected at
recall. The fact reached the narrator through memory, but through a
**restatement's** memory, not the planting-era one.
**The same memory 1 against reference queries** (`nomic-embed-text`):
| Query | Rank | Similarity | Selected |
| --- | --- | --- | --- |
| the recall turn's production query | 10 / 19 | 0.866 | no |
| paraphrase: "… the little brass dial that tells the hour" | 6 / 19 | 0.593 | no |
| unrelated: "… what rope costs at the landing this season" | 12 / 19 | 0.435 | no |
Two properties of the real embedder matter to B.2:
1. **A high floor.** An unrelated query still scores 0.44.
2. **A crowded top.** The bank is full of near-identical "Aldric and Mara step out,
the silver key's weight in his pocket" memories, so 0.866 was not enough to reach
the top 4.
*A side finding outside B.1's scope.* The export's protocol-leak counter flags
**1 of 105** stored turns: action 153, depth 143. Mid-reply, the narrator echoed the
length hint and the state reminder, with an `Events: [...]` line, and then continued
the story. WP-A2's cleanup rules act only at the end of a reply, so an instruction
echo with story after it stays. It is recorded here for the v1.1 backlog. B.1 did not
touch it.
---
## L. First failing stage
**Deterministic, with a best-case summariser and a concept embedder, identical on
v1.0.0 and HEAD:**
| Condition | First failing stage |
| --- | --- |
| a bank under capacity, blocks under 2,000 tokens | **none**: created, retained, ranked (2 of 17) and injected |
| a bank past `memory_bank_capacity` | **retention**: evicted by least-recently-used order after recent narration stops retrieving it (§F) |
| the fact early in a block over 2,000 tokens | **creation**: the fact is in the block, but not in the summariser's excerpt (§E) |
**Real model, attempt 1** (isolation not met, mechanism only): created yes, retained
yes, **ranking failed first**. The planting-era memory scored 0.866 but ranked 10th of
19 behind later, near-identical memories, and outside `top_k` 4.
**Across all the evidence, the first stage an early fact fails at a real campaign's
length is ranking.** In both the deterministic default and the real run, the memory
exists and is retained at 100 turns, where the bank is below capacity. What decides
whether the narrator is shown it is its rank against a recent-narration query in a
bank of similar memories.
- **Retention (eviction)** is a second, later failure. It is certain once a campaign
outgrows capacity: about 480 actions at the defaults.
- **Creation (truncation)** is a third, conditional one. It needs blocks longer than
2,000 tokens, which the 16,384 window with 500-token replies did not produce
(688 tokens).
---
## M. Evidence for likely root cause
1. **The retrieval query is recent narration, not the question** (§B.9, §G). The
query is the last 4 actions cut to 600 tokens, so a one-line recall question is
outweighed by three turns of prose. The effect is deterministic: the question's
own similarity of 0.708 fell to 0.241 in the production query. In the real run,
recent prose about the same tavern, key and people made every memory look
similar, and ten ranked above the planting-era one.
2. **Ranking has no term that favours the planting-era record** (§B.9;
CONTEXT-AND-MEMORY §20). There is cosine only: no lexical match on the question's
rare terms ("sundial", "teapot"), no importance, and no preference for the earliest
or a coverage-distinct source. Later restatement memories carry the same words in
more familiar company, and outrank the original.
3. **Eviction is purely least-recently-used** (§B.7, §F). Retrieval is driven by
recent narration, so exactly the facts nothing recent mentions lose recency and go
first. Recall itself cannot rescue them, because they are no longer retrieved.
4. **The summariser reads only the last 2,000 tokens of a block** (§B.6, §E). This is
proven deterministically. It did not bite at the real run's block sizes.
5. **v1's M04 never tested memory** (§B finding). The clue was planted as state, so
the long-standing "memory does not keep the fact" observation was never a
measurement of memory.
---
## N. What B.2 is allowed to change
B.2 is allowed only the smallest changes the evidence supports, one mechanism at a
time, each with a failing test first. In order of the evidence:
1. **Ranking** (the first failing stage). Candidates:
- build the retrieval query so the player's newest input is not drowned out, for
example by giving the newest player action its own weight or its own query;
- and/or add one inspectable ranking term from CONTEXT-AND-MEMORY §20, most
directly a lexical match on the query's rare terms.
Either must keep `replica_matches_stored_selection` meaningful: a
diagnostic-visible score, recorded in `memories.used`.
2. **Eviction** (the certain second failure). Stop least-recently-used eviction from
discarding a never-again-retrieved early memory first. For example, weight eviction
by coverage, keeping the only memory of a story range, or by age, instead of recency
alone. The frozen-bank protection must be kept.
3. **Creation** (conditional). Choose the summariser's excerpt so that a fact early in
a long block is not cut. For example, the head and tail, or the whole block up to a
larger bound.
Each change turns one of B.1's diagnostics into a passing result:
- `past_capacity` and `long_block_fact_early` flip their strict xfails;
- a real-model re-run shows `ranked: yes` for the planting-era memory.
## O. What B.2 must not change
- **Lineage safety.** `tree.attach_memory`, the path clause, `forget_node` and E02
stay as they are. An abandoned line's memory stays stored and ineligible (§I).
- **Summary lineage** (E03) and summary content policy.
- **Authority.** Memory never writes state and is never framed as canon (F07, §J).
- **Imported-knowledge authority and retrieval.**
- **Pins.** Pinned memories stay always-selected and never evicted.
- **The frozen-bank fix.** A memory is never evicted by the pass that created it.
- **The single-commit use counter** (M11 §O.7).
- **F01-F08, E01-E04, the M04 verdicts, the bundle format and the schema**, unless a
migration is separately justified.
- **The deterministic diagnostic itself.** B.2 flips the strict xfails. It does not
weaken the scenarios or the isolation checks.
---
## P. Tests / regression
| Run | Result |
| --- | --- |
| `test_v11_b1_memory_diagnostic.py` on HEAD | **24 passed, 2 xfailed (strict)** |
| `test_v11_b1_memory_diagnostic.py` on v1.0.0 (§D) | **24 passed, 2 xfailed (strict)**, identical |
| `test_v11_b1_long_run_verdict.py`, `test_m11_long_run_memory.py`, `test_m11_long_run_resume.py` | **52 passed** |
| **Full backend suite**, HEAD plus the B.1 files | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,323 s). The 17 skips are the tests that need a real model, the same 17 as before. The 2 strict xfails are the two diagnosed retention criteria (§D). |
| **Full backend suite, final re-run at staging** (after the real-model attempt; the staged tree) | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,912 s), identical |
No frontend file was changed, so the frontend suite, lint and build are not affected.
---
## Q. Compatibility
| | Result |
| --- | --- |
| Application code changed | **none.** `git diff --stat HEAD -- backend/app frontend` is empty |
| Schema migration | none |
| Bundle format | unchanged |
| Stored campaign behaviour | unchanged |
| Memory behaviour | **unchanged.** Creation, eviction, ranking, pins and lineage are all as in v1.0.0. `memorybank.py` is byte-identical to the tag |
| What changed | Diagnostic tooling (`tools/memory_diagnostic.py`, `tools/v11_b1_memory.py`), an opt-in harness mode (`m11_long_run.py --independent-fact`, with every existing behaviour and M04 verdict unchanged when the flag is off) and tests |
| New strict xfails | 2 (§D). They document the two diagnosed defects, and the suite stays green. B.2 must remove them deliberately when it fixes the mechanisms |
## R. Security / local-only
Checked against the B.1 diff: the `m11_long_run.py` changes plus the four new files.
| | Result |
| --- | --- |
| New network client, endpoint, URL or TLS setting | **none.** The only address in the new code is `http://127.0.0.1:9/v1`, a refused loopback port the deterministic scenarios configure so nothing is contacted |
| External embedding service or remote vector store | **none.** The deterministic runs use `ConceptEmbedder` in-process. The real-model run uses the configured Ollama embedding model through the application's existing provider |
| Endpoint policy (`endpoints.py`, ADR 011) and TLS (`tlstrust.py`) | unchanged; not in the diff |
| Network calls in the real-model run | only the configured Ollama host on the trusted LAN, through the application's own provider and probe paths |
| Real identifiers in committed files | none. Checked again at staging |
| Offline container regression | **not run.** The harness builds and runs a Docker container, and this session's permission policy refused it. B.1 changes no runtime code and nothing in the image (the image carries `backend/app` and `frontend/dist` only), so the container would be byte-identical to the A1/A2 image. That image passed 23/23 on the corrective tree (`V1.1-WP-A1-A2-REPORT.md` §R.3) |
---
## S. Final decision
**What was established.** Every B.1 requirement was carried out except the clean
real-model run, which was attempted:
- the current mechanism (§B);
- a deterministic, isolated harness with four stage outputs and a verdict at recall
depth ≥ 100 (§C);
- the v1.0.0 baseline (§D);
- capacity and eviction (§F), the creation window (§E) and ranking (§G);
- the lineage and authority negative controls (§I, §J);
- the `recovered_through_memory_independent` long-run verdict, with the M04 verdicts
unchanged;
- tests and regression (§P), compatibility (§Q) and security (§R).
**The real-model gap.** The one real-model attempt ran cleanly, but failed isolation
(`precondition_failed:absent_from_summary`). The narrator and summariser restated the
fact, so a clean memory-only result on a real model was **not obtained**. The owner
decided not to make a second attempt, because the same narrator behaviour would very
likely repeat. The attempt is reported in full, and its stage diagnosis is used as
mechanism evidence only (§K).
**B.1 DIAGNOSTIC: COMPLETE**
**FIRST FAILING STAGE: RANKING.** In a real 100-turn campaign the early fact's memory
was created (it carries the fact) and retained (active, 33 of 80), but ranked 10th of
19 (similarity 0.866) against `top_k` 4. Later memories with near-identical wording
outranked it, under a query made of recent narration (§K, §L). Two further failures
are proven deterministically, identically on v1.0.0:
- **retention**, past `memory_bank_capacity`: least-recently-used eviction removes an
early memory first;
- **creation**, for a fact early in a block over 2,000 tokens: the summariser's
last-2,000-token excerpt drops it.
**WP-B.2 RECOMMENDED CHANGE: retrieval ranking first.** Build the retrieval query so
the newest player input is not diluted by three turns of recent narration, and add one
inspectable lexical term for the query's rare words to cosine ranking. The score must
be recorded in `memories.used`. Acceptance: a real-model re-run shows `ranked: yes`
for the planting-era memory.
Then, in separate test-first steps:
1. make eviction coverage-aware instead of purely least-recently-used, so the only
memory of an early range is not discarded first (flips
`test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity`);
2. choose the summariser excerpt so a fact early in a long block is kept (flips
`test_acceptance_a_fact_early_in_a_long_block_is_remembered`).
Everything in §O stays unchanged. B.2 has not been started.