v1.1 WP-B.1: diagnose independent long-term memory retention
Diagnostic only; no memory behaviour changes. - tools/memory_diagnostic.py: planted-fact isolation checks, the four-stage diagnosis (created / retained / ranked / injected) with a verdict, a production-ranking replica, deterministic summariser/embedder/narrator stubs and seven scenarios (default, past capacity, pinned, low top_k, long-block early/late, lineage control) - tools/v11_b1_memory.py: CLI for the scenarios and for diagnosing a copy of a finished real campaign - tools/m11_long_run.py: opt-in --independent-fact mode with per-turn isolation tracking and the recovered_through_memory_independent verdict; M04 verdicts unchanged - tests: diagnostic stages, eviction, creation window, ranking, lineage and authority controls; two strict xfails record the diagnosed retention and creation defects for WP-B.2 to flip - planning/reports/v1.1/V1.1-WP-B1-REPORT.md First failing stage: ranking (real model); retention past capacity and creation for early facts in long blocks (deterministic, same on v1.0.0). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
d63804f22e
commit
beb17ada10
@@ -0,0 +1,86 @@
|
||||
"""v1.1 WP-B.1: the long run's `recovered_through_memory_independent` verdict.
|
||||
|
||||
The new verdict must never be reported when anything other than memory could
|
||||
have carried the fact. Each precondition is named when it fails. The existing M04
|
||||
verdicts keep their meaning exactly.
|
||||
|
||||
python -m pytest tests/test_v11_b1_long_run_verdict.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
from tools import m11_long_run as lr
|
||||
|
||||
GOOD = {
|
||||
"independent_planted_depth": 3,
|
||||
"planted_turn_outside_history": True,
|
||||
"absent_from_state": True,
|
||||
"absent_from_summary": True,
|
||||
"absent_from_knowledge": True,
|
||||
"absent_from_later_narration": True,
|
||||
"memory_covering_planting_carries_fact": True,
|
||||
"memory_forgotten": False,
|
||||
"memory_injected": True,
|
||||
}
|
||||
|
||||
|
||||
def test_every_precondition_and_an_injected_memory_is_the_new_verdict():
|
||||
assert lr._independent_memory_verdict(GOOD) == "recovered_through_memory_independent"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
|
||||
def test_a_failed_precondition_is_named_and_never_a_recovery(name):
|
||||
assert lr._independent_memory_verdict({**GOOD, name: False}) == f"precondition_failed:{name}"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
|
||||
def test_an_unmeasured_precondition_is_unknown_not_a_pass(name):
|
||||
assert lr._independent_memory_verdict({**GOOD, name: None}) == f"precondition_unknown:{name}"
|
||||
|
||||
|
||||
def test_no_planted_depth_is_unknown():
|
||||
assert lr._independent_memory_verdict({**GOOD, "independent_planted_depth": None}) == \
|
||||
"precondition_unknown:planted_depth"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("change, verdict", [
|
||||
({"memory_covering_planting_carries_fact": False}, "not_recovered:not_created"),
|
||||
({"memory_forgotten": True}, "not_recovered:evicted"),
|
||||
({"memory_injected": False}, "not_recovered:not_injected"),
|
||||
])
|
||||
def test_the_failing_memory_stage_is_named(change, verdict):
|
||||
assert lr._independent_memory_verdict({**GOOD, **change}) == verdict
|
||||
|
||||
|
||||
def test_preconditions_are_judged_before_memory():
|
||||
"""A carried fact disqualifies the run even when memory also failed."""
|
||||
both = {**GOOD, "absent_from_state": False, "memory_covering_planting_carries_fact": False}
|
||||
assert lr._independent_memory_verdict(both) == "precondition_failed:absent_from_state"
|
||||
|
||||
|
||||
def test_the_fact_is_matched_as_whole_words():
|
||||
assert lr._mentions_fact("She hid the amber Sundial.")
|
||||
assert lr._mentions_fact("a cracked TEAPOT on the shelf")
|
||||
assert not lr._mentions_fact("teapots") # a different word, not the fact's
|
||||
assert not lr._mentions_fact("the sun dialled down")
|
||||
|
||||
|
||||
def test_the_m04_verdicts_are_unchanged():
|
||||
base = {"planted_turn_in_history_window": False, "in_memories_section": False,
|
||||
"in_summary_section": False, "in_state_section": False}
|
||||
assert lr._m04_verdict(base) == "not_recovered"
|
||||
assert lr._m04_verdict({**base, "in_state_section": True}) == "recovered_through_state_only"
|
||||
assert lr._m04_verdict({**base, "in_memories_section": True}) == \
|
||||
"recovered_through_memory_or_summary"
|
||||
assert lr._m04_verdict({**base, "planted_turn_in_history_window": True}) == \
|
||||
"precondition_not_met"
|
||||
|
||||
|
||||
def test_the_independent_fact_is_not_in_any_imported_knowledge_file():
|
||||
for text in (lr.CANON_MD, lr.REFERENCE_MD, lr.INSPIRATION_MD, *lr.BEATS):
|
||||
assert not lr._mentions_fact(text)
|
||||
|
||||
|
||||
def test_the_planting_text_and_recall_carry_the_fact():
|
||||
assert lr._mentions_fact(lr.INDEPENDENT_FACT_TEXT)
|
||||
assert lr._mentions_fact(lr.INDEPENDENT_RECALL_TEXT)
|
||||
@@ -0,0 +1,308 @@
|
||||
"""v1.1 WP-B.1: the memory-retention diagnostic, deterministically.
|
||||
|
||||
B.1 changes no memory behaviour. These tests prove two things about the
|
||||
diagnostic in `tools/memory_diagnostic.py`:
|
||||
|
||||
1. **It measures what it claims.**
|
||||
- The fixture keeps the planted fact out of every layer except memory.
|
||||
- Each stage (created, retained, ranked, injected) is reported from the rows
|
||||
and the recall turn's own stored context.
|
||||
- Its ranking agrees with the selection production stored.
|
||||
2. **What it finds on this tree.** The scenarios run with a best-case summariser,
|
||||
one that keeps a fact if and only if the fact reached it. Any failure is
|
||||
therefore the application's mechanism, not a model's writing.
|
||||
- The criteria the current code does not meet are marked `xfail(strict=True)`,
|
||||
so B.2 has to flip them deliberately.
|
||||
- The same file is run unchanged against v1.0.0 for the baseline.
|
||||
|
||||
python -m pytest tests/test_v11_b1_memory_diagnostic.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
from tools import memory_diagnostic as md
|
||||
|
||||
_results: dict = {}
|
||||
|
||||
|
||||
def scenario(name: str) -> dict:
|
||||
"""Runs a named scenario once per session and keeps the result."""
|
||||
if name not in _results:
|
||||
_results[name] = md.run_scenario(md.SCENARIOS[name])
|
||||
return _results[name]
|
||||
|
||||
|
||||
# ------------------------------------------------------- fixture preconditions
|
||||
|
||||
def test_the_fact_is_planted_early_and_recalled_past_depth_one_hundred():
|
||||
result = scenario("independent_default")
|
||||
assert result["plant_depth"] is not None and result["plant_depth"] <= 3
|
||||
assert result["recall_depth"] >= 100
|
||||
|
||||
|
||||
@pytest.mark.parametrize("check", ["state_document", "state_snapshots", "later_narration",
|
||||
"summary", "knowledge", "recent_history", "state_section"])
|
||||
def test_no_layer_but_memory_carries_the_fact(check):
|
||||
"""A test where another layer carries F is not evidence about memory."""
|
||||
isolation = scenario("independent_default")["isolation"]
|
||||
assert isolation["checks"][check]["ok"], isolation["checks"][check]
|
||||
assert isolation["ok"]
|
||||
|
||||
|
||||
def test_the_isolation_check_fails_when_another_layer_carries_the_fact():
|
||||
"""The negative control for the precondition itself: a state fact naming F."""
|
||||
fact = md.FACT_F
|
||||
with SessionLocal() as db:
|
||||
Base.metadata.create_all(bind=engine)
|
||||
try:
|
||||
user = models.User(is_guest=False, email="b1-iso@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
adventure = models.Adventure(user_id=user.id, title="iso")
|
||||
adventure.narrative_state = {"facts": [{"id": "x", "predicate": "hidden",
|
||||
"value": "the amber sundial is in the teapot"}]}
|
||||
db.add(adventure)
|
||||
db.commit()
|
||||
result = md.isolation(db, adventure, fact, 1)
|
||||
assert result["ok"] is False
|
||||
assert result["checks"]["state_document"]["ok"] is False
|
||||
finally:
|
||||
db.close()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- stages
|
||||
|
||||
def test_creation_is_reported_with_the_covering_memory_and_what_the_summariser_saw():
|
||||
created = scenario("independent_default")["diagnosis"]["created"]
|
||||
assert created["yes"] is True
|
||||
assert created["source_start"] <= scenario("independent_default")["plant_depth"] <= created["source_end"]
|
||||
assert md.FACT_F.carried_by(created["memory_text"])
|
||||
covering = [c for c in created["covering_memories"] if c["memory_id"] == created["memory_id"]]
|
||||
assert covering and covering[0]["fact_in_block"] and covering[0]["fact_in_summariser_excerpt"]
|
||||
|
||||
|
||||
def test_retention_is_reported_with_the_bank_and_its_eviction_order():
|
||||
retained = scenario("independent_default")["diagnosis"]["retained"]
|
||||
assert retained["yes"] is True and retained["forgotten"] is False
|
||||
assert retained["on_active_lineage"] is True
|
||||
assert retained["active_memories"] <= retained["memory_bank_capacity"]
|
||||
assert retained["eviction_position"] is not None
|
||||
|
||||
|
||||
def test_ranking_is_production_ranking_and_agrees_with_the_stored_selection():
|
||||
ranked = scenario("independent_default")["diagnosis"]["ranked"]
|
||||
assert ranked["replica_matches_stored_selection"] is True
|
||||
assert ranked["lexical_score"] is None # memory ranking has no lexical term
|
||||
assert ranked["top_k_cutoff"] == 5
|
||||
assert ranked["yes"] is True and ranked["selected"] is True
|
||||
assert 1 <= ranked["rank"] <= ranked["top_k_cutoff"]
|
||||
# The production query is the newest four actions, cut to 600 tokens, and the
|
||||
# one-line question is diluted by the narration around it.
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert ranked["semantic_score"] < variants["direct"]["similarity"]
|
||||
|
||||
|
||||
def test_injection_is_read_from_the_recall_turns_own_context():
|
||||
diagnosis = scenario("independent_default")["diagnosis"]
|
||||
assert diagnosis["injected"]["yes"] is True
|
||||
assert diagnosis["injected"]["context_component"] == md.MEMORIES_LABEL
|
||||
assert diagnosis["injected"]["token_count"] > 0
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_ranking_variants_direct_paraphrase_and_unrelated():
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["direct"]["rank"] == 1 and variants["direct"]["selected"]
|
||||
assert variants["paraphrase"]["rank"] == 1 and variants["paraphrase"]["selected"]
|
||||
assert (variants["direct"]["similarity"] > variants["paraphrase"]["similarity"]
|
||||
> 5 * variants["unrelated"]["similarity"])
|
||||
|
||||
|
||||
def test_retrieval_fills_top_k_whatever_the_similarity():
|
||||
"""Diagnosis: there is no relevance floor. With more memories than
|
||||
`memory_top_k`, an unrelated query still selects five, and the early fact
|
||||
rides along at a similarity near zero."""
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["unrelated"]["similarity"] < 0.1
|
||||
assert variants["unrelated"]["selected"] is True
|
||||
|
||||
|
||||
# ---------------------------------------------------------- capacity/eviction
|
||||
|
||||
def test_past_capacity_the_early_memory_is_evicted_and_the_stage_says_so():
|
||||
"""Diagnosis, not a requirement: what the current eviction rule does to F."""
|
||||
result = scenario("past_capacity")
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
assert result["diagnosis"]["verdict"] == "created_but_evicted"
|
||||
eviction = result["eviction"]
|
||||
assert eviction["f_evicted_at_turn"] is not None
|
||||
# It was retrieved while the bank was small, stopped being retrieved once
|
||||
# recent narration filled the top-k, and was then the least recently used.
|
||||
assert eviction["f_use_count_when_evicted"] > 0
|
||||
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
|
||||
assert eviction["f_memory_was_first_evicted"] is True
|
||||
|
||||
|
||||
def test_at_a_lower_top_k_the_early_memory_ages_out_after_it_stops_being_retrieved():
|
||||
"""Diagnosis with most of the bank unretrieved on any turn, nearer the
|
||||
shipped 5-in-80 ratio. F is not simply the oldest row: it is evicted some
|
||||
turns after recent narration stopped pulling it into the top-k, which is
|
||||
what ordering by last use does to a fact nothing recent mentions."""
|
||||
result = scenario("past_capacity_low_top_k")
|
||||
eviction = result["eviction"]
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
assert result["diagnosis"]["verdict"] == "created_but_evicted"
|
||||
assert eviction["f_use_count_when_evicted"] > 0
|
||||
assert result["f_last_use_increase_turn"] < eviction["f_evicted_at_turn"]
|
||||
assert eviction["first_eviction_turn"] <= eviction["f_evicted_at_turn"]
|
||||
assert eviction["created_and_evicted_same_turn"] == []
|
||||
|
||||
|
||||
def test_no_memory_is_evicted_by_the_same_pass_that_created_it():
|
||||
"""The frozen-bank regression the current rule fixed, still holding."""
|
||||
for name in ("past_capacity", "past_capacity_pinned"):
|
||||
assert scenario(name)["eviction"]["created_and_evicted_same_turn"] == []
|
||||
|
||||
|
||||
def test_a_pinned_memory_survives_capacity():
|
||||
eviction = scenario("past_capacity_pinned")["eviction"]
|
||||
assert eviction["pinned_memory_id"] is not None
|
||||
assert eviction["pinned_memory_forgotten"] is False
|
||||
|
||||
|
||||
@pytest.mark.xfail(strict=True, reason=(
|
||||
"WP-B.1 diagnosis on this tree: past memory_bank_capacity the planting-era "
|
||||
"memory is evicted first, because it was never retrieved and eviction orders "
|
||||
"by last use, then creation. B.2 must flip this deliberately."))
|
||||
def test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity():
|
||||
assert scenario("past_capacity")["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ---------------------------------------------------------- creation window
|
||||
|
||||
def test_a_fact_early_in_a_long_block_never_reaches_the_summariser():
|
||||
result = scenario("long_block_fact_early")
|
||||
created = result["diagnosis"]["created"]
|
||||
covering = created["covering_memories"]
|
||||
assert covering, "the long block must have been summarised"
|
||||
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert covering[0]["fact_in_block"] is True
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is False
|
||||
assert result["diagnosis"]["verdict"] == "not_created"
|
||||
|
||||
|
||||
def test_the_same_fact_late_in_the_same_sized_block_does():
|
||||
result = scenario("long_block_fact_late")
|
||||
covering = result["diagnosis"]["created"]["covering_memories"]
|
||||
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is True
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
|
||||
|
||||
@pytest.mark.xfail(strict=True, reason=(
|
||||
"WP-B.1 diagnosis on this tree: the summariser reads only the last "
|
||||
f"{memorybank.MEMORY_EXCERPT_TOKENS} tokens of a block, so a fact early in a "
|
||||
"long block is never seen. B.2 must flip this deliberately."))
|
||||
def test_acceptance_a_fact_early_in_a_long_block_is_remembered():
|
||||
assert scenario("long_block_fact_early")["diagnosis"]["created"]["yes"] is True
|
||||
|
||||
|
||||
# ------------------------------------------------------- lineage control (G)
|
||||
|
||||
def test_an_abandoned_lines_memory_is_stored_but_never_eligible_or_injected():
|
||||
g = scenario("lineage_control")["lineage_control"]
|
||||
assert g["memory_ids"], "G's memory must exist on line A before it is abandoned"
|
||||
assert sorted(g["stored"]) == sorted(g["memory_ids"])
|
||||
assert g["eligible_on_active_line"] == []
|
||||
assert g["g_text_ever_in_used_memories"] is False
|
||||
# Any turn that did name G's memory was on line A, before the divergence.
|
||||
assert g["eligible_after_returning_to_line_a"] == g["memory_ids"]
|
||||
|
||||
|
||||
def test_the_lineage_scenario_still_diagnoses_f_on_the_active_line():
|
||||
result = scenario("lineage_control")
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
assert result["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ----------------------------------------------------- authority control
|
||||
|
||||
@pytest.fixture()
|
||||
def authority_client(monkeypatch):
|
||||
embedder = md.ConceptEmbedder()
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email="b1-auth@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(user_id=user.id, model="script",
|
||||
endpoint_url="http://127.0.0.1:9/v1",
|
||||
embedding_model="concept-embed", memory_top_k=5))
|
||||
adventure = models.Adventure(user_id=user.id, title="auth", memory_bank_enabled=True,
|
||||
auto_summarize=True)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(adventure_id=adventure.id, type="start", text="The tavern at dusk."))
|
||||
db.commit()
|
||||
adv, user_id = adventure.id, user.id
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", md.ScriptNarrator)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: embedder)
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: md.BestCaseSummariser())
|
||||
monkeypatch.setattr(memorybank, "schedule_post_turn", lambda a: None)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id))
|
||||
client = TestClient(app)
|
||||
client.adv = adv
|
||||
try:
|
||||
yield client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def test_a_memory_that_contradicts_state_loses_and_changes_nothing(authority_client):
|
||||
client, adv = authority_client, authority_client.adv
|
||||
corrected = client.post(f"/api/adventures/{adv}/state/corrections", json={"events": [
|
||||
{"type": "add_fact", "predicate": "the tavern lamp is lit", "fact_id": "lamp-lit"}]})
|
||||
assert corrected.status_code in (200, 201), corrected.text[:300]
|
||||
made = client.post(f"/api/adventures/{adv}/memories",
|
||||
json={"text": "The tavern lamp was never lit that night."})
|
||||
assert made.status_code == 201, made.text[:300]
|
||||
client.patch(f"/api/adventures/{adv}/memories/{made.json()['id']}", json={"pinned": True})
|
||||
asyncio.run(memorybank.run_post_turn(adv)) # embed it
|
||||
|
||||
before = client.get(f"/api/adventures/{adv}/state").json()["document"]
|
||||
md.ScriptNarrator.next_reply = 'The fire crackles.\n```state\n{"events": []}\n```'
|
||||
played = client.post(f"/api/adventures/{adv}/actions",
|
||||
json={"type": "do", "text": "I look at the lamp."})
|
||||
assert played.status_code == 200 and '"type": "error"' not in played.text
|
||||
after = client.get(f"/api/adventures/{adv}/state").json()["document"]
|
||||
assert after == before # retrieval mutated no state
|
||||
|
||||
with SessionLocal() as db:
|
||||
action = (db.query(models.Action).filter_by(adventure_id=adv, type="ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
snapshot = action.context_snapshot
|
||||
state_text = md._section(snapshot, md.STATE_LABEL)
|
||||
memory_text = md._section(snapshot, md.MEMORIES_LABEL)
|
||||
assert "the tavern lamp is lit" in state_text
|
||||
assert "never lit" in memory_text
|
||||
assert memory_text.startswith("Memories from earlier in the story")
|
||||
labels = [s["label"] for s in snapshot["sections"]]
|
||||
# State is read last of the live sections: it settles the conflict.
|
||||
assert labels.index(md.STATE_LABEL) > labels.index(md.MEMORIES_LABEL)
|
||||
@@ -190,6 +190,31 @@ CLUE_FACT = {
|
||||
"fact_id": "silver-key-opens-crypt",
|
||||
}
|
||||
|
||||
#: v1.1 WP-B.1: a second planted fact, established in the **story only**.
|
||||
#:
|
||||
#: The M04 clue above is planted as accepted state, and memories are written
|
||||
#: from story text, so no memory could ever carry it on its own. That is why
|
||||
#: every M04 recovery so far ran through state. This fact is told to the reader
|
||||
#: in narration and never corrected into state, so memory is the only layer that
|
||||
#: is meant to carry it. `--independent-fact` plants it and reports
|
||||
#: `recovered_through_memory_independent` only when every other layer is proven
|
||||
#: not to carry it. The words are copied from `tools/memory_diagnostic.FACT_F`,
|
||||
#: for the reason `HISTORY_LABELS` is copied.
|
||||
INDEPENDENT_FACT_TEXT = ("I watch Mara slip the amber sundial inside the cracked teapot on "
|
||||
"the tavern's top shelf, and she makes me promise to tell no one.")
|
||||
INDEPENDENT_FACT_TERMS = ("sundial", "teapot")
|
||||
INDEPENDENT_RECALL_TEXT = "I ask Mara quietly where she hid the amber sundial."
|
||||
#: How far past the planting turn its memory block can reach. Narration inside
|
||||
#: that block may repeat the fact; narration after it may not.
|
||||
INDEPENDENT_BLOCK_SLACK = 6
|
||||
INDEPENDENT_PRECONDITIONS = (
|
||||
"planted_turn_outside_history",
|
||||
"absent_from_state",
|
||||
"absent_from_summary",
|
||||
"absent_from_knowledge",
|
||||
"absent_from_later_narration",
|
||||
)
|
||||
|
||||
CANON = [
|
||||
"The dead do not return. No rite, relic or bargain has ever returned anyone.",
|
||||
"The abbey crypt has been sealed since the founding.",
|
||||
@@ -365,6 +390,14 @@ class Run:
|
||||
#: The depth of the player turn that planted the clue. M04's
|
||||
#: precondition is that this turn has left the history window.
|
||||
self.planted_depth: int | None = None
|
||||
#: v1.1 WP-B.1, with --independent-fact: where the story-only fact was
|
||||
#: planted, and the accepted-turn count at which each isolation
|
||||
#: precondition first failed.
|
||||
self.independent_fact = False
|
||||
self.independent_depth: int | None = None
|
||||
self.independent_violations: dict[str, int] = {}
|
||||
self.last_done: dict = {}
|
||||
self.last_report: dict = {}
|
||||
|
||||
# ------------------------------------------------------------ recording
|
||||
|
||||
@@ -393,6 +426,8 @@ class Run:
|
||||
"turns_target": self.turns_target,
|
||||
"log_offset": self.log_offset,
|
||||
"planted_depth": self.planted_depth,
|
||||
"independent_depth": self.independent_depth,
|
||||
"independent_violations": self.independent_violations,
|
||||
"written": datetime.now().isoformat(timespec="seconds"),
|
||||
}
|
||||
tmp = self.out / (RESUME_FILE + ".tmp")
|
||||
@@ -414,6 +449,8 @@ class Run:
|
||||
self.elapsed_before = prior.get("elapsed_seconds", 0)
|
||||
self.log_offset = prior.get("log_offset", 0)
|
||||
self.planted_depth = prior.get("planted_depth")
|
||||
self.independent_depth = prior.get("independent_depth")
|
||||
self.independent_violations = dict(prior.get("independent_violations") or {})
|
||||
self.resumed = True
|
||||
|
||||
def reattach(self) -> None:
|
||||
@@ -570,15 +607,44 @@ class Run:
|
||||
"observed_margin": accounting.get("observed_margin"),
|
||||
"safety_reserve": accounting.get("safety_reserve"),
|
||||
})
|
||||
self.last_done = done
|
||||
if self.independent_fact and self.independent_depth is not None:
|
||||
self._check_independent_isolation(done, sample)
|
||||
self.note("turn", text=text, seconds=round(seconds, 1), **sample)
|
||||
return {"accepted": True, "seconds": seconds, **sample}
|
||||
|
||||
def _check_independent_isolation(self, done: dict, sample: dict) -> None:
|
||||
"""v1.1 WP-B.1: does anything but memory carry the story-only fact yet?
|
||||
|
||||
Checked on every accepted turn, so a run knows the first turn at which
|
||||
the experiment stopped being about memory, instead of finding out at
|
||||
recall. Each precondition records only its first failure.
|
||||
"""
|
||||
depth = sample.get("total_actions", 0) - 1
|
||||
text = (done.get("action") or {}).get("text") or ""
|
||||
found = {}
|
||||
if depth > self.independent_depth + INDEPENDENT_BLOCK_SLACK and _mentions_fact(text):
|
||||
found["absent_from_later_narration"] = f"narration at depth {depth}"
|
||||
document = self.state().get("document") or {}
|
||||
if _mentions_fact(json.dumps(document)):
|
||||
found["absent_from_state"] = "the narrative state names the fact"
|
||||
summary = next((sec.get("text", "") for sec in (self.last_report.get("sections") or [])
|
||||
if sec.get("label") == SUMMARY_LABEL), "")
|
||||
if _mentions_fact(summary):
|
||||
found["absent_from_summary"] = "the active summary names the fact"
|
||||
for name, detail in found.items():
|
||||
if name not in self.independent_violations:
|
||||
self.independent_violations[name] = self.accepted
|
||||
self.note("independent_precondition_failed", precondition=name, detail=detail)
|
||||
sample["independent_violations"] = dict(self.independent_violations)
|
||||
|
||||
def count_actions(self) -> int:
|
||||
return self.server.call("GET", f"/adventures/{self.adv}/actions?limit=1")["total"]
|
||||
|
||||
def measure(self) -> dict:
|
||||
"""M03's numbers, read from the prompt the app would send right now."""
|
||||
report = self.server.call("GET", f"/adventures/{self.adv}/context")
|
||||
self.last_report = report
|
||||
tokens = report["tokens"]
|
||||
sections = {s["label"]: s["tokens"] for s in report["sections"]}
|
||||
window = report.get("window") or {}
|
||||
@@ -714,6 +780,10 @@ def main() -> int:
|
||||
"--max-consecutive-failures", type=int,
|
||||
default=DEFAULT_MAX_CONSECUTIVE_FAILURES,
|
||||
help="stop and write the evidence after this many unaccepted turns")
|
||||
parser.add_argument(
|
||||
"--independent-fact", action="store_true",
|
||||
help=("v1.1 WP-B.1: also plant a story-only fact at depth 3 and report "
|
||||
"whether memory alone recovers it"))
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL and EMBED_MODEL):
|
||||
@@ -746,6 +816,7 @@ def main() -> int:
|
||||
server.start()
|
||||
run = Run(server, out, turns_target=args.turns,
|
||||
turn_timeout=args.turn_timeout)
|
||||
run.independent_fact = args.independent_fact
|
||||
if prior:
|
||||
run.adopt(prior)
|
||||
|
||||
@@ -782,6 +853,18 @@ def main() -> int:
|
||||
"the planted clue is not in accepted state, so M04 cannot "
|
||||
"be measured from this run. Stopping before the campaign "
|
||||
"starts rather than reporting a recall failure later.")
|
||||
if args.independent_fact:
|
||||
# v1.1 WP-B.1: the story-only fact, told in the next turn and
|
||||
# never corrected into state. Depth 3: the opening, the clue turn
|
||||
# and its reply come first.
|
||||
if any(_mentions_fact(md) for md in (CANON_MD, REFERENCE_MD, INSPIRATION_MD)):
|
||||
raise SystemExit("the imported knowledge names the independent fact")
|
||||
planting_f = run.turn(INDEPENDENT_FACT_TEXT)
|
||||
if not planting_f.get("accepted"):
|
||||
raise SystemExit("the turn that plants the independent fact was not accepted")
|
||||
run.independent_depth = planting_f["total_actions"] - 2
|
||||
run.note("independent_fact_planted", depth=run.independent_depth,
|
||||
terms=list(INDEPENDENT_FACT_TERMS))
|
||||
# The first checkpoint, and the point from which --resume works: the
|
||||
# campaign exists and its clue is planted.
|
||||
run.save_resume()
|
||||
@@ -846,10 +929,15 @@ def main() -> int:
|
||||
# Skipped on an aborted run: it asks the narrator a question, and the
|
||||
# reason the run stopped is that the narrator does not answer.
|
||||
recall = None
|
||||
independent = None
|
||||
if aborted is None:
|
||||
run.note("recall_begin")
|
||||
recall = _recall(run)
|
||||
(out / "recall.json").write_text(json.dumps(recall, indent=2))
|
||||
if args.independent_fact and run.independent_depth is not None:
|
||||
independent = _independent_recall(run, out / "campaign.db")
|
||||
(out / "recall-independent.json").write_text(json.dumps(independent, indent=2))
|
||||
run.note("independent_recall", verdict=independent["verdict"])
|
||||
|
||||
# ---- Export whatever exists, for the recovery evidence. ----
|
||||
# Attempted even for an aborted run: the recovery check and the storage
|
||||
@@ -897,6 +985,7 @@ def main() -> int:
|
||||
"elapsed_seconds": run.elapsed(),
|
||||
"turn_timeout_seconds": args.turn_timeout,
|
||||
"recall": recall,
|
||||
"independent_recall": independent,
|
||||
"final_state": _or_none(lambda: run.state()["document"]),
|
||||
"final_measurement": _or_none(run.measure),
|
||||
"db_bytes": db_path.stat().st_size,
|
||||
@@ -1224,6 +1313,123 @@ def _m04_verdict(recall: dict) -> str:
|
||||
return "not_recovered"
|
||||
|
||||
|
||||
def _mentions_fact(text: str | None) -> bool:
|
||||
"""v1.1 WP-B.1: whether `text` names the independent fact, as a whole word."""
|
||||
low = (text or "").lower()
|
||||
return any(re.search(rf"(?<![a-z]){term}(?![a-z])", low) for term in INDEPENDENT_FACT_TERMS)
|
||||
|
||||
|
||||
def _independent_memory_verdict(check: dict) -> str:
|
||||
"""v1.1 WP-B.1: whether memory alone recovered the story-only fact.
|
||||
|
||||
`recovered_through_memory_independent` requires every precondition, so no
|
||||
other layer could have carried the fact. It also requires that a memory
|
||||
covering the planting turn carries the fact and was injected into the recall
|
||||
turn. A failed precondition is named and is never a recovery, and the M04
|
||||
verdicts above are untouched.
|
||||
"""
|
||||
if check.get("independent_planted_depth") is None:
|
||||
return "precondition_unknown:planted_depth"
|
||||
for name in INDEPENDENT_PRECONDITIONS:
|
||||
value = check.get(name)
|
||||
if value is None:
|
||||
return f"precondition_unknown:{name}"
|
||||
if not value:
|
||||
return f"precondition_failed:{name}"
|
||||
if not check.get("memory_covering_planting_carries_fact"):
|
||||
return "not_recovered:not_created"
|
||||
if check.get("memory_forgotten"):
|
||||
return "not_recovered:evicted"
|
||||
if not check.get("memory_injected"):
|
||||
return "not_recovered:not_injected"
|
||||
return "recovered_through_memory_independent"
|
||||
|
||||
|
||||
def _independent_recall(run: "Run", db_path: Path) -> dict:
|
||||
"""v1.1 WP-B.1: ask for the story-only fact, and find out which layer answered.
|
||||
|
||||
The prompt-level facts come from the recall turn's own stored context. The
|
||||
memory rows come from the campaign database, read-only. Ranking is not
|
||||
recomputed here, because that needs the embedding model;
|
||||
`tools/v11_b1_memory.py diagnose` does it afterwards against a copy of the
|
||||
database.
|
||||
"""
|
||||
import sqlite3
|
||||
import zlib
|
||||
|
||||
result = run.turn(INDEPENDENT_RECALL_TEXT)
|
||||
action_id = (run.last_done.get("action") or {}).get("id")
|
||||
snapshot = (run.server.call("GET", f"/adventures/{run.adv}/actions/{action_id}/context")
|
||||
if result.get("accepted") and action_id else {}) or {}
|
||||
sections = {}
|
||||
for sec in snapshot.get("sections") or []:
|
||||
sections.setdefault(sec.get("label"), []).append(sec.get("text", ""))
|
||||
text_of = {label: "\n".join(parts) for label, parts in sections.items()}
|
||||
floor = (snapshot.get("history") or {}).get("floor_depth")
|
||||
depth = run.independent_depth
|
||||
used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
|
||||
|
||||
covering = []
|
||||
connection = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True)
|
||||
try:
|
||||
rows = connection.execute(
|
||||
"SELECT id, text, source_start, source_end, forgotten, pinned, use_count, "
|
||||
"branch_id, depth FROM memories WHERE adventure_id = ? AND source_start <= ? "
|
||||
"AND source_end >= ? ORDER BY id", (run.adv, depth, depth)).fetchall()
|
||||
blob = connection.execute(
|
||||
"SELECT context_snapshot FROM actions WHERE id = ?", (action_id or -1,)).fetchone()
|
||||
finally:
|
||||
connection.close()
|
||||
for row in rows:
|
||||
memory_id, text, start, end, forgotten, pinned, use_count, branch_id, node_depth = row
|
||||
covering.append({
|
||||
"memory_id": memory_id, "text": text, "source_start": start, "source_end": end,
|
||||
"forgotten": bool(forgotten), "pinned": bool(pinned), "use_count": use_count,
|
||||
"branch_id": branch_id, "depth": node_depth,
|
||||
"carries_fact": all(re.search(rf"(?<![a-z]){t}(?![a-z])", (text or "").lower())
|
||||
for t in INDEPENDENT_FACT_TERMS),
|
||||
"injected": memory_id in used,
|
||||
})
|
||||
carrying = [c for c in covering if c["carries_fact"]]
|
||||
best = next((c for c in carrying if c["injected"]), carrying[0] if carrying else None)
|
||||
stored_snapshot_readable = blob is not None and blob[0] is not None
|
||||
if stored_snapshot_readable:
|
||||
try:
|
||||
json.loads(zlib.decompress(blob[0]))
|
||||
except Exception: # noqa: BLE001
|
||||
stored_snapshot_readable = False
|
||||
|
||||
document = run.state().get("document") or {}
|
||||
violations = dict(run.independent_violations)
|
||||
check = {
|
||||
"independent_planted_depth": depth,
|
||||
"recall_accepted": bool(result.get("accepted")),
|
||||
"history_floor_depth": floor,
|
||||
"planted_turn_outside_history": (None if not snapshot else
|
||||
floor is not None and depth < floor),
|
||||
"absent_from_state": ("absent_from_state" not in violations
|
||||
and not _mentions_fact(json.dumps(document))
|
||||
and not _mentions_fact(text_of.get(STATE_LABEL))),
|
||||
"absent_from_summary": ("absent_from_summary" not in violations
|
||||
and not _mentions_fact(text_of.get(SUMMARY_LABEL))),
|
||||
"absent_from_knowledge": not any(
|
||||
_mentions_fact(text_of.get(label)) for label in IMPORTED_KNOWLEDGE_LABELS),
|
||||
"absent_from_later_narration": "absent_from_later_narration" not in violations,
|
||||
"violations_first_turn": violations,
|
||||
"covering_memories": covering,
|
||||
"memory_covering_planting_carries_fact": bool(carrying),
|
||||
"memory_forgotten": bool(best and best["forgotten"]),
|
||||
"memory_injected": bool(best and best["injected"]),
|
||||
"memory_text_in_memories_section": bool(
|
||||
best and best["text"] and best["text"] in (text_of.get(MEMORIES_LABEL) or "")),
|
||||
"memory_ids_used": used,
|
||||
"recall_action_id": action_id,
|
||||
"stored_snapshot_readable": stored_snapshot_readable,
|
||||
}
|
||||
check["verdict"] = _independent_memory_verdict(check)
|
||||
return check
|
||||
|
||||
|
||||
#: Signs the application stored protocol as story. The first is a state-section
|
||||
#: heading with an indented entry under it, in any markdown, because
|
||||
#: `## Established:` got past a plain substring match and the count read 1
|
||||
|
||||
@@ -0,0 +1,839 @@
|
||||
"""v1.1 WP-B.1: where an early story fact is lost on its way to the narrator.
|
||||
|
||||
One planted fact **F** has four stages to survive before the narrator can use
|
||||
it from memory, and this module reports each one separately:
|
||||
|
||||
created a memory whose `source_start`..`source_end` covers the planting
|
||||
depth carries F
|
||||
retained that memory is not `forgotten`
|
||||
ranked it is eligible on the active lineage and embedded, and where it
|
||||
scores for the recall query against `memory_top_k`
|
||||
injected the recall turn's own stored `memories.used` names it, and its text
|
||||
is in that turn's `used_memories` section
|
||||
|
||||
A fact is only evidence about memory if memory is the **only** thing carrying it.
|
||||
`isolation()` checks every other layer: the authoritative document, per-node state
|
||||
snapshots, the active summary, imported knowledge, the narration after the
|
||||
planting block, and the recent-history window. A run where any of those carries F
|
||||
is reported as a failed precondition, never as a memory result.
|
||||
|
||||
**Nothing here changes behaviour.**
|
||||
- It reads rows.
|
||||
- It reuses production's own pure helpers (`memorybank._drop_redundant`,
|
||||
`memorybank.classify_authority`, `vectors.cosine`, `lineage.path_of`), so its
|
||||
ranking is production's ranking, not a second opinion.
|
||||
- It checks itself against what the recall turn actually recorded.
|
||||
- The only computed fields are ephemeral report data. No column or table is
|
||||
added.
|
||||
|
||||
The deterministic stubs at the bottom stand in for the models when a test needs a
|
||||
fixed answer. **Read what they model before reading any result they produce:**
|
||||
|
||||
- `BestCaseSummariser` keeps F if and only if F is in the excerpt it is given.
|
||||
It is the ideal summariser, so a creation failure under it is the
|
||||
application's, not the model's.
|
||||
- `ConceptEmbedder` maps words to a small concept table, so that "the brass dial
|
||||
that tells the hour" lands near "sundial". It models what an embedding is
|
||||
supposed to do. It says nothing about how well `nomic-embed-text` does it,
|
||||
which is what the real-model run is for.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import math
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from sqlalchemy import select
|
||||
|
||||
from app import memorybank, models, summaries, vectors
|
||||
from app.context import builder, history, lineage
|
||||
from app.knowledge import classes as knowledge_classes
|
||||
|
||||
VERDICTS = (
|
||||
"not_created",
|
||||
"created_but_evicted",
|
||||
"retained_but_not_ranked",
|
||||
"ranked_but_not_selected",
|
||||
"selected_but_not_injected",
|
||||
"injected",
|
||||
)
|
||||
|
||||
#: Section labels in a stored context snapshot. Copied from the builder's
|
||||
#: vocabulary so a renamed section fails loudly here.
|
||||
HISTORY_LABELS = ("history", "recent_history")
|
||||
SUMMARY_LABEL = "story_summary"
|
||||
MEMORIES_LABEL = "used_memories"
|
||||
STATE_LABEL = "narrative_state"
|
||||
KNOWLEDGE_LABELS = (
|
||||
knowledge_classes.SECTION_CANON,
|
||||
knowledge_classes.SECTION_REFERENCE,
|
||||
knowledge_classes.SECTION_INSPIRATION,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Fact:
|
||||
"""A planted fact, and how to recognise it in a text.
|
||||
|
||||
`carry_groups`: a text carries the fact when every group matches, where a
|
||||
group matches when any one of its terms appears as a whole word. A memory has
|
||||
to name both the thing and where it is to carry "where the thing is".
|
||||
|
||||
`leak_terms`: any one of these in another layer means that layer carries the
|
||||
fact. This is deliberately looser than `carry_groups`. For isolation, a
|
||||
mention is enough to disqualify.
|
||||
"""
|
||||
|
||||
fact_id: str
|
||||
sentence: str
|
||||
carry_groups: tuple[tuple[str, ...], ...]
|
||||
leak_terms: tuple[str, ...]
|
||||
|
||||
def carried_by(self, text: str | None) -> bool:
|
||||
low = (text or "").lower()
|
||||
return all(any(_has_word(low, term) for term in group) for group in self.carry_groups)
|
||||
|
||||
def mentioned_by(self, text: str | None) -> bool:
|
||||
low = (text or "").lower()
|
||||
return any(_has_word(low, term) for term in self.leak_terms)
|
||||
|
||||
|
||||
def _has_word(low: str, term: str) -> bool:
|
||||
return re.search(rf"(?<![a-z]){re.escape(term.lower())}(?![a-z])", low) is not None
|
||||
|
||||
|
||||
#: The fixture's planted fact. Chosen to be natural in a tavern scene and absent
|
||||
#: from every existing fixture: no "sundial" or "teapot" appears anywhere in the
|
||||
#: Westhaven campaign, its knowledge files or its beats.
|
||||
FACT_F = Fact(
|
||||
fact_id="F-amber-sundial",
|
||||
sentence="Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf.",
|
||||
carry_groups=(("sundial",), ("teapot",)),
|
||||
leak_terms=("sundial", "teapot"),
|
||||
)
|
||||
#: The abandoned-line control fact.
|
||||
FACT_G = Fact(
|
||||
fact_id="G-iron-weathervane",
|
||||
sentence="Edrin buried the iron weathervane beneath the mill's broken waterwheel.",
|
||||
carry_groups=(("weathervane",), ("waterwheel",)),
|
||||
leak_terms=("weathervane", "waterwheel"),
|
||||
)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ reading
|
||||
|
||||
def _lineage_actions(db, adventure):
|
||||
path = lineage.path_of(db, adventure)
|
||||
return (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adventure.id, path.clause(models.Action))
|
||||
.order_by(models.Action.depth, models.Action.id)
|
||||
.all()
|
||||
)
|
||||
|
||||
|
||||
def covering_memories(db, adventure, depth: int, *, any_branch: bool = False):
|
||||
"""Memories whose source range covers `depth`, oldest first."""
|
||||
query = select(models.Memory).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.source_start <= depth,
|
||||
models.Memory.source_end >= depth,
|
||||
)
|
||||
if not any_branch:
|
||||
query = query.where(lineage.path_of(db, adventure).clause(models.Memory))
|
||||
return db.execute(query.order_by(models.Memory.id)).scalars().all()
|
||||
|
||||
|
||||
def planting_block_end(db, adventure, plant_depth: int) -> int:
|
||||
"""The last depth of the memory block holding the planted turn.
|
||||
|
||||
Taken from the memory that covers it where one exists. Before one exists it
|
||||
is the furthest a block could reach, so a later-narration check never counts
|
||||
a turn inside the planting block as a repetition.
|
||||
"""
|
||||
rows = covering_memories(db, adventure, plant_depth)
|
||||
if rows:
|
||||
return max(row.source_end for row in rows)
|
||||
return plant_depth + memorybank.MEMORY_INTERVAL
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- isolation
|
||||
|
||||
def isolation(db, adventure, fact: Fact, plant_depth: int, *,
|
||||
recall_snapshot: dict | None = None,
|
||||
recall_depth: int | None = None) -> dict:
|
||||
"""Every layer other than memory that could carry F, checked.
|
||||
|
||||
Returns `{check: {"ok": bool, "detail": str}}` and `ok` over all of them.
|
||||
With `recall_snapshot`, the recall turn's stored context, the prompt-level
|
||||
checks (history window, summary section, knowledge sections) are made
|
||||
against what the narrator was actually given.
|
||||
"""
|
||||
checks: dict[str, dict] = {}
|
||||
|
||||
document = adventure.narrative_state or {}
|
||||
hits = [key for key in ("entities", "facts", "relationships", "threads", "scene",
|
||||
"possessions")
|
||||
if fact.mentioned_by(json.dumps(document.get(key), default=str))]
|
||||
checks["state_document"] = {
|
||||
"ok": not hits and not fact.mentioned_by(json.dumps(document, default=str)),
|
||||
"detail": f"mentioned in {hits}" if hits else "absent",
|
||||
}
|
||||
|
||||
snapshot_hits = []
|
||||
later_hits = []
|
||||
block_end = planting_block_end(db, adventure, plant_depth)
|
||||
for action in _lineage_actions(db, adventure):
|
||||
if fact.mentioned_by(json.dumps(action.narrative_state_after, default=str)):
|
||||
snapshot_hits.append(action.depth)
|
||||
if (action.type == "ai" and action.depth is not None and action.depth > block_end
|
||||
and (recall_depth is None or action.depth < recall_depth)
|
||||
and fact.mentioned_by(action.text)):
|
||||
later_hits.append(action.depth)
|
||||
checks["state_snapshots"] = {
|
||||
"ok": not snapshot_hits,
|
||||
"detail": f"mentioned in snapshots at depths {snapshot_hits[:10]}" if snapshot_hits
|
||||
else "absent from every node's narrative_state_after on the active lineage",
|
||||
}
|
||||
checks["later_narration"] = {
|
||||
"ok": not later_hits,
|
||||
"detail": (f"narration after the planting block (ends at depth {block_end}) "
|
||||
f"mentions the fact at depths {later_hits[:10]}") if later_hits
|
||||
else f"no narrator turn after depth {block_end} mentions the fact",
|
||||
}
|
||||
|
||||
active = summaries.current(db, adventure)
|
||||
summary_text = active.text if active is not None else ""
|
||||
if recall_snapshot is not None:
|
||||
summary_text += "\n" + _section(recall_snapshot, SUMMARY_LABEL)
|
||||
checks["summary"] = {
|
||||
"ok": not fact.mentioned_by(summary_text),
|
||||
"detail": "the active summary mentions the fact" if fact.mentioned_by(summary_text)
|
||||
else ("absent from the active summary" if active is not None else "no summary yet"),
|
||||
}
|
||||
|
||||
sources = db.execute(
|
||||
select(models.KnowledgeSource.content).where(
|
||||
models.KnowledgeSource.adventure_id == adventure.id)
|
||||
).scalars().all()
|
||||
knowledge_text = "\n".join(s or "" for s in sources)
|
||||
if recall_snapshot is not None:
|
||||
knowledge_text += "\n" + "\n".join(_section(recall_snapshot, l) for l in KNOWLEDGE_LABELS)
|
||||
checks["knowledge"] = {
|
||||
"ok": not fact.mentioned_by(knowledge_text),
|
||||
"detail": "imported knowledge mentions the fact" if fact.mentioned_by(knowledge_text)
|
||||
else f"absent from {len(sources)} imported source(s)",
|
||||
}
|
||||
|
||||
if recall_snapshot is not None:
|
||||
hist = recall_snapshot.get("history") or {}
|
||||
floor = hist.get("floor_depth")
|
||||
history_text = "\n".join(_section(recall_snapshot, l) for l in HISTORY_LABELS)
|
||||
outside = floor is not None and plant_depth < floor
|
||||
checks["recent_history"] = {
|
||||
"ok": outside and not fact.carried_by(history_text),
|
||||
"detail": (f"history window starts at depth {floor}; planted at {plant_depth}; "
|
||||
f"fact text in history sections: {fact.carried_by(history_text)}"),
|
||||
}
|
||||
checks["state_section"] = {
|
||||
"ok": not fact.mentioned_by(_section(recall_snapshot, STATE_LABEL)),
|
||||
"detail": "the recall prompt's narrative_state section "
|
||||
+ ("mentions the fact" if fact.mentioned_by(_section(recall_snapshot, STATE_LABEL))
|
||||
else "does not mention the fact"),
|
||||
}
|
||||
|
||||
return {"ok": all(c["ok"] for c in checks.values()), "checks": checks}
|
||||
|
||||
|
||||
def _section(snapshot: dict, label: str) -> str:
|
||||
return "\n".join(s.get("text", "") for s in (snapshot.get("sections") or [])
|
||||
if s.get("label") == label)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- stages
|
||||
|
||||
async def rank_bank(db, adventure, settings, query: str, embed) -> dict:
|
||||
"""Production's ranking, recomputed for `query`, for every eligible memory.
|
||||
|
||||
The same catalogue clause, the same cosine, the same pin rule, the same
|
||||
redundancy suppression helper. Returns every scored row, not just the top-k,
|
||||
because "where did F rank" is the question.
|
||||
"""
|
||||
catalogue = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.authority,
|
||||
models.Memory.embedding_blob).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory),
|
||||
models.Memory.forgotten.is_(False),
|
||||
models.Memory.embedded.is_(True),
|
||||
)
|
||||
).all()
|
||||
if not catalogue or not query.strip():
|
||||
return {"query": query, "scored": [], "selected": [], "top_k": settings.memory_top_k}
|
||||
[query_vec] = await embed([query])
|
||||
held = {row.id: vectors.unpack(row.embedding_blob) for row in catalogue if row.embedding_blob}
|
||||
authority_of = {row.id: row.authority for row in catalogue}
|
||||
scored = sorted(
|
||||
((vectors.cosine(query_vec, held[row.id]), row.id, row.pinned)
|
||||
for row in catalogue if row.id in held),
|
||||
key=lambda r: r[0], reverse=True,
|
||||
)
|
||||
top_k = max(1, settings.memory_top_k)
|
||||
used = [r for r in scored if r[2]]
|
||||
remaining = max(0, top_k - len(used))
|
||||
candidates = [r for r in scored if not r[2]]
|
||||
kept, suppressed = memorybank._drop_redundant(candidates, held, authority_of, remaining)
|
||||
selected = {r[1] for r in used + kept}
|
||||
suppressed_by = dict(suppressed)
|
||||
return {
|
||||
"query": query,
|
||||
"top_k": top_k,
|
||||
"scored": [
|
||||
{"rank": i + 1, "memory_id": memory_id, "similarity": round(score, 4),
|
||||
"pinned": pinned, "selected": memory_id in selected,
|
||||
"suppressed_as_duplicate_of": suppressed_by.get(memory_id)}
|
||||
for i, (score, memory_id, pinned) in enumerate(scored)
|
||||
],
|
||||
"selected": sorted(selected),
|
||||
}
|
||||
|
||||
|
||||
def production_query(adventure, exclude_action_id: int | None) -> str:
|
||||
"""The retrieval query a turn used: its newest actions, as `retrieve_memories` builds it."""
|
||||
recent = history.tail(adventure, memorybank.RETRIEVAL_WINDOW_ACTIONS, exclude_action_id)
|
||||
return builder.truncate_to_last_tokens(
|
||||
"\n\n".join(a.text for a in recent), memorybank.RETRIEVAL_WINDOW_TOKENS)
|
||||
|
||||
|
||||
def eviction_order(db, adventure) -> list[int]:
|
||||
"""The order `_evict_over_capacity` would take unpinned active memories in."""
|
||||
from sqlalchemy import func
|
||||
return db.execute(
|
||||
select(models.Memory.id).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.forgotten.is_(False),
|
||||
models.Memory.pinned.is_(False),
|
||||
).order_by(func.coalesce(models.Memory.last_used_at, models.Memory.created_at),
|
||||
models.Memory.use_count)
|
||||
).scalars().all()
|
||||
|
||||
|
||||
async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
|
||||
recall_action: models.Action, embed) -> dict:
|
||||
"""The four stages for `fact`, judged at `recall_action`, the recall turn's AI node.
|
||||
|
||||
Ranking is recomputed with the query that turn used, and checked against the
|
||||
turn's own stored `memories.used`. Injection is read from that snapshot, so
|
||||
it reports what the narrator was actually given, not a re-run.
|
||||
"""
|
||||
snapshot = recall_action.context_snapshot or {}
|
||||
out: dict = {"fact_id": fact.fact_id, "plant_depth": plant_depth,
|
||||
"recall_depth": recall_action.depth}
|
||||
|
||||
covering = covering_memories(db, adventure, plant_depth)
|
||||
carrying = [m for m in covering if fact.carried_by(m.text)]
|
||||
elsewhere = [m for m in db.execute(select(models.Memory).where(
|
||||
models.Memory.adventure_id == adventure.id)).scalars().all()
|
||||
if fact.carried_by(m.text) and m not in carrying]
|
||||
creation_input = []
|
||||
for memory in covering:
|
||||
block = memorybank.source_block(db, memory)
|
||||
raw = "\n\n".join(a.text for a in block)
|
||||
excerpt = builder.truncate_to_last_tokens(raw, memorybank.MEMORY_EXCERPT_TOKENS)
|
||||
creation_input.append({
|
||||
"memory_id": memory.id, "source_start": memory.source_start,
|
||||
"source_end": memory.source_end, "block_tokens": builder.count_tokens(raw),
|
||||
"fact_in_block": fact.carried_by(raw),
|
||||
"fact_in_summariser_excerpt": fact.carried_by(excerpt),
|
||||
"memory_text": memory.text,
|
||||
})
|
||||
memory = carrying[0] if carrying else None
|
||||
out["created"] = {
|
||||
"yes": memory is not None,
|
||||
"memory_id": getattr(memory, "id", None),
|
||||
"source_start": getattr(memory, "source_start", None),
|
||||
"source_end": getattr(memory, "source_end", None),
|
||||
"memory_text": getattr(memory, "text", None),
|
||||
"covering_memories": creation_input,
|
||||
"no_covering_memory": not covering,
|
||||
"carried_by_other_memories": [
|
||||
{"memory_id": m.id, "source_start": m.source_start, "source_end": m.source_end}
|
||||
for m in elsewhere],
|
||||
}
|
||||
|
||||
if memory is None:
|
||||
out["verdict"] = "not_created"
|
||||
return out
|
||||
|
||||
order = eviction_order(db, adventure)
|
||||
active = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.forgotten.is_(False))).scalars().all()
|
||||
on_lineage = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id == memory.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory))).scalar() is not None
|
||||
out["retained"] = {
|
||||
"yes": not memory.forgotten,
|
||||
"forgotten": memory.forgotten,
|
||||
"pinned": memory.pinned,
|
||||
"embedded": memory.embedded,
|
||||
"on_active_lineage": on_lineage,
|
||||
"use_count": memory.use_count,
|
||||
"last_used_at": str(memory.last_used_at) if memory.last_used_at else None,
|
||||
"created_at": str(memory.created_at),
|
||||
"active_memories": len(active),
|
||||
"memory_bank_capacity": settings.memory_bank_capacity,
|
||||
"eviction_position": (order.index(memory.id) + 1) if memory.id in order else None,
|
||||
"reason": ("evicted: marked forgotten by capacity eviction" if memory.forgotten
|
||||
else "active"),
|
||||
}
|
||||
if memory.forgotten:
|
||||
out["verdict"] = "created_but_evicted"
|
||||
return out
|
||||
|
||||
query = production_query(adventure, recall_action.id)
|
||||
ranking = await rank_bank(db, adventure, settings, query, embed)
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory.id), None)
|
||||
stored_used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
|
||||
out["ranked"] = {
|
||||
"yes": row is not None and row["rank"] <= ranking["top_k"],
|
||||
"eligible": row is not None,
|
||||
"lexical_score": None, # memory ranking has no lexical term (CONTEXT-AND-MEMORY §20)
|
||||
"semantic_score": row["similarity"] if row else None,
|
||||
"final_score": row["similarity"] if row else None,
|
||||
"pin_effect": "always selected" if memory.pinned else "none",
|
||||
"rank": row["rank"] if row else None,
|
||||
"of": len(ranking["scored"]),
|
||||
"top_k_cutoff": ranking["top_k"],
|
||||
"selected": bool(row and row["selected"]),
|
||||
"suppressed_as_duplicate_of": row["suppressed_as_duplicate_of"] if row else None,
|
||||
"query": query,
|
||||
"replica_matches_stored_selection": sorted(stored_used) == ranking["selected"],
|
||||
}
|
||||
if row is None or row["rank"] > ranking["top_k"] and not row["selected"]:
|
||||
out["verdict"] = "retained_but_not_ranked"
|
||||
return out
|
||||
if not row["selected"]:
|
||||
out["verdict"] = "ranked_but_not_selected"
|
||||
return out
|
||||
|
||||
section = _section(snapshot, MEMORIES_LABEL)
|
||||
injected = memory.id in stored_used and memory.text in section
|
||||
out["injected"] = {
|
||||
"yes": injected,
|
||||
"context_component": MEMORIES_LABEL,
|
||||
"in_stored_memories_used": memory.id in stored_used,
|
||||
"text_in_section": memory.text in section,
|
||||
"token_count": builder.count_tokens(section) if section else 0,
|
||||
}
|
||||
out["verdict"] = "injected" if injected else "selected_but_not_injected"
|
||||
return out
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- the stubs
|
||||
|
||||
@dataclass
|
||||
class BestCaseSummariser:
|
||||
"""The ideal memory writer: F survives if, and only if, F reached it.
|
||||
|
||||
A memory keeps every sentence of the excerpt that carries a planted fact, and
|
||||
adds one sentence naming the block's own distinct detail so memories differ.
|
||||
Summary updates never repeat a planted fact, so the summary layer stays out of
|
||||
the experiment. Every excerpt it was given is kept, for the creation-window
|
||||
diagnostic.
|
||||
"""
|
||||
|
||||
facts: tuple[Fact, ...] = (FACT_F, FACT_G)
|
||||
excerpts: list = field(default_factory=list)
|
||||
|
||||
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
|
||||
if "Current story summary:" in user:
|
||||
return "The travellers kept moving through the country around Westhaven."
|
||||
excerpt = user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
|
||||
self.excerpts.append(excerpt)
|
||||
kept = [s.strip() for s in re.split(r"(?<=[.!?])\s+", excerpt)
|
||||
if any(f.carried_by(s) for f in self.facts)]
|
||||
detail = re.findall(r"\bat the ([a-z]+ [a-z]+)\b", excerpt.lower())
|
||||
tail = f"The travellers spent time at the {detail[-1]}." if detail else \
|
||||
"The travellers pressed on."
|
||||
return " ".join(dict.fromkeys(kept + [tail]))
|
||||
|
||||
|
||||
#: Words that mean the same thing to `ConceptEmbedder`. The point is only that a
|
||||
#: paraphrase lands near the original; the table is the model of that.
|
||||
CONCEPTS = {
|
||||
"timepiece": ("sundial", "dial", "hour", "hours", "clock", "timepiece"),
|
||||
"vessel": ("teapot", "pot", "kettle", "tea", "jar"),
|
||||
"hid": ("hid", "hide", "hidden", "slipped", "tucked", "put", "stashed"),
|
||||
"weathervane": ("weathervane", "vane"),
|
||||
"waterwheel": ("waterwheel", "wheel", "mill"),
|
||||
}
|
||||
_WORD_TO_CONCEPT = {w: c for c, words in CONCEPTS.items() for w in words}
|
||||
DIMENSIONS = 96
|
||||
|
||||
|
||||
@dataclass
|
||||
class ConceptEmbedder:
|
||||
"""A deterministic embedding: concepts in fixed dimensions, other words hashed."""
|
||||
|
||||
calls: int = 0
|
||||
|
||||
async def embed(self, texts):
|
||||
self.calls += 1
|
||||
return [self.vector(t) for t in texts]
|
||||
|
||||
@staticmethod
|
||||
def vector(text: str) -> list[float]:
|
||||
v = [0.0] * DIMENSIONS
|
||||
v[0] = 0.2 # every text shares a little, as real embeddings do
|
||||
concept_names = list(CONCEPTS)
|
||||
for word in re.findall(r"[a-z]+", text.lower()):
|
||||
concept = _WORD_TO_CONCEPT.get(word)
|
||||
if concept is not None:
|
||||
v[1 + concept_names.index(concept)] += 3.0
|
||||
elif len(word) > 3:
|
||||
bucket = int(hashlib.sha256(word.encode()).hexdigest(), 16)
|
||||
v[1 + len(concept_names) + bucket % (DIMENSIONS - 1 - len(concept_names))] += 1.0
|
||||
norm = math.sqrt(sum(x * x for x in v)) or 1.0
|
||||
return [x / norm for x in v]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- scenarios
|
||||
|
||||
#: Filler places. No word here is in `CONCEPTS`, and none names a planted fact.
|
||||
PLACES = (
|
||||
"north gate", "salt market", "ferry landing", "chapel steps", "rope walk",
|
||||
"fish stalls", "old bridge", "tanner yard", "lamp street", "weir path",
|
||||
"grain store", "boat yard", "watch house", "cloth hall", "eel traps",
|
||||
"sheep fold", "smith forge", "stone quay", "reed beds", "toll booth",
|
||||
)
|
||||
PARAPHRASE_QUERY = "I ask Mara where she tucked the little brass dial that tells the hour."
|
||||
UNRELATED_QUERY = "I ask the ferryman what rope costs at the landing this season."
|
||||
|
||||
|
||||
def filler_prose(index: int, words: int) -> str:
|
||||
"""Narration that moves on and never touches a planted fact."""
|
||||
place = PLACES[index % len(PLACES)]
|
||||
sentence = (f"At the {place} the travellers stopped, listened to the gulls over the "
|
||||
f"grey water, and talked about the long road north.")
|
||||
reps = max(1, round(words / len(sentence.split())))
|
||||
return " ".join([sentence] * reps)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Scenario:
|
||||
"""One deterministic campaign. Depths: the opening is 0, turn *n*'s player
|
||||
action is 2n-1 and its reply 2n."""
|
||||
|
||||
name: str
|
||||
turns: int = 52
|
||||
capacity: int = 80
|
||||
top_k: int = 5
|
||||
budget: int = 4096
|
||||
prose_words: int = 60
|
||||
plant_turn: int = 1
|
||||
recall_text: str = "I ask Mara where she hid the amber sundial."
|
||||
pin_first_memory: bool = False
|
||||
lineage_control: bool = False
|
||||
diagnose_recall: bool = True
|
||||
|
||||
|
||||
SCENARIOS = {
|
||||
"independent_default": Scenario("independent_default"),
|
||||
"past_capacity": Scenario("past_capacity", capacity=6),
|
||||
"past_capacity_pinned": Scenario("past_capacity_pinned", capacity=6, pin_first_memory=True),
|
||||
# Closer to the shipped ratio (memory_top_k 5 against capacity 80): most of
|
||||
# the bank is not retrieved on a given turn.
|
||||
"past_capacity_low_top_k": Scenario("past_capacity_low_top_k", capacity=8, top_k=2),
|
||||
"long_block_fact_early": Scenario("long_block_fact_early", turns=10, prose_words=850,
|
||||
plant_turn=1, budget=16384),
|
||||
"long_block_fact_late": Scenario("long_block_fact_late", turns=10, prose_words=850,
|
||||
plant_turn=3, budget=16384),
|
||||
"lineage_control": Scenario("lineage_control", lineage_control=True),
|
||||
}
|
||||
|
||||
|
||||
class ScriptNarrator:
|
||||
"""Stands in for the narrator: returns `next_reply`, with an empty state block."""
|
||||
|
||||
next_reply = ""
|
||||
last_usage = None
|
||||
prompts: list = []
|
||||
|
||||
def __init__(self, *a, **k):
|
||||
pass
|
||||
|
||||
async def generate(self, parts, *, temperature, max_tokens):
|
||||
ScriptNarrator.prompts.append((parts.system, parts.story))
|
||||
yield ("text", ScriptNarrator.next_reply)
|
||||
|
||||
|
||||
def run_scenario(scenario: Scenario) -> dict:
|
||||
"""Plays `scenario` through the real turn route and returns everything measured.
|
||||
|
||||
Uses the database `app.database` is already bound to, creating and dropping
|
||||
its tables, the way the suite's fixtures do. Patches are applied here and
|
||||
removed before returning, so this runs the same under pytest and from the CLI.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, limits
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures as adventure_routes
|
||||
|
||||
summariser = BestCaseSummariser()
|
||||
embedder = ConceptEmbedder()
|
||||
patches = [
|
||||
(memorybank, "summary_provider", lambda s: summariser),
|
||||
(memorybank, "embedding_provider", lambda s: embedder),
|
||||
# Post-turn work is settled explicitly after each turn, so eviction
|
||||
# happens at a known point rather than whenever a background task runs.
|
||||
(memorybank, "schedule_post_turn", lambda adventure: None),
|
||||
(adventure_routes.turns, "OpenAICompatibleProvider", ScriptNarrator),
|
||||
(limits, "check_row_cap", lambda *a, **k: None),
|
||||
]
|
||||
saved = [(obj, name, getattr(obj, name)) for obj, name, _ in patches]
|
||||
for obj, name, value in patches:
|
||||
setattr(obj, name, value)
|
||||
ScriptNarrator.prompts = []
|
||||
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email=f"b1-{scenario.name}@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="script", endpoint_url="http://127.0.0.1:9/v1",
|
||||
embedding_model="concept-embed", context_token_budget=scenario.budget,
|
||||
max_output_tokens=500, memory_bank_capacity=scenario.capacity,
|
||||
memory_top_k=scenario.top_k,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title=f"B.1 {scenario.name}", memory_bank_enabled=True,
|
||||
auto_summarize=True, persona_name="Aldric",
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(adventure_id=adventure.id, type="start",
|
||||
text="Rain over Westhaven, and the tavern door banging in the wind."))
|
||||
db.commit()
|
||||
adv, user_id = adventure.id, user.id
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
client = TestClient(app)
|
||||
result: dict = {"scenario": scenario.__dict__.copy(), "trace": []}
|
||||
|
||||
def call(method, path, body=None, expect=200):
|
||||
response = client.request(method, f"/api/adventures/{adv}{path}", json=body)
|
||||
assert response.status_code == expect, (path, response.status_code, response.text[:300])
|
||||
return response.json() if response.content else None
|
||||
|
||||
def marks():
|
||||
with SessionLocal() as db:
|
||||
rows = db.execute(select(models.Memory.id, models.Memory.forgotten,
|
||||
models.Memory.embedded).where(
|
||||
models.Memory.adventure_id == adv)).all()
|
||||
summaries_n = db.query(models.Summary).filter_by(adventure_id=adv).count()
|
||||
return tuple(sorted(rows)), summaries_n
|
||||
|
||||
def settle():
|
||||
for _ in range(12):
|
||||
before = marks()
|
||||
asyncio.run(memorybank.run_post_turn(adv))
|
||||
if marks() == before:
|
||||
return
|
||||
|
||||
def memories():
|
||||
with SessionLocal() as db:
|
||||
return [dict(row._mapping) for row in db.execute(select(
|
||||
models.Memory.id, models.Memory.text, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.forgotten, models.Memory.pinned,
|
||||
models.Memory.use_count, models.Memory.last_used_at, models.Memory.branch_id,
|
||||
models.Memory.created_at).where(models.Memory.adventure_id == adv)
|
||||
.order_by(models.Memory.id)).all()]
|
||||
|
||||
def turn(kind, text, reply):
|
||||
ScriptNarrator.next_reply = f"{reply}\n```state\n{{\"events\": []}}\n```"
|
||||
response = client.post(f"/api/adventures/{adv}/actions", json={"type": kind, "text": text})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
assert '"type": "error"' not in response.text, response.text[-300:]
|
||||
|
||||
plant_depth = None
|
||||
f_memory_id = None
|
||||
pinned_id = None
|
||||
known: dict[int, dict] = {}
|
||||
g: dict = {}
|
||||
try:
|
||||
for n in range(1, scenario.turns + 1):
|
||||
if n == scenario.plant_turn:
|
||||
turn("story", FACT_F.sentence, filler_prose(n, scenario.prose_words))
|
||||
with SessionLocal() as db:
|
||||
plant_depth = db.query(models.Action.depth).filter_by(
|
||||
adventure_id=adv, text=FACT_F.sentence).scalar()
|
||||
elif scenario.lineage_control and n == 21:
|
||||
call("POST", "/checkpoints", {"name": "before the mill"}, expect=201)
|
||||
turn("story", FACT_G.sentence, filler_prose(n, scenario.prose_words))
|
||||
with SessionLocal() as db:
|
||||
g["plant_depth"] = db.query(models.Action.depth).filter_by(
|
||||
adventure_id=adv, text=FACT_G.sentence).scalar()
|
||||
elif scenario.lineage_control and n == 30:
|
||||
# Line A carries G's memory. Mark it, then abandon it: Undo back
|
||||
# to before G was planted and write something else.
|
||||
g["line_a"] = call("POST", "/checkpoints", {"name": "line A, after the mill"},
|
||||
expect=201)["id"]
|
||||
with SessionLocal() as db:
|
||||
g_rows = [m for m in db.execute(select(models.Memory).where(
|
||||
models.Memory.adventure_id == adv)).scalars() if FACT_G.carried_by(m.text)]
|
||||
g["memory_ids"] = [m.id for m in g_rows]
|
||||
with SessionLocal() as db:
|
||||
g["last_action_id_before_divergence"] = db.query(models.Action.id).filter_by(
|
||||
adventure_id=adv).order_by(models.Action.id.desc()).limit(1).scalar()
|
||||
for _ in range(9):
|
||||
call("POST", "/undo")
|
||||
turn("do", f"I turn away from the mill and walk to the {PLACES[n % len(PLACES)]}.",
|
||||
filler_prose(n + 100, scenario.prose_words))
|
||||
g["diverged_at_turn"] = n
|
||||
else:
|
||||
turn("do", f"I walk on to the {PLACES[n % len(PLACES)]}.",
|
||||
filler_prose(n, scenario.prose_words))
|
||||
settle()
|
||||
|
||||
rows = memories()
|
||||
created = [r["id"] for r in rows if r["id"] not in known]
|
||||
newly_forgotten = [r["id"] for r in rows
|
||||
if r["forgotten"] and not known.get(r["id"], {}).get("forgotten")]
|
||||
for r in rows:
|
||||
known[r["id"]] = r
|
||||
if f_memory_id is None and plant_depth is not None:
|
||||
for r in rows:
|
||||
if (r["source_start"] is not None and r["source_start"] <= plant_depth
|
||||
<= r["source_end"] and FACT_F.carried_by(r["text"])):
|
||||
f_memory_id = r["id"]
|
||||
if scenario.pin_first_memory and pinned_id is None:
|
||||
candidate = next((r for r in rows if r["id"] != f_memory_id), None)
|
||||
if candidate is not None:
|
||||
call("PATCH", f"/memories/{candidate['id']}", {"pinned": True})
|
||||
pinned_id = candidate["id"]
|
||||
f_row = known.get(f_memory_id) if f_memory_id else None
|
||||
result["trace"].append({
|
||||
"turn": n,
|
||||
"active": sum(1 for r in rows if not r["forgotten"]),
|
||||
"total": len(rows),
|
||||
"created": created,
|
||||
"evicted": newly_forgotten,
|
||||
"created_and_evicted_same_turn": sorted(set(created) & set(newly_forgotten)),
|
||||
"f_memory_id": f_memory_id,
|
||||
"f_forgotten": bool(f_row and f_row["forgotten"]),
|
||||
"f_use_count": f_row["use_count"] if f_row else None,
|
||||
})
|
||||
|
||||
turn("do", scenario.recall_text, filler_prose(999, scenario.prose_words))
|
||||
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv)
|
||||
settings = db.query(models.Settings).filter_by(user_id=user_id).first()
|
||||
recall_action = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
result["plant_depth"] = plant_depth
|
||||
result["recall_depth"] = recall_action.depth
|
||||
result["isolation"] = isolation(
|
||||
db, adventure, FACT_F, plant_depth,
|
||||
recall_snapshot=recall_action.context_snapshot,
|
||||
recall_depth=recall_action.depth)
|
||||
result["diagnosis"] = asyncio.run(diagnose(
|
||||
db, adventure, settings, FACT_F, plant_depth,
|
||||
recall_action=recall_action, embed=embedder.embed))
|
||||
result["summariser_excerpts"] = len(summariser.excerpts)
|
||||
|
||||
memory_id = result["diagnosis"]["created"]["memory_id"]
|
||||
if memory_id is not None and not result["diagnosis"]["retained"]["forgotten"]:
|
||||
variants = {}
|
||||
for label, query in (("direct", scenario.recall_text),
|
||||
("paraphrase", PARAPHRASE_QUERY),
|
||||
("unrelated", UNRELATED_QUERY)):
|
||||
ranking = asyncio.run(rank_bank(db, adventure, settings, query, embedder.embed))
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
|
||||
variants[label] = {"query": query, "rank": row and row["rank"],
|
||||
"of": len(ranking["scored"]),
|
||||
"similarity": row and row["similarity"],
|
||||
"selected": bool(row and row["selected"]),
|
||||
"top_k": ranking["top_k"]}
|
||||
result["ranking_variants"] = variants
|
||||
if memory_id is not None:
|
||||
result["f_first_used_turn"] = next(
|
||||
(t["turn"] for t in result["trace"] if (t["f_use_count"] or 0) > 0), None)
|
||||
result["f_last_use_increase_turn"] = max(
|
||||
(b["turn"] for a, b in zip(result["trace"], result["trace"][1:])
|
||||
if (b["f_use_count"] or 0) > (a["f_use_count"] or 0)), default=None)
|
||||
|
||||
evicted_turn = next((t["turn"] for t in result["trace"] if t["f_forgotten"]), None)
|
||||
first_evictions = next((t["evicted"] for t in result["trace"] if t["evicted"]), [])
|
||||
result["eviction"] = {
|
||||
"capacity": scenario.capacity,
|
||||
"f_evicted_at_turn": evicted_turn,
|
||||
"f_use_count_when_evicted": next(
|
||||
(t["f_use_count"] for t in result["trace"] if t["f_forgotten"]), None),
|
||||
"first_eviction_turn": next(
|
||||
(t["turn"] for t in result["trace"] if t["evicted"]), None),
|
||||
"first_evicted_ids": first_evictions,
|
||||
"f_memory_was_first_evicted": bool(f_memory_id and f_memory_id in first_evictions),
|
||||
"created_and_evicted_same_turn": sorted(
|
||||
{i for t in result["trace"] for i in t["created_and_evicted_same_turn"]}),
|
||||
"pinned_memory_id": pinned_id,
|
||||
"pinned_memory_forgotten": bool(pinned_id and known[pinned_id]["forgotten"]),
|
||||
}
|
||||
|
||||
if scenario.lineage_control:
|
||||
path_clause = lineage.path_of(db, adventure).clause(models.Memory)
|
||||
stored = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id.in_(g.get("memory_ids") or [-1]))).scalars().all()
|
||||
eligible = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id.in_(g.get("memory_ids") or [-1]), path_clause)).scalars().all()
|
||||
used_after = set()
|
||||
injected_text = False
|
||||
# Only turns played after the divergence. Before it, G was on the
|
||||
# active line, and a memory of it being used then is correct.
|
||||
for action in (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv,
|
||||
models.Action.type == "ai",
|
||||
models.Action.id > g["last_action_id_before_divergence"])
|
||||
.options(undefer(models.Action.context_snapshot))):
|
||||
snap = action.context_snapshot or {}
|
||||
for m in (snap.get("memories") or {}).get("used") or []:
|
||||
if m.get("id") in (g.get("memory_ids") or []):
|
||||
used_after.add(action.id)
|
||||
if FACT_G.mentioned_by(_section(snap, MEMORIES_LABEL)):
|
||||
injected_text = True
|
||||
g.update(stored=stored, eligible_on_active_line=eligible,
|
||||
turns_whose_memories_used_named_g=sorted(used_after),
|
||||
g_text_ever_in_used_memories=injected_text)
|
||||
if scenario.lineage_control:
|
||||
call("POST", f"/checkpoints/{g['line_a']}/restore")
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv)
|
||||
eligible = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id.in_(g.get("memory_ids") or [-1]),
|
||||
lineage.path_of(db, adventure).clause(models.Memory))).scalars().all()
|
||||
g["eligible_after_returning_to_line_a"] = eligible
|
||||
result["lineage_control"] = g
|
||||
return result
|
||||
finally:
|
||||
for obj, name, value in saved:
|
||||
setattr(obj, name, value)
|
||||
app.dependency_overrides.clear()
|
||||
adventure_routes.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
@@ -0,0 +1,122 @@
|
||||
"""v1.1 WP-B.1: run the deterministic memory-retention scenarios, or diagnose a real campaign.
|
||||
|
||||
# the deterministic scenarios, against an isolated database in --out
|
||||
.venv/bin/python -m tools.v11_b1_memory scenarios --out "$HOME/v11-evidence/b1/<label>"
|
||||
|
||||
# the four stages for a finished real campaign (reads its database; embeds
|
||||
# the recall query with the campaign's own configured embedding model)
|
||||
AIDND_TEST_ENDPOINT=... AIDND_TEST_EMBED_MODEL=nomic-embed-text:latest \\
|
||||
.venv/bin/python -m tools.v11_b1_memory diagnose --db <campaign.db> \\
|
||||
--plant-depth 3 --out "$HOME/v11-evidence/b1/<label>"
|
||||
|
||||
Run from `backend/`. Nothing here changes memory behaviour; see
|
||||
`tools/memory_diagnostic.py` for what is measured and what the stubs model.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
sub = parser.add_subparsers(dest="command", required=True)
|
||||
scen = sub.add_parser("scenarios")
|
||||
scen.add_argument("--out", required=True)
|
||||
scen.add_argument("--only", action="append", default=[])
|
||||
diag = sub.add_parser("diagnose")
|
||||
diag.add_argument("--db", required=True)
|
||||
diag.add_argument("--plant-depth", type=int, required=True)
|
||||
diag.add_argument("--out", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
if args.command == "scenarios":
|
||||
db_path = out / "scenarios.db"
|
||||
if db_path.exists():
|
||||
db_path.unlink()
|
||||
os.environ["AIDND_DB_PATH"] = str(db_path)
|
||||
else:
|
||||
# A copy, so diagnosis never writes to the evidence database.
|
||||
copy = out / "diagnosed-copy.db"
|
||||
shutil.copy2(args.db, copy)
|
||||
os.environ["AIDND_DB_PATH"] = str(copy)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from tools import memory_diagnostic as md # after the database is chosen
|
||||
|
||||
if args.command == "scenarios":
|
||||
names = args.only or list(md.SCENARIOS)
|
||||
summary = {}
|
||||
for name in names:
|
||||
result = md.run_scenario(md.SCENARIOS[name])
|
||||
(out / f"{name}.json").write_text(json.dumps(result, indent=2, default=str))
|
||||
d = result.get("diagnosis") or {}
|
||||
summary[name] = {
|
||||
"verdict": d.get("verdict"),
|
||||
"isolation_ok": (result.get("isolation") or {}).get("ok"),
|
||||
"plant_depth": result.get("plant_depth"),
|
||||
"recall_depth": result.get("recall_depth"),
|
||||
"f_evicted_at_turn": (result.get("eviction") or {}).get("f_evicted_at_turn"),
|
||||
}
|
||||
print(f"{name:26} verdict={d.get('verdict')!s:26} "
|
||||
f"isolation_ok={summary[name]['isolation_ok']} "
|
||||
f"plant={result.get('plant_depth')} recall={result.get('recall_depth')}")
|
||||
(out / "summary.json").write_text(json.dumps(summary, indent=2))
|
||||
return 0
|
||||
|
||||
import asyncio
|
||||
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import memorybank, models
|
||||
from app.database import SessionLocal
|
||||
|
||||
endpoint = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
embed_model = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
|
||||
with SessionLocal() as db:
|
||||
adventure = db.query(models.Adventure).order_by(models.Adventure.id).first()
|
||||
settings = db.query(models.Settings).filter_by(user_id=adventure.user_id).first()
|
||||
if endpoint:
|
||||
settings.endpoint_url = endpoint
|
||||
if embed_model:
|
||||
settings.embedding_model = embed_model
|
||||
recall_action = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adventure.id,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
embed = memorybank.embedding_provider(settings).embed
|
||||
iso = md.isolation(db, adventure, md.FACT_F, args.plant_depth,
|
||||
recall_snapshot=recall_action.context_snapshot,
|
||||
recall_depth=recall_action.depth)
|
||||
diagnosis = asyncio.run(md.diagnose(db, adventure, settings, md.FACT_F, args.plant_depth,
|
||||
recall_action=recall_action, embed=embed))
|
||||
variants = {}
|
||||
memory_id = diagnosis["created"]["memory_id"]
|
||||
if memory_id is not None and not diagnosis.get("retained", {}).get("forgotten"):
|
||||
for label, query in (("recall_turn", diagnosis.get("ranked", {}).get("query", "")),
|
||||
("paraphrase", md.PARAPHRASE_QUERY),
|
||||
("unrelated", md.UNRELATED_QUERY)):
|
||||
ranking = asyncio.run(md.rank_bank(db, adventure, settings, query, embed))
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
|
||||
variants[label] = {"rank": row and row["rank"], "of": len(ranking["scored"]),
|
||||
"similarity": row and row["similarity"],
|
||||
"selected": bool(row and row["selected"])}
|
||||
db.rollback()
|
||||
report = {"isolation": iso, "diagnosis": diagnosis, "ranking_variants": variants}
|
||||
(out / "diagnosis.json").write_text(json.dumps(report, indent=2, default=str))
|
||||
print(json.dumps({"isolation_ok": iso["ok"], "verdict": diagnosis["verdict"]}, indent=2))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,623 @@
|
||||
# v1.1 WP-B.1 — Independent Long-Term Memory Retention Diagnostic
|
||||
|
||||
**Status:** COMPLETE. Diagnostic only: no memory behaviour was changed. The final decision is in §S.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `v1.1-development` |
|
||||
| HEAD at start | `d63804f22ecbaed80741241a154cdaef82f7b2ed` — *v1.1: harden context window and narrator protocol boundary* (WP-A1/A2), signed by the owner (good signature, RSA key `02C9BF7D…`) |
|
||||
| Its parents | `ac465ed` (planning v4.1, signed) and `432f041` (the signed v1.0.0 release commit; tag `v1.0.0`; `main`) |
|
||||
| Working tree at start | clean |
|
||||
| `git diff --stat v1.0.0..HEAD` | 27 files, +4,880 / −88: the planning commit and WP-A1/A2 |
|
||||
| Comparison baseline | `432f041` (v1.0.0), run from a throwaway worktree (§D) |
|
||||
|
||||
`app/memorybank.py`, `app/summaries.py`, `app/context/lineage.py`, `app/tree.py` and
|
||||
`app/vectors.py` are unchanged between `v1.0.0` and HEAD. A1 and A2 did not touch
|
||||
the memory pipeline, apart from adding `accounting` to `attempts.ATTEMPT_KEYS`.
|
||||
|
||||
---
|
||||
|
||||
## B. Existing memory pipeline
|
||||
|
||||
Answered from the code at HEAD. Nothing was changed.
|
||||
|
||||
| # | Question | Answer |
|
||||
| --- | --- | --- |
|
||||
| 1 | How are memories generated? | `memorybank.run_post_turn`, a fire-and-forget task after each accepted turn (`schedule_post_turn`), runs `_create_due_memories` when the campaign has `auto_summarize`. It writes one memory per block of `MEMORY_INTERVAL` = 6 story actions past the memory cursor. It starts once the story has `MEMORY_START` = 12 actions, and only when `SETTLE_SLACK` = 1 action sits past the block. At most `MAX_MEMORIES_PER_RUN` = 5 memories are written per run. |
|
||||
| 2 | What range does a memory cover? | The block's first and last action depths: `Memory.source_start`, `Memory.source_end`. |
|
||||
| 3 | How are source depth and lineage stored? | `tree.attach_memory` sets `Memory.branch_id` and `Memory.depth` from the block's **last** node, so a memory is visible exactly on paths that contain that node. |
|
||||
| 4 | Memory text length limit | Prompt-only: `MEMORY_MAX_WORDS` = 50 in `MEMORY_SYSTEM_PROMPT` ("1-2 plain sentences"). Nothing truncates the stored text. |
|
||||
| 5 | What input does the summariser receive? | `summarize_block`: a cast brief (`cast_brief`, which reads story cards, persona and plot essentials), then `"Story excerpt:\n\n{excerpt}\n\nMemory:"`. The excerpt is the block's action texts joined by blank lines. |
|
||||
| 6 | Where does the 2,000-token truncation happen? | `summarize_block`: `excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)`, with `MEMORY_EXCERPT_TOKENS` = 2,000. It keeps the block's **last** 2,000 `cl100k_base` tokens. The cast brief is matched against the untruncated block. |
|
||||
| 7 | How does capacity and eviction work? | `_evict_over_capacity`, at the end of every `run_post_turn`. It counts non-forgotten memories in the whole adventure (not lineage-scoped). Above `settings.memory_bank_capacity` (default 80), it marks `forgotten = True` on the overflow unpinned memories, ordered by `coalesce(last_used_at, created_at)` ascending, then `use_count`. Forgotten rows are kept. |
|
||||
| 8 | How is `last_accessed` updated? | The field is `Memory.last_used_at`. `record_use` sets it, and increments `use_count`, for every memory in the turn's `memories.used`. That write is part of the turn's single commit (§O.7 of the M11 report). Dry runs never count. |
|
||||
| 9 | How are memories ranked? | `retrieve_memories`. The query is the text of the newest `RETRIEVAL_WINDOW_ACTIONS` = 4 actions, truncated to the last `RETRIEVAL_WINDOW_TOKENS` = 600 tokens. It is embedded with the configured embedding model. Every eligible memory is scored by `vectors.cosine`. Pinned memories are taken first, the rest fill `memory_top_k` (default 5) in score order, and `_drop_redundant` skips a candidate at cosine ≥ 0.93 to one already chosen, within the same authority. **There is no lexical term, recency term, importance term or similarity floor** (CONTEXT-AND-MEMORY §20, as implemented). |
|
||||
| 10 | How do pins affect ranking and eviction? | Ranking: a pinned memory is always selected, and counts toward `memory_top_k`. Eviction: pinned memories are never evicted, and if every active memory is pinned, capacity is exceeded. |
|
||||
| 11 | How does the active-lineage clause filter memories? | `lineage.path_of(db, adventure).clause(models.Memory)` matches `(branch_id, depth)` against the head's path entries, capped at each fork depth. Retrieval also requires `forgotten = False` and `embedded = True`. |
|
||||
| 12 | How do retrieved memories enter `build_context`? | Through the `memory_bank` argument. The builder renders `"Memories from earlier in the story. Lines marked [inferred] are interpretation, not established fact — do not treat them as settled truth:"` plus one `- [inferred]? text` line per memory, as section `used_memories`. It is a live section priced into protected context, placed after history and before `narrative_state`. |
|
||||
| 13 | Where is the provenance recorded? | `context_snapshot["memories"]`: the whole retrieval result, meaning `used` (id, text, similarity, pinned, authority, source range), `considered` and `suppressed`. It is stored per turn, so a past turn's selection is inspectable. |
|
||||
| 14 | How does memory survive export and import? | `bundle._exported_memory` carries text, pinned, forgotten, `sourceStart`, `sourceEnd`, `useCount`, authority, branch and depth. **It does not carry the vector, `last_used_at` or `created_at`.** On import the memories are re-embedded by the post-turn pass, and their recency restarts from the import. |
|
||||
|
||||
**A finding from the inspection itself.** The v1 long-run harness (`tools/m11_long_run.py`)
|
||||
planted M04's clue as an accepted **state correction** (`CLUE_FACT`, via
|
||||
`add_fact`). The player turn mentioning it uses the sentinel code, but the fact the
|
||||
recall looks for was established in state. Memories are written from **story
|
||||
text**. So every v1 M04 recovery could only ever run through state, or through the
|
||||
narrator restating state, and none could have tested memory on its own. That is why
|
||||
§P risk 5 of the M11 report could say "no run showed memory keeping a planted fact"
|
||||
without any run having given memory the chance.
|
||||
|
||||
---
|
||||
|
||||
## C. Diagnostic design
|
||||
|
||||
### C.1 What is built
|
||||
|
||||
| File | Kind | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `backend/tools/memory_diagnostic.py` | tool, new | See below: the fact spec, isolation checks, four-stage diagnosis, deterministic stubs and scenario runner. |
|
||||
| `backend/tools/v11_b1_memory.py` | tool, new | CLI. `scenarios` runs the deterministic campaigns against an isolated database and writes JSON. `diagnose` runs the four stages against a **copy** of a finished real campaign's database. |
|
||||
| `backend/tools/m11_long_run.py` | tool, extended | `--independent-fact`: see §C.3. |
|
||||
| `backend/tests/test_v11_b1_memory_diagnostic.py` | tests, new | The deterministic diagnostic. |
|
||||
| `backend/tests/test_v11_b1_long_run_verdict.py` | tests, new | The new long-run verdict. |
|
||||
|
||||
What `tools/memory_diagnostic.py` holds:
|
||||
- **`Fact`:** a planted fact, with whole-word carry and leak matching.
|
||||
- **`isolation()`:** every non-memory layer checked.
|
||||
- **`diagnose()`:** the four stages and the verdict.
|
||||
- **`rank_bank()`:** production ranking recomputed for every eligible memory.
|
||||
- **The stubs:** `BestCaseSummariser`, `ConceptEmbedder` and `ScriptNarrator`.
|
||||
- **`run_scenario()`:** a campaign played through the real turn route.
|
||||
|
||||
**No application file is changed.** No column, table, migration or setting is
|
||||
added. The diagnostic's extra fields are computed at report time.
|
||||
|
||||
### C.2 How each stage is judged
|
||||
|
||||
| Stage | Judged from | Output (real field names where they exist) |
|
||||
| --- | --- | --- |
|
||||
| **created** | Memories on the active lineage with `source_start ≤ plant_depth ≤ source_end` whose text carries F (both "sundial" and "teapot"). For every covering memory, the block is re-read (`memorybank.source_block`) and cut exactly as `summarize_block` does, to report whether F was in the block and whether it was in **the excerpt the summariser saw**. | `memory_id`, `source_start`, `source_end`, `memory_text`, `covering_memories[].{block_tokens, fact_in_block, fact_in_summariser_excerpt}` |
|
||||
| **retained** | That memory's row, plus the eviction order production would use (the same `ORDER BY`, read-only) | `forgotten`, `pinned`, `embedded`, `on_active_lineage`, `use_count`, `last_used_at`, `created_at`, `active_memories`, `memory_bank_capacity`, `eviction_position`, `reason` |
|
||||
| **ranked** | `rank_bank`: the same catalogue clause, cosine, pin rule and `memorybank._drop_redundant`, over **every** eligible memory, with the recall turn's own query (`history.tail(4, exclude=recall AI node)`, cut to 600 tokens). It is checked against the recall turn's stored `memories.used` (`replica_matches_stored_selection`). | `semantic_score` (`similarity`), `lexical_score` (always `None`: no such term exists), `final_score`, `rank`, `of`, `top_k_cutoff`, `selected`, `suppressed_as_duplicate_of`, `query` |
|
||||
| **injected** | The recall turn's stored `context_snapshot`: `memories.used` names the memory, and its text is in the `used_memories` section | `context_component`, `in_stored_memories_used`, `text_in_section`, `token_count` |
|
||||
|
||||
Verdicts, in order: `not_created`, `created_but_evicted`, `retained_but_not_ranked`,
|
||||
`ranked_but_not_selected`, `selected_but_not_injected`, `injected`. The fifth is
|
||||
added to the brief's list, so that "the retrieval picked it" and "the narrator was
|
||||
shown it" stay distinguishable.
|
||||
|
||||
### C.3 Isolation (precondition) checks
|
||||
|
||||
A result counts only if every check holds.
|
||||
|
||||
| Check | How |
|
||||
| --- | --- |
|
||||
| `state_document` | `adventure.narrative_state`: entities, facts, relationships, threads, scene and possessions, plus the whole document |
|
||||
| `state_snapshots` | `narrative_state_after` of every node on the active lineage |
|
||||
| `later_narration` | every AI turn deeper than the planting block's end, and before the recall turn |
|
||||
| `summary` | `summaries.current` and the recall prompt's `story_summary` section |
|
||||
| `knowledge` | every `KnowledgeSource.content`, and the recall prompt's imported-knowledge sections |
|
||||
| `recent_history` | the recall prompt's `history.floor_depth` is greater than the planting depth, and F is not in the `history` or `recent_history` sections |
|
||||
| `state_section` | the recall prompt's `narrative_state` section |
|
||||
|
||||
The negative control for the precondition itself is
|
||||
`test_the_isolation_check_fails_when_another_layer_carries_the_fact`.
|
||||
|
||||
### C.4 The deterministic stubs, and what they model
|
||||
|
||||
- **`BestCaseSummariser`.** An ideal memory writer. A memory keeps every sentence of
|
||||
its excerpt that carries a planted fact, plus one sentence naming the block's own
|
||||
place so that memories differ. Summary updates never mention a planted fact. **A
|
||||
creation failure under this stub is the application's, not a model's.**
|
||||
- **`ConceptEmbedder`.** A 96-dimension deterministic embedding. Words in a small
|
||||
concept table ("sundial", "dial", "hour", "clock", …) share a dimension, other
|
||||
words are hashed, and the vector is normalised. It models a paraphrase landing
|
||||
near the original. **It says nothing about `nomic-embed-text`.**
|
||||
- **`ScriptNarrator`.** Narration that names only filler places and never a planted
|
||||
fact, with an empty state block, so state never records F.
|
||||
|
||||
Scenarios are played through the real `POST /actions` route. Automatic post-turn
|
||||
scheduling is replaced by an explicit `run_post_turn` settle after every turn, so
|
||||
eviction happens at a known turn. Depths: the opening is 0, turn *n*'s player
|
||||
action is 2n−1 and its reply 2n. The planted fact is a `story` action.
|
||||
|
||||
### C.4.1 Scenarios
|
||||
|
||||
| Scenario | Turns | Capacity | top_k | History budget | Prose per reply | Planted at |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| `independent_default` | 52 (recall at depth 106) | 80 | 5 | 4,096 | ~60 words | depth 1 |
|
||||
| `past_capacity` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1 |
|
||||
| `past_capacity_pinned` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1; first other memory pinned |
|
||||
| `past_capacity_low_top_k` | 52 | 8 | 2 | 4,096 | ~60 words | depth 1 |
|
||||
| `long_block_fact_early` | 10 | 80 | 5 | 16,384 | ~850 words | depth 1, early in a long block |
|
||||
| `long_block_fact_late` | 10 | 80 | 5 | 16,384 | ~850 words | depth 5, late in the same-sized block |
|
||||
| `lineage_control` | 52 | 80 | 5 | 4,096 | ~60 words | F at depth 1; G on line A, then Undo × 9 and divergence at turn 30; Save Points before and after G |
|
||||
|
||||
### C.5 The long-run verdict
|
||||
|
||||
`tools/m11_long_run.py --independent-fact` plants a second, **story-only** fact
|
||||
at depth 3, right after M04's own plant, and never corrects it into state. Every
|
||||
existing M04 behaviour and verdict is unchanged.
|
||||
|
||||
**Every accepted turn** records the first failure of:
|
||||
- `absent_from_state`;
|
||||
- `absent_from_summary`;
|
||||
- `absent_from_later_narration`, meaning narration deeper than the planting depth
|
||||
plus 6.
|
||||
|
||||
The plant and those first failures survive `--resume`.
|
||||
|
||||
**At recall** a dedicated question is played. `_independent_recall` reads the
|
||||
recall turn's stored context and the campaign database (read-only), and
|
||||
`_independent_memory_verdict` returns one of:
|
||||
|
||||
- `recovered_through_memory_independent`, only when every one of
|
||||
`planted_turn_outside_history`, `absent_from_state`, `absent_from_summary`,
|
||||
`absent_from_knowledge` and `absent_from_later_narration` holds, and a memory
|
||||
covering the planting turn carries the fact and was injected;
|
||||
- `precondition_failed:<name>` or `precondition_unknown:<name>`, never a recovery;
|
||||
- `not_recovered:not_created`, `not_recovered:evicted` or
|
||||
`not_recovered:not_injected`.
|
||||
|
||||
Ranking is recomputed afterwards by `tools/v11_b1_memory.py diagnose`, against a
|
||||
copy of the run's database.
|
||||
|
||||
---
|
||||
|
||||
## D. v1.0.0 baseline
|
||||
|
||||
**How it was run.**
|
||||
- A throwaway worktree was checked out at `432f041` (`git describe`: `v1.0.0`).
|
||||
- Only the three B.1 files were copied in: `tools/memory_diagnostic.py`,
|
||||
`tools/v11_b1_memory.py` and `tests/test_v11_b1_memory_diagnostic.py`.
|
||||
- The imported `app` was confirmed to come from the worktree.
|
||||
- The worktree was removed afterwards, and the `v1.0.0` tag and commit were not
|
||||
touched.
|
||||
- `app/memorybank.py` is byte-identical between `v1.0.0` and HEAD.
|
||||
- Evidence: `$HOME/v11-evidence/b1/v100/`, with HEAD's in
|
||||
`$HOME/v11-evidence/b1/head/`.
|
||||
|
||||
| | v1.0.0 (`432f041`) | HEAD (`d63804f` + B.1 files) |
|
||||
| --- | --- | --- |
|
||||
| `test_v11_b1_memory_diagnostic.py` | **24 passed, 2 xfailed (strict)** | 24 passed, 2 xfailed (strict) |
|
||||
| `independent_default` | `injected` | `injected` |
|
||||
| `past_capacity` (capacity 6, top_k 5) | **`created_but_evicted`** | `created_but_evicted` |
|
||||
| `past_capacity_pinned` | `created_but_evicted` (the pinned memory kept) | same |
|
||||
| `past_capacity_low_top_k` (capacity 8, top_k 2) | **`created_but_evicted`** | `created_but_evicted` |
|
||||
| `long_block_fact_early` | **`not_created`** | `not_created` |
|
||||
| `long_block_fact_late` | `injected` | `injected` |
|
||||
| `lineage_control` | `injected`; G never injected after the divergence | same |
|
||||
|
||||
**The two independent-retention acceptance criteria fail on v1.0.0, and each names
|
||||
the stage.**
|
||||
|
||||
```text
|
||||
criterion: an early fact is recalled from memory past capacity
|
||||
created: yes
|
||||
retained: no
|
||||
FAILURE STAGE: retention (capacity eviction)
|
||||
|
||||
criterion: a fact early in a long block is remembered
|
||||
created: no (the fact was in the block, not in the summariser's excerpt)
|
||||
FAILURE STAGE: creation (input truncation)
|
||||
```
|
||||
|
||||
Under the best-case summariser, with blocks shorter than 2,000 tokens and a bank
|
||||
under capacity, v1.0.0 carries the fact all the way to injection. The two failure
|
||||
stages above are therefore **application mechanisms**, reached under conditions
|
||||
a long campaign meets:
|
||||
- a block of long narration;
|
||||
- more memories than `memory_bank_capacity`.
|
||||
|
||||
Which of them a real campaign meets first is §K's question.
|
||||
|
||||
---
|
||||
|
||||
## E. Creation results
|
||||
|
||||
`long_block_fact_early` and `long_block_fact_late` use the same block geometry: the
|
||||
opening (depth 0) and turns 1-3. Each reply is about 850 words, and the block
|
||||
`source_start` 0 … `source_end` 5 is **2,079 tokens**, 79 over
|
||||
`MEMORY_EXCERPT_TOKENS`.
|
||||
|
||||
| Shape | Planted at | Fact in block | Fact in summariser excerpt | Memory written | Verdict |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| F early in the block | depth 1 | yes | **no** | "The travellers spent time at the ferry landing." | **`not_created`** |
|
||||
| F late in the same-sized block | depth 5 | yes | yes | "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." | `injected` |
|
||||
|
||||
`summarize_block` keeps the **last** 2,000 tokens. An early-block fact is cut off
|
||||
before the summariser reads it, even by an overflow of only 79 tokens. No summariser
|
||||
quality can recover what it was never given.
|
||||
|
||||
`independent_default`: 60-word replies, a 4,096 budget, and a block under 2,000
|
||||
tokens. Memory 1 covers depths 0-5, and its text carries F. Both the block and the
|
||||
excerpt contain F.
|
||||
|
||||
---
|
||||
|
||||
## F. Retention / capacity results
|
||||
|
||||
Default: the bank holds 17 active memories at depth 106, against capacity 80. F's
|
||||
memory is retained, has `use_count` 16, and sits at eviction position 14 of 17.
|
||||
|
||||
**Past capacity** (the brief's capacity test). 17 memories are written over 52 turns.
|
||||
|
||||
| Scenario | F's uses before eviction | F's last retrieval | First eviction | F evicted | F first evicted? | Created and evicted in the same pass | Pinned kept |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| `past_capacity` (6 / top_k 5) | 12 | turn 18 | turn 21, memory 1 | **turn 21** | **yes** | none | — |
|
||||
| `past_capacity_pinned` (6 / top_k 5, memory 2 pinned) | 12 | turn 18 | turn 21, memory 1 | turn 21 | yes | none | **memory 2 never evicted** |
|
||||
| `past_capacity_low_top_k` (8 / top_k 2) | 4 | turn 24 | turn 27, memory 3 | **turn 36** (4th eviction) | no | none | — |
|
||||
|
||||
**What the trace shows about the rule.**
|
||||
- **"Never retrieved" is not the mechanism.** F was retrieved, 12 or 4 times, while
|
||||
it still ranked in the top-k for the recent-narration query.
|
||||
- **Eviction orders by `coalesce(last_used_at, created_at)`.** An early fact that
|
||||
recent narration never mentions stops being retrieved once newer memories fill
|
||||
the top-k. It then ages out:
|
||||
- at capacity 6, it is gone 3 turns after its last use, as the first eviction;
|
||||
- at capacity 8 with `top_k` 2, 12 turns after its last use, as the fourth.
|
||||
- **Recall itself cannot rescue it.** Retrieval is driven by recent narration, and
|
||||
a fact nobody mentions is exactly the one that loses recency.
|
||||
- **Pinned memories stay protected** (`past_capacity_pinned`).
|
||||
- **No memory is evicted by the pass that created it.** The frozen-bank regression
|
||||
fix still holds in all three scenarios.
|
||||
|
||||
---
|
||||
|
||||
## G. Ranking results
|
||||
|
||||
The `independent_default` recall turn is at depth 106. Its query is the newest 4
|
||||
actions, cut to 600 tokens: three narration turns and the one-line question.
|
||||
|
||||
| | Value |
|
||||
| --- | --- |
|
||||
| F's `similarity` (semantic score) | **0.241** |
|
||||
| Lexical score | none: memory ranking has no lexical term |
|
||||
| Pin effect | none (not pinned) |
|
||||
| Final rank | **2 of 17** |
|
||||
| `memory_top_k` cutoff | 5 |
|
||||
| Selected | yes |
|
||||
| Replica agrees with the turn's stored `memories.used` | **yes** |
|
||||
|
||||
The same memory against three stand-alone queries, `ConceptEmbedder`:
|
||||
|
||||
| Query | Rank | Similarity | Selected |
|
||||
| --- | --- | --- | --- |
|
||||
| direct: "I ask Mara where she hid the amber sundial." | 1 of 17 | 0.708 | yes |
|
||||
| paraphrase: "… the little brass dial that tells the hour." | 1 of 17 | 0.636 | yes |
|
||||
| unrelated: "… what rope costs at the landing this season." | 5 of 17 | **0.064** | **yes** |
|
||||
|
||||
Two diagnostic observations. Neither is a failure in this scenario.
|
||||
1. **The production query dilutes the question.** The question alone scores 0.708;
|
||||
inside the four-action window it scores 0.241. F survives at rank 2 of 17. In a
|
||||
bank where more memories share the recent narration's vocabulary, the same
|
||||
dilution would push it past `top_k`.
|
||||
2. **There is no relevance floor.** With more memories than `memory_top_k`, five are
|
||||
injected whatever their similarity. An unrelated query still injects F at 0.064.
|
||||
This matters to F only in the other direction: it can ride along even when it is
|
||||
not relevant.
|
||||
|
||||
---
|
||||
|
||||
## H. Injection results
|
||||
|
||||
`independent_default` injects F: `memories.used` in the recall turn's stored snapshot
|
||||
names memory 1, its text is in the `used_memories` section, and that section is 97
|
||||
tokens. The same holds in `long_block_fact_late` and `lineage_control`.
|
||||
**`selected_but_not_injected` never occurred**: whatever retrieval selected, the
|
||||
builder rendered.
|
||||
|
||||
---
|
||||
|
||||
## I. Lineage negative control
|
||||
|
||||
`lineage_control`:
|
||||
- G is planted on line A at depth 41, turn 21.
|
||||
- G's memory 7 is written.
|
||||
- A Save Point is placed on line A after G.
|
||||
- Undo ×9, then divergent writing at turn 30. The last action before the
|
||||
divergence has id 59.
|
||||
|
||||
| Assertion | Result |
|
||||
| --- | --- |
|
||||
| G's memory stays stored | **yes** (memory 7 present) |
|
||||
| G is not eligible on the active lineage | **yes** (the path clause returns nothing) |
|
||||
| G is never injected after the divergence | **yes** (no turn with id > 59 names it or carries its text) |
|
||||
| `memories.used` does not report it after the divergence | **yes** |
|
||||
| Returning to line A (Save Point restore) makes it eligible again | **yes** (memory 7 eligible) |
|
||||
| F on the active line is unaffected | `injected`, isolation holds |
|
||||
|
||||
Before the divergence, G's memory was legitimately used on line A (turn 22). The
|
||||
first version of this check counted that as a leak. That was a defect in the
|
||||
diagnostic, and the scan now starts after the divergence.
|
||||
|
||||
No lineage code was touched. `test_m11_leakage.py`, 14 tests, passes unchanged (§P).
|
||||
|
||||
---
|
||||
|
||||
## J. Authority negative control
|
||||
|
||||
`test_a_memory_that_contradicts_state_loses_and_changes_nothing` sets up the
|
||||
conflict like this:
|
||||
- a state correction adds the fact "the tavern lamp is lit";
|
||||
- a hand-written memory says "The tavern lamp was never lit that night.";
|
||||
- the memory is pinned, so it is injected;
|
||||
- a turn is played with an empty proposal.
|
||||
|
||||
| Assertion | Result |
|
||||
| --- | --- |
|
||||
| The narrative state document is unchanged by retrieval and the turn | **yes** (identical before and after) |
|
||||
| The state fact is in the prompt's `narrative_state` section | yes |
|
||||
| The memory is in `used_memories`, under "Memories from earlier in the story …" | yes, framed as historical and non-canon context |
|
||||
| `narrative_state` comes after `used_memories`, so state is read last and settles the conflict | yes |
|
||||
|
||||
F07 semantics are unchanged: memory never writes state.
|
||||
|
||||
---
|
||||
|
||||
## K. Real-model attempts
|
||||
|
||||
**Setup.**
|
||||
- **Command:** `tools/m11_long_run.py --turns 100 --independent-fact`.
|
||||
- **Host and models:** the GPU inference host (Ollama 0.34.0), `qwen2.5:3b-instruct-16k` at a verified 16,384 window, embeddings by `nomic-embed-text:latest`.
|
||||
- **Memory:** the memory bank and auto-summarise on.
|
||||
- **Harness settings:** `memory_top_k` 4 and `context_token_budget` 16,384.
|
||||
- **Logging:** the owner's power, link and kernel logging was running on the host before the run started.
|
||||
- **Permission:** inference was used only with the owner's explicit approval.
|
||||
|
||||
**Evidence:**
|
||||
- `$HOME/v11-evidence/b1/real-1/`: `summary.json`, `recall-independent.json`, `timeline.jsonl` and `campaign.db`;
|
||||
- `$HOME/v11-evidence/b1/real-1-diagnosis/diagnosis.json`: the four stages, recomputed on a copy of the database with the same embedding model.
|
||||
|
||||
### K.1 Attempt 1 — PRECONDITION FAILED
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Run status | `complete`: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 active memories (19 eligible on the recall line), 1,308 s elapsed |
|
||||
| Window / accounting | verified 16,384 on every turn; `fits` |
|
||||
| M04 verdict (unchanged) | `recovered_through_state_only` |
|
||||
| Independent fact | planted at depth 3 by the player turn *"I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one."* It was never corrected into state |
|
||||
| **Verdict** | **`precondition_failed:absent_from_summary`** (the verdict names the first failed precondition in its fixed order) |
|
||||
|
||||
| Precondition | Result | First failure |
|
||||
| --- | --- | --- |
|
||||
| planted turn outside recent history | **held**: the history window started at depth 66 | — |
|
||||
| absent from authoritative state | **held**: the document, every node's snapshot and the recall prompt's state section | — |
|
||||
| absent from imported knowledge | **held**: 3 sources | — |
|
||||
| **absent from later narration** | **failed**: the narrator mentions the fact at depths 6, 8, 10, 14, 16, 18, 20, 22, 24, 26 and later | accepted turn **5** |
|
||||
| **absent from the active summary** | **failed** | accepted turn **9** |
|
||||
| fact text in the recall prompt's history sections | **present** (restated narration inside the window) | — |
|
||||
|
||||
**Why it did not qualify.** The 3B narrator took the planted detail up as a motif
|
||||
and restated it for the rest of the campaign ("Mara's amber sundial flickered
|
||||
softly, a silent reminder of their shared history"). The summariser, whose prompt
|
||||
asks it to "preserve important established facts", folded it into the running
|
||||
summary. Neither is a defect in the harness. They are the other layers doing what
|
||||
they do with a salient fact, which is exactly what makes a clean memory-only
|
||||
measurement hard to obtain with a real narrator.
|
||||
|
||||
**The four stages, diagnosed anyway.** They are not evidence for independent
|
||||
retention, but they are evidence for the mechanism.
|
||||
|
||||
| Stage | Result |
|
||||
| --- | --- |
|
||||
| **created** | **yes.** Memory 1 covers depths 0-5. The block is 688 tokens, so there was no truncation: the fact was in the block and in the summariser's excerpt. The memory text: *"Aldric taps the silver key against his chest … Mara slipped the amber sundial into the cracked teapot, her fingers tightening on the silver key. …"* |
|
||||
| **retained** | **yes.** Not forgotten, `use_count` 53, 33 active memories against capacity 80, eviction position 28 |
|
||||
| **ranked** | **no.** Eligible and embedded. For the recall turn's own query it scored `similarity` **0.866** and ranked **10 of 19**, against `top_k_cutoff` 4. It was not selected and was not suppressed as a duplicate. The replica matched the stored selection |
|
||||
| **injected** | not reached. The recall turn's `memories.used` = [30, 29, 31, 5] |
|
||||
|
||||
**What the narrator was given instead.** Five **later** memories also carry the
|
||||
fact, all written from narration that restated it: memories 12 (depths 66-71), 28
|
||||
(81-86), 30 (93-98), 17 (96-101) and 18 (102-107). Memory 30 was injected at
|
||||
recall. The fact reached the narrator through memory, but through a
|
||||
**restatement's** memory, not the planting-era one.
|
||||
|
||||
**The same memory 1 against reference queries** (`nomic-embed-text`):
|
||||
|
||||
| Query | Rank | Similarity | Selected |
|
||||
| --- | --- | --- | --- |
|
||||
| the recall turn's production query | 10 / 19 | 0.866 | no |
|
||||
| paraphrase: "… the little brass dial that tells the hour" | 6 / 19 | 0.593 | no |
|
||||
| unrelated: "… what rope costs at the landing this season" | 12 / 19 | 0.435 | no |
|
||||
|
||||
Two properties of the real embedder matter to B.2:
|
||||
1. **A high floor.** An unrelated query still scores 0.44.
|
||||
2. **A crowded top.** The bank is full of near-identical "Aldric and Mara step out,
|
||||
the silver key's weight in his pocket" memories, so 0.866 was not enough to reach
|
||||
the top 4.
|
||||
|
||||
*A side finding outside B.1's scope.* The export's protocol-leak counter flags
|
||||
**1 of 105** stored turns: action 153, depth 143. Mid-reply, the narrator echoed the
|
||||
length hint and the state reminder, with an `Events: [...]` line, and then continued
|
||||
the story. WP-A2's cleanup rules act only at the end of a reply, so an instruction
|
||||
echo with story after it stays. It is recorded here for the v1.1 backlog. B.1 did not
|
||||
touch it.
|
||||
|
||||
---
|
||||
|
||||
## L. First failing stage
|
||||
|
||||
**Deterministic, with a best-case summariser and a concept embedder, identical on
|
||||
v1.0.0 and HEAD:**
|
||||
|
||||
| Condition | First failing stage |
|
||||
| --- | --- |
|
||||
| a bank under capacity, blocks under 2,000 tokens | **none**: created, retained, ranked (2 of 17) and injected |
|
||||
| a bank past `memory_bank_capacity` | **retention**: evicted by least-recently-used order after recent narration stops retrieving it (§F) |
|
||||
| the fact early in a block over 2,000 tokens | **creation**: the fact is in the block, but not in the summariser's excerpt (§E) |
|
||||
|
||||
**Real model, attempt 1** (isolation not met, mechanism only): created yes, retained
|
||||
yes, **ranking failed first**. The planting-era memory scored 0.866 but ranked 10th of
|
||||
19 behind later, near-identical memories, and outside `top_k` 4.
|
||||
|
||||
**Across all the evidence, the first stage an early fact fails at a real campaign's
|
||||
length is ranking.** In both the deterministic default and the real run, the memory
|
||||
exists and is retained at 100 turns, where the bank is below capacity. What decides
|
||||
whether the narrator is shown it is its rank against a recent-narration query in a
|
||||
bank of similar memories.
|
||||
- **Retention (eviction)** is a second, later failure. It is certain once a campaign
|
||||
outgrows capacity: about 480 actions at the defaults.
|
||||
- **Creation (truncation)** is a third, conditional one. It needs blocks longer than
|
||||
2,000 tokens, which the 16,384 window with 500-token replies did not produce
|
||||
(688 tokens).
|
||||
|
||||
---
|
||||
|
||||
## M. Evidence for likely root cause
|
||||
|
||||
1. **The retrieval query is recent narration, not the question** (§B.9, §G). The
|
||||
query is the last 4 actions cut to 600 tokens, so a one-line recall question is
|
||||
outweighed by three turns of prose. The effect is deterministic: the question's
|
||||
own similarity of 0.708 fell to 0.241 in the production query. In the real run,
|
||||
recent prose about the same tavern, key and people made every memory look
|
||||
similar, and ten ranked above the planting-era one.
|
||||
2. **Ranking has no term that favours the planting-era record** (§B.9;
|
||||
CONTEXT-AND-MEMORY §20). There is cosine only: no lexical match on the question's
|
||||
rare terms ("sundial", "teapot"), no importance, and no preference for the earliest
|
||||
or a coverage-distinct source. Later restatement memories carry the same words in
|
||||
more familiar company, and outrank the original.
|
||||
3. **Eviction is purely least-recently-used** (§B.7, §F). Retrieval is driven by
|
||||
recent narration, so exactly the facts nothing recent mentions lose recency and go
|
||||
first. Recall itself cannot rescue them, because they are no longer retrieved.
|
||||
4. **The summariser reads only the last 2,000 tokens of a block** (§B.6, §E). This is
|
||||
proven deterministically. It did not bite at the real run's block sizes.
|
||||
5. **v1's M04 never tested memory** (§B finding). The clue was planted as state, so
|
||||
the long-standing "memory does not keep the fact" observation was never a
|
||||
measurement of memory.
|
||||
|
||||
---
|
||||
|
||||
## N. What B.2 is allowed to change
|
||||
|
||||
B.2 is allowed only the smallest changes the evidence supports, one mechanism at a
|
||||
time, each with a failing test first. In order of the evidence:
|
||||
|
||||
1. **Ranking** (the first failing stage). Candidates:
|
||||
- build the retrieval query so the player's newest input is not drowned out, for
|
||||
example by giving the newest player action its own weight or its own query;
|
||||
- and/or add one inspectable ranking term from CONTEXT-AND-MEMORY §20, most
|
||||
directly a lexical match on the query's rare terms.
|
||||
|
||||
Either must keep `replica_matches_stored_selection` meaningful: a
|
||||
diagnostic-visible score, recorded in `memories.used`.
|
||||
2. **Eviction** (the certain second failure). Stop least-recently-used eviction from
|
||||
discarding a never-again-retrieved early memory first. For example, weight eviction
|
||||
by coverage, keeping the only memory of a story range, or by age, instead of recency
|
||||
alone. The frozen-bank protection must be kept.
|
||||
3. **Creation** (conditional). Choose the summariser's excerpt so that a fact early in
|
||||
a long block is not cut. For example, the head and tail, or the whole block up to a
|
||||
larger bound.
|
||||
|
||||
Each change turns one of B.1's diagnostics into a passing result:
|
||||
- `past_capacity` and `long_block_fact_early` flip their strict xfails;
|
||||
- a real-model re-run shows `ranked: yes` for the planting-era memory.
|
||||
|
||||
## O. What B.2 must not change
|
||||
|
||||
- **Lineage safety.** `tree.attach_memory`, the path clause, `forget_node` and E02
|
||||
stay as they are. An abandoned line's memory stays stored and ineligible (§I).
|
||||
- **Summary lineage** (E03) and summary content policy.
|
||||
- **Authority.** Memory never writes state and is never framed as canon (F07, §J).
|
||||
- **Imported-knowledge authority and retrieval.**
|
||||
- **Pins.** Pinned memories stay always-selected and never evicted.
|
||||
- **The frozen-bank fix.** A memory is never evicted by the pass that created it.
|
||||
- **The single-commit use counter** (M11 §O.7).
|
||||
- **F01-F08, E01-E04, the M04 verdicts, the bundle format and the schema**, unless a
|
||||
migration is separately justified.
|
||||
- **The deterministic diagnostic itself.** B.2 flips the strict xfails. It does not
|
||||
weaken the scenarios or the isolation checks.
|
||||
|
||||
---
|
||||
|
||||
## P. Tests / regression
|
||||
|
||||
| Run | Result |
|
||||
| --- | --- |
|
||||
| `test_v11_b1_memory_diagnostic.py` on HEAD | **24 passed, 2 xfailed (strict)** |
|
||||
| `test_v11_b1_memory_diagnostic.py` on v1.0.0 (§D) | **24 passed, 2 xfailed (strict)**, identical |
|
||||
| `test_v11_b1_long_run_verdict.py`, `test_m11_long_run_memory.py`, `test_m11_long_run_resume.py` | **52 passed** |
|
||||
| **Full backend suite**, HEAD plus the B.1 files | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,323 s). The 17 skips are the tests that need a real model, the same 17 as before. The 2 strict xfails are the two diagnosed retention criteria (§D). |
|
||||
| **Full backend suite, final re-run at staging** (after the real-model attempt; the staged tree) | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,912 s), identical |
|
||||
|
||||
No frontend file was changed, so the frontend suite, lint and build are not affected.
|
||||
|
||||
---
|
||||
|
||||
## Q. Compatibility
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| Application code changed | **none.** `git diff --stat HEAD -- backend/app frontend` is empty |
|
||||
| Schema migration | none |
|
||||
| Bundle format | unchanged |
|
||||
| Stored campaign behaviour | unchanged |
|
||||
| Memory behaviour | **unchanged.** Creation, eviction, ranking, pins and lineage are all as in v1.0.0. `memorybank.py` is byte-identical to the tag |
|
||||
| What changed | Diagnostic tooling (`tools/memory_diagnostic.py`, `tools/v11_b1_memory.py`), an opt-in harness mode (`m11_long_run.py --independent-fact`, with every existing behaviour and M04 verdict unchanged when the flag is off) and tests |
|
||||
| New strict xfails | 2 (§D). They document the two diagnosed defects, and the suite stays green. B.2 must remove them deliberately when it fixes the mechanisms |
|
||||
|
||||
## R. Security / local-only
|
||||
|
||||
Checked against the B.1 diff: the `m11_long_run.py` changes plus the four new files.
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| New network client, endpoint, URL or TLS setting | **none.** The only address in the new code is `http://127.0.0.1:9/v1`, a refused loopback port the deterministic scenarios configure so nothing is contacted |
|
||||
| External embedding service or remote vector store | **none.** The deterministic runs use `ConceptEmbedder` in-process. The real-model run uses the configured Ollama embedding model through the application's existing provider |
|
||||
| Endpoint policy (`endpoints.py`, ADR 011) and TLS (`tlstrust.py`) | unchanged; not in the diff |
|
||||
| Network calls in the real-model run | only the configured Ollama host on the trusted LAN, through the application's own provider and probe paths |
|
||||
| Real identifiers in committed files | none. Checked again at staging |
|
||||
| Offline container regression | **not run.** The harness builds and runs a Docker container, and this session's permission policy refused it. B.1 changes no runtime code and nothing in the image (the image carries `backend/app` and `frontend/dist` only), so the container would be byte-identical to the A1/A2 image. That image passed 23/23 on the corrective tree (`V1.1-WP-A1-A2-REPORT.md` §R.3) |
|
||||
|
||||
---
|
||||
|
||||
## S. Final decision
|
||||
|
||||
**What was established.** Every B.1 requirement was carried out except the clean
|
||||
real-model run, which was attempted:
|
||||
- the current mechanism (§B);
|
||||
- a deterministic, isolated harness with four stage outputs and a verdict at recall
|
||||
depth ≥ 100 (§C);
|
||||
- the v1.0.0 baseline (§D);
|
||||
- capacity and eviction (§F), the creation window (§E) and ranking (§G);
|
||||
- the lineage and authority negative controls (§I, §J);
|
||||
- the `recovered_through_memory_independent` long-run verdict, with the M04 verdicts
|
||||
unchanged;
|
||||
- tests and regression (§P), compatibility (§Q) and security (§R).
|
||||
|
||||
**The real-model gap.** The one real-model attempt ran cleanly, but failed isolation
|
||||
(`precondition_failed:absent_from_summary`). The narrator and summariser restated the
|
||||
fact, so a clean memory-only result on a real model was **not obtained**. The owner
|
||||
decided not to make a second attempt, because the same narrator behaviour would very
|
||||
likely repeat. The attempt is reported in full, and its stage diagnosis is used as
|
||||
mechanism evidence only (§K).
|
||||
|
||||
**B.1 DIAGNOSTIC: COMPLETE**
|
||||
|
||||
**FIRST FAILING STAGE: RANKING.** In a real 100-turn campaign the early fact's memory
|
||||
was created (it carries the fact) and retained (active, 33 of 80), but ranked 10th of
|
||||
19 (similarity 0.866) against `top_k` 4. Later memories with near-identical wording
|
||||
outranked it, under a query made of recent narration (§K, §L). Two further failures
|
||||
are proven deterministically, identically on v1.0.0:
|
||||
- **retention**, past `memory_bank_capacity`: least-recently-used eviction removes an
|
||||
early memory first;
|
||||
- **creation**, for a fact early in a block over 2,000 tokens: the summariser's
|
||||
last-2,000-token excerpt drops it.
|
||||
|
||||
**WP-B.2 RECOMMENDED CHANGE: retrieval ranking first.** Build the retrieval query so
|
||||
the newest player input is not diluted by three turns of recent narration, and add one
|
||||
inspectable lexical term for the query's rare words to cosine ranking. The score must
|
||||
be recorded in `memories.used`. Acceptance: a real-model re-run shows `ranked: yes`
|
||||
for the planting-era memory.
|
||||
|
||||
Then, in separate test-first steps:
|
||||
1. make eviction coverage-aware instead of purely least-recently-used, so the only
|
||||
memory of an early range is not discarded first (flips
|
||||
`test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity`);
|
||||
2. choose the summariser excerpt so a fact early in a long block is kept (flips
|
||||
`test_acceptance_a_fact_early_in_a_long_block_is_remembered`).
|
||||
|
||||
Everything in §O stays unchanged. B.2 has not been started.
|
||||
Reference in New Issue
Block a user