M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
co-authored by
Claude Opus 5
parent
b7005e6fdd
commit
a6e9c7a32b
@@ -20,6 +20,7 @@ exactly as `test_branch_clause.py` builds it.
|
||||
python -m pytest tests/test_memory_nodes.py -v
|
||||
"""
|
||||
import asyncio
|
||||
import math
|
||||
|
||||
import pytest
|
||||
|
||||
@@ -138,11 +139,26 @@ def forked():
|
||||
nodes[f"C{depth}"] = add_node(db, adventure, c, depth, "C")
|
||||
db.flush()
|
||||
|
||||
# Distinct vectors, equally similar to the query.
|
||||
#
|
||||
# These tests are about *lineage visibility* — which memories a branch can
|
||||
# see. They used to store the same vector in every memory, which was
|
||||
# harmless until M6 added redundancy suppression: four identical vectors are
|
||||
# four copies of one statement as far as retrieval is concerned, so they
|
||||
# collapsed to one and the lineage assertions could no longer be read.
|
||||
#
|
||||
# Each vector below sits at the same angle from the query `(1, 0, 0)`, so
|
||||
# ranking between them is unchanged, and far enough apart from each other
|
||||
# (pairwise cosine -0.28 to 0.36) that none suppresses another.
|
||||
memories = {
|
||||
"shared": add_memory(db, adventure, "on the shared trunk", nodes["A3"]),
|
||||
"sibling": add_memory(db, adventure, "on A's own continuation", nodes["A5"]),
|
||||
"b": add_memory(db, adventure, "on B", nodes["B5"]),
|
||||
"c": add_memory(db, adventure, "on C", nodes["C7"]),
|
||||
"shared": add_memory(db, adventure, "on the shared trunk", nodes["A3"],
|
||||
vector=(0.6, 0.8, 0.0)),
|
||||
"sibling": add_memory(db, adventure, "on A's own continuation", nodes["A5"],
|
||||
vector=(0.6, -0.8, 0.0)),
|
||||
"b": add_memory(db, adventure, "on B", nodes["B5"],
|
||||
vector=(0.6, 0.0, 0.8)),
|
||||
"c": add_memory(db, adventure, "on C", nodes["C7"],
|
||||
vector=(0.6, 0.0, -0.8)),
|
||||
}
|
||||
adventure.head_branch_id = c.id
|
||||
adventure.head_depth = 7
|
||||
@@ -162,6 +178,21 @@ def switch_to(db, adventure, branch_id, tip):
|
||||
db.commit()
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def restore_embedding_provider():
|
||||
"""Puts `memorybank.embedding_provider` back after every test here.
|
||||
|
||||
`retrieved` below replaces it by assignment. Until M6 nothing restored it,
|
||||
so a stub outlived the module and was still installed when a later file ran
|
||||
(`tests/test_provider_wiring.py`, which asserts on the real factory).
|
||||
"""
|
||||
real = memorybank.embedding_provider
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
memorybank.embedding_provider = real
|
||||
|
||||
|
||||
def retrieved(adventure, settings) -> set[str]:
|
||||
memorybank.embedding_provider = lambda s: StubEmbedder()
|
||||
result = asyncio.run(
|
||||
@@ -381,6 +412,8 @@ def deeply_forked():
|
||||
memory_top_k=50,
|
||||
))
|
||||
|
||||
_spread = 2 * math.pi / 14
|
||||
|
||||
def story(title, forks):
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title=title, script_state={}, memory_bank_enabled=True,
|
||||
@@ -404,8 +437,20 @@ def deeply_forked():
|
||||
nodes.append(add_node(db, adventure, branch, depth, "n"))
|
||||
depth += 1
|
||||
db.flush()
|
||||
_placed: list = []
|
||||
for node in nodes[5::6]: # one memory per six actions, as the pass makes them
|
||||
add_memory(db, adventure, f"memory at {node.depth}", node)
|
||||
# A distinct direction per memory, all at the same angle from the
|
||||
# query, so ranking between them is unaffected and M6's redundancy
|
||||
# suppression does not collapse fourteen distinct memories into one.
|
||||
# This test measures bytes fetched, not deduplication.
|
||||
#
|
||||
# Fourteen directions spread evenly around the circle orthogonal to
|
||||
# the query are 2*pi/14 apart; the small shared component keeps the
|
||||
# closest pair at cosine ~0.905, below the 0.93 redundancy bar.
|
||||
angle = _spread * len(_placed)
|
||||
_placed.append(node)
|
||||
add_memory(db, adventure, f"memory at {node.depth}", node,
|
||||
vector=(0.2, 0.98 * math.cos(angle), 0.98 * math.sin(angle)))
|
||||
adventure.head_branch_id = branch.id
|
||||
adventure.head_depth = depth - 1
|
||||
return adventure
|
||||
|
||||
Reference in New Issue
Block a user