v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
beb17ada10
commit
0c1ba836ba
+450
-76
@@ -17,9 +17,10 @@ database session. It does three things:
|
||||
then evicts the bank down to its capacity. Evicted memories are marked as
|
||||
forgotten and kept so that the UI can still show them.
|
||||
|
||||
When the app generates a turn, `retrieve_memories` embeds the recent story text
|
||||
and ranks the bank by cosine similarity. The highest-ranked memories become the
|
||||
Memories section of the context.
|
||||
When the app generates a turn, `retrieve_memories` embeds the player's input
|
||||
and the current scene, and ranks the bank by a fixed mix of cosine similarity
|
||||
and rarity-weighted word overlap with the input (v1.1 WP-B.2). The
|
||||
highest-ranked memories become the Memories section of the context.
|
||||
|
||||
Every AI call in this module is best-effort. A failure is logged to the debug
|
||||
page and retried on a later turn, because the cursors advance only after a call
|
||||
@@ -28,6 +29,7 @@ succeeds.
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import math
|
||||
from array import array
|
||||
from collections import OrderedDict
|
||||
|
||||
@@ -36,6 +38,7 @@ from sqlalchemy.orm import Session, defer, object_session
|
||||
|
||||
from . import derived, models, summaries, tree, vectors
|
||||
from .context import (
|
||||
count_tokens,
|
||||
cursors,
|
||||
history,
|
||||
lineage,
|
||||
@@ -43,8 +46,11 @@ from .context import (
|
||||
story_actions,
|
||||
truncate_to_last_tokens,
|
||||
)
|
||||
from .context.builder import _encoding as _token_encoding
|
||||
from .database import SessionLocal
|
||||
from .knowledge import embeddings as knowledge_embeddings
|
||||
from .knowledge import fts
|
||||
from .narrative import model as narrative_model
|
||||
from .providers import OpenAICompatibleProvider, ProviderError
|
||||
from .vectors import cosine # re-exported: the ranking lives here, the maths there
|
||||
|
||||
@@ -55,10 +61,14 @@ MEMORY_START = 12 # first memory once the adventure reaches this many actions
|
||||
SUMMARY_INTERVAL = 15 # actions between Story Summary updates
|
||||
MAX_MEMORIES_PER_RUN = 5 # cap catch-up work (e.g. imported adventures) per turn
|
||||
MAX_EMBED_BATCH = 32
|
||||
RETRIEVAL_WINDOW_TOKENS = 600 # recent story text used as the similarity query
|
||||
RETRIEVAL_WINDOW_ACTIONS = 4 # ...taken from this many of the newest actions
|
||||
SUMMARY_MAX_WORDS = 250
|
||||
MEMORY_EXCERPT_TOKENS = 2000 # of the block, when a block is longer than this
|
||||
MEMORY_EXCERPT_TOKENS = 2000 # the most of a block the summariser is shown
|
||||
|
||||
# v1.1 WP-B.2: what stands between the two parts of a block too long to send
|
||||
# whole. It says a part is missing, so the summariser does not read the end as
|
||||
# following straight on from the opening, and `summarize_block` removes it from
|
||||
# anything the model repeats back.
|
||||
EXCERPT_OMISSION_MARKER = "[… the middle of this stretch of story is left out here …]"
|
||||
|
||||
# How much story has to sit past a block before that block is summarized.
|
||||
#
|
||||
@@ -278,9 +288,10 @@ def set_vector(memory: models.Memory, vector: list[float] | None) -> None:
|
||||
"""
|
||||
memory.embedding_blob = None if vector is None else vectors.pack(vector)
|
||||
memory.embedded = vector is not None
|
||||
cached = _vector_cache.get(memory.adventure_id)
|
||||
if cached is not None:
|
||||
cached.pop(memory.id, None)
|
||||
for cache in (_vector_cache, _terms_cache):
|
||||
cached = cache.get(memory.adventure_id)
|
||||
if cached is not None:
|
||||
cached.pop(memory.id, None)
|
||||
|
||||
|
||||
# ---------- The vector cache ----------
|
||||
@@ -309,6 +320,37 @@ _vector_cache: OrderedDict[int, dict[int, array]] = OrderedDict()
|
||||
VECTOR_CACHE_ADVENTURES = 8 # ~600 KB each at a 100-memory bank
|
||||
|
||||
|
||||
# v1.1 WP-B.2: each memory's lexical terms, held the same way and by the same
|
||||
# two rules as its vector. `set_vector` is also where a memory's text changes
|
||||
# (an edit clears the vector to re-embed it), so dropping the entry there covers
|
||||
# a rewritten text as well as a rewritten vector. Text is read only for memories
|
||||
# not already held, and only on a turn whose input has words to match.
|
||||
_terms_cache: OrderedDict[int, dict[int, frozenset[str]]] = OrderedDict()
|
||||
|
||||
|
||||
def _terms_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, frozenset[str]]:
|
||||
"""The lexical terms for `ids`, reading text only for the ones not already held."""
|
||||
cached = _terms_cache.get(adventure_id)
|
||||
if cached is None:
|
||||
cached = _terms_cache[adventure_id] = {}
|
||||
_terms_cache.move_to_end(adventure_id)
|
||||
while len(_terms_cache) > VECTOR_CACHE_ADVENTURES:
|
||||
_terms_cache.popitem(last=False)
|
||||
|
||||
wanted = set(ids)
|
||||
for gone in set(cached) - wanted:
|
||||
del cached[gone]
|
||||
missing = [memory_id for memory_id in ids if memory_id not in cached]
|
||||
if missing:
|
||||
rows = db.execute(
|
||||
select(models.Memory.id, models.Memory.text)
|
||||
.where(models.Memory.id.in_(missing))
|
||||
).all()
|
||||
for memory_id, text in rows:
|
||||
cached[memory_id] = lexical_terms(text or "")
|
||||
return cached
|
||||
|
||||
|
||||
def forget_cached_vectors(adventure_id: int) -> None:
|
||||
"""Drops an adventure's cached vectors.
|
||||
|
||||
@@ -316,6 +358,7 @@ def forget_cached_vectors(adventure_id: int) -> None:
|
||||
corrects itself, as described in the comment above.
|
||||
"""
|
||||
_vector_cache.pop(adventure_id, None)
|
||||
_terms_cache.pop(adventure_id, None)
|
||||
|
||||
|
||||
def _vectors_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, array]:
|
||||
@@ -535,6 +578,199 @@ def cast_brief(adventure: models.Adventure, text: str) -> str:
|
||||
|
||||
# ---------- Retrieval (runs inside the turn, before build_context) ----------
|
||||
|
||||
# v1.1 WP-B.2: what the retrieval query is made of, and how a memory is scored
|
||||
# against it (CONTEXT-AND-MEMORY §18, §20).
|
||||
#
|
||||
# WP-B.1 measured the v1.0.0 query, the newest four actions cut to 600 tokens,
|
||||
# against a planted early fact. The player's one-line question arrived after
|
||||
# three turns of narration, so the embedding mostly described the narration: a
|
||||
# direct question about the fact fell from cosine 0.708 on its own to 0.241 in
|
||||
# that query, and a real 100-turn campaign ranked the only memory of the fact
|
||||
# 10th of 19 against a `memory_top_k` of 4.
|
||||
#
|
||||
# The query is now two short texts, embedded in one call:
|
||||
#
|
||||
# input the player's own action this turn, when there is one
|
||||
# context the current scene from the authoritative state (summary, location,
|
||||
# who is present), then the end of the newest narration
|
||||
#
|
||||
# The context is still there because a question often cannot be read without
|
||||
# it ("I ask her where she hid it"), and §18 says retrieval must not rely on raw
|
||||
# input alone. It is bounded so it can resolve a reference but cannot outweigh
|
||||
# the question by sheer length.
|
||||
#
|
||||
# A memory's score is
|
||||
#
|
||||
# semantic_score = INPUT_WEIGHT * cos(input, memory)
|
||||
# + (1 - INPUT_WEIGHT) * cos(context, memory)
|
||||
# lexical_score = rarity-weighted share of the input's words the memory holds
|
||||
# final_score = semantic_score + LEXICAL_WEIGHT * lexical_score
|
||||
#
|
||||
# With no player input (a continue, or a dry run from Insights) the semantic
|
||||
# score is the context cosine alone and the lexical score is 0. Pins are
|
||||
# unchanged: a pinned memory is always used and counts toward `memory_top_k`.
|
||||
INPUT_TYPES = ("do", "say", "story") # player actions that carry words to search for
|
||||
QUERY_INPUT_TOKENS = 200 # of the player's action; a long `story` entry is cut
|
||||
QUERY_SCENE_TOKENS = 60 # of the state's scene line
|
||||
QUERY_NARRATION_TOKENS = 120 # from the end of the newest narration
|
||||
INPUT_WEIGHT = 0.6
|
||||
# Chosen by sweep (0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5) over the deterministic
|
||||
# ranking fixtures, recorded in the WP-B.2 report (§C, §D). The two-part query
|
||||
# alone already ranks the planting-era memory first; 0.15 is the smallest weight
|
||||
# at which the lexical term by itself also lifts it into `memory_top_k` against
|
||||
# the v1.0.0 narration-filled query, and no rare-word negative control put an
|
||||
# unrelated memory above it. At 0.5 an incidental shared word was enough to
|
||||
# select it for an unrelated question, which is the failure a larger weight buys.
|
||||
LEXICAL_WEIGHT = 0.15
|
||||
# `fts.terms` drops these already; the plural fold below is the only stemming.
|
||||
_MIN_FOLD_LENGTH = 5
|
||||
|
||||
|
||||
def _fold(word: str) -> str:
|
||||
"""One term, reduced so "shelves'" and "shelf" do not meet, but "teapots"
|
||||
and "teapot" do. Possessives lose their `'s`, and a trailing `s` goes from a
|
||||
word long enough to be a plural and not ending in `ss`. Deliberately no more
|
||||
than that: a stemmer is a dependency, and a wrong fold merges two words."""
|
||||
word = word.split("'", 1)[0]
|
||||
if len(word) >= _MIN_FOLD_LENGTH and word.endswith("s") and not word.endswith("ss"):
|
||||
word = word[:-1]
|
||||
return word
|
||||
|
||||
|
||||
def lexical_terms(text: str) -> frozenset[str]:
|
||||
"""The words of `text` that lexical matching compares, folded.
|
||||
|
||||
The tokenizer and stop list are imported knowledge's (`knowledge.fts`), so
|
||||
the two retrieval paths agree on what a word is.
|
||||
"""
|
||||
return frozenset(t for t in (_fold(w) for w in fts.terms(text)) if len(t) >= fts.MIN_TERM_LENGTH)
|
||||
|
||||
|
||||
def lexical_scores(input_terms: frozenset[str], terms_of: dict[int, frozenset[str]]) -> dict[int, float]:
|
||||
"""Each candidate's share of the input's rarity, in [0, 1].
|
||||
|
||||
A term's weight is `ln((N + 1) / (df + 1))`: N candidates, df of them holding
|
||||
it. A word every candidate holds weighs exactly 0, so a protagonist's name or
|
||||
a word the whole bank shares moves nothing, and a word no candidate holds
|
||||
weighs the most. The share is taken over **all** the input's terms, so a
|
||||
memory that happens to hold one rare word of a longer question gets that
|
||||
word's part of the question, not the whole of it. The weights live only for
|
||||
this call, over this candidate set: no index, no stored field.
|
||||
"""
|
||||
if not input_terms or not terms_of:
|
||||
return {memory_id: 0.0 for memory_id in terms_of}
|
||||
n = len(terms_of)
|
||||
weight = {
|
||||
term: math.log((n + 1) / (sum(1 for terms in terms_of.values() if term in terms) + 1))
|
||||
for term in input_terms
|
||||
}
|
||||
total = sum(weight.values())
|
||||
if total <= 0:
|
||||
return {memory_id: 0.0 for memory_id in terms_of}
|
||||
return {
|
||||
memory_id: min(1.0, sum(w for term, w in weight.items() if term in terms) / total)
|
||||
for memory_id, terms in terms_of.items()
|
||||
}
|
||||
|
||||
|
||||
def _scene_text(state) -> str:
|
||||
"""The scene as the authoritative state has it: summary, location, who is present.
|
||||
|
||||
Names only, read straight off the document. The full entity list is left
|
||||
out on purpose: a campaign with a large cast would turn every query into a
|
||||
search for everyone.
|
||||
"""
|
||||
if not isinstance(state, dict):
|
||||
return ""
|
||||
scene = state.get("scene")
|
||||
if not isinstance(scene, dict):
|
||||
return ""
|
||||
pieces: list[str] = []
|
||||
summary = scene.get("summary")
|
||||
if isinstance(summary, str) and summary.strip():
|
||||
pieces.append(summary.strip())
|
||||
location = scene.get("location")
|
||||
if isinstance(location, str) and location.strip():
|
||||
pieces.append(narrative_model.entity_name(state, location.strip()))
|
||||
present = scene.get("present")
|
||||
if isinstance(present, list):
|
||||
names = [narrative_model.entity_name(state, key) for key in present[:8]
|
||||
if isinstance(key, str) and key.strip()]
|
||||
if names:
|
||||
pieces.append(", ".join(names))
|
||||
return truncate_to_last_tokens(". ".join(pieces), QUERY_SCENE_TOKENS)
|
||||
|
||||
|
||||
def retrieval_query(adventure: models.Adventure, exclude_action_id: int | None = None) -> dict:
|
||||
"""The two texts a turn's memory retrieval embeds, and the words it matches.
|
||||
|
||||
Returns `{"input", "context", "input_terms"}`. `input` is empty when the
|
||||
newest action is not a player action with text, which is a continue turn or a
|
||||
dry run. `context` is empty only for a story with no scene and no narration.
|
||||
"""
|
||||
recent = history.tail(adventure, 2, exclude_action_id)
|
||||
newest = recent[-1] if recent else None
|
||||
player_input = ""
|
||||
if newest is not None and newest.type in INPUT_TYPES:
|
||||
player_input = truncate_to_last_tokens(newest.text.strip(), QUERY_INPUT_TOKENS)
|
||||
narration = recent[0].text if len(recent) > 1 else ""
|
||||
else:
|
||||
narration = newest.text if newest is not None else ""
|
||||
context = "\n".join(part for part in (
|
||||
_scene_text(adventure.narrative_state),
|
||||
truncate_to_last_tokens(narration.strip(), QUERY_NARRATION_TOKENS),
|
||||
) if part.strip())
|
||||
return {
|
||||
"input": player_input,
|
||||
"context": context,
|
||||
"input_terms": sorted(lexical_terms(player_input)),
|
||||
}
|
||||
|
||||
|
||||
def score_candidates(
|
||||
ids: list[int],
|
||||
held: dict,
|
||||
terms_of: dict[int, frozenset[str]],
|
||||
input_vec,
|
||||
context_vec,
|
||||
input_terms,
|
||||
) -> list[tuple[float, int, float, float]]:
|
||||
"""`(final_score, memory_id, semantic_score, lexical_score)`, best first.
|
||||
|
||||
Ties on the final score are broken by id, so the order never depends on the
|
||||
order the database returned rows in.
|
||||
"""
|
||||
lexical = lexical_scores(frozenset(input_terms), {i: terms_of.get(i, frozenset()) for i in ids})
|
||||
rows = []
|
||||
for memory_id in ids:
|
||||
vector = held[memory_id]
|
||||
if input_vec is not None and context_vec is not None:
|
||||
semantic = (INPUT_WEIGHT * cosine(input_vec, vector)
|
||||
+ (1.0 - INPUT_WEIGHT) * cosine(context_vec, vector))
|
||||
else:
|
||||
semantic = cosine(input_vec if input_vec is not None else context_vec, vector)
|
||||
lex = lexical.get(memory_id, 0.0)
|
||||
rows.append((semantic + LEXICAL_WEIGHT * lex, memory_id, semantic, lex))
|
||||
rows.sort(key=lambda row: (-row[0], row[1]))
|
||||
return rows
|
||||
|
||||
|
||||
def select_memories(scored, pinned_of, held, authority_of, top_k):
|
||||
"""Pins first, then the best-scoring rest, skipping repeats (§22).
|
||||
|
||||
Returns `(used, suppressed)`. `used` is `(final_score, memory_id, pinned)`
|
||||
rows, best first.
|
||||
"""
|
||||
rows = [(final, memory_id, pinned_of[memory_id]) for final, memory_id, _, _ in scored]
|
||||
used = [row for row in rows if row[2]]
|
||||
remaining = max(0, top_k - len(used))
|
||||
candidates = [row for row in rows if not row[2]]
|
||||
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
|
||||
used += kept
|
||||
used.sort(key=lambda row: (-row[0], row[1]))
|
||||
return used, suppressed
|
||||
|
||||
|
||||
async def retrieve_memories(
|
||||
adventure: models.Adventure,
|
||||
settings: models.Settings,
|
||||
@@ -544,16 +780,17 @@ async def retrieve_memories(
|
||||
"""Returns the memories to inject, or None when the bank is off.
|
||||
|
||||
The result is a dict of the form
|
||||
`{"used": [{id, text, similarity, pinned}], "error": str | None}`. It is
|
||||
None when the memory bank is disabled for this adventure.
|
||||
`{"used": [{id, text, similarity, semantic_score, lexical_score,
|
||||
final_score, pinned, authority, source}], "query": {...}, "error": str | None}`.
|
||||
`similarity` is the semantic score, under the name the inspector has always
|
||||
shown. It is None when the memory bank is disabled for this adventure.
|
||||
|
||||
This only reads. A turn counts the memories it used with `record_use`, just
|
||||
before the commit that saves the turn; see that function for why the count
|
||||
cannot be written here.
|
||||
|
||||
`exclude_action_id` removes the action being retried from the similarity
|
||||
query, so that a discarded attempt cannot influence which memories are
|
||||
returned.
|
||||
`exclude_action_id` removes the action being retried from the query, so that
|
||||
a discarded attempt cannot influence which memories are returned.
|
||||
"""
|
||||
if not adventure.memory_bank_enabled:
|
||||
return None
|
||||
@@ -585,39 +822,33 @@ async def retrieve_memories(
|
||||
if not catalogue:
|
||||
return {"used": [], "error": None}
|
||||
|
||||
recent = history.tail(adventure, RETRIEVAL_WINDOW_ACTIONS, exclude_action_id)
|
||||
query = truncate_to_last_tokens(
|
||||
"\n\n".join(a.text for a in recent), RETRIEVAL_WINDOW_TOKENS
|
||||
)
|
||||
if not query.strip():
|
||||
query = retrieval_query(adventure, exclude_action_id)
|
||||
texts = [t for t in (query["input"], query["context"]) if t.strip()]
|
||||
if not texts:
|
||||
return {"used": [], "error": None}
|
||||
|
||||
try:
|
||||
[query_vec] = await embedding_provider(settings).embed([query])
|
||||
embedded = await embedding_provider(settings).embed(texts)
|
||||
except ProviderError as exc:
|
||||
return {"used": [], "error": str(exc)}
|
||||
vectors_by_text = dict(zip(texts, embedded))
|
||||
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
|
||||
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
|
||||
|
||||
held = _vectors_for(db, adventure.id, [memory_id for memory_id, _, _ in catalogue])
|
||||
ids = [memory_id for memory_id, _, _ in catalogue]
|
||||
held = _vectors_for(db, adventure.id, ids)
|
||||
# Memory text is read only when there are input words to match against, and
|
||||
# then only for memories not already held (see `_terms_for`).
|
||||
terms_of = _terms_for(db, adventure.id, ids) if query["input_terms"] else {}
|
||||
authority_of = {memory_id: authority for memory_id, _, authority in catalogue}
|
||||
scored = sorted(
|
||||
(
|
||||
(cosine(query_vec, held[memory_id]), memory_id, pinned)
|
||||
for memory_id, pinned, _ in catalogue
|
||||
if memory_id in held
|
||||
),
|
||||
key=lambda row: row[0],
|
||||
reverse=True,
|
||||
pinned_of = {memory_id: pinned for memory_id, pinned, _ in catalogue}
|
||||
scored = score_candidates(
|
||||
[memory_id for memory_id in ids if memory_id in held],
|
||||
held, terms_of, input_vec, context_vec, query["input_terms"],
|
||||
)
|
||||
# Pinned memories are always used, and they count toward `top_k`, so the
|
||||
# injected set stays within the budget unless the pinned memories alone
|
||||
# exceed it.
|
||||
top_k = max(1, settings.memory_top_k)
|
||||
used = [row for row in scored if row[2]]
|
||||
remaining = max(0, top_k - len(used))
|
||||
candidates = [row for row in scored if not row[2]]
|
||||
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
|
||||
used += kept
|
||||
used.sort(key=lambda row: row[0], reverse=True)
|
||||
components = {memory_id: (semantic, lex) for _, memory_id, semantic, lex in scored}
|
||||
used, suppressed = select_memories(
|
||||
scored, pinned_of, held, authority_of, max(1, settings.memory_top_k))
|
||||
if not used:
|
||||
return {"used": [], "error": None}
|
||||
|
||||
@@ -638,14 +869,19 @@ async def retrieve_memories(
|
||||
).where(models.Memory.id.in_(used_ids))
|
||||
).all()
|
||||
}
|
||||
texts = {memory_id: row.text for memory_id, row in detail.items()}
|
||||
texts_of = {memory_id: row.text for memory_id, row in detail.items()}
|
||||
|
||||
return {
|
||||
"used": [
|
||||
{
|
||||
"id": memory_id,
|
||||
"text": texts.get(memory_id, ""),
|
||||
"similarity": round(score, 4),
|
||||
"text": texts_of.get(memory_id, ""),
|
||||
"similarity": round(components[memory_id][0], 4),
|
||||
# v1.1 WP-B.2: the parts of the score, so an inspector can see
|
||||
# why this memory beat the ones below it.
|
||||
"semantic_score": round(components[memory_id][0], 4),
|
||||
"lexical_score": round(components[memory_id][1], 4),
|
||||
"final_score": round(final, 4),
|
||||
"pinned": pinned,
|
||||
# M6: what weight this carries, and where it came from.
|
||||
"authority": getattr(detail.get(memory_id), "authority", ACCEPTED_STORY),
|
||||
@@ -656,7 +892,7 @@ async def retrieve_memories(
|
||||
"source_end": getattr(detail.get(memory_id), "source_end", None),
|
||||
},
|
||||
}
|
||||
for score, memory_id, pinned in used
|
||||
for final, memory_id, pinned in used
|
||||
],
|
||||
"considered": len(catalogue),
|
||||
# M6: how many candidates were set aside as repeating one already
|
||||
@@ -665,6 +901,15 @@ async def retrieve_memories(
|
||||
{"id": memory_id, "duplicate_of": kept_id}
|
||||
for memory_id, kept_id in suppressed
|
||||
],
|
||||
# v1.1 WP-B.2: what was searched for. Recorded per turn, like the rest.
|
||||
"query": {
|
||||
"input": query["input"],
|
||||
"context": query["context"],
|
||||
"input_terms": query["input_terms"],
|
||||
"input_weight": INPUT_WEIGHT if input_vec is not None and context_vec is not None
|
||||
else (1.0 if input_vec is not None else 0.0),
|
||||
"lexical_weight": LEXICAL_WEIGHT,
|
||||
},
|
||||
"error": None,
|
||||
}
|
||||
|
||||
@@ -822,6 +1067,61 @@ async def _guarded(db: Session, adventure_id: int, kind: str, coro) -> None:
|
||||
db.commit()
|
||||
|
||||
|
||||
def _excerpt_encoding():
|
||||
return _token_encoding()
|
||||
|
||||
|
||||
def excerpt_split(budget: int = MEMORY_EXCERPT_TOKENS) -> tuple[int, int]:
|
||||
"""`(head_tokens, tail_tokens)` for a block longer than `budget`.
|
||||
|
||||
The marker and the blank lines around it are paid for first; what is left is
|
||||
halved, and an odd token goes to the tail, the most recent part. So the two
|
||||
parts plus the marker come to exactly `budget`.
|
||||
"""
|
||||
room = max(0, budget - count_tokens(f"\n\n{EXCERPT_OMISSION_MARKER}\n\n"))
|
||||
head = room // 2
|
||||
return head, room - head
|
||||
|
||||
|
||||
def memory_excerpt(raw: str, budget: int = MEMORY_EXCERPT_TOKENS) -> str:
|
||||
"""What the summariser is shown of one block.
|
||||
|
||||
v1.1 WP-B.2. A block that fits in `budget` tokens is sent whole, exactly as
|
||||
before. A longer block used to be cut to its last `budget` tokens, and B.1
|
||||
showed that a fact near its start then never reached the summariser at all.
|
||||
It is now sent as its opening and its end, in order, with
|
||||
`EXCERPT_OMISSION_MARKER` between them, still inside `budget`.
|
||||
|
||||
Rejoining two token runs can tokenise a little differently at the seams, so
|
||||
the result is measured, and the head gives up tokens until it fits. A fact in
|
||||
the middle of a very long block is still left out: this bounds the input, it
|
||||
does not summarise everything.
|
||||
"""
|
||||
enc = _excerpt_encoding()
|
||||
tokens = enc.encode(raw)
|
||||
if len(tokens) <= budget:
|
||||
return raw
|
||||
head_n, tail_n = excerpt_split(budget)
|
||||
while True:
|
||||
excerpt = (f"{enc.decode(tokens[:head_n]).rstrip()}\n\n{EXCERPT_OMISSION_MARKER}\n\n"
|
||||
f"{enc.decode(tokens[-tail_n:]).lstrip()}" if tail_n else
|
||||
enc.decode(tokens[:head_n]))
|
||||
over = count_tokens(excerpt) - budget
|
||||
if over <= 0 or head_n == 0:
|
||||
return excerpt
|
||||
head_n = max(0, head_n - over)
|
||||
|
||||
|
||||
def memory_user_prompt(brief: str, excerpt: str) -> str:
|
||||
"""The user message of a memory call: the cast brief, then the excerpt.
|
||||
|
||||
Kept apart from `summarize_block` so an evaluation can send a model exactly
|
||||
what the application sends (v1.1 WP-B.2, `tools/memory_fidelity.py`).
|
||||
"""
|
||||
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
|
||||
return f"{brief}\n\n{prompt}" if brief else prompt
|
||||
|
||||
|
||||
async def summarize_block(
|
||||
adventure: models.Adventure,
|
||||
provider: OpenAICompatibleProvider,
|
||||
@@ -840,15 +1140,16 @@ async def summarize_block(
|
||||
old text in place and moves on.
|
||||
"""
|
||||
raw = "\n\n".join(a.text for a in block)
|
||||
excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)
|
||||
excerpt = memory_excerpt(raw)
|
||||
# Match the cast against the untruncated block. The excerpt is what the
|
||||
# model reads, but a character named in the part that was trimmed is still
|
||||
# model reads, but a character named in the part that was left out is still
|
||||
# one the memory may have to name.
|
||||
brief = cast_brief(adventure, raw)
|
||||
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
|
||||
return await provider.complete(
|
||||
MEMORY_SYSTEM_PROMPT, f"{brief}\n\n{prompt}" if brief else prompt
|
||||
)
|
||||
text = await provider.complete(MEMORY_SYSTEM_PROMPT, memory_user_prompt(brief, excerpt))
|
||||
# The marker is an instruction to the summariser, never a fact of the story.
|
||||
if text and EXCERPT_OMISSION_MARKER in text:
|
||||
text = " ".join(text.replace(EXCERPT_OMISSION_MARKER, " ").split())
|
||||
return text
|
||||
|
||||
|
||||
async def _create_due_memories(
|
||||
@@ -1028,11 +1329,93 @@ async def _embed_pending(
|
||||
return len(pending)
|
||||
|
||||
|
||||
def eviction_order(rows, limit: int) -> list[int]:
|
||||
"""The ids eviction would take, first to last, at most `limit` of them.
|
||||
|
||||
v1.1 WP-B.2. `rows` are the active memories of one adventure, each with
|
||||
`id`, `pinned`, `source_start`, `source_end`, `last_used_at`, `created_at`
|
||||
and `use_count`. Nothing here reads a vector or the database, so the same
|
||||
function is what the eviction pass runs and what a diagnostic reports.
|
||||
|
||||
WP-B.1 showed what pure least-recently-used order does to a long campaign.
|
||||
Retrieval is steered by the present scene, so a memory of an early stretch
|
||||
nothing recent resembles stops being used. It then becomes the least
|
||||
recently used row, and it goes first, while the bank keeps several memories
|
||||
of the last few scenes that the history window still holds in full. The
|
||||
rule below keeps the bank spread over the whole story instead.
|
||||
|
||||
**Coverage.** Memories with a source range say which stretch of the story
|
||||
they describe. A memory is judged by the hole its removal would leave: the
|
||||
number of depths between the end of the nearest memory before it and the
|
||||
start of the nearest memory after it. The smallest hole goes first, so the
|
||||
bank thins where it is densest. A memory whose start another memory shares
|
||||
(a retried or re-played stretch, or a sibling line) leaves no hole, and is
|
||||
the first kind to go. Pinned memories count as coverage, since they stay.
|
||||
|
||||
**Boundaries.** The earliest and the latest memory by position leave a hole
|
||||
with no memory on one side: removing the first loses the only record of the
|
||||
opening, and removing the last loses the only record of the most recent
|
||||
stretch, which is also what keeps a memory written this turn from being
|
||||
evicted by the pass that wrote it (the frozen bank, below). Boundaries are
|
||||
not coverage candidates.
|
||||
|
||||
**Recency.** Among memories whose removal leaves the same hole, the least
|
||||
recently used goes first (`coalesce(last_used_at, created_at)`), then the
|
||||
less used, then the lower id. Ties are therefore never left to the order the
|
||||
database returned rows in.
|
||||
|
||||
**Fallback.** When no memory is a coverage candidate — memories typed by the
|
||||
player or migrated from before coordinates have no range, and a bank can be
|
||||
all boundaries — the rest are taken least recently used first, exactly as
|
||||
v1.0.0 did. The bank stays bounded either way. Pinned memories are never
|
||||
taken; if every active memory is pinned, capacity yields to the pins.
|
||||
|
||||
Recomputed after each pick, because removing one memory widens the holes
|
||||
of its neighbours.
|
||||
"""
|
||||
remaining = {row.id: row for row in rows if not row.pinned}
|
||||
coverers = {row.id: row for row in rows
|
||||
if row.source_start is not None and row.source_end is not None}
|
||||
|
||||
def recency(row):
|
||||
return (row.last_used_at or row.created_at, row.use_count or 0, row.id)
|
||||
|
||||
order: list[int] = []
|
||||
while remaining and len(order) < limit:
|
||||
spans = sorted(coverers.values(), key=lambda r: (r.source_start, r.source_end, r.id))
|
||||
starts: dict[int, int] = {}
|
||||
for row in spans:
|
||||
starts[row.source_start] = starts.get(row.source_start, 0) + 1
|
||||
best = None
|
||||
furthest_end = None # the largest source_end before index i
|
||||
for i, row in enumerate(spans):
|
||||
if row.id in remaining:
|
||||
if starts[row.source_start] > 1:
|
||||
cost = 0
|
||||
elif i == 0 or i == len(spans) - 1:
|
||||
cost = None # a boundary
|
||||
else:
|
||||
cost = max(0, spans[i + 1].source_start - furthest_end - 1)
|
||||
if cost is not None:
|
||||
key = (cost, *recency(row))
|
||||
if best is None or key < best[0]:
|
||||
best = (key, row.id)
|
||||
furthest_end = row.source_end if furthest_end is None else max(furthest_end, row.source_end)
|
||||
if best is None:
|
||||
victim = min(remaining.values(), key=recency).id
|
||||
else:
|
||||
victim = best[1]
|
||||
order.append(victim)
|
||||
del remaining[victim]
|
||||
coverers.pop(victim, None)
|
||||
return order
|
||||
|
||||
|
||||
def _evict_over_capacity(
|
||||
adventure: models.Adventure, settings: models.Settings, db: Session
|
||||
) -> None:
|
||||
# The database performs both the count and the ranking, and returns neither
|
||||
# the rows nor the vectors. Counting by walking `adventure.memories` fetched
|
||||
# The database performs the count, and the rows read for ordering carry
|
||||
# neither text nor vectors. Counting by walking `adventure.memories` fetched
|
||||
# every vector in the bank on every turn, whether or not the bank was over
|
||||
# capacity.
|
||||
in_this_bank = (models.Memory.adventure_id == adventure.id,
|
||||
@@ -1043,32 +1426,23 @@ def _evict_over_capacity(
|
||||
overflow = active - max(1, settings.memory_bank_capacity)
|
||||
if overflow <= 0:
|
||||
return
|
||||
# Evict the least recently used memory first, and use the use count only to
|
||||
# break ties.
|
||||
# v1.1 WP-B.2: the order is `eviction_order`, coverage first and recency
|
||||
# second. It replaces least recently used alone; see that function.
|
||||
#
|
||||
# Ordering by use count first froze the bank. A memory written on this turn
|
||||
# has never been used, so once every other memory had been retrieved at
|
||||
# least once, the new memory held the lowest count in the bank. The same
|
||||
# post-turn run that wrote it then evicted it, one pass after embedding it.
|
||||
# Use counts only increase, so the bank never recovered. An adventure kept
|
||||
# whatever memories it held when the bank first filled, and every later
|
||||
# memory was summarized, marked as forgotten, and never ranked.
|
||||
#
|
||||
# Ordering by recency avoids that. A new memory carries the newest
|
||||
# timestamp, so it is the last row to be evicted rather than the first, and
|
||||
# it remains until other memories are used. Demoting the use count costs
|
||||
# little, because retrieving a useful memory also makes it recent. The two
|
||||
# orderings differ only for memories that were used once and have not been
|
||||
# retrieved since, which are the rows a full bank should evict.
|
||||
doomed = db.execute(
|
||||
select(models.Memory.id)
|
||||
.where(*in_this_bank, models.Memory.pinned.is_(False))
|
||||
.order_by(
|
||||
func.coalesce(models.Memory.last_used_at, models.Memory.created_at),
|
||||
models.Memory.use_count,
|
||||
)
|
||||
.limit(overflow)
|
||||
).scalars().all()
|
||||
# What the old ordering fixed still holds. Ordering by use count first froze
|
||||
# the bank: a memory written on this turn has never been used, so once every
|
||||
# other memory had been retrieved at least once, the new memory held the
|
||||
# lowest count in the bank, and the same post-turn run that wrote it evicted
|
||||
# it. Counts only increase, so the bank never recovered. Under the coverage
|
||||
# rule the newest memory is the latest boundary, so it is not a coverage
|
||||
# candidate, and in the fallback it carries the newest timestamp.
|
||||
rows = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.last_used_at,
|
||||
models.Memory.created_at, models.Memory.use_count)
|
||||
.where(*in_this_bank)
|
||||
).all()
|
||||
doomed = eviction_order(rows, overflow)
|
||||
if not doomed:
|
||||
return # Every active memory is pinned, so the pins override capacity.
|
||||
db.execute(
|
||||
|
||||
Reference in New Issue
Block a user