v1.1 WP-B.2: independent long-term memory retention

Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-15 11:21:53 -04:00
co-authored by Claude Opus 5
parent beb17ada10
commit 0c1ba836ba
15 changed files with 3782 additions and 186 deletions
+450 -76
View File
@@ -17,9 +17,10 @@ database session. It does three things:
then evicts the bank down to its capacity. Evicted memories are marked as
forgotten and kept so that the UI can still show them.
When the app generates a turn, `retrieve_memories` embeds the recent story text
and ranks the bank by cosine similarity. The highest-ranked memories become the
Memories section of the context.
When the app generates a turn, `retrieve_memories` embeds the player's input
and the current scene, and ranks the bank by a fixed mix of cosine similarity
and rarity-weighted word overlap with the input (v1.1 WP-B.2). The
highest-ranked memories become the Memories section of the context.
Every AI call in this module is best-effort. A failure is logged to the debug
page and retried on a later turn, because the cursors advance only after a call
@@ -28,6 +29,7 @@ succeeds.
import asyncio
import logging
import math
from array import array
from collections import OrderedDict
@@ -36,6 +38,7 @@ from sqlalchemy.orm import Session, defer, object_session
from . import derived, models, summaries, tree, vectors
from .context import (
count_tokens,
cursors,
history,
lineage,
@@ -43,8 +46,11 @@ from .context import (
story_actions,
truncate_to_last_tokens,
)
from .context.builder import _encoding as _token_encoding
from .database import SessionLocal
from .knowledge import embeddings as knowledge_embeddings
from .knowledge import fts
from .narrative import model as narrative_model
from .providers import OpenAICompatibleProvider, ProviderError
from .vectors import cosine # re-exported: the ranking lives here, the maths there
@@ -55,10 +61,14 @@ MEMORY_START = 12 # first memory once the adventure reaches this many actions
SUMMARY_INTERVAL = 15 # actions between Story Summary updates
MAX_MEMORIES_PER_RUN = 5 # cap catch-up work (e.g. imported adventures) per turn
MAX_EMBED_BATCH = 32
RETRIEVAL_WINDOW_TOKENS = 600 # recent story text used as the similarity query
RETRIEVAL_WINDOW_ACTIONS = 4 # ...taken from this many of the newest actions
SUMMARY_MAX_WORDS = 250
MEMORY_EXCERPT_TOKENS = 2000 # of the block, when a block is longer than this
MEMORY_EXCERPT_TOKENS = 2000 # the most of a block the summariser is shown
# v1.1 WP-B.2: what stands between the two parts of a block too long to send
# whole. It says a part is missing, so the summariser does not read the end as
# following straight on from the opening, and `summarize_block` removes it from
# anything the model repeats back.
EXCERPT_OMISSION_MARKER = "[… the middle of this stretch of story is left out here …]"
# How much story has to sit past a block before that block is summarized.
#
@@ -278,9 +288,10 @@ def set_vector(memory: models.Memory, vector: list[float] | None) -> None:
"""
memory.embedding_blob = None if vector is None else vectors.pack(vector)
memory.embedded = vector is not None
cached = _vector_cache.get(memory.adventure_id)
if cached is not None:
cached.pop(memory.id, None)
for cache in (_vector_cache, _terms_cache):
cached = cache.get(memory.adventure_id)
if cached is not None:
cached.pop(memory.id, None)
# ---------- The vector cache ----------
@@ -309,6 +320,37 @@ _vector_cache: OrderedDict[int, dict[int, array]] = OrderedDict()
VECTOR_CACHE_ADVENTURES = 8 # ~600 KB each at a 100-memory bank
# v1.1 WP-B.2: each memory's lexical terms, held the same way and by the same
# two rules as its vector. `set_vector` is also where a memory's text changes
# (an edit clears the vector to re-embed it), so dropping the entry there covers
# a rewritten text as well as a rewritten vector. Text is read only for memories
# not already held, and only on a turn whose input has words to match.
_terms_cache: OrderedDict[int, dict[int, frozenset[str]]] = OrderedDict()
def _terms_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, frozenset[str]]:
"""The lexical terms for `ids`, reading text only for the ones not already held."""
cached = _terms_cache.get(adventure_id)
if cached is None:
cached = _terms_cache[adventure_id] = {}
_terms_cache.move_to_end(adventure_id)
while len(_terms_cache) > VECTOR_CACHE_ADVENTURES:
_terms_cache.popitem(last=False)
wanted = set(ids)
for gone in set(cached) - wanted:
del cached[gone]
missing = [memory_id for memory_id in ids if memory_id not in cached]
if missing:
rows = db.execute(
select(models.Memory.id, models.Memory.text)
.where(models.Memory.id.in_(missing))
).all()
for memory_id, text in rows:
cached[memory_id] = lexical_terms(text or "")
return cached
def forget_cached_vectors(adventure_id: int) -> None:
"""Drops an adventure's cached vectors.
@@ -316,6 +358,7 @@ def forget_cached_vectors(adventure_id: int) -> None:
corrects itself, as described in the comment above.
"""
_vector_cache.pop(adventure_id, None)
_terms_cache.pop(adventure_id, None)
def _vectors_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, array]:
@@ -535,6 +578,199 @@ def cast_brief(adventure: models.Adventure, text: str) -> str:
# ---------- Retrieval (runs inside the turn, before build_context) ----------
# v1.1 WP-B.2: what the retrieval query is made of, and how a memory is scored
# against it (CONTEXT-AND-MEMORY §18, §20).
#
# WP-B.1 measured the v1.0.0 query, the newest four actions cut to 600 tokens,
# against a planted early fact. The player's one-line question arrived after
# three turns of narration, so the embedding mostly described the narration: a
# direct question about the fact fell from cosine 0.708 on its own to 0.241 in
# that query, and a real 100-turn campaign ranked the only memory of the fact
# 10th of 19 against a `memory_top_k` of 4.
#
# The query is now two short texts, embedded in one call:
#
# input the player's own action this turn, when there is one
# context the current scene from the authoritative state (summary, location,
# who is present), then the end of the newest narration
#
# The context is still there because a question often cannot be read without
# it ("I ask her where she hid it"), and §18 says retrieval must not rely on raw
# input alone. It is bounded so it can resolve a reference but cannot outweigh
# the question by sheer length.
#
# A memory's score is
#
# semantic_score = INPUT_WEIGHT * cos(input, memory)
# + (1 - INPUT_WEIGHT) * cos(context, memory)
# lexical_score = rarity-weighted share of the input's words the memory holds
# final_score = semantic_score + LEXICAL_WEIGHT * lexical_score
#
# With no player input (a continue, or a dry run from Insights) the semantic
# score is the context cosine alone and the lexical score is 0. Pins are
# unchanged: a pinned memory is always used and counts toward `memory_top_k`.
INPUT_TYPES = ("do", "say", "story") # player actions that carry words to search for
QUERY_INPUT_TOKENS = 200 # of the player's action; a long `story` entry is cut
QUERY_SCENE_TOKENS = 60 # of the state's scene line
QUERY_NARRATION_TOKENS = 120 # from the end of the newest narration
INPUT_WEIGHT = 0.6
# Chosen by sweep (0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5) over the deterministic
# ranking fixtures, recorded in the WP-B.2 report (§C, §D). The two-part query
# alone already ranks the planting-era memory first; 0.15 is the smallest weight
# at which the lexical term by itself also lifts it into `memory_top_k` against
# the v1.0.0 narration-filled query, and no rare-word negative control put an
# unrelated memory above it. At 0.5 an incidental shared word was enough to
# select it for an unrelated question, which is the failure a larger weight buys.
LEXICAL_WEIGHT = 0.15
# `fts.terms` drops these already; the plural fold below is the only stemming.
_MIN_FOLD_LENGTH = 5
def _fold(word: str) -> str:
"""One term, reduced so "shelves'" and "shelf" do not meet, but "teapots"
and "teapot" do. Possessives lose their `'s`, and a trailing `s` goes from a
word long enough to be a plural and not ending in `ss`. Deliberately no more
than that: a stemmer is a dependency, and a wrong fold merges two words."""
word = word.split("'", 1)[0]
if len(word) >= _MIN_FOLD_LENGTH and word.endswith("s") and not word.endswith("ss"):
word = word[:-1]
return word
def lexical_terms(text: str) -> frozenset[str]:
"""The words of `text` that lexical matching compares, folded.
The tokenizer and stop list are imported knowledge's (`knowledge.fts`), so
the two retrieval paths agree on what a word is.
"""
return frozenset(t for t in (_fold(w) for w in fts.terms(text)) if len(t) >= fts.MIN_TERM_LENGTH)
def lexical_scores(input_terms: frozenset[str], terms_of: dict[int, frozenset[str]]) -> dict[int, float]:
"""Each candidate's share of the input's rarity, in [0, 1].
A term's weight is `ln((N + 1) / (df + 1))`: N candidates, df of them holding
it. A word every candidate holds weighs exactly 0, so a protagonist's name or
a word the whole bank shares moves nothing, and a word no candidate holds
weighs the most. The share is taken over **all** the input's terms, so a
memory that happens to hold one rare word of a longer question gets that
word's part of the question, not the whole of it. The weights live only for
this call, over this candidate set: no index, no stored field.
"""
if not input_terms or not terms_of:
return {memory_id: 0.0 for memory_id in terms_of}
n = len(terms_of)
weight = {
term: math.log((n + 1) / (sum(1 for terms in terms_of.values() if term in terms) + 1))
for term in input_terms
}
total = sum(weight.values())
if total <= 0:
return {memory_id: 0.0 for memory_id in terms_of}
return {
memory_id: min(1.0, sum(w for term, w in weight.items() if term in terms) / total)
for memory_id, terms in terms_of.items()
}
def _scene_text(state) -> str:
"""The scene as the authoritative state has it: summary, location, who is present.
Names only, read straight off the document. The full entity list is left
out on purpose: a campaign with a large cast would turn every query into a
search for everyone.
"""
if not isinstance(state, dict):
return ""
scene = state.get("scene")
if not isinstance(scene, dict):
return ""
pieces: list[str] = []
summary = scene.get("summary")
if isinstance(summary, str) and summary.strip():
pieces.append(summary.strip())
location = scene.get("location")
if isinstance(location, str) and location.strip():
pieces.append(narrative_model.entity_name(state, location.strip()))
present = scene.get("present")
if isinstance(present, list):
names = [narrative_model.entity_name(state, key) for key in present[:8]
if isinstance(key, str) and key.strip()]
if names:
pieces.append(", ".join(names))
return truncate_to_last_tokens(". ".join(pieces), QUERY_SCENE_TOKENS)
def retrieval_query(adventure: models.Adventure, exclude_action_id: int | None = None) -> dict:
"""The two texts a turn's memory retrieval embeds, and the words it matches.
Returns `{"input", "context", "input_terms"}`. `input` is empty when the
newest action is not a player action with text, which is a continue turn or a
dry run. `context` is empty only for a story with no scene and no narration.
"""
recent = history.tail(adventure, 2, exclude_action_id)
newest = recent[-1] if recent else None
player_input = ""
if newest is not None and newest.type in INPUT_TYPES:
player_input = truncate_to_last_tokens(newest.text.strip(), QUERY_INPUT_TOKENS)
narration = recent[0].text if len(recent) > 1 else ""
else:
narration = newest.text if newest is not None else ""
context = "\n".join(part for part in (
_scene_text(adventure.narrative_state),
truncate_to_last_tokens(narration.strip(), QUERY_NARRATION_TOKENS),
) if part.strip())
return {
"input": player_input,
"context": context,
"input_terms": sorted(lexical_terms(player_input)),
}
def score_candidates(
ids: list[int],
held: dict,
terms_of: dict[int, frozenset[str]],
input_vec,
context_vec,
input_terms,
) -> list[tuple[float, int, float, float]]:
"""`(final_score, memory_id, semantic_score, lexical_score)`, best first.
Ties on the final score are broken by id, so the order never depends on the
order the database returned rows in.
"""
lexical = lexical_scores(frozenset(input_terms), {i: terms_of.get(i, frozenset()) for i in ids})
rows = []
for memory_id in ids:
vector = held[memory_id]
if input_vec is not None and context_vec is not None:
semantic = (INPUT_WEIGHT * cosine(input_vec, vector)
+ (1.0 - INPUT_WEIGHT) * cosine(context_vec, vector))
else:
semantic = cosine(input_vec if input_vec is not None else context_vec, vector)
lex = lexical.get(memory_id, 0.0)
rows.append((semantic + LEXICAL_WEIGHT * lex, memory_id, semantic, lex))
rows.sort(key=lambda row: (-row[0], row[1]))
return rows
def select_memories(scored, pinned_of, held, authority_of, top_k):
"""Pins first, then the best-scoring rest, skipping repeats (§22).
Returns `(used, suppressed)`. `used` is `(final_score, memory_id, pinned)`
rows, best first.
"""
rows = [(final, memory_id, pinned_of[memory_id]) for final, memory_id, _, _ in scored]
used = [row for row in rows if row[2]]
remaining = max(0, top_k - len(used))
candidates = [row for row in rows if not row[2]]
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
used += kept
used.sort(key=lambda row: (-row[0], row[1]))
return used, suppressed
async def retrieve_memories(
adventure: models.Adventure,
settings: models.Settings,
@@ -544,16 +780,17 @@ async def retrieve_memories(
"""Returns the memories to inject, or None when the bank is off.
The result is a dict of the form
`{"used": [{id, text, similarity, pinned}], "error": str | None}`. It is
None when the memory bank is disabled for this adventure.
`{"used": [{id, text, similarity, semantic_score, lexical_score,
final_score, pinned, authority, source}], "query": {...}, "error": str | None}`.
`similarity` is the semantic score, under the name the inspector has always
shown. It is None when the memory bank is disabled for this adventure.
This only reads. A turn counts the memories it used with `record_use`, just
before the commit that saves the turn; see that function for why the count
cannot be written here.
`exclude_action_id` removes the action being retried from the similarity
query, so that a discarded attempt cannot influence which memories are
returned.
`exclude_action_id` removes the action being retried from the query, so that
a discarded attempt cannot influence which memories are returned.
"""
if not adventure.memory_bank_enabled:
return None
@@ -585,39 +822,33 @@ async def retrieve_memories(
if not catalogue:
return {"used": [], "error": None}
recent = history.tail(adventure, RETRIEVAL_WINDOW_ACTIONS, exclude_action_id)
query = truncate_to_last_tokens(
"\n\n".join(a.text for a in recent), RETRIEVAL_WINDOW_TOKENS
)
if not query.strip():
query = retrieval_query(adventure, exclude_action_id)
texts = [t for t in (query["input"], query["context"]) if t.strip()]
if not texts:
return {"used": [], "error": None}
try:
[query_vec] = await embedding_provider(settings).embed([query])
embedded = await embedding_provider(settings).embed(texts)
except ProviderError as exc:
return {"used": [], "error": str(exc)}
vectors_by_text = dict(zip(texts, embedded))
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
held = _vectors_for(db, adventure.id, [memory_id for memory_id, _, _ in catalogue])
ids = [memory_id for memory_id, _, _ in catalogue]
held = _vectors_for(db, adventure.id, ids)
# Memory text is read only when there are input words to match against, and
# then only for memories not already held (see `_terms_for`).
terms_of = _terms_for(db, adventure.id, ids) if query["input_terms"] else {}
authority_of = {memory_id: authority for memory_id, _, authority in catalogue}
scored = sorted(
(
(cosine(query_vec, held[memory_id]), memory_id, pinned)
for memory_id, pinned, _ in catalogue
if memory_id in held
),
key=lambda row: row[0],
reverse=True,
pinned_of = {memory_id: pinned for memory_id, pinned, _ in catalogue}
scored = score_candidates(
[memory_id for memory_id in ids if memory_id in held],
held, terms_of, input_vec, context_vec, query["input_terms"],
)
# Pinned memories are always used, and they count toward `top_k`, so the
# injected set stays within the budget unless the pinned memories alone
# exceed it.
top_k = max(1, settings.memory_top_k)
used = [row for row in scored if row[2]]
remaining = max(0, top_k - len(used))
candidates = [row for row in scored if not row[2]]
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
used += kept
used.sort(key=lambda row: row[0], reverse=True)
components = {memory_id: (semantic, lex) for _, memory_id, semantic, lex in scored}
used, suppressed = select_memories(
scored, pinned_of, held, authority_of, max(1, settings.memory_top_k))
if not used:
return {"used": [], "error": None}
@@ -638,14 +869,19 @@ async def retrieve_memories(
).where(models.Memory.id.in_(used_ids))
).all()
}
texts = {memory_id: row.text for memory_id, row in detail.items()}
texts_of = {memory_id: row.text for memory_id, row in detail.items()}
return {
"used": [
{
"id": memory_id,
"text": texts.get(memory_id, ""),
"similarity": round(score, 4),
"text": texts_of.get(memory_id, ""),
"similarity": round(components[memory_id][0], 4),
# v1.1 WP-B.2: the parts of the score, so an inspector can see
# why this memory beat the ones below it.
"semantic_score": round(components[memory_id][0], 4),
"lexical_score": round(components[memory_id][1], 4),
"final_score": round(final, 4),
"pinned": pinned,
# M6: what weight this carries, and where it came from.
"authority": getattr(detail.get(memory_id), "authority", ACCEPTED_STORY),
@@ -656,7 +892,7 @@ async def retrieve_memories(
"source_end": getattr(detail.get(memory_id), "source_end", None),
},
}
for score, memory_id, pinned in used
for final, memory_id, pinned in used
],
"considered": len(catalogue),
# M6: how many candidates were set aside as repeating one already
@@ -665,6 +901,15 @@ async def retrieve_memories(
{"id": memory_id, "duplicate_of": kept_id}
for memory_id, kept_id in suppressed
],
# v1.1 WP-B.2: what was searched for. Recorded per turn, like the rest.
"query": {
"input": query["input"],
"context": query["context"],
"input_terms": query["input_terms"],
"input_weight": INPUT_WEIGHT if input_vec is not None and context_vec is not None
else (1.0 if input_vec is not None else 0.0),
"lexical_weight": LEXICAL_WEIGHT,
},
"error": None,
}
@@ -822,6 +1067,61 @@ async def _guarded(db: Session, adventure_id: int, kind: str, coro) -> None:
db.commit()
def _excerpt_encoding():
return _token_encoding()
def excerpt_split(budget: int = MEMORY_EXCERPT_TOKENS) -> tuple[int, int]:
"""`(head_tokens, tail_tokens)` for a block longer than `budget`.
The marker and the blank lines around it are paid for first; what is left is
halved, and an odd token goes to the tail, the most recent part. So the two
parts plus the marker come to exactly `budget`.
"""
room = max(0, budget - count_tokens(f"\n\n{EXCERPT_OMISSION_MARKER}\n\n"))
head = room // 2
return head, room - head
def memory_excerpt(raw: str, budget: int = MEMORY_EXCERPT_TOKENS) -> str:
"""What the summariser is shown of one block.
v1.1 WP-B.2. A block that fits in `budget` tokens is sent whole, exactly as
before. A longer block used to be cut to its last `budget` tokens, and B.1
showed that a fact near its start then never reached the summariser at all.
It is now sent as its opening and its end, in order, with
`EXCERPT_OMISSION_MARKER` between them, still inside `budget`.
Rejoining two token runs can tokenise a little differently at the seams, so
the result is measured, and the head gives up tokens until it fits. A fact in
the middle of a very long block is still left out: this bounds the input, it
does not summarise everything.
"""
enc = _excerpt_encoding()
tokens = enc.encode(raw)
if len(tokens) <= budget:
return raw
head_n, tail_n = excerpt_split(budget)
while True:
excerpt = (f"{enc.decode(tokens[:head_n]).rstrip()}\n\n{EXCERPT_OMISSION_MARKER}\n\n"
f"{enc.decode(tokens[-tail_n:]).lstrip()}" if tail_n else
enc.decode(tokens[:head_n]))
over = count_tokens(excerpt) - budget
if over <= 0 or head_n == 0:
return excerpt
head_n = max(0, head_n - over)
def memory_user_prompt(brief: str, excerpt: str) -> str:
"""The user message of a memory call: the cast brief, then the excerpt.
Kept apart from `summarize_block` so an evaluation can send a model exactly
what the application sends (v1.1 WP-B.2, `tools/memory_fidelity.py`).
"""
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
return f"{brief}\n\n{prompt}" if brief else prompt
async def summarize_block(
adventure: models.Adventure,
provider: OpenAICompatibleProvider,
@@ -840,15 +1140,16 @@ async def summarize_block(
old text in place and moves on.
"""
raw = "\n\n".join(a.text for a in block)
excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)
excerpt = memory_excerpt(raw)
# Match the cast against the untruncated block. The excerpt is what the
# model reads, but a character named in the part that was trimmed is still
# model reads, but a character named in the part that was left out is still
# one the memory may have to name.
brief = cast_brief(adventure, raw)
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
return await provider.complete(
MEMORY_SYSTEM_PROMPT, f"{brief}\n\n{prompt}" if brief else prompt
)
text = await provider.complete(MEMORY_SYSTEM_PROMPT, memory_user_prompt(brief, excerpt))
# The marker is an instruction to the summariser, never a fact of the story.
if text and EXCERPT_OMISSION_MARKER in text:
text = " ".join(text.replace(EXCERPT_OMISSION_MARKER, " ").split())
return text
async def _create_due_memories(
@@ -1028,11 +1329,93 @@ async def _embed_pending(
return len(pending)
def eviction_order(rows, limit: int) -> list[int]:
"""The ids eviction would take, first to last, at most `limit` of them.
v1.1 WP-B.2. `rows` are the active memories of one adventure, each with
`id`, `pinned`, `source_start`, `source_end`, `last_used_at`, `created_at`
and `use_count`. Nothing here reads a vector or the database, so the same
function is what the eviction pass runs and what a diagnostic reports.
WP-B.1 showed what pure least-recently-used order does to a long campaign.
Retrieval is steered by the present scene, so a memory of an early stretch
nothing recent resembles stops being used. It then becomes the least
recently used row, and it goes first, while the bank keeps several memories
of the last few scenes that the history window still holds in full. The
rule below keeps the bank spread over the whole story instead.
**Coverage.** Memories with a source range say which stretch of the story
they describe. A memory is judged by the hole its removal would leave: the
number of depths between the end of the nearest memory before it and the
start of the nearest memory after it. The smallest hole goes first, so the
bank thins where it is densest. A memory whose start another memory shares
(a retried or re-played stretch, or a sibling line) leaves no hole, and is
the first kind to go. Pinned memories count as coverage, since they stay.
**Boundaries.** The earliest and the latest memory by position leave a hole
with no memory on one side: removing the first loses the only record of the
opening, and removing the last loses the only record of the most recent
stretch, which is also what keeps a memory written this turn from being
evicted by the pass that wrote it (the frozen bank, below). Boundaries are
not coverage candidates.
**Recency.** Among memories whose removal leaves the same hole, the least
recently used goes first (`coalesce(last_used_at, created_at)`), then the
less used, then the lower id. Ties are therefore never left to the order the
database returned rows in.
**Fallback.** When no memory is a coverage candidate — memories typed by the
player or migrated from before coordinates have no range, and a bank can be
all boundaries — the rest are taken least recently used first, exactly as
v1.0.0 did. The bank stays bounded either way. Pinned memories are never
taken; if every active memory is pinned, capacity yields to the pins.
Recomputed after each pick, because removing one memory widens the holes
of its neighbours.
"""
remaining = {row.id: row for row in rows if not row.pinned}
coverers = {row.id: row for row in rows
if row.source_start is not None and row.source_end is not None}
def recency(row):
return (row.last_used_at or row.created_at, row.use_count or 0, row.id)
order: list[int] = []
while remaining and len(order) < limit:
spans = sorted(coverers.values(), key=lambda r: (r.source_start, r.source_end, r.id))
starts: dict[int, int] = {}
for row in spans:
starts[row.source_start] = starts.get(row.source_start, 0) + 1
best = None
furthest_end = None # the largest source_end before index i
for i, row in enumerate(spans):
if row.id in remaining:
if starts[row.source_start] > 1:
cost = 0
elif i == 0 or i == len(spans) - 1:
cost = None # a boundary
else:
cost = max(0, spans[i + 1].source_start - furthest_end - 1)
if cost is not None:
key = (cost, *recency(row))
if best is None or key < best[0]:
best = (key, row.id)
furthest_end = row.source_end if furthest_end is None else max(furthest_end, row.source_end)
if best is None:
victim = min(remaining.values(), key=recency).id
else:
victim = best[1]
order.append(victim)
del remaining[victim]
coverers.pop(victim, None)
return order
def _evict_over_capacity(
adventure: models.Adventure, settings: models.Settings, db: Session
) -> None:
# The database performs both the count and the ranking, and returns neither
# the rows nor the vectors. Counting by walking `adventure.memories` fetched
# The database performs the count, and the rows read for ordering carry
# neither text nor vectors. Counting by walking `adventure.memories` fetched
# every vector in the bank on every turn, whether or not the bank was over
# capacity.
in_this_bank = (models.Memory.adventure_id == adventure.id,
@@ -1043,32 +1426,23 @@ def _evict_over_capacity(
overflow = active - max(1, settings.memory_bank_capacity)
if overflow <= 0:
return
# Evict the least recently used memory first, and use the use count only to
# break ties.
# v1.1 WP-B.2: the order is `eviction_order`, coverage first and recency
# second. It replaces least recently used alone; see that function.
#
# Ordering by use count first froze the bank. A memory written on this turn
# has never been used, so once every other memory had been retrieved at
# least once, the new memory held the lowest count in the bank. The same
# post-turn run that wrote it then evicted it, one pass after embedding it.
# Use counts only increase, so the bank never recovered. An adventure kept
# whatever memories it held when the bank first filled, and every later
# memory was summarized, marked as forgotten, and never ranked.
#
# Ordering by recency avoids that. A new memory carries the newest
# timestamp, so it is the last row to be evicted rather than the first, and
# it remains until other memories are used. Demoting the use count costs
# little, because retrieving a useful memory also makes it recent. The two
# orderings differ only for memories that were used once and have not been
# retrieved since, which are the rows a full bank should evict.
doomed = db.execute(
select(models.Memory.id)
.where(*in_this_bank, models.Memory.pinned.is_(False))
.order_by(
func.coalesce(models.Memory.last_used_at, models.Memory.created_at),
models.Memory.use_count,
)
.limit(overflow)
).scalars().all()
# What the old ordering fixed still holds. Ordering by use count first froze
# the bank: a memory written on this turn has never been used, so once every
# other memory had been retrieved at least once, the new memory held the
# lowest count in the bank, and the same post-turn run that wrote it evicted
# it. Counts only increase, so the bank never recovered. Under the coverage
# rule the newest memory is the latest boundary, so it is not a coverage
# candidate, and in the fallback it carries the newest timestamp.
rows = db.execute(
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
models.Memory.source_end, models.Memory.last_used_at,
models.Memory.created_at, models.Memory.use_count)
.where(*in_this_bank)
).all()
doomed = eviction_order(rows, overflow)
if not doomed:
return # Every active memory is pinned, so the pins override capacity.
db.execute(