Files
JesseMarkowitzandClaude Opus 5 0c1ba836ba v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 11:21:53 -04:00

1458 lines
66 KiB
Python

"""Phase 6: automatic summarization and the embedding memory bank.
This follows AI Dungeon's memory system. See
help.aidungeon.com/faq/the-memory-system.
After each turn, `run_post_turn` runs as a fire-and-forget task with its own
database session. It does three things:
- Every `MEMORY_INTERVAL` actions, starting once the adventure reaches
`MEMORY_START` actions, it summarizes each uncovered block of actions into a
short memory. A block waits until `SETTLE_SLACK` actions sit past it, so a
memory never ends on the action a retry could replace.
- Every `SUMMARY_INTERVAL` actions, it rewrites the story summary to include the
new memories. The rewrite always starts from the text the user edited and
never discards it.
- It embeds new memories through an OpenAI-compatible `/v1/embeddings` endpoint,
then evicts the bank down to its capacity. Evicted memories are marked as
forgotten and kept so that the UI can still show them.
When the app generates a turn, `retrieve_memories` embeds the player's input
and the current scene, and ranks the bank by a fixed mix of cosine similarity
and rarity-weighted word overlap with the input (v1.1 WP-B.2). The
highest-ranked memories become the Memories section of the context.
Every AI call in this module is best-effort. A failure is logged to the debug
page and retried on a later turn, because the cursors advance only after a call
succeeds.
"""
import asyncio
import logging
import math
from array import array
from collections import OrderedDict
from sqlalchemy import func, select, update
from sqlalchemy.orm import Session, defer, object_session
from . import derived, models, summaries, tree, vectors
from .context import (
count_tokens,
cursors,
history,
lineage,
match_cards,
story_actions,
truncate_to_last_tokens,
)
from .context.builder import _encoding as _token_encoding
from .database import SessionLocal
from .knowledge import embeddings as knowledge_embeddings
from .knowledge import fts
from .narrative import model as narrative_model
from .providers import OpenAICompatibleProvider, ProviderError
from .vectors import cosine # re-exported: the ranking lives here, the maths there
log = logging.getLogger(__name__)
MEMORY_INTERVAL = 6 # actions per memory
MEMORY_START = 12 # first memory once the adventure reaches this many actions
SUMMARY_INTERVAL = 15 # actions between Story Summary updates
MAX_MEMORIES_PER_RUN = 5 # cap catch-up work (e.g. imported adventures) per turn
MAX_EMBED_BATCH = 32
SUMMARY_MAX_WORDS = 250
MEMORY_EXCERPT_TOKENS = 2000 # the most of a block the summariser is shown
# v1.1 WP-B.2: what stands between the two parts of a block too long to send
# whole. It says a part is missing, so the summariser does not read the end as
# following straight on from the opening, and `summarize_block` removes it from
# anything the model repeats back.
EXCERPT_OMISSION_MARKER = "[… the middle of this stretch of story is left out here …]"
# How much story has to sit past a block before that block is summarized.
#
# This is not the pre-SP4 holdback returning. That rule made the newest action
# invisible to the summarizer everywhere, and it existed because a retry
# rewrote an action's text in place, so a memory covering the tip could end up
# describing narration the player had already replaced. Sibling attempts and
# `forget_node` settled that, and correctness does not rest on the number
# below: a delete or an undo anywhere in the story is still repaired by
# withdrawing what the coordinate produced.
#
# This is a cost rule, and it is only about the tip. Retry and take-switching
# both refuse anything but the newest action (`takes.retry_action`,
# `takes.switch_take`), and both withdraw the derived work at the coordinate
# they change. So a memory whose block ends on the tip is the one memory a
# player can still throw away, and every retry of that turn pays for it twice:
# once to write it, once to write it again. A block closes every
# MEMORY_INTERVAL actions and a normal turn writes two of them, so that is one
# turn in three.
#
# One action of slack moves the block end out of reach of both endpoints, and
# it buys nothing back in context quality: the block that just closed is still
# in the history window in full, so a memory of it tells the model what it can
# already read. Memories earn their place once the raw text has scrolled out,
# which is never the turn the block closed.
SETTLE_SLACK = 1
# ---- The cast brief (Phase 18b) ----
# How many characters the brief names, and how much of each description it
# carries. The cast is authored content rather than generated, so it is small in
# practice; these are ceilings against a scenario with a very large cast, not a
# budget anyone is expected to reach.
MAX_CAST_MEMBERS = 8
CAST_ENTRY_CHARS = 240
SETTING_TOKENS = 300 # of `adventure.memory`, the plot essentials
# A ceiling on one memory, in words. Measured, not guessed: with only
# "1-2 plain sentences" to go on, a real model wrote 34 words for one block and
# 105 for the next, and a 105-word memory is a paragraph. Five of those are
# injected per turn at the default `memory_top_k`, so the bank's cost is set
# here. A stated number also holds the length steady between memories, which is
# the same consistency the framing rule buys for the wording.
MEMORY_MAX_WORDS = 50
# The framing rule is the larger half of this prompt, and it is worth the
# tokens. Without it the model chooses a person per call, so one bank ends up
# holding "You entered the crypt", "The player entered the crypt" and "He
# entered the crypt" for the same kind of event.
#
# The reason is stated, not just the rule. A memory is retrieved in isolation
# months of story later, and a model told why bare pronouns fail complies far
# more consistently than one handed a bare instruction.
#
# The protagonist's name lives in the user message, in the Cast, rather than
# here. That keeps this prompt constant across every adventure and every call.
MEMORY_SYSTEM_PROMPT = (
"You compress interactive-fiction story excerpts into memories. Respond with "
"1-2 plain sentences in past tense stating the concrete facts and events "
f"(names, places, items, promises, injuries), in at most {MEMORY_MAX_WORDS} "
"words. Keep the details a later scene could turn on — a name, a promise, an "
"injury, where something is — and drop the ones it could not.\n\n"
"Write in the third person. The excerpt is written in the second person: "
'"you" is the protagonist, who is named in the Cast. Refer to the '
'protagonist by that name, never as "you". If the Cast gives no name for '
'them, call them "the player". Name the other characters too rather than '
'writing "he", "she" or "they" on their own — this memory will be read on '
"its own, much later, with nothing around it to say who a pronoun meant.\n\n"
"No preamble, no commentary."
)
SUMMARY_SYSTEM_PROMPT = (
"You maintain the running summary of an interactive-fiction story. Respond "
"with only the updated summary: a single plain-prose overview of the plot "
f"so far, at most {SUMMARY_MAX_WORDS} words. Preserve important established "
"facts; compress older events harder than recent ones.\n\n"
"Write in the third person, and name the characters. Refer to the "
'protagonist by the name given in the Cast, or as "the player" if the Cast '
'gives no name. Never address them as "you".'
)
# Adventures with a post-turn task currently running (single-process app).
_running: set[int] = set()
# Strong references to tasks that are still running. The event loop holds only
# weak references, so without this set a fire-and-forget task can be garbage
# collected before it finishes.
_tasks: set[asyncio.Task] = set()
# Both factories read the endpoint and the model names straight off `Settings`.
# They used to also read an API key, which is gone: Ollama does not use one and
# M2 removed cloud providers. `summary_model` and `embedding_model` fall back to
# the narrator model when the user has not named a separate one.
# M6: the words that mark a memory as an interpretation rather than a record.
#
# The application owns this classification, not the model
# (`CONTEXT-AND-MEMORY.md` §14, §15). The extractor writes prose; this decides
# what weight the narrator is told to give it. The list is deliberately short
# and readable: a memory that hedges is a reading of the story, not a fact the
# story established, and the narrator must not be able to promote it to canon.
#
# Being wrong in the cautious direction is cheap — a hedged record labelled
# heuristic is still retrieved and still useful. Being wrong the other way is
# what turns a guess into canon, which is the failure §14 exists to prevent.
HEURISTIC_MARKERS = (
"seemed", "seems", "appeared to", "appears to", "apparently", "perhaps",
"maybe", "might have", "may have", "possibly", "presumably", "suggested that",
"suggests that", "implied", "implies", "as if", "likely", "probably",
"seemingly", "hinted", "hints that", "suspects", "suspected", "believes",
"believed", "wondered whether", "wonders whether",
)
ACCEPTED_STORY = "accepted_story"
HEURISTIC = "heuristic"
# M6 corrective (review finding M6-F2). How close two memories have to be before
# the second one is treated as saying nothing new.
#
# The value is measured, not guessed. Against the configured local embedding
# model, on a fixture of near-identical "the party walks the muddy road"
# memories and a set of genuinely distinct ones:
#
# redundant pairs cosine 0.938 - 0.996
# distinct pairs cosine 0.349 - 0.906
#
# 0.93 sits in that gap. The same measurement ruled out the more obvious
# lexical test: word overlap fires hardest on exactly the pair that must NOT be
# merged — "Mara promised to return before dawn" against "Aldric promised to
# return before dawn" shares 71% of its words while meaning something else —
# and is weakest (27%) on filler that plainly repeats itself. Wording is a poor
# proxy for sameness of fact; the embedding is a better one.
#
# The threshold is model-dependent by nature. A different embedding model may
# need a different number, which is why the measurement is written down here
# rather than the value alone.
REDUNDANT_SIMILARITY = 0.93
def _drop_redundant(candidates, vectors, authority_of, limit):
"""Fills `limit` slots, skipping memories that repeat one already chosen.
Greedy over the ranked list, so the highest-scoring statement of a fact is
the one kept and its provenance is the provenance that survives. Two rules
keep this from losing information:
* **Authority is never crossed.** An inference and a record are different
kinds of claim even when they read alike, so a `heuristic` memory can
never suppress an `accepted_story` one or the reverse.
* **The bar is high.** Missing a duplicate costs some budget; dropping a
distinct fact costs the narrator something it needed. The threshold is
set where the measurement says distinct facts stop appearing.
Returns `(kept, suppressed)`, the second for the inspector — a reader
should be able to see that memories were considered and set aside rather
than never retrieved.
"""
kept: list = []
suppressed: list = []
for row in candidates:
if len(kept) >= limit:
break
_score, memory_id, _pinned = row
vector = vectors.get(memory_id)
duplicate_of = None
if vector is not None:
for _kept_score, kept_id, _ in kept:
if authority_of.get(kept_id) != authority_of.get(memory_id):
continue
other = vectors.get(kept_id)
if other is not None and cosine(vector, other) >= REDUNDANT_SIMILARITY:
duplicate_of = kept_id
break
if duplicate_of is None:
kept.append(row)
else:
suppressed.append((memory_id, duplicate_of))
return kept, suppressed
def classify_authority(text: str) -> str:
"""Returns `accepted_story` or `heuristic` for one memory's text.
Hedged language is the signal. "Aldric promised Mara he would return before
dawn" is something the story established; "Mara seemed uneasy when Captain
Vale was mentioned" is an inference about it, and the prompt has to say so.
"""
lowered = text.lower()
return HEURISTIC if any(m in lowered for m in HEURISTIC_MARKERS) else ACCEPTED_STORY
def summary_provider(settings: models.Settings) -> OpenAICompatibleProvider:
return OpenAICompatibleProvider(
settings.endpoint_url,
settings.summary_model or settings.model,
settings.api_mode,
settings.model_timeout_seconds,
)
def embedding_provider(settings: models.Settings) -> OpenAICompatibleProvider:
return OpenAICompatibleProvider(
settings.endpoint_url, settings.embedding_model
)
def set_vector(memory: models.Memory, vector: list[float] | None) -> None:
"""Store (or clear) a memory's embedding.
This function updates both columns that describe the vector. The ranking
reads `embedding_blob`, and everything else reads the `embedded` flag.
Routing every write through one function keeps the two in step, and it makes
this the only place a stored vector changes. That is why the cache below can
be invalidated here and nowhere else.
One caller cannot use this function: the bulk clear in
`routers/settings.py` that runs when the embedding model changes. It sets
the same two columns directly. See the comment there.
"""
memory.embedding_blob = None if vector is None else vectors.pack(vector)
memory.embedded = vector is not None
for cache in (_vector_cache, _terms_cache):
cached = cache.get(memory.adventure_id)
if cached is not None:
cached.pop(memory.id, None)
# ---------- The vector cache ----------
# Maps an adventure id to a dict of memory id to vector, with the
# most recently used adventure last.
#
# Turns for one adventure arrive one after another, and the bank changes little
# between them. Reading every vector on each turn fetches the same 600 KB
# repeatedly. This cache stores vectors as `array("f")`, which uses 4 bytes per
# component and matches the 6 KB the column holds. A list of Python floats would
# use eight times as much.
#
# Two rules keep the cache correct. Any code that changes a vector calls
# `set_vector`, which removes that entry. Any code that removes a memory from
# play removes it from the catalogue query below, and the next read discards
# entries that the catalogue no longer lists. Eviction, deletion, pruning, and
# an edit that clears the vector all work this way. No code path has to remember
# to invalidate the cache, which is the error this design avoids.
#
# The cache lives in the process, so it assumes one worker, which is what the
# deploy runs. With two workers, each keeps its own copy. Both stay correct
# about eviction and deletion, but a vector rewritten by one worker can remain
# stale in the other until that memory leaves the catalogue.
_vector_cache: OrderedDict[int, dict[int, array]] = OrderedDict()
VECTOR_CACHE_ADVENTURES = 8 # ~600 KB each at a 100-memory bank
# v1.1 WP-B.2: each memory's lexical terms, held the same way and by the same
# two rules as its vector. `set_vector` is also where a memory's text changes
# (an edit clears the vector to re-embed it), so dropping the entry there covers
# a rewritten text as well as a rewritten vector. Text is read only for memories
# not already held, and only on a turn whose input has words to match.
_terms_cache: OrderedDict[int, dict[int, frozenset[str]]] = OrderedDict()
def _terms_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, frozenset[str]]:
"""The lexical terms for `ids`, reading text only for the ones not already held."""
cached = _terms_cache.get(adventure_id)
if cached is None:
cached = _terms_cache[adventure_id] = {}
_terms_cache.move_to_end(adventure_id)
while len(_terms_cache) > VECTOR_CACHE_ADVENTURES:
_terms_cache.popitem(last=False)
wanted = set(ids)
for gone in set(cached) - wanted:
del cached[gone]
missing = [memory_id for memory_id in ids if memory_id not in cached]
if missing:
rows = db.execute(
select(models.Memory.id, models.Memory.text)
.where(models.Memory.id.in_(missing))
).all()
for memory_id, text in rows:
cached[memory_id] = lexical_terms(text or "")
return cached
def forget_cached_vectors(adventure_id: int) -> None:
"""Drops an adventure's cached vectors.
Call this only when the adventure itself is deleted. Every other case
corrects itself, as described in the comment above.
"""
_vector_cache.pop(adventure_id, None)
_terms_cache.pop(adventure_id, None)
def _vectors_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, array]:
"""The vectors for `ids`, reading only the ones not already held."""
cached = _vector_cache.get(adventure_id)
if cached is None:
cached = _vector_cache[adventure_id] = {}
_vector_cache.move_to_end(adventure_id)
while len(_vector_cache) > VECTOR_CACHE_ADVENTURES:
_vector_cache.popitem(last=False)
wanted = set(ids)
for gone in set(cached) - wanted:
del cached[gone]
missing = [memory_id for memory_id in ids if memory_id not in cached]
if missing:
rows = db.execute(
select(models.Memory.id, models.Memory.embedding_blob)
.where(models.Memory.id.in_(missing))
).all()
for memory_id, blob in rows:
if blob:
cached[memory_id] = vectors.unpack(blob)
return cached
def forget_node(db: Session, adventure: models.Adventure, action: models.Action) -> int:
"""Withdraw what a node produced, because the node is being removed.
Call this before deleting `action`, which undo and the delete-action
endpoint both do. A memory attaches to the node its block ends on, so
finding the memories that describe a node is a lookup on `(branch_id,
depth)`. The earlier `prune_dangling_memories` instead scanned for rows
whose covered range no longer existed, so it could only detect the problem
after it occurred.
Deleting the memory is half the work. The stretch of story it covered still
sits behind the cursors. Without a rewind, those actions count as
summarized while nothing describes them, and nothing reports the problem for
the rest of the adventure. `source_start` records where that stretch began,
so the anchor moves to the node before it. That depth is valid whether or
not a node still occupies it.
The opening node is the one exception, because migration 62 placed the whole
pre-coordinate bank on it. See the comment on `lineage.ROOT_DEPTH`.
Returns the number of memories withdrawn.
"""
if action.branch_id is None or action.depth is None:
return 0 # A pre-tree row. No path contains it, so nothing refers to it.
doomed = (
db.query(models.Memory)
.filter(
models.Memory.adventure_id == adventure.id,
models.Memory.branch_id == action.branch_id,
models.Memory.depth == action.depth,
)
.all()
)
if action.depth == lineage.ROOT_DEPTH:
# Special case for the opening node, and only for memories that
# describe no stretch of story. Migration 62 placed every memory written
# before memories had coordinates at depth 0. That choice preserved
# every memory, but it also placed them all on one node, so withdrawing
# that node would delete a player's entire bank in one action.
#
# A memory with no `source_start` was either typed by the player or
# migrated. It describes no actions, so no deletion can invalidate it,
# and it stays. A memory that genuinely summarizes a block ending here is
# still withdrawn, because the text it describes is being deleted.
doomed = [m for m in doomed if m.source_start is not None]
if not doomed:
return 0
starts = [m.source_start for m in doomed if m.source_start is not None]
for memory in doomed:
db.delete(memory)
if starts:
cursors.rewind_all(adventure, action.branch_id, min(starts) - 1)
return len(doomed)
def source_block(db: Session, memory: models.Memory) -> list[models.Action]:
"""The actions a memory was written from, oldest first.
The inverse of what `_create_due_memories` recorded. `source_start` and
`source_end` are depths, and `branch_id` says which path they are depths
on — that branch's own lineage, not the adventure's current path. A memory
written before a fork must still read back from the branch it was written
on, whichever branch the adventure has since moved to.
Returns `[]` for a memory that describes no stretch of story. Those are
hand-written, or migrated from before memories had coordinates (see
`lineage.ROOT_DEPTH`), and there is no block to read.
The range may come back shorter than `MEMORY_INTERVAL`. An action inside it
can have been deleted since, and a memory whose block is now partial still
describes the actions that remain.
"""
if memory.source_start is None or memory.source_end is None:
return []
if memory.branch_id is None:
return [] # A pre-tree row: no path contains it.
branch = db.get(models.Branch, memory.branch_id)
if branch is None:
return []
path = lineage.Path(lineage.entries_of(branch))
rows = (
db.query(models.Action)
.filter(
models.Action.adventure_id == memory.adventure_id,
# This clause excludes the sibling attempts at a retried turn, so
# the block holds the one text the story used.
path.clause(models.Action),
models.Action.depth >= memory.source_start,
models.Action.depth <= memory.source_end,
)
# `id` breaks a tie on `depth`, as everywhere else that orders actions.
.order_by(models.Action.depth, models.Action.id)
# Reasoning traces are never part of an excerpt and can outweigh the
# narration on a reasoning model.
.options(defer(models.Action.reasoning))
.all()
)
return [a for a in rows if history.is_story_text(a.text)]
# ---------- The cast brief ----------
def _cast_line(name: str, entry: str, *, protagonist: bool = False) -> str:
"""One roster line: who they are, and nothing about where they stand now."""
who = f"{name} — the protagonist" if protagonist else f"{name} —"
entry = " ".join(entry.split()) # collapse newlines: this is a one-line roster
if len(entry) > CAST_ENTRY_CHARS:
entry = entry[:CAST_ENTRY_CHARS].rsplit(" ", 1)[0] + "…"
if protagonist:
return f"- {who}." if not entry else f"- {who}. {entry}"
return f"- {name} — {entry}" if entry else f"- {name}"
def cast_brief(adventure: models.Adventure, text: str) -> str:
"""Returns who appears in `text`, and what the story is about.
This is the context the summarizer never had. It was handed six actions of
second-person prose and nothing else, so the only honest memory it could
write for `You push the door open. She grabs your arm.` was "You entered a
room and she stopped you." — which, retrieved forty turns later, names
nobody.
Three rules hold this together.
**Fixed descriptions only, never live values.** It is tempting to add
`Gwen: trust 40 (wary)`. That would make the same event summarized at two
different times come out framed differently, which is the fault this whole
change exists to remove.
**The cast comes from the story cards, not from `stat_schema`.** Every NPC a
scenario defines is already turned into a story card at adventure creation
(`scenario_text.scenario_card_specs`), deduplicated against the hand-written
ones by name. Reading the cards therefore covers the schema NPCs, the
author's own cards, and an adventure with no RPG layer at all, through one
path instead of three.
**Keyword matching alone is not enough here, which is why the roster is
topped up.** The turn prompt includes a card only when its trigger words
appear, and that is right for lore: a card nobody mentioned is not relevant
to the next sentence. It is wrong for this brief. The block that most needs
a cast is exactly the one written in bare pronouns — "she grabs your arm"
matches no keyword, and the summarizer is then left guessing at precisely
the moment it was given this brief to stop guessing. So matched cards come
first, and any remaining slots are filled with the other **character**
cards. Places and items are not topped up: an unmentioned tavern is not
who "she" was.
Walking `adventure.story_cards` is a relationship load, which this module
otherwise avoids. It is affordable here for two reasons the memory bank's
own reads were not: a card is five short text columns with no vector, and
this runs once per `MEMORY_INTERVAL` actions in the background task rather
than on every turn. `build_context` already walks the same collection.
"""
lines: list[str] = []
name = adventure.persona_name.strip()
pronouns = adventure.persona_pronouns.strip()
if name or adventure.persona_desc.strip():
who = f"{name} ({pronouns})" if name and pronouns else (name or "The player")
lines.append(_cast_line(who, adventure.persona_desc.strip(), protagonist=True))
seen = {name.lower()} if name else set()
def add(card_name: str, entry: str) -> bool:
"""Adds one roster line. Returns False once the roster is full."""
card_name = (card_name or "").strip()
if card_name and card_name.lower() not in seen:
seen.add(card_name.lower())
lines.append(_cast_line(card_name, (entry or "").strip()))
return len(lines) < MAX_CAST_MEMBERS
room = True
for card in match_cards(adventure.story_cards, text):
room = add(card["name"], card["entry"])
if not room:
break
if room:
for card in adventure.story_cards:
if (card.type or "").strip().lower() != "character":
continue
if not add(card.name, card.entry):
break
parts = []
if lines:
parts.append("Cast:\n" + "\n".join(lines))
setting = adventure.memory.strip()
if setting:
parts.append("Setting:\n" + truncate_to_last_tokens(setting, SETTING_TOKENS))
return "\n\n".join(parts)
# ---------- Retrieval (runs inside the turn, before build_context) ----------
# v1.1 WP-B.2: what the retrieval query is made of, and how a memory is scored
# against it (CONTEXT-AND-MEMORY §18, §20).
#
# WP-B.1 measured the v1.0.0 query, the newest four actions cut to 600 tokens,
# against a planted early fact. The player's one-line question arrived after
# three turns of narration, so the embedding mostly described the narration: a
# direct question about the fact fell from cosine 0.708 on its own to 0.241 in
# that query, and a real 100-turn campaign ranked the only memory of the fact
# 10th of 19 against a `memory_top_k` of 4.
#
# The query is now two short texts, embedded in one call:
#
# input the player's own action this turn, when there is one
# context the current scene from the authoritative state (summary, location,
# who is present), then the end of the newest narration
#
# The context is still there because a question often cannot be read without
# it ("I ask her where she hid it"), and §18 says retrieval must not rely on raw
# input alone. It is bounded so it can resolve a reference but cannot outweigh
# the question by sheer length.
#
# A memory's score is
#
# semantic_score = INPUT_WEIGHT * cos(input, memory)
# + (1 - INPUT_WEIGHT) * cos(context, memory)
# lexical_score = rarity-weighted share of the input's words the memory holds
# final_score = semantic_score + LEXICAL_WEIGHT * lexical_score
#
# With no player input (a continue, or a dry run from Insights) the semantic
# score is the context cosine alone and the lexical score is 0. Pins are
# unchanged: a pinned memory is always used and counts toward `memory_top_k`.
INPUT_TYPES = ("do", "say", "story") # player actions that carry words to search for
QUERY_INPUT_TOKENS = 200 # of the player's action; a long `story` entry is cut
QUERY_SCENE_TOKENS = 60 # of the state's scene line
QUERY_NARRATION_TOKENS = 120 # from the end of the newest narration
INPUT_WEIGHT = 0.6
# Chosen by sweep (0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5) over the deterministic
# ranking fixtures, recorded in the WP-B.2 report (§C, §D). The two-part query
# alone already ranks the planting-era memory first; 0.15 is the smallest weight
# at which the lexical term by itself also lifts it into `memory_top_k` against
# the v1.0.0 narration-filled query, and no rare-word negative control put an
# unrelated memory above it. At 0.5 an incidental shared word was enough to
# select it for an unrelated question, which is the failure a larger weight buys.
LEXICAL_WEIGHT = 0.15
# `fts.terms` drops these already; the plural fold below is the only stemming.
_MIN_FOLD_LENGTH = 5
def _fold(word: str) -> str:
"""One term, reduced so "shelves'" and "shelf" do not meet, but "teapots"
and "teapot" do. Possessives lose their `'s`, and a trailing `s` goes from a
word long enough to be a plural and not ending in `ss`. Deliberately no more
than that: a stemmer is a dependency, and a wrong fold merges two words."""
word = word.split("'", 1)[0]
if len(word) >= _MIN_FOLD_LENGTH and word.endswith("s") and not word.endswith("ss"):
word = word[:-1]
return word
def lexical_terms(text: str) -> frozenset[str]:
"""The words of `text` that lexical matching compares, folded.
The tokenizer and stop list are imported knowledge's (`knowledge.fts`), so
the two retrieval paths agree on what a word is.
"""
return frozenset(t for t in (_fold(w) for w in fts.terms(text)) if len(t) >= fts.MIN_TERM_LENGTH)
def lexical_scores(input_terms: frozenset[str], terms_of: dict[int, frozenset[str]]) -> dict[int, float]:
"""Each candidate's share of the input's rarity, in [0, 1].
A term's weight is `ln((N + 1) / (df + 1))`: N candidates, df of them holding
it. A word every candidate holds weighs exactly 0, so a protagonist's name or
a word the whole bank shares moves nothing, and a word no candidate holds
weighs the most. The share is taken over **all** the input's terms, so a
memory that happens to hold one rare word of a longer question gets that
word's part of the question, not the whole of it. The weights live only for
this call, over this candidate set: no index, no stored field.
"""
if not input_terms or not terms_of:
return {memory_id: 0.0 for memory_id in terms_of}
n = len(terms_of)
weight = {
term: math.log((n + 1) / (sum(1 for terms in terms_of.values() if term in terms) + 1))
for term in input_terms
}
total = sum(weight.values())
if total <= 0:
return {memory_id: 0.0 for memory_id in terms_of}
return {
memory_id: min(1.0, sum(w for term, w in weight.items() if term in terms) / total)
for memory_id, terms in terms_of.items()
}
def _scene_text(state) -> str:
"""The scene as the authoritative state has it: summary, location, who is present.
Names only, read straight off the document. The full entity list is left
out on purpose: a campaign with a large cast would turn every query into a
search for everyone.
"""
if not isinstance(state, dict):
return ""
scene = state.get("scene")
if not isinstance(scene, dict):
return ""
pieces: list[str] = []
summary = scene.get("summary")
if isinstance(summary, str) and summary.strip():
pieces.append(summary.strip())
location = scene.get("location")
if isinstance(location, str) and location.strip():
pieces.append(narrative_model.entity_name(state, location.strip()))
present = scene.get("present")
if isinstance(present, list):
names = [narrative_model.entity_name(state, key) for key in present[:8]
if isinstance(key, str) and key.strip()]
if names:
pieces.append(", ".join(names))
return truncate_to_last_tokens(". ".join(pieces), QUERY_SCENE_TOKENS)
def retrieval_query(adventure: models.Adventure, exclude_action_id: int | None = None) -> dict:
"""The two texts a turn's memory retrieval embeds, and the words it matches.
Returns `{"input", "context", "input_terms"}`. `input` is empty when the
newest action is not a player action with text, which is a continue turn or a
dry run. `context` is empty only for a story with no scene and no narration.
"""
recent = history.tail(adventure, 2, exclude_action_id)
newest = recent[-1] if recent else None
player_input = ""
if newest is not None and newest.type in INPUT_TYPES:
player_input = truncate_to_last_tokens(newest.text.strip(), QUERY_INPUT_TOKENS)
narration = recent[0].text if len(recent) > 1 else ""
else:
narration = newest.text if newest is not None else ""
context = "\n".join(part for part in (
_scene_text(adventure.narrative_state),
truncate_to_last_tokens(narration.strip(), QUERY_NARRATION_TOKENS),
) if part.strip())
return {
"input": player_input,
"context": context,
"input_terms": sorted(lexical_terms(player_input)),
}
def score_candidates(
ids: list[int],
held: dict,
terms_of: dict[int, frozenset[str]],
input_vec,
context_vec,
input_terms,
) -> list[tuple[float, int, float, float]]:
"""`(final_score, memory_id, semantic_score, lexical_score)`, best first.
Ties on the final score are broken by id, so the order never depends on the
order the database returned rows in.
"""
lexical = lexical_scores(frozenset(input_terms), {i: terms_of.get(i, frozenset()) for i in ids})
rows = []
for memory_id in ids:
vector = held[memory_id]
if input_vec is not None and context_vec is not None:
semantic = (INPUT_WEIGHT * cosine(input_vec, vector)
+ (1.0 - INPUT_WEIGHT) * cosine(context_vec, vector))
else:
semantic = cosine(input_vec if input_vec is not None else context_vec, vector)
lex = lexical.get(memory_id, 0.0)
rows.append((semantic + LEXICAL_WEIGHT * lex, memory_id, semantic, lex))
rows.sort(key=lambda row: (-row[0], row[1]))
return rows
def select_memories(scored, pinned_of, held, authority_of, top_k):
"""Pins first, then the best-scoring rest, skipping repeats (§22).
Returns `(used, suppressed)`. `used` is `(final_score, memory_id, pinned)`
rows, best first.
"""
rows = [(final, memory_id, pinned_of[memory_id]) for final, memory_id, _, _ in scored]
used = [row for row in rows if row[2]]
remaining = max(0, top_k - len(used))
candidates = [row for row in rows if not row[2]]
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
used += kept
used.sort(key=lambda row: (-row[0], row[1]))
return used, suppressed
async def retrieve_memories(
adventure: models.Adventure,
settings: models.Settings,
*,
exclude_action_id: int | None = None,
) -> dict | None:
"""Returns the memories to inject, or None when the bank is off.
The result is a dict of the form
`{"used": [{id, text, similarity, semantic_score, lexical_score,
final_score, pinned, authority, source}], "query": {...}, "error": str | None}`.
`similarity` is the semantic score, under the name the inspector has always
shown. It is None when the memory bank is disabled for this adventure.
This only reads. A turn counts the memories it used with `record_use`, just
before the commit that saves the turn; see that function for why the count
cannot be written here.
`exclude_action_id` removes the action being retried from the query, so that
a discarded attempt cannot influence which memories are returned.
"""
if not adventure.memory_bank_enabled:
return None
if not settings.embedding_model.strip():
return {"used": [], "error": "No embedding model configured in Settings."}
db = object_session(adventure)
if db is None:
return {"used": [], "error": None}
# Select which memories are in play, and nothing else about them. This code
# used to walk `adventure.memories`, which loaded every row of the bank,
# including its vector. That cost about 31 KB per memory and about 3 MB per
# turn, which was 96% of everything a turn read. An id and a flag come to
# about eight bytes per row.
#
# The branch clause uses the whole lineage here rather than the window the
# story is read through. Retrieval exists to recall events from far back in
# the story, such as what happened forty turns ago. The full lineage stays
# affordable because memories are sparse, at roughly one per six actions, so
# even a heavily forked story returns only tens of small rows.
catalogue = db.execute(
select(models.Memory.id, models.Memory.pinned, models.Memory.authority).where(
models.Memory.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Memory),
models.Memory.forgotten.is_(False),
models.Memory.embedded.is_(True),
)
).all()
if not catalogue:
return {"used": [], "error": None}
query = retrieval_query(adventure, exclude_action_id)
texts = [t for t in (query["input"], query["context"]) if t.strip()]
if not texts:
return {"used": [], "error": None}
try:
embedded = await embedding_provider(settings).embed(texts)
except ProviderError as exc:
return {"used": [], "error": str(exc)}
vectors_by_text = dict(zip(texts, embedded))
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
ids = [memory_id for memory_id, _, _ in catalogue]
held = _vectors_for(db, adventure.id, ids)
# Memory text is read only when there are input words to match against, and
# then only for memories not already held (see `_terms_for`).
terms_of = _terms_for(db, adventure.id, ids) if query["input_terms"] else {}
authority_of = {memory_id: authority for memory_id, _, authority in catalogue}
pinned_of = {memory_id: pinned for memory_id, pinned, _ in catalogue}
scored = score_candidates(
[memory_id for memory_id in ids if memory_id in held],
held, terms_of, input_vec, context_vec, query["input_terms"],
)
components = {memory_id: (semantic, lex) for _, memory_id, semantic, lex in scored}
used, suppressed = select_memories(
scored, pinned_of, held, authority_of, max(1, settings.memory_top_k))
if not used:
return {"used": [], "error": None}
# Fetch the text only now, and only for the `top_k` rows that were chosen.
#
# M6 adds authority and provenance to this same read rather than to a second
# one. The columns are narrow, the row set is `top_k`, and fetching them
# here is what keeps "why did the narrator remember this?" answerable
# without a query per memory (F06, and the N+1 discipline M5 restored).
used_ids = [memory_id for _, memory_id, _ in used]
detail = {
row.id: row
for row in db.execute(
select(
models.Memory.id, models.Memory.text, models.Memory.authority,
models.Memory.branch_id, models.Memory.depth,
models.Memory.source_start, models.Memory.source_end,
).where(models.Memory.id.in_(used_ids))
).all()
}
texts_of = {memory_id: row.text for memory_id, row in detail.items()}
return {
"used": [
{
"id": memory_id,
"text": texts_of.get(memory_id, ""),
"similarity": round(components[memory_id][0], 4),
# v1.1 WP-B.2: the parts of the score, so an inspector can see
# why this memory beat the ones below it.
"semantic_score": round(components[memory_id][0], 4),
"lexical_score": round(components[memory_id][1], 4),
"final_score": round(final, 4),
"pinned": pinned,
# M6: what weight this carries, and where it came from.
"authority": getattr(detail.get(memory_id), "authority", ACCEPTED_STORY),
"source": {
"branch_id": getattr(detail.get(memory_id), "branch_id", None),
"depth": getattr(detail.get(memory_id), "depth", None),
"source_start": getattr(detail.get(memory_id), "source_start", None),
"source_end": getattr(detail.get(memory_id), "source_end", None),
},
}
for final, memory_id, pinned in used
],
"considered": len(catalogue),
# M6: how many candidates were set aside as repeating one already
# chosen. Visible so that "why is that memory not here?" has an answer.
"suppressed": [
{"id": memory_id, "duplicate_of": kept_id}
for memory_id, kept_id in suppressed
],
# v1.1 WP-B.2: what was searched for. Recorded per turn, like the rest.
"query": {
"input": query["input"],
"context": query["context"],
"input_terms": query["input_terms"],
"input_weight": INPUT_WEIGHT if input_vec is not None and context_vec is not None
else (1.0 if input_vec is not None else 0.0),
"lexical_weight": LEXICAL_WEIGHT,
},
"error": None,
}
def record_use(db: Session, memory_bank: dict | None) -> None:
"""Counts the memories a turn was given, as part of that turn's commit.
Call this immediately before the commit that saves the turn, and never
before the model call. This counter used to be written during retrieval, and
the UPDATE opened a write transaction that stayed open for the whole reply,
because the turn commits only once the narration has streamed. SQLite has
one writer. Every post-turn memory, summary and status write that arrived
during the reply waited out the driver's five-second timeout and failed with
`database is locked`. Recording those failures also needs a write, so it
failed the same way, and derived status kept reporting `idle`. A 26-turn
run on a GPU host wrote two memories and no summary while every turn was
accepted.
Only real turns count, never Insights' dry runs. A turn that fails before
its commit counts nothing, because nothing was used.
Pass `synchronize_session=False` because nothing in this request reads the
counters back. Matching the UPDATE against loaded objects would require
loading those objects, which is the cost this code path exists to avoid.
"""
used_ids = [m["id"] for m in (memory_bank or {}).get("used") or []]
if not used_ids:
return
db.execute(
update(models.Memory)
.where(models.Memory.id.in_(used_ids))
.values(use_count=models.Memory.use_count + 1, last_used_at=models.utcnow())
.execution_options(synchronize_session=False)
)
# ---------- Post-turn background work ----------
def schedule_post_turn(adventure: models.Adventure) -> None:
"""Fire-and-forget summarization/embedding work after a turn is saved.
M7 adds a third reason to run: imported passages that still need vectors.
Without it a campaign that plays with story memory switched off would never
catch up an import whose embedding failed, and the only repair would be an
explicit Reindex.
"""
if not (
adventure.auto_summarize
or adventure.memory_bank_enabled
or adventure.knowledge_sources
):
return
if adventure.id in _running:
return
task = asyncio.get_running_loop().create_task(run_post_turn(adventure.id))
_tasks.add(task)
task.add_done_callback(_tasks.discard)
async def run_post_turn(adventure_id: int) -> None:
if adventure_id in _running:
return
_running.add(adventure_id)
db = SessionLocal()
try:
adventure = db.get(models.Adventure, adventure_id)
if adventure is None:
return
# Settings are per-user (Phase 8): use the adventure owner's row.
settings = (
db.query(models.Settings)
.filter(models.Settings.user_id == adventure.user_id)
.first()
)
if settings is None:
return
# This code no longer clamps the cursors. Undo can leave the story
# shorter than the mark. When the mark was a position, a value past the
# end of the list stalled the pass until the story grew back, so every
# post-turn run clamped it. That clamp introduced its own error, because
# clamping to the settled count rewound a caught-up adventure by one
# step and covered an action twice.
#
# An anchor past the tip is not an invalid value. `settled_after`
# reports that there is nothing to do, and once the story grows past the
# anchor the pass resumes where it stopped.
# M6. Each kind runs inside its own recorder, so one failing pass
# neither hides the others nor takes the turn down with it. The accepted
# narration, its state events, the authoritative document and the head
# were all committed before this task started; nothing here may undo
# them, and nothing here may fail without leaving a record
# (`BUILD-MILESTONES.md`, note from M2).
if adventure.auto_summarize:
await _guarded(db, adventure_id, derived.MEMORY,
_create_due_memories(adventure, settings, db))
await _guarded(db, adventure_id, derived.SUMMARY,
_update_story_summary(adventure, settings, db))
if adventure.memory_bank_enabled and settings.embedding_model.strip():
await _guarded(db, adventure_id, derived.EMBEDDING,
_embed_pending(adventure, settings, db))
# M7: the imported knowledge library's own vectors, caught up here.
#
# Import embeds what it can at the moment the file arrives. This is what
# happens when that failed, when the endpoint was down, when the reader
# configured an embedding model afterwards, or when a library was large
# enough that one pass did not finish it. It is not conditioned on
# `memory_bank_enabled`: the knowledge library is a separate subsystem
# and a reader who turned story memory off did not thereby ask for their
# imported Canon to stop being searchable.
#
# `embed_pending` records its own outcome, per source and per campaign,
# and never raises — so unlike the passes above it needs no guard, and
# wrapping it in one would overwrite the finer-grained record it just
# wrote with a coarser one.
if settings.embedding_model.strip():
await knowledge_embeddings.embed_pending(db, adventure, settings)
db.commit()
_evict_over_capacity(adventure, settings, db)
except BaseException as exc: # noqa: BLE001 - the task boundary
# Anything the per-kind guards did not catch: a failure in the shared
# setup above, or in eviction. M2's lesson is that the one thing this
# may not do is vanish. Re-raising would only feed an unobserved task.
try:
# Roll back first. The failure is often a flush or commit that
# failed, which leaves the session unusable until it is rolled
# back, and the record then fails with `PendingRollbackError`
# instead of being written. `_guarded` already does this.
db.rollback()
derived.failed(db, adventure_id, derived.MEMORY, exc)
db.commit()
except BaseException: # noqa: BLE001 - the recorder must not mask it
log.exception("could not record derived-work failure for %s", adventure_id)
finally:
db.close()
_running.discard(adventure_id)
async def _guarded(db: Session, adventure_id: int, kind: str, coro) -> None:
"""Runs one derived pass, recording whether it worked.
The pass keeps whatever it committed before it failed — a memory written
two blocks ago stays written — because derived work is additive and
partial progress is still progress. What must not survive is an
uncommitted, half-written unit of work, so the session is rolled back to
the last commit before the failure is recorded.
"""
try:
did_work = await coro
except BaseException as exc: # noqa: BLE001 - one kind must not stop another
db.rollback()
derived.failed(db, adventure_id, kind, exc)
db.commit()
else:
derived.succeeded(db, adventure_id, kind, did_work=bool(did_work))
db.commit()
def _excerpt_encoding():
return _token_encoding()
def excerpt_split(budget: int = MEMORY_EXCERPT_TOKENS) -> tuple[int, int]:
"""`(head_tokens, tail_tokens)` for a block longer than `budget`.
The marker and the blank lines around it are paid for first; what is left is
halved, and an odd token goes to the tail, the most recent part. So the two
parts plus the marker come to exactly `budget`.
"""
room = max(0, budget - count_tokens(f"\n\n{EXCERPT_OMISSION_MARKER}\n\n"))
head = room // 2
return head, room - head
def memory_excerpt(raw: str, budget: int = MEMORY_EXCERPT_TOKENS) -> str:
"""What the summariser is shown of one block.
v1.1 WP-B.2. A block that fits in `budget` tokens is sent whole, exactly as
before. A longer block used to be cut to its last `budget` tokens, and B.1
showed that a fact near its start then never reached the summariser at all.
It is now sent as its opening and its end, in order, with
`EXCERPT_OMISSION_MARKER` between them, still inside `budget`.
Rejoining two token runs can tokenise a little differently at the seams, so
the result is measured, and the head gives up tokens until it fits. A fact in
the middle of a very long block is still left out: this bounds the input, it
does not summarise everything.
"""
enc = _excerpt_encoding()
tokens = enc.encode(raw)
if len(tokens) <= budget:
return raw
head_n, tail_n = excerpt_split(budget)
while True:
excerpt = (f"{enc.decode(tokens[:head_n]).rstrip()}\n\n{EXCERPT_OMISSION_MARKER}\n\n"
f"{enc.decode(tokens[-tail_n:]).lstrip()}" if tail_n else
enc.decode(tokens[:head_n]))
over = count_tokens(excerpt) - budget
if over <= 0 or head_n == 0:
return excerpt
head_n = max(0, head_n - over)
def memory_user_prompt(brief: str, excerpt: str) -> str:
"""The user message of a memory call: the cast brief, then the excerpt.
Kept apart from `summarize_block` so an evaluation can send a model exactly
what the application sends (v1.1 WP-B.2, `tools/memory_fidelity.py`).
"""
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
return f"{brief}\n\n{prompt}" if brief else prompt
async def summarize_block(
adventure: models.Adventure,
provider: OpenAICompatibleProvider,
block: list[models.Action],
) -> str:
"""Writes one memory from one block of story.
Both callers come through here, which is the point of the function. The
pass below writes a memory as the story reaches it; `tools/rewrite_memories`
rewrites one an older prompt produced. Assembling the prompt in two places
would mean a rewritten memory was written by a prompt that never shipped,
and nothing would report the difference.
Raises `ProviderError`, which each caller handles its own way: the pass
below leaves the cursor alone and retries next turn, and the tool leaves the
old text in place and moves on.
"""
raw = "\n\n".join(a.text for a in block)
excerpt = memory_excerpt(raw)
# Match the cast against the untruncated block. The excerpt is what the
# model reads, but a character named in the part that was left out is still
# one the memory may have to name.
brief = cast_brief(adventure, raw)
text = await provider.complete(MEMORY_SYSTEM_PROMPT, memory_user_prompt(brief, excerpt))
# The marker is an instruction to the summariser, never a fact of the story.
if text and EXCERPT_OMISSION_MARKER in text:
text = " ".join(text.replace(EXCERPT_OMISSION_MARKER, " ").split())
return text
async def _create_due_memories(
adventure: models.Adventure, settings: models.Settings, db: Session
) -> int:
"""Writes the memories that are due. Returns how many it wrote (M6-F5)."""
provider = summary_provider(settings)
written = 0
for _ in range(MAX_MEMORIES_PER_RUN):
# Re-read the anchor on every pass. Committing a memory does not change
# the story, but this loop is the only code that moves the anchor, so
# both numbers must be current.
anchor = cursors.MEMORY.depth(db, adventure)
if history.count_after(adventure, anchor) < MEMORY_INTERVAL + SETTLE_SLACK:
return written # No settled block of story sits past the mark. The block
# itself is still MEMORY_INTERVAL actions; the slack asks
# for story past its end. See `SETTLE_SLACK`.
if history.count(adventure) < MEMORY_START:
return written # The adventure is too short to have started summarizing.
# The order of those two checks is deliberate. The usual answer is that
# no memory is due, and the first check settles that without measuring
# the length of the whole story.
block = history.after(adventure, anchor, MEMORY_INTERVAL)
if len(block) < MEMORY_INTERVAL:
return written
# A provider failure is no longer caught here. `_guarded` records it
# against this campaign, and the cursor is unchanged either way, so the
# next accepted turn retries this same block (M6).
text = await summarize_block(adventure, provider, block)
if not text:
return written
memory = models.Memory(
adventure_id=adventure.id,
text=text,
source_start=block[0].depth,
source_end=block[-1].depth,
authority=classify_authority(text),
)
# Attach the memory to the node it summarizes, so that a fork inherits
# the memories of the path it forked from and no others. Then move the
# mark to that same node. Both values record how far this pass has
# reached, and taking them from one row keeps them in step even when the
# depths have gaps.
tree.attach_memory(memory, block[-1])
db.add(memory)
cursors.MEMORY.anchor_at(adventure, block[-1])
db.commit()
written += 1
return written
async def _update_story_summary(
adventure: models.Adventure, settings: models.Settings, db: Session
) -> bool:
"""Rolls the summary forward when enough new story has settled.
Returns whether it wrote one (M6-F5)."""
anchor = cursors.SUMMARY.depth(db, adventure)
uncovered = history.count_after(adventure, anchor)
if uncovered < SUMMARY_INTERVAL:
return False
# Where the summary stands once this run succeeds. Read this before the AI
# call rather than after it. The mark records the end of the story as this
# pass saw it, and a turn that arrives during the call must not be counted
# as read.
caught_up = history.newest(adventure)
if caught_up is None:
return False
# Include the memories for the stretch that the summary has not read, which
# means every memory attached to a node past the anchor. The marks and the
# memories are now depths on one path, so no coordinate conversion remains.
# If memory creation has fallen behind, for example because the last attempt
# failed, this falls back to the raw story text.
new_events = db.execute(
select(models.Memory.text)
.where(
models.Memory.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Memory),
models.Memory.depth > anchor,
)
.order_by(models.Memory.depth)
).scalars().all()
if new_events:
events_text = "\n".join(f"- {t}" for t in new_events)
else:
block = history.after(adventure, anchor, uncovered)
events_text = truncate_to_last_tokens("\n\n".join(a.text for a in block), 2000)
# M6 corrective (review finding M6-F1). The previous summary this one
# builds on has to be a summary that is *valid where the story now stands*,
# not merely the last one written.
#
# Seeding from `adventure.story_summary` — a campaign-global column with no
# lineage — is what broke E03. After a divergence that column still held the
# abandoned line's prose, so the summariser was handed it and asked to
# update it. The row it produced was correctly anchored to the new branch
# and was therefore *reported* as lineage-safe, while its sentences
# described a story the reader had left. The row was anchored; the content
# was not.
#
# `summaries.current` answers the same question the context builder asks —
# which summary is eligible at the head — so the input and the output are
# now scoped by one rule. Where no eligible summary exists, the new line
# starts from nothing, which is the truthful starting point for a story
# that has not been summarised yet.
eligible = summaries.current(db, adventure)
current = eligible.text.strip() if eligible is not None else ""
# The summary is built from the memories, so it inherits their framing for
# free once they are named and third-person. It still gets the brief of its
# own, because the fallback above hands it raw second-person story text
# whenever memory creation has fallen behind.
brief = cast_brief(adventure, f"{current}\n\n{events_text}")
user_prompt = (
f"Current story summary:\n{current or '(none yet)'}\n\n"
f"New events since the last update:\n{events_text}\n\n"
"Updated summary:"
)
if brief:
user_prompt = f"{brief}\n\n{user_prompt}"
text = await summary_provider(settings).complete(
SUMMARY_SYSTEM_PROMPT, user_prompt, max_tokens=600
)
if not text:
return False
# M6: anchored to the story it summarizes rather than written into a single
# column. `caught_up` is the last node it covers, so the row is eligible on
# exactly the lineages that contain that node, and an Undo or a divergence
# makes it ineligible without deleting it (E03, `summaries` module).
summaries.record(
db, adventure, text,
node=caught_up,
source_start=anchor + 1 if anchor is not None else None,
trigger="interval",
model_name=(settings.summary_model or settings.model or ""),
)
cursors.SUMMARY.anchor_at(adventure, caught_up)
db.commit()
return True
async def _embed_pending(
adventure: models.Adventure, settings: models.Settings, db: Session
) -> int:
"""Embeds memories that have no vector. Returns how many (M6-F5)."""
# Use a query rather than walking `adventure.memories`. That walk ran on
# every turn and loaded the whole bank's vectors in order to find the few
# rows with none.
#
# Neither this query nor the eviction below applies a branch clause, and
# that is deliberate. Whether a row is embedded is a fact about the row, not
# about the path being played. Skipping a sibling branch's memories would
# only postpone the work until someone switched branches and needed them
# ranked. Capacity works the same way. The bank belongs to the adventure,
# and the memories of a branch nobody is reading are the right ones to evict
# first.
pending = (
db.query(models.Memory)
.filter(
models.Memory.adventure_id == adventure.id,
models.Memory.embedded.is_(False),
models.Memory.forgotten.is_(False),
)
.order_by(models.Memory.id)
.limit(MAX_EMBED_BATCH)
.all()
)
if not pending:
return 0
try:
new = await embedding_provider(settings).embed([m.text for m in pending])
except ProviderError:
return 0
for memory, vector in zip(pending, new):
set_vector(memory, vector)
db.commit()
return len(pending)
def eviction_order(rows, limit: int) -> list[int]:
"""The ids eviction would take, first to last, at most `limit` of them.
v1.1 WP-B.2. `rows` are the active memories of one adventure, each with
`id`, `pinned`, `source_start`, `source_end`, `last_used_at`, `created_at`
and `use_count`. Nothing here reads a vector or the database, so the same
function is what the eviction pass runs and what a diagnostic reports.
WP-B.1 showed what pure least-recently-used order does to a long campaign.
Retrieval is steered by the present scene, so a memory of an early stretch
nothing recent resembles stops being used. It then becomes the least
recently used row, and it goes first, while the bank keeps several memories
of the last few scenes that the history window still holds in full. The
rule below keeps the bank spread over the whole story instead.
**Coverage.** Memories with a source range say which stretch of the story
they describe. A memory is judged by the hole its removal would leave: the
number of depths between the end of the nearest memory before it and the
start of the nearest memory after it. The smallest hole goes first, so the
bank thins where it is densest. A memory whose start another memory shares
(a retried or re-played stretch, or a sibling line) leaves no hole, and is
the first kind to go. Pinned memories count as coverage, since they stay.
**Boundaries.** The earliest and the latest memory by position leave a hole
with no memory on one side: removing the first loses the only record of the
opening, and removing the last loses the only record of the most recent
stretch, which is also what keeps a memory written this turn from being
evicted by the pass that wrote it (the frozen bank, below). Boundaries are
not coverage candidates.
**Recency.** Among memories whose removal leaves the same hole, the least
recently used goes first (`coalesce(last_used_at, created_at)`), then the
less used, then the lower id. Ties are therefore never left to the order the
database returned rows in.
**Fallback.** When no memory is a coverage candidate — memories typed by the
player or migrated from before coordinates have no range, and a bank can be
all boundaries — the rest are taken least recently used first, exactly as
v1.0.0 did. The bank stays bounded either way. Pinned memories are never
taken; if every active memory is pinned, capacity yields to the pins.
Recomputed after each pick, because removing one memory widens the holes
of its neighbours.
"""
remaining = {row.id: row for row in rows if not row.pinned}
coverers = {row.id: row for row in rows
if row.source_start is not None and row.source_end is not None}
def recency(row):
return (row.last_used_at or row.created_at, row.use_count or 0, row.id)
order: list[int] = []
while remaining and len(order) < limit:
spans = sorted(coverers.values(), key=lambda r: (r.source_start, r.source_end, r.id))
starts: dict[int, int] = {}
for row in spans:
starts[row.source_start] = starts.get(row.source_start, 0) + 1
best = None
furthest_end = None # the largest source_end before index i
for i, row in enumerate(spans):
if row.id in remaining:
if starts[row.source_start] > 1:
cost = 0
elif i == 0 or i == len(spans) - 1:
cost = None # a boundary
else:
cost = max(0, spans[i + 1].source_start - furthest_end - 1)
if cost is not None:
key = (cost, *recency(row))
if best is None or key < best[0]:
best = (key, row.id)
furthest_end = row.source_end if furthest_end is None else max(furthest_end, row.source_end)
if best is None:
victim = min(remaining.values(), key=recency).id
else:
victim = best[1]
order.append(victim)
del remaining[victim]
coverers.pop(victim, None)
return order
def _evict_over_capacity(
adventure: models.Adventure, settings: models.Settings, db: Session
) -> None:
# The database performs the count, and the rows read for ordering carry
# neither text nor vectors. Counting by walking `adventure.memories` fetched
# every vector in the bank on every turn, whether or not the bank was over
# capacity.
in_this_bank = (models.Memory.adventure_id == adventure.id,
models.Memory.forgotten.is_(False))
active = db.execute(
select(func.count(models.Memory.id)).where(*in_this_bank)
).scalar() or 0
overflow = active - max(1, settings.memory_bank_capacity)
if overflow <= 0:
return
# v1.1 WP-B.2: the order is `eviction_order`, coverage first and recency
# second. It replaces least recently used alone; see that function.
#
# What the old ordering fixed still holds. Ordering by use count first froze
# the bank: a memory written on this turn has never been used, so once every
# other memory had been retrieved at least once, the new memory held the
# lowest count in the bank, and the same post-turn run that wrote it evicted
# it. Counts only increase, so the bank never recovered. Under the coverage
# rule the newest memory is the latest boundary, so it is not a coverage
# candidate, and in the fallback it carries the newest timestamp.
rows = db.execute(
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
models.Memory.source_end, models.Memory.last_used_at,
models.Memory.created_at, models.Memory.use_count)
.where(*in_this_bank)
).all()
doomed = eviction_order(rows, overflow)
if not doomed:
return # Every active memory is pinned, so the pins override capacity.
db.execute(
update(models.Memory)
.where(models.Memory.id.in_(doomed))
.values(forgotten=True)
.execution_options(synchronize_session=False)
)
db.commit()
# The bulk UPDATE bypassed the loaded objects, so code that still holds the
# collection would otherwise see the evicted memories as active.
db.expire(adventure, ["memories"])