Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
1458 lines
66 KiB
Python
1458 lines
66 KiB
Python
"""Phase 6: automatic summarization and the embedding memory bank.
|
|
|
|
This follows AI Dungeon's memory system. See
|
|
help.aidungeon.com/faq/the-memory-system.
|
|
|
|
After each turn, `run_post_turn` runs as a fire-and-forget task with its own
|
|
database session. It does three things:
|
|
|
|
- Every `MEMORY_INTERVAL` actions, starting once the adventure reaches
|
|
`MEMORY_START` actions, it summarizes each uncovered block of actions into a
|
|
short memory. A block waits until `SETTLE_SLACK` actions sit past it, so a
|
|
memory never ends on the action a retry could replace.
|
|
- Every `SUMMARY_INTERVAL` actions, it rewrites the story summary to include the
|
|
new memories. The rewrite always starts from the text the user edited and
|
|
never discards it.
|
|
- It embeds new memories through an OpenAI-compatible `/v1/embeddings` endpoint,
|
|
then evicts the bank down to its capacity. Evicted memories are marked as
|
|
forgotten and kept so that the UI can still show them.
|
|
|
|
When the app generates a turn, `retrieve_memories` embeds the player's input
|
|
and the current scene, and ranks the bank by a fixed mix of cosine similarity
|
|
and rarity-weighted word overlap with the input (v1.1 WP-B.2). The
|
|
highest-ranked memories become the Memories section of the context.
|
|
|
|
Every AI call in this module is best-effort. A failure is logged to the debug
|
|
page and retried on a later turn, because the cursors advance only after a call
|
|
succeeds.
|
|
"""
|
|
|
|
import asyncio
|
|
import logging
|
|
import math
|
|
from array import array
|
|
from collections import OrderedDict
|
|
|
|
from sqlalchemy import func, select, update
|
|
from sqlalchemy.orm import Session, defer, object_session
|
|
|
|
from . import derived, models, summaries, tree, vectors
|
|
from .context import (
|
|
count_tokens,
|
|
cursors,
|
|
history,
|
|
lineage,
|
|
match_cards,
|
|
story_actions,
|
|
truncate_to_last_tokens,
|
|
)
|
|
from .context.builder import _encoding as _token_encoding
|
|
from .database import SessionLocal
|
|
from .knowledge import embeddings as knowledge_embeddings
|
|
from .knowledge import fts
|
|
from .narrative import model as narrative_model
|
|
from .providers import OpenAICompatibleProvider, ProviderError
|
|
from .vectors import cosine # re-exported: the ranking lives here, the maths there
|
|
|
|
log = logging.getLogger(__name__)
|
|
|
|
MEMORY_INTERVAL = 6 # actions per memory
|
|
MEMORY_START = 12 # first memory once the adventure reaches this many actions
|
|
SUMMARY_INTERVAL = 15 # actions between Story Summary updates
|
|
MAX_MEMORIES_PER_RUN = 5 # cap catch-up work (e.g. imported adventures) per turn
|
|
MAX_EMBED_BATCH = 32
|
|
SUMMARY_MAX_WORDS = 250
|
|
MEMORY_EXCERPT_TOKENS = 2000 # the most of a block the summariser is shown
|
|
|
|
# v1.1 WP-B.2: what stands between the two parts of a block too long to send
|
|
# whole. It says a part is missing, so the summariser does not read the end as
|
|
# following straight on from the opening, and `summarize_block` removes it from
|
|
# anything the model repeats back.
|
|
EXCERPT_OMISSION_MARKER = "[… the middle of this stretch of story is left out here …]"
|
|
|
|
# How much story has to sit past a block before that block is summarized.
|
|
#
|
|
# This is not the pre-SP4 holdback returning. That rule made the newest action
|
|
# invisible to the summarizer everywhere, and it existed because a retry
|
|
# rewrote an action's text in place, so a memory covering the tip could end up
|
|
# describing narration the player had already replaced. Sibling attempts and
|
|
# `forget_node` settled that, and correctness does not rest on the number
|
|
# below: a delete or an undo anywhere in the story is still repaired by
|
|
# withdrawing what the coordinate produced.
|
|
#
|
|
# This is a cost rule, and it is only about the tip. Retry and take-switching
|
|
# both refuse anything but the newest action (`takes.retry_action`,
|
|
# `takes.switch_take`), and both withdraw the derived work at the coordinate
|
|
# they change. So a memory whose block ends on the tip is the one memory a
|
|
# player can still throw away, and every retry of that turn pays for it twice:
|
|
# once to write it, once to write it again. A block closes every
|
|
# MEMORY_INTERVAL actions and a normal turn writes two of them, so that is one
|
|
# turn in three.
|
|
#
|
|
# One action of slack moves the block end out of reach of both endpoints, and
|
|
# it buys nothing back in context quality: the block that just closed is still
|
|
# in the history window in full, so a memory of it tells the model what it can
|
|
# already read. Memories earn their place once the raw text has scrolled out,
|
|
# which is never the turn the block closed.
|
|
SETTLE_SLACK = 1
|
|
|
|
# ---- The cast brief (Phase 18b) ----
|
|
# How many characters the brief names, and how much of each description it
|
|
# carries. The cast is authored content rather than generated, so it is small in
|
|
# practice; these are ceilings against a scenario with a very large cast, not a
|
|
# budget anyone is expected to reach.
|
|
MAX_CAST_MEMBERS = 8
|
|
CAST_ENTRY_CHARS = 240
|
|
SETTING_TOKENS = 300 # of `adventure.memory`, the plot essentials
|
|
|
|
# A ceiling on one memory, in words. Measured, not guessed: with only
|
|
# "1-2 plain sentences" to go on, a real model wrote 34 words for one block and
|
|
# 105 for the next, and a 105-word memory is a paragraph. Five of those are
|
|
# injected per turn at the default `memory_top_k`, so the bank's cost is set
|
|
# here. A stated number also holds the length steady between memories, which is
|
|
# the same consistency the framing rule buys for the wording.
|
|
MEMORY_MAX_WORDS = 50
|
|
|
|
# The framing rule is the larger half of this prompt, and it is worth the
|
|
# tokens. Without it the model chooses a person per call, so one bank ends up
|
|
# holding "You entered the crypt", "The player entered the crypt" and "He
|
|
# entered the crypt" for the same kind of event.
|
|
#
|
|
# The reason is stated, not just the rule. A memory is retrieved in isolation
|
|
# months of story later, and a model told why bare pronouns fail complies far
|
|
# more consistently than one handed a bare instruction.
|
|
#
|
|
# The protagonist's name lives in the user message, in the Cast, rather than
|
|
# here. That keeps this prompt constant across every adventure and every call.
|
|
MEMORY_SYSTEM_PROMPT = (
|
|
"You compress interactive-fiction story excerpts into memories. Respond with "
|
|
"1-2 plain sentences in past tense stating the concrete facts and events "
|
|
f"(names, places, items, promises, injuries), in at most {MEMORY_MAX_WORDS} "
|
|
"words. Keep the details a later scene could turn on — a name, a promise, an "
|
|
"injury, where something is — and drop the ones it could not.\n\n"
|
|
"Write in the third person. The excerpt is written in the second person: "
|
|
'"you" is the protagonist, who is named in the Cast. Refer to the '
|
|
'protagonist by that name, never as "you". If the Cast gives no name for '
|
|
'them, call them "the player". Name the other characters too rather than '
|
|
'writing "he", "she" or "they" on their own — this memory will be read on '
|
|
"its own, much later, with nothing around it to say who a pronoun meant.\n\n"
|
|
"No preamble, no commentary."
|
|
)
|
|
SUMMARY_SYSTEM_PROMPT = (
|
|
"You maintain the running summary of an interactive-fiction story. Respond "
|
|
"with only the updated summary: a single plain-prose overview of the plot "
|
|
f"so far, at most {SUMMARY_MAX_WORDS} words. Preserve important established "
|
|
"facts; compress older events harder than recent ones.\n\n"
|
|
"Write in the third person, and name the characters. Refer to the "
|
|
'protagonist by the name given in the Cast, or as "the player" if the Cast '
|
|
'gives no name. Never address them as "you".'
|
|
)
|
|
|
|
# Adventures with a post-turn task currently running (single-process app).
|
|
_running: set[int] = set()
|
|
# Strong references to tasks that are still running. The event loop holds only
|
|
# weak references, so without this set a fire-and-forget task can be garbage
|
|
# collected before it finishes.
|
|
_tasks: set[asyncio.Task] = set()
|
|
|
|
|
|
# Both factories read the endpoint and the model names straight off `Settings`.
|
|
# They used to also read an API key, which is gone: Ollama does not use one and
|
|
# M2 removed cloud providers. `summary_model` and `embedding_model` fall back to
|
|
# the narrator model when the user has not named a separate one.
|
|
# M6: the words that mark a memory as an interpretation rather than a record.
|
|
#
|
|
# The application owns this classification, not the model
|
|
# (`CONTEXT-AND-MEMORY.md` §14, §15). The extractor writes prose; this decides
|
|
# what weight the narrator is told to give it. The list is deliberately short
|
|
# and readable: a memory that hedges is a reading of the story, not a fact the
|
|
# story established, and the narrator must not be able to promote it to canon.
|
|
#
|
|
# Being wrong in the cautious direction is cheap — a hedged record labelled
|
|
# heuristic is still retrieved and still useful. Being wrong the other way is
|
|
# what turns a guess into canon, which is the failure §14 exists to prevent.
|
|
HEURISTIC_MARKERS = (
|
|
"seemed", "seems", "appeared to", "appears to", "apparently", "perhaps",
|
|
"maybe", "might have", "may have", "possibly", "presumably", "suggested that",
|
|
"suggests that", "implied", "implies", "as if", "likely", "probably",
|
|
"seemingly", "hinted", "hints that", "suspects", "suspected", "believes",
|
|
"believed", "wondered whether", "wonders whether",
|
|
)
|
|
|
|
ACCEPTED_STORY = "accepted_story"
|
|
HEURISTIC = "heuristic"
|
|
|
|
|
|
# M6 corrective (review finding M6-F2). How close two memories have to be before
|
|
# the second one is treated as saying nothing new.
|
|
#
|
|
# The value is measured, not guessed. Against the configured local embedding
|
|
# model, on a fixture of near-identical "the party walks the muddy road"
|
|
# memories and a set of genuinely distinct ones:
|
|
#
|
|
# redundant pairs cosine 0.938 - 0.996
|
|
# distinct pairs cosine 0.349 - 0.906
|
|
#
|
|
# 0.93 sits in that gap. The same measurement ruled out the more obvious
|
|
# lexical test: word overlap fires hardest on exactly the pair that must NOT be
|
|
# merged — "Mara promised to return before dawn" against "Aldric promised to
|
|
# return before dawn" shares 71% of its words while meaning something else —
|
|
# and is weakest (27%) on filler that plainly repeats itself. Wording is a poor
|
|
# proxy for sameness of fact; the embedding is a better one.
|
|
#
|
|
# The threshold is model-dependent by nature. A different embedding model may
|
|
# need a different number, which is why the measurement is written down here
|
|
# rather than the value alone.
|
|
REDUNDANT_SIMILARITY = 0.93
|
|
|
|
|
|
def _drop_redundant(candidates, vectors, authority_of, limit):
|
|
"""Fills `limit` slots, skipping memories that repeat one already chosen.
|
|
|
|
Greedy over the ranked list, so the highest-scoring statement of a fact is
|
|
the one kept and its provenance is the provenance that survives. Two rules
|
|
keep this from losing information:
|
|
|
|
* **Authority is never crossed.** An inference and a record are different
|
|
kinds of claim even when they read alike, so a `heuristic` memory can
|
|
never suppress an `accepted_story` one or the reverse.
|
|
* **The bar is high.** Missing a duplicate costs some budget; dropping a
|
|
distinct fact costs the narrator something it needed. The threshold is
|
|
set where the measurement says distinct facts stop appearing.
|
|
|
|
Returns `(kept, suppressed)`, the second for the inspector — a reader
|
|
should be able to see that memories were considered and set aside rather
|
|
than never retrieved.
|
|
"""
|
|
kept: list = []
|
|
suppressed: list = []
|
|
for row in candidates:
|
|
if len(kept) >= limit:
|
|
break
|
|
_score, memory_id, _pinned = row
|
|
vector = vectors.get(memory_id)
|
|
duplicate_of = None
|
|
if vector is not None:
|
|
for _kept_score, kept_id, _ in kept:
|
|
if authority_of.get(kept_id) != authority_of.get(memory_id):
|
|
continue
|
|
other = vectors.get(kept_id)
|
|
if other is not None and cosine(vector, other) >= REDUNDANT_SIMILARITY:
|
|
duplicate_of = kept_id
|
|
break
|
|
if duplicate_of is None:
|
|
kept.append(row)
|
|
else:
|
|
suppressed.append((memory_id, duplicate_of))
|
|
return kept, suppressed
|
|
|
|
|
|
def classify_authority(text: str) -> str:
|
|
"""Returns `accepted_story` or `heuristic` for one memory's text.
|
|
|
|
Hedged language is the signal. "Aldric promised Mara he would return before
|
|
dawn" is something the story established; "Mara seemed uneasy when Captain
|
|
Vale was mentioned" is an inference about it, and the prompt has to say so.
|
|
"""
|
|
lowered = text.lower()
|
|
return HEURISTIC if any(m in lowered for m in HEURISTIC_MARKERS) else ACCEPTED_STORY
|
|
|
|
|
|
def summary_provider(settings: models.Settings) -> OpenAICompatibleProvider:
|
|
return OpenAICompatibleProvider(
|
|
settings.endpoint_url,
|
|
settings.summary_model or settings.model,
|
|
settings.api_mode,
|
|
settings.model_timeout_seconds,
|
|
)
|
|
|
|
|
|
def embedding_provider(settings: models.Settings) -> OpenAICompatibleProvider:
|
|
return OpenAICompatibleProvider(
|
|
settings.endpoint_url, settings.embedding_model
|
|
)
|
|
|
|
|
|
def set_vector(memory: models.Memory, vector: list[float] | None) -> None:
|
|
"""Store (or clear) a memory's embedding.
|
|
|
|
This function updates both columns that describe the vector. The ranking
|
|
reads `embedding_blob`, and everything else reads the `embedded` flag.
|
|
Routing every write through one function keeps the two in step, and it makes
|
|
this the only place a stored vector changes. That is why the cache below can
|
|
be invalidated here and nowhere else.
|
|
|
|
One caller cannot use this function: the bulk clear in
|
|
`routers/settings.py` that runs when the embedding model changes. It sets
|
|
the same two columns directly. See the comment there.
|
|
"""
|
|
memory.embedding_blob = None if vector is None else vectors.pack(vector)
|
|
memory.embedded = vector is not None
|
|
for cache in (_vector_cache, _terms_cache):
|
|
cached = cache.get(memory.adventure_id)
|
|
if cached is not None:
|
|
cached.pop(memory.id, None)
|
|
|
|
|
|
# ---------- The vector cache ----------
|
|
|
|
# Maps an adventure id to a dict of memory id to vector, with the
|
|
# most recently used adventure last.
|
|
#
|
|
# Turns for one adventure arrive one after another, and the bank changes little
|
|
# between them. Reading every vector on each turn fetches the same 600 KB
|
|
# repeatedly. This cache stores vectors as `array("f")`, which uses 4 bytes per
|
|
# component and matches the 6 KB the column holds. A list of Python floats would
|
|
# use eight times as much.
|
|
#
|
|
# Two rules keep the cache correct. Any code that changes a vector calls
|
|
# `set_vector`, which removes that entry. Any code that removes a memory from
|
|
# play removes it from the catalogue query below, and the next read discards
|
|
# entries that the catalogue no longer lists. Eviction, deletion, pruning, and
|
|
# an edit that clears the vector all work this way. No code path has to remember
|
|
# to invalidate the cache, which is the error this design avoids.
|
|
#
|
|
# The cache lives in the process, so it assumes one worker, which is what the
|
|
# deploy runs. With two workers, each keeps its own copy. Both stay correct
|
|
# about eviction and deletion, but a vector rewritten by one worker can remain
|
|
# stale in the other until that memory leaves the catalogue.
|
|
_vector_cache: OrderedDict[int, dict[int, array]] = OrderedDict()
|
|
VECTOR_CACHE_ADVENTURES = 8 # ~600 KB each at a 100-memory bank
|
|
|
|
|
|
# v1.1 WP-B.2: each memory's lexical terms, held the same way and by the same
|
|
# two rules as its vector. `set_vector` is also where a memory's text changes
|
|
# (an edit clears the vector to re-embed it), so dropping the entry there covers
|
|
# a rewritten text as well as a rewritten vector. Text is read only for memories
|
|
# not already held, and only on a turn whose input has words to match.
|
|
_terms_cache: OrderedDict[int, dict[int, frozenset[str]]] = OrderedDict()
|
|
|
|
|
|
def _terms_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, frozenset[str]]:
|
|
"""The lexical terms for `ids`, reading text only for the ones not already held."""
|
|
cached = _terms_cache.get(adventure_id)
|
|
if cached is None:
|
|
cached = _terms_cache[adventure_id] = {}
|
|
_terms_cache.move_to_end(adventure_id)
|
|
while len(_terms_cache) > VECTOR_CACHE_ADVENTURES:
|
|
_terms_cache.popitem(last=False)
|
|
|
|
wanted = set(ids)
|
|
for gone in set(cached) - wanted:
|
|
del cached[gone]
|
|
missing = [memory_id for memory_id in ids if memory_id not in cached]
|
|
if missing:
|
|
rows = db.execute(
|
|
select(models.Memory.id, models.Memory.text)
|
|
.where(models.Memory.id.in_(missing))
|
|
).all()
|
|
for memory_id, text in rows:
|
|
cached[memory_id] = lexical_terms(text or "")
|
|
return cached
|
|
|
|
|
|
def forget_cached_vectors(adventure_id: int) -> None:
|
|
"""Drops an adventure's cached vectors.
|
|
|
|
Call this only when the adventure itself is deleted. Every other case
|
|
corrects itself, as described in the comment above.
|
|
"""
|
|
_vector_cache.pop(adventure_id, None)
|
|
_terms_cache.pop(adventure_id, None)
|
|
|
|
|
|
def _vectors_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, array]:
|
|
"""The vectors for `ids`, reading only the ones not already held."""
|
|
cached = _vector_cache.get(adventure_id)
|
|
if cached is None:
|
|
cached = _vector_cache[adventure_id] = {}
|
|
_vector_cache.move_to_end(adventure_id)
|
|
while len(_vector_cache) > VECTOR_CACHE_ADVENTURES:
|
|
_vector_cache.popitem(last=False)
|
|
|
|
wanted = set(ids)
|
|
for gone in set(cached) - wanted:
|
|
del cached[gone]
|
|
missing = [memory_id for memory_id in ids if memory_id not in cached]
|
|
if missing:
|
|
rows = db.execute(
|
|
select(models.Memory.id, models.Memory.embedding_blob)
|
|
.where(models.Memory.id.in_(missing))
|
|
).all()
|
|
for memory_id, blob in rows:
|
|
if blob:
|
|
cached[memory_id] = vectors.unpack(blob)
|
|
return cached
|
|
|
|
|
|
def forget_node(db: Session, adventure: models.Adventure, action: models.Action) -> int:
|
|
"""Withdraw what a node produced, because the node is being removed.
|
|
|
|
Call this before deleting `action`, which undo and the delete-action
|
|
endpoint both do. A memory attaches to the node its block ends on, so
|
|
finding the memories that describe a node is a lookup on `(branch_id,
|
|
depth)`. The earlier `prune_dangling_memories` instead scanned for rows
|
|
whose covered range no longer existed, so it could only detect the problem
|
|
after it occurred.
|
|
|
|
Deleting the memory is half the work. The stretch of story it covered still
|
|
sits behind the cursors. Without a rewind, those actions count as
|
|
summarized while nothing describes them, and nothing reports the problem for
|
|
the rest of the adventure. `source_start` records where that stretch began,
|
|
so the anchor moves to the node before it. That depth is valid whether or
|
|
not a node still occupies it.
|
|
|
|
The opening node is the one exception, because migration 62 placed the whole
|
|
pre-coordinate bank on it. See the comment on `lineage.ROOT_DEPTH`.
|
|
|
|
Returns the number of memories withdrawn.
|
|
"""
|
|
if action.branch_id is None or action.depth is None:
|
|
return 0 # A pre-tree row. No path contains it, so nothing refers to it.
|
|
doomed = (
|
|
db.query(models.Memory)
|
|
.filter(
|
|
models.Memory.adventure_id == adventure.id,
|
|
models.Memory.branch_id == action.branch_id,
|
|
models.Memory.depth == action.depth,
|
|
)
|
|
.all()
|
|
)
|
|
if action.depth == lineage.ROOT_DEPTH:
|
|
# Special case for the opening node, and only for memories that
|
|
# describe no stretch of story. Migration 62 placed every memory written
|
|
# before memories had coordinates at depth 0. That choice preserved
|
|
# every memory, but it also placed them all on one node, so withdrawing
|
|
# that node would delete a player's entire bank in one action.
|
|
#
|
|
# A memory with no `source_start` was either typed by the player or
|
|
# migrated. It describes no actions, so no deletion can invalidate it,
|
|
# and it stays. A memory that genuinely summarizes a block ending here is
|
|
# still withdrawn, because the text it describes is being deleted.
|
|
doomed = [m for m in doomed if m.source_start is not None]
|
|
if not doomed:
|
|
return 0
|
|
starts = [m.source_start for m in doomed if m.source_start is not None]
|
|
for memory in doomed:
|
|
db.delete(memory)
|
|
if starts:
|
|
cursors.rewind_all(adventure, action.branch_id, min(starts) - 1)
|
|
return len(doomed)
|
|
|
|
|
|
def source_block(db: Session, memory: models.Memory) -> list[models.Action]:
|
|
"""The actions a memory was written from, oldest first.
|
|
|
|
The inverse of what `_create_due_memories` recorded. `source_start` and
|
|
`source_end` are depths, and `branch_id` says which path they are depths
|
|
on — that branch's own lineage, not the adventure's current path. A memory
|
|
written before a fork must still read back from the branch it was written
|
|
on, whichever branch the adventure has since moved to.
|
|
|
|
Returns `[]` for a memory that describes no stretch of story. Those are
|
|
hand-written, or migrated from before memories had coordinates (see
|
|
`lineage.ROOT_DEPTH`), and there is no block to read.
|
|
|
|
The range may come back shorter than `MEMORY_INTERVAL`. An action inside it
|
|
can have been deleted since, and a memory whose block is now partial still
|
|
describes the actions that remain.
|
|
"""
|
|
if memory.source_start is None or memory.source_end is None:
|
|
return []
|
|
if memory.branch_id is None:
|
|
return [] # A pre-tree row: no path contains it.
|
|
branch = db.get(models.Branch, memory.branch_id)
|
|
if branch is None:
|
|
return []
|
|
path = lineage.Path(lineage.entries_of(branch))
|
|
rows = (
|
|
db.query(models.Action)
|
|
.filter(
|
|
models.Action.adventure_id == memory.adventure_id,
|
|
# This clause excludes the sibling attempts at a retried turn, so
|
|
# the block holds the one text the story used.
|
|
path.clause(models.Action),
|
|
models.Action.depth >= memory.source_start,
|
|
models.Action.depth <= memory.source_end,
|
|
)
|
|
# `id` breaks a tie on `depth`, as everywhere else that orders actions.
|
|
.order_by(models.Action.depth, models.Action.id)
|
|
# Reasoning traces are never part of an excerpt and can outweigh the
|
|
# narration on a reasoning model.
|
|
.options(defer(models.Action.reasoning))
|
|
.all()
|
|
)
|
|
return [a for a in rows if history.is_story_text(a.text)]
|
|
|
|
|
|
# ---------- The cast brief ----------
|
|
|
|
def _cast_line(name: str, entry: str, *, protagonist: bool = False) -> str:
|
|
"""One roster line: who they are, and nothing about where they stand now."""
|
|
who = f"{name} — the protagonist" if protagonist else f"{name} —"
|
|
entry = " ".join(entry.split()) # collapse newlines: this is a one-line roster
|
|
if len(entry) > CAST_ENTRY_CHARS:
|
|
entry = entry[:CAST_ENTRY_CHARS].rsplit(" ", 1)[0] + "…"
|
|
if protagonist:
|
|
return f"- {who}." if not entry else f"- {who}. {entry}"
|
|
return f"- {name} — {entry}" if entry else f"- {name}"
|
|
|
|
|
|
def cast_brief(adventure: models.Adventure, text: str) -> str:
|
|
"""Returns who appears in `text`, and what the story is about.
|
|
|
|
This is the context the summarizer never had. It was handed six actions of
|
|
second-person prose and nothing else, so the only honest memory it could
|
|
write for `You push the door open. She grabs your arm.` was "You entered a
|
|
room and she stopped you." — which, retrieved forty turns later, names
|
|
nobody.
|
|
|
|
Three rules hold this together.
|
|
|
|
**Fixed descriptions only, never live values.** It is tempting to add
|
|
`Gwen: trust 40 (wary)`. That would make the same event summarized at two
|
|
different times come out framed differently, which is the fault this whole
|
|
change exists to remove.
|
|
|
|
**The cast comes from the story cards, not from `stat_schema`.** Every NPC a
|
|
scenario defines is already turned into a story card at adventure creation
|
|
(`scenario_text.scenario_card_specs`), deduplicated against the hand-written
|
|
ones by name. Reading the cards therefore covers the schema NPCs, the
|
|
author's own cards, and an adventure with no RPG layer at all, through one
|
|
path instead of three.
|
|
|
|
**Keyword matching alone is not enough here, which is why the roster is
|
|
topped up.** The turn prompt includes a card only when its trigger words
|
|
appear, and that is right for lore: a card nobody mentioned is not relevant
|
|
to the next sentence. It is wrong for this brief. The block that most needs
|
|
a cast is exactly the one written in bare pronouns — "she grabs your arm"
|
|
matches no keyword, and the summarizer is then left guessing at precisely
|
|
the moment it was given this brief to stop guessing. So matched cards come
|
|
first, and any remaining slots are filled with the other **character**
|
|
cards. Places and items are not topped up: an unmentioned tavern is not
|
|
who "she" was.
|
|
|
|
Walking `adventure.story_cards` is a relationship load, which this module
|
|
otherwise avoids. It is affordable here for two reasons the memory bank's
|
|
own reads were not: a card is five short text columns with no vector, and
|
|
this runs once per `MEMORY_INTERVAL` actions in the background task rather
|
|
than on every turn. `build_context` already walks the same collection.
|
|
"""
|
|
lines: list[str] = []
|
|
name = adventure.persona_name.strip()
|
|
pronouns = adventure.persona_pronouns.strip()
|
|
if name or adventure.persona_desc.strip():
|
|
who = f"{name} ({pronouns})" if name and pronouns else (name or "The player")
|
|
lines.append(_cast_line(who, adventure.persona_desc.strip(), protagonist=True))
|
|
|
|
seen = {name.lower()} if name else set()
|
|
|
|
def add(card_name: str, entry: str) -> bool:
|
|
"""Adds one roster line. Returns False once the roster is full."""
|
|
card_name = (card_name or "").strip()
|
|
if card_name and card_name.lower() not in seen:
|
|
seen.add(card_name.lower())
|
|
lines.append(_cast_line(card_name, (entry or "").strip()))
|
|
return len(lines) < MAX_CAST_MEMBERS
|
|
|
|
room = True
|
|
for card in match_cards(adventure.story_cards, text):
|
|
room = add(card["name"], card["entry"])
|
|
if not room:
|
|
break
|
|
if room:
|
|
for card in adventure.story_cards:
|
|
if (card.type or "").strip().lower() != "character":
|
|
continue
|
|
if not add(card.name, card.entry):
|
|
break
|
|
|
|
parts = []
|
|
if lines:
|
|
parts.append("Cast:\n" + "\n".join(lines))
|
|
setting = adventure.memory.strip()
|
|
if setting:
|
|
parts.append("Setting:\n" + truncate_to_last_tokens(setting, SETTING_TOKENS))
|
|
return "\n\n".join(parts)
|
|
|
|
|
|
# ---------- Retrieval (runs inside the turn, before build_context) ----------
|
|
|
|
# v1.1 WP-B.2: what the retrieval query is made of, and how a memory is scored
|
|
# against it (CONTEXT-AND-MEMORY §18, §20).
|
|
#
|
|
# WP-B.1 measured the v1.0.0 query, the newest four actions cut to 600 tokens,
|
|
# against a planted early fact. The player's one-line question arrived after
|
|
# three turns of narration, so the embedding mostly described the narration: a
|
|
# direct question about the fact fell from cosine 0.708 on its own to 0.241 in
|
|
# that query, and a real 100-turn campaign ranked the only memory of the fact
|
|
# 10th of 19 against a `memory_top_k` of 4.
|
|
#
|
|
# The query is now two short texts, embedded in one call:
|
|
#
|
|
# input the player's own action this turn, when there is one
|
|
# context the current scene from the authoritative state (summary, location,
|
|
# who is present), then the end of the newest narration
|
|
#
|
|
# The context is still there because a question often cannot be read without
|
|
# it ("I ask her where she hid it"), and §18 says retrieval must not rely on raw
|
|
# input alone. It is bounded so it can resolve a reference but cannot outweigh
|
|
# the question by sheer length.
|
|
#
|
|
# A memory's score is
|
|
#
|
|
# semantic_score = INPUT_WEIGHT * cos(input, memory)
|
|
# + (1 - INPUT_WEIGHT) * cos(context, memory)
|
|
# lexical_score = rarity-weighted share of the input's words the memory holds
|
|
# final_score = semantic_score + LEXICAL_WEIGHT * lexical_score
|
|
#
|
|
# With no player input (a continue, or a dry run from Insights) the semantic
|
|
# score is the context cosine alone and the lexical score is 0. Pins are
|
|
# unchanged: a pinned memory is always used and counts toward `memory_top_k`.
|
|
INPUT_TYPES = ("do", "say", "story") # player actions that carry words to search for
|
|
QUERY_INPUT_TOKENS = 200 # of the player's action; a long `story` entry is cut
|
|
QUERY_SCENE_TOKENS = 60 # of the state's scene line
|
|
QUERY_NARRATION_TOKENS = 120 # from the end of the newest narration
|
|
INPUT_WEIGHT = 0.6
|
|
# Chosen by sweep (0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5) over the deterministic
|
|
# ranking fixtures, recorded in the WP-B.2 report (§C, §D). The two-part query
|
|
# alone already ranks the planting-era memory first; 0.15 is the smallest weight
|
|
# at which the lexical term by itself also lifts it into `memory_top_k` against
|
|
# the v1.0.0 narration-filled query, and no rare-word negative control put an
|
|
# unrelated memory above it. At 0.5 an incidental shared word was enough to
|
|
# select it for an unrelated question, which is the failure a larger weight buys.
|
|
LEXICAL_WEIGHT = 0.15
|
|
# `fts.terms` drops these already; the plural fold below is the only stemming.
|
|
_MIN_FOLD_LENGTH = 5
|
|
|
|
|
|
def _fold(word: str) -> str:
|
|
"""One term, reduced so "shelves'" and "shelf" do not meet, but "teapots"
|
|
and "teapot" do. Possessives lose their `'s`, and a trailing `s` goes from a
|
|
word long enough to be a plural and not ending in `ss`. Deliberately no more
|
|
than that: a stemmer is a dependency, and a wrong fold merges two words."""
|
|
word = word.split("'", 1)[0]
|
|
if len(word) >= _MIN_FOLD_LENGTH and word.endswith("s") and not word.endswith("ss"):
|
|
word = word[:-1]
|
|
return word
|
|
|
|
|
|
def lexical_terms(text: str) -> frozenset[str]:
|
|
"""The words of `text` that lexical matching compares, folded.
|
|
|
|
The tokenizer and stop list are imported knowledge's (`knowledge.fts`), so
|
|
the two retrieval paths agree on what a word is.
|
|
"""
|
|
return frozenset(t for t in (_fold(w) for w in fts.terms(text)) if len(t) >= fts.MIN_TERM_LENGTH)
|
|
|
|
|
|
def lexical_scores(input_terms: frozenset[str], terms_of: dict[int, frozenset[str]]) -> dict[int, float]:
|
|
"""Each candidate's share of the input's rarity, in [0, 1].
|
|
|
|
A term's weight is `ln((N + 1) / (df + 1))`: N candidates, df of them holding
|
|
it. A word every candidate holds weighs exactly 0, so a protagonist's name or
|
|
a word the whole bank shares moves nothing, and a word no candidate holds
|
|
weighs the most. The share is taken over **all** the input's terms, so a
|
|
memory that happens to hold one rare word of a longer question gets that
|
|
word's part of the question, not the whole of it. The weights live only for
|
|
this call, over this candidate set: no index, no stored field.
|
|
"""
|
|
if not input_terms or not terms_of:
|
|
return {memory_id: 0.0 for memory_id in terms_of}
|
|
n = len(terms_of)
|
|
weight = {
|
|
term: math.log((n + 1) / (sum(1 for terms in terms_of.values() if term in terms) + 1))
|
|
for term in input_terms
|
|
}
|
|
total = sum(weight.values())
|
|
if total <= 0:
|
|
return {memory_id: 0.0 for memory_id in terms_of}
|
|
return {
|
|
memory_id: min(1.0, sum(w for term, w in weight.items() if term in terms) / total)
|
|
for memory_id, terms in terms_of.items()
|
|
}
|
|
|
|
|
|
def _scene_text(state) -> str:
|
|
"""The scene as the authoritative state has it: summary, location, who is present.
|
|
|
|
Names only, read straight off the document. The full entity list is left
|
|
out on purpose: a campaign with a large cast would turn every query into a
|
|
search for everyone.
|
|
"""
|
|
if not isinstance(state, dict):
|
|
return ""
|
|
scene = state.get("scene")
|
|
if not isinstance(scene, dict):
|
|
return ""
|
|
pieces: list[str] = []
|
|
summary = scene.get("summary")
|
|
if isinstance(summary, str) and summary.strip():
|
|
pieces.append(summary.strip())
|
|
location = scene.get("location")
|
|
if isinstance(location, str) and location.strip():
|
|
pieces.append(narrative_model.entity_name(state, location.strip()))
|
|
present = scene.get("present")
|
|
if isinstance(present, list):
|
|
names = [narrative_model.entity_name(state, key) for key in present[:8]
|
|
if isinstance(key, str) and key.strip()]
|
|
if names:
|
|
pieces.append(", ".join(names))
|
|
return truncate_to_last_tokens(". ".join(pieces), QUERY_SCENE_TOKENS)
|
|
|
|
|
|
def retrieval_query(adventure: models.Adventure, exclude_action_id: int | None = None) -> dict:
|
|
"""The two texts a turn's memory retrieval embeds, and the words it matches.
|
|
|
|
Returns `{"input", "context", "input_terms"}`. `input` is empty when the
|
|
newest action is not a player action with text, which is a continue turn or a
|
|
dry run. `context` is empty only for a story with no scene and no narration.
|
|
"""
|
|
recent = history.tail(adventure, 2, exclude_action_id)
|
|
newest = recent[-1] if recent else None
|
|
player_input = ""
|
|
if newest is not None and newest.type in INPUT_TYPES:
|
|
player_input = truncate_to_last_tokens(newest.text.strip(), QUERY_INPUT_TOKENS)
|
|
narration = recent[0].text if len(recent) > 1 else ""
|
|
else:
|
|
narration = newest.text if newest is not None else ""
|
|
context = "\n".join(part for part in (
|
|
_scene_text(adventure.narrative_state),
|
|
truncate_to_last_tokens(narration.strip(), QUERY_NARRATION_TOKENS),
|
|
) if part.strip())
|
|
return {
|
|
"input": player_input,
|
|
"context": context,
|
|
"input_terms": sorted(lexical_terms(player_input)),
|
|
}
|
|
|
|
|
|
def score_candidates(
|
|
ids: list[int],
|
|
held: dict,
|
|
terms_of: dict[int, frozenset[str]],
|
|
input_vec,
|
|
context_vec,
|
|
input_terms,
|
|
) -> list[tuple[float, int, float, float]]:
|
|
"""`(final_score, memory_id, semantic_score, lexical_score)`, best first.
|
|
|
|
Ties on the final score are broken by id, so the order never depends on the
|
|
order the database returned rows in.
|
|
"""
|
|
lexical = lexical_scores(frozenset(input_terms), {i: terms_of.get(i, frozenset()) for i in ids})
|
|
rows = []
|
|
for memory_id in ids:
|
|
vector = held[memory_id]
|
|
if input_vec is not None and context_vec is not None:
|
|
semantic = (INPUT_WEIGHT * cosine(input_vec, vector)
|
|
+ (1.0 - INPUT_WEIGHT) * cosine(context_vec, vector))
|
|
else:
|
|
semantic = cosine(input_vec if input_vec is not None else context_vec, vector)
|
|
lex = lexical.get(memory_id, 0.0)
|
|
rows.append((semantic + LEXICAL_WEIGHT * lex, memory_id, semantic, lex))
|
|
rows.sort(key=lambda row: (-row[0], row[1]))
|
|
return rows
|
|
|
|
|
|
def select_memories(scored, pinned_of, held, authority_of, top_k):
|
|
"""Pins first, then the best-scoring rest, skipping repeats (§22).
|
|
|
|
Returns `(used, suppressed)`. `used` is `(final_score, memory_id, pinned)`
|
|
rows, best first.
|
|
"""
|
|
rows = [(final, memory_id, pinned_of[memory_id]) for final, memory_id, _, _ in scored]
|
|
used = [row for row in rows if row[2]]
|
|
remaining = max(0, top_k - len(used))
|
|
candidates = [row for row in rows if not row[2]]
|
|
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
|
|
used += kept
|
|
used.sort(key=lambda row: (-row[0], row[1]))
|
|
return used, suppressed
|
|
|
|
|
|
async def retrieve_memories(
|
|
adventure: models.Adventure,
|
|
settings: models.Settings,
|
|
*,
|
|
exclude_action_id: int | None = None,
|
|
) -> dict | None:
|
|
"""Returns the memories to inject, or None when the bank is off.
|
|
|
|
The result is a dict of the form
|
|
`{"used": [{id, text, similarity, semantic_score, lexical_score,
|
|
final_score, pinned, authority, source}], "query": {...}, "error": str | None}`.
|
|
`similarity` is the semantic score, under the name the inspector has always
|
|
shown. It is None when the memory bank is disabled for this adventure.
|
|
|
|
This only reads. A turn counts the memories it used with `record_use`, just
|
|
before the commit that saves the turn; see that function for why the count
|
|
cannot be written here.
|
|
|
|
`exclude_action_id` removes the action being retried from the query, so that
|
|
a discarded attempt cannot influence which memories are returned.
|
|
"""
|
|
if not adventure.memory_bank_enabled:
|
|
return None
|
|
if not settings.embedding_model.strip():
|
|
return {"used": [], "error": "No embedding model configured in Settings."}
|
|
db = object_session(adventure)
|
|
if db is None:
|
|
return {"used": [], "error": None}
|
|
|
|
# Select which memories are in play, and nothing else about them. This code
|
|
# used to walk `adventure.memories`, which loaded every row of the bank,
|
|
# including its vector. That cost about 31 KB per memory and about 3 MB per
|
|
# turn, which was 96% of everything a turn read. An id and a flag come to
|
|
# about eight bytes per row.
|
|
#
|
|
# The branch clause uses the whole lineage here rather than the window the
|
|
# story is read through. Retrieval exists to recall events from far back in
|
|
# the story, such as what happened forty turns ago. The full lineage stays
|
|
# affordable because memories are sparse, at roughly one per six actions, so
|
|
# even a heavily forked story returns only tens of small rows.
|
|
catalogue = db.execute(
|
|
select(models.Memory.id, models.Memory.pinned, models.Memory.authority).where(
|
|
models.Memory.adventure_id == adventure.id,
|
|
lineage.path_of(db, adventure).clause(models.Memory),
|
|
models.Memory.forgotten.is_(False),
|
|
models.Memory.embedded.is_(True),
|
|
)
|
|
).all()
|
|
if not catalogue:
|
|
return {"used": [], "error": None}
|
|
|
|
query = retrieval_query(adventure, exclude_action_id)
|
|
texts = [t for t in (query["input"], query["context"]) if t.strip()]
|
|
if not texts:
|
|
return {"used": [], "error": None}
|
|
|
|
try:
|
|
embedded = await embedding_provider(settings).embed(texts)
|
|
except ProviderError as exc:
|
|
return {"used": [], "error": str(exc)}
|
|
vectors_by_text = dict(zip(texts, embedded))
|
|
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
|
|
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
|
|
|
|
ids = [memory_id for memory_id, _, _ in catalogue]
|
|
held = _vectors_for(db, adventure.id, ids)
|
|
# Memory text is read only when there are input words to match against, and
|
|
# then only for memories not already held (see `_terms_for`).
|
|
terms_of = _terms_for(db, adventure.id, ids) if query["input_terms"] else {}
|
|
authority_of = {memory_id: authority for memory_id, _, authority in catalogue}
|
|
pinned_of = {memory_id: pinned for memory_id, pinned, _ in catalogue}
|
|
scored = score_candidates(
|
|
[memory_id for memory_id in ids if memory_id in held],
|
|
held, terms_of, input_vec, context_vec, query["input_terms"],
|
|
)
|
|
components = {memory_id: (semantic, lex) for _, memory_id, semantic, lex in scored}
|
|
used, suppressed = select_memories(
|
|
scored, pinned_of, held, authority_of, max(1, settings.memory_top_k))
|
|
if not used:
|
|
return {"used": [], "error": None}
|
|
|
|
# Fetch the text only now, and only for the `top_k` rows that were chosen.
|
|
#
|
|
# M6 adds authority and provenance to this same read rather than to a second
|
|
# one. The columns are narrow, the row set is `top_k`, and fetching them
|
|
# here is what keeps "why did the narrator remember this?" answerable
|
|
# without a query per memory (F06, and the N+1 discipline M5 restored).
|
|
used_ids = [memory_id for _, memory_id, _ in used]
|
|
detail = {
|
|
row.id: row
|
|
for row in db.execute(
|
|
select(
|
|
models.Memory.id, models.Memory.text, models.Memory.authority,
|
|
models.Memory.branch_id, models.Memory.depth,
|
|
models.Memory.source_start, models.Memory.source_end,
|
|
).where(models.Memory.id.in_(used_ids))
|
|
).all()
|
|
}
|
|
texts_of = {memory_id: row.text for memory_id, row in detail.items()}
|
|
|
|
return {
|
|
"used": [
|
|
{
|
|
"id": memory_id,
|
|
"text": texts_of.get(memory_id, ""),
|
|
"similarity": round(components[memory_id][0], 4),
|
|
# v1.1 WP-B.2: the parts of the score, so an inspector can see
|
|
# why this memory beat the ones below it.
|
|
"semantic_score": round(components[memory_id][0], 4),
|
|
"lexical_score": round(components[memory_id][1], 4),
|
|
"final_score": round(final, 4),
|
|
"pinned": pinned,
|
|
# M6: what weight this carries, and where it came from.
|
|
"authority": getattr(detail.get(memory_id), "authority", ACCEPTED_STORY),
|
|
"source": {
|
|
"branch_id": getattr(detail.get(memory_id), "branch_id", None),
|
|
"depth": getattr(detail.get(memory_id), "depth", None),
|
|
"source_start": getattr(detail.get(memory_id), "source_start", None),
|
|
"source_end": getattr(detail.get(memory_id), "source_end", None),
|
|
},
|
|
}
|
|
for final, memory_id, pinned in used
|
|
],
|
|
"considered": len(catalogue),
|
|
# M6: how many candidates were set aside as repeating one already
|
|
# chosen. Visible so that "why is that memory not here?" has an answer.
|
|
"suppressed": [
|
|
{"id": memory_id, "duplicate_of": kept_id}
|
|
for memory_id, kept_id in suppressed
|
|
],
|
|
# v1.1 WP-B.2: what was searched for. Recorded per turn, like the rest.
|
|
"query": {
|
|
"input": query["input"],
|
|
"context": query["context"],
|
|
"input_terms": query["input_terms"],
|
|
"input_weight": INPUT_WEIGHT if input_vec is not None and context_vec is not None
|
|
else (1.0 if input_vec is not None else 0.0),
|
|
"lexical_weight": LEXICAL_WEIGHT,
|
|
},
|
|
"error": None,
|
|
}
|
|
|
|
|
|
def record_use(db: Session, memory_bank: dict | None) -> None:
|
|
"""Counts the memories a turn was given, as part of that turn's commit.
|
|
|
|
Call this immediately before the commit that saves the turn, and never
|
|
before the model call. This counter used to be written during retrieval, and
|
|
the UPDATE opened a write transaction that stayed open for the whole reply,
|
|
because the turn commits only once the narration has streamed. SQLite has
|
|
one writer. Every post-turn memory, summary and status write that arrived
|
|
during the reply waited out the driver's five-second timeout and failed with
|
|
`database is locked`. Recording those failures also needs a write, so it
|
|
failed the same way, and derived status kept reporting `idle`. A 26-turn
|
|
run on a GPU host wrote two memories and no summary while every turn was
|
|
accepted.
|
|
|
|
Only real turns count, never Insights' dry runs. A turn that fails before
|
|
its commit counts nothing, because nothing was used.
|
|
|
|
Pass `synchronize_session=False` because nothing in this request reads the
|
|
counters back. Matching the UPDATE against loaded objects would require
|
|
loading those objects, which is the cost this code path exists to avoid.
|
|
"""
|
|
used_ids = [m["id"] for m in (memory_bank or {}).get("used") or []]
|
|
if not used_ids:
|
|
return
|
|
db.execute(
|
|
update(models.Memory)
|
|
.where(models.Memory.id.in_(used_ids))
|
|
.values(use_count=models.Memory.use_count + 1, last_used_at=models.utcnow())
|
|
.execution_options(synchronize_session=False)
|
|
)
|
|
|
|
|
|
# ---------- Post-turn background work ----------
|
|
|
|
def schedule_post_turn(adventure: models.Adventure) -> None:
|
|
"""Fire-and-forget summarization/embedding work after a turn is saved.
|
|
|
|
M7 adds a third reason to run: imported passages that still need vectors.
|
|
Without it a campaign that plays with story memory switched off would never
|
|
catch up an import whose embedding failed, and the only repair would be an
|
|
explicit Reindex.
|
|
"""
|
|
if not (
|
|
adventure.auto_summarize
|
|
or adventure.memory_bank_enabled
|
|
or adventure.knowledge_sources
|
|
):
|
|
return
|
|
if adventure.id in _running:
|
|
return
|
|
task = asyncio.get_running_loop().create_task(run_post_turn(adventure.id))
|
|
_tasks.add(task)
|
|
task.add_done_callback(_tasks.discard)
|
|
|
|
|
|
async def run_post_turn(adventure_id: int) -> None:
|
|
if adventure_id in _running:
|
|
return
|
|
_running.add(adventure_id)
|
|
db = SessionLocal()
|
|
try:
|
|
adventure = db.get(models.Adventure, adventure_id)
|
|
if adventure is None:
|
|
return
|
|
# Settings are per-user (Phase 8): use the adventure owner's row.
|
|
settings = (
|
|
db.query(models.Settings)
|
|
.filter(models.Settings.user_id == adventure.user_id)
|
|
.first()
|
|
)
|
|
if settings is None:
|
|
return
|
|
# This code no longer clamps the cursors. Undo can leave the story
|
|
# shorter than the mark. When the mark was a position, a value past the
|
|
# end of the list stalled the pass until the story grew back, so every
|
|
# post-turn run clamped it. That clamp introduced its own error, because
|
|
# clamping to the settled count rewound a caught-up adventure by one
|
|
# step and covered an action twice.
|
|
#
|
|
# An anchor past the tip is not an invalid value. `settled_after`
|
|
# reports that there is nothing to do, and once the story grows past the
|
|
# anchor the pass resumes where it stopped.
|
|
# M6. Each kind runs inside its own recorder, so one failing pass
|
|
# neither hides the others nor takes the turn down with it. The accepted
|
|
# narration, its state events, the authoritative document and the head
|
|
# were all committed before this task started; nothing here may undo
|
|
# them, and nothing here may fail without leaving a record
|
|
# (`BUILD-MILESTONES.md`, note from M2).
|
|
if adventure.auto_summarize:
|
|
await _guarded(db, adventure_id, derived.MEMORY,
|
|
_create_due_memories(adventure, settings, db))
|
|
await _guarded(db, adventure_id, derived.SUMMARY,
|
|
_update_story_summary(adventure, settings, db))
|
|
if adventure.memory_bank_enabled and settings.embedding_model.strip():
|
|
await _guarded(db, adventure_id, derived.EMBEDDING,
|
|
_embed_pending(adventure, settings, db))
|
|
# M7: the imported knowledge library's own vectors, caught up here.
|
|
#
|
|
# Import embeds what it can at the moment the file arrives. This is what
|
|
# happens when that failed, when the endpoint was down, when the reader
|
|
# configured an embedding model afterwards, or when a library was large
|
|
# enough that one pass did not finish it. It is not conditioned on
|
|
# `memory_bank_enabled`: the knowledge library is a separate subsystem
|
|
# and a reader who turned story memory off did not thereby ask for their
|
|
# imported Canon to stop being searchable.
|
|
#
|
|
# `embed_pending` records its own outcome, per source and per campaign,
|
|
# and never raises — so unlike the passes above it needs no guard, and
|
|
# wrapping it in one would overwrite the finer-grained record it just
|
|
# wrote with a coarser one.
|
|
if settings.embedding_model.strip():
|
|
await knowledge_embeddings.embed_pending(db, adventure, settings)
|
|
db.commit()
|
|
_evict_over_capacity(adventure, settings, db)
|
|
except BaseException as exc: # noqa: BLE001 - the task boundary
|
|
# Anything the per-kind guards did not catch: a failure in the shared
|
|
# setup above, or in eviction. M2's lesson is that the one thing this
|
|
# may not do is vanish. Re-raising would only feed an unobserved task.
|
|
try:
|
|
# Roll back first. The failure is often a flush or commit that
|
|
# failed, which leaves the session unusable until it is rolled
|
|
# back, and the record then fails with `PendingRollbackError`
|
|
# instead of being written. `_guarded` already does this.
|
|
db.rollback()
|
|
derived.failed(db, adventure_id, derived.MEMORY, exc)
|
|
db.commit()
|
|
except BaseException: # noqa: BLE001 - the recorder must not mask it
|
|
log.exception("could not record derived-work failure for %s", adventure_id)
|
|
finally:
|
|
db.close()
|
|
_running.discard(adventure_id)
|
|
|
|
|
|
async def _guarded(db: Session, adventure_id: int, kind: str, coro) -> None:
|
|
"""Runs one derived pass, recording whether it worked.
|
|
|
|
The pass keeps whatever it committed before it failed — a memory written
|
|
two blocks ago stays written — because derived work is additive and
|
|
partial progress is still progress. What must not survive is an
|
|
uncommitted, half-written unit of work, so the session is rolled back to
|
|
the last commit before the failure is recorded.
|
|
"""
|
|
try:
|
|
did_work = await coro
|
|
except BaseException as exc: # noqa: BLE001 - one kind must not stop another
|
|
db.rollback()
|
|
derived.failed(db, adventure_id, kind, exc)
|
|
db.commit()
|
|
else:
|
|
derived.succeeded(db, adventure_id, kind, did_work=bool(did_work))
|
|
db.commit()
|
|
|
|
|
|
def _excerpt_encoding():
|
|
return _token_encoding()
|
|
|
|
|
|
def excerpt_split(budget: int = MEMORY_EXCERPT_TOKENS) -> tuple[int, int]:
|
|
"""`(head_tokens, tail_tokens)` for a block longer than `budget`.
|
|
|
|
The marker and the blank lines around it are paid for first; what is left is
|
|
halved, and an odd token goes to the tail, the most recent part. So the two
|
|
parts plus the marker come to exactly `budget`.
|
|
"""
|
|
room = max(0, budget - count_tokens(f"\n\n{EXCERPT_OMISSION_MARKER}\n\n"))
|
|
head = room // 2
|
|
return head, room - head
|
|
|
|
|
|
def memory_excerpt(raw: str, budget: int = MEMORY_EXCERPT_TOKENS) -> str:
|
|
"""What the summariser is shown of one block.
|
|
|
|
v1.1 WP-B.2. A block that fits in `budget` tokens is sent whole, exactly as
|
|
before. A longer block used to be cut to its last `budget` tokens, and B.1
|
|
showed that a fact near its start then never reached the summariser at all.
|
|
It is now sent as its opening and its end, in order, with
|
|
`EXCERPT_OMISSION_MARKER` between them, still inside `budget`.
|
|
|
|
Rejoining two token runs can tokenise a little differently at the seams, so
|
|
the result is measured, and the head gives up tokens until it fits. A fact in
|
|
the middle of a very long block is still left out: this bounds the input, it
|
|
does not summarise everything.
|
|
"""
|
|
enc = _excerpt_encoding()
|
|
tokens = enc.encode(raw)
|
|
if len(tokens) <= budget:
|
|
return raw
|
|
head_n, tail_n = excerpt_split(budget)
|
|
while True:
|
|
excerpt = (f"{enc.decode(tokens[:head_n]).rstrip()}\n\n{EXCERPT_OMISSION_MARKER}\n\n"
|
|
f"{enc.decode(tokens[-tail_n:]).lstrip()}" if tail_n else
|
|
enc.decode(tokens[:head_n]))
|
|
over = count_tokens(excerpt) - budget
|
|
if over <= 0 or head_n == 0:
|
|
return excerpt
|
|
head_n = max(0, head_n - over)
|
|
|
|
|
|
def memory_user_prompt(brief: str, excerpt: str) -> str:
|
|
"""The user message of a memory call: the cast brief, then the excerpt.
|
|
|
|
Kept apart from `summarize_block` so an evaluation can send a model exactly
|
|
what the application sends (v1.1 WP-B.2, `tools/memory_fidelity.py`).
|
|
"""
|
|
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
|
|
return f"{brief}\n\n{prompt}" if brief else prompt
|
|
|
|
|
|
async def summarize_block(
|
|
adventure: models.Adventure,
|
|
provider: OpenAICompatibleProvider,
|
|
block: list[models.Action],
|
|
) -> str:
|
|
"""Writes one memory from one block of story.
|
|
|
|
Both callers come through here, which is the point of the function. The
|
|
pass below writes a memory as the story reaches it; `tools/rewrite_memories`
|
|
rewrites one an older prompt produced. Assembling the prompt in two places
|
|
would mean a rewritten memory was written by a prompt that never shipped,
|
|
and nothing would report the difference.
|
|
|
|
Raises `ProviderError`, which each caller handles its own way: the pass
|
|
below leaves the cursor alone and retries next turn, and the tool leaves the
|
|
old text in place and moves on.
|
|
"""
|
|
raw = "\n\n".join(a.text for a in block)
|
|
excerpt = memory_excerpt(raw)
|
|
# Match the cast against the untruncated block. The excerpt is what the
|
|
# model reads, but a character named in the part that was left out is still
|
|
# one the memory may have to name.
|
|
brief = cast_brief(adventure, raw)
|
|
text = await provider.complete(MEMORY_SYSTEM_PROMPT, memory_user_prompt(brief, excerpt))
|
|
# The marker is an instruction to the summariser, never a fact of the story.
|
|
if text and EXCERPT_OMISSION_MARKER in text:
|
|
text = " ".join(text.replace(EXCERPT_OMISSION_MARKER, " ").split())
|
|
return text
|
|
|
|
|
|
async def _create_due_memories(
|
|
adventure: models.Adventure, settings: models.Settings, db: Session
|
|
) -> int:
|
|
"""Writes the memories that are due. Returns how many it wrote (M6-F5)."""
|
|
provider = summary_provider(settings)
|
|
written = 0
|
|
for _ in range(MAX_MEMORIES_PER_RUN):
|
|
# Re-read the anchor on every pass. Committing a memory does not change
|
|
# the story, but this loop is the only code that moves the anchor, so
|
|
# both numbers must be current.
|
|
anchor = cursors.MEMORY.depth(db, adventure)
|
|
if history.count_after(adventure, anchor) < MEMORY_INTERVAL + SETTLE_SLACK:
|
|
return written # No settled block of story sits past the mark. The block
|
|
# itself is still MEMORY_INTERVAL actions; the slack asks
|
|
# for story past its end. See `SETTLE_SLACK`.
|
|
if history.count(adventure) < MEMORY_START:
|
|
return written # The adventure is too short to have started summarizing.
|
|
# The order of those two checks is deliberate. The usual answer is that
|
|
# no memory is due, and the first check settles that without measuring
|
|
# the length of the whole story.
|
|
block = history.after(adventure, anchor, MEMORY_INTERVAL)
|
|
if len(block) < MEMORY_INTERVAL:
|
|
return written
|
|
# A provider failure is no longer caught here. `_guarded` records it
|
|
# against this campaign, and the cursor is unchanged either way, so the
|
|
# next accepted turn retries this same block (M6).
|
|
text = await summarize_block(adventure, provider, block)
|
|
if not text:
|
|
return written
|
|
memory = models.Memory(
|
|
adventure_id=adventure.id,
|
|
text=text,
|
|
source_start=block[0].depth,
|
|
source_end=block[-1].depth,
|
|
authority=classify_authority(text),
|
|
)
|
|
# Attach the memory to the node it summarizes, so that a fork inherits
|
|
# the memories of the path it forked from and no others. Then move the
|
|
# mark to that same node. Both values record how far this pass has
|
|
# reached, and taking them from one row keeps them in step even when the
|
|
# depths have gaps.
|
|
tree.attach_memory(memory, block[-1])
|
|
db.add(memory)
|
|
cursors.MEMORY.anchor_at(adventure, block[-1])
|
|
db.commit()
|
|
written += 1
|
|
return written
|
|
|
|
|
|
async def _update_story_summary(
|
|
adventure: models.Adventure, settings: models.Settings, db: Session
|
|
) -> bool:
|
|
"""Rolls the summary forward when enough new story has settled.
|
|
|
|
Returns whether it wrote one (M6-F5)."""
|
|
anchor = cursors.SUMMARY.depth(db, adventure)
|
|
uncovered = history.count_after(adventure, anchor)
|
|
if uncovered < SUMMARY_INTERVAL:
|
|
return False
|
|
# Where the summary stands once this run succeeds. Read this before the AI
|
|
# call rather than after it. The mark records the end of the story as this
|
|
# pass saw it, and a turn that arrives during the call must not be counted
|
|
# as read.
|
|
caught_up = history.newest(adventure)
|
|
if caught_up is None:
|
|
return False
|
|
|
|
# Include the memories for the stretch that the summary has not read, which
|
|
# means every memory attached to a node past the anchor. The marks and the
|
|
# memories are now depths on one path, so no coordinate conversion remains.
|
|
# If memory creation has fallen behind, for example because the last attempt
|
|
# failed, this falls back to the raw story text.
|
|
new_events = db.execute(
|
|
select(models.Memory.text)
|
|
.where(
|
|
models.Memory.adventure_id == adventure.id,
|
|
lineage.path_of(db, adventure).clause(models.Memory),
|
|
models.Memory.depth > anchor,
|
|
)
|
|
.order_by(models.Memory.depth)
|
|
).scalars().all()
|
|
if new_events:
|
|
events_text = "\n".join(f"- {t}" for t in new_events)
|
|
else:
|
|
block = history.after(adventure, anchor, uncovered)
|
|
events_text = truncate_to_last_tokens("\n\n".join(a.text for a in block), 2000)
|
|
|
|
# M6 corrective (review finding M6-F1). The previous summary this one
|
|
# builds on has to be a summary that is *valid where the story now stands*,
|
|
# not merely the last one written.
|
|
#
|
|
# Seeding from `adventure.story_summary` — a campaign-global column with no
|
|
# lineage — is what broke E03. After a divergence that column still held the
|
|
# abandoned line's prose, so the summariser was handed it and asked to
|
|
# update it. The row it produced was correctly anchored to the new branch
|
|
# and was therefore *reported* as lineage-safe, while its sentences
|
|
# described a story the reader had left. The row was anchored; the content
|
|
# was not.
|
|
#
|
|
# `summaries.current` answers the same question the context builder asks —
|
|
# which summary is eligible at the head — so the input and the output are
|
|
# now scoped by one rule. Where no eligible summary exists, the new line
|
|
# starts from nothing, which is the truthful starting point for a story
|
|
# that has not been summarised yet.
|
|
eligible = summaries.current(db, adventure)
|
|
current = eligible.text.strip() if eligible is not None else ""
|
|
# The summary is built from the memories, so it inherits their framing for
|
|
# free once they are named and third-person. It still gets the brief of its
|
|
# own, because the fallback above hands it raw second-person story text
|
|
# whenever memory creation has fallen behind.
|
|
brief = cast_brief(adventure, f"{current}\n\n{events_text}")
|
|
user_prompt = (
|
|
f"Current story summary:\n{current or '(none yet)'}\n\n"
|
|
f"New events since the last update:\n{events_text}\n\n"
|
|
"Updated summary:"
|
|
)
|
|
if brief:
|
|
user_prompt = f"{brief}\n\n{user_prompt}"
|
|
text = await summary_provider(settings).complete(
|
|
SUMMARY_SYSTEM_PROMPT, user_prompt, max_tokens=600
|
|
)
|
|
if not text:
|
|
return False
|
|
# M6: anchored to the story it summarizes rather than written into a single
|
|
# column. `caught_up` is the last node it covers, so the row is eligible on
|
|
# exactly the lineages that contain that node, and an Undo or a divergence
|
|
# makes it ineligible without deleting it (E03, `summaries` module).
|
|
summaries.record(
|
|
db, adventure, text,
|
|
node=caught_up,
|
|
source_start=anchor + 1 if anchor is not None else None,
|
|
trigger="interval",
|
|
model_name=(settings.summary_model or settings.model or ""),
|
|
)
|
|
cursors.SUMMARY.anchor_at(adventure, caught_up)
|
|
db.commit()
|
|
return True
|
|
|
|
|
|
async def _embed_pending(
|
|
adventure: models.Adventure, settings: models.Settings, db: Session
|
|
) -> int:
|
|
"""Embeds memories that have no vector. Returns how many (M6-F5)."""
|
|
# Use a query rather than walking `adventure.memories`. That walk ran on
|
|
# every turn and loaded the whole bank's vectors in order to find the few
|
|
# rows with none.
|
|
#
|
|
# Neither this query nor the eviction below applies a branch clause, and
|
|
# that is deliberate. Whether a row is embedded is a fact about the row, not
|
|
# about the path being played. Skipping a sibling branch's memories would
|
|
# only postpone the work until someone switched branches and needed them
|
|
# ranked. Capacity works the same way. The bank belongs to the adventure,
|
|
# and the memories of a branch nobody is reading are the right ones to evict
|
|
# first.
|
|
pending = (
|
|
db.query(models.Memory)
|
|
.filter(
|
|
models.Memory.adventure_id == adventure.id,
|
|
models.Memory.embedded.is_(False),
|
|
models.Memory.forgotten.is_(False),
|
|
)
|
|
.order_by(models.Memory.id)
|
|
.limit(MAX_EMBED_BATCH)
|
|
.all()
|
|
)
|
|
if not pending:
|
|
return 0
|
|
try:
|
|
new = await embedding_provider(settings).embed([m.text for m in pending])
|
|
except ProviderError:
|
|
return 0
|
|
for memory, vector in zip(pending, new):
|
|
set_vector(memory, vector)
|
|
db.commit()
|
|
return len(pending)
|
|
|
|
|
|
def eviction_order(rows, limit: int) -> list[int]:
|
|
"""The ids eviction would take, first to last, at most `limit` of them.
|
|
|
|
v1.1 WP-B.2. `rows` are the active memories of one adventure, each with
|
|
`id`, `pinned`, `source_start`, `source_end`, `last_used_at`, `created_at`
|
|
and `use_count`. Nothing here reads a vector or the database, so the same
|
|
function is what the eviction pass runs and what a diagnostic reports.
|
|
|
|
WP-B.1 showed what pure least-recently-used order does to a long campaign.
|
|
Retrieval is steered by the present scene, so a memory of an early stretch
|
|
nothing recent resembles stops being used. It then becomes the least
|
|
recently used row, and it goes first, while the bank keeps several memories
|
|
of the last few scenes that the history window still holds in full. The
|
|
rule below keeps the bank spread over the whole story instead.
|
|
|
|
**Coverage.** Memories with a source range say which stretch of the story
|
|
they describe. A memory is judged by the hole its removal would leave: the
|
|
number of depths between the end of the nearest memory before it and the
|
|
start of the nearest memory after it. The smallest hole goes first, so the
|
|
bank thins where it is densest. A memory whose start another memory shares
|
|
(a retried or re-played stretch, or a sibling line) leaves no hole, and is
|
|
the first kind to go. Pinned memories count as coverage, since they stay.
|
|
|
|
**Boundaries.** The earliest and the latest memory by position leave a hole
|
|
with no memory on one side: removing the first loses the only record of the
|
|
opening, and removing the last loses the only record of the most recent
|
|
stretch, which is also what keeps a memory written this turn from being
|
|
evicted by the pass that wrote it (the frozen bank, below). Boundaries are
|
|
not coverage candidates.
|
|
|
|
**Recency.** Among memories whose removal leaves the same hole, the least
|
|
recently used goes first (`coalesce(last_used_at, created_at)`), then the
|
|
less used, then the lower id. Ties are therefore never left to the order the
|
|
database returned rows in.
|
|
|
|
**Fallback.** When no memory is a coverage candidate — memories typed by the
|
|
player or migrated from before coordinates have no range, and a bank can be
|
|
all boundaries — the rest are taken least recently used first, exactly as
|
|
v1.0.0 did. The bank stays bounded either way. Pinned memories are never
|
|
taken; if every active memory is pinned, capacity yields to the pins.
|
|
|
|
Recomputed after each pick, because removing one memory widens the holes
|
|
of its neighbours.
|
|
"""
|
|
remaining = {row.id: row for row in rows if not row.pinned}
|
|
coverers = {row.id: row for row in rows
|
|
if row.source_start is not None and row.source_end is not None}
|
|
|
|
def recency(row):
|
|
return (row.last_used_at or row.created_at, row.use_count or 0, row.id)
|
|
|
|
order: list[int] = []
|
|
while remaining and len(order) < limit:
|
|
spans = sorted(coverers.values(), key=lambda r: (r.source_start, r.source_end, r.id))
|
|
starts: dict[int, int] = {}
|
|
for row in spans:
|
|
starts[row.source_start] = starts.get(row.source_start, 0) + 1
|
|
best = None
|
|
furthest_end = None # the largest source_end before index i
|
|
for i, row in enumerate(spans):
|
|
if row.id in remaining:
|
|
if starts[row.source_start] > 1:
|
|
cost = 0
|
|
elif i == 0 or i == len(spans) - 1:
|
|
cost = None # a boundary
|
|
else:
|
|
cost = max(0, spans[i + 1].source_start - furthest_end - 1)
|
|
if cost is not None:
|
|
key = (cost, *recency(row))
|
|
if best is None or key < best[0]:
|
|
best = (key, row.id)
|
|
furthest_end = row.source_end if furthest_end is None else max(furthest_end, row.source_end)
|
|
if best is None:
|
|
victim = min(remaining.values(), key=recency).id
|
|
else:
|
|
victim = best[1]
|
|
order.append(victim)
|
|
del remaining[victim]
|
|
coverers.pop(victim, None)
|
|
return order
|
|
|
|
|
|
def _evict_over_capacity(
|
|
adventure: models.Adventure, settings: models.Settings, db: Session
|
|
) -> None:
|
|
# The database performs the count, and the rows read for ordering carry
|
|
# neither text nor vectors. Counting by walking `adventure.memories` fetched
|
|
# every vector in the bank on every turn, whether or not the bank was over
|
|
# capacity.
|
|
in_this_bank = (models.Memory.adventure_id == adventure.id,
|
|
models.Memory.forgotten.is_(False))
|
|
active = db.execute(
|
|
select(func.count(models.Memory.id)).where(*in_this_bank)
|
|
).scalar() or 0
|
|
overflow = active - max(1, settings.memory_bank_capacity)
|
|
if overflow <= 0:
|
|
return
|
|
# v1.1 WP-B.2: the order is `eviction_order`, coverage first and recency
|
|
# second. It replaces least recently used alone; see that function.
|
|
#
|
|
# What the old ordering fixed still holds. Ordering by use count first froze
|
|
# the bank: a memory written on this turn has never been used, so once every
|
|
# other memory had been retrieved at least once, the new memory held the
|
|
# lowest count in the bank, and the same post-turn run that wrote it evicted
|
|
# it. Counts only increase, so the bank never recovered. Under the coverage
|
|
# rule the newest memory is the latest boundary, so it is not a coverage
|
|
# candidate, and in the fallback it carries the newest timestamp.
|
|
rows = db.execute(
|
|
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
|
|
models.Memory.source_end, models.Memory.last_used_at,
|
|
models.Memory.created_at, models.Memory.use_count)
|
|
.where(*in_this_bank)
|
|
).all()
|
|
doomed = eviction_order(rows, overflow)
|
|
if not doomed:
|
|
return # Every active memory is pinned, so the pins override capacity.
|
|
db.execute(
|
|
update(models.Memory)
|
|
.where(models.Memory.id.in_(doomed))
|
|
.values(forgotten=True)
|
|
.execution_options(synchronize_session=False)
|
|
)
|
|
db.commit()
|
|
# The bulk UPDATE bypassed the loaded objects, so code that still holds the
|
|
# collection would otherwise see the evicted memories as active.
|
|
db.expire(adventure, ["memories"])
|