M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
co-authored by
Claude Opus 5
parent
b7005e6fdd
commit
a6e9c7a32b
+10
-2
@@ -34,9 +34,9 @@ agreed in every group.
|
||||
|
||||
import copy
|
||||
|
||||
from sqlalchemy.orm import Session, undefer
|
||||
from sqlalchemy.orm import Session, object_session, undefer
|
||||
|
||||
from . import models
|
||||
from . import models, summaries
|
||||
from .context import lineage
|
||||
from .narrative import model as narrative_model
|
||||
|
||||
@@ -191,6 +191,14 @@ def restore_state(adventure: models.Adventure, node: models.Action | None) -> No
|
||||
if isinstance(node.narrative_state_after, dict)
|
||||
else narrative_model.empty()
|
||||
)
|
||||
# M6: the reader-facing summary mirror follows the head too. It is a
|
||||
# convenience column with no lineage of its own, so without this it would go
|
||||
# on showing a summary belonging to a position the story has left. Nothing
|
||||
# authoritative reads it — the prompt takes its summary from
|
||||
# `summaries.current` — but the Plot panel and the export bundle do.
|
||||
session = object_session(adventure)
|
||||
if session is not None:
|
||||
summaries.refresh_mirror(session, adventure)
|
||||
# Legacy, and deliberately still restored: a pre-M5 campaign's numbers stay
|
||||
# coherent with the position being read, so an old save is not left showing
|
||||
# a future's values. Nothing consults them to decide anything.
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
from . import history
|
||||
from .builder import (
|
||||
ContextOverflow,
|
||||
build_context,
|
||||
count_tokens,
|
||||
match_cards,
|
||||
@@ -9,7 +10,7 @@ from .builder import (
|
||||
from .history import story_actions
|
||||
|
||||
__all__ = [
|
||||
"build_context",
|
||||
"ContextOverflow", "build_context",
|
||||
"count_tokens",
|
||||
"history",
|
||||
"match_cards",
|
||||
|
||||
@@ -21,8 +21,9 @@ block and the live sections in `build_context`.
|
||||
from dataclasses import dataclass
|
||||
|
||||
import tiktoken
|
||||
from sqlalchemy.orm import object_session
|
||||
|
||||
from .. import models, narrative, worldstate
|
||||
from .. import derived, models, narrative, summaries, worldstate
|
||||
from . import encoding, history
|
||||
|
||||
AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
|
||||
@@ -63,6 +64,24 @@ MAX_LENGTH_FLOOR_WORDS = 300
|
||||
# Built from the table vendored in `encoding.py`, not fetched: the upstream
|
||||
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
|
||||
# is called on every turn.
|
||||
# M6: added to the configured reply budget when reserving output space. It
|
||||
# absorbs the section separators added after budgeting and the drift between
|
||||
# this tokenizer and the serving model's. Fixed rather than proportional: what
|
||||
# it covers does not grow with the size of the budget.
|
||||
OUTPUT_SAFETY_MARGIN = 64
|
||||
|
||||
|
||||
class ContextOverflow(RuntimeError):
|
||||
"""Raised when protected context alone cannot fit in the token budget.
|
||||
|
||||
Protected means the narrator rules, the campaign canon, the authoritative
|
||||
narrative state, the reader's own input, and the reserve for the reply
|
||||
(`CONTEXT-AND-MEMORY.md` §30). None of those may be dropped to make room for
|
||||
old prose, so when they do not fit there is no prompt to build and saying so
|
||||
is the only honest answer.
|
||||
"""
|
||||
|
||||
|
||||
def _encoding() -> tiktoken.Encoding:
|
||||
return encoding.get_encoding()
|
||||
|
||||
@@ -187,6 +206,12 @@ def _history_text(action: models.Action) -> str:
|
||||
return action.text
|
||||
|
||||
|
||||
def _memory_line(memory: dict) -> str:
|
||||
"""One retrieved memory, marked with its authority (M6)."""
|
||||
mark = " [inferred]" if memory.get("authority") == "heuristic" else ""
|
||||
return f"-{mark} {memory['text']}"
|
||||
|
||||
|
||||
def _canon_section(adventure: models.Adventure) -> str:
|
||||
"""The campaign's own rules, rendered for the system block.
|
||||
|
||||
@@ -317,15 +342,33 @@ def build_context(
|
||||
# memories change on most turns, and the stat values change on nearly every
|
||||
# turn. `world_lore` is added below, because the history window determines
|
||||
# which cards trigger and that window is not known yet.
|
||||
# M6: the summary the *current lineage* is entitled to, not whatever was
|
||||
# written last. A summary is derived data anchored to the story it covers,
|
||||
# so an Undo or a divergence makes an old one ineligible rather than
|
||||
# leaking it into a story it does not describe (E03, `app/summaries.py`).
|
||||
db = object_session(adventure)
|
||||
summary_row = summaries.current(db, adventure) if db is not None else None
|
||||
summary_text = summary_row.text.strip() if summary_row is not None else ""
|
||||
summary_section = (
|
||||
Section("story_summary", f"Story summary:\n{adventure.story_summary.strip()}")
|
||||
if adventure.story_summary.strip()
|
||||
Section("story_summary", f"Story summary:\n{summary_text}")
|
||||
if summary_text
|
||||
else None
|
||||
)
|
||||
memories_section = None
|
||||
if memory_bank and memory_bank.get("used"):
|
||||
lines_text = "\n".join(f"- {m['text']}" for m in memory_bank["used"])
|
||||
memories_section = Section("used_memories", f"Memories:\n{lines_text}")
|
||||
# M6: an inference must not read as a record. A heuristic memory is
|
||||
# marked in the prompt itself, because the narrator decides what to
|
||||
# treat as established from what it is shown, and an unlabelled guess
|
||||
# sitting beside accepted history is how a guess becomes canon
|
||||
# (`CONTEXT-AND-MEMORY.md` §14). Authoritative state changes still come
|
||||
# only from the M5 event path, whatever a memory says.
|
||||
lines_text = "\n".join(_memory_line(m) for m in memory_bank["used"])
|
||||
memories_section = Section(
|
||||
"used_memories",
|
||||
"Memories from earlier in the story. Lines marked [inferred] are "
|
||||
"interpretation, not established fact — do not treat them as "
|
||||
f"settled truth:\n{lines_text}",
|
||||
)
|
||||
world_state_section = None
|
||||
refusal_note = ""
|
||||
# M5: the authoritative narrative state, as the model is shown it. Read from
|
||||
@@ -370,7 +413,35 @@ def build_context(
|
||||
+ count_tokens(narrative.extract.EMIT_REMINDER)
|
||||
+ count_tokens(refusal_note)
|
||||
)
|
||||
available = max(256, settings.context_token_budget - reserved)
|
||||
|
||||
# ----- M6: the output reserve, and what happens when it does not fit -----
|
||||
#
|
||||
# `context_token_budget` is the whole window the model is given, so the
|
||||
# narrator's reply has to be subtracted from it before any history is
|
||||
# chosen. Until M6 it was not: the builder spent the entire budget on input
|
||||
# and left the reply to fit in whatever the endpoint had left, which is a
|
||||
# truncated turn on a model whose window is the budget
|
||||
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
|
||||
#
|
||||
# The margin covers what is added after this arithmetic — the separators
|
||||
# between sections, and the difference between our tokenizer's count and the
|
||||
# serving model's. It is small and fixed rather than proportional, because
|
||||
# what it absorbs does not scale with the budget.
|
||||
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
|
||||
protected = reserved + output_reserve
|
||||
if protected >= settings.context_token_budget:
|
||||
# Failing here is the point. The alternative — carrying on with a token
|
||||
# or two of history — builds a prompt that is known to overflow, and
|
||||
# the reader gets a truncated reply with no explanation. §32: "fail
|
||||
# gracefully if protected context alone is too large."
|
||||
raise ContextOverflow(
|
||||
f"The protected context needs {protected} tokens "
|
||||
f"({reserved} of prompt plus {output_reserve} reserved for the "
|
||||
f"reply) but the context budget is {settings.context_token_budget}. "
|
||||
"Raise the context budget, lower the maximum reply length, or "
|
||||
"shorten the campaign's canon, instructions and persona."
|
||||
)
|
||||
available = settings.context_token_budget - protected
|
||||
|
||||
# Only the newest actions can reach the prompt, because the code below
|
||||
# either truncates the text to `available` tokens or stops at the budget.
|
||||
@@ -475,12 +546,28 @@ def build_context(
|
||||
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
|
||||
],
|
||||
"prompt": {"system": system_text, "story": story_text},
|
||||
# M6: the numbers the reader needs to answer "how much did each part
|
||||
# cost, and what was left for the reply?" (F04, F05). `available` is
|
||||
# what the history was actually allowed to spend after everything
|
||||
# protected was subtracted.
|
||||
"tokens": {
|
||||
"total": count_tokens(system_text) + count_tokens(story_text),
|
||||
"budget": settings.context_token_budget,
|
||||
"output_reserve": output_reserve,
|
||||
"protected": reserved,
|
||||
"available_for_history": available,
|
||||
"history_spent": spent,
|
||||
},
|
||||
"cards": card_records,
|
||||
"memories": memory_bank,
|
||||
# M6: which summary was used, and which stretch of story it covers, so
|
||||
# "what history did that summary cover?" is answerable from the record
|
||||
# rather than by guessing (F05, F06).
|
||||
"summary": summaries.provenance(summary_row),
|
||||
# M6: whether background derived work is currently failing for this
|
||||
# campaign. A dead memory bank is visible here rather than only in a log
|
||||
# nobody reads (F08).
|
||||
"derived": derived.report(db, adventure.id) if db is not None else [],
|
||||
"history": {
|
||||
"included": len(included_actions),
|
||||
# The count covers the whole story rather than the window fetched
|
||||
|
||||
@@ -0,0 +1,113 @@
|
||||
"""M6: recording whether background derived work succeeded, and why not.
|
||||
|
||||
M2 shipped with the entire memory bank dead and the full test suite green. The
|
||||
summariser and the embedder raised `AttributeError` inside a fire-and-forget
|
||||
task: no user-visible error, no log a player would look at, and no failing test,
|
||||
because every memory test stubbed the provider factories out
|
||||
(`BUILD-MILESTONES.md`, note from M2; `M2-IMPLEMENTATION-REPORT.md` §A.1).
|
||||
|
||||
Two rules follow, and they pull in opposite directions:
|
||||
|
||||
* **Derived work must fail softly.** A memory that could not be written, a
|
||||
summary that could not be generated, an embedding the endpoint refused —
|
||||
none of these may roll back the accepted narration, the accepted state
|
||||
events, the authoritative document, the head, or the transcript. The story
|
||||
turn already happened; the derived work is a commentary on it.
|
||||
* **It must fail visibly.** Soft failure without a record is what M2 shipped.
|
||||
|
||||
So each attempt writes its outcome to one row per (campaign, kind), and that row
|
||||
is readable through the API. This is deliberately not a job framework: it holds
|
||||
what happened last, not a queue. Retrying is just running the pass again, which
|
||||
the ordinary post-turn path already does on the next accepted turn.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from . import models
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# The kinds of derived work. Each is independent: embeddings can be failing
|
||||
# while summaries succeed, and a reader should be able to see exactly that.
|
||||
MEMORY = "memory"
|
||||
SUMMARY = "summary"
|
||||
EMBEDDING = "embedding"
|
||||
KINDS = (MEMORY, SUMMARY, EMBEDDING)
|
||||
|
||||
|
||||
def _row(db: Session, adventure_id: int, kind: str) -> models.DerivedStatus:
|
||||
row = db.execute(
|
||||
select(models.DerivedStatus).where(
|
||||
models.DerivedStatus.adventure_id == adventure_id,
|
||||
models.DerivedStatus.kind == kind,
|
||||
)
|
||||
).scalars().first()
|
||||
if row is None:
|
||||
row = models.DerivedStatus(adventure_id=adventure_id, kind=kind)
|
||||
db.add(row)
|
||||
return row
|
||||
|
||||
|
||||
def succeeded(db: Session, adventure_id: int, kind: str, *, did_work: bool = True) -> None:
|
||||
"""Records a clean run, clearing any standing failure.
|
||||
|
||||
`did_work` separates a pass that produced something from one that found
|
||||
nothing to do (M6 review finding M6-F5). Both are healthy, and neither is a
|
||||
failure, but reporting "ok" for a pass that has never actually run reads as
|
||||
"embeddings are working" when nothing has been embedded. `idle` says the
|
||||
true thing: it ran, and there was nothing pending.
|
||||
"""
|
||||
row = _row(db, adventure_id, kind)
|
||||
row.status = "ok" if did_work else "idle"
|
||||
row.detail = ""
|
||||
row.failures = 0
|
||||
row.last_attempt_at = models.utcnow()
|
||||
if did_work:
|
||||
row.last_success_at = row.last_attempt_at
|
||||
|
||||
|
||||
def failed(db: Session, adventure_id: int, kind: str, exc: BaseException) -> None:
|
||||
"""Records a failed run, keeping the reason where someone can find it.
|
||||
|
||||
The detail is the exception's type and message rather than a traceback: it
|
||||
is shown to a reader in the Insights panel, and `ProviderError: connection
|
||||
refused` is the part that tells them what to do. The traceback goes to the
|
||||
log for a maintainer.
|
||||
"""
|
||||
row = _row(db, adventure_id, kind)
|
||||
row.status = "failed"
|
||||
row.detail = f"{type(exc).__name__}: {exc}"[:2000]
|
||||
row.failures = (row.failures or 0) + 1
|
||||
row.last_attempt_at = models.utcnow()
|
||||
log.exception("derived %s work failed for adventure %s", kind, adventure_id)
|
||||
|
||||
|
||||
def report(db: Session, adventure_id: int) -> list[dict]:
|
||||
"""Every kind's last outcome, for the API and the prompt inspector."""
|
||||
rows = db.execute(
|
||||
select(models.DerivedStatus)
|
||||
.where(models.DerivedStatus.adventure_id == adventure_id)
|
||||
.order_by(models.DerivedStatus.kind)
|
||||
).scalars().all()
|
||||
return [
|
||||
{
|
||||
"kind": row.kind,
|
||||
"status": row.status,
|
||||
"detail": row.detail,
|
||||
"failures": row.failures,
|
||||
"last_attempt_at": row.last_attempt_at.isoformat() if row.last_attempt_at else None,
|
||||
"last_success_at": row.last_success_at.isoformat() if row.last_success_at else None,
|
||||
}
|
||||
for row in rows
|
||||
]
|
||||
|
||||
|
||||
def failing(db: Session, adventure_id: int) -> list[str]:
|
||||
"""The kinds currently in a failed state, for a compact UI badge."""
|
||||
return [entry["kind"] for entry in report(db, adventure_id)
|
||||
if entry["status"] == "failed"]
|
||||
+246
-40
@@ -27,13 +27,14 @@ succeeds.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from array import array
|
||||
from collections import OrderedDict
|
||||
|
||||
from sqlalchemy import func, select, update
|
||||
from sqlalchemy.orm import Session, defer, object_session
|
||||
|
||||
from . import models, tree, vectors
|
||||
from . import derived, models, summaries, tree, vectors
|
||||
from .context import (
|
||||
cursors,
|
||||
history,
|
||||
@@ -46,6 +47,8 @@ from .database import SessionLocal
|
||||
from .providers import OpenAICompatibleProvider, ProviderError
|
||||
from .vectors import cosine # re-exported: the ranking lives here, the maths there
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
MEMORY_INTERVAL = 6 # actions per memory
|
||||
MEMORY_START = 12 # first memory once the adventure reaches this many actions
|
||||
SUMMARY_INTERVAL = 15 # actions between Story Summary updates
|
||||
@@ -146,6 +149,104 @@ _tasks: set[asyncio.Task] = set()
|
||||
# They used to also read an API key, which is gone: Ollama does not use one and
|
||||
# M2 removed cloud providers. `summary_model` and `embedding_model` fall back to
|
||||
# the narrator model when the user has not named a separate one.
|
||||
# M6: the words that mark a memory as an interpretation rather than a record.
|
||||
#
|
||||
# The application owns this classification, not the model
|
||||
# (`CONTEXT-AND-MEMORY.md` §14, §15). The extractor writes prose; this decides
|
||||
# what weight the narrator is told to give it. The list is deliberately short
|
||||
# and readable: a memory that hedges is a reading of the story, not a fact the
|
||||
# story established, and the narrator must not be able to promote it to canon.
|
||||
#
|
||||
# Being wrong in the cautious direction is cheap — a hedged record labelled
|
||||
# heuristic is still retrieved and still useful. Being wrong the other way is
|
||||
# what turns a guess into canon, which is the failure §14 exists to prevent.
|
||||
HEURISTIC_MARKERS = (
|
||||
"seemed", "seems", "appeared to", "appears to", "apparently", "perhaps",
|
||||
"maybe", "might have", "may have", "possibly", "presumably", "suggested that",
|
||||
"suggests that", "implied", "implies", "as if", "likely", "probably",
|
||||
"seemingly", "hinted", "hints that", "suspects", "suspected", "believes",
|
||||
"believed", "wondered whether", "wonders whether",
|
||||
)
|
||||
|
||||
ACCEPTED_STORY = "accepted_story"
|
||||
HEURISTIC = "heuristic"
|
||||
|
||||
|
||||
# M6 corrective (review finding M6-F2). How close two memories have to be before
|
||||
# the second one is treated as saying nothing new.
|
||||
#
|
||||
# The value is measured, not guessed. Against the configured local embedding
|
||||
# model, on a fixture of near-identical "the party walks the muddy road"
|
||||
# memories and a set of genuinely distinct ones:
|
||||
#
|
||||
# redundant pairs cosine 0.938 - 0.996
|
||||
# distinct pairs cosine 0.349 - 0.906
|
||||
#
|
||||
# 0.93 sits in that gap. The same measurement ruled out the more obvious
|
||||
# lexical test: word overlap fires hardest on exactly the pair that must NOT be
|
||||
# merged — "Mara promised to return before dawn" against "Aldric promised to
|
||||
# return before dawn" shares 71% of its words while meaning something else —
|
||||
# and is weakest (27%) on filler that plainly repeats itself. Wording is a poor
|
||||
# proxy for sameness of fact; the embedding is a better one.
|
||||
#
|
||||
# The threshold is model-dependent by nature. A different embedding model may
|
||||
# need a different number, which is why the measurement is written down here
|
||||
# rather than the value alone.
|
||||
REDUNDANT_SIMILARITY = 0.93
|
||||
|
||||
|
||||
def _drop_redundant(candidates, vectors, authority_of, limit):
|
||||
"""Fills `limit` slots, skipping memories that repeat one already chosen.
|
||||
|
||||
Greedy over the ranked list, so the highest-scoring statement of a fact is
|
||||
the one kept and its provenance is the provenance that survives. Two rules
|
||||
keep this from losing information:
|
||||
|
||||
* **Authority is never crossed.** An inference and a record are different
|
||||
kinds of claim even when they read alike, so a `heuristic` memory can
|
||||
never suppress an `accepted_story` one or the reverse.
|
||||
* **The bar is high.** Missing a duplicate costs some budget; dropping a
|
||||
distinct fact costs the narrator something it needed. The threshold is
|
||||
set where the measurement says distinct facts stop appearing.
|
||||
|
||||
Returns `(kept, suppressed)`, the second for the inspector — a reader
|
||||
should be able to see that memories were considered and set aside rather
|
||||
than never retrieved.
|
||||
"""
|
||||
kept: list = []
|
||||
suppressed: list = []
|
||||
for row in candidates:
|
||||
if len(kept) >= limit:
|
||||
break
|
||||
_score, memory_id, _pinned = row
|
||||
vector = vectors.get(memory_id)
|
||||
duplicate_of = None
|
||||
if vector is not None:
|
||||
for _kept_score, kept_id, _ in kept:
|
||||
if authority_of.get(kept_id) != authority_of.get(memory_id):
|
||||
continue
|
||||
other = vectors.get(kept_id)
|
||||
if other is not None and cosine(vector, other) >= REDUNDANT_SIMILARITY:
|
||||
duplicate_of = kept_id
|
||||
break
|
||||
if duplicate_of is None:
|
||||
kept.append(row)
|
||||
else:
|
||||
suppressed.append((memory_id, duplicate_of))
|
||||
return kept, suppressed
|
||||
|
||||
|
||||
def classify_authority(text: str) -> str:
|
||||
"""Returns `accepted_story` or `heuristic` for one memory's text.
|
||||
|
||||
Hedged language is the signal. "Aldric promised Mara he would return before
|
||||
dawn" is something the story established; "Mara seemed uneasy when Captain
|
||||
Vale was mentioned" is an inference about it, and the prompt has to say so.
|
||||
"""
|
||||
lowered = text.lower()
|
||||
return HEURISTIC if any(m in lowered for m in HEURISTIC_MARKERS) else ACCEPTED_STORY
|
||||
|
||||
|
||||
def summary_provider(settings: models.Settings) -> OpenAICompatibleProvider:
|
||||
return OpenAICompatibleProvider(
|
||||
settings.endpoint_url,
|
||||
@@ -473,7 +574,7 @@ async def retrieve_memories(
|
||||
# affordable because memories are sparse, at roughly one per six actions, so
|
||||
# even a heavily forked story returns only tens of small rows.
|
||||
catalogue = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned).where(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.authority).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory),
|
||||
models.Memory.forgotten.is_(False),
|
||||
@@ -495,11 +596,12 @@ async def retrieve_memories(
|
||||
except ProviderError as exc:
|
||||
return {"used": [], "error": str(exc)}
|
||||
|
||||
held = _vectors_for(db, adventure.id, [memory_id for memory_id, _ in catalogue])
|
||||
held = _vectors_for(db, adventure.id, [memory_id for memory_id, _, _ in catalogue])
|
||||
authority_of = {memory_id: authority for memory_id, _, authority in catalogue}
|
||||
scored = sorted(
|
||||
(
|
||||
(cosine(query_vec, held[memory_id]), memory_id, pinned)
|
||||
for memory_id, pinned in catalogue
|
||||
for memory_id, pinned, _ in catalogue
|
||||
if memory_id in held
|
||||
),
|
||||
key=lambda row: row[0],
|
||||
@@ -511,19 +613,31 @@ async def retrieve_memories(
|
||||
top_k = max(1, settings.memory_top_k)
|
||||
used = [row for row in scored if row[2]]
|
||||
remaining = max(0, top_k - len(used))
|
||||
used += [row for row in scored if not row[2]][:remaining]
|
||||
candidates = [row for row in scored if not row[2]]
|
||||
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
|
||||
used += kept
|
||||
used.sort(key=lambda row: row[0], reverse=True)
|
||||
if not used:
|
||||
return {"used": [], "error": None}
|
||||
|
||||
# Fetch the text only now, and only for the `top_k` rows that were chosen.
|
||||
#
|
||||
# M6 adds authority and provenance to this same read rather than to a second
|
||||
# one. The columns are narrow, the row set is `top_k`, and fetching them
|
||||
# here is what keeps "why did the narrator remember this?" answerable
|
||||
# without a query per memory (F06, and the N+1 discipline M5 restored).
|
||||
used_ids = [memory_id for _, memory_id, _ in used]
|
||||
texts = dict(
|
||||
db.execute(
|
||||
select(models.Memory.id, models.Memory.text)
|
||||
.where(models.Memory.id.in_(used_ids))
|
||||
detail = {
|
||||
row.id: row
|
||||
for row in db.execute(
|
||||
select(
|
||||
models.Memory.id, models.Memory.text, models.Memory.authority,
|
||||
models.Memory.branch_id, models.Memory.depth,
|
||||
models.Memory.source_start, models.Memory.source_end,
|
||||
).where(models.Memory.id.in_(used_ids))
|
||||
).all()
|
||||
)
|
||||
}
|
||||
texts = {memory_id: row.text for memory_id, row in detail.items()}
|
||||
|
||||
if update_stats:
|
||||
# Pass `synchronize_session=False` because nothing in this request
|
||||
@@ -539,10 +653,29 @@ async def retrieve_memories(
|
||||
|
||||
return {
|
||||
"used": [
|
||||
{"id": memory_id, "text": texts.get(memory_id, ""),
|
||||
"similarity": round(score, 4), "pinned": pinned}
|
||||
{
|
||||
"id": memory_id,
|
||||
"text": texts.get(memory_id, ""),
|
||||
"similarity": round(score, 4),
|
||||
"pinned": pinned,
|
||||
# M6: what weight this carries, and where it came from.
|
||||
"authority": getattr(detail.get(memory_id), "authority", ACCEPTED_STORY),
|
||||
"source": {
|
||||
"branch_id": getattr(detail.get(memory_id), "branch_id", None),
|
||||
"depth": getattr(detail.get(memory_id), "depth", None),
|
||||
"source_start": getattr(detail.get(memory_id), "source_start", None),
|
||||
"source_end": getattr(detail.get(memory_id), "source_end", None),
|
||||
},
|
||||
}
|
||||
for score, memory_id, pinned in used
|
||||
],
|
||||
"considered": len(catalogue),
|
||||
# M6: how many candidates were set aside as repeating one already
|
||||
# chosen. Visible so that "why is that memory not here?" has an answer.
|
||||
"suppressed": [
|
||||
{"id": memory_id, "duplicate_of": kept_id}
|
||||
for memory_id, kept_id in suppressed
|
||||
],
|
||||
"error": None,
|
||||
}
|
||||
|
||||
@@ -587,17 +720,55 @@ async def run_post_turn(adventure_id: int) -> None:
|
||||
# An anchor past the tip is not an invalid value. `settled_after`
|
||||
# reports that there is nothing to do, and once the story grows past the
|
||||
# anchor the pass resumes where it stopped.
|
||||
# M6. Each kind runs inside its own recorder, so one failing pass
|
||||
# neither hides the others nor takes the turn down with it. The accepted
|
||||
# narration, its state events, the authoritative document and the head
|
||||
# were all committed before this task started; nothing here may undo
|
||||
# them, and nothing here may fail without leaving a record
|
||||
# (`BUILD-MILESTONES.md`, note from M2).
|
||||
if adventure.auto_summarize:
|
||||
await _create_due_memories(adventure, settings, db)
|
||||
await _update_story_summary(adventure, settings, db)
|
||||
await _guarded(db, adventure_id, derived.MEMORY,
|
||||
_create_due_memories(adventure, settings, db))
|
||||
await _guarded(db, adventure_id, derived.SUMMARY,
|
||||
_update_story_summary(adventure, settings, db))
|
||||
if adventure.memory_bank_enabled and settings.embedding_model.strip():
|
||||
await _embed_pending(adventure, settings, db)
|
||||
await _guarded(db, adventure_id, derived.EMBEDDING,
|
||||
_embed_pending(adventure, settings, db))
|
||||
_evict_over_capacity(adventure, settings, db)
|
||||
except BaseException as exc: # noqa: BLE001 - the task boundary
|
||||
# Anything the per-kind guards did not catch: a failure in the shared
|
||||
# setup above, or in eviction. M2's lesson is that the one thing this
|
||||
# may not do is vanish. Re-raising would only feed an unobserved task.
|
||||
try:
|
||||
derived.failed(db, adventure_id, derived.MEMORY, exc)
|
||||
db.commit()
|
||||
except BaseException: # noqa: BLE001 - the recorder must not mask it
|
||||
log.exception("could not record derived-work failure for %s", adventure_id)
|
||||
finally:
|
||||
db.close()
|
||||
_running.discard(adventure_id)
|
||||
|
||||
|
||||
async def _guarded(db: Session, adventure_id: int, kind: str, coro) -> None:
|
||||
"""Runs one derived pass, recording whether it worked.
|
||||
|
||||
The pass keeps whatever it committed before it failed — a memory written
|
||||
two blocks ago stays written — because derived work is additive and
|
||||
partial progress is still progress. What must not survive is an
|
||||
uncommitted, half-written unit of work, so the session is rolled back to
|
||||
the last commit before the failure is recorded.
|
||||
"""
|
||||
try:
|
||||
did_work = await coro
|
||||
except BaseException as exc: # noqa: BLE001 - one kind must not stop another
|
||||
db.rollback()
|
||||
derived.failed(db, adventure_id, kind, exc)
|
||||
db.commit()
|
||||
else:
|
||||
derived.succeeded(db, adventure_id, kind, did_work=bool(did_work))
|
||||
db.commit()
|
||||
|
||||
|
||||
async def summarize_block(
|
||||
adventure: models.Adventure,
|
||||
provider: OpenAICompatibleProvider,
|
||||
@@ -629,37 +800,39 @@ async def summarize_block(
|
||||
|
||||
async def _create_due_memories(
|
||||
adventure: models.Adventure, settings: models.Settings, db: Session
|
||||
) -> None:
|
||||
) -> int:
|
||||
"""Writes the memories that are due. Returns how many it wrote (M6-F5)."""
|
||||
provider = summary_provider(settings)
|
||||
written = 0
|
||||
for _ in range(MAX_MEMORIES_PER_RUN):
|
||||
# Re-read the anchor on every pass. Committing a memory does not change
|
||||
# the story, but this loop is the only code that moves the anchor, so
|
||||
# both numbers must be current.
|
||||
anchor = cursors.MEMORY.depth(db, adventure)
|
||||
if history.count_after(adventure, anchor) < MEMORY_INTERVAL + SETTLE_SLACK:
|
||||
return # No settled block of story sits past the mark. The block
|
||||
return written # No settled block of story sits past the mark. The block
|
||||
# itself is still MEMORY_INTERVAL actions; the slack asks
|
||||
# for story past its end. See `SETTLE_SLACK`.
|
||||
if history.count(adventure) < MEMORY_START:
|
||||
return # The adventure is too short to have started summarizing.
|
||||
return written # The adventure is too short to have started summarizing.
|
||||
# The order of those two checks is deliberate. The usual answer is that
|
||||
# no memory is due, and the first check settles that without measuring
|
||||
# the length of the whole story.
|
||||
block = history.after(adventure, anchor, MEMORY_INTERVAL)
|
||||
if len(block) < MEMORY_INTERVAL:
|
||||
return
|
||||
try:
|
||||
text = await summarize_block(adventure, provider, block)
|
||||
except ProviderError:
|
||||
return # Logged on the debug page. The cursor is unchanged, so the
|
||||
# next turn retries this block.
|
||||
return written
|
||||
# A provider failure is no longer caught here. `_guarded` records it
|
||||
# against this campaign, and the cursor is unchanged either way, so the
|
||||
# next accepted turn retries this same block (M6).
|
||||
text = await summarize_block(adventure, provider, block)
|
||||
if not text:
|
||||
return
|
||||
return written
|
||||
memory = models.Memory(
|
||||
adventure_id=adventure.id,
|
||||
text=text,
|
||||
source_start=block[0].depth,
|
||||
source_end=block[-1].depth,
|
||||
authority=classify_authority(text),
|
||||
)
|
||||
# Attach the memory to the node it summarizes, so that a fork inherits
|
||||
# the memories of the path it forked from and no others. Then move the
|
||||
@@ -670,22 +843,27 @@ async def _create_due_memories(
|
||||
db.add(memory)
|
||||
cursors.MEMORY.anchor_at(adventure, block[-1])
|
||||
db.commit()
|
||||
written += 1
|
||||
return written
|
||||
|
||||
|
||||
async def _update_story_summary(
|
||||
adventure: models.Adventure, settings: models.Settings, db: Session
|
||||
) -> None:
|
||||
) -> bool:
|
||||
"""Rolls the summary forward when enough new story has settled.
|
||||
|
||||
Returns whether it wrote one (M6-F5)."""
|
||||
anchor = cursors.SUMMARY.depth(db, adventure)
|
||||
uncovered = history.count_after(adventure, anchor)
|
||||
if uncovered < SUMMARY_INTERVAL:
|
||||
return
|
||||
return False
|
||||
# Where the summary stands once this run succeeds. Read this before the AI
|
||||
# call rather than after it. The mark records the end of the story as this
|
||||
# pass saw it, and a turn that arrives during the call must not be counted
|
||||
# as read.
|
||||
caught_up = history.newest(adventure)
|
||||
if caught_up is None:
|
||||
return
|
||||
return False
|
||||
|
||||
# Include the memories for the stretch that the summary has not read, which
|
||||
# means every memory attached to a node past the anchor. The marks and the
|
||||
@@ -707,7 +885,25 @@ async def _update_story_summary(
|
||||
block = history.after(adventure, anchor, uncovered)
|
||||
events_text = truncate_to_last_tokens("\n\n".join(a.text for a in block), 2000)
|
||||
|
||||
current = adventure.story_summary.strip()
|
||||
# M6 corrective (review finding M6-F1). The previous summary this one
|
||||
# builds on has to be a summary that is *valid where the story now stands*,
|
||||
# not merely the last one written.
|
||||
#
|
||||
# Seeding from `adventure.story_summary` — a campaign-global column with no
|
||||
# lineage — is what broke E03. After a divergence that column still held the
|
||||
# abandoned line's prose, so the summariser was handed it and asked to
|
||||
# update it. The row it produced was correctly anchored to the new branch
|
||||
# and was therefore *reported* as lineage-safe, while its sentences
|
||||
# described a story the reader had left. The row was anchored; the content
|
||||
# was not.
|
||||
#
|
||||
# `summaries.current` answers the same question the context builder asks —
|
||||
# which summary is eligible at the head — so the input and the output are
|
||||
# now scoped by one rule. Where no eligible summary exists, the new line
|
||||
# starts from nothing, which is the truthful starting point for a story
|
||||
# that has not been summarised yet.
|
||||
eligible = summaries.current(db, adventure)
|
||||
current = eligible.text.strip() if eligible is not None else ""
|
||||
# The summary is built from the memories, so it inherits their framing for
|
||||
# free once they are named and third-person. It still gets the brief of its
|
||||
# own, because the fallback above hands it raw second-person story text
|
||||
@@ -720,22 +916,31 @@ async def _update_story_summary(
|
||||
)
|
||||
if brief:
|
||||
user_prompt = f"{brief}\n\n{user_prompt}"
|
||||
try:
|
||||
text = await summary_provider(settings).complete(
|
||||
SUMMARY_SYSTEM_PROMPT, user_prompt, max_tokens=600
|
||||
)
|
||||
except ProviderError:
|
||||
return
|
||||
text = await summary_provider(settings).complete(
|
||||
SUMMARY_SYSTEM_PROMPT, user_prompt, max_tokens=600
|
||||
)
|
||||
if not text:
|
||||
return
|
||||
adventure.story_summary = text
|
||||
return False
|
||||
# M6: anchored to the story it summarizes rather than written into a single
|
||||
# column. `caught_up` is the last node it covers, so the row is eligible on
|
||||
# exactly the lineages that contain that node, and an Undo or a divergence
|
||||
# makes it ineligible without deleting it (E03, `summaries` module).
|
||||
summaries.record(
|
||||
db, adventure, text,
|
||||
node=caught_up,
|
||||
source_start=anchor + 1 if anchor is not None else None,
|
||||
trigger="interval",
|
||||
model_name=(settings.summary_model or settings.model or ""),
|
||||
)
|
||||
cursors.SUMMARY.anchor_at(adventure, caught_up)
|
||||
db.commit()
|
||||
return True
|
||||
|
||||
|
||||
async def _embed_pending(
|
||||
adventure: models.Adventure, settings: models.Settings, db: Session
|
||||
) -> None:
|
||||
) -> int:
|
||||
"""Embeds memories that have no vector. Returns how many (M6-F5)."""
|
||||
# Use a query rather than walking `adventure.memories`. That walk ran on
|
||||
# every turn and loaded the whole bank's vectors in order to find the few
|
||||
# rows with none.
|
||||
@@ -759,14 +964,15 @@ async def _embed_pending(
|
||||
.all()
|
||||
)
|
||||
if not pending:
|
||||
return
|
||||
return 0
|
||||
try:
|
||||
new = await embedding_provider(settings).embed([m.text for m in pending])
|
||||
except ProviderError:
|
||||
return
|
||||
return 0
|
||||
for memory, vector in zip(pending, new):
|
||||
set_vector(memory, vector)
|
||||
db.commit()
|
||||
return len(pending)
|
||||
|
||||
|
||||
def _evict_over_capacity(
|
||||
|
||||
@@ -402,6 +402,15 @@ MIGRATIONS: list[tuple[int, str | dict[str, str]]] = [
|
||||
# so that every position an existing campaign can be restored to has a
|
||||
# snapshot. See `_backfill_narrative_snapshots`.
|
||||
(88, "-- narrative snapshot backfill (data pass only)"),
|
||||
# M6. `create_all` builds the two new tables — `summaries` and
|
||||
# `derived_status` — as it did `state_events` and `checkpoints`. These are
|
||||
# the columns it cannot add to a table that already exists, plus the data
|
||||
# pass that moves an existing campaign's summary onto the lineage.
|
||||
(89, "ALTER TABLE memories ADD COLUMN authority VARCHAR(20) "
|
||||
"NOT NULL DEFAULT 'accepted_story'"),
|
||||
(90, "CREATE INDEX IF NOT EXISTS ix_summaries_adventure "
|
||||
"ON summaries (adventure_id, depth)"),
|
||||
(91, "-- move the existing story summary onto the lineage (data pass only)"),
|
||||
]
|
||||
|
||||
LATEST_VERSION = max((v for v, _ in MIGRATIONS), default=1)
|
||||
@@ -416,6 +425,7 @@ CURSOR_ANCHOR_VERSION = 56
|
||||
SIBLING_SPLIT_VERSION = 60
|
||||
PARENT_BACKFILL_VERSION = 64
|
||||
NARRATIVE_SNAPSHOT_VERSION = 88
|
||||
SUMMARY_LINEAGE_VERSION = 91
|
||||
|
||||
# An adventure with no actions has no tip. A value of -1 keeps the rule that the
|
||||
# next node goes at `head_depth + 1` true without a special case. This matches
|
||||
@@ -433,6 +443,49 @@ SNAPSHOT_BATCH = 50
|
||||
BACKFILL_BATCH = 200
|
||||
|
||||
|
||||
def _backfill_summary_lineage(conn) -> None:
|
||||
"""Moves each campaign's existing summary onto the lineage that produced it.
|
||||
|
||||
Before M6 the rolling summary lived in `adventures.story_summary` with a
|
||||
separate `(branch_id, depth)` cursor recording how far it had read. The
|
||||
cursor is exactly the coordinate the summary belongs at, so the existing
|
||||
text becomes a `summaries` row anchored there and keeps working — including
|
||||
becoming ineligible after an Undo or a divergence, which is what it could
|
||||
not do before.
|
||||
|
||||
A campaign whose cursor never moved (`summary_cursor_branch_id` NULL) has a
|
||||
summary somebody typed rather than one the pass produced. That anchors at
|
||||
the head instead, which is where a hand-written summary belongs.
|
||||
|
||||
One statement, no row loop. The column is left in place: it is the Plot
|
||||
panel's edit surface and the export bundle's field, and it now mirrors
|
||||
whichever summary is eligible.
|
||||
"""
|
||||
conn.execute(text(
|
||||
"""
|
||||
INSERT INTO summaries (
|
||||
adventure_id, text, branch_id, depth, source_start, source_end,
|
||||
trigger, model_name, created_at
|
||||
)
|
||||
SELECT
|
||||
a.id,
|
||||
a.story_summary,
|
||||
COALESCE(a.summary_cursor_branch_id, a.head_branch_id),
|
||||
CASE WHEN a.summary_cursor_branch_id IS NULL
|
||||
THEN a.head_depth ELSE a.summary_cursor_depth END,
|
||||
NULL,
|
||||
CASE WHEN a.summary_cursor_branch_id IS NULL
|
||||
THEN a.head_depth ELSE a.summary_cursor_depth END,
|
||||
CASE WHEN a.summary_cursor_branch_id IS NULL
|
||||
THEN 'manual' ELSE 'interval' END,
|
||||
'',
|
||||
CURRENT_TIMESTAMP
|
||||
FROM adventures a
|
||||
WHERE TRIM(COALESCE(a.story_summary, '')) <> ''
|
||||
"""
|
||||
))
|
||||
|
||||
|
||||
def _backfill_narrative_snapshots(conn) -> None:
|
||||
"""Gives every pre-M5 action the empty narrative document as its outcome.
|
||||
|
||||
@@ -1198,5 +1251,7 @@ def bootstrap(engine: Engine, through: int = LATEST_VERSION) -> None:
|
||||
_backfill_parents(conn)
|
||||
if version == NARRATIVE_SNAPSHOT_VERSION:
|
||||
_backfill_narrative_snapshots(conn)
|
||||
if version == SUMMARY_LINEAGE_VERSION:
|
||||
_backfill_summary_lineage(conn)
|
||||
current = version
|
||||
_set_version(conn, current)
|
||||
|
||||
+104
-1
@@ -2,7 +2,7 @@ from datetime import datetime, timezone
|
||||
|
||||
from sqlalchemy import (
|
||||
JSON, Boolean, DateTime, Float, ForeignKey, Index, Integer, LargeBinary,
|
||||
String, Text, event,
|
||||
String, Text, UniqueConstraint, event,
|
||||
)
|
||||
from sqlalchemy.orm import Mapped, Session, mapped_column, relationship
|
||||
|
||||
@@ -98,6 +98,19 @@ class Adventure(Base):
|
||||
memory: Mapped[str] = mapped_column(Text, default="")
|
||||
authors_note: Mapped[str] = mapped_column(Text, default="")
|
||||
ai_instructions: Mapped[str] = mapped_column(Text, default="")
|
||||
# A convenience mirror of whichever summary is eligible at the current
|
||||
# position, and **never** an input to anything authoritative (M6 corrective,
|
||||
# review finding M6-F1).
|
||||
#
|
||||
# It exists because the Plot panel lets a reader read and edit the summary
|
||||
# and the export bundle carries it. It is not a store: `summaries` rows are,
|
||||
# and `summaries.current` decides which one the story is entitled to. This
|
||||
# column has no lineage of its own, so anything that reads it as truth
|
||||
# inherits whatever was written last, on whatever line — which is exactly
|
||||
# how abandoned prose reached an active prompt before the correction.
|
||||
#
|
||||
# Kept in step by `summaries.record` when one is written and by
|
||||
# `attempts.restore_state` when the head moves.
|
||||
story_summary: Mapped[str] = mapped_column(Text, default="")
|
||||
# Phase 18: who the player is playing as. The AI never writes these — they
|
||||
# are user-only, which is what lets them sit in the cached system block
|
||||
@@ -198,6 +211,17 @@ class Adventure(Base):
|
||||
cascade="all, delete-orphan",
|
||||
order_by="Memory.id",
|
||||
)
|
||||
# M6: the lineage-anchored generated summaries, newest last.
|
||||
summaries: Mapped[list["Summary"]] = relationship(
|
||||
back_populates="adventure",
|
||||
cascade="all, delete-orphan",
|
||||
order_by="Summary.id",
|
||||
)
|
||||
derived_status: Mapped[list["DerivedStatus"]] = relationship(
|
||||
back_populates="adventure",
|
||||
cascade="all, delete-orphan",
|
||||
order_by="DerivedStatus.id",
|
||||
)
|
||||
|
||||
|
||||
class Branch(Base):
|
||||
@@ -467,6 +491,14 @@ class Memory(Base):
|
||||
# current. Readers need only the yes-or-no answer, and fetching six
|
||||
# kilobytes of vector to get it is too expensive.
|
||||
embedded: Mapped[bool] = mapped_column(Boolean, default=False)
|
||||
# M6: how much weight the narrator should give this memory
|
||||
# (`CONTEXT-AND-MEMORY.md` §14). `accepted_story` is something the story
|
||||
# actually established; `heuristic` is an interpretation of it. The
|
||||
# application owns this classification — the extractor may hint, but
|
||||
# `memorybank.classify_authority` decides — so a guess can never become
|
||||
# canon merely by being written down. Authoritative state changes still go
|
||||
# only through the M5 event path (ADR 013).
|
||||
authority: Mapped[str] = mapped_column(String(20), default="accepted_story")
|
||||
pinned: Mapped[bool] = mapped_column(Boolean, default=False)
|
||||
forgotten: Mapped[bool] = mapped_column(Boolean, default=False) # evicted, kept for UI
|
||||
use_count: Mapped[int] = mapped_column(Integer, default=0)
|
||||
@@ -476,6 +508,77 @@ class Memory(Base):
|
||||
adventure: Mapped[Adventure] = relationship(back_populates="memories")
|
||||
|
||||
|
||||
class Summary(Base):
|
||||
"""M6: one generated rolling summary, anchored to the story it summarizes.
|
||||
|
||||
The inherited design kept the summary in a single `adventures.story_summary`
|
||||
column with a lineage cursor recording how far it had read. The cursor was
|
||||
lineage-aware; the prose was not. After an Undo and a divergence the column
|
||||
still held sentences describing the abandoned line, and the context builder
|
||||
injected it unconditionally — the leak `STORY-BRANCH-SEMANTICS.md` §32 and
|
||||
acceptance test E03 forbid.
|
||||
|
||||
A summary is therefore a row on a path, exactly as a `Memory` is, and it is
|
||||
filtered through the same `lineage.Path.clause` chokepoint. `branch_id` and
|
||||
`depth` are the coordinate it was written at; `source_start`/`source_end`
|
||||
are the stretch of story it covers. A summary whose coordinate is not on the
|
||||
active capped lineage is not eligible, and is never deleted for it — the
|
||||
abandoned line keeps its own derived data (§11).
|
||||
"""
|
||||
|
||||
__tablename__ = "summaries"
|
||||
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
|
||||
text: Mapped[str] = mapped_column(Text, default="")
|
||||
# The coordinate this summary was written at: the last node it covers.
|
||||
branch_id: Mapped[int | None] = mapped_column(
|
||||
ForeignKey("branches.id", ondelete="CASCADE"), nullable=True
|
||||
)
|
||||
depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
|
||||
# The stretch of story it summarizes, as depths on `branch_id`.
|
||||
source_start: Mapped[int | None] = mapped_column(Integer, nullable=True)
|
||||
source_end: Mapped[int | None] = mapped_column(Integer, nullable=True)
|
||||
# Why it was generated: "interval" for the automatic pass, "manual" when the
|
||||
# reader wrote or edited it themselves.
|
||||
trigger: Mapped[str] = mapped_column(String(20), default="interval")
|
||||
model_name: Mapped[str] = mapped_column(String(200), default="")
|
||||
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
|
||||
|
||||
adventure: Mapped[Adventure] = relationship(back_populates="summaries")
|
||||
|
||||
|
||||
class DerivedStatus(Base):
|
||||
"""M6: the outcome of one kind of background derived work, per campaign.
|
||||
|
||||
M2 shipped with the whole memory bank dead and the suite green: the
|
||||
summariser and the embedder raised inside a fire-and-forget task, and
|
||||
nothing recorded it (`BUILD-MILESTONES.md`, note from M2). Derived work is
|
||||
allowed to fail — the accepted turn, the state and the head must all
|
||||
survive it — but it is not allowed to fail *invisibly*.
|
||||
|
||||
One row per (adventure, kind), rewritten in place. This is deliberately not
|
||||
a job queue: it records what happened last, so a reader can see that
|
||||
memories stopped being written and why, and so a maintainer can retry.
|
||||
"""
|
||||
|
||||
__tablename__ = "derived_status"
|
||||
__table_args__ = (UniqueConstraint("adventure_id", "kind", name="uq_derived_kind"),)
|
||||
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
|
||||
# "memory", "summary" or "embedding".
|
||||
kind: Mapped[str] = mapped_column(String(20))
|
||||
# "ok" (did work), "idle" (ran, nothing pending) or "failed".
|
||||
status: Mapped[str] = mapped_column(String(20), default="ok")
|
||||
detail: Mapped[str] = mapped_column(Text, default="")
|
||||
failures: Mapped[int] = mapped_column(Integer, default=0)
|
||||
last_attempt_at: Mapped[datetime | None] = mapped_column(DateTime, nullable=True)
|
||||
last_success_at: Mapped[datetime | None] = mapped_column(DateTime, nullable=True)
|
||||
|
||||
adventure: Mapped[Adventure] = relationship(back_populates="derived_status")
|
||||
|
||||
|
||||
class StoryCard(Base):
|
||||
"""Owned by either a scenario or an adventure (exactly one set)."""
|
||||
|
||||
|
||||
@@ -10,7 +10,8 @@ from sqlalchemy.orm import Session
|
||||
from sqlalchemy.orm.attributes import set_committed_value
|
||||
|
||||
from ... import (
|
||||
attempts, head, images, limits, memorybank, models, schemas, tree, worldstate,
|
||||
attempts, head, images, limits, memorybank, models, schemas, summaries, tree,
|
||||
worldstate,
|
||||
)
|
||||
from ...database import get_db
|
||||
|
||||
@@ -294,8 +295,18 @@ def update_adventure(
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
for field, value in payload.model_dump(exclude_unset=True).items():
|
||||
fields = payload.model_dump(exclude_unset=True)
|
||||
for field, value in fields.items():
|
||||
setattr(adventure, field, value)
|
||||
# M6: a summary the reader typed is still a summary, so it is anchored to
|
||||
# the position they typed it at rather than left in a column with no
|
||||
# lineage. Otherwise a hand-written summary would survive an Undo and a
|
||||
# divergence that its generated equivalent correctly does not (E03).
|
||||
if "story_summary" in fields:
|
||||
typed = (fields["story_summary"] or "").strip()
|
||||
held = summaries.current(db, adventure)
|
||||
if typed and (held is None or held.text.strip() != typed):
|
||||
summaries.record(db, adventure, typed, trigger="manual")
|
||||
db.commit()
|
||||
return adventure
|
||||
|
||||
|
||||
@@ -7,8 +7,8 @@ returns the prompt a turn was actually generated from. Neither writes anything.
|
||||
from fastapi import Depends, HTTPException
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from ... import memorybank, models
|
||||
from ...context import build_context
|
||||
from ... import derived, memorybank, models, summaries
|
||||
from ...context import ContextOverflow, build_context
|
||||
from ...database import get_db
|
||||
from ..settings import get_settings
|
||||
|
||||
@@ -24,10 +24,55 @@ async def dry_run_context(
|
||||
"""Returns what the app would send to the AI if the player continued now."""
|
||||
settings = get_settings(db, user)
|
||||
memories = await memorybank.retrieve_memories(adventure, settings, update_stats=False)
|
||||
_, _, report = build_context(adventure, settings, memories)
|
||||
try:
|
||||
_, _, report = build_context(adventure, settings, memories)
|
||||
except ContextOverflow as exc:
|
||||
# M6: a dry run of a prompt that cannot be built is still an answer, and
|
||||
# a more useful one than a 500. The reader opened this panel to find out
|
||||
# what would be sent; "nothing, because the protected context does not
|
||||
# fit, and here is by how much" is exactly that.
|
||||
raise HTTPException(422, str(exc)) from exc
|
||||
return report
|
||||
|
||||
|
||||
@router.get("/{adventure_id}/derived")
|
||||
def derived_status(
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
"""M6: whether background memory, summary and embedding work is healthy.
|
||||
|
||||
The surface that makes a dead memory bank findable. M2 shipped with the
|
||||
whole bank failing inside a fire-and-forget task and nothing anywhere said
|
||||
so — not the UI, not a log a player would read, not a failing test
|
||||
(`BUILD-MILESTONES.md`, note from M2). This endpoint is where that now
|
||||
shows.
|
||||
"""
|
||||
# Resolved once, not once per row: which summary the current head is
|
||||
# entitled to. Asking inside the comprehension would be one query per
|
||||
# summary, which is the shape M5 spent a finding removing.
|
||||
eligible = summaries.current(db, adventure)
|
||||
eligible_id = eligible.id if eligible is not None else None
|
||||
status = derived.report(db, adventure.id)
|
||||
return {
|
||||
"status": status,
|
||||
"failing": [row["kind"] for row in status if row["status"] == "failed"],
|
||||
"summaries": [
|
||||
{
|
||||
"id": row.id,
|
||||
"branch_id": row.branch_id,
|
||||
"depth": row.depth,
|
||||
"trigger": row.trigger,
|
||||
"model": row.model_name,
|
||||
"eligible": row.id == eligible_id,
|
||||
"created_at": row.created_at.isoformat() if row.created_at else None,
|
||||
"preview": row.text[:200],
|
||||
}
|
||||
for row in summaries.all_for(db, adventure)
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
@router.get("/{adventure_id}/actions/{action_id}/context")
|
||||
def action_context(
|
||||
adventure_id: int,
|
||||
|
||||
@@ -16,7 +16,7 @@ from ... import (
|
||||
attempts, head, limits, memorybank, models, narrative, schemas, tree,
|
||||
worldstate,
|
||||
)
|
||||
from ...context import build_context, cursors
|
||||
from ...context import ContextOverflow, build_context, cursors
|
||||
from ...database import get_db
|
||||
from ...providers import OpenAICompatibleProvider, PromptParts, ProviderError
|
||||
from ...sse import SSE_HEADERS, sse, turn_error
|
||||
@@ -164,9 +164,18 @@ async def _generate_turn(
|
||||
memories = await memorybank.retrieve_memories(
|
||||
adventure, settings, update_stats=True, exclude_action_id=replacing_id
|
||||
)
|
||||
system_text, story_text, snapshot = build_context(
|
||||
adventure, settings, memories, exclude_action_id=replacing_id
|
||||
)
|
||||
try:
|
||||
system_text, story_text, snapshot = build_context(
|
||||
adventure, settings, memories, exclude_action_id=replacing_id
|
||||
)
|
||||
except ContextOverflow as exc:
|
||||
# M6: the protected context does not fit in the configured budget, so
|
||||
# there is no prompt to send. This is a settings problem the reader can
|
||||
# fix, and the message says how — reporting it as a failed turn keeps
|
||||
# the story intact and tells them what to change, where building the
|
||||
# prompt anyway would return a silently truncated reply.
|
||||
yield turn_error(str(exc))
|
||||
return
|
||||
|
||||
parts = PromptParts(system=system_text, story=story_text)
|
||||
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
"""M6: the rolling story summary, anchored to the story it summarizes.
|
||||
|
||||
A summary is compressed derived history. It is never the source of truth — the
|
||||
retained transcript is (`CONTEXT-AND-MEMORY.md` §9) — and it is never allowed to
|
||||
describe a story the reader is not on.
|
||||
|
||||
The inherited design kept one `adventures.story_summary` column and a lineage
|
||||
cursor recording how far the summariser had read. The cursor was lineage-aware;
|
||||
the prose it produced was not. After an Undo and a divergence the column still
|
||||
held sentences about the abandoned line, and the context builder injected it
|
||||
with no eligibility check at all — acceptance test E03, and measured failing
|
||||
against the M5 baseline before this module existed.
|
||||
|
||||
The fix is not a new lineage system. A summary is a row with a coordinate, the
|
||||
way a `Memory` already is, and it is filtered through the same
|
||||
`lineage.Path.clause` chokepoint every other read of the story goes through. So:
|
||||
|
||||
eligible == its coordinate is on the active, head-capped lineage
|
||||
|
||||
which gives the four behaviours the milestone asks for, without a rule of its
|
||||
own for any of them:
|
||||
|
||||
A -> B -> C -> D, summary covers A..C, head at D eligible
|
||||
Undo to B not eligible
|
||||
Redo to D eligible again
|
||||
diverge from B onto X -> Y not eligible
|
||||
|
||||
Nothing is deleted when a line is abandoned. The abandoned line keeps its own
|
||||
summaries, and they become eligible again if the reader returns to it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from . import models
|
||||
from .context import lineage
|
||||
|
||||
|
||||
def record(
|
||||
db: Session,
|
||||
adventure: models.Adventure,
|
||||
text: str,
|
||||
*,
|
||||
node: models.Action | None = None,
|
||||
source_start: int | None = None,
|
||||
trigger: str = "interval",
|
||||
model_name: str = "",
|
||||
) -> models.Summary:
|
||||
"""Stores one summary at the coordinate the story has reached.
|
||||
|
||||
`node` is the last action the summary covers, which is where the row is
|
||||
anchored. Without one the summary anchors at the head, which is what a
|
||||
summary the reader typed themselves covers.
|
||||
"""
|
||||
branch_id = adventure.head_branch_id
|
||||
depth = adventure.head_depth
|
||||
if node is not None and node.depth is not None:
|
||||
branch_id, depth = node.branch_id, node.depth
|
||||
row = models.Summary(
|
||||
adventure_id=adventure.id,
|
||||
text=text.strip(),
|
||||
branch_id=branch_id,
|
||||
depth=depth,
|
||||
source_start=source_start,
|
||||
source_end=depth,
|
||||
trigger=trigger,
|
||||
model_name=model_name,
|
||||
)
|
||||
db.add(row)
|
||||
mirror(adventure, row.text)
|
||||
return row
|
||||
|
||||
|
||||
def mirror(adventure: models.Adventure, text: str) -> None:
|
||||
"""Points `adventures.story_summary` at the summary now in force.
|
||||
|
||||
That column is a reader-facing convenience — the Plot panel edits it, the
|
||||
export bundle carries it — and nothing authoritative may read it. It has no
|
||||
lineage, so it holds whatever was written last on whatever line, and the M6
|
||||
review found the summariser seeding itself from exactly that: after a
|
||||
divergence it was handed the abandoned line's prose and asked to update it
|
||||
(finding M6-F1).
|
||||
|
||||
The fix was to seed generation from `current()` instead. This function keeps
|
||||
the column honest as well, so what a reader sees in the Plot panel and what
|
||||
an export carries is the summary the narrator is actually being given.
|
||||
"""
|
||||
adventure.story_summary = text or ""
|
||||
|
||||
|
||||
def refresh_mirror(db: Session, adventure: models.Adventure) -> None:
|
||||
"""Re-points the mirror after the head has moved.
|
||||
|
||||
Called from `attempts.restore_state`, which every Undo, Redo, take switch
|
||||
and Save Point restore goes through. Without it the column would keep
|
||||
showing a summary the story has moved away from.
|
||||
"""
|
||||
row = current(db, adventure)
|
||||
mirror(adventure, row.text if row is not None else "")
|
||||
|
||||
|
||||
def current(db: Session, adventure: models.Adventure) -> models.Summary | None:
|
||||
"""The newest summary eligible for the position being read, or None.
|
||||
|
||||
Eligibility is the capped lineage clause and nothing else. Ordering by
|
||||
depth then id takes the newest summary on the path, so a fresher summary
|
||||
written on a shallower branch does not outrank the deep one it was
|
||||
superseded by.
|
||||
"""
|
||||
return db.execute(
|
||||
select(models.Summary)
|
||||
.where(
|
||||
models.Summary.adventure_id == adventure.id,
|
||||
lineage.path_of(db, adventure).clause(models.Summary),
|
||||
)
|
||||
.order_by(models.Summary.depth.desc(), models.Summary.id.desc())
|
||||
.limit(1)
|
||||
).scalars().first()
|
||||
|
||||
|
||||
def text_for_prompt(db: Session, adventure: models.Adventure) -> str:
|
||||
"""The summary the narrator should be shown, or an empty string."""
|
||||
row = current(db, adventure)
|
||||
return row.text if row is not None and row.text.strip() else ""
|
||||
|
||||
|
||||
def provenance(row: models.Summary | None) -> dict | None:
|
||||
"""What the inspector shows about where a summary came from."""
|
||||
if row is None:
|
||||
return None
|
||||
return {
|
||||
"id": row.id,
|
||||
"branch_id": row.branch_id,
|
||||
"depth": row.depth,
|
||||
"source_start": row.source_start,
|
||||
"source_end": row.source_end,
|
||||
"trigger": row.trigger,
|
||||
"model": row.model_name,
|
||||
"created_at": row.created_at.isoformat() if row.created_at else None,
|
||||
}
|
||||
|
||||
|
||||
def all_for(db: Session, adventure: models.Adventure) -> list[models.Summary]:
|
||||
"""Every stored summary, eligible or not, newest first.
|
||||
|
||||
Abandoned summaries are retained rather than deleted, so this is how a
|
||||
reader or a maintainer sees that they still exist.
|
||||
"""
|
||||
return list(db.execute(
|
||||
select(models.Summary)
|
||||
.where(models.Summary.adventure_id == adventure.id)
|
||||
.order_by(models.Summary.id.desc())
|
||||
).scalars().all())
|
||||
Reference in New Issue
Block a user