M6: branch-safe context, summaries and long-term story memory

Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-06 03:00:33 -04:00
co-authored by Claude Opus 5
parent b7005e6fdd
commit a6e9c7a32b
32 changed files with 4040 additions and 84 deletions
+2 -1
View File
@@ -1,5 +1,6 @@
from . import history
from .builder import (
ContextOverflow,
build_context,
count_tokens,
match_cards,
@@ -9,7 +10,7 @@ from .builder import (
from .history import story_actions
__all__ = [
"build_context",
"ContextOverflow", "build_context",
"count_tokens",
"history",
"match_cards",
+93 -6
View File
@@ -21,8 +21,9 @@ block and the live sections in `build_context`.
from dataclasses import dataclass
import tiktoken
from sqlalchemy.orm import object_session
from .. import models, narrative, worldstate
from .. import derived, models, narrative, summaries, worldstate
from . import encoding, history
AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
@@ -63,6 +64,24 @@ MAX_LENGTH_FLOOR_WORDS = 300
# Built from the table vendored in `encoding.py`, not fetched: the upstream
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
# is called on every turn.
# M6: added to the configured reply budget when reserving output space. It
# absorbs the section separators added after budgeting and the drift between
# this tokenizer and the serving model's. Fixed rather than proportional: what
# it covers does not grow with the size of the budget.
OUTPUT_SAFETY_MARGIN = 64
class ContextOverflow(RuntimeError):
"""Raised when protected context alone cannot fit in the token budget.
Protected means the narrator rules, the campaign canon, the authoritative
narrative state, the reader's own input, and the reserve for the reply
(`CONTEXT-AND-MEMORY.md` §30). None of those may be dropped to make room for
old prose, so when they do not fit there is no prompt to build and saying so
is the only honest answer.
"""
def _encoding() -> tiktoken.Encoding:
return encoding.get_encoding()
@@ -187,6 +206,12 @@ def _history_text(action: models.Action) -> str:
return action.text
def _memory_line(memory: dict) -> str:
"""One retrieved memory, marked with its authority (M6)."""
mark = " [inferred]" if memory.get("authority") == "heuristic" else ""
return f"-{mark} {memory['text']}"
def _canon_section(adventure: models.Adventure) -> str:
"""The campaign's own rules, rendered for the system block.
@@ -317,15 +342,33 @@ def build_context(
# memories change on most turns, and the stat values change on nearly every
# turn. `world_lore` is added below, because the history window determines
# which cards trigger and that window is not known yet.
# M6: the summary the *current lineage* is entitled to, not whatever was
# written last. A summary is derived data anchored to the story it covers,
# so an Undo or a divergence makes an old one ineligible rather than
# leaking it into a story it does not describe (E03, `app/summaries.py`).
db = object_session(adventure)
summary_row = summaries.current(db, adventure) if db is not None else None
summary_text = summary_row.text.strip() if summary_row is not None else ""
summary_section = (
Section("story_summary", f"Story summary:\n{adventure.story_summary.strip()}")
if adventure.story_summary.strip()
Section("story_summary", f"Story summary:\n{summary_text}")
if summary_text
else None
)
memories_section = None
if memory_bank and memory_bank.get("used"):
lines_text = "\n".join(f"- {m['text']}" for m in memory_bank["used"])
memories_section = Section("used_memories", f"Memories:\n{lines_text}")
# M6: an inference must not read as a record. A heuristic memory is
# marked in the prompt itself, because the narrator decides what to
# treat as established from what it is shown, and an unlabelled guess
# sitting beside accepted history is how a guess becomes canon
# (`CONTEXT-AND-MEMORY.md` §14). Authoritative state changes still come
# only from the M5 event path, whatever a memory says.
lines_text = "\n".join(_memory_line(m) for m in memory_bank["used"])
memories_section = Section(
"used_memories",
"Memories from earlier in the story. Lines marked [inferred] are "
"interpretation, not established fact — do not treat them as "
f"settled truth:\n{lines_text}",
)
world_state_section = None
refusal_note = ""
# M5: the authoritative narrative state, as the model is shown it. Read from
@@ -370,7 +413,35 @@ def build_context(
+ count_tokens(narrative.extract.EMIT_REMINDER)
+ count_tokens(refusal_note)
)
available = max(256, settings.context_token_budget - reserved)
# ----- M6: the output reserve, and what happens when it does not fit -----
#
# `context_token_budget` is the whole window the model is given, so the
# narrator's reply has to be subtracted from it before any history is
# chosen. Until M6 it was not: the builder spent the entire budget on input
# and left the reply to fit in whatever the endpoint had left, which is a
# truncated turn on a model whose window is the budget
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
#
# The margin covers what is added after this arithmetic — the separators
# between sections, and the difference between our tokenizer's count and the
# serving model's. It is small and fixed rather than proportional, because
# what it absorbs does not scale with the budget.
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
protected = reserved + output_reserve
if protected >= settings.context_token_budget:
# Failing here is the point. The alternative — carrying on with a token
# or two of history — builds a prompt that is known to overflow, and
# the reader gets a truncated reply with no explanation. §32: "fail
# gracefully if protected context alone is too large."
raise ContextOverflow(
f"The protected context needs {protected} tokens "
f"({reserved} of prompt plus {output_reserve} reserved for the "
f"reply) but the context budget is {settings.context_token_budget}. "
"Raise the context budget, lower the maximum reply length, or "
"shorten the campaign's canon, instructions and persona."
)
available = settings.context_token_budget - protected
# Only the newest actions can reach the prompt, because the code below
# either truncates the text to `available` tokens or stops at the budget.
@@ -475,12 +546,28 @@ def build_context(
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
],
"prompt": {"system": system_text, "story": story_text},
# M6: the numbers the reader needs to answer "how much did each part
# cost, and what was left for the reply?" (F04, F05). `available` is
# what the history was actually allowed to spend after everything
# protected was subtracted.
"tokens": {
"total": count_tokens(system_text) + count_tokens(story_text),
"budget": settings.context_token_budget,
"output_reserve": output_reserve,
"protected": reserved,
"available_for_history": available,
"history_spent": spent,
},
"cards": card_records,
"memories": memory_bank,
# M6: which summary was used, and which stretch of story it covers, so
# "what history did that summary cover?" is answerable from the record
# rather than by guessing (F05, F06).
"summary": summaries.provenance(summary_row),
# M6: whether background derived work is currently failing for this
# campaign. A dead memory bank is visible here rather than only in a log
# nobody reads (F08).
"derived": derived.report(db, adventure.id) if db is not None else [],
"history": {
"included": len(included_actions),
# The count covers the whole story rather than the window fetched