M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
co-authored by
Claude Opus 5
parent
b7005e6fdd
commit
a6e9c7a32b
@@ -21,8 +21,9 @@ block and the live sections in `build_context`.
|
||||
from dataclasses import dataclass
|
||||
|
||||
import tiktoken
|
||||
from sqlalchemy.orm import object_session
|
||||
|
||||
from .. import models, narrative, worldstate
|
||||
from .. import derived, models, narrative, summaries, worldstate
|
||||
from . import encoding, history
|
||||
|
||||
AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
|
||||
@@ -63,6 +64,24 @@ MAX_LENGTH_FLOOR_WORDS = 300
|
||||
# Built from the table vendored in `encoding.py`, not fetched: the upstream
|
||||
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
|
||||
# is called on every turn.
|
||||
# M6: added to the configured reply budget when reserving output space. It
|
||||
# absorbs the section separators added after budgeting and the drift between
|
||||
# this tokenizer and the serving model's. Fixed rather than proportional: what
|
||||
# it covers does not grow with the size of the budget.
|
||||
OUTPUT_SAFETY_MARGIN = 64
|
||||
|
||||
|
||||
class ContextOverflow(RuntimeError):
|
||||
"""Raised when protected context alone cannot fit in the token budget.
|
||||
|
||||
Protected means the narrator rules, the campaign canon, the authoritative
|
||||
narrative state, the reader's own input, and the reserve for the reply
|
||||
(`CONTEXT-AND-MEMORY.md` §30). None of those may be dropped to make room for
|
||||
old prose, so when they do not fit there is no prompt to build and saying so
|
||||
is the only honest answer.
|
||||
"""
|
||||
|
||||
|
||||
def _encoding() -> tiktoken.Encoding:
|
||||
return encoding.get_encoding()
|
||||
|
||||
@@ -187,6 +206,12 @@ def _history_text(action: models.Action) -> str:
|
||||
return action.text
|
||||
|
||||
|
||||
def _memory_line(memory: dict) -> str:
|
||||
"""One retrieved memory, marked with its authority (M6)."""
|
||||
mark = " [inferred]" if memory.get("authority") == "heuristic" else ""
|
||||
return f"-{mark} {memory['text']}"
|
||||
|
||||
|
||||
def _canon_section(adventure: models.Adventure) -> str:
|
||||
"""The campaign's own rules, rendered for the system block.
|
||||
|
||||
@@ -317,15 +342,33 @@ def build_context(
|
||||
# memories change on most turns, and the stat values change on nearly every
|
||||
# turn. `world_lore` is added below, because the history window determines
|
||||
# which cards trigger and that window is not known yet.
|
||||
# M6: the summary the *current lineage* is entitled to, not whatever was
|
||||
# written last. A summary is derived data anchored to the story it covers,
|
||||
# so an Undo or a divergence makes an old one ineligible rather than
|
||||
# leaking it into a story it does not describe (E03, `app/summaries.py`).
|
||||
db = object_session(adventure)
|
||||
summary_row = summaries.current(db, adventure) if db is not None else None
|
||||
summary_text = summary_row.text.strip() if summary_row is not None else ""
|
||||
summary_section = (
|
||||
Section("story_summary", f"Story summary:\n{adventure.story_summary.strip()}")
|
||||
if adventure.story_summary.strip()
|
||||
Section("story_summary", f"Story summary:\n{summary_text}")
|
||||
if summary_text
|
||||
else None
|
||||
)
|
||||
memories_section = None
|
||||
if memory_bank and memory_bank.get("used"):
|
||||
lines_text = "\n".join(f"- {m['text']}" for m in memory_bank["used"])
|
||||
memories_section = Section("used_memories", f"Memories:\n{lines_text}")
|
||||
# M6: an inference must not read as a record. A heuristic memory is
|
||||
# marked in the prompt itself, because the narrator decides what to
|
||||
# treat as established from what it is shown, and an unlabelled guess
|
||||
# sitting beside accepted history is how a guess becomes canon
|
||||
# (`CONTEXT-AND-MEMORY.md` §14). Authoritative state changes still come
|
||||
# only from the M5 event path, whatever a memory says.
|
||||
lines_text = "\n".join(_memory_line(m) for m in memory_bank["used"])
|
||||
memories_section = Section(
|
||||
"used_memories",
|
||||
"Memories from earlier in the story. Lines marked [inferred] are "
|
||||
"interpretation, not established fact — do not treat them as "
|
||||
f"settled truth:\n{lines_text}",
|
||||
)
|
||||
world_state_section = None
|
||||
refusal_note = ""
|
||||
# M5: the authoritative narrative state, as the model is shown it. Read from
|
||||
@@ -370,7 +413,35 @@ def build_context(
|
||||
+ count_tokens(narrative.extract.EMIT_REMINDER)
|
||||
+ count_tokens(refusal_note)
|
||||
)
|
||||
available = max(256, settings.context_token_budget - reserved)
|
||||
|
||||
# ----- M6: the output reserve, and what happens when it does not fit -----
|
||||
#
|
||||
# `context_token_budget` is the whole window the model is given, so the
|
||||
# narrator's reply has to be subtracted from it before any history is
|
||||
# chosen. Until M6 it was not: the builder spent the entire budget on input
|
||||
# and left the reply to fit in whatever the endpoint had left, which is a
|
||||
# truncated turn on a model whose window is the budget
|
||||
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
|
||||
#
|
||||
# The margin covers what is added after this arithmetic — the separators
|
||||
# between sections, and the difference between our tokenizer's count and the
|
||||
# serving model's. It is small and fixed rather than proportional, because
|
||||
# what it absorbs does not scale with the budget.
|
||||
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
|
||||
protected = reserved + output_reserve
|
||||
if protected >= settings.context_token_budget:
|
||||
# Failing here is the point. The alternative — carrying on with a token
|
||||
# or two of history — builds a prompt that is known to overflow, and
|
||||
# the reader gets a truncated reply with no explanation. §32: "fail
|
||||
# gracefully if protected context alone is too large."
|
||||
raise ContextOverflow(
|
||||
f"The protected context needs {protected} tokens "
|
||||
f"({reserved} of prompt plus {output_reserve} reserved for the "
|
||||
f"reply) but the context budget is {settings.context_token_budget}. "
|
||||
"Raise the context budget, lower the maximum reply length, or "
|
||||
"shorten the campaign's canon, instructions and persona."
|
||||
)
|
||||
available = settings.context_token_budget - protected
|
||||
|
||||
# Only the newest actions can reach the prompt, because the code below
|
||||
# either truncates the text to `available` tokens or stops at the budget.
|
||||
@@ -475,12 +546,28 @@ def build_context(
|
||||
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
|
||||
],
|
||||
"prompt": {"system": system_text, "story": story_text},
|
||||
# M6: the numbers the reader needs to answer "how much did each part
|
||||
# cost, and what was left for the reply?" (F04, F05). `available` is
|
||||
# what the history was actually allowed to spend after everything
|
||||
# protected was subtracted.
|
||||
"tokens": {
|
||||
"total": count_tokens(system_text) + count_tokens(story_text),
|
||||
"budget": settings.context_token_budget,
|
||||
"output_reserve": output_reserve,
|
||||
"protected": reserved,
|
||||
"available_for_history": available,
|
||||
"history_spent": spent,
|
||||
},
|
||||
"cards": card_records,
|
||||
"memories": memory_bank,
|
||||
# M6: which summary was used, and which stretch of story it covers, so
|
||||
# "what history did that summary cover?" is answerable from the record
|
||||
# rather than by guessing (F05, F06).
|
||||
"summary": summaries.provenance(summary_row),
|
||||
# M6: whether background derived work is currently failing for this
|
||||
# campaign. A dead memory bank is visible here rather than only in a log
|
||||
# nobody reads (F08).
|
||||
"derived": derived.report(db, adventure.id) if db is not None else [],
|
||||
"history": {
|
||||
"included": len(included_actions),
|
||||
# The count covers the whole story rather than the window fetched
|
||||
|
||||
Reference in New Issue
Block a user