Files
interactive-story/backend/app/summaries.py
T
JesseMarkowitzandClaude Opus 5 a6e9c7a32b M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-06 03:00:33 -04:00

156 lines
5.8 KiB
Python

"""M6: the rolling story summary, anchored to the story it summarizes.
A summary is compressed derived history. It is never the source of truth — the
retained transcript is (`CONTEXT-AND-MEMORY.md` §9) — and it is never allowed to
describe a story the reader is not on.
The inherited design kept one `adventures.story_summary` column and a lineage
cursor recording how far the summariser had read. The cursor was lineage-aware;
the prose it produced was not. After an Undo and a divergence the column still
held sentences about the abandoned line, and the context builder injected it
with no eligibility check at all — acceptance test E03, and measured failing
against the M5 baseline before this module existed.
The fix is not a new lineage system. A summary is a row with a coordinate, the
way a `Memory` already is, and it is filtered through the same
`lineage.Path.clause` chokepoint every other read of the story goes through. So:
eligible == its coordinate is on the active, head-capped lineage
which gives the four behaviours the milestone asks for, without a rule of its
own for any of them:
A -> B -> C -> D, summary covers A..C, head at D eligible
Undo to B not eligible
Redo to D eligible again
diverge from B onto X -> Y not eligible
Nothing is deleted when a line is abandoned. The abandoned line keeps its own
summaries, and they become eligible again if the reader returns to it.
"""
from __future__ import annotations
from sqlalchemy import select
from sqlalchemy.orm import Session
from . import models
from .context import lineage
def record(
db: Session,
adventure: models.Adventure,
text: str,
*,
node: models.Action | None = None,
source_start: int | None = None,
trigger: str = "interval",
model_name: str = "",
) -> models.Summary:
"""Stores one summary at the coordinate the story has reached.
`node` is the last action the summary covers, which is where the row is
anchored. Without one the summary anchors at the head, which is what a
summary the reader typed themselves covers.
"""
branch_id = adventure.head_branch_id
depth = adventure.head_depth
if node is not None and node.depth is not None:
branch_id, depth = node.branch_id, node.depth
row = models.Summary(
adventure_id=adventure.id,
text=text.strip(),
branch_id=branch_id,
depth=depth,
source_start=source_start,
source_end=depth,
trigger=trigger,
model_name=model_name,
)
db.add(row)
mirror(adventure, row.text)
return row
def mirror(adventure: models.Adventure, text: str) -> None:
"""Points `adventures.story_summary` at the summary now in force.
That column is a reader-facing convenience — the Plot panel edits it, the
export bundle carries it — and nothing authoritative may read it. It has no
lineage, so it holds whatever was written last on whatever line, and the M6
review found the summariser seeding itself from exactly that: after a
divergence it was handed the abandoned line's prose and asked to update it
(finding M6-F1).
The fix was to seed generation from `current()` instead. This function keeps
the column honest as well, so what a reader sees in the Plot panel and what
an export carries is the summary the narrator is actually being given.
"""
adventure.story_summary = text or ""
def refresh_mirror(db: Session, adventure: models.Adventure) -> None:
"""Re-points the mirror after the head has moved.
Called from `attempts.restore_state`, which every Undo, Redo, take switch
and Save Point restore goes through. Without it the column would keep
showing a summary the story has moved away from.
"""
row = current(db, adventure)
mirror(adventure, row.text if row is not None else "")
def current(db: Session, adventure: models.Adventure) -> models.Summary | None:
"""The newest summary eligible for the position being read, or None.
Eligibility is the capped lineage clause and nothing else. Ordering by
depth then id takes the newest summary on the path, so a fresher summary
written on a shallower branch does not outrank the deep one it was
superseded by.
"""
return db.execute(
select(models.Summary)
.where(
models.Summary.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Summary),
)
.order_by(models.Summary.depth.desc(), models.Summary.id.desc())
.limit(1)
).scalars().first()
def text_for_prompt(db: Session, adventure: models.Adventure) -> str:
"""The summary the narrator should be shown, or an empty string."""
row = current(db, adventure)
return row.text if row is not None and row.text.strip() else ""
def provenance(row: models.Summary | None) -> dict | None:
"""What the inspector shows about where a summary came from."""
if row is None:
return None
return {
"id": row.id,
"branch_id": row.branch_id,
"depth": row.depth,
"source_start": row.source_start,
"source_end": row.source_end,
"trigger": row.trigger,
"model": row.model_name,
"created_at": row.created_at.isoformat() if row.created_at else None,
}
def all_for(db: Session, adventure: models.Adventure) -> list[models.Summary]:
"""Every stored summary, eligible or not, newest first.
Abandoned summaries are retained rather than deleted, so this is how a
reader or a maintainer sees that they still exist.
"""
return list(db.execute(
select(models.Summary)
.where(models.Summary.adventure_id == adventure.id)
.order_by(models.Summary.id.desc())
).scalars().all())