Files
interactive-story/backend/app/derived.py
T
JesseMarkowitzandClaude Opus 5 a6e9c7a32b M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-06 03:00:33 -04:00

114 lines
4.4 KiB
Python

"""M6: recording whether background derived work succeeded, and why not.
M2 shipped with the entire memory bank dead and the full test suite green. The
summariser and the embedder raised `AttributeError` inside a fire-and-forget
task: no user-visible error, no log a player would look at, and no failing test,
because every memory test stubbed the provider factories out
(`BUILD-MILESTONES.md`, note from M2; `M2-IMPLEMENTATION-REPORT.md` §A.1).
Two rules follow, and they pull in opposite directions:
* **Derived work must fail softly.** A memory that could not be written, a
summary that could not be generated, an embedding the endpoint refused —
none of these may roll back the accepted narration, the accepted state
events, the authoritative document, the head, or the transcript. The story
turn already happened; the derived work is a commentary on it.
* **It must fail visibly.** Soft failure without a record is what M2 shipped.
So each attempt writes its outcome to one row per (campaign, kind), and that row
is readable through the API. This is deliberately not a job framework: it holds
what happened last, not a queue. Retrying is just running the pass again, which
the ordinary post-turn path already does on the next accepted turn.
"""
from __future__ import annotations
import logging
from sqlalchemy import select
from sqlalchemy.orm import Session
from . import models
log = logging.getLogger(__name__)
# The kinds of derived work. Each is independent: embeddings can be failing
# while summaries succeed, and a reader should be able to see exactly that.
MEMORY = "memory"
SUMMARY = "summary"
EMBEDDING = "embedding"
KINDS = (MEMORY, SUMMARY, EMBEDDING)
def _row(db: Session, adventure_id: int, kind: str) -> models.DerivedStatus:
row = db.execute(
select(models.DerivedStatus).where(
models.DerivedStatus.adventure_id == adventure_id,
models.DerivedStatus.kind == kind,
)
).scalars().first()
if row is None:
row = models.DerivedStatus(adventure_id=adventure_id, kind=kind)
db.add(row)
return row
def succeeded(db: Session, adventure_id: int, kind: str, *, did_work: bool = True) -> None:
"""Records a clean run, clearing any standing failure.
`did_work` separates a pass that produced something from one that found
nothing to do (M6 review finding M6-F5). Both are healthy, and neither is a
failure, but reporting "ok" for a pass that has never actually run reads as
"embeddings are working" when nothing has been embedded. `idle` says the
true thing: it ran, and there was nothing pending.
"""
row = _row(db, adventure_id, kind)
row.status = "ok" if did_work else "idle"
row.detail = ""
row.failures = 0
row.last_attempt_at = models.utcnow()
if did_work:
row.last_success_at = row.last_attempt_at
def failed(db: Session, adventure_id: int, kind: str, exc: BaseException) -> None:
"""Records a failed run, keeping the reason where someone can find it.
The detail is the exception's type and message rather than a traceback: it
is shown to a reader in the Insights panel, and `ProviderError: connection
refused` is the part that tells them what to do. The traceback goes to the
log for a maintainer.
"""
row = _row(db, adventure_id, kind)
row.status = "failed"
row.detail = f"{type(exc).__name__}: {exc}"[:2000]
row.failures = (row.failures or 0) + 1
row.last_attempt_at = models.utcnow()
log.exception("derived %s work failed for adventure %s", kind, adventure_id)
def report(db: Session, adventure_id: int) -> list[dict]:
"""Every kind's last outcome, for the API and the prompt inspector."""
rows = db.execute(
select(models.DerivedStatus)
.where(models.DerivedStatus.adventure_id == adventure_id)
.order_by(models.DerivedStatus.kind)
).scalars().all()
return [
{
"kind": row.kind,
"status": row.status,
"detail": row.detail,
"failures": row.failures,
"last_attempt_at": row.last_attempt_at.isoformat() if row.last_attempt_at else None,
"last_success_at": row.last_success_at.isoformat() if row.last_success_at else None,
}
for row in rows
]
def failing(db: Session, adventure_id: int) -> list[str]:
"""The kinds currently in a failed state, for a compact UI badge."""
return [entry["kind"] for entry in report(db, adventure_id)
if entry["status"] == "failed"]