M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
co-authored by
Claude Opus 5
parent
b7005e6fdd
commit
a6e9c7a32b
@@ -0,0 +1,113 @@
|
||||
"""M6: recording whether background derived work succeeded, and why not.
|
||||
|
||||
M2 shipped with the entire memory bank dead and the full test suite green. The
|
||||
summariser and the embedder raised `AttributeError` inside a fire-and-forget
|
||||
task: no user-visible error, no log a player would look at, and no failing test,
|
||||
because every memory test stubbed the provider factories out
|
||||
(`BUILD-MILESTONES.md`, note from M2; `M2-IMPLEMENTATION-REPORT.md` §A.1).
|
||||
|
||||
Two rules follow, and they pull in opposite directions:
|
||||
|
||||
* **Derived work must fail softly.** A memory that could not be written, a
|
||||
summary that could not be generated, an embedding the endpoint refused —
|
||||
none of these may roll back the accepted narration, the accepted state
|
||||
events, the authoritative document, the head, or the transcript. The story
|
||||
turn already happened; the derived work is a commentary on it.
|
||||
* **It must fail visibly.** Soft failure without a record is what M2 shipped.
|
||||
|
||||
So each attempt writes its outcome to one row per (campaign, kind), and that row
|
||||
is readable through the API. This is deliberately not a job framework: it holds
|
||||
what happened last, not a queue. Retrying is just running the pass again, which
|
||||
the ordinary post-turn path already does on the next accepted turn.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from . import models
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# The kinds of derived work. Each is independent: embeddings can be failing
|
||||
# while summaries succeed, and a reader should be able to see exactly that.
|
||||
MEMORY = "memory"
|
||||
SUMMARY = "summary"
|
||||
EMBEDDING = "embedding"
|
||||
KINDS = (MEMORY, SUMMARY, EMBEDDING)
|
||||
|
||||
|
||||
def _row(db: Session, adventure_id: int, kind: str) -> models.DerivedStatus:
|
||||
row = db.execute(
|
||||
select(models.DerivedStatus).where(
|
||||
models.DerivedStatus.adventure_id == adventure_id,
|
||||
models.DerivedStatus.kind == kind,
|
||||
)
|
||||
).scalars().first()
|
||||
if row is None:
|
||||
row = models.DerivedStatus(adventure_id=adventure_id, kind=kind)
|
||||
db.add(row)
|
||||
return row
|
||||
|
||||
|
||||
def succeeded(db: Session, adventure_id: int, kind: str, *, did_work: bool = True) -> None:
|
||||
"""Records a clean run, clearing any standing failure.
|
||||
|
||||
`did_work` separates a pass that produced something from one that found
|
||||
nothing to do (M6 review finding M6-F5). Both are healthy, and neither is a
|
||||
failure, but reporting "ok" for a pass that has never actually run reads as
|
||||
"embeddings are working" when nothing has been embedded. `idle` says the
|
||||
true thing: it ran, and there was nothing pending.
|
||||
"""
|
||||
row = _row(db, adventure_id, kind)
|
||||
row.status = "ok" if did_work else "idle"
|
||||
row.detail = ""
|
||||
row.failures = 0
|
||||
row.last_attempt_at = models.utcnow()
|
||||
if did_work:
|
||||
row.last_success_at = row.last_attempt_at
|
||||
|
||||
|
||||
def failed(db: Session, adventure_id: int, kind: str, exc: BaseException) -> None:
|
||||
"""Records a failed run, keeping the reason where someone can find it.
|
||||
|
||||
The detail is the exception's type and message rather than a traceback: it
|
||||
is shown to a reader in the Insights panel, and `ProviderError: connection
|
||||
refused` is the part that tells them what to do. The traceback goes to the
|
||||
log for a maintainer.
|
||||
"""
|
||||
row = _row(db, adventure_id, kind)
|
||||
row.status = "failed"
|
||||
row.detail = f"{type(exc).__name__}: {exc}"[:2000]
|
||||
row.failures = (row.failures or 0) + 1
|
||||
row.last_attempt_at = models.utcnow()
|
||||
log.exception("derived %s work failed for adventure %s", kind, adventure_id)
|
||||
|
||||
|
||||
def report(db: Session, adventure_id: int) -> list[dict]:
|
||||
"""Every kind's last outcome, for the API and the prompt inspector."""
|
||||
rows = db.execute(
|
||||
select(models.DerivedStatus)
|
||||
.where(models.DerivedStatus.adventure_id == adventure_id)
|
||||
.order_by(models.DerivedStatus.kind)
|
||||
).scalars().all()
|
||||
return [
|
||||
{
|
||||
"kind": row.kind,
|
||||
"status": row.status,
|
||||
"detail": row.detail,
|
||||
"failures": row.failures,
|
||||
"last_attempt_at": row.last_attempt_at.isoformat() if row.last_attempt_at else None,
|
||||
"last_success_at": row.last_success_at.isoformat() if row.last_success_at else None,
|
||||
}
|
||||
for row in rows
|
||||
]
|
||||
|
||||
|
||||
def failing(db: Session, adventure_id: int) -> list[str]:
|
||||
"""The kinds currently in a failed state, for a compact UI badge."""
|
||||
return [entry["kind"] for entry in report(db, adventure_id)
|
||||
if entry["status"] == "failed"]
|
||||
Reference in New Issue
Block a user