M6: branch-safe context, summaries and long-term story memory

Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-06 03:00:33 -04:00
co-authored by Claude Opus 5
parent b7005e6fdd
commit a6e9c7a32b
32 changed files with 4040 additions and 84 deletions
@@ -46,6 +46,12 @@ function InsightsPanel({ advId, inspectActionId, onClearInspect, refreshKey }) {
)}
<div className={`token-total ${overBudget ? 'over' : ''}`}>
{tokens.total} / {tokens.budget} tokens
{/* M6: the reply has to fit in the same window, so say how much of it
is being held back for one. Without this the reader can see that
the prompt fits and still get a truncated turn. */}
{tokens.output_reserve > 0 && (
<span className="dim"> · {tokens.output_reserve} reserved for the reply</span>
)}
</div>
<TokenBreakdown
sections={sections}
@@ -76,15 +82,65 @@ function InsightsPanel({ advId, inspectActionId, onClearInspect, refreshKey }) {
)}
{report.memories.used?.map((m, i) => (
<div key={i}>
▸ memory retrieved ({m.pinned ? 'pinned' : `similarity ${m.similarity.toFixed(2)}`}):
{' '}{m.text.length > 90 ? m.text.slice(0, 90) + '…' : m.text}
▸ memory retrieved ({m.pinned ? 'pinned' : `similarity ${m.similarity.toFixed(2)}`})
{/* M6: what weight it carries, and where it came from. An
inference must not read like a record. */}
{m.authority === 'heuristic'
? <span className="mem-heuristic"> · inferred</span>
: <span className="dim"> · from the story</span>}
{m.source?.depth != null && (
<span className="dim">
{' '}· turn {m.source.source_start === m.source.source_end
? m.source.depth
: `${m.source.source_start}–${m.source.source_end}`}
</span>
)}
: {m.text.length > 90 ? m.text.slice(0, 90) + '…' : m.text}
</div>
))}
{report.memories.considered != null && (
<div className="dim">
{report.memories.considered} memories eligible on this line of the story.
</div>
)}
{!report.memories.error && report.memories.used?.length === 0 && (
<div className="dim">Memory bank on — no memories retrieved.</div>
)}
</div>
)}
{/* M6: which summary was used, and what stretch of story it covers, so
"what history did that summary cover?" is answerable here. */}
{report.summary ? (
<div className="insights-cards">
<div>
▸ summary in use
<span className="dim">
{' '}· covers {report.summary.source_start != null
? `turns ${report.summary.source_start}–${report.summary.source_end}`
: `up to turn ${report.summary.source_end ?? report.summary.depth}`}
{report.summary.trigger === 'manual' ? ' · written by you' : ' · generated'}
</span>
</div>
</div>
) : (
<div className="insights-cards dim">▸ no summary is eligible for this point in the story.</div>
)}
{/* M6: a dead memory bank used to be invisible. It is not any more. */}
{report.derived?.some((d) => d.status === 'failed') && (
<div className="insights-cards">
{report.derived.filter((d) => d.status === 'failed').map((d) => (
<div key={d.kind} className="dropped">
⚠ Background {d.kind} work is failing ({d.failures}
{d.failures === 1 ? ' attempt' : ' attempts'}): {d.detail}
</div>
))}
<div className="dim">
The story itself is unaffected — this only stops new {' '}
{report.derived.filter((d) => d.status === 'failed')
.map((d) => d.kind).join(' and ')} work being written.
</div>
</div>
)}
</div>
{sections.map((s, i) => (
+7
View File
@@ -500,3 +500,10 @@
color: var(--text-dim);
font-variant-numeric: tabular-nums;
}
/* M6: a retrieved memory that is an interpretation rather than a record. The
colour is the same one the state panel uses for a reader's own correction,
because both mean "weigh this differently from accepted story". */
.mem-heuristic {
color: #d8a657;
}