M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
co-authored by
Claude Opus 5
parent
b7005e6fdd
commit
a6e9c7a32b
@@ -46,6 +46,12 @@ function InsightsPanel({ advId, inspectActionId, onClearInspect, refreshKey }) {
|
||||
)}
|
||||
<div className={`token-total ${overBudget ? 'over' : ''}`}>
|
||||
{tokens.total} / {tokens.budget} tokens
|
||||
{/* M6: the reply has to fit in the same window, so say how much of it
|
||||
is being held back for one. Without this the reader can see that
|
||||
the prompt fits and still get a truncated turn. */}
|
||||
{tokens.output_reserve > 0 && (
|
||||
<span className="dim"> · {tokens.output_reserve} reserved for the reply</span>
|
||||
)}
|
||||
</div>
|
||||
<TokenBreakdown
|
||||
sections={sections}
|
||||
@@ -76,15 +82,65 @@ function InsightsPanel({ advId, inspectActionId, onClearInspect, refreshKey }) {
|
||||
)}
|
||||
{report.memories.used?.map((m, i) => (
|
||||
<div key={i}>
|
||||
▸ memory retrieved ({m.pinned ? 'pinned' : `similarity ${m.similarity.toFixed(2)}`}):
|
||||
{' '}{m.text.length > 90 ? m.text.slice(0, 90) + '…' : m.text}
|
||||
▸ memory retrieved ({m.pinned ? 'pinned' : `similarity ${m.similarity.toFixed(2)}`})
|
||||
{/* M6: what weight it carries, and where it came from. An
|
||||
inference must not read like a record. */}
|
||||
{m.authority === 'heuristic'
|
||||
? <span className="mem-heuristic"> · inferred</span>
|
||||
: <span className="dim"> · from the story</span>}
|
||||
{m.source?.depth != null && (
|
||||
<span className="dim">
|
||||
{' '}· turn {m.source.source_start === m.source.source_end
|
||||
? m.source.depth
|
||||
: `${m.source.source_start}–${m.source.source_end}`}
|
||||
</span>
|
||||
)}
|
||||
: {m.text.length > 90 ? m.text.slice(0, 90) + '…' : m.text}
|
||||
</div>
|
||||
))}
|
||||
{report.memories.considered != null && (
|
||||
<div className="dim">
|
||||
{report.memories.considered} memories eligible on this line of the story.
|
||||
</div>
|
||||
)}
|
||||
{!report.memories.error && report.memories.used?.length === 0 && (
|
||||
<div className="dim">Memory bank on — no memories retrieved.</div>
|
||||
)}
|
||||
</div>
|
||||
)}
|
||||
{/* M6: which summary was used, and what stretch of story it covers, so
|
||||
"what history did that summary cover?" is answerable here. */}
|
||||
{report.summary ? (
|
||||
<div className="insights-cards">
|
||||
<div>
|
||||
▸ summary in use
|
||||
<span className="dim">
|
||||
{' '}· covers {report.summary.source_start != null
|
||||
? `turns ${report.summary.source_start}–${report.summary.source_end}`
|
||||
: `up to turn ${report.summary.source_end ?? report.summary.depth}`}
|
||||
{report.summary.trigger === 'manual' ? ' · written by you' : ' · generated'}
|
||||
</span>
|
||||
</div>
|
||||
</div>
|
||||
) : (
|
||||
<div className="insights-cards dim">▸ no summary is eligible for this point in the story.</div>
|
||||
)}
|
||||
{/* M6: a dead memory bank used to be invisible. It is not any more. */}
|
||||
{report.derived?.some((d) => d.status === 'failed') && (
|
||||
<div className="insights-cards">
|
||||
{report.derived.filter((d) => d.status === 'failed').map((d) => (
|
||||
<div key={d.kind} className="dropped">
|
||||
⚠ Background {d.kind} work is failing ({d.failures}
|
||||
{d.failures === 1 ? ' attempt' : ' attempts'}): {d.detail}
|
||||
</div>
|
||||
))}
|
||||
<div className="dim">
|
||||
The story itself is unaffected — this only stops new {' '}
|
||||
{report.derived.filter((d) => d.status === 'failed')
|
||||
.map((d) => d.kind).join(' and ')} work being written.
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
|
||||
{sections.map((s, i) => (
|
||||
|
||||
@@ -500,3 +500,10 @@
|
||||
color: var(--text-dim);
|
||||
font-variant-numeric: tabular-nums;
|
||||
}
|
||||
|
||||
/* M6: a retrieved memory that is an interpretation rather than a record. The
|
||||
colour is the same one the state panel uses for a reader's own correction,
|
||||
because both mean "weigh this differently from accepted story". */
|
||||
.mem-heuristic {
|
||||
color: #d8a657;
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user