M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
co-authored by
Claude Opus 5
parent
b7005e6fdd
commit
a6e9c7a32b
@@ -394,6 +394,79 @@ any distance, in either direction, and identical whether the position is reached
|
||||
from in front of it or from behind. This is the property §10.4's hybrid storage
|
||||
must preserve.
|
||||
|
||||
### M6 — derived context: summaries, memory and budgeting
|
||||
|
||||
Four things future milestones rely on, all built on the lineage machinery M3-M5
|
||||
established rather than beside it.
|
||||
|
||||
**Summary lineage — both halves.** A summary is a `summaries` row carrying
|
||||
`(branch_id, depth)` for the last node it covers plus a
|
||||
`source_start`/`source_end` range. The invariant M6 holds is:
|
||||
|
||||
> Both summary eligibility and the prior-summary input to the summarizer are
|
||||
> lineage-scoped.
|
||||
|
||||
*Eligibility* is `lineage.Path.clause` over the row's coordinate — the same
|
||||
capped-path clause that filters actions and memories — so Undo, Redo, Save Point
|
||||
restore and divergence need no summary-specific rule. *Input* is
|
||||
`summaries.current`, the same question the context builder asks, so a summary is
|
||||
only ever built on top of one that is valid where the story now stands; where
|
||||
none is, generation starts from nothing.
|
||||
|
||||
The second half is not decorative. The first M6 implementation had only the
|
||||
first, seeding generation from `adventures.story_summary`, and the review
|
||||
demonstrated abandoned prose reaching an active prompt inside a row that was
|
||||
itself correctly anchored. Anchoring the output does not make the content safe.
|
||||
|
||||
Nothing is deleted when a line is abandoned. `adventures.story_summary` survives
|
||||
as a reader-facing convenience only — the Plot panel edits it, the export bundle
|
||||
carries it — mirroring whichever summary is eligible, kept in step by
|
||||
`summaries.record` and by `attempts.restore_state` when the head moves. Nothing
|
||||
authoritative reads it.
|
||||
|
||||
**Memory lineage and provenance.** Unchanged from what M3 built and M6 verified:
|
||||
a memory carries `(branch_id, depth)` and a source range, and retrieval filters
|
||||
through the capped path. M6 adds provenance to the *retrieval result*, in the
|
||||
same query that fetches the text, so the inspector can answer "where did this
|
||||
come from?" without a query per memory.
|
||||
|
||||
**Memory authority.** `Memory.authority` is `accepted_story` or `heuristic`,
|
||||
decided by the application in `memorybank.classify_authority`, and rendered into
|
||||
the prompt as an explicit mark. Retrieval never writes state; the M5 typed-event
|
||||
path remains the only route to an authoritative change.
|
||||
|
||||
**Retrieval ranking and redundancy.** Ranking is cosine similarity plus an
|
||||
explicit pin; the other factors `CONTEXT-AND-MEMORY.md` §20 contemplates are not
|
||||
implemented. Before the final top-k cut, retrieval drops a candidate that
|
||||
repeats one already chosen, never across authority classes, at a threshold
|
||||
measured against the configured embedding model
|
||||
(`memorybank.REDUNDANT_SIMILARITY`). Suppressed candidates are reported so the
|
||||
selection stays inspectable. Without this, a stretch of repetitive story fills
|
||||
the whole memory budget with near-copies and evicts the one memory that
|
||||
mattered — which the review measured happening.
|
||||
|
||||
**Context budgeting.** The reply is reserved out of `context_token_budget`
|
||||
before history is selected, with a fixed 64-token margin. Protected content —
|
||||
narrator rules, canon, authoritative state, the reader's input, the reply
|
||||
reserve — is never dropped to fit older prose; history is the elastic part and
|
||||
is filled newest-first until the remaining budget is spent. If the protected
|
||||
part alone exceeds the budget, `build_context` raises `ContextOverflow` rather
|
||||
than assembling a prompt known to overflow.
|
||||
|
||||
**Background failure observability.** Derived work (memory extraction, summary
|
||||
generation, embedding) runs in a fire-and-forget task and must not take an
|
||||
accepted turn down with it. Each pass is wrapped so that a failure rolls back
|
||||
only its own uncommitted work and writes a `derived_status` row naming the kind,
|
||||
the error and the attempt count. That row is served by
|
||||
`GET /adventures/{id}/derived` and shown in the Insights panel. M2 shipped with
|
||||
the whole memory bank dead and the suite green; this is the mechanism that makes
|
||||
the same failure visible.
|
||||
|
||||
**Prompt inspection.** The context report carries per-section token counts, the
|
||||
budget, the output reserve, the protected total, the history allowance, the
|
||||
summary's provenance, each retrieved memory's authority and source coordinate,
|
||||
and the derived-work status.
|
||||
|
||||
**Divergence is a property of the lineage, not a flag.** The first write below a
|
||||
moved-back head forks; Undo alone never does. After the fork, the displaced
|
||||
future is no longer on the lineage being read, so ordinary Redo finds nothing
|
||||
|
||||
Reference in New Issue
Block a user