M6: branch-safe context, summaries and long-term story memory

Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-06 03:00:33 -04:00
co-authored by Claude Opus 5
parent b7005e6fdd
commit a6e9c7a32b
32 changed files with 4040 additions and 84 deletions
+80 -1
View File
@@ -1,6 +1,6 @@
# Adventure Storyteller — Production Build Milestones
**Status:** In implementation. M1-M5 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04, after an independent review and a corrective pass); M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory — next to brief
**Status:** In implementation. M1-M6 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6: 2026-09-06, each of the last two after an independent review and a corrective pass); M7 — First-Class Imported Knowledge Library — next to brief
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
## 1. Purpose
@@ -699,6 +699,85 @@ Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §A.1
---
## M6 — Outcome
**Complete and accepted 2026-09-06**, after an independent review that found E03
still failing and a corrective pass that fixed it. Report, including the review
findings and the corrective addendum:
`planning/reports/M6-IMPLEMENTATION-REPORT.md`.
**M7 is authorized**: the corrected M6 evidence passes, including E03 end to end
against a real local summariser.
What was inherited and kept, rather than rebuilt:
- **Memory lineage was already correct.** Memories carry `(branch_id, depth)`
and retrieval filters them through the capped-path clause. The ten-step
negative control was measured passing against the M5 baseline *before* any M6
change, and is now pinned by tests. M6 added provenance to the retrieval
result and an authority classification; it did not rewrite the bank.
- The lineage chokepoint, the cursors, and the post-turn pass are unchanged in
shape.
What M6 changed:
- **Summary lineage.** Summaries move from a single `adventures.story_summary`
column to `summaries` rows carrying a coordinate and a source range, filtered
by the same lineage clause as memories. E03 was measured leaking at the M5
baseline and is now closed.
- **Memory authority.** `accepted_story` vs `heuristic`, classified by the
application, marked in the prompt.
- **Context budgeting.** The reply is reserved out of the context budget, and an
impossible configuration raises `ContextOverflow` rather than producing a
prompt known to overflow. Nothing was reserved before M6.
- **Background failure observability.** `derived_status` rows, an API endpoint
and an Insights surface, so the M2 failure — the whole memory bank dead with a
green suite — is visible if it recurs.
- **Real provider-wiring tests**, which mock no factory, plus two pre-existing
test-suite leaks they exposed and which are fixed.
What the independent review found, and what the corrective pass did:
- **E03 was still failing (M6-F1).** Summary *rows* were lineage-anchored, but
generation was seeded from `adventures.story_summary`, a campaign-global
column with no lineage — so a summary generated after a divergence inherited
the abandoned line's prose inside a correctly anchored row. Fixed by seeding
from `summaries.current`. The lesson is recorded in `V1-ACCEPTANCE-TESTS.md`:
a valid E03 test must **regenerate** a summary after diverging, not merely
check that the old row went ineligible.
- **F02 passed by one slot (M6-F2).** With a real embedding model, four
near-identical memories crowded out the one that mattered; the clue survived
only because the default `memory_top_k` is 5. Redundancy suppression now runs
before the final cut, and the clue is retrieved at top_k 5, 4 and 3.
- **Two test defects (M6-F3, M6-F4).** The unit fixture made real network calls
to the default endpoint, and the E03 test and browser check shared the blind
spot above. Both fixed; the browser suite now has a dedicated E03 scenario
that regenerates a summary after divergence.
- **A misleading status (M6-F5).** Derived work that had nothing to do reported
`ok`; it now reports `idle`.
Debt carried forward, deliberately:
- **Retrieval ranking** is cosine similarity plus an explicit pin. Importance,
entity overlap, recency and thread overlap are contemplated by
`CONTEXT-AND-MEMORY.md` §20 and are not implemented. Redundancy suppression
covers the failure mode M6 measured; the richer ranking is open, and matters
to M7 because imported material will compete for the same budget.
- **Cross-layer duplication** (§22) is unimplemented: the same fact can appear
in state, memory and history at once. Bounded and legible, but not the
"highest-authority concise representation" the plan asks for.
- **M7:** imported knowledge is not implemented, and nothing was built to fill
its inspector section.
- **M8:** the Insights additions are functional, not designed; the broader UX
pass remains M8's.
- **M9:** derived rows are not carried in an export bundle, so an imported
campaign starts with no summaries or memories and rebuilds them. Authoritative
history is unaffected.
- **M11:** the long-context evidence here is a bounded fixture, not the M01
100-turn campaign.
---
# M7 — First-Class Imported Knowledge Library
## Objective