M6: branch-safe context, summaries and long-term story memory

Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-06 03:00:33 -04:00
co-authored by Claude Opus 5
parent b7005e6fdd
commit a6e9c7a32b
32 changed files with 4040 additions and 84 deletions
+46 -3
View File
@@ -1,8 +1,51 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v2.6
- **Revision date:** 2026-09-03
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M4 implemented and accepted**; M5 is next to brief.
- **Package:** Adventure Storyteller Planning Package v2.7
- **Revision date:** 2026-09-06
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M6 implemented and accepted**; M7 is next to brief.
## v2.7 — M5 and M6 Closeout (2026-09-06)
Two milestones, and one lesson they share: an independent review found a real
defect in each after the implementation reported success, and in both cases the
defect was invisible to the tests the implementation had written for itself.
**M5 — Genre-Neutral Authoritative Narrative State.** Accepted 2026-09-04. Its
review returned *PASS WITH CORRECTIVE WORK REQUIRED*: editing a narrator turn
rewound the campaign's live state while the head stayed at the tip, breaking the
`transcript position == head == authoritative state` invariant. The corrective
pass rebuilt narrator editing on the §§14-15 fork semantics. Recorded here late:
M5's closeout did not add a package revision entry, and this one covers it.
**M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory.** Accepted
2026-09-06. Sequence:
1. **Implementation.** Summaries moved onto lineage-anchored rows; memory
authority, provenance, an output reserve, and observable derived-work
failure were added. Reported as passing, including E03.
2. **Independent review — E03 still failed.** Summary *rows* were anchored, but
generation was seeded from `adventures.story_summary`, a campaign-global
column with no lineage. After a divergence the summariser was handed the
abandoned line's prose and asked to update it, so the new summary carried
abandoned content inside a correctly anchored row. The review also found F02
passing by a single retrieval slot, two test defects, and a misleading
derived-work status.
3. **Corrective pass.** Generation is seeded from `summaries.current`;
redundancy suppression was added before the retrieval cut; the test fixture
no longer makes real network calls; the browser suite gained a scenario that
regenerates a summary after divergence.
**What the planning package learned from M6, recorded in the active documents:**
- `CONTEXT-AND-MEMORY.md` §11 — lineage safety has two halves. Anchoring the
output row is not enough; the *input* to the summarizer must be scoped by the
same rule.
- `V1-ACCEPTANCE-TESTS.md` E03 — a valid test must regenerate a summary after
diverging. Checking only that the old row went ineligible passes while the
defect is live.
- `CONTEXT-AND-MEMORY.md` §20/§22 — what M6 implements (redundancy suppression
between memories) is now distinguished from what it does not (importance,
entity, recency and thread ranking; cross-layer duplication).
## v2.6 — M4 Closeout (2026-09-03)