M6: branch-safe context, summaries and long-term story memory

Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-06 03:00:33 -04:00
co-authored by Claude Opus 5
parent b7005e6fdd
commit a6e9c7a32b
32 changed files with 4040 additions and 84 deletions
+106
View File
@@ -290,6 +290,42 @@ If the user restores or diverges before part of that source history:
A summary from abandoned history must never leak into active context.
### As implemented (M6)
Lineage safety here has **two** halves, and the M6 review found that having only
the first is not enough.
**The row must be eligible.** A summary is a row with a coordinate, exactly as a
memory is: `branch_id` and `depth` name the last node it covers,
`source_start`/`source_end` the stretch. Eligibility is one question — is that
coordinate on the active, head-capped lineage? — answered by
`lineage.Path.clause`, the same chokepoint every read of the story goes through.
Undo, Redo, Save Point restore and divergence all fall out of that without a
rule of their own.
**The input must be eligible too.** Summary generation is seeded only from a
summary that is itself valid on the current head-capped lineage
(`summaries.current`). Where none is, the new line starts from no previous
summary.
The second half was missing in the first M6 implementation and E03 failed
because of it: the summariser seeded itself from `adventures.story_summary`, a
campaign-global column with no lineage, so after a divergence it was handed the
abandoned line's prose and asked to update it. The row it produced was correctly
anchored to the new branch and therefore *looked* lineage-safe while its
sentences described a story the reader had left. Anchoring the output is not
enough; the input has to be scoped by the same rule.
Abandoned summaries are retained, never deleted, and become eligible again if
the reader returns to the line that produced them.
`adventures.story_summary` remains, as a reader-facing convenience only: the
Plot panel edits it and the export bundle carries it. It is a mirror of whichever
summary is currently eligible — kept in step when one is written and when the
head moves — and **nothing authoritative may read it**. A summary the reader
types is recorded as a row anchored at the position they typed it at, so it
behaves like any other.
## 12. Story Memory
Long-term story memory should retrieve important older information that is not present in recent history or current summary.
@@ -372,6 +408,19 @@ The narrator should be told which memories are:
This avoids turning guesses into canon.
### As implemented (M6)
Two values, `accepted_story` and `heuristic`, on `Memory.authority`. The
**application** classifies, not the model: `memorybank.classify_authority`
reads the memory's own text for hedging ("seemed", "appeared to", "probably"),
so an extractor cannot promote a guess by asserting it confidently. The prompt
marks a heuristic memory `[inferred]` and says in words that such lines are
interpretation rather than established fact.
Retrieval never writes state. A memory of either authority is something the
narrator is shown; the only path to an authoritative change remains the M5 typed
event pipeline (ADR 013).
## 15. Memory Creation
Memories may be generated after accepted turns.
@@ -494,6 +543,16 @@ Potential ranking factors:
The final ranking formula should be simple and inspectable.
### As implemented (M6)
Ranking is **cosine similarity against the retrieval query, plus an explicit
pin**. The other factors listed above — importance, lexical match, recency,
entity overlap, story-thread overlap — are **not implemented**, and remain
future work rather than something M6 delivered.
What M6 does implement, because similarity alone proved insufficient, is
redundancy suppression before the final selection: see §22.
## 21. Memory Budget
Retrieved memories should have a bounded token budget.
@@ -524,6 +583,31 @@ there is little value in also including three memories that say the same thing.
Context builder should prefer the highest-authority concise representation.
### As implemented (M6)
**Between memories, yes.** Retrieval walks the ranked candidates and skips one
that repeats a memory already chosen, keeping the highest-ranked statement of a
fact and its provenance. Two rules bound it: authority is never crossed, so an
inference can never suppress a record or the reverse; and the bar is high, set
where measurement showed distinct facts stop appearing. The number of
suppressed candidates is reported in the context record, so a memory that was
considered and set aside can be told from one that was never eligible.
The threshold is a measured property of the embedding model in use, not a
universal constant, and `memorybank.REDUNDANT_SIMILARITY` records the
measurement beside the value.
The same measurement ruled out the more obvious test. Word overlap fires hardest
on exactly the pair that must not be merged — "Mara promised to return before
dawn" against "Aldric promised to return before dawn" shares most of its words
and means something else — and is weakest on filler that plainly repeats itself.
Wording is a poor proxy for sameness of fact.
**Across layers, not yet.** A fact can still appear in the authoritative state,
in a memory and in recent history at once. That is bounded and legible, but it
is not the "highest-authority concise representation" this section asks for, and
it remains open.
## 23. Imported Knowledge Categories
Imported local files must be classified as:
@@ -672,6 +756,28 @@ Recommended behavior:
- add safety margin,
- fail gracefully if protected context alone is too large.
### As implemented (M6)
`context_token_budget` is the whole window, so the reply is subtracted from it
before any history is selected:
```text
available for history = context_token_budget
- (system, canon, state, summary, memories, user input)
- (max_output_tokens + 64)
```
The 64-token margin covers the separators added after budgeting and the drift
between the app's tokenizer and the serving model's; it is fixed rather than
proportional because what it absorbs does not scale with the budget.
When the protected part alone does not fit, `build_context` raises
`ContextOverflow` naming both figures and what to change, and the turn reports
that as a failed turn. It does not build a prompt it knows will overflow.
Before M6 nothing was reserved: the builder spent the entire budget on input and
left the reply to fit in whatever the endpoint had left.
## 33. Context Snapshot
Every narrator generation should preserve enough information to reconstruct the effective prompt.