M6: branch-safe context, summaries and long-term story memory

Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-06 03:00:33 -04:00
co-authored by Claude Opus 5
parent b7005e6fdd
commit a6e9c7a32b
32 changed files with 4040 additions and 84 deletions
+123
View File
@@ -916,6 +916,15 @@ Continue enough turns to exercise long-term memory retrieval.
### Pass
Narrator does not retrieve/use the discarded revelation as active-history truth.
### Result — PASS (M6, 2026-09-05)
The ten-step negative control is `test_e02_the_ten_step_memory_negative_control`,
with the Save Point variant beside it. Both assert on the assembled prompt and
on the eligibility clause, not on the narration.
Measured passing against the M5 baseline *before* any M6 change: memory lineage
was inherited correct, and M6's contribution here is the test that pins it.
---
## E03 — Abandoned Summary Cannot Leak
@@ -932,6 +941,42 @@ Narrator does not retrieve/use the discarded revelation as active-history truth.
### Pass
Old summary content from abandoned future is not applied.
### Result — PASS (M6 corrective pass, 2026-09-06)
Failed twice before it passed, and the history matters because it defines the
shape a valid E03 test has to have.
1. At the M5 baseline the summary was a single column with no coordinate and
survived Undo plus divergence into the active prompt.
2. The first M6 implementation made summaries lineage-anchored rows, and the
test written for it checked that the **old row** became ineligible. The
independent review then found E03 still failing end to end: generation was
seeded from `adventures.story_summary`, so the summary produced *on the new
line* inherited the abandoned line's prose inside a correctly anchored row.
3. The corrective pass seeds generation from `summaries.current`.
**A valid E03 test must regenerate a summary after the divergence.** Checking
only that the old row is ineligible passes while the defect is live. The
regression now required is:
```text
path A: enough history for a real summary, sentinel established on it
POSITIVE CONTROL — the sentinel is in the path-A summary and prompt
move the head below the sentinel, diverge
path B: play far enough that a NEW summary is generated
prove a new summary row exists and is not path A's
prove no path-A action is on path B's lineage
prove the sentinel is absent from the new summary
prove the sentinel is absent from the complete active prompt
prove the old row is retained but ineligible
```
Evidence: `test_e03_a_summary_generated_after_divergence_carries_no_abandoned_content`
(deterministic, fails against the pre-corrective implementation); the same
sequence against a real local summariser; and a dedicated browser scenario that
regenerates a summary after diverging rather than repeating the old blind spot.
---
## E04 — Scene State Is Lineage-Safe
@@ -960,6 +1005,12 @@ Conduct a multi-turn conversation with Mara.
### Pass
Narrator remembers immediately preceding dialogue and actions.
### Result — PASS (M6, 2026-09-05)
`test_f01_recent_turns_stay_in_the_prompt`: the preceding turns and the reader's
own input are present in the assembled prompt, asserted on the context report
rather than on the narration.
---
## F02 — Old Important Event Retrieval
@@ -974,6 +1025,27 @@ Narrator remembers immediately preceding dialogue and actions.
### Pass
Relevant old clue can be recovered through summary/memory/state.
### Result — PASS (M6 corrective pass, 2026-09-06)
A distinctive clue is planted, long turns are played over it, and the clue is
*absent* from the verbatim history sections and *present* through a retrieved
memory, with `history.included < history.total` — so recovery does not come from
sending the transcript.
The independent review found this passing only by a one-slot margin: with a real
embedding model, four near-identical filler memories scored 0.79–0.82 against
the clue's 0.61, so the clue placed fifth and survived only because the default
`memory_top_k` is 5. At 4 it was evicted and F02 failed.
The corrective pass added redundancy suppression before the final selection.
On the same fixture and the same real embedding model, 3 of 5 candidates are now
suppressed as repetitions and the clue is retrieved at `top_k` 5, 4 **and** 3.
Ranking itself remains cosine similarity plus an explicit pin; the further
factors `CONTEXT-AND-MEMORY.md` §20 contemplates are not implemented and are
recorded there as future work.
---
## F03 — Prompt Remains Bounded
@@ -986,6 +1058,13 @@ Generate a long story.
### Pass
Application does not continually append full transcript until context overflows.
### Result — PASS (M6, 2026-09-05)
`test_f03_the_prompt_stays_bounded_as_the_story_grows` and
`test_the_context_size_stops_growing_once_the_budget_is_reached`. The second
measures from a story that already fills the budget, then triples it: the
prompt does not move, while the action count does.
---
## F04 — Output Token Reserve
@@ -995,6 +1074,14 @@ Application does not continually append full transcript until context overflows.
### Pass
Context builder leaves sufficient room for narrator output and does not regularly fail because input consumes entire context.
### Result — PASS (M6, 2026-09-05)
The reply is reserved out of the context budget before history is selected, and
the report exposes it (`tokens.output_reserve`). Three tests: the reserve
survives a long story; an impossible budget raises `ContextOverflow` naming both
figures; and that refusal reaches the reader as a failed turn without disturbing
the stored story. Nothing was reserved before M6.
---
## F05 — Prompt Inspector
@@ -1017,6 +1104,18 @@ User can determine at least:
Exact UI may vary.
### Result — PARTIAL, complete for the components M6 owns (2026-09-05)
`test_f05_the_inspector_shows_every_component_m6_owns` asserts the report
carries narrator rules, authoritative state, the summary in use with its source
coverage, retrieved memories with authority and provenance, recent history, the
reader's input, model and settings, and per-component token costs alongside the
budget, the reply reserve and the history allowance. All of this is rendered in
the Insights panel and was verified in a real browser.
"Retrieved knowledge" is M7's imported-document section and is not implemented;
nothing was built to fill it.
---
## F06 — Retrieval Provenance
@@ -1026,6 +1125,13 @@ Exact UI may vary.
### Pass
A retrieved memory or imported chunk can be traced to its source record/file.
### Result — PASS for story memory (M6, 2026-09-05)
`test_f06_a_retrieved_memory_is_traceable_to_its_source`: every retrieved memory
carries `branch_id`, `depth` and its source range, and the test resolves that
coordinate back to a real action of the campaign's accepted history. Imported
chunks are M7's half of this criterion and are not implemented.
---
## F07 — Heuristic Memory Is Not Canon
@@ -1048,6 +1154,15 @@ Mara is definitely working against Captain Vale.
as authoritative fact.
### Result — PASS (M6, 2026-09-05)
`Memory.authority` is `accepted_story` or `heuristic`, classified by the
application rather than the model. The prompt marks an inference `[inferred]`
and says such lines are not established fact.
`test_f07_a_heuristic_memory_is_labelled_and_is_not_state` also asserts the
inference did not become an authoritative fact: retrieval never writes state,
and the M5 typed-event path remains the only route to one.
---
## F08 — Memory Failure Is Non-Fatal
@@ -1060,6 +1175,14 @@ Cause embedding/memory extraction failure if test harness supports it.
### Pass
Accepted turn persists and story can continue; derived memory may be retried later.
### Result — PASS (M6, 2026-09-05)
With the summariser and embedder both raising, the accepted narration, the
authoritative state, the head and the transcript all survive, the next turn
still plays, and the failure is recorded per kind in `derived_status` and served
by `GET /adventures/{id}/derived`. A later healthy run clears it. This is the
M2 failure — the whole memory bank dead with a green suite — made visible.
---
# G. Imported Knowledge