M6: branch-safe context, summaries and long-term story memory

Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-06 03:00:33 -04:00
co-authored by Claude Opus 5
parent b7005e6fdd
commit a6e9c7a32b
32 changed files with 4040 additions and 84 deletions
+80 -1
View File
@@ -1,6 +1,6 @@
# Adventure Storyteller — Production Build Milestones
**Status:** In implementation. M1-M5 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04, after an independent review and a corrective pass); M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory — next to brief
**Status:** In implementation. M1-M6 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6: 2026-09-06, each of the last two after an independent review and a corrective pass); M7 — First-Class Imported Knowledge Library — next to brief
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
## 1. Purpose
@@ -699,6 +699,85 @@ Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §A.1
---
## M6 — Outcome
**Complete and accepted 2026-09-06**, after an independent review that found E03
still failing and a corrective pass that fixed it. Report, including the review
findings and the corrective addendum:
`planning/reports/M6-IMPLEMENTATION-REPORT.md`.
**M7 is authorized**: the corrected M6 evidence passes, including E03 end to end
against a real local summariser.
What was inherited and kept, rather than rebuilt:
- **Memory lineage was already correct.** Memories carry `(branch_id, depth)`
and retrieval filters them through the capped-path clause. The ten-step
negative control was measured passing against the M5 baseline *before* any M6
change, and is now pinned by tests. M6 added provenance to the retrieval
result and an authority classification; it did not rewrite the bank.
- The lineage chokepoint, the cursors, and the post-turn pass are unchanged in
shape.
What M6 changed:
- **Summary lineage.** Summaries move from a single `adventures.story_summary`
column to `summaries` rows carrying a coordinate and a source range, filtered
by the same lineage clause as memories. E03 was measured leaking at the M5
baseline and is now closed.
- **Memory authority.** `accepted_story` vs `heuristic`, classified by the
application, marked in the prompt.
- **Context budgeting.** The reply is reserved out of the context budget, and an
impossible configuration raises `ContextOverflow` rather than producing a
prompt known to overflow. Nothing was reserved before M6.
- **Background failure observability.** `derived_status` rows, an API endpoint
and an Insights surface, so the M2 failure — the whole memory bank dead with a
green suite — is visible if it recurs.
- **Real provider-wiring tests**, which mock no factory, plus two pre-existing
test-suite leaks they exposed and which are fixed.
What the independent review found, and what the corrective pass did:
- **E03 was still failing (M6-F1).** Summary *rows* were lineage-anchored, but
generation was seeded from `adventures.story_summary`, a campaign-global
column with no lineage — so a summary generated after a divergence inherited
the abandoned line's prose inside a correctly anchored row. Fixed by seeding
from `summaries.current`. The lesson is recorded in `V1-ACCEPTANCE-TESTS.md`:
a valid E03 test must **regenerate** a summary after diverging, not merely
check that the old row went ineligible.
- **F02 passed by one slot (M6-F2).** With a real embedding model, four
near-identical memories crowded out the one that mattered; the clue survived
only because the default `memory_top_k` is 5. Redundancy suppression now runs
before the final cut, and the clue is retrieved at top_k 5, 4 and 3.
- **Two test defects (M6-F3, M6-F4).** The unit fixture made real network calls
to the default endpoint, and the E03 test and browser check shared the blind
spot above. Both fixed; the browser suite now has a dedicated E03 scenario
that regenerates a summary after divergence.
- **A misleading status (M6-F5).** Derived work that had nothing to do reported
`ok`; it now reports `idle`.
Debt carried forward, deliberately:
- **Retrieval ranking** is cosine similarity plus an explicit pin. Importance,
entity overlap, recency and thread overlap are contemplated by
`CONTEXT-AND-MEMORY.md` §20 and are not implemented. Redundancy suppression
covers the failure mode M6 measured; the richer ranking is open, and matters
to M7 because imported material will compete for the same budget.
- **Cross-layer duplication** (§22) is unimplemented: the same fact can appear
in state, memory and history at once. Bounded and legible, but not the
"highest-authority concise representation" the plan asks for.
- **M7:** imported knowledge is not implemented, and nothing was built to fill
its inspector section.
- **M8:** the Insights additions are functional, not designed; the broader UX
pass remains M8's.
- **M9:** derived rows are not carried in an export bundle, so an imported
campaign starts with no summaries or memories and rebuilds them. Authoritative
history is unaffected.
- **M11:** the long-context evidence here is a bounded fixture, not the M01
100-turn campaign.
---
# M7 — First-Class Imported Knowledge Library
## Objective
+106
View File
@@ -290,6 +290,42 @@ If the user restores or diverges before part of that source history:
A summary from abandoned history must never leak into active context.
### As implemented (M6)
Lineage safety here has **two** halves, and the M6 review found that having only
the first is not enough.
**The row must be eligible.** A summary is a row with a coordinate, exactly as a
memory is: `branch_id` and `depth` name the last node it covers,
`source_start`/`source_end` the stretch. Eligibility is one question — is that
coordinate on the active, head-capped lineage? — answered by
`lineage.Path.clause`, the same chokepoint every read of the story goes through.
Undo, Redo, Save Point restore and divergence all fall out of that without a
rule of their own.
**The input must be eligible too.** Summary generation is seeded only from a
summary that is itself valid on the current head-capped lineage
(`summaries.current`). Where none is, the new line starts from no previous
summary.
The second half was missing in the first M6 implementation and E03 failed
because of it: the summariser seeded itself from `adventures.story_summary`, a
campaign-global column with no lineage, so after a divergence it was handed the
abandoned line's prose and asked to update it. The row it produced was correctly
anchored to the new branch and therefore *looked* lineage-safe while its
sentences described a story the reader had left. Anchoring the output is not
enough; the input has to be scoped by the same rule.
Abandoned summaries are retained, never deleted, and become eligible again if
the reader returns to the line that produced them.
`adventures.story_summary` remains, as a reader-facing convenience only: the
Plot panel edits it and the export bundle carries it. It is a mirror of whichever
summary is currently eligible — kept in step when one is written and when the
head moves — and **nothing authoritative may read it**. A summary the reader
types is recorded as a row anchored at the position they typed it at, so it
behaves like any other.
## 12. Story Memory
Long-term story memory should retrieve important older information that is not present in recent history or current summary.
@@ -372,6 +408,19 @@ The narrator should be told which memories are:
This avoids turning guesses into canon.
### As implemented (M6)
Two values, `accepted_story` and `heuristic`, on `Memory.authority`. The
**application** classifies, not the model: `memorybank.classify_authority`
reads the memory's own text for hedging ("seemed", "appeared to", "probably"),
so an extractor cannot promote a guess by asserting it confidently. The prompt
marks a heuristic memory `[inferred]` and says in words that such lines are
interpretation rather than established fact.
Retrieval never writes state. A memory of either authority is something the
narrator is shown; the only path to an authoritative change remains the M5 typed
event pipeline (ADR 013).
## 15. Memory Creation
Memories may be generated after accepted turns.
@@ -494,6 +543,16 @@ Potential ranking factors:
The final ranking formula should be simple and inspectable.
### As implemented (M6)
Ranking is **cosine similarity against the retrieval query, plus an explicit
pin**. The other factors listed above — importance, lexical match, recency,
entity overlap, story-thread overlap — are **not implemented**, and remain
future work rather than something M6 delivered.
What M6 does implement, because similarity alone proved insufficient, is
redundancy suppression before the final selection: see §22.
## 21. Memory Budget
Retrieved memories should have a bounded token budget.
@@ -524,6 +583,31 @@ there is little value in also including three memories that say the same thing.
Context builder should prefer the highest-authority concise representation.
### As implemented (M6)
**Between memories, yes.** Retrieval walks the ranked candidates and skips one
that repeats a memory already chosen, keeping the highest-ranked statement of a
fact and its provenance. Two rules bound it: authority is never crossed, so an
inference can never suppress a record or the reverse; and the bar is high, set
where measurement showed distinct facts stop appearing. The number of
suppressed candidates is reported in the context record, so a memory that was
considered and set aside can be told from one that was never eligible.
The threshold is a measured property of the embedding model in use, not a
universal constant, and `memorybank.REDUNDANT_SIMILARITY` records the
measurement beside the value.
The same measurement ruled out the more obvious test. Word overlap fires hardest
on exactly the pair that must not be merged — "Mara promised to return before
dawn" against "Aldric promised to return before dawn" shares most of its words
and means something else — and is weakest on filler that plainly repeats itself.
Wording is a poor proxy for sameness of fact.
**Across layers, not yet.** A fact can still appear in the authoritative state,
in a memory and in recent history at once. That is bounded and legible, but it
is not the "highest-authority concise representation" this section asks for, and
it remains open.
## 23. Imported Knowledge Categories
Imported local files must be classified as:
@@ -672,6 +756,28 @@ Recommended behavior:
- add safety margin,
- fail gracefully if protected context alone is too large.
### As implemented (M6)
`context_token_budget` is the whole window, so the reply is subtracted from it
before any history is selected:
```text
available for history = context_token_budget
- (system, canon, state, summary, memories, user input)
- (max_output_tokens + 64)
```
The 64-token margin covers the separators added after budgeting and the drift
between the app's tokenizer and the serving model's; it is fixed rather than
proportional because what it absorbs does not scale with the budget.
When the protected part alone does not fit, `build_context` raises
`ContextOverflow` naming both figures and what to change, and the turn reports
that as a failed turn. It does not build a prompt it knows will overflow.
Before M6 nothing was reserved: the builder spent the entire budget on input and
left the reply to fit in whatever the endpoint had left.
## 33. Context Snapshot
Every narrator generation should preserve enough information to reconstruct the effective prompt.
+34
View File
@@ -453,6 +453,40 @@ story_thread:
These are narrative continuity tools, not RPG quests.
## 16A. Derived Context Tables (M6, as implemented)
Two tables and one column carry M6's derived context. All three are derived
data: deleting them changes no accepted history, no authoritative state and no
head position.
```text
summaries
id, adventure_id
text
branch_id, depth the coordinate of the last node covered
source_start, source_end the stretch of story summarized
trigger "interval" (generated) or "manual" (reader-written)
model_name, created_at
derived_status one row per (adventure, kind)
kind "memory" | "summary" | "embedding"
status "ok" (did work) | "idle" (nothing pending) | "failed"
detail, failures
last_attempt_at, last_success_at
memories.authority "accepted_story" | "heuristic"
```
`summaries` mirrors the shape `memories` already had, deliberately: both are
derived rows anchored to a coordinate on a path, and both are filtered by the
same lineage clause — and both are also the *input* to the next round of derived
work, which is why summary generation reads `summaries.current` rather than any
campaign-global field.
`adventures.story_summary` is retained as the reader's edit surface and the
export field, mirroring whichever summary is eligible. It carries no lineage of
its own and nothing authoritative reads it.
## 17. State Version
The system must reconstruct authoritative state at any retained turn.
+11 -9
View File
@@ -3,9 +3,10 @@
**This file is the index. Start here.**
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
milestones **M1, M2, M3 and M4 implemented and accepted** (M3 and M4:
2026-09-03).
**Next:** **M5 — Genre-Neutral Authoritative Narrative State.** Its brief has not
milestones **M1 through M6 implemented and accepted** (M3 and M4: 2026-09-03;
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
independent review found a real defect and a corrective pass fixed it.
**Next:** **M7 — First-Class Imported Knowledge Library.** Its brief has not
been written yet, and writing it is the current action.
**Package version:** see `VERSION.md`, which records what each revision changed
@@ -245,13 +246,14 @@ M5-M11, one at a time see BUILD-MILESTONES.md
**One milestone at a time. Do not begin a milestone before its brief exists.**
**No M5 brief has been prepared.** Writing one is the current action, informed by
the M4 report's §U readiness assessment and by the note `BUILD-MILESTONES.md`
attaches to M5 — in particular that M4 added 55 tests using the inherited
world-state values as deterministic instrumentation, which M5 must **move rather
than delete**.
**No M7 brief has been prepared.** Writing one is the current action, informed by
the M6 report's §W readiness assessment and by the retrieval debt
`BUILD-MILESTONES.md` records against M6 — in particular that ranking is
similarity plus a pin, so imported material will compete for the same memory
budget as story memory, and that cross-layer duplication
(`CONTEXT-AND-MEMORY.md` §22) is still open.
**No conditions remain open on M1-M4.** The browser smoke condition that M3 and
**No conditions remain open on M1-M6.** The browser smoke condition that M3 and
M4 both carried was satisfied at M4 closeout: a real Firefox exercised the
rendered DOM for both milestones' controls, 44/44 checks passing. The M3 report's
§M.2 and the M4 report's §M record the condition as it stood; the M4 report's §W
+73
View File
@@ -394,6 +394,79 @@ any distance, in either direction, and identical whether the position is reached
from in front of it or from behind. This is the property §10.4's hybrid storage
must preserve.
### M6 — derived context: summaries, memory and budgeting
Four things future milestones rely on, all built on the lineage machinery M3-M5
established rather than beside it.
**Summary lineage — both halves.** A summary is a `summaries` row carrying
`(branch_id, depth)` for the last node it covers plus a
`source_start`/`source_end` range. The invariant M6 holds is:
> Both summary eligibility and the prior-summary input to the summarizer are
> lineage-scoped.
*Eligibility* is `lineage.Path.clause` over the row's coordinate — the same
capped-path clause that filters actions and memories — so Undo, Redo, Save Point
restore and divergence need no summary-specific rule. *Input* is
`summaries.current`, the same question the context builder asks, so a summary is
only ever built on top of one that is valid where the story now stands; where
none is, generation starts from nothing.
The second half is not decorative. The first M6 implementation had only the
first, seeding generation from `adventures.story_summary`, and the review
demonstrated abandoned prose reaching an active prompt inside a row that was
itself correctly anchored. Anchoring the output does not make the content safe.
Nothing is deleted when a line is abandoned. `adventures.story_summary` survives
as a reader-facing convenience only — the Plot panel edits it, the export bundle
carries it — mirroring whichever summary is eligible, kept in step by
`summaries.record` and by `attempts.restore_state` when the head moves. Nothing
authoritative reads it.
**Memory lineage and provenance.** Unchanged from what M3 built and M6 verified:
a memory carries `(branch_id, depth)` and a source range, and retrieval filters
through the capped path. M6 adds provenance to the *retrieval result*, in the
same query that fetches the text, so the inspector can answer "where did this
come from?" without a query per memory.
**Memory authority.** `Memory.authority` is `accepted_story` or `heuristic`,
decided by the application in `memorybank.classify_authority`, and rendered into
the prompt as an explicit mark. Retrieval never writes state; the M5 typed-event
path remains the only route to an authoritative change.
**Retrieval ranking and redundancy.** Ranking is cosine similarity plus an
explicit pin; the other factors `CONTEXT-AND-MEMORY.md` §20 contemplates are not
implemented. Before the final top-k cut, retrieval drops a candidate that
repeats one already chosen, never across authority classes, at a threshold
measured against the configured embedding model
(`memorybank.REDUNDANT_SIMILARITY`). Suppressed candidates are reported so the
selection stays inspectable. Without this, a stretch of repetitive story fills
the whole memory budget with near-copies and evicts the one memory that
mattered — which the review measured happening.
**Context budgeting.** The reply is reserved out of `context_token_budget`
before history is selected, with a fixed 64-token margin. Protected content —
narrator rules, canon, authoritative state, the reader's input, the reply
reserve — is never dropped to fit older prose; history is the elastic part and
is filled newest-first until the remaining budget is spent. If the protected
part alone exceeds the budget, `build_context` raises `ContextOverflow` rather
than assembling a prompt known to overflow.
**Background failure observability.** Derived work (memory extraction, summary
generation, embedding) runs in a fire-and-forget task and must not take an
accepted turn down with it. Each pass is wrapped so that a failure rolls back
only its own uncommitted work and writes a `derived_status` row naming the kind,
the error and the attempt count. That row is served by
`GET /adventures/{id}/derived` and shown in the Insights panel. M2 shipped with
the whole memory bank dead and the suite green; this is the mechanism that makes
the same failure visible.
**Prompt inspection.** The context report carries per-section token counts, the
budget, the output reserve, the protected total, the history allowance, the
summary's provenance, each retrieved memory's authority and source coordinate,
and the derived-work status.
**Divergence is a property of the lineage, not a flag.** The first write below a
moved-back head forks; Undo alone never does. After the fork, the displaced
future is no longer on the lineage being read, so ordinary Redo finds nothing
+123
View File
@@ -916,6 +916,15 @@ Continue enough turns to exercise long-term memory retrieval.
### Pass
Narrator does not retrieve/use the discarded revelation as active-history truth.
### Result — PASS (M6, 2026-09-05)
The ten-step negative control is `test_e02_the_ten_step_memory_negative_control`,
with the Save Point variant beside it. Both assert on the assembled prompt and
on the eligibility clause, not on the narration.
Measured passing against the M5 baseline *before* any M6 change: memory lineage
was inherited correct, and M6's contribution here is the test that pins it.
---
## E03 — Abandoned Summary Cannot Leak
@@ -932,6 +941,42 @@ Narrator does not retrieve/use the discarded revelation as active-history truth.
### Pass
Old summary content from abandoned future is not applied.
### Result — PASS (M6 corrective pass, 2026-09-06)
Failed twice before it passed, and the history matters because it defines the
shape a valid E03 test has to have.
1. At the M5 baseline the summary was a single column with no coordinate and
survived Undo plus divergence into the active prompt.
2. The first M6 implementation made summaries lineage-anchored rows, and the
test written for it checked that the **old row** became ineligible. The
independent review then found E03 still failing end to end: generation was
seeded from `adventures.story_summary`, so the summary produced *on the new
line* inherited the abandoned line's prose inside a correctly anchored row.
3. The corrective pass seeds generation from `summaries.current`.
**A valid E03 test must regenerate a summary after the divergence.** Checking
only that the old row is ineligible passes while the defect is live. The
regression now required is:
```text
path A: enough history for a real summary, sentinel established on it
POSITIVE CONTROL — the sentinel is in the path-A summary and prompt
move the head below the sentinel, diverge
path B: play far enough that a NEW summary is generated
prove a new summary row exists and is not path A's
prove no path-A action is on path B's lineage
prove the sentinel is absent from the new summary
prove the sentinel is absent from the complete active prompt
prove the old row is retained but ineligible
```
Evidence: `test_e03_a_summary_generated_after_divergence_carries_no_abandoned_content`
(deterministic, fails against the pre-corrective implementation); the same
sequence against a real local summariser; and a dedicated browser scenario that
regenerates a summary after diverging rather than repeating the old blind spot.
---
## E04 — Scene State Is Lineage-Safe
@@ -960,6 +1005,12 @@ Conduct a multi-turn conversation with Mara.
### Pass
Narrator remembers immediately preceding dialogue and actions.
### Result — PASS (M6, 2026-09-05)
`test_f01_recent_turns_stay_in_the_prompt`: the preceding turns and the reader's
own input are present in the assembled prompt, asserted on the context report
rather than on the narration.
---
## F02 — Old Important Event Retrieval
@@ -974,6 +1025,27 @@ Narrator remembers immediately preceding dialogue and actions.
### Pass
Relevant old clue can be recovered through summary/memory/state.
### Result — PASS (M6 corrective pass, 2026-09-06)
A distinctive clue is planted, long turns are played over it, and the clue is
*absent* from the verbatim history sections and *present* through a retrieved
memory, with `history.included < history.total` — so recovery does not come from
sending the transcript.
The independent review found this passing only by a one-slot margin: with a real
embedding model, four near-identical filler memories scored 0.79–0.82 against
the clue's 0.61, so the clue placed fifth and survived only because the default
`memory_top_k` is 5. At 4 it was evicted and F02 failed.
The corrective pass added redundancy suppression before the final selection.
On the same fixture and the same real embedding model, 3 of 5 candidates are now
suppressed as repetitions and the clue is retrieved at `top_k` 5, 4 **and** 3.
Ranking itself remains cosine similarity plus an explicit pin; the further
factors `CONTEXT-AND-MEMORY.md` §20 contemplates are not implemented and are
recorded there as future work.
---
## F03 — Prompt Remains Bounded
@@ -986,6 +1058,13 @@ Generate a long story.
### Pass
Application does not continually append full transcript until context overflows.
### Result — PASS (M6, 2026-09-05)
`test_f03_the_prompt_stays_bounded_as_the_story_grows` and
`test_the_context_size_stops_growing_once_the_budget_is_reached`. The second
measures from a story that already fills the budget, then triples it: the
prompt does not move, while the action count does.
---
## F04 — Output Token Reserve
@@ -995,6 +1074,14 @@ Application does not continually append full transcript until context overflows.
### Pass
Context builder leaves sufficient room for narrator output and does not regularly fail because input consumes entire context.
### Result — PASS (M6, 2026-09-05)
The reply is reserved out of the context budget before history is selected, and
the report exposes it (`tokens.output_reserve`). Three tests: the reserve
survives a long story; an impossible budget raises `ContextOverflow` naming both
figures; and that refusal reaches the reader as a failed turn without disturbing
the stored story. Nothing was reserved before M6.
---
## F05 — Prompt Inspector
@@ -1017,6 +1104,18 @@ User can determine at least:
Exact UI may vary.
### Result — PARTIAL, complete for the components M6 owns (2026-09-05)
`test_f05_the_inspector_shows_every_component_m6_owns` asserts the report
carries narrator rules, authoritative state, the summary in use with its source
coverage, retrieved memories with authority and provenance, recent history, the
reader's input, model and settings, and per-component token costs alongside the
budget, the reply reserve and the history allowance. All of this is rendered in
the Insights panel and was verified in a real browser.
"Retrieved knowledge" is M7's imported-document section and is not implemented;
nothing was built to fill it.
---
## F06 — Retrieval Provenance
@@ -1026,6 +1125,13 @@ Exact UI may vary.
### Pass
A retrieved memory or imported chunk can be traced to its source record/file.
### Result — PASS for story memory (M6, 2026-09-05)
`test_f06_a_retrieved_memory_is_traceable_to_its_source`: every retrieved memory
carries `branch_id`, `depth` and its source range, and the test resolves that
coordinate back to a real action of the campaign's accepted history. Imported
chunks are M7's half of this criterion and are not implemented.
---
## F07 — Heuristic Memory Is Not Canon
@@ -1048,6 +1154,15 @@ Mara is definitely working against Captain Vale.
as authoritative fact.
### Result — PASS (M6, 2026-09-05)
`Memory.authority` is `accepted_story` or `heuristic`, classified by the
application rather than the model. The prompt marks an inference `[inferred]`
and says such lines are not established fact.
`test_f07_a_heuristic_memory_is_labelled_and_is_not_state` also asserts the
inference did not become an authoritative fact: retrieval never writes state,
and the M5 typed-event path remains the only route to one.
---
## F08 — Memory Failure Is Non-Fatal
@@ -1060,6 +1175,14 @@ Cause embedding/memory extraction failure if test harness supports it.
### Pass
Accepted turn persists and story can continue; derived memory may be retried later.
### Result — PASS (M6, 2026-09-05)
With the summariser and embedder both raising, the accepted narration, the
authoritative state, the head and the transcript all survive, the next turn
still plays, and the failure is recorded per kind in `derived_status` and served
by `GET /adventures/{id}/derived`. A later healthy run clears it. This is the
M2 failure — the whole memory bank dead with a green suite — made visible.
---
# G. Imported Knowledge
+46 -3
View File
@@ -1,8 +1,51 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v2.6
- **Revision date:** 2026-09-03
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M4 implemented and accepted**; M5 is next to brief.
- **Package:** Adventure Storyteller Planning Package v2.7
- **Revision date:** 2026-09-06
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M6 implemented and accepted**; M7 is next to brief.
## v2.7 — M5 and M6 Closeout (2026-09-06)
Two milestones, and one lesson they share: an independent review found a real
defect in each after the implementation reported success, and in both cases the
defect was invisible to the tests the implementation had written for itself.
**M5 — Genre-Neutral Authoritative Narrative State.** Accepted 2026-09-04. Its
review returned *PASS WITH CORRECTIVE WORK REQUIRED*: editing a narrator turn
rewound the campaign's live state while the head stayed at the tip, breaking the
`transcript position == head == authoritative state` invariant. The corrective
pass rebuilt narrator editing on the §§14-15 fork semantics. Recorded here late:
M5's closeout did not add a package revision entry, and this one covers it.
**M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory.** Accepted
2026-09-06. Sequence:
1. **Implementation.** Summaries moved onto lineage-anchored rows; memory
authority, provenance, an output reserve, and observable derived-work
failure were added. Reported as passing, including E03.
2. **Independent review — E03 still failed.** Summary *rows* were anchored, but
generation was seeded from `adventures.story_summary`, a campaign-global
column with no lineage. After a divergence the summariser was handed the
abandoned line's prose and asked to update it, so the new summary carried
abandoned content inside a correctly anchored row. The review also found F02
passing by a single retrieval slot, two test defects, and a misleading
derived-work status.
3. **Corrective pass.** Generation is seeded from `summaries.current`;
redundancy suppression was added before the retrieval cut; the test fixture
no longer makes real network calls; the browser suite gained a scenario that
regenerates a summary after divergence.
**What the planning package learned from M6, recorded in the active documents:**
- `CONTEXT-AND-MEMORY.md` §11 — lineage safety has two halves. Anchoring the
output row is not enough; the *input* to the summarizer must be scoped by the
same rule.
- `V1-ACCEPTANCE-TESTS.md` E03 — a valid test must regenerate a summary after
diverging. Checking only that the old row went ineligible passes while the
defect is live.
- `CONTEXT-AND-MEMORY.md` §20/§22 — what M6 implements (redundancy suppression
between memories) is now distinguished from what it does not (importance,
entity, recency and thread ranking; cross-layer duplication).
## v2.6 — M4 Closeout (2026-09-03)
File diff suppressed because it is too large Load Diff