M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
co-authored by
Claude Opus 5
parent
b7005e6fdd
commit
a6e9c7a32b
@@ -1,6 +1,6 @@
|
||||
# Adventure Storyteller — Production Build Milestones
|
||||
|
||||
**Status:** In implementation. M1-M5 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04, after an independent review and a corrective pass); M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory — next to brief
|
||||
**Status:** In implementation. M1-M6 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6: 2026-09-06, each of the last two after an independent review and a corrective pass); M7 — First-Class Imported Knowledge Library — next to brief
|
||||
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
||||
|
||||
## 1. Purpose
|
||||
@@ -699,6 +699,85 @@ Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §A.1
|
||||
|
||||
---
|
||||
|
||||
## M6 — Outcome
|
||||
|
||||
**Complete and accepted 2026-09-06**, after an independent review that found E03
|
||||
still failing and a corrective pass that fixed it. Report, including the review
|
||||
findings and the corrective addendum:
|
||||
`planning/reports/M6-IMPLEMENTATION-REPORT.md`.
|
||||
|
||||
**M7 is authorized**: the corrected M6 evidence passes, including E03 end to end
|
||||
against a real local summariser.
|
||||
|
||||
What was inherited and kept, rather than rebuilt:
|
||||
|
||||
- **Memory lineage was already correct.** Memories carry `(branch_id, depth)`
|
||||
and retrieval filters them through the capped-path clause. The ten-step
|
||||
negative control was measured passing against the M5 baseline *before* any M6
|
||||
change, and is now pinned by tests. M6 added provenance to the retrieval
|
||||
result and an authority classification; it did not rewrite the bank.
|
||||
- The lineage chokepoint, the cursors, and the post-turn pass are unchanged in
|
||||
shape.
|
||||
|
||||
What M6 changed:
|
||||
|
||||
- **Summary lineage.** Summaries move from a single `adventures.story_summary`
|
||||
column to `summaries` rows carrying a coordinate and a source range, filtered
|
||||
by the same lineage clause as memories. E03 was measured leaking at the M5
|
||||
baseline and is now closed.
|
||||
- **Memory authority.** `accepted_story` vs `heuristic`, classified by the
|
||||
application, marked in the prompt.
|
||||
- **Context budgeting.** The reply is reserved out of the context budget, and an
|
||||
impossible configuration raises `ContextOverflow` rather than producing a
|
||||
prompt known to overflow. Nothing was reserved before M6.
|
||||
- **Background failure observability.** `derived_status` rows, an API endpoint
|
||||
and an Insights surface, so the M2 failure — the whole memory bank dead with a
|
||||
green suite — is visible if it recurs.
|
||||
- **Real provider-wiring tests**, which mock no factory, plus two pre-existing
|
||||
test-suite leaks they exposed and which are fixed.
|
||||
|
||||
What the independent review found, and what the corrective pass did:
|
||||
|
||||
- **E03 was still failing (M6-F1).** Summary *rows* were lineage-anchored, but
|
||||
generation was seeded from `adventures.story_summary`, a campaign-global
|
||||
column with no lineage — so a summary generated after a divergence inherited
|
||||
the abandoned line's prose inside a correctly anchored row. Fixed by seeding
|
||||
from `summaries.current`. The lesson is recorded in `V1-ACCEPTANCE-TESTS.md`:
|
||||
a valid E03 test must **regenerate** a summary after diverging, not merely
|
||||
check that the old row went ineligible.
|
||||
- **F02 passed by one slot (M6-F2).** With a real embedding model, four
|
||||
near-identical memories crowded out the one that mattered; the clue survived
|
||||
only because the default `memory_top_k` is 5. Redundancy suppression now runs
|
||||
before the final cut, and the clue is retrieved at top_k 5, 4 and 3.
|
||||
- **Two test defects (M6-F3, M6-F4).** The unit fixture made real network calls
|
||||
to the default endpoint, and the E03 test and browser check shared the blind
|
||||
spot above. Both fixed; the browser suite now has a dedicated E03 scenario
|
||||
that regenerates a summary after divergence.
|
||||
- **A misleading status (M6-F5).** Derived work that had nothing to do reported
|
||||
`ok`; it now reports `idle`.
|
||||
|
||||
Debt carried forward, deliberately:
|
||||
|
||||
- **Retrieval ranking** is cosine similarity plus an explicit pin. Importance,
|
||||
entity overlap, recency and thread overlap are contemplated by
|
||||
`CONTEXT-AND-MEMORY.md` §20 and are not implemented. Redundancy suppression
|
||||
covers the failure mode M6 measured; the richer ranking is open, and matters
|
||||
to M7 because imported material will compete for the same budget.
|
||||
- **Cross-layer duplication** (§22) is unimplemented: the same fact can appear
|
||||
in state, memory and history at once. Bounded and legible, but not the
|
||||
"highest-authority concise representation" the plan asks for.
|
||||
- **M7:** imported knowledge is not implemented, and nothing was built to fill
|
||||
its inspector section.
|
||||
- **M8:** the Insights additions are functional, not designed; the broader UX
|
||||
pass remains M8's.
|
||||
- **M9:** derived rows are not carried in an export bundle, so an imported
|
||||
campaign starts with no summaries or memories and rebuilds them. Authoritative
|
||||
history is unaffected.
|
||||
- **M11:** the long-context evidence here is a bounded fixture, not the M01
|
||||
100-turn campaign.
|
||||
|
||||
---
|
||||
|
||||
# M7 — First-Class Imported Knowledge Library
|
||||
|
||||
## Objective
|
||||
|
||||
@@ -290,6 +290,42 @@ If the user restores or diverges before part of that source history:
|
||||
|
||||
A summary from abandoned history must never leak into active context.
|
||||
|
||||
### As implemented (M6)
|
||||
|
||||
Lineage safety here has **two** halves, and the M6 review found that having only
|
||||
the first is not enough.
|
||||
|
||||
**The row must be eligible.** A summary is a row with a coordinate, exactly as a
|
||||
memory is: `branch_id` and `depth` name the last node it covers,
|
||||
`source_start`/`source_end` the stretch. Eligibility is one question — is that
|
||||
coordinate on the active, head-capped lineage? — answered by
|
||||
`lineage.Path.clause`, the same chokepoint every read of the story goes through.
|
||||
Undo, Redo, Save Point restore and divergence all fall out of that without a
|
||||
rule of their own.
|
||||
|
||||
**The input must be eligible too.** Summary generation is seeded only from a
|
||||
summary that is itself valid on the current head-capped lineage
|
||||
(`summaries.current`). Where none is, the new line starts from no previous
|
||||
summary.
|
||||
|
||||
The second half was missing in the first M6 implementation and E03 failed
|
||||
because of it: the summariser seeded itself from `adventures.story_summary`, a
|
||||
campaign-global column with no lineage, so after a divergence it was handed the
|
||||
abandoned line's prose and asked to update it. The row it produced was correctly
|
||||
anchored to the new branch and therefore *looked* lineage-safe while its
|
||||
sentences described a story the reader had left. Anchoring the output is not
|
||||
enough; the input has to be scoped by the same rule.
|
||||
|
||||
Abandoned summaries are retained, never deleted, and become eligible again if
|
||||
the reader returns to the line that produced them.
|
||||
|
||||
`adventures.story_summary` remains, as a reader-facing convenience only: the
|
||||
Plot panel edits it and the export bundle carries it. It is a mirror of whichever
|
||||
summary is currently eligible — kept in step when one is written and when the
|
||||
head moves — and **nothing authoritative may read it**. A summary the reader
|
||||
types is recorded as a row anchored at the position they typed it at, so it
|
||||
behaves like any other.
|
||||
|
||||
## 12. Story Memory
|
||||
|
||||
Long-term story memory should retrieve important older information that is not present in recent history or current summary.
|
||||
@@ -372,6 +408,19 @@ The narrator should be told which memories are:
|
||||
|
||||
This avoids turning guesses into canon.
|
||||
|
||||
### As implemented (M6)
|
||||
|
||||
Two values, `accepted_story` and `heuristic`, on `Memory.authority`. The
|
||||
**application** classifies, not the model: `memorybank.classify_authority`
|
||||
reads the memory's own text for hedging ("seemed", "appeared to", "probably"),
|
||||
so an extractor cannot promote a guess by asserting it confidently. The prompt
|
||||
marks a heuristic memory `[inferred]` and says in words that such lines are
|
||||
interpretation rather than established fact.
|
||||
|
||||
Retrieval never writes state. A memory of either authority is something the
|
||||
narrator is shown; the only path to an authoritative change remains the M5 typed
|
||||
event pipeline (ADR 013).
|
||||
|
||||
## 15. Memory Creation
|
||||
|
||||
Memories may be generated after accepted turns.
|
||||
@@ -494,6 +543,16 @@ Potential ranking factors:
|
||||
|
||||
The final ranking formula should be simple and inspectable.
|
||||
|
||||
### As implemented (M6)
|
||||
|
||||
Ranking is **cosine similarity against the retrieval query, plus an explicit
|
||||
pin**. The other factors listed above — importance, lexical match, recency,
|
||||
entity overlap, story-thread overlap — are **not implemented**, and remain
|
||||
future work rather than something M6 delivered.
|
||||
|
||||
What M6 does implement, because similarity alone proved insufficient, is
|
||||
redundancy suppression before the final selection: see §22.
|
||||
|
||||
## 21. Memory Budget
|
||||
|
||||
Retrieved memories should have a bounded token budget.
|
||||
@@ -524,6 +583,31 @@ there is little value in also including three memories that say the same thing.
|
||||
|
||||
Context builder should prefer the highest-authority concise representation.
|
||||
|
||||
### As implemented (M6)
|
||||
|
||||
**Between memories, yes.** Retrieval walks the ranked candidates and skips one
|
||||
that repeats a memory already chosen, keeping the highest-ranked statement of a
|
||||
fact and its provenance. Two rules bound it: authority is never crossed, so an
|
||||
inference can never suppress a record or the reverse; and the bar is high, set
|
||||
where measurement showed distinct facts stop appearing. The number of
|
||||
suppressed candidates is reported in the context record, so a memory that was
|
||||
considered and set aside can be told from one that was never eligible.
|
||||
|
||||
The threshold is a measured property of the embedding model in use, not a
|
||||
universal constant, and `memorybank.REDUNDANT_SIMILARITY` records the
|
||||
measurement beside the value.
|
||||
|
||||
The same measurement ruled out the more obvious test. Word overlap fires hardest
|
||||
on exactly the pair that must not be merged — "Mara promised to return before
|
||||
dawn" against "Aldric promised to return before dawn" shares most of its words
|
||||
and means something else — and is weakest on filler that plainly repeats itself.
|
||||
Wording is a poor proxy for sameness of fact.
|
||||
|
||||
**Across layers, not yet.** A fact can still appear in the authoritative state,
|
||||
in a memory and in recent history at once. That is bounded and legible, but it
|
||||
is not the "highest-authority concise representation" this section asks for, and
|
||||
it remains open.
|
||||
|
||||
## 23. Imported Knowledge Categories
|
||||
|
||||
Imported local files must be classified as:
|
||||
@@ -672,6 +756,28 @@ Recommended behavior:
|
||||
- add safety margin,
|
||||
- fail gracefully if protected context alone is too large.
|
||||
|
||||
### As implemented (M6)
|
||||
|
||||
`context_token_budget` is the whole window, so the reply is subtracted from it
|
||||
before any history is selected:
|
||||
|
||||
```text
|
||||
available for history = context_token_budget
|
||||
- (system, canon, state, summary, memories, user input)
|
||||
- (max_output_tokens + 64)
|
||||
```
|
||||
|
||||
The 64-token margin covers the separators added after budgeting and the drift
|
||||
between the app's tokenizer and the serving model's; it is fixed rather than
|
||||
proportional because what it absorbs does not scale with the budget.
|
||||
|
||||
When the protected part alone does not fit, `build_context` raises
|
||||
`ContextOverflow` naming both figures and what to change, and the turn reports
|
||||
that as a failed turn. It does not build a prompt it knows will overflow.
|
||||
|
||||
Before M6 nothing was reserved: the builder spent the entire budget on input and
|
||||
left the reply to fit in whatever the endpoint had left.
|
||||
|
||||
## 33. Context Snapshot
|
||||
|
||||
Every narrator generation should preserve enough information to reconstruct the effective prompt.
|
||||
|
||||
@@ -453,6 +453,40 @@ story_thread:
|
||||
|
||||
These are narrative continuity tools, not RPG quests.
|
||||
|
||||
## 16A. Derived Context Tables (M6, as implemented)
|
||||
|
||||
Two tables and one column carry M6's derived context. All three are derived
|
||||
data: deleting them changes no accepted history, no authoritative state and no
|
||||
head position.
|
||||
|
||||
```text
|
||||
summaries
|
||||
id, adventure_id
|
||||
text
|
||||
branch_id, depth the coordinate of the last node covered
|
||||
source_start, source_end the stretch of story summarized
|
||||
trigger "interval" (generated) or "manual" (reader-written)
|
||||
model_name, created_at
|
||||
|
||||
derived_status one row per (adventure, kind)
|
||||
kind "memory" | "summary" | "embedding"
|
||||
status "ok" (did work) | "idle" (nothing pending) | "failed"
|
||||
detail, failures
|
||||
last_attempt_at, last_success_at
|
||||
|
||||
memories.authority "accepted_story" | "heuristic"
|
||||
```
|
||||
|
||||
`summaries` mirrors the shape `memories` already had, deliberately: both are
|
||||
derived rows anchored to a coordinate on a path, and both are filtered by the
|
||||
same lineage clause — and both are also the *input* to the next round of derived
|
||||
work, which is why summary generation reads `summaries.current` rather than any
|
||||
campaign-global field.
|
||||
|
||||
`adventures.story_summary` is retained as the reader's edit surface and the
|
||||
export field, mirroring whichever summary is eligible. It carries no lineage of
|
||||
its own and nothing authoritative reads it.
|
||||
|
||||
## 17. State Version
|
||||
|
||||
The system must reconstruct authoritative state at any retained turn.
|
||||
|
||||
+11
-9
@@ -3,9 +3,10 @@
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
milestones **M1, M2, M3 and M4 implemented and accepted** (M3 and M4:
|
||||
2026-09-03).
|
||||
**Next:** **M5 — Genre-Neutral Authoritative Narrative State.** Its brief has not
|
||||
milestones **M1 through M6 implemented and accepted** (M3 and M4: 2026-09-03;
|
||||
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
|
||||
independent review found a real defect and a corrective pass fixed it.
|
||||
**Next:** **M7 — First-Class Imported Knowledge Library.** Its brief has not
|
||||
been written yet, and writing it is the current action.
|
||||
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
@@ -245,13 +246,14 @@ M5-M11, one at a time see BUILD-MILESTONES.md
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
|
||||
**No M5 brief has been prepared.** Writing one is the current action, informed by
|
||||
the M4 report's §U readiness assessment and by the note `BUILD-MILESTONES.md`
|
||||
attaches to M5 — in particular that M4 added 55 tests using the inherited
|
||||
world-state values as deterministic instrumentation, which M5 must **move rather
|
||||
than delete**.
|
||||
**No M7 brief has been prepared.** Writing one is the current action, informed by
|
||||
the M6 report's §W readiness assessment and by the retrieval debt
|
||||
`BUILD-MILESTONES.md` records against M6 — in particular that ranking is
|
||||
similarity plus a pin, so imported material will compete for the same memory
|
||||
budget as story memory, and that cross-layer duplication
|
||||
(`CONTEXT-AND-MEMORY.md` §22) is still open.
|
||||
|
||||
**No conditions remain open on M1-M4.** The browser smoke condition that M3 and
|
||||
**No conditions remain open on M1-M6.** The browser smoke condition that M3 and
|
||||
M4 both carried was satisfied at M4 closeout: a real Firefox exercised the
|
||||
rendered DOM for both milestones' controls, 44/44 checks passing. The M3 report's
|
||||
§M.2 and the M4 report's §M record the condition as it stood; the M4 report's §W
|
||||
|
||||
@@ -394,6 +394,79 @@ any distance, in either direction, and identical whether the position is reached
|
||||
from in front of it or from behind. This is the property §10.4's hybrid storage
|
||||
must preserve.
|
||||
|
||||
### M6 — derived context: summaries, memory and budgeting
|
||||
|
||||
Four things future milestones rely on, all built on the lineage machinery M3-M5
|
||||
established rather than beside it.
|
||||
|
||||
**Summary lineage — both halves.** A summary is a `summaries` row carrying
|
||||
`(branch_id, depth)` for the last node it covers plus a
|
||||
`source_start`/`source_end` range. The invariant M6 holds is:
|
||||
|
||||
> Both summary eligibility and the prior-summary input to the summarizer are
|
||||
> lineage-scoped.
|
||||
|
||||
*Eligibility* is `lineage.Path.clause` over the row's coordinate — the same
|
||||
capped-path clause that filters actions and memories — so Undo, Redo, Save Point
|
||||
restore and divergence need no summary-specific rule. *Input* is
|
||||
`summaries.current`, the same question the context builder asks, so a summary is
|
||||
only ever built on top of one that is valid where the story now stands; where
|
||||
none is, generation starts from nothing.
|
||||
|
||||
The second half is not decorative. The first M6 implementation had only the
|
||||
first, seeding generation from `adventures.story_summary`, and the review
|
||||
demonstrated abandoned prose reaching an active prompt inside a row that was
|
||||
itself correctly anchored. Anchoring the output does not make the content safe.
|
||||
|
||||
Nothing is deleted when a line is abandoned. `adventures.story_summary` survives
|
||||
as a reader-facing convenience only — the Plot panel edits it, the export bundle
|
||||
carries it — mirroring whichever summary is eligible, kept in step by
|
||||
`summaries.record` and by `attempts.restore_state` when the head moves. Nothing
|
||||
authoritative reads it.
|
||||
|
||||
**Memory lineage and provenance.** Unchanged from what M3 built and M6 verified:
|
||||
a memory carries `(branch_id, depth)` and a source range, and retrieval filters
|
||||
through the capped path. M6 adds provenance to the *retrieval result*, in the
|
||||
same query that fetches the text, so the inspector can answer "where did this
|
||||
come from?" without a query per memory.
|
||||
|
||||
**Memory authority.** `Memory.authority` is `accepted_story` or `heuristic`,
|
||||
decided by the application in `memorybank.classify_authority`, and rendered into
|
||||
the prompt as an explicit mark. Retrieval never writes state; the M5 typed-event
|
||||
path remains the only route to an authoritative change.
|
||||
|
||||
**Retrieval ranking and redundancy.** Ranking is cosine similarity plus an
|
||||
explicit pin; the other factors `CONTEXT-AND-MEMORY.md` §20 contemplates are not
|
||||
implemented. Before the final top-k cut, retrieval drops a candidate that
|
||||
repeats one already chosen, never across authority classes, at a threshold
|
||||
measured against the configured embedding model
|
||||
(`memorybank.REDUNDANT_SIMILARITY`). Suppressed candidates are reported so the
|
||||
selection stays inspectable. Without this, a stretch of repetitive story fills
|
||||
the whole memory budget with near-copies and evicts the one memory that
|
||||
mattered — which the review measured happening.
|
||||
|
||||
**Context budgeting.** The reply is reserved out of `context_token_budget`
|
||||
before history is selected, with a fixed 64-token margin. Protected content —
|
||||
narrator rules, canon, authoritative state, the reader's input, the reply
|
||||
reserve — is never dropped to fit older prose; history is the elastic part and
|
||||
is filled newest-first until the remaining budget is spent. If the protected
|
||||
part alone exceeds the budget, `build_context` raises `ContextOverflow` rather
|
||||
than assembling a prompt known to overflow.
|
||||
|
||||
**Background failure observability.** Derived work (memory extraction, summary
|
||||
generation, embedding) runs in a fire-and-forget task and must not take an
|
||||
accepted turn down with it. Each pass is wrapped so that a failure rolls back
|
||||
only its own uncommitted work and writes a `derived_status` row naming the kind,
|
||||
the error and the attempt count. That row is served by
|
||||
`GET /adventures/{id}/derived` and shown in the Insights panel. M2 shipped with
|
||||
the whole memory bank dead and the suite green; this is the mechanism that makes
|
||||
the same failure visible.
|
||||
|
||||
**Prompt inspection.** The context report carries per-section token counts, the
|
||||
budget, the output reserve, the protected total, the history allowance, the
|
||||
summary's provenance, each retrieved memory's authority and source coordinate,
|
||||
and the derived-work status.
|
||||
|
||||
**Divergence is a property of the lineage, not a flag.** The first write below a
|
||||
moved-back head forks; Undo alone never does. After the fork, the displaced
|
||||
future is no longer on the lineage being read, so ordinary Redo finds nothing
|
||||
|
||||
@@ -916,6 +916,15 @@ Continue enough turns to exercise long-term memory retrieval.
|
||||
### Pass
|
||||
Narrator does not retrieve/use the discarded revelation as active-history truth.
|
||||
|
||||
|
||||
### Result — PASS (M6, 2026-09-05)
|
||||
The ten-step negative control is `test_e02_the_ten_step_memory_negative_control`,
|
||||
with the Save Point variant beside it. Both assert on the assembled prompt and
|
||||
on the eligibility clause, not on the narration.
|
||||
|
||||
Measured passing against the M5 baseline *before* any M6 change: memory lineage
|
||||
was inherited correct, and M6's contribution here is the test that pins it.
|
||||
|
||||
---
|
||||
|
||||
## E03 — Abandoned Summary Cannot Leak
|
||||
@@ -932,6 +941,42 @@ Narrator does not retrieve/use the discarded revelation as active-history truth.
|
||||
### Pass
|
||||
Old summary content from abandoned future is not applied.
|
||||
|
||||
|
||||
### Result — PASS (M6 corrective pass, 2026-09-06)
|
||||
|
||||
Failed twice before it passed, and the history matters because it defines the
|
||||
shape a valid E03 test has to have.
|
||||
|
||||
1. At the M5 baseline the summary was a single column with no coordinate and
|
||||
survived Undo plus divergence into the active prompt.
|
||||
2. The first M6 implementation made summaries lineage-anchored rows, and the
|
||||
test written for it checked that the **old row** became ineligible. The
|
||||
independent review then found E03 still failing end to end: generation was
|
||||
seeded from `adventures.story_summary`, so the summary produced *on the new
|
||||
line* inherited the abandoned line's prose inside a correctly anchored row.
|
||||
3. The corrective pass seeds generation from `summaries.current`.
|
||||
|
||||
**A valid E03 test must regenerate a summary after the divergence.** Checking
|
||||
only that the old row is ineligible passes while the defect is live. The
|
||||
regression now required is:
|
||||
|
||||
```text
|
||||
path A: enough history for a real summary, sentinel established on it
|
||||
POSITIVE CONTROL — the sentinel is in the path-A summary and prompt
|
||||
move the head below the sentinel, diverge
|
||||
path B: play far enough that a NEW summary is generated
|
||||
prove a new summary row exists and is not path A's
|
||||
prove no path-A action is on path B's lineage
|
||||
prove the sentinel is absent from the new summary
|
||||
prove the sentinel is absent from the complete active prompt
|
||||
prove the old row is retained but ineligible
|
||||
```
|
||||
|
||||
Evidence: `test_e03_a_summary_generated_after_divergence_carries_no_abandoned_content`
|
||||
(deterministic, fails against the pre-corrective implementation); the same
|
||||
sequence against a real local summariser; and a dedicated browser scenario that
|
||||
regenerates a summary after diverging rather than repeating the old blind spot.
|
||||
|
||||
---
|
||||
|
||||
## E04 — Scene State Is Lineage-Safe
|
||||
@@ -960,6 +1005,12 @@ Conduct a multi-turn conversation with Mara.
|
||||
### Pass
|
||||
Narrator remembers immediately preceding dialogue and actions.
|
||||
|
||||
|
||||
### Result — PASS (M6, 2026-09-05)
|
||||
`test_f01_recent_turns_stay_in_the_prompt`: the preceding turns and the reader's
|
||||
own input are present in the assembled prompt, asserted on the context report
|
||||
rather than on the narration.
|
||||
|
||||
---
|
||||
|
||||
## F02 — Old Important Event Retrieval
|
||||
@@ -974,6 +1025,27 @@ Narrator remembers immediately preceding dialogue and actions.
|
||||
### Pass
|
||||
Relevant old clue can be recovered through summary/memory/state.
|
||||
|
||||
|
||||
### Result — PASS (M6 corrective pass, 2026-09-06)
|
||||
|
||||
A distinctive clue is planted, long turns are played over it, and the clue is
|
||||
*absent* from the verbatim history sections and *present* through a retrieved
|
||||
memory, with `history.included < history.total` — so recovery does not come from
|
||||
sending the transcript.
|
||||
|
||||
The independent review found this passing only by a one-slot margin: with a real
|
||||
embedding model, four near-identical filler memories scored 0.79–0.82 against
|
||||
the clue's 0.61, so the clue placed fifth and survived only because the default
|
||||
`memory_top_k` is 5. At 4 it was evicted and F02 failed.
|
||||
|
||||
The corrective pass added redundancy suppression before the final selection.
|
||||
On the same fixture and the same real embedding model, 3 of 5 candidates are now
|
||||
suppressed as repetitions and the clue is retrieved at `top_k` 5, 4 **and** 3.
|
||||
|
||||
Ranking itself remains cosine similarity plus an explicit pin; the further
|
||||
factors `CONTEXT-AND-MEMORY.md` §20 contemplates are not implemented and are
|
||||
recorded there as future work.
|
||||
|
||||
---
|
||||
|
||||
## F03 — Prompt Remains Bounded
|
||||
@@ -986,6 +1058,13 @@ Generate a long story.
|
||||
### Pass
|
||||
Application does not continually append full transcript until context overflows.
|
||||
|
||||
|
||||
### Result — PASS (M6, 2026-09-05)
|
||||
`test_f03_the_prompt_stays_bounded_as_the_story_grows` and
|
||||
`test_the_context_size_stops_growing_once_the_budget_is_reached`. The second
|
||||
measures from a story that already fills the budget, then triples it: the
|
||||
prompt does not move, while the action count does.
|
||||
|
||||
---
|
||||
|
||||
## F04 — Output Token Reserve
|
||||
@@ -995,6 +1074,14 @@ Application does not continually append full transcript until context overflows.
|
||||
### Pass
|
||||
Context builder leaves sufficient room for narrator output and does not regularly fail because input consumes entire context.
|
||||
|
||||
|
||||
### Result — PASS (M6, 2026-09-05)
|
||||
The reply is reserved out of the context budget before history is selected, and
|
||||
the report exposes it (`tokens.output_reserve`). Three tests: the reserve
|
||||
survives a long story; an impossible budget raises `ContextOverflow` naming both
|
||||
figures; and that refusal reaches the reader as a failed turn without disturbing
|
||||
the stored story. Nothing was reserved before M6.
|
||||
|
||||
---
|
||||
|
||||
## F05 — Prompt Inspector
|
||||
@@ -1017,6 +1104,18 @@ User can determine at least:
|
||||
|
||||
Exact UI may vary.
|
||||
|
||||
|
||||
### Result — PARTIAL, complete for the components M6 owns (2026-09-05)
|
||||
`test_f05_the_inspector_shows_every_component_m6_owns` asserts the report
|
||||
carries narrator rules, authoritative state, the summary in use with its source
|
||||
coverage, retrieved memories with authority and provenance, recent history, the
|
||||
reader's input, model and settings, and per-component token costs alongside the
|
||||
budget, the reply reserve and the history allowance. All of this is rendered in
|
||||
the Insights panel and was verified in a real browser.
|
||||
|
||||
"Retrieved knowledge" is M7's imported-document section and is not implemented;
|
||||
nothing was built to fill it.
|
||||
|
||||
---
|
||||
|
||||
## F06 — Retrieval Provenance
|
||||
@@ -1026,6 +1125,13 @@ Exact UI may vary.
|
||||
### Pass
|
||||
A retrieved memory or imported chunk can be traced to its source record/file.
|
||||
|
||||
|
||||
### Result — PASS for story memory (M6, 2026-09-05)
|
||||
`test_f06_a_retrieved_memory_is_traceable_to_its_source`: every retrieved memory
|
||||
carries `branch_id`, `depth` and its source range, and the test resolves that
|
||||
coordinate back to a real action of the campaign's accepted history. Imported
|
||||
chunks are M7's half of this criterion and are not implemented.
|
||||
|
||||
---
|
||||
|
||||
## F07 — Heuristic Memory Is Not Canon
|
||||
@@ -1048,6 +1154,15 @@ Mara is definitely working against Captain Vale.
|
||||
|
||||
as authoritative fact.
|
||||
|
||||
|
||||
### Result — PASS (M6, 2026-09-05)
|
||||
`Memory.authority` is `accepted_story` or `heuristic`, classified by the
|
||||
application rather than the model. The prompt marks an inference `[inferred]`
|
||||
and says such lines are not established fact.
|
||||
`test_f07_a_heuristic_memory_is_labelled_and_is_not_state` also asserts the
|
||||
inference did not become an authoritative fact: retrieval never writes state,
|
||||
and the M5 typed-event path remains the only route to one.
|
||||
|
||||
---
|
||||
|
||||
## F08 — Memory Failure Is Non-Fatal
|
||||
@@ -1060,6 +1175,14 @@ Cause embedding/memory extraction failure if test harness supports it.
|
||||
### Pass
|
||||
Accepted turn persists and story can continue; derived memory may be retried later.
|
||||
|
||||
|
||||
### Result — PASS (M6, 2026-09-05)
|
||||
With the summariser and embedder both raising, the accepted narration, the
|
||||
authoritative state, the head and the transcript all survive, the next turn
|
||||
still plays, and the failure is recorded per kind in `derived_status` and served
|
||||
by `GET /adventures/{id}/derived`. A later healthy run clears it. This is the
|
||||
M2 failure — the whole memory bank dead with a green suite — made visible.
|
||||
|
||||
---
|
||||
|
||||
# G. Imported Knowledge
|
||||
|
||||
+46
-3
@@ -1,8 +1,51 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v2.6
|
||||
- **Revision date:** 2026-09-03
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M4 implemented and accepted**; M5 is next to brief.
|
||||
- **Package:** Adventure Storyteller Planning Package v2.7
|
||||
- **Revision date:** 2026-09-06
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M6 implemented and accepted**; M7 is next to brief.
|
||||
|
||||
## v2.7 — M5 and M6 Closeout (2026-09-06)
|
||||
|
||||
Two milestones, and one lesson they share: an independent review found a real
|
||||
defect in each after the implementation reported success, and in both cases the
|
||||
defect was invisible to the tests the implementation had written for itself.
|
||||
|
||||
**M5 — Genre-Neutral Authoritative Narrative State.** Accepted 2026-09-04. Its
|
||||
review returned *PASS WITH CORRECTIVE WORK REQUIRED*: editing a narrator turn
|
||||
rewound the campaign's live state while the head stayed at the tip, breaking the
|
||||
`transcript position == head == authoritative state` invariant. The corrective
|
||||
pass rebuilt narrator editing on the §§14-15 fork semantics. Recorded here late:
|
||||
M5's closeout did not add a package revision entry, and this one covers it.
|
||||
|
||||
**M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory.** Accepted
|
||||
2026-09-06. Sequence:
|
||||
|
||||
1. **Implementation.** Summaries moved onto lineage-anchored rows; memory
|
||||
authority, provenance, an output reserve, and observable derived-work
|
||||
failure were added. Reported as passing, including E03.
|
||||
2. **Independent review — E03 still failed.** Summary *rows* were anchored, but
|
||||
generation was seeded from `adventures.story_summary`, a campaign-global
|
||||
column with no lineage. After a divergence the summariser was handed the
|
||||
abandoned line's prose and asked to update it, so the new summary carried
|
||||
abandoned content inside a correctly anchored row. The review also found F02
|
||||
passing by a single retrieval slot, two test defects, and a misleading
|
||||
derived-work status.
|
||||
3. **Corrective pass.** Generation is seeded from `summaries.current`;
|
||||
redundancy suppression was added before the retrieval cut; the test fixture
|
||||
no longer makes real network calls; the browser suite gained a scenario that
|
||||
regenerates a summary after divergence.
|
||||
|
||||
**What the planning package learned from M6, recorded in the active documents:**
|
||||
|
||||
- `CONTEXT-AND-MEMORY.md` §11 — lineage safety has two halves. Anchoring the
|
||||
output row is not enough; the *input* to the summarizer must be scoped by the
|
||||
same rule.
|
||||
- `V1-ACCEPTANCE-TESTS.md` E03 — a valid test must regenerate a summary after
|
||||
diverging. Checking only that the old row went ineligible passes while the
|
||||
defect is live.
|
||||
- `CONTEXT-AND-MEMORY.md` §20/§22 — what M6 implements (redundancy suppression
|
||||
between memories) is now distinguished from what it does not (importance,
|
||||
entity, recency and thread ranking; cross-layer duplication).
|
||||
|
||||
## v2.6 — M4 Closeout (2026-09-03)
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user