Files
JesseMarkowitzandClaude Opus 5 0c1ba836ba v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 11:21:53 -04:00

1487 lines
43 KiB
Markdown

# Adventure Storyteller — Context and Memory
**Status:** v1.0 — aligned to Phase 0B findings
**Purpose:** Define what information is supplied to the narrator on each turn, how long-term memory works, and how authority, lineage, retrieval, summaries, and imported knowledge interact.
## 1. Design Goal
The language model should never be expected to remember an entire long-running story by itself.
The application should construct a bounded context for every narrator turn using:
- durable narrator rules,
- campaign canon,
- current authoritative state,
- branch-safe summaries,
- relevant older story memories,
- relevant imported local knowledge,
- recent active-lineage turns,
- the user's current input.
The full campaign history remains stored locally even when only a subset is sent to Ollama.
The core rule is:
> The application decides what the narrator is allowed to rely on; the model does not decide what counts as canon.
## 2. Context Layers
The narrator context should be assembled from distinct layers.
Recommended order:
```text
1. System / narrator rules
2. Campaign profile
3. Authoritative canon and world rules
4. Current authoritative narrative state
5. High-level campaign / arc summary
6. Relevant older story memories
7. Relevant imported local knowledge
8. Recent active-lineage turns
9. User's current input
```
Not every layer must appear on every turn.
Each layer should have:
- explicit authority,
- source provenance,
- token budget,
- branch/lineage rules where applicable.
## 3. Authority Levels
The narrator must distinguish between authoritative and suggestive information.
Recommended authority hierarchy:
### Level 1 — Explicit Campaign Canon
Highest authority.
Examples:
- FTL does not exist.
- Mara is Edrin's sister.
- Magic cannot resurrect the dead.
- The story takes place in 1892.
Sources:
- campaign setup,
- user-authored canon documents,
- manual canon corrections.
If narration conflicts with Level 1 canon, Level 1 wins.
### Level 2 — Accepted Story Facts / Current State
Facts established by accepted active-lineage story history.
Examples:
- Aldric currently possesses the silver key.
- Mara has already met the protagonist.
- The eastern bridge was destroyed.
- The Persephone is docked at Ceres Station.
These are authoritative unless later invalidated or corrected.
### Level 3 — Accepted Historical Events
Important events that occurred earlier on the active lineage.
Examples:
- Mara warned Aldric not to trust Captain Vale.
- The crew discovered a signal beneath Europa's ice.
- Aldric promised to return before sunrise.
These are historical truth for the active story.
### Level 4 — Derived / Heuristic Memory
Useful inferred information that may help continuity but must not be treated as hard canon.
Examples:
- Mara seemed nervous when Captain Vale arrived.
- Aldric probably distrusts the city guard.
- The abandoned station may be dangerous.
The prompt should identify these as inference or interpretation.
### Level 5 — Imported Reference Material
Supporting local information.
Examples:
- medieval tavern construction,
- orbital mechanics reference notes,
- technical description of fusion drives,
- historical clothing references.
Reference material may guide detail and plausibility but does not override campaign canon.
### Level 6 — Imported Inspiration Material
Lowest authority.
Examples:
- public-domain fantasy passages,
- science-fiction stories,
- descriptive prose samples,
- atmosphere/style excerpts.
Inspiration can influence tone, imagery, pacing, or ideas.
It must never be treated as proof that something exists in the current story world.
## 4. Authority Conflict Rule
When two context items conflict:
```text
higher authority wins
```
Example:
Campaign Canon:
```text
FTL travel does not exist.
```
Imported Reference:
```text
A fictional source describes warp drives.
```
Narrator behavior:
```text
Do not introduce warp drive as established technology.
```
The reference may still inspire descriptive language if relevant, but cannot override canon.
## 5. Pretrained Model Knowledge
The local language model contains pretrained knowledge that cannot be erased.
The application should instruct the narrator:
> Pretrained knowledge may help with language, general plausibility, and invention, but it is not authoritative story canon.
The narrator must not silently import:
- named characters,
- locations,
- technologies,
- factions,
- magic systems,
- plot facts
from unrelated outside works unless the campaign context explicitly establishes them.
The application cannot guarantee perfect suppression of pretrained knowledge, but it can make authority boundaries explicit and inspectable.
## 6. Current Authoritative State
Every turn should include the minimum current state necessary for continuity.
Potential categories:
- current location,
- current scene,
- characters present,
- active relationships,
- possessions,
- important conditions,
- active story threads,
- unresolved facts,
- current organization/faction relationships,
- world constraints relevant to the scene.
The state context should be concise and structured.
Do not dump the entire database into every prompt.
## 7. State Selection
State should be selected based on relevance.
Always include:
- protagonist identity,
- current location,
- current scene,
- key current conditions,
- globally critical canon rules.
Conditionally include:
- nearby characters,
- relevant items,
- related factions,
- thread-specific facts,
- location-specific rules,
- technology/magic constraints relevant to the action.
## 8. Recent History
Recent active-lineage turns should normally be included verbatim.
Purpose:
- local conversational continuity,
- dialogue coherence,
- immediate action continuity,
- writing rhythm.
Recommended policy:
- include as many recent turns as fit within the recent-history token allocation,
- prefer complete turn boundaries,
- never include abandoned/disposable future history,
- preserve speaker/role metadata.
The exact number of turns should be token-based rather than fixed.
## 9. Rolling Summary
Older active-lineage material should be compressed into summaries.
A rolling summary should contain:
- major events,
- current goals,
- important discoveries,
- relationship changes,
- unresolved threads,
- durable consequences.
A summary should not preserve every stylistic detail.
Important rule:
> A summary is derived data, not authoritative history.
The original transcript remains the source of truth.
## 10. Summary Scope
Potential summary levels:
### Campaign Summary
Very compressed overview of the story so far.
### Arc / Chapter Summary
More detailed representation of a recent story segment.
### Turn-Range Summary
Derived from a bounded range of turns.
Recommended v1:
- one campaign-level rolling summary,
- optional turn-range/chapter summaries if inherited architecture supports them cleanly.
## 11. Summary Lineage Safety
Every summary must be associated with the history it summarizes.
If the user restores or diverges before part of that source history:
- the invalid portion must not be reused,
- unaffected ancestral summaries may remain valid,
- new summaries should be created for the new continuation.
A summary from abandoned history must never leak into active context.
### As implemented (M6)
Lineage safety here has **two** halves, and the M6 review found that having only
the first is not enough.
**The row must be eligible.** A summary is a row with a coordinate, exactly as a
memory is: `branch_id` and `depth` name the last node it covers,
`source_start`/`source_end` the stretch. Eligibility is one question — is that
coordinate on the active, head-capped lineage? — answered by
`lineage.Path.clause`, the same chokepoint every read of the story goes through.
Undo, Redo, Save Point restore and divergence all fall out of that without a
rule of their own.
**The input must be eligible too.** Summary generation is seeded only from a
summary that is itself valid on the current head-capped lineage
(`summaries.current`). Where none is, the new line starts from no previous
summary.
The second half was missing in the first M6 implementation and E03 failed
because of it: the summariser seeded itself from `adventures.story_summary`, a
campaign-global column with no lineage, so after a divergence it was handed the
abandoned line's prose and asked to update it. The row it produced was correctly
anchored to the new branch and therefore *looked* lineage-safe while its
sentences described a story the reader had left. Anchoring the output is not
enough; the input has to be scoped by the same rule.
Abandoned summaries are retained, never deleted, and become eligible again if
the reader returns to the line that produced them.
`adventures.story_summary` remains, as a reader-facing convenience only: the
Plot panel edits it and the export bundle carries it. It is a mirror of whichever
summary is currently eligible — kept in step when one is written and when the
head moves — and **nothing authoritative may read it**. A summary the reader
types is recorded as a row anchored at the position they typed it at, so it
behaves like any other.
## 12. Story Memory
Long-term story memory should retrieve important older information that is not present in recent history or current summary.
Examples:
- a promise made 80 turns ago,
- a minor character encountered much earlier,
- the origin of an item,
- a clue from a distant chapter,
- a prior argument between two characters.
Memory exists to restore specific detail that broad summaries may omit.
## 13. Memory Types
Recommended memory categories:
### Event Memory
Something happened.
Example:
```text
Turn 42: Mara hid a letter beneath the hearthstone.
```
### Character Memory
Important information about a character.
Example:
```text
Captain Vale strongly dislikes being touched unexpectedly.
```
### Relationship Memory
A meaningful interaction or relationship change.
Example:
```text
Aldric broke his promise to Mara.
```
### Discovery Memory
A clue or learned fact.
Example:
```text
The silver key bears the same symbol as the old abbey crypt.
```
### Promise / Commitment Memory
Future-relevant obligation.
Example:
```text
Aldric promised to return before dawn.
```
### Location Memory
Important prior detail about a place.
### Heuristic Memory
Interpretive information that may be useful but is not hard canon.
## 14. Memory Authority
Every memory should carry an authority classification.
Examples:
```text
accepted_story
current_state
heuristic
```
The narrator should be told which memories are:
- factual,
- inferred,
- uncertain.
This avoids turning guesses into canon.
### As implemented (M6)
Two values, `accepted_story` and `heuristic`, on `Memory.authority`. The
**application** classifies, not the model: `memorybank.classify_authority`
reads the memory's own text for hedging ("seemed", "appeared to", "probably"),
so an extractor cannot promote a guess by asserting it confidently. The prompt
marks a heuristic memory `[inferred]` and says in words that such lines are
interpretation rather than established fact.
Retrieval never writes state. A memory of either authority is something the
narrator is shown; the only path to an authoritative change remains the M5 typed
event pipeline (ADR 013).
## 15. Memory Creation
Memories may be generated after accepted turns.
Potential pipeline:
```text
accepted turn
|
v
memory extractor
|
v
candidate memories
|
v
application validation / classification
|
v
stored memory records
```
Not every turn needs a permanent memory.
The memory system should favor:
- importance,
- future usefulness,
- uniqueness,
- continuity relevance.
### As implemented (v1.1 WP-B.2): what the summariser is shown
A memory is written from one block of `MEMORY_INTERVAL` (6) story actions. The
summariser is given the cast brief, then the block, inside a budget of
`MEMORY_EXCERPT_TOKENS` (2,000).
- **A block that fits** is sent whole, exactly as v1.0.0 sent it.
- **A longer block** was cut to its last 2,000 tokens in v1.0.0, so a fact near
its start never reached the summariser (WP-B.1). It is now sent as its
opening and its end, in order, with a visible marker between them
(`EXCERPT_OMISSION_MARKER`, "[… the middle of this stretch of story is left
out here …]"). The marker and its blank lines are paid for first, and the rest
is halved, the odd token going to the end: 992 + 993 + 15 = 2,000 tokens. The
rejoined text is measured, and the opening gives up tokens until the whole is
within budget.
- **The marker is never stored.** `summarize_block` removes it from anything the
model repeats back.
- **Limit.** A fact in the middle of a very long block is still left out. The
input stays bounded; it is not a summary of everything.
Existing memories are not rewritten. `tools/rewrite_memories.py`, which is
opt-in, uses the same function.
### As implemented (v1.1): what a memory can be relied on to keep
The memory prompt is v1.0.0's, unchanged. v1.1 changed what the summariser is
shown (above), not what it is told.
**Known limitation, accepted for v1.1.** The application now delivers the whole
relevant block to the summariser, and keeps, ranks and injects the memory it
writes (§18, §20, §21). But the reference summariser, `qwen2.5:3b-instruct-16k`,
can still be given a block that states a distinctive fact and write a memory
that:
- omits the fact, or the specific objects in it;
- attributes it to the wrong character;
- prefers the generic narration that follows it.
So **independent recovery from memory is proven for the application's
mechanisms, and is not guaranteed with the reference model.** A bounded prompt
change aimed at this was tried and rejected
(`reports/v1.1/V1.1-WP-B2-REPORT.md` §T). A stronger dedicated summariser,
structured fact extraction, or separate factual and narrative memory are
future options, and none is implemented.
## 16. Memory Retrieval
Retrieval should be local.
Potential mechanisms:
- lexical search,
- semantic embeddings,
- hybrid search.
Current preference:
> Hybrid local retrieval if practical.
Reason:
- lexical retrieval is transparent and precise for names/terms,
- semantic retrieval is useful for conceptually related old events.
Phase 0B confirmed useful local semantic memory in AI-DnD and deterministic lexical lore in ai-adventure. The selected production direction is hybrid local retrieval, implemented incrementally with a lexical path that remains usable when semantic embeddings are unavailable.
## 17. Embeddings
If semantic retrieval is used:
- embeddings must be generated locally,
- preferably through Ollama,
- no remote embedding API,
- embeddings are derived data,
- embeddings must be rebuildable.
Potential local embedding model should be selected later based on actual hardware and model quality.
## 18. Memory Retrieval Query
The retrieval query may include:
- current user input,
- current scene,
- active story thread names,
- entities mentioned,
- current location,
- current goals.
Do not rely only on raw user input.
Example:
User:
```text
I ask Mara whether she recognizes the symbol.
```
Retrieval query may include:
```text
Mara
symbol
silver key
abbey crypt
prior discoveries
```
### As implemented (v1.1 WP-B.2)
v1.0.0 embedded the newest four actions cut to 600 tokens, so the player's
one-line question arrived after three turns of narration and barely moved the
embedding (WP-B.1: a direct question's cosine fell from 0.708 alone to 0.241 in
that query). The query is now two short texts, embedded in one call:
| Component | What it is | Bound |
| --- | --- | --- |
| **input** | the player's own action this turn (`do`, `say` or `story`) | last 200 tokens |
| **context** | the scene from the authoritative state (its summary, the location's name, the names of who is present), then the end of the newest narration | 60 + 120 tokens |
- A continue turn, and the Insights dry run, have no input; the context alone is
searched.
- A retry searches with the input being retried; the discarded attempt is not in
the context.
- The full entity list, threads and older narration are deliberately left out,
so a long scene or a large cast cannot outweigh the question by length.
- What was searched for is recorded per turn in the context snapshot
(`memories.query`: input, context, input terms, weights).
## 19. Memory Retrieval Filtering
Before ranking memories, filter by:
- campaign,
- active lineage,
- allowed authority,
- source validity,
- enabled status.
Never retrieve memories from:
- abandoned future paths,
- deleted campaigns,
- unrelated campaigns.
## 20. Memory Ranking
Potential ranking factors:
- semantic similarity,
- lexical match,
- recency,
- importance,
- entity overlap,
- story-thread overlap,
- authority,
- explicit user pinning.
The final ranking formula should be simple and inspectable.
### As implemented (M6)
Ranking is **cosine similarity against the retrieval query, plus an explicit
pin**. The other factors listed above — importance, lexical match, recency,
entity overlap, story-thread overlap — are **not implemented**, and remain
future work rather than something M6 delivered.
What M6 does implement, because similarity alone proved insufficient, is
redundancy suppression before the final selection: see §22.
### As implemented (v1.1 WP-B.2)
One transparent lexical term is added to similarity. For each eligible memory:
```text
semantic_score = 0.6 * cos(input, memory) + 0.4 * cos(context, memory)
(either cosine alone when the other text is empty)
lexical_score = sum of w(t) over the input's terms the memory holds
/ sum of w(t) over all the input's terms in [0, 1]
w(t) = ln((N + 1) / (df(t) + 1)) N eligible memories, df holding t
final_score = semantic_score + 0.15 * lexical_score
```
- **Terms** are the knowledge path's tokenizer and stop list (`knowledge.fts`),
with possessives dropped and a plural `s` folded. No stemmer, no dependency.
- **Rarity** is computed per turn over the eligible candidates only. There is no
index and no stored field. A word every candidate holds (a protagonist's
name) weighs 0; a word the question shares with one memory weighs most.
- **Only the player's input** is matched lexically, never the context.
- **The weight** was chosen by sweep over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5:
0.15 is the smallest at which the lexical term alone lifts the planting-era
memory into `memory_top_k` against the v1.0.0 query, while no rare-word
negative control lets an unrelated memory pass a semantically relevant one.
A memory can gain at most 0.15 from wording, so it cannot pass one more than
0.15 ahead of it in meaning.
- **Pins** are unchanged: always used, counted toward `memory_top_k`.
- **Ties** on the final score are broken by memory id.
- **Provenance.** Each used memory records `semantic_score`, `lexical_score` and
`final_score`; `similarity` keeps its v1.0.0 meaning, the semantic score, so
the inspector's "closeness" is unchanged.
Importance, recency, entity overlap and story-thread overlap remain
unimplemented. Redundancy suppression (§22) is unchanged and runs over the final
order.
## 21. Memory Budget
Retrieved memories should have a bounded token budget.
Recommended behavior:
- retrieve more candidates than will be used,
- rerank locally,
- include only the highest-value items that fit,
- preserve source IDs for inspection.
### As implemented (v1.1 WP-B.2): selection and the bank's capacity
**Selection** is unchanged in shape: every eligible, embedded memory on the
active lineage is scored (§20), pinned memories are taken first, the rest fill
`memory_top_k` (default 5) best first, skipping repeats (§22). The Memories
section is priced into the protected context like any other live section, so
its budget is unchanged (F03). There is still no relevance floor: a full
`memory_top_k` is used whenever the bank holds that many.
**Capacity** (`memory_bank_capacity`, default 80) is unchanged. Eviction still
runs after each post-turn pass, over the whole adventure rather than one
lineage, and marks rows `forgotten` rather than deleting them. What changed is
the order (`memorybank.eviction_order`), because least-recently-used alone
discarded the only memory of an early stretch first (WP-B.1):
- **Coverage signal.** Memories with a source range say which stretch they
describe. Each is judged by the hole its removal would leave between the end
of the memory before it and the start of the memory after it. The smallest
hole goes first, so the bank thins where it is densest. A memory whose start
another memory shares leaves no hole.
- **Boundaries.** The earliest and the latest memory by position are not
coverage candidates: they are the only records of the opening and of the most
recent stretch. This also keeps a memory written this turn from being evicted
by the pass that wrote it (the frozen bank).
- **Recency signal.** Among equal holes, the least recently used goes first
(`coalesce(last_used_at, created_at)`), then the less used, then the lower id.
- **Pinned rows** are never taken, and count as coverage.
- **Fallback.** Memories with no range (typed by the player, or migrated) and
boundaries are taken least recently used first, as in v1.0.0, once no coverage
candidate remains. The bank stays bounded either way; only an all-pinned bank
may exceed capacity.
Measured on banks where nothing is ever retrieved, the kept bank starts at the
opening and its largest uncovered stretch stays within about 1.3 times the
average spacing (story length / capacity). The v1.0.0 order kept only the newest
stretch. The rule reads no text and no vectors, and lineage eligibility is
unaffected: it decides only `forgotten`.
## 22. Duplicate Suppression
Do not include the same fact repeatedly through:
- current state,
- summary,
- memory,
- imported canon.
If the current state already says:
```text
Aldric possesses the silver key.
```
there is little value in also including three memories that say the same thing.
Context builder should prefer the highest-authority concise representation.
### As implemented (M6)
**Between memories, yes.** Retrieval walks the ranked candidates and skips one
that repeats a memory already chosen, keeping the highest-ranked statement of a
fact and its provenance. Two rules bound it: authority is never crossed, so an
inference can never suppress a record or the reverse; and the bar is high, set
where measurement showed distinct facts stop appearing. The number of
suppressed candidates is reported in the context record, so a memory that was
considered and set aside can be told from one that was never eligible.
The threshold is a measured property of the embedding model in use, not a
universal constant, and `memorybank.REDUNDANT_SIMILARITY` records the
measurement beside the value.
The same measurement ruled out the more obvious test. Word overlap fires hardest
on exactly the pair that must not be merged — "Mara promised to return before
dawn" against "Aldric promised to return before dawn" shares most of its words
and means something else — and is weakest on filler that plainly repeats itself.
Wording is a poor proxy for sameness of fact.
**Across layers, not yet.** A fact can still appear in the authoritative state,
in a memory and in recent history at once. That is bounded and legible, but it
is not the "highest-authority concise representation" this section asks for, and
it remains open.
## 23. Imported Knowledge Categories
Imported local files must be classified as:
### Canon
Authoritative campaign truth.
### Reference
Supporting factual/descriptive material.
### Inspiration
Optional creative influence.
The classification must be visible and editable by the user.
## 24. Imported Canon
Imported Canon should be treated similarly to manually entered campaign canon.
Examples:
- a setting bible,
- technology rules,
- faction descriptions,
- map/location notes,
- character bible.
Imported Canon may be retrieved selectively rather than fully included every turn.
## 25. Imported Reference
Reference material supports plausibility or detail.
Examples:
- medieval medicine,
- orbital mechanics,
- 19th-century railroad practice,
- astronomy notes.
It should be labeled in context as reference, not story truth.
## 26. Imported Inspiration
Inspiration should be optional and low-authority.
Potential behavior:
- retrieve only when enabled,
- use a small budget,
- prefer scene/style relevance,
- never allow inspiration to override canon.
## 27. Source Provenance
Every retrieved imported chunk should retain:
- source file ID,
- source title,
- classification,
- chunk ID,
- retrieval score/method.
The prompt inspector should be able to show:
```text
Reference used:
Orbital Mechanics Notes.md
Chunk 17
```
## 28. User-Pinned Knowledge
The user should eventually be able to force certain knowledge into context.
Examples:
- always include this world rule,
- include this character note for the next scene,
- pin this reference document temporarily.
For v1, campaign canon rules may serve as the primary pinned mechanism.
A general pinning UI may be deferred.
## 29. Context Budget
The application must explicitly budget tokens.
Conceptual allocation:
```text
System / narrator rules fixed reserve
Campaign canon protected
Current state protected
Summary medium reserve
Retrieved story memories bounded
Imported knowledge bounded
Recent history elastic
Current user input protected
Output generation reserve protected
```
Exact percentages should be configurable or derived from model context size.
### As implemented (M7, for imported knowledge)
The knowledge budget is a share of what is left after everything protected and
the reply reserve are subtracted, and it is spent in authority order:
```text
always-included Canon protected. Counted with the system block, before any
history is chosen, and capped at 20% of the whole
context budget. If it cannot fit alongside the other
protected sections and the reply reserve, the turn fails
with `ContextOverflow` rather than sending a prompt
known to overflow. What does not fit is reported as
dropped, with its token cost.
retrieved knowledge 33% of what is left, filled Canon first, then Reference
(capped at half the knowledge budget), then Inspiration
(capped at a quarter). Whatever is not spent returns to
the story history rather than being lost.
```
So Reference and Inspiration cannot crowd out retrieved Canon, and none of the
three can reach the current authoritative state, the reader's input, the narrator
rules, critical Canon or the output reserve — all of which are priced before the
knowledge budget exists.
Every included passage's token cost is in the context report, and so is every
passage there was no budget for.
## 30. Protected vs Elastic Context
### Protected
Should not be dropped casually:
- system rules,
- critical canon,
- current state,
- current user input,
- output token reserve.
### Elastic
Can be reduced when context is tight:
- old recent-history turns,
- lower-ranked memories,
- reference material,
- inspiration material,
- verbose summaries.
## 31. Context Reduction Order
When the prompt is too large, recommended removal order:
1. lowest-ranked inspiration chunks,
2. lowest-ranked reference chunks,
3. lowest-ranked heuristic memories,
4. low-importance accepted memories already reflected elsewhere,
5. oldest recent-history turns,
6. compress or shorten summaries,
7. trim noncritical state detail.
Do not drop:
- critical narrator rules,
- hard campaign canon needed for the scene,
- current user input,
- core current state.
## 32. Output Reserve
The context builder must leave space for narrator output.
It should not fill the entire model context window with input.
Recommended behavior:
- reserve a configurable maximum response budget,
- add safety margin,
- fail gracefully if protected context alone is too large.
### As implemented (M6)
`context_token_budget` is the whole window, so the reply is subtracted from it
before any history is selected:
```text
available for history = context_token_budget
- (system, canon, state, summary, memories, user input)
- (max_output_tokens + 64)
```
The 64-token margin covers the separators added after budgeting and the drift
between the app's tokenizer and the serving model's; it is fixed rather than
proportional because what it absorbs does not scale with the budget.
When the protected part alone does not fit, `build_context` raises
`ContextOverflow` naming both figures and what to change, and the turn reports
that as a failed turn. It does not build a prompt it knows will overflow.
Before M6 nothing was reserved: the builder spent the entire budget on input and
left the reply to fit in whatever the endpoint had left.
## 33. Context Snapshot
Every narrator generation should preserve enough information to reconstruct the effective prompt.
At minimum record:
- context component IDs,
- rendered text or reproducible form,
- token counts,
- ranking scores,
- model/settings,
- active branch/head.
This allows debugging.
## 34. Prompt Inspector
The UI should eventually allow the user to inspect:
- narrator rules,
- current state,
- summary,
- retrieved memories,
- imported knowledge,
- recent history,
- total token usage.
This is particularly important when the narrator behaves unexpectedly.
## 35. User Override
The user should be able to manually correct context-driving data.
Examples:
- fix canon,
- disable a bad memory,
- disable a knowledge source,
- correct a character record,
- mark a heuristic memory as wrong.
The system should not force the user to manipulate raw embeddings or SQL.
## 36. Memory Correction
If a stored memory is wrong:
Potential actions:
- delete/disable memory,
- downgrade authority,
- correct text,
- replace with authoritative fact.
Corrections should preserve provenance when practical.
## 37. Memory vs Fact
Important distinction:
### Fact
Structured authoritative assertion.
### Memory
Retrievable narrative representation.
Example:
Fact:
```text
Mara knows the location of the key.
```
Memory:
```text
During the tavern conversation, Aldric accidentally revealed where the key was hidden.
```
The memory may provide richer narrative context.
The fact provides concise state authority.
## 38. Memory vs Summary
### Summary
Broad compression of a range of story history.
### Memory
Specific retrievable detail.
They solve different problems and should coexist.
## 39. Memory vs Imported Knowledge
### Story Memory
Comes from this campaign's accepted history.
### Imported Knowledge
Comes from user-supplied local material.
Story memory must be lineage-aware.
Imported knowledge is normally campaign-wide and not branch-specific.
## 40. Knowledge Retrieval Query
Imported-knowledge retrieval may use:
- current scene,
- entities,
- user input,
- active story threads,
- campaign genre/profile.
Example:
Scene:
```text
The crew is approaching Europa.
```
Potential reference retrieval:
- radiation environment,
- orbital dynamics,
- ice crust,
- local campaign technology rules.
## 41. Canon Retrieval
Canonical material should not rely solely on similarity search.
Critical canon rules may need:
- always-on inclusion,
- entity-linked retrieval,
- tag-based retrieval,
- explicit rule triggers.
Example:
```text
FTL does not exist.
```
This should not disappear just because the current user input does not semantically resemble "FTL".
### As implemented (M7)
A Canon source may be marked `always_include`. Its passages are supplied on every
turn whatever the scene is, in their own protected section framed as standing
rules of the world. The flag is **Canon's alone** — it bypasses relevance
entirely, and asserting unranked Reference on every turn would spend a protected
budget on material that establishes nothing — and it is enforced both ways: a
source reclassified away from Canon loses the flag.
Always-included Canon does not set the relevance floor for the passages that had
to earn their place, because it did not earn its own; letting it do so would let
one standing rule silence everything the scene actually turned up.
Entity-linked and tag-based retrieval are **not** implemented. Entity names do
reach the query — it is built partly from the authoritative state, so the
characters and places in play are among the search terms — but there is no
explicit link from a source to an entity, and no tags. Deferred.
## 42. Global Canon
Some canon should always be active.
Examples:
- setting era,
- hard technology constraints,
- magic existence/nonexistence,
- narrator/player-control rules,
- protagonist identity.
Global canon should remain small.
## 43. Conditional Canon
Other canon may be retrieved when relevant.
Examples:
- details of a distant city,
- a faction's internal structure,
- a specific ship subsystem,
- an NPC's background.
## 44. Character Knowledge
Potential future refinement:
The narrator may need to distinguish:
- objective world truth,
- what the protagonist knows,
- what an NPC knows.
This is useful for secrets and mystery stories.
For v1, support at least:
- objective canon/state,
- optional `knows` relationships/facts where important.
A full separate-mind simulation like Sonder Engine is not required.
## 45. Secrets
Secrets should not automatically be shown to the player-facing prose merely because they exist in authoritative state.
The narrator can know secrets necessary to run the story.
The prompt architecture may need separate labels such as:
```text
GM-only canon
player-known facts
character-known facts
```
This should be considered in v1 if the selected base supports it cheaply.
## 46. Spoiler-Safe Context
The application must distinguish:
> narrator knowledge
from:
> text that should be revealed to the user.
The narrator may receive hidden information while being instructed not to reveal it until narratively appropriate.
This is a prompt discipline requirement.
### As implemented (M7)
Source-level, and treated as prompt discipline exactly as this section says. A
source marked `hidden` is retrieved and supplied to the narrator like any other,
and two things mark it: the passage itself carries `[narrator only]` on its
provenance line, and the knowledge rule in the system block says what that means
— the protagonist does not know it, must not be told it, must not act on it, and
a direct question about it is answered from what the protagonist actually knows.
The marker travels on the passage rather than only in the preamble because a
passage is read where it sits. Per-chunk visibility is deferred (§69 of
`IMPORTED-KNOWLEDGE-DESIGN.md` asks only for source level in v1).
"Hidden" is about the protagonist, not about the person running the campaign:
the source is fully readable in the knowledge panel.
## 47. Story Style Memory
Some user preferences may be durable within a campaign:
- preferred prose length,
- dialogue density,
- violence level,
- descriptive richness,
- pacing,
- point of view.
These should live in campaign configuration rather than be inferred repeatedly from history.
## 48. Temporary Direction
User out-of-character direction may be:
- one-turn only,
- scene-level,
- durable campaign guidance.
The UI should eventually distinguish these.
Example:
```text
For this scene, keep the pacing tense and fast.
```
should not necessarily become permanent campaign canon.
## 49. Memory Extraction Timing
Possible strategies:
### Every turn
Simple but potentially expensive.
### Threshold/batch
Extract after several turns.
### Selective
Only when state extractor flags something important.
Recommended initial approach:
> Reuse the selected base's proven mechanism if it is local and correct; otherwise perform lightweight extraction after each accepted turn and allow later optimization.
## 50. Summary Generation Timing
Possible triggers:
- token threshold,
- number of turns,
- scene/chapter boundary,
- manual request.
Recommended:
- automatic token/turn threshold,
- preserve source lineage.
## 51. Failure Handling
If memory extraction fails:
- accepted story turn remains valid,
- no corrupted memory should be stored,
- retry can occur later.
If summary generation fails:
- story continues,
- older direct history may temporarily remain longer,
- failure must not corrupt authoritative state.
If embedding generation fails:
- lexical retrieval should remain possible if implemented.
Derived-memory failures must not block story persistence.
**As implemented (M11).** Two more rules, both learned from a long run in which
every turn was accepted while the memory bank quietly stopped filling:
- **A failure must be recordable.** The record is itself a write. It is made after
the failed session has been rolled back, so the status the reader sees
(`derived_status`) never reads `idle` over a failure.
- **The turn must not cause it.** A turn that held SQLite's write lock through the
model call locked every post-turn write out for the length of the reply. Nothing
in a turn writes before the model call (`TECHNICAL-DESIGN.md`, "Background
failure observability").
## 52. Offline Operation
All context and memory operations must work locally.
Permitted v1 data flow:
```text
Local browser
-> local application
-> local SQLite/files
-> local Ollama
```
No:
- remote vector DB,
- remote embedding API,
- cloud search,
- automatic web retrieval.
## 53. Context Construction Example
User input:
```text
I ask Mara whether she recognizes the symbol on the key.
```
Possible assembled context:
```text
SYSTEM
You are the narrator...
Do not override canon...
GLOBAL CANON
Magic is rare.
The dead cannot be resurrected.
CURRENT STATE
Location: Crooked Lantern Tavern
Aldric possesses the silver key.
Mara is present.
Mara trusts Aldric cautiously.
STORY SUMMARY
Aldric is searching for Edrin...
RELEVANT STORY MEMORIES
[Accepted event] Turn 38: Edrin's desk contained the silver key.
[Accepted discovery] Turn 51: The key bears a symbol matching the abbey crypt.
[Heuristic] Mara appeared uneasy when the abbey was mentioned.
REFERENCE
Source: Abbey Notes.md
The symbol is historically associated with...
RECENT HISTORY
...
USER
I ask Mara whether she recognizes the symbol on the key.
```
## 54. Context Construction Example — Science Fiction
User:
```text
Can the Persephone reach Europa before the storm hits?
```
Context:
```text
GLOBAL CANON
FTL does not exist.
Persephone uses a fusion torch drive.
CURRENT STATE
Persephone is departing Ceres Station.
Fuel reserve: established as limited.
Crew has detected a radiation storm.
REFERENCE
Orbital Mechanics.md
Relevant local transfer notes...
STORY MEMORY
Turn 112: Chief Engineer stated maximum sustained acceleration...
RECENT HISTORY
...
USER
Can the Persephone reach Europa before the storm hits?
```
## 55. What Should Not Be Sent
Avoid routinely sending:
- entire campaign transcript,
- all entities,
- all imported documents,
- abandoned branch memories,
- irrelevant character biographies,
- duplicate facts,
- old low-value heuristics,
- raw embeddings,
- internal database metadata not useful to narration.
## 56. Context Debugging
If the narrator makes an unexpected choice, the system should support questions such as:
- Which memory caused this?
- Which canon rule was present?
- Was the relevant fact omitted?
- Did an abandoned branch leak in?
- Did an inspiration passage overpower canon?
- Was the recent-history window too short?
This is why context provenance is a first-class requirement.
## 57. Phase 0B Findings Applied
The selected AI-DnD base was exercised with real local Ollama embeddings and demonstrated useful lineage behavior:
- branch-scoped memory retrieval worked,
- a memory from a later/abandoned depth was excluded after moving the active head backward,
- the same memory became eligible again after Redo/return to the applicable lineage,
- switching to a different branch prevented abandoned-branch terms from appearing in the assembled prompt,
- the context/Insights path exposes labeled prompt sections and token costs.
Planning consequences:
1. Preserve AI-DnD's common lineage filtering/chokepoint rather than replacing memory from scratch.
2. Make summaries explicitly lineage/turn-range anchored; never rely on a positional message watermark.
3. Keep accepted story memory distinct from heuristic/inferred memory.
4. Imported knowledge remains a separate subsystem; do not promote Story Cards into the knowledge store merely because they are prompt-injection primitives.
5. Use local Ollama embeddings for semantic story memory where enabled; retain a lexical/deterministic path for imported knowledge.
6. Context/state extraction must be tested at realistic prompt length because Phase 0B showed model protocol adherence can degrade under full application context.
7. Derived memory/summary failure must not corrupt or roll back an otherwise valid authoritative story commit.
## 58. Acceptance Criteria
The final implementation must satisfy:
- full campaign history remains stored even when not in context,
- only active-lineage story history influences the narrator,
- campaign canon outranks all other context,
- accepted story facts outrank heuristic memory,
- reference material cannot override canon,
- inspiration cannot silently become canon,
- old important events can be retrieved beyond the recent-history window,
- retrieval works locally,
- no remote embeddings/search are required,
- abandoned branch memories do not leak,
- moving the active head backward excludes memories derived after that head,
- Redo/returning to the valid lineage can make those memories eligible again,
- context remains bounded,
- output space is reserved,
- prompt composition is inspectable,
- retrieved sources retain provenance,
- summaries are lineage-safe,
- memory failures do not corrupt authoritative story state.
## 59. Selected Context / Memory Design
Use a layered, authority-aware context builder:
```text
PROTECTED
|
v
System / Narrator Rules
+
Global Canon
+
Current State
|
v
----------------
+
Lineage-Anchored Summary
+
Relevant Story Memories
+
Relevant Local Knowledge
+
Recent Active-Lineage Turns
+
Current Input
|
v
OLLAMA
```
Retrieval must be:
```text
local
+ lineage-aware where derived from story history
+ provenance-preserving
+ authority-aware
+ token-bounded
```
The application treats context construction as a deterministic, independently testable subsystem. AI-DnD's lineage-aware memory implementation is the starting point; the project's own authority and imported-knowledge rules define the target behavior.