v1.1 WP-B.2: independent long-term memory retention

Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-15 11:21:53 -04:00
co-authored by Claude Opus 5
parent beb17ada10
commit 0c1ba836ba
15 changed files with 3782 additions and 186 deletions
+138
View File
@@ -451,6 +451,50 @@ The memory system should favor:
- uniqueness,
- continuity relevance.
### As implemented (v1.1 WP-B.2): what the summariser is shown
A memory is written from one block of `MEMORY_INTERVAL` (6) story actions. The
summariser is given the cast brief, then the block, inside a budget of
`MEMORY_EXCERPT_TOKENS` (2,000).
- **A block that fits** is sent whole, exactly as v1.0.0 sent it.
- **A longer block** was cut to its last 2,000 tokens in v1.0.0, so a fact near
its start never reached the summariser (WP-B.1). It is now sent as its
opening and its end, in order, with a visible marker between them
(`EXCERPT_OMISSION_MARKER`, "[… the middle of this stretch of story is left
out here …]"). The marker and its blank lines are paid for first, and the rest
is halved, the odd token going to the end: 992 + 993 + 15 = 2,000 tokens. The
rejoined text is measured, and the opening gives up tokens until the whole is
within budget.
- **The marker is never stored.** `summarize_block` removes it from anything the
model repeats back.
- **Limit.** A fact in the middle of a very long block is still left out. The
input stays bounded; it is not a summary of everything.
Existing memories are not rewritten. `tools/rewrite_memories.py`, which is
opt-in, uses the same function.
### As implemented (v1.1): what a memory can be relied on to keep
The memory prompt is v1.0.0's, unchanged. v1.1 changed what the summariser is
shown (above), not what it is told.
**Known limitation, accepted for v1.1.** The application now delivers the whole
relevant block to the summariser, and keeps, ranks and injects the memory it
writes (§18, §20, §21). But the reference summariser, `qwen2.5:3b-instruct-16k`,
can still be given a block that states a distinctive fact and write a memory
that:
- omits the fact, or the specific objects in it;
- attributes it to the wrong character;
- prefers the generic narration that follows it.
So **independent recovery from memory is proven for the application's
mechanisms, and is not guaranteed with the reference model.** A bounded prompt
change aimed at this was tried and rejected
(`reports/v1.1/V1.1-WP-B2-REPORT.md` §T). A stronger dedicated summariser,
structured fact extraction, or separate factual and narrative memory are
future options, and none is implemented.
## 16. Memory Retrieval
Retrieval should be local.
@@ -513,6 +557,27 @@ abbey crypt
prior discoveries
```
### As implemented (v1.1 WP-B.2)
v1.0.0 embedded the newest four actions cut to 600 tokens, so the player's
one-line question arrived after three turns of narration and barely moved the
embedding (WP-B.1: a direct question's cosine fell from 0.708 alone to 0.241 in
that query). The query is now two short texts, embedded in one call:
| Component | What it is | Bound |
| --- | --- | --- |
| **input** | the player's own action this turn (`do`, `say` or `story`) | last 200 tokens |
| **context** | the scene from the authoritative state (its summary, the location's name, the names of who is present), then the end of the newest narration | 60 + 120 tokens |
- A continue turn, and the Insights dry run, have no input; the context alone is
searched.
- A retry searches with the input being retried; the discarded attempt is not in
the context.
- The full entity list, threads and older narration are deliberately left out,
so a long scene or a large cast cannot outweigh the question by length.
- What was searched for is recorded per turn in the context snapshot
(`memories.query`: input, context, input terms, weights).
## 19. Memory Retrieval Filtering
Before ranking memories, filter by:
@@ -553,6 +618,41 @@ future work rather than something M6 delivered.
What M6 does implement, because similarity alone proved insufficient, is
redundancy suppression before the final selection: see §22.
### As implemented (v1.1 WP-B.2)
One transparent lexical term is added to similarity. For each eligible memory:
```text
semantic_score = 0.6 * cos(input, memory) + 0.4 * cos(context, memory)
(either cosine alone when the other text is empty)
lexical_score = sum of w(t) over the input's terms the memory holds
/ sum of w(t) over all the input's terms in [0, 1]
w(t) = ln((N + 1) / (df(t) + 1)) N eligible memories, df holding t
final_score = semantic_score + 0.15 * lexical_score
```
- **Terms** are the knowledge path's tokenizer and stop list (`knowledge.fts`),
with possessives dropped and a plural `s` folded. No stemmer, no dependency.
- **Rarity** is computed per turn over the eligible candidates only. There is no
index and no stored field. A word every candidate holds (a protagonist's
name) weighs 0; a word the question shares with one memory weighs most.
- **Only the player's input** is matched lexically, never the context.
- **The weight** was chosen by sweep over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5:
0.15 is the smallest at which the lexical term alone lifts the planting-era
memory into `memory_top_k` against the v1.0.0 query, while no rare-word
negative control lets an unrelated memory pass a semantically relevant one.
A memory can gain at most 0.15 from wording, so it cannot pass one more than
0.15 ahead of it in meaning.
- **Pins** are unchanged: always used, counted toward `memory_top_k`.
- **Ties** on the final score are broken by memory id.
- **Provenance.** Each used memory records `semantic_score`, `lexical_score` and
`final_score`; `similarity` keeps its v1.0.0 meaning, the semantic score, so
the inspector's "closeness" is unchanged.
Importance, recency, entity overlap and story-thread overlap remain
unimplemented. Redundancy suppression (§22) is unchanged and runs over the final
order.
## 21. Memory Budget
Retrieved memories should have a bounded token budget.
@@ -564,6 +664,44 @@ Recommended behavior:
- include only the highest-value items that fit,
- preserve source IDs for inspection.
### As implemented (v1.1 WP-B.2): selection and the bank's capacity
**Selection** is unchanged in shape: every eligible, embedded memory on the
active lineage is scored (§20), pinned memories are taken first, the rest fill
`memory_top_k` (default 5) best first, skipping repeats (§22). The Memories
section is priced into the protected context like any other live section, so
its budget is unchanged (F03). There is still no relevance floor: a full
`memory_top_k` is used whenever the bank holds that many.
**Capacity** (`memory_bank_capacity`, default 80) is unchanged. Eviction still
runs after each post-turn pass, over the whole adventure rather than one
lineage, and marks rows `forgotten` rather than deleting them. What changed is
the order (`memorybank.eviction_order`), because least-recently-used alone
discarded the only memory of an early stretch first (WP-B.1):
- **Coverage signal.** Memories with a source range say which stretch they
describe. Each is judged by the hole its removal would leave between the end
of the memory before it and the start of the memory after it. The smallest
hole goes first, so the bank thins where it is densest. A memory whose start
another memory shares leaves no hole.
- **Boundaries.** The earliest and the latest memory by position are not
coverage candidates: they are the only records of the opening and of the most
recent stretch. This also keeps a memory written this turn from being evicted
by the pass that wrote it (the frozen bank).
- **Recency signal.** Among equal holes, the least recently used goes first
(`coalesce(last_used_at, created_at)`), then the less used, then the lower id.
- **Pinned rows** are never taken, and count as coverage.
- **Fallback.** Memories with no range (typed by the player, or migrated) and
boundaries are taken least recently used first, as in v1.0.0, once no coverage
candidate remains. The bank stays bounded either way; only an all-pinned bank
may exceed capacity.
Measured on banks where nothing is ever retrieved, the kept bank starts at the
opening and its largest uncovered stretch stays within about 1.3 times the
average spacing (story length / capacity). The v1.0.0 order kept only the newest
stretch. The rule reads no text and no vectors, and lineage eligibility is
unaffected: it decides only `forgotten`.
## 22. Duplicate Suppression
Do not include the same fact repeatedly through:
+4 -2
View File
@@ -3,8 +3,10 @@
**This file is the index. Start here.**
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: WP-A1
and WP-A2 are implemented and staged for owner review**
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`).
and WP-A2 are committed (`d63804f`), the WP-B.1 memory diagnostic is committed
(`beb17ad`), and WP-B.2, the memory-retention fix, is complete and staged for the owner's
signed commit, accepted with a documented reference-model memory limitation** (`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`,
`reports/v1.1/V1.1-WP-B1-REPORT.md`, `reports/v1.1/V1.1-WP-B2-REPORT.md`).
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
+30 -4
View File
@@ -3,10 +3,19 @@
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
signed v1.0.0 release commit `432f041`.
**WP-A1 and WP-A2 are implemented and staged for owner review** (2026-09-14). They
are reported together, and kept separate, in
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`. No other work package has started. No
v1.1 version or tag exists.
**WP-A1 and WP-A2** are committed and signed as `d63804f`
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`). **WP-B.1**, the memory-retention
diagnostic, is committed and signed as `beb17ad`
(`reports/v1.1/V1.1-WP-B1-REPORT.md`). **WP-B.2**, the retention fix, is complete and
staged for the owner's signed commit. It is **accepted with a documented
real-model limitation** (owner decision, 2026-09-15):
- B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship;
- deterministic independent-memory recovery passes;
- the one isolation-valid reference-model run failed at memory creation;
- the B2.4 prompt experiment did not fix that and was reverted.
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
WP-C, WP-D and WP-E have not started. No v1.1 version or tag exists.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
@@ -812,6 +821,23 @@ candidate tree:
A1's headroom table shows the documented reserve on every re-counted prompt,
and every turn is `fits`. A2's leak count is 0. B's independent-retention
verdict is recorded.
12. **WP-B's memory limitation is reported, not summarised away.** The v1.1
release report states each of these, and never shortens them to "WP-B
passed":
- deterministic independent-memory recovery: **PASS**;
- reference-model independent-memory recovery: **FAIL** on the
precondition-valid attempt;
- the failing stage: **memory creation**, the summariser's content
selection;
- the owner's decision to accept that limitation for v1.1.
The release long run's independent-retention verdict (item 6) is read
against it. A recovery there is reported as evidence, not as a reversal of
the limitation, unless it meets every isolation precondition.
13. **Carried residuals are listed with their status:**
- the mid-reply narrator instruction echo that A2's trailing cleanup does
not remove (WP-B.1 §K);
- the doubled full stop in the memory-search scene text (WP-B.2 §R 10).
7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0
protocol shapes in stored narration.
8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all
+23 -1
View File
@@ -2,7 +2,29 @@
- **Package:** Adventure Storyteller Planning Package v4.2
- **Revision date:** 2026-09-14
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): **WP-A1 and WP-A2 are implemented and staged for owner review.** No other work package has started, and no v1.1 version or tag exists.
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1 and WP-A2 are committed (`d63804f`) and WP-B.1 is committed (`beb17ad`). **WP-B.2 is complete and staged for the owner's signed commit, accepted with a documented real-model limitation: B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship; deterministic independent recovery passes; the reference model failed at memory creation on the precondition-valid attempt, and the B2.4 prompt experiment was rejected and reverted** (`reports/v1.1/V1.1-WP-B2-REPORT.md` §S, §T). WP-C, WP-D and WP-E have not started, and no v1.1 version or tag exists.
## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15)
WP-B.1's diagnostic (`beb17ad`) placed three memory deficiencies; WP-B.2 corrects
exactly those, one at a time, each verified before the next. No requirement or
acceptance test changed, and no schema, bundle format or setting default
changed. Evidence and the WP-B decision are in
`reports/v1.1/V1.1-WP-B2-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1 WP-B.2)**: a block longer than 2,000 tokens is shown to the summariser as its opening and its end with an omission marker, inside the same budget; a block that fits is unchanged; the marker is never stored. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` §18 | **As implemented**: the retrieval query is the player's input plus a bounded scene context, recorded per turn. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` §20 | **As implemented**: `final = semantic + 0.15 × lexical`, the rarity-weighted lexical term over the input, the sweep that chose the weight, pins and ties. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1)**: the memory prompt is unchanged from v1.0.0, and the accepted limitation: the reference summariser can omit or misattribute a fact from a block it was given whole. | as-implemented record |
| `DEVELOPMENT.md` | The GPU-host kernel/Ollama watch command corrected: `-k -u ollama` matched nothing; the OR form records both. | developer docs |
| `CONTEXT-AND-MEMORY.md` §21 | **As implemented**: selection unchanged in shape; eviction ordered by coverage first (boundaries kept, smallest hole first), recency second, v1.0.0 order as fallback. | as-implemented record |
| `V1.1-PLAN.md` | Status: WP-B accepted with a documented real-model limitation. §11 release criteria 12 and 13: the v1.1 release report must state the deterministic PASS and reference-model FAIL at memory creation, and list the carried residuals. | status, release gate |
| `planning/README.md` | Current state. | index |
| `reports/v1.1/V1.1-WP-B2-REPORT.md` | **New.** B2.1-B2.3 designs and evidence, full deterministic acceptance, real-model attempts, the rejected B2.4 prompt experiment, compatibility, offline, and the WP-B disposition. | work-package report |
**Requirement changes: zero.**
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
File diff suppressed because it is too large Load Diff