v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
beb17ada10
commit
0c1ba836ba
@@ -451,6 +451,50 @@ The memory system should favor:
|
||||
- uniqueness,
|
||||
- continuity relevance.
|
||||
|
||||
### As implemented (v1.1 WP-B.2): what the summariser is shown
|
||||
|
||||
A memory is written from one block of `MEMORY_INTERVAL` (6) story actions. The
|
||||
summariser is given the cast brief, then the block, inside a budget of
|
||||
`MEMORY_EXCERPT_TOKENS` (2,000).
|
||||
|
||||
- **A block that fits** is sent whole, exactly as v1.0.0 sent it.
|
||||
- **A longer block** was cut to its last 2,000 tokens in v1.0.0, so a fact near
|
||||
its start never reached the summariser (WP-B.1). It is now sent as its
|
||||
opening and its end, in order, with a visible marker between them
|
||||
(`EXCERPT_OMISSION_MARKER`, "[… the middle of this stretch of story is left
|
||||
out here …]"). The marker and its blank lines are paid for first, and the rest
|
||||
is halved, the odd token going to the end: 992 + 993 + 15 = 2,000 tokens. The
|
||||
rejoined text is measured, and the opening gives up tokens until the whole is
|
||||
within budget.
|
||||
- **The marker is never stored.** `summarize_block` removes it from anything the
|
||||
model repeats back.
|
||||
- **Limit.** A fact in the middle of a very long block is still left out. The
|
||||
input stays bounded; it is not a summary of everything.
|
||||
|
||||
Existing memories are not rewritten. `tools/rewrite_memories.py`, which is
|
||||
opt-in, uses the same function.
|
||||
|
||||
### As implemented (v1.1): what a memory can be relied on to keep
|
||||
|
||||
The memory prompt is v1.0.0's, unchanged. v1.1 changed what the summariser is
|
||||
shown (above), not what it is told.
|
||||
|
||||
**Known limitation, accepted for v1.1.** The application now delivers the whole
|
||||
relevant block to the summariser, and keeps, ranks and injects the memory it
|
||||
writes (§18, §20, §21). But the reference summariser, `qwen2.5:3b-instruct-16k`,
|
||||
can still be given a block that states a distinctive fact and write a memory
|
||||
that:
|
||||
- omits the fact, or the specific objects in it;
|
||||
- attributes it to the wrong character;
|
||||
- prefers the generic narration that follows it.
|
||||
|
||||
So **independent recovery from memory is proven for the application's
|
||||
mechanisms, and is not guaranteed with the reference model.** A bounded prompt
|
||||
change aimed at this was tried and rejected
|
||||
(`reports/v1.1/V1.1-WP-B2-REPORT.md` §T). A stronger dedicated summariser,
|
||||
structured fact extraction, or separate factual and narrative memory are
|
||||
future options, and none is implemented.
|
||||
|
||||
## 16. Memory Retrieval
|
||||
|
||||
Retrieval should be local.
|
||||
@@ -513,6 +557,27 @@ abbey crypt
|
||||
prior discoveries
|
||||
```
|
||||
|
||||
### As implemented (v1.1 WP-B.2)
|
||||
|
||||
v1.0.0 embedded the newest four actions cut to 600 tokens, so the player's
|
||||
one-line question arrived after three turns of narration and barely moved the
|
||||
embedding (WP-B.1: a direct question's cosine fell from 0.708 alone to 0.241 in
|
||||
that query). The query is now two short texts, embedded in one call:
|
||||
|
||||
| Component | What it is | Bound |
|
||||
| --- | --- | --- |
|
||||
| **input** | the player's own action this turn (`do`, `say` or `story`) | last 200 tokens |
|
||||
| **context** | the scene from the authoritative state (its summary, the location's name, the names of who is present), then the end of the newest narration | 60 + 120 tokens |
|
||||
|
||||
- A continue turn, and the Insights dry run, have no input; the context alone is
|
||||
searched.
|
||||
- A retry searches with the input being retried; the discarded attempt is not in
|
||||
the context.
|
||||
- The full entity list, threads and older narration are deliberately left out,
|
||||
so a long scene or a large cast cannot outweigh the question by length.
|
||||
- What was searched for is recorded per turn in the context snapshot
|
||||
(`memories.query`: input, context, input terms, weights).
|
||||
|
||||
## 19. Memory Retrieval Filtering
|
||||
|
||||
Before ranking memories, filter by:
|
||||
@@ -553,6 +618,41 @@ future work rather than something M6 delivered.
|
||||
What M6 does implement, because similarity alone proved insufficient, is
|
||||
redundancy suppression before the final selection: see §22.
|
||||
|
||||
### As implemented (v1.1 WP-B.2)
|
||||
|
||||
One transparent lexical term is added to similarity. For each eligible memory:
|
||||
|
||||
```text
|
||||
semantic_score = 0.6 * cos(input, memory) + 0.4 * cos(context, memory)
|
||||
(either cosine alone when the other text is empty)
|
||||
lexical_score = sum of w(t) over the input's terms the memory holds
|
||||
/ sum of w(t) over all the input's terms in [0, 1]
|
||||
w(t) = ln((N + 1) / (df(t) + 1)) N eligible memories, df holding t
|
||||
final_score = semantic_score + 0.15 * lexical_score
|
||||
```
|
||||
|
||||
- **Terms** are the knowledge path's tokenizer and stop list (`knowledge.fts`),
|
||||
with possessives dropped and a plural `s` folded. No stemmer, no dependency.
|
||||
- **Rarity** is computed per turn over the eligible candidates only. There is no
|
||||
index and no stored field. A word every candidate holds (a protagonist's
|
||||
name) weighs 0; a word the question shares with one memory weighs most.
|
||||
- **Only the player's input** is matched lexically, never the context.
|
||||
- **The weight** was chosen by sweep over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5:
|
||||
0.15 is the smallest at which the lexical term alone lifts the planting-era
|
||||
memory into `memory_top_k` against the v1.0.0 query, while no rare-word
|
||||
negative control lets an unrelated memory pass a semantically relevant one.
|
||||
A memory can gain at most 0.15 from wording, so it cannot pass one more than
|
||||
0.15 ahead of it in meaning.
|
||||
- **Pins** are unchanged: always used, counted toward `memory_top_k`.
|
||||
- **Ties** on the final score are broken by memory id.
|
||||
- **Provenance.** Each used memory records `semantic_score`, `lexical_score` and
|
||||
`final_score`; `similarity` keeps its v1.0.0 meaning, the semantic score, so
|
||||
the inspector's "closeness" is unchanged.
|
||||
|
||||
Importance, recency, entity overlap and story-thread overlap remain
|
||||
unimplemented. Redundancy suppression (§22) is unchanged and runs over the final
|
||||
order.
|
||||
|
||||
## 21. Memory Budget
|
||||
|
||||
Retrieved memories should have a bounded token budget.
|
||||
@@ -564,6 +664,44 @@ Recommended behavior:
|
||||
- include only the highest-value items that fit,
|
||||
- preserve source IDs for inspection.
|
||||
|
||||
### As implemented (v1.1 WP-B.2): selection and the bank's capacity
|
||||
|
||||
**Selection** is unchanged in shape: every eligible, embedded memory on the
|
||||
active lineage is scored (§20), pinned memories are taken first, the rest fill
|
||||
`memory_top_k` (default 5) best first, skipping repeats (§22). The Memories
|
||||
section is priced into the protected context like any other live section, so
|
||||
its budget is unchanged (F03). There is still no relevance floor: a full
|
||||
`memory_top_k` is used whenever the bank holds that many.
|
||||
|
||||
**Capacity** (`memory_bank_capacity`, default 80) is unchanged. Eviction still
|
||||
runs after each post-turn pass, over the whole adventure rather than one
|
||||
lineage, and marks rows `forgotten` rather than deleting them. What changed is
|
||||
the order (`memorybank.eviction_order`), because least-recently-used alone
|
||||
discarded the only memory of an early stretch first (WP-B.1):
|
||||
|
||||
- **Coverage signal.** Memories with a source range say which stretch they
|
||||
describe. Each is judged by the hole its removal would leave between the end
|
||||
of the memory before it and the start of the memory after it. The smallest
|
||||
hole goes first, so the bank thins where it is densest. A memory whose start
|
||||
another memory shares leaves no hole.
|
||||
- **Boundaries.** The earliest and the latest memory by position are not
|
||||
coverage candidates: they are the only records of the opening and of the most
|
||||
recent stretch. This also keeps a memory written this turn from being evicted
|
||||
by the pass that wrote it (the frozen bank).
|
||||
- **Recency signal.** Among equal holes, the least recently used goes first
|
||||
(`coalesce(last_used_at, created_at)`), then the less used, then the lower id.
|
||||
- **Pinned rows** are never taken, and count as coverage.
|
||||
- **Fallback.** Memories with no range (typed by the player, or migrated) and
|
||||
boundaries are taken least recently used first, as in v1.0.0, once no coverage
|
||||
candidate remains. The bank stays bounded either way; only an all-pinned bank
|
||||
may exceed capacity.
|
||||
|
||||
Measured on banks where nothing is ever retrieved, the kept bank starts at the
|
||||
opening and its largest uncovered stretch stays within about 1.3 times the
|
||||
average spacing (story length / capacity). The v1.0.0 order kept only the newest
|
||||
stretch. The rule reads no text and no vectors, and lineage eligibility is
|
||||
unaffected: it decides only `forgotten`.
|
||||
|
||||
## 22. Duplicate Suppression
|
||||
|
||||
Do not include the same fact repeatedly through:
|
||||
|
||||
+4
-2
@@ -3,8 +3,10 @@
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** **v1.0.0 released on 2026-09-14. v1.1 is in progress: WP-A1
|
||||
and WP-A2 are implemented and staged for owner review**
|
||||
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`).
|
||||
and WP-A2 are committed (`d63804f`), the WP-B.1 memory diagnostic is committed
|
||||
(`beb17ad`), and WP-B.2, the memory-retention fix, is complete and staged for the owner's
|
||||
signed commit, accepted with a documented reference-model memory limitation** (`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`,
|
||||
`reports/v1.1/V1.1-WP-B1-REPORT.md`, `reports/v1.1/V1.1-WP-B2-REPORT.md`).
|
||||
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
|
||||
through M11 complete and closed**. M11 was accepted at its closeout
|
||||
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
|
||||
|
||||
+30
-4
@@ -3,10 +3,19 @@
|
||||
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
|
||||
signed v1.0.0 release commit `432f041`.
|
||||
|
||||
**WP-A1 and WP-A2 are implemented and staged for owner review** (2026-09-14). They
|
||||
are reported together, and kept separate, in
|
||||
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`. No other work package has started. No
|
||||
v1.1 version or tag exists.
|
||||
**WP-A1 and WP-A2** are committed and signed as `d63804f`
|
||||
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`). **WP-B.1**, the memory-retention
|
||||
diagnostic, is committed and signed as `beb17ad`
|
||||
(`reports/v1.1/V1.1-WP-B1-REPORT.md`). **WP-B.2**, the retention fix, is complete and
|
||||
staged for the owner's signed commit. It is **accepted with a documented
|
||||
real-model limitation** (owner decision, 2026-09-15):
|
||||
- B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship;
|
||||
- deterministic independent-memory recovery passes;
|
||||
- the one isolation-valid reference-model run failed at memory creation;
|
||||
- the B2.4 prompt experiment did not fix that and was reverted.
|
||||
|
||||
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
|
||||
WP-C, WP-D and WP-E have not started. No v1.1 version or tag exists.
|
||||
|
||||
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
|
||||
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
|
||||
@@ -812,6 +821,23 @@ candidate tree:
|
||||
A1's headroom table shows the documented reserve on every re-counted prompt,
|
||||
and every turn is `fits`. A2's leak count is 0. B's independent-retention
|
||||
verdict is recorded.
|
||||
12. **WP-B's memory limitation is reported, not summarised away.** The v1.1
|
||||
release report states each of these, and never shortens them to "WP-B
|
||||
passed":
|
||||
- deterministic independent-memory recovery: **PASS**;
|
||||
- reference-model independent-memory recovery: **FAIL** on the
|
||||
precondition-valid attempt;
|
||||
- the failing stage: **memory creation**, the summariser's content
|
||||
selection;
|
||||
- the owner's decision to accept that limitation for v1.1.
|
||||
|
||||
The release long run's independent-retention verdict (item 6) is read
|
||||
against it. A recovery there is reported as evidence, not as a reversal of
|
||||
the limitation, unless it meets every isolation precondition.
|
||||
13. **Carried residuals are listed with their status:**
|
||||
- the mid-reply narrator instruction echo that A2's trailing cleanup does
|
||||
not remove (WP-B.1 §K);
|
||||
- the doubled full stop in the memory-search scene text (WP-B.2 §R 10).
|
||||
7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0
|
||||
protocol shapes in stored narration.
|
||||
8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all
|
||||
|
||||
+23
-1
@@ -2,7 +2,29 @@
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v4.2
|
||||
- **Revision date:** 2026-09-14
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): **WP-A1 and WP-A2 are implemented and staged for owner review.** No other work package has started, and no v1.1 version or tag exists.
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1 and WP-A2 are committed (`d63804f`) and WP-B.1 is committed (`beb17ad`). **WP-B.2 is complete and staged for the owner's signed commit, accepted with a documented real-model limitation: B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship; deterministic independent recovery passes; the reference model failed at memory creation on the precondition-valid attempt, and the B2.4 prompt experiment was rejected and reverted** (`reports/v1.1/V1.1-WP-B2-REPORT.md` §S, §T). WP-C, WP-D and WP-E have not started, and no v1.1 version or tag exists.
|
||||
|
||||
## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15)
|
||||
|
||||
WP-B.1's diagnostic (`beb17ad`) placed three memory deficiencies; WP-B.2 corrects
|
||||
exactly those, one at a time, each verified before the next. No requirement or
|
||||
acceptance test changed, and no schema, bundle format or setting default
|
||||
changed. Evidence and the WP-B decision are in
|
||||
`reports/v1.1/V1.1-WP-B2-REPORT.md`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1 WP-B.2)**: a block longer than 2,000 tokens is shown to the summariser as its opening and its end with an omission marker, inside the same budget; a block that fits is unchanged; the marker is never stored. | as-implemented record |
|
||||
| `CONTEXT-AND-MEMORY.md` §18 | **As implemented**: the retrieval query is the player's input plus a bounded scene context, recorded per turn. | as-implemented record |
|
||||
| `CONTEXT-AND-MEMORY.md` §20 | **As implemented**: `final = semantic + 0.15 × lexical`, the rarity-weighted lexical term over the input, the sweep that chose the weight, pins and ties. | as-implemented record |
|
||||
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1)**: the memory prompt is unchanged from v1.0.0, and the accepted limitation: the reference summariser can omit or misattribute a fact from a block it was given whole. | as-implemented record |
|
||||
| `DEVELOPMENT.md` | The GPU-host kernel/Ollama watch command corrected: `-k -u ollama` matched nothing; the OR form records both. | developer docs |
|
||||
| `CONTEXT-AND-MEMORY.md` §21 | **As implemented**: selection unchanged in shape; eviction ordered by coverage first (boundaries kept, smallest hole first), recency second, v1.0.0 order as fallback. | as-implemented record |
|
||||
| `V1.1-PLAN.md` | Status: WP-B accepted with a documented real-model limitation. §11 release criteria 12 and 13: the v1.1 release report must state the deterministic PASS and reference-model FAIL at memory creation, and list the carried residuals. | status, release gate |
|
||||
| `planning/README.md` | Current state. | index |
|
||||
| `reports/v1.1/V1.1-WP-B2-REPORT.md` | **New.** B2.1-B2.3 designs and evidence, full deterministic acceptance, real-model attempts, the rejected B2.4 prompt experiment, compatibility, offline, and the WP-B disposition. | work-package report |
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user