v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
beb17ada10
commit
0c1ba836ba
@@ -451,6 +451,50 @@ The memory system should favor:
|
||||
- uniqueness,
|
||||
- continuity relevance.
|
||||
|
||||
### As implemented (v1.1 WP-B.2): what the summariser is shown
|
||||
|
||||
A memory is written from one block of `MEMORY_INTERVAL` (6) story actions. The
|
||||
summariser is given the cast brief, then the block, inside a budget of
|
||||
`MEMORY_EXCERPT_TOKENS` (2,000).
|
||||
|
||||
- **A block that fits** is sent whole, exactly as v1.0.0 sent it.
|
||||
- **A longer block** was cut to its last 2,000 tokens in v1.0.0, so a fact near
|
||||
its start never reached the summariser (WP-B.1). It is now sent as its
|
||||
opening and its end, in order, with a visible marker between them
|
||||
(`EXCERPT_OMISSION_MARKER`, "[… the middle of this stretch of story is left
|
||||
out here …]"). The marker and its blank lines are paid for first, and the rest
|
||||
is halved, the odd token going to the end: 992 + 993 + 15 = 2,000 tokens. The
|
||||
rejoined text is measured, and the opening gives up tokens until the whole is
|
||||
within budget.
|
||||
- **The marker is never stored.** `summarize_block` removes it from anything the
|
||||
model repeats back.
|
||||
- **Limit.** A fact in the middle of a very long block is still left out. The
|
||||
input stays bounded; it is not a summary of everything.
|
||||
|
||||
Existing memories are not rewritten. `tools/rewrite_memories.py`, which is
|
||||
opt-in, uses the same function.
|
||||
|
||||
### As implemented (v1.1): what a memory can be relied on to keep
|
||||
|
||||
The memory prompt is v1.0.0's, unchanged. v1.1 changed what the summariser is
|
||||
shown (above), not what it is told.
|
||||
|
||||
**Known limitation, accepted for v1.1.** The application now delivers the whole
|
||||
relevant block to the summariser, and keeps, ranks and injects the memory it
|
||||
writes (§18, §20, §21). But the reference summariser, `qwen2.5:3b-instruct-16k`,
|
||||
can still be given a block that states a distinctive fact and write a memory
|
||||
that:
|
||||
- omits the fact, or the specific objects in it;
|
||||
- attributes it to the wrong character;
|
||||
- prefers the generic narration that follows it.
|
||||
|
||||
So **independent recovery from memory is proven for the application's
|
||||
mechanisms, and is not guaranteed with the reference model.** A bounded prompt
|
||||
change aimed at this was tried and rejected
|
||||
(`reports/v1.1/V1.1-WP-B2-REPORT.md` §T). A stronger dedicated summariser,
|
||||
structured fact extraction, or separate factual and narrative memory are
|
||||
future options, and none is implemented.
|
||||
|
||||
## 16. Memory Retrieval
|
||||
|
||||
Retrieval should be local.
|
||||
@@ -513,6 +557,27 @@ abbey crypt
|
||||
prior discoveries
|
||||
```
|
||||
|
||||
### As implemented (v1.1 WP-B.2)
|
||||
|
||||
v1.0.0 embedded the newest four actions cut to 600 tokens, so the player's
|
||||
one-line question arrived after three turns of narration and barely moved the
|
||||
embedding (WP-B.1: a direct question's cosine fell from 0.708 alone to 0.241 in
|
||||
that query). The query is now two short texts, embedded in one call:
|
||||
|
||||
| Component | What it is | Bound |
|
||||
| --- | --- | --- |
|
||||
| **input** | the player's own action this turn (`do`, `say` or `story`) | last 200 tokens |
|
||||
| **context** | the scene from the authoritative state (its summary, the location's name, the names of who is present), then the end of the newest narration | 60 + 120 tokens |
|
||||
|
||||
- A continue turn, and the Insights dry run, have no input; the context alone is
|
||||
searched.
|
||||
- A retry searches with the input being retried; the discarded attempt is not in
|
||||
the context.
|
||||
- The full entity list, threads and older narration are deliberately left out,
|
||||
so a long scene or a large cast cannot outweigh the question by length.
|
||||
- What was searched for is recorded per turn in the context snapshot
|
||||
(`memories.query`: input, context, input terms, weights).
|
||||
|
||||
## 19. Memory Retrieval Filtering
|
||||
|
||||
Before ranking memories, filter by:
|
||||
@@ -553,6 +618,41 @@ future work rather than something M6 delivered.
|
||||
What M6 does implement, because similarity alone proved insufficient, is
|
||||
redundancy suppression before the final selection: see §22.
|
||||
|
||||
### As implemented (v1.1 WP-B.2)
|
||||
|
||||
One transparent lexical term is added to similarity. For each eligible memory:
|
||||
|
||||
```text
|
||||
semantic_score = 0.6 * cos(input, memory) + 0.4 * cos(context, memory)
|
||||
(either cosine alone when the other text is empty)
|
||||
lexical_score = sum of w(t) over the input's terms the memory holds
|
||||
/ sum of w(t) over all the input's terms in [0, 1]
|
||||
w(t) = ln((N + 1) / (df(t) + 1)) N eligible memories, df holding t
|
||||
final_score = semantic_score + 0.15 * lexical_score
|
||||
```
|
||||
|
||||
- **Terms** are the knowledge path's tokenizer and stop list (`knowledge.fts`),
|
||||
with possessives dropped and a plural `s` folded. No stemmer, no dependency.
|
||||
- **Rarity** is computed per turn over the eligible candidates only. There is no
|
||||
index and no stored field. A word every candidate holds (a protagonist's
|
||||
name) weighs 0; a word the question shares with one memory weighs most.
|
||||
- **Only the player's input** is matched lexically, never the context.
|
||||
- **The weight** was chosen by sweep over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5:
|
||||
0.15 is the smallest at which the lexical term alone lifts the planting-era
|
||||
memory into `memory_top_k` against the v1.0.0 query, while no rare-word
|
||||
negative control lets an unrelated memory pass a semantically relevant one.
|
||||
A memory can gain at most 0.15 from wording, so it cannot pass one more than
|
||||
0.15 ahead of it in meaning.
|
||||
- **Pins** are unchanged: always used, counted toward `memory_top_k`.
|
||||
- **Ties** on the final score are broken by memory id.
|
||||
- **Provenance.** Each used memory records `semantic_score`, `lexical_score` and
|
||||
`final_score`; `similarity` keeps its v1.0.0 meaning, the semantic score, so
|
||||
the inspector's "closeness" is unchanged.
|
||||
|
||||
Importance, recency, entity overlap and story-thread overlap remain
|
||||
unimplemented. Redundancy suppression (§22) is unchanged and runs over the final
|
||||
order.
|
||||
|
||||
## 21. Memory Budget
|
||||
|
||||
Retrieved memories should have a bounded token budget.
|
||||
@@ -564,6 +664,44 @@ Recommended behavior:
|
||||
- include only the highest-value items that fit,
|
||||
- preserve source IDs for inspection.
|
||||
|
||||
### As implemented (v1.1 WP-B.2): selection and the bank's capacity
|
||||
|
||||
**Selection** is unchanged in shape: every eligible, embedded memory on the
|
||||
active lineage is scored (§20), pinned memories are taken first, the rest fill
|
||||
`memory_top_k` (default 5) best first, skipping repeats (§22). The Memories
|
||||
section is priced into the protected context like any other live section, so
|
||||
its budget is unchanged (F03). There is still no relevance floor: a full
|
||||
`memory_top_k` is used whenever the bank holds that many.
|
||||
|
||||
**Capacity** (`memory_bank_capacity`, default 80) is unchanged. Eviction still
|
||||
runs after each post-turn pass, over the whole adventure rather than one
|
||||
lineage, and marks rows `forgotten` rather than deleting them. What changed is
|
||||
the order (`memorybank.eviction_order`), because least-recently-used alone
|
||||
discarded the only memory of an early stretch first (WP-B.1):
|
||||
|
||||
- **Coverage signal.** Memories with a source range say which stretch they
|
||||
describe. Each is judged by the hole its removal would leave between the end
|
||||
of the memory before it and the start of the memory after it. The smallest
|
||||
hole goes first, so the bank thins where it is densest. A memory whose start
|
||||
another memory shares leaves no hole.
|
||||
- **Boundaries.** The earliest and the latest memory by position are not
|
||||
coverage candidates: they are the only records of the opening and of the most
|
||||
recent stretch. This also keeps a memory written this turn from being evicted
|
||||
by the pass that wrote it (the frozen bank).
|
||||
- **Recency signal.** Among equal holes, the least recently used goes first
|
||||
(`coalesce(last_used_at, created_at)`), then the less used, then the lower id.
|
||||
- **Pinned rows** are never taken, and count as coverage.
|
||||
- **Fallback.** Memories with no range (typed by the player, or migrated) and
|
||||
boundaries are taken least recently used first, as in v1.0.0, once no coverage
|
||||
candidate remains. The bank stays bounded either way; only an all-pinned bank
|
||||
may exceed capacity.
|
||||
|
||||
Measured on banks where nothing is ever retrieved, the kept bank starts at the
|
||||
opening and its largest uncovered stretch stays within about 1.3 times the
|
||||
average spacing (story length / capacity). The v1.0.0 order kept only the newest
|
||||
stretch. The rule reads no text and no vectors, and lineage eligibility is
|
||||
unaffected: it decides only `forgotten`.
|
||||
|
||||
## 22. Duplicate Suppression
|
||||
|
||||
Do not include the same fact repeatedly through:
|
||||
|
||||
Reference in New Issue
Block a user