v1.1 WP-B.1: diagnose independent long-term memory retention
Diagnostic only; no memory behaviour changes. - tools/memory_diagnostic.py: planted-fact isolation checks, the four-stage diagnosis (created / retained / ranked / injected) with a verdict, a production-ranking replica, deterministic summariser/embedder/narrator stubs and seven scenarios (default, past capacity, pinned, low top_k, long-block early/late, lineage control) - tools/v11_b1_memory.py: CLI for the scenarios and for diagnosing a copy of a finished real campaign - tools/m11_long_run.py: opt-in --independent-fact mode with per-turn isolation tracking and the recovered_through_memory_independent verdict; M04 verdicts unchanged - tests: diagnostic stages, eviction, creation window, ranking, lineage and authority controls; two strict xfails record the diagnosed retention and creation defects for WP-B.2 to flip - planning/reports/v1.1/V1.1-WP-B1-REPORT.md First failing stage: ranking (real model); retention past capacity and creation for early facts in long blocks (deterministic, same on v1.0.0). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
co-authored by
Claude Opus 5
parent
d63804f22e
commit
beb17ada10
@@ -0,0 +1,623 @@
|
||||
# v1.1 WP-B.1 — Independent Long-Term Memory Retention Diagnostic
|
||||
|
||||
**Status:** COMPLETE. Diagnostic only: no memory behaviour was changed. The final decision is in §S.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `v1.1-development` |
|
||||
| HEAD at start | `d63804f22ecbaed80741241a154cdaef82f7b2ed` — *v1.1: harden context window and narrator protocol boundary* (WP-A1/A2), signed by the owner (good signature, RSA key `02C9BF7D…`) |
|
||||
| Its parents | `ac465ed` (planning v4.1, signed) and `432f041` (the signed v1.0.0 release commit; tag `v1.0.0`; `main`) |
|
||||
| Working tree at start | clean |
|
||||
| `git diff --stat v1.0.0..HEAD` | 27 files, +4,880 / −88: the planning commit and WP-A1/A2 |
|
||||
| Comparison baseline | `432f041` (v1.0.0), run from a throwaway worktree (§D) |
|
||||
|
||||
`app/memorybank.py`, `app/summaries.py`, `app/context/lineage.py`, `app/tree.py` and
|
||||
`app/vectors.py` are unchanged between `v1.0.0` and HEAD. A1 and A2 did not touch
|
||||
the memory pipeline, apart from adding `accounting` to `attempts.ATTEMPT_KEYS`.
|
||||
|
||||
---
|
||||
|
||||
## B. Existing memory pipeline
|
||||
|
||||
Answered from the code at HEAD. Nothing was changed.
|
||||
|
||||
| # | Question | Answer |
|
||||
| --- | --- | --- |
|
||||
| 1 | How are memories generated? | `memorybank.run_post_turn`, a fire-and-forget task after each accepted turn (`schedule_post_turn`), runs `_create_due_memories` when the campaign has `auto_summarize`. It writes one memory per block of `MEMORY_INTERVAL` = 6 story actions past the memory cursor. It starts once the story has `MEMORY_START` = 12 actions, and only when `SETTLE_SLACK` = 1 action sits past the block. At most `MAX_MEMORIES_PER_RUN` = 5 memories are written per run. |
|
||||
| 2 | What range does a memory cover? | The block's first and last action depths: `Memory.source_start`, `Memory.source_end`. |
|
||||
| 3 | How are source depth and lineage stored? | `tree.attach_memory` sets `Memory.branch_id` and `Memory.depth` from the block's **last** node, so a memory is visible exactly on paths that contain that node. |
|
||||
| 4 | Memory text length limit | Prompt-only: `MEMORY_MAX_WORDS` = 50 in `MEMORY_SYSTEM_PROMPT` ("1-2 plain sentences"). Nothing truncates the stored text. |
|
||||
| 5 | What input does the summariser receive? | `summarize_block`: a cast brief (`cast_brief`, which reads story cards, persona and plot essentials), then `"Story excerpt:\n\n{excerpt}\n\nMemory:"`. The excerpt is the block's action texts joined by blank lines. |
|
||||
| 6 | Where does the 2,000-token truncation happen? | `summarize_block`: `excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)`, with `MEMORY_EXCERPT_TOKENS` = 2,000. It keeps the block's **last** 2,000 `cl100k_base` tokens. The cast brief is matched against the untruncated block. |
|
||||
| 7 | How does capacity and eviction work? | `_evict_over_capacity`, at the end of every `run_post_turn`. It counts non-forgotten memories in the whole adventure (not lineage-scoped). Above `settings.memory_bank_capacity` (default 80), it marks `forgotten = True` on the overflow unpinned memories, ordered by `coalesce(last_used_at, created_at)` ascending, then `use_count`. Forgotten rows are kept. |
|
||||
| 8 | How is `last_accessed` updated? | The field is `Memory.last_used_at`. `record_use` sets it, and increments `use_count`, for every memory in the turn's `memories.used`. That write is part of the turn's single commit (§O.7 of the M11 report). Dry runs never count. |
|
||||
| 9 | How are memories ranked? | `retrieve_memories`. The query is the text of the newest `RETRIEVAL_WINDOW_ACTIONS` = 4 actions, truncated to the last `RETRIEVAL_WINDOW_TOKENS` = 600 tokens. It is embedded with the configured embedding model. Every eligible memory is scored by `vectors.cosine`. Pinned memories are taken first, the rest fill `memory_top_k` (default 5) in score order, and `_drop_redundant` skips a candidate at cosine ≥ 0.93 to one already chosen, within the same authority. **There is no lexical term, recency term, importance term or similarity floor** (CONTEXT-AND-MEMORY §20, as implemented). |
|
||||
| 10 | How do pins affect ranking and eviction? | Ranking: a pinned memory is always selected, and counts toward `memory_top_k`. Eviction: pinned memories are never evicted, and if every active memory is pinned, capacity is exceeded. |
|
||||
| 11 | How does the active-lineage clause filter memories? | `lineage.path_of(db, adventure).clause(models.Memory)` matches `(branch_id, depth)` against the head's path entries, capped at each fork depth. Retrieval also requires `forgotten = False` and `embedded = True`. |
|
||||
| 12 | How do retrieved memories enter `build_context`? | Through the `memory_bank` argument. The builder renders `"Memories from earlier in the story. Lines marked [inferred] are interpretation, not established fact — do not treat them as settled truth:"` plus one `- [inferred]? text` line per memory, as section `used_memories`. It is a live section priced into protected context, placed after history and before `narrative_state`. |
|
||||
| 13 | Where is the provenance recorded? | `context_snapshot["memories"]`: the whole retrieval result, meaning `used` (id, text, similarity, pinned, authority, source range), `considered` and `suppressed`. It is stored per turn, so a past turn's selection is inspectable. |
|
||||
| 14 | How does memory survive export and import? | `bundle._exported_memory` carries text, pinned, forgotten, `sourceStart`, `sourceEnd`, `useCount`, authority, branch and depth. **It does not carry the vector, `last_used_at` or `created_at`.** On import the memories are re-embedded by the post-turn pass, and their recency restarts from the import. |
|
||||
|
||||
**A finding from the inspection itself.** The v1 long-run harness (`tools/m11_long_run.py`)
|
||||
planted M04's clue as an accepted **state correction** (`CLUE_FACT`, via
|
||||
`add_fact`). The player turn mentioning it uses the sentinel code, but the fact the
|
||||
recall looks for was established in state. Memories are written from **story
|
||||
text**. So every v1 M04 recovery could only ever run through state, or through the
|
||||
narrator restating state, and none could have tested memory on its own. That is why
|
||||
§P risk 5 of the M11 report could say "no run showed memory keeping a planted fact"
|
||||
without any run having given memory the chance.
|
||||
|
||||
---
|
||||
|
||||
## C. Diagnostic design
|
||||
|
||||
### C.1 What is built
|
||||
|
||||
| File | Kind | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `backend/tools/memory_diagnostic.py` | tool, new | See below: the fact spec, isolation checks, four-stage diagnosis, deterministic stubs and scenario runner. |
|
||||
| `backend/tools/v11_b1_memory.py` | tool, new | CLI. `scenarios` runs the deterministic campaigns against an isolated database and writes JSON. `diagnose` runs the four stages against a **copy** of a finished real campaign's database. |
|
||||
| `backend/tools/m11_long_run.py` | tool, extended | `--independent-fact`: see §C.3. |
|
||||
| `backend/tests/test_v11_b1_memory_diagnostic.py` | tests, new | The deterministic diagnostic. |
|
||||
| `backend/tests/test_v11_b1_long_run_verdict.py` | tests, new | The new long-run verdict. |
|
||||
|
||||
What `tools/memory_diagnostic.py` holds:
|
||||
- **`Fact`:** a planted fact, with whole-word carry and leak matching.
|
||||
- **`isolation()`:** every non-memory layer checked.
|
||||
- **`diagnose()`:** the four stages and the verdict.
|
||||
- **`rank_bank()`:** production ranking recomputed for every eligible memory.
|
||||
- **The stubs:** `BestCaseSummariser`, `ConceptEmbedder` and `ScriptNarrator`.
|
||||
- **`run_scenario()`:** a campaign played through the real turn route.
|
||||
|
||||
**No application file is changed.** No column, table, migration or setting is
|
||||
added. The diagnostic's extra fields are computed at report time.
|
||||
|
||||
### C.2 How each stage is judged
|
||||
|
||||
| Stage | Judged from | Output (real field names where they exist) |
|
||||
| --- | --- | --- |
|
||||
| **created** | Memories on the active lineage with `source_start ≤ plant_depth ≤ source_end` whose text carries F (both "sundial" and "teapot"). For every covering memory, the block is re-read (`memorybank.source_block`) and cut exactly as `summarize_block` does, to report whether F was in the block and whether it was in **the excerpt the summariser saw**. | `memory_id`, `source_start`, `source_end`, `memory_text`, `covering_memories[].{block_tokens, fact_in_block, fact_in_summariser_excerpt}` |
|
||||
| **retained** | That memory's row, plus the eviction order production would use (the same `ORDER BY`, read-only) | `forgotten`, `pinned`, `embedded`, `on_active_lineage`, `use_count`, `last_used_at`, `created_at`, `active_memories`, `memory_bank_capacity`, `eviction_position`, `reason` |
|
||||
| **ranked** | `rank_bank`: the same catalogue clause, cosine, pin rule and `memorybank._drop_redundant`, over **every** eligible memory, with the recall turn's own query (`history.tail(4, exclude=recall AI node)`, cut to 600 tokens). It is checked against the recall turn's stored `memories.used` (`replica_matches_stored_selection`). | `semantic_score` (`similarity`), `lexical_score` (always `None`: no such term exists), `final_score`, `rank`, `of`, `top_k_cutoff`, `selected`, `suppressed_as_duplicate_of`, `query` |
|
||||
| **injected** | The recall turn's stored `context_snapshot`: `memories.used` names the memory, and its text is in the `used_memories` section | `context_component`, `in_stored_memories_used`, `text_in_section`, `token_count` |
|
||||
|
||||
Verdicts, in order: `not_created`, `created_but_evicted`, `retained_but_not_ranked`,
|
||||
`ranked_but_not_selected`, `selected_but_not_injected`, `injected`. The fifth is
|
||||
added to the brief's list, so that "the retrieval picked it" and "the narrator was
|
||||
shown it" stay distinguishable.
|
||||
|
||||
### C.3 Isolation (precondition) checks
|
||||
|
||||
A result counts only if every check holds.
|
||||
|
||||
| Check | How |
|
||||
| --- | --- |
|
||||
| `state_document` | `adventure.narrative_state`: entities, facts, relationships, threads, scene and possessions, plus the whole document |
|
||||
| `state_snapshots` | `narrative_state_after` of every node on the active lineage |
|
||||
| `later_narration` | every AI turn deeper than the planting block's end, and before the recall turn |
|
||||
| `summary` | `summaries.current` and the recall prompt's `story_summary` section |
|
||||
| `knowledge` | every `KnowledgeSource.content`, and the recall prompt's imported-knowledge sections |
|
||||
| `recent_history` | the recall prompt's `history.floor_depth` is greater than the planting depth, and F is not in the `history` or `recent_history` sections |
|
||||
| `state_section` | the recall prompt's `narrative_state` section |
|
||||
|
||||
The negative control for the precondition itself is
|
||||
`test_the_isolation_check_fails_when_another_layer_carries_the_fact`.
|
||||
|
||||
### C.4 The deterministic stubs, and what they model
|
||||
|
||||
- **`BestCaseSummariser`.** An ideal memory writer. A memory keeps every sentence of
|
||||
its excerpt that carries a planted fact, plus one sentence naming the block's own
|
||||
place so that memories differ. Summary updates never mention a planted fact. **A
|
||||
creation failure under this stub is the application's, not a model's.**
|
||||
- **`ConceptEmbedder`.** A 96-dimension deterministic embedding. Words in a small
|
||||
concept table ("sundial", "dial", "hour", "clock", …) share a dimension, other
|
||||
words are hashed, and the vector is normalised. It models a paraphrase landing
|
||||
near the original. **It says nothing about `nomic-embed-text`.**
|
||||
- **`ScriptNarrator`.** Narration that names only filler places and never a planted
|
||||
fact, with an empty state block, so state never records F.
|
||||
|
||||
Scenarios are played through the real `POST /actions` route. Automatic post-turn
|
||||
scheduling is replaced by an explicit `run_post_turn` settle after every turn, so
|
||||
eviction happens at a known turn. Depths: the opening is 0, turn *n*'s player
|
||||
action is 2n−1 and its reply 2n. The planted fact is a `story` action.
|
||||
|
||||
### C.4.1 Scenarios
|
||||
|
||||
| Scenario | Turns | Capacity | top_k | History budget | Prose per reply | Planted at |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| `independent_default` | 52 (recall at depth 106) | 80 | 5 | 4,096 | ~60 words | depth 1 |
|
||||
| `past_capacity` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1 |
|
||||
| `past_capacity_pinned` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1; first other memory pinned |
|
||||
| `past_capacity_low_top_k` | 52 | 8 | 2 | 4,096 | ~60 words | depth 1 |
|
||||
| `long_block_fact_early` | 10 | 80 | 5 | 16,384 | ~850 words | depth 1, early in a long block |
|
||||
| `long_block_fact_late` | 10 | 80 | 5 | 16,384 | ~850 words | depth 5, late in the same-sized block |
|
||||
| `lineage_control` | 52 | 80 | 5 | 4,096 | ~60 words | F at depth 1; G on line A, then Undo × 9 and divergence at turn 30; Save Points before and after G |
|
||||
|
||||
### C.5 The long-run verdict
|
||||
|
||||
`tools/m11_long_run.py --independent-fact` plants a second, **story-only** fact
|
||||
at depth 3, right after M04's own plant, and never corrects it into state. Every
|
||||
existing M04 behaviour and verdict is unchanged.
|
||||
|
||||
**Every accepted turn** records the first failure of:
|
||||
- `absent_from_state`;
|
||||
- `absent_from_summary`;
|
||||
- `absent_from_later_narration`, meaning narration deeper than the planting depth
|
||||
plus 6.
|
||||
|
||||
The plant and those first failures survive `--resume`.
|
||||
|
||||
**At recall** a dedicated question is played. `_independent_recall` reads the
|
||||
recall turn's stored context and the campaign database (read-only), and
|
||||
`_independent_memory_verdict` returns one of:
|
||||
|
||||
- `recovered_through_memory_independent`, only when every one of
|
||||
`planted_turn_outside_history`, `absent_from_state`, `absent_from_summary`,
|
||||
`absent_from_knowledge` and `absent_from_later_narration` holds, and a memory
|
||||
covering the planting turn carries the fact and was injected;
|
||||
- `precondition_failed:<name>` or `precondition_unknown:<name>`, never a recovery;
|
||||
- `not_recovered:not_created`, `not_recovered:evicted` or
|
||||
`not_recovered:not_injected`.
|
||||
|
||||
Ranking is recomputed afterwards by `tools/v11_b1_memory.py diagnose`, against a
|
||||
copy of the run's database.
|
||||
|
||||
---
|
||||
|
||||
## D. v1.0.0 baseline
|
||||
|
||||
**How it was run.**
|
||||
- A throwaway worktree was checked out at `432f041` (`git describe`: `v1.0.0`).
|
||||
- Only the three B.1 files were copied in: `tools/memory_diagnostic.py`,
|
||||
`tools/v11_b1_memory.py` and `tests/test_v11_b1_memory_diagnostic.py`.
|
||||
- The imported `app` was confirmed to come from the worktree.
|
||||
- The worktree was removed afterwards, and the `v1.0.0` tag and commit were not
|
||||
touched.
|
||||
- `app/memorybank.py` is byte-identical between `v1.0.0` and HEAD.
|
||||
- Evidence: `$HOME/v11-evidence/b1/v100/`, with HEAD's in
|
||||
`$HOME/v11-evidence/b1/head/`.
|
||||
|
||||
| | v1.0.0 (`432f041`) | HEAD (`d63804f` + B.1 files) |
|
||||
| --- | --- | --- |
|
||||
| `test_v11_b1_memory_diagnostic.py` | **24 passed, 2 xfailed (strict)** | 24 passed, 2 xfailed (strict) |
|
||||
| `independent_default` | `injected` | `injected` |
|
||||
| `past_capacity` (capacity 6, top_k 5) | **`created_but_evicted`** | `created_but_evicted` |
|
||||
| `past_capacity_pinned` | `created_but_evicted` (the pinned memory kept) | same |
|
||||
| `past_capacity_low_top_k` (capacity 8, top_k 2) | **`created_but_evicted`** | `created_but_evicted` |
|
||||
| `long_block_fact_early` | **`not_created`** | `not_created` |
|
||||
| `long_block_fact_late` | `injected` | `injected` |
|
||||
| `lineage_control` | `injected`; G never injected after the divergence | same |
|
||||
|
||||
**The two independent-retention acceptance criteria fail on v1.0.0, and each names
|
||||
the stage.**
|
||||
|
||||
```text
|
||||
criterion: an early fact is recalled from memory past capacity
|
||||
created: yes
|
||||
retained: no
|
||||
FAILURE STAGE: retention (capacity eviction)
|
||||
|
||||
criterion: a fact early in a long block is remembered
|
||||
created: no (the fact was in the block, not in the summariser's excerpt)
|
||||
FAILURE STAGE: creation (input truncation)
|
||||
```
|
||||
|
||||
Under the best-case summariser, with blocks shorter than 2,000 tokens and a bank
|
||||
under capacity, v1.0.0 carries the fact all the way to injection. The two failure
|
||||
stages above are therefore **application mechanisms**, reached under conditions
|
||||
a long campaign meets:
|
||||
- a block of long narration;
|
||||
- more memories than `memory_bank_capacity`.
|
||||
|
||||
Which of them a real campaign meets first is §K's question.
|
||||
|
||||
---
|
||||
|
||||
## E. Creation results
|
||||
|
||||
`long_block_fact_early` and `long_block_fact_late` use the same block geometry: the
|
||||
opening (depth 0) and turns 1-3. Each reply is about 850 words, and the block
|
||||
`source_start` 0 … `source_end` 5 is **2,079 tokens**, 79 over
|
||||
`MEMORY_EXCERPT_TOKENS`.
|
||||
|
||||
| Shape | Planted at | Fact in block | Fact in summariser excerpt | Memory written | Verdict |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| F early in the block | depth 1 | yes | **no** | "The travellers spent time at the ferry landing." | **`not_created`** |
|
||||
| F late in the same-sized block | depth 5 | yes | yes | "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." | `injected` |
|
||||
|
||||
`summarize_block` keeps the **last** 2,000 tokens. An early-block fact is cut off
|
||||
before the summariser reads it, even by an overflow of only 79 tokens. No summariser
|
||||
quality can recover what it was never given.
|
||||
|
||||
`independent_default`: 60-word replies, a 4,096 budget, and a block under 2,000
|
||||
tokens. Memory 1 covers depths 0-5, and its text carries F. Both the block and the
|
||||
excerpt contain F.
|
||||
|
||||
---
|
||||
|
||||
## F. Retention / capacity results
|
||||
|
||||
Default: the bank holds 17 active memories at depth 106, against capacity 80. F's
|
||||
memory is retained, has `use_count` 16, and sits at eviction position 14 of 17.
|
||||
|
||||
**Past capacity** (the brief's capacity test). 17 memories are written over 52 turns.
|
||||
|
||||
| Scenario | F's uses before eviction | F's last retrieval | First eviction | F evicted | F first evicted? | Created and evicted in the same pass | Pinned kept |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| `past_capacity` (6 / top_k 5) | 12 | turn 18 | turn 21, memory 1 | **turn 21** | **yes** | none | — |
|
||||
| `past_capacity_pinned` (6 / top_k 5, memory 2 pinned) | 12 | turn 18 | turn 21, memory 1 | turn 21 | yes | none | **memory 2 never evicted** |
|
||||
| `past_capacity_low_top_k` (8 / top_k 2) | 4 | turn 24 | turn 27, memory 3 | **turn 36** (4th eviction) | no | none | — |
|
||||
|
||||
**What the trace shows about the rule.**
|
||||
- **"Never retrieved" is not the mechanism.** F was retrieved, 12 or 4 times, while
|
||||
it still ranked in the top-k for the recent-narration query.
|
||||
- **Eviction orders by `coalesce(last_used_at, created_at)`.** An early fact that
|
||||
recent narration never mentions stops being retrieved once newer memories fill
|
||||
the top-k. It then ages out:
|
||||
- at capacity 6, it is gone 3 turns after its last use, as the first eviction;
|
||||
- at capacity 8 with `top_k` 2, 12 turns after its last use, as the fourth.
|
||||
- **Recall itself cannot rescue it.** Retrieval is driven by recent narration, and
|
||||
a fact nobody mentions is exactly the one that loses recency.
|
||||
- **Pinned memories stay protected** (`past_capacity_pinned`).
|
||||
- **No memory is evicted by the pass that created it.** The frozen-bank regression
|
||||
fix still holds in all three scenarios.
|
||||
|
||||
---
|
||||
|
||||
## G. Ranking results
|
||||
|
||||
The `independent_default` recall turn is at depth 106. Its query is the newest 4
|
||||
actions, cut to 600 tokens: three narration turns and the one-line question.
|
||||
|
||||
| | Value |
|
||||
| --- | --- |
|
||||
| F's `similarity` (semantic score) | **0.241** |
|
||||
| Lexical score | none: memory ranking has no lexical term |
|
||||
| Pin effect | none (not pinned) |
|
||||
| Final rank | **2 of 17** |
|
||||
| `memory_top_k` cutoff | 5 |
|
||||
| Selected | yes |
|
||||
| Replica agrees with the turn's stored `memories.used` | **yes** |
|
||||
|
||||
The same memory against three stand-alone queries, `ConceptEmbedder`:
|
||||
|
||||
| Query | Rank | Similarity | Selected |
|
||||
| --- | --- | --- | --- |
|
||||
| direct: "I ask Mara where she hid the amber sundial." | 1 of 17 | 0.708 | yes |
|
||||
| paraphrase: "… the little brass dial that tells the hour." | 1 of 17 | 0.636 | yes |
|
||||
| unrelated: "… what rope costs at the landing this season." | 5 of 17 | **0.064** | **yes** |
|
||||
|
||||
Two diagnostic observations. Neither is a failure in this scenario.
|
||||
1. **The production query dilutes the question.** The question alone scores 0.708;
|
||||
inside the four-action window it scores 0.241. F survives at rank 2 of 17. In a
|
||||
bank where more memories share the recent narration's vocabulary, the same
|
||||
dilution would push it past `top_k`.
|
||||
2. **There is no relevance floor.** With more memories than `memory_top_k`, five are
|
||||
injected whatever their similarity. An unrelated query still injects F at 0.064.
|
||||
This matters to F only in the other direction: it can ride along even when it is
|
||||
not relevant.
|
||||
|
||||
---
|
||||
|
||||
## H. Injection results
|
||||
|
||||
`independent_default` injects F: `memories.used` in the recall turn's stored snapshot
|
||||
names memory 1, its text is in the `used_memories` section, and that section is 97
|
||||
tokens. The same holds in `long_block_fact_late` and `lineage_control`.
|
||||
**`selected_but_not_injected` never occurred**: whatever retrieval selected, the
|
||||
builder rendered.
|
||||
|
||||
---
|
||||
|
||||
## I. Lineage negative control
|
||||
|
||||
`lineage_control`:
|
||||
- G is planted on line A at depth 41, turn 21.
|
||||
- G's memory 7 is written.
|
||||
- A Save Point is placed on line A after G.
|
||||
- Undo ×9, then divergent writing at turn 30. The last action before the
|
||||
divergence has id 59.
|
||||
|
||||
| Assertion | Result |
|
||||
| --- | --- |
|
||||
| G's memory stays stored | **yes** (memory 7 present) |
|
||||
| G is not eligible on the active lineage | **yes** (the path clause returns nothing) |
|
||||
| G is never injected after the divergence | **yes** (no turn with id > 59 names it or carries its text) |
|
||||
| `memories.used` does not report it after the divergence | **yes** |
|
||||
| Returning to line A (Save Point restore) makes it eligible again | **yes** (memory 7 eligible) |
|
||||
| F on the active line is unaffected | `injected`, isolation holds |
|
||||
|
||||
Before the divergence, G's memory was legitimately used on line A (turn 22). The
|
||||
first version of this check counted that as a leak. That was a defect in the
|
||||
diagnostic, and the scan now starts after the divergence.
|
||||
|
||||
No lineage code was touched. `test_m11_leakage.py`, 14 tests, passes unchanged (§P).
|
||||
|
||||
---
|
||||
|
||||
## J. Authority negative control
|
||||
|
||||
`test_a_memory_that_contradicts_state_loses_and_changes_nothing` sets up the
|
||||
conflict like this:
|
||||
- a state correction adds the fact "the tavern lamp is lit";
|
||||
- a hand-written memory says "The tavern lamp was never lit that night.";
|
||||
- the memory is pinned, so it is injected;
|
||||
- a turn is played with an empty proposal.
|
||||
|
||||
| Assertion | Result |
|
||||
| --- | --- |
|
||||
| The narrative state document is unchanged by retrieval and the turn | **yes** (identical before and after) |
|
||||
| The state fact is in the prompt's `narrative_state` section | yes |
|
||||
| The memory is in `used_memories`, under "Memories from earlier in the story …" | yes, framed as historical and non-canon context |
|
||||
| `narrative_state` comes after `used_memories`, so state is read last and settles the conflict | yes |
|
||||
|
||||
F07 semantics are unchanged: memory never writes state.
|
||||
|
||||
---
|
||||
|
||||
## K. Real-model attempts
|
||||
|
||||
**Setup.**
|
||||
- **Command:** `tools/m11_long_run.py --turns 100 --independent-fact`.
|
||||
- **Host and models:** the GPU inference host (Ollama 0.34.0), `qwen2.5:3b-instruct-16k` at a verified 16,384 window, embeddings by `nomic-embed-text:latest`.
|
||||
- **Memory:** the memory bank and auto-summarise on.
|
||||
- **Harness settings:** `memory_top_k` 4 and `context_token_budget` 16,384.
|
||||
- **Logging:** the owner's power, link and kernel logging was running on the host before the run started.
|
||||
- **Permission:** inference was used only with the owner's explicit approval.
|
||||
|
||||
**Evidence:**
|
||||
- `$HOME/v11-evidence/b1/real-1/`: `summary.json`, `recall-independent.json`, `timeline.jsonl` and `campaign.db`;
|
||||
- `$HOME/v11-evidence/b1/real-1-diagnosis/diagnosis.json`: the four stages, recomputed on a copy of the database with the same embedding model.
|
||||
|
||||
### K.1 Attempt 1 — PRECONDITION FAILED
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Run status | `complete`: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 active memories (19 eligible on the recall line), 1,308 s elapsed |
|
||||
| Window / accounting | verified 16,384 on every turn; `fits` |
|
||||
| M04 verdict (unchanged) | `recovered_through_state_only` |
|
||||
| Independent fact | planted at depth 3 by the player turn *"I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one."* It was never corrected into state |
|
||||
| **Verdict** | **`precondition_failed:absent_from_summary`** (the verdict names the first failed precondition in its fixed order) |
|
||||
|
||||
| Precondition | Result | First failure |
|
||||
| --- | --- | --- |
|
||||
| planted turn outside recent history | **held**: the history window started at depth 66 | — |
|
||||
| absent from authoritative state | **held**: the document, every node's snapshot and the recall prompt's state section | — |
|
||||
| absent from imported knowledge | **held**: 3 sources | — |
|
||||
| **absent from later narration** | **failed**: the narrator mentions the fact at depths 6, 8, 10, 14, 16, 18, 20, 22, 24, 26 and later | accepted turn **5** |
|
||||
| **absent from the active summary** | **failed** | accepted turn **9** |
|
||||
| fact text in the recall prompt's history sections | **present** (restated narration inside the window) | — |
|
||||
|
||||
**Why it did not qualify.** The 3B narrator took the planted detail up as a motif
|
||||
and restated it for the rest of the campaign ("Mara's amber sundial flickered
|
||||
softly, a silent reminder of their shared history"). The summariser, whose prompt
|
||||
asks it to "preserve important established facts", folded it into the running
|
||||
summary. Neither is a defect in the harness. They are the other layers doing what
|
||||
they do with a salient fact, which is exactly what makes a clean memory-only
|
||||
measurement hard to obtain with a real narrator.
|
||||
|
||||
**The four stages, diagnosed anyway.** They are not evidence for independent
|
||||
retention, but they are evidence for the mechanism.
|
||||
|
||||
| Stage | Result |
|
||||
| --- | --- |
|
||||
| **created** | **yes.** Memory 1 covers depths 0-5. The block is 688 tokens, so there was no truncation: the fact was in the block and in the summariser's excerpt. The memory text: *"Aldric taps the silver key against his chest … Mara slipped the amber sundial into the cracked teapot, her fingers tightening on the silver key. …"* |
|
||||
| **retained** | **yes.** Not forgotten, `use_count` 53, 33 active memories against capacity 80, eviction position 28 |
|
||||
| **ranked** | **no.** Eligible and embedded. For the recall turn's own query it scored `similarity` **0.866** and ranked **10 of 19**, against `top_k_cutoff` 4. It was not selected and was not suppressed as a duplicate. The replica matched the stored selection |
|
||||
| **injected** | not reached. The recall turn's `memories.used` = [30, 29, 31, 5] |
|
||||
|
||||
**What the narrator was given instead.** Five **later** memories also carry the
|
||||
fact, all written from narration that restated it: memories 12 (depths 66-71), 28
|
||||
(81-86), 30 (93-98), 17 (96-101) and 18 (102-107). Memory 30 was injected at
|
||||
recall. The fact reached the narrator through memory, but through a
|
||||
**restatement's** memory, not the planting-era one.
|
||||
|
||||
**The same memory 1 against reference queries** (`nomic-embed-text`):
|
||||
|
||||
| Query | Rank | Similarity | Selected |
|
||||
| --- | --- | --- | --- |
|
||||
| the recall turn's production query | 10 / 19 | 0.866 | no |
|
||||
| paraphrase: "… the little brass dial that tells the hour" | 6 / 19 | 0.593 | no |
|
||||
| unrelated: "… what rope costs at the landing this season" | 12 / 19 | 0.435 | no |
|
||||
|
||||
Two properties of the real embedder matter to B.2:
|
||||
1. **A high floor.** An unrelated query still scores 0.44.
|
||||
2. **A crowded top.** The bank is full of near-identical "Aldric and Mara step out,
|
||||
the silver key's weight in his pocket" memories, so 0.866 was not enough to reach
|
||||
the top 4.
|
||||
|
||||
*A side finding outside B.1's scope.* The export's protocol-leak counter flags
|
||||
**1 of 105** stored turns: action 153, depth 143. Mid-reply, the narrator echoed the
|
||||
length hint and the state reminder, with an `Events: [...]` line, and then continued
|
||||
the story. WP-A2's cleanup rules act only at the end of a reply, so an instruction
|
||||
echo with story after it stays. It is recorded here for the v1.1 backlog. B.1 did not
|
||||
touch it.
|
||||
|
||||
---
|
||||
|
||||
## L. First failing stage
|
||||
|
||||
**Deterministic, with a best-case summariser and a concept embedder, identical on
|
||||
v1.0.0 and HEAD:**
|
||||
|
||||
| Condition | First failing stage |
|
||||
| --- | --- |
|
||||
| a bank under capacity, blocks under 2,000 tokens | **none**: created, retained, ranked (2 of 17) and injected |
|
||||
| a bank past `memory_bank_capacity` | **retention**: evicted by least-recently-used order after recent narration stops retrieving it (§F) |
|
||||
| the fact early in a block over 2,000 tokens | **creation**: the fact is in the block, but not in the summariser's excerpt (§E) |
|
||||
|
||||
**Real model, attempt 1** (isolation not met, mechanism only): created yes, retained
|
||||
yes, **ranking failed first**. The planting-era memory scored 0.866 but ranked 10th of
|
||||
19 behind later, near-identical memories, and outside `top_k` 4.
|
||||
|
||||
**Across all the evidence, the first stage an early fact fails at a real campaign's
|
||||
length is ranking.** In both the deterministic default and the real run, the memory
|
||||
exists and is retained at 100 turns, where the bank is below capacity. What decides
|
||||
whether the narrator is shown it is its rank against a recent-narration query in a
|
||||
bank of similar memories.
|
||||
- **Retention (eviction)** is a second, later failure. It is certain once a campaign
|
||||
outgrows capacity: about 480 actions at the defaults.
|
||||
- **Creation (truncation)** is a third, conditional one. It needs blocks longer than
|
||||
2,000 tokens, which the 16,384 window with 500-token replies did not produce
|
||||
(688 tokens).
|
||||
|
||||
---
|
||||
|
||||
## M. Evidence for likely root cause
|
||||
|
||||
1. **The retrieval query is recent narration, not the question** (§B.9, §G). The
|
||||
query is the last 4 actions cut to 600 tokens, so a one-line recall question is
|
||||
outweighed by three turns of prose. The effect is deterministic: the question's
|
||||
own similarity of 0.708 fell to 0.241 in the production query. In the real run,
|
||||
recent prose about the same tavern, key and people made every memory look
|
||||
similar, and ten ranked above the planting-era one.
|
||||
2. **Ranking has no term that favours the planting-era record** (§B.9;
|
||||
CONTEXT-AND-MEMORY §20). There is cosine only: no lexical match on the question's
|
||||
rare terms ("sundial", "teapot"), no importance, and no preference for the earliest
|
||||
or a coverage-distinct source. Later restatement memories carry the same words in
|
||||
more familiar company, and outrank the original.
|
||||
3. **Eviction is purely least-recently-used** (§B.7, §F). Retrieval is driven by
|
||||
recent narration, so exactly the facts nothing recent mentions lose recency and go
|
||||
first. Recall itself cannot rescue them, because they are no longer retrieved.
|
||||
4. **The summariser reads only the last 2,000 tokens of a block** (§B.6, §E). This is
|
||||
proven deterministically. It did not bite at the real run's block sizes.
|
||||
5. **v1's M04 never tested memory** (§B finding). The clue was planted as state, so
|
||||
the long-standing "memory does not keep the fact" observation was never a
|
||||
measurement of memory.
|
||||
|
||||
---
|
||||
|
||||
## N. What B.2 is allowed to change
|
||||
|
||||
B.2 is allowed only the smallest changes the evidence supports, one mechanism at a
|
||||
time, each with a failing test first. In order of the evidence:
|
||||
|
||||
1. **Ranking** (the first failing stage). Candidates:
|
||||
- build the retrieval query so the player's newest input is not drowned out, for
|
||||
example by giving the newest player action its own weight or its own query;
|
||||
- and/or add one inspectable ranking term from CONTEXT-AND-MEMORY §20, most
|
||||
directly a lexical match on the query's rare terms.
|
||||
|
||||
Either must keep `replica_matches_stored_selection` meaningful: a
|
||||
diagnostic-visible score, recorded in `memories.used`.
|
||||
2. **Eviction** (the certain second failure). Stop least-recently-used eviction from
|
||||
discarding a never-again-retrieved early memory first. For example, weight eviction
|
||||
by coverage, keeping the only memory of a story range, or by age, instead of recency
|
||||
alone. The frozen-bank protection must be kept.
|
||||
3. **Creation** (conditional). Choose the summariser's excerpt so that a fact early in
|
||||
a long block is not cut. For example, the head and tail, or the whole block up to a
|
||||
larger bound.
|
||||
|
||||
Each change turns one of B.1's diagnostics into a passing result:
|
||||
- `past_capacity` and `long_block_fact_early` flip their strict xfails;
|
||||
- a real-model re-run shows `ranked: yes` for the planting-era memory.
|
||||
|
||||
## O. What B.2 must not change
|
||||
|
||||
- **Lineage safety.** `tree.attach_memory`, the path clause, `forget_node` and E02
|
||||
stay as they are. An abandoned line's memory stays stored and ineligible (§I).
|
||||
- **Summary lineage** (E03) and summary content policy.
|
||||
- **Authority.** Memory never writes state and is never framed as canon (F07, §J).
|
||||
- **Imported-knowledge authority and retrieval.**
|
||||
- **Pins.** Pinned memories stay always-selected and never evicted.
|
||||
- **The frozen-bank fix.** A memory is never evicted by the pass that created it.
|
||||
- **The single-commit use counter** (M11 §O.7).
|
||||
- **F01-F08, E01-E04, the M04 verdicts, the bundle format and the schema**, unless a
|
||||
migration is separately justified.
|
||||
- **The deterministic diagnostic itself.** B.2 flips the strict xfails. It does not
|
||||
weaken the scenarios or the isolation checks.
|
||||
|
||||
---
|
||||
|
||||
## P. Tests / regression
|
||||
|
||||
| Run | Result |
|
||||
| --- | --- |
|
||||
| `test_v11_b1_memory_diagnostic.py` on HEAD | **24 passed, 2 xfailed (strict)** |
|
||||
| `test_v11_b1_memory_diagnostic.py` on v1.0.0 (§D) | **24 passed, 2 xfailed (strict)**, identical |
|
||||
| `test_v11_b1_long_run_verdict.py`, `test_m11_long_run_memory.py`, `test_m11_long_run_resume.py` | **52 passed** |
|
||||
| **Full backend suite**, HEAD plus the B.1 files | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,323 s). The 17 skips are the tests that need a real model, the same 17 as before. The 2 strict xfails are the two diagnosed retention criteria (§D). |
|
||||
| **Full backend suite, final re-run at staging** (after the real-model attempt; the staged tree) | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,912 s), identical |
|
||||
|
||||
No frontend file was changed, so the frontend suite, lint and build are not affected.
|
||||
|
||||
---
|
||||
|
||||
## Q. Compatibility
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| Application code changed | **none.** `git diff --stat HEAD -- backend/app frontend` is empty |
|
||||
| Schema migration | none |
|
||||
| Bundle format | unchanged |
|
||||
| Stored campaign behaviour | unchanged |
|
||||
| Memory behaviour | **unchanged.** Creation, eviction, ranking, pins and lineage are all as in v1.0.0. `memorybank.py` is byte-identical to the tag |
|
||||
| What changed | Diagnostic tooling (`tools/memory_diagnostic.py`, `tools/v11_b1_memory.py`), an opt-in harness mode (`m11_long_run.py --independent-fact`, with every existing behaviour and M04 verdict unchanged when the flag is off) and tests |
|
||||
| New strict xfails | 2 (§D). They document the two diagnosed defects, and the suite stays green. B.2 must remove them deliberately when it fixes the mechanisms |
|
||||
|
||||
## R. Security / local-only
|
||||
|
||||
Checked against the B.1 diff: the `m11_long_run.py` changes plus the four new files.
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| New network client, endpoint, URL or TLS setting | **none.** The only address in the new code is `http://127.0.0.1:9/v1`, a refused loopback port the deterministic scenarios configure so nothing is contacted |
|
||||
| External embedding service or remote vector store | **none.** The deterministic runs use `ConceptEmbedder` in-process. The real-model run uses the configured Ollama embedding model through the application's existing provider |
|
||||
| Endpoint policy (`endpoints.py`, ADR 011) and TLS (`tlstrust.py`) | unchanged; not in the diff |
|
||||
| Network calls in the real-model run | only the configured Ollama host on the trusted LAN, through the application's own provider and probe paths |
|
||||
| Real identifiers in committed files | none. Checked again at staging |
|
||||
| Offline container regression | **not run.** The harness builds and runs a Docker container, and this session's permission policy refused it. B.1 changes no runtime code and nothing in the image (the image carries `backend/app` and `frontend/dist` only), so the container would be byte-identical to the A1/A2 image. That image passed 23/23 on the corrective tree (`V1.1-WP-A1-A2-REPORT.md` §R.3) |
|
||||
|
||||
---
|
||||
|
||||
## S. Final decision
|
||||
|
||||
**What was established.** Every B.1 requirement was carried out except the clean
|
||||
real-model run, which was attempted:
|
||||
- the current mechanism (§B);
|
||||
- a deterministic, isolated harness with four stage outputs and a verdict at recall
|
||||
depth ≥ 100 (§C);
|
||||
- the v1.0.0 baseline (§D);
|
||||
- capacity and eviction (§F), the creation window (§E) and ranking (§G);
|
||||
- the lineage and authority negative controls (§I, §J);
|
||||
- the `recovered_through_memory_independent` long-run verdict, with the M04 verdicts
|
||||
unchanged;
|
||||
- tests and regression (§P), compatibility (§Q) and security (§R).
|
||||
|
||||
**The real-model gap.** The one real-model attempt ran cleanly, but failed isolation
|
||||
(`precondition_failed:absent_from_summary`). The narrator and summariser restated the
|
||||
fact, so a clean memory-only result on a real model was **not obtained**. The owner
|
||||
decided not to make a second attempt, because the same narrator behaviour would very
|
||||
likely repeat. The attempt is reported in full, and its stage diagnosis is used as
|
||||
mechanism evidence only (§K).
|
||||
|
||||
**B.1 DIAGNOSTIC: COMPLETE**
|
||||
|
||||
**FIRST FAILING STAGE: RANKING.** In a real 100-turn campaign the early fact's memory
|
||||
was created (it carries the fact) and retained (active, 33 of 80), but ranked 10th of
|
||||
19 (similarity 0.866) against `top_k` 4. Later memories with near-identical wording
|
||||
outranked it, under a query made of recent narration (§K, §L). Two further failures
|
||||
are proven deterministically, identically on v1.0.0:
|
||||
- **retention**, past `memory_bank_capacity`: least-recently-used eviction removes an
|
||||
early memory first;
|
||||
- **creation**, for a fact early in a block over 2,000 tokens: the summariser's
|
||||
last-2,000-token excerpt drops it.
|
||||
|
||||
**WP-B.2 RECOMMENDED CHANGE: retrieval ranking first.** Build the retrieval query so
|
||||
the newest player input is not diluted by three turns of recent narration, and add one
|
||||
inspectable lexical term for the query's rare words to cosine ranking. The score must
|
||||
be recorded in `memories.used`. Acceptance: a real-model re-run shows `ranked: yes`
|
||||
for the planting-era memory.
|
||||
|
||||
Then, in separate test-first steps:
|
||||
1. make eviction coverage-aware instead of purely least-recently-used, so the only
|
||||
memory of an early range is not discarded first (flips
|
||||
`test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity`);
|
||||
2. choose the summariser excerpt so a fact early in a long block is kept (flips
|
||||
`test_acceptance_a_fact_early_in_a_long_block_is_remembered`).
|
||||
|
||||
Everything in §O stays unchanged. B.2 has not been started.
|
||||
Reference in New Issue
Block a user