v1.1 WP-B.2: independent long-term memory retention

Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
This commit is contained in:
JesseMarkowitz
2026-09-15 11:21:53 -04:00
co-authored by Claude Opus 5
parent beb17ada10
commit 0c1ba836ba
15 changed files with 3782 additions and 186 deletions
+30 -4
View File
@@ -3,10 +3,19 @@
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
signed v1.0.0 release commit `432f041`.
**WP-A1 and WP-A2 are implemented and staged for owner review** (2026-09-14). They
are reported together, and kept separate, in
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`. No other work package has started. No
v1.1 version or tag exists.
**WP-A1 and WP-A2** are committed and signed as `d63804f`
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`). **WP-B.1**, the memory-retention
diagnostic, is committed and signed as `beb17ad`
(`reports/v1.1/V1.1-WP-B1-REPORT.md`). **WP-B.2**, the retention fix, is complete and
staged for the owner's signed commit. It is **accepted with a documented
real-model limitation** (owner decision, 2026-09-15):
- B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship;
- deterministic independent-memory recovery passes;
- the one isolation-valid reference-model run failed at memory creation;
- the B2.4 prompt experiment did not fix that and was reverted.
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
WP-C, WP-D and WP-E have not started. No v1.1 version or tag exists.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
@@ -812,6 +821,23 @@ candidate tree:
A1's headroom table shows the documented reserve on every re-counted prompt,
and every turn is `fits`. A2's leak count is 0. B's independent-retention
verdict is recorded.
12. **WP-B's memory limitation is reported, not summarised away.** The v1.1
release report states each of these, and never shortens them to "WP-B
passed":
- deterministic independent-memory recovery: **PASS**;
- reference-model independent-memory recovery: **FAIL** on the
precondition-valid attempt;
- the failing stage: **memory creation**, the summariser's content
selection;
- the owner's decision to accept that limitation for v1.1.
The release long run's independent-retention verdict (item 6) is read
against it. A recovery there is reported as evidence, not as a reversal of
the limitation, unless it meets every isolation precondition.
13. **Carried residuals are listed with their status:**
- the mid-reply narrator instruction echo that A2's trailing cleanup does
not remove (WP-B.1 §K);
- the doubled full stop in the memory-search scene text (WP-B.2 §R 10).
7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0
protocol shapes in stored narration.
8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all