Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.
- B2.1 ranking: the retrieval query is the player's input plus a bounded
scene context (state scene + end of the newest narration), embedded in
one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
where lexical is a rarity-weighted share of the input's words, computed
per turn over the candidates with no index. Scores and the query are
recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
and newest memories are kept, the smallest coverage hole goes first,
least-recently-used breaks ties and remains the fallback. Bounded; pins
never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
to the summariser as head + tail with an omission marker, inside the
same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
experiment was measured on the reference model, showed no reliable
improvement for the target failure (0/5 under both prompts, with new
"Memory:"-prefix, second-person and length regressions), and was
reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
the failed block, a deterministic fidelity checker, and a real-model
shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
ranking_crowded, ranking_context_dependent and independent_full
fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
eviction and excerpt tests; summariser acceptance tests kept apart from
diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
release criteria 12-13), planning README, VERSION v4.3,
reports/v1.1/V1.1-WP-B2-REPORT.md.
Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
Diagnostic only; no memory behaviour changes.
- tools/memory_diagnostic.py: planted-fact isolation checks, the four-stage
diagnosis (created / retained / ranked / injected) with a verdict, a
production-ranking replica, deterministic summariser/embedder/narrator
stubs and seven scenarios (default, past capacity, pinned, low top_k,
long-block early/late, lineage control)
- tools/v11_b1_memory.py: CLI for the scenarios and for diagnosing a copy of
a finished real campaign
- tools/m11_long_run.py: opt-in --independent-fact mode with per-turn
isolation tracking and the recovered_through_memory_independent verdict;
M04 verdicts unchanged
- tests: diagnostic stages, eviction, creation window, ranking, lineage and
authority controls; two strict xfails record the diagnosed retention and
creation defects for WP-B.2 to flip
- planning/reports/v1.1/V1.1-WP-B1-REPORT.md
First failing stage: ranking (real model); retention past capacity and
creation for early facts in long blocks (deterministic, same on v1.0.0).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY