"""v1.1 WP-B.2: how faithfully a real model's memories keep the facts of their block. **Diagnostic only. Nothing here is imported by the application, and nothing here is a release gate.** Why it exists: the first isolation-valid real-model WP-B run failed at creation. The whole planting block reached the summariser, and the memory it wrote left the planted fact out (`V1.1-WP-B2-REPORT.md` §L.2). B2.4 tried the plan's bounded remedy, a memory prompt instructing the model to keep named facts and objects. It was measured with this module, did not correct the failure, and was **not shipped** (§T). The module stays so the limitation can be measured again, on this model or a different one. - **Fixtures.** Short story blocks, each built around a fact a later scene could turn on, with the ordinary texture a real block carries around it. They are genre-neutral (office, contemporary, a science-fiction-neutral station), plus the attempt-2 planting block itself. Each names the facts a memory must keep, whom each belongs to, and what it must not invent. - **The checker** (`evaluate`) is deterministic and reads only the memory text. It is a heuristic, and says so: - a fact counts as kept when one sentence names every part of it; - attribution is the nearest named character before the fact's verb; - it also reports word count, a leading "Memory:" and second-person "you". - **The comparison** (`compare`) sends each fixture to a real model through the application's own provider and `memorybank.memory_user_prompt`. It scores memories under the shipped prompt and under the rejected B2.4 experiment. A scripted summariser cannot show what a prompt makes a model do. So the deterministic tests prove only that the checker is right and the fixtures reach the summariser; model quality is measured here, with inference, and reported. # the real-model measurement (inference: ask first) .venv/bin/python -m tools.memory_fidelity --endpoint \\ --model qwen2.5:3b-instruct-16k --samples 5 --out "$HOME/v11-evidence/