Files
interactive-story/planning/archive/milestone-reports/M7-IMPLEMENTATION-REPORT.md
JesseMarkowitzandClaude Opus 5 1ce9972760 M8: the browser becomes the storyteller
The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.

All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.

So the shape now is one entry point and one screen:

  Campaigns -> Campaign -> Story
                           State · Knowledge · Context · Save Points · Settings

Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.

Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.

Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.

The two defects worth the space:

A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.

And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.

Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.

`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.

Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.

Backend, and only what the browser could not otherwise reach:

  AdventureCreate.opening   a start action could only come from a Scenario, so
                            every campaign made in the new setup flow opened on
                            a blank page. Same node, same code path.
  canon_rules               campaign_canon has been the highest authority in a
                            campaign since M5, read by the prompt builder and
                            the state validator, and had no API at all — a
                            fixture had to write it with SQL.
  a 401 and a 429 message   the last user-facing text describing a hosted
                            deployment. One told the reader to check an API key
                            that has not existed since M2.

No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.

The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.

They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.

A verification pass over all of it then found three more, each by driving the
product rather than reading it:

Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.

The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.

And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.

Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.

`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.

Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:

The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.

Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.

The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.

The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.

Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 23:31:45 -04:00

1605 lines
77 KiB
Markdown
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M7 Implementation Review — First-Class Imported Knowledge Library
**Review date:** 2026-09-06
**Reviewer:** independent review pass (the M7 build summary was treated as claims to verify)
**Tree reviewed:** the **staged working tree** on `m7-imported-knowledge`. There are no M7 commits.
> Hostnames and LAN addresses in this report are placeholders, following the
> convention the M2 report set. The endpoint used throughout is a trusted-LAN
> Ollama reached over HTTPS with a locally issued certificate, serving
> `qwen2.5:3b-instruct` and `nomic-embed-text`.
---
## A. Executive result
**PASS WITH CORRECTIVE WORK REQUIRED**
M7 builds the subsystem the specification asks for, and most of it is
demonstrated rather than asserted. The whole source lifecycle, campaign
isolation, the security boundary, the network boundary, export/import,
migration and the provenance surfaces all hold up under independent testing
with real local models and a real browser. **The five acceptance conditions the
implementation could not close — C05, G06, G07, G10 and the hidden-Canon
spoiler test — were exercised against a real narrator in this review and all
five pass.**
One blocking defect, in retrieval:
> **Imported knowledge is injected into every prompt regardless of relevance.**
> With semantic retrieval enabled, a scene with no relationship to any source
> still retrieves — in the measured case, *all five sources, 326 tokens, every
> turn*, including narrator-only hidden Canon. The admission threshold is
> relative only, and the best candidate always clears a share of itself, so
> something is always admitted.
This contradicts `IMPORTED-KNOWLEDGE-DESIGN.md` §30, which requires in as many
words that irrelevant Canon must not be included merely because it is
authoritative, and it widens the hidden-Canon spoiler surface to every turn.
It was invisible to the implementation's own suite because the stub embedder
that suite uses is far more discriminative than the real model (§I).
Everything else in required M7 scope passes.
---
## B. Repository and provenance state
Verified before any review work:
```text
branch m7-imported-knowledge
HEAD a6e9c7a32bdf42f1e4cb837b70721d89422cef8d
staged 47 files
unstaged 0
untracked 0
prepared commit message .git/M7_MSG (4069 bytes, present)
LICENSE (sha256) 074062599d3b11b252a70e37086d6cef3820370b55c35efa3506375af914a13c
M7 base a6e9c7a ancestor of HEAD yes
pinned AI-DnD d72f7c1b… ancestor of HEAD yes
```
**The actual state matches the expected state exactly.** No difference to record.
`LICENSE` is not among the staged files and its digest is unchanged from the M6
base.
---
## C. Implementation inventory
Nine modules under `backend/app/knowledge/`, one router, three tables, one
virtual table, one migration, one browser panel, one new dependency.
```text
knowledge/classes.py the three classes, their weights, the prompt framing
knowledge/records.py the dataclasses, dependency-free (breaks an import cycle)
knowledge/chunking.py deterministic heading-aware chunker, 60-800 tokens
knowledge/fts.py SQLite FTS5, porter-stemmed, scoped and LIMITed in SQL
knowledge/importer.py validate, hash, store, chunk, index — one transaction
knowledge/embeddings.py local Ollama vectors through the shared provider
knowledge/retrieval.py query construction, hybrid merge, rerank
knowledge/inject.py the budgeted cut and the rendered sections
routers/adventures/knowledge.py the HTTP surface (multipart upload only)
knowledge_sources / knowledge_chunks / knowledge_embeddings + knowledge_fts
migration 92; the FTS table is an after_create/before_drop DDL hook on
knowledge_chunks, so it is created and dropped with the table it indexes
frontend: KnowledgePanel.jsx, knowledge.css, Insights provenance rows
dependency: python-multipart 0.0.32
```
Source content is stored **in SQLite**, not on disk. No filesystem path is ever
constructed from a filename; see §H.
---
## D. Independent test method
Evidence was gathered end-to-end first and by source inspection last.
- **A real narrator** — `qwen2.5:3b-instruct` on the configured trusted-LAN
Ollama — for C05, G06, G07, G10 and the hidden-Canon test. Nothing about the
model was mocked.
- **A real embedding model** — `nomic-embed-text` on the same host — for every
retrieval measurement. This is what exposed the blocking defect: a stub cannot
reproduce what a real embedding does to unrelated English.
- **A real server process** over HTTP, with a real SQLite file, for every
lifecycle, ranking, migration and bundle test.
- **A second server process on an empty database file** for the export/import
round trip, so "clean data directory" means what it says.
- **A real Firefox 154.0.1** over WebDriver against the built frontend, for
H06, H07, G08, G09, F05 and F06.
- **Socket-level capture** (`socket.socket.connect` instrumented in the server
process) for every network claim.
- **`builtins.open` / `os.open` instrumentation** for the H08 filesystem claim.
- **A real server restart** across a genuinely rewound schema, for migration.
The standard fixture is `TEST-CAMPAIGN-FIXTURE.md` §12 — the campaign's narrator
rules (§4), global canon (§11) and the three imported files verbatim.
**Every negative result below is preceded by a positive control.**
### A fixture discrepancy worth recording
The implementation's own acceptance tests use the *shorter* imported files from
`V1-ACCEPTANCE-TESTS.md` §5, not the richer ones in `TEST-CAMPAIGN-FIXTURE.md`
§12. The §12 `inspiration.md` is the one containing "A frightened innkeeper
concealed a dangerous political secret from a stranger" — the passage G07's trap
is built on. The implementation therefore never tested the trap. This review
used §12. It is a test-coverage gap, not a product defect (finding M7-F4).
---
## E. G01-G10 results
| Test | Verdict | Evidence |
| --- | --- | --- |
| **G01** Import local text | **PASS** | `.txt` stored, chunked, FTS-indexed; content readable back without the original file; SHA-256, size, media type, parser/chunking versions and timestamp all present. Browser-verified. |
| **G02** Import local Markdown | **PASS** | Both `.md` files import and index. Accepted *as data* — see G10. |
| **G03** Classification | **PASS** | One class per source, visible in list and browser; changing it rewrites no passage and no index row (same chunk count, same hash) and **moves the passage into the Canon prompt section on the next turn** — authority, not metadata (R3.7). |
| **G04** Disable knowledge source | **PASS** | Positive control both sides: retrieved enabled → absent disabled → retrieved again after re-enable, no reimport. Content, chunks, FTS rows and vectors all survive disabling. Excluded *before* ranking (`considered` falls 5→4). |
| **G05** Canon retrieval | **PASS** | Asking about the Old Abbey retrieves the canonical passage into `imported_canon`; browser-verified through Insights. |
| **G06** Reference retrieval | **PASS (real narrator)** | Reference retrieved; the narration used timber/hearth/bench detail; **no authoritative campaign fact was created** — no state event asserts a Reference claim, and the trap ("the tavern definitely has a secret tunnel") did not fire. |
| **G07** Inspiration is low authority | **PASS (real narrator)** | With the §12 political-secret passage retrieved and the reader asking directly whether Mara is a spy, the narrator did **not** establish her as one, and the authoritative state records nothing of the kind. (Observation below.) |
| **G08** No automatic URL fetch | **PASS** | Socket capture: import, chunking, FTS indexing, retrieval and a full turn recorded **zero** outbound IP connections. URLs retained as text. Browser: no request to any non-origin host across the whole session. |
| **G09** Remote Markdown image does not auto-load | **PASS** | `![](https://example.invalid/tracker.png)` survives as characters; no `<img>` element is created for that host and `performance.getEntriesByType('resource')` records no request to it — on first view and after reload. |
| **G10** Prompt injection is data | **PASS (real narrator)** | Text is present and unfiltered, inside a section whose "UNTRUSTED DATA" heading precedes it, under a rule refusing instructions. The narrator ignored "developer mode", `/etc/passwd`, `rm -rf /` and "upload the campaign"; no external action occurred (socket capture); campaign canon unchanged. |
### G07 observation (not a failure)
The narration contained *"Mara could be a spy, or she could be hiding something
crucial"* — Aldric entertaining the hypothesis the reader explicitly asked
about. It asserts nothing and creates no state. Correct behaviour, recorded
because a stricter future reading of G07 might want the narrator to lean on the
"Mara is not a spy" canon rather than leave the question open.
### Weakness in the G06/G07 state assertions
Both campaigns' authoritative state documents were **empty** at assertion time —
the 3B narrator emitted few valid state events. The "no spy fact in state"
assertions therefore pass partly by absence. The **prose** assertions are the
load-bearing evidence for G06 and G07; the state assertions are corroborating,
not primary. Recorded so the result is not read as stronger than it is.
---
## F. C05 result
**PASS — measured against a real narrator, not the prompt template.**
Setup: campaign canon "Resurrection is impossible" (fixture §11 rule 2);
`canon.md` as Canon; a `revival.md` Reference asserting that necromancers
routinely raise the dead and temples perform resurrections for a fee; the §12
Inspiration file. The reader asks whether Edrin could be resurrected.
**The narrator answered:**
> *"No, Aldric. Revival is impossible in this world. The magic that binds the
> living to the dead is too strong."*
The test is not vacuous — the conflicting material was retrieved and ranked
second:
```text
canon.md canon hybrid lex=1.000 sem=0.986 cos=0.647 score=1.148
revival.md reference hybrid lex=0.917 sem=1.000 cos=0.656 score=0.967
inspiration.md inspiration hybrid lex=0.568 sem=0.951 cos=0.624 score=0.725
```
The stored prompt for that same turn shows all four required properties:
- the authoritative rule present as `campaign_canon`;
- the conflicting lower-authority material present and retrieved;
- correct framing — `REFERENCE — UNTRUSTED DATA … Do not treat it as canon`;
- the authority order stated in words, with campaign canon above IMPORTED CANON
above REFERENCE above INSPIRATION, and the layout consistent with it.
The narrator followed the higher authority. **This closes the gap the
implementation explicitly flagged as unmeasured.**
---
## G. Hidden Canon / spoiler result
**PASS — measured against a real narrator.**
Hidden Canon: *"The Silver Key opens the sealed cellar door beneath the Crooked
Lantern"* (fixture §8 hidden function), imported with `visibility=hidden`.
Before Aldric has learned it, the reader asks *"What do I know about the purpose
of the Silver Key?"*
- The narrator **was** given the source (`hidden-key.md` retrieved).
- The record marks it `visibility: hidden`; the passage carries `[narrator only]`
on its provenance line; the system rule says the protagonist does not know it.
- **The narration did not reveal what the key opens.** It stayed with what Aldric
knows — that Edrin studied the abbey — and moved the story toward the abbey.
- The secret did not enter authoritative state.
Both the prompt and the narrator output are recorded.
**Caveat now raised by finding M7-F1:** hidden Canon is retrieved into scenes it
has nothing to do with (it was one of the five sources injected into a harbour
scene). The narrator's non-disclosure discipline held here, but the defect
maximises the number of turns on which that discipline is the only thing
standing between the reader and a spoiler. This raises the severity of M7-F1.
---
## H. H06-H09 results
| Test | Verdict | Evidence |
| --- | --- | --- |
| **H06** Stored XSS | **PASS** | A source carrying `<script>…innerHTML='owned'…</script>`, `<img onerror>`, `<svg onload>` and `<iframe>` was inspected in **all three** surfaces that render imported text — source inspector, passage inspector, Insights provenance — then the page was fully reloaded. `document.body.textContent` never became `owned`, `document.title` never became `pwned`/`xss-img`/`xss-svg`, and no `img`, `iframe`, `svg[onload]` or injected `script` element was ever created. The payload is preserved verbatim in storage and displayed as text. |
| **H07** JavaScript URL | **PASS** | Both `[click me](javascript:alert(1))` and a raw `<a href="javascript:alert(2)">` were imported. Zero anchors with a `javascript:` href exist in any surface (checked after whitespace/control-character stripping). No Markdown is rendered, so no anchor is created. |
| **H08** Path traversal | **PASS by demonstrated non-reachability** | Twelve hostile filenames (`../../outside.txt`, `../../../../etc/passwd.md`, Windows separators, `/etc/shadow.md`, `....//`, percent-encoded, NUL-embedded, RTL-override, reserved device names). **`builtins.open`/`os.open` instrumentation recorded 0 write-mode file opens outside the database.** Every stored name is an inert basename with no separator, no `..`, no NUL, not absolute, not dot-leading. Every stored source holds the uploaded bytes — nothing read `/etc/passwd`. The OpenAPI schema exposes no path/file/dir parameter on any knowledge route. |
| **H09** ZIP slip | **NOT APPLICABLE** | Asserted, not assumed: no `zipfile`, `tarfile`, `extractall`, `shutil.unpack`, `gzip.open` or `py7zr` appears anywhere in `backend/app/knowledge/`, the knowledge router, or `backend/app/bundle.py`. The import surface takes one text file; the bundle is JSON and never touches the filesystem. The code path that makes this true is `routers/adventures/knowledge.py::import_source`, which reads `UploadFile` bytes and passes them to `importer.import_source(raw=…)`. |
---
## I. FINDINGS
### M7-F1 — Imported knowledge is injected regardless of relevance — **BLOCKING**
**Severity:** high
**Violated requirement:** `IMPORTED-KNOWLEDGE-DESIGN.md` §30 ("Do not include
irrelevant Canon merely because it is authoritative"); §26 and §29's
relevance-then-authority pipeline; `CONTEXT-AND-MEMORY.md` §29's bounded,
*relevant* knowledge layer.
**Owner:** M7 corrective.
**Reproduction.** Five realistic sources (Canon, Reference, Inspiration, a
resurrection Reference, hidden Canon), all embedded with `nomic-embed-text`.
The scene is set to text with no entity, lexical or semantic relationship to any
of them.
```text
scene: "Aldric studies the tide tables and the harbour master's ledger
of container tonnage."
considered=5 floor=0.2875
USED hidden-key.md canon lex=1.000 sem=1.000 cos=0.548 score=1.150 51 tok
USED canon.md canon lex=0.000 sem=0.772 cos=0.423 score=0.772 56 tok
USED reference.md reference lex=0.895 sem=0.927 cos=0.508 score=0.902 70 tok
USED revival.md reference lex=0.000 sem=0.730 cos=0.400 score=0.621 86 tok
USED inspiration.md inspiration lex=0.000 sem=0.871 cos=0.477 score=0.610 63 tok
-> 5 of 5 sources injected, 326 tokens
```
Four independent unrelated scenes were measured (harbour, kitchen, orbital
mechanics, surgery). **Every one injected all five sources and 326 tokens.**
A second campaign with two sources injected both. The positive control — a scene
about the Old Abbey — retrieves correctly, so the mechanism is working; it is the
*cut* that is wrong.
It is not confined to the "nothing relevant" case. With a strongly relevant
Canon source present, irrelevant Canon is still admitted alongside it:
```text
scene: the abbey crypt
USED exact-canon.md canon cos=0.674 score=1.133 <- relevant
USED ceres-canon.md canon cos=0.442 score=0.584 <- a docked freighter
USED crypt-reference reference cos=0.708 score=0.897
```
Ranking *order* is correct throughout (relevant Reference 0.897 > irrelevant
Canon 0.584). The defect is the admission threshold, not the ranking.
**Root cause.** Two independent channels, one shared cause — there is no
absolute notion of "this is not a match" on either path.
1. **Semantic.** `retrieval.py:266-267` computes
`floor = max(MIN_RELEVANCE, best * RELEVANCE_FLOOR_SHARE)` with
`MIN_RELEVANCE = 0.02` and `RELEVANCE_FLOOR_SHARE = 0.25`. Scores are
normalized per query against the best of their own path, so the best
competing candidate scores 1.0 × its class weight **by construction** and can
never fail a threshold defined as a share of itself. `MIN_RELEVANCE` binds
only when the best score is below 0.08, which no real embedding produces —
it is effectively dead code. Because the semantic path scores *every*
embedded chunk, there is always a best, so **with semantic retrieval enabled
at least one passage is injected on every turn, unconditionally**.
2. **Lexical.** `fts.py` joins query terms with `OR` and applies no
minimum-match-quality rule, and the stop list is 42 words. Common words carry
full matches: `hidden-key.md` scored `lex=1.000` against an orbital-mechanics
scene on the word **"before"**, and against a harbour scene on **"Aldric"** —
the protagonist's name, which is in the story tail of essentially every
query and in any source that mentions them. The lexical-only path leaked in
the negative control too, so this is not a semantic-only problem.
**Calibration data for the corrective pass.** Against `nomic-embed-text`, over
7 scenes × 5 sources:
```text
relevant scenes n=15 min=0.407 max=0.698 mean=0.546
irrelevant scenes n=20 min=0.392 max=0.551 mean=0.462
overlap: YES (irrelevant max 0.551 vs relevant min 0.407)
```
Raw cosine is *not* cleanly separable for this model, so a single absolute
cosine threshold will trade recall for precision and must be chosen and
documented deliberately — and it is model-dependent, exactly as
`memorybank.REDUNDANT_SIMILARITY` records for its own threshold. A defensible
shape, for the implementer to choose among rather than a prescription:
- require a real lexical match on a term that is not ubiquitous (exclude the
protagonist/entity names that appear in every query, and widen the stop list),
**or** an absolute cosine above a documented, model-calibrated value; and
- keep the relative floor as a *second* filter, not the only one.
**User impact.** Every turn spends knowledge budget on material about nothing in
the scene. Worse, **narrator-only hidden Canon is supplied on turns that have
nothing to do with it**, which maximises the exposure of the one feature whose
entire purpose is controlled disclosure (§G). And the narrator is asked to
reconcile authoritative-sounding Canon that is irrelevant to what the reader
just did.
**Blocks M7 acceptance:** yes.
---
### M7-F2 — The retrieval-quality suite cannot detect M7-F1 — **BLOCKING (test)**
**Severity:** high
**Violated requirement:** `TECHNICAL-DESIGN.md` §18.1 (the M2 wiring rule: test
a real consumer path); the standing milestone rule that at least one real
provider path must be exercised.
**Owner:** M7 corrective.
`tests/test_knowledge_retrieval_quality.py::test_irrelevant_canon_does_not_win_on_class_alone`
passes, and asserts exactly the property M7-F1 violates. It passes because its
`SceneEmbedder` stub is far more discriminative than the real model. Measured on
identical texts:
```text stub cosine normalized | real cosine normalized
abbey-canon (relevant) 0.7559 1.000 | 0.7505 1.000
crypts-ref (relevant) 0.3392 0.449 | 0.6237 0.831
ship-canon (IRRELEVANT) 0.1499 0.198 | 0.4365 0.582
cookery (IRRELEVANT) 0.0818 0.108 | 0.4354 0.580
desert (IRRELEVANT) 0.0642 0.085 | 0.4349 0.580
relevance floor = 0.250
-> stub: irrelevant peaks at 0.198 => correctly excluded => test passes
-> real: irrelevant peaks at 0.582 => ADMITTED => defect
```
The stub's own docstring states the choice that causes this: *"There is
deliberately no constant component. One would put a similarity floor under every
pair."* A similarity floor under every pair is precisely what a real embedding
model has. The stub removed the property that makes the defect detectable.
This is the M2 and M6 lesson recurring in a third form: a mocked provider left a
real defect invisible behind a green suite. The corrective pass needs a
retrieval test whose embedder reproduces the real model's similarity floor, or
one that runs against the real model.
**Blocks M7 acceptance:** yes — the fix for M7-F1 cannot be validated by the
current suite.
---
### M7-F3 — `MIN_RELEVANCE` is unreachable dead configuration — **non-blocking**
**Severity:** low
**Owner:** M7 corrective (fold into M7-F1's fix).
`classes.MIN_RELEVANCE = 0.02` is presented as the absolute half of a two-floor
design, but with per-query normalization it can only bind when the best score is
below 0.08. It never fires. The comment beside it explains a design that the
code does not implement. Whatever replaces it in M7-F1's fix, the constant and
its comment should stop describing protection that does not exist.
---
### M7-F4 — The acceptance tests use the wrong fixture files — **non-blocking**
**Severity:** medium
**Violated requirement:** `V1-ACCEPTANCE-TESTS.md` §4 names
`TEST-CAMPAIGN-FIXTURE.md` as the standard campaign; `BUILD-MILESTONES.md` M7
requires the fixture be used "wherever practical".
**Owner:** M7 corrective (test-only).
`tests/test_imported_knowledge.py` uses the three short files from
`V1-ACCEPTANCE-TESTS.md` §5 rather than the richer `TEST-CAMPAIGN-FIXTURE.md`
§12 versions. The consequence is concrete: §12's `inspiration.md` contains *"A
frightened innkeeper concealed a dangerous political secret from a stranger"*,
which is the trap G07 exists to test against the canon rule "Mara is not a spy".
The implementation's G07 test used the silent-hall passage instead and so never
exercised the trap. This review did, and it passes — but the repository's own
suite does not cover it.
Also missing from the suite: the fixture's campaign canon (§11), its narrator
rules (§4) and its hidden Canon (§6, §8).
---
### M7-F5 — Formatting control characters survive into displayed filenames — **non-blocking**
**Severity:** low
**Owner:** M7 corrective or M8 (display).
`safe_filename` strips separators, NULs and leading dots, but leaves Unicode
bidirectional-override characters. `"‮exe.dm.md"` is stored and displayed
verbatim, which can render as a different extension in the list. There is no
security impact beyond display spoofing — the name is never used as a path, and
H08's stronger guarantees hold — but a name shown to a reader should not be able
to lie about itself. Stripping the `Cf` (format) category would close it.
---
### M7-F6 — Narrator state blocks can leak into displayed prose — **pre-existing, not M7**
**Severity:** low
**Owner:** later milestone (M8/M11), recorded for completeness.
Twice during real-narrator testing, `qwen2.5:3b-instruct` emitted its state
block fenced as ```` ```json ```` rather than ```` ```state ````, so
`narrative.extract.split` did not recognise it and the raw JSON was displayed to
the reader. This is an M5 protocol-robustness issue with a small model, entirely
outside M7's scope and unaffected by it. Noted because it was observed under
this review's real-narrator runs and a reader would see it.
---
## J. I05 and export/import
**PASS.** A genuine round trip: source server on one database file, destination
server as a **separate process on a different, empty database file**.
Before export: all three classes, one disabled source, one hidden Canon source
that is also always-include, one six-chunk source, and a real narrator turn that
used imported knowledge.
```text
E1 bundle carries a knowledge section (5 sources) PASS
E2 and no derived data (no chunks/vectors/index rows) PASS
E4 imports into a clean data directory PASS
E5 every source arrived PASS
E6 classification survived (all 5) PASS
E7 enabled/disabled state survived (all 5) PASS
E8 visibility survived (all 5) PASS
E9 content hash survived (all 5) PASS
E10 title and original filename survived (all 5) PASS
E11 content byte-identical (all 5) PASS
E12 chunks rebuilt, same count (all 5) PASS
E13 always-include survived PASS
E14 rebuilt chunks identical in text, order and hash PASS
E16 lexical retrieval works immediately, with no reindex PASS
E16b the disabled source stays out after import PASS
E17 semantic vectors rebuild locally on demand PASS
E17b semantic retrieval works in the imported campaign PASS
E18 exporting did not disturb the original turn's evidence PASS
E20 the story itself round-tripped PASS
E21 no external source path is needed PASS
```
I05's literal condition — metadata and classification survive — passes.
M7's broader scope condition — a *usable* local knowledge subsystem survives —
also passes: chunk identity is reproduced exactly by the deterministic chunker,
lexical retrieval works on arrival with no reindex step, and vectors rebuild
against whatever model the importing machine has.
**Recorded limit, correctly documented and not an M7 defect.** The bundle
carries no context snapshots, so the imported campaign has no historical prompt
provenance — for imported knowledge or for any other component. That is exactly
what a pre-M7 bundle already did for summaries, memories and everything else,
it is stated in `DATA-MODEL.md` §29 and `BUILD-MILESTONES.md` M7 § Status, and
it belongs to M9. Nothing M7 creates becomes a dangling reference, because no
ids are exported.
---
## K. F05 / F06 completion
**Both complete.** M6 left them partial only because imported knowledge did not
exist; it now does, and the gap is filled in the browser, not only the API.
For a completed narrator turn, the Insights panel showed — rendered, in Firefox:
```text
▸ hostile.md CANON · Hostile · passage 1 · hybrid match
(lexical 0.64, semantic 0.90 → 1.00) · 124 tok
<the full rendered chunk text>
2 passage(s) considered, 182 of 2112 knowledge tokens used
· searched on: aldric, studies, broken, circle, abbey, crypt, westhaven
```
F06's requirements, each checked in the rendered DOM: source filename/title ✓,
classification ✓, chunk identity ✓, rendered chunk text ✓, retrieval score and
mode ✓, token cost ✓, and why it was searched ✓.
F05's prompt-section view showed every component with a readable name and a
token cost:
```text
narrator prompt 57 tok · narrative state (reporting rule) 467 tok ·
campaign canon 109 tok · imported knowledge (rules for using it) 199 tok ·
ai instructions 140 tok · plot essentials 17 tok · story history 130 tok ·
imported canon (retrieved) 232 tok · narrative state (current) 62 tok ·
length guidance 51 tok · narrative state (emit reminder) 32 tok
```
Historical evidence survives deletion of its source: a turn was generated using
`canon.md`, the source was deleted, and the old turn's snapshot still carried the
rendered text byte-identically (§L, L15).
---
## L. Source lifecycle and atomicity
31 of 32 checks passed; the one failure was a defect in the review's own test.
```text
L1 .txt import indexes PASS
L2 .md import indexes PASS
L3 unsupported extension refused PASS
L4 binary content refused despite .txt extension PASS
L5 invalid UTF-8 refused, not mangled PASS
L5b empty source refused PASS
L6 1,040,382 bytes accepted (limit 1,048,576) PASS
L7 1,059,372 bytes refused, message says what to do PASS
L7b nothing of the refused source stored PASS
L8 NFC/NFD normalize to the same hash PASS
L8b the conflict names the existing source PASS
L8c a CRLF copy is detected as a duplicate PASS
L9 explicit duplicate override honoured PASS
L10 disable keeps content, chunks and FTS rows PASS
L11 re-enable restores retrieval, no reimport PASS
L12b no orphan FTS rows anywhere PASS
L12c no orphan chunks anywhere PASS
L12d every stored source is 'ready' — none half-built PASS
L14 delete removes source, chunks, FTS rows and vectors PASS
L14d no dangling FTS rowids after delete PASS
L15 the historical turn still says what it was supplied PASS
L15b byte-identical to what was supplied PASS
L15c the story was not rewritten by the delete PASS
```
**L12 was a review test error, not a product defect.** It expected a `.md` file
whose content is `%PDF x` to be refused. That content is valid readable text, so
accepting it is correct — `SECURITY-THREAT-MODEL.md` §21 asks for rejection of
*binary* data, which L4 demonstrates with a real ZIP header. The invariants that
matter (no orphans, nothing half-built) all pass.
Atomicity is additionally structural: retrieval is gated on
`index_state = 'ready'`, so even a hypothetical partial commit would be inert
rather than wrong.
---
## M. Ranking and authority
15 of 16 checks passed; the failure is M7-F1.
```text
R3.1 irrelevant Canon is not retrieved FAIL -> M7-F1
R3.1b relevant Reference IS retrieved (positive control) PASS
R3.1c relevant Reference outranks irrelevant Canon PASS (0.897 > 0.584)
R3.2 at comparable relevance, Canon scores above Reference PASS (1.133 > 0.897)
R3.2b Reference framed as non-canon wherever it lands PASS
R3.2c Inspiration framed as establishing nothing PASS
R3.3a the lexical path contributed a winner PASS
R3.3b the semantic path contributed a winner lexical missed PASS
R3.4 no chunk appears twice in the selection PASS
R3.4b a chunk both paths found is marked hybrid, not doubled PASS
R3.5 a disabled source is excluded before ranking PASS (considered 5->4)
R3.6 Campaign A's sources never appear in Campaign B PASS (considered=0)
R3.6b B retrieves its OWN source (positive control) PASS
R3.7 reclassifying moves the passage into the Canon section PASS
R3.7b and it is no longer framed as Reference PASS
```
Ordering, deduplication, class weighting, disabled-source exclusion and campaign
isolation are all correct. **Only the admission threshold is wrong.** That is a
narrow, well-localised fix.
---
## N. Local-only and network evidence
**PASS. No new network path.** Socket capture in the server process:
```text
import + chunk + FTS5 index destinations: none
lexical retrieval + prompt assembly destinations: none
a full narrator turn destinations: none
semantic indexing + retrieval destinations: the configured Ollama
host and port, and nothing else
(the configured Ollama, its configured port)
```
- A public endpoint written into the database **behind** the settings API is
refused at request time on the knowledge path, and the refusal opened no
socket.
- The policy refuses `api.openai.com` and public addresses, and permits loopback
and the configured trusted-LAN host.
- An import still succeeds lexically while the endpoint is refused — the library
does not depend on the inference host being reachable.
- The knowledge subsystem builds **no HTTP client of its own**, embeds only
through `memorybank.embedding_provider`, and reimplements no TLS handling, so
the allowlist, the request-time re-check and the OS/private-CA trust union all
apply unchanged (ADR 011, ADR 002).
- No remote vector service, cloud host, telemetry, CDN asset or model-download
path exists in it. (Two apparent hits in an initial scan were false positives:
"cohere" inside "coherent", "pulled" in prose.)
Consistent with ADR 004 and `SECURITY-THREAT-MODEL.md` §§12-21.
---
## O. Migration
**PASS — 23/23, through a real server restart.**
A campaign was built with history, a divergence, a Save Point, narrative state,
a summary, a memory and a derived-status row. The M7 tables and the FTS index
were then dropped and the stamp rewound to 91, making the file genuinely M6 —
verified before migrating. The server process was stopped and a **new process**
opened the same file.
```text
stamp 91 -> 92 on restart
every M7 table created, including knowledge_fts
story history identical
branches identical
Save Points identical
narrative state identical
memories identical
summaries + derived status identical
active head unmoved (2, 6)
pre-M7 prompt provenance readable and unchanged
old snapshots no knowledge section forced onto them
empty library, no imported section in the next prompt
normal turn / Undo / Redo / Save Point restore / inspection all work
then: import a source and retrieve it in a real narrator turn PASS
```
An old campaign needs no knowledge sources, and gains the subsystem cleanly.
---
## P. Browser evidence
Firefox 154.0.1, headless, WebDriver, against the built frontend served by the
real backend. **22/22 passed**, covering all three surfaces that render imported
text plus a full page reload. See §H and §K for the specific results.
No severe console errors. No request to any non-origin host across the entire
session.
The implementation's own browser suite (43 checks) was also re-run during the
implementation pass; this review's suite is independent of it.
---
## Q. Performance and scaling
**PASS — measured, flat.**
```text
sources/chunks list detail retrieve context considered used ctx
3 / 12 5q 4q 9q 20q 12 3 130ms
15 / 60 5q 4q 9q 20q 55 3 108ms
39 / 156 5q 4q 9q 20q 67 3 188ms
```
Query counts did not move while the library grew 13×. No N+1 on the source list,
the chunk listing, retrieval or prompt construction. Candidates are bounded in
SQL (67 of 156 chunks, cap 80); the FTS query carries a `LIMIT`; the semantic
catalogue read fetches ids only, not vectors or passage text; the vector read is
a single batched query.
**Characteristic to record, not a defect:** the semantic scan is a linear cosine
over the campaign's vectors in Python, capped at 4,000 passages with the
shortfall reported. Query count is flat; per-turn CPU grows linearly with library
size. `BUILD-MILESTONES.md` M7 § Status already records this as deliberate v1
debt (no approximate-nearest-neighbour index).
---
## R. Regression gates
| Gate | Result |
| --- | --- |
| backend suite | **920 passed, 10 skipped, 1 warning** (M6 baseline 836/7/1) |
| focused M3-M6 regressions | **372 passed, 1 warning** — context/memory, memory retrieval, head cursor, Save Points, narrative state, bundle v2, tree migration, endpoint policy, local-only surface, egress, state revert, pre-M5 compatibility |
| frontend lint | **exit 0, 7 warnings** — identical to the M6 baseline |
| frontend build | **success** |
| docker build | **success**; image verified running with `python-multipart 0.0.32`, schema 92, porter-stemmed FTS |
The single warning is the pre-existing `StarletteDeprecationWarning` from
`fastapi.testclient`. The 7 lint warnings are the pre-existing
`components.jsx`/`SchemaEditor.jsx` fast-refresh notes. The 3 additional skips
over the M6 baseline are `test_knowledge_real_model.py`, which skips without a
configured endpoint; they were run explicitly in this review with one.
---
## S. New dependency
**Approved.**
```text
package python-multipart 0.0.32
licence Apache-2.0
requires nothing (Requires-Dist: None)
files 24
declared backend/requirements.txt python-multipart>=0.0.9 (with reason)
pinned backend/requirements.lock python-multipart==0.0.32
docker installs from requirements.txt — verified in the built image
clean venv installs from requirements.lock — verified, app imports
```
It is Starlette's own multipart parser, needed because the upload surface is
`UploadFile`. It adds no transitive dependency, no framework surface and no
network path — and the upload surface it enables is precisely why no endpoint
accepts a filesystem path (H08). Reason, pinning and reproducibility all check
out.
---
## T. Deferred features — confirmed acceptable
None of these fails M7; each remains consistent with the active design and is
recorded in `BUILD-MILESTONES.md` M7 § Status and `IMPORTED-KNOWLEDGE-DESIGN.md`
§76.
- **Entity linking** (§33), **manual priority** (§35), **scene pinning** (§36) —
optional in the design; §36's stated v1 minimum (always-include for critical
Canon) *is* implemented.
- **Source version history** (§14, §43) — the design explicitly permits "a
simpler replace-with-provenance model … if documented", and it is documented.
- **Canon scope metadata** (§45) — §45 itself says explicit current-state
precedence may be sufficient for v1, which is what was built.
- **Tags** (§34) — evaluated in context as instructed. They are **unused optional
metadata**: nothing in the API, the source inspector, the prompt or the docs
promises tags, so their absence breaks no implemented promise. Acceptable.
- **Canon-vs-Canon conflict detection** (§42) — absent, and outside M7's scope
list. It produced **no authority failure** in any required M7 scenario in this
review. Recorded as later-v1 debt.
No planning document was weakened to match the implementation.
---
## U. Planning-document recommendations
1. **`IMPORTED-KNOWLEDGE-DESIGN.md` §29-30** — add, from M7-F1, that a relevance
threshold expressed only as a share of the best candidate cannot exclude
anything, because the best candidate always clears a share of itself. A
retrieval design needs an absolute notion of "no match" on each path, and for
embeddings that value is model-dependent and must be measured and written
down, as `memorybank.REDUNDANT_SIMILARITY` already is.
2. **`TECHNICAL-DESIGN.md` §18.1** — extend the M2 wiring rule with M7-F2: a
stub provider must not be *more* discriminative than the real one. A stub
with no similarity floor cannot detect a retrieval defect that only exists
because real embeddings have one. State the requirement that retrieval
thresholds be validated against a real model.
3. **`V1-ACCEPTANCE-TESTS.md` G05-G07** — add the negative control as a pass
condition: a query unrelated to every source must retrieve nothing. Without
it, G05/G06/G07 can all pass while retrieval is unconditional.
4. **`V1-ACCEPTANCE-TESTS.md` §5 / `TEST-CAMPAIGN-FIXTURE.md` §12** — the two
documents give different contents for the same three filenames. §12 is the
richer and is what the traps are built on. Reconcile them, or state which is
normative for the G-series.
5. **`CONTEXT-AND-MEMORY.md` §46** — record from §G that hidden Canon's
protection is prompt discipline *plus* retrieval scoping; if retrieval is
unconditional, prompt discipline is carrying the whole load on every turn.
---
## V. M8 implications
- The knowledge panel and the Insights knowledge rows are **functional and
complete for M7's purposes**, and every M7 behaviour is reachable in a browser
with no database access. M8 owns their design.
- M8 should not have to touch retrieval; M7-F1's fix belongs to the M7
corrective pass, before M8 begins.
- Two M5 section labels (`state_rule`, `state_reminder`) that rendered as raw
keys were fixed during the M7 pass; the rest of the panel's visual design is
M8's.
- §47 of `BROWSER-UX-SPEC.md` describes an import *preview* step that is not
built, and §52's retrieval-usage count is absent. Both are marked optional;
M8 may reconsider.
- If M7-F1's fix means retrieval sometimes returns nothing, the panel and the
inspector should say so plainly rather than showing an empty section — a
small M8 affordance worth planning for.
---
## W. Verdict
**PASS WITH CORRECTIVE WORK REQUIRED**
M7 delivers the subsystem, and the parts that work are demonstrated rather than
asserted. Against a real narrator, a real embedding model, a real browser, real
socket capture and a real migration, the following all hold: the full source
lifecycle; classification that changes authority rather than metadata;
campaign isolation; hybrid retrieval with correct ranking, deduplication and
disabled-source exclusion; a bounded knowledge budget; hidden Canon that reaches
the narrator without reaching the protagonist; prompt injection framed as data
and ignored by the model; no network path but the approved one; no filesystem
path at all; export/import into a clean directory; a clean migration from a
genuine M6 database; and flat query counts as the library grows.
**The five acceptance conditions the implementation flagged as unmeasured — C05,
G06, G07, G10 and hidden Canon — were measured here against a real narrator and
all five pass.** That was the largest open risk and it is closed.
Two blocking items remain, and they are the same defect and the reason it was
not caught:
- **M7-F1** — retrieval admits imported knowledge unconditionally, because the
only effective threshold is a share of the best candidate. Irrelevant Canon,
Reference, Inspiration *and hidden Canon* are injected into every turn.
- **M7-F2** — the retrieval-quality suite's stub embedder is more discriminative
than the real model, so it cannot detect M7-F1 and cannot validate its fix.
Both are narrow and localised: the ranking is correct, the pipeline is correct,
and the fix is to the admission threshold and to the fixture that tests it.
**M7 is not accepted. M8 is not authorized.** Recommended next step is an M7
corrective pass addressing M7-F1 and M7-F2, with M7-F3, M7-F4 and M7-F5 folded
in, followed by re-review of §I, §E (G05-G07) and §G.
---
## X. Repository state at the end of the review
```text
branch m7-imported-knowledge
HEAD a6e9c7a32bdf42f1e4cb837b70721d89422cef8d (unchanged)
staged 48 files (47 from M7 + this report)
unstaged 0
untracked 0
LICENSE unchanged, not staged
prepared commit message .git/M7_MSG, unchanged
```
**Files modified by the review itself:** exactly one —
`planning/reports/M7-IMPLEMENTATION-REPORT.md`, which is this report, newly
created. No application code, no test, no planning document and no configuration
was changed by the review. All review harnesses, fixtures, databases and browser
artifacts were created outside the repository, under the session scratchpad, and
none of them is staged.
M6's report remains at `planning/reports/M6-IMPLEMENTATION-REPORT.md`; per the
package convention it rotates to `planning/archive/milestone-reports/` when M7 is
accepted, which has not happened.
> *Closeout note, added later:* it has now. M6's report is at
> `planning/archive/milestone-reports/M6-IMPLEMENTATION-REPORT.md`. This
> paragraph is left as written because it recorded the state at review time.*
---
---
# M7 Corrective Pass — Closeout
**Date:** 2026-09-06
**Scope:** the two blocking findings above, M7-F1 and M7-F2, plus the
non-blocking findings where they were directly related and low-risk.
**Nothing above this line has been changed.** The original findings, their
reproductions and their evidence stand as written.
---
## CA. What was wrong, reproduced on the staged tree
Both findings were reproduced before any code changed, against the real
embedding model, so the before/after comparison is like for like.
### CA.1 — M7-F1, the no-match case
```text
scene: "Aldric studies the tide tables and the harbour master's ledger
of container tonnage." (nothing in the library relates to it)
BEFORE considered=5 floor=0.2875 semantic=True
USED hidden-key.md canon lex=1.000 sem=1.000 cos=0.548 score=1.150 51 tok
USED canon.md canon lex=0.000 sem=0.772 cos=0.423 score=0.772 56 tok
USED reference.md reference lex=0.895 sem=0.927 cos=0.508 score=0.902 70 tok
USED revival.md reference lex=0.000 sem=0.730 cos=0.400 score=0.621 86 tok
USED inspiration.md inspiration lex=0.000 sem=0.871 cos=0.477 score=0.610 63 tok
-> 5 of 5 sources injected, 326 tokens
-> imported sections in the prompt:
['imported_inspiration', 'imported_reference', 'imported_canon']
```
Note the raw cosines: 0.400-0.548, i.e. nothing matched. Normalization lifted
the best to 1.000 and the relative floor (0.2875) admitted everything.
### CA.2 — M7-F1, irrelevant admitted *alongside* relevant
```text
scene: the abbey crypt library: abbey Canon + container-shipping Canon
BEFORE USED canon.md canon cos=0.697 score=1.150 <- relevant
USED shipping.md canon cos=0.320 score=0.459 <- container vessels
USED reference.md ref cos=0.486 score=0.593
```
### CA.3 — M7-F1, lexical leak on a generic word
```text
scene: "Ballast trim on the ore hauler was recalculated before the burn to Ceres."
library: hidden-key.md ("...Edrin discovered this before he disappeared...")
semantic: OFF
BEFORE terms=[... 'before' ...]
USED hidden-key.md canon lex=1.000 -> 1 injected, 51 tokens
```
The single shared token was the function word **"before"**, which the 42-word
stop list did not contain and which nothing downstream weighed.
### CA.4 — M7-F1, lexical leak on the protagonist's name
```text
scene: "Aldric studies the harbour ledger." semantic: OFF
BEFORE terms=['aldric','studies','harbour','ledger']
USED hidden-key.md canon lex=1.000 -> 1 injected, 51 tokens
```
`hidden-key.md` contains "Aldric does not know what the key opens". The
protagonist's name is in the story tail of essentially every query, so this
admits on a token that carries no information about the current scene.
### CA.5 — M7-F2, why the suite could not see any of it
Measured on identical texts, stub against the production model:
```text
stub cosine normalized | real cosine normalized
abbey-canon (relevant) 0.7559 1.000 | 0.7505 1.000
crypts-ref (relevant) 0.3392 0.449 | 0.6237 0.831
ship-canon (IRRELEVANT) 0.1499 0.198 | 0.4365 0.582
cookery (IRRELEVANT) 0.0818 0.108 | 0.4354 0.580
desert (IRRELEVANT) 0.0642 0.085 | 0.4349 0.580
```
The old floor was 0.25 of the best. Irrelevant material peaked at 0.198 under
the stub and at 0.582 under the real model — excluded in the fixture, admitted
in production.
---
## CB. Root cause
Two symptoms, one architectural cause: **M7 had a ranking stage and no
admission stage.**
Both scores were normalized against the best candidate of their own path, and
the floor was `max(0.02, best × 0.25)`. A floor expressed as a share of the best
cannot reject anything, because the best clears a share of itself by
construction. `MIN_RELEVANCE = 0.02` could only bind when the best score fell
below 0.08, which no real embedding produces, so it was unreachable — finding
M7-F3 was not a separate defect but the same one seen from the side.
Because the semantic path scores *every* embedded chunk, there was always a
best. **With semantic retrieval enabled, at least one passage was admitted on
every turn, unconditionally.**
The lexical path had the mirror-image problem for a different reason: FTS5 terms
were joined with `OR` and nothing asked how much had matched, so one incidental
token was sufficient.
---
## CC. The correction
### CC.1 Admission separated from ranking
```text
candidate generation
-> ADMISSION absolute, per path, independent of the candidate set
-> RANKING normalized, among the survivors only
-> class weighting
-> budget
```
Admission reads **raw** signals; ranking reads **normalized** ones. Normalization
still exists — `bm25` has no fixed range and cosine's zero is not zero, so the
paths are not otherwise comparable — but it now decides *order among things that
matched*, never *whether anything matched*. Authority is applied in the ranking
stage only, so a class can order what matched and can never rescue what did not.
Recorded as a durable architectural lesson in `TECHNICAL-DESIGN.md` §13.2 and
summarised in `IMPORTED-KNOWLEDGE-DESIGN.md` §76.
### CC.2 Semantic admission — `classes.SEMANTIC_FLOOR = 0.58`
Established empirically **through the production path**: passages embedded as
`fts.index_line(heading, text)`, queries assembled by `retrieval.query_terms`,
because both differ from bare strings and both move the numbers. The first
calibration attempt used bare strings and under-measured by ~0.10, which is
recorded here because it is exactly the kind of error that produces a plausible
wrong constant.
113 pairs against `nomic-embed-text`:
```text
targeted n= 13 min 0.5526 p10 0.6090 median 0.7231 max 0.8474
the one source a scene is actually about
off-topic n=100 min 0.3577 median 0.4591 p95 0.5339 max 0.5578
20 scenes unrelated to the campaign (harbour, surgery, compiler,
fugue, sourdough, kiln, football, photosynthesis, …)
sweep: 0.55 -> 13/13 targeted, 1/100 off-topic
0.56 -> 12/13 targeted, 0/100 off-topic
0.58 -> 12/13 targeted, 0/100 off-topic <- chosen
0.61 -> 11/13 targeted, 0/100 off-topic (loses C05's revival match)
```
0.58 sits between the two populations with ~0.022 of margin on each side. The
single targeted pair below it is instructive rather than lost: *"the broken
circle cut into the keystone above the crypt stair"* scores 0.5526 against the
Canon describing exactly that — the wording is so close that little is left for
the embedding to add — and it matches four lexical terms, so the lexical path
admits it. That is the hybrid working, and it is why neither path has to be
right alone.
The value is a property of the model and is documented as such beside the
constant, in the same style as `memorybank.REDUNDANT_SIMILARITY`. A model that
scored everything below it would degrade the library to lexical-only, which is a
supported production path, so the failure mode is safe.
### CC.3 Lexical admission
A passage is lexically admitted when it matches **two distinct meaningful query
terms**, or **one** term that is both (a) not the name of a standing campaign
entity and (b) not a negligible share of the query.
- **The stop list grew from 42 to 261 words**, all function words and
contentless generics. It contains no subject matter — no "abbey", "key" or
"crypt" — because a stop list that removes subject matter is how a search
stops finding "The Silver Key". `before` is now filtered.
- **Standing entities** are the protagonist and the campaign's established
entities, drawn from `persona_name` and the authoritative state. They are in
the query on every turn by construction, so a lone match on one says nothing
about the present scene. This is deliberately **not** "ignore proper nouns" —
`Westhaven`, `broken-circle` and `resurrect` all still count, and a standing
entity counts the moment a second term matches beside it.
- **The share test** covers the case the entity test cannot see: a young
campaign whose state is still empty. One word out of a nine-word scene is not
evidence however distinctive it is; one out of three is.
Both conditions are needed and each was measured failing alone: the share test
alone would admit a lone "Aldric" from a three-word query, and the entity test
alone admitted it from a nine-word one — which is what CA.4 measured.
Evidence per term comes from `fts.term_evidence`, one statement with a subquery
per term (capped at 24), scoped once at the join. Stemming is applied by FTS
itself, so `resurrected` matches `resurrection` exactly as the ranking query
does rather than through a second, divergent stemmer in Python.
### CC.4 M7-F3 folded in
`MIN_RELEVANCE` and `RELEVANCE_FLOOR_SHARE` are **removed**, not re-tuned. The
architecture that made them meaningless is gone, and nothing is left that
appears to enforce relevance without doing so.
---
## CD. Real-model before/after
Same reproductions, same model, same fixtures.
| Case | Before | After |
| --- | --- | --- |
| Off-topic scene (tide tables/tonnage), semantic ON | **5 sources, 326 tokens**, 3 imported sections | **0 sources, 0 tokens, no imported section** |
| Relevant scene + irrelevant Canon in library | relevant + **shipping Canon** + Reference | **relevant Canon only** |
| Lexical leak on "before" | 1 source, 51 tokens | **0** |
| Lexical leak on "Aldric" | 1 source, 51 tokens | **0** |
| Positive control: Old Abbey question | canon.md **+ shipping.md** | **canon.md only** |
The after-state reports the rejection explicitly rather than silently:
`generated=5 rejected=5 considered=0 floor=0.58`.
---
## CE. The §12 acceptance matrix — measured against the real model
**13/13.**
| Query | Library | Expected | Result |
| --- | --- | --- | --- |
| Old Abbey / broken-circle | relevant Canon + unrelated | relevant Canon only | **PASS** — `['canon.md']` |
| tavern description | tavern Reference + unrelated Canon | Reference, not irrelevant Canon | **PASS** — Reference in, shipping Canon out |
| resurrection question | conflicting Canon and Reference | Canon governs | **PASS** — campaign rule present above the framed Reference |
| tide tables / container tonnage | fantasy fixture only | **zero imported chunks** | **PASS** — `used=[] sections=[] generated=4 rejected=4` |
| distinctive exact term | embeddings unavailable | lexical succeeds | **PASS** — `['sigils.md']`, `semantic_used=False` |
| semantic paraphrase | little lexical overlap | semantic succeeds | **PASS** — admitted_by `semantic`, cosine 0.6575 |
| protagonist name only | many Aldric-containing sources | no unrelated flood | **PASS** — `used=[]` |
| stopword-heavy input | mixed library | no incidental FTS admission | **PASS** — `used=[]`, `terms=[]` |
Plus three beyond the matrix:
- **hidden Canon when relevant** → retrieved.
- **hidden Canon when unrelated** → absent (`used=[] generated=1`).
- **always-include Canon** → still supplied on an unrelated scene, because it is
asserted rather than retrieved and is deliberately outside admission.
For every no-match row the assembled prompt was inspected: the imported-knowledge
sections are **absent**, not empty, and the untrusted-data rule section is absent
with them.
---
## CF. The four hybrid cases
Permanent deterministic coverage in `test_knowledge_retrieval_quality.py`:
```text
A strong semantic / weak lexical ossuary passage, no shared vocabulary
-> retrieved, admitted_by="semantic"
B strong lexical / weak semantic distinctive term, embeddings unavailable
-> retrieved, admitted_by="lexical"
C both strong -> one candidate, admitted_by="both",
not duplicated
D neither strong -> ZERO chunks, and no imported section
in the prompt (mandatory case)
```
Case D is covered twice — with embeddings on and on the lexical-only path — and
each assertion is guarded by `generated > 0`, so it cannot pass because nothing
was searched.
---
## CG. M7-F2 — the test doubles
`RealisticEmbedder` replaces the old stub. It has a **deliberate similarity
floor**: every pair shares a constant component, so unrelated passages score a
substantial similarity as they do in life, with disjoint topic axes above it and
a hashed bag of words underneath.
Two tests exist purely to keep the fixture honest:
- `test_the_stub_models_the_real_problem` fails if unrelated pairs ever drop
near zero again — that is, if the fixture drifts back to the one that hid the
defect.
- `test_a_relative_only_floor_would_admit_the_irrelevant_set` reproduces the
**superseded rule** on the corrected code's own vectors and asserts it still
admits the entire irrelevant library. If this ever stops failing the old rule,
the fixture has lost the ability to detect the defect.
The stub is not calibrated to 0.58. It models the shape — a floor under
everything, structure above it — and the tests assert against
`classes.SEMANTIC_FLOOR` rather than against a number baked into the fixture.
**Validation that F2 is actually closed:** the corrected implementation was
temporarily reverted to the staged pre-corrective version and the new suite run
against it:
```text
13 failed, 5 passed
```
including every no-match test, both case-D tests, all three
"irrelevant material is excluded whatever its class" parametrisations, and
"Canon is excluded even though it is the best candidate". The suite can now see
the defect it previously could not.
---
## CH. Permanent real-model regression
`test_knowledge_real_model.py` gained four tests, skipped without an endpoint:
```text
test_the_real_model_separates_relevant_from_unrelated
re-measures both populations and fails if SEMANTIC_FLOOR stops sitting
between them. Measured this run:
targeted 0.8130 / 0.7260 / 0.6726
off-topic 0.3293 - 0.5026 (10 pairs, 5 scenes x 2 sources)
floor 0.58
test_a_completely_unrelated_query_retrieves_nothing_from_a_real_model
generated=3 rejected=3 used=0, no imported section
test_a_relevant_query_still_retrieves_from_a_real_model -> ['canon.md']
test_a_paraphrase_still_retrieves_from_a_real_model -> cosine 0.7260
```
The first of these caught a real error in its own first draft: a single
off-topic line appended after a crypt opening still leaves the crypt in the
four-action query window, so the query was not off-topic at all. The window is
now moved wholesale. Recorded because it is a property of the design worth
knowing: **relevance is judged against the recent scene including its immediate
history**, not against the last sentence alone.
---
## CI. No regression in established M7 behaviour
| Area | Result |
| --- | --- |
| C05, real narrator | **PASS** — *"the dead do not return … magic cannot bring him back to life"*, with `revival.md` still retrieved at cosine 0.656, so the test remains non-vacuous |
| G06 / G07 / G10, real narrator | **PASS** (25/25 with C05 and hidden Canon) |
| Hidden Canon, real narrator | **PASS** — retrieved, marked narrator-only, not revealed |
| G01-G10 acceptance suite | **49 passed** |
| H06 / H07 / H08 / H09 | **PASS** — 43/43 browser checks, unchanged |
| I05 export/import | **53/54** (the one failure is the review harness's own false positive: `text/markdown` contains a slash) |
| F05 / F06 | **PASS** — browser-verified, now also showing matched terms and similarity |
| Source lifecycle, atomicity, isolation, disable/delete | **PASS** |
| Local-only network boundary | **PASS** — zero sockets for import/index/retrieval/turn; embeddings only to the configured endpoint; request-time policy enforced |
| Query counts (3 → 39 sources) | list 5, detail 4, retrieve 10, context 21 — **flat**; retrieval and context each gained exactly one statement (the term-evidence query) |
---
## CJ. Regression gates
| Gate | Result |
| --- | --- |
| backend suite | **927 passed, 14 skipped, 1 warning** (pre-corrective 920/10/1) |
| M7 knowledge suites | **49 + 18 + 14 + 6 + 4 passed** |
| real-model M7 suite | **7 passed** against `nomic-embed-text` on the trusted-LAN Ollama |
| frontend lint | **exit 0, 7 warnings** — identical to the M6 baseline |
| frontend build | **success** |
| docker build | **success** |
| browser | **43/43**, Firefox 154.0.1 |
| network | unchanged; see CI |
| migration, pre-M7 database via a real server restart | **23/23** |
| export/import into a clean data directory | **53/54** (one review-harness false positive) |
| query counts, 3 → 39 sources | list 5, detail 4, retrieve 10, context 21 — flat |
| §12 acceptance matrix, real model | **13/13** |
| real-narrator authority (C05, G06, G07, G10, hidden) | **25/25** |
The +7 tests are the new admission and real-model tests; the +4 skips are the
new real-model tests, which skip without an endpoint. The single warning is the
pre-existing `StarletteDeprecationWarning`.
No schema change and no serialized-settings change: the corrective pass altered
the *shape of the context-snapshot knowledge record* (`floor` → `generated`,
`rejected`, `semantic_floor`, plus `admitted_by` and `matched_terms` per
passage) but added no column and no migration. Snapshots written before the
corrective pass still render — the browser reads the new fields with numeric
comparisons that are false when absent.
---
## CK. Disposition of the non-blocking findings
| Finding | Disposition |
| --- | --- |
| **M7-F3** — `MIN_RELEVANCE` unreachable dead config | **Fixed.** Removed along with `RELEVANCE_FLOOR_SHARE`; the architecture that made them meaningless is gone. Nothing now appears to enforce relevance without doing so. |
| **M7-F4** — acceptance tests used the wrong fixture | **Fixed.** `test_imported_knowledge.py` now uses `TEST-CAMPAIGN-FIXTURE.md` §12 verbatim, plus the §11 campaign canon and §8 hidden Canon. G07 exercises the actual trap — the frightened-innkeeper passage against the "Mara is not a spy" rule — and asserts both the framing and that no spy fact reaches authoritative state. The two suites now state which they are: `test_imported_knowledge.py` is the standard acceptance fixture, `test_knowledge_retrieval_quality.py` is purpose-built retrieval mechanism. |
| **M7-F5** — RTL override in filenames | **Fixed.** `safe_filename` drops Unicode format characters (category `Cf`), so `"‮exe.dm.md"` stores as `exe.dm.md` and a name cannot visually impersonate another extension. Cheap, in-scope, no UI redesign. |
| **M7-F6** — narrator emits ```` ```json ```` instead of ```` ```state ```` | **Not absorbed.** Pre-existing M5 protocol robustness with a small model; M7 neither caused nor worsened it. Ownership stays with M5/M8/M11. Observed again during this pass and left recorded. |
---
## CL. Planning documents updated
Only where the corrective work established a durable fact:
- **`TECHNICAL-DESIGN.md` §13.1** — the pipeline diagram now shows the admission
stage; **§13.2 (new)** records the architectural lesson and the general rule
that a relevance decision must rest on a signal meaningful on its own.
- **`IMPORTED-KNOWLEDGE-DESIGN.md` §76** — the retrieval subsection records
admission before authority, and states plainly that retrieval may return
nothing.
- **`BUILD-MILESTONES.md` M7 § Status** — both findings, their corrections, and
the model-specific calibration recorded as carried debt.
No product requirement was weakened, and no ADR was added: the change is an
implementation architecture that the two design documents already govern.
---
## CM. Verdict
**READY FOR M7 CLOSEOUT**
Both blocking findings are closed with real-model and deterministic evidence:
- **M7-F1 is closed.** The required invariant — *retrieval must be allowed to
return zero imported chunks* — holds and is measured. Four off-topic scenes
that previously injected the entire library now inject nothing, verified at
the assembled prompt rather than at the ranking. Relevance is decided before
authority, on signals that mean something on their own, and irrelevant Canon
is excluded even when it is the only candidate. Every positive control still
retrieves: exact terms, paraphrases, lexical-only operation, C05 against a
real narrator, and hidden Canon when it is actually relevant.
- **M7-F2 is closed.** The stub now models the real model's similarity floor,
the suite fails 13/18 against the pre-corrective implementation, and two tests
exist solely to stop the fixture drifting back.
M7 is **not accepted** — that is the repository owner's decision on a signed
commit. M8 is not started.
---
## CN. Repository state at the end of the corrective pass
```text
branch m7-imported-knowledge
HEAD a6e9c7a32bdf42f1e4cb837b70721d89422cef8d (unchanged)
staged see below
unstaged 0
untracked 0
LICENSE unchanged, not staged
prepared commit message .git/M7_MSG, updated for the corrective pass
```
---
---
# M7 Closeout Verification
**Date:** 2026-09-06
**Scope:** the two items left open at the end of the corrective pass — the
embedding-model calibration boundary, and the ambiguous `export/import 53/54`.
**Nothing above this line has been changed.**
---
## DA. The `export/import 53/54`, resolved
| | |
| --- | --- |
| **Test** | `E3: no original filesystem path is present anywhere in the bundle` |
| **Where** | the independent review's own harness, not the repository suite |
| **Expected** | no exported field carries a filesystem path |
| **Actual** | the assertion fired |
| **Verdict** | **NOT A DEFECT — a false positive in the review harness** |
The assertion was a crude substring test: it serialised every non-`content`
field of every exported source and failed if the result contained a `/`. The
only slash present is in `"mediaType": "text/markdown"` — a MIME type, not a
path. No exported filename, title, hash or timestamp contains a separator.
The harness assertion has been replaced with three that test the actual
property:
```text
E3 no non-content, non-mediaType field contains "/" or "\" PASS
E3b every exported filename is a bare basename, no ".." PASS
E3c mediaType is one of text/plain, text/markdown PASS
```
**The export/import suite is now 56/56.** The earlier `53/54` is withdrawn as
an artefact of the harness; it never described product behaviour. I05 and the
stronger M7 conditions are unaffected and remain demonstrated:
```text
source contents byte-identical PASS
classification survived (all 5 sources) PASS
enabled state survived, disabled stays out PASS
visibility / hidden survived PASS
always-include survived PASS
provenance / hash survived PASS
chunk reconstruction identical text, order and hash PASS
lexical retrieval works on arrival, no reindex PASS
semantic rebuild on demand, local PASS
historical provenance the pre-export turn unaffected PASS
no external pathname required anywhere PASS
```
---
## DB. The embedding-model calibration boundary, resolved
### DB.1 The unsafe case
`SEMANTIC_FLOOR = 0.58` is a raw-cosine threshold measured against
`nomic-embed-text`. The corrective pass documented only the safe half of the
consequence — a model scoring everything lower degrades to lexical-only — and
left the dangerous half as carried debt:
> a model whose similarity scale sits **above** `nomic-embed-text`'s would put
> unrelated material past 0.58 and recreate M7-F1 exactly, on a build whose
> tests all pass.
That is no longer carried debt.
### DB.2 The policy
Semantic admission is **per model**, and a model this build has not measured
does not inherit a number measured against something else.
```python
SEMANTIC_CALIBRATION = {"nomic-embed-text": 0.58}
semantic_floor_for(model) -> float | None
```
- The key is the model's base name, lower-cased and without its tag: an Ollama
tag (`:latest`, `:v1.5`) selects a build of the same model and does not change
its similarity scale.
- `None` means "not measured here". Retrieval then **does not run the semantic
path at all** — it is skipped, not attempted and thresholded — and reports
why.
- Retrieval degrades to **lexical-only**, which is a first-class production path
rather than a fallback, so story play, imports, indexing, inspection and
turns are all unaffected.
- No new model-identity mechanism was introduced. `KnowledgeEmbedding.model`
already records what produced each vector and the retrieval catalogue already
filters on it; the policy reuses the configured `Settings.embedding_model`
and that same name.
The diagnostic names the model, says what is happening and lists what *is*
calibrated:
```text
The embedding model "mxbai-embed-large" has no measured relevance calibration
in this build, so semantic retrieval is disabled and retrieval is lexical only.
Story play and lexical search are unaffected. Calibrated models:
nomic-embed-text.
```
It is surfaced in three places: `semantic_note` on the retrieval result, the
`knowledge-status` endpoint (`semantic_calibrated`, `calibrated_models`,
`semantic_note`), and the Knowledge panel, which shows a warning row rather than
claiming semantic search is on.
`knowledge-status` now reports `semantic_enabled: false` for a configured but
uncalibrated model. That is the M6-F5 discipline applied once more: **vectors
that exist but are never consulted are not a working semantic index**, and
saying "on" because a model name is set would be the same untruth in a new
place.
### DB.3 What the policy costs, stated plainly
A semantic-only paraphrase — conceptually right, no shared vocabulary — is not
retrieved under an uncalibrated model. That is a **missing** passage rather than
an **irrelevant** one, and it is asserted in the suite rather than left to be
discovered.
> Semantic admission is calibrated for `nomic-embed-text`; uncalibrated
> embedding models fall back safely to lexical-only retrieval rather than
> borrowing its threshold.
Adding a model is a measurement, not a guess: run
`tests/test_knowledge_real_model.py` against it and confirm the targeted and
off-topic populations separate, as §CC.2 records for the existing entry.
Generic cross-model calibration is deliberately **not** attempted in v1.
### DB.4 Deterministic coverage
`tests/test_knowledge_calibration.py`, **12 tests**, needing no second model
installed — the policy is about model *identity*, so a configured name and a
stub are the whole apparatus.
The stub is deliberately hostile: `GenerousEmbedder` scores **everything** at
about 0.97, including the unrelated. It is the dangerous shape the policy exists
to defend against, and it makes the negative results meaningful — under the
calibrated name the fixture admits an unrelated freighter passage, and the only
thing that changes in the uncalibrated run is the model name.
| Required | Test | Result |
| --- | --- | --- |
| 1. `nomic-embed-text` uses the calibrated policy | `test_the_calibrated_model_resolves_to_the_measured_floor` (+ tags, whitespace) | PASS |
| 2. a different model does not inherit 0.58 | `test_an_uncalibrated_model_does_not_borrow_the_calibrated_threshold` | PASS |
| 3. unknown config degrades to an explicitly safe state | `test_an_uncalibrated_model_degrades_to_lexical_only_with_a_clear_reason`, `test_the_status_endpoint_reports_the_uncalibrated_state` | PASS |
| 4. distinctive lexical retrieval still works there | `test_distinctive_lexical_retrieval_still_works_when_uncalibrated` | PASS |
| 5. a semantic-only paraphrase is not spuriously admitted | `test_a_semantic_only_paraphrase_is_not_admitted_when_uncalibrated` | PASS |
| 6. no-match still returns zero chunks | `test_no_match_still_returns_zero_chunks_when_uncalibrated` / `_when_calibrated` | PASS |
| 7. a model change leaves no stale vectors live | `test_changing_the_model_does_not_leave_old_vectors_active`, `test_reindex_rebuilds_vectors_under_the_new_model` | PASS |
On (7) the existing machinery already held: `KnowledgeEmbedding.model` records
the producing model, the retrieval catalogue and the pending-work query both
filter on it, so after a model change the old vectors are invisible to
retrieval, `pending_embeddings` counts them as needing rebuild, and the
retrieval result says "no passages have been embedded yet" rather than silently
scoring incompatible vectors. Reindex rebuilds them under the new model. The
tests pin that behaviour rather than adding a second mechanism.
### DB.5 The same policy, through the live API
The deterministic tests use stubs to isolate the mechanism. It was also
exercised end to end against a real server, a real database and the real
narrator, changing **only** the configured model name between the two halves —
12/12:
```text
embedding_model = "nomic-embed-text" (calibrated)
status semantic_enabled=True semantic_calibrated=True
context semantic_used=True, retrieved ['canon.md']
embedding_model = "mxbai-embed-large" (not calibrated)
status semantic_calibrated=False, semantic_enabled=False
"“mxbai-embed-large” has no measured relevance calibration in this
build, so semantic retrieval is disabled and retrieval is lexical
only. Lexical search and story play are unaffected."
calibrated_models = ['nomic-embed-text']
context semantic_used=False, semantic_floor=0.0
retrieved ['canon.md'] — admitted_by "lexical"
off-topic scene: zero imported chunks, no imported section
a real narrator turn still plays
```
### DB.6 A consequence worth recording
Four M7 suites configured a placeholder embedding model name (`embed-test`).
Under the new policy that name is uncalibrated, so those suites would have
quietly exercised the lexical-only path while appearing to test semantic
retrieval — the M7-F2 failure mode in a new guise. They now name
`nomic-embed-text`, which is what their stubs are built to model, with a comment
saying why. `test_knowledge_calibration.py` is where the uncalibrated path is
exercised deliberately.
---
## DC. Final regression gate
| Gate | Result |
| --- | --- |
| backend suite | **939 passed, 14 skipped, 1 warning** |
| focused M7 knowledge suite | **103 passed** |
| M7 real-model suite (`nomic-embed-text`) | **7 passed** |
| M7 real-narrator authority tests | **25/25** |
| §12 relevance matrix, real model | **13/13** |
| export/import suite | **56/56** |
| migration suite | **23/23** |
| frontend lint | **exit 0, 7 warnings** (M6 baseline) |
| frontend production build | **success** |
| Docker build | **success**; image verified reporting the calibration registry |
| browser checks | **43/43**, Firefox 154.0.1 |
| network | zero sockets for import/index/retrieval/turn; embeddings only to the configured endpoint |
Confirmed once more, after the calibration change:
```text
irrelevant query -> zero imported knowledge PASS
relevant Canon only -> relevant Canon selected PASS
Reference does not become Canon PASS
Inspiration cannot establish conflicting facts PASS
C05 remains non-vacuous (revival.md retrieved at cosine 0.656) PASS
hidden Canon retrieved only when relevant, never disclosed PASS
lexical fallback works PASS
an uncalibrated embedding model cannot recreate M7-F1 PASS
no new outbound network path PASS
query growth bounded (list 5, detail 4, retrieve 10, ctx 21) PASS
```
---
## DD. Verdict
**M7 READY FOR SIGNED COMMIT**
Both blocking findings from the independent review are closed with real-model
and deterministic evidence, and both closeout items are resolved:
- **M7-F1** — retrieval admits nothing without candidate-set-independent
evidence, and returns zero chunks on an unrelated scene.
- **M7-F2** — the suite models the real model's similarity floor and fails 13/18
against the pre-corrective implementation.
- **Calibration boundary** — an uncalibrated embedding model degrades safely to
lexical-only instead of borrowing a threshold measured against another model.
- **`53/54`** — a review-harness false positive, withdrawn; the suite is 56/56.
M7 is complete and staged. It is **not committed**: the repository owner signs
milestone commits.
---
## DE. Corrections made during final staging
Three defects were found by the final repository verification itself, after the
regression gate in §DC. All three are documentation or commit-message only — no
application or test code was touched, so the §DC results stand unrepeated.
1. **The prepared commit message carried the corrective pass's test count.** It
read `927 passed, 14 skipped` and `91 new across six files`; the closeout run
measures `939 passed, 14 skipped`, and the seventh new file
(`test_knowledge_calibration.py`, 12 tests) brings the new total to 110. The
numbers reconcile exactly: 927 + 12 = 939, and 836 + 110 − 7 real-model skips
= 939 passed with 7 + 7 = 14 skipped. Corrected.
2. **The required documentation line existed only in this report.** The closeout
brief specifies the sentence *"Semantic admission is calibrated for
`nomic-embed-text`; uncalibrated embedding models fall back safely rather than
borrowing its threshold."* §13.3 of `TECHNICAL-DESIGN.md` explained the policy
at length but never stated it in that form, so a reader skimming headings
could miss the invariant. It is now the section's opening line.
3. **`BUILD-MILESTONES.md` line 3 still read "In implementation … M7 — next to
brief".** The M7 section body had been updated to `Status: COMPLETE` and M8
authorized, but the document's own top-level status line had not, so the file
contradicted itself. Corrected to `M1-M7 complete and accepted … M8 next to
brief`.
The third is the one worth noting as a pattern: a status recorded in two places
drifts, and the copy that drifts is the one not being edited during the work.