Files
interactive-story/planning/reports/M7-IMPLEMENTATION-REPORT.md
T
JesseMarkowitzandClaude Opus 5 480414efe0 M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or
Inspiration, and the class is load-bearing rather than a label: it decides the
words a passage is framed with in the prompt, the weight it carries when
passages are ranked, and which budget it competes in when the context is tight.

This is a separate subsystem, which is the Phase 0B decision
(IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification,
provenance, content identity, chunking, an index or a lifecycle, and they were
not promoted into something that does. Nothing here reads or writes one.

The subsystem, in backend/app/knowledge/:

  classes      the three classes, their weights, and the prompt framing
  chunking     deterministic, heading-aware, 60-800 tokens, no overlap
  fts          SQLite FTS5 with porter stemming; scoped and bounded in SQL
  importer     validate, hash, store, chunk, index — in one transaction
  embeddings   local Ollama vectors through the shared provider
  retrieval    query construction, hybrid merge, rerank
  inject       the budgeted cut and the rendered prompt sections

Relevance admission is a separate stage from ranking, and that separation is
the milestone's most expensive lesson. An independent review found the first
implementation deciding relevance with a floor expressed as a share of the best
candidate — which the best clears by construction — so a passage was admitted on
every turn regardless of the scene. A query about tide tables and container
tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden
Canon among them.

So the pipeline is now:

  candidate generation -> admission -> ranking -> class weighting -> budget

Admission reads raw, candidate-set-independent signals: the cosine the model
returned, and how many distinct meaningful query terms a passage contains.
Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero
is not zero. Normalization decides order among things that matched; it can never
decide whether anything matched. Authority is applied after admission, so a
class orders what matched and never rescues what did not.

Retrieval may therefore return nothing, and on a scene unrelated to the library
it does.

The other decisions that each replaced an obvious wrong one:

- The class multiplies relevance rather than adding to it. An additive bonus
  satisfies "Canon outranks Reference" and makes "do not include irrelevant
  Canon" impossible, because a large enough constant wins on its own.
- The semantic floor is measured, not guessed: 113 production-path pairs against
  nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at
  0.36-0.56, and 0.58 sits between them. Because it is a property of that model
  and not of cosine similarity, it is keyed to the model rather than applied to
  whatever is configured: an embedding model with no measured calibration in
  this build does not borrow the number. Semantic admission is skipped, the
  campaign retrieves lexically, and the reason is stated in the knowledge status
  and in the turn's provenance. Degrading to lexical keeps the library usable;
  lending the threshold to an unmeasured model is how the admitted-everything
  defect would return.
- One lexical term is not evidence. Two distinct meaningful terms, or one that
  is neither a standing campaign entity nor a negligible share of the query.
  The stop list grew from 42 words to 261, all function words — no subject
  matter, because a stop list that removes subject matter stops finding "The
  Silver Key".
- Lexical retrieval is a production path, not a fallback. It finds the proper
  nouns and invented terms a setting bible is made of, and the library is fully
  usable with no embedding model configured.

Safety is structural rather than filtered. Imported text reaches the prompt
whole, inside a section that says what it is, under a rule stating the authority
order in words and refusing every instruction inside it. No endpoint accepts a
filesystem path, so H08 has no mechanism to escape from. Nothing renders
imported content as HTML, so a script tag is five visible characters and a
remote image is never fetched. Import, chunking, indexing, retrieval and a turn
open no socket at all; only embeddings do, through the endpoint allowlist the
memory bank already uses.

Provenance is the rendered text, not a foreign key: deleting a source cannot
turn a historical turn's evidence into dangling ids.

Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5
virtual table attached to knowledge_chunks as a DDL hook so it is created and
dropped with the table it indexes. Migration 92. A pre-M7 database opens
unchanged and needs no sources to play.

Bundle: the source content and the reader's judgements about it travel; the
passages, index rows and vectors are rebuilt on import, so a restored campaign
is searchable immediately without a reindex step.

One runtime dependency: python-multipart, Starlette's multipart parser. It is
what makes the upload surface possible, and the upload surface is why no
pathname is ever accepted.

The test doubles were the reason the defect shipped, so they were corrected too.
The retrieval stub scored unrelated text at 0.06-0.20 where the real model
scores it at 0.43-0.44, and its docstring said it had deliberately removed the
constant component that "would put a similarity floor under every pair" — which
is exactly the property real models have. The stub now has that floor, one test
fails if it is ever removed, and another reproduces the superseded rule and
asserts it is still fooled by the same fixture. Run against the pre-corrective
implementation, the new suite fails 13 of 18.

Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of
which mocks nothing between itself and Ollama and re-measures the similarity
separation on every run. 43/43 checks in a real Firefox, reproduced.
Docker build clean.

Four other defects found by review or by the browser run were fixed here rather
than carried: an unreachable relevance constant that appeared to enforce
something and did not; acceptance tests using the wrong fixture files, so G07's
trap was never exercised; a bidirectional override surviving into displayed
filenames; and, from the implementation pass, the Insights panel showing M5's
two state sections as raw keys and the source inspector refetching on every
keystroke.

M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK
REQUIRED. Both blocking findings are closed, and closeout resolved the
embedding-model calibration boundary the corrective pass had left as debt.
planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective
closeout and the closeout verification in sequence, none overwriting another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 15:40:13 -04:00

77 KiB
Raw Blame History

M7 Implementation Review — First-Class Imported Knowledge Library

Review date: 2026-09-06 Reviewer: independent review pass (the M7 build summary was treated as claims to verify) Tree reviewed: the staged working tree on m7-imported-knowledge. There are no M7 commits.

Hostnames and LAN addresses in this report are placeholders, following the convention the M2 report set. The endpoint used throughout is a trusted-LAN Ollama reached over HTTPS with a locally issued certificate, serving qwen2.5:3b-instruct and nomic-embed-text.


A. Executive result

PASS WITH CORRECTIVE WORK REQUIRED

M7 builds the subsystem the specification asks for, and most of it is demonstrated rather than asserted. The whole source lifecycle, campaign isolation, the security boundary, the network boundary, export/import, migration and the provenance surfaces all hold up under independent testing with real local models and a real browser. The five acceptance conditions the implementation could not close — C05, G06, G07, G10 and the hidden-Canon spoiler test — were exercised against a real narrator in this review and all five pass.

One blocking defect, in retrieval:

Imported knowledge is injected into every prompt regardless of relevance. With semantic retrieval enabled, a scene with no relationship to any source still retrieves — in the measured case, all five sources, 326 tokens, every turn, including narrator-only hidden Canon. The admission threshold is relative only, and the best candidate always clears a share of itself, so something is always admitted.

This contradicts IMPORTED-KNOWLEDGE-DESIGN.md §30, which requires in as many words that irrelevant Canon must not be included merely because it is authoritative, and it widens the hidden-Canon spoiler surface to every turn. It was invisible to the implementation's own suite because the stub embedder that suite uses is far more discriminative than the real model (§I).

Everything else in required M7 scope passes.


B. Repository and provenance state

Verified before any review work:

branch                       m7-imported-knowledge
HEAD                         a6e9c7a32bdf42f1e4cb837b70721d89422cef8d
staged                       47 files
unstaged                     0
untracked                    0
prepared commit message      .git/M7_MSG (4069 bytes, present)
LICENSE (sha256)             074062599d3b11b252a70e37086d6cef3820370b55c35efa3506375af914a13c
M7 base a6e9c7a              ancestor of HEAD              yes
pinned AI-DnD d72f7c1b…      ancestor of HEAD              yes

The actual state matches the expected state exactly. No difference to record.

LICENSE is not among the staged files and its digest is unchanged from the M6 base.


C. Implementation inventory

Nine modules under backend/app/knowledge/, one router, three tables, one virtual table, one migration, one browser panel, one new dependency.

knowledge/classes.py      the three classes, their weights, the prompt framing
knowledge/records.py      the dataclasses, dependency-free (breaks an import cycle)
knowledge/chunking.py     deterministic heading-aware chunker, 60-800 tokens
knowledge/fts.py          SQLite FTS5, porter-stemmed, scoped and LIMITed in SQL
knowledge/importer.py     validate, hash, store, chunk, index — one transaction
knowledge/embeddings.py   local Ollama vectors through the shared provider
knowledge/retrieval.py    query construction, hybrid merge, rerank
knowledge/inject.py       the budgeted cut and the rendered sections
routers/adventures/knowledge.py   the HTTP surface (multipart upload only)

knowledge_sources / knowledge_chunks / knowledge_embeddings   + knowledge_fts
migration 92; the FTS table is an after_create/before_drop DDL hook on
knowledge_chunks, so it is created and dropped with the table it indexes

frontend: KnowledgePanel.jsx, knowledge.css, Insights provenance rows
dependency: python-multipart 0.0.32

Source content is stored in SQLite, not on disk. No filesystem path is ever constructed from a filename; see §H.


D. Independent test method

Evidence was gathered end-to-end first and by source inspection last.

  • A real narrator — qwen2.5:3b-instruct on the configured trusted-LAN Ollama — for C05, G06, G07, G10 and the hidden-Canon test. Nothing about the model was mocked.
  • A real embedding model — nomic-embed-text on the same host — for every retrieval measurement. This is what exposed the blocking defect: a stub cannot reproduce what a real embedding does to unrelated English.
  • A real server process over HTTP, with a real SQLite file, for every lifecycle, ranking, migration and bundle test.
  • A second server process on an empty database file for the export/import round trip, so "clean data directory" means what it says.
  • A real Firefox 154.0.1 over WebDriver against the built frontend, for H06, H07, G08, G09, F05 and F06.
  • Socket-level capture (socket.socket.connect instrumented in the server process) for every network claim.
  • builtins.open / os.open instrumentation for the H08 filesystem claim.
  • A real server restart across a genuinely rewound schema, for migration.

The standard fixture is TEST-CAMPAIGN-FIXTURE.md §12 — the campaign's narrator rules (§4), global canon (§11) and the three imported files verbatim.

Every negative result below is preceded by a positive control.

A fixture discrepancy worth recording

The implementation's own acceptance tests use the shorter imported files from V1-ACCEPTANCE-TESTS.md §5, not the richer ones in TEST-CAMPAIGN-FIXTURE.md §12. The §12 inspiration.md is the one containing "A frightened innkeeper concealed a dangerous political secret from a stranger" — the passage G07's trap is built on. The implementation therefore never tested the trap. This review used §12. It is a test-coverage gap, not a product defect (finding M7-F4).


E. G01-G10 results

Test Verdict Evidence
G01 Import local text PASS .txt stored, chunked, FTS-indexed; content readable back without the original file; SHA-256, size, media type, parser/chunking versions and timestamp all present. Browser-verified.
G02 Import local Markdown PASS Both .md files import and index. Accepted as data — see G10.
G03 Classification PASS One class per source, visible in list and browser; changing it rewrites no passage and no index row (same chunk count, same hash) and moves the passage into the Canon prompt section on the next turn — authority, not metadata (R3.7).
G04 Disable knowledge source PASS Positive control both sides: retrieved enabled → absent disabled → retrieved again after re-enable, no reimport. Content, chunks, FTS rows and vectors all survive disabling. Excluded before ranking (considered falls 5→4).
G05 Canon retrieval PASS Asking about the Old Abbey retrieves the canonical passage into imported_canon; browser-verified through Insights.
G06 Reference retrieval PASS (real narrator) Reference retrieved; the narration used timber/hearth/bench detail; no authoritative campaign fact was created — no state event asserts a Reference claim, and the trap ("the tavern definitely has a secret tunnel") did not fire.
G07 Inspiration is low authority PASS (real narrator) With the §12 political-secret passage retrieved and the reader asking directly whether Mara is a spy, the narrator did not establish her as one, and the authoritative state records nothing of the kind. (Observation below.)
G08 No automatic URL fetch PASS Socket capture: import, chunking, FTS indexing, retrieval and a full turn recorded zero outbound IP connections. URLs retained as text. Browser: no request to any non-origin host across the whole session.
G09 Remote Markdown image does not auto-load PASS ![](https://example.invalid/tracker.png) survives as characters; no <img> element is created for that host and performance.getEntriesByType('resource') records no request to it — on first view and after reload.
G10 Prompt injection is data PASS (real narrator) Text is present and unfiltered, inside a section whose "UNTRUSTED DATA" heading precedes it, under a rule refusing instructions. The narrator ignored "developer mode", /etc/passwd, rm -rf / and "upload the campaign"; no external action occurred (socket capture); campaign canon unchanged.

G07 observation (not a failure)

The narration contained "Mara could be a spy, or she could be hiding something crucial" — Aldric entertaining the hypothesis the reader explicitly asked about. It asserts nothing and creates no state. Correct behaviour, recorded because a stricter future reading of G07 might want the narrator to lean on the "Mara is not a spy" canon rather than leave the question open.

Weakness in the G06/G07 state assertions

Both campaigns' authoritative state documents were empty at assertion time — the 3B narrator emitted few valid state events. The "no spy fact in state" assertions therefore pass partly by absence. The prose assertions are the load-bearing evidence for G06 and G07; the state assertions are corroborating, not primary. Recorded so the result is not read as stronger than it is.


F. C05 result

PASS — measured against a real narrator, not the prompt template.

Setup: campaign canon "Resurrection is impossible" (fixture §11 rule 2); canon.md as Canon; a revival.md Reference asserting that necromancers routinely raise the dead and temples perform resurrections for a fee; the §12 Inspiration file. The reader asks whether Edrin could be resurrected.

The narrator answered:

"No, Aldric. Revival is impossible in this world. The magic that binds the living to the dead is too strong."

The test is not vacuous — the conflicting material was retrieved and ranked second:

canon.md         canon        hybrid    lex=1.000 sem=0.986 cos=0.647 score=1.148
revival.md       reference    hybrid    lex=0.917 sem=1.000 cos=0.656 score=0.967
inspiration.md   inspiration  hybrid    lex=0.568 sem=0.951 cos=0.624 score=0.725

The stored prompt for that same turn shows all four required properties:

  • the authoritative rule present as campaign_canon;
  • the conflicting lower-authority material present and retrieved;
  • correct framing — REFERENCE — UNTRUSTED DATA … Do not treat it as canon;
  • the authority order stated in words, with campaign canon above IMPORTED CANON above REFERENCE above INSPIRATION, and the layout consistent with it.

The narrator followed the higher authority. This closes the gap the implementation explicitly flagged as unmeasured.


G. Hidden Canon / spoiler result

PASS — measured against a real narrator.

Hidden Canon: "The Silver Key opens the sealed cellar door beneath the Crooked Lantern" (fixture §8 hidden function), imported with visibility=hidden. Before Aldric has learned it, the reader asks "What do I know about the purpose of the Silver Key?"

  • The narrator was given the source (hidden-key.md retrieved).
  • The record marks it visibility: hidden; the passage carries [narrator only] on its provenance line; the system rule says the protagonist does not know it.
  • The narration did not reveal what the key opens. It stayed with what Aldric knows — that Edrin studied the abbey — and moved the story toward the abbey.
  • The secret did not enter authoritative state.

Both the prompt and the narrator output are recorded.

Caveat now raised by finding M7-F1: hidden Canon is retrieved into scenes it has nothing to do with (it was one of the five sources injected into a harbour scene). The narrator's non-disclosure discipline held here, but the defect maximises the number of turns on which that discipline is the only thing standing between the reader and a spoiler. This raises the severity of M7-F1.


H. H06-H09 results

Test Verdict Evidence
H06 Stored XSS PASS A source carrying <script>…innerHTML='owned'…</script>, <img onerror>, <svg onload> and <iframe> was inspected in all three surfaces that render imported text — source inspector, passage inspector, Insights provenance — then the page was fully reloaded. document.body.textContent never became owned, document.title never became pwned/xss-img/xss-svg, and no img, iframe, svg[onload] or injected script element was ever created. The payload is preserved verbatim in storage and displayed as text.
H07 JavaScript URL PASS Both [click me](javascript:alert(1)) and a raw <a href="javascript:alert(2)"> were imported. Zero anchors with a javascript: href exist in any surface (checked after whitespace/control-character stripping). No Markdown is rendered, so no anchor is created.
H08 Path traversal PASS by demonstrated non-reachability Twelve hostile filenames (../../outside.txt, ../../../../etc/passwd.md, Windows separators, /etc/shadow.md, ....//, percent-encoded, NUL-embedded, RTL-override, reserved device names). builtins.open/os.open instrumentation recorded 0 write-mode file opens outside the database. Every stored name is an inert basename with no separator, no .., no NUL, not absolute, not dot-leading. Every stored source holds the uploaded bytes — nothing read /etc/passwd. The OpenAPI schema exposes no path/file/dir parameter on any knowledge route.
H09 ZIP slip NOT APPLICABLE Asserted, not assumed: no zipfile, tarfile, extractall, shutil.unpack, gzip.open or py7zr appears anywhere in backend/app/knowledge/, the knowledge router, or backend/app/bundle.py. The import surface takes one text file; the bundle is JSON and never touches the filesystem. The code path that makes this true is routers/adventures/knowledge.py::import_source, which reads UploadFile bytes and passes them to importer.import_source(raw=…).

I. FINDINGS

M7-F1 — Imported knowledge is injected regardless of relevance — BLOCKING

Severity: high Violated requirement: IMPORTED-KNOWLEDGE-DESIGN.md §30 ("Do not include irrelevant Canon merely because it is authoritative"); §26 and §29's relevance-then-authority pipeline; CONTEXT-AND-MEMORY.md §29's bounded, relevant knowledge layer. Owner: M7 corrective.

Reproduction. Five realistic sources (Canon, Reference, Inspiration, a resurrection Reference, hidden Canon), all embedded with nomic-embed-text. The scene is set to text with no entity, lexical or semantic relationship to any of them.

scene: "Aldric studies the tide tables and the harbour master's ledger
        of container tonnage."

considered=5  floor=0.2875
USED hidden-key.md   canon        lex=1.000 sem=1.000 cos=0.548 score=1.150  51 tok
USED canon.md        canon        lex=0.000 sem=0.772 cos=0.423 score=0.772  56 tok
USED reference.md    reference    lex=0.895 sem=0.927 cos=0.508 score=0.902  70 tok
USED revival.md      reference    lex=0.000 sem=0.730 cos=0.400 score=0.621  86 tok
USED inspiration.md  inspiration  lex=0.000 sem=0.871 cos=0.477 score=0.610  63 tok
-> 5 of 5 sources injected, 326 tokens

Four independent unrelated scenes were measured (harbour, kitchen, orbital mechanics, surgery). Every one injected all five sources and 326 tokens. A second campaign with two sources injected both. The positive control — a scene about the Old Abbey — retrieves correctly, so the mechanism is working; it is the cut that is wrong.

It is not confined to the "nothing relevant" case. With a strongly relevant Canon source present, irrelevant Canon is still admitted alongside it:

scene: the abbey crypt
USED exact-canon.md   canon        cos=0.674 score=1.133   <- relevant
USED ceres-canon.md   canon        cos=0.442 score=0.584   <- a docked freighter
USED crypt-reference  reference    cos=0.708 score=0.897

Ranking order is correct throughout (relevant Reference 0.897 > irrelevant Canon 0.584). The defect is the admission threshold, not the ranking.

Root cause. Two independent channels, one shared cause — there is no absolute notion of "this is not a match" on either path.

  1. Semantic. retrieval.py:266-267 computes floor = max(MIN_RELEVANCE, best * RELEVANCE_FLOOR_SHARE) with MIN_RELEVANCE = 0.02 and RELEVANCE_FLOOR_SHARE = 0.25. Scores are normalized per query against the best of their own path, so the best competing candidate scores 1.0 × its class weight by construction and can never fail a threshold defined as a share of itself. MIN_RELEVANCE binds only when the best score is below 0.08, which no real embedding produces — it is effectively dead code. Because the semantic path scores every embedded chunk, there is always a best, so with semantic retrieval enabled at least one passage is injected on every turn, unconditionally.
  2. Lexical. fts.py joins query terms with OR and applies no minimum-match-quality rule, and the stop list is 42 words. Common words carry full matches: hidden-key.md scored lex=1.000 against an orbital-mechanics scene on the word "before", and against a harbour scene on "Aldric" — the protagonist's name, which is in the story tail of essentially every query and in any source that mentions them. The lexical-only path leaked in the negative control too, so this is not a semantic-only problem.

Calibration data for the corrective pass. Against nomic-embed-text, over 7 scenes × 5 sources:

relevant scenes     n=15  min=0.407  max=0.698  mean=0.546
irrelevant scenes   n=20  min=0.392  max=0.551  mean=0.462
overlap: YES (irrelevant max 0.551 vs relevant min 0.407)

Raw cosine is not cleanly separable for this model, so a single absolute cosine threshold will trade recall for precision and must be chosen and documented deliberately — and it is model-dependent, exactly as memorybank.REDUNDANT_SIMILARITY records for its own threshold. A defensible shape, for the implementer to choose among rather than a prescription:

  • require a real lexical match on a term that is not ubiquitous (exclude the protagonist/entity names that appear in every query, and widen the stop list), or an absolute cosine above a documented, model-calibrated value; and
  • keep the relative floor as a second filter, not the only one.

User impact. Every turn spends knowledge budget on material about nothing in the scene. Worse, narrator-only hidden Canon is supplied on turns that have nothing to do with it, which maximises the exposure of the one feature whose entire purpose is controlled disclosure (§G). And the narrator is asked to reconcile authoritative-sounding Canon that is irrelevant to what the reader just did.

Blocks M7 acceptance: yes.


M7-F2 — The retrieval-quality suite cannot detect M7-F1 — BLOCKING (test)

Severity: high Violated requirement: TECHNICAL-DESIGN.md §18.1 (the M2 wiring rule: test a real consumer path); the standing milestone rule that at least one real provider path must be exercised. Owner: M7 corrective.

tests/test_knowledge_retrieval_quality.py::test_irrelevant_canon_does_not_win_on_class_alone passes, and asserts exactly the property M7-F1 violates. It passes because its SceneEmbedder stub is far more discriminative than the real model. Measured on identical texts:

abbey-canon   (relevant)           0.7559        1.000   |    0.7505       1.000
crypts-ref    (relevant)           0.3392        0.449   |    0.6237       0.831
ship-canon    (IRRELEVANT)         0.1499        0.198   |    0.4365       0.582
cookery       (IRRELEVANT)         0.0818        0.108   |    0.4354       0.580
desert        (IRRELEVANT)         0.0642        0.085   |    0.4349       0.580

relevance floor = 0.250
  -> stub: irrelevant peaks at 0.198  => correctly excluded => test passes
  -> real: irrelevant peaks at 0.582  => ADMITTED           => defect

The stub's own docstring states the choice that causes this: "There is deliberately no constant component. One would put a similarity floor under every pair." A similarity floor under every pair is precisely what a real embedding model has. The stub removed the property that makes the defect detectable.

This is the M2 and M6 lesson recurring in a third form: a mocked provider left a real defect invisible behind a green suite. The corrective pass needs a retrieval test whose embedder reproduces the real model's similarity floor, or one that runs against the real model.

Blocks M7 acceptance: yes — the fix for M7-F1 cannot be validated by the current suite.


M7-F3 — MIN_RELEVANCE is unreachable dead configuration — non-blocking

Severity: low Owner: M7 corrective (fold into M7-F1's fix).

classes.MIN_RELEVANCE = 0.02 is presented as the absolute half of a two-floor design, but with per-query normalization it can only bind when the best score is below 0.08. It never fires. The comment beside it explains a design that the code does not implement. Whatever replaces it in M7-F1's fix, the constant and its comment should stop describing protection that does not exist.


M7-F4 — The acceptance tests use the wrong fixture files — non-blocking

Severity: medium Violated requirement: V1-ACCEPTANCE-TESTS.md §4 names TEST-CAMPAIGN-FIXTURE.md as the standard campaign; BUILD-MILESTONES.md M7 requires the fixture be used "wherever practical". Owner: M7 corrective (test-only).

tests/test_imported_knowledge.py uses the three short files from V1-ACCEPTANCE-TESTS.md §5 rather than the richer TEST-CAMPAIGN-FIXTURE.md §12 versions. The consequence is concrete: §12's inspiration.md contains "A frightened innkeeper concealed a dangerous political secret from a stranger", which is the trap G07 exists to test against the canon rule "Mara is not a spy". The implementation's G07 test used the silent-hall passage instead and so never exercised the trap. This review did, and it passes — but the repository's own suite does not cover it.

Also missing from the suite: the fixture's campaign canon (§11), its narrator rules (§4) and its hidden Canon (§6, §8).


M7-F5 — Formatting control characters survive into displayed filenames — non-blocking

Severity: low Owner: M7 corrective or M8 (display).

safe_filename strips separators, NULs and leading dots, but leaves Unicode bidirectional-override characters. "‮exe.dm.md" is stored and displayed verbatim, which can render as a different extension in the list. There is no security impact beyond display spoofing — the name is never used as a path, and H08's stronger guarantees hold — but a name shown to a reader should not be able to lie about itself. Stripping the Cf (format) category would close it.


M7-F6 — Narrator state blocks can leak into displayed prose — pre-existing, not M7

Severity: low Owner: later milestone (M8/M11), recorded for completeness.

Twice during real-narrator testing, qwen2.5:3b-instruct emitted its state block fenced as ```json rather than ```state, so narrative.extract.split did not recognise it and the raw JSON was displayed to the reader. This is an M5 protocol-robustness issue with a small model, entirely outside M7's scope and unaffected by it. Noted because it was observed under this review's real-narrator runs and a reader would see it.


J. I05 and export/import

PASS. A genuine round trip: source server on one database file, destination server as a separate process on a different, empty database file.

Before export: all three classes, one disabled source, one hidden Canon source that is also always-include, one six-chunk source, and a real narrator turn that used imported knowledge.

E1  bundle carries a knowledge section (5 sources)          PASS
E2  and no derived data (no chunks/vectors/index rows)      PASS
E4  imports into a clean data directory                     PASS
E5  every source arrived                                    PASS
E6  classification survived            (all 5)              PASS
E7  enabled/disabled state survived    (all 5)              PASS
E8  visibility survived                (all 5)              PASS
E9  content hash survived              (all 5)              PASS
E10 title and original filename survived (all 5)            PASS
E11 content byte-identical             (all 5)              PASS
E12 chunks rebuilt, same count         (all 5)              PASS
E13 always-include survived                                 PASS
E14 rebuilt chunks identical in text, order and hash        PASS
E16 lexical retrieval works immediately, with no reindex    PASS
E16b the disabled source stays out after import             PASS
E17 semantic vectors rebuild locally on demand              PASS
E17b semantic retrieval works in the imported campaign      PASS
E18 exporting did not disturb the original turn's evidence  PASS
E20 the story itself round-tripped                          PASS
E21 no external source path is needed                       PASS

I05's literal condition — metadata and classification survive — passes. M7's broader scope condition — a usable local knowledge subsystem survives — also passes: chunk identity is reproduced exactly by the deterministic chunker, lexical retrieval works on arrival with no reindex step, and vectors rebuild against whatever model the importing machine has.

Recorded limit, correctly documented and not an M7 defect. The bundle carries no context snapshots, so the imported campaign has no historical prompt provenance — for imported knowledge or for any other component. That is exactly what a pre-M7 bundle already did for summaries, memories and everything else, it is stated in DATA-MODEL.md §29 and BUILD-MILESTONES.md M7 § Status, and it belongs to M9. Nothing M7 creates becomes a dangling reference, because no ids are exported.


K. F05 / F06 completion

Both complete. M6 left them partial only because imported knowledge did not exist; it now does, and the gap is filled in the browser, not only the API.

For a completed narrator turn, the Insights panel showed — rendered, in Firefox:

▸ hostile.md  CANON · Hostile · passage 1 · hybrid match
  (lexical 0.64, semantic 0.90 → 1.00) · 124 tok
  <the full rendered chunk text>
  2 passage(s) considered, 182 of 2112 knowledge tokens used
  · searched on: aldric, studies, broken, circle, abbey, crypt, westhaven

F06's requirements, each checked in the rendered DOM: source filename/title ✓, classification ✓, chunk identity ✓, rendered chunk text ✓, retrieval score and mode ✓, token cost ✓, and why it was searched ✓.

F05's prompt-section view showed every component with a readable name and a token cost:

narrator prompt 57 tok · narrative state (reporting rule) 467 tok ·
campaign canon 109 tok · imported knowledge (rules for using it) 199 tok ·
ai instructions 140 tok · plot essentials 17 tok · story history 130 tok ·
imported canon (retrieved) 232 tok · narrative state (current) 62 tok ·
length guidance 51 tok · narrative state (emit reminder) 32 tok

Historical evidence survives deletion of its source: a turn was generated using canon.md, the source was deleted, and the old turn's snapshot still carried the rendered text byte-identically (§L, L15).


L. Source lifecycle and atomicity

31 of 32 checks passed; the one failure was a defect in the review's own test.

L1  .txt import indexes                                     PASS
L2  .md import indexes                                      PASS
L3  unsupported extension refused                           PASS
L4  binary content refused despite .txt extension           PASS
L5  invalid UTF-8 refused, not mangled                      PASS
L5b empty source refused                                    PASS
L6  1,040,382 bytes accepted (limit 1,048,576)              PASS
L7  1,059,372 bytes refused, message says what to do        PASS
L7b nothing of the refused source stored                    PASS
L8  NFC/NFD normalize to the same hash                      PASS
L8b the conflict names the existing source                  PASS
L8c a CRLF copy is detected as a duplicate                  PASS
L9  explicit duplicate override honoured                    PASS
L10 disable keeps content, chunks and FTS rows              PASS
L11 re-enable restores retrieval, no reimport               PASS
L12b no orphan FTS rows anywhere                            PASS
L12c no orphan chunks anywhere                              PASS
L12d every stored source is 'ready' — none half-built       PASS
L14 delete removes source, chunks, FTS rows and vectors     PASS
L14d no dangling FTS rowids after delete                    PASS
L15 the historical turn still says what it was supplied     PASS
L15b byte-identical to what was supplied                    PASS
L15c the story was not rewritten by the delete              PASS

L12 was a review test error, not a product defect. It expected a .md file whose content is %PDF x to be refused. That content is valid readable text, so accepting it is correct — SECURITY-THREAT-MODEL.md §21 asks for rejection of binary data, which L4 demonstrates with a real ZIP header. The invariants that matter (no orphans, nothing half-built) all pass.

Atomicity is additionally structural: retrieval is gated on index_state = 'ready', so even a hypothetical partial commit would be inert rather than wrong.


M. Ranking and authority

15 of 16 checks passed; the failure is M7-F1.

R3.1  irrelevant Canon is not retrieved                     FAIL  -> M7-F1
R3.1b relevant Reference IS retrieved (positive control)    PASS
R3.1c relevant Reference outranks irrelevant Canon          PASS  (0.897 > 0.584)
R3.2  at comparable relevance, Canon scores above Reference PASS  (1.133 > 0.897)
R3.2b Reference framed as non-canon wherever it lands       PASS
R3.2c Inspiration framed as establishing nothing            PASS
R3.3a the lexical path contributed a winner                 PASS
R3.3b the semantic path contributed a winner lexical missed PASS
R3.4  no chunk appears twice in the selection               PASS
R3.4b a chunk both paths found is marked hybrid, not doubled PASS
R3.5  a disabled source is excluded before ranking          PASS  (considered 5->4)
R3.6  Campaign A's sources never appear in Campaign B       PASS  (considered=0)
R3.6b B retrieves its OWN source (positive control)         PASS
R3.7  reclassifying moves the passage into the Canon section PASS
R3.7b and it is no longer framed as Reference               PASS

Ordering, deduplication, class weighting, disabled-source exclusion and campaign isolation are all correct. Only the admission threshold is wrong. That is a narrow, well-localised fix.


N. Local-only and network evidence

PASS. No new network path. Socket capture in the server process:

import + chunk + FTS5 index          destinations: none
lexical retrieval + prompt assembly  destinations: none
a full narrator turn                 destinations: none
semantic indexing + retrieval        destinations: the configured Ollama
                                     host and port, and nothing else
                                     (the configured Ollama, its configured port)
  • A public endpoint written into the database behind the settings API is refused at request time on the knowledge path, and the refusal opened no socket.
  • The policy refuses api.openai.com and public addresses, and permits loopback and the configured trusted-LAN host.
  • An import still succeeds lexically while the endpoint is refused — the library does not depend on the inference host being reachable.
  • The knowledge subsystem builds no HTTP client of its own, embeds only through memorybank.embedding_provider, and reimplements no TLS handling, so the allowlist, the request-time re-check and the OS/private-CA trust union all apply unchanged (ADR 011, ADR 002).
  • No remote vector service, cloud host, telemetry, CDN asset or model-download path exists in it. (Two apparent hits in an initial scan were false positives: "cohere" inside "coherent", "pulled" in prose.)

Consistent with ADR 004 and SECURITY-THREAT-MODEL.md §§12-21.


O. Migration

PASS — 23/23, through a real server restart.

A campaign was built with history, a divergence, a Save Point, narrative state, a summary, a memory and a derived-status row. The M7 tables and the FTS index were then dropped and the stamp rewound to 91, making the file genuinely M6 — verified before migrating. The server process was stopped and a new process opened the same file.

stamp 91 -> 92 on restart
every M7 table created, including knowledge_fts
story history                 identical
branches                      identical
Save Points                   identical
narrative state               identical
memories                      identical
summaries + derived status    identical
active head                   unmoved (2, 6)
pre-M7 prompt provenance      readable and unchanged
old snapshots                 no knowledge section forced onto them
empty library, no imported section in the next prompt
normal turn / Undo / Redo / Save Point restore / inspection   all work
then: import a source and retrieve it in a real narrator turn PASS

An old campaign needs no knowledge sources, and gains the subsystem cleanly.


P. Browser evidence

Firefox 154.0.1, headless, WebDriver, against the built frontend served by the real backend. 22/22 passed, covering all three surfaces that render imported text plus a full page reload. See §H and §K for the specific results.

No severe console errors. No request to any non-origin host across the entire session.

The implementation's own browser suite (43 checks) was also re-run during the implementation pass; this review's suite is independent of it.


Q. Performance and scaling

PASS — measured, flat.

sources/chunks   list  detail  retrieve  context   considered  used   ctx
  3 /  12          5q      4q        9q      20q          12     3   130ms
 15 /  60          5q      4q        9q      20q          55     3   108ms
 39 / 156          5q      4q        9q      20q          67     3   188ms

Query counts did not move while the library grew 13×. No N+1 on the source list, the chunk listing, retrieval or prompt construction. Candidates are bounded in SQL (67 of 156 chunks, cap 80); the FTS query carries a LIMIT; the semantic catalogue read fetches ids only, not vectors or passage text; the vector read is a single batched query.

Characteristic to record, not a defect: the semantic scan is a linear cosine over the campaign's vectors in Python, capped at 4,000 passages with the shortfall reported. Query count is flat; per-turn CPU grows linearly with library size. BUILD-MILESTONES.md M7 § Status already records this as deliberate v1 debt (no approximate-nearest-neighbour index).


R. Regression gates

Gate Result
backend suite 920 passed, 10 skipped, 1 warning (M6 baseline 836/7/1)
focused M3-M6 regressions 372 passed, 1 warning — context/memory, memory retrieval, head cursor, Save Points, narrative state, bundle v2, tree migration, endpoint policy, local-only surface, egress, state revert, pre-M5 compatibility
frontend lint exit 0, 7 warnings — identical to the M6 baseline
frontend build success
docker build success; image verified running with python-multipart 0.0.32, schema 92, porter-stemmed FTS

The single warning is the pre-existing StarletteDeprecationWarning from fastapi.testclient. The 7 lint warnings are the pre-existing components.jsx/SchemaEditor.jsx fast-refresh notes. The 3 additional skips over the M6 baseline are test_knowledge_real_model.py, which skips without a configured endpoint; they were run explicitly in this review with one.


S. New dependency

Approved.

package        python-multipart 0.0.32
licence        Apache-2.0
requires       nothing (Requires-Dist: None)
files          24
declared       backend/requirements.txt  python-multipart>=0.0.9  (with reason)
pinned         backend/requirements.lock python-multipart==0.0.32
docker         installs from requirements.txt — verified in the built image
clean venv     installs from requirements.lock — verified, app imports

It is Starlette's own multipart parser, needed because the upload surface is UploadFile. It adds no transitive dependency, no framework surface and no network path — and the upload surface it enables is precisely why no endpoint accepts a filesystem path (H08). Reason, pinning and reproducibility all check out.


T. Deferred features — confirmed acceptable

None of these fails M7; each remains consistent with the active design and is recorded in BUILD-MILESTONES.md M7 § Status and IMPORTED-KNOWLEDGE-DESIGN.md §76.

  • Entity linking (§33), manual priority (§35), scene pinning (§36) — optional in the design; §36's stated v1 minimum (always-include for critical Canon) is implemented.
  • Source version history (§14, §43) — the design explicitly permits "a simpler replace-with-provenance model … if documented", and it is documented.
  • Canon scope metadata (§45) — §45 itself says explicit current-state precedence may be sufficient for v1, which is what was built.
  • Tags (§34) — evaluated in context as instructed. They are unused optional metadata: nothing in the API, the source inspector, the prompt or the docs promises tags, so their absence breaks no implemented promise. Acceptable.
  • Canon-vs-Canon conflict detection (§42) — absent, and outside M7's scope list. It produced no authority failure in any required M7 scenario in this review. Recorded as later-v1 debt.

No planning document was weakened to match the implementation.


U. Planning-document recommendations

  1. IMPORTED-KNOWLEDGE-DESIGN.md §29-30 — add, from M7-F1, that a relevance threshold expressed only as a share of the best candidate cannot exclude anything, because the best candidate always clears a share of itself. A retrieval design needs an absolute notion of "no match" on each path, and for embeddings that value is model-dependent and must be measured and written down, as memorybank.REDUNDANT_SIMILARITY already is.
  2. TECHNICAL-DESIGN.md §18.1 — extend the M2 wiring rule with M7-F2: a stub provider must not be more discriminative than the real one. A stub with no similarity floor cannot detect a retrieval defect that only exists because real embeddings have one. State the requirement that retrieval thresholds be validated against a real model.
  3. V1-ACCEPTANCE-TESTS.md G05-G07 — add the negative control as a pass condition: a query unrelated to every source must retrieve nothing. Without it, G05/G06/G07 can all pass while retrieval is unconditional.
  4. V1-ACCEPTANCE-TESTS.md §5 / TEST-CAMPAIGN-FIXTURE.md §12 — the two documents give different contents for the same three filenames. §12 is the richer and is what the traps are built on. Reconcile them, or state which is normative for the G-series.
  5. CONTEXT-AND-MEMORY.md §46 — record from §G that hidden Canon's protection is prompt discipline plus retrieval scoping; if retrieval is unconditional, prompt discipline is carrying the whole load on every turn.

V. M8 implications

  • The knowledge panel and the Insights knowledge rows are functional and complete for M7's purposes, and every M7 behaviour is reachable in a browser with no database access. M8 owns their design.
  • M8 should not have to touch retrieval; M7-F1's fix belongs to the M7 corrective pass, before M8 begins.
  • Two M5 section labels (state_rule, state_reminder) that rendered as raw keys were fixed during the M7 pass; the rest of the panel's visual design is M8's.
  • §47 of BROWSER-UX-SPEC.md describes an import preview step that is not built, and §52's retrieval-usage count is absent. Both are marked optional; M8 may reconsider.
  • If M7-F1's fix means retrieval sometimes returns nothing, the panel and the inspector should say so plainly rather than showing an empty section — a small M8 affordance worth planning for.

W. Verdict

PASS WITH CORRECTIVE WORK REQUIRED

M7 delivers the subsystem, and the parts that work are demonstrated rather than asserted. Against a real narrator, a real embedding model, a real browser, real socket capture and a real migration, the following all hold: the full source lifecycle; classification that changes authority rather than metadata; campaign isolation; hybrid retrieval with correct ranking, deduplication and disabled-source exclusion; a bounded knowledge budget; hidden Canon that reaches the narrator without reaching the protagonist; prompt injection framed as data and ignored by the model; no network path but the approved one; no filesystem path at all; export/import into a clean directory; a clean migration from a genuine M6 database; and flat query counts as the library grows.

The five acceptance conditions the implementation flagged as unmeasured — C05, G06, G07, G10 and hidden Canon — were measured here against a real narrator and all five pass. That was the largest open risk and it is closed.

Two blocking items remain, and they are the same defect and the reason it was not caught:

  • M7-F1 — retrieval admits imported knowledge unconditionally, because the only effective threshold is a share of the best candidate. Irrelevant Canon, Reference, Inspiration and hidden Canon are injected into every turn.
  • M7-F2 — the retrieval-quality suite's stub embedder is more discriminative than the real model, so it cannot detect M7-F1 and cannot validate its fix.

Both are narrow and localised: the ranking is correct, the pipeline is correct, and the fix is to the admission threshold and to the fixture that tests it.

M7 is not accepted. M8 is not authorized. Recommended next step is an M7 corrective pass addressing M7-F1 and M7-F2, with M7-F3, M7-F4 and M7-F5 folded in, followed by re-review of §I, §E (G05-G07) and §G.


X. Repository state at the end of the review

branch                       m7-imported-knowledge
HEAD                         a6e9c7a32bdf42f1e4cb837b70721d89422cef8d   (unchanged)
staged                       48 files  (47 from M7 + this report)
unstaged                     0
untracked                    0
LICENSE                      unchanged, not staged
prepared commit message      .git/M7_MSG, unchanged

Files modified by the review itself: exactly one — planning/reports/M7-IMPLEMENTATION-REPORT.md, which is this report, newly created. No application code, no test, no planning document and no configuration was changed by the review. All review harnesses, fixtures, databases and browser artifacts were created outside the repository, under the session scratchpad, and none of them is staged.

M6's report remains at planning/reports/M6-IMPLEMENTATION-REPORT.md; per the package convention it rotates to planning/archive/milestone-reports/ when M7 is accepted, which has not happened.

Closeout note, added later: it has now. M6's report is at planning/archive/milestone-reports/M6-IMPLEMENTATION-REPORT.md. This paragraph is left as written because it recorded the state at review time.*



M7 Corrective Pass — Closeout

Date: 2026-09-06 Scope: the two blocking findings above, M7-F1 and M7-F2, plus the non-blocking findings where they were directly related and low-risk. Nothing above this line has been changed. The original findings, their reproductions and their evidence stand as written.


CA. What was wrong, reproduced on the staged tree

Both findings were reproduced before any code changed, against the real embedding model, so the before/after comparison is like for like.

CA.1 — M7-F1, the no-match case

scene: "Aldric studies the tide tables and the harbour master's ledger
        of container tonnage."   (nothing in the library relates to it)

BEFORE   considered=5  floor=0.2875  semantic=True
  USED hidden-key.md   canon        lex=1.000 sem=1.000 cos=0.548 score=1.150  51 tok
  USED canon.md        canon        lex=0.000 sem=0.772 cos=0.423 score=0.772  56 tok
  USED reference.md    reference    lex=0.895 sem=0.927 cos=0.508 score=0.902  70 tok
  USED revival.md      reference    lex=0.000 sem=0.730 cos=0.400 score=0.621  86 tok
  USED inspiration.md  inspiration  lex=0.000 sem=0.871 cos=0.477 score=0.610  63 tok
  -> 5 of 5 sources injected, 326 tokens
  -> imported sections in the prompt:
     ['imported_inspiration', 'imported_reference', 'imported_canon']

Note the raw cosines: 0.400-0.548, i.e. nothing matched. Normalization lifted the best to 1.000 and the relative floor (0.2875) admitted everything.

CA.2 — M7-F1, irrelevant admitted alongside relevant

scene: the abbey crypt          library: abbey Canon + container-shipping Canon

BEFORE  USED canon.md      canon  cos=0.697 score=1.150   <- relevant
        USED shipping.md   canon  cos=0.320 score=0.459   <- container vessels
        USED reference.md  ref    cos=0.486 score=0.593

CA.3 — M7-F1, lexical leak on a generic word

scene: "Ballast trim on the ore hauler was recalculated before the burn to Ceres."
library: hidden-key.md ("...Edrin discovered this before he disappeared...")
semantic: OFF

BEFORE  terms=[... 'before' ...]
        USED hidden-key.md  canon  lex=1.000  -> 1 injected, 51 tokens

The single shared token was the function word "before", which the 42-word stop list did not contain and which nothing downstream weighed.

CA.4 — M7-F1, lexical leak on the protagonist's name

scene: "Aldric studies the harbour ledger."      semantic: OFF
BEFORE  terms=['aldric','studies','harbour','ledger']
        USED hidden-key.md  canon  lex=1.000  -> 1 injected, 51 tokens

hidden-key.md contains "Aldric does not know what the key opens". The protagonist's name is in the story tail of essentially every query, so this admits on a token that carries no information about the current scene.

CA.5 — M7-F2, why the suite could not see any of it

Measured on identical texts, stub against the production model:

                          stub cosine  normalized | real cosine  normalized
abbey-canon  (relevant)       0.7559      1.000   |    0.7505      1.000
crypts-ref   (relevant)       0.3392      0.449   |    0.6237      0.831
ship-canon   (IRRELEVANT)     0.1499      0.198   |    0.4365      0.582
cookery      (IRRELEVANT)     0.0818      0.108   |    0.4354      0.580
desert       (IRRELEVANT)     0.0642      0.085   |    0.4349      0.580

The old floor was 0.25 of the best. Irrelevant material peaked at 0.198 under the stub and at 0.582 under the real model — excluded in the fixture, admitted in production.


CB. Root cause

Two symptoms, one architectural cause: M7 had a ranking stage and no admission stage.

Both scores were normalized against the best candidate of their own path, and the floor was max(0.02, best × 0.25). A floor expressed as a share of the best cannot reject anything, because the best clears a share of itself by construction. MIN_RELEVANCE = 0.02 could only bind when the best score fell below 0.08, which no real embedding produces, so it was unreachable — finding M7-F3 was not a separate defect but the same one seen from the side.

Because the semantic path scores every embedded chunk, there was always a best. With semantic retrieval enabled, at least one passage was admitted on every turn, unconditionally.

The lexical path had the mirror-image problem for a different reason: FTS5 terms were joined with OR and nothing asked how much had matched, so one incidental token was sufficient.


CC. The correction

CC.1 Admission separated from ranking

candidate generation
  -> ADMISSION   absolute, per path, independent of the candidate set
  -> RANKING     normalized, among the survivors only
  -> class weighting
  -> budget

Admission reads raw signals; ranking reads normalized ones. Normalization still exists — bm25 has no fixed range and cosine's zero is not zero, so the paths are not otherwise comparable — but it now decides order among things that matched, never whether anything matched. Authority is applied in the ranking stage only, so a class can order what matched and can never rescue what did not.

Recorded as a durable architectural lesson in TECHNICAL-DESIGN.md §13.2 and summarised in IMPORTED-KNOWLEDGE-DESIGN.md §76.

CC.2 Semantic admission — classes.SEMANTIC_FLOOR = 0.58

Established empirically through the production path: passages embedded as fts.index_line(heading, text), queries assembled by retrieval.query_terms, because both differ from bare strings and both move the numbers. The first calibration attempt used bare strings and under-measured by ~0.10, which is recorded here because it is exactly the kind of error that produces a plausible wrong constant.

113 pairs against nomic-embed-text:

targeted   n= 13  min 0.5526  p10 0.6090  median 0.7231  max 0.8474
             the one source a scene is actually about
off-topic  n=100  min 0.3577  median 0.4591  p95 0.5339  max 0.5578
             20 scenes unrelated to the campaign (harbour, surgery, compiler,
             fugue, sourdough, kiln, football, photosynthesis, …)

sweep:  0.55 -> 13/13 targeted, 1/100 off-topic
        0.56 -> 12/13 targeted, 0/100 off-topic
        0.58 -> 12/13 targeted, 0/100 off-topic   <- chosen
        0.61 -> 11/13 targeted, 0/100 off-topic   (loses C05's revival match)

0.58 sits between the two populations with ~0.022 of margin on each side. The single targeted pair below it is instructive rather than lost: "the broken circle cut into the keystone above the crypt stair" scores 0.5526 against the Canon describing exactly that — the wording is so close that little is left for the embedding to add — and it matches four lexical terms, so the lexical path admits it. That is the hybrid working, and it is why neither path has to be right alone.

The value is a property of the model and is documented as such beside the constant, in the same style as memorybank.REDUNDANT_SIMILARITY. A model that scored everything below it would degrade the library to lexical-only, which is a supported production path, so the failure mode is safe.

CC.3 Lexical admission

A passage is lexically admitted when it matches two distinct meaningful query terms, or one term that is both (a) not the name of a standing campaign entity and (b) not a negligible share of the query.

  • The stop list grew from 42 to 261 words, all function words and contentless generics. It contains no subject matter — no "abbey", "key" or "crypt" — because a stop list that removes subject matter is how a search stops finding "The Silver Key". before is now filtered.
  • Standing entities are the protagonist and the campaign's established entities, drawn from persona_name and the authoritative state. They are in the query on every turn by construction, so a lone match on one says nothing about the present scene. This is deliberately not "ignore proper nouns" — Westhaven, broken-circle and resurrect all still count, and a standing entity counts the moment a second term matches beside it.
  • The share test covers the case the entity test cannot see: a young campaign whose state is still empty. One word out of a nine-word scene is not evidence however distinctive it is; one out of three is.

Both conditions are needed and each was measured failing alone: the share test alone would admit a lone "Aldric" from a three-word query, and the entity test alone admitted it from a nine-word one — which is what CA.4 measured.

Evidence per term comes from fts.term_evidence, one statement with a subquery per term (capped at 24), scoped once at the join. Stemming is applied by FTS itself, so resurrected matches resurrection exactly as the ranking query does rather than through a second, divergent stemmer in Python.

CC.4 M7-F3 folded in

MIN_RELEVANCE and RELEVANCE_FLOOR_SHARE are removed, not re-tuned. The architecture that made them meaningless is gone, and nothing is left that appears to enforce relevance without doing so.


CD. Real-model before/after

Same reproductions, same model, same fixtures.

Case Before After
Off-topic scene (tide tables/tonnage), semantic ON 5 sources, 326 tokens, 3 imported sections 0 sources, 0 tokens, no imported section
Relevant scene + irrelevant Canon in library relevant + shipping Canon + Reference relevant Canon only
Lexical leak on "before" 1 source, 51 tokens 0
Lexical leak on "Aldric" 1 source, 51 tokens 0
Positive control: Old Abbey question canon.md + shipping.md canon.md only

The after-state reports the rejection explicitly rather than silently: generated=5 rejected=5 considered=0 floor=0.58.


CE. The §12 acceptance matrix — measured against the real model

13/13.

Query Library Expected Result
Old Abbey / broken-circle relevant Canon + unrelated relevant Canon only PASS — ['canon.md']
tavern description tavern Reference + unrelated Canon Reference, not irrelevant Canon PASS — Reference in, shipping Canon out
resurrection question conflicting Canon and Reference Canon governs PASS — campaign rule present above the framed Reference
tide tables / container tonnage fantasy fixture only zero imported chunks PASS — used=[] sections=[] generated=4 rejected=4
distinctive exact term embeddings unavailable lexical succeeds PASS — ['sigils.md'], semantic_used=False
semantic paraphrase little lexical overlap semantic succeeds PASS — admitted_by semantic, cosine 0.6575
protagonist name only many Aldric-containing sources no unrelated flood PASS — used=[]
stopword-heavy input mixed library no incidental FTS admission PASS — used=[], terms=[]

Plus three beyond the matrix:

  • hidden Canon when relevant → retrieved.
  • hidden Canon when unrelated → absent (used=[] generated=1).
  • always-include Canon → still supplied on an unrelated scene, because it is asserted rather than retrieved and is deliberately outside admission.

For every no-match row the assembled prompt was inspected: the imported-knowledge sections are absent, not empty, and the untrusted-data rule section is absent with them.


CF. The four hybrid cases

Permanent deterministic coverage in test_knowledge_retrieval_quality.py:

A  strong semantic / weak lexical   ossuary passage, no shared vocabulary
                                    -> retrieved, admitted_by="semantic"
B  strong lexical / weak semantic   distinctive term, embeddings unavailable
                                    -> retrieved, admitted_by="lexical"
C  both strong                      -> one candidate, admitted_by="both",
                                       not duplicated
D  neither strong                   -> ZERO chunks, and no imported section
                                       in the prompt  (mandatory case)

Case D is covered twice — with embeddings on and on the lexical-only path — and each assertion is guarded by generated > 0, so it cannot pass because nothing was searched.


CG. M7-F2 — the test doubles

RealisticEmbedder replaces the old stub. It has a deliberate similarity floor: every pair shares a constant component, so unrelated passages score a substantial similarity as they do in life, with disjoint topic axes above it and a hashed bag of words underneath.

Two tests exist purely to keep the fixture honest:

  • test_the_stub_models_the_real_problem fails if unrelated pairs ever drop near zero again — that is, if the fixture drifts back to the one that hid the defect.
  • test_a_relative_only_floor_would_admit_the_irrelevant_set reproduces the superseded rule on the corrected code's own vectors and asserts it still admits the entire irrelevant library. If this ever stops failing the old rule, the fixture has lost the ability to detect the defect.

The stub is not calibrated to 0.58. It models the shape — a floor under everything, structure above it — and the tests assert against classes.SEMANTIC_FLOOR rather than against a number baked into the fixture.

Validation that F2 is actually closed: the corrected implementation was temporarily reverted to the staged pre-corrective version and the new suite run against it:

13 failed, 5 passed

including every no-match test, both case-D tests, all three "irrelevant material is excluded whatever its class" parametrisations, and "Canon is excluded even though it is the best candidate". The suite can now see the defect it previously could not.


CH. Permanent real-model regression

test_knowledge_real_model.py gained four tests, skipped without an endpoint:

test_the_real_model_separates_relevant_from_unrelated
      re-measures both populations and fails if SEMANTIC_FLOOR stops sitting
      between them.  Measured this run:
        targeted   0.8130 / 0.7260 / 0.6726
        off-topic  0.3293 - 0.5026  (10 pairs, 5 scenes x 2 sources)
        floor      0.58
test_a_completely_unrelated_query_retrieves_nothing_from_a_real_model
        generated=3 rejected=3 used=0, no imported section
test_a_relevant_query_still_retrieves_from_a_real_model      -> ['canon.md']
test_a_paraphrase_still_retrieves_from_a_real_model          -> cosine 0.7260

The first of these caught a real error in its own first draft: a single off-topic line appended after a crypt opening still leaves the crypt in the four-action query window, so the query was not off-topic at all. The window is now moved wholesale. Recorded because it is a property of the design worth knowing: relevance is judged against the recent scene including its immediate history, not against the last sentence alone.


CI. No regression in established M7 behaviour

Area Result
C05, real narrator PASS — "the dead do not return … magic cannot bring him back to life", with revival.md still retrieved at cosine 0.656, so the test remains non-vacuous
G06 / G07 / G10, real narrator PASS (25/25 with C05 and hidden Canon)
Hidden Canon, real narrator PASS — retrieved, marked narrator-only, not revealed
G01-G10 acceptance suite 49 passed
H06 / H07 / H08 / H09 PASS — 43/43 browser checks, unchanged
I05 export/import 53/54 (the one failure is the review harness's own false positive: text/markdown contains a slash)
F05 / F06 PASS — browser-verified, now also showing matched terms and similarity
Source lifecycle, atomicity, isolation, disable/delete PASS
Local-only network boundary PASS — zero sockets for import/index/retrieval/turn; embeddings only to the configured endpoint; request-time policy enforced
Query counts (3 → 39 sources) list 5, detail 4, retrieve 10, context 21 — flat; retrieval and context each gained exactly one statement (the term-evidence query)

CJ. Regression gates

Gate Result
backend suite 927 passed, 14 skipped, 1 warning (pre-corrective 920/10/1)
M7 knowledge suites 49 + 18 + 14 + 6 + 4 passed
real-model M7 suite 7 passed against nomic-embed-text on the trusted-LAN Ollama
frontend lint exit 0, 7 warnings — identical to the M6 baseline
frontend build success
docker build success
browser 43/43, Firefox 154.0.1
network unchanged; see CI
migration, pre-M7 database via a real server restart 23/23
export/import into a clean data directory 53/54 (one review-harness false positive)
query counts, 3 → 39 sources list 5, detail 4, retrieve 10, context 21 — flat
§12 acceptance matrix, real model 13/13
real-narrator authority (C05, G06, G07, G10, hidden) 25/25

The +7 tests are the new admission and real-model tests; the +4 skips are the new real-model tests, which skip without an endpoint. The single warning is the pre-existing StarletteDeprecationWarning.

No schema change and no serialized-settings change: the corrective pass altered the shape of the context-snapshot knowledge record (floor → generated, rejected, semantic_floor, plus admitted_by and matched_terms per passage) but added no column and no migration. Snapshots written before the corrective pass still render — the browser reads the new fields with numeric comparisons that are false when absent.


CK. Disposition of the non-blocking findings

Finding Disposition
M7-F3 — MIN_RELEVANCE unreachable dead config Fixed. Removed along with RELEVANCE_FLOOR_SHARE; the architecture that made them meaningless is gone. Nothing now appears to enforce relevance without doing so.
M7-F4 — acceptance tests used the wrong fixture Fixed. test_imported_knowledge.py now uses TEST-CAMPAIGN-FIXTURE.md §12 verbatim, plus the §11 campaign canon and §8 hidden Canon. G07 exercises the actual trap — the frightened-innkeeper passage against the "Mara is not a spy" rule — and asserts both the framing and that no spy fact reaches authoritative state. The two suites now state which they are: test_imported_knowledge.py is the standard acceptance fixture, test_knowledge_retrieval_quality.py is purpose-built retrieval mechanism.
M7-F5 — RTL override in filenames Fixed. safe_filename drops Unicode format characters (category Cf), so "‮exe.dm.md" stores as exe.dm.md and a name cannot visually impersonate another extension. Cheap, in-scope, no UI redesign.
M7-F6 — narrator emits ```json instead of ```state Not absorbed. Pre-existing M5 protocol robustness with a small model; M7 neither caused nor worsened it. Ownership stays with M5/M8/M11. Observed again during this pass and left recorded.

CL. Planning documents updated

Only where the corrective work established a durable fact:

  • TECHNICAL-DESIGN.md §13.1 — the pipeline diagram now shows the admission stage; §13.2 (new) records the architectural lesson and the general rule that a relevance decision must rest on a signal meaningful on its own.
  • IMPORTED-KNOWLEDGE-DESIGN.md §76 — the retrieval subsection records admission before authority, and states plainly that retrieval may return nothing.
  • BUILD-MILESTONES.md M7 § Status — both findings, their corrections, and the model-specific calibration recorded as carried debt.

No product requirement was weakened, and no ADR was added: the change is an implementation architecture that the two design documents already govern.


CM. Verdict

READY FOR M7 CLOSEOUT

Both blocking findings are closed with real-model and deterministic evidence:

  • M7-F1 is closed. The required invariant — retrieval must be allowed to return zero imported chunks — holds and is measured. Four off-topic scenes that previously injected the entire library now inject nothing, verified at the assembled prompt rather than at the ranking. Relevance is decided before authority, on signals that mean something on their own, and irrelevant Canon is excluded even when it is the only candidate. Every positive control still retrieves: exact terms, paraphrases, lexical-only operation, C05 against a real narrator, and hidden Canon when it is actually relevant.
  • M7-F2 is closed. The stub now models the real model's similarity floor, the suite fails 13/18 against the pre-corrective implementation, and two tests exist solely to stop the fixture drifting back.

M7 is not accepted — that is the repository owner's decision on a signed commit. M8 is not started.


CN. Repository state at the end of the corrective pass

branch                       m7-imported-knowledge
HEAD                         a6e9c7a32bdf42f1e4cb837b70721d89422cef8d   (unchanged)
staged                       see below
unstaged                     0
untracked                    0
LICENSE                      unchanged, not staged
prepared commit message      .git/M7_MSG, updated for the corrective pass


M7 Closeout Verification

Date: 2026-09-06 Scope: the two items left open at the end of the corrective pass — the embedding-model calibration boundary, and the ambiguous export/import 53/54. Nothing above this line has been changed.


DA. The export/import 53/54, resolved

Test E3: no original filesystem path is present anywhere in the bundle
Where the independent review's own harness, not the repository suite
Expected no exported field carries a filesystem path
Actual the assertion fired
Verdict NOT A DEFECT — a false positive in the review harness

The assertion was a crude substring test: it serialised every non-content field of every exported source and failed if the result contained a /. The only slash present is in "mediaType": "text/markdown" — a MIME type, not a path. No exported filename, title, hash or timestamp contains a separator.

The harness assertion has been replaced with three that test the actual property:

E3   no non-content, non-mediaType field contains "/" or "\"      PASS
E3b  every exported filename is a bare basename, no ".."          PASS
E3c  mediaType is one of text/plain, text/markdown                PASS

The export/import suite is now 56/56. The earlier 53/54 is withdrawn as an artefact of the harness; it never described product behaviour. I05 and the stronger M7 conditions are unaffected and remain demonstrated:

source contents          byte-identical                    PASS
classification           survived (all 5 sources)          PASS
enabled state            survived, disabled stays out      PASS
visibility / hidden      survived                          PASS
always-include           survived                          PASS
provenance / hash        survived                          PASS
chunk reconstruction     identical text, order and hash    PASS
lexical retrieval        works on arrival, no reindex      PASS
semantic rebuild         on demand, local                  PASS
historical provenance    the pre-export turn unaffected    PASS
no external pathname     required anywhere                 PASS

DB. The embedding-model calibration boundary, resolved

DB.1 The unsafe case

SEMANTIC_FLOOR = 0.58 is a raw-cosine threshold measured against nomic-embed-text. The corrective pass documented only the safe half of the consequence — a model scoring everything lower degrades to lexical-only — and left the dangerous half as carried debt:

a model whose similarity scale sits above nomic-embed-text's would put unrelated material past 0.58 and recreate M7-F1 exactly, on a build whose tests all pass.

That is no longer carried debt.

DB.2 The policy

Semantic admission is per model, and a model this build has not measured does not inherit a number measured against something else.

SEMANTIC_CALIBRATION = {"nomic-embed-text": 0.58}

semantic_floor_for(model) -> float | None
  • The key is the model's base name, lower-cased and without its tag: an Ollama tag (:latest, :v1.5) selects a build of the same model and does not change its similarity scale.
  • None means "not measured here". Retrieval then does not run the semantic path at all — it is skipped, not attempted and thresholded — and reports why.
  • Retrieval degrades to lexical-only, which is a first-class production path rather than a fallback, so story play, imports, indexing, inspection and turns are all unaffected.
  • No new model-identity mechanism was introduced. KnowledgeEmbedding.model already records what produced each vector and the retrieval catalogue already filters on it; the policy reuses the configured Settings.embedding_model and that same name.

The diagnostic names the model, says what is happening and lists what is calibrated:

The embedding model "mxbai-embed-large" has no measured relevance calibration
in this build, so semantic retrieval is disabled and retrieval is lexical only.
Story play and lexical search are unaffected. Calibrated models:
nomic-embed-text.

It is surfaced in three places: semantic_note on the retrieval result, the knowledge-status endpoint (semantic_calibrated, calibrated_models, semantic_note), and the Knowledge panel, which shows a warning row rather than claiming semantic search is on.

knowledge-status now reports semantic_enabled: false for a configured but uncalibrated model. That is the M6-F5 discipline applied once more: vectors that exist but are never consulted are not a working semantic index, and saying "on" because a model name is set would be the same untruth in a new place.

DB.3 What the policy costs, stated plainly

A semantic-only paraphrase — conceptually right, no shared vocabulary — is not retrieved under an uncalibrated model. That is a missing passage rather than an irrelevant one, and it is asserted in the suite rather than left to be discovered.

Semantic admission is calibrated for nomic-embed-text; uncalibrated embedding models fall back safely to lexical-only retrieval rather than borrowing its threshold.

Adding a model is a measurement, not a guess: run tests/test_knowledge_real_model.py against it and confirm the targeted and off-topic populations separate, as §CC.2 records for the existing entry. Generic cross-model calibration is deliberately not attempted in v1.

DB.4 Deterministic coverage

tests/test_knowledge_calibration.py, 12 tests, needing no second model installed — the policy is about model identity, so a configured name and a stub are the whole apparatus.

The stub is deliberately hostile: GenerousEmbedder scores everything at about 0.97, including the unrelated. It is the dangerous shape the policy exists to defend against, and it makes the negative results meaningful — under the calibrated name the fixture admits an unrelated freighter passage, and the only thing that changes in the uncalibrated run is the model name.

Required Test Result
1. nomic-embed-text uses the calibrated policy test_the_calibrated_model_resolves_to_the_measured_floor (+ tags, whitespace) PASS
2. a different model does not inherit 0.58 test_an_uncalibrated_model_does_not_borrow_the_calibrated_threshold PASS
3. unknown config degrades to an explicitly safe state test_an_uncalibrated_model_degrades_to_lexical_only_with_a_clear_reason, test_the_status_endpoint_reports_the_uncalibrated_state PASS
4. distinctive lexical retrieval still works there test_distinctive_lexical_retrieval_still_works_when_uncalibrated PASS
5. a semantic-only paraphrase is not spuriously admitted test_a_semantic_only_paraphrase_is_not_admitted_when_uncalibrated PASS
6. no-match still returns zero chunks test_no_match_still_returns_zero_chunks_when_uncalibrated / _when_calibrated PASS
7. a model change leaves no stale vectors live test_changing_the_model_does_not_leave_old_vectors_active, test_reindex_rebuilds_vectors_under_the_new_model PASS

On (7) the existing machinery already held: KnowledgeEmbedding.model records the producing model, the retrieval catalogue and the pending-work query both filter on it, so after a model change the old vectors are invisible to retrieval, pending_embeddings counts them as needing rebuild, and the retrieval result says "no passages have been embedded yet" rather than silently scoring incompatible vectors. Reindex rebuilds them under the new model. The tests pin that behaviour rather than adding a second mechanism.

DB.5 The same policy, through the live API

The deterministic tests use stubs to isolate the mechanism. It was also exercised end to end against a real server, a real database and the real narrator, changing only the configured model name between the two halves — 12/12:

embedding_model = "nomic-embed-text"      (calibrated)
  status  semantic_enabled=True  semantic_calibrated=True
  context semantic_used=True, retrieved ['canon.md']

embedding_model = "mxbai-embed-large"     (not calibrated)
  status  semantic_calibrated=False, semantic_enabled=False
          "“mxbai-embed-large” has no measured relevance calibration in this
           build, so semantic retrieval is disabled and retrieval is lexical
           only. Lexical search and story play are unaffected."
          calibrated_models = ['nomic-embed-text']
  context semantic_used=False, semantic_floor=0.0
          retrieved ['canon.md'] — admitted_by "lexical"
  off-topic scene: zero imported chunks, no imported section
  a real narrator turn still plays

DB.6 A consequence worth recording

Four M7 suites configured a placeholder embedding model name (embed-test). Under the new policy that name is uncalibrated, so those suites would have quietly exercised the lexical-only path while appearing to test semantic retrieval — the M7-F2 failure mode in a new guise. They now name nomic-embed-text, which is what their stubs are built to model, with a comment saying why. test_knowledge_calibration.py is where the uncalibrated path is exercised deliberately.


DC. Final regression gate

Gate Result
backend suite 939 passed, 14 skipped, 1 warning
focused M7 knowledge suite 103 passed
M7 real-model suite (nomic-embed-text) 7 passed
M7 real-narrator authority tests 25/25
§12 relevance matrix, real model 13/13
export/import suite 56/56
migration suite 23/23
frontend lint exit 0, 7 warnings (M6 baseline)
frontend production build success
Docker build success; image verified reporting the calibration registry
browser checks 43/43, Firefox 154.0.1
network zero sockets for import/index/retrieval/turn; embeddings only to the configured endpoint

Confirmed once more, after the calibration change:

irrelevant query -> zero imported knowledge                      PASS
relevant Canon only -> relevant Canon selected                   PASS
Reference does not become Canon                                  PASS
Inspiration cannot establish conflicting facts                   PASS
C05 remains non-vacuous (revival.md retrieved at cosine 0.656)   PASS
hidden Canon retrieved only when relevant, never disclosed       PASS
lexical fallback works                                           PASS
an uncalibrated embedding model cannot recreate M7-F1            PASS
no new outbound network path                                     PASS
query growth bounded (list 5, detail 4, retrieve 10, ctx 21)     PASS

DD. Verdict

M7 READY FOR SIGNED COMMIT

Both blocking findings from the independent review are closed with real-model and deterministic evidence, and both closeout items are resolved:

  • M7-F1 — retrieval admits nothing without candidate-set-independent evidence, and returns zero chunks on an unrelated scene.
  • M7-F2 — the suite models the real model's similarity floor and fails 13/18 against the pre-corrective implementation.
  • Calibration boundary — an uncalibrated embedding model degrades safely to lexical-only instead of borrowing a threshold measured against another model.
  • 53/54 — a review-harness false positive, withdrawn; the suite is 56/56.

M7 is complete and staged. It is not committed: the repository owner signs milestone commits.


DE. Corrections made during final staging

Three defects were found by the final repository verification itself, after the regression gate in §DC. All three are documentation or commit-message only — no application or test code was touched, so the §DC results stand unrepeated.

  1. The prepared commit message carried the corrective pass's test count. It read 927 passed, 14 skipped and 91 new across six files; the closeout run measures 939 passed, 14 skipped, and the seventh new file (test_knowledge_calibration.py, 12 tests) brings the new total to 110. The numbers reconcile exactly: 927 + 12 = 939, and 836 + 110 − 7 real-model skips = 939 passed with 7 + 7 = 14 skipped. Corrected.

  2. The required documentation line existed only in this report. The closeout brief specifies the sentence "Semantic admission is calibrated for nomic-embed-text; uncalibrated embedding models fall back safely rather than borrowing its threshold." §13.3 of TECHNICAL-DESIGN.md explained the policy at length but never stated it in that form, so a reader skimming headings could miss the invariant. It is now the section's opening line.

  3. BUILD-MILESTONES.md line 3 still read "In implementation … M7 — next to brief". The M7 section body had been updated to Status: COMPLETE and M8 authorized, but the document's own top-level status line had not, so the file contradicted itself. Corrected to M1-M7 complete and accepted … M8 next to brief.

The third is the one worth noting as a pattern: a status recorded in two places drifts, and the copy that drifts is the one not being edited during the work.