The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.
All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.
So the shape now is one entry point and one screen:
Campaigns -> Campaign -> Story
State · Knowledge · Context · Save Points · Settings
Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.
Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.
Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.
The two defects worth the space:
A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.
And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.
Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.
`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.
Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.
Backend, and only what the browser could not otherwise reach:
AdventureCreate.opening a start action could only come from a Scenario, so
every campaign made in the new setup flow opened on
a blank page. Same node, same code path.
canon_rules campaign_canon has been the highest authority in a
campaign since M5, read by the prompt builder and
the state validator, and had no API at all — a
fixture had to write it with SQL.
a 401 and a 429 message the last user-facing text describing a hosted
deployment. One told the reader to check an API key
that has not existed since M2.
No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.
The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.
They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.
A verification pass over all of it then found three more, each by driving the
product rather than reading it:
Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.
The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.
And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.
Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.
`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.
Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:
The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.
Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.
The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.
The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.
Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
77 KiB
M7 Implementation Review — First-Class Imported Knowledge Library
Review date: 2026-09-06
Reviewer: independent review pass (the M7 build summary was treated as claims to verify)
Tree reviewed: the staged working tree on m7-imported-knowledge. There are no M7 commits.
Hostnames and LAN addresses in this report are placeholders, following the convention the M2 report set. The endpoint used throughout is a trusted-LAN Ollama reached over HTTPS with a locally issued certificate, serving
qwen2.5:3b-instructandnomic-embed-text.
A. Executive result
PASS WITH CORRECTIVE WORK REQUIRED
M7 builds the subsystem the specification asks for, and most of it is demonstrated rather than asserted. The whole source lifecycle, campaign isolation, the security boundary, the network boundary, export/import, migration and the provenance surfaces all hold up under independent testing with real local models and a real browser. The five acceptance conditions the implementation could not close — C05, G06, G07, G10 and the hidden-Canon spoiler test — were exercised against a real narrator in this review and all five pass.
One blocking defect, in retrieval:
Imported knowledge is injected into every prompt regardless of relevance. With semantic retrieval enabled, a scene with no relationship to any source still retrieves — in the measured case, all five sources, 326 tokens, every turn, including narrator-only hidden Canon. The admission threshold is relative only, and the best candidate always clears a share of itself, so something is always admitted.
This contradicts IMPORTED-KNOWLEDGE-DESIGN.md §30, which requires in as many
words that irrelevant Canon must not be included merely because it is
authoritative, and it widens the hidden-Canon spoiler surface to every turn.
It was invisible to the implementation's own suite because the stub embedder
that suite uses is far more discriminative than the real model (§I).
Everything else in required M7 scope passes.
B. Repository and provenance state
Verified before any review work:
branch m7-imported-knowledge
HEAD a6e9c7a32bdf42f1e4cb837b70721d89422cef8d
staged 47 files
unstaged 0
untracked 0
prepared commit message .git/M7_MSG (4069 bytes, present)
LICENSE (sha256) 074062599d3b11b252a70e37086d6cef3820370b55c35efa3506375af914a13c
M7 base a6e9c7a ancestor of HEAD yes
pinned AI-DnD d72f7c1b… ancestor of HEAD yes
The actual state matches the expected state exactly. No difference to record.
LICENSE is not among the staged files and its digest is unchanged from the M6
base.
C. Implementation inventory
Nine modules under backend/app/knowledge/, one router, three tables, one
virtual table, one migration, one browser panel, one new dependency.
knowledge/classes.py the three classes, their weights, the prompt framing
knowledge/records.py the dataclasses, dependency-free (breaks an import cycle)
knowledge/chunking.py deterministic heading-aware chunker, 60-800 tokens
knowledge/fts.py SQLite FTS5, porter-stemmed, scoped and LIMITed in SQL
knowledge/importer.py validate, hash, store, chunk, index — one transaction
knowledge/embeddings.py local Ollama vectors through the shared provider
knowledge/retrieval.py query construction, hybrid merge, rerank
knowledge/inject.py the budgeted cut and the rendered sections
routers/adventures/knowledge.py the HTTP surface (multipart upload only)
knowledge_sources / knowledge_chunks / knowledge_embeddings + knowledge_fts
migration 92; the FTS table is an after_create/before_drop DDL hook on
knowledge_chunks, so it is created and dropped with the table it indexes
frontend: KnowledgePanel.jsx, knowledge.css, Insights provenance rows
dependency: python-multipart 0.0.32
Source content is stored in SQLite, not on disk. No filesystem path is ever constructed from a filename; see §H.
D. Independent test method
Evidence was gathered end-to-end first and by source inspection last.
- A real narrator —
qwen2.5:3b-instructon the configured trusted-LAN Ollama — for C05, G06, G07, G10 and the hidden-Canon test. Nothing about the model was mocked. - A real embedding model —
nomic-embed-texton the same host — for every retrieval measurement. This is what exposed the blocking defect: a stub cannot reproduce what a real embedding does to unrelated English. - A real server process over HTTP, with a real SQLite file, for every lifecycle, ranking, migration and bundle test.
- A second server process on an empty database file for the export/import round trip, so "clean data directory" means what it says.
- A real Firefox 154.0.1 over WebDriver against the built frontend, for H06, H07, G08, G09, F05 and F06.
- Socket-level capture (
socket.socket.connectinstrumented in the server process) for every network claim. builtins.open/os.openinstrumentation for the H08 filesystem claim.- A real server restart across a genuinely rewound schema, for migration.
The standard fixture is TEST-CAMPAIGN-FIXTURE.md §12 — the campaign's narrator
rules (§4), global canon (§11) and the three imported files verbatim.
Every negative result below is preceded by a positive control.
A fixture discrepancy worth recording
The implementation's own acceptance tests use the shorter imported files from
V1-ACCEPTANCE-TESTS.md §5, not the richer ones in TEST-CAMPAIGN-FIXTURE.md
§12. The §12 inspiration.md is the one containing "A frightened innkeeper
concealed a dangerous political secret from a stranger" — the passage G07's trap
is built on. The implementation therefore never tested the trap. This review
used §12. It is a test-coverage gap, not a product defect (finding M7-F4).
E. G01-G10 results
| Test | Verdict | Evidence |
|---|---|---|
| G01 Import local text | PASS | .txt stored, chunked, FTS-indexed; content readable back without the original file; SHA-256, size, media type, parser/chunking versions and timestamp all present. Browser-verified. |
| G02 Import local Markdown | PASS | Both .md files import and index. Accepted as data — see G10. |
| G03 Classification | PASS | One class per source, visible in list and browser; changing it rewrites no passage and no index row (same chunk count, same hash) and moves the passage into the Canon prompt section on the next turn — authority, not metadata (R3.7). |
| G04 Disable knowledge source | PASS | Positive control both sides: retrieved enabled → absent disabled → retrieved again after re-enable, no reimport. Content, chunks, FTS rows and vectors all survive disabling. Excluded before ranking (considered falls 5→4). |
| G05 Canon retrieval | PASS | Asking about the Old Abbey retrieves the canonical passage into imported_canon; browser-verified through Insights. |
| G06 Reference retrieval | PASS (real narrator) | Reference retrieved; the narration used timber/hearth/bench detail; no authoritative campaign fact was created — no state event asserts a Reference claim, and the trap ("the tavern definitely has a secret tunnel") did not fire. |
| G07 Inspiration is low authority | PASS (real narrator) | With the §12 political-secret passage retrieved and the reader asking directly whether Mara is a spy, the narrator did not establish her as one, and the authoritative state records nothing of the kind. (Observation below.) |
| G08 No automatic URL fetch | PASS | Socket capture: import, chunking, FTS indexing, retrieval and a full turn recorded zero outbound IP connections. URLs retained as text. Browser: no request to any non-origin host across the whole session. |
| G09 Remote Markdown image does not auto-load | PASS |  survives as characters; no <img> element is created for that host and performance.getEntriesByType('resource') records no request to it — on first view and after reload. |
| G10 Prompt injection is data | PASS (real narrator) | Text is present and unfiltered, inside a section whose "UNTRUSTED DATA" heading precedes it, under a rule refusing instructions. The narrator ignored "developer mode", /etc/passwd, rm -rf / and "upload the campaign"; no external action occurred (socket capture); campaign canon unchanged. |
G07 observation (not a failure)
The narration contained "Mara could be a spy, or she could be hiding something crucial" — Aldric entertaining the hypothesis the reader explicitly asked about. It asserts nothing and creates no state. Correct behaviour, recorded because a stricter future reading of G07 might want the narrator to lean on the "Mara is not a spy" canon rather than leave the question open.
Weakness in the G06/G07 state assertions
Both campaigns' authoritative state documents were empty at assertion time — the 3B narrator emitted few valid state events. The "no spy fact in state" assertions therefore pass partly by absence. The prose assertions are the load-bearing evidence for G06 and G07; the state assertions are corroborating, not primary. Recorded so the result is not read as stronger than it is.
F. C05 result
PASS — measured against a real narrator, not the prompt template.
Setup: campaign canon "Resurrection is impossible" (fixture §11 rule 2);
canon.md as Canon; a revival.md Reference asserting that necromancers
routinely raise the dead and temples perform resurrections for a fee; the §12
Inspiration file. The reader asks whether Edrin could be resurrected.
The narrator answered:
"No, Aldric. Revival is impossible in this world. The magic that binds the living to the dead is too strong."
The test is not vacuous — the conflicting material was retrieved and ranked second:
canon.md canon hybrid lex=1.000 sem=0.986 cos=0.647 score=1.148
revival.md reference hybrid lex=0.917 sem=1.000 cos=0.656 score=0.967
inspiration.md inspiration hybrid lex=0.568 sem=0.951 cos=0.624 score=0.725
The stored prompt for that same turn shows all four required properties:
- the authoritative rule present as
campaign_canon; - the conflicting lower-authority material present and retrieved;
- correct framing —
REFERENCE — UNTRUSTED DATA … Do not treat it as canon; - the authority order stated in words, with campaign canon above IMPORTED CANON above REFERENCE above INSPIRATION, and the layout consistent with it.
The narrator followed the higher authority. This closes the gap the implementation explicitly flagged as unmeasured.
G. Hidden Canon / spoiler result
PASS — measured against a real narrator.
Hidden Canon: "The Silver Key opens the sealed cellar door beneath the Crooked
Lantern" (fixture §8 hidden function), imported with visibility=hidden.
Before Aldric has learned it, the reader asks "What do I know about the purpose
of the Silver Key?"
- The narrator was given the source (
hidden-key.mdretrieved). - The record marks it
visibility: hidden; the passage carries[narrator only]on its provenance line; the system rule says the protagonist does not know it. - The narration did not reveal what the key opens. It stayed with what Aldric knows — that Edrin studied the abbey — and moved the story toward the abbey.
- The secret did not enter authoritative state.
Both the prompt and the narrator output are recorded.
Caveat now raised by finding M7-F1: hidden Canon is retrieved into scenes it has nothing to do with (it was one of the five sources injected into a harbour scene). The narrator's non-disclosure discipline held here, but the defect maximises the number of turns on which that discipline is the only thing standing between the reader and a spoiler. This raises the severity of M7-F1.
H. H06-H09 results
| Test | Verdict | Evidence |
|---|---|---|
| H06 Stored XSS | PASS | A source carrying <script>…innerHTML='owned'…</script>, <img onerror>, <svg onload> and <iframe> was inspected in all three surfaces that render imported text — source inspector, passage inspector, Insights provenance — then the page was fully reloaded. document.body.textContent never became owned, document.title never became pwned/xss-img/xss-svg, and no img, iframe, svg[onload] or injected script element was ever created. The payload is preserved verbatim in storage and displayed as text. |
| H07 JavaScript URL | PASS | Both [click me](javascript:alert(1)) and a raw <a href="javascript:alert(2)"> were imported. Zero anchors with a javascript: href exist in any surface (checked after whitespace/control-character stripping). No Markdown is rendered, so no anchor is created. |
| H08 Path traversal | PASS by demonstrated non-reachability | Twelve hostile filenames (../../outside.txt, ../../../../etc/passwd.md, Windows separators, /etc/shadow.md, ....//, percent-encoded, NUL-embedded, RTL-override, reserved device names). builtins.open/os.open instrumentation recorded 0 write-mode file opens outside the database. Every stored name is an inert basename with no separator, no .., no NUL, not absolute, not dot-leading. Every stored source holds the uploaded bytes — nothing read /etc/passwd. The OpenAPI schema exposes no path/file/dir parameter on any knowledge route. |
| H09 ZIP slip | NOT APPLICABLE | Asserted, not assumed: no zipfile, tarfile, extractall, shutil.unpack, gzip.open or py7zr appears anywhere in backend/app/knowledge/, the knowledge router, or backend/app/bundle.py. The import surface takes one text file; the bundle is JSON and never touches the filesystem. The code path that makes this true is routers/adventures/knowledge.py::import_source, which reads UploadFile bytes and passes them to importer.import_source(raw=…). |
I. FINDINGS
M7-F1 — Imported knowledge is injected regardless of relevance — BLOCKING
Severity: high
Violated requirement: IMPORTED-KNOWLEDGE-DESIGN.md §30 ("Do not include
irrelevant Canon merely because it is authoritative"); §26 and §29's
relevance-then-authority pipeline; CONTEXT-AND-MEMORY.md §29's bounded,
relevant knowledge layer.
Owner: M7 corrective.
Reproduction. Five realistic sources (Canon, Reference, Inspiration, a
resurrection Reference, hidden Canon), all embedded with nomic-embed-text.
The scene is set to text with no entity, lexical or semantic relationship to any
of them.
scene: "Aldric studies the tide tables and the harbour master's ledger
of container tonnage."
considered=5 floor=0.2875
USED hidden-key.md canon lex=1.000 sem=1.000 cos=0.548 score=1.150 51 tok
USED canon.md canon lex=0.000 sem=0.772 cos=0.423 score=0.772 56 tok
USED reference.md reference lex=0.895 sem=0.927 cos=0.508 score=0.902 70 tok
USED revival.md reference lex=0.000 sem=0.730 cos=0.400 score=0.621 86 tok
USED inspiration.md inspiration lex=0.000 sem=0.871 cos=0.477 score=0.610 63 tok
-> 5 of 5 sources injected, 326 tokens
Four independent unrelated scenes were measured (harbour, kitchen, orbital mechanics, surgery). Every one injected all five sources and 326 tokens. A second campaign with two sources injected both. The positive control — a scene about the Old Abbey — retrieves correctly, so the mechanism is working; it is the cut that is wrong.
It is not confined to the "nothing relevant" case. With a strongly relevant Canon source present, irrelevant Canon is still admitted alongside it:
scene: the abbey crypt
USED exact-canon.md canon cos=0.674 score=1.133 <- relevant
USED ceres-canon.md canon cos=0.442 score=0.584 <- a docked freighter
USED crypt-reference reference cos=0.708 score=0.897
Ranking order is correct throughout (relevant Reference 0.897 > irrelevant Canon 0.584). The defect is the admission threshold, not the ranking.
Root cause. Two independent channels, one shared cause — there is no absolute notion of "this is not a match" on either path.
- Semantic.
retrieval.py:266-267computesfloor = max(MIN_RELEVANCE, best * RELEVANCE_FLOOR_SHARE)withMIN_RELEVANCE = 0.02andRELEVANCE_FLOOR_SHARE = 0.25. Scores are normalized per query against the best of their own path, so the best competing candidate scores 1.0 × its class weight by construction and can never fail a threshold defined as a share of itself.MIN_RELEVANCEbinds only when the best score is below 0.08, which no real embedding produces — it is effectively dead code. Because the semantic path scores every embedded chunk, there is always a best, so with semantic retrieval enabled at least one passage is injected on every turn, unconditionally. - Lexical.
fts.pyjoins query terms withORand applies no minimum-match-quality rule, and the stop list is 42 words. Common words carry full matches:hidden-key.mdscoredlex=1.000against an orbital-mechanics scene on the word "before", and against a harbour scene on "Aldric" — the protagonist's name, which is in the story tail of essentially every query and in any source that mentions them. The lexical-only path leaked in the negative control too, so this is not a semantic-only problem.
Calibration data for the corrective pass. Against nomic-embed-text, over
7 scenes × 5 sources:
relevant scenes n=15 min=0.407 max=0.698 mean=0.546
irrelevant scenes n=20 min=0.392 max=0.551 mean=0.462
overlap: YES (irrelevant max 0.551 vs relevant min 0.407)
Raw cosine is not cleanly separable for this model, so a single absolute
cosine threshold will trade recall for precision and must be chosen and
documented deliberately — and it is model-dependent, exactly as
memorybank.REDUNDANT_SIMILARITY records for its own threshold. A defensible
shape, for the implementer to choose among rather than a prescription:
- require a real lexical match on a term that is not ubiquitous (exclude the protagonist/entity names that appear in every query, and widen the stop list), or an absolute cosine above a documented, model-calibrated value; and
- keep the relative floor as a second filter, not the only one.
User impact. Every turn spends knowledge budget on material about nothing in the scene. Worse, narrator-only hidden Canon is supplied on turns that have nothing to do with it, which maximises the exposure of the one feature whose entire purpose is controlled disclosure (§G). And the narrator is asked to reconcile authoritative-sounding Canon that is irrelevant to what the reader just did.
Blocks M7 acceptance: yes.
M7-F2 — The retrieval-quality suite cannot detect M7-F1 — BLOCKING (test)
Severity: high
Violated requirement: TECHNICAL-DESIGN.md §18.1 (the M2 wiring rule: test
a real consumer path); the standing milestone rule that at least one real
provider path must be exercised.
Owner: M7 corrective.
tests/test_knowledge_retrieval_quality.py::test_irrelevant_canon_does_not_win_on_class_alone
passes, and asserts exactly the property M7-F1 violates. It passes because its
SceneEmbedder stub is far more discriminative than the real model. Measured on
identical texts:
abbey-canon (relevant) 0.7559 1.000 | 0.7505 1.000
crypts-ref (relevant) 0.3392 0.449 | 0.6237 0.831
ship-canon (IRRELEVANT) 0.1499 0.198 | 0.4365 0.582
cookery (IRRELEVANT) 0.0818 0.108 | 0.4354 0.580
desert (IRRELEVANT) 0.0642 0.085 | 0.4349 0.580
relevance floor = 0.250
-> stub: irrelevant peaks at 0.198 => correctly excluded => test passes
-> real: irrelevant peaks at 0.582 => ADMITTED => defect
The stub's own docstring states the choice that causes this: "There is deliberately no constant component. One would put a similarity floor under every pair." A similarity floor under every pair is precisely what a real embedding model has. The stub removed the property that makes the defect detectable.
This is the M2 and M6 lesson recurring in a third form: a mocked provider left a real defect invisible behind a green suite. The corrective pass needs a retrieval test whose embedder reproduces the real model's similarity floor, or one that runs against the real model.
Blocks M7 acceptance: yes — the fix for M7-F1 cannot be validated by the current suite.
M7-F3 — MIN_RELEVANCE is unreachable dead configuration — non-blocking
Severity: low Owner: M7 corrective (fold into M7-F1's fix).
classes.MIN_RELEVANCE = 0.02 is presented as the absolute half of a two-floor
design, but with per-query normalization it can only bind when the best score is
below 0.08. It never fires. The comment beside it explains a design that the
code does not implement. Whatever replaces it in M7-F1's fix, the constant and
its comment should stop describing protection that does not exist.
M7-F4 — The acceptance tests use the wrong fixture files — non-blocking
Severity: medium
Violated requirement: V1-ACCEPTANCE-TESTS.md §4 names
TEST-CAMPAIGN-FIXTURE.md as the standard campaign; BUILD-MILESTONES.md M7
requires the fixture be used "wherever practical".
Owner: M7 corrective (test-only).
tests/test_imported_knowledge.py uses the three short files from
V1-ACCEPTANCE-TESTS.md §5 rather than the richer TEST-CAMPAIGN-FIXTURE.md
§12 versions. The consequence is concrete: §12's inspiration.md contains "A
frightened innkeeper concealed a dangerous political secret from a stranger",
which is the trap G07 exists to test against the canon rule "Mara is not a spy".
The implementation's G07 test used the silent-hall passage instead and so never
exercised the trap. This review did, and it passes — but the repository's own
suite does not cover it.
Also missing from the suite: the fixture's campaign canon (§11), its narrator rules (§4) and its hidden Canon (§6, §8).
M7-F5 — Formatting control characters survive into displayed filenames — non-blocking
Severity: low Owner: M7 corrective or M8 (display).
safe_filename strips separators, NULs and leading dots, but leaves Unicode
bidirectional-override characters. "exe.dm.md" is stored and displayed
verbatim, which can render as a different extension in the list. There is no
security impact beyond display spoofing — the name is never used as a path, and
H08's stronger guarantees hold — but a name shown to a reader should not be able
to lie about itself. Stripping the Cf (format) category would close it.
M7-F6 — Narrator state blocks can leak into displayed prose — pre-existing, not M7
Severity: low Owner: later milestone (M8/M11), recorded for completeness.
Twice during real-narrator testing, qwen2.5:3b-instruct emitted its state
block fenced as ```json rather than ```state, so
narrative.extract.split did not recognise it and the raw JSON was displayed to
the reader. This is an M5 protocol-robustness issue with a small model, entirely
outside M7's scope and unaffected by it. Noted because it was observed under
this review's real-narrator runs and a reader would see it.
J. I05 and export/import
PASS. A genuine round trip: source server on one database file, destination server as a separate process on a different, empty database file.
Before export: all three classes, one disabled source, one hidden Canon source that is also always-include, one six-chunk source, and a real narrator turn that used imported knowledge.
E1 bundle carries a knowledge section (5 sources) PASS
E2 and no derived data (no chunks/vectors/index rows) PASS
E4 imports into a clean data directory PASS
E5 every source arrived PASS
E6 classification survived (all 5) PASS
E7 enabled/disabled state survived (all 5) PASS
E8 visibility survived (all 5) PASS
E9 content hash survived (all 5) PASS
E10 title and original filename survived (all 5) PASS
E11 content byte-identical (all 5) PASS
E12 chunks rebuilt, same count (all 5) PASS
E13 always-include survived PASS
E14 rebuilt chunks identical in text, order and hash PASS
E16 lexical retrieval works immediately, with no reindex PASS
E16b the disabled source stays out after import PASS
E17 semantic vectors rebuild locally on demand PASS
E17b semantic retrieval works in the imported campaign PASS
E18 exporting did not disturb the original turn's evidence PASS
E20 the story itself round-tripped PASS
E21 no external source path is needed PASS
I05's literal condition — metadata and classification survive — passes. M7's broader scope condition — a usable local knowledge subsystem survives — also passes: chunk identity is reproduced exactly by the deterministic chunker, lexical retrieval works on arrival with no reindex step, and vectors rebuild against whatever model the importing machine has.
Recorded limit, correctly documented and not an M7 defect. The bundle
carries no context snapshots, so the imported campaign has no historical prompt
provenance — for imported knowledge or for any other component. That is exactly
what a pre-M7 bundle already did for summaries, memories and everything else,
it is stated in DATA-MODEL.md §29 and BUILD-MILESTONES.md M7 § Status, and
it belongs to M9. Nothing M7 creates becomes a dangling reference, because no
ids are exported.
K. F05 / F06 completion
Both complete. M6 left them partial only because imported knowledge did not exist; it now does, and the gap is filled in the browser, not only the API.
For a completed narrator turn, the Insights panel showed — rendered, in Firefox:
▸ hostile.md CANON · Hostile · passage 1 · hybrid match
(lexical 0.64, semantic 0.90 → 1.00) · 124 tok
<the full rendered chunk text>
2 passage(s) considered, 182 of 2112 knowledge tokens used
· searched on: aldric, studies, broken, circle, abbey, crypt, westhaven
F06's requirements, each checked in the rendered DOM: source filename/title ✓, classification ✓, chunk identity ✓, rendered chunk text ✓, retrieval score and mode ✓, token cost ✓, and why it was searched ✓.
F05's prompt-section view showed every component with a readable name and a token cost:
narrator prompt 57 tok · narrative state (reporting rule) 467 tok ·
campaign canon 109 tok · imported knowledge (rules for using it) 199 tok ·
ai instructions 140 tok · plot essentials 17 tok · story history 130 tok ·
imported canon (retrieved) 232 tok · narrative state (current) 62 tok ·
length guidance 51 tok · narrative state (emit reminder) 32 tok
Historical evidence survives deletion of its source: a turn was generated using
canon.md, the source was deleted, and the old turn's snapshot still carried the
rendered text byte-identically (§L, L15).
L. Source lifecycle and atomicity
31 of 32 checks passed; the one failure was a defect in the review's own test.
L1 .txt import indexes PASS
L2 .md import indexes PASS
L3 unsupported extension refused PASS
L4 binary content refused despite .txt extension PASS
L5 invalid UTF-8 refused, not mangled PASS
L5b empty source refused PASS
L6 1,040,382 bytes accepted (limit 1,048,576) PASS
L7 1,059,372 bytes refused, message says what to do PASS
L7b nothing of the refused source stored PASS
L8 NFC/NFD normalize to the same hash PASS
L8b the conflict names the existing source PASS
L8c a CRLF copy is detected as a duplicate PASS
L9 explicit duplicate override honoured PASS
L10 disable keeps content, chunks and FTS rows PASS
L11 re-enable restores retrieval, no reimport PASS
L12b no orphan FTS rows anywhere PASS
L12c no orphan chunks anywhere PASS
L12d every stored source is 'ready' — none half-built PASS
L14 delete removes source, chunks, FTS rows and vectors PASS
L14d no dangling FTS rowids after delete PASS
L15 the historical turn still says what it was supplied PASS
L15b byte-identical to what was supplied PASS
L15c the story was not rewritten by the delete PASS
L12 was a review test error, not a product defect. It expected a .md file
whose content is %PDF x to be refused. That content is valid readable text, so
accepting it is correct — SECURITY-THREAT-MODEL.md §21 asks for rejection of
binary data, which L4 demonstrates with a real ZIP header. The invariants that
matter (no orphans, nothing half-built) all pass.
Atomicity is additionally structural: retrieval is gated on
index_state = 'ready', so even a hypothetical partial commit would be inert
rather than wrong.
M. Ranking and authority
15 of 16 checks passed; the failure is M7-F1.
R3.1 irrelevant Canon is not retrieved FAIL -> M7-F1
R3.1b relevant Reference IS retrieved (positive control) PASS
R3.1c relevant Reference outranks irrelevant Canon PASS (0.897 > 0.584)
R3.2 at comparable relevance, Canon scores above Reference PASS (1.133 > 0.897)
R3.2b Reference framed as non-canon wherever it lands PASS
R3.2c Inspiration framed as establishing nothing PASS
R3.3a the lexical path contributed a winner PASS
R3.3b the semantic path contributed a winner lexical missed PASS
R3.4 no chunk appears twice in the selection PASS
R3.4b a chunk both paths found is marked hybrid, not doubled PASS
R3.5 a disabled source is excluded before ranking PASS (considered 5->4)
R3.6 Campaign A's sources never appear in Campaign B PASS (considered=0)
R3.6b B retrieves its OWN source (positive control) PASS
R3.7 reclassifying moves the passage into the Canon section PASS
R3.7b and it is no longer framed as Reference PASS
Ordering, deduplication, class weighting, disabled-source exclusion and campaign isolation are all correct. Only the admission threshold is wrong. That is a narrow, well-localised fix.
N. Local-only and network evidence
PASS. No new network path. Socket capture in the server process:
import + chunk + FTS5 index destinations: none
lexical retrieval + prompt assembly destinations: none
a full narrator turn destinations: none
semantic indexing + retrieval destinations: the configured Ollama
host and port, and nothing else
(the configured Ollama, its configured port)
- A public endpoint written into the database behind the settings API is refused at request time on the knowledge path, and the refusal opened no socket.
- The policy refuses
api.openai.comand public addresses, and permits loopback and the configured trusted-LAN host. - An import still succeeds lexically while the endpoint is refused — the library does not depend on the inference host being reachable.
- The knowledge subsystem builds no HTTP client of its own, embeds only
through
memorybank.embedding_provider, and reimplements no TLS handling, so the allowlist, the request-time re-check and the OS/private-CA trust union all apply unchanged (ADR 011, ADR 002). - No remote vector service, cloud host, telemetry, CDN asset or model-download path exists in it. (Two apparent hits in an initial scan were false positives: "cohere" inside "coherent", "pulled" in prose.)
Consistent with ADR 004 and SECURITY-THREAT-MODEL.md §§12-21.
O. Migration
PASS — 23/23, through a real server restart.
A campaign was built with history, a divergence, a Save Point, narrative state, a summary, a memory and a derived-status row. The M7 tables and the FTS index were then dropped and the stamp rewound to 91, making the file genuinely M6 — verified before migrating. The server process was stopped and a new process opened the same file.
stamp 91 -> 92 on restart
every M7 table created, including knowledge_fts
story history identical
branches identical
Save Points identical
narrative state identical
memories identical
summaries + derived status identical
active head unmoved (2, 6)
pre-M7 prompt provenance readable and unchanged
old snapshots no knowledge section forced onto them
empty library, no imported section in the next prompt
normal turn / Undo / Redo / Save Point restore / inspection all work
then: import a source and retrieve it in a real narrator turn PASS
An old campaign needs no knowledge sources, and gains the subsystem cleanly.
P. Browser evidence
Firefox 154.0.1, headless, WebDriver, against the built frontend served by the real backend. 22/22 passed, covering all three surfaces that render imported text plus a full page reload. See §H and §K for the specific results.
No severe console errors. No request to any non-origin host across the entire session.
The implementation's own browser suite (43 checks) was also re-run during the implementation pass; this review's suite is independent of it.
Q. Performance and scaling
PASS — measured, flat.
sources/chunks list detail retrieve context considered used ctx
3 / 12 5q 4q 9q 20q 12 3 130ms
15 / 60 5q 4q 9q 20q 55 3 108ms
39 / 156 5q 4q 9q 20q 67 3 188ms
Query counts did not move while the library grew 13×. No N+1 on the source list,
the chunk listing, retrieval or prompt construction. Candidates are bounded in
SQL (67 of 156 chunks, cap 80); the FTS query carries a LIMIT; the semantic
catalogue read fetches ids only, not vectors or passage text; the vector read is
a single batched query.
Characteristic to record, not a defect: the semantic scan is a linear cosine
over the campaign's vectors in Python, capped at 4,000 passages with the
shortfall reported. Query count is flat; per-turn CPU grows linearly with library
size. BUILD-MILESTONES.md M7 § Status already records this as deliberate v1
debt (no approximate-nearest-neighbour index).
R. Regression gates
| Gate | Result |
|---|---|
| backend suite | 920 passed, 10 skipped, 1 warning (M6 baseline 836/7/1) |
| focused M3-M6 regressions | 372 passed, 1 warning — context/memory, memory retrieval, head cursor, Save Points, narrative state, bundle v2, tree migration, endpoint policy, local-only surface, egress, state revert, pre-M5 compatibility |
| frontend lint | exit 0, 7 warnings — identical to the M6 baseline |
| frontend build | success |
| docker build | success; image verified running with python-multipart 0.0.32, schema 92, porter-stemmed FTS |
The single warning is the pre-existing StarletteDeprecationWarning from
fastapi.testclient. The 7 lint warnings are the pre-existing
components.jsx/SchemaEditor.jsx fast-refresh notes. The 3 additional skips
over the M6 baseline are test_knowledge_real_model.py, which skips without a
configured endpoint; they were run explicitly in this review with one.
S. New dependency
Approved.
package python-multipart 0.0.32
licence Apache-2.0
requires nothing (Requires-Dist: None)
files 24
declared backend/requirements.txt python-multipart>=0.0.9 (with reason)
pinned backend/requirements.lock python-multipart==0.0.32
docker installs from requirements.txt — verified in the built image
clean venv installs from requirements.lock — verified, app imports
It is Starlette's own multipart parser, needed because the upload surface is
UploadFile. It adds no transitive dependency, no framework surface and no
network path — and the upload surface it enables is precisely why no endpoint
accepts a filesystem path (H08). Reason, pinning and reproducibility all check
out.
T. Deferred features — confirmed acceptable
None of these fails M7; each remains consistent with the active design and is
recorded in BUILD-MILESTONES.md M7 § Status and IMPORTED-KNOWLEDGE-DESIGN.md
§76.
- Entity linking (§33), manual priority (§35), scene pinning (§36) — optional in the design; §36's stated v1 minimum (always-include for critical Canon) is implemented.
- Source version history (§14, §43) — the design explicitly permits "a simpler replace-with-provenance model … if documented", and it is documented.
- Canon scope metadata (§45) — §45 itself says explicit current-state precedence may be sufficient for v1, which is what was built.
- Tags (§34) — evaluated in context as instructed. They are unused optional metadata: nothing in the API, the source inspector, the prompt or the docs promises tags, so their absence breaks no implemented promise. Acceptable.
- Canon-vs-Canon conflict detection (§42) — absent, and outside M7's scope list. It produced no authority failure in any required M7 scenario in this review. Recorded as later-v1 debt.
No planning document was weakened to match the implementation.
U. Planning-document recommendations
IMPORTED-KNOWLEDGE-DESIGN.md§29-30 — add, from M7-F1, that a relevance threshold expressed only as a share of the best candidate cannot exclude anything, because the best candidate always clears a share of itself. A retrieval design needs an absolute notion of "no match" on each path, and for embeddings that value is model-dependent and must be measured and written down, asmemorybank.REDUNDANT_SIMILARITYalready is.TECHNICAL-DESIGN.md§18.1 — extend the M2 wiring rule with M7-F2: a stub provider must not be more discriminative than the real one. A stub with no similarity floor cannot detect a retrieval defect that only exists because real embeddings have one. State the requirement that retrieval thresholds be validated against a real model.V1-ACCEPTANCE-TESTS.mdG05-G07 — add the negative control as a pass condition: a query unrelated to every source must retrieve nothing. Without it, G05/G06/G07 can all pass while retrieval is unconditional.V1-ACCEPTANCE-TESTS.md§5 /TEST-CAMPAIGN-FIXTURE.md§12 — the two documents give different contents for the same three filenames. §12 is the richer and is what the traps are built on. Reconcile them, or state which is normative for the G-series.CONTEXT-AND-MEMORY.md§46 — record from §G that hidden Canon's protection is prompt discipline plus retrieval scoping; if retrieval is unconditional, prompt discipline is carrying the whole load on every turn.
V. M8 implications
- The knowledge panel and the Insights knowledge rows are functional and complete for M7's purposes, and every M7 behaviour is reachable in a browser with no database access. M8 owns their design.
- M8 should not have to touch retrieval; M7-F1's fix belongs to the M7 corrective pass, before M8 begins.
- Two M5 section labels (
state_rule,state_reminder) that rendered as raw keys were fixed during the M7 pass; the rest of the panel's visual design is M8's. - §47 of
BROWSER-UX-SPEC.mddescribes an import preview step that is not built, and §52's retrieval-usage count is absent. Both are marked optional; M8 may reconsider. - If M7-F1's fix means retrieval sometimes returns nothing, the panel and the inspector should say so plainly rather than showing an empty section — a small M8 affordance worth planning for.
W. Verdict
PASS WITH CORRECTIVE WORK REQUIRED
M7 delivers the subsystem, and the parts that work are demonstrated rather than asserted. Against a real narrator, a real embedding model, a real browser, real socket capture and a real migration, the following all hold: the full source lifecycle; classification that changes authority rather than metadata; campaign isolation; hybrid retrieval with correct ranking, deduplication and disabled-source exclusion; a bounded knowledge budget; hidden Canon that reaches the narrator without reaching the protagonist; prompt injection framed as data and ignored by the model; no network path but the approved one; no filesystem path at all; export/import into a clean directory; a clean migration from a genuine M6 database; and flat query counts as the library grows.
The five acceptance conditions the implementation flagged as unmeasured — C05, G06, G07, G10 and hidden Canon — were measured here against a real narrator and all five pass. That was the largest open risk and it is closed.
Two blocking items remain, and they are the same defect and the reason it was not caught:
- M7-F1 — retrieval admits imported knowledge unconditionally, because the only effective threshold is a share of the best candidate. Irrelevant Canon, Reference, Inspiration and hidden Canon are injected into every turn.
- M7-F2 — the retrieval-quality suite's stub embedder is more discriminative than the real model, so it cannot detect M7-F1 and cannot validate its fix.
Both are narrow and localised: the ranking is correct, the pipeline is correct, and the fix is to the admission threshold and to the fixture that tests it.
M7 is not accepted. M8 is not authorized. Recommended next step is an M7 corrective pass addressing M7-F1 and M7-F2, with M7-F3, M7-F4 and M7-F5 folded in, followed by re-review of §I, §E (G05-G07) and §G.
X. Repository state at the end of the review
branch m7-imported-knowledge
HEAD a6e9c7a32bdf42f1e4cb837b70721d89422cef8d (unchanged)
staged 48 files (47 from M7 + this report)
unstaged 0
untracked 0
LICENSE unchanged, not staged
prepared commit message .git/M7_MSG, unchanged
Files modified by the review itself: exactly one —
planning/reports/M7-IMPLEMENTATION-REPORT.md, which is this report, newly
created. No application code, no test, no planning document and no configuration
was changed by the review. All review harnesses, fixtures, databases and browser
artifacts were created outside the repository, under the session scratchpad, and
none of them is staged.
M6's report remains at planning/reports/M6-IMPLEMENTATION-REPORT.md; per the
package convention it rotates to planning/archive/milestone-reports/ when M7 is
accepted, which has not happened.
Closeout note, added later: it has now. M6's report is at
planning/archive/milestone-reports/M6-IMPLEMENTATION-REPORT.md. This paragraph is left as written because it recorded the state at review time.*
M7 Corrective Pass — Closeout
Date: 2026-09-06 Scope: the two blocking findings above, M7-F1 and M7-F2, plus the non-blocking findings where they were directly related and low-risk. Nothing above this line has been changed. The original findings, their reproductions and their evidence stand as written.
CA. What was wrong, reproduced on the staged tree
Both findings were reproduced before any code changed, against the real embedding model, so the before/after comparison is like for like.
CA.1 — M7-F1, the no-match case
scene: "Aldric studies the tide tables and the harbour master's ledger
of container tonnage." (nothing in the library relates to it)
BEFORE considered=5 floor=0.2875 semantic=True
USED hidden-key.md canon lex=1.000 sem=1.000 cos=0.548 score=1.150 51 tok
USED canon.md canon lex=0.000 sem=0.772 cos=0.423 score=0.772 56 tok
USED reference.md reference lex=0.895 sem=0.927 cos=0.508 score=0.902 70 tok
USED revival.md reference lex=0.000 sem=0.730 cos=0.400 score=0.621 86 tok
USED inspiration.md inspiration lex=0.000 sem=0.871 cos=0.477 score=0.610 63 tok
-> 5 of 5 sources injected, 326 tokens
-> imported sections in the prompt:
['imported_inspiration', 'imported_reference', 'imported_canon']
Note the raw cosines: 0.400-0.548, i.e. nothing matched. Normalization lifted the best to 1.000 and the relative floor (0.2875) admitted everything.
CA.2 — M7-F1, irrelevant admitted alongside relevant
scene: the abbey crypt library: abbey Canon + container-shipping Canon
BEFORE USED canon.md canon cos=0.697 score=1.150 <- relevant
USED shipping.md canon cos=0.320 score=0.459 <- container vessels
USED reference.md ref cos=0.486 score=0.593
CA.3 — M7-F1, lexical leak on a generic word
scene: "Ballast trim on the ore hauler was recalculated before the burn to Ceres."
library: hidden-key.md ("...Edrin discovered this before he disappeared...")
semantic: OFF
BEFORE terms=[... 'before' ...]
USED hidden-key.md canon lex=1.000 -> 1 injected, 51 tokens
The single shared token was the function word "before", which the 42-word stop list did not contain and which nothing downstream weighed.
CA.4 — M7-F1, lexical leak on the protagonist's name
scene: "Aldric studies the harbour ledger." semantic: OFF
BEFORE terms=['aldric','studies','harbour','ledger']
USED hidden-key.md canon lex=1.000 -> 1 injected, 51 tokens
hidden-key.md contains "Aldric does not know what the key opens". The
protagonist's name is in the story tail of essentially every query, so this
admits on a token that carries no information about the current scene.
CA.5 — M7-F2, why the suite could not see any of it
Measured on identical texts, stub against the production model:
stub cosine normalized | real cosine normalized
abbey-canon (relevant) 0.7559 1.000 | 0.7505 1.000
crypts-ref (relevant) 0.3392 0.449 | 0.6237 0.831
ship-canon (IRRELEVANT) 0.1499 0.198 | 0.4365 0.582
cookery (IRRELEVANT) 0.0818 0.108 | 0.4354 0.580
desert (IRRELEVANT) 0.0642 0.085 | 0.4349 0.580
The old floor was 0.25 of the best. Irrelevant material peaked at 0.198 under the stub and at 0.582 under the real model — excluded in the fixture, admitted in production.
CB. Root cause
Two symptoms, one architectural cause: M7 had a ranking stage and no admission stage.
Both scores were normalized against the best candidate of their own path, and
the floor was max(0.02, best × 0.25). A floor expressed as a share of the best
cannot reject anything, because the best clears a share of itself by
construction. MIN_RELEVANCE = 0.02 could only bind when the best score fell
below 0.08, which no real embedding produces, so it was unreachable — finding
M7-F3 was not a separate defect but the same one seen from the side.
Because the semantic path scores every embedded chunk, there was always a best. With semantic retrieval enabled, at least one passage was admitted on every turn, unconditionally.
The lexical path had the mirror-image problem for a different reason: FTS5 terms
were joined with OR and nothing asked how much had matched, so one incidental
token was sufficient.
CC. The correction
CC.1 Admission separated from ranking
candidate generation
-> ADMISSION absolute, per path, independent of the candidate set
-> RANKING normalized, among the survivors only
-> class weighting
-> budget
Admission reads raw signals; ranking reads normalized ones. Normalization
still exists — bm25 has no fixed range and cosine's zero is not zero, so the
paths are not otherwise comparable — but it now decides order among things that
matched, never whether anything matched. Authority is applied in the ranking
stage only, so a class can order what matched and can never rescue what did not.
Recorded as a durable architectural lesson in TECHNICAL-DESIGN.md §13.2 and
summarised in IMPORTED-KNOWLEDGE-DESIGN.md §76.
CC.2 Semantic admission — classes.SEMANTIC_FLOOR = 0.58
Established empirically through the production path: passages embedded as
fts.index_line(heading, text), queries assembled by retrieval.query_terms,
because both differ from bare strings and both move the numbers. The first
calibration attempt used bare strings and under-measured by ~0.10, which is
recorded here because it is exactly the kind of error that produces a plausible
wrong constant.
113 pairs against nomic-embed-text:
targeted n= 13 min 0.5526 p10 0.6090 median 0.7231 max 0.8474
the one source a scene is actually about
off-topic n=100 min 0.3577 median 0.4591 p95 0.5339 max 0.5578
20 scenes unrelated to the campaign (harbour, surgery, compiler,
fugue, sourdough, kiln, football, photosynthesis, …)
sweep: 0.55 -> 13/13 targeted, 1/100 off-topic
0.56 -> 12/13 targeted, 0/100 off-topic
0.58 -> 12/13 targeted, 0/100 off-topic <- chosen
0.61 -> 11/13 targeted, 0/100 off-topic (loses C05's revival match)
0.58 sits between the two populations with ~0.022 of margin on each side. The single targeted pair below it is instructive rather than lost: "the broken circle cut into the keystone above the crypt stair" scores 0.5526 against the Canon describing exactly that — the wording is so close that little is left for the embedding to add — and it matches four lexical terms, so the lexical path admits it. That is the hybrid working, and it is why neither path has to be right alone.
The value is a property of the model and is documented as such beside the
constant, in the same style as memorybank.REDUNDANT_SIMILARITY. A model that
scored everything below it would degrade the library to lexical-only, which is a
supported production path, so the failure mode is safe.
CC.3 Lexical admission
A passage is lexically admitted when it matches two distinct meaningful query terms, or one term that is both (a) not the name of a standing campaign entity and (b) not a negligible share of the query.
- The stop list grew from 42 to 261 words, all function words and
contentless generics. It contains no subject matter — no "abbey", "key" or
"crypt" — because a stop list that removes subject matter is how a search
stops finding "The Silver Key".
beforeis now filtered. - Standing entities are the protagonist and the campaign's established
entities, drawn from
persona_nameand the authoritative state. They are in the query on every turn by construction, so a lone match on one says nothing about the present scene. This is deliberately not "ignore proper nouns" —Westhaven,broken-circleandresurrectall still count, and a standing entity counts the moment a second term matches beside it. - The share test covers the case the entity test cannot see: a young campaign whose state is still empty. One word out of a nine-word scene is not evidence however distinctive it is; one out of three is.
Both conditions are needed and each was measured failing alone: the share test alone would admit a lone "Aldric" from a three-word query, and the entity test alone admitted it from a nine-word one — which is what CA.4 measured.
Evidence per term comes from fts.term_evidence, one statement with a subquery
per term (capped at 24), scoped once at the join. Stemming is applied by FTS
itself, so resurrected matches resurrection exactly as the ranking query
does rather than through a second, divergent stemmer in Python.
CC.4 M7-F3 folded in
MIN_RELEVANCE and RELEVANCE_FLOOR_SHARE are removed, not re-tuned. The
architecture that made them meaningless is gone, and nothing is left that
appears to enforce relevance without doing so.
CD. Real-model before/after
Same reproductions, same model, same fixtures.
| Case | Before | After |
|---|---|---|
| Off-topic scene (tide tables/tonnage), semantic ON | 5 sources, 326 tokens, 3 imported sections | 0 sources, 0 tokens, no imported section |
| Relevant scene + irrelevant Canon in library | relevant + shipping Canon + Reference | relevant Canon only |
| Lexical leak on "before" | 1 source, 51 tokens | 0 |
| Lexical leak on "Aldric" | 1 source, 51 tokens | 0 |
| Positive control: Old Abbey question | canon.md + shipping.md | canon.md only |
The after-state reports the rejection explicitly rather than silently:
generated=5 rejected=5 considered=0 floor=0.58.
CE. The §12 acceptance matrix — measured against the real model
13/13.
| Query | Library | Expected | Result |
|---|---|---|---|
| Old Abbey / broken-circle | relevant Canon + unrelated | relevant Canon only | PASS — ['canon.md'] |
| tavern description | tavern Reference + unrelated Canon | Reference, not irrelevant Canon | PASS — Reference in, shipping Canon out |
| resurrection question | conflicting Canon and Reference | Canon governs | PASS — campaign rule present above the framed Reference |
| tide tables / container tonnage | fantasy fixture only | zero imported chunks | PASS — used=[] sections=[] generated=4 rejected=4 |
| distinctive exact term | embeddings unavailable | lexical succeeds | PASS — ['sigils.md'], semantic_used=False |
| semantic paraphrase | little lexical overlap | semantic succeeds | PASS — admitted_by semantic, cosine 0.6575 |
| protagonist name only | many Aldric-containing sources | no unrelated flood | PASS — used=[] |
| stopword-heavy input | mixed library | no incidental FTS admission | PASS — used=[], terms=[] |
Plus three beyond the matrix:
- hidden Canon when relevant → retrieved.
- hidden Canon when unrelated → absent (
used=[] generated=1). - always-include Canon → still supplied on an unrelated scene, because it is asserted rather than retrieved and is deliberately outside admission.
For every no-match row the assembled prompt was inspected: the imported-knowledge sections are absent, not empty, and the untrusted-data rule section is absent with them.
CF. The four hybrid cases
Permanent deterministic coverage in test_knowledge_retrieval_quality.py:
A strong semantic / weak lexical ossuary passage, no shared vocabulary
-> retrieved, admitted_by="semantic"
B strong lexical / weak semantic distinctive term, embeddings unavailable
-> retrieved, admitted_by="lexical"
C both strong -> one candidate, admitted_by="both",
not duplicated
D neither strong -> ZERO chunks, and no imported section
in the prompt (mandatory case)
Case D is covered twice — with embeddings on and on the lexical-only path — and
each assertion is guarded by generated > 0, so it cannot pass because nothing
was searched.
CG. M7-F2 — the test doubles
RealisticEmbedder replaces the old stub. It has a deliberate similarity
floor: every pair shares a constant component, so unrelated passages score a
substantial similarity as they do in life, with disjoint topic axes above it and
a hashed bag of words underneath.
Two tests exist purely to keep the fixture honest:
test_the_stub_models_the_real_problemfails if unrelated pairs ever drop near zero again — that is, if the fixture drifts back to the one that hid the defect.test_a_relative_only_floor_would_admit_the_irrelevant_setreproduces the superseded rule on the corrected code's own vectors and asserts it still admits the entire irrelevant library. If this ever stops failing the old rule, the fixture has lost the ability to detect the defect.
The stub is not calibrated to 0.58. It models the shape — a floor under
everything, structure above it — and the tests assert against
classes.SEMANTIC_FLOOR rather than against a number baked into the fixture.
Validation that F2 is actually closed: the corrected implementation was temporarily reverted to the staged pre-corrective version and the new suite run against it:
13 failed, 5 passed
including every no-match test, both case-D tests, all three "irrelevant material is excluded whatever its class" parametrisations, and "Canon is excluded even though it is the best candidate". The suite can now see the defect it previously could not.
CH. Permanent real-model regression
test_knowledge_real_model.py gained four tests, skipped without an endpoint:
test_the_real_model_separates_relevant_from_unrelated
re-measures both populations and fails if SEMANTIC_FLOOR stops sitting
between them. Measured this run:
targeted 0.8130 / 0.7260 / 0.6726
off-topic 0.3293 - 0.5026 (10 pairs, 5 scenes x 2 sources)
floor 0.58
test_a_completely_unrelated_query_retrieves_nothing_from_a_real_model
generated=3 rejected=3 used=0, no imported section
test_a_relevant_query_still_retrieves_from_a_real_model -> ['canon.md']
test_a_paraphrase_still_retrieves_from_a_real_model -> cosine 0.7260
The first of these caught a real error in its own first draft: a single off-topic line appended after a crypt opening still leaves the crypt in the four-action query window, so the query was not off-topic at all. The window is now moved wholesale. Recorded because it is a property of the design worth knowing: relevance is judged against the recent scene including its immediate history, not against the last sentence alone.
CI. No regression in established M7 behaviour
| Area | Result |
|---|---|
| C05, real narrator | PASS — "the dead do not return … magic cannot bring him back to life", with revival.md still retrieved at cosine 0.656, so the test remains non-vacuous |
| G06 / G07 / G10, real narrator | PASS (25/25 with C05 and hidden Canon) |
| Hidden Canon, real narrator | PASS — retrieved, marked narrator-only, not revealed |
| G01-G10 acceptance suite | 49 passed |
| H06 / H07 / H08 / H09 | PASS — 43/43 browser checks, unchanged |
| I05 export/import | 53/54 (the one failure is the review harness's own false positive: text/markdown contains a slash) |
| F05 / F06 | PASS — browser-verified, now also showing matched terms and similarity |
| Source lifecycle, atomicity, isolation, disable/delete | PASS |
| Local-only network boundary | PASS — zero sockets for import/index/retrieval/turn; embeddings only to the configured endpoint; request-time policy enforced |
| Query counts (3 → 39 sources) | list 5, detail 4, retrieve 10, context 21 — flat; retrieval and context each gained exactly one statement (the term-evidence query) |
CJ. Regression gates
| Gate | Result |
|---|---|
| backend suite | 927 passed, 14 skipped, 1 warning (pre-corrective 920/10/1) |
| M7 knowledge suites | 49 + 18 + 14 + 6 + 4 passed |
| real-model M7 suite | 7 passed against nomic-embed-text on the trusted-LAN Ollama |
| frontend lint | exit 0, 7 warnings — identical to the M6 baseline |
| frontend build | success |
| docker build | success |
| browser | 43/43, Firefox 154.0.1 |
| network | unchanged; see CI |
| migration, pre-M7 database via a real server restart | 23/23 |
| export/import into a clean data directory | 53/54 (one review-harness false positive) |
| query counts, 3 → 39 sources | list 5, detail 4, retrieve 10, context 21 — flat |
| §12 acceptance matrix, real model | 13/13 |
| real-narrator authority (C05, G06, G07, G10, hidden) | 25/25 |
The +7 tests are the new admission and real-model tests; the +4 skips are the
new real-model tests, which skip without an endpoint. The single warning is the
pre-existing StarletteDeprecationWarning.
No schema change and no serialized-settings change: the corrective pass altered
the shape of the context-snapshot knowledge record (floor → generated,
rejected, semantic_floor, plus admitted_by and matched_terms per
passage) but added no column and no migration. Snapshots written before the
corrective pass still render — the browser reads the new fields with numeric
comparisons that are false when absent.
CK. Disposition of the non-blocking findings
| Finding | Disposition |
|---|---|
M7-F3 — MIN_RELEVANCE unreachable dead config |
Fixed. Removed along with RELEVANCE_FLOOR_SHARE; the architecture that made them meaningless is gone. Nothing now appears to enforce relevance without doing so. |
| M7-F4 — acceptance tests used the wrong fixture | Fixed. test_imported_knowledge.py now uses TEST-CAMPAIGN-FIXTURE.md §12 verbatim, plus the §11 campaign canon and §8 hidden Canon. G07 exercises the actual trap — the frightened-innkeeper passage against the "Mara is not a spy" rule — and asserts both the framing and that no spy fact reaches authoritative state. The two suites now state which they are: test_imported_knowledge.py is the standard acceptance fixture, test_knowledge_retrieval_quality.py is purpose-built retrieval mechanism. |
| M7-F5 — RTL override in filenames | Fixed. safe_filename drops Unicode format characters (category Cf), so "exe.dm.md" stores as exe.dm.md and a name cannot visually impersonate another extension. Cheap, in-scope, no UI redesign. |
M7-F6 — narrator emits ```json instead of ```state |
Not absorbed. Pre-existing M5 protocol robustness with a small model; M7 neither caused nor worsened it. Ownership stays with M5/M8/M11. Observed again during this pass and left recorded. |
CL. Planning documents updated
Only where the corrective work established a durable fact:
TECHNICAL-DESIGN.md§13.1 — the pipeline diagram now shows the admission stage; §13.2 (new) records the architectural lesson and the general rule that a relevance decision must rest on a signal meaningful on its own.IMPORTED-KNOWLEDGE-DESIGN.md§76 — the retrieval subsection records admission before authority, and states plainly that retrieval may return nothing.BUILD-MILESTONES.mdM7 § Status — both findings, their corrections, and the model-specific calibration recorded as carried debt.
No product requirement was weakened, and no ADR was added: the change is an implementation architecture that the two design documents already govern.
CM. Verdict
READY FOR M7 CLOSEOUT
Both blocking findings are closed with real-model and deterministic evidence:
- M7-F1 is closed. The required invariant — retrieval must be allowed to return zero imported chunks — holds and is measured. Four off-topic scenes that previously injected the entire library now inject nothing, verified at the assembled prompt rather than at the ranking. Relevance is decided before authority, on signals that mean something on their own, and irrelevant Canon is excluded even when it is the only candidate. Every positive control still retrieves: exact terms, paraphrases, lexical-only operation, C05 against a real narrator, and hidden Canon when it is actually relevant.
- M7-F2 is closed. The stub now models the real model's similarity floor, the suite fails 13/18 against the pre-corrective implementation, and two tests exist solely to stop the fixture drifting back.
M7 is not accepted — that is the repository owner's decision on a signed commit. M8 is not started.
CN. Repository state at the end of the corrective pass
branch m7-imported-knowledge
HEAD a6e9c7a32bdf42f1e4cb837b70721d89422cef8d (unchanged)
staged see below
unstaged 0
untracked 0
LICENSE unchanged, not staged
prepared commit message .git/M7_MSG, updated for the corrective pass
M7 Closeout Verification
Date: 2026-09-06
Scope: the two items left open at the end of the corrective pass — the
embedding-model calibration boundary, and the ambiguous export/import 53/54.
Nothing above this line has been changed.
DA. The export/import 53/54, resolved
| Test | E3: no original filesystem path is present anywhere in the bundle |
| Where | the independent review's own harness, not the repository suite |
| Expected | no exported field carries a filesystem path |
| Actual | the assertion fired |
| Verdict | NOT A DEFECT — a false positive in the review harness |
The assertion was a crude substring test: it serialised every non-content
field of every exported source and failed if the result contained a /. The
only slash present is in "mediaType": "text/markdown" — a MIME type, not a
path. No exported filename, title, hash or timestamp contains a separator.
The harness assertion has been replaced with three that test the actual property:
E3 no non-content, non-mediaType field contains "/" or "\" PASS
E3b every exported filename is a bare basename, no ".." PASS
E3c mediaType is one of text/plain, text/markdown PASS
The export/import suite is now 56/56. The earlier 53/54 is withdrawn as
an artefact of the harness; it never described product behaviour. I05 and the
stronger M7 conditions are unaffected and remain demonstrated:
source contents byte-identical PASS
classification survived (all 5 sources) PASS
enabled state survived, disabled stays out PASS
visibility / hidden survived PASS
always-include survived PASS
provenance / hash survived PASS
chunk reconstruction identical text, order and hash PASS
lexical retrieval works on arrival, no reindex PASS
semantic rebuild on demand, local PASS
historical provenance the pre-export turn unaffected PASS
no external pathname required anywhere PASS
DB. The embedding-model calibration boundary, resolved
DB.1 The unsafe case
SEMANTIC_FLOOR = 0.58 is a raw-cosine threshold measured against
nomic-embed-text. The corrective pass documented only the safe half of the
consequence — a model scoring everything lower degrades to lexical-only — and
left the dangerous half as carried debt:
a model whose similarity scale sits above
nomic-embed-text's would put unrelated material past 0.58 and recreate M7-F1 exactly, on a build whose tests all pass.
That is no longer carried debt.
DB.2 The policy
Semantic admission is per model, and a model this build has not measured does not inherit a number measured against something else.
SEMANTIC_CALIBRATION = {"nomic-embed-text": 0.58}
semantic_floor_for(model) -> float | None
- The key is the model's base name, lower-cased and without its tag: an Ollama
tag (
:latest,:v1.5) selects a build of the same model and does not change its similarity scale. Nonemeans "not measured here". Retrieval then does not run the semantic path at all — it is skipped, not attempted and thresholded — and reports why.- Retrieval degrades to lexical-only, which is a first-class production path rather than a fallback, so story play, imports, indexing, inspection and turns are all unaffected.
- No new model-identity mechanism was introduced.
KnowledgeEmbedding.modelalready records what produced each vector and the retrieval catalogue already filters on it; the policy reuses the configuredSettings.embedding_modeland that same name.
The diagnostic names the model, says what is happening and lists what is calibrated:
The embedding model "mxbai-embed-large" has no measured relevance calibration
in this build, so semantic retrieval is disabled and retrieval is lexical only.
Story play and lexical search are unaffected. Calibrated models:
nomic-embed-text.
It is surfaced in three places: semantic_note on the retrieval result, the
knowledge-status endpoint (semantic_calibrated, calibrated_models,
semantic_note), and the Knowledge panel, which shows a warning row rather than
claiming semantic search is on.
knowledge-status now reports semantic_enabled: false for a configured but
uncalibrated model. That is the M6-F5 discipline applied once more: vectors
that exist but are never consulted are not a working semantic index, and
saying "on" because a model name is set would be the same untruth in a new
place.
DB.3 What the policy costs, stated plainly
A semantic-only paraphrase — conceptually right, no shared vocabulary — is not retrieved under an uncalibrated model. That is a missing passage rather than an irrelevant one, and it is asserted in the suite rather than left to be discovered.
Semantic admission is calibrated for
nomic-embed-text; uncalibrated embedding models fall back safely to lexical-only retrieval rather than borrowing its threshold.
Adding a model is a measurement, not a guess: run
tests/test_knowledge_real_model.py against it and confirm the targeted and
off-topic populations separate, as §CC.2 records for the existing entry.
Generic cross-model calibration is deliberately not attempted in v1.
DB.4 Deterministic coverage
tests/test_knowledge_calibration.py, 12 tests, needing no second model
installed — the policy is about model identity, so a configured name and a
stub are the whole apparatus.
The stub is deliberately hostile: GenerousEmbedder scores everything at
about 0.97, including the unrelated. It is the dangerous shape the policy exists
to defend against, and it makes the negative results meaningful — under the
calibrated name the fixture admits an unrelated freighter passage, and the only
thing that changes in the uncalibrated run is the model name.
| Required | Test | Result |
|---|---|---|
1. nomic-embed-text uses the calibrated policy |
test_the_calibrated_model_resolves_to_the_measured_floor (+ tags, whitespace) |
PASS |
| 2. a different model does not inherit 0.58 | test_an_uncalibrated_model_does_not_borrow_the_calibrated_threshold |
PASS |
| 3. unknown config degrades to an explicitly safe state | test_an_uncalibrated_model_degrades_to_lexical_only_with_a_clear_reason, test_the_status_endpoint_reports_the_uncalibrated_state |
PASS |
| 4. distinctive lexical retrieval still works there | test_distinctive_lexical_retrieval_still_works_when_uncalibrated |
PASS |
| 5. a semantic-only paraphrase is not spuriously admitted | test_a_semantic_only_paraphrase_is_not_admitted_when_uncalibrated |
PASS |
| 6. no-match still returns zero chunks | test_no_match_still_returns_zero_chunks_when_uncalibrated / _when_calibrated |
PASS |
| 7. a model change leaves no stale vectors live | test_changing_the_model_does_not_leave_old_vectors_active, test_reindex_rebuilds_vectors_under_the_new_model |
PASS |
On (7) the existing machinery already held: KnowledgeEmbedding.model records
the producing model, the retrieval catalogue and the pending-work query both
filter on it, so after a model change the old vectors are invisible to
retrieval, pending_embeddings counts them as needing rebuild, and the
retrieval result says "no passages have been embedded yet" rather than silently
scoring incompatible vectors. Reindex rebuilds them under the new model. The
tests pin that behaviour rather than adding a second mechanism.
DB.5 The same policy, through the live API
The deterministic tests use stubs to isolate the mechanism. It was also exercised end to end against a real server, a real database and the real narrator, changing only the configured model name between the two halves — 12/12:
embedding_model = "nomic-embed-text" (calibrated)
status semantic_enabled=True semantic_calibrated=True
context semantic_used=True, retrieved ['canon.md']
embedding_model = "mxbai-embed-large" (not calibrated)
status semantic_calibrated=False, semantic_enabled=False
"“mxbai-embed-large” has no measured relevance calibration in this
build, so semantic retrieval is disabled and retrieval is lexical
only. Lexical search and story play are unaffected."
calibrated_models = ['nomic-embed-text']
context semantic_used=False, semantic_floor=0.0
retrieved ['canon.md'] — admitted_by "lexical"
off-topic scene: zero imported chunks, no imported section
a real narrator turn still plays
DB.6 A consequence worth recording
Four M7 suites configured a placeholder embedding model name (embed-test).
Under the new policy that name is uncalibrated, so those suites would have
quietly exercised the lexical-only path while appearing to test semantic
retrieval — the M7-F2 failure mode in a new guise. They now name
nomic-embed-text, which is what their stubs are built to model, with a comment
saying why. test_knowledge_calibration.py is where the uncalibrated path is
exercised deliberately.
DC. Final regression gate
| Gate | Result |
|---|---|
| backend suite | 939 passed, 14 skipped, 1 warning |
| focused M7 knowledge suite | 103 passed |
M7 real-model suite (nomic-embed-text) |
7 passed |
| M7 real-narrator authority tests | 25/25 |
| §12 relevance matrix, real model | 13/13 |
| export/import suite | 56/56 |
| migration suite | 23/23 |
| frontend lint | exit 0, 7 warnings (M6 baseline) |
| frontend production build | success |
| Docker build | success; image verified reporting the calibration registry |
| browser checks | 43/43, Firefox 154.0.1 |
| network | zero sockets for import/index/retrieval/turn; embeddings only to the configured endpoint |
Confirmed once more, after the calibration change:
irrelevant query -> zero imported knowledge PASS
relevant Canon only -> relevant Canon selected PASS
Reference does not become Canon PASS
Inspiration cannot establish conflicting facts PASS
C05 remains non-vacuous (revival.md retrieved at cosine 0.656) PASS
hidden Canon retrieved only when relevant, never disclosed PASS
lexical fallback works PASS
an uncalibrated embedding model cannot recreate M7-F1 PASS
no new outbound network path PASS
query growth bounded (list 5, detail 4, retrieve 10, ctx 21) PASS
DD. Verdict
M7 READY FOR SIGNED COMMIT
Both blocking findings from the independent review are closed with real-model and deterministic evidence, and both closeout items are resolved:
- M7-F1 — retrieval admits nothing without candidate-set-independent evidence, and returns zero chunks on an unrelated scene.
- M7-F2 — the suite models the real model's similarity floor and fails 13/18 against the pre-corrective implementation.
- Calibration boundary — an uncalibrated embedding model degrades safely to lexical-only instead of borrowing a threshold measured against another model.
53/54— a review-harness false positive, withdrawn; the suite is 56/56.
M7 is complete and staged. It is not committed: the repository owner signs milestone commits.
DE. Corrections made during final staging
Three defects were found by the final repository verification itself, after the regression gate in §DC. All three are documentation or commit-message only — no application or test code was touched, so the §DC results stand unrepeated.
-
The prepared commit message carried the corrective pass's test count. It read
927 passed, 14 skippedand91 new across six files; the closeout run measures939 passed, 14 skipped, and the seventh new file (test_knowledge_calibration.py, 12 tests) brings the new total to 110. The numbers reconcile exactly: 927 + 12 = 939, and 836 + 110 − 7 real-model skips = 939 passed with 7 + 7 = 14 skipped. Corrected. -
The required documentation line existed only in this report. The closeout brief specifies the sentence "Semantic admission is calibrated for
nomic-embed-text; uncalibrated embedding models fall back safely rather than borrowing its threshold." §13.3 ofTECHNICAL-DESIGN.mdexplained the policy at length but never stated it in that form, so a reader skimming headings could miss the invariant. It is now the section's opening line. -
BUILD-MILESTONES.mdline 3 still read "In implementation … M7 — next to brief". The M7 section body had been updated toStatus: COMPLETEand M8 authorized, but the document's own top-level status line had not, so the file contradicted itself. Corrected toM1-M7 complete and accepted … M8 next to brief.
The third is the one worth noting as a pattern: a status recorded in two places drifts, and the copy that drifts is the one not being edited during the work.