M7: a first-class imported knowledge library

A campaign can import local .txt and .md files as Canon, Reference or
Inspiration, and the class is load-bearing rather than a label: it decides the
words a passage is framed with in the prompt, the weight it carries when
passages are ranked, and which budget it competes in when the context is tight.

This is a separate subsystem, which is the Phase 0B decision
(IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification,
provenance, content identity, chunking, an index or a lifecycle, and they were
not promoted into something that does. Nothing here reads or writes one.

The subsystem, in backend/app/knowledge/:

  classes      the three classes, their weights, and the prompt framing
  chunking     deterministic, heading-aware, 60-800 tokens, no overlap
  fts          SQLite FTS5 with porter stemming; scoped and bounded in SQL
  importer     validate, hash, store, chunk, index — in one transaction
  embeddings   local Ollama vectors through the shared provider
  retrieval    query construction, hybrid merge, rerank
  inject       the budgeted cut and the rendered prompt sections

Relevance admission is a separate stage from ranking, and that separation is
the milestone's most expensive lesson. An independent review found the first
implementation deciding relevance with a floor expressed as a share of the best
candidate — which the best clears by construction — so a passage was admitted on
every turn regardless of the scene. A query about tide tables and container
tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden
Canon among them.

So the pipeline is now:

  candidate generation -> admission -> ranking -> class weighting -> budget

Admission reads raw, candidate-set-independent signals: the cosine the model
returned, and how many distinct meaningful query terms a passage contains.
Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero
is not zero. Normalization decides order among things that matched; it can never
decide whether anything matched. Authority is applied after admission, so a
class orders what matched and never rescues what did not.

Retrieval may therefore return nothing, and on a scene unrelated to the library
it does.

The other decisions that each replaced an obvious wrong one:

- The class multiplies relevance rather than adding to it. An additive bonus
  satisfies "Canon outranks Reference" and makes "do not include irrelevant
  Canon" impossible, because a large enough constant wins on its own.
- The semantic floor is measured, not guessed: 113 production-path pairs against
  nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at
  0.36-0.56, and 0.58 sits between them. Because it is a property of that model
  and not of cosine similarity, it is keyed to the model rather than applied to
  whatever is configured: an embedding model with no measured calibration in
  this build does not borrow the number. Semantic admission is skipped, the
  campaign retrieves lexically, and the reason is stated in the knowledge status
  and in the turn's provenance. Degrading to lexical keeps the library usable;
  lending the threshold to an unmeasured model is how the admitted-everything
  defect would return.
- One lexical term is not evidence. Two distinct meaningful terms, or one that
  is neither a standing campaign entity nor a negligible share of the query.
  The stop list grew from 42 words to 261, all function words — no subject
  matter, because a stop list that removes subject matter stops finding "The
  Silver Key".
- Lexical retrieval is a production path, not a fallback. It finds the proper
  nouns and invented terms a setting bible is made of, and the library is fully
  usable with no embedding model configured.

Safety is structural rather than filtered. Imported text reaches the prompt
whole, inside a section that says what it is, under a rule stating the authority
order in words and refusing every instruction inside it. No endpoint accepts a
filesystem path, so H08 has no mechanism to escape from. Nothing renders
imported content as HTML, so a script tag is five visible characters and a
remote image is never fetched. Import, chunking, indexing, retrieval and a turn
open no socket at all; only embeddings do, through the endpoint allowlist the
memory bank already uses.

Provenance is the rendered text, not a foreign key: deleting a source cannot
turn a historical turn's evidence into dangling ids.

Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5
virtual table attached to knowledge_chunks as a DDL hook so it is created and
dropped with the table it indexes. Migration 92. A pre-M7 database opens
unchanged and needs no sources to play.

Bundle: the source content and the reader's judgements about it travel; the
passages, index rows and vectors are rebuilt on import, so a restored campaign
is searchable immediately without a reindex step.

One runtime dependency: python-multipart, Starlette's multipart parser. It is
what makes the upload surface possible, and the upload surface is why no
pathname is ever accepted.

The test doubles were the reason the defect shipped, so they were corrected too.
The retrieval stub scored unrelated text at 0.06-0.20 where the real model
scores it at 0.43-0.44, and its docstring said it had deliberately removed the
constant component that "would put a similarity floor under every pair" — which
is exactly the property real models have. The stub now has that floor, one test
fails if it is ever removed, and another reproduces the superseded rule and
asserts it is still fooled by the same fixture. Run against the pre-corrective
implementation, the new suite fails 13 of 18.

Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of
which mocks nothing between itself and Ollama and re-measures the similarity
separation on every run. 43/43 checks in a real Firefox, reproduced.
Docker build clean.

Four other defects found by review or by the browser run were fixed here rather
than carried: an unreachable relevance constant that appeared to enforce
something and did not; acceptance tests using the wrong fixture files, so G07's
trap was never exercised; a bidirectional override surviving into displayed
filenames; and, from the implementation pass, the Insights panel showing M5's
two state sections as raw keys and the source inspector refetching on every
keystroke.

M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK
REQUIRED. Both blocking findings are closed, and closeout resolved the
embedding-model calibration boundary the corrective pass had left as debt.
planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective
closeout and the closeout verification in sequence, none overwriting another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
This commit is contained in:
JesseMarkowitz
2026-09-06 15:40:13 -04:00
co-authored by Claude Opus 5
parent a6e9c7a32b
commit 480414efe0
52 changed files with 10894 additions and 52 deletions
+286
View File
@@ -0,0 +1,286 @@
"""M7: local vectors for imported passages, and what happens when there are none.
The semantic half of retrieval. It uses the **existing** provider — the same
`OpenAICompatibleProvider` the memory bank builds through
`memorybank.embedding_provider` — and that is not a convenience. That path is
where the endpoint allowlist is re-checked before every request, where the
OS/private-CA trust store is unioned into verification, and where timeouts and
error shapes are decided (ADR 011, `endpoints.py`, `tlstrust.py`). A second HTTP
client here would be a second policy, and the one thing a local-only product
cannot afford is two answers to "where may this connect".
## Failure is normal and must be visible
Ollama is not running; the embedding model is not pulled; the LAN host is
asleep. None of these may cost the reader their import. So:
the source stays — content and classification are
not derived from anything
lexical retrieval keeps working — FTS5 is local SQLite and never
touched the network
the failure is recorded on the source — `embed_state`, `embed_detail`
and on the campaign — `derived_status`, kind "knowledge"
a retry fixes it — the next turn, or Reindex
The campaign-level record reuses M6's `derived.py` rather than inventing a
second status system. The per-source
columns exist alongside it because "which file failed" is not a question a
per-campaign row can answer, and it is the question a reader actually has.
`derived.KNOWLEDGE` is its own kind rather than folded into `derived.EMBEDDING`.
The memory bank's embeddings and the knowledge library's embeddings fail
independently and are fixed by different actions, and M6's finding M6-F5 —
reporting `ok` for work that never ran — is the same mistake as reporting one
health for two subsystems.
"""
from __future__ import annotations
import logging
from sqlalchemy import select
from sqlalchemy.orm import Session
from .. import derived, memorybank, models, vectors
from ..providers import ProviderError
from . import fts
log = logging.getLogger(__name__)
#: Passages per embedding request. Matches the memory bank's batch size; the
#: endpoint is the same one.
MAX_BATCH = 32
#: How many passages one pass will embed. A first import of a large library
#: would otherwise hold a turn's background task open for a long time; the
#: remainder is picked up by the next pass, and `pending_count` says how many
#: are left, so the state is legible rather than merely eventual.
MAX_PER_RUN = 512
def model_name(settings: models.Settings) -> str:
return (settings.embedding_model or "").strip()
def enabled(settings: models.Settings) -> bool:
"""Whether semantic retrieval is configured at all.
No embedding model is not a failure — it is a supported configuration in
which retrieval is lexical. Reporting it as a failure would be M6-F5 again
in the other direction: an alarm about a thing nobody asked for.
"""
return bool(model_name(settings))
def pending_chunks(
db: Session, adventure_id: int, model: str, limit: int
) -> list[models.KnowledgeChunk]:
"""Passages of enabled, ready sources that have no current vector.
"Current" means a vector from *this* embedding model at *this* parser and
chunking version. A model change invalidates every vector, which is why the
comparison is on the row's own metadata rather than on its presence.
"""
return list(
db.execute(
select(models.KnowledgeChunk)
.join(
models.KnowledgeSource,
models.KnowledgeSource.id == models.KnowledgeChunk.source_id,
)
.outerjoin(
models.KnowledgeEmbedding,
models.KnowledgeEmbedding.chunk_id == models.KnowledgeChunk.id,
)
.where(
models.KnowledgeChunk.adventure_id == adventure_id,
models.KnowledgeSource.enabled.is_(True),
models.KnowledgeSource.index_state == "ready",
(models.KnowledgeEmbedding.id.is_(None))
| (models.KnowledgeEmbedding.model != model),
)
.order_by(models.KnowledgeChunk.id)
.limit(limit)
).scalars().all()
)
def pending_count(db: Session, adventure_id: int, model: str) -> int:
"""How many passages are still waiting for a vector."""
return len(pending_chunks(db, adventure_id, model, MAX_PER_RUN + 1))
async def embed_pending(
db: Session, adventure: models.Adventure, settings: models.Settings
) -> int:
"""Embeds what is missing. Returns how many vectors were written.
Records its own outcome on every source it touched and on the campaign, and
never raises: an embedding failure is not allowed to reach the turn that
scheduled it.
"""
model = model_name(settings)
if not model:
derived.succeeded(db, adventure.id, derived.KNOWLEDGE, did_work=False)
return 0
chunks = pending_chunks(db, adventure.id, model, MAX_PER_RUN)
if not chunks:
derived.succeeded(db, adventure.id, derived.KNOWLEDGE, did_work=False)
_settle_sources(db, adventure.id, model)
return 0
provider = memorybank.embedding_provider(settings)
written = 0
try:
for start in range(0, len(chunks), MAX_BATCH):
batch = chunks[start:start + MAX_BATCH]
payload = [fts.index_line(c.heading_path, c.text) for c in batch]
produced = await provider.embed(payload)
for chunk_row, vector in zip(batch, produced):
_store(db, chunk_row, vector, model)
written += 1
except ProviderError as exc:
# Soft failure, loudly recorded. The chunks keep no vector, so the next
# pass retries exactly them; the sources keep their content and their
# lexical index, so the library still answers queries.
derived.failed(db, adventure.id, derived.KNOWLEDGE, exc)
_mark_sources(db, {c.source_id for c in chunks}, "failed", str(exc))
return written
except Exception as exc: # pragma: no cover - defensive
derived.failed(db, adventure.id, derived.KNOWLEDGE, exc)
_mark_sources(db, {c.source_id for c in chunks}, "failed", str(exc))
return written
derived.succeeded(db, adventure.id, derived.KNOWLEDGE, did_work=written > 0)
_settle_sources(db, adventure.id, model)
return written
def _store(
db: Session, chunk_row: models.KnowledgeChunk, vector: list[float], model: str
) -> None:
"""Writes or replaces one passage's vector, with the metadata to date it."""
row = db.execute(
select(models.KnowledgeEmbedding).where(
models.KnowledgeEmbedding.chunk_id == chunk_row.id
)
).scalars().first()
if row is None:
row = models.KnowledgeEmbedding(
chunk_id=chunk_row.id, adventure_id=chunk_row.adventure_id
)
db.add(row)
row.vector = vectors.pack(vector)
row.model = model
row.dimensions = len(vector)
row.parser_version = chunk_row.source.parser_version if chunk_row.source else 1
row.chunking_version = chunk_row.source.chunking_version if chunk_row.source else 1
row.created_at = models.utcnow()
forget_cached(chunk_row.adventure_id)
def _mark_sources(db: Session, source_ids: set[int], state: str, detail: str) -> None:
if not source_ids:
return
db.query(models.KnowledgeSource).filter(
models.KnowledgeSource.id.in_(source_ids)
).update(
{"embed_state": state, "embed_detail": detail[:2000]},
synchronize_session=False,
)
def _settle_sources(db: Session, adventure_id: int, model: str) -> None:
"""Marks each source `ok` or `pending` according to what it actually holds.
Run after a successful pass so a source that was failing and has now been
embedded stops saying so. A source with passages still waiting reports
`pending` rather than `ok`, because `MAX_PER_RUN` can leave a large library
part-way through and "ok" would be untrue.
The flush is load-bearing. This session does not autoflush, so the rows
`_store` just added are still pending in it, and the query below would not
see them — every source would report `pending` immediately after being
embedded, which is exactly the misleading status M6-F5 was about.
"""
db.flush()
outstanding = {
chunk.source_id
for chunk in pending_chunks(db, adventure_id, model, MAX_PER_RUN + 1)
}
sources = db.execute(
select(models.KnowledgeSource).where(
models.KnowledgeSource.adventure_id == adventure_id
)
).scalars().all()
for source in sources:
if not source.enabled or source.index_state != "ready":
continue
if source.id in outstanding:
source.embed_state = "pending"
source.embed_detail = ""
else:
source.embed_state = "ok"
source.embed_detail = ""
def clear_vectors(db: Session, adventure_id: int) -> int:
"""Drops every vector in one campaign, so the next pass rebuilds them.
This is the semantic half of Reindex. It touches no source, no passage, no
story row, which is what `IMPORTED-KNOWLEDGE-DESIGN.md` §55 requires of a
reindex — and it is the reason `KnowledgeEmbedding` is a table of its own.
"""
removed = db.query(models.KnowledgeEmbedding).filter(
models.KnowledgeEmbedding.adventure_id == adventure_id
).delete(synchronize_session=False)
db.query(models.KnowledgeSource).filter(
models.KnowledgeSource.adventure_id == adventure_id
).update({"embed_state": "idle", "embed_detail": ""}, synchronize_session=False)
forget_cached(adventure_id)
return removed or 0
# ---------------------------------------------------------- the vector cache
#
# The same idea as the memory bank's, and for the same measured reason: turns
# for one campaign arrive one after another, the library changes rarely between
# them, and re-reading every vector on every turn is the largest read a turn
# makes. `array("f")` holds four bytes a component, matching the column.
#
# Correctness rests on one rule: **every write to a vector calls
# `forget_cached`.** There are three of them and they are all in this module.
# Reads reconcile against the catalogue they were given, so a deletion needs no
# invalidation at all — a chunk that is no longer listed is dropped from the
# cache on the next read.
_cache: dict[int, dict[int, object]] = {}
CACHE_ADVENTURES = 8
def forget_cached(adventure_id: int) -> None:
_cache.pop(adventure_id, None)
def vectors_for(
db: Session, adventure_id: int, chunk_ids: list[int]
) -> dict[int, object]:
"""The vectors for `chunk_ids`, reading only the ones not already held."""
held = _cache.get(adventure_id)
if held is None:
while len(_cache) >= CACHE_ADVENTURES:
_cache.pop(next(iter(_cache)))
held = _cache[adventure_id] = {}
wanted = set(chunk_ids)
for gone in set(held) - wanted:
del held[gone]
missing = [chunk_id for chunk_id in chunk_ids if chunk_id not in held]
if missing:
rows = db.execute(
select(models.KnowledgeEmbedding.chunk_id, models.KnowledgeEmbedding.vector)
.where(models.KnowledgeEmbedding.chunk_id.in_(missing))
).all()
for chunk_id, blob in rows:
if blob:
held[chunk_id] = vectors.unpack(blob)
return held