A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
582 lines
25 KiB
Python
582 lines
25 KiB
Python
"""M7: choosing which imported passages a narrator turn should be shown.
|
||
|
||
query terms ──┬──▶ FTS5 lexical candidates ─┐
|
||
│ ├─▶ merge ─▶ dedupe ─▶
|
||
└──▶ semantic candidates ─┘
|
||
(when an embedding model is configured)
|
||
|
||
─▶ authority × relevance rerank ─▶ ranked candidates ─▶ inject.py
|
||
|
||
The cut against the token budget is **not** here. It is in `inject.py`, which is
|
||
the only module that knows what the context builder has left. This module's job
|
||
ends at a ranked, deduplicated, campaign-scoped list with every score on it, so
|
||
that "why did that passage win?" is answerable from the record rather than
|
||
reconstructed.
|
||
|
||
## The query is not the user's sentence
|
||
|
||
`IMPORTED-KNOWLEDGE-DESIGN.md` §27 and `CONTEXT-AND-MEMORY.md` §40 both say so,
|
||
for the same reason: "I open the door" retrieves nothing, and the material that would help is
|
||
about the room the door is in. So the query is assembled from what the
|
||
application already knows is active — the recent story, the current scene and
|
||
location, the entities present, the open threads.
|
||
|
||
Two constraints on where those terms may come from, and they are the same
|
||
constraint twice:
|
||
|
||
* The story terms come from `context.history.tail`, which reads through the
|
||
**head-capped lineage clause**. An Undo followed by a divergence leaves the
|
||
abandoned turns in the database, and they must not reach this query — a
|
||
retrieval influenced by a story the reader walked away from is the M6 leak
|
||
wearing different clothes.
|
||
* The state terms come from `adventure.narrative_state`, which head movement
|
||
repoints at the position being read. Same property, different table.
|
||
|
||
Neither reads the uncapped `actions` table, and nothing here queries by "the
|
||
newest rows".
|
||
|
||
## Admission, then ranking
|
||
|
||
These are two stages and the order is the point.
|
||
|
||
candidate generation
|
||
-> ADMISSION absolute signals, independent of the candidate set
|
||
-> RANKING normalized among the survivors only
|
||
-> class weighting
|
||
-> budget
|
||
|
||
**Admission** asks whether a passage matched *at all*, using signals that mean
|
||
something on their own: the raw cosine the model returned, and how many distinct
|
||
meaningful query terms the passage actually contains. Neither is computed by
|
||
comparison with the other candidates, so a set in which everything is bad
|
||
produces nothing.
|
||
|
||
M7's first implementation had no such stage. It normalized both scores against
|
||
the best of their own path and then applied a floor defined as a *share of the
|
||
best* — which the best candidate clears by construction, every time. With the
|
||
semantic path scoring every embedded chunk there was always a best, so something
|
||
was admitted on every turn regardless of the scene. Review finding M7-F1
|
||
measured the consequence: a query about tide tables and container tonnage
|
||
retrieved all five sources of a fantasy campaign, hidden Canon among them.
|
||
|
||
**Ranking** then runs over the survivors, and only there does normalization
|
||
appear. It is still needed, because `bm25` has no fixed range and cosine's zero
|
||
is not zero, so the two paths cannot be blended raw. But it now decides *order
|
||
among things that matched*, never *whether anything matched*.
|
||
|
||
relevance = max(lexical, semantic) + AGREEMENT × min(lexical, semantic)
|
||
score = relevance × CLASS_WEIGHTS[classification]
|
||
|
||
`max` rather than a weighted sum, because the two paths answer different
|
||
questions and a passage found by only one of them is not thereby worse: an exact
|
||
name match the embedding missed is a good hit, and so is a conceptual match with
|
||
no shared words. The small agreement term breaks ties towards passages both
|
||
paths liked, which is the useful thing a hybrid actually buys.
|
||
|
||
The class multiplies relevance and is applied *after* admission, so authority
|
||
can order what matched and can never rescue what did not. That is what makes
|
||
`IMPORTED-KNOWLEDGE-DESIGN.md` §30's two consecutive sentences —
|
||
`Canon > Reference > Inspiration`, and "do not include irrelevant Canon merely
|
||
because it is authoritative" — both true at once.
|
||
|
||
There is deliberately no model-based reranker. It would be a second inference
|
||
call per turn, and it would be opaque to the inspector — which
|
||
`IMPORTED-KNOWLEDGE-DESIGN.md` §29 rules out in as many words: "keep formula
|
||
simple and inspectable".
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from sqlalchemy import select
|
||
from sqlalchemy.orm import Session, object_session
|
||
|
||
from .. import memorybank, models
|
||
from ..context import history, truncate_to_last_tokens
|
||
from ..providers import ProviderError
|
||
from ..vectors import cosine
|
||
from . import classes, embeddings, fts
|
||
from .records import Candidate, Result # re-exported: callers name these
|
||
|
||
#: How many of the newest actions the query reads. The same window the memory
|
||
#: bank uses, for the same reason: further back is the summary's job.
|
||
QUERY_ACTIONS = 4
|
||
#: A ceiling on the story text that becomes query terms.
|
||
QUERY_TOKENS = 600
|
||
#: Terms taken from the current authoritative state — entity names, the scene,
|
||
#: the location, open threads. Bounded so a campaign with a large cast does not
|
||
#: turn every query into a search for everything.
|
||
STATE_TERMS = 40
|
||
#: The largest number of terms the FTS expression carries.
|
||
MAX_TERMS = 60
|
||
|
||
#: Candidates each path may return before the merge. Both are enforced in the
|
||
#: database, so the Python-side ranking never sees an unbounded set.
|
||
LEXICAL_CANDIDATES = 40
|
||
SEMANTIC_CANDIDATES = 40
|
||
#: The most passages whose vectors are scored in one turn. A campaign larger
|
||
#: than this is ranked over its first N passages by id and the shortfall is
|
||
#: reported on the result, rather than the turn quietly getting slower and
|
||
#: slower. v1 has no approximate-nearest-neighbour index; this is the honest
|
||
#: bound in its place.
|
||
SEMANTIC_SCAN_LIMIT = 4000
|
||
|
||
#: How much agreement between the two paths is worth, when ordering survivors.
|
||
AGREEMENT = 0.15
|
||
|
||
#: How many (term, chunk) evidence rows the admission query may return. Bounded
|
||
#: for the same reason the candidate caps are: nothing about admission may grow
|
||
#: with the size of the library.
|
||
EVIDENCE_ROWS = 2000
|
||
|
||
#: Two passages this close are treated as saying the same thing.
|
||
#:
|
||
#: The value and the reasoning are the memory bank's (`memorybank.py`,
|
||
#: M6 finding M6-F2), measured against the same local embedding model: redundant
|
||
#: pairs scored 0.938-0.996 and genuinely distinct ones 0.349-0.906. The same
|
||
#: measurement ruled out the lexical alternative, which fires hardest on the
|
||
#: pair that must *not* merge — "Mara promised Aldric" against "Aldric promised
|
||
#: Mara" shares most of its words and means the opposite.
|
||
REDUNDANT_SIMILARITY = 0.93
|
||
|
||
|
||
# ------------------------------------------------------------ the query
|
||
|
||
|
||
def query_terms(
|
||
adventure: models.Adventure, *, exclude_action_id: int | None = None
|
||
) -> tuple[list[str], str]:
|
||
"""The search terms for the position the story is being read at.
|
||
|
||
Returns the terms and the raw text they came from — the text is what the
|
||
semantic side embeds, because a bag of words is a poor thing to hand an
|
||
embedding model even when it is the right thing to hand an inverted index.
|
||
"""
|
||
recent = history.tail(adventure, QUERY_ACTIONS, exclude_action_id)
|
||
story = truncate_to_last_tokens("\n\n".join(a.text for a in recent), QUERY_TOKENS)
|
||
state = _state_text(adventure.narrative_state)
|
||
text = "\n".join(part for part in (state, story) if part.strip())
|
||
words = fts.terms(text)[:MAX_TERMS]
|
||
return words, text
|
||
|
||
|
||
def _state_text(state) -> str:
|
||
"""Scene, location, entities and open threads, as searchable words.
|
||
|
||
Read straight off the authoritative document rather than through
|
||
`narrative.render`, whose output is shaped for a model to read and carries
|
||
prose this has no use for. Only the names are wanted here.
|
||
"""
|
||
if not isinstance(state, dict):
|
||
return ""
|
||
pieces: list[str] = []
|
||
scene = state.get("scene")
|
||
if isinstance(scene, dict):
|
||
for key in ("summary", "location"):
|
||
value = scene.get(key)
|
||
if isinstance(value, str) and value.strip():
|
||
pieces.append(value.strip())
|
||
entities = state.get("entities")
|
||
if isinstance(entities, dict):
|
||
for key, entity in list(entities.items())[:STATE_TERMS]:
|
||
pieces.append(str(key))
|
||
if isinstance(entity, dict):
|
||
name = entity.get("name")
|
||
if isinstance(name, str) and name.strip():
|
||
pieces.append(name.strip())
|
||
for alias in (entity.get("aliases") or [])[:3]:
|
||
if isinstance(alias, str) and alias.strip():
|
||
pieces.append(alias.strip())
|
||
threads = state.get("threads")
|
||
if isinstance(threads, dict):
|
||
for key, thread in list(threads.items())[:STATE_TERMS]:
|
||
if isinstance(thread, dict) and thread.get("status") not in (
|
||
"resolved", "abandoned"
|
||
):
|
||
title = thread.get("title")
|
||
pieces.append(str(title) if isinstance(title, str) else str(key))
|
||
return " ".join(pieces)
|
||
|
||
|
||
def standing_entity_terms(adventure: models.Adventure) -> set[str]:
|
||
"""The words that are in the retrieval query on *every* turn.
|
||
|
||
The protagonist's name and the campaign's established entities — their keys,
|
||
names and aliases. The query is built partly from the authoritative state,
|
||
so these are present whatever the scene is, which means a passage that
|
||
matched only one of them has told us nothing about the present moment. That
|
||
is exactly how `hidden-key.md` was admitted into a harbour scene on the word
|
||
"Aldric" (review finding M7-F1).
|
||
|
||
This is **not** "ignore proper nouns". A place name that is not a standing
|
||
entity — `Westhaven`, `broken-circle` — is among the strongest lexical
|
||
signals there is, and a standing entity still counts the moment a second
|
||
term matches alongside it. Only the lone-standing-entity match is refused.
|
||
"""
|
||
words: set[str] = set()
|
||
for value in (adventure.persona_name or "",):
|
||
words.update(fts.terms(value))
|
||
state = adventure.narrative_state
|
||
if isinstance(state, dict):
|
||
entities = state.get("entities")
|
||
if isinstance(entities, dict):
|
||
for key, entity in list(entities.items())[:STATE_TERMS]:
|
||
words.update(fts.terms(str(key)))
|
||
if isinstance(entity, dict):
|
||
words.update(fts.terms(str(entity.get("name") or "")))
|
||
for alias in (entity.get("aliases") or [])[:3]:
|
||
words.update(fts.terms(str(alias)))
|
||
return words
|
||
|
||
|
||
def lexical_admits(
|
||
matched: frozenset[int], words: list[str], standing: set[str]
|
||
) -> bool:
|
||
"""Whether the lexical evidence for one passage is enough to admit it.
|
||
|
||
Two distinct meaningful terms, or one distinctive term — see
|
||
`classes.LEXICAL_MIN_TERMS` and `classes.LEXICAL_SINGLE_TERM_SHARE` for why
|
||
the single-term case needs both a "not a standing entity" test and a share
|
||
test. Common English words never reach here; `fts.terms` removed them.
|
||
"""
|
||
if not words or not matched:
|
||
return False
|
||
if len(matched) >= classes.LEXICAL_MIN_TERMS:
|
||
return True
|
||
(index,) = tuple(matched)
|
||
if not (0 <= index < len(words)):
|
||
return False
|
||
if words[index] in standing:
|
||
return False
|
||
return 1 / len(words) >= classes.LEXICAL_SINGLE_TERM_SHARE
|
||
|
||
|
||
# ------------------------------------------------------------ the retrieval
|
||
|
||
|
||
async def retrieve(
|
||
adventure: models.Adventure,
|
||
settings: models.Settings,
|
||
*,
|
||
exclude_action_id: int | None = None,
|
||
) -> Result:
|
||
"""The ranked passages this campaign's library offers for this position.
|
||
|
||
Never raises for an inference failure. A dead endpoint costs the semantic
|
||
half and is reported on the result; it does not cost the turn.
|
||
"""
|
||
db = object_session(adventure)
|
||
if db is None:
|
||
return Result()
|
||
|
||
always = _always_included(db, adventure.id)
|
||
words, text = query_terms(adventure, exclude_action_id=exclude_action_id)
|
||
result = Result(terms=words)
|
||
|
||
scored: dict[int, Candidate] = {}
|
||
standing = standing_entity_terms(adventure)
|
||
|
||
# ---------------- candidate generation ----------------
|
||
lexical = fts.search(db, adventure.id, words, LEXICAL_CANDIDATES)
|
||
evidence = fts.term_evidence(db, adventure.id, words, EVIDENCE_ROWS)
|
||
|
||
semantic: list[tuple[int, float]] = []
|
||
model = embeddings.model_name(settings)
|
||
floor = classes.semantic_floor_for(model)
|
||
result.embedding_model = model
|
||
result.semantic_calibrated = floor is not None
|
||
result.semantic_floor = floor or 0.0
|
||
if not embeddings.enabled(settings):
|
||
result.semantic_note = (
|
||
"No embedding model is configured, so retrieval is lexical only."
|
||
)
|
||
elif floor is None:
|
||
# The model-aware policy. An admission threshold measured against one
|
||
# embedding model says nothing about another's scale, and borrowing it
|
||
# is how a model that scores unrelated text higher would silently
|
||
# readmit everything. Lexical retrieval is a first-class path, so this
|
||
# costs recall rather than correctness and never costs a turn.
|
||
result.semantic_note = (
|
||
f"The embedding model “{model}” has no measured relevance "
|
||
"calibration in this build, so semantic retrieval is disabled and "
|
||
"retrieval is lexical only. Story play and lexical search are "
|
||
"unaffected. Calibrated models: "
|
||
+ ", ".join(sorted(classes.SEMANTIC_CALIBRATION)) + "."
|
||
)
|
||
elif not text.strip():
|
||
result.semantic_note = "Nothing in the current scene to search on."
|
||
else:
|
||
semantic, note, truncated = await _semantic(db, adventure, settings, text)
|
||
result.semantic_note = note
|
||
result.scan_truncated = truncated
|
||
result.semantic_used = not note
|
||
|
||
# ---------------- ADMISSION ----------------
|
||
#
|
||
# Absolute, per path, and computed before anything is compared with anything
|
||
# else. Each path answers "did this passage match?" on its own terms; a
|
||
# passage is admitted if either says yes. Nothing here consults the class,
|
||
# the other candidates, or the best score — which is the whole correction.
|
||
semantic_raw = dict(semantic)
|
||
lexical_raw = dict(lexical)
|
||
|
||
admitted: dict[int, dict] = {}
|
||
for chunk_id, similarity in semantic:
|
||
# `floor` is None for an uncalibrated model, and `semantic` is then
|
||
# empty, so this loop does not run. The check is written against the
|
||
# resolved floor rather than the module constant so there is exactly one
|
||
# place a threshold can come from.
|
||
if floor is not None and similarity >= floor:
|
||
admitted.setdefault(chunk_id, {})["semantic"] = similarity
|
||
for chunk_id in lexical_raw:
|
||
matched = evidence.get(chunk_id, frozenset())
|
||
if lexical_admits(matched, words, standing):
|
||
admitted.setdefault(chunk_id, {})["lexical"] = matched
|
||
|
||
result.generated = len(set(lexical_raw) | set(semantic_raw))
|
||
result.rejected = result.generated - len(admitted)
|
||
|
||
wanted = set(admitted) | {chunk.id for chunk in always}
|
||
if not wanted:
|
||
# The result this whole stage exists to make reachable: the library was
|
||
# searched, nothing matched, and nothing is supplied.
|
||
return result
|
||
|
||
for chunk_id, candidate in _load(db, adventure.id, sorted(wanted)).items():
|
||
scored[chunk_id] = candidate
|
||
|
||
# ---------------- RANKING, among the survivors only ----------------
|
||
#
|
||
# Normalization returns here, and only here. Both paths are normalized
|
||
# against the best *admitted* value of their own path, because bm25 has no
|
||
# fixed range and cosine's zero is not zero, so the two are not otherwise
|
||
# comparable. This decides order; it no longer decides membership.
|
||
survivors = [c for c in scored if c in admitted]
|
||
lexical_top = max((lexical_raw.get(c, 0.0) for c in survivors), default=0.0)
|
||
semantic_top = max((semantic_raw.get(c, 0.0) for c in survivors), default=0.0)
|
||
|
||
for chunk_id, candidate in scored.items():
|
||
how = admitted.get(chunk_id)
|
||
if how is None:
|
||
continue # an always-included passage
|
||
if "lexical" in how:
|
||
raw = lexical_raw.get(chunk_id, 0.0)
|
||
candidate.lexical = raw / lexical_top if lexical_top else 0.0
|
||
candidate.matched_terms = sorted(
|
||
words[i] for i in how["lexical"] if 0 <= i < len(words)
|
||
)
|
||
if "semantic" in how:
|
||
raw = semantic_raw.get(chunk_id, 0.0)
|
||
candidate.cosine = raw
|
||
candidate.semantic = raw / semantic_top if semantic_top else 0.0
|
||
candidate.admitted_by = (
|
||
"both" if len(how) == 2 else next(iter(how))
|
||
)
|
||
|
||
for chunk in always:
|
||
candidate = scored.get(chunk.id)
|
||
if candidate is not None:
|
||
candidate.always_include = True
|
||
|
||
result.considered = len(scored)
|
||
for candidate in scored.values():
|
||
high, low = max(candidate.lexical, candidate.semantic), min(
|
||
candidate.lexical, candidate.semantic
|
||
)
|
||
candidate.relevance = high + AGREEMENT * low
|
||
candidate.score = candidate.relevance * classes.CLASS_WEIGHTS.get(
|
||
candidate.classification, 1.0
|
||
)
|
||
|
||
ranked = list(scored.values())
|
||
ranked.sort(key=lambda c: (c.always_include, c.score), reverse=True)
|
||
kept, suppressed = _drop_redundant(db, adventure.id, ranked)
|
||
result.candidates = kept
|
||
result.suppressed = suppressed
|
||
return result
|
||
|
||
|
||
def _always_included(db: Session, adventure_id: int) -> list[models.KnowledgeChunk]:
|
||
"""Every passage of every enabled, ready, always-include Canon source."""
|
||
return list(
|
||
db.execute(
|
||
select(models.KnowledgeChunk)
|
||
.join(
|
||
models.KnowledgeSource,
|
||
models.KnowledgeSource.id == models.KnowledgeChunk.source_id,
|
||
)
|
||
.where(
|
||
models.KnowledgeSource.adventure_id == adventure_id,
|
||
models.KnowledgeSource.enabled.is_(True),
|
||
models.KnowledgeSource.index_state == "ready",
|
||
models.KnowledgeSource.always_include.is_(True),
|
||
models.KnowledgeSource.classification == classes.CANON,
|
||
)
|
||
.order_by(models.KnowledgeChunk.source_id, models.KnowledgeChunk.chunk_index)
|
||
).scalars().all()
|
||
)
|
||
|
||
|
||
async def _semantic(
|
||
db: Session,
|
||
adventure: models.Adventure,
|
||
settings: models.Settings,
|
||
text: str,
|
||
) -> tuple[list[tuple[int, float]], str, bool]:
|
||
"""Cosine-ranked passages, or an empty list and the reason there are none."""
|
||
model = embeddings.model_name(settings)
|
||
catalogue = db.execute(
|
||
select(models.KnowledgeEmbedding.chunk_id)
|
||
.join(
|
||
models.KnowledgeChunk,
|
||
models.KnowledgeChunk.id == models.KnowledgeEmbedding.chunk_id,
|
||
)
|
||
.join(
|
||
models.KnowledgeSource,
|
||
models.KnowledgeSource.id == models.KnowledgeChunk.source_id,
|
||
)
|
||
.where(
|
||
models.KnowledgeSource.adventure_id == adventure.id,
|
||
models.KnowledgeSource.enabled.is_(True),
|
||
models.KnowledgeSource.index_state == "ready",
|
||
# A vector from another embedding model would score plausible
|
||
# nonsense against this query. `cosine` catches a width change; it
|
||
# cannot catch a same-width model change, so the model name is the
|
||
# check that matters.
|
||
models.KnowledgeEmbedding.model == model,
|
||
)
|
||
.order_by(models.KnowledgeEmbedding.chunk_id)
|
||
.limit(SEMANTIC_SCAN_LIMIT + 1)
|
||
).scalars().all()
|
||
if not catalogue:
|
||
return [], "No passages have been embedded yet, so retrieval is lexical only.", False
|
||
truncated = len(catalogue) > SEMANTIC_SCAN_LIMIT
|
||
catalogue = list(catalogue[:SEMANTIC_SCAN_LIMIT])
|
||
|
||
try:
|
||
# The shared provider, never a client of this module's own. That is
|
||
# where the endpoint allowlist is re-checked and where the private-CA
|
||
# trust store is honoured (ADR 011).
|
||
[query_vector] = await memorybank.embedding_provider(settings).embed([text])
|
||
except ProviderError as exc:
|
||
return [], f"Semantic retrieval unavailable: {exc}", truncated
|
||
|
||
held = embeddings.vectors_for(db, adventure.id, catalogue)
|
||
ranked = sorted(
|
||
(
|
||
(chunk_id, cosine(query_vector, held[chunk_id]))
|
||
for chunk_id in catalogue
|
||
if chunk_id in held
|
||
),
|
||
key=lambda row: row[1],
|
||
reverse=True,
|
||
)
|
||
# Bounded here, and the bound is applied to the *ranked* list, so the
|
||
# strongest similarities survive to face admission. Anything below the floor
|
||
# would be refused there anyway; cutting first only keeps the set small.
|
||
return ranked[:SEMANTIC_CANDIDATES], "", truncated
|
||
|
||
|
||
def _load(
|
||
db: Session, adventure_id: int, chunk_ids: list[int]
|
||
) -> dict[int, Candidate]:
|
||
"""The passages named, with their source metadata, in one query.
|
||
|
||
One query for the whole candidate set, not one per candidate. The N+1
|
||
discipline M5 restored and M6 kept applies here too, and the join is what
|
||
re-applies campaign scope, enabled state and index state to a set of ids
|
||
that came out of an index rather than out of a scoped read.
|
||
"""
|
||
rows = db.execute(
|
||
select(
|
||
models.KnowledgeChunk.id,
|
||
models.KnowledgeChunk.source_id,
|
||
models.KnowledgeChunk.chunk_index,
|
||
models.KnowledgeChunk.heading_path,
|
||
models.KnowledgeChunk.text,
|
||
models.KnowledgeChunk.token_count,
|
||
models.KnowledgeSource.title,
|
||
models.KnowledgeSource.original_filename,
|
||
models.KnowledgeSource.classification,
|
||
models.KnowledgeSource.visibility,
|
||
models.KnowledgeSource.always_include,
|
||
)
|
||
.join(
|
||
models.KnowledgeSource,
|
||
models.KnowledgeSource.id == models.KnowledgeChunk.source_id,
|
||
)
|
||
.where(
|
||
models.KnowledgeChunk.id.in_(chunk_ids),
|
||
models.KnowledgeSource.adventure_id == adventure_id,
|
||
models.KnowledgeSource.enabled.is_(True),
|
||
models.KnowledgeSource.index_state == "ready",
|
||
)
|
||
).all()
|
||
return {
|
||
row.id: Candidate(
|
||
chunk_id=row.id,
|
||
source_id=row.source_id,
|
||
title=row.title,
|
||
filename=row.original_filename,
|
||
classification=row.classification,
|
||
visibility=row.visibility,
|
||
chunk_index=row.chunk_index,
|
||
heading_path=row.heading_path,
|
||
text=row.text,
|
||
token_count=row.token_count,
|
||
)
|
||
for row in rows
|
||
}
|
||
|
||
|
||
def _drop_redundant(
|
||
db: Session, adventure_id: int, ranked: list[Candidate]
|
||
) -> tuple[list[Candidate], list[Candidate]]:
|
||
"""Sets aside passages that repeat one already kept.
|
||
|
||
**Before** the budget cut, not after — M6's finding M6-F2 was that four
|
||
near-identical entries crowded out the one that mattered, and suppression
|
||
that runs after the cut cannot give the freed slot to anything.
|
||
|
||
Two rules, both inherited from that finding and both load-bearing:
|
||
|
||
* **Class is never crossed.** A Reference passage may not suppress a Canon
|
||
one, or the reverse. They are different kinds of claim even when they
|
||
read alike, and collapsing across them erases exactly the distinction this
|
||
subsystem exists to keep.
|
||
* **Wording is not evidence.** Suppression needs vectors. Without them the
|
||
only thing suppressed is an exact repetition of the same passage text,
|
||
which is a fact rather than a judgement. Word-overlap merging was measured
|
||
wrong for this in M6 and is not used here either.
|
||
"""
|
||
kept: list[Candidate] = []
|
||
suppressed: list[Candidate] = []
|
||
held = embeddings.vectors_for(
|
||
db, adventure_id, [c.chunk_id for c in ranked]
|
||
)
|
||
seen_text: dict[tuple[str, str], int] = {}
|
||
for candidate in ranked:
|
||
duplicate_of = None
|
||
identity = (candidate.classification, candidate.text.strip())
|
||
if identity in seen_text:
|
||
duplicate_of = seen_text[identity]
|
||
else:
|
||
vector = held.get(candidate.chunk_id)
|
||
if vector is not None:
|
||
for other in kept:
|
||
if other.classification != candidate.classification:
|
||
continue
|
||
other_vector = held.get(other.chunk_id)
|
||
if (
|
||
other_vector is not None
|
||
and cosine(vector, other_vector) >= REDUNDANT_SIMILARITY
|
||
):
|
||
duplicate_of = other.chunk_id
|
||
break
|
||
if duplicate_of is None:
|
||
seen_text.setdefault(identity, candidate.chunk_id)
|
||
kept.append(candidate)
|
||
else:
|
||
candidate.duplicate_of = duplicate_of
|
||
suppressed.append(candidate)
|
||
return kept, suppressed
|