A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
377 lines
16 KiB
Python
377 lines
16 KiB
Python
"""M7: accepting a local file into a campaign's knowledge library.
|
|
|
|
One function does the whole job — validate, hash, store, chunk, index — and it
|
|
does it inside one transaction, because the alternative is the state
|
|
`IMPORTED-KNOWLEDGE-DESIGN.md` §57 forbids: a source presented as usable while
|
|
only half its passages exist.
|
|
|
|
## The transactional boundary
|
|
|
|
validate -> no row is written at all; the caller gets a 4xx and the
|
|
reader's file is untouched
|
|
build -> source row, every chunk row, every FTS row, and
|
|
index_state='ready' all commit together, or none of them do
|
|
|
|
`index_state` is the belt to that braces. Retrieval reads only sources marked
|
|
`ready`, so even a hypothetical partial commit could not be retrieved from — it
|
|
would be a stored source that never answers a query, which is inert rather than
|
|
wrong. A failure after validation leaves `failed` with the reason on the row.
|
|
|
|
Embeddings are deliberately *outside* that boundary. They need a network call to
|
|
Ollama, and a knowledge library that cannot be imported while the inference host
|
|
is down would be a worse product than one whose semantic index lags. So the
|
|
import commits lexically complete and the vectors are filled in afterwards, by
|
|
`embeddings.py`, at import time and again after any later turn.
|
|
|
|
## Path safety
|
|
|
|
There is none to get wrong, and that is the design. The only import surface is
|
|
an HTTP upload: the router takes `UploadFile`, and this module takes bytes and a
|
|
filename *string*. No caller anywhere accepts a server-side pathname, so there
|
|
is no path to canonicalize, no root to compare against, and no symlink to
|
|
resolve. `H08` is satisfied by the absence of the mechanism rather than by a
|
|
check that could later be bypassed — and `safe_filename` below still strips
|
|
every separator and traversal segment, because the name is displayed and stored
|
|
and a `../../etc/passwd` in a title is at best confusing.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import unicodedata
|
|
|
|
from sqlalchemy import select
|
|
from sqlalchemy.orm import Session
|
|
|
|
from .. import models
|
|
from . import chunking, classes, fts
|
|
|
|
# ---------------------------------------------------------------- the limits
|
|
#
|
|
# Every one of these is enforced here, on the server, and each raises a message
|
|
# that says what to do. Nothing is silently truncated: a source is accepted
|
|
# whole or refused with a reason (`SECURITY-THREAT-MODEL.md` §20-21,
|
|
# `IMPORTED-KNOWLEDGE-DESIGN.md` §59-60).
|
|
|
|
#: The largest file accepted, in bytes. One mebibyte of prose is roughly a
|
|
#: 150,000-word book — far past any setting bible — and it sits comfortably
|
|
#: under `limits.MAX_BODY_BYTES` (2 MiB), which the multipart request as a whole
|
|
#: still has to fit inside. Raising this past that ceiling would produce a
|
|
#: confusing 413 from the middleware instead of the message below.
|
|
MAX_SOURCE_BYTES = 1024 * 1024
|
|
|
|
#: The most passages one source may produce. At the chunker's floor of 60 tokens
|
|
#: a megabyte cannot reach this, so in practice it is a guard against a future
|
|
#: chunker change rather than against a user, and it fails loudly if one is ever
|
|
#: made that fragments badly.
|
|
MAX_CHUNKS_PER_SOURCE = 4000
|
|
|
|
#: The most sources one campaign may hold. Bounds the retrieval scan and the
|
|
#: export bundle.
|
|
MAX_SOURCES_PER_ADVENTURE = 200
|
|
|
|
ALLOWED_EXTENSIONS = (".txt", ".md")
|
|
MEDIA_TYPES = {".txt": "text/plain", ".md": "text/markdown"}
|
|
|
|
#: Control characters that no text file legitimately contains. Tab, newline and
|
|
#: carriage return are excluded because they plainly do. A file carrying any of
|
|
#: these is binary that happened to decode, and it is refused.
|
|
_BINARY_CONTROLS = frozenset(
|
|
chr(c) for c in list(range(0, 9)) + [11, 12] + list(range(14, 32)) + [127]
|
|
)
|
|
|
|
|
|
class ImportError_(ValueError):
|
|
"""A file that cannot be accepted, with the reason a reader needs.
|
|
|
|
Named with a trailing underscore so it cannot be confused with the builtin
|
|
of the same name, which means something else entirely.
|
|
"""
|
|
|
|
def __init__(self, message: str, *, conflict: dict | None = None):
|
|
super().__init__(message)
|
|
#: Set when the refusal is a duplicate rather than a fault, so the
|
|
#: router can answer 409 and name the source already holding the
|
|
#: content instead of a flat "rejected".
|
|
self.conflict = conflict
|
|
|
|
|
|
# ------------------------------------------------------------- validation
|
|
|
|
|
|
DEFAULT_FILENAME = "imported.txt"
|
|
|
|
|
|
def safe_filename(name: str) -> str:
|
|
"""The displayable basename of an uploaded filename.
|
|
|
|
A *metadata* cleaner, not a path check — nothing downstream opens anything,
|
|
so there is no path here for a check to protect. What this protects is the
|
|
stored string: a name that reads as a path, carries a traversal segment, or
|
|
smuggles a NUL or a newline into a list screen would be confusing at best
|
|
and misleading at worst.
|
|
|
|
The rule is "take the basename", because that is what an uploaded filename
|
|
*is*. Everything before the last separator described a directory on the
|
|
sender's machine, which this one does not have and will never look for, so
|
|
`../../../../etc/passwd.md` stores as `passwd.md`. Leading dots then go, so
|
|
a stored name can never be `..`, `.` or a hidden file.
|
|
"""
|
|
name = unicodedata.normalize("NFC", name or "").replace("\x00", "")
|
|
for separator in ("\\", "/"):
|
|
name = name.rsplit(separator, 1)[-1]
|
|
# Drop Unicode format characters (category Cf), which are invisible and
|
|
# include the bidirectional overrides. `U+202E` before "exe.dm.md" renders
|
|
# as "dm.exe" in most UIs, so a name could otherwise lie about its own
|
|
# extension on the screen it is displayed on (review finding M7-F5). They
|
|
# carry no information in a filename, so removing them costs nothing.
|
|
name = "".join(c for c in name if unicodedata.category(c) != "Cf")
|
|
name = " ".join(name.split()).lstrip(". ")
|
|
return (name or DEFAULT_FILENAME)[:255]
|
|
|
|
|
|
def extension_of(filename: str) -> str:
|
|
lowered = safe_filename(filename).lower()
|
|
for extension in ALLOWED_EXTENSIONS:
|
|
if lowered.endswith(extension):
|
|
return extension
|
|
return ""
|
|
|
|
|
|
def decode(raw: bytes, filename: str) -> str:
|
|
"""Bytes to text, or a refusal that says which rule was broken.
|
|
|
|
Three checks, in the order a wrong file is most likely to fail them:
|
|
|
|
* **Size**, first, so a huge file is refused before it is decoded.
|
|
* **Encoding**, strictly UTF-8. `SECURITY-THREAT-MODEL.md` §21 and
|
|
`IMPORTED-KNOWLEDGE-DESIGN.md` §60 both ask for a clear rejection over a
|
|
silent mangling, so there is no `errors="replace"` here and no charset
|
|
guessing. A UTF-8 BOM is accepted and stripped, because Windows editors
|
|
write one and it is not a different encoding.
|
|
* **Content**, because an extension is not evidence. §21: "do not trust file
|
|
extensions alone... verify readable text content, reject obvious binary
|
|
data." A NUL byte or a scattering of C0 controls is what a `.txt`-renamed
|
|
binary looks like after it fails to be anything else.
|
|
"""
|
|
if len(raw) > MAX_SOURCE_BYTES:
|
|
raise ImportError_(
|
|
f"“{safe_filename(filename)}” is "
|
|
f"{len(raw) / 1024 / 1024:.1f} MB. The limit for one knowledge "
|
|
f"source is {MAX_SOURCE_BYTES // 1024 // 1024} MB — split the file "
|
|
"and import the parts, so nothing is silently left out."
|
|
)
|
|
if not raw.strip():
|
|
raise ImportError_(f"“{safe_filename(filename)}” is empty.")
|
|
if raw.startswith(b"\xef\xbb\xbf"):
|
|
raw = raw[3:]
|
|
try:
|
|
text = raw.decode("utf-8")
|
|
except UnicodeDecodeError as exc:
|
|
raise ImportError_(
|
|
f"“{safe_filename(filename)}” is not valid UTF-8 text (byte "
|
|
f"{exc.start} is not part of a valid character). Save it as UTF-8 "
|
|
"and import it again — the file has not been changed."
|
|
) from None
|
|
controls = sum(1 for character in text if character in _BINARY_CONTROLS)
|
|
if controls:
|
|
raise ImportError_(
|
|
f"“{safe_filename(filename)}” contains {controls} control "
|
|
"character(s) that do not belong in a text file. It looks like "
|
|
"binary data rather than text, and only .txt and .md are supported."
|
|
)
|
|
return text
|
|
|
|
|
|
def validate(
|
|
raw: bytes,
|
|
filename: str,
|
|
classification: str,
|
|
visibility: str = classes.NORMAL,
|
|
) -> tuple[str, str, str]:
|
|
"""Everything checked before a row is written. Returns (text, extension, title)."""
|
|
extension = extension_of(filename)
|
|
if not extension:
|
|
raise ImportError_(
|
|
f"“{safe_filename(filename)}” is not a supported file type. This "
|
|
"version imports .txt and .md files."
|
|
)
|
|
if not classes.is_class(classification):
|
|
raise ImportError_(
|
|
f"“{classification}” is not a knowledge class. Choose Canon, "
|
|
"Reference or Inspiration."
|
|
)
|
|
if not classes.is_visibility(visibility):
|
|
raise ImportError_(f"“{visibility}” is not a visibility.")
|
|
text = decode(raw, filename)
|
|
clean = safe_filename(filename)
|
|
return text, extension, clean[: -len(extension)] or clean
|
|
|
|
|
|
# ----------------------------------------------------------------- importing
|
|
|
|
|
|
def import_source(
|
|
db: Session,
|
|
adventure: models.Adventure,
|
|
*,
|
|
raw: bytes,
|
|
filename: str,
|
|
classification: str,
|
|
title: str = "",
|
|
visibility: str = classes.NORMAL,
|
|
always_include: bool = False,
|
|
allow_duplicate: bool = False,
|
|
) -> models.KnowledgeSource:
|
|
"""Validates, stores, chunks and indexes one file. All of it, or none of it.
|
|
|
|
The caller commits. Nothing here commits or rolls back, so an exception
|
|
leaves the session dirty and the router's error path discards it — which is
|
|
what makes "no active partial source, no half-built FTS rows, no half-valid
|
|
chunk set" true by construction rather than by cleanup.
|
|
"""
|
|
text, extension, derived_title = validate(raw, filename, classification, visibility)
|
|
clean_name = safe_filename(filename)
|
|
|
|
existing = db.execute(
|
|
select(models.KnowledgeSource).where(
|
|
models.KnowledgeSource.adventure_id == adventure.id
|
|
).limit(MAX_SOURCES_PER_ADVENTURE + 1)
|
|
).scalars().all()
|
|
if len(existing) >= MAX_SOURCES_PER_ADVENTURE:
|
|
raise ImportError_(
|
|
f"This campaign already holds {len(existing)} knowledge sources, "
|
|
f"which is the limit of {MAX_SOURCES_PER_ADVENTURE}. Delete one to "
|
|
"make room."
|
|
)
|
|
|
|
# Duplicate detection, over the normalized text, within this campaign only.
|
|
# §13 forbids silently creating a second copy and indexing it twice; it does
|
|
# not forbid the reader deciding they want one anyway, which is what
|
|
# `allow_duplicate` is. A deliberately simple v1 model: no versioning UI, no
|
|
# supersession chain, and the refusal names the source that already holds
|
|
# the content so the choice is an informed one.
|
|
content_hash = chunking.digest(text)
|
|
if not allow_duplicate:
|
|
twin = next((s for s in existing if s.content_hash == content_hash), None)
|
|
if twin is not None:
|
|
raise ImportError_(
|
|
f"This campaign already holds identical content, imported as "
|
|
f"“{twin.title}”. Import it again only if you want a second "
|
|
"copy with its own classification.",
|
|
conflict={
|
|
"source_id": twin.id,
|
|
"title": twin.title,
|
|
"classification": twin.classification,
|
|
"content_hash": content_hash,
|
|
},
|
|
)
|
|
|
|
if classification != classes.CANON:
|
|
# Always-include is a Canon-only mechanism (`IMPORTED-KNOWLEDGE-DESIGN.md`
|
|
# §32, `CONTEXT-AND-MEMORY.md` §41-42). The reason is that the flag
|
|
# bypasses relevance entirely: asserting unranked Reference on every
|
|
# turn would spend a protected budget on material that establishes
|
|
# nothing.
|
|
always_include = False
|
|
|
|
source = models.KnowledgeSource(
|
|
adventure_id=adventure.id,
|
|
title=(title.strip() or derived_title)[:200],
|
|
original_filename=clean_name,
|
|
classification=classification,
|
|
visibility=visibility,
|
|
always_include=always_include,
|
|
enabled=True,
|
|
content=text,
|
|
content_hash=content_hash,
|
|
byte_size=len(raw),
|
|
media_type=MEDIA_TYPES[extension],
|
|
parser_version=chunking.PARSER_VERSION,
|
|
chunking_version=chunking.CHUNKING_VERSION,
|
|
index_state="pending",
|
|
)
|
|
db.add(source)
|
|
db.flush() # the chunks need the source's id
|
|
build_index(db, source, markdown=extension == ".md")
|
|
return source
|
|
|
|
|
|
def build_index(
|
|
db: Session, source: models.KnowledgeSource, *, markdown: bool | None = None
|
|
) -> int:
|
|
"""(Re)builds one source's passages and its lexical index. Returns the count.
|
|
|
|
This is both half of an import and the whole of a lexical reindex, which is
|
|
the point: there is one code path that turns content into passages, so a
|
|
reindexed source is byte-identical to a freshly imported one. It leaves the
|
|
source `ready` or raises, and it does not touch the source's content,
|
|
classification, visibility or enabled state.
|
|
"""
|
|
if markdown is None:
|
|
markdown = source.media_type == "text/markdown"
|
|
clear_index(db, source)
|
|
passages = chunking.chunk(source.content, markdown=markdown)
|
|
if len(passages) > MAX_CHUNKS_PER_SOURCE:
|
|
raise ImportError_(
|
|
f"“{source.original_filename}” splits into {len(passages)} "
|
|
f"passages, past the limit of {MAX_CHUNKS_PER_SOURCE}."
|
|
)
|
|
for passage in passages:
|
|
chunk_row = models.KnowledgeChunk(
|
|
source_id=source.id,
|
|
adventure_id=source.adventure_id,
|
|
chunk_index=passage.index,
|
|
heading_path=passage.heading_path,
|
|
text=passage.text,
|
|
token_count=passage.token_count,
|
|
content_hash=passage.content_hash,
|
|
)
|
|
db.add(chunk_row)
|
|
db.flush() # the FTS rowid is the chunk's primary key
|
|
fts.add(db, chunk_row.id, passage.heading_path, passage.text)
|
|
source.parser_version = chunking.PARSER_VERSION
|
|
source.chunking_version = chunking.CHUNKING_VERSION
|
|
source.index_state = "ready"
|
|
source.index_detail = ""
|
|
return len(passages)
|
|
|
|
|
|
def clear_index(db: Session, source: models.KnowledgeSource) -> None:
|
|
"""Removes a source's passages, its FTS rows and its vectors.
|
|
|
|
The FTS rows go first, by id, while the ids still exist. Deleting the chunk
|
|
rows first would leave the index holding rowids that point at nothing, and
|
|
a search would then return chunk ids that no longer resolve.
|
|
"""
|
|
chunk_ids = list(
|
|
db.execute(
|
|
select(models.KnowledgeChunk.id).where(
|
|
models.KnowledgeChunk.source_id == source.id
|
|
)
|
|
).scalars().all()
|
|
)
|
|
if not chunk_ids:
|
|
return
|
|
fts.remove_chunks(db, chunk_ids)
|
|
db.query(models.KnowledgeEmbedding).filter(
|
|
models.KnowledgeEmbedding.chunk_id.in_(chunk_ids)
|
|
).delete(synchronize_session=False)
|
|
db.query(models.KnowledgeChunk).filter(
|
|
models.KnowledgeChunk.source_id == source.id
|
|
).delete(synchronize_session=False)
|
|
db.expire(source, ["chunks"])
|
|
|
|
|
|
def delete_source(db: Session, source: models.KnowledgeSource) -> None:
|
|
"""Removes a source and everything derived from it.
|
|
|
|
What it does **not** remove is the evidence of what old narrator turns were
|
|
given. That lives in each turn's own context snapshot as rendered text, not
|
|
as a reference to a live chunk row, so deleting a source cannot turn a
|
|
historical prompt into a set of dangling ids
|
|
(`IMPORTED-KNOWLEDGE-DESIGN.md` §49-50, `DATA-MODEL.md` §25). Story history,
|
|
the head and the authoritative state are untouched.
|
|
"""
|
|
clear_index(db, source)
|
|
db.delete(source)
|