A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
420 lines
17 KiB
Python
420 lines
17 KiB
Python
"""M7: turning an imported file into retrievable passages, deterministically.
|
|
|
|
Chunking is derived data, and the whole subsystem leans on that being true: an
|
|
export carries the source text alone, an import rebuilds the passages, and
|
|
"reindex" is "throw the chunks away and run this again". None of that is safe
|
|
unless the same bytes always produce the same passages, in the same order, with
|
|
the same identities. So this module is pure, takes no clock and no randomness,
|
|
and every decision it makes is a function of the text.
|
|
|
|
## What it produces
|
|
|
|
A passage carries the Markdown heading trail above it. That is not decoration:
|
|
"Old Abbey > The Crypt" is most of what tells a narrator — and a lexical index —
|
|
what a paragraph is about, and a heading is the one piece of structure a plain
|
|
paragraph split throws away.
|
|
|
|
## Sizing
|
|
|
|
`IMPORTED-KNOWLEDGE-DESIGN.md` §16 sets the initial target at roughly 300-800
|
|
tokens, and the tokenizer here is the one the context builder budgets with, so
|
|
the numbers below mean the same thing at both ends. Paragraphs under one heading
|
|
are packed together until adding the next would cross `TARGET_MAX`; a paragraph
|
|
that alone exceeds `TARGET_MAX` is split on sentence boundaries. Two failure
|
|
modes are guarded explicitly, because `IMPORTED-KNOWLEDGE-DESIGN.md` §15 names
|
|
both of them as what chunking has to avoid:
|
|
|
|
* **No fragments.** A heading with one short line under it would otherwise
|
|
become a chunk of nine tokens, costing an index row and a rerank slot to carry
|
|
almost nothing — and a reference document is mostly such headings. So a
|
|
heading boundary only *closes* a passage once the passage has reached
|
|
`MIN_TOKENS`. Below that the packing runs straight through the boundary and
|
|
writes every heading it crosses — including the one the passage opened under —
|
|
into the text as it goes, so a run of short sections becomes one passage that
|
|
still says which section each part came from. The passage's own `heading_path`
|
|
becomes the deepest trail all its parts share, which for unrelated siblings is
|
|
nothing; the headings themselves are never lost, only moved inside.
|
|
* **No giants.** A 4,000-token section does not become one chunk merely because
|
|
its author wrote no second heading. `TARGET_MAX` is a ceiling on the packing
|
|
loop and `_split_long` is the escape hatch beneath it.
|
|
|
|
## Overlap
|
|
|
|
There is none, and that is a decision rather than an omission. §15 permits
|
|
"limited overlap"; §16 calls it optional. Overlap buys continuity across a
|
|
boundary and costs the same text twice in a bounded budget — and this build has
|
|
a redundancy suppressor sitting downstream whose job is to notice two passages
|
|
saying the same thing, which is exactly what overlap manufactures. The heading
|
|
path gives each passage its context without duplicating any of it. If retrieval
|
|
quality ever argues for overlap, `CHUNKING_VERSION` is how the change is rolled
|
|
out: bump it, and every source is reprocessed and re-embedded on reindex.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import hashlib
|
|
import re
|
|
import unicodedata
|
|
from dataclasses import dataclass, field
|
|
|
|
from ..context import count_tokens
|
|
|
|
# Bumped when this module's output changes for the same input. Stored on the
|
|
# source, the chunk's embedding row, and nothing else needs to guess.
|
|
PARSER_VERSION = 1
|
|
CHUNKING_VERSION = 1
|
|
|
|
# The packing ceiling: adding a paragraph that would take a group past this
|
|
# closes the group instead.
|
|
TARGET_MAX = 800
|
|
# The floor a finished group has to clear before it is allowed to stand alone.
|
|
MIN_TOKENS = 60
|
|
# A single paragraph longer than TARGET_MAX is cut into pieces no larger than
|
|
# this. Slightly under the ceiling so a piece plus its heading line still fits.
|
|
HARD_MAX = 760
|
|
|
|
_ATX_HEADING = re.compile(r"^(#{1,6})\s+(.*?)\s*#*\s*$")
|
|
_FENCE = re.compile(r"^\s{0,3}(`{3,}|~{3,})")
|
|
# Sentence-ish boundaries, for splitting a paragraph that is too long on its
|
|
# own. Deliberately crude: this runs on the rare oversized paragraph, and a
|
|
# clever splitter would be one more thing whose output has to stay stable.
|
|
_SENTENCE_END = re.compile(r"(?<=[.!?])\s+")
|
|
|
|
|
|
@dataclass
|
|
class Passage:
|
|
"""One chunk, before it becomes a row."""
|
|
|
|
index: int
|
|
heading_path: str
|
|
text: str
|
|
token_count: int
|
|
content_hash: str
|
|
|
|
|
|
@dataclass
|
|
class _Block:
|
|
"""A paragraph, with the heading trail that was open above it."""
|
|
|
|
heading_path: str
|
|
text: str
|
|
tokens: int = 0
|
|
|
|
|
|
@dataclass
|
|
class _Group:
|
|
"""A passage under construction.
|
|
|
|
`heading_path` narrows to the common trail as parts from different sections
|
|
are packed in; `last_heading` is what the text most recently declared, so
|
|
the packer knows when to write a new heading line.
|
|
"""
|
|
|
|
heading_path: str
|
|
parts: list[str] = field(default_factory=list)
|
|
tokens: int = 0
|
|
last_heading: str = ""
|
|
#: Whether this passage has already been written across a heading boundary.
|
|
#: It decides whether the opening heading still needs writing into the text.
|
|
mixed: bool = False
|
|
|
|
|
|
def normalize(text: str) -> str:
|
|
"""The canonical form used for hashing, duplicate detection and indexing.
|
|
|
|
`IMPORTED-KNOWLEDGE-DESIGN.md` §61 asks for consistent normalization for
|
|
exactly those three, and for the original to be preserved for display. That
|
|
is what happens: `KnowledgeSource.content` holds the text as decoded, and
|
|
this form is never stored — it is computed where an identity or an index
|
|
entry is needed.
|
|
|
|
NFC, because two spellings of the same accented character are the same word
|
|
to a reader and to a search. Line endings are unified, because a file that
|
|
travelled through Windows is not a different file. Trailing whitespace goes,
|
|
because it is invisible and would otherwise make two identical documents
|
|
hash differently.
|
|
"""
|
|
text = unicodedata.normalize("NFC", text)
|
|
text = text.replace("\r\n", "\n").replace("\r", "\n")
|
|
return "\n".join(line.rstrip() for line in text.split("\n")).strip()
|
|
|
|
|
|
def digest(text: str) -> str:
|
|
"""SHA-256 of the normalized text, as hex. The content identity (§12)."""
|
|
return hashlib.sha256(normalize(text).encode("utf-8")).hexdigest()
|
|
|
|
|
|
def chunk(text: str, *, markdown: bool = True) -> list[Passage]:
|
|
"""Splits a source into passages, deterministically.
|
|
|
|
`markdown` decides only whether `#` lines open a heading and whether fenced
|
|
code is protected from being read as one. Plain text takes the same
|
|
paragraph packing with an empty heading path throughout, which is what §14
|
|
`IMPORTED-KNOWLEDGE-DESIGN.md` §15 asks for — coherent bounded groups of
|
|
paragraphs — rather than a second algorithm.
|
|
"""
|
|
blocks = _blocks(normalize(text), markdown=markdown)
|
|
groups = _pack(blocks)
|
|
passages: list[Passage] = []
|
|
for group in groups:
|
|
body = "\n\n".join(group.parts).strip()
|
|
if not body:
|
|
continue
|
|
passages.append(
|
|
Passage(
|
|
index=len(passages),
|
|
heading_path=group.heading_path,
|
|
text=body,
|
|
token_count=count_tokens(body),
|
|
# The chunk's own identity, over the heading and the body
|
|
# together. Two identical paragraphs under different headings
|
|
# are different passages, because the heading is part of what
|
|
# is retrieved and part of what reaches the prompt.
|
|
content_hash=hashlib.sha256(
|
|
f"{group.heading_path}\n{body}".encode("utf-8")
|
|
).hexdigest(),
|
|
)
|
|
)
|
|
return passages
|
|
|
|
|
|
def _blocks(text: str, *, markdown: bool) -> list[_Block]:
|
|
"""Paragraphs, each tagged with the heading trail open above it."""
|
|
stack: list[tuple[int, str]] = [] # (level, title)
|
|
blocks: list[_Block] = []
|
|
buffer: list[str] = []
|
|
fence: str | None = None
|
|
|
|
def flush() -> None:
|
|
body = "\n".join(buffer).strip()
|
|
buffer.clear()
|
|
if body:
|
|
blocks.append(_Block(_path(stack), body, count_tokens(body)))
|
|
|
|
for line in text.split("\n"):
|
|
if markdown:
|
|
fence_match = _FENCE.match(line)
|
|
if fence_match:
|
|
# A fence toggles. Inside one, `#` is code and `` is not a
|
|
# paragraph break — a code block is one block, whole, because
|
|
# splitting it mid-listing produces two passages neither of
|
|
# which is readable.
|
|
marker = fence_match.group(1)[0]
|
|
if fence is None:
|
|
fence = marker
|
|
elif marker == fence:
|
|
fence = None
|
|
buffer.append(line)
|
|
continue
|
|
if fence is None:
|
|
heading = _ATX_HEADING.match(line)
|
|
if heading is not None:
|
|
flush()
|
|
level = len(heading.group(1))
|
|
title = heading.group(2).strip()
|
|
while stack and stack[-1][0] >= level:
|
|
stack.pop()
|
|
if title:
|
|
stack.append((level, title))
|
|
continue
|
|
if fence is None and not line.strip():
|
|
flush()
|
|
continue
|
|
buffer.append(line)
|
|
flush()
|
|
return blocks
|
|
|
|
|
|
def _path(stack: list[tuple[int, str]]) -> str:
|
|
return " > ".join(title for _level, title in stack)
|
|
|
|
|
|
def _pack(blocks: list[_Block]) -> list[_Group]:
|
|
"""Groups paragraphs into passages, respecting headings and the ceiling.
|
|
|
|
Two rules, and the interaction between them is the whole design:
|
|
|
|
* The ceiling always closes a passage. Nothing packs past `TARGET_MAX`.
|
|
* A heading boundary closes a passage only once it has reached
|
|
`MIN_TOKENS`. A substantial section therefore becomes its own passage
|
|
with its own heading trail, which is what makes "Old Abbey" retrievable;
|
|
a run of one-line sections is packed together instead of becoming a
|
|
handful of unusable fragments.
|
|
|
|
When the packer does run through a boundary it writes the new heading into
|
|
the passage text, so nothing about the document's structure is lost — the
|
|
heading is simply inside the passage rather than beside it — and it narrows
|
|
the passage's own trail to the deepest one its parts share.
|
|
"""
|
|
groups: list[_Group] = []
|
|
current: _Group | None = None
|
|
|
|
for block in blocks:
|
|
pieces = [block] if block.tokens <= TARGET_MAX else _split_long(block)
|
|
for piece in pieces:
|
|
if current is not None:
|
|
changed = piece.heading_path != current.last_heading
|
|
over = current.tokens + piece.tokens > TARGET_MAX
|
|
if over or (changed and current.tokens >= MIN_TOKENS):
|
|
groups.append(current)
|
|
current = None
|
|
if current is None:
|
|
current = _Group(piece.heading_path, last_heading=piece.heading_path)
|
|
elif piece.heading_path != current.last_heading:
|
|
# The passage is about to hold parts from more than one section,
|
|
# so its own trail narrows to what they share — which can be
|
|
# nothing. Before that happens, write the heading this passage
|
|
# *opened* under into the text, or it would be the one heading
|
|
# in the document that survives nowhere: every later one is
|
|
# written in below, and this one is about to stop being the
|
|
# trail. Done once, on the first crossing, guarded by the flag.
|
|
if not current.mixed:
|
|
opening = _heading_line(current.heading_path)
|
|
if opening:
|
|
current.parts.insert(0, opening)
|
|
current.tokens += count_tokens(opening)
|
|
current.mixed = True
|
|
line = _heading_line(piece.heading_path)
|
|
if line:
|
|
current.parts.append(line)
|
|
current.tokens += count_tokens(line)
|
|
current.last_heading = piece.heading_path
|
|
current.heading_path = _common_path(
|
|
current.heading_path, piece.heading_path
|
|
)
|
|
current.parts.append(piece.text)
|
|
current.tokens += piece.tokens
|
|
if current is not None:
|
|
groups.append(current)
|
|
return _absorb_trailing(groups)
|
|
|
|
|
|
def _heading_line(path: str) -> str:
|
|
"""How a heading appears when it is written into a passage rather than beside it."""
|
|
return f"## {path}" if path else ""
|
|
|
|
|
|
def _common_path(a: str, b: str) -> str:
|
|
"""The deepest heading trail both paths share, or an empty string."""
|
|
if a == b:
|
|
return a
|
|
left, right = a.split(" > ") if a else [], b.split(" > ") if b else []
|
|
shared: list[str] = []
|
|
for one, other in zip(left, right):
|
|
if one != other:
|
|
break
|
|
shared.append(one)
|
|
return " > ".join(shared)
|
|
|
|
|
|
def _split_long(block: _Block) -> list[_Block]:
|
|
"""Cuts one oversized paragraph into pieces at sentence boundaries.
|
|
|
|
A sentence longer than the ceiling on its own — a wall of text with no
|
|
punctuation, which is what a pathological import looks like — is cut on
|
|
whitespace, and then, if even that leaves a piece too long, on characters.
|
|
Every branch terminates, which is the property that matters: a source is
|
|
accepted or rejected, never accepted and then chunked forever.
|
|
"""
|
|
pieces: list[_Block] = []
|
|
buffer: list[str] = []
|
|
tokens = 0
|
|
|
|
def flush() -> None:
|
|
nonlocal tokens
|
|
body = " ".join(buffer).strip()
|
|
buffer.clear()
|
|
tokens = 0
|
|
if body:
|
|
pieces.append(_Block(block.heading_path, body, count_tokens(body)))
|
|
|
|
for sentence in _units(block.text):
|
|
cost = count_tokens(sentence)
|
|
if buffer and tokens + cost > HARD_MAX:
|
|
flush()
|
|
buffer.append(sentence)
|
|
tokens += cost
|
|
flush()
|
|
return pieces or [block]
|
|
|
|
|
|
def _units(text: str) -> list[str]:
|
|
"""Sentences, or words, or fixed slices — whichever is small enough."""
|
|
units: list[str] = []
|
|
for sentence in _SENTENCE_END.split(text):
|
|
sentence = sentence.strip()
|
|
if not sentence:
|
|
continue
|
|
if count_tokens(sentence) <= HARD_MAX:
|
|
units.append(sentence)
|
|
continue
|
|
words = sentence.split()
|
|
if len(words) > 1:
|
|
# Rebuild the sentence in word runs that fit. Recursing on the
|
|
# halves would be shorter and would not terminate on a single
|
|
# enormous token.
|
|
run: list[str] = []
|
|
run_tokens = 0
|
|
for word in words:
|
|
cost = count_tokens(word + " ")
|
|
if run and run_tokens + cost > HARD_MAX:
|
|
units.append(" ".join(run))
|
|
run, run_tokens = [], 0
|
|
run.append(word)
|
|
run_tokens += cost
|
|
if run:
|
|
units.append(" ".join(run))
|
|
continue
|
|
# One word longer than the ceiling: a base64 blob, or a language this
|
|
# tokenizer does not segment. Cut it by characters. The slice width is
|
|
# in characters and the ceiling is in tokens, so it is deliberately
|
|
# conservative — a token is at least one character, so this can only
|
|
# undershoot.
|
|
#
|
|
# This is the one branch that does not preserve the text byte for byte:
|
|
# the slices are rejoined with a space, because everything above this
|
|
# point is joining words. Every character survives and the boundary
|
|
# moves. Prose never reaches here — it takes the sentence or the word
|
|
# branch above — so the cost falls only on input that had no word
|
|
# boundaries to respect in the first place.
|
|
units.extend(sentence[i:i + HARD_MAX] for i in range(0, len(sentence), HARD_MAX))
|
|
return units
|
|
|
|
|
|
def _absorb_trailing(groups: list[_Group]) -> list[_Group]:
|
|
"""Folds a final passage too small to stand into the one before it.
|
|
|
|
The packing loop above cannot reach this case: it decides whether to close a
|
|
passage when the *next* piece arrives, and for the last passage there is no
|
|
next piece. So a document ending in a two-line section leaves one fragment,
|
|
and this is where it goes.
|
|
|
|
Only backward, and only when the result still fits. A document that is
|
|
*entirely* short keeps its single passage — a nine-token source is a
|
|
nine-token passage, and there is nothing wrong with that.
|
|
"""
|
|
if len(groups) < 2:
|
|
return groups
|
|
last = groups[-1]
|
|
if last.tokens >= MIN_TOKENS:
|
|
return groups
|
|
previous = groups[-2]
|
|
if previous.tokens + last.tokens > TARGET_MAX:
|
|
return groups
|
|
if last.heading_path != previous.last_heading:
|
|
if not previous.mixed:
|
|
opening = _heading_line(previous.heading_path)
|
|
if opening:
|
|
previous.parts.insert(0, opening)
|
|
previous.tokens += count_tokens(opening)
|
|
previous.mixed = True
|
|
line = _heading_line(last.heading_path)
|
|
if line:
|
|
previous.parts.append(line)
|
|
previous.tokens += count_tokens(line)
|
|
previous.heading_path = _common_path(previous.heading_path, last.heading_path)
|
|
previous.parts += last.parts
|
|
previous.tokens += last.tokens
|
|
previous.last_heading = last.last_heading
|
|
return groups[:-1]
|