"""M7: turning an imported file into retrievable passages, deterministically. Chunking is derived data, and the whole subsystem leans on that being true: an export carries the source text alone, an import rebuilds the passages, and "reindex" is "throw the chunks away and run this again". None of that is safe unless the same bytes always produce the same passages, in the same order, with the same identities. So this module is pure, takes no clock and no randomness, and every decision it makes is a function of the text. ## What it produces A passage carries the Markdown heading trail above it. That is not decoration: "Old Abbey > The Crypt" is most of what tells a narrator — and a lexical index — what a paragraph is about, and a heading is the one piece of structure a plain paragraph split throws away. ## Sizing `IMPORTED-KNOWLEDGE-DESIGN.md` §16 sets the initial target at roughly 300-800 tokens, and the tokenizer here is the one the context builder budgets with, so the numbers below mean the same thing at both ends. Paragraphs under one heading are packed together until adding the next would cross `TARGET_MAX`; a paragraph that alone exceeds `TARGET_MAX` is split on sentence boundaries. Two failure modes are guarded explicitly, because `IMPORTED-KNOWLEDGE-DESIGN.md` §15 names both of them as what chunking has to avoid: * **No fragments.** A heading with one short line under it would otherwise become a chunk of nine tokens, costing an index row and a rerank slot to carry almost nothing — and a reference document is mostly such headings. So a heading boundary only *closes* a passage once the passage has reached `MIN_TOKENS`. Below that the packing runs straight through the boundary and writes every heading it crosses — including the one the passage opened under — into the text as it goes, so a run of short sections becomes one passage that still says which section each part came from. The passage's own `heading_path` becomes the deepest trail all its parts share, which for unrelated siblings is nothing; the headings themselves are never lost, only moved inside. * **No giants.** A 4,000-token section does not become one chunk merely because its author wrote no second heading. `TARGET_MAX` is a ceiling on the packing loop and `_split_long` is the escape hatch beneath it. ## Overlap There is none, and that is a decision rather than an omission. §15 permits "limited overlap"; §16 calls it optional. Overlap buys continuity across a boundary and costs the same text twice in a bounded budget — and this build has a redundancy suppressor sitting downstream whose job is to notice two passages saying the same thing, which is exactly what overlap manufactures. The heading path gives each passage its context without duplicating any of it. If retrieval quality ever argues for overlap, `CHUNKING_VERSION` is how the change is rolled out: bump it, and every source is reprocessed and re-embedded on reindex. """ from __future__ import annotations import hashlib import re import unicodedata from dataclasses import dataclass, field from ..context import count_tokens # Bumped when this module's output changes for the same input. Stored on the # source, the chunk's embedding row, and nothing else needs to guess. PARSER_VERSION = 1 CHUNKING_VERSION = 1 # The packing ceiling: adding a paragraph that would take a group past this # closes the group instead. TARGET_MAX = 800 # The floor a finished group has to clear before it is allowed to stand alone. MIN_TOKENS = 60 # A single paragraph longer than TARGET_MAX is cut into pieces no larger than # this. Slightly under the ceiling so a piece plus its heading line still fits. HARD_MAX = 760 _ATX_HEADING = re.compile(r"^(#{1,6})\s+(.*?)\s*#*\s*$") _FENCE = re.compile(r"^\s{0,3}(`{3,}|~{3,})") # Sentence-ish boundaries, for splitting a paragraph that is too long on its # own. Deliberately crude: this runs on the rare oversized paragraph, and a # clever splitter would be one more thing whose output has to stay stable. _SENTENCE_END = re.compile(r"(?<=[.!?])\s+") @dataclass class Passage: """One chunk, before it becomes a row.""" index: int heading_path: str text: str token_count: int content_hash: str @dataclass class _Block: """A paragraph, with the heading trail that was open above it.""" heading_path: str text: str tokens: int = 0 @dataclass class _Group: """A passage under construction. `heading_path` narrows to the common trail as parts from different sections are packed in; `last_heading` is what the text most recently declared, so the packer knows when to write a new heading line. """ heading_path: str parts: list[str] = field(default_factory=list) tokens: int = 0 last_heading: str = "" #: Whether this passage has already been written across a heading boundary. #: It decides whether the opening heading still needs writing into the text. mixed: bool = False def normalize(text: str) -> str: """The canonical form used for hashing, duplicate detection and indexing. `IMPORTED-KNOWLEDGE-DESIGN.md` §61 asks for consistent normalization for exactly those three, and for the original to be preserved for display. That is what happens: `KnowledgeSource.content` holds the text as decoded, and this form is never stored — it is computed where an identity or an index entry is needed. NFC, because two spellings of the same accented character are the same word to a reader and to a search. Line endings are unified, because a file that travelled through Windows is not a different file. Trailing whitespace goes, because it is invisible and would otherwise make two identical documents hash differently. """ text = unicodedata.normalize("NFC", text) text = text.replace("\r\n", "\n").replace("\r", "\n") return "\n".join(line.rstrip() for line in text.split("\n")).strip() def digest(text: str) -> str: """SHA-256 of the normalized text, as hex. The content identity (§12).""" return hashlib.sha256(normalize(text).encode("utf-8")).hexdigest() def chunk(text: str, *, markdown: bool = True) -> list[Passage]: """Splits a source into passages, deterministically. `markdown` decides only whether `#` lines open a heading and whether fenced code is protected from being read as one. Plain text takes the same paragraph packing with an empty heading path throughout, which is what §14 `IMPORTED-KNOWLEDGE-DESIGN.md` §15 asks for — coherent bounded groups of paragraphs — rather than a second algorithm. """ blocks = _blocks(normalize(text), markdown=markdown) groups = _pack(blocks) passages: list[Passage] = [] for group in groups: body = "\n\n".join(group.parts).strip() if not body: continue passages.append( Passage( index=len(passages), heading_path=group.heading_path, text=body, token_count=count_tokens(body), # The chunk's own identity, over the heading and the body # together. Two identical paragraphs under different headings # are different passages, because the heading is part of what # is retrieved and part of what reaches the prompt. content_hash=hashlib.sha256( f"{group.heading_path}\n{body}".encode("utf-8") ).hexdigest(), ) ) return passages def _blocks(text: str, *, markdown: bool) -> list[_Block]: """Paragraphs, each tagged with the heading trail open above it.""" stack: list[tuple[int, str]] = [] # (level, title) blocks: list[_Block] = [] buffer: list[str] = [] fence: str | None = None def flush() -> None: body = "\n".join(buffer).strip() buffer.clear() if body: blocks.append(_Block(_path(stack), body, count_tokens(body))) for line in text.split("\n"): if markdown: fence_match = _FENCE.match(line) if fence_match: # A fence toggles. Inside one, `#` is code and `` is not a # paragraph break — a code block is one block, whole, because # splitting it mid-listing produces two passages neither of # which is readable. marker = fence_match.group(1)[0] if fence is None: fence = marker elif marker == fence: fence = None buffer.append(line) continue if fence is None: heading = _ATX_HEADING.match(line) if heading is not None: flush() level = len(heading.group(1)) title = heading.group(2).strip() while stack and stack[-1][0] >= level: stack.pop() if title: stack.append((level, title)) continue if fence is None and not line.strip(): flush() continue buffer.append(line) flush() return blocks def _path(stack: list[tuple[int, str]]) -> str: return " > ".join(title for _level, title in stack) def _pack(blocks: list[_Block]) -> list[_Group]: """Groups paragraphs into passages, respecting headings and the ceiling. Two rules, and the interaction between them is the whole design: * The ceiling always closes a passage. Nothing packs past `TARGET_MAX`. * A heading boundary closes a passage only once it has reached `MIN_TOKENS`. A substantial section therefore becomes its own passage with its own heading trail, which is what makes "Old Abbey" retrievable; a run of one-line sections is packed together instead of becoming a handful of unusable fragments. When the packer does run through a boundary it writes the new heading into the passage text, so nothing about the document's structure is lost — the heading is simply inside the passage rather than beside it — and it narrows the passage's own trail to the deepest one its parts share. """ groups: list[_Group] = [] current: _Group | None = None for block in blocks: pieces = [block] if block.tokens <= TARGET_MAX else _split_long(block) for piece in pieces: if current is not None: changed = piece.heading_path != current.last_heading over = current.tokens + piece.tokens > TARGET_MAX if over or (changed and current.tokens >= MIN_TOKENS): groups.append(current) current = None if current is None: current = _Group(piece.heading_path, last_heading=piece.heading_path) elif piece.heading_path != current.last_heading: # The passage is about to hold parts from more than one section, # so its own trail narrows to what they share — which can be # nothing. Before that happens, write the heading this passage # *opened* under into the text, or it would be the one heading # in the document that survives nowhere: every later one is # written in below, and this one is about to stop being the # trail. Done once, on the first crossing, guarded by the flag. if not current.mixed: opening = _heading_line(current.heading_path) if opening: current.parts.insert(0, opening) current.tokens += count_tokens(opening) current.mixed = True line = _heading_line(piece.heading_path) if line: current.parts.append(line) current.tokens += count_tokens(line) current.last_heading = piece.heading_path current.heading_path = _common_path( current.heading_path, piece.heading_path ) current.parts.append(piece.text) current.tokens += piece.tokens if current is not None: groups.append(current) return _absorb_trailing(groups) def _heading_line(path: str) -> str: """How a heading appears when it is written into a passage rather than beside it.""" return f"## {path}" if path else "" def _common_path(a: str, b: str) -> str: """The deepest heading trail both paths share, or an empty string.""" if a == b: return a left, right = a.split(" > ") if a else [], b.split(" > ") if b else [] shared: list[str] = [] for one, other in zip(left, right): if one != other: break shared.append(one) return " > ".join(shared) def _split_long(block: _Block) -> list[_Block]: """Cuts one oversized paragraph into pieces at sentence boundaries. A sentence longer than the ceiling on its own — a wall of text with no punctuation, which is what a pathological import looks like — is cut on whitespace, and then, if even that leaves a piece too long, on characters. Every branch terminates, which is the property that matters: a source is accepted or rejected, never accepted and then chunked forever. """ pieces: list[_Block] = [] buffer: list[str] = [] tokens = 0 def flush() -> None: nonlocal tokens body = " ".join(buffer).strip() buffer.clear() tokens = 0 if body: pieces.append(_Block(block.heading_path, body, count_tokens(body))) for sentence in _units(block.text): cost = count_tokens(sentence) if buffer and tokens + cost > HARD_MAX: flush() buffer.append(sentence) tokens += cost flush() return pieces or [block] def _units(text: str) -> list[str]: """Sentences, or words, or fixed slices — whichever is small enough.""" units: list[str] = [] for sentence in _SENTENCE_END.split(text): sentence = sentence.strip() if not sentence: continue if count_tokens(sentence) <= HARD_MAX: units.append(sentence) continue words = sentence.split() if len(words) > 1: # Rebuild the sentence in word runs that fit. Recursing on the # halves would be shorter and would not terminate on a single # enormous token. run: list[str] = [] run_tokens = 0 for word in words: cost = count_tokens(word + " ") if run and run_tokens + cost > HARD_MAX: units.append(" ".join(run)) run, run_tokens = [], 0 run.append(word) run_tokens += cost if run: units.append(" ".join(run)) continue # One word longer than the ceiling: a base64 blob, or a language this # tokenizer does not segment. Cut it by characters. The slice width is # in characters and the ceiling is in tokens, so it is deliberately # conservative — a token is at least one character, so this can only # undershoot. # # This is the one branch that does not preserve the text byte for byte: # the slices are rejoined with a space, because everything above this # point is joining words. Every character survives and the boundary # moves. Prose never reaches here — it takes the sentence or the word # branch above — so the cost falls only on input that had no word # boundaries to respect in the first place. units.extend(sentence[i:i + HARD_MAX] for i in range(0, len(sentence), HARD_MAX)) return units def _absorb_trailing(groups: list[_Group]) -> list[_Group]: """Folds a final passage too small to stand into the one before it. The packing loop above cannot reach this case: it decides whether to close a passage when the *next* piece arrives, and for the last passage there is no next piece. So a document ending in a two-line section leaves one fragment, and this is where it goes. Only backward, and only when the result still fits. A document that is *entirely* short keeps its single passage — a nine-token source is a nine-token passage, and there is nothing wrong with that. """ if len(groups) < 2: return groups last = groups[-1] if last.tokens >= MIN_TOKENS: return groups previous = groups[-2] if previous.tokens + last.tokens > TARGET_MAX: return groups if last.heading_path != previous.last_heading: if not previous.mixed: opening = _heading_line(previous.heading_path) if opening: previous.parts.insert(0, opening) previous.tokens += count_tokens(opening) previous.mixed = True line = _heading_line(last.heading_path) if line: previous.parts.append(line) previous.tokens += count_tokens(line) previous.heading_path = _common_path(previous.heading_path, last.heading_path) previous.parts += last.parts previous.tokens += last.tokens previous.last_heading = last.last_heading return groups[:-1]