M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
This commit is contained in:
co-authored by
Claude Opus 5
parent
a6e9c7a32b
commit
480414efe0
+208
-1
@@ -1,13 +1,14 @@
|
||||
from datetime import datetime, timezone
|
||||
|
||||
from sqlalchemy import (
|
||||
JSON, Boolean, DateTime, Float, ForeignKey, Index, Integer, LargeBinary,
|
||||
DDL, JSON, Boolean, DateTime, Float, ForeignKey, Index, Integer, LargeBinary,
|
||||
String, Text, UniqueConstraint, event,
|
||||
)
|
||||
from sqlalchemy.orm import Mapped, Session, mapped_column, relationship
|
||||
|
||||
from .compression import CompressedJSON
|
||||
from .database import Base
|
||||
from .knowledge import fts as knowledge_fts
|
||||
|
||||
|
||||
def utcnow() -> datetime:
|
||||
@@ -222,6 +223,13 @@ class Adventure(Base):
|
||||
cascade="all, delete-orphan",
|
||||
order_by="DerivedStatus.id",
|
||||
)
|
||||
# M7: the imported knowledge library. Campaign-scoped by construction —
|
||||
# there is no path from one campaign's sources to another's.
|
||||
knowledge_sources: Mapped[list["KnowledgeSource"]] = relationship(
|
||||
back_populates="adventure",
|
||||
cascade="all, delete-orphan",
|
||||
order_by="KnowledgeSource.id",
|
||||
)
|
||||
|
||||
|
||||
class Branch(Base):
|
||||
@@ -579,6 +587,205 @@ class DerivedStatus(Base):
|
||||
adventure: Mapped[Adventure] = relationship(back_populates="derived_status")
|
||||
|
||||
|
||||
class KnowledgeSource(Base):
|
||||
"""M7: one local file the reader imported as campaign knowledge.
|
||||
|
||||
A first-class record rather than a Story Card. Phase 0B found Story Cards
|
||||
could not carry what an imported-knowledge system needs — classification,
|
||||
provenance, a content identity, a lifecycle, chunking, or an index — and
|
||||
`IMPORTED-KNOWLEDGE-DESIGN.md` §73 settles that they are not the production
|
||||
store. Nothing here writes a Story Card and nothing reads one.
|
||||
|
||||
Two things about a source are **not** derivable and must survive anything:
|
||||
the accepted content and its classification. Everything else here is either
|
||||
metadata about where it came from or a description of derived work that can
|
||||
be rebuilt (`chunks`, the FTS rows, `KnowledgeEmbedding`).
|
||||
|
||||
## Why the content is in the column
|
||||
|
||||
`IMPORTED-KNOWLEDGE-DESIGN.md` §11 requires the campaign to stop depending
|
||||
on the original file the moment the import succeeds. Two designs satisfy
|
||||
that: copy the bytes into an application-owned directory with the database
|
||||
as metadata authority, or store the text here. This build stores the text.
|
||||
It is the simpler of the two by some distance — one transaction covers the
|
||||
source, its chunks and its index, so a failed import cannot leave a file
|
||||
behind with no row or a row with no file; export carries the content with no
|
||||
second archive format; and there is no directory whose contents can drift
|
||||
away from the rows describing them. Sources are capped at
|
||||
`knowledge.MAX_SOURCE_BYTES`, so the column stays small enough for that to
|
||||
be the right trade.
|
||||
|
||||
`original_filename` is metadata and nothing else. **It is never used as a
|
||||
path.** The import surface is an HTTP upload, so no backend pathname is ever
|
||||
accepted in the first place (H08); see `knowledge/importer.py`.
|
||||
"""
|
||||
|
||||
__tablename__ = "knowledge_sources"
|
||||
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
# Campaign-scoped, and only campaign-scoped: `IMPORTED-KNOWLEDGE-DESIGN.md`
|
||||
# §65-66 make cross-campaign retrieval a defect, not a missing feature.
|
||||
# There is deliberately no branch coordinate. An imported file is campaign
|
||||
# source material; it does not become a different file because the story
|
||||
# forked (`CONTEXT-AND-MEMORY.md` §39). Nothing in M7 derives a knowledge
|
||||
# record from story history, which is the only case that would need one.
|
||||
adventure_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
|
||||
)
|
||||
title: Mapped[str] = mapped_column(String(200), default="")
|
||||
original_filename: Mapped[str] = mapped_column(String(255), default="")
|
||||
# "canon", "reference" or "inspiration". Exactly one, always set, editable
|
||||
# without reimport. This is semantic, not cosmetic: it decides the framing
|
||||
# the chunk is given in the prompt, the weight it carries in ranking, and
|
||||
# which budget it competes in.
|
||||
classification: Mapped[str] = mapped_column(String(20), default="reference")
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, default=True)
|
||||
# "normal" or "hidden". Hidden is narrator-only knowledge — the secret a
|
||||
# mystery turns on. It is not a permission system: the person who imported
|
||||
# the file can always read it here. It means the protagonist does not know
|
||||
# it, and the prompt says so (`IMPORTED-KNOWLEDGE-DESIGN.md` §67-69).
|
||||
visibility: Mapped[str] = mapped_column(String(20), default="normal")
|
||||
# Canon that must be considered whether or not it resembles the query —
|
||||
# "resurrection is impossible" does not stop applying because nobody said
|
||||
# the word (`CONTEXT-AND-MEMORY.md` §41-42). Canon only, and it still costs
|
||||
# measured budget and still appears in provenance.
|
||||
always_include: Mapped[bool] = mapped_column(Boolean, default=False)
|
||||
# SHA-256 of the normalized text. Identity, and the duplicate test.
|
||||
content_hash: Mapped[str] = mapped_column(String(64), default="", index=True)
|
||||
# The accepted source text, exactly as it was decoded. Not the normalized
|
||||
# form: the reader inspects what they imported.
|
||||
content: Mapped[str] = mapped_column(Text, default="")
|
||||
byte_size: Mapped[int] = mapped_column(Integer, default=0)
|
||||
media_type: Mapped[str] = mapped_column(String(80), default="text/plain")
|
||||
# What produced the chunks now on disk, so a later parser change can be
|
||||
# detected rather than guessed at.
|
||||
parser_version: Mapped[int] = mapped_column(Integer, default=1)
|
||||
chunking_version: Mapped[int] = mapped_column(Integer, default=1)
|
||||
# The lexical half: "ready" once chunks and FTS rows are committed,
|
||||
# "failed" if building them raised. A source is retrievable only when this
|
||||
# is "ready", which is what makes a half-built import unreachable rather
|
||||
# than ambiguous (`IMPORTED-KNOWLEDGE-DESIGN.md` §57).
|
||||
index_state: Mapped[str] = mapped_column(String(20), default="pending")
|
||||
index_detail: Mapped[str] = mapped_column(Text, default="")
|
||||
# The semantic half, kept separate on purpose. Lexical retrieval is a
|
||||
# supported production path, not a fallback, so a source whose embeddings
|
||||
# failed still says "lexical available, semantic failed" rather than
|
||||
# reporting one health for both.
|
||||
embed_state: Mapped[str] = mapped_column(String(20), default="idle")
|
||||
embed_detail: Mapped[str] = mapped_column(Text, default="")
|
||||
notes: Mapped[str] = mapped_column(Text, default="")
|
||||
imported_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
|
||||
updated_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow, onupdate=utcnow)
|
||||
|
||||
adventure: Mapped[Adventure] = relationship(back_populates="knowledge_sources")
|
||||
chunks: Mapped[list["KnowledgeChunk"]] = relationship(
|
||||
back_populates="source",
|
||||
cascade="all, delete-orphan",
|
||||
order_by="KnowledgeChunk.chunk_index",
|
||||
)
|
||||
|
||||
|
||||
class KnowledgeChunk(Base):
|
||||
"""M7: one retrievable passage of an imported source.
|
||||
|
||||
Derived data. Deleting every chunk of a source and rebuilding it from
|
||||
`KnowledgeSource.content` must produce the same chunks in the same order —
|
||||
the chunker is deterministic — which is what makes reindexing safe and what
|
||||
lets an export carry the source alone.
|
||||
|
||||
`adventure_id` is denormalized from the source. Retrieval filters by
|
||||
campaign on every query, and carrying the column here means the FTS join
|
||||
reaches the campaign scope without a third table in the hot path.
|
||||
"""
|
||||
|
||||
__tablename__ = "knowledge_chunks"
|
||||
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
source_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("knowledge_sources.id", ondelete="CASCADE"), index=True
|
||||
)
|
||||
adventure_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
|
||||
)
|
||||
chunk_index: Mapped[int] = mapped_column(Integer, default=0)
|
||||
# The Markdown heading trail above this passage, joined with " > ". Empty
|
||||
# for plain text and for a passage above the first heading. It is carried
|
||||
# into the prompt, because "Old Abbey > The Crypt" is most of what tells the
|
||||
# narrator what the passage is about.
|
||||
heading_path: Mapped[str] = mapped_column(Text, default="")
|
||||
text: Mapped[str] = mapped_column(Text, default="")
|
||||
token_count: Mapped[int] = mapped_column(Integer, default=0)
|
||||
content_hash: Mapped[str] = mapped_column(String(64), default="")
|
||||
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
|
||||
|
||||
source: Mapped[KnowledgeSource] = relationship(back_populates="chunks")
|
||||
embedding: Mapped["KnowledgeEmbedding | None"] = relationship(
|
||||
back_populates="chunk", cascade="all, delete-orphan", uselist=False
|
||||
)
|
||||
|
||||
|
||||
# M7: the FTS5 lexical index travels with the table it indexes.
|
||||
#
|
||||
# An FTS5 table is a virtual table, and SQLAlchemy's metadata has no way to
|
||||
# describe one — so left to itself, `create_all` would build every knowledge
|
||||
# table and no index, and `drop_all` would leave the index behind holding
|
||||
# rowids for chunks that no longer exist. Hanging the DDL off
|
||||
# `knowledge_chunks` fixes both ends at once: the index is created with the
|
||||
# table it points at, and dropped before it, on every path that builds or tears
|
||||
# down a schema — a fresh install, an existing database gaining the M7 tables,
|
||||
# and a test's setup and teardown.
|
||||
#
|
||||
# `execute_if(dialect="sqlite")` because FTS5 is SQLite's. This build stores
|
||||
# campaigns in SQLite and nothing else; the Postgres branches elsewhere in the
|
||||
# tree are inherited from upstream and unused (`DEVELOPMENT.md`).
|
||||
event.listen(
|
||||
KnowledgeChunk.__table__,
|
||||
"after_create",
|
||||
DDL(knowledge_fts.DDL).execute_if(dialect="sqlite"),
|
||||
)
|
||||
event.listen(
|
||||
KnowledgeChunk.__table__,
|
||||
"before_drop",
|
||||
DDL(f"DROP TABLE IF EXISTS {knowledge_fts.TABLE}").execute_if(dialect="sqlite"),
|
||||
)
|
||||
|
||||
|
||||
class KnowledgeEmbedding(Base):
|
||||
"""M7: the vector for one chunk, with enough metadata to distrust it.
|
||||
|
||||
A separate table rather than a column on the chunk, for one reason: it makes
|
||||
the rebuildable boundary a table boundary. "Rebuild the semantic index" is
|
||||
`DELETE FROM knowledge_embeddings`, and nothing about the source, its
|
||||
classification or its chunks is in the blast radius.
|
||||
|
||||
`model` and `dimensions` are what make a stale vector detectable rather than
|
||||
silently wrong. `vectors.cosine` already refuses to score two vectors of
|
||||
different lengths, but a same-width vector from a different model would
|
||||
score plausible nonsense, so retrieval checks the model name too.
|
||||
"""
|
||||
|
||||
__tablename__ = "knowledge_embeddings"
|
||||
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
chunk_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("knowledge_chunks.id", ondelete="CASCADE"), unique=True, index=True
|
||||
)
|
||||
adventure_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
|
||||
)
|
||||
# Little-endian float32, the same packing the memory bank uses (vectors.py).
|
||||
vector: Mapped[bytes] = mapped_column(LargeBinary)
|
||||
model: Mapped[str] = mapped_column(String(200), default="")
|
||||
dimensions: Mapped[int] = mapped_column(Integer, default=0)
|
||||
# What the vector was computed against. A parser or chunker change moves the
|
||||
# text under the vector, and these say so without re-reading the chunk.
|
||||
parser_version: Mapped[int] = mapped_column(Integer, default=1)
|
||||
chunking_version: Mapped[int] = mapped_column(Integer, default=1)
|
||||
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
|
||||
|
||||
chunk: Mapped[KnowledgeChunk] = relationship(back_populates="embedding")
|
||||
|
||||
|
||||
class StoryCard(Base):
|
||||
"""Owned by either a scenario or an adventure (exactly one set)."""
|
||||
|
||||
|
||||
Reference in New Issue
Block a user