M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
This commit is contained in:
co-authored by
Claude Opus 5
parent
a6e9c7a32b
commit
480414efe0
@@ -1,6 +1,6 @@
|
||||
# Adventure Storyteller — Production Build Milestones
|
||||
|
||||
**Status:** In implementation. M1-M6 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6: 2026-09-06, each of the last two after an independent review and a corrective pass); M7 — First-Class Imported Knowledge Library — next to brief
|
||||
**Status:** In implementation. M1-M7 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6 and M7: 2026-09-06, each of the last three after an independent review and a corrective pass); M8 — Finished v1 browser experience — next to brief
|
||||
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
||||
|
||||
## 1. Purpose
|
||||
@@ -285,7 +285,7 @@ geckodriver, exercised M3's controls in the rendered application: Undo enabled
|
||||
and Redo disabled at the tip, two Undos moving the transcript back, Redo becoming
|
||||
enabled and returning the original tip exactly, Retry and the take pager, and a
|
||||
divergent write retiring Redo with no stale old-future text on screen. It passed.
|
||||
Evidence: `planning/reports/M4-IMPLEMENTATION-REPORT.md` §W.7.
|
||||
Evidence: `planning/archive/milestone-reports/M4-IMPLEMENTATION-REPORT.md` §W.7.
|
||||
|
||||
**Debt carried forward, none of it blocking M4:** full narrator-edit state
|
||||
re-evaluation is deferred to M5 (`STORY-BRANCH-SEMANTICS.md` §14A records the
|
||||
@@ -357,7 +357,7 @@ divergence in a place where the two paths would silently disagree about what
|
||||
|
||||
## Status: COMPLETE
|
||||
|
||||
Accepted 2026-09-03. Evidence: `planning/reports/M4-IMPLEMENTATION-REPORT.md`,
|
||||
Accepted 2026-09-03. Evidence: `planning/archive/milestone-reports/M4-IMPLEMENTATION-REPORT.md`,
|
||||
including its §W closeout addendum. The Definition of Done is met, and — for the
|
||||
first time in this project — **verified in a real browser**.
|
||||
|
||||
@@ -604,7 +604,7 @@ refusal is gone for narrator turns and remains only for a player's own input
|
||||
## M5 — Outcome
|
||||
|
||||
**Complete and accepted, 2026-09-04**, after an independent implementation
|
||||
review (`planning/reports/M5-IMPLEMENTATION-REPORT.md`) and the corrective pass
|
||||
review (`planning/archive/milestone-reports/M5-IMPLEMENTATION-REPORT.md`) and the corrective pass
|
||||
recorded in that report's addendum.
|
||||
|
||||
Delivered:
|
||||
@@ -704,7 +704,7 @@ Evidence: `planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md` §A.1
|
||||
**Complete and accepted 2026-09-06**, after an independent review that found E03
|
||||
still failing and a corrective pass that fixed it. Report, including the review
|
||||
findings and the corrective addendum:
|
||||
`planning/reports/M6-IMPLEMENTATION-REPORT.md`.
|
||||
`planning/archive/milestone-reports/M6-IMPLEMENTATION-REPORT.md`.
|
||||
|
||||
**M7 is authorized**: the corrected M6 evidence passes, including E03 end to end
|
||||
against a real local summariser.
|
||||
@@ -822,6 +822,120 @@ Implement the separate local knowledge subsystem required by the specification r
|
||||
|
||||
A campaign can import local Canon/Reference/Inspiration files, retrieve them locally with provenance, and maintain authority boundaries.
|
||||
|
||||
## Status: COMPLETE
|
||||
|
||||
**Accepted 2026-09-06**, after an independent review that returned *PASS WITH
|
||||
CORRECTIVE WORK REQUIRED*, a corrective pass that closed both blocking findings,
|
||||
and a closeout verification that resolved the calibration boundary the
|
||||
corrective pass had left as debt. Report, including the original findings, the
|
||||
corrective closeout and the closeout verification, all preserved in sequence:
|
||||
`reports/M7-IMPLEMENTATION-REPORT.md`.
|
||||
|
||||
**M8 is authorized.**
|
||||
|
||||
**Capabilities M7 delivered, which later milestones inherit rather than build:**
|
||||
|
||||
- a **first-class imported knowledge library** — `.txt`/`.md` import, Canon /
|
||||
Reference / Inspiration classification that decides prompt framing, ranking
|
||||
weight and budget rather than labelling a list, enable/disable/delete, content
|
||||
hashing, deterministic heading-aware chunking, SQLite FTS5, local Ollama
|
||||
embeddings, hybrid retrieval, campaign isolation and a source inspector;
|
||||
- **relevance admission separated from ranking**, so retrieval can return
|
||||
nothing (§13.2 of `TECHNICAL-DESIGN.md`);
|
||||
- **prompt provenance that survives its source** — the rendered text travels in
|
||||
the turn's snapshot, so deleting a source cannot orphan a historical prompt;
|
||||
- **imported text framed as untrusted data** with the authority order stated in
|
||||
words, verified against a real narrator;
|
||||
- **an import surface that accepts no filesystem path at all**, so H08 is
|
||||
satisfied by the absence of the mechanism.
|
||||
|
||||
The two blocking findings the review raised, both closed:
|
||||
|
||||
- **M7-F1 — retrieval had no effective no-match gate.** Relevance was decided by
|
||||
a floor expressed as a share of the best candidate, which the best clears by
|
||||
construction, so a passage was admitted on every turn regardless of the scene.
|
||||
A query about tide tables retrieved all five sources of a fantasy campaign,
|
||||
hidden Canon among them. Corrected by separating **relevance admission** from
|
||||
**ranking**: admission now uses raw, candidate-set-independent signals, and
|
||||
retrieval may return nothing. `TECHNICAL-DESIGN.md` §13.2 records the lesson.
|
||||
- **M7-F2 — the retrieval suite could not detect F1.** Its stub embedder scored
|
||||
unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, so
|
||||
the broken gate passed. Corrected with a stub that has a deliberate similarity
|
||||
floor, plus a test that fails if the floor is ever removed and one that
|
||||
demonstrates the superseded rule still being fooled by the same fixture.
|
||||
|
||||
What was built, beyond the scope list above:
|
||||
|
||||
- **The classification is load-bearing, not a label.** It decides the framing a
|
||||
passage is given in the prompt, the weight it carries in ranking, and which
|
||||
budget it competes in. `classes.py` is the single place all three read.
|
||||
- **Lexical retrieval is a production path.** SQLite FTS5 with porter stemming,
|
||||
campaign- and enabled-scoped in SQL, bounded by a `LIMIT` before any Python
|
||||
ranking runs. The library is fully usable with no embedding model configured,
|
||||
and a dead inference host costs the semantic half and nothing else.
|
||||
- **Relevance admission is a separate stage from ranking** (added by the
|
||||
corrective pass). Admission reads raw signals — the cosine the model returned,
|
||||
and how many distinct meaningful query terms a passage contains — so it can
|
||||
answer "nothing matched". Ranking reads normalized ones, because `bm25` has no
|
||||
fixed range and a real embedding model scores any two pieces of English around
|
||||
0.3-0.6. The semantic floor is measured against the production model and
|
||||
recorded beside the constant.
|
||||
- **The class multiplies relevance rather than adding to it**, which is what
|
||||
makes "relevant Canon outranks equally relevant Reference" and "irrelevant
|
||||
Canon does not win on class alone" both true.
|
||||
- **Provenance is the rendered text, not a foreign key.** A turn's knowledge
|
||||
record carries what the narrator was actually shown, so deleting a source
|
||||
cannot turn historical evidence into dangling ids.
|
||||
- **The import surface accepts no pathname at all**, so H08 is satisfied by the
|
||||
absence of the mechanism rather than by a check that could be bypassed.
|
||||
|
||||
Debt carried forward, deliberately:
|
||||
|
||||
- **Entity linking, tags, manual priority and scene pinning** are not
|
||||
implemented (`IMPORTED-KNOWLEDGE-DESIGN.md` §33-36). Entity and place names do
|
||||
reach the retrieval query, because it is built partly from the authoritative
|
||||
state, but there is no explicit source-to-entity link.
|
||||
- **Conflict detection between two Canon sources** (§42) is not implemented.
|
||||
Two Canon sources that disagree are both retrieved and both framed as Canon.
|
||||
- **Source versioning** (§14, §43) is not implemented. A duplicate is refused
|
||||
with a conflict, or imported deliberately as a second source; there is no
|
||||
supersession chain.
|
||||
- **Canon scope metadata** — `invariant / initial / descriptive / historical`
|
||||
(§45) — is not implemented. Current-state precedence stands in its place,
|
||||
which §45 itself permits for v1.
|
||||
- **The semantic scan is linear** over the campaign's vectors, capped at 4,000
|
||||
passages, with the shortfall reported rather than hidden. There is no
|
||||
approximate-nearest-neighbour index in v1.
|
||||
- **The semantic admission floor is calibrated for one embedding model, and the
|
||||
product now knows that.** `nomic-embed-text` was measured over 113
|
||||
production-path pairs. An uncalibrated model does **not** inherit the number:
|
||||
semantic retrieval is skipped for it and the library degrades to lexical-only
|
||||
with the reason reported (`TECHNICAL-DESIGN.md` §13.3). What remains open is
|
||||
only the *enhancement* — calibrating further models, each a measurement rather
|
||||
than a guess. The cost meanwhile is a conceptual-only paraphrase going
|
||||
unretrieved under an uncalibrated model, which is a missing passage rather
|
||||
than an irrelevant one.
|
||||
- **M8:** the knowledge panel and the Insights knowledge rows are functional,
|
||||
not designed. Two labels missing from the Insights section table since M5 were
|
||||
added while M7 was in that file; the rest of the panel's design is M8's. M8
|
||||
should also consider how a "nothing was relevant enough" result and an
|
||||
uncalibrated-model warning should look — both are surfaced plainly today.
|
||||
|
||||
Two defects were found by the browser run and fixed in this pass rather than
|
||||
carried:
|
||||
|
||||
- The Insights panel rendered `state_rule` and `state_reminder` as raw keys,
|
||||
because M5's two sections were never added to the label table.
|
||||
- The open source inspector refetched the source on every render of the Play
|
||||
screen, which re-renders on every keystroke in the story box — 38 needless
|
||||
requests for a 38-character sentence. The reporter callback was in the
|
||||
effect's dependencies and arrives as a fresh function each render. The
|
||||
browser suite gained a check that types and counts requests; it was verified
|
||||
to fail against the unfixed code before the fix was kept.
|
||||
- **M9:** the bundle carries knowledge sources but still carries no context
|
||||
snapshots, so an imported campaign has no historical prompt provenance for any
|
||||
component — which is what a pre-M7 bundle already did for every other one.
|
||||
|
||||
---
|
||||
|
||||
# M8 — Browser UX Completion for v1 Story Operations
|
||||
|
||||
Reference in New Issue
Block a user