480414efe082a4bfe0600a19fe23961f6bddd925
17
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
480414efe0 |
M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b |
||
|
|
a6e9c7a32b |
M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against
|
||
|
|
b7005e6fdd |
M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.
This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:
visible active transcript position == stored head == authoritative state
Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)
A narrator edit no longer rewrites a row. It returns to the state before the
turn, takes the reader's exact text as the accepted narration, re-derives the
state that text implies, and becomes a new active continuation — while the
original narration keeps its words, its live flag and its whole future as
retained history. At the tip the correction is another take; with story below
it, it forks. No new history machinery: this is the existing fork/take/head
path with the reader's text in place of a generated reply. The §14A refusal
is therefore gone for narrator turns, and remains only for player input.
Pre-M5 positions
Migration 88 backfills the empty narrative document onto every action written
before M5, and a missing snapshot now restores the empty document instead of
leaving the previous position's state standing. Restoring to an old Save
Point no longer leaves a later position's entities and facts on screen.
Narrator context
Replayed history carries prose only; the machine-readable block is no longer
reconstructed into past turns, where it contradicted the authoritative state
in the same prompt. A fact withdrawn by a manual correction is now named as
no longer true, with the reader's reason, rather than silently dropped.
Also
- state_changes joins the action-list bulk read, removing one query per row.
- Extraction takes only the application's own protocol payload: an ordinary
```json or ```python block in a story survives, and a mangled proposal
still does not reach the reader.
Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
|
||
|
|
62a997f364 |
M4: close out Save Points, with browser verification
Closes M4. The review's three findings are fixed, the durability rule the specification always implied is now enforced, and M3's and M4's browser behaviour has been verified in a real browser for the first time. B-1 -- the Save Point list was an N+1 that loaded whole Action rows, narration included, to answer "does a row exist here". It is now one bulk two-column coordinate query plus one lineage: 53 SELECTs for 25 Save Points became 5, and the count no longer grows with the list. The clause is an OR of exact (branch, depth) pairs rather than two IN lists, because the cross product would report a Save Point resolved on the strength of another one's depth existing on this one's branch. A test builds exactly that trap. B-2 -- reclassified during closeout from "missing warning" to a behaviour defect, and fixed as one. STORY-BRANCH-SEMANTICS §19 says a named checkpoint remains until explicitly deleted, and §28 already required future cleanup to retain checkpoint-referenced paths; a cascade that silently removed Save Points with a branch violated both, and a warning would only have documented the violation. A branch a Save Point names can no longer be deleted. The request is refused with the offending Save Points named, the user deletes them explicitly -- which deletes no story -- and the branch then goes. The scope is the subtree, because deleting a branch takes its descendants. Both delete controls disable and explain. Recorded as a new §19.1; models.py, TECHNICAL-DESIGN §8.8 and DATA-MODEL §8 had all recorded the cascade as the rule and now record the refusal. An earlier pass in this same closeout had kept the cascade and added a warning. That was the wrong fix and its tests were replaced rather than left standing, since they pinned the defect. B-3 -- the D11/L03 automation never left one process, so it could not distinguish durable state from a live Python object. It now spawns real server processes, kills the first, and reads the campaign back with the second. C-5 -- creating a Save Point takes the campaign's turn lock. "Save where I am" has to name one committed position, and the head is what a turn in flight is about to move. Rename and Delete deliberately do not take it. The architecture is untouched: a Save Point is still name + note + (branch, depth), and restore is still coordinate -> head.move_to_node -> head.move_to -> attempts.restore_state. No second restore path, no state copied into a checkpoint, no fork on restore. Browser verification -- the first in this project, and it covers both milestones. Firefox 154.0.1 through geckodriver over the W3C WebDriver protocol, driving the rendered DOM: 47/47 checks, twice, on independent databases, no console errors. M3's Undo/Redo enable states, transcript movement, Retry and the take pager, divergence retiring Redo; M4's whole Save Point lifecycle, both confirmations, and the new branch-delete refusal including its recovery. No dependency was added: the WebDriver client is stdlib HTTP. No application defect was found by the browser. Four failures occurred, all in the harness -- a wrong SPA route, a wait comparing transcript length when the empty-story placeholder is longer than the first turn, a fixture deleting the branch it was reading, and a reload assertion that sampled once instead of waiting. The last was checked against the app before being called a harness bug. Tests: 698 backend pass (was 680), 60 M4, 94 M3 history, 66 export/ migrations, 93 security/local-only. Frontend lint and build clean, Docker build clean, loopback binding unchanged. No assertion weakened, no skip added. Planning: STORY-BRANCH-SEMANTICS §19.1 is the only behavioural change and it strengthens §19. V1-ACCEPTANCE-TESTS records D11-D14, I04, L03 and the E-series, keeping automated, live-runtime and browser evidence distinct, and weakens no pass condition. DATA-MODEL records the coordinate with the retry measurement that settles it. BROWSER-UX-SPEC rules for Moment over Turn. BUILD-MILESTONES marks M4 COMPLETE, closes M3's browser condition, and lists what M5 inherits. VERSION adds v2.6. No new ADR: ADR 005 already decides that history is preserved rather than overwritten, and §19.1 is that decision applied to checkpoint-referenced history. M4 is closed. M5 may now be briefed; it has not been started. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2 |
||
|
|
279a871a77 |
Planning: add the M4 implementation review report and rotate M3's
Reporting pass only. No application code, no test, and no product requirement changes. Result: PASS WITH CORRECTIVE WORK REQUIRED. M4's Definition of Done is met and demonstrated at the API level, including across a real two-process restart. The load-bearing constraint holds under inspection rather than assertion: the only head-field assignment M4 added anywhere in the backend is one line in head.py, and an exhaustive grep of the diff finds no second mechanism that forks, prunes memories, reconstructs state, filters the transcript or recomputes Redo. Three corrective items, all M4's own, none in the head model: - GET /checkpoints is an N+1 fetching whole Action rows including prose -- measured at 53 SELECTs for 25 Save Points against 4 for the branch panel, in a codebase that keeps test_egress.py for this exact class of mistake; - deleting a branch silently deletes Save Points naming it, and the branch panel's confirmation does not say so. M4 added the consequence to an existing destructive action without updating its warning; - the shipped D11/L03 tests restart a client, not a process, so the suite is weaker than the acceptance items it is named for. Both pass here only because the report re-ran them across a real process boundary by hand. The browser smoke test is NOT PERFORMED, for M4 and still for M3. Firefox is a snap that hangs past 90s on a trivial headless screenshot; there is no Xvfb, no display, no driver library. Two consecutive milestones now carry an unperformed browser requirement, which the report raises as a standing acceptance risk rather than a defect in either milestone's code. Evidence recorded: 680 backend tests pass (42 M4, 130 M3 invariants, 93 security/local-only), frontend lint and build clean, Docker build clean, migration 80 verified against a representative pre-M4 database with both cascades and zero possible orphans, and I04 verified through a real round trip with branch ids remapped 1->3 and 2->4. M4 is NOT accepted by this report, and M5 is NOT authorized. That decision belongs to whoever reviews this. Rotation: planning/reports/M3-IMPLEMENTATION-REPORT.md moves to planning/archive/milestone-reports/ as a pure rename, contents unedited (git reports 100% similarity, 0 insertions, 0 deletions). Six path references in five active documents are updated because the path changed and for no other reason -- no status claim, no wording change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2 |
||
|
|
e08d49c3eb |
M4: add durable named Save Points
A Save Point is a name for a story position, and restoring one is head movement. That is the whole architecture, and it is what ADR 012 and BUILD-MILESTONES' note on M4 asked for: M3 made the head a stored (branch, depth) and made arriving at one a row lookup plus a state restore, so a Save Point needs no restore machinery of its own. What the user gets: - Name the moment they are reading, keep playing, restart the app, and come back to it. Restoring moves the story back and deletes nothing: the later turns stay, Redo still walks forward into them, and writing something different is what starts a new line while the old one is kept. - Rename, delete, and a list, in a Save Points panel beside the branch panel, with a Save Point button next to Undo and Redo. Both confirmations say what is *not* destroyed, because that is the part the screen cannot show. - Save Points survive export and import. What was deliberately not built: - No second restore path. `head.move_to_node` is the only new movement: its depth half is M3's `head.move_to` unchanged, and its branch half is the single assignment `switch_branch` already makes. No head field is written in the checkpoint router, nothing reconstructs state, nothing prunes a memory, nothing copies or deletes a turn, and restore never forks — the first write below the restored head does, through `fork_if_behind_head`. - No automatic cleanup. A Save Point behind the head, or naming a line the story left, is doing its job (STORY-BRANCH-SEMANTICS §19). The one removal is a cascade: deleting a branch takes its Save Points, as it takes its memories, because the story they named went with it. - No new ADR. ADR 012 already decides the architecture, and a table is not a decision. The one call the planning package did not already make: restore moves the branch half of the head only when the coordinate is off the path being read. Doing it unconditionally would quietly hand back an abandoned continuation whenever a Save Point in a shared prefix was restored; never doing it would make a Save Point on a departed line unrestorable, which contradicts §19. TECHNICAL-DESIGN §8.8 records it. Schema: a `checkpoints` table holding a name, an optional note and a (branch, depth) coordinate — no copy of any story. `create_all` builds it as it did `memories` and `branches`; migration 80 adds the index. No backfill, because nobody had named a position before M4. The coordinate is deliberately not an action id: one coordinate holds every attempt at a turn and exactly one is live, so a coordinate follows a retry where a row id would pin a take the story no longer tells. Tests: 680 pass (638 before). 42 new in tests/test_save_points.py covering D11-D14, I04, L03, E-series lineage and memory isolation after restore and divergence, the edge cases, and an M3-database migration. One pre-existing fixture in test_tree_migration.py needed `checkpoints` added to its drop list — SQLite refuses to drop a table another table references. Not verified: the browser. No session has had a usable one, so the Save Point panel's DOM behaviour is unobserved — as M3's Redo control still is. The twenty-step sequence was driven over HTTP against a live server with a real process restart instead, and all seventeen checks pass. M4 is implemented, not accepted: no review has been written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2 |
||
|
|
3c8e91f644 |
Docs: correct post-M3 status and Ollama configuration
Three stale claims found in active documentation after the post-M3 consolidation: - BUILD-MILESTONES.md still opened "M1 and M2 complete; M3 next", which contradicted its own M3 "Status: COMPLETE" block, planning/README.md and VERSION.md. It now states M1, M2 and M3 complete and accepted, M4 next. - DEVELOPMENT.md described the settings row as "endpoint, model and (unused) API key" and sent "api_key":"" in its curl example. M2 removed the field; schemas.SettingsUpdate has no api_key. The prose and the example now match the real request shape. No application code was changed. - README.md listed LM Studio as a supported local endpoint. ADR 002 and ADR 011 make Ollama the only v1 backend; LM Studio is a rejected alternative there. The row is removed and the surrounding wording now says Ollama is the supported backend, same-host is the default, trusted-LAN Ollama is supported, public/cloud is prohibited, and the OpenAI-compatible adapter is an implementation detail rather than a support promise. The claude_shim section stays, relabelled "(development only)". Documentation only; no code, schema or test changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2 |
||
|
|
d27ee34901 |
Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was authoritative. Phase 0 execution prompts sat beside the specification; four completed milestone reports sat beside the current one; and upstream AI-DnD's own `plan/` build log and `docs/` project site still described a hosted, scripted, multi-user product with accounts — every screenshot in it showed a Scripts tab and a Sign up button, none of which has existed since M2. `planning/archive/` now holds the history and says so in its own README: `phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied. `planning/reports/` holds only the current milestone's report, because that is the one M4 planning has to read; it moves to the archive when M4's replaces it. Deleted rather than archived: the Phase 0B execution prompts and the handoff/status/summary documents, the Phase 0A discovery and triage reports, upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's template boilerplate. All of it is in Git history, and the two recommendation reports carry every conclusion the deleted research reached. Archived documents are kept verbatim. Paths written inside them point at where those files were when the document was written, which is the point: an evidence record that has been quietly edited is no longer evidence. Active documentation is corrected where it pointed at the removed trees or described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not touch" list had gone stale at M2 and claimed QuickJS scripting was still tested; its test count was 604 against an actual 638. `README.md` loses the upstream CI badge, which reported upstream's pipeline rather than this fork's, and a reference to `backend/app/worldstate/engine.py`, a file that does not exist. `planning/README.md` is rewritten as the documentation index. New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the manifest of what belongs in the ChatGPT project's Sources. Source comments referring to the deleted trees are reworded; no behaviour changes. 638 backend tests pass, frontend lints and builds, and a reference scan over all 48 tracked Markdown files reports no unresolved path in active documentation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu |
||
|
|
c8755c21c2 |
Planning: close M3 and record active-head architecture
M3's review recommended planning changes and, following the M2 pattern, reported rather than applied them. This applies them, and adds the ADR the review asked for. ADR 012 records the architecture rather than the requirement. ADR 005 already says that going backward must preserve abandoned history and that the user sees Undo/Redo/Retry rather than branch management; it names a movable active head as the direction and stops. What M3 settled is the shape: the head is stored rather than derived, every read of the story is capped at it in one place, one mechanism moves it, the state of a position comes off the node rather than from a replay, the first write below a moved-back head is the divergence, and whether Redo exists is decided by the lineage rather than by a flag that could be stale. The last of those is the property worth keeping — a flag can be wrong and make the story wrong; a lineage cannot. Two semantics are ratified in STORY-BRANCH-SEMANTICS.md, both of them reversals or narrowings that a reader would otherwise take for bugs. Undo now crosses fork points and continues to the campaign opening, because refusing at the fork was a consequence of deleting rows the parent line was also reading, and nothing is deleted any more. And the system refuses to switch which take is live while a later story is off screen, because doing it quietly would leave retained history continuing from words the story no longer says. A new §14A covers editing in place. §14-15 describe the finished behaviour — the edit becomes authoritative, the state it implies is re-evaluated, a new continuation is created, the original is retained — and that requirement is intact and explicitly not weakened here. It is also not built, because re-evaluating state from prose a user typed needs M5's extraction pass. §14A says what exists in the meantime and why refusing is the minimum that holds the invariant rather than the destination. TECHNICAL-DESIGN.md gains §8.7 and §9.1, recording the implemented model and the bundle behaviour as fact in the way §5.2 records M1 and M2. §10.4 gains a constraint that is easy to lose: the snapshot half of the hybrid state model is a requirement, not an optimization. Head movement is a row lookup plus a restore, which is why Undo, Redo and Save Point restore cost the same at any distance into a campaign; a state model recoverable only by replaying from the opening would make all three proportional to campaign length, on exactly the long campaigns this product is for. DATA-MODEL.md records the head as stored on the campaign rather than derived from its newest turn — two campaigns holding identical turns can be read at different places, and nothing about the turns can tell them apart — and the branch disposition as implemented: the depth a divergent write left the branch at, deliberately advisory, and carried through export because every row of an abandoned line is exported either way. BUILD-MILESTONES.md marks M3 complete and states the one condition still open. M4 is told a Save Point is a durable pointer and that restoring one is head movement with a bounds check, not a restore system: a second mover is the specific failure to avoid, because the two paths would silently disagree about what restore means. M5 gets three constraints — keep state efficiently recoverable, move the test instrumentation rather than the assertions when the world-state protocol goes, and finish the narrator edit §14A defers. V1-ACCEPTANCE-TESTS.md clarifies ownership without lowering a bar. D10 keeps all three pass conditions and is explicitly recorded as *not* satisfied at the end of M3; what changed is that the document now says which milestone delivers which condition. D03's result is recorded as a full pass rather than the partial the text allowed for, I07 gains the pre-M3 bundle clause, and L01 gains the note that resolves its apparent conflict with A05 — a failed turn does advance the head by one, onto the player's retained input, and that is A05 working rather than L01 failing. README.md described a different application: a hosted demo, guest accounts, cloud providers, Postgres, a Render blueprint, an analytics dashboard, a QuickJS scripting engine, and 549 tests. M2 removed all of that and the README was never updated — a gap M2's own debt table missed. It now describes what this fork is, including the endpoint policy and the TLS behaviour, and the numbers in it are the current ones. M3's report is included here as its own evidence record: no separate baseline report was produced, so it carries the raw counts and runtime observations as well as the review, and §W records this closeout. SPECIFICATION.md and SECURITY-THREAT-MODEL.md are unchanged. M3 altered no product requirement and touched no path in the threat model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u |
||
|
|
2fdd2547f0 |
Planning: record M2 closeout decisions
M2's review reported six planning recommendations rather than applying them, three marked before M3. All six are applied here, plus three additions drawn from the same evidence. No implementation file is touched. The endpoint policy was the gap that mattered. It is the most consequential setting in the application — the storyteller sends the player's prose, the context, the memories and the embedding inputs to whatever address it names — and it existed only as a module docstring. It is now ADR 011 and a new §10A in the threat model, which also retires the assumption in §71A that the inherited guard was a starting point. It was not: AI-DnD's SSRF guard blocked private addresses to stop a hosted server reaching its own internal network, which is the exact opposite of what a local storyteller needs. It was removed, not adapted. Both documents state the rule as implemented — an allowlist of explicit local-network CIDRs, every resolved address checked, enforced on save and again before every outbound request, TLS never traded against it — and both state the two residual limits plainly rather than implying they are covered: a hostile host already on the trusted LAN is inside the permitted boundary, and a rebinding interval exists between the policy's resolution and the client's connection. Accepted risks, not M3 work. The CIDRs are spelled out rather than derived from is_private/is_reserved, and the ADR records why: is_private is true of the documentation ranges and 0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it refuses an ordinary same-host Ollama on [::1]. TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening items. A new §5.2 records the M1/M2 architecture as fact rather than intention, so later milestones inherit what the code does. A new §18.1 carries the lesson of M2's two regressions: when removing a setting, test a real consumer construction path; when adding one, prove it reaches the component that uses it. Both defects hid behind a green suite because the tests at that boundary were mocks. BUILD-MILESTONES records M2 complete, with the capabilities later milestones inherit and the debt carried forward. Two notes go to milestones that would otherwise misread what M2 left them. M5 is told that eight rollback tests now use the world-state engine as instrumentation and not as endorsement — the instrumentation moves when the protocol does, and those tests are reworked rather than deleted. M6 is told that the memory bank died silently under a green suite, so background failure must be observable and at least one real provider-construction path must be tested. The security contract gains what M2 demonstrated. H10 now names the two conditions that were defects during M2: a wildcard origin must be refused at startup, and an unknown /api path must 404 rather than returning the SPA with 200. New H12 covers endpoint enforcement, and its fourth pass condition is the one that matters — a public endpoint written into the database behind the settings API must still be refused at the wire. A build passing the first three and failing that one has configuration validation only. SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement; it removed capability the specification never asked for. The two M2 reports gain appended closeout notes rather than edits. Their original wording about an uncommitted working tree was true when written, and the note records what happened afterwards: the six-file correction is |
||
|
|
8652fe7cd8 |
M2 review: two regressions the green suite hid, and the reports
The post-implementation review of M2, plus the three corrections it took
to make the evidence true. Reports:
planning/reports/M2-BASELINE-REPORT.md 868 lines, the measurements
planning/reports/M2-IMPLEMENTATION-REPORT.md 758 lines, the reading of them
Verdict is PASS, accept with non-blocking debt, proceed to M3. Every M2
requirement is met and the ones that matter were tested by running the
build rather than reading it: a cloud endpoint written straight into
SQLite with sqlite3, behind the API's back, still refused at the wire;
trusted-LAN HTTPS against the real second machine with verification on;
captures showing zero packets outside loopback and the approved host.
Three defects, all found by running the shipped image.
The memory bank was dead. M2 removed Settings.api_key_plain with the API
key, and memorybank's two provider factories still read it. It failed
inside a fire-and-forget task, so no user error, no log anyone would
read, and no test — every memory test stubs those factories. All 604
tests passed with summaries and embeddings silently not happening.
The configurable model timeout never reached the turn engine. Stored,
validated, exposed in the API, rendered in the UI, and not passed to the
provider. M2's own exit criterion was half met: the constant had moved
but the setting did nothing.
And requirements.lock still pinned quickjs, psycopg and cryptography, so
the setup path DEVELOPMENT.md gives a new developer would have
reinstalled all three.
Both code defects now have the test that would have caught them: one
constructs every provider factory from a real Settings row, one drives
the turn endpoint, the chat endpoint and the summariser and asserts the
configured timeout arrives at each. That is the lesson worth keeping from
this milestone — after removing an attribute, build each consumer from a
real object; after adding a setting, prove it lands. Both failures were
in background or plumbing paths, which is exactly where a subtractive
change cannot see itself.
606 tests pass, up from 604. Lint, build and image are clean. Every
runtime result in the baseline report came from an image built after
these fixes; the reports say plainly that commit
|
||
|
|
1a28a9a708 | Apply post-M1 corrections to the planning package | ||
|
|
645f07f06d | Add the M1 implementation review report | ||
|
|
c1a73b3d77 |
M1: make the first story turn work with no Internet
Phase 0B ran the upstream application on a network with no route out and the first turn died in tiktoken, which downloads its BPE table the first time anything counts a token. The browser separately fetched three font families from Google on every page load. Neither is visible on a machine that has been online once, which is why both now have tests. The tokenizer table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it directly, verifying its SHA-256 against the digest tiktoken itself pins for that URL. No code path in the tokenizer can reach the network any more — not a warm cache, not an environment variable a deployment could forget. The encoding was checked token for token against tiktoken's own. The three font families are self-hosted as variable fonts under frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in frontend/src/styles/fonts.css, and re-vendored by frontend/tools/vendor_fonts.py. Their OFL licences ship beside them. With no remote asset left, the CSP drops both Google hosts and gains object-src, base-uri and form-action; woff2 also gets its real media type, which Python's table lacks on a slim image. A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx verifies against the certifi bundle, so an endpoint whose certificate comes from a CA the user installed on their own machines — a StartOS server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED while curl and the browser on the same host accepted it. app/tlstrust.py builds one context that unions the platform CA store with certifi's, and all four outbound clients use it. A union rather than a swap, so an image with an empty system store cannot start failing on endpoints that worked before. Verification itself is untouched: CERT_REQUIRED, hostname checking on, and no insecure escape hatch. The storyteller listener is now loopback by explicit statement rather than by inheriting uvicorn's default: start.sh, start.ps1, and docker-compose.yml, which publishes to 127.0.0.1 rather than every interface. Reaching an Ollama on another machine is outbound and needs none of that inbound exposure. backend/requirements.lock pins the exact tested closure; requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the same-host and trusted-LAN Ollama configurations, and how to re-run the offline proof. PROVENANCE.md records the upstream commit, the MIT terms, and both vendored assets. Verified, not just compiled. On an --internal Docker network with 1.1.1.1 unreachable and no name resolving, a campaign was created and played for six turns through same-host Ollama, restarted, and resumed. A second run played ten turns through Ollama on a separate physical machine on the LAN over verified HTTPS, summaries and embeddings included, with the storyteller's default route deleted so the LAN was reachable and the Internet was not. Its capture: 893 packets to the approved host, 730 loopback, zero anywhere else, and zero DNS queries. Two induced model failures left the accepted story bit-identical. The inherited SPA was opened in a browser and a campaign read back from it. Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the findings that did not belong in this change. 648 backend tests pass, up from the inherited 632; frontend lint and build are clean; the image builds. No M2 work is included: the hosted, cloud, analytics, Postgres and scripting surfaces are untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL |
||
|
|
717670afe0 | Update planning package after Phase 0B | ||
|
|
ba737de9b4 |
Add Phase 0B local validation findings and recommendation
Validates the three finalists by clone, build, test run and live local Ollama inference, then answers the fork question with measurements rather than static review. Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but it was wrong that AI-DnD's undo is non-destructive — retry preserves the replaced take, undo hard-deletes it. A follow-up spike fixed that in 3 files (+130/-31): undo now moves a head cursor, redo round-trips, writing below a moved-back head forks and keeps the abandoned line, branch-scoped memory isolation survives, suite 627/632 with all 5 failures asserting the deleted-row behaviour that was replaced. Findings that change the plan: - AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its encoding from a CDN. Proven on an internal Docker network, proven fixed by vendoring the file. - ai-adventure needs zero code for Ollama — two config lines — and its turn/head/checkpoint schema is the target model to build to. - Open Dungeon has zero automated tests and a positional summary watermark, making its branch retrofit larger than Phase 0A costed. - The world-state referee takes relative deltas; a 3B model sent absolute values under full context, so a wounded player ended at full health. Validation cannot catch this, so prefer ai-adventure's typed-event vocabulary when generalising narrative state. - Export/import recomputes head depth, so a round-trip silently undoes an undo. Must be fixed alongside the undo work. Docs only; no production code. Working tree from the runs stays untracked under phase0b/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu |
||
|
|
f011362494 | Add initial planning files from ChatGPT research here |