Commit Graph
16 Commits
Author SHA1 Message Date
JesseMarkowitzandClaude Opus 5 ef25b0a876 Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
2026-09-10 06:13:55 -04:00
JesseMarkowitzandClaude Opus 5 144406cd48 M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was
whether any of the earlier evidence meant what it said. M8 measured a deployment
enforcing a 4,096-token input window while the application budgeted 16,384.
Every request returned 200. What Ollama does with the excess is drop the oldest
tokens, and the oldest tokens here are the system block — the narrator's rules
and the campaign canon. A hundred-turn certification against that server would
have looked perfect and proved nothing, which is why this milestone could not
begin with a hundred turns.

So the application asks now. Ollama's window is a property of how a model was
loaded rather than of the request — sending num_ctx is accepted, ignored, and
worse, reloads the model at the server's own default — so the only honest move
is to find out and then tell the truth about it. /api/ps reports what a resident
model is being served with, /api/show what an unloaded one will load with, both
on the same host inference already uses, through the same endpoint policy and
the same TLS trust store. A verified window is a ceiling on the budget; an
unverified one leaves the budget alone and is recorded as unverified in the
turn's own provenance, so an old turn can be asked afterwards whether it was
built against a checked window. There is no third behaviour, and in particular
no hard-coded 4,096: a number the server did not say would be right on one
machine and wrong on the next.

The proof that this is doing something is a campaign whose canon sits at the
front of the prompt, 120 turns of history, and a 4,096-token window. The canon
is still there afterwards and the oldest history is gone. The same campaign
built the old way produces a prompt more than twice the window — the defect,
reproduced, so the fix is measured against it rather than asserted.

Two defects the validation found on its own, and they are the same defect twice:
something was true and nobody was told. A manual state correction of four
changes with one bad reference applied three, returned 201, and said nothing —
while recording the refusal on the audit row nobody reads. It came to light
because the identity diagnostic's own fixture was refused that way and the whole
run proceeded on a campaign with no scene, which would have read as a model
failure. And the narration-length setting moved no number: brief, medium and
long each became one English sentence, while the numeric hint the model actually
reads was derived from the global reply cap and said the same thing for all
three. Both now say what they did.

The other two post-M8 findings are closed as well. The tab said AI D&D, which no
document had ever claimed it did not; it says Interactive Story now, with the
open campaign first, and the name is the owner's decision rather than a
find-and-replace to something narrower than the engine. After an Undo the reader
could not tell where they had landed; the control row now ends with
"Moment 11 · later story ahead", from the server's own answer, in the word the
transcript already uses, with none of head, branch or depth anywhere near it.

The identity diagnostic exists and the root cause does not. That campaign was
destroyed, so no cause can be established — what M11 owes the finding is
something that can classify the next occurrence, and a diagnostic that makes only
the judgements a program can honestly make: duplicate keys, shared names,
protagonist drift, state and context disagreeing. Whether prose misattributed a
line is left to a person reading it beside its prompt, because a regex cannot
read dialogue and one that pretended to would produce exactly the confident wrong
answer this finding is about. Its detectors are proved to fire against a planted
second Alice.

Two entities may still share a display name. That was checked first, as the
finding asked, and left permitted: a mother and a daughter, or a stranger giving
a false name, are ordinary fiction, and refusing them to guard against a model
mistake would refuse the wrong thing. What was missing was that it happened
silently. It is reported now.

Evidence, not inference: a hundred accepted turns against a real narrator with
genuine process restarts; a real browser against the built SPA; a container with
no network at all; a campaign moved into a data directory that never existed.
Each was discarded and re-run whenever the product changed under it, and the runs
that were thrown away are listed in the report with the reason, along with ten
defects in the harnesses themselves — because a harness that has only ever
agreed with itself is not evidence, and two of M8's five harness defects were
masking real ones.

No dependency was added, removed or upgraded. No acceptance test was retired,
relaxed or reclassified. M11 is implemented and verified; it is not accepted, and
there is no release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 14:01:20 -04:00
JesseMarkowitzandClaude Opus 5 1013c94eb1 M10: the seam for media, and no media
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
The media extension contract asks for a scene snapshot a future image or video
provider could be handed: location, who is present, what they hold, what must
stay true, and where in the story it sits. Building one was the milestone's
obvious first task, and it was the wrong one. That snapshot has existed since
M5. `narrative_state["scene"]` holds the summary, the location, the cast and the
coordinate it was written at; a validated `set_scene` event writes it, every
position snapshots it, and every head move restores it. It survives Undo, Redo,
Retry, divergence, Save Point restore and a process restart because it is the
authoritative state rather than a copy of it.

So there is no scenes table here. A second scene store would have been a second
answer to "where is the story now", with its own lineage rules to get wrong —
and the lineage rules are the expensive part, which is the argument for reusing
the ones that already work rather than against it. The Scene Packet is derived
on read, and its identity is computed from the campaign and the position rather
than allocated: the same position yields the same id in another process, after a
restart, and after the packet is thrown away and rebuilt, with no row to keep in
step. That is the part of a future media_assets table that would be expensive to
retrofit, so it is fixed now even though the table is not built.

One table, then: visual_profiles, the only thing the contract's scene list asks
for that nothing already stored. Campaign-scoped and not per-position, because a
character does not change appearance when the story forks — a reader who
diverged would otherwise lose their cast, and the same descriptors would land in
every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn
campaign to say something that never varies. Keyed by the M5 entity key rather
than a new identity namespace, and one table for characters, locations and items
alike, because a location is an entity with a type and splitting them would
reintroduce the genre shape M5 spent a milestone removing.

What the packet leaves out is the more interesting half. Not the transcript, and
not imported knowledge — none of it, not merely the sources marked hidden. The
rule is what the story established at this position, not everything the narrator
was told, and drawing it by class is what makes it hold for a secret nobody
thought to mark. A hidden Canon source proves it, with a positive control
showing the narrator did receive the sentinel the packet does not carry. Once a
validated event puts the observer in the room, the observer is in the packet:
that is no longer narrator-only knowledge, and a packet that hid it would be
hiding the story from itself.

The providers are contracts and nothing else. Protocols for image, video, audio,
speech and transcription, an empty registry, no adapter, no dependency, no
socket, and no media setting to point anywhere — a setting that exists can be
pointed at a cloud by mistake. A future provider endpoint must be loopback,
stricter than narration's trusted-LAN allowance, because a picture of a scene
carries the scene with it. Transcription returns an editable draft with no
commit method, so STT structurally cannot bypass the authoritative path.

Nothing here can write the story. Not by convention: no module under media/
imports the code that writes state, no media event type exists in the state
vocabulary, and every test in the authority suite compares the authoritative
document byte for byte either side of a media operation — including one where a
provider insists Alice is in a red coat in a corridor, and the campaign goes on
disagreeing.

One defect, found by the milestone's own tests. M10 first added a migration
creating an index that create_all already builds from the column, so an upgraded
database ended up with two indexes and a fresh install with one. Comparing the
two schemas is what caught it; neither database examined alone would have. The
migration is gone rather than renamed, and the right number of migrations for a
new table whose indexes are declared on its columns is zero.

Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145
passed. Lint, production build and Docker build clean. No frontend file changed:
M10 adds no reader-facing surface, and ordinary play — turns, state, memory,
knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media
configuration, no warning, no connection attempt and no media row written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 03:41:04 -04:00
JesseMarkowitzandClaude Opus 5 44edece67e M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 01:55:45 -04:00
JesseMarkowitzandClaude Opus 5 1ce9972760 M8: the browser becomes the storyteller
The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.

All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.

So the shape now is one entry point and one screen:

  Campaigns -> Campaign -> Story
                           State · Knowledge · Context · Save Points · Settings

Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.

Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.

Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.

The two defects worth the space:

A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.

And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.

Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.

`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.

Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.

Backend, and only what the browser could not otherwise reach:

  AdventureCreate.opening   a start action could only come from a Scenario, so
                            every campaign made in the new setup flow opened on
                            a blank page. Same node, same code path.
  canon_rules               campaign_canon has been the highest authority in a
                            campaign since M5, read by the prompt builder and
                            the state validator, and had no API at all — a
                            fixture had to write it with SQL.
  a 401 and a 429 message   the last user-facing text describing a hosted
                            deployment. One told the reader to check an API key
                            that has not existed since M2.

No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.

The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.

They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.

A verification pass over all of it then found three more, each by driving the
product rather than reading it:

Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.

The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.

And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.

Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.

`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.

Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:

The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.

Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.

The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.

The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.

Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 23:31:45 -04:00
JesseMarkowitzandClaude Opus 5 480414efe0 M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or
Inspiration, and the class is load-bearing rather than a label: it decides the
words a passage is framed with in the prompt, the weight it carries when
passages are ranked, and which budget it competes in when the context is tight.

This is a separate subsystem, which is the Phase 0B decision
(IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification,
provenance, content identity, chunking, an index or a lifecycle, and they were
not promoted into something that does. Nothing here reads or writes one.

The subsystem, in backend/app/knowledge/:

  classes      the three classes, their weights, and the prompt framing
  chunking     deterministic, heading-aware, 60-800 tokens, no overlap
  fts          SQLite FTS5 with porter stemming; scoped and bounded in SQL
  importer     validate, hash, store, chunk, index — in one transaction
  embeddings   local Ollama vectors through the shared provider
  retrieval    query construction, hybrid merge, rerank
  inject       the budgeted cut and the rendered prompt sections

Relevance admission is a separate stage from ranking, and that separation is
the milestone's most expensive lesson. An independent review found the first
implementation deciding relevance with a floor expressed as a share of the best
candidate — which the best clears by construction — so a passage was admitted on
every turn regardless of the scene. A query about tide tables and container
tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden
Canon among them.

So the pipeline is now:

  candidate generation -> admission -> ranking -> class weighting -> budget

Admission reads raw, candidate-set-independent signals: the cosine the model
returned, and how many distinct meaningful query terms a passage contains.
Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero
is not zero. Normalization decides order among things that matched; it can never
decide whether anything matched. Authority is applied after admission, so a
class orders what matched and never rescues what did not.

Retrieval may therefore return nothing, and on a scene unrelated to the library
it does.

The other decisions that each replaced an obvious wrong one:

- The class multiplies relevance rather than adding to it. An additive bonus
  satisfies "Canon outranks Reference" and makes "do not include irrelevant
  Canon" impossible, because a large enough constant wins on its own.
- The semantic floor is measured, not guessed: 113 production-path pairs against
  nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at
  0.36-0.56, and 0.58 sits between them. Because it is a property of that model
  and not of cosine similarity, it is keyed to the model rather than applied to
  whatever is configured: an embedding model with no measured calibration in
  this build does not borrow the number. Semantic admission is skipped, the
  campaign retrieves lexically, and the reason is stated in the knowledge status
  and in the turn's provenance. Degrading to lexical keeps the library usable;
  lending the threshold to an unmeasured model is how the admitted-everything
  defect would return.
- One lexical term is not evidence. Two distinct meaningful terms, or one that
  is neither a standing campaign entity nor a negligible share of the query.
  The stop list grew from 42 words to 261, all function words — no subject
  matter, because a stop list that removes subject matter stops finding "The
  Silver Key".
- Lexical retrieval is a production path, not a fallback. It finds the proper
  nouns and invented terms a setting bible is made of, and the library is fully
  usable with no embedding model configured.

Safety is structural rather than filtered. Imported text reaches the prompt
whole, inside a section that says what it is, under a rule stating the authority
order in words and refusing every instruction inside it. No endpoint accepts a
filesystem path, so H08 has no mechanism to escape from. Nothing renders
imported content as HTML, so a script tag is five visible characters and a
remote image is never fetched. Import, chunking, indexing, retrieval and a turn
open no socket at all; only embeddings do, through the endpoint allowlist the
memory bank already uses.

Provenance is the rendered text, not a foreign key: deleting a source cannot
turn a historical turn's evidence into dangling ids.

Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5
virtual table attached to knowledge_chunks as a DDL hook so it is created and
dropped with the table it indexes. Migration 92. A pre-M7 database opens
unchanged and needs no sources to play.

Bundle: the source content and the reader's judgements about it travel; the
passages, index rows and vectors are rebuilt on import, so a restored campaign
is searchable immediately without a reindex step.

One runtime dependency: python-multipart, Starlette's multipart parser. It is
what makes the upload surface possible, and the upload surface is why no
pathname is ever accepted.

The test doubles were the reason the defect shipped, so they were corrected too.
The retrieval stub scored unrelated text at 0.06-0.20 where the real model
scores it at 0.43-0.44, and its docstring said it had deliberately removed the
constant component that "would put a similarity floor under every pair" — which
is exactly the property real models have. The stub now has that floor, one test
fails if it is ever removed, and another reproduces the superseded rule and
asserts it is still fooled by the same fixture. Run against the pre-corrective
implementation, the new suite fails 13 of 18.

Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of
which mocks nothing between itself and Ollama and re-measures the similarity
separation on every run. 43/43 checks in a real Firefox, reproduced.
Docker build clean.

Four other defects found by review or by the browser run were fixed here rather
than carried: an unreachable relevance constant that appeared to enforce
something and did not; acceptance tests using the wrong fixture files, so G07's
trap was never exercised; a bidirectional override surviving into displayed
filenames; and, from the implementation pass, the Insights panel showing M5's
two state sections as raw keys and the source inspector refetching on every
keystroke.

M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK
REQUIRED. Both blocking findings are closed, and closeout resolved the
embedding-model calibration boundary the corrective pass had left as debt.
planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective
closeout and the closeout verification in sequence, none overwriting another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 15:40:13 -04:00
JesseMarkowitzandClaude Opus 5 a6e9c7a32b M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-06 03:00:33 -04:00
JesseMarkowitzandClaude Opus 5 b7005e6fdd M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.

This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:

    visible active transcript position == stored head == authoritative state

Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)

  A narrator edit no longer rewrites a row. It returns to the state before the
  turn, takes the reader's exact text as the accepted narration, re-derives the
  state that text implies, and becomes a new active continuation — while the
  original narration keeps its words, its live flag and its whole future as
  retained history. At the tip the correction is another take; with story below
  it, it forks. No new history machinery: this is the existing fork/take/head
  path with the reader's text in place of a generated reply. The §14A refusal
  is therefore gone for narrator turns, and remains only for player input.

Pre-M5 positions

  Migration 88 backfills the empty narrative document onto every action written
  before M5, and a missing snapshot now restores the empty document instead of
  leaving the previous position's state standing. Restoring to an old Save
  Point no longer leaves a later position's entities and facts on screen.

Narrator context

  Replayed history carries prose only; the machine-readable block is no longer
  reconstructed into past turns, where it contradicted the authoritative state
  in the same prompt. A fact withdrawn by a manual correction is now named as
  no longer true, with the reader's reason, rather than silently dropped.

Also

  - state_changes joins the action-list bulk read, removing one query per row.
  - Extraction takes only the application's own protocol payload: an ordinary
    ```json or ```python block in a story survives, and a mangled proposal
    still does not reach the reader.

Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-05 07:01:50 -04:00
JesseMarkowitzandClaude Opus 5 62a997f364 M4: close out Save Points, with browser verification
Closes M4. The review's three findings are fixed, the durability rule the
specification always implied is now enforced, and M3's and M4's browser
behaviour has been verified in a real browser for the first time.

B-1 -- the Save Point list was an N+1 that loaded whole Action rows,
narration included, to answer "does a row exist here". It is now one bulk
two-column coordinate query plus one lineage: 53 SELECTs for 25 Save Points
became 5, and the count no longer grows with the list. The clause is an OR
of exact (branch, depth) pairs rather than two IN lists, because the cross
product would report a Save Point resolved on the strength of another one's
depth existing on this one's branch. A test builds exactly that trap.

B-2 -- reclassified during closeout from "missing warning" to a behaviour
defect, and fixed as one. STORY-BRANCH-SEMANTICS §19 says a named checkpoint
remains until explicitly deleted, and §28 already required future cleanup to
retain checkpoint-referenced paths; a cascade that silently removed Save
Points with a branch violated both, and a warning would only have documented
the violation. A branch a Save Point names can no longer be deleted. The
request is refused with the offending Save Points named, the user deletes
them explicitly -- which deletes no story -- and the branch then goes. The
scope is the subtree, because deleting a branch takes its descendants. Both
delete controls disable and explain. Recorded as a new §19.1; models.py,
TECHNICAL-DESIGN §8.8 and DATA-MODEL §8 had all recorded the cascade as the
rule and now record the refusal.

An earlier pass in this same closeout had kept the cascade and added a
warning. That was the wrong fix and its tests were replaced rather than left
standing, since they pinned the defect.

B-3 -- the D11/L03 automation never left one process, so it could not
distinguish durable state from a live Python object. It now spawns real
server processes, kills the first, and reads the campaign back with the
second.

C-5 -- creating a Save Point takes the campaign's turn lock. "Save where I
am" has to name one committed position, and the head is what a turn in
flight is about to move. Rename and Delete deliberately do not take it.

The architecture is untouched: a Save Point is still name + note +
(branch, depth), and restore is still coordinate -> head.move_to_node ->
head.move_to -> attempts.restore_state. No second restore path, no state
copied into a checkpoint, no fork on restore.

Browser verification -- the first in this project, and it covers both
milestones. Firefox 154.0.1 through geckodriver over the W3C WebDriver
protocol, driving the rendered DOM: 47/47 checks, twice, on independent
databases, no console errors. M3's Undo/Redo enable states, transcript
movement, Retry and the take pager, divergence retiring Redo; M4's whole
Save Point lifecycle, both confirmations, and the new branch-delete refusal
including its recovery. No dependency was added: the WebDriver client is
stdlib HTTP.

No application defect was found by the browser. Four failures occurred, all
in the harness -- a wrong SPA route, a wait comparing transcript length when
the empty-story placeholder is longer than the first turn, a fixture
deleting the branch it was reading, and a reload assertion that sampled
once instead of waiting. The last was checked against the app before being
called a harness bug.

Tests: 698 backend pass (was 680), 60 M4, 94 M3 history, 66 export/
migrations, 93 security/local-only. Frontend lint and build clean, Docker
build clean, loopback binding unchanged. No assertion weakened, no skip
added.

Planning: STORY-BRANCH-SEMANTICS §19.1 is the only behavioural change and it
strengthens §19. V1-ACCEPTANCE-TESTS records D11-D14, I04, L03 and the
E-series, keeping automated, live-runtime and browser evidence distinct, and
weakens no pass condition. DATA-MODEL records the coordinate with the retry
measurement that settles it. BROWSER-UX-SPEC rules for Moment over Turn.
BUILD-MILESTONES marks M4 COMPLETE, closes M3's browser condition, and lists
what M5 inherits. VERSION adds v2.6.

No new ADR: ADR 005 already decides that history is preserved rather than
overwritten, and §19.1 is that decision applied to checkpoint-referenced
history.

M4 is closed. M5 may now be briefed; it has not been started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-04 06:34:56 -04:00
JesseMarkowitzandClaude Opus 5 279a871a77 Planning: add the M4 implementation review report and rotate M3's
Reporting pass only. No application code, no test, and no product
requirement changes.

Result: PASS WITH CORRECTIVE WORK REQUIRED.

M4's Definition of Done is met and demonstrated at the API level, including
across a real two-process restart. The load-bearing constraint holds under
inspection rather than assertion: the only head-field assignment M4 added
anywhere in the backend is one line in head.py, and an exhaustive grep of the
diff finds no second mechanism that forks, prunes memories, reconstructs
state, filters the transcript or recomputes Redo.

Three corrective items, all M4's own, none in the head model:

- GET /checkpoints is an N+1 fetching whole Action rows including prose --
  measured at 53 SELECTs for 25 Save Points against 4 for the branch panel,
  in a codebase that keeps test_egress.py for this exact class of mistake;
- deleting a branch silently deletes Save Points naming it, and the branch
  panel's confirmation does not say so. M4 added the consequence to an
  existing destructive action without updating its warning;
- the shipped D11/L03 tests restart a client, not a process, so the suite is
  weaker than the acceptance items it is named for. Both pass here only
  because the report re-ran them across a real process boundary by hand.

The browser smoke test is NOT PERFORMED, for M4 and still for M3. Firefox is
a snap that hangs past 90s on a trivial headless screenshot; there is no
Xvfb, no display, no driver library. Two consecutive milestones now carry an
unperformed browser requirement, which the report raises as a standing
acceptance risk rather than a defect in either milestone's code.

Evidence recorded: 680 backend tests pass (42 M4, 130 M3 invariants, 93
security/local-only), frontend lint and build clean, Docker build clean,
migration 80 verified against a representative pre-M4 database with both
cascades and zero possible orphans, and I04 verified through a real round
trip with branch ids remapped 1->3 and 2->4.

M4 is NOT accepted by this report, and M5 is NOT authorized. That decision
belongs to whoever reviews this.

Rotation: planning/reports/M3-IMPLEMENTATION-REPORT.md moves to
planning/archive/milestone-reports/ as a pure rename, contents unedited
(git reports 100% similarity, 0 insertions, 0 deletions). Six path
references in five active documents are updated because the path changed
and for no other reason -- no status claim, no wording change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-03 19:15:53 -04:00
JesseMarkowitzandClaude Opus 5 e08d49c3eb M4: add durable named Save Points
A Save Point is a name for a story position, and restoring one is head
movement. That is the whole architecture, and it is what ADR 012 and
BUILD-MILESTONES' note on M4 asked for: M3 made the head a stored
(branch, depth) and made arriving at one a row lookup plus a state restore,
so a Save Point needs no restore machinery of its own.

What the user gets:

- Name the moment they are reading, keep playing, restart the app, and come
  back to it. Restoring moves the story back and deletes nothing: the later
  turns stay, Redo still walks forward into them, and writing something
  different is what starts a new line while the old one is kept.
- Rename, delete, and a list, in a Save Points panel beside the branch panel,
  with a Save Point button next to Undo and Redo. Both confirmations say what
  is *not* destroyed, because that is the part the screen cannot show.
- Save Points survive export and import.

What was deliberately not built:

- No second restore path. `head.move_to_node` is the only new movement: its
  depth half is M3's `head.move_to` unchanged, and its branch half is the
  single assignment `switch_branch` already makes. No head field is written
  in the checkpoint router, nothing reconstructs state, nothing prunes a
  memory, nothing copies or deletes a turn, and restore never forks — the
  first write below the restored head does, through `fork_if_behind_head`.
- No automatic cleanup. A Save Point behind the head, or naming a line the
  story left, is doing its job (STORY-BRANCH-SEMANTICS §19). The one removal
  is a cascade: deleting a branch takes its Save Points, as it takes its
  memories, because the story they named went with it.
- No new ADR. ADR 012 already decides the architecture, and a table is not a
  decision.

The one call the planning package did not already make: restore moves the
branch half of the head only when the coordinate is off the path being read.
Doing it unconditionally would quietly hand back an abandoned continuation
whenever a Save Point in a shared prefix was restored; never doing it would
make a Save Point on a departed line unrestorable, which contradicts §19.
TECHNICAL-DESIGN §8.8 records it.

Schema: a `checkpoints` table holding a name, an optional note and a
(branch, depth) coordinate — no copy of any story. `create_all` builds it as
it did `memories` and `branches`; migration 80 adds the index. No backfill,
because nobody had named a position before M4.

The coordinate is deliberately not an action id: one coordinate holds every
attempt at a turn and exactly one is live, so a coordinate follows a retry
where a row id would pin a take the story no longer tells.

Tests: 680 pass (638 before). 42 new in tests/test_save_points.py covering
D11-D14, I04, L03, E-series lineage and memory isolation after restore and
divergence, the edge cases, and an M3-database migration. One pre-existing
fixture in test_tree_migration.py needed `checkpoints` added to its drop
list — SQLite refuses to drop a table another table references.

Not verified: the browser. No session has had a usable one, so the Save
Point panel's DOM behaviour is unobserved — as M3's Redo control still is.
The twenty-step sequence was driven over HTTP against a live server with a
real process restart instead, and all seventeen checks pass. M4 is
implemented, not accepted: no review has been written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-03 18:48:54 -04:00
JesseMarkowitzandClaude Opus 5 d27ee34901 Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was
authoritative. Phase 0 execution prompts sat beside the specification; four
completed milestone reports sat beside the current one; and upstream AI-DnD's
own `plan/` build log and `docs/` project site still described a hosted,
scripted, multi-user product with accounts — every screenshot in it showed a
Scripts tab and a Sign up button, none of which has existed since M2.

`planning/archive/` now holds the history and says so in its own README:
`phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and
M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied.
`planning/reports/` holds only the current milestone's report, because that is
the one M4 planning has to read; it moves to the archive when M4's replaces it.

Deleted rather than archived: the Phase 0B execution prompts and the
handoff/status/summary documents, the Phase 0A discovery and triage reports,
upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's
template boilerplate. All of it is in Git history, and the two recommendation
reports carry every conclusion the deleted research reached.

Archived documents are kept verbatim. Paths written inside them point at where
those files were when the document was written, which is the point: an evidence
record that has been quietly edited is no longer evidence.

Active documentation is corrected where it pointed at the removed trees or
described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not
touch" list had gone stale at M2 and claimed QuickJS scripting was still tested;
its test count was 604 against an actual 638. `README.md` loses the upstream CI
badge, which reported upstream's pipeline rather than this fork's, and a
reference to `backend/app/worldstate/engine.py`, a file that does not exist.
`planning/README.md` is rewritten as the documentation index.

New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
manifest of what belongs in the ChatGPT project's Sources.

Source comments referring to the deleted trees are reworded; no behaviour
changes. 638 backend tests pass, frontend lints and builds, and a reference scan
over all 48 tracked Markdown files reports no unresolved path in active
documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
2026-09-03 14:33:07 -04:00
JesseMarkowitzandClaude Opus 5 c8755c21c2 Planning: close M3 and record active-head architecture
M3's review recommended planning changes and, following the M2 pattern,
reported rather than applied them. This applies them, and adds the ADR the
review asked for.

ADR 012 records the architecture rather than the requirement. ADR 005
already says that going backward must preserve abandoned history and that
the user sees Undo/Redo/Retry rather than branch management; it names a
movable active head as the direction and stops. What M3 settled is the
shape: the head is stored rather than derived, every read of the story is
capped at it in one place, one mechanism moves it, the state of a position
comes off the node rather than from a replay, the first write below a
moved-back head is the divergence, and whether Redo exists is decided by
the lineage rather than by a flag that could be stale. The last of those
is the property worth keeping — a flag can be wrong and make the story
wrong; a lineage cannot.

Two semantics are ratified in STORY-BRANCH-SEMANTICS.md, both of them
reversals or narrowings that a reader would otherwise take for bugs. Undo
now crosses fork points and continues to the campaign opening, because
refusing at the fork was a consequence of deleting rows the parent line
was also reading, and nothing is deleted any more. And the system refuses
to switch which take is live while a later story is off screen, because
doing it quietly would leave retained history continuing from words the
story no longer says.

A new §14A covers editing in place. §14-15 describe the finished
behaviour — the edit becomes authoritative, the state it implies is
re-evaluated, a new continuation is created, the original is retained —
and that requirement is intact and explicitly not weakened here. It is
also not built, because re-evaluating state from prose a user typed needs
M5's extraction pass. §14A says what exists in the meantime and why
refusing is the minimum that holds the invariant rather than the
destination.

TECHNICAL-DESIGN.md gains §8.7 and §9.1, recording the implemented model
and the bundle behaviour as fact in the way §5.2 records M1 and M2. §10.4
gains a constraint that is easy to lose: the snapshot half of the hybrid
state model is a requirement, not an optimization. Head movement is a row
lookup plus a restore, which is why Undo, Redo and Save Point restore cost
the same at any distance into a campaign; a state model recoverable only
by replaying from the opening would make all three proportional to
campaign length, on exactly the long campaigns this product is for.

DATA-MODEL.md records the head as stored on the campaign rather than
derived from its newest turn — two campaigns holding identical turns can
be read at different places, and nothing about the turns can tell them
apart — and the branch disposition as implemented: the depth a divergent
write left the branch at, deliberately advisory, and carried through
export because every row of an abandoned line is exported either way.

BUILD-MILESTONES.md marks M3 complete and states the one condition still
open. M4 is told a Save Point is a durable pointer and that restoring one
is head movement with a bounds check, not a restore system: a second
mover is the specific failure to avoid, because the two paths would
silently disagree about what restore means. M5 gets three constraints —
keep state efficiently recoverable, move the test instrumentation rather
than the assertions when the world-state protocol goes, and finish the
narrator edit §14A defers.

V1-ACCEPTANCE-TESTS.md clarifies ownership without lowering a bar. D10
keeps all three pass conditions and is explicitly recorded as *not*
satisfied at the end of M3; what changed is that the document now says
which milestone delivers which condition. D03's result is recorded as a
full pass rather than the partial the text allowed for, I07 gains the
pre-M3 bundle clause, and L01 gains the note that resolves its apparent
conflict with A05 — a failed turn does advance the head by one, onto the
player's retained input, and that is A05 working rather than L01 failing.

README.md described a different application: a hosted demo, guest
accounts, cloud providers, Postgres, a Render blueprint, an analytics
dashboard, a QuickJS scripting engine, and 549 tests. M2 removed all of
that and the README was never updated — a gap M2's own debt table missed.
It now describes what this fork is, including the endpoint policy and the
TLS behaviour, and the numbers in it are the current ones.

M3's report is included here as its own evidence record: no separate
baseline report was produced, so it carries the raw counts and runtime
observations as well as the review, and §W records this closeout.

SPECIFICATION.md and SECURITY-THREAT-MODEL.md are unchanged. M3 altered no
product requirement and touched no path in the threat model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u
2026-09-03 13:54:44 -04:00
JesseMarkowitzandClaude Opus 5 2fdd2547f0 Planning: record M2 closeout decisions
M2's review reported six planning recommendations rather than applying them,
three marked before M3. All six are applied here, plus three additions drawn
from the same evidence. No implementation file is touched.

The endpoint policy was the gap that mattered. It is the most consequential
setting in the application — the storyteller sends the player's prose, the
context, the memories and the embedding inputs to whatever address it names —
and it existed only as a module docstring. It is now ADR 011 and a new §10A in
the threat model, which also retires the assumption in §71A that the inherited
guard was a starting point. It was not: AI-DnD's SSRF guard blocked private
addresses to stop a hosted server reaching its own internal network, which is
the exact opposite of what a local storyteller needs. It was removed, not
adapted.

Both documents state the rule as implemented — an allowlist of explicit
local-network CIDRs, every resolved address checked, enforced on save and again
before every outbound request, TLS never traded against it — and both state the
two residual limits plainly rather than implying they are covered: a hostile
host already on the trusted LAN is inside the permitted boundary, and a
rebinding interval exists between the policy's resolution and the client's
connection. Accepted risks, not M3 work.

The CIDRs are spelled out rather than derived from is_private/is_reserved, and
the ADR records why: is_private is true of the documentation ranges and
0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it
refuses an ordinary same-host Ollama on [::1].

TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening
items. A new §5.2 records the M1/M2 architecture as fact rather than intention,
so later milestones inherit what the code does. A new §18.1 carries the lesson
of M2's two regressions: when removing a setting, test a real consumer
construction path; when adding one, prove it reaches the component that uses
it. Both defects hid behind a green suite because the tests at that boundary
were mocks.

BUILD-MILESTONES records M2 complete, with the capabilities later milestones
inherit and the debt carried forward. Two notes go to milestones that would
otherwise misread what M2 left them. M5 is told that eight rollback tests now
use the world-state engine as instrumentation and not as endorsement — the
instrumentation moves when the protocol does, and those tests are reworked
rather than deleted. M6 is told that the memory bank died silently under a
green suite, so background failure must be observable and at least one real
provider-construction path must be tested.

The security contract gains what M2 demonstrated. H10 now names the two
conditions that were defects during M2: a wildcard origin must be refused at
startup, and an unknown /api path must 404 rather than returning the SPA with
200. New H12 covers endpoint enforcement, and its fourth pass condition is the
one that matters — a public endpoint written into the database behind the
settings API must still be refused at the wire. A build passing the first three
and failing that one has configuration validation only.

SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement;
it removed capability the specification never asked for.

The two M2 reports gain appended closeout notes rather than edits. Their
original wording about an uncommitted working tree was true when written, and
the note records what happened afterwards: the six-file correction is 8652fe7,
8c65ae9 remains the implementation commit, and the two were never squashed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
2026-09-03 01:52:03 -04:00
JesseMarkowitz 1a28a9a708 Apply post-M1 corrections to the planning package 2026-09-02 06:03:10 -04:00
JesseMarkowitz 717670afe0 Update planning package after Phase 0B 2026-09-01 20:41:23 -04:00