d63804f22ecbaed80741241a154cdaef82f7b2ed
137
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d63804f22e |
v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
|
||
|
|
432f04100b |
M11 closeout: accept v1 release validation
The browser, offline and identity runs had last been taken on |
||
|
|
96c1bf5ded |
Measure M04 by where the planted turn is, and catch a section of the model's own
The M04 re-run on |
||
|
|
0c7316f951 |
Keep the state section and its proposal out of the story
The first complete M01 run with the memory bank on (
|
||
|
|
f8d401029f |
Stop a turn locking out its own memory bank, and let the long run notice
The first M01 trial with the memory bank on was 26 turns on a GPU host. It
accepted every turn and reported "complete". It also wrote two memories and
no summary, and logged 180 `database is locked` errors, while derived status
still read `idle`.
The cause was a single uncommitted UPDATE. Retrieval bumped each used
memory's counter before the model call, and the turn commits only after the
reply has streamed. SQLite has one writer, so the turn held the write lock for
the whole reply. Every post-turn memory, summary and status write in that
window waited out the five-second timeout and failed. Recording the failure
needed a write as well, and without a rollback first it raised
PendingRollbackError. The loss therefore reached the log and never reached
the status the Insights panel reads, which F08 forbids. The draco run never
hit this because the bank was off there.
- `retrieve_memories` now only reads. `record_use` writes the counters in the
turn's single commit, so a turn that never lands counts nothing.
- The post-turn task's outer handler rolls back before it records a failure.
The harness could not have caught any of this. It read three prompt sections
under names the builder does not use: `memories` (really `used_memories`),
`story_history` (really `history`/`recent_history`), and a `knowledge` prefix
that matched the fixed instruction section instead of the imported passages.
Memory tokens read 0 whatever the prompt held, and the in-history and
in-memories recall checks could never come out true. The labels are now
constants, pinned by a test against a prompt the real builder assembled.
The harness also stops at the first sign of failed post-turn work. It checks
/derived and new server.log lines after every turn, keeps its log position
across --resume, and waits for background work to settle before its final
checks. A run with no memories or no summaries now ends "failed", not
"complete".
Both new application tests fail on
|
||
|
|
fec46f66bb |
Turn the memory bank on for the long run, and refuse one that cannot use it
The first complete hundred-turn campaign did not exercise M01's "summary/memory activation" step. Memory bank and auto-summarize are per-campaign switches that default to off, and m11_long_run never turned them on: summary_tokens and memory_tokens were 0 on every turn, memories_used was empty, and M04's clue was recalled through narrative state alone. The retrieval path M6 built was never asked, and nothing in the evidence said so except a row of zeros. setup now PATCHes both switches on, reads the campaign back, and stops before the first turn if either did not take. memories_in_bank is recorded on every turn, in the final summary and in the recall, and the recall also says whether a summary exists, so which of the two recall paths succeeded is stated rather than implied. The embedding model is now required. Without one the summary pass still writes memories, but memorybank.retrieve answers "No embedding model configured" and returns none -- the same unexercised path in a fuller bank. The harness refuses before it starts a server or claims --out. tests/test_m11_long_run_memory.py drives setup against the real application in-process: the switches are on afterwards, a server that ignores the PATCH is refused before any state is written, the bank count comes from the application and reads -1 rather than raising when it cannot, and a run with no embedding model is refused. The four that exercise setup and the bank count were run against the previous harness and fail there; the premise test (a fresh campaign has both switches off) passes on both, as it should. The 2026-09-10 run in ~/m11-evidence/m01 therefore does not count as M01. It has to be run again on this harness. Backend 1,382 passed, 18 skipped, 0 failed. The eighteenth skip is test_built_spa_fetches_no_fonts_remotely, which wants a built frontend/dist this worktree does not have; it is an environment condition, not a change here. The frontend is untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XKWHt2DXuvqP83cAk6Zq88 |
||
|
|
ef25b0a876 |
Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB |
||
|
|
144406cd48 |
M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was whether any of the earlier evidence meant what it said. M8 measured a deployment enforcing a 4,096-token input window while the application budgeted 16,384. Every request returned 200. What Ollama does with the excess is drop the oldest tokens, and the oldest tokens here are the system block — the narrator's rules and the campaign canon. A hundred-turn certification against that server would have looked perfect and proved nothing, which is why this milestone could not begin with a hundred turns. So the application asks now. Ollama's window is a property of how a model was loaded rather than of the request — sending num_ctx is accepted, ignored, and worse, reloads the model at the server's own default — so the only honest move is to find out and then tell the truth about it. /api/ps reports what a resident model is being served with, /api/show what an unloaded one will load with, both on the same host inference already uses, through the same endpoint policy and the same TLS trust store. A verified window is a ceiling on the budget; an unverified one leaves the budget alone and is recorded as unverified in the turn's own provenance, so an old turn can be asked afterwards whether it was built against a checked window. There is no third behaviour, and in particular no hard-coded 4,096: a number the server did not say would be right on one machine and wrong on the next. The proof that this is doing something is a campaign whose canon sits at the front of the prompt, 120 turns of history, and a 4,096-token window. The canon is still there afterwards and the oldest history is gone. The same campaign built the old way produces a prompt more than twice the window — the defect, reproduced, so the fix is measured against it rather than asserted. Two defects the validation found on its own, and they are the same defect twice: something was true and nobody was told. A manual state correction of four changes with one bad reference applied three, returned 201, and said nothing — while recording the refusal on the audit row nobody reads. It came to light because the identity diagnostic's own fixture was refused that way and the whole run proceeded on a campaign with no scene, which would have read as a model failure. And the narration-length setting moved no number: brief, medium and long each became one English sentence, while the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three. Both now say what they did. The other two post-M8 findings are closed as well. The tab said AI D&D, which no document had ever claimed it did not; it says Interactive Story now, with the open campaign first, and the name is the owner's decision rather than a find-and-replace to something narrower than the engine. After an Undo the reader could not tell where they had landed; the control row now ends with "Moment 11 · later story ahead", from the server's own answer, in the word the transcript already uses, with none of head, branch or depth anywhere near it. The identity diagnostic exists and the root cause does not. That campaign was destroyed, so no cause can be established — what M11 owes the finding is something that can classify the next occurrence, and a diagnostic that makes only the judgements a program can honestly make: duplicate keys, shared names, protagonist drift, state and context disagreeing. Whether prose misattributed a line is left to a person reading it beside its prompt, because a regex cannot read dialogue and one that pretended to would produce exactly the confident wrong answer this finding is about. Its detectors are proved to fire against a planted second Alice. Two entities may still share a display name. That was checked first, as the finding asked, and left permitted: a mother and a daughter, or a stranger giving a false name, are ordinary fiction, and refusing them to guard against a model mistake would refuse the wrong thing. What was missing was that it happened silently. It is reported now. Evidence, not inference: a hundred accepted turns against a real narrator with genuine process restarts; a real browser against the built SPA; a container with no network at all; a campaign moved into a data directory that never existed. Each was discarded and re-run whenever the product changed under it, and the runs that were thrown away are listed in the report with the reason, along with ten defects in the harnesses themselves — because a harness that has only ever agreed with itself is not evidence, and two of M8's five harness defects were masking real ones. No dependency was added, removed or upgraded. No acceptance test was retired, relaxed or reclassified. M11 is implemented and verified; it is not accepted, and there is no release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B |
||
|
|
1013c94eb1 |
M10: the seam for media, and no media
The media extension contract asks for a scene snapshot a future image or video provider could be handed: location, who is present, what they hold, what must stay true, and where in the story it sits. Building one was the milestone's obvious first task, and it was the wrong one. That snapshot has existed since M5. `narrative_state["scene"]` holds the summary, the location, the cast and the coordinate it was written at; a validated `set_scene` event writes it, every position snapshots it, and every head move restores it. It survives Undo, Redo, Retry, divergence, Save Point restore and a process restart because it is the authoritative state rather than a copy of it. So there is no scenes table here. A second scene store would have been a second answer to "where is the story now", with its own lineage rules to get wrong — and the lineage rules are the expensive part, which is the argument for reusing the ones that already work rather than against it. The Scene Packet is derived on read, and its identity is computed from the campaign and the position rather than allocated: the same position yields the same id in another process, after a restart, and after the packet is thrown away and rebuilt, with no row to keep in step. That is the part of a future media_assets table that would be expensive to retrofit, so it is fixed now even though the table is not built. One table, then: visual_profiles, the only thing the contract's scene list asks for that nothing already stored. Campaign-scoped and not per-position, because a character does not change appearance when the story forks — a reader who diverged would otherwise lose their cast, and the same descriptors would land in every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn campaign to say something that never varies. Keyed by the M5 entity key rather than a new identity namespace, and one table for characters, locations and items alike, because a location is an entity with a type and splitting them would reintroduce the genre shape M5 spent a milestone removing. What the packet leaves out is the more interesting half. Not the transcript, and not imported knowledge — none of it, not merely the sources marked hidden. The rule is what the story established at this position, not everything the narrator was told, and drawing it by class is what makes it hold for a secret nobody thought to mark. A hidden Canon source proves it, with a positive control showing the narrator did receive the sentinel the packet does not carry. Once a validated event puts the observer in the room, the observer is in the packet: that is no longer narrator-only knowledge, and a packet that hid it would be hiding the story from itself. The providers are contracts and nothing else. Protocols for image, video, audio, speech and transcription, an empty registry, no adapter, no dependency, no socket, and no media setting to point anywhere — a setting that exists can be pointed at a cloud by mistake. A future provider endpoint must be loopback, stricter than narration's trusted-LAN allowance, because a picture of a scene carries the scene with it. Transcription returns an editable draft with no commit method, so STT structurally cannot bypass the authoritative path. Nothing here can write the story. Not by convention: no module under media/ imports the code that writes state, no media event type exists in the state vocabulary, and every test in the authority suite compares the authoritative document byte for byte either side of a media operation — including one where a provider insists Alice is in a red coat in a corridor, and the campaign goes on disagreeing. One defect, found by the milestone's own tests. M10 first added a migration creating an index that create_all already builds from the column, so an upgraded database ended up with two indexes and a fresh install with one. Comparing the two schemas is what caught it; neither database examined alone would have. The migration is gone rather than renamed, and the right number of migrations for a new table whose indexes are declared on its columns is zero. Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145 passed. Lint, production build and Docker build clean. No frontend file changed: M10 adds no reader-facing surface, and ordinary play — turns, state, memory, knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media configuration, no warning, no connection attempt and no media row written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B |
||
|
|
44edece67e |
M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B |
||
|
|
1ce9972760 |
M8: the browser becomes the storyteller
The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.
All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.
So the shape now is one entry point and one screen:
Campaigns -> Campaign -> Story
State · Knowledge · Context · Save Points · Settings
Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.
Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.
Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.
The two defects worth the space:
A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.
And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.
Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.
`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.
Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.
Backend, and only what the browser could not otherwise reach:
AdventureCreate.opening a start action could only come from a Scenario, so
every campaign made in the new setup flow opened on
a blank page. Same node, same code path.
canon_rules campaign_canon has been the highest authority in a
campaign since M5, read by the prompt builder and
the state validator, and had no API at all — a
fixture had to write it with SQL.
a 401 and a 429 message the last user-facing text describing a hosted
deployment. One told the reader to check an API key
that has not existed since M2.
No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.
The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.
They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.
A verification pass over all of it then found three more, each by driving the
product rather than reading it:
Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.
The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.
And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.
Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.
`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.
Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:
The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.
Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.
The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.
The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.
Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
|
||
|
|
480414efe0 |
M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b |
||
|
|
a6e9c7a32b |
M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against
|
||
|
|
b7005e6fdd |
M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.
This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:
visible active transcript position == stored head == authoritative state
Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)
A narrator edit no longer rewrites a row. It returns to the state before the
turn, takes the reader's exact text as the accepted narration, re-derives the
state that text implies, and becomes a new active continuation — while the
original narration keeps its words, its live flag and its whole future as
retained history. At the tip the correction is another take; with story below
it, it forks. No new history machinery: this is the existing fork/take/head
path with the reader's text in place of a generated reply. The §14A refusal
is therefore gone for narrator turns, and remains only for player input.
Pre-M5 positions
Migration 88 backfills the empty narrative document onto every action written
before M5, and a missing snapshot now restores the empty document instead of
leaving the previous position's state standing. Restoring to an old Save
Point no longer leaves a later position's entities and facts on screen.
Narrator context
Replayed history carries prose only; the machine-readable block is no longer
reconstructed into past turns, where it contradicted the authoritative state
in the same prompt. A fact withdrawn by a manual correction is now named as
no longer true, with the reader's reason, rather than silently dropped.
Also
- state_changes joins the action-list bulk read, removing one query per row.
- Extraction takes only the application's own protocol payload: an ordinary
```json or ```python block in a story survives, and a mangled proposal
still does not reach the reader.
Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
|
||
|
|
62a997f364 |
M4: close out Save Points, with browser verification
Closes M4. The review's three findings are fixed, the durability rule the specification always implied is now enforced, and M3's and M4's browser behaviour has been verified in a real browser for the first time. B-1 -- the Save Point list was an N+1 that loaded whole Action rows, narration included, to answer "does a row exist here". It is now one bulk two-column coordinate query plus one lineage: 53 SELECTs for 25 Save Points became 5, and the count no longer grows with the list. The clause is an OR of exact (branch, depth) pairs rather than two IN lists, because the cross product would report a Save Point resolved on the strength of another one's depth existing on this one's branch. A test builds exactly that trap. B-2 -- reclassified during closeout from "missing warning" to a behaviour defect, and fixed as one. STORY-BRANCH-SEMANTICS §19 says a named checkpoint remains until explicitly deleted, and §28 already required future cleanup to retain checkpoint-referenced paths; a cascade that silently removed Save Points with a branch violated both, and a warning would only have documented the violation. A branch a Save Point names can no longer be deleted. The request is refused with the offending Save Points named, the user deletes them explicitly -- which deletes no story -- and the branch then goes. The scope is the subtree, because deleting a branch takes its descendants. Both delete controls disable and explain. Recorded as a new §19.1; models.py, TECHNICAL-DESIGN §8.8 and DATA-MODEL §8 had all recorded the cascade as the rule and now record the refusal. An earlier pass in this same closeout had kept the cascade and added a warning. That was the wrong fix and its tests were replaced rather than left standing, since they pinned the defect. B-3 -- the D11/L03 automation never left one process, so it could not distinguish durable state from a live Python object. It now spawns real server processes, kills the first, and reads the campaign back with the second. C-5 -- creating a Save Point takes the campaign's turn lock. "Save where I am" has to name one committed position, and the head is what a turn in flight is about to move. Rename and Delete deliberately do not take it. The architecture is untouched: a Save Point is still name + note + (branch, depth), and restore is still coordinate -> head.move_to_node -> head.move_to -> attempts.restore_state. No second restore path, no state copied into a checkpoint, no fork on restore. Browser verification -- the first in this project, and it covers both milestones. Firefox 154.0.1 through geckodriver over the W3C WebDriver protocol, driving the rendered DOM: 47/47 checks, twice, on independent databases, no console errors. M3's Undo/Redo enable states, transcript movement, Retry and the take pager, divergence retiring Redo; M4's whole Save Point lifecycle, both confirmations, and the new branch-delete refusal including its recovery. No dependency was added: the WebDriver client is stdlib HTTP. No application defect was found by the browser. Four failures occurred, all in the harness -- a wrong SPA route, a wait comparing transcript length when the empty-story placeholder is longer than the first turn, a fixture deleting the branch it was reading, and a reload assertion that sampled once instead of waiting. The last was checked against the app before being called a harness bug. Tests: 698 backend pass (was 680), 60 M4, 94 M3 history, 66 export/ migrations, 93 security/local-only. Frontend lint and build clean, Docker build clean, loopback binding unchanged. No assertion weakened, no skip added. Planning: STORY-BRANCH-SEMANTICS §19.1 is the only behavioural change and it strengthens §19. V1-ACCEPTANCE-TESTS records D11-D14, I04, L03 and the E-series, keeping automated, live-runtime and browser evidence distinct, and weakens no pass condition. DATA-MODEL records the coordinate with the retry measurement that settles it. BROWSER-UX-SPEC rules for Moment over Turn. BUILD-MILESTONES marks M4 COMPLETE, closes M3's browser condition, and lists what M5 inherits. VERSION adds v2.6. No new ADR: ADR 005 already decides that history is preserved rather than overwritten, and §19.1 is that decision applied to checkpoint-referenced history. M4 is closed. M5 may now be briefed; it has not been started. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2 |
||
|
|
e08d49c3eb |
M4: add durable named Save Points
A Save Point is a name for a story position, and restoring one is head movement. That is the whole architecture, and it is what ADR 012 and BUILD-MILESTONES' note on M4 asked for: M3 made the head a stored (branch, depth) and made arriving at one a row lookup plus a state restore, so a Save Point needs no restore machinery of its own. What the user gets: - Name the moment they are reading, keep playing, restart the app, and come back to it. Restoring moves the story back and deletes nothing: the later turns stay, Redo still walks forward into them, and writing something different is what starts a new line while the old one is kept. - Rename, delete, and a list, in a Save Points panel beside the branch panel, with a Save Point button next to Undo and Redo. Both confirmations say what is *not* destroyed, because that is the part the screen cannot show. - Save Points survive export and import. What was deliberately not built: - No second restore path. `head.move_to_node` is the only new movement: its depth half is M3's `head.move_to` unchanged, and its branch half is the single assignment `switch_branch` already makes. No head field is written in the checkpoint router, nothing reconstructs state, nothing prunes a memory, nothing copies or deletes a turn, and restore never forks — the first write below the restored head does, through `fork_if_behind_head`. - No automatic cleanup. A Save Point behind the head, or naming a line the story left, is doing its job (STORY-BRANCH-SEMANTICS §19). The one removal is a cascade: deleting a branch takes its Save Points, as it takes its memories, because the story they named went with it. - No new ADR. ADR 012 already decides the architecture, and a table is not a decision. The one call the planning package did not already make: restore moves the branch half of the head only when the coordinate is off the path being read. Doing it unconditionally would quietly hand back an abandoned continuation whenever a Save Point in a shared prefix was restored; never doing it would make a Save Point on a departed line unrestorable, which contradicts §19. TECHNICAL-DESIGN §8.8 records it. Schema: a `checkpoints` table holding a name, an optional note and a (branch, depth) coordinate — no copy of any story. `create_all` builds it as it did `memories` and `branches`; migration 80 adds the index. No backfill, because nobody had named a position before M4. The coordinate is deliberately not an action id: one coordinate holds every attempt at a turn and exactly one is live, so a coordinate follows a retry where a row id would pin a take the story no longer tells. Tests: 680 pass (638 before). 42 new in tests/test_save_points.py covering D11-D14, I04, L03, E-series lineage and memory isolation after restore and divergence, the edge cases, and an M3-database migration. One pre-existing fixture in test_tree_migration.py needed `checkpoints` added to its drop list — SQLite refuses to drop a table another table references. Not verified: the browser. No session has had a usable one, so the Save Point panel's DOM behaviour is unobserved — as M3's Redo control still is. The twenty-step sequence was driven over HTTP against a live server with a real process restart instead, and all seventeen checks pass. M4 is implemented, not accepted: no review has been written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2 |
||
|
|
d27ee34901 |
Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was authoritative. Phase 0 execution prompts sat beside the specification; four completed milestone reports sat beside the current one; and upstream AI-DnD's own `plan/` build log and `docs/` project site still described a hosted, scripted, multi-user product with accounts — every screenshot in it showed a Scripts tab and a Sign up button, none of which has existed since M2. `planning/archive/` now holds the history and says so in its own README: `phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied. `planning/reports/` holds only the current milestone's report, because that is the one M4 planning has to read; it moves to the archive when M4's replaces it. Deleted rather than archived: the Phase 0B execution prompts and the handoff/status/summary documents, the Phase 0A discovery and triage reports, upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's template boilerplate. All of it is in Git history, and the two recommendation reports carry every conclusion the deleted research reached. Archived documents are kept verbatim. Paths written inside them point at where those files were when the document was written, which is the point: an evidence record that has been quietly edited is no longer evidence. Active documentation is corrected where it pointed at the removed trees or described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not touch" list had gone stale at M2 and claimed QuickJS scripting was still tested; its test count was 604 against an actual 638. `README.md` loses the upstream CI badge, which reported upstream's pipeline rather than this fork's, and a reference to `backend/app/worldstate/engine.py`, a file that does not exist. `planning/README.md` is rewritten as the documentation index. New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the manifest of what belongs in the ChatGPT project's Sources. Source comments referring to the deleted trees are reworded; no behaviour changes. 638 backend tests pass, frontend lints and builds, and a reference scan over all 48 tracked Markdown files reports no unresolved path in active documentation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu |
||
|
|
7f082b61d8 |
M3: complete non-destructive history and active-head export
The head-cursor model landed in
|
||
|
|
903fa7a74f |
M3: move the story's head instead of deleting its turns
Undo deleted. It removed the trailing AI action and the player action in front of it, pruned the memories covering them, and let the tip fall back to whatever survived. That made it the one operation in the application that destroyed accepted story, and it was why there was no Redo: the turns to move forward into no longer existed. Phase 0B demonstrated the head-cursor alternative in a disposable spike; this is that concept as production code. backend/app/head.py is the whole of it. Three questions that used to be one — where the story is being read, how far it is retained, and where it opens — are now three functions, and every caller that moves the head or asks about it goes through this module. The spike put the fork check in the write path and left Retry and Add-take on the old one; sharing the rules is what stops that divergence coming back. lineage.Path now caps every entry at the head, so hiding the retained future costs nothing at the call sites: the transcript, the context builder, attempts.preceding and memory retrieval already funnelled through path_of and narrow together. Path.uncapped() is the deliberate exception, and only Redo and the fork check may use it. The memory bank needs no pruning for the same reason — a memory carries the coordinate of the node its block ends on, so one derived past the head falls outside the capped clause and becomes retrievable again on Redo without having been deleted and re-embedded. Undo alone does not fork. Moving the head is not a decision to abandon anything, since the user may be reading or about to Redo; the first write below the head is where the story states which continuation it means. A head already at the tip forks nothing, so a story that is never undone forks exactly as often as it did before and the branch table does not fill up with one branch per turn. Redo follows the lineage rather than choosing among branches, which is what invalidates it after a divergence with no flag to set or clear. Migrations 78 and 79 give a branch superseded_at and superseded_depth. Nothing reads them to decide behaviour — Redo is decided by the lineage, so a stale or hand-edited value here cannot make the story wrong. They exist so the cleanup and discarded-history features left to a later version have something to select on, and so a divergence is observable in a test. Deleting an action no longer drags a moved-back head forward to the recomputed tip, which would have silently redone the story. can_undo and can_redo ride on AdventureOut and ActionPage because the client can work out neither for itself: the campaign opening may be off the top of the loaded window, and the retained future is never sent to it. This is a checkpoint, not the finished milestone. 601 backend tests pass. Five still assert the destructive contract — they count rows after an undo and expect the story to be shorter — and need rewriting against the new one; the world-state assertions inside them already pass. Export and import do not yet carry the head coordinate, so a bundle still reopens at the deepest node and can silently redo an undone story, which is the Phase 0B finding this milestone exists to close. The browser has no Redo control yet. None of the M3 acceptance coverage (D01-D10, E01-E04, I01-I03, I07, L01-L02) is written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u |
||
|
|
8652fe7cd8 |
M2 review: two regressions the green suite hid, and the reports
The post-implementation review of M2, plus the three corrections it took
to make the evidence true. Reports:
planning/reports/M2-BASELINE-REPORT.md 868 lines, the measurements
planning/reports/M2-IMPLEMENTATION-REPORT.md 758 lines, the reading of them
Verdict is PASS, accept with non-blocking debt, proceed to M3. Every M2
requirement is met and the ones that matter were tested by running the
build rather than reading it: a cloud endpoint written straight into
SQLite with sqlite3, behind the API's back, still refused at the wire;
trusted-LAN HTTPS against the real second machine with verification on;
captures showing zero packets outside loopback and the approved host.
Three defects, all found by running the shipped image.
The memory bank was dead. M2 removed Settings.api_key_plain with the API
key, and memorybank's two provider factories still read it. It failed
inside a fire-and-forget task, so no user error, no log anyone would
read, and no test — every memory test stubs those factories. All 604
tests passed with summaries and embeddings silently not happening.
The configurable model timeout never reached the turn engine. Stored,
validated, exposed in the API, rendered in the UI, and not passed to the
provider. M2's own exit criterion was half met: the constant had moved
but the setting did nothing.
And requirements.lock still pinned quickjs, psycopg and cryptography, so
the setup path DEVELOPMENT.md gives a new developer would have
reinstalled all three.
Both code defects now have the test that would have caught them: one
constructs every provider factory from a real Settings row, one drives
the turn endpoint, the chat endpoint and the summariser and asserts the
configured timeout arrives at each. That is the lesson worth keeping from
this milestone — after removing an attribute, build each consumer from a
real object; after adding a setting, prove it lands. Both failures were
in background or plumbing paths, which is exactly where a subtractive
change cannot see itself.
606 tests pass, up from 604. Lint, build and image are clean. Every
runtime result in the baseline report came from an image built after
these fixes; the reports say plainly that commit
|
||
|
|
8c65ae99de |
M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo. |
||
|
|
c1a73b3d77 |
M1: make the first story turn work with no Internet
Phase 0B ran the upstream application on a network with no route out and the first turn died in tiktoken, which downloads its BPE table the first time anything counts a token. The browser separately fetched three font families from Google on every page load. Neither is visible on a machine that has been online once, which is why both now have tests. The tokenizer table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it directly, verifying its SHA-256 against the digest tiktoken itself pins for that URL. No code path in the tokenizer can reach the network any more — not a warm cache, not an environment variable a deployment could forget. The encoding was checked token for token against tiktoken's own. The three font families are self-hosted as variable fonts under frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in frontend/src/styles/fonts.css, and re-vendored by frontend/tools/vendor_fonts.py. Their OFL licences ship beside them. With no remote asset left, the CSP drops both Google hosts and gains object-src, base-uri and form-action; woff2 also gets its real media type, which Python's table lacks on a slim image. A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx verifies against the certifi bundle, so an endpoint whose certificate comes from a CA the user installed on their own machines — a StartOS server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED while curl and the browser on the same host accepted it. app/tlstrust.py builds one context that unions the platform CA store with certifi's, and all four outbound clients use it. A union rather than a swap, so an image with an empty system store cannot start failing on endpoints that worked before. Verification itself is untouched: CERT_REQUIRED, hostname checking on, and no insecure escape hatch. The storyteller listener is now loopback by explicit statement rather than by inheriting uvicorn's default: start.sh, start.ps1, and docker-compose.yml, which publishes to 127.0.0.1 rather than every interface. Reaching an Ollama on another machine is outbound and needs none of that inbound exposure. backend/requirements.lock pins the exact tested closure; requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the same-host and trusted-LAN Ollama configurations, and how to re-run the offline proof. PROVENANCE.md records the upstream commit, the MIT terms, and both vendored assets. Verified, not just compiled. On an --internal Docker network with 1.1.1.1 unreachable and no name resolving, a campaign was created and played for six turns through same-host Ollama, restarted, and resumed. A second run played ten turns through Ollama on a separate physical machine on the LAN over verified HTTPS, summaries and embeddings included, with the storyteller's default route deleted so the LAN was reachable and the Internet was not. Its capture: 893 packets to the approved host, 730 loopback, zero anywhere else, and zero DNS queries. Two induced model failures left the accepted story bit-identical. The inherited SPA was opened in a browser and a campaign read back from it. Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the findings that did not belong in this change. 648 backend tests pass, up from the inherited 632; frontend lint and build are clean; the image builds. No M2 work is included: the hosted, cloud, analytics, Postgres and scripting surfaces are untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL |
||
|
|
d72f7c1bda |
Stop paying twice for a block a retry can still throw away
A memory whose block ends on the newest action is the one memory a player can reach: retry and take-switching both refuse anything else. Each retry of that turn withdrew the memory and wrote it again, and a block closes every six actions while a normal turn writes two, so that was one turn in three. SETTLE_SLACK asks for one action past a block before the block is summarized. The block is still MEMORY_INTERVAL actions; only the moment moves. This is not the pre-SP4 holdback returning: that one was about a retry rewriting text in place, which sibling attempts and forget_node settled, and correctness still rests on the withdrawal rather than on the slack. Undo and delete can carry a summarized node back to the tip, so the withdrawal path stays reachable, just rarely. The slack buys nothing back. The block that just closed is still in the history window in full, so a memory of it says what the model can already read. Tests: the settling suite asserts the new rule and that a retry at the tip finds nothing to withdraw; the rewrite suite builds thirteen actions so both of its blocks settle; the one path test that ended on a block boundary sets the slack to zero, because it is about which actions a block is read from rather than about when a block forms. 632 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NcrxCJjqgDvAkamKeLWdn |
||
|
|
745a4ea9f3 |
Stop printing the database password
The report opened with the connection string it was about to work on, taken straight from AIDND_DATABASE_URL. On the hosted deploy that string carries the Neon password, so the first line of every run put a live credential into the console — and from there into scrollback, a screenshot, or a pasted bug report. Nothing in the output said it was there to notice. It now prints scheme, user, host and database name. The query string goes whole: sslmode is the only part worth reading, and some drivers accept a password there as well. 631 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
eeab8ef5bb |
Survive the console this will actually be run from
The report prints an arrow between the old and new word counts, and it prints the memories themselves, which are model-written prose. A Windows console defaults to cp1252 and cannot encode either, so the run dies on UnicodeEncodeError partway through — after some memories have been rewritten and committed, which is the worst place to stop. STATUS already carries the same warning for tools/dbmeter.py. stdout is reconfigured to UTF-8 with errors="replace" instead, so the report degrades a character at a time rather than failing. The usage notes were also written for a POSIX shell, and this project is developed in PowerShell, where `VAR=x command` is not a thing. Both docs now give the PowerShell form first, via Read-Host so the Neon URL and the secret stay out of the history file, and name the venv's Python. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
4c94dd0fc0 |
Keep a production run to the adventures you meant
The hosted database is not the local one. It holds other people's stories, and `summary_provider` builds from the adventure owner's Settings, so an unfiltered `--write` against it would spend other people's money rewriting memories they never asked about. `--adventure` could already hold a run down, but only if you knew the ids. `--email` names accounts instead, and every line of the report now says who owns the adventure, so a dry run answers "whose keys would this spend" before anything is written. Guests have no email and stay reachable only by id, which is the right amount of friction for touching a stranger's bank. The other half is written down rather than built: stored API keys are encrypted with AIDND_SECRET_KEY, so a run from a checkout against the Neon database needs the same value the web service holds. With a different one `decrypt_secret` returns "" instead of failing, and every adventure is reported as having no key — a run that looks like it worked and did nothing. The Dockerfile copies backend/app alone, so tools/ is not on the box either way; the recipe is a checkout pointed at AIDND_DATABASE_URL. 629 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
0633cb624e |
Run the new memory prompt back over an old bank
The prompt change only reaches memories written after it. plan/18 decided to leave the existing ones alone and let eviction age them out at memory_bank_capacity, on the grounds that re-summarizing would duplicate whatever was still in the bank because nothing deletes the old rows. That was wrong about the only option. A memory can be rewritten in place. The row carries more than its text — whether it is pinned, how often it has been retrieved, and the node it hangs off, which is what makes a fork inherit the right memories — and rewriting `text` keeps all of it. Deleting the bank and rewinding the cursor would lose that, and would trickle memories back at MAX_MEMORIES_PER_RUN per turn, so an adventure nobody is playing would never recover. tools/rewrite_memories.py does it. Without --write it makes no model calls and only reports the scope; --write rewrites, --embed re-embeds in the run rather than leaving it to the app's post-turn pass. It reads whichever database the app reads, so it works against the hosted Postgres as well as a local file. Two things it needed from the app. `summarize_block` is now the one place a memory prompt is assembled, and the post-turn pass calls it too — a backfill that built its own prompt would be writing memories with a prompt that never shipped, and nothing would report the drift. `source_block` reads a memory's block back out of the story, which nothing has ever had to do: it reads on the lineage of the branch the memory was written on, not the branch being played, because after a fork the same depths hold different actions on each side and a read through the adventure's path would summarize the wrong story silently. It also excludes the discarded attempts at a retried turn, and tolerates a block an action has since been deleted from. Left alone: a memory with no source range, which is hand-written or migrated by 62 and may be the player's own words; a memory whose actions are gone; and an adventure whose owner has no API key, because summarization spends the user's own key by construction and never the shared demo key. --api-key/--model/ --endpoint override that, the last of them aiming a run at claude_shim.py. The vector is cleared for every rewrite, because the stored one describes wording that no longer exists. Re-embedding always uses the owner's own embedding model, never --endpoint: a vector only means anything against the vectors it is ranked beside. 17 tests, 627 green. The fork case is the one that would fail quietly, so the test builds a fork whose depths hold different actions on each side. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
712ef44f57 |
Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and would have been gone with the container. The numbers in plan/18 were therefore assertions nobody could check. plan/18-appendix-memory-ab-run.md now carries the whole transcript: both memories, both summaries, and the thirteen-action story they were written from. backend/tools/memory_ab.py reproduces it. It replaces the throwaway script the first run used, and differs in two ways that matter. It goes through OpenAICompatibleProvider rather than calling a model directly, so a run exercises the provider, the streaming path and complete() instead of a stub. And it reads the control prompt out of git at the commit given to --before, so the thing being compared against cannot drift from what actually shipped. There was already a claude_shim.py serving an OpenAI-compatible endpoint backed by the CLI, which is exactly what the throwaway script had reinvented. memory_ab.py points at it by default, so a run spends a Claude subscription rather than API credit, and --endpoint aims it at the provider the deployed app really uses — which is the one question this whole exercise could not answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
07192767d8 |
Give a memory a word ceiling, after measuring one
Ran the memory prompt end to end against a real model as a controlled A/B: one story generated through the app's own build_context a turn at a time, then both prompts run over the same blocks, so the story is held constant and the prompt is the only variable. Fresh process per call, so neither arm sees the other and the model is never told what is being tested. The control is the exact prompt from 9cdcb55. The reported fault reproduced. Two consecutive memories written from one story minutes apart came back in two different persons — "You crept low through the mist" and "The player asked Gwen to". With the cast brief both named Kaelen. The control also inverted who acted on a move whose player text was "grab her wrist and pull her down", which is the failure the brief predicts: with no cast there is nothing to say whose wrist "her wrist" was. It also found something the prompt review had not. "1-2 plain sentences" is not a length, and the same model wrote 34 words for one block and 105 for the next. A 105-word memory is a paragraph, and `memory_top_k` injects five every turn, so the bank's running cost was set by a number nobody had ever stated. MEMORY_MAX_WORDS states it, and the prompt now says which details to keep when trimming: the ones a later scene could turn on. Re-run over the identical story, the same two blocks came back at 32 and 58 words, still named, still third person, still carrying the camp map, the strongbox behind the second tent, and the strap frayed near through. Variance is the real gain — 34..105 became 32..58. Overshooting 50 slightly is expected. Models exceed word budgets, which is why builder.length_hint already carries a buffer for the same reason. What this does not show: the run used a Claude model, and the app talks to an OpenAI-compatible endpoint whose weaker models are why worldstate/parse.py tolerates trailing commas. The prompt is followable and the brief supplies the missing information; a weaker model is not proven to comply as well. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
1b5e4c56cd |
Tell the summarizer who the characters are
`_create_due_memories` sent six actions of second-person prose and nothing else — no protagonist, no cast, no setting, and no instruction about what person to write in. So for `You push the door open. She grabs your arm.` the only honest memory was "You entered a room and she stopped you", which names nobody when it is retrieved forty turns later. The framing wandered too: with no rule, the model picked a person per call, and one bank ended up holding "You entered the crypt", "The player entered the crypt" and "He entered the crypt" for the same kind of event. Both prompts now carry a cast brief and a framing rule: third person, the protagonist by name, other characters named rather than left as bare pronouns. The rule states its reason, because a memory really is read in isolation much later and a model told why complies far more consistently than one handed a bare instruction. The cast comes from the story cards, not from `stat_schema`. Every schema NPC is already turned into a card at adventure creation, deduplicated against the hand-written ones by name, so the cards cover schema NPCs, an author's own cards, and an adventure with no RPG layer at all through one path instead of three. Keyword matching alone was not enough, and finding that out changed the design. Built that way first, the brief for "She grabs your arm" listed the protagonist and nobody else: the block that most needs a cast is exactly the one written in bare pronouns, and Gwen's trigger keys include "her" but the text says "she". So matched cards come first and the remaining slots are filled with the other character cards. Places and items are not topped up — an unmentioned tavern is not who "she" was — though a place that is mentioned still matches normally. The asymmetry with the turn prompt is deliberate: an untriggered card is wrong as lore and right in a roster, because the roster answers "who could these pronouns be" rather than "what is on stage". Fixed descriptions only, never live values. `Gwen: trust 40 (wary)` in the brief would make the same event summarized at two different times come out framed differently, which is the fault this removes. An adventure with no persona still gets the cast and the setting, and the model is told to write "the player". One with nothing to say sends byte for byte the prompt it sent before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
9ee052c51e |
Give the adventure a persona, so the protagonist has a name
The player had stats but no identity. `stat_schema.player` carried hp and
mana beside `npc.gwen.trust`, but where an NPC has a name and a
description the player had neither, so the block rendered as
`You: hp 100/100` and nothing in the prompt said who "you" was.
Three columns on `adventures`: name, pronouns, description. All
user-only, all optional, and an empty name means the app behaves exactly
as it did before — no backfill, no special case for an adventure that
predates the migration.
They are adventure columns rather than part of `stat_schema` for two
reasons. An adventure with no RPG layer still has a protagonist, and
that is the case this was added for. And `worldstate.schema._initials`
treats every dict inside a stat section as a stat definition, so a
persona placed there would be instantiated, rendered in the guide, and
handed an `initial` value as though it were one.
The paths do not change. `player.hp` stays `player.hp`; only the label
moves, to `Kaelen (player): hp 100/100`, the same way NPC lines already
print a display name beside the id. A path carrying the persona's name
would break the moment a player renamed their character, because
`_history_text` replays every past turn's stored delta into the prompt
and those blobs hold literal `player.hp` strings.
The section sits in the system block. Only the user can edit it, so it
never changes mid-story and stays inside the cached prefix. That is what
makes it free, and it is why the AI must not be able to move it — a
delta aimed at `persona.*` is already refused by `_resolve`, and there
is now a test holding that in place.
The modal that used to appear only for scenarios with `${Placeholder}`
tokens now always opens, and is where the character is named. Persona
and placeholders stay independent: a scenario asking for `${Name}` is
asking its own question. No scenario in the repo uses placeholders at
all, so the overlap is hypothetical.
Phase 2, which feeds the persona and the cast to the summarizer, is
written up in plan/18 and not started. That is where the memory-quality
problem actually gets fixed; this change is what gives it a name to use.
Not yet driven in a browser — plan/18 lists what to check by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
|
||
|
|
6a5d8e98f9 |
Give the stress fixture back its sibling attempts
`place_action` used to read `depth` from `index`, so the fixture's retry attempts all landed on the turn's own depth and formed one sibling group. Now that depth comes from the head, each attempt moved the head one step and every retry became a turn of its own. The fixture's own check caught it: a coordinate holding a single superseded attempt has no live row, and that turn disappears from the story. The attempts at one turn share that turn's coordinate, so only the first one goes through `place_action` and the rest copy its placement. This is what `attempts.add_attempt` does. The fixture cannot call it directly, because it makes the newest attempt live and the fixture needs a live attempt that is often not the newest. Also drop the `variant_index` and `variant_count` arguments, which raise `TypeError` now, and read the two mark positions from a helper instead of from the adventure columns migrations 72 and 73 removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV |
||
|
|
949f58a054 |
Sweep the last three uses of the columns SP8 dropped
`seed_demo.py` still passed `index=0` to `models.Action`, which raises `TypeError` now that the attribute is gone. Every call site inside `app/` was updated when the column went, but the seed script sits outside the package and was missed. `bundle.settle` wrote `memory_cursor` and `summary_cursor`, which are no longer mapped, so the assignments only set transient Python attributes. The version 2 branch existed solely to make those assignments, and it ran two queries per cursor to do it, so it goes. `settle` now returns early when the bundle carries anchors. The `db` parameter is unused after that. Also move the comment about deleted memories next to the `db.delete(branch)` it describes. Splitting the router package left it after the `finally`, where it read as attached to nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV |
||
|
|
20f0753277 |
Merge main into the Phase 17 refactor
`main` changed the four files this branch split into packages, so all four came back as modify/delete conflicts. Each change is ported to the module that now holds the code, unchanged in behavior: - `delete_action` moves to `routers/adventures/actions.py`. It reaches the turn lock as `turns.acquire_turn_lock` and `turns._active_turns`, which is the rule this branch set: the package root no longer re-exports the names a test rebinds. - `EMIT_RULE` moves to `worldstate/parse.py`, and `_describe_stat` and `render_reference` to `worldstate/render.py`. - The chip-wrap rules move to `styles/schema-editor.css`, and `StateChangeChips` to `pages/Play/reports.jsx`. `index.css` keeps this branch's import list. `test_delete_state.py` needed four edits to run here. It drops the eight-line database prologue, because `conftest.py` does that once now. It imports the shared `ScriptedProvider` from `fakes.py` instead of carrying a copy. It patches `adventures.turns.OpenAICompatibleProvider` and reaches the lock the same way. And it no longer passes `index=0` when it builds the start action, because migration 71 drops that column. 564 tests pass. `npm run lint` and `npm run build` are clean, and the CSS bundle is 56.71 kB, which is the pre-merge size plus main's new rules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
91907a30bd |
Name every stat in the guide by the path the AI has to write
The referee refused `npc.<id>.status` on a character that has a status stat under another name, and the refusal was right: the model was reaching for a path the guide never gave it. The guide is the fixed list of everything a scenario tracks, and it named stats in prose. A player stat and a world stat both read as a bare name, and an NPC's stats read as the display name plus the stat — "Trainer Milo active_status". The model had to turn that back into `npc.milo.active_status` itself, and in the Pokemon demo five other characters carry a stat called `status`, so the path it built was `npc.milo.status`. The live values do state the paths, but only for the NPCs a scene has mentioned, so an NPC off screen was addressable only by guesswork. Each line now leads with the path: `player.potions`, `npc.milo.active_status`, `flags.sandstorm_active`. An NPC's header is written even when it has no description, because it is the one line that ties a display name to its id. `EMIT_RULE` points at the guide for paths rather than at the live values alone. A free-text stat is marked `(free text)` beside its path, which is the wording `EMIT_RULE` already used to describe it — the guide had been writing "free text" at the end of the line instead. Such a stat now also gets a line when it has no description, where before it was dropped and the model was left to send a number for a stat that holds a string. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST |
||
|
|
41ed2f63b5 |
Put the shared state back when a turn is deleted
Deleting an AI reply removed the text and left everything the turn did to the numbers standing. `script_state` and `world_state` belong to the adventure rather than to the node, so the row going away takes nothing back. Undo, retry, a take and a branch switch all restore them; this endpoint never did. The visible symptom was the cooldown clock, which lives in `world_state._meta.last_changed` and holds a depth. A deleted turn left its depth there, and the turn played in its place is played at that same depth, so the referee refused the change as one that had happened this very turn — on a turn the story no longer contains. Delete the reply because you did not like the stat change it proposed, press Continue, and the same change comes back marked "changed too recently". A script's state stacked the same way: the gold a deleted turn paid out stayed paid, and the replacement turn paid it again. The restore reads the tip's own outcome rather than the deleted node's neighbour, which is what `switch` does. Deleting the newest turn then rewinds to the node in front of it, and deleting one from the middle of the story restores the state the adventure is already in, so the text goes and the numbers stay. The endpoint now takes the turn lock too, for the reason undo takes it: it writes the shared state, and a turn that is still generating is about to write it as well. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST |
||
|
|
f1bebe18d0 |
Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`, `variant_count`, `state_before`, and `world_state_before`, plus `adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in SQLite, so migration 71 quotes it. Nothing outside the migrations read these. `models.py`, `schemas.py`, and `ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders by `id`, and `attempts.renumber`, `context.history.max_action_index`, and `nodes.next_index` are deleted. Two changes keep the migration replayable on a `create_all` database: - `_split_variants_into_siblings` wrote through the live ORM table, so it stopped compiling once migration 66 removed five of its columns. It now writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`. - Five data passes read columns these migrations drop. Each now calls `_has_columns` and returns early when the columns are absent. `bootstrap` takes a `through` version so a migration test can stop at the schema it asserts on. 555 tests pass, up from 549. Eight of the new cases assert each column is gone after a real schema-45 database migrates all the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK |
||
|
|
e0bf2b61d9 |
Add a useDebouncedSave hook and delete Settings.stream
Two items from Stage 2 of `plan/17-refactor.md`. `frontend/src/hooks/useDebouncedSave.js` replaces three copies of the same debounce. `PlotPanel` and `ScenarioEditor` held identical per-key timer maps. `ScriptEditor` held a single shared timer, so editing two fields inside the same 600 ms window canceled the first save. The hook gives every key its own timer, which fixes that. `Settings.stream` was dead state. Nothing read it and every turn streams. This removes the column, both schema fields, and adds migration 65 to drop it. It is item S1 in `docs/self-review.md`. Migration 65 needs a new guard. `_column_already_gone` is the counterpart to `_column_already_there`: `create_all` builds the current schema, which is already missing every dropped column, so a fixture that stamps an old version and replays would fail on a column that is not there. Verified: 549 backend tests pass, lint and build are clean, and migration 65 runs both ways, once against a database that still has the column and once against one that does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
2c57b1ceab |
Remove three kinds of duplication in the backend
Stage 2, items 1, 4, and 5 of `plan/17-refactor.md`. **One path resolver in `worldstate`.** `apply_delta` and `apply_override` routed `flags.<name>`, `milestones.<id>`, `world.<stat>`, `player.<stat>`, and `npc.<id>.<stat>` with parallel code, about 100 lines each. `_resolve` now says what a path points at and returns either a target or the rejection to report. Each function keeps its own write rule, because the rules genuinely differ: an override sets a number rather than adding to it, ignores `cooldown`, `max_delta_per_turn`, and the rule that a counter only counts up, and can un-set a milestone. A differential check ran both implementations over 3960 payloads: twenty paths, fourteen values, three starting states, plus every three-path combination. The results are identical except that 674 rejections from `apply_override` now carry a `fix` string. `apply_delta` already worded those, and the world-state editor renders them, so an override that names an unknown flag now explains itself the way a delta does. **`sse`, `SSE_HEADERS`, and `turn_error` move to `app/sse.py`.** Two routers stream, and `chat.py` had to import from `routers.adventures` to reach them. **`get_adventure_or_404` becomes the `current_adventure` dependency.** All 32 handlers repeated the call as their first statement. The ownership check now reads in the signature and runs before the body. FastAPI caches a dependency for one request, so the handler's `db` is the session the adventure came from. The generated OpenAPI document is byte-identical except on `rename_branch`, where `branch_id` is now listed before `adventure_id`, because that handler no longer names `adventure_id` itself. Parameter order in the document is cosmetic. Six tests in `test_state_revert.py` call `undo_turn` and `retry_action` directly rather than over HTTP. They pass the adventure they already hold instead of an id. 549 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
0623dd8b78 |
Split Play.jsx into a directory of twelve files
`frontend/src/pages/Play.jsx` was 2280 lines holding 27 components. It is now `frontend/src/pages/Play/`: `index.jsx` with the page component, five panels, two drawers, the reports, the take pager, the refresh dialog, and the shared formatting helpers. Every line moved verbatim. A coverage check confirms every non-blank line of the original appears exactly once, in order, across the twelve files, and a name-resolution check confirms every identifier each file references is defined or imported there, with no unused imports. The line split stranded a comment at six of the boundaries. A leading comment sits above the section it describes, so each cut left one at the end of the file before it. All six moved to the section they describe, rewritten in the house style. `usePlaySession.js` is not here. The page component still owns all of the session state. Moving eighteen `useState` calls and seven `useEffect` calls is a rewrite rather than a move, and no frontend test would catch a mistake in it today, so it waits for Stage 5. Four stand-in providers under `backend/tools/` now define `last_usage = None`. The turn engine reads that attribute after every call, and the fixtures never defined it, so `tools.tree_fixture` crashed. That break predates this branch. Verified by 549 passing tests, a clean `npm run lint` and `npm run build`, and by driving the Play screen: the story view, all five panels, both drawers including the world-state edit form, the branch map, the refresh dialog, and the take pager stepping onto a take that lives on another branch. No console errors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
8422ff24f6 |
Split the world-state engine into four modules
`backend/app/worldstate/engine.py` held 918 lines covering four separate jobs: reading a scenario's schema, parsing the block the model writes, applying a change within the schema's limits, and rendering state as prompt text. Each is now its own module, the largest 418 lines. An AST comparison against the old file confirms all 29 definitions are identical. No call site changes, because `worldstate/__init__.py` exports the same names it did before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
2fa812c056 |
Split the adventures router into a package
`backend/app/routers/adventures.py` held 2353 lines and 35 endpoints. It is now a package of 14 modules, the largest 443 lines. The split moves text rather than rewriting it. An AST comparison against the old file confirms all 86 definitions are identical, and the OpenAPI schema still lists the same 35 operations. Names a test replaces now live in `turns.py` only, and other modules reach them as `turns.<name>`. Rebinding a re-exported alias changes the alias and leaves every caller reading the original, so the package root does not re-export them. A patch aimed at the old target raises `AttributeError` instead of passing while doing nothing. Tests and the fixtures in `backend/tools/` say `adventures.turns.<name>`. The same rule keeps the turn lock working. One module owns `_active_turns`, so one lock guards one set. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
32cd7c1077 |
Give the tests one setup instead of thirty-five
Every test module carried the same eight-line prologue redirecting the database to a temp file. Only the first one to be imported ever took effect: `app.database` reads `AIDND_DB_PATH` at import and builds `engine` from it once, so by the time the second module ran the engine already existed. The other 34 copies created a temp file that nothing opened and nothing deleted, and leaked one per module per run. `conftest.py` now does it once, which is early enough because pytest imports conftest before any test module. It also deletes the file when the run ends. The tests still share one database, exactly as they already did: each `client` fixture calls `create_all` on setup and `drop_all` on teardown, so no test sees another test's rows. `tests/fakes.py` holds the one `ScriptedProvider`. Nine modules each had a copy, and the copies had drifted into four feature sets, so a test that needed to raise a provider error had to be written in one of the files whose copy supported that. The shared one is the superset. The two `FakeProvider` copies were the same class with a fixed reply, so they use it too. `test_chat.py` keeps its own, which implements `chat` rather than `generate` and records what it was constructed with. An autouse fixture resets the fake's class state between tests, so a stale reply list can no longer reach the next test. 435 lines out of the suite. 549 tests pass. Verified live by sabotage: breaking the shared fake fails 13 tests across four modules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
b1772c6e21 |
Plan the readability refactor, and clear the tree for it
Phase 17 splits the four files that hold most of the code, finishes the schema migration SP8 left half done, and stops the published guide from drifting away from its Markdown source. `plan/17-refactor.md` carries the plan and the progress table, and `plan/STATUS.md` points at it. Stage 0 is hygiene only. Both abandoned worktrees are gone, which freed about 104 MB. Removing `sp7-tree-ui` needed one extra step: a Vite dev server had been running out of it since 2026-08-18, holding `frontend/.vite` open and owning port 5173, and serving a tree 54 commits behind `main`. The three stale `.db` files are deleted; `data.db` is not. `AIDND_TRUSTED_PROXY_HOPS` is now documented. It was read at `limits.py:55` and named in no `.env.example`, README, or blueprint. It sets how many proxy hops the rate limiter trusts in `X-Forwarded-For`, so a deployment that adds a hop without setting it gets the bucket-rotation bypass back. The 19 squash-landed branches are still there. `git branch -D` is blocked by the permission classifier; the verified command is in the plan file. 549 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
9398c13da5 |
Delete seeded scenarios no seed file claims any more
`previous_titles` stops a rename stranding the row it left behind, but the rows already stranded still had to be deleted by hand on every deployment. The seeder now removes them on the next boot, which retires the stale "Road to the Champion" demo without a database console. Only rows with a NULL owner and `is_public` are considered, and a player's own scenario is neither, so nothing anybody created is reachable. An adventure started from a deleted demo survives: `adventures.scenario_id` is `ON DELETE SET NULL`, and the adventure holds its own copies of the cards and scripts, so it loses only the inherited cover art. Two cases skip the sweep, because neither is an instruction to remove live content: a seed file that fails to parse claims no title, and an empty seed directory reads as a packaging failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
c2d3f0d8b9 |
Show the starter adventure the artwork of the demo it came from
An adventure has no cover art of its own and inherits its scenario's, and a bundle carries no scenario id, because an id means nothing in another database. The starter card therefore fell back to a monogram tile while the demo beside it showed the Pokeball. The starter file names its source under `scenarioTitle`, and the copy is linked to the seeded scenario with that title. If no seed answers to the name, the adventure keeps a NULL `scenario_id`, which is the state every imported bundle is in and costs only the artwork. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
ef4da7fda2 |
Keep the local Claude test rig, and record what it found
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by the local `claude` command line tool, so a demo can be played against a real model with no API key. Each request spawns one `claude --print` process, which suits the turn engine: the app assembles the whole prompt every turn and expects a stateless endpoint. Playing the Pokemon demo through it made all five of the world-state changes in `plan/16` work, and the refusal loop ran end to end for the first time: turn 3 clamped to nothing, turn 4's assembled prompt carried the correction verbatim, and the model's next delta was right. Two failures previously blamed on the code were the demo model. One is a schema fault and is still open: Milo's three Pokemon share one `active_hp` stat, so a switch leaves the newcomer at 0 HP. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
6cbf6d996f |
Give every new guest an adventure that is already played
An empty account gives a visitor nothing to read, and the daily demo turns are limited, so learning what the app does cost one of them. `app/starter.py` now copies a shipped export bundle into each new guest at the point the row is created. The bundle is two exchanges of the Pokemon demo, which ends on a knockout and shows an applied change, a refused one, and a milestone. The guest row is committed before the copy is attempted, so a failure there still leaves them with an account, and the copy runs inside a savepoint. The row building that `POST /adventures/import` did inline moved into `bundle.materialize`, which both callers use. The rate and size checks stayed in the endpoint: the starter writes a file the server ships, so it has no untrusted list to cap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
ae39514d1e |
Name the Pokemon demo, and land a seed rename on its own row
`seed.py` matches a seed file to its scenario by title, so renaming one inserted a second public scenario and stranded the first. The stranded row stays public forever and has to be deleted by hand on every deployment, which is what happened to "Road to the Champion". A seed file now lists its old titles under `previous_titles`, and the rename updates the existing row. The cover art is a PNG data URI. `app/images.py` accepts raster formats only, because SVG can carry script and the bytes are served from the app's own origin, so an SVG stores but yields an empty `image_url`. `tools/make_pokeball.py` draws the ball with `zlib` alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
ab546d6f3c |
Ask the Pokemon demo for changes, not for totals
The model sent a total rather than a change for almost every number: `npc.ivysaur.hp: 96`, `npc.milo.active_hp: 88`, `player.potions: 2`. Because every `hp` starts at its maximum, each one clamped back to where it started, and the potion count rose when the player spent one. `EMIT_RULE` does say "deltas (not new totals)", but the scenario contradicted it at closer range. `milo.active_hp`'s description said "Reset this to the newcomer's full HP", which asks for an absolute and is injected every turn. The bullets said "drop the HP, and raise it when healed", naming a direction but no sign. Five HP descriptions said only whose HP it was. The one line that said "not a delta" covered `player.active_pokemon`, so naming the exception made the rule look optional. `pokemon_fainted` is the control: its description says "add 1 each time", and it is the only number that behaved. Every stat description now states the sign, and the lead-in gives a worked example. `milo.active_hp.max_delta_per_turn` goes 65 to 98, because a switch moves that stat a full bar from 0 and the old cap made the reset unreachable in one turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |