d63804f22ecbaed80741241a154cdaef82f7b2ed
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b7005e6fdd |
M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.
This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:
visible active transcript position == stored head == authoritative state
Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)
A narrator edit no longer rewrites a row. It returns to the state before the
turn, takes the reader's exact text as the accepted narration, re-derives the
state that text implies, and becomes a new active continuation — while the
original narration keeps its words, its live flag and its whole future as
retained history. At the tip the correction is another take; with story below
it, it forks. No new history machinery: this is the existing fork/take/head
path with the reader's text in place of a generated reply. The §14A refusal
is therefore gone for narrator turns, and remains only for player input.
Pre-M5 positions
Migration 88 backfills the empty narrative document onto every action written
before M5, and a missing snapshot now restores the empty document instead of
leaving the previous position's state standing. Restoring to an old Save
Point no longer leaves a later position's entities and facts on screen.
Narrator context
Replayed history carries prose only; the machine-readable block is no longer
reconstructed into past turns, where it contradicted the authoritative state
in the same prompt. A fact withdrawn by a manual correction is now named as
no longer true, with the reader's reason, rather than silently dropped.
Also
- state_changes joins the action-list bulk read, removing one query per row.
- Extraction takes only the application's own protocol payload: an ordinary
```json or ```python block in a story survives, and a mangled proposal
still does not reach the reader.
Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
|
||
|
|
8c65ae99de |
M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo. |
||
|
|
f1bebe18d0 |
Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`, `variant_count`, `state_before`, and `world_state_before`, plus `adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in SQLite, so migration 71 quotes it. Nothing outside the migrations read these. `models.py`, `schemas.py`, and `ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders by `id`, and `attempts.renumber`, `context.history.max_action_index`, and `nodes.next_index` are deleted. Two changes keep the migration replayable on a `create_all` database: - `_split_variants_into_siblings` wrote through the live ORM table, so it stopped compiling once migration 66 removed five of its columns. It now writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`. - Five data passes read columns these migrations drop. Each now calls `_has_columns` and returns early when the columns are absent. `bootstrap` takes a `through` version so a migration test can stop at the schema it asserts on. 555 tests pass, up from 549. Eight of the new cases assert each column is gone after a real schema-45 database migrates all the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK |
||
|
|
32cd7c1077 |
Give the tests one setup instead of thirty-five
Every test module carried the same eight-line prologue redirecting the database to a temp file. Only the first one to be imported ever took effect: `app.database` reads `AIDND_DB_PATH` at import and builds `engine` from it once, so by the time the second module ran the engine already existed. The other 34 copies created a temp file that nothing opened and nothing deleted, and leaked one per module per run. `conftest.py` now does it once, which is early enough because pytest imports conftest before any test module. It also deletes the file when the run ends. The tests still share one database, exactly as they already did: each `client` fixture calls `create_all` on setup and `drop_all` on teardown, so no test sees another test's rows. `tests/fakes.py` holds the one `ScriptedProvider`. Nine modules each had a copy, and the copies had drifted into four feature sets, so a test that needed to raise a provider error had to be written in one of the files whose copy supported that. The shared one is the superset. The two `FakeProvider` copies were the same class with a fixed reply, so they use it too. `test_chat.py` keeps its own, which implements `chat` rather than `generate` and records what it was constructed with. An autouse fixture resets the fake's class state between tests, so a stale reply list can no longer reach the next test. 435 lines out of the suite. 549 tests pass. Verified live by sabotage: breaking the shared fake fails 13 tests across four modules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
47c7800903 |
Report the world-state changes the engine refuses
`apply_delta` records three outcomes for every change the model sends: `applied`, `clamped`, and `rejected`. Everything downstream read only `applied`. A refused change reached the player as an ordinary chip, and reached the model on the next turn as a change that had succeeded. Five parts: - `Action.world_changes` reads `clamped` and `rejected` beside `applied`. Accepted stats carry a `clamped` flag; refusals become `kind: "rejected"` entries. The `fix` key is present only when the engine wrote one, because this property runs for every action of every list response. - The UI separates the three outcomes. A clamp to a standstill reads `no change - at its limit` on a dashed chip, a partial clamp is marked `(limited)`, and a rejection carries its reason. Dashed and dimmed rather than red: a refused change means the rules are working. - The goals line names the milestone id, as `milestones.<id>`. The ids appeared nowhere in the prompt before, so the model could not send one. - Each rejection, and each clamp that moved nothing, builds a `fix` string from the stat definition at the point of refusal. `render_refusals()` renders them into the next prompt above `EMIT_REMINDER`. - `_history_text` replays `applied_delta()` instead of the sent delta, so a past turn's state block shows only what the engine accepted. A clamp that reduced a change but still moved the value reports nothing. If you tell a model its 80 damage became 30, it can treat the shortfall as a debt and send the remaining 50 next turn, which is the swing `max_delta_per_turn` prevents. In the demo scenario, `pokemon_left` becomes `pokemon_fainted` (`type: counter`, `initial: 0`). Starting at the ceiling turned a wrong-signed delta into a silent no-op; counting up puts the wrong sign on the counter rule, which refuses it out loud. The instructions also now ask for `world.turn`, which sat at 0 for a whole playtest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
0a12d9cd47 |
Make a retry a node, not a rewrite
Every attempt at a turn is now its own row at the same (branch, depth), with `live` naming the one the story tells. The JSON repeating group on `actions.variants` is read one last time, by a migration that writes it out as the sibling rows it always described, and then goes unread. The snapshots turn around with it: an action carries the state it left behind rather than the state it started from, because attempts at one turn share a starting position and differ exactly in their outcome. Rolling back is "what the node in front left behind", one lookup on the path, and it is what undo and retry now both read. And the memory holdback goes. It existed because retry rewrote a row under a mark that had already moved past it; a retry writes a sibling now, and replacing what a coordinate says withdraws what was derived from it — the same repair undo and delete already made. The assembled prompt is still stored once per turn: it moves with the live flag, so a superseded attempt keeps only the few hundred bytes that were its own. Measured on the 600-action fixture: 700 rows for the same 600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB and page load 62.7 kB unmoved. 347 tests green. `tests/test_story_tree_baseline.py` and `tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move the baseline for the variant-count semantics and did not need to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
ae6e5af6c7 |
Store context_snapshot compressed
One column is 89% of the database and the free tier allows 512 MB. Reads were already solved -- the column is deferred, so a page load never touches it and one screen fetches one row at a time -- but nothing had costed storage, and storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per action, so the ceiling arrives around 5,400 actions and 944 are stored. Postgres already compresses it and only gets 1.7x. pglz is tuned for fast decompression of data a query might filter on, and nothing has ever filtered on an assembled prompt -- it is written once and read whole, rarely, by the Insights viewer. zlib gets 3.5x on the same text for a decompress on a request that already made an LLM call. Done as a TypeDecorator rather than a second column, so every call site still writes a dict and reads a dict back, and deferred/undefer/load_only keep naming the same attribute. Only the storage format moves. Migrations 43-45: add the bytea, convert into it, drop the original, rename. The backfill is the one destructive step in the file -- 44 removes the only other copy -- so it decompresses every row and compares it against what went in, and a row that fails aborts the run. The whole loop is one transaction, so an abort rolls the DROP back and the prompts are still there. Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column came out named context_snapshot, every snapshot compared equal and the one NULL stayed NULL. Postgres does not return the disk by itself: DROP COLUMN only marks the column gone and the backfill leaves a dead tuple per row, so the table peaks near twice its size before settling. The deploy needs one VACUUM FULL to collect it; the migration comment says so. The egress fixture's snapshots are prose now rather than "x" * 20_000, and the prose generator moved to tools/fakeprose.py so the harness and the tests share one definition. A repeated character compresses a thousandfold: against the old fixture a compressed column looked free and the byte ceilings would have been guarding nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
a6cb49293c |
Name the columns a list response carries
deferred=True keeps the four heavy Action columns out of a bulk read, but it makes narrowness the thing a future column has to remember to ask for -- and both egress blowouts this project has had were a column nobody remembered. Listing what each list response renders inverts the default: a new column costs nothing on these paths until someone adds it to the tuple. The adventures index was not merely a future risk. It loaded whole Adventure entities to render a title, a stamp and a snippet, and an Adventure carries script_state, world_state, placeholders, story_summary, memory, authors_note and ai_instructions -- ~15 kB a row in production, none of it on that screen, all of it fetched once per adventure on every index load. Measured on six adventures with 78 kB of body each: 469.7 kB entity-loaded against 318 B projected. The memories drawer stops walking adventure.memories. The walk is what retrieval used to do and the reason a turn cost megabytes; a relationship load takes whole entities, so it picks up whatever the model happens to grow. Nothing changes today -- embedding_blob is already deferred -- which is the point. world_delta stays on the action list because ActionOut.world_changes is computed from it. Leaving it off would not save the bytes, it would spend them one row at a time as a lazy load. Two tests cover the index: one asserts the listing query names none of the body columns, one puts a byte ceiling on six adventures carrying 80 kB apiece. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
12d57afdac |
Put a number on what an endpoint may fetch, not just a column list
test_egress.py asserted which columns a statement names, which is the shape both of this project's egress blowouts took. It would all still pass if a response grew tenfold within the columns it is allowed to read -- and a story that keeps getting longer does exactly that. Production's longest adventure is 607 actions where the plan assumed 200. So dbmeter, which was built to be importable from tests and was not yet used by any, now backs four byte ceilings: the page load, the action list, and one action's snapshot fetched on demand. Budgets are per action rather than absolute, so they mean the same thing whatever size the fixture is set to, and generous -- 3 kB against a real 994 B. They are there to catch an order of magnitude, not to freeze a byte count. The fourth test is the one that keeps the other three honest. A ceiling proves nothing unless the thing it excludes would breach it, so it undefers the snapshot on purpose and asserts the same twelve rows cost more than ten times the budget. If the fixture ever shrinks below the point where that holds, that test fails rather than the ceilings quietly passing on nothing. Meter grows detach() and a context manager. A script exits and takes the wrapping with it; a test does not, and one test leaving the shared engine metered would charge bytes to a scope nobody opened. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
47e33fa311 |
Read a window of the story per turn instead of all of it
Two reads still grew without bound after the snapshot fix. `Action.variants` holds every discarded retry attempt, but a list response only needs how many there are — so each retry permanently added ~5 KB to every later load of that adventure. Defer the column and keep the count beside it (migration 37, backfilled server-side), with set_variants() as the one write path that keeps the two in step. `story_actions()` walked adventure.actions, then every caller threw almost all of it away: the builder concatenates the story and immediately cuts it back to the token budget, the NPC check looks at the last 6, retrieval at the last 4, the cursor clamp only wants a count. A turn on a 200-action adventure read 839 KB to use ~70 KB, and grew with every turn played. app/context/ history.py serves those shapes from SQL; window_covering() measures the actions it fetched and projects how many more it needs, fetching only the part it does not already hold. Memorybank cursors move to position_of_index() and settled_count()/settled_slice() — same arithmetic, no full list. The scripting pipeline still receives the whole history per AI Dungeon's API, and every helper reuses adventure.actions when it is already loaded, so a scripted adventure pays what it always did and never twice. Measured at production shape: retry tax 5.1 KB -> 0; turn 200 839 KB -> 129 KB and flat from ~turn 50; a 200-turn playthrough 84.5 MB -> 23.0 MB; a delete 115 KB -> 5 KB. Verified the window builds a byte-identical prompt to the full story across budgets from 1K to 100K tokens, with and without the retry exclusion - this is a cost change and nothing else. Cursor helpers checked against the old list arithmetic, including after deleting a middle action. Counts are real SELECT count(...): Query.count() wraps the entity select in a subquery, so the SQL named every deferred column and the egress guard could not tell it apart from a bulk fetch. 139 tests pass; the four new guards verified by sabotage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet |
||
|
|
f1bd099ec8 |
Cut database egress 189x by deferring the prompt snapshot
The free-tier 5 GB/month network transfer allowance ran out, which blocks connections outright. The database is only ~55 MB, so 5 GB meant the whole thing was being pulled roughly 90 times over. Cause: actions is 39 MB of that 55 MB -- 541 rows at ~74 KB each, almost entirely context_snapshot, which stores the whole assembled prompt for a turn. Every adventure load and every turn fetched all of it in order to read two small things out of it: the world-change chips under an AI message (Action.world_changes) and the emit block re-attached when replaying history to the model (_history_text). The Insights viewer is the only consumer that wants the whole snapshot, and it asks for one action at a time. Lifts that slice into its own small actions.world_delta column (migration 36) and marks context_snapshot, state_before and world_state_before deferred, so they load only when something touches the attribute -- Insights, undo and retry, all single-action paths. The backfill runs server-side, dialect-specific (json_extract on SQLite, #> on Postgres), because pulling 39 MB of snapshots into Python to rewrite a slice of each would defeat the purpose. Measured at production shape (541 actions, 72 KB snapshots), one adventure load goes from 38.46 MB to 0.20 MB. The traffic that consumed 5 GB would now be about 27 MB. Deliberately not included: limiting the history query to recent actions, and removing the redundant db.refresh(adventure) calls. Both were sized against the old numbers; against a 0.20 MB load they would take ~27 MB a month down to ~10 MB, which is not worth the complexity. tests/test_egress.py hooks before_cursor_execute and asserts the emitted SQL never names the deferred columns during a bulk load, so this cannot regress silently. 123 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet |