One column is 89% of the database and the free tier allows 512 MB. Reads were
already solved -- the column is deferred, so a page load never touches it and
one screen fetches one row at a time -- but nothing had costed storage, and
storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per
action, so the ceiling arrives around 5,400 actions and 944 are stored.
Postgres already compresses it and only gets 1.7x. pglz is tuned for fast
decompression of data a query might filter on, and nothing has ever filtered
on an assembled prompt -- it is written once and read whole, rarely, by the
Insights viewer. zlib gets 3.5x on the same text for a decompress on a request
that already made an LLM call.
Done as a TypeDecorator rather than a second column, so every call site still
writes a dict and reads a dict back, and deferred/undefer/load_only keep
naming the same attribute. Only the storage format moves.
Migrations 43-45: add the bytea, convert into it, drop the original, rename.
The backfill is the one destructive step in the file -- 44 removes the only
other copy -- so it decompresses every row and compares it against what went
in, and a row that fails aborts the run. The whole loop is one transaction, so
an abort rolls the DROP back and the prompts are still there.
Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway
Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column
came out named context_snapshot, every snapshot compared equal and the one
NULL stayed NULL.
Postgres does not return the disk by itself: DROP COLUMN only marks the column
gone and the backfill leaves a dead tuple per row, so the table peaks near
twice its size before settling. The deploy needs one VACUUM FULL to collect
it; the migration comment says so.
The egress fixture's snapshots are prose now rather than "x" * 20_000, and the
prose generator moved to tools/fakeprose.py so the harness and the tests share
one definition. A repeated character compresses a thousandfold: against the
old fixture a compressed column looked free and the byte ceilings would have
been guarding nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Migration 38 left memories.embedding in place so a rollback could still find
the vectors. Production has since been verified reading from embedding_blob,
so migration 42 drops it: 4 MB of a 99.6 MB database holding nothing anyone
reads.
Removing it surfaced a live bug. Changing your embedding model is supposed to
throw the bank's vectors away and let the post-turn pass rebuild them, because
two models' vectors are not comparable. The settings route did that by nulling
memories.embedding -- correct until 38 moved the vectors, after which it
cleared the dead column and left the blob intact with `embedded` still true.
_embed_pending filters on `embedded IS FALSE`, so it never saw those rows and
the bank went on ranking against the old model's vectors permanently.
Nothing would have reported it. cosine returns 0.0 on a width mismatch, so a
different-width model scores every memory zero and retrieval returns whichever
rows happen to sort first; a same-width model scores plausible garbage.
The bulk clear now sets both columns. It stays a bulk UPDATE rather than going
through set_vector -- loading the rows is the cost that whole path exists to
avoid -- so set_vector's docstring now names it as the one caller that
legitimately writes those columns by hand. No cache invalidation is added:
clearing `embedded` drops the rows out of the catalogue query, and set_vector
evicts each entry as the re-embed puts it back.
test_embedding_blob.py now rebuilds the pre-38 schema by hand where it tests
the backfill, since create_all no longer produces the column it converts from,
and asserts 42 removes it at the end of a full bootstrap -- 38 reads that
column and 42 drops it, so an upgrade that reordered them would arrive with an
empty bank.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.
Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.
one turn 3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
run_post_turn 3,139.1 kB -> 0.7 kB
Insights 3,223.7 kB -> 117.9 kB
Memories drawer ~3.1 MB -> 23.6 kB
A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.
The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.
memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.
Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.
It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.
Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.
The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.
Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
Two reads still grew without bound after the snapshot fix.
`Action.variants` holds every discarded retry attempt, but a list response
only needs how many there are — so each retry permanently added ~5 KB to
every later load of that adventure. Defer the column and keep the count
beside it (migration 37, backfilled server-side), with set_variants() as the
one write path that keeps the two in step.
`story_actions()` walked adventure.actions, then every caller threw almost
all of it away: the builder concatenates the story and immediately cuts it
back to the token budget, the NPC check looks at the last 6, retrieval at the
last 4, the cursor clamp only wants a count. A turn on a 200-action adventure
read 839 KB to use ~70 KB, and grew with every turn played. app/context/
history.py serves those shapes from SQL; window_covering() measures the
actions it fetched and projects how many more it needs, fetching only the
part it does not already hold. Memorybank cursors move to position_of_index()
and settled_count()/settled_slice() — same arithmetic, no full list.
The scripting pipeline still receives the whole history per AI Dungeon's API,
and every helper reuses adventure.actions when it is already loaded, so a
scripted adventure pays what it always did and never twice.
Measured at production shape: retry tax 5.1 KB -> 0; turn 200 839 KB -> 129 KB
and flat from ~turn 50; a 200-turn playthrough 84.5 MB -> 23.0 MB; a delete
115 KB -> 5 KB.
Verified the window builds a byte-identical prompt to the full story across
budgets from 1K to 100K tokens, with and without the retry exclusion - this
is a cost change and nothing else. Cursor helpers checked against the old list
arithmetic, including after deleting a middle action. Counts are real
SELECT count(...): Query.count() wraps the entity select in a subquery, so the
SQL named every deferred column and the egress guard could not tell it apart
from a bulk fetch. 139 tests pass; the four new guards verified by sabotage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
The free-tier 5 GB/month network transfer allowance ran out, which blocks
connections outright. The database is only ~55 MB, so 5 GB meant the whole
thing was being pulled roughly 90 times over.
Cause: actions is 39 MB of that 55 MB -- 541 rows at ~74 KB each, almost
entirely context_snapshot, which stores the whole assembled prompt for a
turn. Every adventure load and every turn fetched all of it in order to
read two small things out of it: the world-change chips under an AI
message (Action.world_changes) and the emit block re-attached when
replaying history to the model (_history_text). The Insights viewer is
the only consumer that wants the whole snapshot, and it asks for one
action at a time.
Lifts that slice into its own small actions.world_delta column
(migration 36) and marks context_snapshot, state_before and
world_state_before deferred, so they load only when something touches
the attribute -- Insights, undo and retry, all single-action paths.
The backfill runs server-side, dialect-specific (json_extract on SQLite,
#> on Postgres), because pulling 39 MB of snapshots into Python to
rewrite a slice of each would defeat the purpose.
Measured at production shape (541 actions, 72 KB snapshots), one
adventure load goes from 38.46 MB to 0.20 MB. The traffic that consumed
5 GB would now be about 27 MB.
Deliberately not included: limiting the history query to recent actions,
and removing the redundant db.refresh(adventure) calls. Both were sized
against the old numbers; against a 0.20 MB load they would take ~27 MB a
month down to ~10 MB, which is not worth the complexity.
tests/test_egress.py hooks before_cursor_execute and asserts the emitted
SQL never names the deferred columns during a bulk load, so this cannot
regress silently. 123 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Retry used to delete the last AI action and generate a replacement, so the
discarded narration was simply gone. The row survives now: each attempt is
appended to actions.variants with variant_index naming the live one, and a
ChatGPT-style pager under the message browses them.
Action.text still mirrors the active variant, so the context builder, memory
bank, summarizer and export needed no changes. A variant carries only what
differs between attempts -- the text, the reasoning, and the state it
produced -- never the assembled prompt, which is identical across attempts of
one turn and is the bulk of context_snapshot.
Only the last message can be switched, restoring the script/world state that
attempt produced; earlier turns were written as a continuation of whatever is
active there, so theirs are read-only previews.
Three things that would otherwise bite:
- generate_turn now wraps _generate_turn and watches for a save sentinel. If
the generator ends without it (provider error, empty reply, script stop,
client hangup) it re-applies the previous variant -- otherwise a failed
retry leaves rolled-back stats under un-rolled-back text.
- Retry reuses the turn's own index rather than next_index, or the clock the
world-state cooldowns run on advances on a re-run of the same turn.
- Editing a message rewrites the active variant too, or paging away and back
silently reverts the edit.
Migrations 34/35 verified as an upgrade against a populated database, not
just a fresh schema. test_state_revert's retry test asserted the old delete
behaviour and was rewritten.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.
Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.
Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
An adventure copies its scenario's plot text and story cards at creation so
later authoring never disturbs a story in progress. This is the explicit
opt-out, alongside the existing per-script "Sync from library".
GET /adventures/{id}/refresh returns a plan (per-field old/new diff, card
add/update/remove, world-state added/removed paths, and any ${...} answers
still needed); POST applies it under the turn lock so it can't race a
generating turn. The Plot panel shows the plan in a confirm modal first.
Overwrites the plot fields and scenario-derived cards. Deliberately left
alone: the opening `start` action (the story is built on it, and it is baked
into memories and the summary), the adventure's own title and summary,
player-authored story cards, and the live value of every stat the schema
still defines.
Two enablers were needed:
- adventures.placeholders (migration 32). ${...} answers were used once at
creation and discarded, so re-copying scenario text would have re-injected
a literal ${Hero}. Adventures predating the column re-prompt once via the
existing modal, then the answers are saved.
- story_cards.source_ref (migration 33), "card:<id>" / "npc:<key>", NULL for
player-authored. Adventure cards had no link back to their source, so a
rename read as delete-plus-add and player cards would have been clobbered.
Legacy cards name-match once, then adopt the ref.
World state goes through a new worldstate.reconcile(): keep values the schema
still defines, add missing ones at their initial, drop removed ones and their
cooldown bookkeeping. Not instantiate(), which would heal the player to full
and wipe their milestones.
The confirm modal is portalled to <body>: .side-panel's panel-in animation
has fill mode `both`, which makes it the containing block for position:fixed
descendants, so an overlay rendered in place was trapped in the 420px panel
and clipped by its overflow.
14 new tests in backend/tests/test_scenario_refresh.py; 85 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XzBCyXH4hBEVHqEaertcq4
Rework the home page into a single landing surface and give the app a
consistent visual language, chosen as "illuminated tome" over two other
pitched directions because it builds on the existing Cinzel + gold identity
instead of replacing it.
Home is now Continue (up to 4 in-progress stories, each showing where you
left off) over a scenario shelf, each section with a "See all" link. The
full adventure list moves to /adventures.
Scenario cover art has three tiers, in precedence order: an uploaded picture
(downscaled client-side to 400px WebP before storing), an emoji, or gradient
art generated from a hash of the title so no card is ever an empty box.
Adventures inherit their scenario's art. Images live in the row rather than
on disk because Render's free tier has no persistent volume, and it keeps
export bundles self-contained; list responses carry a cacheable
/api/scenarios/{id}/image URL rather than the base64.
Also: ambient drifting motes behind the app, loading skeletons, staggered
card entrance, ornamental scene breaks and a drop cap in the story, a
"Weaving" thinking indicator, and a toast system replacing every alert().
Two fixes found along the way:
- Importing a scenario bundle with no "tags" key returned a 500. Column
defaults are not applied until flush, so the attribute was still None
when the width clamp sliced it.
- Anything meaning "the story's latest narration" was missing action type
"start", which is the only text a freshly created adventure has, so new
adventures looked empty. Collected as NARRATION_TYPES.
Migrations 30 and 31 add scenarios.image and scenarios.icon; both are
additive with a '' default and were verified against a database stamped at
29. vite.config.js now reads AIDND_API_PORT so the recurring port-8000
clash with another local app needs no file edit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
Each AI action now carries a compact world_changes summary (derived from its
stored snapshot), rendered as small chips beneath the message: numeric stats
show a signed delta (green up / red down), flags show on/off, milestones show
a check. Gives an at-a-glance "what changed" without opening Insights.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Structured world/player/NPC stats, two-way flags, and sticky milestones
per scenario (stat_schema). The AI proposes a per-turn delta; a Python
engine referees it (clamp to min/max, per-turn cap, cooldown, counters).
Band word-labels plus a fixed stat guide (descriptions + full ranges)
keep the model grounded. World State drawer + Insights delta report;
undo/retry roll it back via the Phase 11 snapshot pattern.
- migrations 26-28 (scenarios.stat_schema, adventures.world_state,
actions.world_state_before); all nullable, additive, safe on existing rows
- migration 29 raises the default context budget 4096 -> 16384
(custom values preserved)
- seeded demo scenario 04-rpg-world-state.json (Bandit Camp)
- 19 new tests (33 total pass)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The shared per-adventure script_state ("scoreboard" scripts write to) was
never reverted by undo, and retry re-ran the output hook on top of the already-
mutated state, double-applying its changes (e.g. "+10 gold" became +20).
Each action now snapshots script_state as it was immediately before its own
hooks ran (new Action.state_before column, migration 25):
- undo restores the turn's first-action snapshot, prunes memories that
summarized the removed actions, and takes the turn lock against races.
- retry restores the AI action's snapshot before regenerating.
Story-card mutations are not reverted (documented limit). Adds the project's
first test suite: unit + full HTTP integration through the real scripting
engine (14 tests). See plan/11-state-revert-and-retry-fix.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adventure scripts are frozen snapshots taken at creation. Add an opt-in
way to pull the latest library code into a live adventure:
- AdventureScript.source_script_id links a copy to its library Script
(migration 24; set at adventure creation)
- list endpoint flags out_of_date by diffing copy vs library
- POST /adventures/{id}/scripts/{sid}/sync overwrites the copy's code,
preserving enabled/position/script_state
- legacy copies with no link fall back to a name match, then adopt the link
- Play page shows a Sync button only when a copy differs from its library
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
- When a model streams reasoning but no story text, return a specific,
actionable error (raise Max output tokens / set a reasoning cap / use a
non-reasoning model) instead of the opaque "empty response".
- Bump default max_output_tokens 400 -> 800: 400 truncated scenes and
left reasoning models with no room after thinking. Affects new settings
rows; existing users keep their value.
- README: add a "Try it live" link to the Render deployment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs
unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me,
register upgrades the guest in place, login/logout, per-IP rate limits.
Every router scoped by user_id; Settings become per-user with the API
key Fernet-encrypted at rest and write-only through the API. Users
without a key get a server-funded demo key (OpenRouter free models,
20 turns/day, memory bank disabled on demo turns). Public read-only
demo scenarios (seed_demo.py); debug log restricted to local mode.
Frontend: auth modal + guest nudge, 401 re-establish/retry, demo
banner and key management in Settings.
Migrations 13-23 adopt existing data under a local user and encrypt
stored keys. Verified: migration on a copy of real data.db, two-session
isolation + register/login via curl and Chrome, demo cap 429, live
OpenRouter turn through the encrypted-key path, vite build + oxlint.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg