Two plan documents from designing the branching story tree. The tree work
turned up that memory retrieval fetches the whole bank's embeddings every
turn -- 96% of a turn's database traffic -- and that is the one read a tree
cannot window, so it lands first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
The app had no cleanup of any kind: in multi-user mode every first visit
mints a users row, so the demo has been accumulating one permanent account
per visitor along with everything they generated.
cleanup.py sweeps guests idle for AIDND_GUEST_RETENTION_DAYS (default 5),
once at startup and then every few hours. Startup is the load-bearing
trigger — the free tier sleeps after ~15 minutes, so a long timer rarely
gets to fire.
Idle is COALESCE(last_seen_at, created_at), not last_seen_at: _touch only
writes that column hourly, and a guest minted by /auth/me has it NULL until
its second request, so the simpler query would have deleted brand-new
visitors mid-session.
It's one Core DELETE rather than db.delete(user), which would SELECT every
adventure, action and memory into Python purely to delete them — the same
egress pattern as the 189x fix. Every FK from users down is ON DELETE
CASCADE, so the database does the whole graph and returns a count.
The filter requires is_guest AND email IS NULL, so registered users (who
upgrade in place) and local mode's implicit user are both out of reach, and
is_public is output-only so a guest can never own content another user can
see. Session cookies have no expiry and can outlive a swept row; that path
401s and the frontend's existing retry re-mints a session.
Guests are told: /auth/me serves guest_retention_days and the signup modal
states the window, sourced from the server so it can't drift from what is
enforced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
The per-IP rate limits could be bypassed entirely: uvicorn ran with
--forwarded-allow-ips "*", which trusts the leftmost X-Forwarded-For value
(client-controlled), and Render forwards the inbound header rather than
stripping it. Rotating the header handed out a fresh rate-limit bucket per
request, so the login/register limit (10/5min) and guest-minting limit
(30/5min) were no throttle at all — unbounded password guessing and guest-row
creation. Confirmed live: fixed IP -> 429 after 10; rotating spoofed header ->
no 429 across 14 attempts.
Two-layer fix:
- limits._client_ip now derives the client IP from the hop the trusted edge
appends (rightmost of X-Forwarded-For), which a client can't spoof past;
tunable via AIDND_TRUSTED_PROXY_HOPS. Dropped --forwarded-allow-ips "*".
- New per-account login throttle (email-keyed, 8 fails / 15 min, cleared on
success): stops distributed guessing against one account that a per-IP limit
can't, since it can't be diluted across many source addresses.
Also close an SSRF on the BYOK endpoint_url (hosted mode only): the connection
test and turn/chat streams now refuse a URL that resolves to a non-public
address (private/loopback/link-local metadata/reserved), checked at request
time so it resists a DNS record flipping to a private IP. No-op locally, where
reaching localhost Ollama is intended.
Tests: test_ratelimit_hardening.py (8), test_netguard.py (13). 172 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
The demo sleeps on Render's free tier, so a cold link looks broken to
anyone who won't wait 30-60s. A static page on GitHub Pages is never
asleep: it shows the screenshots immediately and sets the expectation
before the visitor clicks through to the demo.
Served from main:/docs, reusing the screenshots already committed there.
Palette and type match the app so the two read as one product.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015sygnUWH8WnaoKPDp7hXM1
The README claimed screenshots were "coming" and stopped describing the
project around phase 10, so the world-state engine, the retry/undo state
revert and the egress work — the most interesting parts — were invisible.
- Screenshots of the play screen, Insights, the schema editor and the
script editor, captured from the running app.
- Document the world-state engine, undo/retry rewind, and the two measured
performance fixes; correct the context assembly order to match
context/builder.py; drop the stale "phases 1-6 complete" note.
- Add CI (backend pytest, frontend lint + build, Docker image build) and
its badge. Nothing ran the 151 tests but me.
- Move CODE_REVIEW_FINDINGS.md to docs/self-review.md and say up front
that every correctness finding is resolved.
- Drop a dead import that the newly-wired lint flagged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015sygnUWH8WnaoKPDp7hXM1
max_output_tokens is a hard wall the endpoint enforces mid-sentence. The
```state block is emitted after the narration, so a long turn hits the wall
partway through the block and the deltas are lost — silently, since nothing
reads finish_reason.
builder.length_hint() derives a word limit from the cap ((cap - 50 headroom)
* 0.75 words/token * 0.90 buffer) and injects it just above EMIT_REMINDER,
which keeps the last slot it needs. Reserved in build_context like the
reminder is.
Phrased as a ceiling, not a budget. Measured against gemma-4-26b at cap 800,
n=5 per arm: no hint 174 words, "keep this turn under about N words" 246,
"hard limit ... a typical turn is much shorter" 170. A budget reads as a
target to fill — every budget run was longer than every unhinted one, pushing
turns toward the wall the hint exists to avoid. Ceiling phrasing still works
at tight caps: at 250, unhinted hit finish_reason=length 2/6, hinted 0/6.
tests/test_length_hint.py, 11 tests; each mechanism verified by sabotage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
The global input styling in index.css enumerates types, and `email` was
never in the list, so the Email box fell through to browser defaults while
the Password box right below it got the app treatment. Added it to both
copies of that selector list -- the base rule and the 16px iOS-zoom rule in
the mobile block; missing the second would revert the fix on phones.
While in there, the modal itself: log in / sign up are now segmented tabs
instead of a link crammed into the button row, fields carry autoComplete so
password managers work at all, a show/hide toggle on the password, Esc to
dismiss, and the error is a real role=alert box rather than the shared
.test-error text.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Two reads still grew without bound after the snapshot fix.
`Action.variants` holds every discarded retry attempt, but a list response
only needs how many there are — so each retry permanently added ~5 KB to
every later load of that adventure. Defer the column and keep the count
beside it (migration 37, backfilled server-side), with set_variants() as the
one write path that keeps the two in step.
`story_actions()` walked adventure.actions, then every caller threw almost
all of it away: the builder concatenates the story and immediately cuts it
back to the token budget, the NPC check looks at the last 6, retrieval at the
last 4, the cursor clamp only wants a count. A turn on a 200-action adventure
read 839 KB to use ~70 KB, and grew with every turn played. app/context/
history.py serves those shapes from SQL; window_covering() measures the
actions it fetched and projects how many more it needs, fetching only the
part it does not already hold. Memorybank cursors move to position_of_index()
and settled_count()/settled_slice() — same arithmetic, no full list.
The scripting pipeline still receives the whole history per AI Dungeon's API,
and every helper reuses adventure.actions when it is already loaded, so a
scripted adventure pays what it always did and never twice.
Measured at production shape: retry tax 5.1 KB -> 0; turn 200 839 KB -> 129 KB
and flat from ~turn 50; a 200-turn playthrough 84.5 MB -> 23.0 MB; a delete
115 KB -> 5 KB.
Verified the window builds a byte-identical prompt to the full story across
budgets from 1K to 100K tokens, with and without the retry exclusion - this
is a cost change and nothing else. Cursor helpers checked against the old list
arithmetic, including after deleting a middle action. Counts are real
SELECT count(...): Query.count() wraps the entity select in a subquery, so the
SQL named every deferred column and the egress guard could not tell it apart
from a bulk fetch. 139 tests pass; the four new guards verified by sabotage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
The free-tier 5 GB/month network transfer allowance ran out, which blocks
connections outright. The database is only ~55 MB, so 5 GB meant the whole
thing was being pulled roughly 90 times over.
Cause: actions is 39 MB of that 55 MB -- 541 rows at ~74 KB each, almost
entirely context_snapshot, which stores the whole assembled prompt for a
turn. Every adventure load and every turn fetched all of it in order to
read two small things out of it: the world-change chips under an AI
message (Action.world_changes) and the emit block re-attached when
replaying history to the model (_history_text). The Insights viewer is
the only consumer that wants the whole snapshot, and it asks for one
action at a time.
Lifts that slice into its own small actions.world_delta column
(migration 36) and marks context_snapshot, state_before and
world_state_before deferred, so they load only when something touches
the attribute -- Insights, undo and retry, all single-action paths.
The backfill runs server-side, dialect-specific (json_extract on SQLite,
#> on Postgres), because pulling 39 MB of snapshots into Python to
rewrite a slice of each would defeat the purpose.
Measured at production shape (541 actions, 72 KB snapshots), one
adventure load goes from 38.46 MB to 0.20 MB. The traffic that consumed
5 GB would now be about 27 MB.
Deliberately not included: limiting the history query to recent actions,
and removing the redundant db.refresh(adventure) calls. Both were sized
against the old numbers; against a 0.20 MB load they would take ~27 MB a
month down to ~10 MB, which is not worth the complexity.
tests/test_egress.py hooks before_cursor_execute and asserts the emitted
SQL never names the deferred columns during a bulk load, so this cannot
regress silently. 123 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Three fallout bugs from keeping the retried action row alive (906ba42),
plus two long-standing cursor bugs the same investigation turned up.
Retry context leak: the row being regenerated is still attached to the
adventure, so it was replayed as established story and the model wrote a
continuation of the attempt it was meant to replace — the story visibly
blended both takes. It leaked into four places, not one: history replay,
story-card trigger matching, in-scene NPC detection, and the memory-bank
similarity query. Adds a shared context.story_actions(exclude_action_id),
threaded through build_context and retrieve_memories.
Memory holdback: a memory could summarize the just-generated turn; retry
rewrites Action.text but memory_cursor has already advanced, so the memory
was never regenerated and went on describing narration no longer in the
story. settled_story_actions() holds the newest action back one turn —
only the last action is retryable, so that makes it unreachable. The
settled list is always a prefix, so cursors stay valid and nothing is
skipped. The run_post_turn clamp deliberately still uses the full count:
clamping to settled rewinds legacy adventures a step and double-covers an
action.
Cursor bookkeeping: memory_cursor is a position into story_actions() while
Memory.source_* are Action.index values, and the two diverge as soon as
anything is deleted. Deleting a middle action slid a never-summarized
action into the covered range, skipping it forever; and pruning a memory
left the actions it covered stranded behind the cursor. Adds
note_action_removed() (called before the delete in delete_action and
undo_turn) and a rewind in prune_dangling_memories. delete_action also
now prunes at all, which it never did.
Not addressed: editing an already-summarized action still leaves its
memory stale, and the cumulative story summary can't have one fact
un-mixed from it.
117 backend tests pass, including new test_memory_settling.py (12) and
two retry-context regression tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Retry used to delete the last AI action and generate a replacement, so the
discarded narration was simply gone. The row survives now: each attempt is
appended to actions.variants with variant_index naming the live one, and a
ChatGPT-style pager under the message browses them.
Action.text still mirrors the active variant, so the context builder, memory
bank, summarizer and export needed no changes. A variant carries only what
differs between attempts -- the text, the reasoning, and the state it
produced -- never the assembled prompt, which is identical across attempts of
one turn and is the bulk of context_snapshot.
Only the last message can be switched, restoring the script/world state that
attempt produced; earlier turns were written as a continuation of whatever is
active there, so theirs are read-only previews.
Three things that would otherwise bite:
- generate_turn now wraps _generate_turn and watches for a save sentinel. If
the generator ends without it (provider error, empty reply, script stop,
client hangup) it re-applies the previous variant -- otherwise a failed
retry leaves rolled-back stats under un-rolled-back text.
- Retry reuses the turn's own index rather than next_index, or the clock the
world-state cooldowns run on advances on a re-run of the same turn.
- Editing a message rewrites the active variant too, or paging away and back
silently reverts the edit.
Migrations 34/35 verified as an upgrade against a populated database, not
just a fresh schema. test_state_revert's retry test asserted the old delete
behaviour and was rewritten.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
NPCs are still part of stat_schema, but editing them inside the World State
panel meant a nested dashed sub-box among four flat stat lists. They now get
a top-level "Cast & NPCs" section: one card per NPC (avatar, name, npc.<id>
address, triggers, description, its own stats) in a responsive grid, saved
through the same schema path via the new NpcEditor/addNpc exports.
The panel's real inconsistency was CSS, not layout: the global field rule
keys off input[type="..."], which never matches the editor's typeless key and
description inputs, so those rendered with browser defaults beside properly
styled number boxes. One scoped base rule now covers the whole editor and
every entry is the same tile with a caption over each control.
Two bugs fell out of that: .se-band-n (0,1,0) always lost to
input[type="number"] (0,1,1), so band bounds rendered full-width; and the
exclusion has to be :not(:where([type="checkbox"])) because a plain :not()
inherits its argument's specificity and swallows every override.
In Play, NPCs group under a single "Cast" heading as compact plates instead
of each becoming another top-level group in the rail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.
Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.
Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
An adventure copies its scenario's plot text and story cards at creation so
later authoring never disturbs a story in progress. This is the explicit
opt-out, alongside the existing per-script "Sync from library".
GET /adventures/{id}/refresh returns a plan (per-field old/new diff, card
add/update/remove, world-state added/removed paths, and any ${...} answers
still needed); POST applies it under the turn lock so it can't race a
generating turn. The Plot panel shows the plan in a confirm modal first.
Overwrites the plot fields and scenario-derived cards. Deliberately left
alone: the opening `start` action (the story is built on it, and it is baked
into memories and the summary), the adventure's own title and summary,
player-authored story cards, and the live value of every stat the schema
still defines.
Two enablers were needed:
- adventures.placeholders (migration 32). ${...} answers were used once at
creation and discarded, so re-copying scenario text would have re-injected
a literal ${Hero}. Adventures predating the column re-prompt once via the
existing modal, then the answers are saved.
- story_cards.source_ref (migration 33), "card:<id>" / "npc:<key>", NULL for
player-authored. Adventure cards had no link back to their source, so a
rename read as delete-plus-add and player cards would have been clobbered.
Legacy cards name-match once, then adopt the ref.
World state goes through a new worldstate.reconcile(): keep values the schema
still defines, add missing ones at their initial, drop removed ones and their
cooldown bookkeeping. Not instantiate(), which would heal the player to full
and wipe their milestones.
The confirm modal is portalled to <body>: .side-panel's panel-in animation
has fill mode `both`, which makes it the containing block for position:fixed
descendants, so an overlay rendered in place was trapped in the 420px panel
and clipped by its overflow.
14 new tests in backend/tests/test_scenario_refresh.py; 85 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XzBCyXH4hBEVHqEaertcq4
Editing a beat opened a fixed 110px box. Player actions are one line, but AI
beats are several paragraphs, so editing one meant working through a keyhole.
New AutoTextarea sizes to its content on mount and on every change; CSS keeps a
110px floor so a short player action doesn't collapse, and a 65vh ceiling that
scrolls internally so Save/Cancel can't be pushed off the screen.
On phones both composers are one flex row, which left almost nothing to type
into: measured at a 390px composer, the Play textarea got 138px (Do/Say/Story
eat ~180px, Send ~75px) and the chat textarea 310px. Both now wrap, with the
buttons on their own line and the textarea taking the full 390px — Play keeps
Do/Say/Story and Send together on the top row, chat drops Send beneath. Play
stays two rows, so nothing gets taller. Desktop is unchanged; every rule is
inside the existing max-width:720px block.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
The demo-key backstop added in 1dd3108 keyed its check on
`api_key == DEMO_API_KEY` rather than on `using_demo`. That looked stricter but
was wrong: the demo key is an ordinary OpenRouter key, so a user can
legitimately paste that same value into their own Settings as BYOK. The guard
then raised on every resolve_provider_config() call for that account.
Because me_payload() resolves a provider config, this 500'd GET /api/auth/me —
the SPA's bootstrap call — so the frontend's `me` never resolved and the nav
(including the AI Chat link) never rendered, on top of chat itself failing.
`using_demo` is the flag that actually means "the server is paying", and only
resolve_provider_config's demo branch sets it, so the pinning guarantee is
unchanged: server-funded turns still can't reach an off-whitelist model.
Adds a regression test for a BYOK user whose key equals the demo key value, and
corrects the test that had asserted the buggy behaviour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
AI Chat is a plain scratchpad for talking to a model directly — no story
context, scripts or world state — for poking at models, prompts and endpoints
without starting an adventure. Power users only: the router 404s (rather than
403s) for everyone else and the nav link is hidden. The conversation lives in
localStorage, so there's no new table or migration.
is_power_user() now also returns True in local mode: it's the operator's own
machine and their own key, the same reasoning that makes the provider debug log
local-only.
Alongside that, the rule keeping the shared demo key off paid models now lives
in exactly one place. It had been duplicated into the chat router, which is how
one copy eventually drifts:
- resolve_provider_config() takes an optional model_override and is the only
place the whitelist is applied, so turns, AI Chat and the connection test all
inherit it. An override is a per-request preference, never a grant.
- ProviderConfig.__post_init__ refuses to exist when api_key is the demo key
and the model isn't whitelisted. It keys on the key itself rather than the
using_demo flag, so a mislabelled config can't slip past, and it raises so a
future path that bypasses the resolver fails loudly instead of billing.
- The demo branch still pins endpoint_url too — a user-controlled endpoint
would leak the key itself, which is worse than spending it.
Provider gained chat(messages, ...) beside generate(), both delegating to a
shared _stream(url, body); completion-mode endpoints get the messages flattened
into a labelled transcript. Settings' /models fetch moved to
list_endpoint_models() and is shared with /api/chat/config.
Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub
resolve_provider_config — the point is to exercise the real BYOK-vs-demo
decision and assert on what the provider actually received: off-whitelist
override pinned, off-whitelist Settings.model pinned, redirected endpoint
pinned, BYOK passed through untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
Rework the home page into a single landing surface and give the app a
consistent visual language, chosen as "illuminated tome" over two other
pitched directions because it builds on the existing Cinzel + gold identity
instead of replacing it.
Home is now Continue (up to 4 in-progress stories, each showing where you
left off) over a scenario shelf, each section with a "See all" link. The
full adventure list moves to /adventures.
Scenario cover art has three tiers, in precedence order: an uploaded picture
(downscaled client-side to 400px WebP before storing), an emoji, or gradient
art generated from a hash of the title so no card is ever an empty box.
Adventures inherit their scenario's art. Images live in the row rather than
on disk because Render's free tier has no persistent volume, and it keeps
export bundles self-contained; list responses carry a cacheable
/api/scenarios/{id}/image URL rather than the base64.
Also: ambient drifting motes behind the app, loading skeletons, staggered
card entrance, ornamental scene breaks and a drop cap in the story, a
"Weaving" thinking indicator, and a toast system replacing every alert().
Two fixes found along the way:
- Importing a scenario bundle with no "tags" key returned a 500. Column
defaults are not applied until flush, so the attribute was still None
when the width clamp sliced it.
- Anything meaning "the story's latest narration" was missing action type
"start", which is the only text a freshly created adventure has, so new
adventures looked empty. Collected as NARRATION_TYPES.
Migrations 30 and 31 add scenarios.image and scenarios.icon; both are
additive with a '' default and were verified against a database stamped at
29. vite.config.js now reads AIDND_API_PORT so the recurring port-8000
clash with another local app needs no file edit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
The AI would sometimes stop emitting the `state` delta block once it missed
a turn. Two compounding causes: the emit rule sat only in the system block
(far from where the model generates), and the block was stripped before
storage — so every replayed history turn looked blockless, biasing the model
by imitation to stop emitting too.
- EMIT_REMINDER: a one-line reminder appended last in the prompt (strongest
recency slot), gated on has_ws and counted against the token budget.
- render_delta_block + _history_text: re-attach each past AI turn's own delta
block in replayed history (reconstructed from the stored snapshot delta), so
the model always sees its emit format. Action.text stays clean, so UI,
embeddings, and card/NPC trigger-matching are unaffected. History budgeting
counts the augmented text so it can't overflow.
Undo/retry untouched (read the separate world_state_before column). 46 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UWVyFKvqJGjfbXdibLgkMe
The stylesheet had no real mobile design (one media query that only hid a
nav hint), so on phones the Play screen's four side-by-side columns
overflowed horizontally. Add a max-width:720px layer that collapses the
Play layout to a single story column, turns the World/Script State rails
into left-edge tabs that open as slide-in overlays, and makes the
Plot/Memory/Scripts/Insights side-panel a full-screen overlay. Also add a
hamburger dropdown for the top nav, switch the play-layout heights from vh
to dvh (fixes the mobile address-bar jump), and bump inputs to 16px to stop
iOS zoom-on-focus.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UWVyFKvqJGjfbXdibLgkMe
Players/authors can now directly correct the live values of stats the
schema already defines (health, trust, flags, milestones, etc.) without
waiting for the AI to emit a delta. New apply_override() sets values
absolutely rather than adding deltas, and — unlike the AI-facing
apply_delta() — bypasses cooldown/max_delta_per_turn and lets
milestones be un-set, since this is a deliberate correction rather
than a turn to police. Exposed via PUT /adventures/{id}/world-state
and an edit toggle in the World State drawer. Also fixes the
schema-editor stat-kind dropdown and free-text initial-value input
to size consistently with the numeric fields next to them.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The new kind <select> fell through to the generic input/select CSS
rule instead of the compact .se-num sizing, making it visually
inconsistent with the min/max/initial number boxes next to it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Numeric stats couldn't represent things like worn armor or held items.
A stat can now be marked type: "text" — the AI reports the new full
value instead of a delta, with no clamping/bands (cooldown still
applies). Updated the scenario editor's stat kind selector, the
World State drawer's stat display, and the emit-rule prompt; added
tests and an outfit stat to the RPG demo scenario.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Tell the model to read each stat's range and band labels and keep changes
proportionate — minor events nudge a value, while large jumps or hitting a
min/max are reserved for pivotal moments. Examples are relationship/story
based (not combat) so it fits any scenario.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reword the state-emit rule to be scenario-agnostic: record any tracked
change — health/resources, time, relationships/mood, status, progress,
items/info — not just fight outcomes. Soften the demo's AI instructions to
cover relationships and objectives alongside combat.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rewrite the global emit rule to be directive: treat narration as
authoritative and record numeric changes (damage, healing, time, emotion
shifts, fights intensifying), not just easy on/off flags — updating every
stat the scene affected. Also firm up the demo scenario's AI instructions
to reflect combat outcomes in the numbers each turn. Helps weaker free
models emit numeric deltas; the emit rule applies live so all RPG
adventures benefit immediately.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Each AI action now carries a compact world_changes summary (derived from its
stored snapshot), rendered as small chips beneath the message: numeric stats
show a signed delta (green up / red down), flags show on/off, milestones show
a check. Gives an at-a-glance "what changed" without opening Insights.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Store the model's literal reply (before the world-state block is stripped)
in each turn's snapshot, and render it in the Insights panel when inspecting
a past AI message (the 🔍 button). Makes it possible to see exactly what the
model emitted — including whether it sent a state delta and with what paths —
which is the fastest way to debug under-reported or misrouted stat changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON view no longer parses/saves on every keystroke — you type freely
and click "Save JSON" to apply. Invalid JSON (or a non-object) shows the
parse error inline instead of auto-rejecting; valid JSON is pretty-printed,
applied to the form/preview, and saved. The form editor still auto-saves.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Scenario editor now has a "World State (RPG)" section with an Editor/JSON
toggle. The Editor is a form UI to add/edit/remove world & player stats
(min/max/initial/±per-turn/cooldown/counter/desc/bands), NPCs (name, keys,
description, and their own stats), flags, and milestones — no JSON required.
Both views edit the same parsed schema and stay in sync; empty clears the
RPG layer. Keeps the read-only structured preview under the JSON view.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replace the single shared `npc` stat template (+ npc_card_types) with an
`npcs` section: each NPC keyed by a stable id, carrying its own name,
description, trigger keys, and its OWN stats block. The AI addresses NPCs
as npc.<id>.<stat> (id shown in context), which also fixes the old
card-id-guessing problem. On adventure creation each NPC auto-creates a
story card (name/keys/desc) for lore + in-scene detection, unless a
same-name card already exists. All NPCs instantiate up front.
- engine: npcs instantiate/apply/render/reference, npc_name/npc_triggers
- builder: _visible_npcs matches each NPC's own keys
- create_adventure: auto-create story cards from npcs
- WorldStateDrawer: render defined NPCs with their own stats + desc tooltip
- demo seed: Gwen (health/trust) + Bandit Leader (health/aggression)
- tests updated (34 pass); no new migration (npcs lives in stat_schema JSON)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Structured world/player/NPC stats, two-way flags, and sticky milestones
per scenario (stat_schema). The AI proposes a per-turn delta; a Python
engine referees it (clamp to min/max, per-turn cap, cooldown, counters).
Band word-labels plus a fixed stat guide (descriptions + full ranges)
keep the model grounded. World State drawer + Insights delta report;
undo/retry roll it back via the Phase 11 snapshot pattern.
- migrations 26-28 (scenarios.stat_schema, adventures.world_state,
actions.world_state_before); all nullable, additive, safe on existing rows
- migration 29 raises the default context budget 4096 -> 16384
(custom values preserved)
- seeded demo scenario 04-rpg-world-state.json (Bandit Camp)
- 19 new tests (33 total pass)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The shared per-adventure script_state ("scoreboard" scripts write to) was
never reverted by undo, and retry re-ran the output hook on top of the already-
mutated state, double-applying its changes (e.g. "+10 gold" became +20).
Each action now snapshots script_state as it was immediately before its own
hooks ran (new Action.state_before column, migration 25):
- undo restores the turn's first-action snapshot, prunes memories that
summarized the removed actions, and takes the turn lock against races.
- retry restores the AI action's snapshot before regenerating.
Story-card mutations are not reverted (documented limit). Adds the project's
first test suite: unit + full HTTP integration through the real scripting
engine (14 tests). See plan/11-state-revert-and-retry-fix.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Swap the single-line action <input> for a <textarea> that grows with its
content (capped at ~4 lines, then scrolls) and shrinks back after send.
Enter still sends; Shift+Enter inserts a newline. Buttons stay bottom-aligned
as it grows.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
The State drawer stringified non-string values with JSON.stringify, so
nested objects/arrays showed as raw JSON. Add a recursive StateValue /
StateTree renderer: primitives are typed and coloured, objects/arrays are
collapsible and indented so deep state shows its structure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Map HTTP 429 in _friendly_http_error: OpenRouter's free-models-per-day cap
gets a "daily limit, resets 00:00 UTC" message; other rate limits get a
generic "too many requests, try again" note. Avoids leaking the raw error
JSON (incl. user_id) to players.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Adventure scripts are frozen snapshots taken at creation. Add an opt-in
way to pull the latest library code into a live adventure:
- AdventureScript.source_script_id links a copy to its library Script
(migration 24; set at adventure creation)
- list endpoint flags out_of_date by diffing copy vs library
- POST /adventures/{id}/scripts/{sid}/sync overwrites the copy's code,
preserving enabled/position/script_state
- legacy copies with no link fall back to a name match, then adopt the link
- Play page shows a Sync button only when a copy differs from its library
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Trusted testers listed in AIDND_POWER_USERS get unmetered turns on the
shared demo key: demo_turns_left reports the full cap and count_demo_turn
skips them. Registered accounts only, matched case-insensitively.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Each script in the panel can now be downloaded as an import-compatible
JSON bundle (matching the /scripts export format), so demo/server-owned
scripts can be forked into a user's own library via the Scripts page's
Import.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
The panel only showed a name + on/off toggle, so demo (server-owned)
scripts — which never appear on the personal Scripts page — had no
in-app way to inspect their code. AdventureScriptOut already ships the
hook sources, so add a "View code" expander that renders the non-empty
hooks (library / onInput / onModelContext / onOutput) read-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
The reasoning model truncated replies before the end-of-reply [HP-N]
tag, so damage was never reported and HP stuck at 100. Move the tag to
the FIRST line of the reply so it survives truncation, and clarify the
wording (concrete example numbers instead of the literal "[HP-N]" that
confused the model).
Also make the seeder the source of truth for demo content: it now
reconciles an existing seeded scenario in place when a seed file
changes (keeping the scenario row and its adventure FKs), instead of
only inserting when missing — otherwise fixes to already-seeded demos
never reach production. No-op when content already matches, so it stays
cheap on every boot.
Verified end-to-end through the quickjs engine, including a truncated
reply still applying damage, and the in-place update path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Adds "[Demo] The Goblin Ambush (HP script)": a combat scene whose script
tells the AI the current HP and to tag damage/healing as [HP-N]/[HP+N];
the AI decides the amounts, the onOutput hook applies them to state.hp
(clamped 0..100), hides the tags, and flags death at 0. Showcases the
new left State drawer with live variables. Verified through the real
quickjs engine: a simulated attack drops HP, a potion heals (clamped),
a fatal blow sets state.dead.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Exposes the scripting `state` object (every variable scripts read/write
via state.x, persisted per-adventure) in a collapsible left rail that
refreshes after each turn. New owner-scoped GET
/api/adventures/{id}/script-state endpoint; the drawer only fetches
while open and shows an empty-state hint until a script sets a variable.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Add GET /api/story-cards/export and POST /api/story-cards/import for a
scenario or adventure, using the AI Dungeon world-info array shape
(keys/value/type/title/description/useForCharacterCreation) mapped to our
columns. Import accepts a bare array or {cards|storyCards|worldInfo:[...]},
enforces the per-owner cap (existing + incoming) and import rate limit.
Export/Import buttons added to the scenario editor and the in-play Plot
panel.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
- When a model streams reasoning but no story text, return a specific,
actionable error (raise Max output tokens / set a reasoning cap / use a
non-reasoning model) instead of the opaque "empty response".
- Bump default max_output_tokens 400 -> 800: 400 truncated scenes and
left reasoning models with no room after thinking. Affects new settings
rows; existing users keep their value.
- README: add a "Try it live" link to the Render deployment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
- Pin the play header (Plot/Memory/Scripts/Insights toggles) just below
the top nav so they stay reachable without scrolling back up to the
first message.
- Fix streaming autoscroll: snap to the real document bottom instantly
instead of smooth-scrolling to storyEndRef. That ref sits above the
sticky composer, so block:'end' stopped short — hiding the latest line
behind the input bar and yanking the reader back up when they scrolled
down. Smooth behavior also never settled at per-token speed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
list_adventures selected Scenario.title while grouping only by
Adventure.id. SQLite tolerates selecting an ungrouped column, so it
worked locally, but Postgres (Neon) rejects it — the adventures list
endpoint 500'd on the live deploy, so saved adventures could not be
listed or resumed. Add Scenario.title to the GROUP BY.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Idempotent seeder inserts public, unowned demo scenarios from
backend/app/seed_data/*.json if a public scenario with the same title
is missing, so a fresh Neon database (and any clone) gets demo content
guests can play immediately. Ships two scenarios: the Sunken Crypt of
Vharos (all four script hooks) and Signal from the Derelict (sci-fi,
context-engine focused).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Config via env, abuse/resource limits, and production serving so the app
is safe to expose publicly:
- Fail-fast on missing SECRET_KEY when MULTI_USER=true
- quickjs per-execution time/memory limits (while(true) can't hang server)
- Per-user/per-IP rate limiting on turn/script/auth endpoints
- Request body size limit + per-user row caps
- Security headers (CSP, X-Frame-Options, nosniff, referrer-policy) incl. SSE
- Debug router 403 and /docs disabled in multi-user mode
- DATABASE_URL support (defaults to Neon Postgres) alongside SQLite
- Documented all env vars in backend/.env.example
Verified locally via uvicorn (MULTI_USER=1, SQLite); see plan/09-phase-hardening.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs
unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me,
register upgrades the guest in place, login/logout, per-IP rate limits.
Every router scoped by user_id; Settings become per-user with the API
key Fernet-encrypted at rest and write-only through the API. Users
without a key get a server-funded demo key (OpenRouter free models,
20 turns/day, memory bank disabled on demo turns). Public read-only
demo scenarios (seed_demo.py); debug log restricted to local mode.
Frontend: auth modal + guest nudge, 401 re-establish/retry, demo
banner and key management in Settings.
Migrations 13-23 adopt existing data under a local user and encrypt
stored keys. Verified: migration on a copy of real data.db, two-session
isolation + register/login via curl and Chrome, demo cap 429, live
OpenRouter turn through the encrypted-key path, vite build + oxlint.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
Backend:
- provider: fall back to parsing a plain JSON body when a server ignores
stream=true (was: silent empty turn); error if response has no text (#5)
- memorybank: clamp cursors after undo/retry shrinks the action list, and
translate summary_cursor (list position) to an Action.index boundary
before comparing with Memory.source_end (#7)
- memorybank: pinned memories now count toward the top_k budget (#8)
- memorybank: cosine() returns 0.0 on dimension mismatch; changing the
embedding model clears stored vectors so they re-embed (#9)
- scripting: MAX_STORY_CARDS cap now counts cards inserted during the
hook, so a script can't add unbounded cards in one turn (#10)
- settings: /test tolerates non-dict JSON from /models (#13)
- scenarios: import accepts worldInformation as a story-card source (#14)
Frontend:
- per-key debounce timers in PlotPanel and ScenarioEditor — editing two
things within 600ms no longer drops the first save (#15, #16)
- Continue button no longer discards typed input (#17)
- failed retry resyncs actions from the server instead of leaving the
removed action missing (#18)
- Settings save/test surface errors instead of hanging on Testing… (#19)
- InsightsPanel ignores stale responses from superseded requests (#20)
- placeholder scan includes story-card trigger keys (#21)
addStoryCard returning the 0-based index (falsy for the first card) matches
real AI Dungeon per the scripting guidebook — kept, documented (#11).
Statuses updated in CODE_REVIEW_FINDINGS.md; stale entries for previously
fixed items (#1-4, #6, #12) corrected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
- LICENSE: MIT
- README: rewritten as portfolio-grade docs (features with code pointers,
architecture, Docker/Windows/macOS quick starts, BYO-model table)
- Dockerfile (3-stage: SPA build, pip wheels, slim runtime) + compose with
a /data volume; .dockerignore keeps secrets and local data out
- database.py: AIDND_DB_PATH env override so deployments can relocate the
SQLite file; documented in backend/.env.example
- start.sh: macOS/Linux dev script with first-run setup
Verified locally: production SPA build served by the backend (deep links
OK), DB created at the override path. Docker image itself untested here
(Docker not installed); flagged in plan/07.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
Phases: public repo & Docker (7), guest-first optional accounts (8),
production hardening (9), Render deploy (10). Confirmed decisions and
per-phase open questions recorded in each file.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg