Commit Graph
21 Commits
Author SHA1 Message Date
parththakkar106andClaude Opus 5 1940e222a7 Pin down what a tree must not change
Phase 14 replaces the action list with a tree, which touches the memory bank,
the context builder, undo, retry and the UI at once. Before any of that moves,
write down what "unaffected" means and prove it holds today.

tests/test_story_tree_baseline.py drives the product over HTTP the way a player
does and asserts only on API responses, never on how anything is stored: the
turn engine, retry and variant switching, undo, paging all the way back to the
start without gaps or repeats, edit/delete, memories, export/import round-trip,
world state. 283 green with it added, from 259.

It must pass unmodified through the schema change, the branch clause and the
move of memories onto nodes. A linear story is a tree with one branch, so those
subphases have no licence to change behaviour, and this is what says so.

The plan itself records four decisions that were open: structural-first with no
runtime flag, legacy columns kept one release, full tree visualisation, and a
v2 bundle with a v1 reader. Plus four traps found by reading rather than
running -- history._from_memory and the scripting pipeline both slice
adventure.actions, which under a tree is every branch rather than the path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
ParthandClaude Opus 5 0f1e05c808 Keep the fixture, so there is something to scroll (#5)
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.

--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:

  - create_all() builds the current schema but leaves the version stamp at its
    default, and bootstrap() reads a populated-but-unstamped database as
    ancient — it replays every migration against a schema that already has the
    columns, and fails on the first.
  - the fixture's user is a registered one, but local mode looks for the row
    with email IS NULL and is_guest false, so without clearing the email the
    app opens on an empty library.

--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.

Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.


Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 16:50:41 +05:30
parththakkar106andClaude Opus 5 6bf56ba272 Measure what the vacuum actually left behind
The post-vacuum sizes had never been read. They were 144.2 MB - above the
99.6 MB the compression work started from, not the ~53 MB it projected. The
earlier VACUUM FULL ran before migrations 42-45, which then rewrote every row
and refilled the table with dead space that only another VACUUM FULL returns.

Ran it: 144.2 -> 65.0 MB, 79.2 MB reclaimed in 5.5s, health green after.
actions is now 52.1 MB holding 50.4 MB of live bytes. The projection was right;
nothing had reclaimed what the compression freed.

Also records where those bytes are. context_snapshot is 94% of the table at
104 kB a row, and the per-action state columns a tree would sit beside are
rounding errors - which is the number phase 14 wants before it adds more.

And a trap: n_live_tup on the TOAST relation implied half the real bloat. It
is a leftover ANALYZE estimate, stale in the flattering direction right after
a migration. Size from sum(octet_length(col)) instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 16:16:45 +05:30
parththakkar106andClaude Opus 5 1030a7f234 Record what shipped, and what is still only believed
STATUS was written before the deploy and still asked for a VACUUM that has
since been run. Replaces that with what was actually verified against the
running service, and with the two things a next session would otherwise have
to rediscover.

The post-vacuum sizes were never measured, so the projection stays a
projection; the query to settle it is in the file. And the scroll gap is
promoted from a footnote to the one real open item, because re-reading that
path after shipping found a bug in it -- which is the argument that reading it
again is not the fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 15:26:29 +05:30
parththakkar106andClaude Opus 5 4f067516de Close plan/13 and point STATUS at the tree
Records what the six changes measured, and replaces the pick-up item -- which
was step 6 -- with plan/14, since there is nothing left in 13.

Two things a future session needs and cannot infer from the code. The VACUUM:
migration 43 compresses context_snapshot but Postgres does not return the disk
by itself, so until `VACUUM FULL actions` runs the storage win exists only on
paper and the table is temporarily larger, not smaller. And the gap: nothing
exercises the scroll behaviour in a browser, because the frontend has no test
runner, and prepend-and-restore-scroll is the part most likely to feel wrong
even when it is correct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 14:25:48 +05:30
parththakkar106andClaude Opus 5 85b188977e Correct the production URL: the deploy is ai-dnd-1gmp, not ai-dnd
The previous commit recorded the service as suspended. It is not. Render
appended a suffix to the `ai-dnd` service name in render.yaml, and plain
ai-dnd.onrender.com is a different, suspended service whose 503 page reads
"suspended by its owner" -- indistinguishable from this deploy being down
unless you notice the host is wrong.

The real host answers {"ok":true} on /api/health in 0.6s, mints a guest,
returns an empty adventure list for that guest and serves the starter
scenarios. The authoritative link is the one docs/index.html points at, not
the service name in the blueprint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 12:43:06 +05:30
parththakkar106andClaude Opus 5 70d024da62 Check the egress work against the database it actually runs on
Everything measured so far ran on SQLite against a synthetic fixture, so two
claims were still on trust: that migration 38 spells BYTEA correctly for a real
server, and that the byte figures survive psycopg's encodings.

Both hold. Production reads schema_version 41 with embedding_blob bytea and
embedded boolean present, the backfill is complete at 134/134, and the packed
vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a
memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and
against a throwaway Neon database every shape lands within 0.5% of the SQLite
run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of
it.

The harness writes, so it refuses any target whose name does not say stress or
scratch -- pointed at the production database it stops rather than seeding it
with a fake user and 200 fake turns. It also empties a Postgres target before
building, which a fresh SQLite temp file never needed.

Two corrections fall out, both recorded in plan/13. The page-load model has the
wrong shape: real actions are half the fixture's weight but real stories run to
607 actions, not 200, so the worst real page load is 589.5 kB. And the decision
to leave context_snapshot in the database costed egress but never storage --
it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the
ceiling this deploy will hit first.

Measured with counts and octet_length sums only. No user content was read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 12:37:23 +05:30
parththakkar106andClaude Opus 5 970b13998a Write down where the project stands
plan/STATUS.md: what the last session changed, what to pick up next, and the
handful of things that were learned the hard way and would otherwise have to
be rediscovered. Linked from the overview, which now also lists plans 13 and
14 alongside the earlier phases.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:31:50 +05:30
parththakkar106andClaude Opus 5 b7e53ae581 Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.

Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.

    one turn      3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
    run_post_turn 3,139.1 kB -> 0.7 kB
    Insights      3,223.7 kB -> 117.9 kB
    Memories drawer  ~3.1 MB -> 23.6 kB

A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.

The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.

memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.

Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:28:00 +05:30
parththakkar106andClaude Opus 5 7ee5ceea6c Measure database egress in bytes, from inside the repo
Both egress blowouts this project has had were one query fetching a column
nobody read, and a statement count would have shown nothing wrong in either.
So the meter counts bytes, at the DBAPI cursor -- everything that crosses
that line crossed the wire.

tools/stress_session.py drives a production-shaped adventure through the real
routes with only the network faked. It reproduces both figures measured
directly on production: 426.7 kB for a 200-action page load against 423 KB,
and 3,258.7 kB for one turn against 3,153 kB.

The memory bank is on by default, which is the whole point -- the previous
harness ran without an embedding model, so retrieval returned early and the
heaviest read in a turn never happened. --no-embeddings reproduces that
deliberately, and the gap is 29x.

It also turned up two callers the production SQL could not see: run_post_turn
walks the whole bank again every turn, and Insights pays for it a third time.
A played turn costs ~6.4 MB, not 3.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:11:56 +05:30
parththakkar106andClaude Opus 5 03c6707ca7 Write up the embedding-cost fix and the story-tree design
Two plan documents from designing the branching story tree. The tree work
turned up that memory retrieval fetches the whole bank's embeddings every
turn -- 96% of a turn's database traffic -- and that is the one read a tree
cannot window, so it lands first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:02:20 +05:30
parththakkar106andClaude Opus 4.8 cf464a48b0 Make NPCs a dedicated section with per-NPC stats
Replace the single shared `npc` stat template (+ npc_card_types) with an
`npcs` section: each NPC keyed by a stable id, carrying its own name,
description, trigger keys, and its OWN stats block. The AI addresses NPCs
as npc.<id>.<stat> (id shown in context), which also fixes the old
card-id-guessing problem. On adventure creation each NPC auto-creates a
story card (name/keys/desc) for lore + in-scene detection, unless a
same-name card already exists. All NPCs instantiate up front.

- engine: npcs instantiate/apply/render/reference, npc_name/npc_triggers
- builder: _visible_npcs matches each NPC's own keys
- create_adventure: auto-create story cards from npcs
- WorldStateDrawer: render defined NPCs with their own stats + desc tooltip
- demo seed: Gwen (health/trust) + Bandit Leader (health/aggression)
- tests updated (34 pass); no new migration (npcs lives in stat_schema JSON)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 15:17:47 +05:30
parththakkar106andClaude Opus 4.8 4dac445f90 Add Phase 12: native RPG world state
Structured world/player/NPC stats, two-way flags, and sticky milestones
per scenario (stat_schema). The AI proposes a per-turn delta; a Python
engine referees it (clamp to min/max, per-turn cap, cooldown, counters).
Band word-labels plus a fixed stat guide (descriptions + full ranges)
keep the model grounded. World State drawer + Insights delta report;
undo/retry roll it back via the Phase 11 snapshot pattern.

- migrations 26-28 (scenarios.stat_schema, adventures.world_state,
  actions.world_state_before); all nullable, additive, safe on existing rows
- migration 29 raises the default context budget 4096 -> 16384
  (custom values preserved)
- seeded demo scenario 04-rpg-world-state.json (Bandit Camp)
- 19 new tests (33 total pass)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 14:47:29 +05:30
parththakkar106andClaude Opus 4.8 dab1807118 Roll back script_state on undo/retry; fix retry double-apply
The shared per-adventure script_state ("scoreboard" scripts write to) was
never reverted by undo, and retry re-ran the output hook on top of the already-
mutated state, double-applying its changes (e.g. "+10 gold" became +20).

Each action now snapshots script_state as it was immediately before its own
hooks ran (new Action.state_before column, migration 25):
- undo restores the turn's first-action snapshot, prunes memories that
  summarized the removed actions, and takes the turn lock against races.
- retry restores the AI action's snapshot before regenerating.

Story-card mutations are not reverted (documented limit). Adds the project's
first test suite: unit + full HTTP integration through the real scripting
engine (14 tests). See plan/11-state-revert-and-retry-fix.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 02:58:38 +05:30
parththakkar106andClaude Opus 4.8 68c165585c Phase 10: Render deploy blueprint
- render.yaml: single Docker web service (SPA + API same-origin), free tier,
  Neon Postgres via AIDND_DATABASE_URL, generated AIDND_SECRET_KEY,
  /api/health check, us-east region, auto-deploy on main.
- README: "Deploy (Render)" section.
- Record verified Postgres path (real Neon, PG 18.4) and Phase 10 decisions
  (Neon, free tier, cloud Docker build) in plan/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
2026-07-07 14:07:08 +05:30
parththakkar106andClaude Opus 4.8 4772171b6c Phase 9: production hardening
Config via env, abuse/resource limits, and production serving so the app
is safe to expose publicly:

- Fail-fast on missing SECRET_KEY when MULTI_USER=true
- quickjs per-execution time/memory limits (while(true) can't hang server)
- Per-user/per-IP rate limiting on turn/script/auth endpoints
- Request body size limit + per-user row caps
- Security headers (CSP, X-Frame-Options, nosniff, referrer-policy) incl. SSE
- Debug router 403 and /docs disabled in multi-user mode
- DATABASE_URL support (defaults to Neon Postgres) alongside SQLite
- Documented all env vars in backend/.env.example

Verified locally via uvicorn (MULTI_USER=1, SQLite); see plan/09-phase-hardening.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
2026-07-07 12:30:06 +05:30
parththakkar106andClaude Fable 5 de4db373f2 Phase 8: optional accounts, per-user data, shared demo key
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs
unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me,
register upgrades the guest in place, login/logout, per-IP rate limits.
Every router scoped by user_id; Settings become per-user with the API
key Fernet-encrypted at rest and write-only through the API. Users
without a key get a server-funded demo key (OpenRouter free models,
20 turns/day, memory bank disabled on demo turns). Public read-only
demo scenarios (seed_demo.py); debug log restricted to local mode.
Frontend: auth modal + guest nudge, 401 re-establish/retry, demo
banner and key management in Settings.

Migrations 13-23 adopt existing data under a local user and encrypt
stored keys. Verified: migration on a copy of real data.db, two-session
isolation + register/login via curl and Chrome, demo cap 429, live
OpenRouter turn through the encrypted-key path, vite build + oxlint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 23:04:03 +05:30
parththakkar106andClaude Fable 5 3ee653a631 Plan: record successful Docker test (WSL2)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 17:06:49 +05:30
parththakkar106andClaude Fable 5 22630c8411 Phase 7: MIT license, portfolio README, Docker, cross-platform start
- LICENSE: MIT
- README: rewritten as portfolio-grade docs (features with code pointers,
  architecture, Docker/Windows/macOS quick starts, BYO-model table)
- Dockerfile (3-stage: SPA build, pip wheels, slim runtime) + compose with
  a /data volume; .dockerignore keeps secrets and local data out
- database.py: AIDND_DB_PATH env override so deployments can relocate the
  SQLite file; documented in backend/.env.example
- start.sh: macOS/Linux dev script with first-run setup

Verified locally: production SPA build served by the backend (deep links
OK), DB created at the override path. Docker image itself untested here
(Docker not installed); flagged in plan/07.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:51:48 +05:30
parththakkar106andClaude Fable 5 466d7bc0e4 Add public-release plan (phases 7-10)
Phases: public repo & Docker (7), guest-first optional accounts (8),
production hardening (9), Render deploy (10). Confirmed decisions and
per-phase open questions recorded in each file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:48:13 +05:30
parththakkar106andClaude Fable 5 db9f904222 Initial commit: AI Dungeon clone (FastAPI backend + React frontend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:08:19 +05:30