Commit Graph
24 Commits
Author SHA1 Message Date
parththakkar106andClaude Opus 5 05a2a77e4c Resolve the head once a flush, not once a node
The identity map holds weak references. A branch row nobody keeps a strong
reference to is collected between two nodes, so resolving the head inside the
placement loop read it back from the database for every node in the flush: 201
SELECTs on `branches` to write 200 actions, and 36 s -> 45 s on the same 297
tests. Nothing about any result changed, which is why only a stopwatch found
it, and why there is now a test counting the reads.

Also: the two reads left un-pathed on purpose say so where they live —
`max_action_index` allocates the legacy `index` and must stay adventure-wide
or two branches issue the same number, and export is a flat v1 bundle whose
reader has no idea branches exist. And `Adventure.actions` keeps its `index`
ordering, because ordering the collection by depth would not make it a story:
it is every branch's actions, and a path is a selection out of it.

318 tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 d3756abdaa Give every action a branch and a depth
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a
`branches` table, `branch_id`/`depth` on actions and memories, a head pointer
on adventures, migrations 46-52, and a server-side backfill that re-reads every
existing adventure as a tree with one branch. `depth` holds the number `index`
already held, gaps included, so no story changes — a linear story *is* a tree
with one branch, which is what makes the SP0 baseline passing unmodified the
pass condition rather than a hope.

The writer had to come with it. No migration will ever visit a row written
after it ran, so columns backfilled today and populated next subphase would
leave a hole exactly the width of one deploy, and from SP2 on a row without a
branch is a row no read can see. `app/tree.py` owns that: one module, because a
node written without a branch fails by disappearing rather than by raising.

Three things the schema itself insisted on:

- `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing
  both ways makes the two tables a cycle create_all cannot order, and its
  escape hatch needs an ALTER SQLite does not have. It is a cache, and a head
  naming a branch that is gone recovers onto the root.
- `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a
  second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`,
  because Postgres `json` has no equality operator.
- SQLite will not drop a column a foreign key names, which is how two existing
  tests broke: they simulated an old database by rewinding the stamp while
  leaving the new columns in place. Every ADD COLUMN migration is now
  idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by
  rebuilding three tables from frozen DDL so the real ALTERs run.

297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn;
page load and index are byte-identical to the recorded figures.

The deploy that ships this needs one `VACUUM FULL actions;` on the direct
endpoint afterwards — it rewrites every row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 5c1bcf7305 Build a fixture that can witness what a tree changes
The stress fixture is sized from production and exists to weigh bytes, so it
leaves every column it does not weigh at its default. Checked against a freshly
built one, that is exactly the set phase 14 has to migrate: state_before and
world_state_before NULL on all 600 rows, no scenario and so no RPG layer, no
adventure scripts, both cursors 0, and a hundred retry histories whose two
attempts carry byte-identical text with variant_index pinned at 0.

That last one matters most. "Which attempt is live?" is the question SP4's
migration answers when it decides which sibling node becomes the head of a
turn, and against that fixture the question had no observable answer -- a
migration that picked wrong would have looked correct.

--rich fills in those columns and no others. A real RPG scenario read from the
seed data rather than invented, with a world state played forward so hp
declines and flags flip; monotonic per-action state snapshots, so a rollback
that does not happen reads as a wrong number instead of as nothing; a gold
script, story cards, non-zero cursors, pinned and forgotten memories; retry
attempts with distinct texts, counts of two and three, and a live attempt that
is often not the last written. Plus a second adventure, because a branch clause
that forgot its adventure still looks right on a database holding one.

The invariant SP4 reads -- text mirrors variants[variant_index] -- is asserted
at build time rather than assumed.

The plain fixture is untouched and re-measured unchanged at 1.8 kB, so the
egress ceilings stay comparable. All six shapes run clean on --rich, including
the turn path with the RPG layer and script pipeline now live. 283 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 1940e222a7 Pin down what a tree must not change
Phase 14 replaces the action list with a tree, which touches the memory bank,
the context builder, undo, retry and the UI at once. Before any of that moves,
write down what "unaffected" means and prove it holds today.

tests/test_story_tree_baseline.py drives the product over HTTP the way a player
does and asserts only on API responses, never on how anything is stored: the
turn engine, retry and variant switching, undo, paging all the way back to the
start without gaps or repeats, edit/delete, memories, export/import round-trip,
world state. 283 green with it added, from 259.

It must pass unmodified through the schema change, the branch clause and the
move of memories onto nodes. A linear story is a tree with one branch, so those
subphases have no licence to change behaviour, and this is what says so.

The plan itself records four decisions that were open: structural-first with no
runtime flag, legacy columns kept one release, full tree visualisation, and a
v2 bundle with a v1 reader. Plus four traps found by reading rather than
running -- history._from_memory and the scripting pipeline both slice
adventure.actions, which under a tree is every branch rather than the path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
ParthandClaude Opus 5 0f1e05c808 Keep the fixture, so there is something to scroll (#5)
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.

--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:

  - create_all() builds the current schema but leaves the version stamp at its
    default, and bootstrap() reads a populated-but-unstamped database as
    ancient — it replays every migration against a schema that already has the
    columns, and fails on the first.
  - the fixture's user is a registered one, but local mode looks for the row
    with email IS NULL and is_guest false, so without clearing the email the
    app opens on an empty library.

--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.

Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.


Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 16:50:41 +05:30
parththakkar106andClaude Opus 5 6bf56ba272 Measure what the vacuum actually left behind
The post-vacuum sizes had never been read. They were 144.2 MB - above the
99.6 MB the compression work started from, not the ~53 MB it projected. The
earlier VACUUM FULL ran before migrations 42-45, which then rewrote every row
and refilled the table with dead space that only another VACUUM FULL returns.

Ran it: 144.2 -> 65.0 MB, 79.2 MB reclaimed in 5.5s, health green after.
actions is now 52.1 MB holding 50.4 MB of live bytes. The projection was right;
nothing had reclaimed what the compression freed.

Also records where those bytes are. context_snapshot is 94% of the table at
104 kB a row, and the per-action state columns a tree would sit beside are
rounding errors - which is the number phase 14 wants before it adds more.

And a trap: n_live_tup on the TOAST relation implied half the real bloat. It
is a leftover ANALYZE estimate, stale in the flattering direction right after
a migration. Size from sum(octet_length(col)) instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 16:16:45 +05:30
parththakkar106andClaude Opus 5 1030a7f234 Record what shipped, and what is still only believed
STATUS was written before the deploy and still asked for a VACUUM that has
since been run. Replaces that with what was actually verified against the
running service, and with the two things a next session would otherwise have
to rediscover.

The post-vacuum sizes were never measured, so the projection stays a
projection; the query to settle it is in the file. And the scroll gap is
promoted from a footnote to the one real open item, because re-reading that
path after shipping found a bug in it -- which is the argument that reading it
again is not the fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 15:26:29 +05:30
parththakkar106andClaude Opus 5 4f067516de Close plan/13 and point STATUS at the tree
Records what the six changes measured, and replaces the pick-up item -- which
was step 6 -- with plan/14, since there is nothing left in 13.

Two things a future session needs and cannot infer from the code. The VACUUM:
migration 43 compresses context_snapshot but Postgres does not return the disk
by itself, so until `VACUUM FULL actions` runs the storage win exists only on
paper and the table is temporarily larger, not smaller. And the gap: nothing
exercises the scroll behaviour in a browser, because the frontend has no test
runner, and prepend-and-restore-scroll is the part most likely to feel wrong
even when it is correct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 14:25:48 +05:30
parththakkar106andClaude Opus 5 85b188977e Correct the production URL: the deploy is ai-dnd-1gmp, not ai-dnd
The previous commit recorded the service as suspended. It is not. Render
appended a suffix to the `ai-dnd` service name in render.yaml, and plain
ai-dnd.onrender.com is a different, suspended service whose 503 page reads
"suspended by its owner" -- indistinguishable from this deploy being down
unless you notice the host is wrong.

The real host answers {"ok":true} on /api/health in 0.6s, mints a guest,
returns an empty adventure list for that guest and serves the starter
scenarios. The authoritative link is the one docs/index.html points at, not
the service name in the blueprint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 12:43:06 +05:30
parththakkar106andClaude Opus 5 70d024da62 Check the egress work against the database it actually runs on
Everything measured so far ran on SQLite against a synthetic fixture, so two
claims were still on trust: that migration 38 spells BYTEA correctly for a real
server, and that the byte figures survive psycopg's encodings.

Both hold. Production reads schema_version 41 with embedding_blob bytea and
embedded boolean present, the backfill is complete at 134/134, and the packed
vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a
memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and
against a throwaway Neon database every shape lands within 0.5% of the SQLite
run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of
it.

The harness writes, so it refuses any target whose name does not say stress or
scratch -- pointed at the production database it stops rather than seeding it
with a fake user and 200 fake turns. It also empties a Postgres target before
building, which a fresh SQLite temp file never needed.

Two corrections fall out, both recorded in plan/13. The page-load model has the
wrong shape: real actions are half the fixture's weight but real stories run to
607 actions, not 200, so the worst real page load is 589.5 kB. And the decision
to leave context_snapshot in the database costed egress but never storage --
it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the
ceiling this deploy will hit first.

Measured with counts and octet_length sums only. No user content was read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 12:37:23 +05:30
parththakkar106andClaude Opus 5 970b13998a Write down where the project stands
plan/STATUS.md: what the last session changed, what to pick up next, and the
handful of things that were learned the hard way and would otherwise have to
be rediscovered. Linked from the overview, which now also lists plans 13 and
14 alongside the earlier phases.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:31:50 +05:30
parththakkar106andClaude Opus 5 b7e53ae581 Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.

Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.

    one turn      3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
    run_post_turn 3,139.1 kB -> 0.7 kB
    Insights      3,223.7 kB -> 117.9 kB
    Memories drawer  ~3.1 MB -> 23.6 kB

A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.

The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.

memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.

Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:28:00 +05:30
parththakkar106andClaude Opus 5 7ee5ceea6c Measure database egress in bytes, from inside the repo
Both egress blowouts this project has had were one query fetching a column
nobody read, and a statement count would have shown nothing wrong in either.
So the meter counts bytes, at the DBAPI cursor -- everything that crosses
that line crossed the wire.

tools/stress_session.py drives a production-shaped adventure through the real
routes with only the network faked. It reproduces both figures measured
directly on production: 426.7 kB for a 200-action page load against 423 KB,
and 3,258.7 kB for one turn against 3,153 kB.

The memory bank is on by default, which is the whole point -- the previous
harness ran without an embedding model, so retrieval returned early and the
heaviest read in a turn never happened. --no-embeddings reproduces that
deliberately, and the gap is 29x.

It also turned up two callers the production SQL could not see: run_post_turn
walks the whole bank again every turn, and Insights pays for it a third time.
A played turn costs ~6.4 MB, not 3.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:11:56 +05:30
parththakkar106andClaude Opus 5 03c6707ca7 Write up the embedding-cost fix and the story-tree design
Two plan documents from designing the branching story tree. The tree work
turned up that memory retrieval fetches the whole bank's embeddings every
turn -- 96% of a turn's database traffic -- and that is the one read a tree
cannot window, so it lands first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:02:20 +05:30
parththakkar106andClaude Opus 4.8 cf464a48b0 Make NPCs a dedicated section with per-NPC stats
Replace the single shared `npc` stat template (+ npc_card_types) with an
`npcs` section: each NPC keyed by a stable id, carrying its own name,
description, trigger keys, and its OWN stats block. The AI addresses NPCs
as npc.<id>.<stat> (id shown in context), which also fixes the old
card-id-guessing problem. On adventure creation each NPC auto-creates a
story card (name/keys/desc) for lore + in-scene detection, unless a
same-name card already exists. All NPCs instantiate up front.

- engine: npcs instantiate/apply/render/reference, npc_name/npc_triggers
- builder: _visible_npcs matches each NPC's own keys
- create_adventure: auto-create story cards from npcs
- WorldStateDrawer: render defined NPCs with their own stats + desc tooltip
- demo seed: Gwen (health/trust) + Bandit Leader (health/aggression)
- tests updated (34 pass); no new migration (npcs lives in stat_schema JSON)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 15:17:47 +05:30
parththakkar106andClaude Opus 4.8 4dac445f90 Add Phase 12: native RPG world state
Structured world/player/NPC stats, two-way flags, and sticky milestones
per scenario (stat_schema). The AI proposes a per-turn delta; a Python
engine referees it (clamp to min/max, per-turn cap, cooldown, counters).
Band word-labels plus a fixed stat guide (descriptions + full ranges)
keep the model grounded. World State drawer + Insights delta report;
undo/retry roll it back via the Phase 11 snapshot pattern.

- migrations 26-28 (scenarios.stat_schema, adventures.world_state,
  actions.world_state_before); all nullable, additive, safe on existing rows
- migration 29 raises the default context budget 4096 -> 16384
  (custom values preserved)
- seeded demo scenario 04-rpg-world-state.json (Bandit Camp)
- 19 new tests (33 total pass)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 14:47:29 +05:30
parththakkar106andClaude Opus 4.8 dab1807118 Roll back script_state on undo/retry; fix retry double-apply
The shared per-adventure script_state ("scoreboard" scripts write to) was
never reverted by undo, and retry re-ran the output hook on top of the already-
mutated state, double-applying its changes (e.g. "+10 gold" became +20).

Each action now snapshots script_state as it was immediately before its own
hooks ran (new Action.state_before column, migration 25):
- undo restores the turn's first-action snapshot, prunes memories that
  summarized the removed actions, and takes the turn lock against races.
- retry restores the AI action's snapshot before regenerating.

Story-card mutations are not reverted (documented limit). Adds the project's
first test suite: unit + full HTTP integration through the real scripting
engine (14 tests). See plan/11-state-revert-and-retry-fix.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 02:58:38 +05:30
parththakkar106andClaude Opus 4.8 68c165585c Phase 10: Render deploy blueprint
- render.yaml: single Docker web service (SPA + API same-origin), free tier,
  Neon Postgres via AIDND_DATABASE_URL, generated AIDND_SECRET_KEY,
  /api/health check, us-east region, auto-deploy on main.
- README: "Deploy (Render)" section.
- Record verified Postgres path (real Neon, PG 18.4) and Phase 10 decisions
  (Neon, free tier, cloud Docker build) in plan/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
2026-07-07 14:07:08 +05:30
parththakkar106andClaude Opus 4.8 4772171b6c Phase 9: production hardening
Config via env, abuse/resource limits, and production serving so the app
is safe to expose publicly:

- Fail-fast on missing SECRET_KEY when MULTI_USER=true
- quickjs per-execution time/memory limits (while(true) can't hang server)
- Per-user/per-IP rate limiting on turn/script/auth endpoints
- Request body size limit + per-user row caps
- Security headers (CSP, X-Frame-Options, nosniff, referrer-policy) incl. SSE
- Debug router 403 and /docs disabled in multi-user mode
- DATABASE_URL support (defaults to Neon Postgres) alongside SQLite
- Documented all env vars in backend/.env.example

Verified locally via uvicorn (MULTI_USER=1, SQLite); see plan/09-phase-hardening.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
2026-07-07 12:30:06 +05:30
parththakkar106andClaude Fable 5 de4db373f2 Phase 8: optional accounts, per-user data, shared demo key
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs
unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me,
register upgrades the guest in place, login/logout, per-IP rate limits.
Every router scoped by user_id; Settings become per-user with the API
key Fernet-encrypted at rest and write-only through the API. Users
without a key get a server-funded demo key (OpenRouter free models,
20 turns/day, memory bank disabled on demo turns). Public read-only
demo scenarios (seed_demo.py); debug log restricted to local mode.
Frontend: auth modal + guest nudge, 401 re-establish/retry, demo
banner and key management in Settings.

Migrations 13-23 adopt existing data under a local user and encrypt
stored keys. Verified: migration on a copy of real data.db, two-session
isolation + register/login via curl and Chrome, demo cap 429, live
OpenRouter turn through the encrypted-key path, vite build + oxlint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 23:04:03 +05:30
parththakkar106andClaude Fable 5 3ee653a631 Plan: record successful Docker test (WSL2)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 17:06:49 +05:30
parththakkar106andClaude Fable 5 22630c8411 Phase 7: MIT license, portfolio README, Docker, cross-platform start
- LICENSE: MIT
- README: rewritten as portfolio-grade docs (features with code pointers,
  architecture, Docker/Windows/macOS quick starts, BYO-model table)
- Dockerfile (3-stage: SPA build, pip wheels, slim runtime) + compose with
  a /data volume; .dockerignore keeps secrets and local data out
- database.py: AIDND_DB_PATH env override so deployments can relocate the
  SQLite file; documented in backend/.env.example
- start.sh: macOS/Linux dev script with first-run setup

Verified locally: production SPA build served by the backend (deep links
OK), DB created at the override path. Docker image itself untested here
(Docker not installed); flagged in plan/07.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:51:48 +05:30
parththakkar106andClaude Fable 5 466d7bc0e4 Add public-release plan (phases 7-10)
Phases: public repo & Docker (7), guest-first optional accounts (8),
production hardening (9), Render deploy (10). Confirmed decisions and
per-phase open questions recorded in each file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:48:13 +05:30
parththakkar106andClaude Fable 5 db9f904222 Initial commit: AI Dungeon clone (FastAPI backend + React frontend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:08:19 +05:30