Commit Graph
62 Commits
Author SHA1 Message Date
parththakkar106andClaude Opus 5 20f0753277 Merge main into the Phase 17 refactor
`main` changed the four files this branch split into packages, so all four
came back as modify/delete conflicts. Each change is ported to the module
that now holds the code, unchanged in behavior:

- `delete_action` moves to `routers/adventures/actions.py`. It reaches the
  turn lock as `turns.acquire_turn_lock` and `turns._active_turns`, which is
  the rule this branch set: the package root no longer re-exports the names a
  test rebinds.
- `EMIT_RULE` moves to `worldstate/parse.py`, and `_describe_stat` and
  `render_reference` to `worldstate/render.py`.
- The chip-wrap rules move to `styles/schema-editor.css`, and
  `StateChangeChips` to `pages/Play/reports.jsx`. `index.css` keeps this
  branch's import list.

`test_delete_state.py` needed four edits to run here. It drops the eight-line
database prologue, because `conftest.py` does that once now. It imports the
shared `ScriptedProvider` from `fakes.py` instead of carrying a copy. It
patches `adventures.turns.OpenAICompatibleProvider` and reaches the lock the
same way. And it no longer passes `index=0` when it builds the start action,
because migration 71 drops that column.

564 tests pass. `npm run lint` and `npm run build` are clean, and the CSS
bundle is 56.71 kB, which is the pre-merge size plus main's new rules.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-30 18:03:36 +05:30
Claude 91907a30bd Name every stat in the guide by the path the AI has to write
The referee refused `npc.<id>.status` on a character that has a status stat
under another name, and the refusal was right: the model was reaching for a
path the guide never gave it.

The guide is the fixed list of everything a scenario tracks, and it named
stats in prose. A player stat and a world stat both read as a bare name, and
an NPC's stats read as the display name plus the stat — "Trainer Milo
active_status". The model had to turn that back into `npc.milo.active_status`
itself, and in the Pokemon demo five other characters carry a stat called
`status`, so the path it built was `npc.milo.status`. The live values do state
the paths, but only for the NPCs a scene has mentioned, so an NPC off screen
was addressable only by guesswork.

Each line now leads with the path: `player.potions`, `npc.milo.active_status`,
`flags.sandstorm_active`. An NPC's header is written even when it has no
description, because it is the one line that ties a display name to its id.
`EMIT_RULE` points at the guide for paths rather than at the live values
alone.

A free-text stat is marked `(free text)` beside its path, which is the wording
`EMIT_RULE` already used to describe it — the guide had been writing "free
text" at the end of the line instead. Such a stat now also gets a line when it
has no description, where before it was dropped and the model was left to send
a number for a stat that holds a string.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST
2026-08-30 06:46:12 +00:00
Claude 41ed2f63b5 Put the shared state back when a turn is deleted
Deleting an AI reply removed the text and left everything the turn did to the
numbers standing. `script_state` and `world_state` belong to the adventure
rather than to the node, so the row going away takes nothing back. Undo, retry,
a take and a branch switch all restore them; this endpoint never did.

The visible symptom was the cooldown clock, which lives in
`world_state._meta.last_changed` and holds a depth. A deleted turn left its
depth there, and the turn played in its place is played at that same depth, so
the referee refused the change as one that had happened this very turn — on a
turn the story no longer contains. Delete the reply because you did not like
the stat change it proposed, press Continue, and the same change comes back
marked "changed too recently". A script's state stacked the same way: the gold
a deleted turn paid out stayed paid, and the replacement turn paid it again.

The restore reads the tip's own outcome rather than the deleted node's
neighbour, which is what `switch` does. Deleting the newest turn then rewinds
to the node in front of it, and deleting one from the middle of the story
restores the state the adventure is already in, so the text goes and the
numbers stay.

The endpoint now takes the turn lock too, for the reason undo takes it: it
writes the shared state, and a turn that is still generating is about to write
it as well.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017imUKVPhqophVUZZwJK7ST
2026-08-30 11:30:17 +05:30
parththakkar106andClaude Opus 5 f1bebe18d0 Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`,
`variant_count`, `state_before`, and `world_state_before`, plus
`adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in
SQLite, so migration 71 quotes it.

Nothing outside the migrations read these. `models.py`, `schemas.py`, and
`ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders
by `id`, and `attempts.renumber`, `context.history.max_action_index`, and
`nodes.next_index` are deleted.

Two changes keep the migration replayable on a `create_all` database:

- `_split_variants_into_siblings` wrote through the live ORM table, so it
  stopped compiling once migration 66 removed five of its columns. It now
  writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`.
- Five data passes read columns these migrations drop. Each now calls
  `_has_columns` and returns early when the columns are absent.

`bootstrap` takes a `through` version so a migration test can stop at the
schema it asserts on.

555 tests pass, up from 549. Eight of the new cases assert each column is
gone after a real schema-45 database migrates all the way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK
2026-08-29 17:15:49 +05:30
parththakkar106andClaude Opus 5 2c57b1ceab Remove three kinds of duplication in the backend
Stage 2, items 1, 4, and 5 of `plan/17-refactor.md`.

**One path resolver in `worldstate`.** `apply_delta` and `apply_override` routed
`flags.<name>`, `milestones.<id>`, `world.<stat>`, `player.<stat>`, and
`npc.<id>.<stat>` with parallel code, about 100 lines each. `_resolve` now says
what a path points at and returns either a target or the rejection to report.
Each function keeps its own write rule, because the rules genuinely differ: an
override sets a number rather than adding to it, ignores `cooldown`,
`max_delta_per_turn`, and the rule that a counter only counts up, and can un-set
a milestone.

A differential check ran both implementations over 3960 payloads: twenty paths,
fourteen values, three starting states, plus every three-path combination. The
results are identical except that 674 rejections from `apply_override` now carry
a `fix` string. `apply_delta` already worded those, and the world-state editor
renders them, so an override that names an unknown flag now explains itself the
way a delta does.

**`sse`, `SSE_HEADERS`, and `turn_error` move to `app/sse.py`.** Two routers
stream, and `chat.py` had to import from `routers.adventures` to reach them.

**`get_adventure_or_404` becomes the `current_adventure` dependency.** All 32
handlers repeated the call as their first statement. The ownership check now
reads in the signature and runs before the body. FastAPI caches a dependency for
one request, so the handler's `db` is the session the adventure came from.

The generated OpenAPI document is byte-identical except on `rename_branch`,
where `branch_id` is now listed before `adventure_id`, because that handler no
longer names `adventure_id` itself. Parameter order in the document is
cosmetic.

Six tests in `test_state_revert.py` call `undo_turn` and `retry_action`
directly rather than over HTTP. They pass the adventure they already hold
instead of an id.

549 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
2026-08-29 02:04:02 +05:30
parththakkar106andClaude Opus 5 2fa812c056 Split the adventures router into a package
`backend/app/routers/adventures.py` held 2353 lines and 35 endpoints. It is now
a package of 14 modules, the largest 443 lines.

The split moves text rather than rewriting it. An AST comparison against the
old file confirms all 86 definitions are identical, and the OpenAPI schema
still lists the same 35 operations.

Names a test replaces now live in `turns.py` only, and other modules reach them
as `turns.<name>`. Rebinding a re-exported alias changes the alias and leaves
every caller reading the original, so the package root does not re-export them.
A patch aimed at the old target raises `AttributeError` instead of passing while
doing nothing. Tests and the fixtures in `backend/tools/` say
`adventures.turns.<name>`.

The same rule keeps the turn lock working. One module owns `_active_turns`, so
one lock guards one set.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
2026-08-29 01:05:50 +05:30
parththakkar106andClaude Opus 5 32cd7c1077 Give the tests one setup instead of thirty-five
Every test module carried the same eight-line prologue redirecting the
database to a temp file. Only the first one to be imported ever took
effect: `app.database` reads `AIDND_DB_PATH` at import and builds `engine`
from it once, so by the time the second module ran the engine already
existed. The other 34 copies created a temp file that nothing opened and
nothing deleted, and leaked one per module per run.

`conftest.py` now does it once, which is early enough because pytest
imports conftest before any test module. It also deletes the file when the
run ends. The tests still share one database, exactly as they already did:
each `client` fixture calls `create_all` on setup and `drop_all` on
teardown, so no test sees another test's rows.

`tests/fakes.py` holds the one `ScriptedProvider`. Nine modules each had a
copy, and the copies had drifted into four feature sets, so a test that
needed to raise a provider error had to be written in one of the files
whose copy supported that. The shared one is the superset. The two
`FakeProvider` copies were the same class with a fixed reply, so they use
it too. `test_chat.py` keeps its own, which implements `chat` rather than
`generate` and records what it was constructed with.

An autouse fixture resets the fake's class state between tests, so a stale
reply list can no longer reach the next test.

435 lines out of the suite. 549 tests pass. Verified live by sabotage:
breaking the shared fake fails 13 tests across four modules.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
2026-08-29 00:47:46 +05:30
parththakkar106andClaude Opus 5 9398c13da5 Delete seeded scenarios no seed file claims any more
`previous_titles` stops a rename stranding the row it left behind, but the rows
already stranded still had to be deleted by hand on every deployment. The
seeder now removes them on the next boot, which retires the stale "Road to the
Champion" demo without a database console.

Only rows with a NULL owner and `is_public` are considered, and a player's own
scenario is neither, so nothing anybody created is reachable. An adventure
started from a deleted demo survives: `adventures.scenario_id` is `ON DELETE
SET NULL`, and the adventure holds its own copies of the cards and scripts, so
it loses only the inherited cover art.

Two cases skip the sweep, because neither is an instruction to remove live
content: a seed file that fails to parse claims no title, and an empty seed
directory reads as a packaging failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 19:19:23 +05:30
parththakkar106andClaude Opus 5 c2d3f0d8b9 Show the starter adventure the artwork of the demo it came from
An adventure has no cover art of its own and inherits its scenario's, and a
bundle carries no scenario id, because an id means nothing in another database.
The starter card therefore fell back to a monogram tile while the demo beside
it showed the Pokeball.

The starter file names its source under `scenarioTitle`, and the copy is linked
to the seeded scenario with that title. If no seed answers to the name, the
adventure keeps a NULL `scenario_id`, which is the state every imported bundle
is in and costs only the artwork.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 19:19:23 +05:30
parththakkar106andClaude Opus 5 6cbf6d996f Give every new guest an adventure that is already played
An empty account gives a visitor nothing to read, and the daily demo turns are
limited, so learning what the app does cost one of them. `app/starter.py` now
copies a shipped export bundle into each new guest at the point the row is
created. The bundle is two exchanges of the Pokemon demo, which ends on a
knockout and shows an applied change, a refused one, and a milestone.

The guest row is committed before the copy is attempted, so a failure there
still leaves them with an account, and the copy runs inside a savepoint.

The row building that `POST /adventures/import` did inline moved into
`bundle.materialize`, which both callers use. The rate and size checks stayed
in the endpoint: the starter writes a file the server ships, so it has no
untrusted list to cap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 18:42:15 +05:30
parththakkar106andClaude Opus 5 1988979aeb Store the refusals, so the chips and the model can see them
`world_delta_of` wrote `delta` and `applied` only. Every consumer that tells a
refused change from a successful one reads the two lists it dropped:
`Action.world_changes` marks a clamped chip from `clamped` and builds its
refusal chips from `rejected`, and `worldstate.refusals` reads both. So no chip
could report a limit, no rejection chip could appear, and no correction ever
reached the next prompt. The three mechanisms merged last session were live in
the code and unreachable in production.

Found by playing the Pokemon demo. Turn 3's snapshot held two clamped entries
with correct `fix` text, both chips came back `clamped: false`, and turn 4's
prompt carried no correction, so the model repeated the same mistake.

The 21 tests passed because `action()` built the column by hand with every list
present. It now fills the column through `world_delta_of`. Removing the two new
lines fails 9 of the 23 tests; that was checked by sabotage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 17:22:02 +05:30
parththakkar106andClaude Opus 5 47c7800903 Report the world-state changes the engine refuses
`apply_delta` records three outcomes for every change the model sends:
`applied`, `clamped`, and `rejected`. Everything downstream read only
`applied`. A refused change reached the player as an ordinary chip, and
reached the model on the next turn as a change that had succeeded.

Five parts:

- `Action.world_changes` reads `clamped` and `rejected` beside `applied`.
  Accepted stats carry a `clamped` flag; refusals become `kind: "rejected"`
  entries. The `fix` key is present only when the engine wrote one, because
  this property runs for every action of every list response.
- The UI separates the three outcomes. A clamp to a standstill reads
  `no change - at its limit` on a dashed chip, a partial clamp is marked
  `(limited)`, and a rejection carries its reason. Dashed and dimmed rather
  than red: a refused change means the rules are working.
- The goals line names the milestone id, as `milestones.<id>`. The ids
  appeared nowhere in the prompt before, so the model could not send one.
- Each rejection, and each clamp that moved nothing, builds a `fix` string
  from the stat definition at the point of refusal. `render_refusals()`
  renders them into the next prompt above `EMIT_REMINDER`.
- `_history_text` replays `applied_delta()` instead of the sent delta, so a
  past turn's state block shows only what the engine accepted.

A clamp that reduced a change but still moved the value reports nothing. If
you tell a model its 80 damage became 30, it can treat the shortfall as a
debt and send the remaining 50 next turn, which is the swing
`max_delta_per_turn` prevents.

In the demo scenario, `pokemon_left` becomes `pokemon_fainted`
(`type: counter`, `initial: 0`). Starting at the ceiling turned a wrong-signed
delta into a silent no-op; counting up puts the wrong sign on the counter
rule, which refuses it out loud. The instructions also now ask for
`world.turn`, which sat at 0 for a whole playtest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 16:26:34 +05:30
Parth e7d75c3b05 Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style

Rewrite the comments and docstrings across the backend core modules so they
read plainly. The previous prose was accurate but dense and figurative, which
made it slow to skim.

Applies the Google developer documentation style guide: short sentences, active
voice, present tense, American spelling, and no metaphors, idioms, or
rhetorical asides. Replaces em-dash chains with separate sentences.
2026-08-26 15:37:25 +05:30
parththakkar106andClaude Opus 5 a408c7b6f7 Lay the prompt out so the endpoint can cache most of it
Prompt caching bills on a shared prefix: the endpoint reuses the request up
to the first byte that differs from last time and no further. The live
world-state block sat third from the top of the system message, so every turn
re-priced the instructions, the plot essentials and the whole story history
underneath it. The retrieved memories and the rewritten summary did it again.

Everything fixed is emitted first now, and everything that moves goes after
the history, ordered least-volatile first — which is also where recency serves
it best, the reasoning that already put the emit reminder last. The three tail
sections that are last for their own reasons stay last. The moved sections are
still charged to the token budget; only their position changed.

Two smaller halves of the same problem. OpenRouter serves a model from
whichever upstream is free and each upstream holds its own cache, so a
deepseek model now names deepseek as its preferred upstream — a preference,
not a restriction, so a turn still runs if that upstream is down. And the
endpoint's usage block is read back off the response and kept per attempt, so
the hit rate shows up in Insights and the debug log instead of being assumed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-23 06:31:30 +05:30
parththakkar106andClaude Opus 5 041f9e25f3 Count the visits, and say whether anyone got anywhere
A hosted demo raises a question a local app never does: is anyone using it,
and do they reach the part that matters? `/analytics` answers it — visitors,
pages, referrers, countries, devices, which shared scenarios get played, turns
and demo-key spend, API and turn errors, and a funnel from visited to played a
turn to signed up.

Not a third-party script, for reasons specific to this one. The CSP allows
`script-src 'self'`, so a tracker means loosening it; adblockers eat the
popular ones, which silently biases exactly the technical audience this
project gets shown to; and none of them can see the measurement that actually
matters here, which is a turn, not a pageview.

**A visit is a write and never a read.** After the 189x egress fix it would be
perverse to add a feature that reads rows per request, so counts accumulate in
a process-local dict and flush every 60s as UPSERTs. Storage is a generic
`(day, metric, label) -> hits` counter, so measuring something new later costs
a constant rather than a migration, plus one row per visitor per day for the
funnel flags. Every dashboard query is a GROUP BY returning tens of rows
however much traffic sits behind it; a month reads back in a few kilobytes.
The buffer's cost is that a hard restart can lose up to a minute — the flusher
also runs on shutdown, and a tier that sleeps when idle sleeps on an empty
buffer anyway.

**The counters are anonymous; the access log beside them is not, on purpose.**
A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars —
one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined
back to `users`, and keyed, so no client can compute one. Story content never
reaches that module, and the only content it ever names is a seeded public
scenario's title; a player's own titles are theirs. `accesslog.py` is the
identifying half and is a separate module writing a separate table so that
separation is a property of the code rather than a convention: `access_events`
records sessions, sign-ins, registrations and failed attempts with address,
email and device, read on a second tab of the same page behind the same gate.

Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An
unmetered tester is not automatically someone who should see the traffic. The
route 404s and the nav link is absent for everyone else, the same treatment
AI Chat gets; unset in a hosted deploy means nobody sees it, including me.

Three things came out of building it that a test would not have suggested.

**A failed turn is an HTTP 200 with a bad ending.** The status-code middleware
cannot see one, so a demo whose model had started refusing every request would
look perfectly healthy from outside. All five SSE error paths in
`_generate_turn` now go through a `turn_error()` helper that counts on the way
out. Error buckets elsewhere are labelled by the matched route template rather
than the requested path — one bucket per endpoint instead of one per adventure
id, and, the reason it isn't merely tidier, an unmatched path is entirely
attacker-chosen, so labelling by it would let anyone mint rows.

**The funnel counts people, not clicks.** A player who starts six adventures
is one person who started an adventure. That is the whole reason the
per-visitor-day table exists; its flags only ever turn on, and `is_new` is
settled by the first write of a visitor's first day.

**The tests run on SQLite and production is Neon.** A flush that raises is
caught and logged, so a dialect mistake in the UPSERTs would have stayed
invisible until the dashboard quietly never filled.
`test_the_upserts_compile_for_postgres` compiles both statements against the
Postgres dialect without connecting to one.

Two things this leans on elsewhere. `limits._client_ip` is now public
`client_ip`: the access log needs the same answer, and two functions both
deciding which hop is the caller's is how one of them ends up trusting a
header it shouldn't. And the cleanup sweeper now starts if *either* job has
work — a deployment can keep every guest forever and still want its
visitor-day rows aged out.

No migration. Both tables are new and `bootstrap()` calls `create_all` on
existing databases too, the route `branches` took in Phase 14, so
`LATEST_VERSION` is still 64.

497 tests green, frontend lint and build clean, driven by hand against a
synthetic 90-day fixture at 1568px. The narrow-screen layout follows the
existing 720px block but is unverified: `resize_window` is ignored on a
maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized
iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run
through the test names themselves, leaving `testclient_ip_*` — still collected
by pytest, which is why it passed unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-22 16:24:42 +05:30
parththakkar106andClaude Opus 5 3b9e6b3d50 Give the length hint a floor, not just a wall
A ceiling alone is a one-sided instruction, and models read it in opposite
directions. A verbose one is held back by it; a terse one has nothing to act
on except "write only as much as the moment needs -- a typical turn is much
shorter" and collapses to two paragraphs. Same prompt, wildly different turn
lengths depending on which model is behind it.

State a floor as well, so the guidance is a band. The two bounds are
deliberately asymmetric -- "must not exceed" for the wall the endpoint
enforces, "should not stop short of" for the floor -- so neither reads as a
number to hit, which is the property the earlier A/B says decides whether
this hint helps or hurts. "Prefer the lower end" inherits the anti-overshoot
job the deleted "much shorter" line was doing, but now with a number under
it, so a terse model lands on the floor instead of at forty words.

Below MIN_LENGTH_FLOOR_WORDS the floor is dropped and the tight-cap wording
is left byte-identical: at a tight cap a short turn is the correct turn, and
that phrasing is the one measured to keep the state block alive (0/6
truncations at cap 250 against 2/6 unhinted). So this only moves loose caps.
MAX_LENGTH_FLOOR_WORDS keeps the share from demanding 555 words minimum at
cap 2400 -- a big cap means long turns are allowed, not compulsory.

Shipped without an A/B run, deliberately. Two things to watch live: whether a
stated range invites landing mid-range on verbose models (drop the share to
~0.25 if so), and whether the state block still survives -- nothing reads
finish_reason yet, so truncation is silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-21 00:14:56 +05:30
parththakkar106andClaude Opus 5 ff190d2453 Edit the take you are reading, not the one it replaced
The pager parks a turn on take 2 of 4 and the row draws that take's words,
but the row itself is keyed by the live node — the take the story tells. The
edit button seeded from that node and saved back to it, so opening the editor
on a 2/4 turn showed 4/4's text and saving overwrote 4/4. The take being read
was never reachable.

It has an id of its own, carried on the preview, and the edit endpoint takes
any row by id whether or not the path runs through it. So an edit opened over
a preview carries the take's id, the transcript matches the editor to the row
through the preview rather than the node, and the saved text goes back into
the preview because there is no row in `actions` to put it in. The pager
caches the take list it fetched, so it is told to drop it — otherwise
stepping away and back showed the words from before the edit. Leaving the
take at all drops a half-typed edit with it.

The fork button had the same seed and now takes its text from the screen too.
Its id stays the live node's: a take branches just above the turn, and the
server only accepts a turn that is on the path.

The tests are on the promise the fix leans on — a take is an ordinary row to
the edit endpoint, addressed by its own id, and the group listing says so
afterwards. That already held; nothing in the backend changed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 23:37:36 +05:30
parththakkar106andClaude Opus 5 c0cd6fa7ce Stop a full memory bank from shutting itself
Eviction ranked on use_count first. A memory written this turn has never
been used, so once every survivor in a full bank had been retrieved even
once, the newborn was the lowest row in the bank and was marked forgotten
by the same post-turn run that wrote it — one pass after the one that
embedded it, before retrieval ever saw it.

That state never ends, because counts only go up. The bank an adventure
happened to hold when it first filled is the bank it keeps for good, and
everything the story does afterwards is summarized, evicted and never
ranked. A test now plays four turns against a full bank and asserts the
survivors are the new memories, not the opening ones; before this it kept
the opening three and forgot all four.

Order on last-touch instead, with the count as the tiebreak. A new memory
carries the newest timestamp there is, so it is the safest row in the bank
rather than the most doomed, and it has until something outlives it to
prove itself — the "protect the newest" behaviour falls out of the
ordering rather than being a count of rows to spare. Little is given up:
a memory that is genuinely used stays recently-used by being retrieved,
so the two orders only disagree about memories that mattered once and
have not been wanted since.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 23:22:00 +05:30
parththakkar106andClaude Opus 5 c4936f75f5 Fix two things only a person clicking could find
Driving SP9 by hand found both, and neither was reachable from a test that
did not exist yet.

**The pager arrived one page load late.** A retry's reply *is* the second take
of its turn, so it comes back needing a pager -- but the SSE stream builds its
own ActionOut and so never got the annotation. The numbers only appeared once
the page was reloaded, which is the one moment nobody reloads. That is the
third place this codebase builds an action payload a different way; the other
two were patched when `take_count` was added, and this one was missed because
nothing read it from the wire.

**"> You > You follow the pulse down the corridor."** The editor is seeded from
the stored text, which is already the formatted form -- the same text plain
edit puts in the box and writes back verbatim. Running it through the formatter
again doubles the prefix, so `run_player_turn` learns `preformatted`.

And the question the driving raised: does the shared state follow the branch?
`test_take_state.py` answers it with a script that adds ten gold a turn, which
is the clearest possible witness -- a take that stacks instead of replacing
says so in one digit. Three of its five fail with `roll_back_before` removed;
the other two pass either way on purpose, one being the premise and one going
through `stand_on`, which has restored state since SP5.

Confirmed against the running app too, on the HP-script demo: a path crossing
three branches carried exactly the damage of the nine AI nodes on it, and none
of the 66 points sitting on takes those branches never told.

433 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 21:34:12 +05:30
parththakkar106andClaude Opus 5 ea7336e5d3 Let any turn be played again, not just the newest
`retry` only ever saw the last action: mid-story there was no way to ask for
another take at all. Now the take endpoint takes either kind of node, and the
kind decides what happens -- an AI turn regenerates, a player turn takes the
text supplied. The client asks the same way for both.

The tip is the only case that needs no branch, and only for an AI turn, where
the takes are still leaves nobody has built on. That is the existing retry,
reached by another road. Everywhere else `branch_at` leaves the path just
before the turn so the story after it keeps the take it was written for.

Both paths roll the shared state back to before the turn ran, which the fork
endpoint was already doing and the player-take path was not -- a script would
otherwise have stacked this take's output mutations on the one being replaced.

426 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 21:34:12 +05:30
parththakkar106andClaude Opus 5 c3c9b310eb Put the pager's numbers on the page
`2/4` needs the shape of a turn's take group for every message on screen.
Asking `attempts.group` per row would put a query behind each one -- the exact
cost `variant_count` was cached to avoid, and the reason SP8 could not just
drop that column and be done. So one query per page, keyed on the parents the
page mentions, and the count and ordinal land on the rows before they are
serialised.

`take_count` / `take_index` rather than reusing the old pair, because they do
not mean the same thing: the old ones cache a coordinate's siblings and say 0
for a turn nobody retook, these count the *turn's* takes across whatever
branches they ended up on and say 1/1. SP8 still drops `variant_count` and
`variant_index`.

Two traps, one of them mine. `parent_id` was not in ACTION_LIST_COLUMNS, so
reading it off a windowed row would have been a lazy load per action -- an
N+1 hidden behind the thing `load_only` exists to prevent. And the adventure
GET does not build `ActionOut` at all: it hands the window to the relationship
with `set_committed_value` and lets Pydantic walk it, so patching the three
places that do build ActionOut missed the one path every page load takes.
The test caught it by reading the wire instead of the helper.

424 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 21:34:12 +05:30
parththakkar106andClaude Opus 5 2726840c60 Let a player's own turn be played again
Retry has always given an AI turn another take. Nothing gave one to the
message the player wrote, so the only way to change something you had typed
was to overwrite it -- destructively, losing whatever the story made of it.
That is the half of the tree the player asked for first, and it never existed:
`POST /actions/{id}/fork` answers 400 unless a second take is already there,
so "branch from here" was not a thing the API could do.

`tree.branch_at` is the missing piece. `fork` moves a node that already exists
onto a line of its own; this is the same branch with nothing in it yet, for a
take that has not been written. The head lands one depth short, so the next
node written is the new take -- same depth, same parent, resolved from the
path by `place_action` without being told. The line being left keeps its node,
keeps it live, and keeps everything played after it. Nothing is copied.

Two things the tests said rather than confirmed. A player turn is never the
tip: the reply to it is, so retaking even the newest thing typed still owes a
branch, and the guard against forking for nothing only fires for a player
action with no reply under it. And player text is stored formatted, so a test
that compares it whole is testing the formatter.

421 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 21:34:12 +05:30
parththakkar106andClaude Opus 5 e126bda387 Make the branch happen when you write, not when you look
Stepping between takes used to be two controls and two meanings. At the tip a
chip switched; above the tip it only *previewed*, and taking that line needed a
second button next to it. Which one you got depended on where you were standing,
which is the thing that made the tree unusable when it was driven by hand.

So the server stops caring that anyone is looking. Reading a take the story
moved past changes nothing and creates nothing. `ActionCreate.after_id` names
the node a turn is played after, and naming a take the story left is the first
moment the player has said which line they mean -- so that is where the fork
happens, and only there.

`stand_on` is the old fork endpoint's body, lifted out whole. It already knew
the two cases and got them right: at the tip the takes are leaves nobody built
on, so it is a switch and no branch is made; past the tip the line being left
keeps every turn it has, so the take needs a branch. Both callers now share it,
which is the point -- a fork asked for and a fork arrived at are the same move.

Six tests. Five fail with the grouping reverted to the coordinate, and the one
that does not is deliberate: naming the tip in `after_id` must stay an ordinary
turn that forks nothing, which guards against over-correcting rather than
against the original bug. The nesting test passes both ways too and is kept for
what it says, not for what it catches -- a coordinate separates C1's takes from
C2's by accident, because the fork has already put them on different branches.

415 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 21:34:12 +05:30
parththakkar106andClaude Opus 5 4af6e17406 Answer the review, and keep the opening node's bank
Nine findings from a review of the phase-14 stack. The one about a retry
withdrawing a memory is not a bug — a memory anchored to a node describes that
node, and it goes when the node goes. The root is the exception, and it is the
only one: migration 62 parked every memory written before memories had
coordinates on depth 0, so withdrawing the opening node would retire a whole
bank nobody attached there. A memory with no source range covers no story and
now stays; a summary that genuinely ends there is still withdrawn.

The rest are repairs.

* The adventure list quoted whichever attempt was written last rather than the
  one the story tells, so switching back left the index disagreeing with the
  page.
* A v1 import gave a typed memory no depth, rebuilding the NULL the migration
  exists to remove — invisible until the imported adventure forked.
* The action cap counted a v1 file's turns, and a turn expands into a row per
  saved attempt, so a file inside the cap could write a multiple of it.
* Forking a live node on a borrowed ancestor promoted a sibling on a branch the
  caller never named. It is a branch switch, and now says so.
* Switching attempts left the state, status and memory panels reading the
  previous take: the story does not change length, so nothing keyed on its
  length noticed. Same class as the branch-switch bug this phase already fixed.
* A retry after switching back numbered the new attempt into the middle of the
  group instead of the end.
* The cursor backfill numbered every action in the table once per adventure;
  correlated to the adventure being updated, it is an index lookup instead.
* Renaming a branch answered own_actions=0.

And one behaviour change recorded rather than repaired: script-visible history
and actionCount no longer count blank-text rows. That is the right shape and
there is no reading compatible with both, so plan/14 says so.

409 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 2d38a162d4 Give every memory a node, and show only the ones on your path
A hand-written memory used to carry a NULL depth, described in the model as
"belongs to the adventure rather than to a path". That sounds harmless and
is not: a NULL is a coordinate no fork can cap, so a note typed on one line
followed the reader onto branches whose story it never described. It takes
the head now — the story you were reading when you wrote it — and obeys
exactly the rule a summarised memory obeys.

The unanchored escape clause in lineage.Path.clause existed for that single
case and is deleted rather than left unused. Its docstring argued that a
capped depth would drop a typed memory the moment its branch stopped being
the newest entry; anchoring answers the same worry better, because the
memory is not exempt from the path, it is on one.

The drawer now shows the path being read and nothing else, filtered by the
clause retrieval itself uses, so the bank you can see is the bank the model
can see. Nothing is stranded: a memory lives on a branch, switching to that
branch shows it, and deleting the branch deletes it. Pinning decides order,
the path decides existence.

Migration 62 lands existing NULL-depth memories at depth 0 of their branch
rather than at the tip. 0 is at or before every fork point, so every memory
stays visible from exactly the paths it is visible from today — nobody's
bank loses a row on deploy. The tip is the tidier-sounding choice and would
have emptied them out of every branch forked earlier than they were typed.

This supersedes the on_path flag and the "another branch" badge from
earlier today; anchoring makes them redundant, and they are removed.

Four tests changed because they asserted the old contract, not because
they broke. The one worth reading is the pair replacing
test_a_hand_written_memory_is_not_lost_at_the_first_fork: typed on shared
trunk it still survives a fork, and typed on ground the fork never
travelled it no longer follows you.

402 tests. Verified on tools/branch_fixture.py: each branch's drawer holds
its own memory and not the other's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 1d1367ce3a Say which memories the model can actually see
Retrieval has been path-scoped since SP3: a memory on a branch this story
never travelled is never sent to the AI. The drawer listed every memory
alike, so on a fork you read "Fell down the cellar stairs; badly hurt" and
reasonably concluded the model knew it. It does not. That is worse than
either hiding the row or retrieving it — it is the screen claiming
something the engine contradicts.

Hiding them is not the answer either; that was the original comment's
point, and it stands. A memory nobody can list is a memory nobody can
delete, in a phase whose rule is that nothing is removed automatically.

So the whole bank still lists, and the rows off the current path are set
back, dashed, and labelled "another branch". MemoryOut.on_path carries it,
computed from the predicate retrieval itself uses rather than a second
spelling of the same idea — two spellings drift, and the failure mode here
is a badge that says the opposite of what the model gets. Pinning does not
override it: the path clause runs before pinning is considered.

The relationship is asymmetric and there is now a test that says so. A fork
borrows its ancestors, so a memory written on the parent is on the fork's
path too; the reverse never is. Worth pinning before somebody makes it
symmetric on the grounds that it looks wrong.

399 tests, three new, egress ceilings intact — the flag costs one id-only
query. Verified in a browser on tools/branch_fixture.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 cf3d52171e Let a branch be named, and thrown away
SP7 needs three branch operations and SP5 built one. Switching exists;
naming and deleting had no column and no route between them.

A name is stored because a player chose it. An unnamed branch keeps NULL
rather than a generated "branch 4" — a generated label is derived, and it
would go stale the moment a branch before it is deleted and the ordinals
shift underneath. The client draws those from the fork depth, which
nothing can shift. The v2 bundle carries the name for the same reason it
carries the fork points and leaves `lineage` out: it is a decision, not
something computed from one.

Delete is what stands between a tree and unbounded growth, since nothing
prunes one on its own. It refuses two branches: the root, which holds the
turns every other branch borrows, and the one being read — including any
branch the head was forked from, which is the same mistake in disguise and
the one that would cascade the head away and leave head_branch_id pointing
at nothing. The nodes, memories and descendants go through the foreign
keys that already cascade.

A cursor standing on a deleted branch is cleared. On Postgres a stale
branch id would simply never resolve; SQLite hands the freed id to the
next fork, and then the anchor resolves onto a branch it has never seen
and calls a stretch of story already summarized.

396 tests, 15 new.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 a7bf47a35e Let a backup carry a story that went two ways
A bundle had one list and a forked adventure has two stories, so export was
emitting every branch's turns interleaved by index — a mangled story rather
than lost data, and unreachable only because forking has no UI yet.
`ai-dnd-adventure-v2` carries the branches, the depth each one left its
parent at, which attempt at every turn is the story, and what each node
left behind. That last one is not decoration: the after-snapshots are what
a branch switch puts back, and a bundle without them imports a tree nobody
can switch inside.

`app/bundle.py` owns both formats and nothing else knows either. The v1
reader stays — those files are already on people's disks — and it is now
the only place a `variants` array exists anywhere.

The rule the module is built on is that a bundle carries what was chosen
and never what is derived. The head branch, the fork points, the live flags
and the anchors are decisions somebody made. The lineage, the head depth,
the legacy `index` and the variant ordinals are computed from those and are
rebuilt on the way in, because a bundle is a text file anybody can edit and
a derived field shipped beside its source is a chance for the file to
disagree with itself where no read would report it.

`index` is the one that stops being academic here. It agreed with `depth`
until SP5, and this is the first writer that has to fill it for a forked
story, where two branches both hold a node at depth 4. It is allocated one
per turn instead: siblings share it, no two coordinates do.

Everything a hand-edited file can get wrong about the shape of a tree is a
400 raised before the adventure row exists, because a half-applied import
is exactly the failure this phase exists to end — a story that goes quiet.
A file wrong about which attempt is live is corrected rather than refused;
that is an invariant of the database, not of the format.

Measured on the 600-action fixture: 587 kB to 911 kB, and all of the
increase is the outcomes at 489 B a node — the coordinates themselves save
57.5 B a node against the old turn-and-variants shape. Twenty forks add
660 B. 4.3% of the import body cap.

381 tests green, 16 of them new in test_bundle_v2.py. No migration, no
vacuum owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 ffb2fd5b0e Let a story go two ways at once
Attempts pile up at the tip as leaves and cost nothing. The moment the
player takes the story down one the line has already moved past, the two
futures have to coexist — so `tree.fork` gives that attempt a branch of
its own, forked at the depth just before it, and the line it leaves
keeps every turn it has.

One row is inserted and one row is moved. Nothing is copied: everything
before the fork is borrowed through the lineage cached on the branch
row. Measured on a 40-turn story forked twenty times — 21 branches, 140
rows, an 80-action story — the page load costs 31,652 B against the
31,433 B the same story flat costs, and a branch is 103 B of ancestry.

Nothing derived moves either, and that is the part worth keeping: a
memory hangs off the coordinate the parent's attempt still occupies, and
the lineage caps the parent one depth short of it. The fork simply
cannot see it, so it resummarizes that ground from the text it actually
tells, without a line of bookkeeping.

Two things had to change underneath. A new node's depth now comes from
the tip of its branch rather than from the adventure-wide `index`, which
would have left a hole in a fork's path the width of the other branch.
And undo stops at the fork — the turns before it belong to the branch
this one grew out of.

`GET /branches`, `POST /branches/{id}/switch` and
`POST /actions/{id}/fork` are the endpoints SP7's tree view is drawn on.

365 tests green, 18 of them new in test_branch_forking.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 0a12d9cd47 Make a retry a node, not a rewrite
Every attempt at a turn is now its own row at the same (branch, depth),
with `live` naming the one the story tells. The JSON repeating group on
`actions.variants` is read one last time, by a migration that writes it
out as the sibling rows it always described, and then goes unread.

The snapshots turn around with it: an action carries the state it left
behind rather than the state it started from, because attempts at one
turn share a starting position and differ exactly in their outcome.
Rolling back is "what the node in front left behind", one lookup on the
path, and it is what undo and retry now both read.

And the memory holdback goes. It existed because retry rewrote a row
under a mark that had already moved past it; a retry writes a sibling
now, and replacing what a coordinate says withdraws what was derived
from it — the same repair undo and delete already made.

The assembled prompt is still stored once per turn: it moves with the
live flag, so a superseded attempt keeps only the few hundred bytes that
were its own. Measured on the 600-action fixture: 700 rows for the same
600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB
and page load 62.7 kB unmoved.

347 tests green. `tests/test_story_tree_baseline.py` and
`tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move
the baseline for the variant-count semantics and did not need to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 c51531709d Mark the story with a node, not with a count
The memory bank and the story summary each kept a cursor: how many story
actions they had already covered. A count is a position in a list, and this
list moves — delete an action in front of the mark and every later one slides
down a slot, so the mark now covers one it has never read. All the cursor
bookkeeping existed to patch that up.

Both marks are now (branch_id, depth): the node up to and including which the
work is done. A depth is a coordinate along a path, not an offset into a list,
so nothing in front of it can move it. That deletes rather than rewrites
`position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
`prune_dangling_memories` and the every-pass clamp in `run_post_turn`.

A memory hangs off the node its block ends on, so a fork inherits its
ancestors' memories without copying any, and retrieval selects through the
branch clause over the *whole* lineage — recall is long-range by definition and
cannot be windowed. Measured: 1,807 B on a story forked twenty times against
1,823 B on a flat one of the same length.

Migrations 53-56 translate the old counts into nodes. They rewrite `adventures`
and not `actions`, so this one needs no VACUUM FULL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 05a2a77e4c Resolve the head once a flush, not once a node
The identity map holds weak references. A branch row nobody keeps a strong
reference to is collected between two nodes, so resolving the head inside the
placement loop read it back from the database for every node in the flush: 201
SELECTs on `branches` to write 200 actions, and 36 s -> 45 s on the same 297
tests. Nothing about any result changed, which is why only a stopwatch found
it, and why there is now a test counting the reads.

Also: the two reads left un-pathed on purpose say so where they live —
`max_action_index` allocates the legacy `index` and must stay adventure-wide
or two branches issue the same number, and export is a flat v1 bundle whose
reader has no idea branches exist. And `Adventure.actions` keeps its `index`
ordering, because ordering the collection by depth would not make it a story:
it is every branch's actions, and a path is a selection out of it.

318 tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 563b9af9cf Read one story, and know which one
Every read of an action now goes through a single module. `context/lineage.py`
turns a branch's stored lineage into the OR-of-ranges that is "this story", and
history, paging, the newest-action lookups, the index screen and the scripting
history API all select through it. A forgotten clause does not raise — it
quietly assembles a page, or a prompt, out of two different stories — so the
clause lives in one place rather than in a convention.

The read that mattered most was the shortcut: `_from_memory` sliced
`adventure.actions`, which is every branch's actions, not the path. It now cuts
the loaded collection down with the same predicate the SQL uses. Same trap one
layer up, and user-visible: `pipeline._history()` hands user scripts the story,
and was handing them the collection.

Tail reads window the lineage as well as the rows: the newest few entries cover
the context budget, so a story forked twenty times reads its tail with one
clause and costs 1.07x what an unforked story of the same length costs. The
estimate is depth arithmetic, and where a deleted action leaves a gap the read
notices it came up short and widens to the whole ancestry.

Ordering moves from `index` to `depth`, with `id` breaking ties. The two hold
the same numbers until retry stops mutating rows in SP4, but only one of them
is a position along a path.

One thing SP1 did not anticipate: wiring the writers was not enough. From here
a row without a branch is a row no read can see, and "every writer remembers"
has to hold for every fixture, script and test ever written — including the
SP0 baseline, which writes its actions straight to the database and must pass
unmodified. So the session enforces it: `tree.place_new_nodes` runs from
before_flush and places anything unplaced. The call sites keep their explicit
calls, because a node placed at the call site is placed before the code around
it reads it back.

316 tests green: the 297 from SP1, plus 19 in test_branch_clause.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 d3756abdaa Give every action a branch and a depth
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a
`branches` table, `branch_id`/`depth` on actions and memories, a head pointer
on adventures, migrations 46-52, and a server-side backfill that re-reads every
existing adventure as a tree with one branch. `depth` holds the number `index`
already held, gaps included, so no story changes — a linear story *is* a tree
with one branch, which is what makes the SP0 baseline passing unmodified the
pass condition rather than a hope.

The writer had to come with it. No migration will ever visit a row written
after it ran, so columns backfilled today and populated next subphase would
leave a hole exactly the width of one deploy, and from SP2 on a row without a
branch is a row no read can see. `app/tree.py` owns that: one module, because a
node written without a branch fails by disappearing rather than by raising.

Three things the schema itself insisted on:

- `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing
  both ways makes the two tables a cycle create_all cannot order, and its
  escape hatch needs an ALTER SQLite does not have. It is a cache, and a head
  naming a branch that is gone recovers onto the root.
- `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a
  second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`,
  because Postgres `json` has no equality operator.
- SQLite will not drop a column a foreign key names, which is how two existing
  tests broke: they simulated an old database by rewinding the stamp while
  leaving the new columns in place. Every ADD COLUMN migration is now
  idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by
  rebuilding three tables from frozen DDL so the real ALTERs run.

297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn;
page load and index are byte-identical to the recorded figures.

The deploy that ships this needs one `VACUUM FULL actions;` on the direct
endpoint afterwards — it rewrites every row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 1940e222a7 Pin down what a tree must not change
Phase 14 replaces the action list with a tree, which touches the memory bank,
the context builder, undo, retry and the UI at once. Before any of that moves,
write down what "unaffected" means and prove it holds today.

tests/test_story_tree_baseline.py drives the product over HTTP the way a player
does and asserts only on API responses, never on how anything is stored: the
turn engine, retry and variant switching, undo, paging all the way back to the
start without gaps or repeats, edit/delete, memories, export/import round-trip,
world state. 283 green with it added, from 259.

It must pass unmodified through the schema change, the branch clause and the
move of memories onto nodes. A linear story is a tree with one branch, so those
subphases have no licence to change behaviour, and this is what says so.

The plan itself records four decisions that were open: structural-first with no
runtime flag, legacy columns kept one release, full tree visualisation, and a
v2 bundle with a v1 reader. Plus four traps found by reading rather than
running -- history._from_memory and the scripting pipeline both slice
adventure.actions, which under a tree is every branch rather than the path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 cf8ec22e8b Open an adventure on a window of the story, and page upward
Opening a finished adventure fetched every action in one response: 589.5 kB on
production's longest, and nothing about that curve bends on its own, because a
story only ever gets longer. The page load now brings the newest 60 actions
and the reader pages up from there. On the harness's 600-action fixture that
is 606.0 kB down to 62.6 kB, and -- the part that matters -- it no longer
depends on how long the story is.

Paged by anchor, not by offset. `before_id` is the oldest action the caller
holds; the server returns what precedes it. An offset counted back from the
newest would shift every older position the moment a turn lands, which is
exactly when someone is likely to be scrolling, and the reader would get one
action twice and never see another. It also keeps working when the story stops
being a flat list: comparing indices to order a branch survives the story tree,
treating them as positions does not.

`has_more` comes from fetching one row past the window rather than from
counting. A deleted anchor -- undo, mid-scroll -- reports the end rather than
guessing and serving a page the reader already has.

GET /{id}/actions and POST /{id}/undo now return {actions, total, has_more}
instead of a bare list. Undo is the action most likely to be repeated several
times running, so having it re-fetch the whole story would have undone the
paging on the worst case.

The adventure payload gets its window through set_committed_value rather than
by assignment: the actions relationship cascades delete-orphan, so assigning a
60-item list to it would delete everything outside the window on the next
flush.

In Play.jsx the prepend is followed by a useLayoutEffect that restores the
scroll position, before paint, so the story does not jump. Loading starts 400px
from the top rather than at it, guarded by a ref because scroll fires far
faster than React re-renders. There is a button as well as the scroll trigger:
on a short viewport the transcript may not be tall enough to scroll at all, and
a reader who cannot scroll must still be able to reach the beginning.

Verified against a running backend and a 220-action adventure: the page load
returns 60 of 220 ending on the newest, and walking back from an anchor returns
exactly the actions before it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 14:23:03 +05:30
parththakkar106andClaude Opus 5 ae6e5af6c7 Store context_snapshot compressed
One column is 89% of the database and the free tier allows 512 MB. Reads were
already solved -- the column is deferred, so a page load never touches it and
one screen fetches one row at a time -- but nothing had costed storage, and
storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per
action, so the ceiling arrives around 5,400 actions and 944 are stored.

Postgres already compresses it and only gets 1.7x. pglz is tuned for fast
decompression of data a query might filter on, and nothing has ever filtered
on an assembled prompt -- it is written once and read whole, rarely, by the
Insights viewer. zlib gets 3.5x on the same text for a decompress on a request
that already made an LLM call.

Done as a TypeDecorator rather than a second column, so every call site still
writes a dict and reads a dict back, and deferred/undefer/load_only keep
naming the same attribute. Only the storage format moves.

Migrations 43-45: add the bytea, convert into it, drop the original, rename.
The backfill is the one destructive step in the file -- 44 removes the only
other copy -- so it decompresses every row and compares it against what went
in, and a row that fails aborts the run. The whole loop is one transaction, so
an abort rolls the DROP back and the prompts are still there.

Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway
Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column
came out named context_snapshot, every snapshot compared equal and the one
NULL stayed NULL.

Postgres does not return the disk by itself: DROP COLUMN only marks the column
gone and the backfill leaves a dead tuple per row, so the table peaks near
twice its size before settling. The deploy needs one VACUUM FULL to collect
it; the migration comment says so.

The egress fixture's snapshots are prose now rather than "x" * 20_000, and the
prose generator moved to tools/fakeprose.py so the harness and the tests share
one definition. A repeated character compresses a thousandfold: against the
old fixture a compressed column looked free and the byte ceilings would have
been guarding nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 14:08:23 +05:30
parththakkar106andClaude Opus 5 a6cb49293c Name the columns a list response carries
deferred=True keeps the four heavy Action columns out of a bulk read, but it
makes narrowness the thing a future column has to remember to ask for -- and
both egress blowouts this project has had were a column nobody remembered.
Listing what each list response renders inverts the default: a new column
costs nothing on these paths until someone adds it to the tuple.

The adventures index was not merely a future risk. It loaded whole Adventure
entities to render a title, a stamp and a snippet, and an Adventure carries
script_state, world_state, placeholders, story_summary, memory, authors_note
and ai_instructions -- ~15 kB a row in production, none of it on that screen,
all of it fetched once per adventure on every index load. Measured on six
adventures with 78 kB of body each: 469.7 kB entity-loaded against 318 B
projected.

The memories drawer stops walking adventure.memories. The walk is what
retrieval used to do and the reason a turn cost megabytes; a relationship load
takes whole entities, so it picks up whatever the model happens to grow.
Nothing changes today -- embedding_blob is already deferred -- which is the
point.

world_delta stays on the action list because ActionOut.world_changes is
computed from it. Leaving it off would not save the bytes, it would spend them
one row at a time as a lazy load.

Two tests cover the index: one asserts the listing query names none of the
body columns, one puts a byte ceiling on six adventures carrying 80 kB apiece.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 13:56:07 +05:30
parththakkar106andClaude Opus 5 12d57afdac Put a number on what an endpoint may fetch, not just a column list
test_egress.py asserted which columns a statement names, which is the shape
both of this project's egress blowouts took. It would all still pass if a
response grew tenfold within the columns it is allowed to read -- and a story
that keeps getting longer does exactly that. Production's longest adventure is
607 actions where the plan assumed 200.

So dbmeter, which was built to be importable from tests and was not yet used
by any, now backs four byte ceilings: the page load, the action list, and one
action's snapshot fetched on demand. Budgets are per action rather than
absolute, so they mean the same thing whatever size the fixture is set to, and
generous -- 3 kB against a real 994 B. They are there to catch an order of
magnitude, not to freeze a byte count.

The fourth test is the one that keeps the other three honest. A ceiling proves
nothing unless the thing it excludes would breach it, so it undefers the
snapshot on purpose and asserts the same twelve rows cost more than ten times
the budget. If the fixture ever shrinks below the point where that holds, that
test fails rather than the ceilings quietly passing on nothing.

Meter grows detach() and a context manager. A script exits and takes the
wrapping with it; a test does not, and one test leaving the shared engine
metered would charge bytes to a scope nobody opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 13:51:10 +05:30
parththakkar106andClaude Opus 5 2c5909a268 Drop the JSON vector column, and fix what was hiding behind it
Migration 38 left memories.embedding in place so a rollback could still find
the vectors. Production has since been verified reading from embedding_blob,
so migration 42 drops it: 4 MB of a 99.6 MB database holding nothing anyone
reads.

Removing it surfaced a live bug. Changing your embedding model is supposed to
throw the bank's vectors away and let the post-turn pass rebuild them, because
two models' vectors are not comparable. The settings route did that by nulling
memories.embedding -- correct until 38 moved the vectors, after which it
cleared the dead column and left the blob intact with `embedded` still true.
_embed_pending filters on `embedded IS FALSE`, so it never saw those rows and
the bank went on ranking against the old model's vectors permanently.

Nothing would have reported it. cosine returns 0.0 on a width mismatch, so a
different-width model scores every memory zero and retrieval returns whichever
rows happen to sort first; a same-width model scores plausible garbage.

The bulk clear now sets both columns. It stays a bulk UPDATE rather than going
through set_vector -- loading the rows is the cost that whole path exists to
avoid -- so set_vector's docstring now names it as the one caller that
legitimately writes those columns by hand. No cache invalidation is added:
clearing `embedded` drops the rows out of the catalogue query, and set_vector
evicts each entry as the re-embed puts it back.

test_embedding_blob.py now rebuilds the pre-38 schema by hand where it tests
the backfill, since create_all no longer produces the column it converts from,
and asserts 42 removes it at the end of a full bootstrap -- 38 reads that
column and 42 drops it, so an upgrade that reordered them would arrive with an
empty bank.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 13:44:59 +05:30
parththakkar106andClaude Opus 5 b7e53ae581 Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.

Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.

    one turn      3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
    run_post_turn 3,139.1 kB -> 0.7 kB
    Insights      3,223.7 kB -> 117.9 kB
    Memories drawer  ~3.1 MB -> 23.6 kB

A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.

The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.

memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.

Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:28:00 +05:30
parththakkar106andClaude Opus 5 c56864877a Store embeddings as packed float32 instead of a JSON list
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.

It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.

Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.

The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.

Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:17:44 +05:30
parththakkar106andClaude Opus 5 bbcb07c6be Delete guest accounts left idle for five days
The app had no cleanup of any kind: in multi-user mode every first visit
mints a users row, so the demo has been accumulating one permanent account
per visitor along with everything they generated.

cleanup.py sweeps guests idle for AIDND_GUEST_RETENTION_DAYS (default 5),
once at startup and then every few hours. Startup is the load-bearing
trigger — the free tier sleeps after ~15 minutes, so a long timer rarely
gets to fire.

Idle is COALESCE(last_seen_at, created_at), not last_seen_at: _touch only
writes that column hourly, and a guest minted by /auth/me has it NULL until
its second request, so the simpler query would have deleted brand-new
visitors mid-session.

It's one Core DELETE rather than db.delete(user), which would SELECT every
adventure, action and memory into Python purely to delete them — the same
egress pattern as the 189x fix. Every FK from users down is ON DELETE
CASCADE, so the database does the whole graph and returns a count.

The filter requires is_guest AND email IS NULL, so registered users (who
upgrade in place) and local mode's implicit user are both out of reach, and
is_public is output-only so a guest can never own content another user can
see. Session cookies have no expiry and can outlive a swept row; that path
401s and the frontend's existing retry re-mints a session.

Guests are told: /auth/me serves guest_retention_days and the signup modal
states the window, sourced from the server so it can't drift from what is
enforced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-15 20:14:12 +05:30
parththakkar106andClaude Opus 4.8 c500203270 Harden auth against a forwarded-header rate-limit bypass, and guard BYOK SSRF
The per-IP rate limits could be bypassed entirely: uvicorn ran with
--forwarded-allow-ips "*", which trusts the leftmost X-Forwarded-For value
(client-controlled), and Render forwards the inbound header rather than
stripping it. Rotating the header handed out a fresh rate-limit bucket per
request, so the login/register limit (10/5min) and guest-minting limit
(30/5min) were no throttle at all — unbounded password guessing and guest-row
creation. Confirmed live: fixed IP -> 429 after 10; rotating spoofed header ->
no 429 across 14 attempts.

Two-layer fix:
- limits._client_ip now derives the client IP from the hop the trusted edge
  appends (rightmost of X-Forwarded-For), which a client can't spoof past;
  tunable via AIDND_TRUSTED_PROXY_HOPS. Dropped --forwarded-allow-ips "*".
- New per-account login throttle (email-keyed, 8 fails / 15 min, cleared on
  success): stops distributed guessing against one account that a per-IP limit
  can't, since it can't be diluted across many source addresses.

Also close an SSRF on the BYOK endpoint_url (hosted mode only): the connection
test and turn/chat streams now refuse a URL that resolves to a non-public
address (private/loopback/link-local metadata/reserved), checked at request
time so it resists a DNS record flipping to a private IP. No-op locally, where
reaching localhost Ollama is intended.

Tests: test_ratelimit_hardening.py (8), test_netguard.py (13). 172 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-15 13:47:14 +05:30
parththakkar106andClaude Opus 5 f295893204 Ask the model for a turn that fits inside the output cap
max_output_tokens is a hard wall the endpoint enforces mid-sentence. The
```state block is emitted after the narration, so a long turn hits the wall
partway through the block and the deltas are lost — silently, since nothing
reads finish_reason.

builder.length_hint() derives a word limit from the cap ((cap - 50 headroom)
* 0.75 words/token * 0.90 buffer) and injects it just above EMIT_REMINDER,
which keeps the last slot it needs. Reserved in build_context like the
reminder is.

Phrased as a ceiling, not a budget. Measured against gemma-4-26b at cap 800,
n=5 per arm: no hint 174 words, "keep this turn under about N words" 246,
"hard limit ... a typical turn is much shorter" 170. A budget reads as a
target to fill — every budget run was longer than every unhinted one, pushing
turns toward the wall the hint exists to avoid. Ceiling phrasing still works
at tight caps: at 250, unhinted hit finish_reason=length 2/6, hinted 0/6.

tests/test_length_hint.py, 11 tests; each mechanism verified by sabotage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
2026-08-08 12:20:26 +05:30
parththakkar106andClaude Opus 5 47e33fa311 Read a window of the story per turn instead of all of it
Two reads still grew without bound after the snapshot fix.

`Action.variants` holds every discarded retry attempt, but a list response
only needs how many there are — so each retry permanently added ~5 KB to
every later load of that adventure. Defer the column and keep the count
beside it (migration 37, backfilled server-side), with set_variants() as the
one write path that keeps the two in step.

`story_actions()` walked adventure.actions, then every caller threw almost
all of it away: the builder concatenates the story and immediately cuts it
back to the token budget, the NPC check looks at the last 6, retrieval at the
last 4, the cursor clamp only wants a count. A turn on a 200-action adventure
read 839 KB to use ~70 KB, and grew with every turn played. app/context/
history.py serves those shapes from SQL; window_covering() measures the
actions it fetched and projects how many more it needs, fetching only the
part it does not already hold. Memorybank cursors move to position_of_index()
and settled_count()/settled_slice() — same arithmetic, no full list.

The scripting pipeline still receives the whole history per AI Dungeon's API,
and every helper reuses adventure.actions when it is already loaded, so a
scripted adventure pays what it always did and never twice.

Measured at production shape: retry tax 5.1 KB -> 0; turn 200 839 KB -> 129 KB
and flat from ~turn 50; a 200-turn playthrough 84.5 MB -> 23.0 MB; a delete
115 KB -> 5 KB.

Verified the window builds a byte-identical prompt to the full story across
budgets from 1K to 100K tokens, with and without the retry exclusion - this
is a cost change and nothing else. Cursor helpers checked against the old list
arithmetic, including after deleting a middle action. Counts are real
SELECT count(...): Query.count() wraps the entity select in a subquery, so the
SQL named every deferred column and the egress guard could not tell it apart
from a bulk fetch. 139 tests pass; the four new guards verified by sabotage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
2026-08-06 15:25:27 +05:30
parththakkar106andClaude Opus 5 f1bd099ec8 Cut database egress 189x by deferring the prompt snapshot
The free-tier 5 GB/month network transfer allowance ran out, which blocks
connections outright. The database is only ~55 MB, so 5 GB meant the whole
thing was being pulled roughly 90 times over.

Cause: actions is 39 MB of that 55 MB -- 541 rows at ~74 KB each, almost
entirely context_snapshot, which stores the whole assembled prompt for a
turn. Every adventure load and every turn fetched all of it in order to
read two small things out of it: the world-change chips under an AI
message (Action.world_changes) and the emit block re-attached when
replaying history to the model (_history_text). The Insights viewer is
the only consumer that wants the whole snapshot, and it asks for one
action at a time.

Lifts that slice into its own small actions.world_delta column
(migration 36) and marks context_snapshot, state_before and
world_state_before deferred, so they load only when something touches
the attribute -- Insights, undo and retry, all single-action paths.
The backfill runs server-side, dialect-specific (json_extract on SQLite,
#> on Postgres), because pulling 39 MB of snapshots into Python to
rewrite a slice of each would defeat the purpose.

Measured at production shape (541 actions, 72 KB snapshots), one
adventure load goes from 38.46 MB to 0.20 MB. The traffic that consumed
5 GB would now be about 27 MB.

Deliberately not included: limiting the history query to recent actions,
and removing the redundant db.refresh(adventure) calls. Both were sized
against the old numbers; against a 0.20 MB load they would take ~27 MB a
month down to ~10 MB, which is not worth the complexity.

tests/test_egress.py hooks before_cursor_execute and asserts the emitted
SQL never names the deferred columns during a bulk load, so this cannot
regress silently. 123 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
2026-08-03 20:35:54 +05:30
parththakkar106andClaude Opus 5 7c538c8235 Stop retried and deleted actions from corrupting story context and memories
Three fallout bugs from keeping the retried action row alive (906ba42),
plus two long-standing cursor bugs the same investigation turned up.

Retry context leak: the row being regenerated is still attached to the
adventure, so it was replayed as established story and the model wrote a
continuation of the attempt it was meant to replace — the story visibly
blended both takes. It leaked into four places, not one: history replay,
story-card trigger matching, in-scene NPC detection, and the memory-bank
similarity query. Adds a shared context.story_actions(exclude_action_id),
threaded through build_context and retrieve_memories.

Memory holdback: a memory could summarize the just-generated turn; retry
rewrites Action.text but memory_cursor has already advanced, so the memory
was never regenerated and went on describing narration no longer in the
story. settled_story_actions() holds the newest action back one turn —
only the last action is retryable, so that makes it unreachable. The
settled list is always a prefix, so cursors stay valid and nothing is
skipped. The run_post_turn clamp deliberately still uses the full count:
clamping to settled rewinds legacy adventures a step and double-covers an
action.

Cursor bookkeeping: memory_cursor is a position into story_actions() while
Memory.source_* are Action.index values, and the two diverge as soon as
anything is deleted. Deleting a middle action slid a never-summarized
action into the covered range, skipping it forever; and pruning a memory
left the actions it covered stranded behind the cursor. Adds
note_action_removed() (called before the delete in delete_action and
undo_turn) and a rewind in prune_dangling_memories. delete_action also
now prunes at all, which it never did.

Not addressed: editing an already-summarized action still leaves its
memory stale, and the cumulative story summary can't have one fact
un-mixed from it.

117 backend tests pass, including new test_memory_settling.py (12) and
two retry-context regression tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
2026-08-03 15:23:59 +05:30
parththakkar106andClaude Opus 5 906ba423d8 Keep every retried attempt instead of deleting it
Retry used to delete the last AI action and generate a replacement, so the
discarded narration was simply gone. The row survives now: each attempt is
appended to actions.variants with variant_index naming the live one, and a
ChatGPT-style pager under the message browses them.

Action.text still mirrors the active variant, so the context builder, memory
bank, summarizer and export needed no changes. A variant carries only what
differs between attempts -- the text, the reasoning, and the state it
produced -- never the assembled prompt, which is identical across attempts of
one turn and is the bulk of context_snapshot.

Only the last message can be switched, restoring the script/world state that
attempt produced; earlier turns were written as a continuation of whatever is
active there, so theirs are read-only previews.

Three things that would otherwise bite:

- generate_turn now wraps _generate_turn and watches for a save sentinel. If
  the generator ends without it (provider error, empty reply, script stop,
  client hangup) it re-applies the previous variant -- otherwise a failed
  retry leaves rolled-back stats under un-rolled-back text.
- Retry reuses the turn's own index rather than next_index, or the clock the
  world-state cooldowns run on advances on a re-run of the same turn.
- Editing a message rewrites the active variant too, or paging away and back
  silently reverts the edit.

Migrations 34/35 verified as an upgrade against a populated database, not
just a fresh schema. test_state_revert's retry test asserted the old delete
behaviour and was rewritten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
2026-08-02 20:44:06 +05:30
parththakkar106andClaude Opus 5 d66fd1d6a2 Add a reasoning-off setting (-1 reasoning budget)
Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.

Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.

Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
2026-08-02 11:10:35 +05:30