Give every action a branch and a depth

Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a
`branches` table, `branch_id`/`depth` on actions and memories, a head pointer
on adventures, migrations 46-52, and a server-side backfill that re-reads every
existing adventure as a tree with one branch. `depth` holds the number `index`
already held, gaps included, so no story changes — a linear story *is* a tree
with one branch, which is what makes the SP0 baseline passing unmodified the
pass condition rather than a hope.

The writer had to come with it. No migration will ever visit a row written
after it ran, so columns backfilled today and populated next subphase would
leave a hole exactly the width of one deploy, and from SP2 on a row without a
branch is a row no read can see. `app/tree.py` owns that: one module, because a
node written without a branch fails by disappearing rather than by raising.

Three things the schema itself insisted on:

- `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing
  both ways makes the two tables a cycle create_all cannot order, and its
  escape hatch needs an ALTER SQLite does not have. It is a cache, and a head
  naming a branch that is gone recovers onto the root.
- `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a
  second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`,
  because Postgres `json` has no equality operator.
- SQLite will not drop a column a foreign key names, which is how two existing
  tests broke: they simulated an old database by rewinding the stamp while
  leaving the new columns in place. Every ADD COLUMN migration is now
  idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by
  rebuilding three tables from frozen DDL so the real ALTERs run.

297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn;
page load and index are byte-identical to the recorded figures.

The deploy that ships this needs one `VACUUM FULL actions;` on the direct
endpoint afterwards — it rewrites every row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
This commit is contained in:
parththakkar106
2026-08-18 19:14:07 +05:30
committed by Parth
co-authored by Claude Opus 5
parent 5c1bcf7305
commit d3756abdaa
13 changed files with 1026 additions and 38 deletions
+47
View File
@@ -0,0 +1,47 @@
"""Make a database look like an older schema version, so a migration can run.
`create_all` always builds the *current* schema. A test that wants to watch a
migration happen therefore has to take the newer columns back off before it
stamps an older version — otherwise the migration meets a table that already
has its column and dies on a duplicate.
Rewinding the stamp alone was enough for a while, which is why two test files
did exactly that. It stopped being enough the moment another `ADD COLUMN`
landed after theirs: the replay then runs migrations they never meant to
exercise, against columns `create_all` had already made. This module is that
rewind done properly, in one place, so appending a migration means adding its
inverse here rather than discovering three unrelated test failures.
SQLite only — every test that replays migrations runs on a temp file, and
`PRAGMA user_version` is where the stamp lives there. Migrations that change a
column's *type* (43–45, JSON to compressed bytes) have no clean inverse and are
not listed: they get replayed as-is, which is what the tests using them already
relied on.
"""
from sqlalchemy import text
from sqlalchemy.engine import Engine
# (version that added it, statements that take it back off), newest first.
#
# Phase 14's `branch_id` columns are deliberately absent: SQLite refuses to drop
# a column a foreign key names ("unknown column in foreign key definition"), so
# a current-schema database cannot be rewound past them at all. That is what
# `migrations._column_already_there` is for — the replay skips DDL that has
# already happened, so the tree migrations run their backfill against a schema
# that already has the columns, which is exactly the situation here.
_UNDO: list[tuple[int, tuple[str, ...]]] = [
# Packed float32 vectors and the flag beside them.
(39, ("ALTER TABLE memories DROP COLUMN embedded",)),
(38, ("ALTER TABLE memories DROP COLUMN embedding_blob",)),
]
def rewind_to(engine: Engine, version: int) -> None:
"""Drop everything added after `version`, then stamp the database at it."""
with engine.begin() as conn:
for added_at, statements in _UNDO:
if added_at > version:
for sql in statements:
conn.execute(text(sql))
conn.execute(text(f"PRAGMA user_version = {version}"))