Resolve the head once a flush, not once a node
The identity map holds weak references. A branch row nobody keeps a strong reference to is collected between two nodes, so resolving the head inside the placement loop read it back from the database for every node in the flush: 201 SELECTs on `branches` to write 200 actions, and 36 s -> 45 s on the same 297 tests. Nothing about any result changed, which is why only a stopwatch found it, and why there is now a test counting the reads. Also: the two reads left un-pathed on purpose say so where they live — `max_action_index` allocates the legacy `index` and must stay adventure-wide or two branches issue the same number, and export is a flat v1 bundle whose reader has no idea branches exist. And `Adventure.actions` keeps its `index` ordering, because ordering the collection by depth would not make it a story: it is every branch's actions, and a path is a selection out of it. 318 tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
This commit is contained in:
committed by
Parth
co-authored by
Claude Opus 5
parent
563b9af9cf
commit
05a2a77e4c
@@ -293,6 +293,52 @@ doc's own example reads back as `A0 A1 A2 A3 B4 B5 C6 C7`, and that a sibling's
|
||||
invisible. Egress: a 20-fork fixture costs within a small factor of a 1-fork one —
|
||||
clause count is bounded by the context window, not by fork count.
|
||||
|
||||
**Done, 2026-08-17** (branch `sp2-branch-clause`). **317 tests green**, the 297 from SP1
|
||||
plus 20 in `test_branch_clause.py`. The baseline contract passes **unmodified**, which
|
||||
was the pass condition. Six things worth not rediscovering:
|
||||
|
||||
- **The contract forced the write side, not the read side.** The SP0 baseline writes its
|
||||
actions straight to the database and must pass unmodified — so "every writer calls
|
||||
`place_action`" could not be the invariant, because the baseline is a writer and does
|
||||
not. Neither did any of the eleven other test fixtures. The alternative was a read
|
||||
tolerant of a NULL branch, which is the quiet-wrong-story failure this subphase exists
|
||||
to make impossible. So the session enforces it instead: `tree.place_new_nodes` runs
|
||||
from `Session.before_flush` and places anything unplaced, registered in `models.py` so
|
||||
that importing the models arms it. The call sites keep their explicit calls — a node
|
||||
placed at the call site is placed *before* the code around it reads the row back.
|
||||
- **A branch could no longer be created with a flush.** `root_branch` did
|
||||
`db.add(); db.flush()` to get the id its lineage names, and a nested flush inside
|
||||
`before_flush` raises. It inserts through Core and reads the row back — same
|
||||
transaction, three statements, once per adventure ever.
|
||||
- **The identity map holds weak references, and it cost 25 % of the suite.** Resolving
|
||||
the head branch per node re-read the `branches` row for every node in the flush, because
|
||||
nothing held a strong reference between two calls: 201 SELECTs to write 200 actions,
|
||||
and 36 s → 45 s on the same 297 tests. The head is now resolved once per adventure per
|
||||
flush (2 SELECTs), which put the suite back at 38.8 s *with* 20 more tests. Pinned by a
|
||||
test, because the symptom is only ever a stopwatch.
|
||||
- **The index screen is the one read scoped by head branch rather than by lineage.**
|
||||
`_latest_narration` picks one row per adventure for a hundred adventures at once, and a
|
||||
lineage clause each would put hundreds of OR-terms on that query. The two answers differ
|
||||
only for a branch with no nodes of its own, which cannot exist — a branch is created by
|
||||
playing a turn onto it.
|
||||
- **Two reads are deliberately left un-pathed**, both documented where they live.
|
||||
`max_action_index` allocates the legacy `index`, which must stay adventure-wide or two
|
||||
branches issue the same number; and export is a flat v1 bundle whose reader has no idea
|
||||
branches exist, which is why SP6 replaces the format rather than widening the query.
|
||||
A third is a known divergence, not a decision: the index screen's `action_count` counts
|
||||
the tree, and will overstate a branched story until SP5.
|
||||
- **`Adventure.actions` was left ordered by `index` on purpose.** Ordering the collection
|
||||
by depth would not make it a story — it is every branch's actions, and a path is a
|
||||
selection out of it. What the relationship is still for is ownership and the
|
||||
delete-orphan cascade.
|
||||
|
||||
Measured: a story forked **20 times reads its newest 32-action window for 3,178 B against
|
||||
the 2,961 B an unforked story of the same length costs (1.07×)**, naming one branch of its
|
||||
22 lineage entries. The pre-tree 600-action `--keep` fixture was migrated and then driven
|
||||
over HTTP end to end: index **1,840 B**, page load **64,149 B** — the same shapes as
|
||||
before the phase — and scrolling to the start took 9 pages and saw all 600 actions exactly
|
||||
once.
|
||||
|
||||
### SP3 — Memories and summary attach to nodes
|
||||
|
||||
Cursors stop being positions in a shifting list. `memory_cursor`/`summary_cursor` become
|
||||
|
||||
+47
-14
@@ -78,26 +78,29 @@ needed; nothing requires reading a row of anyone's story.
|
||||
|
||||
## Pick up here
|
||||
|
||||
**`plan/14-phase-story-tree.md`, SP2 — the branch clause.** SP0 (the regression contract
|
||||
and the `--rich` fixture) and SP1 (schema, migration, and the writer that keeps new rows
|
||||
on the tree) are done and green; nothing is deployed yet. SP2 is where the reads move
|
||||
onto `(branch_id, depth)`, and where the highest-risk line in the whole phase lives:
|
||||
`history._from_memory()` slices `adventure.actions`, which under a tree is *every
|
||||
branch's* actions rather than the path. Make that shortcut branch-aware or delete it.
|
||||
**`plan/14-phase-story-tree.md`, SP3 — memories and the summary attach to nodes.** SP0
|
||||
(the regression contract and the `--rich` fixture), SP1 (schema, migration, and the writer
|
||||
that keeps new rows on the tree) and SP2 (the branch clause: every action read now selects
|
||||
on `(branch_id, depth)` through `app/context/lineage.py`) are done and green; nothing is
|
||||
deployed yet. SP3 turns `memory_cursor`/`summary_cursor` from positions in a shifting list
|
||||
into node anchors, and deletes the cursor-position machinery that goes with them —
|
||||
`position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
|
||||
`prune_dangling_memories`. **`settled_story_actions` and the holdback stay until SP4**:
|
||||
they exist because retry mutates a row in place, and retry stops doing that in SP4, not
|
||||
in SP3. Deleting them early reopens the exact bug they were written for.
|
||||
|
||||
**The schema is live in code but not on production.** When SP1 ships, the deploy needs
|
||||
one `VACUUM FULL actions;` on the direct (non-`-pooler`) endpoint afterwards — it rewrites
|
||||
every row. See the 144 MB lesson at the top of this file.
|
||||
|
||||
Two things from the egress work are worth carrying into it:
|
||||
Two things to carry into it:
|
||||
|
||||
- **Paging already anticipates the tree.** `action_window` in `routers/adventures.py`
|
||||
anchors on an action id and orders by comparing `Action.index`, never by treating
|
||||
index as a position. A branch changes which actions are on the path, not how two of
|
||||
them order, so the anchor survives; anything counting offsets would not.
|
||||
- **Weigh new columns in bytes.** A tree adds parent/branch columns to `actions`, which
|
||||
is already the table that fills the disk. `tests/test_egress.py` has byte ceilings
|
||||
now — they will tell you.
|
||||
- **Every action read goes through `context/lineage.py`.** A memory read has to as well —
|
||||
the same `Path` builds a clause over `Memory` — and it reads the *full* lineage, not a
|
||||
window: retrieval is long-range recall and cannot be windowed. It stays affordable
|
||||
because memories are sparse, so assert the byte cost on a deep fork.
|
||||
- **Weigh new columns in bytes.** `actions` is already the table that fills the disk.
|
||||
`tests/test_egress.py` has byte ceilings now — they will tell you.
|
||||
|
||||
**After any migration that rewrites `actions`:** one `VACUUM FULL actions;`. That is the
|
||||
lesson of the 144 MB above — a rewrite doubles the table and only a `VACUUM FULL` gives
|
||||
@@ -109,6 +112,36 @@ drive it before rewriting it.
|
||||
|
||||
---
|
||||
|
||||
## What happened on 2026-08-17, part five — the tree, SP2
|
||||
|
||||
Every read of an action now goes through one module. `app/context/lineage.py` turns a
|
||||
branch's stored lineage into the OR-of-ranges that is "this story", and history, paging,
|
||||
the newest-action lookups, the index screen and the scripting history API all select
|
||||
through it. Ordering moved from `index` to `depth`. **317 tests green**, and the SP0
|
||||
baseline still passes unmodified, which was the pass condition. Branch `sp2-branch-clause`.
|
||||
|
||||
Three things to carry forward:
|
||||
|
||||
- **A read-side invariant needs a write-side floor.** From SP2 a row without a branch is a
|
||||
row no read can see, and it fails by *disappearing*. Wiring every writer was not enough,
|
||||
because the SP0 baseline and eleven other fixtures write actions straight to the database
|
||||
and never call `place_action` — and the baseline may not be edited. `tree.place_new_nodes`
|
||||
now runs from `Session.before_flush`, so nothing can be written unplaced. That is a
|
||||
better invariant than the one SP1 shipped, and it was the contract that forced it.
|
||||
- **The SQLAlchemy identity map is weak, and that is a performance cliff.** Resolving the
|
||||
head branch once per node re-read the row from the database for every node in a flush —
|
||||
201 SELECTs to write 200 actions, and a 25 % slower suite (36 s → 45 s). Nothing about
|
||||
the results changed; only a stopwatch could see it. Hoist the lookup out of the loop and
|
||||
hold the reference for the length of the call. Now pinned by a test.
|
||||
- **Clause count is bounded by the window, and it is now measured.** A story forked 20
|
||||
times reads its newest 32 actions naming *one* branch, for 1.07× what an unforked story
|
||||
of the same length costs. Reading the tail widens the lineage only when a deleted action
|
||||
leaves the estimate short.
|
||||
|
||||
The 600-action `--keep` fixture — a genuine pre-tree database — was migrated and then
|
||||
driven over HTTP: index 1,840 B, page load 64,149 B (both unchanged), and scrolling to the
|
||||
start took 9 pages and saw every action exactly once.
|
||||
|
||||
## What happened on 2026-08-17, part four — the tree, SP0 and SP1
|
||||
|
||||
No behaviour change, and none intended: a linear story is a tree with one branch, so
|
||||
|
||||
Reference in New Issue
Block a user