Mark the story with a node, not with a count

The memory bank and the story summary each kept a cursor: how many story
actions they had already covered. A count is a position in a list, and this
list moves — delete an action in front of the mark and every later one slides
down a slot, so the mark now covers one it has never read. All the cursor
bookkeeping existed to patch that up.

Both marks are now (branch_id, depth): the node up to and including which the
work is done. A depth is a coordinate along a path, not an offset into a list,
so nothing in front of it can move it. That deletes rather than rewrites
`position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
`prune_dangling_memories` and the every-pass clamp in `run_post_turn`.

A memory hangs off the node its block ends on, so a fork inherits its
ancestors' memories without copying any, and retrieval selects through the
branch clause over the *whole* lineage — recall is long-range by definition and
cannot be windowed. Measured: 1,807 B on a story forked twenty times against
1,823 B on a flat one of the same length.

Migrations 53-56 translate the old counts into nodes. They rewrite `adventures`
and not `actions`, so this one needs no VACUUM FULL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
This commit is contained in:
parththakkar106
2026-08-18 19:14:07 +05:30
committed by Parth
co-authored by Claude Opus 5
parent c7b6a46a8a
commit c51531709d
16 changed files with 1334 additions and 260 deletions
+146
View File
@@ -0,0 +1,146 @@
"""Phase 14 — how far along a story the derived work has got.
Two things are built from the story and stored beside it: the memories, and the
Story Summary. Both need to know where they left off, and that mark used to be
a *count* — "the first 12 story actions are covered". A count is a position in
a list, and this list moves: delete an action from in front of the mark and
every later action slides down a slot, so the mark now covers one it has never
seen. Every rule in `memorybank` about sliding cursors, rewinding them and
translating between positions and `Action.index` existed to patch that up, and
each was a separate chance to get it wrong in a way nothing reports.
A cursor here is an **anchor**: `(branch_id, depth)`, the node up to and
including which the work is done. Deleting an action does not move it, because
a depth is not a position — it is a coordinate along a path. "What is not
covered yet" becomes `history.count_after(anchor)`, which is a question about
the story rather than about a list index, and it answers correctly whatever has
been deleted from in front of it.
The branch half is what makes it survive forking. A depth alone is ambiguous
once two branches have a node 41; the anchor says which one, and
`Path.depth_on` reads it back as a depth on whatever story is being played —
capped at the fork, or "nothing covered" if the anchor sits on ground this path
never travelled. Until forking ships there is one branch and that is always a
no-op, which is the point: the coordinate system is right before anything needs
it to be.
`NO_DEPTH` (-1) is "nothing covered", so a fresh adventure needs no special
case: every node is deeper than -1.
"""
from sqlalchemy.orm import Session
from .. import models
from . import history, lineage
NO_DEPTH = lineage.NO_DEPTH
class Cursor:
"""One anchor on the adventure row: the memory bank's, or the summary's.
A pair of columns rather than a foreign key to the node. The node can be
deleted — that is most of what undo does — and the boundary is still
meaningful afterwards, so a pointer that has to resolve would be a pointer
that keeps not resolving.
"""
def __init__(self, name: str):
self.name = name
self.branch_field = f"{name}_cursor_branch_id"
self.depth_field = f"{name}_cursor_depth"
# ------------------------------------------------------------- reading
def stored(self, adventure: models.Adventure) -> tuple[int | None, int]:
"""The anchor exactly as written, unread by any path."""
depth = getattr(adventure, self.depth_field)
return getattr(adventure, self.branch_field), (
NO_DEPTH if depth is None else depth
)
def depth(self, db: Session, adventure: models.Adventure) -> int:
"""The anchor as a depth on the story currently being played."""
branch_id, depth = self.stored(adventure)
return lineage.path_of(db, adventure).depth_on(branch_id, depth)
# ------------------------------------------------------------- writing
def anchor_at(self, adventure: models.Adventure, node: models.Action) -> None:
"""Mark the work done up to and including `node`.
Takes the node's own branch, not the adventure's head: a block of six
actions can end before the fork this branch was made at, and the
coverage belongs where the ground is.
"""
setattr(adventure, self.branch_field, node.branch_id)
setattr(adventure, self.depth_field, lineage.NO_DEPTH
if node.depth is None else node.depth)
def rewind_to(
self, adventure: models.Adventure, branch_id: int | None, depth: int
) -> None:
"""Move the anchor back to `depth` if it is past it; never forward.
The one direction that is safe without knowing what else has happened:
re-covering ground costs a summarizer call, skipping it loses a stretch
of story out of the memories for good.
"""
_, current = self.stored(adventure)
if current <= depth:
return
setattr(adventure, self.branch_field, branch_id)
setattr(adventure, self.depth_field, max(depth, NO_DEPTH))
MEMORY = Cursor("memory")
SUMMARY = Cursor("summary")
ALL = (MEMORY, SUMMARY)
def rewind_all(
adventure: models.Adventure, branch_id: int | None, depth: int
) -> None:
"""Hand a stretch of story back to *both* passes.
They move together because they cover the same ground from different sides:
the summary folds in the memories, so a memory withdrawn without rewinding
the summary leaves the summary claiming to have read something no longer
there.
"""
for cursor in ALL:
cursor.rewind_to(adventure, branch_id, depth)
def anchor_at_position(
adventure: models.Adventure, cursor: Cursor, position: int
) -> None:
"""Set `cursor` from a count of covered story actions — a v1 bundle's mark,
or a database written before the anchors existed.
The position-th story action in depth order is the node that says the same
thing, and goes on saying it once something in front of it is deleted. A
position past the end of the story is not a bad value: an adventure caught
up under the older rule can carry one, and it means the same thing the tip
does, so that is where it lands.
The SQL half of this rule is `migrations._backfill_cursor_anchors`, which
has to do it for every adventure at once without loading any of them; the
two must agree.
"""
if position <= 0:
return
covered = history.slice_(adventure, position - 1, 1) or history.tail(adventure, 1)
if covered:
cursor.anchor_at(adventure, covered[0])
def position_of(adventure: models.Adventure, depth: int) -> int:
"""How many story actions lie at or before `depth` — an anchor read back as
a count.
The v1 export bundle stores the cursors as positions, and a v1 bundle is
read by builds that have never heard of a depth. This is the one place that
still speaks that coordinate system, and SP6's v2 format retires it.
"""
return max(history.count(adventure) - history.count_after(adventure, depth), 0)
+80 -17
View File
@@ -14,10 +14,9 @@ story.
Three rules hold everything together:
* **One definition of "story action".** The cursors in memorybank are
*positions* in this filtered, depth-ordered list, so SQL and Python must
agree on membership exactly or a cursor silently points at a different
action. `_STORY_TEXT` and `is_story_text()` are that one definition, written
* **One definition of "story action".** Membership decides what a reader sees
and what the summarizer is handed, so SQL and Python must agree on it
exactly. `_STORY_TEXT` and `is_story_text()` are that one definition, written
twice; keep them in step.
* **Never load twice.** If `adventure.actions` is already in memory (the
scripting pipeline hands the whole history to user scripts, as AI Dungeon
@@ -32,6 +31,12 @@ Three rules hold everything together:
Ordering is by `depth` now, not `index`. The two hold the same numbers until
retry stops mutating rows (SP4), but only one of them is a position along a
path.
SP3 added the reads that count *from a node* rather than from the start —
`count_after`, `after`, `newest_settled`. The memory bank used to ask for
"positions 12 to 18 of the story", which is a question whose answer moves when
an action is deleted from in front of it. It now asks for "the six actions
after depth 41", which is the same question a fork has to answer anyway.
"""
from sqlalchemy import func, inspect as sa_inspect
@@ -162,6 +167,7 @@ def _count_query(
adventure: models.Adventure,
path: lineage.Path,
exclude_action_id: int | None,
entries: int | None = None,
):
"""A real `SELECT count(...)`.
@@ -172,7 +178,7 @@ def _count_query(
that greps the SQL cannot tell the two apart.
"""
return db.query(func.count(models.Action.id)).filter(
*_filters(adventure, path, exclude_action_id)
*_filters(adventure, path, exclude_action_id, entries)
)
@@ -311,30 +317,87 @@ def slice_(
)
def position_of_index(adventure: models.Adventure, index: int) -> int:
"""The position the story action with `Action.index == index` occupies —
i.e. how many story actions come before it.
def depth_of(action: models.Action) -> int:
"""`action.depth`, with the no-depth case spelled once.
Translates between the two coordinate systems that keep tripping this code
up: cursors are positions, `Memory.source_start/_end` are `Action.index`
values, and the two diverge the moment anything is deleted.
A row with no depth is a pre-tree row, which no path contains — so it can
only turn up in an already-loaded collection, and it sorts before the story
rather than after it.
"""
in_memory = _from_memory(adventure, None)
return action.depth if action.depth is not None else lineage.NO_DEPTH
def count_after(
adventure: models.Adventure, depth: int, exclude_action_id: int | None = None
) -> int:
"""How many story actions lie past `depth` on the path.
The node-anchored replacement for "the story is N long and the cursor is at
M". Deleting an action from in front of the boundary makes this number
smaller, which is true; it does not make the boundary point somewhere else,
which is the bug the positions had.
`covering_after` says exactly which lineage entries can hold a node deeper
than the boundary, so a cursor near the tip names one branch however many
forks are below it.
"""
in_memory = _from_memory(adventure, exclude_action_id)
if in_memory is not None:
return next(
(i for i, a in enumerate(in_memory) if a.index >= index), len(in_memory)
)
return sum(1 for a in in_memory if depth_of(a) > depth)
db = _session(adventure)
if db is None:
return 0
path = _path(db, adventure)
return (
_count_query(db, adventure, _path(db, adventure), None)
.filter(models.Action.index < index)
_count_query(
db, adventure, path, exclude_action_id, path.covering_after(depth)
)
.filter(models.Action.depth > depth)
.scalar()
or 0
)
def after(
adventure: models.Adventure,
depth: int,
limit: int,
exclude_action_id: int | None = None,
) -> list[models.Action]:
"""The oldest `limit` story actions past `depth`, oldest first.
"The next block the summarizer has not seen", asked as a fact about the
story rather than as an offset into a list that shifts underneath it.
"""
if limit <= 0:
return []
in_memory = _from_memory(adventure, exclude_action_id)
if in_memory is not None:
return [a for a in in_memory if depth_of(a) > depth][:limit]
db = _session(adventure)
if db is None:
return []
path = _path(db, adventure)
return (
_query(db, adventure, path, exclude_action_id, path.covering_after(depth))
.filter(models.Action.depth > depth)
.order_by(*_OLDEST_FIRST)
.limit(limit)
.all()
)
def newest_settled(adventure: models.Adventure) -> models.Action | None:
"""The newest story action that is not the newest one — see
`memorybank.settled_story_actions` for why one is always held back.
Two rows, not a count and an offset: this is the node an anchor moves to
when derived work catches up with the settled end of the story.
"""
rows = tail(adventure, 2)
return rows[0] if len(rows) == 2 else None
def max_action_index(adventure: models.Adventure) -> int:
"""Highest `Action.index` in the adventure, story text or not. -1 if empty.
+63 -4
View File
@@ -83,13 +83,25 @@ class Path:
# ---------------------------------------------------------------- SQL
def clause(self, model=models.Action, count: int | None = None):
def clause(
self,
model=models.Action,
count: int | None = None,
unanchored: bool = False,
):
"""The branch clause, over `model` (`Action` or `Memory`).
`count` limits it to the newest `count` lineage entries — the windowed
read. `None` is the whole lineage, which is what anything counting from
the *oldest* end (a slice, a total) has to use.
`unanchored` keeps rows with no depth. Only memories ever have one: a
hand-written memory summarises no node, so it has a branch but no
depth, and a capped `depth <= n` would drop it the moment its branch
stopped being the newest entry — a memory vanishing at the first fork
after it was typed. An action with no depth is a pre-tree row that no
read should see, so actions never pass this.
An empty path yields `false`, not "no filter": an adventure whose nodes
carry no branch has no story, and the loud version of that is an empty
page, not every branch at once.
@@ -97,13 +109,16 @@ class Path:
entries = self.entries if count is None else self.entries[:count]
if not entries:
return false()
return or_(*[self._entry_clause(model, b, d) for b, d in entries])
return or_(*[self._entry_clause(model, b, d, unanchored) for b, d in entries])
@staticmethod
def _entry_clause(model, branch_id: int, max_depth: int | None):
def _entry_clause(model, branch_id: int, max_depth: int | None, unanchored=False):
if max_depth is None:
return model.branch_id == branch_id
return and_(model.branch_id == branch_id, model.depth <= max_depth)
within = model.depth <= max_depth
if unanchored:
within = or_(within, model.depth.is_(None))
return and_(model.branch_id == branch_id, within)
# ------------------------------------------------------------- Python
@@ -158,6 +173,50 @@ class Path:
return i + 1
return total
def covering_after(self, depth: int) -> int:
"""How many lineage entries can hold a node deeper than `depth`.
The counterpart to `prefix_covering`, and unlike it this is exact
rather than an estimate: entry *i* holds nothing deeper than its own
cap, and the caps descend, so the first entry capped at or below
`depth` ends the search — it and everything older is behind the
boundary. Reading "the story after the cursor" therefore names one
branch on any story whose cursor is on its newest branch, however
often it has forked.
"""
for i, (_, max_depth) in enumerate(self.entries):
if max_depth is not None and max_depth <= depth:
return i
return len(self.entries)
def depth_on(self, branch_id: int | None, depth: int) -> int:
"""A stored `(branch_id, depth)` anchor, read as a depth on *this* path.
An anchor is how far along a story some derived work has got — which
memories cover, what the summary has folded in. It names a node, so
moving to another path has to be answered rather than assumed:
* the anchor's branch is on this path — the depth stands, capped at the
fork the path takes off that branch, because nothing past the fork is
on this story;
* the branch is not on this path at all — the work was done on ground
this story never travelled, so nothing here is covered.
The second case cannot arise while an adventure has one branch: the
anchor is always set from a node on it. It exists because the fallback
for "I don't know" must be to redo the work, not to skip it.
"""
if depth <= NO_DEPTH:
return NO_DEPTH
if branch_id is None:
# A pre-tree anchor, or one set by hand. There is one story, so the
# depth is a position in it and means what it says.
return depth
for entry_branch, max_depth in self.entries:
if entry_branch == branch_id:
return depth if max_depth is None else min(depth, max_depth)
return NO_DEPTH
def branch_of(db: Session, adventure: models.Adventure) -> models.Branch | None:
"""The branch this adventure is being read at, or None if it has none.