Files
interactive-story/backend/app/attempts.py
T
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00

318 lines
14 KiB
Python

"""Phase 14, SP4: the attempts at one turn.
A retry used to rewrite the AI action in place and append the discarded attempt
to a JSON list on the same row. Seven separate bugs came from that arrangement.
The row's `text` duplicated one entry of a repeating group, a second column
duplicated its length, and every reader that touched the story during a retry
had to be told to ignore the row.
Now an attempt is a node. A retry writes a sibling at the same `(branch_id,
depth)` and marks it live. The previous attempt stays as it was written, at the
same coordinate, with `live = False`. Nothing is duplicated, so nothing can
diverge.
Two invariants hold the arrangement together, and this module is the only place
that maintains either one:
* Exactly one sibling in a group is live. `lineage.Path.clause` selects on it, so
the other attempts are invisible to every read of the story, and none of those
reads has to know that attempts exist.
* The assembled prompt is stored once per turn, on the live sibling. A
`context_snapshot` is about 163 kB of prompt that every attempt at a turn
shares, plus a few hundred bytes that differ, listed in `ATTEMPT_KEYS`. Giving
each sibling its own copy would make a retry a permanent multiplier on the
largest column in the database, which is what the JSON list was invented to
avoid. The prompt therefore moves with the live flag, and a superseded sibling
keeps only its own slices.
Ordering inside a group comes from `id`, not from `created_at`. Two attempts made
in the same second still have to page in the order they were made, and `id`
increases with every insert. SP8 dropped `variant_index`, an explicit ordinal
that carried the same order, once a run of the suite confirmed that the two
agreed in every group.
"""
import copy
from sqlalchemy.orm import Session, object_session, undefer
from . import models, summaries
from .context import lineage
from .narrative import model as narrative_model
# The slices of a context snapshot that belong to one attempt rather than to the
# turn. They are the world-state delta the attempt proposed and what the engine
# did with it, the model's literal reply, and the endpoint's
# token accounting. Each attempt is its own API call, and a retry is the call
# most likely to read the prompt back out of cache. Everything else in a snapshot
# is the prompt, which is assembled once per turn.
#
# v1.1 WP-A1: `accounting` is one attempt's too. It compares the server's count
# for *that* call with the turn's estimate. Left out of this tuple, it was
# treated as part of the shared prompt, so moving the live flag handed the
# superseded attempt's accounting to the new live one and threw the new one's
# away. Found by the A2 long run: two retries and one take selection left three
# attempts reporting no accounting, or another attempt's.
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage", "accounting")
# ------------------------------------------------------------------ reading
def group(db: Session, action: models.Action) -> list[models.Action]:
"""Returns every attempt at `action`'s turn, oldest first.
The query keys on the parent rather than on the coordinate (SP9). The two
agree until an attempt is forked onto its own branch. That attempt keeps its
parent but leaves the `(branch, depth)` its siblings are still at, so a
coordinate would report it as the only attempt at its turn, showing `1/1`
where the player should see `1/3`.
The parent also nests groups correctly without extra work. Attempts under C1
and attempts under C2 share a depth, and until one of them forks they share a
branch. Only the parent separates them, which is what makes a pager under C2
read `2/2` rather than count C1's three as well.
There are two fallbacks, and both mean the row predates the key being asked
about. A node with no branch is a pre-tree row that no path contains, and a
node with no parent is a pre-SP9 row the backfill could not place. Under the
rule each was written with, both are the only attempt at their turn.
"""
if action.branch_id is None or action.depth is None:
return [action]
if action.parent_id is None:
# This row is pre-SP9, and the coordinate is the key those rows were
# written under. A root node also reaches this branch and is genuinely
# alone, because nothing is an attempt at the opening of a story.
return (
db.query(models.Action)
.filter(
models.Action.adventure_id == action.adventure_id,
models.Action.branch_id == action.branch_id,
models.Action.depth == action.depth,
models.Action.parent_id.is_(None),
)
.order_by(models.Action.id)
.all()
)
return (
db.query(models.Action)
.filter(
models.Action.adventure_id == action.adventure_id,
models.Action.parent_id == action.parent_id,
)
.order_by(models.Action.id)
.all()
)
def on_branch(rows: list[models.Action], node: models.Action) -> list[models.Action]:
"""Returns the attempts in `rows` that are on `node`'s own branch.
`group` reports which attempts belong to this turn, and since SP9 that spans
branches. An attempt forked onto its own line is still an attempt at the same
turn, which is the reason for keying on the parent.
Deletion is the one caller that must not follow a group across branches. An
attempt on another branch is reachable through that branch and belongs to the
story someone is telling there. Removing it because a turn was undone here
would delete a line nobody asked about. The same parent and the same branch
together are the coordinate, which is what every attempt at this turn meant
before a fork could move one out of it.
"""
return [row for row in rows if row.branch_id == node.branch_id]
def live_in(rows: list[models.Action]) -> models.Action | None:
for row in rows:
if row.live:
return row
return None
def preceding(
db: Session, adventure: models.Adventure, node: models.Action
) -> models.Action | None:
"""Returns the node the story tells immediately before `node`.
This reads "before this turn" as a fact about the path rather than as a
snapshot taken from inside the turn, which is what makes the after-snapshots
sufficient on their own. The query undefers both of them, because the only
reason to fetch this row is to restore what it left behind.
"""
if node.depth is None:
return None
return (
db.query(models.Action)
.filter(
models.Action.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Action),
models.Action.depth < node.depth,
)
.options(
undefer(models.Action.state_after),
undefer(models.Action.world_state_after),
)
.order_by(models.Action.depth.desc(), models.Action.id.desc())
.first()
)
# ------------------------------------------------------------------ writing
def restore_state(adventure: models.Adventure, node: models.Action | None) -> None:
"""Restores the state that `node` left behind.
This is what makes Undo, Redo, a branch switch and a Save Point restore cost
the same at any distance: the destination node carries its own outcome, so
arriving is a row read rather than a replay (`TECHNICAL-DESIGN.md` §10.4).
M5 changed what is restored, not how — the narrative state document takes
the place the RPG world state held, through the same single function.
The two columns follow **different** rules about a NULL, and the difference
is not an oversight.
For the narrative document, a NULL means *this position established
nothing*, and it is restored as the empty document. Leaving the live state
alone instead is what the M5 review caught (Finding 3): arriving at a
migrated pre-M5 node left a later position's entities, facts and threads
standing, so the transcript said depth 2 while the state described depth 6.
The invariant this module exists to hold is that the visible position, the
head and the authoritative state agree, and "keep whatever was there" cannot
hold it. An empty document at an old position is honest — the narrative
state system knew nothing then, because it did not exist — where retained
state from elsewhere is a claim about a story that had not been told yet.
Migration backfills those rows explicitly, so this fallback is the belt to
that pair of braces: it also covers a node arriving from an older export,
which the migration never sees.
For the legacy RPG world state a NULL still means leave it alone. Those rows
predate SP4, nothing consults the values to decide anything, and overwriting
a running adventure's numbers with an empty dict would be worse than doing
nothing.
"""
if node is None:
return
adventure.narrative_state = (
copy.deepcopy(node.narrative_state_after)
if isinstance(node.narrative_state_after, dict)
else narrative_model.empty()
)
# M6: the reader-facing summary mirror follows the head too. It is a
# convenience column with no lineage of its own, so without this it would go
# on showing a summary belonging to a position the story has left. Nothing
# authoritative reads it — the prompt takes its summary from
# `summaries.current` — but the Plot panel and the export bundle do.
session = object_session(adventure)
if session is not None:
summaries.refresh_mirror(session, adventure)
# Legacy, and deliberately still restored: a pre-M5 campaign's numbers stay
# coherent with the position being read, so an old save is not left showing
# a future's values. Nothing consults them to decide anything.
if isinstance(node.world_state_after, dict):
adventure.world_state = copy.deepcopy(node.world_state_after)
def snapshot_outcome(adventure: models.Adventure, node: models.Action) -> None:
"""Records on `node` the state of the adventure now that the node has played.
Every node, including a player's action that changed nothing. A position
without a snapshot is a position the head cannot be restored to, and the
head can rest on any node.
"""
world = adventure.world_state if isinstance(adventure.world_state, dict) else {}
# `state_after` held the scripting engine's shared state, which M2 removed.
# The column stays for schema compatibility and is written empty.
node.state_after = {}
node.world_state_after = copy.deepcopy(world)
narrative = adventure.narrative_state
node.narrative_state_after = copy.deepcopy(
narrative if isinstance(narrative, dict) else narrative_model.empty()
)
def roll_back_before(
db: Session, adventure: models.Adventure, node: models.Action
) -> None:
"""Rewinds the shared state to what it was before `node` was played."""
restore_state(adventure, preceding(db, adventure, node))
def add_attempt(
db: Session,
adventure: models.Adventure,
previous: models.Action,
replacement: models.Action,
) -> None:
"""Places `replacement` next to `previous` as the newer attempt at that turn.
The placement is done here rather than through `tree.place_action`, which
moves the head. A sibling is not a new turn. It is another attempt at the
turn the head is already on.
"""
replacement.branch_id = previous.branch_id
replacement.depth = previous.depth
# Copy the parent rather than resolve it from the path. An attempt belongs
# to the turn it is an attempt at, and that is what `group` keys on.
# Resolving it here would ask which node is live one depth back. That is the
# same node right now, and it stops being the same node once the story forks
# away from this turn.
replacement.parent_id = previous.parent_id
replacement.live = True
# The replacement takes its place at the end of the group, because `group`
# orders by `id` and this row has no id yet. Switching a three-attempt turn
# back to attempt 1 and retrying therefore still pages 1, 2, 3, 4, which is
# the order the attempts were made in.
previous.live = False
# The replacement was assembled with a fresh snapshot, so the prompt for
# this turn is now the one it carries. The superseded attempt keeps only the
# slices that were its own.
keep_own_slices(previous)
def make_live(
db: Session, adventure: models.Adventure, node: models.Action
) -> list[models.Action]:
"""Makes `node` the attempt the story tells, and restores its outcome.
Returns the group, so that a caller reporting on it does not read it twice.
"""
rows = group(db, node)
previous = live_in(rows)
if previous is not None and previous is not node:
hand_over_the_prompt(previous, node)
for row in rows:
row.live = row is node
restore_state(adventure, node)
return rows
# ------------------------------------------------- the prompt, stored once
def keep_own_slices(node: models.Action) -> None:
"""Reduces `node`'s snapshot to the slices that are only its own."""
snapshot = node.context_snapshot
if not isinstance(snapshot, dict):
return
node.context_snapshot = {
key: snapshot[key] for key in ATTEMPT_KEYS if key in snapshot
} or None
def hand_over_the_prompt(giver: models.Action, taker: models.Action) -> None:
"""Moves the turn's assembled prompt from one attempt to another.
The caller runs this when the live flag moves, so that the row in the story
is always the row the Insights viewer can explain. Nothing is copied. The
prompt exists once before and once after, on whichever sibling is being read.
"""
held = giver.context_snapshot if isinstance(giver.context_snapshot, dict) else {}
shared = {k: v for k, v in held.items() if k not in ATTEMPT_KEYS}
if not shared:
return
keep_own_slices(giver)
own = taker.context_snapshot if isinstance(taker.context_snapshot, dict) else {}
taker.context_snapshot = shared | {
k: v for k, v in own.items() if k in ATTEMPT_KEYS
}