Files
interactive-story/backend/tests/test_process_restart.py
T
JesseMarkowitzandClaude Opus 5 62a997f364 M4: close out Save Points, with browser verification
Closes M4. The review's three findings are fixed, the durability rule the
specification always implied is now enforced, and M3's and M4's browser
behaviour has been verified in a real browser for the first time.

B-1 -- the Save Point list was an N+1 that loaded whole Action rows,
narration included, to answer "does a row exist here". It is now one bulk
two-column coordinate query plus one lineage: 53 SELECTs for 25 Save Points
became 5, and the count no longer grows with the list. The clause is an OR
of exact (branch, depth) pairs rather than two IN lists, because the cross
product would report a Save Point resolved on the strength of another one's
depth existing on this one's branch. A test builds exactly that trap.

B-2 -- reclassified during closeout from "missing warning" to a behaviour
defect, and fixed as one. STORY-BRANCH-SEMANTICS §19 says a named checkpoint
remains until explicitly deleted, and §28 already required future cleanup to
retain checkpoint-referenced paths; a cascade that silently removed Save
Points with a branch violated both, and a warning would only have documented
the violation. A branch a Save Point names can no longer be deleted. The
request is refused with the offending Save Points named, the user deletes
them explicitly -- which deletes no story -- and the branch then goes. The
scope is the subtree, because deleting a branch takes its descendants. Both
delete controls disable and explain. Recorded as a new §19.1; models.py,
TECHNICAL-DESIGN §8.8 and DATA-MODEL §8 had all recorded the cascade as the
rule and now record the refusal.

An earlier pass in this same closeout had kept the cascade and added a
warning. That was the wrong fix and its tests were replaced rather than left
standing, since they pinned the defect.

B-3 -- the D11/L03 automation never left one process, so it could not
distinguish durable state from a live Python object. It now spawns real
server processes, kills the first, and reads the campaign back with the
second.

C-5 -- creating a Save Point takes the campaign's turn lock. "Save where I
am" has to name one committed position, and the head is what a turn in
flight is about to move. Rename and Delete deliberately do not take it.

The architecture is untouched: a Save Point is still name + note +
(branch, depth), and restore is still coordinate -> head.move_to_node ->
head.move_to -> attempts.restore_state. No second restore path, no state
copied into a checkpoint, no fork on restore.

Browser verification -- the first in this project, and it covers both
milestones. Firefox 154.0.1 through geckodriver over the W3C WebDriver
protocol, driving the rendered DOM: 47/47 checks, twice, on independent
databases, no console errors. M3's Undo/Redo enable states, transcript
movement, Retry and the take pager, divergence retiring Redo; M4's whole
Save Point lifecycle, both confirmations, and the new branch-delete refusal
including its recovery. No dependency was added: the WebDriver client is
stdlib HTTP.

No application defect was found by the browser. Four failures occurred, all
in the harness -- a wrong SPA route, a wait comparing transcript length when
the empty-story placeholder is longer than the first turn, a fixture
deleting the branch it was reading, and a reload assertion that sampled
once instead of waiting. The last was checked against the app before being
called a harness bug.

Tests: 698 backend pass (was 680), 60 M4, 94 M3 history, 66 export/
migrations, 93 security/local-only. Frontend lint and build clean, Docker
build clean, loopback binding unchanged. No assertion weakened, no skip
added.

Planning: STORY-BRANCH-SEMANTICS §19.1 is the only behavioural change and it
strengthens §19. V1-ACCEPTANCE-TESTS records D11-D14, I04, L03 and the
E-series, keeping automated, live-runtime and browser evidence distinct, and
weakens no pass condition. DATA-MODEL records the coordinate with the retry
measurement that settles it. BROWSER-UX-SPEC rules for Moment over Turn.
BUILD-MILESTONES marks M4 COMPLETE, closes M3's browser condition, and lists
what M5 inherits. VERSION adds v2.6.

No new ADR: ADR 005 already decides that history is preserved rather than
overwritten, and §19.1 is that decision applied to checkpoint-referenced
history.

M4 is closed. M5 may now be briefed; it has not been started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-04 06:34:56 -04:00

326 lines
13 KiB
Python

"""D11 and L03 across a genuine OS process boundary.
The rest of the suite runs the app in-process through `TestClient`, which is the
right tool for almost everything and the wrong one for exactly one claim:
*durability*. A Save Point that survived only because a Python object was still
alive would pass a same-process test and fail a user's restart. `TestClient`
cannot tell those apart, so the M4 review recorded the shipped D11/L03 tests as
weaker than the acceptance items they were named for (§R B-3).
This module closes that. It starts the real application as a **subprocess**,
plays a story over HTTP, kills the process, starts a **second** process against
the same database file, and only then asks whether the Save Point is still
there. Everything crossing the boundary crosses it as bytes on disk.
Deterministic and local: the spawned server replaces the model with a scripted
provider (`_restart_server.py`), so there is no Ollama, no network and no
sleep-and-hope — readiness is probed, not waited for.
python -m pytest tests/test_process_restart.py -v
"""
import json
import os
import socket
import subprocess
import sys
import tempfile
import time
import urllib.error
import urllib.request
from pathlib import Path
import pytest
from fakes import GOLD_PER_TURN
HERE = Path(__file__).resolve().parent
SERVER = HERE / "_restart_server.py"
# How long a spawned server may take to answer before the test gives up. The
# process imports the app and runs migrations on a fresh file, which is well
# under a second on this project; the ceiling is for a loaded machine.
STARTUP_TIMEOUT = 60.0
def _free_port() -> int:
"""Returns a port nothing is listening on.
Bind, read, release. There is a race between releasing and the server
claiming it, which is why the caller probes for readiness rather than
assuming success — a lost race shows up as a startup timeout, not as a
silent pass.
"""
with socket.socket() as s:
s.bind(("127.0.0.1", 0))
return s.getsockname()[1]
class Server:
"""One storyteller process, and the HTTP calls the test makes against it."""
def __init__(self, db_path: str, port: int):
self.port = port
self.proc = subprocess.Popen(
[sys.executable, str(SERVER), db_path, str(port)],
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
# Never inherit the parent's database redirection; the child is told
# which file to open on its command line.
env={**os.environ, "AIDND_DB_PATH": db_path},
)
# ------------------------------------------------------------ lifecycle
def wait_until_ready(self) -> None:
deadline = time.monotonic() + STARTUP_TIMEOUT
while time.monotonic() < deadline:
if self.proc.poll() is not None:
raise AssertionError(
f"server exited early ({self.proc.returncode}):\n{self._output()}"
)
try:
self.call("GET", "/settings")
return
except (urllib.error.URLError, ConnectionError, OSError):
time.sleep(0.05)
raise AssertionError(f"server never became ready:\n{self._output()}")
def stop(self) -> None:
"""Ends the process, and does not return until it is actually gone."""
if self.proc.poll() is None:
self.proc.terminate()
try:
self.proc.wait(timeout=15)
except subprocess.TimeoutExpired:
self.proc.kill()
self.proc.wait(timeout=15)
if self.proc.stdout is not None:
self.proc.stdout.close()
def _output(self) -> str:
if self.proc.stdout is None:
return "(no output captured)"
try:
return self.proc.stdout.read().decode(errors="replace")[-2000:]
except Exception:
return "(output unreadable)"
def is_listening(self) -> bool:
try:
self.call("GET", "/settings")
return True
except Exception:
return False
# ---------------------------------------------------------------- HTTP
def call(self, method: str, path: str, payload=None, expect: int | None = None):
data = json.dumps(payload).encode() if payload is not None else None
request = urllib.request.Request(
f"http://127.0.0.1:{self.port}/api{path}",
data=data,
method=method,
headers={"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(request, timeout=60) as response:
body, status = response.read(), response.status
except urllib.error.HTTPError as exc: # a real answer, not a failure
body, status = exc.read(), exc.code
if expect is not None and status != expect:
raise AssertionError(f"{method} {path} -> {status}: {body[:400]!r}")
return json.loads(body) if body and status != 204 else None
def play(self, adventure_id: int, text: str) -> None:
"""Plays one turn through the streaming endpoint, to completion."""
request = urllib.request.Request(
f"http://127.0.0.1:{self.port}/api/adventures/{adventure_id}/actions",
data=json.dumps({"type": "do", "text": text}).encode(),
method="POST",
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(request, timeout=120) as response:
response.read()
# ------------------------------------------------------------- reading
def transcript(self, adventure_id: int) -> list[str]:
page = self.call("GET", f"/adventures/{adventure_id}", expect=200)
return [a["text"] for a in page["actions"]]
def gold(self, adventure_id: int) -> int:
state = self.call("GET", f"/adventures/{adventure_id}/world-state", expect=200)
return state["state"]["player"]["gold"]
def total_rows(self, adventure_id: int) -> int:
"""Every row of the whole tree, head or no head.
`action_count` is scoped to the path being read, so it falls when the
head moves back — which is the feature, not a deletion. The export
carries the entire tree whatever the head is doing, so it is what
"nothing was deleted" has to be measured against.
"""
bundle = self.call("GET", f"/adventures/{adventure_id}/export", expect=200)
return len(bundle["actions"])
@pytest.fixture()
def workspace():
"""A database file, and whichever servers a test starts against it."""
directory = tempfile.mkdtemp(prefix="m4-restart-")
db_path = os.path.join(directory, "campaign.db")
started: list[Server] = []
def start() -> Server:
server = Server(db_path, _free_port())
started.append(server)
server.wait_until_ready()
return server
try:
yield start
finally:
# Every child dies even if the test failed part way through, and each
# stop() waits, so a later test cannot inherit a live listener.
for server in started:
server.stop()
def _campaign_with_a_save_point(server: Server):
"""Three turns, a Save Point on the third, then four more turns.
Returns everything the second process has to be able to reproduce.
"""
scenario = server.call("POST", "/scenarios", {
"title": "Abbey",
"stat_schema": {"player": {"gold": {"initial": 0, "min": 0, "max": 9999}}},
}, expect=201)
adventure = server.call("POST", "/adventures", {
"title": "The abbey", "scenario_id": scenario["id"],
}, expect=201)
adventure_id = adventure["id"]
for n in range(3):
server.play(adventure_id, f"turn {n}")
at_save = {
"transcript": server.transcript(adventure_id),
"gold": server.gold(adventure_id),
}
save_point = server.call("POST", f"/adventures/{adventure_id}/checkpoints",
{"name": "Before entering the abbey"}, expect=201)
for n in range(4):
server.play(adventure_id, f"later {n}")
at_tip = {
"transcript": server.transcript(adventure_id),
"gold": server.gold(adventure_id),
"rows": server.total_rows(adventure_id),
}
return adventure_id, save_point, at_save, at_tip
def test_d11_l03_a_save_point_survives_a_real_process_restart(workspace):
"""D11 and L03 together, across a boundary a same-process test cannot cross.
The first process writes the campaign and exits. The second process is a
different interpreter with an empty session, an empty identity map and no
memory of anything — everything it knows, it reads off the disk.
"""
first = workspace()
adventure_id, save_point, at_save, at_tip = _campaign_with_a_save_point(first)
assert at_tip["gold"] == at_save["gold"] + 4 * GOLD_PER_TURN
assert len(at_tip["transcript"]) == len(at_save["transcript"]) + 8
# --- the boundary -----------------------------------------------------
first.stop()
assert first.proc.poll() is not None, "the first server did not actually exit"
assert not first.is_listening(), "the first server is still answering"
second = workspace()
assert second.proc.pid != first.proc.pid
# --- D11: the Save Point is still there -------------------------------
listed = second.call("GET", f"/adventures/{adventure_id}/checkpoints", expect=200)
assert [c["id"] for c in listed] == [save_point["id"]]
assert listed[0]["name"] == "Before entering the abbey"
assert (listed[0]["branch_id"], listed[0]["depth"]) == (
save_point["branch_id"], save_point["depth"]
)
assert listed[0]["resolved"] is True
# The story came back too, at the position the first process left it.
assert second.transcript(adventure_id) == at_tip["transcript"]
rows_before_restore = second.total_rows(adventure_id)
assert rows_before_restore == at_tip["rows"]
# --- L03: restoring reconstructs the historical position --------------
page = second.call(
"POST", f"/adventures/{adventure_id}/checkpoints/{save_point['id']}/restore",
expect=200,
)
assert [a["text"] for a in page["actions"]] == at_save["transcript"]
assert second.transcript(adventure_id) == at_save["transcript"]
assert second.gold(adventure_id) == at_save["gold"]
# --- restore deleted nothing, and Redo is still available -------------
assert second.total_rows(adventure_id) == rows_before_restore
assert page["can_redo"] is True
assert second.call("GET", f"/adventures/{adventure_id}", expect=200)["can_redo"] is True
def test_the_retained_continuation_is_reachable_after_a_restart(workspace):
"""The later history is not merely present in the database after a restart —
it is still the continuation the story tells, walkable by ordinary Redo."""
first = workspace()
adventure_id, save_point, at_save, at_tip = _campaign_with_a_save_point(first)
first.stop()
second = workspace()
second.call("POST",
f"/adventures/{adventure_id}/checkpoints/{save_point['id']}/restore",
expect=200)
assert second.transcript(adventure_id) == at_save["transcript"]
steps = 0
while second.call("GET", f"/adventures/{adventure_id}", expect=200)["can_redo"]:
second.call("POST", f"/adventures/{adventure_id}/redo", expect=200)
steps += 1
assert steps <= 10, "Redo never stopped"
assert steps == 4
assert second.transcript(adventure_id) == at_tip["transcript"]
assert second.gold(adventure_id) == at_tip["gold"]
assert second.total_rows(adventure_id) == at_tip["rows"]
def test_a_divergent_write_after_a_restart_keeps_the_displaced_future(workspace):
"""D13's second half, across the boundary: the fork still happens on the
write rather than on the restore, and the displaced turns keep their rows."""
first = workspace()
adventure_id, save_point, at_save, at_tip = _campaign_with_a_save_point(first)
first.stop()
second = workspace()
second.call("POST",
f"/adventures/{adventure_id}/checkpoints/{save_point['id']}/restore",
expect=200)
branches_before = second.call("GET", f"/adventures/{adventure_id}/branches",
expect=200)
assert len(branches_before) == 1, "restore must not fork"
second.play(adventure_id, "go around the back")
branches_after = second.call("GET", f"/adventures/{adventure_id}/branches",
expect=200)
assert len(branches_after) == 2, "the write should have forked"
# Nothing was deleted to achieve it: the tree grew by the new turn only.
assert second.total_rows(adventure_id) == at_tip["rows"] + 2
# Ordinary Redo no longer offers the displaced future.
assert second.call("GET", f"/adventures/{adventure_id}", expect=200)["can_redo"] is False
# And the Save Point still names the position it always did.
listed = second.call("GET", f"/adventures/{adventure_id}/checkpoints", expect=200)
assert (listed[0]["branch_id"], listed[0]["depth"]) == (
save_point["branch_id"], save_point["depth"]
)