A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
381 lines
14 KiB
Python
381 lines
14 KiB
Python
"""End-to-end HTTP tests for retry history: every attempt at an AI turn is kept
|
|
as a variant of the same action, browsable and (for the last message)
|
|
switchable, restoring the world/script state that attempt produced.
|
|
|
|
python -m pytest tests/test_retry_variants.py -v
|
|
"""
|
|
import pytest
|
|
from fastapi import Depends
|
|
from fastapi.testclient import TestClient
|
|
|
|
from app import auth, limits, models
|
|
from app.database import Base, SessionLocal, engine, get_db
|
|
from app.main import app
|
|
from app.providers import ProviderError
|
|
from app.routers import adventures
|
|
|
|
from fakes import GOLD_SCHEMA, ScriptedProvider, gold_replies, gold_reply, tally_of, tally_reply
|
|
|
|
SCHEMA = GOLD_SCHEMA
|
|
|
|
# Each turn spends 10 gold, so a double-applied or un-rolled-back attempt shows.
|
|
|
|
|
|
@pytest.fixture()
|
|
def client(monkeypatch):
|
|
Base.metadata.create_all(bind=engine)
|
|
setup = SessionLocal()
|
|
user = models.User(is_guest=False, email="variants@example.com")
|
|
setup.add(user)
|
|
setup.flush()
|
|
setup.add(models.Settings(user_id=user.id, api_key="enc:dummy", model="test-model"))
|
|
scenario = models.Scenario(user_id=user.id, title="S", stat_schema=SCHEMA)
|
|
setup.add(scenario)
|
|
setup.flush()
|
|
adv = models.Adventure(
|
|
user_id=user.id, title="Cave", scenario_id=scenario.id,
|
|
world_state={"player": {"hp": 100, "gold": 0}},
|
|
)
|
|
setup.add(adv)
|
|
setup.flush()
|
|
setup.add(models.Action(adventure_id=adv.id, type="start", text="You enter a cave."))
|
|
setup.commit()
|
|
adv_id, user_id = adv.id, user.id
|
|
setup.close()
|
|
|
|
ScriptedProvider.replies = ["Attempt one."]
|
|
ScriptedProvider.calls = 0
|
|
ScriptedProvider.prompts = []
|
|
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
|
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
|
|
|
def _current_user(db=Depends(get_db)):
|
|
return db.get(models.User, user_id)
|
|
|
|
app.dependency_overrides[auth.get_current_user] = _current_user
|
|
c = TestClient(app)
|
|
c.adv_id = adv_id
|
|
try:
|
|
yield c
|
|
finally:
|
|
app.dependency_overrides.clear()
|
|
adventures.turns._active_turns.clear()
|
|
Base.metadata.drop_all(bind=engine)
|
|
|
|
|
|
def _adv(adv_id):
|
|
"""The instrument, and the whole document behind it.
|
|
|
|
M5 moved the instrument from an RPG stat to a typed narrative fact; the
|
|
tuple shape is kept so the call sites read the same. `[0]["gold"]` is the
|
|
tally, and `[1]` is the authoritative state document.
|
|
"""
|
|
db = SessionLocal()
|
|
try:
|
|
adv = db.get(models.Adventure, adv_id)
|
|
state = adv.narrative_state or {}
|
|
return {"gold": tally_of(state)}, state
|
|
finally:
|
|
db.close()
|
|
|
|
|
|
def _actions(client):
|
|
return client.get(f"/api/adventures/{client.adv_id}").json()["actions"]
|
|
|
|
|
|
def _play(client, text="look around"):
|
|
r = client.post(f"/api/adventures/{client.adv_id}/actions",
|
|
json={"type": "do", "text": text})
|
|
assert r.status_code == 200, r.text
|
|
|
|
|
|
def _retry(client):
|
|
r = client.post(f"/api/adventures/{client.adv_id}/retry")
|
|
assert r.status_code == 200, r.text
|
|
return r
|
|
|
|
|
|
# ---------------------------------------------------------------- keeping them
|
|
|
|
def test_retry_keeps_the_discarded_attempt(client):
|
|
ScriptedProvider.replies = ["Attempt one.", "Attempt two."]
|
|
_play(client)
|
|
assert _actions(client)[-1]["text"] == "Attempt one."
|
|
|
|
_retry(client)
|
|
actions = _actions(client)
|
|
# One AI action still, not two. The retry replaced the live text in place.
|
|
assert [a["type"] for a in actions] == ["start", "do", "ai"]
|
|
last = actions[-1]
|
|
assert last["text"] == "Attempt two."
|
|
assert last["take_count"] == 2
|
|
assert last["take_index"] == 1
|
|
|
|
r = client.get(f"/api/adventures/{client.adv_id}/actions/{last['id']}/variants")
|
|
assert r.status_code == 200, r.text
|
|
assert [v["text"] for v in r.json()] == ["Attempt one.", "Attempt two."]
|
|
assert [v["active"] for v in r.json()] == [False, True]
|
|
|
|
|
|
def test_retry_context_excludes_the_attempt_being_replaced(client):
|
|
"""A retry produces a fresh take on the same turn. The row survives the
|
|
retry because it holds the variant history, so it is still in
|
|
`adventure.actions` while the replacement context is assembled. The
|
|
context builder must filter it out, or the model continues past the
|
|
attempt it is replacing and writes a sequel that blends both."""
|
|
ScriptedProvider.replies = ["Attempt one.", "Attempt two."]
|
|
_play(client)
|
|
_retry(client)
|
|
|
|
retry_story = ScriptedProvider.prompts[-1][1]
|
|
assert "Attempt one." not in retry_story
|
|
# The turn's own player action must still be there. It is what the model responds to.
|
|
assert "look around" in retry_story
|
|
assert "You enter a cave." in retry_story
|
|
|
|
|
|
def test_retry_context_keeps_earlier_ai_turns(client):
|
|
"""Only the action being retried is dropped, not AI history in general."""
|
|
ScriptedProvider.replies = ["First turn.", "Second turn.", "Second, again."]
|
|
_play(client, "go north")
|
|
_play(client, "go south")
|
|
_retry(client)
|
|
|
|
retry_story = ScriptedProvider.prompts[-1][1]
|
|
assert "First turn." in retry_story
|
|
assert "Second turn." not in retry_story
|
|
|
|
|
|
def test_never_retried_action_has_no_variants(client):
|
|
_play(client)
|
|
last = _actions(client)[-1]
|
|
assert last["take_count"] == 1 # itself, and nothing to page to
|
|
assert client.get(
|
|
f"/api/adventures/{client.adv_id}/actions/{last['id']}/variants").json() == []
|
|
|
|
|
|
def test_three_attempts_all_kept_in_order(client):
|
|
ScriptedProvider.replies = ["One.", "Two.", "Three."]
|
|
_play(client)
|
|
_retry(client)
|
|
_retry(client)
|
|
last = _actions(client)[-1]
|
|
assert last["take_count"] == 3
|
|
assert last["take_index"] == 2
|
|
variants = client.get(
|
|
f"/api/adventures/{client.adv_id}/actions/{last['id']}/variants").json()
|
|
assert [v["text"] for v in variants] == ["One.", "Two.", "Three."]
|
|
|
|
|
|
# ---------------------------------------------------------------- switching
|
|
|
|
def test_switching_back_restores_that_attempt_state(client):
|
|
"""Each attempt records its own total, so switching between them shows
|
|
whether the state followed the narration back."""
|
|
ScriptedProvider.replies = [
|
|
tally_reply("You take a scratch.", 95),
|
|
tally_reply("You take a beating.", 60),
|
|
]
|
|
_play(client)
|
|
assert _adv(client.adv_id)[0]["gold"] == 95
|
|
_retry(client)
|
|
player, _document = _adv(client.adv_id)
|
|
assert player["gold"] == 60, "the retake's own state, not the one it replaced"
|
|
|
|
last = _actions(client)[-1]
|
|
r = client.post(
|
|
f"/api/adventures/{client.adv_id}/actions/{last['id']}/variant", json={"index": 0})
|
|
assert r.status_code == 200, r.text
|
|
assert r.json()["text"].startswith("You take a scratch")
|
|
assert r.json()["take_index"] == 0
|
|
# The state follows the narration back.
|
|
assert _adv(client.adv_id)[0]["gold"] == 95
|
|
|
|
# And forward again.
|
|
client.post(f"/api/adventures/{client.adv_id}/actions/{last['id']}/variant",
|
|
json={"index": 1})
|
|
|
|
|
|
def test_switching_updates_the_state_summary_chips(client):
|
|
"""The chip under a message describes the take on screen.
|
|
|
|
M5 replaced the RPG world-change chips with the narrative-state summary;
|
|
what this test guards is unchanged — switching takes must change what the
|
|
chip says, or the reader is shown one attempt's prose beside another's
|
|
consequences.
|
|
"""
|
|
ScriptedProvider.replies = [
|
|
tally_reply("A scratch.", 5),
|
|
tally_reply("A beating.", 40),
|
|
]
|
|
_play(client)
|
|
_retry(client)
|
|
last = _actions(client)[-1]
|
|
assert any("40" in line for line in last["state_summary"]), last["state_summary"]
|
|
|
|
client.post(f"/api/adventures/{client.adv_id}/actions/{last['id']}/variant",
|
|
json={"index": 0})
|
|
switched = _actions(client)[-1]["state_summary"]
|
|
assert any("5" in line for line in switched), switched
|
|
assert not any("40" in line for line in switched)
|
|
|
|
|
|
def test_cannot_switch_a_turn_the_story_moved_past(client):
|
|
ScriptedProvider.replies = ["One.", "Two.", "Three."]
|
|
_play(client)
|
|
_retry(client)
|
|
retried = _actions(client)[-1]
|
|
_play(client) # story continues from "Two."
|
|
|
|
r = client.post(
|
|
f"/api/adventures/{client.adv_id}/actions/{retried['id']}/variant", json={"index": 0})
|
|
assert r.status_code == 400
|
|
assert "latest message" in r.json()["detail"]
|
|
# The variant is still readable. Keeping every attempt browsable is why it still exists.
|
|
variants = client.get(
|
|
f"/api/adventures/{client.adv_id}/actions/{retried['id']}/variants").json()
|
|
assert [v["text"] for v in variants] == ["One.", "Two."]
|
|
|
|
|
|
def test_switching_to_a_missing_index_is_rejected(client):
|
|
ScriptedProvider.replies = ["One.", "Two."]
|
|
_play(client)
|
|
_retry(client)
|
|
last = _actions(client)[-1]
|
|
r = client.post(
|
|
f"/api/adventures/{client.adv_id}/actions/{last['id']}/variant", json={"index": 7})
|
|
assert r.status_code == 400
|
|
|
|
|
|
# ---------------------------------------------------------------- edge cases
|
|
|
|
def test_failed_retry_leaves_the_previous_attempt_in_charge(client):
|
|
"""A provider error mid-retry must undo the rollback, or the stats on
|
|
screen would silently disagree with the text still shown."""
|
|
ScriptedProvider.replies = [gold_reply("Attempt one."), ProviderError("upstream is down")]
|
|
_play(client)
|
|
assert _adv(client.adv_id)[0]["gold"] == 10
|
|
|
|
client.post(f"/api/adventures/{client.adv_id}/retry")
|
|
|
|
actions = _actions(client)
|
|
assert actions[-1]["text"] == "Attempt one." # text never lost
|
|
assert _adv(client.adv_id)[0]["gold"] == 10 # and the state still matches it
|
|
|
|
|
|
def test_undo_removes_the_action_and_its_history(client):
|
|
ScriptedProvider.replies = [gold_reply(t) for t in ["One.", "Two."]]
|
|
_play(client)
|
|
_retry(client)
|
|
r = client.post(f"/api/adventures/{client.adv_id}/undo")
|
|
assert r.status_code == 200, r.text
|
|
# Undo returns the newest window now, not the whole story.
|
|
assert [a["type"] for a in r.json()["actions"]] == ["start"]
|
|
assert _adv(client.adv_id)[0]["gold"] == 0
|
|
|
|
|
|
def test_editing_a_narrator_take_adds_a_take_and_keeps_the_original(client):
|
|
"""M5 corrective pass: a narrator correction is a new take, not a rewrite.
|
|
|
|
§15.5 requires the original narration to be retained. At the tip that means
|
|
the correction joins the turn's attempts rather than replacing the words of
|
|
one, so the pager still reaches what was there before.
|
|
"""
|
|
ScriptedProvider.replies = ["One.", "Two."]
|
|
_play(client)
|
|
_retry(client)
|
|
last = _actions(client)[-1]
|
|
client.patch(f"/api/adventures/{client.adv_id}/actions/{last['id']}",
|
|
json={"text": "Two, but better."})
|
|
|
|
live = _actions(client)[-1]
|
|
assert live["text"] == "Two, but better."
|
|
assert live["take_count"] == 3, "the correction is a third attempt"
|
|
assert live["take_index"] == 2
|
|
|
|
# The original is still reachable, and the edit survives paging away.
|
|
client.post(f"/api/adventures/{client.adv_id}/actions/{live['id']}/variant",
|
|
json={"index": 1})
|
|
assert _actions(client)[-1]["text"] == "Two."
|
|
client.post(f"/api/adventures/{client.adv_id}/actions/{live['id']}/variant",
|
|
json={"index": 2})
|
|
assert _actions(client)[-1]["text"] == "Two, but better."
|
|
|
|
|
|
def _live_depth(client) -> int:
|
|
"""Returns the depth of the newest action the story tells."""
|
|
db = SessionLocal()
|
|
try:
|
|
return (
|
|
db.query(models.Action.depth)
|
|
.filter(models.Action.adventure_id == client.adv_id,
|
|
models.Action.live.is_(True))
|
|
.order_by(models.Action.depth.desc())
|
|
.first()[0]
|
|
)
|
|
finally:
|
|
db.close()
|
|
|
|
|
|
def test_retry_keeps_the_turn_depth(client):
|
|
"""A retry re-runs the same turn, so the world-state clock (which drives
|
|
cooldowns) must not advance.
|
|
|
|
The depth is read from the row rather than the payload. It is a coordinate
|
|
on one branch, and the client pages by id, so it is not exposed.
|
|
"""
|
|
ScriptedProvider.replies = ["One.", "Two."]
|
|
_play(client)
|
|
before = _live_depth(client)
|
|
_retry(client)
|
|
assert _live_depth(client) == before
|
|
|
|
|
|
def test_export_and_import_round_trips_variants(client):
|
|
ScriptedProvider.replies = ["One.", "Two."]
|
|
_play(client)
|
|
_retry(client)
|
|
bundle = client.get(f"/api/adventures/{client.adv_id}/export").json()
|
|
# SP6: the attempts are nodes in the bundle too, sharing one coordinate,
|
|
# and `live` says which of them the story tells. The `variants` array
|
|
# survives only in the v1 reader. See the hand-edited bundle below.
|
|
ai = [a for a in bundle["actions"] if a["type"] == "ai"]
|
|
assert [(a["text"], a["live"]) for a in ai] == [("One.", False), ("Two.", True)]
|
|
assert len({(a["branch"], a["depth"]) for a in ai}) == 1
|
|
|
|
r = client.post("/api/adventures/import", json=bundle)
|
|
assert r.status_code == 201, r.text
|
|
imported = r.json()["id"]
|
|
actions = client.get(f"/api/adventures/{imported}").json()["actions"]
|
|
assert [a["type"] for a in actions] == ["start", "do", "ai"]
|
|
assert actions[-1]["text"] == "Two."
|
|
|
|
# The pager reads 2/2 on the copy, as it does on the original.
|
|
#
|
|
# It read 1/1 until M9, and this test recorded that as a gap in the import
|
|
# rather than in the export: the attempts were both there at one coordinate
|
|
# and `GET .../variants` listed them, but the import wrote no `parent_id`,
|
|
# so `annotate_takes` grouped on the coordinate instead. That is right for a
|
|
# plain retry and wrong the moment two takes of one turn each have takes of
|
|
# their own beneath them, which is why M9 carried the parentage rather than
|
|
# leaving the pager to a fallback. See `bundle._link_take_parents`.
|
|
assert actions[-1]["take_count"] == 2
|
|
assert actions[-1]["take_index"] == 1
|
|
variants = client.get(
|
|
f"/api/adventures/{imported}/actions/{actions[-1]['id']}/variants").json()
|
|
assert [v["text"] for v in variants] == ["One.", "Two."]
|
|
|
|
|
|
def test_import_clamps_an_out_of_range_variant_index(client):
|
|
bundle = {
|
|
"format": "ai-dnd-adventure-v1", "title": "Hand-edited",
|
|
"actions": [{
|
|
"index": 0, "type": "ai", "text": "Only.",
|
|
"variants": [{"text": "Only."}], "variantIndex": 9,
|
|
}],
|
|
}
|
|
r = client.post("/api/adventures/import", json=bundle)
|
|
assert r.status_code == 201, r.text
|
|
actions = client.get(f"/api/adventures/{r.json()['id']}").json()["actions"]
|
|
assert actions[0]["take_index"] == 0
|