Rewrite Python comments in Google developer documentation style (#12)

* Rewrite comments in Google developer documentation style

Rewrite the comments and docstrings across the backend core modules so they
read plainly. The previous prose was accurate but dense and figurative, which
made it slow to skim.

Applies the Google developer documentation style guide: short sentences, active
voice, present tense, American spelling, and no metaphors, idioms, or
rhetorical asides. Replaces em-dash chains with separate sentences.
This commit is contained in:
Parth
2026-08-26 15:37:25 +05:30
committed by GitHub
parent cf6161a5ee
commit e7d75c3b05
83 changed files with 4605 additions and 3988 deletions
+22 -22
View File
@@ -1,35 +1,35 @@
"""Make a database look like an older schema version, so a migration can run.
`create_all` always builds the *current* schema. A test that wants to watch a
migration happen therefore has to take the newer columns back off before it
stamps an older version — otherwise the migration meets a table that already
has its column and dies on a duplicate.
`create_all` always builds the current schema. A test that wants to watch a
migration run must remove the newer columns first and then stamp an older
version. Otherwise the migration finds a column that already exists and
fails on a duplicate.
Rewinding the stamp alone was enough for a while, which is why two test files
did exactly that. It stopped being enough the moment another `ADD COLUMN`
landed after theirs: the replay then runs migrations they never meant to
exercise, against columns `create_all` had already made. This module is that
rewind done properly, in one place, so appending a migration means adding its
inverse here rather than discovering three unrelated test failures.
Rewinding the stamp alone worked for a while, so two test files did exactly
that. It stopped working when another `ADD COLUMN` migration landed. Without
this rewind, the replay runs migrations the tests never intended to
exercise, against columns `create_all` already added. This module rewinds
properly in one place. Adding a migration now means adding its inverse here,
instead of tracking down failures in three unrelated test files.
SQLite only — every test that replays migrations runs on a temp file, and
`PRAGMA user_version` is where the stamp lives there. Migrations that change a
column's *type* (43–45, JSON to compressed bytes) have no clean inverse and are
not listed: they get replayed as-is, which is what the tests using them already
relied on.
This module supports SQLite only. Every test that replays migrations runs on
a temp file, and SQLite stores the stamp in `PRAGMA user_version`.
Migrations that change a column's type (43-45, JSON to compressed bytes)
have no clean inverse, so this list omits them. Those migrations replay
as-is, which is what the tests using them already expect.
"""
from sqlalchemy import text
from sqlalchemy.engine import Engine
# (version that added it, statements that take it back off), newest first.
# Each entry is (version that added the column, statements that remove it).
# The list is ordered newest first.
#
# Phase 14's `branch_id` columns are deliberately absent: SQLite refuses to drop
# a column a foreign key names ("unknown column in foreign key definition"), so
# a current-schema database cannot be rewound past them at all. That is what
# `migrations._column_already_there` is for — the replay skips DDL that has
# already happened, so the tree migrations run their backfill against a schema
# that already has the columns, which is exactly the situation here.
# Phase 14's `branch_id` columns are missing on purpose. SQLite refuses to
# drop a column that a foreign key references, so a current-schema database
# cannot be rewound past them. `migrations._column_already_there` handles
# this case. It skips DDL that already ran, so the tree migrations run their
# backfill against a schema that already has the columns.
_UNDO: list[tuple[int, tuple[str, ...]]] = [
# Packed float32 vectors and the flag beside them.
(39, ("ALTER TABLE memories DROP COLUMN embedded",)),
+19 -17
View File
@@ -1,11 +1,11 @@
"""The access log — app/accesslog.py and GET /api/analytics/access.
"""The access log: app/accesslog.py and GET /api/analytics/access.
This is the half of the analytics work that identifies people on purpose, so
the things worth pinning are the ones that would quietly make it wrong: that
the address recorded is the hardened one and not a header a client chose, that
session rows are thinned instead of written per page load, and that a row
outlives the account it describes — guest cleanup runs on a schedule, and a log
that deletes itself is not a log.
This is the half of the analytics work that identifies people on purpose,
so these tests pin the details that would quietly make it wrong. The
address recorded must be the hardened one, not a header a client chose.
Session rows must be thinned instead of written on every page load. And a
row must outlive the account it describes, because guest cleanup deletes
accounts on a schedule, and a log that deletes itself is not a log.
python -m pytest tests/test_accesslog.py -v
"""
@@ -57,7 +57,7 @@ def client(monkeypatch):
monkeypatch.setattr(auth, "ANALYTICS_EMAILS", {"owner@example.com"})
# /auth/me resolves its own session, so the cookie flow below is the real
# one; every other endpoint goes through get_current_user, and `act_as`
# one. Every other endpoint goes through get_current_user, and `act_as`
# decides who that is.
acting = {"id": ids["owner"]}
@@ -111,9 +111,9 @@ def test_a_new_session_is_logged(client):
def test_the_address_is_the_hardened_one_not_the_clients(client):
visit(client)
# The client prepended its own value; only the hop the edge appended counts.
# Recording the leftmost would make every row forgeable, which for a log is
# worse than having no log.
# The client prepended its own value. Only the hop the edge appended counts.
# Recording the leftmost value would make every row forgeable, which is
# worse for a log than having no log at all.
assert rows()[0].ip == EDGE
@@ -141,15 +141,16 @@ def test_sign_in_and_failure_are_both_logged(client):
assert accesslog.LOGIN_FAILED in kinds and accesslog.LOGIN in kinds
failure = rows(accesslog.LOGIN_FAILED)[0]
# The address tried, not the account that owns it: a run against an address
# with no account behind it is exactly what this row is for.
# This records the address that was tried, not the account it belongs to.
# A failed attempt against an address with no matching account is
# exactly what this row exists to capture.
assert failure.who == "player@example.com"
assert failure.user_id is None
assert rows(accesslog.LOGIN)[0].user_id == client.ids["member"]
def test_registering_is_logged_against_the_upgraded_account(client):
visit(client) # mints the guest whose session registers
visit(client) # creates the guest whose session then registers
client.act_as(rows()[0].user_id)
client.post("/api/auth/register", json={"email": "new@example.com", "password": "hunter2long"})
entry = rows(accesslog.REGISTER)[0]
@@ -165,8 +166,9 @@ def test_a_row_outlives_the_account_it_describes(client):
db.commit()
finally:
db.close()
# No foreign key, and `who` is a snapshot — guest cleanup deletes accounts
# on a schedule, and a log that vanishes with them is not a log.
# There is no foreign key, and `who` is a snapshot. Guest cleanup deletes
# accounts on a schedule, and a log that vanishes along with them is not
# a log.
survivor = rows()[0]
assert survivor.who == entry.who and survivor.ip == EDGE
@@ -178,7 +180,7 @@ def test_a_long_user_agent_is_truncated(client):
def test_a_logging_failure_does_not_break_the_request(client, monkeypatch):
monkeypatch.setattr(accesslog, "_client_ip", lambda request: 1 / 0)
# The log watches sign-in; it must not be able to stand in its way.
# The log observes sign-in. A logging failure must not block the request.
assert visit(client).status_code == 200
+21 -18
View File
@@ -1,12 +1,13 @@
"""Opening an adventure fetches a window, not the whole story.
A story only ever gets longer. Production's longest is 607 actions and 589.5 kB
in one response, and nothing about that curve bends on its own — so the page
load returns the newest ACTION_PAGE and the reader pages upward.
A story only ever gets longer. Production's longest is 607 actions and
589.5 kB in one response, and that number never decreases on its own. The
page load returns the newest `ACTION_PAGE` window, and the reader pages
upward from there.
The paging anchors on an action id rather than an offset, and these tests are
mostly about why. An offset counted back from the newest shifts every older
position the moment a turn lands, which is precisely when a reader is likely
The paging anchors on an action id rather than an offset, and these tests
cover why. An offset counted back from the newest shifts every older
position the moment a turn lands, which is exactly when a reader is likely
to be scrolling. An anchor means the same thing before and after.
python -m pytest tests/test_action_paging.py -v
@@ -104,7 +105,7 @@ def test_the_page_load_returns_only_the_newest_window(client):
body = r.json()
assert len(body["actions"]) == ACTION_PAGE
assert body["action_count"] == TOTAL
# ...and it is the *newest* window, ending on the last action.
# It is the newest window, ending on the last action.
assert body["actions"][-1]["index"] == TOTAL - 1
assert body["actions"][0]["index"] == TOTAL - ACTION_PAGE
@@ -123,8 +124,8 @@ def test_a_short_story_is_returned_whole(client):
def test_the_page_load_does_not_grow_with_the_story(client):
"""The point of the change. Whatever the story's length, opening it costs
a window."""
"""Confirm that opening a story costs a window, regardless of the
story's length."""
meter = dbmeter.Meter()
meter.attach(engine)
try:
@@ -134,8 +135,9 @@ def test_the_page_load_does_not_grow_with_the_story(client):
finally:
meter.detach()
# Each action carries ~1 kB of text and there are 187 of them; a window is
# 60. Generous ceiling, but far below the whole story.
# Each action carries about 1 KB of text, and there are 187 of them. A
# window holds 60 actions. This ceiling is generous but still far below
# the size of the whole story.
assert windowed < ACTION_PAGE * 2_000, f"{windowed:,} B for one window"
assert windowed < TOTAL * 500, (
f"{windowed:,} B — that is the whole story, not a window"
@@ -182,9 +184,10 @@ def test_each_page_is_ordered_oldest_first(client):
# ------------------------------------------------- the reason for the anchor
def test_a_turn_arriving_mid_scroll_does_not_shift_the_next_page(client):
"""The failure an offset would have. Read the newest page, let a turn land,
then page up: the reader must get exactly what precedes what they hold —
no duplicate, no skipped action."""
"""Reproduce the failure an offset-based scheme would have. Read the
newest page, let a turn land, then page up. The reader must get exactly
what precedes the actions they already hold, with no duplicate and no
skipped action."""
first = page(client)
oldest_held = first["actions"][0]
@@ -194,13 +197,13 @@ def test_a_turn_arriving_mid_scroll_does_not_shift_the_next_page(client):
assert older["actions"][-1]["index"] == oldest_held["index"] - 1, \
"the page shifted when a turn landed"
assert all(a["index"] < oldest_held["index"] for a in older["actions"])
# The new turn moved the total, which is fine — it must not move the window.
# The new turn changes the total, which is expected. It must not move the window.
assert older["total"] == TOTAL + 1
def test_a_deleted_anchor_reports_the_end_rather_than_a_duplicate_page(client):
"""Undo can remove the action a slow scroll was anchored to. Better to stop
than to hand back a page the reader already has."""
"""Undo can remove the action a slow scroll was anchored to. The endpoint
must stop instead of returning a page the reader already has."""
body = page(client)
anchor = body["actions"][0]
@@ -220,7 +223,7 @@ def test_a_deleted_anchor_reports_the_end_rather_than_a_duplicate_page(client):
def test_limit_is_honoured_and_capped(client):
assert len(page(client, limit=5)["actions"]) == 5
# A client asking for the whole story does not get to undo the paging.
# A client that asks for the whole story cannot bypass the paging cap.
assert len(page(client, limit=100_000)["actions"]) <= ACTION_PAGE * 4
+19 -17
View File
@@ -1,11 +1,11 @@
"""Visit analytics — app/analytics.py and the two endpoints in front of it.
"""Visit analytics: app/analytics.py and the two endpoints in front of it.
Three things are worth testing here and the rest is arithmetic. That the
counters survive the buffer/UPSERT round trip (a flush must add to what is
already stored, not replace it, or every number is only ever the last minute).
That the funnel counts *people* rather than clicks, which is the only reason
the visitor-day table exists. And that the gate holds: a stranger cannot read
the dashboard, and cannot inflate what it says beyond hitting the page.
This file tests three things, and the rest is arithmetic. The counters must
survive the buffer/UPSERT round trip: a flush adds to what is already
stored instead of replacing it, or every number would show only the last
minute. The funnel counts people rather than clicks, which is the only
reason the visitor-day table exists. The gate holds: a stranger cannot read
the dashboard, and cannot inflate what it reports beyond hitting the page.
python -m pytest tests/test_analytics.py -v
"""
@@ -102,7 +102,7 @@ def test_label_cardinality_is_capped(db):
analytics.flush(db)
labels = db.query(models.AnalyticsDaily).filter_by(metric=analytics.M_REFERRER).count()
# Everything past the cap is folded into one bucket, so a referrer flood
# cannot mint rows without limit.
# cannot create unlimited rows.
assert labels == analytics.MAX_LABELS_PER_METRIC + 1
assert counter(db, analytics.M_REFERRER, analytics.OTHER) == 25
@@ -115,8 +115,9 @@ def test_visitor_id_is_stable_and_keyed(db, monkeypatch):
assert handle == analytics.visitor_id(user) # a returning visitor
assert handle != analytics.visitor_id(make_user(db)) # is still one visitor
assert len(handle) == 32 and int(handle, 16) >= 0 # opaque hex, not an id
# Keyed on the app secret, not a bare hash of the user id: otherwise anyone
# holding this table could rebuild the mapping by hashing 1, 2, 3, …
# Keyed on the app secret, not a bare hash of the user id. Otherwise
# anyone holding this table could rebuild the mapping by hashing
# sequential ids.
monkeypatch.setattr(analytics.security, "SECRET_KEY", b"a-different-secret")
assert analytics.visitor_id(user) != handle
@@ -128,8 +129,8 @@ def test_a_repeat_visitor_is_new_only_once(db):
rows = db.query(models.AnalyticsVisitorDay).all()
assert len(rows) == 1 and rows[0].is_new
# Same visitor, a later day: seen before, so not new — and not merged into
# the first day's row either.
# Same visitor, a later day: seen before, so not new. The row is not
# merged into the first day's row either.
tomorrow = (models.utcnow().date() + timedelta(days=1)).isoformat()
analytics._visits[(tomorrow, analytics.visitor_id(user))] = set()
analytics.flush(db)
@@ -234,7 +235,7 @@ def test_summary_counts_people_once_per_step(db):
steps = {row["step"]: row["count"] for row in result["funnel"]}
assert steps["Visited"] == 2
assert steps["Opened a scenario"] == 2
assert steps["Played a turn"] == 1 # not 3 — one person, three turns
assert steps["Played a turn"] == 1 # not 3, because one person made three turns
assert steps["Signed up"] == 0
# Raw event totals still count every occurrence.
assert result["totals"]["turns"] == 3
@@ -361,10 +362,11 @@ def test_api_errors_are_counted_by_route(client):
# ---------- The dialect the tests never run on ----------
def test_the_upserts_compile_for_postgres():
"""Prod is Neon; these tests are SQLite, and a failed flush is caught and
logged rather than raised. A dialect mistake would therefore be invisible
until the dashboard quietly stayed empty — so compile both statements
against Postgres without ever connecting to one.
"""Prod runs on Neon, but these tests run on SQLite, and a failed flush
is caught and logged instead of raised. A dialect mistake would
therefore stay invisible until the dashboard quietly stayed empty. This
test compiles both statements against Postgres without connecting to
one.
"""
from sqlalchemy import create_engine
from sqlalchemy.dialects import postgresql
+45 -38
View File
@@ -1,10 +1,10 @@
"""Phase 14 SP4 — a retry writes a sibling node instead of rewriting a row.
"""Phase 14 SP4: a retry writes a sibling node instead of rewriting a row.
`test_retry_variants.py` is the behavioural contract, unchanged since before
the tree, and it still passes: the same URLs, the same payload shape, the same
outcomes. This file asserts the things that are *only* true of the new storage
— that a turn can be several rows, that exactly one of them is the story, and
that the arrangement costs neither an extra prompt nor an extra turn.
`test_retry_variants.py` is the behavioral contract from before the tree
existed, and it still passes unchanged: the same URLs, the same payload
shape, the same outcomes. This file asserts the things that are true only of
the new storage. A turn can be several rows, exactly one of them is the
story, and the arrangement costs neither an extra prompt nor an extra turn.
python -m pytest tests/test_attempt_siblings.py -v
"""
@@ -122,8 +122,9 @@ def _page(client) -> dict:
def _rows(adv_id) -> list[models.Action]:
"""Every action row of the adventure, story or not, live or not.
Undeferred, because the session is closed before the caller looks: the
columns this file is about are exactly the ones a page load never loads.
The query undefers these columns because the session closes before the
caller reads the result. This file specifically tests the columns that a
page load never loads.
"""
db = SessionLocal()
try:
@@ -156,7 +157,7 @@ def test_a_retry_writes_a_second_row_at_the_same_coordinate(client):
assert [a.text for a in ai] == ["Attempt one.", "Attempt two."]
# Exactly one of them is the story, and it is the newer take.
assert [a.live for a in ai] == [False, True]
# ...and the discarded attempt is untouched, not a copy of anything.
# The discarded attempt is untouched, not a copy of anything.
assert ai[0].state_after is not None
@@ -173,9 +174,10 @@ def test_the_story_shows_and_counts_the_turn_once(client):
def test_a_discarded_attempt_never_reaches_the_prompt(client):
"""The trap the branch clause exists to close, at sibling scale: the losing
attempt sits at the same branch and depth as the live one, so anything
reading the story by coordinate alone would replay both."""
"""This is the failure case the branch clause exists to prevent, at
sibling scale. The losing attempt sits at the same branch and depth as
the live one, so anything reading the story by coordinate alone would
replay both."""
ScriptedProvider.replies = ["Attempt one.", "Attempt two.", "Next turn."]
_play(client)
_retry(client)
@@ -198,21 +200,22 @@ def test_switching_moves_the_story_onto_the_other_row(client):
r = client.post(
f"/api/adventures/{client.adv_id}/actions/{newest_id}/variant", json={"index": 0})
assert r.status_code == 200, r.text
# A different row answers — that is the whole change.
# A different row answers the request. That is the only change.
assert r.json()["id"] != newest_id
assert r.json()["text"].startswith("A scratch")
rows = _rows(client.adv_id)
ai = [a for a in rows if a.type == "ai"]
assert [a.live for a in ai] == [True, False]
# Both takes are still there, byte for byte.
# Both takes remain unchanged in the database.
assert [a.text.split(".")[0] for a in ai] == ["A scratch", "A beating"]
def test_the_assembled_prompt_is_stored_once_per_turn(client):
"""A snapshot is ~160 kB of prompt every attempt at a turn shares. Giving
each sibling a copy would have made retry a permanent multiplier on the
biggest column in the database, so the prompt moves with the live flag."""
"""A snapshot holds about 160 kB of prompt that every attempt at a turn
shares. Giving each sibling its own copy would make retry multiply the
size of the largest column in the database. Instead, the prompt moves
with the live flag."""
ScriptedProvider.replies = ["Attempt one.", "Attempt two."]
_play(client)
_retry(client)
@@ -278,12 +281,12 @@ def test_deleting_a_turn_through_a_discarded_attempt_still_takes_the_turn(client
def test_retrying_withdraws_the_memory_the_turn_produced(client):
"""Why summarization no longer holds the newest action back.
A memory covering the newest turn used to be unreachable-by-construction:
the summarizer stopped one action short, because a retry rewrote the row
under a mark that had already moved past it. Now the mark and the memory
both name the node, and replacing what a node says withdraws them — the
same repair undo and delete already made, so the holdback was the only
thing left that a retry needed.
A memory covering the newest turn used to be unreachable by
construction. The summarizer stopped one action short, because a retry
rewrote the row under a mark that had already moved past it. Now the
mark and the memory both name the node, so replacing what a node says
withdraws them. Undo and delete already had this repair; the holdback
was the only gap a retry still needed to close.
"""
ScriptedProvider.replies = ["One.", "Two."]
_play(client)
@@ -311,8 +314,8 @@ def test_retrying_withdraws_the_memory_the_turn_produced(client):
try:
adventure = db.get(models.Adventure, client.adv_id)
assert db.query(models.Memory).count() == 0, "the withdrawn memory is gone"
# ...and the ground it covered is handed back, so the block is summarized
# again from where it began rather than silently skipped.
# The depth range it covered is released, so the block is
# summarized again from where it began instead of being skipped.
assert cursors.MEMORY.depth(db, adventure) == 0
assert cursors.SUMMARY.depth(db, adventure) == 0
assert covered_depth > 0
@@ -382,12 +385,14 @@ def test_attempts_module_agrees_with_the_endpoint(client):
def test_export_carries_every_attempt_as_its_own_node(client):
"""SP6 changed the answer here, and the reason is the whole of that subphase.
"""SP6 changed the answer here, for the same reasons as the rest of that
subphase.
A v1 bundle had one entry per turn and folded the group back into a
`variants` array, because the format had nowhere else to put a second take.
A v2 bundle has coordinates, so an attempt is a node in the file exactly as
it is a node in the database, and `live` says which one is the story.
`variants` array, because the format had nowhere else to put a second
take. A v2 bundle has coordinates, so an attempt is a node in the file
exactly as it is a node in the database, and `live` says which one is
the story.
"""
ScriptedProvider.replies = ["One.", "Two."]
_play(client)
@@ -399,7 +404,7 @@ def test_export_carries_every_attempt_as_its_own_node(client):
assert len({(a["branch"], a["depth"]) for a in ai}) == 1, "one turn, two takes"
assert "variants" not in ai[0], "nothing writes the repeating group any more"
# ...and importing it puts the group back exactly as it stood.
# Importing the bundle puts the group back exactly as it stood.
imported = client.post("/api/adventures/import", json=bundle).json()["id"]
rows = _rows(imported)
ai_rows = [a for a in rows if a.type == "ai"]
@@ -410,11 +415,12 @@ def test_export_carries_every_attempt_as_its_own_node(client):
def test_a_retry_after_switching_back_files_the_new_attempt_last(client):
"""The group stays in the order the attempts were made.
`add_attempt` used to number a new take one past the take it replaced, which
is the end of the group only when the story is standing on the newest one.
Switch a three-take turn back to the first and retry, and the new attempt
collided with take 2 — `renumber` then broke the tie by id and filed it
*between* takes 2 and 3, so the pager walked them in an order nobody played.
`add_attempt` used to number a new take one past the take it replaced.
That numbering is correct only when the story is standing on the newest
take. Switch a three-take turn back to the first and retry, and the new
attempt collides with take 2. `renumber` broke the tie by id and placed
the new attempt between takes 2 and 3, so the pager listed them in an
order the player never produced.
"""
ScriptedProvider.replies = ["One.", "Two.", "Three.", "Four."]
_play(client)
@@ -439,9 +445,10 @@ def test_a_retry_after_switching_back_files_the_new_attempt_last(client):
def test_the_adventure_list_quotes_the_take_the_story_tells(client):
"""The index screen and the story have to agree.
Siblings share a depth and the newest of them has the highest id, so a
snippet ordered by `(depth, id)` alone quotes whichever attempt was written
last — which, after switching back, is the one the player threw away.
Siblings share a depth, and the newest of them has the highest id. A
snippet ordered by `(depth, id)` alone quotes whichever attempt was
written last. After switching back, that attempt is the one the player
discarded.
"""
ScriptedProvider.replies = ["One.", "Two."]
_play(client)
+61 -55
View File
@@ -1,21 +1,23 @@
"""Phase 14 SP2 — a read sees one story, and knows which one.
"""Phase 14 SP2: a read sees one story, and knows which one.
These tests build the fork by hand: three branch rows and their nodes, written
straight to the database, arranged as the design doc's own worked example. That
was the only way to build one when this file was written (nothing forked until
SP5) and it stays that way now that `tree.fork` exists — a fixture that agreed
with the code under test could not catch it being wrong. The two are checked
against each other in `test_branch_forking.py`.
These tests build the fork by hand. Three branch rows and their nodes are
written straight to the database, arranged as the design doc's own worked
example. That was the only way to build one when this file was written,
because nothing forked until SP5. It stays that way now that `tree.fork`
exists, because a fixture built with the same code under test could not
catch that code being wrong. `test_branch_forking.py` checks the two
against each other.
branch C, tip at depth 7, lineage [(C, 7), (B, 5), (A, 3)]
→ A0 A1 A2 A3 B4 B5 C6 C7
-> A0 A1 A2 A3 B4 B5 C6 C7
The point of building it by hand is that every read in the app is supposed to
go through one module, and a forgotten clause does not raise — it quietly shows
a story assembled out of two different ones. So the fixture deliberately leaves
nodes lying where a forgotten clause would pick them up: A kept playing past
the fork (A4, A5), B kept playing past its own (B6), and a second adventure
holds a whole story of its own. None of them may appear on C.
The point of building the fixture by hand is that every read in the app
must go through one module. A forgotten clause does not raise an error. It
silently shows a story assembled out of two different branches. The
fixture deliberately leaves nodes where a forgotten clause would pick them
up: A kept playing past the fork (A4, A5), B kept playing past its own
(B6), and a second adventure holds a whole story of its own. None of these
nodes may appear on C.
python -m pytest tests/test_branch_clause.py -v
"""
@@ -46,8 +48,9 @@ from tools import dbmeter
def make_branch(db, adventure, parent=None, fork_depth=None):
"""A branch row whose lineage is its parent's, capped, plus itself.
The same computation SP5 will do at fork time; written out here so the
fixture cannot pass by agreeing with a bug in the code under test.
This function performs the same computation SP5 does at fork time. It
is written out here so the fixture cannot pass by repeating a bug in
the code under test.
"""
branch = models.Branch(
adventure_id=adventure.id,
@@ -89,7 +92,7 @@ def make_adventure(db, user, title):
@pytest.fixture()
def forked():
"""The worked example, plus everything a forgotten clause would sweep up."""
"""The worked example, plus everything a forgotten clause would expose."""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
user = models.User(is_guest=False, email="branch@example.com")
@@ -158,9 +161,10 @@ def test_the_worked_example_reads_back_as_the_design_doc_says(forked):
def test_a_siblings_nodes_are_invisible(forked):
db, adventure, _ = forked
seen = labels(history.story_actions(adventure))
# A4/A5 are A's own continuation past B's fork; B6 is B's past C's.
# A4 and A5 are A's own continuation past B's fork. B6 is B's own
# continuation past C's fork.
assert "A4" not in seen and "A5" not in seen and "B6" not in seen
# And nothing from the adventure next door.
# Nothing from the other adventure appears either.
assert not [text for text in seen if text.startswith("X")]
@@ -213,11 +217,12 @@ def test_a_window_that_reaches_past_the_fork_still_reads_in_order(forked):
# ------------------------------------------- the loaded-collection short cut
def test_an_already_loaded_collection_is_cut_down_to_the_path(forked):
"""`history._from_memory`'s shortcut, which was the highest-risk line here.
"""Tests `history._from_memory`'s shortcut, the highest-risk line here.
`adventure.actions` is every branch's actions. Slicing it without the path
would assemble a prompt out of two different stories, and nothing would
raise — so load it deliberately and check the answer is the path anyway.
`adventure.actions` returns every branch's actions. Slicing it without
the path would assemble a prompt out of two different stories, and
nothing would raise an error. This test loads the collection
deliberately and checks that the answer is still the path.
"""
db, adventure, _ = forked
loaded = list(adventure.actions) # every branch, ordered by index
@@ -230,8 +235,8 @@ def test_an_already_loaded_collection_is_cut_down_to_the_path(forked):
def test_user_scripts_are_handed_the_path(forked):
"""The same trap, one layer up and user-visible: `pipeline._history()` is
the documented scripting history API."""
"""The same risk one layer up, in code visible to users:
`pipeline._history()` is the documented scripting history API."""
db, adventure, _ = forked
list(adventure.actions) # the pipeline's caller has usually loaded these
pipeline = ScriptPipeline(adventure, db)
@@ -268,8 +273,8 @@ def test_the_page_the_reader_opens_is_the_path(client):
assert [a["text"] for a in body["actions"]] == [
"A0", "A1", "A2", "A3", "B4", "B5", "C6", "C7"
]
# `action_count` is what tells the reader there is more above, so it counts
# the path too — 8, not the 13 rows the adventure holds.
# `action_count` tells the reader whether more actions exist above. It
# counts the path too: 8, not the 13 rows the adventure holds.
assert body["action_count"] == 8
@@ -324,11 +329,11 @@ def test_the_index_screen_quotes_the_branch_being_played(client):
# ------------------------------------------------------- the flush guard
def test_a_node_written_without_a_branch_is_placed_anyway(forked):
"""SP1 wired the writers; from SP2 an unplaced node is an invisible one.
"""SP1 wired the writers. Since SP2, an unplaced node is an invisible one.
This is what lets a fixture, a script or a test built straight through the
ORM keep working — and it is why the baseline contract still passes with
its actions written directly to the database.
This behavior lets a fixture, a script, or a test built straight through
the ORM keep working. It is also why the baseline contract still passes
with its actions written directly to the database.
"""
db, adventure, ids = forked
written = models.Action(
@@ -356,11 +361,11 @@ def test_a_memory_written_without_a_branch_is_placed_anyway(forked):
def test_placing_a_flush_of_nodes_reads_the_branch_once(forked, emitted_sql):
"""The guard resolves the head once per flush, not once per node.
The identity map holds weak references, so a branch row nobody keeps a
strong reference to is collected between two nodes and read back for the
next one. Writing two hundred actions in one flush was two hundred SELECTs
on `branches` before the head was hoisted out of the loop, and nothing
about the result would have told you.
The identity map holds weak references. A branch row with no strong
reference gets collected between two nodes and read back again for the
next one. Writing two hundred actions in one flush ran two hundred
SELECTs on `branches` before the head lookup moved outside the loop.
The test result alone would not have shown this.
"""
db, adventure, _ = forked
emitted_sql.clear()
@@ -376,10 +381,11 @@ def test_placing_a_flush_of_nodes_reads_the_branch_once(forked, emitted_sql):
def test_an_adventure_with_no_branch_at_all_reads_as_empty(forked):
"""The loud version of a missing branch: nothing, rather than everything.
"""A missing branch must fail loudly: nothing, rather than everything.
A row with no branch cannot be shown without guessing which story it is
on, and a guess here is how a sibling's turns end up in a prompt.
A row with no branch cannot be shown without guessing which story it
belongs to, and a wrong guess here puts a sibling's turns into a
prompt.
"""
db, adventure, ids = forked
stray = make_adventure(db, db.get(models.User, ids["user"]), "Stray")
@@ -400,9 +406,9 @@ def test_an_adventure_with_no_branch_at_all_reads_as_empty(forked):
def deeply_forked():
"""A story forked twenty times, then played forty turns past the last one.
The shape the design is betting on: reading the tail of this must cost what
reading the tail of an unforked story costs, because the window is covered
long before the ancestry runs out.
This shape tests the design's core assumption: reading the tail of this
story must cost the same as reading the tail of an unforked story,
because the window is covered long before the ancestry runs out.
"""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
@@ -477,7 +483,7 @@ def test_a_window_reaching_past_the_forks_names_only_what_it_needs(
rows = history.tail(adventure, 41) # 40 on the tip branch, one older
assert len(rows) == 41
selects = [s for s in emitted_sql if "FROM actions" in s and branch_terms(s)]
# Two lineage entries reach 41 deep (40 + 2); the other twenty stay
# Two lineage entries reach 41 deep (40 + 2). The other twenty stay
# unnamed. Every fork past the window costs the query nothing.
assert max(branch_terms(s) for s in selects) == 2
@@ -495,13 +501,13 @@ def test_the_estimate_is_arithmetic_not_a_query(deeply_forked):
def test_forking_twenty_times_costs_the_same_bytes_as_never_forking(
deeply_forked,
):
"""The design's bet, in bytes.
"""The design's cost assumption, measured in bytes.
Two stories of the same length, one played straight through and one forked
twenty times, read their newest window for the same money — because the
window is covered by the newest lineage entry either way, and the ancestry
is never named. The forked read pays for one extra row: the branch it read
the lineage off.
Two stories of the same length, one played straight through and one
forked twenty times, cost the same to read their newest window. The
window is covered by the newest lineage entry either way, and the
ancestry is never named. The forked read pays for one extra row: the
branch it reads the lineage from.
"""
db, forked_adventure = deeply_forked
flat = make_adventure(db, db.get(models.User, forked_adventure.user_id), "Flat")
@@ -511,10 +517,10 @@ def test_forking_twenty_times_costs_the_same_bytes_as_never_forking(
flat.head_branch_id = branch.id
flat.head_depth = 83
flat_id, forked_id = flat.id, forked_adventure.id
# Commit and let go of the connection: the meter wraps the pool's factory,
# so a connection checked out before it attaches is a connection it never
# sees. Building the fixture is a write path nobody plays, and is not
# charged to either scope.
# Commit and release the connection. The meter wraps the connection
# pool's factory, so a connection checked out before the meter attaches
# is never visible to it. Building the fixture is a write path the test
# does not measure, and it is not charged to either scope.
db.commit()
db.expire_all()
@@ -541,8 +547,8 @@ def test_forking_twenty_times_costs_the_same_bytes_as_never_forking(
def test_a_gap_in_the_story_widens_the_read_rather_than_shortening_it(
deeply_forked, emitted_sql
):
"""The estimate counts depths, and a deleted action leaves a depth with no
row behind it. The read has to notice it came up short and widen."""
"""The estimate counts depths, and a deleted action leaves a depth with
no row behind it. The read must notice it came up short and widen."""
db, adventure = deeply_forked
victim = (
db.query(models.Action)
+79 -71
View File
@@ -1,14 +1,14 @@
"""Phase 14 SP5 — continuing from a discarded attempt forks a branch.
"""Phase 14 SP5: continuing from a discarded attempt forks a branch.
SP4 made every attempt at a turn a node. While the attempts sit at the tip they
are leaves and cost nothing: switching between them just moves the `live` flag.
The moment the player takes the story down one the line has already moved past,
the two futures have to coexist — and that is a branch.
SP4 made every attempt at a turn a node. While the attempts sit at the
tip, they are leaves and cost nothing: switching between them just moves
the `live` flag. The moment the player continues from an attempt the line
has already moved past, the two futures must coexist. That is a branch.
What this file is really watching is the claim the whole design rests on: **a
fork inserts one row and moves one row, whatever the story behind it is worth.**
Everything before the fork is borrowed, not copied, and the arithmetic that
makes borrowing readable is `lineage`.
This file tests the claim the whole design rests on: a fork inserts one
row and moves one row, no matter how large the story behind it is.
Everything before the fork is borrowed, not copied, and the arithmetic
that makes borrowing possible lives in `lineage`.
python -m pytest tests/test_branch_forking.py -v
"""
@@ -32,8 +32,8 @@ from app.main import app
from app.providers import PromptParts
from app.routers import adventures
# `hp` moves freely; `mana` has a cooldown of 2 turns, so a clock that advances
# when it should not shows up as a change the referee should have rejected.
# `hp` moves freely. `mana` has a cooldown of 2 turns, so an incorrect
# advance shows up as a change the referee should have rejected.
SCHEMA = {
"player": {
"hp": {"min": 0, "max": 100, "initial": 100},
@@ -164,12 +164,12 @@ def _rows(adv_id) -> list[models.Action]:
def _divergent_story(client):
"""A story that retried turn 2, continued from the newer take, and left the
older one behind as a leaf.
"""A story that retried turn 2, continued from the newer take, and left
the older one behind as a leaf.
start · do · [attempt one | ATTEMPT TWO] · do · next turn
start > do > [attempt one | ATTEMPT TWO] > do > next turn
Returns the id of the attempt nobody built on.
Returns the id of the discarded attempt.
"""
ScriptedProvider.replies = ["Attempt one.", "Attempt two.", "Next turn."]
_play(client)
@@ -194,9 +194,10 @@ def test_a_fork_inserts_one_branch_row_and_copies_no_actions(client):
forked = [b for b in branches if b["parent_branch_id"] is not None][0]
assert forked["is_head"] is True
assert forked["own_actions"] == 1, "the promoted attempt, and nothing else"
# The fork point is the depth just before the attempt, stored rather than
# inferred: inferring it from where two branches first differ would be a
# guess, and a wrong one as soon as an attempt repeats its parent's text.
# The fork point is the depth just before the attempt. The code stores
# this value instead of inferring it from where two branches first
# differ, because that inference would guess wrong as soon as an
# attempt repeats its parent's text.
assert forked["fork_depth"] == forked["depth"] - 1
@@ -225,7 +226,7 @@ def test_both_branches_read_independently(client):
assert _texts(client) == [
"You enter a cave.", "> You look around.", "Attempt one.",
]
# And the line it left is exactly as it was, turns after the fork included.
# The branch it left behind is unchanged, including turns after the fork.
r = client.post(f"/api/adventures/{client.adv_id}/branches/{parent}/switch")
assert r.status_code == 200, r.text
assert _texts(client) == [
@@ -235,8 +236,9 @@ def test_both_branches_read_independently(client):
def test_the_parent_keeps_a_live_attempt_where_the_fork_left(client):
"""Promoting the loser must not leave the parent with a hole in its story:
a coordinate with no live node is a turn that disappears from the read."""
"""Promoting the other attempt must not leave the parent with a gap in
its story. A coordinate with no live node is a turn that disappears
from the read."""
discarded = _divergent_story(client)
_fork(client, discarded)
@@ -251,8 +253,8 @@ def test_the_parent_keeps_a_live_attempt_where_the_fork_left(client):
for (branch_id, depth), group in per_coordinate.items():
live = [a for a in group if a.live]
assert len(live) == 1, f"branch {branch_id} depth {depth}"
# The parent's turn 2 is now a single take, so the pager stops offering
# a page through attempts that have gone their own way.
# The parent's turn 2 is now a single take, so the pager no longer
# offers a page through attempts that diverged onto another branch.
parent_turn = per_coordinate[(parent_id, 2)]
assert len(parent_turn) == 1
assert parent_turn[0].variant_count == 0
@@ -261,9 +263,10 @@ def test_the_parent_keeps_a_live_attempt_where_the_fork_left(client):
def test_playing_on_a_fork_continues_that_branchs_depths(client):
"""A depth is a position along *this* story. Numbering the next node from
the adventure-wide index would leave a hole where the other branch's turns
are, which every windowing estimate then has to work around."""
"""A depth is a position along this story. Numbering the next node from
the adventure-wide index would leave a gap where the other branch's
turns are, and every windowing estimate would then have to work around
that gap."""
discarded = _divergent_story(client)
_fork(client, discarded)
ScriptedProvider.replies = ["Onward."]
@@ -283,8 +286,8 @@ def test_playing_on_a_fork_continues_that_branchs_depths(client):
# ------------------------------------------------------------- not a fork
def test_forking_at_the_tip_switches_without_making_a_branch(client):
"""Attempts nobody has built on stay leaves — that is what keeps the
lineage a list of divergences rather than of every retry ever."""
"""Attempts nobody has built on stay leaves. This is what keeps the
lineage a list of divergences instead of a list of every retry."""
ScriptedProvider.replies = ["Attempt one.", "Attempt two."]
_play(client)
_retry(client)
@@ -297,8 +300,8 @@ def test_forking_at_the_tip_switches_without_making_a_branch(client):
def test_forking_the_attempt_already_in_the_story_does_nothing(client):
"""Idempotent, because a client that has lost track of which take is live
must not be able to fork a branch per click."""
"""This call is idempotent. A client that has lost track of which take
is live must not create a new branch on every click."""
discarded = _divergent_story(client)
_fork(client, discarded)
promoted = [a.id for a in _rows(client.adv_id) if a.type == "ai" and a.live
@@ -323,10 +326,11 @@ def test_forking_a_turn_that_is_already_the_story_is_a_no_op(client):
def test_forking_a_live_node_on_another_branch_is_refused(client):
"""A live node off the path is another line's story, not an attempt going
spare — so the refusal names the tool that would actually do it. It used to
answer "only one take", which was true of the group and no help at all: the
caller does not want another take, it wants the branch this one is on."""
"""A live node off the path belongs to another branch's story, not to a
spare attempt on this one. The refusal names the tool that actually
switches branches. It used to answer "only one take", which was true
of the attempt group but useless here: the caller does not want
another take, it wants the branch this node is on."""
discarded = _divergent_story(client)
_fork(client, discarded)
parent_id = [b for b in _branches(client) if b["parent_branch_id"] is None][0]["id"]
@@ -336,8 +340,8 @@ def test_forking_a_live_node_on_another_branch_is_refused(client):
r = _fork(client, stranded)
assert r.status_code == 400
assert "another branch" in r.json()["detail"]
# And refusing left the tree alone — the bug this guards is a fork that
# promotes a sibling on the branch it was called against.
# Refusing must leave the tree alone. The bug this guards against is a
# fork that promotes a sibling on the branch it was called against.
assert len(_branches(client)) == 2
@@ -366,10 +370,10 @@ def test_switching_restores_the_script_and_world_state(client):
def test_the_cooldown_clock_travels_with_the_branch(client):
"""The world-state clock is a depth, and depths repeat across branches — so
it can only be right if each branch carries its own. It does, for free: the
clock lives inside `_meta.last_changed`, which is part of the world state a
switch restores."""
"""The world-state clock is a depth, and depths repeat across branches,
so it can only be correct if each branch carries its own. It does,
without extra work: the clock lives inside `_meta.last_changed`, which
is part of the world state a switch restores."""
ScriptedProvider.replies = [
"Drained.\n```state\n{\"player.mana\": -10}\n```",
"Untouched.",
@@ -393,9 +397,9 @@ def test_the_cooldown_clock_travels_with_the_branch(client):
def test_a_retry_does_not_advance_the_cooldown_clock(client):
"""SP5's one carried-over open item. A retry re-runs the *same* turn, so
the clock the cooldown rules read must not move — it used to be the reused
`index` that guaranteed this, and it is the reused depth now."""
"""SP5's one carried-over open item. A retry re-runs the same turn, so
the clock the cooldown rules read must not move. The reused `index`
used to guarantee this; the reused depth guarantees it now."""
ScriptedProvider.replies = [
"Drained.\n```state\n{\"player.mana\": -10}\n```",
"Drained again.\n```state\n{\"player.mana\": -10}\n```",
@@ -404,18 +408,19 @@ def test_a_retry_does_not_advance_the_cooldown_clock(client):
first = _state(client.adv_id)[1]["_meta"]["last_changed"]["player.mana"]
_retry(client)
assert _state(client.adv_id)[1]["_meta"]["last_changed"]["player.mana"] == first
# ...and the second attempt's drain landed, rather than being rejected for
# a cooldown it should never have been measured against.
# The second attempt's drain must land, instead of being rejected for a
# cooldown it was never actually subject to.
assert _state(client.adv_id)[1]["player"]["mana"] == 40
# --------------------------------------------------------- derived work
def test_a_memory_on_the_line_left_behind_is_out_of_range_on_the_fork(client):
"""Nothing is moved or withdrawn when a branch forks. The memory hangs off
the coordinate the *parent's* attempt still occupies, and the lineage caps
the parent one depth short of it — so the fork simply cannot see it, and
resummarizes that ground from the text it actually tells."""
"""Nothing is moved or removed when a branch forks. The memory attaches
to the coordinate the parent's attempt still occupies, and the lineage
caps the parent one depth short of it. The fork therefore cannot see
this memory, and resummarizes that span from the text it actually
contains."""
from app import tree
discarded = _divergent_story(client)
@@ -441,8 +446,8 @@ def test_a_memory_on_the_line_left_behind_is_out_of_range_on_the_fork(client):
db = SessionLocal()
try:
adventure = db.get(models.Adventure, client.adv_id)
# Still there, untouched — it describes the parent's story, which is
# unchanged.
# The memory is still there, untouched, because it describes the
# parent's story, which is unchanged.
assert [m.text for m in db.query(models.Memory).all()] == ["Attempt two happened."]
path = lineage.path_of(db, adventure)
visible = db.query(models.Memory).filter(
@@ -450,8 +455,8 @@ def test_a_memory_on_the_line_left_behind_is_out_of_range_on_the_fork(client):
path.clause(models.Memory),
).all()
assert visible == [], "a sibling's memory reached this branch"
# ...and the mark reads as one depth short of it, so the block is due
# again on this branch rather than silently claimed as read.
# The cursor reads one depth short of the memory, so this branch
# treats the block as due again instead of silently marking it read.
assert cursors.MEMORY.depth(db, adventure) == 1
finally:
db.close()
@@ -460,20 +465,22 @@ def test_a_memory_on_the_line_left_behind_is_out_of_range_on_the_fork(client):
# ------------------------------------------------------------------- undo
def test_undo_stops_at_the_fork(client):
"""Taking back a turn on a fork must never reach across into the line it
was forked from — those turns are that branch's story too."""
"""Undoing a turn on a fork must never reach into the branch it forked
from. Those turns belong to that branch's story too."""
discarded = _divergent_story(client)
_fork(client, discarded)
rows_before = len(_rows(client.adv_id))
r = client.post(f"/api/adventures/{client.adv_id}/undo")
assert r.status_code == 200, r.text
# The promoted attempt goes, and the player action in front of it stays:
# it is on the parent, and the parent still tells it.
# The promoted attempt is removed, and the player action before it
# stays, because that action belongs to the parent and the parent
# still has it.
assert len(_rows(client.adv_id)) == rows_before - 1
assert _texts(client) == ["You enter a cave.", "> You look around."]
# Nothing left of this branch's own: refuse rather than eat the parent's.
# Nothing is left of this branch's own turns, so undo must refuse
# instead of removing the parent's turns.
r = client.post(f"/api/adventures/{client.adv_id}/undo")
assert r.status_code == 400
assert "forked from" in r.json()["detail"]
@@ -497,10 +504,11 @@ def test_the_branch_list_is_the_tree(client):
def test_fork_agrees_with_the_lineage_computed_by_hand(client):
"""`test_branch_clause.make_branch` has computed a fork's lineage since SP2,
by hand, precisely so the fixture could not pass by agreeing with a bug in
the code under test. SP5 is when that code exists — so check the two
against each other rather than letting them drift apart."""
"""`test_branch_clause.make_branch` has computed a fork's lineage by
hand since SP2, precisely so the fixture could not pass by repeating a
bug in the code under test. SP5 introduces that code, so this test
checks the two against each other instead of letting them drift
apart."""
from tests.test_branch_clause import make_branch
discarded = _divergent_story(client)
@@ -512,8 +520,8 @@ def test_fork_agrees_with_the_lineage_computed_by_hand(client):
real = lineage.branch_of(db, adventure)
parent = db.get(models.Branch, real.parent_branch_id)
by_hand = make_branch(db, adventure, parent=parent, fork_depth=real.fork_depth)
# Same shape, its own id: compare the ancestry, which is the part that
# is arithmetic rather than allocation.
# The two branches share a shape but not an id, so compare only the
# ancestry, which is the arithmetic part rather than the allocated id.
assert lineage.entries_of(real)[1:] == lineage.entries_of(by_hand)[1:]
assert lineage.entries_of(real)[0] == (real.id, None)
finally:
@@ -521,9 +529,9 @@ def test_fork_agrees_with_the_lineage_computed_by_hand(client):
def test_a_deep_fork_chain_reads_for_what_one_branch_costs(client):
"""Clause count is bounded by the window, not by fork count — the property
the whole lineage cache exists for, now measured through real forks rather
than hand-built rows."""
"""Clause count is bounded by the window, not by fork count. This is
the property the whole lineage cache exists for, now measured through
real forks instead of hand-built rows."""
from tools import dbmeter
ScriptedProvider.replies = ["First take.", "Second take.", "Onward."]
@@ -555,9 +563,9 @@ def test_a_deep_fork_chain_reads_for_what_one_branch_costs(client):
try:
adventure = db.get(models.Adventure, client.adv_id)
entries = lineage.entries_of(lineage.branch_of(db, adventure))
# The whole ancestry is there to be named...
# The whole ancestry is available to be named.
assert len(entries) == len(branches)
# ...and the windowed read names as few of them as the window needs.
# The windowed read names as few of them as the window needs.
path = lineage.path_of(db, adventure)
assert path.prefix_covering(60) <= len(entries)
finally:
@@ -565,8 +573,8 @@ def test_a_deep_fork_chain_reads_for_what_one_branch_costs(client):
def test_switching_to_a_branch_of_another_adventure_is_a_404(client):
"""A branch id names one adventure, so the two ids in the URL have to
agree — otherwise a guessed number reads somebody else's story."""
"""A branch id names one adventure, so the two ids in the URL must
agree. Otherwise a guessed number reads somebody else's story."""
other = client.post("/api/adventures", json={"title": "Elsewhere"}).json()["id"]
ScriptedProvider.replies = ["Elsewhere."]
r = client.post(f"/api/adventures/{other}/actions", json={"type": "do", "text": "wait"})
+58 -50
View File
@@ -1,21 +1,21 @@
"""Phase 14 SP7 — naming a branch, and throwing one away.
"""Phase 14 SP7: renaming a branch and deleting one.
SP5 gave the tree a fork and a switch. Neither of them ever removes anything,
and nothing in the design prunes a tree on its own, so an adventure that is
retried and forked enough grows without a ceiling. Delete is what stands
between the tree and that, which is why it ships with the view that first makes
a fork reachable rather than in some later subphase.
SP5 gave the tree a fork and a switch. Neither operation removes anything,
and nothing in the design prunes a tree automatically, so an adventure that
is retried and forked enough grows without a limit. Delete is what limits
that growth. It ships with the same view that first makes a fork reachable,
rather than in a later subphase.
Two rules carry most of this file:
* **A name is chosen, so it is stored; a label is derived, so it is not.** An
unnamed branch keeps NULL and the client draws it from its fork depth. A
generated "branch 4" in the column would be a lie the moment branch 3 is
deleted.
* **The delete may never take the ground under the reader.** Refusing the head
is the obvious half; refusing an *ancestor* of the head is the same mistake
wearing a disguise, and it is the one that would leave `head_branch_id`
pointing at a row the cascade removed.
* A name is chosen, so the database stores it. A label is derived, so the
database does not. An unnamed branch keeps NULL, and the client draws the
label from its fork depth. A generated label such as "branch 4" stored in
the column would become wrong the moment branch 3 is deleted.
* Delete must never remove a branch the reader depends on. Refusing to
delete the head is the obvious case. Refusing to delete an ancestor of the
head is the same rule applied one level up: deleting it would leave
`head_branch_id` pointing at a row the cascade removed.
python -m pytest tests/test_branch_management.py -v
"""
@@ -136,7 +136,7 @@ def _texts(client) -> list[str]:
def _discarded_on(adv_id, branch_id=None) -> int:
"""An AI attempt nobody built on, optionally restricted to one branch."""
"""An AI attempt with no turn built on it, optionally restricted to one branch."""
db = SessionLocal()
try:
q = db.query(models.Action).filter(
@@ -152,7 +152,7 @@ def _discarded_on(adv_id, branch_id=None) -> int:
def _forked(client):
"""A story with one fork. Returns (root id, forked id); the fork is head.
"""A story with one fork. Returns (root id, forked id). The fork is head.
start · do · [attempt one | ATTEMPT TWO] · do · next turn
└── forked here
@@ -184,10 +184,10 @@ def _counts(adv_id, branch_ids):
# ------------------------------------------------------------------ naming
def test_a_branch_starts_unnamed(client):
"""NULL, not a generated label — the client draws one from the fork depth.
"""NULL, not a generated label. The client draws a label from the fork depth.
A name written here would go stale the moment a branch before it is
deleted and the ordinals shift under it.
A name stored here would go stale the moment a branch before it is
deleted and the ordinals shift.
"""
root, forked = _forked(client)
assert [b["name"] for b in _branches(client)] == [None, None]
@@ -206,9 +206,8 @@ def test_a_name_is_stored_and_read_back(client):
def test_a_blank_name_goes_back_to_unnamed(client):
"""A name of spaces is not a name anyone chose.
Storing one would give the client an empty label to draw where it would
otherwise fall back to the fork depth — a branch that looks nameless and
reads as broken.
Storing it would give the client an empty label instead of falling back
to the fork depth. The branch would look nameless and appear broken.
"""
root, forked = _forked(client)
_rename(client, forked, "briefly named")
@@ -226,9 +225,10 @@ def test_a_name_longer_than_the_column_is_refused(client):
def test_a_rename_hands_back_the_row_the_listing_would_give(client):
"""Renaming a branch does not change how many turns are on it.
`own_actions` was hard-coded to 0 in this response, which only stayed
invisible because the panel throws the body away and refetches. Anything
that trusted the reply would draw a branch that had just lost its turns.
`own_actions` was hard-coded to 0 in this response. That bug stayed
invisible because the panel discards the response body and refetches.
Anything that trusted the reply would show a branch that had just lost
its turns.
"""
root, forked = _forked(client)
listed = {b["id"]: b for b in _branches(client)}
@@ -278,10 +278,11 @@ def test_the_branch_being_read_cannot_be_deleted(client):
def test_an_ancestor_of_the_branch_being_read_cannot_be_deleted(client):
"""The same mistake as deleting the head, wearing a disguise.
"""Deleting an ancestor of the head is the same mistake as deleting the head.
`parent_branch_id` cascades, so deleting a branch the head was forked from
would take the head with it and leave `head_branch_id` pointing at nothing.
`parent_branch_id` cascades, so deleting a branch the head was forked
from would delete the head too, and leave `head_branch_id` pointing at
nothing.
"""
root, forked = _forked(client)
# A fork of the fork, so `forked` is an ancestor of the head rather than
@@ -310,7 +311,8 @@ def test_deleting_a_branch_leaves_the_line_it_forked_from_untouched(client):
def test_deleting_a_branch_takes_its_nodes_and_its_descendants(client):
"""One statement, however deep the subtree — the cascade does the walking."""
"""One statement deletes the whole subtree, regardless of depth. The
cascade performs the traversal."""
root, forked = _forked(client)
_retry(client)
_play(client, "press on")
@@ -333,10 +335,11 @@ def test_deleting_a_branch_takes_its_nodes_and_its_descendants(client):
def test_deleting_a_branch_clears_a_cursor_that_stood_on_it(client):
"""Harmless on Postgres, a real bug on SQLite.
Postgres never reuses a branch id, so a stale anchor simply never resolves.
SQLite hands the freed id to the next fork, at which point the anchor
resolves onto a branch it has never seen and reports a stretch of story as
already summarized — which loses it from the memories for good.
Postgres never reuses a branch id, so a stale anchor simply never
resolves. SQLite assigns the freed id to the next fork. The anchor then
resolves onto a branch it never saw, and reports a stretch of story as
already summarized. That stretch is then permanently excluded from the
memories.
"""
root, forked = _forked(client)
db = SessionLocal()
@@ -355,7 +358,7 @@ def test_deleting_a_branch_clears_a_cursor_that_stood_on_it(client):
try:
adventure = db.get(models.Adventure, client.adv_id)
assert cursors.MEMORY.stored(adventure) == (None, cursors.NO_DEPTH)
# The one standing on ground that survived is left exactly where it was.
# The cursor on a branch that still exists is left exactly where it was.
assert cursors.SUMMARY.stored(adventure) == (root, 1)
finally:
db.close()
@@ -398,11 +401,11 @@ def test_a_hand_written_memory_is_anchored_where_it_was_written(client):
def test_the_drawer_shows_the_path_being_read_and_nothing_else(client):
"""The bank you can see is the bank the model can see.
"""The memories a reader can see match the memories the model can retrieve.
An adventure-wide list would show memories from branches this story never
went down — which are never retrieved — and a reader cannot tell those from
the ones actually in play.
An adventure-wide list would include memories from branches this story
never took. Those memories are never retrieved, and a reader could not
distinguish them from the ones actually in play.
"""
root, forked = _forked(client)
on_the_fork = _add_memory(client, "Took the other door.")
@@ -412,9 +415,10 @@ def test_the_drawer_shows_the_path_being_read_and_nothing_else(client):
assert {m["id"] for m in _memories(client)} == {on_the_root}, \
"the fork's memory is not on this story"
# Switching to the fork shows its own memory — and the root's, because a
# fork borrows its ancestors up to the point it left them. The relationship
# is asymmetric on purpose; the parent never went down the fork.
# Switching to the fork shows its own memory and the root's. A fork
# borrows its ancestors up to the point where it diverged from them.
# The relationship is asymmetric on purpose: the parent never took the
# fork's path.
_switch(client, forked)
listed = {m["id"] for m in _memories(client)}
assert on_the_fork in listed
@@ -422,8 +426,9 @@ def test_the_drawer_shows_the_path_being_read_and_nothing_else(client):
def test_the_drawer_and_retrieval_agree_on_what_is_visible(client):
"""One predicate, so a memory can never be listed but unretrievable (or the
reverse). Two spellings of "on this path" would eventually drift."""
"""One predicate decides visibility, so a memory can never be listed but
unretrievable, or the reverse. Two separate definitions of "on this
path" would eventually diverge."""
root, forked = _forked(client)
_add_memory(client, "Took the other door.")
_switch(client, root)
@@ -447,10 +452,11 @@ def test_the_drawer_and_retrieval_agree_on_what_is_visible(client):
def test_deleting_a_branch_deletes_the_memories_written_on_it(client):
"""Not merely out of view — the row goes with the branch, through the
cascade. That is what keeps "the drawer shows only your path" from
stranding anything: a memory you cannot see is on a branch you can still
switch to, and deleting that branch takes it for good."""
"""Deleting a branch removes its memories from the database, not just
from view: the cascade deletes the row along with the branch. This
keeps a hidden memory from becoming unreachable in a different way: a
memory you cannot currently see is on a branch you can still switch to,
and deleting that branch deletes the memory permanently."""
root, forked = _forked(client)
doomed = _add_memory(client, "Took the other door.")
_switch(client, root)
@@ -473,10 +479,12 @@ def test_deleting_a_branch_deletes_the_memories_written_on_it(client):
# ------------------------------------------------------------------- backup
def test_a_bundle_carries_the_name_a_player_chose(client):
"""A name is a decision, so it travels — the rule the v2 format is built on.
"""A name is a decision, so the export includes it. This is the rule the
v2 format is built on.
`lineage` and the head depth stay out because they are computed from what
the file already carries; a name is computed from nothing.
`lineage` and the head depth stay out of the export because the
importer can compute them from what the file already carries. A name
is computed from nothing, so the export must carry it.
"""
root, forked = _forked(client)
_rename(client, root, "the long way")
+91 -78
View File
@@ -1,24 +1,25 @@
"""Phase 14 SP6 — the export bundle carries the tree.
"""Phase 14 SP6: the export bundle carries the tree.
A bundle is the only part of this phase a migration can never reach: the file
is already on somebody's disk. So there are two formats, and the two halves of
this file watch different things.
A bundle is the only part of this phase a migration can never reach: the
file already exists on somebody's disk. So there are two formats, and the
two halves of this file check different properties.
**v2 has to be lossless for a story that forked**, which v1 could not be — it
had one list and there were two stories, so it interleaved them by `index` and
read as a mangled story. Losslessness here means the *tree*: every branch, the
fork point it left its parent at, which attempt at each turn is the story, and
what each node left behind — because that last one is what a branch switch puts
back, and a tree nobody can switch inside is not the tree that was exported.
v2 must be lossless for a story that forked. v1 could not be, because it
stored one list for two stories, interleaved by `index`, which read back
as a mangled story. Losslessness here means the tree: every branch, the
fork point it left on its parent, which attempt at each turn is the
story, and what each node left behind. That last item is what a branch
switch restores, and a tree nobody can switch inside is not the tree
that was exported.
**v1 has to still import**, because a backup that stops importing is not a
v1 must still import, because a backup that stops importing is not a
backup.
Underneath both is the rule the module is built on: a bundle carries what was
*chosen* and never what is *derived*. The lineage, the head depth, the legacy
`index` and the variant ordinals are all rebuilt on the way in, so a
hand-edited file cannot disagree with itself — and the tests that matter most
here are the ones that hand it a file which does.
Both formats follow one rule: a bundle carries what was chosen and never
what is derived. The lineage, the head depth, the legacy `index`, and the
variant ordinals are all rebuilt on the way in, so a hand-edited file
cannot disagree with itself. The tests that matter most here hand the
importer a file that does disagree with itself.
python -m pytest tests/test_bundle_v2.py -v
"""
@@ -44,8 +45,8 @@ from app.routers import adventures
SCHEMA = {"player": {"hp": {"min": 0, "max": 100, "initial": 100}}}
# Ten gold a turn, so the script scoreboard is a number that says how many turns
# the story behind it has — which makes an after-snapshot visible from outside.
# Ten gold a turn, so the stored gold total tells how many turns the
# story behind it played. This makes an after-snapshot visible from outside.
GOLD_SCRIPT = """
const modifier = (text) => {
state.gold = (state.gold || 0) + 10;
@@ -161,7 +162,7 @@ def _texts(client, adv_id) -> list[str]:
def _every_branch_story(client, adv_id) -> list[list[str]]:
"""What each branch tells, in branch order — the whole tree as text."""
"""What each branch tells, in branch order. This is the whole tree as text."""
stories = []
for branch in _branches(client, adv_id):
_switch(client, adv_id, branch["id"])
@@ -213,13 +214,14 @@ def _script_state(adv_id) -> dict:
def _forked_story(client) -> int:
"""A story that went two ways, and stayed both.
"""A story that went two ways and stayed both.
root: start · do · [attempt two] · do · next turn
fork: [ATTEMPT ONE] · do · elsewhere
root: start > do > [attempt two] > do > next turn
fork: [ATTEMPT ONE] > do > elsewhere
The *discarded* attempt is the one that gets promoted, because a fork moves
the take you are leaving for and leaves the line you came from untouched.
The discarded attempt is the one this function promotes, because a
fork moves the attempt being left for and leaves the line it came
from untouched.
Returns the adventure id, with the head on the fork.
"""
@@ -237,11 +239,11 @@ def _forked_story(client) -> int:
# ------------------------------------------------------- the round trip (v2)
def test_a_forked_story_survives_the_round_trip(client):
"""The headline: both futures come back, and both are still readable.
"""The main claim: both futures come back, and both are still readable.
This is the thing v1 could not do. The check is not "the same rows" — the
ids are new — but "the same stories", read the way a player reads them: by
switching to a branch and looking at what it says.
This is the thing v1 could not do. The check is not "the same rows",
since the ids are new, but "the same stories", read the way a player
reads them: by switching to a branch and looking at what it says.
"""
original = _forked_story(client)
before = _every_branch_story(client, original)
@@ -254,11 +256,11 @@ def test_a_forked_story_survives_the_round_trip(client):
def test_the_fork_point_comes_back_where_it_was_put(client):
"""`fork_depth` is stored, never inferred — including through a file.
"""`fork_depth` is stored, never inferred, including through a file.
Inferring it from where two branches' nodes first differ would be a guess
about how the story was played, and a wrong one the moment an attempt
happens to repeat its parent's text.
Inferring it from where two branches' nodes first differ would guess
at how the story was played, and that guess fails as soon as an
attempt happens to repeat its parent's text.
"""
original = _forked_story(client)
before = [(b["parent_branch_id"] is None, b["fork_depth"]) for b in _branches(client, original)]
@@ -275,22 +277,23 @@ def test_the_head_comes_back_on_the_branch_it_was_left_on(client):
copy = _imported(client, _export(client, original))
assert [b["is_head"] for b in _branches(client, copy)] == head_before
# And the tip it sits at is derived from the nodes that arrived, not read
# out of the file — the bundle never says how deep a branch goes.
# The tip it sits at is derived from the nodes that arrived, not
# read from the file. The bundle never states how deep a branch goes.
assert _texts(client, copy) == _texts(client, original)
def test_a_switch_in_the_copy_restores_what_that_branch_left_behind(client):
"""The after-snapshots are why the bundle carries them.
"""This test justifies why the bundle carries after-snapshots.
The gold script adds ten a turn, so the scoreboard is a count of the story
behind it. A bundle that carried the actions but not the outcomes would
import a tree that reads correctly and switches wrong.
The gold script adds ten a turn, so the stored gold total counts the
turns behind it. A bundle that carried the actions but not the
outcomes would import a tree that reads correctly but switches to the
wrong state.
"""
original = _forked_story(client)
# One more turn on the fork, so the two tips are genuinely different
# numbers: played turn for turn, the branches earn the same gold and a
# switch that restored nothing at all would still look right.
# Play one more turn on the fork, so the two tips end up at
# genuinely different totals. Turn for turn, both branches earn the
# same gold, so a switch that restored nothing would still look right.
ScriptedProvider.replies = ["Further still."]
_play(client, original, "press on")
@@ -340,12 +343,12 @@ def test_a_memory_comes_back_on_the_node_it_hangs_off(client):
# --------------------------------------------------- what is not in the file
def test_the_lineage_is_rebuilt_rather_than_carried(client):
"""A cache of `parent` + `fork_depth` is not a second thing to ship.
"""A cache of `parent` plus `fork_depth` is not a second thing to ship.
The file says where each branch forked; the ancestry that makes the fork
readable is computed from that on the way in, capped at the fork exactly as
`tree.fork` caps it. Shipping the cache too would put two sources of truth
for one fact in a file anybody can hand-edit.
The file states where each branch forked. The ancestry that makes the
fork readable is computed from that value on the way in, capped at
the fork exactly as `tree.fork` caps it. Shipping the cache too would
put two sources of truth for one fact in a file anybody can hand-edit.
"""
original = _forked_story(client)
exported = _export(client, original)
@@ -354,18 +357,19 @@ def test_the_lineage_is_rebuilt_rather_than_carried(client):
root, forked = _branch_rows(_imported(client, exported))
assert root.lineage == [[root.id, None]]
assert forked.lineage == [[forked.id, None], [root.id, forked.fork_depth]]
# Which is the arithmetic the reader depends on: the parent is capped one
# depth short of the attempt that was promoted, so the fork cannot see it.
# This is the arithmetic the reader depends on. The parent is capped
# one depth short of the attempt that was promoted, so the fork
# cannot see it.
assert lineage.entries_of(forked) == [(forked.id, None), (root.id, forked.fork_depth)]
def test_the_legacy_index_is_reissued_so_two_branches_never_share_one(client):
"""`index` is a fact about the adventure, and `depth` is one about a path.
"""`index` describes the adventure. `depth` describes a path.
Two branches have a node at depth 3, so depth cannot be the number
`max_action_index` hands out next. The import allocates one per turn
instead: siblings share it, the way SP4 leaves them, and no two coordinates
do.
Two branches can each have a node at depth 3, so depth cannot be the
number `max_action_index` hands out next. The import allocates one
index per turn instead. Siblings share an index, the way SP4 leaves
them, and no two coordinates share one.
"""
copy = _imported(client, _export(client, _forked_story(client)))
rows = _rows(copy)
@@ -376,8 +380,9 @@ def test_the_legacy_index_is_reissued_so_two_branches_never_share_one(client):
"one index per turn, whatever branch it is on"
assert len(by_index) == len({(r.branch_id, r.depth) for r in rows})
# The case that makes the rule necessary: the fork and the line it left
# both hold a turn at depth 2, and they are not the same turn.
# This is the case that makes the rule necessary: the fork and the
# branch it left both hold a turn at depth 2, and they are not the
# same turn.
at_depth_2 = [r for r in rows if r.depth == 2]
assert len({r.branch_id for r in at_depth_2}) == 2
assert len({r.index for r in at_depth_2}) == 2, "same depth, different turns"
@@ -386,8 +391,9 @@ def test_the_legacy_index_is_reissued_so_two_branches_never_share_one(client):
# ------------------------------------------------------- a file that is wrong
def test_a_node_naming_a_branch_the_file_does_not_list_is_refused(client):
"""Refused, not half-applied. A tree missing a branch is a story that
silently stops, which is the failure this whole phase exists to end."""
"""The import must refuse this file rather than half-apply it. A tree
missing a branch is a story that silently stops, which is the failure
this whole phase exists to prevent."""
payload = _export(client, _forked_story(client))
payload["branches"] = payload["branches"][:1]
before = _adventure_count()
@@ -399,8 +405,9 @@ def test_a_node_naming_a_branch_the_file_does_not_list_is_refused(client):
def test_a_branch_forking_from_one_listed_after_it_is_refused(client):
"""The ordering rule buys acyclicity for the price of a comparison — and a
cycle would be an import that never returns rather than one that fails."""
"""The ordering rule guarantees no cycles at the cost of one
comparison. Without it, a cycle would produce an import that never
returns instead of one that fails cleanly."""
payload = _export(client, _forked_story(client))
payload["branches"] = [{"parent": 1, "forkDepth": 0}, {"parent": None, "forkDepth": None}]
before = _adventure_count()
@@ -439,8 +446,9 @@ def test_more_branches_than_the_cap_is_refused(client, monkeypatch):
def test_a_turn_the_file_gives_no_live_attempt_still_tells_one(client):
"""A coordinate with nothing live is a turn no read can see.
The file is allowed to be wrong about this — it is a text file — so the
import picks the first attempt rather than importing a story with a hole.
The file is allowed to be wrong about this, since it is a text file
someone can edit. The import picks the first attempt instead of
importing a story with a gap.
"""
payload = _export(client, _forked_story(client))
for node in payload["actions"]:
@@ -456,7 +464,8 @@ def test_a_turn_the_file_gives_no_live_attempt_still_tells_one(client):
# ------------------------------------------------------------- the v1 reader
def test_a_v1_bundle_still_imports(client):
"""The reader stays after the writer goes: those files are already saved."""
"""The v1 reader must remain even after the v1 writer is gone, because
those files already exist and are saved."""
payload = {
"format": bundle.LEGACY_FORMAT,
"title": "Old backup",
@@ -474,8 +483,8 @@ def test_a_v1_bundle_still_imports(client):
copy = _imported(client, payload)
assert _texts(client, copy) == [OPENING, "> You go north.", "Two."]
# One branch, and the `variants` array split back into the sibling group it
# always described.
# The import produces one branch, and the `variants` array splits
# back into the sibling group it always described.
assert len(_branches(client, copy)) == 1
ai = [r for r in _rows(copy) if r.type == "ai"]
assert [(r.text, r.live) for r in ai] == [("One.", False), ("Two.", True)]
@@ -483,8 +492,9 @@ def test_a_v1_bundle_still_imports(client):
def test_a_v1_bundle_with_a_cursor_lands_it_on_a_node(client):
"""v1 counts covered actions; the tree anchors them. The translation needs
the nodes to exist, so it happens after they are written."""
"""v1 counts covered actions. The tree anchors them to a node instead.
The translation needs the nodes to exist first, so it runs after they
are written."""
payload = {
"format": bundle.LEGACY_FORMAT, "title": "Old backup",
"memoryCursor": 2, "summaryCursor": 2,
@@ -507,7 +517,8 @@ def test_a_v1_bundle_with_a_cursor_lands_it_on_a_node(client):
def test_a_v2_bundle_brings_its_anchors_back(client):
"""The other direction: v2 carries the anchor and the count is read off it."""
"""The other direction: v2 carries the anchor directly, and the legacy
count is derived from it."""
original = _forked_story(client)
tip = [a for a in _rows(original) if a.live][-1]
db = SessionLocal()
@@ -535,13 +546,14 @@ def test_a_v2_bundle_brings_its_anchors_back(client):
def test_a_v1_memory_that_summarises_nothing_lands_on_the_root(client):
"""The import has to answer the question migration 62 answered.
"""The import must answer the question migration 62 already answered.
A v1 file has no depths, and a memory the player typed has no `sourceEnd`
to derive one from — so it used to come back with a NULL depth, which is the
exact state SP7 removed from the schema. `Path._entry_clause` compares
`depth <= max_depth` and a NULL fails it, so the memory would read fine
until the imported adventure was forked and then vanish from the new branch.
A v1 file has no depths, and a memory the player typed has no
`sourceEnd` to derive one from. It used to come back with a NULL
depth, the exact state SP7 removed from the schema.
`Path._entry_clause` compares `depth <= max_depth`, and a NULL fails
that comparison. The memory would read fine until the imported
adventure forked, then vanish from the new branch.
"""
payload = {
"format": bundle.LEGACY_FORMAT, "title": "Old backup",
@@ -572,12 +584,13 @@ def test_a_v1_memory_that_summarises_nothing_lands_on_the_root(client):
def test_the_action_cap_counts_the_rows_a_v1_file_expands_into(client, monkeypatch):
"""The cap has to count what gets written, not what the file lists.
"""The cap must count what gets written, not what the file lists.
A v1 turn carries its retries in a `variants` array, and SP4 made every
attempt a row — so one entry can become ten. Counting entries lets a file
inside the cap write a multiple of it, and the body limit is no help: the
text is tiny, it is the row count that is the cost.
A v1 turn carries its retries in a `variants` array, and SP4 made
every attempt a row, so one entry can expand into ten. Counting
entries instead of rows would let a file inside the cap write a
multiple of it. The body-size limit does not help here: the text is
tiny, and the row count is the actual cost.
"""
monkeypatch.setattr(auth, "MULTI_USER", True)
monkeypatch.setattr(limits, "MAX_ACTIONS_PER_ADVENTURE", 6)
+23 -18
View File
@@ -1,7 +1,8 @@
"""HTTP tests for the AI Chat scratchpad (power users only).
Covers the access gate, the streamed reply, and the demo-key model pinning —
the part that must not let a public visitor reach paid models through this page.
Covers the access gate, the streamed reply, and the demo-key model pinning.
This pinning must not let a public visitor reach paid models through this
page.
python -m pytest tests/test_chat.py -v
"""
@@ -59,13 +60,14 @@ def client(monkeypatch):
monkeypatch.setattr(chat, "OpenAICompatibleProvider", FakeProvider)
monkeypatch.setattr(limits, "rate_limit", lambda *a, **k: None)
# Multi-user mode is what makes the power-user gate meaningful (local mode
# trusts everyone); the allowlist is set per-test.
# Multi-user mode is what makes the power-user gate meaningful, because
# local mode trusts everyone. The allowlist is set per test.
monkeypatch.setattr(auth, "MULTI_USER", True)
monkeypatch.setattr(auth, "POWER_USERS", {"power@example.com"})
# These tests deliberately do NOT stub resolve_provider_config: the point is
# to exercise the real BYOK-vs-demo decision, since that is what keeps the
# shared key off paid models. Each test picks a mode with _byok/_demo below.
# These tests deliberately do not stub resolve_provider_config. The
# point is to exercise the real BYOK-vs-demo decision, since that
# decision is what keeps the shared key off paid models. Each test
# picks a mode with _byok/_demo below.
def _current_user(db=Depends(get_db)):
return db.get(models.User, user_id)
@@ -152,8 +154,9 @@ def test_demo_key_pins_model_to_whitelist(client, monkeypatch):
def test_demo_key_ignores_an_off_whitelist_settings_model(client, monkeypatch):
"""The override isn't the only untrusted input — Settings.model is user-set
too, and it must be pinned the same way when there's no BYOK key."""
"""The override is not the only untrusted input. `Settings.model` is
also user-set, and it must be pinned the same way when there is no
BYOK key."""
_demo(monkeypatch)
db = SessionLocal()
try:
@@ -167,8 +170,8 @@ def test_demo_key_ignores_an_off_whitelist_settings_model(client, monkeypatch):
def test_demo_key_endpoint_cannot_be_redirected(client, monkeypatch):
"""A user-controlled endpoint_url would leak the key itself, which is worse
than spending it — the demo branch pins the URL too."""
"""A user-controlled `endpoint_url` would leak the key itself, which is
worse than spending it. The demo branch pins the URL too."""
_demo(monkeypatch)
db = SessionLocal()
try:
@@ -183,8 +186,8 @@ def test_demo_key_endpoint_cannot_be_redirected(client, monkeypatch):
def test_provider_config_refuses_server_funded_paid_model(monkeypatch):
"""The structural backstop: a hand-built config (a future code path that
forgets to go through resolve_provider_config) can't run a server-funded
turn on an off-whitelist model."""
forgets to go through resolve_provider_config) cannot run a
server-funded turn on an off-whitelist model."""
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_MODELS", ["free/allowed"])
with pytest.raises(ValueError):
@@ -196,9 +199,10 @@ def test_provider_config_refuses_server_funded_paid_model(monkeypatch):
def test_byok_user_may_reuse_the_demo_keys_value(client, monkeypatch):
"""Regression: the demo key is just an OpenRouter key, so a user can paste
that same value into their own Settings. That's BYOK — they're paying — and
it must not trip the guard. It used to raise on every resolution, which
500'd GET /auth/me and took the whole SPA down (no nav, no chat)."""
that same value into their own Settings. That is still BYOK, because
the user is paying, and it must not trip the guard. It used to raise
on every resolution, which returned a 500 from `GET /auth/me` and
broke the entire SPA (no nav, no chat)."""
monkeypatch.setattr(auth, "demo_enabled", lambda: True)
monkeypatch.setattr(auth, "DEMO_API_KEY", "shared-key")
monkeypatch.setattr(auth, "DEMO_ENDPOINT_URL", "http://demo")
@@ -221,14 +225,15 @@ def test_byok_user_may_reuse_the_demo_keys_value(client, monkeypatch):
def test_resolve_provider_config_is_the_single_choke_point(monkeypatch):
"""Turns, AI Chat and the connection test all resolve through this one
"""Turns, AI Chat, and the connection test all resolve through this one
function, so pinning it here pins every caller. No DB or HTTP needed."""
monkeypatch.setattr(auth, "demo_enabled", lambda: True)
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_ENDPOINT_URL", "http://demo")
monkeypatch.setattr(auth, "DEMO_MODELS", ["free/allowed"])
# No key of their own: endpoint AND model are pinned, whatever they set.
# No key of their own: both endpoint and model are pinned, regardless
# of what they set.
no_key = models.Settings(endpoint_url="http://mine/v1", api_key="", model="expensive/paid")
assert auth.resolve_provider_config(no_key) == auth.ProviderConfig(
"http://demo", "demo-key", "free/allowed", True)
+49 -42
View File
@@ -1,19 +1,19 @@
"""Guards on how much the database is asked for.
context_snapshot holds the entire assembled prompt for a turn — 163 KB a row
averaged over production, 232 KB on the longest adventure, and 89% of the
database. It used to be pulled for every action on every adventure load and
every turn, to read two tiny things out of it. These tests fail if that
regresses.
`context_snapshot` holds the entire assembled prompt for a turn. That is 163 kB
per row averaged over production, 232 kB on the longest adventure, and 89% of the
database. It used to be fetched for every action on every adventure load and
every turn, to read two small values out of it. These tests fail if that
returns.
Two kinds of guard live here, and both are needed:
* **column guards** assert which columns a statement names. That is the shape
both of this project's egress blowouts took — one query quietly carrying a
column nobody read.
* **byte ceilings** assert what a request actually costs. Every column guard
would still pass if a response grew tenfold within the columns it is allowed
to read, which is what a story that keeps getting longer does.
* Column guards assert which columns a statement names. That is the shape both
of this project's egress regressions took: one query carrying a column nobody
read.
* Byte ceilings assert what a request costs. Every column guard would still pass
if a response grew tenfold within the columns it is allowed to read, which is
what a story that keeps getting longer does.
python -m pytest tests/test_egress.py -v
"""
@@ -46,10 +46,10 @@ from tools.fakeprose import prose
# column enormous, plus the small world_state slice the UI actually needs.
#
# Varied text, not `"x" * 20_000`. The column is stored compressed now
# (migration 43), and a repeated character compresses about a thousandfold —
# which would make the byte ceilings below pass against a fixture that costs
# nothing, testing nothing. Prose-shaped filler compresses like the prompts
# this stands in for.
# by migration 43, and a repeated character compresses about a thousandfold.
# That would make the byte ceilings below pass against a fixture that costs
# nothing, which tests nothing. Prose-shaped filler compresses like the prompts
# it stands in for.
_SNAPSHOT_RNG = random.Random(20_260_817)
BIG_SNAPSHOT = {
"system": prose(_SNAPSHOT_RNG, 20_000),
@@ -149,8 +149,8 @@ def test_the_state_snapshots_are_not_fetched_in_bulk(client, sql_log):
"""All four are rollback snapshots, only ever needed for the single node
being undone, retried past or switched to.
The `_after` pair is the live one since SP4 and the `_before` pair is dead
weight until SP8 drops it — a page load must pay for neither.
The `_after` pair is the live one since SP4, and the `_before` pair is
unused until SP8 drops it. A page load must pay for neither.
"""
client.get(f"/api/adventures/{client.adv_id}")
selects = action_selects(sql_log)
@@ -195,8 +195,11 @@ def test_variant_count_survives_variants_being_deferred(client):
def test_counting_actions_does_not_name_the_deferred_columns(client, sql_log):
"""A count that wraps the entity select in a subquery names every column in
the emitted SQL — no bytes come back, but the database still reads them and
the guard above cannot tell it apart from a real bulk fetch."""
the emitted SQL.
No bytes come back, but the database still reads them, and the guard above
cannot distinguish that from a real bulk fetch.
"""
db = SessionLocal()
try:
adventure = db.get(models.Adventure, client.adv_id)
@@ -214,7 +217,8 @@ def test_counting_actions_does_not_name_the_deferred_columns(client, sql_log):
def test_snapshot_is_still_reachable_on_demand(client):
"""Deferred means lazy, not gone — Insights still gets the full thing."""
"""Deferred means lazy rather than absent. Insights still gets the whole
snapshot."""
r = client.get(f"/api/adventures/{client.adv_id}")
action_id = r.json()["actions"][0]["id"]
r = client.get(f"/api/adventures/{client.adv_id}/actions/{action_id}/context")
@@ -231,9 +235,9 @@ def as_json_snapshot_column(db) -> None:
Migration 36 lifts world_delta out of the snapshot with SQL JSON
functions, so it can only run while the column still *is* JSON. In a real
upgrade it always is — 36 runs seven migrations before 43 compresses the
column into a BLOB — but `create_all` builds today's schema, so a test
calling that backfill has to rebuild the schema it was written against.
upgrade it always is, because 36 runs seven migrations before 43 compresses
the column into a BLOB. `create_all` builds today's schema, so a test that
calls that backfill has to rebuild the schema it was written against.
"""
db.execute(text("ALTER TABLE actions DROP COLUMN context_snapshot"))
db.execute(text("ALTER TABLE actions ADD COLUMN context_snapshot JSON"))
@@ -269,9 +273,11 @@ def test_backfill_populates_world_delta_from_existing_snapshots(client):
def test_backfill_populates_variant_count_from_existing_variants(client):
"""Migration 37 counts the lists server-side — reading them into Python to
count them would mean pulling the column across the wire once to stop
pulling it across forever."""
"""Migration 37 counts the lists on the server.
Reading them into Python to count them would fetch the column over the wire
once in order to stop fetching it on every request.
"""
db = SessionLocal()
try:
db.execute(text("UPDATE actions SET variant_count = 0"))
@@ -304,20 +310,20 @@ def test_backfill_leaves_actions_without_world_state_alone(client):
# ---------------------------------------------------------------- byte ceilings
#
# The tests above assert which *columns* a statement names, which is the shape
# both of this project's egress blowouts took. They would all still pass if a
# response quietly grew tenfold within the columns it is allowed to read — and
# a story that keeps getting longer does exactly that. These put a number on it.
# The tests above assert which columns a statement names, which is the shape both
# of this project's egress regressions took. They would all still pass if a
# response grew tenfold within the columns it is allowed to read, and a story
# that keeps getting longer does that. These tests put a number on it.
#
# Ceilings are per action rather than absolute, so they mean the same thing
# whatever size the fixture is set to, and they are generous: the point is to
# catch a tenfold regression, not to freeze today's byte count.
# The ceilings are per action rather than absolute, so they mean the same thing
# whatever size the fixture is, and they are generous. They exist to catch a
# tenfold regression rather than to freeze today's byte count.
ACTIONS_IN_FIXTURE = 12
# 3 kB an action against a real 994 B, measured on production 2026-08-17.
# Anything that pulls a deferred column blows past this by two orders of
# magnitude — see test_the_ceiling_discriminates below.
# magnitude. See `test_the_ceiling_discriminates` below.
PAGE_LOAD_BYTES_PER_ACTION = 3_000
@@ -325,8 +331,9 @@ PAGE_LOAD_BYTES_PER_ACTION = 3_000
def meter():
"""A byte meter on the shared engine, removed again afterwards.
Requested *after* `client` in a test's arguments so that building the
fixture — a write path nobody plays — is not charged to any scope.
A test requests this after `client` in its arguments, so that building the
fixture, which is a write path no player takes, is not charged to any
scope.
"""
m = dbmeter.Meter()
m.attach(engine)
@@ -364,8 +371,8 @@ def test_the_action_list_stays_under_its_byte_ceiling(client, meter):
def test_reading_one_action_does_not_cost_the_whole_story(client, meter):
"""The snapshot is reachable on demand, and that request should pay for
one row's worth — not the adventure's."""
"""The snapshot is reachable on demand, and that request pays for one row
rather than for the whole adventure."""
db = SessionLocal()
try:
action_id = db.query(models.Action.id).order_by(models.Action.id).first()[0]
@@ -459,10 +466,10 @@ def test_the_ceiling_discriminates(client, meter):
budget. If this ever stops exceeding it, the fixture has gone too small for
the tests above to mean anything.
The margin used to be a hundredfold and is now about six. That is not the
guard weakening — it is migration 43 compressing the column, and the
fixture text being prose-shaped so it compresses like a real prompt rather
than like a repeated character.
The margin used to be a hundredfold and is now about six. The guard has not
weakened. Migration 43 compresses the column, and the fixture text is
prose-shaped, so it compresses like a real prompt rather than like a repeated
character.
"""
budget = ACTIONS_IN_FIXTURE * PAGE_LOAD_BYTES_PER_ACTION
db = SessionLocal()
+28 -26
View File
@@ -1,9 +1,9 @@
"""Migration 38: embeddings move from a JSON list to packed float32.
The conversion has to be exact, because nothing re-embeds — a memory whose
vector shifts is silently ranked wrong forever, with no error anywhere to say
so. So these tests check the numbers survive the round trip bit for bit, and
that the migration reaches every row however many there are.
The conversion has to be exact, because nothing re-embeds. A memory whose
vector shifts is silently ranked wrong forever, with no error anywhere to
report it. These tests check that the numbers survive the round trip bit
for bit, and that the migration reaches every row, however many there are.
python -m pytest tests/test_embedding_blob.py -v
"""
@@ -29,8 +29,8 @@ from tests import schema_rewind
def float32(value: float) -> float:
"""`value` as the double nearest to its float32 truncation — what an
embedding endpoint's JSON actually holds."""
"""`value` as the double nearest to its float32 truncation. This is
what an embedding endpoint's JSON actually holds."""
return struct.unpack("<f", struct.pack("<f", value))[0]
@@ -71,7 +71,7 @@ def test_pack_round_trips_exactly():
def test_packed_vector_is_four_bytes_per_dimension():
"""The whole point: 1536 dims is 6 KB here against ~31 KB as JSON."""
"""1536 dims takes 6 KB packed, against about 31 KB as JSON."""
vector = sample_vector(random.Random(2))
blob = vectors.pack(vector)
assert len(blob) == 1536 * 4
@@ -84,9 +84,9 @@ def test_pack_handles_the_extremes():
def test_unpack_returns_a_compact_array():
"""These are held in memory between turns, so the container matters: an
array("f") is the 4 bytes a component the column is, a list of Python
floats is eight times that."""
"""These vectors stay in memory between turns, so the container type
matters. An array("f") stores each component in 4 bytes, the same width
as the column. A list of Python floats uses eight times that."""
vector = sample_vector(random.Random(9))
unpacked = vectors.unpack(vectors.pack(vector))
assert unpacked.typecode == "f"
@@ -95,15 +95,16 @@ def test_unpack_returns_a_compact_array():
def test_pack_rounds_a_value_float32_cannot_hold():
"""The guarantee is exactness for vectors that came from an embedding
model, which computes in float32 — not for arbitrary doubles. Worth
pinning down, because it is the line the round-trip claim sits on."""
"""The exactness guarantee applies only to vectors that came from an
embedding model, which computes in float32, not to arbitrary doubles.
This test pins down that boundary, because it is where the round-trip
claim holds."""
assert vectors.unpack(vectors.pack([1e-38]))[0] != 1e-38
assert vectors.unpack(vectors.pack([1e-38]))[0] == pytest.approx(1e-38)
def test_cosine_moved_but_still_reachable_from_memorybank():
"""Callers import it from memorybank; the maths lives in vectors."""
"""Callers import it from memorybank. The math lives in vectors."""
assert memorybank.cosine is vectors.cosine
assert vectors.cosine([1.0, 0.0], [1.0, 0.0]) == pytest.approx(1.0)
assert vectors.cosine([1.0, 0.0], [0.0, 1.0]) == pytest.approx(0.0)
@@ -149,11 +150,11 @@ def test_set_vector_none_clears_both(db, adventure):
def add_legacy_json_column(db) -> None:
"""Put `memories.embedding` back for the length of a test.
Migration 42 dropped it and the model no longer declares it, so
`create_all` does not produce it — but everything below is testing the
upgrade *from* a database that still has it, which is the only state in
which the backfill has any work to do. Re-adding it by hand is what keeps
these tests honest about the schema they claim to be starting from.
Migration 42 dropped it, and the model no longer declares it, so
`create_all` does not produce it. Everything below tests the upgrade
from a database that still has the column, which is the only state
where the backfill has any work to do. Re-adding it by hand keeps these
tests honest about the schema they claim to start from.
"""
db.execute(text("ALTER TABLE memories ADD COLUMN embedding JSON"))
db.commit()
@@ -193,8 +194,9 @@ def test_backfill_converts_every_existing_vector(db, adventure):
def test_backfill_reaches_past_one_batch(db, adventure):
"""It loops on id, and an off-by-one there would silently leave the tail
of a big bank unconverted — which reads as "not embedded yet"."""
"""It loops on id, and an off-by-one there would silently leave the
tail of a big bank unconverted. That failure reads as "not embedded
yet"."""
count = migrations.BACKFILL_BATCH * 2 + 3
expected = seed_json_only(db, adventure, count=count, dims=4)
@@ -237,8 +239,8 @@ def test_backfill_is_idempotent(db, adventure):
def test_backfill_skips_a_malformed_row_without_stopping(db, adventure):
"""One bad row must not strand every row after it — the loop orders by id,
so an exception here would leave the rest of the bank unconverted."""
"""One bad row must not strand every row after it. The loop orders by
id, so an exception here would leave the rest of the bank unconverted."""
expected = seed_json_only(db, adventure, count=2)
broken = models.Memory(adventure_id=adventure.id, text="broken")
db.add(broken)
@@ -288,9 +290,9 @@ def test_bootstrap_adds_the_columns_and_backfills_them(db, adventure):
# embedded must still read as not embedded afterwards.
assert by_id[unembedded_id] == (None, False)
# ...and migration 42, at the end of the same run, takes the JSON column
# away. Ordering matters: 38 reads it, 42 drops it, and an upgrade that
# ran them the other way round would arrive with an empty bank.
# Migration 42, at the end of the same run, removes the JSON column.
# Ordering matters: 38 reads it, 42 drops it, and an upgrade that ran
# them in the other order would arrive with an empty bank.
with engine.begin() as conn:
columns = {row[1] for row in conn.execute(text("PRAGMA table_info(memories)"))}
assert "embedding" not in columns
+18 -17
View File
@@ -1,21 +1,21 @@
"""Switching embedding models must re-embed the bank.
Vectors from two different models are not comparable — different space, often
different width — so changing the model has to throw the stored ones away and
let the post-turn pass rebuild them.
Vectors from two different models are not comparable. They live in
different spaces and often have different widths. Changing the model must
discard the stored vectors and let the post-turn pass rebuild them.
That worked while the vectors lived in `memories.embedding`: the settings
route nulled that column and the embed queue picked the rows up. Migration 38
moved the vectors to `embedding_blob` with an `embedded` flag beside them, and
the bulk clear kept nulling the old column alone. The blob survived, the flag
stayed true, `_embed_pending` (which looks for `embedded IS FALSE`) never saw
the rows, and the bank went on ranking against the previous model's vectors
for good.
This worked while the vectors lived in `memories.embedding`. The settings
route nulled that column, and the embed queue picked up the rows. Migration
38 moved the vectors to `embedding_blob` and added an `embedded` flag beside
them, but the bulk clear kept nulling only the old column. The blob
survived, the flag stayed true, and `_embed_pending` (which filters on
`embedded IS FALSE`) never saw the rows. The bank kept ranking against the
previous model's vectors.
Nothing reports this. `cosine` returns 0.0 on a width mismatch, so a
different-width model scores every memory zero and retrieval quietly returns
whichever rows sort first; a same-width model scores plausible-looking
garbage.
Nothing reports this failure. `cosine` returns 0.0 on a width mismatch, so a
different-width model scores every memory zero, and retrieval silently
returns whichever rows sort first. A same-width model scores plausible
garbage instead.
python -m pytest tests/test_embedding_model_switch.py -v
"""
@@ -114,8 +114,9 @@ def test_changing_the_model_clears_every_vector(client):
def test_cleared_memories_are_queued_for_re_embedding(client):
"""The flag is not cosmetic: it is the only thing `_embed_pending` filters
on, so this is the assertion that the bank actually recovers."""
"""The `embedded` flag is not cosmetic. It is the only condition
`_embed_pending` filters on, so this test confirms the bank actually
recovers."""
client.put("/api/settings", json={"embedding_model": "model-b"})
db = SessionLocal()
@@ -155,7 +156,7 @@ def test_retrieval_uses_no_stale_vector_after_the_switch(client, monkeypatch):
def test_an_unrelated_settings_change_keeps_the_vectors(client):
"""Only an embedding-model change may clear the bank — re-embedding costs
"""Only an embedding-model change may clear the bank. Re-embedding costs
an API call per memory."""
r = client.put("/api/settings", json={"model": "some-other-chat-model"})
assert r.status_code == 200, r.text
+15 -13
View File
@@ -1,7 +1,8 @@
"""Guest retention policy — app/cleanup.py.
"""Guest retention policy: app/cleanup.py.
Covers the two things that matter: that idle guests and their whole data
graph actually go, and that nothing else ever does.
These tests cover the two things that matter. Idle guests and their whole
data graph must actually be deleted, and nothing else must ever be
deleted.
python -m pytest tests/test_guest_cleanup.py -v
"""
@@ -30,8 +31,8 @@ def db(tmp_path):
@event.listens_for(engine, "connect")
def _fk(dbapi_connection, _record):
# The whole policy leans on ON DELETE CASCADE; SQLite ignores every
# one of them unless this is set (same as database.py does).
# The whole policy relies on ON DELETE CASCADE. SQLite ignores every
# one of them unless this is set, the same as database.py does.
cur = dbapi_connection.cursor()
cur.execute("PRAGMA foreign_keys=ON")
cur.close()
@@ -86,7 +87,7 @@ def test_keeps_guest_inside_the_window(db):
def test_boundary_is_not_yet_stale(db):
# Exactly 5 days survives; the comparison is strict.
# Exactly 5 days survives. The comparison is strict.
user = make_user(db, days_idle=cleanup.RETENTION_DAYS)
assert sweep(db) == 0
assert alive(db, user.id)
@@ -112,8 +113,9 @@ def test_recent_visit_beats_an_old_created_at(db):
# ---------- what must never go ----------
def test_spares_registered_users(db):
"""Registering upgrades the guest row in place, so an idle account here is
a real user with real data — the whole point of signing up."""
"""Registering upgrades the guest row in place. An idle account here is
a real user with real data, which is what signing up is meant to
protect."""
user = make_user(db, days_idle=400, guest=False, email="a@b.com")
assert sweep(db) == 0
assert alive(db, user.id)
@@ -127,7 +129,7 @@ def test_spares_the_local_mode_user(db):
def test_spares_a_guest_flagged_row_that_has_an_email(db):
# Shouldn't exist, but both clauses are checked so it can't be collected.
# This row should not exist, but both clauses are checked so it cannot be collected.
user = make_user(db, days_idle=400, guest=True, email="odd@b.com")
assert sweep(db) == 0
assert alive(db, user.id)
@@ -164,10 +166,10 @@ def test_enabled_requires_multi_user(monkeypatch):
# ---------- the cascade ----------
def test_deletes_the_whole_data_graph(db):
"""One DELETE has to take the adventure, its actions and memories, the
story cards and the settings row with it — nothing is loaded into Python,
so if the FK cascade isn't reaching, rows are silently orphaned (or the
statement errors) rather than tidied."""
"""One DELETE must remove the adventure, its actions and memories, the
story cards, and the settings row. Nothing is loaded into Python, so if
the FK cascade does not reach a table, its rows are silently orphaned,
or the statement fails, instead of being removed."""
user = make_user(db, days_idle=30)
scenario = models.Scenario(user_id=user.id, title="S")
db.add(scenario)
+35 -31
View File
@@ -1,17 +1,18 @@
"""The context builder reads a window of the story, not all of it.
Walking `adventure.actions` every turn made a turn cost O(story length), so a
long adventure read hundreds of KB to use the tail of it — and the cost grew
with every turn played. `app.context.history` serves tails, slices and counts
from SQL instead.
Walking `adventure.actions` every turn made the turn cost O(story length).
A long adventure read hundreds of KB to use only the tail of it, and the
cost grew with every turn played. `app.context.history` serves tails,
slices, and counts from SQL instead.
Two things have to hold, and both are easy to break by accident:
Two things must hold, and both are easy to break by accident:
* the window must produce **exactly** the prompt the full story produced, or
this is a behaviour change wearing an optimization's clothes;
* the helpers must agree with the old list arithmetic, because memorybank's
cursors are *positions* in that list and a cursor off by one silently
summarizes the wrong actions.
* The window must produce exactly the prompt the full story produced.
Otherwise, the change alters behavior even though it looks like a pure
optimization.
* The helpers must agree with the old list arithmetic, because
memorybank's cursors are positions in that list. A cursor off by one
silently summarizes the wrong actions.
python -m pytest tests/test_history_window.py -v
"""
@@ -92,16 +93,16 @@ def story():
def full_window(adventure, budget_tokens, token_counter, exclude_action_id=None):
"""Stand-in for window_covering that hands back the entire story, i.e. the
behaviour this module replaced."""
"""Stand-in for `window_covering` that returns the entire story. This is
the behavior this module replaced."""
return history.story_actions(adventure, exclude_action_id)
@pytest.fixture()
def actions_loaded():
"""Counts Action rows the ORM materializes, i.e. how much of the story was
actually fetched. rowcount is meaningless for SELECT on SQLite, so count
the objects the mapper builds instead."""
"""Counts the `Action` rows the ORM materializes, which shows how much of
the story was actually fetched. `rowcount` is meaningless for a SELECT
on SQLite, so this counts the objects the mapper builds instead."""
loaded = {"n": 0}
def on_load(target, context):
@@ -133,7 +134,7 @@ def test_window_builds_the_same_prompt_as_the_whole_story(story, budget, monkeyp
def test_window_matches_on_the_retry_shape(story, monkeypatch):
"""Retry excludes the action being regenerated; the exclusion has to reach
"""Retry excludes the action being regenerated. The exclusion must reach
the window query, not just the in-memory filter."""
db, adventure, settings = story
last = history.tail(adventure, 1)[0]
@@ -148,7 +149,8 @@ def test_window_matches_on_the_retry_shape(story, monkeypatch):
def test_reported_total_is_the_whole_story_not_the_window(story):
"""Insights says "N of M actions included"; M must not become the window."""
"""Insights reports "N of M actions included." M must not become the
window size."""
db, adventure, settings = story
settings.context_token_budget = 4096
report = builder.build_context(adventure, settings)[2]
@@ -160,7 +162,7 @@ def test_reported_total_is_the_whole_story_not_the_window(story):
def test_building_context_reads_far_less_than_the_whole_story(story, actions_loaded):
db, adventure, settings = story
# Expire first: expiring afterwards would discard the unflushed change and
# Expire first. Expiring afterward would discard the unflushed change and
# silently put the budget back to its default.
db.expire_all()
# Small enough that the budget, not the length of the story, decides.
@@ -171,9 +173,10 @@ def test_building_context_reads_far_less_than_the_whole_story(story, actions_loa
included = report["history"]["included"]
assert included < ACTION_COUNT, "fixture is too short to prove anything"
# The window aims a margin past the budget and re-asks if it fell short, so
# it reads somewhat more than it includes. What matters is that the read is
# a function of the token budget, not of how long the story has got.
# The window targets a margin past the budget and requests more if it
# falls short, so it reads somewhat more than it includes. What matters
# is that the read depends on the token budget, not on the length of
# the story.
assert actions_loaded["n"] < ACTION_COUNT // 2, (
f"read {actions_loaded['n']} action rows out of {ACTION_COUNT} to "
f"include {included} — the window is not bounding the read"
@@ -213,13 +216,14 @@ def test_helpers_agree_with_the_full_list(story):
def test_a_depth_boundary_survives_a_middle_action_being_deleted(story):
"""The case that has broken the cursors twice before, and the reason they
are depths now.
"""The case that broke the cursors twice before. This is the reason the
cursors are depths now.
A *position* answers "how much story is past this point?" by counting from
the start, so deleting anything in front of the mark changes which action
the mark names. A depth names the same node either way — the only thing
that changes is the count of what comes after, which is what did change.
A position answers "how much story is past this point?" by counting
from the start. Deleting anything in front of the mark changes which
action the mark names. A depth names the same node either way. The
only thing that changes is the count of what comes after, and that
count is the one thing that should change here.
"""
db, adventure, settings = story
actions = history.story_actions(adventure)
@@ -236,9 +240,9 @@ def test_a_depth_boundary_survives_a_middle_action_being_deleted(story):
assert history.count_after(adventure, mark) == before, "the mark moved"
assert [a.id for a in history.after(adventure, mark, 3)] == next_three
# ...and deleting something *after* it is the one thing that does change
# the count, because that is a fact about the story rather than about the
# coordinate system.
# Deleting something after the mark is the one change that does affect
# the count, because that count reflects the story, not the coordinate
# system.
db.delete(history.after(adventure, mark, 1)[0])
db.commit()
db.expire(adventure)
@@ -258,7 +262,7 @@ def test_blank_actions_are_excluded_the_same_way_in_sql_and_python(story):
# SQL path (relationship not loaded)
from_sql = history.count(adventure)
# Python path (relationship loaded)
adventure.actions # noqa: B018 — force the collection into memory
adventure.actions # noqa: B018 - force the collection into memory
from_python = history.count(adventure)
assert from_sql == from_python == ACTION_COUNT
+35 -30
View File
@@ -1,17 +1,17 @@
"""The turn prompt asks for a turn that fits inside `max_output_tokens`.
`max_output_tokens` is a hard wall the endpoint enforces mid-sentence. The state
block is emitted *after* the narration, so a long turn hits the wall partway
through the block and the deltas are lost — silently, since nothing reads
`finish_reason`. The prompt now carries a word budget derived from the cap so
the model lands just inside it.
`max_output_tokens` is a hard limit the endpoint enforces mid-sentence. The
state block is emitted after the narration, so a long turn hits the limit
partway through the block, and the deltas are lost. Nothing reads
`finish_reason`, so this loss happens silently. The prompt now carries a
word budget derived from the cap, so the model lands just inside it.
Two things are easy to break here:
* the hint must be stated in **words**, not tokens — a model cannot count its
own tokens, and a hint it cannot follow is just wasted budget;
* it must not displace `EMIT_REMINDER` from the last position, which is the
whole mechanism keeping the state block emitted at all (see
* The hint must be stated in words, not tokens. A model cannot count its
own tokens, and a hint it cannot follow is wasted budget.
* The hint must not displace `EMIT_REMINDER` from the last position, which
is the whole mechanism that keeps the state block emitted at all (see
test_worldstate.py and the emit-reliability fix).
python -m pytest tests/test_length_hint.py -v
@@ -100,7 +100,7 @@ def asked_words(cap):
def test_buffer_leaves_room_for_overshoot():
"""The stated number must sit meaningfully under the real ceiling, or an
on-target-but-slightly-long turn still hits the wall."""
on-target-but-slightly-long turn still hits the limit."""
for cap in (400, 800, 1500, 2400):
asked = asked_words(cap)
ceiling = (cap - builder.LENGTH_HEADROOM) * builder.WORDS_PER_TOKEN
@@ -109,9 +109,10 @@ def test_buffer_leaves_room_for_overshoot():
def test_hint_is_phrased_as_a_ceiling_not_a_budget():
"""Measured: budget phrasing ("keep this turn under about N words") reads as a
target to fill and moved the mean turn from 174 to 246 words — toward the wall
it exists to avoid. The limit framing must survive future prompt edits."""
"""Measured: budget phrasing ("keep this turn under about N words")
reads as a target to fill. It moved the mean turn from 174 to 246
words, toward the limit it exists to avoid. The limit framing must
survive future prompt edits."""
hint = builder.length_hint(800, has_ws=True)
assert "must not exceed" in hint
assert "under about" not in hint
@@ -119,15 +120,15 @@ def test_hint_is_phrased_as_a_ceiling_not_a_budget():
def test_hint_states_a_floor_as_well_as_a_ceiling():
"""A ceiling alone is one-sided: a terse model has nothing to act on but the
"only as much as the moment needs" clause and collapses to two paragraphs.
The floor is what makes the same prompt land in the same place across models
that lean opposite ways."""
"""A ceiling alone is one-sided: a terse model has nothing to act on but
the "only as much as the moment needs" clause and produces only two
paragraphs. The floor is what makes the same prompt produce a similar
length across models with different tendencies."""
hint = builder.length_hint(800, has_ws=True)
assert "506" in hint and "177" in hint
assert "should not stop short of" in hint
# Asymmetric on purpose: the wall is a wall, the floor is a floor, and neither
# is phrased as a number to hit.
# Asymmetric on purpose: the ceiling is a hard limit and the floor is a
# soft target, and neither is phrased as a specific number to reach.
assert hint.index("must not exceed") < hint.index("should not stop short of")
@@ -139,9 +140,10 @@ def test_floor_stays_well_under_the_ceiling():
def test_floor_is_dropped_when_the_cap_is_too_tight_for_one():
"""At a tight cap a short turn is the correct turn, and the tight-cap wording
is the one measured to keep the state block alive (0/6 truncations at cap 250
against 2/6 unhinted) — so it is left exactly as it was."""
"""At a tight cap a short turn is the correct turn. The tight-cap
wording is the one measured to keep the state block from being
truncated (0/6 truncations at cap 250, against 2/6 unhinted), so it
is left exactly as it was."""
hint = builder.length_hint(250, has_ws=True)
assert "should not stop short of" not in hint
assert "much shorter" in hint
@@ -199,13 +201,15 @@ def test_emit_reminder_keeps_the_last_word(story):
def test_prompt_stays_inside_the_budget_on_a_long_story(story):
"""Regression guard: the hint is appended after history has already spent
the budget, so it must be reserved up front like EMIT_REMINDER is.
"""Regression guard: the hint is appended after history has already
spent the budget, so it must be reserved up front like `EMIT_REMINDER`
is.
Weak on purpose — the history loop stops *before* crossing its budget, so it
leaves about one action of slack and the ~30-token hint hides inside it.
This catches a hint that grows large, not a missing reservation; the
reservation itself is not observable from the outside."""
This check is weak on purpose. The history loop stops before crossing
its budget, so it leaves about one action of slack, and the roughly
30-token hint fits inside that slack. This catches a hint that grows
large, not a missing reservation. The reservation itself is not
observable from the outside."""
db, adventure, settings, _ = story
for i in range(4, 120):
db.add(models.Action(
@@ -225,8 +229,9 @@ def test_prompt_stays_inside_the_budget_on_a_long_story(story):
def test_hint_is_counted_in_the_reported_totals(story):
"""Insights reports what the turn actually costs; a section that reaches the
model but not the accounting makes that number a lie."""
"""Insights reports what the turn actually costs. A section that
reaches the model but not the accounting makes that reported cost
inaccurate."""
db, adventure, settings, _ = story
settings.max_output_tokens = 800
+79 -70
View File
@@ -1,20 +1,21 @@
"""Phase 14 SP3 — memories hang off nodes, and the marks are nodes too.
"""Phase 14 SP3: memories attach to nodes, and the marks are nodes too.
Two claims, and neither of them fails loudly if it is wrong:
Two claims, and neither fails loudly if it is wrong:
* **A memory belongs to the path that produced it.** A memory made on branch B
* A memory belongs to the path that produced it. A memory made on branch B
must be invisible from A, and the memories of a shared ancestor must be
visible from both — without anything being copied when a fork happens. The
failure mode is a prompt quietly carrying a summary of a story the player
abandoned.
* **Retrieval reads the *whole* lineage, and that stays affordable.** The story
is read through a window, but recall is long-range by definition and cannot
be — so the clause names every ancestor, and the bet is that memories are
sparse enough (one per six actions) for that to be tens of small rows even
twenty forks deep. Measured below rather than asserted.
visible from both, without anything being copied when a fork happens. The
failure mode is a prompt that quietly carries a summary of a story the
player abandoned.
* Retrieval reads the whole lineage, and that stays affordable. The story
is read through a window, but recall is long-range by definition and
cannot use one. So the clause names every ancestor. The bet is that
memories are sparse enough, one per six actions, for that to stay tens of
small rows even twenty forks deep. This file measures that bet below
rather than asserting it.
Nothing in the product forks yet, so the fork is built by hand, exactly as
`test_branch_clause.py` builds it.
Nothing in the product forks yet, so this file builds the fork by hand,
exactly as `test_branch_clause.py` builds it.
python -m pytest tests/test_memory_nodes.py -v
"""
@@ -50,9 +51,10 @@ class StubEmbedder:
# --------------------------------------------------------------- the fixture
def make_branch(db, adventure, parent=None, fork_depth=None):
"""A branch row whose lineage is its parent's, capped, plus itself — the
computation SP5 will do at fork time, written out so the fixture cannot
pass by agreeing with a bug in the code under test."""
"""A branch row whose lineage is its parent's lineage, capped, plus
itself. This is the computation SP5 performs at fork time. The fixture
reimplements it here so it cannot pass by agreeing with a bug in the
code under test."""
branch = models.Branch(
adventure_id=adventure.id,
parent_branch_id=parent.id if parent else None,
@@ -106,12 +108,13 @@ def add_memory(db, adventure, text, node, vector=(1.0, 0.0, 0.0), **kwargs):
@pytest.fixture()
def forked():
"""A0..A3, then B4 B5 off A3, then C6 C7 off B5 — with a memory hung off
one node of each branch, and A playing on past the fork it was left at.
"""A0..A3, then B4 B5 off A3, then C6 C7 off B5, with a memory attached
to one node of each branch. A keeps playing past the fork point where B
left it.
The head is C, so the story is A0 A1 A2 A3 B4 B5 C6 C7 and the memories in
play are A's and B's and C's — but not the one on A5, which is on a sibling
of B4 and belongs to a story nobody is reading.
The head is C, so the story is A0 A1 A2 A3 B4 B5 C6 C7, and the
memories in play are A's, B's, and C's. The one on A5 is excluded: it
is on a sibling of B4 and belongs to a story nobody is reading.
"""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
@@ -181,8 +184,9 @@ def retrieved(adventure, settings) -> set[str]:
# ------------------------------------------------------------- the isolation
def test_a_memory_on_a_sibling_is_not_retrieved(forked):
"""The whole point. A5 is a node of the story that was abandoned when B
forked, and the memory hanging off it must not reach a prompt on C."""
"""The whole point of this file. A5 is a node of the story that was
abandoned when B forked, and the memory attached to it must not reach a
prompt on C."""
db, adventure, settings, ids = forked
assert retrieved(adventure, settings) == {
"on the shared trunk", "on B", "on C"
@@ -196,30 +200,30 @@ def test_a_shared_ancestor_is_visible_from_both_branches(forked):
switch_to(db, adventure, ids["a"], 5)
from_a = retrieved(adventure, settings)
assert "on the shared trunk" in from_a
# ...and from A, the branches taken off it are the ones out of reach.
# From A, the branches taken off it are the ones out of reach.
assert from_a == {"on the shared trunk", "on A's own continuation"}
def test_the_lineage_is_read_whole_not_windowed(forked):
"""The story is read through a window; recall is not. The trunk memory is
four nodes and two forks back, and is still a candidate."""
"""The story is read through a window. Recall is not. The trunk memory
is four nodes and two forks back, and is still a candidate."""
db, adventure, settings, ids = forked
path = lineage.path_of(db, adventure)
assert len(path) == 3
# The window a *story* read would use here names one entry. Retrieval names
# all three, which is the difference this test exists to pin.
# The window a story read would use here names one entry. Retrieval
# names all three, which is the difference this test checks.
assert path.prefix_covering(2) == 1
assert "on the shared trunk" in retrieved(adventure, settings)
def test_a_hand_written_memory_is_anchored_where_it_was_typed(forked):
"""SP7: a typed memory takes the head, so it obeys the same rule as a
summarised one.
summarized one.
It used to carry no depth, which sounded like "belongs to the whole
adventure" and behaved like "cannot be capped at a fork" — it followed the
reader onto branches whose story it never described. Anchoring it makes the
bank answer one question rather than two.
adventure" and behaved like "cannot be capped at a fork." It followed
the reader onto branches whose story it never described. Anchoring it
makes the bank answer one question instead of two.
"""
db, adventure, settings, ids = forked
switch_to(db, adventure, ids["a"], 5)
@@ -228,14 +232,14 @@ def test_a_hand_written_memory_is_anchored_where_it_was_typed(forked):
def test_a_typed_memory_survives_a_fork_of_the_ground_it_was_typed_on(forked):
"""The half of the old behaviour that was right, kept.
"""The half of the old behavior that was correct, kept.
Typed on the shared trunk it is still there after forking away — but
because the fork's path goes through that node, not because the memory was
exempt from being capped.
A memory typed on the shared trunk is still there after forking away,
but only because the fork's path goes through that node, not because
the memory is exempt from being capped.
"""
db, adventure, settings, ids = forked
switch_to(db, adventure, ids["a"], 3) # the trunk B, and so C, branch from
switch_to(db, adventure, ids["a"], 3) # the node B, and so C, forked from
add_memory(db, adventure, "typed on the trunk", None)
switch_to(db, adventure, ids["c"], 7)
@@ -243,12 +247,12 @@ def test_a_typed_memory_survives_a_fork_of_the_ground_it_was_typed_on(forked):
def test_a_typed_memory_does_not_follow_you_onto_a_path_it_is_not_on(forked):
"""And the half that was wrong, fixed.
"""The other half of the old behavior, which was wrong, is now fixed.
A5 is A's own continuation past the point B left it, so it is a sibling of
the story C tells — precisely where the `sibling` memory sits, and excluded
for precisely the same reason. Typing rather than summarising buys no
exemption from the path.
A5 is A's own continuation past the point where B left it, so it is a
sibling of the story C tells. This is exactly where the `sibling`
memory sits, and it is excluded for the same reason. Typing a memory
instead of summarizing it grants no exemption from the path rule.
"""
db, adventure, settings, ids = forked
switch_to(db, adventure, ids["a"], 5)
@@ -271,10 +275,10 @@ def test_a_mark_moves_to_the_node_the_memory_covers(forked):
def test_a_mark_from_a_sibling_reads_as_nothing_covered(forked):
"""A mark is a node, so moving to another story has to be answered rather
than assumed. Ground this path never travelled is not covered ground, and
the fallback for 'I don't know' has to be redoing the work, not skipping
it."""
"""A mark is a node, so switching to another story must resolve the
mark's meaning rather than assume it. A path segment this story never
took is not covered, and the fallback for "not covered" must be redoing
the work, not skipping it."""
db, adventure, settings, ids = forked
cursors.MEMORY.anchor_at(adventure, ids["nodes"]["C7"])
db.commit()
@@ -301,9 +305,10 @@ def test_a_mark_never_moves_forward_on_a_rewind(forked):
# ---------------------------------------------------- what the passes read
def test_the_summary_folds_in_only_the_path_it_is_on(forked, monkeypatch):
"""`_update_story_summary` gathers the memories past its mark. On C that is
B's and C's — never the one on A's own continuation, whose depth would
otherwise put it squarely inside the range."""
"""`_update_story_summary` gathers the memories past its mark. On C
that is B's and C's memories. It never includes the one on A's own
continuation, even though that memory's depth would otherwise put it
inside the range."""
db, adventure, settings, ids = forked
monkeypatch.setattr(memorybank, "SUMMARY_INTERVAL", 1)
@@ -360,7 +365,7 @@ def test_a_block_is_summarized_from_the_path_and_hung_off_its_last_node(
assert ["B4", "B5", "C6", "C7"] == [line for line in second.split() if line[0] in "ABC"]
made = db.query(models.Memory).filter_by(text="Memory 1.").one()
assert (made.branch_id, made.depth) == (ids["a"], 3)
# The mark ends up on the node the *second* block hangs off — the tip.
# The mark ends up on the node the second block attaches to, which is the tip.
assert cursors.MEMORY.stored(adventure) == (ids["c"], 7)
@@ -368,8 +373,8 @@ def test_a_block_is_summarized_from_the_path_and_hung_off_its_last_node(
@pytest.fixture()
def deeply_forked():
"""A story forked twenty times, with a memory every six actions — the
density the post-turn pass actually produces."""
"""A story forked twenty times, with a memory every six actions. This
is the density the post-turn pass actually produces."""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
user = models.User(is_guest=False, email="deepmem@example.com")
@@ -423,10 +428,11 @@ def deeply_forked():
def test_retrieving_from_a_deep_fork_costs_what_a_flat_story_costs(deeply_forked):
"""The bet, in bytes. Retrieval names all twenty-two branches instead of
one — but it is fetching an id and a flag per memory, and there are the
same fourteen either way, so the clause is where the difference is and the
clause is not what crosses the wire."""
"""The bet from the module docstring, measured in bytes. Retrieval
names all twenty-two branches instead of one, but it fetches only an id
and a flag per memory, and both stories return the same fourteen
memories. The clause is where the difference shows up, and the clause
is not what crosses the wire."""
db, flat_story, forked_story = deeply_forked
settings = db.query(models.Settings).one()
flat_id, forked_id = flat_story.id, forked_story.id
@@ -445,8 +451,8 @@ def test_retrieving_from_a_deep_fork_costs_what_a_flat_story_costs(deeply_forked
finally:
meter.detach()
# Measured 2026-08-18: 1,807 B against 1,823 B — the same fourteen rows,
# named through twenty-two branch terms instead of one.
# Measured 2026-08-18: 1,807 B against 1,823 B. Both figures cover the
# same fourteen rows, named through twenty-two branch terms instead of one.
assert flat_bytes > 0, "the meter saw nothing; it is measuring the wrong connection"
assert forked_bytes < flat_bytes * 1.5, (
f"retrieval on a 20-fork story cost {forked_bytes:,} B against the "
@@ -457,17 +463,19 @@ def test_retrieving_from_a_deep_fork_costs_what_a_flat_story_costs(deeply_forked
# ------------------------------------------------------- the opening node
def test_a_typed_memory_on_the_opening_node_survives_that_node_going(forked):
"""The one place a node and its memories part company.
"""The one exception where a node and its memories are not withdrawn
together.
A memory anchored to a node is withdrawn with the node, which is the rule
and is deliberate: it described that turn, and the turn is leaving. But
migration 62 parked *every* memory written before memories had coordinates
on depth 0 — the only landing spot visible from every branch — so the
opening node carries a whole bank it never produced. Withdrawing it would
retire all of that in one click, for every adventure predating the tree.
A memory anchored to a node is withdrawn with the node. This is the
rule, and it is deliberate: the memory described that turn, and the
turn is leaving. But migration 62 parked every memory written before
memories had coordinates on depth 0, the only landing spot visible from
every branch. As a result, the opening node carries a whole bank of
memories it never produced. Withdrawing it would delete all of those
memories at once, for every adventure that predates the tree.
A memory with no `source_start` covers no stretch of story, so nothing about
it can go stale. It stays.
A memory with no `source_start` covers no stretch of story, so nothing
about it can go stale. It stays.
"""
db, adventure, settings, ids = forked
typed = models.Memory(
@@ -489,9 +497,10 @@ def test_a_typed_memory_on_the_opening_node_survives_that_node_going(forked):
def test_a_summary_of_the_opening_node_is_still_withdrawn(forked):
"""The exception is about memories that describe nothing, not about depth 0.
A summary that genuinely ends on the opening node describes text that is
going, so it goes too — otherwise the root would collect exactly the
dangling rows `forget_node` replaced `prune_dangling_memories` to prevent.
A summary that genuinely ends on the opening node describes text that
is being removed, so the summary is removed too. Otherwise the root
would collect exactly the dangling rows that `forget_node` replaced
`prune_dangling_memories` to prevent.
"""
db, adventure, settings, ids = forked
derived = add_memory(db, adventure, "the opening, summarised", ids["nodes"]["A0"])
+33 -29
View File
@@ -1,11 +1,12 @@
"""Ranking the memory bank without reading the memory bank.
Retrieval used to walk `adventure.memories`, which loaded every row *with its
vector* — 96% of everything a turn read. It now asks SQL which memories are in
play, holds their vectors in process, and fetches text for the five it picks.
Retrieval used to walk `adventure.memories`, which loaded every row with
its vector. That vector data was 96% of everything a turn read. Retrieval
now asks SQL which memories are in play, holds their vectors in process,
and fetches text for only the five it picks.
Three things have to stay true for that to be safe, and each is a separate
failure that no error message would ever report:
failure that no error message would report:
* the ranking picks the same memories it always did;
* nothing bulk-reads a vector column again;
@@ -80,8 +81,8 @@ def adventure(db, settings):
)
db.add(adv)
db.flush()
# Retrieval builds its query from the newest actions; with none, it returns
# before ranking anything.
# Retrieval builds its query from the newest actions. With none, it
# returns before ranking anything.
for i in range(2):
db.add(models.Action(
adventure_id=adv.id, index=i, type="ai", text=f"Something happened {i}."
@@ -189,8 +190,8 @@ def test_update_stats_bumps_only_the_used(db, adventure, settings, bank):
def test_dry_runs_do_not_bump_the_counters(db, adventure, settings, bank):
"""Insights assembles a context without spending a turn; it must not look
like the memories were used."""
"""Insights assembles a context without spending a turn. It must not
look like the memories were used."""
retrieve(adventure, settings, StubEmbedder(), update_stats=False)
db.commit()
db.expire_all()
@@ -207,18 +208,19 @@ def memory_selects(statements):
def test_the_json_column_is_gone(db):
"""`memories.embedding` held the vectors before migration 38 and nothing
read it afterwards; migration 42 dropped it. Bringing it back would restore
4 MB of dead weight and a second place vectors can be written from — which
is how the model-switch bug happened (test_embedding_model_switch.py)."""
"""`memories.embedding` held the vectors before migration 38, and
nothing read it afterward. Migration 42 dropped it. Restoring it would
bring back 4 MB of dead weight and a second place vectors can be
written from. That second place is how the model-switch bug happened
(test_embedding_model_switch.py)."""
columns = {c["name"] for c in sa_inspect(engine).get_columns("memories")}
assert "embedding" not in columns
assert {"embedding_blob", "embedded"} <= columns
def test_the_catalogue_query_carries_no_vectors(db, adventure, settings, bank, sql_log):
"""The query that decides *which* memories are in play must stay tiny —
this is the one that used to drag the whole bank across."""
"""The query that decides which memories are in play must stay tiny.
This is the query that used to pull the whole bank across the wire."""
retrieve(adventure, settings, StubEmbedder())
catalogue = [s for s in memory_selects(sql_log) if "memories.pinned" in s]
assert catalogue, "expected a catalogue query"
@@ -261,9 +263,10 @@ def test_only_top_k_texts_are_fetched(db, adventure, settings, bank, sql_log):
# ---------------------------------------------------------------- staleness
def test_a_rewritten_vector_is_not_served_from_cache(db, adventure, settings, bank):
"""The cache's one genuine hazard: a memory keeps its id while its vector
changes, so an id-set check alone would go on serving the old one. Editing
a memory's text and re-embedding it does exactly that.
"""The cache's one genuine hazard: a memory keeps its id while its
vector changes, so an id-set check alone would continue serving the
old vector. Editing a memory's text and re-embedding it does exactly
that.
"""
settings.memory_top_k = 1
db.commit()
@@ -325,9 +328,9 @@ def test_eviction_marks_the_least_recently_used(db, adventure, settings):
def test_eviction_breaks_ties_on_use_count(db, adventure, settings):
"""Two memories last wanted at the same moment: the one the story has
leaned on less goes. Only a tiebreak — ranking on the count first is what
used to freeze the bank (see below)."""
"""Two memories last used at the same moment: the one the story has
used less is the one that goes. This must be only a tiebreak. Ranking
on the count first is what used to freeze the bank (see below)."""
settings.memory_bank_capacity = 1
db.commit()
now = models.utcnow()
@@ -344,12 +347,13 @@ def test_eviction_breaks_ties_on_use_count(db, adventure, settings):
def test_a_newborn_is_not_evicted_by_the_bank_it_joins(db, adventure, settings):
"""The bank used to shut itself. Eviction ranked on use_count first, and a
memory written this turn has never been used, so the moment every survivor
had been retrieved even once the newborn was the lowest row in the bank and
was retired in the same post-turn run that wrote it — before retrieval ever
saw it. That state is absorbing: counts only go up, so no memory written
after it could ever get in either."""
"""The bank used to stop accepting new memories. Eviction ranked on
use_count first, and a memory written this turn has never been used.
Once every existing memory had been retrieved even once, the newborn
became the lowest-ranked row in the bank. Eviction then removed it in
the same post-turn run that wrote it, before retrieval ever saw it.
That state never recovers: counts only go up, so no memory written
after it could get in either."""
settings.memory_bank_capacity = 3
db.commit()
now = models.utcnow()
@@ -370,9 +374,9 @@ def test_a_newborn_is_not_evicted_by_the_bank_it_joins(db, adventure, settings):
def test_a_full_bank_still_turns_over(db, adventure, settings):
"""The same failure seen over several turns: a bank at capacity has to keep
taking on what the story is doing now, or the adventure stops remembering
anything past the point it filled up."""
"""The same failure seen over several turns: a bank at capacity must
keep accepting new memories, or the adventure stops remembering
anything past the point where it filled up."""
settings.memory_bank_capacity = 3
now = models.utcnow()
db.commit()
+40 -36
View File
@@ -1,22 +1,24 @@
"""Memories must never describe narration that is no longer in the story, and
must never skip a stretch of it.
For six phases the answer was a **holdback**: summarization stopped one action
short of the newest, because only the last action was retryable and a retry
rewrote `Action.text` under a mark that had already moved past it. SP4 ended
that — a retry writes a sibling node and the coordinate's derived work is
withdrawn as it does, which is the same repair undo and delete already made.
So the holdback is gone, and the first half of this file now asserts the
property that replaced it: a block forms as soon as there is a block, and
changing what a coordinate says takes back what was derived from it.
For six phases, the answer was a holdback. Summarization stopped one action
short of the newest, because only the last action was retryable, and a
retry rewrote `Action.text` under a mark that had already moved past it.
SP4 ended that: a retry writes a sibling node, and the coordinate's derived
work is withdrawn as it happens, using the same repair that undo and delete
already made. The holdback is gone, so the first half of this file now
asserts the property that replaced it. A block forms as soon as there is a
block, and changing what a coordinate says takes back what was derived from
it.
Phase 14 SP3 changed what the mark *is*. It used to be a count of covered story
actions, and the second half of this file is the price of that: deleting an
action from in front of a position slid a never-summarized action into the
covered range, so every delete had to slide the cursors too. The mark is a node
now — `(branch_id, depth)` — and a node does not move when something in front
of it is deleted, so those tests assert that nothing happens where they used to
assert that the right correction happened.
Phase 14 SP3 changed what the mark is. It used to be a count of covered
story actions, and the second half of this file is the cost of that.
Deleting an action from in front of a position slid a never-summarized
action into the covered range, so every delete had to slide the cursors
too. The mark is a node now, `(branch_id, depth)`, and a node does not move
when something in front of it is deleted. Those tests now assert that
nothing happens, where they used to assert that the right correction
happened.
python -m pytest tests/test_memory_settling.py -v
"""
@@ -110,13 +112,14 @@ def run_memories(db, adventure, stub, monkeypatch):
# --------------------------------------------------- no holdback, since SP4
def test_a_block_forms_as_soon_as_the_story_holds_one(db, monkeypatch):
"""Covered to action 5 with 12 actions: block 6-11 ends on the *newest*
"""Covered to action 5 with 12 actions: block 6-11 ends on the newest
action, and is summarized now rather than a turn later.
This is exactly the case the holdback existed to refuse. What makes it safe
is no longer that the block stops short — it is that a retry of node 11
would withdraw this memory on its way past (see
`test_deleting_a_summarized_node_withdraws_its_memory`, the same repair).
This is exactly the case the holdback existed to refuse. What makes it
safe is no longer that the block stops short. It is that a retry of
node 11 would withdraw this memory on its way past (see
`test_deleting_a_summarized_node_withdraws_its_memory`, the same
repair).
"""
adventure = make_adventure(db, 12)
cover(db, adventure, 6)
@@ -128,8 +131,8 @@ def test_a_block_forms_as_soon_as_the_story_holds_one(db, monkeypatch):
assert "Action 11." in stub.excerpts[0]
memory = db.query(models.Memory).one()
assert (memory.source_start, memory.source_end) == (6, 11)
# The mark and the memory name the same node — that is what keeps them from
# drifting apart however gappy the depths underneath are.
# The mark and the memory name the same node. That is what keeps them
# from drifting apart, however gappy the underlying depths are.
assert (memory.branch_id, memory.depth) == cursors.MEMORY.stored(adventure)
assert covered_depth(db, adventure) == 11
@@ -155,13 +158,13 @@ def test_the_first_memory_lands_at_memory_start(db, monkeypatch):
def test_legacy_caught_up_adventure_is_not_rewound(db, monkeypatch):
"""An adventure summarized under the OLD rule carries a cursor equal to its
action count — one past the end of the story. That used to need a clamp on
every post-turn pass, and clamping it to the settled count re-covered an
action.
"""An adventure summarized under the old rule carries a cursor equal to
its action count, one past the end of the story. That used to require a
clamp on every post-turn pass, and clamping it to the settled count
re-covered an action.
A mark that names a node has no such edge: the newest action is the node,
and "everything after it" is empty until the story grows.
A mark that names a node has no such edge. The newest action is the
node, and "everything after it" is empty until the story grows.
"""
adventure = make_adventure(db, 12)
db.add(models.Memory(adventure_id=adventure.id, text="A", source_start=0, source_end=5))
@@ -238,9 +241,9 @@ def test_deleting_a_middle_action_leaves_the_mark_where_it_was(db):
with node 5 gone.
"""
adventure = summarized_adventure(db)
# Node 4 is inside memory A's block but is not the node it hangs off, so
# nothing is withdrawn — the same reading the old code had, where only a
# memory whose *end* had fallen off the story was pruned.
# Node 4 is inside memory A's block but is not the node it hangs off,
# so nothing is withdrawn. The old code read it the same way: only a
# memory whose end had fallen off the story was pruned.
victim = db.query(models.Action).filter_by(adventure_id=adventure.id, index=4).one()
assert memorybank.forget_node(db, adventure, victim) == 0
@@ -267,12 +270,13 @@ def test_deleting_a_later_action_leaves_the_mark_alone(db):
def test_deleting_a_summarized_node_withdraws_its_memory(db):
"""Discarding the memory isn't enough — the story it covered is still
behind the mark, so the mark has to come back to where that block began.
"""Discarding the memory is not enough. The story it covered is still
behind the mark, so the mark has to move back to where that block began.
Memory B ends on node 11, so deleting node 11 is what withdraws it. The old
code found this by scanning for a memory whose covered range had fallen off
the end of the story; the memory hangs off the node now, so it is a lookup.
Memory B ends on node 11, so deleting node 11 withdraws it. The old code
found this by scanning for a memory whose covered range had fallen off
the end of the story. Now the memory hangs off the node, so finding it
is a lookup.
"""
adventure = summarized_adventure(db)
victim = db.query(models.Action).filter_by(adventure_id=adventure.id, index=11).one()
+2 -2
View File
@@ -22,7 +22,7 @@ def hosted(monkeypatch):
def _resolves_to(monkeypatch, ip: str):
"""Pin getaddrinfo so we test the address decision, not real DNS."""
"""Pin `getaddrinfo` so the test exercises the address decision, not real DNS."""
monkeypatch.setattr(
netguard.socket, "getaddrinfo",
lambda *a, **k: [(2, 1, 6, "", (ip, 443))],
@@ -64,6 +64,6 @@ def test_unresolvable_host_is_blocked(hosted, monkeypatch):
def test_noop_in_local_mode(monkeypatch):
monkeypatch.setattr(auth, "MULTI_USER", False)
# Local installs legitimately reach localhost (Ollama) — never blocked.
# Local installs must reach localhost (Ollama). The guard never blocks local mode.
assert netguard.endpoint_block_reason("http://localhost:11434/v1") is None
assert netguard.endpoint_block_reason("http://127.0.0.1:11434/v1") is None
+35 -31
View File
@@ -1,25 +1,26 @@
"""Prompt caching: the prompt has to start with the same bytes every turn.
Every endpoint that caches prompts caches a *prefix* — it reuses the request up
to the first byte that differs from last time and no further. So the cost of a
turn is decided by layout: one section that changes each turn, placed near the
top, re-prices everything underneath it, and underneath it is the story
history, which is most of the prompt.
Every endpoint that caches prompts caches a prefix. It reuses the request
up to the first byte that differs from last time, and no further. The cost
of a turn is therefore decided by layout: one section that changes each
turn, placed near the top, re-prices everything underneath it, and
underneath it is the story history, which makes up most of the prompt.
Three things have to hold, and each is easy to undo by accident:
Three things must hold, and each is easy to undo by accident:
* the static block is byte-identical across turns — adding a section that moves
(live stats, retrieved memories, a rewritten summary) to `system_sections` is
the mistake this file exists to catch;
* the sections that move sit *after* the history, but still *before* the tail
that is last for its own reasons (front memory, the length hint, and
EMIT_REMINDER, which is what keeps the state block emitted at all);
* moving a section out of the system block does not drop it from the token
budget — it is still in the prompt.
* The static block is byte-identical across turns. Adding a section that
moves (live stats, retrieved memories, a rewritten summary) to
`system_sections` is the mistake this file exists to catch.
* The sections that move sit after the history, but still before the tail
that is last for its own reasons: front memory, the length hint, and
`EMIT_REMINDER`, which is what keeps the state block emitted at all.
* Moving a section out of the system block does not drop it from the
token budget. It is still in the prompt.
Plus the two request-level halves: preferring one OpenRouter upstream, since
each upstream holds its own cache, and reading back the usage the endpoint
reports so the hit rate is measurable rather than assumed.
This file also covers two request-level concerns: preferring one
OpenRouter upstream, since each upstream holds its own cache, and reading
back the usage the endpoint reports, so the hit rate is measurable rather
than assumed.
python -m pytest tests/test_prompt_caching.py -v
"""
@@ -68,8 +69,8 @@ def test_fallbacks_stay_on():
def test_non_openrouter_endpoints_get_no_provider_field():
"""Ollama and friends reject fields they do not know — the same trap the
`reasoning` param is written around."""
"""Ollama and other providers reject fields they do not know. This is
the same problem the `reasoning` param works around."""
body = _routed("http://localhost:11434/v1", "deepseek/deepseek-v4-flash-0731")
assert "provider" not in body
@@ -85,8 +86,9 @@ def test_unknown_vendors_are_left_alone():
# ------------------------------------------------------ reading usage back
def test_usage_is_recorded_from_a_final_chunk():
"""In a stream the usage block rides on a last chunk carrying no choices,
which is why it is read separately from the text extraction."""
"""In a stream, the usage block arrives in a final chunk that carries
no choices, which is why it is read separately from the text
extraction."""
provider = OpenAICompatibleProvider("https://openrouter.ai/api/v1", "k", "m")
assert provider.last_usage is None
provider._record_usage({"choices": [{"delta": {"content": "hi"}}]})
@@ -109,8 +111,8 @@ def test_a_later_chunk_without_usage_does_not_erase_it():
# ------------------------------------------------------------ prompt layout
def _with_hp(world_state, hp):
"""`world_state` is nested by group, and the JSON column only notices a
whole new object — so build one rather than mutating in place."""
"""`world_state` is nested by group, and the JSON column only detects a
whole new object. Build a new one instead of mutating in place."""
return {**world_state, "player": {**world_state["player"], "hp": hp}}
@@ -153,8 +155,9 @@ def story():
def test_changing_a_stat_leaves_the_static_block_untouched(story):
"""The whole point. Live values used to sit third from the top, so a single
point of damage re-priced the instructions, the plot and the history."""
"""The whole point. Live values used to sit third from the top, so a
single point of damage re-priced the instructions, the plot, and the
history."""
db, adventure, settings = story
before, _, _ = builder.build_context(adventure, settings)
adventure.world_state = _with_hp(adventure.world_state, 40)
@@ -169,8 +172,8 @@ def test_the_static_block_holds_the_things_that_do_not_move(story):
system_text, story_text, _ = builder.build_context(adventure, settings)
for fixed in ("Write in second person.", "The hero hunts bandits."):
assert fixed in system_text
# The stat *guide* is derived from the schema and so is fixed; the live
# values it describes are not, and belong to the story text.
# The stat guide is derived from the schema, so it is fixed. The live
# values it describes are not fixed, and belong to the story text.
assert "Stat guide" in system_text
for moves in ("The hero left the village.", "hp 100/100"):
assert moves not in system_text
@@ -186,8 +189,9 @@ def test_volatile_sections_sit_after_the_history(story):
def test_the_tail_stays_the_tail(story):
"""front memory, the length hint and EMIT_REMINDER are last for reasons of
their own, and the live sections must not have displaced them."""
"""Front memory, the length hint, and `EMIT_REMINDER` are last for
reasons of their own, and the live sections must not have displaced
them."""
db, adventure, settings = story
_, story_text, report = builder.build_context(adventure, settings)
labels = [s["label"] for s in report["sections"]]
@@ -199,8 +203,8 @@ def test_the_tail_stays_the_tail(story):
def test_live_sections_are_still_charged_to_the_budget(story):
"""They moved out of `system_sections`, so it would be easy to stop
counting them in `reserved` — and then the history, which is budgeted with
what is left over, would quietly overrun."""
counting them in `reserved`. If that happened, the history, which is
budgeted with what is left over, would quietly overrun."""
db, adventure, settings = story
for i in range(6, 90):
db.add(models.Action(
+12 -12
View File
@@ -1,11 +1,11 @@
"""Regression tests for the X-Forwarded-For rate-limit bypass and the
per-account login throttle added to close it.
Background: uvicorn's --forwarded-allow-ips "*" trusted the LEFTMOST
X-Forwarded-For entry, which the client controls, so rotating the header
handed out a fresh rate-limit bucket per request. client_ip now reads the
hop the trusted edge appends (rightmost), and login has an email-keyed throttle
that no IP trick can dilute.
Background: uvicorn's `--forwarded-allow-ips "*"` trusted the leftmost
`X-Forwarded-For` entry. The client controls that entry, so rotating the
header issued a fresh rate-limit bucket on every request. `client_ip` now
reads the hop the trusted edge appends, which is the rightmost one. Login
also has an email-keyed throttle that no IP trick can weaken.
python -m pytest tests/test_ratelimit_hardening.py -v
"""
@@ -35,15 +35,15 @@ class _Req:
def test_client_ip_takes_appended_rightmost_hop(monkeypatch):
monkeypatch.setattr(limits, "TRUSTED_PROXY_HOPS", 1)
# Attacker prepends a fake IP; the edge appends the real one on the right.
# An attacker prepends a fake IP. The edge appends the real one on the right.
req = _Req("203.0.113.9, 198.51.100.77")
assert limits.client_ip(req) == "198.51.100.77"
def test_client_ip_ignores_spoofed_leftmost(monkeypatch):
monkeypatch.setattr(limits, "TRUSTED_PROXY_HOPS", 1)
# Whatever the client stuffs to the left, the keyed IP stays the real hop —
# so rotating it no longer mints a new bucket.
# The keyed IP stays the real hop regardless of what the client adds on
# the left, so rotating that value no longer creates a new bucket.
a = limits.client_ip(_Req("1.1.1.1, 198.51.100.77"))
b = limits.client_ip(_Req("2.2.2.2, 198.51.100.77"))
c = limits.client_ip(_Req("evil, junk, 198.51.100.77"))
@@ -74,11 +74,11 @@ def _multi_user(monkeypatch):
def test_login_throttle_blocks_after_limit():
email = "victim@example.com"
# Up to the limit: allowed, each a recorded failure.
# Each attempt up to the limit is allowed and recorded as a failure.
for _ in range(limits.LOGIN_FAIL_LIMIT):
limits.check_login_allowed(email) # does not raise
limits.note_login_failure(email)
# One more crosses the line.
# One more failure exceeds the limit.
with pytest.raises(limits.HTTPException) as exc:
limits.check_login_allowed(email)
assert exc.value.status_code == 429
@@ -89,7 +89,7 @@ def test_login_throttle_is_per_account():
limits.note_login_failure("a@example.com")
with pytest.raises(limits.HTTPException):
limits.check_login_allowed("a@example.com")
# A different account is unaffected — this is not an IP bucket.
# A different account is unaffected because the throttle keys on email, not IP.
limits.check_login_allowed("b@example.com") # must not raise
@@ -98,7 +98,7 @@ def test_successful_login_clears_the_streak():
for _ in range(limits.LOGIN_FAIL_LIMIT):
limits.note_login_failure(email)
limits.note_login_success(email)
limits.check_login_allowed(email) # streak wiped — must not raise
limits.check_login_allowed(email) # failure streak cleared, must not raise
def test_throttle_is_noop_in_local_mode(monkeypatch):
+2 -2
View File
@@ -17,7 +17,7 @@ def _body(reasoning_max_tokens, api_mode="chat", max_tokens=1000):
def test_zero_sends_nothing():
"""Ollama and friends reject unknown fields — 0 must stay silent."""
"""Ollama and other providers reject unknown fields. Sending 0 must not add a `reasoning` field."""
assert "reasoning" not in _body(0)
@@ -36,7 +36,7 @@ def test_negative_turns_reasoning_off():
def test_off_is_not_merely_excluded():
"""`exclude: true` still thinks and still bills; we want it actually off."""
"""`exclude: true` still generates and bills for reasoning tokens. The off setting must omit the field entirely instead of relying on `exclude`."""
assert _body(-1)["reasoning"].get("exclude") is None
+9 -9
View File
@@ -137,7 +137,7 @@ def test_retry_keeps_the_discarded_attempt(client):
_retry(client)
actions = _actions(client)
# One AI action still, not two — the retry replaced the live text in place.
# One AI action still, not two. The retry replaced the live text in place.
assert [a["type"] for a in actions] == ["start", "do", "ai"]
last = actions[-1]
assert last["text"] == "Attempt two."
@@ -151,18 +151,18 @@ def test_retry_keeps_the_discarded_attempt(client):
def test_retry_context_excludes_the_attempt_being_replaced(client):
"""The whole point of a retry is a fresh take on the *same* turn. The row
now survives the retry (it holds the variant history), so it is still in
`adventure.actions` while the replacement context is assembled — it must be
filtered out, or the model is asked to continue *past* the attempt it is
supposed to be replacing and writes a sequel that blends both."""
"""A retry produces a fresh take on the same turn. The row survives the
retry because it holds the variant history, so it is still in
`adventure.actions` while the replacement context is assembled. The
context builder must filter it out, or the model continues past the
attempt it is replacing and writes a sequel that blends both."""
ScriptedProvider.replies = ["Attempt one.", "Attempt two."]
_play(client)
_retry(client)
retry_story = ScriptedProvider.prompts[-1][1]
assert "Attempt one." not in retry_story
# The turn's own player action must still be there — it's what to respond to.
# The turn's own player action must still be there. It is what the model responds to.
assert "look around" in retry_story
assert "You enter a cave." in retry_story
@@ -257,7 +257,7 @@ def test_cannot_switch_a_turn_the_story_moved_past(client):
f"/api/adventures/{client.adv_id}/actions/{retried['id']}/variant", json={"index": 0})
assert r.status_code == 400
assert "latest message" in r.json()["detail"]
# Still readable, though — that's the whole point of keeping them.
# The variant is still readable. Keeping every attempt browsable is why it still exists.
variants = client.get(
f"/api/adventures/{client.adv_id}/actions/{retried['id']}/variants").json()
assert [v["text"] for v in variants] == ["One.", "Two."]
@@ -333,7 +333,7 @@ def test_export_and_import_round_trips_variants(client):
bundle = client.get(f"/api/adventures/{client.adv_id}/export").json()
# SP6: the attempts are nodes in the bundle too, sharing one coordinate,
# and `live` says which of them the story tells. The `variants` array
# survives only in the v1 *reader* — see the hand-edited bundle below.
# survives only in the v1 reader. See the hand-edited bundle below.
ai = [a for a in bundle["actions"] if a["type"] == "ai"]
assert [(a["text"], a["live"]) for a in ai] == [("One.", False), ("Two.", True)]
assert len({(a["branch"], a["depth"]) for a in ai}) == 1
+14 -12
View File
@@ -1,9 +1,10 @@
"""Scenario cover art + the Continue-card snippet.
"""Scenario cover art and the Continue-card snippet.
Unit tests for the data-URI handling in app/images.py, then HTTP tests that the
list endpoints advertise a cacheable `image_url` (never the inline base64), that
the image route serves real bytes, and that an adventure's snippet comes from
the latest *narration* rather than the player's last line.
Unit tests for the data-URI handling in app/images.py. Then HTTP tests
confirm that the list endpoints advertise a cacheable `image_url` instead
of the inline base64, that the image route serves real bytes, and that an
adventure's snippet comes from the latest narration rather than the
player's last line.
python -m pytest tests/test_scenario_art.py -v
"""
@@ -41,13 +42,14 @@ def test_decode_returns_bytes_and_content_type():
def test_decode_tolerates_wrapped_base64():
"""A hand-pasted URI can carry newlines; b64decode(validate=True) won't."""
"""A hand-pasted URI can carry newlines. `b64decode(validate=True)` rejects them."""
wrapped = "data:image/png;base64," + "\n".join(
base64.b64encode(PNG_BYTES).decode()[i:i + 24] for i in range(0, 100, 24)
)
# Only asserting it doesn't raise and doesn't silently return a partial
# decode of a truncated payload — the wrapped prefix here is not the whole
# image, so a None result is also acceptable; what matters is no exception.
# This test only confirms that the call does not raise and does not
# silently return a partial decode of a truncated payload. The wrapped
# prefix here is not the whole image, so a None result is also
# acceptable. What matters is that no exception occurs.
images.decode(wrapped)
@@ -59,7 +61,7 @@ def test_decode_rejects_non_data_uris_and_garbage():
def test_decode_rejects_svg():
"""SVG can carry script and these bytes are served from our own origin."""
"""SVG can carry script, and the app serves these bytes from its own origin."""
svg = "data:image/svg+xml;base64," + base64.b64encode(b"<svg/>").decode()
assert images.decode(svg) is None
@@ -151,7 +153,7 @@ def test_scenario_list_advertises_a_url_and_hides_the_base64(client):
pictured = rows["Pictured"]
assert pictured["image_url"].startswith(f"/api/scenarios/{client.ids['scenario']}/image?v=")
# The whole point: a list response must never carry the inline image.
# A list response must never carry the inline image.
assert "image" not in pictured
assert rows["Emoji"]["image_url"] == ""
@@ -180,7 +182,7 @@ def test_single_scenario_still_returns_the_raw_uri_for_editing(client):
def test_adventure_list_carries_snippet_and_inherited_art(client):
row = next(r for r in client.get("/api/adventures").json()
if r["id"] == client.ids["adventure"])
# Latest narration, whitespace collapsed — not the player's "I draw my sword."
# Latest narration, whitespace collapsed. Not the player's line, "I draw my sword."
assert row["snippet"] == "Steel rings. The corridor answers."
assert row["image_url"].startswith(f"/api/scenarios/{client.ids['scenario']}/image?v=")
assert row["action_count"] == 3
+6 -5
View File
@@ -75,12 +75,13 @@ def make_scenario(client, **kwargs):
def edit_scenario(scenario_id, cards=None, **fields):
"""Author-side edit, straight to the DB (the scenario API is tested elsewhere).
"""Author-side edit, straight to the database. The scenario API is
tested elsewhere.
`cards` is the scenario's full card list afterwards. Cards are matched to
existing rows by name and edited in place, exactly as ScenarioEditor does
(PATCH /story-cards/{id}) — card ids are stable across authoring, which is
what source_ref tracking relies on.
`cards` is the scenario's full card list afterwards. Cards are matched
to existing rows by name and edited in place, exactly as ScenarioEditor
does (PATCH /story-cards/{id}). Card ids stay stable across authoring,
which is what source_ref tracking relies on.
"""
db = SessionLocal()
try:
+17 -16
View File
@@ -1,18 +1,18 @@
"""context_snapshot, stored compressed.
The column is 89% of the database and the free tier allows 512 MB. Reads were
solved by deferring it; this is about the storage ceiling. Postgres already
TOASTs it and only gets 1.7x, because pglz is tuned for fast decompression of
data a query might filter on — and nothing ever filters on an assembled
prompt.
The column makes up 89% of the database, and the free tier allows only
512 MB. Deferring the column already solved the read cost, so this is
about the storage ceiling. Postgres already TOASTs the column and gets
only a 1.7x ratio, because pglz favors fast decompression for data a query
might filter on. Nothing ever filters on an assembled prompt.
Three things have to hold, and only the first is obvious:
Three things must hold, and only the first is obvious:
* what goes in comes back out, exactly, including a snapshot written before
the conversion and one that is NULL;
* the model still hands callers a dict, so no call site changes;
* migration 44 drops the original column, so the backfill is the one
destructive step in this file — it must convert every row or abort.
* What goes in comes back out exactly, including a snapshot written before
the conversion and one that is NULL.
* The model still hands callers a dict, so no call site changes.
* Migration 44 drops the original column, so the backfill is the one
destructive step in this file. It must convert every row or abort.
python -m pytest tests/test_snapshot_compression.py -v
"""
@@ -147,8 +147,9 @@ def test_null_stays_null(db, adventure):
def test_an_unreadable_snapshot_reads_as_none_rather_than_raising(db, adventure):
"""One corrupt row must not 500 the turn that happens to load it. The
snapshot is a debugging view; the story is the thing that matters."""
"""One corrupt row must not return a 500 error for the turn that loads
it. The snapshot is a debugging view. The story is what actually
matters."""
action = models.Action(
adventure_id=adventure.id, index=0, type="ai", text="t",
context_snapshot={"a": "b"},
@@ -236,9 +237,9 @@ def test_bootstrap_leaves_the_column_named_context_snapshot(db, adventure):
def test_the_backfill_aborts_rather_than_dropping_unconvertible_data(
db, adventure, monkeypatch
):
"""Migration 44 destroys the original. If anything cannot be converted the
whole run has to roll back with the column still there — the alternative is
losing somebody's prompts to a bug in this file."""
"""Migration 44 destroys the original column. If anything fails to
convert, the whole run must roll back with the column still there.
Otherwise a bug in this file could lose someone's prompts."""
seed_pre_43(db, adventure)
db.close()
+20 -18
View File
@@ -1,20 +1,22 @@
"""Tests for undo/retry rolling back the shared script_state scoreboard
"""Tests for undo and retry rolling back the shared `script_state`
(plan/11-state-revert-and-retry-fix.md).
Phase 14 SP4 turned the snapshots around. An action used to carry the state as
it stood *before* it ran, and rolling back read the snapshot off the action
being removed. It carries what it left *behind* now, and rolling back reads it
off the node in front — which is the same number arrived at from the other
side, and the only version a retry can use: attempts at one turn share a
starting position and differ precisely in their outcome.
Phase 14 SP4 reversed the snapshots. An action used to carry the state as
it stood before it ran, and rolling back read the snapshot off the action
being removed. Now it carries the state it left behind, and rolling back
reads that state off the node in front of it. This is the same value
reached from the other direction, and it is the only version a retry can
use, because attempts at one turn share a starting position and differ
only in their outcome.
Run from the backend dir: python -m pytest tests/test_state_revert.py -v
"""
import os
import tempfile
# Point the app at a throwaway SQLite file BEFORE importing anything that binds
# the engine at import time (app.database reads AIDND_DB_PATH on import).
# Point the app at a throwaway SQLite file before importing anything that
# binds the engine at import time. `app.database` reads `AIDND_DB_PATH`
# on import.
_tmp = tempfile.NamedTemporaryFile(suffix=".db", delete=False)
_tmp.close()
os.environ["AIDND_DB_PATH"] = _tmp.name
@@ -77,9 +79,9 @@ def _forget_snapshots(db, adv):
# ---------------------------------------------------------------- undo
def test_undo_reverts_state_to_before_the_turn(db):
# A turn took the scoreboard from {gold:0} -> {gold:10}. The node in front
# of the turn is what says where it started; current state is the mutated
# one.
# A turn moved script_state from {gold:0} to {gold:10}. The node in
# front of the turn records where it started. The current state is
# the mutated one.
user, adv = _make_adventure(db, {"gold": 10})
_add(db, adv, 0, "start", state_after={"gold": 0})
_add(db, adv, 1, "do", state_after={"gold": 0})
@@ -165,9 +167,9 @@ def test_undo_prunes_memory_covering_removed_actions(db):
# -------------------------------------------------------- withdrawing a node
def test_forget_node_withdraws_only_what_that_node_produced(db):
"""Phase 14 SP3: a memory hangs off the node its block ends on, so removing
a node is a lookup rather than a scan for memories that have fallen off the
end of the story."""
"""Phase 14 SP3: a memory attaches to the node where its block ends, so
removing a node is a lookup rather than a scan for memories that
reference actions the story no longer has."""
user, adv = _make_adventure(db, {})
_add(db, adv, 0, "do")
second = _add(db, adv, 1, "ai")
@@ -216,9 +218,9 @@ def test_restore_state_ignores_a_node_with_no_outcome(db):
# ---------------------------------------------------------------- retry
def test_retry_restores_the_state_the_turn_started_from(db, monkeypatch):
# Retry must roll the scoreboard back to what the node in front of the AI
# action left behind, so regeneration doesn't stack output mutations on top
# of the attempt being replaced.
# Retry must roll script_state back to what the node in front of the AI
# action left behind, so regeneration does not stack output mutations
# on top of the attempt being replaced.
user, adv = _make_adventure(db, {"gold": 20}) # 20 = double-applied bug value
_add(db, adv, 0, "start", state_after={"gold": 0})
_add(db, adv, 1, "do", state_after={"gold": 10})
+34 -32
View File
@@ -1,20 +1,21 @@
"""Phase 14 SP0 — the regression contract for the story tree.
"""Phase 14 SP0: the regression contract for the story tree.
This file exists to answer one question, over and over, as the storage model is
replaced underneath the app: **do existing adventures still behave exactly as
they did?**
This file exists to answer one question repeatedly, as the storage model is
replaced underneath the app: do existing adventures still behave exactly as
they did?
So it drives the product the way a player does — over HTTP, asserting only on
API responses — and never reaches into the ORM to check how something is
It drives the product the way a player does, over HTTP, asserting only on
API responses. It never reaches into the ORM to check how something is
stored. Everything it asserts is true of the linear implementation today and
must stay true of the tree, because a linear story is a tree with one branch.
python -m pytest tests/test_story_tree_baseline.py -v
**It must pass UNMODIFIED through SP1 (schema), SP2 (branch clause) and SP3
(memories on nodes).** If a change here looks necessary in one of those
subphases, the change is wrong, not the test. SP4 is the first subphase allowed
to move it, and only for the variant-count semantics called out in plan/14.
This file must pass unmodified through SP1 (schema), SP2 (branch clause),
and SP3 (memories on nodes). If a change here looks necessary in one of
those subphases, the change is wrong, not the test. SP4 is the first
subphase allowed to move it, and only for the variant-count semantics
called out in plan/14.
"""
import os
import tempfile
@@ -73,9 +74,9 @@ class ScriptedProvider:
def _make_world(monkeypatch, *, seeded_actions: int = 0):
"""A user + scenario + adventure, with `seeded_actions` extra story actions
written straight to the database (paging wants more turns than it is worth
playing one at a time)."""
"""Create a user, a scenario, and an adventure, with `seeded_actions`
extra story actions written straight to the database. Paging tests need
more turns than it is practical to play one at a time."""
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="baseline@example.com")
@@ -181,8 +182,8 @@ def _state(adv_id):
# ------------------------------------------------------- opening the story
def test_adventure_opens_on_a_window_with_a_total(long_client):
"""The payload is the newest window plus the whole story's length — that is
how the client knows there is more above."""
"""The payload includes the newest window and the whole story's length.
That is how the client knows more actions exist above it."""
payload = _open(long_client)
seeded = adventures.ACTION_PAGE + 11 # + the start action
assert payload["action_count"] == seeded
@@ -247,7 +248,7 @@ def test_retry_replaces_the_text_and_keeps_the_attempt(client):
assert r.status_code == 200, r.text
actions = _actions(client)
# Still one AI action for this turn — a retry is a new take, not a new turn.
# This turn still has one AI action. A retry creates a new take, not a new turn.
assert [a["type"] for a in actions] == ["start", "do", "ai"]
assert actions[-1]["text"] == "Attempt two."
@@ -279,14 +280,14 @@ def test_switching_back_to_an_earlier_attempt_restores_it(client):
json={"index": 0})
assert r.status_code == 200, r.text
assert _texts(client)[-1] == "Attempt one."
# And the script state that attempt produced comes back with it.
# The script state that attempt produced comes back with it.
script_state, _ = _state(client.adv_id)
assert script_state["gold"] == 10
def test_only_the_newest_turn_can_be_switched(client):
"""An older turn's alternatives stay readable but not selectable — the
story after it was written as a continuation of what is live."""
"""An older turn's alternatives stay readable but not selectable. The
story after it continues from what is live."""
ScriptedProvider.replies = ["One.", "Again.", "Two."]
_play(client)
client.post(f"/api/adventures/{client.adv_id}/retry")
@@ -313,7 +314,7 @@ def test_undo_removes_the_whole_turn_and_rolls_state_back(client):
# Both the AI action and the player action that prompted it are gone.
assert [a["type"] for a in page["actions"]] == ["start", "do", "ai"]
assert page["total"] == 3
# ...and the second turn's ten gold went with them.
# The second turn's ten gold is also gone.
script_state, _ = _state(client.adv_id)
assert script_state["gold"] == 10
@@ -338,9 +339,9 @@ def test_the_opening_cannot_be_undone(client):
# ------------------------------------------------------------------ paging
def test_paging_walks_back_to_the_start_without_gaps_or_repeats(long_client):
"""The property that matters: page all the way up and you have seen every
action exactly once, in order. This is the invariant a tree must preserve —
the anchor is an action, never a position."""
"""The property that matters: paging all the way up must show every
action exactly once, in order. A tree must preserve this invariant. The
anchor is always an action, never a position."""
payload = _open(long_client)
total = payload["action_count"]
seen = [a["id"] for a in payload["actions"]]
@@ -376,7 +377,7 @@ def test_paging_past_the_start_reports_the_end(long_client):
r = long_client.get(f"/api/adventures/{long_client.adv_id}/actions",
params={"limit": 5})
oldest = r.json()["actions"][0]["id"]
# Walk right off the front.
# Continue paging until it reaches the oldest action.
seen_all = False
cursor = oldest
for _ in range(100):
@@ -399,8 +400,9 @@ def test_editing_an_action_sticks(client):
r = client.patch(f"/api/adventures/{client.adv_id}/actions/{action_id}",
json={"text": "Edited."})
assert r.status_code == 200, r.text
# Re-read rather than trusting the response: an edit that only lives in the
# response has silently reverted for anyone who reloads.
# Re-read from the database instead of trusting the response. An edit
# that only appears in the response has already reverted for anyone who
# reloads.
assert _texts(client)[-1] == "Edited."
@@ -450,13 +452,13 @@ def test_memories_are_created_listed_and_deleted(client):
# ------------------------------------------------------------------ export
def test_export_carries_the_whole_story(client):
"""A bundle is a backup: every action, not the window.
"""A bundle is a backup: it contains every action, not just the window.
SP6 changed the format assertion below, as this note said it would. It also
changed one more line in this file than the note allowed for — the
`variants` array in `test_export_keeps_retry_attempts`, which is the same
fact seen from the other side: a bundle that has coordinates has no use for
a repeating group. Everything else here still passes unmodified.
SP6 changed the format assertion below, as this note said it would. It
also changed one more line than the note allowed for: the `variants`
array in `test_export_keeps_retry_attempts`. That change reflects the
same fact: a bundle that stores coordinates has no use for a repeating
group. Everything else here still passes unmodified.
"""
ScriptedProvider.replies = ["One.", "Two."]
_play(client, "go north")
+25 -21
View File
@@ -1,20 +1,22 @@
"""Phase 14 SP9 — editing the take you are actually reading.
"""Phase 14 SP9: editing the take you are actually reading.
The pager can park a turn on take 2 of 4. The transcript row it sits in is
still keyed by the *live* take, because that is the row the story tells and the
one the window carries; the take being read is only in the client's hand, by
its own node id.
The pager can display a turn on take 2 of 4. The transcript row for that
turn stays keyed by the live take, because that is the row the transcript
reports and the one the window carries. The take currently displayed is
known only to the client, by its own node id.
So "edit this" has two ids to choose from, and the page shipped choosing the
wrong one: it opened the editor on the live take's text and saved over it, from
a row that was showing take 2. That is a client bug and it is fixed in
Play.jsx, but the fix rests on something only the server can promise —
So "edit this" has two ids to choose from, and the page shipped with the
wrong one. It opened the editor on the live take's text and saved over it,
even from a row that was displaying take 2. This is a client bug, and it
is fixed in Play.jsx. The fix depends on a guarantee only the server can
make:
a take is an ordinary row to the edit endpoint, addressed by its own id,
whether or not it is the one the path runs through
A take is an ordinary row to the edit endpoint, addressed by its own id,
whether or not it is the one the path runs through.
— and on the group listing telling the truth about it afterwards. Both are
asserted here so the client's fix cannot be quietly undermined.
The fix also depends on the group listing reporting this correctly
afterward. This file asserts both guarantees, so the client's fix cannot
be silently broken.
python -m pytest tests/test_take_edit.py -v
"""
@@ -119,7 +121,7 @@ def _edit(client, action_id, text):
def _live_ai(client):
"""The AI row the story currently tells, as the transcript reports it."""
"""The AI row the transcript currently reports as live."""
adv = client.get(f"/api/adventures/{client.adv_id}").json()
return [a for a in adv["actions"] if a["type"] == "ai"][-1]
@@ -145,8 +147,9 @@ def test_the_pager_reads_four_distinct_takes(client):
def test_editing_a_take_that_is_not_live_edits_that_take(client):
"""The bug, at the level the client's fix depends on.
Saving against take 2's own id must land on take 2 — not be refused for
being off the path, and not be redirected onto the live row.
Saving against take 2's own id must land on take 2. The request must
not be refused for being off the path, and must not be redirected onto
the live row.
"""
row, takes = _four_takes(client)
second = takes[1]
@@ -171,12 +174,13 @@ def test_editing_a_take_leaves_the_live_one_alone(client):
def test_a_take_edit_survives_paging_away_and_back(client):
"""The listing is the pager's only source, so the edit has to be in it.
"""The listing is the pager's only source, so the edit must appear in it.
(The client caches this list per message; the fix drops that cache after an
edit. If the server ever started answering from a copy of its own, stepping
away and back would show the words before the edit and nobody would see it
here — hence the round trip.)
The client caches this list per message, and the fix drops that cache
after an edit. If the server ever started answering from a stale copy
of its own, stepping away and back would show the text before the
edit, and this test would not catch it. That is why this test makes
the round trip.
"""
row, takes = _four_takes(client)
_edit(client, takes[1]["id"], "Take 2, rewritten.")
+38 -33
View File
@@ -1,18 +1,21 @@
"""Phase 14 SP9 — a turn's takes are grouped by their parent, not by where they sit.
"""Phase 14 SP9: a turn's takes are grouped by their parent, not by where
they sit.
SP4 made every take a node at the same (branch, depth). That coordinate answers
"which takes belong to this turn" right up until one of them is forked onto its
own branch — at which point it *leaves* the coordinate and reads as the only
take of its turn, with its siblings unreachable from the line it was taken on.
SP4 made every take a node at the same (branch, depth). That coordinate
answers "which takes belong to this turn" until one of them forks onto its
own branch. At that point it leaves the coordinate and reads as the only
take of its turn, with its siblings unreachable from the line it was taken
on.
The parent does not move when a branch does, which is the whole of the fix. It
also gets the nesting right for free: takes under C1 and takes under C2 share a
depth and, until one forks, a branch. Only the parent separates them.
The parent does not move when a branch does, which is the whole fix. It
also gets the nesting right without extra work: takes under C1 and takes
under C2 share a depth and, until one forks, a branch. Only the parent
separates them.
And the branch itself is lazy now. Stepping between takes creates nothing —
looking is free. The fork happens on the first thing *written* below a take the
story moved past, which is the first moment the player has said which line they
mean.
The branch itself is lazy now. Stepping between takes creates nothing:
looking is free. The fork happens on the first write below a take the
story moved past, which is the first moment the player has said which line
they mean.
python -m pytest tests/test_take_parentage.py -v
"""
@@ -124,7 +127,8 @@ def _branch_count(adv_id) -> int:
def _ai_rows(adv_id) -> list[models.Action]:
"""Every AI node ever written, oldest first — live or not, any branch."""
"""Every AI node ever written, oldest first. Includes both live and
discarded nodes, on any branch."""
db = SessionLocal()
try:
return (
@@ -152,7 +156,7 @@ def _group_size(action_id: int) -> int:
# ------------------------------------------------------------------ tests
def test_retaken_turn_groups_all_its_takes(client):
"""The baseline the rest of the file leans on: three takes, one turn."""
"""The baseline the rest of the file relies on: three takes, one turn."""
_play(client)
_retry(client)
_retry(client)
@@ -178,7 +182,7 @@ def test_a_forked_take_keeps_its_siblings(client):
first_take = _ai_rows(client.adv_id)[0]
assert first_take.live is False
# Writing below it is what forks -- see the next test.
# Writing below it is what forks. See the next test.
_play(client, "go back and try this instead", after_id=first_take.id)
assert _group_size(first_take.id) == 3, (
@@ -219,9 +223,9 @@ def test_writing_below_a_passed_take_forks_exactly_once(client):
def test_takes_under_one_parent_do_not_count_takes_under_its_sibling(client):
"""The player's own example: 3/3 on one line, 2/2 on the other.
C1 and C2 are takes of the same turn. What is played *below* each of them
is a different turn, and the two must not pool -- they share a depth, and
until the fork they share a branch too.
C1 and C2 are takes of the same turn. What is played below each of them
is a different turn, and the two turns must not merge. They share a
depth, and until the fork they share a branch too.
"""
_play(client)
_retry(client) # two takes at this turn: C1, C2 (C2 live)
@@ -276,12 +280,12 @@ def _done_action(response) -> dict:
def test_the_streamed_action_carries_the_pager_too(client):
"""Found by driving it, not by testing it.
"""Found while exercising the app manually, not by a targeted test.
A retry's reply *is* the second take of its turn, so it arrives needing a
pager. The stream builds its own ActionOut and so missed the annotation:
the pager appeared only once the page was reloaded, which is exactly the
moment nobody reloads.
A retry's reply is the second take of its turn, so it arrives needing a
pager. The stream builds its own ActionOut and missed the annotation.
The pager appeared only once the page was reloaded, which is exactly
the moment nobody reloads.
"""
_play(client)
r = client.post(f"/api/adventures/{client.adv_id}/retry")
@@ -365,11 +369,12 @@ def test_a_players_own_turn_can_be_played_again(client):
def test_a_retaken_player_turn_is_not_formatted_twice(client):
"""Found by driving it: "> You > You open the door."
"""Found while exercising the app manually: "> You > You open the door."
The editor is seeded from the stored text, which is already the formatted
form — the same text plain edit puts in the box and writes back verbatim.
Running it through the formatter again doubles the prefix.
The editor is seeded from the stored text, which is already in the
formatted form. Plain edit puts that same text in the box and writes
it back verbatim. Running it through the formatter again doubles the
prefix.
"""
_play(client, "open the door")
_play(client, "press on")
@@ -423,7 +428,7 @@ def test_an_ai_turn_at_the_tip_takes_no_branch(client):
def test_an_ai_turn_the_story_moved_past_takes_a_branch(client):
"""Retry could never reach here at all — it only ever saw the newest turn."""
"""Retry could never reach this case: it only ever saw the newest turn."""
_play(client)
_play(client, "press on")
before = _branch_count(client.adv_id)
@@ -445,7 +450,7 @@ def test_the_old_line_still_has_its_continuation(client):
first_ai = _ai_rows(client.adv_id)[0]
_take(client, first_ai.id, "")
# The new take is what this branch tells; "press on" belonged to the other.
# The new take is what this branch tells. "press on" belonged to the other.
blob = "\n".join(_path_texts(client))
assert "press on" not in blob
kept = "\n".join(a.text for a in _user_rows(client.adv_id))
@@ -471,10 +476,10 @@ def test_the_opening_has_no_other_take(client):
def test_retaking_even_the_newest_player_turn_forks(client):
"""A player turn is never the tip: the reply to it is.
Retaking the *last* thing the player typed still has a story to protect —
the AI answered it, and that answer was written for the old text. So the
branch is owed here too, and the guard against forking for nothing only
ever fires for a player action with no reply under it.
Retaking the last thing the player typed still puts existing story at
risk: the AI answered it, and that answer was written for the old
text. So this case must fork too. The guard against forking for
nothing only fires for a player action with no reply under it.
"""
_play(client, "open the door")
before = _branch_count(client.adv_id)
+23 -21
View File
@@ -1,15 +1,17 @@
"""Phase 14 SP9 — what a take does to the shared state.
"""Phase 14 SP9: what a take does to the shared state.
A turn does not only write text. A script mutates `script_state`, the referee
mutates `world_state`, and both are *shared* — they belong to the adventure, not
to the node. So playing a turn again has to put them back to where they were
before that turn ran, or the new take stacks its mutations on top of the one it
replaces and the numbers drift every time the player asks for another take.
A turn does not only write text. A script mutates `script_state`, and the
referee mutates `world_state`. Both are shared: they belong to the
adventure, not to the node. Playing a turn again must put them back to
where they were before that turn ran. Otherwise the new take stacks its
mutations on top of the one it replaces, and the numbers drift every time
the player asks for another take.
`retry` has done this since SP4 (`attempts.roll_back_before`). These are the
same guarantee for the two roads SP9 opened: a take of a turn the story moved
past, and a take of the player's own turn. Both create a branch, which is the
interesting part — the rollback has to survive leaving the line it was on.
`retry` has provided this guarantee since SP4 (`attempts.roll_back_before`).
These tests confirm the same guarantee for the two roads SP9 opened: a take
of a turn the story moved past, and a take of the player's own turn. Both
create a branch, which matters because the rollback must survive leaving
the line it was on.
python -m pytest tests/test_take_state.py -v
"""
@@ -34,9 +36,9 @@ from app.routers import adventures
SCHEMA = {"player": {"hp": {"min": 0, "max": 100, "initial": 100}}}
# Ten gold a turn, every turn. A number that only ever goes up is the clearest
# possible witness to a rollback: if a take stacks instead of replacing, it says
# so in one digit.
# Ten gold a turn, every turn. A number that only ever increases makes a
# rollback failure obvious: if a take stacks instead of replacing, the gold
# total is off by exactly one turn's worth.
GOLD_SCRIPT = """
const modifier = (text) => {
state.gold = (state.gold || 0) + 10;
@@ -146,7 +148,7 @@ def _rows(adv_id, type_):
def test_the_script_runs_once_a_turn(client):
"""The premise the rest of the file rests on."""
"""This test establishes the baseline the rest of the file depends on."""
_play(client)
assert _gold(client.adv_id) == 10
_play(client, "press on")
@@ -156,9 +158,9 @@ def test_the_script_runs_once_a_turn(client):
def test_a_take_of_a_past_ai_turn_does_not_stack_its_script(client):
"""Two turns played, then the first one taken again.
The take leaves the path just before turn one, so the state it starts from
is the state turn one started from — nothing, not the twenty that two turns
had accumulated. Then its own run adds ten.
The take leaves the path just before turn one. The state it starts from
is the state turn one started from: zero gold, not the twenty that two
turns accumulated. The take's own run then adds ten.
"""
_play(client)
_play(client, "press on")
@@ -182,11 +184,11 @@ def test_a_take_of_a_player_turn_does_not_stack_its_script(client):
def test_writing_below_a_passed_take_starts_from_that_take_s_state(client):
"""The `after_id` road, which forks on the way to writing.
"""The `after_id` path, which forks while writing.
The take being written under produced the state its own turn left behind —
ten — and the turn played on top of it adds the next ten. The twenty the
abandoned line reached has nothing to do with this branch.
The take being written under produced a state of ten gold, from its own
turn. The turn played on top of it adds another ten. The twenty gold
that the abandoned line reached has no effect on this branch.
"""
_play(client)
r = client.post(f"/api/adventures/{client.adv_id}/retry")
+101 -84
View File
@@ -1,16 +1,17 @@
"""Phase 14 SP1 — every existing adventure becomes a tree with one branch.
"""Phase 14 SP1: every existing adventure becomes a tree with one branch.
The migration this file watches is the one that cannot be re-run: it reads
`index` and writes `depth`, and from SP2 on the reads follow `depth`. If it
mis-maps a row, that row does not error — it *disappears from the story*, which
is why the assertions here are about every row rather than about a sample.
The migration this file tests cannot be re-run. It reads `index` and
writes `depth`, and from SP2 on, the reads follow `depth`. If it mis-maps
a row, that row does not raise an error. It disappears from the story,
which is why the assertions here check every row instead of a sample.
The fixture is a genuine **schema 45** database, not a current one with an old
stamp. `create_all` always builds the current schema, so the three tables the
tree touches are dropped and rebuilt from frozen pre-tree DDL below; the
migration then runs its real ALTERs against them, including the one that adds a
foreign key. A pre-migration database built any other way (stamp rewound,
columns left in place) would quietly skip the DDL and test half the change.
The fixture is a genuine schema-45 database, not a current one with an
old stamp. `create_all` always builds the current schema, so the three
tables the tree touches are dropped and rebuilt from the frozen pre-tree
DDL below. The migration then runs its real ALTER statements against
them, including the one that adds a foreign key. A pre-migration database
built any other way, such as a rewound stamp with the columns left in
place, would silently skip the DDL and test only half the change.
python -m pytest tests/test_tree_migration.py -v
"""
@@ -34,10 +35,11 @@ from app.context import history
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
# The three tables as they stood at schema 45, frozen. This is a snapshot of a
# past schema and must NOT be updated to track models.py — the whole point is
# that it lacks what SP1 adds. SQLite spelling only; the migration's Postgres
# half is exercised against a real server at deploy time (see plan/14).
# The three tables as they stood at schema 45, frozen. This is a snapshot
# of a past schema. It must not be updated to track `models.py`, because
# the whole point is that it lacks what SP1 adds. This DDL uses SQLite
# syntax only. The migration's Postgres half is exercised against a real
# server at deploy time (see plan/14).
PRE_TREE_DDL = (
"""
CREATE TABLE adventures (
@@ -96,15 +98,17 @@ PRE_TREE_DDL = (
""",
)
# The story of adventure "Gapped": index 3 is missing, because deleting a middle
# action never renumbered the ones after it. The gap has to survive as a gap.
# The story of adventure "Gapped": index 3 is missing, because deleting a
# middle action never renumbered the ones after it. The migration must
# preserve the gap.
GAPPED_INDEXES = (0, 1, 2, 4)
STRAIGHT_INDEXES = (0, 1)
# "Blank" holds an action whose text is nothing but whitespace. It is a row of
# the adventure but not of the *story*, so a cursor counting covered actions
# never counted it — and migration 56 has to skip it the same way, using a
# frozen copy of the story-text predicate. This is the one duplicated
# definition in the change, so it gets the one case that can tell.
# "Blank" holds an action whose text is nothing but whitespace. It is a
# row of the adventure but not of the story, so a cursor counting covered
# actions never counted it. Migration 56 must skip it the same way, using
# a frozen copy of the story-text predicate. This predicate is the one
# duplicated definition in the change, so this is the one test case that
# can catch it drifting.
BLANK_INDEXES = (0, 1, 2, 3)
BLANK_AT = 2
@@ -113,15 +117,16 @@ BLANK_AT = 2
def pre_tree():
"""A schema-45 database with three adventures in it, returned as the ids
(gapped, straight, empty) their stories were written under."""
# Every test in this module shares one temp file, and a setup that fails
# before its yield never reaches a teardown — so start from empty rather
# than from whatever the last one left.
# Every test in this module shares one temp file, and a setup that
# fails before its yield never reaches a teardown. Start from empty
# instead of from whatever the previous test left.
Base.metadata.drop_all(bind=engine)
Base.metadata.create_all(bind=engine)
with engine.begin() as conn:
# `branches` and the six new columns never existed at 45. Dropping the
# tables is the only way to lose the columns: SQLite refuses to drop a
# column a foreign key names, which is exactly the case for branch_id.
# `branches` and the six new columns never existed at schema 45.
# Dropping the tables is the only way to remove the columns.
# SQLite refuses to drop a column that a foreign key references,
# and that is exactly the case for `branch_id`.
for table in ("actions", "memories", "branches", "adventures"):
conn.execute(text(f"DROP TABLE IF EXISTS {table}"))
for ddl in PRE_TREE_DDL:
@@ -132,11 +137,12 @@ def pre_tree():
))
# The cursors as schema 45 held them: counts of covered story actions.
# Gapped's story is 0,1,2,4 — so "3 covered" is the node at depth 2 and
# "4 covered" is the node at depth 4, which is the whole reason a count
# and a depth are not the same number. Straight is caught up past its
# own end (5 covered, 2 actions), which is a state the older rule left
# behind and the clamp used to paper over every post-turn pass.
# Gapped's story is 0, 1, 2, 4, so "3 covered" is the node at depth
# 2 and "4 covered" is the node at depth 4. This is the whole
# reason a count and a depth are not the same number. Straight is
# caught up past its own end (5 covered, 2 actions), a state the
# older rule left behind that a clamp used to mask on every
# post-turn pass.
cursors_at = {"Gapped": (3, 4), "Straight": (5, 0), "Empty": (0, 0),
"Blank": (3, 0)}
ids = {}
@@ -276,18 +282,19 @@ def test_memories_attach_to_the_node_they_summarised(pre_tree):
assert depth == source_end, "the memory hangs off the last action it covered"
assert branch_id is not None
# A hand-written memory summarised no node, so SP7's migration 62 lands it
# at depth 0 of its branch. 0 is at or before every fork point, so it stays
# visible from exactly the paths it was visible from before — anchoring
# takes nothing out of anybody's existing bank.
# A hand-written memory summarised no node, so SP7's migration 62
# lands it at depth 0 of its branch. Depth 0 is at or before every
# fork point, so the memory stays visible from exactly the paths it
# was visible from before. Anchoring it this way removes nothing from
# anybody's existing memory bank.
manual = rows("SELECT depth, branch_id FROM memories WHERE source_end IS NULL")
assert manual and all(depth == 0 and branch is not None for depth, branch in manual)
def test_the_cursors_become_the_nodes_they_named(pre_tree):
"""SP3, migration 56. A count of covered actions and a depth are different
numbers the moment the story has a gap in it, which every adventure anyone
has ever deleted from does."""
"""SP3, migration 56. A count of covered actions and a depth become
different numbers as soon as the story has a gap in it, and every
adventure with a deleted action has one."""
migrations.bootstrap(engine)
def marks(title):
@@ -298,9 +305,9 @@ def test_the_cursors_become_the_nodes_they_named(pre_tree):
)
return row
# Gapped's story is 0,1,2,4. "3 covered" is the *third* action, at depth 2 —
# reading the count as a depth would have handed the summarizer node 3,
# which does not exist, and quietly skipped node 4 forever.
# Gapped's story is 0, 1, 2, 4. "3 covered" is the third action, at
# depth 2. Reading the count as a depth would hand the summarizer node
# 3, which does not exist, and silently skip node 4 forever.
memory_depth, summary_depth, memory_branch, summary_branch = marks("Gapped")
assert (memory_depth, summary_depth) == (2, 4)
root = scalar(
@@ -319,10 +326,11 @@ def test_the_cursors_become_the_nodes_they_named(pre_tree):
# Nothing covered stays nothing covered, and names no branch.
assert marks("Empty") == (migrations.NO_DEPTH, migrations.NO_DEPTH, None, None)
# A whitespace-only action is a row but not a story action, so it was never
# counted — "3 covered" of 0,1,[blank],3 is the node at depth 3, not 2. The
# migration's copy of the story-text predicate is the only place that rule
# is written twice, so this is the case that catches it drifting.
# A whitespace-only action is a row but not a story action, so it was
# never counted. "3 covered" of 0, 1, [blank], 3 is the node at depth
# 3, not 2. The migration's copy of the story-text predicate is the
# only place that rule is written twice, so this is the test case
# that catches it drifting.
assert marks("Blank")[0] == 3
# The legacy columns are left exactly as they were: a rolled-back build
@@ -333,8 +341,9 @@ def test_the_cursors_become_the_nodes_they_named(pre_tree):
def test_the_branch_clause_index_exists(pre_tree):
"""SP2's reads are only cheap if this exists — and `create_all` does not add
an index to a table it did not create, which is what migration 52 is for."""
"""SP2's reads are only cheap if this index exists. `create_all` does
not add an index to a table it did not create, which is what
migration 52 handles."""
migrations.bootstrap(engine)
assert scalar(
@@ -354,10 +363,10 @@ def test_running_it_again_changes_nothing(pre_tree):
rows("SELECT id, branch_id, depth FROM memories ORDER BY id"),
)
# Twice through the deploy path, then the data pass on its own — the stamp
# stops the first, the NULL guards stop the second, and a migration that
# only survives because of the stamp is one bad rescue away from doubling
# every branch.
# Run the deploy path twice, then the data pass on its own. The stamp
# stops the first run, the NULL guards stop the second, and a
# migration that only survives because of the stamp is one bad rescue
# away from doubling every branch.
migrations.bootstrap(engine)
with engine.begin() as conn:
migrations._backfill_tree(conn)
@@ -379,8 +388,8 @@ def test_running_it_again_changes_nothing(pre_tree):
def client(monkeypatch):
"""The app on a migrated database, so new rows go through the real writers.
Everything the migration fixes is only half the job: no migration will ever
visit a row written after it ran, and a row without a branch is a row no
Fixing existing rows is only half the job. No migration ever revisits
a row written after it ran, and a row without a branch is a row no
read can see.
"""
Base.metadata.drop_all(bind=engine)
@@ -442,11 +451,12 @@ def test_a_blank_adventure_has_a_branch_before_anything_is_played(client):
def test_a_hand_written_memory_is_anchored_at_the_head(client):
"""SP7: nothing carries a NULL depth any more.
"""SP7: nothing carries a NULL depth anymore.
On an adventure with no story yet the head is NO_DEPTH (-1), which reads as
"before the first node" and so is in range of every branch — right for a
note written before anything has happened.
On an adventure with no story yet, the head is `NO_DEPTH` (-1), which
reads as "before the first node" and so falls in range of every
branch. This is correct for a note written before anything has
happened.
"""
adventure_id = client.post("/api/adventures", json={}).json()["id"]
@@ -466,9 +476,10 @@ def test_a_hand_written_memory_is_anchored_at_the_head(client):
def test_deleting_a_branch_takes_its_nodes_with_it(client):
"""`ON DELETE CASCADE` on both `branch_id` columns, so the database removes a
branch's nodes rather than any code remembering to. SP7 ships delete-a-branch
on top of exactly this, and nothing else has to load a branch to do it."""
"""`ON DELETE CASCADE` on both `branch_id` columns means the database
removes a branch's nodes instead of relying on application code to
remember to. SP7 builds delete-a-branch directly on top of this, and
no other code needs to load a branch to do it."""
adventure_id = client.post(
"/api/adventures", json={"scenario_id": client.scenario_id}
).json()["id"]
@@ -509,8 +520,8 @@ def test_deleting_an_adventure_takes_its_branch_with_it(client):
def test_deleting_the_newest_action_moves_the_head_back(client):
"""The head is a cache, and a cache that only ever moves forward is wrong
the first time someone undoes a turn."""
"""The head is a cache. A cache that only ever moves forward becomes
wrong the first time someone undoes a turn."""
adventure_id = client.post(
"/api/adventures", json={"scenario_id": client.scenario_id}
).json()["id"]
@@ -540,10 +551,11 @@ def test_deleting_the_newest_action_moves_the_head_back(client):
# ---------------------------------------- SP4: variants become sibling rows
# One turn's retry history as schema 45 stored it: a JSON array on the AI row,
# with `variant_index` naming the entry `text` mirrors. The live one is
# deliberately not the last written — a migration that assumed it was would
# look right on every fixture where the player never went back.
# One turn's retry history as schema 45 stored it: a JSON array on the
# AI row, with `variant_index` naming the entry `text` mirrors. The live
# entry is deliberately not the last one written. A migration that
# assumed it was would still pass on every fixture where the player
# never paged back.
RETRY_VARIANTS = [
{"text": "Attempt one.", "reasoning": None,
"script_state": {"gold": 10}, "created_at": "2026-01-01T00:00:00",
@@ -563,10 +575,10 @@ RETRY_VARIANTS = [
]
LIVE_VARIANT = 1
# The whole turn's assembled prompt, stored once. The attempts differ only in
# the three slices above, which is the arrangement SP4 has to preserve — giving
# each sibling a copy of this would multiply the biggest column in the database
# by the retry count.
# The whole turn's assembled prompt, stored once. The attempts differ
# only in the three slices above, and SP4 must preserve that arrangement.
# Giving each sibling a copy of this prompt would multiply the largest
# column in the database by the retry count.
RETRY_SNAPSHOT = {
"sections": [{"label": "history", "text": "A long prompt.", "tokens": 4}],
"prompt": {"system": "S", "story": "A long prompt."},
@@ -578,11 +590,13 @@ RETRY_SNAPSHOT = {
@pytest.fixture()
def pre_split():
"""A schema-45 adventure with one retried turn, plus a plain turn each side.
"""A schema-45 adventure with one retried turn, plus a plain turn on
each side.
Separate from `pre_tree` so SP1's assertions keep counting what they were
written to count. The story is: 0 start, 1 do, 2 ai (three attempts), 3 do,
and the adventure's live state is the one attempt 1 produced.
This fixture is separate from `pre_tree` so SP1's assertions keep
counting what they were written to count. The story is: 0 start, 1
do, 2 ai (three attempts), 3 do. The adventure's live state is the
one attempt 1 produced.
"""
Base.metadata.drop_all(bind=engine)
Base.metadata.create_all(bind=engine)
@@ -603,8 +617,9 @@ def pre_split():
adventure_id = conn.execute(
text("SELECT id FROM adventures WHERE title = 'Retried'")
).scalar()
# `state_before` on each row: the scoreboard as that action found it.
# SP4 reads them one row along to build the `state_after` pair.
# `state_before` on each row records the script state as that
# action found it. SP4 reads each row's `state_before` from the
# next row to build the `state_after` pair.
for index, kind, before in (
(0, "start", None), (1, "do", {"gold": 0}),
(2, "ai", {"gold": 0}), (3, "do", {"gold": 20}),
@@ -653,7 +668,7 @@ def test_each_attempt_becomes_a_row_at_the_turns_coordinate(pre_split):
'AND "index" = 2', a=pre_split,
)
assert len(coordinates) == 1
# ...and the rest of the story is untouched, still one row per turn.
# The rest of the story is untouched, still one row per turn.
assert scalar("SELECT count(*) FROM actions WHERE adventure_id = :a", a=pre_split) == 6
@@ -703,10 +718,11 @@ def test_each_attempt_keeps_the_outcome_it_produced(pre_split):
]
assert parsed[0][2] == {"player": {"hp": 95}}
assert parsed[1][2] == {"player": {"hp": 60}}
# Attempt three recorded no world state — an adventure with no RPG layer,
# or a take made before the column existed. It stays NULL rather than
# borrowing a neighbour's, and switching to it leaves the RPG layer alone:
# exactly what `apply_variant` did with an entry that had no world state.
# Attempt three recorded no world state: either an adventure with no
# RPG layer, or a take made before the column existed. It stays NULL
# instead of borrowing a neighbor's, so switching to it leaves the RPG
# layer alone. This matches what `apply_variant` did with an entry
# that had no world state.
assert parsed[2][2] is None
@@ -739,7 +755,8 @@ def test_the_split_survives_being_run_again(pre_split):
def test_the_migrated_story_reads_back_as_one_turn(pre_split):
"""The point of all of it: the reads see a four-action story, not six."""
"""This is the point of the whole migration: reads see a four-action
story, not six."""
migrations.bootstrap(engine)
db = SessionLocal()
+9 -7
View File
@@ -1,8 +1,9 @@
"""End-to-end HTTP tests for undo/retry state revert, driving real turns through
the actual routes + scripting engine with only the LLM provider mocked.
"""End-to-end HTTP tests for undo and retry state revert. These tests drive
real turns through the actual routes and scripting engine, with only the LLM
provider mocked.
A script's output hook adds 10 gold each turn; we assert the shared scoreboard
behaves correctly across play / undo / retry.
A script's output hook adds 10 gold each turn. These tests confirm that the
adventure's stored gold total stays correct across play, undo, and retry.
python -m pytest tests/test_turn_flow_integration.py -v
"""
@@ -64,7 +65,7 @@ def client(monkeypatch):
adv_id, user_id = adv.id, user.id
setup.close()
# Force a real (non-demo) turn that uses our fake provider.
# Force a real, non-demo turn that uses the fake provider.
monkeypatch.setattr(adventures, "OpenAICompatibleProvider", FakeProvider)
monkeypatch.setattr(auth, "resolve_provider_config", lambda s: auth.ProviderConfig(
"http://fake", "k", "test-model", False))
@@ -107,7 +108,7 @@ def test_play_then_undo_reverts_gold(client):
r = client.post(f"/api/adventures/{client.adv_id}/undo")
assert r.status_code == 200, r.text
assert _state(client.adv_id) == {} # scoreboard rolled back
assert _state(client.adv_id) == {} # gold reverted to zero
def test_two_turns_then_undo_reverts_only_last(client):
@@ -123,7 +124,8 @@ def test_retry_does_not_double_apply_gold(client):
_play(client)
assert _state(client.adv_id) == {"gold": 10}
# Before the fix this produced 20 (output hook ran twice); now it stays 10.
# Before the fix, this produced 20 because the output hook ran twice.
# Now it stays 10.
r = client.post(f"/api/adventures/{client.adv_id}/retry")
assert r.status_code == 200, r.text
assert _state(client.adv_id) == {"gold": 10}
+13 -8
View File
@@ -1,5 +1,5 @@
"""Unit tests for the RPG world-state engine (Phase 12): delta extraction and
the clamp/cooldown/milestone referee.
the code that enforces clamp, cooldown, and milestone rules.
python -m pytest tests/test_worldstate.py -v
"""
@@ -67,7 +67,8 @@ def test_text_stat_noop_when_unchanged():
def test_per_npc_distinct_stats():
ws, _ = w.apply_delta(fresh(), SCHEMA, {"npc.drake.ferocity": 20}, 1)
assert ws["npc"]["drake"]["ferocity"] == 70
# gwen has no "ferocity" stat, drake has no "trust" — cross paths are rejected.
# gwen has no `ferocity` stat, and drake has no `trust` stat. Cross paths
# are rejected.
ws, report = w.apply_delta(ws, SCHEMA, {"npc.gwen.ferocity": 5, "npc.bogus.trust": 5}, 2)
reasons = {r["reason"] for r in report["rejected"]}
assert reasons == {"unknown npc stat", "unknown npc"}
@@ -134,7 +135,8 @@ def test_flags_toggle_both_ways():
ws, report = w.apply_delta(ws, SCHEMA, {"flags.has_key": True}, 1)
assert ws["flags"]["has_key"] is True
assert report["applied"]
# flip back off — flags are two-way (unlike sticky milestones).
# Setting it back to false works, because flags are two-way, unlike
# sticky milestones.
ws, _ = w.apply_delta(ws, SCHEMA, {"flags.has_key": False}, 2)
assert ws["flags"]["has_key"] is False
# setting to the same value is a no-op.
@@ -150,8 +152,9 @@ def test_flag_rejects_non_bool_and_unknown():
def test_override_sets_absolute_value_bypassing_cap():
# A manual override isn't policed by max_delta_per_turn like a turn is —
# it sets the value directly (still clamped to min/max).
# A manual override does not obey max_delta_per_turn the way a turn
# does. It sets the value directly, though the value is still clamped
# to min/max.
ws, report = w.apply_override(fresh(), SCHEMA, {"player.hp": 10})
assert ws["player"]["hp"] == 10
assert report["applied"] == [{"path": "player.hp", "old": 100, "new": 10}]
@@ -202,9 +205,10 @@ def test_override_rejects_unknown_and_bad_type():
def test_reference_includes_desc_and_bands_independently():
guide = w.render_reference(SCHEMA)
# hp has both a description and a band ladder.
# hp has both a description and a set of value bands.
assert "very weak" in guide and "range 0–100" in guide
# day (a counter here has no desc/bands) contributes nothing; flags show desc.
# A counter like day has no desc or bands, so it contributes nothing
# here. Flags do show their desc.
assert "has_key (flag) — Holds the key." in guide
# NPCs contribute their own description and per-NPC stat lines.
assert "NPC Gwen (gwen) — A loyal ranger." in guide
@@ -232,7 +236,8 @@ def test_extract_no_block():
def test_extract_prose_ending_in_brace_not_eaten():
# A bare object with no dotted keys is not a delta — leave the text alone.
# A bare object with no dotted keys is not a delta, so the text stays
# as is.
clean, delta = w.extract_delta('He said {this}')
assert delta == {}
assert clean == "He said {this}"
+5 -4
View File
@@ -36,8 +36,9 @@ SCHEMA = {
"milestones": {"win": {"desc": "Win the fight"}},
}
# The faked model narrates and appends a delta that exceeds the per-turn cap
# (so we can see the engine clamp it), flips a flag, and completes a milestone.
# The faked model narrates and appends a delta that exceeds the per-turn
# cap, so the test can confirm the engine clamps it. It also flips a flag
# and completes a milestone.
AI_REPLY = (
"The goblin's blade bites deep and Gwen nods at your resolve.\n\n"
'```state\n{"player.hp": -80, "npc.gwen.trust": 15, "flags.alarm": true, "milestones.win": true}\n```'
@@ -131,7 +132,7 @@ def test_turn_applies_clamped_delta_and_strips_block(client):
# The state block is not shown to the player.
assert "```state" not in _last_ai_text(client.adv_id)
assert "goblin's blade" in _last_ai_text(client.adv_id)
# ...but the raw model reply (with the block) is kept for the Insights view.
# The raw model reply, including the block, is kept for the Insights view.
db = SessionLocal()
try:
snap = db.get(models.Adventure, client.adv_id).actions[-1].context_snapshot
@@ -175,7 +176,7 @@ def test_override_world_state_endpoint(client):
# persisted to the DB, not just the response.
assert _world(client.adv_id)["player"]["hp"] == 5
# bypasses max_delta_per_turn (30) — a direct correction, not a turn.
# This bypasses max_delta_per_turn (30) because it is a direct correction, not a turn.
r = client.put(f"/api/adventures/{client.adv_id}/world-state", json={"player.hp": 100})
assert r.json()["state"]["player"]["hp"] == 100