Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences.
This commit is contained in:
+222
-194
@@ -1,4 +1,4 @@
|
||||
"""Drive a production-sized adventure through the real routes and report what
|
||||
"""Drives a production-sized adventure through the real routes and reports what
|
||||
each one costs in database bytes.
|
||||
|
||||
cd backend
|
||||
@@ -6,43 +6,42 @@ each one costs in database bytes.
|
||||
.venv/Scripts/python.exe -m tools.stress_session --actions 200 --memories 100
|
||||
.venv/Scripts/python.exe -m tools.stress_session --no-embeddings
|
||||
|
||||
**The memory bank is ON by default, and that is the point.** The round-two
|
||||
stress harness ran without an embedding model configured, and embedding
|
||||
providers are BYOK-only by construction, so `retrieve_memories` returned early
|
||||
every time — the whole exercise measured the turn loop with its heaviest read
|
||||
switched off, and reported 23 MB for a playthrough that actually costs an order
|
||||
of magnitude more. `--no-embeddings` reproduces that blindness deliberately, to
|
||||
show the gap; it is never the default and it prints a warning.
|
||||
The memory bank is on by default, and that matters. The round-two stress
|
||||
harness ran with no embedding model configured, and embedding providers are
|
||||
BYOK-only by construction, so `retrieve_memories` returned early every time.
|
||||
That run measured the turn loop with its heaviest read disabled and reported
|
||||
23 MB for a playthrough that costs an order of magnitude more.
|
||||
`--no-embeddings` reproduces that configuration on purpose, to show the gap. It
|
||||
is never the default, and it prints a warning.
|
||||
|
||||
Everything here is synthetic. The fixture is generated to production *shape* —
|
||||
1536-dimension embeddings, ~74 KB context snapshots, retry variants — and no
|
||||
real adventure, user or backup is ever read.
|
||||
Everything here is synthetic. The fixture is generated to the shape of
|
||||
production data, with 1536-dimension embeddings, context snapshots of about
|
||||
74 kB, and retry variants. It never reads a real adventure, user, or backup.
|
||||
|
||||
Only the network is faked: the LLM and the embedding endpoint. Routing,
|
||||
sessions, the ORM, the scripting engine and the context builder are the real
|
||||
ones, because the bugs this exists to catch live in exactly the layer a mock
|
||||
would replace.
|
||||
Only the network is faked, which means the LLM and the embedding endpoint.
|
||||
Routing, sessions, the ORM, the scripting engine, and the context builder are
|
||||
the real ones, because the bugs this harness exists to catch are in the layer a
|
||||
mock would replace.
|
||||
|
||||
It runs on a throwaway SQLite file by default. What is being measured is which
|
||||
columns of which rows a code path asks for, and that is decided by the ORM,
|
||||
identically on both dialects. The dialects disagree on how a value is encoded
|
||||
on the wire — JSON especially — so treat the absolute figures as
|
||||
production-shaped rather than production-exact, and compare before against
|
||||
after.
|
||||
It runs on a throwaway SQLite file by default. The measurement is which columns
|
||||
of which rows a code path requests, and the ORM decides that identically on both
|
||||
dialects. The dialects differ in how a value is encoded on the wire, especially
|
||||
JSON, so treat the absolute figures as production-shaped rather than
|
||||
production-exact, and compare one run against another.
|
||||
|
||||
To measure the encodings SQLite cannot reach — bytea for the packed vectors,
|
||||
and json columns psycopg parses before the meter sees them — set
|
||||
AIDND_STRESS_DATABASE_URL to a **throwaway** Postgres database:
|
||||
To measure the encodings SQLite cannot reach, which are bytea for the packed
|
||||
vectors and json columns that psycopg parses before the meter sees them, set
|
||||
`AIDND_STRESS_DATABASE_URL` to a throwaway Postgres database:
|
||||
|
||||
AIDND_STRESS_DATABASE_URL=postgresql://…/stress_scratch \
|
||||
.venv/Scripts/python.exe -m tools.stress_session
|
||||
|
||||
The harness writes, so it refuses any target whose database name does not say
|
||||
'stress' or 'scratch'. Never point it at a database holding real users.
|
||||
The harness writes, so it refuses any target whose database name does not
|
||||
contain 'stress' or 'scratch'. Never point it at a database holding real users.
|
||||
|
||||
Calibration. The fixture is sized from production, re-measured 2026-08-17
|
||||
against the live Neon database (aggregates only — counts and octet_length
|
||||
sums, never row contents):
|
||||
Calibration. The fixture is sized from production and was re-measured on
|
||||
2026-08-17 against the live Neon database, reading aggregates only: counts and
|
||||
`octet_length` sums, never row contents.
|
||||
|
||||
per action, text 886 B -> --narration-bytes 1700, alternating
|
||||
with a one-line player input
|
||||
@@ -51,14 +50,13 @@ sums, never row contents):
|
||||
memory bank, largest 100 memories, 6,144 B a vector
|
||||
|
||||
The previous defaults were wrong in both directions at once and happened to
|
||||
land near the right total: actions were modelled at ~2.1 KB against a real
|
||||
886 B, and stories at 200 actions against a real 607. Width was flattering,
|
||||
length was not, and length is what a page load pays for.
|
||||
land near the right total. Actions were modeled at about 2.1 kB against a real
|
||||
886 B, and stories at 200 actions against a real 607. The width was too large
|
||||
and the length was too small, and the length is what a page load pays for.
|
||||
|
||||
Filler text is generated word by word rather than repeated. A repeated
|
||||
sentence compresses about a hundredfold and prose three- or fourfold, so the
|
||||
old fixture would have made any compression measurement on context_snapshot
|
||||
meaningless.
|
||||
Filler text is generated word by word rather than repeated. A repeated sentence
|
||||
compresses about a hundredfold and prose three- or fourfold, so the old fixture
|
||||
would have made any compression measurement on `context_snapshot` meaningless.
|
||||
"""
|
||||
|
||||
import os
|
||||
@@ -68,11 +66,12 @@ from pathlib import Path
|
||||
|
||||
|
||||
def _early_keep(argv: list[str]) -> str:
|
||||
"""--keep, read before argparse exists.
|
||||
"""Reads `--keep` before argparse exists.
|
||||
|
||||
Where the database lives has to be decided before app.database is
|
||||
imported, and that import is three lines below. argparse still declares
|
||||
the flag, so --help documents it and a typo is still an error."""
|
||||
The database location has to be decided before `app.database` is imported,
|
||||
and that import is a few lines below. argparse still declares the flag, so
|
||||
`--help` documents it and a typo is still an error.
|
||||
"""
|
||||
for i, arg in enumerate(argv):
|
||||
if arg == "--keep" and i + 1 < len(argv):
|
||||
return argv[i + 1]
|
||||
@@ -83,17 +82,17 @@ def _early_keep(argv: list[str]) -> str:
|
||||
|
||||
_keep = _early_keep(sys.argv[1:])
|
||||
|
||||
# Must precede the app import: database.py reads these at module scope.
|
||||
# This has to run before the app import, because `database.py` reads these at
|
||||
# module scope.
|
||||
#
|
||||
# Default is a throwaway SQLite file. AIDND_STRESS_DATABASE_URL points the
|
||||
# The default is a throwaway SQLite file. `AIDND_STRESS_DATABASE_URL` points the
|
||||
# harness at a real Postgres instead, which is the only way to reach the
|
||||
# encodings SQLite cannot exercise: bytea for the packed vectors, and json
|
||||
# columns that psycopg parses into Python before the meter ever sees them.
|
||||
# columns that psycopg parses into Python before the meter sees them.
|
||||
#
|
||||
# The name guard is not paranoia. This harness *writes* — it builds a whole
|
||||
# synthetic adventure — so a URL that happened to point at the production
|
||||
# database would quietly seed it with fake users and fake play. The target
|
||||
# must say it is disposable.
|
||||
# The name guard matters. This harness writes a whole synthetic adventure, so a
|
||||
# URL that pointed at the production database would seed it with fake users and
|
||||
# fake play. The target has to name itself as disposable.
|
||||
_stress_url = os.environ.get("AIDND_STRESS_DATABASE_URL", "").strip()
|
||||
if _stress_url:
|
||||
if _keep:
|
||||
@@ -111,10 +110,10 @@ if _stress_url:
|
||||
os.environ["AIDND_DATABASE_URL"] = _stress_url
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
elif _keep:
|
||||
# A fixture to boot the app against rather than a temp file the report
|
||||
# discards. Rebuilt from empty every run: build_fixture() assumes an empty
|
||||
# database on the SQLite path, and a second run would otherwise stack a
|
||||
# second adventure beside the first.
|
||||
# A fixture to start the app against, rather than a temporary file the
|
||||
# report discards. Every run rebuilds it from empty, because
|
||||
# `build_fixture()` assumes an empty database on the SQLite path and a
|
||||
# second run would otherwise add a second adventure next to the first.
|
||||
_keep_path = Path(_keep).resolve()
|
||||
_keep_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
_keep_path.unlink(missing_ok=True)
|
||||
@@ -149,21 +148,22 @@ from .fakeprose import prose
|
||||
|
||||
EMBEDDING_DIMS = 1536
|
||||
|
||||
# ~232 KB, measured on production's largest adventure (2026-08-17). The old
|
||||
# figure here was 74 KB, taken from the comment in models.py; the real column
|
||||
# averages 163 KB a row across the whole table and 232 KB on the adventure that
|
||||
# matters, because the assembled prompt grows with the story behind it.
|
||||
# About 232 kB, measured on the largest adventure in production on 2026-08-17.
|
||||
# The old figure here was 74 kB, taken from the comment in `models.py`. The real
|
||||
# column averages 163 kB per row across the whole table and 232 kB on the
|
||||
# adventure that matters, because the assembled prompt grows with the story
|
||||
# behind it.
|
||||
#
|
||||
# Built from varied text rather than one sentence repeated. A repeated sentence
|
||||
# compresses about a hundredfold and real prose three- or fourfold, so a
|
||||
# fixture made of repeats would make any compression measurement meaningless —
|
||||
# and shrinking this column is the open question it exists to answer.
|
||||
# The text is varied rather than one sentence repeated. A repeated sentence
|
||||
# compresses about a hundredfold and real prose three- or fourfold, so a fixture
|
||||
# built from repeats would make any compression measurement meaningless, and
|
||||
# shrinking this column is the open question the fixture exists to answer.
|
||||
SNAPSHOT_SYSTEM = None # set by _build_text()
|
||||
SNAPSHOT_STORY = None
|
||||
|
||||
PLAYER_INPUT = "> You crouch and look more closely at the grit on the floor."
|
||||
|
||||
# All three are bound by _build_text() from the fixture arguments.
|
||||
# `_build_text()` sets all three from the fixture arguments.
|
||||
NARRATION = None
|
||||
MEMORY_TEXT = (
|
||||
"You found a bandit camp above the ford and agreed to guide Gwen through "
|
||||
@@ -172,15 +172,15 @@ MEMORY_TEXT = (
|
||||
|
||||
|
||||
def _build_text(args, rng: random.Random) -> None:
|
||||
"""Size the three variable-length fixture strings from the arguments.
|
||||
"""Sizes the three variable-length fixture strings from the arguments.
|
||||
|
||||
Separate from build_fixture so the sizes are decided once, before anything
|
||||
is written, and so a shape's cost is a function of the flags rather than of
|
||||
how many rows happened to be generated first.
|
||||
This is separate from `build_fixture` so that the sizes are decided once,
|
||||
before anything is written, and so that a shape's cost depends on the flags
|
||||
rather than on how many rows were generated first.
|
||||
"""
|
||||
global NARRATION, SNAPSHOT_SYSTEM, SNAPSHOT_STORY
|
||||
NARRATION = prose(rng, args.narration_bytes)
|
||||
# The assembled prompt is a system block and the story so far; the split
|
||||
# The assembled prompt is a system block plus the story so far. The split
|
||||
# is roughly one to five in production.
|
||||
SNAPSHOT_SYSTEM = prose(rng, args.snapshot_bytes // 6)
|
||||
SNAPSHOT_STORY = prose(rng, args.snapshot_bytes - args.snapshot_bytes // 6)
|
||||
@@ -190,7 +190,11 @@ def _build_text(args, rng: random.Random) -> None:
|
||||
|
||||
|
||||
class FakeProvider:
|
||||
"""The LLM. Streams one fixed line; no network, no cost, no variance."""
|
||||
"""Stands in for the LLM, streaming one fixed line.
|
||||
|
||||
It makes no network call, costs nothing, and returns the same text every
|
||||
time.
|
||||
"""
|
||||
|
||||
def __init__(self, *a, **k):
|
||||
pass
|
||||
@@ -203,8 +207,11 @@ class FakeProvider:
|
||||
|
||||
|
||||
class FakeEmbeddings:
|
||||
"""The embedding endpoint. Returns vectors of the real width, so what the
|
||||
turn writes back weighs what production weighs."""
|
||||
"""Stands in for the embedding endpoint.
|
||||
|
||||
It returns vectors of the production width, so what the turn writes back is
|
||||
the size production writes back.
|
||||
"""
|
||||
|
||||
def __init__(self, rng: random.Random):
|
||||
self.rng = rng
|
||||
@@ -218,32 +225,34 @@ class FakeEmbeddings:
|
||||
|
||||
# -------------------------------------------------------------- rich fixture
|
||||
|
||||
# The correctness fixture, as against the scale one.
|
||||
# The correctness fixture, as opposed to the scale fixture.
|
||||
#
|
||||
# The default fixture is sized from production and exists to weigh bytes, so it
|
||||
# leaves every column it does not weigh at its default. That makes it a poor
|
||||
# witness for anything *semantic* — and phase 14 replaces the storage model
|
||||
# underneath all of it. Measured on a freshly built default fixture, the
|
||||
# columns a story tree has to migrate correctly look like this:
|
||||
# The default fixture is sized from production and exists to measure bytes, so it
|
||||
# leaves every column it does not measure at its default. That makes it a poor
|
||||
# test of semantics, and phase 14 replaces the storage model underneath all of
|
||||
# it. On a freshly built default fixture, the columns a story tree has to migrate
|
||||
# correctly hold this:
|
||||
#
|
||||
# state_after / world_state_after the same value on all 600 rows (they
|
||||
# state_after, world_state_after The same value on all 600 rows. They
|
||||
# were state_before, and NULL, until SP4
|
||||
# turned them round). A rollback over
|
||||
# identical snapshots proves nothing.
|
||||
# scenario_id / world_state absent. No RPG layer, so the cooldown
|
||||
# clock SP5 must not advance never runs.
|
||||
# adventure_scripts none. script_state rollback is exactly
|
||||
# what a branch switch reuses.
|
||||
# memory_cursor / summary_cursor both 0. SP3 replaces the cursors.
|
||||
# sibling attempts two of byte-identical text with the
|
||||
# first always live — so "which attempt
|
||||
# is live?", the one question SP4 has to
|
||||
# answer, has no observable answer.
|
||||
# reversed them. A rollback over identical
|
||||
# snapshots tests nothing.
|
||||
# scenario_id, world_state Absent. There is no RPG layer, so the
|
||||
# cooldown clock SP5 must not advance
|
||||
# never runs.
|
||||
# adventure_scripts None. A branch switch reuses the
|
||||
# script_state rollback.
|
||||
# memory_cursor, summary_cursor Both 0. SP3 replaces the cursors.
|
||||
# sibling attempts Two attempts of byte-identical text with
|
||||
# the first always live, so the one
|
||||
# question SP4 has to answer, which
|
||||
# attempt is live, has no observable
|
||||
# answer.
|
||||
#
|
||||
# --rich fills in exactly those and changes nothing else, so the measuring
|
||||
# fixture's numbers stay comparable run to run. It is a *correctness* fixture:
|
||||
# prefer it small (--rich --actions 30), because what it is for is variety per
|
||||
# row, not rows.
|
||||
# `--rich` fills in those columns and changes nothing else, so the measuring
|
||||
# fixture's numbers stay comparable from run to run. It is a correctness fixture,
|
||||
# so keep it small, such as `--rich --actions 30`. It provides variety per row
|
||||
# rather than many rows.
|
||||
|
||||
RICH_SUMMARY = (
|
||||
"You tracked the bandits to a camp above the ford, freed Gwen from the "
|
||||
@@ -260,8 +269,8 @@ RICH_CARDS = [
|
||||
"Taken from the quartermaster. Opens something below the camp."),
|
||||
]
|
||||
|
||||
# Ten gold a turn, so a state snapshot that failed to roll back reads as a
|
||||
# wrong total rather than as nothing. Same shape the retry tests use.
|
||||
# Ten gold a turn, so a state snapshot that failed to roll back shows a wrong
|
||||
# total rather than no value. The retry tests use the same shape.
|
||||
RICH_SCRIPT = """
|
||||
const modifier = (text) => {
|
||||
state.gold = (state.gold || 0) + 10;
|
||||
@@ -272,28 +281,31 @@ modifier(text);
|
||||
|
||||
|
||||
def rich_stat_schema() -> dict:
|
||||
"""The demo RPG schema, read from the seed data rather than invented here.
|
||||
"""Returns the demo RPG schema, read from the seed data rather than invented.
|
||||
|
||||
Using the real one means the fixture exercises bands, cooldowns,
|
||||
max_delta_per_turn and NPC stat blocks as they are actually shaped — an
|
||||
invented schema would drift from the thing it is standing in for.
|
||||
Using the real schema means the fixture exercises bands, cooldowns,
|
||||
`max_delta_per_turn`, and NPC stat blocks in their real shapes. An invented
|
||||
schema would diverge from the one it stands in for.
|
||||
"""
|
||||
path = seed.SEED_DIR / "04-rpg-world-state.json"
|
||||
return json.loads(path.read_text(encoding="utf-8"))["stat_schema"]
|
||||
|
||||
|
||||
def rich_script_state(turn: int) -> dict:
|
||||
"""The scoreboard as of `turn`. Monotonic, so any snapshot identifies the
|
||||
turn it was taken at — which is what makes a bad rollback visible."""
|
||||
"""Returns the script state as of `turn`.
|
||||
|
||||
The values increase with the turn, so any snapshot identifies the turn it was
|
||||
taken at, which is what makes an incorrect rollback visible.
|
||||
"""
|
||||
return {"gold": turn * 10, "turn": turn}
|
||||
|
||||
|
||||
def rich_world_state(schema: dict, turn: int) -> dict:
|
||||
"""A live world state that has actually been played to `turn`.
|
||||
"""Returns a live world state that has been played to `turn`.
|
||||
|
||||
`instantiate` gives the initial picture; a fixture whose every row holds
|
||||
that same picture cannot tell a restored snapshot from an unrestored one.
|
||||
So hp declines, mana drains and a flag flips partway through.
|
||||
`instantiate` returns the initial state, and a fixture whose every row holds
|
||||
that same state cannot distinguish a restored snapshot from an unrestored
|
||||
one. Here hp declines, mana declines, and a flag changes partway through.
|
||||
"""
|
||||
ws = worldstate.instantiate(schema)
|
||||
ws["player"]["hp"] = max(20, 100 - turn)
|
||||
@@ -306,28 +318,28 @@ def rich_world_state(schema: dict, turn: int) -> dict:
|
||||
|
||||
|
||||
def rich_attempts(rng: random.Random, index: int) -> tuple[list[str], int]:
|
||||
"""Distinguishable retry attempts, and which one the story tells.
|
||||
"""Returns distinguishable retry attempts, and which one the story tells.
|
||||
|
||||
Every attempt in the default fixture carries the same text with the live
|
||||
one pinned at 0. That is the one thing SP4 has to get right and the one
|
||||
thing that fixture cannot witness, so here the texts differ, the counts
|
||||
differ, and the live one is often not the last.
|
||||
Every attempt in the default fixture carries the same text, with the live one
|
||||
fixed at index 0. That is the behavior SP4 has to get right and the behavior
|
||||
that fixture cannot test, so here the texts differ, the counts differ, and
|
||||
the live attempt is often not the last.
|
||||
"""
|
||||
count = 3 if index % 12 == 1 else 2
|
||||
texts = [
|
||||
f"[attempt {n + 1} of {count} at turn {index}] {prose(rng, 240)}"
|
||||
for n in range(count)
|
||||
]
|
||||
# Deliberately not always the newest: a player who retried twice and then
|
||||
# went back to the first take is the case that breaks anything assuming the
|
||||
# live attempt is the last one written.
|
||||
# The live attempt is not always the newest. A player who retried twice and
|
||||
# then went back to the first attempt is the case that breaks code assuming
|
||||
# the live attempt is the last one written.
|
||||
return texts, (0 if index % 18 == 1 else count - 1)
|
||||
|
||||
|
||||
def add_rich_extras(db, args, rng: random.Random, user, adventure) -> None:
|
||||
"""Everything --rich adds beside the actions themselves.
|
||||
"""Adds everything `--rich` contributes apart from the actions themselves.
|
||||
|
||||
Called with the adventure already flushed, before the actions are written,
|
||||
The caller flushes the adventure first and writes the actions afterwards,
|
||||
because the actions need the schema to snapshot a world state from.
|
||||
"""
|
||||
schema = rich_stat_schema()
|
||||
@@ -343,11 +355,11 @@ def add_rich_extras(db, args, rng: random.Random, user, adventure) -> None:
|
||||
adventure.scenario_id = scenario.id
|
||||
adventure.world_state = rich_world_state(schema, args.actions)
|
||||
adventure.script_state = rich_script_state(args.actions)
|
||||
# The post-turn passes only do anything when summarization is on and the
|
||||
# marks are somewhere other than the start. Written as positions here and
|
||||
# translated into anchors once the actions exist (see build_fixture) — a
|
||||
# position is what a database being migrated to SP3 still holds, so the
|
||||
# fixture carries both and they have to say the same thing.
|
||||
# The post-turn passes do nothing unless summarization is on and the marks
|
||||
# are past the start. They are written here as positions and translated into
|
||||
# anchors once the actions exist. See `build_fixture`. A database being
|
||||
# migrated to SP3 still holds positions, so the fixture carries both forms
|
||||
# and the two have to agree.
|
||||
adventure.auto_summarize = True
|
||||
adventure.story_summary = RICH_SUMMARY
|
||||
adventure.memory_cursor = max(0, args.actions - 8)
|
||||
@@ -366,10 +378,12 @@ def add_rich_extras(db, args, rng: random.Random, user, adventure) -> None:
|
||||
|
||||
|
||||
def add_second_adventure(db, rng: random.Random, user) -> int:
|
||||
"""A short second adventure, so 'does this leak across adventures?' is a
|
||||
question the fixture can answer. A tree scopes every read by branch, and a
|
||||
branch clause that forgot its adventure would still look right on a
|
||||
database holding exactly one."""
|
||||
"""Adds a short second adventure, so the fixture can detect cross-adventure
|
||||
leaks.
|
||||
|
||||
A tree scopes every read by branch, and a branch clause that omitted its
|
||||
adventure would still look correct on a database holding one adventure.
|
||||
"""
|
||||
other = models.Adventure(
|
||||
user_id=user.id, title="Stress (second)", script_state={},
|
||||
memory_bank_enabled=True, auto_summarize=False,
|
||||
@@ -398,12 +412,12 @@ def add_second_adventure(db, rng: random.Random, user) -> int:
|
||||
|
||||
|
||||
def _assert_live_variant_invariant(db, adventure_id: int) -> None:
|
||||
"""Exactly one attempt per turn is live, on every turn.
|
||||
"""Checks that exactly one attempt per turn is live, on every turn.
|
||||
|
||||
Checked here rather than trusted, because a coordinate with two live
|
||||
siblings tells its story twice and a coordinate with none drops a turn out
|
||||
of it — and both fail by *reading* wrong, never by raising. A fixture that
|
||||
quietly violated it would let wrong code look right.
|
||||
The check runs here rather than being assumed, because a coordinate with two
|
||||
live siblings renders its turn twice and a coordinate with none omits the
|
||||
turn. Both failures produce wrong output rather than an exception, so a
|
||||
fixture that broke the rule would let incorrect code look correct.
|
||||
"""
|
||||
groups: dict[tuple, list[models.Action]] = {}
|
||||
for action in (
|
||||
@@ -430,13 +444,15 @@ def _assert_live_variant_invariant(db, adventure_id: int) -> None:
|
||||
|
||||
|
||||
def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
"""A user, settings and one adventure at production scale. Returns
|
||||
(adventure_id, user_id)."""
|
||||
# A SQLite run gets a brand-new temp file every time, so the fixture can
|
||||
# assume an empty database. A Postgres scratch target persists between
|
||||
# runs, and the second one would collide on the fixture user's unique
|
||||
# email — so empty it first. Only ever reached for a target whose name
|
||||
# passed the 'stress'/'scratch' guard at the top of this module.
|
||||
"""Builds a user, settings, and one adventure at production scale.
|
||||
|
||||
Returns `(adventure_id, user_id)`.
|
||||
"""
|
||||
# A SQLite run gets a new temporary file every time, so the fixture can
|
||||
# assume an empty database. A Postgres scratch target persists between runs,
|
||||
# and a second run would collide on the fixture user's unique email, so empty
|
||||
# it first. This runs only for a target whose name passed the 'stress' or
|
||||
# 'scratch' guard at the top of this module.
|
||||
if _stress_url:
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
Base.metadata.create_all(bind=engine)
|
||||
@@ -450,7 +466,7 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
api_key=security.encrypt_secret("stress-key"),
|
||||
model="stress-model",
|
||||
endpoint_url="https://fake.invalid/v1",
|
||||
# The default this whole tool exists to stop anyone forgetting.
|
||||
# The default this tool exists to keep anyone from forgetting.
|
||||
embedding_model="" if args.no_embeddings else "openai/text-embedding-3-small",
|
||||
memory_bank_capacity=args.capacity,
|
||||
))
|
||||
@@ -459,8 +475,9 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
title="Stress",
|
||||
script_state={},
|
||||
memory_bank_enabled=True,
|
||||
# Off so a turn measures the turn. The post-turn pass is its own
|
||||
# shape below; letting it fire mid-measurement would mix the two.
|
||||
# Off, so that a turn measures only the turn. The post-turn pass
|
||||
# is its own shape below, and running it during the measurement
|
||||
# would combine the two.
|
||||
auto_summarize=False,
|
||||
)
|
||||
db.add(adventure)
|
||||
@@ -472,7 +489,7 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
is_ai = bool(i % 2)
|
||||
retried = is_ai and i % 6 == 1
|
||||
if args.rich:
|
||||
# Distinguishable attempts, and a live one that is often not
|
||||
# Distinguishable attempts, with a live one that is often not
|
||||
# the last written.
|
||||
texts, live = rich_attempts(rng, i) if retried else ([], 0)
|
||||
if not retried:
|
||||
@@ -490,10 +507,10 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
type="ai" if is_ai else "do",
|
||||
text=body,
|
||||
# The turn's assembled prompt is stored once, on the
|
||||
# attempt the story tells; a superseded sibling keeps only
|
||||
# its own slices. That is the invariant `app/attempts.py`
|
||||
# maintains, and a fixture that ignored it would multiply
|
||||
# the biggest column in the database by the retry count.
|
||||
# attempt the story tells. A superseded sibling keeps only
|
||||
# its own slices. `app/attempts.py` maintains that
|
||||
# invariant, and a fixture that ignored it would multiply
|
||||
# the largest column in the database by the retry count.
|
||||
context_snapshot=(
|
||||
{"system": SNAPSHOT_SYSTEM, "story": SNAPSHOT_STORY}
|
||||
if n == live else
|
||||
@@ -506,20 +523,21 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
live=(n == live),
|
||||
variant_index=n,
|
||||
variant_count=len(texts) if len(texts) > 1 else 0,
|
||||
# Monotonic under --rich, so a bad rollback reads as a
|
||||
# wrong number rather than as nothing. Left to
|
||||
# `tree.stamp_outcome` otherwise, which writes the
|
||||
# adventure's (unchanging) state — true, and no witness.
|
||||
# Under `--rich` the value increases with the turn, so an
|
||||
# incorrect rollback shows a wrong number rather than no
|
||||
# value. Otherwise `tree.stamp_outcome` writes the
|
||||
# adventure's state, which does not change and therefore
|
||||
# tests nothing.
|
||||
state_after=rich_script_state(i) if args.rich else None,
|
||||
world_state_after=(
|
||||
rich_world_state(schema, i) if args.rich else None
|
||||
),
|
||||
)
|
||||
# A fresh database is built by create_all and stamped LATEST,
|
||||
# so no migration ever runs against it and the tree backfill
|
||||
# never sees it. The fixture has to stamp its own nodes, or it
|
||||
# would be the one database in the project whose actions have
|
||||
# no branch.
|
||||
# `create_all` builds a fresh database and stamps it LATEST,
|
||||
# so no migration runs against it and the tree backfill never
|
||||
# sees it. The fixture has to stamp its own nodes, or it would
|
||||
# be the one database in the project whose actions have no
|
||||
# branch.
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
|
||||
@@ -529,14 +547,15 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
text=f"{MEMORY_TEXT} ({i})",
|
||||
source_start=i * memorybank.MEMORY_INTERVAL,
|
||||
source_end=i * memorybank.MEMORY_INTERVAL + memorybank.MEMORY_INTERVAL - 1,
|
||||
# A bank in play is not uniformly active: some memories are
|
||||
# pinned (always retrieved) and some evicted but kept for the
|
||||
# UI. Both states have to survive being re-attached to nodes.
|
||||
# A bank in play is not uniformly active. Some memories are
|
||||
# pinned, which means they are always retrieved, and some are
|
||||
# evicted but kept for the UI. Both states have to survive being
|
||||
# re-attached to nodes.
|
||||
pinned=bool(args.rich and i % 9 == 0),
|
||||
forgotten=bool(args.rich and i % 11 == 5),
|
||||
)
|
||||
# Through the same door the app uses, so the fixture cannot end up
|
||||
# storing vectors in a shape production never produces.
|
||||
# Use the same call the app uses, so the fixture cannot store
|
||||
# vectors in a shape production never produces.
|
||||
memorybank.set_vector(
|
||||
memory, [rng.uniform(-1.0, 1.0) for _ in range(EMBEDDING_DIMS)]
|
||||
)
|
||||
@@ -544,11 +563,11 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
|
||||
db.add(memory)
|
||||
|
||||
if args.rich:
|
||||
# The memory/summary marks as nodes, translated from the positions
|
||||
# `add_rich_layers` set now that there are actions to point at —
|
||||
# through the same call the v1 importer uses. A fixture stamped
|
||||
# LATEST never meets a migration, so if this is skipped it is the
|
||||
# one database whose marks are only positions.
|
||||
# Translate the memory mark and the summary mark from the
|
||||
# positions `add_rich_layers` set into nodes, now that actions exist
|
||||
# to point at. This uses the same call the v1 importer uses. A
|
||||
# fixture stamped LATEST never runs a migration, so skipping this
|
||||
# would leave the one database whose marks are only positions.
|
||||
db.flush()
|
||||
cursors.anchor_at_position(adventure, cursors.MEMORY, adventure.memory_cursor)
|
||||
cursors.anchor_at_position(adventure, cursors.SUMMARY, adventure.summary_cursor)
|
||||
@@ -571,8 +590,8 @@ def install_fakes(user_id: int, rng: random.Random) -> None:
|
||||
)
|
||||
limits.rate_limit = lambda *a, **k: None
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
# Fire-and-forget post-turn work would land inside whichever scope happened
|
||||
# to be open. It is measured on purpose, as its own shape.
|
||||
# Post-turn work scheduled in the background would be counted inside
|
||||
# whatever scope was open. It is measured deliberately, as its own shape.
|
||||
memorybank.schedule_post_turn = lambda adventure: None
|
||||
|
||||
def _current_user(db=Depends(get_db)):
|
||||
@@ -585,21 +604,23 @@ def install_fakes(user_id: int, rng: random.Random) -> None:
|
||||
|
||||
|
||||
def shape_list(client, meter, adv_id):
|
||||
"""The adventures index — every adventure's latest narration."""
|
||||
"""Measures the adventures index, which returns each adventure's latest
|
||||
narration."""
|
||||
with meter.scope("GET /adventures (index)"):
|
||||
r = client.get("/api/adventures")
|
||||
_check(r)
|
||||
|
||||
|
||||
def shape_load(client, meter, adv_id):
|
||||
"""Opening a finished adventure: the whole story, in one response."""
|
||||
"""Measures opening a finished adventure, which returns the whole story in
|
||||
one response."""
|
||||
with meter.scope(f"GET /adventures/{{id}} (page load)"):
|
||||
r = client.get(f"/api/adventures/{adv_id}")
|
||||
_check(r)
|
||||
|
||||
|
||||
def shape_turn(client, meter, adv_id):
|
||||
"""One played turn, memory retrieval included."""
|
||||
"""Measures one played turn, including memory retrieval."""
|
||||
with meter.scope("POST /adventures/{id}/actions (one turn)"):
|
||||
r = client.post(
|
||||
f"/api/adventures/{adv_id}/actions",
|
||||
@@ -609,21 +630,24 @@ def shape_turn(client, meter, adv_id):
|
||||
|
||||
|
||||
def shape_insights(client, meter, adv_id):
|
||||
"""The Insights dry run — assembles a context without spending a turn."""
|
||||
"""Measures the Insights dry run, which assembles a context without playing
|
||||
a turn."""
|
||||
with meter.scope("GET /adventures/{id}/context (insights)"):
|
||||
r = client.get(f"/api/adventures/{adv_id}/context")
|
||||
_check(r)
|
||||
|
||||
|
||||
def shape_memories(client, meter, adv_id):
|
||||
"""The Memories drawer — every memory, and none of their vectors."""
|
||||
"""Measures the Memories drawer, which returns every memory and no
|
||||
vectors."""
|
||||
with meter.scope("GET /adventures/{id}/memories (drawer)"):
|
||||
r = client.get(f"/api/adventures/{adv_id}/memories")
|
||||
_check(r)
|
||||
|
||||
|
||||
def shape_post_turn(client, meter, adv_id):
|
||||
"""Summarization, embedding and eviction, after the turn is saved."""
|
||||
"""Measures summarization, embedding, and eviction after the turn is
|
||||
saved."""
|
||||
with meter.scope("run_post_turn (background)"):
|
||||
asyncio.run(memorybank.run_post_turn(adv_id))
|
||||
|
||||
@@ -647,21 +671,24 @@ def _check(response) -> None:
|
||||
|
||||
|
||||
def make_bootable() -> None:
|
||||
"""Two edits that turn a measurement fixture into a database the app will
|
||||
actually serve. Both exist because build_fixture() builds a database for
|
||||
the meter, not for a browser."""
|
||||
"""Applies the two edits that make a measurement fixture servable by the app.
|
||||
|
||||
Both edits are needed because `build_fixture()` builds a database for the
|
||||
meter rather than for a browser.
|
||||
"""
|
||||
from app.migrations import LATEST_VERSION
|
||||
|
||||
with engine.begin() as conn:
|
||||
# create_all() builds the current schema but leaves the stamp at its
|
||||
# default, and bootstrap() reads a stamped-but-not-fresh database as
|
||||
# ancient — it would replay all of the migrations against a schema
|
||||
# that already has every column, and fail on the first one.
|
||||
# `create_all()` builds the current schema but leaves the stamp at its
|
||||
# default, and `bootstrap()` reads an unstamped existing database as
|
||||
# very old. It would replay every migration against a schema that
|
||||
# already has every column, and fail on the first one.
|
||||
conn.execute(text(f"PRAGMA user_version = {LATEST_VERSION}"))
|
||||
# In local mode (AIDND_MULTI_USER unset) get_current_user() looks for
|
||||
# the row with email IS NULL and is_guest false. The fixture's user is
|
||||
# a registered one, so without this nothing owns the adventure and the
|
||||
# app opens on an empty library.
|
||||
# In local mode, which is when `AIDND_MULTI_USER` is unset,
|
||||
# `get_current_user()` looks for the row with `email IS NULL` and
|
||||
# `is_guest` false. The fixture's user is a registered one, so without
|
||||
# this update nothing owns the adventure and the app opens on an empty
|
||||
# library.
|
||||
conn.execute(text("UPDATE users SET email = NULL, is_guest = 0"))
|
||||
|
||||
|
||||
@@ -675,9 +702,9 @@ def print_keep_notes(path: str, actions: int) -> None:
|
||||
print(f" .venv/Scripts/python.exe -m uvicorn app.main:app --port {port}")
|
||||
print(f" cd frontend && AIDND_API_PORT={port} npm run dev")
|
||||
print()
|
||||
# 8000 is the vite proxy's default and another local app squats it, which
|
||||
# shadows this API with its own SPA catch-all and looks like an empty
|
||||
# database rather than a proxy problem.
|
||||
# 8000 is the vite proxy's default, and another local app already listens
|
||||
# there. That app shadows this API with its own SPA catch-all route, which
|
||||
# looks like an empty database rather than a proxy problem.
|
||||
print(f" Port {port} rather than 8000 on purpose; AIDND_API_PORT points vite at it.")
|
||||
|
||||
|
||||
@@ -688,8 +715,8 @@ def parse_args(argv=None):
|
||||
p = argparse.ArgumentParser(
|
||||
prog="tools.stress_session", description=__doc__.splitlines()[0]
|
||||
)
|
||||
# 607 is production's longest adventure as of 2026-08-17, and length is
|
||||
# the dimension the old default (200) got wrong: real actions are lighter
|
||||
# 607 is the longest adventure in production as of 2026-08-17. Length is
|
||||
# the dimension the old default of 200 got wrong. Real actions are smaller
|
||||
# than this fixture used to make them, but real stories run three times
|
||||
# longer, and length is what a page load pays for.
|
||||
p.add_argument("--actions", type=int, default=600,
|
||||
@@ -697,14 +724,14 @@ def parse_args(argv=None):
|
||||
"production's longest adventure is 607)")
|
||||
p.add_argument("--memories", type=int, default=100,
|
||||
help="memories, all embedded (default: 100)")
|
||||
# Deliberately not the app's default (80): a measuring instrument should
|
||||
# hold the fixture at the size asked for rather than evict it mid-run.
|
||||
# This is not the app's default of 80. A measuring instrument holds the
|
||||
# fixture at the requested size rather than evict from it during a run.
|
||||
p.add_argument("--capacity", type=int, default=200,
|
||||
help="Settings.memory_bank_capacity; lower it below "
|
||||
"--memories to exercise eviction (default: 200)")
|
||||
# Production's longest adventure carries 886 B of text per action averaged
|
||||
# over both kinds. AI actions alternate with a one-line player input, so
|
||||
# the AI half has to be about twice that.
|
||||
# The longest adventure in production carries 886 B of text per action,
|
||||
# averaged over both kinds. AI actions alternate with a one-line player
|
||||
# input, so the AI half has to be about twice that.
|
||||
p.add_argument("--narration-bytes", type=int, default=1700,
|
||||
help="length of an AI action's text; alternating with a "
|
||||
"one-line player input this averages ~890 B/action, "
|
||||
@@ -730,9 +757,10 @@ def parse_args(argv=None):
|
||||
"one — prefer it small (--rich --actions 30). Changes "
|
||||
"what the shapes cost, so do not compare a --rich run "
|
||||
"against a plain one")
|
||||
# Read at import time by _early_keep as well — the database location has
|
||||
# to be settled before app.database loads. Declared here so it appears in
|
||||
# --help and an unknown spelling is still rejected.
|
||||
# `_early_keep` also reads this flag at import time, because the database
|
||||
# location has to be decided before `app.database` loads. It is declared
|
||||
# here so that it appears in `--help` and an unknown flag is still
|
||||
# rejected.
|
||||
p.add_argument("--keep", metavar="PATH", default="",
|
||||
help="write the fixture to PATH and leave it bootable, so "
|
||||
"the app can serve it in a browser (default: a temp "
|
||||
@@ -756,8 +784,8 @@ def main(argv=None) -> int:
|
||||
install_fakes(user_id, rng)
|
||||
|
||||
meter = Meter()
|
||||
# After the fixture: building it is a write path nobody plays, and its
|
||||
# bytes would drown everything the shapes report.
|
||||
# Attach after the fixture is built. Building it is a write path no player
|
||||
# takes, and its bytes would hide everything the shapes report.
|
||||
meter.attach(engine)
|
||||
|
||||
print(f"fixture: {args.actions} actions × {args.narration_bytes} B "
|
||||
|
||||
Reference in New Issue
Block a user