@@ -1,4 +1,4 @@
""" Drive a production-sized adventure through the real routes and report what
""" Drives a production-sized adventure through the real routes and reports what
each one costs in database bytes.
cd backend
@@ -6,43 +6,42 @@ each one costs in database bytes.
.venv/Scripts/python.exe -m tools.stress_session --actions 200 --memories 100
.venv/Scripts/python.exe -m tools.stress_session --no-embeddings
** The memory bank is ON by default, and that is the point.** The round-two
stress harness ran without an embedding model configured, and embedding
providers are BYOK-only by construction, so `retrieve_memories` returned early
every time — the whole exercise measured the turn loop with its heaviest read
switched off, and reported 23 MB for a playthrough that actually costs an order
of magnitude more. `--no-embeddings` reproduces that blindness deliberately, to
show the gap; it is never the default and it prints a warning.
The memory bank is on by default, and that matters. The round-two stress
harness ran with no embedding model configured, and embedding providers are
BYOK-only by construction, so `retrieve_memories` returned early every time.
That run measured the turn loop with its heaviest read disabled and reported
23 MB for a playthrough that costs an order of magnitude more.
`--no-embeddings` reproduces that configuration on purpose, to show the gap. It
is never the default, and it prints a warning.
Everything here is synthetic. The fixture is generated to production * shape* —
1536-dimension embeddings, ~74 KB context snapshots, retry variants — and no
real adventure, user or backup is ever read .
Everything here is synthetic. The fixture is generated to the shape of
production data, with 1536-dimension embeddings, context snapshots of about
74 kB, and retry variants. It never reads a real adventure, user, or backup.
Only the network is faked: the LLM and the embedding endpoint. Routing,
sessions, the ORM, the scripting engine and the context builder are the real
ones, because the bugs this exists to catch liv e in exactly the layer a mock
would replace.
Only the network is faked, which means the LLM and the embedding endpoint.
Routing, sessions, the ORM, the scripting engine, and the context builder are
the real ones, because the bugs this harness exists to catch ar e in the layer a
mock would replace.
It runs on a throwaway SQLite file by default. What is being measured is which
columns of which rows a code path asks for , and that is decided by the ORM,
identically on both dialects. The dialects disagree o n how a value is encoded
on the wire — JSON especially — so treat the absolute figures as
production-shaped rather than production-exact, and compare before against
after.
It runs on a throwaway SQLite file by default. The measurement is which columns
of which rows a code path requests , and the ORM decides that identically on both
dialects. The dialects differ i n how a value is encoded on the wire, especially
JSON, so treat the absolute figures as production-shaped rather than
production-exact, and compare one run against another.
To measure the encodings SQLite cannot reach — bytea for the packed vectors,
and json columns psycopg parses before the meter sees them — set
AIDND_STRESS_DATABASE_URL to a ** throwaway** Postgres database:
To measure the encodings SQLite cannot reach, which are bytea for the packed
vectors and json columns that psycopg parses before the meter sees them, set
` AIDND_STRESS_DATABASE_URL` to a throwaway Postgres database:
AIDND_STRESS_DATABASE_URL=postgresql://…/stress_scratch \
.venv/Scripts/python.exe -m tools.stress_session
The harness writes, so it refuses any target whose database name does not say
' stress ' or ' scratch ' . Never point it at a database holding real users.
The harness writes, so it refuses any target whose database name does not
contain ' stress ' or ' scratch ' . Never point it at a database holding real users.
Calibration. The fixture is sized from production, re-measured 2026-08-17
against the live Neon database ( aggregates only — counts and octet_length
sums, never row contents):
Calibration. The fixture is sized from production and was re-measured on
2026-08-17 against the live Neon database, reading aggregates only: counts and
`octet_length` sums, never row contents.
per action, text 886 B -> --narration-bytes 1700, alternating
with a one-line player input
@@ -51,14 +50,13 @@ sums, never row contents):
memory bank, largest 100 memories, 6,144 B a vector
The previous defaults were wrong in both directions at once and happened to
land near the right total: a ctions were modell ed at ~ 2.1 K B against a real
886 B, and stories at 200 actions against a real 607. W idth was flattering,
length was not , and length is what a page load pays for.
land near the right total. A ctions were modeled at about 2.1 k B against a real
886 B, and stories at 200 actions against a real 607. The w idth was too large
and the length was too small , and the length is what a page load pays for.
Filler text is generated word by word rather than repeated. A repeated
sentence compresses about a hundredfold and prose three- or fourfold, so the
old fixture would have made any compression measurement on context_snapshot
meaningless.
Filler text is generated word by word rather than repeated. A repeated sentence
compresses about a hundredfold and prose three- or fourfold, so the old fixture
would have made any compression measurement on ` context_snapshot` meaningless.
"""
import os
@@ -68,11 +66,12 @@ from pathlib import Path
def _early_keep ( argv : list [ str ] ) - > str :
""" --keep, read before argparse exists.
""" Reads `--keep` before argparse exists.
Where t he database lives has to be decided before app.database is
imported, and that import is three lines below. argparse still declares
the flag, so --help documents it and a typo is still an error. """
T he database location has to be decided before ` app.database` is imported,
and that import is a few lines below. argparse still declares the flag, so
` --help` documents it and a typo is still an error.
"""
for i , arg in enumerate ( argv ) :
if arg == " --keep " and i + 1 < len ( argv ) :
return argv [ i + 1 ]
@@ -83,17 +82,17 @@ def _early_keep(argv: list[str]) -> str:
_keep = _early_keep ( sys . argv [ 1 : ] )
# Must preced e the app import: database.py reads these at module scope.
# This has to run befor e the app import, because ` database.py` reads these at
# module scope.
#
# D efault is a throwaway SQLite file. AIDND_STRESS_DATABASE_URL points the
# The d efault is a throwaway SQLite file. ` AIDND_STRESS_DATABASE_URL` points the
# harness at a real Postgres instead, which is the only way to reach the
# encodings SQLite cannot exercise: bytea for the packed vectors, and json
# columns that psycopg parses into Python before the meter ever sees them.
# columns that psycopg parses into Python before the meter sees them.
#
# The name guard is not paranoia . This harness * writes* — it builds a whole
# synthetic adventure — so a URL that happened to point at the production
# database would quietly seed it with fake users and fake play. The target
# must say it is disposable.
# The name guard matters . This harness writes a whole synthetic adventure, so a
# URL that pointed at the production database would seed it with fake users and
# fake play. The target has to name itself as disposable.
_stress_url = os . environ . get ( " AIDND_STRESS_DATABASE_URL " , " " ) . strip ( )
if _stress_url :
if _keep :
@@ -111,10 +110,10 @@ if _stress_url:
os . environ [ " AIDND_DATABASE_URL " ] = _stress_url
os . environ . pop ( " DATABASE_URL " , None )
elif _keep :
# A fixture to boo t the app against rather than a temp file the report
# discards. Rebuilt from empty e very run: build_fixture() assumes an empty
# database on the SQLite path, and a second run would otherwise stack a
# second adventure beside the first.
# A fixture to star t the app against, rather than a temporary file the
# report discards. E very run re builds it from empty, because
# `build_fixture()` assumes an empty database on the SQLite path and a
# second run would otherwise add a second adventure next to the first.
_keep_path = Path ( _keep ) . resolve ( )
_keep_path . parent . mkdir ( parents = True , exist_ok = True )
_keep_path . unlink ( missing_ok = True )
@@ -149,21 +148,22 @@ from .fakeprose import prose
EMBEDDING_DIMS = 1536
# ~ 232 K B, measured on production's largest adventure (2026-08-17). The old
# figure here was 74 K B, taken from the comment in models.py; t he real column
# averages 163 K B a row across the whole table and 232 K B on the adventure that
# matters, because the assembled prompt grows with the story behind it.
# About 232 k B, measured on the largest adventure in production on 2026-08-17.
# The old figure here was 74 k B, taken from the comment in ` models.py`. T he real
# column averages 163 k B per row across the whole table and 232 k B on the
# adventure that matters, because the assembled prompt grows with the story
# behind it.
#
# Built from varied text rather than one sentence repeated. A repeated sentence
# compresses about a hundredfold and real prose three- or fourfold, so a
# fixture made of repeats would make any compression measurement meaningless —
# and shrinking this column is the open question i t exists to answer.
# The text is varied rather than one sentence repeated. A repeated sentence
# compresses about a hundredfold and real prose three- or fourfold, so a fixture
# built from repeats would make any compression measurement meaningless, and
# shrinking this column is the open question the fixture exists to answer.
SNAPSHOT_SYSTEM = None # set by _build_text()
SNAPSHOT_STORY = None
PLAYER_INPUT = " > You crouch and look more closely at the grit on the floor. "
# All three are bound by _build_text() from the fixture arguments.
# `_build_text()` sets all three from the fixture arguments.
NARRATION = None
MEMORY_TEXT = (
" You found a bandit camp above the ford and agreed to guide Gwen through "
@@ -172,15 +172,15 @@ MEMORY_TEXT = (
def _build_text ( args , rng : random . Random ) - > None :
""" Size the three variable-length fixture strings from the arguments.
""" Sizes the three variable-length fixture strings from the arguments.
S eparate from build_fixture so the sizes are decided once, before anything
is written, and so a shape ' s cost is a function of the flags rather than of
how many rows happened to b e generated first.
This is s eparate from ` build_fixture` so that the sizes are decided once,
before anything is written, and so that a shape ' s cost depends on the flags
rather than on how many rows wer e generated first.
"""
global NARRATION , SNAPSHOT_SYSTEM , SNAPSHOT_STORY
NARRATION = prose ( rng , args . narration_bytes )
# The assembled prompt is a system block and the story so far; t he split
# The assembled prompt is a system block plus the story so far. T he split
# is roughly one to five in production.
SNAPSHOT_SYSTEM = prose ( rng , args . snapshot_bytes / / 6 )
SNAPSHOT_STORY = prose ( rng , args . snapshot_bytes - args . snapshot_bytes / / 6 )
@@ -190,7 +190,11 @@ def _build_text(args, rng: random.Random) -> None:
class FakeProvider :
""" T he LLM. S treams one fixed line; no network, no cost, no variance. """
""" Stands in for t he LLM, s treaming one fixed line.
It makes no network call, costs nothing, and returns the same text every
time.
"""
def __init__ ( self , * a , * * k ) :
pass
@@ -203,8 +207,11 @@ class FakeProvider:
class FakeEmbeddings :
""" T he embedding endpoint. Returns vectors of the real width, so what the
turn writes back weighs what production weighs. """
""" Stands in for t he embedding endpoint.
It returns vectors of the production width, so what the turn writes back is
the size production writes back.
"""
def __init__ ( self , rng : random . Random ) :
self . rng = rng
@@ -218,32 +225,34 @@ class FakeEmbeddings:
# -------------------------------------------------------------- rich fixture
# The correctness fixture, as against the scale on e.
# The correctness fixture, as opposed to the scale fixtur e.
#
# The default fixture is sized from production and exists to weigh bytes, so it
# leaves every column it does not weigh at its default. That makes it a poor
# witness for anything * semantic* — and phase 14 replaces the storage model
# underneath all of it. Measured o n a freshly built default fixture, the
# columns a story tree has to migrate correctly look like this:
# The default fixture is sized from production and exists to measure bytes, so it
# leaves every column it does not measure at its default. That makes it a poor
# test of semantics, and phase 14 replaces the storage model underneath all of
# it. O n a freshly built default fixture, the columns a story tree has to migrate
# correctly hold this:
#
# state_after / world_state_after t he same value on all 600 rows (t hey
# state_after, world_state_after T he same value on all 600 rows. T hey
# were state_before, and NULL, until SP4
# turn ed them round) . A rollback over
# identical snapshots prove s nothing.
# scenario_id / world_state a bsent. N o RPG layer, so the cooldown
# clock SP5 must not advance never runs.
# adventure_scripts none. script_state rollback is exactly
# what a branch switch reuses.
# memory_cursor / summary_cursor both 0. SP3 replaces the cursors .
# sibling attempts two of byte-identical text with the
# first always live — so "which attempt
# i s live?", the one question SP4 has to
# answer, has no observable answer.
# revers ed them. A rollback over identical
# snapshots test s nothing.
# scenario_id, world_state A bsent. There is n o RPG layer, so the
# cooldown clock SP5 must not advance
# never runs.
# adventure_scripts None. A branch switch reuses the
# script_state rollback .
# memory_cursor, summary_cursor Both 0. SP3 replaces the cursors.
# sibling attempts Two attempts of byte-identical text with
# the first alway s live, so the one
# question SP4 has to answer, which
# attempt is live, has no observable
# answer.
#
# --rich fills in exactly those and changes nothing else, so the measuring
# fixture's numbers stay comparable run to run. It is a * correctness* fixture:
# prefer it small ( --rich --actions 30), because what it is for i s variety per
# row, not rows.
# ` --rich` fills in those columns and changes nothing else, so the measuring
# fixture's numbers stay comparable from run to run. It is a correctness fixture,
# so keep it small, such as ` --rich --actions 30`. It provide s variety per row
# rather than many rows.
RICH_SUMMARY = (
" You tracked the bandits to a camp above the ford, freed Gwen from the "
@@ -260,8 +269,8 @@ RICH_CARDS = [
" Taken from the quartermaster. Opens something below the camp. " ) ,
]
# Ten gold a turn, so a state snapshot that failed to roll back reads as a
# wrong total rather than as nothing. Same shape t he retry tests use.
# Ten gold a turn, so a state snapshot that failed to roll back shows a wrong
# total rather than no value. T he retry tests use the same shape .
RICH_SCRIPT = """
const modifier = (text) => {
state.gold = (state.gold || 0) + 10;
@@ -272,28 +281,31 @@ modifier(text);
def rich_stat_schema ( ) - > dict :
""" T he demo RPG schema, read from the seed data rather than invented here .
""" Returns t he demo RPG schema, read from the seed data rather than invented.
Using the real one means the fixture exercises bands, cooldowns,
max_delta_per_turn and NPC stat blocks as they are actually shaped — an
invented schema would drift from the thing it is standing in for.
Using the real schema means the fixture exercises bands, cooldowns,
` max_delta_per_turn`, and NPC stat blocks in their real shapes. An invented
schema would diverge from the one it stands in for.
"""
path = seed . SEED_DIR / " 04-rpg-world-state.json "
return json . loads ( path . read_text ( encoding = " utf-8 " ) ) [ " stat_schema " ]
def rich_script_state ( turn : int ) - > dict :
""" The scoreboard as of `turn`. Monotonic, so any snapshot identifies the
turn it was taken at — which is what makes a bad rollback visible. """
""" Returns the script state as of `turn`.
The values increase with the turn, so any snapshot identifies the turn it was
taken at, which is what makes an incorrect rollback visible.
"""
return { " gold " : turn * 10 , " turn " : turn }
def rich_world_state ( schema : dict , turn : int ) - > dict :
""" A live world state that has actually been played to `turn`.
""" Returns a live world state that has been played to `turn`.
`instantiate` give s the initial picture; a fixture whose every row holds
that same pictur e cannot tell a restored snapshot from an unrestored one.
S o hp declines, mana drains and a flag flip s partway through.
`instantiate` return s the initial state, and a fixture whose every row holds
that same stat e cannot distinguish a restored snapshot from an unrestored
one. Here hp declines, mana declines, and a flag change s partway through.
"""
ws = worldstate . instantiate ( schema )
ws [ " player " ] [ " hp " ] = max ( 20 , 100 - turn )
@@ -306,28 +318,28 @@ def rich_world_state(schema: dict, turn: int) -> dict:
def rich_attempts ( rng : random . Random , index : int ) - > tuple [ list [ str ] , int ] :
""" D istinguishable retry attempts, and which one the story tells.
""" Returns d istinguishable retry attempts, and which one the story tells.
Every attempt in the default fixture carries the same text with the live
one pinned at 0. That is the one thing SP4 has to get right and the one
thing that fixture cannot witn ess , so here the texts differ, the counts
differ, and the live one is often not the last.
Every attempt in the default fixture carries the same text, with the live one
fixed at index 0. That is the behavior SP4 has to get right and the behavior
that fixture cannot t est , so here the texts differ, the counts differ, and
the live attempt is often not the last.
"""
count = 3 if index % 12 == 1 else 2
texts = [
f " [attempt { n + 1 } of { count } at turn { index } ] { prose ( rng , 240 ) } "
for n in range ( count )
]
# Deliberately not always the newest: a player who retried twice and then
# went back to the first take is the case that breaks anything assuming the
# live attempt is the last one written.
# The live attempt is not always the newest. A player who retried twice and
# then went back to the first attempt is the case that breaks code assuming
# the live attempt is the last one written.
return texts , ( 0 if index % 18 == 1 else count - 1 )
def add_rich_extras ( db , args , rng : random . Random , user , adventure ) - > None :
""" E verything --rich adds beside the actions themselves.
""" Adds e verything ` --rich` contributes apart from the actions themselves.
Called with the adventure already flushed, before the actions are written ,
The caller flushes the adventure first and writes the actions afterwards ,
because the actions need the schema to snapshot a world state from.
"""
schema = rich_stat_schema ( )
@@ -343,11 +355,11 @@ def add_rich_extras(db, args, rng: random.Random, user, adventure) -> None:
adventure . scenario_id = scenario . id
adventure . world_state = rich_world_state ( schema , args . actions )
adventure . script_state = rich_script_state ( args . actions )
# The post-turn passes only do anything when summarization is on and the
# marks are somewhere other than the start. W ritten as positions here and
# translated into anchors once the actions exist (s ee build_fixture) — a
# position is what a database being migrated to SP3 still holds, so the
# fixture carries both and they have to say the same thing .
# The post-turn passes do nothing unless summarization is on and the marks
# are past the start. They are w ritten here as positions and translated into
# anchors once the actions exist. S ee ` build_fixture`. A database being
# migrated to SP3 still holds positions, so the fixture carries both forms
# and the two have to agree .
adventure . auto_summarize = True
adventure . story_summary = RICH_SUMMARY
adventure . memory_cursor = max ( 0 , args . actions - 8 )
@@ -366,10 +378,12 @@ def add_rich_extras(db, args, rng: random.Random, user, adventure) -> None:
def add_second_adventure ( db , rng : random . Random , user ) - > int :
""" A short second adventure, so ' does this leak a cross adventures? ' is a
question the fixture can answer. A tree scopes every read by branch, and a
branch clause that forgot its adventure would still look right on a
database holding exactly one. """
""" Adds a short second adventure, so the fixture can detect cross- adventure
leaks.
A tree scopes every read by branch, and a branch clause that omitted its
adventure would still look correct on a database holding one adventure.
"""
other = models . Adventure (
user_id = user . id , title = " Stress (second) " , script_state = { } ,
memory_bank_enabled = True , auto_summarize = False ,
@@ -398,12 +412,12 @@ def add_second_adventure(db, rng: random.Random, user) -> int:
def _assert_live_variant_invariant ( db , adventure_id : int ) - > None :
""" E xactly one attempt per turn is live, on every turn.
""" Checks that e xactly one attempt per turn is live, on every turn.
Checked here rather than trust ed, because a coordinate with two live
siblings tell s its story twice and a coordinate with none drops a turn out
of it — and both fail by *reading* wrong, never by raising. A fixture that
quietly violated it would let wrong code look righ t.
The check runs here rather than being assum ed, because a coordinate with two
live siblings render s its turn twice and a coordinate with none omits the
turn. Both failures produce wrong output rather than an exception, so a
fixture that broke the rule would let incorrect code look correc t.
"""
groups : dict [ tuple , list [ models . Action ] ] = { }
for action in (
@@ -430,13 +444,15 @@ def _assert_live_variant_invariant(db, adventure_id: int) -> None:
def build_fixture ( args , rng : random . Random ) - > tuple [ int , int ] :
""" A user, settings and one adventure at production scale. Returns
(adventure_id, user_id). """
# A SQLite run gets a brand-new temp file every time, so the fixture can
# assume an empty database. A Postgres scratch target persists between
# runs, and the second one would collide on the fixture user's unique
# email — so empty it first. Only ever reached for a target whose name
# passed the 'stress'/'scratch' guard at the top of this module.
""" Builds a user, settings, and one adventure at production scale.
Returns `(adventure_id, user_id)`.
"""
# A SQLite run gets a new temporary file every time, so the fixture can
# assume an empty database. A Postgres scratch target persists between runs,
# and a second run would collide on the fixture user's unique email, so empty
# it first. This runs only for a target whose name passed the 'stress' or
# 'scratch' guard at the top of this module.
if _stress_url :
Base . metadata . drop_all ( bind = engine )
Base . metadata . create_all ( bind = engine )
@@ -450,7 +466,7 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
api_key = security . encrypt_secret ( " stress-key " ) ,
model = " stress-model " ,
endpoint_url = " https://fake.invalid/v1 " ,
# The default this whole tool exists to sto p anyone forgetting.
# The default this tool exists to kee p anyone from forgetting.
embedding_model = " " if args . no_embeddings else " openai/text-embedding-3-small " ,
memory_bank_capacity = args . capacity ,
) )
@@ -459,8 +475,9 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
title = " Stress " ,
script_state = { } ,
memory_bank_enabled = True ,
# Off so a turn measures the turn. The post-turn pass is its own
# shape below; letting it fire mid-measurement would mix the two.
# Off, so that a turn measures only the turn. The post-turn pass
# is its own shape below, and running it during the measurement
# would combine the two.
auto_summarize = False ,
)
db . add ( adventure )
@@ -472,7 +489,7 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
is_ai = bool ( i % 2 )
retried = is_ai and i % 6 == 1
if args . rich :
# Distinguishable attempts, and a live one that is often not
# Distinguishable attempts, with a live one that is often not
# the last written.
texts , live = rich_attempts ( rng , i ) if retried else ( [ ] , 0 )
if not retried :
@@ -490,10 +507,10 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
type = " ai " if is_ai else " do " ,
text = body ,
# The turn's assembled prompt is stored once, on the
# attempt the story tells; a superseded sibling keeps only
# its own slices. That is the invariant `app/attempts.py`
# maintains , and a fixture that ignored it would multiply
# the big gest column in the database by the retry count.
# attempt the story tells. A superseded sibling keeps only
# its own slices. `app/attempts.py` maintains that
# invariant , and a fixture that ignored it would multiply
# the lar gest column in the database by the retry count.
context_snapshot = (
{ " system " : SNAPSHOT_SYSTEM , " story " : SNAPSHOT_STORY }
if n == live else
@@ -506,20 +523,21 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
live = ( n == live ) ,
variant_index = n ,
variant_count = len ( texts ) if len ( texts ) > 1 else 0 ,
# Monotonic u nder --rich, so a bad rollback reads as a
# wrong number rather than as nothing. Left t o
# `tree.stamp_outcome` otherwise, which writes the
# adventure's (unchanging) state — true, and no witness.
# U nder ` --rich` the value increases with the turn, so an
# incorrect rollback shows a wrong number rather than n o
# value. Otherwise `tree.stamp_outcome` writes the
# adventure's state, which does not change and therefore
# tests nothing.
state_after = rich_script_state ( i ) if args . rich else None ,
world_state_after = (
rich_world_state ( schema , i ) if args . rich else None
) ,
)
# A fresh database is built by create_all and stamped LATEST,
# so no migration ever runs against it and the tree backfill
# never sees it. The fixture has to stamp its own nodes, or it
# would be the one database in the project whose actions have
# no branch.
# `create_all` builds a fresh database and stamps it LATEST,
# so no migration runs against it and the tree backfill never
# sees it. The fixture has to stamp its own nodes, or it would
# be the one database in the project whose actions have no
# branch.
tree . place_action ( db , adventure , action )
db . add ( action )
@@ -529,14 +547,15 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
text = f " { MEMORY_TEXT } ( { i } ) " ,
source_start = i * memorybank . MEMORY_INTERVAL ,
source_end = i * memorybank . MEMORY_INTERVAL + memorybank . MEMORY_INTERVAL - 1 ,
# A bank in play is not uniformly active: s ome memories are
# pinned ( always retrieved) and some evicted but kept for th e
# UI. Both states have to survive being re-attached to nodes.
# A bank in play is not uniformly active. S ome memories are
# pinned, which means they are always retrieved, and some ar e
# evicted but kept for the UI. Both states have to survive being
# re-attached to nodes.
pinned = bool ( args . rich and i % 9 == 0 ) ,
forgotten = bool ( args . rich and i % 11 == 5 ) ,
)
# Through the same door the app uses, so the fixture cannot end up
# storing vectors in a shape production never produces.
# Use the same call the app uses, so the fixture cannot store
# vectors in a shape production never produces.
memorybank . set_vector (
memory , [ rng . uniform ( - 1.0 , 1.0 ) for _ in range ( EMBEDDING_DIMS ) ]
)
@@ -544,11 +563,11 @@ def build_fixture(args, rng: random.Random) -> tuple[int, int]:
db . add ( memory )
if args . rich :
# The memory/summary marks as nodes, translated from the positions
# `add_rich_layers` set now that there are actions to point at —
# through the same call the v1 importer uses. A fixture stamped
# LATEST never meet s a migration, so if this is skipped it is the
# one database whose marks are only positions.
# Translate t he memory mark and the summary mark from the
# positions `add_rich_layers` set into nodes, now that actions exist
# to point at. This uses the same call the v1 importer uses. A
# fixture stamped LATEST never run s a migration, so skipping this
# would leave the one database whose marks are only positions.
db . flush ( )
cursors . anchor_at_position ( adventure , cursors . MEMORY , adventure . memory_cursor )
cursors . anchor_at_position ( adventure , cursors . SUMMARY , adventure . summary_cursor )
@@ -571,8 +590,8 @@ def install_fakes(user_id: int, rng: random.Random) -> None:
)
limits . rate_limit = lambda * a , * * k : None
limits . check_row_cap = lambda * a , * * k : None
# Fire-and-forget p ost-turn work wo uld land inside whichever scope happened
# to be open. It is measured on purpose , as its own shape.
# P ost-turn work sched ule d in the background would be counted inside
# whatever scope was open. It is measured deliberately , as its own shape.
memorybank . schedule_post_turn = lambda adventure : None
def _current_user ( db = Depends ( get_db ) ) :
@@ -585,21 +604,23 @@ def install_fakes(user_id: int, rng: random.Random) -> None:
def shape_list ( client , meter , adv_id ) :
""" T he adventures index — every adventure' s latest narration. """
""" Measures t he adventures index, which returns each adventure' s latest
narration. """
with meter . scope ( " GET /adventures (index) " ) :
r = client . get ( " /api/adventures " )
_check ( r )
def shape_load ( client , meter , adv_id ) :
""" O pening a finished adventure: the whole story, in one response. """
""" Measures o pening a finished adventure, which returns the whole story in
one response. """
with meter . scope ( f " GET /adventures/ {{ id }} (page load) " ) :
r = client . get ( f " /api/adventures/ { adv_id } " )
_check ( r )
def shape_turn ( client , meter , adv_id ) :
""" O ne played turn, memory retrieval included ."""
""" Measures o ne played turn, including memory retrieval."""
with meter . scope ( " POST /adventures/ {id} /actions (one turn) " ) :
r = client . post (
f " /api/adventures/ { adv_id } /actions " ,
@@ -609,21 +630,24 @@ def shape_turn(client, meter, adv_id):
def shape_insights ( client , meter , adv_id ) :
""" T he Insights dry run — assembles a context without spending a turn. """
""" Measures t he Insights dry run, which assembles a context without playing
a turn. """
with meter . scope ( " GET /adventures/ {id} /context (insights) " ) :
r = client . get ( f " /api/adventures/ { adv_id } /context " )
_check ( r )
def shape_memories ( client , meter , adv_id ) :
""" T he Memories drawer — every memory, and none of their vectors. """
""" Measures t he Memories drawer, which returns every memory and no
vectors. """
with meter . scope ( " GET /adventures/ {id} /memories (drawer) " ) :
r = client . get ( f " /api/adventures/ { adv_id } /memories " )
_check ( r )
def shape_post_turn ( client , meter , adv_id ) :
""" S ummarization, embedding and eviction, after the turn is saved. """
""" Measures s ummarization, embedding, and eviction after the turn is
saved. """
with meter . scope ( " run_post_turn (background) " ) :
asyncio . run ( memorybank . run_post_turn ( adv_id ) )
@@ -647,21 +671,24 @@ def _check(response) -> None:
def make_bootable ( ) - > None :
""" T wo edits that turn a measurement fixture into a database the app will
actually serve. Both exist because build_fixture() builds a database for
the meter, not for a browser. """
""" Applies the t wo edits that make a measurement fixture servable by the app.
Both edits are needed because `build_fixture()` builds a database for the
meter rather than for a browser.
"""
from app . migrations import LATEST_VERSION
with engine . begin ( ) as conn :
# create_all() builds the current schema but leaves the stamp at its
# default, and bootstrap() reads a stamped-but-not-fresh database as
# ancient — i t would replay all of the migrations against a schema
# that already has every column, and fail on the first one.
# ` create_all()` builds the current schema but leaves the stamp at its
# default, and ` bootstrap()` reads an un stamped existing database as
# very old. I t would replay every migration against a schema that
# already has every column, and fail on the first one.
conn . execute ( text ( f " PRAGMA user_version = { LATEST_VERSION } " ) )
# In local mode (AIDND_MULTI_USER unset) get_current_user() looks for
# the row with email IS NULL and is_guest false. The fixture's user is
# a registered one, so without this nothing owns the adventure and the
# app opens on an empty library.
# In local mode, which is when `AIDND_MULTI_USER` is unset,
# `get_current_user()` looks for the row with ` email IS NULL` and
# `is_guest` false. The fixture's user is a registered one, so without
# this update nothing owns the adventure and the app opens on an empty
# library.
conn . execute ( text ( " UPDATE users SET email = NULL, is_guest = 0 " ) )
@@ -675,9 +702,9 @@ def print_keep_notes(path: str, actions: int) -> None:
print ( f " .venv/Scripts/python.exe -m uvicorn app.main:app --port { port } " )
print ( f " cd frontend && AIDND_API_PORT= { port } npm run dev " )
print ( )
# 8000 is the vite proxy's default and another local app squats it, which
# shadows this API with its own SPA catch-all and looks like an empty
# database rather than a proxy problem.
# 8000 is the vite proxy's default, and another local app already listens
# there. That app shadows this API with its own SPA catch-all route, which
# looks like an empty database rather than a proxy problem.
print ( f " Port { port } rather than 8000 on purpose; AIDND_API_PORT points vite at it. " )
@@ -688,8 +715,8 @@ def parse_args(argv=None):
p = argparse . ArgumentParser (
prog = " tools.stress_session " , description = __doc__ . splitlines ( ) [ 0 ]
)
# 607 is production's longest adventure as of 2026-08-17, and l ength is
# the dimension the old default ( 200) got wrong: r eal actions are light er
# 607 is the longest adventure in production as of 2026-08-17. L ength is
# the dimension the old default of 200 got wrong. R eal actions are small er
# than this fixture used to make them, but real stories run three times
# longer, and length is what a page load pays for.
p . add_argument ( " --actions " , type = int , default = 600 ,
@@ -697,14 +724,14 @@ def parse_args(argv=None):
" production ' s longest adventure is 607) " )
p . add_argument ( " --memories " , type = int , default = 100 ,
help = " memories, all embedded (default: 100) " )
# Deliberately not the app's default (80): a measuring instrument s hou ld
# hold the fixture at the size asked for rather than evict it mid- run.
# This is not the app's default of 80. A measuring instrument holds the
# fixture at the requested size rather than evict from it during a run.
p . add_argument ( " --capacity " , type = int , default = 200 ,
help = " Settings.memory_bank_capacity; lower it below "
" --memories to exercise eviction (default: 200) " )
# Production's longest adventure carries 886 B of text per action averaged
# over both kinds. AI actions alternate with a one-line player input, so
# the AI half has to be about twice that.
# The longest adventure in production carries 886 B of text per action,
# averaged over both kinds. AI actions alternate with a one-line player
# input, so the AI half has to be about twice that.
p . add_argument ( " --narration-bytes " , type = int , default = 1700 ,
help = " length of an AI action ' s text; alternating with a "
" one-line player input this averages ~890 B/action, "
@@ -730,9 +757,10 @@ def parse_args(argv=None):
" one — prefer it small (--rich --actions 30). Changes "
" what the shapes cost, so do not compare a --rich run "
" against a plain one " )
# Read at import time by _early_keep as well — the database location has
# to be settl ed before app.database loads. Declared here so it appears in
# --help and an unknown spellin g is still rejected.
# `_early_keep` also reads this flag at import time, because the database
# location has to be decid ed before ` app.database` loads. It is declared
# here so that it appears in ` --help` and an unknown fla g is still
# rejected.
p . add_argument ( " --keep " , metavar = " PATH " , default = " " ,
help = " write the fixture to PATH and leave it bootable, so "
" the app can serve it in a browser (default: a temp "
@@ -756,8 +784,8 @@ def main(argv=None) -> int:
install_fakes ( user_id , rng )
meter = Meter ( )
# After the fixture: building it is a write path nobody plays, and its
# bytes would drown everything the shapes report.
# Attach a fter the fixture is built. Buil ding it is a write path no player
# takes, and its bytes would hide everything the shapes report.
meter . attach ( engine )
print ( f " fixture: { args . actions } actions × { args . narration_bytes } B "