Check the egress work against the database it actually runs on

Everything measured so far ran on SQLite against a synthetic fixture, so two
claims were still on trust: that migration 38 spells BYTEA correctly for a real
server, and that the byte figures survive psycopg's encodings.

Both hold. Production reads schema_version 41 with embedding_blob bytea and
embedded boolean present, the backfill is complete at 134/134, and the packed
vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a
memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and
against a throwaway Neon database every shape lands within 0.5% of the SQLite
run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of
it.

The harness writes, so it refuses any target whose name does not say stress or
scratch -- pointed at the production database it stops rather than seeding it
with a fake user and 200 fake turns. It also empties a Postgres target before
building, which a fresh SQLite temp file never needed.

Two corrections fall out, both recorded in plan/13. The page-load model has the
wrong shape: real actions are half the fixture's weight but real stories run to
607 actions, not 200, so the worst real page load is 589.5 kB. And the decision
to leave context_snapshot in the database costed egress but never storage --
it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the
ceiling this deploy will hit first.

Measured with counts and octet_length sums only. No user content was read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
This commit is contained in:
parththakkar106
2026-08-17 12:37:23 +05:30
co-authored by Claude Opus 5
parent 970b13998a
commit 70d024da62
3 changed files with 200 additions and 17 deletions
+51 -11
View File
@@ -23,11 +23,22 @@ sessions, the ORM, the scripting engine and the context builder are the real
ones, because the bugs this exists to catch live in exactly the layer a mock
would replace.
It runs on a throwaway SQLite file rather than Postgres. What is being measured
is which columns of which rows a code path asks for, and that is decided by the
ORM, identically on both. The dialects disagree on how a value is encoded on
the wire — JSON especially — so treat the absolute figures as production-shaped
rather than production-exact, and compare before against after.
It runs on a throwaway SQLite file by default. What is being measured is which
columns of which rows a code path asks for, and that is decided by the ORM,
identically on both dialects. The dialects disagree on how a value is encoded
on the wire — JSON especially — so treat the absolute figures as
production-shaped rather than production-exact, and compare before against
after.
To measure the encodings SQLite cannot reach — bytea for the packed vectors,
and json columns psycopg parses before the meter sees them — set
AIDND_STRESS_DATABASE_URL to a **throwaway** Postgres database:
AIDND_STRESS_DATABASE_URL=postgresql://…/stress_scratch \
.venv/Scripts/python.exe -m tools.stress_session
The harness writes, so it refuses any target whose database name does not say
'stress' or 'scratch'. Never point it at a database holding real users.
Calibration, against the two figures measured directly on production
(2026-08-16): a 200-action page load reported 426.7 kB here against 423 KB
@@ -35,19 +46,41 @@ there, and one turn on a 100-memory bank reported 3,258.7 kB against 3,153 kB.
"""
import os
import sys
import tempfile
# Must precede the app import: database.py reads these at module scope.
_tmp = tempfile.NamedTemporaryFile(suffix=".db", delete=False)
_tmp.close()
os.environ["AIDND_DB_PATH"] = _tmp.name
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
#
# Default is a throwaway SQLite file. AIDND_STRESS_DATABASE_URL points the
# harness at a real Postgres instead, which is the only way to reach the
# encodings SQLite cannot exercise: bytea for the packed vectors, and json
# columns that psycopg parses into Python before the meter ever sees them.
#
# The name guard is not paranoia. This harness *writes* — it builds a whole
# synthetic adventure — so a URL that happened to point at the production
# database would quietly seed it with fake users and fake play. The target
# must say it is disposable.
_stress_url = os.environ.get("AIDND_STRESS_DATABASE_URL", "").strip()
if _stress_url:
_dbname = _stress_url.rsplit("/", 1)[-1].split("?")[0]
if not any(mark in _dbname.lower() for mark in ("stress", "scratch")):
sys.exit(
f"refusing to run against database {_dbname!r}.\n"
"This harness writes a synthetic adventure, so its target must be a\n"
"throwaway database with 'stress' or 'scratch' in the name."
)
os.environ["AIDND_DATABASE_URL"] = _stress_url
os.environ.pop("DATABASE_URL", None)
else:
_tmp = tempfile.NamedTemporaryFile(suffix=".db", delete=False)
_tmp.close()
os.environ["AIDND_DB_PATH"] = _tmp.name
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
import argparse
import asyncio
import random
import sys
from fastapi import Depends
from fastapi.testclient import TestClient
@@ -125,6 +158,13 @@ class FakeEmbeddings:
def build_fixture(args, rng: random.Random) -> tuple[int, int]:
"""A user, settings and one adventure at production scale. Returns
(adventure_id, user_id)."""
# A SQLite run gets a brand-new temp file every time, so the fixture can
# assume an empty database. A Postgres scratch target persists between
# runs, and the second one would collide on the fixture user's unique
# email — so empty it first. Only ever reached for a target whose name
# passed the 'stress'/'scratch' guard at the top of this module.
if _stress_url:
Base.metadata.drop_all(bind=engine)
Base.metadata.create_all(bind=engine)
db = SessionLocal()
try: