Measure database egress in bytes, from inside the repo

Both egress blowouts this project has had were one query fetching a column
nobody read, and a statement count would have shown nothing wrong in either.
So the meter counts bytes, at the DBAPI cursor -- everything that crosses
that line crossed the wire.

tools/stress_session.py drives a production-shaped adventure through the real
routes with only the network faked. It reproduces both figures measured
directly on production: 426.7 kB for a 200-action page load against 423 KB,
and 3,258.7 kB for one turn against 3,153 kB.

The memory bank is on by default, which is the whole point -- the previous
harness ran without an embedding model, so retrieval returned early and the
heaviest read in a turn never happened. --no-embeddings reproduces that
deliberately, and the gap is 29x.

It also turned up two callers the production SQL could not see: run_post_turn
walks the whole bank again every turn, and Insights pays for it a third time.
A played turn costs ~6.4 MB, not 3.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
This commit is contained in:
parththakkar106
2026-08-16 21:11:56 +05:30
co-authored by Claude Opus 5
parent 03c6707ca7
commit 7ee5ceea6c
4 changed files with 705 additions and 3 deletions
+4
View File
@@ -0,0 +1,4 @@
"""Developer tools that are not part of the deployed app.
Nothing under `app/` may import from here.
"""
+323
View File
@@ -0,0 +1,323 @@
"""How many bytes the database was actually asked for.
Query *counts* are easy to see and have never been the problem here. Both
egress blowouts this project has had were one query fetching a column nobody
read: `context_snapshot` on every action, then every memory's embedding on
every turn. Counting statements would have shown nothing wrong in either case,
and at development scale — ten rows, no embeddings — so would a stopwatch.
So this meter measures bytes, and it measures them at the only place the truth
is available: the DBAPI cursor, after the driver has decoded a row and before
SQLAlchemy has turned it into anything. Every value that crosses that line
crossed the wire.
Usage:
meter = Meter()
meter.attach(engine)
with meter.scope("turn"):
... # anything that touches the database
print(meter.render())
`attach()` wraps the pool's connection factory, so it reuses the engine the app
already configured rather than rebuilding one beside it — nothing about
connect args, pre-ping or the SQLite foreign-key pragma has to be repeated
here, and drift between the metered engine and the real one is impossible.
Existing pooled connections are dropped first, so a connection opened before
attaching cannot quietly stay unmetered.
Sizes are of the decoded values, not of the protocol framing: the count is what
the payload weighs, and ignores per-row and per-packet overhead. Text and bytes
are exact. Numbers are charged their binary width, which is what the Postgres
binary protocol sends and close enough elsewhere. JSON is the case that matters
most and the one to be careful about: SQLite hands back the raw string (exact),
while psycopg parses `json`/`jsonb` into Python before this sees it, so the
value is re-serialised to size it. Re-serialising is within a byte or two of
the original for machine-written JSON, which is all this project stores.
"""
from __future__ import annotations
import json
import re
from collections import defaultdict
from contextlib import contextmanager
from dataclasses import dataclass, field
# The table a statement is charged to. Only the first match is used: a join
# reads mostly from its driving table, and splitting a row across tables would
# need the result metadata, which is more machinery than the answer is worth.
_TABLE_RE = re.compile(
r"\b(?:FROM|JOIN|INTO|UPDATE)\s+\"?([A-Za-z_][A-Za-z_0-9]*)\"?", re.IGNORECASE
)
_WHITESPACE_RE = re.compile(r"\s+")
def value_bytes(value) -> int:
"""The payload weight of one decoded column value."""
if value is None:
return 0
if isinstance(value, (bytes, bytearray)):
return len(value)
if isinstance(value, memoryview):
return value.nbytes
if isinstance(value, str):
return len(value.encode("utf-8"))
if isinstance(value, bool):
return 1
if isinstance(value, (int, float)):
return 8
if isinstance(value, (dict, list)):
# A JSON column the driver already parsed (psycopg does; SQLite does
# not). Separators match what a database emits — no spaces.
return len(json.dumps(value, separators=(",", ":"), default=str).encode("utf-8"))
return len(str(value).encode("utf-8"))
def row_bytes(row) -> int:
return sum(value_bytes(v) for v in row)
def _table_of(statement: str) -> str:
match = _TABLE_RE.search(statement)
if match:
return match.group(1).lower()
# No table to name — BEGIN, a PRAGMA, a savepoint. Group those under the
# keyword so they stay countable instead of collapsing into one "?" bucket.
head = statement.strip().split(None, 1)
return head[0].lower()[:20] if head else "(empty)"
def _one_line(statement: str, width: int = 132) -> str:
collapsed = _WHITESPACE_RE.sub(" ", statement).strip()
return collapsed if len(collapsed) <= width else collapsed[: width - 1] + "…"
@dataclass
class Tally:
statements: int = 0
rows: int = 0
fetched: int = 0 # bytes
@dataclass
class Scope:
"""What one measured stretch of work asked the database for."""
name: str
total: Tally = field(default_factory=Tally)
by_table: dict[str, Tally] = field(default_factory=lambda: defaultdict(Tally))
by_statement: dict[str, Tally] = field(default_factory=lambda: defaultdict(Tally))
def add_statement(self, statement: str) -> None:
self.total.statements += 1
self.by_table[_table_of(statement)].statements += 1
self.by_statement[statement].statements += 1
def add_rows(self, statement: str, rows: int, nbytes: int) -> None:
for tally in (
self.total,
self.by_table[_table_of(statement)],
self.by_statement[statement],
):
tally.rows += rows
tally.fetched += nbytes
class Meter:
"""Collects what the metered cursors report, grouped by scope.
Scopes nest: a statement is charged to every scope currently open, so an
inner "retrieve memories" and an outer "turn" both see it.
"""
def __init__(self) -> None:
self._open: list[Scope] = []
self.scopes: list[Scope] = []
self._attached_pools: list = []
# -------------------------------------------------------------- recording
@contextmanager
def scope(self, name: str):
scope = Scope(name)
self._open.append(scope)
try:
yield scope
finally:
self._open.pop()
self.scopes.append(scope)
def note_statement(self, statement: str) -> None:
for scope in self._open:
scope.add_statement(statement)
def note_rows(self, statement: str, rows: int, nbytes: int) -> None:
for scope in self._open:
scope.add_rows(statement, rows, nbytes)
# ------------------------------------------------------------- attaching
def attach(self, engine) -> None:
"""Meter every connection this engine opens from now on."""
# Drop pooled connections created before now; they were built by the
# unwrapped creator and would go on reporting nothing.
engine.dispose()
pool = engine.pool
creator = pool._creator # the factory create_engine() built from the URL
if getattr(creator, "_dbmeter", None) is self:
return
meter = self
def metered_creator():
return _MeteredConnection(creator(), meter)
metered_creator._dbmeter = self
pool._creator = metered_creator
self._attached_pools.append(pool)
# ------------------------------------------------------------- reporting
def render(self, *, statements: int = 5) -> str:
return "\n".join(render_scope(s, statements=statements) for s in self.scopes)
# ------------------------------------------------------------ DBAPI wrappers
class _MeteredConnection:
"""A DBAPI connection that hands out metered cursors.
Everything else is delegated: the dialects reach for driver-specific
attributes (`isolation_level` on SQLite, `info` and `autocommit` on
psycopg) and this must stay transparent to all of them.
"""
def __init__(self, connection, meter: Meter) -> None:
object.__setattr__(self, "_connection", connection)
object.__setattr__(self, "_meter", meter)
def cursor(self, *args, **kwargs):
return _MeteredCursor(self._connection.cursor(*args, **kwargs), self._meter)
def __getattr__(self, name):
return getattr(object.__getattribute__(self, "_connection"), name)
def __setattr__(self, name, value):
setattr(self._connection, name, value)
def __enter__(self):
self._connection.__enter__()
return self
def __exit__(self, *exc):
return self._connection.__exit__(*exc)
class _MeteredCursor:
"""Counts the bytes of every row handed back.
The fetch methods are wrapped rather than `execute`, because what a
statement *costs* is not knowable when it is sent — `SELECT * FROM
memories` and `SELECT count(*) FROM memories` look alike going out and
differ by three megabytes coming back.
"""
def __init__(self, cursor, meter: Meter) -> None:
self.__dict__["_cursor"] = cursor
self.__dict__["_meter"] = meter
self.__dict__["_statement"] = ""
# -- execution
def execute(self, statement, *args, **kwargs):
self.__dict__["_statement"] = statement
self._meter.note_statement(statement)
return self._cursor.execute(statement, *args, **kwargs)
def executemany(self, statement, *args, **kwargs):
self.__dict__["_statement"] = statement
self._meter.note_statement(statement)
return self._cursor.executemany(statement, *args, **kwargs)
# -- fetching
def _charge(self, rows) -> None:
self._meter.note_rows(
self._statement, len(rows), sum(row_bytes(r) for r in rows)
)
def fetchone(self):
row = self._cursor.fetchone()
if row is not None:
self._charge([row])
return row
def fetchmany(self, *args, **kwargs):
rows = self._cursor.fetchmany(*args, **kwargs)
self._charge(rows)
return rows
def fetchall(self):
rows = self._cursor.fetchall()
self._charge(rows)
return rows
def __iter__(self):
# Special methods are looked up on the type, so this cannot be left to
# __getattr__ the way the rest of the driver surface is.
for row in self._cursor:
self._charge([row])
yield row
def __getattr__(self, name):
return getattr(self.__dict__["_cursor"], name)
def __setattr__(self, name, value):
setattr(self.__dict__["_cursor"], name, value)
def __enter__(self):
self._cursor.__enter__()
return self
def __exit__(self, *exc):
return self._cursor.__exit__(*exc)
# --------------------------------------------------------------- formatting
def kb(nbytes: int) -> str:
return f"{nbytes / 1024:,.1f} kB"
def render_scope(scope: Scope, *, statements: int = 5) -> str:
lines = [
"",
f"── {scope.name} " + "─" * max(0, 62 - len(scope.name)),
f" {scope.total.statements} statements · {scope.total.rows} rows · "
f"{kb(scope.total.fetched)} fetched",
]
if scope.by_table:
lines += ["", f" {'table':<20}{'stmts':>7}{'rows':>8}{'fetched':>16}"]
ranked = sorted(
scope.by_table.items(), key=lambda kv: kv[1].fetched, reverse=True
)
for table, tally in ranked:
share = tally.fetched / scope.total.fetched if scope.total.fetched else 0
lines.append(
f" {table:<20}{tally.statements:>7}{tally.rows:>8}"
f"{kb(tally.fetched):>16}{share:>7.0%}"
)
heavy = sorted(
scope.by_statement.items(), key=lambda kv: kv[1].fetched, reverse=True
)[:statements]
heavy = [(s, t) for s, t in heavy if t.fetched]
if heavy:
lines += ["", " heaviest statements"]
for statement, tally in heavy:
lines.append(f" {kb(tally.fetched):>14} {tally.statements}x {_one_line(statement)}")
return "\n".join(lines)
+338
View File
@@ -0,0 +1,338 @@
"""Drive a production-sized adventure through the real routes and report what
each one costs in database bytes.
cd backend
.venv/Scripts/python.exe -m tools.stress_session
.venv/Scripts/python.exe -m tools.stress_session --actions 200 --memories 100
.venv/Scripts/python.exe -m tools.stress_session --no-embeddings
**The memory bank is ON by default, and that is the point.** The round-two
stress harness ran without an embedding model configured, and embedding
providers are BYOK-only by construction, so `retrieve_memories` returned early
every time — the whole exercise measured the turn loop with its heaviest read
switched off, and reported 23 MB for a playthrough that actually costs an order
of magnitude more. `--no-embeddings` reproduces that blindness deliberately, to
show the gap; it is never the default and it prints a warning.
Everything here is synthetic. The fixture is generated to production *shape* —
1536-dimension embeddings, ~74 KB context snapshots, retry variants — and no
real adventure, user or backup is ever read.
Only the network is faked: the LLM and the embedding endpoint. Routing,
sessions, the ORM, the scripting engine and the context builder are the real
ones, because the bugs this exists to catch live in exactly the layer a mock
would replace.
It runs on a throwaway SQLite file rather than Postgres. What is being measured
is which columns of which rows a code path asks for, and that is decided by the
ORM, identically on both. The dialects disagree on how a value is encoded on
the wire — JSON especially — so treat the absolute figures as production-shaped
rather than production-exact, and compare before against after.
Calibration, against the two figures measured directly on production
(2026-08-16): a 200-action page load reported 426.7 kB here against 423 KB
there, and one turn on a 100-memory bank reported 3,258.7 kB against 3,153 kB.
"""
import os
import tempfile
# Must precede the app import: database.py reads these at module scope.
_tmp = tempfile.NamedTemporaryFile(suffix=".db", delete=False)
_tmp.close()
os.environ["AIDND_DB_PATH"] = _tmp.name
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
import argparse
import asyncio
import random
import sys
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, memorybank, models, security
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.providers import PromptParts
from app.routers import adventures
from .dbmeter import Meter, kb
EMBEDDING_DIMS = 1536
# ~74 KB, which is what a real snapshot weighs in production: the assembled
# prompt is nearly all of it.
SNAPSHOT_SYSTEM = "You are a masterful storyteller. " * 400
SNAPSHOT_STORY = "The corridor narrows and the torchlight gutters. " * 1200
_PARAGRAPH = (
"The corridor narrows until your shoulders brush wet stone, and the "
"torchlight gutters in a draught that smells of cold iron. Somewhere ahead, "
"water is moving. You count nine paces before the passage opens into a "
"chamber whose ceiling is lost in the dark, and the sound of your own "
"breathing comes back to you a half-second late.\n\n"
"Gwen catches your sleeve without a word and points at the floor, where a "
"line of pale grit has been laid across the threshold in a deliberate arc.\n\n"
)
PLAYER_INPUT = "> You crouch and look more closely at the grit on the floor."
# Rebound by main() to --narration-bytes. An AI action's length is what makes a
# page load expensive, and it is the one fixture dimension that cannot be
# guessed from the schema: production averages ~2.1 KB across all actions,
# which is ~4 KB of narration alternating with a one-line player input.
NARRATION = _PARAGRAPH
MEMORY_TEXT = (
"You found a bandit camp above the ford and agreed to guide Gwen through "
"the tunnels in exchange for the iron key she took from the quartermaster."
)
# --------------------------------------------------------------- fake network
class FakeProvider:
"""The LLM. Streams one fixed line; no network, no cost, no variance."""
def __init__(self, *a, **k):
pass
async def generate(self, parts: PromptParts, *, temperature, max_tokens):
yield ("text", NARRATION)
async def complete(self, system, user, *, max_tokens=None):
return MEMORY_TEXT
class FakeEmbeddings:
"""The embedding endpoint. Returns vectors of the real width, so what the
turn writes back weighs what production weighs."""
def __init__(self, rng: random.Random):
self.rng = rng
async def embed(self, texts: list[str]) -> list[list[float]]:
return [
[self.rng.uniform(-1.0, 1.0) for _ in range(EMBEDDING_DIMS)]
for _ in texts
]
# ------------------------------------------------------------------- fixture
def build_fixture(args, rng: random.Random) -> tuple[int, int]:
"""A user, settings and one adventure at production scale. Returns
(adventure_id, user_id)."""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
try:
user = models.User(is_guest=False, email="stress@example.invalid")
db.add(user)
db.flush()
db.add(models.Settings(
user_id=user.id,
api_key=security.encrypt_secret("stress-key"),
model="stress-model",
endpoint_url="https://fake.invalid/v1",
# The default this whole tool exists to stop anyone forgetting.
embedding_model="" if args.no_embeddings else "openai/text-embedding-3-small",
memory_bank_capacity=args.capacity,
))
adventure = models.Adventure(
user_id=user.id,
title="Stress",
script_state={},
memory_bank_enabled=True,
# Off so a turn measures the turn. The post-turn pass is its own
# shape below; letting it fire mid-measurement would mix the two.
auto_summarize=False,
)
db.add(adventure)
db.flush()
for i in range(args.actions):
is_ai = bool(i % 2)
db.add(models.Action(
adventure_id=adventure.id,
index=i,
type="ai" if is_ai else "do",
text=NARRATION if is_ai else PLAYER_INPUT,
context_snapshot={"system": SNAPSHOT_SYSTEM, "story": SNAPSHOT_STORY},
world_delta={"delta": {"player.hp": -3},
"applied": [{"path": "player.hp", "old": 88, "new": 85}]},
# Every third AI turn was retried once, so the retry history is
# carrying weight a list response must not pay for.
variants=(
[{"text": NARRATION, "reasoning": None, "script_state": {},
"created_at": "2026-01-01T00:00:00"} for _ in range(2)]
if is_ai and i % 6 == 1 else None
),
variant_count=2 if is_ai and i % 6 == 1 else 0,
))
for i in range(args.memories):
db.add(models.Memory(
adventure_id=adventure.id,
text=f"{MEMORY_TEXT} ({i})",
embedding=[rng.uniform(-1.0, 1.0) for _ in range(EMBEDDING_DIMS)],
source_start=i * memorybank.MEMORY_INTERVAL,
source_end=i * memorybank.MEMORY_INTERVAL + memorybank.MEMORY_INTERVAL - 1,
))
db.commit()
return adventure.id, user.id
finally:
db.close()
def install_fakes(user_id: int, rng: random.Random) -> None:
embeddings = FakeEmbeddings(rng)
adventures.OpenAICompatibleProvider = FakeProvider
memorybank.embedding_provider = lambda settings: embeddings
memorybank.summary_provider = lambda settings: FakeProvider()
auth.resolve_provider_config = lambda s, **k: auth.ProviderConfig(
"https://fake.invalid/v1", "stress-key", "stress-model", False
)
limits.rate_limit = lambda *a, **k: None
limits.check_row_cap = lambda *a, **k: None
# Fire-and-forget post-turn work would land inside whichever scope happened
# to be open. It is measured on purpose, as its own shape.
memorybank.schedule_post_turn = lambda adventure: None
def _current_user(db=Depends(get_db)):
return db.get(models.User, user_id)
app.dependency_overrides[auth.get_current_user] = _current_user
# --------------------------------------------------------------------- shapes
def shape_list(client, meter, adv_id):
"""The adventures index — every adventure's latest narration."""
with meter.scope("GET /adventures (index)"):
r = client.get("/api/adventures")
_check(r)
def shape_load(client, meter, adv_id):
"""Opening a finished adventure: the whole story, in one response."""
with meter.scope(f"GET /adventures/{{id}} (page load)"):
r = client.get(f"/api/adventures/{adv_id}")
_check(r)
def shape_turn(client, meter, adv_id):
"""One played turn, memory retrieval included."""
with meter.scope("POST /adventures/{id}/actions (one turn)"):
r = client.post(
f"/api/adventures/{adv_id}/actions",
json={"type": "do", "text": "look more closely at the grit"},
)
_check(r)
def shape_insights(client, meter, adv_id):
"""The Insights dry run — assembles a context without spending a turn."""
with meter.scope("GET /adventures/{id}/context (insights)"):
r = client.get(f"/api/adventures/{adv_id}/context")
_check(r)
def shape_post_turn(client, meter, adv_id):
"""Summarization, embedding and eviction, after the turn is saved."""
with meter.scope("run_post_turn (background)"):
asyncio.run(memorybank.run_post_turn(adv_id))
SHAPES = {
"list": shape_list,
"load": shape_load,
"turn": shape_turn,
"insights": shape_insights,
"post_turn": shape_post_turn,
}
def _check(response) -> None:
if response.status_code >= 400:
sys.exit(f"shape failed: {response.status_code} {response.text[:400]}")
# ----------------------------------------------------------------------- main
def parse_args(argv=None):
p = argparse.ArgumentParser(
prog="tools.stress_session", description=__doc__.splitlines()[0]
)
p.add_argument("--actions", type=int, default=200,
help="story actions in the fixture (default: 200)")
p.add_argument("--memories", type=int, default=100,
help="memories, all embedded (default: 100)")
p.add_argument("--capacity", type=int, default=200,
help="Settings.memory_bank_capacity (default: 200)")
p.add_argument("--narration-bytes", type=int, default=4000,
help="length of an AI action's text; production averages "
"~2.1 KB per action alternating with player input "
"(default: 4000)")
p.add_argument("--shapes", default=",".join(SHAPES),
help=f"comma-separated subset of: {', '.join(SHAPES)}")
p.add_argument("--repeat", type=int, default=1,
help="run each shape this many times (default: 1)")
p.add_argument("--no-embeddings", action="store_true",
help="unset the embedding model — reproduces the round-two "
"blind spot, where the bank's cost is invisible")
p.add_argument("--seed", type=int, default=7)
p.add_argument("--statements", type=int, default=5,
help="heaviest statements to print per shape (default: 5)")
return p.parse_args(argv)
def main(argv=None) -> int:
args = parse_args(argv)
chosen = [s.strip() for s in args.shapes.split(",") if s.strip()]
unknown = [s for s in chosen if s not in SHAPES]
if unknown:
sys.exit(f"unknown shape(s): {', '.join(unknown)}")
global NARRATION
repeats = max(1, -(-args.narration_bytes // len(_PARAGRAPH)))
NARRATION = (_PARAGRAPH * repeats)[: args.narration_bytes]
rng = random.Random(args.seed)
adv_id, user_id = build_fixture(args, rng)
install_fakes(user_id, rng)
meter = Meter()
# After the fixture: building it is a write path nobody plays, and its
# bytes would drown everything the shapes report.
meter.attach(engine)
print(f"fixture: {args.actions} actions × {args.narration_bytes} B · "
f"{args.memories} memories × {EMBEDDING_DIMS} dims · "
f"capacity {args.capacity}")
if args.no_embeddings:
print("WARNING: embedding model unset — memory retrieval will return "
"early and the bank's cost will not appear below.")
else:
print("memory bank: ON (embedding model configured)")
with TestClient(app) as client:
for _ in range(args.repeat):
for name in chosen:
SHAPES[name](client, meter, adv_id)
print(meter.render(statements=args.statements))
print()
print(f"{'total across all shapes':<44}{kb(sum(s.total.fetched for s in meter.scopes)):>16}")
app.dependency_overrides.clear()
adventures._active_turns.clear()
return 0
if __name__ == "__main__":
raise SystemExit(main())
+40 -3
View File
@@ -87,13 +87,15 @@ no denormalisation to fix. It is a *format* problem plus a *fetch-frequency* pro
+ `stress_session.py`. The originals lived outside the repo and are gone. **Default it
to running with an embedding model configured**, since that omission is precisely what
hid this finding. Do this first so every item below is measured, not assumed.
**Done 2026-08-16** — see the baseline below.
2. **Migration 38 — `memories.embedding_blob` (`LargeBinary`).** Backfill in Python
(`struct.pack(f"<{n}f", *vec)`); the conversion cannot be expressed in portable SQL, so
unlike migration 36/37 this one does pay a one-time 4 MB read. Drop the old JSON column
in a follow-up migration once verified, not in the same one.
3. **Read path.** `retrieve_memories` queries `memories` directly with
`forgotten = false AND embedding_blob IS NOT NULL` in **SQL**, not Python. Unpack with
`struct`/`numpy`.
`struct`/`numpy`. Same for `_evict_over_capacity` and `_embed_pending`, which walk the
same relationship for a count and for the unembedded rows (see the baseline above).
4. **Vector cache.** Keyed by `adventure_id`, invalidated on memory create, evict and
delete. Must survive the retry/undo paths that prune memories.
5. **Capacity default 200 → 80.** Existing adventures inherit it, so adventure 25 evicts
@@ -102,6 +104,38 @@ no denormalisation to fix. It is a *format* problem plus a *fetch-frequency* pro
finished adventure — the last open item from round two. Load the newest turns, fetch
older ones as the reader scrolls up.
## Baseline from the harness (2026-08-16)
`python -m tools.stress_session`, 200 actions × 4 KB, 100 memories × 1536 dims:
| shape | fetched | memories' share |
|---|---|---|
| `GET /adventures` (index) | 0.7 kB | — |
| `GET /adventures/{id}` (page load) | 426.7 kB | — |
| `POST /adventures/{id}/actions` (one turn) | **3,258.7 kB** | 96% |
| `GET /adventures/{id}/context` (Insights) | 3,223.7 kB | 97% |
| `run_post_turn` (background) | **3,139.1 kB** | 100% |
It reproduces both production figures independently — 426.7 kB against the measured
423 KB page load, 3,258.7 kB against 3,024 + 129 kB for a turn, and 31.4 KB per
embedding against 31.0 KB. `--no-embeddings` reports 112.0 kB for the same turn, so
the round-two blind spot is now a **29x** gap anyone can see in one flag.
**Two findings the SQL measurement could not have shown**, both the same root cause
in a different caller:
- **`run_post_turn` fetches the whole bank again**, every turn. `_evict_over_capacity`
walks `adventure.memories` to count the active ones, and `_embed_pending` walks it to
find the unembedded ones. So a played turn actually costs ~6.4 MB, not 3.2 — the
original estimate was half the real number.
- **Insights pays it a third time**, on a page the player can open repeatedly without
spending a turn.
So step 3 below is not just `retrieve_memories`: **every walk of
`adventure.memories` has to go**. `_evict_over_capacity` wants a count and an ordering,
`_embed_pending` wants rows where `embedding IS NULL` — neither needs a single vector,
and both are pure SQL.
## Guardrails to add with this work
- **Query-count / byte assertions per endpoint**, extending the `test_egress.py` idea:
@@ -119,8 +153,11 @@ backups or storage start to hurt.
## Verification
- Harness: one turn on adventure 25, memory bank **on**, before and after. Target is
3,024 kB → ~600 kB cold, ~0 warm.
- Harness: `python -m tools.stress_session`, memory bank **on**, before and after,
against the baseline table above. Target for the turn shape is 3,258 kB → ~700 kB
cold, ~130 kB warm (the cache leaves only what a turn reads besides the bank).
`run_post_turn` should fall to roughly nothing: neither of its two walks needs a
vector at all.
- Re-run the round-two shapes with an embedding model configured, so the 200-action
playthrough number is finally honest.
- The existing `test_egress.py` guard must still pass — nothing here should touch the