test_egress.py asserted which columns a statement names, which is the shape
both of this project's egress blowouts took. It would all still pass if a
response grew tenfold within the columns it is allowed to read -- and a story
that keeps getting longer does exactly that. Production's longest adventure is
607 actions where the plan assumed 200.
So dbmeter, which was built to be importable from tests and was not yet used
by any, now backs four byte ceilings: the page load, the action list, and one
action's snapshot fetched on demand. Budgets are per action rather than
absolute, so they mean the same thing whatever size the fixture is set to, and
generous -- 3 kB against a real 994 B. They are there to catch an order of
magnitude, not to freeze a byte count.
The fourth test is the one that keeps the other three honest. A ceiling proves
nothing unless the thing it excludes would breach it, so it undefers the
snapshot on purpose and asserts the same twelve rows cost more than ten times
the budget. If the fixture ever shrinks below the point where that holds, that
test fails rather than the ceilings quietly passing on nothing.
Meter grows detach() and a context manager. A script exits and takes the
wrapping with it; a test does not, and one test leaving the shared engine
metered would charge bytes to a scope nobody opened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Both egress blowouts this project has had were one query fetching a column
nobody read, and a statement count would have shown nothing wrong in either.
So the meter counts bytes, at the DBAPI cursor -- everything that crosses
that line crossed the wire.
tools/stress_session.py drives a production-shaped adventure through the real
routes with only the network faked. It reproduces both figures measured
directly on production: 426.7 kB for a 200-action page load against 423 KB,
and 3,258.7 kB for one turn against 3,153 kB.
The memory bank is on by default, which is the whole point -- the previous
harness ran without an embedding model, so retrieval returned early and the
heaviest read in a turn never happened. --no-embeddings reproduces that
deliberately, and the gap is 29x.
It also turned up two callers the production SQL could not see: run_post_turn
walks the whole bank again every turn, and Insights pays for it a third time.
A played turn costs ~6.4 MB, not 3.2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7