Store context_snapshot compressed
One column is 89% of the database and the free tier allows 512 MB. Reads were already solved -- the column is deferred, so a page load never touches it and one screen fetches one row at a time -- but nothing had costed storage, and storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per action, so the ceiling arrives around 5,400 actions and 944 are stored. Postgres already compresses it and only gets 1.7x. pglz is tuned for fast decompression of data a query might filter on, and nothing has ever filtered on an assembled prompt -- it is written once and read whole, rarely, by the Insights viewer. zlib gets 3.5x on the same text for a decompress on a request that already made an LLM call. Done as a TypeDecorator rather than a second column, so every call site still writes a dict and reads a dict back, and deferred/undefer/load_only keep naming the same attribute. Only the storage format moves. Migrations 43-45: add the bytea, convert into it, drop the original, rename. The backfill is the one destructive step in the file -- 44 removes the only other copy -- so it decompresses every row and compares it against what went in, and a row that fails aborts the run. The whole loop is one transaction, so an abort rolls the DROP back and the prompts are still there. Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column came out named context_snapshot, every snapshot compared equal and the one NULL stayed NULL. Postgres does not return the disk by itself: DROP COLUMN only marks the column gone and the backfill leaves a dead tuple per row, so the table peaks near twice its size before settling. The deploy needs one VACUUM FULL to collect it; the migration comment says so. The egress fixture's snapshots are prose now rather than "x" * 20_000, and the prose generator moved to tools/fakeprose.py so the harness and the tests share one definition. A repeated character compresses a thousandfold: against the old fixture a compressed column looked free and the byte ceilings would have been guarding nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
This commit is contained in:
co-authored by
Claude Opus 5
parent
a6cb49293c
commit
ae6e5af6c7
@@ -22,7 +22,7 @@ import json
|
||||
from sqlalchemy import inspect, text
|
||||
from sqlalchemy.engine import Engine
|
||||
|
||||
from . import vectors
|
||||
from . import compression, vectors
|
||||
from .database import Base
|
||||
|
||||
# (version, SQL to run when upgrading past it) — append only, never reorder.
|
||||
@@ -150,6 +150,34 @@ MIGRATIONS: list[tuple[int, str | dict[str, str]]] = [
|
||||
# now 4 MB of a 99.6 MB database holding nothing anyone reads. DROP COLUMN
|
||||
# is spelled the same on both dialects — SQLite has had it since 3.35.
|
||||
(42, "ALTER TABLE memories DROP COLUMN embedding"),
|
||||
# context_snapshot, compressed. 89% of the database is one column holding
|
||||
# assembled prompts nobody filters on and one screen reads, one row at a
|
||||
# time; Postgres already TOASTs it, but pglz only manages 1.7x and zlib
|
||||
# gets three to four on the same text. Reads were fixed by deferring it —
|
||||
# this is about the 512 MB the free tier allows.
|
||||
#
|
||||
# Three steps because a column cannot portably change type in place: add
|
||||
# the new one, convert into it (_backfill_context_snapshot, which verifies
|
||||
# every row round-trips before the old column goes), then swap the names so
|
||||
# the model keeps calling it context_snapshot.
|
||||
#
|
||||
# **Postgres does not hand the disk back on its own.** DROP COLUMN only
|
||||
# marks the column dropped, and the backfill's UPDATE leaves a dead tuple
|
||||
# per row, so the table gets *bigger* before it gets smaller: peak is
|
||||
# roughly twice the starting size while both columns are live. Plain
|
||||
# autovacuum makes that space reusable but does not shrink the files. The
|
||||
# deploy that ships this should follow it with, once:
|
||||
#
|
||||
# VACUUM FULL actions;
|
||||
#
|
||||
# which needs exclusive access and free space equal to the finished table.
|
||||
# On the 2026-08-17 figures that is 99.6 MB peaking near 200, settling at
|
||||
# about 53 once vacuumed, against a 512 MB tier. Skipping the vacuum is
|
||||
# safe and simply leaves the win unrealised.
|
||||
(43, {"sqlite": "ALTER TABLE actions ADD COLUMN context_snapshot_z BLOB",
|
||||
"default": "ALTER TABLE actions ADD COLUMN context_snapshot_z BYTEA"}),
|
||||
(44, "ALTER TABLE actions DROP COLUMN context_snapshot"),
|
||||
(45, "ALTER TABLE actions RENAME COLUMN context_snapshot_z TO context_snapshot"),
|
||||
]
|
||||
|
||||
LATEST_VERSION = max((v for v, _ in MIGRATIONS), default=1)
|
||||
@@ -158,6 +186,12 @@ LATEST_VERSION = max((v for v, _ in MIGRATIONS), default=1)
|
||||
WORLD_DELTA_VERSION = 36
|
||||
VARIANT_COUNT_VERSION = 37
|
||||
EMBEDDING_BLOB_VERSION = 38
|
||||
SNAPSHOT_COMPRESS_VERSION = 43
|
||||
|
||||
# Snapshots converted per round trip. Deliberately far smaller than
|
||||
# BACKFILL_BATCH: a vector is 6 KB and a snapshot is 232 KB, so 200 of these
|
||||
# would be 46 MB held at once.
|
||||
SNAPSHOT_BATCH = 50
|
||||
|
||||
# Vectors converted per round trip. Small enough that the backfill never holds
|
||||
# more than a few megabytes, large enough that it isn't a query per row.
|
||||
@@ -257,6 +291,52 @@ def _backfill_embedding_blob(conn) -> None:
|
||||
last_id = rows[-1][0]
|
||||
|
||||
|
||||
def _backfill_context_snapshot(conn) -> None:
|
||||
"""Compress actions.context_snapshot into actions.context_snapshot_z.
|
||||
|
||||
Runs between migration 43 and 44, which is the only window where both
|
||||
columns exist. Migration 44 drops the original, so unlike every other
|
||||
backfill here this one is destructive if it is wrong — and every row it
|
||||
converts is somebody's game. So each row is decompressed again and
|
||||
compared against what went in before it counts as converted, and a row
|
||||
that fails to round-trip aborts the whole run rather than being skipped:
|
||||
the transaction rolls back, the DROP never happens, and the prompts are
|
||||
still there to try again.
|
||||
|
||||
Reads the JSON the same defensive way as the vector backfill — SQLite
|
||||
hands back a raw string, psycopg has already parsed it.
|
||||
"""
|
||||
last_id = 0
|
||||
while True:
|
||||
rows = conn.execute(
|
||||
text("""
|
||||
SELECT id, context_snapshot FROM actions
|
||||
WHERE context_snapshot IS NOT NULL
|
||||
AND context_snapshot_z IS NULL AND id > :last
|
||||
ORDER BY id LIMIT :batch
|
||||
"""),
|
||||
{"last": last_id, "batch": SNAPSHOT_BATCH},
|
||||
).all()
|
||||
if not rows:
|
||||
return
|
||||
for row_id, stored in rows:
|
||||
value = json.loads(stored) if isinstance(stored, str) else stored
|
||||
if value is None:
|
||||
continue
|
||||
packed = compression.pack(value)
|
||||
if compression.unpack(packed) != value:
|
||||
raise RuntimeError(
|
||||
f"context_snapshot for action {row_id} did not survive a "
|
||||
"compress/decompress round trip; refusing to drop the "
|
||||
"original column"
|
||||
)
|
||||
conn.execute(
|
||||
text("UPDATE actions SET context_snapshot_z = :z WHERE id = :id"),
|
||||
{"z": packed, "id": row_id},
|
||||
)
|
||||
last_id = rows[-1][0]
|
||||
|
||||
|
||||
def _get_version(conn) -> int:
|
||||
if conn.dialect.name == "sqlite":
|
||||
return conn.execute(text("PRAGMA user_version")).scalar() or 1
|
||||
@@ -301,6 +381,11 @@ def bootstrap(engine: Engine) -> None:
|
||||
_backfill_variant_count(conn)
|
||||
if version == EMBEDDING_BLOB_VERSION:
|
||||
_backfill_embedding_blob(conn)
|
||||
# Must land between 43 (add the column) and 44 (drop the old
|
||||
# one). The loop is one transaction, so if this raises, the
|
||||
# DROP rolls back with it and the prompts are still there.
|
||||
if version == SNAPSHOT_COMPRESS_VERSION:
|
||||
_backfill_context_snapshot(conn)
|
||||
current = version
|
||||
_set_version(conn, current)
|
||||
_encrypt_plaintext_api_keys(conn)
|
||||
|
||||
Reference in New Issue
Block a user