Store embeddings as packed float32 instead of a JSON list

A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.

It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.

Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.

The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.

Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
This commit is contained in:
parththakkar106
2026-08-16 21:17:44 +05:30
co-authored by Claude Opus 5
parent 7ee5ceea6c
commit c56864877a
7 changed files with 404 additions and 23 deletions
+21 -4
View File
@@ -1,7 +1,8 @@
from datetime import datetime, timezone
from sqlalchemy import (
JSON, Boolean, Column, DateTime, Float, ForeignKey, Integer, String, Table, Text,
JSON, Boolean, Column, DateTime, Float, ForeignKey, Integer, LargeBinary, String,
Table, Text,
)
from sqlalchemy.orm import Mapped, mapped_column, relationship
@@ -141,9 +142,15 @@ class Adventure(Base):
class Memory(Base):
"""Phase 6: an auto-summarized (or hand-written) fact about the adventure.
`embedding` is the raw vector as a JSON list (cosine ranking happens in
Python — fine at bank sizes of a few hundred). NULL until embedded, which
also marks it for backfill when an embedding model becomes available.
The vector lives in `embedding_blob` as packed float32 (see vectors.py).
NULL until embedded, which also marks it for backfill when an embedding
model becomes available.
Cosine ranking happens in Python, which means the vectors cross the wire.
The original comment here sized that by count — "fine at a few hundred" —
and it was wrong by the only measure that mattered: a few hundred JSON
vectors is ten megabytes, fetched fresh every turn. Weigh new columns in
bytes.
"""
__tablename__ = "memories"
@@ -151,7 +158,17 @@ class Memory(Base):
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
text: Mapped[str] = mapped_column(Text, default="")
# Superseded by embedding_blob and written alongside it, so the two stay in
# step until a follow-up migration drops this one. Still the column the
# ranking path reads; that moves next.
embedding: Mapped[list | None] = mapped_column(JSON, nullable=True)
# The vector, little-endian float32. Deferred because it is wider than the
# rest of the row put together and exactly one code path wants it: anything
# bulk-loading memories (the Memories drawer, eviction, the embed queue)
# must project the columns it needs rather than load whole entities.
embedding_blob: Mapped[bytes | None] = mapped_column(
LargeBinary, nullable=True, deferred=True
)
# Action index range this memory summarizes (null for manual memories).
source_start: Mapped[int | None] = mapped_column(Integer, nullable=True)
source_end: Mapped[int | None] = mapped_column(Integer, nullable=True)