Store embeddings as packed float32 instead of a JSON list
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.
It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.
Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.
The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.
Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
This commit is contained in:
co-authored by
Claude Opus 5
parent
7ee5ceea6c
commit
c56864877a
+21
-4
@@ -1,7 +1,8 @@
|
||||
from datetime import datetime, timezone
|
||||
|
||||
from sqlalchemy import (
|
||||
JSON, Boolean, Column, DateTime, Float, ForeignKey, Integer, String, Table, Text,
|
||||
JSON, Boolean, Column, DateTime, Float, ForeignKey, Integer, LargeBinary, String,
|
||||
Table, Text,
|
||||
)
|
||||
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
||||
|
||||
@@ -141,9 +142,15 @@ class Adventure(Base):
|
||||
class Memory(Base):
|
||||
"""Phase 6: an auto-summarized (or hand-written) fact about the adventure.
|
||||
|
||||
`embedding` is the raw vector as a JSON list (cosine ranking happens in
|
||||
Python — fine at bank sizes of a few hundred). NULL until embedded, which
|
||||
also marks it for backfill when an embedding model becomes available.
|
||||
The vector lives in `embedding_blob` as packed float32 (see vectors.py).
|
||||
NULL until embedded, which also marks it for backfill when an embedding
|
||||
model becomes available.
|
||||
|
||||
Cosine ranking happens in Python, which means the vectors cross the wire.
|
||||
The original comment here sized that by count — "fine at a few hundred" —
|
||||
and it was wrong by the only measure that mattered: a few hundred JSON
|
||||
vectors is ten megabytes, fetched fresh every turn. Weigh new columns in
|
||||
bytes.
|
||||
"""
|
||||
|
||||
__tablename__ = "memories"
|
||||
@@ -151,7 +158,17 @@ class Memory(Base):
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
|
||||
text: Mapped[str] = mapped_column(Text, default="")
|
||||
# Superseded by embedding_blob and written alongside it, so the two stay in
|
||||
# step until a follow-up migration drops this one. Still the column the
|
||||
# ranking path reads; that moves next.
|
||||
embedding: Mapped[list | None] = mapped_column(JSON, nullable=True)
|
||||
# The vector, little-endian float32. Deferred because it is wider than the
|
||||
# rest of the row put together and exactly one code path wants it: anything
|
||||
# bulk-loading memories (the Memories drawer, eviction, the embed queue)
|
||||
# must project the columns it needs rather than load whole entities.
|
||||
embedding_blob: Mapped[bytes | None] = mapped_column(
|
||||
LargeBinary, nullable=True, deferred=True
|
||||
)
|
||||
# Action index range this memory summarizes (null for manual memories).
|
||||
source_start: Mapped[int | None] = mapped_column(Integer, nullable=True)
|
||||
source_end: Mapped[int | None] = mapped_column(Integer, nullable=True)
|
||||
|
||||
Reference in New Issue
Block a user