Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.
Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.
one turn 3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
run_post_turn 3,139.1 kB -> 0.7 kB
Insights 3,223.7 kB -> 117.9 kB
Memories drawer ~3.1 MB -> 23.6 kB
A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.
The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.
memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.
Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
53 lines
2.0 KiB
Python
53 lines
2.0 KiB
Python
"""Storing and comparing embedding vectors.
|
|
|
|
A 1536-dimension vector written as a JSON list is about 31 KB, because every
|
|
component is spelled out as a decimal string of seventeen-odd digits. The same
|
|
vector as packed float32 is 6,144 bytes — a straight 5x, and the memory bank is
|
|
read in full on every turn, so those bytes are paid over and over.
|
|
|
|
**Float32 is not an approximation here.** The embedding endpoints return
|
|
vectors computed in float32, rendered into JSON as the shortest decimal string
|
|
that round-trips through a double; converting that back to float32 recovers the
|
|
original bits exactly. Nothing is lost that was ever there, which is why the
|
|
conversion needs no re-embedding and carries no retrieval-quality risk.
|
|
|
|
Dimensions are deliberately unchanged. Dropping to 512 or 768 would have saved
|
|
another 3x and cost an API call per stored memory to re-embed, against a bank
|
|
that the packing and the in-process cache together already make cheap.
|
|
"""
|
|
|
|
import math
|
|
import struct
|
|
import sys
|
|
from array import array
|
|
|
|
|
|
def pack(vector) -> bytes:
|
|
"""A vector as little-endian float32."""
|
|
return struct.pack(f"<{len(vector)}f", *vector)
|
|
|
|
|
|
def unpack(blob: bytes) -> array:
|
|
"""The inverse of `pack`. Length is implied: four bytes per component.
|
|
|
|
Returns an `array("f")` rather than a list, because these are held in
|
|
memory between turns: the array is the same 4 bytes a component the column
|
|
is, where a list of Python floats is eight times that. It indexes, zips and
|
|
lens like a list, which is all the ranking needs.
|
|
"""
|
|
vector = array("f")
|
|
vector.frombytes(blob)
|
|
if sys.byteorder != "little":
|
|
vector.byteswap()
|
|
return vector
|
|
|
|
|
|
def cosine(a: list[float], b: list[float]) -> float:
|
|
# Different lengths means the embedding model changed since this vector was
|
|
# stored; zip() would silently score garbage.
|
|
if len(a) != len(b):
|
|
return 0.0
|
|
dot = sum(x * y for x, y in zip(a, b))
|
|
norm = math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b))
|
|
return dot / norm if norm else 0.0
|