Drop the JSON vector column, and fix what was hiding behind it
Migration 38 left memories.embedding in place so a rollback could still find the vectors. Production has since been verified reading from embedding_blob, so migration 42 drops it: 4 MB of a 99.6 MB database holding nothing anyone reads. Removing it surfaced a live bug. Changing your embedding model is supposed to throw the bank's vectors away and let the post-turn pass rebuild them, because two models' vectors are not comparable. The settings route did that by nulling memories.embedding -- correct until 38 moved the vectors, after which it cleared the dead column and left the blob intact with `embedded` still true. _embed_pending filters on `embedded IS FALSE`, so it never saw those rows and the bank went on ranking against the old model's vectors permanently. Nothing would have reported it. cosine returns 0.0 on a width mismatch, so a different-width model scores every memory zero and retrieval returns whichever rows happen to sort first; a same-width model scores plausible garbage. The bulk clear now sets both columns. It stays a bulk UPDATE rather than going through set_vector -- loading the rows is the cost that whole path exists to avoid -- so set_vector's docstring now names it as the one caller that legitimately writes those columns by hand. No cache invalidation is added: clearing `embedded` drops the rows out of the catalogue query, and set_vector evicts each entry as the re-embed puts it back. test_embedding_blob.py now rebuilds the pre-38 schema by hand where it tests the backfill, since create_all no longer produces the column it converts from, and asserts 42 removes it at the end of a full bootstrap -- 38 reads that column and 42 drops it, so an upgrade that reordered them would arrive with an empty bank. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
This commit is contained in:
co-authored by
Claude Opus 5
parent
85b188977e
commit
2c5909a268
@@ -26,7 +26,7 @@ import asyncio
|
||||
from datetime import timedelta
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import event
|
||||
from sqlalchemy import event, inspect as sa_inspect
|
||||
|
||||
from app import memorybank, models
|
||||
from app.database import Base, SessionLocal, engine
|
||||
@@ -206,13 +206,14 @@ def memory_selects(statements):
|
||||
]
|
||||
|
||||
|
||||
def test_the_json_column_is_never_selected(db, adventure, settings, bank, sql_log):
|
||||
"""`memories.embedding` is dead weight kept only until a follow-up
|
||||
migration drops it. If anything still reads it, dropping it breaks."""
|
||||
retrieve(adventure, settings, StubEmbedder())
|
||||
offenders = [s for s in memory_selects(sql_log) if "memories.embedding " in s
|
||||
or s.rstrip().endswith("memories.embedding")]
|
||||
assert offenders == [], f"the JSON column was read:\n{offenders[0][:300]}"
|
||||
def test_the_json_column_is_gone(db):
|
||||
"""`memories.embedding` held the vectors before migration 38 and nothing
|
||||
read it afterwards; migration 42 dropped it. Bringing it back would restore
|
||||
4 MB of dead weight and a second place vectors can be written from — which
|
||||
is how the model-switch bug happened (test_embedding_model_switch.py)."""
|
||||
columns = {c["name"] for c in sa_inspect(engine).get_columns("memories")}
|
||||
assert "embedding" not in columns
|
||||
assert {"embedding_blob", "embedded"} <= columns
|
||||
|
||||
|
||||
def test_the_catalogue_query_carries_no_vectors(db, adventure, settings, bank, sql_log):
|
||||
|
||||
Reference in New Issue
Block a user