Drop the JSON vector column, and fix what was hiding behind it

Migration 38 left memories.embedding in place so a rollback could still find
the vectors. Production has since been verified reading from embedding_blob,
so migration 42 drops it: 4 MB of a 99.6 MB database holding nothing anyone
reads.

Removing it surfaced a live bug. Changing your embedding model is supposed to
throw the bank's vectors away and let the post-turn pass rebuild them, because
two models' vectors are not comparable. The settings route did that by nulling
memories.embedding -- correct until 38 moved the vectors, after which it
cleared the dead column and left the blob intact with `embedded` still true.
_embed_pending filters on `embedded IS FALSE`, so it never saw those rows and
the bank went on ranking against the old model's vectors permanently.

Nothing would have reported it. cosine returns 0.0 on a width mismatch, so a
different-width model scores every memory zero and retrieval returns whichever
rows happen to sort first; a same-width model scores plausible garbage.

The bulk clear now sets both columns. It stays a bulk UPDATE rather than going
through set_vector -- loading the rows is the cost that whole path exists to
avoid -- so set_vector's docstring now names it as the one caller that
legitimately writes those columns by hand. No cache invalidation is added:
clearing `embedded` drops the rows out of the catalogue query, and set_vector
evicts each entry as the re-embed puts it back.

test_embedding_blob.py now rebuilds the pre-38 schema by hand where it tests
the backfill, since create_all no longer produces the column it converts from,
and asserts 42 removes it at the end of a full bootstrap -- 38 reads that
column and 42 drops it, so an upgrade that reordered them would arrive with an
empty bank.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
This commit is contained in:
parththakkar106
2026-08-17 13:44:59 +05:30
co-authored by Claude Opus 5
parent 85b188977e
commit 2c5909a268
7 changed files with 237 additions and 27 deletions
+13 -1
View File
@@ -51,14 +51,26 @@ def update_settings(
# Vectors from the old model have a different dimensionality/space;
# clear them so the post-turn task re-embeds with the new model.
# (This user's adventures only — settings are per-user now.)
#
# Both columns, and the flag. This is the one place that clears vectors
# in bulk rather than through memorybank.set_vector, and when the
# vectors moved to embedding_blob it kept nulling the old JSON column
# alone: the blob survived, `embedded` stayed true, and _embed_pending
# — which looks for embedded IS FALSE — never picked the rows up. The
# bank went on ranking against the previous model's vectors forever.
owned = (
db.query(models.Adventure.id)
.filter(models.Adventure.user_id == user.id)
.scalar_subquery()
)
db.query(models.Memory).filter(models.Memory.adventure_id.in_(owned)).update(
{"embedding": None}, synchronize_session=False
{"embedding_blob": None, "embedded": False}, synchronize_session=False
)
# No cache invalidation needed, and deliberately none added: clearing
# `embedded` drops these rows out of the catalogue query, so retrieval
# stops asking for them, and by the time _embed_pending puts one back
# it has gone through set_vector, which evicts that entry. The rule
# holds — anything that removes a memory from play self-corrects.
db.commit()
return settings