The previous commit recorded the service as suspended. It is not. Render
appended a suffix to the `ai-dnd` service name in render.yaml, and plain
ai-dnd.onrender.com is a different, suspended service whose 503 page reads
"suspended by its owner" -- indistinguishable from this deploy being down
unless you notice the host is wrong.
The real host answers {"ok":true} on /api/health in 0.6s, mints a guest,
returns an empty adventure list for that guest and serves the starter
scenarios. The authoritative link is the one docs/index.html points at, not
the service name in the blueprint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
213 lines
10 KiB
Markdown
213 lines
10 KiB
Markdown
# Where things stand
|
|
|
|
Read this first when picking the project back up. Updated at the end of a working
|
|
session; the per-phase plan files hold the detail, this holds the thread.
|
|
|
|
**Last updated: 2026-08-17.**
|
|
|
|
---
|
|
|
|
## The live URL is not the one in render.yaml
|
|
|
|
**Production is `https://ai-dnd-1gmp.onrender.com`.** Render appended a suffix to the
|
|
`ai-dnd` service name in `render.yaml`, and plain `ai-dnd.onrender.com` belongs to a
|
|
different, suspended service that answers 503 with "suspended by its owner" — which is
|
|
easy to mistake for this deploy being down. The authoritative link is the one the
|
|
project page points at (`docs/index.html`), not the service name in the blueprint.
|
|
`GET /api/health` on the real host returns `{"ok":true}`.
|
|
|
|
---
|
|
|
|
## Needs a human
|
|
|
|
**The free tier's storage ceiling is closer than the egress work suggested.** The Neon
|
|
database is 99.6 MB of a 512 MB allowance and `actions.context_snapshot` is essentially
|
|
all of it. See "Storage, which this plan did not cost" in `plan/13`. This is now the
|
|
most likely thing to break the deploy, ahead of anything on the read path.
|
|
|
|
---
|
|
|
|
## Pick up here
|
|
|
|
**`plan/13-memory-embedding-cost.md`, step 6 — infinite scroll upward in `Play.jsx`.**
|
|
|
|
Opening a finished adventure is comfortably the largest single read in the app — a turn
|
|
is now 122 kB, Insights 118 kB, the Memories drawer 24 kB. Measured on production
|
|
(2026-08-17), the largest real adventure is **607 actions and 589.5 kB in one
|
|
response**; the 426.7 kB the harness reports is a 200-action fixture whose actions are
|
|
about **twice as heavy as real ones** (994 B/action in production). So the fixture
|
|
overstates width and understates length — real stories get *longer* than it models,
|
|
which is the direction that hurts. The backend already has the
|
|
windowing primitives (`context/history.py`: `tail_range`, `slice_`, `count`), and
|
|
`GET /adventures/{id}/actions` exists. What is missing is a paged shape for it and a
|
|
`Play.jsx` that loads the newest turns and fetches older ones as the reader scrolls up.
|
|
|
|
Watch for: the story is a flat list today, and **the story tree replaces it**
|
|
(`plan/14-phase-story-tree.md`). Paging that reads by *position from the end* survives
|
|
that change; paging that assumes `Action.index` is a dense 0..n sequence does not.
|
|
|
|
After that: the tree itself. Its design is settled in `plan/14`; nothing about it has
|
|
been built.
|
|
|
|
---
|
|
|
|
## What happened on 2026-08-16
|
|
|
|
Four commits, all on the egress work that has to land before the tree.
|
|
|
|
### 1. A byte meter, in the repo this time — `7ee5cee`
|
|
|
|
`backend/tools/dbmeter.py` + `backend/tools/stress_session.py`.
|
|
|
|
```
|
|
cd backend
|
|
.venv/Scripts/python.exe -m tools.stress_session
|
|
.venv/Scripts/python.exe -m tools.stress_session --shapes turn --repeat 2
|
|
.venv/Scripts/python.exe -m tools.stress_session --no-embeddings # the old blind spot
|
|
```
|
|
|
|
It counts **bytes at the DBAPI cursor**, not queries — both egress blowouts this project
|
|
has had were one query fetching a column nobody read, and a statement count showed
|
|
nothing wrong in either. It drives a production-shaped synthetic adventure through the
|
|
real routes with only the LLM and the embedding endpoint faked.
|
|
|
|
**The memory bank is ON by default and that is the whole point.** The previous harness
|
|
ran without an embedding model configured; embedding providers are BYOK-only by
|
|
construction, so `retrieve_memories` returned early every time and the heaviest read in
|
|
a turn never happened. `--no-embeddings` reproduces that deliberately — the gap is 29x.
|
|
|
|
It calibrates against the two figures measured directly on production: 426.7 kB for a
|
|
200-action page load against 423 KB, and 3,258.7 kB for one turn against 3,153 kB.
|
|
It runs on SQLite, so treat absolutes as production-*shaped* and compare before/after.
|
|
|
|
### 2. Packed float32 embeddings — `c568648`
|
|
|
|
Migration 38 + `app/vectors.py`. A 1536-dimension vector as a JSON list is ~31 KB; the
|
|
same numbers as float32 are 6,144 bytes. **Not a precision trade** — the endpoints
|
|
compute in float32 and render that into JSON, so converting back is bit-exact. Nothing
|
|
re-embeds, no API calls.
|
|
|
|
The backfill is the one in `migrations.py` that cannot be portable SQL, so it comes
|
|
through Python, batched. Migration SQL can now be a `{dialect: sql}` map (BLOB vs BYTEA
|
|
have no common spelling).
|
|
|
|
### 3. Ranking the bank without reading the bank — `b7e53ae`
|
|
|
|
`retrieve_memories` walked `adventure.memories`, loading every row *with its vector*. It
|
|
now asks SQL which memories are in play (an id and a flag per row), ranks against
|
|
vectors held in process, and fetches text only for the top-K it picks.
|
|
|
|
Two more callers were doing the same thing, and the production SQL could not see either:
|
|
`_evict_over_capacity` walked the bank to count it, `_embed_pending` walked it to find
|
|
rows with no vector. A played turn cost **6.4 MB**, not the 3.2 the plan assumed.
|
|
|
|
| shape | before | cold | warm |
|
|
|---|---|---|---|
|
|
| one turn | 3,258.7 kB | 723.4 kB | **122.3 kB** |
|
|
| `run_post_turn` | 3,139.1 kB | 0.7 kB | 0.7 kB |
|
|
| Insights | 3,223.7 kB | 117.9 kB | 117.9 kB |
|
|
| Memories drawer | ~3.1 MB | 23.7 kB | 23.7 kB |
|
|
|
|
A played turn is turn + post-turn: **6.4 MB → 123 kB**, 52x.
|
|
|
|
Migrations 39/40 add `memories.embedded`, migration 41 drops the capacity default
|
|
200 → 80 for rows still on the old default.
|
|
|
|
---
|
|
|
|
## What happened on 2026-08-17
|
|
|
|
No new behaviour — a verification pass on what shipped the day before, because every
|
|
number in the section above had been measured on SQLite against a synthetic fixture.
|
|
Full write-up in `plan/13` under "Verified on production".
|
|
|
|
**It holds.** `schema_version` is 41 on the live Postgres with `embedding_blob bytea`
|
|
and `embedded boolean` present, so migration 38's dialect map is correct against a real
|
|
server. The backfill is complete (134/134). The packed vectors are **5.04x** smaller
|
|
than the JSON on real data — 30,971 → 6,144 bytes a memory, as predicted.
|
|
|
|
**SQLite was not lying.** `tools.stress_session` can now target Postgres via
|
|
`AIDND_STRESS_DATABASE_URL`, and every shape agrees within 0.5% — the warm turn is
|
|
121.1 kB on Postgres against 122.3 kB on SQLite, with `memories` down to 1.7 kB of it.
|
|
Run it against a **throwaway** database only; the harness writes, so it refuses any
|
|
target whose name does not contain `stress` or `scratch`.
|
|
|
|
**Two corrections came out of it**, both above: the page-load model has the wrong
|
|
shape (too heavy per action, far too short), and the storage ceiling was never costed.
|
|
|
|
## Things worth remembering
|
|
|
|
**The vector cache needs no invalidation callbacks, and that is why it is safe.** A
|
|
stored vector can only change through `memorybank.set_vector`, which drops that one
|
|
entry. Anything that *removes* a memory from play — eviction, deletion, pruning, an edit
|
|
clearing the vector — falls out of the catalogue query, and entries missing from the
|
|
catalogue are dropped on the next read. So there is no hook anyone can forget to call.
|
|
It is in-process and assumes one worker, which is what the deploy runs.
|
|
|
|
**Weigh new columns in bytes, not rows.** The comment on `Memory.embedding` said "fine
|
|
at bank sizes of a few hundred" and was wrong by the only measure that mattered: a few
|
|
hundred JSON vectors is ten megabytes, fetched fresh every turn.
|
|
|
|
**A deferred column needs a cheap flag beside it.** `memories.embedded` exists because
|
|
once the vector is deferred, every "is this embedded?" check becomes a 6 KB lazy load,
|
|
once per row down the Memories drawer. Exactly the same shape as `actions.variant_count`
|
|
beside `actions.variants`. Expect to need this for any future heavy column.
|
|
|
|
**Any egress measurement must run with an embedding model set.** This is the second
|
|
time that omission has hidden the biggest number in the room.
|
|
|
|
**Production has real users on it now. Measure it without reading it.** Counts,
|
|
`sum(octet_length(...))` and `pg_total_relation_size` answer every sizing question
|
|
asked so far, and none of them return anyone's story, memory text or email. When a
|
|
real Postgres is needed for a *write* path, create a throwaway database beside the real
|
|
one and drop it after — never point a harness at the production database.
|
|
|
|
**`octet_length` is the egress number, not the on-disk number.** Postgres TOAST
|
|
compresses big JSON — `context_snapshot` is 150.8 MB uncompressed but ~89 MB stored —
|
|
and decompresses before sending. Size reads with `octet_length`, size the storage bill
|
|
with `pg_total_relation_size`, and do not mix them up.
|
|
|
|
---
|
|
|
|
## Still open from `plan/13`
|
|
|
|
- **Step 6, infinite scroll upward** — the pick-up item above.
|
|
- **Query-count / byte assertions per endpoint**, extending `tests/test_egress.py`
|
|
against production-sized fixtures. `dbmeter` is importable from tests (`from tools
|
|
import dbmeter`) and was built with this in mind; nothing uses it there yet.
|
|
- **Explicit column projections on read paths**, so the next heavy column is opt-**in**.
|
|
Done for the memory paths, not as a general rule.
|
|
- **Drop `memories.embedding`** (the JSON column) in a follow-up migration. It is still
|
|
written by `set_vector` and read by nothing, kept so a rollback finds the vectors.
|
|
`tests/test_memory_retrieval.py` has a guard asserting nothing selects it. Measured
|
|
on production: dropping it reclaims 4.05 MB, 4% of the database.
|
|
- **`context_snapshot` and the 512 MB ceiling** — new, and now the biggest open item.
|
|
See the two sections named above. The egress case for leaving it in the database
|
|
still stands; the storage case does not.
|
|
|
|
Deliberately not taken: moving `context_snapshot` out of the database (~$0.02/mo, costs
|
|
nothing on reads now that it is deferred), and pgvector (breaks the SQLite dev parity
|
|
this codebase protects on purpose).
|
|
|
|
---
|
|
|
|
## Running things
|
|
|
|
```
|
|
cd backend
|
|
.venv/Scripts/python.exe -m pytest tests/ # 225 tests
|
|
.venv/Scripts/python.exe -m tools.stress_session # egress report (SQLite)
|
|
|
|
# Same harness against a real Postgres. The target must be a THROWAWAY database
|
|
# — this writes a synthetic adventure, and it refuses any name without
|
|
# 'stress'/'scratch' in it.
|
|
AIDND_STRESS_DATABASE_URL=postgresql://…/stress_scratch \
|
|
.venv/Scripts/python.exe -m tools.stress_session
|
|
```
|
|
|
|
On Windows the report's box-drawing characters crash the default cp1252 console;
|
|
prefix with `PYTHONIOENCODING=utf-8`.
|
|
|
|
Port 8000 is shared with the job-pipeline app, which will squat it and silently shadow
|
|
the AI-DnD API — free it before running the backend, or move the vite proxy.
|