Files
interactive-story/plan/STATUS.md
T
parththakkar106andClaude Opus 5 85b188977e Correct the production URL: the deploy is ai-dnd-1gmp, not ai-dnd
The previous commit recorded the service as suspended. It is not. Render
appended a suffix to the `ai-dnd` service name in render.yaml, and plain
ai-dnd.onrender.com is a different, suspended service whose 503 page reads
"suspended by its owner" -- indistinguishable from this deploy being down
unless you notice the host is wrong.

The real host answers {"ok":true} on /api/health in 0.6s, mints a guest,
returns an empty adventure list for that guest and serves the starter
scenarios. The authoritative link is the one docs/index.html points at, not
the service name in the blueprint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 12:43:06 +05:30

213 lines
10 KiB
Markdown

# Where things stand
Read this first when picking the project back up. Updated at the end of a working
session; the per-phase plan files hold the detail, this holds the thread.
**Last updated: 2026-08-17.**
---
## The live URL is not the one in render.yaml
**Production is `https://ai-dnd-1gmp.onrender.com`.** Render appended a suffix to the
`ai-dnd` service name in `render.yaml`, and plain `ai-dnd.onrender.com` belongs to a
different, suspended service that answers 503 with "suspended by its owner" — which is
easy to mistake for this deploy being down. The authoritative link is the one the
project page points at (`docs/index.html`), not the service name in the blueprint.
`GET /api/health` on the real host returns `{"ok":true}`.
---
## Needs a human
**The free tier's storage ceiling is closer than the egress work suggested.** The Neon
database is 99.6 MB of a 512 MB allowance and `actions.context_snapshot` is essentially
all of it. See "Storage, which this plan did not cost" in `plan/13`. This is now the
most likely thing to break the deploy, ahead of anything on the read path.
---
## Pick up here
**`plan/13-memory-embedding-cost.md`, step 6 — infinite scroll upward in `Play.jsx`.**
Opening a finished adventure is comfortably the largest single read in the app — a turn
is now 122 kB, Insights 118 kB, the Memories drawer 24 kB. Measured on production
(2026-08-17), the largest real adventure is **607 actions and 589.5 kB in one
response**; the 426.7 kB the harness reports is a 200-action fixture whose actions are
about **twice as heavy as real ones** (994 B/action in production). So the fixture
overstates width and understates length — real stories get *longer* than it models,
which is the direction that hurts. The backend already has the
windowing primitives (`context/history.py`: `tail_range`, `slice_`, `count`), and
`GET /adventures/{id}/actions` exists. What is missing is a paged shape for it and a
`Play.jsx` that loads the newest turns and fetches older ones as the reader scrolls up.
Watch for: the story is a flat list today, and **the story tree replaces it**
(`plan/14-phase-story-tree.md`). Paging that reads by *position from the end* survives
that change; paging that assumes `Action.index` is a dense 0..n sequence does not.
After that: the tree itself. Its design is settled in `plan/14`; nothing about it has
been built.
---
## What happened on 2026-08-16
Four commits, all on the egress work that has to land before the tree.
### 1. A byte meter, in the repo this time — `7ee5cee`
`backend/tools/dbmeter.py` + `backend/tools/stress_session.py`.
```
cd backend
.venv/Scripts/python.exe -m tools.stress_session
.venv/Scripts/python.exe -m tools.stress_session --shapes turn --repeat 2
.venv/Scripts/python.exe -m tools.stress_session --no-embeddings # the old blind spot
```
It counts **bytes at the DBAPI cursor**, not queries — both egress blowouts this project
has had were one query fetching a column nobody read, and a statement count showed
nothing wrong in either. It drives a production-shaped synthetic adventure through the
real routes with only the LLM and the embedding endpoint faked.
**The memory bank is ON by default and that is the whole point.** The previous harness
ran without an embedding model configured; embedding providers are BYOK-only by
construction, so `retrieve_memories` returned early every time and the heaviest read in
a turn never happened. `--no-embeddings` reproduces that deliberately — the gap is 29x.
It calibrates against the two figures measured directly on production: 426.7 kB for a
200-action page load against 423 KB, and 3,258.7 kB for one turn against 3,153 kB.
It runs on SQLite, so treat absolutes as production-*shaped* and compare before/after.
### 2. Packed float32 embeddings — `c568648`
Migration 38 + `app/vectors.py`. A 1536-dimension vector as a JSON list is ~31 KB; the
same numbers as float32 are 6,144 bytes. **Not a precision trade** — the endpoints
compute in float32 and render that into JSON, so converting back is bit-exact. Nothing
re-embeds, no API calls.
The backfill is the one in `migrations.py` that cannot be portable SQL, so it comes
through Python, batched. Migration SQL can now be a `{dialect: sql}` map (BLOB vs BYTEA
have no common spelling).
### 3. Ranking the bank without reading the bank — `b7e53ae`
`retrieve_memories` walked `adventure.memories`, loading every row *with its vector*. It
now asks SQL which memories are in play (an id and a flag per row), ranks against
vectors held in process, and fetches text only for the top-K it picks.
Two more callers were doing the same thing, and the production SQL could not see either:
`_evict_over_capacity` walked the bank to count it, `_embed_pending` walked it to find
rows with no vector. A played turn cost **6.4 MB**, not the 3.2 the plan assumed.
| shape | before | cold | warm |
|---|---|---|---|
| one turn | 3,258.7 kB | 723.4 kB | **122.3 kB** |
| `run_post_turn` | 3,139.1 kB | 0.7 kB | 0.7 kB |
| Insights | 3,223.7 kB | 117.9 kB | 117.9 kB |
| Memories drawer | ~3.1 MB | 23.7 kB | 23.7 kB |
A played turn is turn + post-turn: **6.4 MB → 123 kB**, 52x.
Migrations 39/40 add `memories.embedded`, migration 41 drops the capacity default
200 → 80 for rows still on the old default.
---
## What happened on 2026-08-17
No new behaviour — a verification pass on what shipped the day before, because every
number in the section above had been measured on SQLite against a synthetic fixture.
Full write-up in `plan/13` under "Verified on production".
**It holds.** `schema_version` is 41 on the live Postgres with `embedding_blob bytea`
and `embedded boolean` present, so migration 38's dialect map is correct against a real
server. The backfill is complete (134/134). The packed vectors are **5.04x** smaller
than the JSON on real data — 30,971 → 6,144 bytes a memory, as predicted.
**SQLite was not lying.** `tools.stress_session` can now target Postgres via
`AIDND_STRESS_DATABASE_URL`, and every shape agrees within 0.5% — the warm turn is
121.1 kB on Postgres against 122.3 kB on SQLite, with `memories` down to 1.7 kB of it.
Run it against a **throwaway** database only; the harness writes, so it refuses any
target whose name does not contain `stress` or `scratch`.
**Two corrections came out of it**, both above: the page-load model has the wrong
shape (too heavy per action, far too short), and the storage ceiling was never costed.
## Things worth remembering
**The vector cache needs no invalidation callbacks, and that is why it is safe.** A
stored vector can only change through `memorybank.set_vector`, which drops that one
entry. Anything that *removes* a memory from play — eviction, deletion, pruning, an edit
clearing the vector — falls out of the catalogue query, and entries missing from the
catalogue are dropped on the next read. So there is no hook anyone can forget to call.
It is in-process and assumes one worker, which is what the deploy runs.
**Weigh new columns in bytes, not rows.** The comment on `Memory.embedding` said "fine
at bank sizes of a few hundred" and was wrong by the only measure that mattered: a few
hundred JSON vectors is ten megabytes, fetched fresh every turn.
**A deferred column needs a cheap flag beside it.** `memories.embedded` exists because
once the vector is deferred, every "is this embedded?" check becomes a 6 KB lazy load,
once per row down the Memories drawer. Exactly the same shape as `actions.variant_count`
beside `actions.variants`. Expect to need this for any future heavy column.
**Any egress measurement must run with an embedding model set.** This is the second
time that omission has hidden the biggest number in the room.
**Production has real users on it now. Measure it without reading it.** Counts,
`sum(octet_length(...))` and `pg_total_relation_size` answer every sizing question
asked so far, and none of them return anyone's story, memory text or email. When a
real Postgres is needed for a *write* path, create a throwaway database beside the real
one and drop it after — never point a harness at the production database.
**`octet_length` is the egress number, not the on-disk number.** Postgres TOAST
compresses big JSON — `context_snapshot` is 150.8 MB uncompressed but ~89 MB stored —
and decompresses before sending. Size reads with `octet_length`, size the storage bill
with `pg_total_relation_size`, and do not mix them up.
---
## Still open from `plan/13`
- **Step 6, infinite scroll upward** — the pick-up item above.
- **Query-count / byte assertions per endpoint**, extending `tests/test_egress.py`
against production-sized fixtures. `dbmeter` is importable from tests (`from tools
import dbmeter`) and was built with this in mind; nothing uses it there yet.
- **Explicit column projections on read paths**, so the next heavy column is opt-**in**.
Done for the memory paths, not as a general rule.
- **Drop `memories.embedding`** (the JSON column) in a follow-up migration. It is still
written by `set_vector` and read by nothing, kept so a rollback finds the vectors.
`tests/test_memory_retrieval.py` has a guard asserting nothing selects it. Measured
on production: dropping it reclaims 4.05 MB, 4% of the database.
- **`context_snapshot` and the 512 MB ceiling** — new, and now the biggest open item.
See the two sections named above. The egress case for leaving it in the database
still stands; the storage case does not.
Deliberately not taken: moving `context_snapshot` out of the database (~$0.02/mo, costs
nothing on reads now that it is deferred), and pgvector (breaks the SQLite dev parity
this codebase protects on purpose).
---
## Running things
```
cd backend
.venv/Scripts/python.exe -m pytest tests/ # 225 tests
.venv/Scripts/python.exe -m tools.stress_session # egress report (SQLite)
# Same harness against a real Postgres. The target must be a THROWAWAY database
# — this writes a synthetic adventure, and it refuses any name without
# 'stress'/'scratch' in it.
AIDND_STRESS_DATABASE_URL=postgresql://…/stress_scratch \
.venv/Scripts/python.exe -m tools.stress_session
```
On Windows the report's box-drawing characters crash the default cp1252 console;
prefix with `PYTHONIOENCODING=utf-8`.
Port 8000 is shared with the job-pipeline app, which will squat it and silently shadow
the AI-DnD API — free it before running the backend, or move the vite proxy.