- render.yaml: single Docker web service (SPA + API same-origin), free tier, Neon Postgres via AIDND_DATABASE_URL, generated AIDND_SECRET_KEY, /api/health check, us-east region, auto-deploy on main. - README: "Deploy (Render)" section. - Record verified Postgres path (real Neon, PG 18.4) and Phase 10 decisions (Neon, free tier, cloud Docker build) in plan/. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
5.4 KiB
Phase 9 — Production hardening
Goal: make the app safe and stable to expose to strangers on the internet: config via environment, resource limits on everything user-controlled, and a single-service production build.
Decisions
| Question | Answer |
|---|---|
| Database | Decide at start of this phase. SQLite on a persistent disk (zero code change, but Render disks require the ~$7/mo starter tier) vs Postgres (free/cheap managed options, better resume talking point, needs SQLAlchemy URL + migration tweaks). Revisit with current Render pricing. |
Ask before implementing: the database choice above, and target monthly budget (drives Render tier: free tier sleeps after idle + has no persistent disk).
Configuration
- All config via env vars with sane local defaults:
DATABASE_URL,SECRET_KEY(sessions + API-key encryption),MULTI_USER,CORS_ORIGINS, demo-key vars (Phase 8), port/host. Document each inbackend/.env.example. - Fail fast on missing
SECRET_KEYwhenMULTI_USER=true.
Abuse & resource limits
- quickjs limits: per-execution time limit and memory limit on the scripting engine
(
scripting/engine.py) — user-submitted JS must not be able to hang or OOM the server. - Rate limiting on expensive endpoints (turn generation, script run, auth) — per-user and
per-IP (e.g.
slowapi). - Request size limits (script source length, memory/story-card text lengths, action text).
- Cap per-user row counts (adventures, scenarios, scripts, story cards) with friendly errors.
- Audit debug router (
routers/debug.py) and/docs: admin-only or disabled whenMULTI_USER=true— debug log may contain other users' prompts.
Production serving
- Single service: FastAPI serves the built SPA (fallback already exists) — verify the Docker
image from Phase 7 is production-ready (no
--reload, multiple workers or async-safe single worker; check SQLite + multiple workers interaction before choosing). - CORS locked to the deployed origin (moot if same-origin single service — verify).
- Security headers middleware; cookies
Secure+SameSite. - Streaming (SSE) works behind Render's proxy — verify no buffering issues.
- Structured logging; scrub API keys from all logs and the debug page.
Database (after decision)
- If Postgres: swap
DATABASE_URL, verify JSON-blob columns (embeddings) andmigrations.pywork; test full play loop. - If SQLite-on-disk: confirm WAL mode + single-worker (or serialized writes) is acceptable.
- Backup story: platform DB backups (Postgres) or a scheduled dump of the disk (SQLite).
Exit criteria
Running the production Docker image locally with MULTI_USER=true: a hostile user cannot hang
the server with a while(true) script, cannot see another user's data or the debug log, gets
rate-limited instead of burning the demo key, and the app streams turns normally the whole time.
Verified 2026-07-07 (uvicorn, MULTI_USER=1, fresh SQLite DB, curl)
- Fail-fast secret:
import app.mainwithMULTI_USER=1and noAIDND_SECRET_KEYraises the RuntimeError as designed (won't boot). while(true)script:POST /api/scripts/{id}/teston aninput_jsinfinite loop returnsInternalError: interrupted(engine time limit) — server stays responsive afterward.- Cross-user isolation: guest B sees
[]for scripts, gets 404 on guest A's script id; guest A keeps its own row. No leakage. - Debug log:
GET /api/debug/requests→ 403 in multi-user mode. - Rate limiting: 12 rapid
POST /api/auth/register→ 429 after the 10th (auth scope, 10/300s). - Body size: 3 MB body to
POST /api/scenarios→ 413 (limit 2 MB) via BodySizeLimitMiddleware. - Security headers: CSP,
x-frame-options: DENY,x-content-type-options: nosniff,referrer-policy: same-originon every response — including the SSE stream. - Docs disabled: Swagger UI and OpenAPI schema not served (
/docs,/openapi.jsonfall through to the SPAindex.html; noswagger-ui, no API schema exposed). - SSE streaming:
POST /api/adventures/{id}/actionsstreamstext/event-streamwithx-accel-buffering: no, chunked, incremental events — the pure-ASGI middlewares don't buffer. (No LLM key configured here, so it streams the "No model configured" error event; a live provider turn through this path was verified end-to-end in Phase 8.)
All Phase 9 exit criteria met.
Postgres path verified 2026-07-07 (real Neon, PostgreSQL 18.4)
Pointed the backend at a live Neon DATABASE_URL (pooled, sslmode=require&channel_binding=require):
- Driver/URL:
postgres:///postgresql://normalized topostgresql+psycopg://; raw psycopg3 connect succeeds. - Bootstrap: fresh DB →
create_allbuilds all 11 tables and stampsschema_version = 23(=LATEST_VERSION); the SQLite-only migrations 2–23 are correctly skipped on a fresh DB. - Embedding column:
memories.embeddingis Postgresjson; a float list round-trips intact ([0.1, 0.2, 0.3, -0.4]back as alist). - ORM CRUD: user → scenario → adventure → memory create/read/delete with FK cascades works.
- HTTP round-trip (multi-user, TestClient): guest bootstrap → register → me → create scenario → list → settings → delete, all 200/201/204 against Neon.
Remaining Phase 10 step: the Docker production image specifically (build + run the container).