A hosted demo raises a question a local app never does: is anyone using it, and do they reach the part that matters? `/analytics` answers it — visitors, pages, referrers, countries, devices, which shared scenarios get played, turns and demo-key spend, API and turn errors, and a funnel from visited to played a turn to signed up. Not a third-party script, for reasons specific to this one. The CSP allows `script-src 'self'`, so a tracker means loosening it; adblockers eat the popular ones, which silently biases exactly the technical audience this project gets shown to; and none of them can see the measurement that actually matters here, which is a turn, not a pageview. **A visit is a write and never a read.** After the 189x egress fix it would be perverse to add a feature that reads rows per request, so counts accumulate in a process-local dict and flush every 60s as UPSERTs. Storage is a generic `(day, metric, label) -> hits` counter, so measuring something new later costs a constant rather than a migration, plus one row per visitor per day for the funnel flags. Every dashboard query is a GROUP BY returning tens of rows however much traffic sits behind it; a month reads back in a few kilobytes. The buffer's cost is that a hard restart can lose up to a minute — the flusher also runs on shutdown, and a tier that sleeps when idle sleeps on an empty buffer anyway. **The counters are anonymous; the access log beside them is not, on purpose.** A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars — one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined back to `users`, and keyed, so no client can compute one. Story content never reaches that module, and the only content it ever names is a seeded public scenario's title; a player's own titles are theirs. `accesslog.py` is the identifying half and is a separate module writing a separate table so that separation is a property of the code rather than a convention: `access_events` records sessions, sign-ins, registrations and failed attempts with address, email and device, read on a second tab of the same page behind the same gate. Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An unmetered tester is not automatically someone who should see the traffic. The route 404s and the nav link is absent for everyone else, the same treatment AI Chat gets; unset in a hosted deploy means nobody sees it, including me. Three things came out of building it that a test would not have suggested. **A failed turn is an HTTP 200 with a bad ending.** The status-code middleware cannot see one, so a demo whose model had started refusing every request would look perfectly healthy from outside. All five SSE error paths in `_generate_turn` now go through a `turn_error()` helper that counts on the way out. Error buckets elsewhere are labelled by the matched route template rather than the requested path — one bucket per endpoint instead of one per adventure id, and, the reason it isn't merely tidier, an unmatched path is entirely attacker-chosen, so labelling by it would let anyone mint rows. **The funnel counts people, not clicks.** A player who starts six adventures is one person who started an adventure. That is the whole reason the per-visitor-day table exists; its flags only ever turn on, and `is_new` is settled by the first write of a visitor's first day. **The tests run on SQLite and production is Neon.** A flush that raises is caught and logged, so a dialect mistake in the UPSERTs would have stayed invisible until the dashboard quietly never filled. `test_the_upserts_compile_for_postgres` compiles both statements against the Postgres dialect without connecting to one. Two things this leans on elsewhere. `limits._client_ip` is now public `client_ip`: the access log needs the same answer, and two functions both deciding which hop is the caller's is how one of them ends up trusting a header it shouldn't. And the cleanup sweeper now starts if *either* job has work — a deployment can keep every guest forever and still want its visitor-day rows aged out. No migration. Both tables are new and `bootstrap()` calls `create_all` on existing databases too, the route `branches` took in Phase 14, so `LATEST_VERSION` is still 64. 497 tests green, frontend lint and build clean, driven by hand against a synthetic 90-day fixture at 1568px. The narrow-screen layout follows the existing 720px block but is unverified: `resize_window` is ignored on a maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run through the test names themselves, leaving `testclient_ip_*` — still collected by pytest, which is why it passed unnoticed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
175 lines
6.9 KiB
Python
175 lines
6.9 KiB
Python
"""Retention policy for throwaway guest accounts.
|
|
|
|
In multi-user mode every first visit mints a `users` row (GET /api/auth/me),
|
|
so a public demo accumulates one account per curious visitor — most of whom
|
|
never come back, each leaving behind whatever scenarios, adventures, actions
|
|
and memories they generated. This drops guests that have gone quiet for
|
|
AIDND_GUEST_RETENTION_DAYS (default 5) along with everything they made.
|
|
|
|
Why this is safe to run unattended:
|
|
|
|
- Only rows with `is_guest` AND `email IS NULL` are ever touched, and both
|
|
clauses are checked rather than either alone. Registering upgrades the row
|
|
in place (is_guest -> False), so a guest who signs up keeps everything;
|
|
local mode's implicit single user is also is_guest=False.
|
|
- Idle time is COALESCE(last_seen_at, created_at). `auth._touch` only writes
|
|
last_seen_at once an hour, and a guest minted by /auth/me has NULL until
|
|
its *second* request, so created_at is the honest floor for a brand-new
|
|
visitor — without the coalesce those rows look infinitely old.
|
|
- Nothing a guest owns is reachable by anyone else: `is_public` is an
|
|
output-only field (see schemas.ScenarioBase), so only seeded scenarios —
|
|
which have user_id NULL and are therefore outside this filter entirely —
|
|
are shared. Deleting a guest can't take content away from another user.
|
|
|
|
Why one Core DELETE instead of an ORM cascade: `db.delete(user)` would SELECT
|
|
every adventure, action, memory and story card into Python purely to delete
|
|
them, which on Neon is exactly the egress pattern that has already cost this
|
|
project once. Every foreign key from users downwards is ON DELETE CASCADE
|
|
(users -> scenarios/adventures/scripts/settings -> actions/memories/cards), so
|
|
the database does the whole graph in one statement and ships back a row count.
|
|
|
|
No index is added for the scan: the sweep runs a handful of times a day
|
|
against a table with at most a few thousand rows, which is not worth a
|
|
migration and the schema surface that comes with it.
|
|
"""
|
|
|
|
import asyncio
|
|
import logging
|
|
import os
|
|
from datetime import datetime, timedelta
|
|
|
|
from sqlalchemy import delete, func
|
|
from sqlalchemy.orm import Session
|
|
from starlette.concurrency import run_in_threadpool
|
|
|
|
from . import analytics, auth, models
|
|
from .database import SessionLocal
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
def _int_env(name: str, default: int) -> int:
|
|
try:
|
|
return int(os.environ.get(name, "").strip() or default)
|
|
except ValueError:
|
|
logger.warning("%s is not an integer; using %d.", name, default)
|
|
return default
|
|
|
|
|
|
# Days of inactivity before a guest account is dropped. 0 or less disables the
|
|
# policy entirely, for a deployment that would rather keep everything.
|
|
RETENTION_DAYS = _int_env("AIDND_GUEST_RETENTION_DAYS", 5)
|
|
|
|
# How often a long-lived process re-checks. Hours, not minutes: nothing here is
|
|
# time-critical, and on Render's free tier the service sleeps and cold-starts
|
|
# often enough that the startup sweep does most of the work by itself.
|
|
SWEEP_INTERVAL_SECONDS = _int_env("AIDND_CLEANUP_INTERVAL_HOURS", 6) * 3600
|
|
|
|
|
|
def enabled() -> bool:
|
|
"""Guests only exist in multi-user mode, so local runs skip the sweep
|
|
rather than pointing a DELETE at a database that has nothing to collect."""
|
|
return auth.MULTI_USER and RETENTION_DAYS > 0
|
|
|
|
|
|
def anything_to_sweep() -> bool:
|
|
"""Whether the periodic task is worth starting at all. The two jobs it runs
|
|
are independent: a deployment can keep every guest forever and still want
|
|
its analytics rows aged out, and vice versa."""
|
|
return enabled() or analytics.RETENTION_DAYS > 0
|
|
|
|
|
|
def delete_stale_guests(db: Session, *, now: datetime | None = None) -> int:
|
|
"""Delete guests idle for RETENTION_DAYS or more. Returns the row count.
|
|
|
|
The caller owns error handling; `sweep` is the safe wrapper.
|
|
"""
|
|
if RETENTION_DAYS <= 0:
|
|
return 0
|
|
# Stored timestamps are naive UTC on both backends (SQLite drops tzinfo;
|
|
# Postgres columns are TIMESTAMP WITHOUT TIME ZONE with the session pinned
|
|
# to UTC in database.py). Match that exactly so the comparison can't hinge
|
|
# on how a given dialect renders an aware value.
|
|
reference = now or models.utcnow()
|
|
cutoff = reference.replace(tzinfo=None) - timedelta(days=RETENTION_DAYS)
|
|
|
|
stmt = (
|
|
delete(models.User)
|
|
.where(
|
|
models.User.is_guest.is_(True),
|
|
models.User.email.is_(None),
|
|
func.coalesce(models.User.last_seen_at, models.User.created_at) < cutoff,
|
|
)
|
|
# Without this, "auto" can't evaluate coalesce in Python and falls back
|
|
# to fetching every matching primary key first — a second round trip
|
|
# for nothing, since this session holds no User objects to synchronize.
|
|
.execution_options(synchronize_session=False)
|
|
)
|
|
removed = db.execute(stmt).rowcount or 0
|
|
db.commit()
|
|
return removed
|
|
|
|
|
|
def sweep() -> int:
|
|
"""One pass, with its own session. Never raises: a failed cleanup must not
|
|
be able to take the app down (same rule as seeding). Returns the guest
|
|
count, which is the number worth logging about."""
|
|
if not anything_to_sweep():
|
|
return 0
|
|
db = SessionLocal()
|
|
try:
|
|
# Ages out the per-visitor analytics rows, on its own terms: it is not
|
|
# about guests, and it must still happen on a deployment that has
|
|
# chosen to keep every account it ever minted.
|
|
aged = analytics.purge_old_visitor_days(db)
|
|
if aged:
|
|
logger.info("Aged out %d analytics visitor-day row(s).", aged)
|
|
removed = delete_stale_guests(db) if enabled() else 0
|
|
if removed:
|
|
logger.info(
|
|
"Cleaned up %d guest account(s) idle for %d+ days.",
|
|
removed,
|
|
RETENTION_DAYS,
|
|
)
|
|
return removed
|
|
except Exception:
|
|
db.rollback()
|
|
logger.exception("Guest cleanup failed; continuing without it.")
|
|
return 0
|
|
finally:
|
|
db.close()
|
|
|
|
|
|
async def _sweep_loop() -> None:
|
|
while True:
|
|
# Blocking DB work: keep it off the event loop, which is also serving
|
|
# SSE turn streams.
|
|
await run_in_threadpool(sweep)
|
|
await asyncio.sleep(SWEEP_INTERVAL_SECONDS)
|
|
|
|
|
|
def start_sweeper() -> asyncio.Task | None:
|
|
"""Kick off the periodic sweep; None when there is nothing to sweep."""
|
|
if not enabled():
|
|
logger.info("Guest cleanup disabled (multi_user=%s, retention_days=%d).",
|
|
auth.MULTI_USER, RETENTION_DAYS)
|
|
else:
|
|
logger.info(
|
|
"Guest cleanup on: deleting guests idle %d+ days, every %d hour(s).",
|
|
RETENTION_DAYS,
|
|
SWEEP_INTERVAL_SECONDS // 3600,
|
|
)
|
|
if not anything_to_sweep():
|
|
return None
|
|
return asyncio.create_task(_sweep_loop())
|
|
|
|
|
|
async def stop_sweeper(task: asyncio.Task | None) -> None:
|
|
if task is None:
|
|
return
|
|
task.cancel()
|
|
try:
|
|
await task
|
|
except asyncio.CancelledError:
|
|
pass
|