Files
interactive-story/backend/app/cleanup.py
T
parththakkar106andClaude Opus 5 041f9e25f3 Count the visits, and say whether anyone got anywhere
A hosted demo raises a question a local app never does: is anyone using it,
and do they reach the part that matters? `/analytics` answers it — visitors,
pages, referrers, countries, devices, which shared scenarios get played, turns
and demo-key spend, API and turn errors, and a funnel from visited to played a
turn to signed up.

Not a third-party script, for reasons specific to this one. The CSP allows
`script-src 'self'`, so a tracker means loosening it; adblockers eat the
popular ones, which silently biases exactly the technical audience this
project gets shown to; and none of them can see the measurement that actually
matters here, which is a turn, not a pageview.

**A visit is a write and never a read.** After the 189x egress fix it would be
perverse to add a feature that reads rows per request, so counts accumulate in
a process-local dict and flush every 60s as UPSERTs. Storage is a generic
`(day, metric, label) -> hits` counter, so measuring something new later costs
a constant rather than a migration, plus one row per visitor per day for the
funnel flags. Every dashboard query is a GROUP BY returning tens of rows
however much traffic sits behind it; a month reads back in a few kilobytes.
The buffer's cost is that a hard restart can lose up to a minute — the flusher
also runs on shutdown, and a tier that sleeps when idle sleeps on an empty
buffer anyway.

**The counters are anonymous; the access log beside them is not, on purpose.**
A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars —
one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined
back to `users`, and keyed, so no client can compute one. Story content never
reaches that module, and the only content it ever names is a seeded public
scenario's title; a player's own titles are theirs. `accesslog.py` is the
identifying half and is a separate module writing a separate table so that
separation is a property of the code rather than a convention: `access_events`
records sessions, sign-ins, registrations and failed attempts with address,
email and device, read on a second tab of the same page behind the same gate.

Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An
unmetered tester is not automatically someone who should see the traffic. The
route 404s and the nav link is absent for everyone else, the same treatment
AI Chat gets; unset in a hosted deploy means nobody sees it, including me.

Three things came out of building it that a test would not have suggested.

**A failed turn is an HTTP 200 with a bad ending.** The status-code middleware
cannot see one, so a demo whose model had started refusing every request would
look perfectly healthy from outside. All five SSE error paths in
`_generate_turn` now go through a `turn_error()` helper that counts on the way
out. Error buckets elsewhere are labelled by the matched route template rather
than the requested path — one bucket per endpoint instead of one per adventure
id, and, the reason it isn't merely tidier, an unmatched path is entirely
attacker-chosen, so labelling by it would let anyone mint rows.

**The funnel counts people, not clicks.** A player who starts six adventures
is one person who started an adventure. That is the whole reason the
per-visitor-day table exists; its flags only ever turn on, and `is_new` is
settled by the first write of a visitor's first day.

**The tests run on SQLite and production is Neon.** A flush that raises is
caught and logged, so a dialect mistake in the UPSERTs would have stayed
invisible until the dashboard quietly never filled.
`test_the_upserts_compile_for_postgres` compiles both statements against the
Postgres dialect without connecting to one.

Two things this leans on elsewhere. `limits._client_ip` is now public
`client_ip`: the access log needs the same answer, and two functions both
deciding which hop is the caller's is how one of them ends up trusting a
header it shouldn't. And the cleanup sweeper now starts if *either* job has
work — a deployment can keep every guest forever and still want its
visitor-day rows aged out.

No migration. Both tables are new and `bootstrap()` calls `create_all` on
existing databases too, the route `branches` took in Phase 14, so
`LATEST_VERSION` is still 64.

497 tests green, frontend lint and build clean, driven by hand against a
synthetic 90-day fixture at 1568px. The narrow-screen layout follows the
existing 720px block but is unverified: `resize_window` is ignored on a
maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized
iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run
through the test names themselves, leaving `testclient_ip_*` — still collected
by pytest, which is why it passed unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-22 16:24:42 +05:30

175 lines
6.9 KiB
Python

"""Retention policy for throwaway guest accounts.
In multi-user mode every first visit mints a `users` row (GET /api/auth/me),
so a public demo accumulates one account per curious visitor — most of whom
never come back, each leaving behind whatever scenarios, adventures, actions
and memories they generated. This drops guests that have gone quiet for
AIDND_GUEST_RETENTION_DAYS (default 5) along with everything they made.
Why this is safe to run unattended:
- Only rows with `is_guest` AND `email IS NULL` are ever touched, and both
clauses are checked rather than either alone. Registering upgrades the row
in place (is_guest -> False), so a guest who signs up keeps everything;
local mode's implicit single user is also is_guest=False.
- Idle time is COALESCE(last_seen_at, created_at). `auth._touch` only writes
last_seen_at once an hour, and a guest minted by /auth/me has NULL until
its *second* request, so created_at is the honest floor for a brand-new
visitor — without the coalesce those rows look infinitely old.
- Nothing a guest owns is reachable by anyone else: `is_public` is an
output-only field (see schemas.ScenarioBase), so only seeded scenarios —
which have user_id NULL and are therefore outside this filter entirely —
are shared. Deleting a guest can't take content away from another user.
Why one Core DELETE instead of an ORM cascade: `db.delete(user)` would SELECT
every adventure, action, memory and story card into Python purely to delete
them, which on Neon is exactly the egress pattern that has already cost this
project once. Every foreign key from users downwards is ON DELETE CASCADE
(users -> scenarios/adventures/scripts/settings -> actions/memories/cards), so
the database does the whole graph in one statement and ships back a row count.
No index is added for the scan: the sweep runs a handful of times a day
against a table with at most a few thousand rows, which is not worth a
migration and the schema surface that comes with it.
"""
import asyncio
import logging
import os
from datetime import datetime, timedelta
from sqlalchemy import delete, func
from sqlalchemy.orm import Session
from starlette.concurrency import run_in_threadpool
from . import analytics, auth, models
from .database import SessionLocal
logger = logging.getLogger(__name__)
def _int_env(name: str, default: int) -> int:
try:
return int(os.environ.get(name, "").strip() or default)
except ValueError:
logger.warning("%s is not an integer; using %d.", name, default)
return default
# Days of inactivity before a guest account is dropped. 0 or less disables the
# policy entirely, for a deployment that would rather keep everything.
RETENTION_DAYS = _int_env("AIDND_GUEST_RETENTION_DAYS", 5)
# How often a long-lived process re-checks. Hours, not minutes: nothing here is
# time-critical, and on Render's free tier the service sleeps and cold-starts
# often enough that the startup sweep does most of the work by itself.
SWEEP_INTERVAL_SECONDS = _int_env("AIDND_CLEANUP_INTERVAL_HOURS", 6) * 3600
def enabled() -> bool:
"""Guests only exist in multi-user mode, so local runs skip the sweep
rather than pointing a DELETE at a database that has nothing to collect."""
return auth.MULTI_USER and RETENTION_DAYS > 0
def anything_to_sweep() -> bool:
"""Whether the periodic task is worth starting at all. The two jobs it runs
are independent: a deployment can keep every guest forever and still want
its analytics rows aged out, and vice versa."""
return enabled() or analytics.RETENTION_DAYS > 0
def delete_stale_guests(db: Session, *, now: datetime | None = None) -> int:
"""Delete guests idle for RETENTION_DAYS or more. Returns the row count.
The caller owns error handling; `sweep` is the safe wrapper.
"""
if RETENTION_DAYS <= 0:
return 0
# Stored timestamps are naive UTC on both backends (SQLite drops tzinfo;
# Postgres columns are TIMESTAMP WITHOUT TIME ZONE with the session pinned
# to UTC in database.py). Match that exactly so the comparison can't hinge
# on how a given dialect renders an aware value.
reference = now or models.utcnow()
cutoff = reference.replace(tzinfo=None) - timedelta(days=RETENTION_DAYS)
stmt = (
delete(models.User)
.where(
models.User.is_guest.is_(True),
models.User.email.is_(None),
func.coalesce(models.User.last_seen_at, models.User.created_at) < cutoff,
)
# Without this, "auto" can't evaluate coalesce in Python and falls back
# to fetching every matching primary key first — a second round trip
# for nothing, since this session holds no User objects to synchronize.
.execution_options(synchronize_session=False)
)
removed = db.execute(stmt).rowcount or 0
db.commit()
return removed
def sweep() -> int:
"""One pass, with its own session. Never raises: a failed cleanup must not
be able to take the app down (same rule as seeding). Returns the guest
count, which is the number worth logging about."""
if not anything_to_sweep():
return 0
db = SessionLocal()
try:
# Ages out the per-visitor analytics rows, on its own terms: it is not
# about guests, and it must still happen on a deployment that has
# chosen to keep every account it ever minted.
aged = analytics.purge_old_visitor_days(db)
if aged:
logger.info("Aged out %d analytics visitor-day row(s).", aged)
removed = delete_stale_guests(db) if enabled() else 0
if removed:
logger.info(
"Cleaned up %d guest account(s) idle for %d+ days.",
removed,
RETENTION_DAYS,
)
return removed
except Exception:
db.rollback()
logger.exception("Guest cleanup failed; continuing without it.")
return 0
finally:
db.close()
async def _sweep_loop() -> None:
while True:
# Blocking DB work: keep it off the event loop, which is also serving
# SSE turn streams.
await run_in_threadpool(sweep)
await asyncio.sleep(SWEEP_INTERVAL_SECONDS)
def start_sweeper() -> asyncio.Task | None:
"""Kick off the periodic sweep; None when there is nothing to sweep."""
if not enabled():
logger.info("Guest cleanup disabled (multi_user=%s, retention_days=%d).",
auth.MULTI_USER, RETENTION_DAYS)
else:
logger.info(
"Guest cleanup on: deleting guests idle %d+ days, every %d hour(s).",
RETENTION_DAYS,
SWEEP_INTERVAL_SECONDS // 3600,
)
if not anything_to_sweep():
return None
return asyncio.create_task(_sweep_loop())
async def stop_sweeper(task: asyncio.Task | None) -> None:
if task is None:
return
task.cancel()
try:
await task
except asyncio.CancelledError:
pass