Count the visits, and say whether anyone got anywhere

A hosted demo raises a question a local app never does: is anyone using it,
and do they reach the part that matters? `/analytics` answers it — visitors,
pages, referrers, countries, devices, which shared scenarios get played, turns
and demo-key spend, API and turn errors, and a funnel from visited to played a
turn to signed up.

Not a third-party script, for reasons specific to this one. The CSP allows
`script-src 'self'`, so a tracker means loosening it; adblockers eat the
popular ones, which silently biases exactly the technical audience this
project gets shown to; and none of them can see the measurement that actually
matters here, which is a turn, not a pageview.

**A visit is a write and never a read.** After the 189x egress fix it would be
perverse to add a feature that reads rows per request, so counts accumulate in
a process-local dict and flush every 60s as UPSERTs. Storage is a generic
`(day, metric, label) -> hits` counter, so measuring something new later costs
a constant rather than a migration, plus one row per visitor per day for the
funnel flags. Every dashboard query is a GROUP BY returning tens of rows
however much traffic sits behind it; a month reads back in a few kilobytes.
The buffer's cost is that a hard restart can lose up to a minute — the flusher
also runs on shutdown, and a tier that sleeps when idle sleeps on an empty
buffer anyway.

**The counters are anonymous; the access log beside them is not, on purpose.**
A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars —
one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined
back to `users`, and keyed, so no client can compute one. Story content never
reaches that module, and the only content it ever names is a seeded public
scenario's title; a player's own titles are theirs. `accesslog.py` is the
identifying half and is a separate module writing a separate table so that
separation is a property of the code rather than a convention: `access_events`
records sessions, sign-ins, registrations and failed attempts with address,
email and device, read on a second tab of the same page behind the same gate.

Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An
unmetered tester is not automatically someone who should see the traffic. The
route 404s and the nav link is absent for everyone else, the same treatment
AI Chat gets; unset in a hosted deploy means nobody sees it, including me.

Three things came out of building it that a test would not have suggested.

**A failed turn is an HTTP 200 with a bad ending.** The status-code middleware
cannot see one, so a demo whose model had started refusing every request would
look perfectly healthy from outside. All five SSE error paths in
`_generate_turn` now go through a `turn_error()` helper that counts on the way
out. Error buckets elsewhere are labelled by the matched route template rather
than the requested path — one bucket per endpoint instead of one per adventure
id, and, the reason it isn't merely tidier, an unmatched path is entirely
attacker-chosen, so labelling by it would let anyone mint rows.

**The funnel counts people, not clicks.** A player who starts six adventures
is one person who started an adventure. That is the whole reason the
per-visitor-day table exists; its flags only ever turn on, and `is_new` is
settled by the first write of a visitor's first day.

**The tests run on SQLite and production is Neon.** A flush that raises is
caught and logged, so a dialect mistake in the UPSERTs would have stayed
invisible until the dashboard quietly never filled.
`test_the_upserts_compile_for_postgres` compiles both statements against the
Postgres dialect without connecting to one.

Two things this leans on elsewhere. `limits._client_ip` is now public
`client_ip`: the access log needs the same answer, and two functions both
deciding which hop is the caller's is how one of them ends up trusting a
header it shouldn't. And the cleanup sweeper now starts if *either* job has
work — a deployment can keep every guest forever and still want its
visitor-day rows aged out.

No migration. Both tables are new and `bootstrap()` calls `create_all` on
existing databases too, the route `branches` took in Phase 14, so
`LATEST_VERSION` is still 64.

497 tests green, frontend lint and build clean, driven by hand against a
synthetic 90-day fixture at 1568px. The narrow-screen layout follows the
existing 720px block but is unverified: `resize_window` is ignored on a
maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized
iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run
through the test names themselves, leaving `testclient_ip_*` — still collected
by pytest, which is why it passed unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
This commit is contained in:
parththakkar106
2026-08-22 16:24:42 +05:30
co-authored by Claude Opus 5
parent 3b9e6b3d50
commit 041f9e25f3
24 changed files with 2698 additions and 45 deletions
+44 -2
View File
@@ -7,12 +7,15 @@ from fastapi.middleware.cors import CORSMiddleware
from fastapi.staticfiles import StaticFiles
from starlette.exceptions import HTTPException as StarletteHTTPException
from . import cleanup
from . import analytics, cleanup
from .auth import MULTI_USER
from .database import engine
from .limits import BodySizeLimitMiddleware
from .migrations import bootstrap
from .routers import adventures, auth, chat, debug, scenarios, scripts, settings, story_cards
from .routers import (
adventures, analytics as analytics_router, auth, chat, debug, scenarios, scripts,
settings, story_cards,
)
from .seed import seed_public_scenarios
bootstrap(engine)
@@ -32,10 +35,15 @@ async def lifespan(_app: FastAPI):
# trigger on Render's free tier, where the service sleeps after ~15
# minutes and a long-running timer rarely gets to fire.
sweeper = cleanup.start_sweeper()
# Visit counters are buffered in memory and written in batches; this is
# what turns them into rows, and stop_flusher writes out the last batch so
# a deploy doesn't drop it.
flusher = analytics.start_flusher()
try:
yield
finally:
await cleanup.stop_sweeper(sweeper)
await analytics.stop_flusher(flusher)
# The interactive API docs stay local-only: in multi-user mode they just hand
@@ -95,6 +103,39 @@ class SecurityHeadersMiddleware:
await self.app(scope, receive, send_with_headers)
class ApiErrorMiddleware:
"""Counts failed API responses for the analytics dashboard.
Here rather than in an exception handler because it sees what the client
actually got: a 429 from a rate limiter, a 404 from routing, a 500 from a
handler that never returned, all the same way. Pure ASGI for the same
reason as the headers above — an SSE turn must not be buffered on its way
out. Only /api is watched; a 404 on the SPA mount is a page load, not a
fault.
"""
def __init__(self, app):
self.app = app
async def __call__(self, scope, receive, send):
if scope["type"] != "http" or not scope.get("path", "").startswith("/api"):
return await self.app(scope, receive, send)
async def send_counting(message):
if message["type"] == "http.response.start" and message["status"] >= 400:
# The router has already put the matched route on the scope by
# the time a response starts, so the label can name the
# endpoint rather than the caller's path.
analytics.record(
analytics.M_ERROR,
analytics.api_route_label(scope, message["status"]),
)
await send(message)
await self.app(scope, receive, send_counting)
app.add_middleware(ApiErrorMiddleware)
app.add_middleware(SecurityHeadersMiddleware)
app.include_router(auth.router)
@@ -105,6 +146,7 @@ app.include_router(scripts.router)
app.include_router(settings.router)
app.include_router(chat.router)
app.include_router(debug.router)
app.include_router(analytics_router.router)
@app.get("/api/health")