M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo.
This commit is contained in:
+30
-194
@@ -1,183 +1,35 @@
|
||||
"""Phase 9: abuse guards for hosted, multi-user deployments.
|
||||
"""Resource bounds on what a single request or a single story may cost.
|
||||
|
||||
Rate limits and row caps do nothing in local mode, because a single local player
|
||||
should never be throttled by their own app. The values are hardcoded on purpose.
|
||||
They are generous enough that a legitimate player never notices them, and tight
|
||||
enough that a hostile visitor cannot exhaust the demo key, saturate the CPU, or
|
||||
fill the database.
|
||||
Upstream carried three things here, and only one of them belongs in a local
|
||||
single-user product. Per-IP and per-user **rate limiting**, the login-attempt
|
||||
throttle, and the per-user **quotas** were hosted-service policy: they existed
|
||||
to stop a hostile visitor exhausting a shared demo key or filling a shared
|
||||
database. M2 removed all of it. There are no visitors, and throttling the one
|
||||
person who started the application would be a bug rather than a guard.
|
||||
|
||||
What is left is defensive programming, and it applies whatever the deployment:
|
||||
|
||||
* a ceiling on the **request body**, so a malformed or hostile payload cannot
|
||||
be read into memory before anything looks at it;
|
||||
* ceilings on how large **one adventure** may grow, in actions, memories,
|
||||
story cards and branches. These bound storage and the cost of the queries
|
||||
that walk them. They are per-story, not per-user: nothing here counts how
|
||||
many campaigns a person may have.
|
||||
|
||||
An import is checked against the same per-adventure ceilings that live creation
|
||||
uses, so a bundle cannot carry a story past a limit that play could not reach.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import threading
|
||||
import time
|
||||
from collections import defaultdict, deque
|
||||
|
||||
from fastapi import HTTPException, Request
|
||||
from fastapi import HTTPException
|
||||
from sqlalchemy import func
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from . import auth, models
|
||||
from . import models
|
||||
|
||||
# ---------- Rate limiting ----------
|
||||
# Fixed windows per scope and caller. The windows live in memory, which is
|
||||
# enough for the single-process deployment this app targets. The worst case
|
||||
# after a restart is a brief extra allowance.
|
||||
# ---------- Per-story row caps ----------
|
||||
|
||||
# Maps a scope to (max requests, window seconds).
|
||||
RATE_LIMITS: dict[str, tuple[int, int]] = {
|
||||
"turn": (10, 60), # AI turn generation. The demo key also has a daily cap.
|
||||
"chat": (30, 60), # The AI Chat scratchpad, for power users.
|
||||
"script-test": (30, 60), # Sandboxed, but each run costs up to 2s of CPU.
|
||||
"connection-test": (10, 60), # Outbound HTTP to a user-supplied URL.
|
||||
"import": (30, 60), # Large writes.
|
||||
"auth": (10, 300), # Register and login attempts, per IP.
|
||||
"guest": (30, 300), # New guest users, per IP. Each one is a database row.
|
||||
# Pageview beacons. The limit is generous, because a real reader clicking
|
||||
# around a SPA sends a handful a minute, and it is low enough that nobody
|
||||
# can inflate the traffic numbers faster than by reloading the page.
|
||||
"analytics": (120, 60),
|
||||
}
|
||||
|
||||
_windows: dict[tuple[str, str], deque] = defaultdict(deque)
|
||||
_windows_guard = threading.Lock()
|
||||
|
||||
|
||||
# How many proxy hops sit between the app and the real client. On Render, and on
|
||||
# most platforms, that is one, because the platform's edge appends the connecting
|
||||
# IP to the right of `X-Forwarded-For`. A client can prepend any value on the
|
||||
# left, but it cannot push a value past the edge's own append, so the trustworthy
|
||||
# client IP is the entry that many places from the right rather than uvicorn's
|
||||
# leftmost choice. Trusting the leftmost entry let anyone rotate
|
||||
# `X-Forwarded-For` to get a fresh rate-limit bucket per request and bypass the
|
||||
# auth and guest limits. If the deployment adds more hops, set
|
||||
# `AIDND_TRUSTED_PROXY_HOPS`.
|
||||
TRUSTED_PROXY_HOPS = max(1, int(os.environ.get("AIDND_TRUSTED_PROXY_HOPS", "1") or 1))
|
||||
|
||||
|
||||
def client_ip(request: Request) -> str:
|
||||
"""Returns the real client IP, resisting a spoofed `X-Forwarded-For`.
|
||||
|
||||
The function reads the hop the trusted edge appended, which is the rightmost
|
||||
entry minus any extra trusted hops. If no forwarded header is present, which
|
||||
happens locally, in development, and on a direct connection, it falls back to
|
||||
the socket peer.
|
||||
|
||||
The function is public because the access log needs the same answer. Two
|
||||
functions that each decide which address belongs to the caller is how one of
|
||||
them ends up trusting a header it should not.
|
||||
"""
|
||||
forwarded = request.headers.get("x-forwarded-for")
|
||||
if forwarded:
|
||||
parts = [p.strip() for p in forwarded.split(",") if p.strip()]
|
||||
if parts:
|
||||
return parts[-min(TRUSTED_PROXY_HOPS, len(parts))]
|
||||
return request.client.host if request.client else "unknown"
|
||||
|
||||
|
||||
def rate_limit(scope: str, request: Request, user: models.User | None = None) -> None:
|
||||
"""Raises a 429 when the caller exceeds the scope's window.
|
||||
|
||||
The window is keyed per user when a user is known, because an account
|
||||
survives an IP change, and per IP otherwise.
|
||||
"""
|
||||
if not auth.MULTI_USER:
|
||||
return
|
||||
limit, window_seconds = RATE_LIMITS[scope]
|
||||
key = (scope, f"u{user.id}" if user else f"ip{client_ip(request)}")
|
||||
now = time.time()
|
||||
with _windows_guard:
|
||||
window = _windows[key]
|
||||
while window and window[0] < now - window_seconds:
|
||||
window.popleft()
|
||||
if len(window) >= limit:
|
||||
raise HTTPException(
|
||||
429, "You're doing that too fast — wait a minute and try again."
|
||||
)
|
||||
window.append(now)
|
||||
if len(_windows) > 10_000:
|
||||
_prune(now)
|
||||
|
||||
|
||||
# ---------- Per-account login throttle ----------
|
||||
# This is defense in depth next to the per-IP `auth` limit. A botnet dilutes
|
||||
# that limit, because many real source IPs each get their own bucket, so it
|
||||
# cannot by itself stop a distributed guessing run against one account. This cap
|
||||
# keys on the target email rather than on the caller, so guessing one account's
|
||||
# password stays expensive however many addresses the guesses come from.
|
||||
#
|
||||
# Only failures count, and a correct password clears the record. The window
|
||||
# slides over a short period rather than locking the account, so a user who
|
||||
# mistypes a few times recovers within minutes. The trade-off is that an
|
||||
# attacker can keep a known account throttled, which is an inconvenience and is
|
||||
# preferable to letting the account be brute-forced.
|
||||
LOGIN_FAIL_LIMIT = 8 # Failed attempts per account.
|
||||
LOGIN_FAIL_WINDOW = 900 # The window in seconds, which is 15 minutes.
|
||||
|
||||
_login_fails: dict[str, deque] = defaultdict(deque)
|
||||
_login_guard = threading.Lock()
|
||||
|
||||
|
||||
def check_login_allowed(email: str) -> None:
|
||||
"""Raises a 429 when an account has too many recent failed logins.
|
||||
|
||||
Call this before verifying the password, so that a guess never reaches the
|
||||
hash.
|
||||
"""
|
||||
if not auth.MULTI_USER:
|
||||
return
|
||||
now = time.time()
|
||||
with _login_guard:
|
||||
window = _login_fails[email]
|
||||
while window and window[0] < now - LOGIN_FAIL_WINDOW:
|
||||
window.popleft()
|
||||
if len(window) >= LOGIN_FAIL_LIMIT:
|
||||
raise HTTPException(
|
||||
429,
|
||||
"Too many failed sign-in attempts for this account — "
|
||||
"wait a few minutes and try again.",
|
||||
)
|
||||
|
||||
|
||||
def note_login_failure(email: str) -> None:
|
||||
"""Records one failed attempt against `email`."""
|
||||
if not auth.MULTI_USER:
|
||||
return
|
||||
now = time.time()
|
||||
with _login_guard:
|
||||
_login_fails[email].append(now)
|
||||
if len(_login_fails) > 10_000: # Bound the map against a flood of unique emails.
|
||||
stale = [
|
||||
key for key, window in _login_fails.items()
|
||||
if not window or window[-1] < now - LOGIN_FAIL_WINDOW
|
||||
]
|
||||
for key in stale:
|
||||
del _login_fails[key]
|
||||
|
||||
|
||||
def note_login_success(email: str) -> None:
|
||||
"""Clears the account's failure record after a correct password."""
|
||||
with _login_guard:
|
||||
_login_fails.pop(email, None)
|
||||
|
||||
|
||||
def _prune(now: float) -> None:
|
||||
"""Drops callers whose whole window has expired, so the per-IP dict stays bounded.
|
||||
|
||||
Call this with the guard held.
|
||||
"""
|
||||
longest = max(seconds for _, seconds in RATE_LIMITS.values())
|
||||
stale = [key for key, window in _windows.items()
|
||||
if not window or window[-1] < now - longest]
|
||||
for key in stale:
|
||||
del _windows[key]
|
||||
|
||||
|
||||
# ---------- Per-user row caps ----------
|
||||
|
||||
MAX_ADVENTURES_PER_USER = 100
|
||||
MAX_SCENARIOS_PER_USER = 200
|
||||
MAX_SCRIPTS_PER_USER = 200
|
||||
MAX_STORY_CARDS_PER_OWNER = 200 # Per scenario or per adventure.
|
||||
MAX_MEMORIES_PER_ADVENTURE = 1000
|
||||
MAX_ACTIONS_PER_ADVENTURE = 5000
|
||||
@@ -200,28 +52,14 @@ def check_row_cap(
|
||||
) -> None:
|
||||
"""Raises a 409 when creating one more row of `kind` would exceed its cap.
|
||||
|
||||
The caller has already checked ownership of the scenario or adventure passed
|
||||
in.
|
||||
Only per-story kinds are capped. `adventures` and `scenarios` were per-user
|
||||
quotas and are no longer checked; the callers still pass them, and they are
|
||||
accepted and ignored so that adding a cap back is a change here rather than
|
||||
at every call site.
|
||||
"""
|
||||
if not auth.MULTI_USER:
|
||||
if kind in ("adventures", "scenarios"):
|
||||
return
|
||||
if kind == "adventures":
|
||||
count = _count(db, models.Adventure, models.Adventure.user_id == user.id)
|
||||
cap, subject, hint = (
|
||||
MAX_ADVENTURES_PER_USER, "adventures",
|
||||
"delete one you no longer play to make room",
|
||||
)
|
||||
elif kind == "scenarios":
|
||||
count = _count(db, models.Scenario, models.Scenario.user_id == user.id)
|
||||
cap, subject, hint = (
|
||||
MAX_SCENARIOS_PER_USER, "scenarios", "delete one to make room"
|
||||
)
|
||||
elif kind == "scripts":
|
||||
count = _count(db, models.Script, models.Script.user_id == user.id)
|
||||
cap, subject, hint = (
|
||||
MAX_SCRIPTS_PER_USER, "scripts", "delete one to make room"
|
||||
)
|
||||
elif kind == "story_cards":
|
||||
if kind == "story_cards":
|
||||
owner_filter = (
|
||||
models.StoryCard.scenario_id == scenario_id
|
||||
if scenario_id is not None
|
||||
@@ -273,8 +111,6 @@ def check_bundle_lists(**lists) -> None:
|
||||
The keyword arguments are `story_cards`, `memories`, `actions`, and
|
||||
`branches`.
|
||||
"""
|
||||
if not auth.MULTI_USER:
|
||||
return
|
||||
for name, value in lists.items():
|
||||
cap = _BUNDLE_LIST_CAPS[name]
|
||||
if isinstance(value, list) and len(value) > cap:
|
||||
@@ -286,8 +122,8 @@ def check_bundle_lists(**lists) -> None:
|
||||
|
||||
# ---------- Request body size ----------
|
||||
# The limit is generous enough for the largest legitimate payload, which is an
|
||||
# adventure export holding thousands of actions. It applies in every mode, and no
|
||||
# honest request approaches it.
|
||||
# adventure export holding thousands of actions. No honest request approaches
|
||||
# it.
|
||||
|
||||
MAX_BODY_BYTES = 2 * 1024 * 1024
|
||||
MAX_IMPORT_BODY_BYTES = 20 * 1024 * 1024
|
||||
|
||||
Reference in New Issue
Block a user