M2: cut the hosted product away from the local one

94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The
milestone is subtraction, and what is left is the single-user local
storyteller the specification describes.

Removed in full: campaign scripting and its QuickJS sandbox; multi-user
accounts, guest sessions, login, registration and the shared demo key;
the visitor-analytics tables, dashboard and page beacon; the access log
of sign-ins, addresses and devices; per-IP and per-user rate limiting
and quotas; Render deployment config; Postgres and psycopg; cloud
inference providers, the API-key field and the key encryption that
existed to store it; session-cookie signing. None of it was hidden
behind a flag — the routes are gone and answer 404.

Two things were kept that the brief allowed keeping. The `users` table
and its foreign keys stay as an internal ownership detail, because
rewriting them out means a migration across most of the schema to
delete a column that costs nothing; nothing creates a second user and
no request carries an identity. Five inert tables and four inert
columns stay for the same reason, so an M1 campaign database opens
unchanged.

The one addition is app/endpoints.py, which decides where a story may
be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an
explicit allowlist of networks, not a guess at what `ipaddress` means
by "private", which calls the documentation ranges private and IPv6
loopback reserved. Every address a hostname resolves to must be in it,
so a split answer does not squeak through, and the rule runs both when
the endpoint is saved and before every outbound request, because a name
that resolved to the LAN this morning can resolve elsewhere this
afternoon. Known cloud hosts are named in the refusal so the error says
why rather than looking like broken DNS. TLS is never traded against
it: M1's shared trust context is intact on all four clients and there
is no way to skip verification.

The hardcoded 120-second model timeout is now a setting. That was not
theoretical — on this GPU-less four-core host a cold load of
qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while
turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short
at 10s so a wrong address still fails fast; the read timeout defaults
to 300s and is bounded at 3600, because "wait longer" must stay a
number.

Two defects found while testing and fixed here. An unknown /api path
fell through the SPA catch-all and came back as HTML with status 200,
so a client asking for JSON parsed a web page instead of learning the
route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an
unauthenticated loopback API would hand every page on the Internet a
write handle on the campaign database; it now refuses to start.

Verified rather than assumed. Offline, on a network with no route out
and no DNS: five turns, retry with both takes retained, restart with an
identical transcript digest, a failed model call leaving the accepted
AI-turn count untouched, and a capture with zero non-loopback unicast
packets. Against a real second machine on the LAN over HTTPS with a
private CA: four turns, restart, and a capture showing 289 packets to
the approved host, 344 loopback, zero anywhere else, zero DNS queries.
Cloud and public endpoints refused with their reasons; no API key
settable; every removed route 404.

604 backend tests pass, down from 648 by the fifteen retired with the
subsystems they tested and up by the twenty-nine added for the endpoint
policy and the removed surface. The scripting tests were not deleted:
eight files used a JavaScript counter as instrumentation for the state
snapshot and rollback machinery, which M2 does not touch, so the
counter moved to the world-state engine and those tests still assert
what they always did. Frontend lint and build are clean; the image
builds, and its wheel-building stage is gone with quickjs.

No M3 work. Undo is still destructive and there is still no Redo.
This commit is contained in:
JesseMarkowitz
2026-09-02 11:27:14 -04:00
parent 1a28a9a708
commit 8c65ae99de
94 changed files with 1384 additions and 6567 deletions
+43 -165
View File
@@ -1,8 +1,14 @@
"""HTTP tests for the AI Chat scratchpad (power users only).
"""HTTP tests for the AI Chat scratchpad.
Covers the access gate, the streamed reply, and the demo-key model pinning.
This pinning must not let a public visitor reach paid models through this
page.
Most of this file used to be about the shared demo key: an access gate on a
"power user" email allowlist, and a pinning rule that stopped a public visitor
reaching paid models on a server-funded key. M2 removed the hosted deployment
those defended, so the rules they tested no longer exist to be tested. See
`planning/reports/M2-*` for the accounting.
What remains is what the page still does: stream a reply from the configured
model, honour a system prompt and a per-request model override, and refuse a
conversation too large to send.
python -m pytest tests/test_chat.py -v
"""
@@ -10,25 +16,27 @@ import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, models
from app import auth, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import chat
class FakeProvider:
"""Records what it was constructed with, then streams a fixed reply. Stands
in for the real egress point, so asserting on last_key/last_model is
asserting on exactly what would have gone over the wire."""
"""Records what it was constructed with, then streams a fixed reply.
It stands in for the real egress point, so asserting on `last_endpoint` and
`last_model` is asserting on exactly what would have gone over the wire.
There is no `last_key` any more: the provider takes no API key, because
Ollama does not use one.
"""
last_usage = None
last_model = None
last_key = None
last_endpoint = None
last_messages = None
def __init__(self, endpoint_url, api_key, model, api_mode="chat", reasoning_max_tokens=0):
def __init__(self, endpoint_url, model, api_mode="chat", read_timeout=None):
FakeProvider.last_model = model
FakeProvider.last_key = api_key
FakeProvider.last_endpoint = endpoint_url
async def chat(self, messages, *, temperature, max_tokens):
@@ -41,68 +49,35 @@ class FakeProvider:
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="power@example.com")
user = models.User(is_guest=False)
setup.add(user)
setup.flush()
setup.add(models.Settings(user_id=user.id, api_key="enc:dummy", model="test-model"))
setup.add(models.Settings(user_id=user.id, model="test-model"))
setup.commit()
user_id = user.id
setup.close()
monkeypatch.setattr(chat, "OpenAICompatibleProvider", FakeProvider)
monkeypatch.setattr(limits, "rate_limit", lambda *a, **k: None)
# Multi-user mode is what makes the power-user gate meaningful, because
# local mode trusts everyone. The allowlist is set per test.
monkeypatch.setattr(auth, "MULTI_USER", True)
monkeypatch.setattr(auth, "POWER_USERS", {"power@example.com"})
# These tests deliberately do not stub resolve_provider_config. The
# point is to exercise the real BYOK-vs-demo decision, since that
# decision is what keeps the shared key off paid models. Each test
# picks a mode with _byok/_demo below.
def _current_user(db=Depends(get_db)):
return db.get(models.User, user_id)
app.dependency_overrides[auth.get_current_user] = _current_user
c = TestClient(app)
try:
yield TestClient(app)
yield c
finally:
app.dependency_overrides.clear()
Base.metadata.drop_all(bind=engine)
def _send(client, **extra):
return client.post("/api/chat/stream", json={"messages": [{"role": "user", "content": "hi"}], **extra})
def _send(client, **body):
payload = {"messages": [{"role": "user", "content": "hi"}]}
payload.update(body)
return client.post("/api/chat/stream", json=payload)
def _byok(monkeypatch):
"""The user brought their own key: no demo key in play, any model allowed."""
monkeypatch.setattr(auth, "demo_enabled", lambda: False)
db = SessionLocal()
try:
settings = db.query(models.Settings).first()
settings.api_key = "sk-my-own-key" # legacy-plaintext path: used as-is
db.commit()
finally:
db.close()
def _demo(monkeypatch, whitelist=("free/allowed",)):
"""The user has no key, so turns run on the server-funded demo key."""
monkeypatch.setattr(auth, "demo_enabled", lambda: True)
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_ENDPOINT_URL", "http://demo")
monkeypatch.setattr(auth, "DEMO_MODELS", list(whitelist))
def test_non_power_user_gets_404(client, monkeypatch):
monkeypatch.setattr(auth, "POWER_USERS", set())
assert _send(client).status_code == 404
assert client.get("/api/chat/config").status_code == 404
def test_power_user_streams_a_reply(client, monkeypatch):
_byok(monkeypatch)
def test_the_page_streams_a_reply(client):
resp = _send(client)
assert resp.status_code == 200, resp.text
assert '"type": "reasoning"' in resp.text
@@ -111,136 +86,39 @@ def test_power_user_streams_a_reply(client, monkeypatch):
assert FakeProvider.last_messages == [{"role": "user", "content": "hi"}]
def test_system_prompt_and_model_override_are_honoured(client, monkeypatch):
_byok(monkeypatch)
def test_it_uses_the_configured_endpoint_and_model(client):
_send(client)
assert FakeProvider.last_model == "test-model"
# The default from `models.Settings`, and the only kind of address the
# endpoint policy allows without configuration.
assert FakeProvider.last_endpoint == "http://localhost:11434/v1"
def test_system_prompt_and_model_override_are_honoured(client):
resp = client.post("/api/chat/stream", json={
"messages": [
{"role": "system", "content": "Be terse."},
{"role": "user", "content": "hi"},
],
"model": "some/other-model",
"model": "some-other-model",
})
assert resp.status_code == 200, resp.text
# BYOK: any model the user names is passed straight through, on their key.
assert FakeProvider.last_model == "some/other-model"
assert FakeProvider.last_key == "sk-my-own-key"
assert FakeProvider.last_model == "some-other-model"
assert FakeProvider.last_messages[0] == {"role": "system", "content": "Be terse."}
def test_demo_key_pins_model_to_whitelist(client, monkeypatch):
_demo(monkeypatch)
resp = _send(client, model="expensive/paid-model")
assert resp.status_code == 200, resp.text
# Refused visibly: the whitelisted model runs instead, with a note. The
# paid slug must never reach the wire alongside the server-funded key.
assert FakeProvider.last_model == "free/allowed"
assert FakeProvider.last_key == "demo-key"
assert '"type": "note"' in resp.text
# A whitelisted model is still selectable on the demo key.
_demo(monkeypatch, ["free/allowed", "free/second"])
resp = _send(client, model="free/second")
assert resp.status_code == 200, resp.text
assert FakeProvider.last_model == "free/second"
def test_demo_key_ignores_an_off_whitelist_settings_model(client, monkeypatch):
"""The override is not the only untrusted input. `Settings.model` is
also user-set, and it must be pinned the same way when there is no
BYOK key."""
_demo(monkeypatch)
db = SessionLocal()
try:
db.query(models.Settings).first().model = "expensive/paid-model"
db.commit()
finally:
db.close()
resp = _send(client)
assert resp.status_code == 200, resp.text
assert FakeProvider.last_model == "free/allowed"
def test_demo_key_endpoint_cannot_be_redirected(client, monkeypatch):
"""A user-controlled `endpoint_url` would leak the key itself, which is
worse than spending it. The demo branch pins the URL too."""
_demo(monkeypatch)
db = SessionLocal()
try:
db.query(models.Settings).first().endpoint_url = "http://attacker.example/v1"
db.commit()
finally:
db.close()
assert _send(client).status_code == 200
assert FakeProvider.last_endpoint == "http://demo"
assert FakeProvider.last_key == "demo-key"
def test_provider_config_refuses_server_funded_paid_model(monkeypatch):
"""The structural backstop: a hand-built config (a future code path that
forgets to go through resolve_provider_config) cannot run a
server-funded turn on an off-whitelist model."""
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_MODELS", ["free/allowed"])
with pytest.raises(ValueError):
auth.ProviderConfig("http://demo", "demo-key", "expensive/paid-model", True)
auth.ProviderConfig("http://demo", "demo-key", "free/allowed", True) # whitelisted: fine
# The user's own key with any model stays fine.
auth.ProviderConfig("http://any", "sk-mine", "expensive/paid-model", False)
def test_byok_user_may_reuse_the_demo_keys_value(client, monkeypatch):
"""Regression: the demo key is just an OpenRouter key, so a user can paste
that same value into their own Settings. That is still BYOK, because
the user is paying, and it must not trip the guard. It used to raise
on every resolution, which returned a 500 from `GET /auth/me` and
broke the entire SPA (no nav, no chat)."""
monkeypatch.setattr(auth, "demo_enabled", lambda: True)
monkeypatch.setattr(auth, "DEMO_API_KEY", "shared-key")
monkeypatch.setattr(auth, "DEMO_ENDPOINT_URL", "http://demo")
monkeypatch.setattr(auth, "DEMO_MODELS", ["free/allowed"])
def test_a_request_with_no_model_anywhere_is_refused(client):
db = SessionLocal()
try:
settings = db.query(models.Settings).first()
settings.api_key = "shared-key" # same value, but supplied by the user
settings.model = "expensive/paid-model" # their spend, their choice
settings.model = ""
db.commit()
finally:
db.close()
assert client.get("/api/auth/me").status_code == 200
assert client.get("/api/chat/config").status_code == 200
resp = _send(client)
assert resp.status_code == 200, resp.text
assert FakeProvider.last_model == "expensive/paid-model"
assert FakeProvider.last_key == "shared-key"
assert _send(client).status_code == 400
def test_resolve_provider_config_is_the_single_choke_point(monkeypatch):
"""Turns, AI Chat, and the connection test all resolve through this one
function, so pinning it here pins every caller. No DB or HTTP needed."""
monkeypatch.setattr(auth, "demo_enabled", lambda: True)
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_ENDPOINT_URL", "http://demo")
monkeypatch.setattr(auth, "DEMO_MODELS", ["free/allowed"])
# No key of their own: both endpoint and model are pinned, regardless
# of what they set.
no_key = models.Settings(endpoint_url="http://mine/v1", api_key="", model="expensive/paid")
assert auth.resolve_provider_config(no_key) == auth.ProviderConfig(
"http://demo", "demo-key", "free/allowed", True)
assert auth.resolve_provider_config(
no_key, model_override="expensive/paid").model == "free/allowed"
assert auth.resolve_provider_config(
no_key, model_override="free/allowed").model == "free/allowed"
# Their own key: their endpoint, their key, their choice of model.
byok = models.Settings(endpoint_url="http://mine/v1", api_key="sk-mine", model="expensive/paid")
assert auth.resolve_provider_config(byok) == auth.ProviderConfig(
"http://mine/v1", "sk-mine", "expensive/paid", False)
def test_oversized_conversation_is_refused(client, monkeypatch):
_byok(monkeypatch)
def test_oversized_conversation_is_refused(client):
huge = "x" * 90_000
resp = client.post("/api/chat/stream", json={
"messages": [{"role": "user", "content": huge} for _ in range(5)],