M2: cut the hosted product away from the local one

94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The
milestone is subtraction, and what is left is the single-user local
storyteller the specification describes.

Removed in full: campaign scripting and its QuickJS sandbox; multi-user
accounts, guest sessions, login, registration and the shared demo key;
the visitor-analytics tables, dashboard and page beacon; the access log
of sign-ins, addresses and devices; per-IP and per-user rate limiting
and quotas; Render deployment config; Postgres and psycopg; cloud
inference providers, the API-key field and the key encryption that
existed to store it; session-cookie signing. None of it was hidden
behind a flag — the routes are gone and answer 404.

Two things were kept that the brief allowed keeping. The `users` table
and its foreign keys stay as an internal ownership detail, because
rewriting them out means a migration across most of the schema to
delete a column that costs nothing; nothing creates a second user and
no request carries an identity. Five inert tables and four inert
columns stay for the same reason, so an M1 campaign database opens
unchanged.

The one addition is app/endpoints.py, which decides where a story may
be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an
explicit allowlist of networks, not a guess at what `ipaddress` means
by "private", which calls the documentation ranges private and IPv6
loopback reserved. Every address a hostname resolves to must be in it,
so a split answer does not squeak through, and the rule runs both when
the endpoint is saved and before every outbound request, because a name
that resolved to the LAN this morning can resolve elsewhere this
afternoon. Known cloud hosts are named in the refusal so the error says
why rather than looking like broken DNS. TLS is never traded against
it: M1's shared trust context is intact on all four clients and there
is no way to skip verification.

The hardcoded 120-second model timeout is now a setting. That was not
theoretical — on this GPU-less four-core host a cold load of
qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while
turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short
at 10s so a wrong address still fails fast; the read timeout defaults
to 300s and is bounded at 3600, because "wait longer" must stay a
number.

Two defects found while testing and fixed here. An unknown /api path
fell through the SPA catch-all and came back as HTML with status 200,
so a client asking for JSON parsed a web page instead of learning the
route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an
unauthenticated loopback API would hand every page on the Internet a
write handle on the campaign database; it now refuses to start.

Verified rather than assumed. Offline, on a network with no route out
and no DNS: five turns, retry with both takes retained, restart with an
identical transcript digest, a failed model call leaving the accepted
AI-turn count untouched, and a capture with zero non-loopback unicast
packets. Against a real second machine on the LAN over HTTPS with a
private CA: four turns, restart, and a capture showing 289 packets to
the approved host, 344 loopback, zero anywhere else, zero DNS queries.
Cloud and public endpoints refused with their reasons; no API key
settable; every removed route 404.

604 backend tests pass, down from 648 by the fifteen retired with the
subsystems they tested and up by the twenty-nine added for the endpoint
policy and the removed surface. The scripting tests were not deleted:
eight files used a JavaScript counter as instrumentation for the state
snapshot and rollback machinery, which M2 does not touch, so the
counter moved to the world-state engine and those tests still assert
what they always did. Frontend lint and build are clean; the image
builds, and its wheel-building stage is gone with quickjs.

No M3 work. Undo is still destructive and there is still no Redo.
This commit is contained in:
JesseMarkowitz
2026-09-02 11:27:14 -04:00
parent 1a28a9a708
commit 8c65ae99de
94 changed files with 1384 additions and 6567 deletions
+54 -82
View File
@@ -3,32 +3,27 @@ from typing import AsyncIterator
import httpx
from .. import debuglog, netguard, tlstrust
from .. import debuglog, endpoints, tlstrust
from .base import PromptParts, Provider, ProviderError
# Appended after the story text in chat mode, so a chat-tuned model continues
# the prose rather than replying conversationally.
CHAT_CONTINUE_HINT = "\n\n[Continue the story directly. Output only story text.]"
# OpenRouter serves one model from whichever upstream is available, and every
# upstream holds its own prompt cache, so a request routed somewhere new starts
# with a cold cache however stable the prompt is. Naming a preferred upstream
# makes routing deterministic, which is what allows a cache hit at all.
#
# `allow_fallbacks` stays at its default of true on purpose, because this is a
# preference rather than a restriction. If the named upstream is down, the
# request still goes elsewhere and only misses the cache, which is the behavior
# without this setting.
#
# This is a list rather than a value derived from the model slug. The vendor half
# of a slug is usually the provider slug, such as "deepseek/..." mapping to
# "deepseek", which was verified against /api/v1/providers, but not reliably.
# Google's models are served by "google-ai-studio" and "google-vertex", and there
# is no "google". Look a vendor up on the model's Providers tab before adding it
# here. A slug that does not exist is a routing preference that, at best, does
# nothing.
_OPENROUTER_HOST = "openrouter.ai"
_PREFERRED_UPSTREAM = {"deepseek": "deepseek"}
# A machine that is not listening refuses in milliseconds, so a slow connect
# means the wrong address rather than a busy model.
CONNECT_TIMEOUT = 10.0
# How long to wait for generation when Settings names no value. Upstream
# hardcoded 120s, and M1 measured a *cold* load of a 3B model on a GPU-less
# four-core host exceeding it three times while the same turn took 6-9 seconds
# once the model was resident. 300s covers a cold start on modest hardware and
# is still a number: a wedged endpoint fails rather than hanging forever.
DEFAULT_READ_TIMEOUT = 300.0
# Embeddings are short and never cold-load a large model.
EMBED_READ_TIMEOUT = 60.0
# Completion endpoints have no roles, so a chat has to be flattened into one
@@ -44,74 +39,47 @@ def flatten_messages(messages: list[dict]) -> str:
class OpenAICompatibleProvider(Provider):
"""Adapter for any /v1-style endpoint.
"""Adapter for Ollama's OpenAI-compatible `/v1` API.
This covers Ollama, LM Studio, OpenAI, OpenRouter, vLLM, and Groq, among
others.
The protocol is OpenAI's, which is what the module is named for; the
product speaks it to Ollama and to nothing else. `endpoints.py` decides
which addresses may be reached, and every request re-checks — the shape of
the wire format is not the same thing as permission to use it.
"""
def __init__(
self,
endpoint_url: str,
api_key: str,
model: str,
api_mode: str = "chat",
reasoning_max_tokens: int = 0,
read_timeout: float | None = None,
):
self.base_url = endpoint_url.rstrip("/")
self.api_key = api_key
self.model = model
self.api_mode = api_mode # Either "chat" or "completion".
# The thinking budget for reasoning models, on top of `max_tokens`. A
# value of 0 means the `reasoning` parameter is not sent, because an
# endpoint that does not know the field may reject it. A negative value
# asks the endpoint to turn reasoning off.
self.reasoning_max_tokens = reasoning_max_tokens
# How long to wait for the model, in seconds. Cold-loading a model on a
# CPU-only machine can take minutes, and a fixed short timeout reports
# that as a failure. See `DEFAULT_READ_TIMEOUT`.
self.read_timeout = read_timeout or DEFAULT_READ_TIMEOUT
# The token accounting from the last call, when the endpoint reported
# any. It holds the prompt and completion counts, plus, on OpenRouter,
# `prompt_tokens_details.cached_tokens`, which is the number of prompt
# tokens read from cache rather than billed in full. Every request method
# writes it, so a caller reads it after the call it made. One provider is
# built per request.
# any. Every request method writes it, so a caller reads it after the
# call it made. One provider is built per request.
self.last_usage: dict | None = None
def _headers(self) -> dict:
headers = {"Content-Type": "application/json"}
if self.api_key:
headers["Authorization"] = f"Bearer {self.api_key}"
return headers
# No Authorization header: Ollama does not use one, and this build has
# no cloud provider to carry a key for.
return {"Content-Type": "application/json"}
def _apply_reasoning_budget(self, body: dict) -> None:
"""Gives reasoning models their own thinking budget, in the OpenRouter style.
def _timeout(self, seconds: float | None = None) -> httpx.Timeout:
"""Short to connect, patient to read.
The method raises `max_tokens`, so the output keeps its full budget.
A negative budget does the opposite. It sends `effort: "none"` to turn
reasoning off on a model that reasons by default, such as DeepSeek V4
Flash. That differs from `exclude: true`, which still reasons and still
bills for it while hiding the trace. Zero still means send nothing, so an
endpoint that rejects unknown fields, such as Ollama, keeps working.
A machine that is not listening says so in milliseconds, so a slow
connect is a wrong address rather than a busy model and should fail
fast. Generation is the opposite: the first token can be minutes away
while a model loads.
"""
if self.api_mode != "chat":
return
if self.reasoning_max_tokens < 0:
body["reasoning"] = {"effort": "none"}
elif self.reasoning_max_tokens > 0:
body["reasoning"] = {"max_tokens": self.reasoning_max_tokens}
body["max_tokens"] += self.reasoning_max_tokens
def _apply_provider_routing(self, body: dict) -> None:
"""Prefers one upstream on OpenRouter, so the prompt cache stays warm.
The method does nothing anywhere else. `provider` is an OpenRouter
extension, and Ollama and similar servers reject fields they do not know.
The `reasoning` parameter above is written around the same constraint.
"""
if _OPENROUTER_HOST not in self.base_url:
return
upstream = _PREFERRED_UPSTREAM.get(self.model.split("/", 1)[0].lower())
if upstream:
body["provider"] = {"order": [upstream]}
return httpx.Timeout(seconds or self.read_timeout, connect=CONNECT_TIMEOUT)
def _record_usage(self, payload: dict) -> None:
"""Records the endpoint's own token accounting, if it reported any.
@@ -147,8 +115,6 @@ class OpenAICompatibleProvider(Provider):
"max_tokens": max_tokens,
"stream": True,
}
self._apply_reasoning_budget(body)
self._apply_provider_routing(body)
return url, body
@staticmethod
@@ -227,8 +193,6 @@ class OpenAICompatibleProvider(Provider):
"max_tokens": max_tokens,
"stream": True,
}
self._apply_reasoning_budget(body)
self._apply_provider_routing(body)
async for event in self._stream(url, body):
yield event
@@ -238,17 +202,17 @@ class OpenAICompatibleProvider(Provider):
The method POSTs a streaming request, yields `("text", chunk)` and
`("reasoning", chunk)` pairs, and logs the exchange.
"""
# SSRF guard for hosted mode. A user-supplied `endpoint_url` must not
# point at an internal or metadata address. This does nothing for a
# local install.
reason = netguard.endpoint_block_reason(url)
# Re-checked on every request, not only when the endpoint was saved: a
# hostname that resolved to a LAN address yesterday can resolve
# somewhere else today, and a database row can be edited by hand.
reason = endpoints.rejection_reason(url)
if reason:
raise ProviderError(f"This endpoint can't be used — {reason}.")
log = debuglog.start_entry(url, self.model, body)
received: list[str] = []
try:
async with httpx.AsyncClient(
timeout=httpx.Timeout(120, connect=10), verify=tlstrust.ssl_context()
timeout=self._timeout(), verify=tlstrust.ssl_context()
) as client:
async with client.stream("POST", url, json=body, headers=self._headers()) as resp:
if resp.status_code != 200:
@@ -351,13 +315,16 @@ class OpenAICompatibleProvider(Provider):
"max_tokens": max_tokens,
"stream": False,
}
self._apply_reasoning_budget(body)
self._apply_provider_routing(body)
# Same check as `_stream`: every outbound request re-tests the
# endpoint, so no path reaches an address the policy refuses.
reason = endpoints.rejection_reason(url)
if reason:
raise ProviderError(f"This endpoint can't be used — {reason}.")
log = debuglog.start_entry(url, self.model, body)
try:
async with httpx.AsyncClient(
timeout=httpx.Timeout(120, connect=10), verify=tlstrust.ssl_context()
timeout=self._timeout(), verify=tlstrust.ssl_context()
) as client:
resp = await client.post(url, json=body, headers=self._headers())
except httpx.HTTPError as exc:
@@ -383,10 +350,15 @@ class OpenAICompatibleProvider(Provider):
raise ProviderError("No embedding model configured — set one in Settings.")
url = f"{self.base_url}/embeddings"
body = {"model": self.model, "input": texts}
# Same check as `_stream`: every outbound request re-tests the
# endpoint, so no path reaches an address the policy refuses.
reason = endpoints.rejection_reason(url)
if reason:
raise ProviderError(f"This endpoint can't be used — {reason}.")
log = debuglog.start_entry(url, self.model, body)
try:
async with httpx.AsyncClient(
timeout=httpx.Timeout(60, connect=10), verify=tlstrust.ssl_context()
timeout=self._timeout(EMBED_READ_TIMEOUT), verify=tlstrust.ssl_context()
) as client:
resp = await client.post(url, json=body, headers=self._headers())
except httpx.HTTPError as exc: