Add power-user AI Chat page; centralize demo-key model pinning

AI Chat is a plain scratchpad for talking to a model directly — no story
context, scripts or world state — for poking at models, prompts and endpoints
without starting an adventure. Power users only: the router 404s (rather than
403s) for everyone else and the nav link is hidden. The conversation lives in
localStorage, so there's no new table or migration.

is_power_user() now also returns True in local mode: it's the operator's own
machine and their own key, the same reasoning that makes the provider debug log
local-only.

Alongside that, the rule keeping the shared demo key off paid models now lives
in exactly one place. It had been duplicated into the chat router, which is how
one copy eventually drifts:

- resolve_provider_config() takes an optional model_override and is the only
  place the whitelist is applied, so turns, AI Chat and the connection test all
  inherit it. An override is a per-request preference, never a grant.
- ProviderConfig.__post_init__ refuses to exist when api_key is the demo key
  and the model isn't whitelisted. It keys on the key itself rather than the
  using_demo flag, so a mislabelled config can't slip past, and it raises so a
  future path that bypasses the resolver fails loudly instead of billing.
- The demo branch still pins endpoint_url too — a user-controlled endpoint
  would leak the key itself, which is worse than spending it.

Provider gained chat(messages, ...) beside generate(), both delegating to a
shared _stream(url, body); completion-mode endpoints get the messages flattened
into a labelled transcript. Settings' /models fetch moved to
list_endpoint_models() and is shared with /api/chat/config.

Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub
resolve_provider_config — the point is to exercise the real BYOK-vs-demo
decision and assert on what the provider actually received: off-whitelist
override pinned, off-whitelist Settings.model pinned, redirected endpoint
pinned, BYOK passed through untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
This commit is contained in:
parththakkar106
2026-07-26 20:16:56 +05:30
co-authored by Claude Opus 5
parent 57c24a07d8
commit 1dd31086c1
17 changed files with 969 additions and 20 deletions
@@ -10,6 +10,17 @@ from .base import PromptParts, Provider, ProviderError
# continuing prose instead of replying conversationally.
CHAT_CONTINUE_HINT = "\n\n[Continue the story directly. Output only story text.]"
# Completion endpoints have no roles, so a plain chat has to be flattened into
# one labelled transcript that trails off on "Assistant:" for the model to continue.
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
def flatten_messages(messages: list[dict]) -> str:
turns = "\n\n".join(
f"{_ROLE_LABELS.get(m['role'], m['role'])}: {m['content']}" for m in messages
)
return f"{turns}\n\nAssistant:"
class OpenAICompatibleProvider(Provider):
"""Adapter for any /v1-style endpoint: Ollama, LM Studio, OpenAI, OpenRouter, vLLM, Groq…"""
@@ -110,7 +121,42 @@ class OpenAICompatibleProvider(Provider):
if not self.model:
raise ProviderError("No model configured — set one in Settings.")
url, body = self._request(parts, temperature, max_tokens)
async for event in self._stream(url, body):
yield event
async def chat(
self, messages: list[dict], *, temperature: float, max_tokens: int
) -> AsyncIterator[tuple[str, str]]:
"""Plain multi-turn chat — no story framing, no context assembly. Takes
[{"role", "content"}, ...] straight to the endpoint. Used by the AI Chat
scratchpad; the turn engine uses generate()."""
if not self.model:
raise ProviderError("No model configured — set one in Settings.")
if self.api_mode == "completion":
url = f"{self.base_url}/completions"
body = {
"model": self.model,
"prompt": flatten_messages(messages),
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
}
else:
url = f"{self.base_url}/chat/completions"
body = {
"model": self.model,
"messages": messages,
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
}
self._apply_reasoning_budget(body)
async for event in self._stream(url, body):
yield event
async def _stream(self, url: str, body: dict) -> AsyncIterator[tuple[str, str]]:
"""Shared SSE plumbing for generate()/chat(): POST a streaming request
and yield ("text" | "reasoning", chunk) pairs, logging the exchange."""
log = debuglog.start_entry(url, self.model, body)
received: list[str] = []
try: