Add power-user AI Chat page; centralize demo-key model pinning

AI Chat is a plain scratchpad for talking to a model directly — no story
context, scripts or world state — for poking at models, prompts and endpoints
without starting an adventure. Power users only: the router 404s (rather than
403s) for everyone else and the nav link is hidden. The conversation lives in
localStorage, so there's no new table or migration.

is_power_user() now also returns True in local mode: it's the operator's own
machine and their own key, the same reasoning that makes the provider debug log
local-only.

Alongside that, the rule keeping the shared demo key off paid models now lives
in exactly one place. It had been duplicated into the chat router, which is how
one copy eventually drifts:

- resolve_provider_config() takes an optional model_override and is the only
  place the whitelist is applied, so turns, AI Chat and the connection test all
  inherit it. An override is a per-request preference, never a grant.
- ProviderConfig.__post_init__ refuses to exist when api_key is the demo key
  and the model isn't whitelisted. It keys on the key itself rather than the
  using_demo flag, so a mislabelled config can't slip past, and it raises so a
  future path that bypasses the resolver fails loudly instead of billing.
- The demo branch still pins endpoint_url too — a user-controlled endpoint
  would leak the key itself, which is worse than spending it.

Provider gained chat(messages, ...) beside generate(), both delegating to a
shared _stream(url, body); completion-mode endpoints get the messages flattened
into a labelled transcript. Settings' /models fetch moved to
list_endpoint_models() and is shared with /api/chat/config.

Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub
resolve_provider_config — the point is to exercise the real BYOK-vs-demo
decision and assert on what the provider actually received: off-whitelist
override pinned, off-whitelist Settings.model pinned, redirected endpoint
pinned, BYOK passed through untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
This commit is contained in:
parththakkar106
2026-07-26 20:16:56 +05:30
co-authored by Claude Opus 5
parent 57c24a07d8
commit 1dd31086c1
17 changed files with 969 additions and 20 deletions
+4 -2
View File
@@ -60,8 +60,10 @@ AIDND_DEMO_MODELS=
# Successful AI turns per user per day on the demo key. Default: 20 # Successful AI turns per user per day on the demo key. Default: 20
AIDND_DEMO_TURNS_PER_DAY= AIDND_DEMO_TURNS_PER_DAY=
# Comma-separated emails of "power users" (trusted testers) who bypass the daily # Comma-separated emails of "power users" (trusted testers) who bypass the daily
# demo cap entirely — unmetered turns on the shared demo key. Registered accounts # demo cap entirely — unmetered turns on the shared demo key — and get the AI Chat
# only (guests have no email). Matched case-insensitively. # page (a plain scratchpad for talking to a model, hidden from everyone else).
# Registered accounts only (guests have no email). Matched case-insensitively.
# Local (single-user) installs are always treated as power users.
AIDND_POWER_USERS= AIDND_POWER_USERS=
# The AI endpoint/API key/model are NOT env vars — they are configured at # The AI endpoint/API key/model are NOT env vars — they are configured at
+41 -5
View File
@@ -81,19 +81,50 @@ def demo_enabled() -> bool:
@dataclass @dataclass
class ProviderConfig: class ProviderConfig:
"""What the turn engine should actually connect with, after the """What the turn engine should actually connect with, after the
BYOK-vs-demo decision.""" BYOK-vs-demo decision. Build these with resolve_provider_config()."""
endpoint_url: str endpoint_url: str
api_key: str api_key: str
model: str model: str
using_demo: bool using_demo: bool
def __post_init__(self) -> None:
# Belt and braces around the shared demo key. resolve_provider_config()
# already pins the model, but this makes it a property of the config
# object itself: however it was built, and by whichever caller, the
# server-funded key can never be paired with an off-whitelist (i.e.
# possibly paid) model. Unreachable by design — a 500 here means a new
# code path tried to bypass the pinning, which is worth failing loudly
# rather than silently billing.
if DEMO_API_KEY and self.api_key == DEMO_API_KEY and self.model not in DEMO_MODELS:
raise ValueError(
f"Refusing to use the shared demo key with non-whitelisted model {self.model!r}"
)
def resolve_provider_config(settings: models.Settings) -> ProviderConfig:
def resolve_provider_config(
settings: models.Settings, *, model_override: str | None = None
) -> ProviderConfig:
"""BYOK when the user has their own key, the shared demo key otherwise.
THE security-relevant branch is the demo one, and it is the only place the
whitelist rule lives — every caller must come through here rather than
building a ProviderConfig itself. On the demo key:
- the model is pinned to DEMO_MODELS, so a caller-supplied override (the AI
Chat page) or a hand-edited Settings row cannot aim a server-funded key at
a paid model; anything unrecognised falls back to DEMO_MODELS[0];
- the endpoint is pinned to DEMO_ENDPOINT_URL, so the key itself can't be
redirected to a URL the user controls and harvested.
`model_override` is a per-request preference (never a grant): it's honoured
verbatim under BYOK, and only if whitelisted on the demo key.
"""
key = settings.api_key_plain key = settings.api_key_plain
requested = (model_override or "").strip() or settings.model
if key or not demo_enabled(): if key or not demo_enabled():
return ProviderConfig(settings.endpoint_url, key, settings.model, False) return ProviderConfig(settings.endpoint_url, key, requested, False)
model = settings.model if settings.model in DEMO_MODELS else DEMO_MODELS[0] model = requested if requested in DEMO_MODELS else DEMO_MODELS[0]
return ProviderConfig(DEMO_ENDPOINT_URL, DEMO_API_KEY, model, True) return ProviderConfig(DEMO_ENDPOINT_URL, DEMO_API_KEY, model, True)
@@ -102,7 +133,12 @@ def _today() -> str:
def is_power_user(user: models.User) -> bool: def is_power_user(user: models.User) -> bool:
"""Trusted testers (email allowlist) bypass the demo turn cap.""" """Trusted testers: unmetered demo turns, plus tooling that isn't part of
the game (the AI Chat scratchpad). Local installs are always trusted — it's
the operator's own machine and their own API key, same reasoning as the
provider debug log being local-only."""
if not MULTI_USER:
return True
return bool(user.email) and user.email.lower() in POWER_USERS return bool(user.email) and user.email.lower() in POWER_USERS
+1
View File
@@ -25,6 +25,7 @@ from . import auth, models
# scope -> (max requests, window seconds) # scope -> (max requests, window seconds)
RATE_LIMITS: dict[str, tuple[int, int]] = { RATE_LIMITS: dict[str, tuple[int, int]] = {
"turn": (10, 60), # AI turn generation (demo key also has a daily cap) "turn": (10, 60), # AI turn generation (demo key also has a daily cap)
"chat": (30, 60), # AI Chat scratchpad (power users only)
"script-test": (30, 60), # sandboxed, but each run costs up to 2s CPU "script-test": (30, 60), # sandboxed, but each run costs up to 2s CPU
"connection-test": (10, 60), # outbound HTTP to a user-supplied URL "connection-test": (10, 60), # outbound HTTP to a user-supplied URL
"import": (30, 60), # large writes "import": (30, 60), # large writes
+2 -1
View File
@@ -10,7 +10,7 @@ from .auth import MULTI_USER
from .database import engine from .database import engine
from .limits import BodySizeLimitMiddleware from .limits import BodySizeLimitMiddleware
from .migrations import bootstrap from .migrations import bootstrap
from .routers import adventures, auth, debug, scenarios, scripts, settings, story_cards from .routers import adventures, auth, chat, debug, scenarios, scripts, settings, story_cards
from .seed import seed_public_scenarios from .seed import seed_public_scenarios
bootstrap(engine) bootstrap(engine)
@@ -88,6 +88,7 @@ app.include_router(adventures.router)
app.include_router(story_cards.router) app.include_router(story_cards.router)
app.include_router(scripts.router) app.include_router(scripts.router)
app.include_router(settings.router) app.include_router(settings.router)
app.include_router(chat.router)
app.include_router(debug.router) app.include_router(debug.router)
+6
View File
@@ -54,6 +54,12 @@ _running: set[int] = set()
_tasks: set[asyncio.Task] = set() _tasks: set[asyncio.Task] = set()
# BYOK-only by construction: both factories below take the user's own
# endpoint/key straight from Settings and never auth.DEMO_*, so summarization
# and embedding can't spend the shared demo key (their call sites are also
# skipped when using_demo). Don't "fix" this by passing a ProviderConfig in —
# summary_model/embedding_model are free-form user input and are not on the
# demo whitelist.
def summary_provider(settings: models.Settings) -> OpenAICompatibleProvider: def summary_provider(settings: models.Settings) -> OpenAICompatibleProvider:
return OpenAICompatibleProvider( return OpenAICompatibleProvider(
settings.endpoint_url, settings.endpoint_url,
@@ -10,6 +10,17 @@ from .base import PromptParts, Provider, ProviderError
# continuing prose instead of replying conversationally. # continuing prose instead of replying conversationally.
CHAT_CONTINUE_HINT = "\n\n[Continue the story directly. Output only story text.]" CHAT_CONTINUE_HINT = "\n\n[Continue the story directly. Output only story text.]"
# Completion endpoints have no roles, so a plain chat has to be flattened into
# one labelled transcript that trails off on "Assistant:" for the model to continue.
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
def flatten_messages(messages: list[dict]) -> str:
turns = "\n\n".join(
f"{_ROLE_LABELS.get(m['role'], m['role'])}: {m['content']}" for m in messages
)
return f"{turns}\n\nAssistant:"
class OpenAICompatibleProvider(Provider): class OpenAICompatibleProvider(Provider):
"""Adapter for any /v1-style endpoint: Ollama, LM Studio, OpenAI, OpenRouter, vLLM, Groq…""" """Adapter for any /v1-style endpoint: Ollama, LM Studio, OpenAI, OpenRouter, vLLM, Groq…"""
@@ -110,7 +121,42 @@ class OpenAICompatibleProvider(Provider):
if not self.model: if not self.model:
raise ProviderError("No model configured — set one in Settings.") raise ProviderError("No model configured — set one in Settings.")
url, body = self._request(parts, temperature, max_tokens) url, body = self._request(parts, temperature, max_tokens)
async for event in self._stream(url, body):
yield event
async def chat(
self, messages: list[dict], *, temperature: float, max_tokens: int
) -> AsyncIterator[tuple[str, str]]:
"""Plain multi-turn chat — no story framing, no context assembly. Takes
[{"role", "content"}, ...] straight to the endpoint. Used by the AI Chat
scratchpad; the turn engine uses generate()."""
if not self.model:
raise ProviderError("No model configured — set one in Settings.")
if self.api_mode == "completion":
url = f"{self.base_url}/completions"
body = {
"model": self.model,
"prompt": flatten_messages(messages),
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
}
else:
url = f"{self.base_url}/chat/completions"
body = {
"model": self.model,
"messages": messages,
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
}
self._apply_reasoning_budget(body)
async for event in self._stream(url, body):
yield event
async def _stream(self, url: str, body: dict) -> AsyncIterator[tuple[str, str]]:
"""Shared SSE plumbing for generate()/chat(): POST a streaming request
and yield ("text" | "reasoning", chunk) pairs, logging the exchange."""
log = debuglog.start_entry(url, self.model, body) log = debuglog.start_entry(url, self.model, body)
received: list[str] = [] received: list[str] = []
try: try:
+2
View File
@@ -32,6 +32,8 @@ def me_payload(user: models.User, db: Session) -> dict:
"id": user.id, "id": user.id,
"email": user.email, "email": user.email,
"is_guest": user.is_guest, "is_guest": user.is_guest,
# Trusted testers: unmetered demo turns, plus the AI Chat scratchpad.
"power_user": auth.is_power_user(user),
"demo": { "demo": {
"enabled": auth.demo_enabled(), "enabled": auth.demo_enabled(),
"using_demo": cfg.using_demo, "using_demo": cfg.using_demo,
+159
View File
@@ -0,0 +1,159 @@
"""AI Chat — a plain scratchpad for talking to a model directly.
Power users only (the AIDND_POWER_USERS email allowlist). Deliberately thin:
no story context, no scripts, no world state, and nothing persisted — the
conversation lives in the browser and is posted up whole on each turn. It
exists to poke at models, prompts and endpoints without starting an adventure.
Model choice is free-form when the user brought their own API key. On the
shared demo key it stays pinned to the AIDND_DEMO_MODELS whitelist, exactly as
turns are: the server funds that key, so it must not be able to reach paid
models by way of this page.
"""
from fastapi import APIRouter, Depends, HTTPException, Request
from fastapi.responses import StreamingResponse
from sqlalchemy.orm import Session
from .. import auth, limits, models, schemas
from ..database import get_db
from ..providers import OpenAICompatibleProvider, ProviderError
from .adventures import SSE_HEADERS, sse
from .settings import get_settings, list_endpoint_models
router = APIRouter(prefix="/api/chat", tags=["chat"])
def power_user(
db: Session = Depends(get_db),
user: models.User = Depends(auth.get_current_user),
) -> models.User:
"""Gate for the whole router. 404 rather than 403 so the feature simply
doesn't appear to exist for everyone else."""
if not auth.is_power_user(user):
raise HTTPException(404, "Not found")
return user
PowerUser = Depends(power_user)
def _resolve_model(
settings: models.Settings, requested: str | None
) -> tuple[auth.ProviderConfig, str | None]:
"""Provider config for this chat, plus a note when the requested model was
not honoured. The pinning rule itself lives in resolve_provider_config —
this only reports the substitution it made, so there is exactly one place
that decides what the demo key is allowed to talk to."""
cfg = auth.resolve_provider_config(settings, model_override=requested)
wanted = (requested or "").strip()
if wanted and wanted != cfg.model:
return cfg, (
f"'{wanted}' isn't available on the shared demo key — using "
f"{cfg.model}. Add your own API key in Settings to use any model."
)
return cfg, None
@router.get("/config")
async def chat_config(
db: Session = Depends(get_db),
user: models.User = PowerUser,
):
"""What this page can talk to: the resolved endpoint/model, whether model
choice is pinned to the demo whitelist, and the endpoint's model listing
(best effort — an unreachable endpoint just yields an empty list)."""
settings = get_settings(db, user)
cfg = auth.resolve_provider_config(settings)
listing = await list_endpoint_models(cfg)
return {
"endpoint_url": cfg.endpoint_url,
"model": cfg.model,
"using_demo": cfg.using_demo,
"api_mode": settings.api_mode,
"temperature": settings.temperature,
"max_tokens": settings.max_output_tokens,
# On the demo key the whitelist IS the list of choices; otherwise it's
# whatever the endpoint advertises (suggestions, not a restriction).
"models": auth.DEMO_MODELS if cfg.using_demo else listing.get("models", []),
"models_error": None if listing.get("ok") else listing.get("detail"),
}
async def run_chat(cfg: auth.ProviderConfig, settings: models.Settings, payload: schemas.ChatRequest,
note: str | None, db: Session, user: models.User):
"""SSE generator mirroring the turn stream's event shape: reasoning/chunk
while generating, then done — so the frontend reuses the same plumbing."""
if note:
yield sse({"type": "note", "detail": note})
provider = OpenAICompatibleProvider(
cfg.endpoint_url, cfg.api_key, cfg.model, settings.api_mode,
settings.reasoning_max_tokens,
)
messages = [m.model_dump() for m in payload.messages]
chunks: list[str] = []
reasoning_chunks: list[str] = []
try:
async for kind, chunk in provider.chat(
messages,
temperature=payload.temperature if payload.temperature is not None else settings.temperature,
max_tokens=payload.max_tokens or settings.max_output_tokens,
):
if kind == "reasoning":
reasoning_chunks.append(chunk)
yield sse({"type": "reasoning", "text": chunk})
else:
chunks.append(chunk)
yield sse({"type": "chunk", "text": chunk})
except ProviderError as exc:
yield sse({"type": "error", "detail": str(exc)})
return
text = "".join(chunks).strip()
if not text:
detail = (
"The model used its entire token budget on reasoning and returned no "
"reply — raise max tokens, cap the reasoning budget in Settings, or "
"use a non-reasoning model."
if reasoning_chunks
else "The AI returned an empty response."
)
yield sse({"type": "error", "detail": detail})
return
if cfg.using_demo:
# Unmetered for power users (count_demo_turn is a no-op for them), but
# keep the call so the accounting stays right if the gate ever widens.
auth.count_demo_turn(user)
db.commit()
yield sse({
"type": "done",
"text": text,
"reasoning": "".join(reasoning_chunks).strip() or None,
"model": cfg.model,
})
@router.post("/stream")
def chat_stream(
payload: schemas.ChatRequest,
request: Request,
db: Session = Depends(get_db),
user: models.User = PowerUser,
):
total = sum(len(m.content) for m in payload.messages)
if total > schemas.CHAT_TOTAL_MAX:
raise HTTPException(
413, f"This conversation is too long to send ({total:,} characters) — "
"clear it or start a new one."
)
limits.rate_limit("chat", request, user)
settings = get_settings(db, user)
cfg, note = _resolve_model(settings, payload.model)
if not cfg.model:
raise HTTPException(400, "No model configured — set one in Settings or pick one here.")
return StreamingResponse(
run_chat(cfg, settings, payload, note, db, user),
media_type="text/event-stream",
headers=SSE_HEADERS,
)
+16 -12
View File
@@ -62,18 +62,9 @@ def update_settings(
return settings return settings
@router.post("/test") async def list_endpoint_models(cfg: auth.ProviderConfig) -> dict:
async def test_connection( """GET the endpoint's /models listing. Doubles as a connectivity check, so
request: Request, failures come back as {"ok": False, "detail": ...} rather than raising."""
db: Session = Depends(get_db),
user: models.User = Depends(auth.get_current_user),
):
"""Hit the endpoint's /models listing as a cheap connectivity check.
Tests whatever the turn engine would actually use — including the shared
demo endpoint when the user has no key of their own."""
limits.rate_limit("connection-test", request, user)
settings = get_settings(db, user)
cfg = auth.resolve_provider_config(settings)
url = cfg.endpoint_url.rstrip("/") + "/models" url = cfg.endpoint_url.rstrip("/") + "/models"
headers = {} headers = {}
if cfg.api_key: if cfg.api_key:
@@ -94,3 +85,16 @@ async def test_connection(
except (ValueError, AttributeError, TypeError): except (ValueError, AttributeError, TypeError):
pass # non-JSON or unexpected shape — connectivity is still confirmed pass # non-JSON or unexpected shape — connectivity is still confirmed
return {"ok": True, "models": models_available} return {"ok": True, "models": models_available}
@router.post("/test")
async def test_connection(
request: Request,
db: Session = Depends(get_db),
user: models.User = Depends(auth.get_current_user),
):
"""Cheap connectivity check against whatever the turn engine would actually
use — including the shared demo endpoint when the user has no key."""
limits.rate_limit("connection-test", request, user)
settings = get_settings(db, user)
return await list_endpoint_models(auth.resolve_provider_config(settings))
+22
View File
@@ -319,6 +319,28 @@ class SettingsOut(ORMModel):
ScenarioOut.model_rebuild() ScenarioOut.model_rebuild()
# ---------- AI Chat (power users) ----------
# A scratchpad for talking to a model directly, with no story framing. Nothing
# is persisted server-side, so these caps are purely per-request abuse limits.
CHAT_MESSAGE_MAX = 100_000 # one message
CHAT_TOTAL_MAX = 400_000 # whole conversation sent up per request
CHAT_MESSAGES_MAX = 200 # turns per request
class ChatMessage(BaseModel):
role: Literal["system", "user", "assistant"]
content: Annotated[str, Field(max_length=CHAT_MESSAGE_MAX)]
class ChatRequest(BaseModel):
messages: Annotated[list[ChatMessage], Field(min_length=1, max_length=CHAT_MESSAGES_MAX)]
# Empty/omitted = fall back to the user's configured model.
model: Name | None = None
temperature: Annotated[float, Field(ge=0, le=5)] | None = None
max_tokens: Annotated[int, Field(ge=1, le=100_000)] | None = None
class SettingsUpdate(BaseModel): class SettingsUpdate(BaseModel):
endpoint_url: Annotated[str, Field(max_length=500)] | None = None # VARCHAR(500) endpoint_url: Annotated[str, Field(max_length=500)] | None = None # VARCHAR(500)
# Encryption expands the stored value ~4/3 into the same VARCHAR(500): # Encryption expands the stored value ~4/3 into the same VARCHAR(500):
+226
View File
@@ -0,0 +1,226 @@
"""HTTP tests for the AI Chat scratchpad (power users only).
Covers the access gate, the streamed reply, and the demo-key model pinning —
the part that must not let a public visitor reach paid models through this page.
python -m pytest tests/test_chat.py -v
"""
import os
import tempfile
_tmp = tempfile.NamedTemporaryFile(suffix=".db", delete=False)
_tmp.close()
os.environ["AIDND_DB_PATH"] = _tmp.name
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import chat
class FakeProvider:
"""Records what it was constructed with, then streams a fixed reply. Stands
in for the real egress point, so asserting on last_key/last_model is
asserting on exactly what would have gone over the wire."""
last_model = None
last_key = None
last_endpoint = None
last_messages = None
def __init__(self, endpoint_url, api_key, model, api_mode="chat", reasoning_max_tokens=0):
FakeProvider.last_model = model
FakeProvider.last_key = api_key
FakeProvider.last_endpoint = endpoint_url
async def chat(self, messages, *, temperature, max_tokens):
FakeProvider.last_messages = messages
yield ("reasoning", "hmm")
yield ("text", "Hello back.")
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="power@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(user_id=user.id, api_key="enc:dummy", model="test-model"))
setup.commit()
user_id = user.id
setup.close()
monkeypatch.setattr(chat, "OpenAICompatibleProvider", FakeProvider)
monkeypatch.setattr(limits, "rate_limit", lambda *a, **k: None)
# Multi-user mode is what makes the power-user gate meaningful (local mode
# trusts everyone); the allowlist is set per-test.
monkeypatch.setattr(auth, "MULTI_USER", True)
monkeypatch.setattr(auth, "POWER_USERS", {"power@example.com"})
# These tests deliberately do NOT stub resolve_provider_config: the point is
# to exercise the real BYOK-vs-demo decision, since that is what keeps the
# shared key off paid models. Each test picks a mode with _byok/_demo below.
def _current_user(db=Depends(get_db)):
return db.get(models.User, user_id)
app.dependency_overrides[auth.get_current_user] = _current_user
try:
yield TestClient(app)
finally:
app.dependency_overrides.clear()
Base.metadata.drop_all(bind=engine)
def _send(client, **extra):
return client.post("/api/chat/stream", json={"messages": [{"role": "user", "content": "hi"}], **extra})
def _byok(monkeypatch):
"""The user brought their own key: no demo key in play, any model allowed."""
monkeypatch.setattr(auth, "demo_enabled", lambda: False)
db = SessionLocal()
try:
settings = db.query(models.Settings).first()
settings.api_key = "sk-my-own-key" # legacy-plaintext path: used as-is
db.commit()
finally:
db.close()
def _demo(monkeypatch, whitelist=("free/allowed",)):
"""The user has no key, so turns run on the server-funded demo key."""
monkeypatch.setattr(auth, "demo_enabled", lambda: True)
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_ENDPOINT_URL", "http://demo")
monkeypatch.setattr(auth, "DEMO_MODELS", list(whitelist))
def test_non_power_user_gets_404(client, monkeypatch):
monkeypatch.setattr(auth, "POWER_USERS", set())
assert _send(client).status_code == 404
assert client.get("/api/chat/config").status_code == 404
def test_power_user_streams_a_reply(client, monkeypatch):
_byok(monkeypatch)
resp = _send(client)
assert resp.status_code == 200, resp.text
assert '"type": "reasoning"' in resp.text
assert "Hello back." in resp.text
assert '"type": "done"' in resp.text
assert FakeProvider.last_messages == [{"role": "user", "content": "hi"}]
def test_system_prompt_and_model_override_are_honoured(client, monkeypatch):
_byok(monkeypatch)
resp = client.post("/api/chat/stream", json={
"messages": [
{"role": "system", "content": "Be terse."},
{"role": "user", "content": "hi"},
],
"model": "some/other-model",
})
assert resp.status_code == 200, resp.text
# BYOK: any model the user names is passed straight through, on their key.
assert FakeProvider.last_model == "some/other-model"
assert FakeProvider.last_key == "sk-my-own-key"
assert FakeProvider.last_messages[0] == {"role": "system", "content": "Be terse."}
def test_demo_key_pins_model_to_whitelist(client, monkeypatch):
_demo(monkeypatch)
resp = _send(client, model="expensive/paid-model")
assert resp.status_code == 200, resp.text
# Refused visibly: the whitelisted model runs instead, with a note. The
# paid slug must never reach the wire alongside the server-funded key.
assert FakeProvider.last_model == "free/allowed"
assert FakeProvider.last_key == "demo-key"
assert '"type": "note"' in resp.text
# A whitelisted model is still selectable on the demo key.
_demo(monkeypatch, ["free/allowed", "free/second"])
resp = _send(client, model="free/second")
assert resp.status_code == 200, resp.text
assert FakeProvider.last_model == "free/second"
def test_demo_key_ignores_an_off_whitelist_settings_model(client, monkeypatch):
"""The override isn't the only untrusted input — Settings.model is user-set
too, and it must be pinned the same way when there's no BYOK key."""
_demo(monkeypatch)
db = SessionLocal()
try:
db.query(models.Settings).first().model = "expensive/paid-model"
db.commit()
finally:
db.close()
resp = _send(client)
assert resp.status_code == 200, resp.text
assert FakeProvider.last_model == "free/allowed"
def test_demo_key_endpoint_cannot_be_redirected(client, monkeypatch):
"""A user-controlled endpoint_url would leak the key itself, which is worse
than spending it — the demo branch pins the URL too."""
_demo(monkeypatch)
db = SessionLocal()
try:
db.query(models.Settings).first().endpoint_url = "http://attacker.example/v1"
db.commit()
finally:
db.close()
assert _send(client).status_code == 200
assert FakeProvider.last_endpoint == "http://demo"
assert FakeProvider.last_key == "demo-key"
def test_provider_config_refuses_demo_key_with_paid_model(monkeypatch):
"""The structural backstop: even a hand-built config (a future code path
that forgets to go through resolve_provider_config) can't pair them."""
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_MODELS", ["free/allowed"])
with pytest.raises(ValueError):
auth.ProviderConfig("http://demo", "demo-key", "expensive/paid-model", True)
# Mislabelling it as non-demo doesn't help: the key is what's checked.
with pytest.raises(ValueError):
auth.ProviderConfig("http://demo", "demo-key", "expensive/paid-model", False)
# The user's own key with any model stays fine.
auth.ProviderConfig("http://any", "sk-mine", "expensive/paid-model", False)
def test_resolve_provider_config_is_the_single_choke_point(monkeypatch):
"""Turns, AI Chat and the connection test all resolve through this one
function, so pinning it here pins every caller. No DB or HTTP needed."""
monkeypatch.setattr(auth, "demo_enabled", lambda: True)
monkeypatch.setattr(auth, "DEMO_API_KEY", "demo-key")
monkeypatch.setattr(auth, "DEMO_ENDPOINT_URL", "http://demo")
monkeypatch.setattr(auth, "DEMO_MODELS", ["free/allowed"])
# No key of their own: endpoint AND model are pinned, whatever they set.
no_key = models.Settings(endpoint_url="http://mine/v1", api_key="", model="expensive/paid")
assert auth.resolve_provider_config(no_key) == auth.ProviderConfig(
"http://demo", "demo-key", "free/allowed", True)
assert auth.resolve_provider_config(
no_key, model_override="expensive/paid").model == "free/allowed"
assert auth.resolve_provider_config(
no_key, model_override="free/allowed").model == "free/allowed"
# Their own key: their endpoint, their key, their choice of model.
byok = models.Settings(endpoint_url="http://mine/v1", api_key="sk-mine", model="expensive/paid")
assert auth.resolve_provider_config(byok) == auth.ProviderConfig(
"http://mine/v1", "sk-mine", "expensive/paid", False)
def test_oversized_conversation_is_refused(client, monkeypatch):
_byok(monkeypatch)
huge = "x" * 90_000
resp = client.post("/api/chat/stream", json={
"messages": [{"role": "user", "content": huge} for _ in range(5)],
})
assert resp.status_code == 413, resp.text
+6
View File
@@ -58,6 +58,12 @@ export default function App() {
<NavLink to="/settings" className={({ isActive }) => `navlink${isActive ? ' active' : ''}`}> <NavLink to="/settings" className={({ isActive }) => `navlink${isActive ? ' active' : ''}`}>
Settings Settings
</NavLink> </NavLink>
{/* Power-user tooling, not part of the game — hidden for everyone else. */}
{me?.power_user && (
<NavLink to="/chat" className={({ isActive }) => `navlink${isActive ? ' active' : ''}`}>
AI Chat
</NavLink>
)}
{me?.multi_user && ( {me?.multi_user && (
<div className="nav-account"> <div className="nav-account">
{me.is_guest ? ( {me.is_guest ? (
+4
View File
@@ -138,6 +138,10 @@ export const api = {
exportStoryCards: (owner) => request(`/story-cards/export?${new URLSearchParams(owner)}`), exportStoryCards: (owner) => request(`/story-cards/export?${new URLSearchParams(owner)}`),
importStoryCards: (payload) => request('/story-cards/import', { method: 'POST', body: JSON.stringify(payload) }), importStoryCards: (payload) => request('/story-cards/import', { method: 'POST', body: JSON.stringify(payload) }),
// AI Chat (power users only — 404 for everyone else)
getChatConfig: () => request('/chat/config'),
chatStream: (payload, handlers, signal) => streamSSE('/chat/stream', payload, handlers, signal),
// Debug // Debug
getDebugRequests: () => request('/debug/requests'), getDebugRequests: () => request('/debug/requests'),
+99
View File
@@ -1624,6 +1624,98 @@ button.primary.compact { padding: 3px 12px; font-size: 0.76rem; margin-left: aut
.icon-row button.active { border-color: var(--accent); box-shadow: 0 0 10px var(--accent-glow); } .icon-row button.active { border-color: var(--accent); box-shadow: 0 0 10px var(--accent-glow); }
.art-hint { font-size: 0.76rem; color: var(--text-dim); line-height: 1.45; margin: 0; } .art-hint { font-size: 0.76rem; color: var(--text-dim); line-height: 1.45; margin: 0; }
/* ---------- AI Chat (power users) ----------
A plain chat scratchpad, styled as a transcript rather than as the story
page: UI font, role labels, and a sticky composer at the bottom. */
.chat-page { max-width: 820px; display: flex; flex-direction: column; }
.chat-meta {
font-size: 0.78rem;
margin: -10px 0 16px;
word-break: break-word;
}
.chat-options {
border: 1px solid var(--border);
border-radius: 10px;
padding: 14px 16px 4px;
margin-bottom: 18px;
background: var(--bg-panel);
}
.chat-transcript { flex: 1; display: flex; flex-direction: column; gap: 16px; padding-bottom: 20px; }
.chat-msg {
border: 1px solid var(--border);
border-radius: 10px;
padding: 10px 14px 12px;
background: var(--bg-panel);
}
.chat-msg.user { border-left: 3px solid var(--player); background: var(--bg-input); }
.chat-msg.assistant { border-left: 3px solid var(--accent-dim); }
.chat-msg-head {
display: flex;
align-items: center;
gap: 10px;
margin-bottom: 6px;
font-size: 0.72rem;
letter-spacing: 0.14em;
text-transform: uppercase;
}
.chat-msg .chat-role { color: var(--accent-dim); font-weight: 600; }
.chat-msg.user .chat-role { color: var(--player); }
.chat-model-tag { text-transform: none; letter-spacing: 0; font-size: 0.74rem; }
.chat-msg-actions {
margin-left: auto;
display: flex;
gap: 12px;
opacity: 0;
transition: opacity 0.15s;
}
.chat-msg:hover .chat-msg-actions, .chat-msg:focus-within .chat-msg-actions { opacity: 1; }
.chat-msg-body {
white-space: pre-wrap;
overflow-wrap: anywhere;
line-height: 1.6;
font-size: 0.95rem;
}
.chat-msg .cursor {
color: var(--accent-bright);
text-shadow: 0 0 8px var(--accent-glow);
animation: blink 1s steps(1) infinite;
}
/* Same collapsed thinking block as the story page, re-scoped for the transcript. */
.chat-msg .reasoning { margin: 0 0 8px; font-size: 0.82rem; line-height: 1.5; color: var(--text-dim); }
.chat-msg .reasoning summary { cursor: pointer; user-select: none; opacity: 0.8; }
.chat-msg .reasoning summary:hover { opacity: 1; }
.chat-msg .reasoning .reasoning-text {
margin-top: 6px;
padding: 8px 12px;
border-left: 2px solid var(--border);
white-space: pre-wrap;
opacity: 0.85;
max-height: 260px;
overflow-y: auto;
}
.chat-composer {
position: sticky;
bottom: 0;
display: flex;
gap: 10px;
align-items: flex-end;
padding: 14px 0;
background: linear-gradient(to top, var(--bg) 55%, transparent);
}
.chat-composer textarea {
flex: 1;
min-height: 44px;
max-height: 220px;
resize: none;
overflow-y: auto;
}
.chat-composer button { flex: none; height: 40px; }
@media (prefers-reduced-motion: reduce) { @media (prefers-reduced-motion: reduce) {
.enter, .skeleton-card, .sk, .toast, .thinking i, .enter, .skeleton-card, .sk, .toast, .thinking i,
.story .action, .page { .story .action, .page {
@@ -1793,4 +1885,11 @@ button.primary.compact { padding: 3px 12px; font-size: 0.76rem; margin-left: aut
font-size: 2.5rem; font-size: 2.5rem;
padding: 4px 8px 0 0; padding: 4px 8px 0 0;
} }
/* ---------- AI Chat ---------- */
/* No hover on touch, so the per-message copy/delete links stay visible. */
.chat-msg-actions { opacity: 1; }
.chat-msg { padding: 9px 11px 11px; }
.chat-composer { padding: 10px 0 14px; }
.chat-composer textarea { max-height: 140px; }
} }
+3
View File
@@ -10,6 +10,7 @@ import Play from './pages/Play.jsx'
import Scripts from './pages/Scripts.jsx' import Scripts from './pages/Scripts.jsx'
import ScriptEditor from './pages/ScriptEditor.jsx' import ScriptEditor from './pages/ScriptEditor.jsx'
import Settings from './pages/Settings.jsx' import Settings from './pages/Settings.jsx'
import Chat from './pages/Chat.jsx'
import './index.css' import './index.css'
const router = createBrowserRouter([ const router = createBrowserRouter([
@@ -25,6 +26,8 @@ const router = createBrowserRouter([
{ path: 'scripts', element: <Scripts /> }, { path: 'scripts', element: <Scripts /> },
{ path: 'scripts/:id', element: <ScriptEditor /> }, { path: 'scripts/:id', element: <ScriptEditor /> },
{ path: 'settings', element: <Settings /> }, { path: 'settings', element: <Settings /> },
// Power users only — the page redirects home and the API 404s otherwise.
{ path: 'chat', element: <Chat /> },
], ],
}, },
]) ])
+326
View File
@@ -0,0 +1,326 @@
/* AI Chat — a plain scratchpad for talking to a model, with none of the game's
context assembly in the way. Power users only (the backend 404s the routes
for everyone else, and the nav link is hidden).
Deliberately client-side: the conversation lives in localStorage, not the
database. Nothing here is part of an adventure, so there's nothing worth a
migration — and a refresh still keeps what you were poking at. */
import { useCallback, useEffect, useRef, useState } from 'react'
import { useNavigate, useOutletContext } from 'react-router-dom'
import { api } from '../api'
import { useToast } from '../components'
const STORAGE_KEY = 'aidnd.chat.v1'
function load() {
try {
const saved = JSON.parse(localStorage.getItem(STORAGE_KEY) || '{}')
return {
messages: Array.isArray(saved.messages) ? saved.messages : [],
system: typeof saved.system === 'string' ? saved.system : '',
model: typeof saved.model === 'string' ? saved.model : '',
temperature: saved.temperature ?? '',
maxTokens: saved.maxTokens ?? '',
}
} catch {
return { messages: [], system: '', model: '', temperature: '', maxTokens: '' }
}
}
const ROLE_LABEL = { user: 'You', assistant: 'AI', system: 'System' }
function ReasoningBlock({ text, streaming }) {
if (!text) return null
return (
<details className="reasoning" open={streaming || undefined}>
<summary>💭 Reasoning{streaming ? '…' : ''}</summary>
<div className="reasoning-text">{text}</div>
</details>
)
}
function Message({ message, onDelete }) {
const [copied, setCopied] = useState(false)
const copy = () => {
navigator.clipboard?.writeText(message.content).then(
() => { setCopied(true); setTimeout(() => setCopied(false), 1500) },
() => {},
)
}
return (
<div className={`chat-msg ${message.role}`}>
<div className="chat-msg-head">
<span className="chat-role">{ROLE_LABEL[message.role] || message.role}</span>
{message.model && <span className="dim chat-model-tag">{message.model}</span>}
<span className="chat-msg-actions">
<button className="linklike" onClick={copy}>{copied ? 'copied' : 'copy'}</button>
<button className="linklike" onClick={onDelete}>delete</button>
</span>
</div>
<ReasoningBlock text={message.reasoning} />
<div className="chat-msg-body">{message.content}</div>
</div>
)
}
export default function Chat() {
const { me } = useOutletContext() ?? {}
const navigate = useNavigate()
const toast = useToast()
const initial = useRef(load()).current
const [messages, setMessages] = useState(initial.messages)
const [system, setSystem] = useState(initial.system)
const [model, setModel] = useState(initial.model)
const [temperature, setTemperature] = useState(initial.temperature)
const [maxTokens, setMaxTokens] = useState(initial.maxTokens)
const [input, setInput] = useState('')
const [config, setConfig] = useState(null)
const [showOptions, setShowOptions] = useState(false)
// Streaming reply in progress: null when idle, else the text so far ('' before
// the first token). `busy` covers the whole request, including the wait.
const [streaming, setStreaming] = useState(null)
const [reasoningStream, setReasoningStream] = useState(null)
const [busy, setBusy] = useState(false)
const abortRef = useRef(null)
const inputRef = useRef(null)
const pinnedRef = useRef(true)
// me is null until /auth/me resolves; only bounce once we know.
useEffect(() => {
if (me && !me.power_user) navigate('/', { replace: true })
}, [me, navigate])
useEffect(() => {
api.getChatConfig().then(setConfig).catch(() => setConfig(null))
}, [])
useEffect(() => {
localStorage.setItem(
STORAGE_KEY,
JSON.stringify({ messages, system, model, temperature, maxTokens }),
)
}, [messages, system, model, temperature, maxTokens])
// Grow the composer with its content (CSS caps the height, then it scrolls).
useEffect(() => {
const el = inputRef.current
if (!el) return
el.style.height = 'auto'
el.style.height = `${el.scrollHeight}px`
}, [input])
useEffect(() => {
const onScroll = () => {
pinnedRef.current =
window.innerHeight + window.scrollY >= document.documentElement.scrollHeight - 120
}
window.addEventListener('scroll', onScroll, { passive: true })
return () => window.removeEventListener('scroll', onScroll)
}, [])
useEffect(() => {
if (pinnedRef.current) window.scrollTo({ top: document.documentElement.scrollHeight })
}, [messages, streaming, reasoningStream])
// Abort any in-flight stream when leaving the page.
useEffect(() => () => abortRef.current?.abort(), [])
const send = useCallback(async (history) => {
const controller = new AbortController()
abortRef.current = controller
setBusy(true)
setStreaming('')
setReasoningStream(null)
pinnedRef.current = true
const payload = {
messages: [
...(system.trim() ? [{ role: 'system', content: system.trim() }] : []),
...history.map(({ role, content }) => ({ role, content })),
],
}
if (model.trim()) payload.model = model.trim()
if (temperature !== '' && temperature !== null) payload.temperature = Number(temperature)
if (maxTokens !== '' && maxTokens !== null) payload.max_tokens = Number(maxTokens)
let reasoning = ''
try {
await api.chatStream(payload, (event) => {
if (event.type === 'chunk') {
setStreaming((prev) => (prev ?? '') + event.text)
} else if (event.type === 'reasoning') {
reasoning += event.text
setReasoningStream((prev) => (prev ?? '') + event.text)
} else if (event.type === 'note') {
toast(event.detail)
} else if (event.type === 'done') {
setMessages((prev) => [...prev, {
role: 'assistant',
content: event.text,
reasoning: event.reasoning || reasoning || undefined,
model: event.model,
}])
} else if (event.type === 'error') {
toast(event.detail, 'error')
}
}, controller.signal)
} catch (err) {
if (err.name !== 'AbortError') toast(err.message, 'error')
} finally {
abortRef.current = null
setBusy(false)
setStreaming(null)
setReasoningStream(null)
}
}, [system, model, temperature, maxTokens, toast])
const submit = () => {
const text = input.trim()
if (!text || busy) return
const history = [...messages, { role: 'user', content: text }]
setMessages(history)
setInput('')
send(history)
}
const regenerate = () => {
if (busy) return
// Drop trailing assistant replies and re-send from the last user message.
let history = [...messages]
while (history.length && history[history.length - 1].role === 'assistant') history.pop()
if (!history.length) return
setMessages(history)
send(history)
}
const stop = () => {
abortRef.current?.abort()
// Keep whatever streamed in — a cut-off reply is often the thing you wanted.
const partial = streaming?.trim()
if (partial) {
setMessages((prev) => [...prev, {
role: 'assistant',
content: partial,
reasoning: reasoningStream || undefined,
model: config?.model,
stopped: true,
}])
}
}
const clear = () => {
if (busy || !messages.length) return
setMessages([])
toast('Conversation cleared')
}
const deleteAt = (index) => setMessages((prev) => prev.filter((_, i) => i !== index))
const onKeyDown = (e) => {
if (e.key === 'Enter' && !e.shiftKey) {
e.preventDefault()
submit()
}
}
if (me && !me.power_user) return null
const waitingForFirstToken = streaming === '' && reasoningStream === null
const canRegenerate = !busy && messages.some((m) => m.role === 'user')
return (
<div className="page chat-page">
<div className="page-header">
<h1>AI Chat</h1>
<div style={{ display: 'flex', gap: 10, alignItems: 'center' }}>
<button className="linklike" onClick={() => setShowOptions((o) => !o)}>
{showOptions ? 'hide options' : 'options'}
</button>
<button onClick={regenerate} disabled={!canRegenerate}>Regenerate</button>
<button className="danger" onClick={clear} disabled={busy || !messages.length}>Clear</button>
</div>
</div>
<div className="chat-meta dim">
{config
? <>
{/* The override wins when set, so show what a send would actually use. */}
{model.trim() || config.model || '(no model set)'} · {config.endpoint_url}
{config.using_demo && ' · shared demo key (whitelisted models only)'}
{config.api_mode === 'completion' && ' · completion mode (messages are flattened)'}
</>
: 'Loading provider config…'}
</div>
{showOptions && (
<div className="chat-options">
<label className="field">
<span className="label">System prompt (sent first, every turn — empty = none)</span>
<textarea rows={3} value={system} placeholder="You are a helpful assistant."
onChange={(e) => setSystem(e.target.value)} />
</label>
<div style={{ display: 'flex', gap: 14, flexWrap: 'wrap' }}>
<label className="field" style={{ flex: '2 1 240px' }}>
<span className="label">Model {config?.using_demo ? '(demo whitelist)' : '(empty = Settings default)'}</span>
<input type="text" list="chat-models" value={model} placeholder={config?.model || 'model slug'}
onChange={(e) => setModel(e.target.value)} />
<datalist id="chat-models">
{(config?.models || []).map((m) => <option key={m} value={m} />)}
</datalist>
</label>
<label className="field" style={{ flex: '1 1 110px' }}>
<span className="label">Temperature</span>
<input type="number" step="0.1" min="0" max="5" value={temperature}
placeholder={config?.temperature ?? ''}
onChange={(e) => setTemperature(e.target.value)} />
</label>
<label className="field" style={{ flex: '1 1 130px' }}>
<span className="label">Max tokens</span>
<input type="number" min="1" value={maxTokens}
placeholder={config?.max_tokens ?? ''}
onChange={(e) => setMaxTokens(e.target.value)} />
</label>
</div>
{config?.models_error && (
<div className="dim" style={{ fontSize: '0.82rem' }}>
Couldn't list models from the endpoint: {config.models_error}
</div>
)}
</div>
)}
<div className="chat-transcript">
{!messages.length && streaming === null && (
<div className="empty">
Nothing here yet — no story, no scripts, no world state. Just you and the model.
</div>
)}
{messages.map((m, i) => (
<Message key={i} message={m} onDelete={() => deleteAt(i)} />
))}
{streaming !== null && (
<div className="chat-msg assistant">
<div className="chat-msg-head">
<span className="chat-role">AI</span>
<span className="dim chat-model-tag">{model.trim() || config?.model}</span>
</div>
<ReasoningBlock text={reasoningStream} streaming />
{waitingForFirstToken
? <div className="thinking" role="status"><i /><i /><i /><span>Thinking</span></div>
: <div className="chat-msg-body">{streaming}<span className="cursor">▋</span></div>}
</div>
)}
</div>
<div className="chat-composer">
<textarea ref={inputRef} rows={1} value={input} onChange={(e) => setInput(e.target.value)}
onKeyDown={onKeyDown} placeholder="Message the model… (Enter to send, Shift+Enter for a new line)" />
{busy
? <button className="danger" onClick={stop}>Stop</button>
: <button className="primary" onClick={submit} disabled={!input.trim()}>Send</button>}
</div>
</div>
)
}
+6
View File
@@ -45,4 +45,10 @@ services:
- key: AIDND_DEMO_TURNS_PER_DAY - key: AIDND_DEMO_TURNS_PER_DAY
value: "20" value: "20"
# Comma-separated emails of trusted testers: unmetered demo turns, plus
# the AI Chat page (hidden from everyone else). Set in the dashboard;
# leave unset and nobody gets it. Registered accounts only.
- key: AIDND_POWER_USERS
sync: false
# CORS is unset on purpose: the SPA is served same-origin by FastAPI. # CORS is unset on purpose: the SPA is served same-origin by FastAPI.