The post-implementation review of M2, plus the three corrections it took
to make the evidence true. Reports:
planning/reports/M2-BASELINE-REPORT.md 868 lines, the measurements
planning/reports/M2-IMPLEMENTATION-REPORT.md 758 lines, the reading of them
Verdict is PASS, accept with non-blocking debt, proceed to M3. Every M2
requirement is met and the ones that matter were tested by running the
build rather than reading it: a cloud endpoint written straight into
SQLite with sqlite3, behind the API's back, still refused at the wire;
trusted-LAN HTTPS against the real second machine with verification on;
captures showing zero packets outside loopback and the approved host.
Three defects, all found by running the shipped image.
The memory bank was dead. M2 removed Settings.api_key_plain with the API
key, and memorybank's two provider factories still read it. It failed
inside a fire-and-forget task, so no user error, no log anyone would
read, and no test — every memory test stubs those factories. All 604
tests passed with summaries and embeddings silently not happening.
The configurable model timeout never reached the turn engine. Stored,
validated, exposed in the API, rendered in the UI, and not passed to the
provider. M2's own exit criterion was half met: the constant had moved
but the setting did nothing.
And requirements.lock still pinned quickjs, psycopg and cryptography, so
the setup path DEVELOPMENT.md gives a new developer would have
reinstalled all three.
Both code defects now have the test that would have caught them: one
constructs every provider factory from a real Settings row, one drives
the turn endpoint, the chat endpoint and the summariser and asserts the
configured timeout arrives at each. That is the lesson worth keeping from
this milestone — after removing an attribute, build each consumer from a
real object; after adding a setting, prove it lands. Both failures were
in background or plumbing paths, which is exactly where a subtractive
change cannot see itself.
606 tests pass, up from 604. Lint, build and image are clean. Every
runtime result in the baseline report came from an image built after
these fixes; the reports say plainly that commit 8c65ae9 itself does not
contain them.
Also recorded: 88 test node IDs disappeared and every one is accounted
for — 64 whole files whose subject was removed, 5 replaced by a better
file, 15 individually retired with their features, and 4 renames. No
meaningful coverage was lost, and the eight files that used a JavaScript
counter as instrumentation kept their assertions by moving the counter to
the world-state engine.
Six planning recommendations are reported, not applied. Three are marked
before M3: the threat model still describes the inherited SSRF guard's
opposite rule, TECHNICAL-DESIGN §5.1 still marks two hardening items
open, and the endpoint policy is a load-bearing security decision that
exists only as a module docstring and deserves an ADR.
M3 is clear to start. Its chokepoints are untouched or simplified — the
rollback paths now carry one shared state instead of two — and the Phase
0B undo/redo spike still applies. No M3 work here: Undo still deletes and
there is still no Redo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL
128 lines
4.6 KiB
Python
128 lines
4.6 KiB
Python
"""AI Chat: a plain scratchpad for talking to the configured model directly.
|
|
|
|
Deliberately thin. It adds no story context and no world state, and it persists
|
|
nothing. The conversation lives in the browser and is posted in full on each
|
|
turn. It exists for checking a model, a prompt, or an endpoint without starting
|
|
an adventure — which is exactly the kind of thing a local single-user install
|
|
wants a page for.
|
|
|
|
Upstream gated this behind a "power user" email allowlist and pinned the model
|
|
when a shared demo key was in play. M2 removed both: there is one local user,
|
|
who owns the endpoint, and there is no server-funded key to protect. The model
|
|
this page talks to is the one in Settings, or one the user names per request —
|
|
either way it is their own Ollama.
|
|
"""
|
|
|
|
from fastapi import APIRouter, Depends, HTTPException
|
|
from fastapi.responses import StreamingResponse
|
|
from sqlalchemy.orm import Session
|
|
|
|
from .. import auth, models, schemas
|
|
from ..database import get_db
|
|
from ..providers import OpenAICompatibleProvider, ProviderError
|
|
from ..sse import SSE_HEADERS, sse
|
|
from .settings import get_settings, list_endpoint_models
|
|
|
|
router = APIRouter(prefix="/api/chat", tags=["chat"])
|
|
|
|
|
|
@router.get("/config")
|
|
async def chat_config(
|
|
db: Session = Depends(get_db),
|
|
user: models.User = Depends(auth.get_current_user),
|
|
):
|
|
"""Returns what this page can talk to.
|
|
|
|
The model listing is best effort: an unreachable endpoint returns an empty
|
|
list and the reason, rather than failing the page.
|
|
"""
|
|
settings = get_settings(db, user)
|
|
listing = await list_endpoint_models(settings.endpoint_url)
|
|
return {
|
|
"endpoint_url": settings.endpoint_url,
|
|
"model": settings.model,
|
|
"api_mode": settings.api_mode,
|
|
"temperature": settings.temperature,
|
|
"max_tokens": settings.max_output_tokens,
|
|
# Suggestions from the endpoint, not a restriction.
|
|
"models": listing.get("models", []),
|
|
"models_error": None if listing.get("ok") else listing.get("detail"),
|
|
}
|
|
|
|
|
|
async def run_chat(
|
|
settings: models.Settings, model: str, payload: schemas.ChatRequest
|
|
):
|
|
"""Streams the reply as SSE, using the turn stream's event shape.
|
|
|
|
The generator emits `reasoning` and `chunk` events while generating and then
|
|
a `done` event, so the frontend reuses the same code.
|
|
"""
|
|
provider = OpenAICompatibleProvider(
|
|
settings.endpoint_url, model, settings.api_mode,
|
|
settings.model_timeout_seconds,
|
|
)
|
|
messages = [m.model_dump() for m in payload.messages]
|
|
chunks: list[str] = []
|
|
reasoning_chunks: list[str] = []
|
|
try:
|
|
async for kind, chunk in provider.chat(
|
|
messages,
|
|
temperature=(
|
|
payload.temperature
|
|
if payload.temperature is not None
|
|
else settings.temperature
|
|
),
|
|
max_tokens=payload.max_tokens or settings.max_output_tokens,
|
|
):
|
|
if kind == "reasoning":
|
|
reasoning_chunks.append(chunk)
|
|
yield sse({"type": "reasoning", "text": chunk})
|
|
else:
|
|
chunks.append(chunk)
|
|
yield sse({"type": "chunk", "text": chunk})
|
|
except ProviderError as exc:
|
|
yield sse({"type": "error", "detail": str(exc)})
|
|
return
|
|
|
|
text = "".join(chunks).strip()
|
|
if not text:
|
|
detail = (
|
|
"The model used its entire token budget on reasoning and returned no "
|
|
"reply — raise max tokens or use a non-reasoning model."
|
|
if reasoning_chunks
|
|
else "The AI returned an empty response."
|
|
)
|
|
yield sse({"type": "error", "detail": detail})
|
|
return
|
|
|
|
yield sse({
|
|
"type": "done",
|
|
"text": text,
|
|
"reasoning": "".join(reasoning_chunks).strip() or None,
|
|
"model": model,
|
|
})
|
|
|
|
|
|
@router.post("/stream")
|
|
def chat_stream(
|
|
payload: schemas.ChatRequest,
|
|
db: Session = Depends(get_db),
|
|
user: models.User = Depends(auth.get_current_user),
|
|
):
|
|
total = sum(len(m.content) for m in payload.messages)
|
|
if total > schemas.CHAT_TOTAL_MAX:
|
|
raise HTTPException(
|
|
413, f"This conversation is too long to send ({total:,} characters) — "
|
|
"clear it or start a new one."
|
|
)
|
|
settings = get_settings(db, user)
|
|
model = (payload.model or "").strip() or settings.model
|
|
if not model:
|
|
raise HTTPException(400, "No model configured — set one in Settings or pick one here.")
|
|
return StreamingResponse(
|
|
run_chat(settings, model, payload),
|
|
media_type="text/event-stream",
|
|
headers=SSE_HEADERS,
|
|
)
|