diff --git a/README.md b/README.md index 91427ea..fd88f13 100644 --- a/README.md +++ b/README.md @@ -76,7 +76,9 @@ isn't the live one starts a new branch.* session cookie), can register (email and password) at any point to keep their data, and each user gets isolated data plus their own encrypted-at-rest API key. A server-funded **shared demo key** with a daily turn cap lets people try it without bringing a key - (`backend/app/auth.py`). + (`backend/app/auth.py`). Each new guest is also given a copy of a short pre-played + adventure, so the first screen shows real turns and their world-state changes without + spending a demo turn (`backend/app/starter.py`). ## Screenshots @@ -127,10 +129,36 @@ Open **Settings** in the app and point it at any OpenAI-compatible endpoint: | LM Studio (local) | `http://localhost:1234/v1` | free, private | | OpenRouter | `https://openrouter.ai/api/v1` | `:free` models cost nothing (no embeddings on the free tier) | | OpenAI / Groq / vLLM / … | provider's `/v1` URL | anything speaking `/v1/chat/completions` | +| Claude Code CLI (local) | `http://127.0.0.1:8787/v1` | your Claude subscription instead of an API key; see [Playing against Claude locally](#playing-against-claude-locally) | Model name, API key, generation parameters, and (optionally) summary and embedding models for the Memory Bank are all configured there too. No config files and no rebuild are needed. +### Playing against Claude locally + +`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by the +`claude` command line tool, so you can play the demos against a real model without an +API key. Each request spawns one `claude --print` process, which suits the turn engine: +the app assembles the whole prompt every turn and expects a stateless endpoint. + +```sh +cd backend +.venv/Scripts/python.exe tools/claude_shim.py # listens on 127.0.0.1:8787 +``` + +In Settings, choose the OpenAI-compatible provider, set the base URL to +`http://127.0.0.1:8787/v1`, put any non-empty string in the API key field, and pick +`sonnet`. The shim ignores the key and authenticates as you, through the CLI. Set the +reasoning budget to `0` or `-1`: a positive budget sends a `reasoning.max_tokens` field +that Claude 5 models reject. + +Embeddings are not served. Leave the embedding model blank, or point the Memory Bank at +a real endpoint. + +Run it against a local backend only. The endpoint has no authentication, and anything +reaching it spends your Claude quota. `app/netguard.py` blocks localhost endpoints when +`AIDND_MULTI_USER` is set, so a deployed instance cannot be pointed at it. + ## How a turn works ``` @@ -174,7 +202,7 @@ development, Vite proxies `/api` to FastAPI. ## Tests -497 backend tests: unit tests plus full HTTP integration through the real quickjs scripting +539 backend tests: unit tests plus full HTTP integration through the real quickjs scripting engine, with the LLM provider mocked. CI runs them on every push, alongside the frontend lint/build and a Docker image build. diff --git a/backend/tools/claude_shim.py b/backend/tools/claude_shim.py new file mode 100644 index 0000000..c23e3a8 --- /dev/null +++ b/backend/tools/claude_shim.py @@ -0,0 +1,272 @@ +"""An OpenAI-compatible endpoint backed by the local `claude` CLI. + +Point AI-DnD's Settings at `http://127.0.0.1:8787/v1` and the turn engine talks +to Claude through your Claude Code subscription instead of an API key. Use it to +play the demos against a real model when you do not want to spend API credit, +and to check prompt changes against the model the deployed app actually meets. +The shim speaks the small part of the OpenAI protocol the app uses: a model +listing, and `chat/completions` in both streaming and non-streaming form. + +Each request spawns one `claude --print` process. The CLI holds no conversation +state here, which suits AI-DnD, because the app assembles the whole prompt every +turn and expects a stateless endpoint. + +Embeddings are not served. Leave the embedding model blank in Settings, or point +the memory bank at a real endpoint. + +To use it: + +1. Start the shim from the backend virtualenv, which already has FastAPI and + uvicorn: + + backend/.venv/Scripts/python.exe tools/claude_shim.py + +2. In Settings, set the provider to OpenAI-compatible, the base URL to + `http://127.0.0.1:8787/v1`, and the API key to any non-empty string. The + shim ignores the key. It authenticates as you, through the CLI. + +3. Pick `sonnet` as the model, then run the connection test. + +Set the reasoning budget to `0` or `-1`. A positive budget sends a +`reasoning.max_tokens` field that Claude 5 models reject with a 400. + +Options: + + --port Port to listen on. Defaults to 8787. + --claude Path to the `claude` executable. Defaults to `$AIDND_CLAUDE_BIN`, + then to whatever is on `PATH`, then to the per-user install + under `~/.local/bin`. + +Run this only against a local backend. The endpoint has no authentication, and +anything that reaches it spends your Claude subscription quota. The SSRF guard +in `app/netguard.py` blocks localhost endpoints when `AIDND_MULTI_USER` is set, +so a deployed instance cannot be pointed here. +""" +import argparse +import asyncio +import json +import os +import shutil +import tempfile +import time +import uuid + +import uvicorn +from fastapi import FastAPI, Request +from fastapi.responses import JSONResponse, StreamingResponse + + +def find_claude() -> str: + """Returns the path to the `claude` executable. + + The per-user install is the last resort, because `PATH` is what a developer + controls. Windows needs the `.exe` spelled out, since `create_subprocess_exec` + does no extension search. + """ + override = os.environ.get("AIDND_CLAUDE_BIN") + if override: + return override + found = shutil.which("claude") + if found: + return found + for name in ("claude.exe", "claude"): + candidate = os.path.expanduser(os.path.join("~", ".local", "bin", name)) + if os.path.exists(candidate): + return candidate + return "claude" + + +CLAUDE = find_claude() +DEFAULT_PORT = 8787 + +# The listing the Settings page shows. The CLI accepts both the aliases and the +# full ids, so both are offered. +MODELS = [ + "opus", + "sonnet", + "haiku", + "claude-opus-5", + "claude-sonnet-5", + "claude-haiku-4-5", +] + +# Flags that strip the coding agent down to a text generator. `--system-prompt-file` +# replaces Claude Code's own system prompt, `--restricted` removes the tools that +# run commands, and the empty MCP config keeps the connected servers out of the +# tool list. An empty `--mcp-config` value is rejected: the key has to be present. +BASE_FLAGS = [ + "--print", + "--restricted", + "--strict-mcp-config", + "--mcp-config", '{"mcpServers": {}}', + "--disallowed-tools", + "Read Write Edit Glob Grep WebSearch WebFetch Task Skill ToolSearch", +] + +app = FastAPI() + + +@app.get("/v1/models") +async def list_models(): + return {"object": "list", "data": [{"id": m, "object": "model"} for m in MODELS]} + + +@app.post("/v1/embeddings") +async def embeddings(): + return JSONResponse( + {"error": {"message": "This shim does not serve embeddings. Leave the " + "embedding model blank, or point the memory bank " + "at a real endpoint."}}, + status_code=501, + ) + + +def split_prompt(messages: list[dict]) -> tuple[str, str]: + """Splits an OpenAI message list into a system prompt and a user prompt. + + The system messages become the CLI's `--system-prompt-file`. Everything else + is joined into one prompt sent on stdin. AI-DnD sends exactly one system and + one user message per turn, so the labeling only matters for the AI Chat page. + """ + system_parts, turn_parts = [], [] + for m in messages: + content = m.get("content") or "" + if isinstance(content, list): # Content-block form, if a caller uses it. + content = "".join(b.get("text", "") for b in content if isinstance(b, dict)) + if m.get("role") == "system": + system_parts.append(content) + elif len(messages) <= 2: + turn_parts.append(content) + else: + turn_parts.append(f"{m.get('role', 'user').title()}: {content}") + return "\n\n".join(system_parts), "\n\n".join(turn_parts) + + +async def run_claude(system: str, prompt: str, model: str): + """Yields `(kind, text)` pairs from one `claude --print` run. + + `kind` is `"text"` for narration, `"usage"` for the final token accounting, + and `"error"` for a failure the CLI reported in its result record. + """ + fd, sys_path = tempfile.mkstemp(suffix=".txt", text=True) + with os.fdopen(fd, "w", encoding="utf-8") as fh: + fh.write(system or "You are a helpful assistant.") + try: + args = [ + CLAUDE, *BASE_FLAGS, + "--system-prompt-file", sys_path, + "--model", model, + "--output-format", "stream-json", + "--include-partial-messages", + "--verbose", + ] + proc = await asyncio.create_subprocess_exec( + *args, + stdin=asyncio.subprocess.PIPE, + stdout=asyncio.subprocess.PIPE, + stderr=asyncio.subprocess.PIPE, + cwd=tempfile.gettempdir(), + ) + proc.stdin.write(prompt.encode("utf-8")) + await proc.stdin.drain() + proc.stdin.close() + + async for raw in proc.stdout: + line = raw.decode("utf-8", errors="replace").strip() + if not line: + continue + try: + rec = json.loads(line) + except ValueError: + continue + if rec.get("type") == "stream_event": + event = rec.get("event") or {} + if event.get("type") == "content_block_delta": + delta = event.get("delta") or {} + if delta.get("type") == "text_delta": + yield "text", delta.get("text", "") + elif rec.get("type") == "result": + if rec.get("is_error"): + yield "error", str(rec.get("result") or "claude failed") + usage = rec.get("usage") or {} + yield "usage", { + "prompt_tokens": (usage.get("input_tokens", 0) + + usage.get("cache_read_input_tokens", 0) + + usage.get("cache_creation_input_tokens", 0)), + "completion_tokens": usage.get("output_tokens", 0), + "total_tokens": 0, + "prompt_tokens_details": { + "cached_tokens": usage.get("cache_read_input_tokens", 0), + }, + "cost_usd": rec.get("total_cost_usd"), + } + await proc.wait() + if proc.returncode != 0: + detail = (await proc.stderr.read()).decode("utf-8", errors="replace") + yield "error", f"claude exited {proc.returncode}: {detail[:500]}" + finally: + os.unlink(sys_path) + + +@app.post("/v1/chat/completions") +async def chat_completions(request: Request): + body = await request.json() + model = body.get("model") or "sonnet" + system, prompt = split_prompt(body.get("messages") or []) + cid = f"chatcmpl-{uuid.uuid4().hex[:24]}" + created = int(time.time()) + + if not body.get("stream"): + text, usage, error = "", {}, None + async for kind, value in run_claude(system, prompt, model): + if kind == "text": + text += value + elif kind == "usage": + usage = value + else: + error = value + if error and not text: + return JSONResponse({"error": {"message": error}}, status_code=502) + return { + "id": cid, "object": "chat.completion", "created": created, "model": model, + "choices": [{"index": 0, "finish_reason": "stop", + "message": {"role": "assistant", "content": text}}], + "usage": usage, + } + + async def sse(): + def chunk(delta: dict, finish=None, usage=None) -> str: + payload = { + "id": cid, "object": "chat.completion.chunk", "created": created, + "model": model, + "choices": [{"index": 0, "delta": delta, "finish_reason": finish}], + } + if usage: + payload["usage"] = usage + return f"data: {json.dumps(payload)}\n\n" + + yield chunk({"role": "assistant", "content": ""}) + sent = False + async for kind, value in run_claude(system, prompt, model): + if kind == "text" and value: + sent = True + yield chunk({"content": value}) + elif kind == "usage": + yield chunk({}, finish="stop", usage=value) + elif kind == "error" and not sent: + yield chunk({"content": f"[shim error] {value}"}, finish="stop") + yield "data: [DONE]\n\n" + + return StreamingResponse(sse(), media_type="text/event-stream", + headers={"Cache-Control": "no-cache"}) + + +if __name__ == "__main__": + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--port", type=int, default=DEFAULT_PORT) + parser.add_argument("--claude", default=CLAUDE, help="Path to the claude executable.") + args = parser.parse_args() + CLAUDE = args.claude + print(f"claude binary: {CLAUDE}") + print(f"Point AI-DnD Settings at http://127.0.0.1:{args.port}/v1") + uvicorn.run(app, host="127.0.0.1", port=args.port, log_level="info") diff --git a/docs/GUIDE.md b/docs/GUIDE.md index 0140973..76fe36e 100644 --- a/docs/GUIDE.md +++ b/docs/GUIDE.md @@ -954,6 +954,14 @@ guest survives with no re-parenting and no migration step. Three kinds of row sh users table: local (email NULL, not guest), guest (email NULL, guest), registered (email set). +**Guests start with a story already in progress.** `starter.py` copies a shipped export +bundle into each new guest account at the same point the row is created. An empty account +gives a visitor nothing to read, and the daily demo turns are limited, so learning what +the app does used to cost one of them. The copy is the guest's own from the first moment: +they can edit, branch, delete, or export it, and nothing links it back to the file. The +guest row is committed before the copy is attempted, so a failure there still leaves them +with an account, and the copy itself runs inside a savepoint. + **Guests expire; accounts don't.** One row per curious visitor adds up, so `cleanup.py` deletes guests idle for `AIDND_GUEST_RETENTION_DAYS` (default 5), measured as `COALESCE(last_seen_at, created_at)`, because `_touch` only writes `last_seen_at` hourly diff --git a/plan/16-world-state-refusals.md b/plan/16-world-state-refusals.md index bbf69dc..88e7edd 100644 --- a/plan/16-world-state-refusals.md +++ b/plan/16-world-state-refusals.md @@ -231,3 +231,95 @@ old cap made the reset unreachable in one turn. **Not fixed:** the `initial == max` shape itself. It is now loud rather than silent, and the wording removes the usual cause, but the shape is still there. + +--- + +## Driven against Claude, 2026-08-28 + +The four turns above ran on the free demo model, which ignored two instructions +that were present and correct in the prompt. To separate model behavior from +code, the same demo was played again through `backend/tools/claude_shim.py`, an +OpenAI-compatible endpoint backed by the local `claude` command line tool. The +model was `sonnet`. See the README for how to run it. + +Every one of the five changes worked: + +| Turn | Chips | +|---|---| +| 1 | `turn +1`, `active pokemon → Wartortle`, `wartortle hp -8`, `milo active hp -46`, dashed `type advantage used refused — no such flag` | +| 2 | `turn +1`, `milo active hp -52`, `milo pokemon fainted +1`, `milo active pokemon → Onix`, `✓ graveler defeated` | +| 3 | `turn +1 (limited)`, dashed `milo active hp no change — at its limit` | +| 4 | `turn +1`, `milo active hp +32`, `✓ type advantage used` | + +`graveler_defeated` fired for the first time in five playtests, `world.turn` +moved every turn, `pokemon_fainted` went 0 to 1 at the faint, and every HP number +was a signed change rather than a total. The two failures left open above were +the demo model, not the instructions. + +**The refusal loop is verified end to end.** Turn 3 clamped to nothing. Turn 4's +assembled prompt, read from `GET /api/adventures/1/actions/9/context`, carried +the correction verbatim: + +``` +[Part of your last state block was not applied. Correct it in this turn's block: +- `npc.milo.active_hp` did not move. It is already at its minimum of 0 (it runs + from 0 to 98; it moves at most 98 per turn).] +``` + +The model's next delta was `{"world.turn": 1, "npc.milo.active_hp": 32, +"milestones.type_advantage_used": true}`. That is the loop the whole change +exists for, and it had never been observed running. + +### The one bug that is still a schema fault + +At the faint the model set `npc.milo.active_pokemon` to `Onix` and sent no +positive `active_hp` change, so Onix arrived on the field at 0 of 98 and every +later hit was refused. Sonnet followed the rest of the scenario closely, so this +is the schema rather than the model: `active_hp` is one stat with one `max` of +98, shared by Graveler at 98, Onix at 90, and Kabutops at 88. A switch has to +raise it, and nothing in the schema can raise it on the referee's side. + +The fix is to give Milo's three Pokemon their own NPC entries, which also removes +the `initial == max` shape from this demo. Not done: the session's remaining time +went to the guest starter instead. + +### Cost + +About $0.04 a turn, billed against the Claude subscription rather than a card. +Roughly 13k of each request's 20k prompt tokens is the command line tool's own +overhead; the app's prompt is about 7k. + +--- + +## What shipped alongside it, 2026-08-28 + +**A permanent local test rig.** `backend/tools/claude_shim.py` is now in the +repo. It finds the `claude` binary through `AIDND_CLAUDE_BIN`, then `PATH`, then +the per-user install, and takes `--port` and `--claude`. + +**The demo says Pokemon in its title, and has a Pokeball for cover art.** +`05-league-championship.json` is now `[Demo] Pokemon League Championship: Round +One`. Renaming a seed exposed a fault: `seed.py` matches a seed to its row by +title, so a rename inserted a second public scenario and stranded the first, +which is the same failure this file records for "Road to the Champion". Seeds +now carry `previous_titles`, and `_find_renamed` lands the rename on the +existing row. The scenario kept its id, and no orphan appeared. + +The cover art is a 3.2 kB PNG data URI. An SVG one was tried first and stored as +an empty `image_url`: `app/images.py` accepts raster formats only, on purpose, +because SVG can carry script and these bytes are served from the app's own +origin. `backend/tools/make_pokeball.py` draws the ball with `zlib` alone, so +regenerating it needs no image library. + +**Every new guest gets the played adventure.** `app/starter.py` copies a shipped +export bundle into each new guest account, from the guest mint in +`routers/auth.py`. The bundle is this session's adventure trimmed to its first +two exchanges, which ends on the knockout and shows an applied change, a refused +one, and a milestone. It stops before the Onix bug above, and the state it +leaves has Onix at its own 90 HP so a guest can play on from it. + +The row building that `POST /adventures/import` did inline moved into +`bundle.materialize`, which both callers now use. The limit and rate checks +stayed in the endpoint: the starter writes a file the server ships, so it has no +untrusted list to cap. `tests/test_starter_adventure.py` covers the file, the +copy, the chips, the playable end state, and the two failure paths. diff --git a/plan/STATUS.md b/plan/STATUS.md index f404cbc..d983b04 100644 --- a/plan/STATUS.md +++ b/plan/STATUS.md @@ -86,8 +86,12 @@ nothing still reads `variant_count`/`variant_index` before dropping them, and no is the migration shape that rewrites toasted values, so it owes one `VACUUM FULL actions;` on the direct (non-`-pooler`) endpoint afterwards. -Ahead of that, `plan/16-world-state-refusals.md` records a set of world-state fixes that -are merged but **never driven in a browser**. See that file for what to test. +Ahead of that, `plan/16-world-state-refusals.md` is closed. Its fixes were driven in a +browser on 2026-08-28, first on the demo model and then on Claude Sonnet through the +local shim, and all five changes work. One item in it is still open and is the next +world-state job: **Milo's three Pokemon share one `active_hp` stat**, so a switch leaves +the newcomer at 0 HP and every later hit is refused. Give them their own NPC entries, +which also removes the `initial == max` shape from that demo. **SP9 and SP10 are both on `main`, despite what earlier notes here said.** The branches `sp7b-take-pager` and `sp10-memory-bank-eviction` still exist and still read as unmerged @@ -206,6 +210,47 @@ an automated version; see SP7's entry in `plan/14`. --- +## What happened on 2026-08-28 — world state against a real model + +**The refusal loop ran end to end for the first time.** `plan/16` shipped a correction +that is injected into the next turn's prompt when the engine clamps or rejects a change, +and nothing had ever observed it working. Turn 3 of the Pokemon demo clamped to nothing, +turn 4's assembled prompt carried the note verbatim, and the model's next delta was +correct. The full record, with the four turns and their chips, is at the end of +`plan/16-world-state-refusals.md`. + +**Two failures blamed on the code were the model.** `graveler_defeated` never firing and +`world.turn` never moving both survived four playtests on the free demo model, with the +instructions present and correct in the assembled prompt. Both worked on the first try +against Sonnet. When an instruction is provably in the prompt, the next thing to change +is the model, not the wording. + +**There is now a permanent way to test against a real model.** +`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by the local +`claude` command line tool, so a demo can be played on a real model with no API key. +About $0.04 a turn against the subscription. The README documents it under **Connect a +model**. + +**Renaming a seed scenario used to strand its old row.** `seed.py` matches a seed file to +its scenario by title, so changing a title inserted a second public scenario and left the +first orphaned and public forever. This is the same failure this file already records for +"Road to the Champion". Seed files now carry `previous_titles`, and the rename lands on +the existing row. Verified: the Pokemon demo kept its id. + +**Every new guest is given a pre-played adventure.** `app/starter.py` copies a shipped +export bundle into each new guest account. The point is the first screen: real turns with +their world-state chips, including a refused change and a milestone, before spending any +of the daily demo turns. The row building inside `POST /adventures/import` moved to +`bundle.materialize` so both callers share one writer. + +**An SVG data URI is not usable as scenario art.** `app/images.py` accepts raster formats +only, deliberately, because SVG can carry script and the bytes are served from the app's +own origin. A rejected value is stored but yields an empty `image_url`, which fails +quietly. The Pokeball is a 3.2 kB PNG, drawn by `backend/tools/make_pokeball.py` with +`zlib` alone. + +--- + ## What happened on 2026-08-21 — visit analytics The hosted demo can now answer whether anyone is using it. `/analytics` is a dashboard — @@ -899,7 +944,7 @@ the SQLite dev parity this codebase protects on purpose). ``` cd backend -.venv/Scripts/python.exe -m pytest tests/ # 365 tests (~100s) +.venv/Scripts/python.exe -m pytest tests/ # 539 tests (~180s) .venv/Scripts/python.exe -m tools.stress_session # egress report (SQLite) # Same harness against a real Postgres. The target must be a THROWAWAY database @@ -945,6 +990,26 @@ holds the egress ceilings. On Windows the report's box-drawing characters crash the default cp1252 console; prefix with `PYTHONIOENCODING=utf-8`. +**A real model, without an API key.** `tools/claude_shim.py` serves an OpenAI-compatible +endpoint backed by the local `claude` command line tool. Point Settings at +`http://127.0.0.1:8787/v1`, put any string in the API key field, and pick `sonnet`. Set +the reasoning budget to `0` or `-1`: a positive budget sends `reasoning.max_tokens`, +which Claude 5 models reject with a 400. It costs about $0.04 a turn against the +subscription, and it does not serve embeddings. + +``` +cd backend +.venv/Scripts/python.exe tools/claude_shim.py # 127.0.0.1:8787 +``` + +**To exercise the guest path**, which is off in single-user mode: + +``` +cd backend +AIDND_MULTI_USER=1 AIDND_SECRET_KEY=throwaway AIDND_DB_PATH=$PWD/guest_scratch.db .venv/Scripts/python.exe -m uvicorn app.main:app --port 8001 +curl -c jar.txt http://127.0.0.1:8001/api/auth/me # mints a guest and its starter +``` + Port 8000 is shared with the job-pipeline app, which will squat it and silently shadow the AI-DnD API — free it before running the backend, or move the vite proxy with `AIDND_API_PORT`, which is what the `--keep` recipe above does.