Docs: correct post-M3 status and Ollama configuration

Three stale claims found in active documentation after the post-M3
consolidation:

- BUILD-MILESTONES.md still opened "M1 and M2 complete; M3 next", which
  contradicted its own M3 "Status: COMPLETE" block, planning/README.md and
  VERSION.md. It now states M1, M2 and M3 complete and accepted, M4 next.
- DEVELOPMENT.md described the settings row as "endpoint, model and (unused)
  API key" and sent "api_key":"" in its curl example. M2 removed the field;
  schemas.SettingsUpdate has no api_key. The prose and the example now match
  the real request shape. No application code was changed.
- README.md listed LM Studio as a supported local endpoint. ADR 002 and ADR 011
  make Ollama the only v1 backend; LM Studio is a rejected alternative there.
  The row is removed and the surrounding wording now says Ollama is the
  supported backend, same-host is the default, trusted-LAN Ollama is supported,
  public/cloud is prohibited, and the OpenAI-compatible adapter is an
  implementation detail rather than a support promise. The claude_shim section
  stays, relabelled "(development only)".

Documentation only; no code, schema or test changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
This commit is contained in:
JesseMarkowitz
2026-09-03 16:39:29 -04:00
co-authored by Claude Opus 5
parent d27ee34901
commit 3c8e91f644
3 changed files with 17 additions and 11 deletions
+5 -4
View File
@@ -91,15 +91,16 @@ decision that this project's threat model does not cover
## Pointing the storyteller at Ollama ## Pointing the storyteller at Ollama
The endpoint, model and (unused) API key are **runtime settings stored in the The endpoint, the model and the generation parameters are **runtime settings
database**, not environment variables. Set them on the app's Settings page, or stored in the database**, not environment variables. There is no API key field:
with one request: M2 removed it along with the cloud providers, and Ollama does not use one. Set
them on the app's Settings page, or with one request:
```bash ```bash
curl -X PUT http://127.0.0.1:8000/api/settings \ curl -X PUT http://127.0.0.1:8000/api/settings \
-H 'Content-Type: application/json' \ -H 'Content-Type: application/json' \
-d '{"endpoint_url":"http://127.0.0.1:11434/v1","model":"qwen2.5:3b-instruct", -d '{"endpoint_url":"http://127.0.0.1:11434/v1","model":"qwen2.5:3b-instruct",
"api_mode":"chat","api_key":"","max_output_tokens":200, "api_mode":"chat","max_output_tokens":200,
"context_token_budget":4096}' "context_token_budget":4096}'
``` ```
+11 -6
View File
@@ -121,18 +121,23 @@ Open http://localhost:5173.
## Connect a model ## Connect a model
Open **Settings** in the app and point it at a local Ollama-compatible endpoint: Ollama is the inference backend v1 supports. Open **Settings** in the app and point it at one:
| Where the model runs | Endpoint URL | Notes | | Where Ollama runs | Endpoint URL | Notes |
|---|---|---| |---|---|---|
| Ollama, same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works | | The same machine | `http://localhost:11434/v1` | the default, and the simplest thing that works |
| Ollama, a machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | see below | | A machine on your own network | `http://<host>:11434/v1` or `https://<host>/v1` | explicitly configured; see below |
| LM Studio, same machine | `http://localhost:1234/v1` | |
Model name, generation parameters, and (optionally) summary and embedding models for the Model name, generation parameters, and (optionally) summary and embedding models for the
Memory Bank are configured there too. No config files and no rebuild are needed. There is no Memory Bank are configured there too. No config files and no rebuild are needed. There is no
API key field, because there is nothing to authenticate to. API key field, because there is nothing to authenticate to.
The adapter underneath speaks the OpenAI-compatible protocol, because that is what Ollama
serves. That is an implementation detail, not a promise of support for arbitrary local
servers that happen to speak the same protocol. Public and cloud inference endpoints are
prohibited outright — see `planning/DECISIONS/002-ollama-only-v1.md` and
`planning/DECISIONS/011-local-inference-endpoint-policy.md`.
### What the endpoint policy allows ### What the endpoint policy allows
The address is checked when you save it and again before every request. Only loopback and The address is checked when you save it and again before every request. Only loopback and
@@ -146,7 +151,7 @@ machine serves HTTPS with a certificate from a CA you installed, it works: certi
verified against your operating system's trust store as well as the bundled one. Verification verified against your operating system's trust store as well as the bundled one. Verification
itself is never relaxed, and there is no option to turn it off. itself is never relaxed, and there is no option to turn it off.
### Playing against a local shim ### Playing against a local shim (development only)
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787` `backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint on `127.0.0.1:8787`
backed by a command-line tool, which is useful for testing the turn engine against a stronger backed by a command-line tool, which is useful for testing the turn engine against a stronger
+1 -1
View File
@@ -1,6 +1,6 @@
# Adventure Storyteller — Production Build Milestones # Adventure Storyteller — Production Build Milestones
**Status:** In implementation. M1 and M2 complete (2026-09-02); M3 next **Status:** In implementation. M1, M2 and M3 complete and accepted (M1 and M2: 2026-09-02; M3: 2026-09-03); M4 — Named Save Points / Checkpoints — next
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` **Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
## 1. Purpose ## 1. Purpose