61 lines
2.9 KiB
Markdown
61 lines
2.9 KiB
Markdown
# ADR 002 — Ollama Is the v1 Model Backend
|
|
|
|
**Status:** Accepted; transport/TLS consequence added after M1
|
|
|
|
## Decision
|
|
|
|
v1 will target Ollama inference running on user-controlled local infrastructure. The default endpoint is same-host loopback, but v1 must also support an explicitly configured Ollama instance on a trusted local-area network.
|
|
|
|
## Context
|
|
|
|
The intended production deployment can eventually run the storyteller and Ollama on one machine, but development and testing may place Ollama on a separate machine on the user's LAN. The project prioritizes local control, privacy, predictable integration, and no dependency on Internet/cloud inference.
|
|
|
|
## Alternatives Considered
|
|
|
|
- multiple cloud providers,
|
|
- LM Studio,
|
|
- llama.cpp direct integration,
|
|
- arbitrary OpenAI-compatible endpoints,
|
|
- Ollama.
|
|
|
|
## Reason
|
|
|
|
Ollama is already available locally, provides a simple local API, supports both text-generation and embedding models, and avoids requiring external inference services.
|
|
|
|
## Consequences
|
|
|
|
- candidate forks supporting multiple cloud providers should be simplified or hardened,
|
|
- candidate projects using another local API need an adapter,
|
|
- the storyteller must support both same-host Ollama and an explicitly configured trusted-LAN Ollama endpoint,
|
|
- LAN inference does not imply LAN exposure of the storyteller UI/API; the storyteller should still bind to loopback by default,
|
|
- arbitrary public Internet/cloud model endpoints remain outside normal v1 configuration,
|
|
- future backend abstraction may be added, but v1 should not be delayed to support it.
|
|
|
|
## Transport for a Trusted-LAN Endpoint
|
|
|
|
Added after M1. Same-host Ollama speaks plain HTTP over loopback, and it was
|
|
assumed a LAN endpoint would look the same. It does not have to.
|
|
|
|
A trusted-LAN Ollama may be served over **HTTPS with a certificate issued by a
|
|
private or local CA** rather than a public one — a self-hosted server that
|
|
terminates TLS for everything it exposes is the ordinary case, not an exotic
|
|
one, and it may offer no cleartext port at all. A v1 client that trusts only a
|
|
bundled public-CA list cannot talk to such a host, while `curl` and the user's
|
|
browser on the same machine can.
|
|
|
|
Therefore:
|
|
|
|
- production clients must verify against the **operating system's trusted CA
|
|
store** in addition to any bundled certificate list, so a CA the user has
|
|
installed on their own machine is honoured by this application too;
|
|
- certificate **and hostname** verification remain fully enabled;
|
|
- there must be **no "ignore TLS errors" / "insecure" option**, in the UI,
|
|
in configuration, or as an environment variable. A LAN endpoint the machine
|
|
does not trust is a configuration problem to fix at the OS level, not a check
|
|
to switch off;
|
|
- the endpoint URL must therefore accept `https://` on any port, not only
|
|
`http://…:11434`.
|
|
|
|
M1 implemented this (`backend/app/tlstrust.py`); see
|
|
`planning/reports/M1-IMPLEMENTATION-REPORT.md` §G.
|