Files
interactive-story/planning/DECISIONS/002-ollama-only-v1.md

2.9 KiB

ADR 002 — Ollama Is the v1 Model Backend

Status: Accepted; transport/TLS consequence added after M1

Decision

v1 will target Ollama inference running on user-controlled local infrastructure. The default endpoint is same-host loopback, but v1 must also support an explicitly configured Ollama instance on a trusted local-area network.

Context

The intended production deployment can eventually run the storyteller and Ollama on one machine, but development and testing may place Ollama on a separate machine on the user's LAN. The project prioritizes local control, privacy, predictable integration, and no dependency on Internet/cloud inference.

Alternatives Considered

  • multiple cloud providers,
  • LM Studio,
  • llama.cpp direct integration,
  • arbitrary OpenAI-compatible endpoints,
  • Ollama.

Reason

Ollama is already available locally, provides a simple local API, supports both text-generation and embedding models, and avoids requiring external inference services.

Consequences

  • candidate forks supporting multiple cloud providers should be simplified or hardened,
  • candidate projects using another local API need an adapter,
  • the storyteller must support both same-host Ollama and an explicitly configured trusted-LAN Ollama endpoint,
  • LAN inference does not imply LAN exposure of the storyteller UI/API; the storyteller should still bind to loopback by default,
  • arbitrary public Internet/cloud model endpoints remain outside normal v1 configuration,
  • future backend abstraction may be added, but v1 should not be delayed to support it.

Transport for a Trusted-LAN Endpoint

Added after M1. Same-host Ollama speaks plain HTTP over loopback, and it was assumed a LAN endpoint would look the same. It does not have to.

A trusted-LAN Ollama may be served over HTTPS with a certificate issued by a private or local CA rather than a public one — a self-hosted server that terminates TLS for everything it exposes is the ordinary case, not an exotic one, and it may offer no cleartext port at all. A v1 client that trusts only a bundled public-CA list cannot talk to such a host, while curl and the user's browser on the same machine can.

Therefore:

  • production clients must verify against the operating system's trusted CA store in addition to any bundled certificate list, so a CA the user has installed on their own machine is honoured by this application too;
  • certificate and hostname verification remain fully enabled;
  • there must be no "ignore TLS errors" / "insecure" option, in the UI, in configuration, or as an environment variable. A LAN endpoint the machine does not trust is a configuration problem to fix at the OS level, not a check to switch off;
  • the endpoint URL must therefore accept https:// on any port, not only http://…:11434.

M1 implemented this (backend/app/tlstrust.py); see planning/reports/M1-IMPLEMENTATION-REPORT.md §G.