Files
interactive-story/planning/DECISIONS/002-ollama-only-v1.md
T

33 lines
1.5 KiB
Markdown

# ADR 002 — Ollama Is the v1 Model Backend
**Status:** Accepted
## Decision
v1 will target Ollama inference running on user-controlled local infrastructure. The default endpoint is same-host loopback, but v1 must also support an explicitly configured Ollama instance on a trusted local-area network.
## Context
The intended production deployment can eventually run the storyteller and Ollama on one machine, but development and testing may place Ollama on a separate machine on the user's LAN. The project prioritizes local control, privacy, predictable integration, and no dependency on Internet/cloud inference.
## Alternatives Considered
- multiple cloud providers,
- LM Studio,
- llama.cpp direct integration,
- arbitrary OpenAI-compatible endpoints,
- Ollama.
## Reason
Ollama is already available locally, provides a simple local API, supports both text-generation and embedding models, and avoids requiring external inference services.
## Consequences
- candidate forks supporting multiple cloud providers should be simplified or hardened,
- candidate projects using another local API need an adapter,
- the storyteller must support both same-host Ollama and an explicitly configured trusted-LAN Ollama endpoint,
- LAN inference does not imply LAN exposure of the storyteller UI/API; the storyteller should still bind to loopback by default,
- arbitrary public Internet/cloud model endpoints remain outside normal v1 configuration,
- future backend abstraction may be added, but v1 should not be delayed to support it.