Files
interactive-story/planning/DECISIONS/011-local-inference-endpoint-policy.md
T
JesseMarkowitzandClaude Opus 5 d27ee34901 Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was
authoritative. Phase 0 execution prompts sat beside the specification; four
completed milestone reports sat beside the current one; and upstream AI-DnD's
own `plan/` build log and `docs/` project site still described a hosted,
scripted, multi-user product with accounts — every screenshot in it showed a
Scripts tab and a Sign up button, none of which has existed since M2.

`planning/archive/` now holds the history and says so in its own README:
`phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and
M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied.
`planning/reports/` holds only the current milestone's report, because that is
the one M4 planning has to read; it moves to the archive when M4's replaces it.

Deleted rather than archived: the Phase 0B execution prompts and the
handoff/status/summary documents, the Phase 0A discovery and triage reports,
upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's
template boilerplate. All of it is in Git history, and the two recommendation
reports carry every conclusion the deleted research reached.

Archived documents are kept verbatim. Paths written inside them point at where
those files were when the document was written, which is the point: an evidence
record that has been quietly edited is no longer evidence.

Active documentation is corrected where it pointed at the removed trees or
described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not
touch" list had gone stale at M2 and claimed QuickJS scripting was still tested;
its test count was 604 against an actual 638. `README.md` loses the upstream CI
badge, which reported upstream's pipeline rather than this fork's, and a
reference to `backend/app/worldstate/engine.py`, a file that does not exist.
`planning/README.md` is rewritten as the documentation index.

New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
manifest of what belongs in the ChatGPT project's Sources.

Source comments referring to the deleted trees are reworded; no behaviour
changes. 638 backend tests pass, frontend lints and builds, and a reference scan
over all 48 tracked Markdown files reports no unresolved path in active
documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
2026-09-03 14:33:07 -04:00

5.4 KiB

ADR 011 — Local Inference Endpoint Policy

Status: Accepted
Date: 2026-09-02
Implemented in: M2, backend/app/endpoints.py

Decision

v1 supports Ollama only, and will send a story to an inference endpoint only when every address that endpoint resolves to lies inside an explicit allowlist of local networks.

  • The storyteller UI/API remains loopback-bound by default. Configuring a LAN inference endpoint does not change where the storyteller listens.
  • Inference may use same-host loopback (the default) or an explicitly configured trusted-LAN/local-network address.
  • Public Internet addresses are denied.
  • The policy is address-based, using explicit allowed CIDRs rather than Python's generic is_private / is_reserved classifications.
  • Every resolved address must be allowed; one address outside the allowlist refuses the endpoint.
  • The policy is checked on configuration and again before every request.
  • TLS verification is mandatory for HTTPS and is never traded against this policy.
  • Arbitrary cloud / OpenAI-compatible endpoints are intentionally outside v1.

Context

The endpoint setting is the most consequential in the application. The storyteller sends the player's prose, the assembled context, the retrieved memories and the embedding inputs to whatever address it names: point it somewhere else and the whole campaign goes there.

The inherited AI-DnD guard could not be reused, because its rule is the opposite of this product's. AI-DnD was a hosted service, so its SSRF guard blocked private addresses to stop a user reaching the server's internal network. A local storyteller must do exactly the reverse — permit the private ranges and refuse the public Internet. The guard was removed rather than adapted.

Alternatives Considered

  • Hostname matching / a denylist of cloud providers. Rejected as the primary rule: it is trivially talked around by spelling a name differently, by a CNAME, or by a private DNS entry pointing at a public host. A denylist of known cloud hostnames is retained, but only to make the error message explain why — the address rule already refuses all of them.
  • Python's is_private / is_reserved. Rejected; see below.
  • Loopback-only inference. Rejected: trusted-LAN inference is accepted production behaviour under ADR 002, and a great many users will run Ollama on the one machine in the house that has a GPU.
  • UI-only discouragement. Rejected: a setting that is merely absent from a dropdown is still reachable by editing the database.

Reason for Explicit CIDRs

Generic address classifications do not answer this project's security question, and get it wrong in both directions for addresses this application actually sees:

  • is_private is true of the documentation ranges and of 0.0.0.0/8, neither of which is a user's LAN;
  • is_reserved is true of IPv6 loopback, so a rule written around it refuses http://[::1]:11434/v1 — an ordinary same-host Ollama.

Naming the networks keeps the policy readable, makes it auditable against this document, and makes anything unnamed refused by default:

127.0.0.0/8     this machine          ::1/128     this machine, v6
10.0.0.0/8      RFC1918               fc00::/7    unique-local, v6
172.16.0.0/12   RFC1918               fe80::/10   link-local, v6
192.168.0.0/16  RFC1918
169.254.0.0/16  link-local
100.64.0.0/10   carrier-grade NAT, which mesh VPNs such as Tailscale use

Carrier-grade NAT is included deliberately: it is what a mesh VPN such as Tailscale hands out, and such a network is as user-controlled as a LAN.

Consequences

  • Validation runs in two places — routers/settings.py on save and on the connection test, and providers/openai_compatible.py before the generate, chat and embedding requests. Configuration validation alone is not sufficient, and a database edited behind the settings API must not become a way out.
  • Because a hostname is resolved at check time, an endpoint given as a literal address behaves most predictably.
  • No cloud provider can be configured, so no API-key storage is needed. M2 removed both.
  • The error strings are user-facing and say what to do, not what failed internally.

Accepted Residual Risks

Documented rather than mitigated, and not M3 work:

  1. A hostile host on a trusted LAN is inside the permitted boundary. The policy authorizes an address range, not a machine. If an attacker already controls a device on the user's network and the user points the storyteller at it, the story goes there. The defence is the inference host's own firewall and network policy.
  2. A DNS-rebinding interval exists between the policy's resolution of a hostname and the HTTP client's own connection. The two resolutions are separate, so a name that answers with a LAN address for the check and a public address for the connection is theoretically possible. Using a literal address closes it entirely.

Scope

This ADR covers which inference endpoints the product will talk to. It is not a general remote-access design: it says nothing about exposing the storyteller itself beyond loopback, which remains out of scope for v1.

References

  • SECURITY-THREAT-MODEL.md §10A, §71A item 5, §77
  • TECHNICAL-DESIGN.md §5.1 item 4, §5.2
  • ADR 002 (Ollama-only, and the TLS consequence), ADR 004 (local-only production)
  • planning/archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md §F, §K.1