Planning: record M2 closeout decisions

M2's review reported six planning recommendations rather than applying them,
three marked before M3. All six are applied here, plus three additions drawn
from the same evidence. No implementation file is touched.

The endpoint policy was the gap that mattered. It is the most consequential
setting in the application — the storyteller sends the player's prose, the
context, the memories and the embedding inputs to whatever address it names —
and it existed only as a module docstring. It is now ADR 011 and a new §10A in
the threat model, which also retires the assumption in §71A that the inherited
guard was a starting point. It was not: AI-DnD's SSRF guard blocked private
addresses to stop a hosted server reaching its own internal network, which is
the exact opposite of what a local storyteller needs. It was removed, not
adapted.

Both documents state the rule as implemented — an allowlist of explicit
local-network CIDRs, every resolved address checked, enforced on save and again
before every outbound request, TLS never traded against it — and both state the
two residual limits plainly rather than implying they are covered: a hostile
host already on the trusted LAN is inside the permitted boundary, and a
rebinding interval exists between the policy's resolution and the client's
connection. Accepted risks, not M3 work.

The CIDRs are spelled out rather than derived from is_private/is_reserved, and
the ADR records why: is_private is true of the documentation ranges and
0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it
refuses an ordinary same-host Ollama on [::1].

TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening
items. A new §5.2 records the M1/M2 architecture as fact rather than intention,
so later milestones inherit what the code does. A new §18.1 carries the lesson
of M2's two regressions: when removing a setting, test a real consumer
construction path; when adding one, prove it reaches the component that uses
it. Both defects hid behind a green suite because the tests at that boundary
were mocks.

BUILD-MILESTONES records M2 complete, with the capabilities later milestones
inherit and the debt carried forward. Two notes go to milestones that would
otherwise misread what M2 left them. M5 is told that eight rollback tests now
use the world-state engine as instrumentation and not as endorsement — the
instrumentation moves when the protocol does, and those tests are reworked
rather than deleted. M6 is told that the memory bank died silently under a
green suite, so background failure must be observable and at least one real
provider-construction path must be tested.

The security contract gains what M2 demonstrated. H10 now names the two
conditions that were defects during M2: a wildcard origin must be refused at
startup, and an unknown /api path must 404 rather than returning the SPA with
200. New H12 covers endpoint enforcement, and its fourth pass condition is the
one that matters — a public endpoint written into the database behind the
settings API must still be refused at the wire. A build passing the first three
and failing that one has configuration validation only.

SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement;
it removed capability the specification never asked for.

The two M2 reports gain appended closeout notes rather than edits. Their
original wording about an uncommitted working tree was true when written, and
the note records what happened afterwards: the six-file correction is 8652fe7,
8c65ae9 remains the implementation commit, and the two were never squashed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
This commit is contained in:
JesseMarkowitz
2026-09-03 01:52:03 -04:00
co-authored by Claude Opus 5
parent 8652fe7cd8
commit 2fdd2547f0
9 changed files with 659 additions and 26 deletions
@@ -0,0 +1,118 @@
# ADR 011 — Local Inference Endpoint Policy
**Status:** Accepted
**Date:** 2026-09-02
**Implemented in:** M2, `backend/app/endpoints.py`
## Decision
v1 supports Ollama only, and will send a story to an inference endpoint **only**
when every address that endpoint resolves to lies inside an explicit allowlist
of local networks.
- The storyteller UI/API remains **loopback-bound by default**. Configuring a
LAN inference endpoint does not change where the storyteller listens.
- Inference may use **same-host loopback** (the default) or an **explicitly
configured trusted-LAN/local-network address**.
- **Public Internet addresses are denied.**
- The policy is **address-based**, using explicit allowed CIDRs rather than
Python's generic `is_private` / `is_reserved` classifications.
- **Every** resolved address must be allowed; one address outside the allowlist
refuses the endpoint.
- The policy is checked **on configuration and again before every request**.
- **TLS verification is mandatory** for HTTPS and is never traded against this
policy.
- Arbitrary cloud / OpenAI-compatible endpoints are **intentionally outside v1**.
## Context
The endpoint setting is the most consequential in the application. The
storyteller sends the player's prose, the assembled context, the retrieved
memories and the embedding inputs to whatever address it names: point it
somewhere else and the whole campaign goes there.
The inherited AI-DnD guard could not be reused, because its rule is the
**opposite** of this product's. AI-DnD was a hosted service, so its SSRF guard
blocked *private* addresses to stop a user reaching the server's internal
network. A local storyteller must do exactly the reverse — permit the private
ranges and refuse the public Internet. The guard was removed rather than adapted.
## Alternatives Considered
- **Hostname matching / a denylist of cloud providers.** Rejected as the primary
rule: it is trivially talked around by spelling a name differently, by a
`CNAME`, or by a private DNS entry pointing at a public host. A denylist of
known cloud hostnames is retained, but only to make the *error message*
explain why — the address rule already refuses all of them.
- **Python's `is_private` / `is_reserved`.** Rejected; see below.
- **Loopback-only inference.** Rejected: trusted-LAN inference is accepted
production behaviour under ADR 002, and a great many users will run Ollama on
the one machine in the house that has a GPU.
- **UI-only discouragement.** Rejected: a setting that is merely absent from a
dropdown is still reachable by editing the database.
## Reason for Explicit CIDRs
Generic address classifications do not answer this project's security question,
and get it wrong in both directions for addresses this application actually sees:
- `is_private` is **true** of the documentation ranges and of `0.0.0.0/8`,
neither of which is a user's LAN;
- `is_reserved` is **true** of IPv6 loopback, so a rule written around it
refuses `http://[::1]:11434/v1` — an ordinary same-host Ollama.
Naming the networks keeps the policy readable, makes it auditable against this
document, and makes anything unnamed refused by default:
```text
127.0.0.0/8 this machine ::1/128 this machine, v6
10.0.0.0/8 RFC1918 fc00::/7 unique-local, v6
172.16.0.0/12 RFC1918 fe80::/10 link-local, v6
192.168.0.0/16 RFC1918
169.254.0.0/16 link-local
100.64.0.0/10 carrier-grade NAT, which mesh VPNs such as Tailscale use
```
Carrier-grade NAT is included deliberately: it is what a mesh VPN such as
Tailscale hands out, and such a network is as user-controlled as a LAN.
## Consequences
- Validation runs in two places — `routers/settings.py` on save and on the
connection test, and `providers/openai_compatible.py` before the generate,
chat and embedding requests. Configuration validation alone is not sufficient,
and a database edited behind the settings API must not become a way out.
- Because a hostname is resolved at check time, an endpoint given as a literal
address behaves most predictably.
- No cloud provider can be configured, so no API-key storage is needed. M2
removed both.
- The error strings are user-facing and say what to do, not what failed
internally.
## Accepted Residual Risks
Documented rather than mitigated, and **not** M3 work:
1. **A hostile host on a trusted LAN is inside the permitted boundary.** The
policy authorizes an address range, not a machine. If an attacker already
controls a device on the user's network and the user points the storyteller
at it, the story goes there. The defence is the inference host's own firewall
and network policy.
2. **A DNS-rebinding interval** exists between the policy's resolution of a
hostname and the HTTP client's own connection. The two resolutions are
separate, so a name that answers with a LAN address for the check and a public
address for the connection is theoretically possible. Using a literal address
closes it entirely.
## Scope
This ADR covers which inference endpoints the product will talk to. It is **not**
a general remote-access design: it says nothing about exposing the storyteller
itself beyond loopback, which remains out of scope for v1.
## References
- `SECURITY-THREAT-MODEL.md` §10A, §71A item 5, §77
- `TECHNICAL-DESIGN.md` §5.1 item 4, §5.2
- ADR 002 (Ollama-only, and the TLS consequence), ADR 004 (local-only production)
- `planning/reports/M2-IMPLEMENTATION-REPORT.md` §F, §K.1