Planning: record M2 closeout decisions

M2's review reported six planning recommendations rather than applying them,
three marked before M3. All six are applied here, plus three additions drawn
from the same evidence. No implementation file is touched.

The endpoint policy was the gap that mattered. It is the most consequential
setting in the application — the storyteller sends the player's prose, the
context, the memories and the embedding inputs to whatever address it names —
and it existed only as a module docstring. It is now ADR 011 and a new §10A in
the threat model, which also retires the assumption in §71A that the inherited
guard was a starting point. It was not: AI-DnD's SSRF guard blocked private
addresses to stop a hosted server reaching its own internal network, which is
the exact opposite of what a local storyteller needs. It was removed, not
adapted.

Both documents state the rule as implemented — an allowlist of explicit
local-network CIDRs, every resolved address checked, enforced on save and again
before every outbound request, TLS never traded against it — and both state the
two residual limits plainly rather than implying they are covered: a hostile
host already on the trusted LAN is inside the permitted boundary, and a
rebinding interval exists between the policy's resolution and the client's
connection. Accepted risks, not M3 work.

The CIDRs are spelled out rather than derived from is_private/is_reserved, and
the ADR records why: is_private is true of the documentation ranges and
0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it
refuses an ordinary same-host Ollama on [::1].

TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening
items. A new §5.2 records the M1/M2 architecture as fact rather than intention,
so later milestones inherit what the code does. A new §18.1 carries the lesson
of M2's two regressions: when removing a setting, test a real consumer
construction path; when adding one, prove it reaches the component that uses
it. Both defects hid behind a green suite because the tests at that boundary
were mocks.

BUILD-MILESTONES records M2 complete, with the capabilities later milestones
inherit and the debt carried forward. Two notes go to milestones that would
otherwise misread what M2 left them. M5 is told that eight rollback tests now
use the world-state engine as instrumentation and not as endorsement — the
instrumentation moves when the protocol does, and those tests are reworked
rather than deleted. M6 is told that the memory bank died silently under a
green suite, so background failure must be observable and at least one real
provider-construction path must be tested.

The security contract gains what M2 demonstrated. H10 now names the two
conditions that were defects during M2: a wildcard origin must be refused at
startup, and an unknown /api path must 404 rather than returning the SPA with
200. New H12 covers endpoint enforcement, and its fourth pass condition is the
one that matters — a public endpoint written into the database behind the
settings API must still be refused at the wire. A build passing the first three
and failing that one has configuration validation only.

SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement;
it removed capability the specification never asked for.

The two M2 reports gain appended closeout notes rather than edits. Their
original wording about an uncommitted working tree was true when written, and
the note records what happened afterwards: the six-file correction is 8652fe7,
8c65ae9 remains the implementation commit, and the two were never squashed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
This commit is contained in:
JesseMarkowitz
2026-09-03 01:52:03 -04:00
co-authored by Claude Opus 5
parent 8652fe7cd8
commit 2fdd2547f0
9 changed files with 659 additions and 26 deletions
+105 -4
View File
@@ -1,6 +1,7 @@
# Adventure Storyteller — Security Threat Model
**Status:** v1.0 — local-only hardening requirements informed by Phase 0B
**Status:** v1.1 — local-only hardening requirements informed by Phase 0B, with the
inference endpoint policy recorded as implemented in M2 (§10A, §71A item 5, §77)
**Purpose:** Define the security and privacy boundaries for a local-only interactive storytelling application.
## 1. Security Objective
@@ -231,6 +232,104 @@ http://inferencebox.local:11434
The LAN hostname/address must be an intentional user configuration. Do not infer that every non-loopback endpoint is trusted merely because it resolves.
## 10A. Inference Endpoint Policy As Implemented (M2)
**Status:** implemented in M2, `backend/app/endpoints.py`. §10 above states the
requirement; this section records the rule that now enforces it, and what it
does not cover. ADR 011 records the decision.
This section supersedes the assumption in §71A item 5 that the inherited
network guard was the starting point. It was not reusable: AI-DnD's guard was
an SSRF guard for a *hosted* deployment, and its rule is the **opposite** of
this product's. A hosted server blocks private addresses to stop a user
reaching its internal network; a local storyteller must permit exactly those
addresses and refuse the public Internet. The inherited guard was removed
rather than adapted.
### The path
```text
Browser -> storyteller on loopback
Storyteller -> SQLite / local files
Storyteller -> one approved Ollama endpoint
```
The Ollama endpoint is either same-host loopback (the default) or an explicitly
configured trusted-LAN/local-network endpoint. Configuring a LAN inference host
does not change where the storyteller itself listens: the UI/API remains
loopback-bound by default, and the endpoint setting has no influence on the
bind address.
### The rule
Endpoints are validated **by address against an explicit allowlist of
networks**, not by hostname matching:
```text
127.0.0.0/8 this machine ::1/128 this machine, v6
10.0.0.0/8 RFC1918 fc00::/7 unique-local, v6
172.16.0.0/12 RFC1918 fe80::/10 link-local, v6
192.168.0.0/16 RFC1918
169.254.0.0/16 link-local
100.64.0.0/10 carrier-grade NAT, which mesh VPNs such as Tailscale use
```
- **Every** address the hostname resolves to must fall inside one of these
networks. One address outside is enough to refuse the endpoint, so a name
resolving to both a private and a public address does not squeak through.
- Public Internet addresses are **refused**, not merely discouraged or hidden
from a dropdown.
- The networks are spelled out rather than derived from Python's `is_private` /
`is_reserved` classifications, which do not answer this question: `is_private`
is true of the documentation ranges and of `0.0.0.0/8`, and `is_reserved` is
true of IPv6 loopback — so a rule built on it would refuse `http://[::1]:11434/v1`,
an ordinary same-host Ollama. See ADR 011.
- A short list of known cloud inference hostnames is checked first. The address
rule already refuses all of them; the list exists only so the error explains
*why* rather than leaving the user to suspect broken DNS.
### Where it is enforced
Twice, deliberately — configuration validation alone is not sufficient:
1. **when settings are saved** (`routers/settings.py`), so the user gets an
immediate, specific error, and on the connection-test path;
2. **before every outbound request** (`providers/openai_compatible.py`, on the
generate, chat and embedding paths), because a name that resolved to a LAN
address this morning can resolve elsewhere this afternoon — and because a
database edited behind the settings API must not become a way out.
M2 demonstrated the second at runtime: a cloud endpoint written straight into
SQLite with `sqlite3`, bypassing the API entirely, was still refused at the wire.
### TLS
HTTPS to a trusted-LAN Ollama with a privately issued certificate is supported.
TLS verification is **never traded against** the address policy:
- certificate and hostname verification remain fully enabled,
- trust is the machine's own CA store unioned with certifi (`tlstrust.py`, ADR 002),
- there is **no `verify=False`, no bypass flag, and no "insecure" option** —
however private the address.
### Residual limits
Stated plainly, because the policy does not cover them:
1. **A hostile host on a network the user treats as trusted is inside the
permitted boundary.** The policy authorizes an address range, not a machine.
If an attacker already controls a device on the user's LAN and the user
points the storyteller at it, the story goes there. Defending that is the
inference host's own firewall and network policy (§7), not this rule.
2. **A DNS-rebinding interval exists** between the policy resolving a hostname
and the HTTP client making its own connection. The two resolutions are
separate, so a name that answers with a LAN address for the check and a
public one for the connection is theoretically possible. Using a literal
address rather than a hostname closes it entirely.
Both are **accepted residual risks for v1**, documented rather than mitigated.
Neither is M3 work.
## 11. Imported Files Must Be Data Only
Imported files must never be treated as executable application extensions.
@@ -1108,9 +1207,10 @@ Phase 0B runtime validation found specific inherited behaviors that production m
- executable campaign scripting is outside the v1 trust boundary.
- Remove/disable the engine and replace any test-only instrumentation that depended on it.
5. **Endpoint policy mismatch**
- inherited network guarding is aimed at hosted deployment behavior, not at preventing accidental story-data exfiltration.
5. **Endpoint policy mismatch** — **resolved in M2.**
- inherited network guarding is aimed at hosted deployment behavior, not at preventing accidental story-data exfiltration. Its rule was in fact the *opposite* of this product's, so it was removed rather than adapted.
- Production should default to loopback Ollama, explicitly support a configured trusted-LAN Ollama host, and reject/avoid arbitrary public Internet inference endpoints.
- **Done.** Implemented as an address-based allowlist enforced on save and again before every request; see §10A and ADR 011.
6. **Postgres is removable**
- Phase 0B found no architectural blocker to dropping Postgres support; SQLite remains the v1 store.
@@ -1222,7 +1322,8 @@ local lexical/semantic retrieval
optional explicitly configured local media services in the future
```
Required production defaults:
Required production defaults. **As of M2 every item below is implemented**;
the endpoint rule that enforces the third and fourth is recorded in §10A:
- storyteller binds loopback by default,
- Ollama endpoint is same-host loopback by default,