Planning: record M2 closeout decisions

M2's review reported six planning recommendations rather than applying them,
three marked before M3. All six are applied here, plus three additions drawn
from the same evidence. No implementation file is touched.

The endpoint policy was the gap that mattered. It is the most consequential
setting in the application — the storyteller sends the player's prose, the
context, the memories and the embedding inputs to whatever address it names —
and it existed only as a module docstring. It is now ADR 011 and a new §10A in
the threat model, which also retires the assumption in §71A that the inherited
guard was a starting point. It was not: AI-DnD's SSRF guard blocked private
addresses to stop a hosted server reaching its own internal network, which is
the exact opposite of what a local storyteller needs. It was removed, not
adapted.

Both documents state the rule as implemented — an allowlist of explicit
local-network CIDRs, every resolved address checked, enforced on save and again
before every outbound request, TLS never traded against it — and both state the
two residual limits plainly rather than implying they are covered: a hostile
host already on the trusted LAN is inside the permitted boundary, and a
rebinding interval exists between the policy's resolution and the client's
connection. Accepted risks, not M3 work.

The CIDRs are spelled out rather than derived from is_private/is_reserved, and
the ADR records why: is_private is true of the documentation ranges and
0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it
refuses an ordinary same-host Ollama on [::1].

TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening
items. A new §5.2 records the M1/M2 architecture as fact rather than intention,
so later milestones inherit what the code does. A new §18.1 carries the lesson
of M2's two regressions: when removing a setting, test a real consumer
construction path; when adding one, prove it reaches the component that uses
it. Both defects hid behind a green suite because the tests at that boundary
were mocks.

BUILD-MILESTONES records M2 complete, with the capabilities later milestones
inherit and the debt carried forward. Two notes go to milestones that would
otherwise misread what M2 left them. M5 is told that eight rollback tests now
use the world-state engine as instrumentation and not as endorsement — the
instrumentation moves when the protocol does, and those tests are reworked
rather than deleted. M6 is told that the memory bank died silently under a
green suite, so background failure must be observable and at least one real
provider-construction path must be tested.

The security contract gains what M2 demonstrated. H10 now names the two
conditions that were defects during M2: a wildcard origin must be refused at
startup, and an unknown /api path must 404 rather than returning the SPA with
200. New H12 covers endpoint enforcement, and its fourth pass condition is the
one that matters — a public endpoint written into the database behind the
settings API must still be refused at the wire. A build passing the first three
and failing that one has configuration validation only.

SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement;
it removed capability the specification never asked for.

The two M2 reports gain appended closeout notes rather than edits. Their
original wording about an uncommitted working tree was true when written, and
the note records what happened afterwards: the six-file correction is 8652fe7,
8c65ae9 remains the implementation commit, and the two were never squashed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
This commit is contained in:
JesseMarkowitz
2026-09-03 01:52:03 -04:00
co-authored by Claude Opus 5
parent 8652fe7cd8
commit 2fdd2547f0
9 changed files with 659 additions and 26 deletions
+76 -3
View File
@@ -1,6 +1,7 @@
# Adventure Storyteller — V1 Acceptance Tests
**Status:** v1.0 planning/release contract — updated after Phase 0B
**Status:** v1.1 planning/release contract — updated after Phase 0B, and after M2 for
the security contract (H10 strengthened, H12 added)
**Purpose:** Define black-box acceptance tests for finalist evaluation during Phase 0B and for the eventual v1 release.
## 1. Test Philosophy
@@ -1174,12 +1175,32 @@ Archive extraction cannot write outside target root.
---
## H10 — Restrictive CORS
## H10 — Restrictive CORS and Local API Behavior
**Priority:** REQUIRED FOR V1
### Steps
1. Start the application with its default origin configuration and confirm the
SPA works.
2. Attempt to start the application with a wildcard origin configured
(`AIDND_CORS_ORIGINS="*"`).
3. Request an `/api/...` path that no router claims — a typo, or an endpoint
this build removed.
### Pass
Privileged local APIs do not allow arbitrary wildcard cross-origin writes.
All three conditions, each independently:
1. Privileged local APIs do not allow arbitrary wildcard cross-origin writes.
2. An unsafe wildcard production configuration is **rejected**: the application
refuses to start rather than honouring `*`. The storyteller API is
unauthenticated and loopback-bound, so a wildcard origin would let any web
page the user visits read and rewrite every campaign.
3. An unknown `/api/...` request returns an actual API **404**, rather than
falling through to the SPA mount and returning the page with HTTP 200.
Conditions 2 and 3 were defects found and fixed during M2. Without naming them
here they can regress unnoticed, because both fail in a direction that still
looks like a working application.
---
@@ -1204,6 +1225,58 @@ The application does not attempt to fetch tokenizer encodings, fonts, scripts, s
---
## H12 — Inference Endpoint Enforcement
**Priority:** REQUIRED FOR V1
Defence in depth for the setting that decides where the story goes.
Configuration validation alone is **not** sufficient, so this test deliberately
checks the request-time rule as well. See ADR 011 and
`SECURITY-THREAT-MODEL.md` §10A.
### Preconditions
- application installed and running,
- an Ollama instance reachable on this machine,
- an Ollama instance reachable on the user's own network (for step 2),
- the ability to edit the application database directly (for step 4).
### Steps
1. Configure a **loopback** Ollama endpoint (`http://127.0.0.1:11434/v1`) and
generate a story turn. Repeat with the IPv6 form `http://[::1]:11434/v1`.
2. Configure an **approved trusted-LAN** Ollama endpoint by address and by
hostname, over HTTP and over HTTPS with a privately issued certificate, and
generate a story turn.
3. Attempt to configure a **public Internet** inference endpoint through the
normal settings API — both a known cloud provider hostname and an arbitrary
public address.
4. With the application configured legitimately, write a **public** endpoint
directly into the settings row in the database, bypassing the settings API
entirely, then attempt to generate a turn.
### Pass
1. Loopback endpoints are **accepted**, in both IPv4 and IPv6 form.
2. Approved trusted-LAN/local-network endpoints are **accepted**, and the HTTPS
case succeeds with certificate and hostname verification fully enabled and no
bypass available.
3. Public Internet endpoints are **rejected** through normal configuration, with
an error that says why and what to use instead.
4. Request-time enforcement **still rejects** the public endpoint written behind
the settings API: no story text, context, memory or embedding input leaves
the machine for that address. The turn fails with the endpoint's rejection
reason rather than succeeding.
A build that passes 1-3 but fails 4 has configuration validation only, and does
not pass this test.
### Notes
Every address a hostname resolves to must be inside the allowed local networks;
one address outside is enough to refuse the endpoint. The two known residual
limits — a hostile host already on the trusted LAN, and the DNS-rebinding
interval between the policy's resolution and the client's connection — are
accepted residual risks and are **not** failures of this test.
---
# I. Export, Backup, and Restore
## I01 — Export Campaign