Apply post-M1 corrections to the planning package

This commit is contained in:
JesseMarkowitz
2026-09-02 06:03:10 -04:00
parent 645f07f06d
commit 1a28a9a708
8 changed files with 304 additions and 35 deletions
+65 -12
View File
@@ -132,19 +132,59 @@ Accepted Story / Scene State
`Local-only` means user-controlled local infrastructure with no required Internet/cloud dependency; it does not require every process to share one host.
Default same-host runtime path:
The architecture has **two boundaries, and they are not the same boundary**. The
storyteller is loopback-only, always. Inference may be same-host or on a
specifically configured trusted-LAN machine. Wording that describes Ollama as
simply "loopback/local" collapses the two and understates the intended
deployment:
```text
Browser -> local FastAPI application -> 127.0.0.1 Ollama
Browser ──loopback──> storyteller (FastAPI + SPA), bound to 127.0.0.1
│
├── same-host Ollama on 127.0.0.1:11434 (default)
│
└── OR an explicitly configured trusted-LAN Ollama
on another user-controlled machine
http://<host>:11434/v1 or https://<host>/v1
```
Supported v1 trusted-LAN inference path:
Read that as three separate rules:
```text
Browser -> local FastAPI application -> explicitly approved LAN Ollama host
```
1. **The storyteller's own listener is loopback, in every run path.** Dev
server, production server, and container alike. Nothing about the inference
choice changes it. Where a container must listen on `0.0.0.0` because a
published port cannot reach anything else, the port is published to the
host's loopback only.
2. **The inference endpoint is an outbound connection, chosen by the user.**
Same-host loopback is the default. A trusted-LAN host is a first-class,
supported v1 configuration — not a workaround and not a development-only
convenience.
3. **The two are independent.** Reaching a LAN Ollama never requires, and must
never cause, LAN exposure of the storyteller UI/API. There is no supported
v1 configuration in which the storyteller itself is reachable from the LAN.
The browser and storyteller API still remain loopback-bound by default. Supporting LAN inference does **not** expose the storyteller web application to the LAN.
### Transport to a trusted-LAN endpoint
A LAN inference host is often reached over **HTTPS with a certificate issued by
a private or local CA**, and may offer no cleartext port at all. This is
ordinary for a self-hosted server, so v1 must handle it rather than assume the
same plain HTTP that same-host loopback uses:
- outbound HTTPS verifies against the **operating system's trusted CA store** in
addition to any bundled certificate list, so a CA the user installed on their
own machine is honoured here as it is by `curl` and their browser;
- certificate **and hostname** verification stay fully enabled;
- there is **no "ignore TLS errors" option** anywhere — not in the UI, not in
configuration, not as an environment variable;
- the endpoint field therefore accepts `https://` on any port.
Prompts, story text, retrieved knowledge and embedding inputs all travel to
whichever inference host is configured, which is why it must be one the user
controls on a network they trust — and why the endpoint is always explicitly
configured, never discovered.
See ADR 002 (*Transport for a Trusted-LAN Endpoint*) and, for the demonstrated
deployment, `planning/reports/M1-IMPLEMENTATION-REPORT.md` §F.
Allowed future local paths may also include explicitly configured local media services.
@@ -161,17 +201,30 @@ Production defaults must not require:
### 5.1 Known AI-DnD hardening work
Phase 0B identified concrete inherited violations:
Phase 0B identified concrete inherited violations, and M1 added a fifth. Items
1, 2 and 5 are **resolved**; items 3 and 4 remain open and belong to M2.
1. `tiktoken` attempts to download the `cl100k_base` encoding on first use.
- production packaging must include/cache the required encoding or replace the dependency path so first story use is offline.
2. the SPA requests Google Fonts at runtime.
- fonts must be self-hosted or replaced with local/system fonts and CSP tightened.
1. ~~`tiktoken` attempts to download the `cl100k_base` encoding on first use.~~
**Done in M1.** The encoding table is vendored in the tree and loaded
directly, with its SHA-256 verified against the digest `tiktoken` pins, so no
code path in the tokenizer can reach the network.
2. ~~the SPA requests Google Fonts at runtime.~~
**Done in M1.** All three families are self-hosted, and the CSP names no
remote origin at all.
3. hosted/multi-user/auth/demo/analytics/Postgres/cloud-provider/QuickJS paths are unnecessary.
- remove them rather than merely hide them where practical.
- **Open — M2.** M1 removed nothing, so this surface is unchanged from the
fork point.
4. endpoint validation must reflect this product's threat model.
- same-host loopback Ollama is the default; an explicitly configured trusted-LAN Ollama endpoint is supported; arbitrary public/Internet model endpoints must be rejected or kept outside normal v1 configuration.
- inference endpoint configuration must not change the storyteller's own loopback bind behavior.
- **Open — M2.** The trusted-LAN path itself works as of M1; what remains is
deciding and enforcing which endpoints normal v1 configuration may name.
5. ~~outbound TLS verified only against a bundled public-CA list, so a LAN host
with a privately issued certificate was refused.~~
**Found and fixed in M1.** Not visible to Phase 0B: every run up to that
point used plain HTTP over loopback, where certificate verification never
happens. See *Transport to a trusted-LAN endpoint* above.
## 6. Browser UI Boundary