M2's review reported six planning recommendations rather than applying them, three marked before M3. All six are applied here, plus three additions drawn from the same evidence. No implementation file is touched. The endpoint policy was the gap that mattered. It is the most consequential setting in the application — the storyteller sends the player's prose, the context, the memories and the embedding inputs to whatever address it names — and it existed only as a module docstring. It is now ADR 011 and a new §10A in the threat model, which also retires the assumption in §71A that the inherited guard was a starting point. It was not: AI-DnD's SSRF guard blocked private addresses to stop a hosted server reaching its own internal network, which is the exact opposite of what a local storyteller needs. It was removed, not adapted. Both documents state the rule as implemented — an allowlist of explicit local-network CIDRs, every resolved address checked, enforced on save and again before every outbound request, TLS never traded against it — and both state the two residual limits plainly rather than implying they are covered: a hostile host already on the trusted LAN is inside the permitted boundary, and a rebinding interval exists between the policy's resolution and the client's connection. Accepted risks, not M3 work. The CIDRs are spelled out rather than derived from is_private/is_reserved, and the ADR records why: is_private is true of the documentation ranges and 0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it refuses an ordinary same-host Ollama on [::1]. TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening items. A new §5.2 records the M1/M2 architecture as fact rather than intention, so later milestones inherit what the code does. A new §18.1 carries the lesson of M2's two regressions: when removing a setting, test a real consumer construction path; when adding one, prove it reaches the component that uses it. Both defects hid behind a green suite because the tests at that boundary were mocks. BUILD-MILESTONES records M2 complete, with the capabilities later milestones inherit and the debt carried forward. Two notes go to milestones that would otherwise misread what M2 left them. M5 is told that eight rollback tests now use the world-state engine as instrumentation and not as endorsement — the instrumentation moves when the protocol does, and those tests are reworked rather than deleted. M6 is told that the memory bank died silently under a green suite, so background failure must be observable and at least one real provider-construction path must be tested. The security contract gains what M2 demonstrated. H10 now names the two conditions that were defects during M2: a wildcard origin must be refused at startup, and an unknown /api path must 404 rather than returning the SPA with 200. New H12 covers endpoint enforcement, and its fourth pass condition is the one that matters — a public endpoint written into the database behind the settings API must still be refused at the wire. A build passing the first three and failing that one has configuration validation only. SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement; it removed capability the specification never asked for. The two M2 reports gain appended closeout notes rather than edits. Their original wording about an uncommitted working tree was true when written, and the note records what happened afterwards: the six-file correction is8652fe7,8c65ae9remains the implementation commit, and the two were never squashed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
929 lines
34 KiB
Markdown
929 lines
34 KiB
Markdown
# M2 — Baseline Report (evidence record)
|
||
|
||
**Date:** 2026-09-02
|
||
**Milestone:** M2, *Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface*
|
||
**Companion:** `planning/reports/M2-IMPLEMENTATION-REPORT.md` interprets this file.
|
||
Where the two disagree on a runtime or test fact, **this file is the record.**
|
||
|
||
This is measurement, not commentary. Command output is quoted verbatim.
|
||
Anything inferred rather than observed is labelled **[inferred]**.
|
||
|
||
Hostnames and LAN addresses are **placeholders** — `inference.lan`,
|
||
`192.168.0.50`, port `8443`. The real ones are in the workspace's untracked
|
||
notes, never in this repository. Packet counts, digests, timings and status
|
||
codes are exact.
|
||
|
||
---
|
||
|
||
## 0. Test environment
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Host | Ubuntu 24.04.4 LTS, x86-64, 4 cores, 15 GB RAM, **no GPU** |
|
||
| Python | 3.12.3 |
|
||
| Node / npm | 22.23.1 / 10.9.8 |
|
||
| Docker | 29.7.2 |
|
||
| Ollama (same-host) | `ollama/ollama:latest` in a container, models from a shared volume |
|
||
| Ollama (LAN) | 0.33.0 on a **second physical machine**, HTTPS, certificate from a private CA |
|
||
| Models | `qwen2.5:0.5b`, `qwen2.5:3b-instruct`, `nomic-embed-text` |
|
||
| Image under test | `m2-final`, built from the working tree by `docker build .` |
|
||
|
||
> **Timings in this file are not benchmarks.** This host was shared with
|
||
> unrelated work during the review; `uptime` reported a 1-minute load average
|
||
> ranging from **1.79 to 74.08** across the session on four cores. Where a
|
||
> measurement was taken under load, it says so. Latency figures are recorded
|
||
> because they mattered to a *timeout* result, not as performance data.
|
||
|
||
---
|
||
|
||
## 1. Repository and provenance
|
||
|
||
```console
|
||
$ git rev-parse --abbrev-ref HEAD
|
||
m2-local-only-surface
|
||
|
||
$ git log --format='%H %G? %s' -3
|
||
8c65ae99deda49b22f415d5868987e00bbcb173c G M2: cut the hosted product away from the local one
|
||
1a28a9a708985e7c98dcbcb87c189e7582ec288d G Apply post-M1 corrections to the planning package
|
||
645f07f06d226e274c33d96e71d9c374413ef771 G Add the M1 implementation review report
|
||
```
|
||
|
||
All three commits verify (`%G? = G`). The repository signs every commit; nothing
|
||
in this milestone is unsigned.
|
||
|
||
### Commit chain
|
||
|
||
| Commit | Sig | Parent | Meaning |
|
||
| --- | --- | --- | --- |
|
||
| `8c65ae9` | G | `1a28a9a` | **The M2 commit.** One commit, no merges. |
|
||
| `1a28a9a` | G | `645f07f` | Post-M1 planning corrections — the M1 baseline this milestone started from |
|
||
| `7f182a8` | G | `717670a`, `d72f7c1` | the fork import, two parents |
|
||
|
||
### Upstream ancestry
|
||
|
||
```console
|
||
$ git cat-file -t d72f7c1bda0f34fccd84afb7a25c34eb01c901de
|
||
commit
|
||
$ git merge-base --is-ancestor d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD && echo yes
|
||
yes
|
||
$ git rev-list --count d72f7c1bda0f34fccd84afb7a25c34eb01c901de
|
||
172
|
||
$ git diff --stat d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD -- LICENSE
|
||
(no output)
|
||
```
|
||
|
||
All 172 upstream commits reachable; `LICENSE` byte-identical to upstream; no
|
||
re-import or upstream substitution occurred. `PROVENANCE.md` and
|
||
`DEVELOPMENT.md` are present and were updated for M2.
|
||
|
||
### Private data
|
||
|
||
```console
|
||
$ grep -rn "<the real LAN hostname>\|<the real LAN subnet>\|<the real port>" \
|
||
--include=*.py --include=*.md --include=*.jsx --include=*.js \
|
||
--include=*.yml --include=*.txt backend frontend *.md *.yml Dockerfile
|
||
(no output)
|
||
```
|
||
|
||
No private hostname, LAN address, certificate or credential is committed.
|
||
|
||
### Working tree
|
||
|
||
At the M2 commit the tree was clean. **This review then modified six files**
|
||
(§9), so `git status` is *not* clean as this report is written:
|
||
|
||
```text
|
||
M backend/app/memorybank.py
|
||
M backend/app/models.py
|
||
M backend/app/routers/adventures/turns.py
|
||
M backend/app/routers/chat.py
|
||
M backend/requirements.lock
|
||
M backend/tests/test_local_only_surface.py
|
||
```
|
||
|
||
Every runtime result below was produced by an image built from that modified
|
||
tree, except where explicitly marked as pre-fix.
|
||
|
||
---
|
||
|
||
## 2. Change inventory
|
||
|
||
```console
|
||
$ git diff --shortstat 1a28a9a 8c65ae9
|
||
94 files changed, 1395 insertions(+), 6578 deletions(-)
|
||
```
|
||
|
||
**Added (3):**
|
||
|
||
```text
|
||
backend/app/endpoints.py
|
||
backend/tests/test_endpoint_policy.py
|
||
backend/tests/test_local_only_surface.py
|
||
```
|
||
|
||
**Deleted (24):**
|
||
|
||
```text
|
||
backend/app/accesslog.py backend/tests/test_accesslog.py
|
||
backend/app/analytics.py backend/tests/test_analytics.py
|
||
backend/app/cleanup.py backend/tests/test_guest_cleanup.py
|
||
backend/app/netguard.py backend/tests/test_netguard.py
|
||
backend/app/security.py backend/tests/test_ratelimit_hardening.py
|
||
backend/app/routers/analytics.py backend/tests/test_reasoning_param.py
|
||
backend/app/routers/auth.py
|
||
backend/app/routers/scripts.py frontend/src/pages/Analytics.jsx
|
||
backend/app/routers/adventures/scripts.py frontend/src/pages/Scripts.jsx
|
||
backend/app/scripting/__init__.py frontend/src/pages/ScriptEditor.jsx
|
||
backend/app/scripting/engine.py frontend/src/pages/Play/panels/ScriptsPanel.jsx
|
||
backend/app/scripting/pipeline.py frontend/src/pages/Play/drawers/StatusDrawer.jsx
|
||
render.yaml
|
||
```
|
||
|
||
**Largest modifications:**
|
||
|
||
```text
|
||
+93 -37 backend/app/routers/settings.py
|
||
+54 -82 backend/app/providers/openai_compatible.py
|
||
+44 -166 backend/tests/test_chat.py
|
||
+38 -81 backend/app/routers/chat.py
|
||
+37 -222 backend/app/auth.py
|
||
+32 -69 backend/app/main.py
|
||
+31 -195 backend/app/limits.py
|
||
+28 -178 backend/app/models.py
|
||
+22 -43 backend/app/database.py
|
||
+19 -58 frontend/src/pages/Settings.jsx
|
||
+18 -97 backend/app/routers/adventures/turns.py
|
||
+15 -85 frontend/src/App.jsx
|
||
```
|
||
|
||
---
|
||
|
||
## 3. Surface reduction, measured
|
||
|
||
### API routes (from the OpenAPI schema, not by grep)
|
||
|
||
| Prefix | M1 | M2 |
|
||
| --- | ---: | ---: |
|
||
| `/api/adventures` | 25 | 21 |
|
||
| `/api/analytics` | 3 | **0** |
|
||
| `/api/auth` | 4 | **0** |
|
||
| `/api/scripts` | 5 | **0** |
|
||
| `/api/chat` | 2 | 2 |
|
||
| `/api/scenarios` | 5 | 5 |
|
||
| `/api/settings` | 2 | 2 |
|
||
| `/api/story-cards` | 4 | 4 |
|
||
| `/api/debug`, `/api/health` | 2 | 2 |
|
||
| **Total** | **52** | **36** |
|
||
|
||
### Dependencies
|
||
|
||
```console
|
||
$ diff <(M1 requirements.txt) <(M2 requirements.txt) | grep '^<'
|
||
< quickjs>=1.19
|
||
< cryptography>=42
|
||
< psycopg[binary]>=3.2
|
||
```
|
||
|
||
Installed closure, measured by building a clean venv from `requirements.txt` +
|
||
`requirements-dev.txt`:
|
||
|
||
| | M1 | M2 |
|
||
| --- | ---: | ---: |
|
||
| Python packages installed | 40 | **34** |
|
||
| Packages gone | — | `cffi`, `cryptography`, `psycopg`, `psycopg-binary`, `pycparser`, `quickjs` |
|
||
| npm runtime dependencies | 6 | **3** |
|
||
| npm packages installed (`npm ls --all`) | 53 | **32** |
|
||
|
||
### Code size
|
||
|
||
| | M1 | M2 | Δ |
|
||
| --- | ---: | ---: | ---: |
|
||
| `backend/app` Python lines | 14 298 | 11 632 | −2 666 |
|
||
| `frontend/src` JS/JSX lines | 6 523 | 5 239 | −1 284 |
|
||
| `frontend/src/pages` files | 21 | 16 | −5 |
|
||
|
||
### Environment variables actually read (`os.environ`)
|
||
|
||
```console
|
||
$ grep -rn "os.environ" backend/app/
|
||
backend/app/database.py:18:_env_db_path = os.environ.get("AIDND_DB_PATH")
|
||
backend/app/main.py:30: for o in os.environ.get("AIDND_CORS_ORIGINS", "").split(",")
|
||
```
|
||
|
||
Two, down from ten. `AIDND_MULTI_USER`, `AIDND_SECRET_KEY`, `AIDND_COOKIE_SECURE`,
|
||
`AIDND_DEMO_API_KEY`, `AIDND_DEMO_ENDPOINT_URL`, `AIDND_DEMO_MODELS`,
|
||
`AIDND_DEMO_TURNS_PER_DAY`, `AIDND_POWER_USERS`, `AIDND_ANALYTICS_EMAILS`,
|
||
`AIDND_TRUSTED_PROXY_HOPS`, `AIDND_DATABASE_URL` and `DATABASE_URL` are no
|
||
longer read. Four of those names still appear in the tree **as prose in
|
||
comments** explaining what was removed; `grep` above shows no read.
|
||
|
||
---
|
||
|
||
## 4. Test suite
|
||
|
||
```console
|
||
$ cd backend && .venv/bin/python -m pytest tests/ -q
|
||
606 passed, 1 warning in 126.29s (0:02:06)
|
||
```
|
||
|
||
Zero failed, zero skipped, zero xfailed. The one warning is the inherited
|
||
Starlette/`httpx` deprecation notice, present since M1.
|
||
|
||
### Accounting for every test
|
||
|
||
Counted by diffing collected node IDs between `1a28a9a` (M1) and the working
|
||
tree, not by reading diffs:
|
||
|
||
| | Count |
|
||
| --- | ---: |
|
||
| Distinct test functions, M1 | 610 |
|
||
| Distinct test functions, M2 | 561 |
|
||
| Node IDs gone | 88 |
|
||
| Node IDs new | 39 |
|
||
| Collected tests (with parametrisation), M1 | 648 |
|
||
| Collected tests, M2 | **606** |
|
||
|
||
**Gone, by file:**
|
||
|
||
| File | Gone | Disposition |
|
||
| --- | ---: | --- |
|
||
| `test_analytics.py` | 25 | whole file — subject removed |
|
||
| `test_guest_cleanup.py` | 14 | whole file — subject removed |
|
||
| `test_accesslog.py` | 12 | whole file — subject removed |
|
||
| `test_ratelimit_hardening.py` | 8 | whole file — subject removed |
|
||
| `test_chat.py` | 8 | demo-key pinning and the power-user gate |
|
||
| `test_reasoning_param.py` | 5 | whole file — subject removed |
|
||
| `test_netguard.py` | 5 | whole file — replaced by `test_endpoint_policy.py` |
|
||
| `test_prompt_caching.py` | 4 | OpenRouter upstream routing |
|
||
| `test_memory_rewrite.py` | 4 | 2 retired (Postgres DSN masking), **2 renamed** |
|
||
| `test_delete_state.py` | 1 | **renamed** |
|
||
| `test_branch_forking.py` | 1 | **renamed** |
|
||
| `test_branch_clause.py` | 1 | scripting history API |
|
||
|
||
**Four of the 88 are renames with equivalent coverage**, so 84 tests were
|
||
genuinely retired:
|
||
|
||
```text
|
||
test_branch_forking: test_switching_restores_the_script_and_world_state
|
||
-> test_switching_restores_the_state_a_branch_left_behind
|
||
test_delete_state: test_deleting_the_ai_turn_rewinds_the_script_state
|
||
-> test_deleting_the_ai_turn_rewinds_the_counter
|
||
test_memory_rewrite: test_an_owner_with_no_api_key_is_skipped
|
||
-> test_an_adventure_with_no_model_configured_is_skipped
|
||
test_memory_rewrite: test_an_api_key_on_the_command_line_covers_that_owner
|
||
-> test_a_model_on_the_command_line_covers_that_adventure
|
||
```
|
||
|
||
**New, by file:** `test_local_only_surface.py` 18, `test_endpoint_policy.py`
|
||
14, `test_chat.py` 3, plus the 4 renames. Collected with parametrisation:
|
||
`test_endpoint_policy.py` 31, `test_local_only_surface.py` 32.
|
||
|
||
### Instrumentation conversion, not deletion
|
||
|
||
Eight files used a QuickJS `output` hook (`state.gold += 10`) as deterministic
|
||
instrumentation for the **state snapshot and rollback machinery**, which M2 does
|
||
not touch. The counter moved to the world-state engine — the model emits a
|
||
` ```state ` delta block, the referee applies it — and the assertions are
|
||
unchanged in substance. Affected: `test_take_state`, `test_attempt_siblings`,
|
||
`test_branch_forking`, `test_bundle_v2`, `test_delete_state`,
|
||
`test_retry_variants`, `test_story_tree_baseline`, `test_turn_flow_integration`,
|
||
plus `test_state_revert` converted from `script_state`/`state_after` to
|
||
`world_state`/`world_state_after`.
|
||
|
||
### Frontend and image
|
||
|
||
```console
|
||
$ npm run lint # oxlint
|
||
exit 0 # 7 warnings, all pre-existing react/only-export-components
|
||
$ npm run build
|
||
dist/assets/index-C9v6AJ3R.js 395.41 kB │ gzip: 120.01 kB ✓ built in 496ms
|
||
$ docker build .
|
||
exit 0
|
||
```
|
||
|
||
There are no automated frontend tests in the repository — none existed at M1
|
||
either.
|
||
|
||
Bundle size: **933.69 kB → 395.41 kB** (M1 → M2), from removing CodeMirror with
|
||
the script editor.
|
||
|
||
---
|
||
|
||
## 5. Offline run — same-host Ollama
|
||
|
||
**Topology.** `--internal` Docker network (no NAT, no external DNS). Ollama in
|
||
one container; the `m2-final` image in a second container sharing Ollama's
|
||
network namespace, so Ollama sits on the storyteller's own loopback. tcpdump ran
|
||
in that namespace for the whole session.
|
||
|
||
### 5.1 Isolation, verified before any test
|
||
|
||
```text
|
||
blocked 1.1.1.1:443 OSError
|
||
blocked 140.82.121.4:443 OSError
|
||
no resolution fonts.googleapis.com
|
||
no resolution fonts.gstatic.com
|
||
no resolution openaipublic.blob.core.windows.net
|
||
no resolution openrouter.ai
|
||
no resolution api.openai.com
|
||
no resolution github.com
|
||
```
|
||
|
||
### 5.2 Listeners
|
||
|
||
```text
|
||
LISTEN 127.0.0.1:8000 <- the storyteller
|
||
LISTEN 127.0.0.11:42229 <- Docker's embedded DNS
|
||
LISTEN 127.0.0.1:43093 <- Ollama's model runner
|
||
LISTEN 127.0.0.1:43313 <- Ollama's model runner
|
||
LISTEN [::]:11434 <- Ollama itself
|
||
```
|
||
|
||
Nothing the storyteller owns is bound off loopback.
|
||
|
||
### 5.3 Ollama-only settings and diagnostics
|
||
|
||
```json
|
||
endpoint: http://127.0.0.1:11434/v1 | model: qwen2.5:0.5b
|
||
embed: nomic-embed-text:latest | timeout: 600
|
||
connection test: {"ok": true, "models": ["qwen2.5:7b-instruct",
|
||
"qwen2.5:3b-instruct", "nomic-embed-text:latest", "qwen2.5:0.5b"]}
|
||
```
|
||
|
||
### 5.4 Story generation and streaming
|
||
|
||
```text
|
||
turn 1 ( 1.7s, 90 stream chunks): 'Oh, how unfortunate! A lighthouse keeper indeed! The light'
|
||
turn 2 ( 35.2s, 90 stream chunks): 'I climb the spiral stair to the lamp room. My hand tightly'
|
||
turn 3 ( 7.1s, 90 stream chunks): "I search the keeper's log for the last entry. The last ent"
|
||
turn 4: ERROR {'type': 'error', 'detail': 'The AI endpoint timed out.'}
|
||
```
|
||
|
||
Turn 4 timed out at the configured 600 s. The application logged no error. Ollama
|
||
reported the model still resident. `uptime` at the time: 1-minute load average
|
||
had risen from 1.79 to the tens on four cores from unrelated work on this host.
|
||
**[inferred]** the timeout is host contention rather than an application fault;
|
||
what is *measured* is that no application error was logged and that the same
|
||
build produced turns in 1.7–35.2 s minutes earlier.
|
||
|
||
### 5.5 Local embeddings, end to end through the app
|
||
|
||
```text
|
||
wrote memory id=1, embedded=False
|
||
turn to trigger the pass (101.6s)
|
||
memories: 1 embedded with the local model: 1
|
||
```
|
||
|
||
Direct timing of the embedding endpoint from inside the container:
|
||
|
||
```text
|
||
embedding round trip: 3.4s, dim=768
|
||
second embedding (warm): 0.1s
|
||
```
|
||
|
||
### 5.6 Branch-scoped memory isolation — positive and negative controls
|
||
|
||
```text
|
||
main branch = 1
|
||
memories on main: 0 ALPHA present: False
|
||
forking at AI action 22 (an earlier turn, so this starts a new line)
|
||
fork events: ['chunk', 'chunk', 'done']
|
||
branches now: 2, fork = 5
|
||
|
||
ON THE FORK (0 memories)
|
||
ALPHA visible: False <- ALPHA is anchored at main's head, which is BELOW the
|
||
fork point, so the fork correctly does not inherit it
|
||
BETA written here, visible: True <- positive control
|
||
|
||
BACK ON MAIN (1 memories)
|
||
ALPHA visible: True <- positive control
|
||
BETA visible: False <- NEGATIVE CONTROL, must be False
|
||
```
|
||
|
||
The decisive result is the last line: a memory written on the fork is **not**
|
||
visible on main.
|
||
|
||
### 5.7 A05 — a failed model call, and a hand-edited database
|
||
|
||
```text
|
||
baseline: (8 actions, 3 accepted AI turns, digest 07615da99b014b70)
|
||
|
||
-- invalid local model --
|
||
events: ['player', 'error']
|
||
error: Endpoint or model not found (HTTP 404). Check the endpoint URL and that model 'n…
|
||
after: (7, 3, 'e18de9dff6455e0d')
|
||
|
||
-- endpoint refused by policy at request time --
|
||
(the settings row was edited directly with sqlite3, behind the app's back,
|
||
to https://openrouter.ai/api/v1)
|
||
events: ['player', 'error']
|
||
error: This endpoint can't be used — openrouter.ai is a cloud inference service — this build…
|
||
after: (8, 3, '157c882c60579617')
|
||
```
|
||
|
||
The **accepted AI-turn count stayed at 3** through both failures. Only the
|
||
player's own typed action was added each time, which is M1's documented and
|
||
deliberate behaviour.
|
||
|
||
The second case is the strongest form of the endpoint evidence: the request-time
|
||
check refused a cloud endpoint that had been written straight into SQLite,
|
||
bypassing the API's save-time validation entirely.
|
||
|
||
### 5.8 Restart and resume
|
||
|
||
```text
|
||
before restart: 8 actions, digest 157c882c60579617
|
||
branches: 2 memories: 1
|
||
after restart: 8 actions, digest 157c882c60579617
|
||
branches: 2 memories: 1
|
||
settings preserved: endpoint http://127.0.0.1:11434/v1 timeout 600
|
||
context inspection: ['narrator', 'persona', 'history', 'used_memories']
|
||
```
|
||
|
||
### 5.9 Packet capture — whole offline session
|
||
|
||
```text
|
||
all packets: 5131
|
||
loopback (127.0.0.0/8): 5086
|
||
non-loopback unicast: 0
|
||
TCP connections opened outside loopback: (none)
|
||
```
|
||
|
||
DNS queries, attributed by timestamp rather than assumed:
|
||
|
||
| Name | First seen | Attribution |
|
||
| --- | --- | --- |
|
||
| `fonts.googleapis.com` | 16:44:02.800449 | **the isolation probe** in §5.1 |
|
||
| `openaipublic.blob.core.windows.net` | 16:44:02.802749 | the isolation probe |
|
||
| `openrouter.ai` | 16:44:02.803260 | the isolation probe |
|
||
| `api.openai.com` | 16:44:02.803956 | the isolation probe |
|
||
| `github.com` | 16:44:02.804747 | the isolation probe |
|
||
| `ollama-container` | 18:28:55.684172 | the endpoint-policy edge-case test (§7) |
|
||
| `ollama.com` | 16:53:29 … 18:28:05 | **the Ollama server's own lookup** |
|
||
|
||
```text
|
||
capture window: 16:44:02.800136 .. 18:51:29.080336
|
||
first packet to the storyteller API: 16:44:47.639083
|
||
```
|
||
|
||
Every cloud/font/tokenizer lookup landed within **5 milliseconds of the capture
|
||
starting and 45 seconds before the storyteller received its first request** —
|
||
they are the deliberate probe, not the application. `ollama.com` recurs
|
||
throughout and is issued by the separately installed Ollama service, the same
|
||
observation M1 recorded.
|
||
|
||
---
|
||
|
||
## 6. Offline run — trusted-LAN Ollama over HTTPS
|
||
|
||
**Topology.** Ollama 0.33.0 on a **second physical machine** on the trusted LAN,
|
||
serving HTTPS with a certificate from a private CA. The storyteller ran in a
|
||
container with `NET_ADMIN`, its **default route deleted** and replaced with a
|
||
route to the LAN subnet only, its resolver pointed at nothing, and the host name
|
||
supplied as a static hosts entry.
|
||
|
||
### 6.1 Isolation and binding
|
||
|
||
```text
|
||
=== isolation ===
|
||
blocked 1.1.1.1:443 OSError
|
||
blocked 140.82.121.4:443 OSError
|
||
no resolution github.com
|
||
no resolution openrouter.ai
|
||
no resolution fonts.gstatic.com
|
||
=== the approved LAN host ===
|
||
inference.lan -> 192.168.0.50
|
||
TLS verified, peer CN = inference.lan
|
||
=== listeners ===
|
||
LISTEN 127.0.0.1:8000
|
||
```
|
||
|
||
Nothing listens on the container's own LAN-facing address.
|
||
|
||
### 6.2 Diagnostics and generation
|
||
|
||
```text
|
||
endpoint: https://inference.lan:8443/v1 | timeout: 600 s
|
||
diagnostics: {"ok": true, "models": ["qwen2.5:3b-instruct", "nomic-embed-text:latest"]}
|
||
turn 1 ( 10.7s, 24 chunks): 'You check the oil and wick, their greasy slicks staining your hand'
|
||
turn 2 ( 4.0s, 17 chunks): 'The air is thick with the musty scent of the old building as you a'
|
||
turn 3 ( 4.6s, 26 chunks): 'You find the last entry: "Exhausted... Oil\'s low..." The room feel'
|
||
|
||
memories: 1 embedded via the LAN host: 1
|
||
```
|
||
|
||
### 6.3 Every model request the application made
|
||
|
||
```text
|
||
4 https://inference.lan:8443/v1/chat/completions
|
||
1 https://inference.lan:8443/v1/embeddings
|
||
```
|
||
|
||
Narration and embeddings both went to the configured host over verified TLS. No
|
||
other URL appears.
|
||
|
||
### 6.4 Retry keeps the discarded attempt
|
||
|
||
Measured on the shipped image, against the LAN host:
|
||
|
||
```text
|
||
head AI action before retry: 8 take_count 1
|
||
retry produced (13.7s): 'You jot down a hurried note: "Visited lamp—oil low—oil lamp '
|
||
head AI action after retry: 9 take_count 2 take_index 1
|
||
alternate takes retained: 2
|
||
- 'Your pen touches the cold wood of the log, the friction'
|
||
- 'You jot down a hurried note: "Visited lamp—oil low—oil '
|
||
branches: 1
|
||
```
|
||
|
||
A retry at the tip files the new attempt as a sibling at the same coordinate and
|
||
keeps the one it replaced. No branch is created, which is correct for a retry at
|
||
the head.
|
||
|
||
### 6.5 Restart and resume
|
||
|
||
```text
|
||
before restart: 8 actions, digest 66cc82944f8b6eef
|
||
after restart: 8 actions, digest 66cc82944f8b6eef
|
||
endpoint preserved: https://inference.lan:8443/v1 timeout 600
|
||
memories: 1
|
||
```
|
||
|
||
### 6.6 Packet capture
|
||
|
||
```text
|
||
all packets: 685
|
||
loopback: 364
|
||
to/from the LAN Ollama host: 308
|
||
any other unicast: 0
|
||
DNS queries: none
|
||
```
|
||
|
||
Zero DNS: the endpoint was configured, not discovered.
|
||
|
||
---
|
||
|
||
## 7. Endpoint policy
|
||
|
||
Run inside the shipped image:
|
||
|
||
```text
|
||
ALLOW same-host loopback http://127.0.0.1:11434/v1
|
||
ALLOW localhost name http://localhost:11434/v1
|
||
ALLOW IPv6 loopback http://[::1]:11434/v1
|
||
ALLOW LAN literal http://192.168.1.50:11434/v1
|
||
ALLOW link-local http://169.254.10.5:11434/v1
|
||
ALLOW CGNAT / mesh VPN http://100.64.3.4:11434/v1
|
||
ALLOW IPv6 unique-local http://[fd00::5]:11434/v1
|
||
ALLOW docker internal name http://ollama-container:11434/v1
|
||
REJECT cloud provider https://openrouter.ai/api/v1
|
||
openrouter.ai is a cloud inference service — this build talks to Ollama on your own mach…
|
||
REJECT cloud provider https://api.openai.com/v1
|
||
REJECT public IPv4 http://8.8.8.8:11434/v1
|
||
8.8.8.8 resolves to 8.8.8.8, which is a public Internet address — …
|
||
REJECT public IPv6 http://[2001:4860:4860::8888]/v1
|
||
REJECT unspecified address http://0.0.0.0:11434/v1
|
||
0.0.0.0 … is not on this machine and not on your own network …
|
||
REJECT wrong scheme ftp://127.0.0.1/v1
|
||
```
|
||
|
||
Over the HTTP API, from inside the offline container:
|
||
|
||
```text
|
||
=== cloud and public endpoints are refused ===
|
||
400 https://openrouter.ai/api/v1 That endpoint can't be used — openrouter.ai is a cloud…
|
||
400 https://api.openai.com/v1 …
|
||
400 https://api.groq.com/openai/v1 …
|
||
400 http://8.8.8.8:11434/v1 …8.8.8.8 … is a public Internet address…
|
||
=== local and LAN endpoints are accepted ===
|
||
200 http://127.0.0.1:11434/v1
|
||
200 http://192.168.1.50:11434/v1
|
||
200 http://[::1]:11434/v1
|
||
```
|
||
|
||
Request-time enforcement against a hand-edited database is in §5.7.
|
||
|
||
`test_endpoint_policy.py` (31 collected) resolves hostnames through a stub, so
|
||
it exercises the policy rather than the machine's DNS. It covers the split-horizon
|
||
case: a name resolving to both `192.168.1.50` and a public address is refused.
|
||
|
||
---
|
||
|
||
## 8. TLS trust
|
||
|
||
```console
|
||
$ grep -rn "verify=False\|ssl._create_unverified\|CERT_NONE\|check_hostname = False" backend/app
|
||
(the only match is prose in tlstrust.py saying such an option deliberately does not exist)
|
||
|
||
$ grep -rn "httpx.AsyncClient(\|httpx.Client(" backend/app
|
||
backend/app/routers/settings.py:122
|
||
backend/app/providers/openai_compatible.py:214
|
||
backend/app/providers/openai_compatible.py:326
|
||
backend/app/providers/openai_compatible.py:360
|
||
```
|
||
|
||
Four clients, all four passing `verify=tlstrust.ssl_context()`. Enforced by
|
||
`test_tls_trust.py`, which walks the AST of the two modules that make outbound
|
||
requests and fails if any `httpx.AsyncClient` is constructed without `verify`,
|
||
or if a third module starts making requests:
|
||
|
||
```console
|
||
$ .venv/bin/python -m pytest tests/test_tls_trust.py -q
|
||
6 passed in 0.20s
|
||
```
|
||
|
||
Runtime confirmation of certificate **and** hostname verification against the
|
||
real private-CA host is §6.1 (`TLS verified, peer CN = inference.lan`), and of
|
||
all four paths — settings/test, narration, model listing, embeddings — §6.2–6.3.
|
||
|
||
### Modules that can open an outbound connection at all
|
||
|
||
```console
|
||
$ grep -rln "httpx\|requests\.\|urllib.request\|socket\.\|aiohttp\|websocket" backend/app
|
||
backend/app/endpoints.py # getaddrinfo only, for the policy
|
||
backend/app/providers/openai_compatible.py
|
||
backend/app/routers/settings.py
|
||
backend/app/tlstrust.py # builds the SSL context; opens nothing
|
||
```
|
||
|
||
Remaining absolute URLs anywhere in backend application code:
|
||
|
||
```text
|
||
3 http://127.0.0.1
|
||
2 http://localhost
|
||
1 https://openaipublic.blob.core.windows.net <- a comment and a SOURCE_URL
|
||
constant in encoding.py; never fetched
|
||
```
|
||
|
||
---
|
||
|
||
## 9. Defects found during this review
|
||
|
||
Three, all found by running the shipped build rather than by reading it. Each was
|
||
corrected because the M2 evidence could not otherwise be accurate; each is
|
||
reported rather than absorbed.
|
||
|
||
### 9.1 The memory bank was broken — summaries and embeddings silently stopped
|
||
|
||
```text
|
||
Task exception was never retrieved
|
||
future: <Task finished coro=<run_post_turn() ...> exception=AttributeError(
|
||
"'Settings' object has no attribute 'api_key_plain'")>
|
||
File "/app/backend/app/memorybank.py", line 155, in summary_provider
|
||
settings.api_key_plain,
|
||
```
|
||
|
||
M2 removed `Settings.api_key_plain` with the API key, but `memorybank`'s two
|
||
provider factories still read it. It failed in a fire-and-forget background
|
||
task, so nothing surfaced to the user and no test caught it: every memory test
|
||
stubs those factories out. **All 604 tests passed with this defect present.**
|
||
|
||
Fixed by rewriting both factories. Covered by a new test that constructs every
|
||
provider factory from a real `Settings` row.
|
||
|
||
### 9.2 The configurable model timeout never reached the turn engine
|
||
|
||
`settings.model_timeout_seconds` was stored, validated, exposed in the API and
|
||
rendered in the UI — and not passed to `OpenAICompatibleProvider` in
|
||
`turns.py` or `chat.py`, so generation used the module default. The M2 exit
|
||
criterion "no longer an undocumented hardcoded limitation" was therefore only
|
||
half met: the constant had moved, but the setting was inert.
|
||
|
||
Fixed in both call sites. Covered by a new test that drives the turn endpoint,
|
||
the chat endpoint and the summariser factory and asserts the configured value
|
||
reaches each.
|
||
|
||
### 9.3 `requirements.lock` still pinned the removed packages
|
||
|
||
```console
|
||
$ grep -n "quickjs\|psycopg\|cryptography" backend/requirements.lock
|
||
25:cryptography==50.0.1
|
||
36:psycopg==3.3.5
|
||
37:psycopg-binary==3.3.5
|
||
45:quickjs==1.19.4
|
||
```
|
||
|
||
`DEVELOPMENT.md` tells a new developer to install from the lock, which would have
|
||
reinstalled all three. Regenerated: 40 pins → 34.
|
||
|
||
### Verification after the fixes
|
||
|
||
```console
|
||
$ .venv/bin/python -m pytest tests/ -q
|
||
606 passed, 1 warning in 126.29s
|
||
$ npm run lint # exit 0
|
||
$ npm run build # ✓ built in 496ms
|
||
$ docker build . # exit 0
|
||
```
|
||
|
||
All §5 and §6 runtime evidence above was produced by an image built **after**
|
||
these fixes.
|
||
|
||
---
|
||
|
||
## 10. M1 capability regression checks
|
||
|
||
| M1 capability | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| Vendored tokenizer, no first-turn download | PASS | §10.1 |
|
||
| Self-hosted fonts | PASS | §10.2 |
|
||
| Same-origin runtime assets | PASS | §10.2 |
|
||
| Restrictive CSP | PASS | §10.2 |
|
||
| Loopback storyteller default | PASS | §5.2, §6.1, §11 |
|
||
| Same-host Ollama | PASS | §5.3–5.4 |
|
||
| Trusted-LAN Ollama | PASS | §6.2 |
|
||
| HTTPS / private-CA trusted-LAN | PASS | §6.1–6.3 |
|
||
| Certificate + hostname verification | PASS | §6.1, §8 |
|
||
| Offline story generation | PASS | §5.1, §5.4 |
|
||
| SQLite persistence | PASS | §5.8, §6.4 |
|
||
| Full regression suite green | PASS | §4 |
|
||
|
||
### 10.1 Tokenizer, inside the shipped image, offline
|
||
|
||
```text
|
||
vendored table sha256 matches pin: True
|
||
count_tokens with sockets blocked: 6 tokens
|
||
```
|
||
|
||
### 10.2 Browser asset graph, fetched over loopback with no route out
|
||
|
||
```text
|
||
CSP: default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline';
|
||
font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none';
|
||
base-uri 'none'; form-action 'self'; frame-ancestors 'none'
|
||
|
||
referenced by index.html:
|
||
/favicon.svg LOCAL
|
||
/assets/index-*.js LOCAL
|
||
/assets/index-*.css LOCAL
|
||
|
||
font urls inside the stylesheet:
|
||
/fonts/cinzel-normal-latin-ext.woff2 200 font/woff2 14540
|
||
/fonts/cinzel-normal-latin.woff2 200 font/woff2 25904
|
||
/fonts/crimson-pro-italic-latin-ext.woff2 200 font/woff2 39808
|
||
/fonts/crimson-pro-italic-latin.woff2 200 font/woff2 51432
|
||
/fonts/crimson-pro-normal-latin-ext.woff2 200 font/woff2 37988
|
||
/fonts/crimson-pro-normal-latin.woff2 200 font/woff2 48200
|
||
/fonts/inter-normal-latin-ext.woff2 200 font/woff2 85068
|
||
/fonts/inter-normal-latin.woff2 200 font/woff2 48256
|
||
```
|
||
|
||
Unchanged from M1, including the `object-src`/`base-uri`/`form-action`
|
||
directives M1 added.
|
||
|
||
---
|
||
|
||
## 11. Storyteller network exposure
|
||
|
||
| Launch path | Listener / publish | Verified by |
|
||
| --- | --- | --- |
|
||
| `./start.sh` (native dev) | `uvicorn --host 127.0.0.1 --port 8000` | file assertion in `test_local_only_surface.py` |
|
||
| `start.ps1` (Windows) | `--host 127.0.0.1` | same test |
|
||
| Production native (documented in `DEVELOPMENT.md`) | `--host 127.0.0.1` | documentation |
|
||
| `docker compose up` | `ports: - "127.0.0.1:8000:8000"` | test parses the published mappings and asserts every one starts `127.0.0.1:` |
|
||
| Container process itself | `--host 0.0.0.0` | deliberate; the only address a published port can reach |
|
||
|
||
`DEVELOPMENT.md` states that publishing the port to `0.0.0.0` is a deliberate
|
||
decision this project's threat model does not cover. The `Dockerfile` carries the
|
||
same warning next to `EXPOSE`.
|
||
|
||
Two supporting checks:
|
||
|
||
* `AIDND_CORS_ORIGINS="*"` makes the application **refuse to start**:
|
||
`RuntimeError: AIDND_CORS_ORIGINS must not contain '*'…`
|
||
* An unknown `/api/...` path now returns **404** instead of the SPA's HTML with
|
||
status 200 (verified for ten removed endpoints, §12).
|
||
|
||
---
|
||
|
||
## 12. Removed-surface verification, at runtime
|
||
|
||
From inside the offline container:
|
||
|
||
```text
|
||
=== removed surfaces answer 404 ===
|
||
404 GET /auth/me 404 GET /analytics/summary
|
||
404 GET /auth/login 404 GET /analytics/collect
|
||
404 GET /auth/register 404 GET /analytics/access
|
||
404 GET /auth/logout 404 GET /scripts
|
||
404 GET /adventures/2/scripts
|
||
404 GET /adventures/2/script-state
|
||
|
||
=== no API key in the settings surface ===
|
||
api_key present: False has_api_key present: False
|
||
after trying to set one: False
|
||
```
|
||
|
||
```console
|
||
$ python -c "import app.scripting"
|
||
ImportError # asserted by test_the_application_has_no_scripting_engine
|
||
$ grep -rn "import quickjs" backend/app
|
||
(no output)
|
||
```
|
||
|
||
---
|
||
|
||
## 13. Inert schema retained for compatibility
|
||
|
||
Verified present in the database and unread by the application.
|
||
|
||
| Object | Was | Status |
|
||
| --- | --- | --- |
|
||
| table `scripts` | script library | unmapped; no model, no query |
|
||
| table `adventure_scripts` | per-adventure script copies | unmapped |
|
||
| table `analytics_daily` | visitor counters | unmapped |
|
||
| table `analytics_visitor_days` | visitor funnel | unmapped |
|
||
| table `access_log` | sign-ins, addresses, devices | unmapped |
|
||
| column `adventures.script_state` | scripting shared state | written `{}` only |
|
||
| column `actions.state_after` | scripting state per node | written `{}` only |
|
||
| column `settings.api_key` | encrypted cloud key | never read or written |
|
||
| column `settings.reasoning_max_tokens` | OpenRouter thinking budget | never read |
|
||
| column `users.demo_turns_used` / `_date` | demo cap tally | never read |
|
||
| table `users` + `user_id` FKs | multi-user ownership | **active**, one row, internal identity only |
|
||
|
||
No destructive migration was performed. One additive migration was added:
|
||
|
||
```python
|
||
(77, "ALTER TABLE settings ADD COLUMN model_timeout_seconds INTEGER NOT NULL DEFAULT 300")
|
||
```
|
||
|
||
Existing M1 databases open unchanged: the offline run in §5 used the volume
|
||
carrying campaigns created before these fixes, and §5.8 shows the transcript
|
||
digest surviving both an image replacement and a restart.
|
||
|
||
---
|
||
|
||
## 14. What was not measured
|
||
|
||
Stated so the implementation report does not overclaim.
|
||
|
||
1. **No browser rendered the UI.** No browser automation was available. §10.2
|
||
fetches the complete asset graph the page references, with correct media
|
||
types and the CSP, but nobody looked at the rendered page. This is unchanged
|
||
from M1, where the maintainer confirmed the render by hand.
|
||
2. **Sustained multi-turn play under the memory bank was not completed on this
|
||
host.** Turns 1–3 succeeded; turn 4 timed out under third-party CPU load
|
||
(§5.4). The individual capabilities — narration, streaming, embeddings,
|
||
memory writes, retry, fork, restart — were each measured separately.
|
||
3. **No load, soak or long-campaign testing.** Out of scope for M2.
|
||
|
||
---
|
||
|
||
# 15. Closeout note — appended 2026-09-03
|
||
|
||
**This section was appended after the fact and is not part of the original
|
||
evidence record.** Everything above was written on 2026-09-02 and describes the
|
||
repository as it stood then. Nothing above has been rewritten, including its
|
||
references to a working tree that was dirty at the time.
|
||
|
||
## What the original report correctly said
|
||
|
||
The report above correctly described commit `8c65ae9` — the M2 implementation
|
||
commit — as **not containing** the three fixes this review found. At the time it
|
||
was written, those fixes existed only in the working tree, six files ahead of
|
||
that commit. Every runtime measurement in this report was produced by an image
|
||
built from that fixed tree, which the report states plainly.
|
||
|
||
## What happened afterwards
|
||
|
||
The six-file correction was committed:
|
||
|
||
```text
|
||
8652fe7cd84bca5173abb03b2a692f15fea8a98c M2 review: two regressions the green suite hid, and the reports
|
||
```
|
||
|
||
That commit carries the six implementation/test/lockfile files **and** the two
|
||
M2 reports themselves, which is why this report's own history begins there. The
|
||
provenance distinction the review asked for is preserved regardless:
|
||
`8c65ae9` is the M2 implementation, `8652fe7` is the review correction, and the
|
||
two were never squashed.
|
||
|
||
## Verification at closeout
|
||
|
||
Re-run on 2026-09-03 against the committed tree — not the working tree — so that
|
||
what was verified is exactly what the repository contains:
|
||
|
||
```text
|
||
backend .venv/bin/python -m pytest tests/ -q 606 passed in 133.26s
|
||
frontend npm run lint 7 warnings, 0 errors, exit 0
|
||
frontend npm run build built in 932ms; index.js 395.41 kB
|
||
root docker build -t storyteller-m2-closeout . exit 0
|
||
git status --short clean
|
||
```
|
||
|
||
The 606 figure matches the count this report recorded, from the same tree.
|
||
|
||
Dependency removal was verified in the built image rather than only in the
|
||
lockfile:
|
||
|
||
```text
|
||
$ docker run --rm --entrypoint sh storyteller-m2-closeout -c "pip list | grep -iE 'quickjs|psycopg|cryptography|cffi|pycparser'"
|
||
ABSENT: none of quickjs/psycopg/cryptography/cffi/pycparser installed
|
||
32 packages total
|
||
```
|
||
|
||
The network measurements in §5, §6 and §10 were **not** re-run. The six
|
||
corrected files change provider construction and a timeout value; they do not
|
||
touch the endpoint policy, the TLS trust context, or the bind addresses, so the
|
||
captures above remain the evidence for the network boundary.
|