Files
interactive-story/planning/archive/milestone-reports/M1-BASELINE-REPORT.md
JesseMarkowitzandClaude Opus 5 d27ee34901 Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was
authoritative. Phase 0 execution prompts sat beside the specification; four
completed milestone reports sat beside the current one; and upstream AI-DnD's
own `plan/` build log and `docs/` project site still described a hosted,
scripted, multi-user product with accounts — every screenshot in it showed a
Scripts tab and a Sign up button, none of which has existed since M2.

`planning/archive/` now holds the history and says so in its own README:
`phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and
M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied.
`planning/reports/` holds only the current milestone's report, because that is
the one M4 planning has to read; it moves to the archive when M4's replaces it.

Deleted rather than archived: the Phase 0B execution prompts and the
handoff/status/summary documents, the Phase 0A discovery and triage reports,
upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's
template boilerplate. All of it is in Git history, and the two recommendation
reports carry every conclusion the deleted research reached.

Archived documents are kept verbatim. Paths written inside them point at where
those files were when the document was written, which is the point: an evidence
record that has been quietly edited is no longer evidence.

Active documentation is corrected where it pointed at the removed trees or
described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not
touch" list had gone stale at M2 and claimed QuickJS scripting was still tested;
its test count was 604 against an actual 638. `README.md` loses the upstream CI
badge, which reported upstream's pipeline rather than this fork's, and a
reference to `backend/app/worldstate/engine.py`, a file that does not exist.
`planning/README.md` is rewritten as the documentation index.

New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
manifest of what belongs in the ChatGPT project's Sources.

Source comments referring to the deleted trees are reworded; no behaviour
changes. 638 backend tests pass, frontend lints and builds, and a reference scan
over all 48 tracked Markdown files reports no unresolved path in active
documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
2026-09-03 14:33:07 -04:00

24 KiB
Raw Permalink Blame History

M1 — Production Fork and Offline Baseline: evidence report

Date: 2026-09-01/02 Milestone: M1, planning/BUILD-MILESTONES.md Base: AI-DnD d72f7c1bda0f34fccd84afb7a25c34eb01c901de (see PROVENANCE.md) Scope note: M1 only. No M2 work was started; nothing was removed from the inherited hosted/cloud/scripting surface.

This report records what was run and what was observed. Where a condition was not reproduced exactly as the acceptance test specifies, it says so and says what was reproduced instead. Nothing untested is called a PASS.

Hostnames and LAN addresses below are placeholders — inference.lan, 192.168.0.0/24. The real ones are in the workspace's untracked notes, not in this repository. Everything else, including packet counts and digests, is verbatim.


1. Test environment

Host Ubuntu 24.04.4 LTS, x86-64, 4 cores, 15 GB RAM, no GPU
Python 3.12.3
Node / npm 22.23.1 / 10.9.8
Docker 29.7.2
Ollama 0.33.2 (ollama/ollama@sha256:020e4134285e…)
Narrator model qwen2.5:3b-instruct (qwen2.5:0.5b in one earlier run)
Second machine inference.lan (192.168.0.50), StartOS, Ollama 0.33.0 over HTTPS on 8443
Embedding model nomic-embed-text:latest
Application this fork, working tree on m1-production-baseline
Browser the maintainer's own desktop browser, against the native run (§6)

The lack of a GPU matters and shows up twice below: a cold model load on four CPU cores can exceed the application's hardcoded 120-second model timeout.

2. What was changed

Thirteen upstream files modified; twenty-two added, of which eleven are the vendored font files and their licences. No upstream file was deleted.

Change Files
Fork lineage and licence provenance merge commit 7f182a8, PROVENANCE.md
Tokenizer no longer downloads backend/app/context/encoding.py, backend/app/context/vendor/cl100k_base.tiktoken, backend/app/context/builder.py
Fonts self-hosted frontend/index.html, frontend/src/index.css, frontend/src/styles/fonts.css, frontend/public/fonts/*, frontend/tools/vendor_fonts.py, .gitattributes
CSP narrowed to same-origin backend/app/main.py
TLS verified against the machine's own CA store as well as certifi's backend/app/tlstrust.py, backend/app/providers/openai_compatible.py, backend/app/routers/settings.py, backend/requirements.txt
woff2 served with its real media type backend/app/main.py
Loopback binding stated explicitly start.sh, start.ps1, docker-compose.yml, Dockerfile
Reproducible environment backend/requirements.lock, DEVELOPMENT.md
Regression tests backend/tests/test_offline_assets.py, backend/tests/test_tls_trust.py
Research scratch ignored .gitignore

3. Regression suite

Suite Before M1 After M1
backend/tests 632 passed, 0 failed (184.7 s) 648 passed, 0 failed (212.1 s)
frontend: npm run lint — exit 0 (6 pre-existing only-export-components warnings)
frontend: npm run build — succeeds
docker build — succeeds

The 632-test baseline was captured on the pinned upstream commit before any change, and matches the Phase 0B figure. The sixteen new tests are test_offline_assets.py (10) and test_tls_trust.py (6). There are no known failing tests and no documented exceptions.

The two dist-reading tests in test_offline_assets.py skip if the SPA has not been built; the run above had it built, so they executed.

Proof the tokenizer guard is a real guard

With TIKTOKEN_CACHE_DIR pointed at an empty directory and Python's socket functions replaced:

UPSTREAM PATH raises: AssertionError socket opened
VENDORED PATH: 2 tokens, no socket opened

The vendored table's SHA-256 is 223921b76ee99bde995b7ff738513eef100fb51d18c93597a113bcffe865b2a7, identical to the digest hardcoded in tiktoken_ext/openai_public.py, and the encoding it produces was checked token-for-token against tiktoken.get_encoding over ASCII, accented text, CJK, emoji, CRLF and special-token literals.

4. Run 1 — offline, same-host Ollama

Setup. An --internal Docker network (no NAT, no external DNS). Ollama runs in one container; the production image runs in a second container that shares Ollama's network namespace, so Ollama is genuinely on the storyteller's own loopback and neither has any route out. Both were driven from a client in the same namespace, over 127.0.0.1:8000.

Isolation confirmed from inside the application container, before any test:

blocked    1.1.1.1:443                              OSError: Network is unreachable
blocked    openaipublic.blob.core.windows.net:443   gaierror
blocked    fonts.googleapis.com:443                 gaierror
blocked    fonts.gstatic.com:443                    gaierror
blocked    openrouter.ai:443                        gaierror
blocked    github.com:443                           gaierror

Listening sockets in that namespace:

127.0.0.1:8000        uvicorn  (the storyteller)
[::]:11434            ollama
127.0.0.11:44683      Docker's embedded DNS

Story play. Campaign "Continuity Test (3b)" created and played for six turns, twelve actions, all offline:

turn 1: 'The old lighthouse looms dark against the stormy sky, its silence heavy as the wind…'
turn 2: "Gripping the lantern's heavy brass handle, you flick it on and off, each attempt a futile…"
turn 3: 'The lamp room is dim and musty, the thick air choking your breath…'
turn 4: 'You find the dated handwriting, the last entry noting the storm started three days ago…'
turn 5: 'The night outside is a tempest, waves crashing against the shore with a deafening roar…'
turn 6: 'The cabinet is cold and heavy, the lock stubbornly refusing to budge…'

Browser asset graph, fetched over loopback with no route out:

GET /                        200  text/html
  /favicon.svg               200      9 522 bytes
  /assets/index-*.js         200  933 695 bytes
  /assets/index-*.css        200   59 902 bytes
    /fonts/cinzel-normal-latin.woff2            200 font/woff2   25 904
    /fonts/cinzel-normal-latin-ext.woff2        200 font/woff2   14 540
    /fonts/crimson-pro-normal-latin.woff2       200 font/woff2   48 200
    /fonts/crimson-pro-normal-latin-ext.woff2   200 font/woff2   37 988
    /fonts/crimson-pro-italic-latin.woff2       200 font/woff2   51 432
    /fonts/crimson-pro-italic-latin-ext.woff2   200 font/woff2   39 808
    /fonts/inter-normal-latin.woff2             200 font/woff2   48 256
    /fonts/inter-normal-latin-ext.woff2         200 font/woff2   85 068

Every URL index.html references is same-origin. The response carried:

content-security-policy: default-src 'self'; script-src 'self';
  style-src 'self' 'unsafe-inline'; font-src 'self'; img-src 'self' data:;
  connect-src 'self'; object-src 'none'; base-uri 'none'; form-action 'self';
  frame-ancestors 'none'
x-content-type-options: nosniff   referrer-policy: same-origin   x-frame-options: DENY

Packet capture (tcpdump in the same namespace for the whole run, 6 074 packets):

total packets:            6074
loopback (127.0.0.0/8):   6062
non-loopback unicast:        0
remainder (12):           received mDNS / ICMPv6 router solicitations from the
                          bridge — inbound multicast, not sent by this namespace
TCP connections opened:   127.0.0.1:8000  (storyteller API)
                          127.0.0.1:11434 (Ollama)
                          127.0.0.1:11499 (the deliberately dead port in A05)
                          two ephemeral loopback ports (Ollama's model runner)

One DNS observation, and it is not the storyteller. Four queries appear, all failing:

127.0.0.11.53 > … ServFail  q: A?    ollama.com.
127.0.0.11.53 > … ServFail  q: AAAA? ollama.com.   (×2 each)

ollama.com is queried by the Ollama server itself, in the shared namespace, not by the storyteller. It failed, nothing depended on it, and no story data could have been in it. It is recorded here rather than dismissed: on a deployment with a route to the Internet, the inference server has its own outbound behaviour, and the storyteller's local-only guarantee does not extend to it. Confirming and, if wanted, suppressing that is an Ollama configuration question — worth settling before release, and out of M1's scope.

5. Run 2 — offline, Ollama on a second machine on the LAN

Setup. Ollama 0.33.0 on inference.lan (192.168.0.50), a separate physical machine on the trusted LAN, serving HTTPS on port 8443 with a certificate issued by CN = StartOS Local Intermediate CA. qwen2.5:3b-instruct and nomic-embed-text were installed there for this run.

The storyteller runs in a container on this host with NET_ADMIN, its default route deleted and replaced by a route to 192.168.0.0/24 only, its resolver pointed at nothing, and inference.lan supplied as a static hosts entry. So: the LAN is reachable, the Internet is not, and the endpoint is configured explicitly rather than discovered. Uvicorn binds 127.0.0.1:8000 inside that container.

route table:  172.17.0.0/16 dev eth0 …
              192.168.0.0/24 via 172.17.0.1 dev eth0      (no default route)
listeners:    LISTEN 127.0.0.1:8000                        (nothing on 172.17.0.3)

Isolation, checked from inside before anything else:

outbound Internet
  blocked   1.1.1.1:443         OSError: Network is unreachable
  blocked   140.82.121.4:443    OSError: Network is unreachable
  blocked   104.16.0.1:443      OSError: Network is unreachable
name resolution
  no resolution  github.com / openrouter.ai / fonts.gstatic.com /
                 openaipublic.blob.core.windows.net        (gaierror)
the approved LAN host
  inference.lan -> 192.168.0.50
  TLS OK, peer CN = inference.lan

Model discovery over the LAN endpoint:

{"ok": true, "models": ["qwen2.5:3b-instruct", "nomic-embed-text:latest"]}

Story play. Nine turns, then a tenth after a restart. That machine has faster hardware than this one, and it shows:

turn 1 (13.0s): "The lamp flickers faintly, a single coal barely keeping the structure's shadow at bay."
turn 2 ( 4.9s): "My lantern's light has failed. Checking the wick, I find it's burnt down to a stub."
turn 3 ( 4.4s): 'I turn the wick higher, attempting to coax more flame from the almost-embers.'
turn 4 ( 5.9s): 'The storm shutter blocks out everything but the dark ocean…'
turn 5 ( 6.8s): 'I insert the key and hear the satisfying click as the lock turns…'
turn 6 ( 8.1s): 'Inside the box, I find a small vial of oil and a note warning…'
turn 7 (10.6s) … turn 9 (14.8s)

State extraction, summaries and embeddings ran through the same endpoint: the memory bank produced memories and embedded them with nomic-embed-text on the remote host, so the narrator, the summarizer and the embedder all went over the LAN, not just the narrator.

Restart and resume. The storyteller container was restarted while the inference host was left alone:

before restart: 18 actions, head 18, digest f24caf86a744ab36
after  restart: 18 actions, head 18, digest f24caf86a744ab36
                endpoint still https://inference.lan:8443/v1
                embedding model still nomic-embed-text:latest, memories intact
next turn (6.4s): 'You make your way back to the lighthouse, lantern in hand…'

Packet capture across the whole run:

all packets:                          1633
loopback (127.0.0.0/8):                730
to/from inference.lan 192.168.0.50:    893
any other unicast:                       0

TCP connections opened:  192.168.0.50:8443  (19)
                         127.0.0.1:8000        (17)

DNS queries:  none — no name was looked up at all

Zero packets to anything but the storyteller's own loopback and the approved inference host, and not a single DNS query, because the endpoint was configured rather than resolved.

5.1 What had to be fixed to get here: TLS trust

The first attempt failed, and the failure was the application's:

{"ok": false, "detail": "Connection failed: [SSL: CERTIFICATE_VERIFY_FAILED]
 certificate verify failed: self-signed certificate in certificate chain"}

httpx verifies against the certifi bundle, which carries the public web's CAs and nothing else. The host's certificate comes from a local StartOS CA that the user had already installed at /usr/local/share/ca-certificates/local-ca.crt, which is why curl and the browser accepted the same endpoint on the same machine. Confirmed directly:

certifi bundle (httpx default)     FAIL  SSLCertVerificationError
system trust store                 OK    peer CN=inference.lan

This is not an edge case for this product. A trusted-LAN inference host is a first-class v1 deployment (planning/DECISIONS/002-ollama-only-v1.md), and such a host is unlikely to hold a publicly-issued certificate.

backend/app/tlstrust.py builds one verification context that unions the platform CA store with certifi's bundle, and all four outbound HTTP clients use it. Deliberately a union rather than a swap: using the platform store alone would be a behaviour change, and an image with an empty or stale system store would start failing on endpoints that used to work. A union can only add trust the user has already granted at the operating-system level.

Verification itself is untouched — verify_mode=CERT_REQUIRED, check_hostname=True, and no "insecure" escape hatch was added. Measured after the change:

  OK    inference.lan    (local CA)   CN=inference.lan
  OK    github.com     (public CA)  CN=github.com
  OK    pypi.org       (public CA)  CN=pypi.org
  certifi-only context still rejects inference.lan  (so the union is what changed)

backend/tests/test_tls_trust.py holds the line in both directions: it asserts every certifi root survives in the union, that verification is not weakened, and — by walking the AST of the two modules that make outbound requests — that no fifth HTTP client is ever added without the shared context.

5.2 An earlier container-only run

Before a second machine was available, the same sequence was run with Ollama in a separate container, network namespace and IP on an --internal network. It produced the same result (all 20 model requests to the configured endpoint, zero other unicast packets) and is superseded by the run above, which has the physical separation A06 actually asks for.

6. Run 3 — native (non-Docker) run, same-host Ollama

The production build run directly from the venv, SPA served by FastAPI, Ollama on host loopback:

LISTEN 0 2048  127.0.0.1:8000   users:(("uvicorn",pid=341269,fd=14))
LISTEN 0 4096  127.0.0.1:11434

connect to 192.168.0.10:8000 (this host's LAN address)  ->  Connection refused
GET http://127.0.0.1:8000/                                ->  200

A campaign was created and played on this instance too. This run had Internet available — its purpose was the native listener check and a real end-to-end run outside Docker, not the offline proof, which Runs 1 and 2 carry.

The UI was opened in a real browser. No browser automation was available in this session, so this step was done by hand: the maintainer loaded http://127.0.0.1:8000, saw the campaign list render, clicked into the "Continuity Test (native run)" campaign, and saw the adventure and its transcript. So the inherited SPA loads, runs, and talks to the API from a browser — the last M1 claim that had been resting on inference rather than observation.

Two limits on what that shows, stated so the evidence is not read wider than it is. It was this native run, which had Internet available, so it is not itself offline evidence; the offline proof that the page needs nothing remote is the asset-graph fetch in §4, made with no route out. And the browser's network panel was not inspected, so "the browser requested nothing external" is carried by §4 and by the CSP — which names no remote origin and would block one — rather than by a devtools capture.

7. Acceptance results

ID Result Evidence
A01 Start application offline PASS Run 1: app started, campaign created, six turns generated, with 1.1.1.1 unreachable and no name resolving. The "open UI" step is covered by the asset-graph fetch there and by the browser render in §6.
A02 Storyteller loopback default PASS Run 1 and Run 2 listeners are 127.0.0.1:8000 only; Run 3 refuses connections on the host's LAN address; docker-compose.yml publishes to 127.0.0.1.
A03 No cloud API key PASS api_key empty in every run; connection test, turns, summaries and embeddings all succeeded.
A04 Campaign survives restart PASS Run 1: 12 actions before and after a container restart, identical transcript and head. Run 2: identical digest dd59e2a564df4e62 across restart, then play resumed.
A05 Failed model call does not corrupt story PASS Two induced failures (nonexistent model; dead endpoint port) plus one natural timeout. Accepted-prefix digest 2ca6ab528178e44e unchanged through all of it; AI-action count stayed at 6; recovery by continue produced turn 7 with the prefix still unchanged. See §7.1.
A06 Trusted-LAN Ollama inference PASS Run 2 — Ollama on inference.lan, a second physical machine on the trusted LAN, over verified HTTPS, with the storyteller's Internet route removed and its listener on 127.0.0.1. Nine turns plus a post-restart turn, summaries and embeddings included. Capture: 893 packets to the approved host, 730 loopback, zero elsewhere, zero DNS queries. Required a TLS trust fix first — §5.1.
H01 No unexpected outbound connections PASS for the application Zero non-loopback unicast packets in Run 1; in Run 2, zero packets outside loopback and the approved host and zero DNS queries of any kind. One caveat, not the storyteller's: Ollama itself queried ollama.com (§4).
H02 No telemetry PASS No outbound destination in either capture; the inherited analytics.py writes to two local SQLite tables and opens no socket (Phase 0B static analysis, re-confirmed by the captures).
H03 No cloud provider required PASS as stated Nothing cloud was reachable in Runs 1 and 2 and everything worked. The test's preferred final state — "cloud provider controls are absent, not merely unused" — is not met and is M2's scope by design.
H11 No first-use runtime asset download PASS First turn on a fresh database succeeded with no route out; the whole browser asset graph resolved same-origin; the tokenizer is built from a vendored, digest-checked table.

7.1 A05 in detail, including a real behaviour worth knowing

Baseline: 12 actions, 6 of them AI, digest 2ca6ab528178e44e.

invalid local model   -> events ['player','error']
                         "Endpoint or model not found (HTTP 404) … model
                          'no-such-model-v9' not found"
                         13 actions, still 6 AI actions
endpoint down (:11499)-> events ['player','error']
                         "Could not connect to http://127.0.0.1:11499/v1 —
                          is the AI server running?"
                         14 actions, still 6 AI actions
accepted prefix through the pre-failure head: 12 actions, digest 2ca6ab528178e44e — unchanged
the two added rows:  [28] story 'I strike a match.'
                     [29] story 'I strike a match again.'
recovery (continue)  -> 40 SSE events, done; 15 actions, 7 AI actions
                        prefix digest still 2ca6ab528178e44e

The player's own typed action is committed before the model is called (run_player_turn in backend/app/routers/adventures/turns.py), so a failed turn leaves the player's text at the head with no reply. No AI output is ever partially committed. That satisfies A05 as written — prior story intact, the failed turn not committed as accepted, retry available — and it is deliberate: it means a model failure never eats what the player typed. It is worth stating explicitly because "the story is unchanged" is not literally true; "the accepted story is unchanged" is.

8. Findings and open items

Nothing here blocks M1. Each is recorded because it was observed, not inferred.

  1. The model timeout is 120 s and hardcoded (httpx.Timeout(120, connect=10) in backend/app/providers/openai_compatible.py). On this GPU-less four-core host, a cold load of qwen2.5:3b-instruct — or two models contending after the memory bank pulls in nomic-embed-text — exceeded it three times during these runs. Once warm, a full turn took 9 s. This is presented as an environment/tuning finding, not an application defect; a configurable timeout is a small change, and changing product behaviour was outside M1's scope.
  2. Ollama queries ollama.com on its own account (§4). Outside the storyteller's code, inside the user's trust boundary. Worth settling before release: the local-only claim covers what this application sends, and a user reading a packet capture will see that query.
  3. The memory bank is off per adventure by default (auto_summarize and memory_bank_enabled are false on a new adventure) even when an embedding model is configured globally. Discovered while trying to exercise embeddings; it is inherited behaviour, not a regression.
  4. docs/*.html still links Google Fonts. That is upstream's GitHub Pages project site; it is not served by the application, not part of any build, and not covered by the runtime rule. Left alone deliberately, and the regression tests scope themselves to frontend/ so they do not give a false signal about it.
  5. A trusted-LAN endpoint over HTTPS needed a code change to work at all (§5.1), and it was found only by pointing the application at a real StartOS-hosted Ollama. Static review would not have found it: every candidate report and every local run up to that point used plain HTTP to loopback, where certificate verification never happens. It is fixed and tested, and recorded here because the class of bug — "works for curl, fails for us" — is worth remembering when M2 formalises endpoint policy.
  6. The browser's own network panel was never inspected (§6). The UI has now been rendered by hand and works, and §4 shows the page's whole asset graph resolving same-origin with no route out, so nothing rests on inference. A devtools capture during an offline session would still be the most direct form of that evidence, and costs a minute if anyone wants it.

9. Definition of done

Requirement Status
Starts with outbound Internet blocked after setup met — Runs 1 and 2
Opens the inherited browser UI locally met — rendered in a browser, campaign opened and read (§6)
Generates and persists turns through same-host Ollama met — Run 1
Generates and persists turns through a configured remote Ollama met — Run 2, on a second physical machine
Restarts and resumes the campaign met — Runs 1 and 2
Survives a failed model call without corrupting accepted state met — §7.1
No first-use tokenizer/font/runtime-asset request met — §3, §4
Storyteller UI/API loopback-bound by default met — §4, §5, §6
Regression suite passes, or failures documented met — 648 passed, 0 failed

Every line of the definition of done is met, and none of them rests on an untested assumption. No M2 work has been started.

10. Artefacts not committed

The packet captures (offline-samehost.pcap 2.7 MB, offline-lan.pcap 0.8 MB, lan-remote-host.pcap 0.4 MB) and the throwaway driver scripts live in this session's scratchpad, not in the repository. Everything drawn from them is quoted above; the procedure in DEVELOPMENT.md regenerates them.