# Development and local operation This is the Adventure Storyteller production fork of AI-DnD. `PROVENANCE.md` records where the code came from; `planning/` holds the product specification and milestone plan. Everything here assumes the local-only rule from `planning/DECISIONS/004-local-only-production.md`: after setup, ordinary story play must work with **no Internet access at all**. Setup itself downloads dependencies and models; playing does not. ## Versions this was built and tested on | | | | --- | --- | | OS | Linux (Ubuntu 24.04 userland), x86-64, 4 cores, 15 GB RAM, no GPU | | Python | 3.12.3 | | Node | 22.23.1, npm 10.9.8 (the Dockerfile builds the SPA on Node 24) | | Ollama | `ollama/ollama:latest` in Docker | | Models | `qwen2.5:3b-instruct` (narrator), `nomic-embed-text` (memory bank) | ## Setup ```bash # Backend, from the exact tested dependency closure. python3 -m venv backend/.venv backend/.venv/bin/pip install -r backend/requirements.lock # Frontend. cd frontend && npm ci && cd .. ``` `backend/requirements.lock` pins every version, transitive ones included. `backend/requirements.txt` states the ranges the code actually needs and stays the file you edit; regenerate the lock after a deliberate upgrade (the header in the lock says how). This is the only step that needs the Internet. It downloads Python and npm packages; it does **not** download a tokenizer or a font, because both are vendored in the tree — see "What was made offline-safe" below. You also need the models, once: ```bash ollama pull qwen2.5:3b-instruct ollama pull nomic-embed-text # only if you want the memory bank ``` There is no account to create and nothing to log in to. The application is single-user: whoever can reach it on loopback is its owner. ## Running **Development** — backend on `:8000`, Vite dev server on `:5173`: ```bash ./start.sh ``` **Production-shaped** — one server, SPA served by FastAPI: ```bash cd frontend && npm run build && cd .. cd backend && .venv/bin/uvicorn app.main:app --host 127.0.0.1 --port 8000 ``` Then open . **Docker**: ```bash docker compose up --build ``` ### The listener is loopback, and stays loopback `start.sh`, `start.ps1` and the production command above all pass `--host 127.0.0.1` explicitly. `docker-compose.yml` publishes `127.0.0.1:8000:8000` — the process inside the container listens on `0.0.0.0` because a published port cannot reach anything else, but the port is only bound on the host's loopback. That is a requirement, not a preference. In local mode the storyteller API is single-user and unauthenticated: anything that can reach it can read and rewrite every campaign. Putting Ollama on another machine (below) does **not** change this — it is an outbound connection and needs no inbound exposure. If you publish the port to `0.0.0.0` anyway, you have made a deliberate decision that this project's threat model does not cover (`planning/SECURITY-THREAT-MODEL.md`). ## Pointing the storyteller at Ollama The endpoint, the model and the generation parameters are **runtime settings stored in the database**, not environment variables. There is no API key field: M2 removed it along with the cloud providers, and Ollama does not use one. Set them on the app's Settings page, or with one request: ```bash curl -X PUT http://127.0.0.1:8000/api/settings \ -H 'Content-Type: application/json' \ -d '{"endpoint_url":"http://127.0.0.1:11434/v1","model":"qwen2.5:3b-instruct", "api_mode":"chat","max_output_tokens":200, "context_token_budget":4096}' ``` `POST /api/settings/test` (the **Test connection** button) returns `{"ok": true, "models": [...]}` and is the fastest way to tell a wrong endpoint from a missing model. When it fails it says which kind of failure it was, and they need different things done about them: | `kind` | What it means | | --- | --- | | `rejected` | The endpoint is outside the policy below. Not a network problem. | | `unreachable` | Nothing answered. Ollama is not running there, or the port is wrong. | | `tls` | The certificate did not verify — install the CA (see below). | | `timeout` | It accepted the connection and then said nothing. | | `http` | It answered with an error status; the body is included. | A successful test also warns when the endpoint is reachable but has no model by the configured name, which is the commonest way for a correct endpoint to still fail every turn. ### Which endpoints are allowed `backend/app/endpoints.py` decides, and it is deliberately narrow: **loopback, your own LAN, or nothing.** The allowed networks are `127.0.0.0/8`, the three RFC1918 ranges, link-local, IPv6 loopback and unique-local, and `100.64.0.0/10` (carrier-grade NAT, which is what a mesh VPN such as Tailscale hands out). Every address the endpoint's hostname resolves to must be in one of them. A public address is refused, a name resolving to both a private and a public address is refused, and known cloud inference hosts are refused by name so the error says why rather than looking like a DNS fault. The rule is applied when you save the endpoint *and* again before every outbound request, so a database edited by hand or a hostname that starts resolving somewhere new cannot turn a local install into an exfiltration path. There is no setting to relax it. ### Same host (the default) ```text endpoint_url = http://127.0.0.1:11434/v1 ``` Nothing else to do. Ollama's own default is to listen on loopback. ### An Ollama on another machine on your trusted LAN Supported and explicitly configured — never guessed, never discovered. On the **inference machine**, tell Ollama to accept connections from the LAN, because it binds loopback by default: ```bash OLLAMA_HOST=0.0.0.0:11434 ollama serve ``` On the **storyteller machine**, set the endpoint to that host's address: ```text endpoint_url = http://192.168.1.50:11434/v1 ``` Use an IP address or a name your own network resolves. Then: - the storyteller UI/API stays on `127.0.0.1` — do not change the listener; - prompts, story text, retrieved memories and embedding inputs all travel to that host, so it has to be one you control, on a network you trust; - the inference machine needs the models installed, not the storyteller; - no Internet is involved in either direction. A LAN endpoint is accepted because it is on one of the allowed networks above. Nothing else about it is special. #### If that endpoint is HTTPS with your own CA Some inference hosts are only reachable over TLS. A StartOS server is one: it serves Ollama over HTTPS with a certificate from its own local CA, and plain HTTP redirects to it. Install that CA on the machine running the storyteller, the same way you would for the browser — on Debian and Ubuntu: ```bash sudo cp your-ca.crt /usr/local/share/ca-certificates/ sudo update-ca-certificates ``` then use the `https://` URL and the hostname the certificate is issued for: ```text endpoint_url = https://inference.lan:8443/v1 ``` The application verifies against the machine's CA store **and** the `certifi` bundle (`backend/app/tlstrust.py`), so a CA you installed at the OS level is honoured, exactly as `curl` and your browser honour it. Public certificates keep working unchanged. There is deliberately **no** setting to skip verification. If a connection is refused with `CERTIFICATE_VERIFY_FAILED`, the CA is not installed where the storyteller can see it, or the URL's hostname does not match the certificate — `openssl s_client -connect host:port` will say which. In a container, remember the CA has to be inside the image or bind-mounted; the host's store is not visible from within. ## Tests ```bash cd backend && .venv/bin/python -m pytest tests/ -q # 680 tests cd frontend && npm run lint && npm run build ``` Two files are the M1 regression guards. `test_offline_assets.py` fails if the tokenizer starts fetching its table again, if a remote font or stylesheet comes back, or if the CSP names a remote origin. Two of its checks read the built SPA under `frontend/dist/` and skip when it has not been built, so run `npm run build` before treating a green suite as complete evidence. `test_tls_trust.py` fails if outbound verification is weakened, if a public CA is lost from the union, or if a new HTTP client is added without the shared verification context. M4 added `test_save_points.py`, which fails if restoring a Save Point starts deleting history, stops going through the active head, forks on its own, or lets a Save Point on one campaign be restored through another. M2 added two more. `test_endpoint_policy.py` fails if the set of reachable addresses widens, or if either place the rule is applied stops applying it — it resolves hostnames through a stub, so it tests the policy rather than whatever DNS the machine has. `test_local_only_surface.py` fails if a removed subsystem comes back as a route, if an API key becomes settable again, if the model timeout stops being configurable or becomes unbounded, or if a supported start path stops binding loopback. ## What was made offline-safe, and how to check Two runtime downloads were removed in Milestone M1. Both were invisible on a machine that had been online once, which is exactly why they need tests. **The tokenizer.** `tiktoken.get_encoding("cl100k_base")` downloads a 1.7 MB BPE table on first use, and the context builder counts tokens on every turn, so the first story turn on an air-gapped install died with a `ConnectionError`. The table is vendored at `backend/app/context/vendor/cl100k_base.tiktoken` and `backend/app/context/encoding.py` builds the encoding from it, verifying its SHA-256 against the digest `tiktoken` itself pins. **The fonts.** The SPA linked `fonts.googleapis.com` from `index.html`, so every page load fetched a stylesheet and font files from Google. The three families are self-hosted under `frontend/public/fonts/`, declared in `frontend/src/styles/fonts.css`, and re-vendored by `python3 frontend/tools/vendor_fonts.py`. The CSP in `backend/app/main.py` now names no remote origin at all. To convince yourself on a machine that has already been online, run the app with no route out rather than trusting a cold cache: ```bash docker network create --internal offline docker run -d --name ollama --network offline -v ollama-models:/root/.ollama ollama/ollama docker build -t storyteller . # The app shares Ollama's network namespace, so Ollama is on its loopback and # neither has a route to the Internet. docker run -d --name app --network container:ollama -v story-data:/data \ storyteller uvicorn app.main:app --host 127.0.0.1 --port 8000 docker exec app python -c "import socket; socket.create_connection(('1.1.1.1',443),timeout=4)" # -> OSError: Network is unreachable, and story turns still work ``` `planning/archive/milestone-reports/M1-BASELINE-REPORT.md` records the run this procedure is taken from, including the packet captures. ## Things still inherited from upstream M2 removed the hosted, cloud, account, analytics, Postgres/Render and QuickJS scripting surfaces outright — `PROVENANCE.md` lists exactly what went. What is left of upstream that a newcomer might report as a defect: - **Inert legacy tables and columns.** Five tables and four columns M2 emptied of meaning are still in the schema, unmapped, so an M1-era campaign database opens unchanged. Nothing reads or writes them. A cleanup migration waits for the schema to settle after M5 (`planning/BUILD-MILESTONES.md`). - **Dual-dialect migration code.** `backend/app/migrations.py` still carries SQLite/Postgres branches from upstream, although Postgres support itself is gone and SQLite is the only store. Same cleanup, same milestone. - **`.github/workflows/ci.yml`** is upstream's GitHub Actions pipeline. This repository lives on a self-hosted Gitea; the workflow is kept for provenance and is not what runs the tests here. - **No frontend tests.** `npm run lint && npm run build` is the whole frontend check. A test runner is M8's job.