Files
interactive-story/DEVELOPMENT.md
T
JesseMarkowitzandClaude Opus 5 b7005e6fdd M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.

This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:

    visible active transcript position == stored head == authoritative state

Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)

  A narrator edit no longer rewrites a row. It returns to the state before the
  turn, takes the reader's exact text as the accepted narration, re-derives the
  state that text implies, and becomes a new active continuation — while the
  original narration keeps its words, its live flag and its whole future as
  retained history. At the tip the correction is another take; with story below
  it, it forks. No new history machinery: this is the existing fork/take/head
  path with the reader's text in place of a generated reply. The §14A refusal
  is therefore gone for narrator turns, and remains only for player input.

Pre-M5 positions

  Migration 88 backfills the empty narrative document onto every action written
  before M5, and a missing snapshot now restores the empty document instead of
  leaving the previous position's state standing. Restoring to an old Save
  Point no longer leaves a later position's entities and facts on screen.

Narrator context

  Replayed history carries prose only; the machine-readable block is no longer
  reconstructed into past turns, where it contradicted the authoritative state
  in the same prompt. A fact withdrawn by a manual correction is now named as
  no longer true, with the reader's reason, rather than silently dropped.

Also

  - state_changes joins the action-list bulk read, removing one query per row.
  - Extraction takes only the application's own protocol payload: an ordinary
    ```json or ```python block in a story survives, and a mangled proposal
    still does not reach the reader.

Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-05 07:01:50 -04:00

321 lines
13 KiB
Markdown

# Development and local operation
This is the Adventure Storyteller production fork of AI-DnD. `PROVENANCE.md`
records where the code came from; `planning/` holds the product specification
and milestone plan.
Everything here assumes the local-only rule from
`planning/DECISIONS/004-local-only-production.md`: after setup, ordinary story
play must work with **no Internet access at all**. Setup itself downloads
dependencies and models; playing does not.
## Versions this was built and tested on
| | |
| --- | --- |
| OS | Linux (Ubuntu 24.04 userland), x86-64, 4 cores, 15 GB RAM, no GPU |
| Python | 3.12.3 |
| Node | 22.23.1, npm 10.9.8 (the Dockerfile builds the SPA on Node 24) |
| Ollama | `ollama/ollama:latest` in Docker |
| Models | `qwen2.5:3b-instruct` (narrator), `nomic-embed-text` (memory bank) |
## Setup
```bash
# Backend, from the exact tested dependency closure.
python3 -m venv backend/.venv
backend/.venv/bin/pip install -r backend/requirements.lock
# Frontend.
cd frontend && npm ci && cd ..
```
`backend/requirements.lock` pins every version, transitive ones included.
`backend/requirements.txt` states the ranges the code actually needs and stays
the file you edit; regenerate the lock after a deliberate upgrade (the header in
the lock says how).
This is the only step that needs the Internet. It downloads Python and npm
packages; it does **not** download a tokenizer or a font, because both are
vendored in the tree — see "What was made offline-safe" below.
You also need the models, once:
```bash
ollama pull qwen2.5:3b-instruct
ollama pull nomic-embed-text # only if you want the memory bank
```
There is no account to create and nothing to log in to. The application is
single-user: whoever can reach it on loopback is its owner.
## Running
**Development** — backend on `:8000`, Vite dev server on `:5173`:
```bash
./start.sh
```
**Production-shaped** — one server, SPA served by FastAPI:
```bash
cd frontend && npm run build && cd ..
cd backend && .venv/bin/uvicorn app.main:app --host 127.0.0.1 --port 8000
```
Then open <http://127.0.0.1:8000>.
**Docker**:
```bash
docker compose up --build
```
### The listener is loopback, and stays loopback
`start.sh`, `start.ps1` and the production command above all pass
`--host 127.0.0.1` explicitly. `docker-compose.yml` publishes
`127.0.0.1:8000:8000` — the process inside the container listens on `0.0.0.0`
because a published port cannot reach anything else, but the port is only
bound on the host's loopback.
That is a requirement, not a preference. In local mode the storyteller API is
single-user and unauthenticated: anything that can reach it can read and
rewrite every campaign. Putting Ollama on another machine (below) does **not**
change this — it is an outbound connection and needs no inbound exposure.
If you publish the port to `0.0.0.0` anyway, you have made a deliberate
decision that this project's threat model does not cover
(`planning/SECURITY-THREAT-MODEL.md`).
## Pointing the storyteller at Ollama
The endpoint, the model and the generation parameters are **runtime settings
stored in the database**, not environment variables. There is no API key field:
M2 removed it along with the cloud providers, and Ollama does not use one. Set
them on the app's Settings page, or with one request:
```bash
curl -X PUT http://127.0.0.1:8000/api/settings \
-H 'Content-Type: application/json' \
-d '{"endpoint_url":"http://127.0.0.1:11434/v1","model":"qwen2.5:3b-instruct",
"api_mode":"chat","max_output_tokens":200,
"context_token_budget":4096}'
```
`POST /api/settings/test` (the **Test connection** button) returns
`{"ok": true, "models": [...]}` and is the fastest way to tell a wrong endpoint
from a missing model. When it fails it says which kind of failure it was, and
they need different things done about them:
| `kind` | What it means |
| --- | --- |
| `rejected` | The endpoint is outside the policy below. Not a network problem. |
| `unreachable` | Nothing answered. Ollama is not running there, or the port is wrong. |
| `tls` | The certificate did not verify — install the CA (see below). |
| `timeout` | It accepted the connection and then said nothing. |
| `http` | It answered with an error status; the body is included. |
A successful test also warns when the endpoint is reachable but has no model by
the configured name, which is the commonest way for a correct endpoint to still
fail every turn.
### Which endpoints are allowed
`backend/app/endpoints.py` decides, and it is deliberately narrow: **loopback,
your own LAN, or nothing.** The allowed networks are `127.0.0.0/8`, the three
RFC1918 ranges, link-local, IPv6 loopback and unique-local, and `100.64.0.0/10`
(carrier-grade NAT, which is what a mesh VPN such as Tailscale hands out).
Every address the endpoint's hostname resolves to must be in one of them. A
public address is refused, a name resolving to both a private and a public
address is refused, and known cloud inference hosts are refused by name so the
error says why rather than looking like a DNS fault.
The rule is applied when you save the endpoint *and* again before every
outbound request, so a database edited by hand or a hostname that starts
resolving somewhere new cannot turn a local install into an exfiltration path.
There is no setting to relax it.
### Same host (the default)
```text
endpoint_url = http://127.0.0.1:11434/v1
```
Nothing else to do. Ollama's own default is to listen on loopback.
### An Ollama on another machine on your trusted LAN
Supported and explicitly configured — never guessed, never discovered.
On the **inference machine**, tell Ollama to accept connections from the LAN,
because it binds loopback by default:
```bash
OLLAMA_HOST=0.0.0.0:11434 ollama serve
```
On the **storyteller machine**, set the endpoint to that host's address:
```text
endpoint_url = http://192.168.1.50:11434/v1
```
Use an IP address or a name your own network resolves. Then:
- the storyteller UI/API stays on `127.0.0.1` — do not change the listener;
- prompts, story text, retrieved memories and embedding inputs all travel to
that host, so it has to be one you control, on a network you trust;
- the inference machine needs the models installed, not the storyteller;
- no Internet is involved in either direction.
A LAN endpoint is accepted because it is on one of the allowed networks above.
Nothing else about it is special.
#### If that endpoint is HTTPS with your own CA
Some inference hosts are only reachable over TLS. A StartOS server is one: it
serves Ollama over HTTPS with a certificate from its own local CA, and plain
HTTP redirects to it.
Install that CA on the machine running the storyteller, the same way you would
for the browser — on Debian and Ubuntu:
```bash
sudo cp your-ca.crt /usr/local/share/ca-certificates/
sudo update-ca-certificates
```
then use the `https://` URL and the hostname the certificate is issued for:
```text
endpoint_url = https://inference.lan:8443/v1
```
The application verifies against the machine's CA store **and** the `certifi`
bundle (`backend/app/tlstrust.py`), so a CA you installed at the OS level is
honoured, exactly as `curl` and your browser honour it. Public certificates
keep working unchanged.
There is deliberately **no** setting to skip verification. If a connection is
refused with `CERTIFICATE_VERIFY_FAILED`, the CA is not installed where the
storyteller can see it, or the URL's hostname does not match the certificate —
`openssl s_client -connect host:port` will say which. In a container, remember
the CA has to be inside the image or bind-mounted; the host's store is not
visible from within.
## Tests
```bash
cd backend && .venv/bin/python -m pytest tests/ -q # 756 tests
cd frontend && npm run lint && npm run build
```
Two files are the M1 regression guards.
`test_offline_assets.py` fails if the tokenizer starts fetching its table
again, if a remote font or stylesheet comes back, or if the CSP names a remote
origin. Two of its checks read the built SPA under `frontend/dist/` and skip
when it has not been built, so run `npm run build` before treating a green
suite as complete evidence.
`test_tls_trust.py` fails if outbound verification is weakened, if a public CA
is lost from the union, or if a new HTTP client is added without the shared
verification context.
M5 added `test_narrative_state.py`, which fails if the state stops being
genre-neutral, if an event outside the allowlist is ever applied, if a malformed
proposal mutates anything, if campaign canon stops outranking the narration, or
if a turn's narration and its state can be committed apart from each other.
`test_narrative_realistic.py` is the one suite that needs a real model, and it is
skipped unless you point it at one:
```bash
AIDND_TEST_ENDPOINT=http://127.0.0.1:11434/v1 \
AIDND_TEST_MODEL=qwen2.5:3b-instruct \
backend/.venv/bin/python -m pytest backend/tests/test_narrative_realistic.py -v -s
```
It exists because Phase 0B found that structured-state behaviour can look
correct on a small prompt and fail under a full one — and it has already earned
its place, catching a case where a model echoed its own instruction into the
narration.
M4 added `test_save_points.py`, which fails if restoring a Save Point starts
deleting history, stops going through the active head, forks on its own, lets a
Save Point on one campaign be restored through another, or lets deleting a branch
take a Save Point with it. It also fails if listing Save Points goes back to one
query per Save Point, or starts fetching narration to render the list.
`test_process_restart.py` is the durability guard: it starts the application as a
real subprocess, kills it, and starts a second one against the same database. A
Save Point that survived only because a Python object was still alive would pass
an in-process test and fail a user's restart.
M2 added two more. `test_endpoint_policy.py` fails if the set of reachable
addresses widens, or if either place the rule is applied stops applying it —
it resolves hostnames through a stub, so it tests the policy rather than
whatever DNS the machine has. `test_local_only_surface.py` fails if a removed
subsystem comes back as a route, if an API key becomes settable again, if the
model timeout stops being configurable or becomes unbounded, or if a supported
start path stops binding loopback.
## What was made offline-safe, and how to check
Two runtime downloads were removed in Milestone M1. Both were invisible on a
machine that had been online once, which is exactly why they need tests.
**The tokenizer.** `tiktoken.get_encoding("cl100k_base")` downloads a 1.7 MB
BPE table on first use, and the context builder counts tokens on every turn, so
the first story turn on an air-gapped install died with a `ConnectionError`.
The table is vendored at `backend/app/context/vendor/cl100k_base.tiktoken` and
`backend/app/context/encoding.py` builds the encoding from it, verifying its
SHA-256 against the digest `tiktoken` itself pins.
**The fonts.** The SPA linked `fonts.googleapis.com` from `index.html`, so
every page load fetched a stylesheet and font files from Google. The three
families are self-hosted under `frontend/public/fonts/`, declared in
`frontend/src/styles/fonts.css`, and re-vendored by
`python3 frontend/tools/vendor_fonts.py`. The CSP in `backend/app/main.py` now
names no remote origin at all.
To convince yourself on a machine that has already been online, run the app
with no route out rather than trusting a cold cache:
```bash
docker network create --internal offline
docker run -d --name ollama --network offline -v ollama-models:/root/.ollama ollama/ollama
docker build -t storyteller .
# The app shares Ollama's network namespace, so Ollama is on its loopback and
# neither has a route to the Internet.
docker run -d --name app --network container:ollama -v story-data:/data \
storyteller uvicorn app.main:app --host 127.0.0.1 --port 8000
docker exec app python -c "import socket; socket.create_connection(('1.1.1.1',443),timeout=4)"
# -> OSError: Network is unreachable, and story turns still work
```
`planning/archive/milestone-reports/M1-BASELINE-REPORT.md` records the run this procedure is
taken from, including the packet captures.
## Things still inherited from upstream
M2 removed the hosted, cloud, account, analytics, Postgres/Render and QuickJS
scripting surfaces outright — `PROVENANCE.md` lists exactly what went. What is
left of upstream that a newcomer might report as a defect:
- **Inert legacy tables and columns.** Five tables and four columns M2 emptied
of meaning are still in the schema, unmapped, so an M1-era campaign database
opens unchanged. Nothing reads or writes them. A cleanup migration waits for
the schema to settle after M5 (`planning/BUILD-MILESTONES.md`).
- **Dual-dialect migration code.** `backend/app/migrations.py` still carries
SQLite/Postgres branches from upstream, although Postgres support itself is
gone and SQLite is the only store. Same cleanup, same milestone.
- **`.github/workflows/ci.yml`** is upstream's GitHub Actions pipeline. This
repository lives on a self-hosted Gitea; the workflow is kept for provenance
and is not what runs the tests here.
- **No frontend tests.** `npm run lint && npm run build` is the whole frontend
check. A test runner is M8's job.