Files
JesseMarkowitz 10d8a988ca
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
DEVELOPMENT.md: keep inference-host logs in one directory
2026-10-08 05:34:53 -04:00

40 KiB

Development and local operation

This is the Adventure Storyteller production fork of AI-DnD. PROVENANCE.md records where the code came from; planning/ holds the product specification and milestone plan.

Everything here assumes the local-only rule from planning/DECISIONS/004-local-only-production.md: after setup, ordinary story play must work with no Internet access at all. Setup itself downloads dependencies and models; playing does not.

Versions this was built and tested on

OS Linux (Ubuntu 24.04 userland), x86-64, 4 cores, 15 GB RAM, no GPU
Python 3.12.3
Node 22.23.1, npm 10.9.8 (the Dockerfile builds the SPA on Node 24)
Ollama ollama/ollama:latest in Docker
Models qwen2.5:3b-instruct (narrator), nomic-embed-text (memory bank)

Setup

# Backend, from the exact tested dependency closure.
python3 -m venv backend/.venv
backend/.venv/bin/pip install -r backend/requirements.lock

# Frontend.
cd frontend && npm ci && cd ..

One runtime dependency was added in M7: python-multipart, which is Starlette's multipart form parser and is how a knowledge source is uploaded. It is pure Python, Apache-2.0, and has no dependencies of its own, so it adds nothing to audit beyond itself and no network path at all.

backend/requirements.lock pins every version, transitive ones included. backend/requirements.txt states the ranges the code actually needs and stays the file you edit; regenerate the lock after a deliberate upgrade (the header in the lock says how).

This is the only step that needs the Internet. It downloads Python and npm packages; it does not download a tokenizer or a font, because both are vendored in the tree — see "What was made offline-safe" below.

You also need the models, once:

ollama pull qwen2.5:3b-instruct
ollama pull nomic-embed-text      # only if you want the memory bank

There is no account to create and nothing to log in to. The application is single-user: whoever can reach it on loopback is its owner.

Running

Development — backend on :8000, Vite dev server on :5173:

./start.sh

Production-shaped — one server, SPA served by FastAPI:

cd frontend && npm run build && cd ..
cd backend && .venv/bin/uvicorn app.main:app --host 127.0.0.1 --port 8000

Then open http://127.0.0.1:8000.

Docker:

docker compose up --build

The listener is loopback, and stays loopback

start.sh, start.ps1 and the production command above all pass --host 127.0.0.1 explicitly. docker-compose.yml publishes 127.0.0.1:8000:8000 — the process inside the container listens on 0.0.0.0 because a published port cannot reach anything else, but the port is only bound on the host's loopback.

That is a requirement, not a preference. In local mode the storyteller API is single-user and unauthenticated: anything that can reach it can read and rewrite every campaign. Putting Ollama on another machine (below) does not change this — it is an outbound connection and needs no inbound exposure.

If you publish the port to 0.0.0.0 anyway, you have made a deliberate decision that this project's threat model does not cover (planning/SECURITY-THREAT-MODEL.md).

Pointing the storyteller at Ollama

The endpoint, the model and the generation parameters are runtime settings stored in the database, not environment variables. There is no API key field: M2 removed it along with the cloud providers, and Ollama does not use one. Set them on the app's Settings page, or with one request:

curl -X PUT http://127.0.0.1:8000/api/settings \
  -H 'Content-Type: application/json' \
  -d '{"endpoint_url":"http://127.0.0.1:11434/v1","model":"qwen2.5:3b-instruct",
       "api_mode":"chat","max_output_tokens":200,
       "context_token_budget":4096}'

POST /api/settings/test (the Test connection button) returns {"ok": true, "models": [...]} and is the fastest way to tell a wrong endpoint from a missing model. When it fails it says which kind of failure it was, and they need different things done about them:

kind What it means
rejected The endpoint is outside the policy below. Not a network problem.
unreachable Nothing answered. Ollama is not running there, or the port is wrong.
tls The certificate did not verify — install the CA (see below).
timeout It accepted the connection and then said nothing.
http It answered with an error status; the body is included.

A successful test also warns when the endpoint is reachable but has no model by the configured name, which is the commonest way for a correct endpoint to still fail every turn.

Which endpoints are allowed

backend/app/endpoints.py decides, and it is deliberately narrow: loopback, your own LAN, or nothing. The allowed networks are 127.0.0.0/8, the three RFC1918 ranges, link-local, IPv6 loopback and unique-local, and 100.64.0.0/10 (carrier-grade NAT, which is what a mesh VPN such as Tailscale hands out).

Every address the endpoint's hostname resolves to must be in one of them. A public address is refused, a name resolving to both a private and a public address is refused, and known cloud inference hosts are refused by name so the error says why rather than looking like a DNS fault.

The rule is applied when you save the endpoint and again before every outbound request, so a database edited by hand or a hostname that starts resolving somewhere new cannot turn a local install into an exfiltration path. There is no setting to relax it.

A future media provider would be held to a stricter rule

The same file decides, plus one extra condition. A media endpoint — a local image or speech generator, when one is eventually supported — must be loopback, not merely on your LAN (backend/app/media/providers.py, endpoint_rejection_reason). A picture of a scene carries the scene with it, and a GPU that renders your campaign is a machine you are sitting at.

Nothing to configure today: no media provider ships, the registry is empty, and there is deliberately no media endpoint setting to fill in. The rule exists so that whoever adds the first provider finds it already there.

Same host (the default)

endpoint_url = http://127.0.0.1:11434/v1

Nothing else to do. Ollama's own default is to listen on loopback.

An Ollama on another machine on your trusted LAN

Supported and explicitly configured — never guessed, never discovered.

On the inference machine, tell Ollama to accept connections from the LAN, because it binds loopback by default:

OLLAMA_HOST=0.0.0.0:11434 ollama serve

On the storyteller machine, set the endpoint to that host's address:

endpoint_url = http://192.168.1.50:11434/v1

Use an IP address or a name your own network resolves. Then:

  • the storyteller UI/API stays on 127.0.0.1 — do not change the listener;
  • prompts, story text, retrieved memories and embedding inputs all travel to that host, so it has to be one you control, on a network you trust;
  • the inference machine needs the models installed, not the storyteller;
  • no Internet is involved in either direction.

A LAN endpoint is accepted because it is on one of the allowed networks above. Nothing else about it is special.

If that endpoint is HTTPS with your own CA

Some inference hosts are only reachable over TLS. A StartOS server is one: it serves Ollama over HTTPS with a certificate from its own local CA, and plain HTTP redirects to it.

Install that CA on the machine running the storyteller, the same way you would for the browser — on Debian and Ubuntu:

sudo cp your-ca.crt /usr/local/share/ca-certificates/
sudo update-ca-certificates

then use the https:// URL and the hostname the certificate is issued for:

endpoint_url = https://inference.lan:8443/v1

The application verifies against the machine's CA store and the certifi bundle (backend/app/tlstrust.py), so a CA you installed at the OS level is honoured, exactly as curl and your browser honour it. Public certificates keep working unchanged.

There is deliberately no setting to skip verification. If a connection is refused with CERTIFICATE_VERIFY_FAILED, the CA is not installed where the storyteller can see it, or the URL's hostname does not match the certificate — openssl s_client -connect host:port will say which. In a container, remember the CA has to be inside the image or bind-mounted; the host's store is not visible from within.

Tests

cd backend  && .venv/bin/python -m pytest tests/ -q   # the backend suite
cd frontend && npm test                               # the component suite (M8)
cd frontend && npm run lint && npm run build

Fourteen backend tests skip without something the machine may not have: seven need a second machine or an environment the suite cannot create, and the rest are the real-model tests below.

The suite takes about fifteen minutes. Several files spawn genuine server processes — a restart is only evidence if the process really went away — and those dominate the wall clock.

The frontend component suite

M8 added one, because until M8 there was none — the browser was covered by real Firefox runs at each milestone's closeout and by nothing in between. It is Vitest and Testing Library over jsdom, and it runs in about two seconds:

cd frontend && npm test          # once
cd frontend && npm run test:watch # while working

It covers the deterministic browser behaviour M8 owns: which history controls are enabled and why, the take selector, the Save Point and delete confirmations, what the State panel shows and does not, knowledge classification and semantic status, the context inspector's sections, the model-empty and model-unavailable states, how failures are presented, the dialog focus trap, and that the reserved dictation control never touches the microphone. Several tests assert the absence of branch vocabulary in the surfaces a reader uses.

markdown.test.jsx is the security one. Narrator prose and imported text both reach the renderer, so it is where H06 and H07 are decided: markup in the source never becomes markup in the page, a javascript: URL never becomes an href, and a remote image is a placeholder rather than a request.

It does not replace the real-browser runs. jsdom has no layout, no navigation and no network, so scroll behaviour, streaming, a genuine process restart and the CSP are all outside its reach. Each milestone's closeout drives a real Firefox over WebDriver, and that evidence is recorded in the milestone report.

Two files are the M1 regression guards.

test_offline_assets.py fails if the tokenizer starts fetching its table again, if a remote font or stylesheet comes back, or if the CSP names a remote origin. Two of its checks read the built SPA under frontend/dist/ and skip when it has not been built, so run npm run build before treating a green suite as complete evidence.

test_tls_trust.py fails if outbound verification is weakened, if a public CA is lost from the union, or if a new HTTP client is added without the shared verification context.

M5 added test_narrative_state.py, which fails if the state stops being genre-neutral, if an event outside the allowlist is ever applied, if a malformed proposal mutates anything, if campaign canon stops outranking the narration, or if a turn's narration and its state can be committed apart from each other.

test_narrative_realistic.py is the one suite that needs a real model, and it is skipped unless you point it at one:

AIDND_TEST_ENDPOINT=http://127.0.0.1:11434/v1 \
AIDND_TEST_MODEL=qwen2.5:3b-instruct \
backend/.venv/bin/python -m pytest backend/tests/test_narrative_realistic.py -v -s

It exists because Phase 0B found that structured-state behaviour can look correct on a small prompt and fail under a full one — and it has already earned its place, catching a case where a model echoed its own instruction into the narration.

M7 added five files. test_imported_knowledge.py is the acceptance contract — G01-G10, C05, F05/F06's imported halves, I05, H06-H09, campaign isolation, lexical retrieval without embeddings, a bounded knowledge budget, deletion that preserves historical prompt evidence, hidden Canon, stale Canon against current state, and an abandoned line of story failing to influence the retrieval query. test_knowledge_chunking.py fails if chunking stops being deterministic or starts producing fragments or giants. test_knowledge_retrieval_quality.py fails if class stops settling ties, if irrelevant Canon starts winning on class alone, if the hybrid merge duplicates a passage, or if suppression crosses a class. test_knowledge_performance.py fails if any knowledge read grows a query per source or per passage, or if candidates stop being bounded in SQL. test_knowledge_migration.py fails if a pre-M7 database stops opening, or if the FTS5 index stops travelling with the table it indexes.

test_knowledge_real_model.py is M7's real-provider test and skips without an endpoint. It mocks nothing between itself and Ollama: a real Settings row, the real factory, a real embedding request, real stored vectors, real hybrid retrieval, and a real prompt.

AIDND_TEST_ENDPOINT=https://inference.lan:8443/v1 \
AIDND_TEST_EMBED_MODEL=nomic-embed-text \
backend/.venv/bin/python -m pytest backend/tests/test_knowledge_real_model.py -v -s

M4 added test_save_points.py, which fails if restoring a Save Point starts deleting history, stops going through the active head, forks on its own, lets a Save Point on one campaign be restored through another, or lets deleting a branch take a Save Point with it. It also fails if listing Save Points goes back to one query per Save Point, or starts fetching narration to render the list.

test_process_restart.py is the durability guard: it starts the application as a real subprocess, kills it, and starts a second one against the same database. A Save Point that survived only because a Python object was still alive would pass an in-process test and fail a user's restart.

M2 added two more. test_endpoint_policy.py fails if the set of reachable addresses widens, or if either place the rule is applied stops applying it — it resolves hostnames through a stub, so it tests the policy rather than whatever DNS the machine has. test_local_only_surface.py fails if a removed subsystem comes back as a route, if an API key becomes settable again, if the model timeout stops being configurable or becomes unbounded, or if a supported start path stops binding loopback.

Why a long campaign is not slow in proportion to its length

An inference server caches the prompt it has already processed, keyed on the prefix. While a story only grows at the end, each turn re-uses that cache and pays for its own new tokens alone. Once the context budget is full, though, the history window has to give something up — and a window that gives up its oldest action every turn changes the prompt near the front, which throws the cache away and makes the server re-read almost the whole thing, every turn.

So the window moves in blocks. context/builder.py snaps the oldest included action to a boundary and holds it there for several turns, then steps. Measured against the reference deployment on real builder output, at an 8,192-token budget:

Per turn
Window held, story grew by one action 14-20 s
Window stepped (one turn in three) 333-338 s
Mean over whole cycles 124.0 s
Window sliding every turn, as before 362.4 s

The cost is history depth: right after a step the window holds up to a block fewer actions than the budget would allow. TRIM_FRACTION bounds that at a quarter of the window, and it is the one number to change if you would rather trade recent history for speed, or the reverse.

The saving grows with the block, and the block grows with the budget — so the larger the context window, the more this is worth. history["floor_depth"] and history["trim_block"] are in every context report, and a floor_depth that is the same on two consecutive turns is the prompt's prefix having been preserved.

The release-validation harnesses

M11 added six runnable harnesses under backend/tools/. They are the evidence behind planning/reports/M11-IMPLEMENTATION-REPORT.md, and they live in the repository so a reviewer can re-run them rather than take the report's word for anything. None is part of the application and none is imported by it.

Write their output somewhere durable, never /tmp. --out is required on every harness precisely so the location is a decision rather than a default, and the examples below use $HOME/m11-evidence. A reboot clears /tmp, and a long-run campaign is hours of evidence that cannot be reproduced by re-reading a file. One run was lost exactly that way; §G.6 of the M11 report records it. Snap Firefox independently refuses a WebDriver file path under /tmp and needs one under $HOME, so $HOME is the only location the browser harness works from in any case.

cd backend
mkdir -p "$HOME/m11-evidence"

# The 100-turn release campaign (M01-M04): real narrator, genuine process
# restarts, every history operation. Hours, not minutes.
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 \
AIDND_TEST_MODEL=<model> AIDND_TEST_EMBED_MODEL=<embedding model> \
  .venv/bin/python -m tools.m11_long_run --turns 100 --out "$HOME/m11-evidence/m01"

# The same campaign, carried on after a crash, a reboot or a Ctrl-C. It picks up
# the adventure the checkpoint names, keeps its place in the beat cycle, and does
# not fire a scheduled operation that already fired.
AIDND_TEST_ENDPOINT=... AIDND_TEST_MODEL=... AIDND_TEST_EMBED_MODEL=... \
  .venv/bin/python -m tools.m11_long_run --turns 100 --resume --out "$HOME/m11-evidence/m01"

# What that campaign is worth on a machine that has never seen it (I01-I07).
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"

# The browser release regression and the accessibility measurements: M11's 38
# checks plus v1.1 WP-C's reader workflows (Retry, Save Points, state correction,
# narration length, failed generation, export download). Needs `frontend/dist`
# built, geckodriver on PATH, and --out under $HOME (the downloads land inside
# it). Release evidence needs the narrator over trusted-LAN HTTPS.
AIDND_TEST_ENDPOINT=https://... AIDND_TEST_MODEL=qwen2.5:3b-instruct \
  .venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/browser"
# Without a narrator (a partial smoke run, not evidence), or one scenario while
# developing (--only takes: shell, history, markdown, hidden, context, csp, a11y,
# retry, savepoint, state, length, failure, export).
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/smoke" --no-narrator

# A container with no network at all: the offline run and the packaging path.
.venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline"

# The multi-character identity diagnostic (post-M8 finding D), and the run that
# proves its detectors fire.
.venv/bin/python -m tools.m11_identity --out "$HOME/m11-evidence/identity"
.venv/bin/python -m tools.m11_identity --scripted --inject

# The palette, against WCAG AA.
.venv/bin/python -m tools.contrast_audit

Resuming the long run, and timing it out

A hundred turns is hours of wall clock, and the first release attempt lost one at turn 97 to a host crash. The harness now checkpoints resume.json into --out after the prologue, after every scheduled operation and after every turn, and --resume continues from it. The file is written under a temporary name and renamed, so a crash during the write cannot leave a half-parsed one.

resume.json is operational state rather than evidence: timeline.jsonl stays the append-only record, a resumed session appends to it, and a finished run deletes its resume.json. That makes the file's presence mean exactly one thing — there is an unfinished run in this directory — and the harness refuses to start a fresh campaign on top of one, because two campaigns interleaved in a single timeline and database are worse evidence than none. It refuses a directory holding a campaign.db with no checkpoint for the same reason.

How long a turn takes is the inference host's characteristic, not the application's, so the timeout is an option rather than a constant:

Flag Default What it does
--turn-timeout 1800 Seconds the application waits for one narrator reply — it becomes model_timeout_seconds, so the settings schema's 30..3600 bound applies. The harness waits 300s longer, so the application's own error arrives inside the stream rather than being cut off at the socket.
--max-consecutive-failures 5 Unaccepted turns in a row before the run stops, writes summary.json with status: aborted, and leaves a resume.json that --resume can carry on.

Measure your host before lowering --turn-timeout. On the M11 reference deployment a turn cost 229-291 seconds at the recommended window; a slower host can exceed the 600 seconds this harness used to hard-code, and an overrun turn is a lost turn.

Logging the inference host during a long run

A long run is the heaviest sustained load an inference host sees. In M11 a GPU host dropped its GPU off the PCIe bus (NVRM: Xid 79) half a minute after a 100-turn run finished. Nothing on disk could say whether power, heat or the link caused it (M11 report, §E.1). For every long run against a GPU host, start this logging on that host first and stop it only when the run has finished.

Run each command in its own terminal on the inference host. tee writes each line as it arrives, so what happened in the seconds before a crash or a forced reboot survives on disk. Every log goes into one directory, so a campaign's evidence stays together and is easy to archive or remove afterwards:

mkdir -p "$HOME/inference-host-logs"

# Power, temperature, utilisation and PCIe link state, once a second
nvidia-smi --query-gpu=timestamp,pcie.link.gen.current,pcie.link.width.current,power.draw,temperature.gpu,utilization.gpu \
  --format=csv -l 1 | tee "$HOME/inference-host-logs/gpu-link-$(date +%F-%H%M).csv"

# The driver's own sampling: power, utilisation, clocks, memory, ECC and throttling
nvidia-smi dmon -s pucvmet -d 5 | tee "$HOME/inference-host-logs/gpu-dmon-$(date +%F-%H%M).log"

# Kernel and Ollama messages, live. The `+` is an OR: `journalctl -k -u ollama`
# asks for messages that are both kernel messages and the ollama unit's, which
# is none, and writes an empty log.
journalctl -f -o short-iso _TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service \
  | tee "$HOME/inference-host-logs/ollama-kernel-watch-$(date +%F-%H%M).log"

If the GPU faults, find the moment and then read what the card was doing just before it:

grep -iE 'xid|fallen off|nvrm' "$HOME"/inference-host-logs/ollama-kernel-watch-*.log
awk -F', ' 'NR>1 && $4+0 > max {max=$4+0; at=$1} END {print "peak W", max, "at", at}' "$HOME"/inference-host-logs/gpu-link-*.csv

A fault that follows sustained draw at the card's power limit points to power delivery. A fault with the link below its usual generation under load points to the connection. A fault with neither is still worth recording, because it rules both out. These commands were verified against NVIDIA driver 580 and Ollama 0.34.

tools/m11_webdriver.py is the W3C WebDriver client the browser harness uses. It exists so browser evidence needs no Selenium in the dependency surface, and it documents the one environment quirk that matters here: a snap Firefox will not open a file the driver names under /tmp, but will under $HOME.

Downloads in the browser harness (v1.1 WP-C). The export checks click the real Export controls and wait for the file on disk, so the browser has to save without asking. m11_webdriver.firefox_download_prefs gives the WebDriver session a profile that does that:

  • browser.download.folderList 2, browser.download.dir the run's downloads/ folder, browser.download.useDownloadDir true;
  • no "always ask", and application/json saved to disk.

It works on the snap Firefox this machine has (155.0.1, geckodriver 0.37.1), and no separate Firefox is needed. The same sandbox rule applies as for opening files: the download folder must be under $HOME, and the harness refuses one that is not.

A download counts as finished only when all of these hold at once (m11_webdriver.wait_for_download):

  • a new name has appeared;
  • no *.part file is left;
  • the file is more than zero bytes;
  • its size is the same across consecutive polls.

The toast that says "Campaign exported." is not evidence.

Waiting. Nothing in the harness sleeps before an assertion. Every wait is on something the page, the browser or the filesystem shows. A condition that already holds before the action it waits for does not count as waiting for that action; the harness defects found in M8, M11 and WP-C were all of that shape.

Backing up, and getting a campaign back

There are two recovery tools and they answer different questions. Using the wrong one is the most common way to be surprised later, so they are described together.

Campaign export Database backup
Covers one campaign every campaign, and your settings
Shape a JSON file you can read a copy of the SQLite database
Moves between machines yes — this is the supported way no; it is this machine's database
Taken from Export, on a campaign Settings → Back up everything on this machine
Restored by Import campaign, on the library screen replacing the database file, below

How large an export can get

The importer accepts a request body up to 20 MB (backend/app/limits.py, MAX_IMPORT_BODY_BYTES), and v1.1 does not change it. What that means for a campaign, measured rather than guessed:

  • the M11 evidence campaign came to roughly 13 kB per action in its bundle;
  • M9's conservative estimate from that figure is about 279 turns before a bundle approaches the limit.

Both are measurements of particular campaigns, not a turn limit. What a campaign actually weighs depends on how long its turns are, how much imported knowledge travels with it, and how many attempts each turn kept. A campaign of 400 short turns can be well inside the limit; one of 200 long ones with a large library may not be.

v1.1 (WP-D) makes the individual case visible. Every export reports its own serialised size and whether this version could import it back:

X-Export-Bytes                  the bundle's size, as the importer would weigh it
X-Import-Limit-Bytes            MAX_IMPORT_BODY_BYTES
X-Importable-By-This-Version    true / false
X-Export-Warning                present only when it is false

The export always succeeds and the file is always delivered — it is complete and undamaged; what it exceeds is this version's import ceiling. Both Export controls show the warning when there is one. The size compared is the compact serialisation the browser would POST back, which is smaller than the pretty-printed file on disk.

Raising the limit, or streaming an import past it, is deferred to v1.2.

Exporting and importing a campaign

Export is on each campaign in the library, and in the campaign's own Settings panel. It writes one .json file holding the whole campaign: the story and its entire retained tree, the branch you are on and the exact position you are reading at — including one you undid back to — every alternate take, your Save Points, the authoritative state and its per-position snapshots, the state history that explains it, your imported knowledge with its classifications, the summaries and memories, and the prompt each turn was actually given.

Import is on the library screen and takes that file back, into this or any other installation. Nothing about the file refers to the machine that wrote it: the imported files come back from their content, not from a path, and no setting of yours is changed by importing somebody's campaign.

Two things it deliberately does not carry: your inference endpoint and model settings, which describe your machine rather than the campaign, and the rebuildable search indexes, which are rebuilt from the imported content before the import returns.

A campaign imports whether or not the model that wrote it is installed here. Recovering a campaign and being able to play it on are separate questions; the first never depends on the second.

Backing up the whole database

Settings → Advanced → Back up everything on this machine. It writes a verified copy into a backups/ directory beside the database itself, and tells you where.

It is a real backup rather than a file copy. It uses SQLite's online backup API, so it is safe to take while you are playing — a cp of a live database can read one page before a transaction and another after it, producing a file that opens, reports a schema, and is quietly missing rows. The copy is checked with PRAGMA quick_check before it is kept, an existing backup is never overwritten, and a failure leaves nothing behind.

You can also take one from the command line, or from cron:

curl -s -X POST http://127.0.0.1:8000/api/backups | python3 -m json.tool

Restoring a whole database

There is deliberately no restore button, because restoring means replacing the file the running application has open — which is how you lose both copies at once. It is a three-step procedure and each step needs the application stopped:

# 1. Stop the application. Nothing below is safe while it is running.
#    (Ctrl-C the server, or `docker compose down`.)

# 2. Keep what is there now, whatever state it is in. You may want it back.
mv backend/data.db backend/data.db.before-restore

# 3. Put the backup in its place, and start the application again.
cp backend/backups/adventure-storyteller-20260907-043000.db backend/data.db

Check the file before you trust it, and check it again after starting:

sqlite3 backend/backups/adventure-storyteller-20260907-043000.db 'PRAGMA quick_check;'
# -> ok

The database path is backend/data.db by default, and whatever AIDND_DB_PATH names otherwise — in Docker that is the mounted volume.

There is one file to move and no others: this build leaves SQLite in its default rollback-journal mode, so there are no -wal or -shm companions beside the database (PRAGMA journal_mode reports delete). A build that switched to WAL would have to move those too, and leaving them behind would pair a new database with an old write-ahead log.

Prefer the campaign export for anything smaller than "everything". Restoring a whole database rolls every campaign back to the moment the backup was taken, including the ones you did not mean to touch. To recover one campaign, export it and import it.

The context window your Ollama actually enforces

Check this before a long campaign. The application budgets a prompt up to Settings.context_token_budget (16,384 by default). Ollama enforces its own input window, and when it sees no VRAM it defaults to 4,096:

level=INFO msg="vram-based default context" total_vram="0 B" default_num_ctx=4096

Confirm what yours is:

curl -s http://127.0.0.1:11434/api/ps | python3 -m json.tool | grep context_length

If that number is smaller than your budget, Ollama silently truncates the input — and llama.cpp drops the oldest tokens, which in this application is the system block: the narrator rules and the campaign canon. The symptom is a narrator that forgets canon deep into a long session, with nothing on screen explaining why.

The application now checks, and will not over-budget. Since M11 it asks the server what window your model actually gets — /api/ps for a model that is loaded, /api/show for one that is not — and caps the prompt to that number. A 4,096-token server therefore no longer receives a 16,384-token prompt: the campaign gets less history than the setting asks for, which is a visible, explicable loss rather than a silent one, and Settings' Test connection reports the window it found or says plainly that it could not check.

It also keeps a margin, and checks the server's own count (v1.1). The application counts tokens with cl100k_base, and your model counts them with its own tokenizer. The two disagree slightly, so the prompt is built to leave max(256, 5% of the window) tokens free on top of the reply: 256 at 4,096, and 820 at 16,384. After each turn the server's reported prompt-token count is compared with what was sent. The context inspector shows the result for any past turn:

  • The server read the whole prompt: the ordinary case.
  • The server did not say how much it read: the server reported no usage. Nothing is wrong, and nothing is confirmed either.
  • The server may have cut the start of the prompt: it read far fewer tokens than were sent. Ollama does this, silently, to a prompt larger than the window the model was loaded with. The turn is kept. Check the window with the commands above.
  • The prompt was larger than the server allowed for: its count and the reply together exceed the window. The reply may have been cut short. The turn is kept.

The last two also appear in the server log as a warning.

A model that is not loaded yet is loaded first. Before a turn, if the application cannot read the window because your model isn't in memory, it asks the same Ollama to load it once. That is a POST /api/generate naming only the model, which generates no text. It then reads the window again, so the first turn of a session is built to the window the model really has rather than to your setting. If loading fails, or the window still can't be read, the turn goes ahead exactly as before, unverified, and the check above still applies.

That does not make the window bigger, and the rest of this section is still how you do that.

On a server that is not Ollama, tell the application the window yourself. The check above uses Ollama's native API, which vLLM, llama.cpp's own server and the rest do not serve — so the window comes back unverified and the budget is left at whatever is configured. Set context_window_override in settings to the window you launched that server with:

curl -X PUT http://127.0.0.1:8000/api/settings \
  -H 'Content-Type: application/json' -d '{"context_window_override": 8192}'

Prompts are then capped to it. It is used only when the server could not be asked — a window the server did report always wins, so this can never be a way to over-budget an Ollama that answered — and it does not count as verification: the turn's provenance still records that nothing checked the number. Send null to remove it. Nothing here validates the figure against the server, so an override larger than the real window puts you back to silent truncation; take it from how you started the server, not from the model card.

Setting it per request does not work from this application. Ollama's OpenAI-compatible endpoint accepts num_ctx — nested in options or at the top level — returns HTTP 200 and ignores it. Worse, it reloads the model at its own default, so priming the server with a native /api/chat call first does not help either: the app's next request resets the window.

Bake it into a model instead. The window travels with the model, and this needs no shell access on the Ollama host — it is a normal API call:

curl http://127.0.0.1:11434/api/create -d '{
  "model": "qwen2.5:3b-instruct-16k",
  "from":  "qwen2.5:3b-instruct",
  "parameters": {"num_ctx": 16384}
}'

The derived model shares the base model's blobs, so it costs a manifest. It then appears in /v1/models, which is the listing the Settings model picker reads — select it there and the storyteller gets the full window through its ordinary OpenAI-compatible path. Remove it with POST /api/delete when you are done.

Where you do control the server environment, OLLAMA_CONTEXT_LENGTH=16384 does the same job. Either way a larger window costs roughly proportionally more KV cache.

If you would rather not raise it at all, you no longer need to do anything: the application caps itself to what the server reports. Setting How much story to send to the same number simply makes the intent explicit.

This matters most on the machine you import to. A campaign carries its history, not the window the machine that wrote it had, and a long imported campaign fills a prompt on its very first turn — so a deployment that has applied neither the derived model above nor a matching budget meets its ceiling immediately rather than gradually. Importing succeeds either way, and since M11 the first turn afterwards is capped rather than truncated — so what a small window costs is history, not the canon at the front of the prompt. It is still worth giving the model its window before playing an imported campaign: a 4,096-token context on a hundred-turn story is a much shorter memory than the story was written with.

What was made offline-safe, and how to check

Two runtime downloads were removed in Milestone M1. Both were invisible on a machine that had been online once, which is exactly why they need tests.

The tokenizer. tiktoken.get_encoding("cl100k_base") downloads a 1.7 MB BPE table on first use, and the context builder counts tokens on every turn, so the first story turn on an air-gapped install died with a ConnectionError. The table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it, verifying its SHA-256 against the digest tiktoken itself pins.

The fonts. The SPA linked fonts.googleapis.com from index.html, so every page load fetched a stylesheet and font files from Google. The three families are self-hosted under frontend/public/fonts/, declared in frontend/src/styles/fonts.css, and re-vendored by python3 frontend/tools/vendor_fonts.py. The CSP in backend/app/main.py now names no remote origin at all.

To convince yourself on a machine that has already been online, run the app with no route out rather than trusting a cold cache:

docker network create --internal offline
docker run -d --name ollama --network offline -v ollama-models:/root/.ollama ollama/ollama
docker build -t storyteller .
# The app shares Ollama's network namespace, so Ollama is on its loopback and
# neither has a route to the Internet.
docker run -d --name app --network container:ollama -v story-data:/data \
  storyteller uvicorn app.main:app --host 127.0.0.1 --port 8000
docker exec app python -c "import socket; socket.create_connection(('1.1.1.1',443),timeout=4)"
# -> OSError: Network is unreachable, and story turns still work

planning/archive/milestone-reports/M1-BASELINE-REPORT.md records the run this procedure is taken from, including the packet captures.

Things still inherited from upstream

M2 removed the hosted, cloud, account, analytics, Postgres/Render and QuickJS scripting surfaces outright — PROVENANCE.md lists exactly what went. What is left of upstream that a newcomer might report as a defect:

  • Inert legacy tables and columns. Five tables and four columns M2 emptied of meaning are still in the schema, unmapped, so an M1-era campaign database opens unchanged. Nothing reads or writes them. A cleanup migration waits for the schema to settle after M5 (planning/BUILD-MILESTONES.md).
  • Dual-dialect migration code. backend/app/migrations.py still carries SQLite/Postgres branches from upstream, although Postgres support itself is gone and SQLite is the only store. Same cleanup, same milestone.
  • .github/workflows/ci.yml is upstream's GitHub Actions pipeline. This repository lives on a self-hosted Gitea; the workflow is kept for provenance and is not what runs the tests here.
  • A thin components.jsx. What is left of upstream's shared component module is a toast host, a file picker, a JSON download and an auto-growing textarea. M8 removed the rest with the screens that used them — the scenario art generator, the placeholder modal, the story-card row.

(Removed from this list by M8: no frontend tests. There is a component suite now — see Tests above.)