Files
JesseMarkowitzandClaude Opus 5 d27ee34901 Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was
authoritative. Phase 0 execution prompts sat beside the specification; four
completed milestone reports sat beside the current one; and upstream AI-DnD's
own `plan/` build log and `docs/` project site still described a hosted,
scripted, multi-user product with accounts — every screenshot in it showed a
Scripts tab and a Sign up button, none of which has existed since M2.

`planning/archive/` now holds the history and says so in its own README:
`phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and
M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied.
`planning/reports/` holds only the current milestone's report, because that is
the one M4 planning has to read; it moves to the archive when M4's replaces it.

Deleted rather than archived: the Phase 0B execution prompts and the
handoff/status/summary documents, the Phase 0A discovery and triage reports,
upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's
template boilerplate. All of it is in Git history, and the two recommendation
reports carry every conclusion the deleted research reached.

Archived documents are kept verbatim. Paths written inside them point at where
those files were when the document was written, which is the point: an evidence
record that has been quietly edited is no longer evidence.

Active documentation is corrected where it pointed at the removed trees or
described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not
touch" list had gone stale at M2 and claimed QuickJS scripting was still tested;
its test count was 604 against an actual 638. `README.md` loses the upstream CI
badge, which reported upstream's pipeline rather than this fork's, and a
reference to `backend/app/worldstate/engine.py`, a file that does not exist.
`planning/README.md` is rewritten as the documentation index.

New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
manifest of what belongs in the ChatGPT project's Sources.

Source comments referring to the deleted trees are reworded; no behaviour
changes. 638 backend tests pass, frontend lints and builds, and a reference scan
over all 48 tracked Markdown files reports no unresolved path in active
documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
2026-09-03 14:33:07 -04:00

34 KiB
Raw Permalink Blame History

M2 — Baseline Report (evidence record)

Date: 2026-09-02 Milestone: M2, Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface Companion: planning/reports/M2-IMPLEMENTATION-REPORT.md interprets this file. Where the two disagree on a runtime or test fact, this file is the record.

This is measurement, not commentary. Command output is quoted verbatim. Anything inferred rather than observed is labelled [inferred].

Hostnames and LAN addresses are placeholders — inference.lan, 192.168.0.50, port 8443. The real ones are in the workspace's untracked notes, never in this repository. Packet counts, digests, timings and status codes are exact.


0. Test environment

Host Ubuntu 24.04.4 LTS, x86-64, 4 cores, 15 GB RAM, no GPU
Python 3.12.3
Node / npm 22.23.1 / 10.9.8
Docker 29.7.2
Ollama (same-host) ollama/ollama:latest in a container, models from a shared volume
Ollama (LAN) 0.33.0 on a second physical machine, HTTPS, certificate from a private CA
Models qwen2.5:0.5b, qwen2.5:3b-instruct, nomic-embed-text
Image under test m2-final, built from the working tree by docker build .

Timings in this file are not benchmarks. This host was shared with unrelated work during the review; uptime reported a 1-minute load average ranging from 1.79 to 74.08 across the session on four cores. Where a measurement was taken under load, it says so. Latency figures are recorded because they mattered to a timeout result, not as performance data.


1. Repository and provenance

$ git rev-parse --abbrev-ref HEAD
m2-local-only-surface

$ git log --format='%H %G? %s' -3
8c65ae99deda49b22f415d5868987e00bbcb173c G M2: cut the hosted product away from the local one
1a28a9a708985e7c98dcbcb87c189e7582ec288d G Apply post-M1 corrections to the planning package
645f07f06d226e274c33d96e71d9c374413ef771 G Add the M1 implementation review report

All three commits verify (%G? = G). The repository signs every commit; nothing in this milestone is unsigned.

Commit chain

Commit Sig Parent Meaning
8c65ae9 G 1a28a9a The M2 commit. One commit, no merges.
1a28a9a G 645f07f Post-M1 planning corrections — the M1 baseline this milestone started from
7f182a8 G 717670a, d72f7c1 the fork import, two parents

Upstream ancestry

$ git cat-file -t d72f7c1bda0f34fccd84afb7a25c34eb01c901de
commit
$ git merge-base --is-ancestor d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD && echo yes
yes
$ git rev-list --count d72f7c1bda0f34fccd84afb7a25c34eb01c901de
172
$ git diff --stat d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD -- LICENSE
(no output)

All 172 upstream commits reachable; LICENSE byte-identical to upstream; no re-import or upstream substitution occurred. PROVENANCE.md and DEVELOPMENT.md are present and were updated for M2.

Private data

$ grep -rn "<the real LAN hostname>\|<the real LAN subnet>\|<the real port>" \
    --include=*.py --include=*.md --include=*.jsx --include=*.js \
    --include=*.yml --include=*.txt backend frontend *.md *.yml Dockerfile
(no output)

No private hostname, LAN address, certificate or credential is committed.

Working tree

At the M2 commit the tree was clean. This review then modified six files (§9), so git status is not clean as this report is written:

 M backend/app/memorybank.py
 M backend/app/models.py
 M backend/app/routers/adventures/turns.py
 M backend/app/routers/chat.py
 M backend/requirements.lock
 M backend/tests/test_local_only_surface.py

Every runtime result below was produced by an image built from that modified tree, except where explicitly marked as pre-fix.


2. Change inventory

$ git diff --shortstat 1a28a9a 8c65ae9
 94 files changed, 1395 insertions(+), 6578 deletions(-)

Added (3):

backend/app/endpoints.py
backend/tests/test_endpoint_policy.py
backend/tests/test_local_only_surface.py

Deleted (24):

backend/app/accesslog.py                     backend/tests/test_accesslog.py
backend/app/analytics.py                     backend/tests/test_analytics.py
backend/app/cleanup.py                       backend/tests/test_guest_cleanup.py
backend/app/netguard.py                      backend/tests/test_netguard.py
backend/app/security.py                      backend/tests/test_ratelimit_hardening.py
backend/app/routers/analytics.py             backend/tests/test_reasoning_param.py
backend/app/routers/auth.py
backend/app/routers/scripts.py               frontend/src/pages/Analytics.jsx
backend/app/routers/adventures/scripts.py    frontend/src/pages/Scripts.jsx
backend/app/scripting/__init__.py            frontend/src/pages/ScriptEditor.jsx
backend/app/scripting/engine.py              frontend/src/pages/Play/panels/ScriptsPanel.jsx
backend/app/scripting/pipeline.py            frontend/src/pages/Play/drawers/StatusDrawer.jsx
                                             render.yaml

Largest modifications:

  +93  -37   backend/app/routers/settings.py
  +54  -82   backend/app/providers/openai_compatible.py
  +44 -166   backend/tests/test_chat.py
  +38  -81   backend/app/routers/chat.py
  +37 -222   backend/app/auth.py
  +32  -69   backend/app/main.py
  +31 -195   backend/app/limits.py
  +28 -178   backend/app/models.py
  +22  -43   backend/app/database.py
  +19  -58   frontend/src/pages/Settings.jsx
  +18  -97   backend/app/routers/adventures/turns.py
  +15  -85   frontend/src/App.jsx

3. Surface reduction, measured

API routes (from the OpenAPI schema, not by grep)

Prefix M1 M2
/api/adventures 25 21
/api/analytics 3 0
/api/auth 4 0
/api/scripts 5 0
/api/chat 2 2
/api/scenarios 5 5
/api/settings 2 2
/api/story-cards 4 4
/api/debug, /api/health 2 2
Total 52 36

Dependencies

$ diff <(M1 requirements.txt) <(M2 requirements.txt) | grep '^<'
< quickjs>=1.19
< cryptography>=42
< psycopg[binary]>=3.2

Installed closure, measured by building a clean venv from requirements.txt + requirements-dev.txt:

M1 M2
Python packages installed 40 34
Packages gone — cffi, cryptography, psycopg, psycopg-binary, pycparser, quickjs
npm runtime dependencies 6 3
npm packages installed (npm ls --all) 53 32

Code size

M1 M2 Δ
backend/app Python lines 14 298 11 632 −2 666
frontend/src JS/JSX lines 6 523 5 239 −1 284
frontend/src/pages files 21 16 −5

Environment variables actually read (os.environ)

$ grep -rn "os.environ" backend/app/
backend/app/database.py:18:_env_db_path = os.environ.get("AIDND_DB_PATH")
backend/app/main.py:30:    for o in os.environ.get("AIDND_CORS_ORIGINS", "").split(",")

Two, down from ten. AIDND_MULTI_USER, AIDND_SECRET_KEY, AIDND_COOKIE_SECURE, AIDND_DEMO_API_KEY, AIDND_DEMO_ENDPOINT_URL, AIDND_DEMO_MODELS, AIDND_DEMO_TURNS_PER_DAY, AIDND_POWER_USERS, AIDND_ANALYTICS_EMAILS, AIDND_TRUSTED_PROXY_HOPS, AIDND_DATABASE_URL and DATABASE_URL are no longer read. Four of those names still appear in the tree as prose in comments explaining what was removed; grep above shows no read.


4. Test suite

$ cd backend && .venv/bin/python -m pytest tests/ -q
606 passed, 1 warning in 126.29s (0:02:06)

Zero failed, zero skipped, zero xfailed. The one warning is the inherited Starlette/httpx deprecation notice, present since M1.

Accounting for every test

Counted by diffing collected node IDs between 1a28a9a (M1) and the working tree, not by reading diffs:

Count
Distinct test functions, M1 610
Distinct test functions, M2 561
Node IDs gone 88
Node IDs new 39
Collected tests (with parametrisation), M1 648
Collected tests, M2 606

Gone, by file:

File Gone Disposition
test_analytics.py 25 whole file — subject removed
test_guest_cleanup.py 14 whole file — subject removed
test_accesslog.py 12 whole file — subject removed
test_ratelimit_hardening.py 8 whole file — subject removed
test_chat.py 8 demo-key pinning and the power-user gate
test_reasoning_param.py 5 whole file — subject removed
test_netguard.py 5 whole file — replaced by test_endpoint_policy.py
test_prompt_caching.py 4 OpenRouter upstream routing
test_memory_rewrite.py 4 2 retired (Postgres DSN masking), 2 renamed
test_delete_state.py 1 renamed
test_branch_forking.py 1 renamed
test_branch_clause.py 1 scripting history API

Four of the 88 are renames with equivalent coverage, so 84 tests were genuinely retired:

test_branch_forking:  test_switching_restores_the_script_and_world_state
                   -> test_switching_restores_the_state_a_branch_left_behind
test_delete_state:    test_deleting_the_ai_turn_rewinds_the_script_state
                   -> test_deleting_the_ai_turn_rewinds_the_counter
test_memory_rewrite:  test_an_owner_with_no_api_key_is_skipped
                   -> test_an_adventure_with_no_model_configured_is_skipped
test_memory_rewrite:  test_an_api_key_on_the_command_line_covers_that_owner
                   -> test_a_model_on_the_command_line_covers_that_adventure

New, by file: test_local_only_surface.py 18, test_endpoint_policy.py 14, test_chat.py 3, plus the 4 renames. Collected with parametrisation: test_endpoint_policy.py 31, test_local_only_surface.py 32.

Instrumentation conversion, not deletion

Eight files used a QuickJS output hook (state.gold += 10) as deterministic instrumentation for the state snapshot and rollback machinery, which M2 does not touch. The counter moved to the world-state engine — the model emits a ```state delta block, the referee applies it — and the assertions are unchanged in substance. Affected: test_take_state, test_attempt_siblings, test_branch_forking, test_bundle_v2, test_delete_state, test_retry_variants, test_story_tree_baseline, test_turn_flow_integration, plus test_state_revert converted from script_state/state_after to world_state/world_state_after.

Frontend and image

$ npm run lint          # oxlint
exit 0                  # 7 warnings, all pre-existing react/only-export-components
$ npm run build
dist/assets/index-C9v6AJ3R.js   395.41 kB │ gzip: 120.01 kB     ✓ built in 496ms
$ docker build .
exit 0

There are no automated frontend tests in the repository — none existed at M1 either.

Bundle size: 933.69 kB → 395.41 kB (M1 → M2), from removing CodeMirror with the script editor.


5. Offline run — same-host Ollama

Topology. --internal Docker network (no NAT, no external DNS). Ollama in one container; the m2-final image in a second container sharing Ollama's network namespace, so Ollama sits on the storyteller's own loopback. tcpdump ran in that namespace for the whole session.

5.1 Isolation, verified before any test

  blocked 1.1.1.1:443  OSError
  blocked 140.82.121.4:443  OSError
  no resolution  fonts.googleapis.com
  no resolution  fonts.gstatic.com
  no resolution  openaipublic.blob.core.windows.net
  no resolution  openrouter.ai
  no resolution  api.openai.com
  no resolution  github.com

5.2 Listeners

  LISTEN 127.0.0.1:8000        <- the storyteller
  LISTEN 127.0.0.11:42229      <- Docker's embedded DNS
  LISTEN 127.0.0.1:43093       <- Ollama's model runner
  LISTEN 127.0.0.1:43313       <- Ollama's model runner
  LISTEN [::]:11434            <- Ollama itself

Nothing the storyteller owns is bound off loopback.

5.3 Ollama-only settings and diagnostics

endpoint: http://127.0.0.1:11434/v1 | model: qwen2.5:0.5b
embed: nomic-embed-text:latest | timeout: 600
connection test: {"ok": true, "models": ["qwen2.5:7b-instruct",
  "qwen2.5:3b-instruct", "nomic-embed-text:latest", "qwen2.5:0.5b"]}

5.4 Story generation and streaming

  turn 1 (  1.7s,  90 stream chunks): 'Oh, how unfortunate! A lighthouse keeper indeed! The light'
  turn 2 ( 35.2s,  90 stream chunks): 'I climb the spiral stair to the lamp room. My hand tightly'
  turn 3 (  7.1s,  90 stream chunks): "I search the keeper's log for the last entry. The last ent"
  turn 4: ERROR {'type': 'error', 'detail': 'The AI endpoint timed out.'}

Turn 4 timed out at the configured 600 s. The application logged no error. Ollama reported the model still resident. uptime at the time: 1-minute load average had risen from 1.79 to the tens on four cores from unrelated work on this host. [inferred] the timeout is host contention rather than an application fault; what is measured is that no application error was logged and that the same build produced turns in 1.7–35.2 s minutes earlier.

5.5 Local embeddings, end to end through the app

  wrote memory id=1, embedded=False
  turn to trigger the pass (101.6s)
  memories: 1  embedded with the local model: 1

Direct timing of the embedding endpoint from inside the container:

  embedding round trip: 3.4s, dim=768
  second embedding (warm): 0.1s

5.6 Branch-scoped memory isolation — positive and negative controls

main branch = 1
  memories on main: 0   ALPHA present: False
forking at AI action 22 (an earlier turn, so this starts a new line)
  fork events: ['chunk', 'chunk', 'done']
  branches now: 2, fork = 5

ON THE FORK (0 memories)
  ALPHA visible: False    <- ALPHA is anchored at main's head, which is BELOW the
                             fork point, so the fork correctly does not inherit it
  BETA written here, visible: True    <- positive control

BACK ON MAIN (1 memories)
  ALPHA visible: True    <- positive control
  BETA visible:  False   <- NEGATIVE CONTROL, must be False

The decisive result is the last line: a memory written on the fork is not visible on main.

5.7 A05 — a failed model call, and a hand-edited database

baseline: (8 actions, 3 accepted AI turns, digest 07615da99b014b70)

-- invalid local model --
  events: ['player', 'error']
  error: Endpoint or model not found (HTTP 404). Check the endpoint URL and that model 'n…
  after: (7, 3, 'e18de9dff6455e0d')

-- endpoint refused by policy at request time --
  (the settings row was edited directly with sqlite3, behind the app's back,
   to https://openrouter.ai/api/v1)
  events: ['player', 'error']
  error: This endpoint can't be used — openrouter.ai is a cloud inference service — this build…
  after: (8, 3, '157c882c60579617')

The accepted AI-turn count stayed at 3 through both failures. Only the player's own typed action was added each time, which is M1's documented and deliberate behaviour.

The second case is the strongest form of the endpoint evidence: the request-time check refused a cloud endpoint that had been written straight into SQLite, bypassing the API's save-time validation entirely.

5.8 Restart and resume

before restart: 8 actions, digest 157c882c60579617
  branches: 2  memories: 1
after restart:  8 actions, digest 157c882c60579617
  branches: 2  memories: 1
  settings preserved: endpoint http://127.0.0.1:11434/v1 timeout 600
  context inspection: ['narrator', 'persona', 'history', 'used_memories']

5.9 Packet capture — whole offline session

all packets:               5131
loopback (127.0.0.0/8):    5086
non-loopback unicast:      0
TCP connections opened outside loopback:  (none)

DNS queries, attributed by timestamp rather than assumed:

Name First seen Attribution
fonts.googleapis.com 16:44:02.800449 the isolation probe in §5.1
openaipublic.blob.core.windows.net 16:44:02.802749 the isolation probe
openrouter.ai 16:44:02.803260 the isolation probe
api.openai.com 16:44:02.803956 the isolation probe
github.com 16:44:02.804747 the isolation probe
ollama-container 18:28:55.684172 the endpoint-policy edge-case test (§7)
ollama.com 16:53:29 … 18:28:05 the Ollama server's own lookup
capture window:                       16:44:02.800136 .. 18:51:29.080336
first packet to the storyteller API:  16:44:47.639083

Every cloud/font/tokenizer lookup landed within 5 milliseconds of the capture starting and 45 seconds before the storyteller received its first request — they are the deliberate probe, not the application. ollama.com recurs throughout and is issued by the separately installed Ollama service, the same observation M1 recorded.


6. Offline run — trusted-LAN Ollama over HTTPS

Topology. Ollama 0.33.0 on a second physical machine on the trusted LAN, serving HTTPS with a certificate from a private CA. The storyteller ran in a container with NET_ADMIN, its default route deleted and replaced with a route to the LAN subnet only, its resolver pointed at nothing, and the host name supplied as a static hosts entry.

6.1 Isolation and binding

=== isolation ===
  blocked 1.1.1.1:443 OSError
  blocked 140.82.121.4:443 OSError
  no resolution github.com
  no resolution openrouter.ai
  no resolution fonts.gstatic.com
=== the approved LAN host ===
  inference.lan -> 192.168.0.50
  TLS verified, peer CN = inference.lan
=== listeners ===
  LISTEN 127.0.0.1:8000

Nothing listens on the container's own LAN-facing address.

6.2 Diagnostics and generation

endpoint: https://inference.lan:8443/v1 | timeout: 600 s
diagnostics: {"ok": true, "models": ["qwen2.5:3b-instruct", "nomic-embed-text:latest"]}
turn 1 ( 10.7s, 24 chunks): 'You check the oil and wick, their greasy slicks staining your hand'
turn 2 (  4.0s, 17 chunks): 'The air is thick with the musty scent of the old building as you a'
turn 3 (  4.6s, 26 chunks): 'You find the last entry: "Exhausted... Oil\'s low..." The room feel'

memories: 1  embedded via the LAN host: 1

6.3 Every model request the application made

    4  https://inference.lan:8443/v1/chat/completions
    1  https://inference.lan:8443/v1/embeddings

Narration and embeddings both went to the configured host over verified TLS. No other URL appears.

6.4 Retry keeps the discarded attempt

Measured on the shipped image, against the LAN host:

head AI action before retry: 8   take_count 1
  retry produced (13.7s): 'You jot down a hurried note: "Visited lamp—oil low—oil lamp '
head AI action after retry:  9   take_count 2   take_index 1
alternate takes retained: 2
   - 'Your pen touches the cold wood of the log, the friction'
   - 'You jot down a hurried note: "Visited lamp—oil low—oil '
branches: 1

A retry at the tip files the new attempt as a sibling at the same coordinate and keeps the one it replaced. No branch is created, which is correct for a retry at the head.

6.5 Restart and resume

before restart: 8 actions, digest 66cc82944f8b6eef
after restart:  8 actions, digest 66cc82944f8b6eef
  endpoint preserved: https://inference.lan:8443/v1  timeout 600
  memories: 1

6.6 Packet capture

all packets:                 685
loopback:                    364
to/from the LAN Ollama host: 308
any other unicast:           0
DNS queries:                 none

Zero DNS: the endpoint was configured, not discovered.


7. Endpoint policy

Run inside the shipped image:

  ALLOW   same-host loopback       http://127.0.0.1:11434/v1
  ALLOW   localhost name           http://localhost:11434/v1
  ALLOW   IPv6 loopback            http://[::1]:11434/v1
  ALLOW   LAN literal              http://192.168.1.50:11434/v1
  ALLOW   link-local               http://169.254.10.5:11434/v1
  ALLOW   CGNAT / mesh VPN         http://100.64.3.4:11434/v1
  ALLOW   IPv6 unique-local        http://[fd00::5]:11434/v1
  ALLOW   docker internal name     http://ollama-container:11434/v1
  REJECT  cloud provider           https://openrouter.ai/api/v1
           openrouter.ai is a cloud inference service — this build talks to Ollama on your own mach…
  REJECT  cloud provider           https://api.openai.com/v1
  REJECT  public IPv4              http://8.8.8.8:11434/v1
           8.8.8.8 resolves to 8.8.8.8, which is a public Internet address — …
  REJECT  public IPv6              http://[2001:4860:4860::8888]/v1
  REJECT  unspecified address      http://0.0.0.0:11434/v1
           0.0.0.0 … is not on this machine and not on your own network …
  REJECT  wrong scheme             ftp://127.0.0.1/v1

Over the HTTP API, from inside the offline container:

=== cloud and public endpoints are refused ===
  400  https://openrouter.ai/api/v1     That endpoint can't be used — openrouter.ai is a cloud…
  400  https://api.openai.com/v1        …
  400  https://api.groq.com/openai/v1   …
  400  http://8.8.8.8:11434/v1          …8.8.8.8 … is a public Internet address…
=== local and LAN endpoints are accepted ===
  200  http://127.0.0.1:11434/v1
  200  http://192.168.1.50:11434/v1
  200  http://[::1]:11434/v1

Request-time enforcement against a hand-edited database is in §5.7.

test_endpoint_policy.py (31 collected) resolves hostnames through a stub, so it exercises the policy rather than the machine's DNS. It covers the split-horizon case: a name resolving to both 192.168.1.50 and a public address is refused.


8. TLS trust

$ grep -rn "verify=False\|ssl._create_unverified\|CERT_NONE\|check_hostname = False" backend/app
(the only match is prose in tlstrust.py saying such an option deliberately does not exist)

$ grep -rn "httpx.AsyncClient(\|httpx.Client(" backend/app
backend/app/routers/settings.py:122
backend/app/providers/openai_compatible.py:214
backend/app/providers/openai_compatible.py:326
backend/app/providers/openai_compatible.py:360

Four clients, all four passing verify=tlstrust.ssl_context(). Enforced by test_tls_trust.py, which walks the AST of the two modules that make outbound requests and fails if any httpx.AsyncClient is constructed without verify, or if a third module starts making requests:

$ .venv/bin/python -m pytest tests/test_tls_trust.py -q
6 passed in 0.20s

Runtime confirmation of certificate and hostname verification against the real private-CA host is §6.1 (TLS verified, peer CN = inference.lan), and of all four paths — settings/test, narration, model listing, embeddings — §6.2–6.3.

Modules that can open an outbound connection at all

$ grep -rln "httpx\|requests\.\|urllib.request\|socket\.\|aiohttp\|websocket" backend/app
backend/app/endpoints.py            # getaddrinfo only, for the policy
backend/app/providers/openai_compatible.py
backend/app/routers/settings.py
backend/app/tlstrust.py             # builds the SSL context; opens nothing

Remaining absolute URLs anywhere in backend application code:

      3 http://127.0.0.1
      2 http://localhost
      1 https://openaipublic.blob.core.windows.net   <- a comment and a SOURCE_URL
                                                        constant in encoding.py; never fetched

9. Defects found during this review

Three, all found by running the shipped build rather than by reading it. Each was corrected because the M2 evidence could not otherwise be accurate; each is reported rather than absorbed.

9.1 The memory bank was broken — summaries and embeddings silently stopped

Task exception was never retrieved
future: <Task finished coro=<run_post_turn() ...> exception=AttributeError(
  "'Settings' object has no attribute 'api_key_plain'")>
  File "/app/backend/app/memorybank.py", line 155, in summary_provider
    settings.api_key_plain,

M2 removed Settings.api_key_plain with the API key, but memorybank's two provider factories still read it. It failed in a fire-and-forget background task, so nothing surfaced to the user and no test caught it: every memory test stubs those factories out. All 604 tests passed with this defect present.

Fixed by rewriting both factories. Covered by a new test that constructs every provider factory from a real Settings row.

9.2 The configurable model timeout never reached the turn engine

settings.model_timeout_seconds was stored, validated, exposed in the API and rendered in the UI — and not passed to OpenAICompatibleProvider in turns.py or chat.py, so generation used the module default. The M2 exit criterion "no longer an undocumented hardcoded limitation" was therefore only half met: the constant had moved, but the setting was inert.

Fixed in both call sites. Covered by a new test that drives the turn endpoint, the chat endpoint and the summariser factory and asserts the configured value reaches each.

9.3 requirements.lock still pinned the removed packages

$ grep -n "quickjs\|psycopg\|cryptography" backend/requirements.lock
25:cryptography==50.0.1
36:psycopg==3.3.5
37:psycopg-binary==3.3.5
45:quickjs==1.19.4

DEVELOPMENT.md tells a new developer to install from the lock, which would have reinstalled all three. Regenerated: 40 pins → 34.

Verification after the fixes

$ .venv/bin/python -m pytest tests/ -q
606 passed, 1 warning in 126.29s
$ npm run lint   # exit 0
$ npm run build  # ✓ built in 496ms
$ docker build . # exit 0

All §5 and §6 runtime evidence above was produced by an image built after these fixes.


10. M1 capability regression checks

M1 capability Result Evidence
Vendored tokenizer, no first-turn download PASS §10.1
Self-hosted fonts PASS §10.2
Same-origin runtime assets PASS §10.2
Restrictive CSP PASS §10.2
Loopback storyteller default PASS §5.2, §6.1, §11
Same-host Ollama PASS §5.3–5.4
Trusted-LAN Ollama PASS §6.2
HTTPS / private-CA trusted-LAN PASS §6.1–6.3
Certificate + hostname verification PASS §6.1, §8
Offline story generation PASS §5.1, §5.4
SQLite persistence PASS §5.8, §6.4
Full regression suite green PASS §4

10.1 Tokenizer, inside the shipped image, offline

  vendored table sha256 matches pin: True
  count_tokens with sockets blocked: 6 tokens

10.2 Browser asset graph, fetched over loopback with no route out

CSP: default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline';
     font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none';
     base-uri 'none'; form-action 'self'; frame-ancestors 'none'

referenced by index.html:
  /favicon.svg              LOCAL
  /assets/index-*.js        LOCAL
  /assets/index-*.css       LOCAL

font urls inside the stylesheet:
  /fonts/cinzel-normal-latin-ext.woff2        200 font/woff2   14540
  /fonts/cinzel-normal-latin.woff2            200 font/woff2   25904
  /fonts/crimson-pro-italic-latin-ext.woff2   200 font/woff2   39808
  /fonts/crimson-pro-italic-latin.woff2       200 font/woff2   51432
  /fonts/crimson-pro-normal-latin-ext.woff2   200 font/woff2   37988
  /fonts/crimson-pro-normal-latin.woff2       200 font/woff2   48200
  /fonts/inter-normal-latin-ext.woff2         200 font/woff2   85068
  /fonts/inter-normal-latin.woff2             200 font/woff2   48256

Unchanged from M1, including the object-src/base-uri/form-action directives M1 added.


11. Storyteller network exposure

Launch path Listener / publish Verified by
./start.sh (native dev) uvicorn --host 127.0.0.1 --port 8000 file assertion in test_local_only_surface.py
start.ps1 (Windows) --host 127.0.0.1 same test
Production native (documented in DEVELOPMENT.md) --host 127.0.0.1 documentation
docker compose up ports: - "127.0.0.1:8000:8000" test parses the published mappings and asserts every one starts 127.0.0.1:
Container process itself --host 0.0.0.0 deliberate; the only address a published port can reach

DEVELOPMENT.md states that publishing the port to 0.0.0.0 is a deliberate decision this project's threat model does not cover. The Dockerfile carries the same warning next to EXPOSE.

Two supporting checks:

  • AIDND_CORS_ORIGINS="*" makes the application refuse to start: RuntimeError: AIDND_CORS_ORIGINS must not contain '*'…
  • An unknown /api/... path now returns 404 instead of the SPA's HTML with status 200 (verified for ten removed endpoints, §12).

12. Removed-surface verification, at runtime

From inside the offline container:

=== removed surfaces answer 404 ===
  404  GET /auth/me            404  GET /analytics/summary
  404  GET /auth/login         404  GET /analytics/collect
  404  GET /auth/register      404  GET /analytics/access
  404  GET /auth/logout        404  GET /scripts
                               404  GET /adventures/2/scripts
                               404  GET /adventures/2/script-state

=== no API key in the settings surface ===
  api_key present: False   has_api_key present: False
  after trying to set one: False
$ python -c "import app.scripting"
ImportError    # asserted by test_the_application_has_no_scripting_engine
$ grep -rn "import quickjs" backend/app
(no output)

13. Inert schema retained for compatibility

Verified present in the database and unread by the application.

Object Was Status
table scripts script library unmapped; no model, no query
table adventure_scripts per-adventure script copies unmapped
table analytics_daily visitor counters unmapped
table analytics_visitor_days visitor funnel unmapped
table access_log sign-ins, addresses, devices unmapped
column adventures.script_state scripting shared state written {} only
column actions.state_after scripting state per node written {} only
column settings.api_key encrypted cloud key never read or written
column settings.reasoning_max_tokens OpenRouter thinking budget never read
column users.demo_turns_used / _date demo cap tally never read
table users + user_id FKs multi-user ownership active, one row, internal identity only

No destructive migration was performed. One additive migration was added:

(77, "ALTER TABLE settings ADD COLUMN model_timeout_seconds INTEGER NOT NULL DEFAULT 300")

Existing M1 databases open unchanged: the offline run in §5 used the volume carrying campaigns created before these fixes, and §5.8 shows the transcript digest surviving both an image replacement and a restart.


14. What was not measured

Stated so the implementation report does not overclaim.

  1. No browser rendered the UI. No browser automation was available. §10.2 fetches the complete asset graph the page references, with correct media types and the CSP, but nobody looked at the rendered page. This is unchanged from M1, where the maintainer confirmed the render by hand.
  2. Sustained multi-turn play under the memory bank was not completed on this host. Turns 1–3 succeeded; turn 4 timed out under third-party CPU load (§5.4). The individual capabilities — narration, streaming, embeddings, memory writes, retry, fork, restart — were each measured separately.
  3. No load, soak or long-campaign testing. Out of scope for M2.

15. Closeout note — appended 2026-09-03

This section was appended after the fact and is not part of the original evidence record. Everything above was written on 2026-09-02 and describes the repository as it stood then. Nothing above has been rewritten, including its references to a working tree that was dirty at the time.

What the original report correctly said

The report above correctly described commit 8c65ae9 — the M2 implementation commit — as not containing the three fixes this review found. At the time it was written, those fixes existed only in the working tree, six files ahead of that commit. Every runtime measurement in this report was produced by an image built from that fixed tree, which the report states plainly.

What happened afterwards

The six-file correction was committed:

8652fe7cd84bca5173abb03b2a692f15fea8a98c   M2 review: two regressions the green suite hid, and the reports

That commit carries the six implementation/test/lockfile files and the two M2 reports themselves, which is why this report's own history begins there. The provenance distinction the review asked for is preserved regardless: 8c65ae9 is the M2 implementation, 8652fe7 is the review correction, and the two were never squashed.

Verification at closeout

Re-run on 2026-09-03 against the committed tree — not the working tree — so that what was verified is exactly what the repository contains:

backend    .venv/bin/python -m pytest tests/ -q     606 passed in 133.26s
frontend   npm run lint                             7 warnings, 0 errors, exit 0
frontend   npm run build                            built in 932ms; index.js 395.41 kB
root       docker build -t storyteller-m2-closeout .    exit 0
git        status --short                           clean

The 606 figure matches the count this report recorded, from the same tree.

Dependency removal was verified in the built image rather than only in the lockfile:

$ docker run --rm --entrypoint sh storyteller-m2-closeout -c "pip list | grep -iE 'quickjs|psycopg|cryptography|cffi|pycparser'"
ABSENT: none of quickjs/psycopg/cryptography/cffi/pycparser installed
32 packages total

The network measurements in §5, §6 and §10 were not re-run. The six corrected files change provider construction and a timeout value; they do not touch the endpoint policy, the TLS trust context, or the bind addresses, so the captures above remain the evidence for the network boundary.