Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was authoritative. Phase 0 execution prompts sat beside the specification; four completed milestone reports sat beside the current one; and upstream AI-DnD's own `plan/` build log and `docs/` project site still described a hosted, scripted, multi-user product with accounts — every screenshot in it showed a Scripts tab and a Sign up button, none of which has existed since M2. `planning/archive/` now holds the history and says so in its own README: `phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied. `planning/reports/` holds only the current milestone's report, because that is the one M4 planning has to read; it moves to the archive when M4's replaces it. Deleted rather than archived: the Phase 0B execution prompts and the handoff/status/summary documents, the Phase 0A discovery and triage reports, upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's template boilerplate. All of it is in Git history, and the two recommendation reports carry every conclusion the deleted research reached. Archived documents are kept verbatim. Paths written inside them point at where those files were when the document was written, which is the point: an evidence record that has been quietly edited is no longer evidence. Active documentation is corrected where it pointed at the removed trees or described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not touch" list had gone stale at M2 and claimed QuickJS scripting was still tested; its test count was 604 against an actual 638. `README.md` loses the upstream CI badge, which reported upstream's pipeline rather than this fork's, and a reference to `backend/app/worldstate/engine.py`, a file that does not exist. `planning/README.md` is rewritten as the documentation index. New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the manifest of what belongs in the ChatGPT project's Sources. Source comments referring to the deleted trees are reworded; no behaviour changes. 638 backend tests pass, frontend lints and builds, and a reference scan over all 48 tracked Markdown files reports no unresolved path in active documentation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
This commit is contained in:
co-authored by
Claude Opus 5
parent
c8755c21c2
commit
d27ee34901
@@ -0,0 +1,67 @@
|
||||
# Archive — historical material, not authoritative
|
||||
|
||||
Everything under `planning/archive/` is **evidence and history**. None of it
|
||||
governs current implementation. If an archived document and an active planning
|
||||
document disagree, the active document is right and the archived one records
|
||||
what was believed or measured at the time.
|
||||
|
||||
Do not consult this directory during ordinary milestone work unless an active
|
||||
document sends you here for a specific piece of historical evidence.
|
||||
|
||||
## What is here
|
||||
|
||||
### `phase0/` — why AI-DnD was selected
|
||||
|
||||
Phase 0A static research and Phase 0B local validation, closed 2026-09-01.
|
||||
|
||||
| File | What it is |
|
||||
| --- | --- |
|
||||
| `RESEARCH-PLAN.md` | The question Phase 0 existed to answer, and its final dispositions. |
|
||||
| `PRELIMINARY-RECOMMENDATION.md` | The Phase 0A conclusion, from static review alone. |
|
||||
| `REUSE-MATRIX.md` | What each candidate offered against the required subsystems. |
|
||||
| `AI-DND-ANALYSIS.md` | The selected base, analysed before the fork. |
|
||||
| `AI-ADVENTURE-ANALYSIS.md` | The rejected finalist still cited as the implementation reference for typed state events. |
|
||||
| `OPEN-DUNGEON-ANALYSIS.md` | The rejected finalist still cited as the UX/future-media reference. |
|
||||
| `PHASE-0B-RECOMMENDATION.md` | **The strongest single document.** The measured Phase 0B findings and the fork decision. |
|
||||
| `PHASE-0B-BASELINE.md` | Clean clones, builds and test runs for the three finalists. |
|
||||
| `PHASE-0B-AI-DND-EXPERIMENT.md` | AI-DnD driven against a real local Ollama. |
|
||||
| `PHASE-0B-AI-ADVENTURE-OLLAMA.md` | ai-adventure's Ollama and service-boundary behaviour. |
|
||||
| `PHASE-0B-OPEN-DUNGEON-HISTORY.md` | Whether Open Dungeon could be retrofitted with history. |
|
||||
| `PHASE-0B-OFFLINE-NETWORK.md` | What each candidate reached for with no route to the Internet. |
|
||||
| `PHASE-0B-UNDO-SPIKE.md` | The disposable spike that proved non-destructive undo/redo, and became ADR 012's architecture. |
|
||||
| `PHASE-0B-FOLLOWUP-CHECKS.md` | The world-state protocol, export/head, story-card lineage and Postgres checks. |
|
||||
|
||||
Phase 0A discovery and triage material (the candidate inventory, source index,
|
||||
reference-project list, licensing and static-privacy reviews, the Phase 0A
|
||||
status page) and the Phase 0B execution prompts were deleted in the 2026-09-03
|
||||
documentation cleanup. They are intermediate working documents whose
|
||||
conclusions all reached the two recommendation reports above, and they remain in
|
||||
Git history.
|
||||
|
||||
### `milestone-reports/` — completed milestone evidence
|
||||
|
||||
`M1-BASELINE-REPORT.md`, `M1-IMPLEMENTATION-REPORT.md`,
|
||||
`M2-BASELINE-REPORT.md`, `M2-IMPLEMENTATION-REPORT.md`.
|
||||
|
||||
Every architectural conclusion these reports reached has already been applied to
|
||||
the active planning documents and the ADRs — see `planning/VERSION.md`, which
|
||||
lists the corrections each milestone produced. The reports are kept for their
|
||||
measurements and their reasoning, not as instructions.
|
||||
|
||||
The **current** milestone's report stays in `planning/reports/` while it is
|
||||
still useful for reviewing the next milestone, and moves here when it is not.
|
||||
|
||||
### `decisions/` — superseded or completed ADRs
|
||||
|
||||
`008-phase0-before-build-plan.md` — a process gate ("do not start production
|
||||
work before Phase 0 closes") that Phase 0 satisfied on 2026-09-01. It
|
||||
constrains nothing now. ADR numbering continues from 012 in
|
||||
`planning/DECISIONS/`; 008 is not reused.
|
||||
|
||||
## A note on paths inside these files
|
||||
|
||||
Archived documents are kept **verbatim**. File paths written inside them refer
|
||||
to where those files lived when the document was written — before this archive
|
||||
existed, and in some cases before files were deleted. That is deliberate: an
|
||||
evidence record that has been quietly edited is no longer evidence. Resolve any
|
||||
such path against Git history, not against the current tree.
|
||||
@@ -0,0 +1,32 @@
|
||||
# ADR 008 — Complete Phase 0 Before Detailed Build Planning
|
||||
|
||||
**Status:** Accepted and completed
|
||||
|
||||
## Decision
|
||||
|
||||
The project completed repository research, validation, architecture selection, and critical prototypes before finalizing the detailed production implementation milestone plan.
|
||||
|
||||
## Context
|
||||
|
||||
Multiple candidate open-source projects implemented overlapping parts of the desired system. The correct build sequence depended on which codebase survived local validation.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
- write full implementation plan immediately,
|
||||
- begin coding against the first plausible project,
|
||||
- perform a bounded Phase 0 and then finalize the build plan.
|
||||
|
||||
## Reason
|
||||
|
||||
The bounded Phase 0 prevented false assumptions from becoming production architecture. In particular, runtime testing discovered destructive AI-DnD Undo, an offline tokenizer fetch, remote fonts, Open Dungeon's lack of tests and destructive summary/history assumptions, and the relative-delta state-protocol failure mode.
|
||||
|
||||
## Completion Record
|
||||
|
||||
Phase 0 selected AI-DnD as the production base, validated the non-destructive head-cursor approach, selected explicit typed narrative-state events, established local-only hardening requirements, and produced the production `BUILD-MILESTONES.md`.
|
||||
|
||||
## Consequences
|
||||
|
||||
- the placeholder build plan has been replaced by a production milestone sequence,
|
||||
- `SPECIFICATION.md` and `TECHNICAL-DESIGN.md` are the v1.0 planning baseline,
|
||||
- production coding still requires explicit authorization and a milestone-specific prompt,
|
||||
- the current review intentionally stops before preparing that prompt.
|
||||
@@ -0,0 +1,466 @@
|
||||
# M1 — Production Fork and Offline Baseline: evidence report
|
||||
|
||||
**Date:** 2026-09-01/02
|
||||
**Milestone:** M1, `planning/BUILD-MILESTONES.md`
|
||||
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` (see `PROVENANCE.md`)
|
||||
**Scope note:** M1 only. No M2 work was started; nothing was removed from the
|
||||
inherited hosted/cloud/scripting surface.
|
||||
|
||||
This report records what was run and what was observed. Where a condition was
|
||||
not reproduced exactly as the acceptance test specifies, it says so and says
|
||||
what was reproduced instead. Nothing untested is called a PASS.
|
||||
|
||||
Hostnames and LAN addresses below are **placeholders** — `inference.lan`,
|
||||
`192.168.0.0/24`. The real ones are in the workspace's untracked notes, not in
|
||||
this repository. Everything else, including packet counts and digests, is
|
||||
verbatim.
|
||||
|
||||
---
|
||||
|
||||
## 1. Test environment
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Host | Ubuntu 24.04.4 LTS, x86-64, 4 cores, 15 GB RAM, **no GPU** |
|
||||
| Python | 3.12.3 |
|
||||
| Node / npm | 22.23.1 / 10.9.8 |
|
||||
| Docker | 29.7.2 |
|
||||
| Ollama | 0.33.2 (`ollama/ollama@sha256:020e4134285e…`) |
|
||||
| Narrator model | `qwen2.5:3b-instruct` (`qwen2.5:0.5b` in one earlier run) |
|
||||
| Second machine | `inference.lan` (`192.168.0.50`), StartOS, Ollama 0.33.0 over HTTPS on 8443 |
|
||||
| Embedding model | `nomic-embed-text:latest` |
|
||||
| Application | this fork, working tree on `m1-production-baseline` |
|
||||
| Browser | the maintainer's own desktop browser, against the native run (§6) |
|
||||
|
||||
The lack of a GPU matters and shows up twice below: a cold model load on four
|
||||
CPU cores can exceed the application's hardcoded 120-second model timeout.
|
||||
|
||||
## 2. What was changed
|
||||
|
||||
Thirteen upstream files modified; twenty-two added, of which eleven are the
|
||||
vendored font files and their licences. No upstream file was deleted.
|
||||
|
||||
| Change | Files |
|
||||
| --- | --- |
|
||||
| Fork lineage and licence provenance | merge commit `7f182a8`, `PROVENANCE.md` |
|
||||
| Tokenizer no longer downloads | `backend/app/context/encoding.py`, `backend/app/context/vendor/cl100k_base.tiktoken`, `backend/app/context/builder.py` |
|
||||
| Fonts self-hosted | `frontend/index.html`, `frontend/src/index.css`, `frontend/src/styles/fonts.css`, `frontend/public/fonts/*`, `frontend/tools/vendor_fonts.py`, `.gitattributes` |
|
||||
| CSP narrowed to same-origin | `backend/app/main.py` |
|
||||
| TLS verified against the machine's own CA store as well as certifi's | `backend/app/tlstrust.py`, `backend/app/providers/openai_compatible.py`, `backend/app/routers/settings.py`, `backend/requirements.txt` |
|
||||
| `woff2` served with its real media type | `backend/app/main.py` |
|
||||
| Loopback binding stated explicitly | `start.sh`, `start.ps1`, `docker-compose.yml`, `Dockerfile` |
|
||||
| Reproducible environment | `backend/requirements.lock`, `DEVELOPMENT.md` |
|
||||
| Regression tests | `backend/tests/test_offline_assets.py`, `backend/tests/test_tls_trust.py` |
|
||||
| Research scratch ignored | `.gitignore` |
|
||||
|
||||
## 3. Regression suite
|
||||
|
||||
| Suite | Before M1 | After M1 |
|
||||
| --- | --- | --- |
|
||||
| `backend/tests` | 632 passed, 0 failed (184.7 s) | **648 passed, 0 failed** (212.1 s) |
|
||||
| `frontend`: `npm run lint` | — | exit 0 (6 pre-existing `only-export-components` warnings) |
|
||||
| `frontend`: `npm run build` | — | succeeds |
|
||||
| `docker build` | — | succeeds |
|
||||
|
||||
The 632-test baseline was captured on the pinned upstream commit before any
|
||||
change, and matches the Phase 0B figure. The sixteen new tests are
|
||||
`test_offline_assets.py` (10) and `test_tls_trust.py` (6). **There are no known
|
||||
failing tests and no documented exceptions.**
|
||||
|
||||
The two `dist`-reading tests in `test_offline_assets.py` skip if the SPA has
|
||||
not been built; the run above had it built, so they executed.
|
||||
|
||||
### Proof the tokenizer guard is a real guard
|
||||
|
||||
With `TIKTOKEN_CACHE_DIR` pointed at an empty directory and Python's socket
|
||||
functions replaced:
|
||||
|
||||
```text
|
||||
UPSTREAM PATH raises: AssertionError socket opened
|
||||
VENDORED PATH: 2 tokens, no socket opened
|
||||
```
|
||||
|
||||
The vendored table's SHA-256 is
|
||||
`223921b76ee99bde995b7ff738513eef100fb51d18c93597a113bcffe865b2a7`, identical
|
||||
to the digest hardcoded in `tiktoken_ext/openai_public.py`, and the encoding it
|
||||
produces was checked token-for-token against `tiktoken.get_encoding` over
|
||||
ASCII, accented text, CJK, emoji, CRLF and special-token literals.
|
||||
|
||||
## 4. Run 1 — offline, same-host Ollama
|
||||
|
||||
**Setup.** An `--internal` Docker network (no NAT, no external DNS). Ollama
|
||||
runs in one container; the production image runs in a second container that
|
||||
**shares Ollama's network namespace**, so Ollama is genuinely on the
|
||||
storyteller's own loopback and neither has any route out. Both were driven from
|
||||
a client in the same namespace, over `127.0.0.1:8000`.
|
||||
|
||||
**Isolation confirmed from inside the application container, before any test:**
|
||||
|
||||
```text
|
||||
blocked 1.1.1.1:443 OSError: Network is unreachable
|
||||
blocked openaipublic.blob.core.windows.net:443 gaierror
|
||||
blocked fonts.googleapis.com:443 gaierror
|
||||
blocked fonts.gstatic.com:443 gaierror
|
||||
blocked openrouter.ai:443 gaierror
|
||||
blocked github.com:443 gaierror
|
||||
```
|
||||
|
||||
**Listening sockets in that namespace:**
|
||||
|
||||
```text
|
||||
127.0.0.1:8000 uvicorn (the storyteller)
|
||||
[::]:11434 ollama
|
||||
127.0.0.11:44683 Docker's embedded DNS
|
||||
```
|
||||
|
||||
**Story play.** Campaign "Continuity Test (3b)" created and played for six
|
||||
turns, twelve actions, all offline:
|
||||
|
||||
```text
|
||||
turn 1: 'The old lighthouse looms dark against the stormy sky, its silence heavy as the wind…'
|
||||
turn 2: "Gripping the lantern's heavy brass handle, you flick it on and off, each attempt a futile…"
|
||||
turn 3: 'The lamp room is dim and musty, the thick air choking your breath…'
|
||||
turn 4: 'You find the dated handwriting, the last entry noting the storm started three days ago…'
|
||||
turn 5: 'The night outside is a tempest, waves crashing against the shore with a deafening roar…'
|
||||
turn 6: 'The cabinet is cold and heavy, the lock stubbornly refusing to budge…'
|
||||
```
|
||||
|
||||
**Browser asset graph, fetched over loopback with no route out:**
|
||||
|
||||
```text
|
||||
GET / 200 text/html
|
||||
/favicon.svg 200 9 522 bytes
|
||||
/assets/index-*.js 200 933 695 bytes
|
||||
/assets/index-*.css 200 59 902 bytes
|
||||
/fonts/cinzel-normal-latin.woff2 200 font/woff2 25 904
|
||||
/fonts/cinzel-normal-latin-ext.woff2 200 font/woff2 14 540
|
||||
/fonts/crimson-pro-normal-latin.woff2 200 font/woff2 48 200
|
||||
/fonts/crimson-pro-normal-latin-ext.woff2 200 font/woff2 37 988
|
||||
/fonts/crimson-pro-italic-latin.woff2 200 font/woff2 51 432
|
||||
/fonts/crimson-pro-italic-latin-ext.woff2 200 font/woff2 39 808
|
||||
/fonts/inter-normal-latin.woff2 200 font/woff2 48 256
|
||||
/fonts/inter-normal-latin-ext.woff2 200 font/woff2 85 068
|
||||
```
|
||||
|
||||
Every URL `index.html` references is same-origin. The response carried:
|
||||
|
||||
```text
|
||||
content-security-policy: default-src 'self'; script-src 'self';
|
||||
style-src 'self' 'unsafe-inline'; font-src 'self'; img-src 'self' data:;
|
||||
connect-src 'self'; object-src 'none'; base-uri 'none'; form-action 'self';
|
||||
frame-ancestors 'none'
|
||||
x-content-type-options: nosniff referrer-policy: same-origin x-frame-options: DENY
|
||||
```
|
||||
|
||||
**Packet capture** (tcpdump in the same namespace for the whole run, 6 074
|
||||
packets):
|
||||
|
||||
```text
|
||||
total packets: 6074
|
||||
loopback (127.0.0.0/8): 6062
|
||||
non-loopback unicast: 0
|
||||
remainder (12): received mDNS / ICMPv6 router solicitations from the
|
||||
bridge — inbound multicast, not sent by this namespace
|
||||
TCP connections opened: 127.0.0.1:8000 (storyteller API)
|
||||
127.0.0.1:11434 (Ollama)
|
||||
127.0.0.1:11499 (the deliberately dead port in A05)
|
||||
two ephemeral loopback ports (Ollama's model runner)
|
||||
```
|
||||
|
||||
**One DNS observation, and it is not the storyteller.** Four queries appear,
|
||||
all failing:
|
||||
|
||||
```text
|
||||
127.0.0.11.53 > … ServFail q: A? ollama.com.
|
||||
127.0.0.11.53 > … ServFail q: AAAA? ollama.com. (×2 each)
|
||||
```
|
||||
|
||||
`ollama.com` is queried by **the Ollama server itself**, in the shared
|
||||
namespace, not by the storyteller. It failed, nothing depended on it, and no
|
||||
story data could have been in it. It is recorded here rather than dismissed:
|
||||
on a deployment with a route to the Internet, the inference server has its own
|
||||
outbound behaviour, and the storyteller's local-only guarantee does not extend
|
||||
to it. Confirming and, if wanted, suppressing that is an Ollama configuration
|
||||
question — worth settling before release, and out of M1's scope.
|
||||
|
||||
## 5. Run 2 — offline, Ollama on a second machine on the LAN
|
||||
|
||||
**Setup.** Ollama 0.33.0 on `inference.lan` (`192.168.0.50`), a **separate
|
||||
physical machine** on the trusted LAN, serving HTTPS on port 8443 with a
|
||||
certificate issued by `CN = StartOS Local Intermediate CA`. `qwen2.5:3b-instruct`
|
||||
and `nomic-embed-text` were installed there for this run.
|
||||
|
||||
The storyteller runs in a container on this host with `NET_ADMIN`, its default
|
||||
route **deleted** and replaced by a route to `192.168.0.0/24` only, its
|
||||
resolver pointed at nothing, and `inference.lan` supplied as a static hosts
|
||||
entry. So: the LAN is reachable, the Internet is not, and the endpoint is
|
||||
configured explicitly rather than discovered. Uvicorn binds `127.0.0.1:8000`
|
||||
inside that container.
|
||||
|
||||
```text
|
||||
route table: 172.17.0.0/16 dev eth0 …
|
||||
192.168.0.0/24 via 172.17.0.1 dev eth0 (no default route)
|
||||
listeners: LISTEN 127.0.0.1:8000 (nothing on 172.17.0.3)
|
||||
```
|
||||
|
||||
**Isolation, checked from inside before anything else:**
|
||||
|
||||
```text
|
||||
outbound Internet
|
||||
blocked 1.1.1.1:443 OSError: Network is unreachable
|
||||
blocked 140.82.121.4:443 OSError: Network is unreachable
|
||||
blocked 104.16.0.1:443 OSError: Network is unreachable
|
||||
name resolution
|
||||
no resolution github.com / openrouter.ai / fonts.gstatic.com /
|
||||
openaipublic.blob.core.windows.net (gaierror)
|
||||
the approved LAN host
|
||||
inference.lan -> 192.168.0.50
|
||||
TLS OK, peer CN = inference.lan
|
||||
```
|
||||
|
||||
**Model discovery over the LAN endpoint:**
|
||||
|
||||
```json
|
||||
{"ok": true, "models": ["qwen2.5:3b-instruct", "nomic-embed-text:latest"]}
|
||||
```
|
||||
|
||||
**Story play.** Nine turns, then a tenth after a restart. That machine has
|
||||
faster hardware than this one, and it shows:
|
||||
|
||||
```text
|
||||
turn 1 (13.0s): "The lamp flickers faintly, a single coal barely keeping the structure's shadow at bay."
|
||||
turn 2 ( 4.9s): "My lantern's light has failed. Checking the wick, I find it's burnt down to a stub."
|
||||
turn 3 ( 4.4s): 'I turn the wick higher, attempting to coax more flame from the almost-embers.'
|
||||
turn 4 ( 5.9s): 'The storm shutter blocks out everything but the dark ocean…'
|
||||
turn 5 ( 6.8s): 'I insert the key and hear the satisfying click as the lock turns…'
|
||||
turn 6 ( 8.1s): 'Inside the box, I find a small vial of oil and a note warning…'
|
||||
turn 7 (10.6s) … turn 9 (14.8s)
|
||||
```
|
||||
|
||||
**State extraction, summaries and embeddings** ran through the same endpoint:
|
||||
the memory bank produced memories and embedded them with `nomic-embed-text` on
|
||||
the remote host, so the narrator, the summarizer and the embedder all went
|
||||
over the LAN, not just the narrator.
|
||||
|
||||
**Restart and resume.** The storyteller container was restarted while the
|
||||
inference host was left alone:
|
||||
|
||||
```text
|
||||
before restart: 18 actions, head 18, digest f24caf86a744ab36
|
||||
after restart: 18 actions, head 18, digest f24caf86a744ab36
|
||||
endpoint still https://inference.lan:8443/v1
|
||||
embedding model still nomic-embed-text:latest, memories intact
|
||||
next turn (6.4s): 'You make your way back to the lighthouse, lantern in hand…'
|
||||
```
|
||||
|
||||
**Packet capture** across the whole run:
|
||||
|
||||
```text
|
||||
all packets: 1633
|
||||
loopback (127.0.0.0/8): 730
|
||||
to/from inference.lan 192.168.0.50: 893
|
||||
any other unicast: 0
|
||||
|
||||
TCP connections opened: 192.168.0.50:8443 (19)
|
||||
127.0.0.1:8000 (17)
|
||||
|
||||
DNS queries: none — no name was looked up at all
|
||||
```
|
||||
|
||||
Zero packets to anything but the storyteller's own loopback and the approved
|
||||
inference host, and not a single DNS query, because the endpoint was configured
|
||||
rather than resolved.
|
||||
|
||||
### 5.1 What had to be fixed to get here: TLS trust
|
||||
|
||||
The first attempt failed, and the failure was the application's:
|
||||
|
||||
```text
|
||||
{"ok": false, "detail": "Connection failed: [SSL: CERTIFICATE_VERIFY_FAILED]
|
||||
certificate verify failed: self-signed certificate in certificate chain"}
|
||||
```
|
||||
|
||||
`httpx` verifies against the `certifi` bundle, which carries the public web's
|
||||
CAs and nothing else. The host's certificate comes from a local StartOS CA that
|
||||
the user had already installed at `/usr/local/share/ca-certificates/local-ca.crt`,
|
||||
which is why `curl` and the browser accepted the same endpoint on the same
|
||||
machine. Confirmed directly:
|
||||
|
||||
```text
|
||||
certifi bundle (httpx default) FAIL SSLCertVerificationError
|
||||
system trust store OK peer CN=inference.lan
|
||||
```
|
||||
|
||||
This is not an edge case for this product. A trusted-LAN inference host is a
|
||||
first-class v1 deployment (`planning/DECISIONS/002-ollama-only-v1.md`), and such
|
||||
a host is unlikely to hold a publicly-issued certificate.
|
||||
|
||||
`backend/app/tlstrust.py` builds one verification context that **unions** the
|
||||
platform CA store with certifi's bundle, and all four outbound HTTP clients use
|
||||
it. Deliberately a union rather than a swap: using the platform store alone
|
||||
would be a behaviour change, and an image with an empty or stale system store
|
||||
would start failing on endpoints that used to work. A union can only add trust
|
||||
the user has already granted at the operating-system level.
|
||||
|
||||
Verification itself is untouched — `verify_mode=CERT_REQUIRED`,
|
||||
`check_hostname=True`, and no "insecure" escape hatch was added. Measured after
|
||||
the change:
|
||||
|
||||
```text
|
||||
OK inference.lan (local CA) CN=inference.lan
|
||||
OK github.com (public CA) CN=github.com
|
||||
OK pypi.org (public CA) CN=pypi.org
|
||||
certifi-only context still rejects inference.lan (so the union is what changed)
|
||||
```
|
||||
|
||||
`backend/tests/test_tls_trust.py` holds the line in both directions: it asserts
|
||||
every certifi root survives in the union, that verification is not weakened,
|
||||
and — by walking the AST of the two modules that make outbound requests — that
|
||||
no fifth HTTP client is ever added without the shared context.
|
||||
|
||||
### 5.2 An earlier container-only run
|
||||
|
||||
Before a second machine was available, the same sequence was run with Ollama in
|
||||
a separate container, network namespace and IP on an `--internal` network. It produced the
|
||||
same result (all 20 model requests to the configured endpoint, zero other
|
||||
unicast packets) and is superseded by the run above, which has the physical
|
||||
separation A06 actually asks for.
|
||||
|
||||
## 6. Run 3 — native (non-Docker) run, same-host Ollama
|
||||
|
||||
The production build run directly from the venv, SPA served by FastAPI, Ollama
|
||||
on host loopback:
|
||||
|
||||
```text
|
||||
LISTEN 0 2048 127.0.0.1:8000 users:(("uvicorn",pid=341269,fd=14))
|
||||
LISTEN 0 4096 127.0.0.1:11434
|
||||
|
||||
connect to 192.168.0.10:8000 (this host's LAN address) -> Connection refused
|
||||
GET http://127.0.0.1:8000/ -> 200
|
||||
```
|
||||
|
||||
A campaign was created and played on this instance too. **This run had Internet
|
||||
available** — its purpose was the native listener check and a real end-to-end
|
||||
run outside Docker, not the offline proof, which Runs 1 and 2 carry.
|
||||
|
||||
**The UI was opened in a real browser.** No browser automation was available in
|
||||
this session, so this step was done by hand: the maintainer loaded
|
||||
<http://127.0.0.1:8000>, saw the campaign list render, clicked into the
|
||||
"Continuity Test (native run)" campaign, and saw the adventure and its
|
||||
transcript. So the inherited SPA loads, runs, and talks to the API from a
|
||||
browser — the last M1 claim that had been resting on inference rather than
|
||||
observation.
|
||||
|
||||
Two limits on what that shows, stated so the evidence is not read wider than it
|
||||
is. It was this native run, which **had Internet available**, so it is not
|
||||
itself offline evidence; the offline proof that the page needs nothing remote
|
||||
is the asset-graph fetch in §4, made with no route out. And the browser's
|
||||
network panel was not inspected, so "the browser requested nothing external" is
|
||||
carried by §4 and by the CSP — which names no remote origin and would block one
|
||||
— rather than by a devtools capture.
|
||||
|
||||
## 7. Acceptance results
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| **A01** Start application offline | **PASS** | Run 1: app started, campaign created, six turns generated, with `1.1.1.1` unreachable and no name resolving. The "open UI" step is covered by the asset-graph fetch there and by the browser render in §6. |
|
||||
| **A02** Storyteller loopback default | **PASS** | Run 1 and Run 2 listeners are `127.0.0.1:8000` only; Run 3 refuses connections on the host's LAN address; `docker-compose.yml` publishes to `127.0.0.1`. |
|
||||
| **A03** No cloud API key | **PASS** | `api_key` empty in every run; connection test, turns, summaries and embeddings all succeeded. |
|
||||
| **A04** Campaign survives restart | **PASS** | Run 1: 12 actions before and after a container restart, identical transcript and head. Run 2: identical digest `dd59e2a564df4e62` across restart, then play resumed. |
|
||||
| **A05** Failed model call does not corrupt story | **PASS** | Two induced failures (nonexistent model; dead endpoint port) plus one natural timeout. Accepted-prefix digest `2ca6ab528178e44e` unchanged through all of it; AI-action count stayed at 6; recovery by `continue` produced turn 7 with the prefix still unchanged. See §7.1. |
|
||||
| **A06** Trusted-LAN Ollama inference | **PASS** | Run 2 — Ollama on `inference.lan`, a second physical machine on the trusted LAN, over verified HTTPS, with the storyteller's Internet route removed and its listener on `127.0.0.1`. Nine turns plus a post-restart turn, summaries and embeddings included. Capture: 893 packets to the approved host, 730 loopback, **zero** elsewhere, **zero** DNS queries. Required a TLS trust fix first — §5.1. |
|
||||
| **H01** No unexpected outbound connections | **PASS for the application** | Zero non-loopback unicast packets in Run 1; in Run 2, zero packets outside loopback and the approved host and zero DNS queries of any kind. One caveat, not the storyteller's: Ollama itself queried `ollama.com` (§4). |
|
||||
| **H02** No telemetry | **PASS** | No outbound destination in either capture; the inherited `analytics.py` writes to two local SQLite tables and opens no socket (Phase 0B static analysis, re-confirmed by the captures). |
|
||||
| **H03** No cloud provider required | **PASS as stated** | Nothing cloud was reachable in Runs 1 and 2 and everything worked. The test's *preferred* final state — "cloud provider controls are absent, not merely unused" — is **not** met and is M2's scope by design. |
|
||||
| **H11** No first-use runtime asset download | **PASS** | First turn on a fresh database succeeded with no route out; the whole browser asset graph resolved same-origin; the tokenizer is built from a vendored, digest-checked table. |
|
||||
|
||||
### 7.1 A05 in detail, including a real behaviour worth knowing
|
||||
|
||||
Baseline: 12 actions, 6 of them AI, digest `2ca6ab528178e44e`.
|
||||
|
||||
```text
|
||||
invalid local model -> events ['player','error']
|
||||
"Endpoint or model not found (HTTP 404) … model
|
||||
'no-such-model-v9' not found"
|
||||
13 actions, still 6 AI actions
|
||||
endpoint down (:11499)-> events ['player','error']
|
||||
"Could not connect to http://127.0.0.1:11499/v1 —
|
||||
is the AI server running?"
|
||||
14 actions, still 6 AI actions
|
||||
accepted prefix through the pre-failure head: 12 actions, digest 2ca6ab528178e44e — unchanged
|
||||
the two added rows: [28] story 'I strike a match.'
|
||||
[29] story 'I strike a match again.'
|
||||
recovery (continue) -> 40 SSE events, done; 15 actions, 7 AI actions
|
||||
prefix digest still 2ca6ab528178e44e
|
||||
```
|
||||
|
||||
**The player's own typed action is committed before the model is called**
|
||||
(`run_player_turn` in `backend/app/routers/adventures/turns.py`), so a failed
|
||||
turn leaves the player's text at the head with no reply. No AI output is ever
|
||||
partially committed. That satisfies A05 as written — prior story intact, the
|
||||
failed turn not committed as accepted, retry available — and it is deliberate:
|
||||
it means a model failure never eats what the player typed. It is worth stating
|
||||
explicitly because "the story is unchanged" is not literally true; "the
|
||||
accepted story is unchanged" is.
|
||||
|
||||
## 8. Findings and open items
|
||||
|
||||
Nothing here blocks M1. Each is recorded because it was observed, not inferred.
|
||||
|
||||
1. **The model timeout is 120 s and hardcoded** (`httpx.Timeout(120, connect=10)`
|
||||
in `backend/app/providers/openai_compatible.py`). On this GPU-less
|
||||
four-core host, a *cold* load of `qwen2.5:3b-instruct` — or two models
|
||||
contending after the memory bank pulls in `nomic-embed-text` — exceeded it
|
||||
three times during these runs. Once warm, a full turn took 9 s. This is
|
||||
presented as an environment/tuning finding, not an application defect; a
|
||||
configurable timeout is a small change, and changing product behaviour was
|
||||
outside M1's scope.
|
||||
2. **Ollama queries `ollama.com` on its own account** (§4). Outside the
|
||||
storyteller's code, inside the user's trust boundary. Worth settling before
|
||||
release: the local-only claim covers what this application sends, and a user
|
||||
reading a packet capture will see that query.
|
||||
3. **The memory bank is off per adventure by default** (`auto_summarize` and
|
||||
`memory_bank_enabled` are false on a new adventure) even when an embedding
|
||||
model is configured globally. Discovered while trying to exercise
|
||||
embeddings; it is inherited behaviour, not a regression.
|
||||
4. **`docs/*.html` still links Google Fonts.** That is upstream's GitHub Pages
|
||||
project site; it is not served by the application, not part of any build,
|
||||
and not covered by the runtime rule. Left alone deliberately, and the
|
||||
regression tests scope themselves to `frontend/` so they do not give a false
|
||||
signal about it.
|
||||
5. **A trusted-LAN endpoint over HTTPS needed a code change to work at all**
|
||||
(§5.1), and it was found only by pointing the application at a real
|
||||
StartOS-hosted Ollama. Static review would not have found it: every
|
||||
candidate report and every local run up to that point used plain HTTP to
|
||||
loopback, where certificate verification never happens. It is fixed and
|
||||
tested, and recorded here because the class of bug — "works for curl,
|
||||
fails for us" — is worth remembering when M2 formalises endpoint policy.
|
||||
6. **The browser's own network panel was never inspected** (§6). The UI has
|
||||
now been rendered by hand and works, and §4 shows the page's whole asset
|
||||
graph resolving same-origin with no route out, so nothing rests on
|
||||
inference. A devtools capture during an offline session would still be the
|
||||
most direct form of that evidence, and costs a minute if anyone wants it.
|
||||
|
||||
## 9. Definition of done
|
||||
|
||||
| Requirement | Status |
|
||||
| --- | --- |
|
||||
| Starts with outbound Internet blocked after setup | met — Runs 1 and 2 |
|
||||
| Opens the inherited browser UI locally | met — rendered in a browser, campaign opened and read (§6) |
|
||||
| Generates and persists turns through same-host Ollama | met — Run 1 |
|
||||
| Generates and persists turns through a configured remote Ollama | met — Run 2, on a second physical machine |
|
||||
| Restarts and resumes the campaign | met — Runs 1 and 2 |
|
||||
| Survives a failed model call without corrupting accepted state | met — §7.1 |
|
||||
| No first-use tokenizer/font/runtime-asset request | met — §3, §4 |
|
||||
| Storyteller UI/API loopback-bound by default | met — §4, §5, §6 |
|
||||
| Regression suite passes, or failures documented | met — 648 passed, 0 failed |
|
||||
|
||||
Every line of the definition of done is met, and none of them rests on an
|
||||
untested assumption. **No M2 work has been started.**
|
||||
|
||||
## 10. Artefacts not committed
|
||||
|
||||
The packet captures (`offline-samehost.pcap` 2.7 MB, `offline-lan.pcap` 0.8 MB,
|
||||
`lan-remote-host.pcap` 0.4 MB) and the throwaway driver scripts live in this
|
||||
session's scratchpad, not in the repository. Everything drawn from them is
|
||||
quoted above; the procedure in `DEVELOPMENT.md` regenerates them.
|
||||
@@ -0,0 +1,995 @@
|
||||
# M1 — Implementation Review Report
|
||||
|
||||
**Date:** 2026-09-02
|
||||
**Milestone:** M1, *Establish Production Fork and Offline Baseline*
|
||||
**Audience:** the architecture/design reviewer deciding whether to prepare M2
|
||||
**Companion:** `planning/reports/M1-BASELINE-REPORT.md` holds the raw run logs
|
||||
and packet-capture output this report summarises. Where the two differ in
|
||||
detail, that one is the primary record.
|
||||
|
||||
Hostnames and LAN addresses are **placeholders** (`inference.lan`,
|
||||
`192.168.0.0/24`). The real ones are in the workspace's untracked notes, never
|
||||
in this repository. Packet counts, digests, timings and command output are
|
||||
verbatim.
|
||||
|
||||
---
|
||||
|
||||
# A. Executive Result
|
||||
|
||||
**Overall M1 result: PASS.**
|
||||
|
||||
- **Is the application a playable production baseline?** Yes. A campaign can be
|
||||
created, played, persisted, restarted and resumed from the inherited browser
|
||||
UI, with outbound Internet blocked. This was exercised end to end, not
|
||||
inferred: three separate runs, plus a by-hand browser check in which the
|
||||
campaign list rendered and a campaign was opened and read.
|
||||
- **Were both Ollama paths demonstrated?** Yes, both at runtime. Same-host
|
||||
loopback in an air-gapped container; trusted-LAN against Ollama 0.33.0 on a
|
||||
**second physical machine**, over TLS, with the storyteller's default route
|
||||
deleted so the LAN was reachable and the Internet was not.
|
||||
- **Proceed to M2?** Yes.
|
||||
- **Blockers before M2?** **None.** Section M lists debt and inherited
|
||||
behaviour, none of which blocks M2 and most of which M2 removes by design.
|
||||
|
||||
One qualification worth the reviewer's attention: M1 required a code change
|
||||
that was **not in its planned scope** — outbound TLS verification (§G, §J).
|
||||
Without it the trusted-LAN path did not work at all against a realistic host,
|
||||
so it was in M1's critical path even though the milestone text never mentions
|
||||
it.
|
||||
|
||||
---
|
||||
|
||||
# B. Repository / Provenance
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `m1-production-baseline` |
|
||||
| HEAD | `c1a73b3d77e48491196e8887ee5abc2f818af05e` |
|
||||
| HEAD signature | good (`%G? = G`) |
|
||||
| Working tree | clean apart from one documentation correction, below |
|
||||
|
||||
### M1 commits
|
||||
|
||||
| Commit | Sig | Parents | Subject |
|
||||
| --- | --- | --- | --- |
|
||||
| `c1a73b3` | G | `7f182a8` | M1: make the first story turn work with no Internet |
|
||||
| `7f182a8` | G | `717670a`, `d72f7c1` | Fork AI-DnD at d72f7c1 as the production base |
|
||||
|
||||
`7f182a8` is the fork import: a merge with **two parents** — the planning
|
||||
package's own history (`717670a`) and upstream AI-DnD (`d72f7c1`). `c1a73b3`
|
||||
is all of the M1 work.
|
||||
|
||||
### Upstream
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Project | AI-DnD, <https://github.com/parththakkar106/AI-DnD> |
|
||||
| Pinned commit | `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` |
|
||||
| Subject | "Stop paying twice for a block a retry can still throw away" |
|
||||
| Position | tip of `upstream/main` when the fork was taken, 1 Sep 2026 |
|
||||
| Substituted? | **No.** The pinned commit was fetched and verified before use. |
|
||||
|
||||
### How provenance is preserved
|
||||
|
||||
Upstream history is *in* this repository rather than copied out of it, so the
|
||||
claim is checkable rather than asserted:
|
||||
|
||||
```console
|
||||
$ git cat-file -t d72f7c1bda0f34fccd84afb7a25c34eb01c901de
|
||||
commit
|
||||
$ git merge-base --is-ancestor d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD && echo yes
|
||||
yes
|
||||
$ git rev-list --count d72f7c1bda0f34fccd84afb7a25c34eb01c901de
|
||||
172
|
||||
```
|
||||
|
||||
All 172 upstream commits are reachable, not a squashed snapshot. Upstream paths
|
||||
are unchanged (`backend/`, `frontend/`, `docs/`, …), so a later upstream commit
|
||||
can still be fetched and cherry-picked against matching files.
|
||||
|
||||
### Licence and provenance files
|
||||
|
||||
| File | State |
|
||||
| --- | --- |
|
||||
| `LICENSE` | **unmodified.** `git diff d72f7c1 HEAD -- LICENSE` is empty. MIT, © 2026 Parth Thakkar. |
|
||||
| `PROVENANCE.md` | **new.** Upstream commit, licence terms, re-verification commands, both vendored assets with sources and digests, and the full list of what M1 changed. |
|
||||
| `frontend/public/fonts/OFL-*.txt` | **new.** SIL OFL 1.1 text for each vendored family, shipped beside the fonts as the licence requires. |
|
||||
|
||||
### Final `git status`
|
||||
|
||||
Clean except for one file, which is a **documentation correction, not an
|
||||
implementation change**:
|
||||
|
||||
```text
|
||||
M planning/reports/M1-BASELINE-REPORT.md
|
||||
```
|
||||
|
||||
Signing the fork-import commit re-hashed it from `46d34dc` to `7f182a8`, and
|
||||
the baseline report's change table still cited the pre-signing hash. That one
|
||||
line now cites `7f182a8`. It is uncommitted and needs a signed commit; §N has
|
||||
the command. No other file differs from `c1a73b3`.
|
||||
|
||||
---
|
||||
|
||||
# C. What Changed
|
||||
|
||||
`c1a73b3` — **35 files changed, 102 102 insertions, 23 deletions.** The
|
||||
insertion count is dominated by two vendored assets: the tokenizer table
|
||||
(100 256 lines) and three OFL licence texts (279 lines). Excluding vendored
|
||||
data and the baseline report, M1 is **1 101 inserted lines against 23 deleted**
|
||||
— code, tests, generated CSS and documentation. Of those, roughly 390 are new
|
||||
application code and tests, 352 are documentation, and the rest is the
|
||||
generated font CSS, the lockfile and the vendoring script.
|
||||
|
||||
```text
|
||||
.gitattributes | 4 +
|
||||
.gitignore | 4 +
|
||||
DEVELOPMENT.md | 246 +
|
||||
Dockerfile | 5 +
|
||||
PROVENANCE.md | 106 +
|
||||
backend/app/context/builder.py | 9 +-
|
||||
backend/app/context/encoding.py | 90 +
|
||||
backend/app/context/vendor/cl100k_base.tiktoken | 100256 +++++++++++++++
|
||||
backend/app/main.py | 29 +-
|
||||
backend/app/providers/openai_compatible.py | 14 +-
|
||||
backend/app/routers/settings.py | 4 +-
|
||||
backend/app/tlstrust.py | 47 +
|
||||
backend/requirements.lock | 57 +
|
||||
backend/requirements.txt | 4 +
|
||||
backend/tests/test_offline_assets.py | 160 +
|
||||
backend/tests/test_tls_trust.py | 93 +
|
||||
docker-compose.yml | 7 +-
|
||||
frontend/index.html | 10 +-
|
||||
frontend/public/fonts/OFL-*.txt | 279 +
|
||||
frontend/public/fonts/*.woff2 | Bin 0 -> 351196 bytes
|
||||
frontend/src/index.css | 1 +
|
||||
frontend/src/styles/fonts.css | 90 +
|
||||
frontend/tools/vendor_fonts.py | 137 +
|
||||
planning/reports/M1-BASELINE-REPORT.md | 466 +
|
||||
start.ps1 | 2 +-
|
||||
start.sh | 5 +-
|
||||
```
|
||||
|
||||
**Nothing inherited was deleted.** The 23 deletions are lines replaced in
|
||||
place, not features removed. Removal is M2's job.
|
||||
|
||||
### Tokenizer
|
||||
|
||||
| File | Origin | What changed | Why M1 needed it |
|
||||
| --- | --- | --- | --- |
|
||||
| `backend/app/context/vendor/cl100k_base.tiktoken` | **new** (vendored data) | The `cl100k_base` BPE table, 1.7 MB, SHA-256 `223921b7…65b2a7`. | The download this replaces is what killed the first story turn offline. |
|
||||
| `backend/app/context/encoding.py` | **new** | Builds a `tiktoken.Encoding` from the vendored table, verifying its SHA-256 against the digest `tiktoken` itself pins for the source URL. | Removes the network from the code path entirely, rather than relying on a warm cache or an env var. |
|
||||
| `backend/app/context/builder.py` | **inherited, modified** | `_encoding()` delegates to the new module; its `functools.lru_cache` moves there (one cache instead of two). | Single call site; every turn goes through it. |
|
||||
|
||||
### Fonts and browser assets
|
||||
|
||||
| File | Origin | What changed | Why M1 needed it |
|
||||
| --- | --- | --- | --- |
|
||||
| `frontend/public/fonts/*.woff2` | **new** (vendored data) | Cinzel, Crimson Pro, Inter as variable fonts, Latin + Latin-Ext, 343 KiB total. | The SPA fetched these from Google on every page load. |
|
||||
| `frontend/public/fonts/OFL-*.txt` | **new** | SIL OFL 1.1 licence text per family. | Required by the OFL for redistribution. |
|
||||
| `frontend/src/styles/fonts.css` | **new, generated** | `@font-face` declarations pointing at `/fonts/…`. | Replaces the Google stylesheet. |
|
||||
| `frontend/tools/vendor_fonts.py` | **new** | Regenerates both of the above from the Google Fonts API. | Keeps the vendored bytes reproducible instead of opaque. |
|
||||
| `frontend/index.html` | **inherited, modified** | Two `preconnect` hints and the Google stylesheet `<link>` removed, replaced by a comment pointing at the script. | The actual remote-asset request. |
|
||||
| `frontend/src/index.css` | **inherited, modified** | One `@import` for `fonts.css`, first in a load-bearing cascade order. | Faces must be declared before `tokens.css` names the families. |
|
||||
| `.gitattributes` | **inherited, modified** | `*.woff2`/`*.woff` marked binary. | Prevents line-ending normalisation corrupting a font. |
|
||||
|
||||
### Security headers and listener
|
||||
|
||||
| File | Origin | What changed | Why M1 needed it |
|
||||
| --- | --- | --- | --- |
|
||||
| `backend/app/main.py` | **inherited, modified** | CSP: `fonts.googleapis.com` and `fonts.gstatic.com` dropped, `font-src 'self'` added, plus `object-src 'none'`, `base-uri 'none'`, `form-action 'self'`. Separately, `mimetypes.add_type("font/woff2", …)` so the fonts are served with their real type instead of `application/octet-stream`. | M1's "tighten the CSP so runtime assets are local". The policy now names no remote origin at all. |
|
||||
| `start.sh`, `start.ps1` | **inherited, modified** | `--host 127.0.0.1` stated explicitly rather than inherited from uvicorn's default. | A02 is a requirement, not a default worth inheriting silently. |
|
||||
| `docker-compose.yml` | **inherited, modified** | Publishes `127.0.0.1:8000:8000` instead of `8000:8000`. | `8000:8000` publishes on every host interface — the storyteller on the LAN, unauthenticated. |
|
||||
| `Dockerfile` | **inherited, modified** | Comment only, explaining why the in-container listener is `0.0.0.0` and that the port must be published to loopback. | The 0.0.0.0 bind reads like a contradiction of A02 without it. |
|
||||
|
||||
### Outbound TLS *(unplanned; see §G and §J)*
|
||||
|
||||
| File | Origin | What changed | Why M1 needed it |
|
||||
| --- | --- | --- | --- |
|
||||
| `backend/app/tlstrust.py` | **new** | One cached `SSLContext` unioning the platform CA store with certifi's bundle. | Without it the trusted-LAN path fails against any host with a locally-issued certificate. |
|
||||
| `backend/app/providers/openai_compatible.py` | **inherited, modified** | Three `httpx.AsyncClient(…)` calls take `verify=tlstrust.ssl_context()`. | Turns, completions and embeddings. |
|
||||
| `backend/app/routers/settings.py` | **inherited, modified** | The connection-test client takes the same context. | Otherwise **Test connection** disagrees with what a turn would do. |
|
||||
| `backend/requirements.txt` | **inherited, modified** | `certifi` declared. | It is now imported by name rather than arriving via httpx. |
|
||||
|
||||
### Environment, tests, documentation
|
||||
|
||||
| File | Origin | What changed | Why M1 needed it |
|
||||
| --- | --- | --- | --- |
|
||||
| `backend/requirements.lock` | **new** | The exact tested closure, 40 pins including transitive and dev dependencies. | M1's "reproducible dev/test environment". `requirements.txt` keeps the ranges. |
|
||||
| `backend/tests/test_offline_assets.py` | **new** | 10 tests: vendored-table integrity, tokenizer opens no socket, golden token counts, CSP names no remote origin, no remote URL in markup/CSS, every declared font file exists, built SPA clean. | Both fixed bugs were invisible on a machine that had been online once. |
|
||||
| `backend/tests/test_tls_trust.py` | **new** | 6 tests: verification not weakened, context cached, certifi roots survive the union, and an AST walk asserting every `httpx.AsyncClient` passes `verify=` — including that no third module starts making requests. | Guards the §G change in both directions. |
|
||||
| `DEVELOPMENT.md` | **new** | Setup, run modes, same-host and trusted-LAN Ollama (including the private-CA case), test commands, the offline re-verification procedure, and what M1 deliberately left alone. | M1's environment/configuration documentation. |
|
||||
| `.gitignore` | **inherited, modified** | `/phase0b/` ignored. | Phase 0B research scratch — virtualenvs, databases, downloaded models. |
|
||||
|
||||
---
|
||||
|
||||
# D. User-Visible M1 Capability
|
||||
|
||||
Everything below was done through the running application. The browser step was
|
||||
confirmed by hand by the maintainer: the campaign list rendered at
|
||||
`http://127.0.0.1:8000`, a campaign was clicked, and the adventure and its
|
||||
transcript displayed.
|
||||
|
||||
### What a user can do now
|
||||
|
||||
1. **Start the application locally** — `./start.sh` for development, or a built
|
||||
SPA served by FastAPI on `127.0.0.1:8000` for production. Both bind loopback.
|
||||
2. **Configure a model** — Settings → endpoint URL, model, API mode, output and
|
||||
context budgets. **Test connection** lists the endpoint's models.
|
||||
3. **Create a campaign** and give it a persona.
|
||||
4. **Generate narration** — Do / Say / Story / Continue, streamed over SSE.
|
||||
5. **Persist and resume** — campaigns survive a clean restart with transcript
|
||||
and head position intact; the adventures list reopens them.
|
||||
6. **Use same-host Ollama** — `http://127.0.0.1:11434/v1`, the default endpoint,
|
||||
with no Internet at any point.
|
||||
7. **Use trusted-LAN Ollama** — an explicitly configured endpoint on another
|
||||
machine, `http://…:11434/v1` or `https://…/v1` with a private CA, while the
|
||||
storyteller UI/API stays on loopback.
|
||||
8. **Recover from a failed model call** — a clear error, accepted history
|
||||
untouched, and play continues.
|
||||
|
||||
The inherited surface also still works and is reachable from the nav: Home,
|
||||
Adventures, Scenarios (with the stat-schema and NPC editors), Scripts (the
|
||||
CodeMirror JavaScript editor), Settings, plus AI Chat and the Visitors
|
||||
dashboard, which local installs always see.
|
||||
|
||||
### First-run step worth knowing
|
||||
|
||||
`Settings.model` defaults to `""`, so **a user must pick a model before the
|
||||
first turn**; the endpoint already defaults to `http://localhost:11434/v1`.
|
||||
Nothing tells them this on the way in. Cosmetic in M1, a real onboarding
|
||||
question for M8.
|
||||
|
||||
### Inherited limitations still visible, deferred by design
|
||||
|
||||
| What the user sees | Milestone that addresses it |
|
||||
| --- | --- |
|
||||
| **Undo deletes turns and there is no Redo.** Verified in code, not assumed: `POST /adventures/{id}/undo` deletes the trailing AI action and its player action and prunes covering memories; no redo endpoint or control exists anywhere in the backend or SPA. | M3 |
|
||||
| Account, hosted and cloud-provider surfaces exist in the tree; the endpoint field accepts any URL. | M2 |
|
||||
| RPG world-state machinery — stats, flags, milestones, cast — with relative-delta proposals. | M5 |
|
||||
| JavaScript campaign scripting, sandboxed in QuickJS. | M2 removes it |
|
||||
| The Visitors analytics dashboard (local counters, two SQLite tables, no outbound request). | M2 |
|
||||
| Memory bank and auto-summarisation are **off per adventure by default**, even when an embedding model is configured globally. | M6 |
|
||||
| No named Save Points; no imported knowledge; UI is still AI-DnD's. | M4, M7, M8 |
|
||||
|
||||
---
|
||||
|
||||
# E. Acceptance-Test Results
|
||||
|
||||
Every row is runtime behaviour observed in a running system. Where a check was
|
||||
source-level, the row says so and does not claim PASS on that basis.
|
||||
|
||||
| ID | Result | Procedure | Evidence | Caveat |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| **A01** Start application offline | **PASS** | Run 1 (§F). `--internal` Docker network; app started; campaign created; six turns. | Isolation proven first — `1.1.1.1:443` → `Network is unreachable`, every name → `gaierror`. Six turns generated and persisted. Whole browser asset graph fetched over loopback, all 200. | The UI render was confirmed in a browser on the *native* run, which had Internet (§F). |
|
||||
| **A02** Storyteller loopback default | **PASS** | Listener enumerated from `/proc/net/tcp` in each container; `ss -ltnp` on the native run; a TCP connect to this host's LAN address. | Runs 1, 2, 3 all show `LISTEN 127.0.0.1:8000` and nothing else. Native run: `connect 192.168.0.10:8000 → Connection refused`. `docker-compose.yml` publishes `127.0.0.1:8000:8000`. | In Docker the *in-container* listener is `0.0.0.0`; loopback-only exposure comes from the published port. A hand-run `docker run -p 8000:8000` would defeat it — §M. |
|
||||
| **A03** No cloud API key | **PASS** | `api_key: ""` in every run; connection test, turns, summaries and embeddings all exercised. | `"api_key_set": false` in each run's settings dump; all operations succeeded. | — |
|
||||
| **A04** Campaign survives restart | **PASS** | Run 1: 6 turns → `docker restart` → re-read. Run 2: same, plus a further turn. | Run 1: 12 actions, head 27, identical transcript before and after. Run 2: digest `f24caf86a744ab36` identical across restart; endpoint, embedding model and memories preserved; next turn produced in 6.4 s. | — |
|
||||
| **A05** Failed model call does not corrupt story | **PASS** | Two induced failures (nonexistent model on a live endpoint; dead endpoint port) plus one natural timeout, then recovery. | Accepted-prefix digest `2ca6ab528178e44e` unchanged throughout; AI-action count stayed at 6 across both failures; recovery via `continue` produced turn 7 with the prefix still unchanged. | The player's own typed action **is** committed before the model call, so the row count grows. §H. |
|
||||
| **A06** Trusted-LAN Ollama inference | **PASS** | Run 2 (§F). Ollama 0.33.0 on a second physical machine over HTTPS; storyteller's default route deleted, LAN-only route added, resolver pointed at nothing, host supplied as a static hosts entry. | Model discovery returned both models; nine turns (4–14 s); memories embedded on the remote host; restart and resume; capture shows 893 packets to the approved host, 730 loopback, **0** elsewhere, **0** DNS queries. | Required the §G TLS fix first — the first attempt failed with `CERTIFICATE_VERIFY_FAILED`. |
|
||||
| **H01** No unexpected outbound connections | **PASS for the application** | `tcpdump -i any` inside each run's network namespace, for the whole run. | Run 1: 6 074 packets, 6 062 loopback, **0 non-loopback unicast**. Run 2: 1 633 packets, 730 loopback, 893 to the approved host, **0** other unicast, **0** DNS. | Ollama itself queried `ollama.com` in Run 1 — not the storyteller, and it failed. §K. |
|
||||
| **H02** No telemetry | **PASS** | The same captures, plus reading `analytics.py`. | No outbound destination in either capture. `analytics.py` writes two local SQLite tables, opens no socket, records no IP or user agent, and HMACs the user id. | The dashboard still exists in the UI. M2 removes it. |
|
||||
| **H03** No cloud provider required | **PASS as written** | Runs 1 and 2, with nothing cloud reachable. | Every operation succeeded with no cloud endpoint and no key. | The test's *preferred* final state — "controls are absent, not merely unused" — is **not** met. That is M2's scope by design, not an M1 gap. |
|
||||
| **H11** No first-use runtime asset download | **PASS** | Run 1 on a **fresh database and fresh container** with no route out: first turn generated, then the full asset graph fetched. | First turn succeeded where upstream raised `ConnectionError`. `index.html` references only same-origin URLs; all 8 woff2 files served locally as `font/woff2`; CSP names no remote origin. Independently: with sockets blocked, upstream's code path raises `AssertionError: socket opened` and the vendored path returns a token count. | Browser devtools were not inspected; the claim rests on the server-side asset graph and the CSP. |
|
||||
|
||||
**Not tested, and not claimed:** any acceptance test outside M1's scope (B, C,
|
||||
D, E, F, G, I, J, K series). No result above is inferred from source
|
||||
inspection alone.
|
||||
|
||||
---
|
||||
|
||||
# F. Offline and Network Evidence
|
||||
|
||||
### Run 1 — same-host Ollama, air-gapped
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Topology | Ollama in one container; the production image in a second container **sharing Ollama's network namespace**, so Ollama is genuinely on the storyteller's loopback. Network is Docker `--internal`: no NAT, no external DNS. |
|
||||
| Storyteller listener | `127.0.0.1:8000` |
|
||||
| Ollama endpoint | `http://127.0.0.1:11434/v1` — same-host loopback |
|
||||
| Outbound Internet actually blocked? | **Yes**, proven before testing: `1.1.1.1:443` → `OSError: Network is unreachable`; `openaipublic.blob.core.windows.net`, `fonts.googleapis.com`, `fonts.gstatic.com`, `openrouter.ai`, `github.com` → `gaierror` |
|
||||
| Expected destinations observed | `127.0.0.1:8000` (API), `127.0.0.1:11434` (Ollama), `127.0.0.1:11499` (the deliberately dead port in A05), two ephemeral loopback ports (Ollama's model runner) |
|
||||
| Unexpected attempts | **None from the application.** 6 062 of 6 074 packets loopback; **0 non-loopback unicast**; the remaining 12 are received mDNS/ICMPv6 multicast from the bridge. |
|
||||
| DNS | Four queries, all `ollama.com`, all `ServFail` — issued by **the Ollama server**, which shares the namespace. Not the storyteller. §K. |
|
||||
| First-turn asset download | **None.** Fresh database, fresh container; the first turn narrated instead of raising `ConnectionError`. No font, tokenizer or script request appears in the capture. |
|
||||
|
||||
### Run 2 — trusted-LAN Ollama, Internet blocked
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Topology | Ollama 0.33.0 on `inference.lan` (`192.168.0.50`), a **separate physical machine** on the trusted LAN, HTTPS, certificate from a local StartOS CA. Storyteller in a container on this host with `NET_ADMIN`: default route **deleted**, replaced by a route to `192.168.0.0/24` only; resolver pointed at nothing; the host supplied as a static hosts entry. |
|
||||
| Storyteller listener | `127.0.0.1:8000` — nothing on the container's own LAN-facing address |
|
||||
| Ollama endpoint | `https://inference.lan:8443/v1` — explicitly configured, non-loopback, TLS |
|
||||
| Outbound Internet actually blocked? | **Yes.** By IP: `1.1.1.1`, `140.82.121.4`, `104.16.0.1` → `Network is unreachable`. By name: `github.com`, `openrouter.ai`, `fonts.gstatic.com`, `openaipublic.blob.core.windows.net` → `gaierror`. The LAN host resolved and TLS-verified: `peer CN = inference.lan`. |
|
||||
| Expected destinations observed | `127.0.0.1:8000` (17 connections), `192.168.0.50:8443` (19 connections) |
|
||||
| Unexpected attempts | **None.** 1 633 packets: 730 loopback, 893 to the approved host, **0 other unicast**. |
|
||||
| DNS | **Zero queries of any kind** — the endpoint was configured, not resolved. |
|
||||
| First-turn asset download | **None.** No name resolves at all, and the turn succeeded. |
|
||||
|
||||
Confirmed by the application's own request log: every model request went to
|
||||
`…/v1/chat/completions` and `…/v1/embeddings` on the configured host, and no
|
||||
other URL — narrator, summariser and embedder alike.
|
||||
|
||||
### Run 3 — native, non-Docker
|
||||
|
||||
Production build from the venv, SPA served by FastAPI, Ollama on host loopback.
|
||||
`LISTEN 127.0.0.1:8000` and `127.0.0.1:11434`; the LAN address refuses
|
||||
connections. A campaign was created and played, and this is the instance the
|
||||
browser check used. **This run had Internet available** — its purpose was the
|
||||
native listener check and a real out-of-Docker run, not the offline proof.
|
||||
|
||||
---
|
||||
|
||||
# G. tiktoken, Fonts, CSP, and Runtime Asset Fixes
|
||||
|
||||
### The tokenizer download
|
||||
|
||||
**Cause.** `backend/app/context/builder.py` called
|
||||
`tiktoken.get_encoding("cl100k_base")`. That fetches the BPE table from
|
||||
`openaipublic.blob.core.windows.net` on first use and caches it under the
|
||||
system temp directory. `count_tokens` runs on **every** turn, for context
|
||||
budgeting. On a developer machine that had been online once the cache was warm
|
||||
and the download invisible; on an air-gapped install the first turn died with
|
||||
`ConnectionError` instead of narrating.
|
||||
|
||||
**Fix.** The table is vendored at
|
||||
`backend/app/context/vendor/cl100k_base.tiktoken`, and
|
||||
`backend/app/context/encoding.py` constructs the `Encoding` directly from it —
|
||||
the same merge table, pattern string and special tokens `tiktoken` uses.
|
||||
Nothing in the tokenizer path can reach the network: not a cache that happens
|
||||
to be warm, not an environment variable a deployment could forget.
|
||||
|
||||
Three things make this trustworthy rather than merely working:
|
||||
|
||||
- The file's SHA-256 is `223921b76ee99bde995b7ff738513eef100fb51d18c93597a113bcffe865b2a7`,
|
||||
**identical to the digest `tiktoken_ext/openai_public.py` pins for that URL**,
|
||||
and it is re-checked every time the encoding is built. A truncated checkout
|
||||
or a substituted table fails loudly instead of silently changing every token
|
||||
count the context budget derives from.
|
||||
- A test asserts that digest still appears in `tiktoken`'s own source, so a
|
||||
future upgrade pointing `cl100k_base` at a different table is caught.
|
||||
- The encoding was compared token-for-token against `tiktoken.get_encoding`
|
||||
across ASCII, accented text, CJK, emoji, CRLF and special-token literals.
|
||||
|
||||
**Proof the guard is real**, with `TIKTOKEN_CACHE_DIR` pointed at an empty
|
||||
directory and Python's socket functions replaced:
|
||||
|
||||
```text
|
||||
UPSTREAM PATH raises: AssertionError socket opened
|
||||
VENDORED PATH: 2 tokens, no socket opened
|
||||
```
|
||||
|
||||
### Remote fonts
|
||||
|
||||
**Behaviour.** `frontend/index.html` carried two `preconnect` hints and a
|
||||
stylesheet `<link>` to `fonts.googleapis.com` for Cinzel, Crimson Pro and
|
||||
Inter. Every page load fetched that stylesheet and then font files from
|
||||
`fonts.gstatic.com` — an Internet dependency at runtime, and a third party
|
||||
learning when the story is being read.
|
||||
|
||||
**Fix.** Self-hosted. `frontend/tools/vendor_fonts.py` downloads the same faces
|
||||
once at development time into `frontend/public/fonts/` and generates
|
||||
`frontend/src/styles/fonts.css`. Variable fonts and the Latin + Latin-Ext
|
||||
subsets: 8 files, 343 KiB, covering every weight the design uses. OFL text
|
||||
ships beside them. Greek, Cyrillic and Vietnamese subsets are deliberately not
|
||||
vendored; text in them falls back to the system stack.
|
||||
|
||||
### CSP
|
||||
|
||||
```diff
|
||||
- style-src 'self' 'unsafe-inline' https://fonts.googleapis.com;
|
||||
- font-src https://fonts.gstatic.com;
|
||||
+ style-src 'self' 'unsafe-inline';
|
||||
+ font-src 'self';
|
||||
+ object-src 'none'; base-uri 'none'; form-action 'self';
|
||||
```
|
||||
|
||||
Final policy, as served:
|
||||
|
||||
```text
|
||||
default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline';
|
||||
font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none';
|
||||
base-uri 'none'; form-action 'self'; frame-ancestors 'none'
|
||||
```
|
||||
|
||||
No remote origin remains. `'unsafe-inline'` stays on `style-src` because React
|
||||
writes inline `style` attributes; it is deliberately absent from `script-src`.
|
||||
|
||||
Separately, `woff2` was being served as `application/octet-stream` because
|
||||
Python's mimetypes table has no entry for it on a slim Debian image. Browsers
|
||||
accept it anyway — a `@font-face src` carries its own `format()` hint — but
|
||||
`main.py` now registers the correct type.
|
||||
|
||||
### Outbound TLS *(unplanned — see §J)*
|
||||
|
||||
**Cause.** `httpx` verifies against the `certifi` bundle, which carries the
|
||||
public web's CAs and nothing else. A trusted-LAN Ollama frequently has no
|
||||
public certificate. Against a real StartOS-hosted Ollama the connection test
|
||||
returned:
|
||||
|
||||
```text
|
||||
{"ok": false, "detail": "Connection failed: [SSL: CERTIFICATE_VERIFY_FAILED]
|
||||
certificate verify failed: self-signed certificate in certificate chain"}
|
||||
```
|
||||
|
||||
while `curl` and the browser on the same machine accepted the identical
|
||||
endpoint, because the CA was installed in the **system** store. Measured:
|
||||
|
||||
```text
|
||||
certifi bundle (httpx default) FAIL SSLCertVerificationError
|
||||
system trust store OK peer CN=inference.lan
|
||||
```
|
||||
|
||||
**Fix.** `backend/app/tlstrust.py` builds one cached context that **unions**
|
||||
the platform CA store with certifi's bundle; all four outbound clients use it.
|
||||
A union rather than a swap on purpose: the platform store alone would be a
|
||||
behaviour *change*, and an image with an empty or stale system store would
|
||||
start failing on endpoints that previously worked. A union can only add trust
|
||||
the user already granted at the OS level.
|
||||
|
||||
Verification is untouched — `verify_mode=CERT_REQUIRED`, `check_hostname=True`,
|
||||
and **no "insecure" escape hatch was added**. After the change:
|
||||
|
||||
```text
|
||||
OK inference.lan (local CA) CN=inference.lan
|
||||
OK github.com (public CA) CN=github.com
|
||||
OK pypi.org (public CA) CN=pypi.org
|
||||
certifi-only context still rejects inference.lan (so the union is what changed)
|
||||
```
|
||||
|
||||
### Fresh-data offline first-turn test
|
||||
|
||||
**Passed.** Run 1 used a fresh Docker volume and a fresh container on a network
|
||||
with no route out and no external DNS. The first turn of a newly created
|
||||
campaign generated and persisted. No tokenizer, font or other runtime asset
|
||||
request appears in the 6 074-packet capture.
|
||||
|
||||
---
|
||||
|
||||
# H. Persistence and Failure Recovery
|
||||
|
||||
### Clean restart
|
||||
|
||||
| Run | Before | After |
|
||||
| --- | --- | --- |
|
||||
| 1 (same-host) | 12 actions, head 27 | 12 actions, head 27, transcript identical line for line |
|
||||
| 2 (trusted-LAN) | 18 actions, head 18, digest `f24caf86a744ab36` | 18 actions, head 18, digest `f24caf86a744ab36` |
|
||||
|
||||
Run 2 additionally confirmed the endpoint (`https://inference.lan:8443/v1`),
|
||||
the embedding model and all memories survived, and produced a further turn in
|
||||
6.4 s afterwards.
|
||||
|
||||
### Failed model request
|
||||
|
||||
Baseline: 12 actions, 6 of them AI, accepted digest `2ca6ab528178e44e`.
|
||||
|
||||
```text
|
||||
invalid local model -> events ['player','error']
|
||||
"Endpoint or model not found (HTTP 404) … model
|
||||
'no-such-model-v9' not found"
|
||||
13 actions, still 6 AI actions
|
||||
endpoint down (:11499) -> events ['player','error']
|
||||
"Could not connect to http://127.0.0.1:11499/v1 —
|
||||
is the AI server running?"
|
||||
14 actions, still 6 AI actions
|
||||
accepted prefix through the pre-failure head:
|
||||
12 actions, digest 2ca6ab528178e44e — unchanged
|
||||
rows added by the failures:
|
||||
[28] story 'I strike a match.'
|
||||
[29] story 'I strike a match again.'
|
||||
recovery (continue) -> 40 SSE events, done; 15 actions, 7 AI actions
|
||||
prefix digest still 2ca6ab528178e44e
|
||||
```
|
||||
|
||||
**No partially accepted turn, in either failure.** The AI-action count never
|
||||
moved; the accepted prefix is bit-identical before, during and after.
|
||||
|
||||
**One behaviour the reviewer should know.** The player's own typed action is
|
||||
committed *before* the model is called (`run_player_turn`), so a failed turn
|
||||
leaves the player's text at the head with no reply, and the row count grows.
|
||||
That satisfies A05 as written — prior story intact, failed turn not committed
|
||||
as accepted, retry available — and it is deliberate: a model failure never eats
|
||||
what the player typed. It is stated because "the story is unchanged" is not
|
||||
literally true; "the accepted story is unchanged" is. Whether a dangling player
|
||||
action is the right *user-facing* recovery state is an M3 question.
|
||||
|
||||
---
|
||||
|
||||
# I. Test Suite / Regression Baseline
|
||||
|
||||
| Check | Command | Result |
|
||||
| --- | --- | --- |
|
||||
| Backend, before M1 | `python -m pytest tests/ -q` on the pinned upstream commit | **632 passed, 0 failed**, 184.7 s |
|
||||
| Backend, after M1 | `python -m pytest tests/ -q` | **648 passed, 0 failed, 0 skipped**, 212.1 s |
|
||||
| Backend, **with no Internet** | same suite in a container on an `--internal` network | **648 passed, 0 failed**, 253.6 s |
|
||||
| Frontend lint | `npm run lint` (oxlint) | **exit 0**, 6 warnings, 0 errors |
|
||||
| Frontend build | `npm run build` (vite 8.1.3) | **succeeds**, 1.13 s |
|
||||
| Image | `docker build .` | **succeeds** |
|
||||
| Typecheck | — | none exists; the SPA is plain JSX with no TypeScript config. |
|
||||
|
||||
The 632-test baseline was captured on the pinned upstream commit *before* any
|
||||
change and matches the Phase 0B figure independently.
|
||||
|
||||
### New M1 tests — 16
|
||||
|
||||
**`backend/tests/test_offline_assets.py` (10)** — vendored table present and
|
||||
intact; the pinned digest still matches `tiktoken`'s own; token counting opens
|
||||
no socket (sockets monkeypatched to raise); golden token counts; round-trip
|
||||
over awkward characters; CSP names no remote origin; `index.html` fetches
|
||||
nothing remote; stylesheets fetch nothing remote; every declared font file
|
||||
exists; the built SPA is clean.
|
||||
|
||||
**`backend/tests/test_tls_trust.py` (6)** — verification not weakened; context
|
||||
cached; every certifi root survives the union; and an **AST walk** over the two
|
||||
modules that make outbound requests asserting each `httpx.AsyncClient` passes
|
||||
`verify=`, plus that no third module has started making requests. That last one
|
||||
catches the failure no runtime test would: a *new* client added later, which
|
||||
would work perfectly until someone pointed it at a LAN endpoint.
|
||||
|
||||
### Remaining failures
|
||||
|
||||
**None.** No failing test, no skipped test, no documented exception, and **no
|
||||
M1 regression**. Nothing was inherited red — the 632-test baseline was green on
|
||||
the first attempt on the pinned commit.
|
||||
|
||||
Two `dist`-reading tests in `test_offline_assets.py` skip when
|
||||
`frontend/dist/` has not been built. In the runs above it had been, so they
|
||||
executed and are counted in the 648.
|
||||
|
||||
### Internet requirement
|
||||
|
||||
**No test requires Internet access** — confirmed by running the entire suite in
|
||||
a container with no route out and no DNS. All 648 passed. Two pre-existing
|
||||
warnings persist: a Starlette `httpx` deprecation notice, and a `SyntaxWarning`
|
||||
for an invalid escape sequence in `tools/rewrite_memories.py`, both inherited
|
||||
and both untouched by M1.
|
||||
|
||||
---
|
||||
|
||||
# J. Deviations From the M1 Plan
|
||||
|
||||
### Added scope
|
||||
|
||||
1. **Outbound TLS trust (`backend/app/tlstrust.py`) — the significant one.**
|
||||
Not in M1's scope list. `BUILD-MILESTONES.md` places "explicit Ollama
|
||||
endpoint policy" in M2. But M1's own scope requires *verifying* trusted-LAN
|
||||
generation, and its Definition of Done requires turns through such a host.
|
||||
Against a realistic LAN Ollama — TLS with a locally-issued certificate — that
|
||||
was impossible without this change. The choice was to change it or to report
|
||||
A06 as blocked. It is small (one module, four call sites), does not weaken
|
||||
verification, and adds no configuration surface.
|
||||
|
||||
2. **`woff2` media type** (one line in `main.py`). Found while verifying the
|
||||
self-hosted fonts. Cosmetic, but wrong is wrong.
|
||||
|
||||
3. **`backend/requirements.lock`.** M1 says "reproducible dev/test
|
||||
environment"; the inherited `requirements.txt` uses `>=` throughout, so a
|
||||
fresh checkout resolved to whatever was newest that day. The lock is the
|
||||
concrete deliverable for that line.
|
||||
|
||||
4. **`base-uri`, `object-src`, `form-action` in the CSP.** M1 says "tighten the
|
||||
CSP as needed". Removing the two Google hosts was the requirement; these
|
||||
three cost nothing and were added while the policy was open.
|
||||
|
||||
5. **`/phase0b/` in `.gitignore`** — housekeeping, so `git status` is readable.
|
||||
|
||||
### Omitted scope
|
||||
|
||||
**None.** Every item in M1's scope list was implemented and demonstrated.
|
||||
|
||||
### Changed assumptions
|
||||
|
||||
1. **"Trusted-LAN" was assumed to mean plain HTTP.** Every planning reference
|
||||
writes `http://…:11434/v1`, and Ollama's own default is cleartext. The real
|
||||
LAN host serves **HTTPS only** (plain HTTP 307-redirects), because it is a
|
||||
StartOS server. This is what surfaced the TLS defect. §L.
|
||||
|
||||
2. **A06 was initially reproduced without physical separation.** Before a
|
||||
second machine was available, Run 2 used a separate container, namespace and
|
||||
IP, and was going to be reported as PASS-with-qualification. A real second
|
||||
machine then became available and the run was redone properly — which is
|
||||
what exposed the TLS defect the container stand-in had hidden, since a
|
||||
container endpoint was plain HTTP.
|
||||
|
||||
### Workarounds
|
||||
|
||||
1. **Blocking Internet without root.** No passwordless sudo here, so host
|
||||
firewall rules were unavailable. Run 1 used a Docker `--internal` network.
|
||||
Run 2 used `NET_ADMIN` inside the container to delete its default route and
|
||||
add a LAN-only route — stronger than a firewall rule, since there is no
|
||||
route to drop packets on.
|
||||
|
||||
2. **Packet capture without root.** `tcpdump` in a container joined to the
|
||||
target namespace with `NET_RAW`.
|
||||
|
||||
3. **Browser verification by hand.** No browser automation in the session; the
|
||||
maintainer confirmed the render directly. Devtools were not inspected. §K.
|
||||
|
||||
### Unexpected inherited behaviour
|
||||
|
||||
Four, all in §K: the model timeout, Ollama's own DNS lookup, the per-adventure
|
||||
memory-bank defaults, and the pre-model commit of the player's action.
|
||||
|
||||
---
|
||||
|
||||
# K. Technical Findings / Surprises
|
||||
|
||||
### Affecting local-only security
|
||||
|
||||
1. **Ollama phones home.** In Run 1 the capture shows four DNS queries for
|
||||
`ollama.com`, issued by the **Ollama server**, not the storyteller. They
|
||||
failed and nothing depended on them. But it sits inside the user's trust
|
||||
boundary and outside this application's code, and a user reading a packet
|
||||
capture will see it. The local-only claim covers what *this* application
|
||||
sends; that distinction should be stated in the release material rather than
|
||||
discovered. Suppressing it is an Ollama configuration question, deliberately
|
||||
not investigated here.
|
||||
|
||||
2. **Loopback in Docker depends on the publish flag, not the bind.** The
|
||||
in-container listener is `0.0.0.0`, which is the only address a published
|
||||
port can reach. `docker-compose.yml` publishes `127.0.0.1:8000:8000`, so the
|
||||
shipped path is safe — but a hand-run `docker run -p 8000:8000` puts an
|
||||
unauthenticated storyteller on the LAN. Comments now say so in both files.
|
||||
Whether the app should *refuse* to serve on a non-loopback bind without an
|
||||
explicit opt-in is an M2 policy question.
|
||||
|
||||
### Affecting trusted-LAN inference
|
||||
|
||||
3. **The TLS defect (§G) is the most important finding in M1.** Beyond the fix,
|
||||
the lesson is methodological: it was undetectable by static review and
|
||||
undetectable by every local run that used loopback HTTP, because
|
||||
certificate verification never happens there. It surfaced within minutes of
|
||||
pointing the application at a real host. **M2's endpoint-policy work should
|
||||
be validated against a real LAN host, not a container stand-in.**
|
||||
|
||||
4. **A CA in a container is not the CA on the host.** Running against a
|
||||
private-CA endpoint from Docker requires the CA inside the image or
|
||||
bind-mounted. Documented in `DEVELOPMENT.md`; it will matter for any future
|
||||
packaged distribution.
|
||||
|
||||
5. **The model timeout is 120 s and hardcoded** (`httpx.Timeout(120,
|
||||
connect=10)`). On this GPU-less four-core host a *cold* model load, or two
|
||||
models contending after the memory bank pulls in the embedder, exceeded it
|
||||
three times. Once warm, a full turn took 6–9 s. Presented as a tuning
|
||||
finding rather than a defect — but a first-run user on modest hardware will
|
||||
likely meet it, and a configurable timeout is a small change.
|
||||
|
||||
### Affecting dependency packaging
|
||||
|
||||
6. **The vendored tokenizer table is 100 256 lines / 1.7 MB.** It dominates the
|
||||
M1 diff. The alternative — `TIKTOKEN_CACHE_DIR` — was rejected as a
|
||||
deployment-time promise that a packaging step could forget. The digest check
|
||||
makes the vendored copy auditable.
|
||||
|
||||
7. **`certifi` is now a direct dependency**, declared in `requirements.txt`
|
||||
rather than arriving via httpx.
|
||||
|
||||
8. **Node version.** The Dockerfile builds on Node 24 with a comment saying
|
||||
npm 10 refuses the lockfile; `npm ci` in fact succeeded on Node 22.23.1 /
|
||||
npm 10.9.8. No `.nvmrc` was added, because asserting a constraint that was
|
||||
not tested would be worse than documenting what was.
|
||||
|
||||
### Affecting browser/runtime assets
|
||||
|
||||
9. **`docs/*.html` still links Google Fonts.** Upstream's GitHub Pages project
|
||||
site — not served by the application, not part of any build, not covered by
|
||||
the runtime rule. Deliberately untouched; the regression tests scope
|
||||
themselves to `frontend/` so they do not give a false signal about it.
|
||||
|
||||
10. **Only devtools were left unchecked.** The page's whole asset graph was
|
||||
fetched and verified same-origin with no route out, the source and built
|
||||
bundle contain no remote reference, and the CSP would block one. A devtools
|
||||
capture during an offline session is the one more direct form of this
|
||||
evidence and costs about a minute.
|
||||
|
||||
### Affecting model-provider configuration
|
||||
|
||||
11. **`Settings.model` defaults to `""`.** A new install cannot generate a turn
|
||||
until a model is chosen, and nothing prompts for it (§D).
|
||||
|
||||
12. **`netguard` is inert locally by design.** It rejects non-public addresses
|
||||
only when `AIDND_MULTI_USER=1`, which local installs never set — which is
|
||||
why a LAN endpoint is accepted at all. M2's endpoint policy replaces this,
|
||||
and should keep that asymmetry deliberate rather than inherit it silently.
|
||||
|
||||
### Affecting repository structure and testability
|
||||
|
||||
13. **The fork merge preserves all 172 upstream commits**, so upstream fixes
|
||||
can still be cherry-picked. Worth protecting: a future squash or filter
|
||||
would throw it away.
|
||||
|
||||
14. **The AST-based test in `test_tls_trust.py`** guards a class of regression
|
||||
no runtime test can reach — a new HTTP client added later. The same shape
|
||||
may be worth reusing in M2 when the provider surface is cut down.
|
||||
|
||||
15. **The suite is fully offline-capable** (§I), so CI needs no network beyond
|
||||
dependency install.
|
||||
|
||||
### Affecting future removal of hosted/cloud/scripting code
|
||||
|
||||
16. **M1 removed nothing**, so M2's removal surface is exactly what Phase 0B
|
||||
described. Two M1 changes touch files M2 will edit heavily —
|
||||
`providers/openai_compatible.py` and `routers/settings.py` — but both are
|
||||
one-line `verify=` additions, so the TLS work should survive that
|
||||
refactoring as long as the shared context follows any new client.
|
||||
|
||||
---
|
||||
|
||||
# L. Planning Documents That May Need Revision
|
||||
|
||||
Reported, not edited. No planning architecture document was changed by M1.
|
||||
|
||||
### 1. `planning/DECISIONS/002-ollama-only-v1.md` — Consequences
|
||||
|
||||
**Discrepancy.** The ADR treats a trusted-LAN Ollama as an addressing question
|
||||
only. It does not anticipate that such a host may be reachable **only over
|
||||
TLS**, with a certificate from a private CA. The real test host is exactly
|
||||
that, and the application could not talk to it until M1 changed how outbound
|
||||
TLS is verified.
|
||||
|
||||
**Recommended correction.** Add a consequence: the storyteller must verify
|
||||
against the operating system's CA store as well as the bundled one, so a user
|
||||
who has installed their own CA is honoured; and no option to skip verification
|
||||
should be offered.
|
||||
|
||||
### 2. `planning/DECISIONS/004-local-only-production.md` — Consequences
|
||||
|
||||
**Discrepancy.** "Production packaging must contain all runtime assets required
|
||||
for ordinary story use" is right but understates it. Both M1 findings were
|
||||
*first-use* downloads invisible on any machine that had been online once, and
|
||||
neither was findable by static analysis.
|
||||
|
||||
**Recommended correction.** Add that offline claims must be validated on a
|
||||
network with no route out and a fresh cache, and that vendored runtime assets
|
||||
should carry a verifiable digest.
|
||||
|
||||
### 3. `planning/V1-ACCEPTANCE-TESTS.md` — A06
|
||||
|
||||
**Discrepancy.** A06's preconditions do not say whether the LAN endpoint is
|
||||
HTTP or HTTPS, and every example elsewhere shows `http://`. The HTTPS
|
||||
private-CA case is the one that broke, and A06 as written could be passed
|
||||
against a plain-HTTP host without ever exercising it.
|
||||
|
||||
**Recommended correction.** Add a step covering an HTTPS endpoint with a
|
||||
locally-issued certificate, and a pass condition that verification is performed
|
||||
rather than bypassed.
|
||||
|
||||
### 4. `planning/V1-ACCEPTANCE-TESTS.md` — A05
|
||||
|
||||
**Discrepancy.** "Failed turn is not partially committed as accepted" is
|
||||
satisfied, but the observable state is a dangling player action at the head
|
||||
with no reply (§H). A reader could reasonably expect the row count not to move.
|
||||
|
||||
**Recommended correction.** State that the player's own input is retained by
|
||||
design, and that the invariant is about *accepted* history.
|
||||
|
||||
### 5. `planning/BUILD-MILESTONES.md` — M1 Scope
|
||||
|
||||
**Discrepancy.** M1's scope lists four hardening items but not outbound TLS
|
||||
trust, which turned out to be on its critical path.
|
||||
|
||||
**Recommended correction.** Note retrospectively that M1 also covered it, so
|
||||
M2's "explicit Ollama endpoint policy" is not planned as if it were still open.
|
||||
|
||||
### 6. `planning/V1-ACCEPTANCE-TESTS.md` §3 — Standard Test Environment
|
||||
|
||||
**Discrepancy.** The environment list assumes one local Ollama and does not
|
||||
mention hardware class. Model cold-load time on a GPU-less host exceeded the
|
||||
application's fixed 120 s timeout three times during M1.
|
||||
|
||||
**Recommended correction.** Record hardware class alongside model name, and
|
||||
note that a cold load may exceed the client timeout on modest hardware.
|
||||
|
||||
---
|
||||
|
||||
# M. Risks / Technical Debt Carried Into M2
|
||||
|
||||
### Blockers before M2
|
||||
|
||||
**None.**
|
||||
|
||||
### Acceptable technical debt
|
||||
|
||||
| Item | Risk | Suggested milestone |
|
||||
| --- | --- | --- |
|
||||
| **120 s hardcoded model timeout.** Cold loads on modest hardware exceed it; the user sees "The AI endpoint timed out." | Medium — a first-run user on a slow box may conclude the app is broken. | M2, with the endpoint policy |
|
||||
| **Ollama's own `ollama.com` lookup.** Outside this codebase, inside the user's trust boundary. | Low technically, medium for the local-only claim. | Decide before release; document either way |
|
||||
| **Docker loopback depends on the publish flag.** A hand-run `-p 8000:8000` exposes an unauthenticated API to the LAN. | Medium. | M2 |
|
||||
| **`Settings.model` defaults to `""`.** No first turn until a model is chosen; nothing prompts. | Low. | M8 |
|
||||
| **Memory bank off per adventure by default** even with an embedding model configured. Cost me a wasted verification cycle. | Low, but it makes M6's features look absent. | M6 |
|
||||
| **Vendored tokenizer table, 1.7 MB.** Needs re-vendoring if the encoding ever changes. | Low — digest-checked, and a test catches a `tiktoken` change. | — |
|
||||
| **No frontend typecheck.** Plain JSX, lint only. | Low. | M8 |
|
||||
| **`docs/*.html` links Google Fonts.** Not served by the app. | Very low. | M2 or M8 |
|
||||
| **The `dist` tests skip when the SPA is unbuilt.** A green suite alone is not full evidence. | Low — documented in `DEVELOPMENT.md`. | — |
|
||||
|
||||
### Inherited AI-DnD behaviour intentionally left in place
|
||||
|
||||
All of this was outside M1's scope and is **explicitly non-scope** in the M1
|
||||
brief. It is listed so nobody reports it as an M1 gap:
|
||||
|
||||
- multi-user/account/guest/auth flows, demo-key behaviour, hosted rate limits;
|
||||
- the self-hosted analytics tables and Visitors dashboard (local-only, no
|
||||
outbound request);
|
||||
- Render and Neon deployment paths; Postgres/`psycopg` support;
|
||||
- the OpenRouter default-endpoint constant and attribution-host check;
|
||||
- arbitrary remote model-provider configuration in the Settings UI;
|
||||
- QuickJS campaign scripting and the AI-Dungeon-compatible script surface;
|
||||
- **destructive Undo with no Redo** — verified in code, and the single largest
|
||||
inherited correctness gap. M3;
|
||||
- RPG world-state machinery and relative-delta proposals. M5;
|
||||
- Story Cards as the only imported-knowledge mechanism. M7;
|
||||
- AI-DnD's UI organisation and vocabulary. M8.
|
||||
|
||||
---
|
||||
|
||||
# N. Reproduction / Verification Commands
|
||||
|
||||
### Environment
|
||||
|
||||
```bash
|
||||
git clone <this repo> && cd interactive-story
|
||||
git checkout m1-production-baseline
|
||||
python3 -m venv backend/.venv
|
||||
backend/.venv/bin/pip install -r backend/requirements.lock
|
||||
(cd frontend && npm ci && npm run build)
|
||||
```
|
||||
|
||||
### Provenance
|
||||
|
||||
```bash
|
||||
git remote add upstream https://github.com/parththakkar106/AI-DnD.git
|
||||
git fetch --no-tags upstream
|
||||
git cat-file -t d72f7c1bda0f34fccd84afb7a25c34eb01c901de # -> commit
|
||||
git merge-base --is-ancestor d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD
|
||||
git diff --stat d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD -- LICENSE # -> empty
|
||||
```
|
||||
|
||||
### Tests
|
||||
|
||||
```bash
|
||||
(cd backend && .venv/bin/python -m pytest tests/ -q) # 648 passed
|
||||
(cd frontend && npm run lint && npm run build)
|
||||
docker build -t storyteller .
|
||||
```
|
||||
|
||||
Prove the suite needs no network:
|
||||
|
||||
```bash
|
||||
docker network create --internal offline
|
||||
docker build -t storyteller . && \
|
||||
printf 'FROM storyteller\nRUN pip install --no-cache-dir pytest\n' | docker build -q -t storyteller-test -
|
||||
docker run --rm --network offline -v "$PWD":/src:ro -w /src/backend \
|
||||
-e AIDND_DB_PATH=/tmp/test.db storyteller-test python -m pytest tests/ -q
|
||||
```
|
||||
|
||||
### Run 1 — offline, same-host Ollama
|
||||
|
||||
```bash
|
||||
docker network create --internal offline
|
||||
docker run -d --name ollama --network offline -v ollama-models:/root/.ollama ollama/ollama
|
||||
docker exec ollama ollama pull qwen2.5:3b-instruct
|
||||
|
||||
# The app shares Ollama's network namespace, so Ollama is on its loopback and
|
||||
# neither has a route out.
|
||||
docker run -d --name app --network container:ollama -v story-data:/data \
|
||||
storyteller uvicorn app.main:app --host 127.0.0.1 --port 8000
|
||||
|
||||
# Isolation, before testing anything:
|
||||
docker exec app python -c "import socket; socket.create_connection(('1.1.1.1',443),timeout=4)"
|
||||
# -> OSError: Network is unreachable
|
||||
|
||||
# Listener:
|
||||
docker exec app sh -c "grep -c . /proc/net/tcp" # or the /proc parser in the baseline report
|
||||
```
|
||||
|
||||
Configure and play:
|
||||
|
||||
```bash
|
||||
docker exec app python - <<'PY'
|
||||
import json, urllib.request
|
||||
B = "http://127.0.0.1:8000/api"
|
||||
def call(m, p, b=None):
|
||||
d = json.dumps(b).encode() if b is not None else None
|
||||
r = urllib.request.Request(B+p, data=d, method=m, headers={"Content-Type":"application/json"})
|
||||
with urllib.request.urlopen(r, timeout=900) as f: t = f.read().decode()
|
||||
return json.loads(t) if t.strip() else None
|
||||
call("PUT", "/settings", {"endpoint_url":"http://127.0.0.1:11434/v1",
|
||||
"model":"qwen2.5:3b-instruct","api_mode":"chat","api_key":"",
|
||||
"max_output_tokens":120,"context_token_budget":2048})
|
||||
print(call("POST", "/settings/test"))
|
||||
adv = call("POST", "/adventures", {"title":"Continuity Test","persona_name":"Vale"})["id"]
|
||||
d = json.dumps({"type":"story","text":"I am Vale, a lighthouse keeper."}).encode()
|
||||
r = urllib.request.Request(f"{B}/adventures/{adv}/actions", data=d, method="POST",
|
||||
headers={"Content-Type":"application/json","Accept":"text/event-stream"})
|
||||
with urllib.request.urlopen(r, timeout=900) as f:
|
||||
for line in f:
|
||||
line = line.decode().strip()
|
||||
if line.startswith("data: ") and '"done"' in line:
|
||||
print(json.loads(line[6:])["action"]["text"][:120])
|
||||
PY
|
||||
```
|
||||
|
||||
Restart and resume: `docker restart app`, then re-read
|
||||
`GET /api/adventures/{id}/actions?limit=500` and compare.
|
||||
|
||||
### Run 2 — trusted-LAN Ollama, Internet blocked
|
||||
|
||||
On the **inference machine**, if it serves plain HTTP:
|
||||
|
||||
```bash
|
||||
OLLAMA_HOST=0.0.0.0:11434 ollama serve
|
||||
ollama pull qwen2.5:3b-instruct && ollama pull nomic-embed-text
|
||||
```
|
||||
|
||||
If it serves HTTPS with a private CA, install that CA where the storyteller
|
||||
runs — inside the image for a container:
|
||||
|
||||
```dockerfile
|
||||
FROM storyteller
|
||||
COPY local-ca.crt /usr/local/share/ca-certificates/
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends ca-certificates iproute2 \
|
||||
&& update-ca-certificates && rm -rf /var/lib/apt/lists/*
|
||||
```
|
||||
|
||||
On the **storyteller machine** — LAN reachable, Internet not:
|
||||
|
||||
```bash
|
||||
docker run -d --name app-lan --cap-add=NET_ADMIN \
|
||||
--add-host inference.lan:192.168.0.50 --dns 127.0.0.1 \
|
||||
-v story-data-lan:/data storyteller-lan \
|
||||
uvicorn app.main:app --host 127.0.0.1 --port 8000
|
||||
|
||||
docker exec app-lan ip route del default
|
||||
docker exec app-lan ip route add 192.168.0.0/24 via 172.17.0.1 dev eth0
|
||||
```
|
||||
|
||||
Then set `endpoint_url` to `https://inference.lan:8443/v1` (or
|
||||
`http://192.168.0.50:11434/v1`) and play as above. `POST /api/settings/test`
|
||||
should list the remote host's models.
|
||||
|
||||
### Packet capture (no host root needed)
|
||||
|
||||
```bash
|
||||
printf 'FROM alpine\nRUN apk add --no-cache tcpdump\n' | docker build -q -t tcpdump-img -
|
||||
docker run -d --name cap --network container:app --cap-add=NET_RAW \
|
||||
-v "$PWD/cap":/cap tcpdump-img tcpdump -i any -n -w /cap/run.pcap
|
||||
# … run the campaign …
|
||||
docker stop cap
|
||||
docker run --rm -v "$PWD/cap":/cap tcpdump-img sh -c '
|
||||
tcpdump -r /cap/run.pcap -nn "not net 127.0.0.0/8 and not host 192.168.0.50 and not multicast" | wc -l
|
||||
tcpdump -r /cap/run.pcap -nn -vv "udp port 53" | grep -oE "q: [A-Z]+\? [^ ]+" | sort | uniq -c'
|
||||
```
|
||||
|
||||
### The one uncommitted file
|
||||
|
||||
```bash
|
||||
git add planning/reports/M1-BASELINE-REPORT.md planning/reports/M1-IMPLEMENTATION-REPORT.md
|
||||
git commit -S -m "Add the M1 implementation review report"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# O. Recommendation
|
||||
|
||||
## READY FOR M2 WITH NOTED NON-BLOCKING ISSUES
|
||||
|
||||
M1 delivered its whole scope and its Definition of Done is met on runtime
|
||||
evidence rather than inspection: the application starts and plays with no route
|
||||
to the Internet, through same-host Ollama and through Ollama on a second
|
||||
physical machine; campaigns persist across restart; a failed model call leaves
|
||||
accepted history bit-identical; no runtime asset is fetched; the listener is
|
||||
loopback in every run; and the suite is green at 648 tests, including with no
|
||||
network at all. Nothing was removed from the inherited codebase, so M2 begins
|
||||
against exactly the surface Phase 0B described.
|
||||
|
||||
The issues in §M are non-blocking and mostly land naturally inside milestones
|
||||
that already exist. Two deserve a decision rather than a queue entry: the
|
||||
hardcoded 120 s model timeout, which a first-run user on modest hardware will
|
||||
probably meet before anything else, and Ollama's own `ollama.com` lookup, which
|
||||
is outside this codebase but inside the claim the product makes about itself.
|
||||
|
||||
The finding worth carrying into M2 planning is not a defect but a method. The
|
||||
TLS gap was invisible to static review and to every local run that used
|
||||
loopback HTTP, and it appeared within minutes of pointing the application at a
|
||||
real LAN host. M2 owns the endpoint policy; it should be validated the same
|
||||
way, against a real second machine rather than a container stand-in.
|
||||
@@ -0,0 +1,928 @@
|
||||
# M2 — Baseline Report (evidence record)
|
||||
|
||||
**Date:** 2026-09-02
|
||||
**Milestone:** M2, *Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface*
|
||||
**Companion:** `planning/reports/M2-IMPLEMENTATION-REPORT.md` interprets this file.
|
||||
Where the two disagree on a runtime or test fact, **this file is the record.**
|
||||
|
||||
This is measurement, not commentary. Command output is quoted verbatim.
|
||||
Anything inferred rather than observed is labelled **[inferred]**.
|
||||
|
||||
Hostnames and LAN addresses are **placeholders** — `inference.lan`,
|
||||
`192.168.0.50`, port `8443`. The real ones are in the workspace's untracked
|
||||
notes, never in this repository. Packet counts, digests, timings and status
|
||||
codes are exact.
|
||||
|
||||
---
|
||||
|
||||
## 0. Test environment
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Host | Ubuntu 24.04.4 LTS, x86-64, 4 cores, 15 GB RAM, **no GPU** |
|
||||
| Python | 3.12.3 |
|
||||
| Node / npm | 22.23.1 / 10.9.8 |
|
||||
| Docker | 29.7.2 |
|
||||
| Ollama (same-host) | `ollama/ollama:latest` in a container, models from a shared volume |
|
||||
| Ollama (LAN) | 0.33.0 on a **second physical machine**, HTTPS, certificate from a private CA |
|
||||
| Models | `qwen2.5:0.5b`, `qwen2.5:3b-instruct`, `nomic-embed-text` |
|
||||
| Image under test | `m2-final`, built from the working tree by `docker build .` |
|
||||
|
||||
> **Timings in this file are not benchmarks.** This host was shared with
|
||||
> unrelated work during the review; `uptime` reported a 1-minute load average
|
||||
> ranging from **1.79 to 74.08** across the session on four cores. Where a
|
||||
> measurement was taken under load, it says so. Latency figures are recorded
|
||||
> because they mattered to a *timeout* result, not as performance data.
|
||||
|
||||
---
|
||||
|
||||
## 1. Repository and provenance
|
||||
|
||||
```console
|
||||
$ git rev-parse --abbrev-ref HEAD
|
||||
m2-local-only-surface
|
||||
|
||||
$ git log --format='%H %G? %s' -3
|
||||
8c65ae99deda49b22f415d5868987e00bbcb173c G M2: cut the hosted product away from the local one
|
||||
1a28a9a708985e7c98dcbcb87c189e7582ec288d G Apply post-M1 corrections to the planning package
|
||||
645f07f06d226e274c33d96e71d9c374413ef771 G Add the M1 implementation review report
|
||||
```
|
||||
|
||||
All three commits verify (`%G? = G`). The repository signs every commit; nothing
|
||||
in this milestone is unsigned.
|
||||
|
||||
### Commit chain
|
||||
|
||||
| Commit | Sig | Parent | Meaning |
|
||||
| --- | --- | --- | --- |
|
||||
| `8c65ae9` | G | `1a28a9a` | **The M2 commit.** One commit, no merges. |
|
||||
| `1a28a9a` | G | `645f07f` | Post-M1 planning corrections — the M1 baseline this milestone started from |
|
||||
| `7f182a8` | G | `717670a`, `d72f7c1` | the fork import, two parents |
|
||||
|
||||
### Upstream ancestry
|
||||
|
||||
```console
|
||||
$ git cat-file -t d72f7c1bda0f34fccd84afb7a25c34eb01c901de
|
||||
commit
|
||||
$ git merge-base --is-ancestor d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD && echo yes
|
||||
yes
|
||||
$ git rev-list --count d72f7c1bda0f34fccd84afb7a25c34eb01c901de
|
||||
172
|
||||
$ git diff --stat d72f7c1bda0f34fccd84afb7a25c34eb01c901de HEAD -- LICENSE
|
||||
(no output)
|
||||
```
|
||||
|
||||
All 172 upstream commits reachable; `LICENSE` byte-identical to upstream; no
|
||||
re-import or upstream substitution occurred. `PROVENANCE.md` and
|
||||
`DEVELOPMENT.md` are present and were updated for M2.
|
||||
|
||||
### Private data
|
||||
|
||||
```console
|
||||
$ grep -rn "<the real LAN hostname>\|<the real LAN subnet>\|<the real port>" \
|
||||
--include=*.py --include=*.md --include=*.jsx --include=*.js \
|
||||
--include=*.yml --include=*.txt backend frontend *.md *.yml Dockerfile
|
||||
(no output)
|
||||
```
|
||||
|
||||
No private hostname, LAN address, certificate or credential is committed.
|
||||
|
||||
### Working tree
|
||||
|
||||
At the M2 commit the tree was clean. **This review then modified six files**
|
||||
(§9), so `git status` is *not* clean as this report is written:
|
||||
|
||||
```text
|
||||
M backend/app/memorybank.py
|
||||
M backend/app/models.py
|
||||
M backend/app/routers/adventures/turns.py
|
||||
M backend/app/routers/chat.py
|
||||
M backend/requirements.lock
|
||||
M backend/tests/test_local_only_surface.py
|
||||
```
|
||||
|
||||
Every runtime result below was produced by an image built from that modified
|
||||
tree, except where explicitly marked as pre-fix.
|
||||
|
||||
---
|
||||
|
||||
## 2. Change inventory
|
||||
|
||||
```console
|
||||
$ git diff --shortstat 1a28a9a 8c65ae9
|
||||
94 files changed, 1395 insertions(+), 6578 deletions(-)
|
||||
```
|
||||
|
||||
**Added (3):**
|
||||
|
||||
```text
|
||||
backend/app/endpoints.py
|
||||
backend/tests/test_endpoint_policy.py
|
||||
backend/tests/test_local_only_surface.py
|
||||
```
|
||||
|
||||
**Deleted (24):**
|
||||
|
||||
```text
|
||||
backend/app/accesslog.py backend/tests/test_accesslog.py
|
||||
backend/app/analytics.py backend/tests/test_analytics.py
|
||||
backend/app/cleanup.py backend/tests/test_guest_cleanup.py
|
||||
backend/app/netguard.py backend/tests/test_netguard.py
|
||||
backend/app/security.py backend/tests/test_ratelimit_hardening.py
|
||||
backend/app/routers/analytics.py backend/tests/test_reasoning_param.py
|
||||
backend/app/routers/auth.py
|
||||
backend/app/routers/scripts.py frontend/src/pages/Analytics.jsx
|
||||
backend/app/routers/adventures/scripts.py frontend/src/pages/Scripts.jsx
|
||||
backend/app/scripting/__init__.py frontend/src/pages/ScriptEditor.jsx
|
||||
backend/app/scripting/engine.py frontend/src/pages/Play/panels/ScriptsPanel.jsx
|
||||
backend/app/scripting/pipeline.py frontend/src/pages/Play/drawers/StatusDrawer.jsx
|
||||
render.yaml
|
||||
```
|
||||
|
||||
**Largest modifications:**
|
||||
|
||||
```text
|
||||
+93 -37 backend/app/routers/settings.py
|
||||
+54 -82 backend/app/providers/openai_compatible.py
|
||||
+44 -166 backend/tests/test_chat.py
|
||||
+38 -81 backend/app/routers/chat.py
|
||||
+37 -222 backend/app/auth.py
|
||||
+32 -69 backend/app/main.py
|
||||
+31 -195 backend/app/limits.py
|
||||
+28 -178 backend/app/models.py
|
||||
+22 -43 backend/app/database.py
|
||||
+19 -58 frontend/src/pages/Settings.jsx
|
||||
+18 -97 backend/app/routers/adventures/turns.py
|
||||
+15 -85 frontend/src/App.jsx
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Surface reduction, measured
|
||||
|
||||
### API routes (from the OpenAPI schema, not by grep)
|
||||
|
||||
| Prefix | M1 | M2 |
|
||||
| --- | ---: | ---: |
|
||||
| `/api/adventures` | 25 | 21 |
|
||||
| `/api/analytics` | 3 | **0** |
|
||||
| `/api/auth` | 4 | **0** |
|
||||
| `/api/scripts` | 5 | **0** |
|
||||
| `/api/chat` | 2 | 2 |
|
||||
| `/api/scenarios` | 5 | 5 |
|
||||
| `/api/settings` | 2 | 2 |
|
||||
| `/api/story-cards` | 4 | 4 |
|
||||
| `/api/debug`, `/api/health` | 2 | 2 |
|
||||
| **Total** | **52** | **36** |
|
||||
|
||||
### Dependencies
|
||||
|
||||
```console
|
||||
$ diff <(M1 requirements.txt) <(M2 requirements.txt) | grep '^<'
|
||||
< quickjs>=1.19
|
||||
< cryptography>=42
|
||||
< psycopg[binary]>=3.2
|
||||
```
|
||||
|
||||
Installed closure, measured by building a clean venv from `requirements.txt` +
|
||||
`requirements-dev.txt`:
|
||||
|
||||
| | M1 | M2 |
|
||||
| --- | ---: | ---: |
|
||||
| Python packages installed | 40 | **34** |
|
||||
| Packages gone | — | `cffi`, `cryptography`, `psycopg`, `psycopg-binary`, `pycparser`, `quickjs` |
|
||||
| npm runtime dependencies | 6 | **3** |
|
||||
| npm packages installed (`npm ls --all`) | 53 | **32** |
|
||||
|
||||
### Code size
|
||||
|
||||
| | M1 | M2 | Δ |
|
||||
| --- | ---: | ---: | ---: |
|
||||
| `backend/app` Python lines | 14 298 | 11 632 | −2 666 |
|
||||
| `frontend/src` JS/JSX lines | 6 523 | 5 239 | −1 284 |
|
||||
| `frontend/src/pages` files | 21 | 16 | −5 |
|
||||
|
||||
### Environment variables actually read (`os.environ`)
|
||||
|
||||
```console
|
||||
$ grep -rn "os.environ" backend/app/
|
||||
backend/app/database.py:18:_env_db_path = os.environ.get("AIDND_DB_PATH")
|
||||
backend/app/main.py:30: for o in os.environ.get("AIDND_CORS_ORIGINS", "").split(",")
|
||||
```
|
||||
|
||||
Two, down from ten. `AIDND_MULTI_USER`, `AIDND_SECRET_KEY`, `AIDND_COOKIE_SECURE`,
|
||||
`AIDND_DEMO_API_KEY`, `AIDND_DEMO_ENDPOINT_URL`, `AIDND_DEMO_MODELS`,
|
||||
`AIDND_DEMO_TURNS_PER_DAY`, `AIDND_POWER_USERS`, `AIDND_ANALYTICS_EMAILS`,
|
||||
`AIDND_TRUSTED_PROXY_HOPS`, `AIDND_DATABASE_URL` and `DATABASE_URL` are no
|
||||
longer read. Four of those names still appear in the tree **as prose in
|
||||
comments** explaining what was removed; `grep` above shows no read.
|
||||
|
||||
---
|
||||
|
||||
## 4. Test suite
|
||||
|
||||
```console
|
||||
$ cd backend && .venv/bin/python -m pytest tests/ -q
|
||||
606 passed, 1 warning in 126.29s (0:02:06)
|
||||
```
|
||||
|
||||
Zero failed, zero skipped, zero xfailed. The one warning is the inherited
|
||||
Starlette/`httpx` deprecation notice, present since M1.
|
||||
|
||||
### Accounting for every test
|
||||
|
||||
Counted by diffing collected node IDs between `1a28a9a` (M1) and the working
|
||||
tree, not by reading diffs:
|
||||
|
||||
| | Count |
|
||||
| --- | ---: |
|
||||
| Distinct test functions, M1 | 610 |
|
||||
| Distinct test functions, M2 | 561 |
|
||||
| Node IDs gone | 88 |
|
||||
| Node IDs new | 39 |
|
||||
| Collected tests (with parametrisation), M1 | 648 |
|
||||
| Collected tests, M2 | **606** |
|
||||
|
||||
**Gone, by file:**
|
||||
|
||||
| File | Gone | Disposition |
|
||||
| --- | ---: | --- |
|
||||
| `test_analytics.py` | 25 | whole file — subject removed |
|
||||
| `test_guest_cleanup.py` | 14 | whole file — subject removed |
|
||||
| `test_accesslog.py` | 12 | whole file — subject removed |
|
||||
| `test_ratelimit_hardening.py` | 8 | whole file — subject removed |
|
||||
| `test_chat.py` | 8 | demo-key pinning and the power-user gate |
|
||||
| `test_reasoning_param.py` | 5 | whole file — subject removed |
|
||||
| `test_netguard.py` | 5 | whole file — replaced by `test_endpoint_policy.py` |
|
||||
| `test_prompt_caching.py` | 4 | OpenRouter upstream routing |
|
||||
| `test_memory_rewrite.py` | 4 | 2 retired (Postgres DSN masking), **2 renamed** |
|
||||
| `test_delete_state.py` | 1 | **renamed** |
|
||||
| `test_branch_forking.py` | 1 | **renamed** |
|
||||
| `test_branch_clause.py` | 1 | scripting history API |
|
||||
|
||||
**Four of the 88 are renames with equivalent coverage**, so 84 tests were
|
||||
genuinely retired:
|
||||
|
||||
```text
|
||||
test_branch_forking: test_switching_restores_the_script_and_world_state
|
||||
-> test_switching_restores_the_state_a_branch_left_behind
|
||||
test_delete_state: test_deleting_the_ai_turn_rewinds_the_script_state
|
||||
-> test_deleting_the_ai_turn_rewinds_the_counter
|
||||
test_memory_rewrite: test_an_owner_with_no_api_key_is_skipped
|
||||
-> test_an_adventure_with_no_model_configured_is_skipped
|
||||
test_memory_rewrite: test_an_api_key_on_the_command_line_covers_that_owner
|
||||
-> test_a_model_on_the_command_line_covers_that_adventure
|
||||
```
|
||||
|
||||
**New, by file:** `test_local_only_surface.py` 18, `test_endpoint_policy.py`
|
||||
14, `test_chat.py` 3, plus the 4 renames. Collected with parametrisation:
|
||||
`test_endpoint_policy.py` 31, `test_local_only_surface.py` 32.
|
||||
|
||||
### Instrumentation conversion, not deletion
|
||||
|
||||
Eight files used a QuickJS `output` hook (`state.gold += 10`) as deterministic
|
||||
instrumentation for the **state snapshot and rollback machinery**, which M2 does
|
||||
not touch. The counter moved to the world-state engine — the model emits a
|
||||
` ```state ` delta block, the referee applies it — and the assertions are
|
||||
unchanged in substance. Affected: `test_take_state`, `test_attempt_siblings`,
|
||||
`test_branch_forking`, `test_bundle_v2`, `test_delete_state`,
|
||||
`test_retry_variants`, `test_story_tree_baseline`, `test_turn_flow_integration`,
|
||||
plus `test_state_revert` converted from `script_state`/`state_after` to
|
||||
`world_state`/`world_state_after`.
|
||||
|
||||
### Frontend and image
|
||||
|
||||
```console
|
||||
$ npm run lint # oxlint
|
||||
exit 0 # 7 warnings, all pre-existing react/only-export-components
|
||||
$ npm run build
|
||||
dist/assets/index-C9v6AJ3R.js 395.41 kB │ gzip: 120.01 kB ✓ built in 496ms
|
||||
$ docker build .
|
||||
exit 0
|
||||
```
|
||||
|
||||
There are no automated frontend tests in the repository — none existed at M1
|
||||
either.
|
||||
|
||||
Bundle size: **933.69 kB → 395.41 kB** (M1 → M2), from removing CodeMirror with
|
||||
the script editor.
|
||||
|
||||
---
|
||||
|
||||
## 5. Offline run — same-host Ollama
|
||||
|
||||
**Topology.** `--internal` Docker network (no NAT, no external DNS). Ollama in
|
||||
one container; the `m2-final` image in a second container sharing Ollama's
|
||||
network namespace, so Ollama sits on the storyteller's own loopback. tcpdump ran
|
||||
in that namespace for the whole session.
|
||||
|
||||
### 5.1 Isolation, verified before any test
|
||||
|
||||
```text
|
||||
blocked 1.1.1.1:443 OSError
|
||||
blocked 140.82.121.4:443 OSError
|
||||
no resolution fonts.googleapis.com
|
||||
no resolution fonts.gstatic.com
|
||||
no resolution openaipublic.blob.core.windows.net
|
||||
no resolution openrouter.ai
|
||||
no resolution api.openai.com
|
||||
no resolution github.com
|
||||
```
|
||||
|
||||
### 5.2 Listeners
|
||||
|
||||
```text
|
||||
LISTEN 127.0.0.1:8000 <- the storyteller
|
||||
LISTEN 127.0.0.11:42229 <- Docker's embedded DNS
|
||||
LISTEN 127.0.0.1:43093 <- Ollama's model runner
|
||||
LISTEN 127.0.0.1:43313 <- Ollama's model runner
|
||||
LISTEN [::]:11434 <- Ollama itself
|
||||
```
|
||||
|
||||
Nothing the storyteller owns is bound off loopback.
|
||||
|
||||
### 5.3 Ollama-only settings and diagnostics
|
||||
|
||||
```json
|
||||
endpoint: http://127.0.0.1:11434/v1 | model: qwen2.5:0.5b
|
||||
embed: nomic-embed-text:latest | timeout: 600
|
||||
connection test: {"ok": true, "models": ["qwen2.5:7b-instruct",
|
||||
"qwen2.5:3b-instruct", "nomic-embed-text:latest", "qwen2.5:0.5b"]}
|
||||
```
|
||||
|
||||
### 5.4 Story generation and streaming
|
||||
|
||||
```text
|
||||
turn 1 ( 1.7s, 90 stream chunks): 'Oh, how unfortunate! A lighthouse keeper indeed! The light'
|
||||
turn 2 ( 35.2s, 90 stream chunks): 'I climb the spiral stair to the lamp room. My hand tightly'
|
||||
turn 3 ( 7.1s, 90 stream chunks): "I search the keeper's log for the last entry. The last ent"
|
||||
turn 4: ERROR {'type': 'error', 'detail': 'The AI endpoint timed out.'}
|
||||
```
|
||||
|
||||
Turn 4 timed out at the configured 600 s. The application logged no error. Ollama
|
||||
reported the model still resident. `uptime` at the time: 1-minute load average
|
||||
had risen from 1.79 to the tens on four cores from unrelated work on this host.
|
||||
**[inferred]** the timeout is host contention rather than an application fault;
|
||||
what is *measured* is that no application error was logged and that the same
|
||||
build produced turns in 1.7–35.2 s minutes earlier.
|
||||
|
||||
### 5.5 Local embeddings, end to end through the app
|
||||
|
||||
```text
|
||||
wrote memory id=1, embedded=False
|
||||
turn to trigger the pass (101.6s)
|
||||
memories: 1 embedded with the local model: 1
|
||||
```
|
||||
|
||||
Direct timing of the embedding endpoint from inside the container:
|
||||
|
||||
```text
|
||||
embedding round trip: 3.4s, dim=768
|
||||
second embedding (warm): 0.1s
|
||||
```
|
||||
|
||||
### 5.6 Branch-scoped memory isolation — positive and negative controls
|
||||
|
||||
```text
|
||||
main branch = 1
|
||||
memories on main: 0 ALPHA present: False
|
||||
forking at AI action 22 (an earlier turn, so this starts a new line)
|
||||
fork events: ['chunk', 'chunk', 'done']
|
||||
branches now: 2, fork = 5
|
||||
|
||||
ON THE FORK (0 memories)
|
||||
ALPHA visible: False <- ALPHA is anchored at main's head, which is BELOW the
|
||||
fork point, so the fork correctly does not inherit it
|
||||
BETA written here, visible: True <- positive control
|
||||
|
||||
BACK ON MAIN (1 memories)
|
||||
ALPHA visible: True <- positive control
|
||||
BETA visible: False <- NEGATIVE CONTROL, must be False
|
||||
```
|
||||
|
||||
The decisive result is the last line: a memory written on the fork is **not**
|
||||
visible on main.
|
||||
|
||||
### 5.7 A05 — a failed model call, and a hand-edited database
|
||||
|
||||
```text
|
||||
baseline: (8 actions, 3 accepted AI turns, digest 07615da99b014b70)
|
||||
|
||||
-- invalid local model --
|
||||
events: ['player', 'error']
|
||||
error: Endpoint or model not found (HTTP 404). Check the endpoint URL and that model 'n…
|
||||
after: (7, 3, 'e18de9dff6455e0d')
|
||||
|
||||
-- endpoint refused by policy at request time --
|
||||
(the settings row was edited directly with sqlite3, behind the app's back,
|
||||
to https://openrouter.ai/api/v1)
|
||||
events: ['player', 'error']
|
||||
error: This endpoint can't be used — openrouter.ai is a cloud inference service — this build…
|
||||
after: (8, 3, '157c882c60579617')
|
||||
```
|
||||
|
||||
The **accepted AI-turn count stayed at 3** through both failures. Only the
|
||||
player's own typed action was added each time, which is M1's documented and
|
||||
deliberate behaviour.
|
||||
|
||||
The second case is the strongest form of the endpoint evidence: the request-time
|
||||
check refused a cloud endpoint that had been written straight into SQLite,
|
||||
bypassing the API's save-time validation entirely.
|
||||
|
||||
### 5.8 Restart and resume
|
||||
|
||||
```text
|
||||
before restart: 8 actions, digest 157c882c60579617
|
||||
branches: 2 memories: 1
|
||||
after restart: 8 actions, digest 157c882c60579617
|
||||
branches: 2 memories: 1
|
||||
settings preserved: endpoint http://127.0.0.1:11434/v1 timeout 600
|
||||
context inspection: ['narrator', 'persona', 'history', 'used_memories']
|
||||
```
|
||||
|
||||
### 5.9 Packet capture — whole offline session
|
||||
|
||||
```text
|
||||
all packets: 5131
|
||||
loopback (127.0.0.0/8): 5086
|
||||
non-loopback unicast: 0
|
||||
TCP connections opened outside loopback: (none)
|
||||
```
|
||||
|
||||
DNS queries, attributed by timestamp rather than assumed:
|
||||
|
||||
| Name | First seen | Attribution |
|
||||
| --- | --- | --- |
|
||||
| `fonts.googleapis.com` | 16:44:02.800449 | **the isolation probe** in §5.1 |
|
||||
| `openaipublic.blob.core.windows.net` | 16:44:02.802749 | the isolation probe |
|
||||
| `openrouter.ai` | 16:44:02.803260 | the isolation probe |
|
||||
| `api.openai.com` | 16:44:02.803956 | the isolation probe |
|
||||
| `github.com` | 16:44:02.804747 | the isolation probe |
|
||||
| `ollama-container` | 18:28:55.684172 | the endpoint-policy edge-case test (§7) |
|
||||
| `ollama.com` | 16:53:29 … 18:28:05 | **the Ollama server's own lookup** |
|
||||
|
||||
```text
|
||||
capture window: 16:44:02.800136 .. 18:51:29.080336
|
||||
first packet to the storyteller API: 16:44:47.639083
|
||||
```
|
||||
|
||||
Every cloud/font/tokenizer lookup landed within **5 milliseconds of the capture
|
||||
starting and 45 seconds before the storyteller received its first request** —
|
||||
they are the deliberate probe, not the application. `ollama.com` recurs
|
||||
throughout and is issued by the separately installed Ollama service, the same
|
||||
observation M1 recorded.
|
||||
|
||||
---
|
||||
|
||||
## 6. Offline run — trusted-LAN Ollama over HTTPS
|
||||
|
||||
**Topology.** Ollama 0.33.0 on a **second physical machine** on the trusted LAN,
|
||||
serving HTTPS with a certificate from a private CA. The storyteller ran in a
|
||||
container with `NET_ADMIN`, its **default route deleted** and replaced with a
|
||||
route to the LAN subnet only, its resolver pointed at nothing, and the host name
|
||||
supplied as a static hosts entry.
|
||||
|
||||
### 6.1 Isolation and binding
|
||||
|
||||
```text
|
||||
=== isolation ===
|
||||
blocked 1.1.1.1:443 OSError
|
||||
blocked 140.82.121.4:443 OSError
|
||||
no resolution github.com
|
||||
no resolution openrouter.ai
|
||||
no resolution fonts.gstatic.com
|
||||
=== the approved LAN host ===
|
||||
inference.lan -> 192.168.0.50
|
||||
TLS verified, peer CN = inference.lan
|
||||
=== listeners ===
|
||||
LISTEN 127.0.0.1:8000
|
||||
```
|
||||
|
||||
Nothing listens on the container's own LAN-facing address.
|
||||
|
||||
### 6.2 Diagnostics and generation
|
||||
|
||||
```text
|
||||
endpoint: https://inference.lan:8443/v1 | timeout: 600 s
|
||||
diagnostics: {"ok": true, "models": ["qwen2.5:3b-instruct", "nomic-embed-text:latest"]}
|
||||
turn 1 ( 10.7s, 24 chunks): 'You check the oil and wick, their greasy slicks staining your hand'
|
||||
turn 2 ( 4.0s, 17 chunks): 'The air is thick with the musty scent of the old building as you a'
|
||||
turn 3 ( 4.6s, 26 chunks): 'You find the last entry: "Exhausted... Oil\'s low..." The room feel'
|
||||
|
||||
memories: 1 embedded via the LAN host: 1
|
||||
```
|
||||
|
||||
### 6.3 Every model request the application made
|
||||
|
||||
```text
|
||||
4 https://inference.lan:8443/v1/chat/completions
|
||||
1 https://inference.lan:8443/v1/embeddings
|
||||
```
|
||||
|
||||
Narration and embeddings both went to the configured host over verified TLS. No
|
||||
other URL appears.
|
||||
|
||||
### 6.4 Retry keeps the discarded attempt
|
||||
|
||||
Measured on the shipped image, against the LAN host:
|
||||
|
||||
```text
|
||||
head AI action before retry: 8 take_count 1
|
||||
retry produced (13.7s): 'You jot down a hurried note: "Visited lamp—oil low—oil lamp '
|
||||
head AI action after retry: 9 take_count 2 take_index 1
|
||||
alternate takes retained: 2
|
||||
- 'Your pen touches the cold wood of the log, the friction'
|
||||
- 'You jot down a hurried note: "Visited lamp—oil low—oil '
|
||||
branches: 1
|
||||
```
|
||||
|
||||
A retry at the tip files the new attempt as a sibling at the same coordinate and
|
||||
keeps the one it replaced. No branch is created, which is correct for a retry at
|
||||
the head.
|
||||
|
||||
### 6.5 Restart and resume
|
||||
|
||||
```text
|
||||
before restart: 8 actions, digest 66cc82944f8b6eef
|
||||
after restart: 8 actions, digest 66cc82944f8b6eef
|
||||
endpoint preserved: https://inference.lan:8443/v1 timeout 600
|
||||
memories: 1
|
||||
```
|
||||
|
||||
### 6.6 Packet capture
|
||||
|
||||
```text
|
||||
all packets: 685
|
||||
loopback: 364
|
||||
to/from the LAN Ollama host: 308
|
||||
any other unicast: 0
|
||||
DNS queries: none
|
||||
```
|
||||
|
||||
Zero DNS: the endpoint was configured, not discovered.
|
||||
|
||||
---
|
||||
|
||||
## 7. Endpoint policy
|
||||
|
||||
Run inside the shipped image:
|
||||
|
||||
```text
|
||||
ALLOW same-host loopback http://127.0.0.1:11434/v1
|
||||
ALLOW localhost name http://localhost:11434/v1
|
||||
ALLOW IPv6 loopback http://[::1]:11434/v1
|
||||
ALLOW LAN literal http://192.168.1.50:11434/v1
|
||||
ALLOW link-local http://169.254.10.5:11434/v1
|
||||
ALLOW CGNAT / mesh VPN http://100.64.3.4:11434/v1
|
||||
ALLOW IPv6 unique-local http://[fd00::5]:11434/v1
|
||||
ALLOW docker internal name http://ollama-container:11434/v1
|
||||
REJECT cloud provider https://openrouter.ai/api/v1
|
||||
openrouter.ai is a cloud inference service — this build talks to Ollama on your own mach…
|
||||
REJECT cloud provider https://api.openai.com/v1
|
||||
REJECT public IPv4 http://8.8.8.8:11434/v1
|
||||
8.8.8.8 resolves to 8.8.8.8, which is a public Internet address — …
|
||||
REJECT public IPv6 http://[2001:4860:4860::8888]/v1
|
||||
REJECT unspecified address http://0.0.0.0:11434/v1
|
||||
0.0.0.0 … is not on this machine and not on your own network …
|
||||
REJECT wrong scheme ftp://127.0.0.1/v1
|
||||
```
|
||||
|
||||
Over the HTTP API, from inside the offline container:
|
||||
|
||||
```text
|
||||
=== cloud and public endpoints are refused ===
|
||||
400 https://openrouter.ai/api/v1 That endpoint can't be used — openrouter.ai is a cloud…
|
||||
400 https://api.openai.com/v1 …
|
||||
400 https://api.groq.com/openai/v1 …
|
||||
400 http://8.8.8.8:11434/v1 …8.8.8.8 … is a public Internet address…
|
||||
=== local and LAN endpoints are accepted ===
|
||||
200 http://127.0.0.1:11434/v1
|
||||
200 http://192.168.1.50:11434/v1
|
||||
200 http://[::1]:11434/v1
|
||||
```
|
||||
|
||||
Request-time enforcement against a hand-edited database is in §5.7.
|
||||
|
||||
`test_endpoint_policy.py` (31 collected) resolves hostnames through a stub, so
|
||||
it exercises the policy rather than the machine's DNS. It covers the split-horizon
|
||||
case: a name resolving to both `192.168.1.50` and a public address is refused.
|
||||
|
||||
---
|
||||
|
||||
## 8. TLS trust
|
||||
|
||||
```console
|
||||
$ grep -rn "verify=False\|ssl._create_unverified\|CERT_NONE\|check_hostname = False" backend/app
|
||||
(the only match is prose in tlstrust.py saying such an option deliberately does not exist)
|
||||
|
||||
$ grep -rn "httpx.AsyncClient(\|httpx.Client(" backend/app
|
||||
backend/app/routers/settings.py:122
|
||||
backend/app/providers/openai_compatible.py:214
|
||||
backend/app/providers/openai_compatible.py:326
|
||||
backend/app/providers/openai_compatible.py:360
|
||||
```
|
||||
|
||||
Four clients, all four passing `verify=tlstrust.ssl_context()`. Enforced by
|
||||
`test_tls_trust.py`, which walks the AST of the two modules that make outbound
|
||||
requests and fails if any `httpx.AsyncClient` is constructed without `verify`,
|
||||
or if a third module starts making requests:
|
||||
|
||||
```console
|
||||
$ .venv/bin/python -m pytest tests/test_tls_trust.py -q
|
||||
6 passed in 0.20s
|
||||
```
|
||||
|
||||
Runtime confirmation of certificate **and** hostname verification against the
|
||||
real private-CA host is §6.1 (`TLS verified, peer CN = inference.lan`), and of
|
||||
all four paths — settings/test, narration, model listing, embeddings — §6.2–6.3.
|
||||
|
||||
### Modules that can open an outbound connection at all
|
||||
|
||||
```console
|
||||
$ grep -rln "httpx\|requests\.\|urllib.request\|socket\.\|aiohttp\|websocket" backend/app
|
||||
backend/app/endpoints.py # getaddrinfo only, for the policy
|
||||
backend/app/providers/openai_compatible.py
|
||||
backend/app/routers/settings.py
|
||||
backend/app/tlstrust.py # builds the SSL context; opens nothing
|
||||
```
|
||||
|
||||
Remaining absolute URLs anywhere in backend application code:
|
||||
|
||||
```text
|
||||
3 http://127.0.0.1
|
||||
2 http://localhost
|
||||
1 https://openaipublic.blob.core.windows.net <- a comment and a SOURCE_URL
|
||||
constant in encoding.py; never fetched
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. Defects found during this review
|
||||
|
||||
Three, all found by running the shipped build rather than by reading it. Each was
|
||||
corrected because the M2 evidence could not otherwise be accurate; each is
|
||||
reported rather than absorbed.
|
||||
|
||||
### 9.1 The memory bank was broken — summaries and embeddings silently stopped
|
||||
|
||||
```text
|
||||
Task exception was never retrieved
|
||||
future: <Task finished coro=<run_post_turn() ...> exception=AttributeError(
|
||||
"'Settings' object has no attribute 'api_key_plain'")>
|
||||
File "/app/backend/app/memorybank.py", line 155, in summary_provider
|
||||
settings.api_key_plain,
|
||||
```
|
||||
|
||||
M2 removed `Settings.api_key_plain` with the API key, but `memorybank`'s two
|
||||
provider factories still read it. It failed in a fire-and-forget background
|
||||
task, so nothing surfaced to the user and no test caught it: every memory test
|
||||
stubs those factories out. **All 604 tests passed with this defect present.**
|
||||
|
||||
Fixed by rewriting both factories. Covered by a new test that constructs every
|
||||
provider factory from a real `Settings` row.
|
||||
|
||||
### 9.2 The configurable model timeout never reached the turn engine
|
||||
|
||||
`settings.model_timeout_seconds` was stored, validated, exposed in the API and
|
||||
rendered in the UI — and not passed to `OpenAICompatibleProvider` in
|
||||
`turns.py` or `chat.py`, so generation used the module default. The M2 exit
|
||||
criterion "no longer an undocumented hardcoded limitation" was therefore only
|
||||
half met: the constant had moved, but the setting was inert.
|
||||
|
||||
Fixed in both call sites. Covered by a new test that drives the turn endpoint,
|
||||
the chat endpoint and the summariser factory and asserts the configured value
|
||||
reaches each.
|
||||
|
||||
### 9.3 `requirements.lock` still pinned the removed packages
|
||||
|
||||
```console
|
||||
$ grep -n "quickjs\|psycopg\|cryptography" backend/requirements.lock
|
||||
25:cryptography==50.0.1
|
||||
36:psycopg==3.3.5
|
||||
37:psycopg-binary==3.3.5
|
||||
45:quickjs==1.19.4
|
||||
```
|
||||
|
||||
`DEVELOPMENT.md` tells a new developer to install from the lock, which would have
|
||||
reinstalled all three. Regenerated: 40 pins → 34.
|
||||
|
||||
### Verification after the fixes
|
||||
|
||||
```console
|
||||
$ .venv/bin/python -m pytest tests/ -q
|
||||
606 passed, 1 warning in 126.29s
|
||||
$ npm run lint # exit 0
|
||||
$ npm run build # ✓ built in 496ms
|
||||
$ docker build . # exit 0
|
||||
```
|
||||
|
||||
All §5 and §6 runtime evidence above was produced by an image built **after**
|
||||
these fixes.
|
||||
|
||||
---
|
||||
|
||||
## 10. M1 capability regression checks
|
||||
|
||||
| M1 capability | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| Vendored tokenizer, no first-turn download | PASS | §10.1 |
|
||||
| Self-hosted fonts | PASS | §10.2 |
|
||||
| Same-origin runtime assets | PASS | §10.2 |
|
||||
| Restrictive CSP | PASS | §10.2 |
|
||||
| Loopback storyteller default | PASS | §5.2, §6.1, §11 |
|
||||
| Same-host Ollama | PASS | §5.3–5.4 |
|
||||
| Trusted-LAN Ollama | PASS | §6.2 |
|
||||
| HTTPS / private-CA trusted-LAN | PASS | §6.1–6.3 |
|
||||
| Certificate + hostname verification | PASS | §6.1, §8 |
|
||||
| Offline story generation | PASS | §5.1, §5.4 |
|
||||
| SQLite persistence | PASS | §5.8, §6.4 |
|
||||
| Full regression suite green | PASS | §4 |
|
||||
|
||||
### 10.1 Tokenizer, inside the shipped image, offline
|
||||
|
||||
```text
|
||||
vendored table sha256 matches pin: True
|
||||
count_tokens with sockets blocked: 6 tokens
|
||||
```
|
||||
|
||||
### 10.2 Browser asset graph, fetched over loopback with no route out
|
||||
|
||||
```text
|
||||
CSP: default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline';
|
||||
font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none';
|
||||
base-uri 'none'; form-action 'self'; frame-ancestors 'none'
|
||||
|
||||
referenced by index.html:
|
||||
/favicon.svg LOCAL
|
||||
/assets/index-*.js LOCAL
|
||||
/assets/index-*.css LOCAL
|
||||
|
||||
font urls inside the stylesheet:
|
||||
/fonts/cinzel-normal-latin-ext.woff2 200 font/woff2 14540
|
||||
/fonts/cinzel-normal-latin.woff2 200 font/woff2 25904
|
||||
/fonts/crimson-pro-italic-latin-ext.woff2 200 font/woff2 39808
|
||||
/fonts/crimson-pro-italic-latin.woff2 200 font/woff2 51432
|
||||
/fonts/crimson-pro-normal-latin-ext.woff2 200 font/woff2 37988
|
||||
/fonts/crimson-pro-normal-latin.woff2 200 font/woff2 48200
|
||||
/fonts/inter-normal-latin-ext.woff2 200 font/woff2 85068
|
||||
/fonts/inter-normal-latin.woff2 200 font/woff2 48256
|
||||
```
|
||||
|
||||
Unchanged from M1, including the `object-src`/`base-uri`/`form-action`
|
||||
directives M1 added.
|
||||
|
||||
---
|
||||
|
||||
## 11. Storyteller network exposure
|
||||
|
||||
| Launch path | Listener / publish | Verified by |
|
||||
| --- | --- | --- |
|
||||
| `./start.sh` (native dev) | `uvicorn --host 127.0.0.1 --port 8000` | file assertion in `test_local_only_surface.py` |
|
||||
| `start.ps1` (Windows) | `--host 127.0.0.1` | same test |
|
||||
| Production native (documented in `DEVELOPMENT.md`) | `--host 127.0.0.1` | documentation |
|
||||
| `docker compose up` | `ports: - "127.0.0.1:8000:8000"` | test parses the published mappings and asserts every one starts `127.0.0.1:` |
|
||||
| Container process itself | `--host 0.0.0.0` | deliberate; the only address a published port can reach |
|
||||
|
||||
`DEVELOPMENT.md` states that publishing the port to `0.0.0.0` is a deliberate
|
||||
decision this project's threat model does not cover. The `Dockerfile` carries the
|
||||
same warning next to `EXPOSE`.
|
||||
|
||||
Two supporting checks:
|
||||
|
||||
* `AIDND_CORS_ORIGINS="*"` makes the application **refuse to start**:
|
||||
`RuntimeError: AIDND_CORS_ORIGINS must not contain '*'…`
|
||||
* An unknown `/api/...` path now returns **404** instead of the SPA's HTML with
|
||||
status 200 (verified for ten removed endpoints, §12).
|
||||
|
||||
---
|
||||
|
||||
## 12. Removed-surface verification, at runtime
|
||||
|
||||
From inside the offline container:
|
||||
|
||||
```text
|
||||
=== removed surfaces answer 404 ===
|
||||
404 GET /auth/me 404 GET /analytics/summary
|
||||
404 GET /auth/login 404 GET /analytics/collect
|
||||
404 GET /auth/register 404 GET /analytics/access
|
||||
404 GET /auth/logout 404 GET /scripts
|
||||
404 GET /adventures/2/scripts
|
||||
404 GET /adventures/2/script-state
|
||||
|
||||
=== no API key in the settings surface ===
|
||||
api_key present: False has_api_key present: False
|
||||
after trying to set one: False
|
||||
```
|
||||
|
||||
```console
|
||||
$ python -c "import app.scripting"
|
||||
ImportError # asserted by test_the_application_has_no_scripting_engine
|
||||
$ grep -rn "import quickjs" backend/app
|
||||
(no output)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 13. Inert schema retained for compatibility
|
||||
|
||||
Verified present in the database and unread by the application.
|
||||
|
||||
| Object | Was | Status |
|
||||
| --- | --- | --- |
|
||||
| table `scripts` | script library | unmapped; no model, no query |
|
||||
| table `adventure_scripts` | per-adventure script copies | unmapped |
|
||||
| table `analytics_daily` | visitor counters | unmapped |
|
||||
| table `analytics_visitor_days` | visitor funnel | unmapped |
|
||||
| table `access_log` | sign-ins, addresses, devices | unmapped |
|
||||
| column `adventures.script_state` | scripting shared state | written `{}` only |
|
||||
| column `actions.state_after` | scripting state per node | written `{}` only |
|
||||
| column `settings.api_key` | encrypted cloud key | never read or written |
|
||||
| column `settings.reasoning_max_tokens` | OpenRouter thinking budget | never read |
|
||||
| column `users.demo_turns_used` / `_date` | demo cap tally | never read |
|
||||
| table `users` + `user_id` FKs | multi-user ownership | **active**, one row, internal identity only |
|
||||
|
||||
No destructive migration was performed. One additive migration was added:
|
||||
|
||||
```python
|
||||
(77, "ALTER TABLE settings ADD COLUMN model_timeout_seconds INTEGER NOT NULL DEFAULT 300")
|
||||
```
|
||||
|
||||
Existing M1 databases open unchanged: the offline run in §5 used the volume
|
||||
carrying campaigns created before these fixes, and §5.8 shows the transcript
|
||||
digest surviving both an image replacement and a restart.
|
||||
|
||||
---
|
||||
|
||||
## 14. What was not measured
|
||||
|
||||
Stated so the implementation report does not overclaim.
|
||||
|
||||
1. **No browser rendered the UI.** No browser automation was available. §10.2
|
||||
fetches the complete asset graph the page references, with correct media
|
||||
types and the CSP, but nobody looked at the rendered page. This is unchanged
|
||||
from M1, where the maintainer confirmed the render by hand.
|
||||
2. **Sustained multi-turn play under the memory bank was not completed on this
|
||||
host.** Turns 1–3 succeeded; turn 4 timed out under third-party CPU load
|
||||
(§5.4). The individual capabilities — narration, streaming, embeddings,
|
||||
memory writes, retry, fork, restart — were each measured separately.
|
||||
3. **No load, soak or long-campaign testing.** Out of scope for M2.
|
||||
|
||||
---
|
||||
|
||||
# 15. Closeout note — appended 2026-09-03
|
||||
|
||||
**This section was appended after the fact and is not part of the original
|
||||
evidence record.** Everything above was written on 2026-09-02 and describes the
|
||||
repository as it stood then. Nothing above has been rewritten, including its
|
||||
references to a working tree that was dirty at the time.
|
||||
|
||||
## What the original report correctly said
|
||||
|
||||
The report above correctly described commit `8c65ae9` — the M2 implementation
|
||||
commit — as **not containing** the three fixes this review found. At the time it
|
||||
was written, those fixes existed only in the working tree, six files ahead of
|
||||
that commit. Every runtime measurement in this report was produced by an image
|
||||
built from that fixed tree, which the report states plainly.
|
||||
|
||||
## What happened afterwards
|
||||
|
||||
The six-file correction was committed:
|
||||
|
||||
```text
|
||||
8652fe7cd84bca5173abb03b2a692f15fea8a98c M2 review: two regressions the green suite hid, and the reports
|
||||
```
|
||||
|
||||
That commit carries the six implementation/test/lockfile files **and** the two
|
||||
M2 reports themselves, which is why this report's own history begins there. The
|
||||
provenance distinction the review asked for is preserved regardless:
|
||||
`8c65ae9` is the M2 implementation, `8652fe7` is the review correction, and the
|
||||
two were never squashed.
|
||||
|
||||
## Verification at closeout
|
||||
|
||||
Re-run on 2026-09-03 against the committed tree — not the working tree — so that
|
||||
what was verified is exactly what the repository contains:
|
||||
|
||||
```text
|
||||
backend .venv/bin/python -m pytest tests/ -q 606 passed in 133.26s
|
||||
frontend npm run lint 7 warnings, 0 errors, exit 0
|
||||
frontend npm run build built in 932ms; index.js 395.41 kB
|
||||
root docker build -t storyteller-m2-closeout . exit 0
|
||||
git status --short clean
|
||||
```
|
||||
|
||||
The 606 figure matches the count this report recorded, from the same tree.
|
||||
|
||||
Dependency removal was verified in the built image rather than only in the
|
||||
lockfile:
|
||||
|
||||
```text
|
||||
$ docker run --rm --entrypoint sh storyteller-m2-closeout -c "pip list | grep -iE 'quickjs|psycopg|cryptography|cffi|pycparser'"
|
||||
ABSENT: none of quickjs/psycopg/cryptography/cffi/pycparser installed
|
||||
32 packages total
|
||||
```
|
||||
|
||||
The network measurements in §5, §6 and §10 were **not** re-run. The six
|
||||
corrected files change provider construction and a timeout value; they do not
|
||||
touch the endpoint policy, the TLS trust context, or the bind addresses, so the
|
||||
captures above remain the evidence for the network boundary.
|
||||
@@ -0,0 +1,850 @@
|
||||
# M2 — Implementation Review Report
|
||||
|
||||
**Date:** 2026-09-02
|
||||
**Milestone:** M2, *Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface*
|
||||
**Audience:** the architecture/design reviewer deciding whether to accept M2 and start M3
|
||||
**Evidence:** `planning/reports/M2-BASELINE-REPORT.md`. Section references below
|
||||
(§) point into it, and where this report and that one differ on a runtime or
|
||||
test fact, **the baseline report is the record.**
|
||||
|
||||
Hostnames and LAN addresses are placeholders (`inference.lan`, `192.168.0.50`).
|
||||
|
||||
---
|
||||
|
||||
# A. Executive result
|
||||
|
||||
```text
|
||||
Overall M2 result: PASS
|
||||
```
|
||||
|
||||
**Recommendation: ACCEPT M2 WITH NON-BLOCKING DEBT AND PROCEED TO M3.**
|
||||
|
||||
Directly, in the order asked:
|
||||
|
||||
| Question | Answer |
|
||||
| --- | --- |
|
||||
| Is it genuinely a single-user local storyteller? | **Yes.** No login, no accounts, no sessions; all four `/api/auth/*` routes 404 (§12). |
|
||||
| Is Ollama the only production inference backend? | **Yes.** No cloud provider code, no key, and public addresses are refused at the wire (§7). |
|
||||
| Are hosted/cloud/account/scripting surfaces actually removed? | **Yes** — removed, not hidden. 52 API routes → 36; `/api/auth`, `/api/analytics`, `/api/scripts` gone entirely (§3, §12). |
|
||||
| Does same-host Ollama still work? | **Yes** (§5.3–5.4). |
|
||||
| Does trusted-LAN Ollama still work? | **Yes**, against the real second machine (§6.2). |
|
||||
| Does HTTPS/private-CA still work with verification on? | **Yes** — `TLS verified, peer CN = inference.lan`, all four clients on the shared trust context, no bypass exists (§6.1, §8). |
|
||||
| Does it still run with Internet blocked? | **Yes** (§5.1, §6.1). |
|
||||
| Did any M1 behaviour regress? | **Yes — two, both found by this review and both fixed** (§A.1). |
|
||||
| Should the project proceed to M3? | **Yes.** |
|
||||
| Blockers before M3? | **None.** |
|
||||
|
||||
## A.1 Two M1 regressions that M2 shipped, and this review caught
|
||||
|
||||
These are the most important findings and they are not buried.
|
||||
|
||||
**1. The memory bank was silently broken.** M2 removed
|
||||
`Settings.api_key_plain`, but `memorybank`'s two provider factories still read
|
||||
it. Summaries and embeddings raised `AttributeError` inside a fire-and-forget
|
||||
background task — no user-visible error, no log a player would see, no failing
|
||||
test. **All 604 tests passed with the memory bank dead** (§9.1).
|
||||
|
||||
**2. The configurable model timeout never reached the turn engine.** The
|
||||
setting was stored, validated, exposed in the API and rendered in the UI, and
|
||||
then not passed to the provider. M2's own exit criterion — "the model timeout is
|
||||
no longer an undocumented hardcoded 120-second limitation" — was half met: the
|
||||
constant had moved but the setting was inert (§9.2).
|
||||
|
||||
A third, smaller: `requirements.lock` still pinned `quickjs`, `psycopg` and
|
||||
`cryptography`, so the documented setup path would have reinstalled all three
|
||||
(§9.3).
|
||||
|
||||
All three are corrected in the working tree, with tests that would have caught
|
||||
the first two. **The M2 commit `8c65ae9` does not contain these fixes**; the tree
|
||||
is six files ahead of it and needs a follow-up commit. Every runtime result in
|
||||
the baseline report was produced by an image built from the fixed tree.
|
||||
|
||||
**What this says about the milestone is more useful than the defects
|
||||
themselves:** a subtractive milestone's risk is not what it deletes, it is what
|
||||
still reaches for the deleted thing from a code path no test exercises. Both
|
||||
defects were in *background* or *plumbing* paths. §K returns to this.
|
||||
|
||||
---
|
||||
|
||||
# B. Repository and provenance
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `m2-local-only-surface` |
|
||||
| HEAD | `8c65ae99deda49b22f415d5868987e00bbcb173c` (signed, `%G? = G`) |
|
||||
| M2 commit | `8c65ae9`, one commit, parent `1a28a9a` |
|
||||
| M1 starting point | `1a28a9a` (post-M1 planning corrections) |
|
||||
| Upstream ancestry | `d72f7c1…` is an ancestor of HEAD; 172 upstream commits reachable |
|
||||
| `LICENSE` | byte-identical to upstream |
|
||||
| Private data committed | none (§1) |
|
||||
| Working tree | **six files modified** by this review (§A.1) |
|
||||
|
||||
No unrelated upstream re-import occurred. `PROVENANCE.md` records what M2
|
||||
removed and what it retained; `DEVELOPMENT.md` documents the new surface.
|
||||
|
||||
---
|
||||
|
||||
# C. Implementation inventory
|
||||
|
||||
**94 files changed, +1 395 −6 578.** Three added, twenty-four deleted, sixty-seven
|
||||
modified. Grouped by purpose:
|
||||
|
||||
### 1. Single-user / account removal
|
||||
Deleted `routers/auth.py`, `cleanup.py` (guest-retention sweeper),
|
||||
`accesslog.py`, `security.py` (session signing + key encryption). `auth.py` cut
|
||||
from 259 to 74 lines: it now resolves one implicit local user and nothing else.
|
||||
Frontend: the `/auth/me` bootstrap, `AuthModal`, guest nudge, log-in/sign-up/
|
||||
log-out controls, and the 401-and-retry dance in `api.js`.
|
||||
|
||||
### 2. Analytics removal
|
||||
Deleted `analytics.py`, `routers/analytics.py`, `pages/Analytics.jsx`, the
|
||||
`trackPageview` beacon, the `ApiErrorMiddleware` that fed the error tally, the
|
||||
lifespan flusher, and every `record_event` call site. Three model classes
|
||||
unmapped.
|
||||
|
||||
### 3. Hosted database / deployment removal
|
||||
`render.yaml` deleted. `database.py` reduced to SQLite only —
|
||||
`AIDND_DATABASE_URL`, `DATABASE_URL`, the psycopg URL normaliser and the
|
||||
serverless pre-ping are gone. `psycopg[binary]` removed.
|
||||
|
||||
### 4. Cloud-provider removal
|
||||
`_OPENROUTER_HOST`, `_PREFERRED_UPSTREAM`, `_apply_provider_routing`,
|
||||
`_apply_reasoning_budget`, the `Authorization` header, the `api_key` field and
|
||||
its Fernet encryption. `cryptography` removed.
|
||||
|
||||
### 5. Ollama-only configuration
|
||||
`OpenAICompatibleProvider(endpoint_url, model, api_mode, read_timeout)` — no
|
||||
key, no reasoning budget. `ProviderConfig`/`resolve_provider_config` deleted
|
||||
outright; callers read `Settings` directly.
|
||||
|
||||
### 6. Endpoint policy — **the only addition**
|
||||
`backend/app/endpoints.py` (179 lines) replaces `netguard.py`, inverting its
|
||||
rule (§F).
|
||||
|
||||
### 7. Trusted-LAN / TLS preservation
|
||||
M1's `tlstrust.py` untouched. All four HTTP clients still pass
|
||||
`verify=tlstrust.ssl_context()`; `test_tls_trust.py` enforces it by AST walk.
|
||||
|
||||
### 8. QuickJS / scripting removal
|
||||
`app/scripting/` (3 files), both script routers, `Script`/`AdventureScript`
|
||||
models, the `scenario_scripts` table, script schemas, script export/import,
|
||||
`Scripts.jsx`, `ScriptEditor.jsx`, `ScriptsPanel.jsx`, `StatusDrawer.jsx`,
|
||||
`ScriptReport`. `quickjs` removed; CodeMirror removed from npm.
|
||||
|
||||
### 9. Hosted-policy removal
|
||||
`limits.py` cut from 345 to ~150 lines: rate limiting, `X-Forwarded-For` client
|
||||
IP, login throttling and per-user quotas gone. **Kept**: request body ceiling,
|
||||
per-adventure row caps, import list caps.
|
||||
|
||||
### 10. Settings simplification
|
||||
Removed: API key field, "Remove key" button, demo banner, reasoning budget.
|
||||
Added: Ollama endpoint help text stating the policy, and a model-timeout field.
|
||||
|
||||
### 11. Model timeout
|
||||
`CONNECT_TIMEOUT = 10`, `DEFAULT_READ_TIMEOUT = 300`, `EMBED_READ_TIMEOUT = 60`,
|
||||
plus `Settings.model_timeout_seconds` (migration 77, default 300, bounded
|
||||
30–3600).
|
||||
|
||||
### 12. Loopback protections
|
||||
`docker-compose.yml` publishes `127.0.0.1:8000:8000`; `--proxy-headers` dropped;
|
||||
`AIDND_CORS_ORIGINS="*"` now refuses to start; an unknown `/api/...` path 404s
|
||||
instead of returning the SPA with status 200.
|
||||
|
||||
### 13. Tests
|
||||
Two new files (63 collected). Eight files converted from JS instrumentation to
|
||||
the world-state engine. Six files retired with their subsystems.
|
||||
|
||||
### 14. Documentation
|
||||
`DEVELOPMENT.md` gained the endpoint policy and the diagnostics table;
|
||||
`.env.example` cut from 110 lines to 24; `PROVENANCE.md` records M2.
|
||||
|
||||
## Deliberately retained
|
||||
|
||||
| Retained | Why |
|
||||
| --- | --- |
|
||||
| `users` table and `user_id` foreign keys | The M2 brief permits it. Removing them means a migration across most of the schema to delete a column that costs nothing. One row; nothing creates a second; no request carries an identity. |
|
||||
| Five inert tables, six inert columns | A destructive migration would risk an existing campaign database for tidiness. Unmapped or written-empty; nothing reads them (§13). |
|
||||
| `providers/openai_compatible.py` name | It speaks OpenAI's *protocol* to Ollama. Renaming would churn a file M3 does not touch, for no behaviour change. |
|
||||
| AI Chat scratchpad | Not hosted-only. It is a local tool for checking a model or prompt, and the endpoint policy constrains where it can talk. |
|
||||
| RPG world-state engine | M5's scope. M2 additionally now *depends* on it as test instrumentation (§K.2). |
|
||||
| Destructive Undo, no Redo | M3's scope, explicitly out of M2. |
|
||||
|
||||
---
|
||||
|
||||
# D. Single-user / hosted-account removal
|
||||
|
||||
| Classification | Contents |
|
||||
| --- | --- |
|
||||
| **Code removed** | `routers/auth.py`, `cleanup.py`, `accesslog.py`, `security.py`, 185 of 259 lines of `auth.py`, `AuthModal`, the account nav block, the session-retry logic |
|
||||
| **Code retained but inert** | none in this area |
|
||||
| **Schema retained for compatibility** | `users` + `user_id` FKs (active but single-row); `users.demo_turns_used`/`_date`; `settings.api_key` |
|
||||
| **Production functionality still active** | **none** |
|
||||
|
||||
* **Can a local user meet a login requirement?** No. There is no login UI, no
|
||||
session cookie, and `get_current_user` always succeeds.
|
||||
* **Are account APIs reachable?** No — 404 on all four (§12).
|
||||
* **Are hosted account modules imported at runtime?** No; the files do not exist.
|
||||
* **Does the database retain user identifiers?** Yes, one row.
|
||||
* **Merely internal now?** Yes, and documented as such in `auth.py`'s docstring
|
||||
and `PROVENANCE.md`.
|
||||
* **Did avoiding the schema rewrite reduce risk?** **Materially.** `user_id`
|
||||
appears on scenarios, adventures, settings and memories, and the ownership
|
||||
filters run through the story-tree and memory queries M3 will modify. A
|
||||
schema rewrite would have put a migration under those queries in the same
|
||||
milestone that removed accounts — two risky changes entangled. Keeping the
|
||||
column made M2 a deletion rather than a redesign.
|
||||
|
||||
No active hosted-account functionality remains.
|
||||
|
||||
---
|
||||
|
||||
# E. Analytics / telemetry
|
||||
|
||||
**Removed:** the collection module, the routes, the beacon, the Visitors page and
|
||||
nav link, the error-tally middleware, the batch flusher, and every call site.
|
||||
**Retained inert:** three tables, unmapped, never opened.
|
||||
|
||||
No destructive migration was performed because dropping tables from a live
|
||||
campaign database buys nothing and can fail.
|
||||
|
||||
**Verified at runtime, not asserted:** across a full offline session of 5 131
|
||||
packets — settings, campaign creation, four turns, embeddings, a fork, two
|
||||
induced failures, a restart — there were **zero non-loopback unicast packets**
|
||||
and no TCP connection opened outside loopback (§5.9). The only DNS names that
|
||||
appear are the ones my own isolation probe deliberately looked up, all within
|
||||
5 ms of the capture starting and 45 seconds before the storyteller received its
|
||||
first request, plus `ollama.com` from the Ollama service itself.
|
||||
|
||||
Analytics is removed, not merely hidden: collection does not occur.
|
||||
|
||||
---
|
||||
|
||||
# F. Hosted database / deployment, and the endpoint policy
|
||||
|
||||
## Database
|
||||
|
||||
Postgres, Neon, `psycopg`, the URL normaliser and `render.yaml` are gone;
|
||||
production is SQLite at `AIDND_DB_PATH`. An existing M1 database opens
|
||||
unchanged — the offline run used a volume carrying campaigns created before the
|
||||
fixes, survived an image replacement and a restart with digest
|
||||
`157c882c60579617` unchanged (§5.8). Story-tree persistence is unaffected;
|
||||
branch count, memory count and context inspection all survive restart.
|
||||
|
||||
One Postgres-shaped abstraction is retained: `migrations.py` and `tools/dbmeter.py`
|
||||
carry comments and a `_for_dialect` helper for SQL that differed between
|
||||
SQLite and psycopg. Removing it would churn the migration history for no gain.
|
||||
|
||||
## Endpoint policy — the design centre of M2
|
||||
|
||||
`endpoints.py` **inverts** the rule it replaces. `netguard.py` refused *private*
|
||||
addresses, to stop a hosted visitor making the server fetch an internal service.
|
||||
The threat here is the opposite: the user is trusted, and what must not happen is
|
||||
the story reaching the public Internet. So the new rule refuses *public*
|
||||
addresses.
|
||||
|
||||
The line is drawn by **address against an explicit allowlist of networks**, not
|
||||
by hostname and not by asking `ipaddress` what it thinks is private:
|
||||
|
||||
```text
|
||||
127.0.0.0/8 10.0.0.0/8 172.16.0.0/12 192.168.0.0/16
|
||||
169.254.0.0/16 100.64.0.0/10 ::1/128 fc00::/7 fe80::/10
|
||||
```
|
||||
|
||||
Every address a name resolves to must be in one. Applied twice: on save, for a
|
||||
good error; and before every outbound request, because DNS moves.
|
||||
|
||||
Measured behaviour (§7): loopback v4 and v6, `localhost`, all three RFC1918
|
||||
ranges, link-local, CGNAT, IPv6 ULA and a Docker-internal name are **accepted**;
|
||||
four cloud providers, public IPv4 and IPv6, `0.0.0.0` and a wrong scheme are
|
||||
**refused**, each with a reason naming what to do instead.
|
||||
|
||||
### What it guarantees, and what it does not
|
||||
|
||||
**Guarantees.** No request leaves for an address outside those networks, from any
|
||||
of the four clients, whatever is stored. Demonstrated against a **hand-edited
|
||||
SQLite row** — the settings row was rewritten to `openrouter.ai` with `sqlite3`,
|
||||
bypassing the API entirely, and the next turn was refused at the wire (§5.7).
|
||||
That is the strongest available form of this evidence.
|
||||
|
||||
**Does not guarantee.** The allowlist is *address* scope, not *ownership* scope:
|
||||
if a hostile host sits on the user's own LAN, the policy permits it — as it must,
|
||||
since that is what "trusted LAN" means. A rebinding window exists in principle
|
||||
between the policy's `getaddrinfo` and httpx's own connect; closing it would mean
|
||||
pinning the resolved address into the connection, which is a larger change than
|
||||
M2 warranted. Neither is a v1 concern: both require an attacker already inside
|
||||
the trusted network.
|
||||
|
||||
**One design finding worth flagging** (§K.1): writing this rule around
|
||||
`ipaddress.is_private` / `is_reserved` — the obvious approach — is wrong in a way
|
||||
that is easy to ship. Python classifies IPv6 loopback `::1` as *reserved*, so
|
||||
that version refused `http://[::1]:11434/v1`, an ordinary same-host endpoint. It
|
||||
also calls the documentation ranges *private*, so they would have been allowed.
|
||||
The explicit CIDR list exists because of that, and the reasoning is recorded in
|
||||
the module docstring.
|
||||
|
||||
---
|
||||
|
||||
# G. Ollama-only provider
|
||||
|
||||
**Providers in production code: one.** `OpenAICompatibleProvider`, speaking
|
||||
OpenAI's `/v1` protocol to Ollama.
|
||||
|
||||
* Provider options visible to the user: **none**. There is no selector, because
|
||||
there is nothing to select between.
|
||||
* OpenAI / OpenRouter / Groq / cloud APIs: **gone** — constants, routing,
|
||||
attribution and the reasoning-budget parameter.
|
||||
* Cloud API-key fields: **gone** from the model API, the UI and the request
|
||||
headers. The column is inert.
|
||||
|
||||
**The module name is not the question.** The question the brief poses —
|
||||
*can production use still send story content to a public inference service
|
||||
through normal configuration?* — was tested directly rather than reasoned about:
|
||||
saving a cloud URL is refused with HTTP 400 (§7), and a cloud URL written
|
||||
straight into the database is refused at request time (§5.7). No.
|
||||
|
||||
---
|
||||
|
||||
# H. Trusted-LAN and TLS regression
|
||||
|
||||
Re-verified after M2 on the **real second physical machine** used in M1, over
|
||||
HTTPS with a private CA (§6).
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| Endpoint class | `https://inference.lan:8443/v1`, non-loopback, private CA |
|
||||
| Connection | succeeded; diagnostics listed both installed models |
|
||||
| Certificate validation | performed and passed |
|
||||
| Hostname verification | performed — `peer CN = inference.lan` |
|
||||
| CA trust | the machine's own CA store, unioned with certifi |
|
||||
| Narration | 3 turns, 4.0–10.7 s |
|
||||
| Model listing/discovery | via the same trust path |
|
||||
| Embeddings | 1 memory embedded through the LAN host |
|
||||
| Settings / Test Connection | same trust path |
|
||||
| Retry | new take filed, both attempts retained (§6.4) |
|
||||
| Restart and resume | digest `66cc82944f8b6eef` unchanged; endpoint and timeout preserved |
|
||||
| Capture | 308 packets to the approved host, 364 loopback, **0 elsewhere, 0 DNS** |
|
||||
|
||||
Every model request went to the configured host: 4 `chat/completions`, 1
|
||||
`embeddings`, nothing else.
|
||||
|
||||
**No `verify=False`, insecure fallback or "ignore certificate errors" path
|
||||
exists.** The only textual match for such a thing in the tree is prose in
|
||||
`tlstrust.py` saying it deliberately does not exist. All four clients use the
|
||||
same context, enforced by an AST-walking test that also fails if a *fifth*
|
||||
client appears (§8). No client uses different trust behaviour.
|
||||
|
||||
---
|
||||
|
||||
# I. QuickJS and scripting removal
|
||||
|
||||
Removed: the QuickJS dependency, `app/scripting/` entire, both routers, the
|
||||
`onInput`/`onModelContext`/`onOutput` hooks in the turn engine, the `Script` and
|
||||
`AdventureScript` models, script schemas, script export/import (bundles that
|
||||
carry `scripts` now import with the story intact and the scripts ignored), the
|
||||
Scripts page, the script editor, the in-play Scripts panel, the script-state
|
||||
drawer and the Insights script report.
|
||||
|
||||
**Imported content cannot execute JavaScript.** There is no engine to execute it
|
||||
in — `import app.scripting` raises `ImportError` and no module imports `quickjs`
|
||||
(§12). A scenario or bundle carrying script source is data that nothing reads.
|
||||
|
||||
## Test-coverage accounting
|
||||
|
||||
This is where a subtractive milestone can hide damage behind a green number, so
|
||||
it is spelled out. Of 88 node IDs that disappeared, **4 are renames** and 84 were
|
||||
genuinely retired (§4).
|
||||
|
||||
| Category | Count | Verdict |
|
||||
| --- | ---: | --- |
|
||||
| Whole files whose subject was removed (`analytics`, `guest_cleanup`, `accesslog`, `ratelimit_hardening`, `reasoning_param`) | 64 | Obsolete product behaviour intentionally deleted. No coverage lost — the behaviour is gone. |
|
||||
| `test_netguard.py` | 5 | **Replaced**, not lost: `test_endpoint_policy.py` (31 collected) covers the inverted rule far more thoroughly. |
|
||||
| `test_chat.py` demo-key pinning and power-user gate | 8 | Subject removed. The file keeps and extends what the page still does. |
|
||||
| `test_prompt_caching.py` OpenRouter routing | 4 | Subject removed. The prompt-layout and usage tests, which are the file's point, are untouched. |
|
||||
| `test_memory_rewrite.py` Postgres DSN masking | 2 | Subject removed with Postgres. |
|
||||
| `test_branch_clause.py::test_user_scripts_are_handed_the_path` | 1 | **Equivalent coverage retained.** It asserted the scripting history API saw the branch path; the test directly above it asserts the same invariant one layer down on `history.story_actions`/`count`/`tail`, which is where the branch clause lives. |
|
||||
| Instrumentation conversion, 8 files | 0 retired | **Coverage retained in full.** A JS counter was measuring the *state snapshot and rollback machinery*, which M2 does not touch. The counter moved to the world-state engine; the assertions are unchanged in substance. |
|
||||
|
||||
**No meaningful coverage was lost.** One area gained a great deal: endpoint
|
||||
policy went from 5 tests of the opposite rule to 31 of the current one.
|
||||
|
||||
---
|
||||
|
||||
# J. Settings surface, and the model timeout
|
||||
|
||||
## Settings after M2
|
||||
|
||||
**Remains:** Ollama endpoint (with help text stating the policy), narrator model,
|
||||
embedding model, temperature, max output tokens, context budget, API mode,
|
||||
narrator prompt, memory-bank capacity and top-k, **model timeout**, Test
|
||||
connection, and the provider debug log.
|
||||
|
||||
**Removed:** the API key field and its "Remove key" button, the demo-key banner,
|
||||
the reasoning budget, and the "OpenAI-compatible" framing on the endpoint field.
|
||||
|
||||
**Is it correct for v1?** Yes. Every field maps to something the product does,
|
||||
and nothing offers a capability the product refuses. It is not *polished* — that
|
||||
is M8 — but it is not misleading, which is what M2 asked for.
|
||||
|
||||
One onboarding gap survives from M1 and is not new: `Settings.model` defaults to
|
||||
`""`, so a fresh install cannot generate a turn until a model is named, and
|
||||
nothing prompts. M2 improves the *diagnosis* — the connection test now warns when
|
||||
the endpoint is reachable but has no model by the configured name — without
|
||||
fixing the onboarding, which is M8's.
|
||||
|
||||
## Model timeout
|
||||
|
||||
| | Value |
|
||||
| --- | --- |
|
||||
| Connect timeout | 10 s, constant — a wrong address should fail fast |
|
||||
| Generation/read timeout | `Settings.model_timeout_seconds`, default **300 s** |
|
||||
| Allowed range | 30–3600, enforced by the schema (422 outside) |
|
||||
| UI | a numeric field with help text about cold loads |
|
||||
| Config source | database row, not an environment variable |
|
||||
| Embeddings | separate 60 s constant — short calls that never cold-load |
|
||||
| Connection test | separate 15 s constant — a listing, not a generation |
|
||||
|
||||
**Demonstrated no longer a fixed 120 s constant:** `test_the_timeout_is_configurable`
|
||||
sets 900 through the API and reads it back;
|
||||
`test_the_configured_timeout_reaches_every_generating_client` drives the turn
|
||||
endpoint, the chat endpoint and the summariser factory with a recording provider
|
||||
and asserts the configured value reaches each; `test_an_unusable_timeout_is_refused`
|
||||
rejects 0, −1, 29 and 3601. The offline run used 600 s and the LAN run used 600 s,
|
||||
both read back after restart (§5.8, §6.5).
|
||||
|
||||
The retained fixed timeouts are the three above, each named and each justified by
|
||||
the shape of the call.
|
||||
|
||||
---
|
||||
|
||||
# K. Important surprises
|
||||
|
||||
Four. None invented for completeness.
|
||||
|
||||
## K.1 An address policy written the obvious way is wrong
|
||||
|
||||
Writing the endpoint rule around `ipaddress.is_private` / `is_global` /
|
||||
`is_reserved` — which is what the inherited `netguard.py` did and what any
|
||||
reviewer would expect — produces a rule that **refuses IPv6 loopback**
|
||||
(`::1` is classified reserved) and **accepts the documentation ranges** (they
|
||||
are classified private). The first was caught by a test asserting
|
||||
`http://[::1]:11434/v1` is allowed. This is worth knowing beyond this project:
|
||||
the standard library's classifications answer a different question than "is this
|
||||
on my own network".
|
||||
|
||||
## K.2 Removing scripting made the product depend on the world-state engine — in the tests
|
||||
|
||||
Eight test files used a QuickJS counter as instrumentation for the state
|
||||
rollback machinery. The natural replacement was the world-state engine, so those
|
||||
tests now assert rollback *through* the RPG delta protocol. **M5 plans to replace
|
||||
that protocol** with typed narrative-state events. M5 will therefore have to move
|
||||
this instrumentation a second time. It is cheap — the counter is one schema entry
|
||||
and one reply template in `tests/fakes.py` — but M5's brief should say so rather
|
||||
than discover it.
|
||||
|
||||
## K.3 The largest surviving subsystem is the one M2 could not touch
|
||||
|
||||
`migrations.py` is **1 128 lines**, the biggest file in the backend, and M2
|
||||
reduced it by one line while adding one. It carries the full history of a schema
|
||||
that M2 has now partly orphaned: migrations that create `scripts`,
|
||||
`analytics_daily` and `access_log`, and backfills for `state_after`. This is
|
||||
correct — migration history must not be rewritten — but a reviewer expecting the
|
||||
"trust and maintenance surface" to have shrunk proportionally should know that a
|
||||
tenth of the backend is history that only grows.
|
||||
|
||||
## K.4 Two defects hid in exactly the places a subtractive milestone cannot see
|
||||
|
||||
Both regressions (§A.1) were invisible to a 604-test green suite, for the same
|
||||
reason: the code that still reached for a removed thing sat in a path no test
|
||||
executed with real objects. The memory-bank factories are stubbed in every
|
||||
memory test; the timeout argument was never asserted. **The lesson is a testing
|
||||
one, not a coding one:** after removing an attribute, the cheapest useful test is
|
||||
one that constructs each consumer from a real object, and after adding a setting,
|
||||
one that proves it arrives. Both now exist.
|
||||
|
||||
---
|
||||
|
||||
# L. Acceptance matrix
|
||||
|
||||
| Requirement | Result | Evidence / notes |
|
||||
| --- | --- | --- |
|
||||
| Single-user, no login | PASS | §12 — four auth routes 404 |
|
||||
| Hosted account/guest/demo removed | PASS | §C.1, §D |
|
||||
| Hosted rate-limit policy removed | PASS | `limits.py` 345→~150 lines; resource bounds kept |
|
||||
| Analytics/Visitors removed | PASS | §E |
|
||||
| No telemetry generated | PASS | §5.9 — 0 non-loopback unicast in 5 131 packets |
|
||||
| Render path removed | PASS | `render.yaml` deleted |
|
||||
| Neon path removed | PASS | §F |
|
||||
| Postgres runtime removed | PASS | `psycopg` gone from requirements and closure |
|
||||
| SQLite M1 data still works | PASS | §5.8 — digest survives image replacement and restart |
|
||||
| Ollama-only production provider | PASS | §G |
|
||||
| Cloud API keys removed | PASS | §12 — absent from the API, unsettable |
|
||||
| Public inference endpoints blocked | PASS | §7, §5.7 — including against a hand-edited DB |
|
||||
| Loopback Ollama accepted | PASS | §7 — v4, v6 and `localhost` |
|
||||
| Trusted-LAN Ollama accepted | PASS | §6.2 |
|
||||
| TLS/private-CA support preserved | PASS | §6.1 |
|
||||
| TLS verification preserved | PASS | §8 — no bypass exists; AST-enforced |
|
||||
| QuickJS removed | PASS | §12 |
|
||||
| Executable campaign scripting removed | PASS | §I |
|
||||
| Settings narrowed to v1 model | PASS | §J |
|
||||
| Connection diagnostics appropriate | PASS | five distinguished failure kinds + a model-not-found warning |
|
||||
| Hardcoded 120 s limitation resolved | **PASS after the §9.2 fix** | Was PARTIAL as committed: the setting existed but was inert |
|
||||
| Supported launch paths loopback-only | PASS | §11 — all four paths, test-enforced |
|
||||
| Offline operation preserved | PASS | §5.1, §5.4 |
|
||||
| Story/tree/retry preserved | PASS | §6.4 — `take_count 2`, both attempts retained |
|
||||
| Context inspection preserved | PASS | §5.8 |
|
||||
| Memory isolation preserved | **PASS after the §9.1 fix** | §5.6 — negative control holds. Was FAIL as committed: the bank was dead |
|
||||
| No M3 work started | PASS | Undo still destructive, no Redo, no head cursor |
|
||||
|
||||
Two rows are "PASS after the fix". As committed at `8c65ae9` they were **FAIL**
|
||||
and **PARTIAL**. Both fixes are in the working tree; neither is a blocker,
|
||||
because both are corrected and tested — but the commit itself does not meet the
|
||||
criteria, which is why §N recommends a follow-up commit before M3 begins.
|
||||
|
||||
---
|
||||
|
||||
# M. V1 acceptance-test mapping
|
||||
|
||||
| ID | Result | Notes |
|
||||
| --- | --- | --- |
|
||||
| **A01** offline startup | PASS | §5.1, §5.4 |
|
||||
| **A02** loopback default | PASS | §5.2, §6.1, §11 |
|
||||
| **A03** no cloud API key | PASS | Stronger than at M1: the field no longer exists |
|
||||
| **A04** restart persistence | PASS | §5.8, §6.5 |
|
||||
| **A05** failed model call integrity | PASS | §5.7 — accepted AI turns held at 3 across two induced failures |
|
||||
| **A06** trusted-LAN Ollama | PASS | §6 — real second machine, HTTPS, private CA |
|
||||
| **H01** no unexpected outbound | PASS | §5.9, §6.6 |
|
||||
| **H02** no telemetry | PASS | §E |
|
||||
| **H03** no cloud provider required | **PASS, and now in its preferred form** | H03 says the *preferred* final v1 is "cloud provider controls are absent, not merely unused". M1 met the base condition; **M2 meets the preferred one.** |
|
||||
| **H04** model output cannot execute shell/tools | PASS, strengthened | The only execution surface was QuickJS, now removed. No shell, MCP or tool framework exists. |
|
||||
| **H10** local API/CORS behaviour | PASS, strengthened | `AIDND_CORS_ORIGINS="*"` refuses to start; an unknown `/api` path 404s instead of returning HTML 200 |
|
||||
| **H11** no first-use runtime download | PASS | §10.1, §10.2 |
|
||||
|
||||
Two acceptance tests deserve planning attention (§O): A05's wording, already
|
||||
corrected post-M1, and H10, which M2 has now exceeded in a way the test does not
|
||||
describe.
|
||||
|
||||
---
|
||||
|
||||
# N. Security review of the simplified product
|
||||
|
||||
## Resulting architecture
|
||||
|
||||
```text
|
||||
Browser (loopback only)
|
||||
│ same-origin; CSP names no remote origin
|
||||
▼
|
||||
Storyteller — FastAPI + SPA, bound 127.0.0.1:8000, no authentication
|
||||
│
|
||||
├─► SQLite on the local filesystem (the only persistence)
|
||||
│
|
||||
└─► exactly one approved Ollama endpoint (endpoints.py)
|
||||
same-host loopback OR a user-named address on their own network
|
||||
HTTP, or HTTPS verified against the machine's CA store
|
||||
```
|
||||
|
||||
## Every remaining way story text could leave the machine
|
||||
|
||||
Found by static search plus runtime confirmation, not by assuming dead code is
|
||||
unreachable:
|
||||
|
||||
| Path | Status |
|
||||
| --- | --- |
|
||||
| `providers/openai_compatible.py` — 3 clients | The intended egress. Policy-checked before every request, TLS-verified. |
|
||||
| `routers/settings.py` — 1 client | Model listing. Same policy, same trust context. |
|
||||
| `endpoints.py` — `getaddrinfo` | Resolves names for the policy. Opens no connection. |
|
||||
| Analytics / telemetry | **Gone.** |
|
||||
| Remote runtime assets | **Gone** since M1; re-verified (§10.2). |
|
||||
| Cloud provider constants | **Gone.** |
|
||||
| Hosted auth | **Gone.** |
|
||||
| Arbitrary endpoint use | **Refused** by address (§7). |
|
||||
| Scripting runtimes | **Gone.** |
|
||||
| Shell / MCP / tool frameworks | Never existed. |
|
||||
| Remote database | **Gone.** |
|
||||
|
||||
Only four modules in `backend/app` can open an outbound connection at all, and
|
||||
the only absolute remote URL left in backend code is a `SOURCE_URL` constant in
|
||||
`encoding.py` documenting where the vendored tokenizer table came from — never
|
||||
fetched (§8).
|
||||
|
||||
## Residual risk, stated plainly
|
||||
|
||||
1. The API is **unauthenticated by design**. Loopback binding is the whole
|
||||
control. Publishing the port defeats it; the code refuses the easy mistakes
|
||||
(wildcard CORS, the default compose mapping) and the documentation warns
|
||||
about the deliberate one.
|
||||
2. A hostile host **on the user's own LAN** is permitted by the endpoint policy,
|
||||
because that is what trusted-LAN means.
|
||||
3. The **Ollama service has its own network behaviour** (`ollama.com`), outside
|
||||
this codebase and inside the user's trust boundary. Unchanged from M1 and
|
||||
still worth settling before release.
|
||||
|
||||
---
|
||||
|
||||
# O. Planning-document recommendations
|
||||
|
||||
Reported, not applied. No planning document was edited by this task.
|
||||
|
||||
### 1.
|
||||
```text
|
||||
Document: BUILD-MILESTONES.md, M5
|
||||
Recommended change: Note that eight test files now use the world-state delta
|
||||
protocol as instrumentation for state rollback, and that
|
||||
replacing that protocol means moving the instrumentation
|
||||
(one schema entry and one helper in tests/fakes.py).
|
||||
Why M2 evidence: §K.2 — the QuickJS counter those tests used was replaced
|
||||
with the RPG engine, which M5 plans to replace in turn.
|
||||
Urgency: later (before M5 is briefed)
|
||||
```
|
||||
|
||||
### 2.
|
||||
```text
|
||||
Document: SECURITY-THREAT-MODEL.md
|
||||
Recommended change: Record the endpoint policy as implemented — an explicit
|
||||
allowlist of networks, applied on save and before every
|
||||
request — together with what it does not guarantee: a
|
||||
hostile host on the trusted LAN, and the rebinding window
|
||||
between the policy's resolution and the connection.
|
||||
Why M2 evidence: §F. The threat model predates the policy and describes the
|
||||
inherited SSRF guard's opposite rule.
|
||||
Urgency: before M3
|
||||
```
|
||||
|
||||
### 3.
|
||||
```text
|
||||
Document: V1-ACCEPTANCE-TESTS.md, H10
|
||||
Recommended change: H10 currently reads only "privileged local APIs do not
|
||||
allow arbitrary wildcard cross-origin writes". M2 goes
|
||||
further: a wildcard origin makes the app refuse to start,
|
||||
and an unknown /api path 404s rather than returning the SPA
|
||||
with status 200. State both as pass conditions.
|
||||
Why M2 evidence: §11. Both were defects found and fixed during M2; without
|
||||
a test that names them they can regress unnoticed.
|
||||
Urgency: later
|
||||
```
|
||||
|
||||
### 4.
|
||||
```text
|
||||
Document: TECHNICAL-DESIGN.md §5.1
|
||||
Recommended change: Mark items 3 and 4 of the hardening list resolved. Item 3
|
||||
(hosted/auth/analytics/Postgres/cloud/QuickJS) is done in
|
||||
full; item 4 (endpoint validation reflecting this product's
|
||||
threat model) is done by endpoints.py.
|
||||
Why M2 evidence: §C, §F. The document currently marks both open — M1 left
|
||||
them so.
|
||||
Urgency: before M3
|
||||
```
|
||||
|
||||
### 5.
|
||||
```text
|
||||
Document: DECISIONS/ — a new ADR
|
||||
Recommended change: Record the endpoint policy as a decision: address-based
|
||||
allowlist rather than hostname matching, deny by default,
|
||||
enforced at save and at request time, never traded against
|
||||
TLS verification. Include the ipaddress-classification
|
||||
finding as the reason the CIDRs are spelled out.
|
||||
Why M2 evidence: §F, §K.1. This is a load-bearing security decision that
|
||||
currently exists only as a module docstring.
|
||||
Urgency: before M3
|
||||
```
|
||||
|
||||
### 6.
|
||||
```text
|
||||
Document: SPECIFICATION.md
|
||||
Recommended change: None. M2 changed no product requirement; it removed
|
||||
capability the specification never asked for.
|
||||
Urgency: —
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# P. Technical debt after M2
|
||||
|
||||
| Item | Classification | Note |
|
||||
| --- | --- | --- |
|
||||
| **The three §9 fixes are uncommitted** | **pre-release — do first** | Six files ahead of `8c65ae9`. Until committed, the branch's HEAD contains a dead memory bank. |
|
||||
| `Settings.model` defaults to `""`; nothing prompts | M8 | Inherited from M1. Diagnosis improved, onboarding not. |
|
||||
| Ollama's own `ollama.com` lookup | pre-release | Outside this codebase, inside the product's claim about itself. |
|
||||
| Inert tables and columns | optional cleanup | Safe legacy remnants. A cleanup migration is cheap once the schema settles — after M3/M5, not before. |
|
||||
| `migrations.py` at 1 128 lines | optional cleanup | Dead/inert history, not active behaviour (§K.3). |
|
||||
| 4 pre-existing unused imports | optional cleanup | Present since M1; `pyflakes` is otherwise clean. |
|
||||
| Unused `request: Request` parameters in three routers | optional cleanup | Left by removing the rate limiter. Harmless. |
|
||||
| No frontend tests | M8 | None existed at M1 either. Lint and build only. |
|
||||
| `docs/*.html` links Google Fonts | M8 or pre-release | Upstream's project site; not served by the app. |
|
||||
| Rebinding window in the endpoint policy | optional | §F. Requires an attacker already on the trusted LAN. |
|
||||
|
||||
**Nothing here is a blocker for M3** except committing the fixes, which is
|
||||
housekeeping rather than engineering.
|
||||
|
||||
---
|
||||
|
||||
# Q. M3 readiness
|
||||
|
||||
M3 is the production history work: non-destructive Undo, Redo, active-head
|
||||
movement, divergence, and active-head export/import.
|
||||
|
||||
**Did M2 alter the chokepoints M3 will modify?** Barely, and helpfully:
|
||||
|
||||
| M3 chokepoint | M2's effect |
|
||||
| --- | --- |
|
||||
| `attempts.py` — snapshot/restore/rollback | `ATTEMPT_KEYS` lost `"script"`; `restore_state` and `snapshot_outcome` no longer carry script state. **One less shared state to move.** |
|
||||
| `tree.py` — placement, head, lineage | **Untouched.** |
|
||||
| `context/lineage.py`, `context/history.py` | **Untouched.** |
|
||||
| `routers/adventures/takes.py` — retry, takes, undo | Only the removal of `ScriptPipeline` arguments and the demo cap. `undo_turn` is byte-for-byte the inherited destructive implementation. |
|
||||
| `routers/adventures/turns.py` | Hook calls removed, so `generate_turn` has 4 parameters instead of 5 and no longer branches on script stop. **Simpler to reason about.** |
|
||||
| `bundle.py` — export/import | `scripts`/`scriptState` no longer exported; old bundles importing them are ignored gracefully. M3 adds the active-head field to a smaller format. |
|
||||
|
||||
**Is the Phase 0B undo/redo spike still applicable?** Yes. It touched three
|
||||
backend files — the tree, the attempt machinery and the takes router — and M2
|
||||
changed none of their logic. If anything the spike is easier to apply now: it no
|
||||
longer has to carry `script_state` alongside `world_state` through every
|
||||
rollback.
|
||||
|
||||
**Did the removals simplify or complicate M3?** Simplified, in three concrete
|
||||
ways: one shared state instead of two through the rollback paths; no demo-cap or
|
||||
rate-limit preconditions wrapped around the turn and retry endpoints; and a
|
||||
provider constructed from `Settings` directly rather than through a
|
||||
`ProviderConfig` indirection.
|
||||
|
||||
**Tests M3 should preserve or rewrite.** Preserve as the contract:
|
||||
`test_story_tree_baseline`, `test_retry_variants`, `test_take_state`,
|
||||
`test_take_parentage`, `test_attempt_siblings`, `test_branch_forking`,
|
||||
`test_delete_state`, `test_state_revert`, `test_bundle_v2`. **Note that these now
|
||||
carry the world-state instrumentation** (§K.2) — M3 must not mistake it for RPG
|
||||
coverage. `test_state_revert` and `test_delete_state` assert *destructive* undo
|
||||
semantics and will need rewriting when undo stops deleting; that is expected M3
|
||||
work, and those files should be rewritten rather than deleted.
|
||||
|
||||
**Anything to fix before history changes?** Only the uncommitted fixes. Nothing
|
||||
M2 discovered touches history semantics.
|
||||
|
||||
**Should M3 proceed as planned?** **Yes, unchanged in scope.** Two additions to
|
||||
its brief: it inherits a single shared state rather than two, and it must be told
|
||||
that the state-rollback tests are instrumented through the world-state engine.
|
||||
|
||||
---
|
||||
|
||||
# R. Final recommendation
|
||||
|
||||
```text
|
||||
Recommendation: ACCEPT M2 WITH NON-BLOCKING DEBT AND PROCEED TO M3
|
||||
```
|
||||
|
||||
**Reasons.**
|
||||
|
||||
1. Every M2 requirement is met, and the ones that matter most were tested at
|
||||
runtime rather than read: cloud endpoints refused against a hand-edited
|
||||
database, trusted-LAN HTTPS working against a real second machine with
|
||||
verification on, and captures showing zero packets outside loopback and the
|
||||
approved host.
|
||||
2. The product is materially simpler, not just smaller: 52 API routes → 36, ten
|
||||
environment variables → two, 6 Python packages and 21 npm packages gone,
|
||||
−2 666 backend lines, −1 284 frontend lines, and a 933 kB bundle down to
|
||||
395 kB. The trust surface shrank in the places that carry story data.
|
||||
3. H03's *preferred* v1 condition — cloud controls absent rather than unused —
|
||||
is now met, which M1 could not claim.
|
||||
4. No M1 capability is regressed **in the working tree**. Two were regressed in
|
||||
the commit; both are fixed and both now have the tests that would have caught
|
||||
them.
|
||||
5. M3's chokepoints are untouched or simplified, and the Phase 0B spike still
|
||||
applies.
|
||||
|
||||
**The one thing to do first** is commit the six-file fix (§A.1). Until then the
|
||||
branch head contains a silently broken memory bank, and anyone building from
|
||||
`8c65ae9` would inherit it.
|
||||
|
||||
The three planning corrections marked *before M3* in §O — the threat model, the
|
||||
technical-design hardening list, and a new ADR for the endpoint policy — are
|
||||
documentation of decisions already made, not new work, and can be done alongside
|
||||
the M3 brief.
|
||||
|
||||
---
|
||||
|
||||
# S. Closeout note — appended 2026-09-03
|
||||
|
||||
**Appended after the fact.** Sections A-R above were written on 2026-09-02 and
|
||||
are left as they were, including §A.1's and §P's statements that the fixes were
|
||||
uncommitted at the time. Those statements were true when written and are the
|
||||
reason this note exists rather than an edit.
|
||||
|
||||
## The one thing §R said to do first
|
||||
|
||||
§R closed with: *"The one thing to do first is commit the six-file fix (§A.1).
|
||||
Until then the branch head contains a silently broken memory bank."*
|
||||
|
||||
Done:
|
||||
|
||||
```text
|
||||
8652fe7cd84bca5173abb03b2a692f15fea8a98c M2 review: two regressions the green suite hid, and the reports
|
||||
```
|
||||
|
||||
The commit carries the six implementation/test/lockfile files **and** these two
|
||||
reports. `8c65ae9` remains the M2 implementation commit and was not amended or
|
||||
squashed, so the provenance distinction §R wanted — original implementation
|
||||
versus review-discovered correction — survives in the history.
|
||||
|
||||
Confirmed present in that commit, against `8c65ae9`:
|
||||
|
||||
| Defect | Fix as committed |
|
||||
| --- | --- |
|
||||
| §9.1 dead memory bank | `memorybank.py` — both provider factories stop reading the removed `Settings.api_key_plain`; `summary_provider` now passes `model_timeout_seconds` |
|
||||
| §9.2 inert timeout | `turns.py`, `chat.py` and the summariser factory pass `settings.model_timeout_seconds` into the provider |
|
||||
| §9.3 stale lock | `requirements.lock` drops `quickjs`, `psycopg`, `psycopg-binary`, `cryptography`, and the transitive `cffi` and `pycparser` |
|
||||
| — | `test_local_only_surface.py` gains the two regression tests: every factory built from a real `Settings` row, and the configured timeout arriving at each generating client |
|
||||
|
||||
The separate short timeout classes were kept, as §9.2 required: `CONNECT_TIMEOUT`
|
||||
(10 s), `EMBED_READ_TIMEOUT` (60 s) and the connection-test/model-list path are
|
||||
unchanged, and only the generation read timeout became configurable.
|
||||
|
||||
## Verification at closeout
|
||||
|
||||
Re-run on 2026-09-03 against the **committed** tree, working tree clean:
|
||||
|
||||
```text
|
||||
backend pytest tests/ -q 606 passed in 133.26s
|
||||
tests/test_local_only_surface.py 32 passed
|
||||
test_endpoint_policy + test_tls_trust + test_offline_assets
|
||||
47 passed
|
||||
frontend npm run lint 7 warnings, 0 errors, exit 0
|
||||
frontend npm run build 395.41 kB, exit 0
|
||||
root docker build exit 0
|
||||
image quickjs/psycopg/cryptography/cffi/pycparser absent (32 packages)
|
||||
```
|
||||
|
||||
The packet-capture exercises were not repeated; the corrected files do not touch
|
||||
the network boundary. See §15 of the baseline report.
|
||||
|
||||
## Planning recommendations from §O
|
||||
|
||||
Three were marked *before M3* and are now applied, in commit
|
||||
`Planning: record M2 closeout decisions`:
|
||||
|
||||
- **§O.2** — `SECURITY-THREAT-MODEL.md` §10A records the endpoint policy as
|
||||
implemented, with both residual limits stated; §71A item 5 is marked resolved;
|
||||
§77 notes the required defaults are now met.
|
||||
- **§O.4** — `TECHNICAL-DESIGN.md` §5.1 items 3 and 4 are marked done, and a new
|
||||
§5.2 records the M1/M2 production architecture as fact rather than intention.
|
||||
- **§O.5** — **ADR 011, *Local Inference Endpoint Policy***, records the
|
||||
decision, including the `ipaddress`-classification finding from §K.1 as the
|
||||
reason the CIDRs are spelled out.
|
||||
|
||||
The remaining three are also applied, ahead of the *later* urgency §O gave them,
|
||||
since they are one-paragraph edits: **§O.1** as an M5 note in
|
||||
`BUILD-MILESTONES.md` about the world-state instrumentation, and **§O.3** as a
|
||||
strengthened H10 in `V1-ACCEPTANCE-TESTS.md`. **§O.6** correctly asked for no
|
||||
change to `SPECIFICATION.md`, and none was made.
|
||||
|
||||
Two additions beyond §O, both drawn from evidence in this report:
|
||||
|
||||
- an M6 note in `BUILD-MILESTONES.md` requiring background memory failure to be
|
||||
observable and at least one real provider-construction path to be tested —
|
||||
§A.1 and §9.1 are the argument for it;
|
||||
- **H12, *Inference Endpoint Enforcement***, in `V1-ACCEPTANCE-TESTS.md`, whose
|
||||
fourth pass condition is the database-edited-behind-the-API case this review
|
||||
demonstrated at runtime in §F.
|
||||
|
||||
`BUILD-MILESTONES.md` also gained M2's `## Status: COMPLETE` block, matching M1's.
|
||||
|
||||
## Not done in this closeout
|
||||
|
||||
M3 was not begun. Undo still deletes, and there is still no Redo. The debt table
|
||||
in §P is unchanged apart from its first row, which this note closes.
|
||||
@@ -0,0 +1,141 @@
|
||||
# CaoRuiming/ai-adventure — Static Architecture Analysis
|
||||
|
||||
**Historical status:** Static Phase 0A analysis. Phase 0B promoted ai-adventure to the primary implementation reference for state/event/head/checkpoint semantics, but not the production base.
|
||||
|
||||
**Project name in repository docs:** Local Adventure Engine
|
||||
**Repository:** https://github.com/CaoRuiming/ai-adventure
|
||||
**Date reviewed:** 2026-09-01
|
||||
**Disposition:** Finalist #3; strongest state/privacy reference, possible core candidate.
|
||||
|
||||
## Architectural fit
|
||||
|
||||
This project most closely matches the desired trust boundary:
|
||||
|
||||
> The model proposes narration/events; the application validates and commits authoritative state.
|
||||
|
||||
Its architecture separates:
|
||||
- authored content,
|
||||
- runtime state and pure reducers,
|
||||
- SQLite storage/migrations,
|
||||
- local lore indexing/retrieval,
|
||||
- deterministic bounded context construction,
|
||||
- model provider,
|
||||
- application turn logic,
|
||||
- CLI presentation.
|
||||
|
||||
That separation makes it especially valuable even if it is not the final fork.
|
||||
|
||||
## Turn/commit model
|
||||
|
||||
The documented flow:
|
||||
|
||||
1. load/replay state,
|
||||
2. synchronize lore,
|
||||
3. build deterministic bounded context,
|
||||
4. call local model,
|
||||
5. parse a structured turn proposal,
|
||||
6. validate proposed events,
|
||||
7. apply events in memory,
|
||||
8. atomically append turn/events and move the session head,
|
||||
9. display narration only after commit.
|
||||
|
||||
This is the strongest candidate design for “the model is not the database.”
|
||||
|
||||
## Persistence and recovery
|
||||
|
||||
The project documents:
|
||||
- parent-linked turn history,
|
||||
- append-only state events,
|
||||
- cached reconstructed state,
|
||||
- undo by moving session head,
|
||||
- named checkpoints,
|
||||
- restore,
|
||||
- branching into another session that shares ancestors,
|
||||
- replayable state.
|
||||
|
||||
This satisfies the conceptual checkpoint/branch requirement better than a destructive chat log.
|
||||
|
||||
Potential mismatch:
|
||||
- branches are represented as sessions rather than necessarily one unified visual story tree.
|
||||
- export behavior and cross-branch navigation should be tested for the browser product.
|
||||
|
||||
## Lore and long memory
|
||||
|
||||
The project currently favors deterministic local retrieval:
|
||||
- Markdown lore,
|
||||
- SQLite FTS5/fallback,
|
||||
- bounded context,
|
||||
- summaries.
|
||||
|
||||
It intentionally avoids an embedding/vector dependency in the initial architecture.
|
||||
|
||||
This is attractive for privacy and auditability, but the target project likely also wants optional local semantic retrieval through Ollama for:
|
||||
- old story events,
|
||||
- large imported reference/inspiration libraries.
|
||||
|
||||
The deterministic lexical layer should still be considered as part of a hybrid retriever.
|
||||
|
||||
## Privacy/security fit
|
||||
|
||||
This is the strongest static privacy design among the finalists.
|
||||
|
||||
The project documentation explicitly addresses:
|
||||
- local data directory,
|
||||
- loopback model endpoint by default,
|
||||
- warning for non-loopback endpoints,
|
||||
- no telemetry/cloud account,
|
||||
- no MCP,
|
||||
- no executable plugins,
|
||||
- no shell tools,
|
||||
- parameterized SQL,
|
||||
- bounded imports,
|
||||
- path traversal/symlink restrictions,
|
||||
- local world files treated as data.
|
||||
|
||||
The default provider is LM Studio rather than Ollama, but the provider boundary appears intentionally small.
|
||||
|
||||
## Tests
|
||||
|
||||
Project documentation reports an offline test suite that grew during implementation (later milestone notes report 74 tests). Phase 0B should run the actual current suite and treat it as authoritative.
|
||||
|
||||
## Major gaps for target product
|
||||
|
||||
- terminal UI,
|
||||
- LM Studio rather than Ollama as documented primary provider,
|
||||
- no browser API/UI,
|
||||
- no current media system,
|
||||
- no semantic embedding retrieval,
|
||||
- authored entity/event model may be more rigid than freeform narrative state,
|
||||
- likely more front-end work than either browser finalist.
|
||||
|
||||
## Best reuse case
|
||||
|
||||
Even if it is not the production base, reuse its architectural rules:
|
||||
|
||||
- append-only authoritative events,
|
||||
- model proposals never direct state writes,
|
||||
- validate before commit,
|
||||
- commit narration and state atomically,
|
||||
- deterministic replay,
|
||||
- non-destructive head movement,
|
||||
- imported files are data only,
|
||||
- minimal network surface.
|
||||
|
||||
If selected as base, Phase 0B must prove that adding Ollama + a browser service/UI is smaller than stripping AI-DnD.
|
||||
|
||||
## Phase 0B questions for Codex
|
||||
|
||||
1. Can its provider interface talk to Ollama via compatibility mode with a tiny adapter?
|
||||
2. Can a native Ollama adapter be added without touching turn/state logic?
|
||||
3. How much application code assumes CLI presentation?
|
||||
4. Is the app/service layer clean enough to expose through FastAPI without refactoring state internals?
|
||||
5. How are branches/checkpoints exported and navigated?
|
||||
6. Can generic freeform narrative facts/entities be represented without expanding typed events excessively?
|
||||
7. With Internet blocked, is the only runtime network connection the configured local model endpoint?
|
||||
|
||||
## Primary source links
|
||||
|
||||
- Repository: https://github.com/CaoRuiming/ai-adventure
|
||||
- Architecture: https://github.com/CaoRuiming/ai-adventure/blob/main/docs/architecture.md
|
||||
- Privacy/security: https://github.com/CaoRuiming/ai-adventure/blob/main/docs/privacy-and-security.md
|
||||
- Apache-2.0 license: repository `LICENSE`
|
||||
@@ -0,0 +1,170 @@
|
||||
# AI-DnD — Static Architecture Analysis
|
||||
|
||||
**Historical status:** Static Phase 0A analysis. Phase 0B corrected the shipped Undo assumption and confirmed AI-DnD as the production base after a successful head-cursor spike.
|
||||
|
||||
**Repository:** https://github.com/parththakkar106/AI-DnD
|
||||
**Date reviewed:** 2026-09-01
|
||||
**Disposition:** Preliminary fork recommendation / Finalist #1.
|
||||
|
||||
## Why it moved to first place
|
||||
|
||||
The static review indicates that AI-DnD already implements most of the difficult correctness infrastructure that would otherwise need to be invented:
|
||||
|
||||
- browser UI (React/Vite),
|
||||
- FastAPI backend,
|
||||
- local SQLite,
|
||||
- Ollama via local OpenAI-compatible endpoint,
|
||||
- story as a tree rather than a list,
|
||||
- alternate takes,
|
||||
- branch lineage that borrows ancestors,
|
||||
- state restored when switching branches,
|
||||
- non-destructive retry,
|
||||
- state snapshots,
|
||||
- exact prompt/context snapshots,
|
||||
- automatic summaries,
|
||||
- embedding-based long-term memory,
|
||||
- story cards/world information,
|
||||
- export/import of the complete story tree,
|
||||
- substantial automated backend testing.
|
||||
|
||||
The current README reports 549 backend tests. The design guide contains an older measured-results count of 440, so the clone should treat the live test suite—not prose counts—as authoritative.
|
||||
|
||||
## Story tree
|
||||
|
||||
This is the strongest reason to prefer AI-DnD.
|
||||
|
||||
The project explicitly models:
|
||||
|
||||
- branches,
|
||||
- actions/nodes,
|
||||
- parent/fork lineage,
|
||||
- multiple takes at a turn,
|
||||
- branch-aware context,
|
||||
- state after a node,
|
||||
- retry that preserves the replaced attempt.
|
||||
|
||||
That matches the user's desired “Git for stories” behavior much more closely than Open Dungeon.
|
||||
|
||||
Its documentation also describes measured optimization work so branches do not duplicate the ancestor transcript.
|
||||
|
||||
## Turn pipeline
|
||||
|
||||
The documented flow is close to the target Story Director:
|
||||
|
||||
```text
|
||||
player input
|
||||
-> optional input hook
|
||||
-> retrieve memories
|
||||
-> assemble bounded context
|
||||
-> snapshot exact context
|
||||
-> stream provider output
|
||||
-> extract proposed state delta
|
||||
-> Python referee validates state
|
||||
-> save action + resulting state
|
||||
-> background summarize/embed
|
||||
```
|
||||
|
||||
The target project would simplify this rather than reinvent it.
|
||||
|
||||
## Memory/context
|
||||
|
||||
AI-DnD already includes three useful layers:
|
||||
|
||||
- direct recent history,
|
||||
- AI-generated memories,
|
||||
- running story summary.
|
||||
|
||||
Embedding retrieval pulls old relevant memories back into context and exposes similarity/context details through an Insights UI.
|
||||
|
||||
Story cards provide a mature starting point for lore/world-info injection.
|
||||
|
||||
The main extension needed is a first-class imported document library with explicit authority classes:
|
||||
- Canon,
|
||||
- Reference,
|
||||
- Inspiration.
|
||||
|
||||
## Prompt transparency
|
||||
|
||||
The current project stores the exact prompt sent for a turn and provides an Insights view with context components and token costs. This directly satisfies a stated debugging requirement.
|
||||
|
||||
## What must be removed or generalized
|
||||
|
||||
AI-DnD is not a clean fit out of the box.
|
||||
|
||||
### RPG-specific world state
|
||||
Current world state is designed around stats, bands, flags, milestones, cooldowns, NPC presence, and state deltas.
|
||||
|
||||
Target:
|
||||
- retain the proposal/referee/snapshot pattern,
|
||||
- replace or supplement RPG stats with generic narrative entities/facts/relationships/story threads/scenes.
|
||||
|
||||
### QuickJS scripting
|
||||
The project includes AI-Dungeon-compatible user scripting.
|
||||
|
||||
For this project, executable campaign content conflicts with the desired narrow trust surface. Unless a compelling future use appears, remove or disable scripting in v1.
|
||||
|
||||
### Hosted/multi-user behavior
|
||||
Current code supports:
|
||||
- optional accounts,
|
||||
- guest users,
|
||||
- rate limits,
|
||||
- demo keys,
|
||||
- hosted deployments,
|
||||
- Postgres/Neon,
|
||||
- Render,
|
||||
- remote model providers.
|
||||
|
||||
The target is a single-user local application. These paths should be removed or compiled/configured out rather than merely hidden in the UI.
|
||||
|
||||
### Analytics
|
||||
The project includes its own owner-only aggregate visit analytics for hosted mode. It is not described as a third-party tracker, but it is unnecessary for the local fork and should be removed.
|
||||
|
||||
### Cloud providers
|
||||
OpenRouter/OpenAI/Groq/vLLM support is broader than desired. v1 should retain only the local Ollama path.
|
||||
|
||||
## Security positive
|
||||
|
||||
The local/hosted modes are already explicitly separated, and the code contains network-guard thinking around hosted deployments. This is a better starting point than a project with cloud assumptions scattered everywhere, but static review cannot prove that removal is trivial.
|
||||
|
||||
## Main risk
|
||||
|
||||
The central Phase 0B question is:
|
||||
|
||||
> Are the RPG/cloud/scripting systems modular enough that removing them is less work and less risk than adding correct branching/state/memory to Open Dungeon?
|
||||
|
||||
Static evidence suggests yes, but this must be tested with a local strip-down experiment.
|
||||
|
||||
## Best reuse case
|
||||
|
||||
If selected:
|
||||
- keep story tree,
|
||||
- keep action/state snapshots,
|
||||
- keep context/history windowing,
|
||||
- keep Memory Bank structure,
|
||||
- keep story cards,
|
||||
- keep Insights/prompt snapshots,
|
||||
- keep SQLite and local FastAPI/React split,
|
||||
- keep Ollama adapter path,
|
||||
- remove hosted/auth/analytics/cloud,
|
||||
- remove QuickJS,
|
||||
- generalize world state,
|
||||
- add document ingestion,
|
||||
- add scene/media schema and provider interface,
|
||||
- use Open Dungeon/Gamentic as media UX references.
|
||||
|
||||
## Phase 0B questions for Codex
|
||||
|
||||
1. Can the app run fully local with only Ollama and no Internet?
|
||||
2. Can QuickJS, hosted auth, analytics, Render/Neon, and remote provider paths be removed without destabilizing core tests?
|
||||
3. How tightly does branching depend on RPG world-state fields?
|
||||
4. Can an adventure run with minimal/no stats while branch rollback still passes?
|
||||
5. Can the state snapshot payload be generalized to narrative JSON without rewriting the tree?
|
||||
6. How many tests cover branch/undo/retry/context/memory independently of RPG logic?
|
||||
7. Does current Memory Bank work with a local Ollama embedding model in practice?
|
||||
8. What exact outbound traffic occurs in default local mode?
|
||||
|
||||
## Primary source links
|
||||
|
||||
- Repository / README: https://github.com/parththakkar106/AI-DnD
|
||||
- Design guide: https://github.com/parththakkar106/AI-DnD/blob/main/docs/GUIDE.md
|
||||
- MIT license: repository `LICENSE`
|
||||
@@ -0,0 +1,140 @@
|
||||
# Open Dungeon — Static Architecture Analysis
|
||||
|
||||
**Historical status:** Static Phase 0A analysis. Phase 0B confirmed Open Dungeon is a UX/media reference rather than the production base.
|
||||
|
||||
**Repository:** https://github.com/newideas99/open-dungeon
|
||||
**Date reviewed:** 2026-09-01
|
||||
**Disposition:** Finalist #2; strongest product/UI/media fit, but branch persistence requires material redesign.
|
||||
|
||||
## What maps well to the specification
|
||||
|
||||
Open Dungeon already provides a product very close to the desired interaction model:
|
||||
|
||||
- browser-first UI,
|
||||
- local Ollama text generation,
|
||||
- streaming narration,
|
||||
- Do / Say / Story-style interaction,
|
||||
- Continue / Retry / Erase / Edit controls,
|
||||
- SQLite persistence,
|
||||
- rolling story summary for long conversations,
|
||||
- persistent character records,
|
||||
- local inline image generation,
|
||||
- character portraits/visual continuity feeding image generation,
|
||||
- optional ComfyUI path.
|
||||
|
||||
The future-media requirement is therefore not hypothetical in this codebase. It already has a concept of the narrator requesting an image after prose and a local image backend producing it.
|
||||
|
||||
## Current stack
|
||||
|
||||
From the current package/config:
|
||||
|
||||
- Next.js 16
|
||||
- React 19
|
||||
- TypeScript
|
||||
- `better-sqlite3`
|
||||
- Node.js 22+
|
||||
- Ollama default endpoint at `127.0.0.1:11434`
|
||||
- local image worker defaults to loopback
|
||||
- optional remote OpenAI-compatible/OpenRouter configuration
|
||||
|
||||
## Persistence finding: the major issue
|
||||
|
||||
The current database is fundamentally a linear chat model.
|
||||
|
||||
The reviewed schema contains:
|
||||
|
||||
- chats,
|
||||
- messages,
|
||||
- characters,
|
||||
- app settings,
|
||||
- rolling story summary fields.
|
||||
|
||||
Messages do not expose a parent-turn/branch-lineage model equivalent to AI-DnD or ai-adventure.
|
||||
|
||||
More importantly, the data layer includes an operation named `deleteMessageAndAfter()`. Its own comment says it is used by retry/erase to discard the tail of the story. Prior text can also be updated in place.
|
||||
|
||||
That is directly contrary to the target requirement:
|
||||
|
||||
> Going backward should preserve the abandoned future as an alternate branch.
|
||||
|
||||
This means adding robust branching is not simply a UI feature. It requires changing the persistence semantics and all features that assume a single mutable message sequence, including retry/erase/edit and summary lineage.
|
||||
|
||||
## Memory/context model
|
||||
|
||||
The prompt builder contains a history-packing mechanism with block eviction. Old story material is compressed into a rolling story summary, while recent history stays direct.
|
||||
|
||||
This is a reasonable lightweight storyteller strategy but is below the target design:
|
||||
|
||||
- no established semantic retrieval of old story events,
|
||||
- no explicit Canon / Reference / Inspiration document library,
|
||||
- no branch-aware memory lineage,
|
||||
- no rich authoritative generic story-state graph.
|
||||
|
||||
Those systems would need to be added.
|
||||
|
||||
## Media design
|
||||
|
||||
Open Dungeon is the strongest finalist for immediate media UX.
|
||||
|
||||
The current narrator prompt exposes a `generate_image` tool after a passage and passes selected character IDs so the image path can preserve visual identity. The app can use local FLUX tooling and supports ComfyUI in recent releases.
|
||||
|
||||
Useful ideas to retain even if Open Dungeon is not the base:
|
||||
|
||||
1. media is optional; text play does not depend on it,
|
||||
2. image generation is scene/turn-associated,
|
||||
3. character visual identity is stored rather than reinvented each image,
|
||||
4. local backend is behind a service boundary,
|
||||
5. generated media appears inline in the story.
|
||||
|
||||
For our architecture, image generation should eventually move behind a generic media-provider interface rather than remain hard-coded to one model/workflow.
|
||||
|
||||
## Privacy/static network assessment
|
||||
|
||||
Positive:
|
||||
- local Ollama is the default,
|
||||
- local SQLite is the default,
|
||||
- local image generation is supported,
|
||||
- no telemetry requirement was apparent in the inspected package/config.
|
||||
|
||||
Hardening needed:
|
||||
- remove or disable OpenRouter and arbitrary remote OpenAI-compatible provider options in v1,
|
||||
- review Tailscale/LAN exposure separately from loopback-only default,
|
||||
- verify built frontend has no remote runtime assets,
|
||||
- verify image setup does not make unexpected runtime downloads after installation,
|
||||
- runtime network capture still required.
|
||||
|
||||
## Testing concern
|
||||
|
||||
The inspected `package.json` exposes build/lint/image checks but no obvious comprehensive automated test command comparable to AI-DnD or Gamentic. This must be verified after clone; if accurate, a branch/persistence rewrite would need a new test foundation before implementation.
|
||||
|
||||
## Best reuse case
|
||||
|
||||
If Open Dungeon becomes the base:
|
||||
- preserve the browser experience,
|
||||
- preserve Ollama integration,
|
||||
- preserve image/visual-continuity concepts,
|
||||
- replace/extend linear message persistence with a parent-linked turn graph,
|
||||
- make summary/memory branch-aware,
|
||||
- add authoritative generic narrative state,
|
||||
- add local document ingestion and retrieval,
|
||||
- add prompt/provenance inspection.
|
||||
|
||||
If AI-DnD becomes the base:
|
||||
- use Open Dungeon primarily as a UX and media-generation reference.
|
||||
|
||||
## Phase 0B questions for Codex
|
||||
|
||||
1. How many routes/components assume messages are a single ordered list?
|
||||
2. Can a parent-linked turn/branch layer be introduced without replacing most chat APIs?
|
||||
3. What happens to `story_summary` when retry/erase edits earlier history?
|
||||
4. Can current image records attach cleanly to immutable turn IDs/scene IDs?
|
||||
5. Is there an automated test suite not visible from the package manifest?
|
||||
6. With Internet blocked, does ordinary text + local image play produce only loopback traffic?
|
||||
|
||||
## Primary source links
|
||||
|
||||
- Repository: https://github.com/newideas99/open-dungeon
|
||||
- DB: https://github.com/newideas99/open-dungeon/blob/main/src/lib/db.ts
|
||||
- Prompt builder: https://github.com/newideas99/open-dungeon/blob/main/src/lib/story-prompt.ts
|
||||
- Environment: https://github.com/newideas99/open-dungeon/blob/main/.env.example
|
||||
- Package: https://github.com/newideas99/open-dungeon/blob/main/package.json
|
||||
@@ -0,0 +1,119 @@
|
||||
# Phase 0B — Experiment C: ai-adventure
|
||||
|
||||
## C1. Tests
|
||||
|
||||
```
|
||||
python -m unittest discover -s tests → Ran 76 tests in 1.9s OK
|
||||
```
|
||||
|
||||
## C2. Provider abstraction — the Ollama adapter is zero code
|
||||
|
||||
`llm/backend.py` defines a `ModelBackend` Protocol with a single `generate`
|
||||
method plus typed `ModelRequest`/`ModelResponse`. `llm/lm_studio.py` is not
|
||||
LM-Studio-specific at all: it POSTs `/v1/chat/completions` and reads `/v1/models`
|
||||
— the same OpenAI-compatible surface Ollama serves. `llm/scripted.py` is a
|
||||
deterministic fake used by the tests.
|
||||
|
||||
The only change needed was two lines of `world.toml`:
|
||||
|
||||
```toml
|
||||
base_url = "http://127.0.0.1:11434" # was 127.0.0.1:1234
|
||||
name = "qwen2.5:3b-instruct" # was "replace-with-a-local-model"
|
||||
```
|
||||
|
||||
```
|
||||
python -m local_adventure doctor --world <world>
|
||||
[PASS] SQLite FTS5 available
|
||||
[PASS] Database schema at version 2
|
||||
[PASS] World valid: Ember Hollow
|
||||
[PASS] LM Studio reachable: http://127.0.0.1:11434
|
||||
[PASS] Configured model visible: qwen2.5:3b-instruct
|
||||
```
|
||||
|
||||
A real turn then committed:
|
||||
|
||||
```
|
||||
TURN: (30ad56d4…, 1, 'committed', 'I search the observatory for the brass k',
|
||||
'You rummage through the dusty shelves, your fingers brushing away
|
||||
decades of cobwebs...')
|
||||
MODELCALL: ('lm_studio', 'qwen2.5:3b-instruct', attempt 1, 832 bytes, errors=False)
|
||||
```
|
||||
|
||||
`ModelSettings.backend` is typed `Literal["lm_studio"]`, so adding a second
|
||||
backend means widening one literal and one factory call in `cli.py` — the
|
||||
transport itself already works. Phase 0A's "MODIFY — LM Studio primary provider"
|
||||
overstated this.
|
||||
|
||||
**Caveat:** with `qwen2.5:0.5b` the same turn failed with "The model response
|
||||
could not be validated; no turn was saved." That is the validation contract
|
||||
working as designed — the engine refuses to commit an unvalidated turn — but the
|
||||
typed-event contract needs a reasonably capable model, where the freeform-prose
|
||||
candidates tolerate a weak one.
|
||||
|
||||
## C3. Undo / branch / checkpoint / replay — verified independently
|
||||
|
||||
Run directly against `GameService` with a temporary database, not by trusting the
|
||||
project's own tests:
|
||||
|
||||
```
|
||||
after 2 turns: turns=2 events=2 has_brass_key=False
|
||||
|
||||
UNDO
|
||||
turns 2 → 2 (DELETED 0)
|
||||
state after undo: has_brass_key=True
|
||||
abandoned turn id_3 still in DB: True
|
||||
|
||||
BRANCH from the restored point, then play on the branch
|
||||
branch state: has_brass_key=True gate_open=True
|
||||
ORIGINAL session leaked 'gate_open'? False
|
||||
total turns in DB: 3 (nothing destroyed)
|
||||
|
||||
CHECKPOINT RESTORE ("before_gate")
|
||||
restored has_brass_key=True
|
||||
later turn id_3 still present: True
|
||||
turns in DB after restore: 3
|
||||
```
|
||||
|
||||
This is the specification's history model, working. `undo` sets
|
||||
`session.head_turn_id` to the parent; `branch` inserts a new session sharing the
|
||||
same head; `restore_checkpoint` replays ancestry; `_replay_to` walks
|
||||
`turns.ancestry(head)` and reduces `state_events`.
|
||||
|
||||
## C4. Schema — the target data model
|
||||
|
||||
```
|
||||
sessions(session_id, ..., head_turn_id, initial_state_json)
|
||||
turns(turn_id, session_id, parent_turn_id, turn_number,
|
||||
player_input, narration, status CHECK IN ('committed','failed'), model_call_id)
|
||||
state_events(event_id, turn_id, sequence_number, event_type, payload_json) -- append-only
|
||||
state_cache(session_id, head_turn_id, state_json) -- replay cache
|
||||
model_calls(..., request_hash, response_hash, parsed_response_json,
|
||||
validation_errors_json, prompt_eval_count, eval_count, duration_ms)
|
||||
named_checkpoints(checkpoint_id, session_id, turn_id, name, UNIQUE(session_id,name))
|
||||
summaries(summary_id, session_id, through_turn_id, kind CHECK IN ('scene','campaign'))
|
||||
lore_documents + lore_documents_fts (FTS5)
|
||||
```
|
||||
|
||||
Note `summaries.through_turn_id` — anchored to a turn, not a count. That is
|
||||
exactly the property Open Dungeon lacks.
|
||||
|
||||
## C5. Coupling to the CLI
|
||||
|
||||
The layering is clean: `cli.py` → `app/commands.py` → `app/game_service.py` /
|
||||
`app/turn_service.py` → `storage/repositories.py`. Authoritative state lives in
|
||||
`state/` (events, reducer, validator) and context assembly in `context/`. None of
|
||||
that imports the CLI. A browser/API layer could sit beside `cli.py` and call the
|
||||
same services without moving state logic.
|
||||
|
||||
What it does **not** have, and would all be new work: any HTTP layer, streaming,
|
||||
in-app campaign setup (worlds are hand-authored TOML + Markdown on disk),
|
||||
embeddings or semantic memory, document import, media, and prompt inspection
|
||||
beyond stored hashes.
|
||||
|
||||
## C6. Lore / FTS — reusable
|
||||
|
||||
`lore/indexer.py` + `lore_documents_fts` (FTS5) with content hashing and
|
||||
`modified_ns` for incremental reindex, scoped by `world_id`, and a `kind` column
|
||||
already carrying a classification axis. This is a good starting shape for the
|
||||
Canon/Reference/Inspiration tiers in `IMPORTED-KNOWLEDGE-DESIGN.md`, and it is
|
||||
deterministic and offline.
|
||||
@@ -0,0 +1,146 @@
|
||||
# Phase 0B — Experiment A: AI-DnD
|
||||
|
||||
All results are from a live server on `127.0.0.1:8321` with a fresh SQLite
|
||||
database, generating against local Ollama.
|
||||
|
||||
## A1. Runs with Ollama locally — yes
|
||||
|
||||
Default `endpoint_url` is `http://localhost:11434/v1`. After setting only the
|
||||
model name:
|
||||
|
||||
```
|
||||
POST /api/settings/test → {"ok":true,"models":["qwen2.5:0.5b"]}
|
||||
```
|
||||
|
||||
Three streamed turns produced 101/162/113 SSE events with `player`, `chunk` and
|
||||
`done` types and a coherent transcript.
|
||||
|
||||
## A2. RPG state can be left empty — yes
|
||||
|
||||
`schemas.AdventureCreate.scenario_id` is optional. With it null,
|
||||
`crud.create_adventure` sets `world_state={}` and creates no story cards, no
|
||||
scripts and no opening action, but still calls `tree.head_branch(...)` so the
|
||||
story tree exists from the start. A noir campaign created this way played,
|
||||
branched, summarized and retrieved memories normally. The RPG referee is
|
||||
opt-in per scenario, not a core dependency.
|
||||
|
||||
## A3. Retry is non-destructive; Undo is destructive
|
||||
|
||||
Measured on the same adventure, reading rows straight from SQLite:
|
||||
|
||||
```
|
||||
baseline rows: 6 ids [1,2,3,4,5,6]
|
||||
|
||||
RETRY → new action id 7; rows 7 ids [1..7]
|
||||
/variants at the tip lists BOTH take 0 (id 6) and take 1 (id 7)
|
||||
|
||||
UNDO → rows 7 → 4; DELETED ids = [5, 6, 7]
|
||||
```
|
||||
|
||||
Undo removed the player action (5), the live AI take (6) **and the alternate
|
||||
take that the retry had just preserved (7)**. `nodes.delete_turn` deletes every
|
||||
attempt in the group on that branch and calls `memorybank.forget_node`. There is
|
||||
no redo endpoint anywhere in `routers/`.
|
||||
|
||||
This contradicts `reports/PRELIMINARY-RECOMMENDATION.md`, which credits AI-DnD
|
||||
with "non-destructive retry" (true) and generalizes it to rollback (false).
|
||||
|
||||
## A4. Branching is non-destructive
|
||||
|
||||
Adding an alternate take at a passed turn (`POST /actions/3/takes`) forked
|
||||
automatically:
|
||||
|
||||
```
|
||||
before: ids 1-8, all branch 1, depths 0-7
|
||||
after : ids 1-8 unchanged on branch 1
|
||||
id 9 branch 2 depth 2 (the alternate player action)
|
||||
id 10 branch 2 depth 3 (its AI continuation)
|
||||
|
||||
/branches → [ {id:1, parent:null, fork_depth:null, own_actions:8},
|
||||
{id:2, parent:1, fork_depth:1, own_actions:2, is_head:true} ]
|
||||
```
|
||||
|
||||
Not one row of branch 1 was touched. Note the semantics: `after_id` pointing at a
|
||||
node that is *live and on the path* is a no-op, not a fork. A fork happens only
|
||||
when writing below a take the story has moved past (`nodes.stand_on`). So
|
||||
"restore to an arbitrary earlier turn and branch" is not directly expressible
|
||||
today — it must go through creating a take at that turn.
|
||||
|
||||
## A5. Memory isolation across branches — correct, with negative control
|
||||
|
||||
Memory bank enabled, summary model `qwen2.5:0.5b`, embeddings
|
||||
`nomic-embed-text`, `memory_top_k=5`. Three more turns played on branch 1
|
||||
establishing possessions (brass key, silver revolver, bullets).
|
||||
|
||||
```
|
||||
memories stored: (id 1, branch 1, depth 5)
|
||||
(id 2, branch 1, depth 11)
|
||||
|
||||
on branch 2 (forked at depth 1):
|
||||
memories retrieved = {"used": [], "error": null}
|
||||
leak probes over the FULL assembled prompt:
|
||||
'brass key' False | 'revolver' False | 'lockbox' False
|
||||
'ashtray' False | 'bullets' False
|
||||
|
||||
NEGATIVE CONTROL — switch back to branch 1:
|
||||
memories retrieved = 2, top similarity 0.8087
|
||||
'revolver' present in prompt = True
|
||||
```
|
||||
|
||||
The empty result on branch 2 is genuine isolation, not a dead retrieval path.
|
||||
`memorybank.retrieve_memories` applies `lineage.path_of(db, adventure).clause(models.Memory)`,
|
||||
and `Memory` rows carry `branch_id` plus source depths. Eviction deliberately
|
||||
ignores the branch clause, which is correct — it is a capacity concern.
|
||||
|
||||
History is likewise lineage-scoped: branch 2's prompt contained only its own
|
||||
text.
|
||||
|
||||
## A6. Insights / prompt transparency — already sufficient
|
||||
|
||||
`GET /adventures/{id}/context` returns the exact next prompt broken into labeled
|
||||
sections with token costs:
|
||||
|
||||
```
|
||||
narrator 57 | ai_instructions 9 | persona 15 | plot_essentials 24
|
||||
history 988 | used_memories 143 | length_hint 46
|
||||
```
|
||||
|
||||
Per-action snapshots are stored on `Action.context_snapshot` and served by
|
||||
`GET /actions/{id}/context`. This satisfies specification §11 as-is.
|
||||
|
||||
## A7. Scripting/QuickJS can be removed cheaply
|
||||
|
||||
`quickjs` is imported in exactly one file (`scripting/engine.py`). The whole
|
||||
feature is reachable through two symbols (`run_hook`, `ScriptPipeline`) across
|
||||
four call sites plus `routers/scripts.py`.
|
||||
|
||||
Disposable experiment: made `quickjs` unimportable via a `meta_path` blocker and
|
||||
replaced `run_hook` with a pass-through stub.
|
||||
|
||||
```
|
||||
app.main imports OK with scripting stubbed and quickjs absent
|
||||
19 failed, 613 passed (3.0% of the suite)
|
||||
```
|
||||
|
||||
Every failure is in `test_turn_flow_integration`, `test_take_state`,
|
||||
`test_story_tree_baseline` or `test_retry_variants`, and every one fails for the
|
||||
same reason: those tests **use a JS `output_js` hook as the instrument** to make
|
||||
state observable, e.g. `test_play_then_undo_reverts_gold` seeds a `GOLD_SCRIPT`
|
||||
and asserts `script_state == {"gold": 10}` then `{}` after undo. The behavior
|
||||
under test (per-node state rollback) is intact; only the measuring device is
|
||||
gone. Re-instrumenting against `world_state_after`, which is already snapshotted
|
||||
per node, is the fix.
|
||||
|
||||
## A8. Hosted/cloud features
|
||||
|
||||
`MULTI_USER` is off by default and gates ten modules
|
||||
(`auth`, `limits`, `netguard`, `cleanup`, `security`, `analytics`,
|
||||
`routers/{auth,analytics,debug,story_cards}`). `analytics.py` is a *self-hosted*
|
||||
counter writing to two local tables — no third-party tracker and no outbound
|
||||
call — so it is removable rather than dangerous. `render.yaml`, the demo-key
|
||||
path and `psycopg` are the hosted leftovers.
|
||||
|
||||
`netguard.py` is worth flagging: it refuses endpoints that resolve to
|
||||
**non-public** addresses, and only in multi-user mode. That is SSRF protection
|
||||
for a hosted deployment. Our requirement is the opposite — warn on endpoints that
|
||||
are **not loopback**. See ai-adventure's `_endpoint_warnings` for the shape we want.
|
||||
@@ -0,0 +1,39 @@
|
||||
# Phase 0B — Baseline Results
|
||||
|
||||
**Date:** 2026-09-01. Host: Linux, 4 cores, 15 GB RAM, no GPU, Python 3.12.3, Node 22.23.1.
|
||||
Inference: Ollama in Docker (`ollama/ollama:latest`) on `127.0.0.1:11434`, models
|
||||
`qwen2.5:0.5b`, `qwen2.5:3b-instruct`, `nomic-embed-text`.
|
||||
|
||||
| | AI-DnD | Open Dungeon | ai-adventure |
|
||||
|---|---|---|---|
|
||||
| SHA | `d72f7c1b` | `b0a79f96` | `873ea918` |
|
||||
| License | MIT | MIT | Apache-2.0 |
|
||||
| Repo age / commits / authors | created 2026-07-20, 172, 3 | 2026-06-12, 45, 1 | 2026-07-20, 20, 1 |
|
||||
| Code size | ~30.7k LOC | ~9.1k LOC | ~4.5k LOC |
|
||||
| Stack | FastAPI + SQLAlchemy + React/Vite | Next.js 16 + React 19 | Python stdlib + pydantic, CLI |
|
||||
| Storage | SQLite (Postgres option via `AIDND_DATABASE_URL`) | SQLite (`better-sqlite3`) | SQLite + FTS5 |
|
||||
| Install | `pip install -r requirements*.txt`, `npm install` — clean | `npm install` — clean, 7 advisories (6 high) | `pip install -r requirements.txt` — clean |
|
||||
| Dependencies | 41 pip + 39 npm | 335 npm | **6 pip (1 direct: pydantic)** |
|
||||
| Tests | **632 passed / 236s** | **none — no script, no framework** | **76 passed / 1.9s** |
|
||||
| Ports | 8000 (prod/docker), 5173 + 8000 (dev) | 3000 (app), 7869 (FLUX worker), 8188 (ComfyUI) | none (CLI) |
|
||||
| Model assumption | any OpenAI-compatible; **defaults to `http://localhost:11434/v1`** | Ollama native API, or OpenAI-compatible "custom" | OpenAI-compatible; base_url in `world.toml` |
|
||||
| Docker build | `docker build` succeeds (3-stage, builds SPA) | not attempted (npm path used) | n/a |
|
||||
| Runtime failures | none once running; **offline blocker, see offline report** | none | typed-event validation fails on a weak model (by design) |
|
||||
| Cloud/hosted surface | `render.yaml`, multi-user auth, demo key, analytics, Postgres, OpenRouter default | OpenRouter preset, Next telemetry on by default, donation links | **none** |
|
||||
|
||||
## Notes
|
||||
|
||||
- **AI-DnD** ran its whole suite green on the first attempt with no fixes. Its
|
||||
default settings already point at Ollama; `POST /api/settings/test` returned
|
||||
`{"ok":true,"models":["qwen2.5:0.5b"]}` with no configuration beyond selecting
|
||||
a model name.
|
||||
- **Open Dungeon** has `lint` and several Windows/image smoke scripts, but no
|
||||
unit or integration tests of any kind. Its `local` provider only offers five
|
||||
hard-coded Gemma 4 QAT builds (`src/lib/text-models.ts`), so a non-Gemma local
|
||||
model must go through the "custom" OpenAI-compatible provider; that is how it
|
||||
was driven here.
|
||||
- **ai-adventure** is the only candidate whose full dependency closure is one
|
||||
third-party package. `python -m local_adventure doctor` is a genuine
|
||||
environment self-check (Python version, writable runtime dir, SQLite version,
|
||||
FTS5 availability, schema version, world validity, endpoint reachability,
|
||||
model visibility).
|
||||
@@ -0,0 +1,164 @@
|
||||
# Phase 0B — Follow-Up Checks
|
||||
|
||||
Four open questions settled after the undo/redo spike. Ordered by how much each
|
||||
changes the plan.
|
||||
|
||||
---
|
||||
|
||||
## 1. The world-state referee under a local model
|
||||
|
||||
**Why it mattered:** the referee is the closest thing AI-DnD has to
|
||||
specification §2's "the model proposes, the application owns authoritative
|
||||
state", and §7's structured narrative state. It was never exercised in the main
|
||||
round, because the RPG path was deliberately left empty to prove genre-neutrality.
|
||||
If it only works with a cloud-class model, the principle the architecture rests
|
||||
on is at risk locally.
|
||||
|
||||
**How it works.** `EMIT_RULE` asks for a fenced ` ```state ` block of **relative
|
||||
deltas**, with an example (`{"player.hp": -15, ...}`) and a recency reminder
|
||||
appended at the end of the prompt. `apply.py` computes `new = old + delta`, then
|
||||
clamps by `max_delta_per_turn`, then by `min`/`max`, and reports every change as
|
||||
applied / clamped / rejected with corrective text fed back to the model.
|
||||
|
||||
### Result A — in the full application, `qwen2.5:3b-instruct` sent absolute values
|
||||
|
||||
Playing the seeded "Bandit Camp (RPG world state)" scenario:
|
||||
|
||||
```
|
||||
turn 1 proposed: {"player.hp": 92, "npc.gwen.trust": 25,
|
||||
"player.outfit": "...arrow in shoulder", "npc.gwen.health": 90}
|
||||
applied : player.hp 100 -> 100 "did not move. It is already at its
|
||||
maximum of 100 ... at most 35 per turn"
|
||||
gwen.trust 20 -> 40 (clamped from +25 to +20)
|
||||
gwen.health 100 -> 100 (clamped)
|
||||
outfit updated correctly
|
||||
|
||||
turn 2 proposed: {"npc.gwen.trust": 30}
|
||||
applied : gwen.trust 40 -> 60 (clamped from +30 to +20)
|
||||
```
|
||||
|
||||
The model meant "hp is now 92". The engine read `+92`, clamped it to `+35`, then
|
||||
to the ceiling of 100. **The player took an arrow to the shoulder and finished
|
||||
the turn at full health**, with the prose describing a wound.
|
||||
|
||||
This is the failure mode that matters: it is not corruption, and no validator can
|
||||
catch it, because `+92` is a perfectly legal proposal. The application stayed
|
||||
authoritative — which is the point of the referee — but the authoritative state
|
||||
silently stopped tracking the narration.
|
||||
|
||||
The free-text stat (`player.outfit`) was handled correctly, because it is marked
|
||||
`(free text)` and takes a whole value rather than a delta.
|
||||
|
||||
### Result B — in isolation, the same model gets it right
|
||||
|
||||
An isolated probe using AI-DnD's own `EMIT_RULE` and `extract_delta` verbatim, so
|
||||
that only the model varies. A short stat guide, live values, one wound scene,
|
||||
3 samples each:
|
||||
|
||||
| Model | Emitted a parsable state block | Correct relative (negative) hp |
|
||||
|---|---|---|
|
||||
| `qwen2.5:0.5b` | 0/3 | 0/3 |
|
||||
| `qwen2.5:3b-instruct` | 3/3 | **3/3** (−20, −20, −15) |
|
||||
| `qwen2.5:7b-instruct` | 3/3 | **3/3** (−15, −30, −15) |
|
||||
|
||||
So the delta protocol is **not** beyond a 3B model. The difference between A and
|
||||
B is context: the full application prompt carries scenario prose, persona, a
|
||||
stat guide, live values, story cards and history, and under that load the 3B
|
||||
model degraded to absolute values. In a short prompt it followed the same
|
||||
instruction perfectly.
|
||||
|
||||
> Recorded because it nearly became a wrong finding: the first run of this probe
|
||||
> reported 0/3 for both models. That was a bug in the probe, not the models —
|
||||
> `extract_delta` returns `(clean_text, delta)` and the harness was reading
|
||||
> element 0. The table above is the corrected run.
|
||||
|
||||
Note also `qwen2.5:7b-instruct` sample 2 proposing `npc.gwen.trust: +75`, which
|
||||
the referee would clamp to +20. Even a capable model proposes disproportionate
|
||||
values, which is precisely what the clamp is for.
|
||||
|
||||
### What this means for the design
|
||||
|
||||
1. **Not a blocker for the fork.** The referee is per-scenario and opt-in, and we
|
||||
are generalizing it into narrative state rather than adopting it as-is.
|
||||
2. **Prefer an unambiguous state vocabulary.** A relative-delta protocol has a
|
||||
failure mode that validation cannot detect, and it degrades with context
|
||||
length. ai-adventure's typed events (`set_flag key=… value=…`) are explicit
|
||||
and absolute, so the same mistake is impossible to express. **This is a
|
||||
concrete argument for taking ai-adventure's event vocabulary rather than
|
||||
AI-DnD's delta vocabulary when generalizing to genre-neutral narrative state.**
|
||||
3. **Keep the emission rule close to generation.** AI-DnD already appends
|
||||
`EMIT_REMINDER` at the end of the prompt for exactly this reason; whatever we
|
||||
build should keep that property and be tested at realistic context length,
|
||||
not just in a clean harness.
|
||||
4. **Sample sizes are small** (3 per model; 2 parsed in-app turns). Enough to
|
||||
establish the failure mode exists and that it is context-dependent, not enough
|
||||
to rank models. A proper capability matrix belongs in the build phase.
|
||||
|
||||
---
|
||||
|
||||
## 2. Export/import silently redoes an undone story
|
||||
|
||||
**Status: a real defect the spike introduces into the export path. One-field fix.**
|
||||
|
||||
`bundle.py:_point_the_head` sets the head on import by recomputing it:
|
||||
|
||||
```python
|
||||
depths = [n["depth"] for n in story["nodes"] if n["branch"] == head]
|
||||
if depths:
|
||||
adventure.head_depth = max(depths)
|
||||
```
|
||||
|
||||
The export format carries `headBranch` (line 120) but **not** the head depth.
|
||||
|
||||
With the destructive undo this was correct — undone rows did not exist, so
|
||||
`max(depths)` was the head. With the spike's non-destructive undo, an adventure
|
||||
exported while its head sits behind the tip comes back with every abandoned turn
|
||||
restored. The undo is silently undone by a round-trip.
|
||||
|
||||
**Fix:** add `headDepth` to the bundle (optional field, or a `v3`), and honour it
|
||||
in `_point_the_head`, falling back to `max(depths)` when absent so that existing
|
||||
`ai-dnd-adventure-v2` files still import. Small, but it must land in the same
|
||||
change as the non-destructive undo, not after it.
|
||||
|
||||
---
|
||||
|
||||
## 3. Story cards are not lineage-scoped
|
||||
|
||||
**Status: a design constraint for imported knowledge, not a present defect.**
|
||||
|
||||
`models.StoryCard` is keyed by `scenario_id` **or** `adventure_id`, and carries
|
||||
**no `branch_id` and no `depth`** — unlike `Memory`, which carries both and is
|
||||
therefore filtered by the same capped lineage clause that gives memories their
|
||||
branch isolation.
|
||||
|
||||
`context.builder.match_cards` is pure keyword matching over the adventure's whole
|
||||
card list; `routers/story_cards.py` contains no lineage reference at all.
|
||||
|
||||
This is harmless today, because cards are authored by the user or copied from a
|
||||
scenario — nothing derives a card from story content, so an abandoned branch
|
||||
cannot invent one.
|
||||
|
||||
It stops being harmless the moment imported knowledge or any derived-card feature
|
||||
lands. **Implication for `IMPORTED-KNOWLEDGE-DESIGN.md`:** any knowledge source
|
||||
that can be created or updated from story content must carry `(branch_id, depth)`
|
||||
like `Memory` does, and be read through `lineage.Path.clause`. Do that and it
|
||||
inherits both the branch isolation and the undo isolation for free — which is the
|
||||
strongest argument for a new table rather than extending `story_cards`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Postgres/psycopg is cleanly removable
|
||||
|
||||
**Status: resolved, low effort.**
|
||||
|
||||
The entire Postgres surface is three places:
|
||||
|
||||
- `database.py` — the `AIDND_DATABASE_URL` / `DATABASE_URL` branch and
|
||||
`_normalize_url`. Deleting the branch leaves the SQLite path, which is already
|
||||
the default when neither variable is set.
|
||||
- `migrations.py` — two helpers that tolerate SQLite returning a raw JSON string
|
||||
where psycopg returns a parsed list. They simplify rather than disappear.
|
||||
- `analytics.py` — `from sqlalchemy.dialects.postgresql import insert as pg_insert`,
|
||||
which goes with analytics, already scheduled for removal.
|
||||
|
||||
`psycopg[binary]` then drops out of `requirements.txt`. No blockers.
|
||||
@@ -0,0 +1,106 @@
|
||||
# Phase 0B — Offline and Network Behavior
|
||||
|
||||
Method: a Docker network created with `--internal` (no NAT, no DNS to the
|
||||
outside). The Ollama container was attached to it so the app could reach a model
|
||||
while having no path to the Internet. Isolation was verified from inside the
|
||||
app container before testing:
|
||||
|
||||
```
|
||||
blocked ('1.1.1.1', 443) OSError
|
||||
blocked ('openrouter.ai', 443) gaierror
|
||||
blocked ('fonts.googleapis.com', 443) gaierror
|
||||
```
|
||||
|
||||
## AI-DnD — fails offline as shipped; fine once one file is vendored
|
||||
|
||||
First turn on the isolated network died. The SSE stream emitted the `player`
|
||||
event and stopped. Container log:
|
||||
|
||||
```
|
||||
requests.exceptions.ConnectionError: HTTPSConnectionPool(
|
||||
host='openaipublic.blob.core.windows.net', port=443):
|
||||
Max retries exceeded with url: /encodings/cl100k_base.tiktoken
|
||||
(NameResolutionError ... Temporary failure in name resolution)
|
||||
```
|
||||
|
||||
`tiktoken` downloads its BPE encoding on first use, and AI-DnD calls it for
|
||||
context budgeting on every turn. On the host this was invisible, because the
|
||||
file had already been cached to `/tmp/data-gym-cache/9b5ad71b…` during an
|
||||
earlier online run.
|
||||
|
||||
After copying that 1.7 MB file into the container:
|
||||
|
||||
```
|
||||
OFFLINE TURN GENERATED: "I am Vale, an explorer. I journey alone through the
|
||||
darkened depths of the lighthouse..."
|
||||
```
|
||||
|
||||
So: a hard blocker on a clean air-gapped install, and a packaging fix — vendor
|
||||
the encoding (or pre-seed `TIKTOKEN_CACHE_DIR`, or replace the tokenizer). Worth
|
||||
stressing that **static analysis could not have found this**; it took an actually
|
||||
isolated run.
|
||||
|
||||
Other AI-DnD network surface:
|
||||
|
||||
- **Google Fonts at runtime.** The built SPA's `index.html` still contains
|
||||
`fonts.googleapis.com/css2?family=Cinzel...&family=Crimson+Pro...&family=Inter...`,
|
||||
and `main.py`'s CSP explicitly allows `fonts.googleapis.com` /
|
||||
`fonts.gstatic.com`. Violates specification §12. Fix: self-host three families.
|
||||
- Only other remote host in backend Python is `https://openrouter.ai/api/v1`,
|
||||
used as a default endpoint constant and a header-attribution host check. Both
|
||||
removable with the cloud-provider path.
|
||||
- `analytics.py` is a **self-hosted** counter writing to two local tables. No
|
||||
third-party script, no outbound request. It records no IP, no user agent, and
|
||||
hashes the user id with HMAC. Removable, and not a telemetry leak in the
|
||||
meantime.
|
||||
- Connection test and turns honour the configured endpoint only; no automatic URL
|
||||
retrieval or content fetching was observed.
|
||||
|
||||
## Open Dungeon — runtime clean, build and telemetry are not
|
||||
|
||||
- **Next.js telemetry is on by default** and printed its notice on first start.
|
||||
Needs `NEXT_TELEMETRY_DISABLED=1` or `next telemetry disable` in the
|
||||
production configuration.
|
||||
- **Fonts are fine at runtime.** `layout.tsx` uses `next/font/google`, which
|
||||
downloads at *build* time and self-hosts. The served page contains no
|
||||
`fonts.googleapis.com` reference. The trade-off is that the build needs
|
||||
network access.
|
||||
- Remote hosts referenced in `src/`: `openrouter.ai` (preset provider URL and a
|
||||
model list), plus `ko-fi.com` and `github.com/sponsors` donation links in the
|
||||
UI. Everything else is `127.0.0.1` / `localhost` (13 occurrences).
|
||||
- Local story play needs no cloud service: Ollama for text, a local FLUX worker
|
||||
or the user's own ComfyUI for images.
|
||||
- 335 npm packages with 6 high-severity advisories is the largest supply-chain
|
||||
surface of the three.
|
||||
|
||||
## ai-adventure — verifiably local-only
|
||||
|
||||
The strongest posture by a wide margin, and the easiest to audit:
|
||||
|
||||
- **Exactly one outbound call site in the whole codebase**: `urlopen` in
|
||||
`llm/lm_studio.py`. Nothing else in `local_adventure/` opens a socket.
|
||||
- **No hardcoded remote host anywhere.** The only `http` string in the package is
|
||||
the scheme check in `content/models.py`.
|
||||
- **One direct dependency** (`pydantic`), six packages in the closure.
|
||||
- It is the only candidate that already implements specification §12's
|
||||
non-local-endpoint warning:
|
||||
|
||||
```python
|
||||
def _endpoint_warnings(config):
|
||||
hostname = urlparse(config.model.base_url).hostname
|
||||
if hostname not in {"127.0.0.1", "localhost", "::1"}:
|
||||
return ["model.base_url is not loopback; game prompts and content will
|
||||
be sent to that endpoint. Enable API authentication."]
|
||||
```
|
||||
|
||||
- `audit.store_prompts` defaults to `false`, with prompt *hashes* stored instead.
|
||||
- Tests, world validation, session creation, replay, branching and export are all
|
||||
offline by construction.
|
||||
|
||||
## Summary
|
||||
|
||||
| | Local play works offline | Cloud service required | Telemetry / remote assets | Removal difficulty |
|
||||
|---|---|---|---|---|
|
||||
| AI-DnD | **Only after vendoring the tiktoken encoding** | No | Google Fonts at runtime; local-only analytics tables | Low — one vendored file, three self-hosted fonts, delete analytics |
|
||||
| Open Dungeon | Yes at runtime; build needs network | No | Next.js telemetry on by default | Low — one env var; fonts already self-hosted |
|
||||
| ai-adventure | **Yes, unconditionally** | No | **None** | Nothing to remove |
|
||||
@@ -0,0 +1,105 @@
|
||||
# Phase 0B — Experiment B: Open Dungeon history model
|
||||
|
||||
Driven live on `127.0.0.1:3111` against Ollama through the "custom"
|
||||
OpenAI-compatible provider (`http://127.0.0.1:11434/v1`, `qwen2.5:0.5b`), with
|
||||
its own SQLite file.
|
||||
|
||||
## B1. Schema — no lineage exists
|
||||
|
||||
`src/lib/db.ts` creates four tables: `chats`, `messages`, `characters`,
|
||||
`app_settings`. The message row is:
|
||||
|
||||
```sql
|
||||
CREATE TABLE messages (
|
||||
id TEXT PRIMARY KEY,
|
||||
chat_id TEXT NOT NULL REFERENCES chats(id) ON DELETE CASCADE,
|
||||
role TEXT CHECK (role IN ('user','assistant')),
|
||||
content TEXT NOT NULL,
|
||||
attachments_json ..., image_request_json ..., generated_image_json ...,
|
||||
created_at TEXT NOT NULL
|
||||
);
|
||||
CREATE INDEX idx_messages_chat_created ON messages(chat_id, created_at);
|
||||
```
|
||||
|
||||
No `parent_id`, no branch, no take, no depth, no checkpoint table. Order is
|
||||
`created_at` (a TEXT timestamp), with ties broken by id.
|
||||
|
||||
## B2. Retry / Edit / Erase — all destructive, measured
|
||||
|
||||
```
|
||||
BEFORE (5 rows)
|
||||
a504cb34 assistant "Oh, Vale, you're just sitting in that rundown ..."
|
||||
e349d8b9 user "I pocket a brass key from the ashtray."
|
||||
9984692b assistant "You find a worn, faded brass key on the tabletop..."
|
||||
bb1ceb3f user "I drive to the address on its tag."
|
||||
a6e11115 assistant "Yes, you keep going. There, you can call your ..."
|
||||
|
||||
EDIT e349d8b9 in place → PATCH /api/chats/{id}/messages/{id}
|
||||
original text "brass key from the ashtray" still anywhere in DB? False
|
||||
|
||||
DELETE a6e11115?after=1 (what BOTH Retry and Erase call)
|
||||
rows 5 → 4
|
||||
|
||||
Scan of every column of every table for the deleted tail
|
||||
or the pre-edit text: none
|
||||
```
|
||||
|
||||
- **Edit** is `UPDATE messages SET content = ? WHERE id = ?`. The prior text is
|
||||
unrecoverable.
|
||||
- **Retry** (`retryLastTurn`, `page.tsx`) deletes the last assistant message and
|
||||
everything after it, then regenerates. The discarded attempt is gone.
|
||||
- **Erase** (`eraseLastTurn`) deletes the last exchange and the tail.
|
||||
- Old future turns are therefore always deleted, never retained.
|
||||
|
||||
## B3. What depends on the linear model
|
||||
|
||||
1. **`db.ts`** — 20 exported functions. `addMessage`, `updateMessageContent`,
|
||||
`deleteMessageAndAfter` and `getChat` are the history four.
|
||||
`getChat` returns a flat `messages` array and must become a lineage query.
|
||||
2. **`story-prompt.ts`** — windows by array index:
|
||||
`for (let i = messages.length - 1; ...)`, `messages.slice(dropped)`,
|
||||
`messages.slice(0, dropped)`. Becomes an ancestry walk.
|
||||
3. **The summary watermark is positional — the cost Phase 0A missed.**
|
||||
`chats.story_summary_count` records "how many of the chat's oldest messages
|
||||
the summary already covers", and `api/story/route.ts` compares
|
||||
`evicted.length > stored.coveredCount`, then summarizes
|
||||
`evicted.slice(stored.coveredCount)`. "The first N messages" is meaningless on
|
||||
a branch. Summaries must be re-anchored to a turn id and made per-lineage.
|
||||
4. **`page.tsx` is 3,991 lines** in one component; `api/story/route.ts` is 1,196.
|
||||
Retry/erase/edit live there as array slicing (`messages.slice(0, cutFrom)`).
|
||||
There is no take stepper, branch panel or tree view to build on.
|
||||
5. **No tests.** All of the above would be changed with no regression net.
|
||||
|
||||
## B4. Invasiveness estimate
|
||||
|
||||
To reach parent-linked turns, an active head, retained abandoned history, named
|
||||
checkpoints and lineage-safe summaries, Open Dungeon needs: a schema migration,
|
||||
a rewrite of its persistence read/write path, a rewrite of context assembly, a
|
||||
redesign of summarization, and new branch UI — with a test suite written first
|
||||
to make any of it safe. This is a from-scratch implementation of the hardest
|
||||
part of the specification, not a retrofit.
|
||||
|
||||
By comparison AI-DnD already has all of it except a non-destructive undo, and
|
||||
its reads funnel through one chokepoint (`lineage.Path.clause`).
|
||||
|
||||
## B5. Worth reusing as reference
|
||||
|
||||
- The reading experience: serif prose, prose-size control, the composer's
|
||||
Do/Say/Story modes.
|
||||
- Inline scene images driven by a narrator tool call (`generate_image`), with a
|
||||
local FLUX worker on `127.0.0.1:7869` or a user's ComfyUI on `8188`.
|
||||
- Character portraits reused as reference images for visual continuity, and fed
|
||||
back to the narrator as vision context. This is the most valuable idea here for
|
||||
`MEDIA-EXTENSION-CONTRACT.md`.
|
||||
- The worker HTTP protocol may be portable even though none of the UI is —
|
||||
Open Dungeon is Next.js, AI-DnD is React + FastAPI.
|
||||
|
||||
## B6. Other findings
|
||||
|
||||
- The `local` provider hard-codes five Gemma 4 QAT builds
|
||||
(`src/lib/text-models.ts`, `LOCAL_TEXT_MODELS`) and `/api/health` reports only
|
||||
those as installed. Any other local Ollama model must be reached through the
|
||||
"custom" provider. A genre-agnostic app wanting "pick any installed model"
|
||||
must lift this.
|
||||
- Next.js telemetry is enabled by default and printed its notice on first start.
|
||||
- 335 npm packages; `npm audit` reports 6 high, 1 low.
|
||||
@@ -0,0 +1,502 @@
|
||||
> **Planning review note (2026-09-01):** This is the coding agent's Phase 0B evidence report. It is retained substantially as delivered. Some recommendations/open questions in the report were superseded during planning review: ai-adventure is an implementation reference rather than the project specification; no automatic abandoned-history cleanup is required in v1; Open Dungeon media portability is deferred; and the architecture decision is closed in ADR 009/010 and `TECHNICAL-DESIGN.md`.
|
||||
|
||||
# Phase 0B — Local Validation Findings and Recommendation
|
||||
|
||||
**Date:** 2026-09-01
|
||||
**Status:** Complete, and updated after the follow-up undo/redo spike.
|
||||
Stop condition reached; no production work started.
|
||||
**Method:** Clean clones, pinned SHAs, documented installs, full test runs, real
|
||||
local Ollama inference, and targeted disposable experiments.
|
||||
|
||||
> **Update — the recommended spike was run and it passed.** Section E previously
|
||||
> recommended proving non-destructive Undo/Redo in AI-DnD before committing to
|
||||
> the fork. That spike is now complete: 3 files, +130/-31 lines, undo deletes
|
||||
> zero rows, redo round-trips exactly, writing below a moved-back head forks and
|
||||
> preserves the abandoned line, branch-scoped memory isolation survives, and the
|
||||
> suite goes 627/632 with all 5 failures asserting the deleted-row behavior that
|
||||
> was deliberately replaced. Section E has been rewritten accordingly, and the
|
||||
> open questions it resolved are marked in section F. Full detail:
|
||||
> `reports/PHASE-0B-UNDO-SPIKE.md`.
|
||||
>
|
||||
> **Update 2 — four follow-up checks run.** The world-state referee was exercised
|
||||
> against local models for the first time, and export/import, story-card lineage
|
||||
> safety and Postgres removability were settled. One new defect found (a
|
||||
> round-trip silently undoes an undo) and one design conclusion that changes what
|
||||
> we should build (§H). Detail: `reports/PHASE-0B-FOLLOWUP-CHECKS.md`.
|
||||
|
||||
Pinned commits (cloned 2026-09-01):
|
||||
|
||||
| Repo | SHA | Last commit |
|
||||
|---|---|---|
|
||||
| AI-DnD | `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` | 2026-08-31 |
|
||||
| Open Dungeon | `b0a79f96bf852be7b4e53908dff6a7f7c179da23` | 2026-07-06 |
|
||||
| ai-adventure | `873ea9180d5b611576cddb155921fc16a17ae88b` | 2026-07-21 |
|
||||
|
||||
Inference used for every live test: Ollama in Docker, `qwen2.5:0.5b` and
|
||||
`qwen2.5:3b-instruct` for narration, `nomic-embed-text` for embeddings.
|
||||
CPU only, no GPU.
|
||||
|
||||
---
|
||||
|
||||
## A. Executive Recommendation
|
||||
|
||||
**Fork AI-DnD as the production base. Confidence: high.**
|
||||
|
||||
**The Phase 0A recommendation held up, but for partly different reasons than it
|
||||
gave, and it was wrong about one important thing.**
|
||||
|
||||
What held: AI-DnD really does implement the expensive correctness work, and it
|
||||
survived every entanglement test Phase 0A worried about. The RPG machinery is
|
||||
optional, the scripting sandbox is removable, and the branch-scoped memory
|
||||
isolation — the single hardest requirement in the specification — works
|
||||
correctly against real local embeddings.
|
||||
|
||||
What did not hold: **Phase 0A's claim that AI-DnD's undo is non-destructive is
|
||||
false.** `POST /undo` hard-deletes rows. Retry is non-destructive; undo is not.
|
||||
There is no Redo. This is a direct conflict with `DECISIONS/005-branch-preserving-history.md`
|
||||
and specification §4.
|
||||
|
||||
That discovery does not change the decision, because the comparison is relative:
|
||||
AI-DnD needs one bounded change at a chokepoint that every read already passes
|
||||
through. Open Dungeon needs the entire non-destructive history subsystem built
|
||||
from nothing, under a 3,991-line component, with **zero automated tests**.
|
||||
|
||||
**That "one bounded change" is no longer an estimate.** The follow-up spike
|
||||
implemented it in 3 files and +130/-31 lines, and it works: undo deletes nothing,
|
||||
redo restores exactly, the abandoned tail becomes an ordinary branch, and memory
|
||||
isolation holds. Confidence in the fork decision moves from "high on the evidence"
|
||||
to "high and demonstrated".
|
||||
|
||||
The decision gate from `reports/PRELIMINARY-RECOMMENDATION.md` said to select
|
||||
AI-DnD unless a blocker is confirmed. All four were tested and none is a blocker:
|
||||
|
||||
| Phase 0A blocker | Result |
|
||||
|---|---|
|
||||
| Branching/state inseparable from RPG mechanics | **No.** A scenario-less adventure with empty world state played, branched, and retrieved memories correctly. |
|
||||
| Removing hosted/scripting destabilizes many tests | **No.** 19 of 632 (3%), and every one fails for instrumentation reasons, not coupling. |
|
||||
| Local-only still needs external services | **Partly — but both are packaging fixes.** `tiktoken` downloads its encoding from a Microsoft CDN; the SPA fetches Google Fonts. |
|
||||
| Dependency/security burden worse than expected | **No.** 80 packages total vs Open Dungeon's 335 with 6 high-severity advisories. |
|
||||
|
||||
---
|
||||
|
||||
## B. What We Learned About Each Candidate
|
||||
|
||||
### AI-DnD — recommended base
|
||||
|
||||
**What worked**
|
||||
|
||||
- `632 passed` in 236s, first try, no fixes needed.
|
||||
- Docker image builds clean; `docker compose` path is real.
|
||||
- Default model endpoint is already `http://localhost:11434/v1`. Connection test
|
||||
and model discovery against Ollama worked with no code change.
|
||||
- Streamed a full multi-turn story over SSE against local Ollama.
|
||||
- **Genre-neutrality is real.** An adventure created with `scenario_id: null`
|
||||
gets an empty world state and no RPG scaffolding, and plays normally. Noir
|
||||
prose, no fantasy code paths, no stats.
|
||||
- **Retry is genuinely non-destructive.** Retrying kept the replaced attempt as a
|
||||
sibling take at the same coordinate (rows 6 and 7 both present; both listed by
|
||||
`/variants`).
|
||||
- **Forking is genuinely non-destructive.** Adding an alternate take at a passed
|
||||
turn created branch 2 while every row of branch 1 survived, with
|
||||
`parent_branch_id` and `fork_depth` recorded.
|
||||
- **Branch-scoped memory isolation is correct.** Two memories were written on
|
||||
branch 1 (depths 5 and 11) from real `nomic-embed-text` embeddings. On branch 2,
|
||||
retrieval returned `used: []` and none of `brass key`, `revolver`, `lockbox`,
|
||||
`ashtray` or `bullets` appeared anywhere in the assembled prompt. The negative
|
||||
control passed: switching back to branch 1 returned both memories with
|
||||
similarity scores (0.8087) and the abandoned facts reappeared. Retrieval is
|
||||
working *and* correctly scoped.
|
||||
- **Prompt transparency is already there.** The context endpoint returns labeled
|
||||
sections with token costs: `narrator 57, ai_instructions 9, persona 15,
|
||||
plot_essentials 24, history 988, used_memories 143, length_hint 46`.
|
||||
|
||||
**What failed**
|
||||
|
||||
- **Undo hard-deletes.** Measured: rows 7 → 4, deleting ids 5, 6 **and 7** — that
|
||||
is, both the player action, the live AI take, *and the alternate take the
|
||||
earlier retry had preserved*. `delete_turn` also calls `memorybank.forget_node`.
|
||||
A retry's preserved alternative does not survive a later undo. There is no Redo
|
||||
endpoint. **Resolved in the follow-up spike** — see section G.
|
||||
- **Fails offline on a clean install.** First turn in a network-isolated container
|
||||
died with `NameResolutionError` for
|
||||
`openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken`. `tiktoken`
|
||||
fetches its BPE encoding on first use. Seeding the 1.7 MB file into the image
|
||||
made a full turn generate with no internet at all. Packaging fix, not
|
||||
architecture — but a hard blocker until fixed, and static review missed it.
|
||||
- **Runtime remote asset.** The built SPA still requests Google Fonts; the CSP
|
||||
explicitly allows `fonts.googleapis.com` and `fonts.gstatic.com`.
|
||||
|
||||
**Strongest reusable pieces**
|
||||
|
||||
Immutable parent-linked story tree with lineage entries; alternate takes; branch
|
||||
fork/switch/rename/delete; per-node world-state and script-state snapshots;
|
||||
branch-scoped memories with local embeddings; the Insights prompt snapshot;
|
||||
`ai-dnd-adventure-v2` whole-tree export; the 632-test suite; SSE streaming.
|
||||
|
||||
**Major architectural problems**
|
||||
|
||||
Destructive undo and no redo. No named checkpoints (branch names are the nearest
|
||||
thing). Multi-user/hosted concerns are spread across ten modules. `netguard.py`
|
||||
is the *inverse* of what we need — it blocks non-public endpoints in hosted mode
|
||||
and does nothing locally; we want a warning on non-loopback endpoints.
|
||||
|
||||
**Adaptation required**
|
||||
|
||||
Bounded and mostly subtractive. The one genuinely additive change is
|
||||
non-destructive undo/redo, now measured at 3 files and +130/-31 lines (section G).
|
||||
|
||||
### Open Dungeon — not recommended as base
|
||||
|
||||
**What worked**
|
||||
|
||||
Installed and ran on first try. Plays a real story against Ollama through its
|
||||
"custom" OpenAI-compatible provider. Clean browser UX. No runtime Google Fonts
|
||||
(`next/font` self-hosts at build time). Local SQLite, local image generation,
|
||||
character portrait continuity.
|
||||
|
||||
**What failed**
|
||||
|
||||
- **Every history operation is destructive, confirmed by measurement.**
|
||||
- Edit is an in-place `UPDATE messages SET content = ?`. After editing, the
|
||||
original text `brass key from the ashtray` was **gone from every column of
|
||||
every table**.
|
||||
- Retry and Erase both call `DELETE /messages/{id}?after=1`, which runs
|
||||
`DELETE FROM messages WHERE chat_id = ? AND (created_at > ? OR ...)`. Measured
|
||||
rows 5 → 4, with no archive anywhere.
|
||||
- The database has exactly four tables — `chats`, `messages`, `characters`,
|
||||
`app_settings`. No parent pointer, no branch, no take, no checkpoint.
|
||||
- **Zero automated tests.** No `test` script, and no test framework in
|
||||
`devDependencies` at all. Phase 0A listed this as "BUILD/VERIFY"; it is zero.
|
||||
- **The local provider hard-codes a five-model Gemma 4 allowlist**
|
||||
(`src/lib/text-models.ts`). Arbitrary local Ollama models are only reachable
|
||||
through the "custom" OpenAI-compatible path.
|
||||
- Next.js telemetry is on by default and printed its notice on first start.
|
||||
- 335 npm packages, 6 high-severity advisories.
|
||||
|
||||
**Cost of the branch retrofit — the Experiment B answer**
|
||||
|
||||
Larger than Phase 0A estimated, because of a dependency Phase 0A did not identify:
|
||||
|
||||
1. `messages` needs `parent_id`, `branch_id`, `depth`, and a live/take flag.
|
||||
2. Of 20 exported `db.ts` functions, the four history ones need rewriting, and
|
||||
`getChat` must stop returning a flat list and become a lineage query.
|
||||
3. `story-prompt.ts` windows history by **array index** (`messages.slice(dropped)`,
|
||||
`messages.length - 1`). All of it becomes an ancestry walk.
|
||||
4. **The rolling summary watermark is positional.** `chats.story_summary_count`
|
||||
means "how many of the chat's oldest messages the summary already covers", and
|
||||
`route.ts` compares `evicted.length > stored.coveredCount`. On a branch, "the
|
||||
first N messages" has no meaning. Summaries must be re-anchored to a turn id
|
||||
and made per-lineage. **Phase 0A did not identify this.**
|
||||
5. `page.tsx` is a single 3,991-line component holding retry/erase/edit as array
|
||||
slicing, and there is no take stepper, branch panel, or tree view to extend.
|
||||
6. All of the above lands with no test suite underneath it.
|
||||
|
||||
**Worth reusing as reference:** the reading UI, inline scene images, the
|
||||
character-portrait-as-reference-image continuity idea, and the scene/media
|
||||
architecture. None of it ports directly — Open Dungeon is Next.js and AI-DnD is
|
||||
React + FastAPI.
|
||||
|
||||
### ai-adventure — not the base, but promote it to the architectural reference
|
||||
|
||||
**What worked — this is the strongest result of the round**
|
||||
|
||||
- `Ran 76 tests ... OK` in 1.9s.
|
||||
- **The Ollama "adapter" is zero code.** `LMStudioBackend` is a generic
|
||||
OpenAI-compatible client hitting `/v1/chat/completions` and `/v1/models`. I
|
||||
changed two lines of `world.toml` — `base_url` and model `name` — and the
|
||||
doctor reported `[PASS] LM Studio reachable: http://127.0.0.1:11434` and
|
||||
`[PASS] Configured model visible: qwen2.5:3b-instruct`. A real turn then
|
||||
committed to the database with `status='committed'`, zero validation errors,
|
||||
and a `model_calls` row recording backend, model and attempt. **Phase 0A's
|
||||
"MODIFY — LM Studio primary provider" was wrong; it is a config change.**
|
||||
- **Its history model is exactly what the specification asks for.** Verified
|
||||
independently, not just by trusting its tests:
|
||||
- Undo deleted **0 rows**, rolled `has_brass_key` back from `False` to `True`,
|
||||
and the abandoned turn was still in the database.
|
||||
- Branching from the restored point produced `gate_open` on the branch and
|
||||
**did not leak it into the original session**.
|
||||
- Checkpoint restore rolled state back with the later turn still present.
|
||||
- Total turns in the database never decreased.
|
||||
- Schema is the target data model: `turns.parent_turn_id`, `status`,
|
||||
append-only `state_events`, `state_cache` keyed by `head_turn_id`,
|
||||
`model_calls` with request/response hashes, `named_checkpoints`, `summaries`
|
||||
anchored by `through_turn_id`, and lore in FTS5.
|
||||
- **The cleanest privacy posture by a wide margin.** Exactly one outbound call
|
||||
site in the entire codebase (`lm_studio.py`'s `urlopen`), no hardcoded remote
|
||||
host anywhere, and **one** third-party dependency (`pydantic`). It is the only
|
||||
candidate that already warns when `model.base_url` is not loopback — which is
|
||||
specification §12, implemented.
|
||||
|
||||
**What failed**
|
||||
|
||||
- With `qwen2.5:0.5b` the turn failed validation and **nothing was saved**. That
|
||||
is the model-proposes/app-validates discipline working correctly, but it means
|
||||
the typed-event contract needs a reasonably capable model; the freeform-prose
|
||||
candidates tolerate a weak one. `qwen2.5:3b-instruct` succeeded.
|
||||
- Product distance is real and large: terminal only, worlds authored as
|
||||
hand-written TOML and Markdown files rather than set up in-app, no streaming,
|
||||
no embeddings/semantic memory, no media, no import of user documents.
|
||||
|
||||
**Adaptation required:** building essentially the entire product on top of a
|
||||
correct core.
|
||||
|
||||
---
|
||||
|
||||
## C. Key Technical Findings
|
||||
|
||||
| Requirement | AI-DnD | Open Dungeon | ai-adventure |
|
||||
|---|---|---|---|
|
||||
| History/undo model | Tree + takes; **undo deletes**, no redo (both fixed in the spike, §G) | Flat list; **retry/edit/erase all destroy** | **Parent-linked, head pointer, undo deletes nothing** |
|
||||
| State rollback | Per-node snapshots, verified | None | **Event replay, verified** |
|
||||
| Memory isolation across branches | **Verified correct, with negative control** | N/A (no branches) | **Verified correct** |
|
||||
| Local Ollama | Default endpoint; verified streaming | Verified via "custom" provider; local provider allowlists 5 Gemma builds | **Verified, zero code change** |
|
||||
| Offline | **Blocked by tiktoken CDN fetch** until vendored; then fully offline | Runtime clean; build needs network; Next telemetry on by default | **Verifiably clean — 1 outbound call site, 1 dependency** |
|
||||
| Browser suitability | React SPA + FastAPI, already there | Next.js, already there | **None** |
|
||||
| Prompt inspection | **Per-action snapshot with labeled sections and token costs** | Not exposed | Prompt hashes; `store_prompts` off by default |
|
||||
| Imported knowledge | Story cards (keyword-triggered) — closest starting point | None | **Local FTS5 lore, deterministic** |
|
||||
| Named checkpoints | Absent (branch names only) | Absent | **Present** |
|
||||
| Future media | Absent | **Present and working** | Absent |
|
||||
| Tests | **632** | **0** | 76 |
|
||||
|
||||
---
|
||||
|
||||
## D. Important Surprises
|
||||
|
||||
Ordered by how much they should change the plan.
|
||||
|
||||
1. **AI-DnD's undo is destructive, and it destroys retry's preserved
|
||||
alternatives too.** `reports/PRELIMINARY-RECOMMENDATION.md` and
|
||||
`REUSE-MATRIX.md` both credit AI-DnD with non-destructive rollback. Retry
|
||||
earns that; undo does not. Measured: undo deleted the player action, the live
|
||||
take, and the sibling take a prior retry had kept. The follow-up spike has
|
||||
since replaced it with a head-cursor undo and added Redo (§G), so this is a
|
||||
corrected planning assumption rather than an outstanding defect.
|
||||
2. **AI-DnD cannot take a single turn on an air-gapped machine as shipped.**
|
||||
`tiktoken` fetches `cl100k_base` from a Microsoft CDN on first use. Proven by
|
||||
running it on a Docker `--internal` network, and proven fixed by seeding the
|
||||
file.
|
||||
3. **ai-adventure needs no Ollama adapter at all.** Two config lines. The
|
||||
"LM Studio" name is a misnomer for a generic OpenAI-compatible client.
|
||||
4. **Open Dungeon has zero automated tests** — not few, none, and no framework
|
||||
installed.
|
||||
5. **Open Dungeon's summary watermark is positional**, which adds a summary
|
||||
redesign to its branch retrofit that Phase 0A did not cost.
|
||||
6. **Open Dungeon's "local" provider only offers five Gemma 4 QAT builds.** A
|
||||
genre-agnostic app that lets the user pick any installed Ollama model needs
|
||||
that lifted.
|
||||
7. **AI-DnD fetches Google Fonts at runtime** and its CSP is written to allow it,
|
||||
violating specification §12's "no remote fonts".
|
||||
8. **AI-DnD's `netguard` is the opposite of our requirement** — it protects a
|
||||
hosted deployment from SSRF and is inert locally. ai-adventure has the
|
||||
loopback warning we actually want.
|
||||
9. Repo maturity differs sharply: AI-DnD is 172 commits / 3 authors, Open Dungeon
|
||||
45 / 1, ai-adventure 20 / 1. All three are young. Since we fork, this matters
|
||||
less than it would for a dependency, but there is no upstream to lean on.
|
||||
|
||||
---
|
||||
|
||||
## E. Recommendation for Next Step
|
||||
|
||||
**Fork AI-DnD and begin the controlled strip-down.** The blocking uncertainty is
|
||||
gone: the spike this section previously asked for has been run and passed
|
||||
(section G).
|
||||
|
||||
Recommended order, cheapest and most load-bearing first:
|
||||
|
||||
1. **Close the two offline blockers.** Vendor the `tiktoken` encoding and
|
||||
self-host the three font families, then re-run the air-gapped test. Until this
|
||||
is done the application does not satisfy specification §2, and it is a day's
|
||||
work.
|
||||
2. **Strip the hosted surface.** Remove multi-user/auth, the demo key, analytics,
|
||||
`render.yaml`, the Postgres path, and the quickjs sandbox. Re-instrument the
|
||||
~19 tests that use JS hooks to observe state, moving them onto the
|
||||
`world_state_after` snapshots that already exist.
|
||||
3. **Land the undo/redo work properly**, promoting the spike from a proof to a
|
||||
feature: fix the four rough edges it left (undo floor at the opening node,
|
||||
retry/`add_take` while behind the head, marking abandoned turns disposable,
|
||||
and the frontend Redo control and branch labelling) — **and add `headDepth` to
|
||||
the export bundle in the same change**, or a round-trip silently undoes every
|
||||
undo (§H).
|
||||
4. **Add named checkpoints**, using ai-adventure's
|
||||
`named_checkpoints(session_id, turn_id, name)` as the model. With an explicit
|
||||
head that can sit anywhere, a checkpoint is now just a named head position.
|
||||
5. **Add the loopback-endpoint warning**, modeled on ai-adventure's
|
||||
`_endpoint_warnings`.
|
||||
6. **Then** imported knowledge, and only after that, media.
|
||||
|
||||
Treat **ai-adventure as the written specification for state and history
|
||||
semantics** throughout. Its schema is the target shape and its 76 tests are the
|
||||
behavioral contract worth copying — the spike already reproduced its central
|
||||
property (undo that moves a head instead of deleting) and benefited from having
|
||||
that reference to aim at. §H extends this: when generalizing AI-DnD's world state
|
||||
into genre-neutral narrative state, take ai-adventure's **typed-event vocabulary**
|
||||
rather than AI-DnD's relative-delta vocabulary, because the delta protocol has a
|
||||
failure mode that validation cannot catch and that worsens as context grows.
|
||||
|
||||
One addition to the ordering above: whatever narrative-state engine we build must
|
||||
be **exercised at realistic context length against the models users will actually
|
||||
run**, not only in a clean harness. That is where the referee's weakness showed
|
||||
up, and it would not have appeared in a unit test.
|
||||
|
||||
Do not merge repositories. Open Dungeon remains a UX and media reference only.
|
||||
|
||||
## F. Open Questions
|
||||
|
||||
**Resolved by the undo/redo spike (§G):**
|
||||
|
||||
- ~~Does capping `Path.clause` by head depth break the memory and summary
|
||||
cursors?~~ **No.** `test_memory_settling`, `test_memory_nodes`,
|
||||
`test_memory_retrieval`, `test_memory_rewrite` and `test_history_window` all
|
||||
passed unchanged, and memory retrieval was verified end-to-end against real
|
||||
embeddings with a control.
|
||||
- ~~How should named checkpoints sit on top of takes and branches?~~ **Largely
|
||||
answered.** The spike introduces a head that can sit behind the tip, so a
|
||||
checkpoint becomes a named `(branch_id, depth)` position. Still needs a
|
||||
decision on whether restoring one forks immediately or on first write — the
|
||||
spike chose "on first write" for undo, and consistency argues for the same.
|
||||
|
||||
**Resolved by the follow-up checks (§H):**
|
||||
|
||||
- ~~What is the model capability floor?~~ **Characterized, not a blocker.** The
|
||||
delta protocol is within a 3B model's reach in a clean prompt (3/3 correct) but
|
||||
degraded to absolute values under the full application context. `0.5b` cannot
|
||||
do it at all. See §H for the design consequence.
|
||||
- ~~Can `psycopg`/Postgres be dropped cleanly?~~ **Yes.** Three places, no
|
||||
blockers: one branch in `database.py`, two migration helpers, one import inside
|
||||
analytics.
|
||||
- ~~Do story cards become imported knowledge, or is a new table needed?~~
|
||||
**A new table.** `StoryCard` carries no `branch_id`/`depth`, so it is not
|
||||
lineage-scoped and inherits neither branch nor undo isolation.
|
||||
- ~~Does the export format survive the history rework?~~ **No — and this is now a
|
||||
known defect to fix, not a question.** See §H.
|
||||
|
||||
**Still open:**
|
||||
|
||||
1. **When does abandoned history get cleaned up?** Nothing marks it disposable
|
||||
yet, and undone turns now accumulate rather than being deleted. This is the
|
||||
intended trade, but it needs a retention story.
|
||||
2. **Is Open Dungeon's image path portable at all?** It is Next.js calling a local
|
||||
FLUX worker over HTTP. The worker protocol may be reusable even though none of
|
||||
the UI is.
|
||||
3. **What is the real model recommendation for users?** The probe establishes a
|
||||
floor and a failure mode, not a ranking. A proper capability matrix across the
|
||||
models a user would actually run belongs in the build phase, measured at
|
||||
realistic context length.
|
||||
4. **Not tested in any round:** import/export round-trip *behavior* beyond the
|
||||
head-depth defect found by reading, multi-hour long-story context behavior, or
|
||||
any concurrency beyond the single-turn lock.
|
||||
|
||||
## G. Follow-Up Spike: Non-Destructive Undo/Redo — Passed
|
||||
|
||||
Run after the main round, on a disposable copy of AI-DnD `d72f7c1b`. Full detail
|
||||
in `reports/PHASE-0B-UNDO-SPIKE.md`.
|
||||
|
||||
**The change is three files, +130 / -31 lines**, roughly a fifth of it comments:
|
||||
|
||||
| File | Change | What |
|
||||
|---|---|---|
|
||||
| `context/lineage.py` | +26 / -3 | Read an uncapped lineage entry as capped at `self.tip` (the head), in `clause` and `contains`; add `uncapped()`. |
|
||||
| `routers/adventures/takes.py` | +74 / -27 | `undo_turn` moves `head_depth` instead of deleting; new `redo_turn`. |
|
||||
| `routers/adventures/turns.py` | +30 / -1 | `fork_if_behind_head` calls the existing `tree.branch_at` when a write lands below the head. |
|
||||
|
||||
It is this small because the pieces already existed: every read funnels through
|
||||
`lineage.path_of()`, `adventures.head_depth` was already a stored head, and
|
||||
`tree.branch_at` already created a branch that leaves a path at a depth without
|
||||
touching the line it leaves.
|
||||
|
||||
**Results against the success criteria:**
|
||||
|
||||
| Criterion | Result |
|
||||
|---|---|
|
||||
| Undo deletes zero rows | **Pass.** 6 → 6 rows across three undos; separately, 11 consecutive undos on a 23-action story with 23 actions still present. Clears the "at least five, unlimited preferred" requirement. |
|
||||
| Abandoned tail stays reachable as a branch | **Pass.** Undo-then-write forked branch 2 at fork depth 3; branch 1 kept all six of its rows, live and readable, and the branch switcher reaches it. |
|
||||
| Redo restores | **Pass.** Three redos returned the head from -1 to 5 and the transcript to all six actions, exactly. |
|
||||
| Branch-scoped memory isolation holds | **Pass, with control.** An embedded memory at depth 5 was retrieved at the tip (similarity 0.689), returned `used: []` once the head moved to depth 3, and became eligible again on redo — without being deleted. None of seven terms unique to the hidden turns appeared in the assembled prompt. |
|
||||
| Suite stays green apart from delete-behavior tests | **Pass.** `627 passed, 5 failed` (0.8%). All five assert deleted rows or pruned memories. The state assertions inside those same tests still pass — `assert adv.script_state == {"gold": 0}` succeeds in both `test_state_revert` failures, and the fork-boundary guard in `test_undo_stops_at_the_fork` is untouched. |
|
||||
|
||||
**Rough edges left for the real implementation** (none architectural): undo can
|
||||
walk past a user-written opening to an empty transcript; retry and `add_take`
|
||||
were not routed through the fork check; abandoned turns are not yet marked
|
||||
disposable; and the frontend has no Redo control.
|
||||
|
||||
**What this changes about the recommendation:** nothing about the choice, and the
|
||||
one thing about the plan — the strip-down can start now rather than after another
|
||||
investigation.
|
||||
|
||||
## H. Follow-Up Checks — Referee, Export, Cards, Postgres
|
||||
|
||||
Run after the spike. Detail in `reports/PHASE-0B-FOLLOWUP-CHECKS.md`.
|
||||
|
||||
### The finding that changes what we build
|
||||
|
||||
AI-DnD's world-state referee asks the model for **relative deltas**
|
||||
(`new = old + delta`, then clamped). Playing the seeded RPG scenario with
|
||||
`qwen2.5:3b-instruct`, the model sent **absolute values** — `{"player.hp": 92}`
|
||||
meaning "hp is now 92". The engine read `+92`, clamped to `+35`, then to the
|
||||
ceiling of 100:
|
||||
|
||||
```
|
||||
player.hp 100 -> 100 "did not move. It is already at its maximum of 100"
|
||||
```
|
||||
|
||||
**The player took an arrow to the shoulder and finished the turn at full health.**
|
||||
|
||||
The referee did its job — nothing was corrupted, and it fed corrective text back.
|
||||
But no validator can catch this, because `+92` is a legal proposal. The
|
||||
application stayed authoritative while the authoritative state silently stopped
|
||||
tracking the narration. That is the exact divergence specification §2 exists to
|
||||
prevent.
|
||||
|
||||
An isolated probe using AI-DnD's own `EMIT_RULE` and parser shows the protocol is
|
||||
**not** beyond a small model — `qwen2.5:3b-instruct` and `qwen2.5:7b-instruct`
|
||||
both got 3/3 correct relative deltas, and `qwen2.5:0.5b` emitted no block at all.
|
||||
The difference is context load: in a short prompt the 3B model follows the rule,
|
||||
and under the full application prompt it drifts to absolutes.
|
||||
|
||||
**Consequence for the design:** when we generalize AI-DnD's world state into
|
||||
genre-neutral narrative state, take **ai-adventure's typed-event vocabulary**
|
||||
(`set_flag key=… value=…` — explicit and absolute) rather than AI-DnD's delta
|
||||
vocabulary. The delta protocol has a failure mode that is invisible to validation
|
||||
and worsens with context length; the event protocol cannot express the mistake.
|
||||
This is the second time in this round that ai-adventure's core has turned out to
|
||||
be the right specification to build to.
|
||||
|
||||
### One new defect
|
||||
|
||||
**Export/import silently redoes an undone story.** `bundle.py:_point_the_head`
|
||||
recomputes `head_depth = max(depths)` on import, and the bundle format carries
|
||||
`headBranch` but not the head depth. That was correct when undo deleted rows; with
|
||||
the spike's non-destructive undo, a round-trip restores every abandoned turn. Fix
|
||||
is one optional `headDepth` field plus honouring it on import, falling back to
|
||||
`max(depths)` so existing `v2` files still load. It must land with the undo work,
|
||||
not after it.
|
||||
|
||||
### Two smaller answers
|
||||
|
||||
- **Story cards are not lineage-scoped.** `StoryCard` has no `branch_id` or
|
||||
`depth`, unlike `Memory`. Harmless today because cards are authored rather than
|
||||
derived — but any imported-knowledge source that can be created from story
|
||||
content must carry those coordinates and be read through `lineage.Path.clause`,
|
||||
or it inherits neither branch nor undo isolation. Argues for a new table rather
|
||||
than extending `story_cards`.
|
||||
- **Postgres is cleanly removable.** One branch in `database.py`, two migration
|
||||
helpers, one import inside analytics. `psycopg[binary]` then drops out.
|
||||
|
||||
## Supporting Evidence
|
||||
|
||||
- `reports/PHASE-0B-BASELINE.md` — installs, test runs, ports, storage, versions.
|
||||
- `reports/PHASE-0B-AI-DND-EXPERIMENT.md` — branch/undo/retry measurements,
|
||||
memory isolation with negative control, scripting removal.
|
||||
- `reports/PHASE-0B-OPEN-DUNGEON-HISTORY.md` — destructive-history proof and
|
||||
retrofit cost map.
|
||||
- `reports/PHASE-0B-AI-ADVENTURE-OLLAMA.md` — zero-change Ollama run and the
|
||||
fixture checks.
|
||||
- `reports/PHASE-0B-OFFLINE-NETWORK.md` — offline runs and egress inventory.
|
||||
- `reports/PHASE-0B-UNDO-SPIKE.md` — the non-destructive undo/redo spike: the
|
||||
diff, the measurements, and the rough edges it left.
|
||||
- `reports/PHASE-0B-FOLLOWUP-CHECKS.md` — the world-state referee under local
|
||||
models, the export/import head defect, story-card lineage safety, Postgres.
|
||||
|
||||
Working tree, scripts and databases from both rounds are under `phase0b/`
|
||||
(untracked scratch, safe to delete). The spike itself is `phase0b/spike/`, a
|
||||
disposable copy of the AI-DnD backend — it is a proof, not a branch to merge.
|
||||
@@ -0,0 +1,214 @@
|
||||
# Phase 0B — Spike: Non-Destructive Undo/Redo in AI-DnD
|
||||
|
||||
**Date:** 2026-09-01
|
||||
**Base:** AI-DnD `d72f7c1b`, disposable copy at `phase0b/spike/`
|
||||
**Verdict:** **Passed on every criterion.** The retrofit is small, centralized,
|
||||
and costs 5 of 632 tests — all of which assert the deleted-row behavior that was
|
||||
deliberately replaced.
|
||||
|
||||
## The question
|
||||
|
||||
Can undo move a head cursor backward without deleting rows, does writing below a
|
||||
moved-back head fork a branch, and does Redo work — without breaking
|
||||
branch-scoped memory isolation?
|
||||
|
||||
## What the change turned out to be
|
||||
|
||||
Three files, **+130 / -31 lines**, and roughly a fifth of the additions are
|
||||
comments and the new endpoint's docstring.
|
||||
|
||||
| File | Change |
|
||||
|---|---|
|
||||
| `app/context/lineage.py` | +26 / -3 |
|
||||
| `app/routers/adventures/takes.py` | +74 / -27 |
|
||||
| `app/routers/adventures/turns.py` | +30 / -1 |
|
||||
|
||||
### 1. Read path — cap the lineage at the head
|
||||
|
||||
The stored lineage records the newest branch entry **uncapped** (`max_depth =
|
||||
None`, meaning "through to the tip"), and older entries capped at their fork
|
||||
depths. `Path.clause` and `Path.contains` now read an uncapped entry as capped at
|
||||
`self.tip`, which is `adventure.head_depth`:
|
||||
|
||||
```python
|
||||
def _cap(self, max_depth: int | None) -> int | None:
|
||||
return self.tip if max_depth is None else max_depth
|
||||
```
|
||||
|
||||
This is the whole read-side change. It works because **every read already funnels
|
||||
through `lineage.path_of()`**, so one substitution moves the entire application —
|
||||
transcript paging, context assembly, `attempts.preceding`, and memory retrieval —
|
||||
onto a head that can sit behind the deepest node.
|
||||
|
||||
A `Path.uncapped()` helper was added for the two callers that must deliberately
|
||||
look past the head (redo, and the fork check).
|
||||
|
||||
`prefix_covering` already computed `top = self.tip if max_depth is None else
|
||||
max_depth`, so the windowing arithmetic was consistent with this reading before
|
||||
the change; nothing there needed touching.
|
||||
|
||||
### 2. Undo — move the head instead of deleting
|
||||
|
||||
`undo_turn` keeps its turn lock, its "nothing to undo" guard and its
|
||||
fork-boundary guard, keeps `attempts.preceding` + `attempts.restore_state` for
|
||||
the state rollback, and replaces the two `delete_turn` calls plus
|
||||
`tree.refresh_head` with one assignment:
|
||||
|
||||
```python
|
||||
adventure.head_depth = (first_removed.depth or 0) - 1
|
||||
```
|
||||
|
||||
No `db.delete`. No `memorybank.forget_node`.
|
||||
|
||||
### 3. Write path — fork when writing below the head
|
||||
|
||||
`turns.fork_if_behind_head` runs in `create_action` beside the existing
|
||||
`_move_to_after`. If any live node on the path sits deeper than the head, it calls
|
||||
the **already-existing** `tree.branch_at(db, adventure, adventure.head_depth)`,
|
||||
which creates an empty branch leaving the path at that depth and does not touch
|
||||
the branch being left. When the head is at the tip it does nothing, so a story
|
||||
that is never undone forks exactly as often as before.
|
||||
|
||||
### 4. Redo — walk the head forward
|
||||
|
||||
New `POST /adventures/{id}/redo`. Reads the path *uncapped*, takes the next node
|
||||
deeper than the head, and advances over the whole turn (player action + AI reply)
|
||||
rather than half of it. 400 when there is nothing to redo.
|
||||
|
||||
## Results
|
||||
|
||||
### Undo deletes nothing; redo round-trips exactly
|
||||
|
||||
Six actions across three turns, live Ollama:
|
||||
|
||||
```
|
||||
AFTER 3 TURNS head=(branch 1, depth 5) rows 1..6 all live
|
||||
visible: [1 story, 2 ai, 3 do, 4 ai, 5 do, 6 ai]
|
||||
|
||||
UNDO x1 rows 6 -> 6 DELETED=0 head=(1, 3)
|
||||
visible: [1 story, 2 ai, 3 do, 4 ai]
|
||||
|
||||
UNDO x2 more head=(1, -1)
|
||||
visible: [] rows still 6
|
||||
|
||||
REDO x3 head=(1, 5) rows 6
|
||||
visible: [1 story, 2 ai, 3 do, 4 ai, 5 do, 6 ai]
|
||||
```
|
||||
|
||||
Elsewhere in the run, **11 consecutive undos** were performed on a 23-action
|
||||
story with 0 rows deleted, which clears the specification's "at least five undo
|
||||
operations, unlimited preferred" comfortably.
|
||||
|
||||
### Writing below a moved-back head forks, and the abandoned line survives
|
||||
|
||||
```
|
||||
UNDO once, then write a different continuation:
|
||||
|
||||
1 story br1 d0 live 5 do br1 d4 live <- the abandoned tail,
|
||||
2 ai br1 d1 live 6 ai br1 d5 live still live, still on br1
|
||||
3 do br1 d2 live
|
||||
4 ai br1 d3 live 7 do br2 d4 live <- the new line
|
||||
8 ai br2 d5 live
|
||||
|
||||
branches: {id 1, parent null, fork_depth null, own_actions 6, is_head False}
|
||||
{id 2, parent 1, fork_depth 3, own_actions 2, is_head True}
|
||||
|
||||
visible on the new line: [1, 2, 3, 4, 7, 8]
|
||||
switch to branch 1 : [1, 2, 3, 4, 5, 6]
|
||||
```
|
||||
|
||||
This is `DECISIONS/005-branch-preserving-history.md` behavior: the abandoned
|
||||
future is retained as an alternate branch rather than erased, and it is reachable
|
||||
through the existing branch switcher with no new UI concept.
|
||||
|
||||
### Memory isolation still holds — the important regression check
|
||||
|
||||
A 23-action story with a real memory embedded via `nomic-embed-text` at depth 5.
|
||||
Probed with terms unique to the turns an undo hides, plus terms from the turns it
|
||||
keeps:
|
||||
|
||||
```
|
||||
AT TIP (control) head_depth=22 memories_used=1 (similarity 0.689)
|
||||
hidden-turn terms leaking: ['six bullets','warehouse nine','iron stairs',
|
||||
'ledger','docks','office desk','inside my coat']
|
||||
(correct — nothing is hidden at the tip)
|
||||
visible-turn terms present: ['rainy city','desk drawer']
|
||||
|
||||
11 undos, 0 rows deleted (actions still 23)
|
||||
|
||||
AFTER UNDO head_depth=3 memories_used=0
|
||||
hidden-turn terms leaking: NONE
|
||||
visible-turn terms present: ['rainy city','desk drawer']
|
||||
|
||||
REDO back to tip head_depth=22 memories_used=1 (similarity 0.689)
|
||||
```
|
||||
|
||||
The memory at depth 5 stops being retrieved when the head moves behind it and
|
||||
becomes eligible again on redo — **without being deleted**. That is the property
|
||||
the destructive undo bought by calling `forget_node`, recovered for free, because
|
||||
`Memory` rows carry `branch_id` and `depth` and retrieval already goes through
|
||||
the same capped clause.
|
||||
|
||||
> One false alarm worth recording: an early probe reported `revolver` leaking. It
|
||||
> was legitimate visible history — actions 22 and 23 sit at depths 2 and 3, at or
|
||||
> below the head. A second false alarm came from probing the context *after* the
|
||||
> script had already redone to the tip. Both were resolved by re-probing inside a
|
||||
> single script with an explicit control, which is why the table above reports the
|
||||
> control run alongside the result.
|
||||
|
||||
### Test suite: 627 passed, 5 failed
|
||||
|
||||
```
|
||||
5 failed, 627 passed in 184s (0.8% of the suite)
|
||||
|
||||
tests/test_attempt_siblings.py::test_undo_takes_every_attempt_with_it
|
||||
tests/test_branch_forking.py::test_undo_stops_at_the_fork
|
||||
tests/test_state_revert.py::test_undo_reverts_state_to_before_the_turn
|
||||
tests/test_state_revert.py::test_undo_of_bare_continue_uses_the_node_in_front
|
||||
tests/test_state_revert.py::test_undo_prunes_memory_covering_removed_actions
|
||||
```
|
||||
|
||||
Every one asserts that rows or memories were **deleted**:
|
||||
|
||||
- `assert [a.type for a in adv.actions] == ["start"]` (three of them)
|
||||
- `assert len(_rows(...)) == rows_before - 1`
|
||||
- `assert texts == {"k"}` — the pruned-memory set
|
||||
|
||||
Crucially, the *state* assertions inside those same tests still pass. In both
|
||||
`test_state_revert` cases, `assert adv.script_state == {"gold": 0}` succeeds and
|
||||
only the row-count line fails — state rollback is intact. And
|
||||
`test_undo_stops_at_the_fork` still enforces its real subject: the guard
|
||||
refusing to undo into a parent branch is untouched and still returns 400.
|
||||
|
||||
Also worth noting for the open question raised in the recommendation: the four
|
||||
memory suites — `test_memory_settling`, `test_memory_nodes`,
|
||||
`test_memory_retrieval`, `test_memory_rewrite` — and `test_history_window` all
|
||||
**passed unchanged**. The cursor arithmetic did not need re-deriving.
|
||||
|
||||
## Rough edges found, not fixed
|
||||
|
||||
1. **Undo can empty the story.** The original guard refuses when the newest node
|
||||
is `type == "start"`, but an adventure opened with a user-written `story`
|
||||
action has no `start` node, so undo walks to `head_depth = -1` and the
|
||||
transcript renders empty. Redo recovers it, but the floor should be the
|
||||
opening node rather than `-1`.
|
||||
2. **Retry and `add_take` were left on the old path.** The spike only routed the
|
||||
ordinary write through `fork_if_behind_head`. Retrying while the head is
|
||||
behind the tip is not yet defined and needs a decision — most likely the same
|
||||
fork.
|
||||
3. **No pruning story yet.** Abandoned turns now accumulate. That matches the
|
||||
specification ("retained but marked disposable; cleanup later"), but nothing
|
||||
marks them disposable and no cleanup exists.
|
||||
4. **Frontend untouched.** There is no Redo button and the branch panel does not
|
||||
distinguish a line abandoned by undo from one forked deliberately.
|
||||
|
||||
## Conclusion
|
||||
|
||||
The retrofit lands where the recommendation predicted: one chokepoint
|
||||
(`lineage.Path.clause`), one stored head that already existed
|
||||
(`adventures.head_depth`), and one branch primitive that already existed
|
||||
(`tree.branch_at`). The unknown was whether the depth cap would disturb the
|
||||
memory and summary cursors, and it does not.
|
||||
|
||||
The fork decision is sound. The remaining work on this axis is the four rough
|
||||
edges above plus named checkpoints, not a redesign.
|
||||
@@ -0,0 +1,155 @@
|
||||
# Phase 0A Preliminary Recommendation
|
||||
|
||||
**Historical status:** Superseded by Phase 0B runtime validation and ADR 009. Retained as Phase 0A evidence.
|
||||
|
||||
**Date:** 2026-09-01
|
||||
**Status:** Static recommendation; pending Phase 0B clone/build/runtime experiments.
|
||||
|
||||
## Recommendation
|
||||
|
||||
### First choice to validate: AI-DnD
|
||||
|
||||
Use **AI-DnD as the preliminary production fork candidate**.
|
||||
|
||||
This is a change from the earlier slight preference for Open Dungeon.
|
||||
|
||||
The deciding evidence is not feature count; it is **where the hard architectural work already lives**.
|
||||
|
||||
AI-DnD already implements:
|
||||
- parent/lineage story tree,
|
||||
- alternate takes,
|
||||
- non-destructive retry,
|
||||
- branch-aware state rollback,
|
||||
- prompt snapshots,
|
||||
- context windowing,
|
||||
- memory bank + embeddings,
|
||||
- story cards,
|
||||
- complete tree export/import,
|
||||
- local Ollama,
|
||||
- a substantial automated test suite.
|
||||
|
||||
Those are precisely the systems most dangerous to retrofit after a linear chat application has accumulated behavior.
|
||||
|
||||
## Why Open Dungeon is second
|
||||
|
||||
Open Dungeon remains the best direct match to the desired *product*:
|
||||
- simple browser fiction interface,
|
||||
- Ollama,
|
||||
- local SQLite,
|
||||
- strong local image path,
|
||||
- visual continuity.
|
||||
|
||||
However, its present persistence semantics are linear and destructive:
|
||||
- prior messages can be updated,
|
||||
- retry/erase deletes the selected message and the story tail.
|
||||
|
||||
To satisfy the specification, we would need to introduce turn parentage/branches, branch-specific summaries/state, and non-destructive editing underneath features already written around a list. That is foundational work.
|
||||
|
||||
Open Dungeon should remain the fallback base if Codex proves AI-DnD's RPG/cloud systems are too entangled to remove.
|
||||
|
||||
## Why ai-adventure is third
|
||||
|
||||
ai-adventure has the cleanest *architecture* for state authority and local-only trust:
|
||||
- typed model proposals,
|
||||
- validation,
|
||||
- atomic commit,
|
||||
- append-only events,
|
||||
- replay,
|
||||
- checkpoints,
|
||||
- branches,
|
||||
- deterministic local lore,
|
||||
- explicit minimal network posture.
|
||||
|
||||
Its problem is product distance:
|
||||
- CLI presentation,
|
||||
- LM Studio primary provider,
|
||||
- no browser application,
|
||||
- no media,
|
||||
- less semantic long-memory machinery.
|
||||
|
||||
It should be the architectural control against which the selected browser fork is judged.
|
||||
|
||||
## Do not merge repositories
|
||||
|
||||
The recommendation is not to combine several projects mechanically.
|
||||
|
||||
Fork one project and re-implement selected ideas using compatible patterns/code only where justified.
|
||||
|
||||
A merged codebase would import:
|
||||
- incompatible assumptions,
|
||||
- duplicate persistence models,
|
||||
- different provider abstractions,
|
||||
- unnecessary dependencies,
|
||||
- licensing complexity.
|
||||
|
||||
## Proposed target architecture after Phase 0B
|
||||
|
||||
If AI-DnD passes validation:
|
||||
|
||||
### Retain
|
||||
- React/Vite browser shell,
|
||||
- FastAPI service boundary,
|
||||
- SQLite,
|
||||
- story tree/lineage,
|
||||
- state snapshots,
|
||||
- context budgeting,
|
||||
- Memory Bank,
|
||||
- story cards,
|
||||
- Insights,
|
||||
- Ollama path,
|
||||
- export/import,
|
||||
- relevant tests.
|
||||
|
||||
### Remove
|
||||
- multi-user/hosted auth,
|
||||
- demo keys/rate-limit hosting features,
|
||||
- Render/Neon path,
|
||||
- analytics,
|
||||
- cloud model providers,
|
||||
- QuickJS scripting,
|
||||
- AI-Dungeon compatibility not needed for core stories,
|
||||
- RPG-only presentation/mechanics.
|
||||
|
||||
### Generalize
|
||||
- world-state engine -> narrative state/facts/entities/threads,
|
||||
- Story Cards -> local knowledge sources with authority/provenance,
|
||||
- Memory Bank -> branch-safe story memory with Canon/Scene/Heuristic trust classes,
|
||||
- scenario -> genre-neutral campaign/story profile.
|
||||
|
||||
### Add
|
||||
- local document ingestion,
|
||||
- Canon / Reference / Inspiration source classification,
|
||||
- local lexical + optional Ollama semantic retrieval,
|
||||
- scene snapshots,
|
||||
- visual character/location fields,
|
||||
- media asset/job records,
|
||||
- media-provider interface,
|
||||
- later local image/video adapters.
|
||||
|
||||
## Phase 0B should be narrow
|
||||
|
||||
Codex should not repeat the broad research.
|
||||
|
||||
It should validate three concrete engineering hypotheses:
|
||||
|
||||
### Hypothesis 1 — AI-DnD can be stripped safely
|
||||
Prove local Ollama story/branch/memory operation still works after disabling/removing hosted/cloud/analytics/scripting paths and running with minimal RPG state.
|
||||
|
||||
### Hypothesis 2 — Open Dungeon branch retrofit is materially larger
|
||||
Map exactly how many DB functions/API routes/UI components/summary behaviors must change to make retry/edit non-destructive and branch-aware.
|
||||
|
||||
### Hypothesis 3 — ai-adventure is viable but farther from product
|
||||
Prove Ollama adapter effort is small and estimate the service/browser wrapper effort without starting production UI development.
|
||||
|
||||
Then choose the fork based on measured modification cost.
|
||||
|
||||
## Decision gate
|
||||
|
||||
Select AI-DnD unless Phase 0B finds one of these blockers:
|
||||
|
||||
- branching/state logic is inseparable from RPG mechanics,
|
||||
- removing hosted/scripting paths destabilizes a large percentage of tests,
|
||||
- local-only configuration still requires hard-to-remove external services,
|
||||
- dependency/security burden is materially worse than static review suggests.
|
||||
|
||||
If any blocker is confirmed, select Open Dungeon and explicitly budget a story-tree/persistence rewrite as the first production architecture milestone.
|
||||
@@ -0,0 +1,618 @@
|
||||
# Adventure Storyteller — Phase 0 Research Plan
|
||||
|
||||
**Status:** Complete — Phase 0 closed 2026-09-01
|
||||
**Phase:** 0 — Research, Validation & Architecture
|
||||
**Goal:** Determine what to build, what to fork/reuse, and finalize the technical design before production implementation begins.
|
||||
|
||||
**Outcome:** AI-DnD selected as the production base; non-destructive head-cursor history and explicit typed narrative-state events selected; production milestones now defined in `BUILD-MILESTONES.md`.
|
||||
|
||||
## 1. Why Phase 0 Exists
|
||||
|
||||
The project has several promising open-source starting points. They differ substantially in:
|
||||
|
||||
- browser UX,
|
||||
- persistence model,
|
||||
- branching semantics,
|
||||
- long-term memory,
|
||||
- local knowledge retrieval,
|
||||
- Ollama support,
|
||||
- dependency footprint,
|
||||
- privacy/network behavior,
|
||||
- game-specific assumptions,
|
||||
- licensing,
|
||||
- test quality.
|
||||
|
||||
A detailed production milestone plan written before inspecting the code would rely on guesses.
|
||||
|
||||
Phase 0 therefore ends when we can answer:
|
||||
|
||||
> What exact codebase and architecture should be used for the production storyteller?
|
||||
|
||||
No production feature work should begin before that decision unless explicitly authorized.
|
||||
|
||||
## 1A. Phase 0 Completion Summary
|
||||
|
||||
Phase 0A static research and Phase 0B local validation are complete.
|
||||
|
||||
Final dispositions:
|
||||
|
||||
- AI-DnD — production fork/base at `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`.
|
||||
- ai-adventure — primary implementation reference for authoritative typed state events, checkpoint/head/replay semantics, and narrow local-only behavior.
|
||||
- Open Dungeon — UX and future-media reference only.
|
||||
|
||||
Critical Phase 0B prototype results:
|
||||
|
||||
- non-destructive Undo/Redo with a movable active head was demonstrated on AI-DnD without deleting history and without breaking branch-scoped memory isolation,
|
||||
- export/import must preserve active head position,
|
||||
- AI-DnD's relative-delta state protocol should not be retained as the generic narrative-state contract,
|
||||
- imported knowledge should be a separate subsystem rather than Story Cards,
|
||||
- runtime offline hardening is required for tokenizer data and fonts.
|
||||
|
||||
The original milestone text below is retained as the research execution record.
|
||||
|
||||
## 2. Candidate Repositories
|
||||
|
||||
Initial candidates:
|
||||
|
||||
1. **Open Dungeon**
|
||||
- Repository: `newideas99/open-dungeon`
|
||||
- Interest: browser-first interactive-fiction UX, Ollama, SQLite.
|
||||
|
||||
2. **AI-DnD**
|
||||
- Repository: `parththakkar106/AI-DnD`
|
||||
- Interest: browser UI, story tree, rollback, memory, story cards, prompt inspection.
|
||||
|
||||
3. **Local Adventure Engine / ai-adventure**
|
||||
- Repository: `CaoRuiming/ai-adventure`
|
||||
- Interest: append-only state, checkpoints, branching, privacy-focused architecture, deterministic replay.
|
||||
|
||||
4. **aiMultiFool**
|
||||
- Repository: exact upstream URL to be confirmed during inventory.
|
||||
- Interest: local semantic memory/RAG and context inspection.
|
||||
|
||||
Reference projects:
|
||||
|
||||
- SillyTavern
|
||||
- RisuAI
|
||||
- KoboldAI
|
||||
- Chronicler
|
||||
- other credible projects discovered during Phase 0.
|
||||
|
||||
## 3. Research Workspace
|
||||
|
||||
Create a dedicated workspace such as:
|
||||
|
||||
```text
|
||||
adventure-storyteller-research/
|
||||
├── candidates/
|
||||
│ ├── open-dungeon/
|
||||
│ ├── ai-dnd/
|
||||
│ ├── ai-adventure/
|
||||
│ └── aimultifool/
|
||||
├── notes/
|
||||
├── experiments/
|
||||
├── reports/
|
||||
└── inventory/
|
||||
```
|
||||
|
||||
Do not copy source code from one project into another during initial analysis.
|
||||
|
||||
Each candidate should remain a clean upstream clone or worktree.
|
||||
|
||||
Record:
|
||||
|
||||
- upstream URL,
|
||||
- upstream default branch,
|
||||
- commit SHA examined,
|
||||
- release/tag if applicable,
|
||||
- clone date,
|
||||
- license,
|
||||
- language/framework,
|
||||
- build tooling,
|
||||
- runtime services,
|
||||
- expected local ports.
|
||||
|
||||
## 4. Phase Rules
|
||||
|
||||
During Phase 0:
|
||||
|
||||
- do not begin production feature development,
|
||||
- do not merge candidate codebases,
|
||||
- do not remove features from candidate repos,
|
||||
- do not commit speculative refactors,
|
||||
- small disposable experiments are allowed,
|
||||
- experiments must be isolated and clearly documented,
|
||||
- candidate repos should remain easy to reset to upstream,
|
||||
- every conclusion should cite observed code/config/test behavior.
|
||||
|
||||
## 5. Milestone R0 — Research Workspace and Inventory
|
||||
|
||||
### Objective
|
||||
|
||||
Create the research environment and establish a reproducible inventory of all candidates.
|
||||
|
||||
### Tasks
|
||||
|
||||
- create the research workspace,
|
||||
- clone all initial candidates,
|
||||
- record exact upstream commits,
|
||||
- locate and record licenses,
|
||||
- inventory languages/frameworks,
|
||||
- inventory package managers,
|
||||
- inventory database/storage dependencies,
|
||||
- inventory model/provider dependencies,
|
||||
- inventory frontend/backend separation,
|
||||
- record build/run instructions,
|
||||
- identify existing tests,
|
||||
- identify documentation directories,
|
||||
- identify migrations/schema definitions,
|
||||
- identify obvious telemetry/cloud integrations.
|
||||
|
||||
### Deliverables
|
||||
|
||||
- `inventory/candidates.md`
|
||||
- `inventory/licenses.md`
|
||||
- `inventory/dependencies.md`
|
||||
- `inventory/build-instructions.md`
|
||||
- machine-readable candidate metadata if useful.
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
All serious candidates are locally available and reproducibly identified.
|
||||
|
||||
## 6. Milestone R1 — Build and Run Candidates
|
||||
|
||||
### Objective
|
||||
|
||||
Verify actual behavior rather than relying on README claims.
|
||||
|
||||
### Tasks
|
||||
|
||||
For each serious candidate:
|
||||
|
||||
- install dependencies,
|
||||
- build successfully where applicable,
|
||||
- start locally,
|
||||
- create a minimal story,
|
||||
- confirm persistence after restart,
|
||||
- test Ollama directly where supported,
|
||||
- identify how model configuration works,
|
||||
- record application ports,
|
||||
- identify data locations,
|
||||
- inspect browser developer/network activity for unexpected outbound requests where applicable,
|
||||
- record startup failures or undocumented requirements.
|
||||
|
||||
For projects not supporting Ollama:
|
||||
|
||||
- determine adapter/interface boundary,
|
||||
- do not yet permanently modify the project.
|
||||
|
||||
### Deliverables
|
||||
|
||||
Per candidate:
|
||||
|
||||
```text
|
||||
reports/runtime-<candidate>.md
|
||||
```
|
||||
|
||||
Include:
|
||||
|
||||
- exact commands,
|
||||
- success/failure,
|
||||
- screenshots only if useful,
|
||||
- local services used,
|
||||
- observed storage files,
|
||||
- observed network activity,
|
||||
- known blockers.
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
Each serious candidate has either been run successfully or has a documented reason it cannot reasonably be evaluated.
|
||||
|
||||
## 7. Milestone R2 — Source Architecture Review
|
||||
|
||||
### Objective
|
||||
|
||||
Understand how each candidate actually works internally.
|
||||
|
||||
### Review Areas
|
||||
|
||||
#### Browser/UI
|
||||
- framework,
|
||||
- state management,
|
||||
- streaming,
|
||||
- transcript representation,
|
||||
- campaign navigation,
|
||||
- edit/retry behavior,
|
||||
- extensibility for future media.
|
||||
|
||||
#### Backend/service layer
|
||||
- routing/API design,
|
||||
- model invocation boundary,
|
||||
- background jobs,
|
||||
- validation boundaries.
|
||||
|
||||
#### Persistence
|
||||
- database type,
|
||||
- schema,
|
||||
- migrations,
|
||||
- turn representation,
|
||||
- snapshots,
|
||||
- event log,
|
||||
- transactions,
|
||||
- branch representation.
|
||||
|
||||
#### Story history
|
||||
- linear vs tree,
|
||||
- retry semantics,
|
||||
- undo semantics,
|
||||
- destructive vs non-destructive restore,
|
||||
- branch naming/navigation.
|
||||
|
||||
#### Context
|
||||
- prompt assembly,
|
||||
- recent history,
|
||||
- summaries,
|
||||
- token budgeting,
|
||||
- author notes/system rules.
|
||||
|
||||
#### Memory
|
||||
- summaries,
|
||||
- vector retrieval,
|
||||
- keyword retrieval,
|
||||
- entity state,
|
||||
- old-turn retrieval.
|
||||
|
||||
#### Lore/knowledge
|
||||
- import formats,
|
||||
- chunking,
|
||||
- story cards/world info,
|
||||
- semantic retrieval,
|
||||
- provenance.
|
||||
|
||||
#### Tests
|
||||
- unit tests,
|
||||
- integration tests,
|
||||
- migration tests,
|
||||
- model mocks,
|
||||
- coverage of state/rollback.
|
||||
|
||||
### Deliverables
|
||||
|
||||
- `reports/architecture-open-dungeon.md`
|
||||
- `reports/architecture-ai-dnd.md`
|
||||
- `reports/architecture-ai-adventure.md`
|
||||
- `reports/architecture-aimultifool.md`
|
||||
- `reports/architecture-comparison.md`
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
We can explain each candidate's architecture without relying on marketing descriptions.
|
||||
|
||||
## 8. Milestone R3 — Privacy and Network Review
|
||||
|
||||
### Objective
|
||||
|
||||
Determine what must be removed, disabled, or isolated to satisfy the local-only requirement.
|
||||
|
||||
### Search For
|
||||
|
||||
- OpenAI,
|
||||
- OpenRouter,
|
||||
- Groq,
|
||||
- Anthropic,
|
||||
- Google,
|
||||
- cloud inference,
|
||||
- telemetry,
|
||||
- analytics,
|
||||
- Sentry,
|
||||
- PostHog,
|
||||
- crash reporting,
|
||||
- CDN,
|
||||
- Google Fonts,
|
||||
- remote image hosts,
|
||||
- automatic update checks,
|
||||
- remote database support,
|
||||
- URL retrieval,
|
||||
- external web search,
|
||||
- MCP,
|
||||
- plugins,
|
||||
- arbitrary executable scripts,
|
||||
- third-party auth.
|
||||
|
||||
### Tasks
|
||||
|
||||
- static source search,
|
||||
- dependency review,
|
||||
- environment-variable review,
|
||||
- runtime network observation,
|
||||
- identify outbound requests required vs optional,
|
||||
- identify localhost vs wildcard binds,
|
||||
- identify stored secrets/API keys,
|
||||
- identify browser-side remote resources.
|
||||
|
||||
### Deliverables
|
||||
|
||||
- `reports/privacy-network-review.md`
|
||||
- per-candidate removal/mitigation list.
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
For each candidate, we can state exactly what local-only hardening would be required.
|
||||
|
||||
## 9. Milestone R4 — Feature and Reuse Matrix
|
||||
|
||||
### Objective
|
||||
|
||||
Compare candidates by subsystem rather than declaring one project the winner prematurely.
|
||||
|
||||
### Compare
|
||||
|
||||
- browser UX,
|
||||
- Ollama adapter,
|
||||
- streaming,
|
||||
- SQLite schema,
|
||||
- story tree,
|
||||
- rollback,
|
||||
- checkpoints,
|
||||
- edit/retry semantics,
|
||||
- state extraction,
|
||||
- entity/world state,
|
||||
- summaries,
|
||||
- semantic memory,
|
||||
- lexical memory,
|
||||
- lore/story cards,
|
||||
- source imports,
|
||||
- prompt inspection,
|
||||
- export/import,
|
||||
- tests,
|
||||
- local-only posture,
|
||||
- media extensibility.
|
||||
|
||||
### Rate Each Feature
|
||||
|
||||
Use categories such as:
|
||||
|
||||
- Keep as-is
|
||||
- Keep with modification
|
||||
- Reuse concept only
|
||||
- Replace
|
||||
- Not present
|
||||
- Not wanted
|
||||
|
||||
### Deliverables
|
||||
|
||||
- `reports/reuse-matrix.md`
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
We know which candidate has the best implementation of each required subsystem.
|
||||
|
||||
## 10. Milestone R5 — Licensing and Code-Reuse Review
|
||||
|
||||
### Objective
|
||||
|
||||
Determine what code can legally be copied, modified, linked, or used only as inspiration.
|
||||
|
||||
### Tasks
|
||||
|
||||
- verify repository licenses at the exact commits reviewed,
|
||||
- note third-party code with separate licenses,
|
||||
- note generated/vendor code,
|
||||
- compare compatibility if combining code from multiple projects,
|
||||
- pay special attention to GPL/copyleft candidates,
|
||||
- distinguish:
|
||||
- direct code reuse,
|
||||
- dependency use,
|
||||
- architecture inspiration,
|
||||
- protocol/API reimplementation.
|
||||
|
||||
### Deliverables
|
||||
|
||||
- `reports/licensing-reuse.md`
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
The recommended architecture does not rely on legally ambiguous code mixing.
|
||||
|
||||
## 11. Milestone R6 — Critical Prototypes
|
||||
|
||||
### Objective
|
||||
|
||||
Test only the uncertainties that could change the architecture decision.
|
||||
|
||||
Possible experiments include:
|
||||
|
||||
### Experiment A — Ollama adapter for ai-adventure
|
||||
Determine how difficult it is to replace/extend the LM Studio adapter with Ollama.
|
||||
|
||||
### Experiment B — Branch-safe state in Open Dungeon
|
||||
Determine whether Open Dungeon's current persistence can support immutable branch parentage without invasive rewrite.
|
||||
|
||||
### Experiment C — Strip-down feasibility in AI-DnD
|
||||
Identify whether RPG/cloud systems are modular enough to remove without destabilizing core story-tree/memory behavior.
|
||||
|
||||
### Experiment D — Local semantic retrieval
|
||||
Test a minimal local embedding pipeline using Ollama and a local-only store.
|
||||
|
||||
### Experiment E — Scene extraction
|
||||
Verify that the narrator/state pipeline can produce a neutral scene packet suitable for future media.
|
||||
|
||||
Only run experiments that resolve a documented decision.
|
||||
|
||||
### Deliverables
|
||||
|
||||
Each experiment:
|
||||
|
||||
```text
|
||||
experiments/<name>/README.md
|
||||
```
|
||||
|
||||
Record:
|
||||
|
||||
- question,
|
||||
- hypothesis,
|
||||
- minimal changes,
|
||||
- result,
|
||||
- implications,
|
||||
- whether code should be discarded.
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
No high-impact fork/architecture decision remains based solely on speculation.
|
||||
|
||||
## 12. Milestone R7 — Fork / Build Decision
|
||||
|
||||
### Objective
|
||||
|
||||
Select the production starting strategy.
|
||||
|
||||
### Required Options to Evaluate
|
||||
|
||||
- fork Open Dungeon,
|
||||
- fork AI-DnD,
|
||||
- fork/use ai-adventure core,
|
||||
- clean new shell with reused permissive components,
|
||||
- other candidate if discovered.
|
||||
|
||||
### Decision Criteria
|
||||
|
||||
Weight heavily:
|
||||
|
||||
1. fit with interactive-story product,
|
||||
2. browser-first architecture,
|
||||
3. Ollama fit,
|
||||
4. state/branch correctness,
|
||||
5. local-only hardening effort,
|
||||
6. amount of code to remove,
|
||||
7. maintainability,
|
||||
8. licensing,
|
||||
9. test quality,
|
||||
10. future media extensibility.
|
||||
|
||||
### Deliverables
|
||||
|
||||
- `reports/fork-build-recommendation.md`
|
||||
- ADR documenting the selected strategy.
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
One strategy is approved as the production base.
|
||||
|
||||
## 13. Milestone R8 — Finalize Specification and Technical Design
|
||||
|
||||
### Objective
|
||||
|
||||
Convert assumptions into committed decisions.
|
||||
|
||||
### Tasks
|
||||
|
||||
Update:
|
||||
|
||||
- `SPECIFICATION.md`
|
||||
- `TECHNICAL-DESIGN.md`
|
||||
|
||||
Resolve:
|
||||
|
||||
- base repository,
|
||||
- frontend framework,
|
||||
- backend framework,
|
||||
- storage model,
|
||||
- branch/state model,
|
||||
- model adapter,
|
||||
- memory/retrieval strategy,
|
||||
- local knowledge design,
|
||||
- import formats for v1,
|
||||
- context budgeting strategy,
|
||||
- security boundaries,
|
||||
- export format,
|
||||
- future media interfaces,
|
||||
- test strategy.
|
||||
|
||||
Mark documents v1.0 when approved.
|
||||
|
||||
### Deliverables
|
||||
|
||||
- `SPECIFICATION.md` v1.0
|
||||
- `TECHNICAL-DESIGN.md` v1.0
|
||||
- relevant ADRs.
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
A developer can explain the final architecture without unresolved foundational choices.
|
||||
|
||||
## 14. Milestone R9 — Create Production Build Plan
|
||||
|
||||
### Objective
|
||||
|
||||
Write the detailed implementation milestone plan only after the technical design is stable.
|
||||
|
||||
### Tasks
|
||||
|
||||
Create:
|
||||
|
||||
- `BUILD-MILESTONES.md`
|
||||
|
||||
It must include:
|
||||
|
||||
- milestone dependencies,
|
||||
- exact intended outcomes,
|
||||
- acceptance criteria,
|
||||
- test expectations,
|
||||
- migration steps from selected upstream,
|
||||
- removal/hardening work,
|
||||
- v1 feature sequence,
|
||||
- definition of done.
|
||||
|
||||
### Exit Criteria
|
||||
|
||||
The build plan is specific enough to hand directly to Codex milestone-by-milestone.
|
||||
|
||||
## 15. Phase 0 Final Deliverables
|
||||
|
||||
At Phase 0 completion:
|
||||
|
||||
```text
|
||||
SPECIFICATION.md v1.0
|
||||
TECHNICAL-DESIGN.md v1.0
|
||||
RESEARCH-PLAN.md completed
|
||||
BUILD-MILESTONES.md production-ready
|
||||
DECISIONS/ finalized foundational ADRs
|
||||
reports/ research evidence
|
||||
experiments/ critical prototype evidence
|
||||
inventory/ candidate metadata
|
||||
```
|
||||
|
||||
## 16. Phase 0 Definition of Done
|
||||
|
||||
Phase 0 is complete only when:
|
||||
|
||||
- candidate repositories have been cloned and reviewed,
|
||||
- serious candidates have been run or ruled out with evidence,
|
||||
- network/privacy behavior is documented,
|
||||
- licensing is understood,
|
||||
- critical architectural uncertainties have been tested,
|
||||
- a fork/build strategy has been selected,
|
||||
- the specification is v1.0,
|
||||
- the technical design is v1.0,
|
||||
- the actual implementation milestone plan has been written.
|
||||
|
||||
At that point, production implementation becomes eligible to begin, but still requires explicit approval and a milestone-specific execution prompt.
|
||||
|
||||
|
||||
## 17. Final Phase 0 Closure Record
|
||||
|
||||
Phase 0 definition of done is satisfied for architecture/planning purposes:
|
||||
|
||||
- finalists cloned and run,
|
||||
- test suites measured,
|
||||
- Ollama exercised locally,
|
||||
- offline/network behavior investigated,
|
||||
- critical history/state uncertainties prototyped,
|
||||
- fork/build strategy selected,
|
||||
- specification revised to v1.0,
|
||||
- technical design revised to v1.0,
|
||||
- production build milestones written,
|
||||
- foundational ADRs updated/added.
|
||||
|
||||
No production implementation prompt is part of this research plan.
|
||||
@@ -0,0 +1,64 @@
|
||||
# Reuse Matrix
|
||||
|
||||
**Historical status:** Phase 0A matrix. Some ratings were corrected by Phase 0B; use the v1 technical design and ADRs for current decisions.
|
||||
|
||||
**Date:** 2026-09-01
|
||||
|
||||
Legend:
|
||||
- **KEEP** — candidate implementation is close to target.
|
||||
- **MODIFY** — strong implementation but needs adaptation.
|
||||
- **REFERENCE** — borrow pattern/idea; do not make it the ownership center.
|
||||
- **BUILD** — target capability is substantially absent.
|
||||
|
||||
| Capability | AI-DnD | Open Dungeon | ai-adventure | Best current source |
|
||||
|---|---|---|---|---|
|
||||
| Browser storyteller UI | **MODIFY/KEEP** | **KEEP** | BUILD | Open Dungeon |
|
||||
| Ollama text adapter | **KEEP** | **KEEP** | MODIFY | AI-DnD/Open Dungeon |
|
||||
| SQLite local persistence | **KEEP** | MODIFY | **KEEP** | AI-DnD / ai-adventure |
|
||||
| Immutable turn parentage | **KEEP** | BUILD | **KEEP** | AI-DnD |
|
||||
| Alternate takes | **KEEP** | BUILD | MODIFY | AI-DnD |
|
||||
| Named checkpoints | MODIFY | BUILD | **KEEP** | ai-adventure |
|
||||
| Branch restore | **KEEP** | BUILD | **KEEP** | AI-DnD / ai-adventure |
|
||||
| Complete tree export | **KEEP** | BUILD | MODIFY | AI-DnD |
|
||||
| Exact prompt inspection | **KEEP** | BUILD/MODIFY | audit-oriented | AI-DnD |
|
||||
| Recent-history budgeting | **KEEP** | **KEEP** | **KEEP** | AI-DnD |
|
||||
| Rolling summaries | **KEEP** | **KEEP** | **KEEP** | all |
|
||||
| Semantic old-story retrieval | **KEEP/MODIFY** | BUILD | BUILD/MODIFY | AI-DnD |
|
||||
| Lexical local lore | MODIFY | BUILD | **KEEP** | ai-adventure |
|
||||
| Lore/story cards | **KEEP/MODIFY** | BUILD | MODIFY | AI-DnD |
|
||||
| Canon/Reference/Inspiration authority tiers | BUILD | BUILD | MODIFY | Chronicler/IFF concepts |
|
||||
| Generic narrative state | MODIFY | BUILD | MODIFY | ai-adventure pattern |
|
||||
| Model-proposes/app-validates | **KEEP but RPG-shaped** | BUILD | **KEEP** | ai-adventure |
|
||||
| Atomic state + turn commit | VERIFY | VERIFY | **KEEP** | ai-adventure |
|
||||
| Local-only privacy posture | MODIFY | MODIFY | **KEEP** | ai-adventure |
|
||||
| Local image generation | BUILD | **KEEP** | BUILD | Open Dungeon |
|
||||
| Provider-neutral media | BUILD | MODIFY | BUILD | Gamentic reference |
|
||||
| Scene/visual continuity | BUILD | **KEEP/MODIFY** | BUILD | Open Dungeon |
|
||||
| Large automated test base | **KEEP** | BUILD/VERIFY | **KEEP/VERIFY** | AI-DnD |
|
||||
| Genre-neutral core | MODIFY | **MODIFY/KEEP** | MODIFY | IFF/story-state concepts |
|
||||
|
||||
## Cross-project architecture we should aim for
|
||||
|
||||
Use one primary fork, not a stitched codebase.
|
||||
|
||||
Preferred composition of ideas:
|
||||
|
||||
```text
|
||||
AI-DnD production base
|
||||
+ ai-adventure trust/commit/privacy rules
|
||||
+ Open Dungeon scene/media UX
|
||||
+ Chronicler memory-authority tiers
|
||||
+ IFF Story Bible authority/validation
|
||||
+ Gamentic media-provider abstraction
|
||||
```
|
||||
|
||||
If AI-DnD strip-down proves too invasive, invert the first line:
|
||||
|
||||
```text
|
||||
Open Dungeon production base
|
||||
+ new AI-DnD-style immutable story tree
|
||||
+ ai-adventure event/commit discipline
|
||||
+ local memory/document retrieval
|
||||
```
|
||||
|
||||
That fallback is viable, but static analysis suggests it recreates more hard correctness work.
|
||||
Reference in New Issue
Block a user