Files
interactive-story/planning/reports/PHASE-0B-BASELINE.md
T
JesseMarkowitzandClaude Opus 5 ba737de9b4 Add Phase 0B local validation findings and recommendation
Validates the three finalists by clone, build, test run and live local
Ollama inference, then answers the fork question with measurements rather
than static review.

Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but
it was wrong that AI-DnD's undo is non-destructive — retry preserves the
replaced take, undo hard-deletes it. A follow-up spike fixed that in 3
files (+130/-31): undo now moves a head cursor, redo round-trips, writing
below a moved-back head forks and keeps the abandoned line, branch-scoped
memory isolation survives, suite 627/632 with all 5 failures asserting the
deleted-row behaviour that was replaced.

Findings that change the plan:
- AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its
  encoding from a CDN. Proven on an internal Docker network, proven fixed
  by vendoring the file.
- ai-adventure needs zero code for Ollama — two config lines — and its
  turn/head/checkpoint schema is the target model to build to.
- Open Dungeon has zero automated tests and a positional summary
  watermark, making its branch retrofit larger than Phase 0A costed.
- The world-state referee takes relative deltas; a 3B model sent absolute
  values under full context, so a wounded player ended at full health.
  Validation cannot catch this, so prefer ai-adventure's typed-event
  vocabulary when generalising narrative state.
- Export/import recomputes head depth, so a round-trip silently undoes an
  undo. Must be fixed alongside the undo work.

Docs only; no production code. Working tree from the runs stays untracked
under phase0b/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu
2026-09-01 16:11:38 -04:00

40 lines
2.7 KiB
Markdown

# Phase 0B — Baseline Results
**Date:** 2026-09-01. Host: Linux, 4 cores, 15 GB RAM, no GPU, Python 3.12.3, Node 22.23.1.
Inference: Ollama in Docker (`ollama/ollama:latest`) on `127.0.0.1:11434`, models
`qwen2.5:0.5b`, `qwen2.5:3b-instruct`, `nomic-embed-text`.
| | AI-DnD | Open Dungeon | ai-adventure |
|---|---|---|---|
| SHA | `d72f7c1b` | `b0a79f96` | `873ea918` |
| License | MIT | MIT | Apache-2.0 |
| Repo age / commits / authors | created 2026-07-20, 172, 3 | 2026-06-12, 45, 1 | 2026-07-20, 20, 1 |
| Code size | ~30.7k LOC | ~9.1k LOC | ~4.5k LOC |
| Stack | FastAPI + SQLAlchemy + React/Vite | Next.js 16 + React 19 | Python stdlib + pydantic, CLI |
| Storage | SQLite (Postgres option via `AIDND_DATABASE_URL`) | SQLite (`better-sqlite3`) | SQLite + FTS5 |
| Install | `pip install -r requirements*.txt`, `npm install` — clean | `npm install` — clean, 7 advisories (6 high) | `pip install -r requirements.txt` — clean |
| Dependencies | 41 pip + 39 npm | 335 npm | **6 pip (1 direct: pydantic)** |
| Tests | **632 passed / 236s** | **none — no script, no framework** | **76 passed / 1.9s** |
| Ports | 8000 (prod/docker), 5173 + 8000 (dev) | 3000 (app), 7869 (FLUX worker), 8188 (ComfyUI) | none (CLI) |
| Model assumption | any OpenAI-compatible; **defaults to `http://localhost:11434/v1`** | Ollama native API, or OpenAI-compatible "custom" | OpenAI-compatible; base_url in `world.toml` |
| Docker build | `docker build` succeeds (3-stage, builds SPA) | not attempted (npm path used) | n/a |
| Runtime failures | none once running; **offline blocker, see offline report** | none | typed-event validation fails on a weak model (by design) |
| Cloud/hosted surface | `render.yaml`, multi-user auth, demo key, analytics, Postgres, OpenRouter default | OpenRouter preset, Next telemetry on by default, donation links | **none** |
## Notes
- **AI-DnD** ran its whole suite green on the first attempt with no fixes. Its
default settings already point at Ollama; `POST /api/settings/test` returned
`{"ok":true,"models":["qwen2.5:0.5b"]}` with no configuration beyond selecting
a model name.
- **Open Dungeon** has `lint` and several Windows/image smoke scripts, but no
unit or integration tests of any kind. Its `local` provider only offers five
hard-coded Gemma 4 QAT builds (`src/lib/text-models.ts`), so a non-Gemma local
model must go through the "custom" OpenAI-compatible provider; that is how it
was driven here.
- **ai-adventure** is the only candidate whose full dependency closure is one
third-party package. `python -m local_adventure doctor` is a genuine
environment self-check (Python version, writable runtime dir, SQLite version,
FTS5 availability, schema version, world validity, endpoint reachability,
model visibility).