Files
interactive-story/planning/reports/PHASE-0B-OPEN-DUNGEON-HISTORY.md
T
JesseMarkowitzandClaude Opus 5 ba737de9b4 Add Phase 0B local validation findings and recommendation
Validates the three finalists by clone, build, test run and live local
Ollama inference, then answers the fork question with measurements rather
than static review.

Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but
it was wrong that AI-DnD's undo is non-destructive — retry preserves the
replaced take, undo hard-deletes it. A follow-up spike fixed that in 3
files (+130/-31): undo now moves a head cursor, redo round-trips, writing
below a moved-back head forks and keeps the abandoned line, branch-scoped
memory isolation survives, suite 627/632 with all 5 failures asserting the
deleted-row behaviour that was replaced.

Findings that change the plan:
- AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its
  encoding from a CDN. Proven on an internal Docker network, proven fixed
  by vendoring the file.
- ai-adventure needs zero code for Ollama — two config lines — and its
  turn/head/checkpoint schema is the target model to build to.
- Open Dungeon has zero automated tests and a positional summary
  watermark, making its branch retrofit larger than Phase 0A costed.
- The world-state referee takes relative deltas; a 3B model sent absolute
  values under full context, so a wounded player ended at full health.
  Validation cannot catch this, so prefer ai-adventure's typed-event
  vocabulary when generalising narrative state.
- Export/import recomputes head depth, so a round-trip silently undoes an
  undo. Must be fixed alongside the undo work.

Docs only; no production code. Working tree from the runs stays untracked
under phase0b/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu
2026-09-01 16:11:38 -04:00

4.8 KiB

Phase 0B — Experiment B: Open Dungeon history model

Driven live on 127.0.0.1:3111 against Ollama through the "custom" OpenAI-compatible provider (http://127.0.0.1:11434/v1, qwen2.5:0.5b), with its own SQLite file.

B1. Schema — no lineage exists

src/lib/db.ts creates four tables: chats, messages, characters, app_settings. The message row is:

CREATE TABLE messages (
  id TEXT PRIMARY KEY,
  chat_id TEXT NOT NULL REFERENCES chats(id) ON DELETE CASCADE,
  role TEXT CHECK (role IN ('user','assistant')),
  content TEXT NOT NULL,
  attachments_json ..., image_request_json ..., generated_image_json ...,
  created_at TEXT NOT NULL
);
CREATE INDEX idx_messages_chat_created ON messages(chat_id, created_at);

No parent_id, no branch, no take, no depth, no checkpoint table. Order is created_at (a TEXT timestamp), with ties broken by id.

B2. Retry / Edit / Erase — all destructive, measured

BEFORE (5 rows)
  a504cb34 assistant "Oh, Vale, you're just sitting in that rundown ..."
  e349d8b9 user      "I pocket a brass key from the ashtray."
  9984692b assistant "You find a worn, faded brass key on the tabletop..."
  bb1ceb3f user      "I drive to the address on its tag."
  a6e11115 assistant "Yes, you keep going. There, you can call your ..."

EDIT e349d8b9 in place  → PATCH /api/chats/{id}/messages/{id}
  original text "brass key from the ashtray" still anywhere in DB? False

DELETE a6e11115?after=1 (what BOTH Retry and Erase call)
  rows 5 → 4

Scan of every column of every table for the deleted tail
or the pre-edit text: none
  • Edit is UPDATE messages SET content = ? WHERE id = ?. The prior text is unrecoverable.
  • Retry (retryLastTurn, page.tsx) deletes the last assistant message and everything after it, then regenerates. The discarded attempt is gone.
  • Erase (eraseLastTurn) deletes the last exchange and the tail.
  • Old future turns are therefore always deleted, never retained.

B3. What depends on the linear model

  1. db.ts — 20 exported functions. addMessage, updateMessageContent, deleteMessageAndAfter and getChat are the history four. getChat returns a flat messages array and must become a lineage query.
  2. story-prompt.ts — windows by array index: for (let i = messages.length - 1; ...), messages.slice(dropped), messages.slice(0, dropped). Becomes an ancestry walk.
  3. The summary watermark is positional — the cost Phase 0A missed. chats.story_summary_count records "how many of the chat's oldest messages the summary already covers", and api/story/route.ts compares evicted.length > stored.coveredCount, then summarizes evicted.slice(stored.coveredCount). "The first N messages" is meaningless on a branch. Summaries must be re-anchored to a turn id and made per-lineage.
  4. page.tsx is 3,991 lines in one component; api/story/route.ts is 1,196. Retry/erase/edit live there as array slicing (messages.slice(0, cutFrom)). There is no take stepper, branch panel or tree view to build on.
  5. No tests. All of the above would be changed with no regression net.

B4. Invasiveness estimate

To reach parent-linked turns, an active head, retained abandoned history, named checkpoints and lineage-safe summaries, Open Dungeon needs: a schema migration, a rewrite of its persistence read/write path, a rewrite of context assembly, a redesign of summarization, and new branch UI — with a test suite written first to make any of it safe. This is a from-scratch implementation of the hardest part of the specification, not a retrofit.

By comparison AI-DnD already has all of it except a non-destructive undo, and its reads funnel through one chokepoint (lineage.Path.clause).

B5. Worth reusing as reference

  • The reading experience: serif prose, prose-size control, the composer's Do/Say/Story modes.
  • Inline scene images driven by a narrator tool call (generate_image), with a local FLUX worker on 127.0.0.1:7869 or a user's ComfyUI on 8188.
  • Character portraits reused as reference images for visual continuity, and fed back to the narrator as vision context. This is the most valuable idea here for MEDIA-EXTENSION-CONTRACT.md.
  • The worker HTTP protocol may be portable even though none of the UI is — Open Dungeon is Next.js, AI-DnD is React + FastAPI.

B6. Other findings

  • The local provider hard-codes five Gemma 4 QAT builds (src/lib/text-models.ts, LOCAL_TEXT_MODELS) and /api/health reports only those as installed. Any other local Ollama model must be reached through the "custom" provider. A genre-agnostic app wanting "pick any installed model" must lift this.
  • Next.js telemetry is enabled by default and printed its notice on first start.
  • 335 npm packages; npm audit reports 6 high, 1 low.