Files
interactive-story/planning/PHASE-0B-RECOMMENDATION.md
T

28 KiB

Planning review note (2026-09-01): This is the coding agent's Phase 0B evidence report. It is retained substantially as delivered. Some recommendations/open questions in the report were superseded during planning review: ai-adventure is an implementation reference rather than the project specification; no automatic abandoned-history cleanup is required in v1; Open Dungeon media portability is deferred; and the architecture decision is closed in ADR 009/010 and TECHNICAL-DESIGN.md.

Phase 0B — Local Validation Findings and Recommendation

Date: 2026-09-01 Status: Complete, and updated after the follow-up undo/redo spike. Stop condition reached; no production work started. Method: Clean clones, pinned SHAs, documented installs, full test runs, real local Ollama inference, and targeted disposable experiments.

Update — the recommended spike was run and it passed. Section E previously recommended proving non-destructive Undo/Redo in AI-DnD before committing to the fork. That spike is now complete: 3 files, +130/-31 lines, undo deletes zero rows, redo round-trips exactly, writing below a moved-back head forks and preserves the abandoned line, branch-scoped memory isolation survives, and the suite goes 627/632 with all 5 failures asserting the deleted-row behavior that was deliberately replaced. Section E has been rewritten accordingly, and the open questions it resolved are marked in section F. Full detail: reports/PHASE-0B-UNDO-SPIKE.md.

Update 2 — four follow-up checks run. The world-state referee was exercised against local models for the first time, and export/import, story-card lineage safety and Postgres removability were settled. One new defect found (a round-trip silently undoes an undo) and one design conclusion that changes what we should build (§H). Detail: reports/PHASE-0B-FOLLOWUP-CHECKS.md.

Pinned commits (cloned 2026-09-01):

Repo SHA Last commit
AI-DnD d72f7c1bda0f34fccd84afb7a25c34eb01c901de 2026-08-31
Open Dungeon b0a79f96bf852be7b4e53908dff6a7f7c179da23 2026-07-06
ai-adventure 873ea9180d5b611576cddb155921fc16a17ae88b 2026-07-21

Inference used for every live test: Ollama in Docker, qwen2.5:0.5b and qwen2.5:3b-instruct for narration, nomic-embed-text for embeddings. CPU only, no GPU.


A. Executive Recommendation

Fork AI-DnD as the production base. Confidence: high.

The Phase 0A recommendation held up, but for partly different reasons than it gave, and it was wrong about one important thing.

What held: AI-DnD really does implement the expensive correctness work, and it survived every entanglement test Phase 0A worried about. The RPG machinery is optional, the scripting sandbox is removable, and the branch-scoped memory isolation — the single hardest requirement in the specification — works correctly against real local embeddings.

What did not hold: Phase 0A's claim that AI-DnD's undo is non-destructive is false. POST /undo hard-deletes rows. Retry is non-destructive; undo is not. There is no Redo. This is a direct conflict with DECISIONS/005-branch-preserving-history.md and specification §4.

That discovery does not change the decision, because the comparison is relative: AI-DnD needs one bounded change at a chokepoint that every read already passes through. Open Dungeon needs the entire non-destructive history subsystem built from nothing, under a 3,991-line component, with zero automated tests.

That "one bounded change" is no longer an estimate. The follow-up spike implemented it in 3 files and +130/-31 lines, and it works: undo deletes nothing, redo restores exactly, the abandoned tail becomes an ordinary branch, and memory isolation holds. Confidence in the fork decision moves from "high on the evidence" to "high and demonstrated".

The decision gate from reports/PRELIMINARY-RECOMMENDATION.md said to select AI-DnD unless a blocker is confirmed. All four were tested and none is a blocker:

Phase 0A blocker Result
Branching/state inseparable from RPG mechanics No. A scenario-less adventure with empty world state played, branched, and retrieved memories correctly.
Removing hosted/scripting destabilizes many tests No. 19 of 632 (3%), and every one fails for instrumentation reasons, not coupling.
Local-only still needs external services Partly — but both are packaging fixes. tiktoken downloads its encoding from a Microsoft CDN; the SPA fetches Google Fonts.
Dependency/security burden worse than expected No. 80 packages total vs Open Dungeon's 335 with 6 high-severity advisories.

B. What We Learned About Each Candidate

What worked

  • 632 passed in 236s, first try, no fixes needed.
  • Docker image builds clean; docker compose path is real.
  • Default model endpoint is already http://localhost:11434/v1. Connection test and model discovery against Ollama worked with no code change.
  • Streamed a full multi-turn story over SSE against local Ollama.
  • Genre-neutrality is real. An adventure created with scenario_id: null gets an empty world state and no RPG scaffolding, and plays normally. Noir prose, no fantasy code paths, no stats.
  • Retry is genuinely non-destructive. Retrying kept the replaced attempt as a sibling take at the same coordinate (rows 6 and 7 both present; both listed by /variants).
  • Forking is genuinely non-destructive. Adding an alternate take at a passed turn created branch 2 while every row of branch 1 survived, with parent_branch_id and fork_depth recorded.
  • Branch-scoped memory isolation is correct. Two memories were written on branch 1 (depths 5 and 11) from real nomic-embed-text embeddings. On branch 2, retrieval returned used: [] and none of brass key, revolver, lockbox, ashtray or bullets appeared anywhere in the assembled prompt. The negative control passed: switching back to branch 1 returned both memories with similarity scores (0.8087) and the abandoned facts reappeared. Retrieval is working and correctly scoped.
  • Prompt transparency is already there. The context endpoint returns labeled sections with token costs: narrator 57, ai_instructions 9, persona 15, plot_essentials 24, history 988, used_memories 143, length_hint 46.

What failed

  • Undo hard-deletes. Measured: rows 7 → 4, deleting ids 5, 6 and 7 — that is, both the player action, the live AI take, and the alternate take the earlier retry had preserved. delete_turn also calls memorybank.forget_node. A retry's preserved alternative does not survive a later undo. There is no Redo endpoint. Resolved in the follow-up spike — see section G.
  • Fails offline on a clean install. First turn in a network-isolated container died with NameResolutionError for openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken. tiktoken fetches its BPE encoding on first use. Seeding the 1.7 MB file into the image made a full turn generate with no internet at all. Packaging fix, not architecture — but a hard blocker until fixed, and static review missed it.
  • Runtime remote asset. The built SPA still requests Google Fonts; the CSP explicitly allows fonts.googleapis.com and fonts.gstatic.com.

Strongest reusable pieces

Immutable parent-linked story tree with lineage entries; alternate takes; branch fork/switch/rename/delete; per-node world-state and script-state snapshots; branch-scoped memories with local embeddings; the Insights prompt snapshot; ai-dnd-adventure-v2 whole-tree export; the 632-test suite; SSE streaming.

Major architectural problems

Destructive undo and no redo. No named checkpoints (branch names are the nearest thing). Multi-user/hosted concerns are spread across ten modules. netguard.py is the inverse of what we need — it blocks non-public endpoints in hosted mode and does nothing locally; we want a warning on non-loopback endpoints.

Adaptation required

Bounded and mostly subtractive. The one genuinely additive change is non-destructive undo/redo, now measured at 3 files and +130/-31 lines (section G).

What worked

Installed and ran on first try. Plays a real story against Ollama through its "custom" OpenAI-compatible provider. Clean browser UX. No runtime Google Fonts (next/font self-hosts at build time). Local SQLite, local image generation, character portrait continuity.

What failed

  • Every history operation is destructive, confirmed by measurement.
    • Edit is an in-place UPDATE messages SET content = ?. After editing, the original text brass key from the ashtray was gone from every column of every table.
    • Retry and Erase both call DELETE /messages/{id}?after=1, which runs DELETE FROM messages WHERE chat_id = ? AND (created_at > ? OR ...). Measured rows 5 → 4, with no archive anywhere.
    • The database has exactly four tables — chats, messages, characters, app_settings. No parent pointer, no branch, no take, no checkpoint.
  • Zero automated tests. No test script, and no test framework in devDependencies at all. Phase 0A listed this as "BUILD/VERIFY"; it is zero.
  • The local provider hard-codes a five-model Gemma 4 allowlist (src/lib/text-models.ts). Arbitrary local Ollama models are only reachable through the "custom" OpenAI-compatible path.
  • Next.js telemetry is on by default and printed its notice on first start.
  • 335 npm packages, 6 high-severity advisories.

Cost of the branch retrofit — the Experiment B answer

Larger than Phase 0A estimated, because of a dependency Phase 0A did not identify:

  1. messages needs parent_id, branch_id, depth, and a live/take flag.
  2. Of 20 exported db.ts functions, the four history ones need rewriting, and getChat must stop returning a flat list and become a lineage query.
  3. story-prompt.ts windows history by array index (messages.slice(dropped), messages.length - 1). All of it becomes an ancestry walk.
  4. The rolling summary watermark is positional. chats.story_summary_count means "how many of the chat's oldest messages the summary already covers", and route.ts compares evicted.length > stored.coveredCount. On a branch, "the first N messages" has no meaning. Summaries must be re-anchored to a turn id and made per-lineage. Phase 0A did not identify this.
  5. page.tsx is a single 3,991-line component holding retry/erase/edit as array slicing, and there is no take stepper, branch panel, or tree view to extend.
  6. All of the above lands with no test suite underneath it.

Worth reusing as reference: the reading UI, inline scene images, the character-portrait-as-reference-image continuity idea, and the scene/media architecture. None of it ports directly — Open Dungeon is Next.js and AI-DnD is React + FastAPI.

ai-adventure — not the base, but promote it to the architectural reference

What worked — this is the strongest result of the round

  • Ran 76 tests ... OK in 1.9s.
  • The Ollama "adapter" is zero code. LMStudioBackend is a generic OpenAI-compatible client hitting /v1/chat/completions and /v1/models. I changed two lines of world.toml — base_url and model name — and the doctor reported [PASS] LM Studio reachable: http://127.0.0.1:11434 and [PASS] Configured model visible: qwen2.5:3b-instruct. A real turn then committed to the database with status='committed', zero validation errors, and a model_calls row recording backend, model and attempt. Phase 0A's "MODIFY — LM Studio primary provider" was wrong; it is a config change.
  • Its history model is exactly what the specification asks for. Verified independently, not just by trusting its tests:
    • Undo deleted 0 rows, rolled has_brass_key back from False to True, and the abandoned turn was still in the database.
    • Branching from the restored point produced gate_open on the branch and did not leak it into the original session.
    • Checkpoint restore rolled state back with the later turn still present.
    • Total turns in the database never decreased.
  • Schema is the target data model: turns.parent_turn_id, status, append-only state_events, state_cache keyed by head_turn_id, model_calls with request/response hashes, named_checkpoints, summaries anchored by through_turn_id, and lore in FTS5.
  • The cleanest privacy posture by a wide margin. Exactly one outbound call site in the entire codebase (lm_studio.py's urlopen), no hardcoded remote host anywhere, and one third-party dependency (pydantic). It is the only candidate that already warns when model.base_url is not loopback — which is specification §12, implemented.

What failed

  • With qwen2.5:0.5b the turn failed validation and nothing was saved. That is the model-proposes/app-validates discipline working correctly, but it means the typed-event contract needs a reasonably capable model; the freeform-prose candidates tolerate a weak one. qwen2.5:3b-instruct succeeded.
  • Product distance is real and large: terminal only, worlds authored as hand-written TOML and Markdown files rather than set up in-app, no streaming, no embeddings/semantic memory, no media, no import of user documents.

Adaptation required: building essentially the entire product on top of a correct core.


C. Key Technical Findings

Requirement AI-DnD Open Dungeon ai-adventure
History/undo model Tree + takes; undo deletes, no redo (both fixed in the spike, §G) Flat list; retry/edit/erase all destroy Parent-linked, head pointer, undo deletes nothing
State rollback Per-node snapshots, verified None Event replay, verified
Memory isolation across branches Verified correct, with negative control N/A (no branches) Verified correct
Local Ollama Default endpoint; verified streaming Verified via "custom" provider; local provider allowlists 5 Gemma builds Verified, zero code change
Offline Blocked by tiktoken CDN fetch until vendored; then fully offline Runtime clean; build needs network; Next telemetry on by default Verifiably clean — 1 outbound call site, 1 dependency
Browser suitability React SPA + FastAPI, already there Next.js, already there None
Prompt inspection Per-action snapshot with labeled sections and token costs Not exposed Prompt hashes; store_prompts off by default
Imported knowledge Story cards (keyword-triggered) — closest starting point None Local FTS5 lore, deterministic
Named checkpoints Absent (branch names only) Absent Present
Future media Absent Present and working Absent
Tests 632 0 76

D. Important Surprises

Ordered by how much they should change the plan.

  1. AI-DnD's undo is destructive, and it destroys retry's preserved alternatives too. reports/PRELIMINARY-RECOMMENDATION.md and REUSE-MATRIX.md both credit AI-DnD with non-destructive rollback. Retry earns that; undo does not. Measured: undo deleted the player action, the live take, and the sibling take a prior retry had kept. The follow-up spike has since replaced it with a head-cursor undo and added Redo (§G), so this is a corrected planning assumption rather than an outstanding defect.
  2. AI-DnD cannot take a single turn on an air-gapped machine as shipped. tiktoken fetches cl100k_base from a Microsoft CDN on first use. Proven by running it on a Docker --internal network, and proven fixed by seeding the file.
  3. ai-adventure needs no Ollama adapter at all. Two config lines. The "LM Studio" name is a misnomer for a generic OpenAI-compatible client.
  4. Open Dungeon has zero automated tests — not few, none, and no framework installed.
  5. Open Dungeon's summary watermark is positional, which adds a summary redesign to its branch retrofit that Phase 0A did not cost.
  6. Open Dungeon's "local" provider only offers five Gemma 4 QAT builds. A genre-agnostic app that lets the user pick any installed Ollama model needs that lifted.
  7. AI-DnD fetches Google Fonts at runtime and its CSP is written to allow it, violating specification §12's "no remote fonts".
  8. AI-DnD's netguard is the opposite of our requirement — it protects a hosted deployment from SSRF and is inert locally. ai-adventure has the loopback warning we actually want.
  9. Repo maturity differs sharply: AI-DnD is 172 commits / 3 authors, Open Dungeon 45 / 1, ai-adventure 20 / 1. All three are young. Since we fork, this matters less than it would for a dependency, but there is no upstream to lean on.

E. Recommendation for Next Step

Fork AI-DnD and begin the controlled strip-down. The blocking uncertainty is gone: the spike this section previously asked for has been run and passed (section G).

Recommended order, cheapest and most load-bearing first:

  1. Close the two offline blockers. Vendor the tiktoken encoding and self-host the three font families, then re-run the air-gapped test. Until this is done the application does not satisfy specification §2, and it is a day's work.
  2. Strip the hosted surface. Remove multi-user/auth, the demo key, analytics, render.yaml, the Postgres path, and the quickjs sandbox. Re-instrument the ~19 tests that use JS hooks to observe state, moving them onto the world_state_after snapshots that already exist.
  3. Land the undo/redo work properly, promoting the spike from a proof to a feature: fix the four rough edges it left (undo floor at the opening node, retry/add_take while behind the head, marking abandoned turns disposable, and the frontend Redo control and branch labelling) — and add headDepth to the export bundle in the same change, or a round-trip silently undoes every undo (§H).
  4. Add named checkpoints, using ai-adventure's named_checkpoints(session_id, turn_id, name) as the model. With an explicit head that can sit anywhere, a checkpoint is now just a named head position.
  5. Add the loopback-endpoint warning, modeled on ai-adventure's _endpoint_warnings.
  6. Then imported knowledge, and only after that, media.

Treat ai-adventure as the written specification for state and history semantics throughout. Its schema is the target shape and its 76 tests are the behavioral contract worth copying — the spike already reproduced its central property (undo that moves a head instead of deleting) and benefited from having that reference to aim at. §H extends this: when generalizing AI-DnD's world state into genre-neutral narrative state, take ai-adventure's typed-event vocabulary rather than AI-DnD's relative-delta vocabulary, because the delta protocol has a failure mode that validation cannot catch and that worsens as context grows.

One addition to the ordering above: whatever narrative-state engine we build must be exercised at realistic context length against the models users will actually run, not only in a clean harness. That is where the referee's weakness showed up, and it would not have appeared in a unit test.

Do not merge repositories. Open Dungeon remains a UX and media reference only.

F. Open Questions

Resolved by the undo/redo spike (§G):

  • Does capping Path.clause by head depth break the memory and summary cursors? No. test_memory_settling, test_memory_nodes, test_memory_retrieval, test_memory_rewrite and test_history_window all passed unchanged, and memory retrieval was verified end-to-end against real embeddings with a control.
  • How should named checkpoints sit on top of takes and branches? Largely answered. The spike introduces a head that can sit behind the tip, so a checkpoint becomes a named (branch_id, depth) position. Still needs a decision on whether restoring one forks immediately or on first write — the spike chose "on first write" for undo, and consistency argues for the same.

Resolved by the follow-up checks (§H):

  • What is the model capability floor? Characterized, not a blocker. The delta protocol is within a 3B model's reach in a clean prompt (3/3 correct) but degraded to absolute values under the full application context. 0.5b cannot do it at all. See §H for the design consequence.
  • Can psycopg/Postgres be dropped cleanly? Yes. Three places, no blockers: one branch in database.py, two migration helpers, one import inside analytics.
  • Do story cards become imported knowledge, or is a new table needed? A new table. StoryCard carries no branch_id/depth, so it is not lineage-scoped and inherits neither branch nor undo isolation.
  • Does the export format survive the history rework? No — and this is now a known defect to fix, not a question. See §H.

Still open:

  1. When does abandoned history get cleaned up? Nothing marks it disposable yet, and undone turns now accumulate rather than being deleted. This is the intended trade, but it needs a retention story.
  2. Is Open Dungeon's image path portable at all? It is Next.js calling a local FLUX worker over HTTP. The worker protocol may be reusable even though none of the UI is.
  3. What is the real model recommendation for users? The probe establishes a floor and a failure mode, not a ranking. A proper capability matrix across the models a user would actually run belongs in the build phase, measured at realistic context length.
  4. Not tested in any round: import/export round-trip behavior beyond the head-depth defect found by reading, multi-hour long-story context behavior, or any concurrency beyond the single-turn lock.

G. Follow-Up Spike: Non-Destructive Undo/Redo — Passed

Run after the main round, on a disposable copy of AI-DnD d72f7c1b. Full detail in reports/PHASE-0B-UNDO-SPIKE.md.

The change is three files, +130 / -31 lines, roughly a fifth of it comments:

File Change What
context/lineage.py +26 / -3 Read an uncapped lineage entry as capped at self.tip (the head), in clause and contains; add uncapped().
routers/adventures/takes.py +74 / -27 undo_turn moves head_depth instead of deleting; new redo_turn.
routers/adventures/turns.py +30 / -1 fork_if_behind_head calls the existing tree.branch_at when a write lands below the head.

It is this small because the pieces already existed: every read funnels through lineage.path_of(), adventures.head_depth was already a stored head, and tree.branch_at already created a branch that leaves a path at a depth without touching the line it leaves.

Results against the success criteria:

Criterion Result
Undo deletes zero rows Pass. 6 → 6 rows across three undos; separately, 11 consecutive undos on a 23-action story with 23 actions still present. Clears the "at least five, unlimited preferred" requirement.
Abandoned tail stays reachable as a branch Pass. Undo-then-write forked branch 2 at fork depth 3; branch 1 kept all six of its rows, live and readable, and the branch switcher reaches it.
Redo restores Pass. Three redos returned the head from -1 to 5 and the transcript to all six actions, exactly.
Branch-scoped memory isolation holds Pass, with control. An embedded memory at depth 5 was retrieved at the tip (similarity 0.689), returned used: [] once the head moved to depth 3, and became eligible again on redo — without being deleted. None of seven terms unique to the hidden turns appeared in the assembled prompt.
Suite stays green apart from delete-behavior tests Pass. 627 passed, 5 failed (0.8%). All five assert deleted rows or pruned memories. The state assertions inside those same tests still pass — assert adv.script_state == {"gold": 0} succeeds in both test_state_revert failures, and the fork-boundary guard in test_undo_stops_at_the_fork is untouched.

Rough edges left for the real implementation (none architectural): undo can walk past a user-written opening to an empty transcript; retry and add_take were not routed through the fork check; abandoned turns are not yet marked disposable; and the frontend has no Redo control.

What this changes about the recommendation: nothing about the choice, and the one thing about the plan — the strip-down can start now rather than after another investigation.

H. Follow-Up Checks — Referee, Export, Cards, Postgres

Run after the spike. Detail in reports/PHASE-0B-FOLLOWUP-CHECKS.md.

The finding that changes what we build

AI-DnD's world-state referee asks the model for relative deltas (new = old + delta, then clamped). Playing the seeded RPG scenario with qwen2.5:3b-instruct, the model sent absolute values — {"player.hp": 92} meaning "hp is now 92". The engine read +92, clamped to +35, then to the ceiling of 100:

player.hp  100 -> 100  "did not move. It is already at its maximum of 100"

The player took an arrow to the shoulder and finished the turn at full health.

The referee did its job — nothing was corrupted, and it fed corrective text back. But no validator can catch this, because +92 is a legal proposal. The application stayed authoritative while the authoritative state silently stopped tracking the narration. That is the exact divergence specification §2 exists to prevent.

An isolated probe using AI-DnD's own EMIT_RULE and parser shows the protocol is not beyond a small model — qwen2.5:3b-instruct and qwen2.5:7b-instruct both got 3/3 correct relative deltas, and qwen2.5:0.5b emitted no block at all. The difference is context load: in a short prompt the 3B model follows the rule, and under the full application prompt it drifts to absolutes.

Consequence for the design: when we generalize AI-DnD's world state into genre-neutral narrative state, take ai-adventure's typed-event vocabulary (set_flag key=… value=… — explicit and absolute) rather than AI-DnD's delta vocabulary. The delta protocol has a failure mode that is invisible to validation and worsens with context length; the event protocol cannot express the mistake. This is the second time in this round that ai-adventure's core has turned out to be the right specification to build to.

One new defect

Export/import silently redoes an undone story. bundle.py:_point_the_head recomputes head_depth = max(depths) on import, and the bundle format carries headBranch but not the head depth. That was correct when undo deleted rows; with the spike's non-destructive undo, a round-trip restores every abandoned turn. Fix is one optional headDepth field plus honouring it on import, falling back to max(depths) so existing v2 files still load. It must land with the undo work, not after it.

Two smaller answers

  • Story cards are not lineage-scoped. StoryCard has no branch_id or depth, unlike Memory. Harmless today because cards are authored rather than derived — but any imported-knowledge source that can be created from story content must carry those coordinates and be read through lineage.Path.clause, or it inherits neither branch nor undo isolation. Argues for a new table rather than extending story_cards.
  • Postgres is cleanly removable. One branch in database.py, two migration helpers, one import inside analytics. psycopg[binary] then drops out.

Supporting Evidence

  • reports/PHASE-0B-BASELINE.md — installs, test runs, ports, storage, versions.
  • reports/PHASE-0B-AI-DND-EXPERIMENT.md — branch/undo/retry measurements, memory isolation with negative control, scripting removal.
  • reports/PHASE-0B-OPEN-DUNGEON-HISTORY.md — destructive-history proof and retrofit cost map.
  • reports/PHASE-0B-AI-ADVENTURE-OLLAMA.md — zero-change Ollama run and the fixture checks.
  • reports/PHASE-0B-OFFLINE-NETWORK.md — offline runs and egress inventory.
  • reports/PHASE-0B-UNDO-SPIKE.md — the non-destructive undo/redo spike: the diff, the measurements, and the rough edges it left.
  • reports/PHASE-0B-FOLLOWUP-CHECKS.md — the world-state referee under local models, the export/import head defect, story-card lineage safety, Postgres.

Working tree, scripts and databases from both rounds are under phase0b/ (untracked scratch, safe to delete). The spike itself is phase0b/spike/, a disposable copy of the AI-DnD backend — it is a proof, not a branch to merge.