Validates the three finalists by clone, build, test run and live local Ollama inference, then answers the fork question with measurements rather than static review. Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but it was wrong that AI-DnD's undo is non-destructive — retry preserves the replaced take, undo hard-deletes it. A follow-up spike fixed that in 3 files (+130/-31): undo now moves a head cursor, redo round-trips, writing below a moved-back head forks and keeps the abandoned line, branch-scoped memory isolation survives, suite 627/632 with all 5 failures asserting the deleted-row behaviour that was replaced. Findings that change the plan: - AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its encoding from a CDN. Proven on an internal Docker network, proven fixed by vendoring the file. - ai-adventure needs zero code for Ollama — two config lines — and its turn/head/checkpoint schema is the target model to build to. - Open Dungeon has zero automated tests and a positional summary watermark, making its branch retrofit larger than Phase 0A costed. - The world-state referee takes relative deltas; a 3B model sent absolute values under full context, so a wounded player ended at full health. Validation cannot catch this, so prefer ai-adventure's typed-event vocabulary when generalising narrative state. - Export/import recomputes head depth, so a round-trip silently undoes an undo. Must be fixed alongside the undo work. Docs only; no production code. Working tree from the runs stays untracked under phase0b/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu
4.7 KiB
Phase 0B — Experiment C: ai-adventure
C1. Tests
python -m unittest discover -s tests → Ran 76 tests in 1.9s OK
C2. Provider abstraction — the Ollama adapter is zero code
llm/backend.py defines a ModelBackend Protocol with a single generate
method plus typed ModelRequest/ModelResponse. llm/lm_studio.py is not
LM-Studio-specific at all: it POSTs /v1/chat/completions and reads /v1/models
— the same OpenAI-compatible surface Ollama serves. llm/scripted.py is a
deterministic fake used by the tests.
The only change needed was two lines of world.toml:
base_url = "http://127.0.0.1:11434" # was 127.0.0.1:1234
name = "qwen2.5:3b-instruct" # was "replace-with-a-local-model"
python -m local_adventure doctor --world <world>
[PASS] SQLite FTS5 available
[PASS] Database schema at version 2
[PASS] World valid: Ember Hollow
[PASS] LM Studio reachable: http://127.0.0.1:11434
[PASS] Configured model visible: qwen2.5:3b-instruct
A real turn then committed:
TURN: (30ad56d4…, 1, 'committed', 'I search the observatory for the brass k',
'You rummage through the dusty shelves, your fingers brushing away
decades of cobwebs...')
MODELCALL: ('lm_studio', 'qwen2.5:3b-instruct', attempt 1, 832 bytes, errors=False)
ModelSettings.backend is typed Literal["lm_studio"], so adding a second
backend means widening one literal and one factory call in cli.py — the
transport itself already works. Phase 0A's "MODIFY — LM Studio primary provider"
overstated this.
Caveat: with qwen2.5:0.5b the same turn failed with "The model response
could not be validated; no turn was saved." That is the validation contract
working as designed — the engine refuses to commit an unvalidated turn — but the
typed-event contract needs a reasonably capable model, where the freeform-prose
candidates tolerate a weak one.
C3. Undo / branch / checkpoint / replay — verified independently
Run directly against GameService with a temporary database, not by trusting the
project's own tests:
after 2 turns: turns=2 events=2 has_brass_key=False
UNDO
turns 2 → 2 (DELETED 0)
state after undo: has_brass_key=True
abandoned turn id_3 still in DB: True
BRANCH from the restored point, then play on the branch
branch state: has_brass_key=True gate_open=True
ORIGINAL session leaked 'gate_open'? False
total turns in DB: 3 (nothing destroyed)
CHECKPOINT RESTORE ("before_gate")
restored has_brass_key=True
later turn id_3 still present: True
turns in DB after restore: 3
This is the specification's history model, working. undo sets
session.head_turn_id to the parent; branch inserts a new session sharing the
same head; restore_checkpoint replays ancestry; _replay_to walks
turns.ancestry(head) and reduces state_events.
C4. Schema — the target data model
sessions(session_id, ..., head_turn_id, initial_state_json)
turns(turn_id, session_id, parent_turn_id, turn_number,
player_input, narration, status CHECK IN ('committed','failed'), model_call_id)
state_events(event_id, turn_id, sequence_number, event_type, payload_json) -- append-only
state_cache(session_id, head_turn_id, state_json) -- replay cache
model_calls(..., request_hash, response_hash, parsed_response_json,
validation_errors_json, prompt_eval_count, eval_count, duration_ms)
named_checkpoints(checkpoint_id, session_id, turn_id, name, UNIQUE(session_id,name))
summaries(summary_id, session_id, through_turn_id, kind CHECK IN ('scene','campaign'))
lore_documents + lore_documents_fts (FTS5)
Note summaries.through_turn_id — anchored to a turn, not a count. That is
exactly the property Open Dungeon lacks.
C5. Coupling to the CLI
The layering is clean: cli.py → app/commands.py → app/game_service.py /
app/turn_service.py → storage/repositories.py. Authoritative state lives in
state/ (events, reducer, validator) and context assembly in context/. None of
that imports the CLI. A browser/API layer could sit beside cli.py and call the
same services without moving state logic.
What it does not have, and would all be new work: any HTTP layer, streaming, in-app campaign setup (worlds are hand-authored TOML + Markdown on disk), embeddings or semantic memory, document import, media, and prompt inspection beyond stored hashes.
C6. Lore / FTS — reusable
lore/indexer.py + lore_documents_fts (FTS5) with content hashing and
modified_ns for incremental reindex, scoped by world_id, and a kind column
already carrying a classification axis. This is a good starting shape for the
Canon/Reference/Inspiration tiers in IMPORTED-KNOWLEDGE-DESIGN.md, and it is
deterministic and offline.