Validates the three finalists by clone, build, test run and live local Ollama inference, then answers the fork question with measurements rather than static review. Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but it was wrong that AI-DnD's undo is non-destructive — retry preserves the replaced take, undo hard-deletes it. A follow-up spike fixed that in 3 files (+130/-31): undo now moves a head cursor, redo round-trips, writing below a moved-back head forks and keeps the abandoned line, branch-scoped memory isolation survives, suite 627/632 with all 5 failures asserting the deleted-row behaviour that was replaced. Findings that change the plan: - AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its encoding from a CDN. Proven on an internal Docker network, proven fixed by vendoring the file. - ai-adventure needs zero code for Ollama — two config lines — and its turn/head/checkpoint schema is the target model to build to. - Open Dungeon has zero automated tests and a positional summary watermark, making its branch retrofit larger than Phase 0A costed. - The world-state referee takes relative deltas; a 3B model sent absolute values under full context, so a wounded player ended at full health. Validation cannot catch this, so prefer ai-adventure's typed-event vocabulary when generalising narrative state. - Export/import recomputes head depth, so a round-trip silently undoes an undo. Must be fixed alongside the undo work. Docs only; no production code. Working tree from the runs stays untracked under phase0b/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu
120 lines
4.7 KiB
Markdown
120 lines
4.7 KiB
Markdown
# Phase 0B — Experiment C: ai-adventure
|
|
|
|
## C1. Tests
|
|
|
|
```
|
|
python -m unittest discover -s tests → Ran 76 tests in 1.9s OK
|
|
```
|
|
|
|
## C2. Provider abstraction — the Ollama adapter is zero code
|
|
|
|
`llm/backend.py` defines a `ModelBackend` Protocol with a single `generate`
|
|
method plus typed `ModelRequest`/`ModelResponse`. `llm/lm_studio.py` is not
|
|
LM-Studio-specific at all: it POSTs `/v1/chat/completions` and reads `/v1/models`
|
|
— the same OpenAI-compatible surface Ollama serves. `llm/scripted.py` is a
|
|
deterministic fake used by the tests.
|
|
|
|
The only change needed was two lines of `world.toml`:
|
|
|
|
```toml
|
|
base_url = "http://127.0.0.1:11434" # was 127.0.0.1:1234
|
|
name = "qwen2.5:3b-instruct" # was "replace-with-a-local-model"
|
|
```
|
|
|
|
```
|
|
python -m local_adventure doctor --world <world>
|
|
[PASS] SQLite FTS5 available
|
|
[PASS] Database schema at version 2
|
|
[PASS] World valid: Ember Hollow
|
|
[PASS] LM Studio reachable: http://127.0.0.1:11434
|
|
[PASS] Configured model visible: qwen2.5:3b-instruct
|
|
```
|
|
|
|
A real turn then committed:
|
|
|
|
```
|
|
TURN: (30ad56d4…, 1, 'committed', 'I search the observatory for the brass k',
|
|
'You rummage through the dusty shelves, your fingers brushing away
|
|
decades of cobwebs...')
|
|
MODELCALL: ('lm_studio', 'qwen2.5:3b-instruct', attempt 1, 832 bytes, errors=False)
|
|
```
|
|
|
|
`ModelSettings.backend` is typed `Literal["lm_studio"]`, so adding a second
|
|
backend means widening one literal and one factory call in `cli.py` — the
|
|
transport itself already works. Phase 0A's "MODIFY — LM Studio primary provider"
|
|
overstated this.
|
|
|
|
**Caveat:** with `qwen2.5:0.5b` the same turn failed with "The model response
|
|
could not be validated; no turn was saved." That is the validation contract
|
|
working as designed — the engine refuses to commit an unvalidated turn — but the
|
|
typed-event contract needs a reasonably capable model, where the freeform-prose
|
|
candidates tolerate a weak one.
|
|
|
|
## C3. Undo / branch / checkpoint / replay — verified independently
|
|
|
|
Run directly against `GameService` with a temporary database, not by trusting the
|
|
project's own tests:
|
|
|
|
```
|
|
after 2 turns: turns=2 events=2 has_brass_key=False
|
|
|
|
UNDO
|
|
turns 2 → 2 (DELETED 0)
|
|
state after undo: has_brass_key=True
|
|
abandoned turn id_3 still in DB: True
|
|
|
|
BRANCH from the restored point, then play on the branch
|
|
branch state: has_brass_key=True gate_open=True
|
|
ORIGINAL session leaked 'gate_open'? False
|
|
total turns in DB: 3 (nothing destroyed)
|
|
|
|
CHECKPOINT RESTORE ("before_gate")
|
|
restored has_brass_key=True
|
|
later turn id_3 still present: True
|
|
turns in DB after restore: 3
|
|
```
|
|
|
|
This is the specification's history model, working. `undo` sets
|
|
`session.head_turn_id` to the parent; `branch` inserts a new session sharing the
|
|
same head; `restore_checkpoint` replays ancestry; `_replay_to` walks
|
|
`turns.ancestry(head)` and reduces `state_events`.
|
|
|
|
## C4. Schema — the target data model
|
|
|
|
```
|
|
sessions(session_id, ..., head_turn_id, initial_state_json)
|
|
turns(turn_id, session_id, parent_turn_id, turn_number,
|
|
player_input, narration, status CHECK IN ('committed','failed'), model_call_id)
|
|
state_events(event_id, turn_id, sequence_number, event_type, payload_json) -- append-only
|
|
state_cache(session_id, head_turn_id, state_json) -- replay cache
|
|
model_calls(..., request_hash, response_hash, parsed_response_json,
|
|
validation_errors_json, prompt_eval_count, eval_count, duration_ms)
|
|
named_checkpoints(checkpoint_id, session_id, turn_id, name, UNIQUE(session_id,name))
|
|
summaries(summary_id, session_id, through_turn_id, kind CHECK IN ('scene','campaign'))
|
|
lore_documents + lore_documents_fts (FTS5)
|
|
```
|
|
|
|
Note `summaries.through_turn_id` — anchored to a turn, not a count. That is
|
|
exactly the property Open Dungeon lacks.
|
|
|
|
## C5. Coupling to the CLI
|
|
|
|
The layering is clean: `cli.py` → `app/commands.py` → `app/game_service.py` /
|
|
`app/turn_service.py` → `storage/repositories.py`. Authoritative state lives in
|
|
`state/` (events, reducer, validator) and context assembly in `context/`. None of
|
|
that imports the CLI. A browser/API layer could sit beside `cli.py` and call the
|
|
same services without moving state logic.
|
|
|
|
What it does **not** have, and would all be new work: any HTTP layer, streaming,
|
|
in-app campaign setup (worlds are hand-authored TOML + Markdown on disk),
|
|
embeddings or semantic memory, document import, media, and prompt inspection
|
|
beyond stored hashes.
|
|
|
|
## C6. Lore / FTS — reusable
|
|
|
|
`lore/indexer.py` + `lore_documents_fts` (FTS5) with content hashing and
|
|
`modified_ns` for incremental reindex, scoped by `world_id`, and a `kind` column
|
|
already carrying a classification axis. This is a good starting shape for the
|
|
Canon/Reference/Inspiration tiers in `IMPORTED-KNOWLEDGE-DESIGN.md`, and it is
|
|
deterministic and offline.
|