Add Phase 0B local validation findings and recommendation
Validates the three finalists by clone, build, test run and live local Ollama inference, then answers the fork question with measurements rather than static review. Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but it was wrong that AI-DnD's undo is non-destructive — retry preserves the replaced take, undo hard-deletes it. A follow-up spike fixed that in 3 files (+130/-31): undo now moves a head cursor, redo round-trips, writing below a moved-back head forks and keeps the abandoned line, branch-scoped memory isolation survives, suite 627/632 with all 5 failures asserting the deleted-row behaviour that was replaced. Findings that change the plan: - AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its encoding from a CDN. Proven on an internal Docker network, proven fixed by vendoring the file. - ai-adventure needs zero code for Ollama — two config lines — and its turn/head/checkpoint schema is the target model to build to. - Open Dungeon has zero automated tests and a positional summary watermark, making its branch retrofit larger than Phase 0A costed. - The world-state referee takes relative deltas; a 3B model sent absolute values under full context, so a wounded player ended at full health. Validation cannot catch this, so prefer ai-adventure's typed-event vocabulary when generalising narrative state. - Export/import recomputes head depth, so a round-trip silently undoes an undo. Must be fixed alongside the undo work. Docs only; no production code. Working tree from the runs stays untracked under phase0b/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu
This commit is contained in:
co-authored by
Claude Opus 5
parent
f011362494
commit
ba737de9b4
@@ -0,0 +1,119 @@
|
||||
# Phase 0B — Experiment C: ai-adventure
|
||||
|
||||
## C1. Tests
|
||||
|
||||
```
|
||||
python -m unittest discover -s tests → Ran 76 tests in 1.9s OK
|
||||
```
|
||||
|
||||
## C2. Provider abstraction — the Ollama adapter is zero code
|
||||
|
||||
`llm/backend.py` defines a `ModelBackend` Protocol with a single `generate`
|
||||
method plus typed `ModelRequest`/`ModelResponse`. `llm/lm_studio.py` is not
|
||||
LM-Studio-specific at all: it POSTs `/v1/chat/completions` and reads `/v1/models`
|
||||
— the same OpenAI-compatible surface Ollama serves. `llm/scripted.py` is a
|
||||
deterministic fake used by the tests.
|
||||
|
||||
The only change needed was two lines of `world.toml`:
|
||||
|
||||
```toml
|
||||
base_url = "http://127.0.0.1:11434" # was 127.0.0.1:1234
|
||||
name = "qwen2.5:3b-instruct" # was "replace-with-a-local-model"
|
||||
```
|
||||
|
||||
```
|
||||
python -m local_adventure doctor --world <world>
|
||||
[PASS] SQLite FTS5 available
|
||||
[PASS] Database schema at version 2
|
||||
[PASS] World valid: Ember Hollow
|
||||
[PASS] LM Studio reachable: http://127.0.0.1:11434
|
||||
[PASS] Configured model visible: qwen2.5:3b-instruct
|
||||
```
|
||||
|
||||
A real turn then committed:
|
||||
|
||||
```
|
||||
TURN: (30ad56d4…, 1, 'committed', 'I search the observatory for the brass k',
|
||||
'You rummage through the dusty shelves, your fingers brushing away
|
||||
decades of cobwebs...')
|
||||
MODELCALL: ('lm_studio', 'qwen2.5:3b-instruct', attempt 1, 832 bytes, errors=False)
|
||||
```
|
||||
|
||||
`ModelSettings.backend` is typed `Literal["lm_studio"]`, so adding a second
|
||||
backend means widening one literal and one factory call in `cli.py` — the
|
||||
transport itself already works. Phase 0A's "MODIFY — LM Studio primary provider"
|
||||
overstated this.
|
||||
|
||||
**Caveat:** with `qwen2.5:0.5b` the same turn failed with "The model response
|
||||
could not be validated; no turn was saved." That is the validation contract
|
||||
working as designed — the engine refuses to commit an unvalidated turn — but the
|
||||
typed-event contract needs a reasonably capable model, where the freeform-prose
|
||||
candidates tolerate a weak one.
|
||||
|
||||
## C3. Undo / branch / checkpoint / replay — verified independently
|
||||
|
||||
Run directly against `GameService` with a temporary database, not by trusting the
|
||||
project's own tests:
|
||||
|
||||
```
|
||||
after 2 turns: turns=2 events=2 has_brass_key=False
|
||||
|
||||
UNDO
|
||||
turns 2 → 2 (DELETED 0)
|
||||
state after undo: has_brass_key=True
|
||||
abandoned turn id_3 still in DB: True
|
||||
|
||||
BRANCH from the restored point, then play on the branch
|
||||
branch state: has_brass_key=True gate_open=True
|
||||
ORIGINAL session leaked 'gate_open'? False
|
||||
total turns in DB: 3 (nothing destroyed)
|
||||
|
||||
CHECKPOINT RESTORE ("before_gate")
|
||||
restored has_brass_key=True
|
||||
later turn id_3 still present: True
|
||||
turns in DB after restore: 3
|
||||
```
|
||||
|
||||
This is the specification's history model, working. `undo` sets
|
||||
`session.head_turn_id` to the parent; `branch` inserts a new session sharing the
|
||||
same head; `restore_checkpoint` replays ancestry; `_replay_to` walks
|
||||
`turns.ancestry(head)` and reduces `state_events`.
|
||||
|
||||
## C4. Schema — the target data model
|
||||
|
||||
```
|
||||
sessions(session_id, ..., head_turn_id, initial_state_json)
|
||||
turns(turn_id, session_id, parent_turn_id, turn_number,
|
||||
player_input, narration, status CHECK IN ('committed','failed'), model_call_id)
|
||||
state_events(event_id, turn_id, sequence_number, event_type, payload_json) -- append-only
|
||||
state_cache(session_id, head_turn_id, state_json) -- replay cache
|
||||
model_calls(..., request_hash, response_hash, parsed_response_json,
|
||||
validation_errors_json, prompt_eval_count, eval_count, duration_ms)
|
||||
named_checkpoints(checkpoint_id, session_id, turn_id, name, UNIQUE(session_id,name))
|
||||
summaries(summary_id, session_id, through_turn_id, kind CHECK IN ('scene','campaign'))
|
||||
lore_documents + lore_documents_fts (FTS5)
|
||||
```
|
||||
|
||||
Note `summaries.through_turn_id` — anchored to a turn, not a count. That is
|
||||
exactly the property Open Dungeon lacks.
|
||||
|
||||
## C5. Coupling to the CLI
|
||||
|
||||
The layering is clean: `cli.py` → `app/commands.py` → `app/game_service.py` /
|
||||
`app/turn_service.py` → `storage/repositories.py`. Authoritative state lives in
|
||||
`state/` (events, reducer, validator) and context assembly in `context/`. None of
|
||||
that imports the CLI. A browser/API layer could sit beside `cli.py` and call the
|
||||
same services without moving state logic.
|
||||
|
||||
What it does **not** have, and would all be new work: any HTTP layer, streaming,
|
||||
in-app campaign setup (worlds are hand-authored TOML + Markdown on disk),
|
||||
embeddings or semantic memory, document import, media, and prompt inspection
|
||||
beyond stored hashes.
|
||||
|
||||
## C6. Lore / FTS — reusable
|
||||
|
||||
`lore/indexer.py` + `lore_documents_fts` (FTS5) with content hashing and
|
||||
`modified_ns` for incremental reindex, scoped by `world_id`, and a `kind` column
|
||||
already carrying a classification axis. This is a good starting shape for the
|
||||
Canon/Reference/Inspiration tiers in `IMPORTED-KNOWLEDGE-DESIGN.md`, and it is
|
||||
deterministic and offline.
|
||||
Reference in New Issue
Block a user