Validates the three finalists by clone, build, test run and live local Ollama inference, then answers the fork question with measurements rather than static review. Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but it was wrong that AI-DnD's undo is non-destructive — retry preserves the replaced take, undo hard-deletes it. A follow-up spike fixed that in 3 files (+130/-31): undo now moves a head cursor, redo round-trips, writing below a moved-back head forks and keeps the abandoned line, branch-scoped memory isolation survives, suite 627/632 with all 5 failures asserting the deleted-row behaviour that was replaced. Findings that change the plan: - AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its encoding from a CDN. Proven on an internal Docker network, proven fixed by vendoring the file. - ai-adventure needs zero code for Ollama — two config lines — and its turn/head/checkpoint schema is the target model to build to. - Open Dungeon has zero automated tests and a positional summary watermark, making its branch retrofit larger than Phase 0A costed. - The world-state referee takes relative deltas; a 3B model sent absolute values under full context, so a wounded player ended at full health. Validation cannot catch this, so prefer ai-adventure's typed-event vocabulary when generalising narrative state. - Export/import recomputes head depth, so a round-trip silently undoes an undo. Must be fixed alongside the undo work. Docs only; no production code. Working tree from the runs stays untracked under phase0b/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu
5.8 KiB
Phase 0B — Experiment A: AI-DnD
All results are from a live server on 127.0.0.1:8321 with a fresh SQLite
database, generating against local Ollama.
A1. Runs with Ollama locally — yes
Default endpoint_url is http://localhost:11434/v1. After setting only the
model name:
POST /api/settings/test → {"ok":true,"models":["qwen2.5:0.5b"]}
Three streamed turns produced 101/162/113 SSE events with player, chunk and
done types and a coherent transcript.
A2. RPG state can be left empty — yes
schemas.AdventureCreate.scenario_id is optional. With it null,
crud.create_adventure sets world_state={} and creates no story cards, no
scripts and no opening action, but still calls tree.head_branch(...) so the
story tree exists from the start. A noir campaign created this way played,
branched, summarized and retrieved memories normally. The RPG referee is
opt-in per scenario, not a core dependency.
A3. Retry is non-destructive; Undo is destructive
Measured on the same adventure, reading rows straight from SQLite:
baseline rows: 6 ids [1,2,3,4,5,6]
RETRY → new action id 7; rows 7 ids [1..7]
/variants at the tip lists BOTH take 0 (id 6) and take 1 (id 7)
UNDO → rows 7 → 4; DELETED ids = [5, 6, 7]
Undo removed the player action (5), the live AI take (6) and the alternate
take that the retry had just preserved (7). nodes.delete_turn deletes every
attempt in the group on that branch and calls memorybank.forget_node. There is
no redo endpoint anywhere in routers/.
This contradicts reports/PRELIMINARY-RECOMMENDATION.md, which credits AI-DnD
with "non-destructive retry" (true) and generalizes it to rollback (false).
A4. Branching is non-destructive
Adding an alternate take at a passed turn (POST /actions/3/takes) forked
automatically:
before: ids 1-8, all branch 1, depths 0-7
after : ids 1-8 unchanged on branch 1
id 9 branch 2 depth 2 (the alternate player action)
id 10 branch 2 depth 3 (its AI continuation)
/branches → [ {id:1, parent:null, fork_depth:null, own_actions:8},
{id:2, parent:1, fork_depth:1, own_actions:2, is_head:true} ]
Not one row of branch 1 was touched. Note the semantics: after_id pointing at a
node that is live and on the path is a no-op, not a fork. A fork happens only
when writing below a take the story has moved past (nodes.stand_on). So
"restore to an arbitrary earlier turn and branch" is not directly expressible
today — it must go through creating a take at that turn.
A5. Memory isolation across branches — correct, with negative control
Memory bank enabled, summary model qwen2.5:0.5b, embeddings
nomic-embed-text, memory_top_k=5. Three more turns played on branch 1
establishing possessions (brass key, silver revolver, bullets).
memories stored: (id 1, branch 1, depth 5)
(id 2, branch 1, depth 11)
on branch 2 (forked at depth 1):
memories retrieved = {"used": [], "error": null}
leak probes over the FULL assembled prompt:
'brass key' False | 'revolver' False | 'lockbox' False
'ashtray' False | 'bullets' False
NEGATIVE CONTROL — switch back to branch 1:
memories retrieved = 2, top similarity 0.8087
'revolver' present in prompt = True
The empty result on branch 2 is genuine isolation, not a dead retrieval path.
memorybank.retrieve_memories applies lineage.path_of(db, adventure).clause(models.Memory),
and Memory rows carry branch_id plus source depths. Eviction deliberately
ignores the branch clause, which is correct — it is a capacity concern.
History is likewise lineage-scoped: branch 2's prompt contained only its own text.
A6. Insights / prompt transparency — already sufficient
GET /adventures/{id}/context returns the exact next prompt broken into labeled
sections with token costs:
narrator 57 | ai_instructions 9 | persona 15 | plot_essentials 24
history 988 | used_memories 143 | length_hint 46
Per-action snapshots are stored on Action.context_snapshot and served by
GET /actions/{id}/context. This satisfies specification §11 as-is.
A7. Scripting/QuickJS can be removed cheaply
quickjs is imported in exactly one file (scripting/engine.py). The whole
feature is reachable through two symbols (run_hook, ScriptPipeline) across
four call sites plus routers/scripts.py.
Disposable experiment: made quickjs unimportable via a meta_path blocker and
replaced run_hook with a pass-through stub.
app.main imports OK with scripting stubbed and quickjs absent
19 failed, 613 passed (3.0% of the suite)
Every failure is in test_turn_flow_integration, test_take_state,
test_story_tree_baseline or test_retry_variants, and every one fails for the
same reason: those tests use a JS output_js hook as the instrument to make
state observable, e.g. test_play_then_undo_reverts_gold seeds a GOLD_SCRIPT
and asserts script_state == {"gold": 10} then {} after undo. The behavior
under test (per-node state rollback) is intact; only the measuring device is
gone. Re-instrumenting against world_state_after, which is already snapshotted
per node, is the fix.
A8. Hosted/cloud features
MULTI_USER is off by default and gates ten modules
(auth, limits, netguard, cleanup, security, analytics,
routers/{auth,analytics,debug,story_cards}). analytics.py is a self-hosted
counter writing to two local tables — no third-party tracker and no outbound
call — so it is removable rather than dangerous. render.yaml, the demo-key
path and psycopg are the hosted leftovers.
netguard.py is worth flagging: it refuses endpoints that resolve to
non-public addresses, and only in multi-user mode. That is SSRF protection
for a hosted deployment. Our requirement is the opposite — warn on endpoints that
are not loopback. See ai-adventure's _endpoint_warnings for the shape we want.