Add Phase 0B local validation findings and recommendation
Validates the three finalists by clone, build, test run and live local Ollama inference, then answers the fork question with measurements rather than static review. Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but it was wrong that AI-DnD's undo is non-destructive — retry preserves the replaced take, undo hard-deletes it. A follow-up spike fixed that in 3 files (+130/-31): undo now moves a head cursor, redo round-trips, writing below a moved-back head forks and keeps the abandoned line, branch-scoped memory isolation survives, suite 627/632 with all 5 failures asserting the deleted-row behaviour that was replaced. Findings that change the plan: - AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its encoding from a CDN. Proven on an internal Docker network, proven fixed by vendoring the file. - ai-adventure needs zero code for Ollama — two config lines — and its turn/head/checkpoint schema is the target model to build to. - Open Dungeon has zero automated tests and a positional summary watermark, making its branch retrofit larger than Phase 0A costed. - The world-state referee takes relative deltas; a 3B model sent absolute values under full context, so a wounded player ended at full health. Validation cannot catch this, so prefer ai-adventure's typed-event vocabulary when generalising narrative state. - Export/import recomputes head depth, so a round-trip silently undoes an undo. Must be fixed alongside the undo work. Docs only; no production code. Working tree from the runs stays untracked under phase0b/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu
This commit is contained in:
co-authored by
Claude Opus 5
parent
f011362494
commit
ba737de9b4
@@ -0,0 +1,146 @@
|
||||
# Phase 0B — Experiment A: AI-DnD
|
||||
|
||||
All results are from a live server on `127.0.0.1:8321` with a fresh SQLite
|
||||
database, generating against local Ollama.
|
||||
|
||||
## A1. Runs with Ollama locally — yes
|
||||
|
||||
Default `endpoint_url` is `http://localhost:11434/v1`. After setting only the
|
||||
model name:
|
||||
|
||||
```
|
||||
POST /api/settings/test → {"ok":true,"models":["qwen2.5:0.5b"]}
|
||||
```
|
||||
|
||||
Three streamed turns produced 101/162/113 SSE events with `player`, `chunk` and
|
||||
`done` types and a coherent transcript.
|
||||
|
||||
## A2. RPG state can be left empty — yes
|
||||
|
||||
`schemas.AdventureCreate.scenario_id` is optional. With it null,
|
||||
`crud.create_adventure` sets `world_state={}` and creates no story cards, no
|
||||
scripts and no opening action, but still calls `tree.head_branch(...)` so the
|
||||
story tree exists from the start. A noir campaign created this way played,
|
||||
branched, summarized and retrieved memories normally. The RPG referee is
|
||||
opt-in per scenario, not a core dependency.
|
||||
|
||||
## A3. Retry is non-destructive; Undo is destructive
|
||||
|
||||
Measured on the same adventure, reading rows straight from SQLite:
|
||||
|
||||
```
|
||||
baseline rows: 6 ids [1,2,3,4,5,6]
|
||||
|
||||
RETRY → new action id 7; rows 7 ids [1..7]
|
||||
/variants at the tip lists BOTH take 0 (id 6) and take 1 (id 7)
|
||||
|
||||
UNDO → rows 7 → 4; DELETED ids = [5, 6, 7]
|
||||
```
|
||||
|
||||
Undo removed the player action (5), the live AI take (6) **and the alternate
|
||||
take that the retry had just preserved (7)**. `nodes.delete_turn` deletes every
|
||||
attempt in the group on that branch and calls `memorybank.forget_node`. There is
|
||||
no redo endpoint anywhere in `routers/`.
|
||||
|
||||
This contradicts `reports/PRELIMINARY-RECOMMENDATION.md`, which credits AI-DnD
|
||||
with "non-destructive retry" (true) and generalizes it to rollback (false).
|
||||
|
||||
## A4. Branching is non-destructive
|
||||
|
||||
Adding an alternate take at a passed turn (`POST /actions/3/takes`) forked
|
||||
automatically:
|
||||
|
||||
```
|
||||
before: ids 1-8, all branch 1, depths 0-7
|
||||
after : ids 1-8 unchanged on branch 1
|
||||
id 9 branch 2 depth 2 (the alternate player action)
|
||||
id 10 branch 2 depth 3 (its AI continuation)
|
||||
|
||||
/branches → [ {id:1, parent:null, fork_depth:null, own_actions:8},
|
||||
{id:2, parent:1, fork_depth:1, own_actions:2, is_head:true} ]
|
||||
```
|
||||
|
||||
Not one row of branch 1 was touched. Note the semantics: `after_id` pointing at a
|
||||
node that is *live and on the path* is a no-op, not a fork. A fork happens only
|
||||
when writing below a take the story has moved past (`nodes.stand_on`). So
|
||||
"restore to an arbitrary earlier turn and branch" is not directly expressible
|
||||
today — it must go through creating a take at that turn.
|
||||
|
||||
## A5. Memory isolation across branches — correct, with negative control
|
||||
|
||||
Memory bank enabled, summary model `qwen2.5:0.5b`, embeddings
|
||||
`nomic-embed-text`, `memory_top_k=5`. Three more turns played on branch 1
|
||||
establishing possessions (brass key, silver revolver, bullets).
|
||||
|
||||
```
|
||||
memories stored: (id 1, branch 1, depth 5)
|
||||
(id 2, branch 1, depth 11)
|
||||
|
||||
on branch 2 (forked at depth 1):
|
||||
memories retrieved = {"used": [], "error": null}
|
||||
leak probes over the FULL assembled prompt:
|
||||
'brass key' False | 'revolver' False | 'lockbox' False
|
||||
'ashtray' False | 'bullets' False
|
||||
|
||||
NEGATIVE CONTROL — switch back to branch 1:
|
||||
memories retrieved = 2, top similarity 0.8087
|
||||
'revolver' present in prompt = True
|
||||
```
|
||||
|
||||
The empty result on branch 2 is genuine isolation, not a dead retrieval path.
|
||||
`memorybank.retrieve_memories` applies `lineage.path_of(db, adventure).clause(models.Memory)`,
|
||||
and `Memory` rows carry `branch_id` plus source depths. Eviction deliberately
|
||||
ignores the branch clause, which is correct — it is a capacity concern.
|
||||
|
||||
History is likewise lineage-scoped: branch 2's prompt contained only its own
|
||||
text.
|
||||
|
||||
## A6. Insights / prompt transparency — already sufficient
|
||||
|
||||
`GET /adventures/{id}/context` returns the exact next prompt broken into labeled
|
||||
sections with token costs:
|
||||
|
||||
```
|
||||
narrator 57 | ai_instructions 9 | persona 15 | plot_essentials 24
|
||||
history 988 | used_memories 143 | length_hint 46
|
||||
```
|
||||
|
||||
Per-action snapshots are stored on `Action.context_snapshot` and served by
|
||||
`GET /actions/{id}/context`. This satisfies specification §11 as-is.
|
||||
|
||||
## A7. Scripting/QuickJS can be removed cheaply
|
||||
|
||||
`quickjs` is imported in exactly one file (`scripting/engine.py`). The whole
|
||||
feature is reachable through two symbols (`run_hook`, `ScriptPipeline`) across
|
||||
four call sites plus `routers/scripts.py`.
|
||||
|
||||
Disposable experiment: made `quickjs` unimportable via a `meta_path` blocker and
|
||||
replaced `run_hook` with a pass-through stub.
|
||||
|
||||
```
|
||||
app.main imports OK with scripting stubbed and quickjs absent
|
||||
19 failed, 613 passed (3.0% of the suite)
|
||||
```
|
||||
|
||||
Every failure is in `test_turn_flow_integration`, `test_take_state`,
|
||||
`test_story_tree_baseline` or `test_retry_variants`, and every one fails for the
|
||||
same reason: those tests **use a JS `output_js` hook as the instrument** to make
|
||||
state observable, e.g. `test_play_then_undo_reverts_gold` seeds a `GOLD_SCRIPT`
|
||||
and asserts `script_state == {"gold": 10}` then `{}` after undo. The behavior
|
||||
under test (per-node state rollback) is intact; only the measuring device is
|
||||
gone. Re-instrumenting against `world_state_after`, which is already snapshotted
|
||||
per node, is the fix.
|
||||
|
||||
## A8. Hosted/cloud features
|
||||
|
||||
`MULTI_USER` is off by default and gates ten modules
|
||||
(`auth`, `limits`, `netguard`, `cleanup`, `security`, `analytics`,
|
||||
`routers/{auth,analytics,debug,story_cards}`). `analytics.py` is a *self-hosted*
|
||||
counter writing to two local tables — no third-party tracker and no outbound
|
||||
call — so it is removable rather than dangerous. `render.yaml`, the demo-key
|
||||
path and `psycopg` are the hosted leftovers.
|
||||
|
||||
`netguard.py` is worth flagging: it refuses endpoints that resolve to
|
||||
**non-public** addresses, and only in multi-user mode. That is SSRF protection
|
||||
for a hosted deployment. Our requirement is the opposite — warn on endpoints that
|
||||
are **not loopback**. See ai-adventure's `_endpoint_warnings` for the shape we want.
|
||||
Reference in New Issue
Block a user