The planning package had grown to where a new agent could not tell what was authoritative. Phase 0 execution prompts sat beside the specification; four completed milestone reports sat beside the current one; and upstream AI-DnD's own `plan/` build log and `docs/` project site still described a hosted, scripted, multi-user product with accounts — every screenshot in it showed a Scripts tab and a Sign up button, none of which has existed since M2. `planning/archive/` now holds the history and says so in its own README: `phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied. `planning/reports/` holds only the current milestone's report, because that is the one M4 planning has to read; it moves to the archive when M4's replaces it. Deleted rather than archived: the Phase 0B execution prompts and the handoff/status/summary documents, the Phase 0A discovery and triage reports, upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's template boilerplate. All of it is in Git history, and the two recommendation reports carry every conclusion the deleted research reached. Archived documents are kept verbatim. Paths written inside them point at where those files were when the document was written, which is the point: an evidence record that has been quietly edited is no longer evidence. Active documentation is corrected where it pointed at the removed trees or described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not touch" list had gone stale at M2 and claimed QuickJS scripting was still tested; its test count was 604 against an actual 638. `README.md` loses the upstream CI badge, which reported upstream's pipeline rather than this fork's, and a reference to `backend/app/worldstate/engine.py`, a file that does not exist. `planning/README.md` is rewritten as the documentation index. New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the manifest of what belongs in the ChatGPT project's Sources. Source comments referring to the deleted trees are reworded; no behaviour changes. 638 backend tests pass, frontend lints and builds, and a reference scan over all 48 tracked Markdown files reports no unresolved path in active documentation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
5.8 KiB
Phase 0B — Experiment A: AI-DnD
All results are from a live server on 127.0.0.1:8321 with a fresh SQLite
database, generating against local Ollama.
A1. Runs with Ollama locally — yes
Default endpoint_url is http://localhost:11434/v1. After setting only the
model name:
POST /api/settings/test → {"ok":true,"models":["qwen2.5:0.5b"]}
Three streamed turns produced 101/162/113 SSE events with player, chunk and
done types and a coherent transcript.
A2. RPG state can be left empty — yes
schemas.AdventureCreate.scenario_id is optional. With it null,
crud.create_adventure sets world_state={} and creates no story cards, no
scripts and no opening action, but still calls tree.head_branch(...) so the
story tree exists from the start. A noir campaign created this way played,
branched, summarized and retrieved memories normally. The RPG referee is
opt-in per scenario, not a core dependency.
A3. Retry is non-destructive; Undo is destructive
Measured on the same adventure, reading rows straight from SQLite:
baseline rows: 6 ids [1,2,3,4,5,6]
RETRY → new action id 7; rows 7 ids [1..7]
/variants at the tip lists BOTH take 0 (id 6) and take 1 (id 7)
UNDO → rows 7 → 4; DELETED ids = [5, 6, 7]
Undo removed the player action (5), the live AI take (6) and the alternate
take that the retry had just preserved (7). nodes.delete_turn deletes every
attempt in the group on that branch and calls memorybank.forget_node. There is
no redo endpoint anywhere in routers/.
This contradicts reports/PRELIMINARY-RECOMMENDATION.md, which credits AI-DnD
with "non-destructive retry" (true) and generalizes it to rollback (false).
A4. Branching is non-destructive
Adding an alternate take at a passed turn (POST /actions/3/takes) forked
automatically:
before: ids 1-8, all branch 1, depths 0-7
after : ids 1-8 unchanged on branch 1
id 9 branch 2 depth 2 (the alternate player action)
id 10 branch 2 depth 3 (its AI continuation)
/branches → [ {id:1, parent:null, fork_depth:null, own_actions:8},
{id:2, parent:1, fork_depth:1, own_actions:2, is_head:true} ]
Not one row of branch 1 was touched. Note the semantics: after_id pointing at a
node that is live and on the path is a no-op, not a fork. A fork happens only
when writing below a take the story has moved past (nodes.stand_on). So
"restore to an arbitrary earlier turn and branch" is not directly expressible
today — it must go through creating a take at that turn.
A5. Memory isolation across branches — correct, with negative control
Memory bank enabled, summary model qwen2.5:0.5b, embeddings
nomic-embed-text, memory_top_k=5. Three more turns played on branch 1
establishing possessions (brass key, silver revolver, bullets).
memories stored: (id 1, branch 1, depth 5)
(id 2, branch 1, depth 11)
on branch 2 (forked at depth 1):
memories retrieved = {"used": [], "error": null}
leak probes over the FULL assembled prompt:
'brass key' False | 'revolver' False | 'lockbox' False
'ashtray' False | 'bullets' False
NEGATIVE CONTROL — switch back to branch 1:
memories retrieved = 2, top similarity 0.8087
'revolver' present in prompt = True
The empty result on branch 2 is genuine isolation, not a dead retrieval path.
memorybank.retrieve_memories applies lineage.path_of(db, adventure).clause(models.Memory),
and Memory rows carry branch_id plus source depths. Eviction deliberately
ignores the branch clause, which is correct — it is a capacity concern.
History is likewise lineage-scoped: branch 2's prompt contained only its own text.
A6. Insights / prompt transparency — already sufficient
GET /adventures/{id}/context returns the exact next prompt broken into labeled
sections with token costs:
narrator 57 | ai_instructions 9 | persona 15 | plot_essentials 24
history 988 | used_memories 143 | length_hint 46
Per-action snapshots are stored on Action.context_snapshot and served by
GET /actions/{id}/context. This satisfies specification §11 as-is.
A7. Scripting/QuickJS can be removed cheaply
quickjs is imported in exactly one file (scripting/engine.py). The whole
feature is reachable through two symbols (run_hook, ScriptPipeline) across
four call sites plus routers/scripts.py.
Disposable experiment: made quickjs unimportable via a meta_path blocker and
replaced run_hook with a pass-through stub.
app.main imports OK with scripting stubbed and quickjs absent
19 failed, 613 passed (3.0% of the suite)
Every failure is in test_turn_flow_integration, test_take_state,
test_story_tree_baseline or test_retry_variants, and every one fails for the
same reason: those tests use a JS output_js hook as the instrument to make
state observable, e.g. test_play_then_undo_reverts_gold seeds a GOLD_SCRIPT
and asserts script_state == {"gold": 10} then {} after undo. The behavior
under test (per-node state rollback) is intact; only the measuring device is
gone. Re-instrumenting against world_state_after, which is already snapshotted
per node, is the fix.
A8. Hosted/cloud features
MULTI_USER is off by default and gates ten modules
(auth, limits, netguard, cleanup, security, analytics,
routers/{auth,analytics,debug,story_cards}). analytics.py is a self-hosted
counter writing to two local tables — no third-party tracker and no outbound
call — so it is removable rather than dangerous. render.yaml, the demo-key
path and psycopg are the hosted leftovers.
netguard.py is worth flagging: it refuses endpoints that resolve to
non-public addresses, and only in multi-user mode. That is SSRF protection
for a hosted deployment. Our requirement is the opposite — warn on endpoints that
are not loopback. See ai-adventure's _endpoint_warnings for the shape we want.