Files
interactive-story/planning/archive/phase0/PHASE-0B-RECOMMENDATION.md
JesseMarkowitzandClaude Opus 5 d27ee34901 Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was
authoritative. Phase 0 execution prompts sat beside the specification; four
completed milestone reports sat beside the current one; and upstream AI-DnD's
own `plan/` build log and `docs/` project site still described a hosted,
scripted, multi-user product with accounts — every screenshot in it showed a
Scripts tab and a Sign up button, none of which has existed since M2.

`planning/archive/` now holds the history and says so in its own README:
`phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and
M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied.
`planning/reports/` holds only the current milestone's report, because that is
the one M4 planning has to read; it moves to the archive when M4's replaces it.

Deleted rather than archived: the Phase 0B execution prompts and the
handoff/status/summary documents, the Phase 0A discovery and triage reports,
upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's
template boilerplate. All of it is in Git history, and the two recommendation
reports carry every conclusion the deleted research reached.

Archived documents are kept verbatim. Paths written inside them point at where
those files were when the document was written, which is the point: an evidence
record that has been quietly edited is no longer evidence.

Active documentation is corrected where it pointed at the removed trees or
described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not
touch" list had gone stale at M2 and claimed QuickJS scripting was still tested;
its test count was 604 against an actual 638. `README.md` loses the upstream CI
badge, which reported upstream's pipeline rather than this fork's, and a
reference to `backend/app/worldstate/engine.py`, a file that does not exist.
`planning/README.md` is rewritten as the documentation index.

New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
manifest of what belongs in the ChatGPT project's Sources.

Source comments referring to the deleted trees are reworded; no behaviour
changes. 638 backend tests pass, frontend lints and builds, and a reference scan
over all 48 tracked Markdown files reports no unresolved path in active
documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
2026-09-03 14:33:07 -04:00

502 lines
28 KiB
Markdown

> **Planning review note (2026-09-01):** This is the coding agent's Phase 0B evidence report. It is retained substantially as delivered. Some recommendations/open questions in the report were superseded during planning review: ai-adventure is an implementation reference rather than the project specification; no automatic abandoned-history cleanup is required in v1; Open Dungeon media portability is deferred; and the architecture decision is closed in ADR 009/010 and `TECHNICAL-DESIGN.md`.
# Phase 0B — Local Validation Findings and Recommendation
**Date:** 2026-09-01
**Status:** Complete, and updated after the follow-up undo/redo spike.
Stop condition reached; no production work started.
**Method:** Clean clones, pinned SHAs, documented installs, full test runs, real
local Ollama inference, and targeted disposable experiments.
> **Update — the recommended spike was run and it passed.** Section E previously
> recommended proving non-destructive Undo/Redo in AI-DnD before committing to
> the fork. That spike is now complete: 3 files, +130/-31 lines, undo deletes
> zero rows, redo round-trips exactly, writing below a moved-back head forks and
> preserves the abandoned line, branch-scoped memory isolation survives, and the
> suite goes 627/632 with all 5 failures asserting the deleted-row behavior that
> was deliberately replaced. Section E has been rewritten accordingly, and the
> open questions it resolved are marked in section F. Full detail:
> `reports/PHASE-0B-UNDO-SPIKE.md`.
>
> **Update 2 — four follow-up checks run.** The world-state referee was exercised
> against local models for the first time, and export/import, story-card lineage
> safety and Postgres removability were settled. One new defect found (a
> round-trip silently undoes an undo) and one design conclusion that changes what
> we should build (§H). Detail: `reports/PHASE-0B-FOLLOWUP-CHECKS.md`.
Pinned commits (cloned 2026-09-01):
| Repo | SHA | Last commit |
|---|---|---|
| AI-DnD | `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` | 2026-08-31 |
| Open Dungeon | `b0a79f96bf852be7b4e53908dff6a7f7c179da23` | 2026-07-06 |
| ai-adventure | `873ea9180d5b611576cddb155921fc16a17ae88b` | 2026-07-21 |
Inference used for every live test: Ollama in Docker, `qwen2.5:0.5b` and
`qwen2.5:3b-instruct` for narration, `nomic-embed-text` for embeddings.
CPU only, no GPU.
---
## A. Executive Recommendation
**Fork AI-DnD as the production base. Confidence: high.**
**The Phase 0A recommendation held up, but for partly different reasons than it
gave, and it was wrong about one important thing.**
What held: AI-DnD really does implement the expensive correctness work, and it
survived every entanglement test Phase 0A worried about. The RPG machinery is
optional, the scripting sandbox is removable, and the branch-scoped memory
isolation — the single hardest requirement in the specification — works
correctly against real local embeddings.
What did not hold: **Phase 0A's claim that AI-DnD's undo is non-destructive is
false.** `POST /undo` hard-deletes rows. Retry is non-destructive; undo is not.
There is no Redo. This is a direct conflict with `DECISIONS/005-branch-preserving-history.md`
and specification §4.
That discovery does not change the decision, because the comparison is relative:
AI-DnD needs one bounded change at a chokepoint that every read already passes
through. Open Dungeon needs the entire non-destructive history subsystem built
from nothing, under a 3,991-line component, with **zero automated tests**.
**That "one bounded change" is no longer an estimate.** The follow-up spike
implemented it in 3 files and +130/-31 lines, and it works: undo deletes nothing,
redo restores exactly, the abandoned tail becomes an ordinary branch, and memory
isolation holds. Confidence in the fork decision moves from "high on the evidence"
to "high and demonstrated".
The decision gate from `reports/PRELIMINARY-RECOMMENDATION.md` said to select
AI-DnD unless a blocker is confirmed. All four were tested and none is a blocker:
| Phase 0A blocker | Result |
|---|---|
| Branching/state inseparable from RPG mechanics | **No.** A scenario-less adventure with empty world state played, branched, and retrieved memories correctly. |
| Removing hosted/scripting destabilizes many tests | **No.** 19 of 632 (3%), and every one fails for instrumentation reasons, not coupling. |
| Local-only still needs external services | **Partly — but both are packaging fixes.** `tiktoken` downloads its encoding from a Microsoft CDN; the SPA fetches Google Fonts. |
| Dependency/security burden worse than expected | **No.** 80 packages total vs Open Dungeon's 335 with 6 high-severity advisories. |
---
## B. What We Learned About Each Candidate
### AI-DnD — recommended base
**What worked**
- `632 passed` in 236s, first try, no fixes needed.
- Docker image builds clean; `docker compose` path is real.
- Default model endpoint is already `http://localhost:11434/v1`. Connection test
and model discovery against Ollama worked with no code change.
- Streamed a full multi-turn story over SSE against local Ollama.
- **Genre-neutrality is real.** An adventure created with `scenario_id: null`
gets an empty world state and no RPG scaffolding, and plays normally. Noir
prose, no fantasy code paths, no stats.
- **Retry is genuinely non-destructive.** Retrying kept the replaced attempt as a
sibling take at the same coordinate (rows 6 and 7 both present; both listed by
`/variants`).
- **Forking is genuinely non-destructive.** Adding an alternate take at a passed
turn created branch 2 while every row of branch 1 survived, with
`parent_branch_id` and `fork_depth` recorded.
- **Branch-scoped memory isolation is correct.** Two memories were written on
branch 1 (depths 5 and 11) from real `nomic-embed-text` embeddings. On branch 2,
retrieval returned `used: []` and none of `brass key`, `revolver`, `lockbox`,
`ashtray` or `bullets` appeared anywhere in the assembled prompt. The negative
control passed: switching back to branch 1 returned both memories with
similarity scores (0.8087) and the abandoned facts reappeared. Retrieval is
working *and* correctly scoped.
- **Prompt transparency is already there.** The context endpoint returns labeled
sections with token costs: `narrator 57, ai_instructions 9, persona 15,
plot_essentials 24, history 988, used_memories 143, length_hint 46`.
**What failed**
- **Undo hard-deletes.** Measured: rows 7 → 4, deleting ids 5, 6 **and 7** — that
is, both the player action, the live AI take, *and the alternate take the
earlier retry had preserved*. `delete_turn` also calls `memorybank.forget_node`.
A retry's preserved alternative does not survive a later undo. There is no Redo
endpoint. **Resolved in the follow-up spike** — see section G.
- **Fails offline on a clean install.** First turn in a network-isolated container
died with `NameResolutionError` for
`openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken`. `tiktoken`
fetches its BPE encoding on first use. Seeding the 1.7 MB file into the image
made a full turn generate with no internet at all. Packaging fix, not
architecture — but a hard blocker until fixed, and static review missed it.
- **Runtime remote asset.** The built SPA still requests Google Fonts; the CSP
explicitly allows `fonts.googleapis.com` and `fonts.gstatic.com`.
**Strongest reusable pieces**
Immutable parent-linked story tree with lineage entries; alternate takes; branch
fork/switch/rename/delete; per-node world-state and script-state snapshots;
branch-scoped memories with local embeddings; the Insights prompt snapshot;
`ai-dnd-adventure-v2` whole-tree export; the 632-test suite; SSE streaming.
**Major architectural problems**
Destructive undo and no redo. No named checkpoints (branch names are the nearest
thing). Multi-user/hosted concerns are spread across ten modules. `netguard.py`
is the *inverse* of what we need — it blocks non-public endpoints in hosted mode
and does nothing locally; we want a warning on non-loopback endpoints.
**Adaptation required**
Bounded and mostly subtractive. The one genuinely additive change is
non-destructive undo/redo, now measured at 3 files and +130/-31 lines (section G).
### Open Dungeon — not recommended as base
**What worked**
Installed and ran on first try. Plays a real story against Ollama through its
"custom" OpenAI-compatible provider. Clean browser UX. No runtime Google Fonts
(`next/font` self-hosts at build time). Local SQLite, local image generation,
character portrait continuity.
**What failed**
- **Every history operation is destructive, confirmed by measurement.**
- Edit is an in-place `UPDATE messages SET content = ?`. After editing, the
original text `brass key from the ashtray` was **gone from every column of
every table**.
- Retry and Erase both call `DELETE /messages/{id}?after=1`, which runs
`DELETE FROM messages WHERE chat_id = ? AND (created_at > ? OR ...)`. Measured
rows 5 → 4, with no archive anywhere.
- The database has exactly four tables — `chats`, `messages`, `characters`,
`app_settings`. No parent pointer, no branch, no take, no checkpoint.
- **Zero automated tests.** No `test` script, and no test framework in
`devDependencies` at all. Phase 0A listed this as "BUILD/VERIFY"; it is zero.
- **The local provider hard-codes a five-model Gemma 4 allowlist**
(`src/lib/text-models.ts`). Arbitrary local Ollama models are only reachable
through the "custom" OpenAI-compatible path.
- Next.js telemetry is on by default and printed its notice on first start.
- 335 npm packages, 6 high-severity advisories.
**Cost of the branch retrofit — the Experiment B answer**
Larger than Phase 0A estimated, because of a dependency Phase 0A did not identify:
1. `messages` needs `parent_id`, `branch_id`, `depth`, and a live/take flag.
2. Of 20 exported `db.ts` functions, the four history ones need rewriting, and
`getChat` must stop returning a flat list and become a lineage query.
3. `story-prompt.ts` windows history by **array index** (`messages.slice(dropped)`,
`messages.length - 1`). All of it becomes an ancestry walk.
4. **The rolling summary watermark is positional.** `chats.story_summary_count`
means "how many of the chat's oldest messages the summary already covers", and
`route.ts` compares `evicted.length > stored.coveredCount`. On a branch, "the
first N messages" has no meaning. Summaries must be re-anchored to a turn id
and made per-lineage. **Phase 0A did not identify this.**
5. `page.tsx` is a single 3,991-line component holding retry/erase/edit as array
slicing, and there is no take stepper, branch panel, or tree view to extend.
6. All of the above lands with no test suite underneath it.
**Worth reusing as reference:** the reading UI, inline scene images, the
character-portrait-as-reference-image continuity idea, and the scene/media
architecture. None of it ports directly — Open Dungeon is Next.js and AI-DnD is
React + FastAPI.
### ai-adventure — not the base, but promote it to the architectural reference
**What worked — this is the strongest result of the round**
- `Ran 76 tests ... OK` in 1.9s.
- **The Ollama "adapter" is zero code.** `LMStudioBackend` is a generic
OpenAI-compatible client hitting `/v1/chat/completions` and `/v1/models`. I
changed two lines of `world.toml` — `base_url` and model `name` — and the
doctor reported `[PASS] LM Studio reachable: http://127.0.0.1:11434` and
`[PASS] Configured model visible: qwen2.5:3b-instruct`. A real turn then
committed to the database with `status='committed'`, zero validation errors,
and a `model_calls` row recording backend, model and attempt. **Phase 0A's
"MODIFY — LM Studio primary provider" was wrong; it is a config change.**
- **Its history model is exactly what the specification asks for.** Verified
independently, not just by trusting its tests:
- Undo deleted **0 rows**, rolled `has_brass_key` back from `False` to `True`,
and the abandoned turn was still in the database.
- Branching from the restored point produced `gate_open` on the branch and
**did not leak it into the original session**.
- Checkpoint restore rolled state back with the later turn still present.
- Total turns in the database never decreased.
- Schema is the target data model: `turns.parent_turn_id`, `status`,
append-only `state_events`, `state_cache` keyed by `head_turn_id`,
`model_calls` with request/response hashes, `named_checkpoints`, `summaries`
anchored by `through_turn_id`, and lore in FTS5.
- **The cleanest privacy posture by a wide margin.** Exactly one outbound call
site in the entire codebase (`lm_studio.py`'s `urlopen`), no hardcoded remote
host anywhere, and **one** third-party dependency (`pydantic`). It is the only
candidate that already warns when `model.base_url` is not loopback — which is
specification §12, implemented.
**What failed**
- With `qwen2.5:0.5b` the turn failed validation and **nothing was saved**. That
is the model-proposes/app-validates discipline working correctly, but it means
the typed-event contract needs a reasonably capable model; the freeform-prose
candidates tolerate a weak one. `qwen2.5:3b-instruct` succeeded.
- Product distance is real and large: terminal only, worlds authored as
hand-written TOML and Markdown files rather than set up in-app, no streaming,
no embeddings/semantic memory, no media, no import of user documents.
**Adaptation required:** building essentially the entire product on top of a
correct core.
---
## C. Key Technical Findings
| Requirement | AI-DnD | Open Dungeon | ai-adventure |
|---|---|---|---|
| History/undo model | Tree + takes; **undo deletes**, no redo (both fixed in the spike, §G) | Flat list; **retry/edit/erase all destroy** | **Parent-linked, head pointer, undo deletes nothing** |
| State rollback | Per-node snapshots, verified | None | **Event replay, verified** |
| Memory isolation across branches | **Verified correct, with negative control** | N/A (no branches) | **Verified correct** |
| Local Ollama | Default endpoint; verified streaming | Verified via "custom" provider; local provider allowlists 5 Gemma builds | **Verified, zero code change** |
| Offline | **Blocked by tiktoken CDN fetch** until vendored; then fully offline | Runtime clean; build needs network; Next telemetry on by default | **Verifiably clean — 1 outbound call site, 1 dependency** |
| Browser suitability | React SPA + FastAPI, already there | Next.js, already there | **None** |
| Prompt inspection | **Per-action snapshot with labeled sections and token costs** | Not exposed | Prompt hashes; `store_prompts` off by default |
| Imported knowledge | Story cards (keyword-triggered) — closest starting point | None | **Local FTS5 lore, deterministic** |
| Named checkpoints | Absent (branch names only) | Absent | **Present** |
| Future media | Absent | **Present and working** | Absent |
| Tests | **632** | **0** | 76 |
---
## D. Important Surprises
Ordered by how much they should change the plan.
1. **AI-DnD's undo is destructive, and it destroys retry's preserved
alternatives too.** `reports/PRELIMINARY-RECOMMENDATION.md` and
`REUSE-MATRIX.md` both credit AI-DnD with non-destructive rollback. Retry
earns that; undo does not. Measured: undo deleted the player action, the live
take, and the sibling take a prior retry had kept. The follow-up spike has
since replaced it with a head-cursor undo and added Redo (§G), so this is a
corrected planning assumption rather than an outstanding defect.
2. **AI-DnD cannot take a single turn on an air-gapped machine as shipped.**
`tiktoken` fetches `cl100k_base` from a Microsoft CDN on first use. Proven by
running it on a Docker `--internal` network, and proven fixed by seeding the
file.
3. **ai-adventure needs no Ollama adapter at all.** Two config lines. The
"LM Studio" name is a misnomer for a generic OpenAI-compatible client.
4. **Open Dungeon has zero automated tests** — not few, none, and no framework
installed.
5. **Open Dungeon's summary watermark is positional**, which adds a summary
redesign to its branch retrofit that Phase 0A did not cost.
6. **Open Dungeon's "local" provider only offers five Gemma 4 QAT builds.** A
genre-agnostic app that lets the user pick any installed Ollama model needs
that lifted.
7. **AI-DnD fetches Google Fonts at runtime** and its CSP is written to allow it,
violating specification §12's "no remote fonts".
8. **AI-DnD's `netguard` is the opposite of our requirement** — it protects a
hosted deployment from SSRF and is inert locally. ai-adventure has the
loopback warning we actually want.
9. Repo maturity differs sharply: AI-DnD is 172 commits / 3 authors, Open Dungeon
45 / 1, ai-adventure 20 / 1. All three are young. Since we fork, this matters
less than it would for a dependency, but there is no upstream to lean on.
---
## E. Recommendation for Next Step
**Fork AI-DnD and begin the controlled strip-down.** The blocking uncertainty is
gone: the spike this section previously asked for has been run and passed
(section G).
Recommended order, cheapest and most load-bearing first:
1. **Close the two offline blockers.** Vendor the `tiktoken` encoding and
self-host the three font families, then re-run the air-gapped test. Until this
is done the application does not satisfy specification §2, and it is a day's
work.
2. **Strip the hosted surface.** Remove multi-user/auth, the demo key, analytics,
`render.yaml`, the Postgres path, and the quickjs sandbox. Re-instrument the
~19 tests that use JS hooks to observe state, moving them onto the
`world_state_after` snapshots that already exist.
3. **Land the undo/redo work properly**, promoting the spike from a proof to a
feature: fix the four rough edges it left (undo floor at the opening node,
retry/`add_take` while behind the head, marking abandoned turns disposable,
and the frontend Redo control and branch labelling) — **and add `headDepth` to
the export bundle in the same change**, or a round-trip silently undoes every
undo (§H).
4. **Add named checkpoints**, using ai-adventure's
`named_checkpoints(session_id, turn_id, name)` as the model. With an explicit
head that can sit anywhere, a checkpoint is now just a named head position.
5. **Add the loopback-endpoint warning**, modeled on ai-adventure's
`_endpoint_warnings`.
6. **Then** imported knowledge, and only after that, media.
Treat **ai-adventure as the written specification for state and history
semantics** throughout. Its schema is the target shape and its 76 tests are the
behavioral contract worth copying — the spike already reproduced its central
property (undo that moves a head instead of deleting) and benefited from having
that reference to aim at. §H extends this: when generalizing AI-DnD's world state
into genre-neutral narrative state, take ai-adventure's **typed-event vocabulary**
rather than AI-DnD's relative-delta vocabulary, because the delta protocol has a
failure mode that validation cannot catch and that worsens as context grows.
One addition to the ordering above: whatever narrative-state engine we build must
be **exercised at realistic context length against the models users will actually
run**, not only in a clean harness. That is where the referee's weakness showed
up, and it would not have appeared in a unit test.
Do not merge repositories. Open Dungeon remains a UX and media reference only.
## F. Open Questions
**Resolved by the undo/redo spike (§G):**
- ~~Does capping `Path.clause` by head depth break the memory and summary
cursors?~~ **No.** `test_memory_settling`, `test_memory_nodes`,
`test_memory_retrieval`, `test_memory_rewrite` and `test_history_window` all
passed unchanged, and memory retrieval was verified end-to-end against real
embeddings with a control.
- ~~How should named checkpoints sit on top of takes and branches?~~ **Largely
answered.** The spike introduces a head that can sit behind the tip, so a
checkpoint becomes a named `(branch_id, depth)` position. Still needs a
decision on whether restoring one forks immediately or on first write — the
spike chose "on first write" for undo, and consistency argues for the same.
**Resolved by the follow-up checks (§H):**
- ~~What is the model capability floor?~~ **Characterized, not a blocker.** The
delta protocol is within a 3B model's reach in a clean prompt (3/3 correct) but
degraded to absolute values under the full application context. `0.5b` cannot
do it at all. See §H for the design consequence.
- ~~Can `psycopg`/Postgres be dropped cleanly?~~ **Yes.** Three places, no
blockers: one branch in `database.py`, two migration helpers, one import inside
analytics.
- ~~Do story cards become imported knowledge, or is a new table needed?~~
**A new table.** `StoryCard` carries no `branch_id`/`depth`, so it is not
lineage-scoped and inherits neither branch nor undo isolation.
- ~~Does the export format survive the history rework?~~ **No — and this is now a
known defect to fix, not a question.** See §H.
**Still open:**
1. **When does abandoned history get cleaned up?** Nothing marks it disposable
yet, and undone turns now accumulate rather than being deleted. This is the
intended trade, but it needs a retention story.
2. **Is Open Dungeon's image path portable at all?** It is Next.js calling a local
FLUX worker over HTTP. The worker protocol may be reusable even though none of
the UI is.
3. **What is the real model recommendation for users?** The probe establishes a
floor and a failure mode, not a ranking. A proper capability matrix across the
models a user would actually run belongs in the build phase, measured at
realistic context length.
4. **Not tested in any round:** import/export round-trip *behavior* beyond the
head-depth defect found by reading, multi-hour long-story context behavior, or
any concurrency beyond the single-turn lock.
## G. Follow-Up Spike: Non-Destructive Undo/Redo — Passed
Run after the main round, on a disposable copy of AI-DnD `d72f7c1b`. Full detail
in `reports/PHASE-0B-UNDO-SPIKE.md`.
**The change is three files, +130 / -31 lines**, roughly a fifth of it comments:
| File | Change | What |
|---|---|---|
| `context/lineage.py` | +26 / -3 | Read an uncapped lineage entry as capped at `self.tip` (the head), in `clause` and `contains`; add `uncapped()`. |
| `routers/adventures/takes.py` | +74 / -27 | `undo_turn` moves `head_depth` instead of deleting; new `redo_turn`. |
| `routers/adventures/turns.py` | +30 / -1 | `fork_if_behind_head` calls the existing `tree.branch_at` when a write lands below the head. |
It is this small because the pieces already existed: every read funnels through
`lineage.path_of()`, `adventures.head_depth` was already a stored head, and
`tree.branch_at` already created a branch that leaves a path at a depth without
touching the line it leaves.
**Results against the success criteria:**
| Criterion | Result |
|---|---|
| Undo deletes zero rows | **Pass.** 6 → 6 rows across three undos; separately, 11 consecutive undos on a 23-action story with 23 actions still present. Clears the "at least five, unlimited preferred" requirement. |
| Abandoned tail stays reachable as a branch | **Pass.** Undo-then-write forked branch 2 at fork depth 3; branch 1 kept all six of its rows, live and readable, and the branch switcher reaches it. |
| Redo restores | **Pass.** Three redos returned the head from -1 to 5 and the transcript to all six actions, exactly. |
| Branch-scoped memory isolation holds | **Pass, with control.** An embedded memory at depth 5 was retrieved at the tip (similarity 0.689), returned `used: []` once the head moved to depth 3, and became eligible again on redo — without being deleted. None of seven terms unique to the hidden turns appeared in the assembled prompt. |
| Suite stays green apart from delete-behavior tests | **Pass.** `627 passed, 5 failed` (0.8%). All five assert deleted rows or pruned memories. The state assertions inside those same tests still pass — `assert adv.script_state == {"gold": 0}` succeeds in both `test_state_revert` failures, and the fork-boundary guard in `test_undo_stops_at_the_fork` is untouched. |
**Rough edges left for the real implementation** (none architectural): undo can
walk past a user-written opening to an empty transcript; retry and `add_take`
were not routed through the fork check; abandoned turns are not yet marked
disposable; and the frontend has no Redo control.
**What this changes about the recommendation:** nothing about the choice, and the
one thing about the plan — the strip-down can start now rather than after another
investigation.
## H. Follow-Up Checks — Referee, Export, Cards, Postgres
Run after the spike. Detail in `reports/PHASE-0B-FOLLOWUP-CHECKS.md`.
### The finding that changes what we build
AI-DnD's world-state referee asks the model for **relative deltas**
(`new = old + delta`, then clamped). Playing the seeded RPG scenario with
`qwen2.5:3b-instruct`, the model sent **absolute values** — `{"player.hp": 92}`
meaning "hp is now 92". The engine read `+92`, clamped to `+35`, then to the
ceiling of 100:
```
player.hp 100 -> 100 "did not move. It is already at its maximum of 100"
```
**The player took an arrow to the shoulder and finished the turn at full health.**
The referee did its job — nothing was corrupted, and it fed corrective text back.
But no validator can catch this, because `+92` is a legal proposal. The
application stayed authoritative while the authoritative state silently stopped
tracking the narration. That is the exact divergence specification §2 exists to
prevent.
An isolated probe using AI-DnD's own `EMIT_RULE` and parser shows the protocol is
**not** beyond a small model — `qwen2.5:3b-instruct` and `qwen2.5:7b-instruct`
both got 3/3 correct relative deltas, and `qwen2.5:0.5b` emitted no block at all.
The difference is context load: in a short prompt the 3B model follows the rule,
and under the full application prompt it drifts to absolutes.
**Consequence for the design:** when we generalize AI-DnD's world state into
genre-neutral narrative state, take **ai-adventure's typed-event vocabulary**
(`set_flag key=… value=…` — explicit and absolute) rather than AI-DnD's delta
vocabulary. The delta protocol has a failure mode that is invisible to validation
and worsens with context length; the event protocol cannot express the mistake.
This is the second time in this round that ai-adventure's core has turned out to
be the right specification to build to.
### One new defect
**Export/import silently redoes an undone story.** `bundle.py:_point_the_head`
recomputes `head_depth = max(depths)` on import, and the bundle format carries
`headBranch` but not the head depth. That was correct when undo deleted rows; with
the spike's non-destructive undo, a round-trip restores every abandoned turn. Fix
is one optional `headDepth` field plus honouring it on import, falling back to
`max(depths)` so existing `v2` files still load. It must land with the undo work,
not after it.
### Two smaller answers
- **Story cards are not lineage-scoped.** `StoryCard` has no `branch_id` or
`depth`, unlike `Memory`. Harmless today because cards are authored rather than
derived — but any imported-knowledge source that can be created from story
content must carry those coordinates and be read through `lineage.Path.clause`,
or it inherits neither branch nor undo isolation. Argues for a new table rather
than extending `story_cards`.
- **Postgres is cleanly removable.** One branch in `database.py`, two migration
helpers, one import inside analytics. `psycopg[binary]` then drops out.
## Supporting Evidence
- `reports/PHASE-0B-BASELINE.md` — installs, test runs, ports, storage, versions.
- `reports/PHASE-0B-AI-DND-EXPERIMENT.md` — branch/undo/retry measurements,
memory isolation with negative control, scripting removal.
- `reports/PHASE-0B-OPEN-DUNGEON-HISTORY.md` — destructive-history proof and
retrofit cost map.
- `reports/PHASE-0B-AI-ADVENTURE-OLLAMA.md` — zero-change Ollama run and the
fixture checks.
- `reports/PHASE-0B-OFFLINE-NETWORK.md` — offline runs and egress inventory.
- `reports/PHASE-0B-UNDO-SPIKE.md` — the non-destructive undo/redo spike: the
diff, the measurements, and the rough edges it left.
- `reports/PHASE-0B-FOLLOWUP-CHECKS.md` — the world-state referee under local
models, the export/import head defect, story-card lineage safety, Postgres.
Working tree, scripts and databases from both rounds are under `phase0b/`
(untracked scratch, safe to delete). The spike itself is `phase0b/spike/`, a
disposable copy of the AI-DnD backend — it is a proof, not a branch to merge.