10 KiB
Phase 0B — Codex Local Validation Handoff
Status: COMPLETE / HISTORICAL — do not execute as a production prompt
Purpose: Validate the Phase 0A recommendation using local builds, tests, offline runtime observation, and tightly scoped experiments.
Stop rule: Do not begin production implementation.
1. Read Before Starting
Read the package in the order listed in README.md.
At minimum, before modifying any finalist, read:
SPECIFICATION.mdDATA-MODEL.mdSTORY-BRANCH-SEMANTICS.mdCONTEXT-AND-MEMORY.mdIMPORTED-KNOWLEDGE-DESIGN.mdSECURITY-THREAT-MODEL.mdMEDIA-EXTENSION-CONTRACT.mdBROWSER-UX-SPEC.mdTEST-CAMPAIGN-FIXTURE.mdV1-ACCEPTANCE-TESTS.mdreports/PRELIMINARY-RECOMMENDATION.mdreports/REUSE-MATRIX.md
Treat the detailed behavioral documents and acceptance tests as the target behavior. Treat TECHNICAL-DESIGN.md as provisional.
2. Finalists to Clone
Clone only these three primary finalists for Phase 0B:
- https://github.com/parththakkar106/AI-DnD
- https://github.com/newideas99/open-dungeon
- https://github.com/CaoRuiming/ai-adventure
At clone time record exact commit SHA, branch/tag, date, license, dependency lockfiles, required runtimes, and documented local model/provider assumptions.
Keep each upstream clone clean. Use separate experiment branches/worktrees for disposable changes. Do not merge candidate repositories together.
3. Result Codes
Use consistently:
PASS
PARTIAL
FAIL
NOT IMPLEMENTED
NOT APPLICABLE
Do not convert an untested requirement into a PASS.
4. Validation V0 — Environment and Baseline
For all three:
- install from documented instructions,
- run existing test suite,
- run build/lint/typecheck where applicable,
- record failures,
- record actual current test count,
- record local data paths,
- record listening ports,
- record child processes/services,
- record model/provider configuration,
- record database/storage technology.
Deliver one baseline report per project. Do not rely on README claims for test counts or feature behavior.
5. Validation V1 — Offline and Network Behavior
After dependencies and local models are already installed, block outbound Internet and exercise launch, story creation, 5+ turns, restart/resume, summaries, memory/embeddings if present, Retry, Undo/rewind, checkpoint/branch features if present, .txt/.md import if present, and Open Dungeon local image generation if configured.
Capture open sockets, DNS attempts, HTTP(S)/WebSocket destinations, and which feature caused each request.
Use SECURITY-THREAT-MODEL.md and acceptance groups A, G, and H.
Pass condition for target v1 operation:
Story content, imported knowledge, prompt/context data, and media prompts do not leave loopback or explicitly approved local endpoints.
6. Standard Fixture Use
Use TEST-CAMPAIGN-FIXTURE.md as the standard narrative test bed.
Where a finalist cannot represent the fixture directly, map it as closely as possible, document the mismatch, and do not silently change expected truth/state to suit the candidate.
Important checks:
- Mara's knowledge boundaries,
- Silver Key ownership,
- resurrection canon,
- Canon vs Reference vs Inspiration authority,
- Path A secret followed by restore/divergence into Path B,
- abandoned-history memory isolation,
- long-term memory plant,
- checkpoint persistence,
- science-fiction variant.
7. Acceptance-Test Mapping
Use V1-ACCEPTANCE-TESTS.md as the common comparison contract. Produce a gap matrix rather than forcing each candidate to fully pass v1.
Use these interpretations:
Already passes
Passes with configuration
Small adaptation
Foundational redesign
Not present
Prioritize high-risk groups:
- A01-A05 — local operation/persistence
- C01-C05 — canon/state
- D01-D14 — Undo/Redo/Retry/checkpoints
- E01-E04 — lineage safety
- F01-F08 — memory/context
- G01-G10 — imported knowledge where supported
- H01-H10 — security/privacy
- J01-J03 — genre independence
- K01-K04 — media readiness
- L01-L04 — data integrity
Long-run M01-M04 need not be fully executed against every candidate if disproportionate; identify production risk and existing test coverage instead.
8. Experiment V2 — AI-DnD Strip-Down Feasibility
Do not redesign the application.
Answer:
- Can a scenario run with RPG stats absent, empty, or minimal?
- Do branch/retry/undo/tree tests operate independently of RPG mechanics?
- Can user-facing tree complexity be hidden behind
STORY-BRANCH-SEMANTICS.md? - Disable QuickJS scripting. What breaks?
- Disable/remove hosted multi-user/auth/demo/analytics paths. What breaks locally?
- Configure only local Ollama generation.
- Configure only local embeddings, preferably Ollama/local.
- Verify branch switching restores correct generic state.
- Verify memory retrieval respects active lineage.
- Test whether abandoned Path A facts leak into Path B.
- Map Story Cards/world-info to Canon / Reference / Inspiration.
- Determine whether prompt/context snapshots satisfy Context Inspector requirements.
- Determine whether visual/scene snapshot data can be added without RPG coupling.
- Inventory code coupled to RPG worldstate, scripting, hosted auth, analytics, remote providers, and AI Dungeon compatibility.
Estimate invasiveness by affected files/modules, not hours. Do not merge the experiment.
9. Experiment V3 — Open Dungeon Branch Retrofit Impact
Do not implement full branching.
Trace message CRUD, Retry, Erase, Edit, Continue, summary generation, state/character persistence, image association, and visual continuity. Confirm destructive-tail assumptions.
Design a minimal hypothetical persistence change supporting:
turn/node ID
parent turn ID
active head
alternate narrator takes
retained disposable history
checkpoint pointer
lineage-safe summaries/memories
Use the standard fixture to reason through Undo/restore, Path A -> Path B divergence, stale summary/state risks, and image attachment after divergence.
Also inspect local image/provider patterns for reuse. Measure invasiveness; do not build the branch system.
10. Experiment V4 — ai-adventure Ollama / Service Boundary
- Run existing tests unchanged.
- Identify provider interface.
- Prove one Ollama-backed turn using the smallest disposable adapter possible.
- Identify modules that know about the CLI.
- Determine whether the app/state layer can be wrapped by a browser/API service without moving authoritative logic.
- Verify undo, branch, checkpoint, restore, replay.
- Evaluate event/commit discipline for reuse.
- Evaluate its FTS/lore system against
IMPORTED-KNOWLEDGE-DESIGN.md. - Determine difficulty of adding semantic local retrieval while preserving lexical retrieval.
- Check privacy boundary with the Ollama adapter.
Do not build a browser UI.
11. Validation V5 — Test Quality
For each finalist report actual test count, categories, branch/rollback coverage, state reconstruction, migrations, summary/memory coverage, provider mocks, offline/network tests, browser tests, security tests, flaky/failing tests, and tests requiring Internet.
Highlight which high-risk acceptance requirements already have regression coverage.
12. Validation V6 — Imported Knowledge Gap Analysis
Against IMPORTED-KNOWLEDGE-DESIGN.md, report local .txt/.md ingestion, classifications, campaign isolation, provenance, lexical/semantic search, embedding provider, enable/disable, deletion, export/import, hidden canon, prompt-injection framing, and remote URL/image behavior.
Do not implement a full new RAG subsystem during Phase 0B.
13. Validation V7 — Browser UX Gap Analysis
Against BROWSER-UX-SPEC.md, report story reading/input quality, streaming, Undo/Redo/Retry UI, alternate-take selection, edit behavior, Save Points, state inspection, knowledge management, prompt/context inspection, local-model status, and advanced complexity exposed to the user.
Explicitly identify AI-DnD components worth retaining and Open Dungeon components worth borrowing/reimplementing. Do not redesign the frontend.
14. Validation V8 — Future Media and Speech Readiness
Against MEDIA-EXTENSION-CONTRACT.md, determine whether the architecture can support future local image generation, video generation, audio/ambience, text-to-speech, and speech-to-text.
For STT, verify the architecture can support:
local microphone/audio
->
local STT provider
->
editable draft text
->
normal user submission
STT output must not bypass the normal story commit path.
Do not implement STT/TTS/video during Phase 0B. Open Dungeon local image behavior may be exercised because it already exists.
15. Final Acceptance Gap Matrix
Produce a matrix organized by acceptance-test group covering A, C, D, E, F, G, H, I, J, K, L, plus UX fit. Include production impact for each gap.
16. Final Decision Matrix
Return:
| Question | AI-DnD | Open Dungeon | ai-adventure |
|---|---|---|---|
| Baseline builds | |||
| Existing tests pass | |||
| Runs offline after setup | |||
| Ollama works | |||
| History semantics fit | |||
| State authority fits | |||
| Memory/lineage fits | |||
| Imported knowledge fit | |||
| Prompt inspection fit | |||
| Security/local-only hardening | |||
| Unwanted-code removal scope | |||
| Browser UX fit | |||
| Media extension fit | |||
| Future TTS/STT fit | |||
| Major blockers |
17. Recommendation Report
The final recommendation should answer:
- Which single repository should be the production base?
- Why?
- What are the top architectural risks?
- What must be removed?
- What must be generalized?
- Which concepts/components should be reimplemented from other candidates?
- Does Phase 0B change the preliminary AI-DnD recommendation?
- Which open questions remain before
TECHNICAL-DESIGN.mdv1.0? - Are any v1 requirements likely to need reconsideration because of real technical constraints?
- Is unlimited Undo straightforward? If not, what practical limit exists and why?
Use evidence, not repository popularity or feature count.
18. Stop Condition
Stop after baseline reports, offline/network evidence, three scoped experiments, test-quality report, acceptance-gap matrix, decision matrix, and final recommendation.
Do not start the production fork conversion, implement the complete branch system, build the final browser UI, implement full RAG, add video/TTS/STT, rewrite the production technical design, or write production milestones.
Return reports and experiment diffs/results for review. The fork/architecture decision will be made from those results.