# Phase 0B — Codex Local Validation Handoff **Status:** COMPLETE / HISTORICAL — do not execute as a production prompt **Purpose:** Validate the Phase 0A recommendation using local builds, tests, offline runtime observation, and tightly scoped experiments. **Stop rule:** Do not begin production implementation. ## 1. Read Before Starting Read the package in the order listed in `README.md`. At minimum, before modifying any finalist, read: 1. `SPECIFICATION.md` 2. `DATA-MODEL.md` 3. `STORY-BRANCH-SEMANTICS.md` 4. `CONTEXT-AND-MEMORY.md` 5. `IMPORTED-KNOWLEDGE-DESIGN.md` 6. `SECURITY-THREAT-MODEL.md` 7. `MEDIA-EXTENSION-CONTRACT.md` 8. `BROWSER-UX-SPEC.md` 9. `TEST-CAMPAIGN-FIXTURE.md` 10. `V1-ACCEPTANCE-TESTS.md` 11. `reports/PRELIMINARY-RECOMMENDATION.md` 12. `reports/REUSE-MATRIX.md` Treat the detailed behavioral documents and acceptance tests as the target behavior. Treat `TECHNICAL-DESIGN.md` as provisional. ## 2. Finalists to Clone Clone only these three primary finalists for Phase 0B: 1. https://github.com/parththakkar106/AI-DnD 2. https://github.com/newideas99/open-dungeon 3. https://github.com/CaoRuiming/ai-adventure At clone time record exact commit SHA, branch/tag, date, license, dependency lockfiles, required runtimes, and documented local model/provider assumptions. Keep each upstream clone clean. Use separate experiment branches/worktrees for disposable changes. Do not merge candidate repositories together. ## 3. Result Codes Use consistently: ```text PASS PARTIAL FAIL NOT IMPLEMENTED NOT APPLICABLE ``` Do not convert an untested requirement into a PASS. ## 4. Validation V0 — Environment and Baseline For all three: - install from documented instructions, - run existing test suite, - run build/lint/typecheck where applicable, - record failures, - record actual current test count, - record local data paths, - record listening ports, - record child processes/services, - record model/provider configuration, - record database/storage technology. Deliver one baseline report per project. Do not rely on README claims for test counts or feature behavior. ## 5. Validation V1 — Offline and Network Behavior After dependencies and local models are already installed, block outbound Internet and exercise launch, story creation, 5+ turns, restart/resume, summaries, memory/embeddings if present, Retry, Undo/rewind, checkpoint/branch features if present, `.txt`/`.md` import if present, and Open Dungeon local image generation if configured. Capture open sockets, DNS attempts, HTTP(S)/WebSocket destinations, and which feature caused each request. Use `SECURITY-THREAT-MODEL.md` and acceptance groups A, G, and H. Pass condition for target v1 operation: > Story content, imported knowledge, prompt/context data, and media prompts do not leave loopback or explicitly approved local endpoints. ## 6. Standard Fixture Use Use `TEST-CAMPAIGN-FIXTURE.md` as the standard narrative test bed. Where a finalist cannot represent the fixture directly, map it as closely as possible, document the mismatch, and do not silently change expected truth/state to suit the candidate. Important checks: - Mara's knowledge boundaries, - Silver Key ownership, - resurrection canon, - Canon vs Reference vs Inspiration authority, - Path A secret followed by restore/divergence into Path B, - abandoned-history memory isolation, - long-term memory plant, - checkpoint persistence, - science-fiction variant. ## 7. Acceptance-Test Mapping Use `V1-ACCEPTANCE-TESTS.md` as the common comparison contract. Produce a gap matrix rather than forcing each candidate to fully pass v1. Use these interpretations: ```text Already passes Passes with configuration Small adaptation Foundational redesign Not present ``` Prioritize high-risk groups: - A01-A05 — local operation/persistence - C01-C05 — canon/state - D01-D14 — Undo/Redo/Retry/checkpoints - E01-E04 — lineage safety - F01-F08 — memory/context - G01-G10 — imported knowledge where supported - H01-H10 — security/privacy - J01-J03 — genre independence - K01-K04 — media readiness - L01-L04 — data integrity Long-run M01-M04 need not be fully executed against every candidate if disproportionate; identify production risk and existing test coverage instead. ## 8. Experiment V2 — AI-DnD Strip-Down Feasibility Do not redesign the application. Answer: 1. Can a scenario run with RPG stats absent, empty, or minimal? 2. Do branch/retry/undo/tree tests operate independently of RPG mechanics? 3. Can user-facing tree complexity be hidden behind `STORY-BRANCH-SEMANTICS.md`? 4. Disable QuickJS scripting. What breaks? 5. Disable/remove hosted multi-user/auth/demo/analytics paths. What breaks locally? 6. Configure only local Ollama generation. 7. Configure only local embeddings, preferably Ollama/local. 8. Verify branch switching restores correct generic state. 9. Verify memory retrieval respects active lineage. 10. Test whether abandoned Path A facts leak into Path B. 11. Map Story Cards/world-info to Canon / Reference / Inspiration. 12. Determine whether prompt/context snapshots satisfy Context Inspector requirements. 13. Determine whether visual/scene snapshot data can be added without RPG coupling. 14. Inventory code coupled to RPG worldstate, scripting, hosted auth, analytics, remote providers, and AI Dungeon compatibility. Estimate invasiveness by affected files/modules, not hours. Do not merge the experiment. ## 9. Experiment V3 — Open Dungeon Branch Retrofit Impact Do not implement full branching. Trace message CRUD, Retry, Erase, Edit, Continue, summary generation, state/character persistence, image association, and visual continuity. Confirm destructive-tail assumptions. Design a minimal hypothetical persistence change supporting: ```text turn/node ID parent turn ID active head alternate narrator takes retained disposable history checkpoint pointer lineage-safe summaries/memories ``` Use the standard fixture to reason through Undo/restore, Path A -> Path B divergence, stale summary/state risks, and image attachment after divergence. Also inspect local image/provider patterns for reuse. Measure invasiveness; do not build the branch system. ## 10. Experiment V4 — ai-adventure Ollama / Service Boundary 1. Run existing tests unchanged. 2. Identify provider interface. 3. Prove one Ollama-backed turn using the smallest disposable adapter possible. 4. Identify modules that know about the CLI. 5. Determine whether the app/state layer can be wrapped by a browser/API service without moving authoritative logic. 6. Verify undo, branch, checkpoint, restore, replay. 7. Evaluate event/commit discipline for reuse. 8. Evaluate its FTS/lore system against `IMPORTED-KNOWLEDGE-DESIGN.md`. 9. Determine difficulty of adding semantic local retrieval while preserving lexical retrieval. 10. Check privacy boundary with the Ollama adapter. Do not build a browser UI. ## 11. Validation V5 — Test Quality For each finalist report actual test count, categories, branch/rollback coverage, state reconstruction, migrations, summary/memory coverage, provider mocks, offline/network tests, browser tests, security tests, flaky/failing tests, and tests requiring Internet. Highlight which high-risk acceptance requirements already have regression coverage. ## 12. Validation V6 — Imported Knowledge Gap Analysis Against `IMPORTED-KNOWLEDGE-DESIGN.md`, report local `.txt`/`.md` ingestion, classifications, campaign isolation, provenance, lexical/semantic search, embedding provider, enable/disable, deletion, export/import, hidden canon, prompt-injection framing, and remote URL/image behavior. Do not implement a full new RAG subsystem during Phase 0B. ## 13. Validation V7 — Browser UX Gap Analysis Against `BROWSER-UX-SPEC.md`, report story reading/input quality, streaming, Undo/Redo/Retry UI, alternate-take selection, edit behavior, Save Points, state inspection, knowledge management, prompt/context inspection, local-model status, and advanced complexity exposed to the user. Explicitly identify AI-DnD components worth retaining and Open Dungeon components worth borrowing/reimplementing. Do not redesign the frontend. ## 14. Validation V8 — Future Media and Speech Readiness Against `MEDIA-EXTENSION-CONTRACT.md`, determine whether the architecture can support future local image generation, video generation, audio/ambience, text-to-speech, and speech-to-text. For STT, verify the architecture can support: ```text local microphone/audio -> local STT provider -> editable draft text -> normal user submission ``` STT output must not bypass the normal story commit path. Do not implement STT/TTS/video during Phase 0B. Open Dungeon local image behavior may be exercised because it already exists. ## 15. Final Acceptance Gap Matrix Produce a matrix organized by acceptance-test group covering A, C, D, E, F, G, H, I, J, K, L, plus UX fit. Include production impact for each gap. ## 16. Final Decision Matrix Return: | Question | AI-DnD | Open Dungeon | ai-adventure | |---|---|---|---| | Baseline builds | | | | | Existing tests pass | | | | | Runs offline after setup | | | | | Ollama works | | | | | History semantics fit | | | | | State authority fits | | | | | Memory/lineage fits | | | | | Imported knowledge fit | | | | | Prompt inspection fit | | | | | Security/local-only hardening | | | | | Unwanted-code removal scope | | | | | Browser UX fit | | | | | Media extension fit | | | | | Future TTS/STT fit | | | | | Major blockers | | | | ## 17. Recommendation Report The final recommendation should answer: 1. Which single repository should be the production base? 2. Why? 3. What are the top architectural risks? 4. What must be removed? 5. What must be generalized? 6. Which concepts/components should be reimplemented from other candidates? 7. Does Phase 0B change the preliminary AI-DnD recommendation? 8. Which open questions remain before `TECHNICAL-DESIGN.md` v1.0? 9. Are any v1 requirements likely to need reconsideration because of real technical constraints? 10. Is unlimited Undo straightforward? If not, what practical limit exists and why? Use evidence, not repository popularity or feature count. ## 18. Stop Condition Stop after baseline reports, offline/network evidence, three scoped experiments, test-quality report, acceptance-gap matrix, decision matrix, and final recommendation. Do not start the production fork conversion, implement the complete branch system, build the final browser UI, implement full RAG, add video/TTS/STT, rewrite the production technical design, or write production milestones. Return reports and experiment diffs/results for review. The fork/architecture decision will be made from those results.