7.8 KiB
Phase 0B — Codex Initial Validation Brief
Status: Ready for execution
Purpose: Run a focused first round of local validation on the three finalist repositories and return a recommendation based on what was actually learned.
1. Goal
We are not asking you to build the production application yet.
The goal of this round is to answer one question:
Which existing project is the best starting point for the local interactive-story application, and what important technical facts did we learn that should affect the next design step?
The three finalists are:
-
Open Dungeon
https://github.com/newideas99/open-dungeon -
ai-adventure
https://github.com/CaoRuiming/ai-adventure
2. Important Product Requirements
Use these as the main evaluation criteria.
The eventual application should be:
- browser-first,
- local-only in v1,
- based on local Ollama inference,
- single-user,
- genre-agnostic,
- persistent across restarts,
- able to retain authoritative story state separately from model prose,
- able to Undo/Redo/Retry safely,
- able to preserve named save points/checkpoints,
- able to retain abandoned history without immediately deleting it,
- able to prevent abandoned-history facts/memories from leaking into the active story,
- able to support long-term context/memory,
- able to import local knowledge,
- able to inspect what context was sent to the model,
- architecturally compatible with future local image/video/TTS/STT support.
Do not try to implement all of these now.
This round is about determining which candidate already gives us the strongest foundation.
3. Read Only What You Need
Start with:
README.mdSPECIFICATION.mdreports/PRELIMINARY-RECOMMENDATION.mdreports/REUSE-MATRIX.md
Then use these only when relevant to a specific experiment:
STORY-BRANCH-SEMANTICS.mdCONTEXT-AND-MEMORY.mdSECURITY-THREAT-MODEL.mdTEST-CAMPAIGN-FIXTURE.md
Do not read every planning document up front unless needed.
4. Baseline Work for Each Candidate
For each repository:
- Clone it cleanly.
- Record the exact commit SHA.
- Follow the documented install instructions.
- Run the existing tests.
- Build/start the application.
- Confirm the basic local story flow.
- Record:
- test results,
- storage/database technology,
- model/provider assumptions,
- local ports,
- major runtime failures,
- obvious cloud/hosted dependencies.
Do not spend excessive time fixing unrelated upstream problems.
If a project does not run cleanly, document why and continue.
5. Focused Experiment A — AI-DnD
We want to know whether AI-DnD can realistically serve as the production base.
Test:
- Can it run with Ollama locally?
- Can its story-tree/history system support simple user-facing Undo/Redo/Retry behavior?
- Does state rollback work independently of heavy RPG/stat mechanics?
- Can RPG-specific state be left empty/minimal without breaking the useful history/state architecture?
- Can QuickJS/scripting and hosted/cloud-oriented features be disabled without breaking local story operation?
- Can local memory/embedding behavior work without cloud services?
- Do memories/state respect the active history path?
- Is its context/Insights system useful for showing what was sent to the model?
Use a small disposable experiment if necessary.
Do not start stripping the whole application down.
6. Focused Experiment B — Open Dungeon
We want to know how expensive it would be to fix its history model.
Test:
- Confirm how Retry/Edit/Erase affect stored history.
- Identify whether old future turns are deleted.
- Trace which parts of the application depend on that linear/destructive behavior.
- Estimate how invasive it would be to change to:
- parent-linked turns,
- active head,
- retained abandoned history,
- named checkpoints,
- lineage-safe summaries/state.
Do not implement the complete branch system.
Also record useful existing pieces:
- browser UX,
- local image generation,
- character visual continuity,
- any scene/media architecture worth reusing.
7. Focused Experiment C — ai-adventure
We want to know whether its strong state/privacy architecture can realistically become a browser-based Ollama application.
Test:
- Run the existing tests.
- Confirm Undo/branch/checkpoint/replay behavior.
- Identify the provider abstraction.
- Prove one local Ollama-backed story turn using the smallest practical adapter.
- Determine how tightly the core application logic is coupled to the CLI.
- Assess whether the core could sit behind a browser/API layer without moving authoritative state logic.
- Review its local lore/FTS approach for possible reuse.
Do not build a browser frontend.
8. Offline / Privacy Check
For each candidate, once dependencies/models are installed:
- run it with outbound Internet unavailable or blocked where practical,
- exercise basic story generation,
- note any unexpected network attempts.
We do not need a full penetration test in this round.
We do need to know:
- whether local story use truly works offline,
- whether cloud services are required,
- whether analytics/telemetry/remote assets are present,
- how difficult those paths would be to remove.
9. Use the Standard Fixture Selectively
Use TEST-CAMPAIGN-FIXTURE.md where it helps answer continuity questions.
You do not need to execute the entire fixture against every candidate.
The most important checks are:
- possession/state consistency,
- restore/undo behavior,
- abandoned-path isolation,
- whether an old discarded fact can leak into current memory/context.
10. What Not to Do
Do not:
- build the production fork,
- merge repositories,
- redesign the full UI,
- implement full RAG,
- implement complete branching in Open Dungeon,
- remove all RPG code from AI-DnD,
- build a browser frontend for ai-adventure,
- add image/video/TTS/STT features,
- write the final production milestone plan.
Small disposable code changes are allowed only when needed to answer the evaluation questions.
11. Final Deliverable
The main output from this round should be a single recommendation document:
PHASE-0B-RECOMMENDATION.md
It should summarize what was learned, not just list test logs.
Include:
A. Executive Recommendation
- Which repository should be the production base?
- Confidence level: high / medium / low.
- Did the initial AI-DnD recommendation hold up?
B. What We Learned About Each Candidate
For each:
- what worked,
- what failed,
- strongest reusable pieces,
- major architectural problems,
- likely amount/type of adaptation required.
C. Key Technical Findings
Especially:
- history/undo model,
- state rollback,
- memory isolation,
- local Ollama support,
- offline/privacy behavior,
- browser suitability,
- imported-knowledge potential,
- future media extension potential.
D. Important Surprises
Anything that contradicts the current planning assumptions.
E. Recommendation for Next Step
Do not perform the next step.
Instead recommend what should happen next, such as:
- fork AI-DnD and begin a controlled strip-down,
- perform one additional experiment first,
- reconsider Open Dungeon,
- use ai-adventure as the base instead,
- revise one of the product assumptions.
F. Open Questions
List anything that could not be resolved in this round.
12. Supporting Evidence
You may also create concise supporting notes/logs for:
- baseline results,
- AI-DnD experiment,
- Open Dungeon history analysis,
- ai-adventure Ollama adapter,
- offline/network observations.
Keep them concise.
The recommendation document is the primary deliverable.
13. Stop Condition
When PHASE-0B-RECOMMENDATION.md is complete, stop.
We will take the findings back into the design discussion, re-examine the assumptions, and decide the next step before any production implementation begins.