Files
interactive-story/planning/PHASE-0B-CODEX-BRIEF.md
T

7.8 KiB

Phase 0B — Codex Initial Validation Brief

Status: Ready for execution
Purpose: Run a focused first round of local validation on the three finalist repositories and return a recommendation based on what was actually learned.

1. Goal

We are not asking you to build the production application yet.

The goal of this round is to answer one question:

Which existing project is the best starting point for the local interactive-story application, and what important technical facts did we learn that should affect the next design step?

The three finalists are:

  1. AI-DnD
    https://github.com/parththakkar106/AI-DnD

  2. Open Dungeon
    https://github.com/newideas99/open-dungeon

  3. ai-adventure
    https://github.com/CaoRuiming/ai-adventure

2. Important Product Requirements

Use these as the main evaluation criteria.

The eventual application should be:

  • browser-first,
  • local-only in v1,
  • based on local Ollama inference,
  • single-user,
  • genre-agnostic,
  • persistent across restarts,
  • able to retain authoritative story state separately from model prose,
  • able to Undo/Redo/Retry safely,
  • able to preserve named save points/checkpoints,
  • able to retain abandoned history without immediately deleting it,
  • able to prevent abandoned-history facts/memories from leaking into the active story,
  • able to support long-term context/memory,
  • able to import local knowledge,
  • able to inspect what context was sent to the model,
  • architecturally compatible with future local image/video/TTS/STT support.

Do not try to implement all of these now.

This round is about determining which candidate already gives us the strongest foundation.

3. Read Only What You Need

Start with:

  1. README.md
  2. SPECIFICATION.md
  3. reports/PRELIMINARY-RECOMMENDATION.md
  4. reports/REUSE-MATRIX.md

Then use these only when relevant to a specific experiment:

  • STORY-BRANCH-SEMANTICS.md
  • CONTEXT-AND-MEMORY.md
  • SECURITY-THREAT-MODEL.md
  • TEST-CAMPAIGN-FIXTURE.md

Do not read every planning document up front unless needed.

4. Baseline Work for Each Candidate

For each repository:

  1. Clone it cleanly.
  2. Record the exact commit SHA.
  3. Follow the documented install instructions.
  4. Run the existing tests.
  5. Build/start the application.
  6. Confirm the basic local story flow.
  7. Record:
    • test results,
    • storage/database technology,
    • model/provider assumptions,
    • local ports,
    • major runtime failures,
    • obvious cloud/hosted dependencies.

Do not spend excessive time fixing unrelated upstream problems.

If a project does not run cleanly, document why and continue.

5. Focused Experiment A — AI-DnD

We want to know whether AI-DnD can realistically serve as the production base.

Test:

  • Can it run with Ollama locally?
  • Can its story-tree/history system support simple user-facing Undo/Redo/Retry behavior?
  • Does state rollback work independently of heavy RPG/stat mechanics?
  • Can RPG-specific state be left empty/minimal without breaking the useful history/state architecture?
  • Can QuickJS/scripting and hosted/cloud-oriented features be disabled without breaking local story operation?
  • Can local memory/embedding behavior work without cloud services?
  • Do memories/state respect the active history path?
  • Is its context/Insights system useful for showing what was sent to the model?

Use a small disposable experiment if necessary.

Do not start stripping the whole application down.

6. Focused Experiment B — Open Dungeon

We want to know how expensive it would be to fix its history model.

Test:

  • Confirm how Retry/Edit/Erase affect stored history.
  • Identify whether old future turns are deleted.
  • Trace which parts of the application depend on that linear/destructive behavior.
  • Estimate how invasive it would be to change to:
    • parent-linked turns,
    • active head,
    • retained abandoned history,
    • named checkpoints,
    • lineage-safe summaries/state.

Do not implement the complete branch system.

Also record useful existing pieces:

  • browser UX,
  • local image generation,
  • character visual continuity,
  • any scene/media architecture worth reusing.

7. Focused Experiment C — ai-adventure

We want to know whether its strong state/privacy architecture can realistically become a browser-based Ollama application.

Test:

  • Run the existing tests.
  • Confirm Undo/branch/checkpoint/replay behavior.
  • Identify the provider abstraction.
  • Prove one local Ollama-backed story turn using the smallest practical adapter.
  • Determine how tightly the core application logic is coupled to the CLI.
  • Assess whether the core could sit behind a browser/API layer without moving authoritative state logic.
  • Review its local lore/FTS approach for possible reuse.

Do not build a browser frontend.

8. Offline / Privacy Check

For each candidate, once dependencies/models are installed:

  • run it with outbound Internet unavailable or blocked where practical,
  • exercise basic story generation,
  • note any unexpected network attempts.

We do not need a full penetration test in this round.

We do need to know:

  • whether local story use truly works offline,
  • whether cloud services are required,
  • whether analytics/telemetry/remote assets are present,
  • how difficult those paths would be to remove.

9. Use the Standard Fixture Selectively

Use TEST-CAMPAIGN-FIXTURE.md where it helps answer continuity questions.

You do not need to execute the entire fixture against every candidate.

The most important checks are:

  • possession/state consistency,
  • restore/undo behavior,
  • abandoned-path isolation,
  • whether an old discarded fact can leak into current memory/context.

10. What Not to Do

Do not:

  • build the production fork,
  • merge repositories,
  • redesign the full UI,
  • implement full RAG,
  • implement complete branching in Open Dungeon,
  • remove all RPG code from AI-DnD,
  • build a browser frontend for ai-adventure,
  • add image/video/TTS/STT features,
  • write the final production milestone plan.

Small disposable code changes are allowed only when needed to answer the evaluation questions.

11. Final Deliverable

The main output from this round should be a single recommendation document:

PHASE-0B-RECOMMENDATION.md

It should summarize what was learned, not just list test logs.

Include:

A. Executive Recommendation

  • Which repository should be the production base?
  • Confidence level: high / medium / low.
  • Did the initial AI-DnD recommendation hold up?

B. What We Learned About Each Candidate

For each:

  • what worked,
  • what failed,
  • strongest reusable pieces,
  • major architectural problems,
  • likely amount/type of adaptation required.

C. Key Technical Findings

Especially:

  • history/undo model,
  • state rollback,
  • memory isolation,
  • local Ollama support,
  • offline/privacy behavior,
  • browser suitability,
  • imported-knowledge potential,
  • future media extension potential.

D. Important Surprises

Anything that contradicts the current planning assumptions.

E. Recommendation for Next Step

Do not perform the next step.

Instead recommend what should happen next, such as:

  • fork AI-DnD and begin a controlled strip-down,
  • perform one additional experiment first,
  • reconsider Open Dungeon,
  • use ai-adventure as the base instead,
  • revise one of the product assumptions.

F. Open Questions

List anything that could not be resolved in this round.

12. Supporting Evidence

You may also create concise supporting notes/logs for:

  • baseline results,
  • AI-DnD experiment,
  • Open Dungeon history analysis,
  • ai-adventure Ollama adapter,
  • offline/network observations.

Keep them concise.

The recommendation document is the primary deliverable.

13. Stop Condition

When PHASE-0B-RECOMMENDATION.md is complete, stop.

We will take the findings back into the design discussion, re-examine the assumptions, and decide the next step before any production implementation begins.