273 lines
7.9 KiB
Markdown
273 lines
7.9 KiB
Markdown
# Phase 0B — Codex Initial Validation Brief
|
|
|
|
**Status:** COMPLETE / HISTORICAL — do not execute as a production prompt
|
|
**Purpose:** Run a focused first round of local validation on the three finalist repositories and return a recommendation based on what was actually learned.
|
|
|
|
## 1. Goal
|
|
|
|
We are **not** asking you to build the production application yet.
|
|
|
|
The goal of this round is to answer one question:
|
|
|
|
> Which existing project is the best starting point for the local interactive-story application, and what important technical facts did we learn that should affect the next design step?
|
|
|
|
The three finalists are:
|
|
|
|
1. AI-DnD
|
|
https://github.com/parththakkar106/AI-DnD
|
|
|
|
2. Open Dungeon
|
|
https://github.com/newideas99/open-dungeon
|
|
|
|
3. ai-adventure
|
|
https://github.com/CaoRuiming/ai-adventure
|
|
|
|
## 2. Important Product Requirements
|
|
|
|
Use these as the main evaluation criteria.
|
|
|
|
The eventual application should be:
|
|
|
|
- browser-first,
|
|
- local-only in v1,
|
|
- based on local Ollama inference,
|
|
- single-user,
|
|
- genre-agnostic,
|
|
- persistent across restarts,
|
|
- able to retain authoritative story state separately from model prose,
|
|
- able to Undo/Redo/Retry safely,
|
|
- able to preserve named save points/checkpoints,
|
|
- able to retain abandoned history without immediately deleting it,
|
|
- able to prevent abandoned-history facts/memories from leaking into the active story,
|
|
- able to support long-term context/memory,
|
|
- able to import local knowledge,
|
|
- able to inspect what context was sent to the model,
|
|
- architecturally compatible with future local image/video/TTS/STT support.
|
|
|
|
Do not try to implement all of these now.
|
|
|
|
This round is about determining which candidate already gives us the strongest foundation.
|
|
|
|
## 3. Read Only What You Need
|
|
|
|
Start with:
|
|
|
|
1. `README.md`
|
|
2. `SPECIFICATION.md`
|
|
3. `reports/PRELIMINARY-RECOMMENDATION.md`
|
|
4. `reports/REUSE-MATRIX.md`
|
|
|
|
Then use these only when relevant to a specific experiment:
|
|
|
|
- `STORY-BRANCH-SEMANTICS.md`
|
|
- `CONTEXT-AND-MEMORY.md`
|
|
- `SECURITY-THREAT-MODEL.md`
|
|
- `TEST-CAMPAIGN-FIXTURE.md`
|
|
|
|
Do **not** read every planning document up front unless needed.
|
|
|
|
## 4. Baseline Work for Each Candidate
|
|
|
|
For each repository:
|
|
|
|
1. Clone it cleanly.
|
|
2. Record the exact commit SHA.
|
|
3. Follow the documented install instructions.
|
|
4. Run the existing tests.
|
|
5. Build/start the application.
|
|
6. Confirm the basic local story flow.
|
|
7. Record:
|
|
- test results,
|
|
- storage/database technology,
|
|
- model/provider assumptions,
|
|
- local ports,
|
|
- major runtime failures,
|
|
- obvious cloud/hosted dependencies.
|
|
|
|
Do not spend excessive time fixing unrelated upstream problems.
|
|
|
|
If a project does not run cleanly, document why and continue.
|
|
|
|
## 5. Focused Experiment A — AI-DnD
|
|
|
|
We want to know whether AI-DnD can realistically serve as the production base.
|
|
|
|
Test:
|
|
|
|
- Can it run with Ollama locally?
|
|
- Can its story-tree/history system support simple user-facing Undo/Redo/Retry behavior?
|
|
- Does state rollback work independently of heavy RPG/stat mechanics?
|
|
- Can RPG-specific state be left empty/minimal without breaking the useful history/state architecture?
|
|
- Can QuickJS/scripting and hosted/cloud-oriented features be disabled without breaking local story operation?
|
|
- Can local memory/embedding behavior work without cloud services?
|
|
- Do memories/state respect the active history path?
|
|
- Is its context/Insights system useful for showing what was sent to the model?
|
|
|
|
Use a small disposable experiment if necessary.
|
|
|
|
Do **not** start stripping the whole application down.
|
|
|
|
## 6. Focused Experiment B — Open Dungeon
|
|
|
|
We want to know how expensive it would be to fix its history model.
|
|
|
|
Test:
|
|
|
|
- Confirm how Retry/Edit/Erase affect stored history.
|
|
- Identify whether old future turns are deleted.
|
|
- Trace which parts of the application depend on that linear/destructive behavior.
|
|
- Estimate how invasive it would be to change to:
|
|
- parent-linked turns,
|
|
- active head,
|
|
- retained abandoned history,
|
|
- named checkpoints,
|
|
- lineage-safe summaries/state.
|
|
|
|
Do not implement the complete branch system.
|
|
|
|
Also record useful existing pieces:
|
|
- browser UX,
|
|
- local image generation,
|
|
- character visual continuity,
|
|
- any scene/media architecture worth reusing.
|
|
|
|
## 7. Focused Experiment C — ai-adventure
|
|
|
|
We want to know whether its strong state/privacy architecture can realistically become a browser-based Ollama application.
|
|
|
|
Test:
|
|
|
|
- Run the existing tests.
|
|
- Confirm Undo/branch/checkpoint/replay behavior.
|
|
- Identify the provider abstraction.
|
|
- Prove one local Ollama-backed story turn using the smallest practical adapter.
|
|
- Determine how tightly the core application logic is coupled to the CLI.
|
|
- Assess whether the core could sit behind a browser/API layer without moving authoritative state logic.
|
|
- Review its local lore/FTS approach for possible reuse.
|
|
|
|
Do not build a browser frontend.
|
|
|
|
## 8. Offline / Privacy Check
|
|
|
|
For each candidate, once dependencies/models are installed:
|
|
|
|
- run it with outbound Internet unavailable or blocked where practical,
|
|
- exercise basic story generation,
|
|
- note any unexpected network attempts.
|
|
|
|
We do not need a full penetration test in this round.
|
|
|
|
We do need to know:
|
|
|
|
- whether local story use truly works offline,
|
|
- whether cloud services are required,
|
|
- whether analytics/telemetry/remote assets are present,
|
|
- how difficult those paths would be to remove.
|
|
|
|
## 9. Use the Standard Fixture Selectively
|
|
|
|
Use `TEST-CAMPAIGN-FIXTURE.md` where it helps answer continuity questions.
|
|
|
|
You do not need to execute the entire fixture against every candidate.
|
|
|
|
The most important checks are:
|
|
|
|
- possession/state consistency,
|
|
- restore/undo behavior,
|
|
- abandoned-path isolation,
|
|
- whether an old discarded fact can leak into current memory/context.
|
|
|
|
## 10. What Not to Do
|
|
|
|
Do not:
|
|
|
|
- build the production fork,
|
|
- merge repositories,
|
|
- redesign the full UI,
|
|
- implement full RAG,
|
|
- implement complete branching in Open Dungeon,
|
|
- remove all RPG code from AI-DnD,
|
|
- build a browser frontend for ai-adventure,
|
|
- add image/video/TTS/STT features,
|
|
- write the final production milestone plan.
|
|
|
|
Small disposable code changes are allowed only when needed to answer the evaluation questions.
|
|
|
|
## 11. Final Deliverable
|
|
|
|
The main output from this round should be a single recommendation document:
|
|
|
|
```text
|
|
PHASE-0B-RECOMMENDATION.md
|
|
```
|
|
|
|
It should summarize what was learned, not just list test logs.
|
|
|
|
Include:
|
|
|
|
### A. Executive Recommendation
|
|
|
|
- Which repository should be the production base?
|
|
- Confidence level: high / medium / low.
|
|
- Did the initial AI-DnD recommendation hold up?
|
|
|
|
### B. What We Learned About Each Candidate
|
|
|
|
For each:
|
|
- what worked,
|
|
- what failed,
|
|
- strongest reusable pieces,
|
|
- major architectural problems,
|
|
- likely amount/type of adaptation required.
|
|
|
|
### C. Key Technical Findings
|
|
|
|
Especially:
|
|
- history/undo model,
|
|
- state rollback,
|
|
- memory isolation,
|
|
- local Ollama support,
|
|
- offline/privacy behavior,
|
|
- browser suitability,
|
|
- imported-knowledge potential,
|
|
- future media extension potential.
|
|
|
|
### D. Important Surprises
|
|
|
|
Anything that contradicts the current planning assumptions.
|
|
|
|
### E. Recommendation for Next Step
|
|
|
|
Do **not** perform the next step.
|
|
|
|
Instead recommend what should happen next, such as:
|
|
- fork AI-DnD and begin a controlled strip-down,
|
|
- perform one additional experiment first,
|
|
- reconsider Open Dungeon,
|
|
- use ai-adventure as the base instead,
|
|
- revise one of the product assumptions.
|
|
|
|
### F. Open Questions
|
|
|
|
List anything that could not be resolved in this round.
|
|
|
|
## 12. Supporting Evidence
|
|
|
|
You may also create concise supporting notes/logs for:
|
|
|
|
- baseline results,
|
|
- AI-DnD experiment,
|
|
- Open Dungeon history analysis,
|
|
- ai-adventure Ollama adapter,
|
|
- offline/network observations.
|
|
|
|
Keep them concise.
|
|
|
|
The recommendation document is the primary deliverable.
|
|
|
|
## 13. Stop Condition
|
|
|
|
When `PHASE-0B-RECOMMENDATION.md` is complete, stop.
|
|
|
|
We will take the findings back into the design discussion, re-examine the assumptions, and decide the next step before any production implementation begins.
|