Files
interactive-story/planning/RESEARCH-PLAN.md

619 lines
15 KiB
Markdown

# Adventure Storyteller — Phase 0 Research Plan
**Status:** Complete — Phase 0 closed 2026-09-01
**Phase:** 0 — Research, Validation & Architecture
**Goal:** Determine what to build, what to fork/reuse, and finalize the technical design before production implementation begins.
**Outcome:** AI-DnD selected as the production base; non-destructive head-cursor history and explicit typed narrative-state events selected; production milestones now defined in `BUILD-MILESTONES.md`.
## 1. Why Phase 0 Exists
The project has several promising open-source starting points. They differ substantially in:
- browser UX,
- persistence model,
- branching semantics,
- long-term memory,
- local knowledge retrieval,
- Ollama support,
- dependency footprint,
- privacy/network behavior,
- game-specific assumptions,
- licensing,
- test quality.
A detailed production milestone plan written before inspecting the code would rely on guesses.
Phase 0 therefore ends when we can answer:
> What exact codebase and architecture should be used for the production storyteller?
No production feature work should begin before that decision unless explicitly authorized.
## 1A. Phase 0 Completion Summary
Phase 0A static research and Phase 0B local validation are complete.
Final dispositions:
- AI-DnD — production fork/base at `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`.
- ai-adventure — primary implementation reference for authoritative typed state events, checkpoint/head/replay semantics, and narrow local-only behavior.
- Open Dungeon — UX and future-media reference only.
Critical Phase 0B prototype results:
- non-destructive Undo/Redo with a movable active head was demonstrated on AI-DnD without deleting history and without breaking branch-scoped memory isolation,
- export/import must preserve active head position,
- AI-DnD's relative-delta state protocol should not be retained as the generic narrative-state contract,
- imported knowledge should be a separate subsystem rather than Story Cards,
- runtime offline hardening is required for tokenizer data and fonts.
The original milestone text below is retained as the research execution record.
## 2. Candidate Repositories
Initial candidates:
1. **Open Dungeon**
- Repository: `newideas99/open-dungeon`
- Interest: browser-first interactive-fiction UX, Ollama, SQLite.
2. **AI-DnD**
- Repository: `parththakkar106/AI-DnD`
- Interest: browser UI, story tree, rollback, memory, story cards, prompt inspection.
3. **Local Adventure Engine / ai-adventure**
- Repository: `CaoRuiming/ai-adventure`
- Interest: append-only state, checkpoints, branching, privacy-focused architecture, deterministic replay.
4. **aiMultiFool**
- Repository: exact upstream URL to be confirmed during inventory.
- Interest: local semantic memory/RAG and context inspection.
Reference projects:
- SillyTavern
- RisuAI
- KoboldAI
- Chronicler
- other credible projects discovered during Phase 0.
## 3. Research Workspace
Create a dedicated workspace such as:
```text
adventure-storyteller-research/
├── candidates/
│ ├── open-dungeon/
│ ├── ai-dnd/
│ ├── ai-adventure/
│ └── aimultifool/
├── notes/
├── experiments/
├── reports/
└── inventory/
```
Do not copy source code from one project into another during initial analysis.
Each candidate should remain a clean upstream clone or worktree.
Record:
- upstream URL,
- upstream default branch,
- commit SHA examined,
- release/tag if applicable,
- clone date,
- license,
- language/framework,
- build tooling,
- runtime services,
- expected local ports.
## 4. Phase Rules
During Phase 0:
- do not begin production feature development,
- do not merge candidate codebases,
- do not remove features from candidate repos,
- do not commit speculative refactors,
- small disposable experiments are allowed,
- experiments must be isolated and clearly documented,
- candidate repos should remain easy to reset to upstream,
- every conclusion should cite observed code/config/test behavior.
## 5. Milestone R0 — Research Workspace and Inventory
### Objective
Create the research environment and establish a reproducible inventory of all candidates.
### Tasks
- create the research workspace,
- clone all initial candidates,
- record exact upstream commits,
- locate and record licenses,
- inventory languages/frameworks,
- inventory package managers,
- inventory database/storage dependencies,
- inventory model/provider dependencies,
- inventory frontend/backend separation,
- record build/run instructions,
- identify existing tests,
- identify documentation directories,
- identify migrations/schema definitions,
- identify obvious telemetry/cloud integrations.
### Deliverables
- `inventory/candidates.md`
- `inventory/licenses.md`
- `inventory/dependencies.md`
- `inventory/build-instructions.md`
- machine-readable candidate metadata if useful.
### Exit Criteria
All serious candidates are locally available and reproducibly identified.
## 6. Milestone R1 — Build and Run Candidates
### Objective
Verify actual behavior rather than relying on README claims.
### Tasks
For each serious candidate:
- install dependencies,
- build successfully where applicable,
- start locally,
- create a minimal story,
- confirm persistence after restart,
- test Ollama directly where supported,
- identify how model configuration works,
- record application ports,
- identify data locations,
- inspect browser developer/network activity for unexpected outbound requests where applicable,
- record startup failures or undocumented requirements.
For projects not supporting Ollama:
- determine adapter/interface boundary,
- do not yet permanently modify the project.
### Deliverables
Per candidate:
```text
reports/runtime-<candidate>.md
```
Include:
- exact commands,
- success/failure,
- screenshots only if useful,
- local services used,
- observed storage files,
- observed network activity,
- known blockers.
### Exit Criteria
Each serious candidate has either been run successfully or has a documented reason it cannot reasonably be evaluated.
## 7. Milestone R2 — Source Architecture Review
### Objective
Understand how each candidate actually works internally.
### Review Areas
#### Browser/UI
- framework,
- state management,
- streaming,
- transcript representation,
- campaign navigation,
- edit/retry behavior,
- extensibility for future media.
#### Backend/service layer
- routing/API design,
- model invocation boundary,
- background jobs,
- validation boundaries.
#### Persistence
- database type,
- schema,
- migrations,
- turn representation,
- snapshots,
- event log,
- transactions,
- branch representation.
#### Story history
- linear vs tree,
- retry semantics,
- undo semantics,
- destructive vs non-destructive restore,
- branch naming/navigation.
#### Context
- prompt assembly,
- recent history,
- summaries,
- token budgeting,
- author notes/system rules.
#### Memory
- summaries,
- vector retrieval,
- keyword retrieval,
- entity state,
- old-turn retrieval.
#### Lore/knowledge
- import formats,
- chunking,
- story cards/world info,
- semantic retrieval,
- provenance.
#### Tests
- unit tests,
- integration tests,
- migration tests,
- model mocks,
- coverage of state/rollback.
### Deliverables
- `reports/architecture-open-dungeon.md`
- `reports/architecture-ai-dnd.md`
- `reports/architecture-ai-adventure.md`
- `reports/architecture-aimultifool.md`
- `reports/architecture-comparison.md`
### Exit Criteria
We can explain each candidate's architecture without relying on marketing descriptions.
## 8. Milestone R3 — Privacy and Network Review
### Objective
Determine what must be removed, disabled, or isolated to satisfy the local-only requirement.
### Search For
- OpenAI,
- OpenRouter,
- Groq,
- Anthropic,
- Google,
- cloud inference,
- telemetry,
- analytics,
- Sentry,
- PostHog,
- crash reporting,
- CDN,
- Google Fonts,
- remote image hosts,
- automatic update checks,
- remote database support,
- URL retrieval,
- external web search,
- MCP,
- plugins,
- arbitrary executable scripts,
- third-party auth.
### Tasks
- static source search,
- dependency review,
- environment-variable review,
- runtime network observation,
- identify outbound requests required vs optional,
- identify localhost vs wildcard binds,
- identify stored secrets/API keys,
- identify browser-side remote resources.
### Deliverables
- `reports/privacy-network-review.md`
- per-candidate removal/mitigation list.
### Exit Criteria
For each candidate, we can state exactly what local-only hardening would be required.
## 9. Milestone R4 — Feature and Reuse Matrix
### Objective
Compare candidates by subsystem rather than declaring one project the winner prematurely.
### Compare
- browser UX,
- Ollama adapter,
- streaming,
- SQLite schema,
- story tree,
- rollback,
- checkpoints,
- edit/retry semantics,
- state extraction,
- entity/world state,
- summaries,
- semantic memory,
- lexical memory,
- lore/story cards,
- source imports,
- prompt inspection,
- export/import,
- tests,
- local-only posture,
- media extensibility.
### Rate Each Feature
Use categories such as:
- Keep as-is
- Keep with modification
- Reuse concept only
- Replace
- Not present
- Not wanted
### Deliverables
- `reports/reuse-matrix.md`
### Exit Criteria
We know which candidate has the best implementation of each required subsystem.
## 10. Milestone R5 — Licensing and Code-Reuse Review
### Objective
Determine what code can legally be copied, modified, linked, or used only as inspiration.
### Tasks
- verify repository licenses at the exact commits reviewed,
- note third-party code with separate licenses,
- note generated/vendor code,
- compare compatibility if combining code from multiple projects,
- pay special attention to GPL/copyleft candidates,
- distinguish:
- direct code reuse,
- dependency use,
- architecture inspiration,
- protocol/API reimplementation.
### Deliverables
- `reports/licensing-reuse.md`
### Exit Criteria
The recommended architecture does not rely on legally ambiguous code mixing.
## 11. Milestone R6 — Critical Prototypes
### Objective
Test only the uncertainties that could change the architecture decision.
Possible experiments include:
### Experiment A — Ollama adapter for ai-adventure
Determine how difficult it is to replace/extend the LM Studio adapter with Ollama.
### Experiment B — Branch-safe state in Open Dungeon
Determine whether Open Dungeon's current persistence can support immutable branch parentage without invasive rewrite.
### Experiment C — Strip-down feasibility in AI-DnD
Identify whether RPG/cloud systems are modular enough to remove without destabilizing core story-tree/memory behavior.
### Experiment D — Local semantic retrieval
Test a minimal local embedding pipeline using Ollama and a local-only store.
### Experiment E — Scene extraction
Verify that the narrator/state pipeline can produce a neutral scene packet suitable for future media.
Only run experiments that resolve a documented decision.
### Deliverables
Each experiment:
```text
experiments/<name>/README.md
```
Record:
- question,
- hypothesis,
- minimal changes,
- result,
- implications,
- whether code should be discarded.
### Exit Criteria
No high-impact fork/architecture decision remains based solely on speculation.
## 12. Milestone R7 — Fork / Build Decision
### Objective
Select the production starting strategy.
### Required Options to Evaluate
- fork Open Dungeon,
- fork AI-DnD,
- fork/use ai-adventure core,
- clean new shell with reused permissive components,
- other candidate if discovered.
### Decision Criteria
Weight heavily:
1. fit with interactive-story product,
2. browser-first architecture,
3. Ollama fit,
4. state/branch correctness,
5. local-only hardening effort,
6. amount of code to remove,
7. maintainability,
8. licensing,
9. test quality,
10. future media extensibility.
### Deliverables
- `reports/fork-build-recommendation.md`
- ADR documenting the selected strategy.
### Exit Criteria
One strategy is approved as the production base.
## 13. Milestone R8 — Finalize Specification and Technical Design
### Objective
Convert assumptions into committed decisions.
### Tasks
Update:
- `SPECIFICATION.md`
- `TECHNICAL-DESIGN.md`
Resolve:
- base repository,
- frontend framework,
- backend framework,
- storage model,
- branch/state model,
- model adapter,
- memory/retrieval strategy,
- local knowledge design,
- import formats for v1,
- context budgeting strategy,
- security boundaries,
- export format,
- future media interfaces,
- test strategy.
Mark documents v1.0 when approved.
### Deliverables
- `SPECIFICATION.md` v1.0
- `TECHNICAL-DESIGN.md` v1.0
- relevant ADRs.
### Exit Criteria
A developer can explain the final architecture without unresolved foundational choices.
## 14. Milestone R9 — Create Production Build Plan
### Objective
Write the detailed implementation milestone plan only after the technical design is stable.
### Tasks
Create:
- `BUILD-MILESTONES.md`
It must include:
- milestone dependencies,
- exact intended outcomes,
- acceptance criteria,
- test expectations,
- migration steps from selected upstream,
- removal/hardening work,
- v1 feature sequence,
- definition of done.
### Exit Criteria
The build plan is specific enough to hand directly to Codex milestone-by-milestone.
## 15. Phase 0 Final Deliverables
At Phase 0 completion:
```text
SPECIFICATION.md v1.0
TECHNICAL-DESIGN.md v1.0
RESEARCH-PLAN.md completed
BUILD-MILESTONES.md production-ready
DECISIONS/ finalized foundational ADRs
reports/ research evidence
experiments/ critical prototype evidence
inventory/ candidate metadata
```
## 16. Phase 0 Definition of Done
Phase 0 is complete only when:
- candidate repositories have been cloned and reviewed,
- serious candidates have been run or ruled out with evidence,
- network/privacy behavior is documented,
- licensing is understood,
- critical architectural uncertainties have been tested,
- a fork/build strategy has been selected,
- the specification is v1.0,
- the technical design is v1.0,
- the actual implementation milestone plan has been written.
At that point, production implementation becomes eligible to begin, but still requires explicit approval and a milestone-specific execution prompt.
## 17. Final Phase 0 Closure Record
Phase 0 definition of done is satisfied for architecture/planning purposes:
- finalists cloned and run,
- test suites measured,
- Ollama exercised locally,
- offline/network behavior investigated,
- critical history/state uncertainties prototyped,
- fork/build strategy selected,
- specification revised to v1.0,
- technical design revised to v1.0,
- production build milestones written,
- foundational ADRs updated/added.
No production implementation prompt is part of this research plan.