M01, the hundred-turn campaign, is the one REQUIRED test still outstanding. Everything here is about it finishing, and being worth believing when it does. No requirement changed, no acceptance test was retired or relaxed, and M11 §P.1's "no performance requirement" still stands: what changed is the cost of a turn, not what a turn contains. An inference server caches a prompt by its prefix. The history window gave up its oldest action every turn, which changed the prompt near the front and threw that cache away, so nearly the whole prompt was reprocessed every turn however little had actually changed. The window now snaps the oldest depth to a block and holds it, stepping every few turns. Measured on real builder output at an 8,192-token budget: 124.0s per turn against 362.4s. The cost is history depth, bounded by TRIM_FRACTION at a quarter of the window, which is the dial between recent history and speed. A run that dies no longer starts again from turn one. m11_long_run checkpoints resume.json after the prologue, after every scheduled step and after every turn, and --resume reattaches to the same campaign. A finished run deletes it, so the file's presence means an unfinished run and starting fresh over one is refused. The model timeout is an option rather than a hard-coded 600s, a turn that overruns is a failed turn instead of an unhandled exception that ends the run with no summary, and a run that has stopped producing turns writes its evidence and stops. Two checks could not fail. M04's planted clue went into an add_fact "detail" key that the event does not define, so it was dropped and fact_still_in_state could never be true; it is now in "value" and proved at turn one, which stops a run measuring nothing for hours. m11_browser degraded silently without a narrator into two failures that read exactly like a product regression, and now requires one, with --no-narrator as an explicit opt-out that marks the run partial. Window discovery speaks Ollama's native API, so against vLLM or llama.cpp's own server the window goes unverified and the budget uncapped -- M11's own failure mode reached by another route. context_window_override lets the operator state what they launched the server with, and is used only where discovery left a hole: a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. "verified" still means the server answered, so window_verified in a turn's provenance keeps the meaning M11's report counts on. planning/README.md said the M11 tree was staged rather than committed, in two places; it was committed and signed. Planning package v3.8. Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and build clean. Every M11 harness re-run on this tree: browser 38/0/0, offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a small bundle. M01 itself has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
1384 lines
64 KiB
Markdown
1384 lines
64 KiB
Markdown
# Adventure Storyteller — Technical Design
|
|
|
|
**Status:** v1.0 — architecture selected after Phase 0B
|
|
**Production base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
|
|
|
## 1. Design Objective
|
|
|
|
Build a local-first, browser-based interactive storytelling application in which:
|
|
|
|
- Ollama provides inference on user-controlled local infrastructure, either same-host or on an explicitly approved trusted-LAN machine,
|
|
- the application owns authoritative story state,
|
|
- complete story history is retained,
|
|
- user-facing Undo/Redo/Retry/Save Point behavior is simple,
|
|
- internal history is non-destructive and lineage-aware,
|
|
- long-running context is reconstructed from state, summaries, retrieval, recent turns, and local knowledge,
|
|
- imported knowledge remains local and authority-classified,
|
|
- prompt/context provenance is inspectable,
|
|
- future image/video/audio/TTS/STT providers can be added without coupling them to the story engine.
|
|
|
|
The application is an interactive-story system, not a D&D rules engine.
|
|
|
|
## 2. Production Base Decision
|
|
|
|
AI-DnD is the production fork/base.
|
|
|
|
Retain its high-value foundation:
|
|
|
|
- React/Vite browser application,
|
|
- FastAPI backend,
|
|
- SQLite persistence/migrations,
|
|
- story-tree and lineage machinery,
|
|
- alternate takes,
|
|
- branch switching,
|
|
- per-node state snapshot pattern,
|
|
- local Ollama/OpenAI-compatible integration where appropriate,
|
|
- branch-scoped memory concepts,
|
|
- prompt/context snapshots and Insights concepts,
|
|
- SSE streaming,
|
|
- export/import framework,
|
|
- automated test foundation.
|
|
|
|
Do not inherit candidate behavior merely because it exists upstream. The product specification and detailed behavior documents remain authoritative.
|
|
|
|
## 3. Reference Projects and Their Role
|
|
|
|
### ai-adventure
|
|
|
|
Primary implementation reference for:
|
|
|
|
- explicit typed state events,
|
|
- model-proposes/application-validates discipline,
|
|
- atomic commit behavior,
|
|
- head movement instead of destructive Undo,
|
|
- named checkpoints,
|
|
- replay/state reconstruction,
|
|
- narrow local-only endpoint handling,
|
|
- deterministic lexical lore concepts.
|
|
|
|
It is **not** the written specification and does not override this design.
|
|
|
|
### Open Dungeon
|
|
|
|
Reference for:
|
|
|
|
- focused story-reading UX,
|
|
- inline generated media presentation,
|
|
- character visual continuity,
|
|
- local media-service boundaries.
|
|
|
|
It is not a production-fork candidate.
|
|
|
|
### Other references
|
|
|
|
- Chronicler: memory authority/trust tiers.
|
|
- Interactive Fiction Framework: canon/application-owned-state concepts.
|
|
- Gamentic: provider-neutral asynchronous media boundary.
|
|
|
|
## 4. Selected High-Level Architecture
|
|
|
|
```text
|
|
Local Browser
|
|
|
|
|
v
|
|
+-------------------+
|
|
| React/Vite UI |
|
|
+---------+---------+
|
|
|
|
|
v
|
|
+-------------------+
|
|
| FastAPI Story |
|
|
| Director/API |
|
|
+---+-----------+---+
|
|
| |
|
|
| +-----------------------+
|
|
v v
|
|
+----------------------+ +----------------------+
|
|
| SQLite Authoritative | | Context / Retrieval |
|
|
| Story Store | | local only |
|
|
+----------+-----------+ +----------+-----------+
|
|
| |
|
|
| +----------+-----------+
|
|
| | FTS + local semantic |
|
|
| | retrieval |
|
|
| +----------+-----------+
|
|
| |
|
|
+--------------------+---------------+
|
|
|
|
|
v
|
|
+----------------------+
|
|
| Ollama |
|
|
| localhost by default |
|
|
| or approved LAN host |
|
|
+----------------------+
|
|
```
|
|
|
|
Future optional extension:
|
|
|
|
```text
|
|
Accepted Story / Scene State
|
|
|
|
|
v
|
|
+------------------+
|
|
| Media Coordinator|
|
|
+--+---+---+---+---+
|
|
| | | | |
|
|
v v v v v
|
|
Image Video Audio TTS STT
|
|
local providers only by default
|
|
```
|
|
|
|
## 5. Local-Only Runtime Boundary
|
|
|
|
`Local-only` means user-controlled local infrastructure with no required Internet/cloud dependency; it does not require every process to share one host.
|
|
|
|
The architecture has **two boundaries, and they are not the same boundary**. The
|
|
storyteller is loopback-only, always. Inference may be same-host or on a
|
|
specifically configured trusted-LAN machine. Wording that describes Ollama as
|
|
simply "loopback/local" collapses the two and understates the intended
|
|
deployment:
|
|
|
|
```text
|
|
Browser ──loopback──> storyteller (FastAPI + SPA), bound to 127.0.0.1
|
|
│
|
|
├── same-host Ollama on 127.0.0.1:11434 (default)
|
|
│
|
|
└── OR an explicitly configured trusted-LAN Ollama
|
|
on another user-controlled machine
|
|
http://<host>:11434/v1 or https://<host>/v1
|
|
```
|
|
|
|
Read that as three separate rules:
|
|
|
|
1. **The storyteller's own listener is loopback, in every run path.** Dev
|
|
server, production server, and container alike. Nothing about the inference
|
|
choice changes it. Where a container must listen on `0.0.0.0` because a
|
|
published port cannot reach anything else, the port is published to the
|
|
host's loopback only.
|
|
2. **The inference endpoint is an outbound connection, chosen by the user.**
|
|
Same-host loopback is the default. A trusted-LAN host is a first-class,
|
|
supported v1 configuration — not a workaround and not a development-only
|
|
convenience.
|
|
3. **The two are independent.** Reaching a LAN Ollama never requires, and must
|
|
never cause, LAN exposure of the storyteller UI/API. There is no supported
|
|
v1 configuration in which the storyteller itself is reachable from the LAN.
|
|
|
|
### Transport to a trusted-LAN endpoint
|
|
|
|
A LAN inference host is often reached over **HTTPS with a certificate issued by
|
|
a private or local CA**, and may offer no cleartext port at all. This is
|
|
ordinary for a self-hosted server, so v1 must handle it rather than assume the
|
|
same plain HTTP that same-host loopback uses:
|
|
|
|
- outbound HTTPS verifies against the **operating system's trusted CA store** in
|
|
addition to any bundled certificate list, so a CA the user installed on their
|
|
own machine is honoured here as it is by `curl` and their browser;
|
|
- certificate **and hostname** verification stay fully enabled;
|
|
- there is **no "ignore TLS errors" option** anywhere — not in the UI, not in
|
|
configuration, not as an environment variable;
|
|
- the endpoint field therefore accepts `https://` on any port.
|
|
|
|
Prompts, story text, retrieved knowledge and embedding inputs all travel to
|
|
whichever inference host is configured, which is why it must be one the user
|
|
controls on a network they trust — and why the endpoint is always explicitly
|
|
configured, never discovered.
|
|
|
|
See ADR 002 (*Transport for a Trusted-LAN Endpoint*) and, for the demonstrated
|
|
deployment, `planning/archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md` §F.
|
|
|
|
Allowed future local paths may also include explicitly configured local media services.
|
|
|
|
Production defaults must not require:
|
|
|
|
- cloud model providers,
|
|
- hosted authentication,
|
|
- analytics/telemetry,
|
|
- remote database services,
|
|
- runtime CDNs/fonts/assets,
|
|
- automatic web retrieval,
|
|
- external embeddings/vector stores,
|
|
- general plugins/MCP/shell execution.
|
|
|
|
### 5.1 Known AI-DnD hardening work
|
|
|
|
Phase 0B identified concrete inherited violations, and M1 added a fifth. **All
|
|
five are now resolved** — items 1, 2 and 5 in M1, items 3 and 4 in M2.
|
|
|
|
1. ~~`tiktoken` attempts to download the `cl100k_base` encoding on first use.~~
|
|
**Done in M1.** The encoding table is vendored in the tree and loaded
|
|
directly, with its SHA-256 verified against the digest `tiktoken` pins, so no
|
|
code path in the tokenizer can reach the network.
|
|
2. ~~the SPA requests Google Fonts at runtime.~~
|
|
**Done in M1.** All three families are self-hosted, and the CSP names no
|
|
remote origin at all.
|
|
3. ~~hosted/multi-user/auth/demo/analytics/Postgres/cloud-provider/QuickJS paths are unnecessary.~~
|
|
- remove them rather than merely hide them where practical.
|
|
- **Done in M2, in full.** Removed rather than hidden: 52 API routes fell to
|
|
36, and `/api/auth`, `/api/analytics` and `/api/scripts` are gone entirely
|
|
rather than gated. See §5.2.
|
|
4. ~~endpoint validation must reflect this product's threat model.~~
|
|
- same-host loopback Ollama is the default; an explicitly configured trusted-LAN Ollama endpoint is supported; arbitrary public/Internet model endpoints must be rejected or kept outside normal v1 configuration.
|
|
- inference endpoint configuration must not change the storyteller's own loopback bind behavior.
|
|
- **Done in M2.** `backend/app/endpoints.py` applies an address-based
|
|
allowlist on save and again before every outbound request. Endpoint
|
|
configuration has no influence on the storyteller's own bind address.
|
|
ADR 011; `SECURITY-THREAT-MODEL.md` §10A.
|
|
5. ~~outbound TLS verified only against a bundled public-CA list, so a LAN host
|
|
with a privately issued certificate was refused.~~
|
|
**Found and fixed in M1.** Not visible to Phase 0B: every run up to that
|
|
point used plain HTTP over loopback, where certificate verification never
|
|
happens. See *Transport to a trusted-LAN endpoint* above.
|
|
|
|
### 5.2 Production architecture as established by M1 and M2
|
|
|
|
The architecture below is no longer a selection; it is what the code does. It is
|
|
recorded here so later milestones inherit facts rather than intentions.
|
|
|
|
| | |
|
|
| --- | --- |
|
|
| Production base | AI-DnD, forked at `d72f7c1` (§2, ADR 009) |
|
|
| Persistence | SQLite. Postgres, Neon and the Render deployment path are removed |
|
|
| Inference | Ollama only. No cloud provider code, no API key, no key UI |
|
|
| Storyteller bind | loopback by default, in every run path including the published Docker port |
|
|
| Ollama endpoint | same-host loopback by default; an explicitly configured trusted-LAN endpoint is equally supported |
|
|
| Public endpoints | refused by address, on save and before every request |
|
|
| Trusted-LAN HTTPS | supported, with full certificate and hostname verification against the machine's CA store; no bypass exists |
|
|
| Runtime assets | self-contained. Tokenizer table and fonts are vendored; the CSP names no remote origin |
|
|
|
|
Removed in M2 rather than hidden: hosted accounts and auth, guest/demo
|
|
behaviour, hosted analytics, cloud inference providers, API-key storage and its
|
|
UI, Postgres/Neon/Render support, and QuickJS campaign scripting.
|
|
|
|
**A trusted-LAN Ollama endpoint is accepted production behaviour**, not a
|
|
development convenience. Any statement that the only valid endpoint is literally
|
|
`127.0.0.1` is stale and should be read against §5 and §10A of the threat model.
|
|
|
|
## 6. Browser UI Boundary
|
|
|
|
The browser remains a presentation/control layer, not the owner of story authority.
|
|
|
|
Primary areas:
|
|
|
|
- campaign library/setup,
|
|
- active story transcript,
|
|
- input composer,
|
|
- Undo / Redo / Retry,
|
|
- alternate-take selector,
|
|
- Save Points,
|
|
- current story state inspector,
|
|
- imported-knowledge manager,
|
|
- prompt/context inspector,
|
|
- local-model status/settings,
|
|
- export/import,
|
|
- future media gallery/actions.
|
|
|
|
Normal storytelling should not expose branch IDs, database rows, embeddings, or event logs unless the user opens advanced diagnostics.
|
|
|
|
## 7. Story Director Boundary
|
|
|
|
The FastAPI service owns the turn lifecycle:
|
|
|
|
1. resolve campaign, active branch, and active head,
|
|
2. load authoritative state at that head,
|
|
3. assemble bounded lineage-safe context,
|
|
4. retain prompt/retrieval provenance,
|
|
5. invoke local Ollama narrator,
|
|
6. stream provisional narration,
|
|
7. obtain/parse a structured state proposal,
|
|
8. validate the proposal,
|
|
9. atomically accept the turn plus validated state consequences,
|
|
10. update derived memory/summary/index data without allowing derived failures to corrupt the accepted turn,
|
|
11. expose the new current head to the browser.
|
|
|
|
A failed generation must not partially advance authoritative story state.
|
|
|
|
## 8. Non-Destructive History Model
|
|
|
|
### 8.1 Core rule
|
|
|
|
Undo moves the active head. It does not delete accepted history.
|
|
|
|
Phase 0B demonstrated that AI-DnD already contains the architectural chokepoints needed for this approach:
|
|
|
|
- stored head depth/position,
|
|
- lineage reads through a common path abstraction,
|
|
- branch-at-depth behavior.
|
|
|
|
The disposable spike is evidence, not production code to merge blindly.
|
|
|
|
### 8.2 Active head versus retained tip
|
|
|
|
The design distinguishes:
|
|
|
|
- **retained tip:** newest retained turn on a continuation,
|
|
- **active head:** the story position from which the user is currently reading/continuing.
|
|
|
|
After Undo, the active head may sit behind a retained tip.
|
|
|
|
Redo moves the head forward along the previous active continuation while no divergent write has occurred.
|
|
|
|
### 8.3 Divergence after moving backward
|
|
|
|
If the user writes/retries/edits from a head behind the retained tip:
|
|
|
|
- the new continuation forks on first write,
|
|
- the previous future remains retained,
|
|
- ordinary Redo into that old future is invalidated,
|
|
- the displaced future is marked abandoned/disposable,
|
|
- lineage-sensitive state, summary, and memory selection follows only the new active path.
|
|
|
|
### 8.4 Retry
|
|
|
|
Retry preserves alternate narrator takes for the same user input.
|
|
|
|
Production implementation must ensure retry/add-take while behind the current tip uses the same safe fork/head semantics as other writes.
|
|
|
|
### 8.5 Checkpoints / Save Points
|
|
|
|
A named checkpoint is a durable pointer to a recoverable story position.
|
|
|
|
Conceptually:
|
|
|
|
```text
|
|
campaign_id
|
|
branch_id
|
|
turn_id or equivalent head coordinate
|
|
name
|
|
notes
|
|
created_at
|
|
```
|
|
|
|
Restoring a checkpoint moves the active head to that position. The existing later future remains retained. A new branch is created only if/when the user creates a different continuation.
|
|
|
|
### 8.6 Abandoned history
|
|
|
|
No automatic cleanup policy is required in v1.
|
|
|
|
Abandoned history must:
|
|
|
|
- remain retained,
|
|
- be marked disposable/inactive through implementation-appropriate metadata,
|
|
- stop influencing current state/context/memory/summary,
|
|
- remain available for future recovery/cleanup features.
|
|
|
|
### 8.7 As implemented in M3
|
|
|
|
M3 built this model. The following is fact rather than direction, and ADR 012
|
|
records it as the architectural decision. Sections 8.1-8.6 stand; this says how
|
|
they were realised.
|
|
|
|
**The head is stored, not derived.** A campaign carries a branch and a depth,
|
|
and that pair is the active head. No read may recompute it from the newest row —
|
|
that was the pre-M3 behavior, and it is what made Redo impossible and made an
|
|
export reopen an undone campaign at its tip.
|
|
|
|
**Lineage reads are capped at the head, in one place.** The path abstraction that
|
|
already resolved a branch's ancestry now also limits every entry to the head, so
|
|
the transcript, the assembled narrator context, take/parent resolution and memory
|
|
retrieval narrow together. There is exactly one way to read past the head — a
|
|
named, uncapped view of the same lineage — and only two callers may use it: Redo,
|
|
and the check that decides whether a write must fork. Any new feature that reads
|
|
story rows directly, rather than through the capped lineage, will see retained
|
|
history the story is not telling.
|
|
|
|
**Head movement is one mechanism.** Undo, Redo, and anything later that restores
|
|
a position resolve a target depth and then call a single move operation, which
|
|
sets the coordinate and restores the state recorded at it. Undo and Redo differ
|
|
only in which way they resolve the target. Both step over a whole turn — a
|
|
player's action and the reply to it — so the head never rests between the two
|
|
halves of one turn.
|
|
|
|
**State comes from the node, not from a replay.** Each node records the state it
|
|
left behind, so moving the head is a row lookup plus a restore: the same cost at
|
|
any distance, in either direction, and identical whether the position is reached
|
|
from in front of it or from behind. This is the property §10.4's hybrid storage
|
|
must preserve.
|
|
|
|
### M6 — derived context: summaries, memory and budgeting
|
|
|
|
Four things future milestones rely on, all built on the lineage machinery M3-M5
|
|
established rather than beside it.
|
|
|
|
**Summary lineage — both halves.** A summary is a `summaries` row carrying
|
|
`(branch_id, depth)` for the last node it covers plus a
|
|
`source_start`/`source_end` range. The invariant M6 holds is:
|
|
|
|
> Both summary eligibility and the prior-summary input to the summarizer are
|
|
> lineage-scoped.
|
|
|
|
*Eligibility* is `lineage.Path.clause` over the row's coordinate — the same
|
|
capped-path clause that filters actions and memories — so Undo, Redo, Save Point
|
|
restore and divergence need no summary-specific rule. *Input* is
|
|
`summaries.current`, the same question the context builder asks, so a summary is
|
|
only ever built on top of one that is valid where the story now stands; where
|
|
none is, generation starts from nothing.
|
|
|
|
The second half is not decorative. The first M6 implementation had only the
|
|
first, seeding generation from `adventures.story_summary`, and the review
|
|
demonstrated abandoned prose reaching an active prompt inside a row that was
|
|
itself correctly anchored. Anchoring the output does not make the content safe.
|
|
|
|
Nothing is deleted when a line is abandoned. `adventures.story_summary` survives
|
|
as a reader-facing convenience only — the Plot panel edits it, the export bundle
|
|
carries it — mirroring whichever summary is eligible, kept in step by
|
|
`summaries.record` and by `attempts.restore_state` when the head moves. Nothing
|
|
authoritative reads it.
|
|
|
|
**Memory lineage and provenance.** Unchanged from what M3 built and M6 verified:
|
|
a memory carries `(branch_id, depth)` and a source range, and retrieval filters
|
|
through the capped path. M6 adds provenance to the *retrieval result*, in the
|
|
same query that fetches the text, so the inspector can answer "where did this
|
|
come from?" without a query per memory.
|
|
|
|
**Memory authority.** `Memory.authority` is `accepted_story` or `heuristic`,
|
|
decided by the application in `memorybank.classify_authority`, and rendered into
|
|
the prompt as an explicit mark. Retrieval never writes state; the M5 typed-event
|
|
path remains the only route to an authoritative change.
|
|
|
|
**Retrieval ranking and redundancy.** Ranking is cosine similarity plus an
|
|
explicit pin; the other factors `CONTEXT-AND-MEMORY.md` §20 contemplates are not
|
|
implemented. Before the final top-k cut, retrieval drops a candidate that
|
|
repeats one already chosen, never across authority classes, at a threshold
|
|
measured against the configured embedding model
|
|
(`memorybank.REDUNDANT_SIMILARITY`). Suppressed candidates are reported so the
|
|
selection stays inspectable. Without this, a stretch of repetitive story fills
|
|
the whole memory budget with near-copies and evicts the one memory that
|
|
mattered — which the review measured happening.
|
|
|
|
**Context budgeting.** The reply is reserved out of `context_token_budget`
|
|
before history is selected, with a fixed 64-token margin. Protected content —
|
|
narrator rules, canon, authoritative state, the reader's input, the reply
|
|
reserve — is never dropped to fit older prose; history is the elastic part and
|
|
is filled newest-first until the remaining budget is spent. If the protected
|
|
part alone exceeds the budget, `build_context` raises `ContextOverflow` rather
|
|
than assembling a prompt known to overflow.
|
|
|
|
**Background failure observability.** Derived work (memory extraction, summary
|
|
generation, embedding) runs in a fire-and-forget task and must not take an
|
|
accepted turn down with it. Each pass is wrapped so that a failure rolls back
|
|
only its own uncommitted work and writes a `derived_status` row naming the kind,
|
|
the error and the attempt count. That row is served by
|
|
`GET /adventures/{id}/derived` and shown in the Insights panel. M2 shipped with
|
|
the whole memory bank dead and the suite green; this is the mechanism that makes
|
|
the same failure visible.
|
|
|
|
**Prompt inspection.** The context report carries per-section token counts, the
|
|
budget, the output reserve, the protected total, the history allowance, the
|
|
summary's provenance, each retrieved memory's authority and source coordinate,
|
|
and the derived-work status.
|
|
|
|
**Divergence is a property of the lineage, not a flag.** The first write below a
|
|
moved-back head forks; Undo alone never does. After the fork, the displaced
|
|
future is no longer on the lineage being read, so ordinary Redo finds nothing
|
|
ahead and reports that it has nowhere to go. Nothing has to be invalidated,
|
|
cleared, or kept in step.
|
|
|
|
**A branch the story leaves records the depth and time it was left**, as metadata
|
|
nothing reads to decide behavior (§8.6's "implementation-appropriate metadata").
|
|
It makes a divergence observable and gives later cleanup and recovery features
|
|
something to select on; because no decision depends on it, a stale or hand-edited
|
|
value cannot make the story wrong.
|
|
|
|
**Operations that change what the story says at a position must ask whether
|
|
story descends from that position and is off screen.** Switching which take is
|
|
live, and editing a turn's text in place, both refuse in that situation rather
|
|
than act silently, because retained history must not be made to disagree with
|
|
itself in a way the user cannot see. See `STORY-BRANCH-SEMANTICS.md` §10 and
|
|
§14A.
|
|
|
|
### 8.8 Save Points, as implemented in M4
|
|
|
|
M4 added durable named Save Points and built nothing in §8 that was not already
|
|
there. This records what the milestone establishes as fact.
|
|
|
|
**A Save Point is a name and a coordinate.** The stored row holds the name, an
|
|
optional note, and `(branch, depth)` — the same pair §8.7 calls the head. It
|
|
holds no transcript, no state, no summary, no memory, and no branch contents.
|
|
`DATA-MODEL.md` §8 describes the pointer as naming a turn; the coordinate is
|
|
that turn's address, and `DATA-MODEL.md` §8's implementation note records why
|
|
this project uses the address rather than a row id: one coordinate can hold
|
|
several attempts at a turn, and a retry replaces the live one. "Turn 42 of this
|
|
line" survives a retry; a row id would pin a take the story no longer tells.
|
|
|
|
**Restore is head movement, and nothing else.** It resolves the coordinate,
|
|
refuses it if it no longer names a live turn, and then moves the head — the
|
|
depth through §8.7's single move operation, unchanged. The transcript, the
|
|
assembled context, the state and memory eligibility all arrive together because
|
|
they already read through the one capped lineage. There is no second restore
|
|
path, no state reconstruction, no memory pruning and no separate redo stack:
|
|
D13 is satisfied by the mechanism rather than by code written to satisfy it.
|
|
|
|
**A Save Point may name a position on a line the story has left.** Save Points
|
|
survive divergence, so this is reachable in ordinary use, and the depth half of
|
|
the head cannot reach a branch the current path does not contain. Restore
|
|
therefore moves the branch half as well when, and only when, the coordinate is
|
|
not on the path being read — the same single assignment a branch switch makes.
|
|
The distinction matters in the other direction too: a Save Point in a shared
|
|
prefix must *not* drag the reader onto the ancestor, because which continuation
|
|
follows that turn is exactly what the reader has already chosen.
|
|
|
|
**Restore never forks.** Moving the head is not a decision to abandon anything.
|
|
The first write below the restored head forks, through §8.7's existing check,
|
|
and the displaced future stays retained — so Redo still walks the original
|
|
continuation until the user writes something different, and stops offering it
|
|
once they have.
|
|
|
|
**Nothing removes a Save Point but the user.** There is no cleanup pass, and none
|
|
is wanted: a Save Point pointing behind the head, or into a line the story left,
|
|
is doing its job.
|
|
|
|
That rule is enforced against the one operation that could break it. Deleting a
|
|
branch deletes everything forked from it, so a Save Point naming a position in
|
|
that subtree would go too — silently, since the story is what the user asked to
|
|
delete. **The deletion is therefore refused while any Save Point names that
|
|
subtree**, and the refusal names them. The user deletes the Save Point
|
|
explicitly, which deletes no story, and then the branch. `STORY-BRANCH-SEMANTICS.md`
|
|
§19.1 states the rule; §28 already required a future cleanup feature to retain
|
|
paths a checkpoint references, and this is that requirement applied to the
|
|
deletion path that exists today.
|
|
|
|
## 9. Export / Import and Head Position
|
|
|
|
AI-DnD's current export carries branch information but reconstructs the imported head at the branch tip.
|
|
|
|
That is invalid after non-destructive Undo because an exported campaign can intentionally have:
|
|
|
|
```text
|
|
active head < retained tip
|
|
```
|
|
|
|
The production export format must preserve:
|
|
|
|
- active branch,
|
|
- active head turn/depth/coordinate,
|
|
- retained alternate/disposable history,
|
|
- checkpoints,
|
|
- state/history/provenance required for recovery.
|
|
|
|
For compatibility with earlier bundles, import may fall back to the retained tip only when no explicit active-head field exists.
|
|
|
|
Export/import regression tests must include an undone campaign and verify the imported story reopens at the exact exported head rather than silently redoing later turns.
|
|
|
|
### 9.1 As implemented in M3
|
|
|
|
The bundle carries the active head depth beside the active branch, and the import
|
|
honors it. This moved the head depth across the format's own rule about what a
|
|
bundle carries: a bundle carries what was *chosen* and recomputes what is
|
|
*derived*, and before M3 the head depth was genuinely derived — the newest row was
|
|
the only place a story could be read. It is a decision now, because the same tree
|
|
exports identically whether the user undid three turns or none, so the file has to
|
|
say.
|
|
|
|
A file that does not state a head is opened at the tip of its active branch. That
|
|
is a fallback only in form: such a file was written when the head could not be
|
|
anywhere else, so deriving the tip reproduces the position it actually recorded.
|
|
Pre-tree bundles take the same path. No format version bump was required, because
|
|
an absent field is unambiguous.
|
|
|
|
The head depth is validated before any row is written — a depth past the branch's
|
|
own retained story is a file disagreeing with itself and is refused, while a depth
|
|
*behind* it is the feature.
|
|
|
|
The bundle also carries which branches the story has left, and at what depth.
|
|
Every row of an abandoned line is exported either way, so without that metadata a
|
|
restored campaign could not distinguish abandoned history from active history —
|
|
which is precisely what a later cleanup or recovery feature has to select on.
|
|
|
|
### 9.2 Save Points in the bundle, as implemented in M4
|
|
|
|
Save Points are exported and imported with the campaign, which is `I04`. They
|
|
fall on the "chosen" side of §9.1's rule without argument: a position someone
|
|
named is not recoverable from the rows, since nothing about a turn records that
|
|
a player once bookmarked it.
|
|
|
|
No format version bump. A bundle written before M4 has no `checkpoints` key and
|
|
imports with none, which is what such a campaign had — the same unambiguous
|
|
absence §9.1 relies on for the head depth, and the same treatment the persona
|
|
block and the branch disposition received.
|
|
|
|
The head and the Save Points are independent, deliberately. An import opens the
|
|
campaign where `headDepth` says, never at a Save Point merely because the file
|
|
carries one: the bundle records where the story was being read and, separately,
|
|
which positions were named, and choosing between them is the user's to make
|
|
after the file is open.
|
|
|
|
A Save Point whose coordinate names no turn in the file is dropped rather than
|
|
refusing the import — the opposite of the head depth's treatment, and for a
|
|
stated reason. A misplaced head affects every read in the file; a bookmark
|
|
pointing outside the story affects only itself, and rejecting a whole campaign
|
|
to protect one bookmark would lose the story to save the pointer.
|
|
|
|
### 9.3 As implemented in M9 — the version, and the third category
|
|
|
|
**The format is now `ai-dnd-adventure-v3`, and the bump is the design.** §9.1 and
|
|
§9.2 each declined one, correctly: an absent `headDepth` or `checkpoints` key is
|
|
unambiguous, because a file either states a position or it does not. That
|
|
property fails for what M9 adds. A v2 file with no prompt provenance may have
|
|
been written before M9, when no file could carry any, or by M9 from a campaign
|
|
whose turns predate the column — different facts about the campaign, and a
|
|
reader has to be able to tell them apart. A version number is how a recovery
|
|
file states what it was capable of recording, which is exactly the reasoning
|
|
§9.1 uses to justify opening a pre-M3 file at its tip. The reader keeps every
|
|
version; only the writer moved.
|
|
|
|
**§9.1's two categories became three.** "Chosen travels, derived is recomputed"
|
|
was sufficient until M9 had to decide about stored prompts, which are derived —
|
|
a machine assembled them — and must travel anyway:
|
|
|
|
```text
|
|
chosen the story, the head, the takes, the Save Points, the
|
|
classifications, the canon travels
|
|
evidence the state events and proposals, the per-turn prompt and the
|
|
passages it was shown, the model and generation settings that
|
|
turn ran under travels
|
|
rebuildable knowledge passages, the FTS index, embeddings, the branch
|
|
lineage cache rebuilt on import
|
|
```
|
|
|
|
The test separating the last two is not "could this be recomputed" but "would a
|
|
recomputation answer the same question". A rebuilt FTS index answers the same
|
|
question. A rebuilt prompt does not — it says what the turn *would be told now*,
|
|
from today's canon, today's sources and today's state, which is the opposite of
|
|
what the inspector is for. Historical evidence is not a cache, so M9 does not
|
|
regenerate one on import at any point.
|
|
|
|
`DATA-MODEL.md` §29 lists what v3 carries. Three implementation facts belong
|
|
here rather than there:
|
|
|
|
1. **The snapshots are encoded, not summarised.** A per-turn prompt contains the
|
|
story so far, so one per turn is O(turns²) — measured at 68% of a 9.7 MB file
|
|
at 120 turns, against a 20 MB import ceiling. The snapshot therefore travels
|
|
as `contextSnapshotZ`, zlib-compressed and base64-encoded through the same
|
|
`compression.pack`/`unpack` the database column already uses. The file is
|
|
still JSON and every other section of it is still plain text. The readable
|
|
`contextSnapshot` key is still accepted and wins when both are present, so a
|
|
hand-edited file keeps importing. A residual ceiling remains and is stated in
|
|
the M9 report rather than hidden.
|
|
2. **One more pointer is translated, and only pointers ever are.** Branch
|
|
numbers already were. A restored retrieval record's `source_id` names a row
|
|
on the machine that wrote the file, so it is repointed at the source that
|
|
landed here, or set to `null` when the file carries no such source. The text
|
|
the record holds — the evidence — is never rewritten.
|
|
3. **Import stays two-phase inside one transaction.** `plan` refuses everything
|
|
a hand-edited file can get wrong before a row exists; `materialize` writes,
|
|
and the endpoint commits once and rolls back explicitly otherwise. A
|
|
*rebuildable* index failing after that does not roll the campaign back: it is
|
|
reported on the response as a warning, shown per source in the Knowledge
|
|
panel, and repaired by Reindex. So a caller sees either "the campaign is not
|
|
there" or "the campaign is complete", never a third thing.
|
|
|
|
### 9.4 The database backup, as implemented in M9
|
|
|
|
A second recovery tool, deliberately not merged with the first. The bundle is a
|
|
logical, portable, human-readable copy of **one campaign** and is the supported
|
|
way to move a campaign between installations; the backup is a physical copy of
|
|
**this machine's whole database** and is what you take before an upgrade.
|
|
|
|
`backend/app/backup.py` uses SQLite's online backup API rather than a file copy,
|
|
because a copy taken while the application runs can read one page before a
|
|
transaction and another after it and produce a file that opens, reports a schema
|
|
and is quietly missing rows. It writes to a temporary name beside the
|
|
destination, runs `PRAGMA quick_check` against the finished file, and only then
|
|
renames it into place; it opens the source read-only, never overwrites an
|
|
existing backup, and leaves nothing behind on failure.
|
|
|
|
No path comes from a caller: the destination is derived from the database the
|
|
application already has open and the filename from the clock, so the endpoints
|
|
accept no body at all (H08).
|
|
|
|
**There is no restore endpoint, and that is a decision.** Restoring means
|
|
replacing the file the running process has open, which is how both copies are
|
|
lost at once. The procedure is in `DEVELOPMENT.md` and is a procedure precisely
|
|
because each step needs the application stopped.
|
|
|
|
## 10. Authoritative Narrative State
|
|
|
|
### 10.1 Do not retain the RPG state protocol as the product model
|
|
|
|
AI-DnD's world-state machinery is useful evidence that state snapshots and rollback are structurally separable from RPG presentation, but the production state model must be genre-neutral.
|
|
|
|
Core concepts include:
|
|
|
|
- entities,
|
|
- facts,
|
|
- relationships,
|
|
- locations,
|
|
- possessions,
|
|
- conditions,
|
|
- organizations,
|
|
- story threads,
|
|
- scene state,
|
|
- chronology where needed.
|
|
|
|
### 10.2 Explicit typed state events
|
|
|
|
Use the ADR 010 model:
|
|
|
|
```text
|
|
model proposes explicit typed operation
|
|
-> schema validation
|
|
-> semantic/referential validation
|
|
-> accepted event(s)
|
|
-> state snapshot/cache
|
|
```
|
|
|
|
Prefer explicit absolute semantics for mutable values.
|
|
|
|
Examples:
|
|
|
|
```text
|
|
set_current_location
|
|
set_entity_status
|
|
set_possession
|
|
add_fact
|
|
invalidate_fact
|
|
add_relationship
|
|
end_relationship
|
|
open_story_thread
|
|
resolve_story_thread
|
|
set_scene
|
|
```
|
|
|
|
Avoid one generic relative-delta protocol whose numeric meaning depends primarily on prompt compliance.
|
|
|
|
### 10.3 Validation limitations
|
|
|
|
Typed events remove the delta/absolute ambiguity but do not guarantee semantic truth.
|
|
|
|
Validation should include deterministic checks where possible:
|
|
|
|
- event type allowlist,
|
|
- schema/type validation,
|
|
- entity/reference existence,
|
|
- impossible transitions where explicitly modeled,
|
|
- authority constraints,
|
|
- conflict handling,
|
|
- transaction integrity.
|
|
|
|
The accepted transcript remains available even if derived state extraction must be retried/repaired according to the final turn-acceptance workflow.
|
|
|
|
### 10.4 Hybrid storage
|
|
|
|
Selected direction:
|
|
|
|
> **validated state events + efficient current/historical snapshots/cache**
|
|
|
|
Events provide audit/reconstruction value. Snapshots/cache make normal reads, Undo/Redo, and context construction fast.
|
|
|
|
**Constraint added by M3 (see ADR 012).** The snapshot half is not an
|
|
optimization to be traded away. M3's head movement is a row lookup plus a
|
|
restore, which is why Undo, Redo and — later — Save Point restore cost the same
|
|
at any distance into a campaign's history. A state model that could only be
|
|
reconstructed by replaying events from the campaign opening would make every one
|
|
of those operations proportional to campaign length, on exactly the long
|
|
campaigns this product exists for. Whatever M5 introduces must keep the
|
|
authoritative state at a retained position efficiently recoverable — a per-node
|
|
snapshot, or an equivalent cache with the same property — while adding the typed
|
|
event model.
|
|
|
|
## 11. Context and Memory
|
|
|
|
Retain AI-DnD's useful lineage-aware memory foundation, but align it with the product authority model.
|
|
|
|
Narrator context is assembled in explicit layers:
|
|
|
|
```text
|
|
Narrator/system rules
|
|
Campaign profile
|
|
Global/explicit canon
|
|
Current authoritative state
|
|
Lineage-safe summary
|
|
Relevant older story memories
|
|
Relevant imported knowledge
|
|
Recent active-lineage turns
|
|
Current user input
|
|
```
|
|
|
|
Requirements:
|
|
|
|
- no abandoned future may appear in active recent history,
|
|
- no memory derived solely from an abandoned future may be retrieved,
|
|
- summaries are anchored to source lineage/turn ranges,
|
|
- derived memory/summary never becomes more authoritative than accepted state/canon,
|
|
- prompt snapshot records what was actually supplied,
|
|
- token budgets remain explicit and inspectable.
|
|
|
|
Phase 0B verified AI-DnD branch-scoped memory isolation against real local embeddings with a negative control. Preserve that property through the history rewrite.
|
|
|
|
## 12. Realistic-Context Model Testing
|
|
|
|
The Phase 0B referee failure appeared under full application context even though the same model followed the state protocol correctly in an isolated probe.
|
|
|
|
Therefore structured-output/state tests must include:
|
|
|
|
- realistic narrator/context length,
|
|
- representative state complexity,
|
|
- actual local models likely to be used,
|
|
- repeated runs rather than one clean prompt,
|
|
- malformed/incorrect semantic proposals,
|
|
- validation and recovery behavior.
|
|
|
|
Model capability recommendations are deferred until these measurements exist; this does not block the architecture.
|
|
|
|
## 13. Imported Knowledge Subsystem
|
|
|
|
Do not turn AI-DnD Story Cards into the production imported-knowledge store.
|
|
|
|
Story Cards may remain a useful reference or authored-rule mechanism, but the imported-knowledge requirements need a separate first-class subsystem with:
|
|
|
|
- source records,
|
|
- `.txt` / `.md` import,
|
|
- Canon / Reference / Inspiration classification,
|
|
- enable/disable/delete,
|
|
- source/version/hash provenance,
|
|
- chunk records,
|
|
- campaign scoping,
|
|
- local lexical index (prefer SQLite FTS5),
|
|
- local Ollama embeddings/semantic index where enabled,
|
|
- authority-aware hybrid retrieval,
|
|
- retrieval provenance,
|
|
- export/import support,
|
|
- no automatic URL/image fetching,
|
|
- imported content treated as data, never executable instructions.
|
|
|
|
If a future knowledge item is derived from story history rather than imported as global campaign material, it must carry lineage/source-turn information sufficient to avoid abandoned-path leakage.
|
|
|
|
### 13.1 As implemented in M7
|
|
|
|
Every item above is built, in `backend/app/knowledge/`. Story Cards were not
|
|
promoted into it and are untouched. The pipeline, and where each decision lives:
|
|
|
|
```text
|
|
upload (multipart; no pathname is ever accepted)
|
|
-> validate size, strict UTF-8, real text, allowed extension, class
|
|
-> hash SHA-256 of the normalized text; the duplicate test
|
|
-> store the text in SQLite, under application control
|
|
-> chunk deterministic, heading-aware, 60-800 tokens
|
|
-> index SQLite FTS5, porter-stemmed
|
|
---- one transaction ends here; the source is now `ready` ----
|
|
-> embed local Ollama, best-effort, through the shared provider
|
|
```
|
|
|
|
```text
|
|
query built from the head-capped story tail and the authoritative state
|
|
-> FTS5 lexical candidates (LIMIT in SQL)
|
|
+ semantic candidates (when an embedding model is configured)
|
|
-> ADMISSION, absolute and per path:
|
|
semantic raw cosine >= SEMANTIC_FLOOR
|
|
lexical >= 2 distinct meaningful terms, or 1 that is neither a
|
|
standing entity nor a negligible share of the query
|
|
a passage needs evidence from at least one path, or it is discarded
|
|
-> RANKING, among survivors only:
|
|
normalize each score against the best surviving value of its own path
|
|
relevance = max(lex, sem) + 0.15 x min(lex, sem)
|
|
score = relevance x class weight (canon 1.00, ref 0.85, insp 0.70)
|
|
-> suppress redundancy, never across classes, before the budget cut
|
|
-> fill Canon, then Reference, then Inspiration, each against a cap
|
|
-> render with class framing and per-passage provenance
|
|
```
|
|
|
|
### 13.2 Relevance admission is a separate stage from ranking
|
|
|
|
**This is M7's most expensive lesson and it generalises beyond knowledge
|
|
retrieval.** M7 shipped with only a ranking stage: both scores were normalized
|
|
against the best candidate of their own path, and the relevance floor was
|
|
expressed as a share of that best. A floor defined as a share of the best is
|
|
structurally incapable of rejecting anything, because the best candidate clears
|
|
a share of itself by construction. With the semantic path scoring every embedded
|
|
chunk there was always a best, so **something was admitted on every turn**
|
|
whatever the reader was doing — a query about tide tables and container tonnage
|
|
retrieved all five sources of a fantasy campaign, narrator-only hidden Canon
|
|
among them.
|
|
|
|
The rule that follows:
|
|
|
|
> A relevance decision must be made on a signal that means something on its own.
|
|
> Normalization answers "which of these is best"; it can never answer "is any of
|
|
> these any good". A pipeline that ranks first and cuts second has no way to
|
|
> return nothing.
|
|
|
|
So the two stages are separated, and they consume different quantities:
|
|
|
|
- **Admission** reads the *raw* signals — the cosine the model returned, and how
|
|
many distinct meaningful query terms a passage contains. Neither is computed
|
|
by comparison with the other candidates.
|
|
- **Ranking** reads the *normalized* signals, because `bm25` has no fixed range
|
|
and cosine's zero is not zero, so the two paths are not otherwise comparable.
|
|
It decides order among things that matched.
|
|
|
|
Authority is applied in the second stage only. That is what makes
|
|
`IMPORTED-KNOWLEDGE-DESIGN.md` §30's two consecutive sentences —
|
|
`Canon > Reference > Inspiration`, and "do not include irrelevant Canon merely
|
|
because it is authoritative" — compatible rather than contradictory: the class
|
|
orders what matched and can never rescue what did not.
|
|
|
|
An absolute threshold on an embedding similarity is a property of the model, not
|
|
of the product, so it is measured, written down beside the constant, and
|
|
re-measured by a real-model test on every run that has one — the same discipline
|
|
`memorybank.REDUNDANT_SIMILARITY` already follows.
|
|
|
|
### 13.3 Semantic admission is calibrated per embedding model
|
|
|
|
**Semantic admission is calibrated for `nomic-embed-text`; uncalibrated
|
|
embedding models fall back safely rather than borrowing its threshold.**
|
|
|
|
The threshold is therefore **not portable**, and the two ways a different model
|
|
can break it are not symmetric. A model whose similarity scale sits *below* the
|
|
calibrated one admits nothing and degrades to lexical-only, which is a supported
|
|
path. A model whose scale sits *above* it would put unrelated material past the
|
|
threshold and reproduce the M7-F1 defect on a build whose tests all pass.
|
|
|
|
So the product does not apply a threshold to a model it has not measured:
|
|
|
|
```text
|
|
SEMANTIC_CALIBRATION = {"nomic-embed-text": 0.58}
|
|
|
|
calibrated model -> semantic admission at its measured floor
|
|
uncalibrated model -> semantic retrieval skipped entirely, reason reported,
|
|
retrieval degrades to lexical-only
|
|
```
|
|
|
|
The model's identity is the one already stored on each vector row, so no second
|
|
mechanism was introduced, and an uncalibrated configuration reports
|
|
`semantic_enabled: false` rather than claiming a semantic index that is never
|
|
consulted. Adding a model is a measurement — run the real-model retrieval test
|
|
against it and confirm the targeted and off-topic populations separate — not a
|
|
guess. Generic cross-model calibration is out of scope for v1.
|
|
|
|
The cost is stated rather than hidden: under an uncalibrated model a
|
|
conceptual-only paraphrase is not retrieved. That is a missing passage rather
|
|
than an irrelevant one, which is the direction this product prefers to fail in.
|
|
|
|
Four decisions are worth recording, because each replaced an obvious wrong one:
|
|
|
|
- **The class multiplies relevance; it does not add to it.** An additive class
|
|
bonus satisfies "Canon outranks Reference" and makes "do not include
|
|
irrelevant Canon" impossible, because a large enough constant wins alone.
|
|
- **Both retrieval scores are normalized per query, against the best of their
|
|
own path — for ranking only.** `bm25` has no fixed range; cosine's zero is not zero, and a real
|
|
embedding model scores any two pieces of English around 0.3-0.6. Blended raw,
|
|
a lexical hit beats every semantic hit on every query.
|
|
- **Admission does not use those normalized values at all.** A normalized score
|
|
cannot express "no match", which was M7's blocking defect; the correction is
|
|
the two-stage separation §13.2 records.
|
|
- **Lexical retrieval is a production path**, not a fallback. It is what finds
|
|
proper nouns and invented terms — most of what a setting bible is made of —
|
|
and the library is fully usable with no embedding model at all.
|
|
|
|
Abandoned-path safety is met at the query rather than by a lineage coordinate on
|
|
the source, because an imported file has no lineage: the query is built from
|
|
`context.history.tail`, which reads through the head-capped clause, and from
|
|
`adventures.narrative_state`, which head movement repoints. Nothing reads the
|
|
uncapped action table.
|
|
|
|
## 13.4 The browser is a presentation layer, and M8 kept it one
|
|
|
|
M8 rebuilt the interface without moving a decision into it. Three rules made
|
|
that hold, and each replaced a tempting shortcut:
|
|
|
|
- **Server-authoritative availability.** Whether Undo and Redo have anywhere to
|
|
go is the server's answer, carried on every window it returns. Neither is
|
|
derivable in the browser: Undo can reach past the top of the loaded page, and
|
|
Redo depends on a retained future the transcript is never sent. A React
|
|
component that computed either would be wrong exactly when it mattered.
|
|
- **Reading is not deciding.** Stepping between alternate takes tells the server
|
|
nothing. The decision is made by writing below one, and that is the only
|
|
moment `after_id` is sent. This is why the take pager can be offered on every
|
|
turn without any of them becoming a commitment.
|
|
- **UI state stays UI state** (§92 of `BROWSER-UX-SPEC.md`). Which panel is
|
|
open, whether narrator-only text is revealed, whether a source's detail is
|
|
expanded — none of it is written anywhere, and none of it changes what is
|
|
retrieved, what is prompted, or what the story believes.
|
|
|
|
### Failure has a taxonomy, not a toast
|
|
|
|
`frontend/src/errors.js` sorts a failure into the five kinds §71 names — model,
|
|
generation, state, knowledge, server — because each one has a different thing to
|
|
*do* about it. It reads the message the server actually sent, which is a
|
|
coupling to backend strings, so it is written as a fallback ladder rather than a
|
|
lookup: an unrecognised message still gets a kind, still shows the server's own
|
|
words, and still offers Retry. Nothing is hidden when the match misses.
|
|
|
|
### Markup is never produced from input
|
|
|
|
`frontend/src/markdown.jsx` renders narrator prose, and it is where H06 and H07
|
|
are decided for rendered text. The safety is structural rather than filtered:
|
|
every node it returns is a React element built from parsed text, the text only
|
|
ever becomes a React child, and there is no `dangerouslySetInnerHTML` in the
|
|
file. A sanitizer is not needed to make markup safe if markup is never produced.
|
|
|
|
Link schemes are checked with the URL parser rather than a pattern, because the
|
|
bypasses are all in the parsing — `java\tscript:`, `JaVaScript:` and
|
|
`%6a%61vascript:` are one URL to a browser and three strings to a regex.
|
|
|
|
The knowledge and context panels deliberately do **not** use this renderer.
|
|
They exist to show a reader exactly what is in their file, and rendering is the
|
|
opposite of that; imported text goes into a `<pre>` as a text node.
|
|
|
|
### Campaign canon is configuration, and its provenance is per turn
|
|
|
|
M8 gave `campaign_canon` an API for the first time (`canon_rules`), which raised
|
|
a question the column had never had to answer: what happens when a reader edits
|
|
the campaign's highest authority *after* a story exists?
|
|
|
|
Measured, against a real server with real turns:
|
|
|
|
- **Every turn already played keeps the canon it was actually given.** The canon
|
|
section is part of that turn's stored context snapshot, so a historical turn
|
|
inspected after the edit still shows the old rule and not the new one. This is
|
|
M7's principle — provenance is the rendered text, not a foreign key — doing
|
|
the work without being asked.
|
|
- **The next turn is told the new canon**, which is the point of editing it.
|
|
- **Nothing recorded is rewritten**: accepted story text, the narrative state
|
|
document and the state audit log are all byte-identical across the edit.
|
|
- **The edit itself is not audited**, because canon is *configuration*. It sits
|
|
with `ai_instructions` and the narrator prompt, none of which are audited
|
|
either, and unlike a manual state correction it asserts nothing about the
|
|
story — it changes what the narrator is told from that point on.
|
|
|
|
So M5's `StateEvent` log is deliberately **not** extended to cover it. That log
|
|
audits accepted changes to narrative state; canon is not narrative state, and
|
|
routing a configuration change through it would create a second representation
|
|
of canon — the same duplication `BROWSER-UX-SPEC.md` §38 forbids for hidden
|
|
information, arrived at from a different direction.
|
|
|
|
What M8 added instead is at the editing surface: once a campaign has moments,
|
|
the canon editor says that the change applies from here on, that everything
|
|
already written stays as it is, and where the per-turn record can be seen. The
|
|
guarantee is not "the edit is logged" but "the edit cannot be mistaken for a
|
|
retroactive one, and every turn can prove what it was told".
|
|
|
|
## 14. Prompt and Provenance Inspection
|
|
|
|
Preserve and extend AI-DnD's Insights/context-snapshot capability.
|
|
|
|
For each narrator turn the system should be able to explain:
|
|
|
|
- narrator/system rules used,
|
|
- campaign/canon context,
|
|
- current authoritative state included,
|
|
- summary included,
|
|
- story memories retrieved,
|
|
- knowledge chunks retrieved,
|
|
- recent history included,
|
|
- user input,
|
|
- model/settings,
|
|
- state proposal,
|
|
- validation result,
|
|
- accepted events,
|
|
- source IDs/turn ranges where applicable.
|
|
|
|
## 15. Scene and Future Media Boundary
|
|
|
|
v1 does not require media generation.
|
|
|
|
It does require preserving scene/entity information so future providers do not have to infer continuity from the entire raw transcript.
|
|
|
|
Persist or derive a scene snapshot containing relevant fields such as:
|
|
|
|
- location,
|
|
- participants,
|
|
- significant objects,
|
|
- current actions,
|
|
- time/lighting/environment,
|
|
- mood,
|
|
- visual character/location profiles,
|
|
- continuity constraints,
|
|
- source turn range and lineage.
|
|
|
|
Future media coordinator consumes a normalized scene packet and records local asset provenance.
|
|
|
|
The story engine must remain fully functional with media disabled.
|
|
|
|
STT specifically follows:
|
|
|
|
```text
|
|
microphone/audio -> local STT -> editable draft -> normal user submission
|
|
```
|
|
|
|
STT never bypasses the ordinary authoritative story commit path.
|
|
|
|
### 15.1 As implemented (M10)
|
|
|
|
The boundary above is built. What follows is what it turned out to be.
|
|
|
|
**The scene snapshot was already there.** §15 says "persist or derive", and the
|
|
answer is *derive*, because M5 had persisted it three milestones earlier:
|
|
`narrative_state["scene"]` holds `summary`, `location`, `present[]` and the
|
|
`at: {branch_id, depth}` coordinate, written by the validated `set_scene` event
|
|
and snapshotted per position. It already restores correctly through Undo, Redo,
|
|
Retry, Save Point restore and divergence, because the head move restores the
|
|
whole state document and the scene is part of it. A second scene store would
|
|
have had to reimplement all of that, and would have been a second answer to
|
|
"where is the story now".
|
|
|
|
So the seam is three modules under `app/media/`, and only one of them has a
|
|
table:
|
|
|
|
```text
|
|
app/media/packet.py the normalized scene packet §15 names, built on read
|
|
app/media/profiles.py visual character/location profiles — the one field in
|
|
§15's list that nothing already stored
|
|
app/media/providers.py the contracts a future coordinator implements
|
|
```
|
|
|
|
`GET /api/adventures/{id}/scene-packet` returns the packet; the four
|
|
`visual-profiles` endpoints read and write profiles. Nothing else in the
|
|
application calls either — no turn, no prompt, no context section.
|
|
|
|
**What the packet contains, and the rule behind it.** Location, characters
|
|
present with their profiles, significant objects, an action summary, continuity
|
|
constraints, ambience, the source turn range and the lineage coordinate. What it
|
|
deliberately excludes is the more interesting half: the raw transcript, all
|
|
imported knowledge, memories and summaries. The rule is *what the story
|
|
established at this position*, not *everything the narrator was told* — which is
|
|
what keeps a hidden Canon source out of a depiction without needing a filter
|
|
that someone has to remember to apply to each new secret.
|
|
|
|
**Story engine functional with media disabled** is the milestone's central
|
|
acceptance condition rather than a footnote, and it is tested as one
|
|
(`backend/tests/test_m10_no_media.py`): a whole campaign — turns, state
|
|
extraction, memory and summary activity, knowledge retrieval, Undo, Redo, Retry,
|
|
Save Point restore, restart — with an empty provider registry, no media setting
|
|
in existence, and no media row written. Media readiness is inert until something
|
|
uses it, and nothing does yet.
|
|
|
|
**STT asymmetry.** The `TranscriptionProvider` contract returns a
|
|
`DraftTranscription` with `editable: bool = True` and no commit method, so the
|
|
diagram above is enforced by the shape of the interface: a transcriber can
|
|
produce a draft and cannot submit one. The ordinary authoritative commit path is
|
|
the only way in.
|
|
|
|
**Endpoint policy.** A future media provider endpoint is checked by
|
|
`providers.endpoint_rejection_reason`, which reuses the local-only policy in
|
|
`app/endpoints.py` and then requires loopback in addition — stricter than
|
|
narrator inference, which permits a trusted LAN host. A GPU that renders a
|
|
reader's campaign is a machine that reader is sitting at. No provider
|
|
configuration setting exists to point anywhere, because none is needed yet.
|
|
|
|
### 15.2 The inference window is a ceiling, not an assumption (M11)
|
|
|
|
M8 measured a reference deployment enforcing a **4,096**-token input window while
|
|
the application budgeted **16,384**, and every request returned HTTP 200. The
|
|
consequence is worse than an error: `llama.cpp` drops the *oldest* tokens, and
|
|
the oldest tokens in this design are the system block — the narrator's rules and
|
|
the campaign canon. A long campaign would quietly stop obeying its own canon,
|
|
and every acceptance test that reads a 200 as success would keep passing.
|
|
|
|
M11's rule:
|
|
|
|
> The application must not silently budget more narrator input than the
|
|
> configured Ollama runtime will actually accept.
|
|
|
|
`app/contextwindow.py` asks the server, on the same host and under the same
|
|
endpoint policy as inference: `/api/ps` reports the window a **loaded** model is
|
|
being served with, and `/api/show` reports the `num_ctx` an unloaded one will
|
|
load with plus the architecture's ceiling. The answer is cached per endpoint and
|
|
model, so it costs one short request per session rather than one per turn, and
|
|
it is cleared when either changes.
|
|
|
|
The builder takes the window as a parameter — like retrieved memories and
|
|
imported passages, and for the same reason: the prompt builder makes no network
|
|
calls. A **verified** window is a ceiling on `context_token_budget`; an
|
|
**unverified** one leaves the configured budget standing and is recorded as
|
|
unverified in the turn's stored provenance, on the context report, and on the
|
|
connection test. There is no third behaviour, and in particular there is no
|
|
hard-coded 4,096: guessing a number the server did not say would be right on one
|
|
machine and wrong on the next.
|
|
|
|
What this does not do is change the window. That is an operator action — a model
|
|
with `num_ctx` baked in, or `OLLAMA_CONTEXT_LENGTH` — and `DEVELOPMENT.md` says
|
|
how. What the application owes the reader is not to lie about it.
|
|
|
|
**As implemented: the server that cannot be asked.** Discovery above speaks
|
|
Ollama's *native* API, and nothing restricts `endpoint_url` to Ollama — any
|
|
allowed address serving an OpenAI-compatible `/v1` is accepted. On vLLM, on
|
|
llama.cpp's own server, on anything else, `/api/ps` and `/api/show` are not
|
|
there. Discovery fails as designed and the window is unverified, which is honest
|
|
and leaves the rule above unenforced: the budget stands, and a server with a
|
|
smaller window drops the oldest tokens exactly as before. The rule names Ollama
|
|
because Ollama is what M8 measured; the failure it forbids is not Ollama's.
|
|
|
|
`Settings.context_window_override` is the operator stating the window because
|
|
they know how they launched the server. It is consulted **only where discovery
|
|
left a hole**, in this order:
|
|
|
|
| Discovery | Override | Result |
|
|
| --- | --- | --- |
|
|
| verified | any | the verified window; a declaration cannot raise it |
|
|
| unverified | set | the declared number, source `declared`, and the prompt is capped |
|
|
| unverified | unset | unknown, and the configured budget stands |
|
|
|
|
The first row is the safety property: an operator may lower an unknown ceiling
|
|
into existence and may never raise a known one, so the override cannot become a
|
|
route back to over-budgeting a server that already answered.
|
|
|
|
`verified` keeps its narrow meaning — *the server answered* — so
|
|
`window_verified` in a turn's provenance still counts what §15.2 says it counts
|
|
and a declaration cannot inflate it. `enforceable` is the separate question of
|
|
whether there is a number to cap to at all, and that is what the builder and the
|
|
`capped` flag use. A declared window is therefore enforced and identifiable as a
|
|
declaration everywhere it appears, and the connection test says plainly that
|
|
nothing has checked it against the server.
|
|
|
|
### 15.3 The history window moves in blocks (post-M11)
|
|
|
|
§15.2 makes the window a ceiling. This is about what happens at that ceiling.
|
|
|
|
An inference server caches a prompt by its **prefix**. A history window that
|
|
gives up its oldest action every turn changes the prompt near the front, which
|
|
discards the cache and makes the server re-read nearly all of it every turn. The
|
|
builder's window did exactly that, and it is why a long campaign cost roughly the
|
|
same per turn however little had changed since the last one.
|
|
|
|
`builder.history_floor` snaps the oldest included action's `depth` forward to a
|
|
multiple of `builder.trim_block` and holds it there. The window then steps: a
|
|
run of turns that re-use the cache, then one that pays to re-read.
|
|
|
|
Two properties make it safe rather than merely fast, and both are pinned by
|
|
tests:
|
|
|
|
- **It only ever drops more.** The kept window is a suffix of what "whatever
|
|
fits" would have kept, so §15.2's ceiling and M03's bound are unweakened.
|
|
- **The block comes from configuration, not from the story.** `trim_block` is
|
|
derived from the history budget and `max_output_tokens`. A block size measured
|
|
from the sizes of recent actions would move the boundary it defines, and a
|
|
boundary that moves is the thing this exists to stop.
|
|
|
|
`depth` is the coordinate because it is stable per action and already branch-
|
|
scoped. Rows without one — anything predating the story tree — are not trimmed,
|
|
and behave exactly as they did.
|
|
|
|
The cost is history depth: right after a step the window holds up to a block
|
|
fewer actions than the budget allows. `TRIM_FRACTION` bounds that at a quarter of
|
|
the window and is the dial between recent history and speed. Measured at an
|
|
8,192-token budget: 124.0s per turn against 362.4s with the floor disabled, the
|
|
saving growing with the block and therefore with the budget.
|
|
|
|
**This is a performance change and nothing above it is a requirement.** M11 §P.1
|
|
records that no performance requirement exists and declines to invent one; that
|
|
still holds. What changed is the cost of a turn, not what a turn must contain.
|
|
|
|
## 16. Database Direction
|
|
|
|
SQLite remains the selected v1 authoritative store.
|
|
|
|
Reasons:
|
|
|
|
- already present in the selected base,
|
|
- local and single-user friendly,
|
|
- transactional,
|
|
- portable,
|
|
- supports FTS5,
|
|
- compatible with backup/export tooling,
|
|
- no external service required.
|
|
|
|
Remove Postgres/Neon support from the production fork unless a later explicit requirement reverses this decision.
|
|
|
|
The exact physical schema may evolve through migrations; the conceptual model is in `DATA-MODEL.md`.
|
|
|
|
## 17. Transaction Boundaries
|
|
|
|
Where practical, one accepted turn should atomically establish:
|
|
|
|
- accepted user input/narration relationship,
|
|
- turn/lineage identity,
|
|
- active-head advancement,
|
|
- validated authoritative state events,
|
|
- resulting state snapshot/cache,
|
|
- core prompt/model provenance needed for recovery/audit.
|
|
|
|
Derived work such as embeddings, memory extraction, summary generation, and future media jobs may occur separately, but failure must not corrupt the authoritative commit.
|
|
|
|
## 18. Testing Strategy
|
|
|
|
Use AI-DnD's inherited tests as a foundation, then rewrite/add tests around product semantics.
|
|
|
|
Required categories:
|
|
|
|
- offline startup/story use,
|
|
- no runtime remote assets/tokenizer fetch,
|
|
- same-host and trusted-LAN Ollama endpoint handling,
|
|
- turn persistence/restart,
|
|
- failed generation atomicity,
|
|
- non-destructive Undo/Redo,
|
|
- divergence after Undo,
|
|
- retry/take retention,
|
|
- named checkpoint restore,
|
|
- branch/lineage state reconstruction,
|
|
- abandoned-history memory/summary isolation,
|
|
- active-head export/import round trip,
|
|
- generic narrative state event validation,
|
|
- realistic-context structured state extraction,
|
|
- imported knowledge authority/provenance/isolation,
|
|
- prompt/context inspection,
|
|
- fantasy + science-fiction genre neutrality,
|
|
- 100-turn/long-run acceptance,
|
|
- future-media schema compatibility.
|
|
|
|
Acceptance-test IDs in `V1-ACCEPTANCE-TESTS.md` are the black-box release contract.
|
|
|
|
### 18.1 Wiring rule, from the M2 regressions
|
|
|
|
M2 shipped two defects that a 604-test green suite did not see: a removed
|
|
`Settings` attribute left two provider factories raising `AttributeError` inside
|
|
a background task, and a newly added timeout setting was stored, validated,
|
|
exposed and rendered without ever being passed to the provider that needed it.
|
|
Both were invisible because the tests at that boundary were mocks.
|
|
|
|
> **When removing a setting, attribute or dependency, test at least one real
|
|
> consumer construction path. When adding a setting, test that the configured
|
|
> value reaches the component that uses it. A green suite built entirely around
|
|
> mocks at that boundary is insufficient evidence.**
|
|
|
|
The corollary is where to look: subtractive changes and plumbing changes fail in
|
|
background and fire-and-forget paths, which are exactly the paths that report
|
|
nothing when they break.
|
|
|
|
|
|
## 19. Removal / Migration Strategy From Upstream
|
|
|
|
Production migration should be incremental and test-gated rather than a broad rewrite.
|
|
|
|
Remove or replace in controlled milestones:
|
|
|
|
1. runtime external dependency leaks,
|
|
2. hosted/multi-user/auth/demo/analytics/cloud/Postgres/QuickJS surfaces,
|
|
3. destructive Undo/no-Redo behavior,
|
|
4. RPG-specific state/referee protocol and UI assumptions,
|
|
5. Story Card assumptions where they conflict with the new knowledge subsystem.
|
|
|
|
Preserve upstream provenance and license notices.
|
|
|
|
Do not mechanically merge the ai-adventure or Open Dungeon repositories into the fork.
|
|
|
|
## 20. Deferred Questions That Do Not Block v1 Architecture
|
|
|
|
The following are implementation/release measurements, not unresolved foundational choices:
|
|
|
|
- which narrator/state models should be recommended to users,
|
|
- exact embedding model recommendation,
|
|
- performance of multi-hour stories,
|
|
- concurrency beyond the single-user turn lock,
|
|
- which future local image/video/TTS/STT provider is selected,
|
|
- abandoned-history cleanup policy/UI after v1.
|
|
|
|
## 21. Technical Design v1.0 Exit Status
|
|
|
|
Phase 0 has resolved the foundational choices required for v1.0:
|
|
|
|
- base repository selected,
|
|
- browser/backend stack selected,
|
|
- SQLite selected,
|
|
- non-destructive history/head model demonstrated,
|
|
- typed narrative-state-event direction selected,
|
|
- Ollama validated on local infrastructure, including the required trusted-LAN deployment mode,
|
|
- memory lineage behavior validated,
|
|
- imported-knowledge architecture selected,
|
|
- local-only hardening scope identified,
|
|
- export/import active-head defect understood,
|
|
- future media boundary retained,
|
|
- production milestone sequence defined in `BUILD-MILESTONES.md`.
|
|
|
|
This technical design is therefore the implementation baseline unless revised by a later ADR.
|