Files
interactive-story/planning/VERSION.md
T
JesseMarkowitzandClaude Opus 5 2fdd2547f0 Planning: record M2 closeout decisions
M2's review reported six planning recommendations rather than applying them,
three marked before M3. All six are applied here, plus three additions drawn
from the same evidence. No implementation file is touched.

The endpoint policy was the gap that mattered. It is the most consequential
setting in the application — the storyteller sends the player's prose, the
context, the memories and the embedding inputs to whatever address it names —
and it existed only as a module docstring. It is now ADR 011 and a new §10A in
the threat model, which also retires the assumption in §71A that the inherited
guard was a starting point. It was not: AI-DnD's SSRF guard blocked private
addresses to stop a hosted server reaching its own internal network, which is
the exact opposite of what a local storyteller needs. It was removed, not
adapted.

Both documents state the rule as implemented — an allowlist of explicit
local-network CIDRs, every resolved address checked, enforced on save and again
before every outbound request, TLS never traded against it — and both state the
two residual limits plainly rather than implying they are covered: a hostile
host already on the trusted LAN is inside the permitted boundary, and a
rebinding interval exists between the policy's resolution and the client's
connection. Accepted risks, not M3 work.

The CIDRs are spelled out rather than derived from is_private/is_reserved, and
the ADR records why: is_private is true of the documentation ranges and
0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it
refuses an ordinary same-host Ollama on [::1].

TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening
items. A new §5.2 records the M1/M2 architecture as fact rather than intention,
so later milestones inherit what the code does. A new §18.1 carries the lesson
of M2's two regressions: when removing a setting, test a real consumer
construction path; when adding one, prove it reaches the component that uses
it. Both defects hid behind a green suite because the tests at that boundary
were mocks.

BUILD-MILESTONES records M2 complete, with the capabilities later milestones
inherit and the debt carried forward. Two notes go to milestones that would
otherwise misread what M2 left them. M5 is told that eight rollback tests now
use the world-state engine as instrumentation and not as endorsement — the
instrumentation moves when the protocol does, and those tests are reworked
rather than deleted. M6 is told that the memory bank died silently under a
green suite, so background failure must be observable and at least one real
provider-construction path must be tested.

The security contract gains what M2 demonstrated. H10 now names the two
conditions that were defects during M2: a wildcard origin must be refused at
startup, and an unknown /api path must 404 rather than returning the SPA with
200. New H12 covers endpoint enforcement, and its fourth pass condition is the
one that matters — a public endpoint written into the database behind the
settings API must still be refused at the wire. A build passing the first three
and failing that one has configuration validation only.

SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement;
it removed capability the specification never asked for.

The two M2 reports gain appended closeout notes rather than edits. Their
original wording about an uncommitted working tree was true when written, and
the note records what happened afterwards: the six-file correction is 8652fe7,
8c65ae9 remains the implementation commit, and the two were never squashed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
2026-09-03 01:52:03 -04:00

4.5 KiB

Planning Package Version

Package: Adventure Storyteller Planning Package v2.2
Revision date: 2026-09-03
Status: Phase 0 complete; architecture selected; Milestones M1 and M2 implemented and accepted; M3 not yet briefed.

v2.2 — Post-M2 Closeout (2026-09-03)

M2 removed the hosted, cloud, account and scripting surface and added the inference endpoint policy. Its review recommended six planning changes and reported rather than applied them; all six are applied in this revision, listed in README.md § Post-M2 corrections applied, with the evidence in reports/M2-BASELINE-REPORT.md and reports/M2-IMPLEMENTATION-REPORT.md.

In summary:

  • the inference endpoint policy is recorded as implemented — an address allowlist of explicit local-network CIDRs, every resolved address checked, enforced when settings are saved and again before every outbound request, with TLS verification never traded against it (new ADR 011, SECURITY-THREAT-MODEL.md §10A),
  • its two residual limits are stated rather than mitigated: a hostile host already on the trusted LAN, and the DNS-rebinding interval between the policy's resolution and the client's connection,
  • TECHNICAL-DESIGN.md §5.1 items 3 and 4 are resolved, and a new §5.2 records the M1/M2 production architecture as fact,
  • a wiring rule is added (§18.1): removing a setting requires testing a real consumer path, and adding one requires proving it reaches its component — M2 shipped two defects behind a 604-test green suite because the tests at that boundary were mocks,
  • BUILD-MILESTONES.md records M2 complete, warns M5 that eight rollback tests use the world-state engine as instrumentation rather than as architecture, and requires M6 to make background memory failure observable,
  • the security acceptance contract is strengthened: H10 now names the wildcard-origin and /api 404 conditions, and new H12 covers inference endpoint enforcement including the database-edited-behind-the-API case.

SPECIFICATION.md is unchanged: M2 altered no product requirement. Nothing in the architecture selected in v2 was reversed.

v2.1 — Post-M1 Corrections (2026-09-02)

M1 implementation evidence contradicted or under-specified parts of v2. The corrections are recorded in the documents themselves and listed in README.md § Post-M1 corrections applied; the evidence behind them is in reports/M1-BASELINE-REPORT.md and reports/M1-IMPLEMENTATION-REPORT.md.

In summary:

  • a trusted-LAN Ollama may be HTTPS with a privately issued certificate; clients verify against the operating system's CA store, with full certificate and hostname verification and no bypass option (ADR 002, TECHNICAL-DESIGN.md §5, A06),
  • offline claims require a fresh cache and no route out to be evidence at all, and vendored runtime artifacts should be integrity-verifiable (ADR 004),
  • A05's invariant is about accepted history; the user's submitted text is deliberately retained on a failed turn,
  • A06 requires a real second machine and an HTTPS endpoint; a plain-HTTP LAN test is no longer sufficient evidence,
  • the standard test environment records CPU/GPU/RAM, because cold model load on a CPU-only host exceeded the inherited 120 s client timeout,
  • BUILD-MILESTONES.md records M1 as complete and reframes M2's endpoint work as narrowing an existing capability rather than inventing it,
  • SECURITY-THREAT-MODEL.md §53 distinguishes inbound TLS (still deferred) from outbound certificate verification (required, done in M1).

Nothing in the architecture selected in v2 was reversed.

v2 — Post Phase 0B Revision (2026-09-01)

This v2 package supersedes the earlier planning package produced before the final Phase 0B review and the trusted-LAN Ollama deployment clarification.

Key v2 changes include:

  • AI-DnD selected as the production base at the pinned Phase 0B commit.
  • Non-destructive head-cursor Undo/Redo design selected.
  • Explicit typed/absolute narrative-state events selected for production state handling.
  • Imported knowledge separated from AI-DnD Story Cards.
  • Trusted-LAN Ollama inference supported in v1 while the storyteller UI/API remains loopback-bound by default.
  • Offline first-use dependencies and runtime remote assets identified as M1 hardening work.
  • Production implementation divided into milestones M1-M11.

Historical Phase 0 prompts/reports are retained as evidence and should not be treated as current implementation instructions unless a current milestone prompt explicitly refers to them.