Planning: record M2 closeout decisions
M2's review reported six planning recommendations rather than applying them, three marked before M3. All six are applied here, plus three additions drawn from the same evidence. No implementation file is touched. The endpoint policy was the gap that mattered. It is the most consequential setting in the application — the storyteller sends the player's prose, the context, the memories and the embedding inputs to whatever address it names — and it existed only as a module docstring. It is now ADR 011 and a new §10A in the threat model, which also retires the assumption in §71A that the inherited guard was a starting point. It was not: AI-DnD's SSRF guard blocked private addresses to stop a hosted server reaching its own internal network, which is the exact opposite of what a local storyteller needs. It was removed, not adapted. Both documents state the rule as implemented — an allowlist of explicit local-network CIDRs, every resolved address checked, enforced on save and again before every outbound request, TLS never traded against it — and both state the two residual limits plainly rather than implying they are covered: a hostile host already on the trusted LAN is inside the permitted boundary, and a rebinding interval exists between the policy's resolution and the client's connection. Accepted risks, not M3 work. The CIDRs are spelled out rather than derived from is_private/is_reserved, and the ADR records why: is_private is true of the documentation ranges and 0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it refuses an ordinary same-host Ollama on [::1]. TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening items. A new §5.2 records the M1/M2 architecture as fact rather than intention, so later milestones inherit what the code does. A new §18.1 carries the lesson of M2's two regressions: when removing a setting, test a real consumer construction path; when adding one, prove it reaches the component that uses it. Both defects hid behind a green suite because the tests at that boundary were mocks. BUILD-MILESTONES records M2 complete, with the capabilities later milestones inherit and the debt carried forward. Two notes go to milestones that would otherwise misread what M2 left them. M5 is told that eight rollback tests now use the world-state engine as instrumentation and not as endorsement — the instrumentation moves when the protocol does, and those tests are reworked rather than deleted. M6 is told that the memory bank died silently under a green suite, so background failure must be observable and at least one real provider-construction path must be tested. The security contract gains what M2 demonstrated. H10 now names the two conditions that were defects during M2: a wildcard origin must be refused at startup, and an unknown /api path must 404 rather than returning the SPA with 200. New H12 covers endpoint enforcement, and its fourth pass condition is the one that matters — a public endpoint written into the database behind the settings API must still be refused at the wire. A build passing the first three and failing that one has configuration validation only. SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement; it removed capability the specification never asked for. The two M2 reports gain appended closeout notes rather than edits. Their original wording about an uncommitted working tree was true when written, and the note records what happened afterwards: the six-file correction is8652fe7,8c65ae9remains the implementation commit, and the two were never squashed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
This commit is contained in:
co-authored by
Claude Opus 5
parent
8652fe7cd8
commit
2fdd2547f0
@@ -1,6 +1,6 @@
|
||||
# Adventure Storyteller — Production Build Milestones
|
||||
|
||||
**Status:** In implementation. M1 complete (2026-09-02); M2 next
|
||||
**Status:** In implementation. M1 and M2 complete (2026-09-02); M3 next
|
||||
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
||||
|
||||
## 1. Purpose
|
||||
@@ -166,6 +166,37 @@ Add:
|
||||
|
||||
The codebase has a narrow single-user/local-only surface and the inherited story foundation still passes its relevant regression suite.
|
||||
|
||||
## Status: COMPLETE
|
||||
|
||||
Accepted 2026-09-02. Evidence: `planning/reports/M2-BASELINE-REPORT.md`
|
||||
(measurements) and `planning/reports/M2-IMPLEMENTATION-REPORT.md` (review);
|
||||
verdict *accept with non-blocking debt, proceed to M3*. Implementation is
|
||||
commit `8c65ae9`, and the three defects the review found are commit `8652fe7`
|
||||
— see the closeout note appended to both reports.
|
||||
|
||||
**Capabilities M2 delivered, which later milestones inherit rather than build:**
|
||||
|
||||
- a single-user product with no accounts, sessions or auth — all four
|
||||
`/api/auth/*` routes are gone, not gated; 52 API routes down to 36,
|
||||
- Ollama as the only inference backend, with no cloud provider code and no API
|
||||
key anywhere in the product,
|
||||
- an **address-based inference endpoint policy**, enforced on save and again
|
||||
before every outbound request, that refuses public addresses even when the
|
||||
database is edited behind the settings API (ADR 011),
|
||||
- M1's trusted-LAN and TLS behaviour carried through the removal intact —
|
||||
verified HTTPS against a real second machine, no bypass option,
|
||||
- SQLite as the only store; Postgres, Neon and the Render deployment path removed,
|
||||
- no campaign scripting: the QuickJS engine and `/api/scripts` are gone,
|
||||
- a materially simpler runtime — 10 environment variables down to two, 6 Python
|
||||
and 21 npm packages removed, and a 933 kB bundle down to 395 kB.
|
||||
|
||||
**Debt carried forward, none of it blocking M3:** `Settings.model` still
|
||||
defaults to `""` with nothing prompting for it (M8); inert legacy tables and
|
||||
columns await a cleanup migration once the schema settles, after M3/M5; there
|
||||
are still no frontend tests (M8); `docs/*.html`, upstream's project site and not
|
||||
served by the app, still links Google Fonts. Full table in the implementation
|
||||
report §P.
|
||||
|
||||
---
|
||||
|
||||
# M3 — Production Non-Destructive History, Redo, and Active-Head Export
|
||||
@@ -310,6 +341,34 @@ Remove or demote:
|
||||
|
||||
Accepted story state is genre-neutral, auditable, reconstructable, and no longer depends on ambiguous relative deltas.
|
||||
|
||||
## Note from M2 — eight rollback tests use the world-state engine as instrumentation
|
||||
|
||||
M2 removed QuickJS. Eight existing rollback/history tests had used a JavaScript
|
||||
counter as deterministic instrumentation — a value they could change on a turn
|
||||
and then assert had been rolled back — and they now use the inherited
|
||||
RPG/world-state delta machinery for the same purpose.
|
||||
|
||||
Read those tests correctly before touching them:
|
||||
|
||||
- **what they exercise** is rollback and state reconstruction across undo,
|
||||
retry, takes and divergence;
|
||||
- **their use of the world-state engine is test instrumentation, not an
|
||||
endorsement** of RPG-shaped state as the target architecture. Nothing about
|
||||
them argues against the typed-event model this milestone installs;
|
||||
- **when M5 replaces or generalizes the world-state protocol, the
|
||||
instrumentation must move with it** to the new narrative-state mechanism. In
|
||||
practice that is one schema entry and one helper in `backend/tests/fakes.py`;
|
||||
- **preserve or rework these tests; do not delete them** merely because their
|
||||
current instrumentation is RPG-shaped. The behaviour they pin is exactly the
|
||||
behaviour M5 is most likely to break.
|
||||
|
||||
`test_state_revert` and `test_delete_state` additionally assert *destructive*
|
||||
undo semantics and are expected to be rewritten by M3; that is separate from
|
||||
this instrumentation point, and they should likewise be rewritten rather than
|
||||
dropped.
|
||||
|
||||
Evidence: `planning/reports/M2-IMPLEMENTATION-REPORT.md` §K.2, §Q.
|
||||
|
||||
---
|
||||
|
||||
# M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory
|
||||
@@ -348,6 +407,28 @@ Align inherited AI-DnD memory/context behavior with the final authority and hist
|
||||
|
||||
Long-running story context is lineage-safe, authority-aware, local, inspectable, and bounded.
|
||||
|
||||
## Note from M2 — background memory failure must be observable
|
||||
|
||||
M2 shipped with the memory bank entirely dead, and the full suite stayed green.
|
||||
Summaries and embeddings raised `AttributeError` inside a fire-and-forget task:
|
||||
no user-visible error, no log a player would read, and no failing test, because
|
||||
every memory test stubs the provider factories out.
|
||||
|
||||
M6 therefore additionally requires:
|
||||
|
||||
- **memory/summarization background failures must be observable** — a
|
||||
fire-and-forget task that dies must leave a record a user or maintainer can
|
||||
actually find, rather than being swallowed;
|
||||
- **tests must exercise at least one real provider-construction/wiring path**,
|
||||
not only mocked factories, so that a moved or removed setting surfaces as a
|
||||
test failure (`TECHNICAL-DESIGN.md` §18.1);
|
||||
- **derived-memory failure must not corrupt accepted story state.** Already in
|
||||
the scope list above; M2's evidence is why it stays there. A dead memory bank
|
||||
degraded the storyteller quietly and left the transcript correct, which is the
|
||||
right failure direction — but it must also be a *visible* one.
|
||||
|
||||
Evidence: `planning/reports/M2-IMPLEMENTATION-REPORT.md` §A.1, §9.1.
|
||||
|
||||
---
|
||||
|
||||
# M7 — First-Class Imported Knowledge Library
|
||||
|
||||
Reference in New Issue
Block a user