m2-local-only-surface
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2fdd2547f0 |
Planning: record M2 closeout decisions
M2's review reported six planning recommendations rather than applying them, three marked before M3. All six are applied here, plus three additions drawn from the same evidence. No implementation file is touched. The endpoint policy was the gap that mattered. It is the most consequential setting in the application — the storyteller sends the player's prose, the context, the memories and the embedding inputs to whatever address it names — and it existed only as a module docstring. It is now ADR 011 and a new §10A in the threat model, which also retires the assumption in §71A that the inherited guard was a starting point. It was not: AI-DnD's SSRF guard blocked private addresses to stop a hosted server reaching its own internal network, which is the exact opposite of what a local storyteller needs. It was removed, not adapted. Both documents state the rule as implemented — an allowlist of explicit local-network CIDRs, every resolved address checked, enforced on save and again before every outbound request, TLS never traded against it — and both state the two residual limits plainly rather than implying they are covered: a hostile host already on the trusted LAN is inside the permitted boundary, and a rebinding interval exists between the policy's resolution and the client's connection. Accepted risks, not M3 work. The CIDRs are spelled out rather than derived from is_private/is_reserved, and the ADR records why: is_private is true of the documentation ranges and 0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it refuses an ordinary same-host Ollama on [::1]. TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening items. A new §5.2 records the M1/M2 architecture as fact rather than intention, so later milestones inherit what the code does. A new §18.1 carries the lesson of M2's two regressions: when removing a setting, test a real consumer construction path; when adding one, prove it reaches the component that uses it. Both defects hid behind a green suite because the tests at that boundary were mocks. BUILD-MILESTONES records M2 complete, with the capabilities later milestones inherit and the debt carried forward. Two notes go to milestones that would otherwise misread what M2 left them. M5 is told that eight rollback tests now use the world-state engine as instrumentation and not as endorsement — the instrumentation moves when the protocol does, and those tests are reworked rather than deleted. M6 is told that the memory bank died silently under a green suite, so background failure must be observable and at least one real provider-construction path must be tested. The security contract gains what M2 demonstrated. H10 now names the two conditions that were defects during M2: a wildcard origin must be refused at startup, and an unknown /api path must 404 rather than returning the SPA with 200. New H12 covers endpoint enforcement, and its fourth pass condition is the one that matters — a public endpoint written into the database behind the settings API must still be refused at the wire. A build passing the first three and failing that one has configuration validation only. SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement; it removed capability the specification never asked for. The two M2 reports gain appended closeout notes rather than edits. Their original wording about an uncommitted working tree was true when written, and the note records what happened afterwards: the six-file correction is |
||
|
|
8652fe7cd8 |
M2 review: two regressions the green suite hid, and the reports
The post-implementation review of M2, plus the three corrections it took
to make the evidence true. Reports:
planning/reports/M2-BASELINE-REPORT.md 868 lines, the measurements
planning/reports/M2-IMPLEMENTATION-REPORT.md 758 lines, the reading of them
Verdict is PASS, accept with non-blocking debt, proceed to M3. Every M2
requirement is met and the ones that matter were tested by running the
build rather than reading it: a cloud endpoint written straight into
SQLite with sqlite3, behind the API's back, still refused at the wire;
trusted-LAN HTTPS against the real second machine with verification on;
captures showing zero packets outside loopback and the approved host.
Three defects, all found by running the shipped image.
The memory bank was dead. M2 removed Settings.api_key_plain with the API
key, and memorybank's two provider factories still read it. It failed
inside a fire-and-forget task, so no user error, no log anyone would
read, and no test — every memory test stubs those factories. All 604
tests passed with summaries and embeddings silently not happening.
The configurable model timeout never reached the turn engine. Stored,
validated, exposed in the API, rendered in the UI, and not passed to the
provider. M2's own exit criterion was half met: the constant had moved
but the setting did nothing.
And requirements.lock still pinned quickjs, psycopg and cryptography, so
the setup path DEVELOPMENT.md gives a new developer would have
reinstalled all three.
Both code defects now have the test that would have caught them: one
constructs every provider factory from a real Settings row, one drives
the turn endpoint, the chat endpoint and the summariser and asserts the
configured timeout arrives at each. That is the lesson worth keeping from
this milestone — after removing an attribute, build each consumer from a
real object; after adding a setting, prove it lands. Both failures were
in background or plumbing paths, which is exactly where a subtractive
change cannot see itself.
606 tests pass, up from 604. Lint, build and image are clean. Every
runtime result in the baseline report came from an image built after
these fixes; the reports say plainly that commit
|