M2's review reported six planning recommendations rather than applying them, three marked before M3. All six are applied here, plus three additions drawn from the same evidence. No implementation file is touched. The endpoint policy was the gap that mattered. It is the most consequential setting in the application — the storyteller sends the player's prose, the context, the memories and the embedding inputs to whatever address it names — and it existed only as a module docstring. It is now ADR 011 and a new §10A in the threat model, which also retires the assumption in §71A that the inherited guard was a starting point. It was not: AI-DnD's SSRF guard blocked private addresses to stop a hosted server reaching its own internal network, which is the exact opposite of what a local storyteller needs. It was removed, not adapted. Both documents state the rule as implemented — an allowlist of explicit local-network CIDRs, every resolved address checked, enforced on save and again before every outbound request, TLS never traded against it — and both state the two residual limits plainly rather than implying they are covered: a hostile host already on the trusted LAN is inside the permitted boundary, and a rebinding interval exists between the policy's resolution and the client's connection. Accepted risks, not M3 work. The CIDRs are spelled out rather than derived from is_private/is_reserved, and the ADR records why: is_private is true of the documentation ranges and 0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it refuses an ordinary same-host Ollama on [::1]. TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening items. A new §5.2 records the M1/M2 architecture as fact rather than intention, so later milestones inherit what the code does. A new §18.1 carries the lesson of M2's two regressions: when removing a setting, test a real consumer construction path; when adding one, prove it reaches the component that uses it. Both defects hid behind a green suite because the tests at that boundary were mocks. BUILD-MILESTONES records M2 complete, with the capabilities later milestones inherit and the debt carried forward. Two notes go to milestones that would otherwise misread what M2 left them. M5 is told that eight rollback tests now use the world-state engine as instrumentation and not as endorsement — the instrumentation moves when the protocol does, and those tests are reworked rather than deleted. M6 is told that the memory bank died silently under a green suite, so background failure must be observable and at least one real provider-construction path must be tested. The security contract gains what M2 demonstrated. H10 now names the two conditions that were defects during M2: a wildcard origin must be refused at startup, and an unknown /api path must 404 rather than returning the SPA with 200. New H12 covers endpoint enforcement, and its fourth pass condition is the one that matters — a public endpoint written into the database behind the settings API must still be refused at the wire. A build passing the first three and failing that one has configuration validation only. SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement; it removed capability the specification never asked for. The two M2 reports gain appended closeout notes rather than edits. Their original wording about an uncommitted working tree was true when written, and the note records what happened afterwards: the six-file correction is8652fe7,8c65ae9remains the implementation commit, and the two were never squashed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HsZBU8sWRuYTyLgWsu2oQ6
851 lines
42 KiB
Markdown
851 lines
42 KiB
Markdown
# M2 — Implementation Review Report
|
||
|
||
**Date:** 2026-09-02
|
||
**Milestone:** M2, *Remove Hosted, Cloud, Scripting, and Unneeded Deployment Surface*
|
||
**Audience:** the architecture/design reviewer deciding whether to accept M2 and start M3
|
||
**Evidence:** `planning/reports/M2-BASELINE-REPORT.md`. Section references below
|
||
(§) point into it, and where this report and that one differ on a runtime or
|
||
test fact, **the baseline report is the record.**
|
||
|
||
Hostnames and LAN addresses are placeholders (`inference.lan`, `192.168.0.50`).
|
||
|
||
---
|
||
|
||
# A. Executive result
|
||
|
||
```text
|
||
Overall M2 result: PASS
|
||
```
|
||
|
||
**Recommendation: ACCEPT M2 WITH NON-BLOCKING DEBT AND PROCEED TO M3.**
|
||
|
||
Directly, in the order asked:
|
||
|
||
| Question | Answer |
|
||
| --- | --- |
|
||
| Is it genuinely a single-user local storyteller? | **Yes.** No login, no accounts, no sessions; all four `/api/auth/*` routes 404 (§12). |
|
||
| Is Ollama the only production inference backend? | **Yes.** No cloud provider code, no key, and public addresses are refused at the wire (§7). |
|
||
| Are hosted/cloud/account/scripting surfaces actually removed? | **Yes** — removed, not hidden. 52 API routes → 36; `/api/auth`, `/api/analytics`, `/api/scripts` gone entirely (§3, §12). |
|
||
| Does same-host Ollama still work? | **Yes** (§5.3–5.4). |
|
||
| Does trusted-LAN Ollama still work? | **Yes**, against the real second machine (§6.2). |
|
||
| Does HTTPS/private-CA still work with verification on? | **Yes** — `TLS verified, peer CN = inference.lan`, all four clients on the shared trust context, no bypass exists (§6.1, §8). |
|
||
| Does it still run with Internet blocked? | **Yes** (§5.1, §6.1). |
|
||
| Did any M1 behaviour regress? | **Yes — two, both found by this review and both fixed** (§A.1). |
|
||
| Should the project proceed to M3? | **Yes.** |
|
||
| Blockers before M3? | **None.** |
|
||
|
||
## A.1 Two M1 regressions that M2 shipped, and this review caught
|
||
|
||
These are the most important findings and they are not buried.
|
||
|
||
**1. The memory bank was silently broken.** M2 removed
|
||
`Settings.api_key_plain`, but `memorybank`'s two provider factories still read
|
||
it. Summaries and embeddings raised `AttributeError` inside a fire-and-forget
|
||
background task — no user-visible error, no log a player would see, no failing
|
||
test. **All 604 tests passed with the memory bank dead** (§9.1).
|
||
|
||
**2. The configurable model timeout never reached the turn engine.** The
|
||
setting was stored, validated, exposed in the API and rendered in the UI, and
|
||
then not passed to the provider. M2's own exit criterion — "the model timeout is
|
||
no longer an undocumented hardcoded 120-second limitation" — was half met: the
|
||
constant had moved but the setting was inert (§9.2).
|
||
|
||
A third, smaller: `requirements.lock` still pinned `quickjs`, `psycopg` and
|
||
`cryptography`, so the documented setup path would have reinstalled all three
|
||
(§9.3).
|
||
|
||
All three are corrected in the working tree, with tests that would have caught
|
||
the first two. **The M2 commit `8c65ae9` does not contain these fixes**; the tree
|
||
is six files ahead of it and needs a follow-up commit. Every runtime result in
|
||
the baseline report was produced by an image built from the fixed tree.
|
||
|
||
**What this says about the milestone is more useful than the defects
|
||
themselves:** a subtractive milestone's risk is not what it deletes, it is what
|
||
still reaches for the deleted thing from a code path no test exercises. Both
|
||
defects were in *background* or *plumbing* paths. §K returns to this.
|
||
|
||
---
|
||
|
||
# B. Repository and provenance
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Branch | `m2-local-only-surface` |
|
||
| HEAD | `8c65ae99deda49b22f415d5868987e00bbcb173c` (signed, `%G? = G`) |
|
||
| M2 commit | `8c65ae9`, one commit, parent `1a28a9a` |
|
||
| M1 starting point | `1a28a9a` (post-M1 planning corrections) |
|
||
| Upstream ancestry | `d72f7c1…` is an ancestor of HEAD; 172 upstream commits reachable |
|
||
| `LICENSE` | byte-identical to upstream |
|
||
| Private data committed | none (§1) |
|
||
| Working tree | **six files modified** by this review (§A.1) |
|
||
|
||
No unrelated upstream re-import occurred. `PROVENANCE.md` records what M2
|
||
removed and what it retained; `DEVELOPMENT.md` documents the new surface.
|
||
|
||
---
|
||
|
||
# C. Implementation inventory
|
||
|
||
**94 files changed, +1 395 −6 578.** Three added, twenty-four deleted, sixty-seven
|
||
modified. Grouped by purpose:
|
||
|
||
### 1. Single-user / account removal
|
||
Deleted `routers/auth.py`, `cleanup.py` (guest-retention sweeper),
|
||
`accesslog.py`, `security.py` (session signing + key encryption). `auth.py` cut
|
||
from 259 to 74 lines: it now resolves one implicit local user and nothing else.
|
||
Frontend: the `/auth/me` bootstrap, `AuthModal`, guest nudge, log-in/sign-up/
|
||
log-out controls, and the 401-and-retry dance in `api.js`.
|
||
|
||
### 2. Analytics removal
|
||
Deleted `analytics.py`, `routers/analytics.py`, `pages/Analytics.jsx`, the
|
||
`trackPageview` beacon, the `ApiErrorMiddleware` that fed the error tally, the
|
||
lifespan flusher, and every `record_event` call site. Three model classes
|
||
unmapped.
|
||
|
||
### 3. Hosted database / deployment removal
|
||
`render.yaml` deleted. `database.py` reduced to SQLite only —
|
||
`AIDND_DATABASE_URL`, `DATABASE_URL`, the psycopg URL normaliser and the
|
||
serverless pre-ping are gone. `psycopg[binary]` removed.
|
||
|
||
### 4. Cloud-provider removal
|
||
`_OPENROUTER_HOST`, `_PREFERRED_UPSTREAM`, `_apply_provider_routing`,
|
||
`_apply_reasoning_budget`, the `Authorization` header, the `api_key` field and
|
||
its Fernet encryption. `cryptography` removed.
|
||
|
||
### 5. Ollama-only configuration
|
||
`OpenAICompatibleProvider(endpoint_url, model, api_mode, read_timeout)` — no
|
||
key, no reasoning budget. `ProviderConfig`/`resolve_provider_config` deleted
|
||
outright; callers read `Settings` directly.
|
||
|
||
### 6. Endpoint policy — **the only addition**
|
||
`backend/app/endpoints.py` (179 lines) replaces `netguard.py`, inverting its
|
||
rule (§F).
|
||
|
||
### 7. Trusted-LAN / TLS preservation
|
||
M1's `tlstrust.py` untouched. All four HTTP clients still pass
|
||
`verify=tlstrust.ssl_context()`; `test_tls_trust.py` enforces it by AST walk.
|
||
|
||
### 8. QuickJS / scripting removal
|
||
`app/scripting/` (3 files), both script routers, `Script`/`AdventureScript`
|
||
models, the `scenario_scripts` table, script schemas, script export/import,
|
||
`Scripts.jsx`, `ScriptEditor.jsx`, `ScriptsPanel.jsx`, `StatusDrawer.jsx`,
|
||
`ScriptReport`. `quickjs` removed; CodeMirror removed from npm.
|
||
|
||
### 9. Hosted-policy removal
|
||
`limits.py` cut from 345 to ~150 lines: rate limiting, `X-Forwarded-For` client
|
||
IP, login throttling and per-user quotas gone. **Kept**: request body ceiling,
|
||
per-adventure row caps, import list caps.
|
||
|
||
### 10. Settings simplification
|
||
Removed: API key field, "Remove key" button, demo banner, reasoning budget.
|
||
Added: Ollama endpoint help text stating the policy, and a model-timeout field.
|
||
|
||
### 11. Model timeout
|
||
`CONNECT_TIMEOUT = 10`, `DEFAULT_READ_TIMEOUT = 300`, `EMBED_READ_TIMEOUT = 60`,
|
||
plus `Settings.model_timeout_seconds` (migration 77, default 300, bounded
|
||
30–3600).
|
||
|
||
### 12. Loopback protections
|
||
`docker-compose.yml` publishes `127.0.0.1:8000:8000`; `--proxy-headers` dropped;
|
||
`AIDND_CORS_ORIGINS="*"` now refuses to start; an unknown `/api/...` path 404s
|
||
instead of returning the SPA with status 200.
|
||
|
||
### 13. Tests
|
||
Two new files (63 collected). Eight files converted from JS instrumentation to
|
||
the world-state engine. Six files retired with their subsystems.
|
||
|
||
### 14. Documentation
|
||
`DEVELOPMENT.md` gained the endpoint policy and the diagnostics table;
|
||
`.env.example` cut from 110 lines to 24; `PROVENANCE.md` records M2.
|
||
|
||
## Deliberately retained
|
||
|
||
| Retained | Why |
|
||
| --- | --- |
|
||
| `users` table and `user_id` foreign keys | The M2 brief permits it. Removing them means a migration across most of the schema to delete a column that costs nothing. One row; nothing creates a second; no request carries an identity. |
|
||
| Five inert tables, six inert columns | A destructive migration would risk an existing campaign database for tidiness. Unmapped or written-empty; nothing reads them (§13). |
|
||
| `providers/openai_compatible.py` name | It speaks OpenAI's *protocol* to Ollama. Renaming would churn a file M3 does not touch, for no behaviour change. |
|
||
| AI Chat scratchpad | Not hosted-only. It is a local tool for checking a model or prompt, and the endpoint policy constrains where it can talk. |
|
||
| RPG world-state engine | M5's scope. M2 additionally now *depends* on it as test instrumentation (§K.2). |
|
||
| Destructive Undo, no Redo | M3's scope, explicitly out of M2. |
|
||
|
||
---
|
||
|
||
# D. Single-user / hosted-account removal
|
||
|
||
| Classification | Contents |
|
||
| --- | --- |
|
||
| **Code removed** | `routers/auth.py`, `cleanup.py`, `accesslog.py`, `security.py`, 185 of 259 lines of `auth.py`, `AuthModal`, the account nav block, the session-retry logic |
|
||
| **Code retained but inert** | none in this area |
|
||
| **Schema retained for compatibility** | `users` + `user_id` FKs (active but single-row); `users.demo_turns_used`/`_date`; `settings.api_key` |
|
||
| **Production functionality still active** | **none** |
|
||
|
||
* **Can a local user meet a login requirement?** No. There is no login UI, no
|
||
session cookie, and `get_current_user` always succeeds.
|
||
* **Are account APIs reachable?** No — 404 on all four (§12).
|
||
* **Are hosted account modules imported at runtime?** No; the files do not exist.
|
||
* **Does the database retain user identifiers?** Yes, one row.
|
||
* **Merely internal now?** Yes, and documented as such in `auth.py`'s docstring
|
||
and `PROVENANCE.md`.
|
||
* **Did avoiding the schema rewrite reduce risk?** **Materially.** `user_id`
|
||
appears on scenarios, adventures, settings and memories, and the ownership
|
||
filters run through the story-tree and memory queries M3 will modify. A
|
||
schema rewrite would have put a migration under those queries in the same
|
||
milestone that removed accounts — two risky changes entangled. Keeping the
|
||
column made M2 a deletion rather than a redesign.
|
||
|
||
No active hosted-account functionality remains.
|
||
|
||
---
|
||
|
||
# E. Analytics / telemetry
|
||
|
||
**Removed:** the collection module, the routes, the beacon, the Visitors page and
|
||
nav link, the error-tally middleware, the batch flusher, and every call site.
|
||
**Retained inert:** three tables, unmapped, never opened.
|
||
|
||
No destructive migration was performed because dropping tables from a live
|
||
campaign database buys nothing and can fail.
|
||
|
||
**Verified at runtime, not asserted:** across a full offline session of 5 131
|
||
packets — settings, campaign creation, four turns, embeddings, a fork, two
|
||
induced failures, a restart — there were **zero non-loopback unicast packets**
|
||
and no TCP connection opened outside loopback (§5.9). The only DNS names that
|
||
appear are the ones my own isolation probe deliberately looked up, all within
|
||
5 ms of the capture starting and 45 seconds before the storyteller received its
|
||
first request, plus `ollama.com` from the Ollama service itself.
|
||
|
||
Analytics is removed, not merely hidden: collection does not occur.
|
||
|
||
---
|
||
|
||
# F. Hosted database / deployment, and the endpoint policy
|
||
|
||
## Database
|
||
|
||
Postgres, Neon, `psycopg`, the URL normaliser and `render.yaml` are gone;
|
||
production is SQLite at `AIDND_DB_PATH`. An existing M1 database opens
|
||
unchanged — the offline run used a volume carrying campaigns created before the
|
||
fixes, survived an image replacement and a restart with digest
|
||
`157c882c60579617` unchanged (§5.8). Story-tree persistence is unaffected;
|
||
branch count, memory count and context inspection all survive restart.
|
||
|
||
One Postgres-shaped abstraction is retained: `migrations.py` and `tools/dbmeter.py`
|
||
carry comments and a `_for_dialect` helper for SQL that differed between
|
||
SQLite and psycopg. Removing it would churn the migration history for no gain.
|
||
|
||
## Endpoint policy — the design centre of M2
|
||
|
||
`endpoints.py` **inverts** the rule it replaces. `netguard.py` refused *private*
|
||
addresses, to stop a hosted visitor making the server fetch an internal service.
|
||
The threat here is the opposite: the user is trusted, and what must not happen is
|
||
the story reaching the public Internet. So the new rule refuses *public*
|
||
addresses.
|
||
|
||
The line is drawn by **address against an explicit allowlist of networks**, not
|
||
by hostname and not by asking `ipaddress` what it thinks is private:
|
||
|
||
```text
|
||
127.0.0.0/8 10.0.0.0/8 172.16.0.0/12 192.168.0.0/16
|
||
169.254.0.0/16 100.64.0.0/10 ::1/128 fc00::/7 fe80::/10
|
||
```
|
||
|
||
Every address a name resolves to must be in one. Applied twice: on save, for a
|
||
good error; and before every outbound request, because DNS moves.
|
||
|
||
Measured behaviour (§7): loopback v4 and v6, `localhost`, all three RFC1918
|
||
ranges, link-local, CGNAT, IPv6 ULA and a Docker-internal name are **accepted**;
|
||
four cloud providers, public IPv4 and IPv6, `0.0.0.0` and a wrong scheme are
|
||
**refused**, each with a reason naming what to do instead.
|
||
|
||
### What it guarantees, and what it does not
|
||
|
||
**Guarantees.** No request leaves for an address outside those networks, from any
|
||
of the four clients, whatever is stored. Demonstrated against a **hand-edited
|
||
SQLite row** — the settings row was rewritten to `openrouter.ai` with `sqlite3`,
|
||
bypassing the API entirely, and the next turn was refused at the wire (§5.7).
|
||
That is the strongest available form of this evidence.
|
||
|
||
**Does not guarantee.** The allowlist is *address* scope, not *ownership* scope:
|
||
if a hostile host sits on the user's own LAN, the policy permits it — as it must,
|
||
since that is what "trusted LAN" means. A rebinding window exists in principle
|
||
between the policy's `getaddrinfo` and httpx's own connect; closing it would mean
|
||
pinning the resolved address into the connection, which is a larger change than
|
||
M2 warranted. Neither is a v1 concern: both require an attacker already inside
|
||
the trusted network.
|
||
|
||
**One design finding worth flagging** (§K.1): writing this rule around
|
||
`ipaddress.is_private` / `is_reserved` — the obvious approach — is wrong in a way
|
||
that is easy to ship. Python classifies IPv6 loopback `::1` as *reserved*, so
|
||
that version refused `http://[::1]:11434/v1`, an ordinary same-host endpoint. It
|
||
also calls the documentation ranges *private*, so they would have been allowed.
|
||
The explicit CIDR list exists because of that, and the reasoning is recorded in
|
||
the module docstring.
|
||
|
||
---
|
||
|
||
# G. Ollama-only provider
|
||
|
||
**Providers in production code: one.** `OpenAICompatibleProvider`, speaking
|
||
OpenAI's `/v1` protocol to Ollama.
|
||
|
||
* Provider options visible to the user: **none**. There is no selector, because
|
||
there is nothing to select between.
|
||
* OpenAI / OpenRouter / Groq / cloud APIs: **gone** — constants, routing,
|
||
attribution and the reasoning-budget parameter.
|
||
* Cloud API-key fields: **gone** from the model API, the UI and the request
|
||
headers. The column is inert.
|
||
|
||
**The module name is not the question.** The question the brief poses —
|
||
*can production use still send story content to a public inference service
|
||
through normal configuration?* — was tested directly rather than reasoned about:
|
||
saving a cloud URL is refused with HTTP 400 (§7), and a cloud URL written
|
||
straight into the database is refused at request time (§5.7). No.
|
||
|
||
---
|
||
|
||
# H. Trusted-LAN and TLS regression
|
||
|
||
Re-verified after M2 on the **real second physical machine** used in M1, over
|
||
HTTPS with a private CA (§6).
|
||
|
||
| Check | Result |
|
||
| --- | --- |
|
||
| Endpoint class | `https://inference.lan:8443/v1`, non-loopback, private CA |
|
||
| Connection | succeeded; diagnostics listed both installed models |
|
||
| Certificate validation | performed and passed |
|
||
| Hostname verification | performed — `peer CN = inference.lan` |
|
||
| CA trust | the machine's own CA store, unioned with certifi |
|
||
| Narration | 3 turns, 4.0–10.7 s |
|
||
| Model listing/discovery | via the same trust path |
|
||
| Embeddings | 1 memory embedded through the LAN host |
|
||
| Settings / Test Connection | same trust path |
|
||
| Retry | new take filed, both attempts retained (§6.4) |
|
||
| Restart and resume | digest `66cc82944f8b6eef` unchanged; endpoint and timeout preserved |
|
||
| Capture | 308 packets to the approved host, 364 loopback, **0 elsewhere, 0 DNS** |
|
||
|
||
Every model request went to the configured host: 4 `chat/completions`, 1
|
||
`embeddings`, nothing else.
|
||
|
||
**No `verify=False`, insecure fallback or "ignore certificate errors" path
|
||
exists.** The only textual match for such a thing in the tree is prose in
|
||
`tlstrust.py` saying it deliberately does not exist. All four clients use the
|
||
same context, enforced by an AST-walking test that also fails if a *fifth*
|
||
client appears (§8). No client uses different trust behaviour.
|
||
|
||
---
|
||
|
||
# I. QuickJS and scripting removal
|
||
|
||
Removed: the QuickJS dependency, `app/scripting/` entire, both routers, the
|
||
`onInput`/`onModelContext`/`onOutput` hooks in the turn engine, the `Script` and
|
||
`AdventureScript` models, script schemas, script export/import (bundles that
|
||
carry `scripts` now import with the story intact and the scripts ignored), the
|
||
Scripts page, the script editor, the in-play Scripts panel, the script-state
|
||
drawer and the Insights script report.
|
||
|
||
**Imported content cannot execute JavaScript.** There is no engine to execute it
|
||
in — `import app.scripting` raises `ImportError` and no module imports `quickjs`
|
||
(§12). A scenario or bundle carrying script source is data that nothing reads.
|
||
|
||
## Test-coverage accounting
|
||
|
||
This is where a subtractive milestone can hide damage behind a green number, so
|
||
it is spelled out. Of 88 node IDs that disappeared, **4 are renames** and 84 were
|
||
genuinely retired (§4).
|
||
|
||
| Category | Count | Verdict |
|
||
| --- | ---: | --- |
|
||
| Whole files whose subject was removed (`analytics`, `guest_cleanup`, `accesslog`, `ratelimit_hardening`, `reasoning_param`) | 64 | Obsolete product behaviour intentionally deleted. No coverage lost — the behaviour is gone. |
|
||
| `test_netguard.py` | 5 | **Replaced**, not lost: `test_endpoint_policy.py` (31 collected) covers the inverted rule far more thoroughly. |
|
||
| `test_chat.py` demo-key pinning and power-user gate | 8 | Subject removed. The file keeps and extends what the page still does. |
|
||
| `test_prompt_caching.py` OpenRouter routing | 4 | Subject removed. The prompt-layout and usage tests, which are the file's point, are untouched. |
|
||
| `test_memory_rewrite.py` Postgres DSN masking | 2 | Subject removed with Postgres. |
|
||
| `test_branch_clause.py::test_user_scripts_are_handed_the_path` | 1 | **Equivalent coverage retained.** It asserted the scripting history API saw the branch path; the test directly above it asserts the same invariant one layer down on `history.story_actions`/`count`/`tail`, which is where the branch clause lives. |
|
||
| Instrumentation conversion, 8 files | 0 retired | **Coverage retained in full.** A JS counter was measuring the *state snapshot and rollback machinery*, which M2 does not touch. The counter moved to the world-state engine; the assertions are unchanged in substance. |
|
||
|
||
**No meaningful coverage was lost.** One area gained a great deal: endpoint
|
||
policy went from 5 tests of the opposite rule to 31 of the current one.
|
||
|
||
---
|
||
|
||
# J. Settings surface, and the model timeout
|
||
|
||
## Settings after M2
|
||
|
||
**Remains:** Ollama endpoint (with help text stating the policy), narrator model,
|
||
embedding model, temperature, max output tokens, context budget, API mode,
|
||
narrator prompt, memory-bank capacity and top-k, **model timeout**, Test
|
||
connection, and the provider debug log.
|
||
|
||
**Removed:** the API key field and its "Remove key" button, the demo-key banner,
|
||
the reasoning budget, and the "OpenAI-compatible" framing on the endpoint field.
|
||
|
||
**Is it correct for v1?** Yes. Every field maps to something the product does,
|
||
and nothing offers a capability the product refuses. It is not *polished* — that
|
||
is M8 — but it is not misleading, which is what M2 asked for.
|
||
|
||
One onboarding gap survives from M1 and is not new: `Settings.model` defaults to
|
||
`""`, so a fresh install cannot generate a turn until a model is named, and
|
||
nothing prompts. M2 improves the *diagnosis* — the connection test now warns when
|
||
the endpoint is reachable but has no model by the configured name — without
|
||
fixing the onboarding, which is M8's.
|
||
|
||
## Model timeout
|
||
|
||
| | Value |
|
||
| --- | --- |
|
||
| Connect timeout | 10 s, constant — a wrong address should fail fast |
|
||
| Generation/read timeout | `Settings.model_timeout_seconds`, default **300 s** |
|
||
| Allowed range | 30–3600, enforced by the schema (422 outside) |
|
||
| UI | a numeric field with help text about cold loads |
|
||
| Config source | database row, not an environment variable |
|
||
| Embeddings | separate 60 s constant — short calls that never cold-load |
|
||
| Connection test | separate 15 s constant — a listing, not a generation |
|
||
|
||
**Demonstrated no longer a fixed 120 s constant:** `test_the_timeout_is_configurable`
|
||
sets 900 through the API and reads it back;
|
||
`test_the_configured_timeout_reaches_every_generating_client` drives the turn
|
||
endpoint, the chat endpoint and the summariser factory with a recording provider
|
||
and asserts the configured value reaches each; `test_an_unusable_timeout_is_refused`
|
||
rejects 0, −1, 29 and 3601. The offline run used 600 s and the LAN run used 600 s,
|
||
both read back after restart (§5.8, §6.5).
|
||
|
||
The retained fixed timeouts are the three above, each named and each justified by
|
||
the shape of the call.
|
||
|
||
---
|
||
|
||
# K. Important surprises
|
||
|
||
Four. None invented for completeness.
|
||
|
||
## K.1 An address policy written the obvious way is wrong
|
||
|
||
Writing the endpoint rule around `ipaddress.is_private` / `is_global` /
|
||
`is_reserved` — which is what the inherited `netguard.py` did and what any
|
||
reviewer would expect — produces a rule that **refuses IPv6 loopback**
|
||
(`::1` is classified reserved) and **accepts the documentation ranges** (they
|
||
are classified private). The first was caught by a test asserting
|
||
`http://[::1]:11434/v1` is allowed. This is worth knowing beyond this project:
|
||
the standard library's classifications answer a different question than "is this
|
||
on my own network".
|
||
|
||
## K.2 Removing scripting made the product depend on the world-state engine — in the tests
|
||
|
||
Eight test files used a QuickJS counter as instrumentation for the state
|
||
rollback machinery. The natural replacement was the world-state engine, so those
|
||
tests now assert rollback *through* the RPG delta protocol. **M5 plans to replace
|
||
that protocol** with typed narrative-state events. M5 will therefore have to move
|
||
this instrumentation a second time. It is cheap — the counter is one schema entry
|
||
and one reply template in `tests/fakes.py` — but M5's brief should say so rather
|
||
than discover it.
|
||
|
||
## K.3 The largest surviving subsystem is the one M2 could not touch
|
||
|
||
`migrations.py` is **1 128 lines**, the biggest file in the backend, and M2
|
||
reduced it by one line while adding one. It carries the full history of a schema
|
||
that M2 has now partly orphaned: migrations that create `scripts`,
|
||
`analytics_daily` and `access_log`, and backfills for `state_after`. This is
|
||
correct — migration history must not be rewritten — but a reviewer expecting the
|
||
"trust and maintenance surface" to have shrunk proportionally should know that a
|
||
tenth of the backend is history that only grows.
|
||
|
||
## K.4 Two defects hid in exactly the places a subtractive milestone cannot see
|
||
|
||
Both regressions (§A.1) were invisible to a 604-test green suite, for the same
|
||
reason: the code that still reached for a removed thing sat in a path no test
|
||
executed with real objects. The memory-bank factories are stubbed in every
|
||
memory test; the timeout argument was never asserted. **The lesson is a testing
|
||
one, not a coding one:** after removing an attribute, the cheapest useful test is
|
||
one that constructs each consumer from a real object, and after adding a setting,
|
||
one that proves it arrives. Both now exist.
|
||
|
||
---
|
||
|
||
# L. Acceptance matrix
|
||
|
||
| Requirement | Result | Evidence / notes |
|
||
| --- | --- | --- |
|
||
| Single-user, no login | PASS | §12 — four auth routes 404 |
|
||
| Hosted account/guest/demo removed | PASS | §C.1, §D |
|
||
| Hosted rate-limit policy removed | PASS | `limits.py` 345→~150 lines; resource bounds kept |
|
||
| Analytics/Visitors removed | PASS | §E |
|
||
| No telemetry generated | PASS | §5.9 — 0 non-loopback unicast in 5 131 packets |
|
||
| Render path removed | PASS | `render.yaml` deleted |
|
||
| Neon path removed | PASS | §F |
|
||
| Postgres runtime removed | PASS | `psycopg` gone from requirements and closure |
|
||
| SQLite M1 data still works | PASS | §5.8 — digest survives image replacement and restart |
|
||
| Ollama-only production provider | PASS | §G |
|
||
| Cloud API keys removed | PASS | §12 — absent from the API, unsettable |
|
||
| Public inference endpoints blocked | PASS | §7, §5.7 — including against a hand-edited DB |
|
||
| Loopback Ollama accepted | PASS | §7 — v4, v6 and `localhost` |
|
||
| Trusted-LAN Ollama accepted | PASS | §6.2 |
|
||
| TLS/private-CA support preserved | PASS | §6.1 |
|
||
| TLS verification preserved | PASS | §8 — no bypass exists; AST-enforced |
|
||
| QuickJS removed | PASS | §12 |
|
||
| Executable campaign scripting removed | PASS | §I |
|
||
| Settings narrowed to v1 model | PASS | §J |
|
||
| Connection diagnostics appropriate | PASS | five distinguished failure kinds + a model-not-found warning |
|
||
| Hardcoded 120 s limitation resolved | **PASS after the §9.2 fix** | Was PARTIAL as committed: the setting existed but was inert |
|
||
| Supported launch paths loopback-only | PASS | §11 — all four paths, test-enforced |
|
||
| Offline operation preserved | PASS | §5.1, §5.4 |
|
||
| Story/tree/retry preserved | PASS | §6.4 — `take_count 2`, both attempts retained |
|
||
| Context inspection preserved | PASS | §5.8 |
|
||
| Memory isolation preserved | **PASS after the §9.1 fix** | §5.6 — negative control holds. Was FAIL as committed: the bank was dead |
|
||
| No M3 work started | PASS | Undo still destructive, no Redo, no head cursor |
|
||
|
||
Two rows are "PASS after the fix". As committed at `8c65ae9` they were **FAIL**
|
||
and **PARTIAL**. Both fixes are in the working tree; neither is a blocker,
|
||
because both are corrected and tested — but the commit itself does not meet the
|
||
criteria, which is why §N recommends a follow-up commit before M3 begins.
|
||
|
||
---
|
||
|
||
# M. V1 acceptance-test mapping
|
||
|
||
| ID | Result | Notes |
|
||
| --- | --- | --- |
|
||
| **A01** offline startup | PASS | §5.1, §5.4 |
|
||
| **A02** loopback default | PASS | §5.2, §6.1, §11 |
|
||
| **A03** no cloud API key | PASS | Stronger than at M1: the field no longer exists |
|
||
| **A04** restart persistence | PASS | §5.8, §6.5 |
|
||
| **A05** failed model call integrity | PASS | §5.7 — accepted AI turns held at 3 across two induced failures |
|
||
| **A06** trusted-LAN Ollama | PASS | §6 — real second machine, HTTPS, private CA |
|
||
| **H01** no unexpected outbound | PASS | §5.9, §6.6 |
|
||
| **H02** no telemetry | PASS | §E |
|
||
| **H03** no cloud provider required | **PASS, and now in its preferred form** | H03 says the *preferred* final v1 is "cloud provider controls are absent, not merely unused". M1 met the base condition; **M2 meets the preferred one.** |
|
||
| **H04** model output cannot execute shell/tools | PASS, strengthened | The only execution surface was QuickJS, now removed. No shell, MCP or tool framework exists. |
|
||
| **H10** local API/CORS behaviour | PASS, strengthened | `AIDND_CORS_ORIGINS="*"` refuses to start; an unknown `/api` path 404s instead of returning HTML 200 |
|
||
| **H11** no first-use runtime download | PASS | §10.1, §10.2 |
|
||
|
||
Two acceptance tests deserve planning attention (§O): A05's wording, already
|
||
corrected post-M1, and H10, which M2 has now exceeded in a way the test does not
|
||
describe.
|
||
|
||
---
|
||
|
||
# N. Security review of the simplified product
|
||
|
||
## Resulting architecture
|
||
|
||
```text
|
||
Browser (loopback only)
|
||
│ same-origin; CSP names no remote origin
|
||
▼
|
||
Storyteller — FastAPI + SPA, bound 127.0.0.1:8000, no authentication
|
||
│
|
||
├─► SQLite on the local filesystem (the only persistence)
|
||
│
|
||
└─► exactly one approved Ollama endpoint (endpoints.py)
|
||
same-host loopback OR a user-named address on their own network
|
||
HTTP, or HTTPS verified against the machine's CA store
|
||
```
|
||
|
||
## Every remaining way story text could leave the machine
|
||
|
||
Found by static search plus runtime confirmation, not by assuming dead code is
|
||
unreachable:
|
||
|
||
| Path | Status |
|
||
| --- | --- |
|
||
| `providers/openai_compatible.py` — 3 clients | The intended egress. Policy-checked before every request, TLS-verified. |
|
||
| `routers/settings.py` — 1 client | Model listing. Same policy, same trust context. |
|
||
| `endpoints.py` — `getaddrinfo` | Resolves names for the policy. Opens no connection. |
|
||
| Analytics / telemetry | **Gone.** |
|
||
| Remote runtime assets | **Gone** since M1; re-verified (§10.2). |
|
||
| Cloud provider constants | **Gone.** |
|
||
| Hosted auth | **Gone.** |
|
||
| Arbitrary endpoint use | **Refused** by address (§7). |
|
||
| Scripting runtimes | **Gone.** |
|
||
| Shell / MCP / tool frameworks | Never existed. |
|
||
| Remote database | **Gone.** |
|
||
|
||
Only four modules in `backend/app` can open an outbound connection at all, and
|
||
the only absolute remote URL left in backend code is a `SOURCE_URL` constant in
|
||
`encoding.py` documenting where the vendored tokenizer table came from — never
|
||
fetched (§8).
|
||
|
||
## Residual risk, stated plainly
|
||
|
||
1. The API is **unauthenticated by design**. Loopback binding is the whole
|
||
control. Publishing the port defeats it; the code refuses the easy mistakes
|
||
(wildcard CORS, the default compose mapping) and the documentation warns
|
||
about the deliberate one.
|
||
2. A hostile host **on the user's own LAN** is permitted by the endpoint policy,
|
||
because that is what trusted-LAN means.
|
||
3. The **Ollama service has its own network behaviour** (`ollama.com`), outside
|
||
this codebase and inside the user's trust boundary. Unchanged from M1 and
|
||
still worth settling before release.
|
||
|
||
---
|
||
|
||
# O. Planning-document recommendations
|
||
|
||
Reported, not applied. No planning document was edited by this task.
|
||
|
||
### 1.
|
||
```text
|
||
Document: BUILD-MILESTONES.md, M5
|
||
Recommended change: Note that eight test files now use the world-state delta
|
||
protocol as instrumentation for state rollback, and that
|
||
replacing that protocol means moving the instrumentation
|
||
(one schema entry and one helper in tests/fakes.py).
|
||
Why M2 evidence: §K.2 — the QuickJS counter those tests used was replaced
|
||
with the RPG engine, which M5 plans to replace in turn.
|
||
Urgency: later (before M5 is briefed)
|
||
```
|
||
|
||
### 2.
|
||
```text
|
||
Document: SECURITY-THREAT-MODEL.md
|
||
Recommended change: Record the endpoint policy as implemented — an explicit
|
||
allowlist of networks, applied on save and before every
|
||
request — together with what it does not guarantee: a
|
||
hostile host on the trusted LAN, and the rebinding window
|
||
between the policy's resolution and the connection.
|
||
Why M2 evidence: §F. The threat model predates the policy and describes the
|
||
inherited SSRF guard's opposite rule.
|
||
Urgency: before M3
|
||
```
|
||
|
||
### 3.
|
||
```text
|
||
Document: V1-ACCEPTANCE-TESTS.md, H10
|
||
Recommended change: H10 currently reads only "privileged local APIs do not
|
||
allow arbitrary wildcard cross-origin writes". M2 goes
|
||
further: a wildcard origin makes the app refuse to start,
|
||
and an unknown /api path 404s rather than returning the SPA
|
||
with status 200. State both as pass conditions.
|
||
Why M2 evidence: §11. Both were defects found and fixed during M2; without
|
||
a test that names them they can regress unnoticed.
|
||
Urgency: later
|
||
```
|
||
|
||
### 4.
|
||
```text
|
||
Document: TECHNICAL-DESIGN.md §5.1
|
||
Recommended change: Mark items 3 and 4 of the hardening list resolved. Item 3
|
||
(hosted/auth/analytics/Postgres/cloud/QuickJS) is done in
|
||
full; item 4 (endpoint validation reflecting this product's
|
||
threat model) is done by endpoints.py.
|
||
Why M2 evidence: §C, §F. The document currently marks both open — M1 left
|
||
them so.
|
||
Urgency: before M3
|
||
```
|
||
|
||
### 5.
|
||
```text
|
||
Document: DECISIONS/ — a new ADR
|
||
Recommended change: Record the endpoint policy as a decision: address-based
|
||
allowlist rather than hostname matching, deny by default,
|
||
enforced at save and at request time, never traded against
|
||
TLS verification. Include the ipaddress-classification
|
||
finding as the reason the CIDRs are spelled out.
|
||
Why M2 evidence: §F, §K.1. This is a load-bearing security decision that
|
||
currently exists only as a module docstring.
|
||
Urgency: before M3
|
||
```
|
||
|
||
### 6.
|
||
```text
|
||
Document: SPECIFICATION.md
|
||
Recommended change: None. M2 changed no product requirement; it removed
|
||
capability the specification never asked for.
|
||
Urgency: —
|
||
```
|
||
|
||
---
|
||
|
||
# P. Technical debt after M2
|
||
|
||
| Item | Classification | Note |
|
||
| --- | --- | --- |
|
||
| **The three §9 fixes are uncommitted** | **pre-release — do first** | Six files ahead of `8c65ae9`. Until committed, the branch's HEAD contains a dead memory bank. |
|
||
| `Settings.model` defaults to `""`; nothing prompts | M8 | Inherited from M1. Diagnosis improved, onboarding not. |
|
||
| Ollama's own `ollama.com` lookup | pre-release | Outside this codebase, inside the product's claim about itself. |
|
||
| Inert tables and columns | optional cleanup | Safe legacy remnants. A cleanup migration is cheap once the schema settles — after M3/M5, not before. |
|
||
| `migrations.py` at 1 128 lines | optional cleanup | Dead/inert history, not active behaviour (§K.3). |
|
||
| 4 pre-existing unused imports | optional cleanup | Present since M1; `pyflakes` is otherwise clean. |
|
||
| Unused `request: Request` parameters in three routers | optional cleanup | Left by removing the rate limiter. Harmless. |
|
||
| No frontend tests | M8 | None existed at M1 either. Lint and build only. |
|
||
| `docs/*.html` links Google Fonts | M8 or pre-release | Upstream's project site; not served by the app. |
|
||
| Rebinding window in the endpoint policy | optional | §F. Requires an attacker already on the trusted LAN. |
|
||
|
||
**Nothing here is a blocker for M3** except committing the fixes, which is
|
||
housekeeping rather than engineering.
|
||
|
||
---
|
||
|
||
# Q. M3 readiness
|
||
|
||
M3 is the production history work: non-destructive Undo, Redo, active-head
|
||
movement, divergence, and active-head export/import.
|
||
|
||
**Did M2 alter the chokepoints M3 will modify?** Barely, and helpfully:
|
||
|
||
| M3 chokepoint | M2's effect |
|
||
| --- | --- |
|
||
| `attempts.py` — snapshot/restore/rollback | `ATTEMPT_KEYS` lost `"script"`; `restore_state` and `snapshot_outcome` no longer carry script state. **One less shared state to move.** |
|
||
| `tree.py` — placement, head, lineage | **Untouched.** |
|
||
| `context/lineage.py`, `context/history.py` | **Untouched.** |
|
||
| `routers/adventures/takes.py` — retry, takes, undo | Only the removal of `ScriptPipeline` arguments and the demo cap. `undo_turn` is byte-for-byte the inherited destructive implementation. |
|
||
| `routers/adventures/turns.py` | Hook calls removed, so `generate_turn` has 4 parameters instead of 5 and no longer branches on script stop. **Simpler to reason about.** |
|
||
| `bundle.py` — export/import | `scripts`/`scriptState` no longer exported; old bundles importing them are ignored gracefully. M3 adds the active-head field to a smaller format. |
|
||
|
||
**Is the Phase 0B undo/redo spike still applicable?** Yes. It touched three
|
||
backend files — the tree, the attempt machinery and the takes router — and M2
|
||
changed none of their logic. If anything the spike is easier to apply now: it no
|
||
longer has to carry `script_state` alongside `world_state` through every
|
||
rollback.
|
||
|
||
**Did the removals simplify or complicate M3?** Simplified, in three concrete
|
||
ways: one shared state instead of two through the rollback paths; no demo-cap or
|
||
rate-limit preconditions wrapped around the turn and retry endpoints; and a
|
||
provider constructed from `Settings` directly rather than through a
|
||
`ProviderConfig` indirection.
|
||
|
||
**Tests M3 should preserve or rewrite.** Preserve as the contract:
|
||
`test_story_tree_baseline`, `test_retry_variants`, `test_take_state`,
|
||
`test_take_parentage`, `test_attempt_siblings`, `test_branch_forking`,
|
||
`test_delete_state`, `test_state_revert`, `test_bundle_v2`. **Note that these now
|
||
carry the world-state instrumentation** (§K.2) — M3 must not mistake it for RPG
|
||
coverage. `test_state_revert` and `test_delete_state` assert *destructive* undo
|
||
semantics and will need rewriting when undo stops deleting; that is expected M3
|
||
work, and those files should be rewritten rather than deleted.
|
||
|
||
**Anything to fix before history changes?** Only the uncommitted fixes. Nothing
|
||
M2 discovered touches history semantics.
|
||
|
||
**Should M3 proceed as planned?** **Yes, unchanged in scope.** Two additions to
|
||
its brief: it inherits a single shared state rather than two, and it must be told
|
||
that the state-rollback tests are instrumented through the world-state engine.
|
||
|
||
---
|
||
|
||
# R. Final recommendation
|
||
|
||
```text
|
||
Recommendation: ACCEPT M2 WITH NON-BLOCKING DEBT AND PROCEED TO M3
|
||
```
|
||
|
||
**Reasons.**
|
||
|
||
1. Every M2 requirement is met, and the ones that matter most were tested at
|
||
runtime rather than read: cloud endpoints refused against a hand-edited
|
||
database, trusted-LAN HTTPS working against a real second machine with
|
||
verification on, and captures showing zero packets outside loopback and the
|
||
approved host.
|
||
2. The product is materially simpler, not just smaller: 52 API routes → 36, ten
|
||
environment variables → two, 6 Python packages and 21 npm packages gone,
|
||
−2 666 backend lines, −1 284 frontend lines, and a 933 kB bundle down to
|
||
395 kB. The trust surface shrank in the places that carry story data.
|
||
3. H03's *preferred* v1 condition — cloud controls absent rather than unused —
|
||
is now met, which M1 could not claim.
|
||
4. No M1 capability is regressed **in the working tree**. Two were regressed in
|
||
the commit; both are fixed and both now have the tests that would have caught
|
||
them.
|
||
5. M3's chokepoints are untouched or simplified, and the Phase 0B spike still
|
||
applies.
|
||
|
||
**The one thing to do first** is commit the six-file fix (§A.1). Until then the
|
||
branch head contains a silently broken memory bank, and anyone building from
|
||
`8c65ae9` would inherit it.
|
||
|
||
The three planning corrections marked *before M3* in §O — the threat model, the
|
||
technical-design hardening list, and a new ADR for the endpoint policy — are
|
||
documentation of decisions already made, not new work, and can be done alongside
|
||
the M3 brief.
|
||
|
||
---
|
||
|
||
# S. Closeout note — appended 2026-09-03
|
||
|
||
**Appended after the fact.** Sections A-R above were written on 2026-09-02 and
|
||
are left as they were, including §A.1's and §P's statements that the fixes were
|
||
uncommitted at the time. Those statements were true when written and are the
|
||
reason this note exists rather than an edit.
|
||
|
||
## The one thing §R said to do first
|
||
|
||
§R closed with: *"The one thing to do first is commit the six-file fix (§A.1).
|
||
Until then the branch head contains a silently broken memory bank."*
|
||
|
||
Done:
|
||
|
||
```text
|
||
8652fe7cd84bca5173abb03b2a692f15fea8a98c M2 review: two regressions the green suite hid, and the reports
|
||
```
|
||
|
||
The commit carries the six implementation/test/lockfile files **and** these two
|
||
reports. `8c65ae9` remains the M2 implementation commit and was not amended or
|
||
squashed, so the provenance distinction §R wanted — original implementation
|
||
versus review-discovered correction — survives in the history.
|
||
|
||
Confirmed present in that commit, against `8c65ae9`:
|
||
|
||
| Defect | Fix as committed |
|
||
| --- | --- |
|
||
| §9.1 dead memory bank | `memorybank.py` — both provider factories stop reading the removed `Settings.api_key_plain`; `summary_provider` now passes `model_timeout_seconds` |
|
||
| §9.2 inert timeout | `turns.py`, `chat.py` and the summariser factory pass `settings.model_timeout_seconds` into the provider |
|
||
| §9.3 stale lock | `requirements.lock` drops `quickjs`, `psycopg`, `psycopg-binary`, `cryptography`, and the transitive `cffi` and `pycparser` |
|
||
| — | `test_local_only_surface.py` gains the two regression tests: every factory built from a real `Settings` row, and the configured timeout arriving at each generating client |
|
||
|
||
The separate short timeout classes were kept, as §9.2 required: `CONNECT_TIMEOUT`
|
||
(10 s), `EMBED_READ_TIMEOUT` (60 s) and the connection-test/model-list path are
|
||
unchanged, and only the generation read timeout became configurable.
|
||
|
||
## Verification at closeout
|
||
|
||
Re-run on 2026-09-03 against the **committed** tree, working tree clean:
|
||
|
||
```text
|
||
backend pytest tests/ -q 606 passed in 133.26s
|
||
tests/test_local_only_surface.py 32 passed
|
||
test_endpoint_policy + test_tls_trust + test_offline_assets
|
||
47 passed
|
||
frontend npm run lint 7 warnings, 0 errors, exit 0
|
||
frontend npm run build 395.41 kB, exit 0
|
||
root docker build exit 0
|
||
image quickjs/psycopg/cryptography/cffi/pycparser absent (32 packages)
|
||
```
|
||
|
||
The packet-capture exercises were not repeated; the corrected files do not touch
|
||
the network boundary. See §15 of the baseline report.
|
||
|
||
## Planning recommendations from §O
|
||
|
||
Three were marked *before M3* and are now applied, in commit
|
||
`Planning: record M2 closeout decisions`:
|
||
|
||
- **§O.2** — `SECURITY-THREAT-MODEL.md` §10A records the endpoint policy as
|
||
implemented, with both residual limits stated; §71A item 5 is marked resolved;
|
||
§77 notes the required defaults are now met.
|
||
- **§O.4** — `TECHNICAL-DESIGN.md` §5.1 items 3 and 4 are marked done, and a new
|
||
§5.2 records the M1/M2 production architecture as fact rather than intention.
|
||
- **§O.5** — **ADR 011, *Local Inference Endpoint Policy***, records the
|
||
decision, including the `ipaddress`-classification finding from §K.1 as the
|
||
reason the CIDRs are spelled out.
|
||
|
||
The remaining three are also applied, ahead of the *later* urgency §O gave them,
|
||
since they are one-paragraph edits: **§O.1** as an M5 note in
|
||
`BUILD-MILESTONES.md` about the world-state instrumentation, and **§O.3** as a
|
||
strengthened H10 in `V1-ACCEPTANCE-TESTS.md`. **§O.6** correctly asked for no
|
||
change to `SPECIFICATION.md`, and none was made.
|
||
|
||
Two additions beyond §O, both drawn from evidence in this report:
|
||
|
||
- an M6 note in `BUILD-MILESTONES.md` requiring background memory failure to be
|
||
observable and at least one real provider-construction path to be tested —
|
||
§A.1 and §9.1 are the argument for it;
|
||
- **H12, *Inference Endpoint Enforcement***, in `V1-ACCEPTANCE-TESTS.md`, whose
|
||
fourth pass condition is the database-edited-behind-the-API case this review
|
||
demonstrated at runtime in §F.
|
||
|
||
`BUILD-MILESTONES.md` also gained M2's `## Status: COMPLETE` block, matching M1's.
|
||
|
||
## Not done in this closeout
|
||
|
||
M3 was not begun. Undo still deletes, and there is still no Redo. The debt table
|
||
in §P is unchanged apart from its first row, which this note closes.
|