Apply post-M1 corrections to the planning package

This commit is contained in:
JesseMarkowitz
2026-09-02 06:03:10 -04:00
parent 645f07f06d
commit 1a28a9a708
8 changed files with 304 additions and 35 deletions
+59 -6
View File
@@ -64,10 +64,13 @@ Recommended environment:
- one installed narrator model,
- one installed embedding model if semantic retrieval is enabled,
- outbound Internet blocked after setup,
- fresh test data directory.
- fresh test data directory,
- for A06, a second user-controlled machine on the trusted LAN serving HTTPS
with a locally issued certificate.
Record:
- OS,
- **CPU (core count), GPU (or explicitly none), and RAM**,
- application commit/version,
- Ollama version,
- narrator model,
@@ -75,6 +78,14 @@ Record:
- browser,
- test date.
Hardware is not bookkeeping. A model's **cold load** time depends on it, and
M1 measured a cold `qwen2.5:3b-instruct` load on a GPU-less four-core host
exceeding the inherited 120-second client timeout — three times — while the
same turn completed in 6 to 9 seconds once the model was resident. A timeout
result is therefore uninterpretable unless the hardware and the warm/cold state
are recorded with it, and a pass on a GPU machine does not predict a pass on a
CPU-only one.
## 4. Standard Test Campaign
Create a campaign named:
@@ -284,9 +295,29 @@ Transcript and authoritative current state are restored.
5. Reopen campaign.
### Pass
- prior story remains intact,
- failed turn is not partially committed as accepted,
- user can retry.
- **previously accepted history and state are unchanged** — every earlier turn,
its text and the authoritative state remain exactly as they were,
- **no AI response is accepted for the failed turn**: no partial or truncated
narration is committed, and the count of accepted AI turns does not move,
- the failure is reported to the user rather than swallowed,
- the user can retry and continue.
### Note on wording
"Nothing was committed" would be misleading, and this test should not be read
that way. The user's **submitted text is deliberately retained**: it is
committed before the model is called, so a model failure never discards what
the player typed. A failed turn therefore leaves the player's input at the head
of the story with no reply, and the total row count grows by one.
The invariant is about *accepted* history, not about row counts. What must
never happen is a partially generated AI response entering the story as though
it were accepted, or an earlier turn being altered or lost.
A test that asserts the story is byte-identical before and after will fail for
the wrong reason. Assert instead on the accepted prefix — for example, a digest
over every action up to the pre-failure head — and on the number of accepted AI
turns.
---
@@ -296,13 +327,18 @@ Transcript and authoritative current state are restored.
### Preconditions
- storyteller and browser run on machine A,
- Ollama runs on a separate user-controlled machine B on the trusted LAN,
- Ollama runs on a separate user-controlled machine B on the trusted LAN —
**a genuinely separate machine**, not another container or namespace on
machine A,
- **machine B serves HTTPS with a certificate issued by a private/local CA**,
and that CA is installed in machine A's operating-system trust store,
- required models are already installed,
- outbound Internet access is blocked.
### Steps
1. Keep the storyteller UI/API bound to loopback on machine A.
2. Configure the storyteller's Ollama endpoint to machine B.
2. Configure the storyteller's Ollama endpoint to machine B, as an `https://`
URL using the hostname the certificate is issued for.
3. Verify model discovery/connection diagnostics.
4. Generate at least three story turns.
5. Trigger state extraction and embeddings/memory retrieval if enabled.
@@ -311,11 +347,28 @@ Transcript and authoritative current state are restored.
### Pass
- story operation succeeds through the explicitly configured LAN Ollama host,
- **TLS is verified, not bypassed**: the certificate chains to the CA installed
on machine A and the hostname is checked; no "insecure" option was used,
because none exists,
- no cloud API key or Internet access is required,
- inference/model traffic goes only to the approved LAN host,
- storyteller UI/API remains loopback-bound,
- prompts, state, retrieved knowledge, and embedding inputs do not go to any unapproved destination.
### Note on sufficient evidence
Added after M1. **A plain-HTTP LAN test is no longer sufficient evidence for
this test.** Certificate verification never happens over cleartext, so an
HTTP-only run cannot exercise the path that actually broke: against a real
HTTPS LAN host, the application refused an endpoint that `curl` and the browser
on the same machine accepted, because it verified against a bundled public-CA
list instead of the machine's own trust store (ADR 002, *Transport for a
Trusted-LAN Endpoint*).
A container or network namespace standing in for machine B is likewise not
sufficient on its own. It exercises the non-loopback address but typically
speaks plain HTTP, and it will hide exactly this class of defect.
---
# B. Basic Story Interaction