Files
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00

1184 lines
74 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# v1.1 WP-A1 and WP-A2 — Implementation and Verification Report
**Work packages:** WP-A1 (context-window safety reserve) and WP-A2 (protocol-echo
cleanup and genre-neutral state prompting), authorised together and implemented
in that order. They are reported separately throughout.
**Status:** COMPLETE, including the owner-requested corrective work (§R). Staged,
uncommitted, for the owner's signed commit (2026-09-14). **Final decisions: §R.6.**
WP-A1: PASS. WP-A2: PASS.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| Starting commit | `ac465ed` — *Planning v4.1: record the v1.0.0 release, and plan v1.1*, signed by the owner (`%G?` = `G`) |
| Its parent | `432f041`, the signed v1.0.0 release commit; tag `v1.0.0`; `main` and `origin/main` |
| Working tree at start | clean; nothing staged, nothing unstaged, nothing untracked |
| Planning baseline | committed before any product change, so it is not in this diff |
**Baseline backend suite, before any change** (the untouched `ac465ed` tree,
`python -m pytest tests/ -q`, no `AIDND_TEST_*` set): **1,421 passed, 17 skipped,
0 failed**, 865 s. This matches the v1 closeout's count exactly.
---
## B. WP-A1 implementation
### B.1 What the code did before, answered from the code
| Question | Answer on `ac465ed` |
| --- | --- |
| 1. What defines the effective context window? | `contextwindow.effective_budget(configured, window)`, which is `min(context_token_budget, window.tokens)` when a window is enforceable (verified by `/api/ps` or `/api/show`, or declared through `context_window_override`), and the configured budget otherwise. |
| 2. Where is the fixed 64-token margin applied? | `context/builder.py`: `OUTPUT_SAFETY_MARGIN = 64`, added to `max_output_tokens` as `output_reserve`. That is subtracted with the protected sections before any knowledge or history is priced. |
| 3. How is `max_output_tokens` reserved? | Inside that same `output_reserve`. |
| 4. Which components are protected? | Every system section (narrator prompt, state rule, campaign canon, always-include Canon and the knowledge rule, AI instructions, persona, plot essentials), plus the summary, retrieved memories, narrative state, author's note, front memory, length hint, refusal note and state reminder. |
| 5. Where is history trimmed? | `build_context`: `history.window_covering`, then `history_floor` and `trim_block` (§15.3), then newest-first until `history_budget` is spent. Knowledge is chosen before history, out of its own share. |
| 6. What counts tokens before the request? | `cl100k_base` from the vendored table (`context/encoding.py`), through `builder.count_tokens`. |
| 7. Which response fields give the real prompt-token count? | `usage.prompt_tokens` in the OpenAI-compatible response. **Measured on Ollama 0.33, a stream sends no `usage` unless `stream_options.include_usage` is set.** Chat and completion modes both report it once asked, and non-streaming requests always do. |
| 8. Where is that metadata persisted? | `snapshot["usage"] = provider.last_usage` in `turns.py`, inside the turn's compressed `context_snapshot`. Because streams were never asked, **none of the 514 AI turns in the v1 evidence databases has a stored usage**. |
| 9. Can the status live in provenance without a migration? | Yes. `context_snapshot` is a compressed JSON column, and the per-action context route returns it whole. |
| 10. What happens if protected context alone is too large? | `ContextOverflow` is raised before the model call. The turn route yields it as a turn error, and the dry-run route returns 422. Nothing is written. |
### B.2 What the 64 tokens were, and what replaced them
M6's comment gave the margin two jobs:
- absorb "the separators between sections" that are added after the arithmetic;
- absorb "the difference between our tokenizer's count and the serving model's".
The first job is not drift. It is application text nobody priced: `SEPARATOR`
joins between sections, and `CHAT_CONTINUE_HINT`, which the provider appends to
every chat request. v1.1 prices it exactly as `tokens.transport`:
- one separator per system section and per story section slot
(`STORY_SECTION_SLOTS = 11`), which over-counts by at most two;
- plus the hint.
The second job is what the reserve now does. The reply allocation is exactly
`max_output_tokens`. Nothing is double-counted: the 64 tokens are gone, and
each of their two jobs has one owner.
### B.3 The reserve
`contextwindow.safety_reserve(effective_budget) = max(256, ceil(effective_budget × 5 / 100))`.
The rounding is up, to a whole token, in integer arithmetic.
| Effective window | Reserve |
| --- | --- |
| 4,096 | 256, because 204.8 is below the floor |
| 5,120 | 256, exactly 5% |
| 8,192 | 410, from 409.6 |
| 16,384 | 820, from 819.2 |
| 32,768 | 1,639, from 1,638.4 |
It follows the *effective* budget. A 16,384 setting against a 4,096 server
reserves 256. It is not a setting, and it is not calibrated per model.
Protected context is `sections + transport + max_output_tokens + reserve`.
`ContextOverflow` now names each part, and keeps the M11 advice sentence when
the server's window is what binds.
### B.4 Accounting after the reply
`contextwindow.classify_usage(usage, estimate, budget, max_output_tokens, window_verified)`.
`estimate` is `tokens.estimate`: the assembled text plus what the provider adds
(the chat hint, or the completion separator). The checks run in this order:
| Status | Condition | What it means |
| --- | --- | --- |
| `unknown` | no positive integer `prompt_tokens` | Nothing can be said. Never reported as `fits`. |
| `truncation_suspected` | `server + reserve < estimate` | The server read fewer tokens than were sent, by more than the drift tolerance. This is the measured shape of an over-window prompt: 6,316 sent, 2,050 read. |
| `exceeded` | `server + max_output_tokens > budget` | The drift used up the whole reserve, so the reply may be cut short. |
| `fits` | otherwise | The server read the whole prompt. |
The record also carries `server_prompt_tokens`, `difference` (server minus
estimate), `observed_margin` (`budget - reply - server`), `safety_reserve`,
`window_verified` and a sentence of detail.
**Why `exceeded` is measured against the hard edge and not the reserve line.**
On prompts that fit, the server counted 13 tokens more than the application,
which is chat-template overhead. A full prompt built exactly to the reserve line
would then read as "exceeded" on every turn. The reserve is the tolerance, so
using part of it is `fits` with a smaller `observed_margin`. Using all of it is
`exceeded`.
**Surfacing.**
- The record is stored as `snapshot["accounting"]`, so it is available for every
past turn.
- It is returned on the turn's `done` SSE event.
- `exceeded` and `truncation_suspected` are logged at WARNING.
- The context inspector shows the safety margin on every turn, and the
accounting for a turn that was sent. The two bad states are a `role="alert"`
notice that says the turn is kept.
**The turn is kept.** Accounting happens after the reply has streamed, and it
changes nothing about whether the turn commits.
### B.5 Files
| File | Change |
| --- | --- |
| `backend/app/contextwindow.py` | `SAFETY_RESERVE_FLOOR`, `SAFETY_RESERVE_PERCENT`, `safety_reserve`; the `FITS` / `EXCEEDED` / `TRUNCATION_SUSPECTED` statuses; `classify_usage` |
| `backend/app/context/builder.py` | `OUTPUT_SAFETY_MARGIN` removed; `transport`, the reserve, and an exact reply allocation; `ContextOverflow` names each part; the report gains `transport`, `safety_reserve` and `estimate` |
| `backend/app/providers/openai_compatible.py` | `STREAM_OPTIONS = {"include_usage": True}` on all four streaming bodies |
| `backend/app/routers/adventures/turns.py` | `snapshot["accounting"]`, a WARNING log, and `accounting` on the `done` event |
| `frontend/src/pages/Play/panels/ContextPanel.jsx`, `frontend/src/styles/context.css` | the safety-margin line and the accounting notice |
| `backend/tools/m11_long_run.py` | the timeline records the turn's accounting from the `done` event |
| `backend/tools/v11_window_accounting.py` | **new**: the real-model accounting harness (§D) |
| `planning/TECHNICAL-DESIGN.md` §15.2, `DEVELOPMENT.md`, `README.md` | as-implemented notes. The README's stale Screenshots paragraph and test count are corrected, as the v1.1 plan asked |
No schema change, no migration, and no bundle-format change.
---
## C. WP-A1 tests
`tests/test_v11_context_reserve.py`, new:
| Criterion | Tests |
| --- | --- |
| A1-1 reserve calculation | `test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up` (1,024; 4,096; 5,120; 5,121; 8,192; 16,384; 32,768), and `…far_larger_than_the_v1_margin…` |
| A1-2 budget enforcement | `test_the_prompt_leaves_the_reply_and_the_reserve_free`, on the **assembled text plus the chat hint**, for verified 4,096, 8,192 and 16,384, a declared 6,000, and an unverified configured 12,000, each against a 120-turn story. Also `test_the_report_prices_the_text_the_provider_adds` and `test_the_reserve_follows_the_effective_window_not_the_setting` |
| A1-3 protected overflow | `test_protected_context_that_only_fits_without_the_reserve_fails_explicitly`: a canon grown until v1's 64-token margin would still have built the prompt and the 256-token reserve does not. Also `test_an_overflowing_turn_never_reaches_the_model`: the scripted provider is called **0** times and no AI action is written |
| A1-4 no silent canon dropping | the canon sentinel is asserted present in every A1-2 configuration while the history gives way. M11's `test_the_canon_at_the_front_survives_a_window_far_too_small` and its negative control still pass |
| A1-5 accounting | `test_a_prompt_the_server_read_in_full_fits`; `test_a_turn_records_what_the_server_read`, end to end, including the `done` event |
| A1-6 unknown | `test_no_usable_count_is_unknown_never_fits`: None, `{}`, no `prompt_tokens`, 0, a string and a negative value. Also `test_a_turn_with_no_reported_usage_is_unknown` |
| A1-7 discrepancy recorded and visible | `test_a_suspected_truncation_keeps_the_turn_and_says_so` (snapshot, `done` event, WARNING log, and the per-action context route); `test_a_server_that_read_far_less…`; `test_a_small_undercount_is_tokenizer_drift_not_truncation` (reserve minus 1 is `fits`, reserve plus 1 is suspected); `test_a_server_that_counts_more…_is_exceeded` |
| A1-8 turn preserved | `test_a_suspected_truncation_keeps_the_turn_and_says_so`, `test_an_exceeded_turn_is_also_kept` |
| request | `test_the_stream_asks_the_server_to_report_its_usage`, in chat and completion modes |
| each take's own accounting (added after the A2 long run; §N.2 item 7) | `test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves`, on `keep_own_slices` and `hand_over_the_prompt`. `test_a_retry_leaves_each_take_with_its_own_accounting`, end to end through `POST /retry`: the superseded take keeps its `fits` and no prompt; the live take keeps the prompt and its own `truncation_suspected` |
The file has 35 tests: 33 at the A1 checkpoint, plus the two added above.
Frontend, `panels.test.jsx`, four new tests: the safety margin is shown; nothing
is shown for a turn not yet sent; `truncation_suspected` is an alert naming both
counts; `unknown` makes no claim that the prompt fitted.
**One existing test changed.** It is a fixture calibration, not a weakened
assertion.
- **The failure.** `test_history_block_trim.py::test_the_story_prompt_keeps_its_prefix_across_a_new_turn`
asserted that the *next* turn after its fixture holds the history floor. Its
budget is 2,048 tokens, where the reserve is 256, or 12.5% of the window. The
history budget fell from 1,186 to 969: the reserve's extra 192 over the old
margin, plus 25 tokens of transport. The block is 2 on both trees. The traces:
```text
turn v1.0.0 floor v1.1 floor
<60 54 54
<61 54 56 <- v1.1 steps here, v1.0.0 one turn later
<62 56 56 prefix kept 0.906
<63 56 58
<64 58 58 prefix kept 0.906
```
- **The judgement.** Same cadence and same step size, offset by one turn. §15.3's
property holds, and only the fixture's choice of turn moved.
- **The change.** The test now walks forward until a turn holds. It requires every
move on the way to be exactly one block, and asserts the prefix on the held
pair. A window that slides by one action every turn fails it exactly as
before, and its negative-control companion is unchanged.
---
## D. WP-A1 real-model evidence
**Harness:** `tools/v11_window_accounting.py`, on the A1 tree, with no source edit
during the run.
**Inference:** the trusted-LAN CPU reference host. It runs Ollama 0.33.0 over
HTTPS, with a private CA in this machine's OS trust store and verification on.
**Campaign:** the v1 evidence run's own bundle (`m04-final/bundle.json`, 207
actions), imported so that every turn is assembled against a full window. Memory
bank and auto-summarise are off, because they would only add post-turn calls on
the same host.
**Settings:** configured budget 16,384, `max_output_tokens` 500.
**Evidence:** `$HOME/v11-evidence/a1-accounting/`.
### D.1 The 4,096-window reference configuration (`qwen2.5:3b-instruct`, no `num_ctx`)
| Turn | Window | Effective budget | App estimate | Server `prompt_tokens` | Difference | Reserve | Observed margin | Status |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | **not verified** (model not resident) | 16,384 (configured) | 13,875 | **2,050** | −11,825 | 820 | — | **`truncation_suspected`** |
| 2 | 4,096, verified (`loaded`) | 4,096 | 3,294 | 3,309 | +15 | 256 | **287** | `fits` |
| 3 | 4,096, verified | 4,096 | 3,253 | 3,268 | +15 | 256 | **328** | `fits` |
| 4 | 4,096, verified | 4,096 | 2,831 | 2,846 | +15 | 256 | **750** | `fits` |
"Observed margin" is `window − reply allocation − the server's count`. The v1
evidence's equivalent at the largest prompts was 23-42 tokens.
**Turn 1 is the most important row in this section.** The model was not
resident, so `/api/ps` could not report its window. M11's rule for that case
leaves the configured 16,384 standing and records the window as unverified. The
application sent 13,875 tokens. The server loaded the model at its 4,096 default,
kept 2,050 of them, and answered HTTP 200. That is the silent truncation M8 found
and M11 set out to prevent, still reachable on a cold model's first turn.
v1.0.0 would have stored this turn with no record of it. A1 recorded
`truncation_suspected`, logged a warning ("The server read 2,050 prompt tokens of
the 13,875 sent…") and kept the turn. §N records the cold-model gap itself as a
finding.
**Turns 2-4:** the window was verified and the budget capped to it. The server
counted exactly 15 tokens more than the application on every turn: chat-template
overhead the application cannot see. That used 15 of the 256-token reserve, and
each turn kept 287 tokens or more beside the reply.
*One harness note.* The bundle carried the evidence campaign's memory bank
setting, so turn 1's post-turn memory pass ran and failed as the in-process
client closed (`derived memory work failed`). It is a failure in the harness,
outside the turn and outside what is measured. The tool was then changed to
switch memory and auto-summarise off before importing. The 16,384 run below used
the changed tool.
### D.2 The 16,384-window configuration (`qwen2.5:3b-instruct-16k`, `num_ctx` 16,384 baked in)
| Turn | Window | Effective budget | App estimate | Server `prompt_tokens` | Difference | Reserve | Observed margin | Status | Seconds |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | 16,384, verified (`parameters`) | 16,384 | 13,875 | 13,890 | +15 | 820 | **1,994** | `fits` | 1,028 |
The first run was **stopped after turn 1 at the owner's request**, because the
shared inference host was needed by another project. One turn took 17 minutes on
the CPU host. A second run of three turns was started when the host was released
(§D.3).
The window was verified from the model's own `num_ctx`, so even a cold model had
a ceiling. The prompt held 60 of 120 history actions. The server again counted 15
tokens more than the application. That left 1,994 tokens beside the reply,
against 23-42 in the v1 evidence at the same window and the same model.
### D.3 16,384 window, second run
Started when the host was released, with the A1 code loaded at process start
and before any A2 edit.
| Turn | Window | Effective budget | App estimate | Server `prompt_tokens` | Difference | Reserve | Observed margin | Status | Seconds |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | 16,384, verified (`parameters`) | 16,384 | 13,875 | 13,890 | +15 | 820 | 1,994 | `fits` | 1,058 |
**The run was stopped deliberately after turn 1** (10:03 EDT), to free the shared
host for WP-A2's real-model runs. Each 16k turn took about 17.6 minutes on the CPU
host, and the two remaining turns would have delayed A2's 50-turn run by more than
half an hour.
**Read this row as a reproduction, not a second sample.** Both 16k runs imported
the same campaign, so turn 1 assembled the same prompt. The identical counts show
that the measurement is repeatable. They add no new prompt shape. The prompts that
grow and change across many turns are measured at 16,384 by the integrated v1.1
release run (plan §11). In this report, the prompts that change turn by turn are
the 4,096 rows in §D.1 and the A2 long run in §J.
### D.4 What the evidence shows
| Criterion | Result |
| --- | --- |
| A1's tolerance is materially larger than v1's margin | At 16,384: 1,994 tokens left beside the reply, against 23-42. At 4,096: 287 or more. The reserve is 820 and 256. |
| Real drift between the counters | +15 tokens on every turn measured, at both windows. That is 6% of the smaller reserve and 2% of the larger. |
| Disagreement is observable | Every sent turn carried a server count and a status. A real cold-model truncation was caught, recorded and logged, and the turn was kept. |
---
## E. WP-A1 interim decision
**A1 RESULT: PASS.**
Recorded at the checkpoint before any WP-A2 product code was written.
- **A1-1 to A1-9** pass on the A1 tree:
- backend suite **1,454 passed, 17 skipped, 0 failed** (the 1,421 baseline plus
33 new);
- frontend **165 of 165**;
- lint clean apart from the six warnings the v1 closeout already had.
- **A1-10** passes on real-model evidence. At 4,096 there are four turns,
including a real `truncation_suspected`. At 16,384 there is one verified turn,
with the second run pending (§D.3).
- **One existing test was re-calibrated**, not weakened (§C). The failure was the
fixture's choice of turn. §15.3's property was unaffected, and the v1.0.0 and
v1.1 traces show it.
- **A1-only checkpoint** saved outside the repository, so A1 can be reviewed on
its own:
- `a1-tracked.patch`, sha256 `3b5d679f…`;
- `a1-untracked.tgz`, sha256 `78275d0f…`.
**Later corrective work on A1, found after this checkpoint.** The A2 long run
exposed a defect in A1: a retried or reselected take lost its accounting, or was
shown another take's. It was fixed with two tests (§N.2 item 7). The fix is one
line in `attempts.py`, plus the two tests. The A1-only checkpoint patch above
predates it. The final staged tree includes it. This interim PASS stands on
A1-1 to A1-10 as written, and §O records the decision after the fix.
**Carried to §N, not corrective work for A1:** the cold-model first turn. When
the window cannot be verified, M11 builds to the configured budget. A1 now makes
the resulting truncation visible, but does not prevent it. Preventing it means a
policy change: build conservatively, or refuse to send, when the window is
unknown. That is outside A1's authorised scope ("do not implement model-specific
dynamic calibration"; "preserve the accepted turn") and is the owner's decision.
---
## F. WP-A2 implementation
Begun only after §E was recorded. The A1-only checkpoint was saved first, so A1
can still be reviewed on its own.
| File | Change |
| --- | --- |
| `backend/app/narrative/events.py` | `vocabulary_for_prompt` shows each event as the JSON object to send, not as `name(field, …)`. `_PLACEHOLDER` holds the neutral placeholders. `SPECS` is unchanged, verified against `v1.0.0` by diff. |
| `backend/app/narrative/extract.py` | `EMIT_RULE`'s example now uses `character-1`, `item-1` and `location-1`. `LENGTH_HINT_OPENING` and `LENGTH_HINT_TAIL` are shared with the builder. R1-R4 and `RULE_*` added (§H). `_clean(…, after_block=)`. `explain_removed_line` added for the replay. |
| `backend/app/context/builder.py` | `length_hint` builds its opening and tail from the shared constants. The hint text is byte-identical to v1.0.0. |
| `backend/tools/m11_long_run.py` | The leak count also detects call lines and echoed hints. |
| `backend/tools/v11_replay_extractor.py` | **New**: the historical replay (§I). |
| `backend/tools/v11_compat_check.py` | **New**: the v1.0.0-database compatibility check (§K). |
| `backend/tests/test_v11_protocol_echo.py` | **New**: prompt, observed shapes, adversarial prose, rule attribution. |
| `planning/TECHNICAL-DESIGN.md` §15.4, `DECISIONS/013` | As-implemented notes. |
Unchanged: no event type, field, reference rule or validation. The fence protocol,
proposal recording, history replay and schema are also untouched.
**Two corrections made during A2, before its evidence was taken.**
1. **Misplaced docstring paragraphs.** The new paragraphs in `extract._clean` and
`events.vocabulary_for_prompt` were first placed after each docstring's
closing quotes, so the modules did not import. Collection failed at once and
nothing ran; the paragraphs were moved into the docstrings.
2. **The first vocabulary form cost too much.** `"<key>"`-style placeholders with
spaced separators took the vocabulary from 258 to 456 tokens and `EMIT_RULE`
from 467 to 690. That pushed protected context in an existing budget test
(2,048 tokens with an 800-token reply cap) to 2,055. It also left an
imported-knowledge fixture with no room to retrieve anything. Both failures
were real consequences of the prompt growth, not faulty tests. The full-suite
run that found them was stopped as void, and the form was changed to compact
JSON with ellipsis placeholders (§G). Both tests pass unchanged on the final
form, and no test was edited to accommodate it.
---
## G. Prompt changes
### G.1 Semantic before and after
| | v1.0.0 | v1.1 |
| --- | --- | --- |
| How the vocabulary is shown | `set_possession(item, owner) — gives an item to an owner`, a call notation that is not the wire format | `{"type":"set_possession","item":"…","owner":"…"} — gives an item to an owner`, the object itself |
| Optional fields | `name(field, optional?)` | ` (optional: description, aliases)` after the summary |
| List-valued fields | not distinguished | shown as `["…"]` |
| Identifier guidance | "short lower-case slugs (mara, silver-key, old-abbey)" | "short lower-case slugs … the example's identifiers are placeholders" |
| Worked example | `silver-key` owned by `aldric`, who is at `old-abbey` | `item-1` owned by `character-1`, who is at `location-1` |
| Length hint | unchanged text | unchanged text, built from shared constants |
| Reminder, canon, state rendering | unchanged | unchanged |
### G.2 Cost
| | v1.0.0 | v1.1 | Change |
| --- | --- | --- | --- |
| `events.vocabulary_for_prompt()` | 258 tokens | 378 | +120 |
| `EMIT_RULE` (which includes the vocabulary) | 467 | 588 | **+121** |
`EMIT_RULE` sits in the static system block, so it is priced once into the cached
prefix and into protected context on every turn. At a 4,096 window, +121 tokens is
about 3% of the window, and it comes out of the history. The cheapest option
measured was a plain `name: fields` list at 274 tokens. It was rejected because it
is not the wire format, and because a copy of it in prose is not a proposal the
extractor can recognise and remove. A copied JSON line is.
### G.3 Guard against reintroduction
`test_v11_protocol_echo.py` covers the following:
- **Fixture identifiers.** No identifier from either acceptance fixture, and no
genre noun, appears in `EMIT_RULE`, `EMIT_REMINDER`, the vocabulary, or any
length hint.
- **The worked example.** It is a block the extractor accepts, with allowed
event types.
- **Call notation.** No vocabulary line is a call.
- **Line shape.** Every line parses as the object, with exactly the required
fields.
- **Hint wording.** The length hint carries the constants the extractor
recognises.
---
## H. Cleanup rules (WP-A2)
### H.0 The evidence the rules were designed from
Before any extractor change, the whole v1 corpus was scanned. That covers every
AI turn in every evidence database under `$HOME/m11-evidence`, plus the exported
bundles. Stored text and stored raw replies were counted separately, and
duplicates were removed by content hash. The scan counted the candidate shapes,
and the near-misses a rule must not touch.
| Shape | Unique stored texts | Where |
| --- | --- | --- |
| A line opening with a vocabulary call | 4 | All four closeout identity turns (depths 10, 12, 14 and 20). |
| A vocabulary call anywhere | 4 | The same four lines. The corpus has no call in the middle of a sentence. |
| `[Hard limit` anywhere | 3 | See the three cases below. |
| A rendered `Scene: … (at …)` as the last line | 1 | Closeout browser-run3 action 3, **stored by the v1.0.0 product tree**. |
| `Scene:` lines anywhere | 62 | Mostly pasted state sections from runs taken before M11's `0c7316f` fix. The v1.0.0 extractor already handles those, and the replay (§I) starts from the raw reply. |
The three `[Hard limit` cases:
- **Identity depth 14:** the application's length hint, reworded at the front.
- **m04-rerun action 151:** the hint's own words, with the last sentence
replaced by "This story ends here." An empty dangling ```` ```json ```` fence
follows it.
- **browser-recheck2 action 5:** "[Hard limit: I have adhered to the prescribed
limit of words. Here's the narrated turn and the state update. Send your next
action.]" The model **invented** this; it is not application text.
### H.1 The rules
Each rule is anchored to something the application owns: the event vocabulary,
the length hint's wording, or the renderer's scene-line format. None of them
judges prose.
**R1 — event-call line** (`RULE_EVENT_CALL`)
| | |
| --- | --- |
| Source | `events.vocabulary_for_prompt` printed every event as `name(field, …)`, and the narrator copied the notation. |
| Matcher | A whole line that starts with a name in `events.SPECS` (case-insensitive) followed by `(`. Indentation and a leading `>` are allowed before the name, and spaces before the `(`. The whole line is removed, including anything after the call on that line, such as depth 20's "Adds the silver key…". Lines inside a fenced code block are never examined. |
| Rationale | `create_entity`, `set_possession` and the rest are this protocol's snake_case identifiers. A story line that *starts* with one of them followed by `(` is protocol. |
| False-positive defence | Not applied in the middle of a sentence (the engineer's whiteboard). Not applied to call-shaped lines whose name is not in the vocabulary (`> open_door(north)`). Not applied inside a story's own code block. Not applied to a vocabulary word with no `(` ("Create_entity is a terrible name"). |
**R2 — echoed length hint** (`RULE_LENGTH_HINT`)
| | |
| --- | --- |
| Source | `builder.length_hint`. Its opening `[Hard limit:` and its closing sentence are now named constants shared with the extractor (`LENGTH_HINT_OPENING`, `LENGTH_HINT_TAIL`). |
| Matcher | A bracket at the end of the reply, closed or cut off; the existing trailing loop already isolates it. Its contents must start with `Hard limit:` **and** carry the application's own wording: either the tail phrase "append the state block", or "turn must not exceed *N* words". The check goes in `_is_echoed_instruction`, beside the reminder and continue-hint checks M11 already has. |
| Rationale | Both phrases are the application's instruction text. The 3B narrator reworded the front ("your next turn") and kept the rest. |
| False-positive defence | Only a whole bracket at the end of the reply is removed. These all stay: "[Hard limit: 500 words]" read aloud mid-sentence; an in-world "[Hard limit of the reactor…]"; a final "[Hard limit: forty days, no extensions]". **The model's invented bracket in browser-recheck2 action 5 also stays**, because it carries neither phrase. That is a deliberate miss. |
**R3 — rendered scene line** (`RULE_SCENE_LINE`)
| | |
| --- | --- |
| Source | The first line of `render.for_prompt`: `Scene: <summary>`, plus ` (at <location>)` when the scene has a location. |
| Matcher | After the other trailing cuts, the reply's last line starts with `Scene: ` and meets one of two conditions. Either it ends in the renderer's `(at <location>)`, or protocol was also cut from the end of the same reply (a block, a hint, a call line or a dangling fence). |
| Rationale | The `(at …)` suffix is the renderer's syntax. Without it, a lone scene line directly above removed protocol belongs to the same pasted tail. |
| False-positive defence | Never removed mid-reply: "Scene: a kitchen, late." followed by story stays, and so does M11's existing "Scene: the docks at dawn." case. A screenplay-style last line, "Scene: take two, and nobody moves.", stays when nothing else was removed. **Accepted risk:** a story whose last line exactly imitates `Scene: … (at …)` loses that line. |
**R4 — empty dangling fence** (`RULE_EMPTY_FENCE`)
| | |
| --- | --- |
| Source | A reply that hit the output limit after opening ```` ```json ```` and before writing anything (m04-rerun action 151). |
| Matcher | A ```` ```json ```` or bare ```` ``` ```` opener as the reply's last line, with nothing after it. |
| Rationale | M11's `_is_opening_of_proposal` already cuts a dangling fence that stopped after `{`, on the principle that "a story's own code block is not that short, and one that is holds nothing to lose". An empty opener holds even less. |
| False-positive defence | Applies only when nothing at all follows the opener. A dangling fence with any content keeps M11's existing judgement. |
### H.2 What is deliberately not removed
- A fact restated inside a sentence (§P risk 4). Removing it would mean judging
prose.
- Model-invented headings that are not the renderer's (§P risk 6).
- Model-invented brackets that carry no application phrasing (browser-recheck2
action 5).
- Prose that follows a removed call line, even when it narrates the call ("You
create a new person, Mike, who has", identity depth 12). It is narration,
however poor.
---
## I. Historical replay
**Tool:** `tools/v11_replay_extractor.py`, on the final A2 extractor.
**Inputs:** every AI turn in the 17 evidence databases under `$HOME/m11-evidence`,
plus the closeout identity run's bundle. A turn is replayed from its stored
`raw_output` where one exists, so the comparison starts from what the narrator
actually sent. The identity bundle carries only stored text; for those turns the
v1.0.0 prose is the input itself.
**Comparison:** each input goes through the extractor **as tagged at `v1.0.0`**,
read with `git show`, and through the v1.1 extractor. Lines are compared with
trailing whitespace ignored.
**Evidence:** `$HOME/v11-evidence/a2-replay/replay.json` (sha256 `f04c54ff…`) and
`replay.md`, which holds the full v1.0.0 and v1.1 prose for every changed turn.
| | |
| --- | --- |
| Unique replies replayed | **518** (6 exact duplicates skipped) |
| Unchanged | **509** |
| Changed | **9** |
| Flagged (an unexplained removal, text added or rewritten, or more than 25% removed) | **0** |
| Removed lines by rule | R3 scene line 6, R1 event call 4, R2 length hint 2, R4 empty fence 1 |
The planning pass estimated about 453 turns; the actual count is 518. The
difference is the trial, browser and recheck databases, which the v1 figure of
443 did not count.
### I.1 Every changed turn
| # | Source | Removed | Rule | Share of prose removed | Review |
| --- | --- | --- | --- | --- | --- |
| 1 | browser-final action 3 (raw) | `Scene: Aldric and Mara by the fire in the Crooked Lantern (at The Crooked Lantern).` | R3 | 13.1% | The rendered scene line, directly above a ```` ```json ```` proposal in the raw reply. **Protocol.** |
| 2 | closeout browser-run3 action 3 (raw), **stored by the v1.0.0 product tree** | `Scene: Aldric and Mara by the fire in The Crooked Lantern. (at The Crooked Lantern)` | R3 | 20.7% | The same, in the renderer's exact form. **Protocol.** |
| 3 | m04-rerun action 151 (raw) | `Scene: Aldric and Mara in the Crooked Lantern, rain outside.`; `[Hard limit: this turn must not exceed 180 words, … This story ends here.]`; ```` ```json ```` | R3, R2, R4 | 11.9% | A pasted scene line, the application's length hint with its last sentence reworded, and a proposal opener the output limit cut off. **Protocol.** |
| 4 | m04-rerun action 153 (raw) | `Scene: Aldric and Mara in the Crooked Lantern, rain outside.` | R3 | 3.1% | The scene line directly above a ```` ```json ```` proposal. One trailing space was also trimmed from the story line before it; no word changed. **Protocol.** |
| 5 | m04-rerun action 157 (raw) | `Scene: Aldric and Mara in the Crooked Lantern, rain outside.` | R3 | 2.8% | The scene line directly above ```` ```json\n{ ````. **Protocol.** |
| 6 | identity depth 10 (stored) | `> Create_entity(new_person, "john", "character", "A determined team member …", ["jo…` | R1 | 6.4% | A call cut off mid-argument by the output limit. **Protocol.** |
| 7 | identity depth 12 (stored) | `> Create_entity(new_person, "mike", "character", "A new team member …", ["mike"])` | R1 | 4.8% | The call. The narration after it ("You create a new person, Mike, who has") is **kept**. **Protocol.** |
| 8 | identity depth 14 (stored) | `> Create_entity(mike, …)`; `Scene: Bill, Alice, Roger, John, and Mike at the table; John not yet arrived. (at The meeting room)`; `[Hard limit: your next turn must not exceed 180 words, … append the state block well inside the limit.]` | R1, R3, R2 | 23.4% | The whole pasted tail: the call, the rendered scene line, and the reworded length hint. **Protocol.** |
| 9 | identity depth 20 (stored) | `> set_possession(silver-key, "alice") Adds the silver key to Alice's possession.` | R1 | 4.1% | The call, with the fantasy example slug and a gloss on the same line. **Protocol.** |
**Genuine prose lost: none.** Each of the 13 removed lines is application
protocol or instruction text: a vocabulary call, the length hint, the renderer's
scene line, or an empty fence opener. On a line that also carried narration, only
depth 20's same-line gloss went with the call; that gloss describes the call, not
the story.
**A replay-tool defect, found and fixed before this result.** The first replay
flagged turn 4 as "text rewritten". The trailing space trimmed from the story line
was compared as a changed line. The tool now ignores trailing whitespace and says
why. The flag was the tool's, and no rule changed as a result.
---
## J. WP-A2 real-model evidence
**Inference:** the trusted-LAN CPU reference host. It runs Ollama 0.33.0 over
HTTPS with a private CA and verification on, and serves `qwen2.5:3b-instruct` at
a verified 4,096 window. Embeddings are `nomic-embed-text`. The storyteller was
on loopback.
**Tree:** the final A2 extractor and prompt. The full backend suite passed on
this tree before the long run was allowed to start. The launcher enforced that
gate.
### J.1 Multi-character identity diagnostic
The office meeting fixture (`TEST-CAMPAIGN-FIXTURE.md` Appendix A), 10 beats,
memory off. Evidence: `$HOME/v11-evidence/a2-identity`.
| | v1 closeout (§S.5 of the M11 report) | v1.1 |
| --- | --- | --- |
| Turns accepted | 10 of 10 | 10 of 10 |
| Identity signals | 0 | **0**; verdict "no objective identity defect detected" |
| Proposals naming an example identifier from the fixed prompt | 1 (`silver-key` given to Alice) | **0**. No `silver-key`, `aldric`, `mara` or `old-abbey`, and none of the new placeholders either |
| Proposal outcomes | 1 applied, 5 without a block, 3 unparseable, 1 refused | **10 accepted, 1 rejected** |
| Stored turns with a shape R1-R4 targets | 4 of 10 (depths 10, 12, 14, 20) | **0 of 10** |
| Stored turns with any echoed application instruction | 4 of 10 | **1 of 10**: depth 16, below |
**A leak A2 does not remove.** Depth 16 kept this tail at the end of the story:
```text
Scene: Bill, Alice and Roger at the table; John not yet arrived.
[Hard limit: this is now 180 words.]
[Reminder: end your reply with a `state` block listing the events your narration made true, with absolute values.]
[You don't need to continue; your turn must now be about John entering the room. Continue the story here, directly. Output only story text.]
```
The trailing cleanup works from the end upwards, and the last bracket is not
recognised: it is a reworded continue hint, and the recogniser matches that hint
only by its opening words. So the reminder above it is never at the end, the
reworded length hint carries none of the application's wording, and the scene
line is never last. **This is the conservative design behaving as documented. It
is still a real leak, on 1 of 10 turns.** The harness's leak counter did not see
it: it looks for proposals, pasted headings, calls and the hint's own wording,
not reminders. It was found by reading the stored narration. A narrow follow-up
is in §N.
### J.2 The 50-turn long run
`tools/m11_long_run.py --turns 50`: the fantasy Continuity Test with the full
schedule of history operations and the memory bank on. Evidence:
`$HOME/v11-evidence/a2-long-run`.
| | |
| --- | --- |
| Status | `complete`; no aborted reason, no failed reason |
| Accepted turns | **51**: 50 scheduled and the recall turn |
| Wall clock | 9,324 s. A turn took 47-243 s, median 165 |
| Genuine process restarts | 3, each compared `identical: true` |
| History operations performed | 2 Save Points, Undo/Redo, 2 retries, undo then divergence, take selection, a Save Point restore, and an induced failed call (HTTP 404; state unchanged) |
| Post-turn failures | **0**. 0 tracebacks and 0 `database is locked` in `server.log` |
| Memories and summaries written | 15 memories (9 in the bank at the end); 5 summaries |
| **Protocol leakage** | **0 of 54 stored AI turns**, by the extended harness count (headings, `"events"`, call lines, the echoed hint). A separate scan of every stored turn for reminders, `[Hard limit`, continue-hint wording, calls, a final `Scene:` and `"events"` also found 0 |
| **Parser failures** | 6 proposals `unparseable`, out of 56 recorded |
| **Refused proposals** | 0 |
| **State updates from narration** | 4 events accepted from story (50 proposals accepted, most with an empty events list); 9 more from the harness's manual corrections |
| **Legitimate prose removed** | **none.** All 64 raw replies from this run and the identity run were replayed through the v1.0.0 and v1.1 extractors. One changed: a trailing ```` ```json ```` opener with nothing after it (R4, 0.4% of that reply). 0 flagged |
| M04-style recall | `recovered_through_state_only`, which is WP-B's subject and unchanged from v1 |
**A1 accounting across the run.**
| | |
| --- | --- |
| Window | verified at 4,096 on 51 of 51 turns |
| Status | **`fits` on 51 of 51** |
| Server count minus the application's estimate | **+15 on every turn** |
| Largest server prompt | 3,321 tokens |
| Observed margin beside the reply | 275 minimum, 620 median, 2,297 maximum, against a 256 reserve |
**A defect the long run exposed, fixed in A1 (§N.2 item 7).** Three of the 54
stored AI turns had no accounting: the two retried takes and the take superseded
by a take selection. `attempts.ATTEMPT_KEYS` lists the snapshot slices that
belong to one attempt, and A1 had not added `accounting` to it. When the live
flag moved, the new live take was handed the old take's accounting with the
shared prompt, and its own was discarded. The long run's numbers above come from
the `done` events and are unaffected. It was the *stored* accounting of a
retried turn that could be wrong.
### J.3 Genre fixtures (A2-8)
| Fixture | How | Result |
| --- | --- | --- |
| Fantasy Continuity Test | the 50-turn long run, real model | 0 example identifiers proposed. The fixture's own `silver-key` is legitimately its entity |
| Multi-Character Identity Test (office) | the identity diagnostic, real model | 0 proposals naming any prompt example identifier (v1: 1) |
| Persephone science fiction | `test_m11_scifi.py`, 10 tests, scripted | passes; no genre noun in any fixed instruction (`test_v11_protocol_echo.py`) |
**Not run:** science fiction against a real narrator. No harness drives that
fixture with a real model, and adding one was out of A2's scope.
### J.4 What this does not replace
This is not the integrated v1.1 release run (plan §11). That run covers 100
turns, the 16,384 window, and the GPU host with the owner's logging.
---
## K. Compatibility
### K.1 A real v1.0.0 database opens unchanged
**Source database:** `m04-final/campaign.db`, the v1 evidence run's own campaign.
It was written by `96c1bf5`, whose product code is identical to v1.0.0, and holds
`user_version` 94. Its contents:
- 207 actions on 4 branches, with retained history;
- 2 Save Points and 32 state events;
- 33 memories and 12 summaries;
- 3 imported knowledge sources, one campaign narration-length choice, and
settings.
**Method:** `tools/v11_compat_check.py`. The database was copied, and each copy
opened by one tree:
- `v100`: the `v1.0.0` tag, from a scratch worktree;
- `v11-readonly`: the v1.1 tree.
Each tree's app was verified by the path it was imported from. The same read-only
snapshot was taken through the API for both:
- the full export bundle (which carries no timestamp of its own);
- narrative state and state events;
- Save Points, knowledge, memories and derived status;
- the newest actions and settings;
- the schema and `user_version` before and after opening.
| Comparison | Result |
| --- | --- |
| Schema before opening, v1.0.0 against v1.1 | **identical** |
| Schema after opening | **identical**; `user_version` 94 → 94 on both, so no migration ran |
| API snapshot | **identical, field for field** |
### K.2 The same database, exercised on v1.1
`--exercise`, on a third copy, with the endpoint first pointed at a refused
loopback port so nothing left the machine:
| Operation | Result |
| --- | --- |
| Undo | 119 → 117 moments; Redo becomes available |
| Redo | back to 119; Redo no longer available |
| Restore Save Point *On the ridge* | 69 moments, with Redo available into the later story |
| Context dry run | HTTP 200. One imported passage retrieved. The campaign canon and narrative state sections are present. The summary that applies (moments 49-64) is included. Budget 16,384: 14,577 prompt + 28 transport + 500 reply + 820 reserve = 15,925. The window is recorded as unverified, because the server was unreachable by design. |
| Export, then import into the same install | 207 → 207 actions, the same head depth, Save Points 2 → 2, memories 33 → 33, narrative state identical |
**Not exercised here, and why.** Writing a new memory or summary needs a model
and an embedding model. The real-model long run in §J writes both on the v1.1
tree. Memory *retrieval* with embeddings is covered the same way. No
data-migration path was needed, and none was added.
---
## L. Security and offline
### L.1 What the diff adds, measured against `v1.0.0`
`git diff v1.0.0` over `backend/app` and `frontend/src`, searched for any new
network or remote-resource use. The search covered `httpx`, `requests`, `urllib`,
`socket`, `aiohttp`, literal URLs, `fetch(`, `XMLHttpRequest` and
`tiktoken.get_encoding`.
| Check | Result |
| --- | --- |
| New network client, destination or URL | **none** |
| New imports in product code | `math`, `json`, `logging`, and one internal constant (`providers.openai_compatible.CHAT_CONTINUE_HINT`) |
| Remote or cloud token counting | **none**. The count is the vendored `cl100k_base`, as before. The server's own count arrives in the response the turn already receives |
| Calibration service | **none**. Calibration was not built |
| The one request-level change | `stream_options: {"include_usage": true}` in the body of the request already sent to the already-approved endpoint. It adds no destination and no data about the user |
| Endpoint policy (`endpoints.py`, ADR 011) | unchanged; the file is not in the diff |
| Probe policy (`contextwindow._ask`) | unchanged; A1 added pure functions beside it |
| Storyteller bind | unchanged; `docker-compose.yml`, `Dockerfile` and the start scripts are not in the diff |
### L.2 Regression runs
- **The endpoint-policy and local-only tests pass in the backend suite (§M):**
`test_m11_security.py` (H-series, including a public endpoint refused at
request time), `test_egress.py`, `test_local_only_surface.py` (loopback
publication), `test_m11_context_window.py` (the probe held to the policy).
- **Trusted-LAN inference still works.** Every real-model turn in §D and §J
went to the trusted-LAN host over HTTPS, with a private CA and verification on.
- **The offline container passed 23 of 23.** It ran `tools/m11_offline.py` on the
final tree, with `--network none`, a fresh volume, and an image built with
`docker build --no-cache`. The checks are the same 23 as the v1 closeout:
- no route and no DNS;
- first page load, no remote origin, a CSP, and every asset local;
- campaign creation, state extraction, knowledge import, prompt assembly and
retrieval;
- a turn with no model reachable is reported, with no narration accepted,
the player's words kept, state unchanged and the earlier story intact;
- export and import with no secret;
- the media module inert;
- campaigns survive a container restart.
The evidence is `$HOME/v11-evidence/offline/offline-report.json`. The A1
accounting path is exercised offline by the failed-turn check: a turn that
never reaches a model writes no accounting, because it commits nothing.
---
## M. Full test results
Every backend command below is `.venv/bin/python -m pytest tests/ -q` from
`backend/`, with no `AIDND_TEST_*` set. The 17 skips are the tests that need a
real local model, the same 17 as the v1.0.0 closeout. **No new skip was
introduced.** The three real-model skips in the targeted runs are those same
tests.
| Run | Tree | Result |
| --- | --- | --- |
| Backend, baseline | untouched `ac465ed` | **1,421 passed, 17 skipped, 0 failed**, 866 s |
| Backend, A1 checkpoint | A1 only; the not-yet-implemented A2 test file excluded | **1,454 passed, 17 skipped, 0 failed**, 835 s |
| Backend, A2 first full run | A2 with the first vocabulary form | **stopped and discarded.** Two budget failures, §F |
| Backend, A1 + A2 | before the take-accounting fix | **1,505 passed, 17 skipped, 0 failed**, 786 s |
| **Backend, final** | **A1 + A2 + the take-accounting fix: the staged tree** | **1,507 passed, 17 skipped, 0 failed**, 792 s |
| New tests | `test_v11_context_reserve.py` (35) and `test_v11_protocol_echo.py` (51) | all pass |
| Frontend | A1 (A2 changes no frontend file) | **165 of 165**, 14 files |
| Lint | A1 | exit 0, six `only-export-components` warnings, the same six as the v1 closeout |
| Production build | final | clean. 16 files, `index-rWgm40jJ.js` 394.47 kB, `index-BXtb0iME.css` 48.09 kB |
| Offline container, `--network none`, `docker build --no-cache` | A1 + A2, before the take-accounting fix | **23 passed, 0 failed** (§L.2) |
| Offline container, repeated | **final**, after the take-accounting fix | **23 passed, 0 failed**, from a fresh `--no-cache` image (`$HOME/v11-evidence/offline-final`) |
| Identifier scan of every changed and new file | final | no real hostname, address, domain or personal name |
---
## N. Findings
### N.1 Blockers
*None found.*
### N.2 Corrective work, done within the packages
1. **A1: `stream_options` was required, not optional.** The provider's
docstring said `stream_options` was "deprecated and does nothing", which is
true of OpenRouter. Against Ollama 0.33 a stream sends no `usage` without it,
and none of the 514 v1 evidence turns stored a count. Without the fix, A1's
accounting would have read `unknown` on every real turn.
2. **A1: one test's fixture re-calibrated** (§C). The history-floor prefix test
had assumed which turn holds the floor. Behaviour was unchanged.
3. **A2: the first vocabulary form was too expensive** (§F, §G). It grew
protected context by 223 tokens and failed two existing budget tests. It was
replaced by compact JSON (+121); the tests pass unchanged.
4. **A2: module syntax.** Two docstring paragraphs were misplaced. Collection
failed immediately, and they were corrected before any evidence was taken.
5. **Harness: the replay tool compared trailing whitespace** (§I), which
produced one false "rewritten" flag. Fixed.
6. **Harness: the accounting tool imported the bundle's memory setting** (§D.1).
That caused failing post-turn calls on the shared host. Fixed before the 16k
runs.
7. **A1: a retried or reselected take lost its accounting, or showed another
take's.**
- **Where it was found:** after the A2 long run. Three of 54 stored AI turns
had no `accounting`, and they were exactly the two retries and the take
selection.
- **Cause:** `attempts.ATTEMPT_KEYS` lists the snapshot slices that belong to
one attempt, and A1 had not added `accounting` to it. It was therefore
treated as part of the shared prompt. `hand_over_the_prompt` gave the new
live take the old take's accounting, and discarded its own.
- **Fix:** `accounting` added to `ATTEMPT_KEYS`, with a unit test and an
end-to-end retry test (§C).
- **Exports:** none needed a change. A snapshot travels whole.
- **Migrations:** none. v1.0.0 turns have no accounting to move, and
`migrations._ATTEMPT_KEYS` is a frozen historical copy, left as it is.
- **Evidence:** the long run's accounting figures in §J come from the `done`
event of each turn and are unaffected. The full backend suite and the
offline container were re-run after the fix (§M).
### N.3 Non-blocking residual risks
1. **Cold-model first turn: truncation is detected, not prevented** (§D.1).
When `/api/ps` cannot see the model and `/api/show` finds no `num_ctx`, M11
builds to the configured budget. On a server whose real window is smaller,
the first turn is silently cut. It was observed for real: 13,875 sent, 2,050
read. A1 now records it, logs it and shows it. Preventing it needs a policy
decision, which is the owner's and outside A1's scope. The two options are to
build conservatively when the window is unknown, or to warm the model and
probe again before the first turn.
2. **The drift measured is one model's.** +15 tokens against `cl100k_base`, for
`qwen2.5:3b-instruct` in both window sizes. A model whose tokenizer diverges
by more than 5% will show up as `exceeded`, not as silent truncation. The
reserve is a tolerance, not a guarantee.
3. **`exceeded` and `truncation_suspected` depend on the server reporting
usage.** A server that ignores `include_usage` produces `unknown` on every
turn. That is honest, but it detects nothing.
4. **R3's accepted false-positive risk.** A story whose last line exactly
imitates the renderer's `Scene: … (at …)` loses that line (§H).
5. **Protocol still left in stored narration, deliberately:**
- model-invented brackets that carry no application wording;
- model-invented headings;
- facts restated inside a sentence.
**One real v1.1 turn shows the cost** (§J.1, identity depth 16). A reworded
continue hint was the last line of the reply, so nothing above it became
trailing. The application's reminder, a reworded length hint and a scene line
all stayed.
**Narrow follow-up, not done here.** Recognise `CHAT_CONTINUE_HINT` by its
distinctive application phrase "Output only story text", not only by its
opening words. The reminder above it would then come off by the existing
rule. The reworded hint, with none of the application's wording, and the
scene line would still stay.
**Why it was not done.** It would have been an extractor change after the
long-run evidence was taken, and it needs its own replay. This is the owner's
decision.
6. **A2 costs 121 tokens of protected context** (§G.2). At 4,096 that is about
3% of the window, and it comes out of history.
7. **A2's real-model run is on the CPU host at 4,096, not the v1 evidence
configuration.** v1 used 16,384 on the GPU host. On the CPU host a full 16k
prompt took 17 minutes a turn, and a GPU-host long run requires the owner's
power, link and kernel logging to be started first
(`gpu-host-logging-before-long-runs`). The integrated v1.1 release run
(plan §11) is where 16,384 is re-established.
### N.4 Future ideas
- **A cold-model policy for the first turn**, as in N.3 item 1.
- **Summarising the accounting across a campaign**, for example "3 of 120 turns
suspected". Today the inspector shows it one turn at a time.
- **Showing the accounting status as a chip under the narration**, not only in
the context inspector. This is browser work, for WP-C.
---
## O. Final A1 decision
**A1 RESULT: PASS.**
The final decision is taken on the staged tree, including the take-accounting
fix found after the interim checkpoint (§N.2 item 7).
| Criterion | Result | Evidence |
| --- | --- | --- |
| A1-1 Reserve calculation | PASS | `max(256, ceil(5%))`: 256 at 4,096, 410 at 8,192, 820 at 16,384; 7 windows tested (§B.3, §C) |
| A1-2 Budget enforcement | PASS | assembled text, provider hint, reply and reserve are within the effective window in 5 configurations (§C) |
| A1-3 Protected overflow | PASS | fails before the model call; the provider is called 0 times and nothing is written (§C) |
| A1-4 No silent canon dropping | PASS | the canon sentinel is present in every configuration while history gives way (§C) |
| A1-5 Post-response accounting | PASS | real: 51 of 51 long-run turns, 4 turns at 4,096 and 2 at 16,384 carry a server count and a status (§D, §J) |
| A1-6 Unknown accounting | PASS | no usable count gives `unknown`, never `fits` (§C) |
| A1-7 Unexpected discrepancy visible | PASS | a **real** `truncation_suspected` on a cold model: 13,875 sent, 2,050 read. It was stored, logged, sent on the `done` event and shown in the inspector (§D.1) |
| A1-8 Accepted turn preserved | PASS | the turn is kept for `truncation_suspected` and for `exceeded`; after a retry, each take keeps its own accounting (§C) |
| A1-9 Existing behaviour | PASS | 1,507 passed, 0 failed. One fixture re-calibrated, not weakened (§C) |
| A1-10 Real-model verification | PASS | at 4,096, 275-2,297 tokens left beside the reply across 55 turns; at 16,384, 1,994. Drift is +15 on every turn, against v1's 23-42 tokens of headroom (§D, §J) |
**Carried as residual risk, not failure:** the cold-model first turn is detected,
not prevented. Preventing it is a policy decision for the owner (§N.3 item 1).
## P. Final A2 decision
**A2 RESULT: PASS, with one recorded residual.**
| Criterion | Result | Evidence |
| --- | --- | --- |
| A2-1 Genre neutrality | PASS | no fantasy or science-fiction identifier and no genre noun appears in any fixed instruction; a test guards against reintroduction (§G.3) |
| A2-2 Protocol clarity | PASS | the vocabulary is the wire-format object, not call notation. +121 tokens, the smallest wire-format form measured (§G) |
| A2-3 Observed leaks | PASS | all four v1 shapes are removed (§H). What is deliberately left is documented, including one real v1.1 occurrence (§J.1, §N.3 item 5) |
| A2-4 Historical corpus safety | PASS | 518 real replies replayed, 9 changed, 0 flagged, **no story prose removed**. A further 64 replies from the v1.1 runs gave 1 changed, 0 flagged (§I, §J.2) |
| A2-5 Adversarial prose | PASS | the owner's four sentences and nine more survive untouched (§C of the A2 tests, `test_v11_protocol_echo.py`) |
| A2-6 State extraction | PASS | a valid block still parses and applies when protocol litter surrounds it. `SPECS`, `render.py` and `validate.py` are identical to v1.0.0 (§F, §K) |
| A2-7 Invalid proposals | PASS | the validator is byte-identical and the existing refusal tests pass; the long run recorded 0 refusals and 6 unparseable proposals (§J.2) |
| A2-8 Genre fixtures | PASS | fantasy by real model (51 turns); office identity by real model (10 turns, 0 example identifiers against v1's 1); science fiction scripted only, which is stated (§J.3) |
| A2-9 Real-model sample | PASS | 51 accepted turns: 0 of 54 stored turns carry protocol, 6 parser failures, 0 refused, 4 story state events, no prose removed, and A1 accounting `fits` 51 of 51 (§J.2) |
**The residual.** In the identity run, 1 of 10 stored turns kept a trailing tail
of echoed instructions: a reworded continue hint, the reminder, a reworded length
hint and a scene line. The conservative design cannot see past the reworded last
line. §N.3 item 5 records a narrow follow-up, which is the owner's decision.
## Q. Recommendation
- **A1 is ready for acceptance.** It meets every criterion on real evidence,
including a real silent truncation it caught.
- **A2 is ready for acceptance, with the residual in §P noted.** The owner may
want the narrow continue-hint follow-up (§N.3 item 5) first. It would be a small
extractor change with its own replay. It does not block acceptance of what A2
set out to do.
- **Owner decisions this report raises:**
1. Accept or correct A1.
2. Accept A2, or ask for the continue-hint follow-up first.
3. Whether the cold-model first turn (§N.3 item 1) needs a prevention policy,
and where that policy belongs: an A1 follow-up, or a later package.
- **WP-B may begin once A1 and A2 are accepted.** Nothing in either package blocks
B's diagnosis. A1's accounting and A2's cleaner prose are both in place for B's
long runs. **WP-B has not been started.**
**Git state:** everything is staged and uncommitted for the owner's signed commit.
Nothing was committed, pushed or tagged. A suggested commit message was prepared
outside the repository.
---
## R. Corrective-work addendum (owner review, 2026-09-14)
The owner accepted A1 and A2 in principle, and asked for two corrective items to
be closed before the commit. §O-§Q above record the decisions as they stood
before this addendum. **§R.6 supersedes them.**
### R.1 WP-A1: preventing the cold-model truncation
**The case, as A1 caught it (§D.1, turn 1):**
- `/api/ps` knew nothing, because the model was not resident.
- `/api/show` found no `num_ctx`, so the window was unverified.
- The prompt was built to the configured 16,384.
- Ollama loaded the model at its 4,096 default, read **2,050 of 13,875** tokens,
and answered HTTP 200.
**Mechanism, measured before it was built.** This was measured on Ollama 0.33
on the trusted-LAN host, with the model first unloaded (`keep_alive: 0`) and
`/api/ps` empty.
| Step | Observed |
| --- | --- |
| `POST /api/generate {"model": "qwen2.5:3b-instruct"}`, with no prompt | HTTP 200 in 3.7 s, `"response": ""`, `"done_reason": "load"`. **Nothing was generated.** |
| `/api/ps` afterwards | the model is listed with `context_length` 4,096 |
| An OpenAI-compatible streamed request afterwards | no reload: `/api/ps` still reads 4,096, and `prompt_tokens` was the estimate plus 13 |
**What was built.**
- `contextwindow.Window` gains `reachable`: the server answered a discovery
request, whatever it said.
- `contextwindow.warm(endpoint, model, timeout)` sends one `POST {native base}/api/generate`
with the body `{"model": model}`. There is no prompt, no `options` and no
`keep_alive`. It goes through `endpoints.rejection_reason` and
`tlstrust.ssl_context()`, and it never raises.
- `contextwindow.ensure_window(endpoint, model, declared, warm_timeout)` does
the following:
1. It probes. A verified window is returned untouched.
2. If the window is unverified, the server was reachable, and a model is
configured, it warms once.
3. If the load succeeds, it probes again with `use_cache=False`.
4. It returns the window and a `preflight` record (`attempted`, `loaded`,
`verified_before`, `verified_after`, `detail`).
- `turns._generate_turn` calls `ensure_window` in place of `probe`, with the
settings' model timeout as the bound. It stores
`snapshot["window"]["preflight"]`. The context dry run still only probes:
loading a model from a read-only panel would be a side effect.
**What is unchanged.**
- A window still unverified leaves the configured budget standing.
- There is no guessed window, no hard-coded 4,096, and no retry loop.
- The accounting still classifies the reply.
**The constraints the owner set, each checked:**
- no calibration and no user setting;
- no endpoint-policy change;
- only the configured endpoint is contacted;
- no dummy narration is generated or accepted;
- no accepted turn is discarded.
**Tests:** `tests/test_v11_cold_window.py`, 15 tests.
| Required | Test |
| --- | --- |
| cold model, then preflight, then a verified smaller window, then the real turn built to it | `test_a_cold_model_is_loaded_once_and_its_window_verified` (request order `/api/ps`, `/api/show`, `/api/generate`, `/api/ps`; the load body exactly `{"model": …}`). `test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports`, end to end: budget 4,096, verified, the canon sentinel present, the assembled text plus transport, reply and 256-token reserve within 4,096 |
| negative control | `test_without_the_load_the_same_cold_turn_would_have_been_built_too_large`: probing alone gives an unverified window, budget 16,384, and a prompt more than twice 4,096 |
| already loaded: no warm | `test_an_already_loaded_model_is_not_warmed` |
| load succeeds, verification still unavailable | `test_a_model_that_loads_but_still_cannot_be_read_stays_unverified`: one load, unverified, the configured budget kept |
| load fails | `test_a_failed_load_is_recorded_and_leaves_the_window_unverified` (HTTP 404 and 500; bounded, with no second probe). `test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe`: the error is reported, with no AI action, state event or proposal. `test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer`: accounting `unknown` |
| the load writes nothing | `test_the_load_itself_writes_nothing`: actions, state events, proposals, memories and summaries are counted before and after, and are identical |
| no new destination | `test_the_load_request_obeys_the_endpoint_policy` (public addresses refused with no transport installed). `test_the_load_request_goes_only_to_the_configured_host`. `test_an_unreachable_server_is_not_asked_to_load_anything` |
| also | `test_a_declared_window_does_not_stop_the_server_being_asked`; `test_the_context_dry_run_never_loads_a_model` |
**Real cold-model test:** §R.3.
### R.2 WP-A2: the continue-hint instruction tail
**The exact output.** Identity depth 16 (action 17 in the run's database). The
stored text equals the v1.1 extractor's output of the raw reply, and the raw
reply has no state block. It ends:
```text
Scene: Bill, Alice and Roger at the table; John not yet arrived.
[Hard limit: this is now 180 words.]
[Reminder: end your reply with a `state` block listing the events your narration made true, with absolute values.]
[You don't need to continue; your turn must now be about John entering the room. Continue the story here, directly. Output only story text.]
```
**Why the application-owned material above the last bracket was not removed.**
`_clean` cuts the end of a reply one bracket at a time, and only ever examines
the **last** bracket. That last bracket is the provider's `CHAT_CONTINUE_HINT`,
reworded at the front. `_is_echoed_instruction` recognised that hint only by its
opening words ("continue the story directly"), so it returned false, and the
loop stopped. The reminder above it would have been removed by the existing
`Reminder:` rule, but it was never at the end. The reworded length hint carries
none of the application's wording, and the scene line was never last.
**The rule, R5 (`RULE_INSTRUCTION_TAIL`).** Two narrow changes:
1. **Recognise the continue hint by its own sentence.** A trailing bracket
containing "Output only story text" (`CONTINUE_HINT_PHRASE`) is an echoed
instruction. A test pins the phrase to `CHAT_CONTINUE_HINT`, so the two
cannot drift.
2. **Remove a hint-opened bracket only directly above an echo.** A trailing
bracket opening with `Hard limit:` is removed only when an echoed instruction
bracket has already been cut from the end of the same reply.
The rest follows from existing rules. The reminder is removed by the `Reminder:`
rule once it is last. The scene line is removed by R3 once it is last and
protocol has been cut.
**False-positive defence.** Each of these is a test that must pass unchanged:
| Case | Why it stays |
| --- | --- |
| A final `[Hard limit: forty days, no extensions]` | no echo was cut below it |
| The same bracket above a state block | a block is not an echoed instruction, so a block below does not license it |
| `[The sign on the door reads: Closed]` directly above an echoed continue hint | it does not open the way the hint opens; only the echo goes |
| "output only story text" inside a sentence mid-story | only a trailing bracket is examined |
| A final `[To be continued]` | no application phrase |
The owner's four original adversarial sentences and all earlier preservation
cases also still pass.
**Accepted risk.** A story-world bracket that opens exactly `[Hard limit:` *and*
sits directly above a parroted instruction would be removed with it.
**Positive regression:** `test_the_depth_sixteen_instruction_tail_leaves_entirely`.
The two story paragraphs are shortened; the four trailing lines are verbatim. It
leaves exactly the story.
### R.3 Evidence
| Check | Result |
| --- | --- |
| New and related unit tests (the A1 cold-window, A1 reserve, A2 echo, extractor, window, take and caching tests) | **308 passed, 3 skipped** (the real-model skips) |
| **Replay: the full v1 corpus** | **518 replayed, 509 unchanged, 9 changed, 0 flagged.** The changes are exactly the nine in §I, with the same rules: R5 changes no v1 turn |
| **Replay: the v1.1 runs' raw replies** (the 51-turn run's database and the identity run's database) | **64 replayed, 61 unchanged, 3 changed, 0 flagged:** |
| | long run, action 77: R4, a trailing ```` ```json ```` (0.4%), as in §J.2 |
| | **identity depth 16:** R3 scene line, plus R5 × 3 (the reworded length hint, the reminder, the reworded continue hint). 18.0% of the reply, all of it the tail above |
| | identity depth 18: R4, a trailing ```` ```json ```` (0.8%) |
| **Genuine story prose removed** | **0** |
| **Unreviewed replay flags** | **0** |
| **Real cold-model A1 test** (`v11_window_accounting --unload-first`, same host, model and 207-action bundle as §D.1) | **Prevented.** Model unloaded (`/api/ps` empty). Turn 1: preflight attempted, model loaded, window verified 4,096 (`loaded`), budget 4,096, estimate 3,082, server read **3,097** (+15), margin 499, **`fits`**. Turn 2: no preflight needed (already verified), estimate 3,086, server 3,101, margin 495, `fits`. Before the correction the same turn sent 13,875 against an unverified window and the server read 2,050 |
| **Identity diagnostic, re-run** (office fixture, 10 beats, the same CPU/HTTPS host at 4,096; `$HOME/v11-evidence/a2c-identity`) | **identity signals: 0** (verdict: no objective identity defect). **Prompt example identifiers in proposals: 0** (no `silver-key`, `aldric`, `mara`, `old-abbey`, nor the placeholders). **Stored protocol/instruction shapes: 0/10** (calls, hint and reminder brackets, continue-hint wording, a final scene line, state headings, `"events"`, fences). Proposals: 9 accepted, 1 partially accepted, 1 unparseable. **Read with the replay:** this run's 10 raw replies replay *unchanged* through v1.0.0 and v1.1, so this narrator produced no echo this time. 0/10 here is the required closeout result, not a live exercise of R5. R5's proof is the verbatim depth-16 regression fixture and the replay of the original depth-16 turn above |
| **Backend full suite** | **1,534 passed, 17 skipped, 0 failed**, 1,007 s (the 1,507 before this pass, plus 15 cold-window and 12 corrective A2 tests; the same 17 real-model skips) |
| Frontend suite | **165 of 165** (no frontend file changed in this pass) |
| Lint | **exit 0**; the same pre-existing `only-export-components` warnings |
| Production build | **clean**, 16 files; `index-rWgm40jJ.js` 394.47 kB and `index-BXtb0iME.css` 48.09 kB, the same hashes as before, so the SPA is byte-identical |
| **Offline container** (`tools/m11_offline.py`, `docker build --no-cache`, `--network none`, a fresh volume) | **23 passed, 0 failed** on the corrective tree (`$HOME/v11-evidence/offline-corrective`). The cold-model load is inert here: the failed-turn check has no reachable server, so no load is attempted (`reachable` false), and the turn is reported and commits nothing, as before |
| **Browser and context-inspector regression** (`tools/m11_browser.py`, Firefox 155.0.1 headless, the built SPA served by FastAPI, the real narrator over trusted-LAN HTTPS; `$HOME/v11-evidence/browser-corrective`) | **38 passed, 0 failed, 0 skipped.** This includes F05 (the context inspector shows the assembled prompt), two real turns through the UI, Undo and Redo, hostile narration, G01 import, the dialog and accessibility checks, and CSP. The two browser-played turns' stored snapshots carry the new provenance: the window was verified at 4,096 (`loaded`), the preflight was not attempted because the model was already resident, and accounting was `fits` (server 876 against estimate 861, margin 2,820; server 1,024 against estimate 1,009, margin 2,672). The inspector's accounting and safety-margin display is covered by the component suite (§C), not by this harness |
| **v1.0.0 database compatibility** (a fresh `v1.0.0` worktree against this tree, same database as §K) | **schema before: identical; schema after: identical; API snapshot: identical; `user_version` 94 → 94.** Exercise: Undo 119 → 117, Redo 117 → 119, restore *On the ridge* → 69 with Redo available, dry run 200 with knowledge retrieved, export and import 207 → 207 actions with the same head and identical state |
### R.4 Findings from the corrective pass
**Blockers:** none.
**Corrective work within the pass:**
1. **Harness slip: two background watchers never ran their jobs.** The watchers
were meant to start the offline container after the suite, and the browser
regression after the identity run. Each waited with `pgrep -f` on a command
name that appears in its own command line, so each matched itself and never
proceeded. Nothing ran and nothing was corrupted. Both watchers were stopped,
and both runs were started directly (§R.3). This is the same class of slip
as the earlier `pkill` self-match (§N.2). It sits in session scripting, not
in the repository.
**Non-blocking residual risks:**
1. **R5 accepted false-positive risk.** A story-world bracket that opens exactly
`[Hard limit:` and sits directly above a parroted instruction would be
removed with it (§R.2).
2. **Cold-model load: what it cannot fix.**
- It does nothing for a server that is unreachable, or that does not speak
Ollama's native API. Both keep the unverified path, now recorded in
`window.preflight`.
- It does not change the window. A model loaded at a smaller default is still
that window; the turn is built to it, not to the setting.
- It costs one load request on a cold turn: 3.7 s measured on the CPU host
for a 3B model, bounded by the model timeout.
3. **The identity closeout's 0/10 does not exercise R5 live** (§R.3). The shape
is proven by the verbatim fixture and the replay, not by this run.
**Future ideas:**
- Warming the model when the reader opens a campaign, so the first turn does not
pay for the load. This was deliberately not done, because the dry-run panel
must have no side effects.
### R.5 Closeout against the owner's required results
| Required | Result |
| --- | --- |
| A1: one bounded load of a cold model, no story content, then the window probed again | built and tested (§R.1); real server: HTTP 200 in 3.7 s, `"response": ""`, `"done_reason": "load"` |
| A1: turn built to the verified window when the load makes it known | **real cold test: verified 4,096, 3,082 sent, 3,097 read, `fits`.** Before the correction: unverified, 13,875 sent, 2,050 read |
| A1: window still unknown → no guess, no hard-coded 4,096, the unverified path kept, accounting kept | tested (§R.1) |
| A1: no calibration, no user setting, no endpoint-policy change, only the configured endpoint, no dummy narration, no discarded turn | met (§R.1). The policy and same-host tests pass. The load writes nothing, tested by row counts |
| A2: minimised positive regression, and why the tail survived | `test_the_depth_sixteen_instruction_tail_leaves_entirely`; root cause in §R.2 |
| A2: narrowest rule; adversarial controls; all extractor tests | R5; 12 new tests; the extractor, A2 and state tests pass |
| A2: replay of all available real turns | **518 replayed, 509 unchanged, 9 changed, 0 flagged.** Classifications exactly as §I.1. Plus the 64 v1.1 replies: 3 changed, 0 flagged |
| **genuine story prose removed** | **0** |
| **unreviewed replay flags** | **0** |
| **identity signals** | **0** |
| **prompt example identifiers in proposals** | **0** |
| **stored protocol/instruction shapes** | **0/10** |
| Backend full suite | **1,534 passed, 17 skipped, 0 failed** |
| Frontend suite / lint / production build | **165/165**; exit 0; clean and byte-identical |
| Offline container | **23 passed, 0 failed** |
| Browser and context-inspector regression | **38 passed, 0 failed, 0 skipped** |
| v1.0.0 database compatibility | **identical** schema before and after, identical API snapshot, `user_version` 94 → 94. Exercise passed |
### R.6 Final decisions
```text
WP-A1: PASS
WP-A2: PASS
```
**WP-A1: PASS.**
- **What is met:** every criterion in §O, plus the owner's corrective
requirement.
- **The case now prevented:** the real cold-model truncation that A1 had only
detected. The window becomes discoverable once the model loads, and the turn
is built to it.
- **What still depends on detection:** a server that cannot be asked, or that
will not load the model on request. That falls back to the unverified path,
which is unchanged, recorded, and caught by the accounting.
**WP-A2: PASS.**
- **What is met:** every criterion in §P, plus the owner's corrective
requirement.
- **The residual §P recorded is closed.** The depth-16 instruction tail is
removed by a narrow rule anchored to application-owned text.
- **Safety evidence:** 0 genuine prose removed and 0 unreviewed flags across
582 real replies (518 v1 and 64 v1.1).
- **Identity closeout:** 0 signals, 0 example identifiers, and 0/10 stored
shapes.
**Next step.** WP-B.1 may begin **after the owner signs the A1/A2 commit.** WP-B
has not been started.
**Git state.** Everything is staged and uncommitted: product code, tests, tools,
planning documents and this report. Nothing was committed, pushed or tagged.