Files
interactive-story/planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00

74 KiB
Raw Permalink Blame History

v1.1 WP-A1 and WP-A2 — Implementation and Verification Report

Work packages: WP-A1 (context-window safety reserve) and WP-A2 (protocol-echo cleanup and genre-neutral state prompting), authorised together and implemented in that order. They are reported separately throughout.

Status: COMPLETE, including the owner-requested corrective work (§R). Staged, uncommitted, for the owner's signed commit (2026-09-14). Final decisions: §R.6. WP-A1: PASS. WP-A2: PASS.


A. Repository baseline

Branch v1.1-development
Starting commit ac465ed — Planning v4.1: record the v1.0.0 release, and plan v1.1, signed by the owner (%G? = G)
Its parent 432f041, the signed v1.0.0 release commit; tag v1.0.0; main and origin/main
Working tree at start clean; nothing staged, nothing unstaged, nothing untracked
Planning baseline committed before any product change, so it is not in this diff

Baseline backend suite, before any change (the untouched ac465ed tree, python -m pytest tests/ -q, no AIDND_TEST_* set): 1,421 passed, 17 skipped, 0 failed, 865 s. This matches the v1 closeout's count exactly.


B. WP-A1 implementation

B.1 What the code did before, answered from the code

Question Answer on ac465ed
1. What defines the effective context window? contextwindow.effective_budget(configured, window), which is min(context_token_budget, window.tokens) when a window is enforceable (verified by /api/ps or /api/show, or declared through context_window_override), and the configured budget otherwise.
2. Where is the fixed 64-token margin applied? context/builder.py: OUTPUT_SAFETY_MARGIN = 64, added to max_output_tokens as output_reserve. That is subtracted with the protected sections before any knowledge or history is priced.
3. How is max_output_tokens reserved? Inside that same output_reserve.
4. Which components are protected? Every system section (narrator prompt, state rule, campaign canon, always-include Canon and the knowledge rule, AI instructions, persona, plot essentials), plus the summary, retrieved memories, narrative state, author's note, front memory, length hint, refusal note and state reminder.
5. Where is history trimmed? build_context: history.window_covering, then history_floor and trim_block (§15.3), then newest-first until history_budget is spent. Knowledge is chosen before history, out of its own share.
6. What counts tokens before the request? cl100k_base from the vendored table (context/encoding.py), through builder.count_tokens.
7. Which response fields give the real prompt-token count? usage.prompt_tokens in the OpenAI-compatible response. Measured on Ollama 0.33, a stream sends no usage unless stream_options.include_usage is set. Chat and completion modes both report it once asked, and non-streaming requests always do.
8. Where is that metadata persisted? snapshot["usage"] = provider.last_usage in turns.py, inside the turn's compressed context_snapshot. Because streams were never asked, none of the 514 AI turns in the v1 evidence databases has a stored usage.
9. Can the status live in provenance without a migration? Yes. context_snapshot is a compressed JSON column, and the per-action context route returns it whole.
10. What happens if protected context alone is too large? ContextOverflow is raised before the model call. The turn route yields it as a turn error, and the dry-run route returns 422. Nothing is written.

B.2 What the 64 tokens were, and what replaced them

M6's comment gave the margin two jobs:

  • absorb "the separators between sections" that are added after the arithmetic;
  • absorb "the difference between our tokenizer's count and the serving model's".

The first job is not drift. It is application text nobody priced: SEPARATOR joins between sections, and CHAT_CONTINUE_HINT, which the provider appends to every chat request. v1.1 prices it exactly as tokens.transport:

  • one separator per system section and per story section slot (STORY_SECTION_SLOTS = 11), which over-counts by at most two;
  • plus the hint.

The second job is what the reserve now does. The reply allocation is exactly max_output_tokens. Nothing is double-counted: the 64 tokens are gone, and each of their two jobs has one owner.

B.3 The reserve

contextwindow.safety_reserve(effective_budget) = max(256, ceil(effective_budget × 5 / 100)). The rounding is up, to a whole token, in integer arithmetic.

Effective window Reserve
4,096 256, because 204.8 is below the floor
5,120 256, exactly 5%
8,192 410, from 409.6
16,384 820, from 819.2
32,768 1,639, from 1,638.4

It follows the effective budget. A 16,384 setting against a 4,096 server reserves 256. It is not a setting, and it is not calibrated per model.

Protected context is sections + transport + max_output_tokens + reserve. ContextOverflow now names each part, and keeps the M11 advice sentence when the server's window is what binds.

B.4 Accounting after the reply

contextwindow.classify_usage(usage, estimate, budget, max_output_tokens, window_verified). estimate is tokens.estimate: the assembled text plus what the provider adds (the chat hint, or the completion separator). The checks run in this order:

Status Condition What it means
unknown no positive integer prompt_tokens Nothing can be said. Never reported as fits.
truncation_suspected server + reserve < estimate The server read fewer tokens than were sent, by more than the drift tolerance. This is the measured shape of an over-window prompt: 6,316 sent, 2,050 read.
exceeded server + max_output_tokens > budget The drift used up the whole reserve, so the reply may be cut short.
fits otherwise The server read the whole prompt.

The record also carries server_prompt_tokens, difference (server minus estimate), observed_margin (budget - reply - server), safety_reserve, window_verified and a sentence of detail.

Why exceeded is measured against the hard edge and not the reserve line. On prompts that fit, the server counted 13 tokens more than the application, which is chat-template overhead. A full prompt built exactly to the reserve line would then read as "exceeded" on every turn. The reserve is the tolerance, so using part of it is fits with a smaller observed_margin. Using all of it is exceeded.

Surfacing.

  • The record is stored as snapshot["accounting"], so it is available for every past turn.
  • It is returned on the turn's done SSE event.
  • exceeded and truncation_suspected are logged at WARNING.
  • The context inspector shows the safety margin on every turn, and the accounting for a turn that was sent. The two bad states are a role="alert" notice that says the turn is kept.

The turn is kept. Accounting happens after the reply has streamed, and it changes nothing about whether the turn commits.

B.5 Files

File Change
backend/app/contextwindow.py SAFETY_RESERVE_FLOOR, SAFETY_RESERVE_PERCENT, safety_reserve; the FITS / EXCEEDED / TRUNCATION_SUSPECTED statuses; classify_usage
backend/app/context/builder.py OUTPUT_SAFETY_MARGIN removed; transport, the reserve, and an exact reply allocation; ContextOverflow names each part; the report gains transport, safety_reserve and estimate
backend/app/providers/openai_compatible.py STREAM_OPTIONS = {"include_usage": True} on all four streaming bodies
backend/app/routers/adventures/turns.py snapshot["accounting"], a WARNING log, and accounting on the done event
frontend/src/pages/Play/panels/ContextPanel.jsx, frontend/src/styles/context.css the safety-margin line and the accounting notice
backend/tools/m11_long_run.py the timeline records the turn's accounting from the done event
backend/tools/v11_window_accounting.py new: the real-model accounting harness (§D)
planning/TECHNICAL-DESIGN.md §15.2, DEVELOPMENT.md, README.md as-implemented notes. The README's stale Screenshots paragraph and test count are corrected, as the v1.1 plan asked

No schema change, no migration, and no bundle-format change.


C. WP-A1 tests

tests/test_v11_context_reserve.py, new:

Criterion Tests
A1-1 reserve calculation test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up (1,024; 4,096; 5,120; 5,121; 8,192; 16,384; 32,768), and …far_larger_than_the_v1_margin…
A1-2 budget enforcement test_the_prompt_leaves_the_reply_and_the_reserve_free, on the assembled text plus the chat hint, for verified 4,096, 8,192 and 16,384, a declared 6,000, and an unverified configured 12,000, each against a 120-turn story. Also test_the_report_prices_the_text_the_provider_adds and test_the_reserve_follows_the_effective_window_not_the_setting
A1-3 protected overflow test_protected_context_that_only_fits_without_the_reserve_fails_explicitly: a canon grown until v1's 64-token margin would still have built the prompt and the 256-token reserve does not. Also test_an_overflowing_turn_never_reaches_the_model: the scripted provider is called 0 times and no AI action is written
A1-4 no silent canon dropping the canon sentinel is asserted present in every A1-2 configuration while the history gives way. M11's test_the_canon_at_the_front_survives_a_window_far_too_small and its negative control still pass
A1-5 accounting test_a_prompt_the_server_read_in_full_fits; test_a_turn_records_what_the_server_read, end to end, including the done event
A1-6 unknown test_no_usable_count_is_unknown_never_fits: None, {}, no prompt_tokens, 0, a string and a negative value. Also test_a_turn_with_no_reported_usage_is_unknown
A1-7 discrepancy recorded and visible test_a_suspected_truncation_keeps_the_turn_and_says_so (snapshot, done event, WARNING log, and the per-action context route); test_a_server_that_read_far_less…; test_a_small_undercount_is_tokenizer_drift_not_truncation (reserve minus 1 is fits, reserve plus 1 is suspected); test_a_server_that_counts_more…_is_exceeded
A1-8 turn preserved test_a_suspected_truncation_keeps_the_turn_and_says_so, test_an_exceeded_turn_is_also_kept
request test_the_stream_asks_the_server_to_report_its_usage, in chat and completion modes
each take's own accounting (added after the A2 long run; §N.2 item 7) test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves, on keep_own_slices and hand_over_the_prompt. test_a_retry_leaves_each_take_with_its_own_accounting, end to end through POST /retry: the superseded take keeps its fits and no prompt; the live take keeps the prompt and its own truncation_suspected

The file has 35 tests: 33 at the A1 checkpoint, plus the two added above.

Frontend, panels.test.jsx, four new tests: the safety margin is shown; nothing is shown for a turn not yet sent; truncation_suspected is an alert naming both counts; unknown makes no claim that the prompt fitted.

One existing test changed. It is a fixture calibration, not a weakened assertion.

  • The failure. test_history_block_trim.py::test_the_story_prompt_keeps_its_prefix_across_a_new_turn asserted that the next turn after its fixture holds the history floor. Its budget is 2,048 tokens, where the reserve is 256, or 12.5% of the window. The history budget fell from 1,186 to 969: the reserve's extra 192 over the old margin, plus 25 tokens of transport. The block is 2 on both trees. The traces:

    turn     v1.0.0 floor     v1.1 floor
    <60           54               54
    <61           54               56    <- v1.1 steps here, v1.0.0 one turn later
    <62           56               56    prefix kept 0.906
    <63           56               58
    <64           58               58    prefix kept 0.906
    
  • The judgement. Same cadence and same step size, offset by one turn. §15.3's property holds, and only the fixture's choice of turn moved.

  • The change. The test now walks forward until a turn holds. It requires every move on the way to be exactly one block, and asserts the prefix on the held pair. A window that slides by one action every turn fails it exactly as before, and its negative-control companion is unchanged.


D. WP-A1 real-model evidence

Harness: tools/v11_window_accounting.py, on the A1 tree, with no source edit during the run.

Inference: the trusted-LAN CPU reference host. It runs Ollama 0.33.0 over HTTPS, with a private CA in this machine's OS trust store and verification on.

Campaign: the v1 evidence run's own bundle (m04-final/bundle.json, 207 actions), imported so that every turn is assembled against a full window. Memory bank and auto-summarise are off, because they would only add post-turn calls on the same host.

Settings: configured budget 16,384, max_output_tokens 500.

Evidence: $HOME/v11-evidence/a1-accounting/.

D.1 The 4,096-window reference configuration (qwen2.5:3b-instruct, no num_ctx)

Turn Window Effective budget App estimate Server prompt_tokens Difference Reserve Observed margin Status
1 not verified (model not resident) 16,384 (configured) 13,875 2,050 −11,825 820 — truncation_suspected
2 4,096, verified (loaded) 4,096 3,294 3,309 +15 256 287 fits
3 4,096, verified 4,096 3,253 3,268 +15 256 328 fits
4 4,096, verified 4,096 2,831 2,846 +15 256 750 fits

"Observed margin" is window − reply allocation − the server's count. The v1 evidence's equivalent at the largest prompts was 23-42 tokens.

Turn 1 is the most important row in this section. The model was not resident, so /api/ps could not report its window. M11's rule for that case leaves the configured 16,384 standing and records the window as unverified. The application sent 13,875 tokens. The server loaded the model at its 4,096 default, kept 2,050 of them, and answered HTTP 200. That is the silent truncation M8 found and M11 set out to prevent, still reachable on a cold model's first turn. v1.0.0 would have stored this turn with no record of it. A1 recorded truncation_suspected, logged a warning ("The server read 2,050 prompt tokens of the 13,875 sent…") and kept the turn. §N records the cold-model gap itself as a finding.

Turns 2-4: the window was verified and the budget capped to it. The server counted exactly 15 tokens more than the application on every turn: chat-template overhead the application cannot see. That used 15 of the 256-token reserve, and each turn kept 287 tokens or more beside the reply.

One harness note. The bundle carried the evidence campaign's memory bank setting, so turn 1's post-turn memory pass ran and failed as the in-process client closed (derived memory work failed). It is a failure in the harness, outside the turn and outside what is measured. The tool was then changed to switch memory and auto-summarise off before importing. The 16,384 run below used the changed tool.

D.2 The 16,384-window configuration (qwen2.5:3b-instruct-16k, num_ctx 16,384 baked in)

Turn Window Effective budget App estimate Server prompt_tokens Difference Reserve Observed margin Status Seconds
1 16,384, verified (parameters) 16,384 13,875 13,890 +15 820 1,994 fits 1,028

The first run was stopped after turn 1 at the owner's request, because the shared inference host was needed by another project. One turn took 17 minutes on the CPU host. A second run of three turns was started when the host was released (§D.3).

The window was verified from the model's own num_ctx, so even a cold model had a ceiling. The prompt held 60 of 120 history actions. The server again counted 15 tokens more than the application. That left 1,994 tokens beside the reply, against 23-42 in the v1 evidence at the same window and the same model.

D.3 16,384 window, second run

Started when the host was released, with the A1 code loaded at process start and before any A2 edit.

Turn Window Effective budget App estimate Server prompt_tokens Difference Reserve Observed margin Status Seconds
1 16,384, verified (parameters) 16,384 13,875 13,890 +15 820 1,994 fits 1,058

The run was stopped deliberately after turn 1 (10:03 EDT), to free the shared host for WP-A2's real-model runs. Each 16k turn took about 17.6 minutes on the CPU host, and the two remaining turns would have delayed A2's 50-turn run by more than half an hour.

Read this row as a reproduction, not a second sample. Both 16k runs imported the same campaign, so turn 1 assembled the same prompt. The identical counts show that the measurement is repeatable. They add no new prompt shape. The prompts that grow and change across many turns are measured at 16,384 by the integrated v1.1 release run (plan §11). In this report, the prompts that change turn by turn are the 4,096 rows in §D.1 and the A2 long run in §J.

D.4 What the evidence shows

Criterion Result
A1's tolerance is materially larger than v1's margin At 16,384: 1,994 tokens left beside the reply, against 23-42. At 4,096: 287 or more. The reserve is 820 and 256.
Real drift between the counters +15 tokens on every turn measured, at both windows. That is 6% of the smaller reserve and 2% of the larger.
Disagreement is observable Every sent turn carried a server count and a status. A real cold-model truncation was caught, recorded and logged, and the turn was kept.

E. WP-A1 interim decision

A1 RESULT: PASS.

Recorded at the checkpoint before any WP-A2 product code was written.

  • A1-1 to A1-9 pass on the A1 tree:
    • backend suite 1,454 passed, 17 skipped, 0 failed (the 1,421 baseline plus 33 new);
    • frontend 165 of 165;
    • lint clean apart from the six warnings the v1 closeout already had.
  • A1-10 passes on real-model evidence. At 4,096 there are four turns, including a real truncation_suspected. At 16,384 there is one verified turn, with the second run pending (§D.3).
  • One existing test was re-calibrated, not weakened (§C). The failure was the fixture's choice of turn. §15.3's property was unaffected, and the v1.0.0 and v1.1 traces show it.
  • A1-only checkpoint saved outside the repository, so A1 can be reviewed on its own:
    • a1-tracked.patch, sha256 3b5d679f…;
    • a1-untracked.tgz, sha256 78275d0f….

Later corrective work on A1, found after this checkpoint. The A2 long run exposed a defect in A1: a retried or reselected take lost its accounting, or was shown another take's. It was fixed with two tests (§N.2 item 7). The fix is one line in attempts.py, plus the two tests. The A1-only checkpoint patch above predates it. The final staged tree includes it. This interim PASS stands on A1-1 to A1-10 as written, and §O records the decision after the fix.

Carried to §N, not corrective work for A1: the cold-model first turn. When the window cannot be verified, M11 builds to the configured budget. A1 now makes the resulting truncation visible, but does not prevent it. Preventing it means a policy change: build conservatively, or refuse to send, when the window is unknown. That is outside A1's authorised scope ("do not implement model-specific dynamic calibration"; "preserve the accepted turn") and is the owner's decision.


F. WP-A2 implementation

Begun only after §E was recorded. The A1-only checkpoint was saved first, so A1 can still be reviewed on its own.

File Change
backend/app/narrative/events.py vocabulary_for_prompt shows each event as the JSON object to send, not as name(field, …). _PLACEHOLDER holds the neutral placeholders. SPECS is unchanged, verified against v1.0.0 by diff.
backend/app/narrative/extract.py EMIT_RULE's example now uses character-1, item-1 and location-1. LENGTH_HINT_OPENING and LENGTH_HINT_TAIL are shared with the builder. R1-R4 and RULE_* added (§H). _clean(…, after_block=). explain_removed_line added for the replay.
backend/app/context/builder.py length_hint builds its opening and tail from the shared constants. The hint text is byte-identical to v1.0.0.
backend/tools/m11_long_run.py The leak count also detects call lines and echoed hints.
backend/tools/v11_replay_extractor.py New: the historical replay (§I).
backend/tools/v11_compat_check.py New: the v1.0.0-database compatibility check (§K).
backend/tests/test_v11_protocol_echo.py New: prompt, observed shapes, adversarial prose, rule attribution.
planning/TECHNICAL-DESIGN.md §15.4, DECISIONS/013 As-implemented notes.

Unchanged: no event type, field, reference rule or validation. The fence protocol, proposal recording, history replay and schema are also untouched.

Two corrections made during A2, before its evidence was taken.

  1. Misplaced docstring paragraphs. The new paragraphs in extract._clean and events.vocabulary_for_prompt were first placed after each docstring's closing quotes, so the modules did not import. Collection failed at once and nothing ran; the paragraphs were moved into the docstrings.
  2. The first vocabulary form cost too much. "<key>"-style placeholders with spaced separators took the vocabulary from 258 to 456 tokens and EMIT_RULE from 467 to 690. That pushed protected context in an existing budget test (2,048 tokens with an 800-token reply cap) to 2,055. It also left an imported-knowledge fixture with no room to retrieve anything. Both failures were real consequences of the prompt growth, not faulty tests. The full-suite run that found them was stopped as void, and the form was changed to compact JSON with ellipsis placeholders (§G). Both tests pass unchanged on the final form, and no test was edited to accommodate it.

G. Prompt changes

G.1 Semantic before and after

v1.0.0 v1.1
How the vocabulary is shown set_possession(item, owner) — gives an item to an owner, a call notation that is not the wire format {"type":"set_possession","item":"…","owner":"…"} — gives an item to an owner, the object itself
Optional fields name(field, optional?) (optional: description, aliases) after the summary
List-valued fields not distinguished shown as ["…"]
Identifier guidance "short lower-case slugs (mara, silver-key, old-abbey)" "short lower-case slugs … the example's identifiers are placeholders"
Worked example silver-key owned by aldric, who is at old-abbey item-1 owned by character-1, who is at location-1
Length hint unchanged text unchanged text, built from shared constants
Reminder, canon, state rendering unchanged unchanged

G.2 Cost

v1.0.0 v1.1 Change
events.vocabulary_for_prompt() 258 tokens 378 +120
EMIT_RULE (which includes the vocabulary) 467 588 +121

EMIT_RULE sits in the static system block, so it is priced once into the cached prefix and into protected context on every turn. At a 4,096 window, +121 tokens is about 3% of the window, and it comes out of the history. The cheapest option measured was a plain name: fields list at 274 tokens. It was rejected because it is not the wire format, and because a copy of it in prose is not a proposal the extractor can recognise and remove. A copied JSON line is.

G.3 Guard against reintroduction

test_v11_protocol_echo.py covers the following:

  • Fixture identifiers. No identifier from either acceptance fixture, and no genre noun, appears in EMIT_RULE, EMIT_REMINDER, the vocabulary, or any length hint.
  • The worked example. It is a block the extractor accepts, with allowed event types.
  • Call notation. No vocabulary line is a call.
  • Line shape. Every line parses as the object, with exactly the required fields.
  • Hint wording. The length hint carries the constants the extractor recognises.

H. Cleanup rules (WP-A2)

H.0 The evidence the rules were designed from

Before any extractor change, the whole v1 corpus was scanned. That covers every AI turn in every evidence database under $HOME/m11-evidence, plus the exported bundles. Stored text and stored raw replies were counted separately, and duplicates were removed by content hash. The scan counted the candidate shapes, and the near-misses a rule must not touch.

Shape Unique stored texts Where
A line opening with a vocabulary call 4 All four closeout identity turns (depths 10, 12, 14 and 20).
A vocabulary call anywhere 4 The same four lines. The corpus has no call in the middle of a sentence.
[Hard limit anywhere 3 See the three cases below.
A rendered Scene: … (at …) as the last line 1 Closeout browser-run3 action 3, stored by the v1.0.0 product tree.
Scene: lines anywhere 62 Mostly pasted state sections from runs taken before M11's 0c7316f fix. The v1.0.0 extractor already handles those, and the replay (§I) starts from the raw reply.

The three [Hard limit cases:

  • Identity depth 14: the application's length hint, reworded at the front.
  • m04-rerun action 151: the hint's own words, with the last sentence replaced by "This story ends here." An empty dangling ```json fence follows it.
  • browser-recheck2 action 5: "[Hard limit: I have adhered to the prescribed limit of words. Here's the narrated turn and the state update. Send your next action.]" The model invented this; it is not application text.

H.1 The rules

Each rule is anchored to something the application owns: the event vocabulary, the length hint's wording, or the renderer's scene-line format. None of them judges prose.

R1 — event-call line (RULE_EVENT_CALL)

Source events.vocabulary_for_prompt printed every event as name(field, …), and the narrator copied the notation.
Matcher A whole line that starts with a name in events.SPECS (case-insensitive) followed by (. Indentation and a leading > are allowed before the name, and spaces before the (. The whole line is removed, including anything after the call on that line, such as depth 20's "Adds the silver key…". Lines inside a fenced code block are never examined.
Rationale create_entity, set_possession and the rest are this protocol's snake_case identifiers. A story line that starts with one of them followed by ( is protocol.
False-positive defence Not applied in the middle of a sentence (the engineer's whiteboard). Not applied to call-shaped lines whose name is not in the vocabulary (> open_door(north)). Not applied inside a story's own code block. Not applied to a vocabulary word with no ( ("Create_entity is a terrible name").

R2 — echoed length hint (RULE_LENGTH_HINT)

Source builder.length_hint. Its opening [Hard limit: and its closing sentence are now named constants shared with the extractor (LENGTH_HINT_OPENING, LENGTH_HINT_TAIL).
Matcher A bracket at the end of the reply, closed or cut off; the existing trailing loop already isolates it. Its contents must start with Hard limit: and carry the application's own wording: either the tail phrase "append the state block", or "turn must not exceed N words". The check goes in _is_echoed_instruction, beside the reminder and continue-hint checks M11 already has.
Rationale Both phrases are the application's instruction text. The 3B narrator reworded the front ("your next turn") and kept the rest.
False-positive defence Only a whole bracket at the end of the reply is removed. These all stay: "[Hard limit: 500 words]" read aloud mid-sentence; an in-world "[Hard limit of the reactor…]"; a final "[Hard limit: forty days, no extensions]". The model's invented bracket in browser-recheck2 action 5 also stays, because it carries neither phrase. That is a deliberate miss.

R3 — rendered scene line (RULE_SCENE_LINE)

Source The first line of render.for_prompt: Scene: <summary>, plus (at <location>) when the scene has a location.
Matcher After the other trailing cuts, the reply's last line starts with Scene: and meets one of two conditions. Either it ends in the renderer's (at <location>), or protocol was also cut from the end of the same reply (a block, a hint, a call line or a dangling fence).
Rationale The (at …) suffix is the renderer's syntax. Without it, a lone scene line directly above removed protocol belongs to the same pasted tail.
False-positive defence Never removed mid-reply: "Scene: a kitchen, late." followed by story stays, and so does M11's existing "Scene: the docks at dawn." case. A screenplay-style last line, "Scene: take two, and nobody moves.", stays when nothing else was removed. Accepted risk: a story whose last line exactly imitates Scene: … (at …) loses that line.

R4 — empty dangling fence (RULE_EMPTY_FENCE)

Source A reply that hit the output limit after opening ```json and before writing anything (m04-rerun action 151).
Matcher A ```json or bare ``` opener as the reply's last line, with nothing after it.
Rationale M11's _is_opening_of_proposal already cuts a dangling fence that stopped after {, on the principle that "a story's own code block is not that short, and one that is holds nothing to lose". An empty opener holds even less.
False-positive defence Applies only when nothing at all follows the opener. A dangling fence with any content keeps M11's existing judgement.

H.2 What is deliberately not removed

  • A fact restated inside a sentence (§P risk 4). Removing it would mean judging prose.
  • Model-invented headings that are not the renderer's (§P risk 6).
  • Model-invented brackets that carry no application phrasing (browser-recheck2 action 5).
  • Prose that follows a removed call line, even when it narrates the call ("You create a new person, Mike, who has", identity depth 12). It is narration, however poor.

I. Historical replay

Tool: tools/v11_replay_extractor.py, on the final A2 extractor.

Inputs: every AI turn in the 17 evidence databases under $HOME/m11-evidence, plus the closeout identity run's bundle. A turn is replayed from its stored raw_output where one exists, so the comparison starts from what the narrator actually sent. The identity bundle carries only stored text; for those turns the v1.0.0 prose is the input itself.

Comparison: each input goes through the extractor as tagged at v1.0.0, read with git show, and through the v1.1 extractor. Lines are compared with trailing whitespace ignored.

Evidence: $HOME/v11-evidence/a2-replay/replay.json (sha256 f04c54ff…) and replay.md, which holds the full v1.0.0 and v1.1 prose for every changed turn.

Unique replies replayed 518 (6 exact duplicates skipped)
Unchanged 509
Changed 9
Flagged (an unexplained removal, text added or rewritten, or more than 25% removed) 0
Removed lines by rule R3 scene line 6, R1 event call 4, R2 length hint 2, R4 empty fence 1

The planning pass estimated about 453 turns; the actual count is 518. The difference is the trial, browser and recheck databases, which the v1 figure of 443 did not count.

I.1 Every changed turn

# Source Removed Rule Share of prose removed Review
1 browser-final action 3 (raw) Scene: Aldric and Mara by the fire in the Crooked Lantern (at The Crooked Lantern). R3 13.1% The rendered scene line, directly above a ```json proposal in the raw reply. Protocol.
2 closeout browser-run3 action 3 (raw), stored by the v1.0.0 product tree Scene: Aldric and Mara by the fire in The Crooked Lantern. (at The Crooked Lantern) R3 20.7% The same, in the renderer's exact form. Protocol.
3 m04-rerun action 151 (raw) Scene: Aldric and Mara in the Crooked Lantern, rain outside.; [Hard limit: this turn must not exceed 180 words, … This story ends here.]; ```json R3, R2, R4 11.9% A pasted scene line, the application's length hint with its last sentence reworded, and a proposal opener the output limit cut off. Protocol.
4 m04-rerun action 153 (raw) Scene: Aldric and Mara in the Crooked Lantern, rain outside. R3 3.1% The scene line directly above a ```json proposal. One trailing space was also trimmed from the story line before it; no word changed. Protocol.
5 m04-rerun action 157 (raw) Scene: Aldric and Mara in the Crooked Lantern, rain outside. R3 2.8% The scene line directly above ```json\n{. Protocol.
6 identity depth 10 (stored) > Create_entity(new_person, "john", "character", "A determined team member …", ["jo… R1 6.4% A call cut off mid-argument by the output limit. Protocol.
7 identity depth 12 (stored) > Create_entity(new_person, "mike", "character", "A new team member …", ["mike"]) R1 4.8% The call. The narration after it ("You create a new person, Mike, who has") is kept. Protocol.
8 identity depth 14 (stored) > Create_entity(mike, …); Scene: Bill, Alice, Roger, John, and Mike at the table; John not yet arrived. (at The meeting room); [Hard limit: your next turn must not exceed 180 words, … append the state block well inside the limit.] R1, R3, R2 23.4% The whole pasted tail: the call, the rendered scene line, and the reworded length hint. Protocol.
9 identity depth 20 (stored) > set_possession(silver-key, "alice") Adds the silver key to Alice's possession. R1 4.1% The call, with the fantasy example slug and a gloss on the same line. Protocol.

Genuine prose lost: none. Each of the 13 removed lines is application protocol or instruction text: a vocabulary call, the length hint, the renderer's scene line, or an empty fence opener. On a line that also carried narration, only depth 20's same-line gloss went with the call; that gloss describes the call, not the story.

A replay-tool defect, found and fixed before this result. The first replay flagged turn 4 as "text rewritten". The trailing space trimmed from the story line was compared as a changed line. The tool now ignores trailing whitespace and says why. The flag was the tool's, and no rule changed as a result.


J. WP-A2 real-model evidence

Inference: the trusted-LAN CPU reference host. It runs Ollama 0.33.0 over HTTPS with a private CA and verification on, and serves qwen2.5:3b-instruct at a verified 4,096 window. Embeddings are nomic-embed-text. The storyteller was on loopback.

Tree: the final A2 extractor and prompt. The full backend suite passed on this tree before the long run was allowed to start. The launcher enforced that gate.

J.1 Multi-character identity diagnostic

The office meeting fixture (TEST-CAMPAIGN-FIXTURE.md Appendix A), 10 beats, memory off. Evidence: $HOME/v11-evidence/a2-identity.

v1 closeout (§S.5 of the M11 report) v1.1
Turns accepted 10 of 10 10 of 10
Identity signals 0 0; verdict "no objective identity defect detected"
Proposals naming an example identifier from the fixed prompt 1 (silver-key given to Alice) 0. No silver-key, aldric, mara or old-abbey, and none of the new placeholders either
Proposal outcomes 1 applied, 5 without a block, 3 unparseable, 1 refused 10 accepted, 1 rejected
Stored turns with a shape R1-R4 targets 4 of 10 (depths 10, 12, 14, 20) 0 of 10
Stored turns with any echoed application instruction 4 of 10 1 of 10: depth 16, below

A leak A2 does not remove. Depth 16 kept this tail at the end of the story:

Scene: Bill, Alice and Roger at the table; John not yet arrived.

[Hard limit: this is now 180 words.]

[Reminder: end your reply with a `state` block listing the events your narration made true, with absolute values.]

[You don't need to continue; your turn must now be about John entering the room. Continue the story here, directly. Output only story text.]

The trailing cleanup works from the end upwards, and the last bracket is not recognised: it is a reworded continue hint, and the recogniser matches that hint only by its opening words. So the reminder above it is never at the end, the reworded length hint carries none of the application's wording, and the scene line is never last. This is the conservative design behaving as documented. It is still a real leak, on 1 of 10 turns. The harness's leak counter did not see it: it looks for proposals, pasted headings, calls and the hint's own wording, not reminders. It was found by reading the stored narration. A narrow follow-up is in §N.

J.2 The 50-turn long run

tools/m11_long_run.py --turns 50: the fantasy Continuity Test with the full schedule of history operations and the memory bank on. Evidence: $HOME/v11-evidence/a2-long-run.

Status complete; no aborted reason, no failed reason
Accepted turns 51: 50 scheduled and the recall turn
Wall clock 9,324 s. A turn took 47-243 s, median 165
Genuine process restarts 3, each compared identical: true
History operations performed 2 Save Points, Undo/Redo, 2 retries, undo then divergence, take selection, a Save Point restore, and an induced failed call (HTTP 404; state unchanged)
Post-turn failures 0. 0 tracebacks and 0 database is locked in server.log
Memories and summaries written 15 memories (9 in the bank at the end); 5 summaries
Protocol leakage 0 of 54 stored AI turns, by the extended harness count (headings, "events", call lines, the echoed hint). A separate scan of every stored turn for reminders, [Hard limit, continue-hint wording, calls, a final Scene: and "events" also found 0
Parser failures 6 proposals unparseable, out of 56 recorded
Refused proposals 0
State updates from narration 4 events accepted from story (50 proposals accepted, most with an empty events list); 9 more from the harness's manual corrections
Legitimate prose removed none. All 64 raw replies from this run and the identity run were replayed through the v1.0.0 and v1.1 extractors. One changed: a trailing ```json opener with nothing after it (R4, 0.4% of that reply). 0 flagged
M04-style recall recovered_through_state_only, which is WP-B's subject and unchanged from v1

A1 accounting across the run.

Window verified at 4,096 on 51 of 51 turns
Status fits on 51 of 51
Server count minus the application's estimate +15 on every turn
Largest server prompt 3,321 tokens
Observed margin beside the reply 275 minimum, 620 median, 2,297 maximum, against a 256 reserve

A defect the long run exposed, fixed in A1 (§N.2 item 7). Three of the 54 stored AI turns had no accounting: the two retried takes and the take superseded by a take selection. attempts.ATTEMPT_KEYS lists the snapshot slices that belong to one attempt, and A1 had not added accounting to it. When the live flag moved, the new live take was handed the old take's accounting with the shared prompt, and its own was discarded. The long run's numbers above come from the done events and are unaffected. It was the stored accounting of a retried turn that could be wrong.

J.3 Genre fixtures (A2-8)

Fixture How Result
Fantasy Continuity Test the 50-turn long run, real model 0 example identifiers proposed. The fixture's own silver-key is legitimately its entity
Multi-Character Identity Test (office) the identity diagnostic, real model 0 proposals naming any prompt example identifier (v1: 1)
Persephone science fiction test_m11_scifi.py, 10 tests, scripted passes; no genre noun in any fixed instruction (test_v11_protocol_echo.py)

Not run: science fiction against a real narrator. No harness drives that fixture with a real model, and adding one was out of A2's scope.

J.4 What this does not replace

This is not the integrated v1.1 release run (plan §11). That run covers 100 turns, the 16,384 window, and the GPU host with the owner's logging.


K. Compatibility

K.1 A real v1.0.0 database opens unchanged

Source database: m04-final/campaign.db, the v1 evidence run's own campaign. It was written by 96c1bf5, whose product code is identical to v1.0.0, and holds user_version 94. Its contents:

  • 207 actions on 4 branches, with retained history;
  • 2 Save Points and 32 state events;
  • 33 memories and 12 summaries;
  • 3 imported knowledge sources, one campaign narration-length choice, and settings.

Method: tools/v11_compat_check.py. The database was copied, and each copy opened by one tree:

  • v100: the v1.0.0 tag, from a scratch worktree;
  • v11-readonly: the v1.1 tree.

Each tree's app was verified by the path it was imported from. The same read-only snapshot was taken through the API for both:

  • the full export bundle (which carries no timestamp of its own);
  • narrative state and state events;
  • Save Points, knowledge, memories and derived status;
  • the newest actions and settings;
  • the schema and user_version before and after opening.
Comparison Result
Schema before opening, v1.0.0 against v1.1 identical
Schema after opening identical; user_version 94 → 94 on both, so no migration ran
API snapshot identical, field for field

K.2 The same database, exercised on v1.1

--exercise, on a third copy, with the endpoint first pointed at a refused loopback port so nothing left the machine:

Operation Result
Undo 119 → 117 moments; Redo becomes available
Redo back to 119; Redo no longer available
Restore Save Point On the ridge 69 moments, with Redo available into the later story
Context dry run HTTP 200. One imported passage retrieved. The campaign canon and narrative state sections are present. The summary that applies (moments 49-64) is included. Budget 16,384: 14,577 prompt + 28 transport + 500 reply + 820 reserve = 15,925. The window is recorded as unverified, because the server was unreachable by design.
Export, then import into the same install 207 → 207 actions, the same head depth, Save Points 2 → 2, memories 33 → 33, narrative state identical

Not exercised here, and why. Writing a new memory or summary needs a model and an embedding model. The real-model long run in §J writes both on the v1.1 tree. Memory retrieval with embeddings is covered the same way. No data-migration path was needed, and none was added.


L. Security and offline

L.1 What the diff adds, measured against v1.0.0

git diff v1.0.0 over backend/app and frontend/src, searched for any new network or remote-resource use. The search covered httpx, requests, urllib, socket, aiohttp, literal URLs, fetch(, XMLHttpRequest and tiktoken.get_encoding.

Check Result
New network client, destination or URL none
New imports in product code math, json, logging, and one internal constant (providers.openai_compatible.CHAT_CONTINUE_HINT)
Remote or cloud token counting none. The count is the vendored cl100k_base, as before. The server's own count arrives in the response the turn already receives
Calibration service none. Calibration was not built
The one request-level change stream_options: {"include_usage": true} in the body of the request already sent to the already-approved endpoint. It adds no destination and no data about the user
Endpoint policy (endpoints.py, ADR 011) unchanged; the file is not in the diff
Probe policy (contextwindow._ask) unchanged; A1 added pure functions beside it
Storyteller bind unchanged; docker-compose.yml, Dockerfile and the start scripts are not in the diff

L.2 Regression runs

  • The endpoint-policy and local-only tests pass in the backend suite (§M): test_m11_security.py (H-series, including a public endpoint refused at request time), test_egress.py, test_local_only_surface.py (loopback publication), test_m11_context_window.py (the probe held to the policy).

  • Trusted-LAN inference still works. Every real-model turn in §D and §J went to the trusted-LAN host over HTTPS, with a private CA and verification on.

  • The offline container passed 23 of 23. It ran tools/m11_offline.py on the final tree, with --network none, a fresh volume, and an image built with docker build --no-cache. The checks are the same 23 as the v1 closeout:

    • no route and no DNS;
    • first page load, no remote origin, a CSP, and every asset local;
    • campaign creation, state extraction, knowledge import, prompt assembly and retrieval;
    • a turn with no model reachable is reported, with no narration accepted, the player's words kept, state unchanged and the earlier story intact;
    • export and import with no secret;
    • the media module inert;
    • campaigns survive a container restart.

    The evidence is $HOME/v11-evidence/offline/offline-report.json. The A1 accounting path is exercised offline by the failed-turn check: a turn that never reaches a model writes no accounting, because it commits nothing.


M. Full test results

Every backend command below is .venv/bin/python -m pytest tests/ -q from backend/, with no AIDND_TEST_* set. The 17 skips are the tests that need a real local model, the same 17 as the v1.0.0 closeout. No new skip was introduced. The three real-model skips in the targeted runs are those same tests.

Run Tree Result
Backend, baseline untouched ac465ed 1,421 passed, 17 skipped, 0 failed, 866 s
Backend, A1 checkpoint A1 only; the not-yet-implemented A2 test file excluded 1,454 passed, 17 skipped, 0 failed, 835 s
Backend, A2 first full run A2 with the first vocabulary form stopped and discarded. Two budget failures, §F
Backend, A1 + A2 before the take-accounting fix 1,505 passed, 17 skipped, 0 failed, 786 s
Backend, final A1 + A2 + the take-accounting fix: the staged tree 1,507 passed, 17 skipped, 0 failed, 792 s
New tests test_v11_context_reserve.py (35) and test_v11_protocol_echo.py (51) all pass
Frontend A1 (A2 changes no frontend file) 165 of 165, 14 files
Lint A1 exit 0, six only-export-components warnings, the same six as the v1 closeout
Production build final clean. 16 files, index-rWgm40jJ.js 394.47 kB, index-BXtb0iME.css 48.09 kB
Offline container, --network none, docker build --no-cache A1 + A2, before the take-accounting fix 23 passed, 0 failed (§L.2)
Offline container, repeated final, after the take-accounting fix 23 passed, 0 failed, from a fresh --no-cache image ($HOME/v11-evidence/offline-final)
Identifier scan of every changed and new file final no real hostname, address, domain or personal name

N. Findings

N.1 Blockers

None found.

N.2 Corrective work, done within the packages

  1. A1: stream_options was required, not optional. The provider's docstring said stream_options was "deprecated and does nothing", which is true of OpenRouter. Against Ollama 0.33 a stream sends no usage without it, and none of the 514 v1 evidence turns stored a count. Without the fix, A1's accounting would have read unknown on every real turn.
  2. A1: one test's fixture re-calibrated (§C). The history-floor prefix test had assumed which turn holds the floor. Behaviour was unchanged.
  3. A2: the first vocabulary form was too expensive (§F, §G). It grew protected context by 223 tokens and failed two existing budget tests. It was replaced by compact JSON (+121); the tests pass unchanged.
  4. A2: module syntax. Two docstring paragraphs were misplaced. Collection failed immediately, and they were corrected before any evidence was taken.
  5. Harness: the replay tool compared trailing whitespace (§I), which produced one false "rewritten" flag. Fixed.
  6. Harness: the accounting tool imported the bundle's memory setting (§D.1). That caused failing post-turn calls on the shared host. Fixed before the 16k runs.
  7. A1: a retried or reselected take lost its accounting, or showed another take's.
    • Where it was found: after the A2 long run. Three of 54 stored AI turns had no accounting, and they were exactly the two retries and the take selection.
    • Cause: attempts.ATTEMPT_KEYS lists the snapshot slices that belong to one attempt, and A1 had not added accounting to it. It was therefore treated as part of the shared prompt. hand_over_the_prompt gave the new live take the old take's accounting, and discarded its own.
    • Fix: accounting added to ATTEMPT_KEYS, with a unit test and an end-to-end retry test (§C).
    • Exports: none needed a change. A snapshot travels whole.
    • Migrations: none. v1.0.0 turns have no accounting to move, and migrations._ATTEMPT_KEYS is a frozen historical copy, left as it is.
    • Evidence: the long run's accounting figures in §J come from the done event of each turn and are unaffected. The full backend suite and the offline container were re-run after the fix (§M).

N.3 Non-blocking residual risks

  1. Cold-model first turn: truncation is detected, not prevented (§D.1). When /api/ps cannot see the model and /api/show finds no num_ctx, M11 builds to the configured budget. On a server whose real window is smaller, the first turn is silently cut. It was observed for real: 13,875 sent, 2,050 read. A1 now records it, logs it and shows it. Preventing it needs a policy decision, which is the owner's and outside A1's scope. The two options are to build conservatively when the window is unknown, or to warm the model and probe again before the first turn.

  2. The drift measured is one model's. +15 tokens against cl100k_base, for qwen2.5:3b-instruct in both window sizes. A model whose tokenizer diverges by more than 5% will show up as exceeded, not as silent truncation. The reserve is a tolerance, not a guarantee.

  3. exceeded and truncation_suspected depend on the server reporting usage. A server that ignores include_usage produces unknown on every turn. That is honest, but it detects nothing.

  4. R3's accepted false-positive risk. A story whose last line exactly imitates the renderer's Scene: … (at …) loses that line (§H).

  5. Protocol still left in stored narration, deliberately:

    • model-invented brackets that carry no application wording;
    • model-invented headings;
    • facts restated inside a sentence.

    One real v1.1 turn shows the cost (§J.1, identity depth 16). A reworded continue hint was the last line of the reply, so nothing above it became trailing. The application's reminder, a reworded length hint and a scene line all stayed.

    Narrow follow-up, not done here. Recognise CHAT_CONTINUE_HINT by its distinctive application phrase "Output only story text", not only by its opening words. The reminder above it would then come off by the existing rule. The reworded hint, with none of the application's wording, and the scene line would still stay.

    Why it was not done. It would have been an extractor change after the long-run evidence was taken, and it needs its own replay. This is the owner's decision.

  6. A2 costs 121 tokens of protected context (§G.2). At 4,096 that is about 3% of the window, and it comes out of history.

  7. A2's real-model run is on the CPU host at 4,096, not the v1 evidence configuration. v1 used 16,384 on the GPU host. On the CPU host a full 16k prompt took 17 minutes a turn, and a GPU-host long run requires the owner's power, link and kernel logging to be started first (gpu-host-logging-before-long-runs). The integrated v1.1 release run (plan §11) is where 16,384 is re-established.

N.4 Future ideas

  • A cold-model policy for the first turn, as in N.3 item 1.
  • Summarising the accounting across a campaign, for example "3 of 120 turns suspected". Today the inspector shows it one turn at a time.
  • Showing the accounting status as a chip under the narration, not only in the context inspector. This is browser work, for WP-C.

O. Final A1 decision

A1 RESULT: PASS.

The final decision is taken on the staged tree, including the take-accounting fix found after the interim checkpoint (§N.2 item 7).

Criterion Result Evidence
A1-1 Reserve calculation PASS max(256, ceil(5%)): 256 at 4,096, 410 at 8,192, 820 at 16,384; 7 windows tested (§B.3, §C)
A1-2 Budget enforcement PASS assembled text, provider hint, reply and reserve are within the effective window in 5 configurations (§C)
A1-3 Protected overflow PASS fails before the model call; the provider is called 0 times and nothing is written (§C)
A1-4 No silent canon dropping PASS the canon sentinel is present in every configuration while history gives way (§C)
A1-5 Post-response accounting PASS real: 51 of 51 long-run turns, 4 turns at 4,096 and 2 at 16,384 carry a server count and a status (§D, §J)
A1-6 Unknown accounting PASS no usable count gives unknown, never fits (§C)
A1-7 Unexpected discrepancy visible PASS a real truncation_suspected on a cold model: 13,875 sent, 2,050 read. It was stored, logged, sent on the done event and shown in the inspector (§D.1)
A1-8 Accepted turn preserved PASS the turn is kept for truncation_suspected and for exceeded; after a retry, each take keeps its own accounting (§C)
A1-9 Existing behaviour PASS 1,507 passed, 0 failed. One fixture re-calibrated, not weakened (§C)
A1-10 Real-model verification PASS at 4,096, 275-2,297 tokens left beside the reply across 55 turns; at 16,384, 1,994. Drift is +15 on every turn, against v1's 23-42 tokens of headroom (§D, §J)

Carried as residual risk, not failure: the cold-model first turn is detected, not prevented. Preventing it is a policy decision for the owner (§N.3 item 1).

P. Final A2 decision

A2 RESULT: PASS, with one recorded residual.

Criterion Result Evidence
A2-1 Genre neutrality PASS no fantasy or science-fiction identifier and no genre noun appears in any fixed instruction; a test guards against reintroduction (§G.3)
A2-2 Protocol clarity PASS the vocabulary is the wire-format object, not call notation. +121 tokens, the smallest wire-format form measured (§G)
A2-3 Observed leaks PASS all four v1 shapes are removed (§H). What is deliberately left is documented, including one real v1.1 occurrence (§J.1, §N.3 item 5)
A2-4 Historical corpus safety PASS 518 real replies replayed, 9 changed, 0 flagged, no story prose removed. A further 64 replies from the v1.1 runs gave 1 changed, 0 flagged (§I, §J.2)
A2-5 Adversarial prose PASS the owner's four sentences and nine more survive untouched (§C of the A2 tests, test_v11_protocol_echo.py)
A2-6 State extraction PASS a valid block still parses and applies when protocol litter surrounds it. SPECS, render.py and validate.py are identical to v1.0.0 (§F, §K)
A2-7 Invalid proposals PASS the validator is byte-identical and the existing refusal tests pass; the long run recorded 0 refusals and 6 unparseable proposals (§J.2)
A2-8 Genre fixtures PASS fantasy by real model (51 turns); office identity by real model (10 turns, 0 example identifiers against v1's 1); science fiction scripted only, which is stated (§J.3)
A2-9 Real-model sample PASS 51 accepted turns: 0 of 54 stored turns carry protocol, 6 parser failures, 0 refused, 4 story state events, no prose removed, and A1 accounting fits 51 of 51 (§J.2)

The residual. In the identity run, 1 of 10 stored turns kept a trailing tail of echoed instructions: a reworded continue hint, the reminder, a reworded length hint and a scene line. The conservative design cannot see past the reworded last line. §N.3 item 5 records a narrow follow-up, which is the owner's decision.

Q. Recommendation

  • A1 is ready for acceptance. It meets every criterion on real evidence, including a real silent truncation it caught.
  • A2 is ready for acceptance, with the residual in §P noted. The owner may want the narrow continue-hint follow-up (§N.3 item 5) first. It would be a small extractor change with its own replay. It does not block acceptance of what A2 set out to do.
  • Owner decisions this report raises:
    1. Accept or correct A1.
    2. Accept A2, or ask for the continue-hint follow-up first.
    3. Whether the cold-model first turn (§N.3 item 1) needs a prevention policy, and where that policy belongs: an A1 follow-up, or a later package.
  • WP-B may begin once A1 and A2 are accepted. Nothing in either package blocks B's diagnosis. A1's accounting and A2's cleaner prose are both in place for B's long runs. WP-B has not been started.

Git state: everything is staged and uncommitted for the owner's signed commit. Nothing was committed, pushed or tagged. A suggested commit message was prepared outside the repository.


R. Corrective-work addendum (owner review, 2026-09-14)

The owner accepted A1 and A2 in principle, and asked for two corrective items to be closed before the commit. §O-§Q above record the decisions as they stood before this addendum. §R.6 supersedes them.

R.1 WP-A1: preventing the cold-model truncation

The case, as A1 caught it (§D.1, turn 1):

  • /api/ps knew nothing, because the model was not resident.
  • /api/show found no num_ctx, so the window was unverified.
  • The prompt was built to the configured 16,384.
  • Ollama loaded the model at its 4,096 default, read 2,050 of 13,875 tokens, and answered HTTP 200.

Mechanism, measured before it was built. This was measured on Ollama 0.33 on the trusted-LAN host, with the model first unloaded (keep_alive: 0) and /api/ps empty.

Step Observed
POST /api/generate {"model": "qwen2.5:3b-instruct"}, with no prompt HTTP 200 in 3.7 s, "response": "", "done_reason": "load". Nothing was generated.
/api/ps afterwards the model is listed with context_length 4,096
An OpenAI-compatible streamed request afterwards no reload: /api/ps still reads 4,096, and prompt_tokens was the estimate plus 13

What was built.

  • contextwindow.Window gains reachable: the server answered a discovery request, whatever it said.
  • contextwindow.warm(endpoint, model, timeout) sends one POST {native base}/api/generate with the body {"model": model}. There is no prompt, no options and no keep_alive. It goes through endpoints.rejection_reason and tlstrust.ssl_context(), and it never raises.
  • contextwindow.ensure_window(endpoint, model, declared, warm_timeout) does the following:
    1. It probes. A verified window is returned untouched.
    2. If the window is unverified, the server was reachable, and a model is configured, it warms once.
    3. If the load succeeds, it probes again with use_cache=False.
    4. It returns the window and a preflight record (attempted, loaded, verified_before, verified_after, detail).
  • turns._generate_turn calls ensure_window in place of probe, with the settings' model timeout as the bound. It stores snapshot["window"]["preflight"]. The context dry run still only probes: loading a model from a read-only panel would be a side effect.

What is unchanged.

  • A window still unverified leaves the configured budget standing.
  • There is no guessed window, no hard-coded 4,096, and no retry loop.
  • The accounting still classifies the reply.

The constraints the owner set, each checked:

  • no calibration and no user setting;
  • no endpoint-policy change;
  • only the configured endpoint is contacted;
  • no dummy narration is generated or accepted;
  • no accepted turn is discarded.

Tests: tests/test_v11_cold_window.py, 15 tests.

Required Test
cold model, then preflight, then a verified smaller window, then the real turn built to it test_a_cold_model_is_loaded_once_and_its_window_verified (request order /api/ps, /api/show, /api/generate, /api/ps; the load body exactly {"model": …}). test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports, end to end: budget 4,096, verified, the canon sentinel present, the assembled text plus transport, reply and 256-token reserve within 4,096
negative control test_without_the_load_the_same_cold_turn_would_have_been_built_too_large: probing alone gives an unverified window, budget 16,384, and a prompt more than twice 4,096
already loaded: no warm test_an_already_loaded_model_is_not_warmed
load succeeds, verification still unavailable test_a_model_that_loads_but_still_cannot_be_read_stays_unverified: one load, unverified, the configured budget kept
load fails test_a_failed_load_is_recorded_and_leaves_the_window_unverified (HTTP 404 and 500; bounded, with no second probe). test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe: the error is reported, with no AI action, state event or proposal. test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer: accounting unknown
the load writes nothing test_the_load_itself_writes_nothing: actions, state events, proposals, memories and summaries are counted before and after, and are identical
no new destination test_the_load_request_obeys_the_endpoint_policy (public addresses refused with no transport installed). test_the_load_request_goes_only_to_the_configured_host. test_an_unreachable_server_is_not_asked_to_load_anything
also test_a_declared_window_does_not_stop_the_server_being_asked; test_the_context_dry_run_never_loads_a_model

Real cold-model test: §R.3.

R.2 WP-A2: the continue-hint instruction tail

The exact output. Identity depth 16 (action 17 in the run's database). The stored text equals the v1.1 extractor's output of the raw reply, and the raw reply has no state block. It ends:

Scene: Bill, Alice and Roger at the table; John not yet arrived.

[Hard limit: this is now 180 words.]

[Reminder: end your reply with a `state` block listing the events your narration made true, with absolute values.]

[You don't need to continue; your turn must now be about John entering the room. Continue the story here, directly. Output only story text.]

Why the application-owned material above the last bracket was not removed. _clean cuts the end of a reply one bracket at a time, and only ever examines the last bracket. That last bracket is the provider's CHAT_CONTINUE_HINT, reworded at the front. _is_echoed_instruction recognised that hint only by its opening words ("continue the story directly"), so it returned false, and the loop stopped. The reminder above it would have been removed by the existing Reminder: rule, but it was never at the end. The reworded length hint carries none of the application's wording, and the scene line was never last.

The rule, R5 (RULE_INSTRUCTION_TAIL). Two narrow changes:

  1. Recognise the continue hint by its own sentence. A trailing bracket containing "Output only story text" (CONTINUE_HINT_PHRASE) is an echoed instruction. A test pins the phrase to CHAT_CONTINUE_HINT, so the two cannot drift.
  2. Remove a hint-opened bracket only directly above an echo. A trailing bracket opening with Hard limit: is removed only when an echoed instruction bracket has already been cut from the end of the same reply.

The rest follows from existing rules. The reminder is removed by the Reminder: rule once it is last. The scene line is removed by R3 once it is last and protocol has been cut.

False-positive defence. Each of these is a test that must pass unchanged:

Case Why it stays
A final [Hard limit: forty days, no extensions] no echo was cut below it
The same bracket above a state block a block is not an echoed instruction, so a block below does not license it
[The sign on the door reads: Closed] directly above an echoed continue hint it does not open the way the hint opens; only the echo goes
"output only story text" inside a sentence mid-story only a trailing bracket is examined
A final [To be continued] no application phrase

The owner's four original adversarial sentences and all earlier preservation cases also still pass.

Accepted risk. A story-world bracket that opens exactly [Hard limit: and sits directly above a parroted instruction would be removed with it.

Positive regression: test_the_depth_sixteen_instruction_tail_leaves_entirely. The two story paragraphs are shortened; the four trailing lines are verbatim. It leaves exactly the story.

R.3 Evidence

Check Result
New and related unit tests (the A1 cold-window, A1 reserve, A2 echo, extractor, window, take and caching tests) 308 passed, 3 skipped (the real-model skips)
Replay: the full v1 corpus 518 replayed, 509 unchanged, 9 changed, 0 flagged. The changes are exactly the nine in §I, with the same rules: R5 changes no v1 turn
Replay: the v1.1 runs' raw replies (the 51-turn run's database and the identity run's database) 64 replayed, 61 unchanged, 3 changed, 0 flagged:
long run, action 77: R4, a trailing ```json (0.4%), as in §J.2
identity depth 16: R3 scene line, plus R5 × 3 (the reworded length hint, the reminder, the reworded continue hint). 18.0% of the reply, all of it the tail above
identity depth 18: R4, a trailing ```json (0.8%)
Genuine story prose removed 0
Unreviewed replay flags 0
Real cold-model A1 test (v11_window_accounting --unload-first, same host, model and 207-action bundle as §D.1) Prevented. Model unloaded (/api/ps empty). Turn 1: preflight attempted, model loaded, window verified 4,096 (loaded), budget 4,096, estimate 3,082, server read 3,097 (+15), margin 499, fits. Turn 2: no preflight needed (already verified), estimate 3,086, server 3,101, margin 495, fits. Before the correction the same turn sent 13,875 against an unverified window and the server read 2,050
Identity diagnostic, re-run (office fixture, 10 beats, the same CPU/HTTPS host at 4,096; $HOME/v11-evidence/a2c-identity) identity signals: 0 (verdict: no objective identity defect). Prompt example identifiers in proposals: 0 (no silver-key, aldric, mara, old-abbey, nor the placeholders). Stored protocol/instruction shapes: 0/10 (calls, hint and reminder brackets, continue-hint wording, a final scene line, state headings, "events", fences). Proposals: 9 accepted, 1 partially accepted, 1 unparseable. Read with the replay: this run's 10 raw replies replay unchanged through v1.0.0 and v1.1, so this narrator produced no echo this time. 0/10 here is the required closeout result, not a live exercise of R5. R5's proof is the verbatim depth-16 regression fixture and the replay of the original depth-16 turn above
Backend full suite 1,534 passed, 17 skipped, 0 failed, 1,007 s (the 1,507 before this pass, plus 15 cold-window and 12 corrective A2 tests; the same 17 real-model skips)
Frontend suite 165 of 165 (no frontend file changed in this pass)
Lint exit 0; the same pre-existing only-export-components warnings
Production build clean, 16 files; index-rWgm40jJ.js 394.47 kB and index-BXtb0iME.css 48.09 kB, the same hashes as before, so the SPA is byte-identical
Offline container (tools/m11_offline.py, docker build --no-cache, --network none, a fresh volume) 23 passed, 0 failed on the corrective tree ($HOME/v11-evidence/offline-corrective). The cold-model load is inert here: the failed-turn check has no reachable server, so no load is attempted (reachable false), and the turn is reported and commits nothing, as before
Browser and context-inspector regression (tools/m11_browser.py, Firefox 155.0.1 headless, the built SPA served by FastAPI, the real narrator over trusted-LAN HTTPS; $HOME/v11-evidence/browser-corrective) 38 passed, 0 failed, 0 skipped. This includes F05 (the context inspector shows the assembled prompt), two real turns through the UI, Undo and Redo, hostile narration, G01 import, the dialog and accessibility checks, and CSP. The two browser-played turns' stored snapshots carry the new provenance: the window was verified at 4,096 (loaded), the preflight was not attempted because the model was already resident, and accounting was fits (server 876 against estimate 861, margin 2,820; server 1,024 against estimate 1,009, margin 2,672). The inspector's accounting and safety-margin display is covered by the component suite (§C), not by this harness
v1.0.0 database compatibility (a fresh v1.0.0 worktree against this tree, same database as §K) schema before: identical; schema after: identical; API snapshot: identical; user_version 94 → 94. Exercise: Undo 119 → 117, Redo 117 → 119, restore On the ridge → 69 with Redo available, dry run 200 with knowledge retrieved, export and import 207 → 207 actions with the same head and identical state

R.4 Findings from the corrective pass

Blockers: none.

Corrective work within the pass:

  1. Harness slip: two background watchers never ran their jobs. The watchers were meant to start the offline container after the suite, and the browser regression after the identity run. Each waited with pgrep -f on a command name that appears in its own command line, so each matched itself and never proceeded. Nothing ran and nothing was corrupted. Both watchers were stopped, and both runs were started directly (§R.3). This is the same class of slip as the earlier pkill self-match (§N.2). It sits in session scripting, not in the repository.

Non-blocking residual risks:

  1. R5 accepted false-positive risk. A story-world bracket that opens exactly [Hard limit: and sits directly above a parroted instruction would be removed with it (§R.2).
  2. Cold-model load: what it cannot fix.
    • It does nothing for a server that is unreachable, or that does not speak Ollama's native API. Both keep the unverified path, now recorded in window.preflight.
    • It does not change the window. A model loaded at a smaller default is still that window; the turn is built to it, not to the setting.
    • It costs one load request on a cold turn: 3.7 s measured on the CPU host for a 3B model, bounded by the model timeout.
  3. The identity closeout's 0/10 does not exercise R5 live (§R.3). The shape is proven by the verbatim fixture and the replay, not by this run.

Future ideas:

  • Warming the model when the reader opens a campaign, so the first turn does not pay for the load. This was deliberately not done, because the dry-run panel must have no side effects.

R.5 Closeout against the owner's required results

Required Result
A1: one bounded load of a cold model, no story content, then the window probed again built and tested (§R.1); real server: HTTP 200 in 3.7 s, "response": "", "done_reason": "load"
A1: turn built to the verified window when the load makes it known real cold test: verified 4,096, 3,082 sent, 3,097 read, fits. Before the correction: unverified, 13,875 sent, 2,050 read
A1: window still unknown → no guess, no hard-coded 4,096, the unverified path kept, accounting kept tested (§R.1)
A1: no calibration, no user setting, no endpoint-policy change, only the configured endpoint, no dummy narration, no discarded turn met (§R.1). The policy and same-host tests pass. The load writes nothing, tested by row counts
A2: minimised positive regression, and why the tail survived test_the_depth_sixteen_instruction_tail_leaves_entirely; root cause in §R.2
A2: narrowest rule; adversarial controls; all extractor tests R5; 12 new tests; the extractor, A2 and state tests pass
A2: replay of all available real turns 518 replayed, 509 unchanged, 9 changed, 0 flagged. Classifications exactly as §I.1. Plus the 64 v1.1 replies: 3 changed, 0 flagged
genuine story prose removed 0
unreviewed replay flags 0
identity signals 0
prompt example identifiers in proposals 0
stored protocol/instruction shapes 0/10
Backend full suite 1,534 passed, 17 skipped, 0 failed
Frontend suite / lint / production build 165/165; exit 0; clean and byte-identical
Offline container 23 passed, 0 failed
Browser and context-inspector regression 38 passed, 0 failed, 0 skipped
v1.0.0 database compatibility identical schema before and after, identical API snapshot, user_version 94 → 94. Exercise passed

R.6 Final decisions

WP-A1: PASS
WP-A2: PASS

WP-A1: PASS.

  • What is met: every criterion in §O, plus the owner's corrective requirement.
  • The case now prevented: the real cold-model truncation that A1 had only detected. The window becomes discoverable once the model loads, and the turn is built to it.
  • What still depends on detection: a server that cannot be asked, or that will not load the model on request. That falls back to the unverified path, which is unchanged, recorded, and caught by the accounting.

WP-A2: PASS.

  • What is met: every criterion in §P, plus the owner's corrective requirement.
  • The residual §P recorded is closed. The depth-16 instruction tail is removed by a narrow rule anchored to application-owned text.
  • Safety evidence: 0 genuine prose removed and 0 unreviewed flags across 582 real replies (518 v1 and 64 v1.1).
  • Identity closeout: 0 signals, 0 example identifiers, and 0/10 stored shapes.

Next step. WP-B.1 may begin after the owner signs the A1/A2 commit. WP-B has not been started.

Git state. Everything is staged and uncommitted: product code, tests, tools, planning documents and this report. Nothing was committed, pushed or tagged.