WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.
WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
The status is returned on the done event, logged when bad, and shown in the
context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
window is unverified but the server answered, contextwindow.ensure_window
makes one bounded POST /api/generate naming only the model. It sends no
prompt, generates nothing and writes nothing. It then probes again, and the
turn is built to that answer. If the load fails, or the window is still
unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
- v1 cold turn: sent 13,875, the server read 2,050.
- Same turn after the correction: the window was verified, 3,082 sent,
3,097 read, fits, 499 tokens left beside the reply.
- Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.
WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
- a vocabulary call line;
- an echoed length hint;
- the renderer's scene line left last;
- an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
("Output only story text"). A Hard-limit-opened bracket is removed only
directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
tail is removed.
- Identity diagnostic after the correction:
- 0 identity signals;
- 0 prompt example identifiers proposed;
- 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.
Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.
Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).
One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
74 KiB
v1.1 WP-A1 and WP-A2 — Implementation and Verification Report
Work packages: WP-A1 (context-window safety reserve) and WP-A2 (protocol-echo cleanup and genre-neutral state prompting), authorised together and implemented in that order. They are reported separately throughout.
Status: COMPLETE, including the owner-requested corrective work (§R). Staged, uncommitted, for the owner's signed commit (2026-09-14). Final decisions: §R.6. WP-A1: PASS. WP-A2: PASS.
A. Repository baseline
| Branch | v1.1-development |
| Starting commit | ac465ed — Planning v4.1: record the v1.0.0 release, and plan v1.1, signed by the owner (%G? = G) |
| Its parent | 432f041, the signed v1.0.0 release commit; tag v1.0.0; main and origin/main |
| Working tree at start | clean; nothing staged, nothing unstaged, nothing untracked |
| Planning baseline | committed before any product change, so it is not in this diff |
Baseline backend suite, before any change (the untouched ac465ed tree,
python -m pytest tests/ -q, no AIDND_TEST_* set): 1,421 passed, 17 skipped,
0 failed, 865 s. This matches the v1 closeout's count exactly.
B. WP-A1 implementation
B.1 What the code did before, answered from the code
| Question | Answer on ac465ed |
|---|---|
| 1. What defines the effective context window? | contextwindow.effective_budget(configured, window), which is min(context_token_budget, window.tokens) when a window is enforceable (verified by /api/ps or /api/show, or declared through context_window_override), and the configured budget otherwise. |
| 2. Where is the fixed 64-token margin applied? | context/builder.py: OUTPUT_SAFETY_MARGIN = 64, added to max_output_tokens as output_reserve. That is subtracted with the protected sections before any knowledge or history is priced. |
3. How is max_output_tokens reserved? |
Inside that same output_reserve. |
| 4. Which components are protected? | Every system section (narrator prompt, state rule, campaign canon, always-include Canon and the knowledge rule, AI instructions, persona, plot essentials), plus the summary, retrieved memories, narrative state, author's note, front memory, length hint, refusal note and state reminder. |
| 5. Where is history trimmed? | build_context: history.window_covering, then history_floor and trim_block (§15.3), then newest-first until history_budget is spent. Knowledge is chosen before history, out of its own share. |
| 6. What counts tokens before the request? | cl100k_base from the vendored table (context/encoding.py), through builder.count_tokens. |
| 7. Which response fields give the real prompt-token count? | usage.prompt_tokens in the OpenAI-compatible response. Measured on Ollama 0.33, a stream sends no usage unless stream_options.include_usage is set. Chat and completion modes both report it once asked, and non-streaming requests always do. |
| 8. Where is that metadata persisted? | snapshot["usage"] = provider.last_usage in turns.py, inside the turn's compressed context_snapshot. Because streams were never asked, none of the 514 AI turns in the v1 evidence databases has a stored usage. |
| 9. Can the status live in provenance without a migration? | Yes. context_snapshot is a compressed JSON column, and the per-action context route returns it whole. |
| 10. What happens if protected context alone is too large? | ContextOverflow is raised before the model call. The turn route yields it as a turn error, and the dry-run route returns 422. Nothing is written. |
B.2 What the 64 tokens were, and what replaced them
M6's comment gave the margin two jobs:
- absorb "the separators between sections" that are added after the arithmetic;
- absorb "the difference between our tokenizer's count and the serving model's".
The first job is not drift. It is application text nobody priced: SEPARATOR
joins between sections, and CHAT_CONTINUE_HINT, which the provider appends to
every chat request. v1.1 prices it exactly as tokens.transport:
- one separator per system section and per story section slot
(
STORY_SECTION_SLOTS = 11), which over-counts by at most two; - plus the hint.
The second job is what the reserve now does. The reply allocation is exactly
max_output_tokens. Nothing is double-counted: the 64 tokens are gone, and
each of their two jobs has one owner.
B.3 The reserve
contextwindow.safety_reserve(effective_budget) = max(256, ceil(effective_budget × 5 / 100)).
The rounding is up, to a whole token, in integer arithmetic.
| Effective window | Reserve |
|---|---|
| 4,096 | 256, because 204.8 is below the floor |
| 5,120 | 256, exactly 5% |
| 8,192 | 410, from 409.6 |
| 16,384 | 820, from 819.2 |
| 32,768 | 1,639, from 1,638.4 |
It follows the effective budget. A 16,384 setting against a 4,096 server reserves 256. It is not a setting, and it is not calibrated per model.
Protected context is sections + transport + max_output_tokens + reserve.
ContextOverflow now names each part, and keeps the M11 advice sentence when
the server's window is what binds.
B.4 Accounting after the reply
contextwindow.classify_usage(usage, estimate, budget, max_output_tokens, window_verified).
estimate is tokens.estimate: the assembled text plus what the provider adds
(the chat hint, or the completion separator). The checks run in this order:
| Status | Condition | What it means |
|---|---|---|
unknown |
no positive integer prompt_tokens |
Nothing can be said. Never reported as fits. |
truncation_suspected |
server + reserve < estimate |
The server read fewer tokens than were sent, by more than the drift tolerance. This is the measured shape of an over-window prompt: 6,316 sent, 2,050 read. |
exceeded |
server + max_output_tokens > budget |
The drift used up the whole reserve, so the reply may be cut short. |
fits |
otherwise | The server read the whole prompt. |
The record also carries server_prompt_tokens, difference (server minus
estimate), observed_margin (budget - reply - server), safety_reserve,
window_verified and a sentence of detail.
Why exceeded is measured against the hard edge and not the reserve line.
On prompts that fit, the server counted 13 tokens more than the application,
which is chat-template overhead. A full prompt built exactly to the reserve line
would then read as "exceeded" on every turn. The reserve is the tolerance, so
using part of it is fits with a smaller observed_margin. Using all of it is
exceeded.
Surfacing.
- The record is stored as
snapshot["accounting"], so it is available for every past turn. - It is returned on the turn's
doneSSE event. exceededandtruncation_suspectedare logged at WARNING.- The context inspector shows the safety margin on every turn, and the
accounting for a turn that was sent. The two bad states are a
role="alert"notice that says the turn is kept.
The turn is kept. Accounting happens after the reply has streamed, and it changes nothing about whether the turn commits.
B.5 Files
| File | Change |
|---|---|
backend/app/contextwindow.py |
SAFETY_RESERVE_FLOOR, SAFETY_RESERVE_PERCENT, safety_reserve; the FITS / EXCEEDED / TRUNCATION_SUSPECTED statuses; classify_usage |
backend/app/context/builder.py |
OUTPUT_SAFETY_MARGIN removed; transport, the reserve, and an exact reply allocation; ContextOverflow names each part; the report gains transport, safety_reserve and estimate |
backend/app/providers/openai_compatible.py |
STREAM_OPTIONS = {"include_usage": True} on all four streaming bodies |
backend/app/routers/adventures/turns.py |
snapshot["accounting"], a WARNING log, and accounting on the done event |
frontend/src/pages/Play/panels/ContextPanel.jsx, frontend/src/styles/context.css |
the safety-margin line and the accounting notice |
backend/tools/m11_long_run.py |
the timeline records the turn's accounting from the done event |
backend/tools/v11_window_accounting.py |
new: the real-model accounting harness (§D) |
planning/TECHNICAL-DESIGN.md §15.2, DEVELOPMENT.md, README.md |
as-implemented notes. The README's stale Screenshots paragraph and test count are corrected, as the v1.1 plan asked |
No schema change, no migration, and no bundle-format change.
C. WP-A1 tests
tests/test_v11_context_reserve.py, new:
| Criterion | Tests |
|---|---|
| A1-1 reserve calculation | test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up (1,024; 4,096; 5,120; 5,121; 8,192; 16,384; 32,768), and …far_larger_than_the_v1_margin… |
| A1-2 budget enforcement | test_the_prompt_leaves_the_reply_and_the_reserve_free, on the assembled text plus the chat hint, for verified 4,096, 8,192 and 16,384, a declared 6,000, and an unverified configured 12,000, each against a 120-turn story. Also test_the_report_prices_the_text_the_provider_adds and test_the_reserve_follows_the_effective_window_not_the_setting |
| A1-3 protected overflow | test_protected_context_that_only_fits_without_the_reserve_fails_explicitly: a canon grown until v1's 64-token margin would still have built the prompt and the 256-token reserve does not. Also test_an_overflowing_turn_never_reaches_the_model: the scripted provider is called 0 times and no AI action is written |
| A1-4 no silent canon dropping | the canon sentinel is asserted present in every A1-2 configuration while the history gives way. M11's test_the_canon_at_the_front_survives_a_window_far_too_small and its negative control still pass |
| A1-5 accounting | test_a_prompt_the_server_read_in_full_fits; test_a_turn_records_what_the_server_read, end to end, including the done event |
| A1-6 unknown | test_no_usable_count_is_unknown_never_fits: None, {}, no prompt_tokens, 0, a string and a negative value. Also test_a_turn_with_no_reported_usage_is_unknown |
| A1-7 discrepancy recorded and visible | test_a_suspected_truncation_keeps_the_turn_and_says_so (snapshot, done event, WARNING log, and the per-action context route); test_a_server_that_read_far_less…; test_a_small_undercount_is_tokenizer_drift_not_truncation (reserve minus 1 is fits, reserve plus 1 is suspected); test_a_server_that_counts_more…_is_exceeded |
| A1-8 turn preserved | test_a_suspected_truncation_keeps_the_turn_and_says_so, test_an_exceeded_turn_is_also_kept |
| request | test_the_stream_asks_the_server_to_report_its_usage, in chat and completion modes |
| each take's own accounting (added after the A2 long run; §N.2 item 7) | test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves, on keep_own_slices and hand_over_the_prompt. test_a_retry_leaves_each_take_with_its_own_accounting, end to end through POST /retry: the superseded take keeps its fits and no prompt; the live take keeps the prompt and its own truncation_suspected |
The file has 35 tests: 33 at the A1 checkpoint, plus the two added above.
Frontend, panels.test.jsx, four new tests: the safety margin is shown; nothing
is shown for a turn not yet sent; truncation_suspected is an alert naming both
counts; unknown makes no claim that the prompt fitted.
One existing test changed. It is a fixture calibration, not a weakened assertion.
-
The failure.
test_history_block_trim.py::test_the_story_prompt_keeps_its_prefix_across_a_new_turnasserted that the next turn after its fixture holds the history floor. Its budget is 2,048 tokens, where the reserve is 256, or 12.5% of the window. The history budget fell from 1,186 to 969: the reserve's extra 192 over the old margin, plus 25 tokens of transport. The block is 2 on both trees. The traces:turn v1.0.0 floor v1.1 floor <60 54 54 <61 54 56 <- v1.1 steps here, v1.0.0 one turn later <62 56 56 prefix kept 0.906 <63 56 58 <64 58 58 prefix kept 0.906 -
The judgement. Same cadence and same step size, offset by one turn. §15.3's property holds, and only the fixture's choice of turn moved.
-
The change. The test now walks forward until a turn holds. It requires every move on the way to be exactly one block, and asserts the prefix on the held pair. A window that slides by one action every turn fails it exactly as before, and its negative-control companion is unchanged.
D. WP-A1 real-model evidence
Harness: tools/v11_window_accounting.py, on the A1 tree, with no source edit
during the run.
Inference: the trusted-LAN CPU reference host. It runs Ollama 0.33.0 over HTTPS, with a private CA in this machine's OS trust store and verification on.
Campaign: the v1 evidence run's own bundle (m04-final/bundle.json, 207
actions), imported so that every turn is assembled against a full window. Memory
bank and auto-summarise are off, because they would only add post-turn calls on
the same host.
Settings: configured budget 16,384, max_output_tokens 500.
Evidence: $HOME/v11-evidence/a1-accounting/.
D.1 The 4,096-window reference configuration (qwen2.5:3b-instruct, no num_ctx)
| Turn | Window | Effective budget | App estimate | Server prompt_tokens |
Difference | Reserve | Observed margin | Status |
|---|---|---|---|---|---|---|---|---|
| 1 | not verified (model not resident) | 16,384 (configured) | 13,875 | 2,050 | −11,825 | 820 | — | truncation_suspected |
| 2 | 4,096, verified (loaded) |
4,096 | 3,294 | 3,309 | +15 | 256 | 287 | fits |
| 3 | 4,096, verified | 4,096 | 3,253 | 3,268 | +15 | 256 | 328 | fits |
| 4 | 4,096, verified | 4,096 | 2,831 | 2,846 | +15 | 256 | 750 | fits |
"Observed margin" is window − reply allocation − the server's count. The v1
evidence's equivalent at the largest prompts was 23-42 tokens.
Turn 1 is the most important row in this section. The model was not
resident, so /api/ps could not report its window. M11's rule for that case
leaves the configured 16,384 standing and records the window as unverified. The
application sent 13,875 tokens. The server loaded the model at its 4,096 default,
kept 2,050 of them, and answered HTTP 200. That is the silent truncation M8 found
and M11 set out to prevent, still reachable on a cold model's first turn.
v1.0.0 would have stored this turn with no record of it. A1 recorded
truncation_suspected, logged a warning ("The server read 2,050 prompt tokens of
the 13,875 sent…") and kept the turn. §N records the cold-model gap itself as a
finding.
Turns 2-4: the window was verified and the budget capped to it. The server counted exactly 15 tokens more than the application on every turn: chat-template overhead the application cannot see. That used 15 of the 256-token reserve, and each turn kept 287 tokens or more beside the reply.
One harness note. The bundle carried the evidence campaign's memory bank
setting, so turn 1's post-turn memory pass ran and failed as the in-process
client closed (derived memory work failed). It is a failure in the harness,
outside the turn and outside what is measured. The tool was then changed to
switch memory and auto-summarise off before importing. The 16,384 run below used
the changed tool.
D.2 The 16,384-window configuration (qwen2.5:3b-instruct-16k, num_ctx 16,384 baked in)
| Turn | Window | Effective budget | App estimate | Server prompt_tokens |
Difference | Reserve | Observed margin | Status | Seconds |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 16,384, verified (parameters) |
16,384 | 13,875 | 13,890 | +15 | 820 | 1,994 | fits |
1,028 |
The first run was stopped after turn 1 at the owner's request, because the shared inference host was needed by another project. One turn took 17 minutes on the CPU host. A second run of three turns was started when the host was released (§D.3).
The window was verified from the model's own num_ctx, so even a cold model had
a ceiling. The prompt held 60 of 120 history actions. The server again counted 15
tokens more than the application. That left 1,994 tokens beside the reply,
against 23-42 in the v1 evidence at the same window and the same model.
D.3 16,384 window, second run
Started when the host was released, with the A1 code loaded at process start and before any A2 edit.
| Turn | Window | Effective budget | App estimate | Server prompt_tokens |
Difference | Reserve | Observed margin | Status | Seconds |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 16,384, verified (parameters) |
16,384 | 13,875 | 13,890 | +15 | 820 | 1,994 | fits |
1,058 |
The run was stopped deliberately after turn 1 (10:03 EDT), to free the shared host for WP-A2's real-model runs. Each 16k turn took about 17.6 minutes on the CPU host, and the two remaining turns would have delayed A2's 50-turn run by more than half an hour.
Read this row as a reproduction, not a second sample. Both 16k runs imported the same campaign, so turn 1 assembled the same prompt. The identical counts show that the measurement is repeatable. They add no new prompt shape. The prompts that grow and change across many turns are measured at 16,384 by the integrated v1.1 release run (plan §11). In this report, the prompts that change turn by turn are the 4,096 rows in §D.1 and the A2 long run in §J.
D.4 What the evidence shows
| Criterion | Result |
|---|---|
| A1's tolerance is materially larger than v1's margin | At 16,384: 1,994 tokens left beside the reply, against 23-42. At 4,096: 287 or more. The reserve is 820 and 256. |
| Real drift between the counters | +15 tokens on every turn measured, at both windows. That is 6% of the smaller reserve and 2% of the larger. |
| Disagreement is observable | Every sent turn carried a server count and a status. A real cold-model truncation was caught, recorded and logged, and the turn was kept. |
E. WP-A1 interim decision
A1 RESULT: PASS.
Recorded at the checkpoint before any WP-A2 product code was written.
- A1-1 to A1-9 pass on the A1 tree:
- backend suite 1,454 passed, 17 skipped, 0 failed (the 1,421 baseline plus 33 new);
- frontend 165 of 165;
- lint clean apart from the six warnings the v1 closeout already had.
- A1-10 passes on real-model evidence. At 4,096 there are four turns,
including a real
truncation_suspected. At 16,384 there is one verified turn, with the second run pending (§D.3). - One existing test was re-calibrated, not weakened (§C). The failure was the fixture's choice of turn. §15.3's property was unaffected, and the v1.0.0 and v1.1 traces show it.
- A1-only checkpoint saved outside the repository, so A1 can be reviewed on
its own:
a1-tracked.patch, sha2563b5d679f…;a1-untracked.tgz, sha25678275d0f….
Later corrective work on A1, found after this checkpoint. The A2 long run
exposed a defect in A1: a retried or reselected take lost its accounting, or was
shown another take's. It was fixed with two tests (§N.2 item 7). The fix is one
line in attempts.py, plus the two tests. The A1-only checkpoint patch above
predates it. The final staged tree includes it. This interim PASS stands on
A1-1 to A1-10 as written, and §O records the decision after the fix.
Carried to §N, not corrective work for A1: the cold-model first turn. When the window cannot be verified, M11 builds to the configured budget. A1 now makes the resulting truncation visible, but does not prevent it. Preventing it means a policy change: build conservatively, or refuse to send, when the window is unknown. That is outside A1's authorised scope ("do not implement model-specific dynamic calibration"; "preserve the accepted turn") and is the owner's decision.
F. WP-A2 implementation
Begun only after §E was recorded. The A1-only checkpoint was saved first, so A1 can still be reviewed on its own.
| File | Change |
|---|---|
backend/app/narrative/events.py |
vocabulary_for_prompt shows each event as the JSON object to send, not as name(field, …). _PLACEHOLDER holds the neutral placeholders. SPECS is unchanged, verified against v1.0.0 by diff. |
backend/app/narrative/extract.py |
EMIT_RULE's example now uses character-1, item-1 and location-1. LENGTH_HINT_OPENING and LENGTH_HINT_TAIL are shared with the builder. R1-R4 and RULE_* added (§H). _clean(…, after_block=). explain_removed_line added for the replay. |
backend/app/context/builder.py |
length_hint builds its opening and tail from the shared constants. The hint text is byte-identical to v1.0.0. |
backend/tools/m11_long_run.py |
The leak count also detects call lines and echoed hints. |
backend/tools/v11_replay_extractor.py |
New: the historical replay (§I). |
backend/tools/v11_compat_check.py |
New: the v1.0.0-database compatibility check (§K). |
backend/tests/test_v11_protocol_echo.py |
New: prompt, observed shapes, adversarial prose, rule attribution. |
planning/TECHNICAL-DESIGN.md §15.4, DECISIONS/013 |
As-implemented notes. |
Unchanged: no event type, field, reference rule or validation. The fence protocol, proposal recording, history replay and schema are also untouched.
Two corrections made during A2, before its evidence was taken.
- Misplaced docstring paragraphs. The new paragraphs in
extract._cleanandevents.vocabulary_for_promptwere first placed after each docstring's closing quotes, so the modules did not import. Collection failed at once and nothing ran; the paragraphs were moved into the docstrings. - The first vocabulary form cost too much.
"<key>"-style placeholders with spaced separators took the vocabulary from 258 to 456 tokens andEMIT_RULEfrom 467 to 690. That pushed protected context in an existing budget test (2,048 tokens with an 800-token reply cap) to 2,055. It also left an imported-knowledge fixture with no room to retrieve anything. Both failures were real consequences of the prompt growth, not faulty tests. The full-suite run that found them was stopped as void, and the form was changed to compact JSON with ellipsis placeholders (§G). Both tests pass unchanged on the final form, and no test was edited to accommodate it.
G. Prompt changes
G.1 Semantic before and after
| v1.0.0 | v1.1 | |
|---|---|---|
| How the vocabulary is shown | set_possession(item, owner) — gives an item to an owner, a call notation that is not the wire format |
{"type":"set_possession","item":"…","owner":"…"} — gives an item to an owner, the object itself |
| Optional fields | name(field, optional?) |
(optional: description, aliases) after the summary |
| List-valued fields | not distinguished | shown as ["…"] |
| Identifier guidance | "short lower-case slugs (mara, silver-key, old-abbey)" | "short lower-case slugs … the example's identifiers are placeholders" |
| Worked example | silver-key owned by aldric, who is at old-abbey |
item-1 owned by character-1, who is at location-1 |
| Length hint | unchanged text | unchanged text, built from shared constants |
| Reminder, canon, state rendering | unchanged | unchanged |
G.2 Cost
| v1.0.0 | v1.1 | Change | |
|---|---|---|---|
events.vocabulary_for_prompt() |
258 tokens | 378 | +120 |
EMIT_RULE (which includes the vocabulary) |
467 | 588 | +121 |
EMIT_RULE sits in the static system block, so it is priced once into the cached
prefix and into protected context on every turn. At a 4,096 window, +121 tokens is
about 3% of the window, and it comes out of the history. The cheapest option
measured was a plain name: fields list at 274 tokens. It was rejected because it
is not the wire format, and because a copy of it in prose is not a proposal the
extractor can recognise and remove. A copied JSON line is.
G.3 Guard against reintroduction
test_v11_protocol_echo.py covers the following:
- Fixture identifiers. No identifier from either acceptance fixture, and no
genre noun, appears in
EMIT_RULE,EMIT_REMINDER, the vocabulary, or any length hint. - The worked example. It is a block the extractor accepts, with allowed event types.
- Call notation. No vocabulary line is a call.
- Line shape. Every line parses as the object, with exactly the required fields.
- Hint wording. The length hint carries the constants the extractor recognises.
H. Cleanup rules (WP-A2)
H.0 The evidence the rules were designed from
Before any extractor change, the whole v1 corpus was scanned. That covers every
AI turn in every evidence database under $HOME/m11-evidence, plus the exported
bundles. Stored text and stored raw replies were counted separately, and
duplicates were removed by content hash. The scan counted the candidate shapes,
and the near-misses a rule must not touch.
| Shape | Unique stored texts | Where |
|---|---|---|
| A line opening with a vocabulary call | 4 | All four closeout identity turns (depths 10, 12, 14 and 20). |
| A vocabulary call anywhere | 4 | The same four lines. The corpus has no call in the middle of a sentence. |
[Hard limit anywhere |
3 | See the three cases below. |
A rendered Scene: … (at …) as the last line |
1 | Closeout browser-run3 action 3, stored by the v1.0.0 product tree. |
Scene: lines anywhere |
62 | Mostly pasted state sections from runs taken before M11's 0c7316f fix. The v1.0.0 extractor already handles those, and the replay (§I) starts from the raw reply. |
The three [Hard limit cases:
- Identity depth 14: the application's length hint, reworded at the front.
- m04-rerun action 151: the hint's own words, with the last sentence
replaced by "This story ends here." An empty dangling
```jsonfence follows it. - browser-recheck2 action 5: "[Hard limit: I have adhered to the prescribed limit of words. Here's the narrated turn and the state update. Send your next action.]" The model invented this; it is not application text.
H.1 The rules
Each rule is anchored to something the application owns: the event vocabulary, the length hint's wording, or the renderer's scene-line format. None of them judges prose.
R1 — event-call line (RULE_EVENT_CALL)
| Source | events.vocabulary_for_prompt printed every event as name(field, …), and the narrator copied the notation. |
| Matcher | A whole line that starts with a name in events.SPECS (case-insensitive) followed by (. Indentation and a leading > are allowed before the name, and spaces before the (. The whole line is removed, including anything after the call on that line, such as depth 20's "Adds the silver key…". Lines inside a fenced code block are never examined. |
| Rationale | create_entity, set_possession and the rest are this protocol's snake_case identifiers. A story line that starts with one of them followed by ( is protocol. |
| False-positive defence | Not applied in the middle of a sentence (the engineer's whiteboard). Not applied to call-shaped lines whose name is not in the vocabulary (> open_door(north)). Not applied inside a story's own code block. Not applied to a vocabulary word with no ( ("Create_entity is a terrible name"). |
R2 — echoed length hint (RULE_LENGTH_HINT)
| Source | builder.length_hint. Its opening [Hard limit: and its closing sentence are now named constants shared with the extractor (LENGTH_HINT_OPENING, LENGTH_HINT_TAIL). |
| Matcher | A bracket at the end of the reply, closed or cut off; the existing trailing loop already isolates it. Its contents must start with Hard limit: and carry the application's own wording: either the tail phrase "append the state block", or "turn must not exceed N words". The check goes in _is_echoed_instruction, beside the reminder and continue-hint checks M11 already has. |
| Rationale | Both phrases are the application's instruction text. The 3B narrator reworded the front ("your next turn") and kept the rest. |
| False-positive defence | Only a whole bracket at the end of the reply is removed. These all stay: "[Hard limit: 500 words]" read aloud mid-sentence; an in-world "[Hard limit of the reactor…]"; a final "[Hard limit: forty days, no extensions]". The model's invented bracket in browser-recheck2 action 5 also stays, because it carries neither phrase. That is a deliberate miss. |
R3 — rendered scene line (RULE_SCENE_LINE)
| Source | The first line of render.for_prompt: Scene: <summary>, plus (at <location>) when the scene has a location. |
| Matcher | After the other trailing cuts, the reply's last line starts with Scene: and meets one of two conditions. Either it ends in the renderer's (at <location>), or protocol was also cut from the end of the same reply (a block, a hint, a call line or a dangling fence). |
| Rationale | The (at …) suffix is the renderer's syntax. Without it, a lone scene line directly above removed protocol belongs to the same pasted tail. |
| False-positive defence | Never removed mid-reply: "Scene: a kitchen, late." followed by story stays, and so does M11's existing "Scene: the docks at dawn." case. A screenplay-style last line, "Scene: take two, and nobody moves.", stays when nothing else was removed. Accepted risk: a story whose last line exactly imitates Scene: … (at …) loses that line. |
R4 — empty dangling fence (RULE_EMPTY_FENCE)
| Source | A reply that hit the output limit after opening ```json and before writing anything (m04-rerun action 151). |
| Matcher | A ```json or bare ``` opener as the reply's last line, with nothing after it. |
| Rationale | M11's _is_opening_of_proposal already cuts a dangling fence that stopped after {, on the principle that "a story's own code block is not that short, and one that is holds nothing to lose". An empty opener holds even less. |
| False-positive defence | Applies only when nothing at all follows the opener. A dangling fence with any content keeps M11's existing judgement. |
H.2 What is deliberately not removed
- A fact restated inside a sentence (§P risk 4). Removing it would mean judging prose.
- Model-invented headings that are not the renderer's (§P risk 6).
- Model-invented brackets that carry no application phrasing (browser-recheck2 action 5).
- Prose that follows a removed call line, even when it narrates the call ("You create a new person, Mike, who has", identity depth 12). It is narration, however poor.
I. Historical replay
Tool: tools/v11_replay_extractor.py, on the final A2 extractor.
Inputs: every AI turn in the 17 evidence databases under $HOME/m11-evidence,
plus the closeout identity run's bundle. A turn is replayed from its stored
raw_output where one exists, so the comparison starts from what the narrator
actually sent. The identity bundle carries only stored text; for those turns the
v1.0.0 prose is the input itself.
Comparison: each input goes through the extractor as tagged at v1.0.0,
read with git show, and through the v1.1 extractor. Lines are compared with
trailing whitespace ignored.
Evidence: $HOME/v11-evidence/a2-replay/replay.json (sha256 f04c54ff…) and
replay.md, which holds the full v1.0.0 and v1.1 prose for every changed turn.
| Unique replies replayed | 518 (6 exact duplicates skipped) |
| Unchanged | 509 |
| Changed | 9 |
| Flagged (an unexplained removal, text added or rewritten, or more than 25% removed) | 0 |
| Removed lines by rule | R3 scene line 6, R1 event call 4, R2 length hint 2, R4 empty fence 1 |
The planning pass estimated about 453 turns; the actual count is 518. The difference is the trial, browser and recheck databases, which the v1 figure of 443 did not count.
I.1 Every changed turn
| # | Source | Removed | Rule | Share of prose removed | Review |
|---|---|---|---|---|---|
| 1 | browser-final action 3 (raw) | Scene: Aldric and Mara by the fire in the Crooked Lantern (at The Crooked Lantern). |
R3 | 13.1% | The rendered scene line, directly above a ```json proposal in the raw reply. Protocol. |
| 2 | closeout browser-run3 action 3 (raw), stored by the v1.0.0 product tree | Scene: Aldric and Mara by the fire in The Crooked Lantern. (at The Crooked Lantern) |
R3 | 20.7% | The same, in the renderer's exact form. Protocol. |
| 3 | m04-rerun action 151 (raw) | Scene: Aldric and Mara in the Crooked Lantern, rain outside.; [Hard limit: this turn must not exceed 180 words, … This story ends here.]; ```json |
R3, R2, R4 | 11.9% | A pasted scene line, the application's length hint with its last sentence reworded, and a proposal opener the output limit cut off. Protocol. |
| 4 | m04-rerun action 153 (raw) | Scene: Aldric and Mara in the Crooked Lantern, rain outside. |
R3 | 3.1% | The scene line directly above a ```json proposal. One trailing space was also trimmed from the story line before it; no word changed. Protocol. |
| 5 | m04-rerun action 157 (raw) | Scene: Aldric and Mara in the Crooked Lantern, rain outside. |
R3 | 2.8% | The scene line directly above ```json\n{. Protocol. |
| 6 | identity depth 10 (stored) | > Create_entity(new_person, "john", "character", "A determined team member …", ["jo… |
R1 | 6.4% | A call cut off mid-argument by the output limit. Protocol. |
| 7 | identity depth 12 (stored) | > Create_entity(new_person, "mike", "character", "A new team member …", ["mike"]) |
R1 | 4.8% | The call. The narration after it ("You create a new person, Mike, who has") is kept. Protocol. |
| 8 | identity depth 14 (stored) | > Create_entity(mike, …); Scene: Bill, Alice, Roger, John, and Mike at the table; John not yet arrived. (at The meeting room); [Hard limit: your next turn must not exceed 180 words, … append the state block well inside the limit.] |
R1, R3, R2 | 23.4% | The whole pasted tail: the call, the rendered scene line, and the reworded length hint. Protocol. |
| 9 | identity depth 20 (stored) | > set_possession(silver-key, "alice") Adds the silver key to Alice's possession. |
R1 | 4.1% | The call, with the fantasy example slug and a gloss on the same line. Protocol. |
Genuine prose lost: none. Each of the 13 removed lines is application protocol or instruction text: a vocabulary call, the length hint, the renderer's scene line, or an empty fence opener. On a line that also carried narration, only depth 20's same-line gloss went with the call; that gloss describes the call, not the story.
A replay-tool defect, found and fixed before this result. The first replay flagged turn 4 as "text rewritten". The trailing space trimmed from the story line was compared as a changed line. The tool now ignores trailing whitespace and says why. The flag was the tool's, and no rule changed as a result.
J. WP-A2 real-model evidence
Inference: the trusted-LAN CPU reference host. It runs Ollama 0.33.0 over
HTTPS with a private CA and verification on, and serves qwen2.5:3b-instruct at
a verified 4,096 window. Embeddings are nomic-embed-text. The storyteller was
on loopback.
Tree: the final A2 extractor and prompt. The full backend suite passed on this tree before the long run was allowed to start. The launcher enforced that gate.
J.1 Multi-character identity diagnostic
The office meeting fixture (TEST-CAMPAIGN-FIXTURE.md Appendix A), 10 beats,
memory off. Evidence: $HOME/v11-evidence/a2-identity.
| v1 closeout (§S.5 of the M11 report) | v1.1 | |
|---|---|---|
| Turns accepted | 10 of 10 | 10 of 10 |
| Identity signals | 0 | 0; verdict "no objective identity defect detected" |
| Proposals naming an example identifier from the fixed prompt | 1 (silver-key given to Alice) |
0. No silver-key, aldric, mara or old-abbey, and none of the new placeholders either |
| Proposal outcomes | 1 applied, 5 without a block, 3 unparseable, 1 refused | 10 accepted, 1 rejected |
| Stored turns with a shape R1-R4 targets | 4 of 10 (depths 10, 12, 14, 20) | 0 of 10 |
| Stored turns with any echoed application instruction | 4 of 10 | 1 of 10: depth 16, below |
A leak A2 does not remove. Depth 16 kept this tail at the end of the story:
Scene: Bill, Alice and Roger at the table; John not yet arrived.
[Hard limit: this is now 180 words.]
[Reminder: end your reply with a `state` block listing the events your narration made true, with absolute values.]
[You don't need to continue; your turn must now be about John entering the room. Continue the story here, directly. Output only story text.]
The trailing cleanup works from the end upwards, and the last bracket is not recognised: it is a reworded continue hint, and the recogniser matches that hint only by its opening words. So the reminder above it is never at the end, the reworded length hint carries none of the application's wording, and the scene line is never last. This is the conservative design behaving as documented. It is still a real leak, on 1 of 10 turns. The harness's leak counter did not see it: it looks for proposals, pasted headings, calls and the hint's own wording, not reminders. It was found by reading the stored narration. A narrow follow-up is in §N.
J.2 The 50-turn long run
tools/m11_long_run.py --turns 50: the fantasy Continuity Test with the full
schedule of history operations and the memory bank on. Evidence:
$HOME/v11-evidence/a2-long-run.
| Status | complete; no aborted reason, no failed reason |
| Accepted turns | 51: 50 scheduled and the recall turn |
| Wall clock | 9,324 s. A turn took 47-243 s, median 165 |
| Genuine process restarts | 3, each compared identical: true |
| History operations performed | 2 Save Points, Undo/Redo, 2 retries, undo then divergence, take selection, a Save Point restore, and an induced failed call (HTTP 404; state unchanged) |
| Post-turn failures | 0. 0 tracebacks and 0 database is locked in server.log |
| Memories and summaries written | 15 memories (9 in the bank at the end); 5 summaries |
| Protocol leakage | 0 of 54 stored AI turns, by the extended harness count (headings, "events", call lines, the echoed hint). A separate scan of every stored turn for reminders, [Hard limit, continue-hint wording, calls, a final Scene: and "events" also found 0 |
| Parser failures | 6 proposals unparseable, out of 56 recorded |
| Refused proposals | 0 |
| State updates from narration | 4 events accepted from story (50 proposals accepted, most with an empty events list); 9 more from the harness's manual corrections |
| Legitimate prose removed | none. All 64 raw replies from this run and the identity run were replayed through the v1.0.0 and v1.1 extractors. One changed: a trailing ```json opener with nothing after it (R4, 0.4% of that reply). 0 flagged |
| M04-style recall | recovered_through_state_only, which is WP-B's subject and unchanged from v1 |
A1 accounting across the run.
| Window | verified at 4,096 on 51 of 51 turns |
| Status | fits on 51 of 51 |
| Server count minus the application's estimate | +15 on every turn |
| Largest server prompt | 3,321 tokens |
| Observed margin beside the reply | 275 minimum, 620 median, 2,297 maximum, against a 256 reserve |
A defect the long run exposed, fixed in A1 (§N.2 item 7). Three of the 54
stored AI turns had no accounting: the two retried takes and the take superseded
by a take selection. attempts.ATTEMPT_KEYS lists the snapshot slices that
belong to one attempt, and A1 had not added accounting to it. When the live
flag moved, the new live take was handed the old take's accounting with the
shared prompt, and its own was discarded. The long run's numbers above come from
the done events and are unaffected. It was the stored accounting of a
retried turn that could be wrong.
J.3 Genre fixtures (A2-8)
| Fixture | How | Result |
|---|---|---|
| Fantasy Continuity Test | the 50-turn long run, real model | 0 example identifiers proposed. The fixture's own silver-key is legitimately its entity |
| Multi-Character Identity Test (office) | the identity diagnostic, real model | 0 proposals naming any prompt example identifier (v1: 1) |
| Persephone science fiction | test_m11_scifi.py, 10 tests, scripted |
passes; no genre noun in any fixed instruction (test_v11_protocol_echo.py) |
Not run: science fiction against a real narrator. No harness drives that fixture with a real model, and adding one was out of A2's scope.
J.4 What this does not replace
This is not the integrated v1.1 release run (plan §11). That run covers 100 turns, the 16,384 window, and the GPU host with the owner's logging.
K. Compatibility
K.1 A real v1.0.0 database opens unchanged
Source database: m04-final/campaign.db, the v1 evidence run's own campaign.
It was written by 96c1bf5, whose product code is identical to v1.0.0, and holds
user_version 94. Its contents:
- 207 actions on 4 branches, with retained history;
- 2 Save Points and 32 state events;
- 33 memories and 12 summaries;
- 3 imported knowledge sources, one campaign narration-length choice, and settings.
Method: tools/v11_compat_check.py. The database was copied, and each copy
opened by one tree:
v100: thev1.0.0tag, from a scratch worktree;v11-readonly: the v1.1 tree.
Each tree's app was verified by the path it was imported from. The same read-only snapshot was taken through the API for both:
- the full export bundle (which carries no timestamp of its own);
- narrative state and state events;
- Save Points, knowledge, memories and derived status;
- the newest actions and settings;
- the schema and
user_versionbefore and after opening.
| Comparison | Result |
|---|---|
| Schema before opening, v1.0.0 against v1.1 | identical |
| Schema after opening | identical; user_version 94 → 94 on both, so no migration ran |
| API snapshot | identical, field for field |
K.2 The same database, exercised on v1.1
--exercise, on a third copy, with the endpoint first pointed at a refused
loopback port so nothing left the machine:
| Operation | Result |
|---|---|
| Undo | 119 → 117 moments; Redo becomes available |
| Redo | back to 119; Redo no longer available |
| Restore Save Point On the ridge | 69 moments, with Redo available into the later story |
| Context dry run | HTTP 200. One imported passage retrieved. The campaign canon and narrative state sections are present. The summary that applies (moments 49-64) is included. Budget 16,384: 14,577 prompt + 28 transport + 500 reply + 820 reserve = 15,925. The window is recorded as unverified, because the server was unreachable by design. |
| Export, then import into the same install | 207 → 207 actions, the same head depth, Save Points 2 → 2, memories 33 → 33, narrative state identical |
Not exercised here, and why. Writing a new memory or summary needs a model and an embedding model. The real-model long run in §J writes both on the v1.1 tree. Memory retrieval with embeddings is covered the same way. No data-migration path was needed, and none was added.
L. Security and offline
L.1 What the diff adds, measured against v1.0.0
git diff v1.0.0 over backend/app and frontend/src, searched for any new
network or remote-resource use. The search covered httpx, requests, urllib,
socket, aiohttp, literal URLs, fetch(, XMLHttpRequest and
tiktoken.get_encoding.
| Check | Result |
|---|---|
| New network client, destination or URL | none |
| New imports in product code | math, json, logging, and one internal constant (providers.openai_compatible.CHAT_CONTINUE_HINT) |
| Remote or cloud token counting | none. The count is the vendored cl100k_base, as before. The server's own count arrives in the response the turn already receives |
| Calibration service | none. Calibration was not built |
| The one request-level change | stream_options: {"include_usage": true} in the body of the request already sent to the already-approved endpoint. It adds no destination and no data about the user |
Endpoint policy (endpoints.py, ADR 011) |
unchanged; the file is not in the diff |
Probe policy (contextwindow._ask) |
unchanged; A1 added pure functions beside it |
| Storyteller bind | unchanged; docker-compose.yml, Dockerfile and the start scripts are not in the diff |
L.2 Regression runs
-
The endpoint-policy and local-only tests pass in the backend suite (§M):
test_m11_security.py(H-series, including a public endpoint refused at request time),test_egress.py,test_local_only_surface.py(loopback publication),test_m11_context_window.py(the probe held to the policy). -
Trusted-LAN inference still works. Every real-model turn in §D and §J went to the trusted-LAN host over HTTPS, with a private CA and verification on.
-
The offline container passed 23 of 23. It ran
tools/m11_offline.pyon the final tree, with--network none, a fresh volume, and an image built withdocker build --no-cache. The checks are the same 23 as the v1 closeout:- no route and no DNS;
- first page load, no remote origin, a CSP, and every asset local;
- campaign creation, state extraction, knowledge import, prompt assembly and retrieval;
- a turn with no model reachable is reported, with no narration accepted, the player's words kept, state unchanged and the earlier story intact;
- export and import with no secret;
- the media module inert;
- campaigns survive a container restart.
The evidence is
$HOME/v11-evidence/offline/offline-report.json. The A1 accounting path is exercised offline by the failed-turn check: a turn that never reaches a model writes no accounting, because it commits nothing.
M. Full test results
Every backend command below is .venv/bin/python -m pytest tests/ -q from
backend/, with no AIDND_TEST_* set. The 17 skips are the tests that need a
real local model, the same 17 as the v1.0.0 closeout. No new skip was
introduced. The three real-model skips in the targeted runs are those same
tests.
| Run | Tree | Result |
|---|---|---|
| Backend, baseline | untouched ac465ed |
1,421 passed, 17 skipped, 0 failed, 866 s |
| Backend, A1 checkpoint | A1 only; the not-yet-implemented A2 test file excluded | 1,454 passed, 17 skipped, 0 failed, 835 s |
| Backend, A2 first full run | A2 with the first vocabulary form | stopped and discarded. Two budget failures, §F |
| Backend, A1 + A2 | before the take-accounting fix | 1,505 passed, 17 skipped, 0 failed, 786 s |
| Backend, final | A1 + A2 + the take-accounting fix: the staged tree | 1,507 passed, 17 skipped, 0 failed, 792 s |
| New tests | test_v11_context_reserve.py (35) and test_v11_protocol_echo.py (51) |
all pass |
| Frontend | A1 (A2 changes no frontend file) | 165 of 165, 14 files |
| Lint | A1 | exit 0, six only-export-components warnings, the same six as the v1 closeout |
| Production build | final | clean. 16 files, index-rWgm40jJ.js 394.47 kB, index-BXtb0iME.css 48.09 kB |
Offline container, --network none, docker build --no-cache |
A1 + A2, before the take-accounting fix | 23 passed, 0 failed (§L.2) |
| Offline container, repeated | final, after the take-accounting fix | 23 passed, 0 failed, from a fresh --no-cache image ($HOME/v11-evidence/offline-final) |
| Identifier scan of every changed and new file | final | no real hostname, address, domain or personal name |
N. Findings
N.1 Blockers
None found.
N.2 Corrective work, done within the packages
- A1:
stream_optionswas required, not optional. The provider's docstring saidstream_optionswas "deprecated and does nothing", which is true of OpenRouter. Against Ollama 0.33 a stream sends nousagewithout it, and none of the 514 v1 evidence turns stored a count. Without the fix, A1's accounting would have readunknownon every real turn. - A1: one test's fixture re-calibrated (§C). The history-floor prefix test had assumed which turn holds the floor. Behaviour was unchanged.
- A2: the first vocabulary form was too expensive (§F, §G). It grew protected context by 223 tokens and failed two existing budget tests. It was replaced by compact JSON (+121); the tests pass unchanged.
- A2: module syntax. Two docstring paragraphs were misplaced. Collection failed immediately, and they were corrected before any evidence was taken.
- Harness: the replay tool compared trailing whitespace (§I), which produced one false "rewritten" flag. Fixed.
- Harness: the accounting tool imported the bundle's memory setting (§D.1). That caused failing post-turn calls on the shared host. Fixed before the 16k runs.
- A1: a retried or reselected take lost its accounting, or showed another
take's.
- Where it was found: after the A2 long run. Three of 54 stored AI turns
had no
accounting, and they were exactly the two retries and the take selection. - Cause:
attempts.ATTEMPT_KEYSlists the snapshot slices that belong to one attempt, and A1 had not addedaccountingto it. It was therefore treated as part of the shared prompt.hand_over_the_promptgave the new live take the old take's accounting, and discarded its own. - Fix:
accountingadded toATTEMPT_KEYS, with a unit test and an end-to-end retry test (§C). - Exports: none needed a change. A snapshot travels whole.
- Migrations: none. v1.0.0 turns have no accounting to move, and
migrations._ATTEMPT_KEYSis a frozen historical copy, left as it is. - Evidence: the long run's accounting figures in §J come from the
doneevent of each turn and are unaffected. The full backend suite and the offline container were re-run after the fix (§M).
- Where it was found: after the A2 long run. Three of 54 stored AI turns
had no
N.3 Non-blocking residual risks
-
Cold-model first turn: truncation is detected, not prevented (§D.1). When
/api/pscannot see the model and/api/showfinds nonum_ctx, M11 builds to the configured budget. On a server whose real window is smaller, the first turn is silently cut. It was observed for real: 13,875 sent, 2,050 read. A1 now records it, logs it and shows it. Preventing it needs a policy decision, which is the owner's and outside A1's scope. The two options are to build conservatively when the window is unknown, or to warm the model and probe again before the first turn. -
The drift measured is one model's. +15 tokens against
cl100k_base, forqwen2.5:3b-instructin both window sizes. A model whose tokenizer diverges by more than 5% will show up asexceeded, not as silent truncation. The reserve is a tolerance, not a guarantee. -
exceededandtruncation_suspecteddepend on the server reporting usage. A server that ignoresinclude_usageproducesunknownon every turn. That is honest, but it detects nothing. -
R3's accepted false-positive risk. A story whose last line exactly imitates the renderer's
Scene: … (at …)loses that line (§H). -
Protocol still left in stored narration, deliberately:
- model-invented brackets that carry no application wording;
- model-invented headings;
- facts restated inside a sentence.
One real v1.1 turn shows the cost (§J.1, identity depth 16). A reworded continue hint was the last line of the reply, so nothing above it became trailing. The application's reminder, a reworded length hint and a scene line all stayed.
Narrow follow-up, not done here. Recognise
CHAT_CONTINUE_HINTby its distinctive application phrase "Output only story text", not only by its opening words. The reminder above it would then come off by the existing rule. The reworded hint, with none of the application's wording, and the scene line would still stay.Why it was not done. It would have been an extractor change after the long-run evidence was taken, and it needs its own replay. This is the owner's decision.
-
A2 costs 121 tokens of protected context (§G.2). At 4,096 that is about 3% of the window, and it comes out of history.
-
A2's real-model run is on the CPU host at 4,096, not the v1 evidence configuration. v1 used 16,384 on the GPU host. On the CPU host a full 16k prompt took 17 minutes a turn, and a GPU-host long run requires the owner's power, link and kernel logging to be started first (
gpu-host-logging-before-long-runs). The integrated v1.1 release run (plan §11) is where 16,384 is re-established.
N.4 Future ideas
- A cold-model policy for the first turn, as in N.3 item 1.
- Summarising the accounting across a campaign, for example "3 of 120 turns suspected". Today the inspector shows it one turn at a time.
- Showing the accounting status as a chip under the narration, not only in the context inspector. This is browser work, for WP-C.
O. Final A1 decision
A1 RESULT: PASS.
The final decision is taken on the staged tree, including the take-accounting fix found after the interim checkpoint (§N.2 item 7).
| Criterion | Result | Evidence |
|---|---|---|
| A1-1 Reserve calculation | PASS | max(256, ceil(5%)): 256 at 4,096, 410 at 8,192, 820 at 16,384; 7 windows tested (§B.3, §C) |
| A1-2 Budget enforcement | PASS | assembled text, provider hint, reply and reserve are within the effective window in 5 configurations (§C) |
| A1-3 Protected overflow | PASS | fails before the model call; the provider is called 0 times and nothing is written (§C) |
| A1-4 No silent canon dropping | PASS | the canon sentinel is present in every configuration while history gives way (§C) |
| A1-5 Post-response accounting | PASS | real: 51 of 51 long-run turns, 4 turns at 4,096 and 2 at 16,384 carry a server count and a status (§D, §J) |
| A1-6 Unknown accounting | PASS | no usable count gives unknown, never fits (§C) |
| A1-7 Unexpected discrepancy visible | PASS | a real truncation_suspected on a cold model: 13,875 sent, 2,050 read. It was stored, logged, sent on the done event and shown in the inspector (§D.1) |
| A1-8 Accepted turn preserved | PASS | the turn is kept for truncation_suspected and for exceeded; after a retry, each take keeps its own accounting (§C) |
| A1-9 Existing behaviour | PASS | 1,507 passed, 0 failed. One fixture re-calibrated, not weakened (§C) |
| A1-10 Real-model verification | PASS | at 4,096, 275-2,297 tokens left beside the reply across 55 turns; at 16,384, 1,994. Drift is +15 on every turn, against v1's 23-42 tokens of headroom (§D, §J) |
Carried as residual risk, not failure: the cold-model first turn is detected, not prevented. Preventing it is a policy decision for the owner (§N.3 item 1).
P. Final A2 decision
A2 RESULT: PASS, with one recorded residual.
| Criterion | Result | Evidence |
|---|---|---|
| A2-1 Genre neutrality | PASS | no fantasy or science-fiction identifier and no genre noun appears in any fixed instruction; a test guards against reintroduction (§G.3) |
| A2-2 Protocol clarity | PASS | the vocabulary is the wire-format object, not call notation. +121 tokens, the smallest wire-format form measured (§G) |
| A2-3 Observed leaks | PASS | all four v1 shapes are removed (§H). What is deliberately left is documented, including one real v1.1 occurrence (§J.1, §N.3 item 5) |
| A2-4 Historical corpus safety | PASS | 518 real replies replayed, 9 changed, 0 flagged, no story prose removed. A further 64 replies from the v1.1 runs gave 1 changed, 0 flagged (§I, §J.2) |
| A2-5 Adversarial prose | PASS | the owner's four sentences and nine more survive untouched (§C of the A2 tests, test_v11_protocol_echo.py) |
| A2-6 State extraction | PASS | a valid block still parses and applies when protocol litter surrounds it. SPECS, render.py and validate.py are identical to v1.0.0 (§F, §K) |
| A2-7 Invalid proposals | PASS | the validator is byte-identical and the existing refusal tests pass; the long run recorded 0 refusals and 6 unparseable proposals (§J.2) |
| A2-8 Genre fixtures | PASS | fantasy by real model (51 turns); office identity by real model (10 turns, 0 example identifiers against v1's 1); science fiction scripted only, which is stated (§J.3) |
| A2-9 Real-model sample | PASS | 51 accepted turns: 0 of 54 stored turns carry protocol, 6 parser failures, 0 refused, 4 story state events, no prose removed, and A1 accounting fits 51 of 51 (§J.2) |
The residual. In the identity run, 1 of 10 stored turns kept a trailing tail of echoed instructions: a reworded continue hint, the reminder, a reworded length hint and a scene line. The conservative design cannot see past the reworded last line. §N.3 item 5 records a narrow follow-up, which is the owner's decision.
Q. Recommendation
- A1 is ready for acceptance. It meets every criterion on real evidence, including a real silent truncation it caught.
- A2 is ready for acceptance, with the residual in §P noted. The owner may want the narrow continue-hint follow-up (§N.3 item 5) first. It would be a small extractor change with its own replay. It does not block acceptance of what A2 set out to do.
- Owner decisions this report raises:
- Accept or correct A1.
- Accept A2, or ask for the continue-hint follow-up first.
- Whether the cold-model first turn (§N.3 item 1) needs a prevention policy, and where that policy belongs: an A1 follow-up, or a later package.
- WP-B may begin once A1 and A2 are accepted. Nothing in either package blocks B's diagnosis. A1's accounting and A2's cleaner prose are both in place for B's long runs. WP-B has not been started.
Git state: everything is staged and uncommitted for the owner's signed commit. Nothing was committed, pushed or tagged. A suggested commit message was prepared outside the repository.
R. Corrective-work addendum (owner review, 2026-09-14)
The owner accepted A1 and A2 in principle, and asked for two corrective items to be closed before the commit. §O-§Q above record the decisions as they stood before this addendum. §R.6 supersedes them.
R.1 WP-A1: preventing the cold-model truncation
The case, as A1 caught it (§D.1, turn 1):
/api/psknew nothing, because the model was not resident./api/showfound nonum_ctx, so the window was unverified.- The prompt was built to the configured 16,384.
- Ollama loaded the model at its 4,096 default, read 2,050 of 13,875 tokens, and answered HTTP 200.
Mechanism, measured before it was built. This was measured on Ollama 0.33
on the trusted-LAN host, with the model first unloaded (keep_alive: 0) and
/api/ps empty.
| Step | Observed |
|---|---|
POST /api/generate {"model": "qwen2.5:3b-instruct"}, with no prompt |
HTTP 200 in 3.7 s, "response": "", "done_reason": "load". Nothing was generated. |
/api/ps afterwards |
the model is listed with context_length 4,096 |
| An OpenAI-compatible streamed request afterwards | no reload: /api/ps still reads 4,096, and prompt_tokens was the estimate plus 13 |
What was built.
contextwindow.Windowgainsreachable: the server answered a discovery request, whatever it said.contextwindow.warm(endpoint, model, timeout)sends onePOST {native base}/api/generatewith the body{"model": model}. There is no prompt, nooptionsand nokeep_alive. It goes throughendpoints.rejection_reasonandtlstrust.ssl_context(), and it never raises.contextwindow.ensure_window(endpoint, model, declared, warm_timeout)does the following:- It probes. A verified window is returned untouched.
- If the window is unverified, the server was reachable, and a model is configured, it warms once.
- If the load succeeds, it probes again with
use_cache=False. - It returns the window and a
preflightrecord (attempted,loaded,verified_before,verified_after,detail).
turns._generate_turncallsensure_windowin place ofprobe, with the settings' model timeout as the bound. It storessnapshot["window"]["preflight"]. The context dry run still only probes: loading a model from a read-only panel would be a side effect.
What is unchanged.
- A window still unverified leaves the configured budget standing.
- There is no guessed window, no hard-coded 4,096, and no retry loop.
- The accounting still classifies the reply.
The constraints the owner set, each checked:
- no calibration and no user setting;
- no endpoint-policy change;
- only the configured endpoint is contacted;
- no dummy narration is generated or accepted;
- no accepted turn is discarded.
Tests: tests/test_v11_cold_window.py, 15 tests.
| Required | Test |
|---|---|
| cold model, then preflight, then a verified smaller window, then the real turn built to it | test_a_cold_model_is_loaded_once_and_its_window_verified (request order /api/ps, /api/show, /api/generate, /api/ps; the load body exactly {"model": …}). test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports, end to end: budget 4,096, verified, the canon sentinel present, the assembled text plus transport, reply and 256-token reserve within 4,096 |
| negative control | test_without_the_load_the_same_cold_turn_would_have_been_built_too_large: probing alone gives an unverified window, budget 16,384, and a prompt more than twice 4,096 |
| already loaded: no warm | test_an_already_loaded_model_is_not_warmed |
| load succeeds, verification still unavailable | test_a_model_that_loads_but_still_cannot_be_read_stays_unverified: one load, unverified, the configured budget kept |
| load fails | test_a_failed_load_is_recorded_and_leaves_the_window_unverified (HTTP 404 and 500; bounded, with no second probe). test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe: the error is reported, with no AI action, state event or proposal. test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer: accounting unknown |
| the load writes nothing | test_the_load_itself_writes_nothing: actions, state events, proposals, memories and summaries are counted before and after, and are identical |
| no new destination | test_the_load_request_obeys_the_endpoint_policy (public addresses refused with no transport installed). test_the_load_request_goes_only_to_the_configured_host. test_an_unreachable_server_is_not_asked_to_load_anything |
| also | test_a_declared_window_does_not_stop_the_server_being_asked; test_the_context_dry_run_never_loads_a_model |
Real cold-model test: §R.3.
R.2 WP-A2: the continue-hint instruction tail
The exact output. Identity depth 16 (action 17 in the run's database). The stored text equals the v1.1 extractor's output of the raw reply, and the raw reply has no state block. It ends:
Scene: Bill, Alice and Roger at the table; John not yet arrived.
[Hard limit: this is now 180 words.]
[Reminder: end your reply with a `state` block listing the events your narration made true, with absolute values.]
[You don't need to continue; your turn must now be about John entering the room. Continue the story here, directly. Output only story text.]
Why the application-owned material above the last bracket was not removed.
_clean cuts the end of a reply one bracket at a time, and only ever examines
the last bracket. That last bracket is the provider's CHAT_CONTINUE_HINT,
reworded at the front. _is_echoed_instruction recognised that hint only by its
opening words ("continue the story directly"), so it returned false, and the
loop stopped. The reminder above it would have been removed by the existing
Reminder: rule, but it was never at the end. The reworded length hint carries
none of the application's wording, and the scene line was never last.
The rule, R5 (RULE_INSTRUCTION_TAIL). Two narrow changes:
- Recognise the continue hint by its own sentence. A trailing bracket
containing "Output only story text" (
CONTINUE_HINT_PHRASE) is an echoed instruction. A test pins the phrase toCHAT_CONTINUE_HINT, so the two cannot drift. - Remove a hint-opened bracket only directly above an echo. A trailing
bracket opening with
Hard limit:is removed only when an echoed instruction bracket has already been cut from the end of the same reply.
The rest follows from existing rules. The reminder is removed by the Reminder:
rule once it is last. The scene line is removed by R3 once it is last and
protocol has been cut.
False-positive defence. Each of these is a test that must pass unchanged:
| Case | Why it stays |
|---|---|
A final [Hard limit: forty days, no extensions] |
no echo was cut below it |
| The same bracket above a state block | a block is not an echoed instruction, so a block below does not license it |
[The sign on the door reads: Closed] directly above an echoed continue hint |
it does not open the way the hint opens; only the echo goes |
| "output only story text" inside a sentence mid-story | only a trailing bracket is examined |
A final [To be continued] |
no application phrase |
The owner's four original adversarial sentences and all earlier preservation cases also still pass.
Accepted risk. A story-world bracket that opens exactly [Hard limit: and
sits directly above a parroted instruction would be removed with it.
Positive regression: test_the_depth_sixteen_instruction_tail_leaves_entirely.
The two story paragraphs are shortened; the four trailing lines are verbatim. It
leaves exactly the story.
R.3 Evidence
| Check | Result |
|---|---|
| New and related unit tests (the A1 cold-window, A1 reserve, A2 echo, extractor, window, take and caching tests) | 308 passed, 3 skipped (the real-model skips) |
| Replay: the full v1 corpus | 518 replayed, 509 unchanged, 9 changed, 0 flagged. The changes are exactly the nine in §I, with the same rules: R5 changes no v1 turn |
| Replay: the v1.1 runs' raw replies (the 51-turn run's database and the identity run's database) | 64 replayed, 61 unchanged, 3 changed, 0 flagged: |
long run, action 77: R4, a trailing ```json (0.4%), as in §J.2 |
|
| identity depth 16: R3 scene line, plus R5 × 3 (the reworded length hint, the reminder, the reworded continue hint). 18.0% of the reply, all of it the tail above | |
identity depth 18: R4, a trailing ```json (0.8%) |
|
| Genuine story prose removed | 0 |
| Unreviewed replay flags | 0 |
Real cold-model A1 test (v11_window_accounting --unload-first, same host, model and 207-action bundle as §D.1) |
Prevented. Model unloaded (/api/ps empty). Turn 1: preflight attempted, model loaded, window verified 4,096 (loaded), budget 4,096, estimate 3,082, server read 3,097 (+15), margin 499, fits. Turn 2: no preflight needed (already verified), estimate 3,086, server 3,101, margin 495, fits. Before the correction the same turn sent 13,875 against an unverified window and the server read 2,050 |
Identity diagnostic, re-run (office fixture, 10 beats, the same CPU/HTTPS host at 4,096; $HOME/v11-evidence/a2c-identity) |
identity signals: 0 (verdict: no objective identity defect). Prompt example identifiers in proposals: 0 (no silver-key, aldric, mara, old-abbey, nor the placeholders). Stored protocol/instruction shapes: 0/10 (calls, hint and reminder brackets, continue-hint wording, a final scene line, state headings, "events", fences). Proposals: 9 accepted, 1 partially accepted, 1 unparseable. Read with the replay: this run's 10 raw replies replay unchanged through v1.0.0 and v1.1, so this narrator produced no echo this time. 0/10 here is the required closeout result, not a live exercise of R5. R5's proof is the verbatim depth-16 regression fixture and the replay of the original depth-16 turn above |
| Backend full suite | 1,534 passed, 17 skipped, 0 failed, 1,007 s (the 1,507 before this pass, plus 15 cold-window and 12 corrective A2 tests; the same 17 real-model skips) |
| Frontend suite | 165 of 165 (no frontend file changed in this pass) |
| Lint | exit 0; the same pre-existing only-export-components warnings |
| Production build | clean, 16 files; index-rWgm40jJ.js 394.47 kB and index-BXtb0iME.css 48.09 kB, the same hashes as before, so the SPA is byte-identical |
Offline container (tools/m11_offline.py, docker build --no-cache, --network none, a fresh volume) |
23 passed, 0 failed on the corrective tree ($HOME/v11-evidence/offline-corrective). The cold-model load is inert here: the failed-turn check has no reachable server, so no load is attempted (reachable false), and the turn is reported and commits nothing, as before |
Browser and context-inspector regression (tools/m11_browser.py, Firefox 155.0.1 headless, the built SPA served by FastAPI, the real narrator over trusted-LAN HTTPS; $HOME/v11-evidence/browser-corrective) |
38 passed, 0 failed, 0 skipped. This includes F05 (the context inspector shows the assembled prompt), two real turns through the UI, Undo and Redo, hostile narration, G01 import, the dialog and accessibility checks, and CSP. The two browser-played turns' stored snapshots carry the new provenance: the window was verified at 4,096 (loaded), the preflight was not attempted because the model was already resident, and accounting was fits (server 876 against estimate 861, margin 2,820; server 1,024 against estimate 1,009, margin 2,672). The inspector's accounting and safety-margin display is covered by the component suite (§C), not by this harness |
v1.0.0 database compatibility (a fresh v1.0.0 worktree against this tree, same database as §K) |
schema before: identical; schema after: identical; API snapshot: identical; user_version 94 → 94. Exercise: Undo 119 → 117, Redo 117 → 119, restore On the ridge → 69 with Redo available, dry run 200 with knowledge retrieved, export and import 207 → 207 actions with the same head and identical state |
R.4 Findings from the corrective pass
Blockers: none.
Corrective work within the pass:
- Harness slip: two background watchers never ran their jobs. The watchers
were meant to start the offline container after the suite, and the browser
regression after the identity run. Each waited with
pgrep -fon a command name that appears in its own command line, so each matched itself and never proceeded. Nothing ran and nothing was corrupted. Both watchers were stopped, and both runs were started directly (§R.3). This is the same class of slip as the earlierpkillself-match (§N.2). It sits in session scripting, not in the repository.
Non-blocking residual risks:
- R5 accepted false-positive risk. A story-world bracket that opens exactly
[Hard limit:and sits directly above a parroted instruction would be removed with it (§R.2). - Cold-model load: what it cannot fix.
- It does nothing for a server that is unreachable, or that does not speak
Ollama's native API. Both keep the unverified path, now recorded in
window.preflight. - It does not change the window. A model loaded at a smaller default is still that window; the turn is built to it, not to the setting.
- It costs one load request on a cold turn: 3.7 s measured on the CPU host for a 3B model, bounded by the model timeout.
- It does nothing for a server that is unreachable, or that does not speak
Ollama's native API. Both keep the unverified path, now recorded in
- The identity closeout's 0/10 does not exercise R5 live (§R.3). The shape is proven by the verbatim fixture and the replay, not by this run.
Future ideas:
- Warming the model when the reader opens a campaign, so the first turn does not pay for the load. This was deliberately not done, because the dry-run panel must have no side effects.
R.5 Closeout against the owner's required results
| Required | Result |
|---|---|
| A1: one bounded load of a cold model, no story content, then the window probed again | built and tested (§R.1); real server: HTTP 200 in 3.7 s, "response": "", "done_reason": "load" |
| A1: turn built to the verified window when the load makes it known | real cold test: verified 4,096, 3,082 sent, 3,097 read, fits. Before the correction: unverified, 13,875 sent, 2,050 read |
| A1: window still unknown → no guess, no hard-coded 4,096, the unverified path kept, accounting kept | tested (§R.1) |
| A1: no calibration, no user setting, no endpoint-policy change, only the configured endpoint, no dummy narration, no discarded turn | met (§R.1). The policy and same-host tests pass. The load writes nothing, tested by row counts |
| A2: minimised positive regression, and why the tail survived | test_the_depth_sixteen_instruction_tail_leaves_entirely; root cause in §R.2 |
| A2: narrowest rule; adversarial controls; all extractor tests | R5; 12 new tests; the extractor, A2 and state tests pass |
| A2: replay of all available real turns | 518 replayed, 509 unchanged, 9 changed, 0 flagged. Classifications exactly as §I.1. Plus the 64 v1.1 replies: 3 changed, 0 flagged |
| genuine story prose removed | 0 |
| unreviewed replay flags | 0 |
| identity signals | 0 |
| prompt example identifiers in proposals | 0 |
| stored protocol/instruction shapes | 0/10 |
| Backend full suite | 1,534 passed, 17 skipped, 0 failed |
| Frontend suite / lint / production build | 165/165; exit 0; clean and byte-identical |
| Offline container | 23 passed, 0 failed |
| Browser and context-inspector regression | 38 passed, 0 failed, 0 skipped |
| v1.0.0 database compatibility | identical schema before and after, identical API snapshot, user_version 94 → 94. Exercise passed |
R.6 Final decisions
WP-A1: PASS
WP-A2: PASS
WP-A1: PASS.
- What is met: every criterion in §O, plus the owner's corrective requirement.
- The case now prevented: the real cold-model truncation that A1 had only detected. The window becomes discoverable once the model loads, and the turn is built to it.
- What still depends on detection: a server that cannot be asked, or that will not load the model on request. That falls back to the unverified path, which is unchanged, recorded, and caught by the accounting.
WP-A2: PASS.
- What is met: every criterion in §P, plus the owner's corrective requirement.
- The residual §P recorded is closed. The depth-16 instruction tail is removed by a narrow rule anchored to application-owned text.
- Safety evidence: 0 genuine prose removed and 0 unreviewed flags across 582 real replies (518 v1 and 64 v1.1).
- Identity closeout: 0 signals, 0 example identifiers, and 0/10 stored shapes.
Next step. WP-B.1 may begin after the owner signs the A1/A2 commit. WP-B has not been started.
Git state. Everything is staged and uncommitted: product code, tests, tools, planning documents and this report. Nothing was committed, pushed or tagged.