M11: what the server will actually read

The release-validation milestone, and the thing it had to settle first was
whether any of the earlier evidence meant what it said. M8 measured a deployment
enforcing a 4,096-token input window while the application budgeted 16,384.
Every request returned 200. What Ollama does with the excess is drop the oldest
tokens, and the oldest tokens here are the system block — the narrator's rules
and the campaign canon. A hundred-turn certification against that server would
have looked perfect and proved nothing, which is why this milestone could not
begin with a hundred turns.

So the application asks now. Ollama's window is a property of how a model was
loaded rather than of the request — sending num_ctx is accepted, ignored, and
worse, reloads the model at the server's own default — so the only honest move
is to find out and then tell the truth about it. /api/ps reports what a resident
model is being served with, /api/show what an unloaded one will load with, both
on the same host inference already uses, through the same endpoint policy and
the same TLS trust store. A verified window is a ceiling on the budget; an
unverified one leaves the budget alone and is recorded as unverified in the
turn's own provenance, so an old turn can be asked afterwards whether it was
built against a checked window. There is no third behaviour, and in particular
no hard-coded 4,096: a number the server did not say would be right on one
machine and wrong on the next.

The proof that this is doing something is a campaign whose canon sits at the
front of the prompt, 120 turns of history, and a 4,096-token window. The canon
is still there afterwards and the oldest history is gone. The same campaign
built the old way produces a prompt more than twice the window — the defect,
reproduced, so the fix is measured against it rather than asserted.

Two defects the validation found on its own, and they are the same defect twice:
something was true and nobody was told. A manual state correction of four
changes with one bad reference applied three, returned 201, and said nothing —
while recording the refusal on the audit row nobody reads. It came to light
because the identity diagnostic's own fixture was refused that way and the whole
run proceeded on a campaign with no scene, which would have read as a model
failure. And the narration-length setting moved no number: brief, medium and
long each became one English sentence, while the numeric hint the model actually
reads was derived from the global reply cap and said the same thing for all
three. Both now say what they did.

The other two post-M8 findings are closed as well. The tab said AI D&D, which no
document had ever claimed it did not; it says Interactive Story now, with the
open campaign first, and the name is the owner's decision rather than a
find-and-replace to something narrower than the engine. After an Undo the reader
could not tell where they had landed; the control row now ends with
"Moment 11 · later story ahead", from the server's own answer, in the word the
transcript already uses, with none of head, branch or depth anywhere near it.

The identity diagnostic exists and the root cause does not. That campaign was
destroyed, so no cause can be established — what M11 owes the finding is
something that can classify the next occurrence, and a diagnostic that makes only
the judgements a program can honestly make: duplicate keys, shared names,
protagonist drift, state and context disagreeing. Whether prose misattributed a
line is left to a person reading it beside its prompt, because a regex cannot
read dialogue and one that pretended to would produce exactly the confident wrong
answer this finding is about. Its detectors are proved to fire against a planted
second Alice.

Two entities may still share a display name. That was checked first, as the
finding asked, and left permitted: a mother and a daughter, or a stranger giving
a false name, are ordinary fiction, and refusing them to guard against a model
mistake would refuse the wrong thing. What was missing was that it happened
silently. It is reported now.

Evidence, not inference: a hundred accepted turns against a real narrator with
genuine process restarts; a real browser against the built SPA; a container with
no network at all; a campaign moved into a data directory that never existed.
Each was discarded and re-run whenever the product changed under it, and the runs
that were thrown away are listed in the report with the reason, along with ten
defects in the harnesses themselves — because a harness that has only ever
agreed with itself is not evidence, and two of M8's five harness defects were
masking real ones.

No dependency was added, removed or upgraded. No acceptance test was retired,
relaxed or reclassified. M11 is implemented and verified; it is not accepted, and
there is no release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
JesseMarkowitz
2026-09-07 14:01:20 -04:00
co-authored by Claude Opus 5
parent 1013c94eb1
commit 144406cd48
57 changed files with 7374 additions and 97 deletions
+28
View File
@@ -157,6 +157,34 @@ later story is available — is one candidate among others. Ownership is M11
release polish; `V1-ACCEPTANCE-TESTS.md` §P1 records what must be settled before
this can become an acceptance test.
#### As implemented (M11)
The candidate above, built. A status line sits at the end of the story-control
row and reads `Moment 12`, gaining `· later story ahead` whenever the server
says Redo is available:
```text
Continue Retry Undo Redo Save Point Moment 11 · later story ahead
```
Three properties, because they are what make it answer §8A rather than merely
occupy the corner:
- **It is the server's answer, not the browser's.** The number is the count of
actions on the active line as the server reports it, and "later story ahead"
is `can_redo`. A component that decided either for itself would be wrong
exactly when it mattered — Undo can reach past the loaded window.
- **It changes visibly.** After an Undo the number decreases *and* the clause
appears; both are asserted, in the component suite and in the browser
regression, because a requirement that a working implementation can satisfy
while the reader is lost is the requirement §8A replaced.
- **It uses no implementation vocabulary** (§38): not head, not branch, not
depth. `Moment` is the word the transcript already uses for the same thing
("Read the 12 earlier moments"), so it introduces no new concept.
Evidence: `frontend/src/m11.test.jsx` ("finding B") and the `B` rows of
`tools/m11_browser.py`.
## 9. User Turn Presentation
User messages should support:
+50 -3
View File
@@ -985,7 +985,7 @@ accepted at closeout on 2026-09-06. The independent review returned
the closeout completed: the build-evidence classification in the report's §P and
finding 14's resolution.
`planning/reports/M8-IMPLEMENTATION-REPORT.md` records what was built, what was
`planning/archive/milestone-reports/M8-IMPLEMENTATION-REPORT.md` records what was built, what was
measured and every finding, including the seven product defects verification
found, the five harness defects, and the evidence runs that were discarded.
@@ -1122,7 +1122,7 @@ A campaign can be safely exported, imported into a clean data directory, and reo
Implemented on `m9-recovery` from the signed M8 commit `1ce9972`, measured
before and after against the same fixture, and verified in a real browser
against a real narrator. `planning/reports/M9-IMPLEMENTATION-REPORT.md` is the
against a real narrator. `planning/archive/milestone-reports/M9-IMPLEMENTATION-REPORT.md` is the
implementer's account, written for a reviewer.
**What it delivered, beyond the scope list above:**
@@ -1348,7 +1348,7 @@ Future media providers can be added through defined local interfaces without red
## Status: COMPLETE — 2026-09-07, pending independent review
Implemented on `m10-media-hooks` from the signed M9 commit `44edece`.
`planning/reports/M10-IMPLEMENTATION-REPORT.md` is the implementer's account,
`planning/archive/milestone-reports/M10-IMPLEMENTATION-REPORT.md` is the implementer's account,
written for a reviewer.
**The finding that shaped the milestone: the scene snapshot already existed.**
@@ -1449,6 +1449,53 @@ Release gate in `V1-ACCEPTANCE-TESTS.md`:
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07, awaiting independent review/acceptance
Implemented on `m11-release-validation` from the signed M10 commit `1013c94`.
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package, written
for a release reviewer. **M11 is not marked accepted here**; that is the
reviewer's to record, and no release tag exists.
**The release blocker it was given, and how it was closed.** M8 measured the
reference deployment enforcing a **4,096**-token input window while the
application budgeted **16,384** — every request returning 200, and `llama.cpp`
dropping the *oldest* tokens, which in this design are the narrator's rules and
the campaign canon. A 100-turn certification against that server would have
looked perfect and proved nothing.
M11's invariant: *the application must not silently budget more narrator input
than the runtime will accept.* It now asks the server — `/api/ps` for a loaded
model, `/api/show` for one that is not — under the same endpoint policy and TLS
trust as inference, and caps the prompt to what it finds, or records the window
as unverified in the turn's own provenance. Not a hard-coded 4,096, which would
cripple a correctly configured deployment; not a guess from the model's name.
Measured on the reference server: the plain model reports 4,096 and the budget
caps to it; the `num_ctx`-baked model reports 16,384 and the full budget stands.
**The four post-M8 playtest findings, disposed of:**
| | Disposition |
| --- | --- |
| **A.** The tab read `AI D&D` | **Fixed.** `Interactive Story`, with the open campaign first — a name chosen by the repository owner, and deliberately not "Adventure Storyteller", which is narrower than a genre-agnostic engine. One module owns it. |
| **B.** No orientation after Undo | **Fixed.** `Moment 11 · later story ahead`, from the server's own answer, in the transcript's existing vocabulary, with no implementation words. `BROWSER-UX-SPEC.md` §8A records the implementation. |
| **C.** Narration length had no effect | **Fixed.** The choice is data (`adventures.narration_length`), and the prompt builder turns it into a real word band. The generation budget is deliberately untouched: capping it would truncate prose, and the state block is emitted last. |
| **D.** Character identity confusion | **Diagnostic built; root cause remains unestablished, as it must.** The campaign was destroyed. `tools/m11_identity.py` runs the finding's own scenario, makes only the judgements a program can make honestly, preserves everything on a signal, and proves its detectors fire. The structural fact the finding asked M11 to check first — two entities may share a display name silently — is now **reported** rather than refused, because two people called Alice is ordinary fiction. |
**Two product defects found by the release validation itself**, both the same
family — something true that nobody was told:
1. **A partly refused manual state correction reported success.** Four changes,
one refused, HTTP 201, nothing said. Found because the identity diagnostic's
own fixture was refused that way and ran on a degraded campaign without
noticing. The refusal was already on the audit record; the reader was not
told. Now returned as `refused`, and shown in the State panel.
2. **The narration-length setting** above, which is finding C.
**Also corrected:** M9's residual risk 6 was narrower than recorded — the snap
Firefox refuses a WebDriver file path under `/tmp`, not all paths. Staging under
`$HOME` makes browser file import work, so knowledge import is now proved
end-to-end in a real browser rather than in two labelled halves.
---
## 4. Milestone Dependency Summary
+37
View File
@@ -884,6 +884,43 @@ rather than remembered: nothing in `app/media/` imports the code that writes it.
The reverse direction is the same rule seen from the other side — a depiction
never becomes canon (`MEDIA-EXTENSION-CONTRACT.md` §35, §37).
## 28B. What M11 Added (as implemented)
One column, and one field on a response. Both exist because something that was
happening silently had to become visible.
```text
adventures.narration_length "" | "brief" | "medium" | "long"
```
The campaign's own narration-length choice, and the *only* schema change M11
makes. Until M11 the choice became one English sentence inside
`ai_instructions` and moved no number: the numeric hint the model actually reads
was derived from the global reply cap and said the same thing for all three
settings. Stored as its own field because the prompt builder has to derive a
word range from it (`context.builder.LENGTH_BANDS`), and reading a length back
out of free text would be a parser nobody wants. Empty is not a missing value —
it is a campaign that never chose, which is exactly what every campaign created
before M11 did, so the migration needs no backfill and no existing prompt
changes under it.
**No new table.** The context-window ceiling M11 enforces is *not* stored: it is
asked of the server, cached in the process, and recorded in the turn's context
snapshot as provenance. A stored ceiling would be a second copy of a fact the
server owns, going stale the moment an operator reloads a model — the same
argument M10 made against a scenes table.
**`duplicate_names` is derived, not stored** (`narrative.model.duplicate_names`).
Two entities sharing a display name is permitted — a mother and a daughter, a
stranger giving a false name — and until M11 it was also *invisible*, which is
one of post-M8 finding D's candidate failure modes. It is computed from the
entities on read and reported beside the state.
**`refused` is per-response, not persisted.** A manual correction that is partly
refused now returns which of its changes did not apply and why. The refusals
were already recorded on the proposal row for §8's audit trail; what was missing
was telling the person who wrote them, who until M11 got an unqualified success.
## 29. Export Package
A campaign export should be capable of preserving:
+8 -7
View File
@@ -82,16 +82,17 @@ in `planning/archive/decisions/`.
One file, and it changes as development progresses:
```text
planning/reports/M8-IMPLEMENTATION-REPORT.md
planning/reports/M11-IMPLEMENTATION-REPORT.md
```
M8 is the most recently completed milestone, and M9 is the next to be briefed.
This report is M8's implementation account *and* its closeout record: its §P
classifies the browser evidence by build, its §S holds every finding, and its §V
records the acceptance.
M11 is the most recently completed milestone, and there is no next one to brief:
what follows is independent review and the v1 acceptance decision. This report is
the release-validation evidence package — its §F is the acceptance matrix, its §G
the 100-turn campaign, its §O every defect found, and its §R the release-readiness
answers. It is **not** an acceptance record.
**Replace it, do not accumulate.** When M9's report lands, remove this one from
the project Sources and upload M9's instead. The repository does the same thing:
**Replace it, do not accumulate.** When a later report lands, remove this one
from the project Sources and upload that one instead. The repository does the same thing:
`planning/reports/` holds the current milestone's report and
`planning/archive/milestone-reports/` holds the rest.
+56 -49
View File
@@ -3,8 +3,9 @@
**This file is the index. Start here.**
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
milestones **M1 through M8 implemented and accepted**, and **M9 and M10
implemented and awaiting review**. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
milestones **M1 through M8 implemented and accepted**; **M9 and M10 implemented,
committed and signed**; **M11 implemented and verified, awaiting independent
review and v1 acceptance**. M11 is the last planned milestone. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
independent review found a real defect and a corrective pass fixed it.
@@ -15,28 +16,34 @@ report keeps all three in sequence, and is now in
**M8 — Browser UX Completion for v1 Story Operations — is complete and
accepted** (2026-09-06), after an independent review and a closeout pass.
`reports/M8-IMPLEMENTATION-REPORT.md` is the implementer's account and now
carries the closeout: the build-evidence classification in its §P, finding 14's
operational resolution, and the acceptance record in its §V. The M8 tree is
staged and awaits the repository owner's signed commit.
`archive/milestone-reports/M8-IMPLEMENTATION-REPORT.md` is the implementer's
account and carries the closeout: the build-evidence classification in its §P,
finding 14's operational resolution, and the acceptance record in its §V. It is
committed and signed (`1ce9972`).
**M9 — Export, Backup, Recovery, and Migration Hardening — is implemented and
awaiting independent review** (2026-09-07).
`reports/M9-IMPLEMENTATION-REPORT.md` is the implementer's account, written for
a reviewer: a set of claims with the measurements attached, not yet a record of
acceptance. M8's report has moved to `archive/milestone-reports/`, which is
where a milestone report goes once the next milestone's report replaces it.
**M9 — Export, Backup, Recovery, and Migration Hardening — is committed and
signed** (`44edece`). It made a campaign portable in the way that matters: the
bundle became `ai-dnd-adventure-v3`, prompt provenance and state events travel,
and a verified SQLite backup exists. Its report is in
`archive/milestone-reports/`, along with M8's and M10's.
**M10 — Future Media Extension Hooks Only — is implemented and awaiting
independent review** (2026-09-07). `reports/M10-IMPLEMENTATION-REPORT.md` is the
implementer's account. It built the seam and no media: one `visual_profiles`
table, a scene packet derived on read, provider contracts with an empty
registry, and no dependency added. Its central finding is that the scene
snapshot the media contract asks for **already existed**, built by M5.
**M10 — Future Media Extension Hooks Only — is committed and signed**
(`1013c94`). It built the seam and no media: one `visual_profiles` table, a
scene packet derived on read, provider contracts with an empty registry, and no
dependency added. Its central finding was that the scene snapshot the media
contract asks for **already existed**, built by M5.
**Next: M11 — v1 Security, Long-Run, and Release Validation.** It has not been
started, and no brief for it exists. It also owns the four post-M8 hands-on
playtest findings recorded in `BUILD-MILESTONES.md`.
**M11 — v1 Security, Long-Run, and Release Validation — is implemented and
verified, and awaits independent review** (2026-09-07).
`reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. It closed the
context-window release blocker M8 found, fixed two defects the validation itself
surfaced, and disposed of all four post-M8 playtest findings. **It is not
accepted**, there is no release tag, and the tree is staged rather than
committed.
**After M11 there is no further planned milestone.** What follows is
independent review and the v1 acceptance decision, which is the repository
owner's.
**Package version:** see `VERSION.md`, which records what each revision changed
and why.
@@ -98,7 +105,7 @@ Two standing qualifications:
| Document | What it is for |
| --- | --- |
| `SPECIFICATION.md` | What the product must do. The top of the authority order. |
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M10 built, recorded as fact. |
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M11 built, recorded as fact. |
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the v3 export contract. |
| `STORY-BRANCH-SEMANTICS.md` | Undo/Redo/Retry/branch/take behavior, including the M3 ratifications. |
| `CONTEXT-AND-MEMORY.md` | Prompt assembly, summarization, branch-safe memory. |
@@ -126,9 +133,8 @@ Two standing qualifications:
10. `BROWSER-UX-SPEC.md`
11. `V1-ACCEPTANCE-TESTS.md`
12. `DECISIONS/` — all of them; they are short.
13. `reports/M10-IMPLEMENTATION-REPORT.md` and
`reports/M9-IMPLEMENTATION-REPORT.md`, for what the most recent milestones
actually left behind — read as claims to check, not records, until they are
13. `reports/M11-IMPLEMENTATION-REPORT.md`, for what the release validation
actually found — read as claims to check, not a record, until it is
reviewed. Nothing in `planning/archive/` unless sent there.
## Architectural decisions
@@ -160,19 +166,17 @@ work until Phase 0 closes — which Phase 0 satisfied on 2026-09-01. It is in
`reports/` holds the report for the milestone most recently completed, because
that is the one the next milestone's planning has to consult:
- `reports/M10-IMPLEMENTATION-REPORT.md` — the M10 implementation: the media
seam, everything it deliberately did not build, and the evidence for K01-K04.
Written by the implementer for an independent reviewer, so it is a set of
claims with the measurements attached and **not** a record of acceptance.
- `reports/M9-IMPLEMENTATION-REPORT.md` — the M9 implementation: the measured M8
portability baseline it started from, the final bundle contract, and the
evidence for every acceptance test it claims. Its §W carries the M10-M11
handoff and its §Y holds the post-M8 playtest findings.
- `reports/M11-IMPLEMENTATION-REPORT.md` — the release validation: the
acceptance matrix run end to end, the 100-turn campaign, the offline and
browser evidence, and every defect the run found. Written by the implementer
for an independent reviewer, so it is a set of claims with the measurements
attached and **not** a record of acceptance.
**It stays here rather than moving to the archive**, against the usual
rotation, because M9 has not been accepted yet: a reviewer of either milestone
needs it, since M10 built on M9 and its baseline is M9's. It moves once M9 is
accepted.
M9's and M10's reports both moved to `archive/milestone-reports/` when this one
was written. M9's had been kept here past its turn because M9 was unaccepted;
both are now committed and signed, so the ordinary rotation applies again. M9's
§Y — the post-M8 playtest findings — has a durable copy in `BUILD-MILESTONES.md`
and did not depend on that file staying put.
Completed earlier milestones are in `archive/milestone-reports/`, which M8's
report joined when M9's was written: a milestone report is useful during the
@@ -298,34 +302,37 @@ Milestone M7 COMPLETE (2026-09-06)
|
v
Milestone M8 COMPLETE / ACCEPTED (2026-09-06)
browser UX completion for v1 reports/M8-IMPLEMENTATION-REPORT.md
browser UX completion for v1 archive/milestone-reports/
story operations review + closeout, in sequence
|
v
Milestone M9 COMPLETE — awaiting review (2026-09-07)
export, backup, recovery, reports/M9-IMPLEMENTATION-REPORT.md
export, backup, recovery, archive/milestone-reports/
migration hardening bundle format v3; SQLite online backup
|
v
Milestone M10 COMPLETE — awaiting review (2026-09-07)
future media extension hooks reports/M10-IMPLEMENTATION-REPORT.md
Milestone M10 COMMITTED AND SIGNED (1013c94)
future media extension hooks archive/milestone-reports/
scene packet derived, not stored; no media
|
v
Milestone M11 NEXT — not started
v1 security, long-run, release see BUILD-MILESTONES.md
validation also owns the post-M8 playtest findings
Milestone M11 IMPLEMENTED AND VERIFIED (2026-09-07)
v1 security, long-run, release reports/M11-IMPLEMENTATION-REPORT.md
validation awaiting independent review; no release tag
|
v
v1 acceptance the repository owner's decision
```
## Stop Rule
**One milestone at a time. Do not begin a milestone before its brief exists.**
**No M11 brief has been prepared**, and neither M9 nor M10 is accepted — both
are implemented and awaiting independent review. Writing the M11 brief is the
action after those reviews close, informed by the M9 report's §W, the M10
report's handoff, and the four post-M8 playtest findings in
`BUILD-MILESTONES.md`.
**Every planned milestone is now implemented.** M11 is verified and staged, and
the next action is not another milestone: it is an independent review of
`reports/M11-IMPLEMENTATION-REPORT.md` against the acceptance contract, and then
the owner's v1 acceptance decision. **Do not begin post-v1 work before that
decision**, and do not treat M11's own report as the acceptance record.
All three questions the M8 debt raised against M9 are settled and recorded:
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
+36
View File
@@ -803,6 +803,42 @@ rather than by filtering marked secrets, which is what makes it hold for a secre
nobody thought to mark. Tested with a sentinel in a hidden source, alongside a
positive control proving the narrator did receive it.
## 42B. M11 Release Validation Notes
Three things M11 changed or measured that belong in this document. **None widens
the trust boundary**; two narrow what can happen silently, which is this
document's own concern (§56 story secrets, §69 auditability).
**The context-window probe is a new outbound request, and it is held to §73.**
`app/contextwindow.py` asks the configured Ollama for the window it will give a
model. It is the same host inference already uses, on the sibling native path,
through the same `endpoints.rejection_reason` check and the same TLS trust union
— so it can reach exactly what a turn can reach and nothing else. A probe that
could resolve an address inference may not would have been a hole in ADR 011, and
it is asserted not to be
(`test_m11_context_window.py::test_a_probe_obeys_the_same_endpoint_policy_as_inference`).
**A silently partial state correction was a §69 auditability gap.** A manual
correction of four changes where one was refused returned an unqualified success:
the refusal was recorded on the proposal row for the audit trail, and the person
who made it was told nothing. The correction response now carries `refused`, with
the reason. Partial application itself is unchanged and deliberate — discarding
three good changes because of one typo would be worse — what changed is that the
reader is told.
**H09 is NOT APPLICABLE, and the condition is now a test.** §59's Zip Slip
concern applies "if ZIP import/export is implemented". Nothing in the
application opens an archive — the bundle is JSON, an imported source is a single
file — and `test_m11_security.py::test_h09_the_product_extracts_no_archives`
fails the day that stops being true, at which point H09 becomes required again.
**Measured, not assumed:** every text/background pair in the palette clears WCAG
AA 1.4.3 (`tools/contrast_audit.py`, lowest 5.02:1). Two control-boundary pairs
are below 1.4.11's 3:1 and are recorded rather than failed, because in this
design a control is identified by its visible text label — which is measured and
passes — and not by its edge. That is a judgement stated so a reviewer can
disagree with it, not a threshold quietly lowered.
## 43. Logging
Logs should minimize story-content exposure.
+34
View File
@@ -1156,6 +1156,40 @@ narrator inference, which permits a trusted LAN host. A GPU that renders a
reader's campaign is a machine that reader is sitting at. No provider
configuration setting exists to point anywhere, because none is needed yet.
### 15.2 The inference window is a ceiling, not an assumption (M11)
M8 measured a reference deployment enforcing a **4,096**-token input window while
the application budgeted **16,384**, and every request returned HTTP 200. The
consequence is worse than an error: `llama.cpp` drops the *oldest* tokens, and
the oldest tokens in this design are the system block — the narrator's rules and
the campaign canon. A long campaign would quietly stop obeying its own canon,
and every acceptance test that reads a 200 as success would keep passing.
M11's rule:
> The application must not silently budget more narrator input than the
> configured Ollama runtime will actually accept.
`app/contextwindow.py` asks the server, on the same host and under the same
endpoint policy as inference: `/api/ps` reports the window a **loaded** model is
being served with, and `/api/show` reports the `num_ctx` an unloaded one will
load with plus the architecture's ceiling. The answer is cached per endpoint and
model, so it costs one short request per session rather than one per turn, and
it is cleared when either changes.
The builder takes the window as a parameter — like retrieved memories and
imported passages, and for the same reason: the prompt builder makes no network
calls. A **verified** window is a ceiling on `context_token_budget`; an
**unverified** one leaves the configured budget standing and is recorded as
unverified in the turn's stored provenance, on the context report, and on the
connection test. There is no third behaviour, and in particular there is no
hard-coded 4,096: guessing a number the server did not say would be right on one
machine and wrong on the next.
What this does not do is change the window. That is an operator action — a model
with `num_ctx` baked in, or `OLLAMA_CONTEXT_LENGTH` — and `DEVELOPMENT.md` says
how. What the application owes the reader is not to lie about it.
## 16. Database Direction
SQLite remains the selected v1 authoritative store.
+30
View File
@@ -2611,6 +2611,15 @@ identity, does not misattribute dialogue, and does not have a character refer to
themself as a separate same-named character — and the authoritative state and
the assembled context do not disagree about who anyone is.
**Settled by M11, and the answer was: report, do not refuse.** Two people called
Alice is ordinary fiction — a mother and a daughter, a stranger giving a false
name — and refusing it would refuse legitimate stories to guard against a model
mistake. What was actually missing was a *signal*: it happened silently and
nobody could see it. `narrative.model.duplicate_names` now reports it, the state
API returns it, the State panel shows it, and the identity diagnostic reads it.
The permissive behaviour is pinned by a test so a later milestone changes it
deliberately rather than by accident.
**Settle first:** which of those are **product** guarantees and which are
**model-quality** observations. They are not the same kind of claim and must not
share one verdict:
@@ -2632,3 +2641,24 @@ so a failing campaign can be exported whole and investigated elsewhere.
knowledge, authority, branch leakage and possession, and its on-stage cast is
effectively two people. A companion fixture is proposed in
`TEST-CAMPAIGN-FIXTURE.md`; the established fixture is deliberately unchanged.
### M11 disposition
Built as `backend/tools/m11_identity.py`: the protagonist and three supporting
characters, ten beats that stress pronouns, dialogue attribution, an entrance, an
exit, reference by name and by role, one character speaking about another, and
the protagonist spoken about in the third person. It makes only the judgements a
program can make honestly — duplicate keys, shared display names, protagonist
drift, state/context disagreement, derived contamination — preserves everything
the section above lists on any signal, and says plainly that prose-level
attribution is for a person to read, because a regex cannot read dialogue and a
diagnostic that pretended to would produce exactly the confident wrong answer
this finding is about.
Its detectors are proved to fire (`--scripted --inject` plants a second Alice and
the run reports it), which is the control this kind of tool most often lacks.
**This remains a test-design task and is still not an acceptance test.** The
model-quality half is not a pass/fail property of the application, and M11 does
not make it one. The M11 report records what the diagnostic found on the
reference narrator, including a fixture defect it caught in itself.
+42 -2
View File
@@ -1,8 +1,48 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v3.6
- **Package:** Adventure Storyteller Planning Package v3.7
- **Revision date:** 2026-09-07
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented and awaiting independent review** (2026-09-07). M11 has not been started.
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance.
## v3.7 — M11 implemented: v1 security, long-run and release validation (2026-09-07)
The release-validation milestone. Most of what it changed is evidence rather
than product; the product changes it did make were each forced by something the
validation found.
| Document | Change | Kind |
| --- | --- | --- |
| `TECHNICAL-DESIGN.md` | **New §15.2** — the inference window as a ceiling: how it is discovered, where it is enforced, and why there is no hard-coded 4,096. | as-implemented record |
| `DATA-MODEL.md` | **New §28B** — M11's one column (`narration_length`), and why the window ceiling, `duplicate_names` and `refused` are all deliberately *not* stored. | as-implemented record |
| `BROWSER-UX-SPEC.md` | **§8A gains an implementation note** — the position indicator that answers it, and the three properties that make it answer it. | as-implemented record |
| `SECURITY-THREAT-MODEL.md` | **New §42B** — the probe held to §73, the silent partial correction closed as a §69 gap, H09 recorded as not applicable with a test to keep it honest, and the measured contrast. | boundary + measurement notes |
| `V1-ACCEPTANCE-TESTS.md` | Results for the M11 release run against every REQUIRED test. | acceptance evidence |
| `BUILD-MILESTONES.md` | **M11 status block**, and the disposition of the four post-M8 playtest findings it owned. | milestone status |
| `README.md`, `DEVELOPMENT.md` | The context-window behaviour, `contextwindow.py` in the architecture map, and how to re-run the six release harnesses. | developer docs |
| `reports/M11-IMPLEMENTATION-REPORT.md` | New. The evidence package for independent release review. | milestone report |
| `reports/M10-IMPLEMENTATION-REPORT.md` | Moved to `archive/milestone-reports/`. | report rotation |
**The release blocker M11 was given, and what it cost.** M8 measured a
deployment enforcing 4,096 tokens while the application budgeted 16,384, with
every request returning 200 and `llama.cpp` silently dropping the oldest tokens —
which here are the narrator's rules and the campaign canon. M11 makes the
application ask the server what it will accept and cap itself to that, or say
that it could not check. Not a hard-coded number, not a cloud probe, not a
guess from the model's name: `/api/ps` for a loaded model, `/api/show` for one
that is not, under the same endpoint policy and TLS trust as inference.
**Two product defects found by the validation itself**, both of the same
family — something true that nobody was told:
- a **manual state correction that was partly refused** returned an unqualified
success. Found because the identity diagnostic's own fixture was refused that
way and the run proceeded silently on a degraded campaign;
- the **narration-length setting moved no number** (post-M8 finding C), so the
numeric hint the model reads said the same thing for brief, medium and long.
**Zero requirement weakenings.** No acceptance test was retired, relaxed or
reclassified. H09 is reported NOT APPLICABLE on the condition its own text
states, and that condition is now enforced by a test.
## v3.6 — M10 implemented: future media extension hooks only (2026-09-07)
File diff suppressed because it is too large Load Diff