M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was whether any of the earlier evidence meant what it said. M8 measured a deployment enforcing a 4,096-token input window while the application budgeted 16,384. Every request returned 200. What Ollama does with the excess is drop the oldest tokens, and the oldest tokens here are the system block — the narrator's rules and the campaign canon. A hundred-turn certification against that server would have looked perfect and proved nothing, which is why this milestone could not begin with a hundred turns. So the application asks now. Ollama's window is a property of how a model was loaded rather than of the request — sending num_ctx is accepted, ignored, and worse, reloads the model at the server's own default — so the only honest move is to find out and then tell the truth about it. /api/ps reports what a resident model is being served with, /api/show what an unloaded one will load with, both on the same host inference already uses, through the same endpoint policy and the same TLS trust store. A verified window is a ceiling on the budget; an unverified one leaves the budget alone and is recorded as unverified in the turn's own provenance, so an old turn can be asked afterwards whether it was built against a checked window. There is no third behaviour, and in particular no hard-coded 4,096: a number the server did not say would be right on one machine and wrong on the next. The proof that this is doing something is a campaign whose canon sits at the front of the prompt, 120 turns of history, and a 4,096-token window. The canon is still there afterwards and the oldest history is gone. The same campaign built the old way produces a prompt more than twice the window — the defect, reproduced, so the fix is measured against it rather than asserted. Two defects the validation found on its own, and they are the same defect twice: something was true and nobody was told. A manual state correction of four changes with one bad reference applied three, returned 201, and said nothing — while recording the refusal on the audit row nobody reads. It came to light because the identity diagnostic's own fixture was refused that way and the whole run proceeded on a campaign with no scene, which would have read as a model failure. And the narration-length setting moved no number: brief, medium and long each became one English sentence, while the numeric hint the model actually reads was derived from the global reply cap and said the same thing for all three. Both now say what they did. The other two post-M8 findings are closed as well. The tab said AI D&D, which no document had ever claimed it did not; it says Interactive Story now, with the open campaign first, and the name is the owner's decision rather than a find-and-replace to something narrower than the engine. After an Undo the reader could not tell where they had landed; the control row now ends with "Moment 11 · later story ahead", from the server's own answer, in the word the transcript already uses, with none of head, branch or depth anywhere near it. The identity diagnostic exists and the root cause does not. That campaign was destroyed, so no cause can be established — what M11 owes the finding is something that can classify the next occurrence, and a diagnostic that makes only the judgements a program can honestly make: duplicate keys, shared names, protagonist drift, state and context disagreeing. Whether prose misattributed a line is left to a person reading it beside its prompt, because a regex cannot read dialogue and one that pretended to would produce exactly the confident wrong answer this finding is about. Its detectors are proved to fire against a planted second Alice. Two entities may still share a display name. That was checked first, as the finding asked, and left permitted: a mother and a daughter, or a stranger giving a false name, are ordinary fiction, and refusing them to guard against a model mistake would refuse the wrong thing. What was missing was that it happened silently. It is reported now. Evidence, not inference: a hundred accepted turns against a real narrator with genuine process restarts; a real browser against the built SPA; a container with no network at all; a campaign moved into a data directory that never existed. Each was discarded and re-run whenever the product changed under it, and the runs that were thrown away are listed in the report with the reason, along with ten defects in the harnesses themselves — because a harness that has only ever agreed with itself is not evidence, and two of M8's five harness defects were masking real ones. No dependency was added, removed or upgraded. No acceptance test was retired, relaxed or reclassified. M11 is implemented and verified; it is not accepted, and there is no release tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
This commit is contained in:
co-authored by
Claude Opus 5
parent
1013c94eb1
commit
144406cd48
@@ -157,6 +157,34 @@ later story is available — is one candidate among others. Ownership is M11
|
||||
release polish; `V1-ACCEPTANCE-TESTS.md` §P1 records what must be settled before
|
||||
this can become an acceptance test.
|
||||
|
||||
#### As implemented (M11)
|
||||
|
||||
The candidate above, built. A status line sits at the end of the story-control
|
||||
row and reads `Moment 12`, gaining `· later story ahead` whenever the server
|
||||
says Redo is available:
|
||||
|
||||
```text
|
||||
Continue Retry Undo Redo Save Point Moment 11 · later story ahead
|
||||
```
|
||||
|
||||
Three properties, because they are what make it answer §8A rather than merely
|
||||
occupy the corner:
|
||||
|
||||
- **It is the server's answer, not the browser's.** The number is the count of
|
||||
actions on the active line as the server reports it, and "later story ahead"
|
||||
is `can_redo`. A component that decided either for itself would be wrong
|
||||
exactly when it mattered — Undo can reach past the loaded window.
|
||||
- **It changes visibly.** After an Undo the number decreases *and* the clause
|
||||
appears; both are asserted, in the component suite and in the browser
|
||||
regression, because a requirement that a working implementation can satisfy
|
||||
while the reader is lost is the requirement §8A replaced.
|
||||
- **It uses no implementation vocabulary** (§38): not head, not branch, not
|
||||
depth. `Moment` is the word the transcript already uses for the same thing
|
||||
("Read the 12 earlier moments"), so it introduces no new concept.
|
||||
|
||||
Evidence: `frontend/src/m11.test.jsx` ("finding B") and the `B` rows of
|
||||
`tools/m11_browser.py`.
|
||||
|
||||
## 9. User Turn Presentation
|
||||
|
||||
User messages should support:
|
||||
|
||||
@@ -985,7 +985,7 @@ accepted at closeout on 2026-09-06. The independent review returned
|
||||
the closeout completed: the build-evidence classification in the report's §P and
|
||||
finding 14's resolution.
|
||||
|
||||
`planning/reports/M8-IMPLEMENTATION-REPORT.md` records what was built, what was
|
||||
`planning/archive/milestone-reports/M8-IMPLEMENTATION-REPORT.md` records what was built, what was
|
||||
measured and every finding, including the seven product defects verification
|
||||
found, the five harness defects, and the evidence runs that were discarded.
|
||||
|
||||
@@ -1122,7 +1122,7 @@ A campaign can be safely exported, imported into a clean data directory, and reo
|
||||
|
||||
Implemented on `m9-recovery` from the signed M8 commit `1ce9972`, measured
|
||||
before and after against the same fixture, and verified in a real browser
|
||||
against a real narrator. `planning/reports/M9-IMPLEMENTATION-REPORT.md` is the
|
||||
against a real narrator. `planning/archive/milestone-reports/M9-IMPLEMENTATION-REPORT.md` is the
|
||||
implementer's account, written for a reviewer.
|
||||
|
||||
**What it delivered, beyond the scope list above:**
|
||||
@@ -1348,7 +1348,7 @@ Future media providers can be added through defined local interfaces without red
|
||||
## Status: COMPLETE — 2026-09-07, pending independent review
|
||||
|
||||
Implemented on `m10-media-hooks` from the signed M9 commit `44edece`.
|
||||
`planning/reports/M10-IMPLEMENTATION-REPORT.md` is the implementer's account,
|
||||
`planning/archive/milestone-reports/M10-IMPLEMENTATION-REPORT.md` is the implementer's account,
|
||||
written for a reviewer.
|
||||
|
||||
**The finding that shaped the milestone: the scene snapshot already existed.**
|
||||
@@ -1449,6 +1449,53 @@ Release gate in `V1-ACCEPTANCE-TESTS.md`:
|
||||
|
||||
The build meets the v1 black-box acceptance contract and can be packaged as the first production release.
|
||||
|
||||
## Status: IMPLEMENTED AND VERIFIED — 2026-09-07, awaiting independent review/acceptance
|
||||
|
||||
Implemented on `m11-release-validation` from the signed M10 commit `1013c94`.
|
||||
`planning/reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package, written
|
||||
for a release reviewer. **M11 is not marked accepted here**; that is the
|
||||
reviewer's to record, and no release tag exists.
|
||||
|
||||
**The release blocker it was given, and how it was closed.** M8 measured the
|
||||
reference deployment enforcing a **4,096**-token input window while the
|
||||
application budgeted **16,384** — every request returning 200, and `llama.cpp`
|
||||
dropping the *oldest* tokens, which in this design are the narrator's rules and
|
||||
the campaign canon. A 100-turn certification against that server would have
|
||||
looked perfect and proved nothing.
|
||||
|
||||
M11's invariant: *the application must not silently budget more narrator input
|
||||
than the runtime will accept.* It now asks the server — `/api/ps` for a loaded
|
||||
model, `/api/show` for one that is not — under the same endpoint policy and TLS
|
||||
trust as inference, and caps the prompt to what it finds, or records the window
|
||||
as unverified in the turn's own provenance. Not a hard-coded 4,096, which would
|
||||
cripple a correctly configured deployment; not a guess from the model's name.
|
||||
Measured on the reference server: the plain model reports 4,096 and the budget
|
||||
caps to it; the `num_ctx`-baked model reports 16,384 and the full budget stands.
|
||||
|
||||
**The four post-M8 playtest findings, disposed of:**
|
||||
|
||||
| | Disposition |
|
||||
| --- | --- |
|
||||
| **A.** The tab read `AI D&D` | **Fixed.** `Interactive Story`, with the open campaign first — a name chosen by the repository owner, and deliberately not "Adventure Storyteller", which is narrower than a genre-agnostic engine. One module owns it. |
|
||||
| **B.** No orientation after Undo | **Fixed.** `Moment 11 · later story ahead`, from the server's own answer, in the transcript's existing vocabulary, with no implementation words. `BROWSER-UX-SPEC.md` §8A records the implementation. |
|
||||
| **C.** Narration length had no effect | **Fixed.** The choice is data (`adventures.narration_length`), and the prompt builder turns it into a real word band. The generation budget is deliberately untouched: capping it would truncate prose, and the state block is emitted last. |
|
||||
| **D.** Character identity confusion | **Diagnostic built; root cause remains unestablished, as it must.** The campaign was destroyed. `tools/m11_identity.py` runs the finding's own scenario, makes only the judgements a program can make honestly, preserves everything on a signal, and proves its detectors fire. The structural fact the finding asked M11 to check first — two entities may share a display name silently — is now **reported** rather than refused, because two people called Alice is ordinary fiction. |
|
||||
|
||||
**Two product defects found by the release validation itself**, both the same
|
||||
family — something true that nobody was told:
|
||||
|
||||
1. **A partly refused manual state correction reported success.** Four changes,
|
||||
one refused, HTTP 201, nothing said. Found because the identity diagnostic's
|
||||
own fixture was refused that way and ran on a degraded campaign without
|
||||
noticing. The refusal was already on the audit record; the reader was not
|
||||
told. Now returned as `refused`, and shown in the State panel.
|
||||
2. **The narration-length setting** above, which is finding C.
|
||||
|
||||
**Also corrected:** M9's residual risk 6 was narrower than recorded — the snap
|
||||
Firefox refuses a WebDriver file path under `/tmp`, not all paths. Staging under
|
||||
`$HOME` makes browser file import work, so knowledge import is now proved
|
||||
end-to-end in a real browser rather than in two labelled halves.
|
||||
|
||||
---
|
||||
|
||||
## 4. Milestone Dependency Summary
|
||||
|
||||
@@ -884,6 +884,43 @@ rather than remembered: nothing in `app/media/` imports the code that writes it.
|
||||
The reverse direction is the same rule seen from the other side — a depiction
|
||||
never becomes canon (`MEDIA-EXTENSION-CONTRACT.md` §35, §37).
|
||||
|
||||
## 28B. What M11 Added (as implemented)
|
||||
|
||||
One column, and one field on a response. Both exist because something that was
|
||||
happening silently had to become visible.
|
||||
|
||||
```text
|
||||
adventures.narration_length "" | "brief" | "medium" | "long"
|
||||
```
|
||||
|
||||
The campaign's own narration-length choice, and the *only* schema change M11
|
||||
makes. Until M11 the choice became one English sentence inside
|
||||
`ai_instructions` and moved no number: the numeric hint the model actually reads
|
||||
was derived from the global reply cap and said the same thing for all three
|
||||
settings. Stored as its own field because the prompt builder has to derive a
|
||||
word range from it (`context.builder.LENGTH_BANDS`), and reading a length back
|
||||
out of free text would be a parser nobody wants. Empty is not a missing value —
|
||||
it is a campaign that never chose, which is exactly what every campaign created
|
||||
before M11 did, so the migration needs no backfill and no existing prompt
|
||||
changes under it.
|
||||
|
||||
**No new table.** The context-window ceiling M11 enforces is *not* stored: it is
|
||||
asked of the server, cached in the process, and recorded in the turn's context
|
||||
snapshot as provenance. A stored ceiling would be a second copy of a fact the
|
||||
server owns, going stale the moment an operator reloads a model — the same
|
||||
argument M10 made against a scenes table.
|
||||
|
||||
**`duplicate_names` is derived, not stored** (`narrative.model.duplicate_names`).
|
||||
Two entities sharing a display name is permitted — a mother and a daughter, a
|
||||
stranger giving a false name — and until M11 it was also *invisible*, which is
|
||||
one of post-M8 finding D's candidate failure modes. It is computed from the
|
||||
entities on read and reported beside the state.
|
||||
|
||||
**`refused` is per-response, not persisted.** A manual correction that is partly
|
||||
refused now returns which of its changes did not apply and why. The refusals
|
||||
were already recorded on the proposal row for §8's audit trail; what was missing
|
||||
was telling the person who wrote them, who until M11 got an unqualified success.
|
||||
|
||||
## 29. Export Package
|
||||
|
||||
A campaign export should be capable of preserving:
|
||||
|
||||
@@ -82,16 +82,17 @@ in `planning/archive/decisions/`.
|
||||
One file, and it changes as development progresses:
|
||||
|
||||
```text
|
||||
planning/reports/M8-IMPLEMENTATION-REPORT.md
|
||||
planning/reports/M11-IMPLEMENTATION-REPORT.md
|
||||
```
|
||||
|
||||
M8 is the most recently completed milestone, and M9 is the next to be briefed.
|
||||
This report is M8's implementation account *and* its closeout record: its §P
|
||||
classifies the browser evidence by build, its §S holds every finding, and its §V
|
||||
records the acceptance.
|
||||
M11 is the most recently completed milestone, and there is no next one to brief:
|
||||
what follows is independent review and the v1 acceptance decision. This report is
|
||||
the release-validation evidence package — its §F is the acceptance matrix, its §G
|
||||
the 100-turn campaign, its §O every defect found, and its §R the release-readiness
|
||||
answers. It is **not** an acceptance record.
|
||||
|
||||
**Replace it, do not accumulate.** When M9's report lands, remove this one from
|
||||
the project Sources and upload M9's instead. The repository does the same thing:
|
||||
**Replace it, do not accumulate.** When a later report lands, remove this one
|
||||
from the project Sources and upload that one instead. The repository does the same thing:
|
||||
`planning/reports/` holds the current milestone's report and
|
||||
`planning/archive/milestone-reports/` holds the rest.
|
||||
|
||||
|
||||
+56
-49
@@ -3,8 +3,9 @@
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
milestones **M1 through M8 implemented and accepted**, and **M9 and M10
|
||||
implemented and awaiting review**. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
milestones **M1 through M8 implemented and accepted**; **M9 and M10 implemented,
|
||||
committed and signed**; **M11 implemented and verified, awaiting independent
|
||||
review and v1 acceptance**. M11 is the last planned milestone. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
|
||||
independent review found a real defect and a corrective pass fixed it.
|
||||
|
||||
@@ -15,28 +16,34 @@ report keeps all three in sequence, and is now in
|
||||
|
||||
**M8 — Browser UX Completion for v1 Story Operations — is complete and
|
||||
accepted** (2026-09-06), after an independent review and a closeout pass.
|
||||
`reports/M8-IMPLEMENTATION-REPORT.md` is the implementer's account and now
|
||||
carries the closeout: the build-evidence classification in its §P, finding 14's
|
||||
operational resolution, and the acceptance record in its §V. The M8 tree is
|
||||
staged and awaits the repository owner's signed commit.
|
||||
`archive/milestone-reports/M8-IMPLEMENTATION-REPORT.md` is the implementer's
|
||||
account and carries the closeout: the build-evidence classification in its §P,
|
||||
finding 14's operational resolution, and the acceptance record in its §V. It is
|
||||
committed and signed (`1ce9972`).
|
||||
|
||||
**M9 — Export, Backup, Recovery, and Migration Hardening — is implemented and
|
||||
awaiting independent review** (2026-09-07).
|
||||
`reports/M9-IMPLEMENTATION-REPORT.md` is the implementer's account, written for
|
||||
a reviewer: a set of claims with the measurements attached, not yet a record of
|
||||
acceptance. M8's report has moved to `archive/milestone-reports/`, which is
|
||||
where a milestone report goes once the next milestone's report replaces it.
|
||||
**M9 — Export, Backup, Recovery, and Migration Hardening — is committed and
|
||||
signed** (`44edece`). It made a campaign portable in the way that matters: the
|
||||
bundle became `ai-dnd-adventure-v3`, prompt provenance and state events travel,
|
||||
and a verified SQLite backup exists. Its report is in
|
||||
`archive/milestone-reports/`, along with M8's and M10's.
|
||||
|
||||
**M10 — Future Media Extension Hooks Only — is implemented and awaiting
|
||||
independent review** (2026-09-07). `reports/M10-IMPLEMENTATION-REPORT.md` is the
|
||||
implementer's account. It built the seam and no media: one `visual_profiles`
|
||||
table, a scene packet derived on read, provider contracts with an empty
|
||||
registry, and no dependency added. Its central finding is that the scene
|
||||
snapshot the media contract asks for **already existed**, built by M5.
|
||||
**M10 — Future Media Extension Hooks Only — is committed and signed**
|
||||
(`1013c94`). It built the seam and no media: one `visual_profiles` table, a
|
||||
scene packet derived on read, provider contracts with an empty registry, and no
|
||||
dependency added. Its central finding was that the scene snapshot the media
|
||||
contract asks for **already existed**, built by M5.
|
||||
|
||||
**Next: M11 — v1 Security, Long-Run, and Release Validation.** It has not been
|
||||
started, and no brief for it exists. It also owns the four post-M8 hands-on
|
||||
playtest findings recorded in `BUILD-MILESTONES.md`.
|
||||
**M11 — v1 Security, Long-Run, and Release Validation — is implemented and
|
||||
verified, and awaits independent review** (2026-09-07).
|
||||
`reports/M11-IMPLEMENTATION-REPORT.md` is the evidence package. It closed the
|
||||
context-window release blocker M8 found, fixed two defects the validation itself
|
||||
surfaced, and disposed of all four post-M8 playtest findings. **It is not
|
||||
accepted**, there is no release tag, and the tree is staged rather than
|
||||
committed.
|
||||
|
||||
**After M11 there is no further planned milestone.** What follows is
|
||||
independent review and the v1 acceptance decision, which is the repository
|
||||
owner's.
|
||||
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
and why.
|
||||
@@ -98,7 +105,7 @@ Two standing qualifications:
|
||||
| Document | What it is for |
|
||||
| --- | --- |
|
||||
| `SPECIFICATION.md` | What the product must do. The top of the authority order. |
|
||||
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M10 built, recorded as fact. |
|
||||
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M11 built, recorded as fact. |
|
||||
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the v3 export contract. |
|
||||
| `STORY-BRANCH-SEMANTICS.md` | Undo/Redo/Retry/branch/take behavior, including the M3 ratifications. |
|
||||
| `CONTEXT-AND-MEMORY.md` | Prompt assembly, summarization, branch-safe memory. |
|
||||
@@ -126,9 +133,8 @@ Two standing qualifications:
|
||||
10. `BROWSER-UX-SPEC.md`
|
||||
11. `V1-ACCEPTANCE-TESTS.md`
|
||||
12. `DECISIONS/` — all of them; they are short.
|
||||
13. `reports/M10-IMPLEMENTATION-REPORT.md` and
|
||||
`reports/M9-IMPLEMENTATION-REPORT.md`, for what the most recent milestones
|
||||
actually left behind — read as claims to check, not records, until they are
|
||||
13. `reports/M11-IMPLEMENTATION-REPORT.md`, for what the release validation
|
||||
actually found — read as claims to check, not a record, until it is
|
||||
reviewed. Nothing in `planning/archive/` unless sent there.
|
||||
|
||||
## Architectural decisions
|
||||
@@ -160,19 +166,17 @@ work until Phase 0 closes — which Phase 0 satisfied on 2026-09-01. It is in
|
||||
`reports/` holds the report for the milestone most recently completed, because
|
||||
that is the one the next milestone's planning has to consult:
|
||||
|
||||
- `reports/M10-IMPLEMENTATION-REPORT.md` — the M10 implementation: the media
|
||||
seam, everything it deliberately did not build, and the evidence for K01-K04.
|
||||
Written by the implementer for an independent reviewer, so it is a set of
|
||||
claims with the measurements attached and **not** a record of acceptance.
|
||||
- `reports/M9-IMPLEMENTATION-REPORT.md` — the M9 implementation: the measured M8
|
||||
portability baseline it started from, the final bundle contract, and the
|
||||
evidence for every acceptance test it claims. Its §W carries the M10-M11
|
||||
handoff and its §Y holds the post-M8 playtest findings.
|
||||
- `reports/M11-IMPLEMENTATION-REPORT.md` — the release validation: the
|
||||
acceptance matrix run end to end, the 100-turn campaign, the offline and
|
||||
browser evidence, and every defect the run found. Written by the implementer
|
||||
for an independent reviewer, so it is a set of claims with the measurements
|
||||
attached and **not** a record of acceptance.
|
||||
|
||||
**It stays here rather than moving to the archive**, against the usual
|
||||
rotation, because M9 has not been accepted yet: a reviewer of either milestone
|
||||
needs it, since M10 built on M9 and its baseline is M9's. It moves once M9 is
|
||||
accepted.
|
||||
M9's and M10's reports both moved to `archive/milestone-reports/` when this one
|
||||
was written. M9's had been kept here past its turn because M9 was unaccepted;
|
||||
both are now committed and signed, so the ordinary rotation applies again. M9's
|
||||
§Y — the post-M8 playtest findings — has a durable copy in `BUILD-MILESTONES.md`
|
||||
and did not depend on that file staying put.
|
||||
|
||||
Completed earlier milestones are in `archive/milestone-reports/`, which M8's
|
||||
report joined when M9's was written: a milestone report is useful during the
|
||||
@@ -298,34 +302,37 @@ Milestone M7 COMPLETE (2026-09-06)
|
||||
|
|
||||
v
|
||||
Milestone M8 COMPLETE / ACCEPTED (2026-09-06)
|
||||
browser UX completion for v1 reports/M8-IMPLEMENTATION-REPORT.md
|
||||
browser UX completion for v1 archive/milestone-reports/
|
||||
story operations review + closeout, in sequence
|
||||
|
|
||||
v
|
||||
Milestone M9 COMPLETE — awaiting review (2026-09-07)
|
||||
export, backup, recovery, reports/M9-IMPLEMENTATION-REPORT.md
|
||||
export, backup, recovery, archive/milestone-reports/
|
||||
migration hardening bundle format v3; SQLite online backup
|
||||
|
|
||||
v
|
||||
Milestone M10 COMPLETE — awaiting review (2026-09-07)
|
||||
future media extension hooks reports/M10-IMPLEMENTATION-REPORT.md
|
||||
Milestone M10 COMMITTED AND SIGNED (1013c94)
|
||||
future media extension hooks archive/milestone-reports/
|
||||
scene packet derived, not stored; no media
|
||||
|
|
||||
v
|
||||
Milestone M11 NEXT — not started
|
||||
v1 security, long-run, release see BUILD-MILESTONES.md
|
||||
validation also owns the post-M8 playtest findings
|
||||
Milestone M11 IMPLEMENTED AND VERIFIED (2026-09-07)
|
||||
v1 security, long-run, release reports/M11-IMPLEMENTATION-REPORT.md
|
||||
validation awaiting independent review; no release tag
|
||||
|
|
||||
v
|
||||
v1 acceptance the repository owner's decision
|
||||
```
|
||||
|
||||
## Stop Rule
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
|
||||
**No M11 brief has been prepared**, and neither M9 nor M10 is accepted — both
|
||||
are implemented and awaiting independent review. Writing the M11 brief is the
|
||||
action after those reviews close, informed by the M9 report's §W, the M10
|
||||
report's handoff, and the four post-M8 playtest findings in
|
||||
`BUILD-MILESTONES.md`.
|
||||
**Every planned milestone is now implemented.** M11 is verified and staged, and
|
||||
the next action is not another milestone: it is an independent review of
|
||||
`reports/M11-IMPLEMENTATION-REPORT.md` against the acceptance contract, and then
|
||||
the owner's v1 acceptance decision. **Do not begin post-v1 work before that
|
||||
decision**, and do not treat M11's own report as the acceptance record.
|
||||
|
||||
All three questions the M8 debt raised against M9 are settled and recorded:
|
||||
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
|
||||
|
||||
@@ -803,6 +803,42 @@ rather than by filtering marked secrets, which is what makes it hold for a secre
|
||||
nobody thought to mark. Tested with a sentinel in a hidden source, alongside a
|
||||
positive control proving the narrator did receive it.
|
||||
|
||||
## 42B. M11 Release Validation Notes
|
||||
|
||||
Three things M11 changed or measured that belong in this document. **None widens
|
||||
the trust boundary**; two narrow what can happen silently, which is this
|
||||
document's own concern (§56 story secrets, §69 auditability).
|
||||
|
||||
**The context-window probe is a new outbound request, and it is held to §73.**
|
||||
`app/contextwindow.py` asks the configured Ollama for the window it will give a
|
||||
model. It is the same host inference already uses, on the sibling native path,
|
||||
through the same `endpoints.rejection_reason` check and the same TLS trust union
|
||||
— so it can reach exactly what a turn can reach and nothing else. A probe that
|
||||
could resolve an address inference may not would have been a hole in ADR 011, and
|
||||
it is asserted not to be
|
||||
(`test_m11_context_window.py::test_a_probe_obeys_the_same_endpoint_policy_as_inference`).
|
||||
|
||||
**A silently partial state correction was a §69 auditability gap.** A manual
|
||||
correction of four changes where one was refused returned an unqualified success:
|
||||
the refusal was recorded on the proposal row for the audit trail, and the person
|
||||
who made it was told nothing. The correction response now carries `refused`, with
|
||||
the reason. Partial application itself is unchanged and deliberate — discarding
|
||||
three good changes because of one typo would be worse — what changed is that the
|
||||
reader is told.
|
||||
|
||||
**H09 is NOT APPLICABLE, and the condition is now a test.** §59's Zip Slip
|
||||
concern applies "if ZIP import/export is implemented". Nothing in the
|
||||
application opens an archive — the bundle is JSON, an imported source is a single
|
||||
file — and `test_m11_security.py::test_h09_the_product_extracts_no_archives`
|
||||
fails the day that stops being true, at which point H09 becomes required again.
|
||||
|
||||
**Measured, not assumed:** every text/background pair in the palette clears WCAG
|
||||
AA 1.4.3 (`tools/contrast_audit.py`, lowest 5.02:1). Two control-boundary pairs
|
||||
are below 1.4.11's 3:1 and are recorded rather than failed, because in this
|
||||
design a control is identified by its visible text label — which is measured and
|
||||
passes — and not by its edge. That is a judgement stated so a reviewer can
|
||||
disagree with it, not a threshold quietly lowered.
|
||||
|
||||
## 43. Logging
|
||||
|
||||
Logs should minimize story-content exposure.
|
||||
|
||||
@@ -1156,6 +1156,40 @@ narrator inference, which permits a trusted LAN host. A GPU that renders a
|
||||
reader's campaign is a machine that reader is sitting at. No provider
|
||||
configuration setting exists to point anywhere, because none is needed yet.
|
||||
|
||||
### 15.2 The inference window is a ceiling, not an assumption (M11)
|
||||
|
||||
M8 measured a reference deployment enforcing a **4,096**-token input window while
|
||||
the application budgeted **16,384**, and every request returned HTTP 200. The
|
||||
consequence is worse than an error: `llama.cpp` drops the *oldest* tokens, and
|
||||
the oldest tokens in this design are the system block — the narrator's rules and
|
||||
the campaign canon. A long campaign would quietly stop obeying its own canon,
|
||||
and every acceptance test that reads a 200 as success would keep passing.
|
||||
|
||||
M11's rule:
|
||||
|
||||
> The application must not silently budget more narrator input than the
|
||||
> configured Ollama runtime will actually accept.
|
||||
|
||||
`app/contextwindow.py` asks the server, on the same host and under the same
|
||||
endpoint policy as inference: `/api/ps` reports the window a **loaded** model is
|
||||
being served with, and `/api/show` reports the `num_ctx` an unloaded one will
|
||||
load with plus the architecture's ceiling. The answer is cached per endpoint and
|
||||
model, so it costs one short request per session rather than one per turn, and
|
||||
it is cleared when either changes.
|
||||
|
||||
The builder takes the window as a parameter — like retrieved memories and
|
||||
imported passages, and for the same reason: the prompt builder makes no network
|
||||
calls. A **verified** window is a ceiling on `context_token_budget`; an
|
||||
**unverified** one leaves the configured budget standing and is recorded as
|
||||
unverified in the turn's stored provenance, on the context report, and on the
|
||||
connection test. There is no third behaviour, and in particular there is no
|
||||
hard-coded 4,096: guessing a number the server did not say would be right on one
|
||||
machine and wrong on the next.
|
||||
|
||||
What this does not do is change the window. That is an operator action — a model
|
||||
with `num_ctx` baked in, or `OLLAMA_CONTEXT_LENGTH` — and `DEVELOPMENT.md` says
|
||||
how. What the application owes the reader is not to lie about it.
|
||||
|
||||
## 16. Database Direction
|
||||
|
||||
SQLite remains the selected v1 authoritative store.
|
||||
|
||||
@@ -2611,6 +2611,15 @@ identity, does not misattribute dialogue, and does not have a character refer to
|
||||
themself as a separate same-named character — and the authoritative state and
|
||||
the assembled context do not disagree about who anyone is.
|
||||
|
||||
**Settled by M11, and the answer was: report, do not refuse.** Two people called
|
||||
Alice is ordinary fiction — a mother and a daughter, a stranger giving a false
|
||||
name — and refusing it would refuse legitimate stories to guard against a model
|
||||
mistake. What was actually missing was a *signal*: it happened silently and
|
||||
nobody could see it. `narrative.model.duplicate_names` now reports it, the state
|
||||
API returns it, the State panel shows it, and the identity diagnostic reads it.
|
||||
The permissive behaviour is pinned by a test so a later milestone changes it
|
||||
deliberately rather than by accident.
|
||||
|
||||
**Settle first:** which of those are **product** guarantees and which are
|
||||
**model-quality** observations. They are not the same kind of claim and must not
|
||||
share one verdict:
|
||||
@@ -2632,3 +2641,24 @@ so a failing campaign can be exported whole and investigated elsewhere.
|
||||
knowledge, authority, branch leakage and possession, and its on-stage cast is
|
||||
effectively two people. A companion fixture is proposed in
|
||||
`TEST-CAMPAIGN-FIXTURE.md`; the established fixture is deliberately unchanged.
|
||||
|
||||
### M11 disposition
|
||||
|
||||
Built as `backend/tools/m11_identity.py`: the protagonist and three supporting
|
||||
characters, ten beats that stress pronouns, dialogue attribution, an entrance, an
|
||||
exit, reference by name and by role, one character speaking about another, and
|
||||
the protagonist spoken about in the third person. It makes only the judgements a
|
||||
program can make honestly — duplicate keys, shared display names, protagonist
|
||||
drift, state/context disagreement, derived contamination — preserves everything
|
||||
the section above lists on any signal, and says plainly that prose-level
|
||||
attribution is for a person to read, because a regex cannot read dialogue and a
|
||||
diagnostic that pretended to would produce exactly the confident wrong answer
|
||||
this finding is about.
|
||||
|
||||
Its detectors are proved to fire (`--scripted --inject` plants a second Alice and
|
||||
the run reports it), which is the control this kind of tool most often lacks.
|
||||
|
||||
**This remains a test-design task and is still not an acceptance test.** The
|
||||
model-quality half is not a pass/fail property of the application, and M11 does
|
||||
not make it one. The M11 report records what the diagnostic found on the
|
||||
reference narrator, including a fixture defect it caught in itself.
|
||||
|
||||
+42
-2
@@ -1,8 +1,48 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v3.6
|
||||
- **Package:** Adventure Storyteller Planning Package v3.7
|
||||
- **Revision date:** 2026-09-07
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented and awaiting independent review** (2026-09-07). M11 has not been started.
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented, M10 committed and signed**; **M11 implemented and verified, awaiting independent review/acceptance** (2026-09-07). M11 is the last planned milestone before v1 acceptance.
|
||||
|
||||
## v3.7 — M11 implemented: v1 security, long-run and release validation (2026-09-07)
|
||||
|
||||
The release-validation milestone. Most of what it changed is evidence rather
|
||||
than product; the product changes it did make were each forced by something the
|
||||
validation found.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `TECHNICAL-DESIGN.md` | **New §15.2** — the inference window as a ceiling: how it is discovered, where it is enforced, and why there is no hard-coded 4,096. | as-implemented record |
|
||||
| `DATA-MODEL.md` | **New §28B** — M11's one column (`narration_length`), and why the window ceiling, `duplicate_names` and `refused` are all deliberately *not* stored. | as-implemented record |
|
||||
| `BROWSER-UX-SPEC.md` | **§8A gains an implementation note** — the position indicator that answers it, and the three properties that make it answer it. | as-implemented record |
|
||||
| `SECURITY-THREAT-MODEL.md` | **New §42B** — the probe held to §73, the silent partial correction closed as a §69 gap, H09 recorded as not applicable with a test to keep it honest, and the measured contrast. | boundary + measurement notes |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | Results for the M11 release run against every REQUIRED test. | acceptance evidence |
|
||||
| `BUILD-MILESTONES.md` | **M11 status block**, and the disposition of the four post-M8 playtest findings it owned. | milestone status |
|
||||
| `README.md`, `DEVELOPMENT.md` | The context-window behaviour, `contextwindow.py` in the architecture map, and how to re-run the six release harnesses. | developer docs |
|
||||
| `reports/M11-IMPLEMENTATION-REPORT.md` | New. The evidence package for independent release review. | milestone report |
|
||||
| `reports/M10-IMPLEMENTATION-REPORT.md` | Moved to `archive/milestone-reports/`. | report rotation |
|
||||
|
||||
**The release blocker M11 was given, and what it cost.** M8 measured a
|
||||
deployment enforcing 4,096 tokens while the application budgeted 16,384, with
|
||||
every request returning 200 and `llama.cpp` silently dropping the oldest tokens —
|
||||
which here are the narrator's rules and the campaign canon. M11 makes the
|
||||
application ask the server what it will accept and cap itself to that, or say
|
||||
that it could not check. Not a hard-coded number, not a cloud probe, not a
|
||||
guess from the model's name: `/api/ps` for a loaded model, `/api/show` for one
|
||||
that is not, under the same endpoint policy and TLS trust as inference.
|
||||
|
||||
**Two product defects found by the validation itself**, both of the same
|
||||
family — something true that nobody was told:
|
||||
|
||||
- a **manual state correction that was partly refused** returned an unqualified
|
||||
success. Found because the identity diagnostic's own fixture was refused that
|
||||
way and the run proceeded silently on a degraded campaign;
|
||||
- the **narration-length setting moved no number** (post-M8 finding C), so the
|
||||
numeric hint the model reads said the same thing for brief, medium and long.
|
||||
|
||||
**Zero requirement weakenings.** No acceptance test was retired, relaxed or
|
||||
reclassified. H09 is reported NOT APPLICABLE on the condition its own text
|
||||
states, and that condition is now enforced by a test.
|
||||
|
||||
## v3.6 — M10 implemented: future media extension hooks only (2026-09-07)
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user