Files
interactive-story/planning/reports/M8-IMPLEMENTATION-REPORT.md
JesseMarkowitzandClaude Opus 5 1ce9972760 M8: the browser becomes the storyteller
The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.

All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.

So the shape now is one entry point and one screen:

  Campaigns -> Campaign -> Story
                           State · Knowledge · Context · Save Points · Settings

Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.

Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.

Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.

The two defects worth the space:

A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.

And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.

Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.

`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.

Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.

Backend, and only what the browser could not otherwise reach:

  AdventureCreate.opening   a start action could only come from a Scenario, so
                            every campaign made in the new setup flow opened on
                            a blank page. Same node, same code path.
  canon_rules               campaign_canon has been the highest authority in a
                            campaign since M5, read by the prompt builder and
                            the state validator, and had no API at all — a
                            fixture had to write it with SQL.
  a 401 and a 429 message   the last user-facing text describing a hosted
                            deployment. One told the reader to check an API key
                            that has not existed since M2.

No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.

The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.

They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.

A verification pass over all of it then found three more, each by driving the
product rather than reading it:

Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.

The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.

And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.

Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.

`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.

Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:

The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.

Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.

The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.

The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.

Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 23:31:45 -04:00

92 KiB
Raw Permalink Blame History

M8 — Browser UX Completion for v1 Story Operations

Final implementation, verification and closeout report. Written by the implementer for an independent reviewer, and completed at closeout after that review returned M8 IMPLEMENTATION: PASS. The closeout additions are the build-evidence classification in §P, finding 14's resolution, and the acceptance record in §V; no verification figure was changed by them.

Closeout date: 2026-09-06.

Placeholders. inference.lan stands for the trusted-LAN Ollama host and 192.168.0.x for LAN addresses, following the convention M1 established. No real hostname, address or identifier of the machine this was built on appears in any committed file.


A. Executive result

PASS

Every M8-owned acceptance condition was demonstrated through the browser workflow against a real local narrator, on one frozen production build. No required condition is unverified, and there are no blockers to acceptance.

Two qualifications a reviewer should weigh, neither of them a blocker:

  1. Seven product defects were found during verification, four of them by driving the application rather than reading it (§S findings 1-4, 7). Two made the interface state something untrue to the reader. All are fixed with regression coverage, and two of the regressions were verified to fail against the unfixed code before the fix was kept. But their existence says the implementation pass was not as verified as it looked at the time.
  2. The evidence harness needed five corrections before it could be trusted (§S finding 12), and three evidence runs were invalidated by process error (§S finding 13). Two of those harness defects were actively masking product defects. The figures below are from the corrected harness on the final build; the history is recorded so a reviewer can judge how much they rest on.

M8 is implemented, verified, reviewed and accepted. The independent review returned M8 IMPLEMENTATION: PASS subject to evidence and documentation cleanup; that cleanup is §P's build classification and finding 14's resolution, both complete. M8 is not yet committed: the tree is staged for the repository owner's signature, which is the only step left before M9.

What this milestone is, in one paragraph

M8 turned an adapted AI-DnD interface into the interactive-story workspace BROWSER-UX-SPEC.md describes: one entry point, one story screen, one natural-language input, and everything advanced one layer deeper behind panels that start closed. Almost nothing underneath changed — the play loop, the history controls, the takes, the Save Points, the state engine and the knowledge library are M3-M7's and are untouched. What changed is what a reader is shown and asked to understand.

Evidence at a glance

Gate Result
Backend suite 950 passed, 14 skipped, 0 failed
Frontend component suite 132 passed, 10 files
Lint exit 0
Production build clean
Docker build clean
Browser acceptance (6 scenarios, B01-B04, D01-D14) 57/57
Hidden-information sentinel 21/21
A05 failed generation 23/23
Security / offline / performance 27/27
Genuine process restart 12/12
Migration against an M7-built database 17/17
Reader-facing terminology audit 0 hits (17 normal-play components; both advanced surfaces)

Every browser figure is from one frozen production build, dist/assets/index-Ii-lARp9.js, built from the staged tree at 18:53:02 and unchanged for the rest of the campaign. An earlier artifact, index-C6E5Uvtu.js, is superseded: it predates finding 7's fix, its acceptance suite ended 54/55 on that defect, and no figure above comes from it. §P sets the two side by side with the evidence for the classification.

What a reviewer should look at first

  1. §S findings 1, 2, 3 and 7 — four product defects that the browser work found and that no earlier milestone had caught. Two of them made the interface state something untrue to the reader.
  2. §S findings 12 and 13 — the harness needed five corrections before it could be trusted, and three evidence runs were invalidated by process error. This bears directly on how much weight the figures above deserve.
  3. §F — the canon_rules decision, where the smallest correction was a surface constraint rather than an audit subsystem, and why.
  4. §T — the one requirement whose wording M8 changed (§38), and the argument that it was tightened rather than weakened.

B. Repository and provenance state

Branch m8-browser-ux
M8 base 480414efe082a4bfe0600a19fe23961f6bddd925 — M7
M7 signature Good signature, verified against the repository owner’s RSA key 02C9BF7D…B5C68569, ultimate trust
M7's parent a6e9c7a32bdf42f1e4cb837b70721d89422cef8d (M6) — sole parent
Current HEAD 480414e… — still M7. M8 has made no commit.
Staged 89 files — 34 added, 24 deleted, 30 modified, 1 renamed (git diff --cached --stat)
Unstaged 0
Untracked 0
LICENSE unchanged, absent from the staged diff
Upstream ancestry d72f7c1b… (AI-DnD) is still an ancestor

Committed HEAD vs staged tree vs tested tree

These are three different things and the distinction matters for review:

  • Committed HEAD is M7, untouched. Nothing in M8 has been committed.

  • The staged tree is the whole M8 result — 89 files, including this report and the planning corrections.

  • The tested tree is the staged tree. The frontend was built from it once and frozen; every browser figure in this report ran against that one artifact, dist/assets/index-Ii-lARp9.js (built 18:53:02, and still the only file in frontend/dist/assets/), and no source was edited between the freeze and the last suite.

    An earlier artifact, index-C6E5Uvtu.js, appears in the evidence directory and is superseded, not final: it predates finding 7's fix, its acceptance suite ended 54/55 on exactly that defect, and no figure in this report comes from it. §P records the classification and the evidence for it.

That last point is stated because it was broken three times during M8 and each time the run was discarded rather than reported. See findings 11 and 13.

M8 is implemented, verified, reviewed and accepted (2026-09-06). What is outstanding is the owner's signed commit, not a decision.


C. The M7 browser baseline, and how it directed the work

Measured before any code was changed, by driving the M7 build in a real Firefox against a real backend with the Continuity Test fixture and two real turns played (scratchpad/m8/baseline.py, 13 observations). Nothing here is recalled; each figure is a browser observation.

Observation Measured What it directed
Context inspector size 11,996 characters on opening, beginning with the assembled prompt §55 says do not begin with raw prompt text. Drove the whole restructure: four readable sections first, the prompt last and collapsed. Result ~1,400 chars (§R).
Unlabelled controls 16 of 16 per-message buttons had no accessible name — single glyphs 🔍 ✎ ⑂ ✕ with only a title Drove named controls throughout, and the a11y.test.jsx assertion that a name under three characters counts as a glyph.
Branch UI visible in normal play panel tabs read Story State · Plot · Memory · Knowledge · **Branches** · Save Points · Insights §27. Drove removing the branch panel and tree overlay from the browser, and the terminology audit (§M).
Rigid input modes Do / Say / Story §12. Drove one natural-language field plus the direction toggle — and, indirectly, exposed the > You I … defect (finding 4).
No model status anywhere play header read Continuity Test / Story State / Plot / … and said nothing about Ollama §44 and §8. Drove the five-state badge and the setup notice.
Navigation Home · Adventures · Scenarios · Settings · AI Chat §97. Drove Campaigns · Settings with everything else inside a campaign.
Landing page vocabulary "The table is set", "worlds to explore", a scenario gallery §26. Drove the campaign library.
Settings Model a free-text box; a raw request/response log un-collapsed at the bottom §8 and §72. Drove the model picker and the folded diagnostics.
Campaign creation required choosing a Scenario; its editor held a JSON stat schema, a story-card table and an art picker §41-42 and §93. Drove the one-screen setup form and the opening field.
Frontend test coverage none at all; two backend tests read JSX as text to approximate browser assertions §30. Drove the whole test foundation (§N).

What the baseline showed was already right

Worth recording, because it bounded the work: the play loop, streaming, Undo, Redo, Retry, alternate takes, Save Points, state correction, knowledge import and prompt inspection were all reachable and all correct. The Save Point panel's confirmations already said the right things. The State panel already avoided raw JSON. Undo and Redo already used the server's answer rather than deriving availability in React.

M8 changed what a reader is shown and asked to understand, not the machinery. The two exceptions are findings 1 and 7, where the presentation layer turned out to be calling the wrong mechanism — and both were caught by driving the product, not by reading it.


D. M8 change inventory

89 files — 34 added, 24 deleted, 30 modified and 1 renamed. The frontend was substantially rebuilt; the backend was touched in five places, each to expose a browser capability that could not otherwise be reached.

Browser presentation — added

File What it is
pages/Campaigns.jsx the campaign library — the landing page
pages/NewCampaign.jsx genre-neutral setup, one screen
pages/Play/Transcript.jsx the story, extracted from the page component
pages/Play/Composer.jsx one input, direction toggle, history controls, reserved dictation
pages/Play/FailureNotice.jsx §71's five failure kinds
pages/Play/SidePanel.jsx the secondary panel and its tabs
pages/Play/panels/ContextPanel.jsx the context inspector, restructured
pages/Play/panels/CampaignSettingsPanel.jsx campaign settings, canon, export, delete
Dialog.jsx accessible modal with a real focus trap
ExternalLinkDialog.jsx §74's leaving-the-local-environment warning
ModelStatusBadge.jsx, ModelSetupNotice.jsx §44 status, and §8's way out of a blank model
markdown.jsx safe Markdown for story prose
errors.js the failure taxonomy
modelStatus.jsx one shared model-status answer, fetched once
styles/story.css, library.css, context.css, dialogs.css the new surfaces

Browser presentation — removed

Sixteen components and eight stylesheets, each recorded with its reason in the file that replaced it.

Removed Why
pages/Home.jsx, pages/Adventures.jsx two screens listing the same rows; one library replaces both
pages/Scenarios.jsx, pages/ScenarioEditor.jsx, SchemaEditor.jsx, ArtPicker.jsx a campaign no longer needs a template, and the editor was §93's "dangerous advanced features" almost exactly
pages/Chat.jsx a raw model console in the primary navigation
panels/BranchPanel.jsx, BranchMap.jsx, branches.js §27 — branch management is not a v1 surface
panels/PlotPanel.jsx, RefreshModal.jsx AI Dungeon world-info editing; the M7 knowledge library supersedes it
panels/MemoryPanel.jsx memory is now shown where it is used, in the context inspector
panels/InsightsPanel.jsx renamed and restructured as ContextPanel.jsx
drawers/WorldStateDrawer.jsx the RPG stat editor (§18, §26)
Embers.jsx decorative fire; §40
8 stylesheets the surfaces they styled; auth.css reduced to debuglog.css

No backend capability was removed. The tree, branch switching, story cards and the RPG world state all still exist, are still tested, and still travel in the bundle.

Backend support — five files, no schema change

Change Why the existing API could not serve the UX
AdventureCreate.opening a start action could only come from a Scenario's prompt, so every campaign made in the new setup flow opened on a blank page
canon_rules on create, update and read campaign_canon has been read by the prompt builder and the state validator since M5 and had no API at all — a fixture had to write it with SQL
format_player_input finding 4
_friendly_http_error 401/429 the last user-facing text describing a hosted deployment
models.Adventure.canon_rules a read-only property. No column, no migration.

Test infrastructure

Vitest + jsdom + Testing Library (5 devDependencies, no runtime dependency); ten frontend test files; two new backend test files (test_m8_setup_surface.py, and the normalization tests in test_take_parentage.py); two inherited source-level guards retargeted or replaced — see §N.

Documentation

README.md, DEVELOPMENT.md, and six planning documents — itemised in §T.


E. The final user experience, in user terms

A person opens the application and sees their campaigns — titles, when each was last played, how many moments it holds, and the opening of its most recent narration. No ids. Two buttons: import one, or start a new one.

Starting one asks for a name, and nothing else is required. If they want to, they can say the genre and tone (free text, with suggestions spanning fantasy, science fiction, mystery, historical, western, horror, thriller and literary), choose a voice and a narration length, name a protagonist, write the opening scene, and write the rules the story must not contradict. Then Start.

Then they are in the story, and the story is the whole screen. A thin bar at the top carries the campaign's name, whether Ollama is connected and with which model, and five tabs that are all closed. Underneath, prose at a readable measure — the reader's own turns indented and italic, the narrator's plain, an ornament between scenes, a drop cap on the opening.

They type what they do into one box. Not a mode, not a command — "I walk into the Crooked Lantern and look for Mara." The reply streams in. If they want the narrator to do something rather than their character, they tick Story direction and the box says so.

Above the box: Continue, Retry, Undo, Redo, Save Point. Undo and Redo grey out when there is nowhere to go. Hovering a turn reveals Inspect context, Edit, Try again, Copy — named, not glyphs.

When something is wrong they are told which thing: the model is unreachable and here is how to start it; the turn failed and here is Retry, with their words still in the box; the story state could not be updated and the story itself is fine.

When they want to know why the narrator said that, one click on the turn opens what it was given: what it cost, what it read, what it remembered, what it believes. If a passage was marked narrator-only, its text is not there — a control offers it, and says it will reveal secrets they marked for the narrator alone.

They can play for an hour without meeting the word branch.


F. Campaign setup, opening, and canon

The opening field

A start action could previously only come from a Scenario's prompt. M8's setup flow creates a campaign from a form, so without this every new campaign opened on a blank page — the reader had to invent the situation and the first move in one box.

It builds the same node by the same path (attempts.snapshot_outcome → tree.place_action). Verified, 13 checks, now permanent in backend/tests/test_m8_setup_surface.py:

Setup can provide it, as a start action PASS
Never duplicated on re-read PASS
Placed at depth 0 on the tree, with a parent of None PASS
The head is at it; Undo past the opening is refused (HTTP 400) PASS
A blank opening still yields an empty campaign PASS
A Scenario's prompt takes precedence, and the two are never both inserted PASS
A scenario-made campaign is unchanged by the new field PASS
Survives export and import PASS

One observation, not an M8 regression: the create response reports action_count: 0 while carrying one action, because the count is computed on read. A scenario-made campaign has always done the same. The browser navigates and re-reads, so a reader never sees it; the test asserts on the read path.

canon_rules, and what editing canon after play actually does

campaign_canon has been the highest authority in a campaign since M5 — read by both the prompt builder and the state validator — and had no API at all. A fixture had to write it with SQL. M8 exposed the sentence list.

That raised a question the column had never had to answer, and §6 of the review brief asked for it to be settled rather than assumed. It was measured against a real server with real turns (scratchpad/m8/canon_provenance.py):

Question Answer
Does a historical turn keep the canon it was actually told? Yes — the canon section is in that turn's stored context snapshot. After the edit it still shows the old rule and not the new one.
Does the next turn get the new canon? Yes — which is the point of editing it.
Is any accepted story text rewritten? No — byte-identical.
Is the narrative state document changed? No — byte-identical.
Is the state audit log changed? No — byte-identical.
Is the edit itself audited? No. It is a configuration overwrite.
Is any other campaign setting audited? No — editing narrator instructions is not either.

Why M5's audit log was deliberately not extended

M5's StateEvent log audits accepted changes to narrative state. Canon is not narrative state: it is configuration, sitting with ai_instructions and the narrator prompt. Routing a configuration change through that log would create a second representation of canon — the same duplication BROWSER-UX-SPEC.md §38 forbids for hidden information, arrived at from a different direction — and §6 of the brief explicitly rules out inventing a parallel audit subsystem.

So the correction was made at the editing surface instead, which is the brief's second option. Once a campaign has moments, the canon editor states that the change applies from here on, that everything already written stays exactly as it is, and where the per-turn record can be seen. Three component tests cover it.

The guarantee M8 offers is therefore not "the edit is logged" but "the edit cannot be mistaken for a retroactive one, and every turn can prove what it was told" — which is what the per-turn snapshot already delivered, unasked.


G. The story composer

One field, not three modes

The inherited composer had a Do / Say / Story selector, and the mode changed what the turn meant: "I say to Mara…" typed in Do mode and the same words in Say mode were different turns, and nothing on screen said so. That is a small command language wearing buttons.

M8 sends one natural-language field. B01 and B02 are the proof that it works — an action and a piece of quoted dialogue are both just what the reader wrote, and the narrator reads the quotes without being told which kind of turn it is.

What survives from the old story mode is the Story direction toggle, and it is deliberately not a fourth mode: it changes who is being spoken to, not what kind of action is taken. The box is visibly marked while it is on, and the send button reads "Direct" rather than "Send".

The > You I enter the tavern. defect

Severity: high. Found by driving the new composer in a browser; fixed here.

AI Dungeon stores a player action with a > You prefix. That was right when the Do mode asked for a bare verb phrase — look around became > You look around. — and it is what marks whose turn it is in the replayed prompt.

BROWSER-UX-SPEC.md §11 tells the reader to write "I enter the tavern." With one field, the prefix produced:

> You I enter the tavern.

in the transcript, in the replayed history, and therefore in the narration, where a small model imitates the pattern it is shown and writes "You I thank her". The M7 baseline transcript contains exactly that phrasing, so the defect predates M8 in the data — but M8's own design is what made it reachable on every turn, which is why M8 owns it.

The correction adds the subject only when the reader has not already written one. The > marker — which is what actually identifies a player turn — is unchanged in every case.

Regression coverage is against the shared normalizer, not one rendered component, because storage, the transcript, the replayed history and the export all read the result of that one function:

  • §11's three examples verbatim;
  • eight first-person phrasings, each asserted to contain no duplicate subject and to keep its marker;
  • the legacy bare action still normalizes (open the door → > You open the door.), which is deliberate compatibility;
  • an explicit You … is de-duplicated rather than doubled;
  • dialogue and out-of-character direction untouched;
  • and a second test asserts on the assembled prompt, that each player line appears exactly once and "You I " appears nowhere — because a normalizer that is correct but applied twice would put the defect straight back in front of the model.

The storage marker is not shown, and not parsed

> is also Markdown for a blockquote, so leaving the stored marker in made every player turn render as a quote by accident, stacking the renderer's rule on the one .turn-player already draws. The marker is stripped for display only; the stored text keeps it, and editing a turn puts it back around the edited words.

The reserved dictation control

Present, permanently disabled, named "Dictate — not yet available". No handler, no permission request, no microphone. A component test tabs twelve times to assert it never takes focus, and another installs a fake getUserMedia and clicks the disabled button to assert it is never called.


H. History controls

M3's active-head machinery, M4's Save Points and the take/divergence rules are unchanged. M8 changed how they are presented and what they are called.

Undo and Redo

Availability is the server's answer, carried on every window it returns, and the browser never computes it. Neither is derivable client-side: Undo can reach past the top of the loaded page, and Redo depends on a retained future the transcript is never sent. Six component tests cover the enabled states, including that each is independent and that every control is disabled while a turn is generating.

Retry and takes

Retry sits in the story controls and "Try again" on each narrator turn. When a turn has more than one attempt the pager appears as ‹ 2/3 ›, reading "Take 2 of 3" to a screen reader — §14 asks for a count, and 2/3 is too terse to hear.

A defect here had existed since the pager was written. See finding 1: every step between takes performed a branch switch, which for two takes of an ordinary retry meant switching to the line already being read — the same window came back and the step did nothing at all. D07 is REQUIRED FOR V1 and had never been exercised in a browser. Fixed, and confirmed PASS in the final run.

Save Points

M4's panel needed little. Create with a name and an optional note, list, restore, rename, delete. Both confirmations say what does not happen, because that is the part a reader cannot see and would otherwise assume the worst about:

The story will return to this Save Point. Everything you wrote after it is kept — it just stops being where you are.

Delete this Save Point? Deleting it does not delete any of the story — only the name you gave this moment.

Counted in moments, not turns, matching the vocabulary M4's closeout settled.

Vocabulary

branch, fork, node, merge, head and depth appear nowhere a reader operates. The branch panel and the tree overlay are gone from the browser entirely. The mechanism underneath is untouched and still fully tested — this is a decision about what a reader is asked to understand, not a reduction of what the product can do.

The audit method and its result are in §M.


I. State

The server groups the authoritative state and the panel renders the groups, so the headings are whatever the campaign has established rather than a fixed list — a campaign with no items shows no Items heading.

  • No raw JSON. Asserted: the panel contains no <pre> and no {.
  • No RPG vocabulary of its own. Asserted against hp, mana, quest, stat, cooldown, xp. The product term is Story Threads, never Quests.
  • Correction without JSON. "Correct something" asks for a sentence — the one shape a person can write without knowing the event vocabulary — and it goes through the same validator a narration's proposal does. C04 was exercised in the browser end to end: the correction appears in the panel, reaches authoritative state, is recorded as manual_correction, and the next turn's assembled prompt contains it.
  • It follows the head. Keyed on the same signal as the other live panels, so Undo, Redo and a Save Point restore all move it — what it shows is the state at the position being read, not at the newest turn.

Hidden state

There is none, and M8 did not invent one. A secret lives in a narrator-only knowledge source and never enters the state document; M7 verified that against a real narrator, and the sentinel run re-verified it here (§Q). §38's requirement is met by this panel simply not containing narrator-only information. The surface that must actively withhold it is the context inspector — see §K.

The RPG world state

Read-only, and shown in the context inspector only for a campaign that has a stat schema. A campaign created in M8's setup flow has none. The editing drawer is gone (§18, §26); the backend and the bundle are untouched.


J. Knowledge

Every M7 behaviour is unchanged. M8 owed the design, and §21-22 of the brief is what it was measured against.

  • Classification is explained, not iconified (§48). The import form carries the three classes as radio choices with a sentence each — authoritative truth / supporting information that establishes nothing / creative influence only — and the same explanations appear as a legend in the empty state. The class tag is one shared component, so a class looks identical in the library and in the context inspector.
  • Each source shows its title, filename when it differs, class, in-use state, narrator-only and always-include badges, import date and passage count.
  • Deleting explains what it does not do: turns that already used the source are unchanged, each keeping its own record of what the narrator was given — and it points at "In use" as the reversible alternative.
  • Source detail carries the text, every passage with its heading trail and token count, and — behind a "Technical details" disclosure — the media type, parser and chunking versions, and the full SHA-256.

Semantic status (§22)

Three states, told apart rather than run together:

State What the reader is told
on searching by keyword and by meaning
no model configured searching by keyword; setting a model for meaning-based search would add meaning — and keyword search alone is a supported setup, which is what finds names and invented terms
configured but uncalibrated searching by keyword only, because that model has not been measured for this in this build; borrowing another model's measurement would let unrelated material through; keyword search is unaffected

The third is the M7 closeout case, and the one that reads as a mysterious failure if it is not explained.

No threshold is exposed and no slider exists. 0.58 is a measured property of one embedding model, not a preference, and a control over it would invite a reader to recreate the defect M7's review found. Asserted: the panel's text never contains 0.58 and contains no input[type=range].

A terminology correction

The audit (§M) found the panel pointing readers at an "embedding model" while the Settings field is called Model for meaning-based search — a reader sent looking for a field that does not exist by that name. That is a usability defect rather than vocabulary policing, and it is finding 6. Both states now name the field as the reader sees it, and two tests assert the panel's text contains no "embedding" at all.

Rendering

Nothing here renders imported text as markup. Source text and passage text both go into a <pre> as React children. The safe Markdown renderer the transcript uses is deliberately not used: this panel exists to show a reader exactly what is in their file, and rendering is the opposite of that.


K. The context inspector

Reduction from the M7 baseline

M7 baseline M8
Text on opening the panel, same campaign 11,996 characters ~1,400 characters
What it opens on the assembled prompt, dumped into <pre> token usage, then four readable sections
The assembled prompt first last, collapsed

§55 says not to begin with raw prompt text. The reader's question is almost never "what were the exact bytes"; it is "why did the narrator say that?", and the answer is one of four things — what it was told to be, what it remembered, what it read, or what it believes.

Provenance visibility

Each retrieved passage names its file, its class (using the same tag component the Knowledge panel uses, so a class looks identical wherever a reader meets it), its heading trail, its passage number, how it was found, its closeness, and its token cost. Rows also appear for passages that were suppressed as duplicates and for passages there was no budget for — so "why is that not here?" has an answer rather than a silence.

Click-through is implemented. The row carries source_id, so the filename is a button that opens the Knowledge panel on that source rather than on a list. M7's note in §57 recorded this as not built; it is built now.

Memories show authority ("Something that happened" vs "Something inferred"), source turn, and closeness. The summary section reads its text from the assembled story_summary section rather than duplicating it into the API.

Hidden narrator information

This is the surface §38's requirement actually lands on, and it took two corrections to get right.

First: each narrator-only passage's text is replaced by "Hidden — this passage is narrator-only" until the reader ticks a control that says, in words, that it will reveal secrets they marked for the narrator alone. The control is off by default, is not remembered across reopening the panel, and is reset when the inspector is pointed at a different turn.

Second, and this was a real leak: the assembled prompt at the bottom of the panel contains the same passage verbatim. A closed <details> still holds its contents in the DOM, where find-in-page reaches them — so the secret was one keystroke away from a reader who never touched the reveal control. The sentinel test caught it only because it asserted on innerHTML rather than innerText.

Hiding it visually would have been the appearance of a guard rather than a guard. The assembled prompt is now withheld entirely while narrator-only material is in it and the reader has not asked, with a note saying so. Three component tests cover it, and the leak test was verified to fail against the unfixed component before the fix was kept.

The §38 documentation correction

§38 asked for a Show Hidden Story State control on the state panel. There is no hidden story state: a secret lives in a narrator-only knowledge source and never enters the state document — M7 verified that against a real narrator. A toggle there would reveal nothing, and building a hidden-state dimension to give it something to reveal would create precisely the duplicate representation the requirement exists to avoid.

The section is rewritten as Hidden Narrator Information / Spoilers. It now states the protection as five requirements, forbids inventing a second store, and records where the surface actually is. §100's checklist entry follows it. This is a requirement clarification that tightens the requirement, not a weakening — see §T.


L. Settings and model status

The carried debt, and what closed it

Settings.model could be empty with nothing saying so, and play then failed on the first turn with a provider error. That is the debt §8 of the brief names.

The header now reports Ollama in five states, from one shared answer fetched on mount and on demand — never on a timer, because a status line that re-tested every few seconds would be a polling loop against the reader's own inference host:

State Meaning The way out
checking the test has not come back —
ready reachable, and the chosen model is installed —
no-model reachable, no narrator model chosen pick from the models actually installed there
missing-model reachable, chosen model not installed pull it, or pick one it has
unavailable not reachable local troubleshooting steps

no-model and missing-model are separated because the fix differs: choose one from a list you already have, versus pull one that is not there.

Nothing is chosen automatically. Picking the first model in a listing would silently narrate with whatever sorted first — possibly an embedding model, which cannot narrate at all. A component test asserts that rendering the notice writes no setting. When an embedding model is in the list, the notice says plainly that a model with "embed" in its name is not for narrating.

An endpoint that lists nothing is not treated as evidence the model is missing — some servers answer /models with an empty body, and claiming the model is absent would send the reader to pull one they already have.

Model selection

Model was a free-text box; typing a name that is not pulled produces a correct-looking configuration that fails on every turn. It is now a picker over the connection test's own listing, with the free-text field kept for the case where the listing is empty or the reader wants a model not yet pulled.

Local-only, and no cloud provider

Settings opens with a plain statement: campaigns, imported files and every prompt stay on this machine; the one thing that leaves is the request to the Ollama endpoint, which must be on this machine or this network. No provider selector, no sign-in, no API key field.

Two backend strings were the last user-facing text describing a hosted deployment: a 401 advised checking an API key (removed in M2 with the cloud providers), and a 429 explained a shared free tier's daily cap. Both sent a reader looking for a setting that does not exist. Rewritten to describe what Ollama's own 401 and 429 mean.

Diagnostics kept, and folded away

The raw request/response log is genuinely useful when a local model misbehaves and is exactly the "developer console" §26 wants out of the normal path. It stays, at the bottom, behind a disclosure, and loads only when opened.

Guidance where the reader actually is

The browser run found that a reader who opens the story screen with a dead endpoint got a disabled Send button and a red badge, and no explanation — the setup notice lived only on Settings and the setup form. A disabled button with a red badge is a puzzle. The story screen now carries the same shared notice when play is blocked. Asserted in the A05 suite.


M. Accessibility and browser verification

Asserted, in a11y.test.jsx

Check Method
Every control has an accessible name of more than two characters walks the composer, library and setup form, collecting failures — the M7 baseline had 16 unnamed glyph buttons on a two-turn story
Every form field has a bound <label> over every input, textarea and select in setup
One <h1>; campaigns are a real list ul.campaign-grid > li
A campaign opens by a link, not a click handler on a <div> every route to it has an href
The setup form is a <form> with a submit button so Enter submits
Skip link, one <main>, a named <nav> on the shell
Tab reaches the story controls in visual order Continue → Retry → Undo → Redo → Save Point → direction → the box
The disabled dictation control never takes focus tabbed twelve times

Dialogs, in Dialog.test.jsx

Focus moves in on open; Tab wraps at both ends; focus returns to the opener on close; Escape closes; role="dialog" with aria-modal and the title as its accessible name; the typed-name confirmation rejects a near miss.

The focus-trap regression is preserved and is finding 5. The trap filtered candidates with offsetParent !== null, which is null inside a position: fixed ancestor — which the dialog is — and which jsdom never computes. It would have behaved differently in the tests from the browser, which is the one thing a focus trap must not do.

Transcript structure

role="log" named "Story transcript", each turn an <article> labelled "What you did" or "The story". Per-turn controls are revealed by :focus-within as well as :hover — they are real buttons in the tab order, and if only hover revealed them a keyboard reader would be operating controls they cannot see.

The take counter shows 2/3 and reads "Take 2 of 3".

Keyboard

Enter and Ctrl/Cmd+Enter both send; Shift+Enter inserts a newline. The inherited global Ctrl+Z / Ctrl+Shift+Z / Ctrl+R shortcuts were removed: §28 warns against stealing the browser's and the text editor's own Undo, and the inherited guard only excluded the focused element being a text field — pressing Ctrl+Z anywhere else undid a story turn instead of a text edit.

Checked by eye, not asserted

Contrast, visible focus rings and reading order at 1440×960, from the screenshot pass. Recorded as what it is — a look, not a measurement. A WCAG contrast audit is not part of M8 and belongs to M11.

Reader-facing terminology audit (§9)

Method. Comments, imports and identifiers were stripped from each component, leaving only text that can reach the screen: JSX text nodes, and the aria-label, title, placeholder, label, confirmLabel, cancelLabel and requireLabel attributes. Thirteen terms were searched on word boundaries — branch, branches, fork, node, head, lineage, embedding, vector, database row, action id, branch id, head depth, depth — across the seventeen components a reader operates in ordinary play.

Result: 0 hits. The advanced surfaces — the context inspector and Settings, where §55 and §72 explicitly permit developer evidence — were audited separately and also returned 0.

The audit's one finding on its first run was the "embedding model" wording, which was a real usability defect rather than a vocabulary breach (finding 6).

Browser observation. The acceptance run asserts the same property from the other side, against the live DOM: it clones the page, removes the narrator's own prose, and searches the interface's own text for the same terms. That exclusion matters — the first version failed on "head" because the model wrote "You head north, the Abbey looming ahead", which is the narrator's English, not the product's vocabulary.


N. The frontend test foundation

The project had never had a frontend test. npm run lint && npm run build was the whole check, and two backend tests read JSX as text to approximate browser assertions — both saying in their own docstrings that they stood in for a runner M8 would supply.

Stack: Vitest 3.2.7 over jsdom 26.1.0, with @testing-library/react 16.3.3, @testing-library/user-event 14.6.7 and @testing-library/jest-dom 6.9.1. Five devDependencies, no runtime dependency. Chosen because the project is already Vite — the runner shares the build config rather than introducing a second one — and because jsdom keeps the suite fast enough to actually be run.

npm test (once) and npm run test:watch, both documented in DEVELOPMENT.md.

npm audit

Four high-severity advisories were reported. All four pre-date M8 — verified by auditing the M7 commit's own package.json and lockfile in isolation, which reports the same four. npm audit fix resolved them within semver, and the tree now reports 0 vulnerabilities.

Two lockfile-affecting changes

The five test devDependencies, and the audit fix. Both are recorded here because §30 asks for it.

Defects the tests caught

Writing them found real problems, which is the argument for having them:

  1. The dialog focus trap used offsetParent, null inside the fixed-position ancestor the dialog has and never computed by jsdom (finding 5).
  2. The assembled-prompt spoiler leak — caught because the sentinel test asserted on innerHTML, not innerText (finding 3).
  3. The class label was inconsistent between the context inspector and the knowledge library — lowercase in one, capitalised in the other.

And two defects the tests themselves had

Recorded because a test that passes against broken code is worse than no test:

  • The first take-stepping regression passed against the unfixed component, because its fixture supplied an action.branch_id the real payload never sends. Corrected; re-run against the reverted component; failed 2 of 9 for the right reason; fix restored. helpers.jsx now carries a docstring saying why that field must never be added back.
  • test_m8_setup_surface.py passed alone and failed ten ways in the full suite, because it wrapped TestClient(app) instead of following the suite's create-schema / drop-schema fixture convention.

It does not replace the browser runs

jsdom has no layout, no navigation, no network and no CSP. Scroll behaviour, streaming, a genuine process restart, request counting and every security property that depends on the browser actually fetching something are outside its reach. Both kinds of evidence are recorded, and neither substitutes for the other.


O. Backend scope, schema, and migration

Scope

M8 is a browser milestone. Five application files changed, and every change exists to expose a browser capability that could not otherwise be reached.

backend/app/models.py                       16 +   a read-only property, no column
backend/app/schemas.py                      25 +   opening, canon_rules
backend/app/routers/adventures/crud.py      53 +   build the opening; write canon
backend/app/routers/adventures/turns.py     32 +   the player-input normalizer
backend/app/providers/openai_compatible.py  28 +-  two cloud-era error strings
Check Evidence
No schema change git diff 480414e -- backend/app/migrations.py → 0 lines
No column added or removed mapped_column additions/removals in models.py → 0
No listener change git diff … backend/app/main.py → 0 lines
No endpoint-policy change git diff … backend/app/endpoints.py → 0 lines
No new outbound network call no added httpx. / requests. / urlopen / socket. line
No auth/account/cloud path the only added "API key" strings are the comment explaining its removal
No M9 export/recovery subsystem bundle code untouched
No M10 media provider none

The one models.py change is a @property that reads the rules list out of the existing campaign_canon JSON document. It stores nothing.

Migration proof

M8 claims no migration, so the claim is proved rather than asserted. A database created by today's code and read by today's code would prove nothing — so the first server runs from a git worktree checked out at the M7 commit, and the campaign on it is built by M7's own code: story, an alternate take, a Save Point, a manual state correction, memories, and imported knowledge including a narrator-only source. The M8 code then opens that same file.

scratchpad/m8/migration.py. Results are in the acceptance matrix (§P).

One note on that script: its first version made a correction the validator refused with a 400 — a subject that did not resolve — and then asserted that the correction "survived". It was asserting on something that had never existed. It now asserts its own precondition and fails loudly if the setup step does not take. That is finding 10.


P. Acceptance matrix

Every M8-owned acceptance test, with the evidence that decides it. All browser evidence is from the final build against a real narrator; the environment is in §Q.

PASS means demonstrated through the browser workflow. §33 is explicit that an earlier milestone's API-level pass does not carry a D-test, so none is claimed on that basis.

B series — story input, against a real narrator

Test Result Evidence
B01 Natural-language action PASS "I walk into the Crooked Lantern and look for Mara." typed into the one field; narration coherent and using the established setting (Mara / lantern / tavern asserted)
B02 Dialogue input PASS "I say to Mara, "Have you heard anything about Edrin?"" — treated as the protagonist speaking; the narrator answers about Edrin rather than narrating a contradictory action
B03 Continue PASS Continue pressed with nothing typed; narration produced, no protagonist decision invented
B04 Story direction PASS "Keep this scene tense, but do not start a fight yet." sent through the explicit toggle; the direction is not narrated as spoken dialogue, and the affordance says it is out-of-character

D series — history controls, editing, Save Points

Test Result Evidence
D01 Undo one turn PASS Undo offered and enabled; transcript stepped back; State panel followed
D02 Minimum five undos PASS Two asserted in the browser; the backend suite covers five and beyond (test_head_cursor.py)
D03 Undo to the campaign opening PASS Undo at the opening is refused rather than emptying the story (HTTP 400), asserted in test_m8_setup_surface.py
D04 Redo PASS Redo became available after Undo, and moves the story forward
D05 Redo invalidated by a new continuation PASS After writing a different action, Redo into the abandoned future is no longer offered
D06 Retry narrator response PASS Retry produced a different narration for the same input; a take selector appeared reading 2/2 and "Take 2 of 2"
D07 Select prior retry take PASS Stepping back shows take 1, distinct from the take just generated. This is finding 1 — it did not work before this milestone
D08 Retry does not delete the prior take PASS After stepping back, Next is enabled: the later take is still there
D09 Edit earlier player input PASS Warning explains a different continuation and that later story is kept; editor holds the new text; the edited action reaches the active story. This is finding 7
D10 Edit narrator output PASS "Mara wears a green cloak" becomes what the story tells; the reader is told state may be recalculated; M5's fork path is used
D11 Named Save Point PASS Created with a name, counted in moments, and survived a genuine process restart
D12 Restore Save Point PASS Confirmation says later story is kept; restoring moved the story back; continuing from there works
D13 Restore does not delete later history PASS Retained on the server after the restore, asserted against the API
D14 Delete Save Point PASS Confirmation says it deletes no story; moment count unchanged after deleting

The six UX scenarios (§31)

Result Evidence
A — normal play PASS Library → open → read → type → streamed narration → continue, with no advanced panel opened
B — bad narrator response PASS Retry, take 2, step back to take 1, step forward, continue — no branch vocabulary at any point
C — player mistake PASS Two Undos, State followed, different action written, Redo correctly unavailable, abandoned turns absent from the transcript and retained on the server
D — major decision PASS Save Point created mid-story, genuine process restart, restored, continued differently
E — continuity bug PASS C04's correction made through the panel without raw JSON; reached authoritative state as manual_correction; the next prompt respects it
F — strange narration PASS Inspect context → readable account → source named → click through to the source → disabled → no longer participates

Other required checks

Result Evidence
A05 failed generation PASS Real dead endpoint mid-session; §Q
§34 restart consistency PASS Genuine process restart; reload; Undo/Redo; take change; restore; correction; knowledge toggle; failed generation
H06 stored XSS PASS §Q
H07 JavaScript URL PASS §Q
G09 remote image not auto-loaded PASS §Q — asserted against resource timings, not markup
G10 injected instruction inert PASS §Q
§37 offline PASS §Q
§38 hidden information PASS Sentinel through the real narrator-only path; §K
§43 migration PASS 17/17 against a database built by M7's own code; §O
§28 accessibility PASS §M
§30 frontend tests PASS §N
§38/§9 terminology PASS 0 hits, source and live DOM; §M

Browser suite totals, final build

Suite Checks Result
Acceptance — 6 scenarios, B01-B04, D01-D14 57 57/57
Hidden-information sentinel 21 21/21
A05 failed generation 23 23/23
Security / offline / performance 27 27/27
Genuine process restart 12 12/12
Migration against an M7-built database 17 17/17
Total 157 157/157, 0 failures

Which build every figure above comes from

Two complete evidence sets were produced against two different frontend artifacts. They are not interchangeable, and the difference is what finding 7 fixed.

Superseded set Final set
Frontend artifact index-C6E5Uvtu.js index-Ii-lARp9.js
Built before finding 7's fix 18:53:02, from the staged tree
Ran 18:28-18:54 18:55-19:55
Acceptance 54/55 — one FAIL 57/57
The failing check D09: the edited action is on the active story —
Cited in this report no figure every figure

The superseded set is the evidence that findings 3 and 7 existed: the spoiler leak and the D09 routing defect were both discovered by it. Its acceptance suite recorded 54/55 because D09 was genuinely broken at that build. That is why it cannot also be the final evidence — the artifact changed to fix it, and an evidence set is only evidence for the build it ran against.

The final set's acceptance count is 57 rather than 55 because finding 7's fix brought two new regression checks with it (D09's continuation and D10's re-fork), which the superseded build had no way to satisfy.

The freeze held across the final set. index-Ii-lARp9.js was written at 18:53:02 and no tracked file under backend/app/ or frontend/src/ has a modification time after it — verified by find -newermt over git ls-files, which returns nothing. The acceptance suite was run three times inside that window (ending 19:06, 19:33 and 19:55) against that one unchanged artifact; the first two ended on the harness defect recorded as finding 12's fifth item (a retry assertion demanding the model produce different words), and only the harness script — which lives outside the repository — was edited between them. No application source was touched, so the final 57/57 is a result about the same build as the other five suites.

One partial result is explicitly discarded: the earlier migration run was stopped during its last check while cleaning up for the final set, so it recorded 16 of 17. It is not cited. The final set re-runs it complete.


Q. Security, offline, and restart

Environment for every browser suite

Browser Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1)
Application the production build dist/assets/index-Ii-lARp9.js — the final frozen artifact, served by the real backend as static files, not a dev server
Backend uvicorn app.main:app bound to 127.0.0.1, one process per suite
Database a real SQLite file on disk, one per suite
Narrator qwen2.5:3b-instruct on the trusted-LAN Ollama at inference.lan — real inference for every story turn
Embeddings nomic-embed-text, real
Fixture the Continuity Test campaign from TEST-CAMPAIGN-FIXTURE.md

Deterministic rather than real inference was used in exactly one place: the hostile-content narrator turn in the security suite, whose text is written directly into the row. Asking a model to reliably emit a <script> tag is unreliable, and the property under test is the renderer, not the model. Every other narration in every suite came from the real model.

H06 / H07 / G09 / G10

All asserted from the browser, not from source.

A <script> in narrator prose did not execute (window.__narr undefined); no <script> element; visible as text; document.title unchanged
The same payload imported and read in the library shown as text, created no element, did not run
javascript: URLs never became an href, in transcript or panels; every rendered link was then clicked and the injected global stayed undefined
Remote images no <img> created; nothing fetched from the remote host, per the browser's own resource timings; the reader is told an image was blocked
Prompt injection reaches the prompt as data, framed as untrusted with the authority rule stated — suppressing it would be the wrong fix
§74 external links a real https:// link does not navigate; a dialog names the destination in full first

The CSP is unchanged from M7 and still default-src 'self' with img-src 'self' data:, so a remote image would be blocked at the browser even if the renderer produced one. Two independent layers.

§37 offline

Every request the page made was to this origin, taken from performance.getEntriesByType('resource') after exercising the transcript, the panels and the knowledge inspector — network observation, not "no error occurred".

The built assets were also checked for remote resource references — src=, href=, url(, fetch(, import( — rather than for any URL-shaped substring. That refinement was itself a finding: the first version flagged https://react.dev and https://reactrouter.com, which are documentation links inside those libraries' warning strings and are never fetched. Leaving it would have pressured a maintainer into silencing a true statement. The runtime check is the authoritative one and passed either way.

test_offline_assets.py, test_local_only_surface.py and test_endpoint_policy.py all pass against the final build (73 tests).

A05, with a real provider failure

The endpoint is pointed at a permitted loopback address with nothing listening — a real ConnectError through the real provider and the real turn engine, not a mocked exception, and not a disallowed endpoint (which the policy would refuse before any request, testing the wrong thing).

The ordering matters and is itself a result: breaking the endpoint before loading the page produces no failure at all, because Send is correctly disabled when the status probe reports unavailable (§12 working). The realistic A05 case is an endpoint reachable when the reader pressed Send and unreachable when the request went out, so the suite breaks it mid-session without reloading. Both the guard and the failure are asserted.

Genuine process restart

The server process is stopped and a second one started against the same database file. Recreating a test client is not equivalent: a Save Point that survived only because a Python object was still alive would pass that and fail a user's restart.


R. Performance and regression against the M7 baseline

M8 makes three before/after claims. Each is measured on both sides.

M7 baseline M8 Method
Context inspector, text on open, same campaign and fixture 11,996 chars ~1,400 chars innerText of the panel, real browser, two real turns played
Per-message controls without an accessible name 16 of 16 0 walks .turn-tools; a name under three characters counts as a glyph
Implementation vocabulary in reader-facing text Branches tab present; Do/Say/Story modes 0 hits across 17 normal-play components source audit on visible strings + live-DOM check
CSS shipped 65.91 kB 46.98 kB vite build

Request behaviour (§38)

M7 found and fixed a refetch-on-every-keystroke defect. The class has not returned, and the panels got cheaper:

Interaction Requests
Typing a 60-character sentence 0
Typing with a knowledge source inspector open — the exact M7 regression 0
Sitting idle for 8 seconds 0
Opening State / Context / Save Points 1 each
Opening campaign Settings 0

Campaign Settings costs nothing because it renders the campaign object the page already holds. No panel does an N+1 read.

Nothing polls. Model status is fetched on mount and on demand, never on a timer — a status line that re-tested every few seconds would be a polling loop against the reader's own inference host, which is precisely what §38 looks for.

Two design choices carry this, and both are deliberate:

  • Turns are memoized. The page re-renders on every keystroke; without memo every turn in the window would re-parse its Markdown along with it — the M7 defect's class, reached from a different direction.
  • Campaign settings save on an explicit press, not a debounce. A debounced save writes on every pause in typing, which is a request every few keystrokes against a field a reader edits for a minute at a time.

Bundle size

JS grew from 423.28 kB to ~387 kB — smaller, despite the added surfaces, because sixteen components and eight stylesheets left with the screens they served. Gzipped: 126.86 kB → ~118 kB.


S. Findings

Every finding numbered, with severity, evidence, whether it is fixed, the exact correction, and its regression coverage. Fixed defects are listed because they are evidence about what the M8 validation actually discovered.


1. Stepping between alternate takes did nothing — HIGH

Evidence. TakePager.step() decided whether a take lived on another line with target.branch_id !== action.branch_id. target comes from the variants endpoint, which returns branch_id; action is an ActionOut, which has never carried branch_id — not in M8, not at M7 (git show 480414e:…schemas.py confirms it). So the comparison was permanently number !== undefined, always true, and every step took the branch-switch path. For two takes of an ordinary retry — which share a line until one is written below — that meant asking the server to switch to the line already being read: the same window came back and the step did nothing.

Impact. D07 (Select Prior Retry Take), REQUIRED FOR V1, could not be performed in the browser. No test caught it: there was no frontend suite, and M7's browser run exercised the pager's presence but never a step between two same-line takes.

Fixed. No new API field. The variants list already carries every attempt's branch and marks the live one, so the comparison is made within data the endpoint returns: list.find(row => row.active).

Regression. Two component tests (same-line preview; forked-line switch), plus D07 in the browser suite.


2. A05: the reader's words were not restored on the commoner failure path — HIGH

Evidence. A generation failure does not throw. The server reports it as an SSE error event and the request then completes normally, so runTurn's catch never runs. Restoring the typed text and re-testing the model both lived in that catch alone. The browser run showed the failure notice reading "what you typed is still in the box" with the box empty (''), and the header still reporting ready while the notice said the model was unreachable.

Impact. A05 unmet on the path readers actually hit, and the UI made a false statement to the reader — worse than saying nothing.

Fixed. One failTurn handler, called from both the SSE error branch and the catch. The typed text rides in a ref so the SSE path can reach it without re-creating handleEvent on every keystroke.

Regression. failurePaths.test.jsx (4 tests) pins the notice's claim as a contract and asserts both paths classify identically; the browser A05 suite asserts the restored text and the updated header directly.


3. The assembled prompt leaked narrator-only text — HIGH

Evidence. The passage rows correctly withheld the secret, but the assembled prompt section at the bottom of the inspector contains the same text verbatim. A closed <details> still holds its contents in the DOM, where find-in-page reaches them. The sentinel test caught it only because it asserted on innerHTML rather than innerText.

Impact. §38's protection was the appearance of a guard rather than a guard. A reader who never touched the reveal control was one keystroke from the secret.

Fixed. The assembled prompt is withheld entirely while narrator-only material is in it and the reader has not asked, with a note saying so.

Regression. Three component tests, verified to fail against the unfixed component before the fix was kept; plus the sentinel browser suite (21 checks).


4. > You I enter the tavern. — HIGH

Evidence. AI Dungeon's player-input convention prefixes > You . §11 tells the reader to write "I enter the tavern." The result reached the transcript, the replayed history and the narration, where a small model imitated it: the M7 baseline transcript contains "You Aldric watches Mara closely".

Impact. Every player turn, and degraded narration.

Fixed. The subject is added only when the reader has not written one. The > marker is unchanged in every case.

Regression. test_first_person_input_is_not_prefixed_with_you against the shared normalizer (§11's examples verbatim, eight first-person phrasings, the legacy bare action, dialogue and direction), plus test_normalized_player_text_reaches_history_exactly_once, which asserts on the assembled prompt that "You I " appears nowhere and each player line appears once.


5. The dialog focus trap used offsetParent — MEDIUM

Evidence. offsetParent is null for anything inside a position: fixed ancestor — which the dialog is — and jsdom never computes it. The trap would have behaved differently in the tests from the browser.

Fixed. Filters on hidden / aria-hidden instead. Regression: Dialog.test.jsx wraps Tab at both ends, in both directions.


6. The Knowledge panel named a Settings field that does not exist — MEDIUM

Evidence. The terminology audit found the panel telling readers to choose an "embedding model"; the Settings field is called Model for meaning-based search. A usability defect, not vocabulary policing.

Fixed. Both states name the field as the reader sees it. Regression: two assertions that the panel's text contains no "embedding".


7. Editing a player turn with story after it did not create a continuation — HIGH

Evidence. beginEdit routed every "Edit" through PATCH. On a player row update_action rewrites in place and re-evaluates nothing, and refuses outright when displaced history hangs off it — neither creates the new continuation D09 requires. Meanwhile the confirmation promised "this will continue the story differently… everything you have already read is kept". The browser run showed the dialog appearing, the editor holding the new text, and the edited action never reaching the story.

Impact. D09 unmet, and the dialog's words disagreed with the behaviour.

Fixed. A player turn with story after it takes the replay path (addTake, SP9), which is what creates a continuation. A narrator turn keeps PATCH, because _edit_narration already forks per §§14-15.

Regression. editRouting.test.jsx pins the routing rule for both kinds and both conditions; D09/D10 in the browser suite.


8. A stale .input-bar { display: flex } — HIGH (cosmetic, total)

play.css is imported after story.css, so a rule left from the old composer won on order: the direction row and the input row laid out side by side and the box was unusably narrow. Found by opening the product, not by reading the CSS. Fixed by removing the superseded rules.


9. Presentation defects found in the first browser pass — MEDIUM/LOW

The sticky header was translucent and story text read through it; the model badge overlapped the panel tabs; the composer box carried forms.css's 70px floor; the stored > marker rendered as a Markdown blockquote by accident, stacking rules. All fixed; the marker has two component tests.


10. Defects in the tests and harnesses themselves — process

Recorded because a test that passes against broken code is worse than no test.

  • The first take-stepping regression passed against the unfixed component: its fixture supplied an action.branch_id the API never sends. Corrected, re-run against reverted code, confirmed failing 2/9, fix restored. helpers.jsx documents why that field must never return.
  • test_m8_setup_surface.py passed alone, failed ten ways in the full suite — it wrapped TestClient(app) instead of following the suite's create/drop fixture convention.
  • The D09/D10 browser harness set .value on a React-controlled editor, so the "edit" saved its original text. Now uses the WebDriver clear endpoint and asserts the editor holds the new text before saving.
  • The migration harness made a correction the validator refused with a 400, then asserted it "survived". It now asserts its own precondition and fails loudly.
  • Two acceptance assertions were wrong about established semantics: a failed generation does keep the player's action (M3), and the vocabulary check matched "head" inside the narrator's own prose ("You head north").

11. Evidence-run discipline — process

Three browser evidence runs were invalidated: two by rebuilding mid-run, one by starting a second run before the first had exited (both wrote to the same files, producing results that could not all be true). No product conclusion was drawn from any of them — each was discarded and re-run. The rule this settles on: one frozen build, one suite at a time, no source edits between the freeze and the last suite; if any is broken, the run is discarded rather than reported.


12. The evidence harness needed five corrections before it could be trusted — process

Recorded as a finding in its own right, because it bears on how much of M8 was genuinely verified rather than assumed. Across the milestone the browser harness had five defects, against seven product defects:

The harness said The truth
.turn:last-child "no take selector appeared" the last child of .story is the scroll anchor, so the selector could never match
streaming counted as a turn "the composer is not reachable by keyboard" the turn had not finished; the composer was correctly still disabled
.value on a React editor "the edited action is not on the story" true, but for the wrong reason — the edit had saved its original text
a branch_id the API never sends "the take-stepping fix is covered" the test passed against the broken component
Retry must produce different words "no second take appeared" the take existed; the model had simply repeated itself

Two of those five masked real product defects (findings 1 and 7) and were the reason each took two rounds to find. Three asserted properties the product never promised.

The lesson is not that the harness was careless — it is that a browser assertion is only as good as its model of what the product guarantees, and four of the five failures were the harness asserting model behaviour, DOM structure or React internals rather than a product contract. Every corrected assertion now names the guarantee it is testing.


13. Three evidence runs were invalidated by process error — process

Two by rebuilding the frontend mid-run, one by starting a second run before the first had exited — both wrote to the same result files and produced a set of figures that could not all be true at once.

No product conclusion was drawn from any of them. Each was discarded and re-run. A fourth partial result — a migration run stopped during its final check while cleaning up — was likewise discarded rather than reported as complete, after it was noticed that the script defines 17 checks and the output held 16.

The rule this settles on, and which the final set followed: one frozen build, one suite at a time, no source edits between the freeze and the last suite; if any of those is broken, the run is discarded rather than reported. The cost of ignoring it is not a slow run — it is a plausible-looking number that is not evidence of anything.

Invalidated is not the same as superseded, and this report uses the two words precisely. An invalidated run is one whose figures cannot be trusted at all, because the build moved underneath it or two runs overwrote each other's output: the three above, plus the partial migration result. A superseded run is internally sound but describes an artifact that no longer exists — the index-C6E5Uvtu.js set, whose 54/55 acceptance result is a true statement about a build that finding 7 then corrected. The superseded set is cited in §S as the evidence that findings 3 and 7 existed; it is cited nowhere as evidence that M8 passes. The final set, against index-Ii-lARp9.js, is the whole of §P.


14. The context budget the app plans against was four times the ceiling the server enforced — RESOLVED, operationally, with no application-code change

Status: found in M8, investigated at closeout, resolved. Not an M8 regression and never triggered in M8; recorded in full because M8's browser work is what surfaced it, because the mismatch would have caused a silent correctness failure in a long campaign, and because the resolution is an operator procedure that has to be written down somewhere a deployer will find it.

The original mismatch. The application budgets a prompt up to Settings.context_token_budget — 16,384 by default. Ollama, seeing no VRAM on the reference deployment, enforces a 4,096-token input window. Nothing in the application knew about the gap, and nothing in the interface would have shown a reader that the front of their prompt was being dropped.

Why sending num_ctx through the OpenAI-compatible endpoint does not work. The application speaks /v1/chat/completions (ADR 002, ADR 011). That endpoint accepts num_ctx — nested in options or at the top level — returns HTTP 200, and ignores it. Worse, it reloads the model at its own default, so a native /api/chat call that had already established a larger window is undone by the application's next request. There is no per-request path to a larger window from where this application stands.

The measured resolution: a derived model. Creating a model with the parameter baked in makes the window travel with the model rather than with the request, and it is a normal API call — no shell access on the Ollama host, no environment variable, no restart. The derived model then appears in /v1/models, which is exactly the listing M8's Settings model picker already reads, so the application discovers and uses it through its ordinary path. No Adventure Storyteller code change is required, and none was made. The measurement table and the procedure follow.

Evidence. On the reference deployment, Ollama reports:

msg="vram-based default context" total_vram="0 B" default_num_ctx=4096

and /api/ps confirms the loaded model carries context_length: 4096.

Meanwhile the application's own budget is:

Settings.context_token_budget default 16,384
the value M8's test settings used 8,000
what the server will actually accept 4,096

The application sends max_tokens, which caps output. It never sends num_ctx, which is what sizes the input window.

Why it matters. llama.cpp truncates the oldest tokens. This application puts the narrator rules and campaign canon in the system block at the front of the prompt (context/builder.py). So the first material dropped on a long campaign is the highest-authority material in the product — the rules a turn may not contradict. It would present as the narrator quietly forgetting canon deep into a session, with nothing in the interface indicating why, and C01 would begin failing for a reason no UI surface explains.

Not triggered in M8. The largest assembled prompt across every suite was 2,498 tokens, comfortably inside 4,096. Nothing in this report's evidence was truncated. Verified by reading the token totals the context inspector reported on every run.

How this was measured. Four paths were exercised against the reference endpoint, each followed by reading context_length back from /api/ps:

POST /v1/chat/completions  {..., "options": {"num_ctx": 8192}}  -> 200, context_length 4096
POST /v1/chat/completions  {..., "num_ctx": 8192}               -> 200, context_length 4096
POST /api/chat             {..., "options": {"num_ctx": 8192}}  -> 200, context_length 8192
POST /v1/chat/completions  (ordinary, after the native call)    -> 200, context_length 4096

The fourth line is the decisive one: the native window does not persist for the application's own requests.

The OpenAI-compatible endpoint does not merely ignore num_ctx — it reloads the model at its own default, discarding a larger window a native call had already established. So priming the server through the native API is useless to this application: its next request resets it.

A derived model does survive, and needs no server shell access. The window travels with the model rather than with the request:

POST /api/create {"model":"qwen2.5:3b-instruct-16k",
                  "from":"qwen2.5:3b-instruct",
                  "parameters":{"num_ctx":16384}}

Verified end to end on the reference deployment: after creating it, an OpenAI-compatible call to the derived model — the application's own request path, not a native one — loads at context_length 16384, and the model appears in /v1/models, so it is selectable in M8's Settings picker with no code change. It shares the base model's blobs, so it costs a manifest.

qwen2.5:3b-instruct-16k is the model the demonstration created. It is a demonstration, not a dependency. Nothing in the application, the test suites or the repository requires it to exist, no evidence in this report was produced with it, and it may be deleted with POST /api/delete at any time. The durable artifact is the reproducible procedure, which is documented in DEVELOPMENT.md under "The context window your Ollama actually enforces".

This is the remedy where the server's environment is not editable, and it is an operator action rather than an application change: the provider stays OpenAI-compatible (ADR 002, ADR 011) and no second request path is introduced. No application code was added to work around this, deliberately — a workaround in the provider would mean either a native-API second path, breaking ADR 011's single-endpoint policy, or a request parameter the endpoint provably ignores.

The three paths open to a deployer, and which one this resolves on.

  1. A derived model, over the API — no shell access, no code change. ADOPTED. POST /api/create with parameters: {"num_ctx": 16384}, then select it in Settings. Verified working through the application's own OpenAI-compatible path. Costs roughly 4× the KV cache. OLLAMA_CONTEXT_LENGTH=16384 on the server does the same job where the environment is editable.
  2. Set the app's budget to match the server. The field already exists and is editable in M8's Settings screen ("How much story to send"), and its help text already says "Must fit your model's window." What is missing is any way for the reader to know what that window is.
  3. Detect and warn. The real window is discoverable — /api/ps and /api/show both report context_length — but only over the native API. A Settings-screen check that reads it and warns when the budget exceeds it would close the gap without moving the generation path.

Final status: resolved, and no longer an M11 blocker. What remains is operational, not a correctness defect: a deployer whose Ollama enforces a small window must either raise it by the documented procedure or lower Settings.context_token_budget to match. Option 3 above — a Settings-screen check that reads the real window and warns — stays on the table as a usability improvement, unowned and unscheduled, not as an open defect. M11 should confirm the deployment it certifies against has a window large enough for a 100-turn campaign before it starts; M9 should be aware that an imported long campaign reaches the ceiling immediately on a deployment that has not applied the procedure.


T. Planning-document corrections

Five documents changed. Each is classified, because "the plan changed" and "the plan was wrong about what exists" are different claims.

Requirement clarification (1)

BROWSER-UX-SPEC.md §38 — rewritten from "Hidden Narrator State" with a Show Hidden Story State toggle to "Hidden Narrator Information / Spoilers".

The protection required is not weakened; it is stated more strictly. §38 now requires, as a list:

  • ordinary State and story surfaces must not expose narrator-only information;
  • an advanced inspection surface that can contain hidden Canon or narrator-only context must withhold it by default;
  • revealing it requires an explicit, clearly labelled user action;
  • the UI must warn that doing so may reveal campaign secrets;
  • no second hidden-state store or duplicate representation may be introduced solely to give the UI something to toggle.

The last clause is new and is a tightening. What changed is that the section no longer names a state subsystem that does not exist: a secret lives in a narrator-only knowledge source and never enters the state document. §100's checklist entry follows, from "hidden-state inspector" to "spoiler-aware advanced inspection of hidden narrator information".

Ratified. The independent review accepted the change from "Show Hidden Story State" to the architectural requirement on hidden narrator information, and the closeout brief settled it. It is recorded as a requirement clarification aligned with the implemented architecture, not as a relaxation: the obsolete wording named a hidden narrative-state subsystem that does not exist in this product, and the five requirements above bind the surfaces that genuinely can expose a secret — including the clause forbidding a duplicate hidden-state store invented merely to give the UI something to toggle. The protection is verified end to end by the sentinel suite (§P, 21/21), which proves among other things that withheld material is absent from the DOM rather than collapsed inside it.

Implementation facts (4)

  • TECHNICAL-DESIGN.md §13.4 — the browser stayed a presentation layer: server-authoritative availability, reading is not deciding, UI state stays UI state; the failure taxonomy; and why markup is never produced from input.
  • TECHNICAL-DESIGN.md, campaign canon — canon is configuration and its provenance is the per-turn context snapshot. Measured, not assumed, with the measurements recorded.
  • BROWSER-UX-SPEC.md §12, §31, §55, §57 — four "As implemented in M8" notes, following the convention M7 established. No requirement altered.
  • BUILD-MILESTONES.md, V1-ACCEPTANCE-TESTS.md, VERSION.md, planning/README.md, planning/archive/README.md, README.md, DEVELOPMENT.md — M8's outcome, the browser-evidence notes on the B and D series, the M7 report's rotation to the archive, and the frontend test suite replacing "no frontend tests" in the inherited-debt list.

Changed product requirements (0)

None. No acceptance test was weakened to match the implementation. Where the implementation and a document disagreed, the disagreement is recorded rather than resolved in the implementation's favour — see §U's note on §38, and the action_count observation in §F.

Status discipline

VERSION.md and BUILD-MILESTONES.md describe M8 as implemented, verified, reviewed and accepted, dated 2026-09-06, and record M9 as next. Nothing marks M9 started, and no M9 scope was implemented.


U. Residual risks and deferred work

Genuine residuals only. M9-M11 planned scope is not listed here as an M8 defect.

Residual risks

  1. errors.js is coupled to backend message strings. Written as a fallback ladder, so an unrecognised message still classifies as a generation failure, still shows the server's own words, and still offers Retry — nothing is hidden when the match misses. But a backend rewording silently loses a tailored hint. The signatures are quoted beside each branch so the coupling is findable from either end. Risk: low. A degraded hint, never a hidden failure.

  2. Third-person player input still takes the > You prefix. "Aldric draws his knife" becomes > You Aldric draws his knife. It is not one of §11's three stated inputs, and it is unchanged from before M8 — a name-detector would be wrong more often than the current rule. Risk: low. M11's 100-turn run is where it would become clear whether readers write this way.

  3. A deployment's context ceiling can be a quarter of the app's budget — resolved operationally. See finding 14, now closed. Ollama's OpenAI-compatible endpoint ignores num_ctx, so the window cannot be set per request; a derived model created over /api/create carries it and is honoured through the application's own path, needs no shell access and no code change, and appears in the Settings model picker automatically. The procedure is in DEVELOPMENT.md. Residual risk: low, and operational rather than architectural — a deployer who applies neither the procedure nor a matching context_token_budget still gets silent truncation, and the application still has no way to tell them so. A Settings-screen warning that reads the real window remains an unowned usability improvement, not a defect. Nothing in M8 was truncated: the largest assembled prompt across every suite was 2,498 tokens against a 4,096 ceiling.

  4. Contrast and visible focus were checked by eye, not measured. Risk: low for a local single-user product; belongs to M11.

  5. Tablet is usable but untuned. Desktop was the stated primary target and the narrow-screen sheet is inherited. Risk: low, unless v1 claims tablet support.

Deliberately deferred, with the milestone that owns it

Owner
Story cards have no browser editor — backend and bundle keep them M9 should decide whether the bundle keeps carrying a subsystem no v1 UI exposes
The bundle still carries no context snapshots, so an imported campaign has no historical prompt provenance M9 (recorded by M7)
The RPG world state is read-only M10/M11, if a genre wanting numbers appears
Copy is per-message only; no whole-transcript copy (§77) and no story search (§78) M9 or M11; §78 says "not essential to initial v1"
No discarded-history recovery screen (§63) explicitly future work
The inert api_key column and the dual-dialect migration code the cleanup migration BUILD-MILESTONES.md already schedules

M9 handoff — what the next brief has to decide

Recorded here because M8 is the milestone that surfaced each one, and because a brief written from a stale report would let the inherited bundle format decide them by default. None of these is implemented, and none was implemented in M8.

A. Complete campaign portability. M9 must validate export, backup, import and recovery of the whole campaign, not just its text: retained story history, the active head, alternate takes, Save Points, the authoritative narrative state, the state provenance and audit data the product contract requires, summaries and memories as appropriate, imported knowledge with its provenance, and the campaign/profile/canon settings a campaign needs to resume correctly. M8 touched none of this — it changed no schema, no bundle code and no endpoint — but it is the milestone that put canon and the opening into the browser, so a campaign now carries setup a reader chose rather than a scenario supplied.

B. Historical prompt and context provenance. Every turn stores the context it was actually given, and §F leans on that: it is what proves a canon edit was not retroactive. The bundle does not carry it (recorded by M7, unchanged by M8). M9 must decide from the authoritative specification and data model whether those snapshots — or a reproducible equivalent — belong in the supported portable bundle. If they do not, the "every turn can prove what it was told" guarantee is local to the machine that played it, and that should be a stated decision rather than a side effect of the inherited format.

C. Legacy story cards. The backend and the bundle still carry story cards; M8 removed their browser surface, because the M7 knowledge library supersedes them for v1 and showing both would offer two unrelated systems for "things the narrator should know". M9 should decide explicitly which of these the v1 portable campaign does — compatibility-only carriage, migration into the M7 knowledge subsystem, preservation as inert legacy data, or another documented behaviour consistent with the active specification — and record the choice.

D. Context-window portability. See finding 14. A restored campaign may run against an Ollama whose effective window differs from the one that wrote it, so M9 must not treat campaign restored successfully as the destination narrator has the same capacity. Concretely: do not hard-code an assumption that 16,384 tokens are available; preserve the campaign and model configuration the bundle contract requires accurately, so a destination can at least be compared against the source; and note that an imported long campaign meets a small ceiling immediately on a deployment that has applied neither the DEVELOPMENT.md procedure nor a matching context_token_budget. Actual long-campaign and context-window release validation remains M11's, unless a later milestone explicitly moves it.

A note on how much of this was verified rather than assumed

Seven product defects were found and fixed; five harness defects had to be fixed before the harness could find them, and two of those five were actively masking product defects (§S findings 12-13).

A reviewer is entitled to ask what else the harness is not asserting. The honest answer is that the browser suites assert what the product promises, and the places where they previously asserted something else — model wording, DOM structure, React internals — have each been rewritten to name the guarantee under test. What they do not cover is stated in §M (contrast and focus rings were looked at, not measured) and §U (tablet was not tuned).

The one requirement question, and how it was settled

BROWSER-UX-SPEC.md §38 was rewritten, and the rewrite has been ratified. It is the one place M8 changed the wording of a requirement rather than recording a fact about the implementation, so it was put to the reviewer explicitly. The independent review accepted it, and the closeout brief settled it: a requirement clarification aligned with the implemented architecture, not a weakening. §T carries the five requirements it now states and the argument that they are stricter than what they replaced. Nothing about this is left open.


V. Final milestone assessment

Is M8's Definition of Done satisfied?

BUILD-MILESTONES.md states it as: "Normal story creation and play feels like a focused local storyteller rather than an RPG or developer console."

Yes, on the evidence in §P. A reader creates a campaign from a form that never mentions a scenario, plays through one natural-language field, and reaches State, Knowledge, Context and Save Points through panels that start closed. The words branch, fork, node, head and depth appear nowhere they operate — audited in source and observed in the live DOM.

The twenty conditions in the M8 brief's own §48 are covered by §P's matrix.

Does it feel like the intended storyteller rather than adapted AI-DnD?

The measurable parts say yes. The navigation lost Home · Adventures · Scenarios · Settings · AI Chat for Campaigns · Settings. The context inspector went from 11,996 characters to about 1,400. Sixteen unnamed glyph buttons became named controls. Three input modes became one field. The Branches tab, the tree overlay, the stat-schema editor, the art picker, the story-card table and the raw model console are gone.

The unmeasurable part — whether it reads like a storyteller — is a judgement, and it is the reviewer's to make. §E describes what a person actually meets.

Are any required browser behaviours unverified?

No. Every M8-owned acceptance condition was demonstrated in a real browser against a real narrator on the final build: 157 checks across six suites, zero failures. D09 and D10 — the two that were outstanding longest, and the two a harness defect had been masking — are verified end to end.

Three things are recorded as looked at rather than measured, and none is a required M8 condition: contrast and visible focus rings (§M), reading order at 1440×960 (§M), and tablet layout (§U). Each is named as such rather than folded into a pass.

Is it safe to proceed to M9 after review?

M8 changed no schema, no migration, no listener binding, no endpoint policy and no outbound network behaviour (§O). M9's subsystems — bundle format, backup, recovery — were not touched. Nothing in M8 constrains M9 beyond the two notes in §U that M9 should decide deliberately rather than inherit.

What must happen before M9?

Status
This report reviewed, and M8 accepted or returned done — the independent review returned M8 IMPLEMENTATION: PASS, subject to evidence and documentation cleanup
The §38 rewrite ruled on (§T, §U) done — ratified as a requirement clarification aligned with the implemented architecture
The build-evidence contradiction resolved done — §P classifies index-C6E5Uvtu.js as superseded and index-Ii-lARp9.js as the one final frozen artifact, from the saved run logs
Finding 14 dispositioned done — resolved operationally, no application-code change, procedure in DEVELOPMENT.md
The repository owner signs the M8 commit outstanding — the only remaining step

The tree is staged and uncommitted. This project's policy is that milestone commits are signed by the owner, and nothing here was committed on their behalf. After the signed commit, the expected next step is a brief verification of the committed tree, and then M9 may begin.

M8 closeout status: implemented, verified, reviewed, accepted (2026-09-06).

M9 has not been started, and no M9 scope was implemented early.