The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.
All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.
So the shape now is one entry point and one screen:
Campaigns -> Campaign -> Story
State · Knowledge · Context · Save Points · Settings
Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.
Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.
Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.
The two defects worth the space:
A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.
And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.
Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.
`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.
Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.
Backend, and only what the browser could not otherwise reach:
AdventureCreate.opening a start action could only come from a Scenario, so
every campaign made in the new setup flow opened on
a blank page. Same node, same code path.
canon_rules campaign_canon has been the highest authority in a
campaign since M5, read by the prompt builder and
the state validator, and had no API at all — a
fixture had to write it with SQL.
a 401 and a 429 message the last user-facing text describing a hosted
deployment. One told the reader to check an API key
that has not existed since M2.
No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.
The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.
They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.
A verification pass over all of it then found three more, each by driving the
product rather than reading it:
Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.
The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.
And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.
Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.
`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.
Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:
The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.
Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.
The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.
The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.
Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
92 KiB
M8 — Browser UX Completion for v1 Story Operations
Final implementation, verification and closeout report. Written by the implementer for an independent reviewer, and completed at closeout after that review returned M8 IMPLEMENTATION: PASS. The closeout additions are the build-evidence classification in §P, finding 14's resolution, and the acceptance record in §V; no verification figure was changed by them.
Closeout date: 2026-09-06.
Placeholders.
inference.lanstands for the trusted-LAN Ollama host and192.168.0.xfor LAN addresses, following the convention M1 established. No real hostname, address or identifier of the machine this was built on appears in any committed file.
A. Executive result
PASS
Every M8-owned acceptance condition was demonstrated through the browser workflow against a real local narrator, on one frozen production build. No required condition is unverified, and there are no blockers to acceptance.
Two qualifications a reviewer should weigh, neither of them a blocker:
- Seven product defects were found during verification, four of them by driving the application rather than reading it (§S findings 1-4, 7). Two made the interface state something untrue to the reader. All are fixed with regression coverage, and two of the regressions were verified to fail against the unfixed code before the fix was kept. But their existence says the implementation pass was not as verified as it looked at the time.
- The evidence harness needed five corrections before it could be trusted (§S finding 12), and three evidence runs were invalidated by process error (§S finding 13). Two of those harness defects were actively masking product defects. The figures below are from the corrected harness on the final build; the history is recorded so a reviewer can judge how much they rest on.
M8 is implemented, verified, reviewed and accepted. The independent review
returned M8 IMPLEMENTATION: PASS subject to evidence and documentation
cleanup; that cleanup is §P's build classification and finding 14's resolution,
both complete. M8 is not yet committed: the tree is staged for the
repository owner's signature, which is the only step left before M9.
What this milestone is, in one paragraph
M8 turned an adapted AI-DnD interface into the interactive-story workspace
BROWSER-UX-SPEC.md describes: one entry point, one story screen, one
natural-language input, and everything advanced one layer deeper behind panels
that start closed. Almost nothing underneath changed — the play loop, the
history controls, the takes, the Save Points, the state engine and the knowledge
library are M3-M7's and are untouched. What changed is what a reader is shown
and asked to understand.
Evidence at a glance
| Gate | Result |
|---|---|
| Backend suite | 950 passed, 14 skipped, 0 failed |
| Frontend component suite | 132 passed, 10 files |
| Lint | exit 0 |
| Production build | clean |
| Docker build | clean |
| Browser acceptance (6 scenarios, B01-B04, D01-D14) | 57/57 |
| Hidden-information sentinel | 21/21 |
| A05 failed generation | 23/23 |
| Security / offline / performance | 27/27 |
| Genuine process restart | 12/12 |
| Migration against an M7-built database | 17/17 |
| Reader-facing terminology audit | 0 hits (17 normal-play components; both advanced surfaces) |
Every browser figure is from one frozen production build,
dist/assets/index-Ii-lARp9.js, built from the staged tree at 18:53:02 and
unchanged for the rest of the campaign. An earlier artifact,
index-C6E5Uvtu.js, is superseded: it predates finding 7's fix, its
acceptance suite ended 54/55 on that defect, and no figure above comes from it.
§P sets the two side by side with the evidence for the classification.
What a reviewer should look at first
- §S findings 1, 2, 3 and 7 — four product defects that the browser work found and that no earlier milestone had caught. Two of them made the interface state something untrue to the reader.
- §S findings 12 and 13 — the harness needed five corrections before it could be trusted, and three evidence runs were invalidated by process error. This bears directly on how much weight the figures above deserve.
- §F — the
canon_rulesdecision, where the smallest correction was a surface constraint rather than an audit subsystem, and why. - §T — the one requirement whose wording M8 changed (§38), and the argument that it was tightened rather than weakened.
B. Repository and provenance state
| Branch | m8-browser-ux |
| M8 base | 480414efe082a4bfe0600a19fe23961f6bddd925 — M7 |
| M7 signature | Good signature, verified against the repository owner’s RSA key 02C9BF7D…B5C68569, ultimate trust |
| M7's parent | a6e9c7a32bdf42f1e4cb837b70721d89422cef8d (M6) — sole parent |
| Current HEAD | 480414e… — still M7. M8 has made no commit. |
| Staged | 89 files — 34 added, 24 deleted, 30 modified, 1 renamed (git diff --cached --stat) |
| Unstaged | 0 |
| Untracked | 0 |
LICENSE |
unchanged, absent from the staged diff |
| Upstream ancestry | d72f7c1b… (AI-DnD) is still an ancestor |
Committed HEAD vs staged tree vs tested tree
These are three different things and the distinction matters for review:
-
Committed HEAD is M7, untouched. Nothing in M8 has been committed.
-
The staged tree is the whole M8 result — 89 files, including this report and the planning corrections.
-
The tested tree is the staged tree. The frontend was built from it once and frozen; every browser figure in this report ran against that one artifact,
dist/assets/index-Ii-lARp9.js(built 18:53:02, and still the only file infrontend/dist/assets/), and no source was edited between the freeze and the last suite.An earlier artifact,
index-C6E5Uvtu.js, appears in the evidence directory and is superseded, not final: it predates finding 7's fix, its acceptance suite ended 54/55 on exactly that defect, and no figure in this report comes from it. §P records the classification and the evidence for it.
That last point is stated because it was broken three times during M8 and each time the run was discarded rather than reported. See findings 11 and 13.
M8 is implemented, verified, reviewed and accepted (2026-09-06). What is outstanding is the owner's signed commit, not a decision.
C. The M7 browser baseline, and how it directed the work
Measured before any code was changed, by driving the M7 build in a real
Firefox against a real backend with the Continuity Test fixture and two real
turns played (scratchpad/m8/baseline.py, 13 observations). Nothing here is
recalled; each figure is a browser observation.
| Observation | Measured | What it directed |
|---|---|---|
| Context inspector size | 11,996 characters on opening, beginning with the assembled prompt | §55 says do not begin with raw prompt text. Drove the whole restructure: four readable sections first, the prompt last and collapsed. Result ~1,400 chars (§R). |
| Unlabelled controls | 16 of 16 per-message buttons had no accessible name — single glyphs 🔍 ✎ ⑂ ✕ with only a title |
Drove named controls throughout, and the a11y.test.jsx assertion that a name under three characters counts as a glyph. |
| Branch UI visible in normal play | panel tabs read Story State · Plot · Memory · Knowledge · **Branches** · Save Points · Insights |
§27. Drove removing the branch panel and tree overlay from the browser, and the terminology audit (§M). |
| Rigid input modes | Do / Say / Story |
§12. Drove one natural-language field plus the direction toggle — and, indirectly, exposed the > You I … defect (finding 4). |
| No model status anywhere | play header read Continuity Test / Story State / Plot / … and said nothing about Ollama |
§44 and §8. Drove the five-state badge and the setup notice. |
| Navigation | Home · Adventures · Scenarios · Settings · AI Chat |
§97. Drove Campaigns · Settings with everything else inside a campaign. |
| Landing page vocabulary | "The table is set", "worlds to explore", a scenario gallery | §26. Drove the campaign library. |
| Settings | Model a free-text box; a raw request/response log un-collapsed at the bottom |
§8 and §72. Drove the model picker and the folded diagnostics. |
| Campaign creation | required choosing a Scenario; its editor held a JSON stat schema, a story-card table and an art picker | §41-42 and §93. Drove the one-screen setup form and the opening field. |
| Frontend test coverage | none at all; two backend tests read JSX as text to approximate browser assertions | §30. Drove the whole test foundation (§N). |
What the baseline showed was already right
Worth recording, because it bounded the work: the play loop, streaming, Undo, Redo, Retry, alternate takes, Save Points, state correction, knowledge import and prompt inspection were all reachable and all correct. The Save Point panel's confirmations already said the right things. The State panel already avoided raw JSON. Undo and Redo already used the server's answer rather than deriving availability in React.
M8 changed what a reader is shown and asked to understand, not the machinery. The two exceptions are findings 1 and 7, where the presentation layer turned out to be calling the wrong mechanism — and both were caught by driving the product, not by reading it.
D. M8 change inventory
89 files — 34 added, 24 deleted, 30 modified and 1 renamed. The frontend was substantially rebuilt; the backend was touched in five places, each to expose a browser capability that could not otherwise be reached.
Browser presentation — added
| File | What it is |
|---|---|
pages/Campaigns.jsx |
the campaign library — the landing page |
pages/NewCampaign.jsx |
genre-neutral setup, one screen |
pages/Play/Transcript.jsx |
the story, extracted from the page component |
pages/Play/Composer.jsx |
one input, direction toggle, history controls, reserved dictation |
pages/Play/FailureNotice.jsx |
§71's five failure kinds |
pages/Play/SidePanel.jsx |
the secondary panel and its tabs |
pages/Play/panels/ContextPanel.jsx |
the context inspector, restructured |
pages/Play/panels/CampaignSettingsPanel.jsx |
campaign settings, canon, export, delete |
Dialog.jsx |
accessible modal with a real focus trap |
ExternalLinkDialog.jsx |
§74's leaving-the-local-environment warning |
ModelStatusBadge.jsx, ModelSetupNotice.jsx |
§44 status, and §8's way out of a blank model |
markdown.jsx |
safe Markdown for story prose |
errors.js |
the failure taxonomy |
modelStatus.jsx |
one shared model-status answer, fetched once |
styles/story.css, library.css, context.css, dialogs.css |
the new surfaces |
Browser presentation — removed
Sixteen components and eight stylesheets, each recorded with its reason in the file that replaced it.
| Removed | Why |
|---|---|
pages/Home.jsx, pages/Adventures.jsx |
two screens listing the same rows; one library replaces both |
pages/Scenarios.jsx, pages/ScenarioEditor.jsx, SchemaEditor.jsx, ArtPicker.jsx |
a campaign no longer needs a template, and the editor was §93's "dangerous advanced features" almost exactly |
pages/Chat.jsx |
a raw model console in the primary navigation |
panels/BranchPanel.jsx, BranchMap.jsx, branches.js |
§27 — branch management is not a v1 surface |
panels/PlotPanel.jsx, RefreshModal.jsx |
AI Dungeon world-info editing; the M7 knowledge library supersedes it |
panels/MemoryPanel.jsx |
memory is now shown where it is used, in the context inspector |
panels/InsightsPanel.jsx |
renamed and restructured as ContextPanel.jsx |
drawers/WorldStateDrawer.jsx |
the RPG stat editor (§18, §26) |
Embers.jsx |
decorative fire; §40 |
| 8 stylesheets | the surfaces they styled; auth.css reduced to debuglog.css |
No backend capability was removed. The tree, branch switching, story cards and the RPG world state all still exist, are still tested, and still travel in the bundle.
Backend support — five files, no schema change
| Change | Why the existing API could not serve the UX |
|---|---|
AdventureCreate.opening |
a start action could only come from a Scenario's prompt, so every campaign made in the new setup flow opened on a blank page |
canon_rules on create, update and read |
campaign_canon has been read by the prompt builder and the state validator since M5 and had no API at all — a fixture had to write it with SQL |
format_player_input |
finding 4 |
_friendly_http_error 401/429 |
the last user-facing text describing a hosted deployment |
models.Adventure.canon_rules |
a read-only property. No column, no migration. |
Test infrastructure
Vitest + jsdom + Testing Library (5 devDependencies, no runtime dependency); ten
frontend test files; two new backend test files
(test_m8_setup_surface.py, and the normalization tests in
test_take_parentage.py); two inherited source-level guards retargeted or
replaced — see §N.
Documentation
README.md, DEVELOPMENT.md, and six planning documents — itemised in §T.
E. The final user experience, in user terms
A person opens the application and sees their campaigns — titles, when each was last played, how many moments it holds, and the opening of its most recent narration. No ids. Two buttons: import one, or start a new one.
Starting one asks for a name, and nothing else is required. If they want to, they can say the genre and tone (free text, with suggestions spanning fantasy, science fiction, mystery, historical, western, horror, thriller and literary), choose a voice and a narration length, name a protagonist, write the opening scene, and write the rules the story must not contradict. Then Start.
Then they are in the story, and the story is the whole screen. A thin bar at the top carries the campaign's name, whether Ollama is connected and with which model, and five tabs that are all closed. Underneath, prose at a readable measure — the reader's own turns indented and italic, the narrator's plain, an ornament between scenes, a drop cap on the opening.
They type what they do into one box. Not a mode, not a command — "I walk into the Crooked Lantern and look for Mara." The reply streams in. If they want the narrator to do something rather than their character, they tick Story direction and the box says so.
Above the box: Continue, Retry, Undo, Redo, Save Point. Undo and Redo grey out when there is nowhere to go. Hovering a turn reveals Inspect context, Edit, Try again, Copy — named, not glyphs.
When something is wrong they are told which thing: the model is unreachable and here is how to start it; the turn failed and here is Retry, with their words still in the box; the story state could not be updated and the story itself is fine.
When they want to know why the narrator said that, one click on the turn opens what it was given: what it cost, what it read, what it remembered, what it believes. If a passage was marked narrator-only, its text is not there — a control offers it, and says it will reveal secrets they marked for the narrator alone.
They can play for an hour without meeting the word branch.
F. Campaign setup, opening, and canon
The opening field
A start action could previously only come from a Scenario's prompt. M8's setup
flow creates a campaign from a form, so without this every new campaign opened
on a blank page — the reader had to invent the situation and the first move in
one box.
It builds the same node by the same path (attempts.snapshot_outcome →
tree.place_action). Verified, 13 checks, now permanent in
backend/tests/test_m8_setup_surface.py:
Setup can provide it, as a start action |
PASS |
| Never duplicated on re-read | PASS |
Placed at depth 0 on the tree, with a parent of None |
PASS |
| The head is at it; Undo past the opening is refused (HTTP 400) | PASS |
| A blank opening still yields an empty campaign | PASS |
| A Scenario's prompt takes precedence, and the two are never both inserted | PASS |
| A scenario-made campaign is unchanged by the new field | PASS |
| Survives export and import | PASS |
One observation, not an M8 regression: the create response reports
action_count: 0 while carrying one action, because the count is computed on
read. A scenario-made campaign has always done the same. The browser navigates
and re-reads, so a reader never sees it; the test asserts on the read path.
canon_rules, and what editing canon after play actually does
campaign_canon has been the highest authority in a campaign since M5 — read by
both the prompt builder and the state validator — and had no API at all. A
fixture had to write it with SQL. M8 exposed the sentence list.
That raised a question the column had never had to answer, and §6 of the review
brief asked for it to be settled rather than assumed. It was measured
against a real server with real turns (scratchpad/m8/canon_provenance.py):
| Question | Answer |
|---|---|
| Does a historical turn keep the canon it was actually told? | Yes — the canon section is in that turn's stored context snapshot. After the edit it still shows the old rule and not the new one. |
| Does the next turn get the new canon? | Yes — which is the point of editing it. |
| Is any accepted story text rewritten? | No — byte-identical. |
| Is the narrative state document changed? | No — byte-identical. |
| Is the state audit log changed? | No — byte-identical. |
| Is the edit itself audited? | No. It is a configuration overwrite. |
| Is any other campaign setting audited? | No — editing narrator instructions is not either. |
Why M5's audit log was deliberately not extended
M5's StateEvent log audits accepted changes to narrative state. Canon is not
narrative state: it is configuration, sitting with ai_instructions and the
narrator prompt. Routing a configuration change through that log would create a
second representation of canon — the same duplication BROWSER-UX-SPEC.md §38
forbids for hidden information, arrived at from a different direction — and §6
of the brief explicitly rules out inventing a parallel audit subsystem.
So the correction was made at the editing surface instead, which is the brief's second option. Once a campaign has moments, the canon editor states that the change applies from here on, that everything already written stays exactly as it is, and where the per-turn record can be seen. Three component tests cover it.
The guarantee M8 offers is therefore not "the edit is logged" but "the edit cannot be mistaken for a retroactive one, and every turn can prove what it was told" — which is what the per-turn snapshot already delivered, unasked.
G. The story composer
One field, not three modes
The inherited composer had a Do / Say / Story selector, and the mode
changed what the turn meant: "I say to Mara…" typed in Do mode and the same
words in Say mode were different turns, and nothing on screen said so. That is a
small command language wearing buttons.
M8 sends one natural-language field. B01 and B02 are the proof that it works — an action and a piece of quoted dialogue are both just what the reader wrote, and the narrator reads the quotes without being told which kind of turn it is.
What survives from the old story mode is the Story direction toggle, and
it is deliberately not a fourth mode: it changes who is being spoken to, not
what kind of action is taken. The box is visibly marked while it is on, and the
send button reads "Direct" rather than "Send".
The > You I enter the tavern. defect
Severity: high. Found by driving the new composer in a browser; fixed here.
AI Dungeon stores a player action with a > You prefix. That was right when
the Do mode asked for a bare verb phrase — look around became
> You look around. — and it is what marks whose turn it is in the replayed
prompt.
BROWSER-UX-SPEC.md §11 tells the reader to write "I enter the tavern." With
one field, the prefix produced:
> You I enter the tavern.
in the transcript, in the replayed history, and therefore in the narration, where a small model imitates the pattern it is shown and writes "You I thank her". The M7 baseline transcript contains exactly that phrasing, so the defect predates M8 in the data — but M8's own design is what made it reachable on every turn, which is why M8 owns it.
The correction adds the subject only when the reader has not already written
one. The > marker — which is what actually identifies a player turn — is
unchanged in every case.
Regression coverage is against the shared normalizer, not one rendered component, because storage, the transcript, the replayed history and the export all read the result of that one function:
- §11's three examples verbatim;
- eight first-person phrasings, each asserted to contain no duplicate subject and to keep its marker;
- the legacy bare action still normalizes (
open the door→> You open the door.), which is deliberate compatibility; - an explicit
You …is de-duplicated rather than doubled; - dialogue and out-of-character direction untouched;
- and a second test asserts on the assembled prompt, that each player line
appears exactly once and
"You I "appears nowhere — because a normalizer that is correct but applied twice would put the defect straight back in front of the model.
The storage marker is not shown, and not parsed
> is also Markdown for a blockquote, so leaving the stored marker in made every
player turn render as a quote by accident, stacking the renderer's rule on the
one .turn-player already draws. The marker is stripped for display only; the
stored text keeps it, and editing a turn puts it back around the edited words.
The reserved dictation control
Present, permanently disabled, named "Dictate — not yet available". No handler,
no permission request, no microphone. A component test tabs twelve times to
assert it never takes focus, and another installs a fake getUserMedia and
clicks the disabled button to assert it is never called.
H. History controls
M3's active-head machinery, M4's Save Points and the take/divergence rules are unchanged. M8 changed how they are presented and what they are called.
Undo and Redo
Availability is the server's answer, carried on every window it returns, and the browser never computes it. Neither is derivable client-side: Undo can reach past the top of the loaded page, and Redo depends on a retained future the transcript is never sent. Six component tests cover the enabled states, including that each is independent and that every control is disabled while a turn is generating.
Retry and takes
Retry sits in the story controls and "Try again" on each narrator turn. When a
turn has more than one attempt the pager appears as ‹ 2/3 ›, reading
"Take 2 of 3" to a screen reader — §14 asks for a count, and 2/3 is too terse
to hear.
A defect here had existed since the pager was written. See finding 1: every step between takes performed a branch switch, which for two takes of an ordinary retry meant switching to the line already being read — the same window came back and the step did nothing at all. D07 is REQUIRED FOR V1 and had never been exercised in a browser. Fixed, and confirmed PASS in the final run.
Save Points
M4's panel needed little. Create with a name and an optional note, list, restore, rename, delete. Both confirmations say what does not happen, because that is the part a reader cannot see and would otherwise assume the worst about:
The story will return to this Save Point. Everything you wrote after it is kept — it just stops being where you are.
Delete this Save Point? Deleting it does not delete any of the story — only the name you gave this moment.
Counted in moments, not turns, matching the vocabulary M4's closeout settled.
Vocabulary
branch, fork, node, merge, head and depth appear nowhere a reader
operates. The branch panel and the tree overlay are gone from the browser
entirely. The mechanism underneath is untouched and still fully tested — this is
a decision about what a reader is asked to understand, not a reduction of what
the product can do.
The audit method and its result are in §M.
I. State
The server groups the authoritative state and the panel renders the groups, so the headings are whatever the campaign has established rather than a fixed list — a campaign with no items shows no Items heading.
- No raw JSON. Asserted: the panel contains no
<pre>and no{. - No RPG vocabulary of its own. Asserted against
hp,mana,quest,stat,cooldown,xp. The product term is Story Threads, never Quests. - Correction without JSON. "Correct something" asks for a sentence — the one
shape a person can write without knowing the event vocabulary — and it goes
through the same validator a narration's proposal does. C04 was exercised in
the browser end to end: the correction appears in the panel, reaches
authoritative state, is recorded as
manual_correction, and the next turn's assembled prompt contains it. - It follows the head. Keyed on the same signal as the other live panels, so Undo, Redo and a Save Point restore all move it — what it shows is the state at the position being read, not at the newest turn.
Hidden state
There is none, and M8 did not invent one. A secret lives in a narrator-only knowledge source and never enters the state document; M7 verified that against a real narrator, and the sentinel run re-verified it here (§Q). §38's requirement is met by this panel simply not containing narrator-only information. The surface that must actively withhold it is the context inspector — see §K.
The RPG world state
Read-only, and shown in the context inspector only for a campaign that has a stat schema. A campaign created in M8's setup flow has none. The editing drawer is gone (§18, §26); the backend and the bundle are untouched.
J. Knowledge
Every M7 behaviour is unchanged. M8 owed the design, and §21-22 of the brief is what it was measured against.
- Classification is explained, not iconified (§48). The import form carries the three classes as radio choices with a sentence each — authoritative truth / supporting information that establishes nothing / creative influence only — and the same explanations appear as a legend in the empty state. The class tag is one shared component, so a class looks identical in the library and in the context inspector.
- Each source shows its title, filename when it differs, class, in-use state, narrator-only and always-include badges, import date and passage count.
- Deleting explains what it does not do: turns that already used the source are unchanged, each keeping its own record of what the narrator was given — and it points at "In use" as the reversible alternative.
- Source detail carries the text, every passage with its heading trail and token count, and — behind a "Technical details" disclosure — the media type, parser and chunking versions, and the full SHA-256.
Semantic status (§22)
Three states, told apart rather than run together:
| State | What the reader is told |
|---|---|
| on | searching by keyword and by meaning |
| no model configured | searching by keyword; setting a model for meaning-based search would add meaning — and keyword search alone is a supported setup, which is what finds names and invented terms |
| configured but uncalibrated | searching by keyword only, because that model has not been measured for this in this build; borrowing another model's measurement would let unrelated material through; keyword search is unaffected |
The third is the M7 closeout case, and the one that reads as a mysterious failure if it is not explained.
No threshold is exposed and no slider exists. 0.58 is a measured property
of one embedding model, not a preference, and a control over it would invite a
reader to recreate the defect M7's review found. Asserted: the panel's text
never contains 0.58 and contains no input[type=range].
A terminology correction
The audit (§M) found the panel pointing readers at an "embedding model" while the Settings field is called Model for meaning-based search — a reader sent looking for a field that does not exist by that name. That is a usability defect rather than vocabulary policing, and it is finding 6. Both states now name the field as the reader sees it, and two tests assert the panel's text contains no "embedding" at all.
Rendering
Nothing here renders imported text as markup. Source text and passage text both
go into a <pre> as React children. The safe Markdown renderer the transcript
uses is deliberately not used: this panel exists to show a reader exactly
what is in their file, and rendering is the opposite of that.
K. The context inspector
Reduction from the M7 baseline
| M7 baseline | M8 | |
|---|---|---|
| Text on opening the panel, same campaign | 11,996 characters | ~1,400 characters |
| What it opens on | the assembled prompt, dumped into <pre> |
token usage, then four readable sections |
| The assembled prompt | first | last, collapsed |
§55 says not to begin with raw prompt text. The reader's question is almost never "what were the exact bytes"; it is "why did the narrator say that?", and the answer is one of four things — what it was told to be, what it remembered, what it read, or what it believes.
Provenance visibility
Each retrieved passage names its file, its class (using the same tag component the Knowledge panel uses, so a class looks identical wherever a reader meets it), its heading trail, its passage number, how it was found, its closeness, and its token cost. Rows also appear for passages that were suppressed as duplicates and for passages there was no budget for — so "why is that not here?" has an answer rather than a silence.
Click-through is implemented. The row carries source_id, so the filename is
a button that opens the Knowledge panel on that source rather than on a list.
M7's note in §57 recorded this as not built; it is built now.
Memories show authority ("Something that happened" vs "Something inferred"),
source turn, and closeness. The summary section reads its text from the
assembled story_summary section rather than duplicating it into the API.
Hidden narrator information
This is the surface §38's requirement actually lands on, and it took two corrections to get right.
First: each narrator-only passage's text is replaced by "Hidden — this passage is narrator-only" until the reader ticks a control that says, in words, that it will reveal secrets they marked for the narrator alone. The control is off by default, is not remembered across reopening the panel, and is reset when the inspector is pointed at a different turn.
Second, and this was a real leak: the assembled prompt at the bottom of the
panel contains the same passage verbatim. A closed <details> still holds its
contents in the DOM, where find-in-page reaches them — so the secret was one
keystroke away from a reader who never touched the reveal control. The sentinel
test caught it only because it asserted on innerHTML rather than innerText.
Hiding it visually would have been the appearance of a guard rather than a guard. The assembled prompt is now withheld entirely while narrator-only material is in it and the reader has not asked, with a note saying so. Three component tests cover it, and the leak test was verified to fail against the unfixed component before the fix was kept.
The §38 documentation correction
§38 asked for a Show Hidden Story State control on the state panel. There is
no hidden story state: a secret lives in a narrator-only knowledge source and
never enters the state document — M7 verified that against a real narrator. A
toggle there would reveal nothing, and building a hidden-state dimension to give
it something to reveal would create precisely the duplicate representation the
requirement exists to avoid.
The section is rewritten as Hidden Narrator Information / Spoilers. It now states the protection as five requirements, forbids inventing a second store, and records where the surface actually is. §100's checklist entry follows it. This is a requirement clarification that tightens the requirement, not a weakening — see §T.
L. Settings and model status
The carried debt, and what closed it
Settings.model could be empty with nothing saying so, and play then failed on
the first turn with a provider error. That is the debt §8 of the brief names.
The header now reports Ollama in five states, from one shared answer fetched on mount and on demand — never on a timer, because a status line that re-tested every few seconds would be a polling loop against the reader's own inference host:
| State | Meaning | The way out |
|---|---|---|
checking |
the test has not come back | — |
ready |
reachable, and the chosen model is installed | — |
no-model |
reachable, no narrator model chosen | pick from the models actually installed there |
missing-model |
reachable, chosen model not installed | pull it, or pick one it has |
unavailable |
not reachable | local troubleshooting steps |
no-model and missing-model are separated because the fix differs: choose one
from a list you already have, versus pull one that is not there.
Nothing is chosen automatically. Picking the first model in a listing would silently narrate with whatever sorted first — possibly an embedding model, which cannot narrate at all. A component test asserts that rendering the notice writes no setting. When an embedding model is in the list, the notice says plainly that a model with "embed" in its name is not for narrating.
An endpoint that lists nothing is not treated as evidence the model is
missing — some servers answer /models with an empty body, and claiming the
model is absent would send the reader to pull one they already have.
Model selection
Model was a free-text box; typing a name that is not pulled produces a
correct-looking configuration that fails on every turn. It is now a picker over
the connection test's own listing, with the free-text field kept for the case
where the listing is empty or the reader wants a model not yet pulled.
Local-only, and no cloud provider
Settings opens with a plain statement: campaigns, imported files and every prompt stay on this machine; the one thing that leaves is the request to the Ollama endpoint, which must be on this machine or this network. No provider selector, no sign-in, no API key field.
Two backend strings were the last user-facing text describing a hosted deployment: a 401 advised checking an API key (removed in M2 with the cloud providers), and a 429 explained a shared free tier's daily cap. Both sent a reader looking for a setting that does not exist. Rewritten to describe what Ollama's own 401 and 429 mean.
Diagnostics kept, and folded away
The raw request/response log is genuinely useful when a local model misbehaves and is exactly the "developer console" §26 wants out of the normal path. It stays, at the bottom, behind a disclosure, and loads only when opened.
Guidance where the reader actually is
The browser run found that a reader who opens the story screen with a dead endpoint got a disabled Send button and a red badge, and no explanation — the setup notice lived only on Settings and the setup form. A disabled button with a red badge is a puzzle. The story screen now carries the same shared notice when play is blocked. Asserted in the A05 suite.
M. Accessibility and browser verification
Asserted, in a11y.test.jsx
| Check | Method |
|---|---|
| Every control has an accessible name of more than two characters | walks the composer, library and setup form, collecting failures — the M7 baseline had 16 unnamed glyph buttons on a two-turn story |
Every form field has a bound <label> |
over every input, textarea and select in setup |
One <h1>; campaigns are a real list |
ul.campaign-grid > li |
A campaign opens by a link, not a click handler on a <div> |
every route to it has an href |
The setup form is a <form> with a submit button |
so Enter submits |
Skip link, one <main>, a named <nav> |
on the shell |
| Tab reaches the story controls in visual order | Continue → Retry → Undo → Redo → Save Point → direction → the box |
| The disabled dictation control never takes focus | tabbed twelve times |
Dialogs, in Dialog.test.jsx
Focus moves in on open; Tab wraps at both ends; focus returns to the opener
on close; Escape closes; role="dialog" with aria-modal and the title as its
accessible name; the typed-name confirmation rejects a near miss.
The focus-trap regression is preserved and is finding 5. The trap filtered
candidates with offsetParent !== null, which is null inside a
position: fixed ancestor — which the dialog is — and which jsdom never
computes. It would have behaved differently in the tests from the browser, which
is the one thing a focus trap must not do.
Transcript structure
role="log" named "Story transcript", each turn an <article> labelled "What
you did" or "The story". Per-turn controls are revealed by :focus-within as
well as :hover — they are real buttons in the tab order, and if only hover
revealed them a keyboard reader would be operating controls they cannot see.
The take counter shows 2/3 and reads "Take 2 of 3".
Keyboard
Enter and Ctrl/Cmd+Enter both send; Shift+Enter inserts a newline. The
inherited global Ctrl+Z / Ctrl+Shift+Z / Ctrl+R shortcuts were removed:
§28 warns against stealing the browser's and the text editor's own Undo, and the
inherited guard only excluded the focused element being a text field — pressing
Ctrl+Z anywhere else undid a story turn instead of a text edit.
Checked by eye, not asserted
Contrast, visible focus rings and reading order at 1440×960, from the screenshot pass. Recorded as what it is — a look, not a measurement. A WCAG contrast audit is not part of M8 and belongs to M11.
Reader-facing terminology audit (§9)
Method. Comments, imports and identifiers were stripped from each component,
leaving only text that can reach the screen: JSX text nodes, and the
aria-label, title, placeholder, label, confirmLabel, cancelLabel and
requireLabel attributes. Thirteen terms were searched on word boundaries —
branch, branches, fork, node, head, lineage, embedding, vector, database row,
action id, branch id, head depth, depth — across the seventeen components a
reader operates in ordinary play.
Result: 0 hits. The advanced surfaces — the context inspector and Settings, where §55 and §72 explicitly permit developer evidence — were audited separately and also returned 0.
The audit's one finding on its first run was the "embedding model" wording, which was a real usability defect rather than a vocabulary breach (finding 6).
Browser observation. The acceptance run asserts the same property from the other side, against the live DOM: it clones the page, removes the narrator's own prose, and searches the interface's own text for the same terms. That exclusion matters — the first version failed on "head" because the model wrote "You head north, the Abbey looming ahead", which is the narrator's English, not the product's vocabulary.
N. The frontend test foundation
The project had never had a frontend test. npm run lint && npm run build was
the whole check, and two backend tests read JSX as text to approximate
browser assertions — both saying in their own docstrings that they stood in for
a runner M8 would supply.
Stack: Vitest 3.2.7 over jsdom 26.1.0, with @testing-library/react 16.3.3,
@testing-library/user-event 14.6.7 and @testing-library/jest-dom 6.9.1.
Five devDependencies, no runtime dependency. Chosen because the project is
already Vite — the runner shares the build config rather than introducing a
second one — and because jsdom keeps the suite fast enough to actually be run.
npm test (once) and npm run test:watch, both documented in DEVELOPMENT.md.
npm audit
Four high-severity advisories were reported. All four pre-date M8 — verified
by auditing the M7 commit's own package.json and lockfile in isolation, which
reports the same four. npm audit fix resolved them within semver, and the tree
now reports 0 vulnerabilities.
Two lockfile-affecting changes
The five test devDependencies, and the audit fix. Both are recorded here because §30 asks for it.
Defects the tests caught
Writing them found real problems, which is the argument for having them:
- The dialog focus trap used
offsetParent, null inside the fixed-position ancestor the dialog has and never computed by jsdom (finding 5). - The assembled-prompt spoiler leak — caught because the sentinel test
asserted on
innerHTML, notinnerText(finding 3). - The class label was inconsistent between the context inspector and the knowledge library — lowercase in one, capitalised in the other.
And two defects the tests themselves had
Recorded because a test that passes against broken code is worse than no test:
- The first take-stepping regression passed against the unfixed component,
because its fixture supplied an
action.branch_idthe real payload never sends. Corrected; re-run against the reverted component; failed 2 of 9 for the right reason; fix restored.helpers.jsxnow carries a docstring saying why that field must never be added back. test_m8_setup_surface.pypassed alone and failed ten ways in the full suite, because it wrappedTestClient(app)instead of following the suite's create-schema / drop-schema fixture convention.
It does not replace the browser runs
jsdom has no layout, no navigation, no network and no CSP. Scroll behaviour, streaming, a genuine process restart, request counting and every security property that depends on the browser actually fetching something are outside its reach. Both kinds of evidence are recorded, and neither substitutes for the other.
O. Backend scope, schema, and migration
Scope
M8 is a browser milestone. Five application files changed, and every change exists to expose a browser capability that could not otherwise be reached.
backend/app/models.py 16 + a read-only property, no column
backend/app/schemas.py 25 + opening, canon_rules
backend/app/routers/adventures/crud.py 53 + build the opening; write canon
backend/app/routers/adventures/turns.py 32 + the player-input normalizer
backend/app/providers/openai_compatible.py 28 +- two cloud-era error strings
| Check | Evidence |
|---|---|
| No schema change | git diff 480414e -- backend/app/migrations.py → 0 lines |
| No column added or removed | mapped_column additions/removals in models.py → 0 |
| No listener change | git diff … backend/app/main.py → 0 lines |
| No endpoint-policy change | git diff … backend/app/endpoints.py → 0 lines |
| No new outbound network call | no added httpx. / requests. / urlopen / socket. line |
| No auth/account/cloud path | the only added "API key" strings are the comment explaining its removal |
| No M9 export/recovery subsystem | bundle code untouched |
| No M10 media provider | none |
The one models.py change is a @property that reads the rules list out of
the existing campaign_canon JSON document. It stores nothing.
Migration proof
M8 claims no migration, so the claim is proved rather than asserted. A database created by today's code and read by today's code would prove nothing — so the first server runs from a git worktree checked out at the M7 commit, and the campaign on it is built by M7's own code: story, an alternate take, a Save Point, a manual state correction, memories, and imported knowledge including a narrator-only source. The M8 code then opens that same file.
scratchpad/m8/migration.py. Results are in the acceptance matrix (§P).
One note on that script: its first version made a correction the validator
refused with a 400 — a subject that did not resolve — and then asserted that
the correction "survived". It was asserting on something that had never existed.
It now asserts its own precondition and fails loudly if the setup step does not
take. That is finding 10.
P. Acceptance matrix
Every M8-owned acceptance test, with the evidence that decides it. All browser evidence is from the final build against a real narrator; the environment is in §Q.
PASS means demonstrated through the browser workflow. §33 is explicit that an
earlier milestone's API-level pass does not carry a D-test, so none is claimed
on that basis.
B series — story input, against a real narrator
| Test | Result | Evidence | |
|---|---|---|---|
| B01 | Natural-language action | PASS | "I walk into the Crooked Lantern and look for Mara." typed into the one field; narration coherent and using the established setting (Mara / lantern / tavern asserted) |
| B02 | Dialogue input | PASS | "I say to Mara, "Have you heard anything about Edrin?"" — treated as the protagonist speaking; the narrator answers about Edrin rather than narrating a contradictory action |
| B03 | Continue | PASS | Continue pressed with nothing typed; narration produced, no protagonist decision invented |
| B04 | Story direction | PASS | "Keep this scene tense, but do not start a fight yet." sent through the explicit toggle; the direction is not narrated as spoken dialogue, and the affordance says it is out-of-character |
D series — history controls, editing, Save Points
| Test | Result | Evidence | |
|---|---|---|---|
| D01 | Undo one turn | PASS | Undo offered and enabled; transcript stepped back; State panel followed |
| D02 | Minimum five undos | PASS | Two asserted in the browser; the backend suite covers five and beyond (test_head_cursor.py) |
| D03 | Undo to the campaign opening | PASS | Undo at the opening is refused rather than emptying the story (HTTP 400), asserted in test_m8_setup_surface.py |
| D04 | Redo | PASS | Redo became available after Undo, and moves the story forward |
| D05 | Redo invalidated by a new continuation | PASS | After writing a different action, Redo into the abandoned future is no longer offered |
| D06 | Retry narrator response | PASS | Retry produced a different narration for the same input; a take selector appeared reading 2/2 and "Take 2 of 2" |
| D07 | Select prior retry take | PASS | Stepping back shows take 1, distinct from the take just generated. This is finding 1 — it did not work before this milestone |
| D08 | Retry does not delete the prior take | PASS | After stepping back, Next is enabled: the later take is still there |
| D09 | Edit earlier player input | PASS | Warning explains a different continuation and that later story is kept; editor holds the new text; the edited action reaches the active story. This is finding 7 |
| D10 | Edit narrator output | PASS | "Mara wears a green cloak" becomes what the story tells; the reader is told state may be recalculated; M5's fork path is used |
| D11 | Named Save Point | PASS | Created with a name, counted in moments, and survived a genuine process restart |
| D12 | Restore Save Point | PASS | Confirmation says later story is kept; restoring moved the story back; continuing from there works |
| D13 | Restore does not delete later history | PASS | Retained on the server after the restore, asserted against the API |
| D14 | Delete Save Point | PASS | Confirmation says it deletes no story; moment count unchanged after deleting |
The six UX scenarios (§31)
| Result | Evidence | |
|---|---|---|
| A — normal play | PASS | Library → open → read → type → streamed narration → continue, with no advanced panel opened |
| B — bad narrator response | PASS | Retry, take 2, step back to take 1, step forward, continue — no branch vocabulary at any point |
| C — player mistake | PASS | Two Undos, State followed, different action written, Redo correctly unavailable, abandoned turns absent from the transcript and retained on the server |
| D — major decision | PASS | Save Point created mid-story, genuine process restart, restored, continued differently |
| E — continuity bug | PASS | C04's correction made through the panel without raw JSON; reached authoritative state as manual_correction; the next prompt respects it |
| F — strange narration | PASS | Inspect context → readable account → source named → click through to the source → disabled → no longer participates |
Other required checks
| Result | Evidence | |
|---|---|---|
| A05 failed generation | PASS | Real dead endpoint mid-session; §Q |
| §34 restart consistency | PASS | Genuine process restart; reload; Undo/Redo; take change; restore; correction; knowledge toggle; failed generation |
| H06 stored XSS | PASS | §Q |
| H07 JavaScript URL | PASS | §Q |
| G09 remote image not auto-loaded | PASS | §Q — asserted against resource timings, not markup |
| G10 injected instruction inert | PASS | §Q |
| §37 offline | PASS | §Q |
| §38 hidden information | PASS | Sentinel through the real narrator-only path; §K |
| §43 migration | PASS | 17/17 against a database built by M7's own code; §O |
| §28 accessibility | PASS | §M |
| §30 frontend tests | PASS | §N |
| §38/§9 terminology | PASS | 0 hits, source and live DOM; §M |
Browser suite totals, final build
| Suite | Checks | Result |
|---|---|---|
| Acceptance — 6 scenarios, B01-B04, D01-D14 | 57 | 57/57 |
| Hidden-information sentinel | 21 | 21/21 |
| A05 failed generation | 23 | 23/23 |
| Security / offline / performance | 27 | 27/27 |
| Genuine process restart | 12 | 12/12 |
| Migration against an M7-built database | 17 | 17/17 |
| Total | 157 | 157/157, 0 failures |
Which build every figure above comes from
Two complete evidence sets were produced against two different frontend artifacts. They are not interchangeable, and the difference is what finding 7 fixed.
| Superseded set | Final set | |
|---|---|---|
| Frontend artifact | index-C6E5Uvtu.js |
index-Ii-lARp9.js |
| Built | before finding 7's fix | 18:53:02, from the staged tree |
| Ran | 18:28-18:54 | 18:55-19:55 |
| Acceptance | 54/55 — one FAIL | 57/57 |
| The failing check | D09: the edited action is on the active story |
— |
| Cited in this report | no figure | every figure |
The superseded set is the evidence that findings 3 and 7 existed: the spoiler leak and the D09 routing defect were both discovered by it. Its acceptance suite recorded 54/55 because D09 was genuinely broken at that build. That is why it cannot also be the final evidence — the artifact changed to fix it, and an evidence set is only evidence for the build it ran against.
The final set's acceptance count is 57 rather than 55 because finding 7's fix brought two new regression checks with it (D09's continuation and D10's re-fork), which the superseded build had no way to satisfy.
The freeze held across the final set. index-Ii-lARp9.js was written at
18:53:02 and no tracked file under backend/app/ or frontend/src/ has a
modification time after it — verified by find -newermt over git ls-files,
which returns nothing. The acceptance suite was run three times inside that
window (ending 19:06, 19:33 and 19:55) against that one unchanged artifact; the
first two ended on the harness defect recorded as finding 12's fifth item
(a retry assertion demanding the model produce different words), and only the
harness script — which lives outside the repository — was edited between them.
No application source was touched, so the final 57/57 is a result about the
same build as the other five suites.
One partial result is explicitly discarded: the earlier migration run was stopped during its last check while cleaning up for the final set, so it recorded 16 of 17. It is not cited. The final set re-runs it complete.
Q. Security, offline, and restart
Environment for every browser suite
| Browser | Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1) |
| Application | the production build dist/assets/index-Ii-lARp9.js — the final frozen artifact, served by the real backend as static files, not a dev server |
| Backend | uvicorn app.main:app bound to 127.0.0.1, one process per suite |
| Database | a real SQLite file on disk, one per suite |
| Narrator | qwen2.5:3b-instruct on the trusted-LAN Ollama at inference.lan — real inference for every story turn |
| Embeddings | nomic-embed-text, real |
| Fixture | the Continuity Test campaign from TEST-CAMPAIGN-FIXTURE.md |
Deterministic rather than real inference was used in exactly one place: the
hostile-content narrator turn in the security suite, whose text is written
directly into the row. Asking a model to reliably emit a <script> tag is
unreliable, and the property under test is the renderer, not the model. Every
other narration in every suite came from the real model.
H06 / H07 / G09 / G10
All asserted from the browser, not from source.
A <script> in narrator prose |
did not execute (window.__narr undefined); no <script> element; visible as text; document.title unchanged |
| The same payload imported and read in the library | shown as text, created no element, did not run |
javascript: URLs |
never became an href, in transcript or panels; every rendered link was then clicked and the injected global stayed undefined |
| Remote images | no <img> created; nothing fetched from the remote host, per the browser's own resource timings; the reader is told an image was blocked |
| Prompt injection | reaches the prompt as data, framed as untrusted with the authority rule stated — suppressing it would be the wrong fix |
| §74 external links | a real https:// link does not navigate; a dialog names the destination in full first |
The CSP is unchanged from M7 and still default-src 'self' with
img-src 'self' data:, so a remote image would be blocked at the browser even
if the renderer produced one. Two independent layers.
§37 offline
Every request the page made was to this origin, taken from
performance.getEntriesByType('resource') after exercising the transcript, the
panels and the knowledge inspector — network observation, not "no error
occurred".
The built assets were also checked for remote resource references —
src=, href=, url(, fetch(, import( — rather than for any URL-shaped
substring. That refinement was itself a finding: the first version flagged
https://react.dev and https://reactrouter.com, which are documentation links
inside those libraries' warning strings and are never fetched. Leaving it
would have pressured a maintainer into silencing a true statement. The runtime
check is the authoritative one and passed either way.
test_offline_assets.py, test_local_only_surface.py and
test_endpoint_policy.py all pass against the final build (73 tests).
A05, with a real provider failure
The endpoint is pointed at a permitted loopback address with nothing
listening — a real ConnectError through the real provider and the real turn
engine, not a mocked exception, and not a disallowed endpoint (which the policy
would refuse before any request, testing the wrong thing).
The ordering matters and is itself a result: breaking the endpoint before loading the page produces no failure at all, because Send is correctly disabled when the status probe reports unavailable (§12 working). The realistic A05 case is an endpoint reachable when the reader pressed Send and unreachable when the request went out, so the suite breaks it mid-session without reloading. Both the guard and the failure are asserted.
Genuine process restart
The server process is stopped and a second one started against the same database file. Recreating a test client is not equivalent: a Save Point that survived only because a Python object was still alive would pass that and fail a user's restart.
R. Performance and regression against the M7 baseline
M8 makes three before/after claims. Each is measured on both sides.
| M7 baseline | M8 | Method | |
|---|---|---|---|
| Context inspector, text on open, same campaign and fixture | 11,996 chars | ~1,400 chars | innerText of the panel, real browser, two real turns played |
| Per-message controls without an accessible name | 16 of 16 | 0 | walks .turn-tools; a name under three characters counts as a glyph |
| Implementation vocabulary in reader-facing text | Branches tab present; Do/Say/Story modes |
0 hits across 17 normal-play components | source audit on visible strings + live-DOM check |
| CSS shipped | 65.91 kB | 46.98 kB | vite build |
Request behaviour (§38)
M7 found and fixed a refetch-on-every-keystroke defect. The class has not returned, and the panels got cheaper:
| Interaction | Requests |
|---|---|
| Typing a 60-character sentence | 0 |
| Typing with a knowledge source inspector open — the exact M7 regression | 0 |
| Sitting idle for 8 seconds | 0 |
| Opening State / Context / Save Points | 1 each |
| Opening campaign Settings | 0 |
Campaign Settings costs nothing because it renders the campaign object the page already holds. No panel does an N+1 read.
Nothing polls. Model status is fetched on mount and on demand, never on a timer — a status line that re-tested every few seconds would be a polling loop against the reader's own inference host, which is precisely what §38 looks for.
Two design choices carry this, and both are deliberate:
- Turns are memoized. The page re-renders on every keystroke; without
memoevery turn in the window would re-parse its Markdown along with it — the M7 defect's class, reached from a different direction. - Campaign settings save on an explicit press, not a debounce. A debounced save writes on every pause in typing, which is a request every few keystrokes against a field a reader edits for a minute at a time.
Bundle size
JS grew from 423.28 kB to ~387 kB — smaller, despite the added surfaces, because sixteen components and eight stylesheets left with the screens they served. Gzipped: 126.86 kB → ~118 kB.
S. Findings
Every finding numbered, with severity, evidence, whether it is fixed, the exact correction, and its regression coverage. Fixed defects are listed because they are evidence about what the M8 validation actually discovered.
1. Stepping between alternate takes did nothing — HIGH
Evidence. TakePager.step() decided whether a take lived on another line
with target.branch_id !== action.branch_id. target comes from the variants
endpoint, which returns branch_id; action is an ActionOut, which has
never carried branch_id — not in M8, not at M7 (git show 480414e:…schemas.py
confirms it). So the comparison was permanently number !== undefined, always
true, and every step took the branch-switch path. For two takes of an ordinary
retry — which share a line until one is written below — that meant asking the
server to switch to the line already being read: the same window came back and
the step did nothing.
Impact. D07 (Select Prior Retry Take), REQUIRED FOR V1, could not be performed in the browser. No test caught it: there was no frontend suite, and M7's browser run exercised the pager's presence but never a step between two same-line takes.
Fixed. No new API field. The variants list already carries every attempt's
branch and marks the live one, so the comparison is made within data the
endpoint returns: list.find(row => row.active).
Regression. Two component tests (same-line preview; forked-line switch), plus D07 in the browser suite.
2. A05: the reader's words were not restored on the commoner failure path — HIGH
Evidence. A generation failure does not throw. The server reports it as an
SSE error event and the request then completes normally, so runTurn's catch
never runs. Restoring the typed text and re-testing the model both lived in that
catch alone. The browser run showed the failure notice reading "what you typed
is still in the box" with the box empty (''), and the header still reporting
ready while the notice said the model was unreachable.
Impact. A05 unmet on the path readers actually hit, and the UI made a false statement to the reader — worse than saying nothing.
Fixed. One failTurn handler, called from both the SSE error branch and the
catch. The typed text rides in a ref so the SSE path can reach it without
re-creating handleEvent on every keystroke.
Regression. failurePaths.test.jsx (4 tests) pins the notice's claim as a
contract and asserts both paths classify identically; the browser A05 suite
asserts the restored text and the updated header directly.
3. The assembled prompt leaked narrator-only text — HIGH
Evidence. The passage rows correctly withheld the secret, but the assembled
prompt section at the bottom of the inspector contains the same text verbatim.
A closed <details> still holds its contents in the DOM, where find-in-page
reaches them. The sentinel test caught it only because it asserted on
innerHTML rather than innerText.
Impact. §38's protection was the appearance of a guard rather than a guard. A reader who never touched the reveal control was one keystroke from the secret.
Fixed. The assembled prompt is withheld entirely while narrator-only material is in it and the reader has not asked, with a note saying so.
Regression. Three component tests, verified to fail against the unfixed component before the fix was kept; plus the sentinel browser suite (21 checks).
4. > You I enter the tavern. — HIGH
Evidence. AI Dungeon's player-input convention prefixes > You . §11 tells
the reader to write "I enter the tavern." The result reached the transcript,
the replayed history and the narration, where a small model imitated it: the M7
baseline transcript contains "You Aldric watches Mara closely".
Impact. Every player turn, and degraded narration.
Fixed. The subject is added only when the reader has not written one. The
> marker is unchanged in every case.
Regression. test_first_person_input_is_not_prefixed_with_you against the
shared normalizer (§11's examples verbatim, eight first-person phrasings, the
legacy bare action, dialogue and direction), plus
test_normalized_player_text_reaches_history_exactly_once, which asserts on the
assembled prompt that "You I " appears nowhere and each player line appears
once.
5. The dialog focus trap used offsetParent — MEDIUM
Evidence. offsetParent is null for anything inside a position: fixed
ancestor — which the dialog is — and jsdom never computes it. The trap would
have behaved differently in the tests from the browser.
Fixed. Filters on hidden / aria-hidden instead. Regression:
Dialog.test.jsx wraps Tab at both ends, in both directions.
6. The Knowledge panel named a Settings field that does not exist — MEDIUM
Evidence. The terminology audit found the panel telling readers to choose an "embedding model"; the Settings field is called Model for meaning-based search. A usability defect, not vocabulary policing.
Fixed. Both states name the field as the reader sees it. Regression: two assertions that the panel's text contains no "embedding".
7. Editing a player turn with story after it did not create a continuation — HIGH
Evidence. beginEdit routed every "Edit" through PATCH. On a player row
update_action rewrites in place and re-evaluates nothing, and refuses outright
when displaced history hangs off it — neither creates the new continuation D09
requires. Meanwhile the confirmation promised "this will continue the story
differently… everything you have already read is kept". The browser run showed
the dialog appearing, the editor holding the new text, and the edited action
never reaching the story.
Impact. D09 unmet, and the dialog's words disagreed with the behaviour.
Fixed. A player turn with story after it takes the replay path (addTake,
SP9), which is what creates a continuation. A narrator turn keeps PATCH,
because _edit_narration already forks per §§14-15.
Regression. editRouting.test.jsx pins the routing rule for both kinds and
both conditions; D09/D10 in the browser suite.
8. A stale .input-bar { display: flex } — HIGH (cosmetic, total)
play.css is imported after story.css, so a rule left from the old composer
won on order: the direction row and the input row laid out side by side and the
box was unusably narrow. Found by opening the product, not by reading the CSS.
Fixed by removing the superseded rules.
9. Presentation defects found in the first browser pass — MEDIUM/LOW
The sticky header was translucent and story text read through it; the model
badge overlapped the panel tabs; the composer box carried forms.css's 70px
floor; the stored > marker rendered as a Markdown blockquote by accident,
stacking rules. All fixed; the marker has two component tests.
10. Defects in the tests and harnesses themselves — process
Recorded because a test that passes against broken code is worse than no test.
- The first take-stepping regression passed against the unfixed component:
its fixture supplied an
action.branch_idthe API never sends. Corrected, re-run against reverted code, confirmed failing 2/9, fix restored.helpers.jsxdocuments why that field must never return. test_m8_setup_surface.pypassed alone, failed ten ways in the full suite — it wrappedTestClient(app)instead of following the suite's create/drop fixture convention.- The D09/D10 browser harness set
.valueon a React-controlled editor, so the "edit" saved its original text. Now uses the WebDriver clear endpoint and asserts the editor holds the new text before saving. - The migration harness made a correction the validator refused with a 400, then asserted it "survived". It now asserts its own precondition and fails loudly.
- Two acceptance assertions were wrong about established semantics: a failed generation does keep the player's action (M3), and the vocabulary check matched "head" inside the narrator's own prose ("You head north").
11. Evidence-run discipline — process
Three browser evidence runs were invalidated: two by rebuilding mid-run, one by starting a second run before the first had exited (both wrote to the same files, producing results that could not all be true). No product conclusion was drawn from any of them — each was discarded and re-run. The rule this settles on: one frozen build, one suite at a time, no source edits between the freeze and the last suite; if any is broken, the run is discarded rather than reported.
12. The evidence harness needed five corrections before it could be trusted — process
Recorded as a finding in its own right, because it bears on how much of M8 was genuinely verified rather than assumed. Across the milestone the browser harness had five defects, against seven product defects:
| The harness said | The truth | |
|---|---|---|
.turn:last-child |
"no take selector appeared" | the last child of .story is the scroll anchor, so the selector could never match |
| streaming counted as a turn | "the composer is not reachable by keyboard" | the turn had not finished; the composer was correctly still disabled |
.value on a React editor |
"the edited action is not on the story" | true, but for the wrong reason — the edit had saved its original text |
a branch_id the API never sends |
"the take-stepping fix is covered" | the test passed against the broken component |
| Retry must produce different words | "no second take appeared" | the take existed; the model had simply repeated itself |
Two of those five masked real product defects (findings 1 and 7) and were the reason each took two rounds to find. Three asserted properties the product never promised.
The lesson is not that the harness was careless — it is that a browser assertion is only as good as its model of what the product guarantees, and four of the five failures were the harness asserting model behaviour, DOM structure or React internals rather than a product contract. Every corrected assertion now names the guarantee it is testing.
13. Three evidence runs were invalidated by process error — process
Two by rebuilding the frontend mid-run, one by starting a second run before the first had exited — both wrote to the same result files and produced a set of figures that could not all be true at once.
No product conclusion was drawn from any of them. Each was discarded and re-run. A fourth partial result — a migration run stopped during its final check while cleaning up — was likewise discarded rather than reported as complete, after it was noticed that the script defines 17 checks and the output held 16.
The rule this settles on, and which the final set followed: one frozen build, one suite at a time, no source edits between the freeze and the last suite; if any of those is broken, the run is discarded rather than reported. The cost of ignoring it is not a slow run — it is a plausible-looking number that is not evidence of anything.
Invalidated is not the same as superseded, and this report uses the two words
precisely. An invalidated run is one whose figures cannot be trusted at all,
because the build moved underneath it or two runs overwrote each other's output:
the three above, plus the partial migration result. A superseded run is
internally sound but describes an artifact that no longer exists — the
index-C6E5Uvtu.js set, whose 54/55 acceptance result is a true statement about
a build that finding 7 then corrected. The superseded set is cited in §S as the
evidence that findings 3 and 7 existed; it is cited nowhere as evidence that M8
passes. The final set, against index-Ii-lARp9.js, is the whole of §P.
14. The context budget the app plans against was four times the ceiling the server enforced — RESOLVED, operationally, with no application-code change
Status: found in M8, investigated at closeout, resolved. Not an M8 regression and never triggered in M8; recorded in full because M8's browser work is what surfaced it, because the mismatch would have caused a silent correctness failure in a long campaign, and because the resolution is an operator procedure that has to be written down somewhere a deployer will find it.
The original mismatch. The application budgets a prompt up to
Settings.context_token_budget — 16,384 by default. Ollama, seeing no VRAM on
the reference deployment, enforces a 4,096-token input window. Nothing in the
application knew about the gap, and nothing in the interface would have shown a
reader that the front of their prompt was being dropped.
Why sending num_ctx through the OpenAI-compatible endpoint does not work.
The application speaks /v1/chat/completions (ADR 002, ADR 011). That endpoint
accepts num_ctx — nested in options or at the top level — returns HTTP 200,
and ignores it. Worse, it reloads the model at its own default, so a native
/api/chat call that had already established a larger window is undone by the
application's next request. There is no per-request path to a larger window from
where this application stands.
The measured resolution: a derived model. Creating a model with the
parameter baked in makes the window travel with the model rather than with the
request, and it is a normal API call — no shell access on the Ollama host, no
environment variable, no restart. The derived model then appears in /v1/models,
which is exactly the listing M8's Settings model picker already reads, so the
application discovers and uses it through its ordinary path. No Adventure
Storyteller code change is required, and none was made. The measurement table
and the procedure follow.
Evidence. On the reference deployment, Ollama reports:
msg="vram-based default context" total_vram="0 B" default_num_ctx=4096
and /api/ps confirms the loaded model carries context_length: 4096.
Meanwhile the application's own budget is:
Settings.context_token_budget default |
16,384 |
| the value M8's test settings used | 8,000 |
| what the server will actually accept | 4,096 |
The application sends max_tokens, which caps output. It never sends
num_ctx, which is what sizes the input window.
Why it matters. llama.cpp truncates the oldest tokens. This
application puts the narrator rules and campaign canon in the system block at
the front of the prompt (context/builder.py). So the first material dropped
on a long campaign is the highest-authority material in the product — the rules
a turn may not contradict. It would present as the narrator quietly forgetting
canon deep into a session, with nothing in the interface indicating why, and
C01 would begin failing for a reason no UI surface explains.
Not triggered in M8. The largest assembled prompt across every suite was 2,498 tokens, comfortably inside 4,096. Nothing in this report's evidence was truncated. Verified by reading the token totals the context inspector reported on every run.
How this was measured. Four paths were exercised against the reference
endpoint, each followed by reading context_length back from /api/ps:
POST /v1/chat/completions {..., "options": {"num_ctx": 8192}} -> 200, context_length 4096
POST /v1/chat/completions {..., "num_ctx": 8192} -> 200, context_length 4096
POST /api/chat {..., "options": {"num_ctx": 8192}} -> 200, context_length 8192
POST /v1/chat/completions (ordinary, after the native call) -> 200, context_length 4096
The fourth line is the decisive one: the native window does not persist for the application's own requests.
The OpenAI-compatible endpoint does not merely ignore num_ctx — it reloads
the model at its own default, discarding a larger window a native call had
already established. So priming the server through the native API is useless to
this application: its next request resets it.
A derived model does survive, and needs no server shell access. The window travels with the model rather than with the request:
POST /api/create {"model":"qwen2.5:3b-instruct-16k",
"from":"qwen2.5:3b-instruct",
"parameters":{"num_ctx":16384}}
Verified end to end on the reference deployment: after creating it, an
OpenAI-compatible call to the derived model — the application's own request
path, not a native one — loads at context_length 16384, and the model appears
in /v1/models, so it is selectable in M8's Settings picker with no code
change. It shares the base model's blobs, so it costs a manifest.
qwen2.5:3b-instruct-16k is the model the demonstration created. It is a
demonstration, not a dependency. Nothing in the application, the test suites
or the repository requires it to exist, no evidence in this report was produced
with it, and it may be deleted with POST /api/delete at any time. The durable
artifact is the reproducible procedure, which is documented in DEVELOPMENT.md
under "The context window your Ollama actually enforces".
This is the remedy where the server's environment is not editable, and it is an operator action rather than an application change: the provider stays OpenAI-compatible (ADR 002, ADR 011) and no second request path is introduced. No application code was added to work around this, deliberately — a workaround in the provider would mean either a native-API second path, breaking ADR 011's single-endpoint policy, or a request parameter the endpoint provably ignores.
The three paths open to a deployer, and which one this resolves on.
- A derived model, over the API — no shell access, no code change. ADOPTED.
POST /api/createwithparameters: {"num_ctx": 16384}, then select it in Settings. Verified working through the application's own OpenAI-compatible path. Costs roughly 4× the KV cache.OLLAMA_CONTEXT_LENGTH=16384on the server does the same job where the environment is editable. - Set the app's budget to match the server. The field already exists and is editable in M8's Settings screen ("How much story to send"), and its help text already says "Must fit your model's window." What is missing is any way for the reader to know what that window is.
- Detect and warn. The real window is discoverable —
/api/psand/api/showboth reportcontext_length— but only over the native API. A Settings-screen check that reads it and warns when the budget exceeds it would close the gap without moving the generation path.
Final status: resolved, and no longer an M11 blocker. What remains is
operational, not a correctness defect: a deployer whose Ollama enforces a small
window must either raise it by the documented procedure or lower
Settings.context_token_budget to match. Option 3 above — a Settings-screen
check that reads the real window and warns — stays on the table as a
usability improvement, unowned and unscheduled, not as an open defect. M11
should confirm the deployment it certifies against has a window large enough for
a 100-turn campaign before it starts; M9 should be aware that an imported long
campaign reaches the ceiling immediately on a deployment that has not applied
the procedure.
T. Planning-document corrections
Five documents changed. Each is classified, because "the plan changed" and "the plan was wrong about what exists" are different claims.
Requirement clarification (1)
BROWSER-UX-SPEC.md §38 — rewritten from "Hidden Narrator State" with a
Show Hidden Story State toggle to "Hidden Narrator Information / Spoilers".
The protection required is not weakened; it is stated more strictly. §38 now requires, as a list:
- ordinary State and story surfaces must not expose narrator-only information;
- an advanced inspection surface that can contain hidden Canon or narrator-only context must withhold it by default;
- revealing it requires an explicit, clearly labelled user action;
- the UI must warn that doing so may reveal campaign secrets;
- no second hidden-state store or duplicate representation may be introduced solely to give the UI something to toggle.
The last clause is new and is a tightening. What changed is that the section no longer names a state subsystem that does not exist: a secret lives in a narrator-only knowledge source and never enters the state document. §100's checklist entry follows, from "hidden-state inspector" to "spoiler-aware advanced inspection of hidden narrator information".
Ratified. The independent review accepted the change from "Show Hidden Story State" to the architectural requirement on hidden narrator information, and the closeout brief settled it. It is recorded as a requirement clarification aligned with the implemented architecture, not as a relaxation: the obsolete wording named a hidden narrative-state subsystem that does not exist in this product, and the five requirements above bind the surfaces that genuinely can expose a secret — including the clause forbidding a duplicate hidden-state store invented merely to give the UI something to toggle. The protection is verified end to end by the sentinel suite (§P, 21/21), which proves among other things that withheld material is absent from the DOM rather than collapsed inside it.
Implementation facts (4)
TECHNICAL-DESIGN.md§13.4 — the browser stayed a presentation layer: server-authoritative availability, reading is not deciding, UI state stays UI state; the failure taxonomy; and why markup is never produced from input.TECHNICAL-DESIGN.md, campaign canon — canon is configuration and its provenance is the per-turn context snapshot. Measured, not assumed, with the measurements recorded.BROWSER-UX-SPEC.md§12, §31, §55, §57 — four "As implemented in M8" notes, following the convention M7 established. No requirement altered.BUILD-MILESTONES.md,V1-ACCEPTANCE-TESTS.md,VERSION.md,planning/README.md,planning/archive/README.md,README.md,DEVELOPMENT.md— M8's outcome, the browser-evidence notes on the B and D series, the M7 report's rotation to the archive, and the frontend test suite replacing "no frontend tests" in the inherited-debt list.
Changed product requirements (0)
None. No acceptance test was weakened to match the implementation. Where the
implementation and a document disagreed, the disagreement is recorded rather
than resolved in the implementation's favour — see §U's note on §38, and the
action_count observation in §F.
Status discipline
VERSION.md and BUILD-MILESTONES.md describe M8 as implemented, verified,
reviewed and accepted, dated 2026-09-06, and record M9 as next. Nothing
marks M9 started, and no M9 scope was implemented.
U. Residual risks and deferred work
Genuine residuals only. M9-M11 planned scope is not listed here as an M8 defect.
Residual risks
-
errors.jsis coupled to backend message strings. Written as a fallback ladder, so an unrecognised message still classifies as a generation failure, still shows the server's own words, and still offers Retry — nothing is hidden when the match misses. But a backend rewording silently loses a tailored hint. The signatures are quoted beside each branch so the coupling is findable from either end. Risk: low. A degraded hint, never a hidden failure. -
Third-person player input still takes the
> Youprefix. "Aldric draws his knife" becomes> You Aldric draws his knife.It is not one of §11's three stated inputs, and it is unchanged from before M8 — a name-detector would be wrong more often than the current rule. Risk: low. M11's 100-turn run is where it would become clear whether readers write this way. -
A deployment's context ceiling can be a quarter of the app's budget — resolved operationally. See finding 14, now closed. Ollama's OpenAI-compatible endpoint ignores
num_ctx, so the window cannot be set per request; a derived model created over/api/createcarries it and is honoured through the application's own path, needs no shell access and no code change, and appears in the Settings model picker automatically. The procedure is inDEVELOPMENT.md. Residual risk: low, and operational rather than architectural — a deployer who applies neither the procedure nor a matchingcontext_token_budgetstill gets silent truncation, and the application still has no way to tell them so. A Settings-screen warning that reads the real window remains an unowned usability improvement, not a defect. Nothing in M8 was truncated: the largest assembled prompt across every suite was 2,498 tokens against a 4,096 ceiling. -
Contrast and visible focus were checked by eye, not measured. Risk: low for a local single-user product; belongs to M11.
-
Tablet is usable but untuned. Desktop was the stated primary target and the narrow-screen sheet is inherited. Risk: low, unless v1 claims tablet support.
Deliberately deferred, with the milestone that owns it
| Owner | |
|---|---|
| Story cards have no browser editor — backend and bundle keep them | M9 should decide whether the bundle keeps carrying a subsystem no v1 UI exposes |
| The bundle still carries no context snapshots, so an imported campaign has no historical prompt provenance | M9 (recorded by M7) |
| The RPG world state is read-only | M10/M11, if a genre wanting numbers appears |
| Copy is per-message only; no whole-transcript copy (§77) and no story search (§78) | M9 or M11; §78 says "not essential to initial v1" |
| No discarded-history recovery screen (§63) | explicitly future work |
The inert api_key column and the dual-dialect migration code |
the cleanup migration BUILD-MILESTONES.md already schedules |
M9 handoff — what the next brief has to decide
Recorded here because M8 is the milestone that surfaced each one, and because a brief written from a stale report would let the inherited bundle format decide them by default. None of these is implemented, and none was implemented in M8.
A. Complete campaign portability. M9 must validate export, backup, import and recovery of the whole campaign, not just its text: retained story history, the active head, alternate takes, Save Points, the authoritative narrative state, the state provenance and audit data the product contract requires, summaries and memories as appropriate, imported knowledge with its provenance, and the campaign/profile/canon settings a campaign needs to resume correctly. M8 touched none of this — it changed no schema, no bundle code and no endpoint — but it is the milestone that put canon and the opening into the browser, so a campaign now carries setup a reader chose rather than a scenario supplied.
B. Historical prompt and context provenance. Every turn stores the context it was actually given, and §F leans on that: it is what proves a canon edit was not retroactive. The bundle does not carry it (recorded by M7, unchanged by M8). M9 must decide from the authoritative specification and data model whether those snapshots — or a reproducible equivalent — belong in the supported portable bundle. If they do not, the "every turn can prove what it was told" guarantee is local to the machine that played it, and that should be a stated decision rather than a side effect of the inherited format.
C. Legacy story cards. The backend and the bundle still carry story cards; M8 removed their browser surface, because the M7 knowledge library supersedes them for v1 and showing both would offer two unrelated systems for "things the narrator should know". M9 should decide explicitly which of these the v1 portable campaign does — compatibility-only carriage, migration into the M7 knowledge subsystem, preservation as inert legacy data, or another documented behaviour consistent with the active specification — and record the choice.
D. Context-window portability. See finding 14. A restored campaign may run
against an Ollama whose effective window differs from the one that wrote it, so
M9 must not treat campaign restored successfully as the destination narrator
has the same capacity. Concretely: do not hard-code an assumption that 16,384
tokens are available; preserve the campaign and model configuration the bundle
contract requires accurately, so a destination can at least be compared against
the source; and note that an imported long campaign meets a small ceiling
immediately on a deployment that has applied neither the DEVELOPMENT.md
procedure nor a matching context_token_budget. Actual long-campaign and
context-window release validation remains M11's, unless a later milestone
explicitly moves it.
A note on how much of this was verified rather than assumed
Seven product defects were found and fixed; five harness defects had to be fixed before the harness could find them, and two of those five were actively masking product defects (§S findings 12-13).
A reviewer is entitled to ask what else the harness is not asserting. The honest answer is that the browser suites assert what the product promises, and the places where they previously asserted something else — model wording, DOM structure, React internals — have each been rewritten to name the guarantee under test. What they do not cover is stated in §M (contrast and focus rings were looked at, not measured) and §U (tablet was not tuned).
The one requirement question, and how it was settled
BROWSER-UX-SPEC.md §38 was rewritten, and the rewrite has been ratified.
It is the one place M8 changed the wording of a requirement rather than
recording a fact about the implementation, so it was put to the reviewer
explicitly. The independent review accepted it, and the closeout brief settled
it: a requirement clarification aligned with the implemented architecture, not a
weakening. §T carries the five requirements it now states and the argument that
they are stricter than what they replaced. Nothing about this is left open.
V. Final milestone assessment
Is M8's Definition of Done satisfied?
BUILD-MILESTONES.md states it as: "Normal story creation and play feels like a
focused local storyteller rather than an RPG or developer console."
Yes, on the evidence in §P. A reader creates a campaign from a form that never mentions a scenario, plays through one natural-language field, and reaches State, Knowledge, Context and Save Points through panels that start closed. The words branch, fork, node, head and depth appear nowhere they operate — audited in source and observed in the live DOM.
The twenty conditions in the M8 brief's own §48 are covered by §P's matrix.
Does it feel like the intended storyteller rather than adapted AI-DnD?
The measurable parts say yes. The navigation lost Home · Adventures · Scenarios · Settings · AI Chat for Campaigns · Settings. The context
inspector went from 11,996 characters to about 1,400. Sixteen unnamed glyph
buttons became named controls. Three input modes became one field. The Branches
tab, the tree overlay, the stat-schema editor, the art picker, the story-card
table and the raw model console are gone.
The unmeasurable part — whether it reads like a storyteller — is a judgement, and it is the reviewer's to make. §E describes what a person actually meets.
Are any required browser behaviours unverified?
No. Every M8-owned acceptance condition was demonstrated in a real browser against a real narrator on the final build: 157 checks across six suites, zero failures. D09 and D10 — the two that were outstanding longest, and the two a harness defect had been masking — are verified end to end.
Three things are recorded as looked at rather than measured, and none is a required M8 condition: contrast and visible focus rings (§M), reading order at 1440×960 (§M), and tablet layout (§U). Each is named as such rather than folded into a pass.
Is it safe to proceed to M9 after review?
M8 changed no schema, no migration, no listener binding, no endpoint policy and no outbound network behaviour (§O). M9's subsystems — bundle format, backup, recovery — were not touched. Nothing in M8 constrains M9 beyond the two notes in §U that M9 should decide deliberately rather than inherit.
What must happen before M9?
| Status | |
|---|---|
| This report reviewed, and M8 accepted or returned | done — the independent review returned M8 IMPLEMENTATION: PASS, subject to evidence and documentation cleanup |
| The §38 rewrite ruled on (§T, §U) | done — ratified as a requirement clarification aligned with the implemented architecture |
| The build-evidence contradiction resolved | done — §P classifies index-C6E5Uvtu.js as superseded and index-Ii-lARp9.js as the one final frozen artifact, from the saved run logs |
| Finding 14 dispositioned | done — resolved operationally, no application-code change, procedure in DEVELOPMENT.md |
| The repository owner signs the M8 commit | outstanding — the only remaining step |
The tree is staged and uncommitted. This project's policy is that milestone commits are signed by the owner, and nothing here was committed on their behalf. After the signed commit, the expected next step is a brief verification of the committed tree, and then M9 may begin.
M8 closeout status: implemented, verified, reviewed, accepted (2026-09-06).
M9 has not been started, and no M9 scope was implemented early.