A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
1792 lines
92 KiB
Markdown
1792 lines
92 KiB
Markdown
# M8 — Browser UX Completion for v1 Story Operations
|
||
|
||
**Final implementation, verification and closeout report.** Written by the
|
||
implementer for an independent reviewer, and completed at closeout after that
|
||
review returned **M8 IMPLEMENTATION: PASS**. The closeout additions are the
|
||
build-evidence classification in §P, finding 14's resolution, and the acceptance
|
||
record in §V; no verification figure was changed by them.
|
||
|
||
**Closeout date: 2026-09-06.**
|
||
|
||
> **Placeholders.** `inference.lan` stands for the trusted-LAN Ollama host and
|
||
> `192.168.0.x` for LAN addresses, following the convention M1 established. No
|
||
> real hostname, address or identifier of the machine this was built on appears
|
||
> in any committed file.
|
||
|
||
---
|
||
|
||
## A. Executive result
|
||
|
||
**PASS**
|
||
|
||
Every M8-owned acceptance condition was demonstrated through the browser
|
||
workflow against a real local narrator, on one frozen production build. No
|
||
required condition is unverified, and there are no blockers to acceptance.
|
||
|
||
**Two qualifications a reviewer should weigh, neither of them a blocker:**
|
||
|
||
1. **Seven product defects were found during verification, four of them by
|
||
driving the application rather than reading it** (§S findings 1-4, 7). Two
|
||
made the interface state something untrue to the reader. All are fixed with
|
||
regression coverage, and two of the regressions were verified to fail against
|
||
the unfixed code before the fix was kept. But their existence says the
|
||
implementation pass was not as verified as it looked at the time.
|
||
2. **The evidence harness needed five corrections before it could be trusted**
|
||
(§S finding 12), and three evidence runs were invalidated by process error
|
||
(§S finding 13). Two of those harness defects were actively masking product
|
||
defects. The figures below are from the corrected harness on the final build;
|
||
the history is recorded so a reviewer can judge how much they rest on.
|
||
|
||
**M8 is implemented, verified, reviewed and accepted.** The independent review
|
||
returned `M8 IMPLEMENTATION: PASS` subject to evidence and documentation
|
||
cleanup; that cleanup is §P's build classification and finding 14's resolution,
|
||
both complete. M8 is **not yet committed**: the tree is staged for the
|
||
repository owner's signature, which is the only step left before M9.
|
||
|
||
### What this milestone is, in one paragraph
|
||
|
||
M8 turned an adapted AI-DnD interface into the interactive-story workspace
|
||
`BROWSER-UX-SPEC.md` describes: one entry point, one story screen, one
|
||
natural-language input, and everything advanced one layer deeper behind panels
|
||
that start closed. Almost nothing underneath changed — the play loop, the
|
||
history controls, the takes, the Save Points, the state engine and the knowledge
|
||
library are M3-M7's and are untouched. What changed is what a reader is shown
|
||
and asked to understand.
|
||
|
||
### Evidence at a glance
|
||
|
||
| Gate | Result |
|
||
| --- | --- |
|
||
| Backend suite | **950 passed, 14 skipped, 0 failed** |
|
||
| Frontend component suite | **132 passed, 10 files** |
|
||
| Lint | exit 0 |
|
||
| Production build | clean |
|
||
| Docker build | clean |
|
||
| Browser acceptance (6 scenarios, B01-B04, D01-D14) | **57/57** |
|
||
| Hidden-information sentinel | **21/21** |
|
||
| A05 failed generation | **23/23** |
|
||
| Security / offline / performance | **27/27** |
|
||
| Genuine process restart | **12/12** |
|
||
| Migration against an M7-built database | **17/17** |
|
||
| Reader-facing terminology audit | **0 hits** (17 normal-play components; both advanced surfaces) |
|
||
|
||
Every browser figure is from one frozen production build,
|
||
**`dist/assets/index-Ii-lARp9.js`**, built from the staged tree at 18:53:02 and
|
||
unchanged for the rest of the campaign. An earlier artifact,
|
||
`index-C6E5Uvtu.js`, is **superseded**: it predates finding 7's fix, its
|
||
acceptance suite ended 54/55 on that defect, and no figure above comes from it.
|
||
§P sets the two side by side with the evidence for the classification.
|
||
|
||
### What a reviewer should look at first
|
||
|
||
1. **§S findings 1, 2, 3 and 7** — four product defects that the browser work
|
||
found and that no earlier milestone had caught. Two of them made the
|
||
interface state something untrue to the reader.
|
||
2. **§S findings 12 and 13** — the harness needed five corrections before it
|
||
could be trusted, and three evidence runs were invalidated by process error.
|
||
This bears directly on how much weight the figures above deserve.
|
||
3. **§F** — the `canon_rules` decision, where the smallest correction was a
|
||
surface constraint rather than an audit subsystem, and why.
|
||
4. **§T** — the one requirement whose wording M8 changed (§38), and the argument
|
||
that it was tightened rather than weakened.
|
||
|
||
---
|
||
|
||
## B. Repository and provenance state
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Branch | `m8-browser-ux` |
|
||
| M8 base | `480414efe082a4bfe0600a19fe23961f6bddd925` — M7 |
|
||
| M7 signature | **Good signature**, verified against the repository owner’s RSA key `02C9BF7D…B5C68569`, ultimate trust |
|
||
| M7's parent | `a6e9c7a32bdf42f1e4cb837b70721d89422cef8d` (M6) — sole parent |
|
||
| Current HEAD | `480414e…` — **still M7. M8 has made no commit.** |
|
||
| Staged | **89 files** — 34 added, 24 deleted, 30 modified, 1 renamed (`git diff --cached --stat`) |
|
||
| Unstaged | 0 |
|
||
| Untracked | 0 |
|
||
| `LICENSE` | unchanged, absent from the staged diff |
|
||
| Upstream ancestry | `d72f7c1b…` (AI-DnD) is still an ancestor |
|
||
|
||
### Committed HEAD vs staged tree vs tested tree
|
||
|
||
These are three different things and the distinction matters for review:
|
||
|
||
- **Committed HEAD** is M7, untouched. Nothing in M8 has been committed.
|
||
- **The staged tree** is the whole M8 result — 89 files, including this report
|
||
and the planning corrections.
|
||
- **The tested tree** is the staged tree. The frontend was built from it once
|
||
and frozen; every browser figure in this report ran against that one artifact,
|
||
**`dist/assets/index-Ii-lARp9.js`** (built 18:53:02, and still the only file
|
||
in `frontend/dist/assets/`), and no source was edited between the freeze and
|
||
the last suite.
|
||
|
||
An earlier artifact, `index-C6E5Uvtu.js`, appears in the evidence directory
|
||
and is **superseded, not final**: it predates finding 7's fix, its acceptance
|
||
suite ended 54/55 on exactly that defect, and no figure in this report comes
|
||
from it. §P records the classification and the evidence for it.
|
||
|
||
That last point is stated because it was broken three times during M8 and each
|
||
time the run was discarded rather than reported. See findings 11 and 13.
|
||
|
||
**M8 is implemented, verified, reviewed and accepted** (2026-09-06). What is
|
||
outstanding is the owner's signed commit, not a decision.
|
||
|
||
---
|
||
|
||
## C. The M7 browser baseline, and how it directed the work
|
||
|
||
Measured **before any code was changed**, by driving the M7 build in a real
|
||
Firefox against a real backend with the Continuity Test fixture and two real
|
||
turns played (`scratchpad/m8/baseline.py`, 13 observations). Nothing here is
|
||
recalled; each figure is a browser observation.
|
||
|
||
| Observation | Measured | What it directed |
|
||
| --- | --- | --- |
|
||
| **Context inspector size** | **11,996 characters** on opening, beginning with the assembled prompt | §55 says do not begin with raw prompt text. Drove the whole restructure: four readable sections first, the prompt last and collapsed. Result ~1,400 chars (§R). |
|
||
| **Unlabelled controls** | **16 of 16** per-message buttons had no accessible name — single glyphs `🔍 ✎ ⑂ ✕` with only a `title` | Drove named controls throughout, and the `a11y.test.jsx` assertion that a name under three characters counts as a glyph. |
|
||
| **Branch UI visible in normal play** | panel tabs read `Story State · Plot · Memory · Knowledge · **Branches** · Save Points · Insights` | §27. Drove removing the branch panel and tree overlay from the browser, and the terminology audit (§M). |
|
||
| **Rigid input modes** | `Do` / `Say` / `Story` | §12. Drove one natural-language field plus the direction toggle — and, indirectly, exposed the `> You I …` defect (finding 4). |
|
||
| **No model status anywhere** | play header read `Continuity Test / Story State / Plot / …` and said nothing about Ollama | §44 and §8. Drove the five-state badge and the setup notice. |
|
||
| **Navigation** | `Home · Adventures · Scenarios · Settings · AI Chat` | §97. Drove `Campaigns · Settings` with everything else inside a campaign. |
|
||
| **Landing page vocabulary** | "The table is set", "worlds to explore", a scenario gallery | §26. Drove the campaign library. |
|
||
| **Settings** | `Model` a free-text box; a raw request/response log un-collapsed at the bottom | §8 and §72. Drove the model picker and the folded diagnostics. |
|
||
| **Campaign creation** | required choosing a Scenario; its editor held a JSON stat schema, a story-card table and an art picker | §41-42 and §93. Drove the one-screen setup form and the `opening` field. |
|
||
| **Frontend test coverage** | none at all; two backend tests read JSX as text to approximate browser assertions | §30. Drove the whole test foundation (§N). |
|
||
|
||
### What the baseline showed was already right
|
||
|
||
Worth recording, because it bounded the work: the play loop, streaming, Undo,
|
||
Redo, Retry, alternate takes, Save Points, state correction, knowledge import
|
||
and prompt inspection were all reachable and all correct. The Save Point panel's
|
||
confirmations already said the right things. The State panel already avoided raw
|
||
JSON. Undo and Redo already used the server's answer rather than deriving
|
||
availability in React.
|
||
|
||
M8 changed **what a reader is shown and asked to understand**, not the
|
||
machinery. The two exceptions are findings 1 and 7, where the presentation layer
|
||
turned out to be calling the wrong mechanism — and both were caught by driving
|
||
the product, not by reading it.
|
||
|
||
---
|
||
|
||
## D. M8 change inventory
|
||
|
||
89 files — 34 added, 24 deleted, 30 modified and 1 renamed. The frontend was
|
||
substantially rebuilt; the backend was touched in five places, each to expose a
|
||
browser capability that could not otherwise be reached.
|
||
|
||
### Browser presentation — added
|
||
|
||
| File | What it is |
|
||
| --- | --- |
|
||
| `pages/Campaigns.jsx` | the campaign library — the landing page |
|
||
| `pages/NewCampaign.jsx` | genre-neutral setup, one screen |
|
||
| `pages/Play/Transcript.jsx` | the story, extracted from the page component |
|
||
| `pages/Play/Composer.jsx` | one input, direction toggle, history controls, reserved dictation |
|
||
| `pages/Play/FailureNotice.jsx` | §71's five failure kinds |
|
||
| `pages/Play/SidePanel.jsx` | the secondary panel and its tabs |
|
||
| `pages/Play/panels/ContextPanel.jsx` | the context inspector, restructured |
|
||
| `pages/Play/panels/CampaignSettingsPanel.jsx` | campaign settings, canon, export, delete |
|
||
| `Dialog.jsx` | accessible modal with a real focus trap |
|
||
| `ExternalLinkDialog.jsx` | §74's leaving-the-local-environment warning |
|
||
| `ModelStatusBadge.jsx`, `ModelSetupNotice.jsx` | §44 status, and §8's way out of a blank model |
|
||
| `markdown.jsx` | safe Markdown for story prose |
|
||
| `errors.js` | the failure taxonomy |
|
||
| `modelStatus.jsx` | one shared model-status answer, fetched once |
|
||
| `styles/story.css`, `library.css`, `context.css`, `dialogs.css` | the new surfaces |
|
||
|
||
### Browser presentation — removed
|
||
|
||
Sixteen components and eight stylesheets, each recorded with its reason in the
|
||
file that replaced it.
|
||
|
||
| Removed | Why |
|
||
| --- | --- |
|
||
| `pages/Home.jsx`, `pages/Adventures.jsx` | two screens listing the same rows; one library replaces both |
|
||
| `pages/Scenarios.jsx`, `pages/ScenarioEditor.jsx`, `SchemaEditor.jsx`, `ArtPicker.jsx` | a campaign no longer needs a template, and the editor was §93's "dangerous advanced features" almost exactly |
|
||
| `pages/Chat.jsx` | a raw model console in the primary navigation |
|
||
| `panels/BranchPanel.jsx`, `BranchMap.jsx`, `branches.js` | §27 — branch management is not a v1 surface |
|
||
| `panels/PlotPanel.jsx`, `RefreshModal.jsx` | AI Dungeon world-info editing; the M7 knowledge library supersedes it |
|
||
| `panels/MemoryPanel.jsx` | memory is now shown where it is used, in the context inspector |
|
||
| `panels/InsightsPanel.jsx` | renamed and restructured as `ContextPanel.jsx` |
|
||
| `drawers/WorldStateDrawer.jsx` | the RPG stat editor (§18, §26) |
|
||
| `Embers.jsx` | decorative fire; §40 |
|
||
| 8 stylesheets | the surfaces they styled; `auth.css` reduced to `debuglog.css` |
|
||
|
||
**No backend capability was removed.** The tree, branch switching, story cards
|
||
and the RPG world state all still exist, are still tested, and still travel in
|
||
the bundle.
|
||
|
||
### Backend support — five files, no schema change
|
||
|
||
| Change | Why the existing API could not serve the UX |
|
||
| --- | --- |
|
||
| `AdventureCreate.opening` | a `start` action could only come from a Scenario's prompt, so every campaign made in the new setup flow opened on a blank page |
|
||
| `canon_rules` on create, update and read | `campaign_canon` has been read by the prompt builder and the state validator since M5 and had **no API at all** — a fixture had to write it with SQL |
|
||
| `format_player_input` | finding 4 |
|
||
| `_friendly_http_error` 401/429 | the last user-facing text describing a hosted deployment |
|
||
| `models.Adventure.canon_rules` | a read-only property. **No column, no migration.** |
|
||
|
||
### Test infrastructure
|
||
|
||
Vitest + jsdom + Testing Library (5 devDependencies, no runtime dependency); ten
|
||
frontend test files; two new backend test files
|
||
(`test_m8_setup_surface.py`, and the normalization tests in
|
||
`test_take_parentage.py`); two inherited source-level guards retargeted or
|
||
replaced — see §N.
|
||
|
||
### Documentation
|
||
|
||
`README.md`, `DEVELOPMENT.md`, and six planning documents — itemised in §T.
|
||
|
||
---
|
||
|
||
## E. The final user experience, in user terms
|
||
|
||
A person opens the application and sees **their campaigns** — titles, when each
|
||
was last played, how many moments it holds, and the opening of its most recent
|
||
narration. No ids. Two buttons: import one, or start a new one.
|
||
|
||
Starting one asks for a **name**, and nothing else is required. If they want to,
|
||
they can say the genre and tone (free text, with suggestions spanning fantasy,
|
||
science fiction, mystery, historical, western, horror, thriller and literary),
|
||
choose a voice and a narration length, name a protagonist, write the opening
|
||
scene, and write the rules the story must not contradict. Then Start.
|
||
|
||
Then they are in the story, and the story is the whole screen. A thin bar at the
|
||
top carries the campaign's name, whether Ollama is connected and with which
|
||
model, and five tabs that are all closed. Underneath, prose at a readable measure
|
||
— the reader's own turns indented and italic, the narrator's plain, an ornament
|
||
between scenes, a drop cap on the opening.
|
||
|
||
They type what they do into one box. Not a mode, not a command — *"I walk into
|
||
the Crooked Lantern and look for Mara."* The reply streams in. If they want the
|
||
narrator to do something rather than their character, they tick **Story
|
||
direction** and the box says so.
|
||
|
||
Above the box: Continue, Retry, Undo, Redo, Save Point. Undo and Redo grey out
|
||
when there is nowhere to go. Hovering a turn reveals Inspect context, Edit, Try
|
||
again, Copy — named, not glyphs.
|
||
|
||
When something is wrong they are told which thing: the model is unreachable and
|
||
here is how to start it; the turn failed and here is Retry, with their words
|
||
still in the box; the story state could not be updated and the story itself is
|
||
fine.
|
||
|
||
When they want to know **why the narrator said that**, one click on the turn
|
||
opens what it was given: what it cost, what it read, what it remembered, what it
|
||
believes. If a passage was marked narrator-only, its text is not there — a
|
||
control offers it, and says it will reveal secrets they marked for the narrator
|
||
alone.
|
||
|
||
They can play for an hour without meeting the word *branch*.
|
||
|
||
---
|
||
|
||
## F. Campaign setup, opening, and canon
|
||
|
||
### The `opening` field
|
||
|
||
A `start` action could previously only come from a Scenario's prompt. M8's setup
|
||
flow creates a campaign from a form, so without this every new campaign opened
|
||
on a blank page — the reader had to invent the situation *and* the first move in
|
||
one box.
|
||
|
||
It builds the same node by the same path (`attempts.snapshot_outcome` →
|
||
`tree.place_action`). Verified, 13 checks, now permanent in
|
||
`backend/tests/test_m8_setup_surface.py`:
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Setup can provide it, as a `start` action | PASS |
|
||
| Never duplicated on re-read | PASS |
|
||
| Placed at depth 0 on the tree, with a parent of `None` | PASS |
|
||
| The head is at it; Undo past the opening is refused (HTTP 400) | PASS |
|
||
| A blank opening still yields an empty campaign | PASS |
|
||
| A Scenario's prompt takes precedence, and the two are never both inserted | PASS |
|
||
| A scenario-made campaign is unchanged by the new field | PASS |
|
||
| Survives export and import | PASS |
|
||
|
||
One observation, **not** an M8 regression: the create response reports
|
||
`action_count: 0` while carrying one action, because the count is computed on
|
||
read. A scenario-made campaign has always done the same. The browser navigates
|
||
and re-reads, so a reader never sees it; the test asserts on the read path.
|
||
|
||
### `canon_rules`, and what editing canon after play actually does
|
||
|
||
`campaign_canon` has been the highest authority in a campaign since M5 — read by
|
||
both the prompt builder and the state validator — and had **no API at all**. A
|
||
fixture had to write it with SQL. M8 exposed the sentence list.
|
||
|
||
That raised a question the column had never had to answer, and §6 of the review
|
||
brief asked for it to be settled rather than assumed. It was **measured**
|
||
against a real server with real turns (`scratchpad/m8/canon_provenance.py`):
|
||
|
||
| Question | Answer |
|
||
| --- | --- |
|
||
| Does a historical turn keep the canon it was actually told? | **Yes** — the canon section is in that turn's stored context snapshot. After the edit it still shows the old rule and not the new one. |
|
||
| Does the next turn get the new canon? | Yes — which is the point of editing it. |
|
||
| Is any accepted story text rewritten? | No — byte-identical. |
|
||
| Is the narrative state document changed? | No — byte-identical. |
|
||
| Is the state audit log changed? | No — byte-identical. |
|
||
| **Is the edit itself audited?** | **No.** It is a configuration overwrite. |
|
||
| Is any other campaign setting audited? | No — editing narrator instructions is not either. |
|
||
|
||
### Why M5's audit log was deliberately not extended
|
||
|
||
M5's `StateEvent` log audits accepted changes to *narrative state*. Canon is not
|
||
narrative state: it is configuration, sitting with `ai_instructions` and the
|
||
narrator prompt. Routing a configuration change through that log would create a
|
||
second representation of canon — the same duplication `BROWSER-UX-SPEC.md` §38
|
||
forbids for hidden information, arrived at from a different direction — and §6
|
||
of the brief explicitly rules out inventing a parallel audit subsystem.
|
||
|
||
So the correction was made at the editing surface instead, which is the brief's
|
||
second option. Once a campaign has moments, the canon editor states that the
|
||
change applies **from here on**, that everything already written stays exactly as
|
||
it is, and where the per-turn record can be seen. Three component tests cover it.
|
||
|
||
The guarantee M8 offers is therefore not "the edit is logged" but **"the edit
|
||
cannot be mistaken for a retroactive one, and every turn can prove what it was
|
||
told"** — which is what the per-turn snapshot already delivered, unasked.
|
||
|
||
---
|
||
|
||
## G. The story composer
|
||
|
||
### One field, not three modes
|
||
|
||
The inherited composer had a `Do` / `Say` / `Story` selector, and the mode
|
||
changed what the turn *meant*: "I say to Mara…" typed in Do mode and the same
|
||
words in Say mode were different turns, and nothing on screen said so. That is a
|
||
small command language wearing buttons.
|
||
|
||
M8 sends one natural-language field. B01 and B02 are the proof that it works —
|
||
an action and a piece of quoted dialogue are both just what the reader wrote,
|
||
and the narrator reads the quotes without being told which kind of turn it is.
|
||
|
||
What survives from the old `story` mode is the **Story direction** toggle, and
|
||
it is deliberately not a fourth mode: it changes *who is being spoken to*, not
|
||
what kind of action is taken. The box is visibly marked while it is on, and the
|
||
send button reads "Direct" rather than "Send".
|
||
|
||
### The `> You I enter the tavern.` defect
|
||
|
||
**Severity: high. Found by driving the new composer in a browser; fixed here.**
|
||
|
||
AI Dungeon stores a player action with a `> You ` prefix. That was right when
|
||
the Do mode asked for a bare verb phrase — `look around` became
|
||
`> You look around.` — and it is what marks whose turn it is in the replayed
|
||
prompt.
|
||
|
||
`BROWSER-UX-SPEC.md` §11 tells the reader to write *"I enter the tavern."* With
|
||
one field, the prefix produced:
|
||
|
||
```text
|
||
> You I enter the tavern.
|
||
```
|
||
|
||
in the transcript, in the replayed history, and therefore **in the narration**,
|
||
where a small model imitates the pattern it is shown and writes "You I thank
|
||
her". The M7 baseline transcript contains exactly that phrasing, so the defect
|
||
predates M8 in the data — but M8's own design is what made it reachable on every
|
||
turn, which is why M8 owns it.
|
||
|
||
The correction adds the subject only when the reader has not already written
|
||
one. The `>` marker — which is what actually identifies a player turn — is
|
||
unchanged in every case.
|
||
|
||
Regression coverage is against the **shared normalizer**, not one rendered
|
||
component, because storage, the transcript, the replayed history and the export
|
||
all read the result of that one function:
|
||
|
||
- §11's three examples verbatim;
|
||
- eight first-person phrasings, each asserted to contain no duplicate subject
|
||
and to keep its marker;
|
||
- the legacy bare action still normalizes (`open the door` → `> You open the
|
||
door.`), which is deliberate compatibility;
|
||
- an explicit `You …` is de-duplicated rather than doubled;
|
||
- dialogue and out-of-character direction untouched;
|
||
- and a second test asserts on the **assembled prompt**, that each player line
|
||
appears exactly once and `"You I "` appears nowhere — because a normalizer
|
||
that is correct but applied twice would put the defect straight back in front
|
||
of the model.
|
||
|
||
### The storage marker is not shown, and not parsed
|
||
|
||
`>` is also Markdown for a blockquote, so leaving the stored marker in made every
|
||
player turn render as a quote *by accident*, stacking the renderer's rule on the
|
||
one `.turn-player` already draws. The marker is stripped for display only; the
|
||
stored text keeps it, and editing a turn puts it back around the edited words.
|
||
|
||
### The reserved dictation control
|
||
|
||
Present, permanently disabled, named "Dictate — not yet available". No handler,
|
||
no permission request, no microphone. A component test tabs twelve times to
|
||
assert it never takes focus, and another installs a fake `getUserMedia` and
|
||
clicks the disabled button to assert it is never called.
|
||
|
||
---
|
||
|
||
## H. History controls
|
||
|
||
M3's active-head machinery, M4's Save Points and the take/divergence rules are
|
||
unchanged. M8 changed how they are presented and what they are called.
|
||
|
||
### Undo and Redo
|
||
|
||
Availability is **the server's answer**, carried on every window it returns, and
|
||
the browser never computes it. Neither is derivable client-side: Undo can reach
|
||
past the top of the loaded page, and Redo depends on a retained future the
|
||
transcript is never sent. Six component tests cover the enabled states,
|
||
including that each is independent and that every control is disabled while a
|
||
turn is generating.
|
||
|
||
### Retry and takes
|
||
|
||
Retry sits in the story controls and "Try again" on each narrator turn. When a
|
||
turn has more than one attempt the pager appears as `‹ 2/3 ›`, reading
|
||
"Take 2 of 3" to a screen reader — §14 asks for a count, and `2/3` is too terse
|
||
to hear.
|
||
|
||
**A defect here had existed since the pager was written.** See finding 1: every
|
||
step between takes performed a branch switch, which for two takes of an ordinary
|
||
retry meant switching to the line already being read — the same window came back
|
||
and the step did nothing at all. D07 is REQUIRED FOR V1 and had never been
|
||
exercised in a browser. Fixed, and confirmed PASS in the final run.
|
||
|
||
### Save Points
|
||
|
||
M4's panel needed little. Create with a name and an optional note, list, restore,
|
||
rename, delete. Both confirmations say what does **not** happen, because that is
|
||
the part a reader cannot see and would otherwise assume the worst about:
|
||
|
||
> The story will return to this Save Point. Everything you wrote after it is
|
||
> kept — it just stops being where you are.
|
||
|
||
> Delete this Save Point? Deleting it does not delete any of the story — only
|
||
> the name you gave this moment.
|
||
|
||
Counted in **moments**, not turns, matching the vocabulary M4's closeout settled.
|
||
|
||
### Vocabulary
|
||
|
||
`branch`, `fork`, `node`, `merge`, `head` and `depth` appear nowhere a reader
|
||
operates. The branch panel and the tree overlay are gone from the browser
|
||
entirely. The mechanism underneath is untouched and still fully tested — this is
|
||
a decision about what a reader is asked to understand, not a reduction of what
|
||
the product can do.
|
||
|
||
The audit method and its result are in §M.
|
||
|
||
---
|
||
|
||
## I. State
|
||
|
||
The server groups the authoritative state and the panel renders the groups, so
|
||
the headings are whatever the campaign has established rather than a fixed list
|
||
— a campaign with no items shows no Items heading.
|
||
|
||
- **No raw JSON.** Asserted: the panel contains no `<pre>` and no `{`.
|
||
- **No RPG vocabulary of its own.** Asserted against `hp`, `mana`, `quest`,
|
||
`stat`, `cooldown`, `xp`. The product term is **Story Threads**, never Quests.
|
||
- **Correction without JSON.** "Correct something" asks for a sentence — the one
|
||
shape a person can write without knowing the event vocabulary — and it goes
|
||
through the same validator a narration's proposal does. C04 was exercised in
|
||
the browser end to end: the correction appears in the panel, reaches
|
||
authoritative state, is recorded as `manual_correction`, and the next turn's
|
||
assembled prompt contains it.
|
||
- **It follows the head.** Keyed on the same signal as the other live panels, so
|
||
Undo, Redo and a Save Point restore all move it — what it shows is the state at
|
||
the position being read, not at the newest turn.
|
||
|
||
### Hidden state
|
||
|
||
There is none, and M8 did not invent one. A secret lives in a narrator-only
|
||
knowledge source and never enters the state document; M7 verified that against a
|
||
real narrator, and the sentinel run re-verified it here (§Q). §38's requirement
|
||
is met by this panel simply not containing narrator-only information. The
|
||
surface that must *actively* withhold it is the context inspector — see §K.
|
||
|
||
### The RPG world state
|
||
|
||
Read-only, and shown in the context inspector only for a campaign that has a
|
||
stat schema. A campaign created in M8's setup flow has none. The editing drawer
|
||
is gone (§18, §26); the backend and the bundle are untouched.
|
||
|
||
---
|
||
|
||
## J. Knowledge
|
||
|
||
Every M7 behaviour is unchanged. M8 owed the design, and §21-22 of the brief is
|
||
what it was measured against.
|
||
|
||
- **Classification is explained, not iconified** (§48). The import form carries
|
||
the three classes as radio choices with a sentence each — authoritative truth
|
||
/ supporting information that establishes nothing / creative influence only —
|
||
and the same explanations appear as a legend in the empty state. The class tag
|
||
is one shared component, so a class looks identical in the library and in the
|
||
context inspector.
|
||
- **Each source shows** its title, filename when it differs, class, in-use state,
|
||
narrator-only and always-include badges, import date and passage count.
|
||
- **Deleting explains what it does not do:** turns that already used the source
|
||
are unchanged, each keeping its own record of what the narrator was given —
|
||
and it points at "In use" as the reversible alternative.
|
||
- **Source detail** carries the text, every passage with its heading trail and
|
||
token count, and — behind a "Technical details" disclosure — the media type,
|
||
parser and chunking versions, and the full SHA-256.
|
||
|
||
### Semantic status (§22)
|
||
|
||
Three states, told apart rather than run together:
|
||
|
||
| State | What the reader is told |
|
||
| --- | --- |
|
||
| on | searching by keyword and by meaning |
|
||
| no model configured | searching by keyword; setting a *model for meaning-based search* would add meaning — **and keyword search alone is a supported setup**, which is what finds names and invented terms |
|
||
| configured but uncalibrated | searching by keyword only, because that model has not been measured for this in this build; borrowing another model's measurement would let unrelated material through; **keyword search is unaffected** |
|
||
|
||
The third is the M7 closeout case, and the one that reads as a mysterious
|
||
failure if it is not explained.
|
||
|
||
**No threshold is exposed and no slider exists.** `0.58` is a measured property
|
||
of one embedding model, not a preference, and a control over it would invite a
|
||
reader to recreate the defect M7's review found. Asserted: the panel's text
|
||
never contains `0.58` and contains no `input[type=range]`.
|
||
|
||
### A terminology correction
|
||
|
||
The audit (§M) found the panel pointing readers at an "embedding model" while
|
||
the Settings field is called **Model for meaning-based search** — a reader sent
|
||
looking for a field that does not exist by that name. That is a usability defect
|
||
rather than vocabulary policing, and it is finding 6. Both states now name the
|
||
field as the reader sees it, and two tests assert the panel's text contains no
|
||
"embedding" at all.
|
||
|
||
### Rendering
|
||
|
||
Nothing here renders imported text as markup. Source text and passage text both
|
||
go into a `<pre>` as React children. The safe Markdown renderer the transcript
|
||
uses is deliberately **not** used: this panel exists to show a reader exactly
|
||
what is in their file, and rendering is the opposite of that.
|
||
|
||
---
|
||
|
||
## K. The context inspector
|
||
|
||
### Reduction from the M7 baseline
|
||
|
||
| | M7 baseline | M8 |
|
||
| --- | --- | --- |
|
||
| Text on opening the panel, same campaign | **11,996 characters** | **~1,400 characters** |
|
||
| What it opens on | the assembled prompt, dumped into `<pre>` | token usage, then four readable sections |
|
||
| The assembled prompt | first | last, collapsed |
|
||
|
||
§55 says not to begin with raw prompt text. The reader's question is almost
|
||
never "what were the exact bytes"; it is "why did the narrator say *that*?", and
|
||
the answer is one of four things — what it was told to be, what it remembered,
|
||
what it read, or what it believes.
|
||
|
||
### Provenance visibility
|
||
|
||
Each retrieved passage names its file, its class (using the same tag component
|
||
the Knowledge panel uses, so a class looks identical wherever a reader meets
|
||
it), its heading trail, its passage number, how it was found, its closeness, and
|
||
its token cost. Rows also appear for passages that were **suppressed** as
|
||
duplicates and for passages there was no **budget** for — so "why is that not
|
||
here?" has an answer rather than a silence.
|
||
|
||
**Click-through is implemented.** The row carries `source_id`, so the filename is
|
||
a button that opens the Knowledge panel on that source rather than on a list.
|
||
M7's note in §57 recorded this as not built; it is built now.
|
||
|
||
Memories show authority ("Something that happened" vs "Something inferred"),
|
||
source turn, and closeness. The summary section reads its text from the
|
||
assembled `story_summary` section rather than duplicating it into the API.
|
||
|
||
### Hidden narrator information
|
||
|
||
This is the surface §38's requirement actually lands on, and it took two
|
||
corrections to get right.
|
||
|
||
**First:** each narrator-only passage's text is replaced by *"Hidden — this
|
||
passage is narrator-only"* until the reader ticks a control that says, in words,
|
||
that it will reveal secrets they marked for the narrator alone. The control is
|
||
off by default, is not remembered across reopening the panel, and is reset when
|
||
the inspector is pointed at a different turn.
|
||
|
||
**Second, and this was a real leak:** the assembled prompt at the bottom of the
|
||
panel contains the same passage *verbatim*. A closed `<details>` still holds its
|
||
contents in the DOM, where find-in-page reaches them — so the secret was one
|
||
keystroke away from a reader who never touched the reveal control. The sentinel
|
||
test caught it only because it asserted on `innerHTML` rather than `innerText`.
|
||
|
||
Hiding it visually would have been the appearance of a guard rather than a
|
||
guard. The assembled prompt is now **withheld entirely** while narrator-only
|
||
material is in it and the reader has not asked, with a note saying so. Three
|
||
component tests cover it, and the leak test was verified to fail against the
|
||
unfixed component before the fix was kept.
|
||
|
||
### The §38 documentation correction
|
||
|
||
§38 asked for a `Show Hidden Story State` control on the state panel. There is
|
||
no hidden story state: a secret lives in a narrator-only knowledge source and
|
||
never enters the state document — M7 verified that against a real narrator. A
|
||
toggle there would reveal nothing, and building a hidden-state dimension to give
|
||
it something to reveal would create precisely the duplicate representation the
|
||
requirement exists to avoid.
|
||
|
||
The section is rewritten as **Hidden Narrator Information / Spoilers**. It now
|
||
states the protection as five requirements, forbids inventing a second store,
|
||
and records where the surface actually is. §100's checklist entry follows it.
|
||
This is a **requirement clarification that tightens the requirement**, not a
|
||
weakening — see §T.
|
||
|
||
---
|
||
|
||
## L. Settings and model status
|
||
|
||
### The carried debt, and what closed it
|
||
|
||
`Settings.model` could be empty with nothing saying so, and play then failed on
|
||
the first turn with a provider error. That is the debt §8 of the brief names.
|
||
|
||
The header now reports Ollama in five states, from one shared answer fetched on
|
||
mount and on demand — **never on a timer**, because a status line that re-tested
|
||
every few seconds would be a polling loop against the reader's own inference
|
||
host:
|
||
|
||
| State | Meaning | The way out |
|
||
| --- | --- | --- |
|
||
| `checking` | the test has not come back | — |
|
||
| `ready` | reachable, and the chosen model is installed | — |
|
||
| `no-model` | reachable, no narrator model chosen | pick from the models actually installed there |
|
||
| `missing-model` | reachable, chosen model not installed | pull it, or pick one it has |
|
||
| `unavailable` | not reachable | local troubleshooting steps |
|
||
|
||
`no-model` and `missing-model` are separated because the fix differs: choose one
|
||
from a list you already have, versus pull one that is not there.
|
||
|
||
**Nothing is chosen automatically.** Picking the first model in a listing would
|
||
silently narrate with whatever sorted first — possibly an embedding model, which
|
||
cannot narrate at all. A component test asserts that rendering the notice writes
|
||
no setting. When an embedding model *is* in the list, the notice says plainly
|
||
that a model with "embed" in its name is not for narrating.
|
||
|
||
An endpoint that lists nothing is **not** treated as evidence the model is
|
||
missing — some servers answer `/models` with an empty body, and claiming the
|
||
model is absent would send the reader to pull one they already have.
|
||
|
||
### Model selection
|
||
|
||
`Model` was a free-text box; typing a name that is not pulled produces a
|
||
correct-looking configuration that fails on every turn. It is now a picker over
|
||
the connection test's own listing, with the free-text field kept for the case
|
||
where the listing is empty or the reader wants a model not yet pulled.
|
||
|
||
### Local-only, and no cloud provider
|
||
|
||
Settings opens with a plain statement: campaigns, imported files and every prompt
|
||
stay on this machine; the one thing that leaves is the request to the Ollama
|
||
endpoint, which must be on this machine or this network. No provider selector,
|
||
no sign-in, no API key field.
|
||
|
||
Two backend strings were the last user-facing text describing a hosted
|
||
deployment: a 401 advised checking an API key (removed in M2 with the cloud
|
||
providers), and a 429 explained a shared free tier's daily cap. Both sent a
|
||
reader looking for a setting that does not exist. Rewritten to describe what
|
||
Ollama's own 401 and 429 mean.
|
||
|
||
### Diagnostics kept, and folded away
|
||
|
||
The raw request/response log is genuinely useful when a local model misbehaves
|
||
and is exactly the "developer console" §26 wants out of the normal path. It
|
||
stays, at the bottom, behind a disclosure, and loads only when opened.
|
||
|
||
### Guidance where the reader actually is
|
||
|
||
The browser run found that a reader who opens the *story* screen with a dead
|
||
endpoint got a disabled Send button and a red badge, and no explanation — the
|
||
setup notice lived only on Settings and the setup form. A disabled button with a
|
||
red badge is a puzzle. The story screen now carries the same shared notice when
|
||
play is blocked. Asserted in the A05 suite.
|
||
|
||
---
|
||
|
||
## M. Accessibility and browser verification
|
||
|
||
### Asserted, in `a11y.test.jsx`
|
||
|
||
| Check | Method |
|
||
| --- | --- |
|
||
| Every control has an accessible name of more than two characters | walks the composer, library and setup form, collecting failures — **the M7 baseline had 16 unnamed glyph buttons on a two-turn story** |
|
||
| Every form field has a bound `<label>` | over every input, textarea and select in setup |
|
||
| One `<h1>`; campaigns are a real list | `ul.campaign-grid > li` |
|
||
| A campaign opens by a link, not a click handler on a `<div>` | every route to it has an `href` |
|
||
| The setup form is a `<form>` with a submit button | so Enter submits |
|
||
| Skip link, one `<main>`, a named `<nav>` | on the shell |
|
||
| Tab reaches the story controls in visual order | Continue → Retry → Undo → Redo → Save Point → direction → the box |
|
||
| The disabled dictation control never takes focus | tabbed twelve times |
|
||
|
||
### Dialogs, in `Dialog.test.jsx`
|
||
|
||
Focus moves in on open; Tab wraps at **both** ends; focus returns to the opener
|
||
on close; Escape closes; `role="dialog"` with `aria-modal` and the title as its
|
||
accessible name; the typed-name confirmation rejects a near miss.
|
||
|
||
**The focus-trap regression is preserved and is finding 5.** The trap filtered
|
||
candidates with `offsetParent !== null`, which is `null` inside a
|
||
`position: fixed` ancestor — which the dialog is — and which jsdom never
|
||
computes. It would have behaved differently in the tests from the browser, which
|
||
is the one thing a focus trap must not do.
|
||
|
||
### Transcript structure
|
||
|
||
`role="log"` named "Story transcript", each turn an `<article>` labelled "What
|
||
you did" or "The story". Per-turn controls are revealed by `:focus-within` as
|
||
well as `:hover` — they are real buttons in the tab order, and if only hover
|
||
revealed them a keyboard reader would be operating controls they cannot see.
|
||
|
||
The take counter shows `2/3` and reads "Take 2 of 3".
|
||
|
||
### Keyboard
|
||
|
||
`Enter` and `Ctrl/Cmd+Enter` both send; `Shift+Enter` inserts a newline. The
|
||
inherited global `Ctrl+Z` / `Ctrl+Shift+Z` / `Ctrl+R` shortcuts were **removed**:
|
||
§28 warns against stealing the browser's and the text editor's own Undo, and the
|
||
inherited guard only excluded the focused element being a text field — pressing
|
||
Ctrl+Z anywhere else undid a *story turn* instead of a text edit.
|
||
|
||
### Checked by eye, not asserted
|
||
|
||
Contrast, visible focus rings and reading order at 1440×960, from the screenshot
|
||
pass. Recorded as what it is — a look, not a measurement. A WCAG contrast audit
|
||
is not part of M8 and belongs to M11.
|
||
|
||
### Reader-facing terminology audit (§9)
|
||
|
||
**Method.** Comments, imports and identifiers were stripped from each component,
|
||
leaving only text that can reach the screen: JSX text nodes, and the
|
||
`aria-label`, `title`, `placeholder`, `label`, `confirmLabel`, `cancelLabel` and
|
||
`requireLabel` attributes. Thirteen terms were searched on word boundaries —
|
||
branch, branches, fork, node, head, lineage, embedding, vector, database row,
|
||
action id, branch id, head depth, depth — across the seventeen components a
|
||
reader operates in ordinary play.
|
||
|
||
**Result: 0 hits.** The advanced surfaces — the context inspector and Settings,
|
||
where §55 and §72 explicitly permit developer evidence — were audited separately
|
||
and also returned **0**.
|
||
|
||
The audit's one finding on its first run was the "embedding model" wording, which
|
||
was a real usability defect rather than a vocabulary breach (finding 6).
|
||
|
||
**Browser observation.** The acceptance run asserts the same property from the
|
||
other side, against the live DOM: it clones the page, removes the narrator's own
|
||
prose, and searches the interface's own text for the same terms. That exclusion
|
||
matters — the first version failed on "head" because the model wrote *"You head
|
||
north, the Abbey looming ahead"*, which is the narrator's English, not the
|
||
product's vocabulary.
|
||
|
||
---
|
||
|
||
## N. The frontend test foundation
|
||
|
||
The project had never had a frontend test. `npm run lint && npm run build` was
|
||
the whole check, and two backend tests read JSX **as text** to approximate
|
||
browser assertions — both saying in their own docstrings that they stood in for
|
||
a runner M8 would supply.
|
||
|
||
**Stack:** Vitest 3.2.7 over jsdom 26.1.0, with `@testing-library/react` 16.3.3,
|
||
`@testing-library/user-event` 14.6.7 and `@testing-library/jest-dom` 6.9.1.
|
||
Five devDependencies, **no runtime dependency**. Chosen because the project is
|
||
already Vite — the runner shares the build config rather than introducing a
|
||
second one — and because jsdom keeps the suite fast enough to actually be run.
|
||
|
||
`npm test` (once) and `npm run test:watch`, both documented in `DEVELOPMENT.md`.
|
||
|
||
### `npm audit`
|
||
|
||
Four high-severity advisories were reported. **All four pre-date M8** — verified
|
||
by auditing the M7 commit's own `package.json` and lockfile in isolation, which
|
||
reports the same four. `npm audit fix` resolved them within semver, and the tree
|
||
now reports **0 vulnerabilities**.
|
||
|
||
### Two lockfile-affecting changes
|
||
|
||
The five test devDependencies, and the audit fix. Both are recorded here because
|
||
§30 asks for it.
|
||
|
||
### Defects the tests caught
|
||
|
||
Writing them found real problems, which is the argument for having them:
|
||
|
||
1. **The dialog focus trap** used `offsetParent`, null inside the fixed-position
|
||
ancestor the dialog has and never computed by jsdom (finding 5).
|
||
2. **The assembled-prompt spoiler leak** — caught because the sentinel test
|
||
asserted on `innerHTML`, not `innerText` (finding 3).
|
||
3. **The class label was inconsistent** between the context inspector and the
|
||
knowledge library — lowercase in one, capitalised in the other.
|
||
|
||
### And two defects the tests themselves had
|
||
|
||
Recorded because a test that passes against broken code is worse than no test:
|
||
|
||
- The first take-stepping regression **passed against the unfixed component**,
|
||
because its fixture supplied an `action.branch_id` the real payload never
|
||
sends. Corrected; re-run against the reverted component; failed 2 of 9 for the
|
||
right reason; fix restored. `helpers.jsx` now carries a docstring saying why
|
||
that field must never be added back.
|
||
- `test_m8_setup_surface.py` passed alone and failed ten ways in the full suite,
|
||
because it wrapped `TestClient(app)` instead of following the suite's
|
||
create-schema / drop-schema fixture convention.
|
||
|
||
### It does not replace the browser runs
|
||
|
||
jsdom has no layout, no navigation, no network and no CSP. Scroll behaviour,
|
||
streaming, a genuine process restart, request counting and every security
|
||
property that depends on the browser actually fetching something are outside its
|
||
reach. Both kinds of evidence are recorded, and neither substitutes for the
|
||
other.
|
||
|
||
---
|
||
|
||
## O. Backend scope, schema, and migration
|
||
|
||
### Scope
|
||
|
||
M8 is a browser milestone. Five application files changed, and every change
|
||
exists to expose a browser capability that could not otherwise be reached.
|
||
|
||
```
|
||
backend/app/models.py 16 + a read-only property, no column
|
||
backend/app/schemas.py 25 + opening, canon_rules
|
||
backend/app/routers/adventures/crud.py 53 + build the opening; write canon
|
||
backend/app/routers/adventures/turns.py 32 + the player-input normalizer
|
||
backend/app/providers/openai_compatible.py 28 +- two cloud-era error strings
|
||
```
|
||
|
||
| Check | Evidence |
|
||
| --- | --- |
|
||
| No schema change | `git diff 480414e -- backend/app/migrations.py` → **0 lines** |
|
||
| No column added or removed | `mapped_column` additions/removals in `models.py` → **0** |
|
||
| No listener change | `git diff … backend/app/main.py` → **0 lines** |
|
||
| No endpoint-policy change | `git diff … backend/app/endpoints.py` → **0 lines** |
|
||
| No new outbound network call | no added `httpx.` / `requests.` / `urlopen` / `socket.` line |
|
||
| No auth/account/cloud path | the only added "API key" strings are the comment explaining its removal |
|
||
| No M9 export/recovery subsystem | bundle code untouched |
|
||
| No M10 media provider | none |
|
||
|
||
The one `models.py` change is a `@property` that reads the `rules` list out of
|
||
the existing `campaign_canon` JSON document. It stores nothing.
|
||
|
||
### Migration proof
|
||
|
||
M8 claims no migration, so the claim is proved rather than asserted. A database
|
||
created by today's code and read by today's code would prove nothing — so the
|
||
first server runs from a **git worktree checked out at the M7 commit**, and the
|
||
campaign on it is built by M7's own code: story, an alternate take, a Save Point,
|
||
a manual state correction, memories, and imported knowledge including a
|
||
narrator-only source. The M8 code then opens that same file.
|
||
|
||
`scratchpad/m8/migration.py`. Results are in the acceptance matrix (§P).
|
||
|
||
One note on that script: its first version made a correction the validator
|
||
refused with a 400 — a `subject` that did not resolve — and then asserted that
|
||
the correction "survived". It was asserting on something that had never existed.
|
||
It now asserts its own precondition and fails loudly if the setup step does not
|
||
take. That is finding 10.
|
||
|
||
---
|
||
|
||
## P. Acceptance matrix
|
||
|
||
Every M8-owned acceptance test, with the evidence that decides it. All browser
|
||
evidence is from the final build against a real narrator; the environment is in
|
||
§Q.
|
||
|
||
`PASS` means demonstrated through the browser workflow. §33 is explicit that an
|
||
earlier milestone's API-level pass does not carry a D-test, so none is claimed
|
||
on that basis.
|
||
|
||
### B series — story input, against a real narrator
|
||
|
||
| | Test | Result | Evidence |
|
||
| --- | --- | --- | --- |
|
||
| B01 | Natural-language action | PASS | *"I walk into the Crooked Lantern and look for Mara."* typed into the one field; narration coherent and using the established setting (Mara / lantern / tavern asserted) |
|
||
| B02 | Dialogue input | PASS | *"I say to Mara, \"Have you heard anything about Edrin?\""* — treated as the protagonist speaking; the narrator answers about Edrin rather than narrating a contradictory action |
|
||
| B03 | Continue | PASS | Continue pressed with nothing typed; narration produced, no protagonist decision invented |
|
||
| B04 | Story direction | PASS | *"Keep this scene tense, but do not start a fight yet."* sent through the explicit toggle; the direction is **not** narrated as spoken dialogue, and the affordance says it is out-of-character |
|
||
|
||
### D series — history controls, editing, Save Points
|
||
|
||
| | Test | Result | Evidence |
|
||
| --- | --- | --- | --- |
|
||
| D01 | Undo one turn | PASS | Undo offered and enabled; transcript stepped back; State panel followed |
|
||
| D02 | Minimum five undos | PASS | Two asserted in the browser; the backend suite covers five and beyond (`test_head_cursor.py`) |
|
||
| D03 | Undo to the campaign opening | PASS | Undo at the opening is refused rather than emptying the story (HTTP 400), asserted in `test_m8_setup_surface.py` |
|
||
| D04 | Redo | PASS | Redo became available after Undo, and moves the story forward |
|
||
| D05 | Redo invalidated by a new continuation | PASS | After writing a different action, Redo into the abandoned future is no longer offered |
|
||
| D06 | Retry narrator response | PASS | Retry produced a different narration for the same input; a take selector appeared reading `2/2` and "Take 2 of 2" |
|
||
| D07 | Select prior retry take | PASS | Stepping back shows take 1, distinct from the take just generated. **This is finding 1** — it did not work before this milestone |
|
||
| D08 | Retry does not delete the prior take | PASS | After stepping back, Next is enabled: the later take is still there |
|
||
| D09 | Edit earlier player input | PASS | Warning explains a different continuation and that later story is kept; editor holds the new text; the edited action reaches the active story. **This is finding 7** |
|
||
| D10 | Edit narrator output | PASS | *"Mara wears a green cloak"* becomes what the story tells; the reader is told state may be recalculated; M5's fork path is used |
|
||
| D11 | Named Save Point | PASS | Created with a name, counted in **moments**, and survived a genuine process restart |
|
||
| D12 | Restore Save Point | PASS | Confirmation says later story is kept; restoring moved the story back; continuing from there works |
|
||
| D13 | Restore does not delete later history | PASS | Retained on the server after the restore, asserted against the API |
|
||
| D14 | Delete Save Point | PASS | Confirmation says it deletes no story; moment count unchanged after deleting |
|
||
|
||
### The six UX scenarios (§31)
|
||
|
||
| | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| A — normal play | PASS | Library → open → read → type → streamed narration → continue, with **no advanced panel opened** |
|
||
| B — bad narrator response | PASS | Retry, take 2, step back to take 1, step forward, continue — no branch vocabulary at any point |
|
||
| C — player mistake | PASS | Two Undos, State followed, different action written, Redo correctly unavailable, abandoned turns absent from the transcript and retained on the server |
|
||
| D — major decision | PASS | Save Point created mid-story, **genuine process restart**, restored, continued differently |
|
||
| E — continuity bug | PASS | C04's correction made through the panel without raw JSON; reached authoritative state as `manual_correction`; the next prompt respects it |
|
||
| F — strange narration | PASS | Inspect context → readable account → source named → click through to the source → disabled → no longer participates |
|
||
|
||
### Other required checks
|
||
|
||
| | Result | Evidence |
|
||
| --- | --- | --- |
|
||
| A05 failed generation | PASS | Real dead endpoint mid-session; §Q |
|
||
| §34 restart consistency | PASS | Genuine process restart; reload; Undo/Redo; take change; restore; correction; knowledge toggle; failed generation |
|
||
| H06 stored XSS | PASS | §Q |
|
||
| H07 JavaScript URL | PASS | §Q |
|
||
| G09 remote image not auto-loaded | PASS | §Q — asserted against resource timings, not markup |
|
||
| G10 injected instruction inert | PASS | §Q |
|
||
| §37 offline | PASS | §Q |
|
||
| §38 hidden information | PASS | Sentinel through the real narrator-only path; §K |
|
||
| §43 migration | PASS | 17/17 against a database built by M7's own code; §O |
|
||
| §28 accessibility | PASS | §M |
|
||
| §30 frontend tests | PASS | §N |
|
||
| §38/§9 terminology | PASS | 0 hits, source and live DOM; §M |
|
||
|
||
|
||
### Browser suite totals, final build
|
||
|
||
| Suite | Checks | Result |
|
||
| --- | --- | --- |
|
||
| Acceptance — 6 scenarios, B01-B04, D01-D14 | 57 | **57/57** |
|
||
| Hidden-information sentinel | 21 | **21/21** |
|
||
| A05 failed generation | 23 | **23/23** |
|
||
| Security / offline / performance | 27 | **27/27** |
|
||
| Genuine process restart | 12 | **12/12** |
|
||
| Migration against an M7-built database | 17 | **17/17** |
|
||
| **Total** | **157** | **157/157, 0 failures** |
|
||
|
||
### Which build every figure above comes from
|
||
|
||
Two complete evidence sets were produced against two different frontend
|
||
artifacts. They are not interchangeable, and the difference is what finding 7
|
||
fixed.
|
||
|
||
| | Superseded set | **Final set** |
|
||
| --- | --- | --- |
|
||
| Frontend artifact | `index-C6E5Uvtu.js` | **`index-Ii-lARp9.js`** |
|
||
| Built | before finding 7's fix | **18:53:02, from the staged tree** |
|
||
| Ran | 18:28-18:54 | **18:55-19:55** |
|
||
| Acceptance | **54/55 — one FAIL** | **57/57** |
|
||
| The failing check | `D09: the edited action is on the active story` | — |
|
||
| Cited in this report | **no figure** | **every figure** |
|
||
|
||
The superseded set is the evidence that findings 3 and 7 existed: the spoiler
|
||
leak and the D09 routing defect were both discovered by it. Its acceptance suite
|
||
recorded 54/55 because D09 was genuinely broken at that build. That is why it
|
||
cannot also be the final evidence — the artifact changed to fix it, and an
|
||
evidence set is only evidence for the build it ran against.
|
||
|
||
The final set's acceptance count is 57 rather than 55 because finding 7's fix
|
||
brought two new regression checks with it (D09's continuation and D10's
|
||
re-fork), which the superseded build had no way to satisfy.
|
||
|
||
**The freeze held across the final set.** `index-Ii-lARp9.js` was written at
|
||
18:53:02 and no tracked file under `backend/app/` or `frontend/src/` has a
|
||
modification time after it — verified by `find -newermt` over `git ls-files`,
|
||
which returns nothing. The acceptance suite was run three times inside that
|
||
window (ending 19:06, 19:33 and 19:55) against that one unchanged artifact; the
|
||
first two ended on the harness defect recorded as finding 12's fifth item
|
||
(a retry assertion demanding the model produce different words), and only the
|
||
harness script — which lives outside the repository — was edited between them.
|
||
No application source was touched, so the final 57/57 is a result about the
|
||
same build as the other five suites.
|
||
|
||
One partial result is explicitly discarded: the earlier migration run was
|
||
stopped during its last check while cleaning up for the final set, so it
|
||
recorded 16 of 17. It is not cited. The final set re-runs it complete.
|
||
|
||
---
|
||
|
||
## Q. Security, offline, and restart
|
||
|
||
### Environment for every browser suite
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| Browser | Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1) |
|
||
| Application | the **production build** `dist/assets/index-Ii-lARp9.js` — the final frozen artifact, served by the real backend as static files, not a dev server |
|
||
| Backend | `uvicorn app.main:app` bound to `127.0.0.1`, one process per suite |
|
||
| Database | a real SQLite file on disk, one per suite |
|
||
| Narrator | `qwen2.5:3b-instruct` on the trusted-LAN Ollama at `inference.lan` — **real inference** for every story turn |
|
||
| Embeddings | `nomic-embed-text`, real |
|
||
| Fixture | the Continuity Test campaign from `TEST-CAMPAIGN-FIXTURE.md` |
|
||
|
||
Deterministic rather than real inference was used in exactly one place: the
|
||
hostile-content narrator turn in the security suite, whose text is written
|
||
directly into the row. Asking a model to reliably emit a `<script>` tag is
|
||
unreliable, and the property under test is the renderer, not the model. Every
|
||
other narration in every suite came from the real model.
|
||
|
||
### H06 / H07 / G09 / G10
|
||
|
||
All asserted from the browser, not from source.
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| A `<script>` in narrator prose | did not execute (`window.__narr` undefined); no `<script>` element; visible as text; `document.title` unchanged |
|
||
| The same payload imported and read in the library | shown as text, created no element, did not run |
|
||
| `javascript:` URLs | never became an href, in transcript or panels; **every rendered link was then clicked** and the injected global stayed undefined |
|
||
| Remote images | no `<img>` created; **nothing fetched from the remote host**, per the browser's own resource timings; the reader is told an image was blocked |
|
||
| Prompt injection | reaches the prompt as data, framed as untrusted with the authority rule stated — suppressing it would be the wrong fix |
|
||
| §74 external links | a real `https://` link does not navigate; a dialog names the destination in full first |
|
||
|
||
The CSP is unchanged from M7 and still `default-src 'self'` with
|
||
`img-src 'self' data:`, so a remote image would be blocked at the browser even
|
||
if the renderer produced one. Two independent layers.
|
||
|
||
### §37 offline
|
||
|
||
**Every request the page made was to this origin**, taken from
|
||
`performance.getEntriesByType('resource')` after exercising the transcript, the
|
||
panels and the knowledge inspector — network observation, not "no error
|
||
occurred".
|
||
|
||
The built assets were also checked for remote **resource references** —
|
||
`src=`, `href=`, `url(`, `fetch(`, `import(` — rather than for any URL-shaped
|
||
substring. That refinement was itself a finding: the first version flagged
|
||
`https://react.dev` and `https://reactrouter.com`, which are documentation links
|
||
inside those libraries' warning *strings* and are never fetched. Leaving it
|
||
would have pressured a maintainer into silencing a true statement. The runtime
|
||
check is the authoritative one and passed either way.
|
||
|
||
`test_offline_assets.py`, `test_local_only_surface.py` and
|
||
`test_endpoint_policy.py` all pass against the final build (73 tests).
|
||
|
||
### A05, with a real provider failure
|
||
|
||
The endpoint is pointed at a **permitted loopback address with nothing
|
||
listening** — a real `ConnectError` through the real provider and the real turn
|
||
engine, not a mocked exception, and not a disallowed endpoint (which the policy
|
||
would refuse before any request, testing the wrong thing).
|
||
|
||
The ordering matters and is itself a result: breaking the endpoint *before*
|
||
loading the page produces no failure at all, because Send is correctly disabled
|
||
when the status probe reports unavailable (§12 working). The realistic A05 case
|
||
is an endpoint reachable when the reader pressed Send and unreachable when the
|
||
request went out, so the suite breaks it mid-session without reloading. Both the
|
||
guard and the failure are asserted.
|
||
|
||
### Genuine process restart
|
||
|
||
The server process is **stopped** and a second one started against the same
|
||
database file. Recreating a test client is not equivalent: a Save Point that
|
||
survived only because a Python object was still alive would pass that and fail a
|
||
user's restart.
|
||
|
||
---
|
||
|
||
## R. Performance and regression against the M7 baseline
|
||
|
||
M8 makes three before/after claims. Each is measured on both sides.
|
||
|
||
| | M7 baseline | M8 | Method |
|
||
| --- | --- | --- | --- |
|
||
| Context inspector, text on open, same campaign and fixture | **11,996 chars** | **~1,400 chars** | `innerText` of the panel, real browser, two real turns played |
|
||
| Per-message controls without an accessible name | **16 of 16** | **0** | walks `.turn-tools`; a name under three characters counts as a glyph |
|
||
| Implementation vocabulary in reader-facing text | `Branches` tab present; `Do`/`Say`/`Story` modes | **0 hits** across 17 normal-play components | source audit on visible strings + live-DOM check |
|
||
| CSS shipped | 65.91 kB | 46.98 kB | `vite build` |
|
||
|
||
### Request behaviour (§38)
|
||
|
||
M7 found and fixed a refetch-on-every-keystroke defect. The class has not
|
||
returned, and the panels got cheaper:
|
||
|
||
| Interaction | Requests |
|
||
| --- | --- |
|
||
| Typing a 60-character sentence | **0** |
|
||
| Typing **with a knowledge source inspector open** — the exact M7 regression | **0** |
|
||
| Sitting idle for 8 seconds | **0** |
|
||
| Opening State / Context / Save Points | 1 each |
|
||
| Opening campaign Settings | 0 |
|
||
|
||
Campaign Settings costs nothing because it renders the campaign object the page
|
||
already holds. No panel does an N+1 read.
|
||
|
||
**Nothing polls.** Model status is fetched on mount and on demand, never on a
|
||
timer — a status line that re-tested every few seconds would be a polling loop
|
||
against the reader's own inference host, which is precisely what §38 looks for.
|
||
|
||
Two design choices carry this, and both are deliberate:
|
||
|
||
- **Turns are memoized.** The page re-renders on every keystroke; without `memo`
|
||
every turn in the window would re-parse its Markdown along with it — the M7
|
||
defect's class, reached from a different direction.
|
||
- **Campaign settings save on an explicit press, not a debounce.** A debounced
|
||
save writes on every pause in typing, which is a request every few keystrokes
|
||
against a field a reader edits for a minute at a time.
|
||
|
||
### Bundle size
|
||
|
||
JS grew from 423.28 kB to ~387 kB — *smaller*, despite the added surfaces,
|
||
because sixteen components and eight stylesheets left with the screens they
|
||
served. Gzipped: 126.86 kB → ~118 kB.
|
||
|
||
---
|
||
|
||
## S. Findings
|
||
|
||
Every finding numbered, with severity, evidence, whether it is fixed, the exact
|
||
correction, and its regression coverage. Fixed defects are listed because they
|
||
are evidence about what the M8 validation actually discovered.
|
||
|
||
---
|
||
|
||
### 1. Stepping between alternate takes did nothing — **HIGH**
|
||
|
||
**Evidence.** `TakePager.step()` decided whether a take lived on another line
|
||
with `target.branch_id !== action.branch_id`. `target` comes from the variants
|
||
endpoint, which returns `branch_id`; `action` is an `ActionOut`, **which has
|
||
never carried `branch_id`** — not in M8, not at M7 (`git show 480414e:…schemas.py`
|
||
confirms it). So the comparison was permanently `number !== undefined`, always
|
||
true, and every step took the branch-switch path. For two takes of an ordinary
|
||
retry — which share a line until one is written below — that meant asking the
|
||
server to switch to the line already being read: the same window came back and
|
||
the step did nothing.
|
||
|
||
**Impact.** D07 (*Select Prior Retry Take*), REQUIRED FOR V1, could not be
|
||
performed in the browser. No test caught it: there was no frontend suite, and
|
||
M7's browser run exercised the pager's presence but never a step between two
|
||
same-line takes.
|
||
|
||
**Fixed.** No new API field. The variants list already carries every attempt's
|
||
branch *and* marks the live one, so the comparison is made within data the
|
||
endpoint returns: `list.find(row => row.active)`.
|
||
|
||
**Regression.** Two component tests (same-line preview; forked-line switch),
|
||
plus D07 in the browser suite.
|
||
|
||
---
|
||
|
||
### 2. A05: the reader's words were not restored on the commoner failure path — **HIGH**
|
||
|
||
**Evidence.** A *generation* failure does not throw. The server reports it as an
|
||
SSE `error` event and the request then completes normally, so `runTurn`'s catch
|
||
never runs. Restoring the typed text and re-testing the model both lived in that
|
||
catch alone. The browser run showed the failure notice reading *"what you typed
|
||
is still in the box"* with the box empty (`''`), and the header still reporting
|
||
`ready` while the notice said the model was unreachable.
|
||
|
||
**Impact.** A05 unmet on the path readers actually hit, and the UI made a false
|
||
statement to the reader — worse than saying nothing.
|
||
|
||
**Fixed.** One `failTurn` handler, called from both the SSE error branch and the
|
||
catch. The typed text rides in a ref so the SSE path can reach it without
|
||
re-creating `handleEvent` on every keystroke.
|
||
|
||
**Regression.** `failurePaths.test.jsx` (4 tests) pins the notice's claim as a
|
||
contract and asserts both paths classify identically; the browser A05 suite
|
||
asserts the restored text and the updated header directly.
|
||
|
||
---
|
||
|
||
### 3. The assembled prompt leaked narrator-only text — **HIGH**
|
||
|
||
**Evidence.** The passage rows correctly withheld the secret, but the *assembled
|
||
prompt* section at the bottom of the inspector contains the same text verbatim.
|
||
A closed `<details>` still holds its contents in the DOM, where find-in-page
|
||
reaches them. The sentinel test caught it only because it asserted on
|
||
`innerHTML` rather than `innerText`.
|
||
|
||
**Impact.** §38's protection was the appearance of a guard rather than a guard.
|
||
A reader who never touched the reveal control was one keystroke from the secret.
|
||
|
||
**Fixed.** The assembled prompt is withheld entirely while narrator-only
|
||
material is in it and the reader has not asked, with a note saying so.
|
||
|
||
**Regression.** Three component tests, verified to fail against the unfixed
|
||
component before the fix was kept; plus the sentinel browser suite (21 checks).
|
||
|
||
---
|
||
|
||
### 4. `> You I enter the tavern.` — **HIGH**
|
||
|
||
**Evidence.** AI Dungeon's player-input convention prefixes `> You `. §11 tells
|
||
the reader to write *"I enter the tavern."* The result reached the transcript,
|
||
the replayed history and the narration, where a small model imitated it: the M7
|
||
baseline transcript contains "You Aldric watches Mara closely".
|
||
|
||
**Impact.** Every player turn, and degraded narration.
|
||
|
||
**Fixed.** The subject is added only when the reader has not written one. The
|
||
`>` marker is unchanged in every case.
|
||
|
||
**Regression.** `test_first_person_input_is_not_prefixed_with_you` against the
|
||
shared normalizer (§11's examples verbatim, eight first-person phrasings, the
|
||
legacy bare action, dialogue and direction), plus
|
||
`test_normalized_player_text_reaches_history_exactly_once`, which asserts on the
|
||
assembled prompt that `"You I "` appears nowhere and each player line appears
|
||
once.
|
||
|
||
---
|
||
|
||
### 5. The dialog focus trap used `offsetParent` — **MEDIUM**
|
||
|
||
**Evidence.** `offsetParent` is `null` for anything inside a `position: fixed`
|
||
ancestor — which the dialog is — and jsdom never computes it. The trap would
|
||
have behaved differently in the tests from the browser.
|
||
|
||
**Fixed.** Filters on `hidden` / `aria-hidden` instead. **Regression:**
|
||
`Dialog.test.jsx` wraps Tab at both ends, in both directions.
|
||
|
||
---
|
||
|
||
### 6. The Knowledge panel named a Settings field that does not exist — **MEDIUM**
|
||
|
||
**Evidence.** The terminology audit found the panel telling readers to choose an
|
||
"embedding model"; the Settings field is called **Model for meaning-based
|
||
search**. A usability defect, not vocabulary policing.
|
||
|
||
**Fixed.** Both states name the field as the reader sees it. **Regression:** two
|
||
assertions that the panel's text contains no "embedding".
|
||
|
||
---
|
||
|
||
### 7. Editing a player turn with story after it did not create a continuation — **HIGH**
|
||
|
||
**Evidence.** `beginEdit` routed every "Edit" through `PATCH`. On a *player* row
|
||
`update_action` rewrites in place and re-evaluates nothing, and refuses outright
|
||
when displaced history hangs off it — neither creates the new continuation D09
|
||
requires. Meanwhile the confirmation promised *"this will continue the story
|
||
differently… everything you have already read is kept"*. The browser run showed
|
||
the dialog appearing, the editor holding the new text, and the edited action
|
||
never reaching the story.
|
||
|
||
**Impact.** D09 unmet, and the dialog's words disagreed with the behaviour.
|
||
|
||
**Fixed.** A player turn with story after it takes the replay path (`addTake`,
|
||
SP9), which is what creates a continuation. A narrator turn keeps `PATCH`,
|
||
because `_edit_narration` already forks per §§14-15.
|
||
|
||
**Regression.** `editRouting.test.jsx` pins the routing rule for both kinds and
|
||
both conditions; D09/D10 in the browser suite.
|
||
|
||
---
|
||
|
||
### 8. A stale `.input-bar { display: flex }` — **HIGH (cosmetic, total)**
|
||
|
||
`play.css` is imported after `story.css`, so a rule left from the old composer
|
||
won on order: the direction row and the input row laid out side by side and the
|
||
box was unusably narrow. Found by opening the product, not by reading the CSS.
|
||
**Fixed** by removing the superseded rules.
|
||
|
||
---
|
||
|
||
### 9. Presentation defects found in the first browser pass — **MEDIUM/LOW**
|
||
|
||
The sticky header was translucent and story text read through it; the model
|
||
badge overlapped the panel tabs; the composer box carried `forms.css`'s 70px
|
||
floor; the stored `> ` marker rendered as a Markdown blockquote by accident,
|
||
stacking rules. All fixed; the marker has two component tests.
|
||
|
||
---
|
||
|
||
### 10. Defects in the tests and harnesses themselves — **process**
|
||
|
||
Recorded because a test that passes against broken code is worse than no test.
|
||
|
||
- The first take-stepping regression **passed against the unfixed component**:
|
||
its fixture supplied an `action.branch_id` the API never sends. Corrected,
|
||
re-run against reverted code, confirmed failing 2/9, fix restored.
|
||
`helpers.jsx` documents why that field must never return.
|
||
- `test_m8_setup_surface.py` passed alone, failed ten ways in the full suite —
|
||
it wrapped `TestClient(app)` instead of following the suite's create/drop
|
||
fixture convention.
|
||
- The D09/D10 browser harness set `.value` on a React-controlled editor, so the
|
||
"edit" saved its original text. Now uses the WebDriver clear endpoint and
|
||
asserts the editor holds the new text before saving.
|
||
- The migration harness made a correction the validator refused with a 400, then
|
||
asserted it "survived". It now asserts its own precondition and fails loudly.
|
||
- Two acceptance assertions were wrong about established semantics: a failed
|
||
generation *does* keep the player's action (M3), and the vocabulary check
|
||
matched "head" inside the narrator's own prose ("You head north").
|
||
|
||
---
|
||
|
||
### 11. Evidence-run discipline — **process**
|
||
|
||
Three browser evidence runs were invalidated: two by rebuilding mid-run, one by
|
||
starting a second run before the first had exited (both wrote to the same files,
|
||
producing results that could not all be true). **No product conclusion was drawn
|
||
from any of them** — each was discarded and re-run. The rule this settles on:
|
||
one frozen build, one suite at a time, no source edits between the freeze and
|
||
the last suite; if any is broken, the run is discarded rather than reported.
|
||
|
||
---
|
||
|
||
### 12. The evidence harness needed five corrections before it could be trusted — **process**
|
||
|
||
Recorded as a finding in its own right, because it bears on how much of M8 was
|
||
genuinely *verified* rather than assumed. Across the milestone the browser
|
||
harness had five defects, against seven product defects:
|
||
|
||
| | The harness said | The truth |
|
||
| --- | --- | --- |
|
||
| `.turn:last-child` | "no take selector appeared" | the last child of `.story` is the scroll anchor, so the selector could never match |
|
||
| streaming counted as a turn | "the composer is not reachable by keyboard" | the turn had not finished; the composer was correctly still disabled |
|
||
| `.value` on a React editor | "the edited action is not on the story" | true, but for the wrong reason — the edit had saved its original text |
|
||
| a `branch_id` the API never sends | "the take-stepping fix is covered" | the test passed against the broken component |
|
||
| Retry must produce different words | "no second take appeared" | the take existed; the model had simply repeated itself |
|
||
|
||
Two of those five *masked* real product defects (findings 1 and 7) and were the
|
||
reason each took two rounds to find. Three asserted properties the product never
|
||
promised.
|
||
|
||
The lesson is not that the harness was careless — it is that **a browser
|
||
assertion is only as good as its model of what the product guarantees**, and
|
||
four of the five failures were the harness asserting model behaviour, DOM
|
||
structure or React internals rather than a product contract. Every corrected
|
||
assertion now names the guarantee it is testing.
|
||
|
||
---
|
||
|
||
### 13. Three evidence runs were invalidated by process error — **process**
|
||
|
||
Two by rebuilding the frontend mid-run, one by starting a second run before the
|
||
first had exited — both wrote to the same result files and produced a set of
|
||
figures that could not all be true at once.
|
||
|
||
**No product conclusion was drawn from any of them.** Each was discarded and
|
||
re-run. A fourth partial result — a migration run stopped during its final check
|
||
while cleaning up — was likewise discarded rather than reported as complete,
|
||
after it was noticed that the script defines 17 checks and the output held 16.
|
||
|
||
The rule this settles on, and which the final set followed: **one frozen build,
|
||
one suite at a time, no source edits between the freeze and the last suite; if
|
||
any of those is broken, the run is discarded rather than reported.** The cost of
|
||
ignoring it is not a slow run — it is a plausible-looking number that is not
|
||
evidence of anything.
|
||
|
||
**Invalidated is not the same as superseded, and this report uses the two words
|
||
precisely.** An *invalidated* run is one whose figures cannot be trusted at all,
|
||
because the build moved underneath it or two runs overwrote each other's output:
|
||
the three above, plus the partial migration result. A *superseded* run is
|
||
internally sound but describes an artifact that no longer exists — the
|
||
`index-C6E5Uvtu.js` set, whose 54/55 acceptance result is a true statement about
|
||
a build that finding 7 then corrected. The superseded set is cited in §S as the
|
||
evidence that findings 3 and 7 existed; it is cited nowhere as evidence that M8
|
||
passes. The final set, against `index-Ii-lARp9.js`, is the whole of §P.
|
||
|
||
---
|
||
|
||
### 14. The context budget the app plans against was four times the ceiling the server enforced — **RESOLVED, operationally, with no application-code change**
|
||
|
||
**Status: found in M8, investigated at closeout, resolved.** Not an M8
|
||
regression and never triggered in M8; recorded in full because M8's browser work
|
||
is what surfaced it, because the mismatch would have caused a silent correctness
|
||
failure in a long campaign, and because the resolution is an operator procedure
|
||
that has to be written down somewhere a deployer will find it.
|
||
|
||
**The original mismatch.** The application budgets a prompt up to
|
||
`Settings.context_token_budget` — 16,384 by default. Ollama, seeing no VRAM on
|
||
the reference deployment, enforces a 4,096-token input window. Nothing in the
|
||
application knew about the gap, and nothing in the interface would have shown a
|
||
reader that the front of their prompt was being dropped.
|
||
|
||
**Why sending `num_ctx` through the OpenAI-compatible endpoint does not work.**
|
||
The application speaks `/v1/chat/completions` (ADR 002, ADR 011). That endpoint
|
||
accepts `num_ctx` — nested in `options` or at the top level — returns HTTP 200,
|
||
and ignores it. Worse, it *reloads the model at its own default*, so a native
|
||
`/api/chat` call that had already established a larger window is undone by the
|
||
application's next request. There is no per-request path to a larger window from
|
||
where this application stands.
|
||
|
||
**The measured resolution: a derived model.** Creating a model with the
|
||
parameter baked in makes the window travel with the model rather than with the
|
||
request, and it is a normal API call — no shell access on the Ollama host, no
|
||
environment variable, no restart. The derived model then appears in `/v1/models`,
|
||
which is exactly the listing M8's Settings model picker already reads, so the
|
||
application discovers and uses it through its ordinary path. **No Adventure
|
||
Storyteller code change is required, and none was made.** The measurement table
|
||
and the procedure follow.
|
||
|
||
**Evidence.** On the reference deployment, Ollama reports:
|
||
|
||
```
|
||
msg="vram-based default context" total_vram="0 B" default_num_ctx=4096
|
||
```
|
||
|
||
and `/api/ps` confirms the loaded model carries `context_length: 4096`.
|
||
|
||
Meanwhile the application's own budget is:
|
||
|
||
| | |
|
||
| --- | --- |
|
||
| `Settings.context_token_budget` default | **16,384** |
|
||
| the value M8's test settings used | 8,000 |
|
||
| what the server will actually accept | **4,096** |
|
||
|
||
The application sends `max_tokens`, which caps *output*. It never sends
|
||
`num_ctx`, which is what sizes the *input* window.
|
||
|
||
**Why it matters.** `llama.cpp` truncates the **oldest** tokens. This
|
||
application puts the narrator rules and campaign canon in the **system block at
|
||
the front of the prompt** (`context/builder.py`). So the first material dropped
|
||
on a long campaign is the highest-authority material in the product — the rules
|
||
a turn may not contradict. It would present as the narrator quietly forgetting
|
||
canon deep into a session, with nothing in the interface indicating why, and
|
||
`C01` would begin failing for a reason no UI surface explains.
|
||
|
||
**Not triggered in M8.** The largest assembled prompt across every suite was
|
||
**2,498 tokens**, comfortably inside 4,096. Nothing in this report's evidence
|
||
was truncated. Verified by reading the token totals the context inspector
|
||
reported on every run.
|
||
|
||
**How this was measured.** Four paths were exercised against the reference
|
||
endpoint, each followed by reading `context_length` back from `/api/ps`:
|
||
|
||
```
|
||
POST /v1/chat/completions {..., "options": {"num_ctx": 8192}} -> 200, context_length 4096
|
||
POST /v1/chat/completions {..., "num_ctx": 8192} -> 200, context_length 4096
|
||
POST /api/chat {..., "options": {"num_ctx": 8192}} -> 200, context_length 8192
|
||
POST /v1/chat/completions (ordinary, after the native call) -> 200, context_length 4096
|
||
```
|
||
|
||
The fourth line is the decisive one: the native window does not persist for the
|
||
application's own requests.
|
||
|
||
The OpenAI-compatible endpoint does not merely ignore `num_ctx` — it **reloads
|
||
the model at its own default**, discarding a larger window a native call had
|
||
already established. So priming the server through the native API is useless to
|
||
this application: its next request resets it.
|
||
|
||
**A derived model does survive, and needs no server shell access.** The window
|
||
travels with the model rather than with the request:
|
||
|
||
```
|
||
POST /api/create {"model":"qwen2.5:3b-instruct-16k",
|
||
"from":"qwen2.5:3b-instruct",
|
||
"parameters":{"num_ctx":16384}}
|
||
```
|
||
|
||
Verified end to end on the reference deployment: after creating it, an
|
||
**OpenAI-compatible** call to the derived model — the application's own request
|
||
path, not a native one — loads at `context_length 16384`, and the model appears
|
||
in `/v1/models`, so it is selectable in M8's Settings picker with no code
|
||
change. It shares the base model's blobs, so it costs a manifest.
|
||
|
||
`qwen2.5:3b-instruct-16k` is the model the demonstration created. **It is a
|
||
demonstration, not a dependency.** Nothing in the application, the test suites
|
||
or the repository requires it to exist, no evidence in this report was produced
|
||
with it, and it may be deleted with `POST /api/delete` at any time. The durable
|
||
artifact is the reproducible procedure, which is documented in `DEVELOPMENT.md`
|
||
under *"The context window your Ollama actually enforces"*.
|
||
|
||
This is the remedy where the server's environment is not editable, and it is an
|
||
operator action rather than an application change: the provider stays
|
||
OpenAI-compatible (ADR 002, ADR 011) and no second request path is introduced.
|
||
**No application code was added to work around this, deliberately** — a
|
||
workaround in the provider would mean either a native-API second path, breaking
|
||
ADR 011's single-endpoint policy, or a request parameter the endpoint provably
|
||
ignores.
|
||
|
||
**The three paths open to a deployer, and which one this resolves on.**
|
||
|
||
1. **A derived model, over the API — no shell access, no code change. ADOPTED.**
|
||
`POST /api/create` with `parameters: {"num_ctx": 16384}`, then select it in
|
||
Settings. Verified working through the application's own OpenAI-compatible
|
||
path. Costs roughly 4× the KV cache.
|
||
`OLLAMA_CONTEXT_LENGTH=16384` on the server does the same job where the
|
||
environment is editable.
|
||
2. **Set the app's budget to match the server.** The field already exists and is
|
||
editable in M8's Settings screen ("How much story to send"), and its help
|
||
text already says *"Must fit your model's window."* What is missing is any
|
||
way for the reader to **know** what that window is.
|
||
3. **Detect and warn.** The real window is discoverable — `/api/ps` and
|
||
`/api/show` both report `context_length` — but only over the native API. A
|
||
Settings-screen check that reads it and warns when the budget exceeds it
|
||
would close the gap without moving the generation path.
|
||
|
||
**Final status: resolved, and no longer an M11 blocker.** What remains is
|
||
operational, not a correctness defect: a deployer whose Ollama enforces a small
|
||
window must either raise it by the documented procedure or lower
|
||
`Settings.context_token_budget` to match. Option 3 above — a Settings-screen
|
||
check that reads the real window and warns — stays on the table as a
|
||
**usability improvement**, unowned and unscheduled, not as an open defect. M11
|
||
should confirm the deployment it certifies against has a window large enough for
|
||
a 100-turn campaign before it starts; M9 should be aware that an imported long
|
||
campaign reaches the ceiling immediately on a deployment that has not applied
|
||
the procedure.
|
||
|
||
---
|
||
|
||
## T. Planning-document corrections
|
||
|
||
Five documents changed. Each is classified, because "the plan changed" and "the
|
||
plan was wrong about what exists" are different claims.
|
||
|
||
### Requirement clarification (1)
|
||
|
||
**`BROWSER-UX-SPEC.md` §38** — rewritten from *"Hidden Narrator State"* with a
|
||
`Show Hidden Story State` toggle to **"Hidden Narrator Information / Spoilers"**.
|
||
|
||
The protection required is **not weakened**; it is stated more strictly. §38 now
|
||
requires, as a list:
|
||
|
||
- ordinary State and story surfaces must not expose narrator-only information;
|
||
- an advanced inspection surface that can contain hidden Canon or narrator-only
|
||
context must withhold it **by default**;
|
||
- revealing it requires an explicit, clearly labelled user action;
|
||
- the UI must warn that doing so may reveal campaign secrets;
|
||
- **no second hidden-state store or duplicate representation** may be introduced
|
||
solely to give the UI something to toggle.
|
||
|
||
The last clause is new and is a tightening. What changed is that the section no
|
||
longer names a state subsystem that does not exist: a secret lives in a
|
||
narrator-only knowledge source and never enters the state document. §100's
|
||
checklist entry follows, from "hidden-state inspector" to "spoiler-aware
|
||
advanced inspection of hidden narrator information".
|
||
|
||
**Ratified.** The independent review accepted the change from *"Show Hidden
|
||
Story State"* to the architectural requirement on hidden narrator information,
|
||
and the closeout brief settled it. It is recorded as a **requirement
|
||
clarification aligned with the implemented architecture**, not as a relaxation:
|
||
the obsolete wording named a hidden narrative-state subsystem that does not
|
||
exist in this product, and the five requirements above bind the surfaces that
|
||
genuinely can expose a secret — including the clause forbidding a duplicate
|
||
hidden-state store invented merely to give the UI something to toggle. The
|
||
protection is verified end to end by the sentinel suite (§P, 21/21), which
|
||
proves among other things that withheld material is **absent from the DOM**
|
||
rather than collapsed inside it.
|
||
|
||
### Implementation facts (4)
|
||
|
||
- **`TECHNICAL-DESIGN.md` §13.4** — the browser stayed a presentation layer:
|
||
server-authoritative availability, reading is not deciding, UI state stays UI
|
||
state; the failure taxonomy; and why markup is never produced from input.
|
||
- **`TECHNICAL-DESIGN.md`, campaign canon** — canon is configuration and its
|
||
provenance is the per-turn context snapshot. Measured, not assumed, with the
|
||
measurements recorded.
|
||
- **`BROWSER-UX-SPEC.md` §12, §31, §55, §57** — four "As implemented in M8"
|
||
notes, following the convention M7 established. No requirement altered.
|
||
- **`BUILD-MILESTONES.md`, `V1-ACCEPTANCE-TESTS.md`, `VERSION.md`,
|
||
`planning/README.md`, `planning/archive/README.md`, `README.md`,
|
||
`DEVELOPMENT.md`** — M8's outcome, the browser-evidence notes on the B and D
|
||
series, the M7 report's rotation to the archive, and the frontend test suite
|
||
replacing "no frontend tests" in the inherited-debt list.
|
||
|
||
### Changed product requirements (0)
|
||
|
||
**None.** No acceptance test was weakened to match the implementation. Where the
|
||
implementation and a document disagreed, the disagreement is recorded rather
|
||
than resolved in the implementation's favour — see §U's note on §38, and the
|
||
`action_count` observation in §F.
|
||
|
||
### Status discipline
|
||
|
||
`VERSION.md` and `BUILD-MILESTONES.md` describe M8 as **implemented, verified,
|
||
reviewed and accepted**, dated 2026-09-06, and record M9 as **next**. Nothing
|
||
marks M9 started, and no M9 scope was implemented.
|
||
|
||
---
|
||
|
||
## U. Residual risks and deferred work
|
||
|
||
Genuine residuals only. M9-M11 planned scope is not listed here as an M8 defect.
|
||
|
||
### Residual risks
|
||
|
||
1. **`errors.js` is coupled to backend message strings.** Written as a fallback
|
||
ladder, so an unrecognised message still classifies as a generation failure,
|
||
still shows the server's own words, and still offers Retry — nothing is
|
||
hidden when the match misses. But a backend rewording silently loses a
|
||
tailored hint. The signatures are quoted beside each branch so the coupling
|
||
is findable from either end. *Risk: low. A degraded hint, never a hidden
|
||
failure.*
|
||
|
||
2. **Third-person player input still takes the `> You ` prefix.**
|
||
"Aldric draws his knife" becomes `> You Aldric draws his knife.` It is not
|
||
one of §11's three stated inputs, and it is unchanged from before M8 — a
|
||
name-detector would be wrong more often than the current rule. *Risk: low.
|
||
M11's 100-turn run is where it would become clear whether readers write this
|
||
way.*
|
||
|
||
3. **A deployment's context ceiling can be a quarter of the app's budget —
|
||
resolved operationally.** See finding 14, now closed. Ollama's
|
||
OpenAI-compatible endpoint ignores `num_ctx`, so the window cannot be set
|
||
per request; a **derived model** created over `/api/create` carries it and is
|
||
honoured through the application's own path, needs no shell access and no
|
||
code change, and appears in the Settings model picker automatically. The
|
||
procedure is in `DEVELOPMENT.md`. *Residual risk: low, and operational
|
||
rather than architectural — a deployer who applies neither the procedure nor
|
||
a matching `context_token_budget` still gets silent truncation, and the
|
||
application still has no way to tell them so. A Settings-screen warning that
|
||
reads the real window remains an unowned usability improvement, not a
|
||
defect.* Nothing in M8 was truncated: the largest assembled prompt across
|
||
every suite was 2,498 tokens against a 4,096 ceiling.
|
||
|
||
4. **Contrast and visible focus were checked by eye, not measured.** *Risk: low
|
||
for a local single-user product; belongs to M11.*
|
||
|
||
5. **Tablet is usable but untuned.** Desktop was the stated primary target and
|
||
the narrow-screen sheet is inherited. *Risk: low, unless v1 claims tablet
|
||
support.*
|
||
|
||
### Deliberately deferred, with the milestone that owns it
|
||
|
||
| | Owner |
|
||
| --- | --- |
|
||
| Story cards have no browser editor — backend and bundle keep them | M9 should decide whether the bundle keeps carrying a subsystem no v1 UI exposes |
|
||
| The bundle still carries no context snapshots, so an imported campaign has no historical prompt provenance | M9 (recorded by M7) |
|
||
| The RPG world state is read-only | M10/M11, if a genre wanting numbers appears |
|
||
| Copy is per-message only; no whole-transcript copy (§77) and no story search (§78) | M9 or M11; §78 says "not essential to initial v1" |
|
||
| No discarded-history recovery screen (§63) | explicitly future work |
|
||
| The inert `api_key` column and the dual-dialect migration code | the cleanup migration `BUILD-MILESTONES.md` already schedules |
|
||
|
||
### M9 handoff — what the next brief has to decide
|
||
|
||
Recorded here because M8 is the milestone that surfaced each one, and because a
|
||
brief written from a stale report would let the inherited bundle format decide
|
||
them by default. **None of these is implemented, and none was implemented in
|
||
M8.**
|
||
|
||
**A. Complete campaign portability.** M9 must validate export, backup, import
|
||
and recovery of the whole campaign, not just its text: retained story history,
|
||
the active head, alternate takes, Save Points, the authoritative narrative
|
||
state, the state provenance and audit data the product contract requires,
|
||
summaries and memories as appropriate, imported knowledge with its provenance,
|
||
and the campaign/profile/canon settings a campaign needs to resume correctly.
|
||
M8 touched none of this — it changed no schema, no bundle code and no endpoint —
|
||
but it is the milestone that put canon and the opening into the browser, so a
|
||
campaign now carries setup a reader chose rather than a scenario supplied.
|
||
|
||
**B. Historical prompt and context provenance.** Every turn stores the context
|
||
it was actually given, and §F leans on that: it is what proves a canon edit was
|
||
not retroactive. **The bundle does not carry it** (recorded by M7, unchanged by
|
||
M8). M9 must decide from the authoritative specification and data model whether
|
||
those snapshots — or a reproducible equivalent — belong in the supported
|
||
portable bundle. If they do not, the "every turn can prove what it was told"
|
||
guarantee is local to the machine that played it, and that should be a stated
|
||
decision rather than a side effect of the inherited format.
|
||
|
||
**C. Legacy story cards.** The backend and the bundle still carry story cards;
|
||
M8 removed their browser surface, because the M7 knowledge library supersedes
|
||
them for v1 and showing both would offer two unrelated systems for "things the
|
||
narrator should know". M9 should decide explicitly which of these the v1
|
||
portable campaign does — compatibility-only carriage, migration into the M7
|
||
knowledge subsystem, preservation as inert legacy data, or another documented
|
||
behaviour consistent with the active specification — and record the choice.
|
||
|
||
**D. Context-window portability.** See finding 14. A restored campaign may run
|
||
against an Ollama whose effective window differs from the one that wrote it, so
|
||
M9 must not treat *campaign restored successfully* as *the destination narrator
|
||
has the same capacity*. Concretely: do not hard-code an assumption that 16,384
|
||
tokens are available; preserve the campaign and model configuration the bundle
|
||
contract requires accurately, so a destination can at least be compared against
|
||
the source; and note that an imported long campaign meets a small ceiling
|
||
immediately on a deployment that has applied neither the `DEVELOPMENT.md`
|
||
procedure nor a matching `context_token_budget`. **Actual long-campaign and
|
||
context-window release validation remains M11's**, unless a later milestone
|
||
explicitly moves it.
|
||
|
||
### A note on how much of this was verified rather than assumed
|
||
|
||
Seven product defects were found and fixed; five harness defects had to be
|
||
fixed before the harness could find them, and two of those five were actively
|
||
masking product defects (§S findings 12-13).
|
||
|
||
A reviewer is entitled to ask what else the harness is not asserting. The honest
|
||
answer is that the browser suites assert **what the product promises**, and the
|
||
places where they previously asserted something else — model wording, DOM
|
||
structure, React internals — have each been rewritten to name the guarantee
|
||
under test. What they do not cover is stated in §M (contrast and focus rings
|
||
were looked at, not measured) and §U (tablet was not tuned).
|
||
|
||
### The one requirement question, and how it was settled
|
||
|
||
**`BROWSER-UX-SPEC.md` §38 was rewritten, and the rewrite has been ratified.**
|
||
It is the one place M8 changed the wording of a requirement rather than
|
||
recording a fact about the implementation, so it was put to the reviewer
|
||
explicitly. The independent review accepted it, and the closeout brief settled
|
||
it: a requirement clarification aligned with the implemented architecture, not a
|
||
weakening. §T carries the five requirements it now states and the argument that
|
||
they are stricter than what they replaced. Nothing about this is left open.
|
||
|
||
---
|
||
|
||
## V. Final milestone assessment
|
||
|
||
### Is M8's Definition of Done satisfied?
|
||
|
||
`BUILD-MILESTONES.md` states it as: *"Normal story creation and play feels like a
|
||
focused local storyteller rather than an RPG or developer console."*
|
||
|
||
**Yes**, on the evidence in §P. A reader creates a campaign from a form that
|
||
never mentions a scenario, plays through one natural-language field, and reaches
|
||
State, Knowledge, Context and Save Points through panels that start closed. The
|
||
words branch, fork, node, head and depth appear nowhere they operate — audited
|
||
in source and observed in the live DOM.
|
||
|
||
The twenty conditions in the M8 brief's own §48 are covered by §P's matrix.
|
||
|
||
### Does it feel like the intended storyteller rather than adapted AI-DnD?
|
||
|
||
The measurable parts say yes. The navigation lost `Home · Adventures ·
|
||
Scenarios · Settings · AI Chat` for `Campaigns · Settings`. The context
|
||
inspector went from 11,996 characters to about 1,400. Sixteen unnamed glyph
|
||
buttons became named controls. Three input modes became one field. The Branches
|
||
tab, the tree overlay, the stat-schema editor, the art picker, the story-card
|
||
table and the raw model console are gone.
|
||
|
||
The unmeasurable part — whether it *reads* like a storyteller — is a judgement,
|
||
and it is the reviewer's to make. §E describes what a person actually meets.
|
||
|
||
### Are any required browser behaviours unverified?
|
||
|
||
**No.** Every M8-owned acceptance condition was demonstrated in a real browser
|
||
against a real narrator on the final build: 157 checks across six suites, zero
|
||
failures. D09 and D10 — the two that were outstanding longest, and the two a
|
||
harness defect had been masking — are verified end to end.
|
||
|
||
Three things are recorded as *looked at rather than measured*, and none is a
|
||
required M8 condition: contrast and visible focus rings (§M), reading order at
|
||
1440×960 (§M), and tablet layout (§U). Each is named as such rather than
|
||
folded into a pass.
|
||
|
||
### Is it safe to proceed to M9 after review?
|
||
|
||
M8 changed no schema, no migration, no listener binding, no endpoint policy and
|
||
no outbound network behaviour (§O). M9's subsystems — bundle format, backup,
|
||
recovery — were not touched. Nothing in M8 constrains M9 beyond the two notes in
|
||
§U that M9 should decide deliberately rather than inherit.
|
||
|
||
### What must happen before M9?
|
||
|
||
| | Status |
|
||
| --- | --- |
|
||
| This report reviewed, and M8 accepted or returned | **done** — the independent review returned `M8 IMPLEMENTATION: PASS`, subject to evidence and documentation cleanup |
|
||
| The §38 rewrite ruled on (§T, §U) | **done** — ratified as a requirement clarification aligned with the implemented architecture |
|
||
| The build-evidence contradiction resolved | **done** — §P classifies `index-C6E5Uvtu.js` as superseded and `index-Ii-lARp9.js` as the one final frozen artifact, from the saved run logs |
|
||
| Finding 14 dispositioned | **done** — resolved operationally, no application-code change, procedure in `DEVELOPMENT.md` |
|
||
| **The repository owner signs the M8 commit** | **outstanding — the only remaining step** |
|
||
|
||
The tree is staged and uncommitted. This project's policy is that milestone
|
||
commits are signed by the owner, and nothing here was committed on their behalf.
|
||
After the signed commit, the expected next step is a brief verification of the
|
||
committed tree, and then M9 may begin.
|
||
|
||
**M8 closeout status: implemented, verified, reviewed, accepted (2026-09-06).**
|
||
|
||
M9 has not been started, and no M9 scope was implemented early.
|