Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.
V1.1 RELEASE VALIDATION: PASS
What was run, on this candidate:
- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
identical on all 15 census fields, schema parity at user_version 94, and both
bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
verified, public endpoint refused, a real turn, restart, persistence, and
Firefox rendering the reopened campaign.
Carried residuals, stated rather than summarised away:
- WP-B: deterministic independent-memory recovery PASS; reference-model
independent-memory recovery FAIL at memory creation — the owner-accepted
limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
backlog, reproduced and not fixed during validation.
Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.
Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.
Still the owner's to do: sign the release commit, update main, tag v1.1.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
1019 lines
68 KiB
Markdown
1019 lines
68 KiB
Markdown
# Planning Package Version
|
||
|
||
- **Package:** Adventure Storyteller Planning Package v4.6
|
||
- **Revision date:** 2026-09-16
|
||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C is signed as `59b5ebc`, and **WP-D (recovery honesty) and WP-E (control-boundary contrast) are signed as `87a4032`**. **Integrated v1.1 release validation has run on candidate `87a4032` and PASSED** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): 82 REQUIRED v1 tests hold (81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a `--no-cache` image whose SPA is file-for-file identical to the local build, offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window run passing M01-M04 with every turn `fits` and 0 protocol leaks, identity 0/0, recovery 16/16, a real v1.0.0 upgrade identical on all 15 fields with both bundle directions importing, and a release smoke of 15/15. WP-B's reference-model memory limitation remains an accepted, documented residual. **v1.0.0 is still the released version: no release commit, no `main` update and no `v1.1.0` tag exist** — those are the owner's events.
|
||
|
||
## v4.6 — v1.1 integrated release validation (2026-09-16)
|
||
|
||
Release validation of candidate `87a4032`, not a work package: no requirement,
|
||
acceptance test, schema, bundle format or product code changed. The evidence is
|
||
`reports/v1.1/V1.1-RELEASE-REPORT.md`, sections A-W.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `reports/v1.1/V1.1-RELEASE-REPORT.md` | **New.** The frozen candidate, the 82-row v1 acceptance matrix, every gate's result, the A1 headroom table, the A2 leak count, the WP-B verdict kept in both halves, the residual classification, and the final decision | release report |
|
||
| `reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL` **PENDING → APPROVED**, sourced and dated to the owner's release-validation brief; the signed commit predated the review | correction of record |
|
||
| `V1.1-PLAN.md`, `planning/README.md`, `VERSION.md` | Status: WP-D/WP-E signed `87a4032`; validation passed; the three owner events still outstanding | status |
|
||
| `README.md` | v1.0.0 **remains** released; v1.1 implemented and validated but untagged; schema figure corrected to 94 | product docs |
|
||
| `backend/tools/v11_upgrade_check.py`, `backend/tools/v11_release_smoke.py` | **New**, harness only: the real-v1.0.0 upgrade gate and the release-shaped smoke test | tooling |
|
||
| `backend/tools/m11_identity.py` | Reads `AIDND_TEST_EMBED_MODEL` and enables the memory bank, so the diagnostic can run with memory on as the gate requires | tooling |
|
||
|
||
**Requirement changes: zero. Product-code changes: zero.**
|
||
|
||
**Outcome:** `V1.1 RELEASE VALIDATION: PASS`. Carried residuals: WP-B's
|
||
reference-model memory limitation, the mid-reply instruction echo (still
|
||
reproducible on the stored fixture, absent from release evidence), and the
|
||
doubled full stop. K1 is classified v1.2 backlog. **No release commit, no `main`
|
||
update, no `v1.1.0` tag.**
|
||
|
||
## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15)
|
||
|
||
The last two planned v1.1 packages, implemented and reported separately. No
|
||
requirement, acceptance test, schema or bundle format changed. WP-D's evidence is
|
||
in `reports/v1.1/V1.1-WP-D-REPORT.md`, WP-E's in `V1.1-WP-E-REPORT.md`.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `reports/v1.1/V1.1-WP-D-REPORT.md` | **New.** Full `integrity_check` on the finished backup copy, proved against a fixture `quick_check` calls healthy; an export that says when this version could not import it back, carried in headers because the response body *is* the bundle | work-package report |
|
||
| `reports/v1.1/V1.1-WP-E-REPORT.md` | **New.** Control boundaries raised to clear WCAG 1.4.11 (3:1), the contrast audit turned from a report into a gate, rendered before/after boundary measurements, and the `.slice-7` finding the change itself created | work-package report |
|
||
| `V1.1-PLAN.md` | Status: WP-C signed `59b5ebc`; WP-D and WP-E complete and staged | status |
|
||
| `planning/README.md` | Current state | index |
|
||
| `DEVELOPMENT.md` | How large an export can get: the 20 MB import ceiling, ~13 kB per action, M9's ~279-turn figure and why it is not a turn limit, and the four export headers | developer docs |
|
||
|
||
**Requirement changes: zero.**
|
||
|
||
Two owner decisions are outstanding, both recorded rather than assumed: WP-E's
|
||
before/after screenshots await approval (`OWNER SCREENSHOT APPROVAL: PENDING`),
|
||
and WP-D records that no real-browser click was made on *Back up now* — the
|
||
endpoint that button calls was driven instead.
|
||
|
||
## v4.4 — WP-C browser release coverage (2026-09-15)
|
||
|
||
Harness work, with one narrow product fix it found. No requirement, acceptance
|
||
test, schema or bundle format changed. The evidence is in
|
||
`reports/v1.1/V1.1-WP-C-REPORT.md`.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `reports/v1.1/V1.1-WP-C-REPORT.md` | **New.** The six reader workflows driven in a real browser, the existing 38 checks, the download environment, harness defects J1-J7, and product defects K1 (open) and K2 (fixed) | work-package report |
|
||
| `V1.1-PLAN.md` | Status: WP-B.2 committed; WP-C complete and staged | status |
|
||
| `planning/README.md` | Current state | index |
|
||
| `DEVELOPMENT.md` | The browser harness command and flags; Firefox download preferences and the `$HOME` rule; what counts as a finished download; waiting on conditions, never sleeping | developer docs |
|
||
|
||
**Requirement changes: zero.**
|
||
|
||
## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15)
|
||
|
||
WP-B.1's diagnostic (`beb17ad`) placed three memory deficiencies; WP-B.2 corrects
|
||
exactly those, one at a time, each verified before the next. No requirement or
|
||
acceptance test changed, and no schema, bundle format or setting default
|
||
changed. Evidence and the WP-B decision are in
|
||
`reports/v1.1/V1.1-WP-B2-REPORT.md`.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1 WP-B.2)**: a block longer than 2,000 tokens is shown to the summariser as its opening and its end with an omission marker, inside the same budget; a block that fits is unchanged; the marker is never stored. | as-implemented record |
|
||
| `CONTEXT-AND-MEMORY.md` §18 | **As implemented**: the retrieval query is the player's input plus a bounded scene context, recorded per turn. | as-implemented record |
|
||
| `CONTEXT-AND-MEMORY.md` §20 | **As implemented**: `final = semantic + 0.15 × lexical`, the rarity-weighted lexical term over the input, the sweep that chose the weight, pins and ties. | as-implemented record |
|
||
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1)**: the memory prompt is unchanged from v1.0.0, and the accepted limitation: the reference summariser can omit or misattribute a fact from a block it was given whole. | as-implemented record |
|
||
| `DEVELOPMENT.md` | The GPU-host kernel/Ollama watch command corrected: `-k -u ollama` matched nothing; the OR form records both. | developer docs |
|
||
| `CONTEXT-AND-MEMORY.md` §21 | **As implemented**: selection unchanged in shape; eviction ordered by coverage first (boundaries kept, smallest hole first), recency second, v1.0.0 order as fallback. | as-implemented record |
|
||
| `V1.1-PLAN.md` | Status: WP-B accepted with a documented real-model limitation. §11 release criteria 12 and 13: the v1.1 release report must state the deterministic PASS and reference-model FAIL at memory creation, and list the carried residuals. | status, release gate |
|
||
| `planning/README.md` | Current state. | index |
|
||
| `reports/v1.1/V1.1-WP-B2-REPORT.md` | **New.** B2.1-B2.3 designs and evidence, full deterministic acceptance, real-model attempts, the rejected B2.4 prompt experiment, compatibility, offline, and the WP-B disposition. | work-package report |
|
||
|
||
**Requirement changes: zero.**
|
||
|
||
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
|
||
|
||
Two v1.1 work packages, implemented in sequence. No requirement or acceptance
|
||
test changed, and no schema or bundle format changed. Evidence is in
|
||
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `TECHNICAL-DESIGN.md` §15.2 | **As implemented (v1.1 WP-A1)**, covering: the safety reserve, `max(256, ceil(5%))`; what replaced M6's 64-token margin; the server's own count; the four accounting states; and keeping the turn. | as-implemented record |
|
||
| `TECHNICAL-DESIGN.md` §15.4 | **As implemented (v1.1 WP-A2)**: the vocabulary shown in the wire format, genre-neutral placeholders, and extractor rules R1-R4, each anchored to application-owned text. | as-implemented record |
|
||
| `DECISIONS/013-authoritative-narrative-state-document.md` | An implementation note for v1.1. No event type, field, validation rule or proposal record changed. | as-implemented note |
|
||
| `V1.1-PLAN.md` | Status: A1 and A2 implemented and staged. | status |
|
||
| `planning/README.md` | Current state. | index |
|
||
| `reports/v1.1/V1.1-WP-A1-A2-REPORT.md` | **New.** The combined review package, with A1 and A2 kept separate. | work-package report |
|
||
| `README.md`, `DEVELOPMENT.md` | The reserve and the accounting. The stale backend test count and the Screenshots paragraph are corrected. | developer docs |
|
||
|
||
**Corrective work before commit (owner review, 2026-09-14).** Report §R.
|
||
|
||
| Document | Change |
|
||
| --- | --- |
|
||
| `TECHNICAL-DESIGN.md` §15.2 | A cold model is loaded once before its turn is built (`contextwindow.ensure_window`). |
|
||
| `TECHNICAL-DESIGN.md` §15.4 | R5: the echoed continue hint is recognised by its own sentence, and the application-opened tail above it is removed. |
|
||
| `DEVELOPMENT.md`, `README.md` | The cold-model load, in operator terms. |
|
||
|
||
**Requirement changes: zero.**
|
||
|
||
## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14)
|
||
|
||
Documentation only. No product code, no requirement, no acceptance test and no
|
||
schema changed.
|
||
|
||
**Part 1 — post-release correction.** v4.0 was written before three owner
|
||
events that have since happened: the closeout commit was signed (`432f041`),
|
||
`main` was fast-forwarded to it, and the signed tag `v1.0.0` was created on it
|
||
and pushed. The current-state wording that said otherwise is corrected. The M11
|
||
report is not edited: its §T records the state at closeout, which was true when
|
||
written.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `README.md` | Status: v1.0.0 released; tag and `main` at `432f041`; v1.1 on `v1.1-development`. | developer docs |
|
||
| `planning/README.md` | Current state, the release events, the map (v1.0.0 released, v1.1 planning), the stop rule restated for work packages, `V1.1-PLAN.md` in the document table and reading order. | index |
|
||
| `planning/BUILD-MILESTONES.md` | Header status, **stale since M8** ("M9 is next"), now says complete and closed. A dated post-release note under M11's status, leaving the closeout paragraph as it stood. The *Post-v1 backlog* points at its triage. | milestone status |
|
||
| `planning/VERSION.md` | This header and entry. | package version |
|
||
|
||
**Part 2 — the v1.1 plan.**
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `planning/V1.1-PLAN.md` | **New.** Every post-v1 backlog item and every other recorded v1 residual risk, triaged. Six v1.1 work packages in order (A1, A2, B, C, D, E), each with objective, rationale, scope, non-scope, affected subsystems, acceptance criteria, regression requirements, dependencies, compatibility and risk. The v1.1 release criteria. A v1.2 list and future work. | planning |
|
||
|
||
**Recommended scope: a focused v1.1.** Context-window safety, protocol-echo
|
||
hardening, independent memory retention, browser coverage, recovery honesty and
|
||
control-boundary contrast. Scheduled backups, identity-diagnostic extensions and
|
||
media adapters are deferred, with the reason for each.
|
||
|
||
**First brief to write:** WP-A1, the context-window safety reserve.
|
||
|
||
**Requirement changes: zero.** Every v1.1 package improves the implementation of
|
||
an existing requirement, so `SPECIFICATION.md` and `V1-ACCEPTANCE-TESTS.md` are
|
||
unchanged.
|
||
|
||
## v4.0 — M11 closeout: v1 release validation accepted (2026-09-14)
|
||
|
||
No requirement change, and no product code change. The closeout re-ran the
|
||
black-box evidence on the exact release-candidate tree, `3652dc6`. Its product
|
||
code is identical to `96c1bf5`, where the 100-turn evidence was taken. The
|
||
browser, offline and identity runs had last been run on `ef25b0a`.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `reports/M11-IMPLEMENTATION-REPORT.md` | **New §S**: the closeout's exact-tree verification, including the identity diagnostic's results, which the report had never carried. **New §T**: the acceptance record. **Corrections**: the REQUIRED FOR V1 count is 82, not 85 (the 85 counted four SHOULD rows and left out H09), and §L's browser narrator was `qwen2.5:3b-instruct`, not the `-16k` model. §A, §B, §D.1, §F (A06, G01), §L, §O, §P and §R now reflect the closeout. | milestone report + corrections |
|
||
| `BUILD-MILESTONES.md` | **M11 marked COMPLETE / ACCEPTED.** A *Post-v1 backlog* section records work that is not a milestone. | milestone status |
|
||
| `V1-ACCEPTANCE-TESTS.md` | **§P V1 Release Gate gains its result.** The §P3 M11 disposition now records the diagnostic's run. | acceptance evidence |
|
||
| `planning/README.md` | Status, milestone map, stop rule, and why the M11 report stays in `reports/`. | index |
|
||
| `README.md` | One status sentence: v1 release validation passed; not yet tagged. | developer docs |
|
||
| `backend/tools/m11_browser.py` | **Harness fix.** G01's import wait could not fail, because the campaign's own title contains "hidden". It now waits for the imported source's row. | release-test infrastructure |
|
||
|
||
**What the closeout found.** No release blocker. One harness defect, fixed. Three
|
||
non-blocking observations, all from the identity run on a 4,096 window:
|
||
|
||
- the extractor leaves event-call syntax and a parroted length hint in stored
|
||
narration;
|
||
- the state rule's fixed example carries fantasy slugs, which the model copied;
|
||
- the state lagged the narration.
|
||
|
||
The owner chose on 2026-09-14 to carry the extractor gap as a residual risk
|
||
rather than change product code. Report §S.4.
|
||
|
||
**Requirement changes: zero.**
|
||
|
||
## v3.9 — M11's long-run evidence, and the two defects it took to get it (2026-09-14)
|
||
|
||
No milestone, no requirement change, and no acceptance claim. M01 was the one
|
||
REQUIRED test v3.8 left outstanding. It now has a complete 100-turn run on commit
|
||
`96c1bf5`, and M01-M04 pass on it. Getting there found two product defects and
|
||
five harness defects. The M11 report's revision (`d198806`) is the evidence.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `V1-ACCEPTANCE-TESTS.md` | **Result blocks for M01-M04**, citing the evidence run. **Correction:** v3.7 said this file carried results for the M11 release run against every REQUIRED test. No such blocks were written. The per-test matrix is the M11 report's §F. The M11 disposition under §P3 also said the report records the identity diagnostic's findings; it does not, and the disposition now says so. | acceptance evidence + correction |
|
||
| `BUILD-MILESTONES.md` | **M11 status block** gains the long-run evidence and the two further product defects. | milestone status |
|
||
| `TECHNICAL-DESIGN.md` | **"Background failure observability" gains the write-lock rule.** **New §15.4**: stored narration carries story only. | as-implemented record |
|
||
| `CONTEXT-AND-MEMORY.md` | **§51 gains what M11 found**: a derived failure has to be recordable, and must not be caused by the turn itself. | as-implemented record |
|
||
| `DATA-MODEL.md` | **§28B corrected.** M11 added two columns, not one. `settings.context_window_override` (migration 94, `ef25b0a`) was never recorded here. | correction |
|
||
| `DECISIONS/013-authoritative-narrative-state-document.md` | **Implementation note** on what "the protocol block is separated from the prose" now has to cover. | as-implemented note |
|
||
| `planning/README.md` | Status, milestone map and stop rule: the long-run evidence is complete. | index |
|
||
| `reports/M11-IMPLEMENTATION-REPORT.md` | **Revised** (`d198806`): M01-M04 on the complete run, §G.0 removed. | milestone report |
|
||
| `DEVELOPMENT.md` | Logging a GPU inference host during a long run, so a hardware fault can be tied to power or ruled out. | developer docs |
|
||
|
||
**The long-run evidence.** 101 accepted turns, three genuine restarts, every
|
||
scheduled history operation, zero failed post-turn passes, and recovery onto a
|
||
clean data directory 16 of 16. M04 passes on a positional precondition: the
|
||
planting turn was outside the history window, which is the acceptance text's own
|
||
"without entire transcript in prompt". The fact was recovered through
|
||
authoritative state, and it reached memory only as the narrator's restatement of
|
||
that state. The repository owner accepted state-based recovery on 2026-09-13.
|
||
|
||
**Two defects the memory bank had been hiding.** No long run had ever had it
|
||
switched on.
|
||
|
||
- **A turn locked out its own post-turn work.** It held SQLite's single write lock
|
||
through the model call. Every memory, summary and status write during the reply
|
||
timed out, and so did the record of the failure. A run reported `complete` with
|
||
two memories and no summary.
|
||
- **The narrator's protocol was stored as story.** The model pasted the state
|
||
section and unfenced proposals into its prose on up to 42 of 104 turns, and
|
||
stored text is replayed as history.
|
||
|
||
**Requirement changes: zero.** The M04 precondition the harness measures was
|
||
corrected to the acceptance text's wording; the test itself is unchanged.
|
||
|
||
## v3.8 — Two gaps closed under M11's own rules, and the cost of a long turn (2026-09-10)
|
||
|
||
No milestone, no requirement change, and no acceptance claim. Three pieces of
|
||
work done while M01 was still outstanding: one hole in §15.2's guarantee, one
|
||
stale statement of fact, and the reason a long campaign cost the same per turn
|
||
however little had changed.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `TECHNICAL-DESIGN.md` | **§15.2 gains "the server that cannot be asked"** — discovery speaks Ollama's native API, nothing restricts the endpoint to Ollama, and on any other server the window goes unverified and the budget uncapped. `Settings.context_window_override` lets an operator state it; a verified window always wins, so a declaration can lower an unknown ceiling into existence and never raise a known one. | as-implemented record |
|
||
| `TECHNICAL-DESIGN.md` | **New §15.3** — the history window moves in blocks. What happens *at* §15.2's ceiling: a window that gave up its oldest action every turn changed the prompt near the front and cost a full re-read every turn. Records the two safety properties and that M11 §P.1's "no performance requirement" still stands. | as-implemented record |
|
||
| `planning/README.md` | **Corrected**: said M11's tree was "staged rather than committed" in two places. It was committed and signed on `m11-release-validation` (`fedb714`); `main` is still M10. | correction |
|
||
| `README.md`, `DEVELOPMENT.md` | The window override and how to set it; why a long campaign is not slow in proportion to its length, with the measurements. | developer docs |
|
||
|
||
**The window a server will not tell you.** §15.2 makes the application ask the
|
||
inference server what it will accept and cap itself to the answer. It asks over
|
||
Ollama's *native* API — and nothing restricts `endpoint_url` to Ollama. Against
|
||
vLLM, llama.cpp's own server, or anything else serving an OpenAI-compatible
|
||
`/v1`, `/api/ps` and `/api/show` are simply absent, discovery fails as designed,
|
||
and the budget stands uncapped at whatever is configured. §15.2's rule names
|
||
Ollama because Ollama is what M8 measured; the failure it forbids is not
|
||
Ollama's. The override closes that without weakening what `verified` claims:
|
||
`verified` still means the server answered, so `window_verified` in a turn's
|
||
provenance counts what M11's report says it counts, and a declared window is
|
||
identifiable as a declaration everywhere it appears.
|
||
|
||
**What a long turn was paying for.** An inference server caches a prompt by its
|
||
prefix. The history window gave up its oldest action every turn, which changed
|
||
the prompt near the front and discarded that cache, so nearly the whole prompt
|
||
was reprocessed every turn regardless of how little had changed. The window now
|
||
snaps to a block and holds, stepping every few turns. Measured against the
|
||
reference deployment on real builder output, at an 8,192-token budget: **124.0s
|
||
per turn against 362.4s** with the floor disabled. The cost is history depth —
|
||
up to a block fewer actions right after a step — bounded by `TRIM_FRACTION` at a
|
||
quarter of the window, which is the dial between recent history and speed.
|
||
|
||
**Requirement changes: zero.** Nothing was retired, relaxed or reclassified.
|
||
§15.3 records explicitly that M11 §P.1 declines to set a performance requirement
|
||
and that this does not invent one: what changed is the cost of a turn, not what
|
||
a turn must contain.
|
||
|
||
## v3.7 — M11 implemented: v1 security, long-run and release validation (2026-09-07)
|
||
|
||
The release-validation milestone. Most of what it changed is evidence rather
|
||
than product; the product changes it did make were each forced by something the
|
||
validation found.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `TECHNICAL-DESIGN.md` | **New §15.2** — the inference window as a ceiling: how it is discovered, where it is enforced, and why there is no hard-coded 4,096. | as-implemented record |
|
||
| `DATA-MODEL.md` | **New §28B** — M11's one column (`narration_length`), and why the window ceiling, `duplicate_names` and `refused` are all deliberately *not* stored. | as-implemented record |
|
||
| `BROWSER-UX-SPEC.md` | **§8A gains an implementation note** — the position indicator that answers it, and the three properties that make it answer it. | as-implemented record |
|
||
| `SECURITY-THREAT-MODEL.md` | **New §42B** — the probe held to §73, the silent partial correction closed as a §69 gap, H09 recorded as not applicable with a test to keep it honest, and the measured contrast. | boundary + measurement notes |
|
||
| `V1-ACCEPTANCE-TESTS.md` | Results for the M11 release run against every REQUIRED test. | acceptance evidence |
|
||
| `BUILD-MILESTONES.md` | **M11 status block**, and the disposition of the four post-M8 playtest findings it owned. | milestone status |
|
||
| `README.md`, `DEVELOPMENT.md` | The context-window behaviour, `contextwindow.py` in the architecture map, and how to re-run the six release harnesses. | developer docs |
|
||
| `reports/M11-IMPLEMENTATION-REPORT.md` | New. The evidence package for independent release review. | milestone report |
|
||
| `reports/M10-IMPLEMENTATION-REPORT.md` | Moved to `archive/milestone-reports/`. | report rotation |
|
||
|
||
**The release blocker M11 was given, and what it cost.** M8 measured a
|
||
deployment enforcing 4,096 tokens while the application budgeted 16,384, with
|
||
every request returning 200 and `llama.cpp` silently dropping the oldest tokens —
|
||
which here are the narrator's rules and the campaign canon. M11 makes the
|
||
application ask the server what it will accept and cap itself to that, or say
|
||
that it could not check. Not a hard-coded number, not a cloud probe, not a
|
||
guess from the model's name: `/api/ps` for a loaded model, `/api/show` for one
|
||
that is not, under the same endpoint policy and TLS trust as inference.
|
||
|
||
**Two product defects found by the validation itself**, both of the same
|
||
family — something true that nobody was told:
|
||
|
||
- a **manual state correction that was partly refused** returned an unqualified
|
||
success. Found because the identity diagnostic's own fixture was refused that
|
||
way and the run proceeded silently on a degraded campaign;
|
||
- the **narration-length setting moved no number** (post-M8 finding C), so the
|
||
numeric hint the model reads said the same thing for brief, medium and long.
|
||
|
||
**Zero requirement weakenings.** No acceptance test was retired, relaxed or
|
||
reclassified. H09 is reported NOT APPLICABLE on the condition its own text
|
||
states, and that condition is now enforced by a test.
|
||
|
||
## v3.6 — M10 implemented: future media extension hooks only (2026-09-07)
|
||
|
||
M10 built the media seam and no media. The documentation change is mostly a
|
||
record of what was deliberately **not** built, because that is the part a later
|
||
reader will otherwise re-litigate.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `DATA-MODEL.md` | **New §28A**, the media extension points as implemented: one `visual_profiles` table, no scenes table, no job or asset tables, and why each. | as-implemented record |
|
||
| `TECHNICAL-DESIGN.md` | **New §15.1** under the Scene and Future Media Boundary: what the seam turned out to be, what the packet carries and excludes, the STT asymmetry, the endpoint policy. | as-implemented record |
|
||
| `MEDIA-EXTENSION-CONTRACT.md` | **New §90**, appended rather than woven in so the Phase 0B contract stays readable. Records the three places implementation answered an open question, and the sections left unbuilt. | as-implemented record |
|
||
| `BUILD-MILESTONES.md` | **M10 status block**: the finding that shaped the milestone, what shipped, and the migration defect its own tests caught. | milestone status |
|
||
| `V1-ACCEPTANCE-TESTS.md` | **K01-K04 results.** K01 passes and *was already passing*; K04 is reported on the acceptance text's deferred branch, as its own wording provides for. | acceptance evidence |
|
||
| `SECURITY-THREAT-MODEL.md` | **New §42A.** The trust boundary did not widen; in one place the implementation is deliberately **narrower** than §73 permits, which is still a discrepancy worth recording. | boundary note |
|
||
| `README.md`, `DEVELOPMENT.md` | The `media/` package in the architecture map, the backend test count (1,189), and the stricter rule a future media endpoint will meet. | developer docs |
|
||
| `reports/M10-IMPLEMENTATION-REPORT.md` | New. The implementer's account, written for a reviewer. | milestone report |
|
||
|
||
**The decision behind the whole milestone**: the scene snapshot the media
|
||
contract asks for **already existed** as `narrative_state["scene"]`, built by M5
|
||
and carrying lineage correctly since. Verified with a probe rather than taken
|
||
from an earlier report. So M10 added no scenes table, and derives the scene
|
||
packet on read.
|
||
|
||
**Two decisions recorded rather than assumed:**
|
||
|
||
- **The bundle format stays `ai-dnd-adventure-v3`.** M9's own semantic test —
|
||
does omission create ambiguity about what an older file *could* have recorded?
|
||
— says no: a campaign with no visual profiles is the ordinary case, so an
|
||
absent key unambiguously means "none". Older v3 files still import.
|
||
- **M10 adds no migration.** A `CREATE INDEX` migration was written first and
|
||
removed: `create_all` already builds the table and the index declared on its
|
||
column, so the migration left an upgraded database holding an index a fresh
|
||
install did not have. `LATEST_VERSION` stays 92.
|
||
|
||
**The four post-M8 playtest findings below remain M11's** and are untouched by
|
||
this entry.
|
||
|
||
## v3.5 — Post-M8 hands-on playtest findings recorded (2026-09-07)
|
||
|
||
**Documentation only. No application code changed, and M9's verified result is
|
||
untouched** — see the note at the end of this entry.
|
||
|
||
A real play session against **accepted, signed M8** (real browser, trusted-LAN
|
||
Ollama, `qwen2.5:3b-instruct-16k`, disposable database since destroyed) surfaced
|
||
four product-quality observations. **None is an M9 defect, none was caused by
|
||
M9, and none blocks M9 acceptance.** They are recorded so they cannot be lost
|
||
when the M9 report is archived.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `reports/M9-IMPLEMENTATION-REPORT.md` | **New §Y**, the full write-up with the verification behind each mechanism, plus a pointer from §W. Explicitly labelled as not-M9. | observation record |
|
||
| `BUILD-MILESTONES.md` | **New section before M10**, the durable copy owned by M11; a note bounding M10 out of it; the four items added to M11's scope list. | milestone sequencing |
|
||
| `V1-ACCEPTANCE-TESTS.md` | **New §P**, three items as **test-design tasks, explicitly not acceptance tests**, each naming what must be settled before it could become one. No existing test changed or weakened. | test design |
|
||
| `BROWSER-UX-SPEC.md` | **New §8A.** §8's *"The current endpoint should be clear"* was satisfied while a real reader was lost, so it could not hold the behaviour. States the orientation requirement; **prescribes no wording**. | requirement clarification |
|
||
| `TEST-CAMPAIGN-FIXTURE.md` | **New Appendix A** proposing a companion `Multi-Character Identity Test`. The established deterministic fixture is **unchanged** — altering it would invalidate earlier milestones' comparisons. | test design |
|
||
|
||
**What was verified rather than assumed**, since the playtest campaign no longer
|
||
exists and root causes largely cannot be proven:
|
||
|
||
- The browser title genuinely is `AI D&D` (`frontend/index.html`), and **no
|
||
accepted document ever claimed otherwise** — so this is an uncovered gap, not
|
||
documentation needing correction. Nothing was corrected and nothing was
|
||
renamed: *Adventure Storyteller* is itself narrower than the genre-agnostic
|
||
engine `SPECIFICATION.md` requires, and the naming decision is the owner's.
|
||
- The narration-length setting adds **one English sentence** and changes **no
|
||
generation budget**, while the numeric hint derived from the global
|
||
`max_output_tokens` is **identical for every setting** — measured at the
|
||
default as *"must not exceed 506 words, and it should not stop short of about
|
||
177."* Recorded as a mechanism to check first, **not** as the proven cause.
|
||
- The narrative state **permits two entities to share a display name and
|
||
reports nothing** — `DUPLICATE_ENTITY` rejects a repeated key only. That is
|
||
one of the identity finding's failure modes; it establishes nothing about what
|
||
actually happened.
|
||
|
||
**M9 is unaffected.** No application file changed in this pass, so M9's final
|
||
verified result stands exactly as recorded: **1,102 backend passed / 14 skipped
|
||
/ 0 failed**, 145 frontend, 36/36 browser. The expensive M9 suites were
|
||
deliberately **not** re-run, because only Markdown changed.
|
||
|
||
## v3.4 — M9 Implementation (2026-09-07)
|
||
|
||
M9 — Export, Backup, Recovery, and Migration Hardening — is implemented on
|
||
`m9-recovery` from the signed M8 commit `1ce9972`. This revision records what the
|
||
implementation settled. It is **not** an acceptance: the milestone report is
|
||
written for a reviewer and the tree is staged for the repository owner's signed
|
||
commit.
|
||
|
||
**Requirement corrections:** none. M9 altered no product requirement.
|
||
`SPECIFICATION.md` and `SECURITY-THREAT-MODEL.md` are unchanged — §16 already
|
||
required the export to preserve the exact active position, §6.1 already required
|
||
an exact prompt/context snapshot per turn, and M9 implements both rather than
|
||
redefining either.
|
||
|
||
| Document | Change | Kind |
|
||
| --- | --- | --- |
|
||
| `DATA-MODEL.md` §29 | The v3 format, the three-category rule (chosen / evidence / rebuildable), what each version can be trusted to say, the encoding of the snapshots, and the two pointers the import translates. | implementation fact |
|
||
| `TECHNICAL-DESIGN.md` §9.3 | Why the version was bumped when §9.1 and §9.2 each correctly declined one; the third data category; the two-phase transaction and the warning path for a failed derived rebuild. | implementation fact |
|
||
| `TECHNICAL-DESIGN.md` §9.4 | **New.** The SQLite backup: the online backup API rather than a file copy, the verify-then-rename order, and why there is no restore endpoint. | implementation fact |
|
||
| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **New subsection.** Story Cards settled as compatibility-only legacy data and removed from the narrator's prompt, with the evidence that they were the "alternate untracked path" §73 already forbade. | newly settled design decision |
|
||
| `V1-ACCEPTANCE-TESTS.md` I01-I07, L02-L04 | Results recorded. I05's M7-era limit is marked closed with the original paragraph kept, because the M9 decision is only legible against it. L04 records the defect running it found. | implementation fact |
|
||
| `BUILD-MILESTONES.md` M9 | Marked complete, with what it delivered, the three defects it found, the story-card decision, and the debt carried forward. | implementation fact |
|
||
| `README.md`, `VERSION.md` | Status. | implementation fact |
|
||
| `DEVELOPMENT.md` | **New section**: the two recovery tools and when each applies, taking a backup, and the stop-move-start restore procedure. Plus a note that an imported long campaign meets a small context ceiling on its first turn rather than gradually. | implementation fact |
|
||
|
||
**The four M8 handoff questions, answered**
|
||
|
||
| | Answer |
|
||
| --- | --- |
|
||
| **A. Complete campaign portability** | Every family travels and is measured family by family, before and after, by a tool a reviewer can rerun. |
|
||
| **B. Historical prompt provenance** | **It belongs in the bundle, and it is in it.** An old turn in a restored campaign shows what it was actually given, after the source has been deleted and the canon edited. |
|
||
| **C. Legacy story cards** | Compatibility-only. Carried in both directions; removed from the narrator's prompt; still the summariser's character roster. |
|
||
| **D. Context-window portability** | The campaign travels; the machine's model configuration does not. Importing changes no setting of the destination's, and a campaign imports whether or not any model is installed. The window itself remains M11's. |
|
||
|
||
**What the implementation found rather than assumed**
|
||
|
||
Three defects, all found by running the milestone's own tests rather than by
|
||
reading: an FTS index leak that made an ordinary import fail in an unrelated
|
||
campaign and that Reindex could not repair; an imported node with no state
|
||
snapshot being stamped with the campaign's *head* state; and a snapshot pointer
|
||
that was not being translated because the code mutated a dict in place. The
|
||
first two predate M9.
|
||
|
||
One measurement changed a plan, and then corrected the conclusion drawn from it.
|
||
Carrying per-turn prompts looked like it would halve the length of campaign that
|
||
can be restored. Measured, and after compressing them inside the file, everything
|
||
M9 added costs **12%** of reachable campaign length — the import ceiling moves
|
||
from about 318 turns to about 279, against a 100-turn certification target. The
|
||
dominant cost is not M9's at all: the **per-position narrative state document is
|
||
74% of a bundle**, and v2 already carried it.
|
||
|
||
## v3.3 — M8 Closeout (2026-09-06)
|
||
|
||
M8 is **complete and accepted**. The independent review returned
|
||
`M8 IMPLEMENTATION: PASS` subject to evidence and documentation cleanup; this
|
||
revision is that cleanup. No product requirement changed and no application code
|
||
changed in it.
|
||
|
||
**What M8 leaves the package with**
|
||
|
||
- **The browser is now the intended v1 storyteller surface**, not an adapted
|
||
AI-DnD one. One entry point, one story screen, one natural-language field;
|
||
State, Knowledge, Context, Save Points and campaign Settings one layer deeper
|
||
behind panels that start closed. Branch, fork, node, head and depth appear
|
||
nowhere a reader can see them — audited in source and in the live DOM, 0 hits.
|
||
- **Automated frontend testing exists for the first time** — 132 tests across 10
|
||
files. The project had none before M8.
|
||
- **Hidden narrator information is withheld at the surface that can actually
|
||
expose it** — advanced context inspection — rather than through a fictitious
|
||
separate hidden-state subsystem. Withheld text is absent from the DOM, not
|
||
collapsed inside it.
|
||
- **Final acceptance evidence:** 950 backend passed / 14 skipped / 0 failed; 132
|
||
frontend passed; lint, production build and Docker build clean; **157 browser
|
||
checks across six suites, zero failures**, on one frozen build; no schema
|
||
change, proved against an M7-built database.
|
||
- **M9 — Export, Backup, Recovery, and Migration Hardening — is next.** Its four
|
||
handoff questions are in the M8 report's §U.
|
||
|
||
**What the closeout settled**
|
||
|
||
- **The build-evidence contradiction.** The M8 report named two different
|
||
frontend bundles as the artifact behind its acceptance evidence. The saved run
|
||
logs settle it: `index-Ii-lARp9.js`, built 18:53:02 from the staged tree, is
|
||
the one **final frozen** artifact behind all 157 browser checks.
|
||
`index-C6E5Uvtu.js` is **superseded** — it predates finding 7's fix and its
|
||
acceptance suite ended 54/55 on exactly that defect. The report's §P now sets
|
||
the two side by side, and §B, §Q and §S agree with it.
|
||
- **Finding 14 — resolved, operationally, with no application change.** Ollama's
|
||
OpenAI-compatible endpoint ignores `num_ctx` and reloads the model at its own
|
||
default, so the window cannot be set per request from where this application
|
||
stands. A **derived model** created over `/api/create` carries the parameter,
|
||
is honoured through the application's own OpenAI-compatible path, and appears
|
||
in `/v1/models` — so the existing Settings model picker finds it with no code
|
||
change. Measured end to end. It is now an operational note in
|
||
`DEVELOPMENT.md`, not an unresolved M11 blocker.
|
||
- **`BROWSER-UX-SPEC.md` §38 — ratified.** The change from *"Show Hidden Story
|
||
State"* to the requirement on **hidden narrator information / spoilers** is
|
||
accepted as a **requirement clarification aligned with the implemented
|
||
architecture**, not a weakening: ordinary Story and State surfaces expose no
|
||
narrator-only information; the advanced Context and Knowledge surfaces that
|
||
could withhold it by default; revealing it takes an explicit, warned action;
|
||
withheld material is **absent from the DOM**, not merely collapsed; and no
|
||
duplicate hidden-state subsystem is required to satisfy obsolete UI wording.
|
||
Verified by the sentinel suite, 21/21.
|
||
|
||
**Documents changed**
|
||
|
||
- `reports/M8-IMPLEMENTATION-REPORT.md` — §A, §B, §D, §P, §Q, §S (findings 13
|
||
and 14), §T, §U and §V. §U gains the **M9 handoff**: complete campaign
|
||
portability, historical prompt/context provenance in the bundle, legacy story
|
||
cards, and context-window portability. Verification figures unchanged.
|
||
- `BROWSER-UX-SPEC.md` §38 — one requirement added: withheld material must be
|
||
**absent from the rendered DOM**, not merely visually collapsed. A closed
|
||
`<details>` is still findable by browser search and by the accessibility tree,
|
||
which is how finding 3 leaked.
|
||
- `PROJECT-SOURCES.md`, `V1-ACCEPTANCE-TESTS.md` — three pointers still aimed at
|
||
`reports/M7-IMPLEMENTATION-REPORT.md`, which M8's rotation moved to the
|
||
archive.
|
||
- `BUILD-MILESTONES.md` — M8 marked **COMPLETE / ACCEPTED**, M9 marked next and
|
||
not started, finding 14's carried-forward entry reworded as resolved.
|
||
- `README.md` (planning) — M8 complete and accepted; next is M9.
|
||
- `VERSION.md` — this entry.
|
||
- `DEVELOPMENT.md` — the derived-model procedure (carried from the finding-14
|
||
investigation).
|
||
|
||
**Status discipline.** M8 is implemented, verified, reviewed and accepted. It is
|
||
**not committed**: the tree is staged for the repository owner's signature. M9
|
||
has not been started.
|
||
|
||
---
|
||
|
||
## v3.2 — M8 Final Verification Pass (2026-09-06)
|
||
|
||
The corrective, verification and reporting pass over M8. At the time of this
|
||
revision M8 was **implemented and verified, not accepted**; v3.3 records the
|
||
review outcome and the acceptance.
|
||
|
||
**Documents changed in this pass**
|
||
|
||
- `BROWSER-UX-SPEC.md` §38 — **requirement clarification.** Rewritten from
|
||
"Hidden Narrator State" with a `Show Hidden Story State` toggle to
|
||
"Hidden Narrator Information / Spoilers". The protection required is
|
||
unchanged and, if anything, stated more strictly; what changed is that it
|
||
no longer names a state subsystem that does not exist, and it now forbids
|
||
inventing one. §100's checklist entry follows it.
|
||
- `TECHNICAL-DESIGN.md` — **implementation fact.** Campaign canon is
|
||
configuration; its provenance is the per-turn context snapshot, measured
|
||
rather than assumed.
|
||
- `V1-ACCEPTANCE-TESTS.md`, `BUILD-MILESTONES.md`, `README.md`,
|
||
`DEVELOPMENT.md` — implementation facts and final counts.
|
||
|
||
**No product requirement was weakened.** The §38 change is the only one that
|
||
touches a requirement's wording, and it tightens it.
|
||
|
||
---
|
||
|
||
## v3.1 — M8 Implementation (2026-09-06)
|
||
|
||
M8 turned the adapted AI-DnD interface into the interactive-story workspace
|
||
`BROWSER-UX-SPEC.md` describes. Not yet accepted — the report is written for an
|
||
independent reviewer.
|
||
|
||
**Documents changed**
|
||
|
||
- `BROWSER-UX-SPEC.md` — four "As implemented in M8" notes (§12 one input, §31
|
||
the state panel and why there is no hidden-state toggle there, §55 the context
|
||
inspector's default view, §57 click-through and the narrator-only guard). No
|
||
requirement was altered.
|
||
- `TECHNICAL-DESIGN.md` — new §13.4, recording that the browser stayed a
|
||
presentation layer, the failure taxonomy, and why markup is never produced
|
||
from input.
|
||
- `BUILD-MILESTONES.md` — M8's outcome and the debt it carries forward.
|
||
- `V1-ACCEPTANCE-TESTS.md` — B01-B04 and D01-D14 recorded as browser evidence.
|
||
- `README.md`, `DEVELOPMENT.md` — one input rather than four modes, the context
|
||
inspector, the component test suite, and the removal of "no frontend tests"
|
||
from the inherited-debt list.
|
||
|
||
**What M8 established that the plan did not already say**
|
||
|
||
- The narrative state has **no hidden dimension**, so §38's `Show Hidden Story
|
||
State` has nothing to reveal there. A campaign's secrets live in narrator-only
|
||
knowledge sources, and the only ordinary screen that can surface one is the
|
||
context inspector — which is where the guard was built.
|
||
- `campaign_canon` had no API. It is the highest authority in a campaign, read
|
||
by both the prompt builder and the state validator since M5, and until M8 a
|
||
fixture had to write it with SQL.
|
||
- AI Dungeon's `> You {text}` player-input convention is incompatible with one
|
||
natural-language field: it produced `> You I enter the tavern.`, which a small
|
||
model then imitates. Corrected for first-person input.
|
||
|
||
---
|
||
|
||
## v3.0 — M7 Closeout (2026-09-06)
|
||
|
||
M7 is complete. Two items the corrective pass had left open are resolved.
|
||
|
||
**The embedding-model calibration boundary.** `SEMANTIC_FLOOR = 0.58` was
|
||
measured against `nomic-embed-text`, and the corrective pass documented only the
|
||
safe half of that: a model scoring everything lower degrades to lexical-only. A
|
||
model scoring unrelated material *higher* would have recreated M7-F1 on a build
|
||
whose tests all pass. Semantic admission is now **per model**: an uncalibrated
|
||
model does not inherit the threshold, semantic retrieval is skipped for it with
|
||
the reason reported, and the library degrades to lexical-only. Recorded in
|
||
`TECHNICAL-DESIGN.md` §13.3 and `IMPORTED-KNOWLEDGE-DESIGN.md` §76.
|
||
|
||
**The ambiguous `export/import 53/54`.** The 54th case was a false positive in
|
||
the independent review's own harness — its "no filesystem path" assertion was a
|
||
substring test that fired on `text/markdown`, a MIME type. Replaced with three
|
||
precise checks; the suite is **56/56** and no product behaviour was involved.
|
||
|
||
**What closeout changed in the active documents:**
|
||
|
||
- `TECHNICAL-DESIGN.md` §13.3 — **new.** A similarity threshold is a property of
|
||
the model, the two ways a different model breaks it are not symmetric, and the
|
||
product refuses to apply a threshold to a model it has not measured.
|
||
- `IMPORTED-KNOWLEDGE-DESIGN.md` §76 — the same, in the design's own terms:
|
||
§25's local-Ollama embedding stands; what is added is that a *threshold* must
|
||
be measured before it is trusted.
|
||
- `BUILD-MILESTONES.md` M7 — marked COMPLETE, with the capabilities later
|
||
milestones inherit and the debt carried forward, including that calibrating
|
||
further embedding models is a measurement rather than a guess.
|
||
- `V1-ACCEPTANCE-TESTS.md` — the M7 results promoted from implementation-pass
|
||
evidence to reviewed results.
|
||
|
||
## v2.9 — M7 Independent Review and Corrective Pass (2026-09-06)
|
||
|
||
The review returned *PASS WITH CORRECTIVE WORK REQUIRED*. It closed the five
|
||
acceptance conditions the implementation had flagged as unmeasured — C05, G06,
|
||
G07, G10 and hidden Canon, all exercised against a real narrator and all
|
||
passing — and found two blocking defects, both now corrected.
|
||
|
||
**M7-F1 — imported knowledge was injected regardless of relevance.** Relevance
|
||
was decided by a floor expressed as a share of the best candidate, which the
|
||
best clears by construction. A query about tide tables and container tonnage
|
||
retrieved all five sources of a fantasy campaign, narrator-only hidden Canon
|
||
among them. Corrected by separating relevance **admission** from **ranking**.
|
||
|
||
**M7-F2 — the retrieval suite could not detect it.** Its stub scored unrelated
|
||
text an order of magnitude lower than the real model, so the broken gate passed.
|
||
Corrected with a stub that has the real model's similarity floor, plus a test
|
||
that fails if the floor is removed and one that shows the superseded rule still
|
||
being fooled. The new suite fails 13/18 against the pre-corrective code.
|
||
|
||
**What the corrective pass forced into the active documents:**
|
||
|
||
- `TECHNICAL-DESIGN.md` §13.2 — **new.** Relevance admission is a separate stage
|
||
from ranking, and the general rule behind it: a relevance decision must rest on
|
||
a signal meaningful on its own, because normalization answers "which of these
|
||
is best" and can never answer "is any of these any good". A pipeline that ranks
|
||
first and cuts second has no way to return nothing.
|
||
- `TECHNICAL-DESIGN.md` §13.1 — the pipeline diagram gains the admission stage.
|
||
- `IMPORTED-KNOWLEDGE-DESIGN.md` §76 — retrieval corrected: admission before
|
||
authority, and the plain statement that **retrieval may return nothing**,
|
||
which is what §30 means when every source is irrelevant.
|
||
- `BUILD-MILESTONES.md` M7 § Status — both findings, their corrections, and the
|
||
model-specific calibration recorded as carried debt.
|
||
|
||
Three non-blocking findings were folded in: a relevance constant that could
|
||
never fire was removed rather than re-tuned; the acceptance tests moved to the
|
||
standard `TEST-CAMPAIGN-FIXTURE.md` §12 files so G07's trap is finally
|
||
exercised; and Unicode format characters are stripped from displayed filenames.
|
||
A pre-existing M5 narrator-protocol issue was recorded and deliberately left
|
||
with M5.
|
||
|
||
## v2.8 — M7 Implementation Pass (2026-09-06)
|
||
|
||
**Not a closeout.** M7 is implemented, not accepted, and this revision records
|
||
what the implementation pass built and measured so that an independent review
|
||
has something to verify against. No milestone report was written: the
|
||
convention this package follows puts the report in `reports/` and has the
|
||
*reviewer* write it, treating the build summary as claims to check.
|
||
|
||
**M7 — First-Class Imported Knowledge Library.** A campaign can import local
|
||
`.txt` and `.md` files as Canon, Reference or Inspiration; retrieval is hybrid
|
||
(SQLite FTS5 plus local Ollama embeddings), reranked by relevance × class,
|
||
bounded by its own token budget, framed in the prompt as untrusted data with the
|
||
authority order stated in words, and fully traceable in the Insights panel. It is
|
||
a separate subsystem: AI-DnD's Story Cards were not promoted into it and are
|
||
untouched.
|
||
|
||
**What implementation forced into the active documents:**
|
||
|
||
- `TECHNICAL-DESIGN.md` §13.1 — **new.** The implemented pipeline, and four
|
||
decisions that each replaced an obvious wrong one: the class multiplies
|
||
relevance rather than adding to it; both retrieval scores are normalized per
|
||
query against the best of their own path; the relevance floor is therefore
|
||
relative rather than absolute; and lexical retrieval is a production path
|
||
rather than a fallback.
|
||
- `DATA-MODEL.md` §24A — **new.** The three tables and the FTS5 virtual table,
|
||
and the line between what the reader gave the campaign and what the machine
|
||
derived from it. Only a source's content and its classification are not
|
||
derivable.
|
||
- `DATA-MODEL.md` §25 — the retrieval record is **not** a table. It lives in the
|
||
turn's own context snapshot and carries the *rendered text*, because a table of
|
||
foreign keys would turn every historical turn's evidence into dangling
|
||
references the moment a source were deleted.
|
||
- `DATA-MODEL.md` §29 — what the bundle carries for imported knowledge, and why
|
||
passages, index rows and vectors are rebuilt rather than exported.
|
||
- `CONTEXT-AND-MEMORY.md` §29, §41-42, §46 — the knowledge budget as implemented
|
||
(a protected cap for always-included Canon, a share of the rest filled in
|
||
authority order), always-include as a Canon-only mechanism, and hidden Canon as
|
||
prompt discipline rather than as filtering.
|
||
- `IMPORTED-KNOWLEDGE-DESIGN.md` §76 — **new.** Where this document offered
|
||
options, which was chosen and why; and, named rather than left to be
|
||
discovered, the six things it contemplates that M7 does **not** implement —
|
||
entity linking, tags, manual priority, scene pinning, Canon-versus-Canon
|
||
conflict detection, and source versioning.
|
||
- `V1-ACCEPTANCE-TESTS.md` — results for G01-G10, C05, F05, F06, I05 and
|
||
H06-H09, marked as implementation-pass evidence rather than review findings.
|
||
F05 and F06 move from *PARTIAL / PASS for story memory* to complete. H09 is
|
||
recorded NOT APPLICABLE with a test that fails if an archive extractor is ever
|
||
added to this surface. **C05 is recorded as a pass on the assembled prompt
|
||
with the gap stated**: no real narrator generation was run against it.
|
||
- `BUILD-MILESTONES.md` M7 § Status — **new.** What was built beyond the scope
|
||
list, and the debt carried forward, deliberately.
|
||
|
||
**One runtime dependency was added**: `python-multipart`, Starlette's multipart
|
||
parser. It is what makes the upload surface possible, and the upload surface is
|
||
why no endpoint in the knowledge API accepts a filesystem path.
|
||
|
||
## v2.7 — M5 and M6 Closeout (2026-09-06)
|
||
|
||
Two milestones, and one lesson they share: an independent review found a real
|
||
defect in each after the implementation reported success, and in both cases the
|
||
defect was invisible to the tests the implementation had written for itself.
|
||
|
||
**M5 — Genre-Neutral Authoritative Narrative State.** Accepted 2026-09-04. Its
|
||
review returned *PASS WITH CORRECTIVE WORK REQUIRED*: editing a narrator turn
|
||
rewound the campaign's live state while the head stayed at the tip, breaking the
|
||
`transcript position == head == authoritative state` invariant. The corrective
|
||
pass rebuilt narrator editing on the §§14-15 fork semantics. Recorded here late:
|
||
M5's closeout did not add a package revision entry, and this one covers it.
|
||
|
||
**M6 — Branch-Safe Context, Summaries, and Long-Term Story Memory.** Accepted
|
||
2026-09-06. Sequence:
|
||
|
||
1. **Implementation.** Summaries moved onto lineage-anchored rows; memory
|
||
authority, provenance, an output reserve, and observable derived-work
|
||
failure were added. Reported as passing, including E03.
|
||
2. **Independent review — E03 still failed.** Summary *rows* were anchored, but
|
||
generation was seeded from `adventures.story_summary`, a campaign-global
|
||
column with no lineage. After a divergence the summariser was handed the
|
||
abandoned line's prose and asked to update it, so the new summary carried
|
||
abandoned content inside a correctly anchored row. The review also found F02
|
||
passing by a single retrieval slot, two test defects, and a misleading
|
||
derived-work status.
|
||
3. **Corrective pass.** Generation is seeded from `summaries.current`;
|
||
redundancy suppression was added before the retrieval cut; the test fixture
|
||
no longer makes real network calls; the browser suite gained a scenario that
|
||
regenerates a summary after divergence.
|
||
|
||
**What the planning package learned from M6, recorded in the active documents:**
|
||
|
||
- `CONTEXT-AND-MEMORY.md` §11 — lineage safety has two halves. Anchoring the
|
||
output row is not enough; the *input* to the summarizer must be scoped by the
|
||
same rule.
|
||
- `V1-ACCEPTANCE-TESTS.md` E03 — a valid test must regenerate a summary after
|
||
diverging. Checking only that the old row went ineligible passes while the
|
||
defect is live.
|
||
- `CONTEXT-AND-MEMORY.md` §20/§22 — what M6 implements (redundancy suppression
|
||
between memories) is now distinguished from what it does not (importance,
|
||
entity, recency and thread ranking; cross-layer duplication).
|
||
|
||
## v2.6 — M4 Closeout (2026-09-03)
|
||
|
||
M4 is **accepted**. Its review returned *PASS WITH CORRECTIVE WORK REQUIRED*; the
|
||
corrective work is done, and the browser condition that M3 and M4 both carried is
|
||
closed.
|
||
|
||
**The three review findings, fixed:**
|
||
|
||
- **B-1 — the Save Point list was an N+1.** It resolved each Save Point with its
|
||
own query and loaded whole `Action` rows, narration included, to answer "does a
|
||
row exist here". It is now one bulk two-column coordinate query plus one
|
||
lineage computation: **53 SELECTs for 25 Save Points became 5**, and the count
|
||
no longer grows with the list. Guarded by three tests, including one proving
|
||
the coordinate is matched as a *pair* — an `IN`-list version would report a
|
||
Save Point resolved because another branch has a live row at the same depth.
|
||
- **B-2 — deleting a branch silently deleted its Save Points.** Fixed as a
|
||
**behaviour** defect rather than a missing warning, because
|
||
`STORY-BRANCH-SEMANTICS.md` §19 says a named checkpoint remains until
|
||
explicitly deleted and §28 already required future cleanup to retain
|
||
checkpoint-referenced paths. **A branch a Save Point names can no longer be
|
||
deleted.** The request is refused with the offending Save Points named; the
|
||
user deletes them explicitly, which deletes no story, and the branch then
|
||
goes. Both delete controls disable and explain. Recorded as a new
|
||
**§19.1**.
|
||
- **B-3 — the D11/L03 automation never left one process.** A new module spawns
|
||
real server processes, kills the first, and reads the campaign back with the
|
||
second.
|
||
|
||
**Also fixed (review §S C-5):** creating a Save Point now takes the campaign's
|
||
turn lock, so "save where I am" cannot read a head a turn in flight is about to
|
||
move. Rename and Delete deliberately do not take it, and a test pins that
|
||
decision.
|
||
|
||
**Real-browser verification — the first in this project.** A Firefox 154.0.1
|
||
driven through geckodriver over the W3C WebDriver protocol exercised the rendered
|
||
DOM for **both** milestones: **44/44 checks passed**, no console errors. It
|
||
covered M3's Undo/Redo enable states, transcript movement, Retry and the take
|
||
pager, and divergence retiring Redo; and M4's whole Save Point lifecycle
|
||
including both confirmations and the new branch-delete warning. **The outstanding
|
||
M3 browser condition is therefore closed as well.** No dependency was added to
|
||
the repository: the WebDriver client for the run was written against stdlib HTTP.
|
||
|
||
**Also corrected, found while fixing B-2:** `models.py`, `TECHNICAL-DESIGN.md`
|
||
§8.8 and `DATA-MODEL.md` §8 all described the cascade as the durability rule.
|
||
They now describe the refusal, and record that `checkpoints.branch_id`'s cascade
|
||
survives as referential integrity that the application no longer reaches.
|
||
|
||
**Documents corrected by this closeout:**
|
||
|
||
- `V1-ACCEPTANCE-TESTS.md` records results for **D11-D14, I04, L03** and the
|
||
E-series, and states that the browser-level condition is satisfied. **No pass
|
||
condition was weakened** — and D11/L03 now note that the automation crosses a
|
||
genuine OS process boundary, which is the standard later milestones should
|
||
meet.
|
||
- `DATA-MODEL.md` §8 records the coordinate as implemented, with the retry
|
||
measurement that settles coordinate-versus-turn-id, and three decisions that
|
||
were previously implicit: names are not unique, several Save Points may name
|
||
one position, and the list is newest-created first.
|
||
- `STORY-BRANCH-SEMANTICS.md` gains **§19.1** — a checkpoint protects the
|
||
history it names. This is the only behavioural specification change in the
|
||
closeout, and it strengthens §19 rather than weakening anything.
|
||
- `BROWSER-UX-SPEC.md` §25 rules for the implemented vocabulary: **Moment N**,
|
||
not *Turn N*, because the branch panel and tree overlay already count in
|
||
moments. Vocabulary only; no behaviour changes.
|
||
- `BUILD-MILESTONES.md` marks **M4 COMPLETE**, records the fixes and the browser
|
||
result, and warns M5 that the instrumentation to move is now **55 tests**.
|
||
- `README.md` records M1-M4 accepted and M5 as next to brief.
|
||
- `PROJECT-SOURCES.md` and `project-sources.txt` point at the current report;
|
||
`project-sources.txt` still named the archived M3 report and was corrected.
|
||
|
||
**Report rotation** happened in the reporting pass that preceded this closeout:
|
||
`M3-IMPLEMENTATION-REPORT.md` moved to `archive/milestone-reports/` as a pure
|
||
rename, and `reports/` now holds M4's report, whose **§W** is this closeout's
|
||
evidence record.
|
||
|
||
**No new ADR.** ADR 012 already decides the architecture, and the corrective work
|
||
forced no new architectural decision. `SPECIFICATION.md`,
|
||
`STORY-BRANCH-SEMANTICS.md`, `SECURITY-THREAT-MODEL.md`, `CONTEXT-AND-MEMORY.md`,
|
||
`IMPORTED-KNOWLEDGE-DESIGN.md` and ADRs 003, 005 and 012 are unchanged.
|
||
|
||
**M5 readiness:** ready. Save Points store no state and no checkpoint code reads
|
||
any, so M5 can change what a snapshot contains without touching what a Save Point
|
||
is — provided it keeps state recoverable at a position without replay
|
||
(`TECHNICAL-DESIGN.md` §10.4).
|
||
|
||
## v2.5 — M4 Implementation (2026-09-03)
|
||
|
||
M4 added durable named Save Points. This revision records only what the
|
||
implementation established as fact; **no product requirement changed**, and the
|
||
milestone is **not** marked accepted — its review has not been written.
|
||
|
||
- `TECHNICAL-DESIGN.md` gains **§8.8** and **§9.2**: the Save Point as a name
|
||
plus a coordinate holding no story, restore as head movement with a bounds
|
||
check, the rule that the branch half of the head moves only when the
|
||
coordinate is off the path being read, restore never forking, and the bundle
|
||
carrying Save Points independently of the head.
|
||
- `DATA-MODEL.md` **§8** records the pointer as implemented — `(branch, depth)`
|
||
rather than a turn id, with the reason: one coordinate holds every attempt at
|
||
a turn and exactly one is live, so a coordinate follows a retry where a row id
|
||
would pin a superseded take. **§29** records the checkpoints in the export.
|
||
- `BUILD-MILESTONES.md` **M4** gains a status block: what shipped, the one
|
||
architectural decision the milestone had to make and why it needed no new ADR,
|
||
the test count, and the outstanding browser condition.
|
||
- `README.md` describes Save Points as a user-facing capability.
|
||
|
||
**No ADR was created.** ADR 012 already decides the architecture M4 needed —
|
||
restore reuses active-head movement — and a table is not a decision. The one
|
||
question ADR 012 does not answer, whether restore moves the branch half of the
|
||
head, is that same mechanism applied to a coordinate ADR 012 already defines;
|
||
`TECHNICAL-DESIGN.md` §8.8 records the answer rather than a new ADR asserting it.
|
||
|
||
`SPECIFICATION.md`, `SECURITY-THREAT-MODEL.md`, `STORY-BRANCH-SEMANTICS.md` and
|
||
`V1-ACCEPTANCE-TESTS.md` are unchanged. M4 altered no product requirement, added
|
||
no outbound path, and implemented the checkpoint semantics
|
||
`STORY-BRANCH-SEMANTICS.md` §18-25 already specified rather than amending them.
|
||
|
||
**The browser smoke test remains unperformed, now for both M3 and M4.** No
|
||
session has had a usable browser. See `BUILD-MILESTONES.md` M3 and M4.
|
||
|
||
## v2.4 — Documentation Consolidation (2026-09-03)
|
||
|
||
No product requirement, architecture decision or milestone status changed in
|
||
this revision. It reorganises the documentation so that a new coding agent can
|
||
tell authoritative material from evidence at a glance.
|
||
|
||
- `planning/archive/` is created and is **non-authoritative by declaration**
|
||
(`archive/README.md`). It holds `phase0/` — the research that chose AI-DnD —
|
||
`milestone-reports/` — the completed M1 and M2 reports — and `decisions/`,
|
||
which now holds ADR **008**, the Phase-0-before-build process gate that Phase 0
|
||
satisfied. ADR numbering continues from 012; 008 is not reused.
|
||
- `planning/reports/` now holds **only the current milestone's report**,
|
||
`M3-IMPLEMENTATION-REPORT.md`, because M4 planning has to consult it. It moves
|
||
to the archive when M4's report replaces it.
|
||
- The Phase 0B execution prompts and handoff/status/summary documents
|
||
(`CODEX-HANDOFF-NOTE.md`, `PHASE-0B-CODEX-BRIEF.md`,
|
||
`PHASE-0B-CODEX-HANDOFF.md`, `PLANNING-UPDATE-SUMMARY.md`) and the Phase 0A
|
||
discovery and triage reports were **deleted**: intermediate working documents
|
||
whose conclusions all reached the two recommendation reports, and which remain
|
||
in Git history.
|
||
- Upstream AI-DnD's inherited `plan/` build log and `docs/` project site
|
||
(guides, generated HTML, screenshots) were **deleted**. They documented a
|
||
hosted, scripted, multi-user product with accounts — every screenshot showed a
|
||
Scripts tab and a Sign up button — which M2 removed. Both trees remain in Git
|
||
history and in upstream.
|
||
- `README.md`, `DEVELOPMENT.md` and `PROVENANCE.md` are corrected where they
|
||
pointed at the removed trees or described removed capability as present.
|
||
`DEVELOPMENT.md`'s "things M1 did not touch" section had gone stale at M2 and
|
||
now says what is actually still inherited.
|
||
- `planning/README.md` is rewritten as **the documentation index**: current
|
||
milestone, the three-way active/ADR/archive split, the authority order,
|
||
reading order, where reports live, and what comes next. The milestone
|
||
correction tables are preserved unchanged.
|
||
- New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
|
||
manifest of files to upload as ChatGPT Project Sources.
|
||
|
||
## v2.3 — Post-M3 Closeout (2026-09-03)
|
||
|
||
M3 replaced destructive Undo with a stored active head. Its review is
|
||
`archive/milestone-reports/M3-IMPLEMENTATION-REPORT.md`, which is also M3's primary evidence
|
||
record — no separate baseline report was produced — and whose §W records this
|
||
closeout.
|
||
|
||
In summary:
|
||
|
||
- the architecture is recorded as **ADR 012 — Active-Head Non-Destructive
|
||
History**: the head is stored rather than derived, every read of the story is
|
||
capped at it in one place, one mechanism moves it, state comes from the node
|
||
rather than from a replay, the first write below a moved-back head is the
|
||
divergence, and Redo is decided by the lineage rather than by a flag. ADR 005
|
||
is unchanged: it states the product requirement, and ADR 012 states the
|
||
architecture chosen to implement it,
|
||
- two history semantics are **ratified** in `STORY-BRANCH-SEMANTICS.md`: Undo
|
||
crosses fork points to the campaign opening (§5), and the system refuses to
|
||
switch which take is live while a later story is off screen (§10),
|
||
- a **new §14A** records the interim in-place-editing rule — refuse when story
|
||
descends from the turn and is not on screen — and states explicitly that
|
||
§14-15's full narrator-edit requirement stands and is completed in M5,
|
||
- `TECHNICAL-DESIGN.md` gains **§8.7** and **§9.1** recording the implemented
|
||
model and bundle behaviour as fact, and a constraint on §10.4: the snapshot
|
||
half of the hybrid state model is a requirement, because head movement must
|
||
not become proportional to campaign length,
|
||
- `DATA-MODEL.md` records the head as campaign-stored, the branch disposition as
|
||
implemented and deliberately advisory, and the export as carrying a chosen
|
||
position rather than a derived one,
|
||
- `BUILD-MILESTONES.md` marks **M3 complete**, states the one outstanding
|
||
condition (the browser smoke test), tells **M4** to reuse M3's head movement
|
||
rather than build a second restore path, and gives **M5** three constraints,
|
||
- `V1-ACCEPTANCE-TESTS.md` records D03's full pass, states **D10's milestone
|
||
ownership without weakening any pass condition**, adds I07's pre-M3 bundle
|
||
clause, and resolves the apparent L01/A05 conflict,
|
||
- `README.md` is corrected to describe the current local-only single-user
|
||
application rather than the upstream hosted one.
|
||
|
||
`SPECIFICATION.md` and `SECURITY-THREAT-MODEL.md` are unchanged: M3 altered no
|
||
product requirement and touched no path in the threat model.
|
||
|
||
## v2.2 — Post-M2 Closeout (2026-09-03)
|
||
|
||
M2 removed the hosted, cloud, account and scripting surface and added the
|
||
inference endpoint policy. Its review recommended six planning changes and
|
||
reported rather than applied them; all six are applied in this revision, listed
|
||
in `README.md` § *Post-M2 corrections applied*, with the evidence in
|
||
`archive/milestone-reports/M2-BASELINE-REPORT.md` and `archive/milestone-reports/M2-IMPLEMENTATION-REPORT.md`.
|
||
|
||
In summary:
|
||
|
||
- the **inference endpoint policy is recorded as implemented** — an address
|
||
allowlist of explicit local-network CIDRs, every resolved address checked,
|
||
enforced when settings are saved and again before every outbound request, with
|
||
TLS verification never traded against it (new **ADR 011**,
|
||
`SECURITY-THREAT-MODEL.md` §10A),
|
||
- its two **residual limits are stated rather than mitigated**: a hostile host
|
||
already on the trusted LAN, and the DNS-rebinding interval between the
|
||
policy's resolution and the client's connection,
|
||
- `TECHNICAL-DESIGN.md` §5.1 items 3 and 4 are **resolved**, and a new §5.2
|
||
records the M1/M2 production architecture as fact,
|
||
- a **wiring rule** is added (§18.1): removing a setting requires testing a real
|
||
consumer path, and adding one requires proving it reaches its component — M2
|
||
shipped two defects behind a 604-test green suite because the tests at that
|
||
boundary were mocks,
|
||
- `BUILD-MILESTONES.md` records **M2 complete**, warns M5 that eight rollback
|
||
tests use the world-state engine as instrumentation rather than as
|
||
architecture, and requires M6 to make background memory failure observable,
|
||
- the security acceptance contract is strengthened: **H10** now names the
|
||
wildcard-origin and `/api` 404 conditions, and new **H12** covers inference
|
||
endpoint enforcement including the database-edited-behind-the-API case.
|
||
|
||
`SPECIFICATION.md` is unchanged: M2 altered no product requirement. Nothing in
|
||
the architecture selected in v2 was reversed.
|
||
|
||
## v2.1 — Post-M1 Corrections (2026-09-02)
|
||
|
||
M1 implementation evidence contradicted or under-specified parts of v2. The
|
||
corrections are recorded in the documents themselves and listed in
|
||
`README.md` § *Post-M1 corrections applied*; the evidence behind them is in
|
||
`archive/milestone-reports/M1-BASELINE-REPORT.md` and `archive/milestone-reports/M1-IMPLEMENTATION-REPORT.md`.
|
||
|
||
In summary:
|
||
|
||
- a trusted-LAN Ollama may be **HTTPS with a privately issued certificate**;
|
||
clients verify against the operating system's CA store, with full certificate
|
||
and hostname verification and no bypass option (ADR 002,
|
||
`TECHNICAL-DESIGN.md` §5, A06),
|
||
- offline claims require a **fresh cache and no route out** to be evidence at
|
||
all, and vendored runtime artifacts should be integrity-verifiable
|
||
(ADR 004),
|
||
- A05's invariant is about **accepted** history; the user's submitted text is
|
||
deliberately retained on a failed turn,
|
||
- A06 requires a **real second machine and an HTTPS endpoint**; a plain-HTTP
|
||
LAN test is no longer sufficient evidence,
|
||
- the standard test environment records **CPU/GPU/RAM**, because cold model
|
||
load on a CPU-only host exceeded the inherited 120 s client timeout,
|
||
- `BUILD-MILESTONES.md` records M1 as complete and reframes M2's endpoint work
|
||
as **narrowing** an existing capability rather than inventing it,
|
||
- `SECURITY-THREAT-MODEL.md` §53 distinguishes **inbound** TLS (still deferred)
|
||
from **outbound** certificate verification (required, done in M1).
|
||
|
||
Nothing in the architecture selected in v2 was reversed.
|
||
|
||
## v2 — Post Phase 0B Revision (2026-09-01)
|
||
|
||
This v2 package supersedes the earlier planning package produced before the final Phase 0B review and the trusted-LAN Ollama deployment clarification.
|
||
|
||
Key v2 changes include:
|
||
|
||
- AI-DnD selected as the production base at the pinned Phase 0B commit.
|
||
- Non-destructive head-cursor Undo/Redo design selected.
|
||
- Explicit typed/absolute narrative-state events selected for production state handling.
|
||
- The authoritative narrative-state document — its shape, its authority/provenance fields, and the
|
||
event/document/snapshot split — recorded in ADR 013 after M5 implemented it.
|
||
- Imported knowledge separated from AI-DnD Story Cards.
|
||
- Trusted-LAN Ollama inference supported in v1 while the storyteller UI/API remains loopback-bound by default.
|
||
- Offline first-use dependencies and runtime remote assets identified as M1 hardening work.
|
||
- Production implementation divided into milestones M1-M11.
|
||
|
||
Historical Phase 0 prompts/reports are retained as evidence and should not be treated as current implementation instructions unless a current milestone prompt explicitly refers to them.
|